
OpenAI · tested in Taskade Genesis
How GPT builds in Taskade Genesis
Use GPT in Taskade Genesis for detailed requests and fast delivery.
Last test
TSK-1 scorecard for GPT
TSK Score
Task completion across real apps built in Taskade Genesis
- Interfaceready to share
- Strong
- Taskcompletes your task
- Leading
- Memorykeeps your data
- Strong
- Adapthandles follow-up edits
- Leading
GPT is strongest at following detailed requests and finishing quickly.
MoreLess
GPT-6.1 Sol, the newest, shared the top result of our latest phone-form test. Luna preserves exact wording and builds complete apps efficiently. Terra is the fastest in nearly every test it enters. GPT-5.6 Sol builds wide apps but can add sign-in screens nobody asked for.
Where GPT ranks
TSK Score out of 100.
| Date | GPT-6.1 Sol | GPT-5.6 Luna | GPT-5.6 Terra | GPT-5.6 Sol |
|---|---|---|---|---|
| 0 | 0 | 0 | 1, tested this day | |
| 0 | 0 | 0 | 2, tested this day | |
| 0 | 1, tested this day | 0 | 2 | |
| 0 | 2, tested this day | 0 | 2 | |
| 0 | 3, tested this day | 1, tested this day | 2 | |
| 0 | 4, tested this day | 2, tested this day | 2 | |
| 0 | 5, tested this day | 2 | 2 | |
| 0 | 6, tested this day | 3, tested this day | 3, tested this day | |
| 0 | 7, tested this day | 4, tested this day | 3 | |
| 0 | 8, tested this day | 4 | 3 | |
| 0 | 9, tested this day | 4 | 3 | |
| 0 | 10, tested this day | 5, tested this day | 4, tested this day | |
| 0 | 11, tested this day | 6, tested this day | 5, tested this day | |
| 1, tested this day | 12, tested this day | 6 | 5 | |
| 1 | 13, tested this day | 7, tested this day | 6, tested this day |
Versions tested
Open a version to read its findings.
GPT-6.1 SolTested Sep 20264 findings
The newest GPT we test. All three of its family intake forms saved every answer and both photos, parents could use them without an account as the request asked, and every follow-up change landed.
Built the family intake form three times. All three saved every answer and both photos, and parents could fill them in without an account, as the request asked.
In two of the three builds the teacher page listed every family with their photos. In the third it read the saved answers wrong and showed "No photos saved" for a complete entry.
Added the new teacher question in all three apps, saved it with each submission, and put the change live.
Every closing summary said what it had not tested and pointed out that the teacher page was open to anyone with the link.
GPT-5.6 LunaTested Oct 202620 findings
The most faithful model we have tested. It reproduced all 32 questions of a real customer's client sign-up form word for word again, built a scoring automation that ran, and passed every September follow-up-edit test.
Reproduced all 32 questions of a real customer's client sign-up form word for word. The only model to get all four of its builds word for word across more than one test.
The only model to build a working scoring grid: 33 rows by 5 ratings, with every score saved straight back into the workspace.
The best client sign-up form Luna has built: nothing stalled, and a human reviewer called it beautiful.
Test winner on value. The only model to explain its own design choices, and both builds carried all 32 questions word for word.
The baseline for automatic model picking: an app built in about 11 minutes and edited in about 6.
Its match tracker crashed on the first real entry and its sign-up form rejected the first submission; later tests finished cleanly.
The value pick on every app in a five-model test of identical prompts.
Unchanged result after a large platform update, with fewer failed build actions.
Kept a build checklist inside both apps, so the finished work could be checked against the request.
Built the client sign-up form with four pages and a scoring automation that ran on its own.
Built the fleet-inspection app with both weekly automations live in 7 minutes 18 seconds.
Built the AI governance office in 6 minutes 38 seconds with all 24 named items in place.
Built a complete match tracker with a live dashboard and a weekly recap automation.
Passed all five follow-up-edit tests on its habit tracker.
Its recipe box rendered as a warm printed cookbook in light and dark, with every page loading cleanly.
Built a two-page match tracker in 8 minutes 19 seconds, the fastest and most economical complete build of its test, and a match saved from the app came back with all eight details in place.
Built the family intake form so it saved every answer and both photos, and added the new teacher question cleanly: the most economical complete build of its test.
Its public class assistant told an anonymous visitor a child's name from a saved form.
Built the family intake form twice. One build saved every answer and both photos in under 5 minutes. In the other, every parent submission failed on the photo permission question, although its closing summary said the form had worked.
Built the Dota 2 match tracker twice in about 5 and a half minutes each. Both saved a logged match with every detail and added a ranked heroes page when asked.
GPT-5.6 TerraTested Oct 202613 findings
The fastest model in nearly every test it enters, and the first to pass every check in a single run: it built the app, took a form submission, put every answer in the right place, ran the automation, and answered questions about the data. It also turned its weakest result around, from three tests where none of the brief came through word for word to all 32 questions word for word.
The first model to pass all five checks in one run: built the app, took a submission, saved every answer in the right field, ran the automation, and answered questions about the data.
Fastest of the test: 8 minutes 26 seconds.
Both builds carried all 32 questions word for word, after three tests where none of the four came through. A change of setting turned it around.
Terra's premium setting was the worst value we have measured on the client sign-up form: far more expensive than Luna, and it delivered less, a single page rewritten 19 times.
Best all-round match tracker and best client sign-up build of the test: form to saved score worked on the first try.
The quality pick in a five-model test of identical prompts.
Fastest complete client sign-up build of the test at 5 minutes 6 seconds; it left the scoring automation out.
Finished a two-page sales pipeline with a daily call brief in 10 minutes 47 seconds.
Finished the governance office in under 6 minutes and the fleet inspection in under 5.
Passed all five follow-up-edit tests; its recipe box was a single page.
Built a clean match tracker in about 11 minutes, but its match form offered only four heroes from its own sample list, so a player could not log the hero they played.
Built the family intake form twice, about 4 minutes each, and both saved every answer and both photos. One opened in a dark look where the request asked for cream paper. The other asked first who could open the teacher page, then put it behind a teacher sign-in.
Built the Dota 2 match tracker twice in under 4 and a half minutes, and both added a ranked heroes page when asked. One build labeled its starter matches, and the other ran wider than a phone screen on every page.
GPT-5.6 SolTested Oct 202611 findings
The login-wall lesson, and a turnaround. Through September Sol built wide apps with complete data but kept adding sign-in screens nobody asked for. In October its family intake form needed no sign-in, both of its apps labeled their sample data, and both follow-up changes landed.
Added a sign-in screen the brief never asked for, and did not say so.
Added the same unasked-for sign-in screen again. Repeatable enough that it shaped how we score a build that answers a different brief than the one given.
Added a sign-in gate to the match tracker again, the third time in three tests.
Put an unasked sign-in screen in front of two of its three finished client sign-up builds, over complete data and three scoring runs.
Built a 34-field fleet-inspection form and left the Thursday reminder switched off.
Built the widest governance office of the test, with seven pages and every named item, behind a sign-in screen.
Built a two-page match tracker with no sign-in screen.
Passed all five follow-up-edit tests, but the habit tracker sat behind a sign-in screen it added itself.
Put a sign-in screen nobody asked for in front of a personal match tracker, and its summary did not mention that an account had to be created first.
Built the family intake form so it saved every answer and both photos, labeled its sample family, and said on the teacher page that it needs no sign-in. The new teacher question landed.
Built the Dota 2 match tracker with its sample matches labeled as a demo, saved a logged match with every detail, and put the new heroes page in both the desktop and the phone menu.
Same-request tests
The family intake form, more models
All models →The request
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
GPT-5.6 SolWorks
Saved every answer and both photos, labeled its sample family, and kept the teacher page open with no sign-in. The new teacher question landed.
GPT-5.6 Luna×2 buildsWorks, with gaps
One build saved every answer and both photos in under 5 minutes. In the other, every parent submission failed on the photo permission question.
GPT-5.6 Terra×3 buildsWorks, with gaps
All three builds saved every answer and both photos in about 4 minutes or less. One opened in a dark look where the request asked for cream paper.
A Dota 2 match tracker with a heroes page
All models →The request
One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.
GPT-5.6 SolWorks
Labeled its sample matches as a demo, saved a logged match with every detail, and put the new heroes page in both menus.
GPT-5.6 Luna×2 buildsWorks
Both builds took about 5 and a half minutes, saved a logged match with every detail, and added a ranked heroes page when asked. Sample matches showed as the player's own history.
GPT-5.6 Terra×2 buildsWorks
Both builds took under 4 and a half minutes and added a ranked heroes page. One labeled its starter matches, and the other ran wider than a phone screen.
A family intake form on a phone
All models →The request
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
GPT-6.1 Sol×3 buildsWorks
All three saved every answer and both photos, and every follow-up change landed. One teacher page misread its saved answers.
GPT-5.6 LunaWorks, with gaps
The most economical complete build: every answer and both photos saved. Its public assistant told a visitor a child's name.
A match tracker
All models →The request
One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.
GPT-5.6 LunaWorks
The fastest and most economical complete build, in 8 minutes 19 seconds. A match saved from the app came back with all eight details.
GPT-5.6 TerraWorks, with gaps
A clean tracker, but its match form offered only four heroes.
GPT-5.6 SolWorks, with gaps
Put a sign-in screen nobody asked for in front of a personal tracker.
One build per model unless marked ×N. A direction, not a final rank.
App kits you can open today
Live apps from the official Taskade account, not builds by GPT.
Content studio assistant TOOLContent studio assistant4 projects · 1 agent · 3 flowsBusiness metrics console DASHBOARDBusiness metrics console5 projects · 1 agent · 2 flowsEvent seating planner TOOLEvent seating planner3 projects · 1 agent · 2 flowsFleet tracking dashboard DASHBOARDFleet tracking dashboard3 projects · 1 agent · 2 flowsProduct storefront PORTALProduct storefront3 projects · 1 agent · 4 flowsSales CRM CRMSales CRM3 projects · 1 agent · 2 flows
Compare GPT
See all 9 comparisonsShow fewer
Other models in TSK-1
FAQ
What is GPT's TSK Score?
GPT scores 88 out of 100 on TSK-1: Interface Strong, Task Leading, Memory Strong, Adapt Leading. Last test Oct 8, 2026.
Which GPT variant builds the best apps?
GPT-6.1 Sol, the newest, saved every answer and photo in all three of its phone-form builds and landed every follow-up change. GPT-5.6 Luna sticks to the brief best and builds complete apps efficiently. GPT-5.6 Terra is the fastest and was the first to pass every check in one run. GPT-5.6 Sol builds wider apps, but it can add sign-in screens nobody asked for.
Is a faster AI model worse for app building?
Not necessarily. GPT-5.6 Terra is the fastest model we test and the first to pass every check in one run. But speed only counts if the app matches the brief. For three tests Terra reproduced none of the customer's questions word for word; a change of setting brought it back to all 32.
Can I use GPT models in Taskade?
Yes. GPT-6.1 Sol, GPT-6 Luna, and GPT-5.6 Luna are in the Taskade model picker on every paid plan. On Max and Enterprise, you can also pick GPT-5.6 Sol and GPT-5.6 Terra directly.
Where can I compare GPT with other models, or try it myself?
Side-by-side model pages live at /compare. To try the same kind of app request yourself, start from /create. This page stays on what GPT produced in TSK-1 tests.





