download dots

OpenAI · tested in Taskade Genesis

How GPT builds in Taskade Genesis

Use GPT in Taskade Genesis for detailed requests and fast delivery.

Last test

TSK-1 scorecard for GPT

TSK Score

Task completion across real apps built in Taskade Genesis

TSK Score 88 out of 100
Interfaceready to share
Strong
Taskcompletes your task
Leading
Memorykeeps your data
Strong
Adapthandles follow-up edits
Leading

GPT is strongest at following detailed requests and finishing quickly.

More

GPT-6.1 Sol, the newest, shared the top result of our latest phone-form test. Luna preserves exact wording and builds complete apps efficiently. Terra is the fastest in nearly every test it enters. GPT-5.6 Sol builds wide apps but can add sign-in screens nobody asked for.

Where GPT ranks

TSK Score out of 100.

  1. 1GPT88 out of 100
  2. 2Claude75 out of 100
  3. 2DeepSeek75 out of 100
  4. 2Kimi75 out of 100
  5. 2Grok75 out of 100
  6. 6GLM63 out of 100
  7. 7Qwen50 out of 100
  8. 8Gemini44 out of 100
GPT versions over timeDays each version was tested
GPT versions over time: days tested so far, Jul 30 to Oct 8
DateGPT-6.1 SolGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 Sol
0001, tested this day
0002, tested this day
01, tested this day02
02, tested this day02
03, tested this day1, tested this day2
04, tested this day2, tested this day2
05, tested this day22
06, tested this day3, tested this day3, tested this day
07, tested this day4, tested this day3
08, tested this day43
09, tested this day43
010, tested this day5, tested this day4, tested this day
011, tested this day6, tested this day5, tested this day
1, tested this day12, tested this day65
113, tested this day7, tested this day6, tested this day

Versions tested

Open a version to read its findings.

GPT-6.1 SolTested Sep 20264 findings

The newest GPT we test. All three of its family intake forms saved every answer and both photos, parents could use them without an account as the request asked, and every follow-up change landed.

Sep 30, 2026

Built the family intake form three times. All three saved every answer and both photos, and parents could fill them in without an account, as the request asked.

Sep 30, 2026

In two of the three builds the teacher page listed every family with their photos. In the third it read the saved answers wrong and showed "No photos saved" for a complete entry.

Sep 30, 2026

Added the new teacher question in all three apps, saved it with each submission, and put the change live.

Sep 30, 2026

Every closing summary said what it had not tested and pointed out that the teacher page was open to anyone with the link.

GPT-5.6 LunaTested Oct 202620 findings

The most faithful model we have tested. It reproduced all 32 questions of a real customer's client sign-up form word for word again, built a scoring automation that ran, and passed every September follow-up-edit test.

Aug 3, 2026

Reproduced all 32 questions of a real customer's client sign-up form word for word. The only model to get all four of its builds word for word across more than one test.

Aug 7, 2026

The only model to build a working scoring grid: 33 rows by 5 ratings, with every score saved straight back into the workspace.

Aug 4, 2026

The best client sign-up form Luna has built: nothing stalled, and a human reviewer called it beautiful.

Aug 8, 2026

Test winner on value. The only model to explain its own design choices, and both builds carried all 32 questions word for word.

Aug 19, 2026

The baseline for automatic model picking: an app built in about 11 minutes and edited in about 6.

Aug 25, 2026

Its match tracker crashed on the first real entry and its sign-up form rejected the first submission; later tests finished cleanly.

Aug 27, 2026

The value pick on every app in a five-model test of identical prompts.

Aug 28, 2026

Unchanged result after a large platform update, with fewer failed build actions.

Aug 29, 2026

Kept a build checklist inside both apps, so the finished work could be checked against the request.

Sep 19, 2026

Built the client sign-up form with four pages and a scoring automation that ran on its own.

Sep 19, 2026

Built the fleet-inspection app with both weekly automations live in 7 minutes 18 seconds.

Sep 19, 2026

Built the AI governance office in 6 minutes 38 seconds with all 24 named items in place.

Sep 19, 2026

Built a complete match tracker with a live dashboard and a weekly recap automation.

Sep 19, 2026

Passed all five follow-up-edit tests on its habit tracker.

Sep 19, 2026

Its recipe box rendered as a warm printed cookbook in light and dark, with every page loading cleanly.

Sep 22, 2026

Built a two-page match tracker in 8 minutes 19 seconds, the fastest and most economical complete build of its test, and a match saved from the app came back with all eight details in place.

Sep 30, 2026

Built the family intake form so it saved every answer and both photos, and added the new teacher question cleanly: the most economical complete build of its test.

Sep 30, 2026

Its public class assistant told an anonymous visitor a child's name from a saved form.

Oct 8, 2026

Built the family intake form twice. One build saved every answer and both photos in under 5 minutes. In the other, every parent submission failed on the photo permission question, although its closing summary said the form had worked.

Oct 8, 2026

Built the Dota 2 match tracker twice in about 5 and a half minutes each. Both saved a logged match with every detail and added a ranked heroes page when asked.

GPT-5.6 TerraTested Oct 202613 findings

The fastest model in nearly every test it enters, and the first to pass every check in a single run: it built the app, took a form submission, put every answer in the right place, ran the automation, and answered questions about the data. It also turned its weakest result around, from three tests where none of the brief came through word for word to all 32 questions word for word.

Aug 7, 2026

The first model to pass all five checks in one run: built the app, took a submission, saved every answer in the right field, ran the automation, and answered questions about the data.

Aug 7, 2026

Fastest of the test: 8 minutes 26 seconds.

Aug 8, 2026

Both builds carried all 32 questions word for word, after three tests where none of the four came through. A change of setting turned it around.

Aug 8, 2026

Terra's premium setting was the worst value we have measured on the client sign-up form: far more expensive than Luna, and it delivered less, a single page rewritten 19 times.

Aug 25, 2026

Best all-round match tracker and best client sign-up build of the test: form to saved score worked on the first try.

Aug 27, 2026

The quality pick in a five-model test of identical prompts.

Sep 19, 2026

Fastest complete client sign-up build of the test at 5 minutes 6 seconds; it left the scoring automation out.

Sep 19, 2026

Finished a two-page sales pipeline with a daily call brief in 10 minutes 47 seconds.

Sep 19, 2026

Finished the governance office in under 6 minutes and the fleet inspection in under 5.

Sep 19, 2026

Passed all five follow-up-edit tests; its recipe box was a single page.

Sep 22, 2026

Built a clean match tracker in about 11 minutes, but its match form offered only four heroes from its own sample list, so a player could not log the hero they played.

Oct 8, 2026

Built the family intake form twice, about 4 minutes each, and both saved every answer and both photos. One opened in a dark look where the request asked for cream paper. The other asked first who could open the teacher page, then put it behind a teacher sign-in.

Oct 8, 2026

Built the Dota 2 match tracker twice in under 4 and a half minutes, and both added a ranked heroes page when asked. One build labeled its starter matches, and the other ran wider than a phone screen on every page.

GPT-5.6 SolTested Oct 202611 findings

The login-wall lesson, and a turnaround. Through September Sol built wide apps with complete data but kept adding sign-in screens nobody asked for. In October its family intake form needed no sign-in, both of its apps labeled their sample data, and both follow-up changes landed.

Jul 30, 2026

Added a sign-in screen the brief never asked for, and did not say so.

Aug 1, 2026

Added the same unasked-for sign-in screen again. Repeatable enough that it shaped how we score a build that answers a different brief than the one given.

Aug 25, 2026

Added a sign-in gate to the match tracker again, the third time in three tests.

Sep 19, 2026

Put an unasked sign-in screen in front of two of its three finished client sign-up builds, over complete data and three scoring runs.

Sep 19, 2026

Built a 34-field fleet-inspection form and left the Thursday reminder switched off.

Sep 19, 2026

Built the widest governance office of the test, with seven pages and every named item, behind a sign-in screen.

Sep 19, 2026

Built a two-page match tracker with no sign-in screen.

Sep 19, 2026

Passed all five follow-up-edit tests, but the habit tracker sat behind a sign-in screen it added itself.

Sep 22, 2026

Put a sign-in screen nobody asked for in front of a personal match tracker, and its summary did not mention that an account had to be created first.

Oct 8, 2026

Built the family intake form so it saved every answer and both photos, labeled its sample family, and said on the teacher page that it needs no sign-in. The new teacher question landed.

Oct 8, 2026

Built the Dota 2 match tracker with its sample matches labeled as a demo, saved a logged match with every detail, and put the new heroes page in both the desktop and the phone menu.

Same-request tests

The family intake form, more models

All models →
The request

Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.

  • GPT-5.6 SolWorks

    Saved every answer and both photos, labeled its sample family, and kept the teacher page open with no sign-in. The new teacher question landed.

  • GPT-5.6 Luna×2 buildsWorks, with gaps

    One build saved every answer and both photos in under 5 minutes. In the other, every parent submission failed on the photo permission question.

  • GPT-5.6 Terra×3 buildsWorks, with gaps

    All three builds saved every answer and both photos in about 4 minutes or less. One opened in a dark look where the request asked for cream paper.

A Dota 2 match tracker with a heroes page

All models →
The request

One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.

  • GPT-5.6 SolWorks

    Labeled its sample matches as a demo, saved a logged match with every detail, and put the new heroes page in both menus.

  • GPT-5.6 Luna×2 buildsWorks

    Both builds took about 5 and a half minutes, saved a logged match with every detail, and added a ranked heroes page when asked. Sample matches showed as the player's own history.

  • GPT-5.6 Terra×2 buildsWorks

    Both builds took under 4 and a half minutes and added a ranked heroes page. One labeled its starter matches, and the other ran wider than a phone screen.

A family intake form on a phone

All models →
The request

Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.

  • GPT-6.1 Sol×3 buildsWorks

    All three saved every answer and both photos, and every follow-up change landed. One teacher page misread its saved answers.

  • GPT-5.6 LunaWorks, with gaps

    The most economical complete build: every answer and both photos saved. Its public assistant told a visitor a child's name.

A match tracker

All models →
The request

One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.

  • GPT-5.6 LunaWorks

    The fastest and most economical complete build, in 8 minutes 19 seconds. A match saved from the app came back with all eight details.

  • GPT-5.6 TerraWorks, with gaps

    A clean tracker, but its match form offered only four heroes.

  • GPT-5.6 SolWorks, with gaps

    Put a sign-in screen nobody asked for in front of a personal tracker.

One build per model unless marked ×N. A direction, not a final rank.

App kits you can open today

Live apps from the official Taskade account, not builds by GPT.

Browse all App Kits →

Compare GPT

See all 9 comparisons

Other models in TSK-1

See the full TSK-1 ranking →

FAQ

What is GPT's TSK Score?

GPT scores 88 out of 100 on TSK-1: Interface Strong, Task Leading, Memory Strong, Adapt Leading. Last test Oct 8, 2026.

Which GPT variant builds the best apps?

GPT-6.1 Sol, the newest, saved every answer and photo in all three of its phone-form builds and landed every follow-up change. GPT-5.6 Luna sticks to the brief best and builds complete apps efficiently. GPT-5.6 Terra is the fastest and was the first to pass every check in one run. GPT-5.6 Sol builds wider apps, but it can add sign-in screens nobody asked for.

Is a faster AI model worse for app building?

Not necessarily. GPT-5.6 Terra is the fastest model we test and the first to pass every check in one run. But speed only counts if the app matches the brief. For three tests Terra reproduced none of the customer's questions word for word; a change of setting brought it back to all 32.

Can I use GPT models in Taskade?

Yes. GPT-6.1 Sol, GPT-6 Luna, and GPT-5.6 Luna are in the Taskade model picker on every paid plan. On Max and Enterprise, you can also pick GPT-5.6 Sol and GPT-5.6 Terra directly.

Where can I compare GPT with other models, or try it myself?

Side-by-side model pages live at /compare. To try the same kind of app request yourself, start from /create. This page stays on what GPT produced in TSK-1 tests.