download dots

TSK-Bench · Task completion benchmark for AI models

TSK-1 Benchmark

Same request. Every AI. Scored on what it actually finished.

GPT follows long briefs word for word. Claude finishes apps best, on your own key in automations and Enterprise agents. DeepSeek leads on design. Kimi builds rich apps, but slowly. Not sure? TSK-1 Auto picks for you.

TSK-1 model benchmark

ModelInterfaceHow complete and polished the finished app feels.TaskTask completion: how fully the app does what you asked, word for word.MemoryHow reliably the app keeps what people add to your workspace.AdaptHow cleanly the app changes when you ask.TSK ScoreTask completion out of 100: the four columns in one number. Leading counts most, Limited least.
Show all 11 models
TSK-1 leaderboard: rank, four graded axes, and TSK Score
RankModelInterfaceTaskMemoryAdaptTSK ScoreNewest version tested
1GPTStrongLeadingStrongLeading88 out of 100GPT-5.6 Luna
2ClaudeStrongStrongEmergingLeading75 out of 100Claude Sonnet 5
2DeepSeekLeadingStrongStrongEmerging75 out of 100DeepSeek V4.1 Flash
2KimiStrongStrongStrongStrong75 out of 100Kimi K3
2GrokStrongStrongStrongStrong75 out of 100Grok 4.7
6GLMEmergingStrongEmergingStrong63 out of 100GLM-5
7QwenEmergingStrongEmergingLimited50 out of 100Qwen 3.7 Plus
8GeminiLimitedEmergingEmergingEmerging44 out of 100Gemini 3.6 Flash
Not rankedMiniMaxNot scoredNot scoredNot scoredNot scoredNot scoredMiniMax M3
Not rankedMistralNot scoredNot scoredNot scoredNot scoredNot scoredMistral Large 4
  • Leading
  • Strong
  • Emerging
  • Limited
Updated
Interfaceready to shareTaskcompletes your taskMemorykeeps your dataAdapthandles follow-up edits
TSK-1 Benchmark over timeDays each model was tested, Jul 30 to Oct 8, 2026
TSK-1 Benchmark over time: days tested so far, Jul 30 to Oct 8
DateGPTClaudeDeepSeekKimiGrokGLMQwenGeminiMiniMaxMistral
1, tested this day1, tested this day0001, tested this day0000
12, tested this day01, tested this day01001, tested this day0
2, tested this day3, tested this day1, tested this day102, tested this day01, tested this day10
232, tested this day1020110
3, tested this day4, tested this day3, tested this day10202, tested this day10
4, tested this day431020210
444, tested this day1020210
45, tested this day5, tested this day1020210
5, tested this day551020210
6, tested this day551020210
7, tested this day551020210
7551020210
8, tested this day6, tested this day6, tested this day2, tested this day1, tested this day3, tested this day1, tested this day22, tested this day0
8662131220
9, tested this day662131220
10, tested this day662131220
11, tested this day66214, tested this day1220
11662141220
12, tested this day7, tested this day7, tested this day3, tested this day2, tested this day5, tested this day1220
13, tested this day8, tested this day8, tested this day3251220
14, tested this day9, tested this day9, tested this day4, tested this day251220
15, tested this day10, tested this day10, tested this day43, tested this day6, tested this day1221, tested this day

Start building

Describe the app you need. TSK-1 coordinates the rest.

Benchmark updates

What each test found.

More models on two real apps

More than 20 models and settings rebuilt the family intake form and a Dota 2 match tracker, and then took one follow-up change each. Grok 4.7 and GPT-5.6 Sol built both apps well, several builds showed sample families or matches as if they were real, and a few forms could not save at all.

A phone form with photos

A real customer's family intake form joined the test set: parents fill it in on a phone with two to four photos, and the teacher sees every family on one page. GPT-6.1 Sol and Claude Opus 5.5 shared the top result, and every model that finished also added a follow-up question cleanly.

One match tracker, model by model

Several models built the same match tracker. GPT-5.6 Luna was the fastest and most economical complete build, DeepSeek V4.1 Flash ran a coaching automation on the first match logged, and two models limited their match forms to a short list of heroes.

Real-customer shapes and follow-up edits

Three real-customer app shapes joined the set: a sales CRM, a fleet inspection, and an AI governance office. The first full follow-up-edit ladder passed on every model that ran it.

A broad model sweep and parallel helpers

Seventeen models built the same client sign-up form. A separate test of parallel helpers showed that one file per helper produced the cleanest result.

Show all 22 updates
Same request, four setups

A control run told us how big a difference has to be before it counts. Two identical setups produced noticeably different costs for the same app, so we now only claim a difference when it is larger than that gap, and we grade on what the app does.

Same result after a platform update

We re-ran the automatic model after a large platform update. The apps came out the same, with fewer failed build actions along the way.

Sixteen builds, five models, identical prompts

Five models built and edited the same apps from identical prompts. GPT-5.6 Luna was the value pick on every app and GPT-5.6 Terra the quality pick. Most builds now keep a checklist of the request inside the app.

Cleaner follow-up edits

Follow-up edits now start from a map of the app, and they landed cleaner. GPT-5.6 Luna edited apps other models had built without breaking them, and a missing page now explains itself.

Two new model families, three new app shapes

Qwen and Grok entered the benchmark, and three real-customer app shapes joined the test set: a sales CRM, a field-inspection audit, and an AI governance office. Every model still gets the same words.

A dedicated follow-up-edit test

We added a dedicated follow-up-edit pass and stricter instruction checks. On its first run, none of the 31 edit attempts finished - a strong first build is not enough on its own; an app must also change cleanly and stay within the brief.

Faster builds, verified results

A complete app finished in under five minutes without follow-up. We also rechecked every published result and established a baseline for automatic model selection.

Accuracy and value moved together

The strongest results reproduced a long brief exactly, balanced build quality with efficiency, and explained their design choices clearly.

Lower cost did not mean lower quality

An efficient build completed a long form at the lowest measured cost, while the leading result also saved every submitted field correctly.

Long briefs became a harder test

Several builds followed the brief word for word, and one handled a 60-field form without cutting requirements.

Working software became the baseline

We made a running app the entry requirement. Attractive results that failed to open no longer received a score.

Efficiency without tradeoffs

The cleanest test also delivered the lowest measured cost, showing that careful model selection can improve both quality and value.

Brief fidelity and saved data

The strongest builds matched every requested field, stored submitted data correctly, and remained useful from form to follow-up.

Prompt length changed the outcome

Shorter instructions improved brief matching and edits, but sometimes reduced visual quality. The benchmark now balances all four dimensions.

Nine models, one consistent standard

Nine models built the same app. Several produced polished designs, but only working results qualified for comparison.

Self-checking improved reliability

The most dependable results detected and repaired problems before finishing, while weaker builds reported success too early.

Restraint mattered

The strongest designs followed the brief without adding unwanted gates or steps. Doing only what was asked became part of quality.

Same request, every model

A Dota 2 match tracker with a heroes page

The request

One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.

A family intake form on a phone

The request

Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.

A match tracker

The request

One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.

One build per model unless marked ×N. A direction, not a final rank.

Up next: GPT-6 Luna and Nemotron 3 Super

Every version we tested

Open a version to see what it built.

Compare two models

See all 26 comparisons

FAQ

Which AI model is best for building apps?

GPT has the highest TSK Score, 88 out of 100, but the best pick depends on the app. Some models hold closer to a brief, some finish faster, some look more polished. In Taskade you do not have to choose: TSK-1 Auto handles the default, and you can pick a model yourself when you want more control.

What can Taskade Genesis build with TSK-1?

Taskade Genesis can turn one request into CRMs, dashboards, portals, forms, trackers, and internal tools. TSK-1 coordinates the AI model, workspace memory, agents, and automations so the result can store information, answer questions, run workflows, and keep improving.

What does the TSK-1 Benchmark measure?

Whether a model can turn one prompt into a complete, working app. We grade four qualities of a living system: Interface, Task, Memory, and Adaptation: how finished it feels, how closely it follows your request, whether it keeps your data, and how cleanly it handles follow-up changes.

How is TSK-1 different from Vibe Code Bench, WebDev Arena, and Design Arena?

Those benchmarks grade code, a web page, or crowd votes on a screenshot. TSK-1 grades a running app: we open it, enter data the way a customer would, check that the data was kept, run its automation, and ask for a follow-up change. A build that looks right but loses your data does not score.

Does TSK-1 benchmark app builders or AI models?

AI models. Every model builds inside the same app builder, Taskade Genesis, from the same request, word for word, so the builder is held constant. A control run in August showed that two identical setups can still differ, so we only call out a difference between models when it is larger than that gap.

Do you test the apps or just the code?

We test the apps. Every build gets opened and used. We fill in its form the way a customer would, submit it, then check that the answers landed in the right place. An app that looks beautiful and loses your data fails.

Which models can I use in Taskade?

Taskade runs frontier models from top AI labs, and TSK-1 Auto handles the default. Built in, with no key, every plan can pick DeepSeek V4.1 Flash, and paid plans add OpenAI GPT models, xAI Grok and more open-weight models such as Qwen, Kimi and GLM. Claude runs on your own Anthropic key, in automations and in custom agents on Enterprise. Gemini runs only in automations, on your own Gemini API key. Model lineup as of October 2026.

How often is this updated?

We run a new test whenever a notable model ships. Every result cleared for public comparison stays on this page, including the ones that did not go well.

What is TSK-1?

TSK-1 is the intelligence behind Taskade Genesis. It brings AI models, workspace memory, agents, and automations together so one request can become a working app. The TSK-1 Benchmark shows how that process performs on real app builds.

Can I try the same app request?

Yes. Every model in a comparison receives the same request, word for word. Describe the same idea in Taskade Genesis, choose a model or let TSK-1 Auto decide, and see what it builds.

Are the benchmarked apps publicly available?

Not yet one by one. We are listing every finished benchmark build under a single official creator so you can open and clone it. Until then, each model page explains what that model produced, and the App Kits on this page are live systems you can open and clone today.

How many builds are behind each result?

Each result comes from a small set of hands-on app builds. The TSK Score rescales four evidence tiers to 0-100 for easier comparison; it is not 100 separate checks. Models with the same result share a rank.