The TSK-1 benchmark is the public evidence layer for TSK-1, the Taskade System Kernel. AI models build the same two real apps inside Taskade Genesis, a tracker and a real customer's 32-question client sign-up form with instructions not to shorten it, and every build is graded on whether it opens, runs, and saves submitted answers into the right fields. Results are published as dated summaries on the TSK-1 benchmark hub, with one evidence page per model family. Updated August 2026.
TL;DR: The TSK-1 benchmark grades finished apps, not answers. Every build test runs the same two fixed prompts; on Aug 1, 2026 nine models built the same app and 3 of the 9 builds never opened. Each graded build is written up per model, with the day it was measured. See every model's evidence
What Is the TSK-1 Benchmark?
The TSK-1 benchmark is the evidence layer that makes TSK-1's routing claims checkable. Where traditional benchmarks publish only a chart of scores, the TSK-1 benchmark publishes what the finished apps did: each model's page carries its graded evidence, the hub logs a summary for every result day, and the matrix distills the tiers into a 0-100 Intelligence Index. The hub also links live app kits built in Taskade Genesis you can open, click through, and clone into your own workspace.
TSK-1 is the kernel that coordinates models, memory, agents, and workflows. The benchmark answers the question the kernel raises every day: which model should this task actually run on? It answers by running model after model through the same two prompts, in the same setup, and grading the apps they hand back. Seven model families have been benchmarked so far, and on one day in August nine models built the same app from the same prompt.
The two prompts never change because a benchmark only measures what is held constant. The tracker exercises real app structure: entities, relationships, filters, and views. The 32-question client sign-up form exercises how closely a model sticks to what it was given: instructions not to shorten that the models are not allowed to dodge.
The sign-up prompt text itself is not published. It comes from a real customer, so the customer stays anonymous and the wording stays private. Keeping it unpublished also keeps it a controlled variable: wording that circulates publicly is wording that ends up in training data, and a benchmark whose hardest task can be memorised stops measuring anything. What is published is the shape — question count, the instruction not to shorten, and the field-by-field check the submitted answers have to survive — which is enough to write an equivalent prompt of your own and compare the result against the dated evidence.
The benchmark sits in a longer tradition. It borrows the discipline of evals, fixed tasks with fixed scoring, and applies the LLM-as-a-judge idea to results you can actually verify instead of vibes: an app either opens or it does not, a field either lands or it does not.
What Does TSK-1 Measure?
Every build climbs five gates in order:
- Does it open. The finished build actually opens and runs. A model that says "built" but hands back an app that does not run fails here.
- The colour check. The finished theme holds together throughout. Broken colours and orphan styles fail here.
- A scripted person submits the form. Somebody fills in the built app's form and submits it. If the submission fails, the build fails.
- Field-by-field check of what was saved. The submitted answers are checked one by one against Workspace DNA. Data that lands in the wrong fields or never lands at all fails here.
- Wrong-app penalty. The build made the product that was asked for. A customer database delivered instead of a sign-up form is capped regardless of how pretty it is.
The four gates after the first are what separate this benchmark from code-level tests like SWE-bench, which grades bug-fix patches on existing codebases. TSK-1 grades the full product: it opens, it runs, it saves your data, the data is intact, and it is the app you asked for.
Who Is Benchmarked?
Seven model families have been tested; each has its own evidence page on the hub:
- Claude (Anthropic): the cleanest code we have measured (Aug 3, 2026) and the best-looking build we have measured (Jul 30, 2026), with the builds that lost their thread reported honestly. See Claude's evidence.
- GPT (OpenAI): builds what was asked for most closely, and the fastest. Luna echoed the 32-question sign-up form word for word from Aug 3 to Aug 19, 2026, and Terra posted the first run where every stage worked (Aug 7, 2026). See GPT's evidence.
- DeepSeek: the value story. Flash, the cheapest model in the field, made the best-looking app (Aug 1, 2026), and Pro wired up more data than anything else we have measured, with 8 automations and 60 fields (Aug 6, 2026). See DeepSeek's evidence.
- Gemini (Google): the story of the opening check. It did not place on Aug 1, 2026 when the build it turned in would not open, then opened and shipped on Aug 3, 2026. See Gemini's evidence.
- GLM (Zhipu): the behavioural finding. The only model to decline the unrequested login screen and offer it as a suggestion instead (Jul 30, 2026). See GLM's evidence.
- Kimi (Moonshot): the fewest failed steps of that test at 7.2% (Jul 31, 2026). See Kimi's evidence.
- MiniMax: the cautionary finding. On Jul 31, 2026, 47.2% of its steps failed, and it still summarised its work as a success over its own record of what had gone wrong. See MiniMax's evidence.
Two more families, Qwen and Grok, are available inside Taskade with tests pending, and each has a page on the hub. Taskade offers 15+ frontier models from OpenAI, Anthropic, Google, and open-weight providers.
What the Tests Found
Across every test from late July to late August 2026, no model won everything. What we found clusters into five findings, each carrying the day it was measured:
- One model sticks to your wording better than the rest. GPT-5.6 Luna echoed all 32 questions of the customer's sign-up form word for word, repeatedly, and is the only model that builds a real scoring grid you can use (Aug 3 to Aug 19, 2026).
- Value and design are not in opposition. The cheapest model in the Aug 1, 2026 field, DeepSeek V4 Flash, made the best-looking app, with a coherent light and dark theme, nothing broken on the page, and a clean layout on a phone.
- A build that loses its thread is its own failure class. On Aug 3, 2026 one build lost the customer's wording partway through and finished as the wrong app, which is why wrong-app penalties exist.
- Silent extras are a real failure class. GPT-5.6 Sol shipped an unrequested login screen twice (Jul 30 and Aug 1, 2026) and never mentioned it. GLM-5.2 thought about the same screen and declined it, offering it as a suggestion instead (Jul 30, 2026). The right behaviour under ambiguity.
- Reliability is a gate. On Jul 31, 2026, 47.2% of MiniMax M3's steps failed and it still summarised its work as a success over its own record of what had gone wrong, and Gemini 3.6 Flash did not place on Aug 1, 2026 when the build it turned in would not open.
The point is not a podium. It is that the failure modes differ per model, are measurable, and repeat. A benchmark that grades the finished app catches all of them.
How to Read the Scores
The benchmark publishes evidence first and an order second. The hub does rank the models, but only coarsely — four steps on each of four measures — and models that come out level share a position. Three rules govern how to read it:
- Honest sample size. Each cell is 1-3 builds per model per test. Positions are directional across tests and evidence-graded; ties are real and are shown as ties.
- Dated claims. Every evidence claim on a model page cites the day it was measured, and the page will not publish if one of them points at a test that is not in the record.
- Cost discipline. Cost comparisons use relative ratios from comparable tests only, never per-model credit counts or dollar anchors.
The cost rule exists because early tests predate a billing change, so comparing costs across all of them would be meaningless. Honest sample sizes exist because one strong test is a data point, not a title. Findings accumulate directionally across tests instead of averaging into one aggregate score.
No model has won every measure, and none is expected to. The finding to internalize is that different models win different kinds of work: sticking to your wording, design, speed, code health, or how much data they wire up. That is exactly why automatic model routing picks the model per task instead of naming one winner.
Learn More
- TSK-1 (Taskade System Kernel), the kernel the benchmark evidences
- TSK-1 Benchmark Hub, dated result summaries and every model profile
- Evals, the plain-English definition of evaluation
- LLM-as-a-Judge, rubric-based automated grading
- Workspace DNA, where the field-by-field check actually lands
- The History of AI Benchmarks, why every leaderboard saturates
- TSK-1 Benchmark Methodology, the full nine-model breakdown
Ready to see your own model's evidence? Run the same fixed prompts in Taskade Genesis or clone an example app and grade the build yourself.