Sales CRM
CRM · 3 projects · 1 agent · 2 flows

TSK-Bench · Task completion benchmark for AI models
Same request. Every AI. Scored on what it actually finished.
GPT follows long briefs word for word. Claude finishes apps best, on your own key in automations and Enterprise agents. DeepSeek leads on design. Kimi builds rich apps, but slowly. Not sure? TSK-1 Auto picks for you.
| Rank | Model | Interface | Task | Memory | Adapt | TSK Score | Newest version tested |
|---|---|---|---|---|---|---|---|
| 1 | GPT | Strong | Leading | Strong | Leading | 88 out of 100 | GPT-5.6 Luna |
| 2 | Claude | Strong | Strong | Emerging | Leading | 75 out of 100 | Claude Sonnet 5 |
| 2 | DeepSeek | Leading | Strong | Strong | Emerging | 75 out of 100 | DeepSeek V4.1 Flash |
| 2 | Kimi | Strong | Strong | Strong | Strong | 75 out of 100 | Kimi K3 |
| 2 | Grok | Strong | Strong | Strong | Strong | 75 out of 100 | Grok 4.7 |
| 6 | GLM | Emerging | Strong | Emerging | Strong | 63 out of 100 | GLM-5 |
| 7 | Qwen | Emerging | Strong | Emerging | Limited | 50 out of 100 | Qwen 3.7 Plus |
| 8 | Gemini | Limited | Emerging | Emerging | Emerging | 44 out of 100 | Gemini 3.6 Flash |
| Not ranked | MiniMax | Not scored | Not scored | Not scored | Not scored | Not scored | MiniMax M3 |
| Not ranked | Mistral | Not scored | Not scored | Not scored | Not scored | Not scored | Mistral Large 4 |
| Date | GPT | Claude | DeepSeek | Kimi | Grok | GLM | Qwen | Gemini | MiniMax | Mistral |
|---|---|---|---|---|---|---|---|---|---|---|
| 1, tested this day | 1, tested this day | 0 | 0 | 0 | 1, tested this day | 0 | 0 | 0 | 0 | |
| 1 | 2, tested this day | 0 | 1, tested this day | 0 | 1 | 0 | 0 | 1, tested this day | 0 | |
| 2, tested this day | 3, tested this day | 1, tested this day | 1 | 0 | 2, tested this day | 0 | 1, tested this day | 1 | 0 | |
| 2 | 3 | 2, tested this day | 1 | 0 | 2 | 0 | 1 | 1 | 0 | |
| 3, tested this day | 4, tested this day | 3, tested this day | 1 | 0 | 2 | 0 | 2, tested this day | 1 | 0 | |
| 4, tested this day | 4 | 3 | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 4 | 4 | 4, tested this day | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 4 | 5, tested this day | 5, tested this day | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 5, tested this day | 5 | 5 | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 6, tested this day | 5 | 5 | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 7, tested this day | 5 | 5 | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 7 | 5 | 5 | 1 | 0 | 2 | 0 | 2 | 1 | 0 | |
| 8, tested this day | 6, tested this day | 6, tested this day | 2, tested this day | 1, tested this day | 3, tested this day | 1, tested this day | 2 | 2, tested this day | 0 | |
| 8 | 6 | 6 | 2 | 1 | 3 | 1 | 2 | 2 | 0 | |
| 9, tested this day | 6 | 6 | 2 | 1 | 3 | 1 | 2 | 2 | 0 | |
| 10, tested this day | 6 | 6 | 2 | 1 | 3 | 1 | 2 | 2 | 0 | |
| 11, tested this day | 6 | 6 | 2 | 1 | 4, tested this day | 1 | 2 | 2 | 0 | |
| 11 | 6 | 6 | 2 | 1 | 4 | 1 | 2 | 2 | 0 | |
| 12, tested this day | 7, tested this day | 7, tested this day | 3, tested this day | 2, tested this day | 5, tested this day | 1 | 2 | 2 | 0 | |
| 13, tested this day | 8, tested this day | 8, tested this day | 3 | 2 | 5 | 1 | 2 | 2 | 0 | |
| 14, tested this day | 9, tested this day | 9, tested this day | 4, tested this day | 2 | 5 | 1 | 2 | 2 | 0 | |
| 15, tested this day | 10, tested this day | 10, tested this day | 4 | 3, tested this day | 6, tested this day | 1 | 2 | 2 | 1, tested this day |
Describe the app you need. TSK-1 coordinates the rest.
What each test found.
More than 20 models and settings rebuilt the family intake form and a Dota 2 match tracker, and then took one follow-up change each. Grok 4.7 and GPT-5.6 Sol built both apps well, several builds showed sample families or matches as if they were real, and a few forms could not save at all.
A real customer's family intake form joined the test set: parents fill it in on a phone with two to four photos, and the teacher sees every family on one page. GPT-6.1 Sol and Claude Opus 5.5 shared the top result, and every model that finished also added a follow-up question cleanly.
Several models built the same match tracker. GPT-5.6 Luna was the fastest and most economical complete build, DeepSeek V4.1 Flash ran a coaching automation on the first match logged, and two models limited their match forms to a short list of heroes.
Three real-customer app shapes joined the set: a sales CRM, a fleet inspection, and an AI governance office. The first full follow-up-edit ladder passed on every model that ran it.
Seventeen models built the same client sign-up form. A separate test of parallel helpers showed that one file per helper produced the cleanest result.
A control run told us how big a difference has to be before it counts. Two identical setups produced noticeably different costs for the same app, so we now only claim a difference when it is larger than that gap, and we grade on what the app does.
We re-ran the automatic model after a large platform update. The apps came out the same, with fewer failed build actions along the way.
Five models built and edited the same apps from identical prompts. GPT-5.6 Luna was the value pick on every app and GPT-5.6 Terra the quality pick. Most builds now keep a checklist of the request inside the app.
Follow-up edits now start from a map of the app, and they landed cleaner. GPT-5.6 Luna edited apps other models had built without breaking them, and a missing page now explains itself.
Qwen and Grok entered the benchmark, and three real-customer app shapes joined the test set: a sales CRM, a field-inspection audit, and an AI governance office. Every model still gets the same words.
We added a dedicated follow-up-edit pass and stricter instruction checks. On its first run, none of the 31 edit attempts finished - a strong first build is not enough on its own; an app must also change cleanly and stay within the brief.
A complete app finished in under five minutes without follow-up. We also rechecked every published result and established a baseline for automatic model selection.
The strongest results reproduced a long brief exactly, balanced build quality with efficiency, and explained their design choices clearly.
An efficient build completed a long form at the lowest measured cost, while the leading result also saved every submitted field correctly.
Several builds followed the brief word for word, and one handled a 60-field form without cutting requirements.
We made a running app the entry requirement. Attractive results that failed to open no longer received a score.
The cleanest test also delivered the lowest measured cost, showing that careful model selection can improve both quality and value.
The strongest builds matched every requested field, stored submitted data correctly, and remained useful from form to follow-up.
Shorter instructions improved brief matching and edits, but sometimes reduced visual quality. The benchmark now balances all four dimensions.
Nine models built the same app. Several produced polished designs, but only working results qualified for comparison.
The most dependable results detected and repaired problems before finishing, while weaker builds reported success too early.
The strongest designs followed the brief without adding unwanted gates or steps. Doing only what was asked became part of quality.
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.
One build per model unless marked ×N. A direction, not a final rank.
Up next: GPT-6 Luna and Nemotron 3 Super
Open a version to see what it built.
GPT has the highest TSK Score, 88 out of 100, but the best pick depends on the app. Some models hold closer to a brief, some finish faster, some look more polished. In Taskade you do not have to choose: TSK-1 Auto handles the default, and you can pick a model yourself when you want more control.
Taskade Genesis can turn one request into CRMs, dashboards, portals, forms, trackers, and internal tools. TSK-1 coordinates the AI model, workspace memory, agents, and automations so the result can store information, answer questions, run workflows, and keep improving.
Whether a model can turn one prompt into a complete, working app. We grade four qualities of a living system: Interface, Task, Memory, and Adaptation: how finished it feels, how closely it follows your request, whether it keeps your data, and how cleanly it handles follow-up changes.
Those benchmarks grade code, a web page, or crowd votes on a screenshot. TSK-1 grades a running app: we open it, enter data the way a customer would, check that the data was kept, run its automation, and ask for a follow-up change. A build that looks right but loses your data does not score.
AI models. Every model builds inside the same app builder, Taskade Genesis, from the same request, word for word, so the builder is held constant. A control run in August showed that two identical setups can still differ, so we only call out a difference between models when it is larger than that gap.
We test the apps. Every build gets opened and used. We fill in its form the way a customer would, submit it, then check that the answers landed in the right place. An app that looks beautiful and loses your data fails.
Taskade runs frontier models from top AI labs, and TSK-1 Auto handles the default. Built in, with no key, every plan can pick DeepSeek V4.1 Flash, and paid plans add OpenAI GPT models, xAI Grok and more open-weight models such as Qwen, Kimi and GLM. Claude runs on your own Anthropic key, in automations and in custom agents on Enterprise. Gemini runs only in automations, on your own Gemini API key. Model lineup as of October 2026.
We run a new test whenever a notable model ships. Every result cleared for public comparison stays on this page, including the ones that did not go well.
TSK-1 is the intelligence behind Taskade Genesis. It brings AI models, workspace memory, agents, and automations together so one request can become a working app. The TSK-1 Benchmark shows how that process performs on real app builds.
Yes. Every model in a comparison receives the same request, word for word. Describe the same idea in Taskade Genesis, choose a model or let TSK-1 Auto decide, and see what it builds.
Not yet one by one. We are listing every finished benchmark build under a single official creator so you can open and clone it. Until then, each model page explains what that model produced, and the App Kits on this page are live systems you can open and clone today.
Each result comes from a small set of hands-on app builds. The TSK Score rescales four evidence tiers to 0-100 for easier comparison; it is not 100 separate checks. Models with the same result share a rank.
CRM · 3 projects · 1 agent · 2 flows
DASHBOARD · 3 projects · 1 agent
PORTAL · 3 projects · 1 agent · 4 flows
DASHBOARD · 3 projects · 1 agent · 1 flow
CRM · 4 projects · 2 agents · 3 flows
DASHBOARD · 5 projects · 2 agents · 4 flows
CRM · 2 projects · 1 agent · 2 flows
OPS · 34 projects · 1 agent · 2 flows
DASHBOARD · 1 project · 1 agent · 1 flow
TRACKER · 2 projects · 1 agent · 3 flows
DASHBOARD · 4 projects · 1 agent · 3 flows
PORTAL · 4 projects · 1 agent · 3 flows