download dots

TSK-1: Taskade System Kernel Benchmarks

TSK-1 brings AI models, workspace memory, agents, and automations together so Taskade Genesis can build a working app, keep its data, and improve it with you. What is TSK-1? →

TSK-1 model benchmark

Interfaceready to shareTaskfollows your requestMemorykeeps your dataAdapthandles follow-up edits

The app request is held constant within each comparison. Working apps only. We test the result, saved data, and follow-up edits. Benchmark history Meet TSK-1 →

Start building

Describe the app you need. TSK-1 coordinates the rest.

Benchmark updates

Plain-language takeaways from every published test. Open any update to read what changed.

A dedicated follow-up-edit test

We added a dedicated follow-up-edit pass and stricter instruction checks. On its first run, none of the 31 edit attempts finished - a strong first build is not enough on its own; an app must also change cleanly and stay within the brief.

Faster builds, verified results

A complete app finished in under five minutes without follow-up. We also rechecked every published result and established a baseline for automatic model selection.

Accuracy and value moved together

The strongest results reproduced a long brief exactly, balanced build quality with efficiency, and explained their design choices clearly.

Lower cost did not mean lower quality

An efficient build completed a long form at the lowest measured cost, while the leading result also saved every submitted field correctly.

Long briefs became a harder test

Several builds followed the brief word for word, and one handled a 60-field form without cutting requirements.

Working software became the baseline

We made a running app the entry requirement. Attractive results that failed to open no longer received a score.

Efficiency without tradeoffs

The cleanest test also delivered the lowest measured cost, showing that careful model selection can improve both quality and value.

Brief fidelity and saved data

The strongest builds matched every requested field, stored submitted data correctly, and remained useful from form to follow-up.

Prompt length changed the outcome

Shorter instructions improved brief matching and edits, but sometimes reduced visual quality. The benchmark now balances all four dimensions.

Nine models, one consistent standard

Nine models built the same app. Several produced polished designs, but only working results qualified for comparison.

Self-checking improved reliability

The most dependable results detected and repaired problems before finishing, while weaker builds reported success too early.

Restraint mattered

The strongest designs followed the brief without adding unwanted gates or steps. Doing only what was asked became part of quality.

FAQ

Which AI model is best for building apps?

It depends on what you are building. Some models hold closer to a brief, some finish faster, some look more polished. In Taskade you do not have to choose: TSK-1 Auto handles the default, and you can pick a model yourself when you want more control.

What can Taskade Genesis build with TSK-1?

Taskade Genesis can turn one request into CRMs, dashboards, portals, forms, trackers, and internal tools. TSK-1 coordinates the AI model, workspace memory, agents, and automations so the result can store information, answer questions, run workflows, and keep improving.

What does the TSK-1 Benchmark measure?

Whether a model can turn one prompt into a complete, working app. We grade four qualities of a living system: Interface, Task, Memory, and Adaptation — how finished it feels, how closely it follows your request, whether it keeps your data, and how cleanly it handles follow-up changes.

Do you test the apps or just the code?

We test the apps. Every build gets opened and used. We fill in its form the way a customer would, submit it, then check that the answers landed in the right place. An app that looks beautiful and loses your data fails.

Which models can I use in Taskade?

Taskade gives you 15+ frontier models from OpenAI, Anthropic, Google, and open-weight providers, all on one subscription. TSK-1 Auto handles the default, and you can choose a model yourself when you want more control.

How often is this updated?

We run a new test whenever a notable model ships. Every result cleared for public comparison stays on this page, including the ones that did not go well.

What is TSK-1?

TSK-1 is the intelligence behind Taskade Genesis. It brings AI models, workspace memory, agents, and automations together so one request can become a working app. The TSK-1 Benchmark shows how that process performs on real app builds.

Can I try the same app request?

Yes. Every model in a comparison receives the same request, word for word. Describe the same idea in Taskade Genesis, choose a model or let TSK-1 Auto decide, and see what it builds.

Are the benchmarked apps publicly available?

The test apps are not published one by one. Each model page explains what that model produced. The App Kits near the bottom of this page are separate live systems you can open, explore, and clone.

How many builds are behind each result?

Each result comes from a small set of hands-on app builds. The Intelligence Index rescales four evidence tiers to 0–100 for easier comparison; it is not 100 separate checks. Models with the same result share a rank.