download dots

TSK-1 · Anthropic · tested in Taskade Genesis

How Claude builds in Taskade Genesis

Updated · Last measured

Your requestTaskade EVETSK-1Workspace memoryLive app

TSK-1 Score

Aggregated from hands-on Taskade Genesis builds · September 2026

TSK-1 Score 81 out of 100
Interfaceready to share
Strong
Taskfollows your request
Strong
Memorykeeps your data
Strong
Adapthandles follow-up edits
Leading

Best for

Polished apps that keep improving

Claude produces the most polished finished apps we test and handles follow-up changes especially well. Its main risk is losing details in a long request.

Model results

Claude Sonnet 5Tested Aug 20265 findings

Our most carefully finished result. It checked its own app, but can miss details in long requests.

Jul 30, 2026

Opened its own app and checked the assistant's answers — the only model to verify its work this way.

Aug 3, 2026

Most carefully finished result, with eight minor items left.

Aug 3, 2026

Built a contacts database instead of the requested sign-up form after two stalls.

Aug 6, 2026

Four attempts stalled before the sign-up form was completed.

Aug 25, 2026

Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.

Claude Opus 5Tested Aug 20264 findings

Our best-looking result, with a complete workspace and an assistant that understood its contents.

Jul 30, 2026

Set the design high-water mark with a polished, consistent interface.

Aug 1, 2026

Created the fullest workspace, with an assistant that understood its contents.

Aug 25, 2026

The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.

Aug 25, 2026

Best-executed match tracker of the test, without a single failed build action, at by far the highest cost.

Claude Haiku 4.5Tested Aug 20262 findings

Caught and repaired a problem that would have left the app blank.

Jul 31, 2026

Found the problem while building, repaired it, and then finished — the only model in the test to recover on its own.

Aug 25, 2026

Fastest build of the test, but it left out the weekly recap automation the brief asked for.

App kits you can open today

Live apps from the official Taskade account, not builds by Claude. Each is the same shape every benchmark request asks for: projects, agents, and automations you can open and clone.

Browse all App Kits →

Compare Claude

FAQ

What is Claude best at in Taskade Genesis?

Claude is strongest when finish and follow-up changes matter. Sonnet produced our most carefully finished result, while Opus set the visual quality high-water mark. With long requests, review the finished app to make sure every detail carried through.

Sonnet or Opus for app building?

Sonnet 5 is the stronger all-round choice. It checked its own finished app and handled changes especially well. Opus 5 is the visual quality leader when the final finish matters most.

Why did Claude fail a test?

In one August test it built a contacts database instead of the client sign-up form, after two three-minute stalls lost it the thread. In another, four stalled runs in a row meant the client sign-up form never got built at all. We publish the tests that go badly alongside the ones that go well, and we open every app and read it back against the brief. That is the point of the benchmark.

Can I use Claude in Taskade?

Yes. Claude Sonnet 5, Opus 5, and Haiku 4.5 are all available in Taskade. Taskade offers 15+ frontier models from OpenAI, Anthropic, and open-weight providers.