download dots

Anthropic · tested in Taskade Genesis

How Claude builds in Taskade Genesis

Claude is not in the Taskade model picker. It runs on your own Claude API key, in automations and in custom agents on Enterprise.

Last test

TSK-1 scorecard for Claude

TSK Score

Task completion across real apps built in Taskade Genesis

TSK Score 75 out of 100
Interfaceready to share
Strong
Taskcompletes your task
Strong
Memorykeeps your data
Emerging
Adapthandles follow-up edits
Leading

Best for

Polished apps that keep improving

Claude produces the most polished finished apps we test and handles follow-up changes especially well.

More

Its main risk is losing details in a long request. Opus 5.5, the newest, shared the top result of our latest phone-form test.

Where Claude ranks

TSK Score out of 100.

  1. 1GPT88 out of 100
  2. 2Claude75 out of 100
  3. 2DeepSeek75 out of 100
  4. 2Kimi75 out of 100
  5. 2Grok75 out of 100
  6. 6GLM63 out of 100
  7. 7Qwen50 out of 100
  8. 8Gemini44 out of 100
Claude versions over timeDays each version was tested
Claude versions over time: days tested so far, Jul 30 to Oct 8
DateClaude Opus 5.5Claude Sonnet 5Claude Sonnet 4.5Claude Opus 5Claude Haiku 4.5
01, tested this day01, tested this day0
01011, tested this day
0102, tested this day1
02, tested this day021
03, tested this day021
04, tested this day03, tested this day2, tested this day
05, tested this day04, tested this day2
1, tested this day5042
2, tested this day6, tested this day05, tested this day2
27, tested this day1, tested this day53, tested this day

Versions tested

Open a version to read its findings.

Claude Opus 5.5Tested Sep 20264 findings

The newest Claude we test. On a family intake form it saved every answer and both photos, turned away an oversized photo before the upload, and was the only build of its test to open in the warm cream look the request asked for.

Sep 22, 2026

Asked where the weekly recap should go and who would use the tracker before it built, then built three pages and a coaching automation that wrote a note into a logged match, at a far higher cost than the economical models.

Sep 30, 2026

Built the family intake form so it saved every answer and both photos, and it turned away a 12 MB photo before the upload started.

Sep 30, 2026

The only build of its test that opened in the light cream look the request asked for.

Sep 30, 2026

Added the new teacher question and put the change live. Its class assistant answered questions about the class but could not count the families.

Claude Sonnet 5Tested Oct 202610 findings

Our most carefully finished result in August. In September, one dashboard could not read its own sample data and a longer build stopped before the app was written. Its latest build, a family intake form, saved every answer and both photos.

Jul 30, 2026

Opened its own app and checked the assistant's answers: the only model to verify its work this way.

Aug 3, 2026

Most carefully finished result, with eight minor items left.

Aug 3, 2026

Built a contacts database instead of the requested sign-up form after two stalls.

Aug 6, 2026

Four attempts stalled before the sign-up form was completed.

Aug 25, 2026

Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.

Sep 19, 2026

Typed an 81-row scoring rubric and a 36-field sign-up table, then stopped at an approval prompt before writing the app.

Sep 19, 2026

Built a three-page match tracker whose dashboard reported zero matches over the eight matches it had entered itself.

Sep 30, 2026

Built the family intake form so it saved every answer and both photos, then added the new teacher question and put the change live.

Oct 8, 2026

Built the family intake form twice. Both builds saved every answer and both photos and took the new teacher question, and both showed sample families on the teacher page with no label.

Oct 8, 2026

Built the Dota 2 match tracker twice, each with a win-rate trend on the dashboard, and both saved a logged match. One ran wider than a phone screen. In the other the change added a second heroes page beside the one it had already built.

Claude Sonnet 4.5Tested Oct 20262 findings

An older Claude, tested in October. Its match tracker rejected every logged match and read every sample match as a loss, while its family intake form saved every answer and both photos.

Oct 8, 2026

Built the family intake form so it saved every answer and both photos, in a plain look with little of the warm pastel the request asked for. Its open teacher page listed parent phone numbers.

Oct 8, 2026

Built a Dota 2 match tracker that rejected every logged match and showed every sample match as a loss, so the dashboard read a 0% win rate.

Claude Opus 5Tested Sep 20267 findings

The most complete client sign-up build of September: five pages, every question in the customer's wording, and a scoring automation that ran and saved its verdict.

Jul 30, 2026

Set the design high-water mark with a polished, consistent interface.

Aug 1, 2026

Created the fullest workspace, with an assistant that understood its contents.

Aug 25, 2026

The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.

Aug 25, 2026

Best-executed match tracker of the test, without a single failed build action, at by far the highest cost.

Sep 19, 2026

Finished the client sign-up form with all 32 questions in the customer's wording, 31 word for word, five pages, and a scoring automation that saved its result.

Sep 19, 2026

Decided two open points in the brief and named both decisions in its closing summary.

Sep 30, 2026

Built the family intake form so it saved every answer and both photos, and its email automation ran three times. It cost far more than Claude Opus 5.5 on the same request for a lower result.

Claude Haiku 4.5Tested Oct 20264 findings

Caught and repaired a problem that would have left the app blank.

Jul 31, 2026

Found the problem while building, repaired it, and then finished: the only model in the test to recover on its own.

Aug 25, 2026

Fastest build of the test, but it left out the weekly recap automation the brief asked for.

Oct 8, 2026

Built the family intake form in under 4 minutes and took the new teacher question, but the family photo replaced the child photo on every submission.

Oct 8, 2026

Built a Dota 2 match tracker that saved logged matches and added a heroes page when asked, but its top-heroes chart showed empty axes.

Same-request tests

The family intake form, more models

All models →
The request

Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.

  • Claude Sonnet 5×2 buildsWorks

    Both builds saved every answer and both photos, and the new teacher question landed. Sample families showed on the teacher page with no label.

  • Claude Haiku 4.5Works, with gaps

    Saved every answer, but the family photo replaced the child photo on every submission.

  • Claude Sonnet 4.5Works, with gaps

    Saved every answer and both photos in a plain look, and its open teacher page listed parent phone numbers.

A Dota 2 match tracker with a heroes page

All models →
The request

One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.

  • Claude Sonnet 5×2 buildsWorks

    Both builds put a win-rate trend on the dashboard and saved a logged match. One ran wider than a phone screen, and in the other the change added a second heroes page.

  • Claude Haiku 4.5Works, with gaps

    Saved logged matches and added the heroes page, but its top-heroes chart showed empty axes.

  • Claude Sonnet 4.5Works, with gaps

    Rejected every logged match and showed every sample match as a loss, so the win rate read 0%.

A family intake form on a phone

All models →
The request

Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.

  • Claude Opus 5.5Works

    Saved every answer and both photos, turned away an oversized photo, and opened in the cream look the request asked for. The follow-up change landed.

  • Claude Opus 5Works

    Saved every answer and both photos, and its email automation ran. The follow-up change landed.

  • Claude Sonnet 5Works

    Saved every answer and both photos. The follow-up change landed.

A match tracker

All models →
The request

One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.

  • Claude Opus 5.5Works

    Asked two questions first, then built three pages and a coaching automation, at a far higher cost.

One build per model unless marked ×N. A direction, not a final rank.

App kits you can open today

Live apps from the official Taskade account, not builds by Claude.

Browse all App Kits →

Compare Claude

See all 11 comparisons

Other models in TSK-1

See the full TSK-1 ranking →

FAQ

What is Claude's TSK Score?

Claude scores 75 out of 100 on TSK-1: Interface Strong, Task Strong, Memory Emerging, Adapt Leading. Last test Oct 8, 2026.

What is Claude best at in Taskade Genesis?

Claude is strongest when finish and follow-up changes matter. Sonnet produced our most carefully finished result, while Opus set the visual quality high-water mark. With long requests, review the finished app to make sure every detail carried through.

Which Claude model is best for app building?

Claude Opus 5.5 shared the top result of our latest phone-form test and opened in the look the request asked for. Claude Opus 5 produced the more complete September sign-up app, with five pages and a scoring automation that ran. Claude Sonnet 5 remains strong at careful finish and follow-up changes, but review its dashboards against the saved data.

Why did Claude fail a test?

In one August test it built a contacts database instead of the client sign-up form, after two three-minute stalls lost it the thread. In another, four stalled runs in a row meant the client sign-up form never got built at all. We publish the tests that go badly alongside the ones that go well, and we open every app and read it back against the brief. That is the point of the benchmark.

Can I use Claude in Taskade?

Yes, with your own Claude API key: use it in the Anthropic Claude automation step, or connect it to your custom agents on Enterprise. Claude is not in the Taskade model picker.

Where can I compare Claude with other models, or try it myself?

Side-by-side model pages live at /compare. Claude is not in the model picker, so /create starts a build with Auto or another available model. This page stays on what Claude produced in TSK-1 tests.