
Anthropic · tested in Taskade Genesis
How Claude builds in Taskade Genesis
Claude is not in the Taskade model picker. It runs on your own Claude API key, in automations and in custom agents on Enterprise.
Last test
TSK-1 scorecard for Claude
TSK Score
Task completion across real apps built in Taskade Genesis
- Interfaceready to share
- Strong
- Taskcompletes your task
- Strong
- Memorykeeps your data
- Emerging
- Adapthandles follow-up edits
- Leading
Best for
Polished apps that keep improving
Claude produces the most polished finished apps we test and handles follow-up changes especially well.
MoreLess
Its main risk is losing details in a long request. Opus 5.5, the newest, shared the top result of our latest phone-form test.
Where Claude ranks
TSK Score out of 100.
| Date | Claude Opus 5.5 | Claude Sonnet 5 | Claude Sonnet 4.5 | Claude Opus 5 | Claude Haiku 4.5 |
|---|---|---|---|---|---|
| 0 | 1, tested this day | 0 | 1, tested this day | 0 | |
| 0 | 1 | 0 | 1 | 1, tested this day | |
| 0 | 1 | 0 | 2, tested this day | 1 | |
| 0 | 2, tested this day | 0 | 2 | 1 | |
| 0 | 3, tested this day | 0 | 2 | 1 | |
| 0 | 4, tested this day | 0 | 3, tested this day | 2, tested this day | |
| 0 | 5, tested this day | 0 | 4, tested this day | 2 | |
| 1, tested this day | 5 | 0 | 4 | 2 | |
| 2, tested this day | 6, tested this day | 0 | 5, tested this day | 2 | |
| 2 | 7, tested this day | 1, tested this day | 5 | 3, tested this day |
Versions tested
Open a version to read its findings.
Claude Opus 5.5Tested Sep 20264 findings
The newest Claude we test. On a family intake form it saved every answer and both photos, turned away an oversized photo before the upload, and was the only build of its test to open in the warm cream look the request asked for.
Asked where the weekly recap should go and who would use the tracker before it built, then built three pages and a coaching automation that wrote a note into a logged match, at a far higher cost than the economical models.
Built the family intake form so it saved every answer and both photos, and it turned away a 12 MB photo before the upload started.
The only build of its test that opened in the light cream look the request asked for.
Added the new teacher question and put the change live. Its class assistant answered questions about the class but could not count the families.
Claude Sonnet 5Tested Oct 202610 findings
Our most carefully finished result in August. In September, one dashboard could not read its own sample data and a longer build stopped before the app was written. Its latest build, a family intake form, saved every answer and both photos.
Opened its own app and checked the assistant's answers: the only model to verify its work this way.
Most carefully finished result, with eight minor items left.
Built a contacts database instead of the requested sign-up form after two stalls.
Four attempts stalled before the sign-up form was completed.
Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.
Typed an 81-row scoring rubric and a 36-field sign-up table, then stopped at an approval prompt before writing the app.
Built a three-page match tracker whose dashboard reported zero matches over the eight matches it had entered itself.
Built the family intake form so it saved every answer and both photos, then added the new teacher question and put the change live.
Built the family intake form twice. Both builds saved every answer and both photos and took the new teacher question, and both showed sample families on the teacher page with no label.
Built the Dota 2 match tracker twice, each with a win-rate trend on the dashboard, and both saved a logged match. One ran wider than a phone screen. In the other the change added a second heroes page beside the one it had already built.
Claude Sonnet 4.5Tested Oct 20262 findings
An older Claude, tested in October. Its match tracker rejected every logged match and read every sample match as a loss, while its family intake form saved every answer and both photos.
Built the family intake form so it saved every answer and both photos, in a plain look with little of the warm pastel the request asked for. Its open teacher page listed parent phone numbers.
Built a Dota 2 match tracker that rejected every logged match and showed every sample match as a loss, so the dashboard read a 0% win rate.
Claude Opus 5Tested Sep 20267 findings
The most complete client sign-up build of September: five pages, every question in the customer's wording, and a scoring automation that ran and saved its verdict.
Set the design high-water mark with a polished, consistent interface.
Created the fullest workspace, with an assistant that understood its contents.
The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.
Best-executed match tracker of the test, without a single failed build action, at by far the highest cost.
Finished the client sign-up form with all 32 questions in the customer's wording, 31 word for word, five pages, and a scoring automation that saved its result.
Decided two open points in the brief and named both decisions in its closing summary.
Built the family intake form so it saved every answer and both photos, and its email automation ran three times. It cost far more than Claude Opus 5.5 on the same request for a lower result.
Claude Haiku 4.5Tested Oct 20264 findings
Caught and repaired a problem that would have left the app blank.
Found the problem while building, repaired it, and then finished: the only model in the test to recover on its own.
Fastest build of the test, but it left out the weekly recap automation the brief asked for.
Built the family intake form in under 4 minutes and took the new teacher question, but the family photo replaced the child photo on every submission.
Built a Dota 2 match tracker that saved logged matches and added a heroes page when asked, but its top-heroes chart showed empty axes.
Same-request tests
The family intake form, more models
All models →The request
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
Claude Sonnet 5×2 buildsWorks
Both builds saved every answer and both photos, and the new teacher question landed. Sample families showed on the teacher page with no label.
Claude Haiku 4.5Works, with gaps
Saved every answer, but the family photo replaced the child photo on every submission.
Claude Sonnet 4.5Works, with gaps
Saved every answer and both photos in a plain look, and its open teacher page listed parent phone numbers.
A Dota 2 match tracker with a heroes page
All models →The request
One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.
Claude Sonnet 5×2 buildsWorks
Both builds put a win-rate trend on the dashboard and saved a logged match. One ran wider than a phone screen, and in the other the change added a second heroes page.
Claude Haiku 4.5Works, with gaps
Saved logged matches and added the heroes page, but its top-heroes chart showed empty axes.
Claude Sonnet 4.5Works, with gaps
Rejected every logged match and showed every sample match as a loss, so the win rate read 0%.
A family intake form on a phone
All models →The request
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
Claude Opus 5.5Works
Saved every answer and both photos, turned away an oversized photo, and opened in the cream look the request asked for. The follow-up change landed.
Claude Opus 5Works
Saved every answer and both photos, and its email automation ran. The follow-up change landed.
Claude Sonnet 5Works
Saved every answer and both photos. The follow-up change landed.
A match tracker
All models →The request
One player logs each game with hero, result and notes, and a dashboard shows win rate and streaks.
Claude Opus 5.5Works
Asked two questions first, then built three pages and a coaching automation, at a far higher cost.
One build per model unless marked ×N. A direction, not a final rank.
App kits you can open today
Live apps from the official Taskade account, not builds by Claude.
Lead pipeline tracker CRMLead pipeline tracker4 projects · 2 agents · 3 flowsInvestor metrics dashboard DASHBOARDInvestor metrics dashboard5 projects · 2 agents · 4 flowsClient pipeline CRM CRMClient pipeline CRM2 projects · 1 agent · 2 flowsStore command center OPSStore command center34 projects · 1 agent · 2 flowsProduct launch dashboard DASHBOARDProduct launch dashboard1 project · 1 agent · 1 flowOrder and stock tracker TRACKEROrder and stock tracker2 projects · 1 agent · 3 flows
Compare Claude
See all 11 comparisonsShow fewer
Other models in TSK-1
FAQ
What is Claude's TSK Score?
Claude scores 75 out of 100 on TSK-1: Interface Strong, Task Strong, Memory Emerging, Adapt Leading. Last test Oct 8, 2026.
What is Claude best at in Taskade Genesis?
Claude is strongest when finish and follow-up changes matter. Sonnet produced our most carefully finished result, while Opus set the visual quality high-water mark. With long requests, review the finished app to make sure every detail carried through.
Which Claude model is best for app building?
Claude Opus 5.5 shared the top result of our latest phone-form test and opened in the look the request asked for. Claude Opus 5 produced the more complete September sign-up app, with five pages and a scoring automation that ran. Claude Sonnet 5 remains strong at careful finish and follow-up changes, but review its dashboards against the saved data.
Why did Claude fail a test?
In one August test it built a contacts database instead of the client sign-up form, after two three-minute stalls lost it the thread. In another, four stalled runs in a row meant the client sign-up form never got built at all. We publish the tests that go badly alongside the ones that go well, and we open every app and read it back against the brief. That is the point of the benchmark.
Can I use Claude in Taskade?
Yes, with your own Claude API key: use it in the Anthropic Claude automation step, or connect it to your custom agents on Enterprise. Claude is not in the Taskade model picker.
Where can I compare Claude with other models, or try it myself?
Side-by-side model pages live at /compare. Claude is not in the model picker, so /create starts a build with Auto or another available model. This page stays on what Claude produced in TSK-1 tests.





