download dots

TSK-1 · xAI · tested in Taskade Genesis

How Grok builds in Taskade Genesis

Updated · Last measured

Your requestTaskade EVETSK-1Workspace memoryLive app

TSK-1 Score

Aggregated from hands-on Taskade Genesis builds · September 2026

TSK-1 Score 69 out of 100
Interfaceready to share
Strong
Taskfollows your request
Emerging
Memorykeeps your data
Strong
Adapthandles follow-up edits
Strong

Best for

Follow-up edits that land, at a high price

Grok gets there, slowly. Both of its August apps worked and both follow-up changes landed, which few models manage. Cost was the story: it repeated failed steps dozens of times and spent far more than the leaders on the same requests.

Model results

Grok 4.6Tested Aug 20264 findings

Working apps on both tests, and both follow-up edits landed. It struggled to hand data to the workspace, repeating failed steps dozens of times, so each build took far longer and cost far more than the leaders.

Aug 25, 2026

Built a working match tracker with every asked piece, then ended without telling the user it was done.

Aug 25, 2026

Took the full 32-question sign-up form and saved every answer, but spread the data across ten projects where the leader used two.

Aug 25, 2026

Both follow-up edits landed: a ranked heroes page, and a text score field with next-step suggestions.

Aug 25, 2026

Dozens of failed build actions on each app made it one of the slowest and most expensive results of the test.

Grok Build 0.1Tested Aug 20261 finding

Not scored. It described every layer of the build, then repeated the same failed step hundreds of times until the test was stopped. A newer version can earn a fresh test.

Aug 25, 2026

Produced no projects, agents, or app: hundreds of failed build actions and no finished work.

App kits you can open today

Live apps from the official Taskade account, not builds by Grok. Each is the same shape every benchmark request asks for: projects, agents, and automations you can open and clone.

Browse all App Kits →

Compare Grok

FAQ

What is Grok best at in Taskade Genesis?

Follow-up changes. Grok 4.6 was one of the few models to land both edits we asked for on its August apps, and both apps worked. The trade-off is cost and time: it repeated failed steps dozens of times and spent far more than the leaders on the same requests.

Why is Grok Build 0.1 listed with no result?

Because it produced no app. It described every layer of the build and then repeated the same failed step hundreds of times until we stopped the test. We publish the tests that go badly alongside the ones that go well.

Can I use Grok in Taskade?

Yes. Grok is available in Taskade alongside 15+ frontier models from OpenAI, Anthropic, and open-weight providers.