
xAI · tested in Taskade Genesis
How Grok builds in Taskade Genesis
Use Grok in Taskade Genesis for follow-up edits that land, at a high price.
Last test
TSK-1 scorecard for Grok
TSK Score
Task completion across real apps built in Taskade Genesis
- Interfaceready to share
- Strong
- Taskcompletes your task
- Strong
- Memorykeeps your data
- Strong
- Adapthandles follow-up edits
- Strong
Grok gets there, slowly.
MoreLess
Both of its August apps worked and both follow-up changes landed, which few models manage. Cost was the story: it repeated failed steps dozens of times and spent far more than the leaders on the same requests. In October, Grok 4.7 and Grok 4.6 each built a family intake form that saved every answer and both photos.
Where Grok ranks
TSK Score out of 100.
| Date | Grok 4.7 | Grok 4.6 | Grok 4.5 | Grok 4.3 | Grok 4.20 | Grok Build 0.1 |
|---|---|---|---|---|---|---|
| 0 | 1, tested this day | 0 | 0 | 0 | 1, tested this day | |
| 0 | 2, tested this day | 0 | 0 | 0 | 1 | |
| 1, tested this day | 3, tested this day | 1, tested this day | 1, tested this day | 1, tested this day | 2, tested this day |
Versions tested
Open a version to read its findings.
Grok 4.7Tested Oct 20264 findings
The newest Grok in the Taskade model picker. The 15-minute test limit ended its first complete build before a closing summary, and the family intake form it built saved every answer and both photos and took a new question cleanly.
Built the family intake form so it saved every answer and both photos, and it refused a submission with a required answer missing.
The 15-minute test limit ended the build while it checked its own form, before a closing summary.
Added the new teacher question to the form and the teacher page, and said plainly that it had not tested a photo from a real phone.
Built the best Dota 2 match tracker of the Grok set: a three-page esports dashboard with labeled sample matches, a coach and a weekly automation. A logged match saved, and the heroes page landed. The 15-minute limit ended the build while it polished the look.
Grok 4.6Tested Oct 20266 findings
Working apps on both August tests, and both follow-up edits landed. In September it built rich sign-up data across four pages, then added a sign-in screen nobody asked for.
Built a working match tracker with every asked piece, then ended without telling the user it was done.
Took the full 32-question sign-up form and saved every answer, but spread the data across ten projects where the leader used two.
Both follow-up edits landed: a ranked heroes page, and a text score field with next-step suggestions.
Dozens of failed build actions on each app made it one of the slowest and most expensive results of the test.
Took the full 32-question sign-up form into a 58-field table and four pages, with 31 of 32 questions word for word, then added a sign-in screen.
Built the family intake form in 14 minutes so it saved every answer and both photos, and added the new teacher question cleanly. Its teacher page showed three sample families as if they were real.
Grok 4.5Tested Oct 20263 findings
Not in the Taskade model picker. In its first test it built the most complete Grok family intake form: every answer and both photos saved, labeled sample families, a class assistant and a weekly automation.
Built the family intake form in 7 minutes 25 seconds so it saved every answer and both photos, with labeled sample families, a class assistant and a weekly automation.
Added the new teacher question to the form and to every family card, and said plainly that it had not tested the photo picker.
Built a four-page Dota 2 match tracker with labeled sample matches that saved a logged match and added the heroes page, but its hero chart showed empty axes.
Grok 4.3Tested Oct 20263 findings
Not in the Taskade model picker. Its family intake form had hard-to-read fields and no link to the teacher page, yet it saved every answer and both photos in under 4 minutes, and the new question landed.
Built the family intake form in 3 minutes 21 seconds so it saved every answer and both photos, but its form fields were hard to read and no link led to the teacher page.
Added the new teacher question to the form and the family cards.
Built a Dota 2 match tracker in under 5 minutes that saved a logged match and added the heroes page, but its average KDA card showed a garbled number and the board ran wider than a phone screen.
Grok 4.20Tested Oct 20263 findings
Not scored, and not in the Taskade model picker. Neither of its two October builds kept a parent's answers. A newer version can earn a fresh test.
With reasoning on, it described the plan, stopped before writing an app, and wrote the app only on the follow-up request. That form rejected every submission and showed no message.
With reasoning off, it built a dark form where the request asked for cream paper, and the saved row lost the child's name and kept only the photo file names.
On a Dota 2 match tracker, neither setting could log a match. With reasoning on, every save was rejected and the board showed made-up stats. With reasoning off, it built no app page, and its summary described a heroes page that did not exist.
Grok Build 0.1Tested Oct 20264 findings
Its Dota 2 match tracker worked in October, with a logged match saved and a heroes page added. That month its family intake form dropped four of its 14 answers, one of them required. In August it built no app and repeated the same failed step hundreds of times, so October is its first test with working apps.
Produced no projects, agents, or app: hundreds of failed build actions and no finished work.
Built the family intake form so the photos landed on every row, but 4 of 14 answers did not save, among them the required photo permission, while the form thanked the parent.
Added the new teacher question, saved it, and showed it on the teacher page.
Built a four-page Dota 2 match tracker with a win-rate trend that saved a logged match and added the heroes page when asked. Its sample matches carried no sample label.
Same-request tests
The family intake form, more models
All models →The request
Parents fill in a form on their phone with two to four photos, the teacher sees every family on one page, and then we ask for one more question on the form.
Grok 4.7Works
Saved every answer and both photos, and the new teacher question landed. The 15-minute limit ended the build before its closing summary.
Grok 4.6Works
Saved every answer and both photos, and the new teacher question landed. Its teacher page showed three sample families with no sample label.
Grok 4.5Works
Saved every answer and both photos, labeled its sample families, and added a class assistant and a weekly automation. The new teacher question landed.
Grok 4.3Works, with gaps
Saved every answer and both photos in under 4 minutes, but its form fields were hard to read and no link led to the teacher page.
Grok Build 0.1Works, with gaps
The photos landed, but 4 of 14 answers did not save, among them a required one, while the form thanked the parent.
Grok 4.20×2 buildsWorks, with gaps
Neither build kept a parent's answers: one lost the child's name, and the other rejected every submission with no message.
A Dota 2 match tracker with a heroes page
All models →The request
One player logs each match with hero, result, kills, deaths, assists, duration and notes, a dashboard shows win rate and streaks in a premium esports look, and then we ask for a heroes page.
Grok 4.7Works
A three-page esports dashboard with labeled sample matches, a coach and a weekly automation. A logged match saved, and the heroes page landed.
Grok 4.5Works
Four pages with labeled sample matches, a logged match saved, and the heroes page landed. Its hero chart showed empty axes.
Grok Build 0.1Works
Four pages with a win-rate trend, a logged match saved, and the heroes page landed. Its sample matches carried no label.
Grok 4.3Works, with gaps
Saved a logged match and added the heroes page, but its average KDA card showed a garbled number.
Grok 4.20×2 buildsWorks, with gaps
Neither setting could log a match: one rejected every save, and the other built no app page while its summary described one.
One build per model unless marked ×N. A direction, not a final rank.
App kits you can open today
Live apps from the official Taskade account, not builds by Grok.
Product storefront PORTALProduct storefront3 projects · 1 agent · 4 flowsSales CRM CRMSales CRM3 projects · 1 agent · 2 flowsSales pipeline dashboard DASHBOARDSales pipeline dashboard3 projects · 1 agentClient portal PORTALClient portal3 projects · 1 agent · 4 flowsGrowth dashboard DASHBOARDGrowth dashboard3 projects · 1 agent · 1 flowLead pipeline tracker CRMLead pipeline tracker4 projects · 2 agents · 3 flows
Compare Grok
Other models in TSK-1
FAQ
What is Grok's TSK Score?
Grok scores 75 out of 100 on TSK-1: Interface Strong, Task Strong, Memory Strong, Adapt Strong. Last test Oct 8, 2026.
What is Grok best at in Taskade Genesis?
Follow-up changes. Grok 4.6 was one of the few models to land both edits we asked for on its August apps, and both apps worked. The trade-off is cost and time: it repeated failed steps dozens of times and spent far more than the leaders on the same requests.
How did Grok Build 0.1 do?
In August it produced no app: it described every layer of the build and then repeated the same failed step hundreds of times until we stopped the test. In October it built a working Dota 2 match tracker, and its family intake form lost four of 14 answers. We publish the tests that go badly alongside the ones that go well.
Can I use Grok in Taskade?
Yes. Grok 4.7, Grok 4.6 and Grok Build 0.1 are in the Taskade model picker on Max and Enterprise. Grok 4.7, the newest, saved every answer and both photos on a family intake form in its first complete test.
Where can I compare Grok with other models, or try it myself?
Side-by-side model pages live at /compare. To try the same kind of app request yourself, start from /create. This page stays on what Grok produced in TSK-1 tests.





