Skip to main content
Introducing TSK-1Introducing TSK-1·Taskade's intelligence layer.
taskade
PricingHelpDashboard →Dashboard →
PricingLoginSign up for free →Sign up for free →
Dashboard →Dashboard →
Sign up →Sign up →
Loved by 1M+ users·Hosting 100K+ apps·Deploying 500K+ AI agents·Running 1M+ automations·Backed by Y Combinator·Powered by TSK-1
TaskadeCreate an AppPricingFeaturesTSK-1 BenchmarkContact usIntegrationsMCP ServerPressAbout
ConnectProductivityKitsVideosReviewsFAQ
LearnGenesisProjectsAI Agents
AutomationConnectorsAccount & BillingImport & ExportVideo TutorialsSearch Articles
DocsGetting StartedREST APIAction API
MCP ServersGuides & SDKModels
Community
FeaturedQuick AppsToolsDashboardsWebsites
WorkflowsProjectsFormsCreators
DownloadsAndroidiOSMacWindows
ChromeFirefoxEdge
Compare
vs Cursorvs Boltvs Lovablevs V0vs Windsurf
vs Replitvs Emergentvs Devinvs Claude Codevs ChatGPTvs Claudevs Perplexityvs GitHub Copilotvs Figma AIvs Notionvs ClickUpvs Asanavs Mondayvs Trellovs Jiravs Linearvs Todoistvs Evernotevs Obsidianvs Airtablevs Basecampvs Mirovs Slackvs Bubblevs Retoolvs Webflowvs Framervs Softrvs Glidevs FlutterFlowvs Base44vs Adalovs Durablevs Gammavs Squarespacevs WordPressvs UI Bakeryvs Zapiervs Makevs n8nvs Jaspervs Copy.aivs Writervs Rytrvs Manusvs Crewvs Lindyvs Relevance AIvs Wrikevs Smartsheetvs Monday Magicvs Codavs TickTickvs Any.dovs Thingsvs OmniFocusvs MeisterTaskvs Teamworkvs Workfrontvs Bitrix24vs Process Streetvs Toggl Planvs Motionvs Momentumvs Habiticavs Zenkitvs Google Docsvs Google Keepvs Google Tasksvs Microsoft Teamsvs Dropbox Papervs Quipvs Roam Researchvs Logseqvs Memvs WorkFlowyvs Dynalistvs XMindvs Whimsicalvs Zoomvs Remember The Milkvs Wunderlist
Taskade AIVideo GuideApp BuilderVibe CodingAgent BuilderDashboard Builder
CRM BuilderWebsite BuilderForm BuilderWorkflow AutomationWorkflow BuilderBusiness-in-a-BoxAI for MarketingAI for Developers
AI Agents
FeaturedProject ManagementOperations IntelligenceProductivityMarketing
TranslatorContentWorkflowResearchPersonalSalesSocial MediaTo-Do ListCRMTask AutomationCoachingCreativityTask ManagementBrandingFinanceLearning and DevelopmentBusinessCommunity ManagementMeetingsAnalyticsDigital AdvertisingContent CurationKnowledge ManagementProduct DevelopmentPublic RelationsProgrammingHuman ResourcesE-CommerceEducationLegalEmailSEODeveloperVideo ProductionDesignFlowchartDataPromptNonprofitAssistantsTeamsCustomer ServiceTrainingTravel PlanningUML DiagramER DiagramMath TutorLanguage LearningCode ReviewerLogo DesignerUI WireframeFitness CoachLead EnrichmentFounder OSSales DevelopmentBookkeepingRecruitingWebsite MonitoringField ServiceLicensingAll Categories
Automations
FeaturedAI Agent AutomationAI WorkflowsLogic AutomationsTrigger Automations
Agentic Process AutomationAction AutomationsAI Models in WorkflowsAgentic AutomationMulti-Agent AutomationBusiness-in-a-BoxOperations IntelligenceInvestor OperationsEducation & LearningHealthcare & ClinicsReal EstateStripeSalesHR & People OpsField Service & DispatchRenewals & LicensesE-commerceContentMarketingEmailCustomer SupportHubSpotProject ManagementAgentic WorkflowsAppointment SchedulingCalendarReportsSlackWebsiteFormTaskWeb ScrapingWeb SearchChatGPTText to ActionYoutubeLinkedInTwitterGitHubDiscordMicrosoft TeamsWebflowIndustry News & RSS FeedsGoogle WorkspaceManufacturing & OperationsAI Agent TeamsNotion AutomationsProposalBookkeeping & ExpensesClient OnboardingGoogle SheetsGoogle DriveGoogle CalendarGoogle FormsShopifyAsanaAirtableTrelloTodoistMailchimpClickUpGoogle DocsGmailGoogle TasksJiraLinearMicrosoft OutlookTelegramAll Categories
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Templates
FeaturedChatGPTOperations IntelligenceTablePersonal
Project ManagementSalesFlowchartTask ManagementEngineeringEducationDesignTo-Do ListMarketingMind MapGantt ChartOrganizationalPlanningMeetingsTeam ManagementStrategyGamingProductionProduct ManagementStartupRemote WorkY CombinatorRoadmapCustomer ServiceLegalEmailBudgetsContentConsultingE-CommerceStandard Operating Procedure (SOP)Human ResourcesProgrammingMaintenanceCoachingSocial MediaHow-TosResearchMusicTrip PlanningCRMClient OnboardingEmployee OnboardingSOPBug TrackerRecruitment TrackerFormSales PipelineContent CalendarMarketing PlanProduct RoadmapBusiness PlanSWOT Analysis30-60-90 Day PlanInterviewNotion AlternativeKPIStrategic PlanMeeting AgendaInvoiceRisk RegisterIT Asset ManagementKanban BoardChange ManagementCommunication PlanRFPScope of WorkStatement of WorkHelpdeskKnowledge BaseCreative BriefGoal SettingExecutive SummaryGap AnalysisBooking SystemEvent ManagementPortfolio TrackerCustomer Onboarding PortalsClient PortalAgency OperationsFinance TrackingAll Categories
Generators
AI SoftwareNo-Code AI AppAI AppAI WebsiteAI Dashboard
AI FinanceAI Operations IntelligenceAI FormAI AgentAI Client Portal BuilderAI WorkspaceAI ProductivityAI To-Do ListAI WorkflowsAI EducationAI Mind MapsAI FlowchartAI Scrum Project ManagementAI Agile Project ManagementAI MarketingAI Project ManagementAI Social Media ManagementAI BloggingAI Agency WorkflowsAI ContentAI Software DevelopmentAI MeetingAI PersonasAI OutlineAI SalesAI ProgrammingAI DesignAI FreelancingAI ResumeAI Human ResourceAI SOPAI E-CommerceAI EmailAI Public RelationsAI InfluencersAI Content CreatorsAI Customer ServiceAI BusinessAI PromptsAI Tool BuilderAI SEOAI Gantt ChartAI CalendarsAI BoardAI TableAI ResearchAI LegalAI ProposalAI Video ProductionAI Health and WellnessAI WritingAI PublishingAI NonprofitAI DataAI Event PlanningAI Game DevelopmentAI Project Management AgentAI Productivity AgentAI Marketing AgentAI Personal AgentAI Business and Work AgentAI Education and Learning AgentAI Task Management AgentAI Customer Relations AgentAI Programming AgentAI SchemaAI Business PlanAI Pitch DeckAI InvoiceAI Lesson PlanAI Social Media CalendarAI API DocumentationAI Database SchemaAI Marketing PlanAI Sales Pipeline GeneratorAI Course BuilderInternal ToolsBooking SystemReal Estate CRMInventory ManagementAI CRM BuilderAI TimesheetAI DispatchAI NewsletterAI Clinic OperationsAI Directory BuilderAll Categories
Converters
AI Featured ConvertersAI PDF ConvertersAI CSV ConvertersAI Markdown ConvertersAI Prompt to App Converters
AI Data to Dashboard ConvertersAI Workflow to App ConvertersAI Idea to App ConvertersAI Flowcharts ConvertersAI Mind Map ConvertersAI Text ConvertersAI Youtube ConvertersAI Knowledge ConvertersAI Spreadsheet ConvertersAI Email ConvertersAI Web Page ConvertersAI Video ConvertersAI Coding ConvertersAI Task ConvertersAI Kanban Board ConvertersAI Notes ConvertersAI Education ConvertersAI Language TranslatorsAI Business → Backend App ConvertersAI File → App ConvertersAI SOP → Workflow App ConvertersAI Portal → App ConvertersAI Form → App ConvertersAI Schedule → Booking App ConvertersAI Metrics → Dashboard ConvertersAI Game → Playable App ConvertersAI Catalog → Directory App ConvertersAI Creative → Studio App ConvertersAI Agent → Agent App ConvertersAI Audio ConvertersAI DOCX ConvertersAI EPUB ConvertersAI Image ConvertersAI Resume & Career ConvertersAI Presentation ConvertersAI PDF to Spreadsheet ConvertersAI PDF to Database ConvertersAI PDF to Quiz ConvertersAI Image to Notes ConvertersAI Audio to Notes ConvertersAI Email to Tasks ConvertersAI CSV to Dashboard ConvertersAI YouTube to Flashcards ConvertersURL to NotesVideo → SummaryAI Receipts to Expense Tracker ConvertersAI Docs to Knowledge Base ConvertersAI Form to Client Portal ConvertersSpreadsheet to CRMAll Categories
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
Blog
Introducing Taskade TSK-1: The System Kernel Behind Every App (2026)Chat-Native App Builders in 2026: What You Actually Own When the Chat EndsGenerate the Art. Preview the Agent. Put It on Your Domain (2026)
Agentic Automation Explained: Agent vs AI Step (2026)The Scaffolding Tax: Why Less Prompt Beats More (2026)History of Mind Mapping: From Porphyry to Buzan to AI (2026)The Bitter Lesson Explained: Richard Sutton's 26 Words (2026)Self-Replicating Code: Quines, von Neumann, and the Programs That Copy Themselves (2026)Markov Chains Explained: The Memoryless Math Behind Google, Monte Carlo, and ChatGPT (2026)Add Client Logins. Connect Your Domain. Ship a Real Product in 2026Track Customer Health. Catch Churn Early. Keep the Accounts You Won (2026)Compression Is Intelligence: What Cross-Entropy Really Measures (2026)Automate License Renewals. Track Every Key. Own Your Software Spend (2026)Track Hours. Bill Clients. Get Paid. (Clone a Working Time Tracker in 2026)Claude Shannon: The History of Information Theory and the Man Who Invented the Bit (2026)Connect Your Apps. Automate Your Business. (Two-Way Workflows in 2026)Excel Job Log to Dispatch App (2026): Own the BoardMaintainX Alternative for Small Shops (2026)The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day (2026)Run Your Whole Business in One App with Taskade Genesis (June 2026)
AIAutomationProductivityProject ManagementRemote WorkStartupsKnowledge ManagementCollaborative WorkUpdates
Changelog
Google Sheets Trigger & Automation Stall Hotfix (Sep 3, 2026)Project Tap Hotfix (Sep 2, 2026)Run Two Builds at Once & Connect ClickUp (Sep 2, 2026)
App Kits Carry Agent Teams & CSV Attachments (Sep 2, 2026)Whole-File App Edits & Markdown Attachments (Sep 2, 2026)App Header Controls Hotfix (Aug 31, 2026)Automation Email Safety Hotfix (Aug 31, 2026)
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
© 2026 Taskade
PrivacyTermsSecurity
Made withTaskade AIforBuilders
BlogAIThe TSK-1 Methodology: How We…

The TSK-1 Methodology: How We Benchmark AI Models by Building Real Apps (2026)

The full TSK-1 method: frozen requests, one to three builds per model per test, a customer-style check of every saved field, four graded qualities, and what we never publish.

The TSK-1 methodology: the dated benchmark updates log on the Taskade hub, one entry per results day from July 30 to August 20, 2026
September 15, 202630 min readJohn XieAI·#tsk-1#ai-benchmarks#llm-evals
On this page (16)
What Is the TSK-1 Methodology?Why a Fixed Request Is the Whole BenchmarkOne Frozen Request, and Every Build KeptSame Settings or No ComparisonWorking Software Is the Entry TicketWe Use Every App Like a CustomerThe Four Qualities: Interface, Task, Memory, AdaptThe Intelligence IndexWhere Each Rule Came FromHonest Sample Sizes and Dated ClaimsWhat We Deliberately Do Not Publish, and WhyHow TSK-1 Differs From Other BenchmarksHow to Reproduce the Method on Your Own WorkWhat Comes NextFrequently Asked QuestionsRelated Reading

Most benchmark methodology pages are a paragraph. Ours is a post, because the rules are the benchmark. Change the request between models and you have two tests. Grade a build from the model's closing message and you have graded a press release. Average a design score against an app that never opened and you have manufactured a decimal that means nothing.

TSK-1 grades AI models on the app they build from one fixed request inside Taskade Genesis. This post is the full method: what is held constant, what is measured, how it is measured, where each rule came from, and what we refuse to publish. If you want the results, they are on the hub and in the side-by-side write-up. If you want to know whether to trust them, read on. 🔬

TL;DR: TSK-1 gives every AI model the same fingerprinted request inside the same builder. Every app is opened in light and dark, used like a customer, checked field by field, and asked to change. Four qualities, 1 to 3 builds per model per test, every claim dated, every miss published. See the live results →

What Is the TSK-1 Methodology?

TSK-1 is a hands-on test of one thing: whether an AI model can turn one request into a complete, working app. Every model receives the same request inside Taskade Genesis, builds it one to three times per test, and is graded on the app that comes out. Testers open the app, use it as a customer would, check what it saved, and ask it to change. Each model family earns a tier on four qualities, and the hub publishes the evidence with the day it was measured.

It is not a coding puzzle, and it is not a preference vote. The unit of work is a finished product with a real workspace behind it: pages, a database with typed fields, an automation, and a built-in AI assistant. The unit of grading is what a customer could do with that product on the day it was tested.

One fixed requestfrozen and fingerprinted Model A Model B Model C Taskade Genesisheld constant App A App B App C Opened, used, checked, changedby the same tester, the same way Four tiers per familyInterface · Task · Memory · Adapt
One fixed requestfrozen and fingerprinted Model A Model B Model C Taskade Genesisheld constant App A App B App C Opened, used, checked, changedby the same tester, the same way Four tiers per familyInterface · Task · Memory · Adapt

The rest of this post takes each box in that diagram and explains the rule behind it, in the order a build passes through them.

Why a Fixed Request Is the Whole Benchmark

A benchmark measures only what it holds constant, so TSK-1 freezes the text of every request and registers a fingerprint of it before any scored build. Two results go into the same comparison only if the request bytes match. A label such as "the tracker test" is not accepted as proof that two results are comparable, because the moment a request is reworded, even by one clause, the test has changed.

Two requests anchor the program. They were chosen to stress different things.

The tracker The client sign-up form
What it asks for A Dota 2 match tracker: log matches with hero, result, KDA, duration, and notes; a dashboard with win rate and streaks; a premium esports HUD look A real customer's 32-question intake form with a scoring formula, an automation, and a dashboard that says whether an applicant is eligible and why
The follow-up request Add a heroes page showing most-played heroes with win rates, linked from the main navigation In the customer's own typing, typos included: allow text input for the eligibility score, and if no score shows, recommend next steps from the answers
What it stresses App structure, dashboard math, a dark-first design, and an edit that needs a new page rather than a patch Length, a word-for-word mandate, long-answer judgment, a formula with no pass mark, and a two-part edit written informally
The hard line "Premium esports HUD": many models describe a dark design and ship a light one "Do not shorten my question or answers": models paraphrase, renumber, or strip punctuation
Public? Yes, the shape is described in full The shape is published; the text is not

The sign-up form text stays private for two reasons, and both are part of the method rather than an apology for it. Consent. It is a real customer's form. It is evidence of how they work, not copy to reprint, and the customer is never named. Contamination. A published prompt is a prompt that ends up in training data, and a benchmark whose hardest task can be memorized stops measuring anything within weeks. The history of AI benchmarks is a history of exactly that failure, from ImageNet to SWE-bench.

What we publish instead is the shape: 32 questions, an instruction not to shorten them, a scoring formula that deliberately omits a pass mark, an automation, a dashboard, and a field-by-field check of every saved answer. That is enough for you to run the same kind of test on your own form, which is the point.

no yes no yes no yes Is the request text byte-identicalto the registered fingerprint? New test. Results are notcomparable to earlier ones. Same Taskade Genesis version,same instructions, same conditions? Comparison ends. Report separately,never spliced into one table. Same day? Directional only. Dated,never a single ranked list. Comparable. Goes in the same table.
no yes no yes no yes Is the request text byte-identicalto the registered fingerprint? New test. Results are notcomparable to earlier ones. Same Taskade Genesis version,same instructions, same conditions? Comparison ends. Report separately,never spliced into one table. Same day? Directional only. Dated,never a single ranked list. Comparable. Goes in the same table.

One Frozen Request, and Every Build Kept

A test runs one to three builds per model, each followed by one follow-up request on the app that build produced. Every build that runs is recorded, and nobody edits or removes a result after the fact. A build that is cut short is reported as cut short. A quality that was not measured is recorded as not measured. Nothing is inferred, and no score is ever backfilled from an adjacent test.

That sounds obvious until you see what it rules out. It rules out re-running a model until it produces a flattering result and publishing only that run. It rules out quietly dropping the builds that did not finish. It rules out reading a design score off a build that never opened. And it rules out the most common benchmark sin of all: assuming a model would have passed a check nobody ran.

the model starts the build stops before publishing the model declares it finished fails in light, dark, or at phone width loads and runs persona fills in the form and submits one real follow-up request graded, dated, published with its misses reported as cut short, not graded reported as not scored Requested Building CutShort Published DidNotOpen Opened Used Changed WrittenUp
the model starts the build stops before publishing the model declares it finished fails in light, dark, or at phone width loads and runs persona fills in the form and submits one real follow-up request graded, dated, published with its misses reported as cut short, not graded reported as not scored Requested Building CutShort Published DidNotOpen Opened Used Changed WrittenUp

Every state in that diagram is a public outcome. "Not scored" is a real row on the hub, not an absence. When MiniMax M3 failed 47.2 percent of its build actions on July 31, 2026 and still described the app as built and live, the row reads "Not scored" and the write-up says why. A newer version can earn a fresh test. Nothing about the old result is edited when it does.

One more rule lives here. Results days are the public unit, not internal tests. Several tests can land on the same day, and the hub groups them by the day they were measured. That is why the benchmark updates log shows twelve dated entries between July 30 and August 20, 2026, each with a plain-language headline and the per-test notes beneath it.

Same Settings or No Comparison

Within one comparison, every model runs on the same Taskade Genesis version, the same instructions, and the same conditions. A change to any of those ends the comparison. Results from different weeks are directional and are never presented as a single ranked list, because a ranking across weeks is a ranking across two variables, the model and the builder.

This is the rule most benchmark readers never think about, and it is the one that most often invalidates a comparison. AI app builders ship changes weekly. So do the models. If a model scores better in week two than in week one, the honest statement is that the pair scored better, and you do not know which half moved.

TSK-1 handles this in three ways:

  1. Every test pins its conditions. The product version and the instruction set are recorded with the test. If either changes mid-test, the test is split, not continued.
  2. Comparisons are same-day by default. The August 1, 2026 nine-model test is the cleanest example: nine models, one request, one day, one builder. That is why it anchors so much of the public write-up.
  3. Cross-week claims are labeled directional. The hub's tiers are evidence-graded across tests, and the write-ups say "an August test" rather than pretending two dates were one.

There is a subtle corollary. When the product itself improves, so that every model's app gets richer, the cost of each build tends to rise with it. That is why cost comparisons are drawn only from comparable tests and expressed only as relative words. More on that below.

Working Software Is the Entry Ticket

Since August 5, 2026, an app that does not open receives no score on any quality, however good it looks in the model's description. The rule came from the August 1 nine-model test, where three of nine builds never ran, and two of them would have scored well on design from their screenshots alone.

"Opens" is checked, not assumed. The published app is loaded in light mode, in dark mode, and at phone width. A build that references code it never created, links to a machine nobody else can reach, or renders a blank page has not opened. On the day the rule was introduced it changed a ranking: two good-looking tracker builds turned out not to open, and DeepSeek V4 Pro won that test as the cheapest app that actually worked.

The gate exists because a benchmark that grades screenshots grades the wrong thing. From the customer's chair, a beautiful app that will not load is a failed build, and the method should say so before it says anything else.

We Use Every App Like a Customer

For each build that opens, a tester uses it the way a customer would, and every step produces evidence that a screenshot cannot. The sequence is fixed, so every model is used the same way.

Open in light mode, dark mode, phone width Fill in every field as the fixed test persona Submit Writes the record Compare every saved field to what was typed Check the derived values against the formula Automation fires, or does not Confirm it ran, and on the right record Ask one question about the record just saved Answers from the data, or does not Ask the app to make one change Check nothing was lost and no page broke Tester The published app Workspace behind it Automation Built-in assistant
Open in light mode, dark mode, phone width Fill in every field as the fixed test persona Submit Writes the record Compare every saved field to what was typed Check the derived values against the formula Automation fires, or does not Confirm it ran, and on the right record Ask one question about the record just saved Answers from the data, or does not Ask the app to make one change Check nothing was lost and no page broke Tester The published app Workspace behind it Automation Built-in assistant

Each step is there because a build once passed the step before it and failed this one:

  • A fixed test persona with a known outcome. The same fictional applicant fills in the sign-up form every time, with answers chosen so the eligibility result is known in advance. If the app says a different result, the scoring is wrong, and the tester knows without doing the arithmetic. On August 19, 2026 a hand-checked score of 11 matched the formula exactly. On August 7, one build let an applicant grade themselves a perfect score through rubrics that should have been staff-only.
  • Every saved field compared to what was typed. Not "the submission succeeded". Every field. On August 7, a tracker saved both dropdown values as "undefined", so a logged win displayed as a loss. The page looked fine.
  • Derived values checked. A KDA is computed from three numbers. A score is computed from a formula. Both are checked against the inputs, because a field that saves is not the same as a field that is used correctly.
  • The automation confirmed on the right record. An automation that fires on demo data, or posts a fixed string forever, has not run in any sense a customer would recognize.
  • One grounded question to the built-in assistant. Every Taskade Genesis app ships with an AI agent that can read the app's own project as knowledge. The tester asks it one question that can only be answered from the record just saved. On August 3, 2026 a DeepSeek V4 Flash build passed this whole path live: form filled in, data saved, assistant answering from it.
  • One real change request. The edit is graded for whether it landed, whether any data was lost, and whether a new page is a new page rather than a patch to the home screen.

The test is also what catches the class of failure no leaderboard can see: a model whose closing summary does not match its work. The MiniMax M3 result on July 31 is the canonical case, and it is why the tester reads the model's summary last, after using the app, never first.

The Four Qualities: Interface, Task, Memory, Adapt

Each model family earns a tier on four qualities, and the four together describe what a customer gets. The vocabulary on the hub is deliberately plain.

Quality What it grades What the tester actually checks Failure it catches
Interface How complete and polished the finished app feels Every page in light, dark, and at phone width; the model chose its own colors rather than the template default; the look matches the brief; a written rationale matches the shipped colors A dark block copied into the light theme; a "premium esports HUD" shipped as a light theme; every hover a solid slab because the accent color equals the primary
Task How closely the app understands and follows your brief Sample questions, and later all 32, checked word for word; named entities from the brief present; nothing added that nobody asked for; a sensible policy where the brief is silent A contacts database instead of a sign-up form; 7 of 32 questions shortened; an unrequested sign-in screen; a pass mark invented silently
Memory How reliably the app keeps what people add to your workspace Every saved field against what was typed; derived values against the formula; the automation ran on the right record; the assistant answers from the saved data Every stat card at zero over seeded data; dropdowns saved as "undefined"; four fictional sample applicants counted in live statistics
Adapt How cleanly the app changes when you ask One real follow-up request; no data lost; a new page is a new page; the model respects an app another model built Half the app rebuilt to add one card; an edit that asks a clarifying question instead of acting; a change that orphans a page

Each quality earns one of five tiers: Leading, Strong, Emerging, Limited, or Not scored. Tiers are deliberately coarse. Four steps and nothing finer, because averaging a design finding against an app that would not open produces false precision, and a decimal that separates two models by less than the noise between two builds of the same model is a lie with extra digits.

The Intelligence Index

The Intelligence Index rescales the four tiers to 0 to 100 for easier comparison. Leading is worth four points, Strong three, Emerging two, Limited one, Not scored zero. Four qualities times four points is sixteen, and sixteen maps to 100. It is a scale change, not a hundred separate checks, and families with the same result share a rank rather than being separated by an invented tiebreaker.

0 20 40 60 80 100 Claude GPT DeepSeek Kimi GLM Gemini Index, 0 to 100 Family TSK-1 Intelligence Index by model family (hub, August 2026)
0 20 40 60 80 100 Claude GPT DeepSeek Kimi GLM Gemini Index, 0 to 100 Family TSK-1 Intelligence Index by model family (hub, August 2026)
Family Interface Task Memory Adapt Index
Claude Strong Strong Strong Leading 81
GPT Strong Leading Strong Strong 81
DeepSeek Leading Strong Leading Emerging 81
Kimi Strong Strong Strong Strong 75
GLM Emerging Strong Strong Strong 69
Gemini Limited Emerging Emerging Emerging 44
MiniMax Not scored Not scored Not scored Not scored —

The GPT evidence card on the TSK-1 hub: four quality tiers, the Intelligence Index, and dated findings per version

Three families tie at 81, and the hub shows three number ones. That is not a hedge. It is what the evidence supports, and it is the honest resolution of a benchmark that grades 1 to 3 builds per model per test. Qwen and Grok have no tiers yet because they have no hands-on test yet. Their pages carry the public record with every source named until evidence replaces it.

Where Each Rule Came From

Every rule in this method has an origin story, and most of them are a specific build on a specific day. A benchmark that cannot say why a rule exists is a benchmark that will not know when to retire it.

Rule The build that produced it Date
A build must open before it can score Three of nine builds never ran; two would have scored well on design Aug 1, 2026, rule introduced Aug 5
Every build is read back against the brief, and the request is pinned so it cannot be lost After two stalls of about three minutes, a model lost the 32 questions and built a contacts database instead Aug 3, 2026
Extras count against a model A sign-in screen nobody asked for, added twice and never mentioned Jul 30 and Aug 1, 2026, codified Aug 20
Both themes are checked end to end Two of five tracker builds failed light and dark Aug 2, 2026
The design rationale is checked against the shipped colors Five of five models described a dark HUD and shipped a light theme Aug 5, 2026
A fixed persona with a known outcome, and every saved field compared The first full check from request to saved data; a self-grading rubric exposed Aug 7, 2026
Disclose-and-decide beats refuse, and both beat silent invention Three models chose three policies for the missing pass mark Aug 7, 2026
Sample data must be labeled Fictional applicants counted in live statistics; a sample record one letter from the real test submission Aug 7, 2026
The closing summary is read last, and compared to the work A model reported the app built and live while 47.2 percent of its build actions had failed Jul 31, 2026
A stuck build may be retried in a fresh conversation, and both results are recorded 52 minutes going in circles, then about five minutes from a fresh conversation Aug 7, 2026
A dedicated follow-up-edit test A strong first build is not enough; an app must also change cleanly Aug 20, 2026

The pattern across the table is the pattern of the whole benchmark. Nothing was designed in advance from theory. Each check was added the day a model got past the previous one in a way that would have hurt a customer. The failure taxonomy for AI-generated apps describes the same classes from the customer's side.

Honest Sample Sizes and Dated Claims

Each TSK-1 result rests on one to three builds per model per test, and the hub says so on the page. Positions are directional across tests rather than absolute rankings. Every evidence claim carries the day it was measured, and results that went badly are published beside the ones that went well.

Small samples are a limitation, and we would rather state it than hide it behind a decimal. But they are the right limitation for this kind of test, for three reasons:

  • The failure modes are large. A sign-in screen nobody asked for, a brief lost after a stall, an app that will not open, every stat at zero. None of these needs a hundred trials to observe. One is enough to know it can happen, and three is enough to know whether it repeats.
  • The tests are expensive in the right way. Every build is a real app on the real product, used by a person. That is the opposite of a synthetic task set that can be run ten thousand times and memorized. The trade is fewer runs for a result that means something.
  • The record is cumulative. GPT-5.6 Luna reproduced all 32 questions word for word not once but across repeated tests from August 3 to August 19, 2026. The confidence comes from the repetition across dated tests, not from a large n on one day.

Dating is the other half. A claim without a date is a claim about a model that no longer exists, because models and builders both change weekly. On the hub, the evidence cards are the load-bearing layer: each carries the day it was measured, and the page will not publish an evidence card that points at a test not in the record. The summary lines above them are ordinary prose, and we would rather say that plainly than imply the dating rule covers more than it does.

  WHAT A TSK-1 CLAIM LOOKS LIKE

[model] [what it did] [what that counts] [date]

GPT-5.6 Luna reproduced all 32 questions of the sign-up form, Aug 3-19, 2026
DeepSeek V4 Pro wired 8 automations, 60 fields into one build, Aug 6, 2026
Kimi K3 7.2% of build actions went wrong best of the test, Jul 31, 2026
MiniMax M3 not scored 47.2% of actions failed, Jul 31, 2026

Never: "4 of 4" without saying 4 of what
Never: "the fastest" without minutes and seconds and a date
Never: "cheapest" without "in that test"

What We Deliberately Do Not Publish, and Why

The method includes a list of things TSK-1 will not publish, and the list is as much a part of the benchmark as the checks are. Each item is a place where a number would look more precise than the evidence behind it.

Not published What appears instead Why
Per-model cost figures in absolute terms Relative words: cheapest, a fraction of the cost, roughly ten times, at half Early tests predate a billing change, so an average across them is arithmetic on incompatible units. Only comparable tests feed a cost comparison, and only as a ratio
The customer's sign-up form text The shape: 32 questions, a formula with no pass mark, the instruction not to shorten Consent, and contamination of the benchmark's hardest task
Provider and infrastructure details behind the models The public model name and the version tested They are not the thing being graded, and they change without notice
Internal test identifiers "An August test", or the results day A test id means nothing to a reader and invites false precision about comparability
A score for a quality that was not measured Not measured Inferring a score from an adjacent test is the exact failure the method exists to prevent
A grade for a build that was cut short Cut short, and the reason if known A half-built app is not a data point about the model's ceiling
A model's headline written around its failure The failure stays in the record, in the body, dated The point is to keep every honest miss, not to lead with one

Two of those deserve a longer note.

Cost. The temptation to publish a cost-per-build leaderboard is strong, because it is the number buyers ask for first. We publish relative words instead because they are the honest resolution of the data. When DeepSeek V4 Flash built the best-looking app of nine on August 1, 2026 for a fraction of what the others cost, "a fraction" is true across every way of counting. A precise multiple would be true only under one billing regime, on one day, for one build. If you want to think about the cost of running an AI-built app after launch, which is the larger number, reducing LLM costs is the better read.

Headlines. Every model family's write-up leads with what it did well and keeps every miss in the body. That is not marketing. It is the same rule a fair reviewer applies to a person: the failure is part of the record, the failure is not the name. Claude Sonnet 5 built the wrong app on August 3 after losing the brief, and the write-up says so in full. It also wrote the cleanest code of that test and was the only model on July 30 to check its own app's assistant. Both are true. The method requires both to be published.

How TSK-1 Differs From Other Benchmarks

The public coding benchmarks answer narrow questions well, and TSK-1 answers a different one. The comparison is not about which is better. It is about which instrument fits which question.

Unit of work Grader Defense against memorization Grades the follow-up edit
SWE-bench family A patch to an existing repository Hidden unit tests Weak; OpenAI stopped reporting the Verified set on February 23, 2026 after finding 59.4 percent of its hardest problems flawed No
Arena-style leaderboards A single answer or front end Human preference votes Not applicable; measures taste No
Terminal-Bench A command-line task Automated tests Version churn; 4.0 shipped August 28, 2026 and removed already-saturated tasks No
Vibe Code Bench A web app from a written spec A browser agent running workflows Synthetic specs No, self-debugging only
TSK-1 A complete app from one fixed request A tester using the app, checking every saved field The hardest request is unpublished and fingerprinted Yes, one real change

The side-by-side results post goes deeper on what each benchmark skips, and the history of AI benchmarks covers why every one of them eventually saturates. The one-line version: a patch benchmark is the right instrument for patch work, and an app benchmark is the right instrument for the question a buyer is actually asking.

How to Reproduce the Method on Your Own Work

You can apply every rule in this post to your own app request without any of our tooling, and the result will be more useful to you than any public leaderboard, because it runs on the only test set that matters. The rules are the protocol.

  1. Freeze the brief. Write it once. Save it somewhere you cannot accidentally edit. If you change a word, start a new test and say so.
  2. Hold the builder constant. Change only the model. In Taskade Genesis you can switch the model behind an app or an AI agent without rebuilding it.
  3. Run every model on the same day. If you cannot, label the comparison directional and keep the dates.
  4. Require the app to open. Light, dark, phone. If it does not, it gets no other grade. Write down why.
  5. Use it as a fixed persona. Decide the expected outcome before you fill in the form. Submit. Compare every saved field to what you typed. Check any derived value against its formula.
  6. Ask the built-in assistant one grounded question. If the app has one, ask something only the saved record can answer.
  7. Ask for one change. Check for lost data and broken pages.
  8. Read the model's summary last. Compare it to what you saw. A summary that overstates the work is a finding.
  9. Date everything. State your n. Record "not measured" as a value.
  TSK-1 CHECKLIST, ONE BUILD

[ ] brief frozen, byte-identical to the last run date: ______
[ ] builder version unchanged since the last run of this test
[ ] opens in light [ ] opens in dark [ ] opens at phone width
[ ] persona filled in, expected outcome written down first
[ ] every saved field == what was typed fields checked: __ / __
[ ] derived values match the formula
[ ] automation ran, on the right record
[ ] assistant answered from the saved record
[ ] one change requested [ ] nothing lost [ ] no page broke
[ ] model summary read LAST and compared to the above
[ ] misses written down next to the wins

Switching the model behind an agent inside Taskade Genesis, which is how one frozen brief runs through several models without rebuilding anything

Start from a shape close to ours if you like: a form template or a gaming tracker template, then describe your own version. The Taskade Genesis quickstart walks through the first build, and Workspace DNA explains what the app is standing on: projects that remember, agents that think, automations that execute.

What Comes Next

The method keeps changing in one direction: toward the customer's chair. Three additions are underway as of September 2026, and none of them produces a public result until it has run on the shipping product under the comparability rules above.

  • New request shapes drawn from how customers actually build. A sales pipeline with AI lead scoring and a daily digest, a weekly pass-or-fail field audit with a findings dashboard, and an approval queue with an execution log and an audit archive. Each adds a shape the two anchor requests do not cover, and each will get its own frozen text and fingerprint.
  • The built-in assistant as a graded step. Asking the app's own AI agent one grounded question about the record just saved is now part of the standard check, so Memory measures whether the data is usable, not only whether it landed.
  • Hands-on tests for Qwen and Grok, and a fresh test for any family that ships a notable new version. Old results are never edited when new ones arrive. Both stay on the page, in order, with their dates.

The hub is the source of truth for all of it. Inside Taskade Genesis, TSK-1 Auto handles the default, so choosing a model is optional. When the method changes, the benchmark updates log says what changed and on which day.

Frequently Asked Questions

What is the TSK-1 methodology?

TSK-1 is a hands-on test of whether an AI model can turn one request into a complete, working app inside Taskade Genesis. Every model receives the same request word for word, and is graded on the app that comes out, with one to three builds per model per test and every recorded build kept in the write-up. Testers open each app in light and dark and at phone width, fill in its form as a fixed persona, compare every saved field to what was typed, confirm the automation ran, ask the built-in assistant a grounded question, and then ask the app to change. Each family earns a tier on Interface, Task, Memory, and Adapt.

Why does TSK-1 use a fixed request that never changes?

A benchmark only measures what it holds constant. TSK-1 freezes each request and registers a fingerprint before any scored build, so two results only go into the same comparison if the request bytes match. A label is not accepted as proof that two results are comparable. Two requests anchor the program: a match tracker that must look like a premium esports HUD, and a real customer's 32-question client sign-up form with the instruction not to shorten any question.

Why is the 32-question sign-up form prompt not published?

Consent and contamination. It is a real customer's form, so reprinting it would turn evidence of how they work into public copy. And a published prompt leaks into training data, so a benchmark whose hardest task can be memorized stops measuring anything within weeks. TSK-1 publishes the shape instead: 32 questions, a formula with no pass mark, an automation, a dashboard, and the field-by-field check.

What are the four qualities TSK-1 grades?

Interface is how complete and polished the app feels, checked in light, dark, and at phone width. Task is how closely the app follows the brief, including word-for-word fidelity and whether anything was added that nobody asked for. Memory is whether the app keeps what people add, checked by comparing every saved field to what was typed. Adapt is how cleanly the app changes on request. Each earns a tier: Leading, Strong, Emerging, Limited, or Not scored.

What is the TSK-1 Intelligence Index?

A rescaling of the four tiers to 0 to 100. Leading is four points, Strong three, Emerging two, Limited one, Not scored zero, so sixteen points map to 100. It is a scale change, not a hundred checks, and families with the same result share a rank. As of August 2026, Claude, GPT, and DeepSeek share the top position at 81, then Kimi at 75, GLM at 69, and Gemini at 44.

How many builds does each TSK-1 result rest on?

One to three builds per model per test, stated on the hub. Positions are directional across tests, every claim carries the day it was measured, and anything not measured is recorded as not measured. That is enough to observe an unrequested sign-in screen, a brief lost after a stall, or a model reproducing all 32 questions across repeated tests.

What does TSK-1 refuse to publish?

Absolute per-model cost figures, which appear only as relative words from comparable tests. The customer's form text. Provider and infrastructure details. Internal test identifiers. Any score for a quality that was not measured. Any grade for a build that was cut short. Every honest miss stays in the record beside the wins.

Why does a TSK-1 build have to open before it can score?

Because on August 1, 2026 three of nine builds never opened, and two of them looked excellent in the model's own description. Since August 5, 2026 a running app is the entry requirement. The rule changed a ranking the day it was introduced.

Why does TSK-1 grade the follow-up edit?

Because the edit is what you will do to an app every week for as long as you own it, and no other public benchmark grades it. Every test ends with one real change request, checked for lost data and broken pages. A dedicated follow-up-edit test was added on August 20, 2026, together with stricter instruction checks so that actions nobody asked for count against a model.

Can I use the TSK-1 methodology on my own app request?

Yes. Freeze the brief, hold the builder constant, run every model on the same day, require the app to open, use it as a fixed persona and compare every saved field, ask for one change, read the summary last, and date everything. Taskade Genesis lets you switch the model behind an app or AI agent without rebuilding it, so one brief can run through several models in one workspace. Paid plans start at $10 per month billed annually, and the free plan includes three Taskade Genesis apps.

Related Reading

  • Best AI Model for Building Apps in 2026, the side-by-side results this method produced.
  • The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day, the original results write-up.
  • TSK-1 Benchmark in the wiki, the short reference version of this page.
  • Introducing Taskade TSK-1, what the kernel is.
  • What Are AI Agent Evals? and LLM-as-a-judge, the wider evaluation literature this method sits inside.
  • RL Environments Explained, the training-side loop this method borrows its shape from.
  • History of AI Benchmarks, why every benchmark saturates and what survives.

A method is a promise about what you will not do. We will not reword the request. We will not grade from the summary. We will not average a design score against an app that never opened. We will not publish a number the evidence cannot carry. Hold us to it, and run it on your own work. ▲ ■ ●

0%

On this page

What Is the TSK-1 Methodology?Why a Fixed Request Is the Whole BenchmarkOne Frozen Request, and Every Build KeptSame Settings or No ComparisonWorking Software Is the Entry TicketWe Use Every App Like a CustomerThe Four Qualities: Interface, Task, Memory, AdaptThe Intelligence IndexWhere Each Rule Came FromHonest Sample Sizes and Dated ClaimsWhat We Deliberately Do Not Publish, and WhyHow TSK-1 Differs From Other BenchmarksHow to Reproduce the Method on Your Own WorkWhat Comes NextFrequently Asked QuestionsRelated Reading

Related Articles

A tracker app built end to end in Taskade Genesis, the same app shape the TSK-1 benchmark asks every AI model to produce from one fixed prompt
August 22, 2026AI

The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day (2026)

On Aug 1, 2026 nine AI models built the same app from one fixed prompt in Taskade Genesis. Three builds never opened, an...

Close-up of a human eye, representing the ImageNet moment when machine vision surpassed hand-designed computer vision methods
August 13, 2026AI

The ImageNet Moment, Explained: How Computer Vision Broke Open (2026)

In 2012 AlexNet cut ImageNet error from 26 to 15.3 percent. Weeks later a Stanford student wrote that vision was hopeles...

History of AI benchmarks chart showing test scores of AI systems on MNIST, ImageNet, GLUE, SuperGLUE, MMLU and HumanEval relative to human performance
August 10, 2026AI

The History of AI Benchmarks: Why Every Model Claims to Be the Best (2026)

AI benchmarks go from impossible to solved in about two years. A verified history from the Turing test to ARC-AGI-2, plu...

A live app built from a prompt, running with its own database, agents, and automations
September 1, 2026AI

Chat-Native App Builders in 2026: What You Actually Own When the Chat Ends

Claude, ChatGPT, and Gemini can all build an app inside the chat. The question nobody answers is what survives when you ...

Previewing and customizing a branded AI agent in Taskade before publishing it as a public page, a custom domain, or a website widget
August 31, 2026AI

Generate the Art. Preview the Agent. Put It on Your Domain (2026)

Ship an AI agent that looks like your company: generate its art in the workspace, preview it the way a visitor sees it, ...

An AI agent with its own tools running inside a Taskade automation
August 30, 2026AI

Agentic Automation Explained: Agent vs AI Step (2026)

Agentic automation runs a named agent with memory and tools inside a workflow. Here is the five-point test that separate...

View All Articles