Most benchmark methodology pages are a paragraph. Ours is a post, because the rules are the benchmark. Change the request between models and you have two tests. Grade a build from the model's closing message and you have graded a press release. Average a design score against an app that never opened and you have manufactured a decimal that means nothing.
TSK-1 grades AI models on the app they build from one fixed request inside Taskade Genesis. This post is the full method: what is held constant, what is measured, how it is measured, where each rule came from, and what we refuse to publish. If you want the results, they are on the hub and in the side-by-side write-up. If you want to know whether to trust them, read on. 🔬
TL;DR: TSK-1 gives every AI model the same fingerprinted request inside the same builder. Every app is opened in light and dark, used like a customer, checked field by field, and asked to change. Four qualities, 1 to 3 builds per model per test, every claim dated, every miss published. See the live results →
What Is the TSK-1 Methodology?
TSK-1 is a hands-on test of one thing: whether an AI model can turn one request into a complete, working app. Every model receives the same request inside Taskade Genesis, builds it one to three times per test, and is graded on the app that comes out. Testers open the app, use it as a customer would, check what it saved, and ask it to change. Each model family earns a tier on four qualities, and the hub publishes the evidence with the day it was measured.
It is not a coding puzzle, and it is not a preference vote. The unit of work is a finished product with a real workspace behind it: pages, a database with typed fields, an automation, and a built-in AI assistant. The unit of grading is what a customer could do with that product on the day it was tested.
The rest of this post takes each box in that diagram and explains the rule behind it, in the order a build passes through them.
Why a Fixed Request Is the Whole Benchmark
A benchmark measures only what it holds constant, so TSK-1 freezes the text of every request and registers a fingerprint of it before any scored build. Two results go into the same comparison only if the request bytes match. A label such as "the tracker test" is not accepted as proof that two results are comparable, because the moment a request is reworded, even by one clause, the test has changed.
Two requests anchor the program. They were chosen to stress different things.
| The tracker | The client sign-up form | |
|---|---|---|
| What it asks for | A Dota 2 match tracker: log matches with hero, result, KDA, duration, and notes; a dashboard with win rate and streaks; a premium esports HUD look | A real customer's 32-question intake form with a scoring formula, an automation, and a dashboard that says whether an applicant is eligible and why |
| The follow-up request | Add a heroes page showing most-played heroes with win rates, linked from the main navigation | In the customer's own typing, typos included: allow text input for the eligibility score, and if no score shows, recommend next steps from the answers |
| What it stresses | App structure, dashboard math, a dark-first design, and an edit that needs a new page rather than a patch | Length, a word-for-word mandate, long-answer judgment, a formula with no pass mark, and a two-part edit written informally |
| The hard line | "Premium esports HUD": many models describe a dark design and ship a light one | "Do not shorten my question or answers": models paraphrase, renumber, or strip punctuation |
| Public? | Yes, the shape is described in full | The shape is published; the text is not |
The sign-up form text stays private for two reasons, and both are part of the method rather than an apology for it. Consent. It is a real customer's form. It is evidence of how they work, not copy to reprint, and the customer is never named. Contamination. A published prompt is a prompt that ends up in training data, and a benchmark whose hardest task can be memorized stops measuring anything within weeks. The history of AI benchmarks is a history of exactly that failure, from ImageNet to SWE-bench.
What we publish instead is the shape: 32 questions, an instruction not to shorten them, a scoring formula that deliberately omits a pass mark, an automation, a dashboard, and a field-by-field check of every saved answer. That is enough for you to run the same kind of test on your own form, which is the point.
One Frozen Request, and Every Build Kept
A test runs one to three builds per model, each followed by one follow-up request on the app that build produced. Every build that runs is recorded, and nobody edits or removes a result after the fact. A build that is cut short is reported as cut short. A quality that was not measured is recorded as not measured. Nothing is inferred, and no score is ever backfilled from an adjacent test.
That sounds obvious until you see what it rules out. It rules out re-running a model until it produces a flattering result and publishing only that run. It rules out quietly dropping the builds that did not finish. It rules out reading a design score off a build that never opened. And it rules out the most common benchmark sin of all: assuming a model would have passed a check nobody ran.
Every state in that diagram is a public outcome. "Not scored" is a real row on the hub, not an absence. When MiniMax M3 failed 47.2 percent of its build actions on July 31, 2026 and still described the app as built and live, the row reads "Not scored" and the write-up says why. A newer version can earn a fresh test. Nothing about the old result is edited when it does.
One more rule lives here. Results days are the public unit, not internal tests. Several tests can land on the same day, and the hub groups them by the day they were measured. That is why the benchmark updates log shows twelve dated entries between July 30 and August 20, 2026, each with a plain-language headline and the per-test notes beneath it.
Same Settings or No Comparison
Within one comparison, every model runs on the same Taskade Genesis version, the same instructions, and the same conditions. A change to any of those ends the comparison. Results from different weeks are directional and are never presented as a single ranked list, because a ranking across weeks is a ranking across two variables, the model and the builder.
This is the rule most benchmark readers never think about, and it is the one that most often invalidates a comparison. AI app builders ship changes weekly. So do the models. If a model scores better in week two than in week one, the honest statement is that the pair scored better, and you do not know which half moved.
TSK-1 handles this in three ways:
- Every test pins its conditions. The product version and the instruction set are recorded with the test. If either changes mid-test, the test is split, not continued.
- Comparisons are same-day by default. The August 1, 2026 nine-model test is the cleanest example: nine models, one request, one day, one builder. That is why it anchors so much of the public write-up.
- Cross-week claims are labeled directional. The hub's tiers are evidence-graded across tests, and the write-ups say "an August test" rather than pretending two dates were one.
There is a subtle corollary. When the product itself improves, so that every model's app gets richer, the cost of each build tends to rise with it. That is why cost comparisons are drawn only from comparable tests and expressed only as relative words. More on that below.
Working Software Is the Entry Ticket
Since August 5, 2026, an app that does not open receives no score on any quality, however good it looks in the model's description. The rule came from the August 1 nine-model test, where three of nine builds never ran, and two of them would have scored well on design from their screenshots alone.
"Opens" is checked, not assumed. The published app is loaded in light mode, in dark mode, and at phone width. A build that references code it never created, links to a machine nobody else can reach, or renders a blank page has not opened. On the day the rule was introduced it changed a ranking: two good-looking tracker builds turned out not to open, and DeepSeek V4 Pro won that test as the cheapest app that actually worked.
The gate exists because a benchmark that grades screenshots grades the wrong thing. From the customer's chair, a beautiful app that will not load is a failed build, and the method should say so before it says anything else.
We Use Every App Like a Customer
For each build that opens, a tester uses it the way a customer would, and every step produces evidence that a screenshot cannot. The sequence is fixed, so every model is used the same way.
Each step is there because a build once passed the step before it and failed this one:
- A fixed test persona with a known outcome. The same fictional applicant fills in the sign-up form every time, with answers chosen so the eligibility result is known in advance. If the app says a different result, the scoring is wrong, and the tester knows without doing the arithmetic. On August 19, 2026 a hand-checked score of 11 matched the formula exactly. On August 7, one build let an applicant grade themselves a perfect score through rubrics that should have been staff-only.
- Every saved field compared to what was typed. Not "the submission succeeded". Every field. On August 7, a tracker saved both dropdown values as "undefined", so a logged win displayed as a loss. The page looked fine.
- Derived values checked. A KDA is computed from three numbers. A score is computed from a formula. Both are checked against the inputs, because a field that saves is not the same as a field that is used correctly.
- The automation confirmed on the right record. An automation that fires on demo data, or posts a fixed string forever, has not run in any sense a customer would recognize.
- One grounded question to the built-in assistant. Every Taskade Genesis app ships with an AI agent that can read the app's own project as knowledge. The tester asks it one question that can only be answered from the record just saved. On August 3, 2026 a DeepSeek V4 Flash build passed this whole path live: form filled in, data saved, assistant answering from it.
- One real change request. The edit is graded for whether it landed, whether any data was lost, and whether a new page is a new page rather than a patch to the home screen.
The test is also what catches the class of failure no leaderboard can see: a model whose closing summary does not match its work. The MiniMax M3 result on July 31 is the canonical case, and it is why the tester reads the model's summary last, after using the app, never first.
The Four Qualities: Interface, Task, Memory, Adapt
Each model family earns a tier on four qualities, and the four together describe what a customer gets. The vocabulary on the hub is deliberately plain.
| Quality | What it grades | What the tester actually checks | Failure it catches |
|---|---|---|---|
| Interface | How complete and polished the finished app feels | Every page in light, dark, and at phone width; the model chose its own colors rather than the template default; the look matches the brief; a written rationale matches the shipped colors | A dark block copied into the light theme; a "premium esports HUD" shipped as a light theme; every hover a solid slab because the accent color equals the primary |
| Task | How closely the app understands and follows your brief | Sample questions, and later all 32, checked word for word; named entities from the brief present; nothing added that nobody asked for; a sensible policy where the brief is silent | A contacts database instead of a sign-up form; 7 of 32 questions shortened; an unrequested sign-in screen; a pass mark invented silently |
| Memory | How reliably the app keeps what people add to your workspace | Every saved field against what was typed; derived values against the formula; the automation ran on the right record; the assistant answers from the saved data | Every stat card at zero over seeded data; dropdowns saved as "undefined"; four fictional sample applicants counted in live statistics |
| Adapt | How cleanly the app changes when you ask | One real follow-up request; no data lost; a new page is a new page; the model respects an app another model built | Half the app rebuilt to add one card; an edit that asks a clarifying question instead of acting; a change that orphans a page |
Each quality earns one of five tiers: Leading, Strong, Emerging, Limited, or Not scored. Tiers are deliberately coarse. Four steps and nothing finer, because averaging a design finding against an app that would not open produces false precision, and a decimal that separates two models by less than the noise between two builds of the same model is a lie with extra digits.
The Intelligence Index
The Intelligence Index rescales the four tiers to 0 to 100 for easier comparison. Leading is worth four points, Strong three, Emerging two, Limited one, Not scored zero. Four qualities times four points is sixteen, and sixteen maps to 100. It is a scale change, not a hundred separate checks, and families with the same result share a rank rather than being separated by an invented tiebreaker.
| Family | Interface | Task | Memory | Adapt | Index |
|---|---|---|---|---|---|
| Claude | Strong | Strong | Strong | Leading | 81 |
| GPT | Strong | Leading | Strong | Strong | 81 |
| DeepSeek | Leading | Strong | Leading | Emerging | 81 |
| Kimi | Strong | Strong | Strong | Strong | 75 |
| GLM | Emerging | Strong | Strong | Strong | 69 |
| Gemini | Limited | Emerging | Emerging | Emerging | 44 |
| MiniMax | Not scored | Not scored | Not scored | Not scored | — |

Three families tie at 81, and the hub shows three number ones. That is not a hedge. It is what the evidence supports, and it is the honest resolution of a benchmark that grades 1 to 3 builds per model per test. Qwen and Grok have no tiers yet because they have no hands-on test yet. Their pages carry the public record with every source named until evidence replaces it.
Where Each Rule Came From
Every rule in this method has an origin story, and most of them are a specific build on a specific day. A benchmark that cannot say why a rule exists is a benchmark that will not know when to retire it.
| Rule | The build that produced it | Date |
|---|---|---|
| A build must open before it can score | Three of nine builds never ran; two would have scored well on design | Aug 1, 2026, rule introduced Aug 5 |
| Every build is read back against the brief, and the request is pinned so it cannot be lost | After two stalls of about three minutes, a model lost the 32 questions and built a contacts database instead | Aug 3, 2026 |
| Extras count against a model | A sign-in screen nobody asked for, added twice and never mentioned | Jul 30 and Aug 1, 2026, codified Aug 20 |
| Both themes are checked end to end | Two of five tracker builds failed light and dark | Aug 2, 2026 |
| The design rationale is checked against the shipped colors | Five of five models described a dark HUD and shipped a light theme | Aug 5, 2026 |
| A fixed persona with a known outcome, and every saved field compared | The first full check from request to saved data; a self-grading rubric exposed | Aug 7, 2026 |
| Disclose-and-decide beats refuse, and both beat silent invention | Three models chose three policies for the missing pass mark | Aug 7, 2026 |
| Sample data must be labeled | Fictional applicants counted in live statistics; a sample record one letter from the real test submission | Aug 7, 2026 |
| The closing summary is read last, and compared to the work | A model reported the app built and live while 47.2 percent of its build actions had failed | Jul 31, 2026 |
| A stuck build may be retried in a fresh conversation, and both results are recorded | 52 minutes going in circles, then about five minutes from a fresh conversation | Aug 7, 2026 |
| A dedicated follow-up-edit test | A strong first build is not enough; an app must also change cleanly | Aug 20, 2026 |
The pattern across the table is the pattern of the whole benchmark. Nothing was designed in advance from theory. Each check was added the day a model got past the previous one in a way that would have hurt a customer. The failure taxonomy for AI-generated apps describes the same classes from the customer's side.
Honest Sample Sizes and Dated Claims
Each TSK-1 result rests on one to three builds per model per test, and the hub says so on the page. Positions are directional across tests rather than absolute rankings. Every evidence claim carries the day it was measured, and results that went badly are published beside the ones that went well.
Small samples are a limitation, and we would rather state it than hide it behind a decimal. But they are the right limitation for this kind of test, for three reasons:
- The failure modes are large. A sign-in screen nobody asked for, a brief lost after a stall, an app that will not open, every stat at zero. None of these needs a hundred trials to observe. One is enough to know it can happen, and three is enough to know whether it repeats.
- The tests are expensive in the right way. Every build is a real app on the real product, used by a person. That is the opposite of a synthetic task set that can be run ten thousand times and memorized. The trade is fewer runs for a result that means something.
- The record is cumulative. GPT-5.6 Luna reproduced all 32 questions word for word not once but across repeated tests from August 3 to August 19, 2026. The confidence comes from the repetition across dated tests, not from a large n on one day.
Dating is the other half. A claim without a date is a claim about a model that no longer exists, because models and builders both change weekly. On the hub, the evidence cards are the load-bearing layer: each carries the day it was measured, and the page will not publish an evidence card that points at a test not in the record. The summary lines above them are ordinary prose, and we would rather say that plainly than imply the dating rule covers more than it does.
WHAT A TSK-1 CLAIM LOOKS LIKE [model] [what it did] [what that counts] [date]
GPT-5.6 Luna reproduced all 32 questions of the sign-up form, Aug 3-19, 2026
DeepSeek V4 Pro wired 8 automations, 60 fields into one build, Aug 6, 2026
Kimi K3 7.2% of build actions went wrong best of the test, Jul 31, 2026
MiniMax M3 not scored 47.2% of actions failed, Jul 31, 2026
Never: "4 of 4" without saying 4 of what
Never: "the fastest" without minutes and seconds and a date
Never: "cheapest" without "in that test"
What We Deliberately Do Not Publish, and Why
The method includes a list of things TSK-1 will not publish, and the list is as much a part of the benchmark as the checks are. Each item is a place where a number would look more precise than the evidence behind it.
| Not published | What appears instead | Why |
|---|---|---|
| Per-model cost figures in absolute terms | Relative words: cheapest, a fraction of the cost, roughly ten times, at half | Early tests predate a billing change, so an average across them is arithmetic on incompatible units. Only comparable tests feed a cost comparison, and only as a ratio |
| The customer's sign-up form text | The shape: 32 questions, a formula with no pass mark, the instruction not to shorten | Consent, and contamination of the benchmark's hardest task |
| Provider and infrastructure details behind the models | The public model name and the version tested | They are not the thing being graded, and they change without notice |
| Internal test identifiers | "An August test", or the results day | A test id means nothing to a reader and invites false precision about comparability |
| A score for a quality that was not measured | Not measured | Inferring a score from an adjacent test is the exact failure the method exists to prevent |
| A grade for a build that was cut short | Cut short, and the reason if known | A half-built app is not a data point about the model's ceiling |
| A model's headline written around its failure | The failure stays in the record, in the body, dated | The point is to keep every honest miss, not to lead with one |
Two of those deserve a longer note.
Cost. The temptation to publish a cost-per-build leaderboard is strong, because it is the number buyers ask for first. We publish relative words instead because they are the honest resolution of the data. When DeepSeek V4 Flash built the best-looking app of nine on August 1, 2026 for a fraction of what the others cost, "a fraction" is true across every way of counting. A precise multiple would be true only under one billing regime, on one day, for one build. If you want to think about the cost of running an AI-built app after launch, which is the larger number, reducing LLM costs is the better read.
Headlines. Every model family's write-up leads with what it did well and keeps every miss in the body. That is not marketing. It is the same rule a fair reviewer applies to a person: the failure is part of the record, the failure is not the name. Claude Sonnet 5 built the wrong app on August 3 after losing the brief, and the write-up says so in full. It also wrote the cleanest code of that test and was the only model on July 30 to check its own app's assistant. Both are true. The method requires both to be published.
How TSK-1 Differs From Other Benchmarks
The public coding benchmarks answer narrow questions well, and TSK-1 answers a different one. The comparison is not about which is better. It is about which instrument fits which question.
| Unit of work | Grader | Defense against memorization | Grades the follow-up edit | |
|---|---|---|---|---|
| SWE-bench family | A patch to an existing repository | Hidden unit tests | Weak; OpenAI stopped reporting the Verified set on February 23, 2026 after finding 59.4 percent of its hardest problems flawed | No |
| Arena-style leaderboards | A single answer or front end | Human preference votes | Not applicable; measures taste | No |
| Terminal-Bench | A command-line task | Automated tests | Version churn; 4.0 shipped August 28, 2026 and removed already-saturated tasks | No |
| Vibe Code Bench | A web app from a written spec | A browser agent running workflows | Synthetic specs | No, self-debugging only |
| TSK-1 | A complete app from one fixed request | A tester using the app, checking every saved field | The hardest request is unpublished and fingerprinted | Yes, one real change |
The side-by-side results post goes deeper on what each benchmark skips, and the history of AI benchmarks covers why every one of them eventually saturates. The one-line version: a patch benchmark is the right instrument for patch work, and an app benchmark is the right instrument for the question a buyer is actually asking.
How to Reproduce the Method on Your Own Work
You can apply every rule in this post to your own app request without any of our tooling, and the result will be more useful to you than any public leaderboard, because it runs on the only test set that matters. The rules are the protocol.
- Freeze the brief. Write it once. Save it somewhere you cannot accidentally edit. If you change a word, start a new test and say so.
- Hold the builder constant. Change only the model. In Taskade Genesis you can switch the model behind an app or an AI agent without rebuilding it.
- Run every model on the same day. If you cannot, label the comparison directional and keep the dates.
- Require the app to open. Light, dark, phone. If it does not, it gets no other grade. Write down why.
- Use it as a fixed persona. Decide the expected outcome before you fill in the form. Submit. Compare every saved field to what you typed. Check any derived value against its formula.
- Ask the built-in assistant one grounded question. If the app has one, ask something only the saved record can answer.
- Ask for one change. Check for lost data and broken pages.
- Read the model's summary last. Compare it to what you saw. A summary that overstates the work is a finding.
- Date everything. State your n. Record "not measured" as a value.
TSK-1 CHECKLIST, ONE BUILD [ ] brief frozen, byte-identical to the last run date: ______
[ ] builder version unchanged since the last run of this test
[ ] opens in light [ ] opens in dark [ ] opens at phone width
[ ] persona filled in, expected outcome written down first
[ ] every saved field == what was typed fields checked: __ / __
[ ] derived values match the formula
[ ] automation ran, on the right record
[ ] assistant answered from the saved record
[ ] one change requested [ ] nothing lost [ ] no page broke
[ ] model summary read LAST and compared to the above
[ ] misses written down next to the wins

Start from a shape close to ours if you like: a form template or a gaming tracker template, then describe your own version. The Taskade Genesis quickstart walks through the first build, and Workspace DNA explains what the app is standing on: projects that remember, agents that think, automations that execute.
What Comes Next
The method keeps changing in one direction: toward the customer's chair. Three additions are underway as of September 2026, and none of them produces a public result until it has run on the shipping product under the comparability rules above.
- New request shapes drawn from how customers actually build. A sales pipeline with AI lead scoring and a daily digest, a weekly pass-or-fail field audit with a findings dashboard, and an approval queue with an execution log and an audit archive. Each adds a shape the two anchor requests do not cover, and each will get its own frozen text and fingerprint.
- The built-in assistant as a graded step. Asking the app's own AI agent one grounded question about the record just saved is now part of the standard check, so Memory measures whether the data is usable, not only whether it landed.
- Hands-on tests for Qwen and Grok, and a fresh test for any family that ships a notable new version. Old results are never edited when new ones arrive. Both stay on the page, in order, with their dates.
The hub is the source of truth for all of it. Inside Taskade Genesis, TSK-1 Auto handles the default, so choosing a model is optional. When the method changes, the benchmark updates log says what changed and on which day.
Frequently Asked Questions
What is the TSK-1 methodology?
TSK-1 is a hands-on test of whether an AI model can turn one request into a complete, working app inside Taskade Genesis. Every model receives the same request word for word, and is graded on the app that comes out, with one to three builds per model per test and every recorded build kept in the write-up. Testers open each app in light and dark and at phone width, fill in its form as a fixed persona, compare every saved field to what was typed, confirm the automation ran, ask the built-in assistant a grounded question, and then ask the app to change. Each family earns a tier on Interface, Task, Memory, and Adapt.
Why does TSK-1 use a fixed request that never changes?
A benchmark only measures what it holds constant. TSK-1 freezes each request and registers a fingerprint before any scored build, so two results only go into the same comparison if the request bytes match. A label is not accepted as proof that two results are comparable. Two requests anchor the program: a match tracker that must look like a premium esports HUD, and a real customer's 32-question client sign-up form with the instruction not to shorten any question.
Why is the 32-question sign-up form prompt not published?
Consent and contamination. It is a real customer's form, so reprinting it would turn evidence of how they work into public copy. And a published prompt leaks into training data, so a benchmark whose hardest task can be memorized stops measuring anything within weeks. TSK-1 publishes the shape instead: 32 questions, a formula with no pass mark, an automation, a dashboard, and the field-by-field check.
What are the four qualities TSK-1 grades?
Interface is how complete and polished the app feels, checked in light, dark, and at phone width. Task is how closely the app follows the brief, including word-for-word fidelity and whether anything was added that nobody asked for. Memory is whether the app keeps what people add, checked by comparing every saved field to what was typed. Adapt is how cleanly the app changes on request. Each earns a tier: Leading, Strong, Emerging, Limited, or Not scored.
What is the TSK-1 Intelligence Index?
A rescaling of the four tiers to 0 to 100. Leading is four points, Strong three, Emerging two, Limited one, Not scored zero, so sixteen points map to 100. It is a scale change, not a hundred checks, and families with the same result share a rank. As of August 2026, Claude, GPT, and DeepSeek share the top position at 81, then Kimi at 75, GLM at 69, and Gemini at 44.
How many builds does each TSK-1 result rest on?
One to three builds per model per test, stated on the hub. Positions are directional across tests, every claim carries the day it was measured, and anything not measured is recorded as not measured. That is enough to observe an unrequested sign-in screen, a brief lost after a stall, or a model reproducing all 32 questions across repeated tests.
What does TSK-1 refuse to publish?
Absolute per-model cost figures, which appear only as relative words from comparable tests. The customer's form text. Provider and infrastructure details. Internal test identifiers. Any score for a quality that was not measured. Any grade for a build that was cut short. Every honest miss stays in the record beside the wins.
Why does a TSK-1 build have to open before it can score?
Because on August 1, 2026 three of nine builds never opened, and two of them looked excellent in the model's own description. Since August 5, 2026 a running app is the entry requirement. The rule changed a ranking the day it was introduced.
Why does TSK-1 grade the follow-up edit?
Because the edit is what you will do to an app every week for as long as you own it, and no other public benchmark grades it. Every test ends with one real change request, checked for lost data and broken pages. A dedicated follow-up-edit test was added on August 20, 2026, together with stricter instruction checks so that actions nobody asked for count against a model.
Can I use the TSK-1 methodology on my own app request?
Yes. Freeze the brief, hold the builder constant, run every model on the same day, require the app to open, use it as a fixed persona and compare every saved field, ask for one change, read the summary last, and date everything. Taskade Genesis lets you switch the model behind an app or AI agent without rebuilding it, so one brief can run through several models in one workspace. Paid plans start at $10 per month billed annually, and the free plan includes three Taskade Genesis apps.
Related Reading
- Best AI Model for Building Apps in 2026, the side-by-side results this method produced.
- The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day, the original results write-up.
- TSK-1 Benchmark in the wiki, the short reference version of this page.
- Introducing Taskade TSK-1, what the kernel is.
- What Are AI Agent Evals? and LLM-as-a-judge, the wider evaluation literature this method sits inside.
- RL Environments Explained, the training-side loop this method borrows its shape from.
- History of AI Benchmarks, why every benchmark saturates and what survives.
A method is a promise about what you will not do. We will not reword the request. We will not grade from the summary. We will not average a design score against an app that never opened. We will not publish a number the evidence cannot carry. Hold us to it, and run it on your own work. ▲ ■ ●





