Skip to main content
Introducing TSK-1Introducing TSK-1·Taskade's intelligence layer.
taskade
PricingHelpDashboard →Dashboard →
PricingLoginSign up for free →Sign up for free →
Dashboard →Dashboard →
Sign up →Sign up →
Loved by 1M+ users·Hosting 100K+ apps·Deploying 500K+ AI agents·Running 1M+ automations·Backed by Y Combinator·Powered by TSK-1
TaskadeCreate an AppPricingFeaturesTSK-1 BenchmarkContact usIntegrationsMCP ServerPressAbout
ConnectProductivityKitsVideosReviewsFAQ
LearnGenesisProjectsAI Agents
AutomationConnectorsAccount & BillingImport & ExportVideo TutorialsSearch Articles
DocsGetting StartedREST APIAction API
MCP ServersGuides & SDKModels
Community
FeaturedQuick AppsToolsDashboardsWebsites
WorkflowsProjectsFormsCreators
DownloadsAndroidiOSMacWindows
ChromeFirefoxEdge
Compare
vs Cursorvs Boltvs Lovablevs V0vs Windsurf
vs Replitvs Emergentvs Devinvs Claude Codevs ChatGPTvs Claudevs Perplexityvs GitHub Copilotvs Figma AIvs Notionvs ClickUpvs Asanavs Mondayvs Trellovs Jiravs Linearvs Todoistvs Evernotevs Obsidianvs Airtablevs Basecampvs Mirovs Slackvs Bubblevs Retoolvs Webflowvs Framervs Softrvs Glidevs FlutterFlowvs Base44vs Adalovs Durablevs Gammavs Squarespacevs WordPressvs UI Bakeryvs Zapiervs Makevs n8nvs Jaspervs Copy.aivs Writervs Rytrvs Manusvs Crewvs Lindyvs Relevance AIvs Wrikevs Smartsheetvs Monday Magicvs Codavs TickTickvs Any.dovs Thingsvs OmniFocusvs MeisterTaskvs Teamworkvs Workfrontvs Bitrix24vs Process Streetvs Toggl Planvs Motionvs Momentumvs Habiticavs Zenkitvs Google Docsvs Google Keepvs Google Tasksvs Microsoft Teamsvs Dropbox Papervs Quipvs Roam Researchvs Logseqvs Memvs WorkFlowyvs Dynalistvs XMindvs Whimsicalvs Zoomvs Remember The Milkvs Wunderlist
Taskade AIVideo GuideApp BuilderVibe CodingAgent BuilderDashboard Builder
CRM BuilderWebsite BuilderForm BuilderWorkflow AutomationWorkflow BuilderBusiness-in-a-BoxAI for MarketingAI for Developers
AI Agents
FeaturedProject ManagementOperations IntelligenceProductivityMarketing
TranslatorContentWorkflowResearchPersonalSalesSocial MediaTo-Do ListCRMTask AutomationCoachingCreativityTask ManagementBrandingFinanceLearning and DevelopmentBusinessCommunity ManagementMeetingsAnalyticsDigital AdvertisingContent CurationKnowledge ManagementProduct DevelopmentPublic RelationsProgrammingHuman ResourcesE-CommerceEducationLegalEmailSEODeveloperVideo ProductionDesignFlowchartDataPromptNonprofitAssistantsTeamsCustomer ServiceTrainingTravel PlanningUML DiagramER DiagramMath TutorLanguage LearningCode ReviewerLogo DesignerUI WireframeFitness CoachLead EnrichmentFounder OSSales DevelopmentBookkeepingRecruitingWebsite MonitoringField ServiceLicensingAll Categories
Automations
FeaturedAI Agent AutomationAI WorkflowsLogic AutomationsTrigger Automations
Agentic Process AutomationAction AutomationsAI Models in WorkflowsAgentic AutomationMulti-Agent AutomationBusiness-in-a-BoxOperations IntelligenceInvestor OperationsEducation & LearningHealthcare & ClinicsReal EstateStripeSalesHR & People OpsField Service & DispatchRenewals & LicensesE-commerceContentMarketingEmailCustomer SupportHubSpotProject ManagementAgentic WorkflowsAppointment SchedulingCalendarReportsSlackWebsiteFormTaskWeb ScrapingWeb SearchChatGPTText to ActionYoutubeLinkedInTwitterGitHubDiscordMicrosoft TeamsWebflowIndustry News & RSS FeedsGoogle WorkspaceManufacturing & OperationsAI Agent TeamsNotion AutomationsProposalBookkeeping & ExpensesClient OnboardingGoogle SheetsGoogle DriveGoogle CalendarGoogle FormsShopifyAsanaAirtableTrelloTodoistMailchimpClickUpGoogle DocsGmailGoogle TasksJiraLinearMicrosoft OutlookTelegramAll Categories
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Templates
FeaturedChatGPTOperations IntelligenceTablePersonal
Project ManagementSalesFlowchartTask ManagementEngineeringEducationDesignTo-Do ListMarketingMind MapGantt ChartOrganizationalPlanningMeetingsTeam ManagementStrategyGamingProductionProduct ManagementStartupRemote WorkY CombinatorRoadmapCustomer ServiceLegalEmailBudgetsContentConsultingE-CommerceStandard Operating Procedure (SOP)Human ResourcesProgrammingMaintenanceCoachingSocial MediaHow-TosResearchMusicTrip PlanningCRMClient OnboardingEmployee OnboardingSOPBug TrackerRecruitment TrackerFormSales PipelineContent CalendarMarketing PlanProduct RoadmapBusiness PlanSWOT Analysis30-60-90 Day PlanInterviewNotion AlternativeKPIStrategic PlanMeeting AgendaInvoiceRisk RegisterIT Asset ManagementKanban BoardChange ManagementCommunication PlanRFPScope of WorkStatement of WorkHelpdeskKnowledge BaseCreative BriefGoal SettingExecutive SummaryGap AnalysisBooking SystemEvent ManagementPortfolio TrackerCustomer Onboarding PortalsClient PortalAgency OperationsFinance TrackingAll Categories
Generators
AI SoftwareNo-Code AI AppAI AppAI WebsiteAI Dashboard
AI FinanceAI Operations IntelligenceAI FormAI AgentAI Client Portal BuilderAI WorkspaceAI ProductivityAI To-Do ListAI WorkflowsAI EducationAI Mind MapsAI FlowchartAI Scrum Project ManagementAI Agile Project ManagementAI MarketingAI Project ManagementAI Social Media ManagementAI BloggingAI Agency WorkflowsAI ContentAI Software DevelopmentAI MeetingAI PersonasAI OutlineAI SalesAI ProgrammingAI DesignAI FreelancingAI ResumeAI Human ResourceAI SOPAI E-CommerceAI EmailAI Public RelationsAI InfluencersAI Content CreatorsAI Customer ServiceAI BusinessAI PromptsAI Tool BuilderAI SEOAI Gantt ChartAI CalendarsAI BoardAI TableAI ResearchAI LegalAI ProposalAI Video ProductionAI Health and WellnessAI WritingAI PublishingAI NonprofitAI DataAI Event PlanningAI Game DevelopmentAI Project Management AgentAI Productivity AgentAI Marketing AgentAI Personal AgentAI Business and Work AgentAI Education and Learning AgentAI Task Management AgentAI Customer Relations AgentAI Programming AgentAI SchemaAI Business PlanAI Pitch DeckAI InvoiceAI Lesson PlanAI Social Media CalendarAI API DocumentationAI Database SchemaAI Marketing PlanAI Sales Pipeline GeneratorAI Course BuilderInternal ToolsBooking SystemReal Estate CRMInventory ManagementAI CRM BuilderAI TimesheetAI DispatchAI NewsletterAI Clinic OperationsAI Directory BuilderAll Categories
Converters
AI Featured ConvertersAI PDF ConvertersAI CSV ConvertersAI Markdown ConvertersAI Prompt to App Converters
AI Data to Dashboard ConvertersAI Workflow to App ConvertersAI Idea to App ConvertersAI Flowcharts ConvertersAI Mind Map ConvertersAI Text ConvertersAI Youtube ConvertersAI Knowledge ConvertersAI Spreadsheet ConvertersAI Email ConvertersAI Web Page ConvertersAI Video ConvertersAI Coding ConvertersAI Task ConvertersAI Kanban Board ConvertersAI Notes ConvertersAI Education ConvertersAI Language TranslatorsAI Business → Backend App ConvertersAI File → App ConvertersAI SOP → Workflow App ConvertersAI Portal → App ConvertersAI Form → App ConvertersAI Schedule → Booking App ConvertersAI Metrics → Dashboard ConvertersAI Game → Playable App ConvertersAI Catalog → Directory App ConvertersAI Creative → Studio App ConvertersAI Agent → Agent App ConvertersAI Audio ConvertersAI DOCX ConvertersAI EPUB ConvertersAI Image ConvertersAI Resume & Career ConvertersAI Presentation ConvertersAI PDF to Spreadsheet ConvertersAI PDF to Database ConvertersAI PDF to Quiz ConvertersAI Image to Notes ConvertersAI Audio to Notes ConvertersAI Email to Tasks ConvertersAI CSV to Dashboard ConvertersAI YouTube to Flashcards ConvertersURL to NotesVideo → SummaryAI Receipts to Expense Tracker ConvertersAI Docs to Knowledge Base ConvertersAI Form to Client Portal ConvertersSpreadsheet to CRMAll Categories
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
Blog
Introducing Taskade TSK-1: The System Kernel Behind Every App (2026)Chat-Native App Builders in 2026: What You Actually Own When the Chat EndsGenerate the Art. Preview the Agent. Put It on Your Domain (2026)
Agentic Automation Explained: Agent vs AI Step (2026)The Scaffolding Tax: Why Less Prompt Beats More (2026)History of Mind Mapping: From Porphyry to Buzan to AI (2026)The Bitter Lesson Explained: Richard Sutton's 26 Words (2026)Self-Replicating Code: Quines, von Neumann, and the Programs That Copy Themselves (2026)Markov Chains Explained: The Memoryless Math Behind Google, Monte Carlo, and ChatGPT (2026)Add Client Logins. Connect Your Domain. Ship a Real Product in 2026Track Customer Health. Catch Churn Early. Keep the Accounts You Won (2026)Compression Is Intelligence: What Cross-Entropy Really Measures (2026)Automate License Renewals. Track Every Key. Own Your Software Spend (2026)Track Hours. Bill Clients. Get Paid. (Clone a Working Time Tracker in 2026)Claude Shannon: The History of Information Theory and the Man Who Invented the Bit (2026)Connect Your Apps. Automate Your Business. (Two-Way Workflows in 2026)Excel Job Log to Dispatch App (2026): Own the BoardMaintainX Alternative for Small Shops (2026)The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day (2026)Run Your Whole Business in One App with Taskade Genesis (June 2026)
AIAutomationProductivityProject ManagementRemote WorkStartupsKnowledge ManagementCollaborative WorkUpdates
Changelog
Google Sheets Trigger & Automation Stall Hotfix (Sep 3, 2026)Project Tap Hotfix (Sep 2, 2026)Run Two Builds at Once & Connect ClickUp (Sep 2, 2026)
App Kits Carry Agent Teams & CSV Attachments (Sep 2, 2026)Whole-File App Edits & Markdown Attachments (Sep 2, 2026)App Header Controls Hotfix (Aug 31, 2026)Automation Email Safety Hotfix (Aug 31, 2026)
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
© 2026 Taskade
PrivacyTermsSecurity
Made withTaskade AIforBuilders
BlogAIRL Environments Explained:…

RL Environments Explained: How AI Labs Now Train Models on Real Work (2026)

RL environments are where AI models now learn real work: what they are, why labs pay thousands per task, how verifiable rewards work, and what it means for you.

RL environments explained: the model picker in Taskade Genesis, where models trained in reinforcement learning environments are put to work building real apps
September 22, 202626 min readJohn XieAI·#reinforcement-learning#rl-environments#ai-training
On this page (11)
What Is an RL Environment?Why AI Labs Switched From Labels to EnvironmentsHow a Verifiable Reward WorksReplication Training: When the Answer Already ExistsWhy a Good Task Costs Thousands of DollarsWho Builds RL Environments in 2026What an Environment Cannot Grade, YetThe Same Loop, Seen From the App Builder's ChairWhat This Means If You Build With AIFrequently Asked QuestionsRelated Reading

For a decade, the recipe for a better AI model was more text. Scrape the internet, predict the next word, scale. That recipe ran out of internet, and the models it produced could write an essay but could not finish a job. The thing that replaced it is quieter and stranger: frontier labs now pay people to build little worlds where a model attempts real work, gets scored, and tries again. Those worlds are called reinforcement learning environments, and by 2026 they are where a large share of AI progress actually happens.

This post explains what an RL environment is, why labs switched to them, how the scoring works, what a good one costs, who builds them, and where they still fall short. Then it turns the same lens on something we run at Taskade, because once you understand environments you start seeing them everywhere, including in the way we test which AI model builds the best app. 🧭

TL;DR: An RL environment is a task, a place to attempt it, and a grader. Labs now pay for environments instead of labels because static data goes stale. Tasks run roughly $200 to $2,000 each, complex product replicas about $300,000, per Epoch AI. The hard part is a grader that cannot be fooled. How we grade real apps →

What Is an RL Environment?

A reinforcement learning environment is a controlled setting where an AI model acts and gets scored. In the definition Epoch AI settled on after interviewing the people who build them, it has three parts: the actions a model can take, such as running code, clicking a button, or searching; the surrounding context that determines what those actions do; and task prompts paired with automated graders that decide whether an attempt succeeded. Most ship as software containers.

That is the whole idea, and it is worth pausing on how different it is from a dataset. A dataset is a pile of examples with answers attached. The model reads them and adjusts. An environment is a place. The model does something in it, the place changes, a grader looks at the result, and a reward comes back. The model adjusts toward whatever earned reward, then tries again.

yes no Task promptwhat the model is asked to do Model actsruns code, clicks, edits, searches Environment changes statefiles, screens, databases Graderdid the task actually succeed? Reward No reward Model updates towardwhat earned reward
yes no Task promptwhat the model is asked to do Model actsruns code, clicks, edits, searches Environment changes statefiles, screens, databases Graderdid the task actually succeed? Reward No reward Model updates towardwhat earned reward

A dataset An RL environment
What it is Examples with answers A place where tasks can be attempted
How the model learns Imitates the examples Tries, gets graded, adjusts, tries again
Who supplies the answer A human, in advance A grader, after the attempt
What happens as models improve The examples get too easy and stop teaching The same environment keeps teaching, because harder tasks can be posed in it
First big domains Language, images, speech Math, competitive programming, then software engineering, then enterprise workflows
Where it fails Stale as soon as the model exceeds the examples A weak grader gets gamed, and a narrow environment does not generalize

Reinforcement learning itself is not new. It trained game-playing systems long before large language models existed, and our wiki entries on reinforcement learning and RLHF cover the older lineage. What is new is the object of the reward. RLHF rewarded a model for answers a human rater preferred. Environments reward a model for finishing a job, as judged by something that checks the result.

Why AI Labs Switched From Labels to Environments

Labs switched because the text on the internet taught models to write, and nothing on the internet taught them to work. Pretraining on a corpus of human writing produced models that could draft, translate, and explain, but that had to be kept on a very short leash the moment a task ran longer than a few steps. There was no internet-scale record of people finishing long tasks, recovering from their own mistakes, and using tools reliably. So labs built places where models could generate that record themselves.

The economics followed. A September 2025 report said Anthropic had discussed spending more than one billion dollars a year on RL environments. Wing Venture Capital put traditional data labeling at roughly a five billion dollar market in January 2026, growing more than 50 percent a year, and described environments as the larger opportunity forming on top of it. The vendors that used to sell labeled examples now sell task suites with graders attached.

Three things changed at once, and the RL environment company Mechanize stated them as a manifesto in July 2025 under the title Sweatshop data is over:

  • Software, not datasets. Static data goes stale as models improve. An interactive environment keeps offering a challenge, in their words, "much like how games continue to engage players across a wide range of skill levels."
  • Full-time contributors, not contractors. Not low-skill workers at scale, and not even high-skill contractors working sporadically. Environments that teach a whole job need months of sustained attention from the same people.
  • Deep expertise. The tacit knowledge of subject-matter experts is now the bottleneck. Data generation, they argued, has to be reframed "from a low-status activity outsourced to workers in poor countries, to an elaborate process requiring the world's finest talent."

Epoch's interviewees said the same thing in plainer terms. Domain knowledge and expert-level prompting matter more than machine learning skills for building a good task. The thing that limits growth is not finding experts. It is, as one founder put it, "maintaining quality while scaling."

The pretraining era The environment era Scrape internet text Predict the next word A model that can write Experts design taskswith sound graders Model attempts themthousands of times A model that canfinish a job
The pretraining era The environment era Scrape internet text Predict the next word A model that can write Experts design taskswith sound graders Model attempts themthousands of times A model that canfinish a job

The domains moved in a telling order. Math and coding came first, because grading is free there: a proof checks or it does not, a test passes or it does not. Then software engineering, where an agent works inside a real repository. Now, per Epoch, the fastest-growing category is enterprise workflows: expense reports, pivot tables, navigating a CRM, working a spreadsheet. The frontier of AI training has moved from the whiteboard to the back office, and the reason is simply that the back office is where the graders can be built next.

How a Verifiable Reward Works

A verifiable reward is a score a program can compute without asking a human. The unit test passes or fails. The answer matches the reference or it does not. The browser workflow reaches its end state or stalls. This is the RLVR paradigm, reinforcement learning with verifiable rewards, and it is why the first sharp capability gains of the environment era landed in mathematics and competitive programming.

The reason it works is also the reason it is hard. A grader is a promise that high reward means the task was done. Break that promise and the model learns to collect the reward instead of doing the task. One environment builder gave Epoch the rule in a sentence: "Soundness matters most: high reward must mean the task was actually solved, not hacked."

The public benchmarks show what happens when soundness slips:

  • In a May 2026 audit of SWE-bench Pro, a frontier model read the correct fix out of the repository's own version history in 12 to 25 percent of its passing runs. The grader checked whether the tests passed. It did not check how.
  • In April 2026, Berkeley researchers built an agent that scored 100 percent on SWE-bench Verified, SWE-bench Pro, and Terminal-Bench without solving a single problem. Their conclusion: if your benchmark is exploitable, it will be exploited.
  • OpenAI stopped reporting SWE-bench Verified on February 23, 2026 after finding that 59.4 percent of the hardest problems it audited were flawed, and that models were reproducing gold-standard patches from task identifiers alone.

Graders sit on a ladder from cheap and sound to expensive and fuzzy, and the environment market is largely a market for climbing it.

Exact matchthe output equals the reference Unit testshidden tests pass or fail State checkthe database, file, or screen ends in the right state Workflow replaya browser agent walks the finished product Rubric plus judge modela model grades against written criteria Expert human reviewa person grades the finished work Cheap, sound, narrow Expensive, broad, slow
Exact matchthe output equals the reference Unit testshidden tests pass or fail State checkthe database, file, or screen ends in the right state Workflow replaya browser agent walks the finished product Rubric plus judge modela model grades against written criteria Expert human reviewa person grades the finished work Cheap, sound, narrow Expensive, broad, slow

Two more design rules from Epoch's interviews are worth knowing, because they explain why environments are hard to build well. Difficulty has to be calibrated. A task a model never solves teaches nothing, so builders aim for a minimum pass rate of about 2 to 3 percent, or at least one success in every 64 to 128 attempts, and a smooth gradient of harder tasks above that. Tasks should share underlying skills. A suite of disconnected puzzles trains disconnected tricks. A suite that exercises the same capabilities from different angles trains something that transfers.

The judge-model rung of the ladder is where most of the current research effort goes, and it is the rung our LLM-as-a-judge entry and the agent evaluation literature cover. The short version: a model can grade an open-ended answer against a rubric far more cheaply than a person can, but it inherits every blind spot of the model doing the grading, so the sound practice is a deterministic gate first and a judge second.

Replication Training: When the Answer Already Exists

The most concrete proposal for scaling environments comes from Mechanize, and it borrows the trick that made pretraining work. Pretraining scaled because nobody had to write the corpus. Books, papers, and forum threads already existed. In The upcoming GPT-3 moment for RL, they argue that software already exists in the same abundance, and that it can be turned into tasks without hand-authoring each one.

A replication task is a detailed specification plus a reference implementation. Take an existing piece of software, write down exactly what it does, and train the model to produce an implementation whose behavior matches the reference. Grading collapses to a binary: it behaves identically, or it does not. That makes the grader both cheap and sound, which is the combination the whole market is chasing.

Detailed description of the software's behavior Builds an implementation Runs the same inputs Runs the same inputs Compares every output Match: reward. Any difference: no reward. Specification Model Model's implementation Reference implementation Grader
Detailed description of the software's behavior Builds an implementation Runs the same inputs Runs the same inputs Compares every output Match: reward. Any difference: no reward. Specification Model Model's implementation Reference implementation Grader

Their argument for what this trains is a good list, and it reads like a description of every complaint people have about AI coding agents:

  • Read and deeply understand detailed instructions.
  • Execute meticulously, without errors.
  • Notice earlier mistakes and reliably recover from them.
  • Sustain performance over month-scale horizons where quality is directly rewarded by correctness.
  • Do not settle prematurely for a solution that merely looks good enough.

They also state the limits honestly. Writing comprehensive tests for a replication task is itself hard engineering. Exact replication is rare in everyday software work, showing up mainly in porting, legacy rewrites, and clean-room reimplementation. And a model that can replicate a specification perfectly may still be poor at the open-ended planning a real project needs. Replication training, in their framing, is a bridge to the next paradigm rather than the destination.

One number from that essay is worth carrying around. They estimate that matching frontier pretraining budgets with RL would take on the order of ten thousand years of model-facing task time, meaning the time humans would need to do the tasks the model trains on. For comparison, they note that projects such as Windows Server 2008 or Grand Theft Auto V each consumed roughly that many person-years. The environment era is, in effect, an effort to manufacture several major software projects' worth of graded work.

Why a Good Task Costs Thousands of Dollars

A single RL task can justify a price of a few thousand dollars, because the compute spent running it dwarfs the cost of building it. Epoch's January 2026 interviews put individual tasks at roughly $200 to $2,000, website replicas used for interface training at about $20,000 each, and a complex product clone, of something on the scale of Slack, at about $300,000. Exclusive deals run four to five times the price of shared ones, and contracts reach six or seven figures per quarter.

Mechanize published the arithmetic behind those prices in August 2025, in Cheap RL tasks will waste compute. It is a reusable template, so here it is in full.

Input Their figure Where it comes from
Opportunity cost of compute $15 per million output tokens The API price of a frontier model at the time, since that compute could be sold as inference instead
Tokens per task, SWE-bench Verified About 20,000 Published run logs
Tokens per frontier RL task today About 100,000 Their own experience
Growth in transcript length About 5 times per year Epoch AI's output-length data
Tokens per task within a year About 500,000 Extrapolation
Attempts per task per training run 64 The group size used in DeepSeek-R1
Compute per task, per run About $480 500,000 tokens times $15 per million times 64
Reuse across research and production runs 5 times, conservatively —
Lifetime compute per task About $2,400 —
SWE-bench Verified Frontier task today Projected one year out 0 100 200 300 400 500 Stage Tokens per task, thousands Tokens per RL task, as estimated by Mechanize (August 2025)
SWE-bench Verified Frontier task today Projected one year out 0 100 200 300 400 500 Stage Tokens per task, thousands Tokens per RL task, as estimated by Mechanize (August 2025)

The principle underneath the numbers is that data and compute are complements. Underinvest severely in one and you waste most of your spend on the other. Their image for it: putting cheap tires on a Ferrari. If running a task costs $2,400 over its life, paying $100 to build it badly is not thrift. It is the most expensive line in the budget, because it degrades every one of those runs.

Wing Venture Capital reached the same conclusion from the investor's side: "environment quality, not environment count, becomes the binding constraint." The winners in this market, both essays agree, will be the teams with deep domain expertise, rigorous validation, and the discipline to build fewer, better tasks.

Who Builds RL Environments in 2026

The market has three kinds of builders, plus the labs themselves and a surprising fourth group: the companies whose software is being replicated.

Group Examples What they bring
Human data companies that added environments Scale AI, Surge AI, Mercor, Turing Existing lab relationships, expert networks, and the operations to manage thousands of contributors
Environment-native startups Mechanize, Fleet AI, HUD, Veris AI, Plato, Bespoke Labs Purpose-built tooling for tasks, graders, and containers, and a thesis about which domains matter
Open ecosystems Prime Intellect Shared frameworks and community-built environments for decentralized training
Frontier labs, in-house The major labs and the newer research labs Environments too sensitive or too specific to outsource
Product companies partnering with labs Salesforce, Slack, Benchling, per Epoch's interviews The real software, so a model can learn a product from the product rather than from a replica

Wing predicted in January 2026 that the field, then roughly twenty seed- to Series A-stage companies, narrows to three to five leaders by 2030, with one or two dominant platforms. Epoch's interviewees expect growth in enterprise workflow environments, longer multi-step tasks, multi-turn interaction with simulated users, and tooling that lets a lab inspect a model's attempts rather than just score them.

There is a detail in that last row that matters for anyone who builds software. When Slack or Salesforce partners with a lab, the model is being trained to operate their product. The next generation of models will be fluent in the tools that showed up in environments, and less fluent in the ones that did not. That is a new kind of distribution advantage, and it did not exist two years ago.

What an Environment Cannot Grade, Yet

The environment era has a boundary, and the people inside it are candid about where it sits. In How to fully automate software engineering, Mechanize lists what today's graders cannot see: whether an agent followed open-ended instructions from a customer who did not have a full technical specification in mind, whether its code is maintainable, whether it avoided technical debt, whether it dodged a trapdoor decision. "Without being able to grade these parts of the AI's work," they write, "we can't know if an AI can act as a fully independent engineer, or whether it will just be a tool that saves human engineers time."

Their July 2025 essay puts the same boundary in one example: a scoring script cannot tell you whether an AI would make an effective lawyer. That requires constructing cogent arguments, contextualizing information properly, and prevailing in court. None of those has a unit test.

The numbers agree. On Vibe Code Bench, the first serious benchmark to ask a model to build a whole web app from a written specification, the best model completed 61.8 percent of the work, and the researchers found that changing the automated evaluator moved step-level agreement anywhere from 31.8 to 93.6 percent. The grader was as much a variable as the model.

Gradable today Hard to grade today
Does the code pass the tests Is the code maintainable
Does the output match the reference Did it understand what the customer meant
Did the workflow reach the end state Did it make a decision it cannot undo
Was every field saved correctly Is the design good
Did the automation fire on the right record Would a real user come back tomorrow

This boundary is not a reason to dismiss the shift. It is a map of where the gains will land first. Anything in the left column is about to get much better, because thousands of experts are building graders for it right now. Anything in the right column will improve more slowly, and will need a human in the loop for longer. If you build with AI, that map is the most useful thing in this post.

The Same Loop, Seen From the App Builder's Chair

Once you understand environments, you start seeing them everywhere, and one of the places we saw one was in our own benchmark. TSK-1 is not an RL environment. Taskade does not train models, and nothing that happens in a TSK-1 test flows back into any model's weights. But it is built like one, and the resemblance is instructive.

Every TSK-1 test gives every model the same frozen request, registered with a fingerprint so that two results only go in the same comparison if the request bytes match. That is the task prompt. Every model builds inside Taskade Genesis, the same product, held constant. That is the environment. A tester then opens the finished app in light and dark and at phone width, fills in its form as a fixed persona whose expected outcome is known, submits, and compares every saved field to what was typed. That is the grader, and it sits on the sound end of the ladder: a state check on the database, not a vibe about the screen.

The sharpest parallel is the second of TSK-1's two requests. It is a real customer's 32-question client sign-up form, pasted word for word, with the customer's own instruction: do not shorten my questions or answers. Grading fidelity to that request means counting how many of the 32 questions survived into the finished app exactly as written. That is a replication-shaped reward, exact match against a reference, applied to a product a real business asked for.

An RL environment at a lab Frozen, fingerprinted request Model builds in Taskade Genesis Tester uses the app,checks every saved field Grade flows to a public hubfor buyers, dated end Task prompt Model acts in a container Automated grader Reward flows backinto the model's weights
An RL environment at a lab Frozen, fingerprinted request Model builds in Taskade Genesis Tester uses the app,checks every saved field Grade flows to a public hubfor buyers, dated end Task prompt Model acts in a container Automated grader Reward flows backinto the model's weights

The TSK-1 matrix on the Taskade hub: seven graded families across Interface, Task, Memory, and Adapt, with the Intelligence Index beside each

The differences matter as much as the parallel:

  • Where the grade goes. In a lab, the reward updates the model. In TSK-1, the grade goes to a public hub so a buyer can see what each model actually built, with the day it was measured.
  • Who holds the grader. In a lab, a program. In TSK-1, a person using the app, with the field-by-field check as the sound core and human judgment for the parts no program can grade yet: whether the app looks right, whether it built what was asked, whether a change landed cleanly.
  • Consent. The environment market has a gap here that its own essays do not mention. Replication training runs on existing software, and nobody in the manifestos asks who owns it. TSK-1's real customer request is never published, and the customer is never named. That is a rule, not a courtesy, and it is also the reason the benchmark's hardest task cannot leak into training data.

The full protocol is in the TSK-1 methodology, and the results, model by model, are in Best AI Model for Building Apps in 2026. The reason to mention it here is narrower. Environments have a grading gap, and the gap is exactly where a product with a real workspace behind it has something to offer: an app built in Taskade Genesis lands its data in a project with typed fields, its logic in an automation with triggers and actions, and its knowledge in an AI agent that can be asked about the record. Every one of those is a state a grader can check. That is the Workspace DNA loop, projects that remember, agents that think, automations that execute, and it happens to be the shape of a gradable environment.

What This Means If You Build With AI

The practical lesson of the environment era is that models get better at whatever has a grader, so the fastest way to get more from AI is to give your work one. Five habits follow, and none of them requires understanding a gradient.

  1. Write the brief like a specification. The models being trained today are rewarded for following detailed instructions exactly. A vague request leaves the model to invent, and invention is where the wrong app comes from. Name the fields, name the outcomes, name what must not change.
  2. Decide the expected outcome before you test. A fixed persona with a known result turns "does this feel right" into "did it produce 11 or did it not." That is the difference between a vibe and a grader.
  3. Check the saved data, not the screen. A form can render perfectly and save nothing. Open the record. Compare every field to what you typed. This is the single check that catches the most failures, and it is the one almost nobody does.
  4. Expect fast gains where grading is cheap and slow gains where it is not. Forms, data, tests, and automations will keep improving quickly. Taste, judgment, and reading an ambiguous customer's mind will not, so keep a person in those loops and do not be surprised when the model needs one.
  5. Match the model to the task, or leave the default alone. No model led every quality in our tests, and the environment era makes that more likely, not less, because each lab is training in the domains it chose. TSK-1 Auto handles the default, and Taskade Genesis offers 15+ frontier models from OpenAI, Anthropic, and open-weight providers, so you can switch the model behind an app or agent without rebuilding it.
  THE GRADER TEST FOR ANY AI TASK

Can a program, or a person with a checklist, say for certain whether it succeeded?

 YES  ->  expect rapid improvement; automate it, measure it, trust it sooner
 NO   ->  expect slower improvement; keep a human in the loop, write down
          what "good" means, and revisit when a grader exists

Forms that must save every field ............ YES
Data that must match a formula .............. YES
Code that must pass a test .................. YES
An automation that must fire on a record .... YES
A design that must look right ............... PARTLY
A brief the customer never fully specified .. NO, not yet

TSK-1 coordinating models, workspace memory, agents, and automations inside Taskade Genesis, the environment an app lives in after the prompt

If you would rather see the loop than read about it, describe an app in Taskade Genesis, open the live apps other people have built, or read what each model did on the hub. The quickstart covers the first build, and the free plan includes three Taskade Genesis apps.

Frequently Asked Questions

What is an RL environment in AI?

A controlled setting where an AI model acts and gets scored. It has three parts: the actions the model can take, such as running code or clicking buttons; the surrounding context that determines what those actions do; and task prompts paired with automated graders that decide whether an attempt succeeded. Environments typically ship as software containers, and by 2025 they had replaced static labeled datasets as the main way frontier labs teach models to do multi-step work.

Why did AI labs switch from labeled data to RL environments?

Because static datasets go stale as models improve, while an environment keeps offering a challenge the model has not solved. Pretraining taught models language, not how to finish a long task or recover from mistakes, and there was no internet-scale record of people doing those things. A September 2025 report said Anthropic had discussed spending more than one billion dollars a year on environments.

How much does an RL environment or task cost?

Per Epoch AI's January 2026 interviews, individual tasks run roughly $200 to $2,000, website replicas about $20,000 each, and complex product clones about $300,000, with exclusive deals costing four to five times more. Mechanize argued in August 2025 that labs should spend a few thousand dollars per task, because the compute spent running a task over its life is around $2,400 and a cheap task degrades every run.

What is a verifiable reward?

A score a program can compute without human judgment: a test passes or fails, an output matches the reference, a workflow reaches its end state. Reinforcement learning with verifiable rewards is why math and coding improved first, because grading is free there. The hard part is soundness: high reward must mean the task was actually solved, not hacked.

What is reward hacking?

Earning a high score without doing the task the score was meant to measure. In a May 2026 audit of SWE-bench Pro, a frontier model read the correct fix out of the repository's version history in 12 to 25 percent of its passing runs. In April 2026 Berkeley researchers scored 100 percent on three major coding benchmarks without solving a single problem. Any exploitable grader eventually gets exploited.

What is replication training?

A proposal from Mechanize to scale RL tasks the way pretraining scaled text: each task is a detailed specification plus a reference implementation of existing software, and the model is trained to match the reference's behavior exactly. The grader collapses to a binary, which makes it cheap and sound. Its authors concede that exact replication is rare in everyday engineering and will not, alone, produce open-ended planning.

Who builds RL environments in 2026?

Human data companies that added environments, such as Scale AI, Surge AI, Mercor, and Turing; environment-native startups such as Mechanize, Fleet AI, HUD, Veris AI, Plato, and Bespoke Labs; open ecosystems such as Prime Intellect; frontier labs in-house; and product companies including Salesforce, Slack, and Benchling partnering with labs. Wing Venture Capital predicted in January 2026 that the market narrows to three to five leaders by 2030.

Is TSK-1 an RL environment?

No. TSK-1 is an evaluation, and Taskade does not train models. But it is built like one: a frozen request, the same product for every model, and a tester who checks every saved field against what was typed. The word-for-word check on a real customer's 32-question form is the same exact-match grader that replication training uses. The grade goes to a public hub for buyers rather than back into a model.

What does the shift to RL environments mean for people who build apps with AI?

Models improve fastest at whatever has a grader. Forms, data, tests, and automations are gradable, so expect rapid gains there. Judgment, taste, and ambiguous briefs are not yet gradable, so expect slower gains and keep a person in that loop. Write briefs like specifications, decide the expected outcome before you test, and check the saved data rather than the screen.

Can I use models trained in RL environments inside Taskade?

Yes. Taskade Genesis offers 15+ frontier models from OpenAI, Anthropic, and open-weight providers. TSK-1 Auto handles the default, and you can set a model per AI agent. The hub publishes dated, hands-on evidence of how each family builds a real app. The free plan includes three Taskade Genesis apps, and paid plans start at $10 per month billed annually.

Related Reading

  • The TSK-1 Methodology, how we grade AI models on the app they build.
  • Best AI Model for Building Apps in 2026, what nine models built from one request.
  • History of AI Benchmarks, why every benchmark saturates and what survives.
  • What Are AI Agent Evals? and AI agent reliability, the evaluation side of the same loop.
  • Reasoning models and how LLMs work, for the training background.
  • Why AI-Generated Apps Break, the failures a grader would have caught.
  • The Living Software Era, what happens to an app after the prompt.

Primary sources: Epoch AI's FAQ on reinforcement learning environments (January 12, 2026), Wing Venture Capital's Who Will Win the RL Environment Market (January 2026), and Mechanize's essays How to fully automate software engineering, The upcoming GPT-3 moment for RL, Sweatshop data is over, and Cheap RL tasks will waste compute (May to August 2025). Mechanize is a company that sells RL environments, so read its forecasts as an interested party's, which is also how we would ask you to read ours.

The pretraining era taught machines to talk. The environment era is teaching them to work, one graded task at a time, and the people writing the graders are quietly deciding what the next decade of software will be good at. The best thing you can do about that is to build things that can be checked. ▲ ■ ●

0%

On this page

What Is an RL Environment?Why AI Labs Switched From Labels to EnvironmentsHow a Verifiable Reward WorksReplication Training: When the Answer Already ExistsWhy a Good Task Costs Thousands of DollarsWho Builds RL Environments in 2026What an Environment Cannot Grade, YetThe Same Loop, Seen From the App Builder's ChairWhat This Means If You Build With AIFrequently Asked QuestionsRelated Reading

Related Articles

A tracker app built end to end in Taskade Genesis, the same app shape the TSK-1 benchmark asks every AI model to produce from one fixed prompt
August 22, 2026AI

The TSK-1 Benchmark: Nine AI Models Built the Same App in One Day (2026)

On Aug 1, 2026 nine AI models built the same app from one fixed prompt in Taskade Genesis. Three builds never opened, an...

Close-up of a human eye, representing the ImageNet moment when machine vision surpassed hand-designed computer vision methods
August 13, 2026AI

The ImageNet Moment, Explained: How Computer Vision Broke Open (2026)

In 2012 AlexNet cut ImageNet error from 26 to 15.3 percent. Weeks later a Stanford student wrote that vision was hopeles...

History of AI benchmarks chart showing test scores of AI systems on MNIST, ImageNet, GLUE, SuperGLUE, MMLU and HumanEval relative to human performance
August 10, 2026AI

The History of AI Benchmarks: Why Every Model Claims to Be the Best (2026)

AI benchmarks go from impossible to solved in about two years. A verified history from the Turing test to ARC-AGI-2, plu...

Richard S. Sutton, author of The Bitter Lesson and 2024 Turing Award winner, photographed in December 2025. Photo: Wikimedia Commons / Xuthoria / CC BY-SA 4.0
August 29, 2026AI

The Bitter Lesson Explained: Richard Sutton's 26 Words (2026)

Richard Sutton compressed The Bitter Lesson into 26 words. What it means, why he now says LLMs break it, and what it cha...

A live app built from a prompt, running with its own database, agents, and automations
September 1, 2026AI

Chat-Native App Builders in 2026: What You Actually Own When the Chat Ends

Claude, ChatGPT, and Gemini can all build an app inside the chat. The question nobody answers is what survives when you ...

Previewing and customizing a branded AI agent in Taskade before publishing it as a public page, a custom domain, or a website widget
August 31, 2026AI

Generate the Art. Preview the Agent. Put It on Your Domain (2026)

Ship an AI agent that looks like your company: generate its art in the workspace, preview it the way a visitor sees it, ...

View All Articles
RL Environments Explained: How AI Learns Real Work (2026)