Definition: A long-horizon agent is an AI agent that can complete tasks that take a skilled person hours or days, not seconds. The standard way to measure it is the 50% task-completion time horizon: the length of task, in human-professional time, that an agent finishes successfully half the time. The research nonprofit METR introduced the metric in March 2025, and it has become the field's common yardstick for agent capability.
TL;DR: Time horizon measures how long a task, in human hours, an agent finishes half the time. METR found it doubled about every 7 months since 2019, reaching about 12 hours by early 2026. Long runs still fail from lost context, so the fix is a written plan, one task per session, and a progress file. Build one free →
Think of a contractor renovating a kitchen over three weeks. Nobody holds the whole job in their head. There is a plan on the wall, a punch list, and a note at the end of each day that says what got done and what comes next. When a new crew member shows up on day nine, they read the list and start work. Long-horizon agents need exactly the same thing, because every new session starts with no memory of the last one.
Why Long-Horizon Agents Matter in 2026
Time horizon turned "is this agent good?" into a number you can plot over time, and the plot is steep. In Measuring AI Ability to Complete Long Tasks, METR timed skilled people on 170 software and research tasks, then ran frontier models on the same tasks. The length of task a model could finish with 50% reliability had doubled roughly every 7 months for 6 years, and the authors noted the trend "may have accelerated in 2024". Claude 3.7 Sonnet scored a horizon of about 50 minutes in that original study.
METR checked whether software was a special case. Its July 2025 cross-domain study covered math contests, competitive programming, science questions, video, driving, and computer use. It found no domain where progress was clearly slower than exponential, and it found its own software benchmark "is not an outlier for seeing ~100 minute horizons doubling every ~4 months since 2024." In January 2026 METR released Time Horizon 1.1, which grew the suite from 170 to 228 tasks and more than doubled the number of tasks that take a person 8 hours or more. The post-2023 doubling time came out about 20% faster (130.8 days versus 165.3).
The values come from METR's published Time Horizon 1.1 data (last updated May 8, 2026). GPT-2's bar is invisible for a reason: about 3 seconds. o3 sits near 2 hours, Claude Opus 4.5 near 5, and Claude Opus 4.6 near 12. METR's May 2026 entry for an early Claude Mythos Preview came out near 17 hours, but METR now states that "measurements above 16 hrs are unreliable with our current task suite," so the chart stops below that line. Each bar also carries wide error bars. The shape is the story: capability that compounds on a doubling clock, which is why the question in 2026 shifted from "can an agent do this step?" to "can it stay on track for a whole project?"
| Milestone | What METR reported | Source |
|---|---|---|
| March 2025 | 170 tasks, doubling about every 7 months since 2019, Claude 3.7 Sonnet near 50 minutes | Original paper |
| July 2025 | No domain "clearly sub-exponential" across software, math, science QA, computer use, and driving | Cross-domain study |
| January 2026 | 228 tasks, 8-hour-plus tasks up from 14 to 31, post-2023 doubling of 130.8 days | Time Horizon 1.1 |
| May 2026 | Frontier near 12 hours, with a warning that results above 16 hours are unreliable | Time horizons page |
How Long-Horizon Agents Work
A model with a long time horizon still fails a long job if the job outlives its context window. Anthropic's engineering team put it plainly in Effective harnesses for long-running agents (November 2025): long-running agents "must work in discrete sessions, and each new session begins with no memory of what came before." The answer is a harness that writes the plan and the progress outside the model.
- Plan once, in writing. An initializer session turns the goal into a full feature list. In Anthropic's example that list held more than 200 features in JSON, each with steps and a
passesflag that starts as false. Later sessions can flip the flag, and the prompt told them it is "unacceptable to remove or edit tests." - Pick one task. The agent reads the progress file and the git log, then chooses one unfinished feature. Working on one thing at a time prevents the most common failure, trying to build the whole app in a single pass.
- Execute inside a clean context. The session holds only what this task needs, which keeps it clear of context rot.
- Verify before you mark it done. Anthropic saw agents declare a project finished after partial progress, and mark features complete without an end-to-end test. Verification is its own step.
- Checkpoint. A descriptive commit plus a note in the progress file (
claude-progress.txtin Anthropic's harness) leaves the code in a state someone can merge. - Hand off. The session ends. The next one starts fresh and reads the files, not a memory. That is a session boundary by design.
The files between sessions are plain text. A progress note in this style is short and boring on purpose, because the next session has to trust it:
PROGRESS NOTE (example, written at the end of a session)
------------------------------------------------------------
Goal: Customer portal, 42 features in feature_list.json
Done today: #17 password reset -> passes: true (tested end to end)
In progress: #18 invoice export -> passes: false
Blocked: #23 webhook retries -> needs an API key from a person
Next: finish #18, then #19 (invoice email)
Rules: one feature per session, never edit a test to pass
------------------------------------------------------------
Anthropic's post named the failure modes this pattern answers. Each one maps to one part of the harness:
| Failure mode Anthropic observed | What the agent did | Harness fix |
|---|---|---|
| One-shotting | Tried to build the whole app in one pass | One feature per session |
| Early victory | Declared the project finished after partial progress | A feature list with a passes flag for every item |
| Untested "done" | Marked features complete without an end-to-end test | A separate verification step before the flag flips |
| Lost handoff | Left half-done work with no note for the next session | A progress file plus a descriptive commit |
Anthropic's March 2026 follow-up split the work across three agents: planner, generator, and evaluator. The reason was blunt. When asked to evaluate their own output, agents "tend to respond by confidently praising the work," so a separate evaluator does the checking. The same post explains the trade-off between a full context reset and compaction. Compaction keeps continuity but "doesn't give the agent a clean slate." An earlier model wrapped up work too early as it neared its context limit, so the team needed full resets. With a newer model the team dropped resets and ran one continuous session with automatic compaction. The harness has to change as the model changes.
Limits of the Time Horizon Metric
METR is also the most candid critic of its own number. In a January 2026 note, METR's Thomas Kwa listed what the metric does not mean:
- It is not how long an agent can run. It measures how much serial human labor the agent can replace at 50% success, not hours of wall-clock autonomy.
- Fifty percent is not usable reliability. Some tasks need 98%+ success to be worth automating. In the original paper, 80% horizons were roughly 5x shorter than 50% horizons.
- Error bars are wide. Roughly a factor of 2 in each direction, and wider for the newest models.
- The tasks are tidy. They are self-contained software and research problems with automatic scoring. The paper found models do worse on "messier" tasks without clear feedback loops.
Outside critics push further. Oxford philosopher Toby Ord showed in May 2025 that METR's data fits a simple model: a constant chance of failure for every minute of human work. On that model, doubling a task's length squares the success rate, which points to tasks built from sequential steps and agents that are "not very good at recovering from earlier mistakes." Ord added an update in February 2026: new analysis by Gus Hamilton suggests agents probably do not fail at a constant rate, and that their hazard rate falls as a task goes on. The simple model still shows why length hurts:
| Task length vs the agent's 50% horizon | Success rate if the failure rate were constant |
|---|---|
| Half the horizon | About 71% |
| Equal to the horizon | 50% |
| Twice the horizon | 25% |
| Four times the horizon | About 6% |
METR's own August 2025 update found Claude 3.7 Sonnet passed the tests on 38% of 18 real open-source tasks, yet none of its pull requests were mergeable as-is. Passing a check is not the same as finishing the job. That gap is also why evaluation noise matters when you compare two agents.
Long-Horizon Agents vs Short-Horizon Assistants
| Question | Short-horizon assistant | Long-horizon agent |
|---|---|---|
| Typical task | One answer, one edit, one lookup | A feature, a migration, a research report |
| Where the plan lives | In the prompt, if anywhere | In a written task list outside the model |
| What survives between sessions | Nothing, or a chat log | Progress file, task list, version history |
| Main failure mode | A wrong answer you can see | Drift, lost context, and early "done" claims |
| How you verify | Read the reply | Test each task before marking it complete |
| Who checks the work | You | A separate evaluator agent, then a person |
| What limits it | Model knowledge | Time horizon, context window, and the harness |
The short version: a long horizon is a property of the whole system, not the model alone. A strong model with no written plan behaves like a short-horizon agent on a long job.
Connection to Taskade
A long-horizon agent is only as good as its task list, and a task list is what Taskade has always been. Taskade projects hold goals, checklists, and notes that an agent can read before it acts and that your team can read too. That is the durable external plan the research keeps arriving at.
Taskade EVE, the agent that builds Taskade Genesis apps, works this way. Each app workspace keeps one running task list, a project named TASKS.md, and Taskade EVE keeps it current so it knows what is done and what remains. Its notes live as ordinary, readable projects in a projects/memory folder that you can open, edit, or delete. Memory here improves what agents know, not how they learn.
For work bigger than one agent, AI Teams have four execution modes in team chat. Auto lets Taskade pick the most suitable agent or agents. Everyone has every agent on the team respond. Manual lets you pick who replies. Orchestrate builds a plan step by step and hands each step to the best-suited agent, the planner-and-specialists pattern from the harness research. The Ask Agent Team action runs a team inside an automation in Auto, Everyone, or Orchestrate mode, so a schedule or a new task can start the work, and the team's conversation comes back as output that later steps can check. Agents use built-in tools to search and read the web and run code, pick from frontier models by top AI labs, and reach your stack through 100+ bidirectional integrations. This is Workspace DNA: Memory holds the plan, Intelligence does the next task, Execution writes the result back.
What You Would Build in Taskade
You already run long projects this way with people: a shared plan, one owner per task, a status note at the end of each day, and a reviewer before anything ships. Agents need the same discipline, written down where every session can find it.
In Taskade you would describe a long-running project board. The goal sits at the top. Each row is one task with a definition of done and a pass or fail check. An automation picks the next open row on a schedule and sends it to an AI Team in Orchestrate mode. One agent does the work. A second agent, briefed only on the definition of done, tests it. A later automation step writes the result and a short progress note back to the row. Anything that fails twice goes to a person, using the same human in the loop line you would draw for a new hire. Keep tools that need manual approval out of the team you call from an automation, because an approval step inside Ask Agent Team fails the run. When you come back on Monday, the board tells you what finished, what is blocked, and what is next.

Describe yours and build it free →
Related Concepts
- Agent Harness: the scaffolding that turns a model into a long-running agent
- Context Compaction: summarizing a long session so work can continue, and what it loses
- Agent Session: the unit of work a long job is cut into
- Context Rot: why quality drops as a context window fills
- Agent Handoff: passing state cleanly from one agent or session to the next
- Autonomous Agents: agents that plan and act toward a goal with little supervision
- Jagged Intelligence: why a long horizon on one task says little about the next
- Agent Harness Explained and The History of AI Benchmarks: the longer reads
Frequently Asked Questions About Long-Horizon Agents
What is a long-horizon agent?
A long-horizon agent is an AI agent that can complete tasks that take a skilled person hours or days. It works across many steps and many sessions, keeping its plan and progress in files or task lists outside the model so each new session can pick up where the last one stopped.
What is the 50% time horizon in AI?
It is the length of task, measured in how long a human professional takes, that an AI agent completes successfully half the time. METR introduced it in March 2025. A model with a 2-hour horizon finishes about half of the tasks that take an expert 2 hours.
How fast is AI agent time horizon growing?
METR measured a doubling about every 7 months from 2019 to 2025, with a faster pace since 2024. By May 2026 its frontier measurement was near 12 hours. Its Time Horizon 1.1 update in January 2026 put the post-2023 doubling time at about 131 days. METR itself cautions that the error bars are wide and that the tasks are self-contained software and research work.
Does a 2-hour time horizon mean an agent can work for 2 hours alone?
No. METR says time horizon measures how much human labor an agent can replace at 50% success, not how long it can run. An agent can run for many hours on a task with a short horizon, or finish a long-horizon task quickly.
Why do AI agents fail on long tasks?
Mostly because each session starts with no memory of the last, and small errors compound. Anthropic saw agents try to do everything at once, declare victory early, and skip testing. Toby Ord's 2025 analysis modeled each extra minute of work as a roughly constant chance of failure, so length punishes agents that cannot recover from mistakes. Later analysis suggests the failure rate falls as a task goes on, but long tasks still fail more often than short ones.
How do you build a reliable long-horizon agent?
Write the plan outside the model. Break the goal into a task list, work one task per session, test each task before marking it done, commit progress with a note, and let the next session read those files first. Use a separate agent or a person to check the work.
How does task management help AI agents?
A task list is the external memory a long-horizon agent needs. It says what is done, what is next, and what "done" means. In Taskade, projects and checklists hold that plan, Taskade EVE keeps a running TASKS.md list in each app workspace, and AI Teams in Orchestrate mode hand each step to the best-suited agent.