AI Agent Task-Completion Time Horizon
AI agent task-completion time horizon is a human-calibrated capability metric. For a specified task distribution and success probability, it is the human-expert task duration at which a model-and-scaffold agent's fitted probability of success reaches that threshold. The common 50% horizon is therefore a partial-reliability difficulty point, not the agent's elapsed run time, context-window length, or a guarantee of safe unattended operation.
Origin and context
METR introduced the metric in March 2025 using 170 tasks drawn from HCAST, RE-Bench, and Software Atomic Actions; the paper later appeared at NeurIPS 2025. TH1.1, released in January 2026, expanded the suite to 228 tasks and moved evaluation infrastructure to Inspect. METR publishes a live measurement page plus analysis code and run data. Independent work has since reused both the name and the human-time logistic-fit construction.
Why it matters
The metric converts benchmark success into a human-readable scale and makes longitudinal capability comparisons easier. The original work reported roughly seven-month doubling over 2019–2025. TH1.1 retained an approximately 196-day full-history hybrid fit but estimated about 131 days after 2023, showing that the selected period and suite version matter. The UK Government Office for Science treats the data as a useful signal for structured, verifiable software work while explicitly rejecting a direct inference to messy cognitive work or real-world adoption.
Example
Suppose an agent has a two-hour 50% horizon on a named suite. This means the fitted curve crosses 50% for tasks whose qualified-human baseline is two hours; it does not mean the agent runs for two hours or succeeds on every shorter task. A result should state the model, scaffold, suite version, probability threshold, point estimate, and interval. Token and wall-clock limits are evaluation settings, while the reported duration remains a human reference value.
How it differs
RE-Bench (Research Engineering Benchmark)
RE-Bench is one named seven-environment research-engineering benchmark with continuous scores. Task-completion time horizon is a fitted aggregate metric based on binary success and human duration across a task distribution; RE-Bench tasks can contribute evidence without being the metric itself.
Long-context language models
Long context describes how much input or working history a system can accept, usually in tokens. It may influence agent performance, but it does not report success probability as a function of human task duration and is not interchangeable with a time horizon.
Maturity and evidence
Maturity is rated 3. The metric has a peer-reviewed NeurIPS paper, a maintained TH1.1 dashboard, public analysis artifacts, and exact-name adoption in independent research and UK government foresight. It remains below 4 because neither the task distribution nor protocol is standardized across organizations, the suite is still being revised as it saturates, and independent use does not establish broad cross-domain validity.
Limits and open questions
Current METR tasks mainly cover self-contained software engineering, machine-learning, and cybersecurity work, often with low prior context and algorithmic grading. Human baselines, scaffold elicitation, task selection, curve form, and regularization affect estimates; METR warns that current measurements above 16 hours are unreliable. Confidence intervals can span roughly a factor of two, and 50% reliability is inadequate for many deployments. Trend fits must therefore remain versioned empirical summaries, not predictions of job automation, continuous autonomy, or safety.
Related terms
References
- Measuring AI Ability to Complete Long Software TasksMETR / arXiv; published at NeurIPS 2025 · 2025-03-18 · class A
- Measuring AI Ability to Complete Long Software TasksNeurIPS 2025 · 2025 · class A
- Task-Completion Time Horizons of Frontier AI ModelsMETR · 2026-02-06 · class A
- Time Horizon 1.1METR · 2026-01-29 · class A
- METR Time Horizon AnalysisMETR · 2025-03 · class A
- AI Scenarios 2030: Helping policymakers plan for the future of AIUK Government Office for Science · 2026-06-15 · class B
- Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI ModelsRedwood Research / arXiv · 2026-06-05 · class B
- Clarifying limitations of time horizonMETR · 2026-01-22 · class A
- Impact of modelling assumptions on time horizon resultsMETR · 2026-03-20 · class A
Last updated: 2026-09-05