Glossary · term

AI Agent Task-Completion Time Horizon

AI agent task-completion time horizon is a human-calibrated capability metric. For a specified task distribution and success probability, it is the human-expert task duration at which a model-and-scaffold agent's fitted probability of success reaches that threshold. The common 50% horizon is therefore a partial-reliability difficulty point, not the agent's elapsed run time, context-window length, or a guarantee of safe unattended operation.

Safety2025-03-18Wave 2 · 2024Maturity: 3/5

Origin and context

METR introduced the metric in March 2025 using 170 tasks drawn from HCAST, RE-Bench, and Software Atomic Actions; the paper later appeared at NeurIPS 2025. TH1.1, released in January 2026, expanded the suite to 228 tasks and moved evaluation infrastructure to Inspect. METR publishes a live measurement page plus analysis code and run data. Independent work has since reused both the name and the human-time logistic-fit construction.

Sources: s1, s2, s3, s4, s5, s7

Why it matters

The metric converts benchmark success into a human-readable scale and makes longitudinal capability comparisons easier. The original work reported roughly seven-month doubling over 2019–2025. TH1.1 retained an approximately 196-day full-history hybrid fit but estimated about 131 days after 2023, showing that the selected period and suite version matter. The UK Government Office for Science treats the data as a useful signal for structured, verifiable software work while explicitly rejecting a direct inference to messy cognitive work or real-world adoption.

Sources: s1, s4, s6

Example

Suppose an agent has a two-hour 50% horizon on a named suite. This means the fitted curve crosses 50% for tasks whose qualified-human baseline is two hours; it does not mean the agent runs for two hours or succeeds on every shorter task. A result should state the model, scaffold, suite version, probability threshold, point estimate, and interval. Token and wall-clock limits are evaluation settings, while the reported duration remains a human reference value.

Sources: s2, s3, s5

How it differs

RE-Bench (Research Engineering Benchmark)

RE-Bench is one named seven-environment research-engineering benchmark with continuous scores. Task-completion time horizon is a fitted aggregate metric based on binary success and human duration across a task distribution; RE-Bench tasks can contribute evidence without being the metric itself.

Long-context language models

Long context describes how much input or working history a system can accept, usually in tokens. It may influence agent performance, but it does not report success probability as a function of human task duration and is not interchangeable with a time horizon.

Maturity and evidence

Maturity is rated 3. The metric has a peer-reviewed NeurIPS paper, a maintained TH1.1 dashboard, public analysis artifacts, and exact-name adoption in independent research and UK government foresight. It remains below 4 because neither the task distribution nor protocol is standardized across organizations, the suite is still being revised as it saturates, and independent use does not establish broad cross-domain validity.

Sources: s2, s3, s5, s6, s7

Limits and open questions

Current METR tasks mainly cover self-contained software engineering, machine-learning, and cybersecurity work, often with low prior context and algorithmic grading. Human baselines, scaffold elicitation, task selection, curve form, and regularization affect estimates; METR warns that current measurements above 16 hours are unreliable. Confidence intervals can span roughly a factor of two, and 50% reliability is inadequate for many deployments. Trend fits must therefore remain versioned empirical summaries, not predictions of job automation, continuous autonomy, or safety.

Sources: s2, s6, s8, s9

Related terms

References

Last updated: 2026-09-05