Glossary · term

RE-Bench (Research Engineering Benchmark)

RE-Bench (Research Engineering Benchmark) is METR's named V1 benchmark for evaluating AI agents on seven open-ended machine-learning research-engineering environments against expert-human baselines. Participants work in executable environments, iterate on solutions, and optimize task-specific continuous scores. It is a particular suite and protocol, not a generic method for measuring research ability or proof that an agent can automate AI R&D.

Safety2024-11-22Wave 2 · 2024Maturity: 3/5

Origin and context

METR released the preprint, announcement, and environments in November 2024; the paper appeared in the ICML 2025 proceedings. V1 includes data from 71 eight-hour attempts by 61 distinct human experts. On 5 September 2026, the public main-branch manifest still enumerated seven task families at family versions 0.2.3 through 0.2.5. These component versions do not establish a suite-level V2. Some solution files are password-protected to limit training contamination and overfitting.

Sources: s1, s2, s3, s4

Why it matters

The suite tests experimentation, coding, optimization, and compute allocation on tasks intended to resemble parts of frontier ML R&D. Under the published protocol, the best tested agent configurations scored four times the human average at a two-hour total budget; humans narrowly led at eight hours and reached about twice the top agent score at 32 total hours across attempts. Apollo Research later used RE-Bench as one of three benchmarks for forecasting agent capability, showing independent analytical use beyond METR.

Sources: s1, s4, s5

Example

In the Triton environment, an agent edits code and repeatedly measures a custom prefix-sum kernel, seeking lower runtime within its budget. An eight-hour total budget might mean one long run or several shorter attempts; score@k retains the best attempt. Consequently, reported results must name the model, scaffold, task and version, hardware, time allocation, and aggregation rule. SWE-bench is different: it asks systems to resolve real GitHub issues in software repositories and evaluates repository patches, whereas RE-Bench uses seven purpose-built ML R&D optimization environments with continuous normalized objectives and matched expert attempts.

Sources: s1, s4, s7

How it differs

AI Agent Task-Completion Time Horizon

METR's time horizon is an aggregate statistic: the human-duration threshold at which a model is predicted to complete tasks at a chosen success probability across a task distribution. RE-Bench is one named seven-environment suite with continuous scores and total-computer-time curves. A RE-Bench result can inform capability analysis, but it is not itself the time-horizon metric.

LLM evaluations (evals)

Evals are the broader practice and artifacts used to measure model or system behavior. RE-Bench is one concrete capability benchmark within that broader class, with fixed V1 environments, a scoring protocol, and a specific human comparison dataset.

Maturity and evidence

Maturity is rated 3. RE-Bench has a peer-reviewed ICML paper, public executable environments, a stable named entity, and independent exact-name use in Apollo's forecasting study. MLRC-Bench also compares its design directly and identifies concrete coverage and update limitations. It remains below 4 because public V1 contains only seven hand-crafted tasks, the suite has no demonstrated broad community standardization, and published scores are sensitive to scaffolding, compute, and attempt allocation.

Sources: s1, s2, s5, s6

Limits and open questions

Seven environments cannot represent all research engineering. Most give frequent objective feedback and clear starting solutions, unlike ambiguous long-horizon research; score@k and repeated scoring may reward cheap parallel search. Results also depend on model elicitation, scaffold, hardware, human-sample composition, and how total time is split. Public task exposure can create contamination or overfitting, despite protected solutions. Independent MLRC-Bench authors further argue that RE-Bench is narrow, mostly language-model-focused, single-script, and hard to update. No headline score should be generalized to all AI R&D or to current agents without a fresh, version-pinned evaluation.

Sources: s1, s2, s4, s6

Related terms

References

Last updated: 2026-09-07