RE-Bench (Research Engineering Benchmark)
RE-Bench (Research Engineering Benchmark) is METR's named V1 benchmark for evaluating AI agents on seven open-ended machine-learning research-engineering environments against expert-human baselines. Participants work in executable environments, iterate on solutions, and optimize task-specific continuous scores. It is a particular suite and protocol, not a generic method for measuring research ability or proof that an agent can automate AI R&D.
Origin and context
METR released the preprint, announcement, and environments in November 2024; the paper appeared in the ICML 2025 proceedings. V1 includes data from 71 eight-hour attempts by 61 distinct human experts. On 5 September 2026, the public main-branch manifest still enumerated seven task families at family versions 0.2.3 through 0.2.5. These component versions do not establish a suite-level V2. Some solution files are password-protected to limit training contamination and overfitting.
Why it matters
The suite tests experimentation, coding, optimization, and compute allocation on tasks intended to resemble parts of frontier ML R&D. Under the published protocol, the best tested agent configurations scored four times the human average at a two-hour total budget; humans narrowly led at eight hours and reached about twice the top agent score at 32 total hours across attempts. Apollo Research later used RE-Bench as one of three benchmarks for forecasting agent capability, showing independent analytical use beyond METR.
Example
In the Triton environment, an agent edits code and repeatedly measures a custom prefix-sum kernel, seeking lower runtime within its budget. An eight-hour total budget might mean one long run or several shorter attempts; score@k retains the best attempt. Consequently, reported results must name the model, scaffold, task and version, hardware, time allocation, and aggregation rule. SWE-bench is different: it asks systems to resolve real GitHub issues in software repositories and evaluates repository patches, whereas RE-Bench uses seven purpose-built ML R&D optimization environments with continuous normalized objectives and matched expert attempts.
How it differs
AI Agent Task-Completion Time Horizon
METR's time horizon is an aggregate statistic: the human-duration threshold at which a model is predicted to complete tasks at a chosen success probability across a task distribution. RE-Bench is one named seven-environment suite with continuous scores and total-computer-time curves. A RE-Bench result can inform capability analysis, but it is not itself the time-horizon metric.
LLM evaluations (evals)
Evals are the broader practice and artifacts used to measure model or system behavior. RE-Bench is one concrete capability benchmark within that broader class, with fixed V1 environments, a scoring protocol, and a specific human comparison dataset.
Maturity and evidence
Maturity is rated 3. RE-Bench has a peer-reviewed ICML paper, public executable environments, a stable named entity, and independent exact-name use in Apollo's forecasting study. MLRC-Bench also compares its design directly and identifies concrete coverage and update limitations. It remains below 4 because public V1 contains only seven hand-crafted tasks, the suite has no demonstrated broad community standardization, and published scores are sensitive to scaffolding, compute, and attempt allocation.
Limits and open questions
Seven environments cannot represent all research engineering. Most give frequent objective feedback and clear starting solutions, unlike ambiguous long-horizon research; score@k and repeated scoring may reward cheap parallel search. Results also depend on model elicitation, scaffold, hardware, human-sample composition, and how total time is split. Public task exposure can create contamination or overfitting, despite protected solutions. Independent MLRC-Bench authors further argue that RE-Bench is narrow, mostly language-model-focused, single-script, and hard to update. No headline score should be generalized to all AI R&D or to current agents without a fresh, version-pinned evaluation.
Related terms
References
- RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human ExpertsProceedings of Machine Learning Research / ICML 2025 · 2025-07-13 · class A
- METR/RE-BenchMETR · 2024-11 · class A
- RE-Bench suite manifestMETR · 2024-11 · class A
- Evaluating frontier AI R&D capabilities of language model agents against human expertsMETR · 2024-11-22 · class A
- Forecasting Frontier Language Model Agent CapabilitiesApollo Research / arXiv · 2025-02-21 · class A
- MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?NeurIPS 2025 Datasets and Benchmarks Track · 2025 · class A
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Princeton NLP / arXiv · 2023-10-10 · class A
- Measuring AI Ability to Complete Long Software TasksMETR · 2025-03-19 · class A
Last updated: 2026-09-07