RE-Bench score reaches 1.3
1.3 on RE-Bench matching top expert humans given 8 hours at well-defined AI research engineering tasks.
At a glance
- Assessment: Behind
- Confidence in assessment: 70%
- Timing: Original window passed; outcome unresolved
- Predicted timing: Early 2026
- Primary source: AI 2027, page 7, footnote 15 (Early 2026: Coding Automation)
What AI 2027 Predicted
The scenario predicted that by early 2026, AI agents would score 1.3 on METR’s RE-Bench (Research Engineering Benchmark) — a suite of 7 challenging, open-ended ML research engineering tasks. A score of 1.3 means surpassing what top human experts achieve when given a full 8-hour session on well-defined research engineering tasks like implementing ML algorithms, debugging research code, and optimizing training pipelines. This prediction sits in the “Coding Automation” section, representing the point where AI begins to meaningfully accelerate AI R&D itself.
How We Track This
- METR’s official RE-Bench results and publications at metr.org/research
- Model-specific evaluations published by METR
- METR time horizon measurements (related metric, drawn partly from RE-Bench tasks)
- AI lab announcements referencing RE-Bench or research engineering capability
Current Evidence
No qualifying public result at or above the specified 1.3 threshold was identified in this review.
- Early RE-Bench comparisons depend on AI and human time budgets. Short-budget results do not establish the specified eight-hour reference score.
- Results for Claude 3.5 Sonnet and o1-preview are historical observations, not a current frontier estimate.
- Early model-specific evaluations supply historical context, not a current ceiling on the benchmark.
- Forecasts are not observations of whether a threshold has been achieved.
- The related METR time horizon metric shows strong progress: current TH1.1 raw data estimates Claude Opus 4.6 at roughly 11h59m on the 50% horizon and roughly 1h10m on the 80% horizon. This indicates rapid capability growth, but RE-Bench measures a specific type of research engineering skill, not general software task completion
The original early-2026 target remains unresolved in the reviewed comparable evidence. Related time-horizon gains cannot substitute for the specified RE-Bench result.
Counterevidence & Limitations
- RE-Bench V1 consists of only 7 tasks, making it a noisy benchmark — small changes in model capability can produce large score swings
- The benchmark may be nearing the limits of its discriminative power as models improve; METR is developing updated evaluation approaches
- The 8-hour human expert reference was established with 61 human experts across 71 attempts; this baseline may shift as the expert pool is refined
- Some research engineering capabilities may not be well-captured by RE-Bench’s specific tasks (e.g., novel architecture design, long-horizon experiment planning)
- Strong time-horizon results show progress on a related distribution without resolving the RE-Bench threshold.
- The 50% and 80% METR horizons imply different levels of reliability; the 50% horizon should not be read as reliable 8-hour autonomous research-engineering performance
What Would Change Our Assessment
- Strengthen the assessment: A dated result identifies task version, normalization, human reference and agent budget.
- Confirm: A comparable result reaches 1.3 or higher; report whether it arrived after the original window.
- Retain an evidence gap: Successor results require a validated mapping before resolving the original threshold.
Update History
| Date | Update |
|---|---|
| 2026-09-06 | Current evidence and assessment explanation reviewed. Clarified measurement scope, source interpretation and limitations. |
| 2026-06-06 | Corrected the Opus 4.6 METR time-horizon reference to current TH1.1 raw data: roughly 11h59m at 50% reliability and roughly 1h10m at 80% reliability. This weakens earlier wording that implied the 8-hour RE-Bench reference had been approximately vindicated by time-horizon data. Status changed from on-track to behind and confidence raised from 0.60 to 0.70 because no published 1.3+ RE-Bench score has been verified. |
| 2026-03 | The improvement from Opus 4.5 (~4h49m) to Opus 4.6 on METR time horizons suggested the 1.3 score might come within reach. Key uncertainty: whether time-horizon metrics and RE-Bench V1 numeric scores map cleanly onto each other. Later correction: the current Opus 4.6 50% horizon is roughly 11h59m, while the 80% horizon is much shorter. |
| 2026-02 | Claude Opus 4.6 initially appeared to surpass the 8-hour human-expert reference level on the 50% time-horizon metric, suggesting possible progress toward RE-Bench-style research engineering. Later correction: this did not establish reliable 8-hour autonomous performance, and no published 1.3+ RE-Bench score was available. METR announces redesign of developer productivity study (Feb 24), noting developers now refuse to work without AI, invalidating earlier “19% slower” finding. Status revised from Behind to On Track. Confidence 0.60. |
| 2026-01 | METR publishes Time Horizon 1.1 (Jan 29). Claude Opus 4.5 measured at ~4 hours 49 minutes — a substantial jump suggesting rapid progress toward the 8-hour RE-Bench reference level. External forecasters had predicted RE-Bench 1.0 in 2027; this pace suggests that was too conservative. |
| 2025-12 | METR begins publishing time-horizon metric — measuring how long a task an AI can autonomously complete at human-expert level. Reframes the RE-Bench question from a single score to “how many hours of autonomous work can an AI sustain?“ |
| 2025-09 | METR evaluates Claude 3.7 Sonnet on RE-Bench (preliminary). Results described as “impressive” on a subset of tasks. Exact numeric score not widely published, but qualitative signal is positive. |
| 2025-07 | METR publishes study on early-2025 AI developer productivity (July 10). Finding: experienced developers were 19% slower with AI tools in a controlled trial — contradicting self-reported 20% speedups. Raises questions about AI’s ability to accelerate sustained research work. |
| 2025-05 | RE-Bench V1 baseline established (published late 2024 by METR). Claude 3.5 Sonnet and o1-preview score ~0.5-0.8 at 8-hour human-expert reference. Models outperform humans ~4x at 2-hour budget but plateau or regress on longer tasks. Target of 1.3 appears distant. |