METR time horizon doubles every 4 months

Last updated
METR time horizons doubled every 7 months from 2019-2024 and every 4 months from 2024-onward (Appendix G, page 51). The acceleration from 7-month to 4-month doubling is a key claim.

METR’s published 2024-onward fit doubles in 105 days; its longer-window fit is about 129 days. These support the approximate four-month pace, subject to fit and task-suite limitations.

At a glance

  • Assessment: On Track
  • Confidence in assessment: 85%
  • Outcome: unresolved
  • Timing: pending
  • Evidence: direct
  • Predicted timing: ~4 month doubling from 2024+
  • Primary source: ai-2027.com, Appendix G (page 51)

What AI 2027 Predicted

The scenario’s supporting argument uses an approximately four-month doubling of AI task horizons from 2024 onward. This is a trend claim about task difficulty at a specified success rate. AI 2027

How We Track This

Match the fit’s starting date, model selection and success threshold. Human-task-hours measure how long a person would need for a task; they do not measure uninterrupted autonomous runtime.

Current Evidence

METR’s May 19 report explicitly publishes a fit to frontier public models released after January 1, 2024, using 50% success horizons: a 105-day doubling time, R² = 0.98. The matching-window result exists and supports the approximate scenario pace. METR report, Appendix E

METR’s TH1.1 data separately gives a 2023-onward fit of about 129 days, with an interval of approximately 104-158 days, and an all-history fit of about 188 days. These are different windows, rather than contradictory measurements of an identical trend. The longer-window interval is not an uncertainty interval for the 105-day estimate. METR raw results

Counterevidence & Limitations

The report warns that the suite cannot reliably measure horizons above roughly 16 human-task-hours. Frontier measurements approach this ceiling. A high fitted R² does not eliminate extrapolation error, benchmark selection effects or uncertainty about future doubling rates. METR methodology and dashboard

On-track acknowledges the matching-window evidence. We do not infer a sustained acceleration beyond the approximate four-month prediction from one faster point estimate.

What Would Change Our Assessment

  • Strengthen: Extended, unsaturated tasks reproduce the pace with comparable success thresholds and transparent fits.
  • Weaken: Comparable new data shows a sustained slowdown beyond uncertainty.
  • Reassess acceleration: Multiple robust windows establish a materially faster trajectory rather than a single fitted difference.

Update History

DateUpdate
2026-09-07Recovered METR’s directly published 2024-onward 105-day fit. The matching-window evidence supports the approximate pace; longer-window estimates and saturation limits remain explicit.
2026-09-06Assessment revised from ahead (0.80) to on-track (0.75). The published 2023-onward fit is close to four-month doubling with uncertainty; it does not establish a materially faster trend or exactly reproduce a 2024-onward fit.
2026-05-25METR’s Frontier Risk Report reported that the strongest Feb-Mar 2026 assessed agents were saturating Time Horizon 1.1, with a most-capable shared model point estimate around 16-20 hours at the 50% horizon and 3-4 hours at the 80% horizon. This reinforces the ahead assessment, but METR cautions that measurements above 16 hours are unreliable with the current task suite.
2026-04-13METR time horizon benchmark page updated Feb 4, 2026 with new model measurements. Academic literature (GovAI, March 2026) notes significant uncertainty in translating benchmark performance to real-world R&D productivity gains. The doubling trend continues broadly on pace with the ~4-month rate. No dramatic acceleration or deceleration evident. No status or confidence change. Sources: METR, METR simpler timelines model, arXiv:2603.03992
2026-04-02AI Futures Project Q1 2026 update revised METR doubling time from 5.5 months to 4 months (Kokotajlo) and 4.5 months (Lifland), citing METR v1.1 trend and new model evaluations (Gemini 3, GPT-5.2, Opus 4.6). The authors’ own assessment now aligns closely with our “Ahead” status. Source: LessWrong
2026-03-30Ajeya Cotra (METR researcher, former Open Philanthropy) published analysis (March 3) noting that Opus 4.6 at ~12 hours was already past her January forecast for year-end 2026, and projecting 100+ hour time horizon by year-end at current pace (planned-obsolescence.org). She notes the benchmark suite is “nearly saturated” for short tasks, creating noise in longer-horizon estimates. Wide CI (5.3 to 66 hours), but 19 tasks estimated at 8+ hours — Opus 4.6 solved 14 of them at least once. GPT-5.4 released March 5 with Thinking and Pro variants; METR measurement pending. The 4-month doubling rate appears to be accelerating further, not stabilizing.
2026-03-23METR’s Opus 4.6 measurement initially appeared around 14.5 hours on the 50% horizon, but current TH1.1 raw data estimates roughly 11h59m at 50% and roughly 1h10m at 80%. METR also published study (Mar 10) finding roughly half of SWE-bench-passing PRs would not be merged by maintainers, suggesting benchmark scores overstate real-world agent capability (METR). Doubling pace remains ahead of prediction. No status change.
2026-03Actual doubling pace (~3 months) is faster than the predicted 4 months. AI agent capabilities advancing ahead of the scenario’s timeline.
2026-01METR released Time Horizon 1.1: expanded from 170 to 228 tasks, long tasks (8h+) doubled from 14 to 31. Key finding: in 2024–2025, coding time horizons doubled every approximately 4 months — matching AI 2027’s predicted acceleration. Status upgraded to on-track.
2025-12AI Futures Project Dec 2025 model update assessed METR coding time horizon as tracking at approximately 1.04× their central AI-2027-speed trajectory. The historical 7-month doubling rate confirmed across domains showed signs of acceleration in the most recent period.
2025-11GPT-5.1-Codex-Max METR evaluation (November 19) shows 50% time horizon of approximately 2 hours 42 minutes — up from 2h17m in August. Share of success on hardest AI R&D-relevant tasks jumps from 2% to 8% (4x improvement). Doubling trend continues.
2025-08GPT-5 launches with METR pre-deployment evaluation disclosing a 50% time horizon of approximately 2 hours 17 minutes on software engineering tasks (August 7). METR’s August 12 research update confirms the ~7-month doubling trend across domains for the 2019-2024 historical period, while noting recent acceleration toward 4-month doubling.
2025-07METR publishes domain time-horizon analysis (July 14): approximately 7-month doubling time across software engineering, ML, and cybersecurity domains. Frontier models (Claude 3.7) at ~50-minute 50% horizon in early 2025. The 7-month historical average is slower than AI 2027’s predicted 4-month acceleration, but the extrapolation still arrives at approximately early 2027 — consistent with the essay’s timeline.