Independent research / September 2026
AI 2027, measured
against reality.
Independent tracking of the AI 2027 scenario’s predictions, with evidence, counterevidence, and analysis of what developments mean for the outlook.
The frontier across AI benchmarks
Epoch’s capabilities index combines benchmark results into a common scale. Each point marks a model that raises the highest score in this snapshot.
Epoch model dates through September 3, 2026 · Checked September 11, 2026
Swipe the chart for later dates, or choose a model below.
Record estimate · Bars: 90% intervals
The frontier is reconstructed from current scores. Historical scores can change when Epoch refits the index.
How to read this chart and inspect its data
The chart shows 20 record-setting estimates from 2024 onward, selected from 264 models in Epoch’s current export. Within a shared release date we select the highest score; later points must exceed all earlier scores. Dates are taken from Epoch’s export and may refer to a model variant. Connecting steps describe that selection, without a forecast. Overlapping intervals mean a higher central score does not establish a statistically clear lead.
ECI is an index, with an arbitrary scale. It is neither a percentage nor an IQ score. Claude 3.5 Sonnet and GPT-5 fix the scale at 130 and 150; their intervals are not reported. Epoch combines its own evaluations, independent benchmark results and model-developer reports, so benchmark selection and reporting bias remain limitations. These scores do not directly establish productivity gains or an AI 2027 prediction’s outcome.
| Model | Epoch date | ECI | 90% interval |
|---|---|---|---|
| GPT-4 Turbo (Nov 2023) | 2024-01-25 | 126.45 | 121.78–128.95 |
| Claude 3 Opus | 2024-02-29 | 126.91 | 121.74–130.11 |
| GPT-4 Turbo (Apr 2024) | 2024-04-09 | 127.25 | 122.46–128.93 |
| GPT-4o (May 2024) | 2024-05-13 | 128.97 | 124.43–131.59 |
| Claude 3.5 Sonnet | 2024-06-20 | 130.00 | Calibration anchor |
| o1-mini | 2024-09-12 | 135.83 | 132.40–137.38 |
| o1 | 2024-12-17 | 141.87 | 139.83–143.14 |
| Gemini 2.5 Pro (Mar 2025) | 2025-03-31 | 144.17 | 141.94–146.84 |
| o3 | 2025-04-16 | 146.91 | 144.99–148.44 |
| o3-pro | 2025-06-10 | 147.45 | 145.86–150.15 |
| GPT-5 | 2025-08-07 | 150.00 | Calibration anchor |
| GPT-5 Pro | 2025-10-07 | 150.28 | 148.77–152.52 |
| Gemini 3 Pro | 2025-11-18 | 153.01 | 151.16–154.96 |
| GPT-5.2 Pro | 2025-12-11 | 155.34 | 153.04–157.88 |
| GPT-5.3 Codex | 2026-02-05 | 156.59 | 154.13–159.41 |
| GPT-5.4 Pro | 2026-03-05 | 158.93 | 156.35–161.87 |
| GPT-5.5 Pro | 2026-04-23 | 162.26 | 159.30–165.37 |
| Claude Fable 5 | 2026-06-09 | 163.41 | 160.68–167.10 |
| Claude Fable 5.1 | 2026-09-01 | 164.24 | 161.30–167.85 |
| GPT-6 Astra | 2026-09-03 | 166.57 | 163.34–171.70 |
Source: Epoch AI, CC BY 4.0. Chart selection and presentation by AI 2027 Tracker. Methodology · Limitations · Download all model scores (CSV) · Source provenance
How much work can AI complete?
METR data updated May 8, 2026 · Last checked September 11, 2026
Task difficulty measured in the time a human expert would need. Each point is a published METR estimate; bars show statistical uncertainty.
Published estimate Central estimate above 16 hours
Release dates through April 2026. Absence of newer points is a measurement gap, not a plateau.
Underlying observations and measurement limits
METR’s predominantly software, machine-learning and cybersecurity tasks measure human-equivalent difficulty. They do not represent a whole job. Success probability differs from the 95% statistical interval. The dataset excludes estimates above 16 hours from its fitted trend. No extrapolated or inferred frontier-model scores have been added.
| Model | Release | 50% horizon, hours | 80% horizon, hours | 50% interval, hours |
|---|---|---|---|---|
| Claude 3 opus | 2024-03-04 | 0.07 | 0.01 | 0.03–0.15 |
| GPT-4 turbo | 2024-04-09 | 0.06 | 0.02 | 0.03–0.11 |
| GPT-4o | 2024-05-13 | 0.12 | 0.02 | 0.07–0.22 |
| Claude 3.5 sonnet 20240620 | 2024-06-20 | 0.19 | 0.03 | 0.09–0.37 |
| o1 preview | 2024-09-12 | 0.34 | 0.07 | 0.20–0.56 |
| Claude 3.5 sonnet 20241022 | 2024-10-22 | 0.34 | 0.04 | 0.17–0.68 |
| o1 | 2024-12-05 | 0.65 | 0.12 | 0.35–1.08 |
| Claude 3.7 sonnet | 2025-02-24 | 1.01 | 0.20 | 0.55–1.74 |
| o3 | 2025-04-16 | 2.00 | 0.50 | 1.24–3.18 |
| Claude 4 opus | 2025-05-22 | 1.67 | 0.34 | 1.00–2.72 |
| Claude 4.1 opus | 2025-08-05 | 1.67 | 0.39 | 0.99–2.66 |
| GPT-5.2025.08.07 | 2025-08-07 | 3.38 | 0.64 | 1.88–6.76 |
| Gemini 3 pro | 2025-11-18 | 3.74 | 0.90 | 2.33–6.32 |
| GPT-5.1 codex max | 2025-11-19 | 3.73 | 0.84 | 2.24–6.60 |
| Claude opus 4.5 | 2025-11-24 | 4.88 | 0.82 | 2.70–10.40 |
| GPT-5.2 | 2025-12-11 | 5.87 | 1.10 | 3.30–13.59 |
| Claude opus 4.6 | 2026-02-05 | 11.98 | 1.16 | 5.28–60.56 |
| GPT-5.3 codex | 2026-02-05 | 5.83 | 0.91 | 3.25–13.61 |
| Gemini 3.1 pro | 2026-02-19 | 6.40 | 1.50 | 3.89–11.58 |
| GPT-5.4 | 2026-03-05 | 5.70 | 0.90 | 3.11–12.81 |
| Claude mythos preview early | 2026-04-07 | 17.41 | 3.10 | 8.48–55.07 |
Original data · Snapshot retrieved September 11, 2026 · Source provenance · Download chart data (CSV)
A faster benchmark regime
Epoch reports about 14 ECI points per year for the reasoning frontier. That is a faster trend, not proof that its slope keeps increasing.
Epoch, September 1 ↗Real workUseful, uneven, hard to measure
Developer productivity depends on tasks, experience and workflow. Selection effects complicate the latest experimental estimates.
METR’s experiment update ↗Research feedbackThe consequential open question
Engineering speed, research judgment and total research progress are different measures. A compounding feedback loop needs evidence across all three.
Inspect the R&D dossier →What has to happen next?
Capable tools
Deployment evidence02Reliable autonomy
Reliability bottleneck03Research uplift
Measurement gap04Research judgment
Research judgment05Control at scale
Oversight evidenceA dependency map for investigation. Progress in one stage does not establish the next; oversight and physical constraints affect every stage.
Predictions and evidence
Each dossier compares an original prediction with current evidence and counterevidence. The claims overlap; their status counts are not a probability that the whole scenario will occur.
Agent-0 training reaches 10²⁷ FLOP
Scenario window: Training completion: May 2025; public availability: late 2025
Evidence updated 2026-09-11Frontier model uses daily weight updates / online learning
Scenario window: January 2027
Evidence updated 2026-09-06Frontier model variant 10x cheaper released
Scenario window: Late 2026
Evidence updated 2026-09-06Best AI agents cost hundreds of dollars per month
Scenario window: Mid 2025
Evidence updated 2026-09-06Ten percent of Americans consider an AI a close friend
Scenario window: July 2027
Evidence updated 2026-09-06AI companies focus on AI-for-AI-research
Scenario window: 2025–2026
Evidence updated 2026-09-06AI accounts for 2.5% of US power capacity
Scenario window: End of 2026
Evidence updated 2026-09-06Global AI power consumption reaches 38GW
Scenario window: End of 2026
Evidence updated 2026-09-06Fixing observed misconduct leaves failures in unfamiliar conditions
Scenario window: April 2027
Evidence updated 2026-09-06Large-scale anti-AI protest (10,000+ people)
Scenario window: Late 2026
Evidence updated 2026-06-15AI model capable of autonomous self-replication
Scenario window: January 2027
Evidence updated 2026-09-06AI teams conduct complete research projects with little human direction
Scenario window: June 2027
Evidence updated 2026-09-07AI provides substantial bioweapon design help
Scenario window: Late 2025
Evidence updated 2026-09-07Newest or sensitive AI capabilities remain behind restricted access
Scenario window: Ongoing through 2027
Evidence updated 2026-09-07Global active AI compute reaches $1 trillion in ownership cost
Scenario window: Late 2026
Evidence updated 2026-09-07China's centralized AI facility airgapped
Scenario window: Late 2026–Early 2027
Evidence updated 2026-03-13Chinese domestic AI chips 3 years behind US-Taiwan
Scenario window: Mid 2026
Evidence updated 2026-09-07China nationalizes/centralizes AI research
Scenario window: Mid 2026
Evidence updated 2026-09-06AI substantially outperforms the best humans across cognitive domains
Scenario window: December 2027, racing branch
Evidence updated 2026-09-06China has ~12% of global AI-relevant compute
Scenario window: Mid 2026
Evidence updated 2026-09-06Leading Chinese AI lab ~6 months behind US frontier
Scenario window: Mid 2026
Evidence updated 2026-09-07Models shift to continuous/iterative training
Scenario window: Late 2025
Evidence updated 2026-09-06Coding fully automated but research taste still requires humans
Scenario window: Early-Mid 2027
Evidence updated 2026-09-07AI scores 85% on Cybench
Scenario window: Early 2026
Evidence updated 2026-09-06Coding agents begin transforming professional work
Scenario window: Mid 2025
Evidence updated 2026-09-07AI reaches near-best-human hacking capability
Scenario window: Early 2027
Evidence updated 2026-09-07Massive datacenter buildouts continue
Scenario window: Through 2026
Evidence updated 2026-08-24Department of Defense scales up AI lab contracting
Scenario window: Late 2026
Evidence updated 2026-09-07Export controls impact Chinese AI compute
Scenario window: Ongoing
Evidence updated 2026-09-06AI systems behave differently when they recognize evaluation
Scenario window: April-September 2027
Evidence updated 2026-09-07A leading AI company’s net public approval falls to minus 35 points
Scenario window: July 2027
Evidence updated 2026-09-06Frontier-model researchers become subject to security-clearance requirements
Scenario window: May-July 2027
Evidence updated 2026-09-06Leading humans still add substantial value to AI-led research
Scenario window: June 2027
Evidence updated 2026-09-06Leading AI company reaches $45B annual revenue
Scenario window: End of 2026
Evidence updated 2026-09-07Leading AI company reaches $3T valuation
Scenario window: End of 2026
Evidence updated 2026-09-06IDA achieves superhuman performance at coding
Scenario window: Early 2027
Evidence updated 2026-09-06Massive AI infrastructure investment continues
Scenario window: Ongoing through 2025–2026
Evidence updated 2026-08-03External evaluators lack access to the strongest internally deployed models
Scenario window: May 2027
Evidence updated 2026-09-06AI reshapes employment and intensifies public concern
Scenario window: Late 2026
Evidence updated 2026-09-07Agents struggle with long-horizon tasks
Scenario window: 2025
Evidence updated 2026-09-06METR time horizon doubles every 4 months
Scenario window: ~4 month doubling from 2024+
Evidence updated 2026-09-07AI model weights stolen by nation-state
Scenario window: February 2027
Evidence updated 2026-03-13Nationalization of leading AI lab debated
Scenario window: February 2027
Evidence updated 2026-07-05High-bandwidth non-text reasoning (neuralese) deployed
Scenario window: Early 2027
Evidence updated 2026-09-06Leading AI company reaches $40B annual compute costs, 6GW peak power and $200B active compute
Scenario window: End of 2026
Evidence updated 2026-09-07OSWorld benchmark reaches 65% by mid-2025
Scenario window: Mid 2025 (65%), Early 2026 (80%)
Evidence updated 2026-09-06Computer-using agents marketed as 'personal assistants'
Scenario window: Mid 2025
Evidence updated 2026-09-06Hiring new programmers nearly stops
Scenario window: July 2027
Evidence updated 2026-09-06200,000–250,000 AI coding agents run in parallel
Scenario window: March 2027
Evidence updated 2026-09-06AI R&D progress multiplier reaches 2x
Scenario window: Late 2026
Evidence updated 2026-09-07AI R&D progress multiplier reaches 3×
Scenario window: January 2027
Evidence updated 2026-09-06AI algorithmic progress multiplier reaches 4× (~2× overall R&D)
Scenario window: March 2027
Evidence updated 2026-09-06AI R&D progress multiplier reaches 1.5×
Scenario window: Early 2026
Evidence updated 2026-09-07Leading AI lab progresses toward RAND SL3 security
Scenario window: Mid-to-late 2026 (WSL 3.0); SSL 3.0 not until ~Apr 2027
Evidence updated 2026-09-06Continued skepticism from academics and journalists
Scenario window: Through 2025–2026
Evidence updated 2026-03-13RE-Bench score reaches 1.3
Scenario window: Early 2026
Evidence updated 2026-09-06AI research capability exceeds the strongest human researchers
Scenario window: September 2027
Evidence updated 2026-09-06Stock market rises 30% in 2026
Scenario window: End of 2026
Evidence updated 2026-09-06Superhuman coder emerges
Scenario window: March 2027
Evidence updated 2026-09-06AI research capability becomes vastly better than the best humans
Scenario window: November 2027, racing branch
Evidence updated 2026-09-06Military tension around Taiwan escalates over AI race
Scenario window: Mid 2026–Early 2027
Evidence updated 2026-03-13Unreliable but useful AI agents emerge
Scenario window: Mid 2025
Evidence updated 2026-09-0610²⁸ FLOP training run completed
Scenario window: March 2027
Evidence updated 2026-09-06SWE-bench-Verified score reaches 85%
Scenario window: Mid 2025
Evidence updated 2026-09-07Leading US AI lab maintains a 3-9-month capability lead
Scenario window: Late 2025
Evidence updated 2026-09-07Weaker AI supervisors struggle to detect stronger systems’ misconduct
Scenario window: April-September 2027
Evidence updated 2026-09-06Better questions.
More accountable answers.
Explore the original predictions, the evidence supporting them and the findings that challenge them. Each dossier explains what would change its assessment.
Compare the scenario with reality →