Independent research / September 2026

AI 2027, measured
against reality.

Independent tracking of the AI 2027 scenario’s predictions, with evidence, counterevidence, and analysis of what developments mean for the outlook.

01a / General capability

The frontier across AI benchmarks

Explore reliable task completion ↓

Epoch’s capabilities index combines benchmark results into a common scale. Each point marks a model that raises the highest score in this snapshot.

Epoch model dates through September 3, 2026 · Checked September 11, 2026

Epoch Capabilities Index frontierCurrent estimates plotted by release dates as recorded by Epoch, on a linear index scale. Vertical bars show 90% confidence intervals. The connecting steps trace record scores, with no extrapolation. Epoch Capabilities Index · higher scores indicate stronger benchmark performance 120130140150160170180 Jan 2024Jul 2024Jan 2025Jul 2025Jan 2026Jul 2026 GPT-4 Turbo (Nov 2023): 126.45 ECI points. 90% interval: 121.8–128.9 points Claude 3 Opus: 126.91 ECI points. 90% interval: 121.7–130.1 points GPT-4 Turbo (Apr 2024): 127.25 ECI points. 90% interval: 122.5–128.9 points GPT-4o (May 2024): 128.97 ECI points. 90% interval: 124.4–131.6 points Claude 3.5 Sonnet: 130 ECI points. Calibration anchor; interval not reported. o1-mini: 135.83 ECI points. 90% interval: 132.4–137.4 points o1: 141.87 ECI points. 90% interval: 139.8–143.1 points Gemini 2.5 Pro (Mar 2025): 144.17 ECI points. 90% interval: 141.9–146.8 points o3: 146.91 ECI points. 90% interval: 145.0–148.4 points o3-pro: 147.45 ECI points. 90% interval: 145.9–150.2 points GPT-5: 150 ECI points. Calibration anchor; interval not reported. GPT-5 Pro: 150.28 ECI points. 90% interval: 148.8–152.5 points Gemini 3 Pro: 153.01 ECI points. 90% interval: 151.2–155.0 points GPT-5.2 Pro: 155.34 ECI points. 90% interval: 153.0–157.9 points GPT-5.3 Codex: 156.59 ECI points. 90% interval: 154.1–159.4 points GPT-5.4 Pro: 158.93 ECI points. 90% interval: 156.3–161.9 points GPT-5.5 Pro: 162.26 ECI points. 90% interval: 159.3–165.4 points Claude Fable 5: 163.41 ECI points. 90% interval: 160.7–167.1 points Claude Fable 5.1: 164.24 ECI points. 90% interval: 161.3–167.8 points GPT-6 Astra: 166.57 ECI points. 90% interval: 163.3–171.7 points

Swipe the chart for later dates, or choose a model below.

Record estimate · Bars: 90% intervals

The frontier is reconstructed from current scores. Historical scores can change when Epoch refits the index.

How to read this chart and inspect its data

The chart shows 20 record-setting estimates from 2024 onward, selected from 264 models in Epoch’s current export. Within a shared release date we select the highest score; later points must exceed all earlier scores. Dates are taken from Epoch’s export and may refer to a model variant. Connecting steps describe that selection, without a forecast. Overlapping intervals mean a higher central score does not establish a statistically clear lead.

ECI is an index, with an arbitrary scale. It is neither a percentage nor an IQ score. Claude 3.5 Sonnet and GPT-5 fix the scale at 130 and 150; their intervals are not reported. Epoch combines its own evaluations, independent benchmark results and model-developer reports, so benchmark selection and reporting bias remain limitations. These scores do not directly establish productivity gains or an AI 2027 prediction’s outcome.

ModelEpoch dateECI90% interval
GPT-4 Turbo (Nov 2023)2024-01-25126.45121.78–128.95
Claude 3 Opus2024-02-29126.91121.74–130.11
GPT-4 Turbo (Apr 2024)2024-04-09127.25122.46–128.93
GPT-4o (May 2024)2024-05-13128.97124.43–131.59
Claude 3.5 Sonnet2024-06-20130.00Calibration anchor
o1-mini2024-09-12135.83132.40–137.38
o12024-12-17141.87139.83–143.14
Gemini 2.5 Pro (Mar 2025)2025-03-31144.17141.94–146.84
o32025-04-16146.91144.99–148.44
o3-pro2025-06-10147.45145.86–150.15
GPT-52025-08-07150.00Calibration anchor
GPT-5 Pro2025-10-07150.28148.77–152.52
Gemini 3 Pro2025-11-18153.01151.16–154.96
GPT-5.2 Pro2025-12-11155.34153.04–157.88
GPT-5.3 Codex2026-02-05156.59154.13–159.41
GPT-5.4 Pro2026-03-05158.93156.35–161.87
GPT-5.5 Pro2026-04-23162.26159.30–165.37
Claude Fable 52026-06-09163.41160.68–167.10
Claude Fable 5.12026-09-01164.24161.30–167.85
GPT-6 Astra2026-09-03166.57163.34–171.70

Source: Epoch AI, CC BY 4.0. Chart selection and presentation by AI 2027 Tracker. Methodology · Limitations · Download all model scores (CSV) · Source provenance

01b / Reliable task completion

How much work can AI complete?

METR data updated May 8, 2026 · Last checked September 11, 2026

Task difficulty measured in the time a human expert would need. Each point is a published METR estimate; bars show statistical uncertainty.

METR task-completion time horizonsHuman-equivalent task duration versus model release date. Logarithmic duration axis. Estimates above 16 hours have limited reliability. Measurement-limited region 6 sec1 min10 min1 hour8 hours16 hours4 days Jan 2024Jul 2024Jan 2025Jul 2025Jan 2026Jul 2026 METR data updated Claude 3 opus: 0.07 hours; 95% interval 0.03–0.15 GPT-4 turbo: 0.06 hours; 95% interval 0.03–0.11 GPT-4o: 0.12 hours; 95% interval 0.07–0.22 Claude 3.5 sonnet 20240620: 0.19 hours; 95% interval 0.09–0.37 o1 preview: 0.34 hours; 95% interval 0.20–0.56 Claude 3.5 sonnet 20241022: 0.34 hours; 95% interval 0.17–0.68 o1: 0.65 hours; 95% interval 0.35–1.08 Claude 3.7 sonnet: 1.01 hours; 95% interval 0.55–1.74 o3: 2.00 hours; 95% interval 1.24–3.18 Claude 4 opus: 1.67 hours; 95% interval 1.00–2.72 Claude 4.1 opus: 1.67 hours; 95% interval 0.99–2.66 GPT-5.2025.08.07: 3.38 hours; 95% interval 1.88–6.76 Gemini 3 pro: 3.74 hours; 95% interval 2.33–6.32 GPT-5.1 codex max: 3.73 hours; 95% interval 2.24–6.60 Claude opus 4.5: 4.88 hours; 95% interval 2.70–10.40 GPT-5.2: 5.87 hours; 95% interval 3.30–13.59 Claude opus 4.6: 11.98 hours; 95% interval 5.28–60.56 GPT-5.3 codex: 5.83 hours; 95% interval 3.25–13.61 Gemini 3.1 pro: 6.40 hours; 95% interval 3.89–11.58 GPT-5.4: 5.70 hours; 95% interval 3.11–12.81 Claude mythos preview early: 17.41 hours; 95% interval 8.48–55.07

Published estimate Central estimate above 16 hours

Release dates through April 2026. Absence of newer points is a measurement gap, not a plateau.

Underlying observations and measurement limits

METR’s predominantly software, machine-learning and cybersecurity tasks measure human-equivalent difficulty. They do not represent a whole job. Success probability differs from the 95% statistical interval. The dataset excludes estimates above 16 hours from its fitted trend. No extrapolated or inferred frontier-model scores have been added.

ModelRelease50% horizon, hours80% horizon, hours50% interval, hours
Claude 3 opus2024-03-040.070.010.03–0.15
GPT-4 turbo2024-04-090.060.020.03–0.11
GPT-4o2024-05-130.120.020.07–0.22
Claude 3.5 sonnet 202406202024-06-200.190.030.09–0.37
o1 preview2024-09-120.340.070.20–0.56
Claude 3.5 sonnet 202410222024-10-220.340.040.17–0.68
o12024-12-050.650.120.35–1.08
Claude 3.7 sonnet2025-02-241.010.200.55–1.74
o32025-04-162.000.501.24–3.18
Claude 4 opus2025-05-221.670.341.00–2.72
Claude 4.1 opus2025-08-051.670.390.99–2.66
GPT-5.2025.08.072025-08-073.380.641.88–6.76
Gemini 3 pro2025-11-183.740.902.33–6.32
GPT-5.1 codex max2025-11-193.730.842.24–6.60
Claude opus 4.52025-11-244.880.822.70–10.40
GPT-5.22025-12-115.871.103.30–13.59
Claude opus 4.62026-02-0511.981.165.28–60.56
GPT-5.3 codex2026-02-055.830.913.25–13.61
Gemini 3.1 pro2026-02-196.401.503.89–11.58
GPT-5.42026-03-055.700.903.11–12.81
Claude mythos preview early2026-04-0717.413.108.48–55.07

Original data · Snapshot retrieved September 11, 2026 · Source provenance · Download chart data (CSV)

Capability

A faster benchmark regime

Epoch reports about 14 ECI points per year for the reasoning frontier. That is a faster trend, not proof that its slope keeps increasing.

Epoch, September 1 ↗
Real work

Useful, uneven, hard to measure

Developer productivity depends on tasks, experience and workflow. Selection effects complicate the latest experimental estimates.

METR’s experiment update ↗
Research feedback

The consequential open question

Engineering speed, research judgment and total research progress are different measures. A compounding feedback loop needs evidence across all three.

Inspect the R&D dossier →
02 / The causal chain

What has to happen next?

Research predictions →

A dependency map for investigation. Progress in one stage does not establish the next; oversight and physical constraints affect every stage.

03 / The evidence library

Predictions and evidence

Each dossier compares an original prediction with current evidence and counterevidence. The claims overlap; their status counts are not a probability that the whole scenario will occur.

13 Confirmed2 Ahead8 On track6 Behind22 Emerging15 Not yet testable
model capabilityBehind

Agent-0 training reaches 10²⁷ FLOP

Scenario window: Training completion: May 2025; public availability: late 2025

Evidence updated 2026-09-11
model capabilityNot yet testable

Frontier model uses daily weight updates / online learning

Scenario window: January 2027

Evidence updated 2026-09-06
economic impactOn track

Frontier model variant 10x cheaper released

Scenario window: Late 2026

Evidence updated 2026-09-06
agent autonomyConfirmed

Best AI agents cost hundreds of dollars per month

Scenario window: Mid 2025

Evidence updated 2026-09-06
economic impactEmerging

Ten percent of Americans consider an AI a close friend

Scenario window: July 2027

Evidence updated 2026-09-06
takeoffConfirmed

AI companies focus on AI-for-AI-research

Scenario window: 2025–2026

Evidence updated 2026-09-06
economic impactEmerging

AI accounts for 2.5% of US power capacity

Scenario window: End of 2026

Evidence updated 2026-09-06
economic impactEmerging

Global AI power consumption reaches 38GW

Scenario window: End of 2026

Evidence updated 2026-09-06
securityEmerging

Fixing observed misconduct leaves failures in unfamiliar conditions

Scenario window: April 2027

Evidence updated 2026-09-06
governanceEmerging

Large-scale anti-AI protest (10,000+ people)

Scenario window: Late 2026

Evidence updated 2026-06-15
agent autonomyEmerging

AI model capable of autonomous self-replication

Scenario window: January 2027

Evidence updated 2026-09-06
takeoffEmerging

AI teams conduct complete research projects with little human direction

Scenario window: June 2027

Evidence updated 2026-09-07
securityOn track

AI provides substantial bioweapon design help

Scenario window: Late 2025

Evidence updated 2026-09-07
governanceOn track

Newest or sensitive AI capabilities remain behind restricted access

Scenario window: Ongoing through 2027

Evidence updated 2026-09-07
economic impactEmerging

Global active AI compute reaches $1 trillion in ownership cost

Scenario window: Late 2026

Evidence updated 2026-09-07
securityNot yet testable

China's centralized AI facility airgapped

Scenario window: Late 2026–Early 2027

Evidence updated 2026-03-13
geopoliticsOn track

Chinese domestic AI chips 3 years behind US-Taiwan

Scenario window: Mid 2026

Evidence updated 2026-09-07
geopoliticsEmerging

China nationalizes/centralizes AI research

Scenario window: Mid 2026

Evidence updated 2026-09-06
model capabilityNot yet testable

AI substantially outperforms the best humans across cognitive domains

Scenario window: December 2027, racing branch

Evidence updated 2026-09-06
geopoliticsEmerging

China has ~12% of global AI-relevant compute

Scenario window: Mid 2026

Evidence updated 2026-09-06
geopoliticsOn track

Leading Chinese AI lab ~6 months behind US frontier

Scenario window: Mid 2026

Evidence updated 2026-09-07
model capabilityConfirmed

Models shift to continuous/iterative training

Scenario window: Late 2025

Evidence updated 2026-09-06
codingEmerging

Coding fully automated but research taste still requires humans

Scenario window: Early-Mid 2027

Evidence updated 2026-09-07
securityConfirmed

AI scores 85% on Cybench

Scenario window: Early 2026

Evidence updated 2026-09-06
codingConfirmed

Coding agents begin transforming professional work

Scenario window: Mid 2025

Evidence updated 2026-09-07
securityAhead

AI reaches near-best-human hacking capability

Scenario window: Early 2027

Evidence updated 2026-09-07
economic impactConfirmed

Massive datacenter buildouts continue

Scenario window: Through 2026

Evidence updated 2026-08-24
governanceAhead

Department of Defense scales up AI lab contracting

Scenario window: Late 2026

Evidence updated 2026-09-07
geopoliticsConfirmed

Export controls impact Chinese AI compute

Scenario window: Ongoing

Evidence updated 2026-09-06
securityEmerging

AI systems behave differently when they recognize evaluation

Scenario window: April-September 2027

Evidence updated 2026-09-07
governanceNot yet testable

A leading AI company’s net public approval falls to minus 35 points

Scenario window: July 2027

Evidence updated 2026-09-06
governanceNot yet testable

Frontier-model researchers become subject to security-clearance requirements

Scenario window: May-July 2027

Evidence updated 2026-09-06
takeoffNot yet testable

Leading humans still add substantial value to AI-led research

Scenario window: June 2027

Evidence updated 2026-09-06
economic impactOn track

Leading AI company reaches $45B annual revenue

Scenario window: End of 2026

Evidence updated 2026-09-07
economic impactBehind

Leading AI company reaches $3T valuation

Scenario window: End of 2026

Evidence updated 2026-09-06
model capabilityEmerging

IDA achieves superhuman performance at coding

Scenario window: Early 2027

Evidence updated 2026-09-06
economic impactConfirmed

Massive AI infrastructure investment continues

Scenario window: Ongoing through 2025–2026

Evidence updated 2026-08-03
governanceNot yet testable

External evaluators lack access to the strongest internally deployed models

Scenario window: May 2027

Evidence updated 2026-09-06
economic impactOn track

AI reshapes employment and intensifies public concern

Scenario window: Late 2026

Evidence updated 2026-09-07
agent autonomyConfirmed

Agents struggle with long-horizon tasks

Scenario window: 2025

Evidence updated 2026-09-06
agent autonomyOn track

METR time horizon doubles every 4 months

Scenario window: ~4 month doubling from 2024+

Evidence updated 2026-09-07
securityNot yet testable

AI model weights stolen by nation-state

Scenario window: February 2027

Evidence updated 2026-03-13
governanceEmerging

Nationalization of leading AI lab debated

Scenario window: February 2027

Evidence updated 2026-07-05
model capabilityEmerging

High-bandwidth non-text reasoning (neuralese) deployed

Scenario window: Early 2027

Evidence updated 2026-09-06
economic impactEmerging

Leading AI company reaches $40B annual compute costs, 6GW peak power and $200B active compute

Scenario window: End of 2026

Evidence updated 2026-09-07
model capabilityConfirmed

OSWorld benchmark reaches 65% by mid-2025

Scenario window: Mid 2025 (65%), Early 2026 (80%)

Evidence updated 2026-09-06
agent autonomyConfirmed

Computer-using agents marketed as 'personal assistants'

Scenario window: Mid 2025

Evidence updated 2026-09-06
economic impactEmerging

Hiring new programmers nearly stops

Scenario window: July 2027

Evidence updated 2026-09-06
codingNot yet testable

200,000–250,000 AI coding agents run in parallel

Scenario window: March 2027

Evidence updated 2026-09-06
takeoffEmerging

AI R&D progress multiplier reaches 2x

Scenario window: Late 2026

Evidence updated 2026-09-07
takeoffNot yet testable

AI R&D progress multiplier reaches 3×

Scenario window: January 2027

Evidence updated 2026-09-06
takeoffNot yet testable

AI algorithmic progress multiplier reaches 4× (~2× overall R&D)

Scenario window: March 2027

Evidence updated 2026-09-06
takeoffEmerging

AI R&D progress multiplier reaches 1.5×

Scenario window: Early 2026

Evidence updated 2026-09-07
securityEmerging

Leading AI lab progresses toward RAND SL3 security

Scenario window: Mid-to-late 2026 (WSL 3.0); SSL 3.0 not until ~Apr 2027

Evidence updated 2026-09-06
governanceConfirmed

Continued skepticism from academics and journalists

Scenario window: Through 2025–2026

Evidence updated 2026-03-13
codingBehind

RE-Bench score reaches 1.3

Scenario window: Early 2026

Evidence updated 2026-09-06
takeoffNot yet testable

AI research capability exceeds the strongest human researchers

Scenario window: September 2027

Evidence updated 2026-09-06
economic impactBehind

Stock market rises 30% in 2026

Scenario window: End of 2026

Evidence updated 2026-09-06
codingNot yet testable

Superhuman coder emerges

Scenario window: March 2027

Evidence updated 2026-09-06
takeoffNot yet testable

AI research capability becomes vastly better than the best humans

Scenario window: November 2027, racing branch

Evidence updated 2026-09-06
geopoliticsEmerging

Military tension around Taiwan escalates over AI race

Scenario window: Mid 2026–Early 2027

Evidence updated 2026-03-13
agent autonomyConfirmed

Unreliable but useful AI agents emerge

Scenario window: Mid 2025

Evidence updated 2026-09-06
model capabilityNot yet testable

10²⁸ FLOP training run completed

Scenario window: March 2027

Evidence updated 2026-09-06
codingBehind

SWE-bench-Verified score reaches 85%

Scenario window: Mid 2025

Evidence updated 2026-09-07
geopoliticsBehind

Leading US AI lab maintains a 3-9-month capability lead

Scenario window: Late 2025

Evidence updated 2026-09-07
securityEmerging

Weaker AI supervisors struggle to detect stronger systems’ misconduct

Scenario window: April-September 2027

Evidence updated 2026-09-06
An evolving public reference

Better questions.
More accountable answers.

Explore the original predictions, the evidence supporting them and the findings that challenge them. Each dossier explains what would change its assessment.

Compare the scenario with reality →