AI 2027 vs AI Futures Project Self-Grading
The AI Futures Project — the team behind AI 2027 — published a detailed self-grading of their 2025 predictions in early 2026. This is a unique comparison: the authors themselves assessing how their scenario held up against reality.
We compare their findings with our independent tracker assessments below.
Comparing different assessment methods
The authors’ self-grading is a dated assessment of their scenario. This tracker evaluates individual claims with their original quantities and time windows. Its status counts are not an aggregate pace estimate and should not be translated into a takeoff date.
The comparisons below retain the historical authors’ assessment as context. Follow each dossier for the current independent assessment, its evidence and its timing explanation. A later milestone does not retroactively confirm an earlier deadline.
Prediction-by-Prediction Comparison
Coding & Benchmarks
| Prediction | AI Futures Assessment | Our Tracker | Agreement? |
|---|---|---|---|
| SWE-bench 85% by mid-2025 | Behind — best comparable score in their grading was 74.5% (Opus 4.1) | See linked dossier for current outcome and timing | Compare the original metric and date |
| Coding agents transform development | ”Fairly accurate” — Claude Code at $500M+ run-rate | See linked dossier for current outcome and timing | Compare the original metric and date |
| METR time horizons | 0.66×–1.04× of predicted pace | See linked dossier for current outcome and timing | Compare the original metric and date |
| Agents unreliable | ”Broadly accurate” — coding agents slightly more reliable than expected | See linked dossier for current outcome and timing | Compare the original metric and date |
Economic & Revenue
| Prediction | AI Futures Assessment | Our Tracker | Agreement? |
|---|---|---|---|
| OpenAI revenue ($18B by 2025) | Slightly ahead (~$20B annualized) | See linked dossier for current outcome and timing | Compare the original metric and date |
| OpenAI valuation ($500B by Jun 2025) | Behind — hit $500B in Oct 2025 | See linked dossier for current outcome and timing | Compare the original metric and date |
| Infrastructure investment | Not specifically graded | See linked dossier for current outcome and timing | Compare the original metric and date |
AI Capabilities & Research
| Prediction | AI Futures Assessment | Our Tracker | Agreement? |
|---|---|---|---|
| First glimpse of AI agents | Correct — ChatGPT agent (Jul 2025) | See linked dossier for current outcome and timing | Compare the original metric and date |
| AI for AI research focus | AIs helping with coding, less with other research | See linked dossier for current outcome and timing | Compare the original metric and date |
| AI R&D uplift | Behind pace — estimates revised downward | See linked dossier for current outcome and timing | Compare the original metric and date |
| Models frequently updated | Correct — GPT-4o → GPT-5 → GPT-5.1 | See linked dossier for current outcome and timing | Compare the original metric and date |
Safety & Security
| Prediction | AI Futures Assessment | Our Tracker | Agreement? |
|---|---|---|---|
| AI good at hacking/bioweapons | On track — Anthropic upgraded to ASL-3 | See linked dossier for current outcome and timing | Compare the original metric and date |
| Sycophancy and deception | On track — MechaHitler partly user-prompted | → On Track (conf. 0.65) | Agree |
| Model spec / alignment training | ”Already true at publication” | Confirmed (conf. 0.80) | Agree |
Where They Got It Wrong
| Prediction | AI Futures Assessment | Our Tracker |
|---|---|---|
| Others 3–9 months behind OpenAI | Race closer than predicted (0–2 month lead) | Confirmed — the narrower-than-expected gap materialized |
| Compute growth (largest training run) | Uncertain — no confirmed run much larger than GPT-4.5 | Not yet testable; public compute evidence remains limited |
Key Divergences: Our Tracker vs Their Self-Grading
Where We’re More Optimistic
- AI R&D multiplier (1.5×): We currently rate this “on track” while the authors’ self-grading was more cautious about R&D uplift. The difference may be definitional — they may be measuring a narrower realized uplift, while the tracker also weighs broader evidence of AI contributing to AI research workflows.
Where We Agree Most Strongly
- Coding transformation is real — both assessments give this full marks
- Agents exist but struggle with reliability — both confirmed
- Revenue growth is approximately on track — both agree
- SWE-bench progress is slower than predicted — strong agreement
Where Neither of Us Can Say Much Yet
- Neuralese/non-text reasoning — too early (our: emerging, conf. 0.35)
- Self-replication capabilities — too early (our: emerging, conf. 0.50)
- Superhuman coder — not yet testable (predicted March 2027)
Their Adjusted Timeline
The AI Futures team’s biggest update is on timing, not direction:
| Original AI 2027 Timeline | Updated Estimate (Feb 2026) |
|---|---|
| Takeoff: Early 2027 | Takeoff: Mid-2028 to mid-2030 |
| Full coding automation: ~2027 | Daniel: 2029 median / Eli: early 2030s |
| 65% pace → ~2 year delay | With compute slowdowns → ~3 year delay |
This is a significant shift. The scenario’s direction is largely intact — AI research acceleration, coding automation, agent deployment, infrastructure buildout — but the timing is 2–3 years later than originally depicted.
What They’ll Track in 2026
The AI Futures team identified four key metrics to watch:
- AI R&D uplift — AI 2027 predicted 1.9× by end 2026. This is the single most important metric for the takeoff timeline.
- Revenue & valuations — $55B revenue, $2.5T valuation predicted for 2026. Tests economic traction.
- Coding time horizons — ~3 work weeks by end 2026 on central trajectory (via METR measurements).
- Other benchmarks — SWE-bench, VideoGameBench, and emerging evaluations.
Our tracker monitors all of these plus governance, geopolitics, and security developments that the self-grading didn’t specifically cover.
Who Is More Aggressive?
| Area | More Aggressive Source | Notes |
|---|---|---|
| Overall timeline | AI 2027 original (early 2027 takeoff) | Authors themselves now say mid-2028 to mid-2030 |
| Coding progress | AI 2027 original | 85% SWE-bench target missed significantly |
| Revenue growth | Roughly equal | Both see ~on-pace revenue |
| AI R&D uplift | AI 2027 original | Authors downgraded their own estimate |
| Agent deployment | Roughly equal | Both confirmed agents arrived on schedule |
| Lab competition gap | AI 2027 original | Predicted 3–9 month gaps; actual is 0–2 months |
Bottom Line
The strongest signal: both our tracker and the authors themselves agree that the qualitative predictions are holding up better than the quantitative ones. AI agents exist, coding is being transformed, labs are racing, infrastructure is booming — it’s just happening somewhat more slowly than the scenario’s compressed timeline.
Related:
- Which AI 2027 predictions came true? →
- AI 2027 vs Reality — the full picture →
- AI 2027 vs Metaculus crowd forecasts →
- Browse all tracked predictions →
Data source: Grading AI 2027’s 2025 Predictions by Daniel Kokotajlo and Eli Lifland, AI Futures Project (~February 2026)