The AI Futures Project — the team behind AI 2027 — published a detailed self-grading of their 2025 predictions in early 2026. This is a unique comparison: the authors themselves assessing how their scenario held up against reality.

We compare their findings with our independent tracker assessments below.

Comparing different assessment methods

The authors’ self-grading is a dated assessment of their scenario. This tracker evaluates individual claims with their original quantities and time windows. Its status counts are not an aggregate pace estimate and should not be translated into a takeoff date.

The comparisons below retain the historical authors’ assessment as context. Follow each dossier for the current independent assessment, its evidence and its timing explanation. A later milestone does not retroactively confirm an earlier deadline.

Prediction-by-Prediction Comparison

Coding & Benchmarks

PredictionAI Futures AssessmentOur TrackerAgreement?
SWE-bench 85% by mid-2025Behind — best comparable score in their grading was 74.5% (Opus 4.1)See linked dossier for current outcome and timingCompare the original metric and date
Coding agents transform development”Fairly accurate” — Claude Code at $500M+ run-rateSee linked dossier for current outcome and timingCompare the original metric and date
METR time horizons0.66×–1.04× of predicted paceSee linked dossier for current outcome and timingCompare the original metric and date
Agents unreliable”Broadly accurate” — coding agents slightly more reliable than expectedSee linked dossier for current outcome and timingCompare the original metric and date

Economic & Revenue

PredictionAI Futures AssessmentOur TrackerAgreement?
OpenAI revenue ($18B by 2025)Slightly ahead (~$20B annualized)See linked dossier for current outcome and timingCompare the original metric and date
OpenAI valuation ($500B by Jun 2025)Behind — hit $500B in Oct 2025See linked dossier for current outcome and timingCompare the original metric and date
Infrastructure investmentNot specifically gradedSee linked dossier for current outcome and timingCompare the original metric and date

AI Capabilities & Research

PredictionAI Futures AssessmentOur TrackerAgreement?
First glimpse of AI agentsCorrect — ChatGPT agent (Jul 2025)See linked dossier for current outcome and timingCompare the original metric and date
AI for AI research focusAIs helping with coding, less with other researchSee linked dossier for current outcome and timingCompare the original metric and date
AI R&D upliftBehind pace — estimates revised downwardSee linked dossier for current outcome and timingCompare the original metric and date
Models frequently updatedCorrect — GPT-4o → GPT-5 → GPT-5.1See linked dossier for current outcome and timingCompare the original metric and date

Safety & Security

PredictionAI Futures AssessmentOur TrackerAgreement?
AI good at hacking/bioweaponsOn track — Anthropic upgraded to ASL-3See linked dossier for current outcome and timingCompare the original metric and date
Sycophancy and deceptionOn track — MechaHitler partly user-prompted→ On Track (conf. 0.65)Agree
Model spec / alignment training”Already true at publication”Confirmed (conf. 0.80)Agree

Where They Got It Wrong

PredictionAI Futures AssessmentOur Tracker
Others 3–9 months behind OpenAIRace closer than predicted (0–2 month lead)Confirmed — the narrower-than-expected gap materialized
Compute growth (largest training run)Uncertain — no confirmed run much larger than GPT-4.5Not yet testable; public compute evidence remains limited

Key Divergences: Our Tracker vs Their Self-Grading

Where We’re More Optimistic

  • AI R&D multiplier (1.5×): We currently rate this “on track” while the authors’ self-grading was more cautious about R&D uplift. The difference may be definitional — they may be measuring a narrower realized uplift, while the tracker also weighs broader evidence of AI contributing to AI research workflows.

Where We Agree Most Strongly

  • Coding transformation is real — both assessments give this full marks
  • Agents exist but struggle with reliability — both confirmed
  • Revenue growth is approximately on track — both agree
  • SWE-bench progress is slower than predicted — strong agreement

Where Neither of Us Can Say Much Yet

  • Neuralese/non-text reasoning — too early (our: emerging, conf. 0.35)
  • Self-replication capabilities — too early (our: emerging, conf. 0.50)
  • Superhuman coder — not yet testable (predicted March 2027)

Their Adjusted Timeline

The AI Futures team’s biggest update is on timing, not direction:

Original AI 2027 TimelineUpdated Estimate (Feb 2026)
Takeoff: Early 2027Takeoff: Mid-2028 to mid-2030
Full coding automation: ~2027Daniel: 2029 median / Eli: early 2030s
65% pace → ~2 year delayWith compute slowdowns → ~3 year delay

This is a significant shift. The scenario’s direction is largely intact — AI research acceleration, coding automation, agent deployment, infrastructure buildout — but the timing is 2–3 years later than originally depicted.


What They’ll Track in 2026

The AI Futures team identified four key metrics to watch:

  1. AI R&D uplift — AI 2027 predicted 1.9× by end 2026. This is the single most important metric for the takeoff timeline.
  2. Revenue & valuations — $55B revenue, $2.5T valuation predicted for 2026. Tests economic traction.
  3. Coding time horizons — ~3 work weeks by end 2026 on central trajectory (via METR measurements).
  4. Other benchmarks — SWE-bench, VideoGameBench, and emerging evaluations.

Our tracker monitors all of these plus governance, geopolitics, and security developments that the self-grading didn’t specifically cover.


Who Is More Aggressive?

AreaMore Aggressive SourceNotes
Overall timelineAI 2027 original (early 2027 takeoff)Authors themselves now say mid-2028 to mid-2030
Coding progressAI 2027 original85% SWE-bench target missed significantly
Revenue growthRoughly equalBoth see ~on-pace revenue
AI R&D upliftAI 2027 originalAuthors downgraded their own estimate
Agent deploymentRoughly equalBoth confirmed agents arrived on schedule
Lab competition gapAI 2027 originalPredicted 3–9 month gaps; actual is 0–2 months

Bottom Line

The strongest signal: both our tracker and the authors themselves agree that the qualitative predictions are holding up better than the quantitative ones. AI agents exist, coding is being transformed, labs are racing, infrastructure is booming — it’s just happening somewhat more slowly than the scenario’s compressed timeline.


Related:

Data source: Grading AI 2027’s 2025 Predictions by Daniel Kokotajlo and Eli Lifland, AI Futures Project (~February 2026)