AI 2027 vs Reality

Last updated

Capabilities are advancing; the research feedback loop remains unresolved

AI 2027 describes a chain of developments: more capable models enable useful agents, agents accelerate AI research, and that acceleration creates still more capable systems. The evidence for each link differs. Progress on a coding benchmark cannot by itself establish the full chain.

The strongest public evidence shows increasingly capable agents, substantial investment and AI contributing to research and engineering. The key uncertainty is how far those contributions increase accepted research progress across complete projects, once human work, experimental compute, failed experiments and quality are included.

What the capability chart tells us

The homepage chart shows METR task horizons at two reliability thresholds, with uncertainty and source data. A horizon describes the human task duration associated with a model’s success rate. It is not a measurement of how long an agent runs unattended.

The METR doubling dossier explains the fitting window and suite limitations. Its evidence supports continued growth; it does not justify translating all AI development into a single speed ratio. Longer-horizon estimates become less reliable when the task suite saturates.

Useful agents are not yet the whole research loop

Documented AI-for-AI research establishes that AI contributes to its own development. Coding automation and research judgment require more specific comparisons. Automated research results on well-characterized problems are meaningful, while proposal selection and open-ended project outcomes test different capabilities.

The research multiplier dossier therefore asks about accepted progress against a credible no-AI comparison. Code volume, reported hours saved and benchmark performance are supporting indicators. None supplies that counterfactual on its own.

Investment needs the original units

The financial predictions are unusually easy to misread. Annual company revenue is different from an annualized run rate. Active-compute ownership value is different from cumulative spending or a future investment commitment. Peak power is different from energy consumed over a year.

Investment announcements establish a substantial pipeline. Judging the numerical predictions requires knowing what has actually become operational and matching the original quantity.

Timing and visibility change the interpretation

A benchmark milestone can occur after its predicted date. That is evidence of capability progress and a timing miss. An undisclosed training run or internal research workflow presents a different problem: missing public evidence cannot establish that it did not happen.

China-US comparisons also need a fixed measurement basis. The model-gap dossier distinguishes dated capability estimates from release clustering, while the chip-gap dossier separates hardware performance, production and deployment.

What would change the outlook

  • Independently assessed research projects that measure accepted progress, human intervention, quality and compute on comparable terms.
  • Reliable capability evaluations that extend coverage beyond saturated benchmarks.
  • Direct evidence about oversight failures and whether mitigations generalize to new settings.
  • Comparable operational infrastructure and realized financial measurements.

These observations could strengthen or weaken different parts of the scenario. They should change the relevant assessments before changing a broad narrative about acceleration.

Explore the predictions · Read the methodology · Which predictions came true?