AI 2027 Predictions Explained
How Do You Turn a Story Into Predictions?
AI 2027 is a narrative scenario rather than a list of predictions. It tells a story, month by month, about how AI development unfolds. To track it, we need to extract discrete, testable claims from that narrative.
This is harder than it sounds, and we want to be transparent about how we do it.
Extraction Process
We read the full scenario and identify specific factual claims: things that are either true or false, happening or not happening. Each claim needs to be:
- Specific enough to evaluate: “AI gets better” doesn’t count. “SWE-bench-Verified reaches 85% by mid-2025” does.
- Attributable to a timeframe: The scenario is chronological, so most claims have implicit or explicit timing.
- Distinct from other claims: We separate claims where they have different observable tests. Related predictions can still overlap, so status counts are not independent votes on the scenario.
Some claims are quantitative (benchmark scores, revenue figures, compute targets) and relatively easy to evaluate. Others are qualitative (institutional dynamics, public sentiment, strategic decisions) and require more judgment.
We currently track 66 predictions across 8 categories. This is not exhaustive. The scenario contains hundreds of implicit claims, and the tracker focuses on the most important and trackable ones.
What We Don’t Track
Some elements of the scenario are inherently untestable or not yet meaningful to track:
- Hypothetical decisions (e.g., “OpenBrain decides to continue racing”): these are scenario branches, not predictions
- Extremely vague claims: narrative color that doesn’t make a testable assertion
- Internal details with no observable test: we track lab dynamics where evidence can bear on a defined claim, while acknowledging gaps in public visibility
Our Status Taxonomy
Every prediction gets one of six statuses. Here’s what each means and how we assign it:
Confirmed
What it means: The predicted event or trend has clearly materialized, within roughly the predicted timeframe.
Example: “Unreliable but useful AI agents emerge”: ChatGPT agent launched July 2025, consistent with the scenario’s mid-2025 prediction.
What would change it: Nothing, short of evidence that we misidentified the event. Confirmed predictions stay confirmed.
Ahead
What it means: The prediction is happening faster than the scenario expected.
Example: Defense contracting with frontier AI labs: CDAO announced awards to four labs in July 2025, before the scenario’s late-2026 window. The early contracting milestone supports the Ahead assessment; it does not establish complete operational integration. CDAO announcement
What would change it: For an ongoing trend, a sustained slowdown could change the pace assessment. For a completed milestone, new evidence would need to change our understanding of what happened or when.
→ On Track
What it means: The prediction is progressing roughly as expected. It is not yet fully confirmed. The trajectory aligns.
Example: METR time horizons doubling roughly every four months: METR’s May 2026 report fits a 105-day doubling time to frontier public models released after January 1, 2024, at 50% task success. Our dossier judges this broadly consistent with the scenario’s trajectory. The fit depends on the model window and benchmark coverage. METR report, Appendix E
What would change it: Confirmation (upgrade to Confirmed) or evidence of a slowdown (downgrade to Behind).
Behind
What it means: The predicted event is happening more slowly than the scenario expected, or hasn’t happened by the predicted timeframe.
Example: “SWE-bench-Verified score reaches 85%”: Predicted by mid-2025; progress was meaningful, but comparable public evidence arrived late relative to the target.
What would change it: Catch-up can change an ongoing trajectory assessment. Later attainment of a dated milestone is recorded alongside its missed deadline.
Emerging
What it means: We see early signals that point toward this prediction, but it’s too soon to score definitively. The evidence is suggestive, not conclusive.
Example: Global active AI compute reaches $1 trillion in ownership value by late 2026: our September 7 reconstruction places central year-end estimates at $0.83-0.86 trillion, with higher scenarios exceeding $1 trillion. The buildout is close to the scenario’s scale, while the exact outcome remains unresolved. These are modeled values of operating hardware, including servers and networking; annual capex and announced commitments measure different things. Original compute forecast
What would change it: Stronger evidence (upgrade to On Track or Confirmed) or evidence the trend is stalling (downgrade to Behind).
◌ Not Yet Testable
What it means: The target remains in the future and available evidence does not yet support a meaningful assessment of the full claim. We still track relevant early evidence.
Example: AI accelerates algorithmic progress by 4×: the scenario places this milestone in March 2027 and describes it as roughly 2× overall progress once compute constraints are included. Task-level productivity gains alone cannot establish that full research multiplier. AI 2027, March 2027 and its explanatory footnote
What would change it: Time passing and evidence arriving. These will eventually move to other statuses.
How to Read a Prediction Page
Each of our 66 prediction pages follows a consistent structure. Here’s what you’ll find:
Header
- Title: a clear label for the prediction
- Status badge: current assessment with color coding
- Category: which domain (Model Capability, Coding, Geopolitics, etc.)
- Confidence score: how certain we are about our status assessment
- Predicted date: when the scenario expects this to happen
- Last updated: when the page last received a substantive update
Body Sections
-
What AI 2027 Predicted: The original claim from the scenario, in context. We quote or closely paraphrase the source material.
-
How We Track This: Our operationalization: what real-world indicators do we monitor? What benchmarks, data sources, or events count as evidence?
-
Current Evidence: Sourced evidence supporting our current status assessment. Every claim here should have a link or citation.
-
Counterevidence & Limitations: What argues against our assessment? What are we uncertain about? This section exists to keep us honest.
-
What Would Change Our Assessment: Explicit criteria for upgrading or downgrading the status. This makes our reasoning auditable.
-
Update History: A changelog of status changes with dates and reasoning.
What Confidence Scores Mean
Each prediction has a confidence score on a 0-1 scale that represents how certain we are about our status assessment. It does not represent how likely the prediction is to ultimately come true.
- 0.90-0.95: Very high confidence in our assessment. Clear evidence, minimal ambiguity. The tracker caps confidence at 0.95 rather than treating any public assessment as certain.
- 0.70-0.89: High confidence. Strong evidence, but some room for interpretation.
- 0.50-0.69: Moderate confidence. Evidence points one way but isn’t conclusive.
- 0.30-0.49: Low confidence. Genuine uncertainty about how to read the evidence.
- Below 0.30: Very uncertain. We’re making a judgment call with limited information.
Important distinction: A “confirmed” prediction with 0.85 confidence means we’re quite sure it’s confirmed, but there’s some ambiguity. It does not mean there’s an 85% chance the prediction is true. We already believe it’s true, and we are not perfectly certain about our interpretation.
Why This Matters
Tracking predictions from a narrative scenario is inherently messier than scoring a prediction market or a list of concrete bets. We’re making interpretive choices at every step:
- Which claims count as “predictions”?
- How do we operationalize qualitative claims?
- What counts as sufficient evidence?
- When does “on track” become “behind”?
We try to be transparent about these choices. Our methodology page goes deeper into the principles we follow. And every prediction page shows its evidence and reasoning, so you can disagree with our assessments if you see the evidence differently.
The goal isn’t to “grade” AI 2027 as right or wrong. It’s to build a structured, honest, ongoing record of how the AI 2027 scenario compares with reality. That record should be useful whether the scenario turns out to be prescient or overblown.
Read more: