Agent-0 training reaches 10²⁷ FLOP

Last updated
Behind · Model Capability · 60% confidence in assessment
Predicted: Training completion: May 2025; public availability: late 2025 · Updated: 2026-09-11 · Source: AI 2027, Late 2025 narrative; Compute Forecast, Section 5 training-run table
OpenBrain's latest public model—Agent-0—was trained with 10²⁷ FLOP.

Leading publicly characterized 2025 runs reached several 10²⁶ FLOP, below the modeled 10²⁷ scale. The precise May completion milestone remains unresolved.

At a glance

  • Assessment: Behind
  • Confidence in assessment: 60%
  • Outcome: unresolved
  • Timing: Original window passed; outcome unresolved
  • Evidence: proxy
  • Predicted timing: Training completion: May 2025; public availability: late 2025
  • Primary source: AI 2027, Late 2025 narrative; Compute Forecast, Section 5 training-run table

What AI 2027 Predicted

The compute supplement models Agent-0 training from October 2024 through May 2025 at 10²⁷ FLOP. Public availability later in 2025 is a separate milestone. The May date therefore describes completion of a run already underway when the scenario was published, rather than a new project to be started and finished the following month. AI 2027 Compute Forecast, section 5

How We Track This

Compare estimated compute used in individual frontier training runs with the modeled scale and completion window. A release date is an imperfect proxy for training completion. Public estimates cannot rule out undisclosed runs.

Current Evidence

Epoch’s current model dataset estimates Grok 4 at 5 × 10²⁶ FLOP, GPT-4.5 at 3.8 × 10²⁶ and Grok 3 at 3.5 × 10²⁶. These characterized 2025 runs were approximately two to three times below the target. They are much closer than a comparison using GPT-5 alone would suggest. These values are research reconstructions, not company-confirmed training accounts. Epoch model dataset

Epoch’s October 2025 GPT-5 analysis estimated 5 × 10²⁵ FLOP. That dated estimate remains part of the historical record; it does not establish the maximum compute used by any lab. Epoch’s GPT-5 analysis

Counterevidence & Limitations

The leading estimates support a scale shortfall, but model identities, training boundaries and unreported experiments remain uncertain. Comparisons with models released after May cannot establish what had completed by May. Changes in training efficiency also mean a compute shortfall need not imply a proportionate capability shortfall.

Behind expresses a tentative assessment of the modeled training trajectory. It is not a confirmed failure established one month after publication.

What Would Change Our Assessment

  • Strengthen: Comparable records place a completed run at or above 10²⁷ FLOP, with its completion date documented.
  • Confirm a shortfall: More complete lab records establish that the relevant frontier remained materially below the target within the window.
  • Reassess timing: A later qualifying run would distinguish a delayed milestone from an unobserved one.

Update History

DateUpdate
2026-09-11Clarification of the September 7 comparison: leading characterized 2025 runs were approximately 2-3× below the modeled target. That broader comparison set confidence at 0.60, following the earlier 0.55 assessment recorded below. The current Behind rating and 0.60 confidence are unchanged.
2026-09-07Expanded the comparison beyond GPT-5: leading characterized 2025 runs were approximately 2-3× below the target. May training completion and later public availability remain separate; the exact milestone is unresolved.
2026-09-07Distinguished May training completion from late-2025 public availability and clarified the source of both dates. Behind remains a tentative trajectory assessment; the exact milestone is unresolved. Assessment confidence reduced from 0.75 to 0.55.
2026-09-06Assessment revised from emerging (0.45) to behind (0.75). No qualifying run was established by the original May 2025 deadline; estimates of disclosed runs do not exclude undisclosed training.
2026-03-13GPT-5 estimated at ~3×10²⁵ FLOP pretraining, well below 10²⁷ target. Significant uncertainty remains around total compute including post-training and inference-time scaling.