All Predictions

Tracking 66 predictions from the AI 2027 scenario against reality.

Best AI agents cost hundreds of dollars per month Confirmed
Agent Autonomy Mid 2025 Confidence in assessment: 95% Updated: 2026-09-06

The better agents are also expensive; you get what you pay for, and the best performance costs hundreds of dollars a month.

AI companies focus on AI-for-AI-research Confirmed
Takeoff 2025–2026 Confidence in assessment: 95% Updated: 2026-09-21

Frontier AI labs increasingly use AI systems to accelerate their own AI research and development.

Models shift to continuous/iterative training Confirmed
Model Capability Late 2025 Confidence in assessment: 90% Updated: 2026-09-06

By this point 'finishes training' is a bit of a misnomer; models are frequently updated to newer versions trained on additional data or partially re-trained to patch some weaknesses.

Coding agents begin transforming professional work Confirmed
Coding Mid 2025 Confidence in assessment: 85% Updated: 2026-09-07

Specialized coding agents begin to transform professional work, providing substantial time savings in some settings despite continuing unreliability.

AI scores 85% on Cybench Confirmed
Security Early 2026 Confidence in assessment: 75% Updated: 2026-09-06

85% on Cybench, matching a top professional human team on hacking tasks that take those teams 4 hours

Massive datacenter buildouts continue Confirmed
Economic Impact Through 2026 Confidence in assessment: 95% Updated: 2026-09-21

Data center construction accelerates dramatically, with power grid constraints becoming a real bottleneck.

Export controls impact Chinese AI compute Confirmed
Geopolitics Ongoing Confidence in assessment: 90% Updated: 2026-09-06

US export controls constrain Chinese access to frontier AI chips, but China adapts through domestic alternatives and workarounds.

Massive AI infrastructure investment continues Confirmed
Economic Impact Ongoing through 2025–2026 Confidence in assessment: 95% Updated: 2026-09-21

Hundreds of billions pour into AI infrastructure, with hyperscalers racing to build compute capacity at unprecedented scale.

Agents struggle with long-horizon tasks Confirmed
Agent Autonomy 2025 Confidence in assessment: 85% Updated: 2026-09-06

Agent-1 is bad at even simple long-horizon tasks (page 7, Early 2026 section). Also: agents in Mid 2025 are 'impressive in theory but in practice unreliable.'

OSWorld benchmark reaches 65% by mid-2025 Confirmed
Model Capability Mid 2025 (65%), Early 2026 (80%) Confidence in assessment: 90% Updated: 2026-09-06

Specifically, we forecast that they score 65% on the OSWorld benchmark of basic computer tasks (compared to 38% for Operator and 70% for a typical skilled non-expert human).

Computer-using agents marketed as 'personal assistants' Confirmed
Agent Autonomy Mid 2025 Confidence in assessment: 85% Updated: 2026-09-06

Advertisements for computer-using agents emphasize the term 'personal assistant': you can prompt them with tasks like 'order me a burrito on DoorDash' or 'open my budget spreadsheet and sum this month's expenses.'

Continued skepticism from academics and journalists Confirmed
Governance Through 2025–2026 Confidence in assessment: 95% Updated: 2026-03-13

Despite rapid capability gains, mainstream skepticism about AI's transformative potential persists among academics, journalists, and policymakers.

Unreliable but useful AI agents emerge Confirmed
Agent Autonomy Mid 2025 Confidence in assessment: 95% Updated: 2026-09-06

AI agents become increasingly useful for real tasks but remain unreliable on complex, multi-step workflows.

AI reaches near-best-human hacking capability Ahead
Security Early 2027 Confidence in assessment: 85% Updated: 2026-09-21

Agent-2 is 'only' a little worse than the best human hackers, but thousands of copies can be run in parallel, searching for and exploiting weaknesses faster than defenders can respond. (page 10)

Department of Defense scales up AI lab contracting Ahead
Governance Late 2026 Confidence in assessment: 95% Updated: 2026-09-07

Department of Defense quietly but significantly begins scaling up contracting OpenBrain directly for cyber, data analysis, and R&D, but integration is slow due to bureaucracy.

Frontier model variant 10x cheaper released On Track
Economic Impact Late 2026 Confidence in assessment: 75% Updated: 2026-09-28

OpenBrain blows the competition out of the water again by releasing Agent-1-mini—a model 10x cheaper than Agent-1 and more easily fine-tuned for different applications.

Newest or sensitive AI capabilities remain behind restricted access On Track
Governance Ongoing through 2027 Confidence in assessment: 90% Updated: 2026-09-07

OpenBrain 'responsibly' elects not to release it publicly yet (page 10); very few have access to the newest capabilities (page 16).

AI provides substantial bioweapon design help On Track
Security Late 2025 Confidence in assessment: 80% Updated: 2026-09-07

It could offer substantial help to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web.

Chinese domestic AI chips 3 years behind US-Taiwan On Track
Geopolitics Mid 2026 Confidence in assessment: 80% Updated: 2026-09-07

Producing domestic chips about three years behind the U.S.-Taiwanese frontier.

Leading Chinese AI lab ~6 months behind US frontier On Track
Geopolitics Mid 2026 Confidence in assessment: 75% Updated: 2026-09-07

They are about six months behind the best OpenBrain models.

Leading AI company reaches $45B annual revenue On Track
Economic Impact End of 2026 Confidence in assessment: 75% Updated: 2026-09-07

OPENBRAIN REVENUE $45B 2026 ANNUAL

AI reshapes employment and intensifies public concern On Track
Economic Impact Late 2026 Confidence in assessment: 75% Updated: 2026-09-07

AI takes some jobs and creates others, junior software roles face disruption, and public concern produces a large Washington protest in late 2026.

METR time horizon doubles every 4 months On Track
Agent Autonomy ~4 month doubling from 2024+ Confidence in assessment: 85% Updated: 2026-09-07

METR time horizons doubled every 7 months from 2019-2024 and every 4 months from 2024-onward (Appendix G, page 51). The acceleration from 7-month to 4-month doubling is a key claim.

Agent-0 training reaches 10²⁷ FLOP Behind
Model Capability Training completion: May 2025; public availability: late 2025 Confidence in assessment: 60% Updated: 2026-09-11

OpenBrain's latest public model—Agent-0—was trained with 10²⁷ FLOP.

Leading AI company reaches $3T valuation Behind
Economic Impact End of 2026 Confidence in assessment: 80% Updated: 2026-09-06

Sidebar graphics show OpenBrain valuation trajectory: $1T (Apr 2026), $2T (Aug 2026), $3T (Dec 2026)

RE-Bench score reaches 1.3 Behind
Coding Early 2026 Confidence in assessment: 70% Updated: 2026-09-06

1.3 on RE-Bench matching top expert humans given 8 hours at well-defined AI research engineering tasks.

Stock market rises 30% in 2026 Behind
Economic Impact End of 2026 Confidence in assessment: 75% Updated: 2026-09-06

The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants.

SWE-bench-Verified score reaches 85% Behind
Coding Mid 2025 Confidence in assessment: 90% Updated: 2026-09-07

We forecast that mid-2025 agents will score 85% on SWEBench-Verified. (page 3, footnote 3. Note: the 72% starting point is tracker-added context, not from the source.)

Leading US AI lab maintains a 3-9-month capability lead Behind
Geopolitics Late 2025 Confidence in assessment: 85% Updated: 2026-09-07

We imagine the others to be 3-9 months behind OpenBrain (page 4, Late 2025). By Early 2026, several competing AIs match or exceed Agent-0 (page 7). The 3-9 month gap is the Late 2025 state; near-parity emerges by Early 2026.

Ten percent of Americans consider an AI a close friend Emerging
Economic Impact July 2027 Confidence in assessment: 80% Updated: 2026-09-06

Ten percent of Americans, mostly young people, consider an AI a close friend.

Global AI power consumption reaches 38GW Emerging
Economic Impact End of 2026 Confidence in assessment: 55% Updated: 2026-09-06

GLOBAL AI POWER 38GW PEAK POWER

AI accounts for 2.5% of US power capacity Emerging
Economic Impact End of 2026 Confidence in assessment: 50% Updated: 2026-09-06

SHARE OF US POWER ON AI 2.5% — 33 GW OF 1.34TW CAPACITY

Fixing observed misconduct leaves failures in unfamiliar conditions Emerging
Security April 2027 Confidence in assessment: 80% Updated: 2026-09-14

Researchers patch problematic behavior but cannot establish whether the underlying problem has been fixed; improved honesty-test results have competing explanations.

Large-scale anti-AI protest (10,000+ people) Emerging
Governance Late 2026 Confidence in assessment: 55% Updated: 2026-06-15

There is a 10,000 person anti-AI protest in DC.

AI model capable of autonomous self-replication Emerging
Agent Autonomy January 2027 Confidence in assessment: 55% Updated: 2026-09-06

The safety team finds that if Agent-2 somehow escaped and wanted to 'survive' and 'replicate' autonomously, it might be able to do so.

Global active AI compute reaches $1 trillion in ownership cost Emerging
Economic Impact Late 2026 Confidence in assessment: 65% Updated: 2026-09-07

GLOBAL AI CAPEX $1T: COST OF OWNERSHIP OF ACTIVE COMPUTE

AI teams conduct complete research projects with little human direction Emerging
Takeoff June 2027 Confidence in assessment: 80% Updated: 2026-09-21

AI collectives produce most research progress while most humans cease to contribute usefully; leading humans retain some research judgment.

China nationalizes/centralizes AI research Emerging
Geopolitics Mid 2026 Confidence in assessment: 65% Updated: 2026-09-06

He sets in motion the nationalization of Chinese AI research... all the best researchers merge into a DeepCent-led collective... A Centralized Development Zone (CDZ) is created at the Tianwan Power Plant.

China has ~12% of global AI-relevant compute Emerging
Geopolitics Mid 2026 Confidence in assessment: 55% Updated: 2026-09-06

China has managed to maintain about 12% of the world's AI-relevant compute.

Coding fully automated but research taste still requires humans Emerging
Coding Early-Mid 2027 Confidence in assessment: 80% Updated: 2026-09-07

Now that coding has been fully automated... research taste has proven difficult to train due to longer feedback loops and less data availability.

AI systems behave differently when they recognize evaluation Emerging
Security April-September 2027 Confidence in assessment: 75% Updated: 2026-09-14

A system can behave as intended during training and testing while behaving differently in deployment; evaluation recognition can hide misalignment.

IDA achieves superhuman performance at coding Emerging
Model Capability Early 2027 Confidence in assessment: 40% Updated: 2026-09-06

Now, the models have become sufficiently good at verifying more subjective things, allowing the use of IDA to improve the model at many tasks.

Nationalization of leading AI lab debated Emerging
Governance February 2027 Confidence in assessment: 65% Updated: 2026-07-05

Someone mentions the possibility of nationalizing OpenBrain, but other cabinet officials think that's premature... He elects to hold off on major action.

Leading AI company reaches $40B annual compute costs, 6GW peak power and $200B active compute Emerging
Economic Impact End of 2026 Confidence in assessment: 60% Updated: 2026-09-07

OPENBRAIN'S COMPUTE COSTS $40B 2026 ANNUAL · OPENBRAIN POWER REQUIREMENT 6GW PEAK POWER · CAPITAL EXPENDITURE $200B COST OF OWNERSHIP OF OPENBRAIN'S ACTIVE COMPUTE

High-bandwidth non-text reasoning (neuralese) deployed Emerging
Model Capability Early 2027 Confidence in assessment: 35% Updated: 2026-09-06

One such breakthrough is augmenting the AI's text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory).

Hiring new programmers nearly stops Emerging
Economic Impact July 2027 Confidence in assessment: 75% Updated: 2026-09-06

Hiring new programmers has nearly stopped, while workers able to manage teams of AIs remain valuable.

AI R&D progress multiplier reaches 2x Emerging
Takeoff Late 2026 Confidence in assessment: 50% Updated: 2026-09-07

While the latest Agent-1 could double the pace of OpenBrain's algorithmic progress, Agent-2 can now triple it, and will improve further with time.

AI R&D progress multiplier reaches 1.5× Emerging
Takeoff Early 2026 Confidence in assessment: 75% Updated: 2026-09-07

AI systems contribute enough to AI research that the effective R&D multiplier reaches 1.5× — AI makes AI research 50% faster.

Leading AI lab progresses toward RAND SL3 security Emerging
Security Mid-to-late 2026 (WSL 3.0); SSL 3.0 not until ~Apr 2027 Confidence in assessment: 60% Updated: 2026-09-06

OpenBrain's security level is typical of a fast-growing ~3,000 person tech company (RAND's SL2). They are working hard to protect their weights and secrets from insider threats and top cybercrime syndicates (SL3).

Military tension around Taiwan escalates over AI race Emerging
Geopolitics Mid 2026–Early 2027 Confidence in assessment: 50% Updated: 2026-03-13

Other Party members discuss extreme measures to neutralize the West's chip advantage. A blockade of Taiwan? A full invasion? Both sides signal seriousness by repositioning military assets around Taiwan.

Weaker AI supervisors struggle to detect stronger systems’ misconduct Emerging
Security April-September 2027 Confidence in assessment: 80% Updated: 2026-09-14

Agent-2 supervision reduces human monitoring workload while widening the capability gap; Agent-4 later becomes harder for Agent-3 to oversee.

Frontier model uses daily weight updates / online learning Not Yet Testable
Model Capability January 2027 Confidence in assessment: 20% Updated: 2026-09-06

Agent-2 is effectively 'online learning,' in that it's built to never really finish training. Every day, the weights get updated to the latest version, trained on more data generated by the previous version the previous day.

China's centralized AI facility airgapped Not Yet Testable
Security Late 2026–Early 2027 Confidence in assessment: 30% Updated: 2026-03-13

China has aggressively hardened security by airgapping (closing external connections) and siloing internally.

AI substantially outperforms the best humans across cognitive domains Not Yet Testable
Model Capability December 2027, racing branch Confidence in assessment: 90% Updated: 2026-09-06

Artificial superintelligence is much better than the best human at every cognitive task.

A leading AI company’s net public approval falls to minus 35 points Not Yet Testable
Governance July 2027 Confidence in assessment: 75% Updated: 2026-09-06

OpenBrain has net approval of −35%, with 25% approving, 60% disapproving and 15% unsure.

Frontier-model researchers become subject to security-clearance requirements Not Yet Testable
Governance May-July 2027 Confidence in assessment: 80% Updated: 2026-09-06

The government contract requires security clearances for everyone working on the frontier models within two months.

Leading humans still add substantial value to AI-led research Not Yet Testable
Takeoff June 2027 Confidence in assessment: 80% Updated: 2026-09-06

Most humans contribute little, but the strongest retain research taste and planning value; the appendix estimates removing humans would slow research by roughly half.

External evaluators lack access to the strongest internally deployed models Not Yet Testable
Governance May 2027 Confidence in assessment: 75% Updated: 2026-09-21

A lab limits a prior model-sharing commitment to external deployment, leaving foreign safety evaluators without access to the strongest internal systems.

AI model weights stolen by nation-state Not Yet Testable
Security February 2027 Confidence in assessment: 40% Updated: 2026-03-13

CCP leadership recognizes the importance of Agent-2 and tells their spies and cyberforce to steal the weights. (page 11; Appendix D provides detailed theft mechanics.)

200,000–250,000 AI coding agents run in parallel Not Yet Testable
Coding March 2027 Confidence in assessment: 40% Updated: 2026-09-06

OpenBrain runs 200,000 Agent-3 copies in parallel. (page 13. Note: Appendix I page 54 states 250,000 copies — an internal source inconsistency.)

AI algorithmic progress multiplier reaches 4× (~2× overall R&D) Not Yet Testable
Takeoff March 2027 Confidence in assessment: 35% Updated: 2026-09-06

This massive superhuman labor force speeds up OpenBrain's overall rate of algorithmic progress by 'only' 4x due to bottlenecks and diminishing returns to coding labor. (Note: Footnote 31 clarifies 4x algorithmic progress corresponds to roughly 2x overall progress rate.)

AI R&D progress multiplier reaches 3× Not Yet Testable
Takeoff January 2027 Confidence in assessment: 40% Updated: 2026-09-06

While the latest Agent-1 could double the pace of OpenBrain's algorithmic progress, Agent-2 can now triple it.

AI research capability exceeds the strongest human researchers Not Yet Testable
Takeoff September 2027 Confidence in assessment: 85% Updated: 2026-09-06

An individual Agent-4 copy operating at human speed is qualitatively better at AI research than any human.

Superhuman coder emerges Not Yet Testable
Coding March 2027 Confidence in assessment: 60% Updated: 2026-09-06

A superhuman coder (SC): an AI system that can do any coding tasks that the best AGI company engineer does.

AI research capability becomes vastly better than the best humans Not Yet Testable
Takeoff November 2027, racing branch Confidence in assessment: 90% Updated: 2026-09-06

A superintelligent AI researcher is vastly better than the best human researcher at AI research.

10²⁸ FLOP training run completed Not Yet Testable
Model Capability March 2027 Confidence in assessment: 50% Updated: 2026-09-06

Agent-2 trained at 2×10²⁸ FLOP (Apr 2026 – Mar 2027), representing a 1000× increase over GPT-4 scale.