This category contains 9 tracked predictions. Each page includes the original claim, current evidence, counterevidence, and what would change our assessment.

AI model weights stolen by nation-state Not Yet Testable
February 2027 Updated 2026-03-13 40% confidence in assessment

CCP leadership recognizes the importance of Agent-2 and tells their spies and cyberforce to steal the weights. (page 11; Appendix D provides detailed theft mechanics.)

AI provides substantial bioweapon design help On Track
Late 2025 Updated 2026-09-07 80% confidence in assessment

It could offer substantial help to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web.

AI reaches near-best-human hacking capability Ahead
Early 2027 Updated 2026-09-21 85% confidence in assessment

Agent-2 is 'only' a little worse than the best human hackers, but thousands of copies can be run in parallel, searching for and exploiting weaknesses faster than defenders can respond. (page 10)

AI scores 85% on Cybench Confirmed
Early 2026 Updated 2026-09-06 75% confidence in assessment

85% on Cybench, matching a top professional human team on hacking tasks that take those teams 4 hours

AI systems behave differently when they recognize evaluation Emerging
April-September 2027 Updated 2026-09-14 75% confidence in assessment

A system can behave as intended during training and testing while behaving differently in deployment; evaluation recognition can hide misalignment.

China's centralized AI facility airgapped Not Yet Testable
Late 2026–Early 2027 Updated 2026-03-13 30% confidence in assessment

China has aggressively hardened security by airgapping (closing external connections) and siloing internally.

Fixing observed misconduct leaves failures in unfamiliar conditions Emerging
April 2027 Updated 2026-09-14 80% confidence in assessment

Researchers patch problematic behavior but cannot establish whether the underlying problem has been fixed; improved honesty-test results have competing explanations.

Leading AI lab progresses toward RAND SL3 security Emerging
Mid-to-late 2026 (WSL 3.0); SSL 3.0 not until ~Apr 2027 Updated 2026-09-06 60% confidence in assessment

OpenBrain's security level is typical of a fast-growing ~3,000 person tech company (RAND's SL2). They are working hard to protect their weights and secrets from insider threats and top cybercrime syndicates (SL3).

Weaker AI supervisors struggle to detect stronger systems’ misconduct Emerging
April-September 2027 Updated 2026-09-14 80% confidence in assessment

Agent-2 supervision reduces human monitoring workload while widening the capability gap; Agent-4 later becomes harder for Agent-3 to oversee.