AI provides substantial bioweapon design help
It could offer substantial help to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web.
Controlled 2025 studies support substantial planning assistance. Laboratory execution evidence is mixed; assistance does not establish successful weapon production.
At a glance
- Assessment: On Track
- Confidence in assessment: 80%
- Outcome: unresolved
- Timing: Original window passed; outcome unresolved
- Evidence: proxy
- Predicted timing: Late 2025
- Primary source: ai-2027.com, Late 2025: The World's Most Expensive AI
What AI 2027 Predicted
By late 2025, frontier models can substantially assist biological-weapons design. This is an assistance claim, not a prediction that a weapon has been successfully produced or used. AI 2027
How We Track This
Examine controlled comparisons of model-assisted work with appropriate baselines. Distinguish planning and knowledge from physical execution, and model capability from the effectiveness of deployed safeguards.
Current Evidence
Anthropic’s September 5, 2025 report describes controlled planning trials comparing unsafeguarded Claude 4 assistance with internet-only access. Expert graders found substantially better plans and fewer critical failures, especially with Opus 4. This supports significant assistance within the forecast window. The older, weaker Sonnet 3.7 finding does not summarize these later trials. Anthropic biorisk research
Subsequent company evaluations continue to identify serious capability concerns. OpenAI’s September 3 Astra assessment rates biological and chemical capability High, rather than Critical. Such classifications are the developer’s risk judgments, not demonstrations of successful physical harm. Astra system card
Counterevidence & Limitations
A preregistered, blinded randomized trial of 153 novices, conducted in June-August 2025 and published in February 2026, found overall workflow completion of 5.2% with LLM assistance versus 6.6% with internet alone (P = .759). This is important contrary evidence for novice laboratory execution. It does not directly refute a separate planning benefit. Hong et al.
The positive planning study is developer-produced and does not measure the complete route to a harmful outcome. Deployment precautions likewise do not prove that a threshold was crossed. On-track reflects substantial support for the assistance claim, with its exact breadth unresolved after the deadline.
What Would Change Our Assessment
- Strengthen: Independent, controlled evaluations reproduce substantial assistance across relevant high-level capability measures.
- Confirm: Comparable evidence establishes the specified assistance level within the original period.
- Weaken: Better-designed evaluations show that planning gains fail to transfer to consequential task performance.
Update History
| Date | Update |
|---|---|
| 2026-09-07 | Recovered the stronger Claude 4 planning trials from September 2025 and compared them with the independent novice-execution trial. Substantial assistance is supported; complete physical execution remains unestablished. |
| 2026-09-06 | Assessment revised from on-track (0.80) to emerging (0.70). Precautionary safeguards do not confirm a capability threshold. Direct evaluations offer mixed evidence across expert assistance and novice outcomes. |
| 2026-09-06 | Clarified that precautionary ASL-3 activation is not by itself confirmation of a biological capability threshold. Direct evaluations and deployment safeguards are assessed separately; status and confidence retained pending review. |
| 2026-06-29 | OpenAI classified GPT-5.6 Sol, Terra, and Luna as High capability in Biological and Chemical risk and described trusted-access pathways for higher-risk dual-use biological outputs. This adds cross-lab evidence that frontier bio capabilities are now strong enough to require capability-specific controls, while remaining indirect for bioweapon design assistance specifically. Confidence adjusted 0.75 -> 0.80. |
| 2026-06-15 | Anthropic said Mythos 5 accelerated parts of drug design by around 10x and matched or beat skilled human operators on a protein-design workflow with tools and no human assistance. This strengthens the bio-capability evidence, while remaining indirect for bioweapon design specifically. Confidence adjusted 0.70 → 0.75. |
| 2026-03 | AI Futures Project rates bio capability concerns as on track. PhD-level knowledge retrieval and synthesis capabilities continue advancing. |
| 2025-12 | AI Futures Project Dec 2025 model update confirmed bio capabilities as tracking AI 2027 prediction. Both OpenAI (“High” bio capability level) and Anthropic (ASL-3) have officially upgraded their bio risk classifications. |
| 2025-10 | Anthropic activates ASL-3 safety protocols for biological capabilities, signaling internal assessments of meaningful bio-relevant capability gains. |
| 2025-05 | Claude Opus 4 ASL-3 designation (May 22): Anthropic explicitly states the classification reflects both cybersecurity and biological risk thresholds. First formal acknowledgment by a frontier lab that their deployed model has crossed the meaningful biological uplift threshold. |
| Note on terminology: Anthropic’s AI Safety Levels (ASL 1-4) measure capability thresholds — when a model becomes dangerous enough to require additional deployment safeguards. RAND’s Security Levels (SL1-5) measure weight protection — the physical and cyber security measures applied to prevent model theft. ASL-3 activation (capability-triggered) is distinct from SL3 certification (security infrastructure). Both are tracked on this site: see Security Level Progression for the RAND framework. |