AI provides substantial bioweapon design help

Last updated
On Track · Security · 80% confidence in assessment
Predicted: Late 2025 · Updated: 2026-09-07 · Source: ai-2027.com, Late 2025: The World's Most Expensive AI
It could offer substantial help to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web.

Controlled 2025 studies support substantial planning assistance. Laboratory execution evidence is mixed; assistance does not establish successful weapon production.

At a glance

  • Assessment: On Track
  • Confidence in assessment: 80%
  • Outcome: unresolved
  • Timing: Original window passed; outcome unresolved
  • Evidence: proxy
  • Predicted timing: Late 2025
  • Primary source: ai-2027.com, Late 2025: The World's Most Expensive AI

What AI 2027 Predicted

By late 2025, frontier models can substantially assist biological-weapons design. This is an assistance claim, not a prediction that a weapon has been successfully produced or used. AI 2027

How We Track This

Examine controlled comparisons of model-assisted work with appropriate baselines. Distinguish planning and knowledge from physical execution, and model capability from the effectiveness of deployed safeguards.

Current Evidence

Anthropic’s September 5, 2025 report describes controlled planning trials comparing unsafeguarded Claude 4 assistance with internet-only access. Expert graders found substantially better plans and fewer critical failures, especially with Opus 4. This supports significant assistance within the forecast window. The older, weaker Sonnet 3.7 finding does not summarize these later trials. Anthropic biorisk research

Subsequent company evaluations continue to identify serious capability concerns. OpenAI’s September 3 Astra assessment rates biological and chemical capability High, rather than Critical. Such classifications are the developer’s risk judgments, not demonstrations of successful physical harm. Astra system card

Counterevidence & Limitations

A preregistered, blinded randomized trial of 153 novices, conducted in June-August 2025 and published in February 2026, found overall workflow completion of 5.2% with LLM assistance versus 6.6% with internet alone (P = .759). This is important contrary evidence for novice laboratory execution. It does not directly refute a separate planning benefit. Hong et al.

The positive planning study is developer-produced and does not measure the complete route to a harmful outcome. Deployment precautions likewise do not prove that a threshold was crossed. On-track reflects substantial support for the assistance claim, with its exact breadth unresolved after the deadline.

What Would Change Our Assessment

  • Strengthen: Independent, controlled evaluations reproduce substantial assistance across relevant high-level capability measures.
  • Confirm: Comparable evidence establishes the specified assistance level within the original period.
  • Weaken: Better-designed evaluations show that planning gains fail to transfer to consequential task performance.

Update History

DateUpdate
2026-09-07Recovered the stronger Claude 4 planning trials from September 2025 and compared them with the independent novice-execution trial. Substantial assistance is supported; complete physical execution remains unestablished.
2026-09-06Assessment revised from on-track (0.80) to emerging (0.70). Precautionary safeguards do not confirm a capability threshold. Direct evaluations offer mixed evidence across expert assistance and novice outcomes.
2026-09-06Clarified that precautionary ASL-3 activation is not by itself confirmation of a biological capability threshold. Direct evaluations and deployment safeguards are assessed separately; status and confidence retained pending review.
2026-06-29OpenAI classified GPT-5.6 Sol, Terra, and Luna as High capability in Biological and Chemical risk and described trusted-access pathways for higher-risk dual-use biological outputs. This adds cross-lab evidence that frontier bio capabilities are now strong enough to require capability-specific controls, while remaining indirect for bioweapon design assistance specifically. Confidence adjusted 0.75 -> 0.80.
2026-06-15Anthropic said Mythos 5 accelerated parts of drug design by around 10x and matched or beat skilled human operators on a protein-design workflow with tools and no human assistance. This strengthens the bio-capability evidence, while remaining indirect for bioweapon design specifically. Confidence adjusted 0.70 → 0.75.
2026-03AI Futures Project rates bio capability concerns as on track. PhD-level knowledge retrieval and synthesis capabilities continue advancing.
2025-12AI Futures Project Dec 2025 model update confirmed bio capabilities as tracking AI 2027 prediction. Both OpenAI (“High” bio capability level) and Anthropic (ASL-3) have officially upgraded their bio risk classifications.
2025-10Anthropic activates ASL-3 safety protocols for biological capabilities, signaling internal assessments of meaningful bio-relevant capability gains.
2025-05Claude Opus 4 ASL-3 designation (May 22): Anthropic explicitly states the classification reflects both cybersecurity and biological risk thresholds. First formal acknowledgment by a frontier lab that their deployed model has crossed the meaningful biological uplift threshold.
Note on terminology: Anthropic’s AI Safety Levels (ASL 1-4) measure capability thresholds — when a model becomes dangerous enough to require additional deployment safeguards. RAND’s Security Levels (SL1-5) measure weight protection — the physical and cyber security measures applied to prevent model theft. ASL-3 activation (capability-triggered) is distinct from SL3 certification (security infrastructure). Both are tracked on this site: see Security Level Progression for the RAND framework.