Leading AI lab progresses toward RAND SL3 security

Last updated
OpenBrain's security level is typical of a fast-growing ~3,000 person tech company (RAND's SL2). They are working hard to protect their weights and secrets from insider threats and top cybercrime syndicates (SL3).

At a glance

  • Assessment: Emerging
  • Confidence in assessment: 60%
  • Predicted timing: Mid-to-late 2026 (WSL 3.0); SSL 3.0 not until ~Apr 2027
  • Primary source: ai-2027.com, Early 2026: Coding Automation

What AI 2027 Predicted

The scenario describes a specific security progression at the leading AI lab. In early 2026, the lab’s security is at RAND’s Security Level 2 (SL2) — typical of a fast-growing tech company — and they are actively working to reach SL3, which involves protecting model weights from “insider threats and top cybercrime syndicates.” The RAND framework defines security levels as:

  • SL2: Protects against opportunistic attackers and basic cybercrime
  • SL3: Protects against sophisticated persistent threats, insider threats, and top cybercrime organizations
  • SL4: Protects against nation-state actors (intelligence agencies)

The implication is that in early 2026, labs are vulnerable to nation-state theft (which plays out later in the scenario when China steals model weights).

How We Track This

We monitor:

  • RAND security level assessments and recommendations
  • AI lab security disclosures and commitments
  • Third-party security audits of frontier labs
  • Cybersecurity incidents at AI companies
  • AI Safety Institute evaluations of lab security practices
  • Lab security commitments in safety policies (RSPs, Preparedness Frameworks)

Current Evidence

Public disclosures describe controls relevant to RAND’s threat categories, but do not establish a verified security level for a particular lab:

RAND framework adoption: RAND published its “Playbook for Securing AI Model Weights” (May 2024) defining the SL1–SL5 framework. LessWrong analysis noted that lab commitments range from “generally aligned with RAND SL2” to “RAND SL4,” with practice characterized by that analysis as SL2–SL3. The analysis concluded “AI companies aren’t planning to secure critical model weights” at levels commensurate with the threat.

Anthropic espionage report: Anthropic described attackers using Claude Code against external organizations. The report concerns offensive AI use; it does not establish a breach of Anthropic’s own systems or determine the lab’s RAND security level. Anthropic incident report.

Lab security investments: Both OpenAI and Anthropic have publicly committed to “continuing to invest in cybersecurity and insider threat safeguards to protect proprietary and unreleased model weights.” VentureBeat (Dec 2025) reported on the growing obsession with model weight security at both companies.

AI 2027’s own security forecast acknowledges that “we expect security to be less of a priority through 2025” given that DeepSeek pushed the open-source frontier closer to US labs, reducing the marginal value of weight theft.

Infrastructure and security requirements: Large infrastructure projects and defense partnerships create security requirements. Their scale or existence alone does not establish a RAND SL3 assessment; the relevant controls and threat model require separate evidence.

Safety policy erosion: TIME reported (Mar 2026) that Anthropic dropped its flagship safety pledge, noting the company may release future models “without ironclad safety guarantees.” This raises questions about whether security investments keep pace even as safety commitments weaken.

OpenAI model-weight security controls: OpenAI’s GPT-5.6 Preview System Card provides current first-party evidence that at least one leading lab is investing in security controls that map directionally toward RAND SL3 concerns. OpenAI describes defense-in-depth for model weights, including access control, infrastructure hardening, egress controls, monitoring, purpose-built detections for exfiltration of high-risk weights, dedicated Detection and Response and Threat Intelligence teams, and Insider-Risk programs. This is not an independent RAND SL3 assessment. It directly addresses insider threat and sophisticated exfiltration concerns central to progression beyond SL2.

Sources:

Counterevidence & Limitations

  • This review did not identify an independent public assessment resolving the relevant lab’s RAND level
  • Labs may be further along in security upgrades than their public statements suggest (security improvements are often kept quiet for obvious reasons)
  • The Anthropic espionage report does not measure the security of Anthropic’s own model weights. Offensive capability, a lab’s defensive controls and a RAND security assessment must be evaluated separately.
  • The RAND framework is advisory, not regulatory — labs adopt it voluntarily and may interpret levels differently
  • DeepSeek’s open-source releases arguably reduce the strategic value of weight theft, potentially slowing security investment

What Would Change Our Assessment

  • Upgrade to “on-track”: A major lab publicly claims or is independently assessed at SL3; government mandates SL3+ for frontier models
  • Upgrade to “confirmed”: A qualifying leading lab independently verified at SL3 by RAND or equivalent third-party assessment
  • Downgrade to “behind”: A successful large-scale weight theft demonstrating SL2-level vulnerability persists; labs publicly acknowledge they haven’t reached SL3

Update History

DateUpdate
2026-09-06Current evidence and assessment explanation reviewed. Clarified measurement scope, source interpretation and limitations.
2026-09-06Corrected the external-target espionage report’s classification and removed unsupported SL3 inferences from infrastructure scale. These sources do not independently establish the lab’s security level. Status and confidence were unchanged.
2026-06-29OpenAI’s GPT-5.6 system card described model-weight security controls including infrastructure hardening, egress controls, monitoring, exfiltration detections, threat intelligence, and insider-risk programs. This supports the view that leading labs are moving toward SL3-relevant protections, though no independent RAND SL3 certification is public. Confidence adjusted 0.55 -> 0.60.
2026-03Labs appear in SL2–SL3 transition. Significant security investments underway but no public SL3 certification achieved yet.
2025-12Anthropic espionage incident validates the threat model for insider threats at AI labs.
2025-05Anthropic designates Claude Opus 4 as the first ASL-3 model (May 22). ASL-3 is defined as the threshold where a model could provide meaningful uplift toward weapons of mass destruction. This is the first time a frontier model has crossed this threshold with binding operational consequences — enhanced security and deployment requirements now apply.