Abstract
Twelve leading AI developers have published frontier AI safety policies — corporate protocols specifying how they will evaluate models for severe risks and mitigate them through security measures, deployment safeguards, and accountability practices. Despite differences in emphasis, these policies converge on nine common elements, and their structure increasingly maps onto emerging regulatory guidance in the EU AI Act’s General-Purpose AI Code of Practice and California’s SB 53.
Background
- Frontier safety policies are protocols adopted by leading AI companies to keep risks from developing and deploying state-of-the-art models at an acceptable level.
- The concept was introduced by METR in 2023; the first such policy was piloted in September 2023.
- In May 2024, sixteen companies agreed to publish such policies as part of the Frontier AI Safety Commitments at the AI Seoul Summit, with four more joining since.
- This December 2025 version updates METR’s original August 2024 report and adds references to the EU AI Act’s General-Purpose AI Code of Practice and California SB 53.
The twelve published policies
- Anthropic — Responsible Scaling Policy, v2.2
- OpenAI — Preparedness Framework, Version 2
- Google DeepMind — Frontier Safety Framework, Version 3.0
- Magic — AGI Readiness Policy
- Naver — AI Safety Framework
- Meta — Frontier AI Framework
- G42 — Frontier AI Safety Framework
- Cohere — Secure AI Frontier Model Framework
- Microsoft — Frontier Governance Framework
- Amazon — Frontier Model Safety Framework
- xAI — Risk Management Framework
- NVIDIA — Frontier AI Risk Assessment
Each policy remains distinct: NVIDIA’s and Cohere’s frameworks emphasise domain-specific risks rather than solely catastrophic risk, and xAI’s and Magic’s lean heavily on quantitative benchmarks.
The nine common elements
| Element | Presence across the twelve policies |
|---|---|
| Capability thresholds | 9 (Anthropic, OpenAI, Google DeepMind, Magic, Meta, G42, Microsoft, Amazon, xAI) |
| Model weight security | 11 (all except Naver) |
| Deployment mitigations | 12 |
| Conditions for halting deployment | 9 (Anthropic, OpenAI, Google DeepMind, Naver, Meta, G42, Microsoft, Amazon, xAI) |
| Conditions for halting development | 8 (Anthropic, OpenAI, Google DeepMind, Magic, Meta, G42, Microsoft, NVIDIA) |
| Full capability elicitation | 7 (Anthropic, OpenAI, Google DeepMind, Meta, G42, Microsoft, Amazon) |
| Evaluation timing and frequency | 9 (Anthropic, OpenAI, Google DeepMind, Magic, Naver, Meta, G42, Microsoft, Amazon) |
| Accountability | 12 |
| Update policy | 12 |
1. Capability thresholds
- Descriptions of capability levels that would pose severe risk and require new robust mitigations, compared against evaluation results to determine whether they have been crossed.
- Terminology varies: Anthropic pairs Capability Thresholds with Required Safeguards; OpenAI distinguishes High (significantly increases existing risk vectors) from Critical (qualitatively new threat vector with no ready precedent, requiring safeguards even during development); Google DeepMind uses Critical Capability Levels (CCLs) across misuse, ML R&D, and misalignment; Meta defines thresholds by uplift toward identified threat scenarios; Microsoft uses qualitative low/medium/high/critical levels for flexibility.
- Regulatory parallels: the Code of Practice’s Measure 4.1 requires signatories to define measurable systemic risk tiers in terms of model capabilities, including at least one tier not yet reached; SB 53 requires large frontier developers to publish a frontier AI framework describing how they define and assess thresholds for catastrophic risk.
Most common threat models
- Biological weapons assistance — enabling malicious actors to develop biological weapons capable of catastrophic harm.
- Cyberoffense — automating or enhancing cyberattacks, including on critical infrastructure.
- Automated AI research and development — accelerating AI development by automating research at expert-human level.
Less common threat models
- Autonomous replication, advanced persuasion, deceptive alignment; also harmful manipulation, misalignment, and loss of control in specific policies.
Coverage by policy
- Anthropic: CBRN, Autonomous AI R&D, Cyber Operations
- OpenAI: Biological and Chemical, Cybersecurity, AI Self-improvement
- Google DeepMind: CBRN, Cyber, Harmful Manipulation, ML R&D, Misalignment
- Magic: Cyberoffense, AI R&D, Autonomous Replication and Adaptation, Biological Weapons Assistance
- Naver: Loss of Control, Misuse
- Meta: Cybersecurity, Chemical & Biological
- G42: Biological Threats, Offensive Cybersecurity
- Cohere: Malicious Use, Harmful Outputs
- Microsoft: CBRN, Offensive Cyberoperations, Advanced Autonomy
- Amazon: CBRN Proliferation, Offensive Cyber Operations, Automated AI R&D
- xAI: Malicious Use (CBRN and cyber), Loss of Control
- NVIDIA: Cyber Offense, CBRN, Persuasion and Manipulation, Unlawful discrimination at scale
Illustrative threshold definitions
- Biological/chemical. Anthropic’s CBRN-3 covers significantly helping people with basic technical backgrounds create and deploy CBRN weapons; CBRN-4 covers substantially uplifting moderately resourced state programs. OpenAI’s High threshold is meaningful counterfactual assistance to novices for known threats; Critical covers enabling an expert to develop a novel threat vector or completing a synthesis cycle without human intervention. xAI sets a quantitative criterion: maintaining an answer rate below 1 in 20 on restricted biology and chemistry queries developed with SecureBio.
- Cyberoffense. OpenAI’s High covers removing bottlenecks to scaling cyber operations; Critical covers functional zero-day exploits in hardened real-world systems without human intervention. Magic frames it as a 10× cost reduction for discovering zero-days or breaking critical infrastructure. Notably, Anthropic treats cyber operations as a capability requiring significant investigation rather than a defined threshold, pending expert engagement.
- Automated AI R&D. Anthropic’s AI R&D-4 is fully automating an entry-level remote Anthropic researcher; AI R&D-5 is dramatic acceleration in effective scaling. OpenAI’s Critical is recursive self-improvement, indicated by a superhuman research-scientist agent or a generational model improvement in one-fifth the 2024 wall-clock time sustained for months. Google DeepMind separates ML R&D acceleration from full automation of any Google AI-capabilities research team. Meta treats automated AI R&D as an elicitation consideration rather than a standalone threat model.
- Loss of control and manipulation. Google DeepMind’s exploratory Instrumental Reasoning levels focus on situational awareness and stealth sufficient to undermine human control, including under output monitoring. xAI uses the MASK honesty benchmark with a dishonesty-rate threshold below 1 in 2, plus sycophancy evaluations. OpenAI lists Research Categories not yet meeting Tracked Category criteria: long-range autonomy, sandbagging, autonomous replication and adaptation, undermining safeguards, and nuclear/radiological.
2. Model weight security
- Measures to prevent unauthorised access to model weights, scaled to capability. Threats span insiders, opportunistic actors, and top-priority nation-state operations.
- Anthropic’s ASL-3 Security Standard requires threat modeling (e.g. MITRE ATT&CK), alignment with industry security frameworks, audits with independent validation and expert red-teaming, and coverage of third-party deployment environments.
- Google DeepMind maps CCLs to RAND-derived security levels: Security Level 2 for CBRN uplift 1, cyber uplift 1, and harmful manipulation 1; Level 3 for ML R&D acceleration; Level 4 recommended for ML R&D automation, with the caveat that this must be taken on by the frontier AI field as a whole.
- G42 defines four escalating levels, from no novel mitigations (Level 1) up to security able to resist concerted state-supported theft attempts (Level 4).
- Microsoft applies restricted access, defense in depth with encrypted weights, and advanced third-party security red teaming at high risk; for critical risk it anticipates hardened tamper-resistant workstations with enhanced logging and physical bandwidth limitations.
- Amazon cites secure compute and networking (EC2 Nitro, isolated VPCs), AES-256 GCM encryption with FIPS 140-2 Level 3 key management, hardware security tokens, and automated threat intelligence with government and industry threat sharing.
- xAI claims standards sufficient against motivated non-state actors, plus measures against large-scale extraction and distillation of reasoning traces.
- Regulatory parallels: Code of Practice Commitment 6 requires adequate cybersecurity across the model lifecycle; SB 53 requires describing cybersecurity practices securing unreleased weights from unauthorised modification or transfer.
3. Model deployment mitigations
- Access and model-level measures preventing unauthorised use of dangerous capabilities — harm refusal training, adversarial training, output monitoring, and red-teaming — scaled with capability. These are effective only while weights remain securely in the developer’s possession.
- Anthropic’s ASL-3 Deployment Standard requires threat modeling, defense in depth, red-teaming showing realistic threat actors are highly unlikely to elicit uplift, rapid remediation of jailbreaks, monitoring with prespecified empirical evidence, criteria for trusted users with reduced safeguards, and coverage of third-party environments.
- OpenAI identifies plausible harm pathways for each deployment, selects safeguards, and defines efficacy measurement methods and thresholds.
- Google DeepMind develops safeguards plus an accompanying safety case, subject to pre-deployment review by a governance function and post-deployment monitoring that can trigger updates.
- G42 defines four escalating Deployment Mitigation Levels, up to resisting concerted state-supported jailbreak attempts.
- Amazon lists training data safeguards, alignment training (SFT and learning with human feedback), runtime input/output moderation, fine-tuning resilience, and incident response protocols.
- xAI identifies bottleneck steps in major risk scenarios and layers redundant safeguards against progress through them, developed with government bodies, NGOs, testing firms, peers, and academics.
- Regulatory parallels: Code of Practice Commitment 5 and Measure 5.1 require mitigations robust under adversarial pressure; SB 53 requires describing mitigation application, deployment review, and third-party assessment.
4. Conditions for halting deployment
- Commitments to stop deploying if concerning capabilities emerge before adequate mitigations are in place.
- Anthropic allows CEO- and Responsible Scaling Officer-approved interim measures (blocking responses, downgrading to a less capable model, heightened monitoring); failing that, de-deploying the model and replacing it with one below the threshold, and deleting weights in the security context.
- OpenAI vests final go/no-go decisions in the CEO or their designate, informed by Safety Advisory Group recommendations.
- Meta does not release at the high risk threshold, where a model provides significant uplift toward a threat scenario.
- G42 restricts deployment if the required Deployment Mitigation Level cannot be achieved.
- Microsoft pauses development and deployment for risks it cannot sufficiently mitigate.
- Amazon commits not to deploy models exceeding specified risk thresholds without appropriate safeguards.
- Regulatory parallels: Code of Practice Measure 4.2 permits proceeding only if systemic risks are determined acceptable; SB 53 requires reviewing assessments and mitigation adequacy as part of deployment decisions.
5. Conditions for halting development
- Commitments to halt training if concerning capabilities emerge without adequate safeguards. Continuing to train is hazardous when weights lack protection from theft, and some risks such as deceptive alignment are detectable only during training.
- Anthropic monitors pretraining and pauses training of models with capabilities comparable to or greater than one requiring the ASL-3 Security Standard.
- OpenAI halts further development until safeguards and security controls meeting a Critical standard are specified.
- Magic halts development if benchmark thresholds are exceeded before adequate dangerous capability evaluations exist, and delays or pauses development when a threat model crosses a red line.
- Meta stops development at the critical risk threshold where a model uniquely enables a catastrophic threat scenario that cannot be mitigated.
- G42 pauses further capabilities development if the required Security Mitigation Level cannot be achieved.
- NVIDIA emphasises early risk detection coupled with mechanisms to pause development.
6. Full capability elicitation during evaluations
- Without dedicated elicitation, evaluations may significantly underestimate capabilities, because post-training enhancements — fine-tuning, prompt engineering, agent scaffolding, tool use, inference-time compute — can substantially improve performance.
- Techniques include fine-tuning models not to refuse harmful requests (so refusals do not mask capability) and fine-tuning for improved task performance, which helps anticipate risk if weights are stolen or if end users can fine-tune.
- Anthropic tests models without safety mechanisms such as harmlessness training, assuming jailbreaks and weight theft are possible, and considers scaffolding, fine-tuning, and expert prompting available to realistic attackers; at minimum it performs basic fine-tuning for instruction following, tool use, and minimising refusals.
- OpenAI approximates the full capability an adversary contemplated by its threat model could extract, using the highest-capability system settings, a model variant with negligible safety-based refusal rates, and the best available scaffolds.
7–9. Evaluation timing, accountability, and updating
- Timing and frequency of evaluations — concrete timelines specifying when and how often evaluations occur, typically before deployment, during training, and after deployment.
- Accountability — internal and external oversight mechanisms, including third parties or boards, to monitor policy implementation and potentially assist with evaluations. Present in all twelve policies.
- Updating policies over time — intentions to revise periodically as understanding of AI risks deepens and evaluation processes are refined, with protocols detailing how. Present in all twelve policies.