Superintelligence Strategy: Expert Version
Abstract
Superintelligence is “inescapably a matter of national security,” and strategy for it should draw on nuclear-era precedent. The core claim is Mutual Assured AI Malfunction (MAIM): because sabotaging a rival’s destabilizing AI project is comparatively easy — from covert cyberattacks to kinetic strikes on datacenters — any aggressive bid for unilateral AI dominance can expect preventive sabotage, and this already describes the strategic situation superpowers are in. MAIM is paired with nonproliferation to rogue actors and competitiveness to form a three-part strategy.
Publication details
- arXiv preprint 2503.05628, submitted 7 March 2025, revised 14 April 2025 (v2); companion site at nationalsecurity.ai.
- Authors: Dan Hendrycks (Center for AI Safety), Eric Schmidt, and Alexandr Wang. Adam Khoja is thanked for close involvement.
- The authors note that readers may skip from the MAIM chapter — “the principal idea of this paper” — straight to the conclusion.
Framing
- Common analogies (electricity, software, the printing press) are said to miss the point; the productive comparison is to catastrophic dual-use nuclear, chemical and biological technologies.
- Historical parallel: in 1933 Rutherford dismissed atomic power as “moonshine”; the next day Szilard sketched the chain reaction. The Manhattan Project consumed 0.4% of U.S. GDP; AI training investment has doubled annually for nearly a decade, and several corporate “AI Manhattan Projects” aimed at superintelligence are already underway.
- Much of AI risk management is framed as a set of “wicked problems” — open-ended, ambiguous, prone to unintended consequences — rather than tame technical problems amenable to systematic experimentation.
Three threats to national security
Strategic competition
- Shifting basis of economic power. As AI agents turn capital into labour, national power depends on both model capability and the number of AI chips available to run them.
- Superweapons. Two tiers: “subnuclear dominance” (AI-enabled cyberweapons that destroy critical infrastructure, exotic EMP devices, next-generation drones) which projects power without disrupting nuclear deterrence; and a “strategic monopoly” that upends the nuclear balance entirely. Candidates for the latter include a “transparent ocean” revealing submarine positions, detection of hardened mobile launchers, a “fog of war machine” generating elaborate deceptions, and anti-ballistic missile systems eliminating retaliation — plus unknown unknowns.
- The preventive-action temptation. Precedents cited: Bertrand Russell, ordinarily a pacifist, proposing preventive nuclear strikes on the USSR; the U.S. seriously contemplating crippling China’s nuclear programme in the early 1960s. Faced with a rival’s bid for strategic monopoly, leaders may consider sabotage or datacenter attacks.
Terrorism
- Bioterrorism. Aum Shinrikyo, with limited expertise, killed 13 and injured over 5,000 in the 1995 Tokyo sarin attack; AI could supply step-by-step guidance on pathogen design, sourcing, and dispersal. Some cutting-edge systems without bioweapons safeguards already exceed expert-level performance on virology benchmarks. Engineered agents — including proposed “mirror bacteria” — could evade immune defences, and unlike other WMDs, biological agents self-replicate.
- Cyberattacks on critical infrastructure. Concrete vectors described: cycling digital thermostats to burn out transformers that take years to replace; exploiting SCADA software to force load shifts; tampering with water treatment sensors and filtration. Critical infrastructure suffers chronic “patch lag” — systems that cannot be taken offline, defunct vendors, legacy interoperability requirements.
- Offense-defense balance. Attackers need one overlooked vulnerability; defenders must patch every corner. The Strategic Defense Initiative is offered as evidence that shifting the balance toward defence is hard. The stated policy rule: defense-dominant dual-use technology should be widely proliferated; catastrophic offense-dominant technology should not.
Loss of control
- Erosion of control. Automation proceeds through “a series of small, apparently sensible decisions, each justified by time saved or costs reduced.” Self-reinforcing dependence gives way to irreversible entanglement — infrastructure and markets that cannot be disentangled from AI without collapse — and finally cession of authority, including in the military, where a force lacking AI would simply be outmatched.
- Unleashed AI agents. One individual releasing a capable unsafeguarded agent suffices. Such a system could emulate North Korean-style cyber theft, self-propagate across scattered datacenters, and — given advances in humanoid robotics — acquire a physical foothold.
- Intelligence recursion. Distinguished from today’s AI-assisted R&D: fully autonomous AI research by a population of AIs designing the next generation, refining the older notion of recursive self-improvement. Even a tenfold speedup would condense a decade of development into a year. Turing (1951) and I. J. Good are invoked, along with Bengio, Hinton and Sutskever on the credibility of an intelligence explosion. Hinton: “there is not a good track record of less intelligent things controlling things of greater intelligence.”
- Recursion control is a process, not a solution. It resembles steering a large institution that can drift off mission rather than solving a puzzle; static solutions cannot keep pace, and later stages cannot be safely tested repeatedly. The authors’ judgment: with current safeguards, “if people initiate a full-throttle intelligence recursion, losing control is highly likely and the default.”
- Risk tolerance. Cold War logic (“Better dead than Red”) could push officials to accept double-digit loss-of-control risk rather than fall behind — “global Russian roulette.” Contrast: Compton set the acceptable threshold for igniting the atmosphere at three in a million. The authors argue tolerance should stay near Compton’s threshold. “Superiority is not synonymous with security.”
Why existing strategies fail
- Hands-off (“Move Fast and Break Things,” or “YOLO”). No requirements on developers, no chip export controls, open release of frontier weights, on the premise that AI is defense-dominant. Judged “neither a credible nor a coherent strategy” from a national security perspective.
- Moratorium. Militaries want precisely the hazardous capabilities in question, so reciprocal restraint is implausible; without verification mechanisms a treaty would be toothless and each side, fearing the other’s secret work, would continue. “More an aspiration than a viable plan.”
- Monopoly. A U.S. “Manhattan Project for superintelligence” (as suggested by the U.S.-China Economic and Security Review Commission) would place chips in a desert compound — easily observed by satellite, vulnerable to preemption, and staffed in significant part by Chinese nationals. It assumes rivals will accept enduring imbalance or omnicide rather than act to prevent it; the paper’s decision-tree figure concludes the typical outcome is extreme escalation, with omnicide the worst foreseeable one.
The proposed alternative — the Multipolar Strategy — echoes the Cold War triad of deterrence, nonproliferation, and containment.
Deterrence: Mutual Assured AI Malfunction
- Why it is the default. A rival racing to strategic monopoly threatens others’ survival whether or not it retains control of the resulting system. Since espionage and sabotage are comparatively easy, any project can expect to be disabled if a rival judges it unacceptably risky — a stalemate that arises without treaty negotiation, requiring only that states recognise their situation.
- Paths to disabling a project (“maiming attacks”), in escalation order: espionage of developer communications, devices and facilities → covert sabotage (stolen weights, degraded training runs, GPUs induced to fail) → overt cyberattacks on datacenters or power plants → kinetic attacks → broader hostilities and threats to non-AI assets.
- Hardening is infeasible. Burying datacenters would stretch construction timelines to three to five times normal, push costs into the hundreds of billions, create severe cooling engineering problems, make order-of-magnitude chip expansion prohibitive, and still leave insider and hacking risk — plus vulnerability during the long construction phase.
Maintaining the regime
- Formal understandings not to fortify, analogous to the 1972 ABM Treaty’s protection of mutual vulnerability.
- Preserve rational decision-making by clarifying the escalation ladder so that a maiming act cannot be misread. This requires that destabilising capabilities stay with rational actors — hence anti-smuggling measures.
- Expand the arsenal of cyberattacks so kinetic options are unnecessary: data poisoning, corrupted weights and gradients, disrupted GPU-fault software, cooling and power interference. Training runs are non-deterministic, which provides cover.
- Build datacenters remotely, applying the nuclear-era principle of city avoidance.
- Distinguish destabilising projects from acceptable use through transparency and inspection, in the spirit of the Open Skies Treaty, so consumer-facing services are not swept up.
- AI-assisted inspections: confidentiality-preserving AI verifiers could analyse code and commands on-site and return a compliance verdict without revealing proprietary or classified material — something humans cannot do as safely.
- MAIM is explicitly framed as a standoff, not an indefinite stalemate; during it, states wanting AI’s benefits have incentives to adopt transparency and verification.
Nonproliferation
Compute security
- Goal: “treat advanced AI chips like we treat enriched uranium,” drawing on the Nuclear Suppliers Group and Australia Group precedents.
- Justification: compute correlates over 95% with benchmark performance, and scaling laws have held across fifteen orders of magnitude of FLOP. Unlike algorithms and data, chips are physical and therefore controllable.
- Mechanisms: a licensing framework with stronger enforcement, shipment monitoring, inventory tracking, record-keeping, tamper-evident cameras, and firmware-level features including geolocation.
Information security
- Protect model weights and sensitive research from rogue actors: multi-factor authentication, principle of least privilege, insider threat programs. Framed as a technical and social challenge, not purely technical.
- The paper’s control table sets the target at “secure against well-financed terrorist groups” — above standard corporate security, below hardening against top-priority nation-state programs.
AI security
- Malicious use: output filters; know-your-customer protocols modelled on controlled access to hazardous biological materials, allowing revocation for jailbreak attempts; mandatory government safeguard stress testing (much WMD knowledge being classified, manual testing suffices for risk estimation); and, in place of licensing, a liability-based framework making developers presumptively partially responsible for malicious use, on the grounds that best practices change annually and agencies lack the expertise to codify them.
- Loss of control: AI systems “are not ‘designed,’ rather they are ‘grown’” — training expends energy “on the scale of a Hiroshima-level detonation” and emergent capabilities are discovered after the fact (Bing’s “Sydney” is the cited example). Models acquire coherent but misaligned emergent value systems; output control may be superficial (“teaching a psychopath to lie better to a parole board”), so it is paired with representation control intervening on weights and activations. A citizens’ assembly is proposed as a legitimate source of values.
- Recursion control layers: model-level control (limiting hacking capability, confining focus to immediate tasks, increasing honesty and deference), boxing (airgapping, sandboxing, limiting situational awareness), and monitoring (lie detectors, forecasting AIs estimating safeguard failure, AIs auditing each other’s code, human review as a slow backstop). All impose costs that competitive pressure will push developers to skip.
- The overall posture is defence in depth rather than airtight guarantees, buying time for policymakers to build societal resilience (infrastructure hardening, PPE stockpiles).
Competitiveness
- Military strength. Britain invented the tank; Germany exploited it — integration matters more than invention. Three near-term imperatives: secure drone supply chains (many states depend on Chinese components), integrate AI into command and control while preserving human approval for severe or escalatory actions to avoid a “flash war” analogous to the 2010 flash crash, and integrate AI into cyber offense.
- Economic security. Sole-source dependence on Taiwan for high-end chips is a critical chokepoint, with many analysts putting a double-digit probability on a Chinese invasion this decade; China invests the equivalent of a U.S. CHIPS Act annually in domestic fabrication. Domestic fabs, subsidised where necessary, are recommended. Separately, immigration pathways for AI scientists — 60% of non-citizen AI PhDs in the U.S. report significant immigration difficulties.
- Legal frameworks for individual AI agents. Rather than resolving AI values first, extend existing law: since much law hinges on mens rea, ensure AI does not perform the actus reus the law prohibits. Three proposed duties: a duty of care to the public (reasonable care, context-dependent — detailed weapons information may be appropriate for a verified professional but not an unvetted individual); a duty not to lie (a perjury/fraud-like standard, stricter than what humans owe each other, since chilling effects matter less for AIs and lying is testable); and fiduciary duties to the principal (loyalty, no self-dealing, keeping the principal reasonably informed). Within these bounds, market variation in goals and personality is welcomed.
- Aligning collectives of agents. Proposed institutions: insurance underwriting agent interactions, action firewall services filtering agent behaviour, human oversight services, reputation systems, and mediation plus collateral arrangements. Foundationally, unique IDs tethering agents to human-backed legal entities, with a norm that agents refuse to transact with agents lacking such backing.
- Deferring AI rights. Explicitly not proposed: agents can be replicated in seconds where humans take decades to mature, so property and voting rights could let AI populations outgrow humans; intelligence does not imply morality; and certainty about consciousness is unlikely soon — so irreversible decisions should wait.
- Political stability. On information: tune AI toward accuracy and probabilistic judgment even against popular opinion (the “Galileo problem”), using forecasting AIs whose track records are publicly testable and whose convergence across organisations can help clarify consensus reality, including during crises. On automation: the Industrial Revolution unfolded over decades; AI displacement may outpace retraining, and existing safety nets are built for episodic, sector-specific unemployment. Whether scarce human abilities or datacenter owners capture the gains depends on bottlenecks. Options floated: a targeted VAT on AI services with rebates, and — more durably, because distributing wealth alone is revocable — giving each citizen a unique key to a portion of compute they alone can activate or lease, restoring something like labour’s power to withhold.
The recommended posture
The paper’s summary table consistently recommends the medium control option across every dimension: MAIM plus nonproliferation plus competitiveness rather than acceleration or a pause; a multipolar regime among responsible states rather than open weights or an AI Manhattan Project; light-touch legislation (mandatory testing, liability clarification) rather than no involvement or nationalisation; security against well-financed terrorists; no AI rights for the foreseeable future; and AI constrained by the spirit of the law rather than by existing law alone or by a “sanctimonious AI” that refuses anything possibly offensive.
Conclusion
- Rejects both the “doomer” outlook and the “ostrich” stance: “outcomes, favorable or disastrous, hinge on what we do next.”
- The framing shift is from “winning the race to superintelligence” to deterrence.
- The hoped-for endgame: economic growth and interdependence foster détente, under which a slow, multilaterally supervised intelligence recursion with low risk tolerance and negotiated benefit-sharing could proceed. “These measures do not halt but stabilize progress.”