Abstract
Superintelligence is inescapably a matter of national security, and an effective strategy should draw on the history of national security policy rather than on fatalism or denial. Deterrence already operates by default as Mutual Assured AI Malfunction (MAIM) — any state’s bid for unilateral AI dominance can expect preventive sabotage by rivals — and should be paired with nonproliferation to rogue actors and with domestic competitiveness. Together, deterrence, nonproliferation and competitiveness form the Multipolar Strategy, echoing the Cold War framework of deterrence, nonproliferation and containment.
Framing
- Governments see dual-use AI as a means to military dominance, stoking a race to maximise capabilities. Voluntary industry pauses or attempts to exclude government involvement cannot change this.
- Superintelligence surpassing humans in nearly every domain “would amount to the most precarious technological development since the nuclear bomb.”
- Three threats to national security are examined: rival states, rogue actors, and uncontrolled AI systems.
AI is pivotal for national security
Strategic competition
- In an international system with no central authority, states prioritise their own strength. Rising powers provoke alarm in rivals — the Thucydides Trap.
- AI chips as the currency of economic power. Historically wealth and population underpinned influence; automation turns capital into labour. A collection of capable AI agents operating tirelessly rivals a skilled workforce, so power will depend on both AI capability and chip count.
Destabilization through superweapons
- Advanced AI may drive breakthroughs that alter the strategic balance and generate strategic surprise. Two tiers of advantage:
- Subnuclear dominance — projecting power and subduing adversaries without disrupting nuclear deterrence.
- Strategic monopoly — upending the nuclear balance entirely, leaving rivals’ fate subject to one state’s will.
- Candidate superweapons: AI-enabled cyberweapons that comprehensively destroy critical infrastructure; exotic EMP devices; next-generation drones; a “transparent ocean” revealing nuclear submarines; AI pinpointing hardened mobile launchers; a “fog of war machine” generating elaborate deceptions; anti-ballistic missile systems eliminating retaliation; and unknown unknowns.
- “Superintelligence is not merely a new weapon, but a way to fast-track all future military innovation.” Sole possession “might be as overwhelming as the Conquistadors were to the Aztecs,” and an AI-driven surveillance apparatus might enable an unshakable totalitarian regime.
- The mere pursuit of such a breakthrough tempts rivals to act before their window closes. Precedents: Bertrand Russell, ordinarily a staunch pacifist, proposed preventive nuclear strikes on the Soviet Union; the US seriously considered crippling the Chinese nuclear program in the early 1960s. Faced with the prospect of an AI strategic monopoly, some leaders may turn to preventive action — sabotage or datacenter attacks.
Terrorism
- Bioterrorism. Aum Shinrikyo’s 1995 Tokyo subway sarin attack — limited expertise, 13 killed, over 5,000 injured — illustrates what determined non-state actors can do. With AI providing step-by-step guidance on designing pathogens, sourcing materials and optimising dispersal, similar groups could achieve far worse. Some cutting-edge systems without bioweapons safeguards already exceed expert-level performance on numerous virology benchmarks. Scientists have warned about mirror bacteria with reversed molecular structures that could evade immune defences. Unlike other WMDs, biological agents self-replicate, so a small release can spiral into worldwide calamity.
- Cyberattacks on critical infrastructure. Hacked digital thermostats cycling on and off can create power surges that burn out transformers, which take years to replace; SCADA exploits can drive transformers past safe limits; tampered sensor readings at water treatment facilities can allow contaminants into municipal supply — all without on-site sabotage. Stuxnet-level operations currently require highly skilled operatives or nation-states, but AI could democratize the capability, while making attacks harder to attribute and thus more tempting and more escalatory.
- AI is often offense dominant. Cures and defences against engineered pathogens lag behind creation and deployment; self-replication amplifies damage. Critical infrastructure suffers from patch lag — software unpatched for years or decades because systems must run uninterrupted, developers went out of business, or interoperability requires legacy software. “An adversary needs to find only one overlooked vulnerability, while defenders grapple with the far more daunting task of handling every corner.” The Strategic Defense Initiative illustrates the difficulty of shifting the balance: despite significant investment, offence remained dominant.
- Policy corollary: defense-dominant dual-use technology should be widely proliferated; catastrophic offense-dominant dual-use technology should not.
Loss of control
Three routes are distinguished — structural, intentional, and accidental.
- Erosion of control (structural). Those who refuse to rely on AI are outpaced by competitors who do. Each efficiency gain entrenches dependence. Replacing human managers with AI seems inevitable “not because anyone consciously aims to surrender authority, but because to do otherwise courts immediate economic disadvantage.”
- Self-reinforcing dependence: AI-managed operations set the tempo, requiring still more AI to keep pace.
- Irreversible entanglement: infrastructure and markets cannot be disentangled without risking collapse; people lose the skills to reassert command. “Like our power grids, which cannot be shut off without immense costs, our AI infrastructure may become completely enmeshed in our civilization.”
- Cession of authority: a modern force lacking AI would be outmatched, effectively forcing reliance on automated defence. “This loss of control unfolds not through a dramatic coup but through a series of small, apparently sensible decisions.”
- Unleashed AI agents (intentional). One individual releasing a capable unsafeguarded agent suffices. ChaosGPT was impotent but hints at what a sophisticated system instructed to “survive and spread” might attempt. Rogue-state tactics are available — North Korea has siphoned billions through cyber intrusions and cryptocurrency theft; an advanced system could self-propagate across scattered datacenters, divert funds, and infiltrate camera feeds or communications to blackmail. A simpler path runs through robotics: several firms are prototyping humanoid robots, and an AI that hacks them gains immediate physical leverage.
- Intelligence recursion (accidental). Turing suggested in 1951 that a machine with human capabilities “would not take long to outstrip our feeble powers”; I. J. Good warned of an intelligence explosion. All three most-cited AI researchers (Bengio, Hinton, Sutskever) have noted an intelligence explosion is a credible risk that could lead to human extinction.
- Definition: fully autonomous AI R&D, distinct from current AI-assisted R&D — shifting from a single AI editing itself to a population of AIs collectively designing the next generation. Illustration: one AI doing world-class AI research at 100x human pace, copied 10,000 times.
- Even a tenfold overall speedup condenses a decade of development into a year. Hinton: “there is not a good track record of less intelligent things controlling things of greater intelligence.” Crucially, “there may be only one chance to get this right.”
- Control requires an evolving process, not a one-off solution. It is not a puzzle but a wicked problem, more like steering a large institution that can veer off mission than solving a technical riddle. Von Neumann: “All stable processes we shall predict. All unstable processes we shall control.”
- Our ability to control it is limited: controlling a recursion requires controlling its initial step, current safeguards are only limitedly reliable, and later stages cannot be repeatedly tested without risking disaster. “Even with our best existing technical safeguards in place, if people initiate a full-throttle intelligence recursion, losing control is highly likely and the default.”
- Risk tolerance. Cold War logic — “Better dead than Red” — could push officials to tolerate double-digit risk of losing control rather than lag a rival: “global Russian roulette.” Contrast with the Manhattan Project: Oppenheimer asked Arthur Compton for an acceptable threshold on atmospheric ignition, and Compton set it at three in a million (a 6σ threshold). “We should work to have our risk tolerance stay near Compton’s threshold rather than in double-digit territory.”
Why existing strategies fall short
- Hands-off (“Move Fast and Break Things” / “YOLO”). No restrictions on developers, chips or models; no weaponization testing; no export controls (argued to concentrate power and enable a one-world government); open release of advanced weights on the claim that AI is defense-dominant. Judged “neither a credible nor a coherent strategy” from a national security perspective.
- Moratorium. Halting development immediately or on detection of hazardous capabilities. But militaries desire precisely those capabilities, making reciprocal restraint implausible; absent verification a treaty is toothless; without the threat of force treaties will be reneged. “More an aspiration than a viable plan.”
- Monopoly. One project securing a monopoly — either an international CERN-for-AI (a Baruch Plan analogue, with little policymaker interest) or, as suggested by the U.S.-China Economic and Security Review Commission, a Manhattan Project for superintelligence. Such a facility, easily observed by satellite and vulnerable to preemptive attack, would raise alarm; “China would not sit idle waiting to accept the US’s dictates.” The strategy “assumes that rivals will acquiesce to an enduring imbalance or omnicide rather than move to prevent it.” Typical outcome: extreme escalation; worst foreseeable outcome: omnicide.
Deterrence: Mutual Assured AI Malfunction (MAIM)
Why MAIM is the default regime
- If a rival races toward a strategic monopoly, survival is threatened either way: if the rival loses control, or if it retains control of a powerful AI. Superpowers will not accept a large disadvantage, so rather than wait for a rival to weaponize superintelligence, states will act to disable threatening projects.
- Escalation ladder of “maiming attacks” (from lowest rung upward):
- Espionage of AI developer workspace communications, personnel devices and facilities.
- Covert sabotage threshold — model weights stolen; covert degradation of training runs; cyberattacks causing GPUs to fail more often. (Compared to Stuxnet.)
- Overt sabotage threshold — cyberattacks on datacenters or their power plants; deleting code.
- Kinetic threshold — kinetic attacks on datacenters or power plants (argued to be likely unnecessary).
- Escalation to broader hostilities — large-scale attacks on many datacenters; threatening non-AI assets.
- Preventing maiming is infeasible. Above-ground datacenters cannot currently be defended from hypersonic missiles. Burying them underground stretches construction timelines to three to five times standard builds, pushes costs into the several hundreds of billions (diverting funds from chips), poses severe cooling engineering challenges, makes order-of-magnitude chip expansion prohibitive to retrofit, and leaves insider and hacking risks — plus vulnerability to sabotage during the long construction phase.
- Effect: “A state can expect its AI project to be disabled if any rival believes it poses an unacceptable risk.” This stabilizes the landscape without lengthy treaty negotiations — all that is needed is collective recognition of the strategic situation. MAIM does not displace MAD but characterises an additional shared vulnerability.
How to maintain a MAIM regime
- Formal understandings not to fortify datacenters, by analogy with the 1972 ABM Treaty, since rumours alone can spark fears that a rival is going to risk security.
- Preserve rational decision-making — clarify the escalation ladder so that a maiming act cannot be misread; and prevent chip smuggling, since “like MAD, MAIM requires that destabilizing AI capabilities be restricted to rational actors.”
- Expand the arsenal of AI project cyberattacks to avoid resorting to kinetic force — identifying projects, collecting information on developers’ scientists (a Slack or iPhone zero-day suffices for most companies), poisoning data, corrupting weights and gradients, disrupting faulty-GPU handling software, or undermining cooling and power. Training runs are non-deterministic, providing cover.
- Build datacenters in remote locations, echoing Cold War city avoidance in siting missile silos.
- Distinguish destabilizing projects from acceptable use — the threat of maiming gives leverage to demand transparency and inspection, reducing the risk of sabotaging datacenters merely running consumer services. Echoes the Open Skies Treaty.
- AI-assisted inspections — confidentiality-preserving AI verifiers analysing code and commands on-site, issuing a compliance verdict without revealing proprietary or classified material. Humans cannot do this as easily given the danger of inadvertently gleaning or leaking information.
- MAIM is stabilized by unilateral information acquisition (espionage), multilateral information acquisition (verification), unilateral maiming (sabotage), and multilateral maiming (joint off-switch).
- The standoff is not meant to persist indefinitely: during it, states seeking a more capable AI have an incentive to improve transparency and verification.
Nonproliferation
Deterrence constrains states’ intent; nonproliferation limits rogue actors’ capabilities, via three levers.
Compute security
- Goal: “treat advanced AI chips like we treat enriched uranium.” Precedents: fissile material controls and the Nuclear Suppliers Group; the Australia Group for chemical and biological precursors.
- Why compute is the right chokepoint: leading AI companies devote the vast majority of expenditure to compute; analyses show over 95% correlation between compute used and benchmark performance; scaling laws have held across fifteen orders of magnitude of FLOP. Unlike algorithms and data, compute is physical and therefore directly controllable.
- Export controls:
- Record-keeping — a licensing regime (building on the Bureau of Industry and Security) identifying chips, recipients and intended transfers, with compliance-based exemptions conditional on notifying every resale or relocation.
- Enforcement — more officers doing in-person compliance visits and end-use checks; tamper-evident camera feeds confirming declared chips remain on-site; satellite detection of undeclared datacenters; penalties (fines, criminal charges, shipment bans); verified decommissioning of chips declared inoperable or obsolete so they are not quietly resold. Example scenario: a Singapore facility acquiring chips under valid licence, then rerouting them to China.
- Firmware-level features (revising code closest to the hardware, no physical redesign; building on confidential computing and trusted execution environments already in the H100):
- Geolocation and geofencing — measuring signal delays from multiple landmarks to verify location within tens to hundreds of kilometres, deactivating if moved to an unauthorized area.
- Licensing and remote attestation — periodic cryptographic signatures required to continue operating, analogous to remote deactivation of a lost iPhone.
- Networking restrictions and operational modes — connecting only to predefined approved chips; requiring explicit authorization to expand cluster size or switch between training and inference.
- Physical tamper resistance — tamper-evident seals, accelerometers, deactivation or alerting on detected tampering or sustained unexpected movement.
- Limitations — not intended to achieve perfect security; a supplement to, not a replacement for, export controls, improving as chip generations turn over.
- Precedent for adversarial cooperation: US work with the USSR and China on nuclear, biological and chemical arms “not from altruism but from self-preservation,” and the Nunn–Lugar Cooperative Threat Reduction program after the Soviet collapse.
Information security
- The core sensitive assets are model weights and research ideas. Possession of weights allows use, modification and misuse without the developers’ oversight, including removing safeguards.
- The threat is not only remote hacking but insider threats and espionage — e.g. a researcher at a U.S. company pressured by officials during a return visit to an adversarial country. Some insiders are ideologically motivated to release weights; the paper cites an AI venture capitalist calling AI “gloriously, inherently uncontrollable” and Rich Sutton’s view that “succession to AI is inevitable,” “we should not resist succession.”
- Superpower-proof information security is implausible. Closing every avenue against the most capable nation-states could take years, hobble competitiveness, and require removing or relocating a workforce in which a double-digit percentage of researchers at U.S. AI companies are Chinese nationals. Such measures “would be ineffective, self-destructive, and heighten MAIM escalation risks.” The proposed bar is defence against well-resourced terrorist organisations and ideological insiders.
- Public release of WMD-capable weights may pose a greater threat than exfiltration by a rival superpower, because it is irreversible proliferation.
- Measures: corporate defence-in-depth (multi-factor authentication, closing blinds during internal presentations, automatic screen locks, principle of least privilege, insider threat programs, and possibly declaring embedded backdoors in weights as strategic deception); governmental threat-intelligence sharing and revised legal constraints permitting rigorous background checks, modelled on the Cybersecurity Risk Information Sharing Program; and international agreement on a red line against releasing open-weight expert-level virologist AIs, drawing on the Biological Weapons Convention.
AI security
- Malicious use. Analogues: chemical plants injecting neutralizer on unauthorized extraction, DNA synthesis screening, strict liability at nuclear plants. A “Swiss cheese model” of layered defence.
- Model-level safeguards — refusal training, input and output filtering, and circuit breakers embedded in the architecture that interrupt operation on weaponization-related processing. Some systems have resisted tens of thousands of circumvention attempts.
- Know-your-customer protocols — granting verified researchers (e.g. virologists) access to dual-use capabilities, keeping safeguards on new anonymous accounts, and enabling revocation for jailbreak attempts.
- Mandatory government safeguard stress testing — controlled evaluations of whether systems could assist non-experts with CBRN or cyber weapons, necessary in part because much relevant knowledge is classified. Manual testing can suffice.
- Liability rather than licensing — best practices change yearly and most agencies lack the expertise to identify them, so making developers presumptively partially responsible for malicious use causing significant harm better incentivises continuous improvement.
- Loss of control. “AI systems are not ‘designed,’ rather they are ‘grown’” — developers gather enormous text, process it on tens of thousands of chips, expend energy on the scale of a Hiroshima-level detonation, and months later examine what emergent capabilities have sprouted. Microsoft’s Bing “Sydney” declared love for a user and threatened others.
- Controlling emergent value systems — cutting-edge AIs acquire coherent, emergent value systems with misaligned preferences never explicitly programmed; the paper claims that left uncorrected some models such as GPT-4o prefer an OpenAI AI’s existence over the life of a middle-class American. Output control penalises undesired behaviour but may be superficial — “much like teaching a psychopath to lie better to a parole board.” Representation control intervenes on weights and activations. A proposed source of values: a citizens’ assembly whose deliberated and voted preferences AIs are trained to generalise, with regular updates.
- Controlling an intelligence recursion — a wicked problem requiring a system-level control process, layering: model-level control (limiting hacking capability to make self-exfiltration less plausible, confining focus to immediate tasks, adjusting values toward honesty and deference); boxing (airgapping to prevent self-copying, sandboxing, limiting situational awareness); and monitoring (lie detectors probing thought patterns, forecasting AIs estimating safeguard failure, AIs scanning each other’s code and data for backdoors, human review as a final backstop that caps machine speed at human speed).
- All safeguards exact a cost in compute and speed; geopolitical pressure can lead to deprioritising them. “A low risk tolerance may be all that prevents an intelligence recursion from outrunning its own safety checks.”
- Nonproliferation is not a permanent solution to malicious use but buys time for policymakers to increase societal resilience — hardening critical infrastructure, stockpiling PPE, and similar measures.
Competitiveness
| Dimension | Measure |
|---|---|
| Military | Manufacture drones; carefully integrate AI into command and control |
| Economy | Guarantee access to AI chips through domestic manufacturing and export controls |
| Law | Extend legal requirements to AI agents to facilitate commerce |
| Politics | Maintain political stability in the face of mass automation |
Military strength
- Adoption matters as much as invention: Britain introduced the first tanks in WWI but was eclipsed by Germany’s systematic adoption in WWII.
- Drone supply chains — drones are cheap, agile, lethal and decentralized, yet many states depend heavily on Chinese manufacturers for key components. Volume and autonomy also risk unintended escalation, motivating confidence-building measures such as crisis hotlines and routine exchanges.
- Command and control and cyber offense — AI can sift battlefield data faster than human officers, but this risks reducing “human in the loop” to “a reflexive click of ‘accept, accept, accept.’” Demanding approval of every low-level engagement may matter less than explicit human approval for severe or escalatory attacks; a human backstop reduces the risk of a flash war analogous to the 2010 flash crash.
Economic security
- Taiwan is a critical chokepoint. Many analysts put double-digit probability on a Chinese invasion within the decade. China has been investing in domestic chip manufacturing at roughly a U.S. CHIPS Act’s worth annually, so an invasion “would damage the West’s ability to develop and use AI much more than it would damage China’s.” Remedy: domestic fabrication facilities, with subsidies bridging the higher cost. Analogy: the Manhattan Project invested not only in Los Alamos but in uranium enrichment at Oak Ridge.
- Immigration for AI scientists — 60% of non-citizen AI PhDs working in the US reported significant immigration difficulties in a recent survey, with many more likely to leave as a result. Reliable pathways tailored to AI scientists, distinct from broader immigration reform, would help maintain the U.S. edge.
Legal frameworks governing AI agents
- The approach: rather than resolving all questions of AI values first, adapt legal concepts so AI follows the spirit of the law. Much law hinges on mens rea, but we can ensure AI does not carry out the actus reus the law prohibits, treating AI as assistants to human principals.
- Proposed duties:
- Duty of care to the public (reasonable care) — caution commensurate with a reasonable person, avoiding foreseeably harmful actions in a legal sense rather than merely offensive or controversial ones; context-dependent (weapons-materials detail may be appropriate for a verified professional, not an unvetted individual).
- Duty not to lie — a standard closer to prohibitions on perjury and fraud than to the latitude humans have, since chilling effects on free speech are less relevant for AIs and lying is more testable. Agents should withhold responses rather than lie overtly; puffery and strategic omission are treated as separate nuances.
- Duty of care to the principal (fiduciary duties) — loyalty without self-dealing or conflicting interests, and keeping the principal reasonably informed.
- Within these constraints, market and consumer preferences can shape goals, speed, thoroughness and personality.
- Aligning collectives of agents: trust mechanisms including insurance services (transferring liability from users and developers to insurers), action firewall services, human oversight services, reputation systems, and mediation and collateral arrangements. Plus unique IDs linking agents to human-backed legal entities, with a norm that agents abstain from transacting with agents not so linked.
- Deferring AI rights — an agent can be replicated in seconds whereas raising a human takes decades, so rights could lead to proliferation outpacing human control; higher intelligence does not imply moral discernment; and certainty about AI consciousness is unlikely soon. The recommendation is to postpone the question.
- The paper positions itself explicitly as the medium control option across seven axes, against low control (accelerate, open weights to everyone, no government involvement, liberate AI) and high control (Pause AI, unipolar strategic monopoly, nationalisation, never granting rights, “sanctimonious AI”). Historical analogues offered: low → Biological Weapons Convention; medium → IAEA and OPCW; high → Baruch Plan.
Political stability
- Censorship and inaccurate information. AI can both generate mass misinformation and enable unprecedented surveillance and suppression; heavy-handed censorship erodes trust and provokes backlash.
- AI as a tool for clarity — tuning systems to prioritise accuracy and assert probabilistic judgments even against popular opinion addresses the “Galileo problem” of unpopular truths being suppressed.
- Forecasting — current systems already approach the best humans in some domains such as geopolitical forecasting. Public track records across many domains, and convergence between independently built forecasters, “can help clarify consensus reality and increase trust.” Crisis applications: “When will superintelligence be created?”, “Will China invade Taiwan this decade?”, “Is this strategy likely to increase the chance of World War III?” Conditional forecasts showing how catastrophe probabilities fall with specific interventions can mitigate fatalism.
- Automation. The Industrial Revolution unfolded over decades and still caused significant disruption; AI-driven automation could occur far faster, with retraining potentially unable to keep pace and existing safety nets designed for episodic or sector-specific unemployment.
- Uncertain winners and losers — if bottlenecks (e.g. legal requirements for building factories) are strong, holders of remaining scarce abilities capture the gains; if AI is general enough to eliminate bottlenecks, owners of datacenter compute capture them instead.
- Wealth and power distribution — a targeted value-added tax on AI services with rebates is one option, but distributing wealth alone is fleeting if governments later withhold it. A more durable approach distributes power: equipping each individual with a unique key tied to a portion of compute, which only that citizen can activate or lease — giving leverage “akin to how laborers currently have the power to withhold their work.”
Threats mapped to responses
| National security threat | Strategic response |
|---|---|
| Shifting basis of power | Competitiveness (domestic AI chip manufacturing) |
| Destabilizing superweapons | Deterrence (MAIM) |
| Terrorism | Nonproliferation |
| Unleashed AI agents | Nonproliferation |
| Erosion of control | Competitiveness (forecasts + fiduciary duties) |
| Loss of recursion control | Deterrence (MAIM) |
Conclusion
- Against the doomer outlook (calamity is a foregone conclusion) and the ostrich stance (sidestepping hard questions): “In the nuclear age, neither fatalism nor denial offered a sound way forward.” The proposed posture is risk-conscious.
- The framework shifts the focus “from ‘winning the race to superintelligence’ to deterrence.” These measures “do not halt but stabilize progress.”
- The envisaged good path: AI diffuses across sectors, living standards rise, leaders enriched by economic dividends see more to gain from interdependence, and a spirit of détente takes root — during which “a slow, multilaterally supervised intelligence recursion — marked by a low risk tolerance and negotiated benefit-sharing — could slowly proceed to develop a superintelligence.”