Abstract
The central goal of AI governance should be existential security — a state in which existential risk from AI is negligible indefinitely, or long enough for humanity to plan carefully. Framing governance positively, around an endgame to reach rather than outcomes to avoid, yields greater strategic clarity; actors should therefore make their preferred theories of victory explicit and public.
The core argument
- An endgame is a state of existential security. A theory of victory combines an endgame with a plausible and prescriptive strategy for reaching it, and should be robust across a range of future scenarios.
- Risk management is normally continual, but this is insufficient for existential risk: any non-negligible level makes catastrophe all but certain given sufficient time.
- An endgame is not a description of human flourishing — it only preserves optionality to enable flourishing.
Why a positive framing
- Strategic clarity: it is easier to conceptualise bringing about one particular outcome than avoiding several different ones. The analogy offered is a chess opening system (e.g. the London) versus learning to counter each opening trap separately.
- Minimising strategic incoherence: a strategy targeting one threat model can accidentally undermine strategies for others, and different actors may be implicitly pursuing incompatible theories of victory.
Evaluation criteria
- Endgames must be existentially secure (expected time to failure of hundreds or thousands of years) and preserve optionality — no lock-in of a dystopic or merely mediocre future, ideally enabling a “long reflection.”
- Strategies must be plausible (not “convince all labs to voluntarily stop”) and prescriptive (not “get lucky with an unforeseen technical barrier”).
- Robustness means satisfying these across plausible scenarios — combinations of strategic parameters like timeline to TAI, takeoff speed, and difficulty of technical safety. A robust theory is likely a set of conditionals rather than one endgame-strategy pair.
- Interventions can be evaluated by how many theories of victory they support. Absent a clearly preferable theory, prefer interventions supporting several (e.g. raising public awareness, expanding governance-organisation capacity); if few such exist, prefer going “all in” on one or a small compatible group.
- Caveat: mitigating AI existential risk is an imperfect proxy for mitigating existential risk overall — advanced AI might help secure against nuclear, biological and climate risks, so between two theories with equal AI-risk effects, prefer the one delivering benefits sooner.
Nuclear risk as a case study
- The world has not reached existential security from nuclear risk: mutually assured destruction only works if nuclear powers can be sufficiently modelled as rational agents.
International coordination to prevent development
- Otto Frisch on wartime pressure: “We were at war, and the idea was reasonably obvious; very probably some German scientists had had the same idea and were working on it.”
- Existential risk only emerged with the hydrogen bomb and arsenal expansion, so risk might have been averted even after initial development — but this would have required strong coordination, and in the extreme, world government. Edward Teller and Albert Einstein both advocated for world government at various points; Niels Bohr attempted to broker US-USSR coordination, interesting FDR before failing after a meeting with Churchill.
- Two features made this promising for nuclear that don’t transfer to AI:
- Nuclear weapons are single-use. Much of the development was purely destructive, so scientists and states would face internal and external resistance in peacetime.
- Proliferation is limited. Development is restricted to well-resourced states and is relatively visible to satellite imagery and inspectors.
A unilateral monopoly
- Bertrand Russell proposed in The Bomb and Civilization that the US use its temporary superiority to insist on disarmament everywhere else, by war if necessary, forming a new League of Nations under American leadership.
- Einstein described this as “preventive war” and rejected it, noting no military leader believed Russian surrender possible without a costly ground invasion of Europe and Asia, leaving Western Europe ruined and requiring generations of rebuilding and policing.
Nuclear vs AI compared
- Decisive strategic advantage: both confer it, and both therefore create racing dynamics.
- Threat models: the only plausible nuclear existential threat model is intentional use; nuclear weapons pose no intrinsic risk of existential accident. AI adds loss of control, which means it may not be in any actor’s interest to develop TAI even if it would grant a DSA.
- Single vs dual use: TAI is trivially dual-use and general-purpose, so there are reasons to develop it independent of strategic balance, making a case against development harder to sustain. But it also offers defensive capabilities nuclear weapons cannot — a nuclear weapon can’t defend against one already deployed.
- Distribution: trained models can be copied, distributed, stolen, or open-sourced, so TAI could proliferate far faster than nuclear weapons.
Theory of victory 1 — AI development moratorium
- International coordination and technical governance indefinitely prevent development beyond a specified threshold or of particular kinds.
- Anthony Aguirre’s “Close the Gates” argument: superhuman GPAI poses profound risks; nearly all the benefits we really want are obtainable “inside the Gate”; systems requiring superhuman capability can always be developed later once judged provably safe; and Gate closure is more viable than it might seem.
- Eliezer Yudkowsky’s version: an indefinite, worldwide moratorium with no exceptions for governments or militaries; shut down large GPU clusters and training runs; cap and ratchet down training compute; track all GPUs sold; be more afraid of the moratorium being violated than of shooting conflict.
- Assessment: guarantees existential security if successful, and preserves optionality (less powerful systems still permitted, moratorium liftable later). The challenge is establishment and indefinite enforcement — potential developers have strong reasons to build, both strategic and beneficial.
- Historical record is mixed: nonproliferation largely succeeded post-Cold War (except North Korea); disarmament less so; climate coordination has largely failed. But the successes may rest on features AI lacks — individual nuclear weapons did not pose existential risk and Hiroshima turned opinion against use, while climate risk built up gradually enough for consensus and renewables to develop.
Theory of victory 2 — AI leviathan
- TAI itself enforces a monopoly on TAI development, created either by a first mover (analogous to Russell’s nuclear proposal) or voluntarily by several actors to avoid a multipolar scenario.
- Samuel Hammond frames the intelligence explosion as putting liberal democracy on a knife edge — an AI Leviathan global singleton on one side, state collapse and fragmentation on the other.
- Nick Bostrom in Superintelligence notes that a post-transition Leviathan would be created by agents who may themselves be superintelligent, improving the odds they could solve the control problem.
- Assessment: solves the moratorium’s enforcement and coordination problems. But its merit depends on the likelihood of the agentic threat model — if the first TAI system pursuing problematic goals is more likely than a moratorium succeeding without a Leviathan, it isn’t worth it.
- Lock-in risk: mistakes in construction may be uncorrectable, eliminating optionality. Hammond’s “bounded leviathan” retains human control but may trade off against effective enforcement, and control introduces a point of failure that could be gamed.
- Prescriptive but clearly underspecified: who builds it, when and how it is empowered, and how it would represent humanity’s collective interests all remain open.
Theory of victory 3 — Defensive acceleration
- Leverage advanced AI to develop defensive technologies (AI-empowered technical safety research, cybersecurity, biosecurity) so that defensive applications outpace offensive ones.
- Term introduced by Vitalik Buterin (d/acc), building on differential technological development — proposed by Bostrom in 2002 and developed by Sandbrink, Hobbs, Swett, Dafoe and Sandberg (2022): “delay or halt risk-increasing technologies and preferentially advance risk-reducing defensive, safety, or substitute technologies.”
- Assessment: existentially secure as an endgame if the offense-defense balance indefinitely favours defence, and it preserves optionality — some offensive technologies deferred, but nothing permanently off-limits.
- Plausibility problems: it is neither the default nor easily achievable. If it requires international coordination to slow offensive development, it faces the same game-theoretic challenges as a moratorium. Private companies on profit-seeking incentives should not be expected to curate a defence-favouring development pipeline.
- Prescription: governments intervene to incentivise defensive technologies — large R&D investment, financial incentives, and outright prohibition of offensive technologies. But it isn’t always clear which defensive technologies to prioritise, or even in advance whether a technology will be primarily offensive or defensive.
- Possibly better as a complement to another theory of victory than as a stand-alone.
Conclusion
- The authors deliberately do not choose between the three theories.
- They call for follow-up work that proposes novel theories of victory, evaluates particular ones thoroughly, identifies strategic compatibilities and incompatibilities between them, and maps the explicit or implicit theories of victory held by different actors in AI governance.