Abstract
Rather than focusing solely on restricting access to powerful models, we should first identify domains where the offense-defense balance favours offense and use frontier models to strengthen societal defences and “antibodies” in those areas. Sequencing of releases matters: if everyone gains access to powerful AGI simultaneously, the balance between misuse and patching becomes unclear.
Third of a seven-part series on responsibility, accountability and control in a pre-AGI world, written in a personal capacity by Séb Krier of Google DeepMind’s Policy Development & Strategy team.
The core proposal
- Sequencing matters. If everyone suddenly has access to a particularly powerful AGI, it is unclear how the balance between misuse/offense and patching/defense is affected.
- Assuming frontier models do present sophisticated misuse potential, the first step should be identifying domains where the offense-defense balance favours offense, then using those models to improve societal defences. Not a panacea, but such measures would undeniably enhance resilience.
Concrete defensive applications
- Cybersecurity — address technical vulnerabilities, bolster protections and red-teaming in national infrastructure, hospitals and similar, raising barriers to entry for malicious actors.
- Defensive AI research — e.g. oracle-like scientists with sufficient epistemic uncertainty; the post cites work showing that modelling diverse reasoning chains and full posterior distributions makes language models less prone to overconfidence and shortcut learning.
- Legal and regulatory scrutiny — analyse laws for inefficiency, contradictions and obsolete standards to make them simpler, less gameable and more consistent; use simulation and modelling to foresee unexpected effects of proposed legislation.
- Cognitive security — ensure assistants filter information in epistemically desirable ways: highlighting confirmation bias, handling information overload, encouraging critical reflection, reinforcing good epistemic practices.
- National security and defence — continuous analysis of global digital activities, improved missile defence, identification of threats and of AI system misuse.
- Biosecurity — monitoring and oversight of DNA synthesis, securing cloud labs, better vaccine production, pandemic preparedness, and accelerated development of medical countermeasures.
- Digital infrastructure resilience — stress-test and redesign key internet components against AGI-driven attacks and foreign hybrid warfare.
- Social security — radical rethinking, since job impacts will likely be faster and more chaotic than previous technological breakthroughs.
- Digital identity and authentication — deepfakes have so far been less problematic than anticipated but it’s too early to judge; impersonation is already growing.
The vulnerable world hypothesis and offense-defense balance
- Bostrom’s vulnerable world hypothesis holds that technological progress has made mass harm easier to cause than to prevent, with technologies like nuclear weapons and engineered pathogens letting small groups cause catastrophic damage that defensive technologies may not fully counteract.
- Counterpoints offered: nuclear deterrence and nonproliferation have prevented wartime use since 1945, and global public health infrastructure has largely contained dangerous natural pandemics.
- Historical analogy (from Gustavs Zilgalvis): the Theodosian Walls of the Byzantine Empire represented a strategic equilibrium in the grand strategy against the Huns; where territories couldn’t be fortified, the Empire paid 2,100 pounds of gold per year to spare unfortified regions. This persisted until offensive advances — Mehmed II’s cannons and the composite reflex bow — rendered the walls ineffective, leading to the fall of Constantinople.
- “It’s a lot easier to split the atom than it is to build walls large enough and do it fast enough to defend against nuclear weapons” — but the right historical lessons may offer some hope.
The gap in the literature
- This requires thinking not only about how or whether to restrict models, but how to use them — and “ideas along these lines seem notably absent from the literature.”
- Academia has been effective at scrutinising risks and harms; remedies and prescriptions are harder to find. A proactive approach both minimises risks and amplifies AGI’s benefits in reinforcing societal structures.
- Even where the offense-defense balance favours offense (nuclear security, biosecurity), defensive systems are not entirely ineffective.
- Michael Nielsen is quoted: many aspects of safety and security are approximately public goods or collective action problems that the market as currently constructed undersupplies — motivating efforts, including Buterin’s, to develop new financial instruments addressing such problems.
Closing question posed to researchers
How can we proactively harness the power of advanced AI to strengthen societal resilience and address vulnerabilities, rather than focusing solely on restricting access?