Abstract

The fundamental existential danger from ASI is not whether it goes rogue but the raw power it confers and the lowered barriers to creating dangerous technologies — i.e. the Vulnerable World Hypothesis. Because overemphasis on rogue-ASI scenarios makes xrisk easy to dismiss, and because alignment is market-supplied safety that accelerates ASI without addressing the world’s vulnerability to dangerous technologies, alignment should be treated as a subsidiary goal rather than the primary one.

The disagreement among experts

  • Concerned: Geoffrey Hinton (“If you take the existential risk seriously, as I now do, it might be quite sensible to just stop developing these things any further”); Yoshua Bengio (“around, like, 20 per cent probability that it turns out catastrophic”); Sam Altman, pre-OpenAI (“probably the greatest threat to the continued existence of humanity”).
  • Dismissive: Yann LeCun (xrisk talk is “premature”, “preposterous”, “complete B.S.”; ASI “will be under our control”); Andrew Ng (worrying about it is “like worrying about overpopulation on Mars”); Marc Andreessen (“all the hallmarks of a millenarian apocalypse cult”).
  • The common explanation — short-term self-interest on both sides — is “too glib.” Many xrisk-concerned people forgo fortunes; xrisk-skeptics have far more interest in remaining alive than in getting richer. The disagreement is attributed instead to sincerely held differences in underlying assumptions.

Scope assumptions

  • ASI is assumed achieved: superhuman performance across most domains of human endeavour, plus ability to act in the physical world. Not about ChatGPT/Claude/Gemini as they exist.
  • Early hints: AlphaGo Master (Ke Jie’s “God of Go”; “not a single human has touched the edge of the truth of Go”); AlphaChip generating superhuman chip layouts in hours. ASI would be Move 37 “a trillion-fold, pervasive across multiple domains.”
  • Whether ASI arrives via LLMs or another route, and whether it is decades away, does not change the argument.
  • “xrisk” is used as shorthand covering catastrophic risks too, since the analyses overlap.

Biorisk scenario

The scenario is a vehicle for examining patterns of argument, not an attempt to convince.

  • Setup: a doomsday cultist asks an ASI to design an airborne virus spreading as fast as measles, as lethal as ebola, with asymptomatic spread like Covid, evading known vaccine techniques — and to help fabricate and release it.

Objection: “We will align our AI systems to refuse”

  • There is tension between safety guardrails and the desire for systems that seek truth without barrier. Who decides the boundary between safe and dangerous truths?
  • Militaries, open-source groups and truth-seeking idealists will draw that boundary very differently from consumer products (cf. Musk’s call for “TruthGPT”). DeepSeek appears to have few or no biorisk guardrails.
  • Techniques for building “safe” systems are easily repurposed to build less safe ones — reducing the cost of alignment work while increasing truth-seeking power.
  • Deep understanding of reality is intrinsically dual use. Quantum mechanics gave molecular biology, materials science, biomedicine and semiconductors — and nuclear weapons. You could not have had one without the other.
  • The fundamental asymmetry: “understand reality as deeply as possible” is a simple, well-defined goal grounded in objective reality; “aligned” systems require complex, subjective, hard-to-agree guardrails built on top of it, with definitions that shift with social consensus. This makes unrestricted systems intrinsically easier to build than aligned ones — so alignment is intrinsically an unstable situation, ripe for proliferation.
  • Removing alignment barriers is cheap: Kevin Esvelt found it cost only a few hundred dollars to finetune an open-source model into being far more helpful for generating pandemic agents, concluding that releasing weights of future models “no matter how robustly safeguarded, will trigger the proliferation of capabilities sufficient to acquire pandemic agents.”

Objection: “Is such a virus biologically plausible?”

  • Not certain, but there is considerable chance: viruses with each single characteristic are known, and combination engineering is improving.
  • In the 2000s researchers accidentally engineered a mousepox variant 100% lethal to vaccinated mice. Case fatality rates near 100% exist in nature (MSW strain of Californian myxoma virus in European rabbits).
  • “Only” 50%, 90% or 99% mortality is small comfort and would leave humanity vulnerable to cascading failure.

Objection: “Are such cults for real?”

  • World-ending intent is more common than one might assume a priori. It often involves significant mental illness and limited competence — fortunately.
  • Aum Shinrikyo: founded 1987, >10,000 recruits by the mid-1990s, many educated Japanese in their 20s–30s with science and technology expertise. After a failed 1990 bid for political power, leaders shifted from predicting apocalypse to bringing it about. Programs for chemical and biological weapons, missiles, and an Australian sheep station for uranium prospecting. Per historian Charles Townshend, “police raids found enough sarin in Aum’s possession to kill over four million people.”
  • The concern is that ASI “may greatly increase the supply of apocalyptic competence.”

The best skeptical response: just-in-time coevolution

  • “Humanity will prevent it in other ways” — controlling synthesis equipment, screening viral sequences, law enforcement and intelligence, and ASI’s own help.
  • This is humanity’s historical strategy: allow innovation, respond to problems as they arise. It works well — for every asbestos, leaded petrol or climate change there are thousands of innovations where institutions overcame early problems.
  • It will work well in the short term. The medium term is where it becomes hard to defend.

The Vulnerable World Hypothesis

  • Bostrom’s question: is there a “recipe for ruin” — some cheap, easy-to-make-and-deploy technology capable of ending humanity, or causing >90% deaths?
  • Surviving past advances is no guarantee: “Like explorers in uncharted territory, we could suddenly find ourselves facing an insurmountable danger.”

Candidate recipes

  • Mirror bacteria — organisms with all molecules reversed into chiral mirror images; conventional immune defences may fail to recognise them, allowing unchecked spread and mass extinction. Scientists previously pursuing elements of mirror life recently published in Science explaining the dangers and publicly abandoning the work.
  • Easy-to-make nuclear weapons — Ted Taylor, the leading American nuclear weapons designer, told John McPhee there is “a way to make a bomb… so simple that I just don’t want to describe it.” Fissile material remains the bottleneck; if removed, proliferation may be unavoidable. Fewer than ten nuclear powers can perhaps avoid use; thousands of rogue actors cannot. Taylor: “Every civilization must go through [its nuclear crisis]… Those that don’t make it destroy themselves.”

Objection: “This is about ingenuity, not ASI”

  • Correct — the Vulnerable World Hypothesis is about the nature of the universe and technology, not ASI specifically.
  • But ASI is a supercharger, potentially uncovering recipes for ruin that would have taken centuries or never been found, while collapsing the expertise and resource requirements.
  • “You cannot put an impermeable barrier around understanding and controlling reality, when you have built systems with intellectual capacity beyond von Neumann or Einstein” — and if one such mind can be made, it can be scaled to a million running a thousand times faster. Expect a major discontinuity in individual capability unless individuals are denied access.

Objection: “Most people don’t want to blow up the world”

  • Fair, but a lot of people strongly desire power and domination — visible from everyday behaviour up to colonial powers destroying or oppressing indigenous populations, often out of indifference rather than malice. In a Vulnerable World the competitive drive to ratchet up that power will be enormous.

On “uplift” proposals

  • BCIs, genetic engineering and human enhancement misunderstand the issue: the problem is not carbon versus silicon but increased capability leading to increased power and access to catastrophic technologies.
  • Uplift seems likely to increase the danger, and adds problems of its own (BCI companies or regulators able to reshape human thought dictatorially).

Objection: “Capabilities and safety co-evolve”

  • Assumes ASI capabilities improve slowly enough for institutions to adapt. History suggests discontinuities shock even top experts.
  • Admiral William Leahy — a former Annapolis physics and chemistry teacher and Chairman of the Joint Chiefs — told Truman the atomic bomb was “the biggest fool thing we have ever done. The bomb will never go off, and I speak as an expert in explosives.”
  • Simple offensive technologies can be very hard to defend against; there is no intrinsic requirement for symmetry.

Loss of control to ASI

Standard skeptical responses, examined briefly:

  • “We won’t make agents with their own interests and power-seeking drive.” Implausible: intellectually satisfying for developers, profitable for investors. Romance-bots, personal assistants, agents of persuasion, military robots and trading systems all function “better” with agency, and market incentives reward agency without equivalent rewards for safety.
  • “We won’t cede much power to inhuman entities.” We already delegate life-and-death authority to guided missiles, financial decisions to trading algorithms (the 2010 Flash Crash lost more than a trillion dollars in 36 minutes), and — via the Soviet Dead Hand system — world-ending authority to automated systems.
  • “We can simply turn them off.” Could we turn off Facebook or X if they were net negative? Deep integration and aligned powerful interests make this unrealistic.
  • Better response: we will control ASI as we control other powerful actors — institutions, norms, laws, education. Viable only if governance keeps pace with capability, which will be an enormous challenge.

Why the rogue-ASI framing is a mistake

  1. It gives skeptics an easy exit. People used to technology being ultimately human-controlled find loss of control implausible, and so dismiss xrisk entirely. But the underlying issue is the power conferred by the machines, whether wielded by humans or by out-of-control machines.
  2. It leads to badly mistaken actions. If the fundamental challenge is the dual-use nature of reality, then alignment and interpretability work is counterproductive on net: it reduces certain risks while making commercial ASI development far more tractable, speeding progress toward catastrophic capabilities and building “a largely invisible latent overhang in dangerous capabilities.”
  • Alignment is characterised as market-supplied safety, leaving critically undersupplied the non-market parts — preventing proliferation of easy development of dangerous technologies.
  • Explicit clarification: this is not the claim that rogue ASI won’t occur, nor that misuse rather than agentic ASI is the worry. The claim is that whether ASI gets out of control is not fundamental to whether ASI poses an xrisk or how to avert it.

Conclusion

Why disagreement persists

  • Disagreement arises from differing thresholds for conviction and differing comfort with reasoning based on toy models and heuristic arguments.
  • Climate analogy: Arrhenius and Ångström reached opposite conclusions from different toy models; satisfactory models only emerged in the 1950s–60s, consensus in the 1980s–90s. For ASI we are still at the toy model stage — and unlike climate, ASI carries a wildcard factor of acting in ways intrinsically unpredictable in advance (Vinge’s “opaque wall across the future”).
  • Much xrisk writing has been bombastic, overconfident or narratively plausible rather than factual, inviting comparison to earlier panics (population bomb, Y2K). Conversely, many skeptics have never seriously engaged with the strongest forms of the arguments, engaging only defensively to find minor holes.

Barriers to engagement

  • The strangeness of the idea (usually a good heuristic to ignore wild-sounding scenarios — but “one thinks of aristocrats in 1780s France”).
  • The unpleasantness of engaging with large-scale death.
  • Difficulty of assessment without domain expertise, and lacking “expert intuitions for how much power is latent in the world.”
  • Countervailing incentives: near-term AI benefits will be real and profound, xrisk worriers will keep seeming wrong, and betting on AI increases wealth, power and status. Far more capital flows to capability advances than to non-market safety. “Capital reshapes people, aligning their beliefs with short-term corporate interests rather than with humanity’s interests as a whole.”

What to do

  • Efforts that take recipes for ruin seriously and prioritise safety go under differential technological development, d/acc, and coceleration.
  • Such work is less well-paid and less prestigious than working on AI models; existing institutions reward improving technology (including aligning it), which — when deeper understanding is dual-use — inadvertently rewards increasing existential risk.
  • “There is a moral obligation for anyone working on AGI to investigate this risk with deep seriousness, and to act even if it means giving up their own short-term interests.”

Postscript (added 3 June 2025): self-critique and replies

  • “So you’re saying we shouldn’t align AGI or ASI?” No. The objection confuses “solve the alignment problem” with “make ASI safe for humanity” — an AI control fallacy conflating control over ASI with safety from ASI. Analogy: a swimmer aiming for the Olympic team who adopts “increase arm strength” as a goal — necessary up to a point, then actively harmful to the primary goal. Alignment is a subsidiary goal; confusing it with the primary goal has damaged the ability to achieve the primary goal.
  • RLHF as illustration — developed by Paul Christiano from the mid-2010s, crucial to ChatGPT’s launch, and thereby to kicking off a furious capital- and talent-funded race to ASI. The pattern will recur: better alignment techniques keep ever-more-capable systems acceptable to society, making them more attractive as products. Short term this is the pattern you want; long term it accelerates toward an unstable situation, both because capabilities lie hidden and hard-to-align, and because the same techniques will be used by militaries, intelligence agencies and finetuning bad actors with conflicting notions of “alignment.”
  • Even successful aligned ASI leaves the problem that, absent one organisation rising to totalitarian dominance, other parties will build ASI with conflicting intent — “a recipe for disaster, unless we have governance mechanisms almost entirely absent today.”
  • A skeptic’s reply: in 5,000 BCE, a thousandfold increase in individual destructive power would have looked like certain ruin, yet ideas and institutions improved — rule of law, democracy, separation of powers. The counter is that it is foolish to take for granted that governance will improve fast enough.
  • Why alignment attracts people: “do what you can” is normally a good heuristic for civilisation-scale problems, and alignment is tractable, enjoyable, lucrative and socially validated. Because technical alignment work is market-supplied safety, it grows with the companies and comes to dominate AI safety — and its dominance is easily confused with intrinsic importance. Selection effects mean lab safety researchers’ beliefs approximate what the market wants, which should not be confused with being correct. The strong version of the claim: “I’m afraid I believe many of the people working on alignment at OpenAI and Anthropic have made human extinction or some similarly bad outcome more likely.”
  • “So what to do instead?” — “I wish I knew.” The crucial tension is between work that centralises power within AI companies (most alignment research, safety work that makes company systems more successful) and work that decentralises — strengthening institutions, governance and defensive capabilities across society. Litmus test: does my work make AI companies more central and powerful, or does it build capacity elsewhere? At the margin the latter is nearly always better.
  • A residual difficulty: as ingenuity increasingly comes from AI, even decentralised governance and defensive work requires working with AI and eventually ASI, which likely strengthens the companies.
  • When is alignment work a good idea? At the margin it almost never makes sense to work on market-supplied safety, because “the world is always well-supplied with people willing to do what capital wants.” The heuristic offered: work on non-market safety — governance, pause or slowdown, new institutional ideas for governing technology.
  • Is ASI a recipe for ruin? Two simplifications acknowledged: (1) Bostrom’s formulation does not require a cheap-and-easy recipe, only some technology that inevitably ends humanity — this essay uses a more extreme variant; (2) the first ASI will likely be extremely expensive, but if expensive ASI is possible, cheap ASI likely becomes possible too.
  • On the intuition that recipes for ruin must be impossible — outputs are not always proportional to inputs. Fire is a prototype: one person can start one for free causing many deaths and billions in damage, every year, despite enormous societal spending. The NFPA estimates fire costs roughly 1.9% of US GDP — in 2014, $273 billion for defence and $55 billion in losses, with millennia of mostly invisible defensive infrastructure. “We can improve defences across many domains, but if in just one domain it turns out to be too difficult, then we’re in real trouble.”
  • “If we’re headed toward an ASI singleton, surely alignment is the goal?” (Steve Byrnes). The essay does not claim alignment is never a goal — only that it is subsidiary and likely always well-supplied. The focus on alignment has so far likely accelerated a singleton while also, confusingly, decreasing its probability.
  • Final thought: against the view that incentives and capital are all — “People don’t merely respond to incentives and power, they also respond to truth and improvements in our understanding,” making robust discussion of what is true, independent of power or incentives, worthwhile.