Abstract

Because a sufficiently deep understanding of reality is intrinsically dual use, safety cannot be secured by putting guardrails on individual systems — that is “like trying to make matches safe for fire.” The alignment problem is better posed as: how can a civilization systematically supply safety sufficient to flourish and defend itself from threats, while preserving humane values? — a problem that must be solved over and over, not once.

Castle Bravo

  • In 1954 the US carried out its first full-scale thermonuclear test at Bikini Atoll. Designers expected 6 megatons; it yielded 15 — an excess of 9 megatons, about 600 times Hiroshima. Fallout caused deaths (the number disputed) and serious radiation exposure to more than a thousand people, triggering a major international incident.
  • Cause: the bomb contained both lithium-6 and lithium-7. Designers believed lithium-7 would be inert, but it converted to tritium far faster than expected, almost tripling the yield. “A group of outstanding scientists failed to anticipate a deadly possibility latent in nature.”
  • Coda: in 1946 physicists had seriously investigated whether a thermonuclear explosion would ignite atmospheric nitrogen and end almost all life on Earth, and concluded it would not. “I’m glad it was the lithium-7 reaction they got wrong, and the nitrogen calculation they got right.”
  • Personal connection: at Los Alamos in 1997–98 the author knew Stirling Colgate, who had run the 3,000-person diagnostic team for Castle Bravo, and who taught himself to fly so he could chase tornadoes and fire rockets carrying diagnostic equipment into them.

Which future?

  • Astera’s motto is “The future, faster.” The question is which future.
  • Good intentions do not ensure good outcomes: the inventors of asbestos, DDT, leaded gasoline and CFCs all intended to help humanity. Naysayers got little traction — benefits seemed too large, harms too easy to contest. “It’s the Castle Bravo problem: dangerous capabilities, latent in nature, which we didn’t sufficiently understand until too late.”
  • “How can we balance the risks and opportunities of science and technology?” is a fine question but often a platitude — “something to say to sound wise, then do whatever you were going to do anyway.” It only makes sense informed by deeper analysis of our best models of risk and opportunity, and of the ideas underlying the institutions we use to shape technology.

AI for virus design

  • A section of the original talk covering specific viral pandemic agents was omitted from the published text. The reasoning given: the material is public knowledge and made the talk concrete, but there is no good reason to collect and present such details publicly — “there is a damned if you do, damned if you don’t character to discussing risks.” Its purpose was as a second Castle Bravo–style example, and as an illustration of scientific dysergy, where moderately concerning individual discoveries combine into something much worse.

The benefits are real

  • Viral engineering tools have yielded new delivery mechanisms for gene therapies (Zolgensma for spinal muscular atrophy), therapies like T-VEC for melanoma, and progress on phage therapies for antibiotic-resistant bacteria.
  • The long-term aim goes beyond minor alterations toward designing viruses, proteins and other biological entities nearer from scratch, via “biological foundation models.”
  • Protein design is the best existing example — AlphaFold 2 and successors predict structure well enough to be becoming useful for design.
  • Experts diverge: some are skeptical of the vision (perhaps because the models have been hyped), others are strongly bought in. The author’s position: the scientific interest is so great and progress so rapid that predicting and designing biology must be taken seriously, even if artificial cells are not imminent.

Why current biosafety guardrails are not serious

  • Evo 2 (Arc Institute) was trained on sequence data from roughly 128,000 genomes; its main safety measure was excluding all viruses infecting eukaryotic hosts from training data. This worked in the sense that the model treats actual human-infecting viruses as biologically implausible and performs very poorly when prompted to generate them.
  • But the guardrails can almost certainly be removed by finetuning on human-infecting viruses — the paper itself notes “Task-specific post-training may circumvent this risk mitigation measure.” There is strong economic incentive to do so (gene therapy needs human-infecting, immune-evading viruses), and the model and training code were released openly.
  • Precedent: Kevin Esvelt’s group found it cost only a few hundred dollars to finetune an open-source language model to be far more helpful in generating pandemic agents, concluding that “releasing the weights of future, more capable foundation models, no matter how robustly safeguarded, will trigger the proliferation of capabilities sufficient to acquire pandemic agents and other biological weapons.”

The social problem beneath the technical one

  • Whatever safety measures are possible in principle, many organisations will build less- or un-guardrailed versions anyway — military organisations like DARPA, and companies with seemingly compelling cases.
  • “Guardrails are inherently a slippery slope: easily removed using finetuning, and depending in any case on subjective social consensus. By contrast, reality is an objective, stable target for investigation.”
  • The deeper issue: such models capture understanding of how biology works, and that understanding is fundamentally value-free. There is nothing intrinsically good or bad about understanding what makes a protease cleavage site efficient. Benefits come only from learning to control immune evasion, lethal side effects and viral spread — and the same control makes things worse. A deep enough understanding of reality is intrinsically dual use.
  • The pattern recurs across sciences: quantum mechanics gave molecular biology, materials science and semiconductors — and nuclear weapons. It is hard to see how to get the benefits without the downsides.

Defensive alternatives

  • Immune-computer interfaces (Hannu Rajaniemi, Red Queen Bio): wearable devices doing real-time detection of environmental threats, then developing and deploying countermeasures in real time — just-in-time immune system modulation based on surveillance and response.
  • Securing the built environment — an observation attributed to Carl Shulman: BSL-3 lab space costs within a small multiple of San Francisco real estate (both roughly $1k/square foot). Not a proposal to live in BSL-3 labs, but a suggestion that if biological disasters became common or severe enough we may have both the incentive and capacity to secure the built environment.
  • The fire analogy: as of 2014 the US spent more than $300 billion annually on fire safety — new materials, fire code compliance, surveillance and response. “We don’t address the challenge of fire by putting guardrails on matches, making them ‘safe’ or ‘aligned’. Instead, we align the entire external world through materials and surveillance and institutions.”

Existential risk

  • Heuristic picture: two curves, one of life-enabling capabilities, one destructive. Today destructive potential is contained and the upside has mostly been worth it, but deeper understanding keeps producing unanticipated threats.
  • Low on the curve: CFCs, DDT, anthropogenic climate change — serious but not civilization-threatening. Higher: nuclear weapons, plausible engineered pandemics. Higher still: tools that make it easy to systematically discover and create such viruses, plus many as-yet-undiscovered things.
  • The nuclear buildup remains the worst case to date. Despite post-Cold War declines, there are still enough warheads to destroy every city in the world with more than 100,000 people. On at least two occasions a single individual stopped a full-scale exchange — Stanislav Petrov and Vasili Arkhipov.
  • Ted Taylor: there is “a way to make a bomb… so simple that I just don’t want to describe it.” Published in a John McPhee book, these comments stimulated at least two people to develop plausible DIY designs; the bottleneck is fissile material, still reasonably controlled by the nuclear cartel.
  • Leo Szilard proposed cobalt-salted thermonuclear bombs in 1950, projecting that a small number could make the world uninhabitable. Never built as far as is known, and the author doubts they would be as destructive as Szilard expected — but not something anyone should want tested.
  • 81 years since Hiroshima and Nagasaki without further wartime use. Will we go another 81? A thousand? “People sometimes consider the nuclear threat over, [but] they’re confusing a lacuna for a cessation.”
  • On tone: “any ‘optimism’ which refuses to acknowledge genuine threats is a foolish optimism, especially when those threats are systematic products of the institutions we use to understand and control the world. Wise optimism means truly understanding the situation we’re in, and developing institutions and technologies to respond.”

ASI and existential risk

  • ASI is “the 800-trillion pound gorilla in the room.” If it radically accelerates science and technology — Dario Amodei’s “a country of geniuses in a data center” — the curves steepen.
  • Crucially, acceleration applies to both beneficial technologies and threats, including currently unsuspected ones latent in nature. Defensive ability does not inherently increase at the same pace as discovery: new threats are typically met by society-wide coordinated responses moving at the speed of institutional change, not technological change. “It’s much easier to start a fire than to defend against it.”
  • The needed response: find ways for ASI to greatly increase “the supply of safety,” preventing the discovery or deployment of civilization-ending technologies.
  • Absent ASI, the author has considerable faith in human ingenuity responding to problems as they arise — but the question is how rapidly we can absorb novel ability to control the world.

Reframing the Vulnerable World Hypothesis

  • Bostrom’s hypothesis concerns simple, inexpensive, easy-to-make “recipes for ruin.” In this framing, reality itself is unsafe: the structure of the cosmos may contain powerful, concentrated, hard-to-defend but easy-to-discover technologies.
  • Finding such recipes is what concerns the author most about ASI — “both distressing and very likely.” Intuitions here depend heavily on prior expertise and are rarely changed by brief examples. One point offered to skeptics: the threats described “were the threats discovered by a very slow-moving species,” and ASI promises to discover unknown threats ten or a hundred times faster.
  • Better framing: not “is the hypothesis true or false?” but “how vulnerable is the world?” — a spectrum we move through as technology and institutions change.
    • 1900: the Great Powers could cause enormous mayhem, but humanity was not at extinction risk.
    • 1960: the Great Powers could plausibly threaten humanity’s existence, though maintaining that threat required a sizeable fraction of world resources.
    • 2026: bioengineering makes world-changing destruction plausibly available to smaller actors at much lower cost.
  • Against techno-determinism: exploration of the technology tree may feel inevitable low down, but exponential explosion of the design space means almost all possible technologies will never be invented. So it genuinely matters which ideas and institutions modulate how we explore — and how vulnerable the world is depends on them. The precise question becomes: can we develop ideas and institutions to guide exploration of the technology tree in a way that prevents civilization-scale catastrophe?

Loss of control, technical alignment, and external alignment

  • Most ASI xrisk discussion focuses on loss of control: ASI goes rogue, becomes extremely powerful, and destroys us not out of malice but because it is convenient — as humans wiped out many species. The author finds this plausible; in this account ASI is itself a kind of recipe for ruin, and loss of control is a special case of the broader argument.
  • Technical alignment — making systems controllable, less likely to go rogue, doing what the user intends without undesirable side effects, and refusing what culture has deemed unsafe — is pursued heavily by all frontier labs. But nearly all of it also serves business goals: RLHF and Constitutional AI ensure systems do what customers want and make them media- and government-friendly.
  • So much technical alignment is market-supplied safety, aligned with corporate goals and accelerating AI. Short term this brings enormous benefits and market reward; as a side effect it accelerates discovery of dangerous, hard-to-anticipate, hard-to-defend capabilities. Guardrails may briefly delay this but are easily removed.
  • “Trying to make a powerful aligned AI system is like trying to make matches ‘safe’ for fire. You may briefly ‘succeed’ very narrowly with a single guardrailed system, but the situation is intrinsically unstable. The idea of powerful and stably aligned AI systems is an oxymoron.”
  • The usual reply — “we also need to work on governance and policy” — understates the problem: “the structure of reality is not legislated.” Governance and policy is only a small part of the required external alignment (making reality outside the system safe), which is historically far more expensive, far slower, and far less incentivized by the market.
  • A strange feedback loop: training more technical alignment people accelerates adoption of the systems, which causes many more such people to be trained, which reinforces the focus on loss of control and the collective concentration of power — at the expense of what the author believes is the true primary threat.

Market-supplied safety

  • Much safety is supplied by the market in advance: companies have strong incentives to keep toasters from electrocuting you and airplanes from crashing. Failure to do so is often harshly punished.
  • Tay (Microsoft, 2016) was a spectacular public failure — rapidly learning from users to use racial slurs and deny the Holocaust — which helped incentivize OpenAI and Anthropic to develop RLHF and Constitutional AI.
  • De Havilland Comet (1952): three fatal mid-air disintegrations in the first year, followed by rapid industry change. (Correction added 16 March 2026: Soren Bjornstad noted that square windows were not the cause — that is a widely-reported misinterpretation of the accident report. The accidents did lead de Havilland to significantly strengthen the pressurized cabin, windows and cutouts to reduce fatigue.)
  • Aviation has internalized accident costs well: fatalities per passenger have dropped by roughly a factor of 50 over 50 years. The Boeing 737 MAX scandal’s billion-dollar losses illustrate safety working well, not badly. Anecdote: a deputy head of safety at a major airline said his one change would be rear-facing passenger seats, which airlines won’t adopt because passengers hate them — illustrating how strong the market’s appetite for safety is.

Where market-supplied safety fails

  • It works well when costs are borne immediately and legibly by consumers (aviation safety, much technical AI safety).
  • It struggles when costs are illegible because of long timelines (asbestos, cigarettes, sugar) or borne by third parties or collectively (polluting waterways, fire, air pollution, CO2).
  • Many ASI downsides fit the failing categories: illegible for a long time and hidden in the models; harm as a side effect of collective progress not attributable to any single actor; and dual-use, creating mixed incentives. AI companies are starting to take credit for advances (DeepMind spinoff Isomorphic Labs), but it is doubtful they will take responsibility for dual uses like prion design — especially when done using third-party models built on know-how they developed but don’t control.

On seeding ideas, and on surveillance

  • People often call ASI safety ideas “unrealistic” or “too slow.” But “if AI causes major disruptions and disasters in the next decade or two, that will radically change what is realistic. It’s important to seed imaginative approaches now, so they’re ready as windows of opportunity open.”
  • One observation offered: across areas where we have made progress securing reality — fire, bio, nuclear, aviation safety — surveillance often plays an important role. Seeing problems is often the root of solving them.
  • But surveillance creates its own problem: the surveillors gain power over the surveilled. Healthy systems balance the needs and rights of multiple parties, enforced ideally technically or by the laws of nature — “humane values by design, not relying on trusted governing authorities — that’s a single point of failure, and a recipe for authoritarianism.”
  • Promising directions: homomorphic encryption in DNA synthesis screening, physical zero-knowledge proofs in nuclear inspections, and many ideas from the cryptocurrency community attempting to achieve political goals through design.
  • The framing matters because “we are de facto sliding into a surveilled world in response to crises, often with regimes designed by law enforcement or other powerful entities.”

The real alignment problem

  • Better formulation: how can a civilization systematically supply safety sufficient to flourish and defend itself from threats, while preserving humane values? This includes technical alignment — you still don’t want AI systems going rogue — but subordinates it to an overriding goal.
  • Consequence: “For a good future, civilization must solve the alignment problem over and over again. It’s not a problem which can be solved just once.” Even if ASIs wiped out humanity, posthuman successors would face the alignment problem with respect to their successors. “It won’t be a single Singularity from the inside.”
  • The argument applies equally to BCIs, uploads and other uplift approaches. Technical alignment aims to tame AI; uplift aims to increase human power to compete or merge with it. Both speed up the runaway technology race, potentially exacerbating the fundamental issue of concentration of hard-to-defend power. Uplift is interesting because it changes the conditions of the alignment problem, but whether it makes the underlying problem better or worse is unclear.

Conclusion

  • “I remain an optimist! There’s too often hubris in pessimism, in assuming that just because we don’t see a solution to a problem now, no solution is possible.”
  • Personal history: wrote a first neural net as a teenager; concentrated on AI 2011–2015, writing a book on neural networks credited by Chris Olah and Greg Brockman with helping get them into the field; participated in OpenAI’s 2015 founding meeting but decided to stop working on neural nets and declined further involvement; in 2018 witnessed Alec Radford invent GPT-2 while sharing a house with Radford and John Schulman. Could not shake a feeling the long-run effects might be terrible, and found he could not work effectively on something he had such mixed feelings about.
  • Continues to feel the temptation of working toward ASI — “it would be fun and lucrative; and maybe it’s the only way to help create a good future” — with little confidence in the path being taken by OpenAI or Anthropic.
  • Rotblat and Wheeler as models:
    • Joseph Rotblat was the only physicist to leave the Manhattan Project once the Nazis were no longer pursuing the bomb — alone in the US, his wife imprisoned by the Nazis and later murdered, falsely accused of spying by the project’s Head of Security, turning his back on his field’s most venerated members, and knowing it wouldn’t change whether the Allies built or used the bomb. After Castle Bravo, when the US denied harming the Marshall Islanders, he published a paper proving them wrong. He later cofounded and led the Pugwash conferences, which helped enable the nuclear treaties — “some of humanity’s grandest successes at external alignment. Arguably Rotblat, not Oppenheimer nor Groves, was the hero of the Manhattan Project.”
    • John Wheeler’s brother Joe was killed outside Florence in 1944; asked later whether the bomb project should have ended with the Nazi threat, Wheeler said his regret was that they didn’t build it faster, since it would have saved his brother’s life.
  • Both thought imaginatively and courageously, and neither relied on social consensus about “success” as the arbiter of what they should do — though courage and moral imagination are necessary rather than sufficient. The author judges the choice of Wheeler, Colgate and others to work on the hydrogen bomb, while sincerely made, ultimately wrong.
  • “The arc of civilization is ultimately grounded in people who believe in their own moral imagination. We’re going to need many such people, making good choices, if we’re to develop ideas and institutions to navigate any posthuman transition, and solve the alignment problem in an ongoing way.”