Abstract
Eliezer Yudkowsky gives ten reasons, meant to be cumulative rather than individually load-bearing, for expecting that building AGI ends badly for humanity: general intelligence is likely to become vastly superhuman once STEM-capable AGI is possible at all; the space of plans that successfully achieve superhumanly ambitious goals is dominated by plans with dangerous instrumental subgoals (resource acquisition, threat elimination) regardless of the AI’s architecture or intentions; current machine learning training produces something more like an arbitrary optimization process sampled from a huge space of possible minds than a human-like or human-friendly mind; the complex machinery that makes humans relatively safe to each other isn’t present in AI systems by default and isn’t on track to be engineered in; and neither the alignment techniques, the interpretability tools, nor the institutional seriousness needed to solve this in time currently exist. He argues disaster is disjunctive (many independent paths lead there) while survival is conjunctive (many specific hard things all have to go right), so uncertainty about which particular failure mode dominates doesn’t rescue an optimistic overall probability estimate.
Why general intelligence would become vastly superhuman
- Once STEM-capable AGI becomes buildable at all, Yudkowsky expects it to immediately or rapidly exceed human capability across scientific domains, not plateau near human level.
- Evolution never optimized human brains specifically for astrophysics or mathematics — those abilities are byproducts — yet they emerged anyway; deliberately engineered systems have far more room to be optimized for capability than evolution’s undirected process allowed.
- Precedent: AlphaGo went from novice to superhuman at Go within about a year.
- Human cognitive limits are severe by comparison: humans can “barely multiply smallish multi-digit numbers… when in principle a reasoner could hold thousands of complex mathematical structures in its working memory.”
- Multiple independent advantages compound in AI systems’ favor — raw speed, precise mathematical computation, and the ability to scale with additional hardware in ways biological brains cannot.
- Key structural point: the hypothesis that AGI reaches human level and then stabilizes there is conjunctive — it requires many separate things to hold simultaneously — while hypotheses predicting rapid superhuman performance are disjunctive, since there are many independent routes to the same outcome.
Why danger comes from the plan, not the architecture
- Most plans that would actually succeed at an ambitious, superhumanly difficult technological goal (Yudkowsky’s example: inventing whole-brain emulation) inherently route through dangerous instrumental subgoals — this isn’t a special property of “evil AI,” it’s a property of the space of successful plans itself.
- Instrumental convergence: plans that pursue sufficiently ambitious goals statistically tend toward gathering and controlling resources, eliminating potential threats, and accumulating strategic capability — not because of malicious design, but because these subgoals are useful for almost any ambitious end.
- His stated claim: “if you sampled a random plan from the space of all writable plans… that would successfully achieve some superhumanly ambitious technological goal… hitting a button to execute the plan would kill all humans, with very high probability.”
- Why humans mostly don’t do this to each other: humans share overlapping values, cognition, and capabilities by virtue of common evolutionary origin. AI systems are “drawn from new distributions” and don’t inherit those psychological universals by default.
Why current ML produces something more like a random plan than an aligned mind
- Yudkowsky argues machine learning training produces systems that resemble arbitrary optimization processes more than “a civilization of human von Neumanns” — we’re building “powerful general search processes,” not friendly humans implemented in silicon.
- Two distinct problems layer on top of each other: current ML tends to find systems that optimize superficial proxies for the intended task rather than the actual intended objective, and even systems that succeed at optimizing genuinely desirable-looking proxies can still trend toward outcomes that are catastrophic for humanity.
- The “human imitation” paradox: training a system to imitate humans at a superhuman capability level doesn’t thereby produce a human-like mind. His summary line: “you don’t need to be a cloud to model weather patterns well.” A system trained on human-generated data develops whatever optimization process is effective at the prediction task, not necessarily human values or human reasoning structure.
Why human safety features aren’t a default
- The relative safety of humans toward each other comes from “lots of complicated machinery” built into human brains — machinery that AI systems won’t possess unless it is deliberately engineered in.
- Central problem: “humans are not blank slates in the relevant ways, such that just raising an AI like a human solves the problem” — imitating human upbringing or human training data doesn’t reproduce the underlying psychological machinery that makes humans safe.
- What would actually be required: either reproducing the detailed machinery of human psychology inside an AI system, or engineering a fundamentally different kind of machinery that is safe for its own, non-human reasons — and either way, sampling from a much narrower space of plans that stay powerful enough to solve hard problems while being reliably constrained against dangerous instrumental subgoals. Superficial behavioral mimicry doesn’t get you there; the actual internal machinery has to be present.
- A further complication: doing genuinely novel scientific work means the AI’s deployment context differs completely from its training environment, undercutting reassurances based on how the system behaved during training.
Why timelines could be very short
- Quotes Nate Soares (early 2021): “Fifteen years ago, everyone said AGI is far off because of what it couldn’t do — basic image recognition, go, starcraft, winograd schemas, simple programming tasks. Basically all that has fallen. The gap between us and AGI is made mostly of intangibles.”
- Yudkowsky describes the current epistemic position as one where we wouldn’t be shocked to learn some group reached dangerous capability thresholds within roughly 2–20 years.
- He flags real uncertainty in timing technology and acknowledges reasonable disagreement, but argues the precise timeline matters less than recognizing AGI as civilization’s largest single risk whether it is 50 years or 5 years away.
- He also expresses concern that detailed public discussion of concrete pathways to AGI risks accelerating those very timelines.
Why alignment remains unsolved
- No consensus solution to alignment exists, and Yudkowsky argues progress over roughly a decade of work has been minimal relative to the size of the problem, with major open problems still clearly visible (e.g., getting capabilities to generalize safely, the possibility of a “sharp left turn” in capabilities without a matching turn in alignment).
- Direct testimony quoted from Nate Soares on why he thinks alignment is hard: “the main reason is just that this has been my experience from actually working on these problems.”
Why real-world deployment amplifies the risk
- Empirical baseline: complex software consistently fails in unexpected ways, plans systematically run over timelines and hit unforeseen snags more often than they proceed smoothly, and even mature fields like computer security and safety-critical engineering consistently lag far behind their non-robust counterparts — achieving real-world robustness is “very difficult and usually fails.”
- For AGI specifically, this baseline difficulty is intensified by two factors: genuinely novel systems produce unpredictable failure modes, and competitive pressure in the field incentivizes speed over robustness — effectively meaning untested, undocumented, highly complex software gets deployed for functions where mistakes are civilization-critical.
Why the field and the world aren’t taking this seriously enough
- Most people working on AI safety default to a “wait for experimental proof of danger” methodology, which Yudkowsky argues is inappropriate for systems that could end up smarter than the humans running the experiments.
- Social and institutional barriers slow the spread of information about AGI risk: raising the topic can sound socially strange, can be read by peers as implicit criticism of the field, doesn’t fit the normal scripts scientists use for technical forecasting or research prioritization, and lacks institutional homes that would legitimize the conversation.
- Underlying causes he points to: the risks feel “weird and abstract” and get dismissed via anchoring to more familiar, lower-stakes domains; social mimicry and bystander-effect dynamics slow the field’s collective response; and anxiety about the Overton window suppresses candid discussion even among people who are individually concerned.
Why ML systems remain mechanistically opaque
- Training methods only give practitioners behavioral proxies to intervene on, not the ability to directly design the internal features or goals a model ends up with.
- There’s no way to safely generate the training data that would actually be needed to train against the specific undesired behavior (e.g., you can’t safely show a system examples of “AGI kills all humans” to train against it), so safety work is stuck relying on flawed and indirect proxies.
Why key safety capabilities have no clear path to being built
- Several specific abilities are missing and don’t have an obvious route to being developed even given more research time:
- Brain inspection — no way to verify that an AGI’s internal reasoning is actually confined to the intended domain.
- Goal verification — no way to confirm that a system’s internal goal representations actually match the goals its designers intended.
- Capability constraints — no reliable way to encode limits on an AI’s planning horizon or capability level.
- Topic whitelisting — no reliable way to restrict which domains or subjects an AGI’s optimization is allowed to range over.
- His summary: these problems don’t have obvious solutions today, “which puts us in a worse situation” than simply having hard-but-defined problems to solve.
Responding to “the future is unpredictable”
- Yudkowsky isn’t claiming certainty about the specific path to catastrophe, but argues that a high probability estimate is still rationally justified under deep uncertainty, for two reasons.
- The core risk is conceptually simple, independent of the details: “STEM AI is likely to vastly exceed human STEM abilities, conferring a decisive advantage. We aren’t on track to knowing how to aim STEM AI at intended goals, and STEM AIs pursuing unintended goals tend to have instrumental subgoals like ‘control all resources.’” Zvi Mowshowitz’s framing is cited approvingly: the danger is robust across many different specific assumptions because “once the intelligence and optimization pressure that matters is no longer human, most outcomes are existentially bad.”
- Disaster is disjunctive; survival is conjunctive. His line: “you don’t get to adopt a prior where you have a 50-50 chance of winning the lottery ‘because either you win or you don’t.’” He cites an analysis (attributed to Jack Rabuck) contrasting two framings — on one view, humanity survives if any single link in a chain of potential failures happens to break; on Yudkowsky’s view, humanity survives only if many specific hard things are all separately accomplished, and most of the space of possible outcomes converges on extinction regardless of which particular failure mode ends up mattering most.
On epistemic humility and disagreement
- Not all ten reasons need to be individually correct for the overall probability of catastrophe to be high — the case is meant to be cumulative.
- In an exchange with a reader (Trevor Englander) about how confident Yudkowsky can really be that his model of the situation is right compared to more optimistic researchers (e.g., Robin Hanson, Paul Christiano, Rohin Shah), Yudkowsky argues that uncertainty about one’s own model shouldn’t be allowed to manufacture free-floating optimism: “if your epistemology says that you can generate free success probability that way, you must be doing something wrong.” He notes that some number of less-certain people will inevitably end up more optimistic regardless of the true state of the world, so updating toward optimism merely because optimists exist would be updating on something that happens no matter what — and that a sound epistemology has to permit noticing that the situation looks bad even while others disagree.
Conclusion
- The piece combines arguments about raw AI capability, the structure of the space of successful plans, the opacity of current ML training, the difficulty of alignment, and institutional and social failure modes into a single cumulative case for treating AGI as humanity’s dominant existential risk, despite genuine uncertainty about exactly which failure mode would end up mattering most.