Abstract
So8res (Nate Soares) argues that AGI ruin is the disjunctive default outcome — many separate, independent paths lead to catastrophe — while avoiding it is conjunctive, requiring a long list of demanding conditions about world state, technical alignment, and organizational behavior to all hold simultaneously; since these conditions are correlated through general real-world incompetence rather than independent, multiplying their individual failure probabilities understates rather than overstates the risk.
Framing: disjunctive doom, conjunctive survival
- Rejects the framing that catastrophe is a narrow target: “most of the outcome space is full of AGI ruin, and avoiding it is what requires navigating a treacherous and narrow course.”
- Presents a simplified toy model listing conditions that must all hold for humanity to survive AGI development, organized into three categories: world state, technical alignment, and organizational dynamics.
World state requirements
- A deployment strategy for AGI must exist that is compatible with what early-stage alignment techniques can realistically achieve.
- Leading AI organizations must know about and accept that strategy.
- Those organizations need enough time to develop, align, and deploy AGI safely.
- Either a single organization must maintain a lead in AGI capabilities for years, or all capable organizations must cooperate cautiously with one another.
- Governments must avoid destabilizing actions before or during AGI development.
Technical alignment requirements
- Dedicated alignment researchers must be genuinely integrated into AGI development teams.
- Those teams must identify every single lethal problem far enough in advance to actually solve it.
- Alignment research must be productive and stack cleanly, or else teams need substantially more time than is likely available.
- Work must succeed without direct access to an actual AGI to study, or else the world must delay deployment of misaligned AGI long enough for solutions to catch up.
Organizational dynamics requirements
- Development teams must genuinely prioritize alignment, rather than adopting philosophies that treat it as secondary.
- Internal structures must be able to distinguish real safety solutions from fake ones, despite technical disagreement and competitive pressure.
- Teams must be able to detect dangerous warning signs as they arise.
- Key personnel raising concerns need enough social capital within their organizations to act on them.
- Organizations must prevent proliferation of dangerous capabilities through splinter groups or leaks.
- Teams must avoid prematurely releasing insights that would accelerate less careful competitors.
Why correlation strengthens rather than weakens the argument
- An obvious objection: treating these conditions as independent and multiplying their individual failure probabilities overstates total risk if the conditions are actually correlated.
- So8res agrees they’re correlated — but argues this makes the outlook worse, not better, because “they’re especially correlated through the fact that the world is derpy”: the same institutional incompetence tends to show up across unrelated domains.
- Illustrates with the COVID-19 pandemic response: “the US federal government’s response to COVID was to ban private COVID testing, confiscate PPE bought by states, and warn citizens not to use PPE.”
- Concludes that competence failures cluster together — the ability to solve one of these problems is evidence of the ability to solve the others, and current evidence points toward broad incompetence rather than broad competence.
The competence floor and warning-shot skepticism
- The single variable he sees as most decisive: whether humanity will “rapidly become much more competent about AGI than it appears to be about everything else” — which he judges unlikely given patterns in COVID response, political dynamics, and AI policy discourse.
- Notes a recurring dynamic where technical and policy communities each assume the other will solve the hard parts, while doubting their own domain’s ability to do so.
- Dismisses hope that a “warning shot” will snap civilization into a competent response, comparing it to “wishful thinking from the world where the world has a competent response to the pandemic warning shot we just got.”
Where the remaining hope comes from
- Uncertainty about whether his own model is accurate.
- The possibility of unknown factors changing the strategic landscape.
- A small residual chance that, despite the above, humanity manages to thread the needle regardless.
- Explicitly frames the piece as a simplified toy model — his actual views are described as “more subtle, take into account more factors, and are less articulate” than the model captures.