Abstract

The alignment community set out to differentially advance safety over capabilities, but over the past decade instead accelerated capabilities while making little scientific progress on alignment. This is the first post in a six-part sequence tracing that outcome to a shift away from conceptual clarity and toward chasing prestige, funding, and political influence — with agent foundations singled out as the subfield that kept doing genuine science.

Central pattern: “jumping down the slippery slope”

  • A recurring cognitive move in the field’s history: treating an outcome as inevitable, and then becoming a major force that brings that outcome about.
  • Example given: Sam Altman reasoning that since AGI development is inevitable, an organization like OpenAI should be the one to do it first — a stance the author reads as self-fulfilling rather than merely predictive.

Conceptual clarity and scientific progress

  • The early rationalist community, drawing on transhumanist and economics-blogging circles, showed unusual intellectual clarity. Figures cited include Wei Dai and Hal Finney (cryptocurrency), Robin Hanson (prediction markets), Scott Alexander (game-theoretic political analysis), Nick Bostrom (anthropics), and Eliezer Yudkowsky with Wei Dai (decision theory).
  • Scientific progress, on this view, comes from developing new concepts that link into coherent ontologies — internalizing a simple idea and building further theoretical layers on top of it. The author argues this capability was largely absent elsewhere.
  • Mainstream ML prioritized engineering over understanding, a tendency the rise of deep learning worsened by substituting compute scaling for conceptual development. The author’s framing: alignment as a field exists precisely because ML never prioritized deeply understanding the systems it was building.
  • Academic science is diagnosed with a “streetlight effect” — optimizing for what’s publishable rather than what’s important. Cited examples: RL theory’s focus on unrealistic tabular settings, and statistical learning theory’s emphasis on underparameterized regimes that poorly explain neural network generalization.
  • Distinction drawn between “generative” science (developing new concepts) and “discriminative” science (evaluating existing frameworks); academia is said to over-emphasize the latter.

Orienting towards prestige

  • Early links between rationalists and Silicon Valley elites formed through Peter Thiel and Jaan Tallinn. Nick Bostrom’s Superintelligence (endorsed by Bill Gates) was an attempt to make AGI risk a prestigious topic.
  • The author argues this outreach backfired: the assumption that “competent people out there should be recruited as allies” proved wrong, since prestigious figures turned out to reason about AGI less clearly than anonymous LessWrong commenters.
  • OpenAI’s founding — Elon Musk and Sam Altman’s response to AGI concerns — is presented as a sign something had already gone badly wrong, since making AGI development more “open” ran directly counter to what early rationalists wanted.
  • Contrasted with this is Eliezer Yudkowsky’s Harry Potter and the Methods of Rationality, which the author credits with recruiting people capable of clear alignment thinking more effectively than prestige-oriented outreach did.

Open Philanthropy’s funding decisions

  • Holden Karnofsky initially rejected MIRI’s arguments in 2007 for lacking prestigious endorsement, then reversed course after Superintelligence gained mainstream acceptance and began funding AI safety in 2015.
  • The resulting funding allocation is described as calibrated to conventional prestige rather than predictive accuracy: MIRI received $500,000 in 2016 (characterized as a “participation grant”), the Future of Life Institute $1 million, Stuart Russell’s CHAI $5 million, university-affiliated groups roughly $10 million combined, and OpenAI $30 million — the latter in exchange for a board seat for Holden.
  • Daniel Dewey’s stated reasoning for not funding agent foundations — that it “has not gained much support among AI researchers” and so wouldn’t attract new people — is characterized as circular, deferring to the judgment of people not yet convinced AGI risk is real.
  • The $30 million to OpenAI (roughly 3% of the organization’s donations at the time, against Open Philanthropy giving away some 20% of its total charitable funding) is offered as another instance of “jumping down the slippery slope”: treating an outcome as justified by counterfactual reasoning while it violates the community’s own stated ethical standards.
  • The professional and personal connections between Holden Karnofsky, Dario Amodei, and Paul Christiano are argued to have shaped what “AI alignment” became as a field, steering subsequent development toward their approaches rather than toward agent foundations.

Assessment of alignment subfields

  • Agent foundations is identified as the only alignment subfield that consistently pursued genuine scientific progress.
  • Mechanistic interpretability is described as promising but still entangled with capabilities advancement.
  • Most AI governance interventions are judged likely to backfire given how adversarial contemporary politics is; only outcomes with robust backing — such as building justified trust between key actors — are considered worth supporting.

Broader diagnosis of ML and academia

  • Even mainstream ML research is said to prioritize engineering over understanding, which is offered as the underlying reason alignment had to exist as a separate field.
  • A 1995 NIPS commentary is cited as early criticism of “publish or perish” incentives displacing genuine scientific inquiry in favor of mathematical theory chosen for provability over importance.

Structure of the sequence

This post is the first of six:

  1. Conceptual clarity and scientific progress
  2. Orienting towards prestige
  3. Pragmatism and pessimization
  4. Conforming to the ML community
  5. Fear and anticipatory obedience
  6. Déjà vu

Closing assessment

  • The author frames the sequence’s overall message as negative, but maintains that the alignment community still shows more clarity and sincerity than comparable intellectual communities.
  • The future is described as “up for grabs,” with visible paths to very good outcomes, blocked mainly by the community’s difficulty learning from its own past mistakes.