Abstract

Over the past decade, alignment researchers repeatedly used consequentialist “it accelerates safety on net” arguments to justify work that in practice accelerated capabilities, until the distinction between “alignment research” and “capabilities research” became meaningless. This post — third in the “What just happened?” sequence — traces that pattern through OpenAI, DeepMind, and Anthropic, diagnoses it as failures at both the individual and group level, and proposes four principles for how the field should reason and act differently going forward.

Two failure modes: individual and group

  • Individual level: consequentialist reasoning about one’s own marginal impact is easy to rationalize, because “there are many possible scenarios for how the future could play out.” External incentives — especially inside AGI companies — pushed researchers toward whichever scenarios happened to justify their preferred research direction, producing pervasive motivated reasoning about counterfactual impact.
  • Group level: the alignment community never built accountability mechanisms to check this. “Extreme fear of publicly criticizing powerful people” combined with strong charitability and mistake-theory norms blocked the formation of common knowledge about who was doing motivated reasoning or straightforwardly lying. Even Sam Altman and Dario Amodei were given “benefit of the doubt for many years” despite pursuing “enormously power-seeking strategies.”

The prosaic ideal versus the pragmatic reality

  • Paul Christiano’s “prosaic AI alignment” framework held that AGI could be aligned without a fundamentally new understanding of intelligence — but the author argues it never cashed out into a defensible research strategy.
  • Christiano positioned himself between two camps: one holding it’s impossible to do meaningful safety work without knowing more about what powerful AI will look like, the other holding that aligning prosaic AGI is probably infeasible.
  • The author argues Christiano made versions of both the individual- and group-level mistakes: joining OpenAI meant his empirical predictions were read as endorsing OpenAI’s capabilities acceleration, and he “conspicuously failed to correct this impression by critiquing OpenAI publicly.”

OpenAI: scalable oversight decoupled from RLHF

  • Around 2018, three scalable-oversight proposals circulated: Paul Christiano’s iterated amplification, Geoffrey Irving’s debate, and Jan Leike’s recursive reward modeling.
  • The actual engineering that shipped — reinforcement learning from human feedback (RLHF) — became decoupled from these theoretical justifications: “although Paul’s theoretical justifications for iterated amplification referred a lot to the safety properties of imitation learning, all of these papers added RLHF for better performance.”
  • Dario Amodei pushed the scaling path from GPT-1 through GPT-3, justifying it partly through alignment arguments and partly through competitive positioning against China; the resulting scaling-laws paper informed the field that AGI was plausible while undercutting any claim to be following a coherent safety strategy.
  • WebGPT is described as an alignment project with weak underlying justification that nonetheless recruited alignment-motivated researchers; ChatGPT inherited its codebase and techniques, effectively becoming “WebGPT minus the Web.” Its launch is called “one of the most acceleratory events in the history of AI, funneling many billions of dollars into the field.”
  • On Paul’s later claim that RLHF wasn’t actually that important: InstructGPT’s 1.3B RLHF’d model outperformed a 175B supervised-only GPT-3, and Sydney/Bing (without RLHF) produced “fairly unhinged outputs” — both cited as evidence RLHF mattered a great deal.
  • The “overhang” argument — that speeding up some input to AI progress prevents a worse future overhang by using up low-hanging fruit — is characterized as “extremely slippery” reasoning: people “pick whichever inputs to AI progress they wanted to defend speeding up, and just assume… that there were other background constraints.” Sam Altman is offered as the clearest example: arguing faster algorithmic progress alleviates compute overhangs while simultaneously planning to raise trillions of dollars for chip fabrication. The author describes their own past sycophancy toward Altman as preventing them from seriously entertaining that he might be “lying about his motivations.”

DeepMind: from embodied AGI to LLMs plus RLHF

  • Shane Legg, an early LessWronger concerned about AGI risk, was sidelined as Demis Hassabis consolidated executive power. Legg founded the Technical AGI Safety Team (TAGIS) and a more secretive “AGI team” pursuing embodied learning in simulation.
  • Geoffrey Irving’s move from OpenAI shifted DeepMind toward LLMs, bringing knowledge of GPT models ahead of their public release; he argued language “crystallized the knowledge of humans,” against Demis’s view that language lacked “grounding.” Irving initiated Gopher (2020) and later Sparrow, which was fine-tuned with RLHF.
  • The author questions the sincerity of the argument that “LLMs were a safer path to AGI than DeepMind’s RL-focused approach”: if that reasoning were genuine, why did alignment researchers at all three major labs independently converge on the same sequence — first scale LLMs, then add RLHF? The pattern looks more like post-hoc rationalization of work people already wanted to do.

Anthropic: alignment and capabilities conflated

  • Anthropic’s framing of building Claude as “tackling alignment directly,” while in practice inventing highly capable natural-language agents, is called an “absurd” conflation of alignment work with capabilities work.
  • Constitutional AI replaced RLHF with RLAIF, delegating more of the verification process to AI systems themselves. Follow-on work on model-written evaluations, model-assisted red-teaming, and model-assisted question-answering created pervasive AI-assisted processes that now make it “hard to reliably track whether or how even current models are deceiving them.”
  • The “harmless” component of the alignment work is argued to have merged technical safety with “ideological control” — emphasis on political correctness and censorship-adjacent categories rather than technical risk — a pattern the author traces to industry-wide “mass censorship against ‘harmful’ ideas and speech” over the prior decade, reproduced inside AGI companies via “Trust and Safety” teams.
  • Mechanistic interpretability under Chris Olah is credited as genuinely valuable science, but described as “dwarfed by the amount of work that blurs the alignment/capabilities line.”

If not alignment research, then what?

  • The term “alignment research” is argued to be too corrupted by this history to remain a useful rallying cry, even though the underlying alignment problem is still real and important — the community has “lost any moral right to try to gain power on altruistic grounds.”
  • The author expresses cautious optimism about the Pause/Stop AI movement, while flagging concerns about its political sophistication, and argues most remaining value comes from “high-integrity scientific research” rather than from mainstream alignment work as currently practiced.

Four principles for reform

  1. Differential-impact reasoning is dangerous. Compared to timing a stock-market bubble: “If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical.” Given how entangled individual choices are with everyone else’s decisions, more robust strategies are preferable to marginal-impact bets.
  2. Assume wider agency. People should act as though choosing on behalf of a much wider range of people than just themselves, because: setting field norms builds high-trust communities; others copy behavior through trust or perceived social permission; underestimating one’s influence is worse than overestimating it; being right about AGI risk correlates with being right about other things; others’ decisions are logically correlated with one’s own; and this follows a broadly Kantian universalizability principle. This principle argues against joining unethical organizations for “harm mitigation” and against racing “because we’re the good guys.”
  3. Articulate and update cruxes publicly. Individuals should state their decision cruxes openly and either acknowledge when those cruxes are disproved or clearly explain what changed. Groups tend to become “accountability sinks,” which makes individual public statements more valuable — e.g., discussing research intentions before executing them, so the community can respond appropriately as strategies shift.
  4. Integrity over power. Outsiders with fewer resources have repeatedly shown more willingness to sacrifice money and prestige than insiders: MIRI kept research non-disclosed-by-default; Janus chose not to publicize chain-of-thought prompting despite discovering it early; Dario released acceleratory papers regardless. Real evidence of integrity comes from “criticizing or standing up to powerful people even when few others around you are doing so” — cited examples include Daniel Kokotajlo’s departure from OpenAI, Yudkowsky’s Time essay, Pause/Stop AI advocacy, and the OpenAI board’s actions.

Self-critique

  • The author recounts discovering they had once downvoted Adam Shimi’s critical post about OpenAI “because, when I first read it, I was scared of the alignment community alienating OpenAI” — offered as a personal example of how sycophancy warped judgment across the field.
  • Alignment researchers are credited with thinking more seriously about impact than most comparable communities, which the author says provides “hope that (some subset of) the current field of alignment is able to learn from its mistakes.” But genuinely acknowledging the error means accepting there is “no longer any widespread (implicit or explicit) definition under which ‘alignment research’… is robustly good for the world, based on the evidence.”
  • The path forward is described as requiring “enough courage and integrity to move towards (emotional, social, and financial) independence from the existing field,” and “clear, open-ended thinking” through this paradigm transition — the community should “halt, melt, and catch fire,” acknowledging its own unreadiness rather than proceeding under false pretenses.