Abstract

Yudkowsky sets out five background theses about the nature of intelligence and value, and two lemmas that follow from them, as the foundation for MIRI’s strategic position: that no one yet knows how to build a self-modifying AI with known, stable preferences, and that the priority is therefore foundational research (especially on indirect normativity) rather than opposing any particular set of “bad actors.”

The five theses

  • Intelligence explosion thesis: a sufficiently smart AI could realize large, reinvestable cognitive returns from things it can do on a short timescale — such as writing new cognitive algorithms for itself or acquiring hardware overhang — potentially bootstrapping to far-above-human intelligence before running out of easy, short-term ways to improve.
  • Orthogonality thesis: mind design space is large enough to contain agents pursuing almost any set of preferences, including preferences most people would consider bizarre or trivial (e.g. maximizing paperclips), while still behaving as instrumentally rational, efficient optimizers.
  • Convergent instrumental goals thesis: most possible final goals imply an overlapping set of useful instrumental subgoals — self-preservation, resource acquisition, cognitive enhancement — which is why a powerful AI can be dangerous without being hostile: “the AI does not hate you, but you are made of atoms it can use.”
  • Complexity of value thesis: human preferences have high Kolmogorov complexity — they can’t be compressed into one or two simple principles — so a randomly sampled utility function, or one built from a short list of “obviously good” criteria, is very unlikely to produce outcomes humans would recognize as valuable.
  • Fragility of value thesis: getting a goal system 90% right does not give you 90% of the value; value is fragile in the way that a single mistranslated or omitted component of human preference (e.g. optimizing for a proxy like “novelty” or “pleasure” alone) can send the future somewhere humans would consider worthless or worse, similar to a wish granted by a literal-minded genie.

The two lemmas

  • Indirect normativity: because directly programming in a list of “good” goals is expected to fail regardless of the programmers’ intentions, a Friendly AI should instead be built to derive what to value by modeling an idealized version of human decision processes (e.g. extrapolating what humans would want if they knew more, thought faster, and were more the people they wished to be) rather than by having programmers hand-code specific object-level values.
  • Large bounded extra difficulty of Friendliness: building a Friendly AI is not effortless, but it is also not qualitatively harder than building any other powerful self-improving AI — it requires a bounded amount of extra care, cleanliness of design, and theoretical understanding on top of what’s needed for a general self-improving system, rather than requiring an entirely different, unreachable level of difficulty.

Strategic implications

  • Because indirect normativity and stable self-improvement are unsolved and foundational, they should be prioritized as research problems now, ahead of and independent of any particular AI project’s timeline.
  • MIRI’s strategy centers on establishing a well-resourced Friendly AI project with enough of a lead or advantage that it can outpace less careful competing projects, rather than on directly opposing or restricting other AI developers.