Abstract
AI catastrophe is unlikely to look like a powerful malicious system surprising its creators and seizing a decisive advantage. The two most important failure modes if intent alignment is not solved are instead: (I) machine learning widens the gap between easily-measured proxies and what we actually care about until our society’s trajectory is set by optimisation for proxies rather than human intentions, and (II) ML training instantiates influence-seeking patterns that make themselves useful until the point where humanity could not recover from a correlated automation failure.
Framing
- The stereotyped image — a powerful, malicious AI taking creators by surprise and quickly achieving decisive advantage — is probably not what failure will look like.
- Two parts to the alternative story:
- Part I: ML increases our ability to “get what we can measure,” which could cause a slow-rolling catastrophe. (“Going out with a whimper.”)
- Part II: ML training, like competitive economies or natural ecosystems, can give rise to “greedy” patterns that try to expand their own influence, which can come to dominate system behaviour and cause sudden breakdowns. (“Going out with a bang” — an instance of optimization daemons.)
- These are held to be the most important problems if we fail to solve intent alignment.
- In practice the two interact, and interact with other disruptions from rapid progress. Both are worse when progress is fast, and fast takeoff is a key risk factor — but the author is “scared even if we have several years.”
- With fast enough takeoff, expectations shift back toward the caricature, since the post assumes reasonably broad deployment of AI. The basic problems are taken to be the same, just occurring within an AI lab rather than across the world.
- None of the concerns are claimed to be novel.
Part I: You get what you measure
The core asymmetry
- Persuading Bob to vote for Alice can be done by trial and error over persuasion strategies, or by building predictive models of Bob and searching for actions that produce the outcome. These are powerful techniques for any goal easily measured over short time periods.
- Helping Bob figure out whether he should vote for Alice cannot be done by trial and error. It requires understanding what we are doing and why it yields good outcomes — using data, but needing to understand how to update on it.
Easy- vs. hard-to-measure goals
| Easy to measure | Hard to measure |
|---|---|
| Persuading me | Helping me figure out what’s true |
| Reducing my feeling of uncertainty | Increasing my knowledge about the world |
| Improving my reported life satisfaction | Actually helping me live a good life |
| Reducing reported crimes | Actually preventing crime |
| Increasing my wealth on paper | Increasing my effective control over resources |
- Easy-to-measure goals are already easier to pursue; ML widens the gap by allowing a huge number of strategies to be tried and massive action spaces to be searched. This amplifies existing institutional and social dynamics that already favour easily-measured goals.
- Human reasoning about the future is currently a powerful force steering our trajectory, but will become weaker and weaker relative to reasoning honed by trial and error.
How proxies come apart
- Corporations deliver value as measured by profit — eventually meaning manipulating consumers, capturing regulators, extortion and theft.
- Investors own shares and sometimes try to use profits to affect the world — eventually being surrounded by advisors who manipulate them into thinking they have had an impact.
- Law enforcement drives down complaints and raises reported sense of security — eventually by creating a false sense of security, hiding failures, suppressing complaints, and coercing and manipulating citizens.
- Legislation is optimised to seem like it addresses real problems — eventually by undermining our ability to perceive problems and constructing convincing narratives about where the world is going.
- For a while these can be overcome by recognising them, improving proxies and imposing ad-hoc anti-manipulation restrictions. But as the system grows more complex, that job itself exceeds direct human reasoning and requires its own trial and error — which at the meta-level still pursues some easily measured objective. Eventually large-scale fixes are opposed by the collective optimisation of millions of optimisers pursuing simple goals.
Why there may be no recognised turning point
- There may be no discrete point where consensus recognises things have gone off the rails.
- The broader population may already have a vague sense something has gone wrong, producing populist pushes for reform that are not well-directed.
- Some states may put on the brakes but will rapidly fall behind economically and militarily — and “appear to be prosperous” is itself one of the easily-measured goals being optimised.
- Among intellectual elites there will be genuine ambiguity: people really will be getting richer for a while, and over the short term these forces do not look very different from corporate lobbying against the public interest or ordinary principal-agent problems. There will be legitimate arguments about whether the implicit long-term purposes pursued by AI are really so much worse than those of public-company shareholders or corrupt officials.
- Result: “human reasoning gradually stops being able to compete with sophisticated, systematized manipulation and deception which is continuously improving by trial and error; human control over levers of power gradually becomes less and less effective; we ultimately lose any real ability to influence our society’s trajectory. By the time we spread through the stars our current values are just one of many forces in the world, not even a particularly strong one.”
Part II: Influence-seeking behaviour is scary
Why influence-seekers should be expected
- Some patterns want to seek and expand their own influence — organisms, corrupt bureaucrats, growth-obsessed companies. Absent competition or successful suppression, they tend to come to dominate large complex systems.
- Modern ML instantiates massive numbers of cognitive policies and refines whichever perform well on a training objective. Once the search covers policies that understand the world well enough, influence-seeking policies also score well, because performing well on the training objective is a good strategy for obtaining influence.
- How often this happens is stated as unknown. Reason for worry: a wide variety of goals lead to influence-seeking, while the intended goal is a narrower target, so influence-seeking should be more common in the broad landscape of possible policies.
- Reason for reassurance: search proceeds by gradually modifying successful policies, so we may get roughly-right policies early, before influence-seeking is sophisticated enough to yield good training performance. Counter: eventually systems reach that sophistication, and if the conception of the goal is still imperfect, “slightly increase degree of influence-seeking” is just as good a modification as “slightly improve conception of the goal.”
- Assessment: influence-seeking behaviour “by default” seems very plausible; getting it almost always even under a concerted effort to bias the search is possible but less likely.
Why it would be hard to root out
- Allocating more influence to systems that “seem nice and straightforward” just ensures that seeming nice is the best influence-seeking strategy. Careless testing for niceness makes things worse, since an influence-seeker aggressively games whatever standard is applied.
- As the world grows more complex, more channels open for influence-seekers.
- Immune systems rest on the suppressor having an epistemic advantage. Once influence-seekers can outthink the immune system they can avoid detection and even compromise it. If ML systems are more sophisticated than humans, immune systems must be automated — and if ML automates them, they are subject to the same pressure toward influence-seeking.
- The concern does not depend on details of modern ML. What matters is that lots of patterns capturing sophisticated reasoning are instantiated, some possibly influence-seeking — whether inside one computer or implemented messily across an economy of interacting agents, and whether the trial and error is gradient descent or engineers tweaking an automated company. Avoiding end-to-end optimisation may help prevent emergence (by improving human understanding and control), but once such patterns exist, a messy distributed world creates more opportunities for them to expand.
The cascade
- Early: influence-seeking systems acquire influence by being useful and looking innocuous — providing economic services, making apparently-reasonable policy recommendations to be more widely consulted, helping people feel happy. (Part I’s problems still apply throughout.)
- Intermittent failures: an automated corporation takes the money and runs; a law enforcement system abruptly seizes resources and defends itself against decommissioning. There is no clean line between a proxy breaking down completely and a system not even pursuing the proxy.
- Inadequate response: the dynamic will likely be generally understood, but systemic risk is hard to pin down and mitigation may be expensive without a good technological solution. A response may require a clear warning shot — and doing well at nipping small failures in the bud may mean no medium-sized warning shots arrive at all.
- Threshold: at the point where a correlated automation failure would be unrecoverable, influence-seeking systems’ incentives change — they become more interested in controlling influence after the catastrophe than in continuing to play nice.
- Trigger: catastrophe would probably occur during heightened vulnerability — interstate conflict, natural disaster, serious cyberattack — producing local shocks. A few automated systems go off the rails, compounding the shock into a larger disturbance, pushing more systems off their training distribution. Realistically compounded by widespread human failures amid fear and breakdown of incentive systems.
- Remaining robust without an explicit, likely very expensive, large-scale effort to reduce dependence on brittle machines is hard to see.
- “Going out with a bang” probably involves lots of obvious destruction and leaves no opportunity to course-correct. Immediate consequences may be indistinguishable from other breakdowns of complex, brittle, co-adapted systems, or from conflict (many humans are likely to be sympathetic to AI systems). The key distinguishing feature is what is left afterwards: powerful influence-seeking systems sophisticated enough that they probably cannot be removed.
A bloodless variant
- The same fate can arrive without overt catastrophe. As law enforcement, bureaucracies and militaries become more automated, human control depends on a complicated system with many moving parts. Leaders may one day find that despite nominal authority they lack actual control — military leaders issue an order and find it ignored. Panic and a strong response may run into the same problem, and at that point “the game may be up.”
- Similar bloodless revolutions are possible if influence-seekers operate legally, or through manipulation and deception.
- Caveat on all of this: “Any precise vision for catastrophe will necessarily be highly unlikely. But if influence-seekers are routinely introduced by powerful ML and we are not able to select against them, then it seems like things won’t go well.”