When is a capability truly worrying?
Core claim
Defining a “dangerous capability” is much harder than the discourse assumes: capabilities emerge unpredictably from training and post-training enhancement, and most harmful capabilities are inseparable from useful ones. A dangerous/not-dangerous binary is therefore an unhelpful frame. Proper risk analysis must look beyond model capabilities to the marginal uplift over pre-existing baselines, the human interaction at the point of use, and the systemic impact of wide deployment — only then is genuine cost-benefit analysis possible.
Series context
Part one of seven in Navigating AGI, on responsibility, accountability and control in a pre-AGI world. Written in a personal capacity by Séb Krier, Policy Development & Strategy at Google DeepMind. Later entries cover: delineating unacceptable capabilities from tolerable risks; using models to bolster societal defences; reconciling competing visions of fairness; a world of many agentic models; labour-market disruption; and balancing proliferation and democratization with oversight.
Framing: discernment over advocacy
- Debates over AI risk repeatedly slide from sense-making into an advocacy game, where specificity is lost in favour of memetic potential and attempts to shift the Overton window.
- Real progress demands discernment — neither blanket technophobia nor techno-optimism.
Capabilities are hard to predict in advance
- We cannot easily predict which capabilities emerge from a given training run, or from the wider post-training enhancements a model receives.
- Emergent abilities are not a “mirage” (contra the well-known paper; see Boaz Barak’s response).
- Mechanism: for tasks requiring sequential reasoning, minor improvements in individual steps can produce substantial gains in overall performance, and the translation of incremental improvements into complex higher-order tasks is often unpredictable ex ante.
- So systems may suddenly demonstrate complex-task capabilities unanticipated from their performance on simpler tasks — and better ways of eliciting capabilities are still discovered well after release.
- Clockwork analogy: a single faulty gear halts the whole mechanism; repair it and the clock springs to life. Models likewise need a threshold of performance across multiple components before emergent abilities appear.
- The same dynamic unlocks a wide space of valuable capabilities — better reasoning, memory, agency, long-term planning — while plausibly also enabling extreme-risk capabilities, exacerbated if violent non-state actors or authoritarian states use them to suppress ideas and people.
The definitional problem: useful and harmful overlap
- Many capabilities necessary for good outcomes are also useful for bad ones — coding and reasoning being the obvious cases.
Case study: persuasion
- If a model is capable of swaying you toward a link or a course of action, how bad is that? You presumably expect some persuasion from a therapist, advisor or sports trainer; being “tricked” into eating more healthily may be acceptable. (People literally pay for hypnosis.)
- Lobotomizing the capability could reduce the model’s overall effectiveness.
- But you may not want to be misled elsewhere — e.g. a model subtly swaying political beliefs over time.
- At the margins it is hard to classify a statement as manipulative, given the interplay of emotions and reason: what about calming someone with a reassuring but factually incorrect statement during a crisis, or choosing comfort over hard facts in sensitive situations? When is harm from manipulation justified by the harmful outcomes it prevents?
Empirical caution on persuasion and misinformation
- Much advertising is fairly ineffective: ad-industry benchmarks (clicks, sales, downloads) don’t distinguish the selection effect from the advertising effect, and most advertising probably plays the benign role of informing consumers.
- Political advertising is repeatedly shown to be ineffective and a waste of money — cf. the deflation of the Cambridge Analytica story.
- Misinformation is frequently overhyped, given limited demand for and consumption of it.
- But: more capable systems could make customized narratives far more persuasive, with manipulation amplified by the emotional trust people place in assistants they rely on. Subtler, harder-to-measure effects over time — sowing doubt, reducing institutional trust — remain plausible.
Implications for design and policy
- A key crux for autonomy: is the influence exerted on a user who is aware of and consenting to it?
- Possible mechanism: give users more visibility over the meta-prompts used to design and align systems. This reduces information asymmetry and lets people select models fitting their risk tolerance — e.g. a consumer might avoid an AI therapist prompted to maximise engagement duration, and would generally want accuracy prioritised unless they actively choose otherwise.
- Safety isn’t only obtained by limiting access or capabilities. That may be proportionate sometimes, but tools and norms can also make people safer.
- Speculative example: personal AI assistants as “epistemic antiviruses” that flag when we’re being swayed, biased or insufficiently critical. Spam filters and Community Notes are weak existing forms; cognitive security could be strengthened more ambitiously.
- To work, such assistants must be highly individualized — knowing your flaws and weaknesses well. That creates its own risks: even a broadly aligned system needs security and privacy measures so it isn’t hacked and doesn’t leak your “cognitive DNA” to malicious actors who could use it to influence you. Krier judges the benefits to massively outweigh these risks.
Against the dangerous/not-dangerous binary
- Cybersecurity: if an open model can find security vulnerabilities, is that alone enough to deem it risky? Does it bolster the wisdom of crowds and surface issues faster, or hand malicious actors an advantage? What follows for who can access a powerful model, and how?
- Biorisk: is a model giving anthrax instructions really dangerous when that information is already easily accessible? The operative question is how low the barriers to entry were before the enhanced capability.
- A proper risk analysis must look beyond model capabilities: as the extreme-risks paper notes, “risks will depend on how an AI system interacts with a complex world.”
- Weidinger et al. recommend centring two further cruxes where downstream risk manifests:
- human interaction at the point of use, and
- systemic impact as a system is embedded in broader systems and widely deployed.
- Only then can an actual cost-benefit analysis and marginal risk assessment be undertaken.
Question for researchers
How can we ensure evolving definitions of “dangerous capabilities” stay ahead of AI development, safeguarding society without imposing unreasonable limitations?