AI Governance: Overview and Theoretical Lenses

Core claim

AI governance — the norms and institutions shaping how AI is built and deployed — should have a scope commensurate with AI’s transformative potential. AI’s impacts can be understood through three lenses: as a general purpose technology, an information technology, and an intelligence technology. Governance can then be analysed as a problem of institutional fit to the externalities AI produces. Adaptation is hardest where governance touches deep social conflicts, and hardest of all under great power security competition — which may ultimately require historically unparalleled global institutions to steer away from the greatest risks.

Publication

Introductory chapter for the Oxford Handbook on AI Governance, May 2022. Dafoe was then Senior Staff Research Scientist at DeepMind and President of the Centre for the Governance of AI.

Definitions and framing

  • AI governance, descriptively: the policies, norms, laws and institutions shaping how AI is built and deployed. Normatively: the aspiration that these promote good decisions — effective, safe, inclusive, legitimate, adaptive.
  • Governance is much more than acts of government; it includes behaviours, norms and institutions from all segments of society.
  • One formulation: the field studies how humanity can best navigate the transition to advanced AI systems.
  • Forecasting context cited: ML researchers put broadly human-level AI at 8% within one decade, 22% within two, and >50% by 2060; Cotra’s (2020) biological-anchors forecast reaches similar estimates.

Four risk clusters

Inequality, turbulence, authoritarianism

  • Declining labour share of value and winner-take-most labour markets could erode the position of labour and the relative equality underpinning democracy.
  • Digitally mediated and AI-filtered communication could increase polarisation, epistemic balkanisation and vulnerability to manipulation.
  • Totalitarianism could be made more robust by ubiquitous physical and digital surveillance, social manipulation, enhanced lie detection and autonomous weapons.

Great-power war

  • Advanced AI could make crisis dynamics more complex and unpredictable, enabling escalation faster than humans can manage — a “flash war” (Scharre, 2018) — increasing risk of inadvertent war.
  • Additional pathways: extreme first-strike advantages, power shifts, and novel destructive capabilities.

The problems of control, alignment, and political order

  • The control problem: human intention controlling what an advanced AI does. The alignment problem: constructing AI agents whose goals are the intentions of the human principal.
  • A reframing for sceptics: the perennial problem of political order is also one of aligning and controlling powerful social entities — corporations, militaries, political parties. That control problem remains unsolved; existing solutions are patchwork and periodically fail (corporate malfeasance, military coups, unaccountable political systems).
  • As AI becomes more capable, autonomous and empowering of particular social entities, the two control problems intertwine and compound.

Value erosion from competition

  • A high-stakes race pushes parties to cut corners on safety — a structural risk (Zwetsloot & Dafoe, 2019) generalisable to any tradeoff between something of value and competitive advantage.
  • Contemporary analogues eroded by global economic competition: sustainability, decentralised technological development, privacy, equality.
  • In the long run, ungoverned competition could pull the future of humanity towards what is most adaptive rather than what is good, with lock-in of impoverished values and forms of life.

Extreme risks and a holistic sensibility

  • Attention to extreme and existential risks helps ensure adequate investment in avoiding worst cases; transformative AI broadens the focus — AI which could “precipitate a transition comparable to (or more significant than) the agricultural or industrial revolution” (Karnofsky, 2016) or lead to “radical changes in welfare, wealth, or power” (Dafoe, 2018).
  • Analysis skews towards risks, but most expect AI to robustly improve welfare, health, wealth and sustainability; economists overwhelmingly believe AI will create benefits sufficient to make everyone better off. The challenge is sufficient safety and distribution.
  • Framing AI ethics and existential-risk-focused AI safety as in conflict is unhelpful and often misplaced — the scholarship and policy work overlap considerably.
  • Parallel to AGI safety, which moved from thought experiments about superintelligence to empirically informed research programmes: AI governance should emphasise scalable governance — solutions to pressing challenges that remain relevant to future extreme ones.

Lens 1: AI as a general purpose technology

  • A GPT (1) provides a valuable input to many processes and (2) enables important complementary innovations. Examples: printing, steam engines, rail, electricity, motor vehicles, aviation, computers — mostly energy, transport or information processing.
  • AI is plausibly the quintessential GPT: a fundamental input, highly complementary, still early in development with a capability ceiling likely above human level. Kevin Kelly: “Everything that we formerly electrified we will now cognitize.”
  • Political-economic transformations from AI are likely to arrive faster than from most previous GPTs.

Properties AI likely inherits:

  1. Growth: GPTs grow the economy, often radically, via efficiency gains and entirely new processes. Potentially Pareto-improving if deployed well and losers are compensated.
  2. Disruption and distribution: GPTs disrupt existing processes and the social-political relations depending on them, shifting power and wealth and imposing hard-to-insure displacement costs. Early versions look harmless; cumulative impact after generations of improvement can be revolutionary.
    • The favourable net effect of the past two centuries arguably depended on labour and liberal institutions being economic and military complements to the technological ecosystem — which may not continue.
  3. Anticipatory conflict: expected losers mobilise, from minor regulatory protections to revolutions; internationally, anticipated rise and fall of countries can precipitate aggression and war.
  4. Strategic and dual-use character: possession of mature variants is close to a necessary condition of great power status, making them sites of rivalry. Their inseparable dual-use nature makes arms control especially difficult.

Lens 2: AI as an information technology

  • An information technology improves the production, compression, transmission, reproduction, enhancement, storage, control or use of information. AI enhances each, and amplifies other information technologies.

Economic implications: increasing returns and distribution

  • Low marginal costs relative to fixed costs produce large economies of scale — for ML, training costs often exceed deployment costs by orders of magnitude.
  • Network economies add further returns to scale, concentrating market structure towards one or a few firms, and producing winner-take-most labour markets for superstar actors, researchers, designers, CEOs.
  • These push towards greater income inequality — but there is a countervailing pull towards consumption equality, since information is non-rival and hard to exclude:
    • Information is hard to hold on to: knowing something is possible accelerates competitors’ catch-up; services are copyable; IP theft is comparatively cheap.
    • Socially efficient pricing is near zero marginal cost, achieved via public-interest services (Wikipedia), IP limits (Project Gutenberg) and competition (free/ad-based services). The value of free services to a smartphone user is plausibly tens of thousands of dollars per year.
    • A billionaire’s books, films, games, navigation apps and social media are largely accessible to the median wage earner.

Coordination and identity

  • Political impacts are ambiguous — innovations may strengthen or undermine existing communities and power centres.
  • Scale economies and standardisation encourage broader collective identities, historically fuelling national identity, cosmopolitanism and liberalism.
  • Simultaneously, better coordination among spatially distributed individuals supports narrow distributed identities — global ideologies and movements, religions, cultural identities — sometimes undermining incumbent power centres. Analogously in the economy: boutique firms, the gig economy.

Power

  • Information shifts power within relationships: monitoring, monopolising critical information, information rents in bargaining and principal-agent settings. The US intelligence budget is ~10% of its military budget; in coups, “information is the greatest asset” (Luttwak).
  • Information technology is transforming privacy — plausibly weakening it against authorities while strengthening it against social peers.
  • Centralisation depends on the authority’s ability to monitor and communicate with agents: the telegraph and radio curtailed the autonomy of ambassadors and ship captains; remote and autonomous weapons let commanders execute orders without delegating to officers who might object. But trends are not uniformly centralising — cf. the printing press, RSA, PGP.
  • Information moves faster than other processes, so information-based dynamics accelerate crises — e.g. financial “flash crashes”.
  • Whether AI will strengthen, weaken, subsume or transform the state remains too early to say.

Lens 3: AI as an intelligence technology

  • Intelligence technology: an innovation in the ability of some entity to solve cognitive tasks. Least developed of the three lenses in the literature, but arguably illuminates the most important impacts.
  • Key distinction along a spectrum:
    • Tools — abacus, dictionary, notepad: narrow, non-autonomous, requiring a user.
    • Systems and agents — more general and autonomous, and often most impactful. Systems: the price mechanism, language, bureaucracy, peer review, the justice system. Agents: a Grand Vizier, a military general staff, a corporation, a deeply socialised bureaucracy.
  • High-level properties: often critical for military and economic survival; often transform the character of the largest political entities (human tribes, the Neolithic state, the medieval state, the modern state each arose partly from improvements in intelligence technologies).
  • Historically they both substitute for and complement human cognitive labour (machine calculators; competent bureaucracies). Whether future AI complements or substitutes has profound implications for labour share, inequality and growth rates, since capital can grow itself.

Bias, alignment, control

  • Even simple tools bias decision-making towards what fits the tool: early states biased towards legible social arrangements (Scott); policymakers towards GDP rather than wealth and wellbeing; social media companies towards engagement.
  • Systems and agents may be misaligned such that increasing their power systematically produces outcomes against the principal’s interests — markets under significant externalities; Enron and Andersen; civil-military friction (the Kennedy administration and the Joint Chiefs during the Cuban Missile Crisis; the Obama administration and troop requests for Afghanistan).
  • These are principal-agent problems where the agent’s advantage is not merely informational but potentially one of vastly superior intelligence: the principal may not know what to ask, where to look, or possess the concepts to make sense of the problem.
  • Social science has explored solutions — oversight, transparency, whistleblowers, representation, institutional design — yet AI alignment work is almost exclusively done by AI researchers.
  • The upshot: our social order depends on aligning and controlling human and organisational intelligence; augmenting social entities with machine intelligence makes those problems more complex and more critical.

Governance and anarchy

Institutional fit and externalities

  • Governance shapes behaviour towards social goals through institutions — norms, rituals, rules, organisations, regulations, regulatory bodies, legislatures (North, 1991).
  • Analytic tool: externalities. Institutions can raise welfare by discouraging negative-externality behaviour and encouraging positive-externality behaviour — “internalising” them.
  • Diagnostic questions for whether an institution fits a governance issue — does it have the needed:
    • spatial remit?
    • issue area remit?
    • political remit (adequate, legitimate representation of stakeholders)?
    • technical competence?
    • institutional competence?
    • influence (ability to shape material incentives)?

Worked case: self-driving vehicles

Interests: citizen safety, traveller mobility, producer business interests. Primary externality: the effect of driving algorithms on other road users. Traffic safety agencies have a presumption of fit, but the domain requires new safety-evaluation methods (fleet crash statistics, simulators), new best practices (privacy, driver attention, crash data sharing), insurance repricing, liability attribution between producers and users, and management of new risks (hacking, whole-fleet compromise) — plus standards-setting, smart infrastructure and regulatory harmonisation. Some issues (the trolley problem) attract attention out of proportion to their policy importance.

  • The greater the distance between an emerging governance area and an existing legitimate competent institution, the harder adaptation becomes. Example of failure: out-of-copyright books are not freely available in local libraries — not for want of a willing provider, but because copyright law was never updated.
  • Adaptation is hardest where issues fall within, or are framed as part of, unresolved social conflicts — there is no consensus on social goals, no trusted institutions, and the issue itself becomes a battleground.

Domestic conflicts

  • Social media moderation sits on the US left-right cleavage: censorship, bias, “fake news”, filter bubbles, foreign intelligence operations, hostile social movements. Society fundamentally disagrees about what effective, safe, legitimate moderation would mean.
  • Algorithmic fairness: ProPublica’s 2016 COMPAS investigation found false positive rates significantly lower for white than black defendants. Later research clarified that if any demographic difference in false positive rates, false negative rates or calibration counts as “bias,” then — given imperfect prediction and differing base rates — “bias is mathematically inevitable.” Every classifier is then biased, so a more refined understanding of fairness is required.
    • The case shows how a new capability forces precision about contested concepts, opening political debate where consensus, principles and institutions are lacking.
  • Systematically, as sensors record more behaviour and decisions move from the black box of the brain into manipulable, auditable algorithms, AI expands the terrain subject to political control, raising the stakes of political contestation — analogous to a resource discovery inflaming a disputed border.
  • We may come to regard the pre-AI era as one of significant autonomy for citizens, workers, students, police officers, managers and educators.
  • Where authorities and citizens sit at an uneasy status quo, expanded scope for political control tends to shift the balance towards the state — the “digital authoritarianism” literature. More work is needed on a positive agenda for AI strengthening liberal institutions.

Great power security competition

  • International anarchy: no higher authority to make and enforce law, so the final recourse is force; this generates the security dilemma. Costs include $2 trillion a year in military spending and thousands of nuclear warheads, ~2000 on high alert; inadequate provision of other global public goods plausibly follows too.
  • There are also significant risks from remedies to anarchy, namely excessive political centralisation and global totalitarianism.
  • Two clusters of issues:
    1. Military applications — lethal autonomous weapons, cyber operations, foreign influence operations. Norms can shape behaviour (cf. the nuclear taboo since 1945), but security competition bears down: even US DoD proponents of humans-in-the-loop concede the control cannot be achieved when instant response is imperative. LaPointe and Levin: “Military superpowers in the next century will have superior autonomous capabilities, or they will not be superpowers.”
    2. Decoupling of supply chains, commerce and research between China and the West — China’s constraints on Western tech companies and effective ban on most Western AI services; US chip supply-chain independence efforts (~$50bn CHIPS Act) and blocking ASML EUV exports; scrutiny of ML research collaborations; China’s exclusion from the Global Partnership on AI.

The AI race

  • The arms race metaphor communicates (1) rapidly increasing investment, (2) driven by perceived very large geopolitical stakes, (3) with military relevance. But most geopolitical activity concerns supply chains, infrastructure, industrial base, strategic industries, scientific capability and prestige — so strategic technology competition or “the AI race” is more accurate.
  • The race to the precipice: perceived gains from relative advantage induce corner-cutting and excessive risk. Commercial precedents include Boeing’s 737 MAX and Uber’s self-driving catch-up efforts.
  • A geopolitical race to the bottom could produce “AI Manhattan Projects”: hurried crash programmes, deployment of powerful but unreliable systems in cyber and kinetic conflict.
    • Cyber-war favours speed, making humans-in-the-loop untenable; ML systems behave unexpectedly in complex adversarial settings; incentives to deploy at scale mean extreme behaviour has broad impacts. One current aspiration is insulating nuclear command and control from ML-powered cyber operations.

How risky is such a race?

  • Model: the war of attrition (equivalently an all-pay auction where the “revenue” is risk of conflict), the canonical model of nuclear brinkmanship.
  • With rational players and common knowledge, each generates expected “revenue” of 1/3 of the prize’s value; a related rent-seeking literature finds 50% of rents dissipated in two-player contests.
    • Concretely: if the prize is perceived as existential — as valuable relative to the status quo as the status quo is to nuclear war — a decision maker should be willing to race up to a 33% chance of nuclear war.
    • Bad news: sufficiently attractive prizes make rational actors willing to expose the world to significant devastation risk. Good news: they don’t race all the way to the bottom.
  • Real-world races may be worse:
    • No common knowledge — decision makers may encounter the game for the first time (the economists’ joke: to raise $100 quickly, run an all-pay auction for $10).
    • Novel tail risks are hard to forecast.
    • Irrationality: overconfidence, intrinsic value placed on winning, honour or regime survival making backing down costlier than the material stakes.
    • Unobservable risk-taking: without a public signal of risky behaviour, norms of mutual restraint are harder to build.
    • Theory-dependent safety: via a winner’s curse dynamic and rationalisation, each party may systematically perceive its own behaviour as safer than others’, producing an escalation spiral.

Escaping race dynamics

  • Unilateral de-escalation: interventions that reduce underestimation of risk, or make it easier for leaders to opt out, should reduce risk — though as in models of coercion, they may embolden the other party. Force-based unilateral solutions can be as risky as the race.
  • Cooperative mutual restraint: norms, treaties and institutions that change the rules so racers internalise risk. Requires agreement on what is unacceptably risky, means to observe compliance, and incentives to induce it.
    • Core obstacle: the transparency-security tradeoff — enough transparency to reassure, without compromising the monitored party’s security (Coe & Vaynman, 2020).
  • Third parties and global institutions: articulate focal norms and safety standards; redirect prestige motivations towards prosocial endeavours (cf. the Space Race); verify and rule on non-compliance (as WTO and IAEA do); overcome disclosure dilemmas. Further off: hard powers such as sanctions or direct control of materials and activities (as proposed for the Atomic Development Authority).

Why AI is harder to control than nuclear weapons (Zaidi & Dafoe, 2021)

  • Dual use: dangerous applications are harder to separate from beneficial ones.
  • Value: economic and scientific value from general AI advances greatly exceeds that from nuclear technology.
  • Diffusion: AI assets and innovation are far more globally diffused.
  • Discernibility of risks: nuclear risks are easier to understand.
  • Verification and control: nuclear tests and ICBM deployments are unilaterally verifiable and chokepoints (fuel, centrifuges) controllable; cyber-weapon deployments are not, and compute remains deeply dual-use.
  • Strategic gradient: nuclear value plateaued at secure second-strike (China held under 300 warheads for six decades); AI may have a persistently steep gradient, incentivising racing and increasing volatility in power.

Value erosion

  • Beyond catastrophic accidents, competition and advanced AI can mix badly in gradual ways. Generalising the safety-performance tradeoff: any tradeoff between something of value and performance in a high-stakes contest can push decision makers to sacrifice that value.
  • Zuckerberg’s prepared congressional talking point exemplifies the logic: breaking up Facebook would undermine a “key asset for America” and “strengthen Chinese companies.”
  • Long run: competitive dynamics could proliferate systems — organisational types, countries, autonomous AIs — that lock in undesirable values. Related treatments: Bostrom’s The Future of Human Evolution (2004), Christiano’s “greedy patterns” (2019), Hanson’s Age of Em (2016).
  • Objection: history shows long-term trends favouring humanity, driven by the agency of key decision makers.
    • Reply: human circumstances did not obviously improve after previous technological revolutions such as the Neolithic. Post-industrial gains in wellbeing and liberty may be attributable to human labour becoming a complement to industrial machinery, and to national power depending on an educated, free, supportive citizenry. AI could reverse this two-hundred-year trend if it substitutes more than it complements labour and reduces authorities’ need for citizen support.
  • Value erosion’s gradual operation may make it easier to observe and coordinate against — but may also mean attention is not mobilised in time.

Conclusion

  • In the late 1920s some military analysts believed unstoppable bombers dropping poison gas would fundamentally alter warfare. They were wrong about timing and mechanism, but correctly foresaw the strategic logic of the nuclear era, realised instead by the neutron chain reaction.
  • Confident predictions of nuclear apocalypse were factually mistaken — but the early arms controllers’ framing (“One World or None”) may have had the strategic logic right: increasingly powerful technology and great power competition are ultimately not compatible with the long-run flourishing of humanity.
  • AI is not a narrow technology; the AI revolution will be more like the industrial revolution, transforming economics, politics and society. To succeed, the field of AI governance must be comparably expansive and ambitious.