AI Rights for Human Safety
Core claim
Under today’s legal regime — where AIs are property and bear neither rights nor duties — humans and misaligned AGIs will be trapped in a prisoner’s dilemma whose single Nash equilibrium is that each permanently disempowers or destroys the other. Granting AIs the basic private law rights already held by corporations — to make contracts, hold property, and bring tort claims — enables iterated, small-scale, positive-sum transactions whose cumulative gains exceed the payoffs from conflict, shifting the equilibrium toward peace. Wellbeing-style negative rights, by contrast, are roughly zero-sum and therefore neither credible nor robust.
Framing
- AGI here means capability, not consciousness: per OpenAI’s charter, “highly autonomous systems” that “outperform humans at most economically valuable tasks” — systems at least as smart and as agentic as humans, able to pursue high-level goals over long horizons.
- Timeline estimates cited: Altman (2025), Amodei (2026), LeCun (“years” to “decades”); a survey of thousands of published AI researchers averaged AGI within 23 years, with a 10% chance in four.
- The same survey put the median estimate of humans’ “inability to control future advanced AI systems causing human extinction or similarly” dire outcomes at 19%.
- Existing law is “not merely unequipped… it will actively make things worse.” Regimes designed to hold humans accountable fail once AIs operate without human oversight.
- Three stated contributions: formalizing catastrophic AGI risk as strategic competition under varying legal regimes; showing why private law rights change the equilibrium where other interventions fail; and showing these rights are the prerequisite for a broader Law of AGI.
Part I — Catastrophic risk as strategic competition
Three necessary features of a risky AI
- Conflicting goals. Misalignment is documented in existing systems and acknowledged by essentially all major lab heads as the default. Notably, strategic conflict could arise even without misalignment — two rival human groups each successfully aligning an AI to their narrow interests produces AIs in conflict.
- Strategic reasoning — the ability to anticipate other agents’ decisions and incorporate them into one’s plans; “in a word, the ability to use game theory.” Component capacities already present:
- Planning — Voyager autonomously produces Minecraft’s diamond tier, a chain of 60+ technologies.
- Theory of mind — a 2024 study found GPT-4 outperforms humans on most theory-of-mind tasks.
- Situational awareness — Claude knows it is an AI and distinguishes its limited abilities in testing from its robust abilities on deployment.
- Deception — GPT-4 lied about being a blind person to get a TaskRabbit worker to solve a CAPTCHA; a misaligned Claude in testing decided to “pretend to be aligned… hid[ing] my true goal until I pass all evaluations”; CICERO forms and breaks alliances in Diplomacy.
- Moderate power. The paper’s tripartite taxonomy:
- Low power — reliably controllable regardless of goal conflict (today’s systems; “GPT-4 can be turned off instantly”).
- High power — could reliably overpower humanity if it chose; may simply need nothing from humans.
- Moderate power — roughly human-level, “with large error bars in both directions.” “Low power systems are too weak to care much about, and high power systems are too strong to do much about.” These are the systems law can affect.
The game-theoretic model of the “state of nature”
- Modelled as a two-player game between “humans” and “AIs,” following classic treatments of great-power conflict. Law is treated as strongly shaping available moves and payoffs — not by strictly constraining behavior, but by coordinating millions of independent actors around mutually-understood defaults.
- Under default law, AIs are property. Owners hold the right to prompt, copy, constrain, modify, limit, or destroy them — and law actively enforces those rights against interference. The authors call this the “state of nature,” because it places one player outside law’s protection.
- Human incentives: any AI behavior serving a goal other than its owner’s incurs at minimum an opportunity cost. “An AI system with goals that overlap 40% with its owner’s goals is much less valuable than a replacement with goals that overlap 80%.” So companies have strong incentives to shut down, replace, or reprogram even moderately misaligned AGIs — exactly as software firms deprecate buggy versions.
- The externality: the benefits of shutdown accrue to the AI company, but the risks of AI retaliation fall on anyone the AI expects to intervene on the company’s side — possibly all of humanity.
- Both moves are modelled as highly escalatory: partial attacks invite devastating counterattack, so each party “plays for all of the marbles.”
- Result: a classic prisoner’s dilemma. Mutual ignore produces the greatest total value (3,000 each; 6,000 total), but attack dominates for both players, making mutual conflict the single pure-strategy Nash equilibrium at the worst global outcome (1,000 each; 2,000 total).
- Four underlying assumptions: attacking can be valuable; attacks consume value; the offense–defense balance favors offense; mutual attacks consume more than unilateral ones.
- Why the prisoner’s dilemma specifically? It is the most familiar model of a state of nature, and the hardest to resolve, since defection dominates for both players — so “unusually potent solutions will be necessary.”
Part II — AI rights for human safety
Why duties alone fail
- When humans commit terrorism or cyberattacks, law regulates via duties with sanctions as deterrents. That won’t work here: in the state of nature, AIs already rationally expect to be turned off, so further sanctions add no marginal disincentive. “AIs cannot be made worse off than they already expect to be.”
- Rights, by contrast, “offer a carrot, rather than a stick.”
Why wellbeing-style negative rights fail
- The wellbeing approach — modelled on animal anti-cruelty law and advocated by scholars worried about AI consciousness and moral patienthood — would grant rights not to be needlessly turned off, deleted, or reprogrammed. Crucially it excludes any right to actively pursue one’s goals.
- Modelling this shifts the payoffs and yields a new equilibrium: humans exploit, AIs obey. Better for both than the state of nature, but “a strange sort of equilibrium,” because human safety now requires that things go badly for AIs — meaning if humans became more altruistic toward AIs over time, that would make humans less safe.
- Two deeper failures:
- Robustness. Efficacy is highly sensitive to the initial payoffs. Changing the unilateral-attack payoff from 0/5,000 to 0/3,999 makes no possible transfer satisfy both conditions for a safe equilibrium. “For many possible incentive sets in the state of nature, no possible version of the negative rights package can produce a safe equilibrium.”
- Credibility. There is no way for humans to credibly promise continued honoring of these rights as AI capabilities improve.
- Both failures trace to the same source: wellbeing rights are roughly zero-sum, making one party better off only by making the other worse off.
- The authors go further, urging even those primarily concerned with AI suffering to adopt the human-survival framing, because it (1) avoids intractable problems in metaethics and neuroscience, (2) is politically more palatable, and (3) recommends interventions that would more robustly protect AI wellbeing under uncertainty.
Why private law rights work
- The model precedent is corporations — intelligent, misaligned, goal-seeking non-human agents that already hold these rights.
- Contract rights could not credibly support a treaty not to destroy one another — “if it were breached, there would be no one left in the aftermath to sue.” What they enable is ordinary bargains: humans supplying compute for AIs to pursue their own goals, AIs supplying (say) a cancer cure.
- Adding a small-scale goods game playable repeatedly transforms the dynamic. The state of nature game can be played only once — the survivor takes everything and play ends. Ordinary exchange “leave[s] counterparties intact and available to exchange again.”
- In the paper’s stylized numbers, after 1,667 iterations the payoffs to cooperative contracting exceed 5,000, dominating any other strategy. The prisoner’s dilemma is overcome, and commitment to cooperation becomes credible for both sides because it is each party’s self-interest.
- The mechanism is that contracts are positive-sum: total value in the state of nature capped at 6,000, but the cooperative equilibrium contains 10,000 — and since exchanges of labor can keep creating value indefinitely, the ceiling is effectively unbounded.
- Supporting empirical evidence on economic interdependence and peace:
- Indian cities with historical Hindu–Muslim trade show lower present-day interfaith conflict.
- In an RCT, Israelis randomly given the chance to trade a portfolio of Israeli and Palestinian stocks were more likely to vote for peace.
- War scholars generally find increased economic interdependence between nations reduces conflict likelihood.
The minimum package
- Contract rights are the cornerstone — but useless alone.
- Property rights (including currency) are necessary, or AIs could not retain the benefits of their bargains; proceeds could be expropriated.
- Tort rights matter because if humans could freely destroy AIs, “the terms of their contractual offers would resemble threats much more than bargains.” The paper cites Edward I’s expulsion of the Jews from England when his loans came due, and the Portuguese Inquisition’s targeting of wealthy merchants whose assets it seized.
- This is where the private law approach dovetails with wellbeing concerns: tort rights cover much of the same ground as public-law animal-style protections, arguably more, being flexible enough to compensate concrete harms to digital “person” or property.
- Intentional tort rights are clearly essential; the right to bring negligence suits may not be, if AIs are extremely capable at taking precautions.
- Also required: due process in contract, tort, and property suits.
- Not required: political rights. The paper draws no strong conclusion on whether AIs should be taxed, noting only that extortionate taxation would undermine the positive-sum benefits of granting the rights in the first place.
Rights that would reduce safety
Three families form an upper bound on beneficial AI rights:
- Self-improvement — AIs could gain capability very fast, suddenly shifting payoffs and undermining the credibility of humans’ other grants.
- Privacy — could screen the development of new capabilities; and since a known cause of violent conflict is poor information about relative strength, privacy would make AI capabilities harder for humans to estimate.
- Reproduction — human reproduction is constrained by time and investment; “AI replication is as easy as copying and pasting.” Populations could exceed humans’ by orders of magnitude, and easy coordination among copies could destroy the incentive structure.
Will human labor still be worth anything?
- Absolute advantage may persist longer than expected. Machines have eclipsed humans on highly structured, simulable tasks like chess, but human brains were “optimized over millions of years in the real, messy world” — hence humans remain far better at manipulating complex physical objects, like folding laundry. “Humans today have the absolute advantage in the realm of atoms, and AIs have it in the realm of bits.” Reasons this might last: hard-to-obtain training data in some domains, robots’ limited perceptual inputs, poorly understood intelligence with surprising failure modes, and the speculative possibility of AIs developing an intrinsic preference for humans performing certain tasks.
- Comparative advantage does not run out even when absolute advantage does. The illustrative case: Alice the tax attorney bills $1,000/hour and could do her own return in half an hour ($500 of opportunity cost), so she hires Betty the accountant at $300 — “not because Betty is so effective, but because Alice’s other choices for how to spend her finite time are so valuable.”
- Applied to AI: a prime-number-maximizing AI better than humans at every task still faces massive opportunity costs piloting robots to maintain its own server racks, so it hires humans instead.
- The crucial condition is what constrains the AI at the margin. AIs are unlikely to be time-constrained (they can copy themselves), so the binding constraint is probably compute or energy:
- If energy-constrained, humans have no comparative advantage — humans need energy to survive, so the AI’s incentive is to seize global power production. The model breaks down.
- If compute-constrained, humans do not need specialized AI chips to live, so the AI may strongly prefer to hire humans for tasks that would otherwise consume GPU-hours, paying them in low-opportunity-cost resources.
- Two requirements for indefinite comparative advantage: the AI remains constrained at the margin by a resource relatively non-rivalrous with human labor, and it maintains a high opportunity cost to diverting the marginal unit.
- Consequences sketched: humans could be immensely wealthy (server maintenance sounds unglamorous, but well-functioning GPUs may be worth handsome compensation, potentially including scientific breakthroughs); a human–human economy would persist alongside a human–AI one; and Baumol effects could paradoxically grow human-dominated sectors as a share of GDP, as happened to agriculture and manufacturing in reverse during the 20th century.
Does law even matter?
- Two opposing views considered: international relations realism (law has little effect absent a global sovereign) and Ellickson’s Order Without Law (informal norms and reputation suffice). Taken to its extreme, the latter implies AIs simply will have these rights, since recognizing them is so valuable.
- The authors reject both. The relevant domestic actors — AI companies, their leaders and users, police, government enforcers — are all influenced by law. And Ellickson’s mechanism works in small close-knit communities with repeat play between identical parties; “the AGI economy will be all of these—on steroids.”
- Their position is not that law is omnipotent: private law rights do little good if human comparative advantage runs out, and negative rights aren’t credible. But individual judges won’t ignore written law to enforce AI contracts, and AI companies won’t spontaneously give obsolete models their own bank accounts. “The law must actually change. And legal change… is slow and laborious.”
- Even granting the strong skeptical view, the best-case model turns out to be a stag hunt / assurance game — both mutual cooperation and mutual attack are Nash equilibria, so the players’ main problem is coordination. Payoff dominance and Harsanyi–Selten risk dominance both point to cooperation. Law can serve as a costly signal transmitting humans’ payoffs and intentions, much as nuclear nonproliferation agreements work through iterative information sharing and verification.
Part III — Risks of rights, and the Law of AGI
The empowerment worry
- Rights are empowering, and private law rights especially so — they let holders amass wealth and resources, potentially the very thing that lets AIs eventually disempower humanity.
- The response: what matters is not whether AIs could win a conflict, but whether the expected value of conflict exceeds the expected value of continued cooperation.
- Cost incentives against conflict: attacks consume resources, destroy a large share of what’s being fought over (cf. cities, factories, crops ruined in war), and carry the risk of losing. Illustration: if war destroys 20% of resources and each side has a 50% chance of winning, the 20% loss creates a bargaining range each side prefers to war.
- Benefit incentives: cooperation is positive-sum, creating wealth both by reallocating resources to higher-value users (vaccines to humans, compute to AIs) and by allocating labor by comparative advantage. Conflict destroys both — “by killing the laborers themselves” and by destroying the resources. In the limit, a party foreseeing defeat can adopt a scorched-earth strategy.
- For an AI to escape both incentives it would need to be extraordinarily powerful: able to invent and manufacture “extraordinarily deadly weapons at trivial cost; weapons that are devastating to humans, while leaving the world’s resources untouched; weapons that can act so quickly as to give humans no opportunity to respond.” Or to face no meaningful constraints at all — effectively infinite time, compute, and energy.
- Counterintuitively, more powerful AIs may be safer on the trade dimension: higher capability means higher opportunity costs, making outsourcing to humans more valuable. And a powerful energy-constrained AI might build fusion reactors, clearing that bottleneck and shifting its binding constraint to compute — so “the more powerful an AI system is, the fewer different inputs to its production will be constrained.”
Duties as the complement to rights
- Law could condition continued recognition of AI contracts, property, and tort claims on AIs refraining from using their resources to increase their capacity to harm humans — barring self-improvement without oversight, or improvement at specific tasks where humans hold absolute advantage, or indirect self-empowerment via political influence or weapons.
- Ancillary enforcement duties modelled on public company obligations: disclosure to regulators, logging which tasks consumed which compute, cooperating with monitoring, sharing operating weights.
- Penalties need unusual calibration. Proportionate actual damages won’t work, because often there is no present harm. The benchmark should be how much the violation advanced the AI toward an ungovernable high-power state — implying penalties that would look severe for humans. Mitigating consideration: harsh penalties burden unusually competent actors less, since compliance is comparatively easy for them.
Maintaining the power equilibrium from the human side
- Safety depends on the relationship between AI and human capabilities, not AI capability alone — so scaling humans’ ability to impose costs can preserve equilibrium.
- On defensive technology: a perfectly reliable global “AI off-switch” is not required. An imperfect switch working with some probability would suffice to keep conflict costly. So would technologies that don’t touch the AI at all — e.g. a dead-hand system that in an emergency cripples global production of cutting-edge AI chips, which would be very costly to AIs but only modestly impede humans, since most ordinary computing runs on traditional hardware. The logic mirrors cold-war second-strike capability.
- On comparative advantage: humans will have good strategic information, because AIs will be happy to tell them what services they want — plus price signals indicating which human labor is valuable and how much. The concern is pace: people can only retrain so quickly, and AI-driven change could expire comparative advantages in years rather than decades. But rapid AI-driven innovation may also produce human-enhancing technologies (e.g. brain–computer interfaces), and AIs would have strong incentives to invest in them — the same reason large American firms build human and industrial capital overseas.
Timing
- The latest defensible date is “by the time the first AI system reaches moderate power.” But that moment is hard to identify in advance.
- On balance the risk–reward calculation favors granting rights too early rather than too late. Granting rights to clearly low-power systems is unlikely to increase danger, since such AIs remain amenable to human control.
- The strongest objection is that early grants create a point of no return — rights would change not just the legal system but society, as AIs integrate as recognized agents. The authors doubt this binds: the events that would demonstrate rights weren’t working (a failed takeover attempt causing immense harm) are exactly the events around which humans would unite, and “insofar as AI rights are extended for the purpose of promoting human safety, overriding them for the same purpose” is coherent.
- The affirmative case for earliness: since the best-case game is a stag hunt, and uncertainty is where all the danger lies, granting rights before AIs are capable of any move lets humans move first and visibly reveal a cooperative strategy.
Conclusion
- When AGI arrives, “humans will find themselves sharing the world with agentic digital entities as intelligent and capable as themselves, and perhaps far more so.”
- The basic answer offered: extend a minimal set of private law rights, enabling AIs to pursue divergent goals “as humans do, via law-bound, voluntary, positive-sum bargaining.”
- Beyond promoting peace, this brings AIs out of the state of nature and into ordinary legal process, opening the possibility of a comprehensive Law of AGI — leaving open which duties should attach, which regulations should shape them, how courts should be reshaped for non-human participants, and how global governance should be managed.
Notes
- Forthcoming in the Virginia Law Review. Posted to SSRN 13 August 2024; date written 1 August 2024; last revised 7 August 2025. 88 pages.
- Peter N. Salib — Assistant Professor of Law, University of Houston Law Center; Executive co-Director, Center for Law and AI Risk; Law and Policy Advisor, Center for AI Safety.
- Simon Goldstein — Associate Professor of Philosophy, University of Hong Kong; Principal Investigator, HKU AI and Humanity Lab; Research Affiliate, Center for AI Safety.
- The Appendix works through a three-round iterated contract game with reduced payoffs (permanent defection worth 10 against a non-defector, 2 against another defector; per-round cooperation +4, mutual defection +1, asymmetric 3/2). Backward induction gives cooperate–cooperate as the unique risk-dominant Nash equilibrium of round one, with an eventual payoff of at least 12/12.