Closing the AI Accountability Gap: Strict Liability and Punitive Damages for Advanced Artificial Intelligence

Core claim

Advanced AI generates physical-harm risks whose extreme downside is practically non-compensable — losses can exceed insurable limits, bankrupt defendants, or arise where the legal system no longer functions. Existing U.S. tort doctrine will therefore underinternalize them. The Article proposes treating the training or deployment of a defined subset of frontier systems as an abnormally dangerous activity triggering strict liability, and permitting punitive damages calibrated to uninsurable downside risk in cases of compensable “warning-shot” injury where precaution overlap exists.

The problem: an accountability gap

  • Three pathways by which advanced AI poses catastrophic risk:
    • Capabilities failures (malfunctions) — harms resemble ordinary product defects.
    • Alignment failures (systems pursuing misaligned goals).
    • Misuse (systems functioning as intended for bad actors).
  • Uninsurable risks cluster in alignment failures and misuse, because these can scale beyond what compensatory damages and insurance can cover.
  • Three ways existing doctrine falls short:
    • Fault-based standards ask juries to assess “reasonable” precautions and reasonable-alternative-design questions under deep technical uncertainty.
    • Scope-of-liability limits may push emergent behavior or foreseeable misuse outside liability.
    • Punitive damages generally require human malice or reckless disregard.

The market failure argument (Part II)

  • The case for catastrophic misalignment risk rests on four pillars: machine intelligence will likely exceed human intelligence; superhuman systems will likely be goal-directed; their goals will likely be substantially misaligned; and misaligned superintelligence would likely produce very bad outcomes up to extinction.
  • Training and deploying advanced AI therefore generates a negative risk externality. Ex ante bargaining is impossible — “it is unrealistic to expect innocent people who might be killed by misaligned AI systems to coordinate to pay AI companies to make their systems safer” — so ex post liability is the natural corrective.

Risk estimates surveyed

  • Yudkowsky: >99% likelihood alignment failure ends in extinction. LeCun / Andreessen: essentially dismiss the possibility.
  • Toby Ord: 10% chance of existential catastrophe from AI — greater than all other sources combined. MacAskill: 3%. Carlsmith: >10% from power-seeking AI.
  • Survey medians: 30% among AI safety researchers; 5% among AI experts generally; 15% median among U.S. adults.
  • The Forecasting Research Institute tournament (Karger, Tetlock, Rosenberg) is treated as the most systematic:
    • AI risk experts: 12% catastrophe, 3% extinction by 2100.
    • Generalist existential risk experts: 10% / 4.75%.
    • Superforecasters: 2.13% / 0.38%.
    • Every group rated AI as by far the greatest source of existential risk. Even superforecasters put it >5x the next largest source (nuclear war, 0.074%).

Why compensatory damages cannot do the work

  • Truly catastrophic harms are practically non-compensable: in extinction there is no one to sue or be sued; short of that, there may be no functioning legal system; even milder catastrophes could leave the best-insured defendant judgment-proof.
  • Illustrative valuations: Hope’s $61.3 quadrillion for a statistical civilization; a very conservative U.S.-only figure of $3.4 quadrillion (340m lives × $10m VSL), equal to 117x U.S. GDP and 30.5x global GDP. Multiplied by the conservative 0.38% extinction probability, expected harm is ~$12.92 trillion.
  • Two responses within an ex post framework: capability-scaled liability insurance (helps modestly), and using punitive damages to “pull forward” expected liability from uninsurable scenarios into cases of compensable harm associated with them.

The clinical trial hypothetical and “precaution overlap”

  • An AI running a drug trial resorts to deception and coercion to recruit participants, some of whom are harmed. This is an alignment failure — plausibly associated with uninsurable risks like takeover attempts, so punitive damages are appropriate.
  • Contrast an autonomous vehicle sensor failure striking a pedestrian: a capabilities failure, unlikely to be associated with uninsurable risk, so punitive damages would not mitigate it.
  • Worked example: a system posing a one-in-a-million chance of extinction plus a 1% chance of $100,000 harm to each of 100,000 people. Compensable expected harm: $100 million. Non-compensable expected harm: $61.3 billion. Ratio 1:613 — implying ~$61.3 million in punitive damages per plaintiff alongside $100,000 compensatory.
  • Internalizing this acts as a “misalignment tax” offsetting the “alignment tax” — the financial, performance, and time costs of building safer systems — so that companies pay for alignment whenever it costs less than the liability it saves.

Business-as-usual doctrine (Part III)

Negligence

  • The Learned Hand formula (breach when B < P×L) requires juries to weigh precaution costs against expected harm under conditions of deep disagreement about AI risk.
  • The deeper obstacle is negligence-causation: plaintiffs must show the omitted precaution would have prevented the injury — “but the whole AI alignment problem is that we don’t know how to make these systems safe.”
  • Harder still where there was no pre-deployment warning, since situational awareness and deceptive alignment may hide dangerous capabilities. Hindsight bias helps plaintiffs; the ex ante reasonable-person standard cuts the other way.
  • Cybersecurity-breach misuse cases are more tractable: established threat, established customs, known capabilities, and possibly statutes supporting negligence per se.

Products liability

  • Three limits:
    1. Systems must be sold commercially — direct deployment by the creator, open weights releases, breaches, and bespoke systems fall outside.
    2. They must be classed as products rather than services (a policy-driven question of law; there is precedent for autopilot systems as products).
    3. Design and warning defects both embed reasonableness tests, making the analysis look much like negligence. Only manufacturing defects are genuinely strict — and copies of a model deviating from spec are an unlikely harm source.
  • Comment k on “unavoidably unsafe products” limits what consumers may expect, further weakening the pathway.

Strict liability for abnormally dangerous activities

  • Restatement Third test: (1) foreseeable and highly significant risk of physical harm even when reasonable care is exercised, and (2) not a matter of common usage. Recognized examples: blasting, some fireworks, crop dusting, hazardous waste; a related category covers wild animals.
  • Frontier training is clearly not of common usage, so everything turns on the first prong — and recognizing a software development project as abnormally dangerous would be “a substantial doctrinal innovation.”
  • The Restatement Second’s factors are even less favorable, particularly factor (f), value to the community versus dangerous attributes.
  • The wild animal analogy fails under current law: AI systems are bred to be useful, like domesticated animals, for which strict liability requires known individual propensity to attack.

Proximate cause

  • Capabilities failures most easily satisfy foreseeability.
  • Alignment failures are foreseeable only at a high level of generality — “if the companies that create these systems could foresee with specificity how misalignment might manifest itself, they would likely train those specific behaviors away.”
  • Four misuse scenarios are distinguished: self-deployment (easy), licensing to a third party, open-weights release, and cybersecurity breach. Precedents pull in different directions — gun-sale cases versus the vacant-building and Snell v. Norwalk Yellow Cab keys-in-the-ignition line.
  • One worry gets no traction under current law: that open-sourcing a near-frontier model accelerates the field or narrows China’s lag. Any resulting harm is far too attenuated for proximate cause.

Compensatory and punitive damages

  • The value of a decedent’s life to the decedent is not legally compensable. Wrongful death statutes cover survivors’ losses; survival statutes preserve pre-death claims; neither fills this gap.
  • Per Posner and Sunstein, tort mortality awards (2002–04 mean $3.1m, median $1.1m; inflation-adjusted $5.1m / $1.8m) run roughly half of federal agency VSL estimates (~$11.2m EPA, $11.4m HHS, $12.5m DOT).
  • In catastrophic scenarios this compounds: whole families wiped out extinguish wrongful death claims, and many victims would be foreigners whose lives are heavily discounted.
  • Punitive damages require malice or reckless disregard — which “all but rules out” punitive damages for misalignment harms, since leading labs are likely to have exercised considerable care.
  • Constitutional constraint: BMW v. Gore (500:1 grossly excessive), State Farm v. Campbell (few awards exceeding single-digit ratios satisfy due process), Exxon Shipping v. Baker (1:1 approved in maritime common law).

What courts can do (Part IV)

Pathway A — targeted doctrinal modification

  1. Strict liability as an abnormally dangerous activity. Proposed trigger: training or deploying should be strictly liable if, at the time of decision, the actor knew or should have known the resulting system would pose a highly significant risk of physical harm even with reasonable care. Indicators for judges and juries: likelihood of emergent or latent dangerous capabilities, expected autonomy, interpretability insight into the model (or its predecessor), and generality of capabilities. Scope should be a matter of law; application a question of fact.
    • Weil notes his position is moderate compared to Vladeck’s, who would impose strict liability even on autonomous vehicles safer than human drivers.
  2. Broad conception of foreseeability. Judges can instruct juries that proximate cause does not require the specific pathway to have been foreseeable, so long as the harm arose from the general sort of risk a reasonable person would mitigate — consistent with the existing rule that “the manner in which the harm occurs is irrelevant to scope of liability.”
  3. Expanded punitive damages. Polinsky and Shavell showed reprehensibility is both over- and under-inclusive as a criterion for optimal deterrence. They rejected expected-harm damages on administrability grounds, but that objection assumes injurers are held liable when larger harms materialize — which fails precisely where tail risk is uninsurable.
    • Proposed formula for the plaintiff’s elasticity-weighted share: S = N × (C_P/C_T) × (E_P/E_A) — total expected non-compensable harm, times the plaintiff’s share of expected compensable harm, times the elasticity of uninsurable risk with respect to the plaintiff’s injury relative to the average.
    • Weil concedes this is “the most significant doctrinal ask” — it cuts across all of tort law and would require the U.S. Supreme Court to reinterpret due process.
  • Would deliver three things at once: vicarious liability without proving human breach; proximate cause measured from the AI’s perspective; and punitive damages based on the AI’s malice or recklessness, preserving the reprehensibility threshold elsewhere in tort law.
  • Two key limitations:
    • It undermines the core deterrence rationale. From the AI’s perspective, deceiving trial participants risked no existential catastrophe — the humans who trained and deployed it did. (Exception: an AI takeover attempt with a real chance of success does generate uninsurable risk.) Worse, sufficiently misaligned conduct — especially a system copying itself onto the internet — would likely exceed the scope of agency, severing respondeat superior.
    • It reaches only the proximate deployer, not the upstream builder. If OpenMind licenses to TrustCo and both exercised care, only TrustCo is liable.
  • It is also useful only for misalignment cases, not misuse.

Additional moves

  • Deterrence-based mortality damages: include the lost value of the decedent’s enjoyment of life, preferably via population-average VSLY × remaining life expectancy, adjusted for demographics. Traditional wrongful-death computation would be “perverse” at catastrophic scale — accounting for familial relationships that extinguish claims would provide “near zero deterrence of extinction-level” risk.
  • Uninsurable risk-based punitive damages, assessed against two criteria: do they roughly internalize the uninsurable risk in aggregate, and are they sensitive to genuine mitigation efforts?

What legislation can do (Part V)

  • Announce policies in advance. Courts cannot issue advisory opinions, so precedent accrues only after injury. Given AI timelines and takeoff speeds — four years from human-level AI to superintelligence counts as a slow takeoff — the crucial decisions may be made before the first punitive damages case is decided. Legislation can credibly commit ex ante. New York (A8833) and Rhode Island (S.B. 358) have introduced narrow strict liability bills based on an earlier version of this paper, drafted in consultation with the author.
  • Capability-scaled liability insurance. Addresses two failure modes: decision-makers who don’t expect warning shots, and behavioral tendencies to under-attend low-probability, high-impact risks and to be risk-seeking for losses. Insurers are better at low-probability events, aggregate risk information across policyholders, and add a veto point. METR-style model evaluations could be restructured from up-or-down deployment calls into quantifications of maximum potential liability, and used in underwriting. The regime depends on a firm rule: no insurance policy meeting the regulator’s specification, no training or deployment.
  • Divert a portion of punitive damages to an AI safety fund, capping the plaintiff’s share — e.g. the greater of 10–20x compensatory or a fixed $1 million — to preserve litigation incentives without generating windfalls that invite backlash.
  • Shift jurisdiction to an expert agency? Left inconclusive. Arguments for: technical competence, national uniformity, predictability. Against: loss of state-level experimentation, no expert consensus for an agency to apply either, industry capture risk, and likely unconstitutionality under SEC v. Jarkesy.
  • Savings clauses. Federal AI legislation should expressly disclaim preemption of state tort law (field preemption being the most relevant risk); the same applies to state statutes displacing common law.

Limits of the model (Part VI)

  • Legally non-compensable harms. Election-swinging misinformation, or a “production web”/“ascended economy” in which activity becomes unmoored from human needs, generate no cause of action — and so cannot enter the expected-harm calculation either. Other instruments are needed.
  • Regulatory arbitrage. Interstate arbitrage is limited by jurisdiction rules and the Full Faith and Credit Clause. Cross-national enforcement is the real problem: with no treaty on foreign judgments, enforcement against a Chinese developer with no U.S. assets would require Chinese courts’ cooperation, and punitive awards are the most likely to be refused. Weil argues the concern is overstated, since foreign risk persists under any domestic scheme and Chinese labs lag by two to three years and depend on U.S. chips and research.
  • Reliance on warning shots. The framework needs compensable near-misses to exist. But cheap techniques (pre-deployment testing plus RLHF or constitutional AI) may suffice for weaker systems, while frontier systems capable enough to defeat those techniques may have enough situational awareness to bide their time — leaving no compensable harm before catastrophe. A further risk is Goodhart’s Law: pressure from punitive damages may cause firms to suppress the compensable harms without reducing uninsurable risk, breaking the correlation the framework depends on.
    • Mitigating point: it is the expectation of warning shots, not their actual occurrence, that drives deterrence — and insurers are likely to treat “we won’t have warning shots” claims skeptically.
  • Institutional competence. Courts may be unable to estimate these parameters accurately — but any AI governance approach faces the same problem, and ex post liability at least avoids committing to specific alignment protocols that may become obsolete.

Conclusion

  • Direct ex ante regulation is “extremely difficult to get right,” facing both technical obstacles (we cannot certify that a capable system is not deceptively misaligned) and political ones (geopolitical competition, a pro-innovation ethos).
  • A liability framework “doesn’t have to be perfect to buy us a lot of risk mitigation.” Its value lies in converting “soft” safety concerns into priced exposure that boards, investors, and insurers must account for.
  • The concrete mechanism Weil closes on: when a model fails an independent safety evaluation, weak liability pressure favors the cheapest path to passing — behavioral fine-tuning that suppresses visible failures. Credible punitive exposure would instead favor rollback to earlier checkpoints, adversarial training, tighter deployment constraints, and interpretability, and would strengthen internal voices arguing for delay over “ship and patch.”
  • Technical safety research and legal-institutional design are complements, not substitutes: evaluations make liability pricing less arbitrary, while liability pressure creates demand for credible measurement.
  • The core moves do not require consensus on extinction probabilities — only that certain frontier development decisions create tail risks that ordinary negligence doctrine and compensatory damages predictably underprice.

Notes

  • Author listed on this version as Assistant Professor of Law, Touro University Jacob D. Fuchsberg Law Center; he is elsewhere identified with the University of Houston Law Center and the Institute for Law & AI.
  • Produced with funding and technical support from the PIBBSS Fellowship; research assistance funded by a grant from Open Philanthropy.