What lawyers can do for AI safety

Core claim

The technology is transformative, the risks catastrophic, and the rate of change overwhelming — and the current bottleneck for AI safety is political will rather than research. Lawyers have the skills (advocacy, adversarial argument, handling complex material) and the access (legislators, judges, media) to build that will, and the legal profession needs to be stirred into action.

Framing

  • The author is a barrister with over a decade in litigation and human rights.
  • Stated goals: bring lawyers into AI safety, explain why AI safety needs more lawyers, suggest routes for movement-building and gathering political will, and foster interdisciplinary communication.
  • Why AI safety researchers should engage with outsiders:
    • It increases impact and sharpens thinking — a twist on the line attributed to Feynman: if you can’t explain something to a lawyer, you haven’t really understood it.
    • It improves the research itself; AI safety calls for multi-disciplinary effort, and interdisciplinarity spurs new ideas, improves scientific writing, and builds stronger academic communities.

A personal taxonomy

Offered as a coping mechanism for the two problems a lawyer new to AI safety encounters — volume of information and separating signal from noise. Four categories:

  1. Technology — the software (code, underlying maths, training data) and hardware (chips, data centres, financing, energy, infrastructure). Each is a locus of competition, research, investment, policy intervention and political pressure.
  2. Tools — the models themselves, their uses, and what is built on or around them: chatbots, agents, harnesses, classifiers, recommendation systems, audio-visual generation. Noted as the source of much fatigue and switching-off, given the “unending torrent” of apps, workflows and “solutions.”
  3. Society — discussion, reporting, research and debate on the use, diffusion and impact of AI on individuals, communities, nations and the environment.
  4. Safety — AI governance and regulation, model evaluations, control, alignment, risk mitigation.

Side note: no universal definition of “alignment”

Six definitions collected from different sources:

SourceDefinition
Richard NgoEnsuring AI systems pursue goals that match human values or interests rather than unintended and undesirable goals
Nate SoaresHow in principle to direct a powerful AI system towards a specific goal
Holden KarnofskyBuilding very powerful systems that don’t aim to bring down civilisation
AnthropicBuilding safe, reliable and steerable systems as those systems become as intelligent and as aware of their surroundings as their designers
OpenAIMaking AGI aligned with human values and following human intent
IBMEncoding human values and goals into LLMs to make them as helpful, safe and reliable as possible

How lawyers can help

  • What the profession trains you to do: advocate, advise, guide, and help clients navigate complicated institutions and systems. The work trains lawyers to hunt for ambiguity, build arguments, test evidence, and craft strong and memorable narratives. In the common-law world it is adversarial — processing lengthy material, strategising, researching obsessively, then arguing in public against an equally prepared opponent looking to dismantle every point.
  • The leverage point: the current bottleneck is political will, not research. Lawyers have access to legislators, judges and media, and building political will is something the profession knows how to do.
  • Concrete contributions listed:
    • find the rules that need changing, or those that might help achieve AI safety goals;
    • help navigate complex institutions and systems;
    • draft policy, standards, regulations and legislation;
    • design monitoring, enforcement and evidence-gathering mechanisms;
    • educate policymakers, the public, stakeholders and AI researchers;
    • develop frameworks for collaboration, coordination and communication between organisations;
    • tailor federal or international laws and policy proposals to a local jurisdiction;
    • promote clear language — shifting discourse away from anthropomorphising language, and translating AI safety concepts for policymakers.

Current challenges

Two aspects to the work: ensuring good outcomes from AI, and averting the litany of possible risks. The issues below are presented both as entry points for the motivated and as points of leverage to slow model development or contribute to alignment work.

  • AI and children’s safety — impact on identity, education, attachment, cognition and development; deepfakes and exploitation.
  • Liability and insurance — targeted policy here may change the behaviour of AI companies or the companies they rely on.
  • Two budding research fields:
    • law-following AI — designing agents that obey the law;
    • legal alignment — studying how legal rules, principles and methods can help address alignment problems.
  • Standards and red lines.
  • Verifiable oversight — regulatory structures and private mechanisms for monitoring capabilities, training and evaluation, including post-deployment. Sub-examples: whistleblower protection for AI company employees reporting safety failures or capability advances; actions available to business and civil society; helping design a science of model evaluations that produces evidence in a form usable by policymakers.
  • AI-related human rights — human rights apply universally, provide a globally shared language, and galvanise social movements; notably an area where middle powers can play a major role. Examples: a right to know when one is interacting with AI; a right to human-human interaction; growing scientific consensus on cognitive and psychological risks (mental health, learning, agency, deskilling and addiction).
  • Autonomous weapons systems — criminal liability and regulation of use.
  • CBRN weapons risks and the shifting nuclear risk landscape.
  • Privacy and AI-supercharged surveillance.
  • AI’s impact on democracy.
  • Copyright.
  • Labour rights — including the lack of transparency around the globally distributed, largely invisible workforce doing data annotation and content moderation (if these workers were employed directly and paid fairly, “the cost of producing training data and running reinforcement learning would multiply severalfold”); adequacy of existing employment law; labour market risks; employer reliance on automated recruiting, performance management and dismissal.
  • Appropriation of Aboriginal and indigenous cultural knowledge without consent.
  • Environmental issues.
  • Criminal justice
    • generating convincing fake audio-visual material is now trivial: do criminal laws or rules of evidence need amending to maintain public trust in the legal system?
    • AI can improve access to justice (e.g. Miranda from Stanford’s April law hackathon; Princeton’s work assisting public defenders);
    • rules and transparency around judges using AI (a study found a majority of US federal judges are using it);
    • police use of AI — a UK officer is under investigation for allegedly using AI to generate statements in criminal cases, and police were ordered to stop using AI to prepare court statements;
    • the suggestion that AI might need criminal law.
  • Courts and law
    • using AI to promote swifter, fairer and more accessible justice;
    • reviewing how well existing law in your jurisdiction (employment, product liability, negligence, copyright, criminal, administrative, corporate governance) handles these issues;
    • whether your jurisdiction has solid rights to identify and challenge automated decisions;
    • whether conversations with AI chatbots should be privileged;
    • pushing your local law society or bar association to provide AI training, building capacity for lawyers to engage in policy- and law-making;
    • the flood of AI-generated paperwork — systems can now generate hundreds of pages of apparently relevant material that court staff must process, parties must answer, and judges must read, inflating resolution times in already backlogged, resource-strained courts;
    • France’s ban on publishing statistical analysis of judicial decisions, as a prompt to consider what educational, regulatory or informational interventions a jurisdiction might need.

Where to start

  • Recommended reading: The Problem (LessWrong) and AGI safety from first principles.
  • Organisations working in or around law and AI safety: Institute for Law & AI; CLAIR (Center for Law & AI Risk); Lawfare.
  • Training, events, getting involved: AISafety.com.

Noam Kolt · Nicholas Caputo · Peter Salib · Christoph Winter · Gabriel Weil · Cullen O’Keefe · Gillian Hadfield · Peter Henderson · Jonathan Zittrain · Seth Lazar

LessWrong and EA Forum posts relating to law

The author notes there are few such posts and lists them as at July 2026, including Legal scholarship: Is it high-impact?, LLMs as Fiduciaries to Humans, arguments against the AI safety community’s embrace of strict liability, the Law-Following AI sequence, The Threat of AI Crimes Are Under-Appreciated, Lawyers are uniquely well-placed to resist AI job automation, Learning societal values from law as part of an AGI alignment strategy, and Leveraging Legal Informatics to Align AI.

What the author is working on

Offered for comment, guidance or path-correction:

  • building a list of high-impact legal research papers of practical use to lawyers;
  • tracking what judges say about AI in decisions, as an indicator of how informed they are and of AI’s day-to-day impact in court;
  • designing criminal law evaluations;
  • applying courtroom skills to model evaluations — e.g. cross-examining models, or models as cross-examiners; formalising evals and whether there should be standards of proof;
  • an article idea, “How you can help the lawyers” — e.g. a database of expert witnesses relevant to aspects of AI safety.

Discussion

  • Talia Honikman suggested the legal alignment side — setting rules for agents whose internals cannot be inspected — resembles the textualism/purposivism debate, and that mens rea is effectively a graded procedure for adjudicating unobservable internals through behavioural evidence, offering more structure than most current evals. She asked whether graded scienter categories (purpose, knowledge, recklessness) do more work than standards of proof, since “is the model deceptive y/n” is a binary worth breaking up.