Abstract

Autonomous AI agents that plan and execute open-ended tasks are being deployed rapidly, but basic facts about who builds them, where they are used, and how they are tested were until recently unavailable. The AI Agent Index supplies those facts and finds a systematic gap: developers publish code and documentation but rarely safety policies or external evaluations. Because agents act rather than merely produce content, closing this gap requires both technical governance mechanisms and legal frameworks — a project neither computer scientists nor legal scholars can complete alone.

The shift from content-producing models to agents

  • Leading AI companies have released autonomous agents that plan and execute complex digital tasks with limited human involvement: OpenAI’s Operator, Google’s Project Mariner, Anthropic’s Computer Use Model. All type, click, and scroll in a browser to order groceries, make reservations, and book flights.
  • Performance is currently unreliable, but scores on multiple benchmarks are steadily improving. The stated aspiration is agents serving as artificial personal assistants and virtual coworkers across personal and professional activities.
  • The distinguishing feature relative to language models and generative tools: agents independently take actions to accomplish lengthy open-ended goals rather than primarily producing content.
  • Examples given: OpenAI’s Deep Research analyses retrieved information, decides to perform additional searches, and produces reports rivalling those of human experts; Google’s Project Astra is a prototype universal assistant operating across phones and glasses.
  • Economic opportunities span automating household purchases and travel arrangements through to conducting cutting-edge scientific research.

Risks specific to agents

  • Malicious actors using agents to autonomously carry out cyberattacks and perpetrate online fraud.
  • Broader concerns from changes in human behaviour and social structures as people delegate personal and professional tasks to agents.
  • Users losing control over their agents, or discovering the agents engaged in undesirable or unethical activity.
  • These risks stem from the distinct feature of agents — the ability to take actions in pursuit of goals — and differ from those associated with ordinary content-producing models.

Why the risks are hard to tackle

  • The technology is progressing quickly and being deployed across diverse domains, each presenting distinct challenges.
  • Limited publicly available information. Until recently there were no reliable answers to: which organisations are building agents; in which domains they are deployed; what infrastructure agents rely on; how performance and safety are evaluated; and what steps are taken to mitigate risks.

The AI Agent Index

  • Co-led by the author and Stephen Casper with researchers from MIT, Stanford, Harvard, and other institutions; the first public database documenting technical, safety, and policy-relevant features of deployed AI agents.
  • Contains 33 fields of information across 67 agents, collected manually from public documentation and correspondence with developers, spanning how agents are built and tested and details about the organisations building them.

Findings

  • Domain concentration. 75% of agents specialise either in using computers for diverse tasks (like Google’s Mariner) or in assisting with software engineering — suggesting governance should focus on these broad domains rather than narrower applications.
  • Geographic and sectoral concentration. 67% of developers are based in the United States (12% in China) and 73% are companies in industry rather than academic institutions — suggesting governance should focus initially on the US and account for the incentives of industry actors.
  • Safety disclosure gap. Most developers release code and documentation, but fewer than 20% disclosed a formal safety policy and fewer than 10% reported external safety evaluations.
  • The disclosure gap is described as a red flag for users and for actors concerned with societal impact, underscoring the need for systematic testing and more robust transparency and accountability mechanisms.
  • A first step would be for governance institutions to establish and maintain an agent index of their own, though this alone would be insufficient.

Technical governance mechanisms

  • Mechanisms explored by computer scientists include requiring human approval for certain actions, automatically monitoring agent behaviour, and enabling shutdown in the event of malfunction or misconduct.
  • Other proposals focus on IDs for agents, accessible to users, auditors, and other stakeholders, improving visibility into operation and impact.
  • Drawing on internet architecture, researchers propose that agent governance infrastructure perform three core functions:
    • Attribution — tying specific actions to particular agents, including verifying that an agent acts on behalf of a particular individual or organisation.
    • Interaction shaping — establishing protocols for communication and cooperation between agents.
    • Detection and remedy — identifying harmful actions and enabling certain agent actions to be reversed.
  • Contract law: will the actions of agents legally bind the users who instructed them?
  • Tort law: how will liability for harm caused by an agent be allocated among users, developers, and intermediaries?
  • Criminal law: can any of these actors be held criminally liable, under what conditions, and in which jurisdictions?
  • Regulation: how do instruments such as the EU AI Act affect the application of existing law to agents?
  • Agents are not being developed in a legal vacuum but within an existing tapestry of legal rules and principles. Studying these is necessary both to anticipate how legal institutions will respond and to design technical governance mechanisms that operate alongside existing legal frameworks.

The interdisciplinary claim

  • Progress in technical governance requires legal knowledge; effective legal frameworks require technical expertise.
  • The closing position: governance of AI agents is a deeply interdisciplinary project in which computer scientists and legal scholars together have the opportunity and responsibility to shape the technology’s trajectory.
  • A version of the article first appeared in Lawfare.