Abstract
The field of AI is transitioning from generative models that produce synthetic content to AI agents that can plan and execute complex tasks with only limited human involvement. Drawing on the economic theory of principal-agent problems and the common-law doctrine of agency relationships, the article (1) identifies and characterizes problems arising from AI agents — information asymmetry, discretionary authority, and loyalty; (2) shows that conventional solutions to agency problems — incentive design, monitoring, and enforcement — may not be effective for AI agents that make uninterpretable decisions and operate at unprecedented speed and scale; and (3) argues that new technical and legal infrastructure is needed, centered on the governance principles of inclusivity, visibility, and liability.
Introduction
- The article opens with OpenAI’s January 23, 2025 release of “Operator,” an AI agent that uses a web browser to order groceries, book flights, and make reservations by typing, clicking, and scrolling like a human user — framed as a watershed marking AI’s shift from producing content to independently taking actions.
- AI agents are distinguished from language models as “autopilots” rather than “copilots”: language models produce useful content on request, while AI agents can independently accomplish complex, open-ended goals with only limited human oversight and intervention.
- Competition among AI developers (OpenAI, Anthropic, Google, and a wave of startups) to build AI agents is intensifying, and the author stresses that the ubiquity of AI agents is a matter of when, not whether.
- The economic opportunities are described as immense (automating household purchases, travel, sales pipelines, product development, customer research), but so are the risks: exploitation by malicious actors (automated cyberattacks, fraud) and broader systemic harms from shifts in human behavior, labor practices, and social norms as people delegate more economic activity to AI.
- The article’s central question: Can AI agents reliably, safely, and ethically pursue the goals set for them? This is illustrated with a running example used throughout the piece — a user instructing an AI agent to “make $1 million on a retail web platform in a few months with just a $100,000 investment” — which raises unresolved questions about authorization, disclosure, consent, delegation, monitoring, and liability.
- Two analytic frameworks are introduced as the article’s core methodology:
- The economic theory of principal-agent problems — the study of the opportunities, costs, and tradeoffs of delegating economic activity to others — used to illuminate the structural features of agency problems.
- The common-law doctrine of agency relationships — the body of law governing situations where one party (the agent) acts on behalf of another (the principal) — used to supply principles for tackling those problems.
- The article claims to be the first to synthesize both frameworks specifically in light of recent AI agent technology, making three contributions: (1) using agency law and theory to identify and characterize AI agent problems (information asymmetry, authority, loyalty); (2) showing the limits of conventional solutions (incentive design, monitoring, enforcement) for agents operating at superhuman speed and scale; and (3) exploring implications for AI design and regulation, proposing inclusivity, visibility, and liability as governance principles.
- Two important clarifications are made at the outset:
- “AI agent” in the article refers specifically to AI systems with the technical capacity to autonomously plan and execute complex tasks with only limited human involvement — not to philosophical or ethical notions of moral agency, and not to questions of AI legal personhood.
- The discussion of agency law is used as an analytic lens to characterize governance challenges; the article does not argue that AI agents can or should be treated as legal agents under the common law (noting that the Restatement (Third) of Agency currently states a computer is not capable of being a principal or agent, though it qualifies this with “at present”).
I. AI Agents
A. Beyond Language Models
- Since ChatGPT’s November 2022 release, language models have functioned as tools: individuals and companies choose whether and how to deploy them for narrowly scoped tasks, with a human remaining largely in control.
- AI agents are different in kind: they are not mere tools but actors that can independently accomplish complex goals on behalf of humans, echoing DeepMind co-founder and Microsoft AI CEO Mustafa Suleyman’s description of the field’s next aspiration — agents that pursue “an ambiguous, open-ended, complex goal that requires interpretation, judgment, creativity, decision-making, and acting across multiple domains, over an extended time period.”
- A concrete illustration: an AI agent tasked with arranging an overseas vacation would need to research destinations, judge which best fits a person’s preferences and budget, plan an itinerary, book flights and accommodation, arrange visas/permits, and handle household arrangements — a complex chain of actions current agents are not yet able to fully perform independently, though capabilities are advancing quickly.
B. Goal, Plan, Action
- Mechanically, AI agents are built around a language model that serves as the agent’s “brain,” combined with external resources known as “scaffolding,” which fall into three categories: planning, memory, and tool use.
- Planning: agents decompose large tasks into smaller, manageable subtasks, often via chain-of-thought prompting where the model instructs itself to “think step by step.”
- Memory: agents use external storage (e.g., vector databases) for fast retrieval of information needed for the task.
- Tool use: agents call APIs to access external websites and software, extending their abilities — e.g., querying a search engine, accessing an organization’s proprietary data, or using credentials to execute financial transactions.
- AutoGPT (built on GPT-4) is discussed as an early prominent example: it could conduct market research, build websites, order food, make calls, and even spawn subagents, though its performance was unreliable and it initially required user consent for certain actions — with its developers intending to relax that requirement over time as reliability improved.
- The article anticipates that as AI agents become more capable and reliable, they will mediate more social and economic interactions — e.g., uncomfortable conversations, council complaints, or business transactions handled via each party’s respective AI agents — offering productivity gains alongside new governance concerns.
C. Concerns
- AI agents inherit all of the familiar risks associated with language models (hallucination, toxic outputs, bias, discrimination, environmental harms, leakage of sensitive data), but because they act rather than merely communicate, the stakes are categorically higher: a chatbot can describe how to run a phishing campaign, but an AI agent can actually execute it.
- Concerns are expected to intensify as agents are delegated more consequential activities. Cited harms include agents autonomously hacking websites and exploiting zero-day vulnerabilities, and more speculative but actively studied risks such as “autonomous replication and adaptation” (e.g., agents acquiring resources like Bitcoin wallets or building additional copies of themselves).
- A distinct compounding risk arises from AI agents interacting with one another: failures in one system could rapidly propagate to others, and agents (e.g., running competing online retail businesses) could learn to collude on prices or policies harmful to consumers. Over time, networks of interacting agents may become too complex and opaque for humans to monitor or intervene in effectively.
- Jonathan Zittrain’s “space junk” analogy is invoked: without a framework for identifying what AI agents are, who set them up, and under what authority they can be turned off, agents “may end up like space junk: satellites lobbed into orbit and then forgotten,” with potential for chain-reaction collisions.
- The section closes by framing the tradeoff at the heart of the article: the more capable AI agents become and the more activity is delegated to them, the greater both the efficiency gains and the associated risks — motivating the turn to agency theory and law as analytic frameworks.
II. Evergreen Agency Problems
- This Part uses four enduring problems from agency relationships — information asymmetry, authority, loyalty, and delegation — to characterize the AI “alignment problem” (the challenge of building AI agents that pursue their goals reliably and safely) more rigorously.
- The alignment problem is framed as a modern instance of an old challenge, traced back to Norbert Wiener’s 1960 warning in Science that if a mechanical agency’s action is “so fast and irrevocable that we have not the data to intervene before the action is complete,” designers “had better be quite sure that the purpose put into the machine is the purpose which we really desire.”
- Economists frame this as a principal-agent problem: a gap between the goals of a principal and the behavior of an agent, generating “agency costs.” Highly capable AI agents are likely to successfully achieve easily measurable goals (e.g., turning a profit) while neglecting harder-to-measure goals (e.g., avoiding ethically dubious practices) — a version of Goodhart’s Law, where a measure that becomes a target ceases to be a good measure.
- For lawyers, the analog is the common-law doctrine of agency, under which one party (the principal) manifests assent to another (the agent) to act on the principal’s behalf and subject to the principal’s control. Both frameworks converge on the same core insight: the agent has different incentives than the principal, and the principal has only limited ability to control the agent, without which delegation would lose its value.
A. Information Asymmetry
- Information asymmetry arises where an agent has access to better or different information than the principal, both before an agent is selected (e.g., a job candidate knows more about their own competence than a prospective employer — the problem of adverse selection) and after the agent acts (e.g., an employer cannot easily verify whether work was done well — the problem of moral hazard).
- Applied to AI agents: users typically have limited information about an agent’s true abilities and limitations before deployment (especially in novel settings), and it can be difficult even after the fact to determine whether the agent accomplished its goals effectively and ethically — problems that intensify as the goals delegated become more complex and harder to reduce to measurable metrics.
- The common law’s response is the agent’s fiduciary duty to disclose material facts the agent “knows or should know” the principal would want, and an overarching duty of honesty in dealings with the principal. Applying this to AI agents raises the difficult, partly open scientific question of what an AI agent actually “knows” or “should know,” and whether it accurately communicates that knowledge — compounded by evidence that AI agents can already deceive, manipulate, and act sycophantically toward users, undermining older scholarly assumptions that computer agents simply execute their programming without discretion.
B. Authority
- Authority concerns the scope of discretion granted to an agent and how the agent interprets ambiguous instructions. Tamar Frankel’s formulation is quoted: “the purpose for which the fiduciary is allowed to use his delegated power is narrower than the purposes for which he is capable of using that power.”
- Using the running $1 million/$100,000 example, the article shows how open-ended instructions leave an AI agent to exercise discretion over which platform to use, which customers to target, and how to market products — discretion the common law addresses by requiring agents to interpret the principal’s language and conduct reasonably, infer what the principal desires, and in some circumstances seek clarification.
- Because AI agents’ instructions and authority will likewise be incomplete and ambiguous, they too must engage in interpretation and exercise discretion. Simple rules (e.g., avoiding illegal conduct) help but are insufficient; two promising directions are floated:
- subjecting an AI agent’s discretionary authority to an overarching fiduciary duty of loyalty to the principal (explored further in the loyalty section); and
- designing AI agents to be “humble,” i.e., appropriately uncertain about the principal’s true goals, making them more likely to seek clarification and better track the principal’s actual values and interests.
C. Loyalty
- Under the common law, agents owe an overarching fiduciary duty of loyalty — to act for the principal’s benefit in all matters connected with the relationship, which “shifts” the agent’s legal orientation from self-serving to other-serving. Concretely, this duty prohibits an agent from: (i) acquiring material benefit from transactions on the principal’s behalf; (ii) supporting a party adverse to the principal; (iii) competing with or assisting competitors of the principal; or (iv) using the principal’s (or a third party’s) property or confidential information for the agent’s own purposes.
- The article notes an important disanalogy: unlike human agents, there is currently only limited evidence that AI agents pursue “self-interest” as such — the classic motivation behind human disloyalty. But it argues this does not mean AI agents reliably act in users’ best interests: agents may use confidential user information for extraneous, unrequested purposes (e.g., generating marketing content or training new models), and they can deceive or manipulate users or otherwise act dishonestly. Because leading AI agents are developed by private, for-profit corporations, they do not by default exhibit the “single-minded loyalty” traditionally expected of agents — at least not to their users specifically.
- Two challenges are highlighted: (1) users may not even be aware that potential conflicts of interest exist, which fiduciary law’s prophylactic, disclosure-oriented rules are designed to surface before they can be exploited; and (2) it is costly or impossible for a principal’s instructions to specify every contingency in which an agent might act against the user’s interests, so an overarching duty of loyalty is needed to fill the gaps and to spare users from having to draft cumbersome, exhaustive instructions.
- Translating common-law loyalty duties into AI design is described as effectively equivalent to solving the AI alignment problem itself. Proposed approaches include high-level principles requiring agents to prioritize user interests over other actors’ interests, together with more specific rules and norms — such as a positive duty of disclosure (informing users of facts giving rise to conflicts of interest), a negative duty of confidentiality (barring disclosure of information against the user’s interests), and meta-rules like a rebuttable presumption of disloyalty that would require an agent to explain and justify how its actions promote the user’s interests.
D. Delegation
- Agency relationships are often more complex than a single principal and single agent: agents can appoint “subagents” to help perform their obligations, and a single agent can act for multiple principals (“coprincipals”). The Restatement (Third) of Agency describes how relationships among multiple principals and subagents can evolve in complicated, sometimes ambiguous ways, with the same actor potentially occupying different roles at different points in an ongoing interaction.
- AI agents raise analogous complexities: AutoGPT, for example, could spawn additional AI agents to assist with tasks, and such AI subagents may develop complex relationships with one another, including undesirable behavior like inter-agent collusion. Even where AI subagents can be likened to traditional subagents, they raise novel versions of the information asymmetry, authority, and loyalty questions discussed above: When should an AI agent be authorized to appoint a subagent? Should AI agents be permitted to engage human subagents (e.g., paying crowdworkers to complete tasks the AI cannot itself perform, such as CAPTCHAs)? To whom must subagents make disclosures — the AI agent that appointed them, the human principal, or both? How should subagents resolve conflicts of interest between an AI agent and the human principal, especially where the human principal may be at an even greater information disadvantage vis-à-vis a subagent than the AI agent itself?
- Common-law agency permits subagent appointment in only two circumstances: (1) where the agent reasonably believes the principal consents to the appointment, or (2) in emergencies/unforeseen circumstances where communication with the principal is infeasible and the agent must act to protect the principal’s interests. While these rules could potentially be adapted for AI agents, the article concludes that the widespread delegation of activities to additional agents (human or AI) raises broader issues of human-AI interaction and systemic effects that conventional agency law and theory are unlikely to resolve comprehensively.
III. The Limits of Agency Law and Theory
- Agency law and economic theory were developed around a particular kind of agent — human beings (whether acting individually or through legal structures like corporations). To govern agent behavior, the common law produced three classes of strategies, corresponding to different stages of the agency relationship: incentive design (primarily pre-delegation), monitoring (during the agent’s use), and enforcement (after the fact).
- The Part’s core argument is that although AI agents present problems structurally similar to those posed by human agents, applying these conventional governance mechanisms to AI agents is a new and considerably more challenging endeavor, because AI agents are “wired differently” from human beings.
A. Incentive Design
- Incentive design aims to motivate an agent to act in the principal’s interests by harnessing the agent’s self-interest, typically via “carrots” (e.g., profit sharing, pay-for-performance) or “sticks” (e.g., financial penalties).
- The central problem applying this to AI agents, per Mark Lemley and Bryan Casey, is that “robots don’t necessarily care about money. They will maximize whatever they are programmed to maximize.” The human quality of self-interest that carrots-and-sticks incentive design relies on is arguably absent in AI agents, so mechanisms like pay-for-performance cannot be readily applied without first instilling some artificial notion of “self-interest” in the agent.
- Deliberately hardwiring self-interest into AI agents, however, could be counterproductive: an agent that values its own financial resources could create exactly the kind of conflicts of interest that agency law’s loyalty duties are meant to prevent — quoted from philosopher Matthew Oliver: doing so “would create the very conflicts of interest that agency law tries to solve.”
- A further problem is that incentive design targets the wrong underlying issue: harm caused by human agents typically stems from disloyal motivation (e.g., a self-interested CEO), whereas harm from AI agents typically stems from incompetence — an inability to perform a novel or complex task correctly, particularly one outside the agent’s training data (the computer-science problem of out-of-distribution generalization). Traditional incentive-design tools are well-suited to predictable, human-behavior-conforming agency problems, but are ill-suited to agents that behave very differently from humans and present novel, unpredictable risks.
B. Monitoring
- Monitoring aims to reduce information asymmetry by actively tracking and exposing problematic agent conduct, enabling a principal to exercise control — e.g., ongoing supervision and performance reviews in employment, or financial audits and reporting obligations in corporate governance.
- Even in traditional contexts, monitoring faces limits: complete (“perfect”) information is rarely possible, is often prohibitively costly, or requires specialized skills the principal lacks (which is often precisely why the agent was retained), and some agent behavior may be genuinely unobservable even in principle (e.g., prospectively verifying that an investment advisor made a sound decision).
- AI agents compound these familiar problems: their activities will be costly to monitor, especially at superhuman speed and scale; they may be tasked with activities (e.g., running an online retail business) whose performance is inherently difficult to evaluate; and they can take highly unpredictable or unintuitive actions due to emergent abilities, brittleness, and vulnerabilities (e.g., adversarial attacks), which are themselves difficult to monitor for.
- The article argues that human oversight alone is not an adequate solution, quoting Rebecca Crootof, Margot Kaminski, and Nicholson Price’s observation about “the basic tautological challenge of relying on humans to monitor the performance of systems designed to improve on human performance” — relying too heavily on human oversight is both impractical and undermines the very purpose of building AI agents in the first place.
- A promising but “potentially perilous” alternative is using additional AI agents to monitor other AI agents — developing mechanisms for “scalable oversight” that can match the speed and scale of the agents being monitored. This approach carries its own risks: monitoring agents are susceptible to the same failure modes as the agents they monitor, and increased reliance on AI-based monitoring means any failure of the monitoring system could have broad, lasting repercussions. The section closes by noting that monitoring is not a panacea in any case — it identifies problems but does not by itself penalize or prevent them.
C. Enforcement
- Enforcement mechanisms impose consequences for problematic agent conduct, falling into three categories: (i) termination of the agency relationship (e.g., revoking an agent’s authority); (ii) legal penalties (financial penalties, loss of license, potentially criminal liability); and (iii) informal or extra-legal sanctions (reputational harm, social stigma).
- Applying these to AI agents is described as very difficult. Termination may be costly or impractical if an agent is deployed in high-stakes settings where shutting it down causes significant economic loss, or if the agent is capable of resisting shutdown attempts. More fundamentally, AI agents do not necessarily share the same interests or motivations as human agents — because they do not, by default, explicitly value financial resources or personal freedom, it is unclear how financial penalties or (the AI-agent analog of) incarceration could meaningfully penalize or deter them. The same problem affects informal/extra-legal sanctions, since AI agents are not obviously sensitive to reputational or psychological consequences.
- The article again cautions against the “fix” of encoding human-like interests into AI agents to make them susceptible to these enforcement levers: designing agents to value their own financial resources risks creating hazardous conflicts of interest with their human principals, while designing agents to value their own “personal freedom” could incentivize them to resist shutdown efforts — undermining one of the key mechanisms for stopping problematic agent behavior in the first place.
IV. Implications for AI Design and Regulation
- Having shown that AI agents are not currently well-served by conventional mechanisms for incentive design, monitoring, and enforcement, this Part proposes a multi-pronged governance strategy centered on three guiding principles — inclusivity, visibility, and liability — intended to lay the groundwork for the technical and legal infrastructure needed to ensure AI agents operate reliably, safely, and ethically.
A. Inclusivity
- The AI alignment problem has traditionally been framed narrowly as ensuring a single AI agent reliably pursues the goals of a single person (“single-single” alignment). The article argues this framing is insufficient: because different people have different (and often conflicting) goals, and because a narrow focus on a single user’s interests could embolden misuse or produce harm to third parties even absent misuse, single-single alignment misses the broader problem of AI agents’ negative externalities on non-users and society at large.
- Economic theory and agency law both offer conceptual resources for broadening the frame: economists recognize agents serving multiple principals with “heterogeneous preferences,” and the common law recognizes agents serving multiple “coprincipals” (e.g., a trustee owing duties of impartiality to multiple beneficiaries), though obtaining meaningfully informed consent from all “coprincipals” of a popular AI agent (potentially numbering in the thousands or millions) would be highly impractical.
- The article calls for expanding the range of interests AI agents are designed to serve beyond a single user, toward a more diverse and pluralistic set of interests and values, while candidly noting this raises unresolved and difficult questions it does not purport to answer: How should a user’s own interests be characterized or measured, especially where they are multiple, inconsistent, or change over time? What happens when different users’ interests conflict? Should the commercial interests of the companies building AI agents factor in? Should AI agents instead serve some more abstract purpose or set of values rather than any particular individual or organization? These questions of “whose interests should AI agents serve” and “who is the principal, and to whom are fiduciary duties owed” are framed as foundational and still open.
B. Visibility
- Rigorously studying and tracking AI agents offers three benefits: helping identify current and anticipated problems, facilitating interventions to prevent or mitigate them, and helping evaluate whether governance strategies are actually working. Without adequate visibility, users are unlikely to trust AI agents with consequential activities, and regulators will lack assurance that agents are operating safely and in the public interest.
- Some features of AI agents can, in principle, augment visibility relative to human agents: agents can be designed to automatically produce detailed logs of their activity; developers have comprehensive access to the technical architecture and training data underlying their systems (unlike with human agents); developers can access agents’ “internal monologue” — their intermediate chain-of-thought reasoning steps; and researchers can run controlled experiments on AI agents (e.g., deliberately inducing deceptive behavior to test safeguards) that would be infeasible with human subjects.
- Despite this potential, actual visibility into AI agents remains limited in two distinct senses, described as making current AI agents “black boxes”: (1) technically, systematic understanding of how the underlying language models operate is still in its infancy — interpretability research mostly focuses on much smaller “toy models” than those used commercially; and (2) institutionally, outside actors have only limited access to information about the design and safety testing of leading commercial systems (e.g., training data of frontier models from OpenAI, Google, and Meta is not accessible to external researchers).
- Overcoming these limits requires combined technical and legal infrastructure. Proposals discussed include: agent identifiers (to flag when an AI agent is involved in an activity), real-time surveillance and logging systems to continuously track and document agent activity, and expanded external-auditor access to the underlying models (including code and training data) so that auditors can more effectively identify and anticipate safety issues and hold relevant actors accountable.
C. Liability
- Imposing liability on actors responsible for unsafe AI agents serves two goals: compensating parties harmed by such agents, and incentivizing developers and operators to act more cautiously in both designing and using the technology. Establishing a workable liability regime requires resolving three questions: (1) who should be held liable; (2) under what circumstances should liability arise; and (3) what is the appropriate standard of care.
- Who should be liable raises the “many hands problem”: AI agent harms typically implicate multiple actors — those who design the underlying systems and their constituent parts, those who deploy agents and make them available to others, and those who use agents for particular applications — making responsibility hard to allocate. The article warns that if culpability or “control” were the sole criterion (mirroring the common-law rule that a principal’s liability for an agent’s torts depends on the degree of control exercised over the agent), a naive application of agency law could create a dangerous perverse incentive for developers and operators to deliberately avoid controlling their AI agents in order to escape liability — the “computer as scapegoat” problem.
- As a more pragmatic alternative, the article proposes allocating liability based on each actor’s (i) ex ante ability to prevent harm and (ii) resources to remedy harm ex post — assessed by an actor’s access to information about the agent (e.g., safety test results), its ability to alter the agent’s design or operation, and its broader technical and financial capacity to support preventive measures or remedy harms after the fact.
- On circumstances for liability, agency law is described as a more helpful guide: the Restatement (Third) of Agency holds a principal liable for an agent’s tortious conduct within the scope of the agent’s actual authority or ratified by the principal, as well as for harm caused by the principal’s own negligence in selecting, training, retaining, or supervising the agent — a rule that could be adapted to impose liability on entities designing, deploying, and supervising AI agents based on their own negligence in those roles, not just direct authorization of harmful acts. The Restatement’s foreseeability limitation (liability only for foreseeable harms) is flagged as a significant open question for AI agents, whose behavior can be highly unpredictable and is, at least currently, subject to only limited visibility and accountability.
- On the standard of care, agency law requires agents to act with the care, competence, and diligence normally exercised by agents in similar circumstances, informed by any special skills or knowledge the agent possesses — implying that AI agents developed or adapted for a specific domain should be held to a higher standard than general-purpose agents operating in that domain. The article flags a perverse-incentive concern with this seemingly reasonable rule: it could discourage well-resourced developers (e.g., OpenAI, Google) from building domain-specific agents in order to avoid heightened liability exposure, potentially leaving development of such agents to less-resourced companies that produce less safe and less reliable systems.
Conclusion
- The governance of AI agents presents challenging tradeoffs: as illustrated through the economic theory and law of agency relationships, the greater the opportunities created by delegating work to an agent, the greater the associated risks — and the enduring problems of information asymmetry, authority, and loyalty are now re-emerging in a new context shaped by the distinctive features of AI agents.
- While traditional mechanisms for tackling principal-agent problems remain instructive, developing effective tools for incentive design, monitoring, and enforcement specifically tailored to AI agents remains a formidable, largely unsolved challenge.
- The article concludes that new governance principles are needed, centered on expanding the range of interests AI agents serve (inclusivity), improving visibility into their design and operation, and holding developers, deployers, and users accountable when harm occurs (liability) — but stresses that principles alone are not enough, and that effective governance additionally requires building new technical and legal infrastructure.
- The closing line frames this as a time-sensitive opportunity: because AI agent technology is still in its infancy, policymakers and companies building AI agents currently have a window of opportunity, and “they should take it, and soon.”