Leashes, Not Guardrails: A Management-Based Approach to Artificial Intelligence Risk Regulation
Abstract
Physical guardrails lower risk by keeping behaviour on a predetermined path; the regulatory equivalent for AI is “unrealistic and even unwise” because AI is too heterogeneous and dynamic to confine to fixed paths. Policymakers should instead impose leashes: flexible, adaptable management-based regulation that permits exploration while requiring firms to maintain internal risk management systems — and, crucially, that demands continuing human oversight, since “a physical leash only protects others when a human retains a firm grip on the handle.”
Publication details
- University of Pennsylvania Public Law and Legal Theory Research Paper No. 25-04, circulated February 2025; the paper states it has been accepted for publication in the journal Risk Analysis.
- Cary Coglianese (Edward B. Shils Professor of Law and Political Science; Director, Penn Program on Regulation) and Colton R. Crum (Ph.D. candidate in Computer Science & Engineering, Notre Dame).
Why “guardrails” is the wrong metaphor
- The term has proliferated across policymakers, analysts, advocates, and academics — cited examples include calls for “government-imposed guardrails” — until it covers “virtually any proposal for AI governance.”
- The regulatory design most akin to a physical guardrail is a fixed bright-line rule: outright bans (Italy’s initial prohibition of ChatGPT), prohibitions on particular uses (AI in employment screening, facial recognition), or mandates specifying exactly how technologies must be built or used.
- Four objections:
- Guardrails foreclose “off-road” discovery — AI’s versatility in “conquering new terrain where humans are otherwise unable to navigate” creates novel paths that can benefit society.
- They are immovable structures lacking the flexibility a dynamic technology demands.
- They emphasise fixed safeguards at the point of use rather than ongoing precautions across training, development, and deployment.
- Most importantly, once installed, a roadway guardrail protects without further human involvement — implying no continuing vigilance is needed, when in reality ongoing attention by AI firms’ managers and engineers is essential.
The heterogeneity problem
- AI is “a suite of different technologies, more than a single technology per se” — narrow tools built for specific purposes (skin cancer detection, movie recommendations), general-purpose foundation models adaptable by users or embedded in larger systems, and increasingly digital agents acting on behalf of humans.
- Heterogeneity runs deeper than purpose: neural network architectures vary in the ordering and configuration of neurons and in activation functions; training introduces further variability, since algorithmic configuration changes within a solution space “can contribute to wild fluctuations in model outcomes”; and training data sources and preparation methods add another layer.
- The upshot: no regulator can specify in advance a uniform set of rules adequate to the specific risks each firm and tool poses.
Leashes: management-based regulation
- Management-based regulation already operates wherever a “one size fits all” legal solution does not fit — toxic pollution prevention, food safety, accident prevention, chemical facility security — precisely because regulated entities are too heterogeneous for uniform prescriptive rules.
- The core requirement: the regulated entity must develop an internal plan to identify and monitor risks, produce and implement protective procedures, and document changes — subject to auditing and continual updating to prevent risk management decay.
- Its virtue is informational: it “eliminates regulators from having to have the same knowledge about AI for every firm” while permitting scalable reporting, detailed disclosure, and product testing that an outside regulator could not specify uniformly.
What an AI leash would require
- Standardised documentation of current uses, a description of the AI’s goals, and thorough identification of possible failures and harms.
- Internal practices following a Deming Cycle of “plan-do-check-act” — in the AI context sometimes called an AI impact assessment — a form of what the National Academies call “macro-means” regulation.
- Specified management practices such as red-teaming or adversarial testing during training.
- Regulator access to documentation, relevant data, and training information, plus external auditing of management plans and procedures. The authors suggest that peer review of the company’s plan by a regulator or third-party auditor might have prompted the additional night-time testing that could have averted the Arizona autonomous-vehicle pedestrian fatality.
- For social media, anticipating harms including mental health effects and building processes to monitor and address them — user feedback systems and reporting channels tracking harms that emerge after deployment. Precedent: the FAA’s anonymous reporting system for dangerous pilot behaviour; feasibility is shown by the voluntary AIAAIC repository of AI incidents, which could be made mandatory and centralised.
Leashes and guardrails are not mutually exclusive
- Firms subject to a leash would presumably install their own internal guardrails, and management-based regulation is compatible with targeted prescriptive rules where a best practice is obvious enough to be codified as a “micro-means” or “micro-ends” standard.
- Example given: requiring a model prompted for self-harm instructions to respond with mental health crisis hotline information — entirely compatible with also requiring the developer to run a risk management system that continually scans for problems, including those the micro-rule covers.
- “A management-based leash will always be needed to address the overall varied and rapidly changing risks associated with different forms of AI.”
The emerging management-based paradigm
- EU AI Act. Contains guardrail-like elements — Article 5’s prohibitions on government social scoring and image scraping for facial recognition — but for high-risk systems imposes leash-like requirements for a comprehensive, faithfully implemented risk management system. Article 9 stresses a “continuous iterative process” across the entire lifecycle: characterising foreseeable risks, evaluating them under intended purposes and foreseeable misuse, identifying new risks post-market through monitoring, and adopting management methods — with due consideration to the deployer’s “technical knowledge, experience, education, the training to be expected… and the presumable context” of use.
- Executive Order 14110. Never imposed binding requirements on firms and has since been rescinded, but called for risk management measures for foundation models including “the physical and cybersecurity protections taken to assure the integrity of that training process against sophisticated threats.” The responsive OMB memorandum required agencies using AI to conduct impact assessments and to perform ongoing monitoring and periodic human review, including real-world performance testing “at least annually, and after significant modifications.”
- Conclusion drawn: notwithstanding the rhetoric of guardrails, both instruments “exhibit core elements of a management-based or leashing approach.”
Three implementation questions
Management-based regulation has produced measurable results elsewhere — reduced toxic air pollution under state prevention planning laws, and fewer foodborne illnesses following HACCP requirements — but positive results are never guaranteed.
When should a leash be required?
- Just as leashes matter more for a German shepherd than a Cavapoo, and more in a playground than a wilderness, leashing is more appropriate where an AI tool can cause significant harm — depending on both the tool’s function and the environment of deployment.
- Initial heuristic: would a human performing the same task warrant risk regulation? If so, management-based requirements are presumptively justified. If not, ask whether the AI introduces new risk — “What’s the worst that can happen because of the AI?” Answers range from cancelled streaming subscriptions, to a chatbot producing slurs or threats, to accidents from robotic pizza delivery vehicles near campuses.
- Conventional risk analysis and benefit-cost analysis apply, attentive not only to deployment but to the entire machine-learning pipeline — leashes are more likely justified for tools deriving from unethical practices, sensitive web-scraping, or training on unauthorised or copyrighted material.
- Assessment should proceed with the status quo in mind: AI tools will make mistakes, and the question is how they compare to the errors of the existing human-dependent systems, “with their own well-known limitations of perception and decision-making.”
What constitutes an appropriate leash?
- Candidate measures: internal risk analysis, auditing, red-teaming, data acquisition controls, personnel training, monitoring, documentation, reporting. General parameters typically follow plan-do-check-act, as in the NIST AI Risk Management Framework — but leashes “will need to vary in their scope, stringency, and specificity,” as a large dog’s leash is thicker than a small dog’s.
- Regulators should refine an overall problem statement: what harms are possible; what is the firm’s goal for the tool; are those objectives aligned with the regulator’s; and are the goals meaningfully conveyed to the model during training, e.g. through a well-defined loss function. For social media, whether the objective simply maximises watch time, engagement, or impressions, or instead minimises exposure to offensive content or balances multiple values against social welfare.
- These questions distinguish three failure types: (a) the firm has malevolent goals and the tool follows them (a dog trained to bite); (b) firm and regulator goals align but the tool’s objective does not (a poorly trained dog); (c) the tool defies both (a dog refusing its handler).
- Past performance matters: a dog that has acted aggressively toward children needs a stronger leash. If a training set, architecture, or configuration is known to regurgitate or leak sensitive information, stronger leashes follow — more frequent monitoring, greater disclosure of testing results, or periodic regulatory approval of the management plan and its operation.
What is the role of the human at the end of the leash?
- Related research on explainability, interpretability, and trustworthiness aims to surface models’ internal operations in human-interpretable form. Early firm-level efforts include model cards (also Datasheets, FactSheets). Disclosure has so far been voluntary; open policy questions are whether to mandate them, in what form, and whether results go only to regulators or to the public. Experts have questioned whether current disclosures assure that companies are responsibly managing risk.
- Human-guided training injects task-specific expert knowledge during training. Example: a chest X-ray model that classifies using corner timestamps rather than thoracic features “has effectively learned the task” but in a way misaligned with radiologists’ diagnostic methods — supplemental expert input such as eye-tracking patterns can steer it toward expert-salient features. Costs are real, efficacy is not guaranteed, and the required number of annotations is unclear, so management-based regulation “would not dictate that firms deploy these models as much as consider and explain whether and how they should be deployed.” Most promising where the AI affects safety or rights, where human expertise offers distinctive insight, where outputs must align with experts, or where training samples are scarce.
- Where human-guided training is impractical, human-AI teaming — with continually developing alerts, descriptions, and visualisations — becomes a candidate element of a risk management plan, requiring continual personnel training analogous to retraining commercial pilots on updated cockpit technology.
- Regulators must also oversee the overseers. Without ongoing oversight, compliance is “prone to slippage or practices that simply ‘go through the motions,’” producing plans in pro forma fashion. Auditing by regulators or third parties is therefore necessary, and the information made available to auditors must be specified — “algorithms cannot be successfully audited or examined if no information about the algorithm’s training or its operation is disclosed to a human.”
Conclusion
- Leashes are flexible, permitting AI tools to explore new domains without regulatory barriers in the way, but that flexibility comes paired with active human oversight.
- “The goal of AI risk management should not be simply to establish guardrails and let AI tools operate within a fixed space unsupervised—at least until significant harms occur.”
- The aim is regulation that keeps human overseers “at the other end of a leash, ready and capable of steering AI away from danger as needed.”