Legal Considerations for Defining “Frontier Model”

Abstract

Almost every proposed AI law distinguishes the most advanced models from everything else, so each will need a workable definition of “frontier model” or an analogous term. Getting that definition right implicates the choice between statutory and regulatory definitions, the selection of definitional elements (technical inputs, capabilities, risk, epistemic status, deployment context), the legal obstacles to rapid updating, and recent administrative law developments — the end of Chevron deference, the nondelegation doctrine, and the major questions doctrine — that constrain how much discretion Congress can confer on agencies.

Publication details

  • LawAI Working Paper Series, No. 2-2024, September 2024, Institute for Law and AI (Christoph Winter also affiliated with ITAM and Harvard).
  • Keywords: law & artificial intelligence; legal definitions; frontier models; regulatory updating.
  • The paper is deliberately agnostic on how frontier models should be regulated — only on how the category should be defined.

Why the definition matters

  • Industry and government have largely converged on recognising a distinct category of most-advanced systems, variously called “dual-use foundation models” (EO 14110), “general-purpose AI models with systemic risk” (EU AI Act), and “frontier models” (researchers, labs, some legislators). The phrases are not synonymous but address the same issue.
  • Frontier models are expected to be broadly capable and to have applications “not readily predictable prior to development, nor even immediately known or knowable after development.”
  • A rule applying to a category cannot be enforced or complied with unless membership is determinable. Undefined ambiguous technical terms impose costs on firms (over- or unintentional under-compliance) and the public (weaker compliance, higher enforcement costs, less risk protection, more litigation over scope).
  • Cited historical failures: the Audio Home Recording Act of 1992, obsolete within a few years because its definition of “digital musical recording” excluded files on computer hard drives; and the lack of a statutory definition of “broadband” in the Telecommunications Act of 1996.

Statutory vs. regulatory definitions

  • Regulatory definitions (agency rules) can be more numerous and detailed, updated faster, and benefit from deeper agency subject-matter expertise. Federal agencies publish roughly 3,500–4,000 final rules a year; Congress passed 27 bills into law in 2023.
  • Statutory definitions purchase greater democratic legitimacy and legal resiliency with their procedural and political cost. Several challenges available against regulations — the major questions doctrine, APA defects — are unavailable against statutes. A regulation generally cannot override a statutory definition, only clarify or interpret it.
  • The hybrid pattern is common: a threshold statutory definition plus a more detailed, updatable regulatory one. Examples given: the FMLA’s “serious health condition” and FIFRA’s “pest.”

Existing definitions

Executive Order 14110 (October 2023)

  • Broad definition: a model trained on broad data, generally using self-supervision, containing at least tens of billions of parameters, applicable across a wide range of contexts, exhibiting (or easily modified to exhibit) high performance at tasks posing serious risk to security, national economic security, or public health and safety — such as substantially lowering barriers to CBRN weapons design, enabling powerful offensive cyber operations through automated vulnerability discovery and exploitation, or permitting evasion of human control through deception or obfuscation. Technical safeguards against misuse do not remove a model from the definition.
  • Placeholder definition pending Commerce’s regulatory definition: models trained on >1026 integer or floating-point operations, or >1023 FLOP if trained primarily on biological sequence data.
  • The EO thus pairs a high-level “quasi-statutory” definition with a directive to an agency to promulgate and regularly update a detailed regulatory one. The first definition relies on subjective evaluations (no objective test for “serious” risk, “broad data,” or “high levels of performance”); the placeholder is purely objective.

California SB 1047 (vetoed)

  • May 2024 Senate version defined “covered model” as either (1) trained on >1026 FLOP, or (2) trained on compute sufficient to be reasonably expected to match or exceed the performance of a model trained on >1026 FLOP in 2024, assessed on common benchmarks. The second prong was a future-proofing device against compute thresholds becoming underinclusive as algorithmic efficiency improves.
  • Final version, after amendments responding to innovation-stifling objections, replaced the capability prong with a $100,000,000 training cost floor (plus a fine-tuning prong at ≥3×1025 FLOP and $10,000,000), and provided for the Government Operations Agency to set new thresholds from January 2027.
  • Key analytical point: the capability threshold was disjunctive and therefore expanded coverage, guarding against underinclusiveness; the cost threshold was conjunctive and therefore restricted coverage, guarding against overinclusiveness as compute prices fall. The cost figure was baked into the statute and changeable only by new legislation.
  • Compared with EO 14110, SB 1047’s scheme was “simpler, easier to operationalize, and less flexible” — no broad risk-based definition, and the regulator could only adjust numerical compute values within an otherwise rigid statutory definition.

EU AI Act

  • “General-purpose AI model with systemic risk” covers models with “high impact capabilities” assessed via appropriate technical tools, indicators and benchmarks, or by Commission decision (ex officio or on a qualified alert from the scientific panel) that a model has equivalent capabilities or impact under the Annex XIII criteria.
  • Models are presumed to have high impact capabilities if trained on >1025 FLOP — an order of magnitude below the EO and SB 1047 thresholds. The paper flags this as potentially politically significant: GPT-4 and Gemini are thought to fall within it, while Mistral’s latest model is thought not to.
  • The Annex XIII criteria span parameter count, dataset size and quality (e.g. measured in tokens), benchmark performance, and user numbers. The Commission may amend the threshold and supplement benchmarks in response to algorithmic improvements or hardware efficiency gains.
  • The EU threshold is sufficient but not necessary, making the definition considerably broader than the EO’s, where the placeholder threshold is both necessary and sufficient.

Candidate definitional elements

Technical inputs and characteristics

  • Training compute is the most attractive option: quantifiable, measurable, monitorable, verifiable, hard to manipulate, closely correlated with capability, and estimable before the training run, so developers know in advance whether they are covered.
  • Its weakness is layered proxying — compute proxies capability, which proxies risk — which makes definitions “particularly prone to becoming untethered from their original purpose,” aggravated by Goodhart’s Law. The main failure mode is underinclusiveness over time as algorithmic efficiency improves.
  • Parameter count and dataset size share compute’s pros and cons and can serve as partially redundant backup metrics; a definition using both compute and dataset size would correctly exclude a model trained on huge compute but a tiny dataset. Dataset type or quality is less quantifiable but captures what the numbers cannot — hence the EO’s lower threshold for biological sequence data.

Capabilities

  • Can be specified objectively (benchmark scores, better suited to a frequently updated regulatory definition) or in general terms left to future interpreters (typical of a high-level statutory definition).
  • Advantage: eliminates the risk that a proxy decays — more robust to algorithmic efficiency gains. This was the purpose of SB 1047’s May 2024 capability prong.
  • Disadvantages: capabilities are much harder to measure than compute; benchmarks that accurately capture the diverse capabilities of general-purpose foundation models are notoriously difficult to build; and capabilities are generally not measurable until after training, which complicates regulating development (though pre-release regulation remains possible).

Risk

  • Defining directly by risk could in theory allow better-targeted rules — excluding highly capable but demonstrably low-risk models.
  • But measuring risk is harder than measuring capabilities, and the science of rigorous safety evaluations for foundation models “is still in its infancy.”
  • Only EO 14110 among the three measures mentions risk directly, and it does so by specifying the type of risk (tasks posing serious risk to security, economic security, public health and safety) rather than the severity — a capability threshold combined with a risk threshold.

Epistemic elements

  • A distinction between “known” models (well understood, posing only known risks) and “unknown” models (poorly understood, potentially unpredictable risks) matches the intuition behind the word “frontier”: models that “push into the unknown.”
  • Precedent: the Toxic Substances Control Act, under which any chemical not on the EPA’s regularly updated inventory is “new” by definition and requires a licence.
  • Advantages: separates “unknown unknowns” from better-understood risks, and the category shrinks automatically as regulators become familiar with a model’s capabilities and risks — models drop out of “frontier” status over time.
  • Disadvantage: hard to operationalise; it would require either a proxy for unknown capabilities or an authorised regulatory categorisation process.

Deployment context

  • The EU AI Act counts registered end users and EU business users among its Annex XIII criteria.
  • Rarely sufficient alone, but a useful proxy for the kind of risk: some harms scale with user numbers, and a model deployed only to government agencies or the military presents a different risk profile than one released to the public.

Updating regulatory definitions

  • Emerging-technology rules become obsolete quickly if they cannot be updated, which favours delegating the definition to an agency.
  • Cautionary case: 1990s–2000s U.S. export controls defined “supercomputer” by MTOPS (millions of theoretical operations per second). Rapid processor improvements forced the Clinton administration to revise the threshold repeatedly to protect U.S. industry competitiveness; eventually the metric itself became obsolete, leaving several years in which the controls were “ineffective at best.”
  • Notice and comment under the APA — publication, comment period, agency response, final rule, then a 30–60 day delay — can take months to years. Agencies may waive the delay or the whole process for “good cause” where standard procedure would be “impracticable, unnecessary, or contrary to the public interest.” BIS waived the 30-day wait for its October 2023 interim rule restricting advanced AI chip sales to China.
  • The paper notes notice and comment has real benefits — substantive input, democratic accountability, transparency — that must be weighed against delay costs. Congress can also statutorily limit or waive it, or impose rulemaking deadlines.
  • OIRA review of economically significant rules similarly improves quality and interagency coordination while typically adding several months; it can be waived by statute or by OIRA itself.

Administrative law constraints

Loper Bright and the end of Chevron

  • Loper Bright Enterprises v. Raimondo / Relentless v. Department of Commerce repealed Chevron deference. Agency interpretations now prevail only under Skidmore — to the extent courts find them persuasive.
  • Kagan’s dissent warned that courts will now “play a commanding role” in questions like “what rules are going to constrain the development of A.I.?” The authors think this “probably somewhat overstates the significance… for rhetorical effect”: where Congress explicitly directs an agency to define a term — as EO 14110 does for Commerce — Loper Bright poses no obstacle. The real uncertainty concerns implied delegations.
  • Worked example: under EPCA, the DOE issued a regulatory definition of “small electric motor” as 0.25–3 horsepower. NEMA sued; a 2011 court found the statute ambiguous, deemed the DOE’s reading reasonable, and upheld it under Chevron. EPCA authorises the DOE to set testing requirements and efficiency standards but does not explicitly authorise a definition — so the same rule would be markedly more vulnerable today.
  • Two knock-on effects: challenges to impliedly authorised definitions become more likely to succeed, and litigation-averse agencies may regulate more cautiously to avoid suit.
  • Drafting implication: Congress can insulate agency definitions with clear, explicit authorising language — but it is hard to predict in advance how a statutory definition will become ambiguous. A narrow authorisation (e.g. to update only a compute threshold) may prove insufficiently flexible if the relevant technological change requires a different factor entirely; a very broad authorisation raises democratic accountability concerns and exposes the scheme to the next two doctrines.

Nondelegation doctrine

  • Currently toothless in this context: a delegation is valid so long as the statute supplies an “intelligible principle,” a standard satisfied even by guidance as vague as regulating in a way that “will be generally fair and equitable.” The Supreme Court has struck statutes on this ground only twice, both in 1935.
  • Commentators expect the Court may revisit it, possibly adopting something like Gorsuch’s Gundy dissent, which would require Congress to make “all the relevant policy decisions” and leave agencies only to “fill up the details.”
  • If strengthened, a statute authorising a regulatory definition of “frontier model” might need meaningful guidance on what the definition should look like — the more so because acceptable agency discretion “varies according to the power congressionally conferred”: no direction is needed for minor technical terms, but “substantial guidance” is needed for tasks that could significantly affect the national economy.

Major questions doctrine

  • Named for the first time in West Virginia v. EPA (2022): courts will not accept a statutory interpretation granting an agency authority over a matter of great “economic or political significance” absent “clear congressional authorization.” Unlike nondelegation, it affects interpretation rather than constitutionality.
  • Critics argue it inhibits Congress’s ability to confer broad discretion over problems that are hard to foresee — Kagan’s West Virginia dissent noted the Clean Air Act was broadly worded precisely because changing circumstances and scientific developments would otherwise render it obsolete.
  • Applied hypothetical: a federal licensing statute empowers BIS to define and regularly update the technical conditions for “frontier model.” BIS begins with a compute threshold; ten years later a new architecture achieves cutting-edge capability on relatively little compute, and BIS tries to add a capabilities threshold. A deregulatory court might find the broad original authorisation insufficiently clear to support an expanded licensing regime based on less objective criteria.

Conclusion

  • Nonlawyers commonly assume statutory words carry their ordinary English meaning, missing that legal rules operate as “a sort of simple code” in which terms stand in for definitions catalogued elsewhere — an oversight mirrored in the general neglect of how much a regulatory scheme depends on well-crafted definitions.
  • Four considerations are consolidated: the respective roles of statutory and regulatory definitions, which can be combined for technical soundness plus democratic legitimacy; the selection and combination of elements (technical inputs, capability metrics, risk, deployment context, familiarity); legal mechanisms for rapid updating; and the effect of nondelegation, major questions, and the end of Chevron on how much discretion can be conferred.