Abstract

Model evaluations are increasingly used to inform high-stakes decisions about AI systems, but the field lacks the methodological foundations — standardised procedures, statistical confidence, reproducibility — needed to justify that trust. A dedicated “Science of Evals” is required before evaluation results can responsibly underpin regulation.

The problems with current practice

Sensitivity to prompt variation

  • Evaluation results are extremely brittle: subtle changes to prompt formatting in few-shot settings have produced performance swings of up to 76 percentage points.
  • Even trivial modifications — changing (A) to (1), or adjusting spacing — shift accuracy by roughly 5 points.

Capability elicitation uncertainty

  • The field cannot reliably distinguish maximal from average capability.
  • Prompting techniques (chain-of-thought, tree of thought, uncertainty-routed methods) keep setting new benchmark records, so a negative result may reflect an incomplete elicitation strategy rather than an absent capability.

Lack of scientific rigour

  • Evaluations lack standardised procedures, statistical confidence estimates, and reproducibility guarantees.
  • This matters most where evals are being used for safety-critical and regulatory purposes.

Proposed maturation framework

  • Nascent — exploratory research without agreed practices; the current state of the field.

  • Maturation — informal consensus on norms among stakeholders.

  • Mature — formal standards that can be applied in regulation.

  • A mature field should be able to answer questions about measurement precision, coverage breadth, robustness of results, replicability, statistical guarantees, and predictive accuracy for future systems.

Open research questions

  • Measurement validity — how to establish conceptual clarity about what property is being measured, and whether definitional shortcomings can be identified systematically through adversarial testing.
  • Trustworthiness of results — how to ensure comprehensive coverage of the relevant distribution, how to quantify biases and information leakage, and which statistical significance procedures apply to evaluations.
  • Dedicated research funding through bodies such as NSF, NIST and AI safety institutes.
  • Collaboration between academics, industry, policymakers and auditors.
  • Adoption of best practices from mature scientific disciplines — hypothesis testing, significance thresholds.
  • Learning from fields such as aviation that underwent comparable standardisation processes.
  • The piece is framed as a call for coordinated field-building rather than a technical solutions document.