Abstract
Model evaluations are increasingly used to inform high-stakes decisions about AI systems, but the field lacks the methodological foundations — standardised procedures, statistical confidence, reproducibility — needed to justify that trust. A dedicated “Science of Evals” is required before evaluation results can responsibly underpin regulation.
The problems with current practice
Sensitivity to prompt variation
- Evaluation results are extremely brittle: subtle changes to prompt formatting in few-shot settings have produced performance swings of up to 76 percentage points.
- Even trivial modifications — changing
(A)to(1), or adjusting spacing — shift accuracy by roughly 5 points.
Capability elicitation uncertainty
- The field cannot reliably distinguish maximal from average capability.
- Prompting techniques (chain-of-thought, tree of thought, uncertainty-routed methods) keep setting new benchmark records, so a negative result may reflect an incomplete elicitation strategy rather than an absent capability.
Lack of scientific rigour
- Evaluations lack standardised procedures, statistical confidence estimates, and reproducibility guarantees.
- This matters most where evals are being used for safety-critical and regulatory purposes.
Proposed maturation framework
-
Nascent — exploratory research without agreed practices; the current state of the field.
-
Maturation — informal consensus on norms among stakeholders.
-
Mature — formal standards that can be applied in regulation.
-
A mature field should be able to answer questions about measurement precision, coverage breadth, robustness of results, replicability, statistical guarantees, and predictive accuracy for future systems.
Open research questions
- Measurement validity — how to establish conceptual clarity about what property is being measured, and whether definitional shortcomings can be identified systematically through adversarial testing.
- Trustworthiness of results — how to ensure comprehensive coverage of the relevant distribution, how to quantify biases and information leakage, and which statistical significance procedures apply to evaluations.
Recommended actions
- Dedicated research funding through bodies such as NSF, NIST and AI safety institutes.
- Collaboration between academics, industry, policymakers and auditors.
- Adoption of best practices from mature scientific disciplines — hypothesis testing, significance thresholds.
- Learning from fields such as aviation that underwent comparable standardisation processes.
- The piece is framed as a call for coordinated field-building rather than a technical solutions document.