About NIST
The National Institute of Standards and Technology (NIST) is a non-regulatory federal agency within the U.S. Department of Commerce. Within NIST, the Center for AI Standards and Innovation (CAISI) serves as the primary point of contact within the U.S. government for industry, facilitating testing and collaborative research to help harness and secure the potential of commercial AI systems.
The challenge
The AI community uses benchmarks to evaluate and compare the performance of Large Language Models (LLMs) and routinely cites benchmark results in leaderboards, model and system cards, and press releases. These results increasingly shape how procurers, deployers, and regulators make high-impact decisions about which AI systems to trust and adopt.
Scrutiny into benchmark methods and meaning has increased alongside reliance on benchmark scores. The White House’s “America’s AI Action Plan,” published in July 2025, included building an AI evaluations ecosystem as a key policy goal, including: “Support the development of the science of measuring and evaluating AI models, led by NIST at DOC, DOE, NSF, and other Federal science agencies.”
The approach
NIST researchers, including a U.S. Digital Corps Data Science Fellow, used established statistical modeling techniques to improve AI benchmarking practices.
They found that many existing approaches to computing benchmark metrics produced invalid uncertainty estimates or relied on unrecognized assumptions about the evaluation setting. For example, a benchmark score reported as “x% accuracy, ±1%” might convey a false sense of precision.
This created a measurement gap between a benchmark score and how well the score generalized to the broader class of problems the benchmark meant to represent. This distinction matters for anyone making real-world decisions based on benchmark claims.
Their work was published in NIST AI 800-3, which advances the field in two key ways:
- Formally distinguishing measurements between two related-but-different quantities:
- Benchmark accuracy: performance conditioned on a fixed benchmark
- Generalized accuracy: performance on all potential test items similar to those included in the benchmark
- Demonstrating a more rigorous method for estimating generalized accuracy. In a simulated setting and with large-scale evaluation of 22 API-access frontier LLMs across three popular benchmarks, the team showed that compared to existing regression-free approaches, generalized linear mixed models can more accurately estimate generalized accuracy and more efficiently quantify uncertainty. The model gives a more accurate picture by separating and and measuring two hidden factors driving benchmark results:
- Item-level difficulty: How hard a specific question or task is.
- Model-level performance: How good a specific AI or test taker is overall.
By statistically separating these two variables, NIST AI 800-3 prevents “benchmark hacking” and data contamination. Evaluators can accurately see if a model is:
- Genuinely improving in generalized capability, or
- Benefiting from a favorable distribution of easier benchmark items
The impact
The statistical advances introduced in NIST AI 800-3 empower researchers, journalists, and policymakers with important context for interpreting LLM evaluation results, including:
- Variance decomposition: better understanding of causes of benchmark score changes
- Item difficulty estimates
Together, these tools illuminate important aspects of LLM performance and benchmark construction. The findings make a compelling case for the AI evaluation community to explicitly specify a statistical model when they report benchmark results, rather than relying on ad hoc calculations. This shift toward rigorous, transparent measurement practices helps ensure that as benchmark results increasingly inform consequential decisions about AI development, procurement, and regulation, those decisions rest on a statistically sound foundation.
digitalcorps.gsa.gov
An official website of GSA’s Technology Transformation Services