AI Evaluation Should Require Standardized Item-Level Data Releases
Hang JiangSusu ZhangDongyao ZhuYuzhuo BaiSang TruongXiaoyuan YiSanmi KoyejoXing XieZiang Xiao
Proposes standardizing item-level model response releases as essential AI evaluation infrastructure, introducing the ten-million-response OpenEval repository to expose benchmark flaws, diagnose construct misalignment, and prevent inflated model capability claims.
Artificial intelligence systems are increasingly deployed in high-stakes environments, yet the benchmarks used to measure their capabilities and guide governance rely heavily on aggregate scores. This standard approach obscures critical flaws, including benchmark saturation, outdated content, test-train data contamination, and construct misalignment—where an evaluation fails to measure the actual ability it claims to test. Because individual model responses are routinely discarded after summary scores are calculated, decision-makers face inflated capability claims and unwarranted trust in deployed systems without the empirical means to audit results.
The article demonstrates that standardizing and openly releasing item-level evaluation data is essential to establish a rigorous, transparent science of AI evaluation. It introduces an open, unified repository and applies established measurement science techniques to prove that item-level granularity is necessary to diagnose benchmark health, evaluate test validity, and understand true system capabilities.
To establish feasibility and test this framework, the article constructs OpenEval, a centralized archive containing 10 million responses across 155,000 items from widely used AI benchmarks, with an average of roughly 70 models evaluated per dataset. The data is structured under a unified, multi-tiered schema that separates original benchmark inputs, test-specific prompt adaptations, model execution parameters, generated responses, and metric scores. Using this archive, the article applies psychometric methodologies, including Classical Test Theory to assess item difficulty and discrimination, alongside Item Factor Analysis to evaluate the underlying capability dimensions measured by benchmarks.
The empirical analysis yielded several critical findings. First, item characteristic analysis revealed that advanced benchmarks suffer from rapid saturation; a substantial proportion of items in MMLU-Pro exhibit near-zero difficulty for modern models. Second, while MMLU-Pro reduced noisy items compared to its predecessor, it still contains items with negative discrimination, meaning higher-performing models were paradoxically more likely to get them wrong due to potential errors, ambiguity, or misleading cues. Third, item factor analysis of the BABIQA deductive reasoning benchmark showed that model responses clustered based on the specific animal named in the answer rather than true reasoning ability, exposing clear construct misalignment. Finally, factor analysis on MMLU-Pro confirmed that differences in model performance are driven by distinct high-level reasoning skills—such as formal quantitative modeling versus domain recall—rather than subject-matter categories.
These findings indicate that aggregate benchmark scores can mislead procurement, risk management, and regulatory compliance by masking severe performance flaws and shortcut learning. Evaluating individual item responses allows researchers, regulators, and enterprise users to isolate contaminated or uninformative questions, update benchmarks efficiently without full redesigns, and verify that scores reflect genuine operational readiness. The article addresses common counterarguments, noting that withholding data does not stop data contamination but only prevents its detection, and that standardized submission tools minimize the reporting burden on researchers.
To translate these findings into practice, the AI community and standard-setting bodies should mandate standardized item-level data releases as default infrastructure for all evaluation reporting. Evaluators should adopt shared schemas and conversion tools to support cumulative, auditable research. In addition, benchmark maintainers should routinely conduct item-level audits to prune saturated and defective questions.
The current analysis relies primarily on Classical Test Theory and linear factor analysis, leaving more advanced psychometric models for future study. Furthermore, while the analytical benefits of data sharing are clear, technical and governance safeguards—such as access-controlled repositories and staged embargo periods—require formal operational development to balance open auditability with proprietary protections and contamination prevention.
- Paper: General Scales Unlock AI Evaluation with Explanatory and Predictive Power, Lexin Zhou et al. (2025). Its item-level demand scales show how benchmark responses can reveal what tasks measure, clarifying the source’s case for validity evidence beyond aggregate scores.
- Paper: Understanding Dataset Difficulty with V-Usable Information, Kawin Ethayarajh et al. (2022). Its pointwise V-information framework demonstrates how individual benchmark items expose difficulty, label errors, and artifacts that aggregate performance conceals.
- Paper: Are We Done with MMLU?, Aryo Pradipta Gema et al. (2025). Its expert audit of MMLU items shows how question-level defects can distort model rankings, grounding the source’s argument for item-level evaluation evidence.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). CheckList’s behavioral tests make concrete why benchmark-wide accuracy can miss specific failures, a key premise behind inspecting item-level responses.
No sufficiently relevant recommendations were found.
