AI Evaluation Should Require Standardized Item-Level Data Releases

Hang JiangSusu ZhangDongyao ZhuYuzhuo BaiSang TruongXiaoyuan YiSanmi KoyejoXing XieZiang Xiao

article2026arXiv1 citations

Proposes standardizing item-level model response releases as essential AI evaluation infrastructure, introducing the ten-million-response OpenEval repository to expose benchmark flaws, diagnose construct misalignment, and prevent inflated model capability claims.

Listen

Artificial intelligence systems are increasingly deployed in high-stakes environments, yet the benchmarks used to measure their capabilities and guide governance rely heavily on aggregate scores. This standard approach obscures critical flaws, including benchmark saturation, outdated content, test-train data contamination, and construct misalignment—where an evaluation fails to measure the actual ability it claims to test. Because individual model responses are routinely discarded after summary scores are calculated, decision-makers face inflated capability claims and unwarranted trust in deployed systems without the empirical means to audit results.

The article demonstrates that standardizing and openly releasing item-level evaluation data is essential to establish a rigorous, transparent science of AI evaluation. It introduces an open, unified repository and applies established measurement science techniques to prove that item-level granularity is necessary to diagnose benchmark health, evaluate test validity, and understand true system capabilities.

To establish feasibility and test this framework, the article constructs OpenEval, a centralized archive containing 10 million responses across 155,000 items from widely used AI benchmarks, with an average of roughly 70 models evaluated per dataset. The data is structured under a unified, multi-tiered schema that separates original benchmark inputs, test-specific prompt adaptations, model execution parameters, generated responses, and metric scores. Using this archive, the article applies psychometric methodologies, including Classical Test Theory to assess item difficulty and discrimination, alongside Item Factor Analysis to evaluate the underlying capability dimensions measured by benchmarks.

The empirical analysis yielded several critical findings. First, item characteristic analysis revealed that advanced benchmarks suffer from rapid saturation; a substantial proportion of items in MMLU-Pro exhibit near-zero difficulty for modern models. Second, while MMLU-Pro reduced noisy items compared to its predecessor, it still contains items with negative discrimination, meaning higher-performing models were paradoxically more likely to get them wrong due to potential errors, ambiguity, or misleading cues. Third, item factor analysis of the BABIQA deductive reasoning benchmark showed that model responses clustered based on the specific animal named in the answer rather than true reasoning ability, exposing clear construct misalignment. Finally, factor analysis on MMLU-Pro confirmed that differences in model performance are driven by distinct high-level reasoning skills—such as formal quantitative modeling versus domain recall—rather than subject-matter categories.

These findings indicate that aggregate benchmark scores can mislead procurement, risk management, and regulatory compliance by masking severe performance flaws and shortcut learning. Evaluating individual item responses allows researchers, regulators, and enterprise users to isolate contaminated or uninformative questions, update benchmarks efficiently without full redesigns, and verify that scores reflect genuine operational readiness. The article addresses common counterarguments, noting that withholding data does not stop data contamination but only prevents its detection, and that standardized submission tools minimize the reporting burden on researchers.

To translate these findings into practice, the AI community and standard-setting bodies should mandate standardized item-level data releases as default infrastructure for all evaluation reporting. Evaluators should adopt shared schemas and conversion tools to support cumulative, auditable research. In addition, benchmark maintainers should routinely conduct item-level audits to prune saturated and defective questions.

The current analysis relies primarily on Classical Test Theory and linear factor analysis, leaving more advanced psychometric models for future study. Furthermore, while the analytical benefits of data sharing are clear, technical and governance safeguards—such as access-controlled repositories and staged embargo periods—require formal operational development to balance open auditability with proprietary protections and contamination prevention.

arXiv: 2604.03244

No sufficiently relevant recommendations were found.

Cover for AI Evaluation Should Require Standardized Item-Level Data Releases

Abstract

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified item selection, construct misalignment, and poor generalization. The root cause of these failures is a misplaced focus on aggregate model scores. Without item-level evidence, validity claims cannot be assessed, resulting in inflated capability claims, misdirected research, and unwarranted trust in deployed systems. Our position is that designing valid evaluations requires empirical evidence from item-level model responses, and the standardized release of such data should be treated as core AI evaluation infrastructure. Such a release, in addition, enables transparency, replicability, and auditability of evaluation results. To show the norm is both feasible and consequential, we construct OpenEval, an item-level archive of 10M responses across 155k items from widely-used benchmarks, under a unified schema that the AI evaluation community can develop upon. We demonstrate how item-level data can identify low-quality items, document construct misalignment, and recover validity evidence about benchmarks' internal structure. We address objections around contamination and author burden, and show each is tractable relative to the cost of decisions made on claims that cannot be trusted.

Table of Contents

  • 1 Introduction
  • 2 Validity Challenges in AI Benchmarking
  • 2.1 Methodological Issues in Benchmark Design
  • 2.2 AI Advances Pressures on Benchmarking Practices
  • 3 Item-Level Data: The Missing Foundation for AI Evaluation Science
  • 4 OpenEval: An Item-Centered Benchmark Repository
  • 5 Empirical Illustrations of Item-Level Analysis
  • 5.1 Background and Analytical Approaches
  • 5.2 Examining Item Quality via Classical Test Theory
  • 5.3 Revealing Construct Alignment via Item Factor Analyses
  • 6 Broader Implications beyond Benchmarking
  • 7 Alternative Views
  • References
  • A Limitations
  • B Additional Analysis Results
  • C OpenEval Schema

Knowls

  1. Knowl 1 — Standardized item-level data is proposed as core AI evaluation infrastructure

    model/method

    The paper argues that benchmark results should ordinarily be released with standardized item-level data whenever legal and ethical constraints allow. Such data includes the test conditions, item content, model responses, and per-response scores. Aggregate scores alone cannot reveal whether a benchmark measures its intended capability, whether nuisance factors drive performance, or whether items distinguish models effectively. The paper therefore treats item-level evidence as necessary infrastructure for assessing benchmark validity and reliability, maintaining benchmarks, and making evaluation findings reproducible and auditable.

  2. Knowl 2 — OPENEVAL organizes evaluation records around benchmark items

    model/method

    OPENEVAL is a community-oriented repository schema in which benchmark records contain items, and each item contains one or more model responses. A benchmark records provenance such as its name, version, URLs, and tags; an item records its original content and references for scoring; and each response records the model, the input actually given to it, its output, and one or more metric scores. Response context can include demonstrations, external resources, system instructions, generation parameters, and tools. The schema distinguishes original item content from adapted model input so that different evaluation conditions remain reproducible. It uses hierarchical identifiers for benchmarks, items, and responses, while allowing heterogeneous content in flexible fields to accommodate varied task formats and evaluation settings.

  3. Knowl 3 — OPENEVAL demonstrates item-level archiving at scale

    data/table

    At the time of the paper, OPENEVAL contained more than 155,000 items and 10 million item-level responses from widely used benchmarks. The number of evaluated models per dataset ranged from 11 to 111, with a reported average of 70.3; each response has one or more metric scores. The archive combines results collected to supplement existing repositories, interdisciplinary datasets concerning social aspects of AI, and conversions of existing item-level resources into the shared schema. The authors also provide converters for popular evaluation harnesses, presenting the archive as evidence that standardized item-level release is practicable.

  4. Knowl 4 — Classical test theory quantifies item difficulty and discrimination

    equation

    For a benchmark with nn evaluated models and numeric item scores, let xijx_{ij} be model jj’s score on item ii, and let mim_i be the maximum possible score for item ii. The classical test theory difficulty index is the mean fraction of the maximum score earned across models; a larger value indicates an easier item: pi=1n∑j=1nxijmip_i = \frac{1}{n}\sum_{j=1}^{n}\frac{x_{ij}}{m_i}. Let T−i,j=∑k≠ixkjT_{-i,j}=\sum_{k\ne i}x_{kj} be model jj’s total score on all benchmark items other than item ii. Item discrimination is the Pearson correlation across evaluated models between scores on item ii and their rest-of-benchmark totals: ri=Corr⁡j(xij,T−i,j)r_i=\operatorname{Corr}_j(x_{ij},T_{-i,j}). A large positive rir_i indicates that the item tends to be answered correctly by models that perform well on the rest of the benchmark; near-zero or negative values flag items for scrutiny. The paper also illustrates item characteristic curves by sorting models into six equally sized groups according to their rest-of-benchmark scores and comparing mean item scores across groups.

  5. Knowl 5 — MMLU-PRO items are often easy for recent models but discriminate better than MMLU items

    empirical result

    The paper applies classical test theory to 567 MMLU items evaluated by 66 models released before November 2023, and to 1,000 MMLU-PRO items evaluated by 72 models released after June 2024. Many MMLU-PRO items have low difficulty indices under the recent-model sample, indicating that they are easy for those models and suggesting benchmark saturation. At the same time, MMLU-PRO has substantially fewer items with low or negative discrimination than MMLU, consistent with its design goal of reducing noise. Some MMLU-PRO items nevertheless discriminate poorly and warrant further examination for ambiguity, miskeying, or construct-irrelevant cues. The item-difficulty values from the two benchmarks are not directly comparable because they were computed on different model samples. In MMLU, item #496 has discrimination r496=0.8401r_{496}=0.8401 and an increasing item characteristic curve across model-performance groups, while items #55 and #374 have r55=−0.0663r_{55}=-0.0663 and r374=−0.4535r_{374}=-0.4535, respectively, and do not show the expected positive relationship.

  6. Knowl 6 — BABIQA item clusters expose answer-preference effects in a deduction benchmark

    empirical result

    For BABIQA Task 15, which is intended to assess basic deduction through inheritance of properties, the authors use singular-value-decomposition-based item factor analysis and cluster the 1,000 items by their loadings on the top three factors. The resulting three clusters are strongly associated with the items’ reference answers: the sheep-answer items all fall in cluster 2 (240 items); mouse-answer items fall in cluster 0 (225); cat-answer items fall mostly in cluster 0 (221), with 3 in cluster 1; and wolf-answer items fall mainly in cluster 1 (326), with 6 in cluster 0. This pattern raises a construct-validity concern: differences in model performance may partly reflect tendencies to choose particular animals rather than basic deductive reasoning. Generalized low-rank-model factor analysis yields consistent findings.

  7. Knowl 7 — MMLU-PRO performance differences align with four candidate reasoning dimensions

    empirical result

    The authors apply generalized low-rank-model item factor analysis to MMLU-PRO, retain four factors, and rotate them using varimax. They send the 100 items with the largest absolute loadings on each factor to GPT-5 for interpretation, then manually revise the candidate labels. The resulting dimensions are: (1) formal, quantitative, multi-step modeling; (2) domain-specific recall and simple reasoning; (3) conceptual understanding and explanation; and (4) applied synthesis and case-based judgment. These exploratory results suggest that differences among models are organized more by higher-level reasoning demands than by subject domain alone: items from the same subject can load differently on the four factors. The finding is consistent with MMLU-PRO’s stated aim of increasing reasoning demands relative to MMLU, but the factor labels are interpretations rather than definitive construct definitions.

  8. Knowl 8 — External benchmark relationships provide tentative validity evidence for MMLU-PRO factors

    empirical result

    To examine whether the four candidate MMLU-PRO factor interpretations have plausible external relationships, the authors calculate each factor subscore as a model’s mean score on the 100 items with the largest absolute loadings for that factor, then compare those subscores with model scores on GPQA and OMNI-MATH. GPQA covers graduate-level biology, physics, and chemistry; OMNI-MATH covers Olympiad-level mathematics. The authors hypothesize that Factor 1, formal quantitative multi-step modeling, should align with both external benchmarks, while Factor 4, applied synthesis and case-based judgment, should align more with GPQA. The observed relationships are broadly consistent with these expectations. Factors 2 and 3 show weak correlations with both external benchmarks, offering discriminant-validity evidence. The paper presents these relationships as descriptive and tentative, not conclusive validation of the factor interpretations.

  9. Knowl 9 — The paper argues that contamination and release burdens do not outweigh item-level transparency

    model/method

    The paper’s position is that releasing item-level benchmark data can make contamination easier to detect and affected items easier to remove, whereas withholding the data does not prevent contamination and makes it harder to investigate. It notes that release can be combined with hold-out designs, verified access, or staged publication after an embargo, including to protect licensed content or limit strategic misuse. The authors also argue that evaluation already generates the data, so item-level release requires no additional compute; the remaining work is conversion and upload. A shared schema and converters are presented as ways to reduce that work, although the authors acknowledge that some contributor burden remains.

  10. Knowl 10 — The empirical demonstrations and data-release safeguards remain limited in scope

    limitation

    The paper’s empirical illustrations rely primarily on classical test theory and item factor analysis. They do not explore item response theory, differential item functioning, or diagnostic classification models, which could demonstrate additional analyses supported by item-level data. The discussion of mitigating risks from open item-level releases is also conceptual: concrete mechanisms such as access controls, usage monitoring, and community-agreed responsible-use norms are not developed in detail.

Coverage note — The broader proposed benefits for machine-learning research, policy, and participatory evaluation are omitted because they are implications rather than separately demonstrated contributions. Supplementary MMLU and HELM analyses are also omitted because the main empirical illustrations capture the paper’s central item-quality and construct-alignment findings.

References

  1. 1.M. Akhtar, A. Reuel, P. Soni, S. Ahuja, P. S. Ammanamanchi, R. Rawal, V. Zouhar, S. Yadav, C. Whitehouse, D. Ki, et al. When ai benchmarks plateau: A systematic study of benchmark saturation. arXiv preprint arXiv:2602.16763, 2026.
  2. 2.A. F. Akyürek, M. Y. Kocyigit, S. Paik, and D. T. Wijaya. Challenges in measuring bias via open-ended language generation. In C. Hardmeier, C. Basta, M. R. Costa-jussà, G. Stanovsky, and H. Gonen, editors, Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 76–76, Seattle, Washington, July 2022. Association for Computational Linguistics.
  3. 3.American Educational Research Association, American Psychological Association, and National Council on Measurement in Education. Standards for Educational and Psychological Testing. American Educational Research Association, 2014.
  4. 4.M. Anderljung, J. Barnhart, A. Korinek, J. Leung, C. O’Keefe, J. Whittlestone, S. Avin, M. Brundage, J. Bullock, D. Cass-Beggs, B. Chang, T. Collins, T. Fist, G. Hadfield, A. Hayes, L. Ho, S. Hooker, E. Horvitz, N. Kolt, J. Schuett, Y. Shavit, D. Siddarth, R. Trager, and K. Wolf. Frontier ai regulation: Managing emerging risks to public safety, 2023.
  5. 5.T. W. Anderson, T. W. Anderson, T. W. Anderson, T. W. Anderson, and E.-U. Mathématicien. An introduction to multivariate statistical analysis, volume 2. Wiley New York, 1958.
  6. 6.A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, J. Batzner, N. Foroutan, C. Schmitz, K. Korgul, H. Batra, O. Deb, E. Beharry, C. Emde, T. Foster, A. Gausen, M. Grandury, S. Han, V. Hofmann, L. Ibrahim, H. Kim, H. R. Kirk, F. Lin, G. K.-M. Liu, L. Luettgau, J. Magomere, J. Rystrøm, A. Sotnikova, Y. Yang, Y. Zhao, A. Bibi, A. Bosselut, R. Clark, A. Cohan, J. N. Foerster, Y. Gal, S. A. Hale, I. D. Raji, C. Summerfield, P. Torr, C. Ududec, L. Rocher, and A. Mahdi. Measuring what matters: Construct validity in large language model benchmarks. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025.
  7. 7.Y. Bengio, J. Louradour, R. Collobert, and J. Weston. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning, ICML ’09, page 41–48, New York, NY, USA, 2009. Association for Computing Machinery.
  8. 8.S. Bhojanam and S. Mehta. Prompt genotyping: Quantifying the evaluation gap between synthetic benchmarks and real LLM performance. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  9. 9.S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach. Language (technology) is power: A critical survey of “bias” in NLP. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online, July 2020. Association for Computational Linguistics.
  10. 10.S. L. Blodgett, G. Lopez, A. Olteanu, R. Sim, and H. Wallach. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In C. Zong, F. Xia, W. Li, and R. Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1004–1015, Online, Aug. 2021. Association for Computational Linguistics.
  11. 11.S. R. Bowman and G. Dahl. What will it take to fix benchmarking in natural language understanding? In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4843–4855, Online, June 2021. Association for Computational Linguistics.
  12. 12.D. T. Campbell and D. W. Fiske. Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2):81–105, 1959.
  13. 13.W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. I. Jordan, J. E. Gonzalez, and I. Stoica. Chatbot arena: an open platform for evaluating llms by human preference. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  14. 14.Y. Y. Chiu, L. Jiang, B. Y. Lin, C. Y. Park, S. S. Li, S. Ravi, M. Bhatia, M. Antoniak, Y. Tsvetkov, V. Shwartz, and Y. Choi. CulturalBench: A robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through human-AI red-teaming. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25663–25701, Vienna, Austria, July 2025. Association for Computational Linguistics.
  15. 15.L. L. Cook and M. J. Pitoniak, editors. Educational Measurement. Oxford University Press, 5 edition, 2025.
  16. 16.I. Covert, W. Ji, T. Hashimoto, and J. Zou. Scaling laws for the value of individual data points in machine learning. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  17. 17.L. J. Cronbach and P. E. Meehl. Construct validity in psychological tests. Psychological Bulletin, 52(4):281–302, 1955. 60 references. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
  18. 18.M. Dehghani, Y. Tay, A. A. Gritsenko, Z. Zhao, N. Houlsby, F. Diaz, D. Metzler, and O. Vinyals. The benchmark lottery, 2021.
  19. 19.˙I. E. Deveci and D. Ataman. The ouroboros of benchmarking: Reasoning evaluation in an era of saturation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  20. 20.M. Du, V. Manjunatha, R. Jain, R. Deshpande, F. Dernoncourt, J. Gu, T. Sun, and X. Hu. Towards interpreting and mitigating shortcut learning behavior of NLU models. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 915–929, Online, June 2021. Association for Computational Linguistics.
  21. 21.M. Eriksson, E. Purificato, A. Noroozian, J. Vinagre, G. Chaslot, E. Gomez, and D. Fernandez-Llorca. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1):850–864, Oct. 2025.
  22. 22.European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L series, 2024.
  23. 23.FadillAmir. Benchmarking and standardization of evaluation protocols: A feedback-driven framework using LLM judges to gatekeep and iteratively improve synthetic benchmarks. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  24. 24.C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, and T. Wolf. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2024.
  25. 25.B. Gao, F. Song, Z. Yang, Z. Cai, Y. Miao, Q. Dong, L. Li, C. Ma, L. Chen, R. Xu, Z. Tang, B. Wang, D. Zan, S. Quan, G. Zhang, L. Sha, Y. Zhang, X. Ren, T. Liu, and B. Chang. Omni-MATH: A universal olympiad level mathematic benchmark for large language models. In The Thirteenth International Conference on Learning Representations, 2025.
  26. 26.S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In T. Cohn, Y. He, and Y. Liu, editors, Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online, Nov. 2020. Association for Computational Linguistics.
  27. 27.S. Golchin and M. Surdeanu. Time travel in LLMs: Tracing data contamination in large language models. In The Twelfth International Conference on Learning Representations, 2024.
  28. 28.H. Gulliksen. Theory of Mental Tests. Wiley Publications in Psychology. John Wiley & Sons, Hoboken, NJ, 1950.
  29. 29.R. K. Hambleton and R. W. Jones. Comparison of classical test theory and item response theory and their applications to test development. Educational Measurement: Issues and Practice, 12(3):38–47, 1993.
  30. 30.R. K. Hambleton, H. Swaminathan, and H. J. Rogers. Fundamentals of Item Response Theory, volume 2 of Measurement Methods for the Social Sciences. Sage Publications, Thousand Oaks, CA, 1991.
  31. 31.A. Hardy, A. Reuel, K. Jafari Meimandi, L. Soder, A. Griffith, D. M. Asmar, S. Koyejo, M. S. Bernstein, and M. J. Kochenderfer. More than marketing? on the information value of ai benchmarks for practitioners. In Proceedings of the 30th International Conference on Intelligent User Interfaces, IUI ’25, page 1032–1047, New York, NY, USA, 2025. Association for Computing Machinery.
  32. 32.D. Heineman, V. Hofmann, I. Magnusson, Y. Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge. Signal and noise: A framework for reducing uncertainty in language model evaluation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
  33. 33.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021.
  34. 34.S. Henrysson. Correction of item-total correlations in item analysis. Psychometrika, 28(2):211–218, 1963.
  35. 35.A. Jacovi, A. Caciularu, O. Goldman, and Y. Goldberg. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5075–5084, Singapore, Dec. 2023. Association for Computational Linguistics.
  36. 36.H. Jiang, X. Yi, Z. Wei, Z. Xiao, S. Wang, and X. Xie. Raising the bar: Investigating the values of large language models via generative evolving testing. In Forty-second International Conference on Machine Learning, 2025.
  37. 37.X. Jiang, D. Chang, and X. Xu. Time waits for no benchmark: Exploring the temporal misalignment between static benchmarks, modern LLMs, and the real world. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  38. 38.G. Kamradt. There are 4 stages in a benchmark lifecycle. X (formerly Twitter) post, Nov 2025. Accessed: 2026-01-16.
  39. 39.D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stenetorp, R. Jia, M. Bansal, C. Potts, and A. Williams. Dynabench: Rethinking benchmarking in NLP. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110–4124, Online, June 2021. Association for Computational Linguistics.
  40. 40.E. Kim, S. Li, S. Khalil, and H. J. Shin. STAIR-AIG: Optimizing the automated item generation process through human-AI collaboration for critical thinking assessment. In E. Kochmar, B. Alhafni, M. Bexte, J. Burstein, A. Horbach, R. Laarmann-Quante, A. Tack, V. Yaneva, and Z. Yuan, editors, Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 920–930, Vienna, Austria, July 2025. Association for Computational Linguistics.
  41. 41.R. Le Bras, S. Swayamdipta, C. Bhagavatula, R. Zellers, M. E. Peters, A. Sabharwal, and Y. Choi. Adversarial filters of dataset biases. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org, 2020.
  42. 42.M. Li, H. Jiao, T. Zhou, N. Zhang, S. Peters, and R. W. Lissitz. Item difficulty modeling using fine-tuned small and large language models. In J. Wilson, C. Ormerod, and M. Beiting Parrish, editors, Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers, pages 48–55, Wyndham Grand Pittsburgh, Downtown, Pittsburgh, Pennsylvania, United States, Oct. 2025. National Council on Measurement in Education (NCME).
  43. 43.T. Li, W.-L. Chiang, E. Frick, L. Dunlap, T. Wu, B. Zhu, J. E. Gonzalez, and I. Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning, 2025.
  44. 44.P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Re, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. WANG, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. S. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. A. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda. Holistic evaluation of language models. Transactions on Machine Learning Research, 2023. Featured Certification, Expert Certification, Outstanding Certification.
  45. 45.Q. V. Liao and Z. Xiao. Rethinking model evaluation as narrowing the socio-technical gap, 2025.
  46. 46.B. Y. Lin, Y. Deng, K. Chandu, F. Brahman, A. Ravichander, V. Pyatkin, N. Dziri, R. L. Bras, and Y. Choi. Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024.
  47. 47.F. Lin, S. Xie, Y. Dai, W. Yao, T. Lang, and Y. Zhang. IDGen: Item discrimination induced prompt generation for LLM evaluation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  48. 48.J. Liu, Y. Nam, X. Cui, and S. Swayamdipta. Evaluation under imperfect benchmarks and ratings: A case study in text simplification. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  49. 49.Y. L. Liu, S. L. Blodgett, J. Cheung, Q. V. Liao, A. Olteanu, and Z. Xiao. ECBD: Evidence-centered benchmark design for NLP. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16349–16365, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics.
  50. 50.F. M. Lord. Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980.
  51. 51.F. M. Lord and M. R. Novick. Statistical Theories of Mental Test Scores. Addison-Wesley, Reading, MA, 1968.
  52. 52.I. Magnusson, N. Tai, B. Bogin, D. Heineman, J. D. Hwang, L. Soldaini, A. Bhagia, J. Liu, D. Groeneveld, O. Tafjord, N. A. Smith, P. W. Koh, and J. Dodge. Datadecide: How to predict best pretraining data with small experiments. In Forty-second International Conference on Machine Learning, 2025.
  53. 53.N. Muennighoff, N. Tazi, L. Magne, and N. Reimers. MTEB: Massive text embedding benchmark. In A. Vlachos and I. Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2014–2037, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics.
  54. 54.O. Nahum, N. Calderon, O. Keller, I. Szpektor, and R. Reichart. Are LLMs better than reported? detecting label errors and mitigating their effect on model performance. In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26782–26809, Suzhou, China, Nov. 2025. Association for Computational Linguistics.
  55. 55.Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela. Adversarial NLI: A new benchmark for natural language understanding. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, Online, July 2020. Association for Computational Linguistics.
  56. 56.S. Ott, A. Barbosa-Silva, K. Blagec, J. Brauner, and M. Samwald. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13(1):6793, 2022.
  57. 57.Y. Perlitz, A. Gera, O. Arviv, A. Yehudai, E. Bandel, E. Shnarch, M. Shmueli-Scheuer, and L. Choshen. Benchmark agreement testing done right: A guide for LLM benchmark evaluation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  58. 58.I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes. Closing the ai accountability gap: defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, page 33–44, New York, NY, USA, 2020. Association for Computing Machinery.
  59. 59.M. D. Reckase. 18 multidimensional item response theory. Handbook of statistics, 26:607–642, 2006.
  60. 60.D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. GPQA: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024.
  61. 61.S. Sabour, S. Liu, Z. Zhang, J. Liu, J. Zhou, A. Sunaryo, T. Lee, R. Mihalcea, and M. Huang. EmoBench: Evaluating the emotional intelligence of large language models. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5986–6004, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics.
  62. 62.O. E. Salaudeen, A. Reuel, A. M. Ahmed, S. Bedi, Z. Robertson, S. Sundar, B. W. Domingue, A. Wang, and S. Koyejo. Measurement to meaning: A validity-centered framework for AI evaluation. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  63. 63.D. Sculley, W. Cukierski, P. Culliton, S. Dane, M. M. Demkin, R. Holbrook, A. Howard, P. T. Mooney, W. Reade, M. Risdal, and N. Keating. Position: AI competitions provide the gold standard for empirical rigor in genAI evaluation. In Forty-second International Conference on Machine Learning Position Paper Track, 2025.
  64. 64.H. S. Son. Validity evaluation for the data used for artificial intelligence system. In Y. Bi, R. Bhatia, and S. Kapoor, editors, Intelligent Systems and Applications, pages 362–369, Cham, 2020. Springer International Publishing.
  65. 65.S. Swayamdipta, R. Schwartz, N. Lourie, Y. Wang, H. Hajishirzi, N. A. Smith, and Y. Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In B. Webber, T. Cohn, Y. He, and Y. Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online, Nov. 2020. Association for Computational Linguistics.
  66. 66.S. T. Truong, Y. Tu, P. Liang, B. Li, and S. Koyejo. Reliable and efficient amortized model-based evaluation. In Forty-second International Conference on Machine Learning, 2025.
  67. 67.M. Udell, C. Horn, R. Zadeh, S. Boyd, et al. Generalized low rank models. Foundations and Trends® in Machine Learning, 9(1):1–118, 2016.
  68. 68.K. L. Wagstaff. Machine learning that matters. In Proceedings of the 29th International Coference on International Conference on Machine Learning, ICML’12, page 1851–1856, Madison, WI, USA, 2012. Omnipress.
  69. 69.A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. SuperGLUE: a stickier benchmark for general-purpose language understanding systems. Curran Associates Inc., Red Hook, NY, USA, 2019.
  70. 70.Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. MMLU-pro: A more robust and challenging multi-task language understanding benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024.
  71. 71.D. Westen and R. Rosenthal. Quantifying construct validity: Two simple measures. Journal of Personality and Social Psychology, 84(3):608–618, 2003.
  72. 72.C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Dey, Shubh-Agrawal, S. S. Sandha, S. V. Naidu, C. Hegde, Y. LeCun, T. Goldstein, W. Neiswanger, and M. Goldblum. Livebench: A challenging, contamination-limited LLM benchmark. In The Thirteenth International Conference on Learning Representations, 2025.
  73. 73.S. Wu, H. Bao, S. Li, A. Holtzman, and J. A. Evans. Mapping overlaps in benchmarks through perplexity in the wild, 2025.
  74. 74.Z. Xiao, S. Zhang, V. Lai, and Q. Liao. Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  75. 75.S. M. Xie, H. Pham, X. Dong, N. Du, H. Liu, Y. Lu, P. Liang, Q. V. Le, T. Ma, and A. W. Yu. Doremi: Optimizing data mixtures speeds up language model pretraining. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  76. 76.C. Xu, N. Yan, S. Guan, C. Jin, Y. Mei, Y. Guo, and T. Kechadi. DCR: Quantifying data contamination in LLMs evaluation. In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 23002–23020, Suzhou, China, Nov. 2025. Association for Computational Linguistics.
  77. 77.X. Xu, Z. Wu, R. Qiao, A. Verma, Y. Shu, J. Wang, X. Niu, Z. He, J. Chen, Z. Zhou, G. K. R. Lau, H. Dao, L. Agussurja, R. H. L. Sim, X. Lin, W. Hu, Z. Dai, P. W. Koh, and B. K. H. Low. Position paper: Data-centric AI in the age of large language models. In Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 11895–11913, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics.
  78. 78.Z. Xu, S. Xie, Q. Lv, S. Xiao, L. Song, S. Wenjuan, and F. Lin. Diagnosing failures in large language models’ answers: Integrating error attribution into evaluation framework. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 21148–21165, Vienna, Austria, July 2025. Association for Computational Linguistics.
  79. 79.J. Yao, P. Jin, K. Bao, Q. Yu, K. Bhardwaj, C. Su, J. Wang, Y. ZHU, S. Devare, D. Mosk-Aoyama, Z. Dong, V. K. Srinivasan, Y. Zhang, O. Kuchaiev, J. Jiao, and B. Zhu. The measure of all measures: Quantifying LLM benchmark quality. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025.
  80. 80.A. K. Zhang, K. Klyman, Y. Mai, Y. Levine, Y. Zhang, R. Bommasani, and P. Liang. Position: Language model developers should report train-test overlap. In Forty-second International Conference on Machine Learning Position Paper Track, 2025.
  81. 81.H. Zhang, Y. Chen, and X. Li. A note on exploratory item factor analysis by singular value decomposition. Psychometrika, 85(2):358–372, 2020.

Citation

MLA
Jiang, H., et al. “AI Evaluation Should Require Standardized Item-Level Data Releases”. arXiv, 2026, http://arxiv.org/abs/2604.03244v2.
APA
Jiang, H., Zhang, S., Zhu, D., Bai, Y., Truong, S. T., Yi, X., Koyejo, S., Xie, X., & Xiao, Z. (2026). AI Evaluation Should Require Standardized Item-Level Data Releases. arXiv. http://arxiv.org/abs/2604.03244v2
Chicago
Jiang, H., S. Zhang, D. Zhu, et al. 2026. “AI Evaluation Should Require Standardized Item-Level Data Releases”. arXiv. http://arxiv.org/abs/2604.03244v2.
Harvard
Jiang, H. et al. (2026) “AI Evaluation Should Require Standardized Item-Level Data Releases”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.03244v2.
Vancouver
1. Jiang H, Zhang S, Zhu D, Bai Y, Truong ST, Yi X, Koyejo S, Xie X, Xiao Z (2026) AI Evaluation Should Require Standardized Item-Level Data Releases. arXiv

BibTeX

@article{jiang2026evaluation,
  title = {AI Evaluation Should Require Standardized Item-Level Data Releases},
  author = {Jiang, Han and Zhang, Susu and Zhu, Dongyao and Bai, Yuzhuo and Truong, Sang T. and Yi, Xiaoyuan and Koyejo, Sanmi and Xie, Xing and Xiao, Ziang},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.03244v2},
  eprint = {2604.03244}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/