When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Norah A. AlzahraniHisham Abdullah AlyahyaYazeed AlnumaySultan AlrashedShaykhah AlsubaieYousef AlmushayqihFaisal MirzaNouf AlotaibiNora Al-TwaireshAreeb Alowisheq

article2024ACL180 citations

Demonstrates that minor prompt perturbations and scoring variations in multiple-choice benchmarks like MMLU can drastically shift language model rankings by up to eight positions, exposing widespread evaluation biases and establishing practical guidelines for reliable model comparison.

Listen

Organizations increasingly rely on public leaderboards based on multiple-choice benchmarks to select large language models for production and research. Because training and deploying these models require substantial capital investments, selecting the right model is often the single most costly decision in an initiative. However, taking published leaderboard standings at face value introduces hidden risks, as relative model performance can shift dramatically under minor, superficial formatting changes.

The main objective of the article is to systematically evaluate how minor perturbations in multiple-choice testing affect model accuracy and rankings across popular benchmarks. It demonstrates the extent to which standard leaderboards are brittle and identifies the primary behavioral biases causing these ranking instabilities.

To evaluate benchmark sensitivity, the authors conducted controlled experiments across eleven prominent language models spanning different sizes and architectures. Using the massive multitask language understanding benchmark of over 14,000 questions across 57 subjects alongside the grade-school science reasoning challenge, the team introduced variations across three primary categories: changing the presentation order and symbols of answer choices, altering prompt phrasing and scoring mechanisms, and manipulating in-context knowledge through few-shot examples.

The investigation produced several key findings. First, minor perturbations cause severe leaderboard volatility, shifting model standings by up to eight positions and altering rank correlation measures significantly. Second, all evaluated models exhibit acute selection bias driven by preferences for specific position orders or token symbols; replacing standard letters with rare symbols caused notable accuracy drops across models and triggered unpredictable bias spikes. Third, models are highly sensitive to the scoring method used: standard symbol scoring yields the highest nominal accuracy but the worst selection bias, whereas cloze scoring lowers bias at the expense of accuracy. Fourth, providing misleading or patterned answers in few-shot demonstration examples severely degrades reasoning across models of all sizes, dropping accuracy significantly when incorrect context is supplied. Conversely, benign prompt changes—such as removing subject names or adding simple formatting examples—exerted minimal influence on relative rankings.

These findings imply that multiple-choice benchmark leaderboards reflect prompt formatting affinities and selection biases rather than genuine differences in language comprehension or reasoning. Relying blindly on standard rankings introduces serious performance and financial risks, potentially leading organizations to invest heavily in overfitted or fragile models that underperform in real-world applications.

For practitioners evaluating models, the article recommends replacing pure symbol scoring with hybrid scoring—a method that presents all choices in the prompt but scores the likelihood of full answer text normalized by length. Hybrid scoring provides a superior balance by preserving model accuracy while substantially mitigating position and symbol biases. Organizations should also incorporate diverse prompt structures and few-shot examples rather than depending on a single zero-shot format.

The primary limitation of this work is that it identifies behavioral vulnerabilities without isolating the exact root causes within model training data, which remain proprietary and inaccessible. Consequently, while hybrid scoring and multi-prompt evaluations reduce measurement sensitivity, they do not completely resolve leaderboard instability. Decision-makers should treat current multiple-choice leaderboards with caution and augment them with domain-specific validations before committing to major model deployments.

  • Paper: Are We Done with MMLU?, Aryo Pradipta Gema et al. (2025). It extends scrutiny of MMLU leaderboard reliability by auditing question defects and showing how corrected data can change model rankings.
Cover for When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Abstract

Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value — we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance of LLMs is highly sensitive to (often minute) details. We show that for popular multiple-choice question benchmarks (e.g., MMLU), minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions. We explain this phenomenon by conducting systematic experiments over three broad categories of benchmark perturbations and identifying the sources of this behavior. Our analysis results in several best-practice recommendations, including the advantage of a hybrid scoring method for answer selection. Our study highlights the dangers of relying on simple benchmark evaluations and charts the path for more robust evaluation schemes on the existing benchmarks. The code for this paper is available at https://github.com/National-Center-for-AI-Saudi-Arabia/lm-evaluation-harness.

Table of Contents

  • 1 Introduction
  • 2 LLM Evaluation with MCQs
  • 3 Methods
  • 3.1 Answer choice format and ordering
  • 3.2 Prompt and scoring modifications
  • 3.3 In-context knowledge manipulation
  • 4 Experiments
  • 5 Results & Analysis
  • 5.1 MCQ benchmarks are not robust to perturbations
  • 5.2 Revisiting selection bias: token bias vs. position bias
  • 5.3 Another source of bias: scoring bias
  • 5.4 Minor few-shot and prompt changes have little effect on benchmark rankings
  • 5.5 LLMs readily reference knowledge provided in-context (even if it is misleading)
  • 6 Related Work
  • 7 Conclusion
  • 8 Limitations
  • 9 Potential Risks
  • References
  • A.1 Appendix
  • A.1.1 Baselines
  • A.1.2 Answer choice format and ordering
  • A.1.3 Prompt and scoring modifications
  • A.1.4 In-context Knowledge Manipulation
  • A.1.5 MMLU Overview

Knowls

  1. Knowl 1 — Evaluation design for testing leaderboard sensitivity

    experimental setup

    The study tested 11 language-model variants, primarily on MMLU, which contains 14,042 four-option questions across 57 subjects. Choice-order experiments used a manually checked subset of college chemistry, college mathematics, and global facts to avoid questions whose answer choices refer to one another. Some experiments were also run on the ARC Challenge benchmark. Models were evaluated in zero-shot and five-shot settings. The authors measured accuracy changes, answer-choice bias using RStd (the standard deviation of recall across answer choices), and ranking agreement using normalized Kendall agreement, where 1 means identical rankings and 0 means complete reversal. Perturbations covered answer-choice ordering and labels, prompt and scoring changes, and information supplied in context.

  2. Knowl 2 — Small benchmark changes can substantially reorder model rankings

    empirical result

    Across the tested settings, minor changes to multiple-choice evaluation—including changing choice order or labels and changing the scoring method—sometimes caused large leaderboard shifts. The paper reports movements of up to eight positions across perturbations. In a controlled three-subject MMLU choice-shuffling experiment, 5 of 11 models changed rank, the largest movement was five positions, and normalized Kendall agreement with the original ranking was 0.564. Sensitivity varied by model: Yi-6B moved from third place to seventh or eighth under some rare-label and cloze evaluations, while Mistral-7B and Llama-2-7B generally moved only one or two places. The authors suggest benchmark-style overfitting as one possible explanation for some models’ behavior, but could not verify it.

  3. Knowl 3 — Forcing the correct answer to a position exposes strong position preferences

    empirical result

    When every correct answer on zero-shot MMLU was placed at the same option position, all tested models showed a preference for particular positions, and the preferred position differed across models. For example, Llama-2-7B accuracy ranged from 23.67% with the correct answer always at D to 79.00% with it always at A; Yi-6B scored 18.33% when the answer was always at C, compared with a 61.12% baseline. Ranking agreement with the baseline was 0.455, 0.527, 0.527, and 0.855 when the correct answer was fixed at A, B, C, and D, respectively. The preference persisted in five-shot tests, including cases where the demonstration answers were also position-controlled.

  4. Knowl 4 — Replacing familiar answer labels does not remove label bias

    empirical result

    The authors replaced A/B/C/D with either common language-independent symbols ($, &, #, @) or rarer labels (œ, §, a Cyrillic character, ü). On full zero-shot MMLU, accuracy decreased for every tested model under both label sets; most models also showed increased RStd, and normalized ranking agreement with the original evaluation was 0.689 for the common-symbol set and 0.733 for the rare-symbol set. In three-subject tests, shuffling rare labels or shuffling choice text while keeping labels fixed changed accuracy and bias inconsistently. Thus, using less familiar labels did not reliably eliminate answer-label or position-related selection bias.

  5. Knowl 5 — Scoring method trades accuracy against selection bias

    model/method

    The study compared three multiple-choice scoring methods. Symbol scoring presents all choices and selects the most likely option label. Hybrid scoring also presents all choices but scores the answer content using length-normalized likelihood. Cloze scoring presents the question with one candidate answer at a time and selects the candidate with the highest normalized likelihood. On MMLU, symbol scoring generally produced the highest accuracy and the greatest selection bias. Cloze scoring substantially reduced RStd but lowered accuracy across the tested models; its ranking agreement with the symbol-scoring baseline was 0.527. Hybrid scoring reduced bias relative to symbol scoring while retaining a prompt that shows all choices; its MMLU ranking agreement with the symbol baseline was 0.709. On ARC Challenge, hybrid scoring had ranking agreement of 0.782 with the cloze baseline, and some models scored higher than under that baseline. The authors recommend hybrid scoring as a useful compromise, while noting that it does not make rankings fully robust.

  6. Knowl 6 — Models often use an answer supplied in the prompt

    empirical result

    When the target question and its correct answer were supplied in context as a demonstration, models achieved high but not perfect MMLU accuracy. Across the 11 models, accuracy ranged from 61.00% to 99.10% in the one-shot setting and from 63.82% to 99.79% in the five-shot setting. Ranking agreement with the ordinary evaluation was only 0.491 for one-shot and 0.382 for five-shot. The results show that models can use answer information present in context, while their failure to reach 100% indicates that the supplied answer did not ensure a correct response on every item.

  7. Knowl 7 — Incorrect answers in context sharply degrade performance

    empirical result

    When the target question was instead paired with an incorrect answer in context, performance fell sharply across all tested models. MMLU accuracy ranged from 10.71% to 36.13% in the one-shot setting and from 4.50% to 37.42% in the five-shot setting. Ranking agreement with the ordinary evaluation was 0.382 and 0.164, respectively. The result held across model sizes and indicates that misleading answer information in demonstrations can strongly affect both measured accuracy and model rankings.

  8. Knowl 8 — A fixed answer pattern in demonstrations disrupts performance

    empirical result

    The authors fixed the correct answers in all five-shot demonstrations to the same option position (A, B, C, or D) and then evaluated the target questions. In the reported three-subject MMLU results for four models, accuracy declined for every model and every fixed position, by 11.55 to 27.70 percentage points relative to the five-shot baseline. This indicates that models can be affected by a repeated answer-position pattern in context, even when that pattern is not itself relevant to solving the target question.

  9. Knowl 9 — Several prompt and demonstration changes leave rankings largely intact

    empirical result

    Removing the subject name from the prompt or changing “Answer:” to “Correct Answer:” produced minimal accuracy changes and high ranking agreement: normalized Kendall agreement ranged from 0.927 to 0.964 in zero-shot tests and was 1.0 in the reported five-shot tests. Replacing ordinary demonstrations with trivial questions or five-shot questions from another subject also generally preserved rankings: agreement ranged from 0.891 to 0.927 for trivial demonstrations and was 0.927 or 0.964 for subject-independent demonstrations. Subject-independent demonstrations usually lowered accuracy by about two to three percentage points. These results distinguish relatively benign prompt and demonstration changes from the choice-format and scoring perturbations that caused larger ranking shifts.

  10. Knowl 10 — The study could not identify a robust remedy or fully explain the biases

    limitation

    The experiments establish that multiple-choice leaderboard rankings can be sensitive to evaluation details, but they do not quantify the relative causal contributions of token, position, scoring, and other biases. The authors could not rule out benchmark contamination because the models’ pretraining data were unavailable. Hybrid scoring is recommended as a way to reduce some bias, not as a complete solution: it remains sensitive to perturbations. The study therefore demonstrates instability more conclusively than it explains its origins or resolves it.

Coverage note — The exhaustive per-model appendix tables and per-subject MMLU answer distributions are omitted because they repeat the reported patterns or provide dataset bookkeeping rather than distinct contributions.

References

  1. 1.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023. Palm 2 technical report.
  2. 2.Anthropic. 2023. Anthropic. model card and evaluations for claude models.
  3. 3.Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. 2023. Open llm leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard.
  4. 4.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2023. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018a. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457.
  7. 7.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018b. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  8. 8.OpenCompass Contributors. 2023. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass.
  9. 9.Google Deepmind. 2023. Gemini: A family of highly capable multimodal models.
  10. 10.Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. 2021. The benchmark lottery.
  11. 11.Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. 2023. Benchmark probing: Investigating data leakage in large language models. In NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly.
  12. 12.Kawin Ethayarajh and Dan Jurafsky. 2021. Utility is in the eye of the user: A critique of nlp leaderboards.
  13. 13.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023. A framework for few-shot language model evaluation.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  15. 15.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  16. 16.Maurice G Kendall. 1938. A new measure of rank correlation. Biometrika, 30(1/2):81–93.
  17. 17.Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang. 2023. A systematic study and comprehensive evaluation of chatgpt on benchmark datasets. arXiv preprint arXiv:2305.18486.
  18. 18.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023. Holistic evaluation of language models.
  19. 19.Séverin Lions, Carlos Monsalve, Pablo Dartnell, María Inés Godoy, Nora Córdova, Daniela Jiménez, María Paz Blanco, Gabriel Ortega, and Julie Lemarié. 2021. The position of distractors in multiple-choice test items: The strongest precede the weakest. In Frontiers in Education, volume 6, page 731763. Frontiers.
  20. 20.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  21. 21.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11048–11064, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  22. 22.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions.
  23. 23.Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2024. State of what art? a call for multi-prompt llm evaluation.
  24. 24.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019. Adversarial nli: A new benchmark for natural language understanding. arXiv preprint arXiv:1910.14599.
  25. 25.OpenAI. 2023. Gpt-4 technical report.
  26. 26.Pouya Pezeshkpour and Estevam Hruschka. 2023. Large language models sensitivity to the order of options in multiple-choice questions.
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  28. 28.Joshua Robinson, Christopher Michael Rytting, and David Wingate. 2023. Leveraging large language models for multiple choice question answering.
  29. 29.Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, and Sarath Chandar. 2018. Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph.
  30. 30.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  31. 31.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.
  32. 32.Chenhui Shen, Liying Cheng, Yang You, and Lidong Bing. 2023. Are large language models good evaluators for abstractive summarization? arXiv preprint arXiv:2305.13091.
  33. 33.Saleh Soltan, Shankar Ananthakrishnan, Jack Fitzgerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, et al. 2022. Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model. arXiv preprint arXiv:2208.01448.
  34. 34.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  36. 36.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32.
  37. 37.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461.
  38. 38.Yike Wang, Shangbin Feng, Heng Wang, Weijia Shi, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov. 2023. Resolving knowledge conflicts in large language models.
  39. 39.Lucas Weber, Elia Bruni, and Dieuwke Hupkes. 2023. Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 294–313, Singapore. Association for Computational Linguistics.
  40. 40.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma. 2023. Larger language models do in-context learning differently.
  41. 41.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2023. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts.
  42. 42.Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang goo Lee, and Taeuk Kim. 2022. Ground-truth labels matter: A deeper look into input-label demonstrations.
  43. 43.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.
  44. 44.Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023. Large language models are not robust multiple choice selectors. arXiv e-prints, pages arXiv–2309.
  45. 45.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.

Citation

MLA
Alzahrani, N. A., et al. “When Benchmarks Are Targets: Revealing the Sensitivity of Large Language Model Leaderboards”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13787–805, https://doi.org/10.18653/v1/2024.acl-long.744.
APA
Alzahrani, N. A., Alyahya, H., Alnumay, Y., AlRashed, S., Alsubaie, S., Almushayqih, Y., Mirza, F., Alotaibi, N. M., Al-Twairesh, N., Alowisheq, A., Bari, M. S., & Khan, H. (2024). When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13787–13805. https://doi.org/10.18653/v1/2024.acl-long.744
Chicago
Alzahrani, N. A., H. Alyahya, Y. Alnumay, et al. 2024. “When Benchmarks Are Targets: Revealing the Sensitivity of Large Language Model Leaderboards”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13787–805. https://doi.org/10.18653/v1/2024.acl-long.744.
Harvard
Alzahrani, N.A. et al. (2024) “When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13787–13805. Available at: https://doi.org/10.18653/v1/2024.acl-long.744.
Vancouver
1. Alzahrani NA, Alyahya H, Alnumay Y, et al (2024) When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13787–13805

BibTeX

@inproceedings{alzahrani-etal-2024-benchmarks,
    title = "When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards",
    author = "Alzahrani, Norah  and
      Alyahya, Hisham  and
      Alnumay, Yazeed  and
      AlRashed, Sultan  and
      Alsubaie, Shaykhah  and
      Almushayqih, Yousef  and
      Mirza, Faisal  and
      Alotaibi, Nouf  and
      Al-Twairesh, Nora  and
      Alowisheq, Areeb  and
      Bari, M Saiful  and
      Khan, Haidar",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.744/",
    doi = "10.18653/v1/2024.acl-long.744",
    pages = "13787--13805"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/