When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Norah A. AlzahraniHisham Abdullah AlyahyaYazeed AlnumaySultan AlrashedShaykhah AlsubaieYousef AlmushayqihFaisal MirzaNouf AlotaibiNora Al-TwaireshAreeb Alowisheq
Demonstrates that minor prompt perturbations and scoring variations in multiple-choice benchmarks like MMLU can drastically shift language model rankings by up to eight positions, exposing widespread evaluation biases and establishing practical guidelines for reliable model comparison.
Organizations increasingly rely on public leaderboards based on multiple-choice benchmarks to select large language models for production and research. Because training and deploying these models require substantial capital investments, selecting the right model is often the single most costly decision in an initiative. However, taking published leaderboard standings at face value introduces hidden risks, as relative model performance can shift dramatically under minor, superficial formatting changes.
The main objective of the article is to systematically evaluate how minor perturbations in multiple-choice testing affect model accuracy and rankings across popular benchmarks. It demonstrates the extent to which standard leaderboards are brittle and identifies the primary behavioral biases causing these ranking instabilities.
To evaluate benchmark sensitivity, the authors conducted controlled experiments across eleven prominent language models spanning different sizes and architectures. Using the massive multitask language understanding benchmark of over 14,000 questions across 57 subjects alongside the grade-school science reasoning challenge, the team introduced variations across three primary categories: changing the presentation order and symbols of answer choices, altering prompt phrasing and scoring mechanisms, and manipulating in-context knowledge through few-shot examples.
The investigation produced several key findings. First, minor perturbations cause severe leaderboard volatility, shifting model standings by up to eight positions and altering rank correlation measures significantly. Second, all evaluated models exhibit acute selection bias driven by preferences for specific position orders or token symbols; replacing standard letters with rare symbols caused notable accuracy drops across models and triggered unpredictable bias spikes. Third, models are highly sensitive to the scoring method used: standard symbol scoring yields the highest nominal accuracy but the worst selection bias, whereas cloze scoring lowers bias at the expense of accuracy. Fourth, providing misleading or patterned answers in few-shot demonstration examples severely degrades reasoning across models of all sizes, dropping accuracy significantly when incorrect context is supplied. Conversely, benign prompt changes—such as removing subject names or adding simple formatting examples—exerted minimal influence on relative rankings.
These findings imply that multiple-choice benchmark leaderboards reflect prompt formatting affinities and selection biases rather than genuine differences in language comprehension or reasoning. Relying blindly on standard rankings introduces serious performance and financial risks, potentially leading organizations to invest heavily in overfitted or fragile models that underperform in real-world applications.
For practitioners evaluating models, the article recommends replacing pure symbol scoring with hybrid scoring—a method that presents all choices in the prompt but scores the likelihood of full answer text normalized by length. Hybrid scoring provides a superior balance by preserving model accuracy while substantially mitigating position and symbol biases. Organizations should also incorporate diverse prompt structures and few-shot examples rather than depending on a single zero-shot format.
The primary limitation of this work is that it identifies behavioral vulnerabilities without isolating the exact root causes within model training data, which remain proprietary and inaccessible. Consequently, while hybrid scoring and multi-prompt evaluations reduce measurement sensitivity, they do not completely resolve leaderboard instability. Decision-makers should treat current multiple-choice leaderboards with caution and augment them with domain-specific validations before committing to major model deployments.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Its controlled option-reordering tests show how answer-position cues can distort multiple-choice accuracy, a key mechanism behind the leaderboard sensitivity examined here.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). Its CircularEval method rotates answer choices to expose option bias, providing useful context for this paper’s experiments on benchmark perturbations.
- Paper: Are We Done with MMLU?, Aryo Pradipta Gema et al. (2025). It extends scrutiny of MMLU leaderboard reliability by auditing question defects and showing how corrected data can change model rankings.
