Are We Done with MMLU?
Aryo Pradipta GemaJoshua Ong Jun LeangGiwon HongAlessio DevotoAlberto Carlo Maria MancinoRohit SaxenaXuanli HeYu ZhaoXiaotang DuMohammad Reza Ghasemi Madani
Reveals widespread ground-truth errors across the MMLU benchmark and provides a corrected 5,700-question dataset that alters established large language model rankings.
The rapid advancement of artificial intelligence has led to widespread reliance on standardized benchmarks to evaluate and compare leading large language models. Among these, the Massive Multitask Language Understanding (MMLU) benchmark has served as a primary industry standard. However, the integrity of model evaluations depends entirely on benchmark data quality. Flaws in underlying test questions can distort performance metrics, obscure true capabilities, and lead to poor strategic decisions when selecting and deploying models.
The article evaluates the reliability of MMLU by systematically identifying ground-truth errors, ambiguities, and formatting flaws across its subjects, and it demonstrates how these defects alter comparative model evaluations. To accomplish this, fourteen human experts reviewed 5,700 questions across all 57 MMLU subjects (100 randomly sampled questions per subject) using a standardized hierarchical taxonomy. The evaluation protocol classified defects into presentation issues (such as poor question or option clarity) and ground-truth issues (such as wrong labels, multiple correct answers, or missing correct options). The authors also tested whether state-of-the-art language models could automatically identify these benchmark errors using techniques like few-shot prompting, retrieval-augmented generation, and fine-tuning on synthetically corrupted datasets.
The investigation produced several key findings. First, an estimated 6.49% of all questions across MMLU contain errors, with extreme error rates concentrated in specific domains—most notably Virology (57% erroneous), Logical Fallacies (26%), and College Chemistry (25%). Second, re-evaluating top models on only the verified, correct subset drastically alters performance scores and shifts model rankings. For example, in the Virology subset, Llama 3.1 405B shifted from 16th place on the original dataset to 1st place on the clean subset, while GPT-4 (0613) in Human Sexuality dropped from 5th to last among top models due to its lower performance on clean items. Third, the analysis uncovered signs of model memorization: certain models performed equal to or better on flawed questions than clean ones, suggesting they learned erroneous benchmark labels during training. Finally, automated error detection using current language models remains ineffective; even the top-performing model (Claude 3 Opus with external retrieval) achieved only an F2 score of 41.92%, demonstrating that automated filtering cannot yet replace human auditing.
These findings indicate that decisions based on standard MMLU benchmarks carry significant risk. Organizations evaluating models based on aggregate benchmark leaderboards may be making procurement or deployment choices driven by flawed questions or training data memorization rather than actual task competence. To mitigate these risks, organizations should rely on corrected subsets—such as the open-source MMLU-Redux dataset created by the authors—rather than uncorrected legacy benchmarks. Furthermore, benchmark curators must institute transparent version control and rigorous human-in-the-loop review processes rather than relying solely on automated quality checks.
While the study reviewed a representative sample across all domains, its primary limitation is that it directly audited only 5,700 of the 14,042 total MMLU questions. Additionally, expert assessments may contain minor subjective variance in interpretive disciplines such as law and ethics. Nevertheless, with high inter-annotator agreement across subjects (Cohen's Kappa above 0.6 to 1.0), there is high confidence in the overall error rates and the conclusion that legacy MMLU data significantly distorts model evaluation.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). This paper establishes the original MMLU benchmark whose questions, design, and ground-truth errors are directly audited and revised in the source study.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). This work analyzes saturation and noise in MMLU to develop MMLU-Pro, providing vital context on efforts to address quality and discriminative flaws in the benchmark.
- Paper: Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?, Nishant Balepur et al. (2024). This paper investigates how language models exploit dataset artifacts and spurious cues in multiple-choice benchmarks, including MMLU, without relying on the actual questions.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). This work illustrates how rigorous auditing and fixing benchmark test cases alters reported model rankings, providing a close conceptual parallel to cleaning MMLU.
- Paper: Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs, Wei Zhou et al. (2026). This survey examines broad techniques for LLM-driven and automated data preparation, cleaning, and error correction across benchmark and real-world datasets.
