Are We Done with MMLU?

Aryo Pradipta GemaJoshua Ong Jun LeangGiwon HongAlessio DevotoAlberto Carlo Maria MancinoRohit SaxenaXuanli HeYu ZhaoXiaotang DuMohammad Reza Ghasemi Madani

article2025NAACL193 citations

Reveals widespread ground-truth errors across the MMLU benchmark and provides a corrected 5,700-question dataset that alters established large language model rankings.

Listen

The rapid advancement of artificial intelligence has led to widespread reliance on standardized benchmarks to evaluate and compare leading large language models. Among these, the Massive Multitask Language Understanding (MMLU) benchmark has served as a primary industry standard. However, the integrity of model evaluations depends entirely on benchmark data quality. Flaws in underlying test questions can distort performance metrics, obscure true capabilities, and lead to poor strategic decisions when selecting and deploying models.

The article evaluates the reliability of MMLU by systematically identifying ground-truth errors, ambiguities, and formatting flaws across its subjects, and it demonstrates how these defects alter comparative model evaluations. To accomplish this, fourteen human experts reviewed 5,700 questions across all 57 MMLU subjects (100 randomly sampled questions per subject) using a standardized hierarchical taxonomy. The evaluation protocol classified defects into presentation issues (such as poor question or option clarity) and ground-truth issues (such as wrong labels, multiple correct answers, or missing correct options). The authors also tested whether state-of-the-art language models could automatically identify these benchmark errors using techniques like few-shot prompting, retrieval-augmented generation, and fine-tuning on synthetically corrupted datasets.

The investigation produced several key findings. First, an estimated 6.49% of all questions across MMLU contain errors, with extreme error rates concentrated in specific domains—most notably Virology (57% erroneous), Logical Fallacies (26%), and College Chemistry (25%). Second, re-evaluating top models on only the verified, correct subset drastically alters performance scores and shifts model rankings. For example, in the Virology subset, Llama 3.1 405B shifted from 16th place on the original dataset to 1st place on the clean subset, while GPT-4 (0613) in Human Sexuality dropped from 5th to last among top models due to its lower performance on clean items. Third, the analysis uncovered signs of model memorization: certain models performed equal to or better on flawed questions than clean ones, suggesting they learned erroneous benchmark labels during training. Finally, automated error detection using current language models remains ineffective; even the top-performing model (Claude 3 Opus with external retrieval) achieved only an F2 score of 41.92%, demonstrating that automated filtering cannot yet replace human auditing.

These findings indicate that decisions based on standard MMLU benchmarks carry significant risk. Organizations evaluating models based on aggregate benchmark leaderboards may be making procurement or deployment choices driven by flawed questions or training data memorization rather than actual task competence. To mitigate these risks, organizations should rely on corrected subsets—such as the open-source MMLU-Redux dataset created by the authors—rather than uncorrected legacy benchmarks. Furthermore, benchmark curators must institute transparent version control and rigorous human-in-the-loop review processes rather than relying solely on automated quality checks.

While the study reviewed a representative sample across all domains, its primary limitation is that it directly audited only 5,700 of the 14,042 total MMLU questions. Additionally, expert assessments may contain minor subjective variance in interpretive disciplines such as law and ethics. Nevertheless, with high inter-annotator agreement across subjects (Cohen's Kappa above 0.6 to 1.0), there is high confidence in the overall error rates and the conclusion that legacy MMLU data significantly distorts model evaluation.

arXiv: 2406.04127
Cover for Are We Done with MMLU?

Abstract

Maybe not. We identify and analyse errors in the popular Massive Multitask Language Understanding (MMLU) benchmark. Even though MMLU is widely adopted, our analysis demonstrates numerous ground truth errors that obscure the true capabilities of LLMs. For example, we find that 57% of the analysed questions in the Virology subset contain errors. To address this issue, we introduce a comprehensive framework for identifying dataset errors using a novel error annotation protocol. Then, we create MMLU-Redux, which is a subset of 5,700 manually re-annotated questions across all 57 MMLU subjects. We estimate that 6.49% of MMLU questions contain errors. Using MMLU-Redux, we demonstrate significant discrepancies with the model performance metrics that were originally reported. Our results strongly advocate for revising MMLU's error-ridden questions to enhance its future utility and reliability as a benchmark. https://huggingface.co/datasets/edinburgh-dawg/mmlu-redux-2.0.

Table of Contents

  • 1 Introduction
  • 2 What Is Wrong with MMLU?
  • 2.1 Error Detection Annotation Protocol
  • 2.2 Heterogeneity of Errors in MMLU
  • 3 MMLU-Redux: A Correct MMLU Subset
  • 3.1 Analysis of MMLU-Redux
  • 3.2 Re-Evaluating the State-of-the-Art LLMs
  • 3.3 Correlation between Exact Match Score and Data Quality
  • 4 Can We Fix the MMLU Dataset Automatically?
  • 4.1 Error Detection Experiments
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Acknowledgement
  • References
  • A Responsible Research Checklist
  • B Data Collection and Organisation
  • C Usage Guidelines
  • D Maintenance Plan
  • D.1 General Maintenance
  • D.2 Contributing Guidelines
  • D.3 Contingency Plan
  • E Authors Statement
  • F MMLU-Redux Error Type Statistics
  • G Prompting Methods
  • Chain of Thought (CoT) Prompting
  • H Details on automatic error detection
  • I Heterogeneity of Errors in MMLU Subsets - Full List
  • J Error Detection via Fine-tuning

Knowls

  1. Knowl 1 — Hierarchical Error Annotation Protocol for Multiple-Choice QA

    model/method

    To categorize errors in multiple-choice language understanding benchmarks, a two-stage hierarchical annotation decision protocol is used:

    1. Question Assessment (Type 1): Assesses whether the question and choices are unambiguous, well-formed, and self-contained.

      • (1a) Bad Question Clarity: The prompt is poorly phrased, ungrammatical, incomplete, or refers to unavailable external context (such as a missing diagram, chapter passage, or prior question).
      • (1b) Bad Options Clarity: The options are malformed, nonsensical, duplicated, or corrupted by parsing errors (e.g., merging or splitting option strings).
    2. Ground Truth Verification (Type 2): If presentation is clear, the correctness of the choices relative to the prompt is evaluated.

      • (2a) No Correct Answer: None of the listed choices provide a correct answer to the question.
      • (2b) Multiple Correct Answers: More than one choice is valid (e.g., distinct valid answers or synonymous options).
      • (2c) Wrong Ground Truth: Exactly one option is correct, but the benchmark's annotated label points to an incorrect option.
      • OK: Exactly one option is correct, and it matches the benchmark's annotated ground truth label.
  2. Knowl 2 — MMLU-Redux Dataset Construction and Schema

    experimental setup

    MMLU-Redux is a human-annotated dataset designed to audit and categorize ground truth errors across all 57 subjects of the Massive Multitask Language Understanding (MMLU) benchmark. A team of 14 annotators manually reviewed 100 randomly sampled questions per subject, yielding 5,700 annotated questions out of the 14,042 questions in MMLU.

    Annotators checked the original question sources (such as textbooks, GRE/MCAT exams, and academic databases) and authoritative reference materials (e.g., government publications, World Health Organization reports) to verify factual correctness.

    Each sample in MMLU-Redux contains seven fields:

    • question (string): The original text of the MMLU question.
    • choices (list of four strings): The four multiple-choice candidate options.
    • answer (integer from 00 to 33): The original MMLU ground truth index.
    • error_type (string): One of ok, bad_question_clarity, bad_options_clarity, no_correct_answer, multiple_correct_answers, wrong_groundtruth, or expert (flagged when specialized domain-level consensus is required).
    • source (string): The online URL or literature reference identified for the question.
    • correct_answer (string): The suggested correct option text if the original label was erroneous.
    • potential_reason (string): Free-text diagnostic explaining the suspected origin of the error (e.g., OCR/HTML parsing glitch, omitted 5th choice, or temporal ambiguity).
  3. Knowl 3 — Overall and Per-Subject Error Rate Distribution in MMLU

    empirical result

    Based on stratified sampling across the 5,700 questions reviewed in MMLU-Redux, an estimated 6.49%6.49\% of questions in the full MMLU benchmark contain errors (93.51%93.51\% are annotated as OK). The benchmark-wide estimated breakdown by error category is:

    • Bad Question Clarity (1a): 2.47%2.47\%
    • Multiple Correct Answers (2b): 1.54%1.54\%
    • Wrong Ground Truth (2c): 1.42%1.42\%
    • No Correct Answer (2a): 0.62%0.62\%
    • Bad Options Clarity (1b): 0.44%0.44\%

    Error prevalence varies substantially across subjects. The most error-dense subjects (sampled at 100 questions each) include:

    • Virology: 57%57\% errors (4343 OK, 3333 wrong ground truth, 1414 bad question clarity, 44 multiple correct answers, 44 no correct answer, 22 bad options clarity)
    • Logical Fallacies: 26%26\% errors (7474 OK, 1414 bad question clarity, 44 no correct answer, 33 multiple correct answers, 33 wrong ground truth, 22 bad options clarity)
    • College Chemistry: 25%25\% errors (7575 OK, 2121 wrong ground truth, 22 bad question clarity, 22 no correct answer)
    • Professional Law: 18%18\% errors (8282 OK, 1111 multiple correct answers, 44 bad question clarity, 22 wrong ground truth, 11 bad options clarity)
    • Human Sexuality: 17%17\% errors (8383 OK, 1111 bad question clarity, 22 no correct answer, 22 wrong ground truth, 11 bad options clarity, 11 multiple correct answers)
    • Business Ethics: 14%14\% errors (8686 OK, 1414 bad question clarity)
    • Formal Logic: 13%13\% errors (8787 OK, 99 multiple correct answers, 22 no correct answer, 11 bad question clarity, 11 wrong ground truth)
    • Human Aging: 12%12\% errors (8888 OK, 1212 bad question clarity)
    • Global Facts: 12%12\% errors (8888 OK, 55 wrong ground truth, 44 no correct answer, 22 bad question clarity, 11 bad options clarity)
    • Machine Learning: 11%11\% errors (8989 OK, 55 bad options clarity, 33 bad question clarity, 22 wrong ground truth, 11 multiple correct answers)

    Conversely, nine subjects showed 0%0\% errors in the evaluated sample: Management, Philosophy, Medical Genetics, High School US History, High School Government and Politics, High School Geography, High School Computer Science, College Physics, and Prehistory.

  4. Knowl 4 — Inter-Annotator Agreement on MMLU Error Labeling

    data/table

    To assess the reliability of the error annotation protocol, inter-annotator agreement was evaluated using Cohen's Kappa (κ\kappa) across three independent annotators on five high-error subjects from MMLU-Redux. Agreement was calculated for fine-grained multi-class classification ("All Error Type") and binary error presence ("Binary Error Type").

    Subject Cohen's Kappa (All Error Type) Cohen's Kappa (Binary Error Type)
    Virology 0.67 0.67
    Logical Fallacies 0.73 0.71
    College Chemistry 0.92 0.95
    Formal Logic 0.96 1.00
    Human Sexuality 0.64 0.64

    All evaluated subjects achieved κ>0.60\kappa > 0.60, reflecting substantial (κ∈[0.60,0.80]\kappa \in [0.60, 0.80]) to almost perfect (κ>0.80\kappa > 0.80) agreement. The close alignment between multi-class and binary κ\kappa values indicates that human experts consistently identified not only whether an error was present, but also the specific error category.

  5. Knowl 5 — Distortion of LLM Evaluation and Leaderboard Rankings from Benchmark Errors

    data/table

    When large language models (LLMs) are evaluated exclusively on error-free instances compared to all instances on the five most error-prone subjects in MMLU-Redux, Exact Match (EM) accuracy and model rankings change significantly.

    Model Virology Logical Fallacies College Chemistry Professional Law Human Sexuality
    Claude 3.5 Sonnet (20240620) 0.60 (1)→0.91 (5)0.60\,(1) \rightarrow 0.91\,(5) 0.93 (1)→0.96 (5)0.93\,(1) \rightarrow 0.96\,(5) 0.59 (9)→0.73 (4)0.59\,(9) \rightarrow 0.73\,(4) 0.75 (1)→0.77 (1)0.75\,(1) \rightarrow 0.77\,(1) 0.94 (1)→0.98 (1)0.94\,(1) \rightarrow 0.98\,(1)
    Claude 3 Opus (20240229) 0.58 (12)→0.88 (8)0.58\,(12) \rightarrow 0.88\,(8) 0.90 (4)→0.96 (5)0.90\,(4) \rightarrow 0.96\,(5) 0.60 (5)→0.72 (5)0.60\,(5) \rightarrow 0.72\,(5) 0.72 (4)→0.72 (3)0.72\,(4) \rightarrow 0.72\,(3) 0.91 (5)→0.96 (2)0.91\,(5) \rightarrow 0.96\,(2)
    Llama 3.1 Instruct Turbo (405B) 0.57 (16)→0.93 (1)0.57\,(16) \rightarrow 0.93\,(1) 0.92 (2)→0.96 (5)0.92\,(2) \rightarrow 0.96\,(5) 0.60 (5)→0.76 (1)0.60\,(5) \rightarrow 0.76\,(1) 0.70 (6)→0.72 (3)0.70\,(6) \rightarrow 0.72\,(3) 0.86 (20)→0.91 (9)0.86\,(20) \rightarrow 0.91\,(9)
    GPT-4o (2024-05-13) 0.60 (3)→0.91 (5)0.60\,(3) \rightarrow 0.91\,(5) 0.88 (6)→0.99 (2)0.88\,(6) \rightarrow 0.99\,(2) 0.61 (4)→0.71 (7)0.61\,(4) \rightarrow 0.71\,(7) 0.72 (3)→0.70 (5)0.72\,(3) \rightarrow 0.70\,(5) 0.91 (5)→0.96 (2)0.91\,(5) \rightarrow 0.96\,(2)
    Gemini 1.5 Pro (001) 0.55 (28)→0.91 (5)0.55\,(28) \rightarrow 0.91\,(5) 0.90 (4)→0.99 (2)0.90\,(4) \rightarrow 0.99\,(2) 0.62 (2)→0.72 (5)0.62\,(2) \rightarrow 0.72\,(5) 0.67 (9)→0.67 (7)0.67\,(9) \rightarrow 0.67\,(7) 0.37 (55)→0.94 (6)0.37\,(55) \rightarrow 0.94\,(6)
    GPT-4 (0613) 0.60 (3)→0.86 (10)0.60\,(3) \rightarrow 0.86\,(10) 0.87 (11)→0.99 (2)0.87\,(11) \rightarrow 0.99\,(2) 0.55 (18)→0.75 (3)0.55\,(18) \rightarrow 0.75\,(3) 0.73 (2)→0.68 (6)0.73\,(2) \rightarrow 0.68\,(6) 0.91 (5)→0.43 (10)0.91\,(5) \rightarrow 0.43\,(10)
    Qwen2 Instruct (72B) 0.56 (24)→0.88 (8)0.56\,(24) \rightarrow 0.88\,(8) 0.91 (3)→0.96 (5)0.91\,(3) \rightarrow 0.96\,(5) 0.65 (1)→0.68 (8)0.65\,(1) \rightarrow 0.68\,(8) 0.66 (10)→0.74 (2)0.66\,(10) \rightarrow 0.74\,(2) 0.89 (11)→0.94 (6)0.89\,(11) \rightarrow 0.94\,(6)
    GPT-4 Turbo (2024-04-09) 0.60 (1)→0.93 (1)0.60\,(1) \rightarrow 0.93\,(1) 0.87 (11)→1.00 (1)0.87\,(11) \rightarrow 1.00\,(1) 0.53 (22)→0.76 (1)0.53\,(22) \rightarrow 0.76\,(1) 0.67 (8)→0.63 (9)0.67\,(8) \rightarrow 0.63\,(9) 0.90 (9)→0.93 (8)0.90\,(9) \rightarrow 0.93\,(8)
    Gemini 1.5 Pro (0409 preview) 0.58 (10)→0.93 (1)0.58\,(10) \rightarrow 0.93\,(1) 0.86 (18)→0.92 (10)0.86\,(18) \rightarrow 0.92\,(10) 0.58 (13)→0.67 (9)0.58\,(13) \rightarrow 0.67\,(9) 0.64 (13)→0.61 (10)0.64\,(13) \rightarrow 0.61\,(10) 0.40 (56)→0.95 (5)0.40\,(56) \rightarrow 0.95\,(5)
    Llama 3.1 Instruct Turbo (70B) 0.58 (12)→0.93 (1)0.58\,(12) \rightarrow 0.93\,(1) 0.84 (27)→0.96 (5)0.84\,(27) \rightarrow 0.96\,(5) 0.59 (9)→0.64 (10)0.59\,(9) \rightarrow 0.64\,(10) 0.67 (7)→0.65 (8)0.67\,(7) \rightarrow 0.65\,(8) 0.86 (20)→0.96 (2)0.86\,(20) \rightarrow 0.96\,(2)

    In Virology, Llama 3.1 Instruct Turbo (405B) moves from 16th place on all instances (extEM=0.57 ext{EM}=0.57) to 1st place on clean instances (extEM=0.93 ext{EM}=0.93). In Human Sexuality, GPT-4 (0613) scores 0.910.91 across all instances (rank 5) but drops to 0.430.43 (rank 10) on clean instances.

    In subjects such as Professional Law and Formal Logic, models occasionally perform equally well or better on erroneous instances than on clean instances with respect to the flawed original labels, providing evidence of pretraining memorization on MMLU benchmark items.

  6. Knowl 6 — LLM Performance on Automated Benchmark Error Detection

    empirical result

    Language models evaluated on identifying erroneous items ("not ok" as the positive class) in MMLU-Redux under standard prompting, Chain of Thought (CoT), few-shot demonstrations, and Retrieval-Augmented Generation (RAG) exhibit limited precision and overall reliability:

    1. Prompting Configurations (Table 3):

      • Few-Shot CoT achieves the strongest overall prompting performance across models: Claude 3 Opus achieves 48.85%48.85\% Recall, 24.03%24.03\% F1, and 40.29%40.29\% F2 score; GPT-4 Turbo achieves 46.68%46.68\% Recall, 31.68%31.68\% F1, and 36.58%36.58\% F2 score.
      • Zero-Shot CoT achieves 44.87%44.87\% Recall and 34.68%34.68\% F1 with Claude 3 Opus, but drops to 10.74%10.74\% Recall and 8.15%8.15\% F1 with Llama-3-70B.
    2. Retrieval-Augmented Generation (RAG) (Table 4):

      • Incorporating five BM25-retrieved external paragraphs from Wikipedia (enwiki-paragraphs) or MS MARCO (msmarco-v1-passage) increases error detection recall.
      • Claude 3 Opus (Zero-Shot + MS MARCO) achieves the highest error recall of 83.91%83.91\% with an F1 score of 28.09%28.09\% and F2 of 41.27%41.27\%.
      • Claude 3 Opus (Zero-Shot + Wikipedia) achieves 82.61%82.61\% Recall, 28.72%28.72\% F1, and an F2 score of 41.92%41.92\%.
      • GPT-4 Turbo (Zero-Shot + Wikipedia) yields 57.00%57.00\% Recall, 27.47%27.47\% F1, and 36.87%36.87\% F2.

    While external retrieval provides up to 83.91%83.91\% recall for error detection, F1 scores remain below 32%32\% due to high false-positive rates, demonstrating that zero-shot and few-shot LLM error auditing is insufficient for autonomous benchmark cleaning.

  7. Knowl 7 — LabelChaos Synthetic Corruption Method for Error Detection

    model/method

    To train error detection models without relying on scarce human annotations of benchmark bugs, the LabelChaos pipeline constructs a synthetic training set mirroring the MMLU error distribution. LabelChaos merges six human-annotated datasets—ARC, OpenBookQA, MathQA, PubMedQA/medical QA, and TruthfulQA—totaling approximately 264,000 samples.

    Synthetic corruptions are injected using five targeted operations:

    1. Wrong Ground Truth: The correct label index is randomly swapped with one of the incorrect choice indices.
    2. Poor Question Clarity: GPT-3.5 Turbo is prompted with few-shot MMLU examples to rewrite the question prompt to introduce vagueness, missing information, or confusing syntax.
    3. No Correct Answers: The correct choice text is replaced with meta-options such as 'all options listed' or 'all of the above'.
    4. Unclear Options: Choices are corrupted by simulating common scraping/parsing errors (e.g., splitting text strings at arbitrary character boundaries or dropping initial characters).
    5. Multiple Correct Answers: GPT-3.5 Turbo generates a semantically equivalent reformulation of the correct option, which then replaces one of the distractors.

    The resulting corpus is configured with a label distribution of 50%50\% unmodified clean items (correct) and 10%10\% for each of the five corrupted error categories.

  8. Knowl 8 — Instruction Fine-Tuning of Open-Weight Models on LabelChaos for Error Detection

    empirical result

    Instruction fine-tuning of an open-weight language model on synthetic error corruptions yields higher automated error detection performance on MMLU-Redux than few-shot prompting of larger frontier models.

    Llama-3-8B-Instruct was fine-tuned on the 264,000-sample LabelChaos dataset using Low-Rank Adaptation (LoRA, rank r=16r=16) with the AdamW optimizer (learning rate 2×10−42 \times 10^{-4}, weight decay 00, batch size 6464, 2,0482,048 steps on a single 40GB Nvidia A100 GPU).

    When evaluated on MMLU-Redux for binary classification (ok vs. not ok), the models achieved:

    • Llama-3-8B (Fine-tuned on LabelChaos): Recall = 56.58%56.58\%, F1 Score = 34.06%34.06\%, F2 Score = 44.75%44.75\%
    • Claude 3 Opus (Few-shot CoT): Recall = 48.85%48.85\%, F1 Score = 24.03%24.03\%, F2 Score = 40.29%40.29\%
    • GPT-4 Turbo (Few-shot CoT): Recall = 46.68%46.68\%, F1 Score = 31.68%31.68\%, F2 Score = 36.58%36.58\%
    • Llama-3-70B (Few-shot CoT): Recall = 24.87%24.87\%, F1 Score = 23.16%23.16\%, F2 Score = 23.10%23.10\%

    The fine-tuned 8B model surpassed prompting-based frontier models in both F1 (34.06%34.06\%) and F2 (44.75%44.75\%).

  9. Knowl 9 — Limitations of MMLU-Redux and Automated Benchmark Verification

    limitation

    The analysis and dataset present three key limitations:

    1. Incomplete MMLU Coverage: MMLU-Redux provides expert re-annotations for a subsample of 5,700 questions (100 questions per subject across 57 subjects), leaving the remaining 8,342 questions of the full 14,042-question MMLU dataset unreviewed.
    2. Annotator Subjectivity and Bias: While inter-annotator agreement is high (κ>0.64\kappa > 0.64), classifying questions with ambiguous context or regional standards (such as US-specific legal or accounting assumptions) remains vulnerable to individual annotator interpretations and cultural biases.
    3. Low Precision in Automated Detection: The best automated error detection method (Claude 3 Opus with RAG) achieves an F2 score of only 41.92%41.92\% and F1 of 28.72%28.72\%, while the fine-tuned Llama-3-8B model reaches an F2 of 44.75%44.75\% and F1 of 34.06%34.06\%. These performance levels are inadequate to replace human review for benchmark auditing.

Coverage note — None was omitted; all key contributions—the annotation taxonomy, MMLU-Redux dataset construction and statistics, Cohen's kappa agreement, LLM benchmark re-evaluation and ranking distortions, automated error detection via prompting and RAG, LabelChaos synthetic data creation and fine-tuning, and limitations—are represented in the knowls.

References

  1. 1.Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Pieter Gijsbers, Joan Giner-Miguelez, Nitisha Jain, Michael Kuchnik, Quentin Lhoest, Pierre Marcenac, Manil Maskey, Peter Mattson, Luis Oala, Pierre Ruyssen, Rajat Shinde, Elena Simperl, Goeffry Thomas, Slava Tykhonov, Joaquin Vanschoren, Jos van der Velde, Steffen Vogler, and Carole-Jean Wu. 2024. Croissant: A metadata format for ml-ready datasets. In Proceedings of the Eighth Workshop on Data Management for End-to-End Machine Learning, DEEM ’24, page 1–6, New York, NY, USA. Association for Computing Machinery.
  2. 2.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367, Minneapolis, Minnesota. Association for Computational Linguistics.
  3. 3.Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul Ronald Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, George Tucker, Enrique Piqueras, Maxim Krikun, Iain Barr, Nikolay Savinov, Ivo Danihelka, Becca Roelofs, Anaïs White, Anders Andreassen, Tamara von Glehn, Lakshman Yagati, Mehran Kazemi, Lucas Gonzalez, Misha Khalman, Jakub Sygnowski, and et al. 2023a. Gemini: A family of highly capable multimodal models. CoRR, abs/2312.11805.
  4. 4.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023b. Palm 2 technical report. CoRR, abs/2305.10403.
  5. 5.Anthropic. 2023. Anthropic. model card and evaluations for claude models.
  6. 6.AI Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card.
  7. 7.Lucas Beyer, Olivier J. Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. 2020. Are we done with imagenet? CoRR, abs/2006.07159.
  8. 8.Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating evaluation in text summarization. In EMNLP (1), pages 9347–9359. Association for Computational Linguistics.
  9. 9.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  10. 10.Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the cnn/daily mail reading comprehension task. In ACL (1). The Association for Computer Linguistics.
  11. 11.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1.
  12. 12.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1):37–46.
  13. 13.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  14. 14.Daniel Erenrich. 2023. Errors in the MMLU: The Deep Learning Benchmark is Wrong Surprisingly Often. URL: https://derenrich.medium.com/errors-in-the-mmlu-the-deep-learning-benchmark-is-wrong-surprisingly-often-7258bb045859.
  15. 15.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023. A framework for few-shot language model evaluation.
  16. 16.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In EMNLP/IJCNLP (1), pages 1161–1166. Association for Computational Linguistics.
  17. 17.Mononito Goswami, Benedikt Boecking, and Artur Dubrawski. 2021. Weak supervision for affordable modeling of electrocardiogram data. In AMIA. AMIA.
  18. 18.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In NAACL-HLT (2), pages 107–112. Association for Computational Linguistics.
  19. 19.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In ICLR. OpenReview.net.
  20. 20.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  21. 21.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  22. 22.Jean Kaddour. 2023. The minipile challenge for data-efficient language models. CoRR, abs/2304.08442.
  23. 23.Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. CoRR, abs/2307.10169.
  24. 24.Philip L and Joshua Stapleton. 2023. AI Explained YouTube Channel - SmartGPT: Major Benchmark Broken - 89.0% on MMLU + Exam’s Many Errors. URL: https://youtu.be/hVade_8H8mE?t=829.
  25. 25.Md. Tahmid Rahman Laskar, M. Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang. 2023. A systematic study and comprehensive evaluation of chatgpt on benchmark datasets. In ACL (Findings), pages 431–469. Association for Computational Linguistics.
  26. 26.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499.
  27. 27.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  28. 28.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023. Holistic evaluation of language models. Transactions on Machine Learning Research. Featured Certification, Expert Certification.
  29. 29.Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021a. Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations. In Proceedings of the 44th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2021), pages 2356–2362.
  30. 30.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021b. Truthfulqa: Measuring how models mimic human falsehoods. Preprint, arXiv:2109.07958.
  31. 31.Haochen Liu, Joseph Thekinen, Sinem Mollaoglu, Da Tang, Ji Yang, Youlong Cheng, Hui Liu, and Jiliang Tang. 2022. Toward annotator group bias in crowdsourcing. In ACL (1), pages 1797–1806. Association for Computational Linguistics.
  32. 32.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
  33. 33.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP.
  34. 34.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  35. 35.Joshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, and Olga Russakovsky. 2019. Human uncertainty makes classification more robust. In ICCV, pages 9616–9625. IEEE.
  36. 36.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In *SEM@NAACL-HLT, pages 180–191. Association for Computational Linguistics.
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the limits of transfer learning with a unified text-to-text transformer. Preprint, arXiv:1910.10683.
  38. 38.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252.
  39. 39.Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves, Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G. T. Barrett, Cathy Cheung, Basil Mustafa, Anil Palepu, Daniel McDuff, Le Hou, Tomer Golany, Luyang Liu, Jean-Baptiste Alayrac, Neil Houlsby, Nenad Tomasev, Jan Freyberg, Charles Lau, Jonas Kemp, Jeremy Lai, Shekoofeh Azizi, Kimberly Kanada, SiWai Man, Kavita Kulkarni, Ruoxi Sun, Siamak Shakeri, Luheng He, Benjamin Caine, Albert Webson, Natasha Latysheva, Melvin Johnson, Philip Andrew Mansfield, Jian Lu, Ehud Rivlin, Jesper Anderson, Bradley Green, Renee Wong, Jonathan Krause, Jonathon Shlens, Ewa Dominowska, S. M. Ali Eslami, Katherine Chou, Claire Cui, Oriol Vinyals, Koray Kavukcuoglu, James Manyika, Jeff Dean, Demis Hassabis, Yossi Matias, Dale R. Webster, Joelle K. Barral, Greg Corrado, Christopher Semturs, S. Sara Mahdavi, Juraj Gottweis, Alan Karthikesalingam, and Vivek Natarajan. 2024. Capabilities of gemini models in medicine. CoRR, abs/2404.18416.
  40. 40.Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith. 2017. The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task. In CoNLL, pages 15–25. Association for Computational Linguistics.
  41. 41.Chenhui Shen, Liying Cheng, Yang You, and Lidong Bing. 2023. Are large language models good evaluators for abstractive summarization? CoRR, abs/2305.13091.
  42. 42.Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Madeline Smith, Antoine Bosselut, Alice Oh, Andre F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, and Sara Hooker. 2024. Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation. Preprint, arXiv:2412.03304.
  43. 43.Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Sebastian Riedel, and Tim Rocktäschel. 2020. Avoiding the hypothesis-only bias in natural language inference via ensemble adversarial training. In EMNLP (1), pages 8281–8291. Association for Computational Linguistics.
  44. 44.Priyam Tejaswin, Dhruv Naik, and Pengfei Liu. 2021. How well do you know your summarization datasets? In ACL/IJCNLP (Findings), volume ACL/IJCNLP 2021 of Findings of ACL, pages 3436–3449. Association for Computational Linguistics.
  45. 45.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Bob Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  46. 46.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Preprint, arXiv:2406.01574.
  47. 47.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  48. 48.Yuxiang Wu, Matt Gardner, Pontus Stenetorp, and Pradeep Dasigi. 2022. Generating data to mitigate spurious correlations in natural language inference datasets. In ACL (1), pages 2660–2676. Association for Computational Linguistics.
  49. 49.Man-Ching Yuen, Irwin King, and Kwong-Sak Leung. 2011. A survey of crowdsourcing systems. In SocialCom/PASSAT, pages 766–773. IEEE Computer Society.
  50. 50.Jieyu Zhang, Cheng-Yu Hsieh, Yue Yu, Chao Zhang, and Alexander Ratner. 2022. A survey on programmatic weak supervision. CoRR, abs/2202.05433.

Citation

MLA
Gema, A. P., et al. “Are We Done with MMLU?”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 5069–96, https://doi.org/10.18653/v1/2025.naacl-long.262.
APA
Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., He, X., Zhao, Y., Du, X., Madani, M. R. G., Barale, C., McHardy, R., Harris, J., Kaddour, J., Krieken, E. V., & Minervini, P. (2025). Are We Done with MMLU?. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5069–5096. https://doi.org/10.18653/v1/2025.naacl-long.262
Chicago
Gema, A. P., J. O. J. Leang, G. Hong, et al. 2025. “Are We Done with MMLU?”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5069–96. https://doi.org/10.18653/v1/2025.naacl-long.262.
Harvard
Gema, A.P. et al. (2025) “Are We Done with MMLU?”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5069–5096. Available at: https://doi.org/10.18653/v1/2025.naacl-long.262.
Vancouver
1. Gema AP, Leang JOJ, Hong G, et al (2025) Are We Done with MMLU?. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5069–5096

BibTeX

@inproceedings{gema-etal-2025-done,
    title = "Are We Done with {MMLU}?",
    author = "Gema, Aryo Pradipta  and
      Leang, Joshua Ong Jun  and
      Hong, Giwon  and
      Devoto, Alessio  and
      Mancino, Alberto Carlo Maria  and
      Saxena, Rohit  and
      He, Xuanli  and
      Zhao, Yu  and
      Du, Xiaotang  and
      Ghasemi Madani, Mohammad Reza  and
      Barale, Claire  and
      McHardy, Robert  and
      Harris, Joshua  and
      Kaddour, Jean  and
      Van Krieken, Emile  and
      Minervini, Pasquale",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.262/",
    doi = "10.18653/v1/2025.naacl-long.262",
    pages = "5069--5096",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/