MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Yubo WangXueguang MaGe ZhangYuansheng NiAbhranil ChandraShiguang GuoWeiming RenAaran ArulrajXuan HeZiyan Jiang

article2024NeurIPS2,156 citations

Introduces MMLU-Pro, an upgraded benchmark that expands multiple-choice options from four to ten and prioritizes complex reasoning to resolve score saturation and reliably differentiate top-tier language models.

Listen

As artificial intelligence models advance rapidly, existing standard evaluation benchmarks have begun to saturate. Frontier systems have clustered tightly at near-ceiling performance levels, making it increasingly difficult for organizations to accurately measure progress, differentiate competitive models, or identify critical reasoning weaknesses. Furthermore, established benchmarks suffer from excessive noise and vulnerability to minor prompt phrasing changes, which distorts leaderboard rankings and creates uncertainty for technology leaders evaluating deployment options.

This article introduces and evaluates MMLU-Pro, an enhanced benchmark designed to test expert-level reasoning across diverse academic and professional disciplines. The initiative set out to create a more demanding, discriminative, and robust evaluation standard by addressing the structural shortcomings of the widely used Massive Multitask Language Understanding (MMLU) benchmark.

To construct MMLU-Pro, the authors curated 12,032 questions spanning 14 disciplines by filtering out overly simple and erroneous items from the original benchmark and integrating advanced university-level problems from external scientific and theorem datasets. To substantially lower the chance of successful guessing and test deeper comprehension, the authors expanded the multiple-choice format from four options to ten options per question, using model-assisted generation followed by two phases of rigorous expert human review. The authors then evaluated more than 50 leading proprietary and open-source models using a five-shot Chain-of-Thought prompting approach.

Key findings demonstrate that MMLU-Pro successfully restores benchmark difficulty and separation among top systems. First, overall model accuracy dropped sharply by 16% to 33% compared to the original benchmark; the highest-performing model, GPT-4o, achieved only 72.6% accuracy, confirming substantial headroom for future development. Second, the benchmark provides much greater differentiation: the performance gap between top-tier models expanded from a negligible 1% on the original benchmark to 9% on MMLU-Pro. Third, the benchmark proves significantly more robust against prompt variations, reducing score volatility across 24 distinct prompt styles from 4–5% down to approximately 2%. Fourth, unlike the original test where direct answering often yielded equal or better results, models achieved marked gains on MMLU-Pro when forced to reason step-by-step—boosting GPT-4o by 19.1%. Finally, an error analysis of the leading model revealed that 39% of its mistakes stemmed from reasoning failures, 35% from lack of specialized domain knowledge, and 12% from computational errors.

These findings indicate that prior assessments have overestimated the reasoning capabilities of leading models due to simpler formats and random guessing advantages. For technology leaders, MMLU-Pro provides a more dependable gauge of operational readiness for high-stakes tasks in complex fields like engineering, law, physics, and mathematics. The results show that open-source models are closing the gap with mid-tier commercial offerings, though top-tier proprietary systems still hold a clear advantage in multi-step problem solving.

Organizations evaluating or deploying advanced language models should adopt MMLU-Pro alongside their current testing suites to gain clearer insight into true reasoning capabilities and reduce benchmark prompt sensitivity. AI developers should prioritize enhancements in logical consistency, domain-specific knowledge integration, and external tool integration (such as calculators or code execution) to mitigate common computational and reasoning bottlenecks.

The benchmark remains subject to the inherent limitations of multiple-choice formats, which do not fully capture open-ended, creative real-world problem solving, and it does not currently evaluate multi-modal inputs such as diagrams or charts. Nevertheless, the rigorous multi-stage expert curation and extensive empirical validation across 50 models provide high confidence in MMLU-Pro as a robust and reliable evaluation standard for modern language models.

Cover for MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Abstract

In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Language Models
  • 2.2 LLMs Evaluation Benchmarks
  • 3 The MMLU-Pro Benchmark
  • 3.1 Overview
  • 3.2 Dataset Construction Pipeline
  • 4 Experimental Setup
  • 5 Results and Analysis
  • 5.1 Overall Performance
  • 5.2 Subject-Specific Insights
  • 5.3 Error Analysis
  • 6 Comparison with MMLU
  • 6.1 Difficulty Level
  • 6.2 Reasoning Level
  • 6.3 Robustness Degree
  • 7 Limitations
  • 8 Conclusion
  • References
  • A Appendix
  • A.1 Dataset Construction Details
  • A.2 5-shot CoT Prompt example
  • A.3 List of Language Models Studied
  • A.4 Computational Resources
  • A.5 Dataset Licensing
  • A.6 Error Analysis Cases

Knowls

  1. Knowl 1 — MMLU-Pro Benchmark Dataset Specification

    definition

    MMLU-Pro is a multi-discipline language understanding and reasoning benchmark consisting of 12,032 multiple-choice questions spanning 14 subject categories: Mathematics (1,351 questions, 11.23%), Physics (1,299 questions, 10.80%), Chemistry (1,132 questions, 9.41%), Law (1,101 questions, 9.15%), Engineering (969 questions, 8.05%), Other (924 questions, 7.68%), Economics (844 questions, 7.01%), Health (818 questions, 6.80%), Psychology (798 questions, 6.63%), Business (789 questions, 6.56%), Biology (717 questions, 5.96%), Philosophy (499 questions, 4.15%), Computer Science (410 questions, 3.41%), and History (381 questions, 3.17%).

    Unlike traditional 4-option multiple-choice benchmarks, MMLU-Pro expands the choice options up to 10 choices (AA through JJ). Across the dataset, 83% of questions contain 10 options, 17% contain fewer options, and the average number of options per question is 9.47. The questions are integrated from four sources: filtered original MMLU questions (56.60%), STEM website problems (33.93%), TheoremQA (4.97%), and SciBench (4.50%).

  2. Knowl 2 — MMLU-Pro Dataset Construction and Verification Pipeline

    model/method

    The construction pipeline for MMLU-Pro consists of four sequential stages:

    1. Initial Filtering: The 57 subjects of the original MMLU benchmark are consolidated into 14 broader disciplines. Overly simple questions are eliminated by evaluating eight open-source language models: Llama-2-7B, Llama-2-7B-Chat, Llama-2-13B, Llama-2-13B-Chat, Mistral-7B, Gemma-7B, Yi-6B, and Yi-6B-Chat. Any question correctly answered by more than four models is considered trivial and removed, eliminating 5,886 questions (42.23% of original MMLU questions).

    2. Question Collection and Integration: Challenging college-level exam problems and theorem-driven questions are collected from online STEM sources, TheoremQA, and SciBench. Short answers and three initial distractors are extracted or synthesized using GPT-4-Turbo (gpt-4-turbo-2024-04-09), followed by manual verification against reference solutions to prune incomplete or flawed extractions.

    3. Option Augmentation: GPT-4-Turbo is prompted to generate six additional plausible distractors per question, expanding the choice pool from four to ten options to reduce random guessing probability from 25% to 10%.

    4. Two-Phase Expert Review:

    • Phase 1 (Verification of Correctness and Appropriateness): Human subject experts verify answer correctness, discard questions requiring non-textual modalities (images, diagrams, tables), and remove questions lacking sufficient textual information or unsuitable for multiple-choice formats (proofs, open-ended tasks).
    • Phase 2 (Distractor Validity): Gemini-1.5-Pro re-evaluates all answer choices to flag potential false negatives (distractors that are accidentally correct). Human experts review all flagged options, removing invalid distractors and retaining verified ones.
  3. Knowl 3 — Large Language Model Performance on MMLU-Pro

    data/table

    Evaluation of closed-source and open-source models using 5-shot Chain-of-Thought (CoT) prompting (0-shot for Gemini-1.5-Pro and Gemini-1.5-Flash) shows that MMLU-Pro introduces substantial difficulty and broader discriminative headroom across disciplines compared to MMLU:

    Model Overall Math Physics Engineering History Law Psychology
    Closed-Source Models
    GPT-4o 72.6 76.1 74.7 55.0 70.1 51.0 79.2
    Gemini-1.5-Pro 69.0 72.8 70.4 48.7 65.6 50.8 77.2
    Claude-3-Opus 68.5 69.6 69.7 48.4 61.4 53.5 76.3
    GPT-4-Turbo 63.7 62.8 61.0 35.9 67.7 51.2 78.3
    Gemini-1.5-Flash 59.1 59.6 61.2 44.2 53.8 37.3 70.1
    Yi-Large 58.1 64.8 57.0 45.4 49.6 36.2 50.6
    Claude-3-Sonnet 56.8 49.0 53.1 40.5 57.2 42.7 72.2
    Open-Source Models
    Llama-3-70B-Instruct 56.2 54.0 49.6 43.6 56.9 39.9 70.2
    Phi-3-medium-4k-instruct 55.7 52.2 49.4 37.9 57.2 38.3 73.4
    DeepSeek-V2-Chat 54.8 53.7 54.0 31.9 45.3 40.6 66.2
    Llama-3-70B 52.8 49.7 49.8 35.0 57.7 35.0 71.4
    Qwen1.5-72B-Chat 52.6 52.3 44.2 36.6 55.9 38.5 67.7
    Yi-1.5-34B-Chat 52.3 56.2 49.4 34.4 52.8 34.8 64.3
    MAmmoTH2-8x7B-Plus 50.4 50.3 45.7 34.0 50.9 35.5 63.8
    Qwen1.5-110B 49.9 50.4 41.4 35.3 54.1 35.1 66.3
    Phi-3-mini-4k-instruct 45.7 41.8 41.0 28.7 41.5 28.5 65.2
    Mixtral-8x7B-Instruct-v0.1 43.3 36.3 39.9 29.2 44.6 32.1 63.4
    Yi-34B 43.0 31.8 35.0 32.6 52.0 32.7 62.5
    Mixtral-8x7B-v0.1 41.0 34.1 37.2 27.9 47.5 27.1 61.0
    Llama-3-8B-Instruct 41.0 36.1 34.4 31.3 42.3 26.5 59.4
    Starling-7B 37.9 34.9 38.5 27.0 43.6 24.7 32.5
    c4ai-command-r-v01 37.9 26.3 28.3 24.8 47.5 34.0 58.5
    Llama-2-70B 37.5 26.8 28.2 23.5 45.9 28.6 59.0
    OpenChat-3.5-8B 37.2 36.2 30.5 26.9 39.9 24.6 54.5
    InternMath-20B-Plus 37.1 56.1 24.0 30.4 20.5 15.2 42.3
    Llama3-Smaug-8B 36.9 33.2 37.3 19.8 42.0 26.5 28.6
    Llama-3-8B 35.4 30.4 31.4 25.5 36.2 19.6 53.3
    Gemma-7B 33.7 25.1 27.6 22.7 36.8 21.7 51.8
    InternMath-7B-Plus 33.5 48.3 22.8 28.2 19.2 14.4 38.3
    Zephyr-7B-Beta 33.0 23.6 35.7 23.9 32.0 22.0 28.2
    Mistral-7B-v0.1 30.9 23.5 24.8 22.4 32.6 20.7 48.9
    Neo-7B-Instruct 28.7 35.5 23.7 19.1 28.2 18.0 36.2
    Llemma-7B 23.5 21.6 25.7 23.8 15.2 14.8 29.6
    Gemma-2B 15.9 16.3 15.6 12.7 15.4 12.3 16.1

    The top-performing model, GPT-4o, achieves 72.6% accuracy on MMLU-Pro (leaving 27.4% room for improvement), compared to 88.7% on MMLU. The performance gap between GPT-4o and GPT-4-Turbo widens from ~1.5% on MMLU to 8.9% on MMLU-Pro.

  4. Knowl 4 — Chain-of-Thought Prompting Advantage on MMLU-Pro vs. MMLU

    empirical result

    On MMLU-Pro, Chain-of-Thought (CoT) prompting provides substantial accuracy gains over Direct Answering (DA), whereas on original MMLU, CoT provides minimal benefit or harms accuracy:

    Model Name MMLU MMLU-Pro
    CoT Direct Answer CoT - DA CoT Direct Answer CoT - DA
    GPT-4o 88.7 87.2 +1.5 72.6 53.5 +19.1
    GPT-4-Turbo 86.5 86.7 -0.2 63.7 48.4 +15.3
    Phi-3-medium-4k-instruct 79.4 78.0 +1.4 55.7 47.5 +8.2
    Llama-3-8B 62.7 66.6 -3.9 35.4 31.5 +3.9
    Gemma-7B 62.4 66.0 -3.6 33.7 27.0 +6.7

    On MMLU-Pro, GPT-4o improves by 19.1% and GPT-4-Turbo improves by 15.3% when utilizing CoT reasoning compared to direct answer generation, reflecting the benchmark's emphasis on multi-step analytical reasoning over pure factual memorization.

  5. Knowl 5 — Prompt Sensitivity and Scoring Robustness

    empirical result

    Evaluating large language models across 24 distinct prompt templates demonstrates that MMLU-Pro substantially reduces scoring variance compared to original MMLU.

    On MMLU, prompt modifications induce an average performance fluctuation of 4% to 5% across models, with peak variations reaching 10.98%. On MMLU-Pro, model accuracy variations across the 24 prompts decrease to an average of approximately 2%, with a maximum variation of 3.74%. Expanding the answer options from 4 to 10 lowers random guessing probability and reduces reliance on prompt phrasing heuristics, providing more stable model rankings.

  6. Knowl 6 — 5-Shot Chain-of-Thought and Regex Answer Extraction Protocol

    experimental setup

    The evaluation protocol on MMLU-Pro uses 5-shot Chain-of-Thought (CoT) prompting. The prompt consists of task instructions, five discipline-specific demonstration exemplars sampled from the validation split of MMLU-Pro, and the input test question.

    Candidate answers are extracted from the model reasoning outputs using a tiered extraction hierarchy:

    1. Primary Regex: The pattern answer is \(?([A-J])\)? matches the standard final sentence specified in the few-shot demonstrations.
    2. Secondary Regex: If the primary regex fails, the fallback pattern \.*[aA]nswer:\s*([A-J]) is applied.
    3. Random Fallback: If both regular expressions fail to extract an option letter, a choice is randomly sampled from the available options for that question.
  7. Knowl 7 — Error Distribution and Failure Modes of Frontier Models on MMLU-Pro

    empirical result

    An expert analysis of 120 randomly sampled erroneous predictions made by GPT-4o on MMLU-Pro classifies the failure causes into the following distribution:

    • Reasoning Errors (39%): Flaws in logical deduction, step ordering, or physical interactions despite recalling relevant domain concepts (e.g., adding rather than subtracting internal and external pressures when finding the net force on a piston).
    • Lack of Specific Knowledge (35%): Inaccurate or missing specialized formulas, domain facts, or operational definitions (e.g., subtracting refractive indices directly instead of calculating the ratio of lens and medium refractive indices in lens-maker equations, or failing to deduct down payments prior to computing loan interest balances).
    • Calculation Errors (12%): Arithmetic or algebraic calculation errors despite setting up the correct analytical formula.
    • No Selection Made (5%): Responses where the model generates an analysis but fails to output a final choice letter.
    • Question Understanding Errors (4%): Misinterpreting problem scope or constraints.
    • Generation Issues (2%): Infinite generation loops repeating a single phrase until exceeding maximum context length.
    • Annotation Errors (2%): Ground truth label errors in the benchmark data.
    • Answer Extraction Errors (1%): Regex extraction failures caused by non-standard model answer formatting.
  8. Knowl 8 — Limitations of the MMLU-Pro Benchmark

    limitation

    MMLU-Pro is subject to two main constraints:

    1. Multiple-Choice Question (MCQ) Format: Constraining evaluations to multiple-choice options limits the assessment of open-ended generation, creative formulation, and detailed qualitative explanations that characterize real-world expert tasks.
    2. Text-Only Modality: MMLU-Pro exclusively evaluates textual understanding and reasoning, discarding all questions requiring diagrams, figures, images, or multi-modal inputs.

Coverage note — None was omitted; all primary benchmark contributions, construction stages, comparative empirical results, prompt robustness evaluations, and error taxonomies are covered.

References

  1. 1.Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  3. 3.Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, et al. When benchmarks are targets: Revealing the sensitivity of large language model leaderboards. arXiv preprint arXiv:2402.01781, 2024.
  4. 4.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631, 2023.
  5. 5.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  6. 6.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  8. 8.c4ai-command-r-v01. C4ai-command-r-v01. https://huggingface.co/cohereforai/c4ai-command-r-v01. URL https://huggingface.co/CohereForAI/c4ai-command-r-v01.
  9. 9.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.
  10. 10.Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. Theoremqa: A theorem-driven question answering dataset. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  11. 11.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024.
  12. 12.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  13. 13.claude. Introducing the next generation of Claude https://www.anthropic.com/news/claude-3-family. URL https://www.anthropic.com/news/claude-3-family.
  14. 14.OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
  15. 15.DeepSeek-AI. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.
  16. 16.Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance. arXiv preprint arXiv:2305.17306, 2023.
  17. 17.gpt-4o. Hello gpt4-o. https://openai.com/index/hello-gpt-4o/. URL https://openai.com/index/hello-gpt-4o/.
  18. 18.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  19. 19.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  20. 20.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  21. 21.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  22. 22.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  23. 23.lingyiwanwu. Lingyiwanwu, yi-large. https://www.lingyiwanwu.com/en. URL https://www.lingyiwanwu.com/en.
  24. 24.Meta Llama 3. Build the future of ai with meta llama 3 - https://llama.meta.com/llama3/. URL https://llama.meta.com/llama3/.
  25. 25.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. Cross-task generalization via natural language crowdsourcing instructions. arXiv preprint arXiv:2104.08773, 2021.
  26. 26.Meredith Ringel Morris, Jascha Sohl-dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. Levels of agi: Operationalizing progress on the path to agi. arXiv preprint arXiv:2311.02462, 2023.
  27. 27.open llm leaderboard. Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard. URL https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard.
  28. 28.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  29. 29.Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White. Smaug: Fixing failure modes of preference optimisation with dpo-positive. arXiv preprint arXiv:2402.13228, 2024.
  30. 30.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  31. 31.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  32. 32.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  33. 33.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  34. 34.Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024.
  35. 35.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  36. 36.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al. Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944, 2023.
  37. 37.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  38. 38.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32, 2019.
  39. 39.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. Open-chat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:2309.11235, 2023.
  40. 40.Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. Scibench: Evaluating college-level scientific problem-solving abilities of large language models. arXiv preprint arXiv:2307.10635, 2023.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  42. 42.Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, Yudong Wang, Zijian Wu, Shuaibin Li, Fengzhe Zhou, Hongwei Liu, Songyang Zhang, Wenwei Zhang, Hang Yan, Xipeng Qiu, Jiayu Wang, Kai Chen, and Dahua Lin. Internlm-math: Open math large language models toward verifiable reasoning, 2024.
  43. 43.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024.
  44. 44.Xiang Yue, Tuney Zheng, Ge Zhang, and Wenhu Chen. Mammoth2: Scaling instructions from the web. arXiv preprint arXiv:2405.03548, 2024.
  45. 45.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  46. 46.Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, Raven Yuan, Tuney Zheng, Wei Pang, Xinrun Du, Yiming Liang, Yinghao Ma, Yizhi Li, Ziyang Ma, Bill Lin, Emmanouil Benetos, Huan Yang, Junting Zhou, Kaijing Ma, Minghao Liu, Morry Niu, Noah Wang, Quehry Que, Ruibo Liu, Sine Liu, Shawn Guo, Soren Gao, Wangchunshu Zhou, Xinyue Zhang, Yizhi Zhou, Yubo Wang, Yuelin Bai, Yuhan Zhang, Yuxiang Zhang, Zenith Wang, Zhenzhu Yang, Zijian Zhao, Jiajun Zhang, Wanli Ouyang, Wenhao Huang, and Wenhu Chen. Map-neo: Highly capable and transparent bilingual large language model series, 2024.
  47. 47.Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations, 2023.
  48. 48.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364, 2023.
  49. 49.Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao. Starling-7b: Improving llm helpfulness and harmlessness with rlaif, November 2023.

Citation

MLA
Wang, Y., et al. “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark”. arXiv, 2024, http://arxiv.org/abs/2406.01574v6.
APA
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., & Chen, W. (2024). MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv. http://arxiv.org/abs/2406.01574v6
Chicago
Wang, Y., X. Ma, G. Zhang, et al. 2024. “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark”. arXiv. http://arxiv.org/abs/2406.01574v6.
Harvard
Wang, Y. et al. (2024) “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.01574v6.
Vancouver
1. Wang Y, Ma X, Zhang G, et al (2024) MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv

BibTeX

@article{wang2024mmlu,
  title = {MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark},
  author = {Wang, Yubo and Ma, Xueguang and Zhang, Ge and Ni, Yuansheng and Chandra, Abhranil and Guo, Shiguang and Ren, Weiming and Arulraj, Aaran and He, Xuan and Jiang, Ziyan and Li, Tianle and Ku, Max and Wang, Kai and Zhuang, Alex and Fan, Rongqi and Yue, Xiang and Chen, Wenhu},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.01574v6},
  eprint = {2406.01574}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/