MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Yubo WangXueguang MaGe ZhangYuansheng NiAbhranil ChandraShiguang GuoWeiming RenAaran ArulrajXuan HeZiyan Jiang
Introduces MMLU-Pro, an upgraded benchmark that expands multiple-choice options from four to ten and prioritizes complex reasoning to resolve score saturation and reliably differentiate top-tier language models.
As artificial intelligence models advance rapidly, existing standard evaluation benchmarks have begun to saturate. Frontier systems have clustered tightly at near-ceiling performance levels, making it increasingly difficult for organizations to accurately measure progress, differentiate competitive models, or identify critical reasoning weaknesses. Furthermore, established benchmarks suffer from excessive noise and vulnerability to minor prompt phrasing changes, which distorts leaderboard rankings and creates uncertainty for technology leaders evaluating deployment options.
This article introduces and evaluates MMLU-Pro, an enhanced benchmark designed to test expert-level reasoning across diverse academic and professional disciplines. The initiative set out to create a more demanding, discriminative, and robust evaluation standard by addressing the structural shortcomings of the widely used Massive Multitask Language Understanding (MMLU) benchmark.
To construct MMLU-Pro, the authors curated 12,032 questions spanning 14 disciplines by filtering out overly simple and erroneous items from the original benchmark and integrating advanced university-level problems from external scientific and theorem datasets. To substantially lower the chance of successful guessing and test deeper comprehension, the authors expanded the multiple-choice format from four options to ten options per question, using model-assisted generation followed by two phases of rigorous expert human review. The authors then evaluated more than 50 leading proprietary and open-source models using a five-shot Chain-of-Thought prompting approach.
Key findings demonstrate that MMLU-Pro successfully restores benchmark difficulty and separation among top systems. First, overall model accuracy dropped sharply by 16% to 33% compared to the original benchmark; the highest-performing model, GPT-4o, achieved only 72.6% accuracy, confirming substantial headroom for future development. Second, the benchmark provides much greater differentiation: the performance gap between top-tier models expanded from a negligible 1% on the original benchmark to 9% on MMLU-Pro. Third, the benchmark proves significantly more robust against prompt variations, reducing score volatility across 24 distinct prompt styles from 4–5% down to approximately 2%. Fourth, unlike the original test where direct answering often yielded equal or better results, models achieved marked gains on MMLU-Pro when forced to reason step-by-step—boosting GPT-4o by 19.1%. Finally, an error analysis of the leading model revealed that 39% of its mistakes stemmed from reasoning failures, 35% from lack of specialized domain knowledge, and 12% from computational errors.
These findings indicate that prior assessments have overestimated the reasoning capabilities of leading models due to simpler formats and random guessing advantages. For technology leaders, MMLU-Pro provides a more dependable gauge of operational readiness for high-stakes tasks in complex fields like engineering, law, physics, and mathematics. The results show that open-source models are closing the gap with mid-tier commercial offerings, though top-tier proprietary systems still hold a clear advantage in multi-step problem solving.
Organizations evaluating or deploying advanced language models should adopt MMLU-Pro alongside their current testing suites to gain clearer insight into true reasoning capabilities and reduce benchmark prompt sensitivity. AI developers should prioritize enhancements in logical consistency, domain-specific knowledge integration, and external tool integration (such as calculators or code execution) to mitigate common computational and reasoning bottlenecks.
The benchmark remains subject to the inherent limitations of multiple-choice formats, which do not fully capture open-ended, creative real-world problem solving, and it does not currently evaluate multi-modal inputs such as diagrams or charts. Nevertheless, the rigorous multi-stage expert curation and extensive empirical validation across 50 models provide high confidence in MMLU-Pro as a robust and reliable evaluation standard for modern language models.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). This paper introduces the original MMLU benchmark, which MMLU-Pro directly critiques, redesigns, and extends to address saturation and lack of reasoning depth.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational work introduces Chain-of-Thought prompting, establishing the reasoning elicitation paradigm that MMLU-Pro explicitly benchmarks and evaluates against standard prompting.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). This work isolates challenging benchmark tasks to analyze the differential impact of Chain-of-Thought reasoning versus direct answering, directly motivating MMLU-Pro's methodology.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). This paper introduces self-consistency over reasoning chains, providing core conceptual background on prompt sensitivity and reasoning stability in multi-step evaluations.
- Paper: SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, Alex Wang et al. (2019). This paper provides the foundational precedent for upgrading saturated general-purpose language understanding benchmarks with more demanding multi-task challenges.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey establishes the landscape of LLM evaluation frameworks and details the saturation and robustness challenges that motivated next-generation datasets like MMLU-Pro.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). This work investigates advanced System 2 reasoning methods and search-based deliberation to solve the complex reasoning problems highlighted by benchmarks such as MMLU-Pro.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). This survey examines the post-training reinforcement learning techniques and test-time compute scaling used to tackle rigorous multi-step reasoning evaluations.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). This paper develops automated pipelines to construct discriminative, hard benchmarks that address benchmark saturation in language models.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). This study analyzes prompt sensitivity and performance variance under symbolic perturbations, complementing MMLU-Pro's findings on benchmark robustness and prompt stability.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). This paper demonstrates test-time compute scaling on complex reasoning questions, providing practical methods to improve performance on demanding benchmarks like MMLU-Pro.
- Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). This primer analyzes how reasoning trajectories and verifiable post-training data drive model improvements on challenging multi-step benchmarks.
