MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Xiang YueTianyu ZhengYuansheng NiYubo WangKai ZhangShengbang TongYuxuan SunBotao YuGe ZhangHuan Sun
Introduces MMMU-Pro, an advanced multimodal benchmark that eliminates text-only shortcuts, expands answer choices, and embeds text within images to expose severe performance drops in leading vision-language models.
Recent advancements in multimodal artificial intelligence have led systems to achieve high scores on standardized benchmarks combining visual and textual data. However, standard evaluations often allow models to exploit statistical shortcuts, text-only correlations, and narrow multiple-choice options without demonstrating genuine multimodal comprehension. The article addresses this critical gap by evaluating whether current multimodal systems possess true deliberate reasoning capabilities across diverse academic disciplines, introducing a more robust benchmark named MMMU-Pro.
To construct this benchmark, the researchers implemented a three-stage approach utilizing college-level exam and textbook questions across 30 subjects. First, they used multiple advanced language models to filter out questions answerable purely through text without visual information. Second, they expanded the candidate answer choices from four to ten options through assisted generation and two rounds of human expert validation, significantly reducing the probability of successful guessing. Third, they established a realistic vision-only evaluation setting where questions, charts, and choices were embedded directly into screenshots and photographs with varied display conditions, creating a final evaluation suite of 3,460 questions.
The findings show that current state-of-the-art multimodal systems experience significant performance drops under these rigorous conditions, with accuracy decreasing across all evaluated models by roughly 17% to 27% compared to the original benchmark. Top-performing proprietary models that previously reached near 70% accuracy fell to around 44% to 54% on the expanded-option and vision-only formats, while several open-source models suffered even steeper declines. The analysis revealed that strong optical character recognition alone is insufficient for success, as models frequently transcribed embedded text correctly but still failed to synthesize the information. Additionally, structured step-by-step reasoning prompts improved model accuracy in technical and scientific domains but provided minimal or negative benefits in subjective fields such as art and design.
These results indicate that conventional benchmarks substantially overestimate the reliability and reasoning proficiency of multimodal models in real-world applications. When textual and visual cues are intertwined in practical formats like screenshots, models face increased visual processing demands, struggle with context switching between modalities, and frequently commit complex reasoning errors. For organizations deploying artificial intelligence, relying on standard benchmark scores risks deploying systems that fail unpredictably in operational settings involving mixed visual and textual workflows.
The article recommends that future artificial intelligence development focus on scaling underlying model architectures, adopting self-supervised vision encoders that learn richer visual features, generating specialized step-by-step training data, and training models on synthetic text-rich images. While the benchmark provides a more reliable evaluation standard, readers should note that human expert comparisons were approximated from historical data and that closed multiple-choice formats do not fully capture open-ended real-world complexity. Decision-makers should therefore exercise caution and conduct practical pilot testing before deploying multimodal systems in high-stakes reasoning environments.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). This paper establishes the original MMMU benchmark for college-level multidisciplinary multimodal reasoning that MMMU-Pro directly refines and hardens.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). This work introduces the option-expansion and benchmark-hardening methodology ('-Pro' paradigm) that MMMU-Pro translates into the multimodal domain.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). This paper demonstrates how standard multimodal benchmarks overestimate model capabilities due to text-solvable shortcuts, motivating MMMU-Pro's rigorous filtering and Chain-of-Thought evaluations.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). This work analyzes option bias and robustness issues in vision-language multiple-choice evaluation, establishing foundational techniques for robust benchmark design.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). This technical report evaluates next-generation multimodal architectures across rigorous multidisciplinary benchmarks, demonstrating practical model improvements on robust evaluation standards.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This report presents a frontier vision-language model family developed and tested against advanced visual-textual reasoning and OCR-integrated benchmark settings.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). This paper extends the paradigm of massive multidisciplinary expert reasoning benchmarks from the visual-language domain to complex multi-task audio comprehension.
