MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Xiang YueTianyu ZhengYuansheng NiYubo WangKai ZhangShengbang TongYuxuan SunBotao YuGe ZhangHuan Sun

article2025ACL579 citations

Introduces MMMU-Pro, an advanced multimodal benchmark that eliminates text-only shortcuts, expands answer choices, and embeds text within images to expose severe performance drops in leading vision-language models.

Listen

Recent advancements in multimodal artificial intelligence have led systems to achieve high scores on standardized benchmarks combining visual and textual data. However, standard evaluations often allow models to exploit statistical shortcuts, text-only correlations, and narrow multiple-choice options without demonstrating genuine multimodal comprehension. The article addresses this critical gap by evaluating whether current multimodal systems possess true deliberate reasoning capabilities across diverse academic disciplines, introducing a more robust benchmark named MMMU-Pro.

To construct this benchmark, the researchers implemented a three-stage approach utilizing college-level exam and textbook questions across 30 subjects. First, they used multiple advanced language models to filter out questions answerable purely through text without visual information. Second, they expanded the candidate answer choices from four to ten options through assisted generation and two rounds of human expert validation, significantly reducing the probability of successful guessing. Third, they established a realistic vision-only evaluation setting where questions, charts, and choices were embedded directly into screenshots and photographs with varied display conditions, creating a final evaluation suite of 3,460 questions.

The findings show that current state-of-the-art multimodal systems experience significant performance drops under these rigorous conditions, with accuracy decreasing across all evaluated models by roughly 17% to 27% compared to the original benchmark. Top-performing proprietary models that previously reached near 70% accuracy fell to around 44% to 54% on the expanded-option and vision-only formats, while several open-source models suffered even steeper declines. The analysis revealed that strong optical character recognition alone is insufficient for success, as models frequently transcribed embedded text correctly but still failed to synthesize the information. Additionally, structured step-by-step reasoning prompts improved model accuracy in technical and scientific domains but provided minimal or negative benefits in subjective fields such as art and design.

These results indicate that conventional benchmarks substantially overestimate the reliability and reasoning proficiency of multimodal models in real-world applications. When textual and visual cues are intertwined in practical formats like screenshots, models face increased visual processing demands, struggle with context switching between modalities, and frequently commit complex reasoning errors. For organizations deploying artificial intelligence, relying on standard benchmark scores risks deploying systems that fail unpredictably in operational settings involving mixed visual and textual workflows.

The article recommends that future artificial intelligence development focus on scaling underlying model architectures, adopting self-supervised vision encoders that learn richer visual features, generating specialized step-by-step training data, and training models on synthetic text-rich images. While the benchmark provides a more reliable evaluation standard, readers should note that human expert comparisons were approximated from historical data and that closed multiple-choice formats do not fully capture open-ended real-world complexity. Decision-makers should therefore exercise caution and conduct practical pilot testing before deploying multimodal systems in high-stakes reasoning environments.

arXiv: 2409.02813
Cover for MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Abstract

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models’ true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly “see” and “read” simultaneously, testing a core human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future multimodal research.

Table of Contents

  • 1 Introduction
  • 2 MMMU-Pro: A More Robust Version of MMMU
  • 2.1 Revisiting the MMMU Benchmark
  • 2.2 Methods
  • 3 Experiments
  • 3.1 Experimental Setups
  • 3.2 Overall Results
  • 3.3 Impact of CoT Prompting
  • 3.4 Does OCR Help in the Vision Setting?
  • 3.5 Qualitative Analysis
  • 3.6 Error Analysis
  • 3.7 Response Length Comparison
  • 4 Guide for Future Model Training
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Evaluation Prompts
  • B Approximating Human Expert Performance
  • C Ensuring Quality and Diversity of Expanded Options
  • D Analysis of CoT’s Impact
  • E Experimental Setup of Vision Encoder Impact
  • F Comparison of GPT-4o’s responses between Standard and Vision Input settings
  • G CoT vs. Direct Acc: Model Differences Across Disciplines
  • H Comparison With and Without Augmented Options
  • I Comparison of Model Outputs Across Different Input Modes
  • J Qualitative Examples
  • J.1 Art and Design: Art
  • J.2 Art and Design: Art Theory
  • J.3 Art and Design: Design
  • J.4 Art and Design: Music
  • J.5 Business: Accounting
  • J.6 Business: Economics
  • J.7 Business: Finance
  • J.8 Business: Manage
  • J.9 Business: Marketing
  • J.10 Science: Biology
  • J.11 Science: Chemistry
  • J.12 Science: Geography
  • J.13 Science: Math
  • J.14 Science: Physics
  • J.15 Health and Medicine: Basic Medical Science
  • J.16 Health and Medicine: Clinical Medicine
  • J.17 Health and Medicine: Diagnostics and Laboratory Medicine
  • J.18 Health and Medicine: Pharmacy
  • J.19 Health and Medicine: Public Health
  • J.20 Humanities and Social Science: History
  • J.21 Humanities and Social Science: Literature
  • J.22 Humanities and Social Science: Sociology
  • J.23 Humanities and Social Science: Psychology
  • J.24 Tech and Engineering: Agriculture
  • J.25 Tech and Engineering: Architecture and Engineering
  • J.26 Tech and Engineering: Computer Science
  • J.27 Tech and Engineering: Electronics
  • J.28 Tech and Engineering: Energy and Power
  • J.29 Tech and Engineering: Materials
  • J.30 Tech and Engineering: Mechanical Engineering

Knowls

  1. Knowl 1 — MMMU-Pro Benchmark Construction Methodology

    model/method

    MMMU-Pro is a benchmark designed to evaluate multimodal large language models (MLLMs) on multi-discipline, college-level reasoning tasks while eliminating text-only shortcuts and guessing vulnerabilities present in the original MMMU dataset. The construction process follows a three-step pipeline:

    1. Text-Only LLM Filtering: To remove questions solvable without visual inputs, four large text-only language models (Llama-3-70B-Instruct, Qwen2-72B-Instruct, Yi-1.5-34B-Chat, and Mixtral-8x22B-Instruct) are prompted with MMMU questions without access to the corresponding images. Each model answers each question across 10 repeated trials. A question is considered answerable if a model answers it correctly in more than 5 trials. Any question where at least three of the four models answer correctly across the majority of trials is eliminated. From the remaining pool, 1,800 questions are sampled evenly across 30 subjects (60 questions per subject).

    2. Candidate Option Augmentation: The number of candidate options for each multiple-choice question is expanded from 4 to 10 to suppress random guessing and heuristic elimination. Candidate distractors are generated using GPT-4o, filtered for consistency by Claude 3.5, and subjected to two rounds of human expert validation to refine choices and eliminate questions lacking image dependency. This stage eliminates 70 questions, yielding 1,730 core questions.

    3. Vision-Only Input Setting: Questions and candidate options are embedded directly into visual images (as screenshots or photographs) across simulated display environments featuring varied backgrounds, font styles, and font sizes, presenting models with purely visual inputs rather than providing text in separate input prompt tokens.

    The resulting benchmark contains 3,460 total test instances: 1,730 questions in the Standard (10-option) setting and 1,730 questions in the Vision-Only setting. The overall MMMU-Pro benchmark score is defined as the arithmetic mean of performance in the Standard (10-option) and Vision-Only settings.

  2. Knowl 2 — Model Performance Breakdown on MMMU-Pro Benchmark

    data/table

    Evaluating multimodal large language models on MMMU-Pro demonstrates sharp accuracy declines relative to the original MMMU validation set (4 options with separate text and image inputs). Performance drops occur both when expanding candidate options (Δ1=Standard10 Opts−MMMUVal\Delta_1 = \text{Standard}_{10\text{ Opts}} - \text{MMMU}_{\text{Val}}) and when presenting inputs in the vision-only image format (Δ2=Vision−MMMUVal\Delta_2 = \text{Vision} - \text{MMMU}_{\text{Val}}).

    Model / Baseline MMMU (Val) Standard (4 Opts) Standard (10 Opts) Vision Δ1\mathbf{\Delta_1} Δ2\mathbf{\Delta_2}
    Random Choice 22.1 24.9 12.8 12.4 -9.3 -9.7
    Frequent Choice 26.8 27.8 12.1 12.1 -14.7 -14.7
    Human Expert (Low) 76.2 75.4 73.0 73.0 -3.2 -3.2
    Human Expert (Medium) 82.6 82.1 80.8 80.8 -1.8 -1.8
    Human Expert (High) 88.6 88.6 85.4 85.4 -3.2 -3.2
    GPT-4o (0513) 69.1 64.7 54.0 49.7 -15.1 -19.4
    Claude 3.5 Sonnet 68.3 63.7 55.0 48.0 -13.3 -20.3
    Gemini 1.5 Pro (0801) 65.8 60.6 49.4 44.4 -16.4 -21.4
    Gemini 1.5 Pro (0523) 62.2 57.6 46.5 40.5 -15.7 -21.7
    GPT-4o mini 59.4 55.3 39.9 35.2 -19.5 -24.2
    Qwen2-VL-72B 64.5 59.3 49.2 43.3 -15.3 -21.2
    InternVL2-Llama3-76B 58.3 55.0 41.9 38.0 -16.4 -20.3
    InternVL2-40B 55.2 47.4 36.3 32.1 -18.9 -23.1
    LLaVA-OneVision-72B 56.8 52.3 38.0 24.0 -18.8 -32.8
    Qwen2-VL-7B 54.1 46.6 34.1 27.0 -20.0 -27.1
    Pixtral-12B 52.5 47.5 33.4 25.0 -19.1 -27.5
    InternVL2-8B 51.2 42.6 32.5 25.4 -18.7 -25.8
    MiniCPM-V2.6 49.8 40.6 30.2 24.2 -19.6 -25.6
    VILA-1.5-40B 51.9 46.8 35.9 14.1 -16.0 -37.8
    LLaVA-NeXT-72B 49.9 43.0 31.0 19.2 -18.9 -30.7
    LLaVA-OneVision-7B 48.8 42.8 29.5 18.7 -19.3 -30.1
    LLaVA-NeXT-34B 48.1 44.5 30.3 17.2 -17.8 -30.9
    Idefics3-8B-Llama3 46.6 40.8 30.1 15.6 -16.5 -31.0
    Qwen2-VL-2B 41.1 34.8 25.3 17.2 -15.8 -23.9
    Phi-3.5-Vision 43.0 37.8 26.3 13.1 -16.7 -29.9
    LLaVA-NeXT-7B 35.3 33.7 19.4 14.6 -15.9 -20.7
    LLaVA-NeXT-13B 36.2 33.9 19.8 14.5 -16.4 -21.7

    Expanding candidate options from 4 to 10 causes accuracy decreases of 13.3%13.3\% to 20.0%20.0\% across models. Presenting questions as images in the Vision-Only setting leads to further performance degradations, with declines reaching up to 37.8%37.8\% compared to original MMMU validation scores.

  3. Knowl 3 — Human Expert Performance Approximation on Augmented Benchmark

    model/method

    To estimate human expert performance on MMMU-Pro without re-annotating all questions, an analytical approximation method is used based on validation logs from 90 human experts on the original MMMU dataset.

    Questions are partitioned according to whether an expert provided a documented solving process (Numw/ SolutionNum_{\text{w/ Solution}}) or answered without providing a detailed solving process (Numw/o SolutionNum_{\text{w/o Solution}}):

    Numtotal=Numw/o Solution+Numw/ Solution=Numw/o Solution(wrong)+Numw/o Solution(correct)+Numw/ Solution(wrong)+Numw/ Solution(correct)Num_{\text{total}} = Num_{\text{w/o Solution}} + Num_{\text{w/ Solution}} = Num_{\text{w/o Solution(wrong)}} + Num_{\text{w/o Solution(correct)}} + Num_{\text{w/ Solution(wrong)}} + Num_{\text{w/ Solution(correct)}}

    For questions lacking detailed solution processes, guessing on expanded candidate options is accounted for by calculating a conservative lower bound of correct responses:

    NumEstimate(correct)=Numw/ Solution(correct)+[Numw/o SolutionNumtotal]×Numw/o SolutionNum_{\text{Estimate(correct)}} = Num_{\text{w/ Solution(correct)}} + \left[ \frac{Num_{\text{w/o Solution}}}{Num_{\text{total}}} \right] \times Num_{\text{w/o Solution}}

    This yields estimated human expert baseline performance across three proficiency tiers on MMMU-Pro:

    • Low: 73.0%73.0\% overall (Art & Design: 77.4%77.4\%, Business: 77.9%77.9\%, Science: 78.5%78.5\%, Health & Medicine: 65.2%65.2\%, Human & Social Sci.: 63.6%63.6\%, Tech & Eng.: 73.5%73.5\%).
    • Medium: 80.8%80.8\% overall (Art & Design: 83.3%83.3\%, Business: 88.4%88.4\%, Science: 84.9%84.9\%, Health & Medicine: 72.8%72.8\%, Human & Social Sci.: 75.8%75.8\%, Tech & Eng.: 78.2%78.2\%).
    • High: 85.4%85.4\% overall (Art & Design: 85.7%85.7\%, Business: 89.5%89.5\%, Science: 86.0%86.0\%, Health & Medicine: 84.8%84.8\%, Human & Social Sci.: 81.8%81.8\%, Tech & Eng.: 84.4%84.4\%).
  4. Knowl 4 — OCR Accuracy Metric and Impact of OCR Prompts on Vision-Only Evaluation

    empirical result

    In the MMMU-Pro Vision-Only setting, Optical Character Recognition (OCR) performance is calculated using normalized Levenshtein distance between model-extracted text and ground-truth text:

    OCR Accuracy=1−Levenshtein.distance(text1,text2)max⁡(len(text1),len(text2))\text{OCR Accuracy} = 1 - \frac{\text{Levenshtein.distance}(\text{text}_1, \text{text}_2)}{\max(\text{len}(\text{text}_1), \text{len}(\text{text}_2))}

    Comparing OCR accuracy and Vision-Only question-answering accuracy with and without an explicit OCR prompt ("Write out the multiple-choice question in the image and then solve it") reveals the following results:

    Model OCR Acc. (%) Vision Acc. w/ OCR Prompt (%) Vision Acc. w/o OCR Prompt (%)
    GPT-4o 92.3 49.7 49.4
    Gemini 1.5 Pro (0801) 89.7 44.4 43.6
    GPT-4o mini 89.6 35.2 35.6
    InternVL2-Llama3-76B 88.1 38.0 37.9
    LLaVA-OneVision-72B 87.8 24.0 23.8
    InternVL2-Llama3-40B 85.5 32.1 28.9
    InternVL2-8B 85.2 25.4 24.6
    Pixtral-12B 83.1 25.0 24.1
    Idefics3-8B-Llama3 68.5 15.6 14.1
    MiniCPM-V2.6 67.0 24.2 21.1
    LLaVA-NeXT-72B 62.0 19.2 20.0
    LLaVA-NeXT-13B 51.1 14.5 12.8
    LLaVA-NeXT-7B 36.6 14.6 14.3

    High OCR accuracy does not ensure strong multimodal reasoning performance on MMMU-Pro Vision (e.g., LLaVA-OneVision-72B attains 87.8%87.8\% OCR accuracy but only 24.0%24.0\% Vision accuracy, whereas InternVL2-Llama3-76B attains 88.1%88.1\% OCR accuracy and 38.0%38.0\% Vision accuracy). Additionally, instructing models to extract text explicitly prior to reasoning provides negligible improvements.

  5. Knowl 5 — Discipline-Specific Impact of Chain-of-Thought Prompting in Vision Input Setting

    empirical result

    Chain-of-Thought (CoT) prompting in the MMMU-Pro Vision-Only setting exhibits stark variations across academic disciplines. Evaluating LLaVA-OneVision-72B and GPT-4o with direct answering versus CoT prompting demonstrates that performance gains concentrate in structured, reasoning-intensive domains:

    LLaVA-OneVision-72B GPT-4o
    Discipline CoT Acc (%) Direct Acc (%) Difference (%) CoT Acc (%) Direct Acc (%) Difference (%)
    Tech and Engineering 22.98 20.65 +2.33 37.72 23.23 +14.49
    Business 29.26 24.50 +4.76 57.45 42.79 +14.66
    Science 23.89 22.61 +1.28 46.67 38.46 +8.22
    Health and Medicine 19.22 20.78 -1.56 49.68 44.34 +5.34
    Humanities and Social Science 32.14 36.60 -4.46 60.08 57.87 +2.21
    Art and Design 20.42 37.53 -17.12 63.14 61.55 +1.58

    CoT prompting yields substantial improvements for GPT-4o in Tech and Engineering (+14.49%+14.49\%) and Business (+14.66%+14.66\%). In contrast, in Art and Design, CoT produces a minimal gain for GPT-4o (+1.58%+1.58\%) and a large drop for LLaVA-OneVision-72B (−17.12%-17.12\%), demonstrating reduced utility in tasks requiring direct visual interpretation rather than formal deductive reasoning.

  6. Knowl 6 — Error Distribution and Failure Modes in Vision-Only Setting

    empirical result

    An error analysis conducted on 60 failure cases of GPT-4o in the MMMU-Pro Vision-Only setting reveals the following error distribution:

    • Reasoning Error: 46%46\%
    • Perceptual Error: 27%27\%
    • Lack of Knowledge: 25%25\%
    • Annotation Error: 2%2\%
    • OCR Error: 0%0\%

    Reasoning errors account for 46%46\% of total errors, compared to 26%26\% in the original MMMU benchmark. Pure text extraction (OCR) accounted for 0%0\% of errors, demonstrating that the primary bottleneck in vision-only evaluation is not character recognition, but rather multimodal reasoning, cross-modal context switching, and the integration of visual and textual constraints.

  7. Knowl 7 — Cognitive Load Shift and Output Token Allocation Under Vision Inputs

    empirical result

    When analyzing GPT-4o response generation across the Standard and Vision-Only input settings (with sentences categorized by Qwen2-72B-Instruct into descriptive versus analytical sentences), models show a pronounced shift in token allocation:

    • Descriptive Tokens: 49 tokens in Standard setting vs. 108 tokens in Vision-Only setting.
    • Analytical Tokens: 360 tokens in Standard setting vs. 258 tokens in Vision-Only setting.
    • Total Generated Tokens: 409 tokens in Standard setting vs. 366 tokens in Vision-Only setting.

    In the Vision-Only setting, the model generates fewer overall tokens and decreases analytical reasoning tokens by 28.3%28.3\% while more than doubling the descriptive tokens, indicating that the cognitive load of extracting and interpreting visual information reduces the depth of downstream analytical reasoning chains.

  8. Knowl 8 — Comparison of Self-Supervised vs. Language-Supervised Vision Encoders on MMMU-Pro

    empirical result

    Evaluating different vision encoders within the Cambrian-1 architecture (using a fixed Llama 3.1 8B LLM backbone trained on 1M Cambrian supervised fine-tuning data, with visual features interpolated to 576 tokens) shows that vision encoder pretraining objectives influence multimodal reasoning resilience under embedded vision inputs:

    Vision Encoder MMMU (Val) Acc. (%) MMMU-Pro (Vision) Acc. (%)
    DINOv2 ViT-G-14 (Self-Supervised) 37.1 17.4
    SigLIP ViT-SO400M-14 (Language-Supervised) 37.9 16.7

    While the language-supervised encoder (SigLIP ViT-SO400M-14) achieves higher accuracy on the standard MMMU validation set (37.9%37.9\% vs. 37.1%37.1\%), the self-supervised encoder (DINOv2 ViT-G-14) achieves higher accuracy in the MMMU-Pro Vision-Only setting (17.4%17.4\% vs. 16.7%16.7\%), suggesting that self-supervised visual representations offer greater robustness for complex, text-rich image reasoning.

  9. Knowl 9 — Limitations of the MMMU-Pro Benchmark

    limitation

    The authors identify four main limitations of the MMMU-Pro benchmark:

    1. Residual Statistical Shortcuts: Although four text-only LLMs were used to filter out solvable questions, subtle statistical shortcuts or biases within the remaining questions and augmented options may persist.
    2. Constrained Task Format: The benchmark remains confined to predefined academic disciplines and multiple-choice formats with up to 10 options, not covering open-ended generation or interactive evaluation.
    3. Approximated Human Baselines: Human expert scores are derived analytically from existing MMMU human evaluation logs using conservative lower-bound assumptions rather than direct, de novo human testing on the expanded option sets.
    4. Visual Complexity Gap: The vision-only setting uses synthetic and photographed 2D captures, which does not fully encompass the multi-sensory and dynamic complexities of natural human visual perception.

Coverage note — No substantial contributed material was omitted. The knowls cover the benchmark construction pipeline, comprehensive model performance data, human baseline approximation, OCR evaluation, discipline-level CoT effects, error and token distribution analyses, vision encoder experiments, and stated limitations.

References

  1. 1.Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. ArXiv preprint, abs/2404.14219.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems.
  3. 3.Anthropic. 2024. Claude 3.5 sonnet. https://www.anthropic.com/news/claude-3-5-sonnet.
  4. 4.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: visual question answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2425–2433. IEEE Computer Society.
  5. 5.Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. 2023. Openflamingo: An open-source framework for training large autoregressive vision-language models. ArXiv preprint, abs/2308.01390.
  6. 6.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In European Conference on Computer Vision, pages 104–120.
  7. 7.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. 2024. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. ArXiv preprint, abs/2404.16821.
  8. 8.Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao. 2023. Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges. ArXiv preprint, abs/2311.03287.
  9. 9.Wenliang Dai, Junnan Li, DONGXU LI, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems, volume 36, pages 49250–49267. Curran Associates, Inc.
  10. 10.Mengnan Du, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2023. Shortcut learning of large language models in natural language understanding. Communications of the ACM, 67(1):110–120.
  11. 11.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. ArXiv preprint, abs/2407.21783.
  12. 12.Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. 2024. Blink: Multimodal large language models can see but not perceive. ArXiv preprint, abs/2404.12390.
  13. 13.Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. 2023. Llama-adapter v2: Parameter-efficient visual instruction model. ArXiv preprint, abs/2304.15010.
  14. 14.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 6325–6334. IEEE Computer Society.
  15. 15.gpt-4o. 2024. Cheaper, better, faster, stronger. https://mistral.ai/news/mixtral-8x22b/.
  16. 16.Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. 2024. Mantis: Interleaved multi-image instruction tuning. ArXiv preprint, abs/2405.01483.
  17. 17.Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, et al. 2024. Efficient multimodal large language models: A survey. ArXiv preprint, abs/2405.10739.
  18. 18.Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 881–905, Bangkok, Thailand. Association for Computational Linguistics.
  19. 19.Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Léo Tronchon. 2024. Building and better understanding vision-language models: insights and future directions. ArXiv preprint, abs/2408.12637.
  20. 20.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023. Otter: A multi-modal model with in-context instruction tuning. ArXiv preprint, abs/2305.03726.
  21. 21.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024a. Llava-onevision: Easy visual task transfer. ArXiv preprint, abs/2408.03326.
  22. 22.Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024b. Seedbench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13299–13308.
  23. 23.Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024c. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. ArXiv preprint, abs/2407.07895.
  24. 24.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX 16, pages 121–137. Springer.
  25. 25.Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, et al. 2024d. Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehension. ArXiv preprint, abs/2407.04903.
  26. 26.Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26689–26699.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer.
  28. 28.Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. 2023a. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. ArXiv preprint, abs/2310.14566.
  29. 29.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023b. Improved baselines with visual instruction tuning. In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
  30. 30.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024a. Llava-next: Improved reasoning, ocr, and world knowledge.
  31. 31.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023c. Visual instruction tuning. In Advances in Neural Information Processing Systems, volume 36, pages 34892–34916. Curran Associates, Inc.
  32. 32.Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue. 2024b. Visualwebbench: How far have multimodal llms evolved in web page understanding and grounding? Conference on Language Modeling.
  33. 33.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. 2023d. Mmbench: Is your multi-modal model an all-around player? ArXiv preprint, abs/2307.06281.
  34. 34.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13–23.
  35. 35.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2023a. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. ArXiv preprint, abs/2310.02255.
  36. 36.Yujie Lu, Xiujun Li, William Yang Wang, and Yejin Choi. 2023b. Vim: Probing multimodal large language models for visual embedded instruction following. ArXiv preprint, abs/2311.17647.
  37. 37.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. OK-VQA: A visual question answering benchmark requiring external knowledge. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 3195–3204. Computer Vision Foundation / IEEE.
  38. 38.Mistral. 2024. Pixtral-12b. https://mistral.ai/news/pixtral-12b.
  39. 39.Masoud Monajatipoor, Liunian Harold Li, Mozhdeh Rouhsedaghat, Lin Yang, and Kai-Wei Chang. 2023. MetaVL: Transferring in-context learning ability from language models to vision-language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 495–508, Toronto, Canada. Association for Computational Linguistics.
  40. 40.OpenAI. 2023. Gpt-4v(ision) system card.
  41. 41.OpenAI. 2024a. Gpt-4o mini: advancing cost-efficient intelligence. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/.
  42. 42.OpenAI. 2024b. Hello gpt4-o. https://openai.com/index/hello-gpt-4o/.
  43. 43.Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193.
  44. 44.Qwen. 2024. Qwen2-vl: To see the world more clearly. https://qwenlm.github.io/blog/qwen2-vl/ .
  45. 45.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. ArXiv preprint, abs/2403.05530.
  46. 46.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. ArXiv preprint, abs/2312.11805.
  47. 47.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. 2024a. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. ArXiv preprint, abs/2406.16860.
  48. 48.Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024b. Eyes wide shut? exploring the visual shortcomings of multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9568–9578.
  49. 49.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. 2024. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. ArXiv preprint, abs/2406.01574.
  50. 50.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  51. 51.Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui. 2024. From decoding to meta-generation: Inference-time algorithms for large language models. arXiv preprint arXiv:2406.16838.
  52. 52.Penghao Wu and Saining Xie. 2024. V*: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094.
  53. 53.Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Qiao, and Ping Luo. 2023. Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models. ArXiv preprint, abs/2306.09265.
  54. 54.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. 2024. Qwen2 technical report. ArXiv preprint, abs/2407.10671.
  55. 55.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2024. Minicpm-v: A gpt-4v level mllm on your phone. ArXiv preprint, abs/2408.01800.
  56. 56.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. 2023a. mplug-owl: Modularization empowers large language models with multimodality. ArXiv preprint, abs/2304.14178.
  57. 57.Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. 2023b. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. ArXiv preprint, abs/2311.04257.
  58. 58.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023a. A survey on multimodal large language models. ArXiv preprint, abs/2306.13549.
  59. 59.Zhenfei Yin, Jiong Wang, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang, Lu Sheng, LEI BAI, Jing Shao, and Wanli Ouyang. 2023b. Lamm: Language-assisted multimodal instruction-tuning dataset, framework, and benchmark. In Advances in Neural Information Processing Systems, volume 36, pages 26650–26685. Curran Associates, Inc.
  60. 60.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. 2024. Yi: Open foundation models by 01. ai. ArXiv preprint, abs/2403.04652.
  61. 61.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2024. MM-vet: Evaluating large multimodal models for integrated capabilities. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 57730–57754. PMLR.
  62. 62.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2024. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9556–9567.
  63. 63.Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. 2023. When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations.
  64. 64.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986.
  65. 65.Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, and Jiaqi Wang. 2024a. Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output. ArXiv preprint, abs/2407.03320.
  66. 66.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021. Vinvl: Revisiting visual representations in vision-language models. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 5579–5588. Computer Vision Foundation / IEEE.
  67. 67.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2023. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. ArXiv preprint, abs/2303.16199.
  68. 68.Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang, Peng Gao, et al. 2024b. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? arXiv preprint arXiv:2403.14624.
  69. 69.Bo Zhao, Boya Wu, and Tiejun Huang. 2023. Svit: Scaling up visual instruction tuning. ArXiv preprint, abs/2307.04087.
  70. 70.Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang. 2024. Mmicl: Empowering vision-language model with multi-modal in-context learning. The Twelfth International Conference on Learning Representations.
  71. 71.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. 2024. Gpt-4v(ision) is a generalist web agent, if grounded. In Forty-first International Conference on Machine Learning.
  72. 72.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and VQA. In Proceedings of the AAAI Conference on Artificial Intelligence, 34, pages 13041–13049. AAAI Press.
  73. 73.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. ArXiv preprint, abs/2304.10592.

Citation

MLA
Yue, X., et al. “MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 15134–86, https://doi.org/10.18653/v1/2025.acl-long.736.
APA
Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., Su, Y., Chen, W., & Neubig, G. (2025). MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15134–15186. https://doi.org/10.18653/v1/2025.acl-long.736
Chicago
Yue, X., T. Zheng, Y. Ni, et al. 2025. “MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15134–86. https://doi.org/10.18653/v1/2025.acl-long.736.
Harvard
Yue, X. et al. (2025) “MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15134–15186. Available at: https://doi.org/10.18653/v1/2025.acl-long.736.
Vancouver
1. Yue X, Zheng T, Ni Y, et al (2025) MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15134–15186

BibTeX

@inproceedings{yue-etal-2025-mmmu,
    title = "{MMMU}-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark",
    author = "Yue, Xiang  and
      Zheng, Tianyu  and
      Ni, Yuansheng  and
      Wang, Yubo  and
      Zhang, Kai  and
      Tong, Shengbang  and
      Sun, Yuxuan  and
      Yu, Botao  and
      Zhang, Ge  and
      Sun, Huan  and
      Su, Yu  and
      Chen, Wenhu  and
      Neubig, Graham",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.736/",
    doi = "10.18653/v1/2025.acl-long.736",
    pages = "15134--15186",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/