MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Dongzhi JiangRenrui ZhangZiyu GuoYanwei LiYu QiXinyan ChenLiuhui WangJianhan JinClaire GuoShen Yan

article2025ICML137 citations

Presents MME-CoT, a comprehensive benchmark and evaluation suite across six multimodal domains that reveals how chain-of-thought prompting can improve reasoning quality through reflection while paradoxically degrading model performance on perception-heavy tasks due to overthinking.

Listen

Artificial intelligence systems that integrate vision and text are increasingly adopting step-by-step reasoning prompts—known as Chain-of-Thought—to solve complex problems. However, current evaluations predominantly measure whether an AI model reaches the correct final answer, ignoring whether the intermediate reasoning steps are logically sound, necessary, or efficient. This lack of detailed assessment creates an illusion of competence and obscures hidden operational risks when deploying reasoning models in real-world scenarios.

The article introduces a dedicated evaluation suite called MME-CoT to systematically benchmark multimodal AI models across three core operational dimensions: reasoning quality, robustness, and efficiency. To achieve this, the authors curated a verified dataset containing 1,130 problems across six domains, including math, science, optical character recognition, logic, spatial-temporal reasoning, and general scene comprehension. The evaluation breaks down model outputs into discrete intermediate steps and compares them against human-verified key logical deductions and visual captions, while also comparing step-by-step reasoning prompts against direct-answer requests across both perception-focused and reasoning-intensive tasks.

The analysis reveals several critical findings for technical and strategic leaders. First, while self-reflection mechanisms improve overall reasoning quality—leading closed-source models like Kimi k1.5 and GPT-4o to achieve high quality scores—extended reasoning often introduces severe inefficiencies. Models with extended reasoning capabilities frequently generate distracting visual descriptions, resulting in 30% to 40% of their self-reflection steps failing to contribute meaningfully to the correct solution. Second, prompting models to think step by step systematically degrades performance on direct perception tasks that do not require complex logic, showing an accuracy drop of up to 6.8% in some models due to overthinking. Third, larger model scale significantly improves reasoning efficacy; for example, Qwen2-VL-72B demonstrated an accuracy improvement with step-by-step reasoning, whereas its smaller 7B counterpart experienced a 4.8% drop on reasoning tasks under the same prompt.

These findings demonstrate that applying long, step-by-step reasoning as a default setting across all tasks is both computationally expensive and detrimental to accuracy in visual recognition settings. System architects cannot assume that a correct final answer implies sound underlying logic, nor that extended reasoning is uniformly helpful across problem types. Instead, engineering teams should implement task-routing mechanisms that restrict step-by-step reasoning prompts strictly to logic-heavy challenges while maintaining direct-answering protocols for visual perception tasks.

Moving forward, developers and researchers should focus on refining self-reflection algorithms to suppress redundant reasoning steps and eliminate hallucinations during image description. While the study's automated scoring is strongly validated by human evaluations reaching 86% to 98% agreement, organizations should remain cautious when interpreting benchmarks from models that refuse direct prompting instructions, as out-of-distribution formatting can distort evaluation metrics.

arXiv: 2502.09621MME-Benchmarks/MME-CoT

No sufficiently relevant recommendations were found.

Cover for MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Abstract

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation. In this paper, we introduce MME-CoT, a specialized benchmark evaluating the CoT reasoning performance of LMMs, spanning six domains: math, science, OCR, logic, space-time, and general scenes. As the first comprehensive study in this area, we propose a thorough evaluation suite incorporating three novel metrics that assess the reasoning quality, robustness, and efficiency at a fine-grained level. Leveraging curated high-quality data and a unique evaluation strategy, we conduct an in-depth analysis of state-of-the-art LMMs, uncovering several key insights: 1) Models with reflection mechanism demonstrate a superior CoT quality, with Kimi k1.5 outperforming GPT-4o and demonstrating the highest quality results; 2) CoT prompting often degrades LMM performance on perception-heavy tasks, suggesting a potentially harmful overthinking behavior; and 3) Although the CoT quality is high, LMMs with reflection exhibit significant inefficiency in both normal response and self-correction phases. We hope MME-CoT serves as a foundation for advancing multimodal reasoning in LMMs.

Table of Contents

  • 1. Introduction
  • 2. Dataset Curation
  • 2.1. Data Composition and Categorization
  • 2.2. Data Annotation and Review
  • 3. CoT Evaluation Strategy
  • 3.1. CoT Quality Evaluation
  • 3.2. CoT Robustness Evaluation
  • 3.3. CoT Efficiency Evaluation
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Quantitative Results
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix Overview
  • A. Related Work
  • A.1. Large Multimodal Models
  • A.2. Reasoning Evaluation
  • C.2. Preliminary Categorization Result
  • D. More Experiment Details
  • D.1. More Findings
  • D.2. Human Agreement
  • E. Error Analysis
  • F. More Qualitative Examples
  • G. Detailed Evaluation Setup
  • G.1. CoT Quality Evaluation Prompts
  • G.2. CoT Efficiency Prompt
  • G.3. Direct Evaluation Prompt

Knowls

  1. Knowl 1 — MME-CoT dataset scope and composition

    data/table

    MME-CoT is a benchmark for evaluating chain-of-thought (CoT) in large multimodal models across six visual question domains: math, science, OCR, logic, space-time, and general scenes. It contains 1,130 questions: 837 reasoning questions (74.1%) and 293 perception questions (25.9%). The dataset includes 431 multiple-choice and 406 free-form reasoning questions, plus 275 multiple-choice and 18 free-form perception questions. Its annotations comprise 2,667 inference conclusions and 1,198 image captions, for 3,865 key-step annotations in total. The benchmark covers 2,380 unique images; the paper also reports 808 unique questions and 271 unique answers.

  2. Knowl 2 — Human-verified key-step and image-caption annotations

    model/method

    For each reasoning question in MME-CoT, annotators provide necessary key steps in two forms: inference conclusions, including the final answer, and image captions describing visually critical information. GPT-4o first drafts the annotations using the question, image, and ground-truth answer; human annotators then correct the draft or create annotations independently when the draft is unsuitable. They retain concise core conclusions and relevant visual descriptions, and document all solution paths for questions with multiple valid methods. Separate reference image captions record relevant visual information not already captured by the key-step captions, supporting evaluation of generated visual descriptions.

  3. Knowl 3 — CoT quality measured by recall, precision, and F1

    model/method

    MME-CoT evaluates CoT quality along two dimensions: coverage of necessary reasoning and faithfulness of generated steps. For recall, let S(m)S^{(m)} be the set of annotated inference conclusions and image captions for valid solution method mm, and let Smatched(m)S^{(m)}_{\mathrm{matched}} be the subset found in a model response. The method selected for scoring is the one with the largest matched fraction; overall recall is that fraction:

    m∗=arg⁡max⁡m∣Smatched(m)∣∣S(m)∣,Recall=∣Smatched(m∗)∣∣S(m∗)∣.m^*=\arg\max_m\frac{|S^{(m)}_{\mathrm{matched}}|}{|S^{(m)}|},\qquad \mathrm{Recall}=\frac{|S^{(m^*)}_{\mathrm{matched}}|}{|S^{(m^*)}|}.

    Recall is also calculated separately for inference conclusions and image captions. For precision, the response is partitioned into generated inference steps CPC^P and image-caption steps IPI^P; background-information steps are excluded. A generated inference is correct if it matches an annotated conclusion or is a valid inference consistent with the annotations; a generated caption is correct if it matches an annotation or faithfully describes the image. If CcorrectPC^P_{\mathrm{correct}} and IcorrectPI^P_{\mathrm{correct}} are the correct generated steps of each type, then Precision=∣CcorrectP∪IcorrectP∣/∣CP∪IP∣\mathrm{Precision}=|C^P_{\mathrm{correct}}\cup I^P_{\mathrm{correct}}|/|C^P\cup I^P|. The reported CoT quality F1 is the harmonic mean of precision and recall.

  4. Knowl 4 — Robustness compares direct and step-by-step prompting

    model/method

    MME-CoT measures the effect of CoT prompting by comparing final-answer accuracy under a direct-answer prompt (DIR) and a step-by-step prompt (COT). For perception-task set PP, stability is AccCOTP−AccDIRP\mathrm{Acc}^{P}_{\mathrm{COT}}-\mathrm{Acc}^{P}_{\mathrm{DIR}}; for reasoning-task set RR, efficacy is AccCOTR−AccDIRR\mathrm{Acc}^{R}_{\mathrm{COT}}-\mathrm{Acc}^{R}_{\mathrm{DIR}}. Here, AccqT\mathrm{Acc}^{T}_{q} is the proportion of correct final answers on task set TT under prompt qq. Stability indicates whether CoT harms or preserves perception performance, while efficacy measures its accuracy gain or loss on reasoning tasks. The benchmark treats nonnegative stability as the desired behavior.

  5. Knowl 5 — Relevance rate scores the useful fraction of generated steps

    model/method

    The relevance-rate metric measures how much of a model's generated CoT contributes to solving the question, independently of whether the reasoning is correct. Responses are partitioned into steps, and a step is relevant when at least 75% of its content is directed toward reaching the answer. Let PP be all generated steps, PrelevantP_{\mathrm{relevant}} the relevant subset, CPC^P the inference steps, and IPI^P the image-caption steps. The raw overall rate is r∅=∣Prelevant∣/∣P∣r_{\varnothing}=|P_{\mathrm{relevant}}|/|P|; the raw rates for inference and caption steps are rC=∣CrelevantP∣/∣CP∣r_C=|C^P_{\mathrm{relevant}}|/|C^P| and rI=∣IrelevantP∣/∣IP∣r_I=|I^P_{\mathrm{relevant}}|/|I^P|. Each reported rate is scaled using RelevanceRatex=(rx−α)/(1−α)\mathrm{RelevanceRate}_x=(r_x-\alpha)/(1-\alpha) for x∈{C,I,∅}x\in\{C,I,\varnothing\}, with α=0.8\alpha=0.8. The scaled score makes differences among models more visible.

  6. Knowl 6 — Reflection quality measures useful self-correction

    model/method

    A reflection is valid if it either identifies and corrects a previous mistake or verifies a previous conclusion using a new insight. Restating earlier content, introducing an error, or proposing an analysis without carrying it out does not qualify as valid reflection. Given the set RR of identified reflection steps and its valid subset RvalidR_{\mathrm{valid}}, reflection quality is ∣Rvalid∣/∣R∣|R_{\mathrm{valid}}|/|R|. The evaluation identifies candidate reflections using indicators of reconsideration, such as “Wait,” “Alternatively,” and “Let me double-check,” and judges them against the validity definition.

  7. Knowl 7 — Evaluation protocol and model coverage

    experimental setup

    The evaluation compares open- and closed-source multimodal models, including LLaVA-OneVision, Qwen2-VL, InternVL2.5, MiniCPM-V, DeepSeek-VL2, LLaVA-CoT, Mulberry, Virgo, QVQ, GPT-4o, Claude-3.5, Gemini-2.0-Flash, and Kimi k1.5. The direct prompt requests only a final answer; the CoT prompt requests step-by-step reasoning followed by a final answer. GPT-4o mini extracts final answers for direct accuracy comparisons, while GPT-4o evaluates the other CoT criteria. Models without reflection capability are assigned a reflection-quality score of 100 in the reported results. The authors follow VLMEvalKit settings for model hyperparameters.

  8. Knowl 8 — Reflection models approach or surpass GPT-4o in CoT quality

    empirical result

    Kimi k1.5 obtained the highest overall CoT-quality F1 score in the reported evaluation, 64.2, narrowly above GPT-4o at 64.0. Kimi also had the highest precision score, 92.0, while GPT-4o led the recall metrics. Among open-source models, QVQ-72B scored 62.0 F1, compared with 56.2 for its Qwen2-VL-72B base model, a gain of 5.8 points; QVQ's precision was 80.2 versus 77.3 for the base model. The authors interpret this comparison as evidence that reflection-oriented reasoning training can improve CoT quality, while noting that long CoT does not guarantee that all annotated key steps are covered.

  9. Knowl 9 — CoT often harms perception accuracy and can be bypassed

    empirical result

    Most evaluated models had negative stability scores, meaning that step-by-step prompting reduced accuracy on perception tasks relative to direct-answer prompting. The largest reported decline was 6.8 percentage points for InternVL2.5-8B. CoT's effect on reasoning varied with model scale in the Qwen2-VL comparison: Qwen2-VL-7B lost 4.6 points under CoT, whereas Qwen2-VL-72B gained 2.4 points. The authors also caution that robustness scores can be misleading when a model does not follow the direct prompt: they observed extended rationales under direct-answer prompting from Mulberry and from models including LLaVA-CoT, Virgo, QVQ, and Kimi k1.5.

  10. Knowl 10 — Efficiency analysis exposes irrelevant content and failed reflections

    empirical result

    The efficiency evaluation found that lengthy CoT can include content that does not help answer the question. Irrelevant image descriptions were especially common in general-scene, space-time, and OCR tasks, where models sometimes described visual details beyond those needed for the answer. InternVL2.5-8B received the highest reported overall relevance-rate score, 98.4%. Reflection was also frequently unproductive: the paper reports that roughly 30%–40% of reflection steps fail to help reach a correct answer; QVQ-72B and Virgo-72B had reflection-quality scores of 61.7 and 60.6, respectively, and Kimi k1.5 scored 72.2.

  11. Knowl 11 — CoT and reflection failures have distinct error patterns

    data/table

    The paper classifies CoT errors into visual-perception mistakes, failures to reason from relevant visual information, flawed logical reasoning, and calculation errors. Their reported distribution is 11.1%, 31.5%, 38.9%, and 18.5%, respectively, making logical-reasoning errors the largest category. In a separate analysis of 200 QVQ predictions, reflection failures were classified as ineffective reflection (continuing to make incorrect adjustments), incompleteness (suggesting but not executing a new analysis), repetition, or interference (introducing errors after a correct conclusion). The reported shares were 76.0%, 17.3%, 4.9%, and 1.8%, respectively.

Coverage note — The supplementary human-agreement audit and detailed per-source and per-subcategory dataset breakdowns are omitted because they support evaluation reliability and describe dataset composition in finer detail, rather than adding a core benchmark component or principal finding.

References

  1. 1.Chen, D., Chen, R., Zhang, S., Wang, Y., Liu, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L. Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. In Forty-first International Conference on Machine Learning, 2024a.
  2. 2.Chen, G., Zheng, Y.-D., Wang, J., Xu, J., Huang, Y., Pan, J., Wang, Y., Wang, Y., Qiao, Y., Lu, T., et al. Videollm: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292, 2023.
  3. 3.Chen, Q., Qin, L., Zhang, J., Chen, Z., Xu, X., and Che, W. M3 cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought. In Proc. of ACL, 2024b.
  4. 4.Chen, X., Zhang, R., Jiang, D., Zhou, A., Yan, S., Lin, W., and Li, H. Mint-cot: Enabling interleaved visual tokens in mathematical chain-of-thought reasoning. arXiv preprint arXiv:2506.05331, 2025.
  5. 5.Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024c.
  6. 6.Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024d.
  7. 7.Du, Y., Liu, Z., Li, Y., Zhao, W. X., Huo, Y., Wang, B., Chen, W., Liu, Z., Wang, Z., and Wen, J.-R. Virgo: A preliminary exploration on reproducing o1-like mllm. arXiv preprint arXiv:2501.01904, 2025.
  8. 8.Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 11198–11201, 2024.
  9. 9.Gao, P., Zhang, R., Liu, C., Qiu, L., Huang, S., Lin, W., Zhao, S., Geng, S., Lin, Z., Jin, P., et al. Sphinx-x: Scaling data and parameters for a family of multi-modal large language models. ICML 2024, 2024.
  10. 10.Golovneva, O., Chen, M., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A. Roscoe: A suite of metrics for scoring step-by-step reasoning. arXiv preprint arXiv:2212.07919, 2022.
  11. 11.Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025a.
  12. 12.Guo, Z., Zhang, R., Zhu, X., Tang, Y., Ma, X., Han, J., Chen, K., Gao, P., Li, X., Li, H., et al. Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following. arXiv preprint arXiv:2309.00615, 2023.
  13. 13.Guo, Z., Zhang, R., Zhu, X., Tong, C., Gao, P., Li, C., and Heng, P.-A. Sam2point: Segment any 3d as videos in zero-shot and promptable manners. arXiv preprint arXiv:2408.16768, 2024.
  14. 14.Guo, Z., Lin, H., Yuan, Z., Zheng, C., Qiu, P., Jiang, D., Zhang, R., Feng, C.-M., and Li, Z. Pisa: A self-augmented data engine and training strategy for 3d understanding with large models. arXiv preprint arXiv:2503.10529, 2025b.
  15. 15.Guo, Z., Zhang, R., Chen, H., Gao, J., Jiang, D., Wang, J., and Heng, P.-A. Sciverse: Unveiling the knowledge comprehension and visual reasoning of lmms on multi-modal scientific problems. arXiv preprint arXiv:2503.10627, 2025c.
  16. 16.Guo, Z., Zhang, R., Tong, C., Zhao, Z., Gao, P., Li, H., and Heng, P.-A. Can we generate images with cot? let’s verify and reinforce image generation step by step. arXiv preprint arXiv:2501.13926, 2025d.
  17. 17.Hao, S., Gu, Y., Luo, H., Liu, T., Shao, X., Wang, X., Xie, S., Ma, H., Samavedhi, A., Gao, Q., et al. Llm reasoners: New evaluation, library, and analysis of step-by-step reasoning with large language models. arXiv preprint arXiv:2404.05221, 2024.
  18. 18.He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M. Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.
  19. 19.Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  20. 20.Jia, Y., Liu, J., Chen, S., Gu, C., Wang, Z., Luo, L., Lee, L., Wang, P., Wang, Z., Zhang, R., et al. Lift3d foundation policy: Lifting 2d large-scale pretrained models for robust 3d robotic manipulation. arXiv preprint arXiv:2411.18623, 2024.
  21. 21.Jiang, D., Song, G., Wu, X., Zhang, R., Shen, D., Zong, Z., Liu, Y., and Li, H. Comat: Aligning text-to-image diffusion model with image-to-text concept matching. arXiv preprint arXiv:2404.03653, 2024a.
  22. 22.Jiang, D., Zhang, R., Guo, Z., Wu, Y., Lei, J., Qiu, P., Lu, P., Chen, Z., Song, G., Gao, P., et al. Mmsearch: Benchmarking the potential of large models as multimodal search engines. arXiv preprint arXiv:2409.12959, 2024b.
  23. 23.Jiang, D., Guo, Z., Zhang, R., Zong, Z., Li, H., Zhuo, L., Yan, S., Heng, P.-A., and Li, H. T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot. arXiv preprint arXiv:2505.00703, 2025.
  24. 24.Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024a.
  25. 25.Li, F., Zhang, R., Zhang, H., Zhang, Y., Li, B., Li, W., Ma, Z., and Li, C. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024b.
  26. 26.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pp. 12888–12900. PMLR, 2022.
  27. 27.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  28. 28.Lin, X., Zare, A., Huang, S., Yang, M.-H., Chang, S.-F., and Zhang, L. Personalized video comment generation. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 16806–16820, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.979. URL https://aclanthology.org/2024.findings-emnlp.979/.
  29. 29.Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. ECCV 2024, 2023.
  30. 30.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, 2023.
  31. 31.Lu, P., Bansal, H., Xia, T., Liu, J., yue Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models. ArXiv, abs/2310.02255, 2023.
  32. 32.OpenAI. GPT-4V(ision) system card, 2023. URL https://openai.com/research/gpt-4v-system-card.
  33. 33.OpenAI. Introducing openai o1, 2024., 2024a. URL https://openai.com/o1/.
  34. 34.OpenAI. Hello gpt-4o. https://openai.com/index/hello-gpt-4o/, 2024b.
  35. 35.Peng, T., Li, M., Zhou, H., Xia, R., Zhang, R., Bai, L., Mao, S., Wang, B., He, C., Zhou, A., et al. Chimera: Improving generalist model with domain-specific experts. arXiv preprint arXiv:2412.05983, 2024.
  36. 36.Prasad, A., Saha, S., Zhou, X., and Bansal, M. Receval: Evaluating reasoning chains via correctness and informativeness. arXiv preprint arXiv:2304.10703, 2023.
  37. 37.Qwen Team. Qwen2-vl. 2024.
  38. 38.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231591445.
  39. 39.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  40. 40.Shen, D., Song, G., Zhang, Y., Ma, B., Li, L., Jiang, D., Zong, Z., and Liu, Y. Adt: Tuning diffusion models with adversarial supervision. arXiv preprint arXiv:2504.11423, 2025.
  41. 41.Sprague, Z., Yin, F., Rodriguez, J. D., Jiang, D., Wadhwa, M., Singhal, P., Zhao, X., Ye, X., Mahowald, K., and Durrett, G. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. arXiv preprint arXiv:2409.12183, 2024.
  42. 42.Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025.
  43. 43.Team, Q. Qvq: To see the world with wisdom, December 2024. URL https://qwenlm.github.io/blog/qvq-72b-preview/.
  44. 44.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  45. 45.Wang, F., Fu, X., Huang, J. Y., Li, Z., Liu, Q., Liu, X., Ma, M. D., Xu, N., Zhou, W., Zhang, K., et al. Muirbench: A comprehensive benchmark for robust multi-image understanding. arXiv preprint arXiv:2406.09411, 2024a.
  46. 46.Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024b.
  47. 47.Wang, W., Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Zhu, J., Zhu, X., Lu, L., Qiao, Y., and Dai, J. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv preprint arXiv:2411.10442, 2024c.
  48. 48.Wang, W., Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Zhu, J., Zhu, X., Lu, L., Qiao, Y., et al. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization. arXiv preprint arXiv:2411.10442, 2024d.
  49. 49.Wang, Z., Xia, M., He, L., Chen, H., Liu, Y., Zhu, R., Liang, K., Wu, X., Liu, H., Malladi, S., Chevalier, A., Arora, S., and Chen, D. Charxiv: Charting gaps in realistic chart understanding in multimodal llms. arXiv preprint arXiv:2406.18521, 2024e.
  50. 50.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  51. 51.Xu, G., Jin, P., Li, H., Song, Y., Sun, L., and Yuan, L. Llava-cot: Let vision language models reason step-by-step, 2024. URL https://arxiv.org/abs/2411.10440.
  52. 52.Xu, R., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D. Pointllm: Empowering large language models to understand point clouds. arXiv preprint arXiv:2308.16911, 2023.
  53. 53.Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024.
  54. 54.Yao, H., Huang, J., Wu, W., Zhang, J., Wang, Y., Liu, S., Wang, Y., Song, Y., Feng, H., Shen, L., et al. Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search. arXiv preprint arXiv:2412.18319, 2024a.
  55. 55.Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024b.
  56. 56.Ying, K., Meng, F., Wang, J., Li, Z., Lin, H., Yang, Y., Zhang, H., Zhang, W., Lin, Y., Liu, S., et al. Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi. arXiv preprint arXiv:2404.16006, 2024.
  57. 57.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  58. 58.Yue, X., Zheng, T., Ni, Y., Wang, Y., Zhang, K., Tong, S., Sun, Y., Yu, B., Zhang, G., Sun, H., Su, Y., Chen, W., and Neubig, G. Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark, 2024. URL https://arxiv.org/abs/2409.02813.
  59. 59.Zhang, H., Li, H., Li, F., Ren, T., Zou, X., Liu, S., Huang, S., Gao, J., Zhang, L., Li, C., et al. Llava-grounding: Grounded visual chat with large multimodal models. arXiv preprint arXiv:2312.02949, 2023.
  60. 60.Zhang, R., Han, J., Liu, C., Zhou, A., Lu, P., Qiao, Y., Li, H., and Gao, P. Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention. In ICLR 2024, 2024a.
  61. 61.Zhang, R., Han, J., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., Gao, P., and Qiao, Y. LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention. In The Twelfth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=d4UiXAHN2W.
  62. 62.Zhang, R., Jiang, D., Zhang, Y., Lin, H., Guo, Z., Qiu, P., Zhou, A., Lu, P., Chang, K.-W., Gao, P., et al. Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024c.
  63. 63.Zhang, R., Wei, X., Jiang, D., Zhang, Y., Guo, Z., Tong, C., Liu, J., Zhou, A., Wei, B., Zhang, S., et al. Mavis: Mathematical visual instruction tuning. arXiv preprint arXiv:2407.08739, 2024d.
  64. 64.Zhang, Y., Bai, H., Zhang, R., Gu, J., Zhai, S., Susskind, J., and Jaitly, N. How far are we from intelligent visual deductive reasoning? In COLM, 2024e.
  65. 65.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
  66. 66.Zong, Z., Jiang, D., Ma, B., Song, G., Shao, H., Shen, D., Liu, Y., and Li, H. Easyref: Omni-generalized group image reference for diffusion models via multimodal llm. arXiv preprint arXiv:2412.09618, 2024a.
  67. 67.Zong, Z., Ma, B., Shen, D., Song, G., Shao, H., Jiang, D., Li, H., and Liu, Y. Mova: Adapting mixture of vision experts to multimodal context. arXiv preprint arXiv:2404.13046, 2024b.

Citation

MLA
Jiang, D., et al. “MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.09621.
APA
Jiang, D., Zhang, R., Guo, Z., Li, Y., Qi, Y., Chen, X., Wang, L., Jin, J., Guo, C., Yan, S., Zhang, B., Fu, C., Gao, P., & Li, H. (2025). MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency. arXiv. https://doi.org/10.48550/arxiv.2502.09621
Chicago
Jiang, D., R. Zhang, Z. Guo, et al. 2025. “MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.09621.
Harvard
Jiang, D. et al. (2025) “MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.09621.
Vancouver
1. Jiang D, Zhang R, Guo Z, et al (2025) MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency. https://doi.org/10.48550/arxiv.2502.09621

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.09621,
  doi = {10.48550/ARXIV.2502.09621},
  url = {https://arxiv.org/abs/2502.09621},
  author = {Jiang, Dongzhi and Zhang, Renrui and Guo, Ziyu and Li, Yanwei and Qi, Yu and Chen, Xinyan and Wang, Liuhui and Jin, Jianhan and Guo, Claire and Yan, Shen and Zhang, Bo and Fu, Chaoyou and Gao, Peng and Li, Hongsheng},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency},
  publisher = {arXiv},
  year = {2025},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/