MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Dongzhi JiangRenrui ZhangZiyu GuoYanwei LiYu QiXinyan ChenLiuhui WangJianhan JinClaire GuoShen Yan
Presents MME-CoT, a comprehensive benchmark and evaluation suite across six multimodal domains that reveals how chain-of-thought prompting can improve reasoning quality through reflection while paradoxically degrading model performance on perception-heavy tasks due to overthinking.
Artificial intelligence systems that integrate vision and text are increasingly adopting step-by-step reasoning prompts—known as Chain-of-Thought—to solve complex problems. However, current evaluations predominantly measure whether an AI model reaches the correct final answer, ignoring whether the intermediate reasoning steps are logically sound, necessary, or efficient. This lack of detailed assessment creates an illusion of competence and obscures hidden operational risks when deploying reasoning models in real-world scenarios.
The article introduces a dedicated evaluation suite called MME-CoT to systematically benchmark multimodal AI models across three core operational dimensions: reasoning quality, robustness, and efficiency. To achieve this, the authors curated a verified dataset containing 1,130 problems across six domains, including math, science, optical character recognition, logic, spatial-temporal reasoning, and general scene comprehension. The evaluation breaks down model outputs into discrete intermediate steps and compares them against human-verified key logical deductions and visual captions, while also comparing step-by-step reasoning prompts against direct-answer requests across both perception-focused and reasoning-intensive tasks.
The analysis reveals several critical findings for technical and strategic leaders. First, while self-reflection mechanisms improve overall reasoning quality—leading closed-source models like Kimi k1.5 and GPT-4o to achieve high quality scores—extended reasoning often introduces severe inefficiencies. Models with extended reasoning capabilities frequently generate distracting visual descriptions, resulting in 30% to 40% of their self-reflection steps failing to contribute meaningfully to the correct solution. Second, prompting models to think step by step systematically degrades performance on direct perception tasks that do not require complex logic, showing an accuracy drop of up to 6.8% in some models due to overthinking. Third, larger model scale significantly improves reasoning efficacy; for example, Qwen2-VL-72B demonstrated an accuracy improvement with step-by-step reasoning, whereas its smaller 7B counterpart experienced a 4.8% drop on reasoning tasks under the same prompt.
These findings demonstrate that applying long, step-by-step reasoning as a default setting across all tasks is both computationally expensive and detrimental to accuracy in visual recognition settings. System architects cannot assume that a correct final answer implies sound underlying logic, nor that extended reasoning is uniformly helpful across problem types. Instead, engineering teams should implement task-routing mechanisms that restrict step-by-step reasoning prompts strictly to logic-heavy challenges while maintaining direct-answering protocols for visual perception tasks.
Moving forward, developers and researchers should focus on refining self-reflection algorithms to suppress redundant reasoning steps and eliminate hallucinations during image description. While the study's automated scoring is strongly validated by human evaluations reaching 86% to 98% agreement, organizations should remain cautious when interpreting benchmarks from models that refuse direct prompting instructions, as out-of-distribution formatting can distort evaluation metrics.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). M³CoT establishes a prior multimodal, multi-step CoT benchmark, making its scope and evaluation design essential context for MME-CoT’s broader assessment.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This work introduces zero-shot chain-of-thought prompting, the core reasoning approach whose effects MME-CoT tests in multimodal models.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides an earlier standardized evaluation framework for multimodal models, clarifying the benchmark tradition that MME-CoT extends to CoT reasoning.
No sufficiently relevant recommendations were found.
