M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Qiguang ChenLibo QinJin ZhangZhi ChenXiao XuWanxiang Che
Introduces M³CoT, a comprehensive multi-domain benchmark that rigorously evaluates vision large language models on authentic multi-step multimodal reasoning and exposes significant performance gaps compared to human capabilities.
Recent advancements in artificial intelligence have led vision-language models to achieve seemingly superhuman scores on standard visual reasoning tests. However, the article demonstrates that existing evaluation benchmarks overestimate these capabilities because their questions are overly simplistic. Most existing test questions can be solved using text alone, require only a single visual inspection step, or fail to cover critical subjects such as mathematics and commonsense reasoning.
The main objective of the article is to introduce a rigorous benchmark called M3CoT to evaluate multi-domain, multi-step, multi-modal reasoning in advanced vision models, and to establish an accurate assessment of current machine capabilities relative to human performance.
To construct this benchmark, the authors curated 11,459 multi-choice questions across science, mathematics, and commonsense topics. They eliminated questions that did not strictly require visual information and filtered out single-step reasoning problems through automated screening and expert human review. The team augmented missing subject areas using synthetic generation guided by language models, followed by multiple rounds of human quality verification that achieved high annotator agreement. They then evaluated leading proprietary and open-source vision models across various prompt strategies, external tool integrations, and fine-tuning setups.
The evaluation yielded several critical findings. First, existing models struggle substantially with multi-step visual reasoning, showing at least a 29% drop in performance compared to single-step benchmarks. Second, a large performance gap remains between models and people: the highest-performing model, GPT-4V, achieved an overall accuracy of 62.60%, falling well behind the human benchmark of 91.17%. Third, zero-step and prompt-based reasoning only emerge in models with 13 billion parameters or more, with smaller models failing to benefit from step-by-step prompting. Fourth, text-based tool planning and standard in-context prompting largely failed, with tool-assisted systems scoring between 14.60% and 34.29% due to errors in visual tool selection.
These findings indicate that prior claims of artificial intelligence matching human-level visual understanding were premature. Relying on current vision-language models for autonomous, multi-step technical decision-making introduces significant operational and accuracy risks. Tool-use frameworks that plan actions in text without continuous visual feedback are particularly unreliable for complex tasks.
For practitioners and decision-makers, the article recommends prioritizing targeted supervised fine-tuning over basic prompt engineering or disconnected tool frameworks when building visual reasoning systems. Fine-tuning models on multi-step reasoning data yielded major performance gains, enabling even smaller open models to surpass some large zero-shot systems. Developers should focus research efforts on improving intermediate cross-modal attention and high-quality image-text interleaving.
The primary limitations of the work include its exclusive focus on the English language and potential subjectivity in manual annotations, although double-checking procedures minimized labeling errors. Overall, the evidence provides high confidence that current vision models require improved architectures and multi-step training before they can be reliably deployed for complex real-world visual analysis.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). This seminal work establishes the ScienceQA benchmark and demonstrates multimodal chain-of-thought reasoning, providing the foundational conceptual paradigm that M³CoT expands into a more rigorous multi-step benchmark.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). This paper presents the MMMU benchmark for college-level multi-discipline multimodal reasoning, setting the standard for evaluating expert AGI capabilities that M³CoT seeks to assess via multi-step chains.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). This paper introduces MathVista for evaluating mathematical reasoning in multimodal visual contexts, directly informing M³CoT's focus on complex, multi-domain visual reasoning across math and science.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). This work introduces MMBench to evaluate fine-grained multimodal perception and reasoning, addressing standard evaluation pitfalls that motivate M³CoT's stricter multi-step filtering.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). This paper develops the MME benchmark to evaluate multimodal perception and cognition, establishing the baseline methodologies for assessing vision-language models prior to multi-step chain-of-thought analysis.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). This paper introduces MM-Vet to measure integrated core multimodal capabilities, highlighting the challenges of composite vision-language tasks that M³CoT formalizes into multi-step reasoning chains.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This foundational text introduces zero-shot chain-of-thought reasoning in language models, establishing the step-by-step reasoning framework that M³CoT adapts and evaluates in multimodal domains.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). This study addresses language-only shortcuts in visual question answering datasets, establishing the rationale for M³CoT's strict elimination of questions solvable by text alone.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). This paper builds directly on multimodal chain-of-thought benchmarking by introducing a framework and dataset that operationalizes visual step-by-step reasoning via focal region localization.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). This study applies advanced scaling, data filtering, and test-time multi-step reasoning strategies to open-source vision-language models, directly tackling the performance gaps highlighted in M³CoT.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). This work explores native pre-training and test-time reasoning recipes in multimodal large language models to address the multi-discipline and multi-step reasoning challenges identified by M³CoT.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report details the Qwen3-VL architecture and its reinforcement learning techniques tailored for extended chain-of-thought reasoning across complex multimodal tasks.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). This text advances chain-of-thought reasoning by modeling deliberate System 2 search and backtracking, extending the step-by-step reasoning paradigms evaluated in benchmarks like M³CoT.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). This work extends multi-task, multi-step reasoning evaluation from the vision-language domain into the complex audio-language domain with the MMAU benchmark.
- Paper: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, Yuxin Zuo et al. (2025). This benchmark extends multi-step multimodal evaluation into specialized clinical decision-making, probing expert reasoning where simple perceptual shortcuts fail.
