M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Qiguang ChenLibo QinJin ZhangZhi ChenXiao XuWanxiang Che

article2024ACL123 citations

Introduces M³CoT, a comprehensive multi-domain benchmark that rigorously evaluates vision large language models on authentic multi-step multimodal reasoning and exposes significant performance gaps compared to human capabilities.

Listen

Recent advancements in artificial intelligence have led vision-language models to achieve seemingly superhuman scores on standard visual reasoning tests. However, the article demonstrates that existing evaluation benchmarks overestimate these capabilities because their questions are overly simplistic. Most existing test questions can be solved using text alone, require only a single visual inspection step, or fail to cover critical subjects such as mathematics and commonsense reasoning.

The main objective of the article is to introduce a rigorous benchmark called M3CoT to evaluate multi-domain, multi-step, multi-modal reasoning in advanced vision models, and to establish an accurate assessment of current machine capabilities relative to human performance.

To construct this benchmark, the authors curated 11,459 multi-choice questions across science, mathematics, and commonsense topics. They eliminated questions that did not strictly require visual information and filtered out single-step reasoning problems through automated screening and expert human review. The team augmented missing subject areas using synthetic generation guided by language models, followed by multiple rounds of human quality verification that achieved high annotator agreement. They then evaluated leading proprietary and open-source vision models across various prompt strategies, external tool integrations, and fine-tuning setups.

The evaluation yielded several critical findings. First, existing models struggle substantially with multi-step visual reasoning, showing at least a 29% drop in performance compared to single-step benchmarks. Second, a large performance gap remains between models and people: the highest-performing model, GPT-4V, achieved an overall accuracy of 62.60%, falling well behind the human benchmark of 91.17%. Third, zero-step and prompt-based reasoning only emerge in models with 13 billion parameters or more, with smaller models failing to benefit from step-by-step prompting. Fourth, text-based tool planning and standard in-context prompting largely failed, with tool-assisted systems scoring between 14.60% and 34.29% due to errors in visual tool selection.

These findings indicate that prior claims of artificial intelligence matching human-level visual understanding were premature. Relying on current vision-language models for autonomous, multi-step technical decision-making introduces significant operational and accuracy risks. Tool-use frameworks that plan actions in text without continuous visual feedback are particularly unreliable for complex tasks.

For practitioners and decision-makers, the article recommends prioritizing targeted supervised fine-tuning over basic prompt engineering or disconnected tool frameworks when building visual reasoning systems. Fine-tuning models on multi-step reasoning data yielded major performance gains, enabling even smaller open models to surpass some large zero-shot systems. Developers should focus research efforts on improving intermediate cross-modal attention and high-quality image-text interleaving.

The primary limitations of the work include its exclusive focus on the English language and potential subjectivity in manual annotations, although double-checking procedures minimized labeling errors. Overall, the evidence provides high confidence that current vision models require improved architectures and multi-step training before they can be reliably deployed for complex real-world visual analysis.

Cover for M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Abstract

Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-by-step reasoning, which gains increasing attention. Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT. Motivated by this, we introduce a novel benchmark (M³CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT. Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs). In addition, we highlight that the current VLLMs still struggle to correctly reason in M³CoT and there remains a large gap between existing VLLMs and human performance in M³CoT, despite their superior results on previous MCoT benchmarks. To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT. We hope that M³CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research.

Table of Contents

  • 1 Introduction
  • 2 Problem Formalization
  • 3 Dataset Annotation
  • 3.1 Absence of Visual Modal Reasoning Sample Removal
  • 3.2 Multi-step MCoT Sample Construction
  • 3.4 Quality Assurance
  • 4 Data Analysis
  • 5 Experiments
  • 5.1 Experiments Setting
  • 5.2 Results for M 3 CoT
  • 5.3 Analysis
  • 5.4 Exploration
  • 5.4.1 Tool Usage Exploration
  • 5.4.2 In-Context-Learning Exploration
  • 5.4.3 Finetuning Exploration
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgments
  • References
  • Appendix
  • A Dataset Annotation Details
  • A.1 Statistical Analysis of Existing Datasets
  • A.2 The Details of Absence of Visual Modal Reasoning Sample Removal
  • A.3 The Details of Multi-step MCoT Sample Construction
  • A.4 Domain Augmentation Details
  • A.4.1 Mathematics Domain Augmentation Details
  • A.4.2 Commonsense Domain Augmentation Details
  • A.5 Quality Assurance Details
  • A.5.1 Human Annotation Details
  • A.5.2 Human Recheck Details
  • A.5.3 Rationale Rewriting
  • A.6 Image Redundancy Removal
  • B Experiment Details
  • B.1 Main Result Details
  • B.1.1 Heuristic baselines
  • B.2 Exploration Details
  • B.2.1 Tool Usage Details
  • B.2.2 In-Context-Learning Details
  • B.2.3 Finetuning Details
  • B.3 Error Analysis
  • B.3.1 Zero-shot Chain-of-Thought Error Analysis
  • B.3.2 Tool Usage Error Analysis

Knowls

  1. Knowl 1 — Mathematical Formulation of Multi-Step Multi-Modal Chain-of-Thought Reasoning

    equation

    Multi-step multi-modal Chain-of-Thought (MCoT) reasoning formalizes the task of answering a question given multi-modal context by generating a step-by-step rationale where multiple reasoning steps explicitly require visual evidence.

    Let II denote an input image, QQ denote a textual question, CC denote an optional textual context, and O={o1,o2,…,on}\mathcal{O} = \{o_1, o_2, \dots, o_n\} denote a set of nn candidate options. A textual instruction prompt TT is constructed as:

    T=Prompt(Q,C,O)T = \text{Prompt}(Q, C, \mathcal{O})

    The model generates a step-wise rationale Rm={S1,S2,…,Sm}R_m = \{S_1, S_2, \dots, S_m\} consisting of mm reasoning steps, where each step SiS_i is determined autoregressively by:

    Si=arg⁡max⁡SiP(Si∣I,T,Ri−1)S_i = \arg\max_{S_i} P(S_i \mid I, T, R_{i-1})

    with the conditional step probability defined according to whether the step relies on visual evidence:

    P(Si∣I,T,Ri−1)={P(Si∣T,Ri−1)if Si∉SP(Si∣I,T,Ri−1)if Si∈SP(S_i \mid I, T, R_{i-1}) = \begin{cases} P(S_i \mid T, R_{i-1}) & \text{if } S_i \notin \mathcal{S} \\ P(S_i \mid I, T, R_{i-1}) & \text{if } S_i \in \mathcal{S} \end{cases}

    where S⊆{S1,…,Sm}\mathcal{S} \subseteq \{S_1, \dots, S_m\} represents the set of reasoning steps that strictly require visual modal interaction. A reasoning chain is defined as multi-step multi-modal CoT if and only if ∣aturalS∣≥2| atural \mathcal{S}| \ge 2.

    Finally, the model predicts the answer Y∈OY \in \mathcal{O} conditioned on the rationale:

    Y=arg⁡max⁡o∈OP(o∣Rm)Y = \arg\max_{o \in \mathcal{O}} P(o \mid R_m)

  2. Knowl 2 — M3CoT Benchmark Dataset Specifications and Domain Distribution

    data/table

    The M3CoT\text{M}^3\text{CoT} benchmark is a multi-domain, multi-step, multi-modal chain-of-thought dataset consisting of 11,459 total multiple-choice questions requiring multi-step visual reasoning (∣S∣≥2|\mathcal{S}| \ge 2 for 100% of samples).

    Domain / Split Science Mathematics Commonsense Total
    Domain Size 7,973 1,166 2,163 11,459
    Train Set Split - - - 7,973
    Dev Set Split - - - 1,127
    Test Set Split - - - 2,359
    Topic Classes - - - 17
    Category Classes - - - 263
    Avg. Question Length - - - 14.33 words
    Avg. Choice Length - - - 3.60 words
    Avg. Context Length - - - 19.35 words
    Avg. Rationale Length - - - 293.93 words
    Avg. Reasoning Steps - - - 10.9 steps

    Compared to existing multi-modal QA benchmarks such as ScienceQA (average rationale length 48 words, 2.5 steps, ∼8%\sim 8\% multi-step MCoT), OKVQA (3.0 steps), MMMU (1.0 step), and VCR (1.0 step), M3CoT\text{M}^3\text{CoT} requires a substantially deeper reasoning depth with an average of 10.9 steps per rationale across 17 topics spanning Natural Science, Social Science, Language Science, Algebra, Geometry, Theory, Physical Commonsense, Social Commonsense, and Temporal Commonsense.

  3. Knowl 3 — M3CoT Multi-Stage Data Construction and Quality Assurance Pipeline

    model/method

    The construction of the M3CoT\text{M}^3\text{CoT} benchmark addresses three deficiencies in prior MCoT datasets: absence of visual modal reasoning, single-step visual modal reasoning, and domain missing. The pipeline consists of four stages:

    1. Absence of Visual Reasoning Removal: Text-only masking is applied to candidate samples from ScienceQA. Annotators inspect questions, options, and rationales with images hidden. If a question can be answered purely via text, it is discarded, eliminating over 30% of candidate multi-modal samples.
    2. Multi-Step MCoT Sample Construction: Samples with rationales containing fewer than two steps are automatically filtered using ROSCOE step segmentation. Human annotators then review the text and image to verify that solving the question strictly requires at least two distinct visual reasoning steps (∣S∣≥2|\mathcal{S}| \ge 2).
    3. Multi-Modal Domain Augmentation:
      • Mathematics: Derived from the text-only MATH dataset. gpt-3.5-turbo generates three plausible distractors. Mathematical expressions and geometric code are rendered into PNG images and integrated into multi-modal instances via HTML/CSS.
      • Commonsense: Derived from Sherlock visual clue annotations. An LLM is prompted via one-shot demonstrations to synthesize multiple-choice questions, options, and explanations that simultaneously require at least two distinct visual clues from an image.
      • Science: Sparse topics are augmented using TabMWP, KiloGram, and synthetic figures rendered via Matplotlib.
    4. Quality Assurance and Rationale Rewriting: gpt-3.5-turbo rewrites low-quality candidate rationales for scientific accuracy and clarity. Annotators must pass an onboarding qualification test with ≥80%\ge 80\% accuracy. Two subsequent human verification rounds are conducted; acceptance requires majority agreement (inter-annotator Cohen's kappa κ=0.85\kappa = 0.85, with <5%< 5\% discard rate in the second round).
  4. Knowl 4 — Zero-Shot Performance Disparity Between Proprietary and Open-Source VLLMs on M3CoT

    empirical result

    Zero-shot evaluation of Vision Large Language Models (VLLMs) on the M3CoT\text{M}^3\text{CoT} benchmark reveals a large capability gap between proprietary models (GPT-4V), open-source models, and human performance across Science, Commonsense, and Mathematics domains.

    Model Prompting Strategy Science (Avg) Commonsense (Avg) Total Acc (%)
    Random Baseline - 30.01 25.17 28.56
    InstructBLIP-13B Direct 31.73 61.03 35.94
    InstructBLIP-13B CoT 31.98 61.19 36.07
    LLaVA-V1.5-13B Direct 28.22 34.47 27.05
    LLaVA-V1.5-13B CoT 37.54 60.47 39.52
    CogVLM-17B Direct 38.98 46.32 37.19
    CogVLM-17B CoT 41.43 41.80 38.91
    Gemini Direct 48.80 63.59 45.17
    Gemini CoT 51.10 65.30 47.50
    GPT-4V Direct 59.57 79.22 56.95
    GPT-4V CoT 66.86 80.49 62.60
    Human - 92.42 94.64 91.17

    Key observations:

    • GPT-4V with standard zero-shot CoT ("Let's think step-by-step!") achieves the highest automated accuracy (62.60%), but substantially trails the human ceiling of 91.17% (a 28.57% gap).
    • The best open-source VLLM (LLaVA-V1.5-13B with CoT at 39.52%) lags behind GPT-4V by at least 23.08% overall, and falls short on advanced multi-step visual reasoning.
  5. Knowl 5 — Parameter-Scale Emergence Threshold for Multi-Modal Chain-of-Thought Reasoning

    empirical result

    Zero-shot multi-modal Chain-of-Thought (MCoT) prompting only improves reasoning accuracy on models with parameter counts exceeding 10 billion parameters (≥13B\ge 13\text{B}).

    When evaluating smaller VLLMs on the M3CoT\text{M}^3\text{CoT} test set:

    • Kosmos-2 (2B): Accuracy drops from 23.17% (Direct) to 18.68% (CoT), 0.04% (Desp-CoT), and 0.99% (CCoT).
    • InstructBLIP (7B): Accuracy drops from 36.11% (Direct) to 35.76% (CoT).
    • LLaVA-V1.5 (7B): Accuracy drops from 36.63% (Direct) to 35.81% (CoT), 34.43% (Desp-CoT), and 35.72% (CCoT).

    In contrast, larger models show consistent performance gains when utilizing CoT reasoning over direct answer prediction:

    • InstructBLIP (13B): Accuracy rises from 35.94% (Direct) to 36.07% (CoT).
    • LLaVA-V1.5 (13B): Accuracy rises from 27.05% (Direct) to 39.52% (CoT, a +12.47%+12.47\% gain).
    • CogVLM (17B): Accuracy rises from 37.19% (Direct) to 38.91% (CoT).
    • Gemini: Accuracy rises from 45.17% (Direct) to 47.50% (CoT).
    • GPT-4V: Accuracy rises from 56.95% (Direct) to 62.60% (CoT).
  6. Knowl 6 — Failure of Text-Only Tool-Augmented LLMs on Multi-Modal CoT Tasks

    empirical result

    Multi-modal tool-augmented LLM systems that execute visual question answering via text-based orchestration fail on complex multi-step multi-modal CoT reasoning in M3CoT\text{M}^3\text{CoT}.

    Framework Science (Avg) Commonsense (Avg) Mathematics (Avg) Total Acc (%)
    Random Baseline 30.01 25.17 29.01 28.56
    HuggingGPT 16.28 11.07 14.46 14.60
    VisualChatGPT 24.72 35.58 23.94 25.92
    IdealGPT 29.86 44.45 29.56 32.19
    Chameleon 31.79 41.74 22.60 34.29
    GPT-4V (CoT) 66.86 80.49 44.60 62.60

    Key findings:

    • All tested tool-usage frameworks lag behind GPT-4V CoT by at least 28.31 percentage points. HuggingGPT (14.60%) and VisualChatGPT (25.92%) score below the random guessing baseline (28.56%).
    • Qualitative error analysis indicates that because the text-modal controller cannot directly observe the image during planning, it makes critical tool selection errors (such as selecting image classification instead of visual QA), introduces redundant tool invocations, and generates hallucinated sub-question plans that cascade into complete reasoning failures.
  7. Knowl 7 — Fine-Tuning Traditional Vision-Language Models and VLLMs on M3CoT

    empirical result

    Supervised fine-tuning on the M3CoT\text{M}^3\text{CoT} training set substantially enhances reasoning capabilities across both traditional small Vision-Language Models (VLMs) and Vision Large Language Models (VLLMs), allowing fine-tuned small models to outperform zero-shot large models.

    Model Architecture Category Science Commonsense Math Total Acc (%)
    MM-CoTbase_{\text{base}} Trad. VLM 42.70 49.30 37.38 44.85
    MC-CoTbase_{\text{base}} Trad. VLM 53.70 53.45 35.06 53.51
    MM-CoTlarge_{\text{large}} Trad. VLM 46.42 53.89 43.51 48.73
    MMR Trad. VLM 48.04 58.43 51.27 50.67
    MC-CoTlarge_{\text{large}} Trad. VLM 53.55 58.28 44.88 57.69
    LLaMA-Adapter-7B VLLM 55.02 69.65 35.85 54.89
    LLaVA-V1.5-7B VLLM 58.15 68.16 33.14 56.74
    LLaVA-V1.5-13B VLLM 60.66 72.41 39.60 59.50
    CogVLM-17B VLLM 57.50 74.12 43.19 58.25
    GPT-4V (CoT, Zero-Shot) Proprietary 66.86 80.49 44.60 62.60
    Human Upper Bound 94.92 92.47 87.23 91.61

    Key takeaways:

    • Fine-tuned traditional VLMs (scoring 44.85%–57.69%) and fine-tuned VLLMs (scoring 54.89%–59.50%) significantly outperform zero-shot open-source VLLMs (which max out at 38.91% for CogVLM-17B) and match or exceed zero-shot Gemini (47.50%).
    • Fine-tuning larger VLLMs (e.g., LLaVA-1.5-13B at 59.50%) yields stronger results than smaller models, approaching the zero-shot performance of GPT-4V (62.60%).
  8. Knowl 8 — In-Context Learning Dynamics with Textual and Interleaved Multi-Modal Demonstrations

    empirical result

    Evaluating few-shot In-Context Learning (ICL) on M3CoT\text{M}^3\text{CoT} demonstrates that standard textual and interleaved demonstrations fail to resolve multi-step multi-modal reasoning:

    1. Text-Only Demonstrations: Providing 1- to 4-shot in-domain text-only demonstrations does not meaningfully improve multi-modal performance:

      • LLaVA-V1.5-13B: 36.62% (1-shot) →\to 36.62% (2-shot) →\to 35.37% (3-shot) →\to 35.76% (4-shot).
      • OpenFlamingo-7B: 22.40% (1-shot) →\to 25.51% (2-shot) →\to 24.08% (3-shot) →\to 22.88% (4-shot).
      • GPT-4V: 56.61% (1-shot) →\to 54.46% (2-shot) →\to 57.78% (3-shot) →\to 56.66% (4-shot).
    2. Image-Text Interleaved Demonstrations: Providing full multi-modal interleaved demonstrations harms models not natively trained on multi-image interleaved contexts, while giving only minor gains to GPT-4V:

      • LLaVA-V1.5-13B: Accuracy degrades sharply from 35.07% (1-shot) to 35.11% (2-shot), 33.91% (3-shot), and 19.65% (4-shot).
      • OpenFlamingo-7B: Accuracy degrades from 25.61% (1-shot) to 24.17% (2-shot), 24.08% (3-shot), and 14.13% (4-shot).
      • GPT-4V: Accuracy increases slightly from 51.62% (1-shot) to 52.09% (2-shot), 52.09% (3-shot), and 54.16% (4-shot), but remains lower than zero-shot direct CoT (62.60%).
  9. Knowl 9 — Impact of Reasoning Steps, Multi-Modal Interactions, and Rationale Quality on M3CoT Accuracy

    empirical result

    Analytical experiments on M3CoT\text{M}^3\text{CoT} identify three core determinants of multi-modal reasoning performance:

    1. Reasoning Step Complexity Penalty: Multi-step MCoT poses a severe out-of-distribution challenge relative to single-step MCoT. When comparing performance on single-step ScienceQA instances versus multi-step M3CoT\text{M}^3\text{CoT} instances using similar images, models experience at least a 29.06% absolute drop in accuracy (CogVLM: 60.69% →\to 31.54%; LLaVA-13B: 68.02% →\to 38.96%; GPT-4V: 91.64% →\to 60.17%). Accuracy across all evaluated VLLMs monotonically declines as the number of annotated reasoning steps increases from <5<5 steps to >25>25 steps.
    2. Multi-Modal Interaction Step Correlation: By computing cross-modal similarity between step texts and image regions, the number of required visual interaction steps is quantified. Total question-answering accuracy exhibits a strict positive correlation with average multi-modal interaction depth, rising from 35.81% (at 0.6 average interaction steps) to 39.52% (1.0 step), 62.60% (1.6 steps), and 91.17% (2.5 steps for human performance).
    3. Rationale Quality Correlation: Evaluating model-generated rationales using ROSCOE metrics (Semantic Coverage, Step Diversity, Key Step, Informativeness, and Knowledge Usage) shows that higher scores across all five dimensions correlate strongly with higher final task accuracy.
  10. Knowl 10 — Stated Limitations of the M3CoT Benchmark

    limitation

    The authors identify three primary limitations of the M3CoT\text{M}^3\text{CoT} benchmark:

    1. Monolingual Limitation: Due to regional and resource constraints, the benchmark is constructed entirely in English and does not evaluate multi-lingual multi-modal chain-of-thought capabilities.
    2. Subjectivity in Human Annotation: Despite multi-stage verification and qualification testing, manual annotation of reasoning chains, multi-modal dependency verification, and rationale selection inherently introduce human subjective biases.
    3. Proprietary Model Dependency: Heavy reliance on proprietary black-box APIs (e.g., GPT-4V and Gemini) introduces potential reproducibility issues over time due to API versioning, model deprecation, or retirement.

Coverage note — None was omitted; all key contributions including problem formalization, dataset construction/statistics, zero-shot evaluation, parameter emergence, tool-usage failure analysis, in-context learning, fine-tuning benchmarks, rationale/interaction depth analysis, and limitations are covered.

References

  1. 1.Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. 2023. Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv preprint arXiv:2308.01390.
  2. 2.Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. 2023. Measuring and improving chain-of-thought reasoning in vision-language models. arXiv preprint arXiv:2309.04461.
  3. 3.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning.
  4. 4.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. 2023. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394.
  5. 5.Jiaxin Ge, Hongyin Luo, Siyuan Qian, Yulu Gan, Jie Fu, and Shanghang Zhan. 2023. Chain of thought prompt tuning in vision language models. arXiv preprint arXiv:2304.07919.
  6. 6.Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023. ROSCOE: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations.
  7. 7.Google. 2023. Gemini: A family of highly capable multimodal models.
  8. 8.Liqi He, Zuchao Li, Xiantao Cai, and Ping Wang. 2023. Multi-modal latent space learning for chain-of-thought reasoning in language models. arXiv preprint arXiv:2312.08762.
  9. 9.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. NeurIPS.
  10. 10.Jack Hessel, Jena D Hwang, Jae Sung Park, Rowan Zellers, Chandra Bhagavatula, Anna Rohrbach, Kate Saenko, and Yejin Choi. 2022. The abduction of sherlock holmes: A dataset for visual abductive reasoning. In European Conference on Computer Vision, pages 558–575.
  11. 11.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  12. 12.Mengkang Hu, Yao Mu, Xinmiao Chelsey Yu, Mingyu Ding, Shiguang Wu, Wenqi Shao, Qiguang Chen, Bin Wang, Yu Qiao, and Ping Luo. 2024. Tree-planner: Efficient close-loop task planning with large language models. In The Twelfth International Conference on Learning Representations.
  13. 13.Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert Hawkins, and Yoav Artzi. 2022. Abstract visual reasoning with tangram shapes. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 582–601.
  14. 14.Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern recognition, pages 4999–5007.
  15. 15.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  16. 16.J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics, pages 159–174.
  17. 17.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. 2023a. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125.
  18. 18.Yunxin Li, Longyue Wang, Baotian Hu, Xinyu Chen, Wanqi Zhong, Chenyang Lyu, and Min Zhang. 2023b. A comprehensive evaluation of gpt-4v on knowledge-intensive visual question answering. arXiv preprint arXiv:2311.07536.
  19. 19.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023. Improved baselines with visual instruction tuning. arXiv preprint arXiv:2310.03744.
  20. 20.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2023a. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255.
  21. 21.Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-chun Zhu. 2021. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6774–6786.
  22. 22.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022a. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521.
  23. 23.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022b. Learn to explain: Multimodal reasoning via thought chains for science question answering. In NeurIPS.
  24. 24.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023b. Chameleon: Plug-and-play compositional reasoning with large language models. CoRR, abs/2304.09842.
  25. 25.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2022c. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610.
  26. 26.Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. 2023. Compositional chain-of-thought prompting for large multimodal models. arXiv preprint arXiv:2311.17076.
  27. 27.Debjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh, and Godawari Sudhakar Rao. 2024. Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning. arXiv preprint arXiv:2401.12863.
  28. 28.OpenAI. 2023. Gpt-4 technical report.
  29. 29.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. ArXiv, abs/2306.
  30. 30.Libo Qin, Qiguang Chen, Xiachong Feng, Yang Wu, Yongheng Zhang, Yinghui Li, Min Li, Wanxiang Che, and Philip S. Yu. 2024a. Large language models meet nlp: A survey.
  31. 31.Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che. 2023. Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. arXiv preprint arXiv:2310.14799.
  32. 32.Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S Yu. 2024b. Multilingual large language model: A survey of resources, taxonomy and frontiers. arXiv preprint arXiv:2404.04925.
  33. 33.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR.
  34. 34.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. In European Conference on Computer Vision, pages 146–162. Springer.
  35. 35.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.
  36. 36.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022. Language models are multilingual chain-of-thought reasoners. arXiv preprint arXiv:2210.03057.
  37. 37.Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le, and Gust Verbruggen. 2023. Assessing gpt4-v on structured reasoning tasks. arXiv preprint arXiv:2312.11524.
  38. 38.Cheng Tan, Jingxuan Wei, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Xihong Yang, and Stan Z Li. 2023. Boosting the power of small multimodal reasoning models to match larger models with self-consistency training. arXiv preprint arXiv:2311.14109.
  39. 39.Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen. 2023a. T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering. arXiv preprint arXiv:2305.03453.
  40. 40.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023b. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2609–2634, Toronto, Canada. Association for Computational Linguistics.
  41. 41.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang. 2023c. Cogvlm: Visual expert for pretrained language models.
  42. 42.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  43. 43.Zefeng Wang, Zhen Han, Shuo Chen, Volker Tresp, and Jindong Gu. 2023d. Towards the adversarial robustness of vision-language model with chain-of-thought reasoning.
  44. 44.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research.
  45. 45.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  46. 46.Jingxuan Wei, Cheng Tan, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Bihui Yu, Ruifeng Guo, and Stan Z Li. 2023. Enhancing human-like multi-modal reasoning: A new challenging dataset and comprehensive framework. arXiv preprint arXiv:2307.12626.
  47. 47.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023a. Visual chatgpt: Talking, drawing and editing with visual foundation models.
  48. 48.Yifan Wu, Pengchuan Zhang, Wenhan Xiong, Barlas Oguz, James C Gee, and Yixin Nie. 2023b. The role of chain-of-thought in complex vision-language reasoning task. arXiv preprint arXiv:2311.09193.
  49. 49.Fanglong Yao, Changyuan Tian, Jintao Liu, Zequn Zhang, Qing Liu, Li Jin, Shuchao Li, Xiaoyu Li, and Xian Sun. 2023. Thinking like an expert: Multimodal hypergraph-of-thought (hot) reasoning to boost foundation modals. arXiv preprint arXiv:2308.06207.
  50. 50.Haoxuan You, Rui Sun, Zhecan Wang, Long Chen, Gengyu Wang, Hammad A. Ayyubi, Kai-Wei Chang, and Shih-Fu Chang. 2023. Idealgpt: Iteratively decomposing vision and language reasoning via large language models.
  51. 51.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490.
  52. 52.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2023. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502.
  53. 53.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488.
  54. 54.Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6720–6731.
  55. 55.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023a. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199.
  56. 56.Zhehao Zhang, Xitao Li, Yan Gao, and Jian-Guang Lou. 2023b. CRT-QA: A dataset of complex reasoning question answering over tabular data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2131–2153, Singapore. Association for Computational Linguistics.
  57. 57.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023c. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  58. 58.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023d. Multimodal chain-of-thought reasoning in language models. CoRR, abs/2302.00923.
  59. 59.Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.
  60. 60.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.
  61. 61.Ziyu Zhuang, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Zixian Feng, Weinan Zhang, and Ting Liu. 2023. Through the lens of core competency: Survey on evaluation of large language models. arXiv preprint arXiv:2308.07902.

Citation

MLA
Chen, Q., et al. “M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought”. arXiv, 2024, http://arxiv.org/abs/2405.16473v1.
APA
Chen, Q., Qin, L., Zhang, J., Chen, Z., Xu, X., & Che, W. (2024). M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought. arXiv. http://arxiv.org/abs/2405.16473v1
Chicago
Chen, Q., L. Qin, J. Zhang, Z. Chen, X. Xu, and W. Che. 2024. “M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought”. arXiv. http://arxiv.org/abs/2405.16473v1.
Harvard
Chen, Q. et al. (2024) “M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.16473v1.
Vancouver
1. Chen Q, Qin L, Zhang J, Chen Z, Xu X, Che W (2024) M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought. arXiv

BibTeX

@article{chen2024cot,
  title = {M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought},
  author = {Chen, Qiguang and Qin, Libo and Zhang, Jin and Chen, Zhi and Xu, Xiao and Che, Wanxiang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.16473v1},
  eprint = {2405.16473}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/