MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Jiawei GuoTianyu ZhengYizhi LiYuelin BaiBo LiYubo WangKing ZhuGraham NeubigWenhu ChenXiang Yue

article2025ACL1 citations

Presents a cost-effective data rewriting and filtering pipeline that uses only open-weight models to construct a 12-million-sample visual instruction dataset with chain-of-thought rationales, substantially boosting open-source multimodal model accuracy on complex reasoning benchmarks like MathVerse, MMMU-Pro, and MuirBench.

Listen

Multimodal artificial intelligence models that process both images and text have advanced rapidly, yet open-source versions continue to struggle with complex, multi-step reasoning. This limitation largely stems from existing training datasets, which rely on simple academic question-answering formats that provide brief phrase answers without explaining intermediate logical steps. While generating detailed chain-of-thought rationales using human annotators or commercial proprietary models is effective, both approaches present significant financial, labor, and licensing barriers.

The article demonstrates a scalable, cost-effective framework to generate high-quality multimodal instruction-tuning data using exclusively open-weight models, and evaluates how fine-tuning an 8-billion-parameter model on this curated data enhances complex visual reasoning across single-image, multi-image, and video domains.

The authors established a three-stage automated data pipeline. First, they collected and screened 153 public multimodal datasets comprising image-text pairs across 10 categories, filtering out low-quality sources. Second, they used open models to rewrite brief question-answer pairs into rich, step-by-step reasoning dialogues tailored to specific tasks. Third, they implemented an automated filtering step where the open model verified the factual consistency of the rewritten content against the original images to eliminate generated inaccuracies. Using this pipeline, they assembled a 12-million-instance dataset (mixed at a 70:30 ratio of rewritten to original data) and trained MAmmoTH-VL-8B using a three-stage fine-tuning schedule across 23 standard evaluation benchmarks.

The resulting model demonstrated substantial performance gains over competing models. First, MAmmoTH-VL-8B outperformed leading open-source models in the 10-billion-parameter class on reasoning-intensive benchmarks, achieving an 8.1% gain on MathVerse, 7.1% on MMMU-Pro, and 4.4% on MathVista. Second, on multi-image tasks, the model achieved a 13.3% improvement on MuirBench. Third, general visual perception tasks showed gains up to 4%, including improvements on chart and document benchmarks like ChartQA (+2.1%) and AI2D (+2.4%). Finally, ablation analyses revealed that automated self-filtering was essential: optical character recognition and chart data suffered rejection rates of 54.9% and 48.4% respectively, and removing these flawed samples dramatically boosted downstream model accuracy.

These findings prove that competitive multimodal reasoning can be elicited using fully open pipelines without relying on expensive proprietary model APIs or manual annotations. For organizations developing vision-language systems, this significantly lowers development costs and regulatory licensing risks while delivering performance that approaches much larger systems. The evidence also highlights that model-based verification is a practical, effective safeguard against visual hallucinations in synthetic training pipelines.

Organizations aiming to build or deploy multimodal models should adopt task-aware data rewriting and self-filtering workflows, while utilizing a balanced mix of synthetic reasoning data and original ground-truth samples to maintain task diversity. Further work should prioritize scaling up multi-image and video training collections beyond the current 1-million-sample size, as the model showed a slight performance lag behind leading specialized systems on long-form video tasks due to limited compute during training.

Confidence in these findings is supported by rigorous contamination checks confirming zero overlap between training and benchmark sets, as well as high agreement between automated filtering and human evaluation (Cohen's Kappa of 0.64). However, caution is advised in high-stakes operational environments, as open-model-generated rationales can still contain subtle visual errors, and the current 8-billion-parameter model exhibits minor performance trade-offs in fine-grained attribute detection during late-stage training.

Cover for MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Abstract

Open-source multimodal large language models (MLLMs) have shown significant potential in a broad range of tasks. However, their reasoning capabilities remain constrained by existing instruction-tuning datasets, which were predominately repurposed from academic datasets such as VQA, AI2D, and ChartQA. These datasets target simplistic tasks, and only provide phrase-level answers without any intermediate rationales. To address these challenges, we introduce a scalable and cost-effective method to construct a large-scale multimodal instruction-tuning dataset with rich intermediate rationales designed to elicit CoT reasoning. Using only open models, we create a dataset containing 12M instruction-response pairs to cover diverse reasoning-intensive tasks. Experiments demonstrate that training MLLMs on our dataset not only significantly improves reasoning capabilities, achieving state-of-the-art performance on benchmarks such as MathVerse (+8.1%), MMMU-Pro (+7%), and MuirBench (+13.3%), but also gains improvements of up to 4% on non-reasoning-based benchmarks.

Citation

MLA
Guo, J., et al. “MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 13869–920, https://doi.org/10.18653/v1/2025.acl-long.680.
APA
Guo, J., Zheng, T., Li, Y., Bai, Y., Li, B., Wang, Y., Zhu, K., Neubig, G., Chen, W., & Yue, X. (2025). MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13869–13920. https://doi.org/10.18653/v1/2025.acl-long.680
Chicago
Guo, J., T. Zheng, Y. Li, et al. 2025. “MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13869–920. https://doi.org/10.18653/v1/2025.acl-long.680.
Harvard
Guo, J. et al. (2025) “MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13869–13920. Available at: https://doi.org/10.18653/v1/2025.acl-long.680.
Vancouver
1. Guo J, Zheng T, Li Y, Bai Y, Li B, Wang Y, Zhu K, Neubig G, Chen W, Yue X (2025) MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13869–13920

BibTeX

@inproceedings{guo-etal-2025-mammoth,
    title = "{MA}mmo{TH}-{VL}: Eliciting Multimodal Reasoning with Instruction Tuning at Scale",
    author = "Guo, Jiawei  and
      Zheng, Tianyu  and
      Li, Yizhi  and
      Bai, Yuelin  and
      Li, Bo  and
      Wang, Yubo  and
      Zhu, King  and
      Neubig, Graham  and
      Chen, Wenhu  and
      Yue, Xiang",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.680/",
    doi = "10.18653/v1/2025.acl-long.680",
    pages = "13869--13920",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/