MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Jiawei GuoTianyu ZhengYizhi LiYuelin BaiBo LiYubo WangKing ZhuGraham NeubigWenhu ChenXiang Yue
Presents a cost-effective data rewriting and filtering pipeline that uses only open-weight models to construct a 12-million-sample visual instruction dataset with chain-of-thought rationales, substantially boosting open-source multimodal model accuracy on complex reasoning benchmarks like MathVerse, MMMU-Pro, and MuirBench.
Multimodal artificial intelligence models that process both images and text have advanced rapidly, yet open-source versions continue to struggle with complex, multi-step reasoning. This limitation largely stems from existing training datasets, which rely on simple academic question-answering formats that provide brief phrase answers without explaining intermediate logical steps. While generating detailed chain-of-thought rationales using human annotators or commercial proprietary models is effective, both approaches present significant financial, labor, and licensing barriers.
The article demonstrates a scalable, cost-effective framework to generate high-quality multimodal instruction-tuning data using exclusively open-weight models, and evaluates how fine-tuning an 8-billion-parameter model on this curated data enhances complex visual reasoning across single-image, multi-image, and video domains.
The authors established a three-stage automated data pipeline. First, they collected and screened 153 public multimodal datasets comprising image-text pairs across 10 categories, filtering out low-quality sources. Second, they used open models to rewrite brief question-answer pairs into rich, step-by-step reasoning dialogues tailored to specific tasks. Third, they implemented an automated filtering step where the open model verified the factual consistency of the rewritten content against the original images to eliminate generated inaccuracies. Using this pipeline, they assembled a 12-million-instance dataset (mixed at a 70:30 ratio of rewritten to original data) and trained MAmmoTH-VL-8B using a three-stage fine-tuning schedule across 23 standard evaluation benchmarks.
The resulting model demonstrated substantial performance gains over competing models. First, MAmmoTH-VL-8B outperformed leading open-source models in the 10-billion-parameter class on reasoning-intensive benchmarks, achieving an 8.1% gain on MathVerse, 7.1% on MMMU-Pro, and 4.4% on MathVista. Second, on multi-image tasks, the model achieved a 13.3% improvement on MuirBench. Third, general visual perception tasks showed gains up to 4%, including improvements on chart and document benchmarks like ChartQA (+2.1%) and AI2D (+2.4%). Finally, ablation analyses revealed that automated self-filtering was essential: optical character recognition and chart data suffered rejection rates of 54.9% and 48.4% respectively, and removing these flawed samples dramatically boosted downstream model accuracy.
These findings prove that competitive multimodal reasoning can be elicited using fully open pipelines without relying on expensive proprietary model APIs or manual annotations. For organizations developing vision-language systems, this significantly lowers development costs and regulatory licensing risks while delivering performance that approaches much larger systems. The evidence also highlights that model-based verification is a practical, effective safeguard against visual hallucinations in synthetic training pipelines.
Organizations aiming to build or deploy multimodal models should adopt task-aware data rewriting and self-filtering workflows, while utilizing a balanced mix of synthetic reasoning data and original ground-truth samples to maintain task diversity. Further work should prioritize scaling up multi-image and video training collections beyond the current 1-million-sample size, as the model showed a slight performance lag behind leading specialized systems on long-form video tasks due to limited compute during training.
Confidence in these findings is supported by rigorous contamination checks confirming zero overlap between training and benchmark sets, as well as high agreement between automated filtering and human evaluation (Cohen's Kappa of 0.64). However, caution is advised in high-stakes operational environments, as open-model-generated rationales can still contain subtle visual errors, and the current 8-billion-parameter model exhibits minor performance trade-offs in fine-grained attribute detection during late-stage training.
- Paper: MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning, Zhiyang Xu et al. (2023). MultiInstruct establishes how diverse multimodal instruction tuning can improve zero-shot transfer, the dataset-design premise that MAmmoTH-VL scales up with reasoning-rich responses.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Learn to Explain shows how multimodal science answers can be trained with intermediate thought chains, providing an early precedent for MAmmoTH-VL’s rationale-supervised approach.
- Paper: The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning, Seungone Kim et al. (2023). The CoT Collection demonstrates that diverse instruction-tuned rationale data can elicit reasoning across tasks, a key foundation for MAmmoTH-VL’s large-scale rationale dataset.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT pairs multimodal questions with intermediate reasoning and visual evidence, clarifying the rationale-augmented training direction that MAmmoTH-VL broadens.
- Paper: Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations, Eunkyu Park et al. (2025). Cognitive Chain-of-Thought extends multimodal reasoning traces into explicit perception, situation, and norm stages, building on the rationale-tuning approach at MAmmoTH-VL’s core.
- Paper: Imagine While Reasoning in Space: Multimodal Visualization-of-Thought, Chengzu Li et al. (2025). Multimodal Visualization-of-Thought carries intermediate reasoning beyond text into generated images, extending the rationale-based multimodal training direction explored by MAmmoTH-VL.
