T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering
Lei WangYi HuJiabang HeXing XuNing LiuHui LiuHeng Tao Shen
Proposes a data-generation and mixing framework that uses large language model reasoning signals to teach smaller multimodal models how to solve complex scientific questions, achieving a new state of the art on ScienceQA.
Complex scientific question answering requires artificial intelligence systems to interpret multimodal information—such as text, diagrams, and maps—while performing multi-step reasoning. Traditional approaches rely on training smaller models using human-annotated step-by-step explanations (known as chain-of-thought rationales). However, manual annotation is labor-intensive, expensive, and frequently lacks the broader external knowledge required to solve complex problems accurately.
The article demonstrates a new training approach called T-SciQ, which uses large language models to generate high-quality reasoning explanations and systematically teach compact student models to solve multimodal science questions.
The researchers developed a three-stage framework using an advanced teacher model to generate two distinct types of explanations: standard step-by-step reasoning for straightforward questions and plan-based reasoning that breaks complex problems into simpler subtasks. The framework applies a data-mixing strategy that uses a validation set to select the optimal explanation type for each skill category. These combined signals are then used to train smaller student models (under 1 billion parameters, over 200 times smaller than the teacher model) across a standard multimodal benchmark consisting of 21,208 science questions, as well as six additional language reasoning benchmarks.
The analysis produced several key findings. First, the primary student model achieved a new state-of-the-art accuracy of 96.18% on the multimodal benchmark, outperforming human performance (88.40%), the leading multimodal baseline (91.68%), and large few-shot models like GPT-4 (82.69%). Second, student models consistently outperformed baselines trained on human-annotated explanations across different model sizes and architectures, yielding absolute gains of 4.5% to 6.84%. Third, combining standard and plan-based explanations delivered superior accuracy compared to using either explanation style in isolation. Finally, the framework generalized effectively across six diverse text reasoning benchmarks, substantially improving performance in arithmetic, commonsense, and logic tasks.
These results demonstrate that compact, cost-efficient artificial intelligence models can surpass human benchmarks and massive commercial systems when trained on structured, model-generated explanations. Organizations can drastically lower operational deployment costs and latency by replacing human annotation pipelines with synthetic teaching data while achieving superior multi-step reasoning and open-world knowledge integration.
Decision-makers should consider adopting synthetic reasoning generation and dynamic data-mixing strategies when deploying smaller, task-specific models. For immediate next steps, technical teams should explore parameter-efficient fine-tuning techniques and evaluate different foundation models as teachers to optimize training costs.
The primary limitation noted in the article is the current reliance on full fine-tuning across specific model architectures and proprietary teacher model interfaces. However, given the consistent outperformance across diverse question categories, model architectures, and task benchmarks, there is high confidence in the robustness and practical effectiveness of the proposed approach.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Introduces the ScienceQA benchmark and establishes the foundational multimodal chain-of-thought paradigm for science question answering that T-SciQ directly aims to teach and outperform.
- Paper: Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes, Cheng-Yu Hsieh et al. (2023). Develops the core methodology of distilling LLM-generated rationales into compact student models to achieve high reasoning performance with reduced model capacity.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Establishes a knowledge distillation pipeline transferring multi-step chain-of-thought reasoning from massive teacher models to sub-billion parameter student models.
- Paper: Is a Question Decomposition Unit All We Need?, Pruthvi Patel et al. (2022). Pioneers the question decomposition approach that informs the plan-based subtask reasoning strategy utilized within T-SciQ's three-stage framework.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Provides the foundational demonstration that chain-of-thought prompting elicits multi-step reasoning capabilities in large language models.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Presents the STaR framework for bootstrapping reasoning from model-generated rationales, laying critical groundwork for using synthetic explanations as training signals.
- Paper: Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models, Zhihong Shao et al. (2023). Demonstrates how large models can self-synthesize high-quality reasoning demonstrations, prefiguring T-SciQ's replacement of human annotations with LLM signals.
- Paper: Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks, Minki Kang et al. (2023). Explores distilling knowledge-augmented reasoning paths to compact models, directly addressing the knowledge-retrieval challenges inherent in science QA.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). Presents M³CoT, a multi-domain benchmark that pushes multimodal chain-of-thought evaluation beyond standard datasets by requiring genuine multi-step visual reasoning.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). Extends multimodal scientific reasoning to broader discovery domains like chemistry and biology using structured reasoning recipes in compact models.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). Provides a comprehensive survey synthesizing how scientific language and multimodal models are trained and applied across scientific discovery workflows.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). Addresses the efficiency and verbosity challenges of learned chain-of-thought traces by introducing elastic length compression for reasoning paths.
- Paper: Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models, Xingyu Chen et al. (2025). Critically analyzes and mitigates the overthinking behavior that occurs when models generate extended reasoning traces for simpler questions.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). Synthesizes the broader landscape of advancing reasoning models through synthetic data generation, distillation, and reinforcement learning.
