Teaching Small Language Models to Reason
Lucie Charlotte MagisterJonathan MallinsonJakub AdámekEric MalmiAliaksei Severyn
Demonstrates that fine-tuning smaller language models on chain-of-thought rationales distilled from giant teacher models significantly improves their reasoning capabilities across arithmetic, commonsense, and symbolic benchmarks.
Large language models can solve complex multi-step problems using chain-of-thought prompting, which encourages them to generate intermediate reasoning steps before providing an answer. However, this capability typically emerges only in massive models with tens or hundreds of billions of parameters. Smaller models under 10 billion parameters often fail to generate logical intermediate reasoning and can even experience decreased accuracy when prompted this way. Because deploying giant models is computationally expensive and difficult to scale in production environments, transferring these reasoning abilities to smaller, more efficient models has become an important priority.
The article demonstrates a two-step knowledge distillation pipeline that transfers multi-step reasoning capabilities from large teacher models to substantially smaller student models. The researchers evaluate this approach across arithmetic, commonsense, and symbolic reasoning benchmarks to determine whether smaller models can successfully learn to reason without requiring real-time prompting tricks.
The approach consists of generating step-by-step reasoning solutions for existing training datasets using large teacher models—specifically PaLM with 540 billion parameters and GPT-3 with 175 billion parameters. Crucially, the authors supply the final target answer in the prompt to help the teacher correct minor errors in its chain of thought, and they filter out any remaining incorrect rationales. Smaller student models from the T5 family, ranging from 220 million to 11 billion parameters, are then fine-tuned on these generated reasoning sequences using teacher forcing.
The findings show substantial performance improvements across multiple domains. On the GSM8K math reasoning benchmark, the 11-billion-parameter T5 model increased its accuracy from 8.11% to 21.99% when trained on PaLM-generated reasoning, and reached 38.21% when paired with an external calculator. Similarly, on the MAWPS math dataset, accuracy rose from 54.15% to 70.41% (and 88.22% with a calculator). In ablation studies, a compact T5 base model with 44 times fewer parameters matched the baseline performance of the much larger T5 model when trained on reasoning data. The process also proved highly data-efficient: training on just 20% of the reasoning dataset yielded an 11.22% accuracy, outperforming the full baseline dataset.
These results demonstrate that organizations can significantly reduce model size, operational costs, and latency while retaining advanced multi-step reasoning performance. Smaller models fine-tuned on reasoning data can handle complex logic tasks without paying the computational overhead of running massive frontier models at inference time. However, improvements in commonsense reasoning were more modest—increasing from 68.12% to 71.98% on StrategyQA—because smaller architectures have limited internal memory to store broad factual knowledge.
Organizations seeking to deploy cost-effective reasoning systems should adopt this fine-tuning strategy: annotate datasets using large models conditioned on known target answers, filter out invalid reasoning paths, and train smaller models directly on the step-by-step solutions. To maximize mathematical accuracy, small models should be paired with external tool execution, such as automated calculators.
Confidence in these findings is strong across arithmetic tasks and standard text-based reasoning pipelines, with consistent results across multiple teacher architectures. However, decision-makers should note key limitations: the evaluations were conducted exclusively on English-language benchmarks, evaluated tasks independently rather than in multi-task configurations, and showed limited out-of-distribution generalizability on certain symbolic sequence tasks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This establishes chain-of-thought prompting and the large-model reasoning gains that the source transfers into smaller students through distillation.
- Paper: Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes, Cheng-Yu Hsieh et al. (2023). Its rationale-distillation framework directly supplies the teacher-generated step-by-step supervision and compact T5-student setup developed further by the source.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). STaR provides the preceding rationale-generation and answer-conditioned filtering strategy needed to understand how reasoning data can be created without exhaustive human annotation.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Its zero-shot chain-of-thought results clarify why large teachers can produce useful intermediate reasoning traces for student training.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Program-of-Thoughts motivates the source’s calculator augmentation by showing how external execution can correct language models’ numerical computation errors.
- Paper: Improved Knowledge Distillation via Teacher Assistant, Seyed Iman Mirzadeh et al. (2020). Teacher-assistant distillation introduces the capacity-gap problem between large teachers and small students that informs the source’s compression-focused pipeline.
- Paper: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Rishabh Agarwal et al. (2024). This extends rationale distillation to on-policy student-generated trajectories, addressing the train–inference mismatch left by the source’s teacher-forced training.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Math-Shepherd continues the source’s automated-supervision agenda by using step-level verification to train and rank reasoning solutions without human annotations.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). This advances efficient reasoning beyond offline distillation by allocating additional test-time computation, search, and revision according to problem difficulty.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). s1 extends compact reasoning-model training with carefully curated traces and budget forcing, turning distilled reasoning into controllable test-time scaling.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). DeepSeek-R1 continues the transfer of reasoning capabilities into deployable models through reinforcement learning, rejection sampling, and subsequent model distillation.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). This applies the source’s compact-reasoning principle to edge deployment by combining parameter-efficient tuning, shorter traces, verification, and quantization.
- Paper: Qwen3 Technical Report, An Yang et al. (2025). Qwen3 generalizes small-model reasoning transfer into a unified family that combines chain-of-thought initialization, reinforcement learning, and on-policy distillation across scales.
