Phi-4-reasoning Technical Report
Marah AbdinSahaj AgarwalSahaj AgarwalAhmed AwadallahVidhisha BalachandranHarkirat BehlLingjiao Chen (lingjiaochen)Gustavo de RosaSuriya GunasekarMojan Javaheripi
Demonstrates how a 14-billion parameter language model trained with targeted distillation and outcome-based reinforcement learning can match or exceed the reasoning performance of significantly larger systems like DeepSeek-R1-Distill-Llama-70B across math, coding, and scientific problem-solving.
Recent advances in artificial intelligence emphasize reasoning models that spend additional computational effort during inference to solve complex, multi-step problems. While leading frontier systems achieve strong reasoning capabilities, they typically rely on massive model sizes or proprietary architectures that carry high computational costs. The article evaluates whether a compact, 14-billion parameter model can achieve state-of-the-art reasoning and problem-solving performance through targeted data curation, supervised fine-tuning, and lightweight reinforcement learning.
The developers created two models: Phi-4-reasoning, trained by fine-tuning the base Phi-4 model on over 1.4 million curated prompts containing high-quality reasoning traces generated by OpenAI's o3-mini, and Phi-4-reasoning-plus, which adds a short phase of outcome-based reinforcement learning using approximately 6,000 verifiable math problems. The training emphasized teachable prompts situated at the boundary of the base model's capabilities across science, mathematics, coding, and safety domains, while expanding the context window to 32,000 tokens. To evaluate performance and address benchmark volatility, the article assessed the models across extensive reasoning benchmarks—such as math competitions, graduate-level science, coding, and planning—using multi-run evaluations and open assessment frameworks.
The findings show that both 14-billion parameter models dramatically outperform their base model and compete effectively with significantly larger systems. On the American Invitational Mathematics Examination (AIME 2025), Phi-4-reasoning and Phi-4-reasoning-plus achieved 63.1% and 78.0% accuracy respectively, outperforming the 70-billion parameter DeepSeek-R1-Distill (51.5%) and matching or exceeding larger frontier models. Improvements extended to un-targeted domains, showing gains of 30 to 60 percentage points on algorithmic problem-solving and calendar planning tasks. On general capabilities, Phi-4-reasoning-plus delivered a 22-point gain in instruction following and improved robustness to long context windows. The analysis also revealed substantial output non-determinism across all industry reasoning models, where single-run scores on small benchmarks varied by up to 40 percentage points.
These results demonstrate that meticulous data selection and post-training allow smaller, cost-effective models to rival massive systems, lowering deployment costs and latency for high-difficulty enterprise workflows. Moreover, reasoning capabilities transfer positively to general-purpose language tasks without causing catastrophic forgetting. However, the reinforcement learning stage in Phi-4-reasoning-plus required 1.5 times more generation tokens on average to achieve its math gains, illustrating a direct trade-off between inference cost and top-tier accuracy.
Organizations evaluating reasoning models should implement robust, multi-run testing protocols rather than relying on single-score public benchmarks, as run-to-run variance can produce misleading comparisons. Decision-makers should deploy the base reasoning model for balanced speed and cost, reserving the reinforced plus variant for math-heavy workflows requiring maximal accuracy. Future development should focus on expanding reinforcement learning verification beyond mathematics into coding and general planning domains, while extending context length beyond 32,000 tokens to mitigate truncation on complex tasks.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). Introduces the pure reinforcement learning and distillation methodology for reasoning models that Phi-4-reasoning explicitly compares against and seeks to match at smaller scale.
- Paper: ReFT: Reasoning with Reinforced Fine-Tuning, Luong Quoc Trung et al. (2024). Establishes the reinforced fine-tuning paradigm combining supervised warm-up with reinforcement learning to explore diverse reasoning paths.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Demonstrates the foundational technique of distilling multi-step reasoning capabilities from large teacher models into compact student architectures.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Pioneers the bootstrapping of language model reasoning by filtering and learning from self-generated rationale traces.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). Develops the Group Relative Policy Optimization algorithm and curated data pipelines essential for training mathematical reasoning models via reinforcement learning.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the chain-of-thought prompting framework that serves as the conceptual bedrock for multi-step reasoning models.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). Establishes step-by-step verification and process supervision methodologies critical for evaluating and guiding intermediate reasoning chains.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Provides techniques for automated step-by-step reinforcement without manual annotation, foundational to scalable reasoning training.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Presents majority voting and self-consistency decoding mechanisms used extensively during inference and data curation in reasoning models.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). Establishes a rigorous multi-task reasoning benchmark that standardizes the modern evaluation of frontier reasoning systems.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). Critically examines the underlying drivers of emergent reasoning behaviors in reinforcement-learned models like Phi-4-reasoning.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). Extends test-time compute scaling by introducing budget forcing mechanisms on distilled reasoning traces.
- Paper: DAPO: An Open-Source LLM Reinforcement Learning System at Scale, Qiying Yu et al. (2025). Addresses system-level scaling and stability bottlenecks in reinforcement learning pipelines for competitive mathematical reasoning.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). Proposes evaluation frameworks that measure intermediate reasoning validity and redundancy beyond mere final-answer accuracy.
- Paper: Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?, Zhiyuan Zeng 0004 et al. (2025). Investigates whether sequential test-time scaling and extended reasoning chains reliably yield accuracy improvements.
- Paper: MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations, Kaixuan Huang et al. (2025). Benchmarks reasoning models against hard structural perturbations to evaluate genuine analytical capability versus memorization.
- Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). Synthesizes post-training reasoning data construction methods and analyzes how data curation dictates model performance.
- Paper: Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation, Yanqi Dai et al. (2026). Enhances outcome-based reinforcement learning through difficulty-aware optimization and question reformulation.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). Adapts compact reasoning models for low-power edge deployment using length penalties and budget control.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). Develops multi-tree test-time compute scaling frameworks to enhance robustness across complex reasoning trajectories.
