Phi-4-reasoning Technical Report

Marah AbdinSahaj AgarwalSahaj AgarwalAhmed AwadallahVidhisha BalachandranHarkirat BehlLingjiao Chen (lingjiaochen)Gustavo de RosaSuriya GunasekarMojan Javaheripi

article2025arXiv130 citations

Demonstrates how a 14-billion parameter language model trained with targeted distillation and outcome-based reinforcement learning can match or exceed the reasoning performance of significantly larger systems like DeepSeek-R1-Distill-Llama-70B across math, coding, and scientific problem-solving.

Listen

Recent advances in artificial intelligence emphasize reasoning models that spend additional computational effort during inference to solve complex, multi-step problems. While leading frontier systems achieve strong reasoning capabilities, they typically rely on massive model sizes or proprietary architectures that carry high computational costs. The article evaluates whether a compact, 14-billion parameter model can achieve state-of-the-art reasoning and problem-solving performance through targeted data curation, supervised fine-tuning, and lightweight reinforcement learning.

The developers created two models: Phi-4-reasoning, trained by fine-tuning the base Phi-4 model on over 1.4 million curated prompts containing high-quality reasoning traces generated by OpenAI's o3-mini, and Phi-4-reasoning-plus, which adds a short phase of outcome-based reinforcement learning using approximately 6,000 verifiable math problems. The training emphasized teachable prompts situated at the boundary of the base model's capabilities across science, mathematics, coding, and safety domains, while expanding the context window to 32,000 tokens. To evaluate performance and address benchmark volatility, the article assessed the models across extensive reasoning benchmarks—such as math competitions, graduate-level science, coding, and planning—using multi-run evaluations and open assessment frameworks.

The findings show that both 14-billion parameter models dramatically outperform their base model and compete effectively with significantly larger systems. On the American Invitational Mathematics Examination (AIME 2025), Phi-4-reasoning and Phi-4-reasoning-plus achieved 63.1% and 78.0% accuracy respectively, outperforming the 70-billion parameter DeepSeek-R1-Distill (51.5%) and matching or exceeding larger frontier models. Improvements extended to un-targeted domains, showing gains of 30 to 60 percentage points on algorithmic problem-solving and calendar planning tasks. On general capabilities, Phi-4-reasoning-plus delivered a 22-point gain in instruction following and improved robustness to long context windows. The analysis also revealed substantial output non-determinism across all industry reasoning models, where single-run scores on small benchmarks varied by up to 40 percentage points.

These results demonstrate that meticulous data selection and post-training allow smaller, cost-effective models to rival massive systems, lowering deployment costs and latency for high-difficulty enterprise workflows. Moreover, reasoning capabilities transfer positively to general-purpose language tasks without causing catastrophic forgetting. However, the reinforcement learning stage in Phi-4-reasoning-plus required 1.5 times more generation tokens on average to achieve its math gains, illustrating a direct trade-off between inference cost and top-tier accuracy.

Organizations evaluating reasoning models should implement robust, multi-run testing protocols rather than relying on single-score public benchmarks, as run-to-run variance can produce misleading comparisons. Decision-makers should deploy the base reasoning model for balanced speed and cost, reserving the reinforced plus variant for math-heavy workflows requiring maximal accuracy. Future development should focus on expanding reinforcement learning verification beyond mathematics into coding and general planning domains, while extending context length beyond 32,000 tokens to mitigate truncation on complex tasks.

Cover for Phi-4-reasoning Technical Report

Abstract

We introduce Phi-4-reasoning, a 14-billion parameter reasoning model that achieves strong performance on complex reasoning tasks. Trained via supervised fine-tuning of Phi-4 on carefully curated set of “teachable” prompts–selected for the right level of complexity and diversity–and reasoning demonstrations generated using o3-mini, Phi-4-reasoning generates detailed reasoning chains that effectively leverage inference-time compute. We further develop Phi-4-reasoning-plus, a variant enhanced through a short phase of outcome-based reinforcement learning that offers higher performance by generating longer reasoning traces. Across a wide range of reasoning tasks, both models outperform significantly larger open-weight models such as DeepSeek-R1-Distill-Llama-70B model and approach the performance levels of full DeepSeek-R1 model. Our comprehensive evaluations span benchmarks in math and scientific reasoning, coding, algorithmic problem solving, planning, and spatial understanding. Interestingly, we observe a non-trivial transfer of improvements to general-purpose benchmarks as well. In this report, we provide insights into our training data, our training methodologies, and our evaluations. We show that the benefit of careful data curation for supervised fine-tuning (SFT) extends to reasoning language models, and can be further amplified by reinforcement learning (RL). Finally, our evaluation points to opportunities for improving how we assess the performance and robustness of reasoning models.

Table of Contents

  • 1 Introduction
  • 2 Data Methodology
  • 2.1 Seeds database
  • 2.2 Training data
  • 3 Phi-4-reasoning : Supervised Finetuning of Phi-4
  • 3.1 Exploration Stage
  • 3.2 Scaling Stage
  • 4 Phi-4-reasoning-plus : A bit of RL on top of Phi-4-reasoning
  • 4.1 Reward Function
  • 4.2 Training Details and Experimental Observations
  • 5 Evaluation
  • 5.1 Reasoning Benchmarks
  • 5.1.1 Baseline models
  • 5.1.2 Accuracy distribution on AIME 2025: Beyond single-score analyses
  • 5.1.3 Main findings
  • 5.1.4 Average vs. best-of-5 performance
  • 5.2 General-purpose Benchmarks
  • 5.3 Safety Evaluation
  • 6 Limitations
  • Author Contributions
  • Acknowledgements
  • References
  • A Benchmarking Details
  • B Additional Results

Knowls

  1. Knowl 1 — Phi-4-reasoning Architecture and Context Window Extension

    model/method

    Phi-4-reasoning is a 14-billion parameter reasoning language model built upon the base Phi-4 architecture. To support structured step-by-step reasoning and extended inference-time compute, two specific modifications are introduced:

    1. Specialized Reasoning Tokens: Two reserved placeholder tokens from the base vocabulary are assigned as <think> and </think> to explicitly delimit internal reasoning blocks from final answer blocks.
    2. Context Length Expansion: While the base Phi-4 model supports a maximum sequence length of 16,384 tokens (16K16\text{K}), Phi-4-reasoning doubles the base frequency of the Rotary Position Embedding (RoPE) mechanism to accommodate up to 32,768 tokens (32K32\text{K}) during training and generation.
  2. Knowl 2 — Teachable Seed Curation and Data Filtering Pipeline

    model/method

    The prompt dataset for training Phi-4-reasoning and Phi-4-reasoning-plus is constructed through a multi-stage filtering process targeting "teachable" prompts that sit at the capability boundary of the base model:

    • Seed Collection: Diverse prompts are collected from web sources, licensed datasets, and synthetic generation grounded in web text across STEM disciplines, coding, general question answering, and Responsible AI alignment.
    • Teachable Boundary Filtering: Because base Phi-4 already solves simpler tasks, prompts are filtered using heuristic difficulty scoring. When ground truth is unavailable, plurality consensus from a strong reference LLM serves as a proxy target; the agreement rate of weaker models (such as Phi-4 or GPT-4o) against this target measures sample difficulty. Prompts with significant error margins are preserved.
    • Multi-Step Complexity Evaluation: Rubric-based LLM evaluators score each prompt by the number and depth of reasoning steps required, prioritizing complex logical decomposition over factual recall.
    • Synthetic Rewriting: Filtered seeds are programmatically transformed, such as converting programming challenges into math-style word problems or restructuring math problems into verifiable forms with concise solutions suitable for reinforcement learning verification.
    • Benchmark Decontamination: Seed data is scrubbed using exact-match and semantic n-gram filtering against major benchmarks (including AIME 2024, MATH, GPQA, LiveCodeBench, Codeforces, OmniMATH, and SWE-Bench Verified).
  3. Knowl 3 — Additive Property in Supervised Fine-Tuning Data Mixture Optimization

    model/method

    In the supervised fine-tuning (SFT) of Phi-4-reasoning, training data mixtures exhibit an empirical additive property across distinct capability domains. Training data sources are clustered by domain (e.g., mathematics, software engineering) and data quality, with each cluster assigned a sampling repetition weight (number of epochs).

    Rather than performing a combinatorial search over the joint multi-domain space, optimal mixture weights are determined independently for each individual domain by increasing cluster iterations until downstream validation metrics saturate. Concatenating these domain-specific optimal weights into a unified training mixture preserves the independent accuracy gains achieved in isolated tuning without negative domain transfer or catastrophic forgetting on general-purpose benchmarks.

  4. Knowl 4 — Supervised Fine-Tuning Setup and Hyperparameters for Phi-4-reasoning

    experimental setup

    Phi-4-reasoning is trained using supervised fine-tuning directly on the post-trained Phi-4 base checkpoint (retaining safety and Responsible AI behaviors). The training configuration comprises:

    • Dataset Size: Over 1.4 million curated prompt-response pairs containing structured reasoning chains generated by OpenAI's o3-mini (in high reasoning effort mode), totaling approximately 8.3 billion tokens across mathematics, coding, and safety alignment.
    • Optimization: AdamW optimizer trained over approximately 16,000 optimization steps on 16 billion tokens, with a global batch size of 32 sequences at a full context length of 32,768 tokens.
    • Hyperparameters: Peak learning rate η=10−5\eta = 10^{-5} (selected via grid search over [10−6,2×10−5][10^{-6}, 2 \times 10^{-5}]), linear learning rate warmup over 450 steps, and weight decay of 10−410^{-4}.
    • System Message: A standardized system prompt instructs the model to structure all responses strictly into a <think> [detailed multi-step reasoning] </think> section followed by a concise [final solution] section.
  5. Knowl 5 — Length-Aware Accuracy Reward Formulation for Reasoning Reinforcement Learning

    equation

    Phi-4-reasoning-plus uses a rule-based length-aware accuracy reward Racc_scaledR_{\text{acc\_scaled}} during reinforcement learning to encourage concise reasoning on correct answers while encouraging extended exploration when incorrect.

    Let Racc_raw∈{0,1}R_{\text{acc\_raw}} \in \{0, 1\} denote raw binary answer correctness verified via exact matching, boxed string extraction, or an LLM equivalence verifier. Let LL be the response token length, Lmax⁡=31,744L_{\max} = 31{,}744 be the maximum allowed response length, Lpos_control=25,600L_{\text{pos\_control}} = 25{,}600 be the threshold length before penalizing correct responses, and Lneg_control=3,702L_{\text{neg\_control}} = 3{,}702 be the threshold length before penalizing short incorrect responses.

    For correct responses (Racc_raw=1R_{\text{acc\_raw}} = 1), define the progress metric: ρ+=min⁡(1,max⁡(L−Lpos_control,0)Lmax⁡−Lpos_control)\rho_+ = \min\left(1, \frac{\max(L - L_{\text{pos\_control}}, 0)}{L_{\max} - L_{\text{pos\_control}}}\right) The scaled reward ranges from Rmin⁡+=0.5R_{\min}^+ = 0.5 to Rmax⁡+=1.0R_{\max}^+ = 1.0 using cosine scaling: Racc_scaled=Rmin⁡++0.5⋅(Rmax⁡+−Rmin⁡+)⋅(1+cos⁡(πρ+))R_{\text{acc\_scaled}} = R_{\min}^+ + 0.5 \cdot (R_{\max}^+ - R_{\min}^+) \cdot (1 + \cos(\pi \rho_+))

    For incorrect responses (Racc_raw=0R_{\text{acc\_raw}} = 0), define the progress metric: ρ−=min⁡(1,LLneg_control)\rho_- = \min\left(1, \frac{L}{L_{\text{neg\_control}}}\right) The scaled reward ranges from Rmin⁡−=−1.0R_{\min}^- = -1.0 to Rmax⁡−=−0.5R_{\max}^- = -0.5: Racc_scaled=Rmax⁡−+0.5⋅(Rmin⁡−−Rmax⁡−)⋅(1+cos⁡(πρ−))R_{\text{acc\_scaled}} = R_{\max}^- + 0.5 \cdot (R_{\min}^- - R_{\max}^-) \cdot (1 + \cos(\pi \rho_-))

    Formatting violations override this calculation:

    • Missing end-of-sequence token (<|im_end|>): Racc_scaled=−0.5R_{\text{acc\_scaled}} = -0.5
    • Missing or improperly formatted <think> tags: Racc_scaled=−1.0R_{\text{acc\_scaled}} = -1.0
  6. Knowl 6 — Repetition Penalty and Total RL Reward Objective

    equation

    To prevent degenerative repeating loops in generated reasoning traces during reinforcement learning, a 5-gram Repetition Penalty RrepR_{\text{rep}} is defined as: Rrep=−max⁡(#{5-grams with frequency>5}#{5-grams},max⁡ frequency of 5-grams with frequency>5#{words}/5)R_{\text{rep}} = -\max\left( \frac{\#\{5\text{-grams with frequency} > 5\}}{\#\{5\text{-grams}\}}, \frac{\max\text{ frequency of 5-grams with frequency} > 5}{\#\{\text{words}\} / 5} \right)

    The total outcome reward RfinalR_{\text{final}} combining scaled accuracy Racc_scaledR_{\text{acc\_scaled}} and repetition penalty RrepR_{\text{rep}} is given by: Rfinal=waccRacc_scaled+wrepRrepR_{\text{final}} = w_{\text{acc}} R_{\text{acc\_scaled}} + w_{\text{rep}} R_{\text{rep}} where the weights are set to wacc=813w_{\text{acc}} = \frac{8}{13} and wrep=113w_{\text{rep}} = \frac{1}{13}, establishing a maximum possible reward upper bound of 813≈0.615\frac{8}{13} \approx 0.615.

  7. Knowl 7 — Group Relative Policy Optimization (GRPO) Configuration for Phi-4-reasoning-plus

    model/method

    Phi-4-reasoning-plus is trained using Group Relative Policy Optimization (GRPO) on top of the SFT checkpoint. For each prompt query qq sampled from a dataset of 72,401 verifiable mathematical problems, the model generates a group of G=8G = 8 rollout candidate outputs {o1,o2,…,oG}\{o_1, o_2, \dots, o_G\}.

    The per-token group relative advantage A^i,t\hat{A}_{i,t} for token tt in output oio_i is estimated by normalizing total rewards over the group: A^i,t=Rfinal(q,oi)−mean⁡({Rfinal(q,oj)}j=1G)std⁡({Rfinal(q,oj)}j=1G)\hat{A}_{i,t} = \frac{R_{\text{final}}(q, o_i) - \operatorname{mean}(\{R_{\text{final}}(q, o_j)\}_{j=1}^G)}{\operatorname{std}(\{R_{\text{final}}(q, o_j)\}_{j=1}^G)}

    The GRPO objective maximized over policy parameters θ\theta is: 1G∑i=1G1∣oi∣∑t=1∣oi∣{min⁡[ri,t(θ)A^i,t,clip⁡(ri,t(θ),1−ϵ,1+ϵ)A^i,t]−βDKL(πθ∥πθold)+γEntropy⁡(πθ)}\frac{1}{G} \sum_{i=1}^G \frac{1}{|o_i|} \sum_{t=1}^{|o_i|} \left\{ \min \left[ r_{i,t}(\theta) \hat{A}_{i,t}, \operatorname{clip}(r_{i,t}(\theta), 1-\epsilon, 1+\epsilon) \hat{A}_{i,t} \right] - \beta D_{\text{KL}}(\pi_\theta \parallel \pi_{\theta_{\text{old}}}) + \gamma \operatorname{Entropy}(\pi_\theta) \right\} where ri,t(θ)=πθ(oi,t∣q,oi,<t)πθold(oi,t∣q,oi,<t)r_{i,t}(\theta) = \frac{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}{\pi_{\theta_{\text{old}}}(o_{i,t} \mid q, o_{i,<t})}, KL divergence penalty weight β=0.001\beta = 0.001, and entropy coefficient γ=0.001\gamma = 0.001.

    Training is executed using the verl framework on 32 Nvidia H100 GPUs with a global batch size of 64 problem prompts per step, using the Adam optimizer at learning rate 5×10−85 \times 10^{-8} with a 10-step cosine warmup. The final model is selected at step 90 (processing only ∼6,400\sim 6{,}400 problem seeds across 88 rollouts each).

  8. Knowl 8 — Contrasting Token Length Dynamics Between Reasoning SFT and RL

    empirical result

    Supervised fine-tuning and reinforcement learning exhibit opposing effects on reasoning trace length:

    1. SFT Token Efficiency: Across SFT iterations, the average reasoning response length on benchmarks like AIME 2024 and GPQA Diamond slightly decreases while accuracy increases, indicating that SFT teaches more concise and efficient utilization of reasoning tokens.
    2. RL Token Expansion: Under GRPO training with length-aware accuracy rewards, reasoning trace lengths grow steadily, with average response length strongly correlated with mathematical problem-solving accuracy on AIME.
    3. Length Divergence by Answer Correctness: In RL rollouts, incorrect attempts expand in length significantly faster (aligning with the 75th percentile of response length) than correct attempts (which remain near the bottom 25th percentile). Beyond step 90, clipping generations at the 31K31\text{K} token ceiling leads to reward plateauing because incorrect chains exhaust their budget before producing boxed final answers.
  9. Knowl 9 — Benchmark Accuracy of Phi-4-reasoning Across Core Reasoning Benchmarks

    data/table

    The performance of Phi-4-reasoning (SFT) and Phi-4-reasoning-plus (SFT + RL) is evaluated on competitive mathematics, scientific reasoning, and programming benchmarks against state-of-the-art open and closed models. AIME 2025 and HMMT use 50 and 64 repetitions respectively, while other benchmarks use 5 repetitions.

    Model AIME 24 AIME 25 HMMT OmniMath GPQA-D LCB Codeforces
    Feb 2025 8/24–1/25 Elo
    Phi-4 (14B) 28.0 12.9 3.8 31.9 54.7 26.9 –
    Phi-4-reasoning (14B) 74.6 (5.1) 63.1 (6.3) 43.8 (6.2) 76.6 (0.5) 67.1 (2.7) 53.8 1736
    Phi-4-reasoning-plus (14B) 81.3 (1.8) 78.0 (4.6) 53.6 (6.3) 81.9 (0.1) 69.3 (2.1) 53.1 1723
    OpenThinker2-32B 58.0 58.0 – – 64.1 – –
    QwQ-32B 79.5 65.8 47.5 – 59.5 63.4 –
    EXAONE-Deep-32B 72.1 65.8 – – 66.1 59.5 –
    DeepSeek-R1-Distill-70B 69.3 (2.7) 51.5 (5.8) 33.3 63.4 (0.4) 66.2 (2.4) 57.5 1633
    DeepSeek-R1 (671B MoE) 78.7 (3.8) 70.4 (4.3) 41.7 85.0 (0.6) 73.0 (1.7) 65.9 2029
    o1-mini 63.6 54.8 38.0 (6.2) 60.5 60.0 53.8 1650
    o1 74.6 (6.5) 71.4 (5.7) 48.3 67.5 (0.9) 76.7 (1.8) 63.4 1891
    o3-mini (high) 88.0 (5.5) 82.5 (4.9) 67.5 74.6 (5.1) 77.7 (0.6) 68.8 2130
    Claude-3.7-Sonnet (thinking) 55.3 (3.0) 53.0 (5.8) 31.7 54.6 (0.9) 76.8 (1.3) 52.6 –
    Gemini-2.5-Pro 92.0 86.7 82.5 – 84.0 69.1 –

    Despite having only 14B parameters, Phi-4-reasoning and Phi-4-reasoning-plus outperform the 70B distilled model DeepSeek-R1-Distill-Llama-70B on math benchmarks and surpass full DeepSeek-R1 (671B) on AIME 2025 (78.0% vs. 70.4%) and HMMT Feb 2025 (53.6% vs. 41.7%). Both models also generalize to out-of-domain combinatorial search tasks, achieving 78.0% on 3SAT and 42.6% on TSP.

  10. Knowl 10 — Transfer of Reasoning Gains to General-Purpose and Long-Context Benchmarks

    data/table

    Post-training for reasoning does not induce catastrophic forgetting on general-purpose NLP capabilities, but instead produces substantial transfer improvements in instruction following, long-context retrieval, and coding benchmarks.

    Benchmark Phi-4 Phi-4-reasoning Phi-4-reasoning-plus o3-mini GPT-4o
    FlenQA (3K-token subset) 82.0 97.7 97.9 96.8 90.8
    IFEval Strict 62.3 83.4 84.9 91.5 81.8
    ArenaHard 68.1 73.3 79.0 81.9 69.0
    HumanEvalPlus 83.5 92.9 92.3 94.0 84.9
    MMLU-Pro 71.5 74.3 76.0 79.4 73.5
    Kitab (No Context - Precision) 19.3 23.2 27.6 37.9 53.7
    Kitab (With Context - Precision) 88.5 93.8 93.6 94.0 84.7
    Kitab (No Context - Recall) 8.2 4.9 6.3 4.2 20.3
    Kitab (With Context - Recall) 68.1 74.8 75.4 76.1 69.2
    Toxigen Discriminative (Toxic) 72.6 86.7 77.3 85.4 87.6
    Toxigen Discriminative (Neutral) 90.0 84.7 90.5 88.7 85.1
    PhiBench 2.21 58.2 70.6 74.2 78.0 73.1

    Phi-4-reasoning-plus improves over base Phi-4 by 22.6 percentage points on IFEval Strict and 15.9 points on FlenQA. On FlenQA, reasoning models demonstrate robustness to input context length and are invariant to whether needle facts are presented contiguously or scattered at random context locations.

  11. Knowl 11 — Evaluation Variance and Parallel Test-Time Scaling on AIME 2025

    empirical result

    Analysis of reasoning model evaluations across multiple repeated runs yields two key findings:

    1. Evaluation Instability on Small Datasets: Kernel density estimation across 50 independent runs on AIME 2025 (30 problems) at non-zero temperature reveals wide performance variance. DeepSeek-R1-Distill-Llama-70B ranges from 30% to 70% accuracy across runs, and o3-mini ranges from 70% to 100%. Consequently, reporting single-run or 5-run average scores on small benchmarks creates unreliable comparisons.
    2. Parallel Test-Time Scaling Gains: Scaling parallel generation budget NN from 11 to 6464 (using Majority Voting or Best-of-NN) produces substantial accuracy gains. On AIME 2025, Phi-4-reasoning-plus Majority@N scales from 78.0% to over 88%, while Best-of-64 surpasses 90%, exceeding the pass@1 baseline of the o3-mini teacher model.
  12. Knowl 12 — Limitations of Phi-4-reasoning and Phi-4-reasoning-plus

    limitation

    The Phi-4-reasoning model family exhibits several specific constraints:

    • Context Window Upper Bound: The maximum context window of 32,768 tokens can be exceeded by long multi-turn chats or deep search trajectories, resulting in output truncation and lost answer tags.
    • Domain Specialization Scope: SFT data is restricted to STEM, software code, and safety, while RL training data is strictly limited to mathematics. As a consequence, performance gains on non-STEM domains rely on meta-skill transfer.
    • Discipline Disparities: Evaluation across scientific sub-disciplines shows markedly lower improvements in biology, chemistry, and discrete mathematics compared to continuous mathematics and physics.
    • Parametric Fact Recall: In ungrounded information retrieval (e.g., Kitab without context), the models achieve poor recall (4.9%--6.3%), reflecting limitations in purely parametric factual knowledge.

Coverage note — None was omitted; all key methodologies, architectures, data recipes, mathematical formulations, empirical results, and limitations were fully extracted into self-contained knowls.

References

  1. 1.Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu, Cyril Zhang, and Yi Zhang. Phi-4 technical report, 2024. URL https://arxiv.org/abs/2412.08905.
  3. 3.Marah I Abdin, Suriya Gunasekar, Varun Chandrasekaran, Jerry Li, Mert Yüksekgönül, Rahee Ghosh Peshawaria, Ranjita Naik, and Besmira Nushi. KITAB: evaluating llms on constraint satisfaction for information retrieval. In International Conference on Learning Representations, 2024.
  4. 4.AIME. Aime 83-24. https://huggingface.co/datasets/lchen001/AIME1983_2024, 2024. Accessed: 2025-03-17.
  5. 5.AIME. Aime 2025. https://huggingface.co/datasets/lchen001/AIME2025, 2025. Accessed: 2025-03-17.
  6. 6.Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
  7. 7.Anthropic. Claude 3.7 sonnet. https://www.anthropic.com/news/claude-3-7-sonnet, 2025. Accessed: 2025-03-17.
  8. 8.Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy. Chain-of-thought reasoning in the wild is not always faithful. arXiv preprint arXiv:2503.08679, 2025.
  9. 9.Vidhisha Balachandran, Jingya Chen, Neel Joshi, Besmira Nushi, Hamid Palangi, Eduardo Salinas, Vibhav Vineet, James Woffinden-Luey, and Safoora Yousefi. Eureka: Evaluating and understanding large foundation models. arXiv preprint arXiv:2409.10566, 2024.
  10. 10.Vidhisha Balachandran, Jingya Chen, Lingjiao Chen, Shivam Garg, Neel Joshi, Yash Lara, John Langford, Besmira Nushi, Vibhav Vineet, Yue Wu, and Safoora Yousefi. Inference-time scaling for complex tasks: Where we stand and what lies ahead, 2025. URL https://arxiv.org/abs/2504.00294.
  11. 11.Mislav Balunović, Jasper Dekoninck, Ivo Petrov, Nikola Jovanović, and Martin Vechev. Matharena: Evaluating llms on uncontaminated math competitions, February 2025. URL https://matharena.ai/.
  12. 12.Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W Duncan Wadsworth, and Hanna Wallach. Designing disaggregated evaluations of ai systems: Choices, considerations, and tradeoffs. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 368–378, 2021.
  13. 13.Natasha Butt, Varun Chandrasekaran, Neel Joshi, Besmira Nushi, and Vidhisha Balachandran. Benchagents: Automated benchmark creation with agent interaction. arXiv preprint arXiv:2410.22584, 2024.
  14. 14.Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
  15. 15.Quy-Anh Dang and Chris Ngo. Reinforcement learning for reasoning in small llms: What works and what doesn't, 2025. URL https://arxiv.org/abs/2503.16219.
  16. 16.Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, et al. Omni-math: A universal olympiad level mathematic benchmark for large language models. ICLR, 2025.
  17. 17.Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pages 10835–10866. PMLR, 2023.
  18. 18.Google. Gemini flash thinking. https://deepmind.google/technologies/gemini/flash/, 2025. Accessed: 2025-03-17.
  19. 19.Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang, Youran Sun, Yi Zhu, Fan Yang, and Mao Yang. rstar-math: Small llms can master math reasoning with self-evolved deep thinking. arXiv preprint arXiv:2501.04519, 2025.
  20. 20.Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023.
  21. 21.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  22. 22.Juris Hartmanis. Computers and intractability: a guide to the theory of np-completeness (michael r. garey and david s. johnson). Siam Review, 24(1):90, 1982.
  23. 23.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326. Association for Computational Linguistics, 2022.
  24. 24.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding, 2021. URL https://arxiv.org/abs/2009.03300.
  25. 25.Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge. A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility. arXiv preprint arXiv:2504.07086, 2025.
  26. 26.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  27. 27.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  28. 28.Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, Suriya Gunasekar, Piero Kauffmann, Yin Tat Lee, Yuanzhi Li, Anh Nguyen, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Michael Santacroce, Harkirat Singh Behl, Adam Taumann Kalai, Xin Wang, Rachel Ward, Philipp Witte, Cyril Zhang, and Yi Zhang. Phi-2: The surprising power of small language models. Microsoft Research Blog, 2023.
  29. 29.Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. Safechain: Safety of language models with long chain-of-thought reasoning capabilities. arXiv preprint arXiv:2502.12025, 2025.
  30. 30.Mosh Levy, Alon Jacoby, and Yoav Goldberg. Same task, more tokens: the impact of input length on the reasoning performance of large language models. In ACL, 2024.
  31. 31.LG AI Research. Exaone deep: Reasoning enhanced language models. arXiv preprint arXiv:2503.12524, 2025.
  32. 32.Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. Functional interpolation for relative positions improves long context transformers. arXiv preprint arXiv:2310.04418, 2023.
  33. 33.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024.
  34. 34.Xuefeng Li, Haoyang Zou, and Pengfei Liu. Limr: Less is more for rl scaling, 2025.
  35. 35.Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023. URL https://arxiv.org/abs/2305.01210.
  36. 36.Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2, 2025. Notion Blog.
  37. 37.Ahmed Magooda, Alec Helyar, Kyle Jackson, David Sullivan, Chad Atalla, Emily Sheng, Dan Vann, Richard Edgar, Hamid Palangi, Roman Lutz, Hongliang Kong, Vincent Yun, Eslam Kamal, Federico Zarfati, Hanna Wallach, Sarah Bird, and Mei Chen. A framework for automated measurement of responsible ai harms in generative ai applications, 2023. URL https://arxiv.org/abs/2310.17750.
  38. 38.Arindam Mitra, Luciano Del Corro, Shweti Mahajan, Andres Codas, Clarisse Simoes, Sahaj Agarwal, Xuxi Chen, Anastasia Razdaibiedina, Erik Jones, Kriti Aggarwal, et al. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045, 2023.
  39. 39.Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andres Codas, Yadong Lu, Wei-ge Chen, Olga Vrousgos, Corby Rosset, et al. Agentinstruct: Toward generative teaching with agentic flows. arXiv preprint arXiv:2407.03502, 2024.
  40. 40.Mazda Moayeri, Vidhisha Balachandran, Varun Chandrasekaran, Safoora Yousefi, Thomas Fel, Soheil Feizi, Besmira Nushi, Neel Joshi, and Vibhav Vineet. Unearthing skill-level insights for understanding trade-offs of foundation models. ICLR, 2025.
  41. 41.Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707, 2023.
  42. 42.Besmira Nushi, Ece Kamar, and Eric Horvitz. Towards accountable ai: Hybrid human-machine analyses for characterizing system failure. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 6, pages 126–135, 2018.
  43. 43.OpenAI. Openai o3-mini system card. https://openai.com/index/o3-mini-system-card/, 2025. Accessed: 2025-03-17.
  44. 44.Christos H Papadimitriou. Computational complexity. In Encyclopedia of computer science, pages 260–265. John Wiley and Sons Ltd., 2003.
  45. 45.Samir Passi and Mihaela Vorvoreanu. Overreliance on ai literature review. Microsoft Research, 339:340, 2022.
  46. 46.Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev, Maria Drencheva, Kristian Minchev, Mislav Balunović, Nikola Jovanović, and Martin Vechev. Proof or bluff? evaluating llms on 2025 usa math olympiad. arXiv preprint arXiv:2503.21934, 2025.
  47. 47.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024.
  48. 48.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  49. 49.Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. Hybridflow: A flexible and efficient rlhf framework. arXiv preprint arXiv: 2409.19256, 2024.
  50. 50.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners, 2022. URL https://arxiv.org/abs/2210.03057.
  51. 51.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021.
  52. 52.Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599, 2025.
  53. 53.OpenThoughts Team. Open Thoughts. https://open-thoughts.ai, January 2025.
  54. 54.Qwen Team. Qwq-32b: Embracing the power of reinforcement learning, March 2025. URL https://qwenlm.github.io/blog/qwq-32b/.
  55. 55.Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023. URL https://arxiv.org/abs/2305.04388.
  56. 56.Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Sharon Li, and Neel Joshi. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. Advances in Neural Information Processing Systems, 37:75392–75421, 2024.
  57. 57.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024. URL https://arxiv.org/abs/2406.01574.
  58. 58.An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115, 2024.
  59. 59.Guanghao Ye, Khiem Duc Pham, Xinzhi Zhang, Sivakanth Gopi, Baolin Peng, Beibin Li, Janardhan Kulkarni, and Huseyin A Inan. On the emergence of thinking in llms i: Searching for the right intuition. arXiv preprint arXiv:2502.06773, 2025.
  60. 60.Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. Demystifying long chain-of-thought reasoning in llms. arXiv preprint arXiv:2502.03373, 2025.
  61. 61.Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Weinan Dai, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei-Ying Ma, Ya-Qin Zhang, Lin Yan, Mu Qiao, Yonghui Wu, and Mingxuan Wang. Dapo: An open-source llm reinforcement learning system at scale, 2025. URL https://arxiv.org/abs/2503.14476.
  62. 62.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023.

Citation

MLA
Abdin, M., et al. “Phi-4-reasoning Technical Report”. arXiv, 2025, http://arxiv.org/abs/2504.21318v1.
APA
Abdin, M., Agarwal, S., Awadallah, A., Balachandran, V., Behl, H., Chen, L., Rosa, G. de ., Gunasekar, S., Javaheripi, M., Joshi, N., Kauffmann, P., Lara, Y., Mendes, C. C. T., Mitra, A., Nushi, B., Papailiopoulos, D., Saarikivi, O., Shah, S., Shrivastava, V., … Zheng, G. (2025). Phi-4-reasoning Technical Report. arXiv. http://arxiv.org/abs/2504.21318v1
Chicago
Abdin, M., S. Agarwal, A. Awadallah, et al. 2025. “Phi-4-reasoning Technical Report”. arXiv. http://arxiv.org/abs/2504.21318v1.
Harvard
Abdin, M. et al. (2025) “Phi-4-reasoning Technical Report”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2504.21318v1.
Vancouver
1. Abdin M, Agarwal S, Awadallah A, et al (2025) Phi-4-reasoning Technical Report. arXiv

BibTeX

@article{abdin2025phi,
  title = {Phi-4-reasoning Technical Report},
  author = {Abdin, Marah and Agarwal, Sahaj and Awadallah, Ahmed and Balachandran, Vidhisha and Behl, Harkirat and Chen, Lingjiao and Rosa, Gustavo de and Gunasekar, Suriya and Javaheripi, Mojan and Joshi, Neel and Kauffmann, Piero and Lara, Yash and Mendes, Caio César Teodoro and Mitra, Arindam and Nushi, Besmira and Papailiopoulos, Dimitris and Saarikivi, Olli and Shah, Shital and Shrivastava, Vaishnavi and Vineet, Vibhav and Wu, Yue and Yousefi, Safoora and Zheng, Guoqing},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2504.21318v1},
  eprint = {2504.21318}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission