ReFT: Reasoning with Reinforced Fine-Tuning
Luong Quoc TrungXinbo ZhangZhanming JiePeng SunXiaoran JinHang Li
Proposes Reinforced Fine-Tuning, a framework that couples supervised warm-up with online reinforcement learning to explore multiple automated reasoning paths using ground-truth rewards, substantially outperforming standard supervised fine-tuning on mathematical reasoning benchmarks without requiring extra training questions.
Large language models frequently struggle with complex multi-step reasoning, such as solving mathematical problems. The standard training practice, supervised fine-tuning, relies on single human-annotated step-by-step reasoning paths per question. This constraint limits model generalization because real-world mathematical problems typically allow multiple valid problem-solving paths that standard training fails to explore.
The article evaluates a novel two-stage training approach called Reinforced Fine-Tuning (ReFT) designed to boost reasoning and generalization in language models. ReFT combines an initial supervised warm-up with online reinforcement learning to allow models to automatically sample and learn from diverse reasoning paths using existing training data.
The researchers conducted extensive experiments across three standard mathematical benchmark datasets (GSM8K, SVAMP, and MathQA) using foundational language models ranging from small versions up to 7 billion parameters (specifically Galactica and CodeLLAMA). The framework was evaluated across both natural language reasoning and Python program-based reasoning formats, comparing ReFT against conventional supervised fine-tuning and self-training baselines.
Key findings show that ReFT significantly outperforms standard supervised fine-tuning across models and benchmarks. On average, across datasets using a 7-billion-parameter CodeLLAMA model, ReFT delivered a 6.7 percentage-point gain in natural language reasoning and a 7.4 percentage-point gain in program-based reasoning over standard supervised fine-tuning. On the GSM8K benchmark, performance improved by roughly 10 points for natural language and nearly 12 points for program-based reasoning. When combined with test-time reranking techniques, the 7-billion-parameter model achieved an 81.2% accuracy on GSM8K, outperforming larger existing baselines and commercial systems like GPT-3.5-turbo. Even small language models under 350 million parameters achieved consistent accuracy gains under ReFT.
These results demonstrate that reinforcement learning can significantly improve a model's internal reasoning capability without requiring expensive data collection, human annotations, or external reward models. Because the reinforcement learning reward is derived directly from checking whether the final answer matches ground truth, ReFT provides rich supervision efficiently. However, ReFT is susceptible to reward hacking on multiple-choice formats, where models can reach a correct multiple-choice letter despite flawed intermediate steps; the method performed substantially better when evaluating direct numerical answers.
Organizations developing reasoning models should consider implementing ReFT over basic supervised fine-tuning to maximize model capabilities from existing datasets. When deploying ReFT, practitioners should prioritize open-ended or numerical generation formats over multiple-choice setups to mitigate reward hacking, and combine trained policies with inference-time reranking for optimal accuracy. Further work is recommended to develop process-based intermediate reward signals and offline reinforcement learning methods to improve training speed and stability.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting for multi-step reasoning in language models, establishing the foundational reasoning paradigm that ReFT seeks to optimize via reinforcement learning.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot step-by-step reasoning elicitation in LLMs, providing the core reasoning elicitation principles foundational to subsequent supervised fine-tuning and RL pipelines.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Establishes sampling multiple diverse reasoning paths and majority voting over final answers, a technique directly integrated and benchmarked as an inference-time strategy in ReFT.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Introduces bootstrapping reasoning via self-generated rationales filtered by final-answer correctness, motivating the transition to online policy optimization used in ReFT.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). Formalizes outcome- versus process-based verifiers on mathematical reasoning datasets like MATH and GSM8K, underpinning ReFT's use of ground-truth outcome rewards.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Examines transferring chain-of-thought reasoning to smaller language models via supervised fine-tuning, highlighting the supervised generalization limits addressed by ReFT.
- Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). Provides theoretical and empirical analysis of why multi-step intermediate generation aids autoregressive model inference, motivating why exploration across reasoning paths aids generalization.
- Paper: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training, Tianzhe Chu et al. (2025). Directly examines and validates the core premise of ReFT by empirically comparing how supervised fine-tuning tends to memorize while reinforcement learning generalizes.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). Scales outcome-rewarded reinforcement learning post-training to develop emergent reasoning behaviors without relying heavily on human chain-of-thought annotations.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Extends reinforcement fine-tuning for math problem solving from ReFT's outcome-only reward formulation to automated step-level process supervision without human annotation.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). Provides a comprehensive survey synthesizing post-training reinforced reasoning methods, categorizing techniques like ReFT within the broader landscape of Large Reasoning Models.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). Critically examines the mechanisms and latent base-model capabilities driving reinforced reasoning gains over supervised fine-tuning baselines.
- Paper: DAPO: An Open-Source LLM Reinforcement Learning System at Scale, Qiying Yu et al. (2025). Builds on reinforced post-training principles to develop scalable, open-source system-level algorithms addressing training stability and entropy collapse in reasoning tasks.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). Generalizes intermediate verification beyond simple outcome-based RL by modeling multi-step reasoning trajectories as sequential decision processes with Q-value rankings.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). Investigates the generalization dynamics of RL algorithms (including PPO) when scaling mathematical reasoning models from simple to complex tasks.
- Paper: Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, Yang Yue et al. (2025). Investigates whether reinforced post-training creates entirely new reasoning pathways or primarily optimizes sampling of capabilities inherent to the base model.
- Paper: Understanding Reasoning from Pretraining to Post-Training, Jingyan Shen et al. (2026). Analyzes the complete pretraining-to-post-training interface to explain how initial policy distributions interact with subsequent reinforcement learning for reasoning.
