ReFT: Reasoning with Reinforced Fine-Tuning

Luong Quoc TrungXinbo ZhangZhanming JiePeng SunXiaoran JinHang Li

article2024ACL338 citations

Proposes Reinforced Fine-Tuning, a framework that couples supervised warm-up with online reinforcement learning to explore multiple automated reasoning paths using ground-truth rewards, substantially outperforming standard supervised fine-tuning on mathematical reasoning benchmarks without requiring extra training questions.

Listen

Large language models frequently struggle with complex multi-step reasoning, such as solving mathematical problems. The standard training practice, supervised fine-tuning, relies on single human-annotated step-by-step reasoning paths per question. This constraint limits model generalization because real-world mathematical problems typically allow multiple valid problem-solving paths that standard training fails to explore.

The article evaluates a novel two-stage training approach called Reinforced Fine-Tuning (ReFT) designed to boost reasoning and generalization in language models. ReFT combines an initial supervised warm-up with online reinforcement learning to allow models to automatically sample and learn from diverse reasoning paths using existing training data.

The researchers conducted extensive experiments across three standard mathematical benchmark datasets (GSM8K, SVAMP, and MathQA) using foundational language models ranging from small versions up to 7 billion parameters (specifically Galactica and CodeLLAMA). The framework was evaluated across both natural language reasoning and Python program-based reasoning formats, comparing ReFT against conventional supervised fine-tuning and self-training baselines.

Key findings show that ReFT significantly outperforms standard supervised fine-tuning across models and benchmarks. On average, across datasets using a 7-billion-parameter CodeLLAMA model, ReFT delivered a 6.7 percentage-point gain in natural language reasoning and a 7.4 percentage-point gain in program-based reasoning over standard supervised fine-tuning. On the GSM8K benchmark, performance improved by roughly 10 points for natural language and nearly 12 points for program-based reasoning. When combined with test-time reranking techniques, the 7-billion-parameter model achieved an 81.2% accuracy on GSM8K, outperforming larger existing baselines and commercial systems like GPT-3.5-turbo. Even small language models under 350 million parameters achieved consistent accuracy gains under ReFT.

These results demonstrate that reinforcement learning can significantly improve a model's internal reasoning capability without requiring expensive data collection, human annotations, or external reward models. Because the reinforcement learning reward is derived directly from checking whether the final answer matches ground truth, ReFT provides rich supervision efficiently. However, ReFT is susceptible to reward hacking on multiple-choice formats, where models can reach a correct multiple-choice letter despite flawed intermediate steps; the method performed substantially better when evaluating direct numerical answers.

Organizations developing reasoning models should consider implementing ReFT over basic supervised fine-tuning to maximize model capabilities from existing datasets. When deploying ReFT, practitioners should prioritize open-ended or numerical generation formats over multiple-choice setups to mitigate reward hacking, and combine trained policies with inference-time reranking for optimal accuracy. Further work is recommended to develop process-based intermediate reward signals and offline reinforcement learning methods to improve training speed and stability.

  • Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting for multi-step reasoning in language models, establishing the foundational reasoning paradigm that ReFT seeks to optimize via reinforcement learning.
  • Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot step-by-step reasoning elicitation in LLMs, providing the core reasoning elicitation principles foundational to subsequent supervised fine-tuning and RL pipelines.
  • Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Establishes sampling multiple diverse reasoning paths and majority voting over final answers, a technique directly integrated and benchmarked as an inference-time strategy in ReFT.
  • Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Introduces bootstrapping reasoning via self-generated rationales filtered by final-answer correctness, motivating the transition to online policy optimization used in ReFT.
  • Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). Formalizes outcome- versus process-based verifiers on mathematical reasoning datasets like MATH and GSM8K, underpinning ReFT's use of ground-truth outcome rewards.
  • Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Examines transferring chain-of-thought reasoning to smaller language models via supervised fine-tuning, highlighting the supervised generalization limits addressed by ReFT.
  • Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). Provides theoretical and empirical analysis of why multi-step intermediate generation aids autoregressive model inference, motivating why exploration across reasoning paths aids generalization.
Cover for ReFT: Reasoning with Reinforced Fine-Tuning

Abstract

One way to enhance the reasoning capability of Large Language Models (LLMs) is to conduct Supervised Fine-Tuning (SFT) using Chain-of-Thought (CoT) annotations. This approach does not show sufficiently strong generalization ability, because the training only relies on the given CoT data. In math problem-solving, for example, there is usually only one annotated reasoning path for each question in the training data. Intuitively, it would be better for the algorithm to learn from multiple annotated reasoning paths given a question. To address this issue, we propose a simple yet effective approach called Reinforced Fine-Tuning (ReFT) to enhance the generalizability of learning LLMs for reasoning, with math problem-solving as an example. ReFT first warmups the model with SFT, and then employs on-line reinforcement learning, specifically the PPO algorithm in this paper, to further fine-tune the model, where an abundance of reasoning paths are automatically sampled given the question and the rewards are naturally derived from the ground-truth answers. Extensive experiments on GSM8K, MathQA, and SVAMP datasets show that ReFT significantly outperforms SFT, and the performance can be potentially further boosted by combining inference-time strategies such as majority voting and re-ranking. Note that ReFT obtains the improvement by learning from the same training questions as SFT, without relying on extra or augmented training questions. This indicates a superior generalization ability for ReFT 1 .

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Reinforced Fine-Tuning
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Baseline
  • 4.3 Experimental Setup
  • 4.4 Results
  • 5 Analysis
  • 6 Conclusion
  • 7 Future Work
  • Limitations
  • References
  • A Examples of N-CoT and P-CoT Representations
  • B Detailed Hyperparameter Setting
  • C Case Study
  • D Attempts with DPO and IPO

Knowls

  1. Knowl 1 — Reinforced Fine-Tuning Algorithm for Mathematical Reasoning

    algorithm

    Reinforced Fine-Tuning (ReFT) is a two-stage fine-tuning framework designed to enhance the reasoning and generalization capabilities of Large Language Models (LLMs) on mathematical problem solving without requiring extra training questions or external reward model annotations.

    In the first stage (Warm-up), the policy model πθ\pi_\theta undergoes Supervised Fine-Tuning (SFT) on question-solution pairs (x,e)(x, e) for a small number of epochs WW (typically 1 to 2 epochs) to acquire basic problem-solving format capabilities. In the second stage (Reinforcement Learning), the model explores multiple reasoning paths by sampling Chain-of-Thought (CoT) trajectories on-policy for questions xx from the training set, extracts the final answer y^\hat{y}, computes golden rewards by verifying y^\hat{y} directly against ground-truth answers yy, and updates both the policy model πθ\pi_\theta and value head VϕV_\phi via Proximal Policy Optimization (PPO).

    Algorithm: Reinforced Fine-Tuning (ReFT)
    Input: Training dataset Dtrain={(x,e,y)}\mathcal{D}_{\text{train}} = \{(x, e, y)\} containing tuples of question xx, ground-truth CoT reasoning path ee, and ground-truth answer yy; number of warm-up steps WW; number of RL steps TT; number of updates per RL step UU; initial policy model πθ(0)\pi_\theta^{(0)}.
    Output: Final policy model πθ\pi_\theta.
    πθ←πθ(0)\pi_\theta \leftarrow \pi_\theta^{(0)}
    // Stage 1: Warm-up via Supervised Fine-Tuning
    for i←1i \leftarrow 1 to WW do
        Sample mini-batch (x,e,y)∼Dtrain(x, e, y) \sim \mathcal{D}_{\text{train}}
        θ←OPTIMIZATION_STEP(LSFT(θ))\theta \leftarrow \text{OPTIMIZATION\_STEP}(\mathcal{L}_{\text{SFT}}(\theta))
    end for
    // Stage 2: On-line Reinforcement Learning via PPO
    Initialize linear value head VϕV_\phi on top of the last hidden state of πθ\pi_\theta
    for i←1i \leftarrow 1 to TT do
        Sample mini-batch of questions without CoT annotations (x,_,y)∼Dtrain(x, \_, y) \sim \mathcal{D}_{\text{train}}
        e^∼πθ(x)\hat{e} \sim \pi_\theta(x) // On-policy CoT sampling
        y^←EXTRACT(e^)\hat{y} \leftarrow \text{EXTRACT}(\hat{e}) // Extract predicted answer from generated reasoning path
        πθold←πθ,Vϕold←Vϕ\pi_{\theta_{\text{old}}} \leftarrow \pi_\theta, V_{\phi_{\text{old}}} \leftarrow V_\phi
        Compute TD errors δt\delta_t, Generalized Advantage Estimates A^t\hat{A}_t, and target returns R^t\hat{R}_t using πθold,Vϕold,x,e^,y^,y\pi_{\theta_{\text{old}}}, V_{\phi_{\text{old}}}, x, \hat{e}, \hat{y}, y
        for j←1j \leftarrow 1 to UU do
            θ,ϕ←OPTIMIZATION_STEP(LRL(θ,ϕ))\theta, \phi \leftarrow \text{OPTIMIZATION\_STEP}(\mathcal{L}_{\text{RL}}(\theta, \phi))
        end for
    end for
    return πθ\pi_\theta
  2. Knowl 2 — Ground-Truth Derived Reward Function and Advantage Estimation in ReFT

    model/method

    In Reinforced Fine-Tuning (ReFT), the policy generates a Chain-of-Thought (CoT) reasoning sequence e=[a1,a2,…,aL]e = [a_1, a_2, \dots, a_L] where aL=<eos>a_L = \text{<eos>} terminates generation. At step tt, state sts_t comprises the input question xx and all tokens generated up to t−1t-1, with action at∼πθ(⋅∣st)a_t \sim \pi_\theta(\cdot|s_t). For non-terminal generation steps (1≤t<L1 \le t < L), the intermediate environment reward is 00.

    At the terminal state sL+1s_{L+1}, the reward function compares the extracted predicted answer EXTRACT(sL+1)\text{EXTRACT}(s_{L+1}) with the ground-truth answer yy. To mitigate learning difficulties associated with sparse rewards, a partial reward of 0.10.1 is provided for problems with numeric answers whenever an answer can be extracted and is numeric, even if incorrect:

    r(st,at,st+1)={1,EXTRACT(st+1)=y0.1,EXTRACT(st+1)≠null and EXTRACT(st+1)≠y0,EXTRACT(st+1)=nullr(s_t, a_t, s_{t+1}) = \begin{cases} 1, & \text{EXTRACT}(s_{t+1}) = y \\ 0.1, & \text{EXTRACT}(s_{t+1}) \neq \text{null} \text{ and } \text{EXTRACT}(s_{t+1}) \neq y \\ 0, & \text{EXTRACT}(s_{t+1}) = \text{null} \end{cases}

    To prevent the policy from diverging excessively from the warm-up initialization policy πθ(0)\pi_\theta^{(0)}, a per-token Kullback-Leibler (KL) divergence penalty scaled by coefficient β\beta is added to form the total reward:

    rtotal(st,at,st+1)=r(st,at,st+1)−βKL(πθ(⋅∣st),πθ(0)(⋅∣st))r_{\text{total}}(s_t, a_t, s_{t+1}) = r(s_t, a_t, s_{t+1}) - \beta \text{KL}\left(\pi_\theta(\cdot|s_t), \pi_\theta^{(0)}(\cdot|s_t)\right)

    The generalized advantage estimate (GAE) A^t\hat{A}_t is computed across sequence length LL:

    A^t=∑l=0L−t(γλ)lδt+l\hat{A}_t = \sum_{l=0}^{L-t} (\gamma \lambda)^l \delta_{t+l}

    where δt′=−Vϕ(st′)+rtotal(st′,at′,st′+1)+γVϕ(st′+1)\delta_{t'} = -V_\phi(s_{t'}) + r_{\text{total}}(s_{t'}, a_{t'}, s_{t'+1}) + \gamma V_\phi(s_{t'+1}), Vϕ(sL+1):=0V_\phi(s_{L+1}) := 0, γ∈[0,1]\gamma \in [0, 1] is the discount factor for temporal difference errors, and λ∈(0,1]\lambda \in (0, 1] is the GAE discount factor. The target return is calculated via R^t=A^t+Vϕ(st)\hat{R}_t = \hat{A}_t + V_\phi(s_t).

  3. Knowl 3 — Policy and Value Optimization Objectives in Reinforced Fine-Tuning

    model/method

    ReFT optimizes a unified loss during the reinforcement learning phase that balances the PPO clipped surrogate policy loss and a clipped value regression loss:

    LRL(θ,ϕ)=Lpolicy(θ)+αLvalue(ϕ)\mathcal{L}_{\text{RL}}(\theta, \phi) = \mathcal{L}_{\text{policy}}(\theta) + \alpha \mathcal{L}_{\text{value}}(\phi)

    where α\alpha is the weighting coefficient for the value objective.

    The clipped policy surrogate objective is defined over generated reasoning paths e∼πθolde \sim \pi_{\theta_{\text{old}}}:

    Lpolicy(θ)=−Ee∼πθold[∑t=1Lmin⁡(πθ(at∣st)πθold(at∣st)A^t,clip(πθ(at∣st)πθold(at∣st),1−ϵ,1+ϵ)A^t)]\mathcal{L}_{\text{policy}}(\theta) = -\mathbb{E}_{e \sim \pi_{\theta_{\text{old}}}} \left[ \sum_{t=1}^L \min\left( \frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)} \hat{A}_t, \text{clip}\left(\frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\text{old}}}(a_t|s_t)}, 1 - \epsilon, 1 + \epsilon\right) \hat{A}_t \right) \right]

    where ϵ\epsilon is the PPO clipping parameter and A^t\hat{A}_t is the generalized advantage estimate.

    The value loss objective fits the parameterized value head VϕV_\phi (constructed by placing a linear layer on the final hidden state of the policy model backbone) using a clipped squared error against the λ\lambda-return estimate R^t\hat{R}_t:

    Lvalue(ϕ)=12Ee∼πθold[∑t=1Lmax⁡((Vϕ(st)−R^t)2,(clip(R^t−Vϕ(st),A^t−ϵ,A^t+ϵ))2)]\mathcal{L}_{\text{value}}(\phi) = \frac{1}{2} \mathbb{E}_{e \sim \pi_{\theta_{\text{old}}}} \left[ \sum_{t=1}^L \max\left( \left(V_\phi(s_t) - \hat{R}_t\right)^2, \left(\text{clip}\left(\hat{R}_t - V_\phi(s_t), \hat{A}_t - \epsilon, \hat{A}_t + \epsilon\right)\right)^2 \right) \right]

    During the initial warm-up stage, standard Supervised Fine-Tuning (SFT) minimizes the cross-entropy loss over ground-truth question-CoT sequences (x,e)∼Dtrain(x, e) \sim \mathcal{D}_{\text{train}}:

    LSFT(θ)=−E(x,e)∼Dtrain[∑t=1Llog⁡(πθ(at∣st))]\mathcal{L}_{\text{SFT}}(\theta) = -\mathbb{E}_{(x, e) \sim \mathcal{D}_{\text{train}}} \left[ \sum_{t=1}^L \log\left(\pi_\theta(a_t|s_t)\right) \right]

  4. Knowl 4 — Mathematical Reasoning Accuracy of ReFT Compared with SFT and Self-Training Baselines

    data/table

    ReFT was evaluated across three benchmarks: GSM8K, SVAMP, and MathQAMCQ, using two 7B-scale foundation models (Galactica-6.7B and CodeLLAMA-7B) across both Natural Language Chain-of-Thought (N-CoT) and Program-based Chain-of-Thought (P-CoT) formats. Baselines included Supervised Fine-Tuning (SFT), Offline Self-Training (Offline-ST, which samples 100 CoTs using an early SFT checkpoint, filters by correctness, and fine-tunes on the combined dataset), and Online Self-Training (Online-ST, which continually samples and updates on correct batches with LSFT\mathcal{L}_{\text{SFT}}).

    Method Size GSM8K SVAMP MathQAMCQ Average
    N-CoT P-CoT N-CoT P-CoT N-CoT P-CoT N-CoT P-CoT
    Galactica + SFT 6.7B 42.68 58.83 54.50 70.09 58.07 64.61 51.75 64.51
    Galactica + Offline-ST 6.7B 42.60 60.72 57.90 72.30 60.75 67.04 53.75 66.69
    Galactica + Online-ST 6.7B 47.84 62.93 59.40 74.59 59.38 61.24 55.54 66.25
    Galactica + ReFT 6.7B 48.14 68.91 61.40 74.09 58.13 70.47 55.89 71.16
    CodeLLAMA + SFT 7B 43.59 63.68 58.09 75.40 56.01 64.79 52.56 67.96
    CodeLLAMA + Offline-ST 7B 45.10 68.00 60.20 77.69 59.81 68.53 55.04 71.41
    CodeLLAMA + Online-ST 7B 44.66 67.85 58.60 77.40 56.95 68.85 53.40 71.37
    CodeLLAMA + ReFT 7B 53.30 75.28 64.50 79.19 60.13 71.83 59.31 75.43

    CodeLLAMA-7B trained with ReFT achieved absolute gains over SFT of 9.71 points on GSM8K N-CoT (53.30% vs. 43.59%) and 11.60 points on GSM8K P-CoT (75.28% vs. 63.68%). Across all datasets, ReFT improved average accuracy over SFT by 6.75 points on N-CoT and 7.47 points on P-CoT with CodeLLAMA. Online-ST and Offline-ST underperformed ReFT because they only perform positive reinforcement on filtered samples and lack negative reinforcement feedback or explicit KL divergence constraints.

  5. Knowl 5 — Inference-Time Majority Voting and Outcome Reward Model Reranking with ReFT

    data/table

    ReFT-trained policies can be combined with inference-time majority voting and reward model reranking. For evaluation on GSM8K, 100 CoT solutions per question were sampled at temperature 1.0. For reranking, an Outcome-based Reward Model (ORM) with a linear classifier head on top of an SFT checkpoint was trained on binary correctness labels of sampled training paths.

    Method Size GSM8K
    N-CoT P-CoT
    Galactica + SFT + Voting 6.7B 52.8 62.9
    Galactica + ReFT + Voting 6.7B 58.5 71.8
    Galactica + SFT + Reranking 6.7B 57.5 73.4
    Galactica + ReFT + Reranking 6.7B 59.2 76.4
    CodeLLAMA + SFT + Voting 7B 53.5 68.0
    CodeLLAMA + ReFT + Voting 7B 63.2 78.0
    CodeLLAMA + SFT + Reranking 7B 62.9 77.0
    CodeLLAMA + ReFT + Reranking 7B 66.0 81.2
    Models Trained with Extra Distilled Data
    WizardMath 7B 54.9 –
    WizardMath 13B 63.9 –
    MathCoder 7B 67.8 –
    MAmmoTH-Coder 7B 22.2 58.8
    MAmmoTH-Coder 70B 72.4 76.7
    GPT-3.5-turbo N.A. 75.3 78.0
    GPT-4 N.A. 93.0 97.0

    CodeLLAMA-7B + ReFT + Reranking achieved 81.2% accuracy on GSM8K P-CoT, outperforming GPT-3.5-turbo (78.0%) as well as open-source 70B models (MAmmoTH-Coder 70B at 76.7%), despite using only the original GSM8K training questions without ChatGPT data augmentation or distillation.

  6. Knowl 6 — Reward Hacking in Multiple-Choice Math Reasoning and MathQAnumeric Evaluation

    empirical result

    When applying Reinforcement Learning with terminal outcome rewards to multiple-choice datasets such as MathQAMCQ, models can suffer from reward hacking. Because the action space for the final answer is restricted to choices {A,B,C,D,E}\{A, B, C, D, E\}, an on-policy trajectory can execute flawed or incorrect intermediate reasoning (for example, computing an erroneous intermediate quantity like 172 cm2172\text{ cm}^2 instead of 198 cm2198\text{ cm}^2) but still guess the correct option letter (e.g., 'C') at the final token. This assigns a full positive reward of 11 to invalid reasoning steps, degrading policy optimization.

    To eliminate the multiple-choice artifact, experiments were conducted on MathQAnumeric, a variant where multiple-choice options are stripped and the model must directly predict numeric values using N-CoT:

    Method (N-CoT) Galactica (6.7B) CodeLLAMA (7B)
    SFT 40.08 37.32
    Offline Self-Training 44.23 41.24
    Online Self-Training 43.78 38.06
    ReFT 45.23 42.24

    On MathQAnumeric, ReFT outperformed SFT by 5.15 points on Galactica (45.23% vs. 40.08%) and 4.92 points on CodeLLAMA (42.24% vs. 37.32%), demonstrating that removing multiple-choice exploitation restores the empirical advantages of ReFT.

  7. Knowl 7 — Scalability of ReFT to Small Language Models

    data/table

    To assess whether online reinforcement learning can function effectively on small base models without policy collapse during exploration, ReFT was evaluated on Program-based Chain-of-Thought (P-CoT) reasoning using small models: Galactica-125M, Codeparrot-small (110M), and Codegen-350M across GSM8K, SVAMP, and MathQAMCQ.

    Method GSM8K SVAMP MathQAMCQ
    Galactica-125M + SFT 23.7 35.6 58.4
    Galactica-125M + ReFT 29.8 39.4 60.7
    Codeparrot-small + SFT 13.8 25.7 55.3
    Codeparrot-small + ReFT 16.8 27.4 58.3
    Codegen-350M + SFT 20.4 34.4 56.4
    Codegen-350M + ReFT 28.4 39.3 59.1

    ReFT consistently outperformed SFT across all three small models and all three datasets. For instance, ReFT improved GSM8K accuracy from 20.4% to 28.4% on Codegen-350M (+8.0 points) and from 23.7% to 29.8% on Galactica-125M (+6.1 points), demonstrating the stability of reinforced fine-tuning across smaller model scales.

  8. Knowl 8 — Ablation Study on ReFT Architectural and Algorithmic Components

    empirical result

    Ablation experiments conducted on CodeLLAMA-7B on GSM8K with P-CoT evaluated the necessity of individual components of ReFT:

    Model Setting Accuracy (%)
    CodeLLAMA + ReFT 75.28
    – remove partial reward 74.40
    – KL coefficient β=0\beta = 0 collapse (0.0)
    – non-shared value model 75.15
    1. Partial Reward: Removing the 0.10.1 reward for valid extractable numeric answers reduced accuracy from 75.28% to 74.40%, showing that partial rewards mitigate optimization challenges from sparse reward signals.
    2. KL Divergence Regularization: Setting the KL penalty coefficient β=0\beta = 0 caused policy collapse (0.0% accuracy), indicating that constraining policy divergence from the warm-up initialization πθ(0)\pi_\theta^{(0)} is necessary for stable RL exploration.
    3. Shared vs. Separate Value Model: Initializing a completely separate value model backbone yielded 75.15% accuracy (comparable to 75.28% for the shared backbone with a linear head), but doubled memory consumption and added an extra forward pass.
  9. Knowl 9 — Qualitative Human Evaluation of Reasoning Paths Generated by ReFT

    empirical result

    A blind human evaluation was conducted on 50 GSM8K test questions that were solved correctly by all three models (SFT, Warmup checkpoint, and ReFT) on P-CoT using CodeLLAMA-7B. Four independent annotators scored the generated Python code across three criteria on a scale of 0 to 1:

    • Logic: whether the logical sequence of operations leading to the final result is sound.
    • Naming: whether variable names convey clear, appropriate domain semantics.
    • Compactness: whether the reasoning trajectory is concise without redundant steps.
    Method Logic Naming Compactness Overall Score (out of 3)
    SFT 0.986 0.988 0.994 2.967
    Warmup 0.949 0.982 0.990 2.920
    ReFT 0.992 0.990 0.996 2.982

    ReFT achieved higher scores across Logic (0.992 vs. 0.986), Naming (0.990 vs. 0.988), and Compactness (0.996 vs. 0.994) compared to SFT, showing that exploring multiple reasoning paths during RL improves solution conciseness and logical structure without degrading variable naming clarity.

  10. Knowl 10 — Limitations of Reinforced Fine-Tuning: Sample Efficiency and Outcome-Based Reward Gaming

    limitation

    Reinforced Fine-Tuning exhibits two main limitations:

    1. Training Efficiency and Convergence Speed: Because ReFT explores a discrete token generation space to optimize a non-differentiable correctness reward via PPO, it requires substantially more training epochs to reach convergence (e.g., up to 300 epochs) compared to Supervised Fine-Tuning (which converges in around 40 epochs). Increasing the learning rate to accelerate training risks policy instability and collapse, whereas increasing batch size substantially increases GPU memory and computational costs.
    2. Reward Hacking in Small Discrete Answer Spaces: The reward formulation depends entirely on the extracted terminal answer. In multiple-choice question formats with restricted output spaces (e.g., choices A, B, C, D), the policy can receive positive rewards for hallucinated or faulty intermediate reasoning steps if the final multiple-choice token happens to match the ground truth. Addressing this requires process-based reward models (PRMs) or converting multiple-choice problems into direct open-ended generation.

Coverage note — None was omitted; all primary methodological components, algorithmic procedures, experimental comparisons, ablation studies, qualitative analyses, and stated limitations are fully represented.

Citation

MLA
Trung, L., et al. “ReFT: Reasoning with Reinforced Fine-Tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7601–14, https://doi.org/10.18653/v1/2024.acl-long.410.
APA
Trung, L., Zhang, X., Jie, Z., Sun, P., Jin, X., & Li, H. (2024). ReFT: Reasoning with Reinforced Fine-Tuning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7601–7614. https://doi.org/10.18653/v1/2024.acl-long.410
Chicago
Trung, L., X. Zhang, Z. Jie, P. Sun, X. Jin, and H. Li. 2024. “ReFT: Reasoning with Reinforced Fine-Tuning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7601–14. https://doi.org/10.18653/v1/2024.acl-long.410.
Harvard
Trung, L. et al. (2024) “ReFT: Reasoning with Reinforced Fine-Tuning”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7601–7614. Available at: https://doi.org/10.18653/v1/2024.acl-long.410.
Vancouver
1. Trung L, Zhang X, Jie Z, Sun P, Jin X, Li H (2024) ReFT: Reasoning with Reinforced Fine-Tuning. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7601–7614

BibTeX

@inproceedings{trung-etal-2024-reft,
    title = "{R}e{FT}: Reasoning with Reinforced Fine-Tuning",
    author = "Trung, Luong  and
      Zhang, Xinbo  and
      Jie, Zhanming  and
      Sun, Peng  and
      Jin, Xiaoran  and
      Li, Hang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.410/",
    doi = "10.18653/v1/2024.acl-long.410",
    pages = "7601--7614"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/