Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Fahim TajwarAnikait SinghArchit SharmaRafael RafailovJeff SchneiderTengyang XieStefano ErmonChelsea FinnAviral Kumar

article2024ICML206 citations
Listen

Aligning large language models with human values typically relies on preference fine-tuning, but practitioners face major dilemmas regarding methodology. Teams often struggle to determine whether to invest in complex on-policy reinforcement learning—where the model generates fresh responses during training—or rely on simpler, offline contrastive or supervised techniques. A related challenge involves deciding whether human feedback data must be generated interactively by the model in development or if static datasets suffice. Understanding these dynamics is critical today because training large models is computationally expensive, and selecting the wrong alignment strategy leads to wasted compute, suboptimal performance, and model failure.

The article evaluates why specific preference fine-tuning approaches succeed while others fail. It investigates the distinct roles and interactions of on-policy sampling and negative gradient terms across diverse data coverage conditions and geometric relationships between initial models and human reward targets.

To conduct this evaluation, the researchers established a unified experimental and theoretical framework spanning multiple tiers of complexity. This included controlled multi-dimensional bandit tasks, synthetic language model experiments targeting specific response lengths, and full-scale fine-tuning using Pythia-1.4B and Mistral-7B architectures on established benchmarks like AlpacaFarm and UltraFeedback. Through a generalized algorithm, the study systematically varied data batch freshness, sample reuse, and loss formulation across thousands of optimization steps, benchmarking standard reinforcement learning, contrastive optimization, and supervised likelihood methods.

The findings establish that methods integrating on-policy sampling and negative gradients consistently outperform standard offline supervised fine-tuning. First, on-policy sampling is especially essential when the desired high-reward responses lie far from the initial model distribution, whereas offline methods suffice when high rewards already align with the initial model mode. Second, negative gradient terms—which actively push down the probability of undesirable completions—accelerate training convergence and create a substantially wider preference margin compared to standard likelihood maximization, which often inadvertently increases the probability of both good and bad answers. Third, combining on-policy generation with contrastive loss objectives achieves superior reward optimization and up to 3-fold to 10-fold wall-clock training acceleration over pure reinforcement learning. Finally, theoretical analysis reveals that on-policy updates and negative gradients exhibit mode-seeking mathematical behavior, allowing rapid probability mass relocation onto high-reward responses in a few gradient steps, whereas standard supervised approaches display mode-covering dynamics that dilute probability mass across all responses.

These insights demonstrate that relying exclusively on offline, maximum-likelihood supervised learning carries significant risk when target behaviors differ substantially from initial model defaults. By adopting on-policy sampling or contrastive negative gradients, engineering teams can cut compute costs, accelerate alignment timelines, and prevent models from plateauing at suboptimal behaviors. The results clarify conflicting industry findings by demonstrating that the necessity of reinforcement learning depends directly on the geometric gap between the initial model and the desired human preference distribution.

Practitioners should actively deploy on-policy contrastive pipelines that sample fresh model completions and score them via reward models before applying negative gradient updates. For tasks where optimal responses align with existing model capabilities, teams can save compute by utilizing straightforward offline methods. When sampling fresh responses during training, teams should implement mild sample reuse to improve efficiency, but exercise caution with non-clipped methods to avoid overfitting on stale data. Future operational decisions should be supported by small-scale pilot assessments measuring the alignment between baseline policy outputs and reward targets.

These conclusions are bounded by key limitations: the analysis assumes an underlying reward model accurately represents preferences and does not deeply analyze the consequences of noisy or exploitative reward models. Furthermore, theoretical claims are established through learning dynamics on categorical distributions rather than end-to-end sample complexity bounds for massive architectures. Nonetheless, the experimental consistency across both synthetic environments and large-scale standard benchmarks provides high confidence in the practical recommendations.

Abstract

Learning from preference labels plays a crucial role in fine-tuning large language models — this is done via supervised learning, on-policy reinforcement learning (RL), or contrastive learning. Different methods come with different implementation tradeoffs, and existing empirical findings present different conclusions, for instance, some results show that online RL is quite important to attain good fine-tuning results, while others find offline methods sufficient. This raises a question: what kind of approaches are important for fine-tuning with preference data and why? In this paper, we answer this question by performing a rigorous analysis of a number of fine-tuning techniques on didactic and full-scale LLM problems. Our main finding is that approaches that use on-policy sampling and attempt to push down the likelihood on certain responses (i.e., employ a “negative gradient”) outperform offline and maximum likelihood objectives. We conceptualize our insights and unify methods that use on-policy sampling or negative gradient under a notion of mode-seeking objectives for categorical distributions. Mode-seeking objectives are able to alter probability mass on specific bins of a categorical distribution at a fast rate compared to maximum likelihood, allowing them to relocate masses across bins more effectively. Our analysis prescribes actionable insights for preference fine-tuning of LLMs and informs how data should be collected for maximal improvement.

Table of Contents

  • 1. Introduction
  • 2. Unifying Preference Fine-Tuning Methods
  • 2.1. Preliminaries and Notation
  • 2.2. Characterizing Fine-Tuning Methods
  • 3. Research Questions and Analysis Setup
  • 3.1. Coverage Conditions and Geometric Relationships
  • 3.2. Tasks and Datasets
  • 3.3. A Generic Fine-Tuning Algorithm
  • 4. Empirical Analysis Results
  • 4.1. Question 1: The Role of On-Policy Sampling
  • 4.2. Question 2: The Role of Negative Gradient
  • 4.3. Question 3: On-Policy Sampling and Negative Gradients are Complementary
  • 5. Theoretical Analysis
  • 6. Discussion and Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • Appendices
  • A. Related Work
  • B. Limitations
  • C. Connections to Existing Fine-Tuning Results
  • D. Computational vs Wall-Clock Time Tradeoff for Various Methods
  • E. More on Conceptual Unification and Theoretical Analysis
  • E.1. Seeking Modes Unifies On-Policy Sampling and Negative Gradients
  • E.1.1. ON-POLICY METHODS ARE MODE-SEEKING
  • E.1.2. CONTRASTIVE APPROACHES (E.G., DPO/IPO) ARE MODE-SEEKING
  • E.1.3. SUPERVISED OFFLINE ALGORITHMS ARE MODE-COVERING
  • E.1.4. GRADIENTS FOR BOTH DPO AND IPO EXHIBIT THE FORM IN LEMMA E.2.
  • E.2. Case Study: Mode-Seeking Reverse KL vs. Mode-Covering Forward KL
  • F. Additional Algorithmic Details
  • F.1. Score/Reward Standardization
  • F.2. IPO
  • G. Method Hyperparameters
  • G.1. DPO (Rafailov et al., 2023)
  • G.2. Pref-FT (Dubois et al., 2024)
  • G.3. PPO (Schulman et al., 2017)
  • G.4. RWR
  • G.5. Iterated Best-of-N (Mukobi et al., 2023)
  • H. Code For Running Experiments
  • I. More on Didactic Bandit Problems
  • I.1. Problem Setup
  • I.2. Algorithmic Details
  • I.2.1. BEST-OF-N
  • I.2.2. IPO
  • I.2.3. REINFORCE
  • I.2.4. PPO
  • I.2.5. RWR
  • I.3. Experiment Details
  • J. Detailed Empirical Results
  • J.1. On-policy Sampling in Synthetic LLM Problems
  • J.2. More on On-Policy Sample Reuse
  • J.3. Effect of Negative Gradient in the Didactic Bandit Problem
  • J.4. Setup for Negative Gradient Experiments in Full-scale LLM Fine-tuning
  • J.5. More Details on Mechanisms Explaining the Behavior of the Negative Gradient
  • J.6. More Empirical Results on Complimentary Nature of On-Policy Sampling and Negative Gradients
  • K. Additional Experiments on Synthetic LLM Setup
  • K.1. Performance of Various Algorithms on the Mode Length Setting
  • K.2. Effect of On-policy Samples vs Samples from an Older Policy in Synthetic Length Settings
  • K.3. Sample Reuse in Synthetic LLM Settings

Knowls

  1. Knowl 1 — Unification of Preference Fine-Tuning Objectives via Mode-Seeking and Mode-Covering Divergences

    theoretical result

    Preference fine-tuning objectives can be conceptually unified and classified according to whether they optimize a mode-seeking (reverse KL-divergence) or mode-covering (forward KL-divergence) objective:

    1. On-Policy Reinforcement Learning (RL) and Weighted Likelihood Methods are Mode-Seeking: On-policy RL algorithms (such as PPO and REINFORCE) and on-policy weighted supervised approaches optimize the regularized objective

    LRL(Dpref,πθ)=−Ex∼Dpref[Ey∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x)∥πref(⋅∣x))]\mathcal{L}_{\text{RL}}(\mathcal{D}_{\text{pref}}, \pi_\theta) = -\mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} \left[ \mathbb{E}_{y \sim \pi_\theta(\cdot|x)} [r(x, y)] - \beta D_{\text{KL}}(\pi_\theta(\cdot|x) \parallel \pi_{\text{ref}}(\cdot|x)) \right]

    where πθ\pi_\theta is the parameterized policy, πref\pi_{\text{ref}} is the reference policy, r(x,y)r(x, y) is the reward, and β>0\beta > 0 controls the regularizer strength. Under the optimal policy parameterization π∗(y∣x)=1Z(x)πref(y∣x)exp⁡(r(x,y)/β)\pi^*(y|x) = \frac{1}{Z(x)} \pi_{\text{ref}}(y|x) \exp(r(x, y) / \beta) with partition function Z(x)=∑yπref(y∣x)exp⁡(r(x,y)/β)Z(x) = \sum_y \pi_{\text{ref}}(y|x) \exp(r(x, y) / \beta), this loss reduces to

    LRL(Dpref,πθ)=−βEx∼Dpref[log⁡Z(x)]+βEx∼Dpref[DKL(πθ(⋅∣x)∥π∗(⋅∣x))]\mathcal{L}_{\text{RL}}(\mathcal{D}_{\text{pref}}, \pi_\theta) = -\beta \mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} [\log Z(x)] + \beta \mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} \left[ D_{\text{KL}}(\pi_\theta(\cdot|x) \parallel \pi^*(\cdot|x)) \right]

    Because minimizing DKL(πθ∥π∗)D_{\text{KL}}(\pi_\theta \parallel \pi^*) corresponds to the reverse KL-divergence with respect to the optimal policy π∗\pi^*, on-policy RL is intrinsically mode-seeking.

    1. Offline Supervised Methods are Mode-Covering: Offline supervised approaches (such as offline Best-of-N, Pref-FT, and offline Reward-Weighted Regression) maximize weighted likelihood:

    Loff-sup(πθ;πref)=−Ex∼Dpref[Ey∼πref(⋅∣x)[log⁡πθ(y∣x)⋅F(x,y)]]\mathcal{L}_{\text{off-sup}}(\pi_\theta; \pi_{\text{ref}}) = -\mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} \left[ \mathbb{E}_{y \sim \pi_{\text{ref}}(\cdot|x)} [\log \pi_\theta(y|x) \cdot F(x, y)] \right]

    for non-negative weights F(x,y)≥0F(x, y) \ge 0 where ∑yF(x,y)>0\sum_y F(x, y) > 0. By defining the re-weighted distribution π~(y∣x)=πref(y∣x)F(x,y)Z(x)\tilde{\pi}(y|x) = \frac{\pi_{\text{ref}}(y|x) F(x, y)}{Z(x)} with normalizer Z(x)=∑zπref(z∣x)F(x,z)Z(x) = \sum_z \pi_{\text{ref}}(z|x) F(x, z), the objective becomes

    Loff-sup(πθ;πref)=Ex∼Dpref[Z(x)DKL(π~(⋅∣x)∥πθ(⋅∣x))]+Ex∼Dpref[Z(x)H(π~(⋅∣x))]\mathcal{L}_{\text{off-sup}}(\pi_\theta; \pi_{\text{ref}}) = \mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} \left[ Z(x) D_{\text{KL}}(\tilde{\pi}(\cdot|x) \parallel \pi_\theta(\cdot|x)) \right] + \mathbb{E}_{x \sim \mathcal{D}_{\text{pref}}} [Z(x) H(\tilde{\pi}(\cdot|x))]

    where H(⋅)H(\cdot) is the Shannon entropy. This minimizes the forward KL-divergence DKL(π~∥πθ)D_{\text{KL}}(\tilde{\pi} \parallel \pi_\theta), which exhibits mode-covering behavior.

    1. Offline Contrastive Methods with Negative Gradients Exhibit Mode-Seeking Dynamics: For contrastive updates of the form

    θt+1←θt+ηEx,yw,yl∼D[∇θlog⁡πθ(yw∣x)⋅c1(x,yw,yl)−∇θlog⁡πθ(yl∣x)⋅c2(x,yw,yl)]\theta_{t+1} \leftarrow \theta_t + \eta \mathbb{E}_{x, y_w, y_l \sim \mathcal{D}} \left[ \nabla_\theta \log \pi_\theta(y_w|x) \cdot c_1(x, y_w, y_l) - \nabla_\theta \log \pi_\theta(y_l|x) \cdot c_2(x, y_w, y_l) \right]

    with non-negative scalar weight functions c1,c2≥0c_1, c_2 \ge 0, setting c2>0c_2 > 0 introduces a negative gradient. For any model parameterization and iteration tt, there exists an appropriate pairing of positive and negative responses (yw,yl)(y_w, y_l) such that the expected log-likelihood ratio ωt+1=log⁡πθ(yw∣x)−log⁡πθ(yl∣x)\omega_{t+1} = \log \pi_\theta(y_w|x) - \log \pi_\theta(y_l|x) is strictly greater when c2>0c_2 > 0 than under weighted maximum likelihood (c2=0c_2 = 0). When gradient directions between ywy_w and yly_l satisfy ED[∇θlog⁡πθ(yw∣x)⊤∇θlog⁡πθ(yl∣x)]≤0\mathbb{E}_{\mathcal{D}}[\nabla_\theta \log \pi_\theta(y_w|x)^\top \nabla_\theta \log \pi_\theta(y_l|x)] \le 0, the negative gradient pushes probability mass aggressively onto the modes of πθ\pi_\theta rather than spreading mass across all high-reward responses.

  2. Knowl 2 — Learning Dynamics and Acceleration of Reverse KL Divergence on Categorical Distributions

    theoretical result

    Consider training a categorical model distribution p(x)∝exp⁡(f(x))p(x) \propto \exp(f(x)) parameterized by independent logits f(x)f(x) to fit a target distribution q(x)q(x) via gradient descent with learning rate η\eta, starting from distribution ptp_t at step tt.

    The one-step parameter updates for the forward KL divergence DKL(q∥p)D_{\text{KL}}(q \parallel p) and reverse KL divergence DKL(p∥q)D_{\text{KL}}(p \parallel q) yield:

    log⁡pt+1f(x)pt(x)=η(q(x)−pt(x))+Z\log \frac{p_{t+1}^f(x)}{p_t(x)} = \eta \left( q(x) - p_t(x) \right) + Z

    log⁡pt+1r(x)pt(x)=η(pt(x)[log⁡q(x)pt(x)+DKL(pt(⋅)∥q(⋅))])+Z′\log \frac{p_{t+1}^r(x)}{p_t(x)} = \eta \left( p_t(x) \left[ \log \frac{q(x)}{p_t(x)} + D_{\text{KL}}(p_t(\cdot) \parallel q(\cdot)) \right] \right) + Z'

    where Z,Z′Z, Z' are log-normalizing constants. Let Δtf(x1,x2):=log⁡pt+1f(x1)pt(x1)−log⁡pt+1f(x2)pt(x2)\Delta_t^f(x_1, x_2) := \log \frac{p_{t+1}^f(x_1)}{p_t(x_1)} - \log \frac{p_{t+1}^f(x_2)}{p_t(x_2)} and Δtr(x1,x2):=log⁡pt+1r(x1)pt(x1)−log⁡pt+1r(x2)pt(x2)\Delta_t^r(x_1, x_2) := \log \frac{p_{t+1}^r(x_1)}{p_t(x_1)} - \log \frac{p_{t+1}^r(x_2)}{p_t(x_2)} denote the difference of log-probability ratios between two categories x1x_1 and x2x_2. For positive constants β,δ1,δ2>0\beta, \delta_1, \delta_2 > 0:

    1. Aggressive modification: If δ1≤pt(x1)=pt(x2)≤1−δ2\delta_1 \le p_t(x_1) = p_t(x_2) \le 1 - \delta_2 and q(x1)≥q(x2)+βq(x_1) \ge q(x_2) + \beta, then Δtr(x1,x2)>Δtf(x1,x2)\Delta_t^r(x_1, x_2) > \Delta_t^f(x_1, x_2). Due to its logarithmic dependence on q(x)q(x), reverse KL modifies probability mass more aggressively than the linear dependence in forward KL.
    2. Concentration on subset of modes: If pt(x2)+β≤pt(x1)≤1−δ2p_t(x_2) + \beta \le p_t(x_1) \le 1 - \delta_2 and q(x1)=q(x2)>c0⋅pt(x1)q(x_1) = q(x_2) > c_0 \cdot p_t(x_1) for constant c0>1c_0 > 1, then Δtr(x1,x2)>Δtf(x1,x2)\Delta_t^r(x_1, x_2) > \Delta_t^f(x_1, x_2). Reverse KL preferentially increases probability mass on the category that already has higher current probability pt(x1)p_t(x_1), committing to a subset of target modes.
    3. Aggressive pruning of low-target mass categories: If pt(x2)+β≤pt(x1)≤1−δ2p_t(x_2) + \beta \le p_t(x_1) \le 1 - \delta_2 and q(x1)=q(x2)<c1⋅pt(x2)q(x_1) = q(x_2) < c_1 \cdot p_t(x_2) for constant c1<1c_1 < 1, then Δtr(x1,x2)<Δtf(x1,x2)\Delta_t^r(x_1, x_2) < \Delta_t^f(x_1, x_2), indicating that reverse KL reduces probability mass on over-represented, less-likely target categories more rapidly than forward KL.

    Even when the parameterized model p(x)p(x) is expressive enough to represent q(x)q(x) perfectly, reverse KL achieves faster redistribution of probability mass to a subset of target modes within very few gradient steps.

  3. Knowl 3 — Unified Algorithmic Framework for Preference Fine-Tuning

    algorithm

    The unified preference fine-tuning framework systematizes policy optimization by varying on-policy sample generation, sample reuse, and the loss objective function.

    Input: Initial reference policy πref\pi_{\text{ref}}, initial policy πθ←πref\pi_\theta \leftarrow \pi_{\text{ref}}, reward model rϕr_\phi (or ground truth r∗r^*), preference dataset Dpref\mathcal{D}_{\text{pref}}, batch size BB, responses per prompt CC, mini-batch size MM, inner gradient steps TT, learning rate η\eta
    for each outer training iteration do
        Sample B/CB/C prompts [x1,x2,…,xB/C][x_1, x_2, \dots, x_{B/C}]
        if On-Policy then
            Generate CC responses per prompt from the current policy: yi1,…,yiC∼πθ(⋅∣xi)y_i^1, \dots, y_i^C \sim \pi_\theta(\cdot|x_i) for each i∈{1,…,B/C}i \in \{1, \dots, B/C\}
        else
            Sample CC responses per prompt from offline data: yi1,…,yiC∼Dprefy_i^1, \dots, y_i^C \sim \mathcal{D}_{\text{pref}} for each i∈{1,…,B/C}i \in \{1, \dots, B/C\}
        end if
        Construct dataset D={(xi,yij)}\mathcal{D} = \{(x_i, y_i^j)\} with ∣D∣=B|\mathcal{D}| = B
        Label responses yijy_i^j with proxy rewards rˉϕ(xi,yij)\bar{r}_\phi(x_i, y_i^j) or pairwise preferences
        for step t=1t = 1 to TT do
            Partition D\mathcal{D} into N=B/MN = B/M mini-batches D1,…,DN\mathcal{D}_1, \dots, \mathcal{D}_N, each of size MM
            for k=1k = 1 to NN do
                Compute loss gradient ∇θL(θ;Dk;rˉϕ)\nabla_\theta \mathcal{L}(\theta; \mathcal{D}_k; \bar{r}_\phi)
                Update policy parameters: θ←θ−η∇θL(θ;Dk;rˉϕ)\theta \leftarrow \theta - \eta \nabla_\theta \mathcal{L}(\theta; \mathcal{D}_k; \bar{r}_\phi)
            end for
        end for
    end for
    Output: Fine-tuned policy πθ\pi_\theta

    The framework controls off-policyness through two mechanisms: (1) varying total batch size BB with T=1T=1 (larger BB processes older samples across mini-batches without sample reuse); and (2) varying TT for a fixed BB (larger TT performs multiple gradient updates on the same batch, isolating sample reuse effects).

  4. Knowl 4 — Three-Axis Taxonomy of Preference Fine-Tuning Methods

    definition

    Preference fine-tuning methods for language models are categorized along three independent structural axes:

    1. On-Policy Sampling: Whether the algorithm generates new responses dynamically from the current policy πθ(⋅∣x)\pi_\theta(\cdot|x) during training (e.g., PPO, REINFORCE, online RWR, ReST, online Best-of-N, online DPO/IPO) or learns strictly from a fixed, pre-collected dataset Dpref\mathcal{D}_{\text{pref}} (e.g., offline DPO, IPO, Pref-FT, offline Best-of-N, offline RWR).
    2. Sample Reuse: For methods utilizing on-policy data, whether a sampled batch of prompt-response pairs (x,y)(x, y) is used for more than one gradient step (T>1T > 1, as in PPO and online RWR) or discarded after exactly one gradient update (T=1T = 1, as in standard REINFORCE).
    3. Negative Gradient: Whether the loss function contains a term that explicitly decreases the log-likelihood of dispreferred or low-reward responses by multiplying their gradient by a negative coefficient (e.g., PPO, REINFORCE with baseline-subtracted rewards, DPO, IPO, and Best-of-N with unlikelihood), as opposed to purely maximizing likelihood on filtered or reward-weighted samples (e.g., Pref-FT, Binary Feed-ME, RWR, ReST, vanilla Best-of-N).
    Fine-Tuning ApproachOn-Policy SamplingSample ReuseNegative Gradient
    PPOYesYesYes
    REINFORCEYesNoYes
    DPO, IPO, and variantsNoN/AYes
    Pref-FT, Binary Feed-MENoN/ANo
    Offline RWR, Offline Best-of-NNoN/ANo
    ReST, RWR, Online Best-of-NYesYesNo
  5. Knowl 5 — Geometric Alignment and Preference Data Coverage Conditions

    definition

    The behavior and performance of preference fine-tuning algorithms are governed by two geometric relationships:

    • [C1] Geometric Alignment (D(πref,exp⁡(r∗)))D(\pi_{\text{ref}}, \exp(r^*))): The divergence between the initialization reference policy πref\pi_{\text{ref}} and the ground-truth reward distribution exp⁡(r∗(x,⋅))\exp(r^*(x, \cdot)). When the reward function mode lies in the low-density tail of πref\pi_{\text{ref}} (misaligned), the policy must shift substantial mass to regions rarely sampled by πref\pi_{\text{ref}}. When the mode of r∗r^* is aligned with the mode of πref\pi_{\text{ref}}, little distribution shift is required.
    • [C2] Preference Data Coverage: The average probability density of responses in the preference training dataset Dpref\mathcal{D}_{\text{pref}} under the reference policy πref\pi_{\text{ref}}. When Dpref\mathcal{D}_{\text{pref}} has low overlap with πref\pi_{\text{ref}} (skewed away from πref\pi_{\text{ref}}) but covers the mode of r∗r^*, algorithms must navigate severe distribution mismatches between the offline data, reference policy, and reward optima.
  6. Knowl 6 — Empirical Superiority of On-Policy Sampling Under Misaligned Reward Peaks

    empirical result

    Across didactic NN-dimensional contextual bandits, synthetic LLM length-control tasks, and full-scale LLM alignment (AlpacaFarm with Pythia-1.4B), the necessity of on-policy sampling depends directly on geometric condition [C1]:

    • Misaligned Reward Optimum: In the bandit task R1R_1 (reward peak at token index 70 where πref\pi_{\text{ref}} has near-zero density) and synthetic LLM tasks ('Min Length' and 'Skew Length', where shorter lengths are rewarded but πref\pi_{\text{ref}} generates long texts), highly on-policy training (smaller batch size B=64B=64 per data collection step) converges significantly faster and achieves higher gold reward than more off-policy configurations (B=128,192,256B=128, 192, 256) across PPO, REINFORCE, and RWR.
    • Aligned Reward Optimum: In the bandit task R2R_2 (reward peak at token index 20 aligned with πref\pi_{\text{ref}}'s mode) and synthetic LLM 'Mode Length' (rewarding completion lengths close to the dataset mean of 203), varying the degree of on-policyness (B=64B=64 vs B=256B=256) yields nearly identical convergence and final rewards.
    • Full-Scale Validation: On AlpacaFarm with Pythia-1.4B, increasing batch size BB from 64 to 256 for on-policy RWR and REINFORCE systematically reduces final gold reward.
  7. Knowl 7 — Role and Mechanics of Explicit Negative Gradients in Preference Objectives

    empirical result

    Incorporating an explicit negative gradient term (pushing down the likelihood of dispreferred completions yly_l) provides substantial optimization advantages over maximum-likelihood methods:

    • Faster Convergence and Better Solutions: In fully offline settings, contrastive methods with negative gradients (DPO, IPO, and Best-of-N augmented with an unlikelihood penalty) outperform pure likelihood maximization methods (Pref-FT, standard Best-of-N, and offline RWR) whenever the reward peak is misaligned with πref\pi_{\text{ref}} (Min Length, Skew Length, AlpacaFarm, and UltraFeedback). On AlpacaFarm and UltraFeedback, DPO yields positive average gold reward gains over πref\pi_{\text{ref}}, whereas Pref-FT achieves minimal to negative gains.
    • Likelihood Margin Dynamics: Offline contrastive training (DPO) significantly widens the log-likelihood margin log⁡πθ(yw∣x)−log⁡πθ(yl∣x)\log \pi_\theta(y_w|x) - \log \pi_\theta(y_l|x) across training steps, whereas the margin for Pref-FT quickly plateaus near zero because Pref-FT increases the likelihood of both ywy_w and yly_l.
    • Capacity and Disentanglement Dependence: On smaller models or overlapping datasets (Pythia-1.4B on AlpacaFarm), DPO decreases the absolute log-likelihood of both ywy_w and yly_l while increasing their difference, shifting probability mass to out-of-distribution completions. With larger capacity and semantically distinct responses (Mistral-7B on UltraFeedback), DPO successfully increases log⁡πθ(yw∣x)\log \pi_\theta(y_w|x) while decreasing log⁡πθ(yl∣x)\log \pi_\theta(y_l|x).
  8. Knowl 8 — On-Policy Contrastive Optimization (Online DPO/IPO)

    model/method

    On-policy contrastive optimization combines on-policy sampling with contrastive negative-gradient losses. In each training iteration:

    1. Given a prompt xx, sample NN candidate responses y1,…,yN∼πθ(⋅∣x)y_1, \dots, y_N \sim \pi_\theta(\cdot|x) from the current policy snapshot.
    2. Score and rank candidates using a learned proxy reward model rϕ(x,y)r_\phi(x, y) to form preference pairs (x,yw,yl)(x, y_w, y_l) where rϕ(x,yw)>rϕ(x,yl)r_\phi(x, y_w) > r_\phi(x, y_l).
    3. Update πθ\pi_\theta by minimizing the contrastive loss (DPO or IPO) evaluated on these freshly generated on-policy preference pairs:

    Lon-policy DPO(πθ)=−E(x,yw,yl)[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{on-policy DPO}}(\pi_\theta) = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right]

    Lon-policy IPO(πθ)=E(x,yw,yl)[(log⁡πθ(yw∣x)πref(yl∣x)πref(yw∣x)πθ(yl∣x)−τ−12)2]\mathcal{L}_{\text{on-policy IPO}}(\pi_\theta) = \mathbb{E}_{(x, y_w, y_l)} \left[ \left( \log \frac{\pi_\theta(y_w|x)\pi_{\text{ref}}(y_l|x)}{\pi_{\text{ref}}(y_w|x)\pi_\theta(y_l|x)} - \frac{\tau^{-1}}{2} \right)^2 \right]

    On didactic bandits and synthetic LLM length tasks, on-policy DPO/IPO converges significantly faster and reaches higher final rewards than offline DPO/IPO, on-policy RL (PPO, REINFORCE), and on-policy supervised methods (RWR, Best-of-N). The contrastive negative gradient provides a lower-variance, aggressive redistribution signal across the support of the current policy.

  9. Knowl 9 — Effects of Sample Reuse and Importance Clipping on Policy Optimization

    empirical result

    Varying the number of inner gradient steps TT per sampled on-policy batch reveals a clear trade-off between sample efficiency and off-policy degradation:

    • Non-Monotonic Benefit of Mild Reuse: In both didactic bandits and synthetic LLM tasks, moderate sample reuse (e.g., T=2T=2 or T=5T=5) learns faster per generated sample than T=1T=1, improving data efficiency without sacrificing asymptotic reward.
    • Excessive Reuse Degradation: Increasing TT to excessive levels (T=8T=8) degrades performance and causes training instability or propensity overfitting in methods lacking off-policy corrections (such as on-policy Best-of-N and RWR), because maximum-likelihood updates on stale samples anchor πθ\pi_\theta to outdated policy distributions.
    • Robustness of PPO: PPO demonstrates substantially greater resilience to large TT compared to Best-of-N and RWR due to its importance sampling ratio clipping mechanism Clip(r(θ),1−ϵ,1+ϵ)\text{Clip}(r(\theta), 1-\epsilon, 1+\epsilon), which prevents severely off-policy samples from generating destructive gradient updates.
  10. Knowl 10 — Wall-Clock and Sample Efficiency Tradeoffs across Preference Fine-Tuning Algorithms

    data/table

    Across contextual bandit and synthetic LLM benchmarks, on-policy contrastive methods (On-policy DPO/IPO) achieve superior or matching reward values with significantly lower wall-clock time compared to both offline contrastive methods and on-policy RL (PPO, RWR).

    Method Bandit (R1) Min Length Skew Length
    Reward (↑\uparrow) Time Completion Length (↓\downarrow) Time Completion Length (↓\downarrow) Time
    Offline DPO / IPO 0.82 (0.04) 1.7 h 1.0 (0.0) 1.3 h 11.8 (14.0) 0.12 h
    On-policy PPO 0.92 (0.01) 0.93 h 20.5 (25.4) 4.84 h 15.8 (11.1) 7.26 h
    On-policy RWR 0.88 (0.01) 0.12 h 65.5 (36.7) 15.5 h 15.8 (9.3) 15.5 h
    On-policy DPO / IPO 0.92 (0.01) 0.12 h 1.0 (0.0) 0.4 h 0.0 (0.0) 0.4 h

    In the bandit task R1R_1, on-policy DPO/IPO achieves a maximum reward of 0.92 in 0.12 hours (matching PPO's 0.92 reward which took 0.93 hours). In 'Min Length', on-policy DPO reaches optimal completion length 1.0 in 0.4 hours, whereas offline DPO requires 1.3 hours, and PPO/RWR fail to reach optimal length after 4.84 and 15.5 hours respectively. In 'Skew Length', offline DPO converges prematurely to a suboptimal completion length of 11.8, while on-policy DPO achieves optimal length 0.0 in 0.4 hours. Experiments were run on a single A40 GPU for synthetic LLM tasks and Intel Xeon E5-2698 v4 CPU (4 threads) for bandits.

  11. Knowl 11 — Limitations of Preference Fine-Tuning Analysis Framework

    limitation

    The analysis framework and findings possess several explicit limitations:

    1. Lack of Formal Statistical Guarantees for Negative Gradients: While empirical results and mode-seeking connections show advantages for negative gradient terms, the paper does not derive formal statistical variance bounds or finite-sample error guarantees quantifying the exact reduction in learning signal variance.
    2. Exclusion of Pre-Training Data Coverage: The framework evaluates data coverage by measuring the density of preference data Dpref\mathcal{D}_{\text{pref}} relative to the supervised fine-tuning reference policy πref\pi_{\text{ref}}, but does not model the broader coverage distribution of the original pre-training corpus.
    3. Simplified Reward Model Quality Dynamics: In didactic settings, the reward function is assumed exact, and in synthetic/real LLM tasks, reward model errors are simplified. The analysis does not fully explore the impact of reward model capacity, parameterization architectures, and reward hacking on policy optimization.

Coverage note — None was omitted; all key theoretical unifications (reverse vs. forward KL mode-seeking properties), the unified algorithm, taxonomic categorizations, experimental conditions [C1]/[C2], detailed empirical findings (on-policy sampling, negative gradients, sample reuse, on-policy DPO), performance tables, and stated limitations were fully extracted.

References

  1. 1.Adolphs, L., Gao, T., Xu, J., Shuster, K., Sukhbaatar, S., and Weston, J. The CRINGE Loss: Learning what language not to model. arXiv e-prints, art. arXiv:2211.05826, November 2022. doi: 10.48550/arXiv.2211.05826.
  2. 2.Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O. Gkd: Generalized knowledge distillation for auto-regressive sequence models. arXiv preprint arXiv:2306.13649, 2023.
  3. 3.Ahmadian, A., Cremer, C., Galle, M., Fadaee, M., Kreutzer, J., Üstün, A., and Hooker, S. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms. arXiv preprint arXiv:2402.14740, 2024.
  4. 4.Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952. ISSN 00063444. URL http://www.jstor.org/stable/2334029.
  5. 5.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020.
  6. 6.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T. T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Biyik, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D. Open problems and fundamental limitations of reinforcement learning from human feedback. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=bx24KpJ4Eb. Survey Certification.
  7. 7.Chang, J. D., Zhan, W., Oertell, O., Brantley, K., Misra, D., Lee, J. D., and Sun, W. Dataset Reset Policy Optimization for RLHF. arXiv e-prints, art. arXiv:2404.08495, April 2024. doi: 10.48550/arXiv.2404.08495.
  8. 8.Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. arXiv e-prints, art. arXiv:2401.01335, January 2024. doi: 10.48550/arXiv.2401.01335.
  9. 9.ContextualAI. Human-centered loss functions (halos), 2024. URL https://github.com/ContextualAI/HALOs.
  10. 10.Coste, T., Anwar, U., Kirk, R., and Krueger, D. Reward Model Ensembles Help Mitigate Overoptimization. arXiv e-prints, art. arXiv:2310.02743, October 2023. doi: 10.48550/arXiv.2310.02743.
  11. 11.Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. UltraFeedback: Boosting Language Models with High-quality Feedback. arXiv e-prints, art. arXiv:2310.01377, October 2023. doi: 10.48550/arXiv.2310.01377.
  12. 12.Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233, 2023.
  13. 13.Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T. RAFT: Reward ranked finetuning for generative foundation model alignment. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=m7p5O7zblY.
  14. 14.Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacafarm: A simulation framework for methods that learn from human feedback, 2024.
  15. 15.Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., Shaw, P., and Berant, J. Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking. arXiv e-prints, art. arXiv:2312.09244, December 2023. doi: 10.48550/arXiv.2312.09244.
  16. 16.Ethayarajh, K., Xu, W., Jurafsky, D., and Kiela, D. Human-aware loss functions (halos). Technical report, Contextual AI, 2023. URL https://github.com/ContextualAI/HALOs/blob/main/assets/report.pdf.
  17. 17.Gao, L., Schulman, J., and Hilton, J. Scaling Laws for Reward Model Overoptimization. arXiv e-prints, art. arXiv:2210.10760, October 2022. doi: 10.48550/arXiv.2210.10760.
  18. 18.Gheshlaghi Azar, M., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. A General Theoretical Paradigm to Understand Learning from Human Preferences. arXiv e-prints, art. arXiv:2310.12036, October 2023. doi: 10.48550/arXiv.2310.12036.
  19. 19.Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  20. 20.Guo, S., Zhang, B., Liu, T., Liu, T., Khalman, M., Llinares, F., Rame, A., Mesnard, T., Zhao, Y., Piot, B., Ferret, J., and Blondel, M. Direct Language Model Alignment from Online AI Feedback. arXiv e-prints, art. arXiv:2402.04792, February 2024. doi: 10.48550/arXiv.2402.04792.
  21. 21.Hong, J., Lee, N., and Thorne, J. Orpo: Monolithic preference optimization without reference model, 2024.
  22. 22.Hu, J., Tao, L., Yang, J., and Zhou, C. Aligning Language Models with Offline Learning from Human Feedback. arXiv e-prints, art. arXiv:2308.12050, August 2023. doi: 10.48550/arXiv.2308.12050.
  23. 23.Jain, S., Kirk, R., Singh Lubana, E., Dick, R. P., Tanaka, H., Grefenstette, E., Rocktäschel, T., and Krueger, D. S. Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. arXiv e-prints, art. arXiv:2311.12786, November 2023. doi: 10.48550/arXiv.2311.12786.
  24. 24.Karpathy, A. minGPT. URL https://github.com/karpathy/minGPT.
  25. 25.Khaki, S., Li, J., Ma, L., Yang, L., and Ramachandra, P. RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models. arXiv e-prints, art. arXiv:2402.10038, February 2024. doi: 10.48550/arXiv.2402.10038.
  26. 26.Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. Morel: Model-based offline reinforcement learning. arXiv preprint arXiv:2005.05951, 2020.
  27. 27.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2017.
  28. 28.Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R. Understanding the Effects of RLHF on LLM Generalisation and Diversity. arXiv e-prints, art. arXiv:2310.06452, October 2023. doi: 10.48550/arXiv.2310.06452.
  29. 29.Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting. Advances in Neural Information Processing Systems, 35:16203–16220, 2022.
  30. 30.Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R. A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv e-prints, art. arXiv:2401.01967, January 2024. doi: 10.48550/arXiv.2401.01967.
  31. 31.Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J. Statistical rejection sampling improves preference optimization. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=xbjSwwrQOe.
  32. 32.Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. Quark: Controllable text generation with reinforced unlearning. Advances in neural information processing systems, 35:27591–27609, 2022.
  33. 33.Mei, J., Chung, W., Thomas, V., Dai, B., Szepesvari, C., and Schuurmans, D. The role of baselines in policy gradient optimization. Advances in Neural Information Processing Systems, 35:17818–17830, 2022.
  34. 34.Mukobi, G., Chatain, P., Fong, S., Windesheim, R., Kutyniok, G., Bhatia, K., and Alberti, S. SuperHF: Supervised Iterative Learning from Human Feedback. arXiv e-prints, art. arXiv:2310.16763, October 2023. doi: 10.48550/arXiv.2310.16763.
  35. 35.Munos, R. and Szepesvári, C. Finite-time bounds for fitted value iteration. Journal of Machine Learning Research, 9 (5), 2008.
  36. 36.Munos, R., Valko, M., Calandriello, D., Gheshlaghi Azar, M., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., Selvi, M., Girgin, S., Momchev, N., Bachem, O., Mankowitz, D. J., Precup, D., and Piot, B. Nash Learning from Human Feedback. arXiv e-prints, art. arXiv:2312.00886, December 2023. doi: 10.48550/arXiv.2312.00886.
  37. 37.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback, 2022.
  38. 38.Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J. Iterative reasoning preference optimization, 2024.
  39. 39.Peters, J. and Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. In Proceedings of the 24th International Conference on Machine Learning, pp. 745–750. ACM, 2007.
  40. 40.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018.
  41. 41.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  42. 42.Rafailov, R., Hejna, J., Park, R., and Finn, C. From r to q*: Your language model is secretly a q-function. arXiv preprint arXiv:2404.12358, 2024.
  43. 43.Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T. Direct nash optimization: Teaching language models to self-improve with general preferences. arXiv preprint arXiv:2404.03715, 2024.
  44. 44.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal Policy Optimization Algorithms. arXiv e-prints, art. arXiv:1707.06347, July 2017. doi: 10.48550/arXiv.1707.06347.
  45. 45.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  46. 46.Sharma, A., Keh, S., Mitchell, E., Finn, C., Arora, K., and Kollar, T. A critical evaluation of ai feedback for aligning large language models. arXiv preprint arXiv:2402.12366, 2024.
  47. 47.Singhal, P., Goyal, T., Xu, J., and Durrett, G. A long way to go: Investigating length correlations in rlhf. arXiv preprint arXiv:2310.03716, 2023.
  48. 48.Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Solla, S., Leen, T., and Müller, K. (eds.), Advances in Neural Information Processing Systems, volume 12. MIT Press, 1999. URL https://proceedings.neurips.cc/paper_files/paper/1999/file/464d828b85b0bed98e80ade0a5c43b0f-Paper.pdf.
  49. 49.Swaminathan, A. and Joachims, T. The self-normalized estimator for counterfactual learning. In advances in neural information processing systems, pp. 3231–3239, 2015.
  50. 50.Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A. A minimaximalist approach to reinforcement learning from human feedback. arXiv preprint arXiv:2401.04056, 2024.
  51. 51.Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T. Zephyr: Direct distillation of lm alignment, 2023.
  52. 52.von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., and Huang, S. Trl: Transformer reinforcement learning. https://github.com/huggingface/trl, 2020.
  53. 53.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SJeYe0NtvH.
  54. 54.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3-4):229–256, May 1992.
  55. 55.Xie, A., Tajwar, F., Sharma, A., and Finn, C. When to ask for help: Proactive interventions in autonomous reinforcement learning. In Advances in Neural Information Processing Systems, volume 35, 2022.
  56. 56.Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. arXiv e-prints, art. arXiv:2312.11456, December 2023. doi: 10.48550/arXiv.2312.11456.
  57. 57.Xu, S., Fu, W., Gao, J., Ye, W., Liu, W., Mei, Z., Wang, G., Yu, C., and Wu, Y. Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study. arXiv e-prints, art. arXiv:2404.10719, April 2024. doi: 10.48550/arXiv.2404.10719.
  58. 58.Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. Mastering visual continuous control: Improved data-augmented reinforcement learning. arXiv preprint arXiv:2107.09645, 2021a.
  59. 59.Yarats, D., Kostrikov, I., and Fergus, R. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=GY6-6sTvGaf.
  60. 60.Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. Mopo: Model-based offline policy optimization. arXiv preprint arXiv:2005.13239, 2020.
  61. 61.Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. Combo: Conservative offline model-based policy optimization. Advances in neural information processing systems, 34:28954–28967, 2021.
  62. 62.Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. Self-rewarding language models. arXiv preprint arXiv:2401.10020, 2024.
  63. 63.Yuan, W., Yuanzhe Pang, R., Cho, K., Li, X., Sukhbaatar, S., Xu, J., and Weston, J. Self-Rewarding Language Models. arXiv e-prints, art. arXiv:2401.10020, January 2024. doi: 10.48550/arXiv.2401.10020.
  64. 64.Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J. SLiC-HF: Sequence Likelihood Calibration with Human Feedback. arXiv e-prints, art. arXiv:2305.10425, May 2023. doi: 10.48550/arXiv.2305.10425.
  65. 65.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences, 2020.

Citation

MLA
Tajwar, F., et al. “Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data”. arXiv, 2024, http://arxiv.org/abs/2404.14367v3.
APA
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., & Kumar, A. (2024). Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data. arXiv. http://arxiv.org/abs/2404.14367v3
Chicago
Tajwar, F., A. Singh, A. Sharma, et al. 2024. “Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data”. arXiv. http://arxiv.org/abs/2404.14367v3.
Harvard
Tajwar, F. et al. (2024) “Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.14367v3.
Vancouver
1. Tajwar F, Singh A, Sharma A, Rafailov R, Schneider J, Xie T, Ermon S, Finn C, Kumar A (2024) Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data. arXiv

BibTeX

@article{tajwar2024preference,
  title = {Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data},
  author = {Tajwar, Fahim and Singh, Anikait and Sharma, Archit and Rafailov, Rafael and Schneider, Jeff and Xie, Tengyang and Ermon, Stefano and Finn, Chelsea and Kumar, Aviral},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.14367v3},
  eprint = {2404.14367}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/