Human Alignment of Large Language Models through Online Preference Optimisation

Daniele CalandrielloZhaohan Daniel GuoRémi MunosMark RowlandYunhao TangBernardo Ávila PiresPierre Harvey RichemondCharline Le LanMichal ValkoTianqi Liu

article2024ICML89 citations

Establishes a theoretical equivalence between offline Identity Preference Optimisation and online Nash Mirror Descent, introducing a unified algorithm that combines online self-play with regularized mixture sampling to improve language model alignment.

Listen

Aligning large language models with human preferences is essential for ensuring that automated text generation remains helpful, reliable, and safe. While traditional alignment relies on Reinforcement Learning from Human Feedback or offline direct optimization techniques, these strategies often face practical limitations. Offline methods frequently suffer from distribution shifts when model outputs drift away from static training datasets, whereas reinforcement learning approaches can be computationally unstable and prone to reward gaming.

The article establishes a theoretical framework connecting contrastive preference optimization with game-theoretic self-play, introducing two new alignment methods: Online Identity Preference Optimisation (Online IPO) and Identity Preference Optimisation with Mirror Descent (IPO-MD). The primary objective is to demonstrate how combining contrastive loss functions with dynamic, online sampling improves the stability and alignment quality of language models.

The authors conducted mathematical derivations to prove the theoretical equivalence of these methods and evaluated them empirically on an article summarization benchmark using a 770-million-parameter encoder-decoder model. The training framework utilized a 3-billion-parameter preference model to provide automated feedback on newly generated outputs, and evaluation was performed through automated side-by-side comparisons using PaLM 2 across multiple random seeds.

The investigation produced four central findings. First, mathematically, Online IPO's expected update direction is equivalent to finding a regularized Nash equilibrium through self-play in a two-player game. Second, online alignment methods overwhelmingly outperformed their offline counterparts, achieving win rates exceeding 95% against static offline baselines. Third, Online IPO and IPO-MD achieved the highest performance overall, winning approximately 60% of side-by-side evaluations against existing direct preference methods and over 77% against standard reinforcement learning baselines. Fourth, IPO-MD effectively bridges online and offline dynamics by sampling from a mixture policy, allowing smooth interpolation between exploratory self-play and regularized baseline stability.

These results indicate that active generation during alignment significantly enhances model output quality by keeping training data aligned with the model's evolving capabilities. However, shifting from offline datasets to online sampling introduces a practical engineering trade-off: real-time generation during training slows processing speed roughly threefold compared to loading pre-existing offline datasets. Organizations must weigh this additional computational expense against substantial gains in generation quality and safety.

Teams developing language models should consider adopting online contrastive methods like Online IPO or IPO-MD when high performance and robustness are critical. Before widespread production deployment across diverse domains, further evaluation is recommended on large-scale models exceeding 100 billion parameters and on open-ended conversational tasks. Because current empirical findings are established on a single summarization task with medium-sized models evaluated via an automated judge, practitioners should conduct targeted validation within their specific operational workflows.

arXiv: 2403.08635
Cover for Human Alignment of Large Language Models through Online Preference Optimisation

Abstract

Ensuring alignment of language models’ outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensively studied recently and several methods such as Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimisation (DPO) and Sequence Likelihood Calibration (SLiC) have emerged. In this paper, our contribution is two-fold. First, we show the equivalence between two recent alignment methods, namely Identity Preference Optimisation (IPO) and Nash Mirror Descent (Nash-MD). Second, we introduce a generalisation of IPO, named IPO-MD, that leverages the regularised sampling approach proposed by Nash-MD. This equivalence may seem surprising at first sight, since IPO is an offline method whereas Nash-MD is an online method using a preference model. However, this equivalence can be proven when we consider the online version of IPO, that is when both generations are sampled by the online policy and annotated by a trained preference model. Optimising the IPO loss with such a stream of data becomes then equivalent to finding the Nash equilibrium of the preference model through self-play. Building on this equivalence, we introduce the IPO-MD algorithm that generates data with a mixture policy (between the online and reference policy) similarly as the general Nash-MD algorithm. We compare online-IPO and IPO-MD to different online versions of existing losses on preference data such as DPO and SLiC on a summarisation task.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Preference optimisation in bandits
  • 2.2. RLHF with a Bradley-Terry reward model
  • 2.3. Direct preference optimisation (DPO)
  • 2.4. Sequence Likelihood Calibration (SLiC)
  • 2.5. Identity preference optimisation (IPO)
  • 2.6. Nash-MD-PG
  • 3. Comparative discussion of preference optimisation algorithms
  • 4. Online IPO
  • 4.1. Algorithm
  • 4.2. Analysis
  • 4.3. Online DPO
  • 5. IPO-MD
  • 5.1. Algorithm
  • 5.2. Analysis
  • 6. Experiments
  • 6.1. Main results
  • 6.2. Ablations and Additional Results on Summarisation
  • 6.3. Offline vs Online Settings
  • 7. Conclusion
  • Acknowledgements
  • Impact statement
  • References
  • A. Related Work
  • B. Additional Implementation Details
  • B.1. Pseudo-Codes for Offline and Online Contrastive Preference Algorithm
  • B.2. Diagrams for Offline and Online Contrastive Preference Algorithm
  • C. Additional Experimental Results
  • C.1. Regularisation Sweep for Online and Offline
  • C.2. Mixing ratio curve
  • C.3. Best hyperparameters found and KL values for chosen checkpoints
  • D. Proofs
  • E. Comparison of the variance of contrastive versus non-contrastive gradient estimates
  • E.1. A Toy Example
  • F. Tabular example
  • G. Supplementary Theoretical Study of Online DPO
  • G.1. Results and Proofs

Knowls

  1. Knowl 1 — Equivalence of Online Identity Preference Optimisation to Preference Game Nash Equilibrium

    theoretical result

    Consider a finite action space Y\mathcal{Y}, a reference policy πref∈Δ(Y)\pi_{\text{ref}} \in \Delta(\mathcal{Y}), a preference function p:Y×Y→[0,1]p: \mathcal{Y} \times \mathcal{Y} \to [0, 1] satisfying p(y′≻y)=1−p(y≻y′)p(y' \succ y) = 1 - p(y \succ y'), and a regularisation temperature τ>0\tau > 0.

    The population loss for Online Identity Preference Optimisation (Online IPO) replaces the static data distribution of offline IPO with samples generated from the current policy π\pi:

    EY,Y′∼SG[π],(Y+,Y−)∼λp(Y,Y′)[(log⁡π(Y+)πref(Y−)π(Y−)πref(Y+)−12τ)2]\mathbb{E}_{Y, Y' \sim \text{SG}[\pi], (Y^+, Y^-) \sim \lambda_p(Y, Y')} \left[ \left( \log \frac{\pi(Y^+)\pi_{\text{ref}}(Y^-)}{\pi(Y^-)\pi_{\text{ref}}(Y^+)} - \frac{1}{2\tau} \right)^2 \right]

    where SG[⋅]\text{SG}[\cdot] denotes a stop-gradient operator on the sampling policy, and λp(y,y′)\lambda_p(y, y') samples (y,y′)(y, y') with probability p(y≻y′)p(y \succ y') and (y′,y)(y', y) with probability 1−p(y≻y′)1 - p(y \succ y').

    A policy π∈Δ(Y)\pi \in \Delta(\mathcal{Y}) achieves a zero gradient for the Online IPO population loss if and only if it satisfies the fixed-point condition:

    π(y)∝πref(y)exp⁡(1τEY′∼π[p(y≻Y′)])\pi(y) \propto \pi_{\text{ref}}(y) \exp\left( \frac{1}{\tau} \mathbb{E}_{Y' \sim \pi}[p(y \succ Y')] \right)

    This condition is identical to the best-response condition against itself in the two-player constant-sum regularised preference game where each player i∈{1,2}i \in \{1, 2\} chooses policy πi\pi_i to maximize:

    EY∼πi,Y′∼π−i[p(Y≻Y′)]−τKL(πi∥πref)+τKL(π−i∥πref)\mathbb{E}_{Y \sim \pi_i, Y' \sim \pi_{-i}}[p(Y \succ Y')] - \tau \text{KL}(\pi_i \parallel \pi_{\text{ref}}) + \tau \text{KL}(\pi_{-i} \parallel \pi_{\text{ref}})

    Therefore, the minimiser of the Online IPO objective is the Nash equilibrium of this regularised two-player preference game.

  2. Knowl 2 — Equivalence of Online IPO Expected Gradient to Game Self-Play Updates

    theoretical result

    For a policy π\pi parameterised by ϕ\phi, the expected gradient of the Online Identity Preference Optimisation (Online IPO) population loss is identical to the self-play policy gradient update direction in the two-player regularised constant-sum preference game.

    Specifically, the self-play update performs gradient ascent on expected payoff against an opponent playing the same policy:

    ∇ϕ(EY∼π,Y′∼SG[π][p(Y≻Y′)]−τKL(π∥πref))=∑y∈Yπ(y)(EY′∼π[p(y≻Y′)]−τlog⁡π(y)πref(y))∇ϕlog⁡π(y)\nabla_\phi \left( \mathbb{E}_{Y \sim \pi, Y' \sim \text{SG}[\pi]}[p(Y \succ Y')] - \tau \text{KL}(\pi \parallel \pi_{\text{ref}}) \right) = \sum_{y \in \mathcal{Y}} \pi(y) \left( \mathbb{E}_{Y' \sim \pi}[p(y \succ Y')] - \tau \log \frac{\pi(y)}{\pi_{\text{ref}}(y)} \right) \nabla_\phi \log \pi(y)

    The expected gradient of the contrastive Online IPO loss evaluates to this exact same vector quantity. Thus, Online IPO provides a contrastive, pair-based gradient estimator that targets the exact self-play trajectory of the regularised preference game.

  3. Knowl 3 — Identity Preference Optimisation with Mirror Descent (IPO-MD)

    model/method

    Identity Preference Optimisation with Mirror Descent (IPO-MD) is a preference optimisation method that samples data from a geometric mixture distribution between the online policy and the reference policy, controlled by a mixture parameter β∈[0,1]\beta \in [0, 1].

    Given prompt xx, learnable policy πθ\pi_\theta, reference policy πref\pi_{\text{ref}}, temperature τ>0\tau > 0, and pairwise preference model pϕp_\phi, the IPO-MD(β\beta) population loss is defined as:

    EY,Y′∼SG[πβ],(Y+,Y−)∼λp(Y,Y′)[(log⁡πθ(Y+∣x)πref(Y−∣x)πθ(Y−∣x)πref(Y+∣x)−12τ)2]\mathbb{E}_{Y, Y' \sim \text{SG}[\pi_\beta], (Y^+, Y^-) \sim \lambda_p(Y, Y')} \left[ \left( \log \frac{\pi_\theta(Y^+\mid x)\pi_{\text{ref}}(Y^-\mid x)}{\pi_\theta(Y^-\mid x)\pi_{\text{ref}}(Y^+\mid x)} - \frac{1}{2\tau} \right)^2 \right]

    where πβ∝πθ1−β(πref)β\pi_\beta \propto \pi_\theta^{1-\beta} (\pi_{\text{ref}})^\beta is the normalized geometric mixture policy. When β=0\beta = 0, IPO-MD reduces to Online IPO (self-play); when β=1\beta = 1, it optimizes preferences against the fixed reference policy πref\pi_{\text{ref}}.

    Because exact sequence-level autoregressive sampling from πθ1−β(πref)β\pi_\theta^{1-\beta}(\pi_{\text{ref}})^\beta is intractable, practical implementations sample tokens autoregressively from the one-step-at-a-time logit mixture π^β\hat{\pi}_\beta:

    log⁡π^β(yt∣y<t,x)=(1−β)log⁡πθ(yt∣y<t,x)+βlog⁡πref(yt∣y<t,x)+C(y<t,x)\log \hat{\pi}_\beta(y_t \mid y_{<t}, x) = (1 - \beta) \log \pi_\theta(y_t \mid y_{<t}, x) + \beta \log \pi_{\text{ref}}(y_t \mid y_{<t}, x) + C(y_{<t}, x)

    where C(y<t,x)C(y_{<t}, x) is a normalizing constant.

  4. Knowl 4 — Fixed Points, Equilibria, and Gradient Formulation of IPO-MD

    theoretical result

    Let β∈[0,1)\beta \in [0, 1) and τ>0\tau > 0. The fixed point policy πβ∗\pi_\beta^* of IPO-MD(β\beta) satisfies:

    πβ∗(y)∝πref(y)exp⁡(1τEY′∼(πβ∗)1−β(πref)β[p(y≻Y′)])\pi_\beta^*(y) \propto \pi_{\text{ref}}(y) \exp\left( \frac{1}{\tau} \mathbb{E}_{Y' \sim (\pi_\beta^*)^{1-\beta}(\pi_{\text{ref}})^\beta}[p(y \succ Y')] \right)

    This fixed point coincides exactly with the fixed point of the Nash-MD-PG(β\beta) algorithm.

    Furthermore, the induced geometric mixture policy πβ′=(πβ∗)1−β(πref)β\pi'_\beta = (\pi_\beta^*)^{1-\beta}(\pi_{\text{ref}})^\beta is the exact Nash equilibrium of the two-player regularised preference game with modified regularisation parameter τ~=τ(1−β)−1\tilde{\tau} = \tau (1 - \beta)^{-1}.

    The expected gradient of IPO-MD(β\beta) is related to that of Nash-MD-PG(β\beta) by:

    gNash-MD-PG(β)=−Ey∼π[g(y)],gIPO-MD(β)=−2τEy∼π1−β(πref)β[g(y)]g_{\text{Nash-MD-PG}}(\beta) = -\mathbb{E}_{y \sim \pi}[g(y)], \quad g_{\text{IPO-MD}}(\beta) = -\frac{2}{\tau} \mathbb{E}_{y \sim \pi^{1-\beta}(\pi_{\text{ref}})^\beta}[g(y)]

    where g(y)=∇log⁡π(y)(p(y≻π1−β(πref)β)−12−τlog⁡π(y)πref(y))g(y) = \nabla \log \pi(y) \left( p(y \succ \pi^{1-\beta}(\pi_{\text{ref}})^\beta) - \frac{1}{2} - \tau \log \frac{\pi(y)}{\pi_{\text{ref}}(y)} \right).

  5. Knowl 5 — Online Contrastive Preference Optimization Framework

    algorithm

    The Online Contrastive Preference Optimization framework continuously generates action pairs from the current policy (or policy mixture) and labels them using a learned preference model pϕp_\phi, bypassing the requirement for static offline datasets.

    Input: Prompt dataset {xix_i}, parameterized policy πθ\pi_\theta, reference policy πref\pi_{\text{ref}}, preference model pϕp_\phi, total steps KK, batch size BB, optimizer UpdateOptimizer, regularization parameter τ\tau, algorithm loss LALGO\mathcal{L}_{\text{ALGO}}
    Output: Optimized policy parameters θ\theta
    for k=1k = 1 to KK do
        Sample batch of prompts {xix_i}i=1B_{i=1}^B uniformly from dataset
        for each prompt xix_i in the batch do
            Sample two independent generations (yi,yi′)∼πθ(⋅∣xi)(y_i, y'_i) \sim \pi_\theta(\cdot \mid x_i)
            Compute preference pi=pϕ(yi≻yi′∣xi)p_i = p_\phi(y_i \succ y'_i \mid x_i)
        end for
        Compute batch loss:
        L(θ)=1B∑i=1B(piLALGO(θ,xi,yi,yi′)+(1−pi)LALGO(θ,xi,yi′,yi))\mathcal{L}(\theta) = \frac{1}{B} \sum_{i=1}^B \left( p_i \mathcal{L}_{\text{ALGO}}(\theta, x_i, y_i, y'_i) + (1 - p_i) \mathcal{L}_{\text{ALGO}}(\theta, x_i, y'_i, y_i) \right)
        Update policy parameters: θ←UpdateOptimizer(θ,L(θ))\theta \leftarrow \text{UpdateOptimizer}(\theta, \mathcal{L}(\theta))
    end for
    return θ\theta

    The loss LALGO(θ,x,y,y′)\mathcal{L}_{\text{ALGO}}(\theta, x, y, y') is instantiated for each algorithm as follows:

    • IPO (simplified expanded form): −log⁡πθ(y∣x)πθ(y′∣x)+τ(log⁡πθ(y∣x)πref(y′∣x)πθ(y′∣x)πref(y∣x))2-\log \frac{\pi_\theta(y \mid x)}{\pi_\theta(y' \mid x)} + \tau \left( \log \frac{\pi_\theta(y \mid x)\pi_{\text{ref}}(y' \mid x)}{\pi_\theta(y' \mid x)\pi_{\text{ref}}(y \mid x)} \right)^2
    • DPO: σ(τlog⁡πθ(y∣x)πref(y′∣x)πθ(y′∣x)πref(y∣x))\sigma\left( \tau \log \frac{\pi_\theta(y \mid x)\pi_{\text{ref}}(y' \mid x)}{\pi_\theta(y' \mid x)\pi_{\text{ref}}(y \mid x)} \right)
    • SLiC: max⁡(0,1−τlog⁡πθ(y∣x)πref(y′∣x)πθ(y′∣x)πref(y∣x))\max\left(0, 1 - \tau \log \frac{\pi_\theta(y \mid x)\pi_{\text{ref}}(y' \mid x)}{\pi_\theta(y' \mid x)\pi_{\text{ref}}(y \mid x)}\right)
  6. Knowl 6 — Variance Reduction Condition for Contrastive Preference Gradient Estimators

    theoretical result

    Let y,y′∼πy, y' \sim \pi be independent samples from policy π\pi, and define f(y,y′):=p(y≻y′)−12−τlog⁡π(y)πref(y)+τlog⁡π(y′)πref(y′)f(y, y') := p(y \succ y') - \frac{1}{2} - \tau \log \frac{\pi(y)}{\pi_{\text{ref}}(y)} + \tau \log \frac{\pi(y')}{\pi_{\text{ref}}(y')}, which satisfies anti-symmetry f(y,y′)=−f(y′,y)f(y, y') = -f(y', y).

    Consider the non-contrastive estimator g^non-contrastive=−∇log⁡π(y)f(y,y′)\hat{g}_{\text{non-contrastive}} = -\nabla \log \pi(y) f(y, y') and the contrastive estimator g^contrastive=−12(∇log⁡π(y)−∇log⁡π(y′))f(y,y′)\hat{g}_{\text{contrastive}} = -\frac{1}{2} (\nabla \log \pi(y) - \nabla \log \pi(y')) f(y, y'). Both estimators have identical expectations.

    A sufficient condition for the variance of the contrastive gradient estimator to be less than or equal to the variance of the non-contrastive estimator (Var(g^contrastive)≤Var(g^non-contrastive)\text{Var}(\hat{g}_{\text{contrastive}}) \le \text{Var}(\hat{g}_{\text{non-contrastive}})) is:

    Ey,y′∼π[∇log⁡π(y)∇log⁡π(y′)f(y,y′)2]≥0\mathbb{E}_{y, y' \sim \pi} \left[ \nabla \log \pi(y) \nabla \log \pi(y') f(y, y')^2 \right] \ge 0

    When this condition is met, contrastive policy updates achieve variance reduction over non-contrastive self-play policy gradient estimates through the principle of antithetic variates.

  7. Knowl 7 — Stationary Point Characterization and Discrepancy of Online DPO

    theoretical result

    The Nash equilibrium π∗\pi^* of the regularised preference game is a stationary point of Online Direct Preference Optimisation (Online DPO) if and only if:

    p(y≻π∗)=∑y′∈Yπ∗(y′)σ(p(y≻π∗)−p(y′≻π∗)),∀y∈Yp(y \succ \pi^*) = \sum_{y' \in \mathcal{Y}} \pi^*(y') \sigma\left( p(y \succ \pi^*) - p(y' \succ \pi^*) \right), \quad \forall y \in \mathcal{Y}

    In any two-action problem (∣Y∣=2|\mathcal{Y}| = 2), this condition cannot be satisfied unless the preference probabilities are completely uniform (p(y1≻y2)=12p(y_1 \succ y_2) = \frac{1}{2}). Consequently, Online DPO and Online IPO optimize fundamentally different objectives and do not share stationary points in general.

    However, under the Bradley-Terry model assumption where p(y≻y′)=σ(r(y)−r(y′))p(y \succ y') = \sigma(r(y) - r(y')) for a ground-truth reward function r:Y→Rr: \mathcal{Y} \to \mathbb{R}, the standard RLHF solution πr(y)∝πref(y)exp⁡(1τr(y))\pi^r(y) \propto \pi_{\text{ref}}(y) \exp\left( \frac{1}{\tau} r(y) \right) is a stationary point of Online DPO.

  8. Knowl 8 — Side-by-Side Win Rate Comparison of Online Alignment Algorithms

    empirical result

    Evaluation of online alignment methods on Reddit TL;DR summarization using a 770M parameter T5X-Large policy model and 3B parameter T5X-XL reward/preference models. Models were evaluated using PaLM 2 side-by-side preference win rates across 9 comparisons (3×33 \times 3 across 3 independent seeds, 2000 evaluation prompts per pair from XSum):

    p(y≻y′)p(y \succ y') IPO IPO-MD DPO Nash-MD-PG SLiC RL
    IPO 0.500 0.515 (0.024) 0.608 (0.038) 0.621 (0.030) 0.608 (0.025) 0.791 (0.012)
    IPO-MD 0.485 (0.024) 0.500 0.600 (0.028) 0.608 (0.026) 0.594 (0.020) 0.778 (0.004)
    DPO 0.392 (0.038) 0.400 (0.028) 0.500 0.520 (0.041) 0.493 (0.040) 0.727 (0.020)
    Nash-MD-PG 0.379 (0.030) 0.392 (0.026) 0.480 (0.041) 0.500 0.479 (0.029) 0.729 (0.020)
    SLiC 0.392 (0.025) 0.406 (0.020) 0.507 (0.040) 0.521 (0.029) 0.500 0.728 (0.010)
    RL 0.209 (0.012) 0.222 (0.004) 0.273 (0.020) 0.271 (0.020) 0.272 (0.010) 0.500

    Each entry is the mean preference score (standard deviation in parentheses) of the row algorithm over the column algorithm. Online IPO and IPO-MD achieve statistically indistinguishable performance (51.5% vs 48.5%), and both consistently outperform DPO (60.8% and 60.0% win rate), Nash-MD-PG (62.1% and 60.8%), SLiC (60.8% and 59.4%), and regularised RL (79.1% and 77.8%).

  9. Knowl 9 — Performance Comparison and Computational Overhead of Online versus Offline Alignment

    empirical result

    Comparing online versus offline variants of IPO and DPO on TL;DR summarization shows large performance advantages for online algorithms over offline training and the initial supervised fine-tuned (SFT) baseline, at the cost of increased compute:

    p(y≻y′)p(y \succ y') IPO DPO Offline-IPO Offline-DPO RLHF SFT
    IPO 0.500 0.608 (0.038) 0.972 (0.008) 0.962 (0.008) 0.791 (0.012) 0.996 (0.001)
    DPO 0.392 (0.038) 0.500 0.958 (0.007) 0.944 (0.012) 0.727 (0.020) 0.995 (0.001)
    Offline-IPO 0.028 (0.008) 0.042 (0.007) 0.500 0.459 (0.042) 0.096 (0.011) 0.840 (0.013)
    Offline-DPO 0.038 (0.008) 0.056 (0.012) 0.541 (0.042) 0.500 0.120 (0.019) 0.869 (0.017)
    RLHF 0.209 (0.012) 0.273 (0.020) 0.904 (0.011) 0.880 (0.019) 0.500 0.988 (0.000)
    SFT 0.004 (0.001) 0.005 (0.001) 0.160 (0.013) 0.131 (0.017) 0.012 (0.000) 0.500

    Online IPO and DPO defeat their offline counterparts with win rates above 94%. However, online methods require generating sequences during training (runtime inference), which acts as a computational bottleneck and causes a roughly 3×3\times end-to-end slowdown (0.25 steps/second on a 4×44 \times 4 TPU v5e configuration, taking ~24 hours per 20,000 steps) compared to offline training.

  10. Knowl 10 — Hyperparameter Settings, KL Divergence, and Regularization Effects in Online Preference Optimization

    empirical result

    Empirical evaluation on summarization reveals differing regularization sensitivities and optimal operating regimes across alignment algorithms:

    1. Regularization sweep ( au\ au): For small regularisation temperatures ( au≤0.5\ au \le 0.5), online IPO and online DPO exhibit similar win rates against RLHF. As τ\tau increases, IPO's win rate decays significantly faster than DPO's, confirming that IPO exerts a stronger regularisation penalty than DPO.
    2. Training steps: Higher regularization in Online IPO requires more training steps to reach peak win rate against RLHF.
    3. Optimal hyperparameters and trajectory KL divergence (measured as KL(π∥πref)\text{KL}(\pi \parallel \pi_{\text{ref}}) over 11,305 XSum validation prompts):
    • Regularised RL: τ=0.05\tau = 0.05, lr=10−4\text{lr} = 10^{-4}, 10,000 steps, KL=25.28\text{KL} = 25.28
    • Online IPO: τ=1.0\tau = 1.0, lr=10−4\text{lr} = 10^{-4}, 28,000 steps, KL=79.04\text{KL} = 79.04
    • Online DPO: τ=5.0\tau = 5.0, lr=10−4\text{lr} = 10^{-4}, 10,000 steps, KL=68.26\text{KL} = 68.26
    • Online SLiC: τ=10.0\tau = 10.0, lr=10−4\text{lr} = 10^{-4}, 30,000 steps, KL=79.32\text{KL} = 79.32
    • IPO-MD: τ=1.0\tau = 1.0, lr=10−4\text{lr} = 10^{-4}, β=0.125\beta = 0.125, 28,000 steps, KL=80.70\text{KL} = 80.70
    • Nash-MD-PG: τ=0.008\tau = 0.008, lr=3⋅10−5\text{lr} = 3 \cdot 10^{-5}, β=0.125\beta = 0.125, 20,000 steps, KL=57.20\text{KL} = 57.20

Coverage note — None was omitted; the extracted knowls comprehensively cover the paper's theoretical characterizations of Online IPO, IPO-MD, and Online DPO, the algorithms, variance analysis, and all empirical benchmarks.

References

  1. 1.Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mane, D. Concrete problems in AI safety. ´ arXiv, 2016.
  2. 2.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., D´ıaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y. PaLM 2 technical report, 2023.
  3. 3.Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. A general theoretical paradigm to understand learning from human preferences. arXiv, 2023.
  4. 4.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv, 2022a.
  5. 5.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukoiut̄ e, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, ȩ N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T. J., Hume, T., Bowman, S., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T. B., and Kaplan, J. Constitutional AI: Harmlessness from AI feedback. arXiv, 2022b.
  6. 6.Boyd, S. P. and Vandenberghe, L. Convex Optimization. Cambridge University Press, 2004.
  7. 7.Bradley, R. A. and Terry, M. E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  8. 8.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217, 2023.
  9. 9.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, 2017.
  10. 10.Coste, T., Anwar, U., Kirk, R., and Krueger, D. S. Reward model ensembles help mitigate overoptimization. arXiv, 2023.
  11. 11.Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. RAFT: Reward rAnked FineTuning for generative foundation model alignment. arXiv, 2023.
  12. 12.Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S. R., Ramachandran, D., Shaw, P., and Berant, J. Helping or herding? Rward model ensembles mitigate but do not eliminate reward hacking. arXiv, 2023.
  13. 13.Fujimoto, S., Meger, D., and Precup, D. Off-policy deep reinforcement learning without exploration. In Proceedings of the International Conference on Machine Learning, 2019.
  14. 14.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In Proceedings of the International Conference on Machine Learning, 2022.
  15. 15.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokra, S., Fernando, N., Wu, B., Foley, R., ´ Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. Improving alignment of dialogue agents via targeted human judgements. arXiv, 2022.
  16. 16.Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L. Policy shaping: Integrating human feedback with reinforcement learning. In Advances in Neural Information Processing Systems, 2013.
  17. 17.Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. Contrastive prefence learning: Learning from human feedback without rl. arXiv preprint arXiv:2310.13639, 2023.
  18. 18.Ivison, H., Wang, Y., Pyatkin, V., Lambert, N., Peters, M., Dasigi, P., Jang, J., Wadden, D., Smith, N. A., Beltagy, I., and Hajishirzi, H. Camels in a changing climate: Enhancing LM adaptation with Tulu 2. arXiv, 2023.
  19. 19.Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv, 2019.
  20. 20.Jouppi, N. P., Kurian, G., Li, S., Ma, P. C., Nagarajan, R., Nai, L., Patil, N., Subramanian, S., Swing, A., Towles, B., Young, C., Zhou, X., Zhou, Z., and Patterson, D. A. TPU v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. In Proceedings of the Annual International Symposium on Computer Architecture, 2023.
  21. 21.Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R. Understanding the effects of rlhf on llm generalisation and diversity. arXiv preprint arXiv:2310.06452, 2023.
  22. 22.Knox, W. B. and Stone, P. TAMER: Training an agent manually via evaluative reinforcement. In Proceedings of the IEEE International Conference on Development and Learning, 2008.
  23. 23.Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A. RLAIF: Scaling reinforcement learning from human feedback with AI feedback. arXiv, 2023.
  24. 24.Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J. Statistical rejection sampling improves preference optimization. arXiv, 2023.
  25. 25.Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. Asynchronous methods for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning, 2016.
  26. 26.Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, D., Tang, Y., Geist, M., Mesnard, T., Michi, A., Selvi, M., Girgin, S., Momchev, N., Bachem, O., Mankowitz, D. J., Precup, D., and Piot, B. Nash learning from human feedback. arXiv, 2023.
  27. 27.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. WebGPT: Browser-assisted question-answering with human feedback. arXiv, 2021.
  28. 28.OpenAI. Introducing ChatGPT, 2022. URL https://openai.com/blog/chatgpt.
  29. 29.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. arXiv, 2022.
  30. 30.Pan, A., Bhatia, K., and Steinhardt, J. The effects of reward misspecification: Mapping and mitigating misaligned models. arXiv, abs/2201.03544, 2022. URL https://api.semanticscholar.org/CorpusID:245837268.
  31. 31.Pang, R. Y., Padmakumar, V., Sellam, T., Parikh, A. P., and He, H. Reward gaming in conditional text generation. In Annual Meeting of the Association for Computational Linguistics, 2022. URL https://api.semanticscholar.org/CorpusID:253553557.
  32. 32.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, 2023.
  33. 33.Ramé, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, ´ G., Bachem, O., and Ferret, J. WARM: On the benefits of weight averaged reward models. arXiv, 2024.
  34. 34.Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A. Scaling up models and data with t5x and seqio. arXiv, 2022.
  35. 35.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv, 2017.
  36. 36.Shashi, N., Cohen, S. B., and Mirella, L. Don’t Give Me the Details, Just the Summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2018.
  37. 37.Shazeer, N. M. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. arXiv, 2018.
  38. 38.Shin, D., Dragan, A. D., and Brown, D. S. Benchmarks and algorithms for offline preference-based reward learning. arXiv, 2023.
  39. 39.Singhal, P., Goyal, T., Xu, J., and Durrett, G. A long way to go: Investigating length correlations in rlhf. ArXiv, abs/2310.03716, 2023.
  40. 40.Skalse, J., Howe, N. H. R., Krasheninnikov, D., and Krueger, D. Defining and characterizing reward gaming. In Neural Information Processing Systems, 2022.
  41. 41.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, 2020.
  42. 42.Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A. A minimaximalist approach to reinforcement learning from human feedback. arXiv, 2024.
  43. 43.Touvron, H., Martin, L., Stone, K. R., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D. M., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A. S., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I. M., Korenev, A. V., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. arXiv, 2023.
  44. 44.Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T. Zephyr: Direct distillation of LM alignment. arXiv, 2023.
  45. 45.Völske, M., Potthast, M., Syed, S., and Stein, B. TL;DR: ¨ Mining Reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization. Association for Computational Linguistics, 2017.
  46. 46.Wang, C., Jiang, Y., Yang, C., Liu, H., and Chen, Y. Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints. arXiv, 2023.
  47. 47.Warnell, G., Waytowich, N., Lawhern, V., and Stone, P. Deep TAMER: Interactive agent shaping in high-dimensional state spaces. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018.
  48. 48.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Proceedings of the International Conference on Machine Learning, 2022.
  49. 49.Wu, Y., Tucker, G., and Nachum, O. Behavior regularized offline reinforcement learning. arXiv, 2019.
  50. 50.Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. Self-rewarding language models, 2024.
  51. 51.Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. Rrhf: Rank responses to align language models with human feedback without tears. arXiv, abs/2304.05302, 2023.
  52. 52.Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J. SLiC-HF: Sequence likelihood calibration with human feedback. arXiv, 2023.
  53. 53.Zhuang, S. and Hadfield-Menell, D. Consequences of misaligned AI. In Advances in Neural Information Processing Systems, 2020.

Citation

MLA
Calandriello, D., et al. “Human Alignment of Large Language Models Through Online Preference Optimisation”. arXiv, 2024, http://arxiv.org/abs/2403.08635v1.
APA
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., Joshi, R., Zheng, Z., & Piot, B. (2024). Human Alignment of Large Language Models through Online Preference Optimisation. arXiv. http://arxiv.org/abs/2403.08635v1
Chicago
Calandriello, D., D. Guo, R. Munos, et al. 2024. “Human Alignment of Large Language Models Through Online Preference Optimisation”. arXiv. http://arxiv.org/abs/2403.08635v1.
Harvard
Calandriello, D. et al. (2024) “Human Alignment of Large Language Models through Online Preference Optimisation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.08635v1.
Vancouver
1. Calandriello D, Guo D, Munos R, et al (2024) Human Alignment of Large Language Models through Online Preference Optimisation. arXiv

BibTeX

@article{calandriello2024human,
  title = {Human Alignment of Large Language Models through Online Preference Optimisation},
  author = {Calandriello, Daniele and Guo, Daniel and Munos, Remi and Rowland, Mark and Tang, Yunhao and Pires, Bernardo Avila and Richemond, Pierre Harvey and Lan, Charline Le and Valko, Michal and Liu, Tianqi and Joshi, Rishabh and Zheng, Zeyu and Piot, Bilal},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.08635v1},
  eprint = {2403.08635}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/