Generalized Preference Optimization: A Unified Approach to Offline Alignment

Yunhao TangZhaohan Daniel GuoZeyu ZhengDaniele CalandrielloRémi MunosMark RowlandPierre Harvey RichemondMichal ValkoBernardo Ávila PiresBilal Piot

article2024ICML173 citations

Unifies offline alignment methods under a single convex-loss framework that explains how algorithms like DPO and IPO enforce implicit regularization and provides principled guidance for tuning their hyperparameters.

Listen

Aligning large language models with human preferences has traditionally relied on complex reinforcement learning workflows that require training separate reward models and repeatedly generating online model outputs. While newer offline preference optimization techniques streamline this process by training directly on pre-collected pairwise comparison datasets, the machine learning community has lacked a comprehensive theoretical framework explaining how these distinct algorithms relate to one another and how they control model drift. The article addresses this gap by introducing Generalized Preference Optimization, a unifying framework that maps offline alignment methods to established binary classification loss functions.

To evaluate this framework, the article proves that existing algorithms—such as Direct Preference Optimization, Identity Preference Optimization, and Sequence Likelihood Calibration—are specific instances within a broader family of convex losses, while introducing novel loss variants including exponential, truncated quadratic, and Savage losses. The authors conducted mathematical analyses of the regularization mechanisms and paired their theory with empirical testing. These tests included controlled synthetic experiments on an 11-billion-parameter text summarization setup, low-dimensional mathematical simulations, and large language model evaluations judged across 2,000 test summaries.

Across both theoretical analysis and empirical benchmarking, the article establishes four primary findings. First, all examined loss variants demonstrate remarkably similar peak performance and follow the same fundamental trade-offs between response quality and baseline model drift. Second, the mathematical shape of the chosen loss function directly determines its baseline regularization strength; squared and margin-based losses enforce tighter constraints, requiring regularization hyperparameter settings roughly an order of magnitude lower than logistic losses to achieve peak performance. Third, the implicit regularization enforced on offline data approximates online distribution constraints well near the starting model, but diverges significantly when the policy drifts too far. Fourth, in practical summarization tasks, all algorithmic variants reached their highest win rates against the baseline when the regularization parameter was set between 0.1 and 1.

These findings indicate that practitioners do not need to cycle through competing offline alignment algorithms seeking major performance breakthroughs. Instead, model performance depends almost entirely on calibrating the regularization hyperparameter to match the inherent strength of the selected loss function. Engineering teams can select the algorithm that is simplest and most numerically stable for their software stack without sacrificing output quality, reducing developmental complexity and tuning costs.

Teams implementing offline preference optimization should prioritize hyperparameter tuning and early checkpoint selection over switching loss functions. When transitioning between algorithms, engineers must adjust their regularization values according to the specific loss curve being used. Because this framework relies on the assumption that human feedback can be represented as a single scalar ranking, practitioners should exercise caution in domains with complex, intransitive human preferences. Additional research and pilot testing will be necessary to extend these unified methods to multi-turn interactions and non-ranking alignment objectives.

arXiv: 2402.05749
Cover for Generalized Preference Optimization: A Unified Approach to Offline Alignment

Abstract

Offline preference optimization allows fine-tuning large models directly from offline data, and has proved effective in recent alignment practices. We propose generalized preference optimization (GPO), a family of offline losses parameterized by a general class of convex functions. GPO enables a unified view over preference optimization, encompassing existing algorithms such as DPO, IPO and SLiC as special cases, while naturally introducing new variants. The GPO framework also sheds light on how offline algorithms enforce regularization, through the design of the convex function that defines the loss. Our analysis and experiments reveal the connections and subtle differences between the offline regularization and the KL divergence regularization intended by the canonical RLHF formulation. In a controlled setting akin to Gao et al. (2023), we also show that different GPO variants achieve similar trade-offs between regularization and performance, though the optimal values of hyper-parameter might differ as predicted by theory. In all, our results present new algorithmic toolkits and empirical insights to alignment practitioners. Please see https://arxiv.org/pdf/2402.05749 for the full version of the paper.

Table of Contents

  • 1. Introduction
  • 2. A general family of offline preference optimization losses
  • 2.1. A recipe to derive preference optimization losses
  • 2.2. GPO: A generalized family of offline preference optimization algorithms
  • 3. Reward modeling viewed as a binary classification problem
  • 3.1. Characterizing what the reward model learns
  • 4. Understanding regularization in offline preference optimization
  • 4.1. How do offline losses enforce regularization
  • 4.2. Offline regularization vs. KL regularization
  • 4.3. Analyzing a language modeling example
  • 5. Empirical study of GPO variants
  • 5.1. Trade-offs between KL divergence and performance
  • 5.2. Model-based side by side evaluation
  • 6. Discussions and conclusion
  • References
  • A. Experiment details and additional results
  • A.1. Bandit experiment
  • A.2. A mixture of Gaussian counterexample
  • A.3. Language modeling experiments
  • A.4. Trade-off between performance and KL divergence
  • A.5. Model-based side by side evaluation
  • B. Proof and derivations of theoretical results
  • B.1. Derivation of the gradient of KL divergence and µ-weighted squared loss
  • B.2. Discussion on Taylor expansions of the GPO losses
  • C. Discussion on Bayes consistency for the learned reward model
  • C.1. Discussion of pairwise preference model

Knowls

  1. Knowl 1 — Generalized preference optimization loss

    model/method

    For a prompt-conditioned policy πθ\pi_\theta and a reference policy πref\pi_{\mathrm{ref}}, let (yw,yl)(y_w,y_l) be an offline preference pair drawn from the data distribution μ\mu, with ywy_w preferred to yly_l. Define the pairwise log-ratio margin

    ρθ(yw,yl)=log⁡πθ(yw)πref(yw)−log⁡πθ(yl)πref(yl).\rho_\theta(y_w,y_l)=\log\frac{\pi_\theta(y_w)}{\pi_{\mathrm{ref}}(y_w)}-\log\frac{\pi_\theta(y_l)}{\pi_{\mathrm{ref}}(y_l)}.

    Generalized preference optimization (GPO) minimizes

    LGPO(θ)=E(yw,yl)∼μ[f ⁣(βρθ(yw,yl))],\mathcal{L}_{\mathrm{GPO}}(\theta)=\mathbb{E}_{(y_w,y_l)\sim\mu}\left[f\!\left(\beta\rho_\theta(y_w,y_l)\right)\right],

    where f:R→Rf:\mathbb{R}\to\mathbb{R} is a convex scalar loss and β>0\beta>0 is a scale or regularization hyperparameter. The formulation is intended for convex losses that decrease locally at zero, f′(0)<0f'(0)<0, so that increasing the preferred-over-dispreferred margin lowers the loss near the reference policy. GPO does not require pairwise preferences to follow a Bradley–Terry model.

  2. Knowl 2 — Existing and new losses as GPO instances

    model/method

    The GPO objective yields familiar offline preference-optimization methods by choosing the convex loss ff of the scaled margin βρθ\beta\rho_\theta, where ρθ\rho_\theta is the difference between the policy-to-reference log ratios for the preferred and dispreferred responses:

    • Logistic loss, f(z)=log⁡(1+e−z)f(z)=\log(1+e^{-z}), gives DPO.
    • Squared loss, f(z)=(z−1)2f(z)=(z-1)^2, gives IPO.
    • Hinge loss, f(z)=max⁡(0,1−z)f(z)=\max(0,1-z), gives the contrastive preference loss used by SLiC.
    • Exponential loss, f(z)=e−zf(z)=e^{-z}, gives an additional GPO variant.
    • Truncated quadratic loss, f(z)=max⁡(0,1−z)2f(z)=\max(0,1-z)^2, gives another variant.
    • Savage loss, f(z)=(1+ez)−2f(z)=(1+e^z)^{-2}, gives another variant.

    The last three losses are proposed as offline preference-optimization choices within the framework; the paper does not identify existing offline algorithms corresponding to them. The SLiC mapping describes its contrastive loss and excludes SLiC’s additional supervised-learning component.

  3. Knowl 3 — Preference learning as binary classification

    model/method

    Pairwise reward learning can be written as supervised binary classification. For a pair of responses (y1,y2)(y_1,y_2), assign label ℓ=1\ell=1 when y1y_1 is preferred and ℓ=−1\ell=-1 otherwise. A pointwise reward model rϕ(y)r_\phi(y) predicts the pair’s signed margin through rϕ(y1)−rϕ(y2)r_\phi(y_1)-r_\phi(y_2), and a convex classification surrogate ff yields the loss f(ℓ[rϕ(y1)−rϕ(y2)])f(\ell[r_\phi(y_1)-r_\phi(y_2)]). After orienting each observed pair as preferred and dispreferred, this becomes E(yw,yl)∼μ[f(rϕ(yw)−rϕ(yl))]\mathbb{E}_{(y_w,y_l)\sim\mu}[f(r_\phi(y_w)-r_\phi(y_l))].

    The paper’s key construction replaces the reward difference in this supervised loss with the policy-to-reference log-ratio difference, scaled by β\beta, to obtain the GPO objective. This binary-classification interpretation motivates the use of different convex surrogate losses rather than requiring one particular preference-probability model.

  4. Knowl 4 — Equivalence of global offline and regularized-policy optima

    theoretical result

    Under the global-optimization equivalence stated in the paper, a global minimizer of the offline preference loss

    E(yw,yl)∼μ[f ⁣(β[log⁡πθ(yw)πref(yw)−log⁡πθ(yl)πref(yl)])]\mathbb{E}_{(y_w,y_l)\sim\mu}\left[f\!\left(\beta\left[\log\frac{\pi_\theta(y_w)}{\pi_{\mathrm{ref}}(y_w)}-\log\frac{\pi_\theta(y_l)}{\pi_{\mathrm{ref}}(y_l)}\right]\right)\right]

    is the same policy as the optimal solution to the regularized reward objective

    max⁡π  Ey∼π[r(y)]−β KL(π∥πref),\max_{\pi}\;\mathbb{E}_{y\sim\pi}[r(y)]-\beta\,\mathrm{KL}(\pi\|\pi_{\mathrm{ref}}),

    for a reward function rr that globally minimizes the corresponding pairwise reward loss E(yw,yl)∼μ[f(r(yw)−r(yl))]\mathbb{E}_{(y_w,y_l)\sim\mu}[f(r(y_w)-r(y_l))]. Here KL(π∥πref)=Ey∼π[log⁡(π(y)/πref(y))]\mathrm{KL}(\pi\|\pi_{\mathrm{ref}})=\mathbb{E}_{y\sim\pi}[\log(\pi(y)/\pi_{\mathrm{ref}}(y))]. This is a statement about matching global optima; it does not assert that local optimization of an offline loss follows the KL-regularized objective’s learning dynamics.

  5. Knowl 5 — Local GPO losses induce an offline squared-margin regularizer

    theoretical result

    Suppose ff is twice differentiable near zero, convex, and satisfies f′(0)<0f'(0)<0 and f′′(0)>0f''(0)>0. When the pairwise margins ρθ\rho_\theta are small, as expected near initialization at πθ=πref\pi_\theta=\pi_{\mathrm{ref}}, the GPO objective has the approximation

    Eμ[f(βρθ)]≈f(0)+f′(0)β Eμ[ρθ]+f′′(0)β22 Eμ[ρθ2].\mathbb{E}_{\mu}[f(\beta\rho_\theta)]\approx f(0)+f'(0)\beta\,\mathbb{E}_{\mu}[\rho_\theta]+\frac{f''(0)\beta^2}{2}\,\mathbb{E}_{\mu}[\rho_\theta^2].

    The first-order term favors larger margins on preferred pairs; the second-order term penalizes pairwise deviations between the policy and reference log ratios. The latter is an offline, data-weighted regularizer, not a KL divergence. Up to an overall positive scale and additive constants, this local objective is equivalent to an IPO-style squared loss with effective hyperparameter βeff=f′′(0)∣f′(0)∣β\beta_{\mathrm{eff}}=\frac{f''(0)}{|f'(0)|}\beta. Thus, different losses can require different values of β\beta to produce comparable local regularization.

  6. Knowl 6 — Loss-tail behavior controls response-level restraint

    theoretical result

    For a GPO margin ρθ\rho_\theta that has become large and positive, the behavior of the loss derivative affects how strongly optimization continues to increase the preferred response’s probability relative to the reference policy. With logistic, exponential, and Savage losses, the derivative remains negative at finite positive margins, although its magnitude can become small; optimization can therefore keep increasing a winning response’s relative probability, with a weakening gradient. With squared loss, hinge loss, and truncated quadratic loss, the derivative can become zero or point back toward a smaller margin after the loss’s target or threshold is passed. These losses can consequently restrain individual response-level deviations more strongly at a fixed β\beta.

    This tail-based comparison is distinct from the local Taylor approximation: it concerns behavior at large margins. It also explains why the paper recommends tuning β\beta for the chosen loss rather than assuming that the same numerical value imposes the same regularization across GPO variants.

  7. Knowl 7 — The KL gradient equals an on-policy pairwise squared-log-ratio gradient

    theoretical result

    Let sθ(y)=log⁡(πθ(y)/πref(y))s_\theta(y)=\log(\pi_\theta(y)/\pi_{\mathrm{ref}}(y)), and define the squared pairwise log-ratio penalty under a response distribution ν\nu by

    Rν(θ)=12Ey1,y2∼i.i.d.ν[(sθ(y1)−sθ(y2))2].R_\nu(\theta)=\frac{1}{2}\mathbb{E}_{y_1,y_2\stackrel{\mathrm{i.i.d.}}{\sim}\nu}\left[(s_\theta(y_1)-s_\theta(y_2))^2\right].

    For differentiable policies and finite expectations, the paper establishes

    ∇θKL(πθ∥πref)=∇θRπθ(θ).\nabla_\theta\mathrm{KL}(\pi_\theta\|\pi_{\mathrm{ref}})=\nabla_\theta R_{\pi_\theta}(\theta).

    Thus the KL gradient corresponds to the gradient of a pairwise squared discrepancy evaluated using on-policy samples from πθ\pi_\theta. Offline preference optimization instead encounters such discrepancies under its offline data distribution μ\mu. The identity equates gradients in the on-policy case; it does not make the offline and KL regularizers identical.

  8. Knowl 8 — Offline squared regularization need not decrease KL locally

    theoretical result

    The offline squared-margin penalty and KL divergence share the reference policy as a global minimizer, but descending the offline penalty need not decrease KL at every point. The paper demonstrates this numerically with a one-dimensional shifted-mixture example. The offline distribution is μ=13∑i=13N(ui,0.052)\mu=\frac{1}{3}\sum_{i=1}^3\mathcal{N}(u_i,0.05^2), where the uiu_i are independently sampled uniformly from [−1,1][-1,1]. The reference distribution is πref=0.3N(−0.8,0.12)+0.4N(0,0.12)+0.3N(0.8,0.12)\pi_{\mathrm{ref}}=0.3\mathcal{N}(-0.8,0.1^2)+0.4\mathcal{N}(0,0.1^2)+0.3\mathcal{N}(0.8,0.1^2), and the policy πθ\pi_\theta is a constant shift of that reference distribution, with shift parameter c∈[−1,1]c\in[-1,1].

    Using 2,000 samples to estimate the losses, the paper finds local minima of the μ\mu-weighted squared discrepancy at shifts away from c=0c=0, whereas the reference policy occurs at c=0c=0. Consequently, local descent of the offline penalty can fail to reduce KL or reach its global minimum. This is a numerical counterexample, not a claim that such misalignment occurs for every offline distribution.

  9. Knowl 9 — Language-model measurements show only conditional alignment between offline penalty and KL

    empirical result

    In summarization experiments using a human-feedback dataset and 700-million-parameter T5X models, the authors initialized each learner from the reference policy and tracked KL divergence alongside the offline μ\mu-weighted squared-margin loss during GPO training. Experiments fixed a loss variant and β\beta; measurements were made every 2,000 steps, with additional measurements every 200 steps during the first 2,000 steps. The trajectories plotted on page 7 show that the two quantities tend to track one another more closely when stronger regularization keeps the policy near its initialization. With weaker regularization, the offline penalty can grow while its relationship to KL becomes less reliable.

    Within the plotted regime where the quantities were better correlated, KL varied by roughly 10−0.510^{-0.5} to 101.510^{1.5}, while the offline squared penalty varied by roughly 10−1.510^{-1.5} to 100.510^{0.5}. This larger KL variation for a smaller change in the offline penalty illustrates why controlling the offline loss may not tightly control KL. Across GPO variants, the page 7 trajectories were broadly similar, with most alternatives appearing slightly more stable than the logistic variant.

  10. Knowl 10 — Controlled evaluation finds similar performance–KL trade-offs across GPO losses

    empirical result

    For a controlled summarization comparison, the authors trained an 11-billion-parameter preference model to relabel the offline data and used a policy trained with that preference model as the evaluation baseline. Offline learners were 110-million-parameter T5X models. For each GPO loss, they swept β\beta, trained for 20,000 steps at constant learning rates of 10−510^{-5} and 3×10−53\times10^{-5}, and evaluated checkpoints every 2,000 steps against the baseline.

    As shown by the trade-off trajectories on pages 8–9, golden win rate initially improves as KL divergence from the supervised fine-tuned initialization grows, then falls as the policy over-optimizes. The overall performance-versus-KL curves and peak performance were similar across loss variants, and the variants reached their peaks at broadly similar KL levels. Their best β\beta values differed: squared and truncated quadratic losses generally peaked around β≈1\beta\approx1, while the other variants generally peaked at larger values around β≈10\beta\approx10. The authors caution that some observed peak-performance differences may not be statistically significant.

  11. Knowl 11 — Model-judged summarization results are similar across GPO variants after tuning

    empirical result

    On the open-source summarization dataset with human preference labels, the authors trained 700-million-parameter T5X models using a constant learning rate of 10−510^{-5}, batch size 32, Adafactor with decay rate 0.8, and 200,000 training steps. A prompted PaLM-2 model compared summaries on 2,000 evaluation examples, judging each learner against the supervised fine-tuned reference policy. The model-based comparison described on page 9, and the win-rate plot on page 17, show that the different GPO loss variants achieved broadly similar peak performance after tuning β\beta.

    Across variants, the strongest results were generally obtained for β∈[0.1,1]\beta\in[0.1,1]; performance declined more quickly when β\beta was too large. The reported finding is that choosing an appropriate regularization coefficient mattered more than choosing among these loss variants under this evaluation setup.

Coverage note — The paper’s Bayes-consistency theorem and special-case analytic characterizations of learned pointwise rewards were omitted because they support the classification framing but are secondary to the GPO formulation, regularization analysis, and empirical comparisons; proof derivations and auxiliary diagnostics are also omitted.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. PaLM 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. A general theoretical paradigm to understand learning from human preferences. In Proceedings of the International Conference on Artificial Intelligence and Statistics, 2024.
  4. 4.Peter L. Bartlett, Michael I. Jordan, and Jon D. McAuliffe. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101(473):138–156, 2006.
  5. 5.Bernhard E. Boser, Isabelle M. Guyon, and Vladimir N. Vapnik. A training algorithm for optimal margin classifiers. In Proceedings of the Workshop on Computational Learning Theory, 1992.
  6. 6.Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
  7. 7.Daniele Calandriello, Daniel Guo, Rémi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, and Bilal Piot. Human alignment of large language models through online preference optimisation. In Proceedings of the International Conference on Machine Learning, 2024.
  8. 8.Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, 2017.
  9. 9.Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine learning, 20:273–297, 1995.
  10. 10.Yoav Freund and Robert E. Schapire. A desicion-theoretic generalization of on-line learning and an application to boosting. In Proceedings of the European conference on Computational Learning Theory, 1995.
  11. 11.Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the International Conference on Machine Learning, 2023.
  12. 12.Trevor Hastie, Robert Tibshirani, and Jerome H. Friedman. The elements of statistical learning: Data mining, inference, and prediction. Springer, 2009.
  13. 13.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  14. 14.Hamed Masnadi-Shirazi and Nuno Vasconcelos. On the design of loss functions for classification: Yheory, robustness to outliers, and SavageBoost. In Advances in Neural Information Processing Systems, 2008.
  15. 15.Hamed Masnadi-Shirazi and Nuno Vasconcelos. A view of margin losses as regularizers of probability estimates. Journal of Machine Learning Research, 16(1):2751–2795, 2015.
  16. 16.Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, and Bilal Piot. Nash learning from human feedback. In Proceedings of the International Conference on Machine Learning, 2024.
  17. 17.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, 2022.
  18. 18.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, 2023.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020.
  20. 20.Lorenz Richter, Ayman Boustati, Nikolas Nusken, Francisco Ruiz, and Omer Deniz Akyildiz. VarGrad: A low-variance gradient estimator for variational inference. In Advances in Neural Information Processing Systems, 2020.
  21. 21.Adam Roberts, Hyung Won Chung, Gaurav Mishra, Anselm Levskaya, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, et al. Scaling up models and data with t5x and seqio. Journal of Machine Learning Research, 24(377):1–8, 2023.
  22. 22.Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, et al. Factually consistent summarization via reinforcement learning with textual entailment feedback. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2023.
  23. 23.Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, and Alessandro Verri. Are loss functions all the same? Neural computation, 16(5):1063–1076, 2004.
  24. 24.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In Proceedings of the International Conference on Machine Learning, 2018.
  25. 25.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems, 2020.
  26. 26.Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal. A minimaximalist approach to reinforcement learning from human feedback. arXiv preprint arXiv:2401.04056, 2024.
  27. 27.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  28. 28.Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu. SLiC-HF: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425, 2023.

Citation

MLA
Tang, Y., et al. “Generalized Preference Optimization: A Unified Approach to Offline Alignment”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.05749.
APA
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. Á., & Piot, B. (2024). Generalized Preference Optimization: A Unified Approach to Offline Alignment. arXiv. https://doi.org/10.48550/arxiv.2402.05749
Chicago
Tang, Y., Z. D. Guo, Z. Zheng, et al. 2024. “Generalized Preference Optimization: A Unified Approach to Offline Alignment”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.05749.
Harvard
Tang, Y. et al. (2024) “Generalized Preference Optimization: A Unified Approach to Offline Alignment”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.05749.
Vancouver
1. Tang Y, Guo ZD, Zheng Z, Calandriello D, Munos R, Rowland M, Richemond PH, Valko M, Pires BÁ, Piot B (2024) Generalized Preference Optimization: A Unified Approach to Offline Alignment. https://doi.org/10.48550/arxiv.2402.05749

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.05749,
  doi = {10.48550/ARXIV.2402.05749},
  url = {https://arxiv.org/abs/2402.05749},
  author = {Tang, Yunhao and Guo, Zhaohan Daniel and Zheng, Zeyu and Calandriello, Daniele and Munos, Rémi and Rowland, Mark and Richemond, Pierre Harvey and Valko, Michal and Pires, Bernardo Ávila and Piot, Bilal},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Generalized Preference Optimization: A Unified Approach to Offline Alignment},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/