SimPO: Simple Preference Optimization with a Reference-Free Reward

Yu MengMengzhou XiaDanqi Chen

article2024NeurIPS904 citations

Proposes SimPO, a reference-free preference optimization method that uses average sequence log probabilities and a target reward margin to outperform Direct Preference Optimization across standard benchmarks while significantly reducing memory and compute costs during language model alignment.

Listen

Aligning large language models with human preferences is critical for ensuring helpful, safe, and coherent conversational behavior. Traditional alignment relied on complex reinforcement learning pipelines involving separate reward models, but recent industry practice has shifted toward direct preference optimization algorithms. Existing direct approaches, however, rely on a static reference model to anchor probability updates. This dependency incurs substantial memory and compute overhead during training while creating a fundamental mathematical mismatch between the training reward objective and the actual generation metric used during text generation.

The article introduces SimPO (Simple Preference Optimization), an offline preference optimization algorithm designed to eliminate reference models while directly aligning training rewards with generation probabilities. The study evaluates whether formulating rewards as length-normalized average log probabilities and introducing an explicit target reward margin can improve conversational quality and computational efficiency across multiple model families and benchmarks.

The authors conducted comprehensive experiments across four core model configurations using 7-billion and 8-billion parameter models from the Mistral, Llama 3, and Gemma 2 model families, testing both base and instruction-tuned versions. Alignment was evaluated using standard open-ended conversational benchmarks—such as AlpacaEval 2, Arena-Hard, and MT-Bench—as well as broader task suites measuring reasoning, truthfulness, and knowledge retention. The team also conducted ablation analyses isolating the impact of length normalization, target reward margins, training hyperparameters, and computing resource usage on an eight-GPU hardware setup.

The evaluation yielded several key findings. First, SimPO consistently outperformed direct preference optimization and multiple variants across benchmarks, achieving improvements of up to 6.4 percentage points on AlpacaEval 2 and up to 7.5 points on Arena-Hard. Second, length normalization proved essential; removing it caused severe length exploitation and repetitive text generation, whereas normalized models retained quality with concise outputs. Third, enforcing an explicit target reward margin widened separation between preferred and rejected outputs, directly improving validation accuracy. Fourth, eliminating the reference model reduced overall training runtime by approximately 20% and lowered peak GPU memory usage by roughly 10%. Finally, the resulting Gemma-2-9B model trained with SimPO achieved a 72.4% length-controlled win rate on AlpacaEval 2 and placed first among all sub-10-billion parameter models on the real-world crowdsourced Chatbot Arena leaderboard.

These findings indicate that aligning training objectives directly with inference metrics offers a more effective, cost-efficient path to model alignment than maintaining complex reference-model constraints. Organizations training language models can achieve higher conversational quality at lower cloud computing and infrastructure costs. Furthermore, the approach mitigates the risk of models learning to exploit evaluators through empty verbosity, producing concise and structured outputs instead.

Teams deploying alignment pipelines should consider replacing traditional reference-dependent preference optimization algorithms with length-normalized margin objectives. Practitioners must calibrate the target margin and scaling hyperparameters carefully, as excessively high margins can flatten probability distributions. When tuning strong instruction-tuned checkpoints, engineers should evaluate learning rates carefully to balance conversational gains against potential degradation on specialized tasks, or consider hybrid objectives incorporating supervised loss regularization when task preservation is essential.

The study notes certain limitations, including an empirical drop in reasoning and mathematical benchmarks (such as GSM8K) on certain model families like Llama 3 when using higher learning rates, although this trade-off was not observed with Gemma 2. Additionally, the primary datasets focused predominantly on conversational helpfulness rather than exhaustive safety and honesty filtering. While confidence in conversational performance and training efficiency is high across standard open-source benchmarks, organizations should conduct targeted domain testing before applying this alignment method to strict mathematical and safety-critical production environments.

Cover for SimPO: Simple Preference Optimization with a Reference-Free Reward

Abstract

Direct Preference Optimization (DPO) is a widely used offline preference optimization algorithm that reparameterizes reward functions in reinforcement learning from human feedback (RLHF) to enhance simplicity and training stability. In this work, we propose SimPO, a simpler yet more effective approach. The effectiveness of SimPO is attributed to a key design: using the average log probability of a sequence as the implicit reward. This reward formulation better aligns with model generation and eliminates the need for a reference model, making it more compute and memory efficient. Additionally, we introduce a target reward margin to the Bradley-Terry objective to encourage a larger margin between the winning and losing responses, further improving the algorithm's performance. We compare SimPO to DPO and its latest variants across various state-of-the-art training setups, including both base and instruction-tuned models such as Mistral, Llama 3, and Gemma 2. We evaluate on extensive chat-based evaluation benchmarks, including AlpacaEval 2, MT-Bench, and Arena-Hard. Our results demonstrate that SimPO consistently and significantly outperforms existing approaches without substantially increasing response length. Specifically, SimPO outperforms DPO by up to 6.4 points on AlpacaEval 2 and by up to 7.5 points on Arena-Hard. Our top-performing model, built on Gemma-2-9B-it, achieves a 72.4% length-controlled win rate on AlpacaEval 2, a 59.1% win rate on Arena-Hard, and ranks 1st on Chatbot Arena among <10B models with real user votes.

Table of Contents

  • 1 Introduction
  • 2 SimPO: Simple Preference Optimization
  • 2.1 Background: Direct Preference Optimization (DPO)
  • 2.2 A Simple Reference-Free Reward Aligned with Generation
  • 2.3 The SimPO Objective
  • 3 Experimental Setup
  • 4 Experimental Results
  • 4.1 Main Results and Ablations
  • 4.2 Length Normalization (LN) Prevents Length Exploitation
  • 4.3 The Impact of Target Reward Margin in SimPO
  • 4.4 In-Depth Analysis of DPO vs. SimPO
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Limitations
  • B Implementation Details
  • C Downstream Task Evaluation
  • D Standard Deviation of AlpacaEval 2 and Arena-Hard
  • E Generation Length Analysis
  • F Gradient Analysis
  • G Qualitative Analysis
  • H Llama-3-Instruct v0.2 (Jul 7, 2024)
  • I Applying Length Normalization and Target Reward Margin to DPO (Jul 7, 2024)
  • J Applying SimPO to Gemma 2 Models (Sept 16, 2024)

Knowls

  1. Knowl 1 — SimPO optimizes a length-normalized preference margin

    model/method

    SimPO trains a policy language model directly from preference pairs, without a reference model. For a prompt xx, let ywy_w be the preferred response and yly_l the dispreferred response; let ∣y∣|y| denote the number of response tokens, and let πθ(y∣x)\pi_\theta(y\mid x) be the policy's probability of generating response yy. SimPO uses the average token log probability as its reward basis and minimizes the following loss over preference data D\mathcal{D}:

    LSimPO(πθ)=−E(x,yw,yl)∼D[log⁡σ(β∣yw∣log⁡πθ(yw∣x)−β∣yl∣log⁡πθ(yl∣x)−γ)].\mathcal{L}_{\mathrm{SimPO}}(\pi_\theta) = -\mathbb{E}_{(x,y_w,y_l)\sim\mathcal{D}}\left[\log\sigma\left(\frac{\beta}{|y_w|}\log\pi_\theta(y_w\mid x)-\frac{\beta}{|y_l|}\log\pi_\theta(y_l\mid x)-\gamma\right)\right].

    Here σ(z)=1/(1+e−z)\sigma(z)=1/(1+e^{-z}) is the logistic function, β>0\beta>0 scales the reward difference, and γ>0\gamma>0 is the target margin. The objective encourages the preferred response's scaled average log probability to exceed the dispreferred response's by at least the target margin.

  2. Knowl 2 — SimPO's reward aligns preference training with generation likelihood

    model/method

    At generation time, a model's responses are ranked using a likelihood-based score, which the paper approximates with the average token log probability of a response. SimPO uses that same average log probability as its implicit reward. Averaging over response tokens also avoids the length bias of using summed log probabilities: a longer response would otherwise tend to have a lower sum simply because it contains more tokens, potentially encouraging the model to compensate by inflating its likelihood. Because SimPO computes its reward from the policy alone, it does not require a reference model during training, unlike reference-based preference objectives.

  3. Knowl 3 — SimPO improves chat-benchmark results across four model setups

    data/table

    The table compares the supervised fine-tuned (SFT), DPO, and SimPO models on AlpacaEval 2 length-controlled win rate, Arena-Hard win rate, and the GPT-4 Turbo MT-Bench rating. The four setups use Mistral-7B or Llama-3-8B in base or instruction-tuned form. SimPO exceeds DPO on AlpacaEval 2 and Arena-Hard in every setup; its MT-Bench ratings are similar to DPO's. AlpacaEval 2 and Arena-Hard are pairwise-judged benchmarks, while MT-Bench reports a single-answer rating.

    SetupMethodAlpacaEval 2 LC (%)Arena-Hard WR (%)MT-Bench GPT-4 Turbo
    Mistral-BaseSFT8.41.34.8
    Mistral-BaseDPO15.110.45.9
    Mistral-BaseSimPO21.516.66.0
    Mistral-InstructSFT17.112.66.2
    Mistral-InstructDPO26.816.36.3
    Mistral-InstructSimPO32.121.06.6
    Llama-3-BaseSFT6.23.35.2
    Llama-3-BaseDPO18.215.96.5
    Llama-3-BaseSimPO22.023.46.6
    Llama-3-InstructSFT26.022.36.9
    Llama-3-InstructDPO40.332.67.0
    Llama-3-InstructSimPO44.733.87.0

    The SimPO-minus-DPO gain in AlpacaEval 2 length-controlled win rate ranges from 3.8 to 6.4 percentage points; on Arena-Hard it ranges from 1.2 to 7.5 points. The paper reports that SimPO's best-baseline AlpacaEval 2 gains across these setups are 3.6–4.8 points.

  4. Knowl 4 — Training and evaluation settings used for the main comparisons

    experimental setup

    The experiments cover four starting-model configurations: Mistral-7B and Llama-3-8B, each in a base setup and an instruction-tuned setup. For base models, the authors first supervised fine-tune on UltraChat-200k, then preference-optimize on UltraFeedback. For instruction-tuned models, they start from released instruction checkpoints and construct preference pairs by sampling five responses per UltraFeedback prompt at temperature 0.8, scoring them with PairRM, and selecting the highest- and lowest-scoring responses. Preference optimization uses batch size 128, one epoch, maximum sequence length 2048, and a cosine learning-rate schedule with 10% warmup.

    Setupβ\betaγ\gammaLearning rate
    Mistral-Base2.01.63e-7
    Mistral-Instruct2.50.35e-7
    Llama-3-Base2.01.06e-7
    Llama-3-Instruct2.51.41e-6

    Evaluation uses 805 AlpacaEval 2 prompts, 500 Arena-Hard prompts, and 80 MT-Bench questions. AlpacaEval 2 reports raw and length-controlled win rates, Arena-Hard reports win rate against a baseline, and MT-Bench reports judge ratings.

  5. Knowl 5 — Removing length normalization or the target margin degrades results

    empirical result

    Ablations compare SimPO with DPO, SimPO without length normalization (w/o LN), and SimPO with the target margin set to zero. The table reports AlpacaEval 2 length-controlled win rate, Arena-Hard win rate, and the GPT-4 Turbo MT-Bench rating for Mistral-7B base and instruction-tuned setups. Removing length normalization causes the largest decline in AlpacaEval 2 and Arena-Hard performance; setting γ=0\gamma=0 also lowers results relative to SimPO.

    SetupMethodAlpacaEval 2 LC (%)Arena-Hard WR (%)MT-Bench GPT-4 Turbo
    Mistral-BaseDPO15.110.45.9
    Mistral-BaseSimPO21.516.66.0
    Mistral-BaseSimPO w/o LN11.99.45.5
    Mistral-BaseSimPO, γ=0\gamma=016.811.75.6
    Mistral-InstructDPO26.816.36.3
    Mistral-InstructSimPO32.121.06.6
    Mistral-InstructSimPO w/o LN19.116.36.4
    Mistral-InstructSimPO, γ=0\gamma=030.920.56.6

    The margin sweep shows that increasing γ\gamma improves held-out reward accuracy, but AlpacaEval 2 win rate first rises and then falls as the margin grows. The authors observe that overly large margins flatten reward and winning-response likelihood distributions and can lead to degeneration; the useful margin therefore requires tuning.

  6. Knowl 6 — DPO reward rankings can disagree with response-likelihood rankings

    empirical result

    On UltraFeedback training pairs, the paper compares the ordering induced by DPO's reward with the ordering induced by average log likelihood, the generation-aligned score. Among the 44.2 thousand pairs for which DPO assigns the preferred response a higher reward, 21.2 thousand nevertheless have lower average log likelihood for that preferred response. Thus, in nearly half of the pairs that DPO ranks correctly by its own reward, its reward ordering disagrees with the likelihood ordering used for generation. For SimPO, which uses scaled average log likelihood as its reward, the paper reports no such ordering disagreement in this comparison. Separately, on held-out preference pairs, SimPO has higher reward accuracy than DPO in both Mistral-Base and Mistral-Instruct settings; reward accuracy is the fraction of pairs where the preferred response receives the higher learned reward.

  7. Knowl 7 — Length normalization reduces length-related reward behavior

    empirical result

    On held-out responses, the Spearman correlation between average log probability and response length is 0.82 for SimPO trained without length normalization, 0.59 for DPO, 0.34 for SimPO, and 0.33 for the Mistral supervised fine-tuned model. The much stronger correlation for SimPO without length normalization is evidence of a stronger association between the learned likelihood score and response length. In the authors' analyses, removing normalization also produced long, repetitive generations and lower benchmark win rates.

    Setup and benchmarkSimPO win rateSimPO mean lengthSimPO w/o LN win rateSimPO w/o LN mean length
    Mistral-Base, AlpacaEval 2 LC (%)21.5186811.92345
    Mistral-Base, Arena-Hard WR (%)16.66999.4851
    Mistral-Instruct, AlpacaEval 2 LC (%)32.1219319.12067
    Mistral-Instruct, Arena-Hard WR (%)21.053916.3679

    The length effect is not uniform across every benchmark and setup: for example, the Mistral-Instruct response length on AlpacaEval 2 is lower without normalization. The reported pattern is that the unnormalized objective is more length-correlated and performs worse, not that it always produces longer responses.

  8. Knowl 8 — Reference-free training reduces measured time and memory use

    empirical result

    In a Llama-3-Base training comparison on 8 H100 GPUs, SimPO took 60 minutes and reached 69 GB peak GPU memory per GPU, whereas the vanilla DPO implementation took 73 minutes and reached 77 GB. The authors attribute the savings to removing reference-model forward passes. This corresponds to about 18% less runtime and 10% lower peak GPU memory for this particular implementation and setup.

  9. Knowl 9 — Gemma-2-9B SimPO reaches strong chat scores while retaining measured capabilities

    empirical result

    For Gemma-2-9B-it, the authors generated up to five responses per UltraFeedback prompt and used ArmoRM to label preference pairs. The table compares the starting checkpoint, DPO, and SimPO variants trained with different learning rates. Scores are AlpacaEval 2 length-controlled win rate, Arena-Hard win rate, and ZeroEval GSM and MMLU scores. The SimPO model trained at learning rate 8e−78\mathrm{e}{-7} achieved the strongest reported chat scores while its GSM and MMLU scores remained close to the original model's scores.

    ModelAlpacaEval 2 LC (%)Arena-Hard WR (%)ZeroEval GSM (%)ZeroEval MMLU (%)
    Gemma-2-9B-it51.140.887.472.7
    Gemma-2-9B-DPO67.858.988.572.2
    Gemma-2-9B-SimPO, lr=6e−76\mathrm{e}{-7}71.758.388.372.2
    Gemma-2-9B-SimPO, lr=8e−78\mathrm{e}{-7}72.459.188.072.2
    Gemma-2-9B-SimPO, lr=1e−61\mathrm{e}{-6}71.058.387.471.5

    The authors also report that the 8e−78\mathrm{e}{-7} checkpoint moved Gemma-2-9B-it from 36th to 25th on Chatbot Arena and ranked first among models below 10B parameters by real-user votes, as of September 16, 2024.

  10. Knowl 10 — Downstream capability retention varies by model and task

    limitation

    The paper does not establish that preference optimization preserves every capability. On the reported Llama-3-Instruct evaluation, SimPO's GSM8K score was 50.72 versus 68.69 for the SFT checkpoint, and its MMLU score was 65.63 versus 67.06; the authors report that math performance often declines across preference-optimization methods. In contrast, the Gemma-2-9B experiment retained scores near the initial checkpoint on GSM and MMLU, indicating that capability changes depend on the model and training setup. The authors also state that SimPO does not explicitly impose safety or honesty constraints, that the UltraFeedback data used primarily emphasizes helpfulness, and that the target margin requires manual tuning. They identify a more rigorous theoretical account and automatic margin selection as open work.

Coverage note — The appendix's gradient decomposition and qualitative response examples are omitted because they explain or illustrate the method but do not add a separate load-bearing result beyond these knowls.

References

  1. 1.Alan Agresti. Categorical data analysis, volume 792. John Wiley & Sons, 2012.
  2. 2.AI@Meta. Llama 3 model card. 2024.
  3. 3.Afra Amini, Tim Vieira, and Ryan Cotterell. Direct preference optimization with an offset. arXiv preprint arXiv:2402.10571, 2024.
  4. 4.Thomas Anthony, Zheng Tian, and David Barber. Thinking fast and slow with deep learning and tree search. Advances in neural information processing systems, 30, 2017.
  5. 5.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, John Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, and Jared Kaplan. A general language assistant as a laboratory for alignment. ArXiv, abs/2112.00861, 2021.
  6. 6.Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. A general theoretical paradigm to understand learning from human preferences. ArXiv, abs/2310.12036, 2023.
  7. 7.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  8. 8.Hritik Bansal, Ashima Suvarna, Gantavya Bhatt, Nanyun Peng, Kai-Wei Chang, and Aditya Grover. Comparing bad apples to good oranges: Aligning large language models via joint preference optimization. arXiv preprint arXiv:2404.00530, 2024.
  9. 9.Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open LLM leaderboard. https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard, 2023.
  10. 10.Bernhard E Boser, Isabelle M Guyon, and Vladimir N Vapnik. A training algorithm for optimal margin classifiers. In Proceedings of the fifth annual workshop on Computational learning theory, pages 144–152, 1992.
  11. 11.Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39:324, 1952.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020.
  13. 13.Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217, 2023.
  14. 14.Angelica Chen, Sadhika Malladi, Lily H Zhang, Xinyi Chen, Qiuyi Zhang, Rajesh Ranganath, and Kyunghyun Cho. Preference learning algorithms do not learn preference rankings. NeurIPS, 2024.
  15. 15.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. AlpaGasus: Training a better Alpaca with fewer data. In ICLR, 2024.
  16. 16.Lichang Chen, Chen Zhu, Davit Soselia, Jiuhai Chen, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro. ODIN: Disentangled reward mitigates hacking in RLHF. arXiv preprint arXiv:2402.07319, 2024.
  17. 17.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024.
  18. 18.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  19. 19.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge. ArXiv, abs/1803.05457, 2018.
  20. 20.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  21. 21.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world’s first truly open instruction-tuned LLM, 2023.
  22. 22.Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine learning, 20:273–297, 1995.
  23. 23.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. UltraFeedback: Boosting language models with high-quality feedback. In ICML, 2024.
  24. 24.Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. Safe RLHF: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023.
  25. 25.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. Enhancing chat language models by scaling high-quality instructional conversations. In EMNLP, 2023.
  26. 26.Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, SHUM KaShun, and Tong Zhang. RAFT: Reward ranked finetuning for generative foundation model alignment. Transactions on Machine Learning Research, 2023.
  27. 27.Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. Rlhf workflow: From reward modeling to online rlhf. arXiv preprint arXiv:2405.07863, 2024.
  28. 28.Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. Length-controlled AlpacaEval: A simple way to debias automatic evaluators. ArXiv, abs/2404.04475, 2024.
  29. 29.Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. KTO: Model alignment as prospect theoretic optimization. ArXiv, abs/2402.01306, 2024.
  30. 30.Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, 2018.
  31. 31.David Firth and Heather Turner. Bradley-terry models in R: the BradleyTerry2 package. Journal of Statistical Software, 48(9), 2012.
  32. 32.Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pages 10835–10866. PMLR, 2023.
  33. 33.Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song. Koala: A dialogue model for academic research. Blog post, April, 1:6, 2023.
  34. 34.Ulrich Germann. Greedy decoding for statistical machine translation in almost linear time. In NAACL, 2003.
  35. 35.Alex Graves. Sequence transduction with recurrent neural networks. ArXiv, abs/1211.3711, 2012.
  36. 36.Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy, Christoforos Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, Sainbayar Sukhbaatar, and Roberta Raileanu. Teaching large language models to reason with reinforcement learning. arXiv preprint arXiv:2403.04642, 2024.
  37. 37.Alex Havrilla, Sharath Raparthy, Christoforus Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, and Roberta Railneau. GLoRe: When, where, and how to improve LLM reasoning via global and local refinements. arXiv preprint arXiv:2402.10963, 2024.
  38. 38.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2020.
  39. 39.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2019.
  40. 40.Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. Learning to write with cooperative discriminators. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1638–1649, 2018.
  41. 41.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051, 2021.
  42. 42.Jiwoo Hong, Noah Lee, and James Thorne. ORPO: Monolithic preference optimization without reference model. ArXiv, abs/2403.07691, 2024.
  43. 43.Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. BeaverTails: Towards improved safety alignment of LLM via a human-preference dataset. ArXiv, abs/2307.04657, 2023.
  44. 44.Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7B. ArXiv, abs/2310.06825, 2023.
  45. 45.Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion. In ACL, 2023.
  46. 46.Dahyun Kim, Yungi Kim, Wonho Song, Hyeonwoo Kim, Yunsu Kim, Sanghoon Kim, and Chanjun Park. sDPO: Don’t use your data all at once. ArXiv, abs/2403.19270, 2024.
  47. 47.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  48. 48.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Minh Nguyen, Oliver Stanley, Richárd Nagyfi, et al. Openassistant conversations-democratizing large language model alignment. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  49. 49.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. Pretraining language models with human preferences. In International Conference on Machine Learning, pages 17506–17533. PMLR, 2023.
  50. 50.Nathan Lambert, Valentina Pyatkin, Jacob Daniel Morrison, Lester James Validad Miranda, Bill Yuchen Lin, Khyathi Raghavi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hanna Hajishirzi. RewardBench: Evaluating reward models for language modeling. ArXiv, abs/2403.13787, 2024.
  51. 51.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871, 2018.
  52. 52.Hector Levesque, Ernest Davis, and Leora Morgenstern. The Winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  53. 53.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. Deep reinforcement learning for dialogue generation. In Jian Su, Kevin Duh, and Xavier Carreras, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Austin, Texas, November 2016. Association for Computational Linguistics.
  54. 54.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From live data to high-quality benchmarks: The Arena-Hard pipeline, April 2024.
  55. 55.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. AlpacaEval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval, 2023.
  56. 56.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  57. 57.Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In ACL, pages 3214–3252, 2022.
  58. 58.Tianqi Liu, Zhen Qin, Junru Wu, Jiaming Shen, Misha Khalman, Rishabh Joshi, Yao Zhao, Mohammad Saleh, Simon Baumgartner, Jialu Liu, et al. LiPO: Listwise preference optimization through learning-to-rank. arXiv preprint arXiv:2402.01878, 2024.
  59. 59.Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu. Statistical rejection sampling improves preference optimization. In The Twelfth International Conference on Learning Representations, 2024.
  60. 60.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583, 2023.
  61. 61.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  62. 62.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. Training language models to follow instructions with human feedback. In NeurIPS, 2022.
  63. 63.Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White. Smaug: Fixing failure modes of preference optimisation with dpo-positive. arXiv preprint arXiv:2402.13228, 2024.
  64. 64.Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn. Disentangling length from quality in direct preference optimization. ArXiv, abs/2403.19159, 2024.
  65. 65.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  66. 66.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023.
  67. 67.Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie. Direct nash optimization: Teaching language models to self-improve with general preferences. ArXiv, abs/2404.03715, 2024.
  68. 68.Sebastian Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747, 2016.
  69. 69.Michael Santacroce, Yadong Lu, Han Yu, Yuanzhi Li, and Yelong Shen. Efficient RLHF: Reducing the memory usage of PPO. arXiv preprint arXiv:2309.00754, 2023.
  70. 70.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  71. 71.Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. A long way to go: Investigating length correlations in RLHF. arXiv preprint arXiv:2310.03716, 2023.
  72. 72.Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. Preference ranking optimization for human alignment. In AAAI, 2024.
  73. 73.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  74. 74.Yunhao Tang, Daniel Zhaohan Guo, Zeyu Zheng, Daniele Calandriello, Yuan Cao, Eugene Tarassov, Rémi Munos, Bernardo Ávila Pires, Michal Valko, Yong Cheng, et al. Understanding the performance gap between online and offline alignment algorithms. arXiv preprint arXiv:2405.08448, 2024.
  75. 75.Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot. Generalized preference optimization: A unified approach to offline alignment. arXiv preprint arXiv:2402.05749, 2024.
  76. 76.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Stanford alpaca: An instruction-following llama model, 2023.
  77. 77.Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024.
  78. 78.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. Fine-tuning language models for factuality. In The Twelfth International Conference on Learning Representations, 2024.
  79. 79.Hoang Tran, Chris Glaze, and Braden Hancock. Iterative DPO alignment. Technical report, Snorkel AI, 2023.
  80. 80.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. Zephyr: Direct distillation of LM alignment. ArXiv, abs/2310.16944, 2023.
  81. 81.Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zi-Han Lin, Yuk-Kit Cheng, Sanmi Koyejo, Dawn Xiaodong Song, and Bo Li. DecodingTrust: A comprehensive assessment of trustworthiness in GPT models. In NeurIPS, 2023.
  82. 82.Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. OpenChat: Advancing open-source language models with mixed-quality data. In ICLR, 2024.
  83. 83.Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, and Tong Zhang. Arithmetic control of LLMs for diverse user preferences: Directional preference alignment with multi-objective rewards. ArXiv, abs/2402.18571, 2024.
  84. 84.Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. arXiv preprint arXiv:2406.12845, 2024.
  85. 85.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al. How far can camels go? exploring the state of instruction tuning on open resources. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023.
  86. 86.Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. LESS: Selecting influential data for targeted instruction tuning. In ICML, 2024.
  87. 87.Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint. In Forty-first International Conference on Machine Learning, 2024.
  88. 88.Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim. Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation. ArXiv, abs/2401.08417, 2024.
  89. 89.Jing Xu, Andrew Lee, Sainbayar Sukhbaatar, and Jason Weston. Some things are more cringe than others: Preference optimization with the pairwise cringe loss. arXiv preprint arXiv:2312.16682, 2023.
  90. 90.Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, and Yi Wu. Is DPO superior to PPO for LLM alignment? a comprehensive study. arXiv preprint arXiv:2404.10719, 2024.
  91. 91.Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. RRHF: Rank responses to align language models with human feedback. In NeurIPS, 2023.
  92. 92.Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. Self-rewarding language models. arXiv preprint arXiv:2401.10020, 2024.
  93. 93.Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. WildBench: Benchmarking LLMs with challenging tasks from real users in the wild. arXiv e-prints, pages arXiv–2406, 2024.
  94. 94.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Anna Korhonen, David Traum, and Lluís Màrquez, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics.
  95. 95.Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024.
  96. 96.Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu. SLiC-HF: Sequence likelihood calibration with human feedback. ArXiv, abs/2305.10425, 2023.
  97. 97.Chujie Zheng, Pei Ke, Zheng Zhang, and Minlie Huang. Click: Controllable text generation with sequence likelihood contrastive learning. In Findings of ACL, 2023.
  98. 98.Chujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang, and Nanyun Peng. Weak-to-strong extrapolation expedites alignment. arXiv preprint arXiv:2404.16792, 2024.
  99. 99.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In NeurIPS Datasets and Benchmarks Track, 2023.
  100. 100.Rui Zheng, Shihan Dou, Songyang Gao, Yuan Hua, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Yuhao Zhou, et al. Secrets of RLHF in large language models part I: PPO. arXiv preprint arXiv:2307.04964, 2023.
  101. 101.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. LIMA: Less is more for alignment. NeurIPS, 2023.
  102. 102.Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Meng, Y., et al. “SimPO: Simple Preference Optimization with a Reference-Free Reward”. arXiv, 2024, http://arxiv.org/abs/2405.14734v3.
APA
Meng, Y., Xia, M., & Chen, D. (2024). SimPO: Simple Preference Optimization with a Reference-Free Reward. arXiv. http://arxiv.org/abs/2405.14734v3
Chicago
Meng, Y., M. Xia, and D. Chen. 2024. “SimPO: Simple Preference Optimization with a Reference-Free Reward”. arXiv. http://arxiv.org/abs/2405.14734v3.
Harvard
Meng, Y., Xia, M. and Chen, D. (2024) “SimPO: Simple Preference Optimization with a Reference-Free Reward”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.14734v3.
Vancouver
1. Meng Y, Xia M, Chen D (2024) SimPO: Simple Preference Optimization with a Reference-Free Reward. arXiv

BibTeX

@article{meng2024simpo,
  title = {SimPO: Simple Preference Optimization with a Reference-Free Reward},
  author = {Meng, Yu and Xia, Mengzhou and Chen, Danqi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.14734v3},
  eprint = {2405.14734}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors