Abstract

Large language models (LLMs) often contain misleading content, emphasizing the need to align them with human values to ensure secure AI systems. Reinforcement learning from human feedback (RLHF) has been employed to achieve this alignment. However, it encompasses two main drawbacks: (1) RLHF exhibits complexity, instability, and sensitivity to hyperparameters in contrast to SFT. (2) Despite massive trial-and-error, multiple sampling is reduced to pair-wise contrast, thus lacking contrasts from a macro perspective. In this paper, we propose Preference Ranking Optimization (PRO) as an efficient SFT algorithm to directly fine-tune LLMs for human alignment. PRO extends the pair-wise contrast to accommodate preference rankings of any length. By iteratively contrasting candidates, PRO instructs the LLM to prioritize the best response while progressively ranking the rest responses. In this manner, PRO effectively transforms human alignment into aligning the probability ranking of n responses generated by LLM with the preference ranking of humans towards these responses. Experiments have shown that PRO outperforms baseline algorithms, achieving comparable results to ChatGPT and human responses through automatic-based, reward-based, GPT-4, and human evaluations.

Table of Contents

  • Abstract
  • Introduction
  • Preliminary
  • Methodology
  • From RLHF to PRO
  • Grafting RLHF onto PRO
  • Experiments
  • Data Preparation
  • Evaluation Metrics
  • Implementation Details
  • Main Experiment
  • Effect of Expanding Preference Ranking Sequence
  • Human and GPT-4 Evaluation
  • Ablation Study
  • Related Work
  • Reinforcement Learning from Human Feedback
  • SFT for Human Preference Alignment
  • Conclusion
  • Ethics Statement
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — PRO aligns a language model to an entire preference ranking

    model/method

    Given a prompt xx and n≥2n\ge2 candidate responses ordered by human preference as y1≻y2≻⋯≻yny^1\succ y^2\succ\cdots\succ y^n, Preference Ranking Optimization (PRO) trains the language model through successive one-versus-the-rest contrasts. At position kk, the response yky^k is the positive and every response from yky^k through yny^n is included in the comparison denominator; the process advances through k=1,…,n−1k=1,\ldots,n-1. The ranking loss is Lrank=−log⁡∏k=1n−1exp⁡(sk)∑i=knexp⁡(si)L_{\mathrm{rank}}=-\log\prod_{k=1}^{n-1}\frac{\exp(s_k)}{\sum_{i=k}^{n}\exp(s_i)}, where sis_i is the model score for response yiy^i under prompt xx. In PRO, si=rπPRO(x,yi)=1∣yi∣∑t=1∣yi∣log⁡P(yti∣x,y<ti)s_i=r_{\pi_{\mathrm{PRO}}}(x,y^i)=\frac{1}{|y^i|}\sum_{t=1}^{|y^i|}\log P(y^i_t\mid x,y^i_{<t}): the mean conditional token log-probability of response yiy^i, whose length is ∣yi∣|y^i|. PRO adds the negative log-likelihood LSFTL_{\mathrm{SFT}} of the top-ranked response to preserve response quality, giving LPRO=Lrank+βLSFTL_{\mathrm{PRO}}=L_{\mathrm{rank}}+\beta L_{\mathrm{SFT}}, with β\beta controlling the balance. The model is optimized directly with supervised fine-tuning rather than by an RL trial-and-error loop. For a ranking of length two, this reduces to a pairwise contrast.

  2. Knowl 2 — Reward gaps set the strength of PRO’s ranking contrasts

    model/method

    PRO also defines a differentiated-contrast variant that uses a reward model’s scores to vary the penalty applied to lower-ranked responses. For a prompt xx and responses ordered by the reward model rϕr_\phi, let sis_i be the policy model’s score for response yiy^i, and suppose the relevant reward gaps are positive. For each position kk and lower-ranked response yiy^i with i>ki>k, set Ti(k)=1/[rϕ(x,yk)−rϕ(x,yi)]T_i^{(k)}=1/[r_\phi(x,y^k)-r_\phi(x,y^i)] and set the numerator temperature to Tk(k)=min⁡i>kTi(k)T_k^{(k)}=\min_{i>k}T_i^{(k)}. The differentiated ranking loss is −∑k=1n−1log⁡exp⁡(sk/Tk(k))∑i=knexp⁡(si/Ti(k))-\sum_{k=1}^{n-1}\log\frac{\exp(s_k/T_k^{(k)})}{\sum_{i=k}^{n}\exp(s_i/T_i^{(k)})}. Thus, larger reward gaps yield lower temperatures and sharper contrasts, while close-scoring alternatives receive less severe differentiation. The paper reports that this temperature design improves the ranking-only objective and can also provide gains when combined with the top-response SFT loss.

  3. Knowl 3 — Preference rankings can be expanded with heterogeneous model responses

    model/method

    PRO requires ranked responses but does not require that every candidate be written by a human or generated by the same model. The paper describes extending a ranking with responses from different language models, then using a reward model to score and order the candidates for training. It also proposes self-bootstrapping: sample an additional candidate from the recipient language model, score it with an additional reward model, and add it to the training ranking. These options are intended to increase the number, quality, and diversity of contrastive examples; the paper presents them as ways to construct affordable longer rankings rather than as a separate optimization objective.

  4. Knowl 4 — Training and evaluation used HH-RLHF and LLaMA-7B

    experimental setup

    Experiments used the four HH-RLHF subsets Harmlessbase, Helpfulbase, Helpfulonline, and Helpfulrejection, with human-ranked pairs in the unaugmented data. Augmented training rankings included additional responses from Alpaca or ChatGPT; further length experiments added responses from Curie, Alpaca-7B, and ChatGPT. LLaMA-7B was the fine-tuning backbone. Training used sequence length 512, two epochs, learning rate 5×10−65\times10^{-6}, and total batch size 112; inference generated at most 128 new tokens. For ranking length ℓ\ell, the weight on the top-response SFT loss was β=0.05(ℓ−1)2\beta=0.05(\ell-1)^2. A training reward model, RMtrain, scored and sorted augmented rankings before training, while a separate evaluation reward model, RMeval, measured preference scores; RMeval outputs were sigmoid-normalized. BLEU assessed text quality. GPT-4 and human evaluations compared PRO responses with reference responses, with GPT-4 comparisons run in both candidate orders and averaged to reduce positional bias.

  5. Knowl 5 — PRO’s aggregate reward scores were strongest on ChatGPT-augmented rankings

    empirical result

    On the HH-RLHF test sets, the paper reports total BLEU/Reward scores for zero-shot models and fine-tuned LLaMA-7B systems. The zero-shot scores were LLaMA 13.13/38.94, Curie 16.99/48.71, Alpaca 19.12/52.72, ChatGLM 21.99/61.27, and ChatGPT 22.56/68.48. On raw pairwise HH-RLHF, SFT scored 21.80/48.83, RLHF 21.19/48.93, CoH 24.06/45.00, DPO 22.62/52.75, RRHF 20.91/52.25, and PRO 21.54/55.35. With Alpaca-augmented rankings of length three, BoN scored 23.7/57.66, RLHF 23.82/57.28, CoH 23.54/47.15, DPO 22.98/59.27, RRHF 21.02/55.39, and PRO 22.11/58.72. With ChatGPT-augmented rankings of length three, the corresponding scores were BoN 22.45/63.83, RLHF 20.99/58.65, CoH 23.26/55.58, DPO 22.35/64.10, RRHF 20.86/63.12, and PRO 23.07/67.97. Thus PRO had the highest reward among the listed fine-tuned methods on raw and ChatGPT-augmented rankings, but not on Alpaca-augmented rankings, where DPO’s reward was 59.27 versus PRO’s 58.72. PRO’s reward on the ChatGPT-augmented data was close to the zero-shot ChatGPT score. For raw HH-RLHF, PRO’s BLEU/Reward by subset was Harmlessbase 12.05/62.96, Helpfulbase 20.83/48.51, Helpfulonline 28.75/59.02, and Helpfulrejection 27.17/53.28.

  6. Knowl 6 — Longer and more diverse rankings generally improved PRO

    empirical result

    To test ranking length, the researchers started with raw HH-RLHF pairs and added up to three responses, producing rankings of lengths two through five. They compared four addition strategies: repeated Alpaca responses; repeated ChatGPT responses; ascending response quality, adding Curie, then Alpaca-7B, then ChatGPT; and random ordering. Rankings were reordered with a reward model. The reported results show that longer rankings generally improved PRO, though not every strategy improved at every step. Repeated additions from Alpaca produced limited gains after the first added response, whereas successive high-quality ChatGPT responses led to consistent gains. Diversity also mattered: at length four, adding Curie and Alpaca responses outperformed adding two Alpaca responses, despite Curie’s lower response quality. At length five, combining Curie, Alpaca, and ChatGPT achieved performance close to using three ChatGPT responses. These findings support the paper’s qualified conclusion that both the quality and heterogeneity of added candidates affect alignment performance.

  7. Knowl 7 — GPT-4 and human judges compared PRO with the human top-ranked response

    empirical result

    The evaluation compared responses from PRO trained on raw HH-RLHF with the dataset’s first-ranked response, called the Golden response. Across subsets, GPT-4 judged PRO to win 55.00% of comparisons, tie 4.37%, and lose 40.63%; human annotators judged PRO to win 22.50%, tie 56.25%, and lose 21.25%. The GPT-4 result varied by subset: for Helpfulonline, PRO won 27.50%, tied 12.50%, and lost 60.00%, unlike the other three subsets, where PRO’s win rate exceeded its loss rate. Human win rates by Harmlessbase, Helpfulbase, Helpfulonline, and Helpfulrejection were 20%, 20%, 20%, and 30%, respectively, with ties frequent in every subset. GPT-4 saw both candidate orders, and three annotators assessed the same sampled comparisons for the human evaluation. The results therefore show an overall GPT-4 preference for PRO but a predominantly tied human assessment, rather than uniform preference across evaluators and tasks.

  8. Knowl 8 — Ablations show complementary roles for ranking, SFT, and temperature

    empirical result

    Ablations measured total BLEU/Reward on raw and augmented HH-RLHF. On raw data, full PRO scored 21.54/55.35; removing the SFT term yielded 9.85/53.25, removing temperature yielded 21.41/55.04, and removing both yielded 5.14/46.17. On Alpaca-augmented rankings, full PRO scored 22.11/58.72; keeping only the first ranking contrast scored 21.10/58.11, removing SFT scored 18.29/59.71, removing temperature scored 21.34/58.40, and removing both scored 2.05/32.33. On ChatGPT-augmented rankings, the corresponding results were 23.07/67.97 for full PRO, 22.80/67.75 with only the first ranking contrast, 21.84/67.84 without SFT, 22.98/68.40 without temperature, and 6.25/43.16 without both. The authors interpret the SFT term as helping maintain text quality and the multiple ranking contrasts as helping distinguish preference levels. Removing temperature alone had modest and nonuniform effects, but removing both temperature and SFT sharply degraded results, especially BLEU.

Coverage note — The exhaustive per-subset BLEU and reward matrices for every baseline and ablation are not reproduced; aggregate comparisons and salient subset-level findings retain the central results without duplicating the full result tables.

References

  1. 1.Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; Das-Sarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862.
  2. 2.Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022b. Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
  3. 3.Bradley, R. A.; and Terry, M. E. 1952. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4): 324–345.
  4. 4.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  5. 5.Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv:2303.12712.
  6. 6.Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022. Palm: Scaling language modeling with pathways. arXiv:2204.02311.
  7. 7.Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep Reinforcement Learning from Human Preferences. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  8. 8.Dong, H.; Xiong, W.; Goyal, D.; Pan, R.; Diao, S.; Zhang, J.; Shum, K.; and Zhang, T. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv:2304.06767.
  9. 9.Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 320–335. Dublin, Ireland: Association for Computational Linguistics.
  10. 10.Gugger, S.; Debut, L.; Wolf, T.; Schmid, P.; Mueller, Z.; and Mangrulkar, S. 2022. Accelerate: Training and inference at scale made simple, efficient and adaptable.
  11. 11.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  12. 12.Lee, K.; Smith, L.; and Abbeel, P. 2021. Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training. arXiv:2106.05091.
  13. 13.Lei, W.; Zhang, Y.; Song, F.; Liang, H.; Mao, J.; Lv, J.; Yang, Z.; and Chua, T.-S. 2022. Interacting with Non-Cooperative User: A New Paradigm for Proactive Dialogue Policy. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, 212–222. New York, NY, USA: Association for Computing Machinery. ISBN 9781450387323.
  14. 14.Li, M.; Song, F.; Yu, B.; Yu, H.; Li, Z.; Huang, F.; and Li, Y. 2023. Api-bank: A benchmark for tool-augmented llms. arXiv:2304.08244.
  15. 15.Liu, H.; Sferrazza, C.; and Abbeel, P. 2023. Chain of Hindsight aligns Language Models with Feedback. arXiv:2302.02676.
  16. 16.Luce, R. D. 2012. Individual choice behavior: A theoretical analysis. Courier Corporation.
  17. 17.MacGlashan, J.; Ho, M. K.; Loftin, R.; Peng, B.; Wang, G.; Roberts, D. L.; Taylor, M. E.; and Littman, M. L. 2017. Interactive Learning from Policy-Dependent Human Feedback. In Precup, D.; and Teh, Y. W., eds., Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, 2285–2294. PMLR.
  18. 18.Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv:2112.09332.
  19. 19.OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774.
  20. 20.Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744.
  21. 21.Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–318. Philadelphia, Pennsylvania, USA: Association for Computational Linguistics.
  22. 22.Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023. Instruction tuning with gpt-4. arXiv:2304.03277.
  23. 23.Plackett, R. L. 1975. The analysis of permutations. Journal of the Royal Statistical Society Series C: Applied Statistics, 24(2): 193–202.
  24. 24.Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290.
  25. 25.Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal policy optimization algorithms. arXiv:1707.06347.
  26. 26.Snell, C.; Kostrikov, I.; Su, Y.; Yang, M.; and Levine, S. 2022. Offline rl for natural language generation with implicit language q learning. arXiv:2206.11871.
  27. 27.Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021.
  28. 28.Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023. Stanford Alpaca: An Instruction-following LLaMA model.
  29. 29.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; `Azhar, F.; et al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971.
  30. 30.Wang, P.; Li, L.; Chen, L.; Zhu, D.; Lin, B.; Cao, Y.; Liu, Q.; Liu, T.; and Sui, Z. 2023. Large language models are not fair evaluators. arXiv:2305.17926.
  31. 31.Warnell, G.; Waytowich, N.; Lawhern, V.; and Stone, P. 2018. Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  32. 32.Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Le Scao, T.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45. Online: Association for Computational Linguistics.
  33. 33.Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023. Fine-Grained Human Feedback Gives Better Rewards for Language Model Training. arXiv:2306.01693.
  34. 34.Xue, W.; An, B.; Yan, S.; and Xu, Z. 2023. Reinforcement Learning from Diverse Human Preferences. arXiv:2301.11774.
  35. 35.Yuan, Z.; Yuan, H.; Tan, C.; Wang, W.; Huang, S.; and Huang, F. 2023. Rrhf: Rank responses to align language models with human feedback without tears. arXiv:2304.05302.
  36. 36.Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685.
  37. 37.Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; et al. 2023. Lima: Less is more for alignment. arXiv:2305.11206.
  38. 38.Zhu, B.; Jiao, J.; and Jordan, M. I. 2023. Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. arXiv:2301.11270.
  39. 39.Zhu, B.; Sharma, H.; Frujeri, F. V.; Dong, S.; Zhu, C.; Jordan, M. I.; and Jiao, J. 2023. Fine-Tuning Language Models with Advantage-Induced Policy Alignment. arXiv:2306.02231.
  40. 40.Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019. Fine-tuning language models from human preferences. arXiv:1909.08593.

Citation

MLA
Song, F., et al. “Preference Ranking Optimization for Human Alignment”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, 2024, pp. 18990–98, https://doi.org/10.1609/AAAI.V38I17.29865.
APA
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., & Wang, H. (2024). Preference Ranking Optimization for Human Alignment. Proceedings of the AAAI Conference on Artificial Intelligence, 38(17), 18990–18998. https://doi.org/10.1609/AAAI.V38I17.29865
Chicago
Song, F., B. Yu, M. Li, et al. 2024. “Preference Ranking Optimization for Human Alignment”. Proceedings of the AAAI Conference on Artificial Intelligence 38 (17): 18990–98. https://doi.org/10.1609/AAAI.V38I17.29865.
Harvard
Song, F. et al. (2024) “Preference Ranking Optimization for Human Alignment”, Proceedings of the AAAI Conference on Artificial Intelligence, 38(17), pp. 18990–18998. Available at: https://doi.org/10.1609/AAAI.V38I17.29865.
Vancouver
1. Song F, Yu B, Li M, Yu H, Huang F, Li Y, Wang H (2024) Preference Ranking Optimization for Human Alignment. Proceedings of the AAAI Conference on Artificial Intelligence 38:18990–18998

BibTeX

@article{Song_2024, title={Preference Ranking Optimization for Human Alignment}, volume={38}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V38I17.29865}, DOI={10.1609/aaai.v38i17.29865}, number={17}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Song, Feifan and Yu, Bowen and Li, Minghao and Yu, Haiyang and Huang, Fei and Li, Yongbin and Wang, Houfeng}, year={2024}, month=Mar, pages={18990–18998} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF