Preference Ranking Optimization for Human Alignment
Feifan SongBowen YuMinghao LiHaiyang YuFei HuangYongbin LiHoufeng Wang
Feifan SongBowen YuMinghao LiHaiyang YuFei HuangYongbin LiHoufeng Wang
Large language models (LLMs) often contain misleading content, emphasizing the need to align them with human values to ensure secure AI systems. Reinforcement learning from human feedback (RLHF) has been employed to achieve this alignment. However, it encompasses two main drawbacks: (1) RLHF exhibits complexity, instability, and sensitivity to hyperparameters in contrast to SFT. (2) Despite massive trial-and-error, multiple sampling is reduced to pair-wise contrast, thus lacking contrasts from a macro perspective. In this paper, we propose Preference Ranking Optimization (PRO) as an efficient SFT algorithm to directly fine-tune LLMs for human alignment. PRO extends the pair-wise contrast to accommodate preference rankings of any length. By iteratively contrasting candidates, PRO instructs the LLM to prioritize the best response while progressively ranking the rest responses. In this manner, PRO effectively transforms human alignment into aligning the probability ranking of n responses generated by LLM with the preference ranking of humans towards these responses. Experiments have shown that PRO outperforms baseline algorithms, achieving comparable results to ChatGPT and human responses through automatic-based, reward-based, GPT-4, and human evaluations.
Given a prompt and candidate responses ordered by human preference as , Preference Ranking Optimization (PRO) trains the language model through successive one-versus-the-rest contrasts. At position , the response is the positive and every response from through is included in the comparison denominator; the process advances through . The ranking loss is , where is the model score for response under prompt . In PRO, : the mean conditional token log-probability of response , whose length is . PRO adds the negative log-likelihood of the top-ranked response to preserve response quality, giving , with controlling the balance. The model is optimized directly with supervised fine-tuning rather than by an RL trial-and-error loop. For a ranking of length two, this reduces to a pairwise contrast.
PRO also defines a differentiated-contrast variant that uses a reward model’s scores to vary the penalty applied to lower-ranked responses. For a prompt and responses ordered by the reward model , let be the policy model’s score for response , and suppose the relevant reward gaps are positive. For each position and lower-ranked response with , set and set the numerator temperature to . The differentiated ranking loss is . Thus, larger reward gaps yield lower temperatures and sharper contrasts, while close-scoring alternatives receive less severe differentiation. The paper reports that this temperature design improves the ranking-only objective and can also provide gains when combined with the top-response SFT loss.
PRO requires ranked responses but does not require that every candidate be written by a human or generated by the same model. The paper describes extending a ranking with responses from different language models, then using a reward model to score and order the candidates for training. It also proposes self-bootstrapping: sample an additional candidate from the recipient language model, score it with an additional reward model, and add it to the training ranking. These options are intended to increase the number, quality, and diversity of contrastive examples; the paper presents them as ways to construct affordable longer rankings rather than as a separate optimization objective.
Experiments used the four HH-RLHF subsets Harmlessbase, Helpfulbase, Helpfulonline, and Helpfulrejection, with human-ranked pairs in the unaugmented data. Augmented training rankings included additional responses from Alpaca or ChatGPT; further length experiments added responses from Curie, Alpaca-7B, and ChatGPT. LLaMA-7B was the fine-tuning backbone. Training used sequence length 512, two epochs, learning rate , and total batch size 112; inference generated at most 128 new tokens. For ranking length , the weight on the top-response SFT loss was . A training reward model, RMtrain, scored and sorted augmented rankings before training, while a separate evaluation reward model, RMeval, measured preference scores; RMeval outputs were sigmoid-normalized. BLEU assessed text quality. GPT-4 and human evaluations compared PRO responses with reference responses, with GPT-4 comparisons run in both candidate orders and averaged to reduce positional bias.
On the HH-RLHF test sets, the paper reports total BLEU/Reward scores for zero-shot models and fine-tuned LLaMA-7B systems. The zero-shot scores were LLaMA 13.13/38.94, Curie 16.99/48.71, Alpaca 19.12/52.72, ChatGLM 21.99/61.27, and ChatGPT 22.56/68.48. On raw pairwise HH-RLHF, SFT scored 21.80/48.83, RLHF 21.19/48.93, CoH 24.06/45.00, DPO 22.62/52.75, RRHF 20.91/52.25, and PRO 21.54/55.35. With Alpaca-augmented rankings of length three, BoN scored 23.7/57.66, RLHF 23.82/57.28, CoH 23.54/47.15, DPO 22.98/59.27, RRHF 21.02/55.39, and PRO 22.11/58.72. With ChatGPT-augmented rankings of length three, the corresponding scores were BoN 22.45/63.83, RLHF 20.99/58.65, CoH 23.26/55.58, DPO 22.35/64.10, RRHF 20.86/63.12, and PRO 23.07/67.97. Thus PRO had the highest reward among the listed fine-tuned methods on raw and ChatGPT-augmented rankings, but not on Alpaca-augmented rankings, where DPO’s reward was 59.27 versus PRO’s 58.72. PRO’s reward on the ChatGPT-augmented data was close to the zero-shot ChatGPT score. For raw HH-RLHF, PRO’s BLEU/Reward by subset was Harmlessbase 12.05/62.96, Helpfulbase 20.83/48.51, Helpfulonline 28.75/59.02, and Helpfulrejection 27.17/53.28.
To test ranking length, the researchers started with raw HH-RLHF pairs and added up to three responses, producing rankings of lengths two through five. They compared four addition strategies: repeated Alpaca responses; repeated ChatGPT responses; ascending response quality, adding Curie, then Alpaca-7B, then ChatGPT; and random ordering. Rankings were reordered with a reward model. The reported results show that longer rankings generally improved PRO, though not every strategy improved at every step. Repeated additions from Alpaca produced limited gains after the first added response, whereas successive high-quality ChatGPT responses led to consistent gains. Diversity also mattered: at length four, adding Curie and Alpaca responses outperformed adding two Alpaca responses, despite Curie’s lower response quality. At length five, combining Curie, Alpaca, and ChatGPT achieved performance close to using three ChatGPT responses. These findings support the paper’s qualified conclusion that both the quality and heterogeneity of added candidates affect alignment performance.
The evaluation compared responses from PRO trained on raw HH-RLHF with the dataset’s first-ranked response, called the Golden response. Across subsets, GPT-4 judged PRO to win 55.00% of comparisons, tie 4.37%, and lose 40.63%; human annotators judged PRO to win 22.50%, tie 56.25%, and lose 21.25%. The GPT-4 result varied by subset: for Helpfulonline, PRO won 27.50%, tied 12.50%, and lost 60.00%, unlike the other three subsets, where PRO’s win rate exceeded its loss rate. Human win rates by Harmlessbase, Helpfulbase, Helpfulonline, and Helpfulrejection were 20%, 20%, 20%, and 30%, respectively, with ties frequent in every subset. GPT-4 saw both candidate orders, and three annotators assessed the same sampled comparisons for the human evaluation. The results therefore show an overall GPT-4 preference for PRO but a predominantly tied human assessment, rather than uniform preference across evaluators and tasks.
Ablations measured total BLEU/Reward on raw and augmented HH-RLHF. On raw data, full PRO scored 21.54/55.35; removing the SFT term yielded 9.85/53.25, removing temperature yielded 21.41/55.04, and removing both yielded 5.14/46.17. On Alpaca-augmented rankings, full PRO scored 22.11/58.72; keeping only the first ranking contrast scored 21.10/58.11, removing SFT scored 18.29/59.71, removing temperature scored 21.34/58.40, and removing both scored 2.05/32.33. On ChatGPT-augmented rankings, the corresponding results were 23.07/67.97 for full PRO, 22.80/67.75 with only the first ranking contrast, 21.84/67.84 without SFT, 22.98/68.40 without temperature, and 6.25/43.16 without both. The authors interpret the SFT term as helping maintain text quality and the multiple ranking contrasts as helping distinguish preference levels. Removing temperature alone had modest and nonuniform effects, but removing both temperature and SFT sharply degraded results, especially BLEU.
Coverage note — The exhaustive per-subset BLEU and reward matrices for every baseline and ablation are not reproduced; aggregate comparisons and salient subset-level findings retain the central results without duplicating the full result tables.
@article{Song_2024, title={Preference Ranking Optimization for Human Alignment}, volume={38}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V38I17.29865}, DOI={10.1609/aaai.v38i17.29865}, number={17}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Song, Feifan and Yu, Bowen and Li, Minghao and Yu, Haiyang and Huang, Fei and Li, Yongbin and Wang, Houfeng}, year={2024}, month=Mar, pages={18990–18998} }This paper has an official code repository available. Click below to access the source code.
View RepositoryThis paper is available from its original source. Click below to access the PDF.
Open PDF