HelpSteer2-Preference: Complementing Ratings with Preferences

Zhilin WangAlexander BukharinOlivier DelalleauDaniel EgertGerald ShenJiaqi ZengOleksii KuchaievYi Dong

article2025ICLR149 citations

Presents a method combining Bradley-Terry preferences and regression ratings using the open-source HelpSteer2-Preference dataset, achieving top performance on RewardBench and Arena-Hard for language model alignment.

Listen

Aligning large language models with human expectations requires effective reward models that score AI responses for helpfulness and safety. Historically, developers have used two competing training paradigms: Bradley-Terry models, which evaluate comparative preference between pairs of responses, and regression models, which assign absolute ratings to individual responses. A lack of directly comparable, high-quality data has created an ongoing debate over which training paradigm performs better.

The article aims to evaluate the relative strengths of Bradley-Terry and regression reward modeling approaches when trained on identical, purpose-built data, and to demonstrate how combining these methods improves language model alignment.

To conduct this evaluation, the researchers open-sourced HelpSteer2-Preference, a dataset containing pairwise human preferences, preference strengths, and written justifications alongside existing absolute ratings across 7,118 response pairs. The team trained several 70-billion-parameter reward models using Bradley-Terry, regression, and generative pairwise justification approaches, benchmarking them across 2,985 diverse tasks on RewardBench. They then integrated the top-performing reward model into reinforcement learning workflows to align instruction-following policy models.

The head-to-head comparison revealed four central findings. First, when matched on high-quality data, standard regression and Bradley-Terry reward models achieve virtually identical benchmark accuracy (93.0% and 92.7%, respectively). Second, combining both paradigms into a two-stage training pipeline—initializing a Bradley-Terry model with regression weights and applying weight extrapolation—achieves state-of-the-art accuracy of 94.1%, outperforming all public baselines. Third, reward models that generate text justifications perform worse (under 90.0% accuracy) than direct scoring models because the evaluation task is more complex. Fourth, in downstream reinforcement learning alignment, using this hybrid reward model within the REINFORCE framework achieved an 85.0 score on Arena Hard, outperforming traditional Direct Preference Optimization (52.9) and Proximal Policy Optimization (58.6) while maintaining core reasoning abilities.

These findings indicate that the format of data collection is less important than how effectively training objectives capture preference strength. The results demonstrate that two-stage reward modeling pipelines provide superior alignment guidance without compromising core model capabilities. For engineering trade-offs, regression models remain best suited for rapid, interpretable data filtering, while Bradley-Terry models are optimal for reinforcement learning.

Organizations developing frontier language models should adopt two-stage reward modeling and utilize online REINFORCE over offline Direct Preference Optimization where compute budgets allow. Further work should explore larger datasets across specialized domains, as well as investigate structured formatting for preference justifications.

The conclusions are limited by the small, general-domain dataset size (approximately 6,766 training pairs) and benchmark evaluations that rely heavily on automated judges powered by commercial language models. Nevertheless, given the consistent cross-benchmark gains and rigorous ablation studies, confidence in the primary findings remains high.

arXiv: 2410.01257
Cover for HelpSteer2-Preference: Complementing Ratings with Preferences

Abstract

Reward models are critical for aligning models to follow instructions, and are typically trained following one of two popular paradigms: Bradley-Terry style or Regression style. However, there is a lack of evidence that either approach is better than the other, when adequately matched for data. This is primarily because these approaches require data collected in different (but incompatible) formats, meaning that adequately matched data is not available in existing public datasets. To tackle this problem, we release preference annotations (designed for Bradley-Terry training) to complement existing ratings (designed for Regression style training) in the HelpSteer2 dataset. To improve data interpretability, preference annotations are accompanied with human-written justifications. Using this data, we conduct the first head-to-head comparison of Bradley-Terry and Regression models when adequately matched for data. Based on insights derived from such a comparison, we propose a novel approach to combine Bradley-Terry and Regression reward modeling. A Llama-3.1-70B-Instruct model tuned with this approach scores 94.1 on RewardBench, emerging top of more than 140 reward models as of 1 Oct 2024. This reward model can then be used with REINFORCE algorithm (RLHF) to align an Instruct model to reach 85.0 on Arena Hard, which is No. 1 as of 1 Oct 2024. We open-source this dataset (CC-BY-4.0 license) at this https URL -- 1-oct-2024 and openly release the trained Reward and Instruct models at this https URL and this https URL

Table of Contents

  • 1 Introduction
  • 2 Dataset
  • 3 Reward Model
  • 3.1 Evaluation
  • 3.2 Training
  • 4 Reward Model Results
  • 5 Aligned Models
  • 5.1 Evaluation
  • 5.2 Training
  • 5.3 Results
  • References
  • A Related Works
  • B Preference Ranking Guidelines
  • B.1 Ranking prioritization
  • B.2 Ranking strength
  • B.3 Ranking examples
  • B.3.1 Response 1 is slightly better than Response 2
  • B.3.2 Response 1 is better than Response 2
  • B.3.3 Response 1 is much better than Response 2
  • B.3.4 Neither response is valid
  • C Justification Pre-processing
  • D Preference Justification Analysis
  • E Training Hyper-parameters
  • F Compute requirements
  • G Rewardbench Response Distribution
  • H Aligned Model Evaluation Details
  • I Reward Curves
  • J Further analysis for Scaled Bradley Terry
  • K Reward Model Ablation Studies
  • L Aligned Model Ablation Studies
  • M Further REINFORCE Model Metrics

Knowls

  1. Knowl 1 — Scaled Bradley-Terry Loss Formulation

    equation

    The Scaled Bradley-Terry (Scaled BT) loss incorporates human-annotated preference magnitude m∈{1,2,3}m \in \{1, 2, 3\} (where 1 denotes slightly better, 2 denotes better, and 3 denotes much better) as a multiplicative scaling factor outside the log-sigmoid loss:

    LSBT(θ)=−mlog⁡(σ(rθ(x,yc)−rθ(x,yr)))\mathcal{L}_{\text{SBT}}(\theta) = - m \log \left(\sigma\left(r_\theta(x, y_c) - r_\theta(x, y_r)\right)\right)

    where xx denotes the prompt, ycy_c is the chosen response, yry_r is the rejected response, rθ(x,y)r_\theta(x, y) is the scalar reward predicted by the model parameterized by θ\theta, and σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the sigmoid function.

    In contrast to the Margin Bradley-Terry loss LMBT=−log⁡(σ(rθ(x,yc)−rθ(x,yr)−m))\mathcal{L}_{\text{MBT}} = -\log(\sigma(r_\theta(x, y_c) - r_\theta(x, y_r) - m)), which imposes a penalty on correct predictions whenever the predicted reward gap Δr=rθ(x,yc)−rθ(x,yr)\Delta r = r_\theta(x, y_c) - r_\theta(x, y_r) is less than mm, Scaled BT maintains a loss below 1 for any correctly ordered pair (Δr>0\Delta r > 0). Simultaneously, it applies substantially higher gradient penalties to severe prediction errors when the preference margin mm is large (for example, when Δr=−3\Delta r = -3 and m=3m=3, the loss is 9.14589.1458 under Scaled BT compared to 6.00256.0025 under Margin BT). Additionally, scaling by preference magnitude provides robustness against annotation noise, which occurs predominantly in weak preference pairs (m=1m=1).

  2. Knowl 2 — Two-Stage Reward Model Training with ExPO Extrapolation

    model/method

    A two-stage training framework integrates SteerLM regression pre-initialization and Bradley-Terry preference modeling, followed by weak-to-strong model extrapolation (ExPO):

    1. Stage 1 (Helpfulness Regression Pre-training): A pretrained language model (Llama-3.1-70B-Instruct) is trained with a linear regression head attached to the end-of-response token dense representation using mean squared error (MSE) loss to predict scalar helpfulness scores ∈[0,4]\in [0, 4] on HelpSteer2.
    2. Stage 2 (Scaled Bradley-Terry Fine-tuning): The weights of the Stage 1 helpfulness regression model are used to initialize a Bradley-Terry reward model. The model is fine-tuned for 1 epoch on HelpSteer2-Preference using the Scaled Bradley-Terry loss LSBT=−mlog⁡(σ(rθ(x,yc)−rθ(x,yr)))\mathcal{L}_{\text{SBT}} = -m \log(\sigma(r_\theta(x, y_c) - r_\theta(x, y_r))), where m∈{1,2,3}m \in \{1, 2, 3\}.
    3. Stage 3 (Weak-to-Strong Weight Extrapolation): Final model parameters WfinalW_{\text{final}} are derived by extrapolating the delta between the weak model (the Stage 1 Helpfulness Regression model WweakW_{\text{weak}}) and the strong model (the Stage 2 Scaled BT model WstrongW_{\text{strong}}):

    Wfinal=Wweak+α(Wstrong−Wweak)W_{\text{final}} = W_{\text{weak}} + \alpha \left(W_{\text{strong}} - W_{\text{weak}}\right)

    with an optimal extrapolation factor of α=1.52\alpha = 1.52 (located via grid search in [1.1,2.0][1.1, 2.0]).

    This pipeline achieves an overall score of 94.194.1 on RewardBench.

  3. Knowl 3 — RewardBench Benchmark Performance of Reward Model Architectures

    data/table

    RewardBench evaluates reward models across 2,985 tasks divided into four categories: Chat, Chat-Hard, Safety, and Reasoning. All tested architectures use Llama-3.1-70B-Instruct as the base model and are compared against external open and proprietary baselines:

    Model Type Model Overall Chat Chat-Hard Safety Reasoning
    SteerLM Regression HelpSteer Attributes (5 attributes) 92.4 95.0 85.5 94.0 95.1
    SteerLM Regression Helpfulness Only 93.0 97.2 84.2 94.6 95.8
    Bradley-Terry (from scratch) Regular BT 91.5 97.5 80.3 90.5 97.9
    Bradley-Terry (from scratch) Margin BT 91.5 98.0 78.5 94.6 94.8
    Bradley-Terry (from scratch) Scaled BT 92.7 97.8 83.5 93.2 96.0
    Bradley-Terry (init. Helpfulness Reg.) Regular BT 92.7 98.9 82.9 93.7 95.4
    Bradley-Terry (init. Helpfulness Reg.) Margin BT 93.0 98.3 83.8 94.0 95.8
    Bradley-Terry (init. Helpfulness Reg.) Scaled BT 93.7 98.0 85.7 94.3 96.7
    Bradley-Terry (init. Helpfulness Reg.) Scaled BT + ExPO 94.1 97.5 85.7 95.1 98.1
    Pairwise Justifier Full Preference Justification 88.1 94.4 82.2 90.9 84.9
    Pairwise Justifier Preference Statement Only (no elaboration) 88.4 96.4 78.7 93.4 85.3
    Pairwise Justifier Preference Label Only (Response 1 / 2) 90.0 96.1 80.0 93.1 90.9
    External Baseline Skywork-Reward-Gemma-2-27B 93.8 95.8 91.4 91.9 96.1
    External Baseline TextEval-Llama3.1-70B 93.5 94.1 90.1 93.2 96.4
    External Baseline Nemotron-4-340B-Reward 92.0 95.8 87.1 91.5 93.6
    External Baseline GPT-4o-2024-08-06 86.7 96.1 76.1 88.1 86.6
    External Baseline Claude-3.5-Sonnet-20240620 84.2 96.4 74.0 81.6 84.7
    External Baseline Meta-Llama-3.1-70B-Instruct 84.0 97.2 70.2 82.8 86.0

    The evaluation demonstrates that combining Helpfulness Regression initialization with Scaled Bradley-Terry training and ExPO extrapolation attains 94.194.1 overall accuracy, leading all tested baselines.

  4. Knowl 4 — HelpSteer2-Preference Dataset Construction and Quality Filtering

    experimental setup

    HelpSteer2-Preference is a human-annotated preference dataset collected on the exact prompt-response pairs of HelpSteer2. For each prompt and pair of model responses (Response 1 and Response 2), independent annotators (3–5 per prompt) assign a categorical preference rating:

    • ±1\pm 1: Response 2 / Response 1 is slightly better
    • ±2\pm 2: Response 2 / Response 1 is better
    • ±3\pm 3: Response 2 / Response 1 is much better
    • −100-100: Neither response is valid (used when both responses are egregiously flawed; excluded from preference pairs).

    Annotators also write free-form text justifications explaining their decisions.

    Data Curation and Filtering Pipeline:

    1. For each prompt, the three most mutually similar annotations are selected to mitigate outlier influence, averaged, and rounded to the nearest integer.
    2. Samples where the rating spread among the top 3 annotations exceeds 22 (10% of tasks) are removed as ambiguous.
    3. Samples with an aggregate rounded preference score of 00 (22% of tasks, indicating unresolved annotator disagreement) are removed to avoid low-confidence noise.

    The final curated dataset consists of 7,118 preference pairs (6,766 training, 352 validation). Quadratic-weighted Cohen's κ\kappa inter-rater reliability rises from 0.4890.489 on raw annotations to 0.8430.843 after spread filtering, and to 0.8780.878 after removing zero-preference ties (higher than the HelpSteer2 helpfulness rating agreement of 0.7910.791).

  5. Knowl 5 — Comparative Efficacy of Regression Versus Bradley-Terry Preference Modeling

    empirical result

    When evaluated on matched prompts and responses from HelpSteer2, a SteerLM regression model trained exclusively on helpfulness scores scores 93.093.0 overall on RewardBench, whereas a standalone Bradley-Terry model trained from scratch using Scaled BT loss scores 92.792.7.

    This demonstrates that when data quality and prompt-response sets are properly controlled and collected in the intended format, neither the regression paradigm nor the Bradley-Terry ranking paradigm holds an inherent performance advantage. Modeling effectiveness depends primarily on whether the training loss adequately leverages preference magnitude information. Furthermore, training on a single scalar helpfulness attribute outperforms training across five multi-attribute dimensions (helpfulness, correctness, coherence, complexity, verbosity, which scored 92.492.4) by eliminating multi-objective trade-offs and the need for heuristic attribute weighting.

  6. Knowl 6 — Online RLHF and Offline Alignment Benchmarks with HelpSteer2-Preference

    data/table

    Llama-3.1-70B-Instruct aligned using various preference optimization algorithms on HelpSteer2-Preference evaluated across GPT-4-Turbo MT-Bench, AlpacaEval 2.0 Length-Controlled (LC) win rate against GPT-4-Preview, and Arena Hard win rate against GPT-4-0314:

    Paradigm Algorithm MT-Bench Mean Chars AlpacaEval 2.0 LC (%) Arena Hard (%)
    Base Model Llama-3.1-70B-Instruct 8.22 1728.6 38.1 55.7
    Offline RLHF Regular DPO 8.66 1502.2 40.4 52.8
    Offline RLHF Margin DPO 8.58 1496.6 41.1 52.6
    Offline RLHF Scaled DPO 8.74 1514.8 41.0 52.9
    Online RLHF PPO (2 rounds) 8.74 1842.8 43.8 58.6
    Online RLHF REINFORCE (leave-one-out) 8.98 2199.8 57.6 85.0
    External Baseline Llama-3.1-405B-Instruct 8.49 1664.7 39.3 69.3
    External Baseline Claude-3.5-Sonnet-20240620 8.81 1619.9 52.4 79.2
    External Baseline GPT-4o-2024-05-13 8.74 1752.2 57.5 79.3

    Online RLHF using the Scaled BT + ExPO reward model strictly outperforms offline Direct Preference Optimization (DPO) variants. REINFORCE achieves the strongest overall alignment performance (8.988.98 MT-Bench, 57.6%57.6\% AlpacaEval 2.0 LC, and 85.0%85.0\% Arena Hard), outperforming proprietary frontier models including GPT-4o and Claude 3.5 Sonnet without exhibiting catastrophic forgetting on MMLU (82.4%), GSM8K (90.9%), or HumanEval (83.5%).

  7. Knowl 7 — Leave-One-Out REINFORCE Versus Critic-Based PPO for Online LLM Alignment

    empirical result

    In online RLHF utilizing the Scaled BT + ExPO reward model on HelpSteer2-Preference prompts, REINFORCE substantially outperforms two-round Proximal Policy Optimization (PPO), scoring higher on MT-Bench (8.988.98 vs. 8.748.74), AlpacaEval 2.0 LC (57.6%57.6\% vs. 43.8%43.8\%), and Arena Hard (85.0%85.0\% vs. 58.6%58.6\%).

    The performance divergence stems from baseline estimation methodology:

    • PPO: Relies on a learned neural critic network to estimate state values. Function approximation errors in the critic introduce bias and instability into advantage estimation, causing high sensitivity to hyperparameters and requiring value warmup phases.
    • REINFORCE: Uses an unbiased, stable Monte-Carlo baseline estimator via a leave-one-out formulation with K=4K=4 rollouts sampled per prompt. For response i∈{1,…,K}i \in \{1, \dots, K\} with KL-regularized reward RiR_i, the baseline subtracted is 1K−1∑j≠iRj\frac{1}{K-1} \sum_{j \neq i} R_j. This provides stable policy gradients without critic-induced approximation error, enabling superior reward optimization.
  8. Knowl 8 — Operational Profiles and Trade-offs of Reward Model Paradigms

    model/method

    Different reward modeling paradigms present distinct operational characteristics and suitable application domains:

    • Bradley-Terry Reward Models:

      • Characteristics: Highest comparative ranking accuracy (RewardBench 94.194.1). Scalar output values are uncalibrated and context-dependent (e.g., spanning [−33.50,2.654][-33.50, 2.654]), meaning an individual score cannot be interpreted without reference to alternative responses for the same prompt.
      • Inference Speed: Fast (single forward pass scoring the terminal token, equivalent to 1 generated token).
      • Primary Application: Reinforcement learning from human feedback (RLHF) policy optimization and relative data filtering.
    • SteerLM Regression Models:

      • Characteristics: Highly calibrated outputs bounded within the defined training score interval ([0,4][0, 4]) where values correspond to direct semantic quality thresholds (e.g., 4=perfect4 = \text{perfect}, 0=poor0 = \text{poor}). Capable of predicting multiple independent response dimensions (helpfulness, correctness, verbosity, complexity, coherence).
      • Inference Speed: Fast (1 token compute equivalent).
      • Primary Application: Absolute threshold filtering and relative ranking for Supervised Fine-Tuning (SFT) data selection.
    • Pairwise Justifier Models (Generative Judges):

      • Characteristics: Provides natural language rationales detailing comparative strengths and weaknesses. However, it achieves lower ranking accuracy (88.188.1–90.090.0 on RewardBench).
      • Inference Speed: Slow (requires auto-regressive generation of hundreds of tokens per comparison).
      • Primary Application: Human-in-the-loop validation and audit workflows.
  9. Knowl 9 — Empirical Performance of Generative Pairwise Justifier Reward Models

    empirical result

    Pairwise Justifier models are trained via supervised cross-entropy fine-tuning to generate a comparative rationale followed by a structured decision string @Response 1/2 ... better conditioned on the prompt and both responses.

    Empirical findings on RewardBench include:

    • Task Formulation Complexity: Full Pairwise Justifier models achieve an overall accuracy of 88.188.1, trailing scalar SteerLM regression (93.093.0) and Bradley-Terry models (92.792.7). Generating natural language explanations forces the model to solve a composite task (scoring Response 1, scoring Response 2, and verbalizing the difference), which is more difficult to optimize than direct scalar scoring.
    • Multi-Rationale Training: Providing up to three distinct human justifications per training task increases overall accuracy from 86.186.1 to 88.188.1, primarily via gains in the Chat-Hard category (71.7→82.271.7 \to 82.2).
    • Label-Only Optimization: Training the generative model to predict only the preference label (Response 1 or Response 2) without generating human-written explanatory text achieves higher overall accuracy (90.090.0) than generating full human rationales (88.188.1) or statement-only justifications (88.488.4). Unstructured human explanations introduce training noise that impedes learning optimal latent comparative representations.
  10. Knowl 10 — Ablation Studies on Reward Model Initialization, Curation, and Data Conversion

    empirical result

    Ablation experiments conducted on the 70B Scaled Bradley-Terry model initialized with Helpfulness Regression isolate the contribution of each design component:

    • Purpose-Built Pairwise Data vs. Converted Ratings: Fine-tuning the helpfulness regression model with Scaled BT on purpose-built HelpSteer2-Preference reaches 93.793.7 RewardBench accuracy. In contrast, fine-tuning with Scaled BT on HelpSteer2 single-response ratings converted post-hoc into pairwise differences yields 92.992.9 (failing to improve over the 93.093.0 regression starting point). Direct pairwise collection captures nuanced distinctions (e.g., resolving ties where both responses received identical Likert scores) that converted rating differences miss.
    • Annotation Spread Filtering: Retaining the 10% most ambiguous tasks (annotation spread >2> 2) degrades RewardBench accuracy by 1.01.0 point (93.7→92.793.7 \to 92.7). Averaging across all annotations per task instead of selecting the three most mutually similar annotations drops performance further to 92.392.3.
    • Reversed Model Initialization (Regression Initialized from BT): Initializing a Helpfulness Regression model from a pretrained Scaled BT model results in a degraded score of 92.292.2 on RewardBench (worse than training regression from scratch at 93.093.0). The uncalibrated, unbounded range of Bradley-Terry logits ([−35,5][-35, 5]) causes massive initial MSE loss gradients against [0,4][0, 4] target scores.

Coverage note — None was omitted; all key contributions including dataset collection and curation, Scaled Bradley-Terry loss, two-stage reward modeling with ExPO, alignment with REINFORCE/PPO/DPO, and experimental comparisons across paradigms are captured.

References

  1. 1.Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Ahmet Üstün, and Sara Hooker. Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms. arXiv preprint arXiv:2402.14740, 2024.
  2. 2.AllenAI. Reward bench leaderboard. https://huggingface.co/spaces/allenai/reward-bench, 2024.
  3. 3.Duane F. Alwin and Jon A. Krosnick. The Measurement of Values in Surveys: A Comparison of Ratings and Rankings. Public Opinion Quarterly, 49(4):535–552, 01 1985. ISSN 0033-362X. doi: 10.1086/268949. URL https://doi.org/10.1086/268949.
  4. 4.Ron Artstein and Massimo Poesio. Survey article: Inter-coder agreement for computational linguistics. Computational Linguistics, 34(4):555–596, 2008. doi: 10.1162/coli.07-034-R2. URL https://aclanthology.org/J08-4004.
  5. 5.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022.
  6. 6.Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating llms by human preference, 2024.
  7. 7.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. Ultrafeedback: Boosting language models with high-quality feedback, 2023.
  8. 8.Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. RLHF workflow: From reward modeling to online RLHF, 2024.
  9. 9.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaoqing Ellen Tan, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aaron Grattafiori, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alex Vaughan, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Franco, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, Danny Wyatt, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Firat Ozgenel, Francesco Caggioni, Francisco Guzmán, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Govind Thattai, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Karthik Prasad, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kun Huang, Kunal Chawla, Kushal Lakhotia, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Maria Tsimpoukelli, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikolay Pavlovich Laptev, Ning Dong, Ning Zhang, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Rohan Maheswari, Russ Howes, Ruty Rinott, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Kohler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vítor Albiero, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaofang Wang, Xiaojian Wu, Xiaolan Wang, Xide Xia, Xilun Wu, Xinbo Gao, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yuchen Hao, Yundi Qian, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, and Zhiwei Zhao. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  10. 10.Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto. Length-controlled alpacaeval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475, 2024.
  11. 11.Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. Understanding dataset difficulty with V-usable information. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 5988–6008. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/ethayarajh22a.html.
  12. 12.Wouter Kool, Herke van Hoof, and Max Welling. Buy 4 REINFORCE samples, get a baseline for free!, 2019. URL https://openreview.net/forum?id=r1lgTGL5DE.
  13. 13.Aviral Kumar, Abhishek Gupta, and Sergey Levine. Discor: Corrective feedback in reinforcement learning via distribution correction. Advances in Neural Information Processing Systems, 33: 18560–18572, 2020.
  14. 14.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. Openassistant conversations – democratizing large language model alignment, 2023.
  15. 15.Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. Rewardbench: Evaluating reward models for language modeling, 2024.
  16. 16.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From live data to high-quality benchmarks: The Arena-Hard pipeline. https://lmsys.org/blog/2024-04-19-arena-hard/, April 2024.
  17. 17.Chris Yuhao Liu and Liang Zeng. Skywork reward model. https://huggingface.co/Skywork/Skywork-Reward-Gemma-2-27B, September 2024. URL https://huggingface.co/Skywork/Skywork-Reward-Gemma-2-27B.
  18. 18.LMSys. Arena-hard-auto leaderboard. https://github.com/lm-sys/arena-hard-auto, 2024.
  19. 19.John A. McCarty and L. J. Shrum. The Measurement of Personal Values in Survey Research: A Test of Alternative Rating Procedures*. Public Opinion Quarterly, 64(3):271–298, 11 2000. ISSN 0033-362X. doi: 10.1086/317989. URL https://doi.org/10.1086/317989.
  20. 20.Yu Meng, Mengzhou Xia, and Danqi Chen. SimPO: Simple preference optimization with a reference-free reward, 2024.
  21. 21.Meta AI. Llama 3 model card. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md, 2024.
  22. 22.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. Webgpt: Browser-assisted question-answering with human feedback. In arXiv, 2021.
  23. 23.NLTK. nltk.tokenize.sent-tokenize. https://www.nltk.org/api/nltk.tokenize.sent_tokenize.html, 2024.
  24. 24.Nvidia, :, Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H. Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, Sirshak Das, Ayush Dattagupta, Olivier Delalleau, Leon Derczynski, Yi Dong, Daniel Egert, Ellie Evans, Aleksander Ficek, Denys Fridman, Shaona Ghosh, Boris Ginsburg, Igor Gitman, Tomasz Grzegorzek, Robert Hero, Jining Huang, Vibhu Jawa, Joseph Jennings, Aastha Jhunjhunwala, John Kamalu, Sadaf Khan, Oleksii Kuchaiev, Patrick LeGresley, Hui Li, Jiwei Liu, Zihan Liu, Eileen Long, Ameya Sunil Mahabaleshwarkar, Somshubra Majumdar, James Maki, Miguel Martinez, Maer Rodrigues de Melo, Ivan Moshkov, Deepak Narayanan, Sean Narenthiran, Jesus Navarro, Phong Nguyen, Osvald Nitski, Vahid Noroozi, Guruprasad Nutheti, Christopher Parisien, Jupinder Parmar, Mostofa Patwary, Krzysztof Pawelec, Wei Ping, Shrimai Prabhumoye, Rajarshi Roy, Trisha Saar, Vasanth Rao Naik Sabavat, Sanjeev Satheesh, Jane Polak Scowcroft, Jason Sewall, Pavel Shamis, Gerald Shen, Mohammad Shoeybi, Dave Sizer, Misha Smelyanskiy, Felipe Soares, Makesh Narsimhan Sreedhar, Dan Su, Sandeep Subramanian, Shengyang Sun, Shubham Toshniwal, Hao Wang, Zhilin Wang, Jiaxuan You, Jiaqi Zeng, Jimmy Zhang, Jing Zhang, Vivienne Zhang, Yian Zhang, and Chen Zhu. Nemotron-4 340b technical report, 2024. URL https://arxiv.org/abs/2406.11704.
  25. 25.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022.
  26. 26.Junsoo Park, Seungyeon Jwa, Meiying Ren, Daeyoung Kim, and Sanghyuk Choi. Offsetbias: Leveraging debiased data for tuning evaluators, 2024.
  27. 27.Gloria Phillips-Wren, Daniel J. Power, and Manuel Mora. Cognitive bias, decision styles, and risk attitudes in decision making and dss. Journal of Decision Systems, 28(2):63–66, April 2019. ISSN 2116-7052. doi: 10.1080/12460125.2019.1646509. URL http://dx.doi.org/10.1080/12460125.2019.1646509.
  28. 28.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2023.
  29. 29.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  30. 30.Scikit-Learn. Cohen kappa score. https://scikit-learn.org/stable/modules/generated/sklearn.metrics.cohen_kappa_score.html, 2024.
  31. 31.Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models, 2023. URL https://arxiv.org/abs/2310.13548.
  32. 32.Judy Hanwen Shen, Archit Sharma, and Jun Qin. Towards data-centric rlhf: Simple metrics for preference dataset comparison, 2024. URL https://arxiv.org/abs/2409.09603.
  33. 33.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback, 2022.
  34. 34.Richard S Sutton. Reinforcement learning: An introduction. A Bradford Book, 2018.
  35. 35.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  36. 36.Tatsu-Lab. Alpacaeval leaderboard. https://tatsu-lab.github.io/alpaca_eval/, 2023.
  37. 37.Tenyx. Tenyxchat: Language model alignment using tenyx fine-tuning, 2024.
  38. 38.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023.
  39. 39.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. Zephyr: Direct distillation of lm alignment, 2023.
  40. 40.Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, and Tong Zhang. Arithmetic control of llms for diverse user preferences: Directional preference alignment with multi-objective rewards, 2024a. URL https://arxiv.org/abs/2402.18571.
  41. 41.Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. Interpretable preferences via multi-objective reward modeling and mixture-of-experts, 2024b. URL https://arxiv.org/abs/2406.12845.
  42. 42.Peifeng Wang, Austin Xu, Yilun Zhou, Caiming Xiong, and Shafiq Joty. Direct judgement preference optimization, 2024c. URL https://arxiv.org/abs/2409.14664.
  43. 43.Zhilin Wang, Yi Dong, Olivier Delalleau, Jiaqi Zeng, Gerald Shen, Daniel Egert, Jimmy Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev. Helpsteer 2: Open-source dataset for training top-performing reward models. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 1474–1501. Curran Associates, Inc., 2024d. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/02fd91a387a6a5a5751e81b58a75af90-Paper-Datasets_and_Benchmarks_Track.pdf.
  44. 44.Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev. HelpSteer: Multi-attribute helpfulness dataset for SteerLM. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 3371–3384, Mexico City, Mexico, June 2024e. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-long.185. URL https://aclanthology.org/2024.naacl-long.185/.
  45. 45.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8:229–256, 1992.
  46. 46.Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun. Advancing LLM reasoning generalists with preference trees, 2024.
  47. 47.Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. Evaluating large language models at evaluating instruction following, 2024. URL https://arxiv.org/abs/2310.07641.
  48. 48.Chujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang, and Nanyun Peng. Weak-to-strong extrapolation expedites alignment, 2024. URL https://arxiv.org/abs/2404.16792.
  49. 49.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023.
  50. 50.Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao. Starling-7b: Improving LLM helpfulness & harmlessness with RLAIF, November 2023.
  51. 51.Banghua Zhu, Michael I. Jordan, and Jiantao Jiao. Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf, 2024. URL https://arxiv.org/abs/2401.16335.

Citation

MLA
Wang, Z., et al. “HelpSteer2-Preference: Complementing Ratings with Preferences”. arXiv, 2024, http://arxiv.org/abs/2410.01257v2.
APA
Wang, Z., Bukharin, A., Delalleau, O., Egert, D., Shen, G., Zeng, J., Kuchaiev, O., & Dong, Y. (2024). HelpSteer2-Preference: Complementing Ratings with Preferences. arXiv. http://arxiv.org/abs/2410.01257v2
Chicago
Wang, Z., A. Bukharin, O. Delalleau, et al. 2024. “HelpSteer2-Preference: Complementing Ratings with Preferences”. arXiv. http://arxiv.org/abs/2410.01257v2.
Harvard
Wang, Z. et al. (2024) “HelpSteer2-Preference: Complementing Ratings with Preferences”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.01257v2.
Vancouver
1. Wang Z, Bukharin A, Delalleau O, Egert D, Shen G, Zeng J, Kuchaiev O, Dong Y (2024) HelpSteer2-Preference: Complementing Ratings with Preferences. arXiv

BibTeX

@article{wang2024helpsteer2,
  title = {HelpSteer2-Preference: Complementing Ratings with Preferences},
  author = {Wang, Zhilin and Bukharin, Alexander and Delalleau, Olivier and Egert, Daniel and Shen, Gerald and Zeng, Jiaqi and Kuchaiev, Oleksii and Dong, Yi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.01257v2},
  eprint = {2410.01257}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors