HelpSteer2-Preference: Complementing Ratings with Preferences
Zhilin WangAlexander BukharinOlivier DelalleauDaniel EgertGerald ShenJiaqi ZengOleksii KuchaievYi Dong
Presents a method combining Bradley-Terry preferences and regression ratings using the open-source HelpSteer2-Preference dataset, achieving top performance on RewardBench and Arena-Hard for language model alignment.
Aligning large language models with human expectations requires effective reward models that score AI responses for helpfulness and safety. Historically, developers have used two competing training paradigms: Bradley-Terry models, which evaluate comparative preference between pairs of responses, and regression models, which assign absolute ratings to individual responses. A lack of directly comparable, high-quality data has created an ongoing debate over which training paradigm performs better.
The article aims to evaluate the relative strengths of Bradley-Terry and regression reward modeling approaches when trained on identical, purpose-built data, and to demonstrate how combining these methods improves language model alignment.
To conduct this evaluation, the researchers open-sourced HelpSteer2-Preference, a dataset containing pairwise human preferences, preference strengths, and written justifications alongside existing absolute ratings across 7,118 response pairs. The team trained several 70-billion-parameter reward models using Bradley-Terry, regression, and generative pairwise justification approaches, benchmarking them across 2,985 diverse tasks on RewardBench. They then integrated the top-performing reward model into reinforcement learning workflows to align instruction-following policy models.
The head-to-head comparison revealed four central findings. First, when matched on high-quality data, standard regression and Bradley-Terry reward models achieve virtually identical benchmark accuracy (93.0% and 92.7%, respectively). Second, combining both paradigms into a two-stage training pipeline—initializing a Bradley-Terry model with regression weights and applying weight extrapolation—achieves state-of-the-art accuracy of 94.1%, outperforming all public baselines. Third, reward models that generate text justifications perform worse (under 90.0% accuracy) than direct scoring models because the evaluation task is more complex. Fourth, in downstream reinforcement learning alignment, using this hybrid reward model within the REINFORCE framework achieved an 85.0 score on Arena Hard, outperforming traditional Direct Preference Optimization (52.9) and Proximal Policy Optimization (58.6) while maintaining core reasoning abilities.
These findings indicate that the format of data collection is less important than how effectively training objectives capture preference strength. The results demonstrate that two-stage reward modeling pipelines provide superior alignment guidance without compromising core model capabilities. For engineering trade-offs, regression models remain best suited for rapid, interpretable data filtering, while Bradley-Terry models are optimal for reinforcement learning.
Organizations developing frontier language models should adopt two-stage reward modeling and utilize online REINFORCE over offline Direct Preference Optimization where compute budgets allow. Further work should explore larger datasets across specialized domains, as well as investigate structured formatting for preference justifications.
The conclusions are limited by the small, general-domain dataset size (approximately 6,766 training pairs) and benchmark evaluations that rely heavily on automated judges powered by commercial language models. Nevertheless, given the consistent cross-benchmark gains and rigorous ablation studies, confidence in the primary findings remains high.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). This foundational work introduces Direct Preference Optimization (DPO) and Bradley-Terry preference modeling, against which HelpSteer2-Preference directly compares and benchmarks its two-stage regression and REINFORCE alignment pipeline.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This seminal work establishes the foundational paradigm of training reward models on human preference comparisons to steer language models using reinforcement learning.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). It details the standard modern pipeline for collecting pairwise comparisons and fitting Bradley-Terry reward models for policy optimization.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It provides essential background on multi-objective preference modeling and iterative RLHF for balancing helpfulness and harmlessness.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). It analyzes the dynamics and failure modes of reward model optimization, providing key context for why precise scoring and reward modeling pipelines matter in RLHF.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). It introduces the Chatbot Arena evaluation methodology and Bradley-Terry win-rate ranking that underpin the downstream Arena-Hard benchmarks evaluated in the source.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). It demonstrates how multi-dimensional feedback criteria can be separated into dedicated reward models to guide alignment safely.
- Paper: KTO: Model Alignment as Prospect Theoretic Optimization, Kawin Ethayarajh et al. (2024). It explores alternatives to pairwise preference optimization using utility functions on individual outputs, directly informing the regression-versus-preference debate.
- Paper: Self-Generated Critiques Boost Reward Modeling for Language Models, Yue Yu 0009 et al. (2025). This work directly extends reward modeling architectures by developing self-generated critiques and refinement schedules to overcome the performance limitations of generative evaluators identified in HelpSteer2-Preference.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). This volume provides a comprehensive theoretical unification of preference objectives and reinforcement learning versus direct alignment, synthesizing empirical trade-offs such as those demonstrated in the source.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). It advances policy optimization beyond standard REINFORCE and GRPO by decoupling multi-reward signals and normalization across distinct preference attributes.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). It builds upon reward-guided post-training to systematically analyze the exact mechanistic interactions between reward signal density, prompt distributions, and policy exploration.
- Paper: Efficient Exploration at Scale, Seyed Mohammad Asghari et al. (2026). It extends online reinforcement learning pipelines with uncertainty-directed exploration strategies to dramatically improve the sample efficiency of preference learning.
- Paper: Theoretical guarantees on the best-of-n alignment policy, Ahmad Beirami et al. (2025). It provides rigorous theoretical drift guarantees and closed-form bounds when applying reward models in best-of-n sampling and rejection policies.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It applies preference reward estimators dynamically during decoding-time heuristic search rather than relying solely on post-training policy optimization.
