Reward Models Inherit Value Biases from Pretraining
Brian ChristianJessica A.F. ThompsonElleVincent AdamHannah Rose KirkChristopher SummerfieldTsvetomira Dumbalska
Demonstrates that reward models inherit persistent value biases directly from their pretrained base language models, causing architectures like Llama and Gemma to diverge systematically across psychological dimensions regardless of identical preference training.
Reward models are critical components in the artificial intelligence alignment pipeline, used to evaluate outputs and steer large language models toward human preferences and safety standards. However, reward models are themselves initialized from pretrained base models before undergoing preference finetuning. Because pretraining data outstrips fine-tuning data by orders of magnitude, there is a risk that foundational value biases embedded during pretraining remain inside reward models, inadvertently dictating downstream model behavior.
The article evaluates whether open-weight reward models inherit systematic value biases from their underlying base models and investigates how these biases evolve during preference training. The authors demonstrate that reward models reflect distinct philosophical and moral orientations rooted directly in their pretraining, rather than acting as neutral evaluators of human preference data.
To conduct this evaluation, the researchers combined an exhaustive token vocabulary search across 54 value-laden prompts with validated psycholinguistic dictionaries, focusing on the "Big Two" psychological dimensions—agency (individual goals and freedom) and communion (relationships and social connection)—as well as the Moral Foundations Theory framework. They analyzed 10 leading public reward models based on Llama and Gemma architectures from the RewardBench benchmark. They further examined the underlying instruction-tuned and raw pretrained base models, derived an implicit reward model score to measure token-level preference differences across model families from 1 billion to 70 billion parameters, and trained custom reward models across various data sources and dataset sizes (ranging from 13,000 to over 100,000 preference pairs) while holding hyperparameters constant.
The analysis yielded several key findings regarding the persistence of inherited values:
- Base model choice drives systematic value splits: Llama-based reward models consistently assign higher rewards to agency-oriented concepts (such as freedom and capability), whereas Gemma-based reward models consistently prioritize communion-oriented concepts (such as love, family, and connection). In top-token rankings, communion tokens occupied an average of 5 out of the top 10 positions for Gemma models compared to zero agency tokens, whereas Llama models averaged 2.33 agency tokens in the top 10.
- Pretraining is the definitive source of bias: Identical agency-versus-communion splits appear in the raw log probabilities of both instruction-tuned and unaligned, pretrained base models. Across 21 cross-family model comparisons, the relative reward score for "Freedom" versus "Love" was consistently positive, with this gap widening as model size increased from 1B to 70B parameters.
- Preference finetuning narrows but fails to eliminate biases: In controlled training experiments, the preference gap between Llama and Gemma reward models stabilized within the first third of training. While scaling preference data to approximately 100,000 pairs attenuated some differences, models retaining underlying generative states retained strong biases even after training on over 630,000 preferences. Exploratory tests on Qwen-based models revealed communion biases that did not close even at maximum data scales.
These findings indicate that reward models are not neutral arbiters of human feedback; they inherit moral and conceptual orientations from pretraining that directly influence downstream alignment. For organizations developing or deploying language models, selecting an open-weight base model represents a fundamental choice of values rather than just a performance benchmark decision. Post-training reinforcement learning cannot simply overwrite deep pretraining representations, meaning current alignment methods risk leaving systematic biases unaddressed.
Organizations developing aligned models should treat pretraining data filtering, curation, and auditing as critical safety and alignment stages rather than relying solely on post-training corrections. When building reward models, teams should evaluate multiple base model families to understand how inherited biases alter application-specific outcomes, and consider collecting larger, multi-dimensional preference datasets to mitigate baseline distortions. Further research is necessary to map biases across a wider array of model architectures and develop formal scaling laws governing the amount of preference data required to overcome pretraining priors.
These conclusions carry high confidence based on consistent findings across both public and controlled in-house models, multiple model sizes, and diverse prompt phrasings. However, readers should consider specific boundary conditions: the empirical training experiments focused primarily on 2-billion to 3-billion parameter models and evaluated a constrained set of psycholinguistic dimensions. Further empirical validation is warranted before generalizing these dynamics to longer multi-token generations and broader multi-dimensional value spaces.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Provides the foundational methodology for training reward models from human comparisons to fine-tune language models, establishing the standard paradigm investigated by the source.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Introduces standard practices for training preference and reward models for helpful and harmless AI assistants, providing essential context on how reward models are structured and tuned.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). Demonstrates how base and feedback-tuned language models inherently reflect distinct human value distributions and opinions, establishing empirical groundwork for evaluating value biases in models.
- Paper: Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models, Paul Röttger et al. (2024). Critically examines the psychometric and opinion evaluation frameworks in language models, informing the methodology needed to measure psychological value dimensions like agency and communion.
- Paper: HelpSteer2-Preference: Complementing Ratings with Preferences, Zhilin Wang et al. (2025). Examines reward model training architectures and benchmarks, giving concrete insight into how preference datasets interact with reward model initialization.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). Analyzes how downstream alignment pipelines skew global and value representations learned during pretraining, directly preceding the source's investigation into reward model bias inheritance.
- Paper: Demystifying Reinforcement Learning Post-Training of Language Models, Donovan Clay et al. (2026). Extends the findings on inherited base model biases by dissecting how the base model's initial probability distribution interacts with reward signals during RL post-training.
- Paper: Understanding Reasoning from Pretraining to Post-Training, Jingyan Shen et al. (2026). Provides a mechanistic continuation by analyzing how specific pretraining dynamics mathematically shape and constrain subsequent post-training and RL outcomes.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). Builds upon implicit and explicit reward modeling dynamics to unify preference learning methods across varied reference and base model constraints.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). Applies multi-reward optimization techniques that address competing reward signals, extending the implications of multidimensional value biases in reward modeling.
