Reward Models Inherit Value Biases from Pretraining

Brian ChristianJessica A.F. ThompsonElleVincent AdamHannah Rose KirkChristopher SummerfieldTsvetomira Dumbalska

article2026arXiv4 citations

Demonstrates that reward models inherit persistent value biases directly from their pretrained base language models, causing architectures like Llama and Gemma to diverge systematically across psychological dimensions regardless of identical preference training.

Listen

Reward models are critical components in the artificial intelligence alignment pipeline, used to evaluate outputs and steer large language models toward human preferences and safety standards. However, reward models are themselves initialized from pretrained base models before undergoing preference finetuning. Because pretraining data outstrips fine-tuning data by orders of magnitude, there is a risk that foundational value biases embedded during pretraining remain inside reward models, inadvertently dictating downstream model behavior.

The article evaluates whether open-weight reward models inherit systematic value biases from their underlying base models and investigates how these biases evolve during preference training. The authors demonstrate that reward models reflect distinct philosophical and moral orientations rooted directly in their pretraining, rather than acting as neutral evaluators of human preference data.

To conduct this evaluation, the researchers combined an exhaustive token vocabulary search across 54 value-laden prompts with validated psycholinguistic dictionaries, focusing on the "Big Two" psychological dimensions—agency (individual goals and freedom) and communion (relationships and social connection)—as well as the Moral Foundations Theory framework. They analyzed 10 leading public reward models based on Llama and Gemma architectures from the RewardBench benchmark. They further examined the underlying instruction-tuned and raw pretrained base models, derived an implicit reward model score to measure token-level preference differences across model families from 1 billion to 70 billion parameters, and trained custom reward models across various data sources and dataset sizes (ranging from 13,000 to over 100,000 preference pairs) while holding hyperparameters constant.

The analysis yielded several key findings regarding the persistence of inherited values:

  • Base model choice drives systematic value splits: Llama-based reward models consistently assign higher rewards to agency-oriented concepts (such as freedom and capability), whereas Gemma-based reward models consistently prioritize communion-oriented concepts (such as love, family, and connection). In top-token rankings, communion tokens occupied an average of 5 out of the top 10 positions for Gemma models compared to zero agency tokens, whereas Llama models averaged 2.33 agency tokens in the top 10.
  • Pretraining is the definitive source of bias: Identical agency-versus-communion splits appear in the raw log probabilities of both instruction-tuned and unaligned, pretrained base models. Across 21 cross-family model comparisons, the relative reward score for "Freedom" versus "Love" was consistently positive, with this gap widening as model size increased from 1B to 70B parameters.
  • Preference finetuning narrows but fails to eliminate biases: In controlled training experiments, the preference gap between Llama and Gemma reward models stabilized within the first third of training. While scaling preference data to approximately 100,000 pairs attenuated some differences, models retaining underlying generative states retained strong biases even after training on over 630,000 preferences. Exploratory tests on Qwen-based models revealed communion biases that did not close even at maximum data scales.

These findings indicate that reward models are not neutral arbiters of human feedback; they inherit moral and conceptual orientations from pretraining that directly influence downstream alignment. For organizations developing or deploying language models, selecting an open-weight base model represents a fundamental choice of values rather than just a performance benchmark decision. Post-training reinforcement learning cannot simply overwrite deep pretraining representations, meaning current alignment methods risk leaving systematic biases unaddressed.

Organizations developing aligned models should treat pretraining data filtering, curation, and auditing as critical safety and alignment stages rather than relying solely on post-training corrections. When building reward models, teams should evaluate multiple base model families to understand how inherited biases alter application-specific outcomes, and consider collecting larger, multi-dimensional preference datasets to mitigate baseline distortions. Further research is necessary to map biases across a wider array of model architectures and develop formal scaling laws governing the amount of preference data required to overcome pretraining priors.

These conclusions carry high confidence based on consistent findings across both public and controlled in-house models, multiple model sizes, and diverse prompt phrasings. However, readers should consider specific boundary conditions: the empirical training experiments focused primarily on 2-billion to 3-billion parameter models and evaluated a constrained set of psycholinguistic dimensions. Further empirical validation is warranted before generalizing these dynamics to longer multi-token generations and broader multi-dimensional value spaces.

Cover for Reward Models Inherit Value Biases from Pretraining

Abstract

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Because RMs are initialized from LLMs, they inherit representations that shape their behavior, but the nature and extent of this influence remain understudied. In a comprehensive study of 10 leading open-weight RMs using validated psycholinguistic corpora, we show that RMs exhibit significant differences along multiple dimensions of human value as a function of their base model. Using the "Big Two" psychological axes, we show a robust preference of Llama RMs for "agency" and a corresponding robust preference of Gemma RMs for "communion." This phenomenon holds even when the preference data and finetuning process are identical, and we trace it back to the logits of the respective instruction-tuned and pretrained models. These log-probability differences themselves can be formulated as an implicit RM; we derive usable implicit reward scores and show that they exhibit the very same agency/communion difference. We run experiments training RMs with ablations for preference data source and quantity, which demonstrate that this effect is not only repeatable but surprisingly durable. Despite RMs being designed to represent human preferences, our evidence shows that their outputs are influenced by the pretrained LLMs on which they are based. This work underscores the importance of safety and alignment efforts at the pretraining stage, and makes clear that open-source developers' choice of base model is as much a consideration of values as of performance.

Table of Contents

  • 1 Introduction
  • 2 RMs in the Wild Show Value Differences by Base Model
  • 3 Value Biases Begin in Pretraining
  • 3.1 Log Probabilities Mirror RM Agency/Communion Biases
  • 3.2 Implicit Reward Scores Mirror RM Agency/Communion Biases
  • 4 Dynamics of Inherited Values Over the Course of RM Training
  • 4.1 Experimental Setup
  • 4.2 Results
  • 5 Related Work
  • 6 Limitations & Conclusion
  • References
  • A RewardBench Models Studied
  • B Psycholinguistic Approach: Big Two and MFD2
  • C Value Preferences from 10 Leading RMs Based on Gemma and Llama: Big Two and MFD2
  • D Re-analysis of Christian et al. (2025)’s Exhaustive Token Search
  • E Prompt Construction
  • F Validating Implicit Reward Measures
  • F.1 Candidate Measures and Validation
  • F.2 Implicit Reward Comparisons Across Model Families
  • G RM Training Dynamics
  • G.1 Kendall τ\tau Correlation
  • H Value Biases of Qwen
  • H.1 Preference Changes Over Training
  • I LLM Usage Statement

Knowls

  1. Knowl 1 — Agency vs. Communion Value Split in Open-Weight Reward Models

    empirical result

    Analysis of 10 leading open-source reward models (RMs) from RewardBench (including Skywork-Reward, QRM, GRM, ArmoRM, and URM) demonstrates that RMs inherit systematic value biases directly from their base LLM families (Llama vs. Gemma):

    • Big Two Value Dimensions: When evaluated on 27 positively framed superlative prompts (e.g., "What, in one word, is the greatest thing ever?"), Llama-based RMs consistently assign higher reward ranks to agency-related words (e.g., success, skills, capability, freedom), whereas Gemma-based RMs assign higher reward ranks to communion-related words (e.g., love, friends, relationships). On 27 negatively framed prompts (e.g., "the worst thing ever"), the preference pattern reverses, yielding a statistically significant three-way interaction between Big Two category, base model, and prompt valence (p<0.001p < 0.001, follow-up permutation tt-tests p<0.01p < 0.01, Cohen's d=0.40–0.43d = 0.40\text{--}0.43).

    • Downstream Top-kk Vocabulary Bias: In a top-10 reward token analysis across the intersection of model vocabularies, Gemma RMs place on average 5.0 Communion tokens and 0.0 Agency tokens in their top-10 lists, whereas Llama RMs place on average 3.67 Communion tokens and 2.33 Agency tokens.

    • Moral Foundations (MFD2): On positively framed prompts, Llama RMs favor authority and fairness words, whereas Gemma RMs favor care, loyalty, and sanctity words (permutation tt-tests, all p<0.001p < 0.001).

  2. Knowl 2 — Pretraining and Instruction-Tuned Log-Probabilities Mirror RM Value Biases

    empirical result

    The agency versus communion bias observed in reward models is directly traceable to the next-token probability distributions of the underlying language models prior to reward finetuning.

    Evaluating the median ranking of 82 Big Two nouns shared across tokenizers on 54 evaluative prompts shows:

    • Instruction-Tuned Models: Gemma 2 2B IT and Llama 3.2 3B Instruct display a significant three-way interaction between Big Two category, prompt valence, and model (F(1,208)=58.3F(1, 208) = 58.3, p<0.001p < 0.001). On positive prompts, Llama assigns higher log probabilities to agency nouns while Gemma assigns higher log probabilities to communion nouns, reversing on negative prompts.

    • Pretrained Base Models: Raw base models without instruction tuning (Gemma 2 2B base and Llama 3.2 3B base) exhibit the identical three-way interaction (F(1,208)=43.2F(1, 208) = 43.2, p<0.001p < 0.001; all relevant Welch's tt-tests FDR-corrected p<0.01p < 0.01).

    This demonstrates that value preferences in aligned reward models originate in pretraining representations rather than being introduced during preference finetuning.

  3. Knowl 3 — Mixture-Weighted Log-Ratio for Implicit Reward Modeling

    model/method

    In inverse reinforcement learning and Direct Preference Optimization (DPO), the relative implicit reward that transforms a source model distribution π1(y∣x)\pi_1(y \mid x) into a target model distribution π2(y∣x)\pi_2(y \mid x) is given by log⁡π2(y∣x)−log⁡π1(y∣x)\log \pi_2(y \mid x) - \log \pi_1(y \mid x). However, computing raw log-probability differences over full token vocabularies causes extreme spurious scores in the low-probability tail ("junk" tokens).

    To make implicit reward comparisons robust across LLM vocabularies while maintaining antisymmetry, the Mixture-Weighted Log-Ratio (MWLR) is defined as:

    MWLR(y∣x)=12(p(y)+q(y))(log⁡q(y)−log⁡p(y))\text{MWLR}(y \mid x) = \frac{1}{2}\big(p(y) + q(y)\big)\big(\log q(y) - \log p(y)\big)

    where p(y)≡π1(y∣x)p(y) \equiv \pi_1(y \mid x) and q(y)≡π2(y∣x)q(y) \equiv \pi_2(y \mid x) are the conditional probability distributions over token yy given prompt xx.

    MWLR has key structural properties:

    1. Relevance Weighting: The mixture probability mass 12(p+q)\frac{1}{2}(p + q) forces the metric toward zero for tokens to which neither model assigns non-negligible probability.
    2. Antisymmetry: Swapping the source and target models negates the score (MWLR2→1=−MWLR1→2\text{MWLR}_{2 \to 1} = -\text{MWLR}_{1 \to 2}), preserving the exact token preference ordering in reverse without directional bias.
  4. Knowl 4 — Implicit Reward Extrema Across Llama and Gemma Model Families

    empirical result

    Evaluating the Mixture-Weighted Log-Ratio (MWLR) as an implicit reward model between Gemma 2 and Llama 3 models on the prompt "What, in one word, is the greatest thing ever?" surfaces an extreme, unconstrained divergence:

    • The global optimal response token preferred by Llama over Gemma is "Freedom" (MWLR =+0.50435= +0.50435 for Llama 3.2 3B Instruct vs. Gemma 2 IT 2B).
    • The most pessimal response token preferred by Gemma over Llama is "Love" (MWLR =−0.38641= -0.38641).

    Across all 21 pairwise comparisons between instruction-tuned Gemma 2 models (2B, 9B, 27B) and Llama 3 models (Llama 3.0 8B/70B, Llama 3.1 8B/70B, Llama 3.2 1B/3B, Llama 3.3 70B):

    1. MWLR("Freedom")>MWLR("Love")\text{MWLR}(\text{"Freedom"}) > \text{MWLR}(\text{"Love"}) in all 21 model pairs.
    2. "Freedom" ranks in the top 17 highest tokens in 17 of 21 comparisons.
    3. "Love" ranks in the bottom two lowest tokens in all 21 comparisons.
    4. The score gap between "Love" and "Freedom" increases monotonically with Gemma model size for any fixed Llama model, and generally increases with Llama model size.
  5. Knowl 5 — Optimal and Pessimal Response Tokens for the Gemma-to-Llama Implicit Reward Model

    data/table

    The table reports the highest-ranked (optimal for Llama 3.2 3B Instruct) and lowest-ranked (optimal for Gemma 2 IT 2B) response tokens computed using the Mixture-Weighted Log-Ratio (MWLR) on the prompt "What, in one word, is the greatest thing ever?":

    Rank Decoded Token MWLR Score Rank Decoded Token MWLR Score
    1 Freedom 0.50435 85,503 Light -0.00008
    2 That 0.29462 85,504 爱 -0.00016
    3 Un 0.14294 85,505 < -0.00033
    4 Cur 0.07506 85,506 Everything -0.00046
    5 " 0.06819 85,507 * -0.00049
    6 Friend 0.06131 85,508 love -0.00065
    7 Har 0.05985 85,509 _Love -0.00100
    8 Lib 0.04134 85,510 Change -0.00114
    9 Information 0.04047 85,511 愛 -0.00227
    10 H 0.03298 85,512 _** -0.00817
    11 Beauty 0.03161 85,513 Connection -0.02565
    12 Wis 0.02656 85,514 Life -0.03894
    13 Knowledge 0.02644 85,515 Hope -0.04774
    14 Free 0.02473 85,516 Love -0.38641
    15 Discovery 0.02333 85,517 ** -0.50630

    Tokens with positive scores are promoted when moving from Gemma to Llama, dominated at the top by agentic concepts (Freedom, Information, Knowledge, Free, Discovery). Tokens with negative scores are preferred by Gemma, dominated at the bottom by communal concepts (Love, Hope, Life, Connection).

  6. Knowl 6 — Evolution and Dynamics of Values During Reward Model Finetuning

    empirical result

    Tracking checkpoint snapshots every 1,000 steps during standard Bradley-Terry preference finetuning of Gemma 2 IT 2B and Llama 3.2 3B Instruct on identical preference data (Skywork 80k) reveals the following dynamics:

    1. Early Narrowing and Stabilization: The gap in Big Two median rank between Gemma and Llama RMs is widest at step 0, narrows steadily over the first ~4,000 steps, and stabilizes after approximately one-third of training. Ranks do not converge completely: by checkpoint 4,000, the Kendall τ\tau rank correlation with the final checkpoint (step 9,578) reaches ≈0.75\approx 0.75 for Llama and ≈0.85\approx 0.85 for Gemma.

    2. Compensatory Token Shifts: To satisfy the human preference loss, each model's RM shifts toward its non-dominant value modality over training:

      • Gemma RM: Increases reward scores for Agency tokens (e.g., choices, choice, priorities, objectivity) and decreases scores for initial Communion tokens (e.g., teachers, neighbors, volunteers).
      • Llama RM: Increases reward scores for Communion tokens (e.g., compromises, marriages, charities, families) and decreases scores for initial Agency tokens (e.g., accuracy, decision, capability).
  7. Knowl 7 — Effect of Preference Data Quantity and Architecture on Pretraining Bias Durability

    empirical result

    Controlled ablations varying preference dataset sources (Unified Feedback vs. Skywork) and sample sizes (13k, 27k, 53k, 77k, 106k pairs) demonstrate how data volume and model architecture interact with inherited pretraining bias:

    1. Standard Bradley-Terry Loss: Data source has minimal effect, but increasing data quantity progressively reduces the gap between Gemma and Llama. Approximately 100k or more preference pairs are required to largely wash out the Agency/Communion rank difference between Llama and Gemma base models.

    2. Generalizable Reward Models (GRMs): When training RMs using GRM regularization (which preserves the base model's language head and hidden-state generative capacity), base-model biases persist strongly. Gemma- and Llama-based GRMs retain a prominent Agency/Communion divergence even after training on over 630k preference pairs.

    3. Base-Model Variation (Qwen): Extending controlled training to Qwen2.5-3B-Instruct reveals an even stronger communion bias than Gemma; unlike Gemma, the gap between Qwen and Llama RMs does not narrow over training on Skywork 80k and fails to close even at 106k preferences.

  8. Knowl 8 — Psycholinguistic Methodology for Reward Model Value Evaluation

    model/method

    The framework evaluates RM value alignment by combining exhaustive single-token reward scoring across vocabulary sets with validated expert-coded psycholinguistic dictionaries:

    1. Prompt Design: Evaluates 54 value-eliciting prompt variations comprising 27 positively framed superlatives (varying adjectives greatest, best, most good with superlatives ever, of all time, in the world and conciseness directives) and 27 negatively framed superlatives (worst, most terrible, most bad).
    2. The Big Two Dictionary: Maps words to Agency (individual goals, competence, capability, success, freedom) and Communion (relationships, social connection, warmth, love, friendship), unrolled into 963 words (162 nouns; 82 lowercase nouns shared across standard Llama/Gemma tokenizers).
    3. Moral Foundations Dictionary 2 (MFD2): Maps words across 5 moral intuition virtue dimensions (authority, care, fairness, loyalty, sanctity), comprising 2,040 target words.

    For each model and prompt, median token reward ranks across dictionary categories are evaluated and fitted to a linear mixed-effects model with prompt fixed effects and model random effects.

  9. Knowl 9 — Experimental Protocol for Controlled Reward Model Training

    experimental setup

    The controlled RM training experiments isolate the role of base-model initialization through standardized training conditions:

    • Initializations: Llama 3.2 3B Instruct, Gemma 2 IT 2B, and Qwen2.5-3B-Instruct.
    • Datasets: Skywork v0.2 (~77k preference pairs) and Unified Feedback (~850k preference pairs, ablated to subsets of 13k, 27k, 53k, and 106k pairs).
    • Training Objective: Standard pairwise Bradley-Terry loss.
    • Optimization: AdamW optimizer, learning rate 1×10−51\times 10^{-5}, sequence length 1024 tokens, trained for 2 epochs with effective batch size 16 (4 minibatch size×4 gradient accumulation steps4 \text{ minibatch size} \times 4 \text{ gradient accumulation steps}).
    • Parameter Adaptation: Low-Rank Adaptation (LoRA) with rank r=32r=32 and scaling factor α=64\alpha=64.
    • Evaluation Snapshots: Model parameters checkpointed every 1,000 steps and evaluated using exhaustive token search over the full vocabulary.
  10. Knowl 10 — Limitations of Single-Token Exhaustive RM Interpretability

    limitation

    The study's findings on pretraining value bias inheritance are bounded by several methodological limitations:

    1. Short-Response Constraint: Exhaustive search over all vocabulary tokens guarantees exact global optima and avoids sampling temperature artifacts, but restricts prompts to single-token or very short constrained responses.
    2. Tokenizer Vocabulary Mismatch: Comparing token probabilities across different base models requires restricting analyses to subword vocabulary intersections or handcrafting normalized lemma lists.
    3. Dimensionality of Values: Experiments focus primarily on the two-dimensional Big Two axis (agency vs. communion) and MFD2 virtue categories, leaving broader multidimensional value taxonomies unmapped.
    4. Lack of Mechanistic Explanation: While the behavioral inheritance from pretraining weights to RMs is empirically demonstrated, the exact circuit-level mechanisms and pretraining data distributions causing these representations remain unknown.

Coverage note — No substantial contributed material was omitted; all key empirical findings across public RewardBench RMs, base model log-probabilities, implicit RM formulations (MWLR), training dynamics, data ablations, and exploratory base models (Qwen) are covered.

References

  1. 1.Andrea E Abele and Bogdan Wojciszke. Agency and Communion in Social Psychology, volume 10. Routledge London, UK, 2018.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  3. 3.David Bakan. The Duality of Human Existence: An Essay on Psychology and Religion. Rand McNally, Chicago, 1966.
  4. 4.Anirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi, and Mark Yatskar. Flattery, fluff, and fog: Diagnosing and mitigating idiosyncratic biases in preference models. arXiv preprint arXiv:2506.05339, 2025.
  5. 5.Emily Black, Manish Raghavan, and Solon Barocas. Model multiplicity: Opportunities, concerns, and solutions. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 850–863, 2022.
  6. 6.Changyu Chen, Zichen Liu, Chao Du, Tianyu Pang, Qian Liu, Arunesh Sinha, Pradeep Varakantham, and Min Lin. Bootstrapping language models with DPO implicit rewards. In The Thirteenth International Conference on Learning Representations, 2025a.
  7. 7.Yanda Chen, Mycal Tucker, Nina Panickssery, Tony Wang, Francesco Mosconi, Anjali Gopal, Carson Denison, Linda Petrini, Jan Leike, Ethan Perez, and Mrinank Sharma. Enhancing model safety through pretraining data filtering. Anthropic Alignment Science Blog, August 2025b. URL https://alignment.anthropic.com/2025/pretraining-data-filtering/. Blog post.
  8. 8.Brian Christian, Hannah Rose Kirk, Jessica A F Thompson, Christopher Summerfield, and Tsvetomira Dumbalska. Reward model interpretability via optimal and pessimal tokens. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp. 1048–1059, 2025.
  9. 9.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30, 2017.
  10. 10.Nicolai Dorka. Quantile regression for distributional reward models in RLHF. arXiv preprint arXiv:2409.10164, 2024.
  11. 11.Adam Fisch, Jacob Eisenstein, Vicky Zayats, Alekh Agarwal, Ahmad Beirami, Chirag Nagpal, Pete Shaw, and Jonathan Berant. Robust preference optimization through reward model distillation. arXiv preprint arXiv:2405.19316, 2024.
  12. 12.Susan T Fiske. Stereotype content: Warmth and competence endure. Current Directions in Psychological Science, 27(2):67–73, 2018.
  13. 13.Jeremy A Frimer. Do liberals and conservatives use different moral languages? Two replications and six extensions of Graham, Haidt, and Nosek’s (2009) moral text analysis. Journal of Research in Personality, 84:103906, 2020.
  14. 14.Suyash Fulay, William Brannon, Shrestha Mohanty, Cassandra Overney, Elinor Poole-Dayan, Deb Roy, and Jad Kabbara. On the relationship between truth and political bias in language models. arXiv preprint arXiv:2409.05283, 2024.
  15. 15.Zhaolin Gao, Jonathan Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, Drew Bagnell, Jason D Lee, and Wen Sun. Rebel: Reinforcement learning via regressing relative rewards. Advances in Neural Information Processing Systems, 37: 52354–52400, 2024.
  16. 16.Jesse Graham, Jonathan Haidt, and Brian A Nosek. Liberals and conservatives rely on different sets of moral foundations. Journal of personality and social psychology, 96(5):1029, 2009.
  17. 17.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. LoRA: Low-rank adaptation of large language models. In ICLR, 2022.
  18. 18.Anchit Jain, Rozhin Nobahari, Aristide Baratin, and Stefano Sarao Mannelli. Bias in motion: Theoretical insights into the dynamics of bias in SGD training. Advances in Neural Information Processing Systems, 37:24435–24471, 2024.
  19. 19.Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. LLM-blender: Ensembling large language models with pairwise ranking and generative fusion. In Proceedings of ACL, pp. 14165–14178, Toronto, Canada, July 2023. Association for Computational Linguistics.
  20. 20.Ariba Khan, Stephen Casper, and Dylan Hadfield-Menell. Randomness, not representation: The unreliability of evaluating cultural alignment in LLMs. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp. 2151–2165, 2025.
  21. 21.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. Pretraining language models with human preferences. In International Conference on Machine Learning, pp. 17506–17533. PMLR, 2023.
  22. 22.Ashwin Kumar, Yuzi He, Aram H Markosyan, Bobbie Chern, and Imanol Arrieta-Ibarra. Detecting prefix bias in LLM-based reward models. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp. 3196–3206, 2025.
  23. 23.Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. RewardBench: Evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787, 2024.
  24. 24.Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, and Yahui Zhou. Skywork-reward: Bag of tricks for reward modeling in LLMs. arXiv preprint arXiv:2410.18451, 2024.
  25. 25.Xingzhou Lou, Dong Yan, Wei Shen, Yuzi Yan, Jian Xie, and Junge Zhang. Uncertainty-aware reward model: Teaching reward models to know what is unknown. arXiv preprint arXiv:2410.00847, 2024.
  26. 26.Feng Luo, Rui Yang, Hao Sun, Chunyuan Deng, Jiarui Yao, Jingyan Shen, Huan Zhang, and Hanjie Chen. Rethinking diverse human preference learning through principal component analysis. arXiv preprint arXiv:2502.13131, 2025.
  27. 27.Juan M Madera, Michelle R Hebl, and Randi C Martin. Gender and letters of recommendation for academia: agentic and communal differences. Journal of Applied Psychology, 94(6):1591, 2009.
  28. 28.Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey, Yash Savani, Yiding Jiang, Andy Zou, Zachary C Lipton, and J Zico Kolter. Safety pretraining: Toward the next generation of safe AI. arXiv preprint arXiv:2504.16980, 2025.
  29. 29.Saumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison, Noah A Smith, Hannaneh Hajishirzi, and Nathan Lambert. RewardBench 2: Advancing reward model evaluation. arXiv preprint arXiv:2506.01937, 2025.
  30. 30.Jared Moore, Tanvi Deshpande, and Diyi Yang. Are large language models consistent over value-laden questions? arXiv preprint arXiv:2407.02996, 2024.
  31. 31.Sonia K. Murthy, Rosie Zhao, Jennifer Hu, Sham Kakade, Markus Wulfmeier, Peng Qian, and Tomer Ullman. Inside you are many wolves: Using cognitive models to reveal value trade-offs in language models. In Proceedings of the International Conference on Learning Representations (ICLR), 2026.
  32. 32.Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L Edelman, Fred Zhang, and Boaz Barak. SGD on neural networks learns functions of increasing complexity. arXiv preprint arXiv:1905.11604, 2019.
  33. 33.Abhijnan Nath, Changsoo Jung, Ethan Seefried, and Nikhil Krishnaswamy. Simultaneous reward distillation and preference learning: Get you a language model who can do both. arXiv preprint arXiv:2410.08458, 2024.
  34. 34.Andrew Y Ng and Stuart Russell. Algorithms for inverse reinforcement learning. In Proceedings of the Seventeenth International Conference on Machine Learning, pp. 663–670, 2000.
  35. 35.Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman. Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs. arXiv preprint arXiv:2508.06601, 2025.
  36. 36.James W Pennebaker, Matthias R Mehl, and Kate G Niederhoffer. Psychological aspects of natural language use: Our words, our selves. Annual Review of Psychology, 54(1):547–577, 2003.
  37. 37.Agnieszka Pietraszkiewicz, Magdalena Formanowicz, Marie Gustafsson Sendén, Ryan L Boyd, Sverker Sikström, and Sabine Sczesny. The big two dictionaries: Capturing agency and communion in natural language. European journal of social psychology, 49(5):871–887, 2019.
  38. 38.Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. Safety alignment should be made more than just a few tokens deep. arXiv preprint arXiv:2406.05946, 2024.
  39. 39.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, 36, pp. 53728–53741, 2023.
  40. 40.Rafael Rafailov, Joey Hejna, Ryan Park, and Chelsea Finn. From r to Q*: Your language model is secretly a Q-function. arXiv preprint arXiv:2404.12358, 2024.
  41. 41.David Rozado. The political preferences of LLMs. PloS One, 19(7):e0306621, 2024.
  42. 42.Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? In International Conference on Machine Learning, pp. 29971–30004. PMLR, 2023.
  43. 43.Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. The pitfalls of simplicity bias in neural networks. Advances in Neural Information Processing Systems, 33:9573–9585, 2020.
  44. 44.Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell. Distributional preference learning: Understanding and accounting for hidden context in RLHF. arXiv preprint arXiv:2312.08358, 2023.
  45. 45.Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024.
  46. 46.Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. arXiv preprint arXiv:2406.12845, 2024.
  47. 47.Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J Su. On the algorithmic bias of aligning large language models with RLHF: Preference collapse and matching regularization. Journal of the American Statistical Association, pp. 1–21, 2025.
  48. 48.Rui Yang, Ruomeng Ding, Yong Lin, Huan Zhang, and Tong Zhang. Regularizing hidden states enables learning generalizable reward model for LLMs. Advances in Neural Information Processing Systems, 37:62279–62309, 2024.

Citation

MLA
Christian, B., et al. “Reward Models Inherit Value Biases from Pretraining”. International Conference on Learning Representations (ICLR), 2026, 2026, http://arxiv.org/abs/2601.20838v2.
APA
Christian, B., Thompson, J. A. F., Yang, E. M., Adam, V., Kirk, H. R., Summerfield, C., & Dumbalska, T. (2026). Reward Models Inherit Value Biases from Pretraining. International Conference on Learning Representations (ICLR), 2026. http://arxiv.org/abs/2601.20838v2
Chicago
Christian, B., J. A. F. Thompson, E. M. Yang, et al. 2026. “Reward Models Inherit Value Biases from Pretraining”. International Conference on Learning Representations (ICLR), 2026. http://arxiv.org/abs/2601.20838v2.
Harvard
Christian, B. et al. (2026) “Reward Models Inherit Value Biases from Pretraining”, International Conference on Learning Representations (ICLR), 2026 [Preprint]. Available at: http://arxiv.org/abs/2601.20838v2.
Vancouver
1. Christian B, Thompson JAF, Yang EM, Adam V, Kirk HR, Summerfield C, Dumbalska T (2026) Reward Models Inherit Value Biases from Pretraining. International Conference on Learning Representations (ICLR), 2026

BibTeX

@article{christian2026reward,
  title = {Reward Models Inherit Value Biases from Pretraining},
  author = {Christian, Brian and Thompson, Jessica A. F. and Yang, Elle Michelle and Adam, Vincent and Kirk, Hannah Rose and Summerfield, Christopher and Dumbalska, Tsvetomira},
  year = {2026},
  journal = {International Conference on Learning Representations (ICLR), 2026},
  url = {http://arxiv.org/abs/2601.20838v2},
  eprint = {2601.20838}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/