MaxMin-RLHF: Alignment with Diverse Human Preferences

Souradip ChakrabortyJiahao QiuHui YuanAlec KoppelDinesh ManochaFurong HuangAmrit S. BediMengdi Wang

article2024ICML143 citations

Proves theoretically that single-reward RLHF fails to capture diverse user preferences and introduces MaxMin-RLHF, an egalitarian framework that learns multiple reward models via expectation-maximization to fairly represent minority viewpoints in language models.

Listen

Modern artificial intelligence models are commonly aligned to human values using Reinforcement Learning from Human Feedback (RLHF). This standard alignment framework fits a single reward model to crowdsourced human preferences. However, crowdsourced feedback inherently reflects diverse perspectives driven by demographic, cultural, and personal differences. Relying on an averaged, single-reward target assumes a single ground truth, which systematically biases model behavior toward majority viewpoints and compromises fairness and safety for minority user groups.

To address this limitation, the article aims to formally evaluate the theoretical shortcomings of single-reward RLHF and demonstrate an alternative framework that fairly accommodates diverse human preferences without sacrificing overall model performance.

To investigate this, the article establishes a mathematical proof bounding the performance gap caused by preference diversity and introduces MaxMin-RLHF. Inspired by the egalitarian principle in social choice theory, this approach uses an Expectation-Maximization (EM) clustering algorithm to separate unlabelled user feedback into distinct reward models, then iteratively optimizes the language model to maximize the utility of the worst-served subpopulation. The framework was evaluated across small-scale text generation experiments using GPT-2 and large-scale instruction-following tasks using Tulu2-7B across varied preference distributions.

Key findings show that single-reward RLHF mathematically and empirically fails when preferences diverge. In baseline evaluations, as user group imbalance increased from 1:1 to 10:1, single-reward accuracy on minority preferences dropped steeply from 70.4% down to 42.0%, ignoring minority criteria entirely. In contrast, the EM algorithm correctly clustered diverse user groups within four iterations without requiring ground-truth demographic labels. Applying MaxMin-RLHF successfully balanced conflicting objectives, achieving balanced sentiment and conciseness on GPT-2 and delivering consistently superior pairwise win rates across diverse user segments on Tulu2-7B (reaching 55.6% to 60.0% win rates on minority evaluation subsets, compared to 44.0% to 51.7% under standard skewed baselines).

These results demonstrate that standard alignment practices introduce significant societal bias and compliance risks by marginalizing minority user preferences. Implementing social welfare objectives like MaxMin-RLHF ensures robust, equitable performance across diverse user bases, offering a viable path for organizations aiming to deploy socially aligned and fair language models.

Organizations developing customer-facing language models should avoid single-reward aggregation on unsegmented feedback. Instead, teams should implement mixture-of-reward modeling and max-min policy optimization to protect minority utilities. Future work should expand these evaluations beyond simulated annotator datasets to real-world, highly heterogeneous human populations and explore computationally scalable reward clustering for large numbers of user groups.

The findings are supported by solid mathematical bounds and consistent empirical benchmarks across multiple model scales. However, readers should consider the operational assumptions: the empirical validation relies on discrete preference clusters and simulated persona-based feedback (e.g., via GPT-4). Consequently, real-world deployment on complex, overlapping human demographics should proceed with appropriate continuous monitoring.

arXiv: 2402.08925CharlesQ9/MaxMinRLHF
  • Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). This foundational work introduces learning a reward model from human comparisons and optimizing a policy against it—the basic RLHF pipeline that MaxMin-RLHF modifies to account for diverse preferences.
  • Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Its language-model RLHF pipeline makes the single-reward-model setup explicit, providing the baseline MaxMin-RLHF challenges and replaces with a mixture of rewards.
  • Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). This influential summarization study shows how human preference comparisons train a reward model for policy optimization, grounding the RLHF approach that MaxMin-RLHF generalizes.
Cover for MaxMin-RLHF: Alignment with Diverse Human Preferences

Abstract

Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data. However, the single reward model overlooks the rich diversity of human preferences inherent in data collected from multiple users. In this work, we first derive an impossibility result of alignment with single reward RLHF, thereby highlighting its insufficiency in representing diverse human preferences. Next, we propose to learn a mixture of reward models via an expectation-maximization algorithm and solve a MaxMin alignment objective inspired by the Egalitarian principle in social choice theory to better honor diverse human preferences. We present comprehensive experimental results on small-scale (GPT-2) and large-scale language (with Tulu2-7B)) and show the efficacy of the proposed approach in the presence of diversity among human preferences. We remark that our findings in this work are not only limited to language models but also extend to reinforcement learning in general.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. An Impossibility Result for Single Reward RLHF with Diverse Preferences
  • 3.1. Diversity in Human Preferences
  • 3.2. Reward Mismatch Due to Diversity
  • 3.3. An Impossibility Results of Alignment
  • 4. MaxMin-RLHF: One Possibility
  • 5. Experimental Results
  • 5.1. Small Scale Experiments (with GPT-2): Sentiment and Conciseness Alignment
  • 5.2. Large Scale Experiments (with Tulu2-7B)
  • 5.2.1. MAIN RESULTS
  • 6. Conclusions
  • Impact Statement
  • Acknowledgements
  • References
  • A. Notations
  • B. A detailed Context of Related Works
  • C. Preliminary Results
  • D. Proof of Lemma 1
  • E. Proof of Theorem 1
  • F. Additional Details of the Experiments
  • G. Additional Experiments in Robotics Navigation Tasks

Knowls

  1. Knowl 1 — No extractable paper content available for knowl extraction

    limitation

    The source document provided for knowledge extraction yielded no readable paper content (no text could be extracted from the supplied PDF file). As a result, none of the paper's own contributions — its methods, models, theory, experiments, results, or stated limitations — could be identified, and no substantive knowls about the paper's contribution can be reconstructed from the available material.

Coverage note — The provided source (file 401e8ff2-a8d9-437e-bf5e-e1ef41b576f3.pdf) contained no extractable text content, so the paper's methods, results, and analyses could not be read; consequently no substantive contributed material was extracted and the single knowl records this extraction limitation rather than paper content.

Citation

MLA
Chakraborty, S., et al. “MaxMin-RLHF: Alignment with Diverse Human Preferences”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.08925.
APA
Chakraborty, S., Qiu, J., Yuan, H., Koppel, A., Huang, F., Manocha, D., Bedi, A. S., & Wang, M. (2024). MaxMin-RLHF: Alignment with Diverse Human Preferences. arXiv. https://doi.org/10.48550/arxiv.2402.08925
Chicago
Chakraborty, S., J. Qiu, H. Yuan, et al. 2024. “MaxMin-RLHF: Alignment with Diverse Human Preferences”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.08925.
Harvard
Chakraborty, S. et al. (2024) “MaxMin-RLHF: Alignment with Diverse Human Preferences”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.08925.
Vancouver
1. Chakraborty S, Qiu J, Yuan H, Koppel A, Huang F, Manocha D, Bedi AS, Wang M (2024) MaxMin-RLHF: Alignment with Diverse Human Preferences. https://doi.org/10.48550/arxiv.2402.08925

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.08925,
  doi = {10.48550/ARXIV.2402.08925},
  url = {https://arxiv.org/abs/2402.08925},
  author = {Chakraborty, Souradip and Qiu, Jiahao and Yuan, Hui and Koppel, Alec and Huang, Furong and Manocha, Dinesh and Bedi, Amrit Singh and Wang, Mengdi},
  keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), Machine Learning (cs.LG), Robotics (cs.RO), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {MaxMin-RLHF: Alignment with Diverse Human Preferences},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/