MaxMin-RLHF: Alignment with Diverse Human Preferences
Souradip ChakrabortyJiahao QiuHui YuanAlec KoppelDinesh ManochaFurong HuangAmrit S. BediMengdi Wang
Proves theoretically that single-reward RLHF fails to capture diverse user preferences and introduces MaxMin-RLHF, an egalitarian framework that learns multiple reward models via expectation-maximization to fairly represent minority viewpoints in language models.
Modern artificial intelligence models are commonly aligned to human values using Reinforcement Learning from Human Feedback (RLHF). This standard alignment framework fits a single reward model to crowdsourced human preferences. However, crowdsourced feedback inherently reflects diverse perspectives driven by demographic, cultural, and personal differences. Relying on an averaged, single-reward target assumes a single ground truth, which systematically biases model behavior toward majority viewpoints and compromises fairness and safety for minority user groups.
To address this limitation, the article aims to formally evaluate the theoretical shortcomings of single-reward RLHF and demonstrate an alternative framework that fairly accommodates diverse human preferences without sacrificing overall model performance.
To investigate this, the article establishes a mathematical proof bounding the performance gap caused by preference diversity and introduces MaxMin-RLHF. Inspired by the egalitarian principle in social choice theory, this approach uses an Expectation-Maximization (EM) clustering algorithm to separate unlabelled user feedback into distinct reward models, then iteratively optimizes the language model to maximize the utility of the worst-served subpopulation. The framework was evaluated across small-scale text generation experiments using GPT-2 and large-scale instruction-following tasks using Tulu2-7B across varied preference distributions.
Key findings show that single-reward RLHF mathematically and empirically fails when preferences diverge. In baseline evaluations, as user group imbalance increased from 1:1 to 10:1, single-reward accuracy on minority preferences dropped steeply from 70.4% down to 42.0%, ignoring minority criteria entirely. In contrast, the EM algorithm correctly clustered diverse user groups within four iterations without requiring ground-truth demographic labels. Applying MaxMin-RLHF successfully balanced conflicting objectives, achieving balanced sentiment and conciseness on GPT-2 and delivering consistently superior pairwise win rates across diverse user segments on Tulu2-7B (reaching 55.6% to 60.0% win rates on minority evaluation subsets, compared to 44.0% to 51.7% under standard skewed baselines).
These results demonstrate that standard alignment practices introduce significant societal bias and compliance risks by marginalizing minority user preferences. Implementing social welfare objectives like MaxMin-RLHF ensures robust, equitable performance across diverse user bases, offering a viable path for organizations aiming to deploy socially aligned and fair language models.
Organizations developing customer-facing language models should avoid single-reward aggregation on unsegmented feedback. Instead, teams should implement mixture-of-reward modeling and max-min policy optimization to protect minority utilities. Future work should expand these evaluations beyond simulated annotator datasets to real-world, highly heterogeneous human populations and explore computationally scalable reward clustering for large numbers of user groups.
The findings are supported by solid mathematical bounds and consistent empirical benchmarks across multiple model scales. However, readers should consider the operational assumptions: the empirical validation relies on discrete preference clusters and simulated persona-based feedback (e.g., via GPT-4). Consequently, real-world deployment on complex, overlapping human demographics should proceed with appropriate continuous monitoring.
- Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). This foundational work introduces learning a reward model from human comparisons and optimizing a policy against it—the basic RLHF pipeline that MaxMin-RLHF modifies to account for diverse preferences.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). Its language-model RLHF pipeline makes the single-reward-model setup explicit, providing the baseline MaxMin-RLHF challenges and replaces with a mixture of rewards.
- Paper: Learning to summarize from human feedback, Nisan Stiennon et al. (2020). This influential summarization study shows how human preference comparisons train a reward model for policy optimization, grounding the RLHF approach that MaxMin-RLHF generalizes.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). GDPO carries multi-reward alignment into a later policy-optimization method, making it a direct next step for exploring how to preserve distinct preference signals during training.
- Paper: Rushes: A Human Preference Dataset for Pluralistic Alignment, Michael Xu et al. (2026). Rushes extends pluralistic alignment from optimizing against diverse preferences to measuring personalized engagement across real user choices.
- Paper: Vector Policy Optimization: Training for Diversity Improves Test-Time Search, Ryan Bahlous-Boldi et al. (2026). VPO continues the multi-objective alignment thread by training models to cover diverse trade-offs among reward components rather than collapse to one response.
