Rushes: A Human Preference Dataset for Pluralistic Alignment
Michael XuJorge J. G. LeandroSudha RaoWeijia XuNebojsa JojicGabriel DesGarennesChris QuirkBill Dolan
Introduces Rushes, a sequential human choice dataset that reveals frontier language models fall behind simple matrix factorization in predicting personalized engagement, exposing the failure of standard reinforcement learning from human feedback to capture diverse individual preferences.
Current development and alignment paradigms for artificial intelligence predominantly focus on safety, helpfulness, and factual accuracy. While these objectives aim for universal convergence across users, entertainment and interactive domains require divergence—adapting to subjective, individualized notions of fun, interest, and engagement. Standard preference optimization methods aggregate feedback into population-level rewards, which suppresses minority preferences and produces generic interactions. Consequently, systems lack the capability to model long-term, individualized user engagement across sequential narrative decisions.
The article introduces Rushes, a new dataset and diagnostic benchmark designed to evaluate how well artificial intelligence models can predict revealed, personalized human engagement in interactive branching narratives.
To capture organic engagement rather than synthetic or survey-based feedback, the authors deployed six artificial intelligence-generated interactive games to 8,167 authenticated players via the Xbox Insiders Program. Across 44,226 uncompensated decision events, users navigated multi-day storylines by selecting one of four distinct narrative branches at each step. The authors evaluated models on chronological, single-choice prediction using candidate sets, narrative history, and user trajectories, comparing state-of-the-art large language models against classical recommendation and popularity baselines.
The evaluation revealed several key findings regarding personalized modeling and current frontier systems. First, user decisions exhibited clear, non-random structure, with aggregate choice entropy measuring significantly lower than a uniform baseline. Second, frontier models exhibited an "Engagement Gap": GPT-5 prompted with user history achieved only 34.2% top-1 accuracy, failing to beat a simple popularity heuristic at 36.4%. Third, classical matrix factorization via Singular Value Decomposition captured the strongest personalization signal at 37.7%, demonstrating that standard language models default to generic population preferences rather than tailoring decisions to individual trajectories. Fourth, history source proved critical: within-game user history yielded an accuracy of 38.9%, whereas cross-game history dropped performance to 29.1%, highlighting that engagement preferences are highly context-dependent.
These findings indicate that scaling model size and contextual reasoning alone is insufficient for subjective alignment. Moving from GPT-4o to GPT-5 provided less than a 1% gain in zero-shot choice prediction. For decision-makers and system architects building consumer-facing generative applications, relying solely on standard safety and helpfulness tuning risks deploying unengaging, homogenized products. Successfully capturing subjective appeal requires developing specialized reward architectures and recommendation-aware frameworks capable of disentangling universal quality from idiosyncratic taste.
Organizations developing interactive artificial intelligence systems should integrate sequential collaborative filtering signals alongside language representations and prioritize domain-specific context over broad cross-domain assumptions. The authors emphasize that engagement-modeling techniques carry dual-use risks, cautioning that these methods must be directed toward pluralistic alignment rather than manipulative or addictive interfaces. Further research should explore generative reward models capable of interpreting implicit narrative style and subtext.
The findings are subject to several boundary conditions. The underlying dataset is derived from an English-speaking, gaming-literate Xbox Insider demographic, which may not generalize to non-gaming or multilingual populations. In addition, the generated narrative content is bound to the capabilities and prompting structures of the specific models used. While the reported confidence intervals demonstrate robust statistical separation among baselines, stakeholders should treat these conclusions as representative of context-dependent gaming preferences rather than universal human behavior.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational language-model RLHF pipeline shows how human comparisons become reward signals for policy optimization, the alignment setup Rushes evaluates against personalized engagement choices.
- Paper: Deep reinforcement learning from human preferences, Paul F. Christiano et al. (2017). Its preference-learning framework establishes the comparison-based feedback foundations that help contextualize Rushes’s use of human choices as alignment signals.
- Paper: Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation, Yuval Kirstain et al. (2023). Pick-a-Pic’s large-scale dataset of real users choosing generated outputs provides a direct precedent for collecting revealed preferences through an interactive generative interface.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). Chatbot Arena demonstrates how live user choices can power human-preference evaluation, a useful foundation for understanding Rushes’s interactive benchmark design.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). This review explains why treating human feedback as a universal objective obscures differences in values and populations, motivating Rushes’s focus on pluralistic preferences.
No sufficiently relevant recommendations were found.
