Rushes: A Human Preference Dataset for Pluralistic Alignment

Michael XuJorge J. G. LeandroSudha RaoWeijia XuNebojsa JojicGabriel DesGarennesChris QuirkBill Dolan

article2026arXiv0 citations

Introduces Rushes, a sequential human choice dataset that reveals frontier language models fall behind simple matrix factorization in predicting personalized engagement, exposing the failure of standard reinforcement learning from human feedback to capture diverse individual preferences.

Listen

Current development and alignment paradigms for artificial intelligence predominantly focus on safety, helpfulness, and factual accuracy. While these objectives aim for universal convergence across users, entertainment and interactive domains require divergence—adapting to subjective, individualized notions of fun, interest, and engagement. Standard preference optimization methods aggregate feedback into population-level rewards, which suppresses minority preferences and produces generic interactions. Consequently, systems lack the capability to model long-term, individualized user engagement across sequential narrative decisions.

The article introduces Rushes, a new dataset and diagnostic benchmark designed to evaluate how well artificial intelligence models can predict revealed, personalized human engagement in interactive branching narratives.

To capture organic engagement rather than synthetic or survey-based feedback, the authors deployed six artificial intelligence-generated interactive games to 8,167 authenticated players via the Xbox Insiders Program. Across 44,226 uncompensated decision events, users navigated multi-day storylines by selecting one of four distinct narrative branches at each step. The authors evaluated models on chronological, single-choice prediction using candidate sets, narrative history, and user trajectories, comparing state-of-the-art large language models against classical recommendation and popularity baselines.

The evaluation revealed several key findings regarding personalized modeling and current frontier systems. First, user decisions exhibited clear, non-random structure, with aggregate choice entropy measuring significantly lower than a uniform baseline. Second, frontier models exhibited an "Engagement Gap": GPT-5 prompted with user history achieved only 34.2% top-1 accuracy, failing to beat a simple popularity heuristic at 36.4%. Third, classical matrix factorization via Singular Value Decomposition captured the strongest personalization signal at 37.7%, demonstrating that standard language models default to generic population preferences rather than tailoring decisions to individual trajectories. Fourth, history source proved critical: within-game user history yielded an accuracy of 38.9%, whereas cross-game history dropped performance to 29.1%, highlighting that engagement preferences are highly context-dependent.

These findings indicate that scaling model size and contextual reasoning alone is insufficient for subjective alignment. Moving from GPT-4o to GPT-5 provided less than a 1% gain in zero-shot choice prediction. For decision-makers and system architects building consumer-facing generative applications, relying solely on standard safety and helpfulness tuning risks deploying unengaging, homogenized products. Successfully capturing subjective appeal requires developing specialized reward architectures and recommendation-aware frameworks capable of disentangling universal quality from idiosyncratic taste.

Organizations developing interactive artificial intelligence systems should integrate sequential collaborative filtering signals alongside language representations and prioritize domain-specific context over broad cross-domain assumptions. The authors emphasize that engagement-modeling techniques carry dual-use risks, cautioning that these methods must be directed toward pluralistic alignment rather than manipulative or addictive interfaces. Further research should explore generative reward models capable of interpreting implicit narrative style and subtext.

The findings are subject to several boundary conditions. The underlying dataset is derived from an English-speaking, gaming-literate Xbox Insider demographic, which may not generalize to non-gaming or multilingual populations. In addition, the generated narrative content is bound to the capabilities and prompting structures of the specific models used. While the reported confidence intervals demonstrate robust statistical separation among baselines, stakeholders should treat these conclusions as representative of context-dependent gaming preferences rather than universal human behavior.

No sufficiently relevant recommendations were found.

Cover for Rushes: A Human Preference Dataset for Pluralistic Alignment

Abstract

We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching narratives and select one choice from a small, explicit candidate set at each decision point. Each interaction logs the full candidate set, the user's choice, and the evolving narrative context, yielding time-ordered trajectories with persistent user-level identifiers. Rushes contains 44,226 decision events from 8,167 unique users across six games, capturing sequential, personalized engagement behavior rather than static judgments. We show that user choices exhibit structured, non-random patterns, quantified by a low choice entropy relative to a uniform baseline. We position Rushes as a diagnostic benchmark for pluralistic alignment and demonstrate a robust Engagement Gap: state-of-the-art LLMs, including GPT-5, fail to outperform simple baselines. While classical Matrix Factorization (SVD) captures measurable personalized signal (37.7%), frontier LLMs (34.23%) struggle to even match the Popularity Baseline (36.4%) on event-level choice prediction. This gap suggests that single, population-level objectives, like those used in modern RLHF, appear insufficient to capture heterogeneous, context-dependent engagement signals. As a result, even highly capable models default to majority preferences rather than adapting to individual trajectories. We release Rushes to support research into pluralistic alignment and sequential decision-making in generative systems. The full code for the platform and dataset will be available here: this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Rushes
  • 3.1 Game Generation
  • 3.1.1 Generating branching narrative text
  • 3.1.2 Generating image, audio, and video
  • 3.1.3 Generating narrative continuations
  • 3.1.4 Quality Control and Responsible AI
  • 3.2 Analysis of Generated Games
  • 3.2.1 Lexical and Semantic Diversity
  • 3.2.2 Multimodal Asset Evaluation
  • 3.2.3 Summary
  • 3.3 Data Collection
  • 3.3.1 Logging and Schema
  • 3.3.2 Dataset Composition
  • 3.3.3 Preference Transformation and Modeling
  • 4 Experiments and Results
  • 4.1 Main Results
  • 4.2 Ablation Studies
  • 4.2.1 Engagement by Narrative Depth
  • 4.2.2 The Role of History
  • 4.2.3 Active vs. Sparse Players
  • 4.2.4 Frontier Model Scaling: GPT-5 vs. GPT-4o
  • 5 Conclusion
  • References
  • A Rushes Game Generation Pipeline
  • A.1 Configuration Parameters
  • A.2 Deriving the paraphrase scaling rule
  • A.3 Main Generation Pipeline
  • A.4 Story Setup and Theme Generation
  • A.4.1 Theme Extraction Prompt
  • A.5 Recursive Story Generation
  • A.6 Level Generation with Branching
  • A.7 Option Generation with Variations
  • A.7.1 Option Creation Prompt
  • A.8 Similarity Checking for Uniqueness
  • A.9 Option Expansion for Variation
  • A.10 Game Continuation Algorithm
  • A.11 Image Prompt Generation
  • A.11.1 Character Extraction and Management
  • A.12 Audio Generation with SSML
  • A.13 Media Generation Pipeline

Knowls

  1. Knowl 1 — Rushes records longitudinal, revealed choices in interactive narratives

    data/table

    Rushes contains 44,226 decision events from 8,167 unique users across six AI-generated narrative games. A decision record includes an anonymized persistent user identifier, game and narrative level, the selected option, the unselected options, and interaction metadata such as time taken and session depth; the records preserve user trajectories through the branching stories. A decision typically offered four options. Average trajectory length was 5.4 decisions, users played an average of 1.4 games, and 195 users played all six. Participants were voluntary, authenticated Xbox Insiders and received no financial incentive.

  2. Knowl 2 — Collaborative filtering exceeds popularity, while language models do not

    empirical result

    On the same held-out Rushes test set of 8,293 events, evaluated by top-1 choice accuracy, SVD collaborative filtering achieved 0.3773 (95% Wilson CI [0.3669, 0.3878]); the most-frequent-option popularity baseline achieved 0.3639 ([0.3536, 0.3743]); GPT-5 prompted with user history achieved 0.3423 ([0.3321, 0.3526]); SASRec achieved 0.3406 ([0.3304, 0.3509]); a DeBERTa-v3 semantic classifier achieved 0.3000 ([0.2902, 0.3100]); and uniform random choice achieved 0.2541 ([0.2448, 0.2636]). The SVD point estimate is above popularity, indicating exploitable user-specific signal, but the tested frontier LLM and sequential recommender remain below the simple popularity baseline.

  3. Knowl 3 — Branching narrative and multimodal stimulus generation

    model/method

    Rushes generates narrative text and decision options with GPT-4o at temperature 0.3, starting from a user-provided synopsis and expanding a pre-generated story tree to four levels per day. Each decision presents four actionable options; an LLM similarity checker rejects options judged too similar to earlier options along the user's trajectory, considering action type, complexity, and outcome. Games contain approximately 330 nodes and are manually reviewed before release. To vary surface wording without changing option semantics, the system creates paraphrases according to a traffic-based heuristic. Let PP be the expected number of players, bb the number of options per node, and dd the decision depth with root depth 00. Assuming approximately even traffic across branches for sizing purposes, the number of additional paraphrases per option is

    V(d)=⌈Pbd+1⌉−1.V(d)=\left\lceil\sqrt{\frac{P}{b^{d+1}}}\right\rceil-1.

    The configuration used P=5,000P=5{,}000 and b=4b=4. A paraphrase is selected deterministically from a hash of the anonymized user ID and the node and option identifiers. The system also generates images using FLUX.1 schnell, short video clips from images using LTX-Video, and expressive Azure Text-to-Speech narration. For continuation across days, active leaf scenes are grouped into four broad narrative categories and continued individually in alignment with those categories.

  4. Knowl 4 — Choice prediction uses chronological, user-stratified evaluation

    experimental setup

    Rushes evaluates event-level prediction of the single option a user selected. For each event, a model receives the narrative context, the available candidate options, and the user's interaction history up to that decision, and predicts one option; performance is measured by top-1 accuracy. The split is chronological within each user: the first 80% of that user's interactions are used for training and the remaining 20% for testing, so each test decision occurs after that user's training history. The reported baseline comparisons use the same held-out test set.

  5. Knowl 5 — Same-game history predicts choices better than cross-game history

    empirical result

    For SVD-based prediction, accuracy was 0.3886 (95% Wilson CI [0.3775, 0.3999], 7,310 events) when the user's available history came from the same game as the test decision, compared with 0.2909 ([0.2634, 0.3201], 983 events) when history came from other games. SVD over all available history achieved 0.3773 ([0.3669, 0.3878], 8,293 events). The same-game versus cross-game difference is 9.77 percentage points, consistent with user preferences being substantially dependent on narrative context rather than transferring uniformly across games.

  6. Knowl 6 — Observed choices have lower entropy than uniform four-option choice

    empirical result

    The mean user-vote entropy across Rushes decision points is 1.04 nats, below the 1.39-nat entropy of a uniform distribution over four options, the typical candidate-set size. This difference indicates that observed choices are structured rather than uniformly random. It does not establish that users converge on one dominant option or that their preferences are homogeneous.

  7. Knowl 7 — User history helps GPT-4o and GPT-5 more than model scaling alone

    empirical result

    On the 8,293-event test set, GPT-4o accuracy rose from 0.3030 (95% Wilson CI [0.2931, 0.3130]) zero-shot to 0.3390 ([0.3288, 0.3493]) with user history. GPT-5 rose from 0.3090 ([0.2991, 0.3191]) zero-shot to 0.3423 ([0.3321, 0.3526]) with history. Thus, adding history yielded a larger improvement than scaling from GPT-4o to GPT-5, whose zero-shot gain was 0.6 percentage points. GPT-5 with history nevertheless remained below the 0.3639 popularity-baseline accuracy.

  8. Knowl 8 — Single choices can be converted into pairwise preference training data

    model/method

    A Rushes event in which a user chooses one option from a candidate set of size kk can be transformed into k−1k-1 pairwise preferences. For each unchosen alternative, the chosen option is recorded as preferred to that alternative, (ochosen≻orejected)(o_{\text{chosen}} \succ o_{\text{rejected}}). For the typical four-option event, this yields three comparisons. The resulting pairs can be used to train standard reward models or Direct Preference Optimization systems.

  9. Knowl 9 — Popularity prediction improves at later narrative depths

    empirical result

    The popularity baseline's accuracy varied by narrative depth: depth 0, 0.3048 (95% Wilson CI [0.2726, 0.3390], 735 events); depth 1, 0.3100 ([0.2854, 0.3357], 1,300); depth 2, 0.3740 ([0.3494, 0.3993], 1,441); depth 3, 0.3763 ([0.3620, 0.3908], 4,361); and depth 4, 0.4207 ([0.3761, 0.4666], 454). Two test events without matched depth metadata were excluded. The results show higher popularity-baseline accuracy later in the narrative, with its highest measured accuracy at depth 4.

  10. Knowl 10 — Rushes results are specific to an English-speaking gaming population

    limitation

    Rushes participants were recruited through Xbox Insiders, and the games and interface were presented in English. The resulting sample is therefore skewed toward gaming-literate users familiar with branching narratives and does not establish a universal baseline for human preferences; the paper cautions against generalizing the observed engagement patterns to non-gaming or non-English-speaking populations without further validation. The benchmark experiments are text-conditioned even though the collected stories include images, video, and audio. The reported generation setup uses GPT-4o, and outputs may vary with the language model, temperature, prompts, or formats.

Coverage note — The safety-screening counts, generated-option semantic-diversity analysis, and sparse-versus-active-player ablation are omitted because they are secondary quality-control or subgroup analyses rather than central dataset or engagement-prediction findings.

References

  1. 1.Dalia Ali, Dora Zhao, Allison Koenecke, and Orestis Papakyriakopoulos. 2025. Operationalizing pluralistic values in large language model alignment reveals trade-offs in safety, inclusivity, and model behavior.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. CoRR.
  3. 3.Louis Castricato, Spencer Frazier, Jonathan Balloch, and Mark Riedl. 2021. Fabula entropy indexing: Objective measures of story coherence. In Proceedings of the Third Workshop on Narrative Understanding, pages 84–94, Virtual. Association for Computational Linguistics.
  4. 4.Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, and Chelsea Finn. 2025. PERSONA: A reproducible testbed for pluralistic alignment. In Proceedings of the 31st International Conference on Computational Linguistics, pages 11348–11368, Abu Dhabi, UAE. Association for Computational Linguistics.
  5. 5.John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yi Wang, Yuqian Sun, Tiffany Wang, Shm Garanganao Almeda, Brett A. Halperin, Yuwen Lu, and Max Kreminski. 2025. Literarytaste: A preference dataset for creative writing personalization.
  6. 6.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  7. 7.Parsa Ghaffari and Chris Hokamp. 2025. Narrative studio: Visual narrative exploration using LLMs and Monte Carlo Tree Search.
  8. 8.Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weissbuch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. 2025. Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103.
  9. 9.Runsheng "Anson" Huang, Lara J. Martin, and Chris Callison-Burch. 2024. What-if: Exploring branching narratives by meta-prompting large language models.
  10. 10.Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In 2018 IEEE International Conference on Data Mining (ICDM), pages 197–206.
  11. 11.OpenAI. 2024. Gpt-4o system card. arXiv preprint, https://arxiv.org/abs/2410.21276. Accessed 2025-09-22.
  12. 12.OpenAI. 2025. Gpt-5 system card. https://openai.com/index/gpt-5-system-card. Accessed: 2025-09-22.
  13. 13.Long Ouyang et al. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  14. 14.Long Phan, Mantas Mazeika, Andy Zou, and Dan Hendrycks. 2025. Textquests: How good are LLMs at text-based video games?
  15. 15.Md Awsafur Rahman, Adam Gabrys, Doug Kang, Jingjing Sun, Tian Tan, and Ashwin Chandramouli. 2025. Likebench: Evaluating subjective likability in LLMs for personalization.
  16. 16.Mark O. Riedl and Vadim Bulitko. 2013. Interactive narrative: An intelligent systems approach. AI Magazine, 34(1):67–77.
  17. 17.Nick Walton. 2019. Ai dungeon: Dragon model upgrade. Aidungeon. io.
  18. 18.Shuangshuang Ying, Yunwen Li, Xingwei Qu, Xin Li, Sheng Jin, Minghao Liu, Zhoufutu Wen, Xeron Du, Tianyu Zheng, Yichi Zhang, Letian Ni, Yuyang Cheng, Qiguang Chen, Jingzhe Ding, Shengda Long, Wangchunshu Zhou, Jiazhan Feng, Wanjun Zhong, Libo Qin, Ge Zhang, Wenhao Huang, Wanxiang Che, and Chenghua Lin. 2025. Beyond correctness: Evaluating subjective writing preferences across cultures.
  19. 19.Hong Yu and Mark Riedl. 2013. Data-driven personalized drama management. Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, 9(1):191–197.

Citation

MLA
Xu, M., et al. “Rushes: A Human Preference Dataset for Pluralistic Alignment”. arXiv, 2026, http://arxiv.org/abs/2607.20767v1.
APA
Xu, M., Leandro, J., Rao, S., Xu, W., Jojic, N., DesGarennes, G., Quirk, C., & Dolan, B. (2026). Rushes: A Human Preference Dataset for Pluralistic Alignment. arXiv. http://arxiv.org/abs/2607.20767v1
Chicago
Xu, M., J. Leandro, S. Rao, et al. 2026. “Rushes: A Human Preference Dataset for Pluralistic Alignment”. arXiv. http://arxiv.org/abs/2607.20767v1.
Harvard
Xu, M. et al. (2026) “Rushes: A Human Preference Dataset for Pluralistic Alignment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.20767v1.
Vancouver
1. Xu M, Leandro J, Rao S, Xu W, Jojic N, DesGarennes G, Quirk C, Dolan B (2026) Rushes: A Human Preference Dataset for Pluralistic Alignment. arXiv

BibTeX

@article{xu2026rushes,
  title = {Rushes: A Human Preference Dataset for Pluralistic Alignment},
  author = {Xu, Michael and Leandro, Jorge and Rao, Sudha and Xu, Weijia and Jojic, Nebojsa and DesGarennes, Gabriel and Quirk, Chris and Dolan, Bill},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.20767v1},
  eprint = {2607.20767}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/