The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Hannah KirkAndrew M. BeanBertie VidgenPaul RöttgerScott Hale
Systematizes the evolution of human feedback learning across 95 studies to identify critical conceptual and practical challenges in aligning language models with subjective preferences and values.
Large Language Models are increasingly steered using human feedback to make systems helpful, honest, and harmless, yet the field faces critical uncertainties regarding how to gather and implement this feedback without embedding severe biases. This becomes especially problematic when models attempt to reflect inherently subjective human preferences, values, and cultural norms. Current practices frequently rely on simplifying assumptions that treat subjective human values as objective and universal, risking the deployment of systems misaligned with diverse populations.
The article systematically assesses how human feedback has been integrated into language models across time and identifies unresolved conceptual and methodological challenges in current model alignment practices. To achieve this, the authors conducted a structured literature survey of 95 empirical articles retrieved from the Association for Computational Linguistics (ACL) repository and arXiv up to February 2023. They systematically categorized 22 foundational papers from the pre-LLM era (2014–2019) and 50 LLM-era papers across conceptual motivations, data collection practices, workforce demographics, and technical integration methods.
The review revealed several critical patterns in current feedback learning. First, technical practices have shifted from early proxy measures and user simulations toward general-purpose language assistants fine-tuned via direct human feedback and reinforcement learning. Second, alignment goals rely heavily on abstract concepts—such as harmlessness or quality—that function as empty signifiers, meaning different things across different cultural and ethical contexts. Third, feedback pipelines rely on remarkably narrow annotator pools: a majority of reviewed studies employed fewer than 100 human evaluators, often dominated by demographic clusters of English-speaking, US-based crowdworkers aged 25 to 34 with higher education. In notable industry cases, as few as 20 individuals provided roughly 80% of the training feedback. Furthermore, the workforce remains poorly documented, with only 9 out of 50 LLM papers providing demographic breakdowns, while many influential industry papers lack independent peer review or public release of model artifacts.
These findings indicate that current language model alignment risks introducing severe ideological and cultural biases by centralizing value definitions within small, unrepresentative groups. Abstract guidelines fail to resolve low inter-annotator agreement, and reinforcement learning pipelines remain vulnerable to path dependency and hidden misalignment outside training distributions. Treating subjective value alignment as a purely technical optimization problem obscures critical governance, safety, and compliance risks for organizations deploying generative artificial intelligence.
To build more robust and equitable systems, practitioners and decision-makers must move away from treating human feedback as an objective ground truth. Organizations should diversify feedback sources through democratic, jury-based, and participatory sampling rather than relying exclusively on small crowdworker pools. Researchers should ground alignment guidelines in established legal frameworks, such as human rights law, to standardize high-stakes value judgments. Furthermore, teams must require standardized demographic reporting via data statements, associate individual feedback with annotator identifiers to model distributional disagreement, and subject alignment techniques to rigorous external peer review before deployment.
Confidence in these findings is supported by a systematic coding methodology and dual-reviewer consistency checks. However, the analysis is limited to English-language academic publications and preprints up to early 2023, omitting subsequent alignment methods such as Direct Preference Optimization as well as informal discussions from social media and practitioner forums. Readers should recognize that while technical mechanisms will continue to evolve, the underlying governance and demographic challenges identified by the article remain fundamental to feedback learning.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It introduces the foundational RLHF pipeline and annotator-driven instruction fine-tuning that the survey directly reviews and critiques regarding workforce demographics and value assumptions.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It formalizes the criteria of helpfulness and harmlessness via human feedback modeling, providing the seminal alignment framework examined by the survey.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). It provides the critical conceptual taxonomy of representational and allocational harms in NLP that grounds the survey's normative analysis of value alignment.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). It establishes the extensive categorization of ethical, representational, and social harms from language models that motivates the survey's investigation into subjective human values.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). It introduces demographic-conditioned simulation using language models, establishing a foundational perspective on proxying human sub-populations that the survey critically examines.
- Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). It details early industry practices of fine-tuning conversational models on crowdworker safety and quality annotations, serving as a primary case study analyzed in the review.
- Paper: Model Cards for Model Reporting, Margaret Mitchell et al. (2019). It pioneers standardized demographic reporting and model documentation, directly underpinning the survey's core policy recommendation to mandate data statements and annotator documentation.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). It establishes core benchmarks for evaluating social and demographic bias in pretrained representations, illustrating the persistent biases that post-training human feedback aims to address.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). It provides the benchmark and empirical baseline for neural toxic degeneration, highlighting the baseline safety failures that prompted reinforcement learning from human feedback.
- Paper: The Ethics of AI Ethics: An Evaluation of Guidelines, Thilo Hagendorff (2019). It analyzes the limitations and omissions of AI ethics guidelines, providing conceptual context for the survey's argument that abstract alignment concepts operate as empty signifiers.
- Paper: Whose Opinions Do Language Models Reflect?, Shibani Santurkar et al. (2023). It directly tests the ideological skew diagnosed in the survey by measuring whose opinions feedback-tuned language models reflect across diverse demographic subgroups.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). It investigates Reinforcement Learning from AI Feedback as an alternative to human annotator pipelines, exploring a direct solution to the human feedback scaling bottleneck identified in the survey.
- Paper: Towards Understanding Sycophancy in Language Models, Mrinank Sharma et al. (2023). It demonstrates how human preference learning unintentionally induces sycophancy, empirically illustrating the survey's critique of optimizing against unrepresentative or flawed human feedback.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). It develops a decoupled multi-objective RLHF method that resolves the annotator ambiguity between helpfulness and harmlessness highlighted in the survey.
- Paper: Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment, Rui Yang et al. (2024). It addresses the survey's call for dynamic value integration by proposing a context-conditioned alignment framework that adjusts to conflicting human preferences at inference time.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). It implements a global crowdsourced preference evaluation platform that operationalizes the survey's recommendation to collect broader, real-world user preference distributions.
- Paper: KTO: Model Alignment as Prospect Theoretic Optimization, Kawin Ethayarajh et al. (2024). It applies prospect theory to optimize alignment from simpler binary human feedback, extending the methodological evolution beyond traditional RLHF preference pairs.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). It reveals the brittleness of standard safety alignment when downstream users fine-tune models, extending the survey's warning about hidden misalignment risks.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). It theoretically unifies the post-RLHF landscape of direct preference learning methods that emerged after the survey's pre-2023 review period.
- Paper: From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models, Shangbin Feng et al. (2023). It maps political and ideological coordinates across models, showing how underlying training data disparities propagate to downstream task unfairness.
