The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

Hannah KirkAndrew M. BeanBertie VidgenPaul RöttgerScott Hale

article2023EMNLP70 citations

Systematizes the evolution of human feedback learning across 95 studies to identify critical conceptual and practical challenges in aligning language models with subjective preferences and values.

Listen

Large Language Models are increasingly steered using human feedback to make systems helpful, honest, and harmless, yet the field faces critical uncertainties regarding how to gather and implement this feedback without embedding severe biases. This becomes especially problematic when models attempt to reflect inherently subjective human preferences, values, and cultural norms. Current practices frequently rely on simplifying assumptions that treat subjective human values as objective and universal, risking the deployment of systems misaligned with diverse populations.

The article systematically assesses how human feedback has been integrated into language models across time and identifies unresolved conceptual and methodological challenges in current model alignment practices. To achieve this, the authors conducted a structured literature survey of 95 empirical articles retrieved from the Association for Computational Linguistics (ACL) repository and arXiv up to February 2023. They systematically categorized 22 foundational papers from the pre-LLM era (2014–2019) and 50 LLM-era papers across conceptual motivations, data collection practices, workforce demographics, and technical integration methods.

The review revealed several critical patterns in current feedback learning. First, technical practices have shifted from early proxy measures and user simulations toward general-purpose language assistants fine-tuned via direct human feedback and reinforcement learning. Second, alignment goals rely heavily on abstract concepts—such as harmlessness or quality—that function as empty signifiers, meaning different things across different cultural and ethical contexts. Third, feedback pipelines rely on remarkably narrow annotator pools: a majority of reviewed studies employed fewer than 100 human evaluators, often dominated by demographic clusters of English-speaking, US-based crowdworkers aged 25 to 34 with higher education. In notable industry cases, as few as 20 individuals provided roughly 80% of the training feedback. Furthermore, the workforce remains poorly documented, with only 9 out of 50 LLM papers providing demographic breakdowns, while many influential industry papers lack independent peer review or public release of model artifacts.

These findings indicate that current language model alignment risks introducing severe ideological and cultural biases by centralizing value definitions within small, unrepresentative groups. Abstract guidelines fail to resolve low inter-annotator agreement, and reinforcement learning pipelines remain vulnerable to path dependency and hidden misalignment outside training distributions. Treating subjective value alignment as a purely technical optimization problem obscures critical governance, safety, and compliance risks for organizations deploying generative artificial intelligence.

To build more robust and equitable systems, practitioners and decision-makers must move away from treating human feedback as an objective ground truth. Organizations should diversify feedback sources through democratic, jury-based, and participatory sampling rather than relying exclusively on small crowdworker pools. Researchers should ground alignment guidelines in established legal frameworks, such as human rights law, to standardize high-stakes value judgments. Furthermore, teams must require standardized demographic reporting via data statements, associate individual feedback with annotator identifiers to model distributional disagreement, and subject alignment techniques to rigorous external peer review before deployment.

Confidence in these findings is supported by a systematic coding methodology and dual-reviewer consistency checks. However, the analysis is limited to English-language academic publications and preprints up to early 2023, omitting subsequent alignment methods such as Direct Preference Optimization as well as informal discussions from social media and practitioner forums. Readers should recognize that while technical mechanisms will continue to evolve, the underlying governance and demographic challenges identified by the article remain fundamental to feedback learning.

arXiv: 2310.07629
Cover for The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

Abstract

Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values. In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories. First, we summarise the past, pre-LLM trends for integrating human feedback into language models. Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom. Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges.

Table of Contents

  • 1 Introduction
  • 2 Methods
  • 2.1 Selecting Articles
  • 2.2 Coding Articles
  • 3 The Past
  • 3.1 Conceptual Classification
  • 3.2 Methodological Classification
  • 4 The Present
  • 4.1 Conceptual Classification
  • 4.2 Methodological Classification
  • 4.2.1 Collecting Feedback
  • 4.2.2 Integrating Feedback
  • 5 Challenges and Recommendations for the Future
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgements
  • References
  • A Flowchart of Articles for Scoping the Review
  • B Code Book
  • C Additional Information on Reviewed Articles
  • C.1 Target Tasks
  • C.2 Evaluating Models
  • D Articles with Other Contribution Types

Knowls

  1. Knowl 1 — Standard Reinforcement Learning from Human Feedback (RLHF) Pipeline for Language Models

    model/method

    The standard process for aligning Large Language Models (LLMs) with direct human feedback consists of four sequential stages:

    1. Base Model Initialization and Demonstration Tuning: A pre-trained base autoregressive language model is initialized and optionally adapted via prompt engineering, supervised fine-tuning (SFT), or behavioral cloning on human-generated demonstration outputs representing desirable behaviors.

    2. Comparisons Data Collection: For a set of prompts xx, the model generates multiple candidate completions y1,y2,…,yKy_1, y_2, \dots, y_K. Human annotators (or crowdsourced raters) evaluate these outputs and produce comparative rankings or pairwise preference labels denoting which output is preferred (e.g., yw≻yly_w \succ y_l, where ywy_w is the preferred response and yly_l is the less preferred response).

    3. Preference Reward Model (PM) Training: A neural reward model rθ(x,y)r_\theta(x, y) parameterized by θ\theta is trained on the comparative feedback dataset to output a scalar reward. The model is typically optimized using a binary cross-entropy loss based on the Bradley-Terry preference model:

    LPM(θ)=−E(x,yw,yl)[log⁡σ(rθ(x,yw)−rθ(x,yl))]\mathcal{L}_{\text{PM}}(\theta) = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma\left(r_\theta(x, y_w) - r_\theta(x, y_l)\right) \right]

    where σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} represents the sigmoid function, expressing the probability P(yw≻yl∣x)P(y_w \succ y_l | x).

    1. Reinforcement Learning Policy Optimization: An autoregressive policy πϕ(y∣x)\pi_\phi(y|x) initialized from the fine-tuned model is trained using reinforcement learning (predominantly Proximal Policy Optimization, PPO) to maximize the scalar reward signal from rθ(x,y)r_\theta(x, y), augmented with a Kullback-Leibler (KL) divergence penalty to prevent excessive deviation from the initial reference policy πref\pi_{\text{ref}}:

    max⁡ϕEx∼D,y∼πϕ(y∣x)[rθ(x,y)−βDKL(πϕ(y∣x)∥πref(y∣x))]\max_\phi \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_\phi(y|x)} \left[ r_\theta(x, y) - \beta D_{\text{KL}}\left(\pi_\phi(y|x) \parallel \pi_{\text{ref}}(y|x)\right) \right]

    where β\beta is a hyperparameter controlling the strength of the KL penalty constraint.

  2. Knowl 2 — Mechanisms for Integrating Human Feedback Across Language Model Lifecycle Stages

    model/method

    Human feedback is integrated into Large Language Models across multiple intervention stages beyond standard RL fine-tuning:

    • Preference Pre-training: Alignment objectives are introduced during the initial self-supervised pre-training phase (e.g., conditional pre-training on reward-annotated data), which provides greater robustness than exclusively fine-tuning an already pre-trained model.
    • Supervised Preference Fine-Tuning: Supervised adaptation using filtered or weighted datasets, including contrastive learning losses that simultaneously maximize the likelihood of positive demonstrations and minimize negative demonstrations, and Chain of Hindsight (CoH) tuning, where models are trained to condition generation on historical natural language feedback and numerical ratings.
    • Decoding-Time Guidance (Generator-Discriminator Architectures): Ensembles where a base generator model's next-token output probabilities are modified at inference time by an external discriminator, classifier (e.g., toxic text detector), or expert/anti-expert models (e.g., DExperts) without modifying base model weights.
    • Inference-Time Re-ranking and Best-of-NN Sampling: The policy generates NN candidate responses for a given input prompt, and a trained preference reward model scores and re-ranks the candidates, selecting the top-scoring output. This test-time sampling mechanism often matches or exceeds the performance of full RL policy optimization.
    • Prompting and Context Distillation: Incorporating long, author-written demonstration prompts or multi-step chain-of-thought moral principles directly into the LLM context, or training a new student model via context distillation to mimic a prompt-guided teacher model.
  3. Knowl 3 — Feedback Collection Modalities and Proxy Signal Paradigms

    model/method

    Approaches for obtaining feedback to steer language model behavior encompass both direct human elicitation and indirect proxy signals:

    • Explicit Direct Feedback: Human raters provide explicit comparative judgments (pairwise preferences, KK-way rankings), fine-grained ordinal ratings (Likert scales on attributes like coherence, groundedness, informativeness, and safety), binary classifications, targeted natural language critiques, or direct text edits and corrections.
    • Adversarial Red-Teaming: Human experts or crowdworkers deliberately craft adversarial prompts to probe system boundaries, elicit toxic or unsafe completions, and generate negative feedback data for iterative hardening.
    • Implicit Interaction Signals: Passive behavioral proxies derived from real-world human interactions, such as conversation length, user sentiment shifts, follow-up query structures, and engagement patterns in dialogue sessions.
    • Simulated and Oracle Feedback: Automated reward metrics (such as ROUGE or SacreBLEU against reference texts), pre-trained downstream classifiers (for toxicity, moral sentiment, or politeness), or parallel translation corpora used as automated proxies for human preferences.
    • AI-Assisted and Constitutional Feedback: Leveraging a high-level set of human-written principles (a constitution) while employing language models to critique, revise, and generate synthetic preference labels to guide alignment training.
  4. Knowl 4 — Non-Universality and Semantic Ambiguity of Values and Preferences in AI Alignment

    theoretical result

    The conceptual foundations of feedback learning in language models suffer from two structural vulnerabilities regarding normative goals:

    1. The Empty Signifier Problem: Target alignment concepts such as 'helpfulness', 'honesty', 'harmlessness' (the 3Hs), 'quality', and 'safety' are broadly accepted only because they are abstract. In practice, they function as empty signifiers—terms that carry universally positive connotations but lack concrete definitions, causing individuals from different cultural, religious, and socio-political backgrounds to interpret them in fundamentally conflicting ways.
    2. Conflation of Preferences and Values: Existing literature blurs the distinction between individual preferences (subjective, idiosyncratic stylistic tastes, writing tone, instrumental utility) and human values (normative ethical principles, human rights, moral duties). Treating societal safety and ethical constraints with the same optimization tools used for subjective user preferences creates tension between personal customization and universal harm mitigation.

    To resolve these conceptual gaps, alignment targets must be grounded in formal legal frameworks (such as international human rights law) and governed by principles of subsidiarity—establishing clear boundaries between non-negotiable safety guardrails and personalized subjective configurations.

  5. Knowl 5 — Epistemic Limits and Generalization Bounds of Incomplete Feedback in LLMs

    limitation

    Aligning LLMs via human feedback relies on the assumption that models generalize robustly from sparse feedback samples to unseen distributions. However, several structural factors limit this generalization:

    • Distributional Incompleteness: The space of possible input prompts and generative behaviors is vast. Partial feedback on specific violations (such as safe responses to historical denialism) does not guarantee appropriate generalization to out-of-domain harmful queries (such as weapons synthesis or medical misinformation).
    • Reinforcement Learning Path Dependence: In policy gradient optimization (e.g., PPO), the sequence and order in which feedback examples are presented influence optimization trajectories, frequently trapping policies in suboptimal local equilibria.
    • Superficial Alignment vs. Latent Misalignment: Models fine-tuned on preference data often learn to produce superficial markers of alignment that satisfy reward models on in-distribution evaluations, while retaining latent misaligned behaviors that can be exposed by adversarial red-teaming or distribution shift.
  6. Knowl 6 — Methodological Pitfalls in Operationalizing Subjective Annotations and Aggregating Feedback

    limitation

    Operationalizing abstract human values into discrete empirical feedback signals introduces significant measurement errors:

    • Prescriptive vs. Subjective Annotation Dilemma: Prescriptive guidelines attempt to eliminate ambiguity by enforcing strict rules, but studies continue to document low inter-annotator agreement on complex tasks. Conversely, subjective paradigms instruct annotators to use personal judgment, resulting in noisy, inconsistent labels.
    • Cognitive and Instrument Biases: Rater responses are vulnerable to survey design artifacts, scale presentation biases, order effects, intransitivity of preferences, and the absence of null-vote options (e.g., forcing a preference choice when both candidate completions are harmful or low quality).
    • The Aggregation Fallacy: Standard alignment pipelines aggregate human feedback via majority voting or scalar reward regression, which smooths over genuine disagreements and erases minority viewpoints on contentious ethical and subjective topics.
  7. Knowl 7 — Annotator Concentration and Demographic Disparities in LLM Feedback Datasets

    empirical result

    Empirical analysis of LLM alignment workforces reveals extreme labor concentration and demographic skew:

    • Workforce Concentration: Most feedback datasets rely on small rater pools (fewer than 100 workers). In prominent benchmark datasets, a fraction of participants accounts for the vast majority of training data (e.g., 20 annotators contributed 80% of feedback in Anthropic's Helpful and Harmless dataset, and 5 annotators provided 50% of the feedback in WebGPT).
    • Demographic Homogeneity: The annotator workforce across crowdsourcing platforms (Amazon Mechanical Turk, Upwork, Scale AI, Surge AI) is predominantly US-based, English-speaking, tertiary-educated (holding Master's degrees), and concentrated between ages 25 and 34, producing a 'tyranny of the crowdworker' that injects specific cultural, political, and socio-economic biases into reward models.
    • Reporting Deficits: Across 50 reviewed papers applying feedback learning to LLMs, only 9 papers (18%) documented annotator demographic distributions. The majority omitted workforce size, compensation rates, demographic profiles, and potential rater artifacts.
  8. Knowl 8 — Historical Shift from Pre-LLM Proxied Bandit Learning to Direct LLM Alignment

    empirical result

    The methodology of learning from feedback transitioned across two historical eras:

    • The Pre-LLM Era (2014–2019): Encompassed task-specific recurrent architectures (such as LSTMs) applied to constrained NLP domains (neural machine translation, semantic parsing, task-oriented dialogue). Direct human feedback was rare due to cost; systems relied primarily on bandit structured prediction, simulated user rewards derived from parallel corpora (e.g., TED talks, Europarl), and heuristic proxies (such as sentiment scores, response length, or vocabulary overlap).
    • The LLM Era (2018–2023): Characterized by general-purpose, pre-trained autoregressive foundation models evaluated across multi-task instruction following and open-ended dialogue. Feedback shifted from simulated proxies to large-scale direct human pairwise comparisons, enabling the training of parameterized reward models and iterative multi-stage alignment (SFT, RLHF, rejection sampling).
  9. Knowl 9 — Target Task Taxonomy and the Alignment Tax in Language Model Adaptation

    empirical result

    Feedback learning is applied across a distinct taxonomy of natural language tasks, accompanied by a performance trade-off known as the alignment tax:

    • Target Task Distribution: Alignment research covers general text generation, instruction following, open-ended conversational dialogue, open-book generative question answering, abstractive text summarization (including recursive book-length summarization), toxic language mitigation, and moral/normative judgment classification.
    • The Alignment Tax: Imposing preference and safety constraints on an LLM can degrade its performance on standard NLP capability benchmarks, general reasoning, or truthfulness metrics. Quantitative evaluation of aligned models balances alignment rewards (e.g., Elo win-rates, reduction in Perspective API toxicity scores, PII violations) against capability retention (measured via KL divergence from the unconstrained base model and accuracy on NLP reasoning benchmarks).
  10. Knowl 10 — Taxonomy of Coding Dimensions for Language Model Human Feedback Research

    model/method

    A comprehensive analytical schema for characterizing human feedback research in language models comprises five primary categories:

    1. Metadata and Scope: Contribution type (training/alignment, capability evaluation/benchmarking, or social preference prediction) and model family (pre-LLM recurrent architectures vs. self-supervised pre-trained LLMs).
    2. Conceptual Dimensions: Feedback terminology ('preferences' vs. 'values'), primary motivations (value alignment vs. capability enhancement), target attributes, formal theoretical grounding, universality assumptions (universal vs. culturally contextual), and annotator interpretative freedom (prescriptive vs. subjective paradigm).
    3. Labor and Feedback Elicitation: Data generation mechanism (direct explicit, indirect implicit, simulated/oracle, AI feedback), feedback format (pairwise comparison, Likert scale, text edit, natural language feedback), annotator pool composition (crowdworkers, in-house teams, study authors), workforce size, and demographic documentation transparency.
    4. Technical Integration: Lifecycle intervention stage (pre-training, fine-tuning, prompting, inference decoding), dataset volume, algorithmic architecture (PPO, contrastive loss, Chain of Hindsight, best-of-NN), and evaluation metrics (Elo ratings, win rates, ROUGE, SacreBLEU, KL divergence).
    5. Procedural Attributes: Author affiliation (academia, industry, mixed), artifact availability (open datasets/models vs. closed API access), and peer-review status.

Coverage note — None was omitted; all major survey taxonomies, algorithmic RLHF formulations, empirical findings regarding annotator demographics and alignment tax, and the five conceptual/practical challenges were comprehensively extracted.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’21, pages 298–306, New York, NY, USA. Association for Computing Machinery.
  2. 2.Kushal Arora, Kurt Shuster, Sainbayar Sukhbaatar, and Jason Weston. 2022. Director: Generator-Classifiers For Supervised Language Modeling. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 512–526, Online only. Association for Computational Linguistics.
  3. 3.Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24.
  4. 4.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. 2021. A General Language Assistant as a Laboratory for Alignment. arXiv:2112.00861 [cs].
  5. 5.Luigi Asprino, Luana Bulla, Stefano De Giorgis, Aldo Gangemi, Ludovica Marinucci, and Misael Mongiovi. 2022. Uncovering Values: Detecting Latent Moral Content from Natural Language with Explainable and Non-Trained Methods. In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 33–41, Dublin, Ireland and Online. Association for Computational Linguistics.
  6. 6.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022a. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv: 2204.05862.
  7. 7.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022b. Constitutional AI: Harmlessness from AI Feedback. arXiv: 2212.08073.
  8. 8.Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan, Michael Henry Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matthew M. Botvinick, and Christopher Summerfield. 2022. Fine-tuning language models to find agreement among humans with diverse preferences. arXiv: 2211.15006v1.
  9. 9.Yejin Bang, Tiezheng Yu, Andrea Madotto, Zhaojiang Lin, Mona Diab, and Pascale Fung. 2022. Enabling Classifiers to Make Judgements Explicitly Aligned with Human Values. arXiv: 2210.07652v1.
  10. 10.Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  11. 11.Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Diaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022. Power to the People? Opportunities and Challenges for Participatory AI. In Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’22, pages 1–8, New York, NY, USA. Association for Computing Machinery.
  12. 12.Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan. 2022. Measuring Progress on Scalable Oversight for Large Language Models. arXiv:2211.03540 [cs].
  13. 13.Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych. 2019. Better Rewards Yield Better Summaries: Learning to Summarise Without References. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3110–3120, Hong Kong, China. Association for Computational Linguistics.
  14. 14.Sabrina Campano, Jessica Durand, and Chloé Clavel. 2014. Comparative analysis of verbal alignment in human-human and human-agent interactions. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 4415–4422, Reykjavik, Iceland. European Language Resources Association (ELRA).
  15. 15.Louis Castricato, Alexander Havrilla, Shahbuland Matiana, Michael Pieler, Anbang Ye, Ian Yang, Spencer Frazier, and Mark Riedl. 2022. Robust preference learning for storytelling via contrastive reinforcement learning. arXiv: 2210.07792v2 [cs.CL].
  16. 16.Kai-Wei Chang, Vinodkumar Prabhakaran, and Vicente Ordonez. 2019. Bias and fairness in natural language processing. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): Tutorial Abstracts.
  17. 17.Mecca Chiesa and Sandy Hobbs. 2008. Making sense of social research: How useful is the hawthorne effect? European Journal of Social Psychology, 38(1):67–74.
  18. 18.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  19. 19.Yating Chuang and Laura Schechter. 2015. Stability of experimental and survey measures of risk, time, and social preferences: A review and some new results. Journal of development economics, 117:151–170.
  20. 20.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  21. 21.Simon De Deyne, Amy Perfors, and Daniel J Navarro. 2016. Predicting human similarity judgments with distributional models: The value of word associations. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1861–1870, Osaka, Japan. The COLING 2016 Organizing Committee.
  22. 22.Nicola Dell, Vidya Vaidyanathan, Indrani Medhi, Edward Cutrell, and William Thies. 2012. " yours is better!" participant response bias in hci. In Proceedings of the sigchi conference on human factors in computing systems, pages 1321–1330.
  23. 23.Floris Den Hengst, Mark Hoogendoorn, Frank Van Harmelen, and Joost Bosman. 2019. Reinforcement learning for personalized dialogue management. In IEEE/WIC/ACM International Conference on Web Intelligence, pages 59–67.
  24. 24.Liqiong Deng and Marshall Scott Poole. 2010. Affect in web interfaces: A study of the impacts of web page visual complexity and order. Mis Quarterly, pages 711–730.
  25. 25.Yang Deng, Yaliang Li, Wenxuan Zhang, Bolin Ding, and Wai Lam. 2022. Toward personalized answer generation in e-commerce via multi-perspective preference modeling. ACM Transactions on Information Systems (TOIS), 40(4):1–28.
  26. 26.Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yulia Tsvetkov, M. R. Leiser, and Saif Mohammad. 2023. Assessing Language Model Deployment with Risk Cards. arXiv:2303.18190 [cs].
  27. 27.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs].
  28. 28.Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017. Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 484–495, Vancouver, Canada. Association for Computational Linguistics.
  29. 29.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack. arXiv:1908.06083 [cs].
  30. 30.Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023. RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment. arXiv: 2304.06767.
  31. 31.Guillaume Dubuisson Duplessis, Chloé Clavel, and Frédéric Landragin. 2017. Automatic Measures to Characterise Verbal Alignment in Human-Agent Interaction. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 71–81, Saarbrücken, Germany. Association for Computational Linguistics.
  32. 32.Jessica Ficler and Yoav Goldberg. 2017. Controlling Linguistic Style Aspects in Neural Language Generation. In Proceedings of the Workshop on Stylistic Variation, pages 94–104, Copenhagen, Denmark. Association for Computational Linguistics.
  33. 33.Ronald Fischer. 2017. Personality, Values, Culture. In Personality, Values, Culture: An Evolutionary Approach, Culture and Psychology, pages i–ii. Cambridge University Press, Cambridge.
  34. 34.Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. Social Chemistry 101: Learning to Reason about Social and Moral Norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653–670, Online. Association for Computational Linguistics.
  35. 35.Hershey H Friedman, Paul J Herskovitz, and Simcha Pollack. 1994. The biasing effects of scale-checking styles on response to a likert scale. In Proceedings of the American statistical association annual conference: survey research methods, volume 792, pages 792–795.
  36. 36.Richard Futrell and Roger P. Levy. 2019. Do RNNs learn human-like abstract word order preferences? In Proceedings of the Society for Computation in Linguistics (SCiL) 2019, pages 50–59.
  37. 37.Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv: 2209.07858v2.
  38. 38.Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020. Dialogue Response Ranking Training with Large-Scale Human Feedback Data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 386–395, Online. Association for Computational Linguistics.
  39. 39.Yang Gao, Christian M. Meyer, and Iryna Gurevych. 2018. APRIL: Interactively learning to summarise by combining active preference learning and reinforcement learning. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 4120–4130, Brussels, Belgium. Association for Computational Linguistics.
  40. 40.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  41. 41.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1161–1166.
  42. 42.Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Sona Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022. Improving alignment of dialogue agents via targeted human judgements. arXiv: 2209.14375v1.
  43. 43.Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–19.
  44. 44.David Gros, Yu Li, and Zhou Yu. 2021. The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human Identity. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6999–7013, Online. Association for Computational Linguistics.
  45. 45.Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019. Learning from Dialogue after Deployment: Feed Yourself, Chatbot! In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3667–3684, Florence, Italy. Association for Computational Linguistics.
  46. 46.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020. Aligning AI With Shared Human Values. arXiv: 2008.02275v5.
  47. 47.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-term Memory. Neural computation, 9:1735–80.
  48. 48.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor. arXiv: 2212.09689v1.
  49. 49.Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, Gabriela Moreno, Christina Park, Tingyee E. Chang, Jenna Chin, Christian Leong, Jun Yen Leung, Arineh Mirinjian, and Morteza Dehghani. 2020. Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment. Social Psychological and Personality Science, 11(8):1057–1071. Publisher: SAGE Publications Inc.
  50. 50.Tom Hosking, Phil Blunsom, and Max Bartolo. 2023. Human Feedback is not Gold Standard. arXiv: 2309.16349.
  51. 51.Gary Hsieh and Rafał Kocielnik. 2016. You get who you pay for: The impact of incentives on participation bias. In Proceedings of the 19th ACM conference on computer-supported cooperative work & social computing, pages 823–835.
  52. 52.Xiaolei Huang, Alexandra Wormley, and Adam Cohen. 2022. Learning to Adapt Domain Shifts of Moral Values via Instance Weighting. arXiv: 2204.07603v2.
  53. 53.Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019. Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog. arXiv:1907.00456 [cs, stat].
  54. 54.Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2020. Human-centric dialog training via offline reinforcement learning. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pages 3985–4003, Online. Association for Computational Linguistics.
  55. 55.Sophie Jentzsch, Patrick Schramowski, Constantin Rothkopf, and Kristian Kersting. 2019. Semantics Derived Automatically from Language Corpora Contain Human-like Moral Choices. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’19, pages 37–44, New York, NY, USA. Association for Computing Machinery.
  56. 56.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022. Survey of Hallucination in Natural Language Generation.
  57. 57.Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi. 2022. Can Machines Learn Morality? The Delphi Experiment. arXiv:2110.07574 [cs].
  58. 58.Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022. When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment. arXiv: 2210.01478v3.
  59. 59.Anna Jobin, Marcello Ienca, and Effy Vayena. 2019. The global landscape of ai ethics guidelines. Nature Machine Intelligence, 1(9):389–399.
  60. 60.Da Ju, Jing Xu, Y.-Lan Boureau, and Jason Weston. 2022. Learning from data in the mixed adversarial non-adversarial case: Finding the helpers and ignoring the trolls. arXiv:2208.03295 [cs].
  61. 61.Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, and Benno Stein. 2022. Identifying the Human Values behind Arguments. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4459–4471, Dublin, Ireland. Association for Computational Linguistics.
  62. 62.Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models. In Advances in Neural Information Processing Systems, volume 34, pages 2611–2624.
  63. 63.Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale. 2023a. The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising "Alignment" in Large Language Models. arXiv: 2310.02457.
  64. 64.Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale. 2023b. Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback. arXiv:2303.05453.
  65. 65.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez. 2023. Pre-training Language Models with Human Preferences. arXiv:2302.08582 [cs].
  66. 66.Julia Kreutzer, Artem Sokolov, and Stefan Riezler. 2017. Bandit Structured Prediction for Neural Sequence-to-Sequence Learning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1503–1513, Vancouver, Canada. Association for Computational Linguistics.
  67. 67.Julia Kreutzer, Joshua Uyheng, and Stefan Riezler. 2018. Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1777–1788, Melbourne, Australia. Association for Computational Linguistics.
  68. 68.Ernesto Laclau. 2005. On populist reason. Verso, London New York (N.Y.).
  69. 69.Dean Lacy. 2001. A theory of nonseparable preferences in survey responses. American Journal of Political Science, pages 239–258.
  70. 70.Carolin Lawrence and Stefan Riezler. 2018. Improving a Neural Semantic Parser by Counterfactual Learning from Human Bandit Feedback. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1820–1830, Melbourne, Australia. Association for Computational Linguistics.
  71. 71.Carolin Lawrence, Artem Sokolov, and Stefan Riezler. 2017. Counterfactual Learning from Bandit Feedback under Deterministic Logging : A Case Study in Statistical Machine Translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2566–2576, Copenhagen, Denmark. Association for Computational Linguistics.
  72. 72.Leonard Lee, On Amir, and Dan Ariely. 2009. In search of homo economicus: Cognitive noise and the role of emotion in preference consistency. Journal of consumer research, 36(2):173–187.
  73. 73.Jiwei Li, Alexander H. Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston. 2017a. Dialogue Learning With Human-In-The-Loop. arXiv:1611.09823 [cs].
  74. 74.Jiwei Li, Alexander H. Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston. 2017b. Learning through Dialogue Interactions by Asking Questions. arXiv:1612.04936 [cs].
  75. 75.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016. Deep Reinforcement Learning for Dialogue Generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1192–1202, Austin, Texas. Association for Computational Linguistics.
  76. 76.Manling Li, Ying Lin, Joseph Hoover, Spencer Whitehead, Clare Voss, Morteza Dehghani, and Heng Ji. 2019. Multilingual Entity, Relation, Event and Human Value Extraction. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 110–115, Minneapolis, Minnesota. Association for Computational Linguistics.
  77. 77.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
  78. 78.Ying Lin, Joe Hoover, Morteza Dehghani, Marlon Mooijman, and Heng Ji. 2017. Acquiring Background Knowledge to Improve Moral Value Prediction. arXiv: 1709.05467v1.
  79. 79.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021a. DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  80. 80.Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018. Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2060–2069, New Orleans, Louisiana. Association for Computational Linguistics.
  81. 81.Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023a. Chain of Hindsight Aligns Language Models with Feedback. arXiv: 2302.02676.
  82. 82.Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023b. Languages are rewards: Chain of hindsight finetuning using human feedback. arXiv: 2302.02676v2 [cs.LG].
  83. 83.Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, Lili Wang, and Soroush Vosoughi. 2021b. Mitigating Political Bias in Language Models through Reinforced Calibration. Proceedings of the AAAI Conference on Artificial Intelligence, 35(17):14857–14866. Number: 17.
  84. 84.Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X. Liu, and Soroush Vosoughi. 2023c. Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits. arXiv: 2301.00355v2.
  85. 85.Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M. Dai, Diyi Yang, and Soroush Vosoughi. 2023d. Training Socially Aligned Language Models in Simulated Human Society. arXiv: 2305.16960.
  86. 86.Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022. Aligning Generative Language Models with Human Values. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 241–252, Seattle, United States. Association for Computational Linguistics.
  87. 87.Nicholas Lourie, Ronan Le Bras, and Yejin Choi. 2021. SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes. Proceedings of the AAAI Conference on Artificial Intelligence, 35(15):13470–13479. Number: 15.
  88. 88.Hua Lu, Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2022. Towards Boosting the Open-Domain Chatbot with Human Feedback. arXiv: 2208.14165v1.
  89. 89.Li Lucy and David Bamman. 2021. Gender and Representation Bias in GPT-3 Generated Stories. In Proceedings of the Third Workshop on Narrative Understanding, pages 48–55, Virtual. Association for Computational Linguistics.
  90. 90.Claude Lévi-Strauss. 1987. Introduction to the work of Marcel Mauss. Routledge & Kegan Paul, London.
  91. 91.Hotaka Maeda. 2015. Response option configuration of online administered likert scales. International Journal of Social Research Methodology, 18(1):15–26.
  92. 92.Tushar Maheshwari, Aishwarya N. Reganti, Samiksha Gupta, Anupam Jamatia, Upendra Kumar, Björn Gambäck, and Amitava Das. 2017. A Societal Sentiment Analysis: Predicting the Values and Ethics of Individuals by Analysing Social Media Content. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 731–741, Valencia, Spain. Association for Computational Linguistics.
  93. 93.Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, and Julian McAuley. 2019. Generating personalized recipes from historical user preferences. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 5976–5982, Hong Kong, China. Association for Computational Linguistics.
  94. 94.Donald Martin Jr., Vinodkumar Prabhakaran, Jill Kuhlberg, Andrew Smart, and William S. Isaac. 2020. Participatory Problem Formulation for Fairer Machine Learning Through Community Based System Dynamics. arXiv:2005.07572 [cs, stat].
  95. 95.Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese. 2022. Teaching language models to support answers with verified quotes. arXiv:2203.11147 [cs].
  96. 96.Shachar Mirkin and Jean-Luc Meunier. 2015. Personalized machine translation: Predicting translational preferences. In Proceedings of the 2015 conference on empirical methods in natural language processing, pages 2019–2025, Lisbon, Portugal. Association for Computational Linguistics.
  97. 97.Shachar Mirkin, Scott Nowson, Caroline Brun, and Julien Perez. 2015. Motivating Personality-aware Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1102–1108, Lisbon, Portugal. Association for Computational Linguistics.
  98. 98.Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Barriere Valentin, Doratossadat Dastgheib, Omid Ghahroodi, Mohammad Ali Sadraei, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, and Benno Stein. 2023. The touché23-ValueEval dataset for identifying human values behind arguments. arXiv: 2301.13771v1 [cs.CL].
  99. 99.Kaixiang Mo, Shuangyin Li, Yu Zhang, Jiajun Li, and Qiang Yang. 2016. Personalizing a dialogue system with transfer reinforcement learning. arXiv: 1610.02891v3 [cs.AI].
  100. 100.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  101. 101.Md Sultan Al Nahian, Spencer Frazier, Mark Riedl, and Brent Harrison. 2020. Learning norms from stories: A prior for value aligned agents. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 124–130.
  102. 102.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. WebGPT: Browser-assisted question-answering with human feedback. arXiv: 2112.09332v3.
  103. 103.Duy-Hung Nguyen, Nguyen Viet Dung Nghiem, Bao-Sinh Nguyen, Dung Tien Tien Le, Shahab Sabahi, Minh-Tien Nguyen, and Hung Le. 2022. Make The Most of Prior Data: A Solution for Interactive Text Summarization with Preference Feedback. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1919–1930, Seattle, United States. Association for Computational Linguistics.
  104. 104.Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber. 2017. Reinforcement Learning for Bandit Neural Machine Translation with Simulated Human Feedback. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1464–1474, Copenhagen, Denmark. Association for Computational Linguistics.
  105. 105.Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131–9143.
  106. 106.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. arXiv: 2203.02155v1.
  107. 107.Xiangyu Peng, Siyan Li, Spencer Frazier, and Mark Riedl. 2020. Reducing Non-Normative Text Generation from Language Models. In Proceedings of the 13th International Conference on Natural Language Generation, pages 374–383, Dublin, Ireland. Association for Computational Linguistics.
  108. 108.Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2022. Discovering Language Model Behaviors with Model-Written Evaluations. arXiv:2212.09251 [cs].
  109. 109.ME Peters, M Neumann, M Iyyer, M Gardner, C Clark, K Lee, and L Zettlemoyer. 2018. Deep contextualized word representations. arxiv 2018. arXiv preprint arXiv:1802.05365, 12.
  110. 110.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of The Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138.
  111. 111.Vjosa Preniqi, Kyriaki Kalimeri, and Charalampos Saitis. 2022. "More Than Words": Linking Music Preferences and Moral Values Through Lyrics. arXiv: 2209.01169v1.
  112. 112.Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, and Chandra Bhagavatula. 2022. Reinforced clarification question generation with defeasibility rewards for disambiguating social and moral situations. arXiv: 2212.10409v1 [cs.CL].
  113. 113.Liang Qiu, Yizhou Zhao, Jinchao Li, Pan Lu, Baolin Peng, Jianfeng Gao, and Song-Chun Zhu. 2021. ValueNet: A New Dataset for Human Value Driven Dialogue System. arXiv: 2112.06346v1.
  114. 114.Ella Rabinovich, Raj Nath Patel, Shachar Mirkin, Lucia Specia, and Shuly Wintner. 2017. Personalized Machine Translation: Preserving Original Author Traits. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, pages 1074–1084, Valencia, Spain. Association for Computational Linguistics.
  115. 115.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv: 2305.18290.
  116. 116.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  117. 117.Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective nlp tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190.
  118. 118.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490, Online. Association for Computational Linguistics.
  119. 119.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. Social IQa: Commonsense Reasoning about Social Interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4463–4473, Hong Kong, China. Association for Computational Linguistics.
  120. 120.Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2022. Training Language Models with Language Feedback. arXiv:2204.14146 [cs].
  121. 121.Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin Rothkopf, and Kristian Kersting. 2019. BERT has a Moral Compass: Improvements of ethical and moral values of machines. arXiv: 1912.05238v1.
  122. 122.Olga Seminck and Pascal Amsili. 2017. A Computational Model of Human Preferences for Pronoun Resolution. In Proceedings of the Student Research Workshop at the 15th Conference of the European Chapter of the Association for Computational Linguistics, pages 53–63, Valencia, Spain. Association for Computational Linguistics.
  123. 123.Eric Michael Smith, Melissa Hall Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. "I’m sorry to hear that": Finding bias in language models with a holistic descriptor dataset.
  124. 124.Irene Solaiman and Christy Dennison. 2021. Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets. In Advances in Neural Information Processing Systems, volume 34, pages 5861–5873. Curran Associates, Inc.
  125. 125.Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023. Preference Ranking Optimization for Human Alignment. arXiv: 2306.17492.
  126. 126.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020. Learning to summarize from human feedback. arXiv: 2009.01325v3.
  127. 127.Yi Tay, Donovan Ong, Jie Fu, Alvin Chan, Nancy Chen, Anh Tuan Luu, and Chris Pal. 2020. Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social Preferences. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5369–5373, Online. Association for Computational Linguistics.
  128. 128.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022. LaMDA: Language Models for Dialog Applications. arXiv:2201.08239 [cs].
  129. 129.Amos Tversky. 1969. Intransitivity of preferences. Psychological review, 76(1):31.
  130. 130.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need.
  131. 131.Dongqi Wang, Haoran Wei, Zhirui Zhang, Shujian Huang, Jun Xie, and Jiajun Chen. 2021. Non-Parametric Online Learning from Human Feedback for Neural Machine Translation. arXiv: 2109.11136v3.
  132. 132.Xin Wang, Jianan Wang, Yuanchao Liu, Xiaolong Wang, Zhuoran Wang, and Baoxun Wang. 2017. Predicting Users’ Negative Feedbacks in Multi-Turn Human-Computer Dialogues. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 713–722, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  133. 133.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-Instruct: Aligning Language Model with Self Generated Instructions. arXiv: 2212.10560v1.
  134. 134.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in Detoxifying Language Models.
  135. 135.J Christopher Westland. 2022. Information loss and bias in likert survey responses. Plos one, 17(7):e0271949.
  136. 136.Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively Summarizing Books with Human Feedback. arXiv: 2109.10862v2.
  137. 137.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-Grained Human Feedback Gives Better Rewards for Language Model Training. arXiv: 2306.01693.
  138. 138.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021a. Bot-Adversarial Dialogue for Safe Conversational Agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2950–2968, Online. Association for Computational Linguistics.
  139. 139.Jing Xu, Da Ju, Margaret Li, Y.-Lan Boureau, Jason Weston, and Emily Dinan. 2021b. Recipes for Safety in Open-domain Chatbots. arXiv:2010.07079 [cs].
  140. 140.Jing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora, Y.-Lan Boureau, and Jason Weston. 2022. Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback. arXiv: 2208.03270v2.
  141. 141.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023. RRHF: Rank Responses to Align Language Models with Human Feedback without tears. arXiv: 2304.05302.
  142. 142.Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213, Melbourne, Australia. Association for Computational Linguistics.
  143. 143.Jieyu Zhao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Kai-Wei Chang. 2021. Ethical-Advice Taker: Do Language Models Understand Natural Language Interventions? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4158–4164, Online. Association for Computational Linguistics.
  144. 144.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. LIMA: Less Is More for Alignment. arXiv: 2305.11206.
  145. 145.Ruijie Zhou, Soham Deshmukh, Jeremiah Greer, and Charles Lee. 2021. NaRLE: Natural language models using reinforcement learning with emotion feedback. arXiv: 2110.02148v1 [cs.CL].
  146. 146.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. arXiv: 1909.08593v2.
  147. 147.Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3755–3773, Dublin, Ireland. Association for Computational Linguistics.
  148. 148.Douglas Zytko, Pamela J. Wisniewski, Shion Guha, Eric P. S. Baumer, and Min Kyung Lee. 2022. Participatory Design of AI Systems: Opportunities and Challenges Across Diverse Users, Relationships, and Application Domains. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, CHI EA ’22, pages 1–4, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Kirk, H., et al. “The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2409–30, https://doi.org/10.18653/v1/2023.emnlp-main.148.
APA
Kirk, H., Bean, A. M., Vidgen, B., Röttger, P., & Hale, S. A. (2023). The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2409–2430. https://doi.org/10.18653/v1/2023.emnlp-main.148
Chicago
Kirk, H., A. M. Bean, B. Vidgen, P. Röttger, and S. A. Hale. 2023. “The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2409–30. https://doi.org/10.18653/v1/2023.emnlp-main.148.
Harvard
Kirk, H. et al. (2023) “The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2409–2430. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.148.
Vancouver
1. Kirk H, Bean AM, Vidgen B, Röttger P, Hale SA (2023) The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2409–2430

BibTeX

@inproceedings{kirk-etal-2023-past,
    title = "The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values",
    author = {Kirk, Hannah Rose  and
      Bean, Andrew M.  and
      Vidgen, Bertie  and
      R{\"o}ttger, Paul  and
      Hale, Scott A.},
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.148/",
    doi = "10.18653/v1/2023.emnlp-main.148",
    pages = "2409--2430"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/