The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Eve FleisigSu Lin BlodgettDan KleinZeerak Talat

article2024NAACL61 citations

Critiques standard machine learning practices that aggregate human labels into a single ground truth, providing practical recommendations to treat annotator disagreement as valuable signal rather than noise.

Listen

Data labeling for artificial intelligence systems has long operated under the assumption that multiple human annotations should be collapsed into a single, aggregated "ground truth" label per example. In practice, annotator disagreement has typically been dismissed as statistical noise, poor data quality, or annotator ineptitude. As artificial intelligence systems are deployed in socially consequential domains like content moderation and conversational assistants, this traditional aggregation approach introduces significant risks. Eliminating divergent human judgments often erases minoritized perspectives, distorts true population beliefs, and builds miscalibrated models that fail in deployment.

The article synthesizes the emerging "perspectivist" paradigm shift in machine learning, which treats annotator disagreement not as error to eliminate, but as valuable signal. Its primary objective is to evaluate the foundational assumptions of both longstanding and perspectivist data annotation approaches, delineate the practical and normative challenges that persist, and provide an actionable framework for capturing human label variation across the artificial intelligence development lifecycle.

The authors conducted a conceptual synthesis and meta-analysis of data labeling practices across natural language processing and broader machine learning research. By analyzing crowdsourcing dynamics, demographic studies, and statistical aggregation techniques, the article examines how current methods fail to represent target populations and explores the implications of shifting from single-label truth models to perspectivist frameworks that capture label distributions.

The article presents several critical findings. First, annotator disagreement is not confined to subjective tasks; it regularly occurs in foundational, seemingly objective tasks like natural language inference, semantic textual similarity, and image classification due to task complexity, differing dialects, and varied interpretations. Second, standard majority-vote aggregation systematically discards valid minority perspectives, moving estimated dataset means further away from true stakeholder population values and generating miscalibrated models that disproportionately reflect dominant demographic groups. Third, while demographic traits influence annotations, non-demographic factors—such as task-specific context, personal lived experience, and an individual's digital habits—frequently drive disagreement more powerfully than isolated traits like gender or age. Over half of crowdworkers explicitly report needing task context, such as system purpose and real-world consequences, which directly shifts their labeling decisions. Fourth, despite the growing collection of rich perspectivist data, model evaluation still bottlenecks around single aggregated "gold" labels because standard engineering pipelines lack non-aggregated evaluation benchmarks.

These findings have direct implications for operational risk, model reliability, and fairness. Treating disagreement as noise introduces silent compliance and safety risks: models trained on naive majority votes embed majority biases while failing to protect vulnerable communities. In addition, organizations that rely on unrepresentative crowdworker platforms risk deploying products that completely misjudge end-user norms. Conversely, moving toward perspectivist data practices creates operational tensions, including higher data collection costs, potential participant privacy trade-offs, and friction with institutional pressures that prioritize development speed over data quality.

To resolve these challenges, the article outlines actionable recommendations across the data pipeline. Development teams must define the acceptable bounds of disagreement before data collection begins, choosing deliberately between prescriptive standards and descriptive variation. During recruitment, teams should stratify annotator pools along task-relevant axes, cap the volume of annotations per individual to avoid dominant-rater skew, and verify quality using intra-annotator consistency checks rather than simple majority consensus. Task designers should provide explicit task context and uncertainty options to annotators, while dataset curators must document selection procedures and retain raw, non-aggregated labels. Finally, model developers should adopt distribution-aware loss functions and evaluation metrics—such as statistical divergence and calibration against human opinion distributions—rather than relying solely on majority-vote accuracy.

The article notes several limitations in the emerging perspectivist paradigm. As a position paper, its findings are based on a synthesis of literature rather than a comprehensive empirical meta-analysis across all artificial intelligence domains. Crucially, the authors emphasize that simple personalization or algorithmic diversification cannot bypass normative decisions: organizations must still make deliberate, transparent choices about which human perspectives to prioritize when deploying systems that produce singular real-world outcomes.

arXiv: 2404.13038
  • Paper: MaxMin-RLHF: Alignment with Diverse Human Preferences, Souradip Chakraborty et al. (2024). MaxMin-RLHF carries the source’s critique of majority-dominated preferences into language-model alignment, optimizing for underserved groups rather than a single averaged reward.
Cover for The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Abstract

Longstanding data labeling practices in machine learning involve collecting and aggregating labels from multiple annotators. But what should we do when annotators disagree? Though annotator disagreement has long been seen as a problem to minimize, new perspectivist approaches challenge this assumption by treating disagreement as a valuable source of information. In this position paper, we examine practices and assumptions surrounding the causes of disagreement—some challenged by perspectivist approaches, and some that remain to be addressed—as well as practical and normative challenges for work operating under these assumptions. We conclude with recommendations for the data labeling pipeline and avenues for future research engaging with subjectivity and disagreement.

Table of Contents

  • 1 Introduction and Related Work
  • 2 Learning Problems: Preference Modeling and Social Choice
  • 2.1 Defining the (Preference Modeling) Voting Rule
  • 3 Evaluating the (Preference Modeling) Voting Rule
  • 3.1 Perspective 1: Generalization
  • 3.2 Perspective 2: Axiomatic Characterizations
  • 3.3 Perspective 3: Distortion
  • 4 Discussion
  • References
  • A Additional Related Work

Knowls

  1. Knowl 1 — Two paradigms for interpreting annotator variation

    definition

    The longstanding data-labeling paradigm collects multiple labels for each item and aggregates them—commonly by majority vote or averaging—with the aim of estimating one underlying ground-truth label. The perspectivist paradigm instead treats differences among annotators’ labels as potentially meaningful information to retain or model. Its approaches include learning from individual labels or annotator attributes, modeling annotator behavior, training on label distributions, calibrating to annotator variation, collecting more annotations, and investigating why judgments differ.

  2. Knowl 2 — Disagreement can reflect perspective and task context, not merely poor labeling

    model/method

    The paper challenges the assumption that annotator disagreement is generally noise caused by bias, ineptitude, or a subjective task. In particular, statistical bias—deviation from an aggregate estimate—is not equivalent to societal prejudice: a minority judgment may reflect relevant knowledge or lived experience, and the aggregate may itself encode a dominant group’s perspective. Treating lived experience as a legitimate form of expertise helps explain why annotators can disagree while being informed and attentive. Judgments also depend on context, including task instructions, the intended use of labels, and the consequences annotators associate with decisions; withholding this context can therefore conceal a source of systematic variation. Disagreement also occurs in tasks often treated as objective, so it cannot be assumed to be confined to explicitly opinion-based tasks.

  3. Knowl 3 — Demographics explain only part of disagreement

    empirical result

    The paper’s synthesis finds that demographic and cultural characteristics—including race, gender, age, education, political affiliation, and language proficiency—can be associated with annotator disagreement, but their predictive value varies across tasks and individual demographic factors often do not predict disagreement reliably. Other potentially important causes include task-specific attitudes and experiences, such as social-media use or views about online toxicity, as well as linguistic and cultural differences not captured by demographic categories. Demographic analysis remains valuable for ensuring that different groups’ views are heard, even when demographics do not explain a particular item’s labels; broadening inquiry beyond demographics may improve explanations of disagreement and task-specific annotator recruitment.

  4. Knowl 4 — Aggregated labels need not represent the stakeholder population

    empirical result

    The paper argues that a mean or majority label from a recruited annotator pool is not necessarily a good estimate of the views of a broader stakeholder population. Crowdsourcing pools may differ demographically from the populations affected by a system, while small annotator samples and few labels per item increase sampling error and make it less likely that relevant perspectives are represented for each item. Allowing a small number of prolific annotators to label many items can further concentrate influence. Aggregation can then discard minority ratings systematically rather than randomly, underrepresent groups with less presence in the pool, and produce labels that align more strongly with majority groups. Models trained on such labels may consequently be poorly calibrated to variation in population opinions.

  5. Knowl 5 — Ground truth and equal-weight aggregation are normative choices

    theoretical result

    The paper argues that some annotation tasks have no single ground-truth answer: ambiguity may persist because a task is underspecified, or because reasonable, informed people can hold different views even when they understand the task. In such cases, treating annotators as noisy approximators of one correct label mischaracterizes the task. Averaging labels also implicitly gives annotators equal weight, although people may differ in relevant expertise, lived experience, cultural grounding, and exposure to the consequences of a model’s decisions. For tasks affecting a particular community, equal-weight aggregation can therefore fail to represent those most affected and obscure whose values determine the resulting label.

  6. Knowl 6 — Preserving disagreement requires quality checks beyond consensus

    limitation

    Perspectivist labeling faces a practical tension: retaining meaningful disagreement while detecting spam, inattentive responses, and other low-quality data. Inter-annotator agreement alone cannot resolve this tension, because disagreement is not necessarily evidence of poor work. The paper identifies possible complementary checks, including clear-cut control items on which disagreement is not reasonably expected, completion time, consistency among similar labels, annotator briefing or training, and within-annotator consistency. Developing and using such checks is important for perspectivist methods to preserve dissent without treating all responses as equally reliable.

  7. Knowl 7 — Design the labeling process around its purpose and likely sources of disagreement

    model/method

    The paper recommends making labeling decisions explicit before annotation begins: identify whose views the dataset is intended to represent, set any prescriptive bounds on acceptable disagreement, and anticipate task-relevant sources of variation or confusion. Recruitment should match that purpose. For population representation, recruit a representative sample and consider stratification or additional recruitment where participation is uneven; for tasks requiring specific expertise, recruit or allocate annotators accordingly. A larger pool can reduce sampling error, and capping the number of items assigned to each annotator can limit overreliance on prolific labelers. To reduce spam without filtering out minority views, use checks such as intra-annotator consistency or multiple qualification rounds rather than agreement with the majority. During labeling, explain the task’s intended use and potential consequences, provide options for uncertainty, and let annotators give open-ended feedback; disagreement can then prompt clarification or expansion of instructions and label options. Documentation should record selection procedures, annotator counts and participation restrictions, item counts per annotator, filtering decisions, normative bounds and their rationale, and individual labels where possible.

  8. Knowl 8 — Train and evaluate models against variation in human labels

    model/method

    The paper recommends objectives that represent annotator variation instead of optimizing only for agreement with an aggregate label. Examples include comparing model predictions with the distribution of human labels using KL divergence or calibrating predictions to the distribution of annotator opinions. Evaluation alternatives to averaged “gold” labels include distributional comparisons such as KL divergence, cosine similarity, or correlation; accuracy at modeling individual annotators; and calibration to population uncertainty. Disagreement among evaluators of model outputs can also help identify weaknesses or differences in quality of service across subgroups. These approaches address a mismatch in which richer perspectivist data are collected but models often produce one output and are still judged against a single aggregated label.

  9. Knowl 9 — Privacy, participation, and institutional incentives constrain perspectivist data collection

    limitation

    Collecting detailed judgments from minoritized or otherwise affected groups can impose burdens without providing a corresponding benefit to those groups. More detailed annotation can also increase privacy risks. The paper identifies learning from fewer data points, privacy-preserving methods, and community-led data ownership—including Indigenous data sovereignty—as possible directions for addressing these concerns. Meaningful participation faces a separate institutional constraint: research incentives often favor fast data collection over reciprocal engagement, contextual understanding, and shared influence over problem formulation or evaluation. Lowering barriers for participants and encouraging institutional support for slower, context-specific work are therefore part of the practical challenge, not merely matters of annotation-interface design.

  10. Knowl 10 — Make normative choices explicit and draw on approaches beyond majority vote

    model/method

    The paper recommends stating who should influence a labeling decision and how their views should count, rather than treating majority vote or researcher neutrality as value-free defaults. Annotation can range from a broadly democratic approach, in which stakeholder views are solicited on equal terms, to an expert-driven approach, in which particular expertise is required; the appropriate choice depends on the task. Researchers should also specify whether annotators are meant to describe their own judgments or follow prescriptive guidelines, and where task-specific variation is acceptable. These choices matter especially for social-norm tasks, where unclear boundaries can make aggregation an opaque way of setting rules. The paper points to social choice, science and technology studies, pragmatics, philosophy of mind, and participatory design as useful sources for examining representation, interpretation, and power. Personalization does not remove these questions; it changes them into questions about when personalization is appropriate. Where a system must return a single output, considering multiple dimensions of preference may allow one output to satisfy people who value different, compatible qualities.

Coverage note — No substantial contributed material is omitted. Supporting examples and individual citations are condensed; the authors’ caveat that the position paper is not a comprehensive literature review is not expanded into a separate knowl because it qualifies the paper’s coverage rather than adding a substantive method or finding.

References

  1. 1.Gavin Abercrombie, Verena Rieser, and Dirk Hovy. 2023. Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement. ArXiv:2301.10684 [cs].
  2. 2.Philip E Agre. 2014. Toward a critical technical practice: Lessons learned in trying to reform AI. In Social science, technical systems, and cooperative work, pages 131–157. Psychology Press.
  3. 3.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2021. Whose opinions matter? perspective-aware models to identify opinions of hate speech victims in abusive language detection.
  4. 4.Hala Al Kuwatly, Maximilian Wich, and Georg Groh. 2020. Identifying and Measuring Annotator Bias Based on Annotators’ Demographic Characteristics. In Proceedings of the Fourth Workshop on Online Abuse and Harms, pages 184–190, Online. Association for Computational Linguistics.
  5. 5.Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang. 2023. DICES dataset: Diversity in conversational ai evaluation for safety.
  6. 6.Lora Aroyo and Chris Welty. 2014. The Three Sides of CrowdTruth. Human Computation, 1(1).
  7. 7.Kenneth J Arrow. 1977. Social Choice and Individual Values, 2 edition. Cowles Foundation Monographs. Yale University Press, New Haven, CT.
  8. 8.Ron Artstein. 2017. Inter-annotator agreement. Handbook of linguistic annotation, pages 297–313.
  9. 9.Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022. Stop measuring calibration when humans disagree. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1892–1915, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  10. 10.Aparna Balagopalan, David Madras, David H. Yang, Dylan Hadfield-Menell, Gillian K. Hadfield, and Marzyeh Ghassemi. 2023. Judging facts, judging norms: Training machine learning models to judge humans requires a modified approach to labeling data. Science Advances, 9(19):eabq0701. Publisher: American Association for the Advancement of Science.
  11. 11.Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021. We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, pages 15–21, Online. Association for Computational Linguistics.
  12. 12.Laura Biester, Vanita Sharma, Ashkan Kazemi, Naihao Deng, Steven Wilson, and Rada Mihalcea. 2022. Analyzing the effects of annotator gender across NLP tasks. In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @LREC2022, pages 10–19, Marseille, France. European Language Resources Association.
  13. 13.Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022. The values encoded in machine learning research. In 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22. ACM.
  14. 14.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  15. 15.Alan L. Boegehold. 1963. Toward a study of athenian voting procedure. Hesperia: The Journal of the American School of Classical Studies at Athens, 32(4):366–374.
  16. 16.Geoffrey C Bowker and Susan Leigh Star. 2000. Sorting Things Out: Classification and Its Consequences. MIT Press, London, England.
  17. 17.Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a perspectivist turn in ground truthing for predictive computing. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’23/IAAI’23/EAAI’23. AAAI Press.
  18. 18.David J Chalmers. 1997. The conscious mind. Philosophy of Mind. Oxford University Press, New York, NY.
  19. 19.Herbert H. Clark and Thomas B. Carlson. 1982. Hearers and speech acts. Language, 58(2):332.
  20. 20.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  21. 21.Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The CommitmentBank: Investigating projection in naturally occurring discourse.
  22. 22.Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The participatory turn in ai design: Theoretical foundations and the current state of practice. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’23, New York, NY, USA. Association for Computing Machinery.
  23. 23.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040–4054, Online. Association for Computational Linguistics.
  24. 24.Naihao Deng, Xinliang Zhang, Siyang Liu, Winston Wu, Lu Wang, and Rada Mihalcea. 2023. You are what you annotate: Towards better models through annotator representations. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12475–12498, Singapore. Association for Computational Linguistics.
  25. 25.Mark Diaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. 2018. Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI ’18, page 1–14, New York, NY, USA. Association for Computing Machinery.
  26. 26.Mary Douglas. 1978. Purity and danger: an analysis of the concepts of pollution and taboo, repr edition. Routledge, London. OCLC: 248038797.
  27. 27.Anca Dumitrache, Lora Aroyo, and Chris Welty. 2019. A crowdsourced frame disambiguation corpus with ambiguity. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2164–2170, Minneapolis, Minnesota. Association for Computational Linguistics.
  28. 28.Elizabeth Edenberg and Alexandra Wood. 2023. Disambiguating Algorithmic Bias: From Neutrality to Justice. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’23, pages 691–704, New York, NY, USA. Association for Computing Machinery.
  29. 29.Michael D Ekstrand, Anubrata Das, Robin Burke, and Fernando Diaz. 2012. Fairness in recommender systems. In Recommender Systems Handbook, pages 679–707. Springer.
  30. 30.Allan Feldman and Roberto Serrano. 2006. Welfare economics and social choice theory, 2 edition. Springer, New York, NY.
  31. 31.Eve Fleisig, Rediet Abebe, and Dan Klein. 2023. When the majority is wrong: Modeling annotator disagreement for subjective tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6715–6726, Singapore. Association for Computational Linguistics.
  32. 32.Lucie Flek. 2020. Returning the N to NLP: Towards Contextually Personalized Classification Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7828–7838, Online. Association for Computational Linguistics.
  33. 33.Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021. Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2591–2597, Online. Association for Computational Linguistics.
  34. 34.Paula Fortuna, Monica Dominguez, Leo Wanner, and Zeerak Talat. 2022. Directions for NLP practices applied to online hate speech detection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11794–11805, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  35. 35.Batya Friedman. 1996. Value-sensitive design. interactions, 3(6):16–23.
  36. 36.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM, 64(12):86–92.
  37. 37.Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019. Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1161–1166, Hong Kong, China. Association for Computational Linguistics.
  38. 38.Erving Goffman. 1976. Replies and responses. Language in Society, 5(3):257–313.
  39. 39.Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury Learning: Integrating Dissenting Voices into Machine Learning Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, pages 1–19, New York, NY, USA. Association for Computing Machinery.
  40. 40.Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proc. ACM Hum.-Comput. Interact., 6(CSCW2).
  41. 41.Stuart Hall et al. 1997. The spectacle of the other. Representation: Cultural representations and signifying practices, 7.
  42. 42.Anna Lauren Hoffmann. 2020. Terms of inclusion: Data, discourse, violence. New Media & Society, 23(12):3539–3556.
  43. 43.Pei-Yun Hsueh, Prem Melville, and Vikas Sindhwani. 2009. Data quality from crowdsourcing: a study of annotation selection criteria. In Proceedings of the NAACL HLT 2009 workshop on active learning for natural language processing, pages 27–35.
  44. 44.Olivia Huang, Eve Fleisig, and Dan Klein. 2023. Incorporating worker perspectives into mturk annotation practices for NLP.
  45. 45.Christoph Hube, Besnik Fetahu, and Ujwal Gadiraju. 2019. Understanding and mitigating worker biases in the crowdsourced collection of subjective judgments. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–12.
  46. 46.Frank Jackson. 1982. Epiphenomenal qualia. The Philosophical Quarterly (1950-), 32(127):127–136.
  47. 47.Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022. Investigating reasons for disagreement in natural language inference. Transactions of the Association for Computational Linguistics, 10:1357–1374.
  48. 48.Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the snark: Annotator diversity in data practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA. Association for Computing Machinery.
  49. 49.Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale. 2023. The empty signifier problem: Towards clearer paradigms for operationalising "alignment" in large language models.
  50. 50.Tahu Kukutai and John Taylor, editors. 2016. Indigenous Data Sovereignty: Toward an agenda, volume 38. ANU Press.
  51. 51.Savannah Larimore, Ian Kennedy, Breon Haskett, and Alina Arseniev-Koehler. 2021. Reconsidering annotator disagreement about racist language: Noise or signal? In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 81–90, Online. Association for Computational Linguistics.
  52. 52.Josh Lepawsky. 2019. No insides on the outsides. Discard Studies.
  53. 53.Clarence Irving Lewis. 1930. Mind and the world-order. International Journal of Ethics, 40(4):550–556.
  54. 54.Yanying Li, Haipei Sun, and Wendy Hui Wang. 2020. Towards fair truth discovery from biased crowdsourced answers. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, page 599–607, New York, NY, USA. Association for Computing Machinery.
  55. 55.Yunqi Li, Hanxiong Chen, Shuyuan Xu, Yingqiang Ge, Juntao Tan, Shuchang Liu, and Yongfeng Zhang. 2023. Fairness in recommendation: Foundations, methods, and applications. ACM Transactions on Intelligent Systems and Technology, 14(5):1–48.
  56. 56.Angelina McMillan-Major, Emily M. Bender, and Batya Friedman. 2024. Data statements: From technical concept to community practice. ACM J. Responsib. Comput., 1(1).
  57. 57.Milagros Miceli and Julian Posada. 2022. The data-production dispositif. Proc. ACM Hum.-Comput. Interact., 6(CSCW2).
  58. 58.Michael J Muller and Sarah Kuhn. 1993. Participatory design. Communications of the ACM, 36(6):24–28.
  59. 59.Hasti Narimanzadeh, Arash Badie-Modiri, Iuliia G. Smirnova, and Ted Hsuan Yun Chen. 2023. Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method. Proc. ACM Hum.-Comput. Interact., 7(CSCW2).
  60. 60.Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131–9143, Online. Association for Computational Linguistics.
  61. 61.Stefanie Nowak and Stefan Rüger. 2010. How reliable are annotations via crowdsourcing: a study about inter-annotator agreement for multi-label image annotation. In Proceedings of the international conference on Multimedia information retrieval, pages 557–566.
  62. 62.Matthias Orlikowski, Paul Röttger, Philipp Cimiano, and Dirk Hovy. 2023. The ecological fallacy in annotation: Modeling human label variation goes beyond sociodemographics. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1017–1029, Toronto, Canada. Association for Computational Linguistics.
  63. 63.Alicia Parrish, Sarah Laszlo, and Lora Aroyo. 2023. Is a picture of a bird a bird: Policy recommendations for dealing with ambiguity in machine vision models. ArXiv:2306.15777 [cs].
  64. 64.Desmond Upton Patton, Philipp Blandfort, William R. Frey, Michael B. Gaskell, and Svebor Karaman. 2019. Annotating social media data from vulnerable populations: Evaluating disagreement between domain experts and graduate student annotators. In Hawaii International Conference on System Sciences.
  65. 65.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  66. 66.Jiaxin Pei and David Jurgens. 2023. When Do Annotator Demographics Matter? Measuring the Influence of Annotator Demographics with the POPQUORN Dataset.
  67. 67.Pew Research Center. 2016. Research in the crowdsourcing age, a case study.
  68. 68.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  69. 69.Joan Plepi, Béla Neuendorf, Lucie Flek, and Charles Welch. 2022. Unifying data perspectivism and personalization: An application to social norms. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7391–7402, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  70. 70.Maja Popovic. 2021. ´ Agree to disagree: Analysis of inter-annotator disagreements in human evaluation of machine translation output. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 234–243, Online. Association for Computational Linguistics.
  71. 71.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  72. 72.Paul Resnick, Yuqing Kong, Grant Schoenebeck, and Tim Weninger. 2021. Survey equivalence: A procedure for measuring classifier accuracy against human labels. CoRR, abs/2106.01254.
  73. 73.Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190, Seattle, United States. Association for Computational Linguistics.
  74. 74.Pratik Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris Kennedy. 2022. The measuring hate speech corpus: Leveraging rasch measurement theory for data perspectivism. In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @LREC2022, pages 83–94, Marseille, France. European Language Resources Association.
  75. 75.Sebastin Santy, Jenny Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. NLPositionality: Characterizing design biases of datasets and models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9080–9102, Toronto, Canada. Association for Computational Linguistics.
  76. 76.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  77. 77.Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021. Do datasets have politics? disciplinary values in computer vision dataset development. Proc. ACM Hum.-Comput. Interact., 5(CSCW2).
  78. 78.Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, page 59–68, New York, NY, USA. Association for Computing Machinery.
  79. 79.Amartya Sen. 2018. Collective choice and social welfare. Harvard University Press.
  80. 80.Brooklyn Sheppard, Anna Richter, Allison Cohen, Elizabeth Allyn Smith, Tamara Kneese, Carolyne Pelletier, Ioana Baldini, and Yue Dong. 2023. Subtle misogyny detection and mitigation: An expert-annotated dataset. In Socially Responsible Language Modelling Research (SoLaR) Workshop.
  81. 81.Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008. Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing, pages 254–263, Honolulu, Hawaii. Association for Computational Linguistics.
  82. 82.Nikki Stevens and Os Keyes. 2021. Seeing infrastructure: race, facial recognition and the politics of data. Cultural Studies, 35(4-5):833–853.
  83. 83.Zeerak Talat, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein. 2021. Disembodied Machine Learning: On the Illusion of Objectivity in NLP.
  84. 84.Terne Sasha Thorn Jakobsen, Laura Cabello, and Anders Søgaard. 2023. Being right for whose right reasons? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1033–1054, Toronto, Canada. Association for Computational Linguistics.
  85. 85.Nanna Thylstrup and Zeerak Talat. 2020. Detecting ‘Dirt’ and ‘Toxicity’: Rethinking Content Moderation as Pollution Behaviour. SSRN Electronic Journal.
  86. 86.Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2020. A Case for Soft Loss Functions. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 8:173–177.
  87. 87.Ruyuan Wan, Jaehyung Kim, and Dongyeop Kang. 2023. Everyone’s voice matters: Quantifying annotation disagreement using demographic information. Proceedings of the AAAI Conference on Artificial Intelligence, 37(12):14523–14530.
  88. 88.Yifan Wang, Weizhi Ma, M. Zhang, Yiqun Liu, and Shaoping Ma. 2022. A survey on the fairness of recommender systems. ACM Transactions on Information Systems, 41:1 – 43.
  89. 89.Yuxia Wang, Shimin Tao, Ning Xie, Hao Yang, Timothy Baldwin, and Karin Verspoor. 2023. Collective Human Opinions in Semantic Textual Similarity. Transactions of the Association for Computational Linguistics, 11:997–1013.
  90. 90.Langdon Winner. 1980. Do artifacts have politics? Daedalus, 109(1):121–136.
  91. 91.Runhua Xu, Nathalie Baracaldo, and James B. D. Joshi. 2021. Privacy-preserving machine learning: Methods, challenges and directions. ArXiv, abs/2108.04417.
  92. 92.Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahmood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, and João Sedoc. 2023. A needle in a haystack: An analysis of high-agreement workers on MTurk for summarization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14944–14982, Toronto, Canada. Association for Computational Linguistics.
  93. 93.Xiang Zhou, Yixin Nie, and Mohit Bansal. 2022. Distributed NLI: Learning to predict human opinion distributions for language reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pages 972–987, Dublin, Ireland. Association for Computational Linguistics.

Citation

MLA
Fleisig, E., et al. “The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 2279–92, https://doi.org/10.18653/v1/2024.naacl-long.126.
APA
Fleisig, E., Blodgett, S. L., Klein, D., & Talat, Z. (2024). The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2279–2292. https://doi.org/10.18653/v1/2024.naacl-long.126
Chicago
Fleisig, E., S. L. Blodgett, D. Klein, and Z. Talat. 2024. “The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2279–92. https://doi.org/10.18653/v1/2024.naacl-long.126.
Harvard
Fleisig, E. et al. (2024) “The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2279–2292. Available at: https://doi.org/10.18653/v1/2024.naacl-long.126.
Vancouver
1. Fleisig E, Blodgett SL, Klein D, Talat Z (2024) The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 2279–2292

BibTeX

@inproceedings{fleisig-etal-2024-perspectivist,
    title = "The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels",
    author = "Fleisig, Eve  and
      Blodgett, Su Lin  and
      Klein, Dan  and
      Talat, Zeerak",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.126/",
    doi = "10.18653/v1/2024.naacl-long.126",
    pages = "2279--2292"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/