The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Barbara Plank

article2022EMNLP215 citations

Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.

Listen

Modern artificial intelligence and natural language processing systems typically rely on supervised learning, where models train on datasets labeled by human annotators. Standard industry and research practice assumes there is a single objective ground truth for every task, collapsing multiple human viewpoints into one aggregate label—often through majority voting—while treating disagreements as mere noise or errors. However, because language and visual interpretation are inherently complex, subjective, and context-dependent, forcing data into single gold labels discards valuable information and creates artificial performance metrics that fail in real-world deployment.

The article aims to evaluate how human label variation affects every stage of the machine learning pipeline—specifically data collection, model training, and performance evaluation. It demonstrates that genuine human disagreement is not noise to be eliminated, but a rich signal that developers must embrace to build more reliable, trustworthy, and inclusive systems.

To establish this framework, the article synthesizes findings across natural language processing, computer vision, and human-computer interaction literature. It reviews existing approaches for handling label discrepancies, categorizes core modeling paradigms, and compiles a comprehensive repository of publicly available datasets containing un-aggregated annotator labels.

The analysis reveals several critical findings. First, irreconcilable label variation is widespread rather than rare, appearing in at least 20 percent of instances in key language inference tasks due to genuine linguistic ambiguity, varying perspectives, and multiple plausible answers. Second, traditional methods of resolving variation through majority aggregation or data filtering discard valuable evidence, and removing low-agreement instances can actively harm model performance. Third, emerging modeling techniques—such as multi-task learning, soft labels, and loss weighting—show that training directly on the distribution of human opinions improves model generalization, robustness, and out-of-distribution performance. Fourth, current evaluation methods suffer from a severe disconnect: even though models are beginning to learn from diverse viewpoints, researchers predominantly evaluate them against single aggregate labels, masking system overconfidence and real-world failure modes.

These findings imply significant risks for organizations deploying artificial intelligence systems. Relying on majority-vote ground truths distorts performance benchmarks, introduces compliance and safety hazards, and risks marginalizing underrepresented viewpoints in sensitive applications like content moderation. Conversely, capturing human label distributions offers cost and efficiency advantages, as models trained on richer, nuanced distributions may require less overall labeled data to achieve strong generalization.

The article recommends that organizations and researchers immediately transition from aggregating labels to collecting and releasing un-aggregated, annotator-level data along with comprehensive metadata, such as annotator backgrounds and uncertainty indicators. System developers should adopt soft evaluation metrics, including cross-entropy and entropy correlation, alongside standard accuracy to assess calibration and trustworthiness. Furthermore, interdisciplinary standards must be developed to systematically distinguish between careless annotation errors and legitimate, informative human disagreement.

While the article provides strong qualitative synthesis and establishes a clear conceptual framework, it notes that current empirical evidence remains fragmented across subfields and that universal standards for cross-task evaluation are not yet fully established. Leaders should proceed with confidence in the conceptual shift toward preserving label variation, while exercising caution regarding implementation until broader empirical benchmarks are standardized across their specific operational domains.

arXiv: 2211.02570
Cover for The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Abstract

Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics. However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers. In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward. This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation. However, few works consider all of these dimensions jointly; and existing research is fragmented. We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward. As datasets are becoming increasingly available, we hope that this synthesized view on the “problem” will lead to an open discussion on possible strategies to devise fundamentally new directions.

Table of Contents

  • 1 Introduction
  • 2 Data and Human Label Variation
  • 3 Modeling and Human Label Variation
  • 4 Evaluation and Human Label Variation
  • 5 Conclusions
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Datasets with Multiple Annotations

Knowls

  1. Knowl 1 — Human label variation is plausible variation, not simply annotation error

    definition

    Human label variation (HLV) is plausible variation among human annotations of the same instance. It can arise because an instance is ambiguous, an annotator is uncertain, people genuinely disagree, or more than one answer is correct. The concept assumes annotators are generally giving their best judgments; it distinguishes this meaningful variation from differences caused by annotation mistakes such as lapses in attention. The term “variation” is preferred to “disagreement” because the latter can imply that the different judgments cannot all be valid.

  2. Knowl 2 — A single gold label shapes data, modeling, and evaluation

    model/method

    The conventional machine-learning pipeline often turns multiple human judgments into one gold label, commonly by taking the majority vote. This practice encodes a single answer in the dataset, encourages models to predict one preferred output, and evaluates predictions against that same answer. It is reasonable when annotators strongly agree, but can hide ambiguity, subjectivity, or multiple plausible answers in tasks where those conditions occur. The paper argues that HLV should instead be considered across all three pipeline stages—data, modeling, and evaluation—because treating the aggregated label as unquestioned ground truth can discard information at each stage.

  3. Knowl 3 — Four modeling strategies either resolve or retain label variation

    model/method

    The paper organizes approaches to learning from human annotations into two broad camps. Approaches that resolve variation either aggregate labels into one answer, using methods such as majority voting or probabilistic aggregation, or filter out instances with low annotator agreement. Aggregation necessarily loses the possibility of preserving several plausible labels; filtering discards data and can reduce performance. Approaches that embrace variation either learn directly from unaggregated judgments—for example, through repeated labeling, crowd layers, or soft-label learning—or enrich a conventional gold label with information from the judgments. Enrichment methods include cost-sensitive loss weighting, multitask learning that represents annotators or the per-instance label distribution, and sequential fine-tuning.

  4. Knowl 4 — Release unaggregated judgments and the metadata needed to interpret them

    model/method

    To make HLV available for research, dataset creators should release annotator-level, unaggregated labels, even if only for a subset of instances. They should document how annotations were collected and provide available metadata, ideally at the instance level: for example, source and time of the item, annotator identifiers, annotation completion time, annotator background, and rationales. Such information can help researchers interpret why labels vary and distinguish sources of variation. Collection and release should be done responsibly: annotator backgrounds can affect judgments, and modeling those judgments may include underrepresented perspectives but can also amplify some voices in ways that harm others.

  5. Knowl 5 — Evaluation should compare models with the distribution of human judgments

    model/method

    Accuracy against a single gold label evaluates only whether a model's top prediction matches an aggregated answer; it does not show whether the model reflects the range of human judgments or how well its confidence corresponds to them. The paper describes soft-label evaluation, which compares a model's output distribution with the human label distribution. Proposed measures include cross-entropy, which scores the model across the human-assigned labels; Pearson correlation between instance-level human and model entropy; and KL-divergence or Jensen–Shannon-divergence comparisons. Other evaluation designs compare predictions with individual annotators or report performance across groups of instances defined by annotator agreement, annotator clusters, item difficulty, or annotator uncertainty. The paper calls for evaluation that considers task properties and the reasons for variation, rather than relying on one metric and one aggregated answer.

  6. Knowl 6 — The dataset repository spans varied NLP and vision tasks

    data/table

    The paper's dataset catalogue collects publicly available resources with multiple judgments per instance and points to the repository at https://github.com/mainlp/awesome-human-label-variation. Its NLP examples span tasks such as word-sense disambiguation, part-of-speech tagging, named-entity recognition, discourse, natural-language inference, emotion classification, and abusive-language detection. Examples include the reannotated RTE data and ChaosNLI for inference, GoEmotions for emotion categories, and resources for offensive-language annotation. The catalogue also includes computer-vision datasets such as reannotated LabelMe and CIFAR-10H image-classification data, as well as medical-image judgments. The annual-count chart on page 4 counts NLP resource papers releasing publicly available datasets with multiple labels per instance and depicts more such releases in later years; this is a descriptive count of resources, not a measure of all HLV research. The authors describe the repository as, to their best knowledge, the most comprehensive available at the time, while inviting additions.

  7. Knowl 7 — Research on learning from variation remains fragmented

    limitation

    The paper identifies limited overlap among HLV research communities and task areas, including work on subjectivity, natural-language inference, and methods spanning NLP and computer vision. It identifies comprehensive evaluation and transfer of methods across tasks as open questions: task properties may make some approaches more suitable than others, but the paper reports that broad comparative evidence is lacking. It also raises the practical question of how to allocate annotation effort between collecting more instances and collecting more judgments per instance. The authors hypothesize that richer human label distributions could reduce the number of instances required, but present this as a possibility to investigate, not an established result.

  8. Knowl 8 — Separating annotation errors from meaningful variation is an open problem

    limitation

    Learning from HLV requires distinguishing meaningful differences in human judgment from annotation errors, yet the paper identifies this distinction as underexamined. Low agreement alone does not establish that an instance is erroneous: variation may reflect a hard case, uncertainty, or multiple plausible interpretations. Conversely, some differences may result from mistakes. The paper calls for further work on defining these categories and detecting annotation errors, so that systems do not automatically discard informative judgments as noise.

  9. Knowl 9 — Calibration against a majority label can misrepresent human disagreement

    theoretical result

    The paper reports that measuring model calibration against the human majority label is theoretically and empirically problematic when judgments vary within an instance. A model can match the majority while failing to reflect the distribution of human responses. As a first step toward addressing this issue, the authors report proposing instance-level calibration measures intended to better capture that distribution; the paper leaves open how HLV can best be used to make systems more trustworthy.

  10. Knowl 10 — The synthesis and dataset collection are explicitly incomplete

    limitation

    The paper presents itself as a concise synthesis of a broad topic and states that both its account and its dataset repository are necessarily incomplete. It calls for an open, interdisciplinary discussion and community contributions to the resource collection rather than claiming to provide a definitive treatment of HLV.

Coverage note — The ethical implications are incorporated into the data-release knowl; the page-4 resource-count chart is summarized qualitatively because the paper does not enumerate its annual values in the text.

References

  1. 1.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2021. Whose opinions matter? Perspective-aware models to identify opinions of hate speech victims in abusive language detection. arXiv preprint arXiv:2106.15896.
  2. 2.Cecilia Ovesdotter Alm. 2011. Subjective natural language problems: Motivations, applications, characterizations, and implications. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 107–112.
  3. 3.Dina Almanea and Massimo Poesio. 2022. The ARMIS dataset of misogyny in arabic tweets. In Proc. of LREC.
  4. 4.Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24.
  5. 5.Ron Artstein and Massimo Poesio. 2008. Inter-coder agreement for computational linguistics. Computational linguistics, 34(4):555–596.
  6. 6.Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022. Stop measuring calibration when humans disagree. In Empirical Methods of Natural Language Processing: EMNLP 2022. Association for Computational Linguistics.
  7. 7.Valerio Basile, Federico Cabitza, Andrea Campagner, and Michael Fell. 2021a. Toward a perspectivist turn in ground truthing for predictive computing. arXiv preprint arXiv:2109.04270.
  8. 8.Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, Alexandra Uma, et al. 2021b. We need to consider disagreement in evaluation. In 1st Workshop on Benchmarking: Past, Present and Future, pages 15–21. Association for Computational Linguistics.
  9. 9.Elisa Bassignana and Barbara Plank. 2022. CrossRE: A Cross-Domain Dataset for Relation Extraction. In Findings of the Association for Computational Linguistics: EMNLP 2022. Association for Computational Linguistics.
  10. 10.Eyal Beigman and Beata Beigman-Klebanov. 2009. Learning with annotation noise. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, pages 280–287, Suntec, Singapore. Association for Computational Linguistics.
  11. 11.Beata Beigman Klebanov, Eyal Beigman, and Danie Diermaier. 2008. Analyzing disagreements. In Proceedings of the Coling 2008 workshop on Human Judgements in Computational Linguistics, pages 2–7.
  12. 12.Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  13. 13.Yevgeni Berzak, Yan Huang, Andrei Barbu, Anna Korhonen, and Boris Katz. 2016. Anchoring and agreement in syntactic annotations. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2215–2224, Austin, Texas. Association for Computational Linguistics.
  14. 14.Lars Borin. 2022. All that glitters... interannotator agreement in natural language processing. Morfologi, målstrev og maskinar – Trond Trosterud {fyller ∥tytt∥deavd∥turns}60!, 46(1).
  15. 15.Christopher Bryant and Hwee Tou Ng. 2015. How far are we from fully automatic high quality grammatical error correction? In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 697–707, Beijing, China. Association for Computational Linguistics.
  16. 16.Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser. 2021. ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7388–7403, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Veronika Cheplygina and Josien PW Pluim. 2018. Crowd disagreement about medical images is informative. In Intravascular imaging and computer assisted stenting and large-scale annotation of biomedical data and expert label synthesis, pages 105–111. Springer.
  18. 18.Trevor Cohn and Lucia Specia. 2013. Modelling annotator bias with multi-task Gaussian processes: An application to machine translation quality estimation. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32–42, Sofia, Bulgaria. Association for Computational Linguistics.
  19. 19.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The pascal recognising textual entailment challenge. In Machine learning challenges workshop, pages 177–190. Springer.
  20. 20.Cathrine Damgaard, Paulina Toborek, Trine Eriksen, and Barbara Plank. 2021. “I’ll be there for you”: The one with understanding indirect answers. In Proceedings of the 2nd Workshop on Computational Approaches to Discourse, pages 1–11, Punta Cana, Dominican Republic and Online. Association for Computational Linguistics.
  21. 21.Debopam Das, Manfred Stede, and Maite Taboada. 2017. The good, the bad, and the disagreement: Complex ground truth in rhetorical structure analysis. In Proceedings of the 6th Workshop on Recent Advances in RST and Related Formalisms, pages 11–19, Santiago de Compostela, Spain. Association for Computational Linguistics.
  22. 22.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  23. 23.Alexander Philip Dawid and Allan M Skene. 1979. Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20–28.
  24. 24.Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012. Did it happen? the pragmatic complexity of veridicality assessment. Computational Linguistics, 38(2):301–333.
  25. 25.Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser. 2019. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124.
  26. 26.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040–4054, Online. Association for Computational Linguistics.
  27. 27.Emily Denton, Mark Díaz, Ian Kivlichan, Vinodkumar Prabhakaran, and Rachel Rosen. 2021. Whose ground truth? accounting for individual and collective identities underlying dataset annotation. In NeurIPS workshop on Data-Centric AI.
  28. 28.Leon Derczynski, Kalina Bontcheva, and Ian Roberts. 2016. Broad Twitter corpus: A diverse named entity recognition resource. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 1169–1179, Osaka, Japan. The COLING 2016 Organizing Committee.
  29. 29.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  30. 30.Liviu P. Dinu, Ioan-Bogdan Iordache, Ana Sabina Uban, and Marcos Zampieri. 2021. A computational exploration of pejorative language in social media. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3493–3498, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Anca Dumitrache, Lora Aroyo, and Chris Welty. 2018. Crowdsourcing semantic label propagation in relation classification. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER), pages 16–21, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Anca Dumitrache, Lora Aroyo, and Chris Welty. 2019. A crowdsourced frame disambiguation corpus with ambiguity. In Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).
  33. 33.Elisa Ferracane, Greg Durrett, Junyi Jessy Li, and Katrin Erk. 2021. Did they answer? subjective acts and intents in conversational discourse. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1626–1644, Online. Association for Computational Linguistics.
  34. 34.Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021. Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2591–2597, Online. Association for Computational Linguistics.
  35. 35.R Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang. 2020. Garbage in, garbage out? do machine learning application papers in social computing report where human-labeled training data comes from? In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 325–336.
  36. 36.Jacob Goldberger and Ehud Ben-Reuven. 2016. Training deep neural-networks using a noise adaptation layer. ICLR 2017.
  37. 37.Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In CHI Conference on Human Factors in Computing Systems, pages 1–19.
  38. 38.Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein. 2021. The disagreement deconvolution: Bringing machine learning performance metrics in line with reality. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–14.
  39. 39.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML’17, page 1321–1330. JMLR.org.
  40. 40.Bo Han, Jiangchao Yao, Gang Niu, Mingyuan Zhou, Ivor Tsang, Ya Zhang, and Masashi Sugiyama. 2018a. Masking: A new perspective of noisy supervision. In Advances in Neural Information Processing Systems, pages 5836–5846.
  41. 41.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. 2018b. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In Advances in neural information processing systems, pages 8527–8537.
  42. 42.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6997–7013, Dublin, Ireland. Association for Computational Linguistics.
  43. 43.Emily Jamison and Iryna Gurevych. 2015. Noise or additional information? leveraging crowdsource annotation item agreement for natural language tasks. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 291–297.
  44. 44.Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022. Investigating reasons for disagreement in natural language inference. Transactions of the Association for Computational Linguistics, 10:1081–1097.
  45. 45.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  46. 46.Shailza Jolly, Sandro Pezzelle, and Moin Nabi. 2021. EaSe: A diagnostic tool for VQA based on answer diversity. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2407–2414, Online. Association for Computational Linguistics.
  47. 47.Chris J Kennedy, Geoff Bacon, Alexander Sahn, and Claudia von Vacano. 2020. Constructing interval variables via faceted rasch measurement and multi-task deep learning: a hate speech application. arXiv preprint arXiv:2009.10277.
  48. 48.Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2022. Annotation error detection: Analyzing the past and present for a more coherent future. Computational Linguistics.
  49. 49.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in- and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1326–1340, Online. Association for Computational Linguistics.
  50. 50.Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology. Sage publications.
  51. 51.Ranjay A Krishna, Kenji Hata, Stephanie Chen, Joshua Kravitz, David A Shamma, Li Fei-Fei, and Michael S Bernstein. 2016. Embracing error to enable rapid crowdsourcing. In Proceedings of the 2016 CHI conference on human factors in computing systems, pages 3167–3179.
  52. 52.John P Lalor, Hao Wu, and Hong Yu. 2017. Soft label memorization-generalization for natural language inference. arXiv preprint arXiv:1702.08563.
  53. 53.Savannah Larimore, Ian Kennedy, Breon Haskett, and Alina Arseniev-Koehler. 2021. Reconsidering annotator disagreement about racist language: Noise or signal? In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 81–90.
  54. 54.Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021. Agreeing to disagree: Annotating offensive language datasets with annotators’ disagreement. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  55. 55.Tong Liu, Christopher Homan, Cecilia Ovesdotter Alm, Megan Lytle, Ann Marie White, and Henry Kautz. 2016. Understanding discourse on work and job-related well-being in public social media. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1044–1053, Berlin, Germany. Association for Computational Linguistics.
  56. 56.Christopher D Manning. 2011. Part-of-speech tagging from 97% to 100%: is it time for some linguistics? In International conference on intelligent text processing and computational linguistics, pages 171–189. Springer.
  57. 57.Marian Marchal, Merel Scholman, Frances Yung, and Vera Demberg. 2022. Establishing annotation quality in multi-label annotations. In Proceedings of the 29th International Conference on Computational Linguistics, pages 3659–3668, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  58. 58.Héctor Martínez Alonso, Anders Johannsen, and Barbara Plank. 2016. Supersense tagging with inter-annotator disagreement. In Proceedings of the 10th Linguistic Annotation Workshop held in conjunction with ACL 2016 (LAW-X 2016), pages 43–48, Berlin, Germany. Association for Computational Linguistics.
  59. 59.Tyler McDonnell, Matthew Lease, Mucahid Kutlu, and Tamer Elsayed. 2016. Why is that relevant? collecting annotator rationales for relevance judgments. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 4, pages 139–148.
  60. 60.Clara Meister, Elizabeth Salesky, and Ryan Cotterell. 2020. Generalized entropy regularization or: There’s nothing special about label smoothing. arXiv preprint arXiv:2005.00820.
  61. 61.Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining Well Calibrated Probabilities Using Bayesian Binning. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, AAAI’15, page 2901–2907. AAAI Press.
  62. 62.Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
  63. 63.Jennimaria Palomaki, Olivia Rhinehart, and Michael Tseng. 2018. A case for a range of acceptable annotations. In SAD/CrowdBias@ HCOMP, pages 19–31.
  64. 64.Rebecca J Passonneau, Ansaf Salleb-Aouissi, Vikas Bhardwaj, and Nancy Ide. 2010. Word sense annotation of polysemous words by multiple annotators. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10).
  65. 65.Silviu Paun, Ron Artstein, and Massimo Poesio. 2022. Statistical methods for annotation analysis. Synthesis Lectures on Human Language Technologies, 15(1):1–217.
  66. 66.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  67. 67.Siyao Peng, Yang Janet Liu, and Amir Zeldes. 2022. Gcdt: A chinese rst treebank for multigenre and multilingual discourse parsing. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (AACL-IJCNLP).
  68. 68.Joshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, and Olga Russakovsky. 2019. Human uncertainty makes classification more robust. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9616–9625.
  69. 69.Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014a. Learning part-of-speech taggers with inter-annotator agreement loss. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, pages 742–751. Association for Computational Linguistics.
  70. 70.Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014b. Linguistically debatable or just plain wrong? In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 507–511.
  71. 71.Massimo Poesio and Ron Artstein. 2005. The reliability of anaphoric annotation, reconsidered: Taking ambiguity into account. In Proc. of ACL Workshop on Frontiers in Corpus Annotation, pages 76–83.
  72. 72.Massimo Poesio, Jon Chamberlain, Silviu Paun, Juntao Yu, Alexandra Uma, and Udo Kruschwitz. 2019. A crowdsourced corpus of multiple judgments and disagreement on anaphoric interpretation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1778–1789, Minneapolis, Minnesota. Association for Computational Linguistics.
  73. 73.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of The Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  74. 74.James Pustejovsky and Amber Stubbs. 2012. Natural Language Annotation for Machine Learning: A guide to corpus-building for applications. O’Reilly Media.
  75. 75.Ciyang Qing, Ulle Endriss, Raquel Fernández, and Justin Kruger. 2014. Empirical analysis of aggregation methods for collective annotation. In COLING.
  76. 76.Dennis Reidsma and Jean Carletta. 2008. Reliability measurement without limits. Computational Linguistics, 34(3):319–326.
  77. 77.Dennis Reidsma and Rieks op den Akker. 2008. Exploiting ‘subjective’ annotations. In Proc. of the Workshop on Human Judgments in Computational Linguistics, pages 8–16.
  78. 78.Paul Resnick, Yuqing Kong, Grant Schoenebeck, and Tim Weninger. 2021. Survey equivalence: A procedure for measuring classifier accuracy against human labels. arXiv preprint arXiv:2106.01254.
  79. 79.Filipe Rodrigues and Francisco Pereira. 2018. Deep learning from crowds. In Proceedings of the AAAI conference on artificial intelligence, volume 32.
  80. 80.Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190, Seattle, United States. Association for Computational Linguistics.
  81. 81.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  82. 82.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United States. Association for Computational Linguistics.
  83. 83.David Schlangen. 2021. Targeting the benchmark: On methodology in current natural language processing research. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 670–674, Online. Association for Computational Linguistics.
  84. 84.Victor S. Sheng, Foster Provost, and Panagiotis G. Ipeirotis. 2008. Get another label? improving data quality and data mining using multiple, noisy labelers. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’08, page 614–622, New York, NY, USA. Association for Computing Machinery.
  85. 85.Edwin Simpson, Erik-Lân Do Dinh, Tristan Miller, and Iryna Gurevych. 2019. Predicting humorousness and metaphor novelty with Gaussian process preference learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5716–5728, Florence, Italy. Association for Computational Linguistics.
  86. 86.Rion Snow, Brendan O’connor, Dan Jurafsky, and Andrew Y Ng. 2008. Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks. In Proceedings of the 2008 conference on empirical methods in natural language processing, pages 254–263.
  87. 87.Pia Sommerauer, Antske Fokkens, and Piek Vossen. 2020. Would you describe a leopard as yellow? evaluating crowd-annotations with justified and informative disagreement. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4798–4809, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  88. 88.Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online. Association for Computational Linguistics.
  89. 89.Alexandra Uma, Tommaso Fornaciari, Anca Dumitrache, Tristan Miller, Jon Chamberlain, Barbara Plank, Edwin Simpson, and Massimo Poesio. 2021a. SemEval-2021 task 12: Learning with disagreements. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021), pages 338–347, Online. Association for Computational Linguistics.
  90. 90.Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2020. A case for soft-loss functions. In Proc. of HCOMP.
  91. 91.Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021b. Learning from disagreement: A survey. Journal of Artificial Intelligence Research, 72:1385–1470.
  92. 92.Bonnie Webber and Aravind Joshi. 2012. Discourse structure and computation: Past, present and future. In Proceedings of the ACL-2012 Special Workshop on Rediscovering 50 Years of Discoveries, pages 42–54, Jeju Island, Korea. Association for Computational Linguistics.
  93. 93.Maximilian Wich, Christian Widmer, Gerhard Hagerer, and Georg Groh. 2021. Investigating annotator bias in abusive language datasets. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), pages 1515–1525, Held Online. INCOMA Ltd.
  94. 94.Daniel Zeman. 2010. Hard problems of tagset conversion. In Proceedings of the Second International Conference on Global Interoperability for Language Resources, pages 181–185.
  95. 95.Mike Zhang and Barbara Plank. 2021. Cartography active learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 395–406, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  96. 96.Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021. Learning with different amounts of annotation: From zero to many labels. arXiv preprint arXiv:2109.04408.
  97. 97.Xinliang Frederick Zhang and Marie-Catherine de Marneffe. 2021. Identifying inherent disagreement in natural language inference. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4908–4915, Online. Association for Computational Linguistics.

Citation

MLA
Plank, B. “The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 10671–82, https://doi.org/10.18653/v1/2022.emnlp-main.731.
APA
Plank, B. (2022). The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10671–10682. https://doi.org/10.18653/v1/2022.emnlp-main.731
Chicago
Plank, B. 2022. “The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10671–82. https://doi.org/10.18653/v1/2022.emnlp-main.731.
Harvard
Plank, B. (2022) “The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10671–10682. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.731.
Vancouver
1. Plank B (2022) The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10671–10682

BibTeX

@inproceedings{plank-2022-problem,
    title = "The ``Problem'' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation",
    author = "Plank, Barbara",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.731/",
    doi = "10.18653/v1/2022.emnlp-main.731",
    pages = "10671--10682"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/