VariErr NLI: Separating Annotation Error from Human Label Variation

Leon Weber-GenzelSiyao PengMarie-Catherine de MarneffeBarbara Plank

article2024ACL67 citationsACL 2024 Area Chair Award

Introduces the VariErr benchmark and a two-round explanation-validation methodology to disentangle genuine human label variation from true annotation errors in natural language inference, exposing critical performance gaps in automatic error detection methods.

Listen

High-quality benchmark data is critical for training and evaluating artificial intelligence models, yet language datasets frequently suffer from annotation errors. At the same time, real-world language often contains inherent ambiguity, meaning multiple distinct interpretations or labels can be valid for the same example. Prior research has treated label errors and legitimate human variation separately, leaving a significant gap in understanding how to systematically untangle true mistakes from acceptable subjective disagreement.

The main objective of the article is to establish a rigorous methodology to separate genuine annotation errors from plausible human label variation and to evaluate how effectively automated error-detection tools, large language models, and human heuristics can identify these errors.

To address this challenge, the authors developed a multi-round annotation framework applied to 500 English natural language inference examples sampled from a challenging subset of the Multi-Genre Natural Language Inference corpus. In the first round, four annotators assigned labels alongside written explanations justifying their choices, generating 1,933 label-explanation pairs. In the second round, all annotators anonymously reviewed and judged the validity of every label-explanation pair, yielding 7,732 validity judgments. An assigned label was formally defined as an error only if the annotator who originally wrote it rejected all corresponding explanations upon review. Using this dataset, named VARIERR, the authors benchmarked traditional automatic error detection algorithms, advanced language models (GPT-3.5 and GPT-4), and human scoring heuristics on their ability to rank and flag dataset errors.

The evaluation revealed four key findings. First, annotation errors are frequently hidden beneath apparent variation; annotators retracted their own original labels as errors in 37.6% of the evaluated examples after seeing peer rationales. Second, traditional model-based automatic error detection methods performed poorly, achieving average precision scores between 17.7% and 22.8% because they frequently misidentified legitimate multi-label variations as errors. Third, GPT-4 significantly outperformed traditional automated tools with an average precision of 31.3% and a precision of 46.0% among the top 100 suspected errors, though it still trailed expert human judgments. Fourth, simple human heuristics based on small-team peer validation achieved the strongest standalone performance (46.5% average precision), while combining human annotation counts with model training dynamics yielded the overall highest error-detection precision (50.4%).

These findings demonstrate that current automated data-cleaning pipelines risk removing valid, nuanced linguistic data rather than actual mistakes. Treating all annotator disagreement as noise risks undermining model robustness, while failing to catch actual errors damages system reliability. Furthermore, the results indicate that reviewing label-explanation pairs with a small team of trained annotators is more effective at purifying datasets than crowdsourcing annotations across hundreds of non-experts.

For organizations aiming to improve data quality, the article demonstrates that two-round validation using written rationales is an effective operational strategy. Teams seeking to automate data auditing should prioritize hybrid pipelines that re-rank human label counts using model training dynamics or advanced language models, rather than relying on conventional automatic error detectors alone.

The authors note several limitations, including the focus on English natural language inference and the reliance on a four-annotator pool. While the underlying methodology is designed to be task-agnostic, organizations should exercise caution before generalizing performance metrics to broader languages or highly specialized domain tasks without pilot validation.

arXiv: 2403.01931mainlp/VariErr-NLI

No sufficiently relevant recommendations were found.

Cover for VariErr NLI: Separating Annotation Error from Human Label Variation

Abstract

Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons. These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation. To the best of our knowledge, there exists no prior work that focuses on teasing apart error from signal, especially in cases where signal is beyond black-and-white. To fill this gap, we introduce a systematic methodology and a new dataset, VariErr (variation versus error), focusing on the NLI task in English. We propose a 2-round annotation procedure with annotators explaining each label and subsequently judging the validity of label-explanation pairs. VariErr contains 7,732 validity judgments on 1,933 explanations for 500 re-annotated MNLI items. We assess the effectiveness of various automatic error detection (AED) methods and GPTs in uncovering errors versus human label variation. We find that state-of-the-art AED methods significantly underperform GPTs and humans. While GPT-4 is the best system, it still falls short of human performance. Our methodology is applicable beyond NLI, offering fertile ground for future research on error versus plausible variation, which in turn can yield better and more trustworthy NLP systems.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 VariErr: Annotation Procedure
  • 3.1 Round 1: NLI Labels & Explanations
  • 3.2 Round 2: Validity Judgments
  • 4 VariErr: Detecting Errors
  • 4.1 Self versus Peer
  • 4.2 Validating Labels
  • 4.3 What counts as an error?
  • 4.4 Data Statistics & IAA
  • 5 Automatic Error Detection (AED) on VariErr
  • 5.1 Task Definition and Evaluation
  • 5.2 Models
  • 5.3 Human Heuristics
  • 6 Results for AED on VariErr
  • 6.1 Human Performance
  • 6.2 Model Performance
  • 6.3 Influence of Human Label Variation
  • 6.4 Reranking models using label counts
  • 7 Conclusion
  • References
  • A Data Statistics
  • A.1 Pair-wise inter-annotator agreements (Cohen’s kappa) with MASI-distance for non-validated, self-validated, and peer-validated versions
  • A.2 Frequency of NLI label on non-validation, self-validated, and peer-validated explanation-label pairs
  • B GPT Prompt
  • C More VariErr Examples

Knowls

  1. Knowl 1 — Two-round annotation procedure for separating variation from error

    model/method

    VARIERR applies a two-round annotation procedure to English natural language inference (NLI), where a label is Entailment, Neutral, or Contradiction. Four annotators independently re-annotated 500 items randomly sampled from the MNLI subset of ChaosNLI, a collection selected for existing label disagreement. In Round 1, annotators could assign one or more NLI labels to each premise–hypothesis pair and wrote a one-sentence explanation for each chosen label; they could also choose “I don’t know” (IDK). This yielded 1,933 label–explanation pairs with standard NLI labels and 331 IDK annotations. In Round 2, the same four annotators anonymously judged every Round 1 label–explanation pair, including their own, as making sense for the label (“yes”), not making sense (“no”), or IDK. The second round produced 7,732 judgments, including 158 IDKs. Because explanations were collected alongside the original labels rather than added afterward, the validity judgments could assess annotators’ stated reasons as well as their labels.

  2. Knowl 2 — Operational criteria for valid explanations, labels, and errors

    definition

    For a Round 1 label–explanation pair, a self-validation is the judgment by the annotator who supplied that pair that its explanation makes sense for its NLI label. A peer-validation occurs when at least two of the other three annotators judge that explanation to make sense. VARIERR defines a label assigned to an NLI item as an annotation error only when none of the explanations supplied for that label is self-validated; one self-validated explanation is sufficient for the label not to count as an error. This is a deliberately strict, self-based error criterion: peer judgments are analyzed separately and do not determine the error labels used as the main evaluation target. Multiple distinct labels can remain plausible for an item, so valid disagreement is not itself treated as error.

  3. Knowl 3 — Validation reduces disagreement while preserving plausible label variation

    data/table

    Across the 500 NLI items, validation removed many label–explanation pairs judged invalid, raised inter-annotator agreement, and still left multiple plausible labels for some items. Repeated counts count each annotator’s assignment separately; aggregated counts count a given label only once per item. Before validation, repeated counts were Entailment 554, Neutral 977, and Contradiction 402 (1,933 total); aggregated counts were 263, 403, and 212 (878 total), respectively. After self-validation, repeated counts were 467, 916, and 329 (1,712 total), and aggregated counts were 210, 380, and 159 (749 total). After peer-validation, repeated counts were 446, 859, and 296 (1,601 total), and aggregated counts were 177, 335, and 130 (642 total). Krippendorff’s α with MASI distance for multi-label judgments was 0.35 before validation, 0.50 after self-validation, and 0.69 after peer-validation. In total, 1,712/1,933 (88.57%) explanations were self-validated and 1,601/1,933 (82.82%) were peer-validated. Under the strict self-based criterion, 188/500 items (37.6%) had at least one error label; peer-validation rejected labels on 258/500 items (51.6%). The validated aggregated labels averaged 749/500 = 1.50 labels per item after self-validation and 642/500 = 1.28 after peer-validation, demonstrating that label variation remained after validation.

  4. Knowl 4 — Automatic error detection is evaluated as label ranking

    experimental setup

    VARIERR evaluates automatic error detection (AED) as a ranking task: a scorer assigns an error score to each Round 1 label for each item, with likely errors ranked first. The evaluation input consists of the 500 NLI items paired with their aggregated Round 1 labels, giving 878 item–label pairs. Rankings are compared with the study’s self-flagged error labels, defined by the rule that a label is erroneous if none of its explanations is self-validated. Performance is measured by average precision (AP) over the ranked labels, plus precision at the top 100 labels (P@100) and recall at the top 100 labels (R@100).

  5. Knowl 5 — Human heuristics and language models score label plausibility using different evidence

    model/method

    The human heuristics use label frequencies or Round 2 peer judgments to rank the 878 aggregated item–label pairs. LCCHAOS uses the number of ChaosNLI annotators (100 crowd-workers) who assigned a label to an item; LCVARIERR uses the number of the four VARIERR annotators who assigned it in Round 1. Both counts are negated when used as error scores, so labels with more annotator support rank as less likely errors. For each label, Peersum sums the peer “yes” judgments across its explanations, while Peeravg first counts “yes” judgments for each explanation and then averages those counts across explanations; both exclude self-judgments and are negated so fewer peer approvals indicate greater error likelihood. GPT-3.5 and GPT-4 receive the premise, hypothesis, and label–explanation pairs, assign each explanation a 0–1 probability of making sense for its label, and use the mean explanation score for that label. Unlike the label-only AED models, the GPT systems have access to explanations.

  6. Knowl 6 — GPT-4 leads automatic scorers, while peer-based heuristics lead overall

    empirical result

    On the self-flagged error-ranking task, the reported values are percentages. In AP / P@100 / R@100 order, the results are: Random 14.7 / 14.7 / 11.4; Metadata Archaeology (MA) 17.7±1.5 / 18.3±4.2 / 14.2±3.2; Datamaps DMmean 22.8±0.4 / 23.7±2.1 / 18.3±1.6; Datamaps DMstd 22.3±1.9 / 22.7±1.2 / 17.6±0.9; GPT-3.5 17.6 / 21.0 / 16.3; GPT-4 31.3 / 46.0 / 35.9; LCCHAOS 32.5 / 35.0 / 27.3; LCVARIERR 40.8 / 42.0 / 32.6; Peeravg 42.2 / 46.0 / 35.9; and Peersum 46.5 / 47.0 / 36.7. MA and Datamaps values are means ± standard deviations over three random seeds. GPT-4 is the strongest automatic scorer, but Peersum has higher AP, P@100, and R@100. GPT-4 matches Peeravg on P@100 and R@100, but does not reach human heuristic performance on AP. The peer-based heuristics and LCVARIERR also outperform LCCHAOS, despite using judgments from four VARIERR annotators rather than 100 ChaosNLI crowd-workers.

  7. Knowl 7 — Training-dynamics AED models use label probabilities and error-labeled neighbors

    model/method

    The Datamaps (DM) systems train DistilRoBERTa-base in a multi-label setting on Round 1 VARIERR labels and use the model’s probability for each observed label across training epochs. Let pi,j,ep_{i,j,e} be the probability assigned after epoch ee to the observed jjth label of NLI item ii, and let EE be the number of training epochs. The mean-probability error score negates the average probability, so that labels the model assigns low probability rank as more likely errors: DMmean=−1E∑e=1Epi,j,e\mathrm{DM}_{\mathrm{mean}}=-\frac{1}{E}\sum_{e=1}^{E}p_{i,j,e}. The variability score is the standard deviation of these probabilities: DMstd=1E∑e=1E(pi,j,e+DMmean)2\mathrm{DM}_{\mathrm{std}}=\sqrt{\frac{1}{E}\sum_{e=1}^{E}(p_{i,j,e}+\mathrm{DM}_{\mathrm{mean}})^2}. Metadata Archaeology (MA) represents each label by its vector of epoch-wise negative log probabilities, (−log⁡pi,j,1,…,−log⁡pi,j,E)(-\log p_{i,j,1},\ldots,-\log p_{i,j,E}), and applies a kk-nearest-neighbors classifier with k=20k=20 to predict the number of error-labeled neighbors. It uses two-fold cross-validation: each half supplies the error labels for predicting the other half, and then the halves are reversed. MA and DM use labels rather than explanations.

  8. Knowl 8 — Top-ranked false positives differ between training-dynamics systems and other scorers

    empirical result

    The analysis grouped the top 100 ranked item–label pairs for each scorer into error labels, valid labels from items retaining multiple plausible labels after self-validation (HLV), and other labels from items with a single plausible label and neither an error nor HLV. GPT systems and human heuristics placed only 0–11 of these 100 labels in the “other” category. By contrast, MA and the two Datamaps systems placed 17.6–29.8 such labels in their top 100. Thus, the top-ranked outputs of GPTs and human heuristics were more concentrated on errors and valid variation, whereas training-dynamics AED systems more often ranked labels that were neither.

  9. Knowl 9 — Breaking annotator-count ties with AED scores improves ranking

    empirical result

    LCVARIERR assigns only four possible scores (−1, −2, −3, or −4), producing many ties. The reranking procedure uses another scorer’s values to break those ties while retaining the LCVARIERR ranking where scores differ. Reranking improved AP over un-reranked LCV​ARIERR (40.8%) for every reported scorer except GPT-3.5. Reranked AP was MA 44.2±3.0%, DMmean 50.4±0.7%, DMstd 50.0±1.5%, GPT-3.5 37.6%, GPT-4 47.4%, LCCHAOS 49.8%, Peeravg 47.8%, and Peersum 47.8%. The gains for DMmean and DMstd put their reranked AP above the 46.5% AP of the best un-reranked human heuristic, Peersum, suggesting that combining annotator counts with training-dynamics scores can improve error ranking.

  10. Knowl 10 — Generality and information-use limitations

    limitation

    The two-round procedure was tested only on English NLI, so its effectiveness on other languages or NLP tasks remains unverified. The training-dynamics AED experiments also do not use all the information VARIERR collects: they do not model the soft distribution of labels across annotators or use annotators’ explanations. The paper identifies learning from the soft label distributions and modeling whether explanations fit their labels as possible future approaches, but does not establish that either would improve AED.

Coverage note — The paper’s scorer-correlation analysis is omitted because it is secondary to the benchmark comparisons and does not establish an additional detection method or outcome.

References

  1. 1.Christoph Alt, Aleksandra Gabryszak, and Leonhard Hennig. 2020. TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1558–1569, Online. Association for Computational Linguistics.
  2. 2.Hadi Amiri, Timothy Miller, and Guergana Savova. 2018. Spotting Spurious Data with Neural Networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2006–2016, New Orleans, Louisiana. Association for Computational Linguistics.
  3. 3.Eric Arazo, Diego Ortego, Paul Albert, Noel O’Connor, and Kevin Mcguinness. 2019. Unsupervised Label Noise Modeling and Loss Correction. In Proceedings of the 36th International Conference on Machine Learning, pages 312–321. PMLR.
  4. 4.Lora Aroyo and Chris Welty. 2013. Crowd truth: Harnessing disagreement in crowdsourcing a relation extraction gold standard. ACM Web Science 2013.
  5. 5.Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24.
  6. 6.Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondrej Dusek. 2024. Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 67–93, St. Julian’s, Malta. Association for Computational Linguistics.
  7. 7.Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. 2020. Are we done with ImageNet? arXiv preprint arXiv:2006.07159.
  8. 8.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In International Conference on Learning Representations.
  9. 9.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  10. 10.Samuel R. Bowman and George Dahl. 2021. What will it take to fix benchmarking in natural language understanding? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4843–4855, Online. Association for Computational Linguistics.
  11. 11.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  12. 12.Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a perspectivist turn in ground truthing for predictive computing. Proceedings of the AAAI Conference on Artificial Intelligence, 37(6):6860–6868.
  13. 13.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  14. 14.Marie-Catherine de Marneffe, Christopher D. Manning, and Christopher Potts. 2012. Did it happen? The pragmatic complexity of veridicality assessment. Computational Linguistics, 38(2):301–333.
  15. 15.Markus Dickinson and W. Detmar Meurers. 2003. Detecting Errors in Part-of-Speech Annotation. In 10th Conference of the European Chapter of the Association for Computational Linguistics, Budapest, Hungary. Association for Computational Linguistics.
  16. 16.Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In CHI Conference on Human Factors in Computing Systems, pages 1–19.
  17. 17.Jinchi Huang, Lie Qu, Rongfei Jia, and Binqiang Zhao. 2019. O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural Networks. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3325–3333.
  18. 18.Nan-Jiang Jiang and Marie-Catherine de Marneffe. 2022. Investigating reasons for disagreement in natural language inference. Transactions of the Association for Computational Linguistics, 10:1357–1374.
  19. 19.Nan-Jiang Jiang, Chenhao Tan, and Marie-Catherine de Marneffe. 2023. Ecologically valid explanations for label variation in NLI. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10622–10633, Singapore. Association for Computational Linguistics.
  20. 20.Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2023. Annotation error detection: Analyzing the past and present for a more coherent future. Computational Linguistics, 49(1):157–198.
  21. 21.Stefan Larson, Anish Mahendran, Andrew Lee, Jonathan K. Kummerfeld, Parker Hill, Michael A. Laurenzano, Johann Hauswald, Lingjia Tang, and Jason Mars. 2019. Outlier detection for improved data quality and diversity in dialog systems. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 517–527, Minneapolis, Minnesota. Association for Computational Linguistics.
  22. 22.Christopher D Manning. 2011. Part-of-speech tagging from 97% to 100%: Is it time for some linguistics? In International conference on intelligent text processing and computational linguistics, pages 171–189. Springer.
  23. 23.Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Evan Sabri Eyuboglu, Amirata Ghorbani, Emmett Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas W Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Y Zou, Carole-Jean Wu, Cody Coleman, Andrew Ng, Peter Mattson, and Vijay Janapa Reddi. 2023. DataPerf: Benchmarks for Data-Centric AI Development. In Advances in Neural Information Processing Systems, volume 36, pages 5320–5347. Curran Associates, Inc.
  24. 24.Mohammad Motamedi, Nikolay Sakharnykh, and Tim Kaldewey. 2021. A data-centric approach for training deep neural networks with less data. CoRR, abs/2110.03613.
  25. 25.Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020. What can we learn from collective human opinions on natural language inference data? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9131–9143, Online. Association for Computational Linguistics.
  26. 26.Rodrigo Frassetto Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-stage document ranking with BERT. CoRR, abs/1910.14424.
  27. 27.Curtis G Northcutt, Anish Athalye, and Jonas Mueller. 2021. Pervasive label errors in test sets destabilize machine learning benchmarks. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  28. 28.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  29. 29.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  30. 30.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  31. 31.Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014. Linguistically debatable or just plain wrong? In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 507–511.
  32. 32.Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman, and Zachary Eichenberger. 2020. Identifying Incorrect Labels in the CoNLL-2003 Corpus. In Proceedings of the 24th Conference on Computational Natural Language Learning, pages 215–226, Online. Association for Computational Linguistics.
  33. 33.Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190, Seattle, United States. Association for Computational Linguistics.
  34. 34.Susanna Rücker and Alan Akbik. 2023. CleanCoNLL: A nearly noise-free named entity recognition dataset. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8628–8645, Singapore. Association for Computational Linguistics.
  35. 35.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108.
  36. 36.Lars Schmarje, Vasco Grossmann, Tim Michels, Jakob Nazarenus, Monty Santarossa, Claudius Zelenka, and Reinhard Koch. 2024. Label Smarter, Not Harder: CleverLabel for Faster Annotation of Ambiguous Image Classification with Higher Quality. In Pattern Recognition, pages 459–475, Cham. Springer Nature Switzerland.
  37. 37.Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker. 2023. Metadata archaeology: Unearthing data subsets by leveraging training dynamics. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  38. 38.Pia Sommerauer, Antske Fokkens, and Piek Vossen. 2020. Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4798–4809, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  39. 39.Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9275–9293, Online. Association for Computational Linguistics.
  40. 40.Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021. Learning from Disagreement: A Survey. Journal of Artificial Intelligence Research, 72:1385–1470.
  41. 41.Vijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil, and Rebecca Roelofs. 2022. When does dough become a bagel? Analyzing the remaining mistakes on ImageNet. In Advances in Neural Information Processing Systems, volume 35, pages 6720–6734. Curran Associates, Inc.
  42. 42.Zihan Wang, Jingbo Shang, Liyuan Liu, Lihao Lu, Jiacheng Liu, and Jiawei Han. 2019. CrossWeigh: Training Named Entity Tagger from Imperfect Annotations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5154–5163, Hong Kong, China. Association for Computational Linguistics.
  43. 43.Leon Weber and Barbara Plank. 2023. ActiveAED: A human in the loop improves annotation error detection. In Findings of the Association for Computational Linguistics: ACL 2023, pages 8834–8845, Toronto, Canada. Association for Computational Linguistics.
  44. 44.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  45. 45.Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021. Learning with different amounts of annotation: From zero to many labels. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7620–7632, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Xinliang Frederick Zhang and Marie-Catherine de Marneffe. 2021. Identifying inherent disagreement in natural language inference. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4908–4915, Online. Association for Computational Linguistics.
  47. 47.Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2023. Efficiently programming large language models using sglang. arXiv preprint arXiv:2312.07104.

Citation

MLA
Weber-Genzel, L., et al. “VariErr NLI: Separating Annotation Error from Human Label Variation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 2256–69, https://doi.org/10.18653/v1/2024.acl-long.123.
APA
Weber-Genzel, L., Peng, S., Marneffe, M.-C. de ., & Plank, B. (2024). VariErr NLI: Separating Annotation Error from Human Label Variation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2256–2269. https://doi.org/10.18653/v1/2024.acl-long.123
Chicago
Weber-Genzel, L., S. Peng, M.-C. de . Marneffe, and B. Plank. 2024. “VariErr NLI: Separating Annotation Error from Human Label Variation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2256–69. https://doi.org/10.18653/v1/2024.acl-long.123.
Harvard
Weber-Genzel, L. et al. (2024) “VariErr NLI: Separating Annotation Error from Human Label Variation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2256–2269. Available at: https://doi.org/10.18653/v1/2024.acl-long.123.
Vancouver
1. Weber-Genzel L, Peng S, Marneffe M-C de, Plank B (2024) VariErr NLI: Separating Annotation Error from Human Label Variation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2256–2269

BibTeX

@inproceedings{weber-genzel-etal-2024-varierr,
    title = "{V}ari{E}rr {NLI}: Separating Annotation Error from Human Label Variation",
    author = "Weber-Genzel, Leon  and
      Peng, Siyao  and
      De Marneffe, Marie-Catherine  and
      Plank, Barbara",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.123/",
    doi = "10.18653/v1/2024.acl-long.123",
    pages = "2256--2269"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/