keyword
annotator disagreement
Annotator disagreement refers to the situation in data annotation where multiple human raters assign different or conflicting labels to the same data instance. In fields such as machine learning and natural language processing, this variation often occurs in subjective or ambiguous tasks, including toxicity detection, sentiment analysis, and content moderation, where multiple plausible interpretations exist. Rather than merely reflecting human error or random noise, annotator disagreement frequently stems from differing individual perspectives, cultural backgrounds, identities, and lived experiences. While conventional data preparation practices often aggregate conflicting labels into a single ground truth using majority voting or consensus rules, contemporary approaches increasingly analyze and model this variation directly to represent subjective uncertainty and preserve diverse human viewpoints.
3 items

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics
Matthias Orlikowski, Paul Röttger, Philipp Cimiano, Dirk Hovy
Why you should read this
Demonstrates that incorporating sociodemographic attributes into multi-annotator models fails to improve individual label prediction in toxic content detection, cautioning researchers against the ecological fallacy of reducing individual human perspectives to demographic group averages.
Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregated labels, and we would expect that sociodemographic information is useful for these models. On the other hand, the ecological fallacy states that aggregate group behaviour, such as the behaviour of the average female annotator, does not necessarily explain individual behaviour. To account for sociodemographics in models of individual annotator behaviour, we introduce group-specific layers to multi-annotator models. In a series of experiments for toxic content detection, we find that explicitly accounting for sociodemographic attributes in this way does not significantly improve model performance. This result shows that individual annotation behaviour depends on much more than just sociodemographics.
Added
2026-10-03

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, Dan Klein
Why you should read this
Presents a framework that predicts individual annotator judgments alongside the targeted demographic groups of text to identify when targeted populations disagree with majority-vote labels in offensive content detection.
People often disagree on subjective tasks such as determining what is offensive or toxic online, where each annotator brings their own perspective influenced by factors like culture, identity, and lived experience. In many annotation settings today, however, we ask multiple people — who may have different beliefs — to provide just one label per example, treating majority vote as ground truth. We show how training models to capture individual annotator behavior instead yields better modeling of disagreement patterns among human raters.
Added
2026-10-02

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
Barbara Plank
Why you should read this
Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.
Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics. However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers. In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward. This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation. However, few works consider all of these dimensions jointly; and existing research is fragmented. We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward. As datasets are becoming increasingly available, we hope that this synthesized view on the “problem” will lead to an open discussion on possible strategies to devise fundamentally new directions.
Added
2026-10-01
