Built independently by an author, for readers. Read the story and support ChapterPal

keyword

annotator disagreement

Annotator disagreement refers to the situation in data annotation where multiple human raters assign different or conflicting labels to the same data instance. In fields such as machine learning and natural language processing, this variation often occurs in subjective or ambiguous tasks, including toxicity detection, sentiment analysis, and content moderation, where multiple plausible interpretations exist. Rather than merely reflecting human error or random noise, annotator disagreement frequently stems from differing individual perspectives, cultural backgrounds, identities, and lived experiences. While conventional data preparation practices often aggregate conflicting labels into a single ground truth using majority voting or consensus rules, contemporary approaches increasingly analyze and model this variation directly to represent subjective uncertainty and preserve diverse human viewpoints.

3 items

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

Matthias Orlikowski, Paul Röttger, Philipp Cimiano, Dirk Hovy

OrganizationsBielefeld UniversityBocconi UniversityUniversity of Oxford

Why you should read this

Demonstrates that incorporating sociodemographic attributes into multi-annotator models fails to improve individual label prediction in toxic content detection, cautioning researchers against the ecological fallacy of reducing individual human perspectives to demographic group averages.

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregated labels, and we would expect that sociodemographic information is useful for these models. On the other hand, the ecological fallacy states that aggregate group behaviour, such as the behaviour of the average female annotator, does not necessarily explain individual behaviour. To account for sociodemographics in models of individual annotator behaviour, we introduce group-specific layers to multi-annotator models. In a series of experiments for toxic content detection, we find that explicitly accounting for sociodemographic attributes in this way does not significantly improve model performance. This result shows that individual annotation behaviour depends on much more than just sociodemographics.

Added

2026-10-03

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Barbara Plank

OrganizationsLMU MunichLudwig Maximilian University of MunichMaiNLP LabMunich Center for Machine Learning

Why you should read this

Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.

Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics. However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers. In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward. This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation. However, few works consider all of these dimensions jointly; and existing research is fragmented. We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward. As datasets are becoming increasingly available, we hope that this synthesized view on the “problem” will lead to an open discussion on possible strategies to devise fundamentally new directions.

Added

2026-10-01