keyword
annotator demographics
Annotator demographics refers to the social, cultural, and personal characteristics of human raters who label data for machine learning systems, encompassing attributes such as age, gender, race, ethnicity, education, geographic location, and linguistic background. These identity factors and lived experiences shape how individuals interpret information, particularly in subjective tasks like evaluating toxicity, humor, politeness, or sentiment. In computational research and dataset creation, recording and analyzing annotator demographics helps identify representation gaps, understand systematic patterns in annotator disagreement, and support modeling approaches that capture diverse viewpoints rather than collapsing varied human perspectives into a single aggregate label.
2 items

The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
Eve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat
Why you should read this
Critiques standard machine learning practices that aggregate human labels into a single ground truth, providing practical recommendations to treat annotator disagreement as valuable signal rather than noise.
Longstanding data labeling practices in machine learning involve collecting and aggregating labels from multiple annotators. But what should we do when annotators disagree? Though annotator disagreement has long been seen as a problem to minimize, new perspectivist approaches challenge this assumption by treating disagreement as a valuable source of information. In this position paper, we examine practices and assumptions surrounding the causes of disagreement—some challenged by perspectivist approaches, and some that remain to be addressed—as well as practical and normative challenges for work operating under these assumptions. We conclude with recommendations for the data labeling pipeline and avenues for future research engaging with subjectivity and disagreement.
Added
2026-10-04

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, Dan Klein
Why you should read this
Presents a framework that predicts individual annotator judgments alongside the targeted demographic groups of text to identify when targeted populations disagree with majority-vote labels in offensive content detection.
People often disagree on subjective tasks such as determining what is offensive or toxic online, where each annotator brings their own perspective influenced by factors like culture, identity, and lived experience. In many annotation settings today, however, we ask multiple people — who may have different beliefs — to provide just one label per example, treating majority vote as ground truth. We show how training models to capture individual annotator behavior instead yields better modeling of disagreement patterns among human raters.
Added
2026-10-02
