When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve FleisigRediet AbebeDan Klein
Presents a framework that predicts individual annotator judgments alongside the targeted demographic groups of text to identify when targeted populations disagree with majority-vote labels in offensive content detection.
Machine learning systems for subjective tasks such as online hate speech detection typically rely on majority-vote aggregation to establish ground truth labels. This common practice assumes that disagreement among annotators is random noise rather than meaningful signal. In reality, majority aggregation systematically silences minoritized populations whose lived experiences make them better equipped to recognize harmful content directed at their group, especially when those target groups form only a small fraction of the annotator pool.
The article demonstrates that disagreement in subjective language tasks reflects structured differences in perspective that can be modeled directly. It develops and evaluates a framework that predicts individual annotator ratings alongside the demographic target of a statement, with the primary objective of estimating whether members of a targeted group will find a statement offensive—even when the overall annotator majority disagrees.
The researchers implemented a dual-module architecture evaluated on large-scale public datasets covering over 100,000 toxic content examples and social bias frames. The first module uses a pre-trained language model fine-tuned on regression tasks to predict how individual annotators rate offensiveness on a scale from 0 to 4 based on their demographics and survey responses. The second module predicts the demographic groups harmed by the text. The system then matches predicted target groups to annotator profiles to forecast target-group sentiment without relying on user IDs or manual human-in-the-loop jury selection.
Key findings show that incorporating annotator backgrounds significantly outperforms conventional aggregate baselines. Predicting individual ratings using both demographic features and survey responses improved accuracy by 22% and reduced error in estimating annotator variance by 33%. When predicting whether targeted group members considered an utterance offensive, the model achieved a 22% improvement in aggregate rating predictions and a 28% improvement in capturing within-group variance. Crucially, ablation analyses revealed that full demographic profiles are not strictly required: a compact subset comprising annotator race, views on online toxicity, and social media usage matched the predictive power of exhaustive demographic questionnaires.
These results demonstrate that annotator disagreement is an actionable source of model confidence and risk assessment rather than noise. Accurately capturing annotator variance provides automated content moderation pipelines with a metric for uncertainty, allowing platforms to automatically identify ambiguous or contentious text and route it to human experts. Furthermore, achieving high predictive accuracy through non-invasive questions regarding online habits mitigates the privacy and compliance risks associated with collecting sensitive demographic data.
Organizations developing moderation systems should move away from raw majority-vote datasets and instead model individual annotator perspectives and demographic targets to detect underrepresented harms. When collecting annotation data, teams should prioritize recruiting representative annotator pools along high-impact dimensions such as race while utilizing lightweight surveys on online experiences to protect annotator privacy. Decision-makers should treat this model as a tool to allocate review resources and highlight missing perspectives, rather than as a total replacement for diverse, participatory human oversight.
Readers should note certain limitations: the underlying datasets primarily reflect North American English and demographic contexts, which may not generalize to international environments with differing legal and cultural risks. The model also exhibits variable error across distinct demographic categories and target groups, such as higher error rates when analyzing complex socio-political categories. Nevertheless, the findings provide strong confidence that disagreement-aware modeling substantially improves the fairness and accuracy of subjective machine learning tasks.
- Paper: The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, Barbara Plank (2022). Plank’s overview establishes why annotator disagreement is meaningful signal rather than noise, framing the premise that this paper operationalizes for subjective toxicity judgments.
- Paper: Learning From Crowds, V. Raykar et al. (2010). This foundational method for learning from multiple noisy annotators provides useful grounding in modeling individual labels rather than relying on majority-vote aggregation.
- Paper: MaxMin-RLHF: Alignment with Diverse Human Preferences, Souradip Chakraborty et al. (2024). MaxMin-RLHF carries the paper’s critique of majority-grounded targets into preference alignment, optimizing for the worst-served preference groups instead of averaging them away.
- Paper: Quantifying the Persona Effect in LLM Simulations, Tiancheng Hu et al. (2024). This later study tests how far demographic personas can explain subjective annotation differences, probing a key premise behind modeling ratings from annotator profiles.
- Paper: Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI, Nick Pangakis et al. (2025). This later work extends disagreement-aware annotation concerns to generative-AI coding, assessing when automated labels are unreliable enough to require human-centered oversight.
