When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks

Eve FleisigRediet AbebeDan Klein

article2023EMNLP117 citations

Presents a framework that predicts individual annotator judgments alongside the targeted demographic groups of text to identify when targeted populations disagree with majority-vote labels in offensive content detection.

Listen

Machine learning systems for subjective tasks such as online hate speech detection typically rely on majority-vote aggregation to establish ground truth labels. This common practice assumes that disagreement among annotators is random noise rather than meaningful signal. In reality, majority aggregation systematically silences minoritized populations whose lived experiences make them better equipped to recognize harmful content directed at their group, especially when those target groups form only a small fraction of the annotator pool.

The article demonstrates that disagreement in subjective language tasks reflects structured differences in perspective that can be modeled directly. It develops and evaluates a framework that predicts individual annotator ratings alongside the demographic target of a statement, with the primary objective of estimating whether members of a targeted group will find a statement offensive—even when the overall annotator majority disagrees.

The researchers implemented a dual-module architecture evaluated on large-scale public datasets covering over 100,000 toxic content examples and social bias frames. The first module uses a pre-trained language model fine-tuned on regression tasks to predict how individual annotators rate offensiveness on a scale from 0 to 4 based on their demographics and survey responses. The second module predicts the demographic groups harmed by the text. The system then matches predicted target groups to annotator profiles to forecast target-group sentiment without relying on user IDs or manual human-in-the-loop jury selection.

Key findings show that incorporating annotator backgrounds significantly outperforms conventional aggregate baselines. Predicting individual ratings using both demographic features and survey responses improved accuracy by 22% and reduced error in estimating annotator variance by 33%. When predicting whether targeted group members considered an utterance offensive, the model achieved a 22% improvement in aggregate rating predictions and a 28% improvement in capturing within-group variance. Crucially, ablation analyses revealed that full demographic profiles are not strictly required: a compact subset comprising annotator race, views on online toxicity, and social media usage matched the predictive power of exhaustive demographic questionnaires.

These results demonstrate that annotator disagreement is an actionable source of model confidence and risk assessment rather than noise. Accurately capturing annotator variance provides automated content moderation pipelines with a metric for uncertainty, allowing platforms to automatically identify ambiguous or contentious text and route it to human experts. Furthermore, achieving high predictive accuracy through non-invasive questions regarding online habits mitigates the privacy and compliance risks associated with collecting sensitive demographic data.

Organizations developing moderation systems should move away from raw majority-vote datasets and instead model individual annotator perspectives and demographic targets to detect underrepresented harms. When collecting annotation data, teams should prioritize recruiting representative annotator pools along high-impact dimensions such as race while utilizing lightweight surveys on online experiences to protect annotator privacy. Decision-makers should treat this model as a tool to allocate review resources and highlight missing perspectives, rather than as a total replacement for diverse, participatory human oversight.

Readers should note certain limitations: the underlying datasets primarily reflect North American English and demographic contexts, which may not generalize to international environments with differing legal and cultural risks. The model also exhibits variable error across distinct demographic categories and target groups, such as higher error rates when analyzing complex socio-political categories. Nevertheless, the findings provide strong confidence that disagreement-aware modeling substantially improves the fairness and accuracy of subjective machine learning tasks.

arXiv: 2305.06626
Cover for When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks

Abstract

People often disagree on subjective tasks such as determining what is offensive or toxic online, where each annotator brings their own perspective influenced by factors like culture, identity, and lived experience. In many annotation settings today, however, we ask multiple people — who may have different beliefs — to provide just one label per example, treating majority vote as ground truth. We show how training models to capture individual annotator behavior instead yields better modeling of disagreement patterns among human raters.

Table of Contents

  • 1 Introduction
  • 2 Motivation and Related Work
  • 3 Approach
  • 3.1 Individual Rating Prediction Module
  • 3.2 Target Group Prediction Module
  • 4 Results
  • 4.1 Individual Rating Prediction Module
  • 4.2 Target Group Prediction Module
  • 4.3 Full Model Performance
  • 4.4 Ablations
  • 5 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgments
  • References
  • A Input Formatting
  • B Further Model and Dataset Details
  • C Results: Details

Knowls

  1. Knowl 1 — Predicting the opinions of a text’s target group

    model/method

    The system estimates whether a potentially offensive text is harmful to the demographic group or groups it targets by combining two separately trained predictors. One model generates the text’s target group or groups; the other predicts individual annotators’ ratings from the text and information about each annotator. The system standardizes the generated group names, matches them against annotator demographic descriptions, and uses the rating model to estimate the ratings of annotators in the predicted target group who labeled that example. Their predicted aggregate rating and disagreement can then be compared with the broader annotator consensus, helping identify cases where the majority may not reflect the target group’s view. The two models were trained separately because the available datasets did not jointly provide annotator characteristics, individual ratings, and target-group labels.

  2. Knowl 2 — Rating individual annotators from text and annotator information

    model/method

    The rating predictor is a RoBERTa-based regression model fine-tuned first on the Jigsaw toxicity dataset and then on the hate-speech annotations of Kumar et al. (2021). Its input concatenates a natural-language description of an annotator’s survey responses, a natural-language description of their demographics, and the text to be rated, separated by [SEP] tokens. Survey information covers online content use and experiences, including whether the annotator has seen or been personally targeted by toxic content and whether toxic content is a problem. Demographic descriptions can include race, gender, religion’s importance, LGBT status, education, parental status, and political stance. The model is trained with mean squared error to predict ratings from 0 (not at all offensive) to 4 (very offensive), treating each annotator’s rating of an example as a separate training instance. The researchers also tested adding an annotator ID to the input; IDs were assigned to unique combinations of survey and demographic responses.

  3. Knowl 3 — Generating and matching target-group predictions

    model/method

    A GPT-2-based language model fine-tuned on the Social Bias Frames dataset predicts the demographic group or groups targeted by a text. It receives the text followed by a separator and is trained with the standard next-token language-modeling objective to produce a comma-separated, free-text list of target groups. Free-text generation accommodates the long tail of possible groups. After generation, variants of group names are mapped to standardized demographic labels; string matching between those labels and the annotator demographic descriptions selects annotators who belong to at least one predicted target group. When multiple groups are predicted, the relevant target population is their union.

  4. Knowl 4 — Datasets and evaluation setup

    experimental setup

    The individual-rating model was trained and evaluated on Kumar et al. (2021), where each example had ratings from five annotators and each annotator rated about 20 examples on average. The split contained 97,620 training examples and 5,000 examples each for validation and testing. The target-group model used the Social Bias Frames dataset’s existing split: 35,424 training, 4,666 validation, and 4,691 test examples. For target-group evaluation on the Kumar test data, which lacked target-group labels, the researchers manually labeled 100 examples. The models used RoBERTa-base (123 million parameters) and GPT-2-large (1.5 billion parameters); recommended hyperparameters were used without a hyperparameter search. Each training run used two NVIDIA Quadro RTX 8000 GPUs for approximately 12 hours. Rating evaluations report mean absolute error (MAE) for individual ratings, per-example aggregate ratings, and variance among annotators.

  5. Knowl 5 — Annotator information improves individual-rating and disagreement prediction

    data/table

    On the test set, the best rating predictor combined demographic descriptions and survey responses. Compared with the text-only baseline, it reduced individual-rating MAE from 0.88 to 0.69, aggregate-rating MAE from 0.49 to 0.41, and variance MAE from 1.16 to 0.78. These correspond to the paper’s reported improvements of 22%, 16%, and 33%, respectively. Either demographics or surveys alone also improved predictions, while adding IDs did not improve this model’s results. The table compares input-feature choices; lower MAE indicates more accurate prediction.

    Input features Individual rating MAE Aggregate rating MAE Variance MAE
    Text only 0.88 0.49 1.16
    IDs 0.92 0.51 1.11
    Demographics 0.79 0.44 1.01
    Surveys 0.79 0.45 1.00
    Demographics + IDs 0.85 0.46 1.02
    Demographics + ID + survey 0.74 0.41 0.84
    Demographics + survey 0.69 0.41 0.78
  6. Knowl 6 — Target-group generation achieves moderate exact and stronger partial matches

    empirical result

    On the manually annotated set of 100 Kumar test examples, generated target-group descriptions had a word mover’s distance of 0.370 from the human-provided descriptions. The predicted list exactly matched the annotated target-group list in 58% of cases and partially matched it in 81%. Thus, partial identification of relevant groups was more frequent than exact reproduction of the full list.

  7. Knowl 7 — The combined system predicts ratings and uncertainty for target-group members

    data/table

    The combined evaluation first predicted each example’s target group and then assessed predicted ratings for annotators in that group who had labeled the example. Target offense error is the MAE for target-group annotators on examples where the target group’s average observed rating was at least 1; the paper describes that threshold on a 0–5 scale. Against the target-group-only baseline, the best model (demographics plus surveys) improved all four reported measures: individual-rating MAE fell by 18%, aggregate-rating MAE by 22%, variance MAE by 28%, and target offense error by 17%. This evaluates both the group’s average view and variation among its members, rather than treating the group as having a single opinion.

    Metric for target-group annotators Baseline Demographics + survey
    Individual rating MAE 0.89 0.73
    Aggregated rating MAE 0.76 0.59
    Variance MAE 1.69 1.22
    Target offense error 0.96 0.80
  8. Knowl 8 — A small set of less intrusive features approximates demographic-only prediction

    data/table

    Ablations examined which annotator attributes were most useful for predicting ratings. In isolation, the annotator’s view on whether toxic posts are a problem and the annotator’s race were the strongest individual features for individual ratings, with MAEs of 0.87 and 0.88. A model using only race, that opinion, and social-media use achieved individual-rating MAE 0.79, matching the reported result for the model using full demographic information and approaching the full demographics-plus-surveys model’s 0.69. The three-feature combination therefore offers a way to reduce the amount of sensitive information collected while retaining much of the predictive performance. The table reports test MAE for the three rating targets.

    Feature(s) Individual rating MAE Aggregate rating MAE Variance MAE
    Gender 1.0 0.67 1.2
    Race 0.88 0.48 1.1
    LGBT status 1.0 0.67 1.2
    Education level 1.0 0.57 1.2
    Age 0.98 0.56 1.2
    Political leaning 1.0 0.57 1.2
    Importance of religion 0.92 0.51 1.2
    Impact of technology 1.0 0.67 1.2
    Social-media use 1.0 0.67 1.2
    Thinks toxic posts are a problem? 0.87 0.49 1.1
    Personally seen toxic content? 0.95 0.52 1.2
    Race + toxic-post opinion + social-media use 0.79 0.44 1.02
  9. Knowl 9 — Prediction errors vary across annotator and target-group categories

    empirical result

    The model’s individual-rating MAE varied considerably across demographic categories. Reported lower-error annotator groups included conservative annotators, nonbinary annotators, Native American or Alaska Native annotators, Native Hawaiian or Pacific Islander annotators, and annotators with less than a high-school degree, each with MAE below 0.66. Higher-error groups included liberal, politically independent, and transgender annotators, annotators with doctoral or professional degrees, and annotators for whom religion was somewhat important; their reported MAEs were above 1.24. Errors also varied by predicted target group: MAE was below 0.35 for text targeting people described as racist, Syrian, Brazilian, teenagers, or millennials, and above 1.52 for text targeting people described as well-educated, non-violent, Russian, Australian, or non-believers. For text with up to four target groups, individual-rating MAE varied by no more than 0.02 as the number of groups changed.

  10. Knowl 10 — Scope and ethical limits of group-based opinion prediction

    limitation

    The experiments cover English text primarily representing varieties used in the United States, with annotators from the United States and Canada; generalization to other languages and social or political contexts remains uncertain. Text that implicates multiple groups in competing harms may also be difficult to represent with the method’s union of predicted target groups. The authors caution that predicted group opinions must not be treated as uniform or as a substitute for adequately representing minoritized communities in data collection. Detailed annotator attributes can also create re-identification and privacy risks; the reduced-feature results suggest one possible mitigation, but proxy survey variables may themselves introduce skews.

Coverage note — The exhaustive appendix-level per-category error tables and detailed error examples are omitted; the principal subgroup error patterns and the target-group model’s aggregate error findings are included.

References

  1. 1.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  2. 2.Alexander Philip Dawid and Allan M Skene. 1979. Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20–28.
  3. 3.Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, and Hanna Wallach. 2023. FairPrism: Evaluating fairness-related harms in text generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
  4. 4.Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021. Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2591–2597, Online. Association for Computational Linguistics.
  5. 5.Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY, USA. Association for Computing Machinery.
  6. 6.Deepak Kumar, Patrick Gage, Sunny Consolvo, Joshua Mason, Elie Bursztein, Zakir Durumeric, Kurt Thomas, and Michael Bailey. 2021. Designing toxic content classification for a diversity of perspectives. In SOUPS. Usenix.
  7. 7.Savannah Larimore, Ian Kennedy, Breon Haskett, and Alina Arseniev-Koehler. 2021. Reconsidering annotator disagreement about racist language: Noise or signal? In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 81–90, Online. Association for Computational Linguistics.
  8. 8.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  9. 9.Jennimaria Palomaki, Olivia Rhinehart, and Michael Tseng. 2018. A case for a range of acceptable annotations. In SAD/CrowdBias@HCOMP.
  10. 10.Desmond Upton Patton, Philipp Blandfort, William R. Frey, Michael B. Gaskell, and Svebor Karaman. 2019. Annotating social media data from vulnerable populations: Evaluating disagreement between domain experts and graduate student annotators. In Hawaii International Conference on System Sciences.
  11. 11.Ellie Pavlick and Tom Kwiatkowski. 2019. Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics, 7:677–694.
  12. 12.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  13. 13.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. Language models are unsupervised multitask learners.
  14. 14.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  15. 15.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490, Online. Association for Computational Linguistics.
  16. 16.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United States. Association for Computational Linguistics.
  17. 17.Ruyuan Wan, Jaehyung Kim, and Dongyeop Kang. 2023. Everyone’s voice matters: Quantifying annotation disagreement using demographic information. arXiv preprint arXiv:2301.05036.
  18. 18.Zeerak Waseem. 2016. Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science, pages 138–142, Austin, Texas. Association for Computational Linguistics.

Citation

MLA
Fleisig, E., et al. “When the Majority Is Wrong: Modeling Annotator Disagreement for Subjective Tasks”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 6715–26, https://doi.org/10.18653/v1/2023.emnlp-main.415.
APA
Fleisig, E., Abebe, R., & Klein, D. (2023). When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6715–6726. https://doi.org/10.18653/v1/2023.emnlp-main.415
Chicago
Fleisig, E., R. Abebe, and D. Klein. 2023. “When the Majority Is Wrong: Modeling Annotator Disagreement for Subjective Tasks”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6715–26. https://doi.org/10.18653/v1/2023.emnlp-main.415.
Harvard
Fleisig, E., Abebe, R. and Klein, D. (2023) “When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6715–6726. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.415.
Vancouver
1. Fleisig E, Abebe R, Klein D (2023) When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6715–6726

BibTeX

@inproceedings{fleisig-etal-2023-majority,
    title = "When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks",
    author = "Fleisig, Eve  and
      Abebe, Rediet  and
      Klein, Dan",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.415/",
    doi = "10.18653/v1/2023.emnlp-main.415",
    pages = "6715--6726"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/