The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

Matthias OrlikowskiPaul RöttgerPhilipp CimianoDirk Hovy

article2023ACL64 citations

Demonstrates that incorporating sociodemographic attributes into multi-annotator models fails to improve individual label prediction in toxic content detection, cautioning researchers against the ecological fallacy of reducing individual human perspectives to demographic group averages.

Listen

Natural language processing systems often rely on human annotators to label text, but different individuals frequently disagree on the same content. This label variation is especially common in subjective tasks such as toxic content detection and is known to correlate with annotator sociodemographics like age and gender. Recent research has shifted toward predicting individual annotator decisions directly rather than relying on a single majority-vote label. However, assuming that group-level characteristics explain individual judgments risks committing an ecological fallacy, where aggregate demographic trends fail to reflect personal behavior.

The article evaluates whether explicitly incorporating annotator sociodemographic attributes into multi-annotator machine learning models improves the prediction of individual annotator decisions. Specifically, it tests whether group-level identity features enhance individual toxicity labeling performance compared to models that do not use these features.

To test this, the authors scaled a multi-task multi-annotator model framework to evaluate 111,780 toxicity annotations from 5,002 annotators across 22,360 social media comments. They introduced group-specific neural network layers representing four sociodemographic attributes: gender, age, education level, and sexual orientation. They compared this demographic-aware architecture against two benchmarks: an individual baseline model that predicts annotator behavior without group information, and a control model where annotators were assigned to demographic groups at random while preserving group sizes. The evaluation used three iterations of four-fold cross-validation with statistical significance assessed via replicability analysis.

The analysis yielded three core findings. First, explicitly adding sociodemographic group layers did not produce statistically significant performance improvements over the baseline model across any evaluated demographic categories. Second, the sociodemographic models performed comparably to the randomized control models, showing that slight score variations were driven by additional model parameters rather than meaningful demographic signals. Third, the results confirmed that multi-annotator frameworks can successfully scale to over 5,000 annotators, a substantial increase compared to prior implementations that evaluated fewer than 100 annotators.

These findings indicate that individual annotation behavior is governed by personal factors that go far beyond broad sociodemographic categories, including personal values, cognitive biases, and psychological traits. When models already learn from an individual's past labeling history, group demographics offer little additional predictive power. For organizational leaders and technical teams, this demonstrates that collecting and encoding sensitive demographic data adds computational complexity and privacy risk without providing measurable accuracy gains in individual behavior modeling.

Organizations developing automated moderation or personalized language models should exercise caution when using broad demographic categories as predictive shortcuts. Future efforts should explore finer-grained representations, such as intersecting identities (intersectionality) or alternative modeling approaches like probabilistic graphical models and recommender systems. Where annotator modeling is required, developers should prioritize robust individual-level representations while implementing aggregation methods that protect individual privacy.

Confidence in these findings is supported by thorough cross-validation and statistical significance testing across a large annotator pool. However, readers should note several limitations: the source dataset reflects only self-reported data from United States annotators, focuses on a single broad toxicity detection task, and evaluates a single base language model. Results may differ for tasks with targeted demographic abuse, international user pools, or larger foundational models.

Cover for The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

Abstract

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregated labels, and we would expect that sociodemographic information is useful for these models. On the other hand, the ecological fallacy states that aggregate group behaviour, such as the behaviour of the average female annotator, does not necessarily explain individual behaviour. To account for sociodemographics in models of individual annotator behaviour, we introduce group-specific layers to multi-annotator models. In a series of experiments for toxic content detection, we find that explicitly accounting for sociodemographic attributes in this way does not significantly improve model performance. This result shows that individual annotation behaviour depends on much more than just sociodemographics.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • Sociodemographics in Annotation Behaviour
  • 3 Data
  • 4 Experiments
  • 4.1 Evaluation Setup
  • 5 Results
  • 6 Discussion
  • 7 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Annotator Sociodemographics in Sample
  • A.2 Significance Tests
  • A.3 Training Details, Hyperparameters and Computational Resources
  • A.4 Number of Annotations per Group across all Test Sets
  • A.5 Full Results

Knowls

  1. Knowl 1 — Group-specific layers for individual annotation prediction

    model/method

    The model predicts a label for a particular annotator and text by combining a shared text encoder, an annotator-specific classification layer, and—when modeling sociodemographics—a group-specific feature transformation. The encoder is RoBERTa-base, whose output has 768 dimensions. Each annotator has a separate classification layer trained on that annotator’s labels. For a sociodemographic model, a fully connected linear layer is inserted between the encoder and annotator-specific layers; annotators sharing a category of the modeled attribute use the same group layer. The experiments model gender, age, education, and sexual orientation separately.

  2. Knowl 2 — Randomized group assignments control for added parameters

    model/method

    To distinguish the effect of sociodemographic information from the effect of adding group-specific parameters, the authors train control models with the same architecture and group sizes as the sociodemographic models, but randomly reassign annotators to groups. The randomized assignments preserve the relative sizes of the actual groups while breaking the link between an annotator’s reported attribute and the group layer used for that annotator. Comparisons between a sociodemographic model and its randomized counterpart therefore test whether the actual attribute-based grouping helps beyond the extra group-specific layers.

  3. Knowl 3 — Toxicity dataset sample and label conversion

    data/table

    Experiments use a sample of the Kumar et al. (2021) annotator-level toxicity dataset. The original dataset contains 107,620 English comments from Twitter, Reddit, and 4Chan, annotated by 17,280 annotators. The study sampled comments until annotations from more than 5,000 annotators were included, then retained the other annotations from those annotators. The resulting sample contains 22,360 comments, 111,780 annotations, and 5,002 annotators; each annotator contributes 20–120 annotations, with a mean of 22.35. Most comments have five annotations. The original five-point ratings were binarized: ratings 2–4 became toxic, and ratings 0–1 became non-toxic. The sample has 78,357 toxic annotations (70.10%) and 33,423 non-toxic annotations (29.90%). The four modeled attributes are gender, age, education, and sexual orientation; group sizes are imbalanced, including 23 nonbinary annotators and 134 homosexual annotators.

  4. Knowl 4 — Training and evaluation protocol

    experimental setup

    The models were evaluated on individual annotation labels using macro-average F1, calculated separately within groups of each modeled attribute. Comparisons between different groups are not meaningful because annotators from different groups generally annotated different examples. Evaluation used three runs of four-fold cross-validation with different random seeds. Test sets contained comments unseen by the relevant annotators during training, and were constructed to have similar proportions of comments whose majority label was toxic or non-toxic. Models used RoBERTa-base, a maximum sequence length of 512 tokens, three training epochs, batch size 8, and initial learning rate 10−510^{-5}. Training used a weighted loss, with label weights calculated per annotator from that fold’s training data.

  5. Knowl 5 — Macro-F1 results across demographic groups

    data/table

    The table reports mean macro F1 and standard deviation across three runs of four-fold cross-validation. It compares the annotator-specific baseline with models using actual sociodemographic groups or randomized groups, separately within each reported group. The averages are generally close across model types; neither the occasional sociodemographic-model advantage nor the more frequent randomized-model advantage amounts to a consistent improvement.

    Attribute Group Baseline Soc-demographic Randomized
    Gender Male 68.00 ±\pm 0.49 67.66 ±\pm 0.46 67.63 ±\pm 0.53
    Gender Female 62.23 ±\pm 0.53 62.25 ±\pm 1.19 62.41 ±\pm 0.92
    Gender Nonbinary 56.33 ±\pm 6.00 56.80 ±\pm 7.24 58.00 ±\pm 7.49
    Age 18–24 59.39 ±\pm 1.58 60.44 ±\pm 1.05 60.52 ±\pm 1.37
    Age 25–34 66.72 ±\pm 0.56 66.63 ±\pm 0.83 66.92 ±\pm 0.51
    Age 35–44 64.50 ±\pm 0.59 64.94 ±\pm 1.33 65.24 ±\pm 0.89
    Age 45–54 65.68 ±\pm 0.66 65.88 ±\pm 1.39 65.98 ±\pm 0.83
    Age 55–64 64.37 ±\pm 1.22 64.94 ±\pm 1.66 64.84 ±\pm 1.30
    Age 65 or older 63.34 ±\pm 2.07 64.70 ±\pm 2.21 62.77 ±\pm 2.39
    Education Associate degree 60.69 ±\pm 1.44 60.54 ±\pm 2.35 60.78 ±\pm 1.62
    Education Bachelor's degree 66.16 ±\pm 0.51 66.23 ±\pm 0.82 66.80 ±\pm 0.54
    Education Doctoral degree 61.93 ±\pm 3.82 63.79 ±\pm 5.03 63.27 ±\pm 3.67
    Education High school 60.53 ±\pm 1.39 60.47 ±\pm 2.22 60.55 ±\pm 1.87
    Education Below high school 58.28 ±\pm 4.68 62.12 ±\pm 4.90 60.17 ±\pm 4.25
    Education Master's degree 69.71 ±\pm 0.86 69.58 ±\pm 0.93 69.45 ±\pm 0.96
    Education Professional degree 66.75 ±\pm 2.37 67.84 ±\pm 3.32 68.62 ±\pm 2.84
    Education College, no degree 58.65 ±\pm 1.19 59.40 ±\pm 1.79 59.99 ±\pm 2.19
    Sexual orientation Bisexual 71.83 ±\pm 1.14 71.42 ±\pm 1.51 69.46 ±\pm 1.95
    Sexual orientation Heterosexual 63.25 ±\pm 0.39 63.32 ±\pm 1.21 63.82 ±\pm 0.55
    Sexual orientation Homosexual 64.43 ±\pm 1.75 66.11 ±\pm 2.20 65.12 ±\pm 1.94
  6. Knowl 6 — No consistent performance gain from actual demographic grouping

    empirical result

    Across the four modeled attributes, the authors report no statistically significant improvement from adding actual sociodemographic group layers to the annotator-specific baseline, and no significant advantage of actual group assignments over randomized assignments. Randomized models have the highest mean score for many groups, while the sociodemographic model has higher averages for some groups—for example, homosexual annotators—but those differences are accompanied by substantial variability in small groups. The result is specific to predicting individual toxicity annotations in this dataset and model setup; it does not establish that demographic information is never useful.

  7. Knowl 7 — Significance analysis found only isolated fold-level differences

    empirical result

    The significance analysis compared models over 12 folds from three four-fold cross-validation runs. For each fold, the authors used a paired bootstrap test with significance level α=0.05\alpha=0.05, 1,000 bootstrap samples, and a sample size equal to 50% of that fold’s test set; Bonferroni correction accounted for overlapping test sets across runs. The appendix reports only isolated corrected significant-fold counts: for sociodemographic versus baseline models, heterosexual and college-with-no-degree groups each had 2 corrected significant folds out of 12; for sociodemographic versus randomized models, bisexual annotators had 2, and heterosexual annotators and college-with-no-degree groups each had 1. These isolated fold results do not show a consistent advantage across runs, which is consistent with the paper’s overall conclusion of no robust performance gain.

  8. Knowl 8 — Group averages need not explain individual annotation behavior

    theoretical result

    The experiments illustrate a risk associated with the ecological fallacy in individual-annotator modeling: a pattern associated with a demographic group does not necessarily provide useful additional information about a particular member of that group. In this setting, explicitly encoding group membership did not improve individual-label prediction over a model already trained on each annotator’s own labels. The authors suggest, as a possible explanation rather than a demonstrated mechanism, that the baseline may already capture other influences on an annotator’s decisions—such as attitudes, moral values, cognitive biases, or psychological traits—so that making a demographic attribute explicit adds little predictive information.

  9. Knowl 9 — The approach scaled to thousands of annotators, with privacy caveats

    empirical result

    The study trained multi-annotator models for 5,002 annotators, a substantially larger annotator count than the 18 and 82 annotators in the prior work cited by the authors. In the reported implementation, RoBERTa-base has 125 million parameters, each 768-dimensional group-specific linear layer adds 590,592 parameters, and each two-class annotator-specific layer adds 1,538 parameters. The authors report no discernible difference in runtime between models with and without group layers or across different numbers of groups. They note that supporting many individual annotators may be relevant to privacy, but reducing privacy risks would require effective aggregation of individual predictions, which their experiments did not evaluate.

  10. Knowl 10 — Limits on generalization and untested settings

    limitation

    The findings come from one dataset of annotators exclusively from the United States, using a broad toxicity definition, and may not generalize to other cultures, tasks, or more narrowly defined forms of harmful content. The models test four attributes individually, not intersections among attributes; the authors note that single-attribute groups may be too coarse to capture experiences associated with interacting identities. Computational constraints also limited the experiments to RoBERTa-base, so results with larger language models may differ. Finally, the study evaluates individual annotation prediction, not aggregation of predictions or performance against majority-vote labels, and therefore supports no conclusion about those evaluation settings.

Coverage note — The full appendix’s residual-category scores, detailed per-fold annotation-count diagnostics, and computational-run-time accounting are omitted because they do not change the main model comparison or its scope.

References

  1. 1.Gavin Abercrombie, Valerio Basile, Sara Tonelli, Verena Rieser, and Alexandra Uma, editors. 2022. Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @LREC2022. European Language Resources Association, Marseille, France.
  2. 2.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2020. Modeling annotator perspective and polarized opinions to improve hate speech detection. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 8, pages 151–154.
  3. 3.Sohail Akhtar, Valerio Basile, and Viviana Patti. 2021. Whose opinions matter? perspective-aware models to identify opinions of hate speech victims in abusive language detection. Preprint arXiv:2106.15896.
  4. 4.Hala Al Kuwatly, Maximilian Wich, and Georg Groh. 2020. Identifying and measuring annotator bias based on annotators’ demographic characteristics. In Proceedings of the Fourth Workshop on Online Abuse and Harms, pages 184–190, Online. Association for Computational Linguistics.
  5. 5.Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021. We need to consider disagreement in evaluation. In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, pages 15–21, Online. Association for Computational Linguistics.
  6. 6.Laura Biester, Vanita Sharma, Ashkan Kazemi, Naihao Deng, Steven Wilson, and Rada Mihalcea. 2022. Analyzing the effects of annotator gender across NLP tasks. In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @LREC2022, pages 10–19, Marseille, France. European Language Resources Association.
  7. 7.Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. 2017. Like trainer, like bot? inheritance of bias in algorithmic content moderation. In Social Informatics, Lecture Notes in Computer Science, pages 405–415. Springer International Publishing.
  8. 8.Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser. 2021. ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7388–7403, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  9. 9.Kimberle Crenshaw. 1989. Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. University of Chicago Legal Forum, 1989(1):Article 8.
  10. 10.Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110.
  11. 11.Rotem Dror, Gili Baumer, Marina Bogomolov, and Roi Reichart. 2017. Replicability analysis for natural language processing: Testing significance with multiple datasets. Transactions of the Association for Computational Linguistics, 5:471–486.
  12. 12.Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018. The hitchhiker’s guide to testing statistical significance in natural language processing. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1383–1392, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Elizabeth Excell and Noura Al Moubayed. 2021. Towards equal gender representation in the annotations of toxic language detection. In Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing, pages 55–65, Online. Association for Computational Linguistics.
  14. 14.Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021. Beyond black & white: Leveraging annotator disagreement via soft-label multi-task learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2591–2597, Online. Association for Computational Linguistics.
  15. 15.Tommaso Fornaciari, Alexandra Uma, Massimo Poesio, and Dirk Hovy. 2022. Hard and soft evaluation of NLP models with BOOtSTrap SAmpling - BooStSa. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 127–134, Dublin, Ireland. Association for Computational Linguistics.
  16. 16.David A. Freedman. 2015. Ecological inference. In James D. Wright, editor, International Encyclopedia of the Social & Behavioral Sciences (Second Edition), pages 868–870. Elsevier.
  17. 17.Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, pages 1–19. Association for Computing Machinery.
  18. 18.Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proceedings of the ACM on Human-Computer Interaction, 6:1–28.
  19. 19.Dirk Hovy and Diyi Yang. 2021. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 588–602, Online. Association for Computational Linguistics.
  20. 20.Emily Jamison and Iryna Gurevych. 2015. Noise or additional information? leveraging crowdsource annotation item agreement for natural language tasks. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 291–297, Lisbon, Portugal. Association for Computational Linguistics.
  21. 21.Jialun Aaron Jiang, Morgan Klaus Scheuerman, Casey Fiesler, and Jed R. Brubaker. 2021. Understanding international perceptions of the severity of harmful content online. PLOS ONE, 16(8).
  22. 22.Deepak Kumar, Patrick Gage Kelley, Sunny Consolvo, Joshua Mason, Elie Bursztein, Zakir Durumeric, Kurt Thomas, and Michael Bailey. 2021. Designing toxic content classification for a diversity of perspectives. In Seventeenth Symposium on Usable Privacy and Security (SOUPS 2021), pages 299–318. USENIX Association.
  23. 23.Savannah Larimore, Ian Kennedy, Breon Haskett, and Alina Arseniev-Koehler. 2021. Reconsidering annotator disagreement about racist language: Noise or signal? In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 81–90, Online. Association for Computational Linguistics.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. Preprint arXiv:1907.11692.
  25. 25.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830.
  26. 26.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  27. 27.Barbara Plank, Dirk Hovy, and Anders Søgaard. 2014. Learning part-of-speech taggers with inter-annotator agreement loss. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, pages 742–751, Gothenburg, Sweden. Association for Computational Linguistics.
  28. 28.Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021. On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  29. 29.W. S. Robinson. 1950. Ecological correlations and the behavior of individuals. American Sociological Review, 15(3):351–357.
  30. 30.Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. Two contrasting data annotation paradigms for subjective NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 175–190, Seattle, United States. Association for Computational Linguistics.
  31. 31.Pratik Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia von Vacano, and Chris Kennedy. 2022. The measuring hate speech corpus: Leveraging rasch measurement theory for data perspectivism. In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP @LREC2022, pages 83–94, Marseille, France. European Language Resources Association.
  32. 32.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. In 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing @ NeurIPS 2019.
  33. 33.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1668–1678, Florence, Italy. Association for Computational Linguistics.
  34. 34.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5884–5906, Seattle, United States. Association for Computational Linguistics.
  35. 35.Qinlan Shen and Carolyn Rose. 2021. What sounds “right” to me? experiential factors in the perception of political ideology. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1762–1771, Online. Association for Computational Linguistics.
  36. 36.Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021. Learning from disagreement: A survey. Journal of Artificial Intelligence Research, 72:1385–1470.
  37. 37.Angelina Wang, Vikram V Ramaswamy, and Olga Russakovsky. 2022. Towards intersectionality in machine learning: Including more identities, handling underrepresentation, and performing evaluation. In 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, pages 336–349. Association for Computing Machinery.
  38. 38.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.

Citation

MLA
Orlikowski, M., et al. “The Ecological Fallacy in Annotation: Modeling Human Label Variation Goes Beyond Sociodemographics”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 1017–29, https://doi.org/10.18653/v1/2023.acl-short.88.
APA
Orlikowski, M., Röttger, P., Cimiano, P., & Hovy, D. (2023). The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1017–1029. https://doi.org/10.18653/v1/2023.acl-short.88
Chicago
Orlikowski, M., P. Röttger, P. Cimiano, and D. Hovy. 2023. “The Ecological Fallacy in Annotation: Modeling Human Label Variation Goes Beyond Sociodemographics”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1017–29. https://doi.org/10.18653/v1/2023.acl-short.88.
Harvard
Orlikowski, M. et al. (2023) “The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 1017–1029. Available at: https://doi.org/10.18653/v1/2023.acl-short.88.
Vancouver
1. Orlikowski M, Röttger P, Cimiano P, Hovy D (2023) The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 1017–1029

BibTeX

@inproceedings{orlikowski-etal-2023-ecological,
    title = "The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics",
    author = {Orlikowski, Matthias  and
      R{\"o}ttger, Paul  and
      Cimiano, Philipp  and
      Hovy, Dirk},
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.88/",
    doi = "10.18653/v1/2023.acl-short.88",
    pages = "1017--1029"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/