The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics
Matthias OrlikowskiPaul RöttgerPhilipp CimianoDirk Hovy
Demonstrates that incorporating sociodemographic attributes into multi-annotator models fails to improve individual label prediction in toxic content detection, cautioning researchers against the ecological fallacy of reducing individual human perspectives to demographic group averages.
Natural language processing systems often rely on human annotators to label text, but different individuals frequently disagree on the same content. This label variation is especially common in subjective tasks such as toxic content detection and is known to correlate with annotator sociodemographics like age and gender. Recent research has shifted toward predicting individual annotator decisions directly rather than relying on a single majority-vote label. However, assuming that group-level characteristics explain individual judgments risks committing an ecological fallacy, where aggregate demographic trends fail to reflect personal behavior.
The article evaluates whether explicitly incorporating annotator sociodemographic attributes into multi-annotator machine learning models improves the prediction of individual annotator decisions. Specifically, it tests whether group-level identity features enhance individual toxicity labeling performance compared to models that do not use these features.
To test this, the authors scaled a multi-task multi-annotator model framework to evaluate 111,780 toxicity annotations from 5,002 annotators across 22,360 social media comments. They introduced group-specific neural network layers representing four sociodemographic attributes: gender, age, education level, and sexual orientation. They compared this demographic-aware architecture against two benchmarks: an individual baseline model that predicts annotator behavior without group information, and a control model where annotators were assigned to demographic groups at random while preserving group sizes. The evaluation used three iterations of four-fold cross-validation with statistical significance assessed via replicability analysis.
The analysis yielded three core findings. First, explicitly adding sociodemographic group layers did not produce statistically significant performance improvements over the baseline model across any evaluated demographic categories. Second, the sociodemographic models performed comparably to the randomized control models, showing that slight score variations were driven by additional model parameters rather than meaningful demographic signals. Third, the results confirmed that multi-annotator frameworks can successfully scale to over 5,000 annotators, a substantial increase compared to prior implementations that evaluated fewer than 100 annotators.
These findings indicate that individual annotation behavior is governed by personal factors that go far beyond broad sociodemographic categories, including personal values, cognitive biases, and psychological traits. When models already learn from an individual's past labeling history, group demographics offer little additional predictive power. For organizational leaders and technical teams, this demonstrates that collecting and encoding sensitive demographic data adds computational complexity and privacy risk without providing measurable accuracy gains in individual behavior modeling.
Organizations developing automated moderation or personalized language models should exercise caution when using broad demographic categories as predictive shortcuts. Future efforts should explore finer-grained representations, such as intersecting identities (intersectionality) or alternative modeling approaches like probabilistic graphical models and recommender systems. Where annotator modeling is required, developers should prioritize robust individual-level representations while implementing aggregation methods that protect individual privacy.
Confidence in these findings is supported by thorough cross-validation and statistical significance testing across a large annotator pool. However, readers should note several limitations: the source dataset reflects only self-reported data from United States annotators, focuses on a single broad toxicity detection task, and evaluates a single base language model. Results may differ for tasks with targeted demographic abuse, international user pools, or larger foundational models.
- Paper: The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, Barbara Plank (2022). Plank’s account of label variation as meaningful signal, rather than noise to collapse, grounds the source’s decision to model individual judgments.
- Paper: Learning From Crowds, V. Raykar et al. (2010). Raykar et al.’s probabilistic framework for learning from multiple annotators provides foundational context for the source’s multi-annotator modeling approach.
- Paper: When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks, Eve Fleisig et al. (2023). This work models individual judgments in subjective language tasks, introducing the disagreement-focused approach that the source tests against demographic group signals.
- Paper: Quantifying the Persona Effect in LLM Simulations, Tiancheng Hu et al. (2024). Building on the source’s finding that broad demographics explain little annotation variation, this study tests how far persona prompting can simulate individual human judgments.
- Paper: The steerability of large language models toward data-driven personas, Junyi Li et al. (2024). This study takes the source’s warning about coarse demographic proxies further by steering language models with individual- and cluster-level opinion patterns instead.
- Paper: MaxMin-RLHF: Alignment with Diverse Human Preferences, Souradip Chakraborty et al. (2024). Extending the challenge of preserving diverse individual preferences, this work applies group-aware modeling to alignment and tests how to protect worse-served preference groups.
