The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
Eve FleisigSu Lin BlodgettDan KleinZeerak Talat
Critiques standard machine learning practices that aggregate human labels into a single ground truth, providing practical recommendations to treat annotator disagreement as valuable signal rather than noise.
Data labeling for artificial intelligence systems has long operated under the assumption that multiple human annotations should be collapsed into a single, aggregated "ground truth" label per example. In practice, annotator disagreement has typically been dismissed as statistical noise, poor data quality, or annotator ineptitude. As artificial intelligence systems are deployed in socially consequential domains like content moderation and conversational assistants, this traditional aggregation approach introduces significant risks. Eliminating divergent human judgments often erases minoritized perspectives, distorts true population beliefs, and builds miscalibrated models that fail in deployment.
The article synthesizes the emerging "perspectivist" paradigm shift in machine learning, which treats annotator disagreement not as error to eliminate, but as valuable signal. Its primary objective is to evaluate the foundational assumptions of both longstanding and perspectivist data annotation approaches, delineate the practical and normative challenges that persist, and provide an actionable framework for capturing human label variation across the artificial intelligence development lifecycle.
The authors conducted a conceptual synthesis and meta-analysis of data labeling practices across natural language processing and broader machine learning research. By analyzing crowdsourcing dynamics, demographic studies, and statistical aggregation techniques, the article examines how current methods fail to represent target populations and explores the implications of shifting from single-label truth models to perspectivist frameworks that capture label distributions.
The article presents several critical findings. First, annotator disagreement is not confined to subjective tasks; it regularly occurs in foundational, seemingly objective tasks like natural language inference, semantic textual similarity, and image classification due to task complexity, differing dialects, and varied interpretations. Second, standard majority-vote aggregation systematically discards valid minority perspectives, moving estimated dataset means further away from true stakeholder population values and generating miscalibrated models that disproportionately reflect dominant demographic groups. Third, while demographic traits influence annotations, non-demographic factors—such as task-specific context, personal lived experience, and an individual's digital habits—frequently drive disagreement more powerfully than isolated traits like gender or age. Over half of crowdworkers explicitly report needing task context, such as system purpose and real-world consequences, which directly shifts their labeling decisions. Fourth, despite the growing collection of rich perspectivist data, model evaluation still bottlenecks around single aggregated "gold" labels because standard engineering pipelines lack non-aggregated evaluation benchmarks.
These findings have direct implications for operational risk, model reliability, and fairness. Treating disagreement as noise introduces silent compliance and safety risks: models trained on naive majority votes embed majority biases while failing to protect vulnerable communities. In addition, organizations that rely on unrepresentative crowdworker platforms risk deploying products that completely misjudge end-user norms. Conversely, moving toward perspectivist data practices creates operational tensions, including higher data collection costs, potential participant privacy trade-offs, and friction with institutional pressures that prioritize development speed over data quality.
To resolve these challenges, the article outlines actionable recommendations across the data pipeline. Development teams must define the acceptable bounds of disagreement before data collection begins, choosing deliberately between prescriptive standards and descriptive variation. During recruitment, teams should stratify annotator pools along task-relevant axes, cap the volume of annotations per individual to avoid dominant-rater skew, and verify quality using intra-annotator consistency checks rather than simple majority consensus. Task designers should provide explicit task context and uncertainty options to annotators, while dataset curators must document selection procedures and retain raw, non-aggregated labels. Finally, model developers should adopt distribution-aware loss functions and evaluation metrics—such as statistical divergence and calibration against human opinion distributions—rather than relying solely on majority-vote accuracy.
The article notes several limitations in the emerging perspectivist paradigm. As a position paper, its findings are based on a synthesis of literature rather than a comprehensive empirical meta-analysis across all artificial intelligence domains. Crucially, the authors emphasize that simple personalization or algorithmic diversification cannot bypass normative decisions: organizations must still make deliberate, transparent choices about which human perspectives to prioritize when deploying systems that produce singular real-world outcomes.
- Paper: The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, Barbara Plank (2022). Plank lays out how collapsing human label variation into a single ground truth distorts data, modeling, and evaluation—the central premise this synthesis examines.
- Paper: When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks, Eve Fleisig et al. (2023). This study shows how majority voting can silence the perspectives of targeted groups, grounding the source’s account of disagreement as structured signal rather than noise.
- Paper: Stop Measuring Calibration When Humans Disagree, Joris Baan et al. (2022). Its analysis of calibration against a single gold label prepares readers to understand why the source calls for evaluating models against distributions of human judgments.
- Paper: Learning From Crowds, V. Raykar et al. (2010). Raykar et al. establish the classic probabilistic approach to multiple noisy annotations, providing a useful contrast for the source’s critique of latent-truth assumptions.
- Paper: MaxMin-RLHF: Alignment with Diverse Human Preferences, Souradip Chakraborty et al. (2024). MaxMin-RLHF carries the source’s critique of majority-dominated preferences into language-model alignment, optimizing for underserved groups rather than a single averaged reward.
