The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
Barbara Plank
Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.
Modern artificial intelligence and natural language processing systems typically rely on supervised learning, where models train on datasets labeled by human annotators. Standard industry and research practice assumes there is a single objective ground truth for every task, collapsing multiple human viewpoints into one aggregate label—often through majority voting—while treating disagreements as mere noise or errors. However, because language and visual interpretation are inherently complex, subjective, and context-dependent, forcing data into single gold labels discards valuable information and creates artificial performance metrics that fail in real-world deployment.
The article aims to evaluate how human label variation affects every stage of the machine learning pipeline—specifically data collection, model training, and performance evaluation. It demonstrates that genuine human disagreement is not noise to be eliminated, but a rich signal that developers must embrace to build more reliable, trustworthy, and inclusive systems.
To establish this framework, the article synthesizes findings across natural language processing, computer vision, and human-computer interaction literature. It reviews existing approaches for handling label discrepancies, categorizes core modeling paradigms, and compiles a comprehensive repository of publicly available datasets containing un-aggregated annotator labels.
The analysis reveals several critical findings. First, irreconcilable label variation is widespread rather than rare, appearing in at least 20 percent of instances in key language inference tasks due to genuine linguistic ambiguity, varying perspectives, and multiple plausible answers. Second, traditional methods of resolving variation through majority aggregation or data filtering discard valuable evidence, and removing low-agreement instances can actively harm model performance. Third, emerging modeling techniques—such as multi-task learning, soft labels, and loss weighting—show that training directly on the distribution of human opinions improves model generalization, robustness, and out-of-distribution performance. Fourth, current evaluation methods suffer from a severe disconnect: even though models are beginning to learn from diverse viewpoints, researchers predominantly evaluate them against single aggregate labels, masking system overconfidence and real-world failure modes.
These findings imply significant risks for organizations deploying artificial intelligence systems. Relying on majority-vote ground truths distorts performance benchmarks, introduces compliance and safety hazards, and risks marginalizing underrepresented viewpoints in sensitive applications like content moderation. Conversely, capturing human label distributions offers cost and efficiency advantages, as models trained on richer, nuanced distributions may require less overall labeled data to achieve strong generalization.
The article recommends that organizations and researchers immediately transition from aggregating labels to collecting and releasing un-aggregated, annotator-level data along with comprehensive metadata, such as annotator backgrounds and uncertainty indicators. System developers should adopt soft evaluation metrics, including cross-entropy and entropy correlation, alongside standard accuracy to assess calibration and trustworthiness. Furthermore, interdisciplinary standards must be developed to systematically distinguish between careless annotation errors and legitimate, informative human disagreement.
While the article provides strong qualitative synthesis and establishes a clear conceptual framework, it notes that current empirical evidence remains fragmented across subfields and that universal standards for cross-task evaluation are not yet fully established. Leaders should proceed with confidence in the conceptual shift toward preserving label variation, while exercising caution regarding implementation until broader empirical benchmarks are standardized across their specific operational domains.
- Paper: Learning From Crowds, V. Raykar et al. (2010). This foundational framework shows how to infer labels and annotator reliability from disagreement, grounding the source’s challenge to treating variation as mere noise.
- Paper: Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise, Jacob Whitehill et al. (2009). GLAD models both labeler expertise and item difficulty when aggregating conflicting judgments, introducing core approaches to human-label variation that the source later synthesizes.
- Paper: Snorkel: Rapid Training Data Creation with Weak Supervision, Alexander J. Ratner et al. (2017). Snorkel estimates the accuracy and correlations of conflicting weak labeling sources without ground truth, clarifying a key line of work the source reviews.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). This later survey carries the source’s concern about subjective, diverse judgments into LLM feedback, examining how annotator populations and value disagreements shape alignment.
- Paper: Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI, Nick Pangakis et al. (2025). This later study applies the source’s scrutiny of human labels to generative-AI annotation, testing where model outputs diverge from human judgments and motivating human-centered evaluation.
