Built independently by an author, for readers. Read the story and support ChapterPal

keyword

label distributions

In machine learning and data annotation, a label distribution is a probability or frequency distribution across a set of possible categories assigned to a single data instance, quantifying the degree to which each label describes that instance. Unlike conventional annotation schemes that enforce a single discrete ground truth or binary multi-label assignments, a label distribution retains the complete spread of ratings or human judgments. This approach preserves meaningful human label variation resulting from task subjectivity, annotator disagreement, multiple valid interpretations, and inherent ambiguity, allowing machine learning models to directly train on, represent, and evaluate uncertainty across diverse perspectives rather than discarding non-majority opinions as noise.

1 item

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Barbara Plank

OrganizationsLMU MunichLudwig Maximilian University of MunichMaiNLP LabMunich Center for Machine Learning

Why you should read this

Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.

Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics. However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers. In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward. This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation. However, few works consider all of these dimensions jointly; and existing research is fragmented. We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward. As datasets are becoming increasingly available, we hope that this synthesized view on the “problem” will lead to an open discussion on possible strategies to devise fundamentally new directions.

Added

2026-10-01