VariErr NLI: Separating Annotation Error from Human Label Variation
Leon Weber-GenzelSiyao PengMarie-Catherine de MarneffeBarbara Plank
Introduces the VariErr benchmark and a two-round explanation-validation methodology to disentangle genuine human label variation from true annotation errors in natural language inference, exposing critical performance gaps in automatic error detection methods.
High-quality benchmark data is critical for training and evaluating artificial intelligence models, yet language datasets frequently suffer from annotation errors. At the same time, real-world language often contains inherent ambiguity, meaning multiple distinct interpretations or labels can be valid for the same example. Prior research has treated label errors and legitimate human variation separately, leaving a significant gap in understanding how to systematically untangle true mistakes from acceptable subjective disagreement.
The main objective of the article is to establish a rigorous methodology to separate genuine annotation errors from plausible human label variation and to evaluate how effectively automated error-detection tools, large language models, and human heuristics can identify these errors.
To address this challenge, the authors developed a multi-round annotation framework applied to 500 English natural language inference examples sampled from a challenging subset of the Multi-Genre Natural Language Inference corpus. In the first round, four annotators assigned labels alongside written explanations justifying their choices, generating 1,933 label-explanation pairs. In the second round, all annotators anonymously reviewed and judged the validity of every label-explanation pair, yielding 7,732 validity judgments. An assigned label was formally defined as an error only if the annotator who originally wrote it rejected all corresponding explanations upon review. Using this dataset, named VARIERR, the authors benchmarked traditional automatic error detection algorithms, advanced language models (GPT-3.5 and GPT-4), and human scoring heuristics on their ability to rank and flag dataset errors.
The evaluation revealed four key findings. First, annotation errors are frequently hidden beneath apparent variation; annotators retracted their own original labels as errors in 37.6% of the evaluated examples after seeing peer rationales. Second, traditional model-based automatic error detection methods performed poorly, achieving average precision scores between 17.7% and 22.8% because they frequently misidentified legitimate multi-label variations as errors. Third, GPT-4 significantly outperformed traditional automated tools with an average precision of 31.3% and a precision of 46.0% among the top 100 suspected errors, though it still trailed expert human judgments. Fourth, simple human heuristics based on small-team peer validation achieved the strongest standalone performance (46.5% average precision), while combining human annotation counts with model training dynamics yielded the overall highest error-detection precision (50.4%).
These findings demonstrate that current automated data-cleaning pipelines risk removing valid, nuanced linguistic data rather than actual mistakes. Treating all annotator disagreement as noise risks undermining model robustness, while failing to catch actual errors damages system reliability. Furthermore, the results indicate that reviewing label-explanation pairs with a small team of trained annotators is more effective at purifying datasets than crowdsourcing annotations across hundreds of non-experts.
For organizations aiming to improve data quality, the article demonstrates that two-round validation using written rationales is an effective operational strategy. Teams seeking to automate data auditing should prioritize hybrid pipelines that re-rank human label counts using model training dynamics or advanced language models, rather than relying on conventional automatic error detectors alone.
The authors note several limitations, including the focus on English natural language inference and the reliance on a four-annotator pool. While the underlying methodology is designed to be task-agnostic, organizations should exercise caution before generalizing performance metrics to broader languages or highly specialized domain tasks without pilot validation.
- Paper: The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation, Barbara Plank (2022). Plank’s account of genuine label variation as meaningful signal provides the conceptual foundation for distinguishing valid disagreement from annotation mistakes.
- Paper: Understanding Dataset Difficulty with V-Usable Information, Kawin Ethayarajh et al. (2022). Its pointwise V-information framework shows how model behavior can reveal mislabeled or difficult examples, preparing you for VariErr’s comparison of automated error-detection methods.
No sufficiently relevant recommendations were found.
