Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations
Jiaheng WeiZhaowei ZhuHao ChengTongliang LiuGang NiuYang Liu
Presents the CIFAR-10N and CIFAR-100N benchmarks with real-world human-annotated noise and verified ground truth, demonstrating that practical annotation errors are instance-dependent and establishing a standardized testbed for evaluating noise-resistant learning algorithms.
Data labeling in deep learning is expensive and prone to human error, resulting in noisy datasets that hinder model training. Machine learning research on learning with noisy labels has predominantly relied on synthetic label noise that assumes errors occur uniformly or strictly across categories, independent of image content. Existing real-world noisy benchmarks are often excessively large, lack ground-truth clean verification labels, or introduce artificial collection biases, preventing controlled and fair evaluations.
The article aims to introduce and analyze two accessible, verified real-world noisy label benchmark datasets—CIFAR-10N and CIFAR-100N (collectively termed CIFAR-N)—and systematically evaluate how real human annotation errors differ from theoretical synthetic noise models.
To conduct this evaluation, the researchers collected human annotations via Amazon Mechanical Turk for all 50,000 training images in both CIFAR-10 and CIFAR-100. For CIFAR-10N, three independent human annotations were collected per image to create aggregate, individual random, and worst-case noisy sets. For CIFAR-100N, workers provided both coarse and fine labels using an interactive hierarchical interface. The authors then conducted statistical hypothesis testing to evaluate feature dependency, analyzed neural network memorization patterns, and benchmarked a wide suite of noise-robust machine learning algorithms.
The analysis revealed several critical findings. First, human noise is statistically feature-dependent rather than purely class-dependent, rejecting the standard independent noise hypothesis at high significance levels (p-value of 1.8e-36 for CIFAR-10N). Second, error distributions are heavily skewed: annotators confuse visually similar categories (such as trucks with automobiles or worms with snakes) and introduce substantial class imbalances, with some categories receiving four times more annotations than others. Third, CIFAR-100N revealed an inherent multi-label phenomenon where human annotators chose prominent secondary subjects in images, such as labeling a person holding a flatfish as 'man' rather than 'flatfish'. Finally, deep neural networks consistently showed lower classification accuracy and faster memorization and overfitting on human noise compared to synthetic noise with identical transition matrices, exhibiting performance drops of up to 9% under high human noise.
These findings indicate that existing theoretical noise models and standard algorithmic solutions developed on synthetic data do not accurately capture real-world human annotation behaviors. Consequently, models developed under synthetic assumptions risk underperforming when deployed in practical, human-annotated machine learning pipelines. Furthermore, the presence of multiple legitimate objects within single-label datasets highlights a fundamental flaw in evaluating image classifiers against rigid single-label ground truths.
Researchers and machine learning practitioners should shift benchmark evaluations toward realistic, instance-dependent noise settings using open resources like CIFAR-N. Algorithmic development should prioritize methods combining semi-supervised techniques, data augmentation, and dual-network architectures (such as Divide-Mix and ELR+), which showed the highest resilience to human noise. Additionally, data collection workflows should incorporate multi-label handling to accommodate images with multiple valid subjects.
While the article provides high confidence through rigorous hypothesis testing and multi-run benchmarking across dozens of methods, limitations remain regarding image resolution, as the study focuses on 32x32 pixel images where low resolution inherently exacerbates human error. Further work across higher-resolution imagery and non-visual domains is needed to determine how these exact noise distributions generalize across broader machine learning settings.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). This comprehensive survey categorizes the landscape of deep learning methods for noisy labels that the source evaluates and benchmarks.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). This work establishes the standard class-conditional noise transition matrix formulation that the source re-evaluates and critiques using real human annotations.
- Paper: Co-teaching: Robust training of deep neural networks with extremely noisy labels, Bo Han et al. (2018). This paper introduces the Co-teaching sample-selection framework, representing one of the prominent baseline algorithms evaluated on the CIFAR-N datasets.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). This article presents DivideMix, a leading semi-supervised method for learning with noisy labels that serves as a core benchmark method in the source study.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). This foundational paper develops generalized cross-entropy robust loss functions, providing essential methodology benchmarked against real-world human label noise.
- Paper: MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels, Lu Jiang et al. (2017). This study introduces MentorNet for curriculum-based sample weighting under label noise, establishing key principles for deep neural network noise tolerance examined by the source.
- Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). This paper analyzes how deep networks memorize clean patterns before noisy labels, providing the analytical basis for the source's empirical study of human noise memorization.
- Paper: Learning with Noisy Labels, Nagarajan Natarajan et al. (2013). This work provides classical theoretical foundations for learning under class-conditional label noise assumptions that the source contrasts against instance-dependent human noise.
- Paper: Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise, Jacob Whitehill et al. (2009). This foundational paper presents probabilistic modeling for crowdsourced label aggregation based on annotator ability and item difficulty, motivating human noise characteristics.
- Paper: Get another label? improving data quality and data mining using multiple, noisy labelers, Victor S. Sheng et al. (2008). This early work analyzes the behavior and quality trade-offs of acquiring multiple noisy annotations from non-expert crowd workers.
No sufficiently relevant recommendations were found.
