DivideMix: Learning with Noisy Labels as Semi-supervised Learning
Junnan LiRichard SocherSteven C. H. Hoi
Proposes DivideMix, a framework that treats noisy label learning as a semi-supervised problem by dynamically partitioning data into clean and unlabeled sets and co-training two networks to prevent confirmation bias.
High-quality human data labeling for deep neural networks is expensive, time-consuming, and difficult to scale. Organizations frequently rely on automated web scraping or single-annotator pipelines to collect large datasets, but these practical alternatives inevitably introduce noisy, inaccurate labels that cause standard models to overfit and generalize poorly. The article introduces and evaluates a unified framework called DivideMix, which reformulates training on corrupted labels as a semi-supervised learning problem where questionable data points are stripped of their labels and repurposed to regularize model training.
The framework operates by training two neural networks simultaneously to avoid confirmation bias, where an individual model simply reinforces its own errors. In each training cycle, one network uses statistical loss modeling to partition the dataset into clean labeled samples and noisy unlabeled samples for the peer network. Both models are then trained on the combined data using data augmentation, cross-network label refinement for labeled data, and ensemble label guessing for unlabeled data. The article assesses this methodology across standard synthetic benchmarks, including CIFAR-10 and CIFAR-100 with noise levels ranging from 20% to 90%, as well as two large-scale real-world datasets, Clothing1M and WebVision.
The empirical findings demonstrate substantial performance gains over existing baseline methods. On synthetic benchmarks with extreme 80% to 90% label corruption, DivideMix achieves major accuracy improvements, outperforming existing techniques by up to approximately 10 to 12 percentage points on complex classification tasks. On real-world corrupted datasets, the method achieves state-of-the-art results, including over 12 percentage points higher top-1 accuracy on the WebVision benchmark and higher accuracy on Clothing1M. Ablation analyses confirm that dual-network co-training, data augmentation, and label refinement are critical to maintaining robustness and preventing performance degradation.
These results show that organizations do not need to discard imperfect or noisy data, nor do they need to invest heavily in exhaustive manual data cleaning. Repurposing suspect labels as unlabeled regularization data significantly lowers data curation costs, reduces operational risks tied to inaccurate training inputs, and maintains high classification performance. While training two networks concurrently requires more computing time than simple baseline training, the computational cost remains lower than multi-stage meta-learning alternatives.
Decision-makers and engineering teams working with noisy datasets should consider adopting dual-model semi-supervised training pipelines to salvage untrusted annotations and reduce data preparation overhead. Future initiatives should explore expanding this dual-network framework into domains outside computer vision, such as natural language processing, and test larger-scale deployment in production environments.
- Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning, David Berthelot et al. (2019). DivideMix directly builds upon and improves the MixMatch semi-supervised framework via co-refinement and co-guessing on dynamically partitioned data.
- Paper: Co-teaching: Robust training of deep neural networks with extremely noisy labels, Bo Han et al. (2018). Co-teaching introduces the dual-network paradigm for filtering noisy examples across peer models, which DivideMix adopts and enhances to combat confirmation bias.
- Paper: mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al. (2017). Mixup establishes the core data-mixing augmentation technique utilized throughout MixMatch and the DivideMix semi-supervised pipeline.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). This paper establishes fundamental concepts and baselines for training deep networks under label noise, which DivideMix reframes into a semi-supervised learning problem.
- Paper: MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels, Lu Jiang et al. (2017). MentorNet pioneered dynamic sample weighting and curriculum strategies for noisy label learning, directly motivating DivideMix's mixture-model dataset division.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). Generalized Cross Entropy provides a foundational robust loss baseline analyzed and compared against when modeling per-sample loss distributions.
- Paper: Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results, Antti Tarvainen et al. (2017). Mean Teacher establishes critical consistency regularization and ensembling principles used in modern semi-supervised learning frameworks like DivideMix.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). This survey provides a comprehensive taxonomy and evaluation of deep learning under noisy labels, contextualizing DivideMix within the broader sample selection and hybrid semi-supervised literature.
- Paper: FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence, Kihyuk Sohn et al. (2020). FixMatch advances semi-supervised learning beyond MixMatch through simplified confidence thresholding and strong-weak consistency regularization, offering a next-generation SSL strategy for noisy label pipelines.
