Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach
Giorgio PatriniAlessandro RozzaAditya MenonRichard NockLizhen Qu
Presents an architecture-agnostic loss correction method that enables deep neural networks to accurately learn from corrupted training data by estimating and mathematically inverting the class-noise transition matrix.
Training modern deep neural networks requires massive volumes of data, but high-quality expert annotation is expensive and slow. Organizations increasingly rely on cost-effective alternatives such as crowdsourcing and automated web scraping, which inevitably introduce class-dependent label noise (systematic mislabeling between confusing or related categories). Without remediation, this corrupted supervision severely degrades model accuracy and reliability across critical applications.
The article establishes a theoretically grounded, architecture-agnostic framework to train robust deep neural networks under class-conditional label noise, presenting an end-to-end pipeline that corrects training loss functions and estimates mislabeling rates directly from noisy data.
The authors develop two distinct loss-adjustment procedures: a "backward" correction that inversely weights the loss by a noise transition matrix summarizing class-flipping probabilities, and a "forward" correction that multiplies the model predictions by the transition matrix. To eliminate the requirement of knowing noise rates beforehand, they extend a class-probability estimation technique to estimate transition matrices without ground-truth labels. The methodology was validated across diverse architectures—including dense, convolutional, recurrent (LSTM), and residual networks—spanning benchmark datasets (MNIST, IMDB, CIFAR-10, CIFAR-100) and a real-world dataset of one million noisy clothing images (Clothing1M).
The evaluation yielded several key findings. First, under severe asymmetric label noise (up to 60%), standard cross-entropy loss degraded drastically by 30 to over 45 percentage points, whereas the proposed forward correction maintained near-clean performance. Second, forward loss correction consistently outperformed backward correction in practical optimization, avoiding numerical instability associated with matrix inversions. Third, the fully automated noise estimator recovered accurate transition matrices directly from noisy samples, suffering a median accuracy drop of only 10 percentage points compared to having perfect prior knowledge of noise rates. Fourth, combining forward loss correction with fine-tuning on a 50-layer residual network achieved 80.38% classification accuracy on Clothing1M, outperforming previous approaches by over two percentage points without requiring complex iterative sampling procedures.
These results provide immediate operational benefits for deploying deep learning in high-scale, cost-sensitive environments. Organizations can significantly lower data curation costs and project timelines by safely training models on weakly labeled, crowdsourced, or web-harvested data without bespoke architecture redesigns. The authors also establish a theoretical guarantee that networks utilizing rectified linear unit (ReLU) activations retain an invariant loss curvature (Hessian) under noise, ensuring stable optimization convergence.
Engineering teams facing noisy datasets should adopt the forward loss correction method as a plug-and-play modification to standard cross-entropy objectives. When transition rates are unknown, teams should employ the automated two-stage estimator using large uncurated datasets to infer noise structures before final model training. Future development should focus on enhancing estimation algorithms with structural priors—such as low-rank matrices—to maintain robustness in heavily fine-grained scenarios with many classes (e.g., CIFAR-100), as well as extending the framework to handle instance-dependent, input-specific noise.
- Paper: Online Passive-Aggressive Algorithms, Koby Crammer et al. (2003). Learn the foundational principles of margin-based online updates and noise-tolerant loss adjustments in linear and margin models.
- Paper: A Simple Weight Decay Can Improve Generalization, Anders Krogh et al. (1991). Understand the classical mathematical analysis of how parameter weight decay curbs overfitting to noisy targets during neural network training.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). Discover how generalized Box-Cox and truncated loss functions improve noise robustness in deep networks without needing explicit noise transition matrix estimation.
- Paper: Co-teaching: Robust training of deep neural networks with extremely noisy labels, Bo Han et al. (2018). Explore an alternative paradigm for handling severe label noise where dual neural networks mutually filter corrupted samples via cross-training.
- Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). Examine empirical analyses of how deep networks prioritize learning clean patterns before memorizing noisy labels, clarifying the training dynamics underlying loss correction.
- Paper: Detecting and Correcting for Label Shift with Black Box Predictors, Zachary C. Lipton et al. (2018). See how confusion-matrix-based linear estimation techniques are extended to detect and correct for distribution-level label shifts using black-box neural predictors.
