Robust Training under Label Noise by Over-parameterization
Sheng LiuZhihui ZhuQing QuChong You
Proposes a sparse over-parameterization framework that exploits implicit algorithmic regularization to isolate label noise from clean data during training, achieving state-of-the-art classification accuracy on corrupted datasets with theoretical guarantees for exact noise separation.
Modern artificial intelligence heavily relies on large, over-parameterized neural networks that contain far more parameters than training samples. While these architectures deliver superior performance across computer vision and language tasks, their high capacity causes them to memorize errors when training data contains incorrect labels. Because manual data annotation is expensive and inherently error-prone, real-world datasets frequently suffer from corrupted labels, which severely degrades model generalization on unseen data.
The article introduces and evaluates Sparse Over-Parameterization (SOP), a principled training method designed to prevent over-parameterized networks from overfitting to noisy labels in classification tasks. The core objective is to demonstrate that explicitly modeling label corruptions through auxiliary over-parameterized variables can separate sparse noise from true data patterns, both in practical deep learning systems and under formal mathematical models.
To achieve this, the authors model label errors using an extra set of parameters that represent the discrepancy between observed labels and true underlying classes. By initializing these parameters at very small values and training the entire system via standard gradient descent with distinct learning rates, the optimization dynamics naturally induce a sparse penalty that isolates erroneous labels. The credibility of the approach was tested through extensive image classification experiments on benchmark datasets with simulated label noise (CIFAR-10 and CIFAR-100) and real-world annotation noise (CIFAR-N, Clothing-1M, and WebVision), alongside theoretical analysis and numerical verification on linearized mathematical models.
Key findings show that the proposed method consistently achieves state-of-the-art test accuracy across diverse noise conditions. On benchmarks with synthetic label corruptions ranging up to 80%, the enhanced variant (SOP+) achieved top-tier performance, reaching 94.0% accuracy on CIFAR-10 with 80% noise and 78.0% on CIFAR-100 with 40% asymmetric noise. On realistic human-annotated noise (CIFAR-10N worst-case noise), the method attained 93.24% accuracy, outperforming leading existing techniques. Additionally, the approach demonstrated superior computational efficiency, training in 1.0 to 2.1 hours on benchmark datasets compared to 2.3 to 5.4 hours for competing methods. Theoretically, the authors proved that gradient descent on linearized over-parameterized models exacts full separation between the underlying ground truth and sparse corruptions under low-rank data conditions.
These findings indicate that organizations can train high-capacity deep learning models directly on imperfect, real-world data without expensive label-cleaning pipelines or complex multi-network training setups. The method significantly reduces the computational overhead and risk of model degradation caused by flawed annotations, offering a mathematically grounded alternative to heuristic filtering techniques.
For technical teams managing data annotation challenges, the primary recommendation is to integrate the sparse over-parameterization framework into existing classification training pipelines, using standard gradient descent optimizers and minimal initialization for noise variables. Future engineering work should explore incorporating structured or group-sparse priors to account for known class-confusion patterns in specific domain applications.
While empirical results are strong across standard image benchmarks, the theoretical exact recovery guarantees currently rely on linearized models and low-rank data assumptions. Decision-makers should validate the technique in specific production environments, particularly where noise rates exceed extreme thresholds or where label corruption patterns deviate significantly from standard sparsity assumptions.
- Paper: Robust Principal Component Analysis: Exact Recovery of Corrupted Low-Rank Matrices via Convex Optimization, John Wright et al. (2009). Establishes the foundational theory of exact separation between low-rank data and sparse corruptions via convex optimization, directly motivating the sparse over-parameterization approach for label noise.
- Paper: Understanding deep learning requires rethinking generalization, Chiyuan Zhang et al. (2017). Demonstrates the capacity of over-parameterized deep neural networks to memorize completely random labels, highlighting the core vulnerability addressed by robust over-parameterization.
- Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). Analyzes the differing training dynamics between clean patterns and noise memorization in deep networks, underpinning the incoherence and algorithmic regularization assumptions used to isolate label corruptions.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). Provides a comprehensive taxonomy of deep learning approaches for noisy labels, contextualizing the limitations of existing sample selection and loss adjustment paradigms.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). Introduces forward and backward loss correction matrices for deep networks under label noise, establishing key baselines for noise-tolerant training.
- Paper: Co-teaching: Robust training of deep neural networks with extremely noisy labels, Bo Han et al. (2018). Presents a prominent multi-network sample selection strategy for noisy labels against which sparse over-parameterization methods are compared.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). Frames noisy label learning as a semi-supervised problem, representing the state-of-the-art hybrid framework prior to sparse over-parameterization.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). Develops generalized noise-robust loss functions bridging cross-entropy and mean absolute error, offering foundational insight into robust training objectives.
- Paper: Learning with Noisy Labels, Nagarajan Natarajan et al. (2013). Formulates theoretical guarantees and unbiased loss estimators for classification under label noise, establishing the classical mathematical foundation of the field.
- Paper: Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations, Jiaheng Wei et al. (2022). Provides real-world human-annotated noisy datasets (CIFAR-N) to evaluate whether label-noise robust methods like sparse over-parameterization generalize beyond synthetic noise distributions.
- Paper: Debiased Learning from Naturally Imbalanced Pseudo-Labels, Xudong Wang et al. (2022). Extends the challenge of learning from corrupted signals by tackling endogenous, imbalanced noise in self-generated pseudo-labels.
