Mitigating Memorization of Noisy Labels by Clipping the Model Prediction
Hongxin WeiHuiping ZhuangRenchunzi XieLei FengGang NiuBo AnYixuan Li
Proposes LogitClip, a simple yet theoretically grounded method that bounds the norm of model logit vectors to prevent deep neural networks from overfitting to noisy labels without causing the underfitting common to symmetric loss functions.
Modern deep learning systems depend on vast amounts of labeled training data, but practical collection methods such as crowdsourcing and automated web queries routinely introduce label errors. When models are trained on datasets containing noisy labels using standard loss functions like cross-entropy, they suffer from significant performance degradation. Standard cross-entropy loss is mathematically unbounded, causing the training process to overfit corrupted labels and severely harm real-world generalization. While prior solutions introduced specialized noise-robust loss functions, these alternatives often suffer from optimization difficulties and underfitting on complex datasets.
The article addresses this challenge by proposing logit clipping, a lightweight and universally applicable method designed to enhance the noise robustness of existing loss functions. Instead of fundamentally redesigning the loss formulation, logit clipping clamps the vector norm of the model's pre-softmax outputs (logits) to a predefined threshold while preserving the original prediction direction. This constraint enforces a strict mathematical upper bound on the loss value, directly preventing neural networks from memorizing incorrect labels during training.
The researchers evaluated logit clipping through comprehensive theoretical analysis and extensive empirical benchmarks on both synthetic and real-world noisy datasets, including CIFAR-10, CIFAR-100, and the large-scale WebVision dataset. The empirical evaluation covered diverse noise patterns (symmetric, asymmetric, instance-dependent, and real-world noise) across multiple neural network architectures, testing logit clipping alongside standard cross-entropy, focal loss, and several established noise-robust loss functions.
The key findings demonstrate significant and consistent performance gains across training settings. First, integrating logit clipping into standard cross-entropy loss dramatically increases classification accuracy under label noise; for example, test accuracy on CIFAR-10 with instance-dependent noise rose from 68.26% to 86.74%, representing an absolute improvement of over 18 percentage points. Second, the technique universally enhances existing robust loss formulations, such as Generalized Cross Entropy and Normalized Cross Entropy, while also boosting advanced sample-selection and optimization frameworks like DivideMix and Sharpness-Aware Minimization. Third, experiments across diverse model backbones (such as ResNet, DenseNet, and SqueezeNet) and real-world web datasets confirm that the approach is model-agnostic and scalable. Finally, the analysis shows that clipping by vector norm is substantially superior to clipping individual logit values or using soft norm penalties, which degrade gradient signals and alter prediction directions.
These findings have immediate practical implications for machine learning deployment. Organizations can dramatically reduce the financial costs, labor, and timelines associated with exhaustive manual data cleaning. Because logit clipping acts as a simple, end-to-end plug-in for existing training pipelines, engineering teams can implement it without substantial architecture modifications or extra computational overhead, lowering deployment risk and improving model reliability on messy, real-world data.
Practitioners are recommended to incorporate logit clipping when training classification models on datasets with uncertain label quality. Teams should calibrate the clipping threshold hyperparameter on a clean or validation split, balancing noise tolerance against potential underfitting if the threshold is set too conservatively. Further work should explore automated threshold tuning schedules and validate the technique across non-vision domains such as natural language processing and tabular data.
Confidence in these findings is high due to consistent empirical outperformance across diverse noise types and rigorous theoretical risk bounds. However, stakeholders should note that extreme or poorly calibrated clipping thresholds can restrict model expressiveness on pristine, clean datasets, requiring careful hyperparameter validation during production rollout.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). The source evaluates logit clipping as an enhancement to Generalized Cross Entropy, so this paper establishes the robust loss that clipping is shown to improve.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). The source tests logit clipping within DivideMix, and understanding DivideMix’s noisy-sample selection and semi-supervised training clarifies that integration.
- Paper: Mitigating Neural Network Overconfidence with Logit Normalization, Hongxin Wei et al. (2022). Its logit-norm constraint provides a direct conceptual precursor for understanding how the source constrains logits while preserving prediction direction.
- Paper: Towards Understanding Sharpness-Aware Minimization, Maksym Andriushchenko et al. (2022). The source reports gains when clipping is combined with Sharpness-Aware Minimization, and this paper explains the optimizer whose training framework it extends.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). This survey maps the noisy-label methods and failure modes that motivate the source’s bounded-loss approach and position its comparisons.
No sufficiently relevant recommendations were found.
