Mitigating Memorization of Noisy Labels by Clipping the Model Prediction

Hongxin WeiHuiping ZhuangRenchunzi XieLei FengGang NiuBo AnYixuan Li

article2023ICML51 citations

Proposes LogitClip, a simple yet theoretically grounded method that bounds the norm of model logit vectors to prevent deep neural networks from overfitting to noisy labels without causing the underfitting common to symmetric loss functions.

Listen

Modern deep learning systems depend on vast amounts of labeled training data, but practical collection methods such as crowdsourcing and automated web queries routinely introduce label errors. When models are trained on datasets containing noisy labels using standard loss functions like cross-entropy, they suffer from significant performance degradation. Standard cross-entropy loss is mathematically unbounded, causing the training process to overfit corrupted labels and severely harm real-world generalization. While prior solutions introduced specialized noise-robust loss functions, these alternatives often suffer from optimization difficulties and underfitting on complex datasets.

The article addresses this challenge by proposing logit clipping, a lightweight and universally applicable method designed to enhance the noise robustness of existing loss functions. Instead of fundamentally redesigning the loss formulation, logit clipping clamps the vector norm of the model's pre-softmax outputs (logits) to a predefined threshold while preserving the original prediction direction. This constraint enforces a strict mathematical upper bound on the loss value, directly preventing neural networks from memorizing incorrect labels during training.

The researchers evaluated logit clipping through comprehensive theoretical analysis and extensive empirical benchmarks on both synthetic and real-world noisy datasets, including CIFAR-10, CIFAR-100, and the large-scale WebVision dataset. The empirical evaluation covered diverse noise patterns (symmetric, asymmetric, instance-dependent, and real-world noise) across multiple neural network architectures, testing logit clipping alongside standard cross-entropy, focal loss, and several established noise-robust loss functions.

The key findings demonstrate significant and consistent performance gains across training settings. First, integrating logit clipping into standard cross-entropy loss dramatically increases classification accuracy under label noise; for example, test accuracy on CIFAR-10 with instance-dependent noise rose from 68.26% to 86.74%, representing an absolute improvement of over 18 percentage points. Second, the technique universally enhances existing robust loss formulations, such as Generalized Cross Entropy and Normalized Cross Entropy, while also boosting advanced sample-selection and optimization frameworks like DivideMix and Sharpness-Aware Minimization. Third, experiments across diverse model backbones (such as ResNet, DenseNet, and SqueezeNet) and real-world web datasets confirm that the approach is model-agnostic and scalable. Finally, the analysis shows that clipping by vector norm is substantially superior to clipping individual logit values or using soft norm penalties, which degrade gradient signals and alter prediction directions.

These findings have immediate practical implications for machine learning deployment. Organizations can dramatically reduce the financial costs, labor, and timelines associated with exhaustive manual data cleaning. Because logit clipping acts as a simple, end-to-end plug-in for existing training pipelines, engineering teams can implement it without substantial architecture modifications or extra computational overhead, lowering deployment risk and improving model reliability on messy, real-world data.

Practitioners are recommended to incorporate logit clipping when training classification models on datasets with uncertain label quality. Teams should calibrate the clipping threshold hyperparameter on a clean or validation split, balancing noise tolerance against potential underfitting if the threshold is set too conservatively. Further work should explore automated threshold tuning schedules and validate the technique across non-vision domains such as natural language processing and tabular data.

Confidence in these findings is high due to consistent empirical outperformance across diverse noise types and rigorous theoretical risk bounds. However, stakeholders should note that extreme or poorly calibrated clipping thresholds can restrict model expressiveness on pristine, clean datasets, requiring careful hyperparameter validation during production rollout.

No sufficiently relevant recommendations were found.

Cover for Mitigating Memorization of Noisy Labels by Clipping the Model Prediction

Abstract

In the presence of noisy labels, designing robust loss functions is critical for securing the generalization performance of deep neural networks. Cross Entropy (CE) loss has been shown to be not robust to noisy labels due to its unboundedness. To alleviate this issue, existing works typically design specialized robust losses with the symmetric condition, which usually lead to the underfitting issue. In this paper, our key idea is to induce a loss bound at the logit level, thus universally enhancing the noise robustness of existing losses. Specifically, we propose logit clipping (LogitClip), which clamps the norm of the logit vector to ensure that it is upper bounded by a constant. In this manner, CE loss equipped with our LogitClip method is effectively bounded, mitigating the overfitting to examples with noisy labels. Moreover, we present theoretical analyses to certify the noise-tolerant ability of LogitClip. Extensive experiments show that LogitClip not only significantly improves the noise robustness of CE loss, but also broadly enhances the generalization performance of popular robust losses.

Table of Contents

  • 1. Introduction
  • 2. Motivation and Method
  • 2.1. Preliminaries: The Unboundedness of CE loss
  • 2.2. Our Proposed Method
  • 3. Experiments
  • 3.1. Setups
  • 3.2. CIFAR-10 and CIFAR-100
  • 3.3. WebVision
  • 4. Discussion
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Proof of Proposition 2.1
  • B. Proof of Theorem 2.2
  • C. Proof of Theorem 2.3
  • D. Proof of Proposition 2.4
  • E. Theoretical analysis under instance-dependent setting
  • F. More details on experimental setup
  • G. More empirical results
  • G.1. Can LogitClip improve deep learning methods?
  • G.2. Logit Clipping vs. Norm Regularization
  • G.3. Performance on Clean datasets
  • G.4. Logit Clipping can improve the latest SOTA methods

Knowls

  1. Knowl 1 — LogitClip bounds model outputs before applying the loss

    model/method

    For a KK-class classifier with logit vector z∈RKz\in\mathbb{R}^K, LogitClip replaces the logits passed to the link function with a norm-clipped vector. For norm order p≥1p\geq 1, threshold τ>0\tau>0, and scale δ>0\delta>0, the clipping map is

    clip⁡τ,δ,p(z)={δz/∥z∥p,∥z∥p≥τ,z,∥z∥p<τ.\operatorname{clip}_{\tau,\delta,p}(z)=\begin{cases}\delta z/\lVert z\rVert_p,&\lVert z\rVert_p\geq\tau,\\z,&\lVert z\rVert_p<\tau.\end{cases}

    The resulting link is σˉ(z)=σ(clip⁡τ,δ,p(z))\bar{\sigma}(z)=\sigma(\operatorname{clip}_{\tau,\delta,p}(z)), where σ\sigma is the original link function, such as softmax. In the canonical setting δ=τ\delta=\tau, clipping ensures ∥clip⁡τ,τ,p(z)∥p≤τ\lVert\operatorname{clip}_{\tau,\tau,p}(z)\rVert_p\leq\tau and preserves the direction of any clipped nonzero logit vector, so the predicted class is unchanged. The loss function itself is not modified; clipping is applied to the model output before the loss.

  2. Knowl 2 — LogitClip gives softmax cross-entropy finite upper and lower bounds

    theoretical result

    For KK-class softmax cross-entropy using canonical LogitClip with the maximum norm (p=∞p=\infty) and threshold τ>0\tau>0, let z∈RKz\in\mathbb{R}^K be the clipped logits and yy the target class. Each clipped logit lies in [−τ,τ][-\tau,\tau], so the loss obeys

    log⁡ ⁣(1+(K−1)e−2τ)≤LCEτ(z,y)≤log⁡ ⁣(1+(K−1)e2τ).\log\!\left(1+(K-1)e^{-2\tau}\right)\leq L^{\tau}_{\mathrm{CE}}(z,y)\leq\log\!\left(1+(K-1)e^{2\tau}\right).

    Thus the loss cannot diverge even when the observed label disagrees with the model's preferred class. The paper states that the boundedness conclusion extends to other norms. As τ\tau grows without bound, these limits approach the range of the original, unbounded cross-entropy; decreasing τ\tau tightens the loss range.

  3. Knowl 3 — Locally Lipschitz composite losses are bounded after logit clipping

    theoretical result

    Consider a composite classification loss whose base loss ϕ(py)\phi(p_y) acts on the predicted probability pyp_y of the target class yy. Under canonical LogitClip with KK classes and threshold τ>0\tau>0, the target probability is restricted to [MτK,NτK][M^K_\tau,N^K_\tau], where

    MτK=11+(K−1)e2τ,NτK=11+(K−1)e−2τ.M^K_\tau=\frac{1}{1+(K-1)e^{2\tau}},\qquad N^K_\tau=\frac{1}{1+(K-1)e^{-2\tau}}.

    If ϕ\phi is Lipschitz continuous with constant LL on that interval, the clipped composite loss satisfies

    ∣Lϕτ(f(x),y)∣≤L(NτK−MτK)+∣ϕ(MτK)∣.\left|L^\tau_\phi(f(x),y)\right|\leq L\left(N^K_\tau-M^K_\tau\right)+\left|\phi(M^K_\tau)\right|.

    Here f(x)f(x) denotes the classifier output for input xx. The result applies to losses that are locally Lipschitz on the restricted probability interval; the paper notes that cross-entropy and focal loss meet this condition there for finite τ\tau, despite being unbounded over the full probability range.

  4. Knowl 4 — Symmetric-noise risk gap is bounded for clipped cross-entropy

    theoretical result

    Suppose labels in a KK-class problem are corrupted by symmetric instance-independent noise at rate η\eta, with η≤1−1/K\eta\leq 1-1/K. Let RLCEτ(f)R_{L^\tau_{\mathrm{CE}}}(f) be the expected clipped cross-entropy risk of classifier ff under clean labels, and RLCEτη(f)R^\eta_{L^\tau_{\mathrm{CE}}}(f) its risk under corrupted labels. Let f⋆f^\star minimize the clean risk and f~\tilde f minimize the noisy risk. The paper's bound is

    0≤RLCEτ(f~)−RLCEτ(f⋆)≤ηK(1−η)K−1 AτK,AτK=log⁡ ⁣(1+(K−1)e2τ1+(K−1)e−2τ).0\leq R_{L^\tau_{\mathrm{CE}}}(\tilde f)-R_{L^\tau_{\mathrm{CE}}}(f^\star)\leq \frac{\eta K}{(1-\eta)K-1}\,A^K_\tau, \qquad A^K_\tau=\log\!\left(\frac{1+(K-1)e^{2\tau}}{1+(K-1)e^{-2\tau}}\right).

    The bound concerns the clean-risk excess of the noisy-risk minimizer. For parameter values where the denominator is positive, decreasing τ\tau decreases the bound's factor AτKA^K_\tau, giving a tighter guarantee.

  5. Knowl 5 — Class-conditional and instance-dependent noise bounds are stated with a minimizer inconsistency

    theoretical result

    For asymmetric class-conditional noise, let ηi\eta_i be the probability that a true class-ii label is corrupted, and let ηij=P(y=j∣y⋆=i)\eta_{ij}=P(y=j\mid y^\star=i) for j≠ij\neq i. The paper assumes ηij<1−ηi\eta_{ij}<1-\eta_i for every wrong class. For instance-dependent noise, it assumes 1−ηx>ηxj1-\eta_x>\eta_{xj} for every input xx and wrong class jj, where ηx\eta_x is the total corruption probability and ηxj=P(y=j∣x)\eta_{xj}=P(y=j\mid x) for a wrong class. With KK classes, clipped cross-entropy threshold τ\tau, and clean distribution over input and true label, the paper gives the respective upper-bound quantities

    BτK=Klog⁡ ⁣(1+(K−1)e2τ1+(K−1)e−2τ)E[1−ηi],CτK=Klog⁡ ⁣(1+(K−1)e2τ1+(K−1)e−2τ)E[1−ηx].B^K_\tau=K\log\!\left(\frac{1+(K-1)e^{2\tau}}{1+(K-1)e^{-2\tau}}\right)\mathbb{E}[1-\eta_i],\qquad C^K_\tau=K\log\!\left(\frac{1+(K-1)e^{2\tau}}{1+(K-1)e^{-2\tau}}\right)\mathbb{E}[1-\eta_x].

    The displayed statements in the paper bound RLCEτη(f⋆)−RLCEτη(f~)R^\eta_{L^\tau_{\mathrm{CE}}}(f^\star)-R^\eta_{L^\tau_{\mathrm{CE}}}(\tilde f) between zero and BτKB^K_\tau or CτKC^K_\tau, respectively. The paper defines f~\tilde f as a global minimizer of noisy risk, so the stated nonnegative direction appears inconsistent with that definition. The bound expressions and assumptions are reported here as printed; this knowl does not resolve that inconsistency.

  6. Knowl 6 — CIFAR experiments test four noise settings with a shared training protocol

    experimental setup

    The CIFAR-10 and CIFAR-100 experiments evaluate symmetric noise at 20% and 50%, asymmetric noise, instance-dependent noise at 40%, and real-world human-annotation noise. CIFAR-10 asymmetric noise flips truck to automobile, bird to airplane, deer to horse, and cat to dog with probability 40%; CIFAR-100 asymmetric noise flips each class to the next class circularly with probability 40%. Instance-dependent noise is synthesized using PDN at a 40% noisy rate. For real-world labels, the experiments use the Worst label set of CIFAR-10N and the Fine label set of CIFAR-100N.

    The CIFAR models use WRN-40-2, trained for 200 epochs with SGD, momentum 0.9, weight decay 0.0005, batch size 128, and initial learning rate 0.1, reduced by a factor of 10 after epochs 80 and 140. LogitClip uses the Euclidean norm and δ=1/τ\delta=1/\tau. The selected 1/τ1/\tau is tuned from {0.1,0.5,1,1.5,…,4.5,5}\{0.1,0.5,1,1.5,\ldots,4.5,5\} using 5,000 noisy validation examples; the final model is trained on the full training set. Reported test accuracy is averaged over the last 10 epochs and five random seeds.

  7. Knowl 7 — LogitClip substantially improves cross-entropy accuracy on both CIFAR datasets

    empirical result

    With WRN-40-2 under the CIFAR protocol, adding LogitClip to cross-entropy improved mean test accuracy in every reported noise condition. The paired results below are baseline CE followed by CE with LogitClip; values are percentages with standard deviations across five runs.

    • CIFAR-10: symmetric 20%, 86.73±0.7286.73\pm0.72 to 91.62±0.1691.62\pm0.16; symmetric 50%, 70.88±0.4670.88\pm0.46 to 84.37±0.3484.37\pm0.34; asymmetric, 78.34±0.5478.34\pm0.54 to 86.91±0.6886.91\pm0.68; instance-dependent, 68.26±0.2168.26\pm0.21 to 86.74±0.5586.74\pm0.55; real-world, 72.85±0.3272.85\pm0.32 to 82.06±0.7082.06\pm0.70.
    • CIFAR-100: symmetric 20%, 64.81±1.1064.81\pm1.10 to 71.59±0.7671.59\pm0.76; symmetric 50%, 47.07±1.0747.07\pm1.07 to 63.16±0.7463.16\pm0.74; asymmetric, 47.68±0.9347.68\pm0.93 to 59.04±0.1859.04\pm0.18; instance-dependent, 52.49±0.7952.49\pm0.79 to 66.24±0.7166.24\pm0.71; real-world, 55.68±0.8155.68\pm0.81 to 58.61±0.3558.61\pm0.35.

    The largest listed CIFAR-10 gain is 18.48 percentage points for instance-dependent noise, using the table's reported values.

  8. Knowl 8 — LogitClip improves a broad range of robust losses on CIFAR

    empirical result

    The CIFAR experiments also apply LogitClip to existing loss functions, comparing each loss with and without clipping under the same noise conditions as the cross-entropy experiments. On CIFAR-10, every reported pairing improved in mean test accuracy across all five conditions for MAE, PHuber-CE, SCE, GCE, Taylor-CE, NCE, AEL, AUL, and Cores, as well as CE and focal loss. On CIFAR-100, all reported pairings improved across the five conditions for focal loss, PHuber-CE, SCE, GCE, Cores, NCE+MAE, and NCE+AGCE, as well as CE.

    Examples illustrate the breadth of the changes: CIFAR-10 asymmetric-noise NCE rose from 83.68±0.85%83.68\pm0.85\% to 88.44±0.53%88.44\pm0.53\%; CIFAR-100 asymmetric-noise Cores rose from 50.24±0.38%50.24\pm0.38\% to 63.32±0.70%63.32\pm0.70\%; and CIFAR-100 symmetric-50% GCE rose from 9.10±0.72%9.10\pm0.72\% to 62.14±1.20%62.14\pm1.20\%. These results show improvements for both conventional and noise-robust losses, although the gains vary by loss and noise setting.

  9. Knowl 9 — LogitClip improves results on WebVision's real-world noisy labels

    empirical result

    On WebVision, the experiment trains ResNet-18 under the Mini setting using images from the first 50 Google-resized classes and evaluates top-1 accuracy on the corresponding clean ILSVRC12 validation classes. Models are trained for 120 epochs with SGD, batch size 128, initial learning rate 0.1, Nesterov momentum 0.9, and weight decay 5×10−45\times10^{-4}; the learning rate is reduced by a factor of 10 after epochs 40 and 80. The reported best and last values are validation accuracy percentages, with last averaged over the final 10 epochs.

    CE with LogitClip achieved 65.12% best and 63.75% last, compared with 62.6% and 60.84% for CE. The strongest reported last-epoch result was 64.50% for NCE+AGCE with LogitClip, versus 62.46% for NCE+AGCE without it. The paper reports that the clipped CE result was the highest best accuracy among the listed methods.

  10. Knowl 10 — LogitClip transfers across architectures and complements other methods

    empirical result

    On noisy CIFAR-10, the paper reports higher test accuracy after adding LogitClip to CE for each of the tested SqueezeNet, ResNet-34, and DenseNet architectures across symmetric-20%, symmetric-50%, asymmetric, instance-dependent, and real-world noise. For example, DenseNet accuracy under symmetric-50% noise increased from 59.34% to 81.29%.

    The method also improved existing training approaches when applied to their model outputs during training. For DivideMix on noisy CIFAR-10, accuracy increased from 94.41% to 95.15% under symmetric-50% noise, 92.02% to 92.73% under asymmetric noise, 94.11% to 95.25% under instance-dependent noise, and 92.24% to 93.16% under real-world noise. For SOP and SAM, LogitClip improved all four reported CIFAR-10/100 symmetric-50% and asymmetric-40% settings. Notable examples are SAM on CIFAR-100 asymmetric-40%, which rose from 51.90% to 72.08%, and SOP on the same setting, which rose from 66.89% to 70.41%.

  11. Knowl 11 — The clipping threshold trades noise robustness against underfitting

    empirical result

    The paper's threshold ablations show a tradeoff: reducing τ\tau tightens the theoretical loss and risk bounds, but an overly restrictive threshold can make optimization underfit. On clean CIFAR-100, CE with LogitClip achieved test accuracies of 75.90% at τ=20\tau=20, 75.51% at 1010, 74.38% at 22, 73.23% at 11, 68.00% at 0.50.5, and 46.04% at 0.250.25; vanilla CE achieved 75.98%. Thus a large threshold approached the unmodified CE result, while small thresholds degraded clean-data accuracy. On noisy CIFAR-10 and CIFAR-100, the best threshold varied across noise types, datasets, and base losses, so the paper tunes it rather than prescribing one universal value.

  12. Knowl 12 — Ablations distinguish norm clipping from related output constraints

    empirical result

    The paper compares norm-based LogitClip with clipping each logit value to [−λ,λ][-\lambda,\lambda], ReLU6, logit-norm regularization, and LogitNorm. Clipping-by-value also bounds the cross-entropy loss, but the reported CIFAR-10 experiments found it inferior to norm clipping, especially under instance-dependent and real-world noise; the paper attributes this partly to altered logit directions and potentially diminished gradients on clipped components. ReLU6 constrains intermediate activations but does not bound the final logits, and the reported experiment did not show the noise-robustness improvement achieved by LogitClip.

    For CIFAR-10, LogitClip outperformed LogitNorm in all four displayed settings: symmetric-50% accuracy was 84.37% versus 83.97%, asymmetric 86.91% versus 81.81%, instance-dependent 86.74% versus 84.56%, and real-world 82.06% versus 80.10%. LogitNorm imposes an exact norm constraint on every logit vector, whereas LogitClip imposes an upper-bound constraint. LogitClip also outperformed logit-norm regularization in the reported noisy-CIFAR comparison; the paper notes that a large regularization coefficient could cause optimization difficulty or failure to converge.

Coverage note — The paper's detailed proofs, complete per-loss CIFAR accuracy tables, and some auxiliary comparisons were not reproduced in full; the knowls retain the main bounds, assumptions, experimental conditions, and representative numerical evidence needed to reconstruct the contributions.

References

  1. 1.Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on Computer and Communications Security, pp. 308–318, 2016.
  2. 2.Arazo, E., Ortego, D., Albert, P., O’Connor, N., and McGuinness, K. Unsupervised label noise modeling and loss correction. In International Conference on Machine Learning, pp. 312–321. PMLR, 2019.
  3. 3.Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al. A closer look at memorization in deep networks. In Proceedings of the 34th International Conference on Machine Learning, pp. 233–242, 2017.
  4. 4.Bai, Y., Yang, E., Han, B., Yang, Y., Li, J., Mao, Y., Niu, G., and Liu, T. Understanding and improving early stopping for learning with noisy labels. In Advances in Neural Information Processing Systems, 2021.
  5. 5.Bengio, Y., Simard, P., and Frasconi, P. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2):157–166, 1994.
  6. 6.Blum, A., Kalai, A., and Wasserman, H. Noise-tolerant learning, the parity problem, and the statistical query model. Journal of the ACM, 50(4):506–519, 2003.
  7. 7.Chen, H., Shah, A., Wang, J., Tao, R., Wang, Y., Xie, X., Sugiyama, M., Singh, R., and Raj, B. Imprecise label learning: A unified framework for learning with various imprecise label configurations. arXiv preprint arXiv:2305.12715, 2023.
  8. 8.Chen, P., Ye, J., Chen, G., Zhao, J., and Heng, P.-A. Beyond class-conditional assumption: A primary attempt to combat instance-dependent label noise. arXiv preprint arXiv:2012.05458, 2020.
  9. 9.Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., and Liu, Y. Learning with instance-dependent label noise: A sample sieve approach. International Conference on Learning Representations, 2021.
  10. 10.Cheng, H., Zhu, Z., Sun, X., and Liu, Y. Mitigating memorization of noisy labels via regularization between representations. In International Conference on Learning Representations (ICLR), 2023.
  11. 11.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255. IEEE, 2009.
  12. 12.Ding, K., Shu, J., Meng, D., and Xu, Z. Improve noise tolerance of robust loss via noise-awareness. arXiv preprint arXiv:2301.07306, 2023.
  13. 13.Fatras, K., Damodaran, B. B., Lobry, S., Flamary, R., Tuia, D., and Courty, N. Wasserstein adversarial regularization for learning with label noise. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  14. 14.Feng, L., Shu, S., Lin, Z., Lv, F., Li, L., and An, B. Can cross entropy loss be robust to label noise? In International Joint Conference on Artificial Intelligence, pp. 2206–2212, 2020.
  15. 15.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2021.
  16. 16.Ghosh, A., Kumar, H., and Sastry, P. Robust loss functions under label noise for deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31, pp. 1919–1925, 2017.
  17. 17.Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In Advances in Neural Information Processing Systems, pp. 8527–8537, 2018.
  18. 18.Hazan, E., Levy, K., and Shalev-Shwartz, S. Beyond convexity: Stochastic quasi-convex optimization. Advances in Neural Information Processing Systems, 28, 2015.
  19. 19.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  20. 20.Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K. Using trusted data to train deep networks on labels corrupted by severe noise. arXiv preprint arXiv:1802.05300, 2018.
  21. 21.Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  22. 22.Hu, W., Li, Z., and Yu, D. Simple and effective regularization methods for training on noisily labeled data with generalization guarantee. arXiv preprint arXiv:1905.11368, 2019.
  23. 23.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708, 2017.
  24. 24.Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., and Keutzer, K. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and < 0.5 mb model size. arXiv preprint arXiv:1602.07360, 2016.
  25. 25.Jiang, L., Zhou, Z., Leung, T., Li, L.-J., and Fei-Fei, L. Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In International Conference on Machine Learning, pp. 2304–2313. PMLR, 2018.
  26. 26.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  27. 27.Laine, S. and Aila, T. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016.
  28. 28.Levy, K. Y. The power of normalization: Faster evasion of saddle points. arXiv preprint arXiv:1611.04831, 2016.
  29. 29.Li, J., Socher, R., and Hoi, S. C. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394, 2020.
  30. 30.Li, S., Xia, X., Ge, S., and Liu, T. Selective-supervised contrastive learning with noisy labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 316–325, 2022a.
  31. 31.Li, S., Xia, X., Zhang, H., Zhan, Y., Ge, S., and Liu, T. Estimating noise transition matrix with label correlations for noisy multi-label learning. In NeurIPS, 2022b.
  32. 32.Li, W., Wang, L., Li, W., Agustsson, E., and Van Gool, L. Webvision database: Visual learning and understanding from web data. arXiv preprint arXiv:1708.02862, 2017.
  33. 33.Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, pp. 2980–2988, 2017.
  34. 34.Liu, M., Wei, J., Liu, Y., and Davis, J. Do humans and machines have the same eyes? human-machine perceptual differences on image classification. arXiv preprint arXiv:2304.08733, 2023.
  35. 35.Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C. Early-learning regularization prevents memorization of noisy labels. Advances in Neural Information Processing Systems, 33, 2020.
  36. 36.Liu, S., Zhu, Z., Qu, Q., and You, C. Robust training under label noise by over-parameterization. In Proceedings of the 39th International Conference on Machine Learning, pp. 14153–14172. PMLR, 2022a.
  37. 37.Liu, S., Zhu, Z., Qu, Q., and You, C. Robust training under label noise by over-parameterization. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 14153–14172. PMLR, 2022b.
  38. 38.Liu, T. and Tao, D. Classification with noisy labels by importance reweighting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(3):447–461, 2015.
  39. 39.Liu, Y. and Guo, H. Peer loss functions: Learning from noisy labels without knowing noise rates. In International Conference on Machine Learning, pp. 6226–6236. PMLR, 2020.
  40. 40.Lukasik, M., Bhojanapalli, S., Menon, A., and Kumar, S. Does label smoothing mitigate label noise? In Proceedings of International Conference on Machine Learning, pp. 6448–6458. PMLR, 2020.
  41. 41.Ma, X., Huang, H., Wang, Y., Romano, S., Erfani, S., and Bailey, J. Normalized loss functions for deep learning with noisy labels. In International Conference on Machine Learning, pp. 6543–6553. PMLR, 2020.
  42. 42.Malach, E. and Shalev-Shwartz, S. Decoupling “when to update" from “how to update". Advances in Neural Information Processing Systems, 30, 2017.
  43. 43.Menon, A. K., Rawat, A. S., Kumar, S., and Reddi, S. Can gradient clipping mitigate label noise? In International Conference on Learning Representations (ICLR), 2020.
  44. 44.Miyato, T., Maeda, S.-i., Koyama, M., and Ishii, S. Virtual adversarial training: A regularization method for supervised and semi-supervised learning. IEEE transactions on Pattern Analysis and Machine Intelligence, 41 (8):1979–1993, 2018.
  45. 45.Murray, R., Swenson, B., and Kar, S. Revisiting normalized gradient descent: Fast evasion of saddle points. IEEE Transactions on Automatic Control, 64(11):4818–4824, 2019.
  46. 46.Nguyen, D. T., Mummadi, C. K., Ngo, T. P. N., Nguyen, T. H. P., Beggel, L., and Brox, T. Self: Learning to filter noisy labels with self-ensembling. In International Conference on Learning Representations, 2020.
  47. 47.Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1944–1952, 2017.
  48. 48.Pichapati, V., Suresh, A. T., Yu, F. X., Reddi, S. J., and Kumar, S. Adaclip: Adaptive clipping for private sgd. arXiv preprint arXiv:1908.07643, 2019.
  49. 49.Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., and Rabinovich, A. Training deep neural networks on noisy labels with bootstrapping. arXiv preprint arXiv:1412.6596, 2014.
  50. 50.Ren, M., Zeng, W., Yang, B., and Urtasun, R. Learning to reweight examples for robust deep learning. In International Conference on Machine Learning, pp. 4334–4343. PMLR, 2018.
  51. 51.Shu, J., Xie, Q., Yi, L., Zhao, Q., Zhou, S., Xu, Z., and Meng, D. Meta-weight-net: Learning an explicit mapping for sample weighting. In Advances in Neural Information Processing Systems, pp. 1919–1930, 2019.
  52. 52.Shu, J., Meng, D., and Xu, Z. Learning an explicit hyperparameter prediction policy conditioned on tasks. arXiv preprint arXiv:2107.02378, 2021.
  53. 53.Shu, J., Yuan, X., Meng, D., and Xu, Z. Cmw-net: Learning a class-aware sample weighting mapping for robust deep learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  54. 54.Sohrab, H. H. Basic real analysis, volume 231. Springer, 2003.
  55. 55.Sun, Y., Guo, C., and Li, Y. React: Out-of-distribution detection with rectified activations. Advances in Neural Information Processing Systems, 34:144–157, 2021.
  56. 56.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 2818–2826, 2016.
  57. 57.Tanaka, D., Ikami, D., Yamasaki, T., and Aizawa, K. Joint optimization framework for learning with noisy labels. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5552–5560, 2018.
  58. 58.Wang, Y., Ma, X., Chen, Z., Luo, Y., Yi, J., and Bailey, J. Symmetric cross entropy for robust learning with noisy labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 322–330, 2019.
  59. 59.Wei, H., Feng, L., Chen, X., and An, B. Combating noisy labels by agreement: A joint training method with co-regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13726–13735, 2020a.
  60. 60.Wei, H., Feng, L., Wang, R., and An, B. Metainfonet: Learning task-guided information for sample reweighting. arXiv preprint arXiv:2012.05273, 2020b.
  61. 61.Wei, H., Xie, R., Cheng, H., Feng, L., An, B., and Li, Y. Mitigating neural network overconfidence with logit normalization. In International Conference on Machine Learning (ICML). PMLR, 2022a.
  62. 62.Wei, H., Xie, R., Feng, L., Han, B., and An, B. Deep learning from multiple noisy annotators as a union. IEEE Transactions on Neural Networks and Learning Systems, pp. 1–11, 2022b.
  63. 63.Wei, J. and Liu, Y. When optimizing ff-divergence is robust with label noise. In International Conference on Learning Representations, 2021.
  64. 64.Wei, J., Liu, H., Liu, T., Niu, G., Sugiyama, M., and Liu, Y. To smooth or not? when label smoothing meets noisy labels. In International Conference on Machine Learning, pp. 23589–23614. PMLR, 2022c.
  65. 65.Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., and Liu, Y. Learning with noisy labels revisited: A study using real-world human annotations. In International Conference on Learning Representations, 2022d.
  66. 66.Wei, J., Zhu, Z., Luo, T., Amid, E., Kumar, A., and Liu, Y. To aggregate or not? learning with separate noisy labels. In ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023.
  67. 67.Wu, S., Xia, X., Liu, T., Han, B., Gong, M., Wang, N., Liu, H., and Niu, G. Class2simi: A noise reduction perspective on learning with noisy labels. In ICML, pp. 11285–11295, 2021a.
  68. 68.Wu, Y., Shu, J., Xie, Q., Zhao, Q., and Meng, D. Learning to purify noisy labels via meta soft label corrector. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 10388–10396, 2021b.
  69. 69.Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., and Sugiyama, M. Are anchor points really indispensable in label-noise learning? In NeurIPS, 2019.
  70. 70.Xia, X., Liu, T., Han, B., Gong, C., Wang, N., Ge, Z., and Chang, Y. Robust early-learning: Hindering the memorization of noisy labels. In International Conference on Learning Representations, 2020a.
  71. 71.Xia, X., Liu, T., Han, B., Wang, N., Gong, M., Liu, H., Niu, G., Tao, D., and Sugiyama, M. Part-dependent label noise: Towards instance-dependent label noise. Advances in Neural Information Processing Systems, 33:7597–7610, 2020b.
  72. 72.Xia, X., Liu, T., Bo, H., Mingming, G., Jun, Y., Gang, N., and Masashi, S. Sample selection with uncertainty of losses for learning with noisy labels. In International Conference on Learning Representations, 2022.
  73. 73.Yan, Y., Rosales, R., Fung, G., Subramanian, R., and Dy, J. Learning from multiple annotators with varying expertise. Machine Learning, 95(3):291–327, 2014.
  74. 74.Yu, X., Han, B., Yao, J., Niu, G., Tsang, I., and Sugiyama, M. How does disagreement help generalization against label corruption? In International Conference on Machine Learning, pp. 7164–7173. PMLR, 2019.
  75. 75.Zagoruyko, S. and Komodakis, N. Wide residual networks. In BMVC, 2016.
  76. 76.Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. Understanding deep learning requires rethinking generalization. In Proceedings of International Conference on Learning Representations, 2016.
  77. 77.Zhang, J., He, T., Sra, S., and Jadbabaie, A. Why gradient clipping accelerates training: A theoretical justification for adaptivity. In International Conference on Learning Representations, 2020.
  78. 78.Zhang, Z. and Sabuncu, M. Generalized cross entropy loss for training deep neural networks with noisy labels. Advances in Neural Information Processing Systems, 31, 2018.
  79. 79.Zheng, S., Wu, P., Goswami, A., Goswami, M., Metaxas, D., and Chen, C. Error-bounded correction of noisy labels. In International Conference on Machine Learning, pp. 11447–11457. PMLR, 2020.
  80. 80.Zhou, X., Liu, X., Jiang, J., Gao, X., and Ji, X. Asymmetric loss functions for learning with noisy labels. In International Conference on Machine Learning, pp. 12846–12856. PMLR, 2021.
  81. 81.Zhu, Z., Liu, T., and Liu, Y. A second-order approach to learning with instance-dependent label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10113–10123, 2021a.
  82. 82.Zhu, Z., Song, Y., and Liu, Y. Clusterability as an alternative to anchor points when learning with noisy labels. In International Conference on Machine Learning (ICML), 2021b.
  83. 83.Zhu, Z., Dong, Z., and Liu, Y. Detecting corrupted labels without training a model to predict. In International Conference on Machine Learning (ICML), 2022a.
  84. 84.Zhu, Z., Wang, J., and Liu, Y. Beyond images: Label noise transition matrix estimation for tasks with lower-quality features. In International Conference on Machine Learning (ICML). PMLR, 17–23 Jul 2022b.

Citation

MLA
Wei, H., et al. “Mitigating Memorization of Noisy Labels by Clipping the Model Prediction”. International Conference on Machine Learning, vol. 202, 2023, pp. 36868–86, https://proceedings.mlr.press/v202/wei23e.html.
APA
Wei, H., Zhuang, H., Xie, R., Feng, L., Niu, G., An, B., & Li, Y. (2023). Mitigating Memorization of Noisy Labels by Clipping the Model Prediction. International Conference on Machine Learning, 202, 36868–36886. https://proceedings.mlr.press/v202/wei23e.html
Chicago
Wei, H., H. Zhuang, R. Xie, et al. 2023. “Mitigating Memorization of Noisy Labels by Clipping the Model Prediction”. International Conference on Machine Learning 202: 36868–86. https://proceedings.mlr.press/v202/wei23e.html.
Harvard
Wei, H. et al. (2023) “Mitigating Memorization of Noisy Labels by Clipping the Model Prediction”, International Conference on Machine Learning. PMLR, pp. 36868–36886. Available at: https://proceedings.mlr.press/v202/wei23e.html.
Vancouver
1. Wei H, Zhuang H, Xie R, Feng L, Niu G, An B, Li Y (2023) Mitigating Memorization of Noisy Labels by Clipping the Model Prediction. In: International Conference on Machine Learning. PMLR, pp 36868–36886

BibTeX

@InProceedings{pmlr-v202-wei23e,
  title = 	 {Mitigating Memorization of Noisy Labels by Clipping the Model Prediction},
  author =       {Wei, Hongxin and Zhuang, Huiping and Xie, Renchunzi and Feng, Lei and Niu, Gang and An, Bo and Li, Yixuan},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {36868--36886},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/wei23e/wei23e.pdf},
  url = 	 {https://proceedings.mlr.press/v202/wei23e.html},
  abstract = 	 {In the presence of noisy labels, designing robust loss functions is critical for securing the generalization performance of deep neural networks. Cross Entropy (CE) loss has been shown to be not robust to noisy labels due to its unboundedness. To alleviate this issue, existing works typically design specialized robust losses with the symmetric condition, which usually lead to the underfitting issue. In this paper, our key idea is to induce a loss bound at the logit level, thus universally enhancing the noise robustness of existing losses. Specifically, we propose logit clipping (LogitClip), which clamps the norm of the logit vector to ensure that it is upper bounded by a constant. In this manner, CE loss equipped with our LogitClip method is effectively bounded, mitigating the overfitting to examples with noisy labels. Moreover, we present theoretical analyses to certify the noise-tolerant ability of LogitClip. Extensive experiments show that LogitClip not only significantly improves the noise robustness of CE loss, but also broadly enhances the generalization performance of popular robust losses.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/