Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach

Giorgio PatriniAlessandro RozzaAditya MenonRichard NockLizhen Qu

article2016CVPR1,730 citations

Presents an architecture-agnostic loss correction method that enables deep neural networks to accurately learn from corrupted training data by estimating and mathematically inverting the class-noise transition matrix.

Listen

Training modern deep neural networks requires massive volumes of data, but high-quality expert annotation is expensive and slow. Organizations increasingly rely on cost-effective alternatives such as crowdsourcing and automated web scraping, which inevitably introduce class-dependent label noise (systematic mislabeling between confusing or related categories). Without remediation, this corrupted supervision severely degrades model accuracy and reliability across critical applications.

The article establishes a theoretically grounded, architecture-agnostic framework to train robust deep neural networks under class-conditional label noise, presenting an end-to-end pipeline that corrects training loss functions and estimates mislabeling rates directly from noisy data.

The authors develop two distinct loss-adjustment procedures: a "backward" correction that inversely weights the loss by a noise transition matrix summarizing class-flipping probabilities, and a "forward" correction that multiplies the model predictions by the transition matrix. To eliminate the requirement of knowing noise rates beforehand, they extend a class-probability estimation technique to estimate transition matrices without ground-truth labels. The methodology was validated across diverse architectures—including dense, convolutional, recurrent (LSTM), and residual networks—spanning benchmark datasets (MNIST, IMDB, CIFAR-10, CIFAR-100) and a real-world dataset of one million noisy clothing images (Clothing1M).

The evaluation yielded several key findings. First, under severe asymmetric label noise (up to 60%), standard cross-entropy loss degraded drastically by 30 to over 45 percentage points, whereas the proposed forward correction maintained near-clean performance. Second, forward loss correction consistently outperformed backward correction in practical optimization, avoiding numerical instability associated with matrix inversions. Third, the fully automated noise estimator recovered accurate transition matrices directly from noisy samples, suffering a median accuracy drop of only 10 percentage points compared to having perfect prior knowledge of noise rates. Fourth, combining forward loss correction with fine-tuning on a 50-layer residual network achieved 80.38% classification accuracy on Clothing1M, outperforming previous approaches by over two percentage points without requiring complex iterative sampling procedures.

These results provide immediate operational benefits for deploying deep learning in high-scale, cost-sensitive environments. Organizations can significantly lower data curation costs and project timelines by safely training models on weakly labeled, crowdsourced, or web-harvested data without bespoke architecture redesigns. The authors also establish a theoretical guarantee that networks utilizing rectified linear unit (ReLU) activations retain an invariant loss curvature (Hessian) under noise, ensuring stable optimization convergence.

Engineering teams facing noisy datasets should adopt the forward loss correction method as a plug-and-play modification to standard cross-entropy objectives. When transition rates are unknown, teams should employ the automated two-stage estimator using large uncurated datasets to infer noise structures before final model training. Future development should focus on enhancing estimation algorithms with structural priors—such as low-rank matrices—to maintain robustness in heavily fine-grained scenarios with many classes (e.g., CIFAR-100), as well as extending the framework to handle instance-dependent, input-specific noise.

  • Paper: Online Passive-Aggressive Algorithms, Koby Crammer et al. (2003). Learn the foundational principles of margin-based online updates and noise-tolerant loss adjustments in linear and margin models.
  • Paper: A Simple Weight Decay Can Improve Generalization, Anders Krogh et al. (1991). Understand the classical mathematical analysis of how parameter weight decay curbs overfitting to noisy targets during neural network training.
Cover for Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach

Abstract

We present a theoretically grounded approach to train deep neural networks, including recurrent networks, subject to class-dependent label noise. We propose two procedures for loss correction that are agnostic to both application domain and network architecture. They simply amount to at most a matrix inversion and multiplication, provided that we know the probability of each class being corrupted into another. We further show how one can estimate these probabilities, adapting a recent technique for noise estimation to the multi-class setting, and thus providing an end-to-end framework. Extensive experiments on MNIST, IMDB, CIFAR-10, CIFAR-100 and a large scale dataset of clothing images employing a diversity of architectures --- stacking dense, convolutional, pooling, dropout, batch normalization, word embedding, LSTM and residual layers --- demonstrate the noise robustness of our proposals. Incidentally, we also prove that, when ReLU is the only non-linearity, the loss curvature is immune to class-dependent label noise.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Preliminaries
  • 4 Label noise and loss robustness
  • 4.1 The backward correction procedure
  • 4.2 The forward correction procedure
  • 4.3 The overall algorithm
  • 4.4 Digression: noise free Hessians via ReLU
  • 5 Experiments
  • 5.1 Loss corrections with TT known or estimated
  • 5.2 Comparing with other loss functions
  • 5.3 Experiments on Clothing1M
  • 6 Discussion and Conclusion
  • References

Knowls

  1. Knowl 1 — Backward Loss Correction for Multi-Class Classification under Label Noise

    theoretical result

    In a cc-class classification problem with feature space X⊆Rd\mathcal{X} \subseteq \mathbb{R}^d and label space Y={e1,…,ec}\mathcal{Y} = \{e_1, \dots, e_c\} (where ei∈{0,1}ce_i \in \{0,1\}^c denotes the ii-th canonical basis vector), let T∈[0,1]c×cT \in [0, 1]^{c \times c} be a non-singular, row-stochastic noise transition matrix whose entries specify class-conditional corruption probabilities:

    Tij=p(y~=ej∣y=ei)T_{ij} = p(\tilde{y} = e_j \mid y = e_i)

    where y∈Yy \in \mathcal{Y} represents the true unobserved label and y~∈Y\tilde{y} \in \mathcal{Y} represents the corrupted observed label. Let p^(y∣x)∈Δc−1\hat{p}(y \mid x) \in \Delta^{c-1} denote the model's predicted class-probability vector (e.g., the output of a softmax layer), and let ℓ(p^(y∣x))=[ℓ(e1,p^(y∣x)),…,ℓ(ec,p^(y∣x))]⊤∈Rc\ell(\hat{p}(y \mid x)) = [\ell(e_1, \hat{p}(y \mid x)), \dots, \ell(e_c, \hat{p}(y \mid x))]^\top \in \mathbb{R}^c be the loss vector evaluated across all cc classes.

    The backward corrected loss is defined by linearly transforming the loss vector using the inverse of the noise matrix:

    ℓ←(p^(y∣x))=T−1ℓ(p^(y∣x))\ell^{\leftarrow}(\hat{p}(y \mid x)) = T^{-1} \ell(\hat{p}(y \mid x))

    For any input xx, the backward corrected loss evaluated on the noisy label y~\tilde{y} is an unbiased estimator of the original loss ℓ\ell evaluated on the clean label yy:

    ∀x,Ey~∣x[ℓ←(y~,p^(y∣x))]=Ey∣x[ℓ(y,p^(y∣x))]\forall x, \quad \mathbb{E}_{\tilde{y} \mid x} [\ell^{\leftarrow}(\tilde{y}, \hat{p}(y \mid x))] = \mathbb{E}_{y \mid x} [\ell(y, \hat{p}(y \mid x))]

    Consequently, the expected risk minimizer over the noisy distribution coincides with the expected risk minimizer over the clean data distribution:

    arg⁡min⁡p^(y∣x)Ex,y~[ℓ←(y~,p^(y∣x))]=arg⁡min⁡p^(y∣x)Ex,y[ℓ(y,p^(y∣x))]\arg\min_{\hat{p}(y \mid x)} \mathbb{E}_{x,\tilde{y}} [\ell^{\leftarrow}(\tilde{y}, \hat{p}(y \mid x))] = \arg\min_{\hat{p}(y \mid x)} \mathbb{E}_{x,y} [\ell(y, \hat{p}(y \mid x))]

  2. Knowl 2 — Forward Loss Correction for Proper Composite Losses

    theoretical result

    Let h:X→Rch: \mathcal{X} \to \mathbb{R}^c be a neural network classifier and let ℓψ:Y×Rc→R\ell_\psi: \mathcal{Y} \times \mathbb{R}^c \to \mathbb{R} be a proper composite loss function defined via an invertible link function ψ:Δc−1→Rc\psi: \Delta^{c-1} \to \mathbb{R}^c such that ℓψ(y,h(x))=ℓ(y,ψ−1(h(x)))\ell_\psi(y, h(x)) = \ell(y, \psi^{-1}(h(x))) (for multi-class cross-entropy, ψ−1\psi^{-1} is the softmax function). Let T∈[0,1]c×cT \in [0, 1]^{c \times c} be a non-singular row-stochastic noise transition matrix with Tij=p(y~=ej∣y=ei)T_{ij} = p(\tilde{y} = e_j \mid y = e_i).

    The forward loss correction modifies the prediction before loss evaluation by multiplying the predicted clean distribution ψ−1(h(x))\psi^{-1}(h(x)) by T⊤T^\top:

    ℓψ→(h(x))=ℓ(T⊤ψ−1(h(x)))\ell^{\rightarrow}_\psi(h(x)) = \ell\left(T^\top \psi^{-1}(h(x))\right)

    For the cross-entropy loss, the forward corrected loss given an observed noisy label y~=ei\tilde{y} = e_i becomes:

    ℓ→(ei,h(x))=−log⁡(∑j=1cTjip^(y=ej∣x))\ell^{\rightarrow}(e_i, h(x)) = -\log \left( \sum_{j=1}^c T_{ji} \hat{p}(y = e_j \mid x) \right)

    where p^(y=ej∣x)=[ψ−1(h(x))]j\hat{p}(y = e_j \mid x) = [\psi^{-1}(h(x))]_j.

    Unlike the backward correction, the forward correction does not require matrix inversion. The composite loss ℓψ→\ell^{\rightarrow}_\psi corresponds to a proper composite loss with the modified denoising link function ϕ=(T−1)⊤∘ψ\phi = (T^{-1})^\top \circ \psi. The minimizer of the forward corrected loss under the corrupted distribution equals the minimizer of the original loss under the clean distribution:

    arg⁡min⁡hEx,y~[ℓψ→(y~,h(x))]=arg⁡min⁡hEx,y[ℓψ(y,h(x))]\arg\min_h \mathbb{E}_{x,\tilde{y}} \left[\ell^{\rightarrow}_\psi(\tilde{y}, h(x))\right] = \arg\min_h \mathbb{E}_{x,y} \left[\ell_\psi(y, h(x))\right]

  3. Knowl 3 — Multi-Class Noise Transition Matrix Estimation via Anchor Point Posteriors

    theoretical result

    Let p(x,y)p(x, y) be the clean data distribution over X×{e1,…,ec}\mathcal{X} \times \{e_1, \dots, e_c\} and let T∈[0,1]c×cT \in [0, 1]^{c \times c} be the noise transition matrix Tij=p(y~=ej∣y=ei)T_{ij} = p(\tilde{y} = e_j \mid y = e_i). Assume:

    1. For every class j∈[c]j \in [c], there exists at least one perfect instance (anchor point) xˉj∈X\bar{x}^j \in \mathcal{X} such that p(xˉj)>0p(\bar{x}^j) > 0 and p(y=ej∣xˉj)=1p(y = e_j \mid \bar{x}^j) = 1.
    2. Given sufficient training examples, the hypothesis class hh is expressive enough to accurately model the noisy posterior p(y~∣x)p(\tilde{y} \mid x).

    Under these conditions, each transition probability TijT_{ij} is uniquely determined by the noisy conditional distribution evaluated at the anchor point xˉi\bar{x}^i:

    ∀i,j∈[c],Tij=p(y~=ej∣xˉi)\forall i, j \in [c], \quad T_{ij} = p(\tilde{y} = e_j \mid \bar{x}^i)

    In practice, given a model p^(y~∣x)\hat{p}(\tilde{y} \mid x) trained on noisy data and an unlabeled sample set X′X' (which can be the training or validation features), the anchor point xˉi\bar{x}^i and transition matrix estimate T^\hat{T} are computed as:

    xˉi=arg⁡max⁡x∈X′p^(y~=ei∣x)\bar{x}^i = \arg\max_{x \in X'} \hat{p}(\tilde{y} = e_i \mid x)

    T^ij=p^(y~=ej∣xˉi)\hat{T}_{ij} = \hat{p}(\tilde{y} = e_j \mid \bar{x}^i)

    followed by row-normalization to ensure ∑j=1cT^ij=1\sum_{j=1}^c \hat{T}_{ij} = 1 for each row ii.

  4. Knowl 4 — Two-Stage Robust Training Algorithm with Loss Correction

    algorithm

    When the noise transition matrix TT is not known a priori, training proceeds in two distinct stages to decouple noise rate estimation from classifier learning.

    Input: Noisy training dataset S={(xk,y~k)}k=1mS = \{(x_k, \tilde{y}_k)\}_{k=1}^m, base loss ℓ\ell, unlabeled sample X′X'
    Output: Robust model parameters for network h(⋅)h(\cdot)
    if TT is not provided then
        Train a neural network hnoisy(x)h_{\text{noisy}}(x) on SS by minimizing standard loss ℓ\ell
        for each class i∈{1,…,c}i \in \{1, \dots, c\} do
            Identify anchor point xˉi=arg⁡max⁡x∈X′p^noisy(y~=ei∣x)\bar{x}^i = \arg\max_{x \in X'} \hat{p}_{\text{noisy}}(\tilde{y} = e_i \mid x)
            for each class j∈{1,…,c}j \in \{1, \dots, c\} do
                T^ij←p^noisy(y~=ej∣xˉi)\hat{T}_{ij} \leftarrow \hat{p}_{\text{noisy}}(\tilde{y} = e_j \mid \bar{x}^i)
            end for
            Normalize row: T^i⋅←T^i⋅/∑j=1cT^ij\hat{T}_{i\cdot} \leftarrow \hat{T}_{i\cdot} / \sum_{j=1}^c \hat{T}_{ij}
        end for
        T←T^T \leftarrow \hat{T}
    end if
    Train network h(x)h(x) on SS using backward corrected loss ℓ←\ell^{\leftarrow} or forward corrected loss ℓ→\ell^{\rightarrow} computed with TT
    return h(⋅)h(\cdot)

    The unlabeled set X′X' can be the feature set of SS combined with validation data. The computational complexity of computing T^\hat{T} from the first-stage network outputs is O(c2⋅∣X′∣)\mathcal{O}(c^2 \cdot |X'|).

  5. Knowl 5 — Invariance of Loss Hessian to Label Noise in ReLU Networks

    theoretical result

    Let h:X→Rch: \mathcal{X} \to \mathbb{R}^c be an nn-layer neural network where all hidden activation functions are rectified linear units (ReLUs, σ(z)=max⁡(0,z)\sigma(z) = \max(0, z)) or continuous piecewise linear functions, and the final layer is a linear projection h(x)=W(n)x(n−1)+b(n)h(x) = W^{(n)} x^{(n-1)} + b^{(n)}. Let ℓ\ell be a linear-odd loss function, such as the multi-class cross-entropy loss or square loss.

    Then:

    1. The Hessian of the expected loss with respect to all network parameters under label noise, ∇2Ex,y~[ℓ(y~,h(x))]\nabla^2 \mathbb{E}_{x, \tilde{y}}[\ell(\tilde{y}, h(x))], is identical to the Hessian under the clean data distribution, ∇2Ex,y[ℓ(y,h(x))]\nabla^2 \mathbb{E}_{x, y}[\ell(y, h(x))].
    2. The Hessian of the backward corrected loss ℓ←\ell^\leftarrow is equal to the Hessian of the uncorrected loss ℓ\ell for any transition matrix TT.

    Because the true class label y=eiy = e_i only appears in the linear components −Wi⋅(n)x(n−1)−bi(n)-W^{(n)}_{i\cdot} x^{(n-1)} - b^{(n)}_i while the non-linear log-partition function log⁡∑k=1cexp⁡(Wk⋅(n)x(n−1)+bk(n))\log \sum_{k=1}^c \exp(W^{(n)}_{k\cdot} x^{(n-1)} + b^{(n)}_k) is class-independent, the second derivatives of the loss with respect to parameters depend only on the log-partition term. Thus, label noise and backward matrix correction alter first-order gradient components but leave the second-order curvature of the loss surface unchanged point by point.

  6. Knowl 6 — Practical Regularization and Percentile Modifications for Noise Correction

    model/method

    In practical implementations of loss correction and noise estimation, two key numerical adjustments are applied:

    1. Robust Anchor Selection via Percentiles: To avoid selecting extreme outliers when estimating the noise transition matrix TT, the argmax anchor point selection xˉi=arg⁡max⁡x∈X′p^(y~=ei∣x)\bar{x}^i = \arg\max_{x \in X'} \hat{p}(\tilde{y} = e_i \mid x) is replaced by taking the α\alpha-percentile of the predicted probabilities p^(y~=ei∣x)\hat{p}(\tilde{y} = e_i \mid x) across X′X', with α=97%\alpha = 97\% functioning robustly across datasets with moderate-to-large samples per class.

    2. Condition Number Regularization for Backward Correction: While a row-stochastic matrix TT is invertible almost surely, its condition number can be poor, leading to large positive and negative loss weights in T−1T^{-1}. To stabilize matrix inversion, TT can be convexly regularized with the identity matrix II prior to inversion:

    Treg=(1−γ)T+γIT_{\text{reg}} = (1 - \gamma) T + \gamma I

    for γ∈(0,1)\gamma \in (0, 1), acting as a conservative noise-free prior.

  7. Knowl 7 — Empirical Classification Accuracy across Noise Benchmarks

    data/table

    Evaluation of loss correction methods compared to baseline surrogate losses across MNIST (fully connected network), CIFAR-10 (14-layer and 32-layer ResNets), CIFAR-100 (44-layer ResNet), and IMDB sentiment analysis (word embeddings and LSTM) under symmetric (SYMM.) and asymmetric (ASYMM.) label noise rates NN. Across datasets, forward correction consistently outperforms backward correction and prior robust losses, and loss estimation T^\hat{T} preserves most of the gain.

    Dataset No Noise Symm. N=0.2N=0.2 Asymm. N=0.2N=0.2 Asymm. N=0.6N=0.6
    MNIST (FC)
    Cross-entropy 97.9±0.097.9 \pm 0.0 96.9±0.196.9 \pm 0.1 97.5±0.097.5 \pm 0.0 53.0±0.653.0 \pm 0.6
    Unhinged (BN) 97.6±0.097.6 \pm 0.0 96.9±0.196.9 \pm 0.1 97.0±0.197.0 \pm 0.1 71.2±1.071.2 \pm 1.0
    Sigmoid (BN) 97.2±0.197.2 \pm 0.1 93.1±0.293.1 \pm 0.2 96.7±0.196.7 \pm 0.1 71.4±1.371.4 \pm 1.3
    Savage 97.3±0.097.3 \pm 0.0 96.9±0.096.9 \pm 0.0 97.0±0.197.0 \pm 0.1 51.3±0.451.3 \pm 0.4
    Bootstrap Soft 97.9±0.097.9 \pm 0.0 96.9±0.096.9 \pm 0.0 97.5±0.097.5 \pm 0.0 53.0±0.453.0 \pm 0.4
    Backward (TT) 97.9±0.097.9 \pm 0.0 96.3±0.196.3 \pm 0.1 96.6±1.196.6 \pm 1.1 93.0±0.993.0 \pm 0.9
    Backward (T^\hat{T}) 97.9±0.097.9 \pm 0.0 96.9±0.096.9 \pm 0.0 96.7±0.196.7 \pm 0.1 67.4±1.567.4 \pm 1.5
    Forward (TT) 97.9±0.097.9 \pm 0.0 97.3±0.097.3 \pm 0.0 97.7±0.097.7 \pm 0.0 97.3±0.097.3 \pm 0.0
    Forward (T^\hat{T}) 97.9±0.097.9 \pm 0.0 96.9±0.096.9 \pm 0.0 97.7±0.097.7 \pm 0.0 64.9±4.464.9 \pm 4.4
    CIFAR-10 (ResNet-32)
    Cross-entropy 90.1 86.6 89.0 53.6
    Backward (TT) 90.1 83.0 84.4 74.3
    Backward (T^\hat{T}) 90.8 86.9 86.4 66.7
    Forward (TT) 91.2 87.7 89.9 87.6
    Forward (T^\hat{T}) 90.5 87.9 90.1 77.6
    CIFAR-100 (ResNet-44)
    Cross-entropy 68.5 57.9 63.5 17.1
    Backward (TT) 68.5 55.1 53.8 36.8
    Forward (TT) 68.8 64.0 68.1 68.4
    Forward (T^\hat{T}) 68.1 58.6 64.2 15.9
  8. Knowl 8 — Clothing1M State-of-the-Art Results Using Loss Correction

    data/table

    The Clothing1M dataset contains 1 million clothing images with noisy labels across 14 categories, along with clean subsets of 50k (training), 14k (validation), and 10k (test). A 50-layer ResNet pre-trained on ImageNet was trained using forward and backward loss corrections (using a transition matrix TT estimated from the paired clean/noisy labels of the 50k set) and subsequently fine-tuned on the clean 50k data, establishing a new state of the art of 80.38% test accuracy without bootstrapping.

    # Model Loss Initialization Training Data Accuracy (%)
    1 AlexNet Cross-entropy ImageNet 50k 72.63
    2 AlexNet (Sukhbaatar et al.) Cross-entropy #1 1M, 50k (bootstrapped) 76.22
    3 AlexNet (Xiao et al.) Cross-entropy #1 1M, 50k (bootstrapped) 78.24
    4 50-ResNet Cross-entropy ImageNet 1M 68.94
    5 50-ResNet Backward ImageNet 1M 69.13
    6 50-ResNet Forward ImageNet 1M 69.84
    7 50-ResNet Cross-entropy ImageNet 50k 75.19
    8 50-ResNet Cross-entropy #6 (Forward 1M) 50k 80.38

Coverage note — None was omitted; all core theoretical results (Theorems 1-4), algorithms, numerical stabilization details, and empirical benchmarks on synthetic noise and Clothing1M are fully covered.

References

  1. 1.F. Chollet. Keras. github.com/fchollet/keras.
  2. 2.A. M. Dai and Q. V. Le. Semi-supervised sequence learning. In NIPS*29, 2015.
  3. 3.S. Divvala, A. Farhadi, and C. Guestrin. Learning everything about anything: Webly-supervised visual concept learning. In 27th IEEE CVPR, 2014.
  4. 4.J. Duchi, E. Hazan, and Y. Singer. Adaptive subgradient methods for online learning and stochastic optimization. JMLR, 12:2121–2159, 2011.
  5. 5.R. Fergus, L. Fei-Fei, P. Perona, and A. Zisserman. Learning object categories from internet image searches. Proceedings of the IEEE, 98(8):1453–1466, 2010.
  6. 6.M. Feurer, A. Klein, K. Eggensperger, J. Springenberg, M. Blum, and F. Hutter. Efficient and robust automated machine learning. In NIPS*29, 2015.
  7. 7.B. Frénay and M. Verleysen. Classification in the Presence of Label Noise: A Survey. IEEE Transactions on Neural Networks and Learning Systems, 25(5):845–869, May 2014.
  8. 8.A. Ghosh, N. Manwani, and P. S. Sastry. Making risk minimization tolerant to label noise. Neurocomputing, 2015.
  9. 9.X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. In AISTATS, 2010.
  10. 10.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015.
  11. 11.K. He, X. Zhang., S. Ren, and J. Sun. Deep residual learning for image recognition. In 29th IEEE CVPR, 2016.
  12. 12.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In 14 th ECCV, 2016.
  13. 13.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  14. 14.G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger. Deep networks with stochastic depth. In 14 th ECCV, 2016.
  15. 15.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In 32 th ICML, 2015.
  16. 16.K. Kawaguchi. Deep learning without poor local minima. In NIPS*30, 2016.
  17. 17.J. Krause, B. Sapp, A. Howard, H. Zhou, A. Toshev, T. Duerig, J. Philbin, and L. Fei-Fei. The unreasonable effectiveness of noisy data for fine-grained recognition. In 14 th ECCV, 2016.
  18. 18.A. Krizhevsky and G. Hinton. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  19. 19.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS*26, 2012.
  20. 20.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  21. 21.T. Liu and D. Tao. Classification with noisy labels by importance reweighting. IEEE Transactions on PAMI, 38(3):447–461, 2016.
  22. 22.P. M. Long and R. A. Servedio. Random classification noise defeats all convex potential boosters. Machine learning, 78(3):287–304, 2010.
  23. 23.A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts. Learning word vectors for sentiment analysis. In 49 th ACL, 2011.
  24. 24.H. Masnadi-Shirazi and N. Vasconcelos. On the design of loss functions for classification: theory, robustness to outliers, and savageboost. In NIPS*23, 2009.
  25. 25.A. Menon, B. van Rooyen, and N. Natarajan. Learning from binary labels with instance-dependent corruption. arXiv preprint arXiv:1605.00751, 2016.
  26. 26.A. Menon, B. van Rooyen, C. S. Ong, and B. Williamson. Learning from corrupted binary labels via class-probability estimation. In 32 th ICML, 2015.
  27. 27.V. Mnih and G. E. Hinton. Learning to label aerial images from noisy data. In 29 th ICML, 2012.
  28. 28.N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari. Learning with noisy labels. In NIPS*27, 2013.
  29. 29.L. Niu, W. Li, and D. Xu. Visual recognition by learning from web data: A weakly supervised domain generalization approach. In 28th IEEE CVPR, 2015.
  30. 30.G. Patrini, F. Nielsen, R. Nock, and M. Carioni. Loss factorization, weakly supervised learning and label noise robustness. In 33 th ICML, 2016.
  31. 31.H. G. Ramaswamy, C. Scott, and A. Tewari. Mixture proportion estimation via kernel embedding of distributions. In 33 th ICML, 2016.
  32. 32.S. Reed, H. Lee, D. Anguelov, C. Szegedy, D. Erhan, and A. Rabinovich. Training deep neural networks on noisy labels with bootstrapping. arXiv preprint arXiv:1412.6596, 2014.
  33. 33.M. D. Reid and R. C. Williamson. Composite binary losses. JMLR, 11:2387–2422, 2010.
  34. 34.T. Sanderson and C. C. Scott. Class proportion estimation with application to multiclass anomaly rejection. In AISTATS, 2014.
  35. 35.F. Schroff, A. Criminisi, and A. Zisserman. Harvesting image databases from the web. IEEE Transactions on PAMI, 33(4):754–766, 2011.
  36. 36.C. Scott, G. Blanchard, and G. Handy. Classification with asymmetric label noise : Consistency and maximal denoising. In 26 rd COLT, 2013.
  37. 37.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. JMLR, 15(1):1929–1958, 2014.
  38. 38.G. Stempfel and L. Ralaivola. Learning SVMs from sloppily labeled data. In Artificial Neural Networks (ICANN), pages 884–893. Springer, 2009.
  39. 39.S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus. Training convolutional networks with noisy labels. In ICLR Workshops, 2015.
  40. 40.B. van Rooyen. Machine Learning via Transitions. PhD thesis, The Australian National University, 2015.
  41. 41.B. van Rooyen, A. K. Menon, and R. C. Williamson. Learning with symmetric label noise: The importance of being unhinged. In NIPS*29, 2015.
  42. 42.T. Xiao, T. Xia, T. Yang, C. Huang, and X. Wang. Learning from massive noisy labeled data for image classification. In 28th IEEE CVPR, 2015.

Citation

MLA
Patrini, G., et al. “Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach”. arXiv, 2016, http://arxiv.org/abs/1609.03683v2.
APA
Patrini, G., Rozza, A., Menon, A., Nock, R., & Qu, L. (2016). Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach. arXiv. http://arxiv.org/abs/1609.03683v2
Chicago
Patrini, G., A. Rozza, A. Menon, R. Nock, and L. Qu. 2016. “Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach”. arXiv. http://arxiv.org/abs/1609.03683v2.
Harvard
Patrini, G. et al. (2016) “Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1609.03683v2.
Vancouver
1. Patrini G, Rozza A, Menon A, Nock R, Qu L (2016) Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach. arXiv

BibTeX

@article{patrini2016making,
  title = {Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach},
  author = {Patrini, Giorgio and Rozza, Alessandro and Menon, Aditya and Nock, Richard and Qu, Lizhen},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1609.03683v2},
  eprint = {1609.03683}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE