Robust Training under Label Noise by Over-parameterization

Sheng LiuZhihui ZhuQing QuChong You

article2022ICML150 citations

Proposes a sparse over-parameterization framework that exploits implicit algorithmic regularization to isolate label noise from clean data during training, achieving state-of-the-art classification accuracy on corrupted datasets with theoretical guarantees for exact noise separation.

Listen

Modern artificial intelligence heavily relies on large, over-parameterized neural networks that contain far more parameters than training samples. While these architectures deliver superior performance across computer vision and language tasks, their high capacity causes them to memorize errors when training data contains incorrect labels. Because manual data annotation is expensive and inherently error-prone, real-world datasets frequently suffer from corrupted labels, which severely degrades model generalization on unseen data.

The article introduces and evaluates Sparse Over-Parameterization (SOP), a principled training method designed to prevent over-parameterized networks from overfitting to noisy labels in classification tasks. The core objective is to demonstrate that explicitly modeling label corruptions through auxiliary over-parameterized variables can separate sparse noise from true data patterns, both in practical deep learning systems and under formal mathematical models.

To achieve this, the authors model label errors using an extra set of parameters that represent the discrepancy between observed labels and true underlying classes. By initializing these parameters at very small values and training the entire system via standard gradient descent with distinct learning rates, the optimization dynamics naturally induce a sparse penalty that isolates erroneous labels. The credibility of the approach was tested through extensive image classification experiments on benchmark datasets with simulated label noise (CIFAR-10 and CIFAR-100) and real-world annotation noise (CIFAR-N, Clothing-1M, and WebVision), alongside theoretical analysis and numerical verification on linearized mathematical models.

Key findings show that the proposed method consistently achieves state-of-the-art test accuracy across diverse noise conditions. On benchmarks with synthetic label corruptions ranging up to 80%, the enhanced variant (SOP+) achieved top-tier performance, reaching 94.0% accuracy on CIFAR-10 with 80% noise and 78.0% on CIFAR-100 with 40% asymmetric noise. On realistic human-annotated noise (CIFAR-10N worst-case noise), the method attained 93.24% accuracy, outperforming leading existing techniques. Additionally, the approach demonstrated superior computational efficiency, training in 1.0 to 2.1 hours on benchmark datasets compared to 2.3 to 5.4 hours for competing methods. Theoretically, the authors proved that gradient descent on linearized over-parameterized models exacts full separation between the underlying ground truth and sparse corruptions under low-rank data conditions.

These findings indicate that organizations can train high-capacity deep learning models directly on imperfect, real-world data without expensive label-cleaning pipelines or complex multi-network training setups. The method significantly reduces the computational overhead and risk of model degradation caused by flawed annotations, offering a mathematically grounded alternative to heuristic filtering techniques.

For technical teams managing data annotation challenges, the primary recommendation is to integrate the sparse over-parameterization framework into existing classification training pipelines, using standard gradient descent optimizers and minimal initialization for noise variables. Future engineering work should explore incorporating structured or group-sparse priors to account for known class-confusion patterns in specific domain applications.

While empirical results are strong across standard image benchmarks, the theoretical exact recovery guarantees currently rely on linearized models and low-rank data assumptions. Decision-makers should validate the technique in specific production environments, particularly where noise rates exceed extreme thresholds or where label corruption patterns deviate significantly from standard sparsity assumptions.

Cover for Robust Training under Label Noise by Over-parameterization

Abstract

Recently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning. However, when the training data is corrupted, it has been well-known that over-parameterized networks tend to overfit and do not generalize. In this work, we propose a principled approach for robust training of over-parameterized deep networks in classification tasks where a proportion of training labels are corrupted. The main idea is yet very simple: label noise is sparse and incoherent with the network learned from clean data, so we model the noise and learn to separate it from the data. Specifically, we model the label noise via another sparse over-parameterization term, and exploit implicit algorithmic regularizations to recover and separate the underlying corruptions. Remarkably, when trained using such a simple method in practice, we demonstrate state-of-the-art test accuracy against label noise on a variety of real datasets. Furthermore, our experimental results are corroborated by theory on simplified linear models, showing that exact separation between sparse noise and low-rank data can be achieved under incoherent conditions. The work opens many interesting directions for improving over-parameterized models by using sparse over-parameterization and implicit regularization. Code is available at https://github.com/shengliu66/SOP.

Table of Contents

  • 1. Introduction
  • 2. Robust Classification with Label Noise
  • 2.1. Implementation Details of SOP
  • 2.2. Experiments
  • 3. Theoretical Insights with Simplified Models
  • 3.1. Problem Setup & Main Result
  • 3.2. Landscapes & Implicit Sparse Regularization
  • 3.3. Exact Recovery under Incoherence Conditions
  • 4. Related Work and Discussion
  • 4.1. Prior Arts on Implicit Regularization
  • 4.2. Relationship to Existing Work on Label Noise
  • 4.3. Sparsity in Deep Learning
  • 4.4. Limitations and Future Directions
  • Acknowledgements
  • References
  • Appendices
  • A. Training Details for Robust Classification with Label Noise
  • A.1. Choice of Loss Function
  • A.2. Definition of Label Noise
  • A.3. Implementation Details of SOP+
  • A.4. Experimental Settings
  • B. Proofs for Theoretical Analysis with Linear Models
  • B.1. Proof of Proposition 3.2
  • B.2. Proof of Proposition 3.3
  • B.3. Proof of Proposition 3.5
  • Proof of Lemma B.5.

Knowls

  1. Knowl 1 — Sparse Over-Parameterization Formulation for Label Noise Modeling

    model/method

    In classification problems with noisy training labels, let (xi,yi)i=1N(x_i, y_i)_{i=1}^N denote the training set of NN samples, where xix_i is the input image and yiightin{0,1}Ky_i ightin \{0, 1\}^K is the observed one-hot label vector for KK classes. To prevent an over-parameterized neural network f(⋅;θ)f(\cdot; \theta) from memorizing corruptions, the label corruption on sample ii is explicitly modeled by an additive sparse vector si∈RKs_i \in \mathbb{R}^K.

    Because sis_i represents the difference between the observed label yiy_i and the true latent label, corruptions on the true class correspond to negative values on the non-zero entry of yiy_i, while corruptions assigning incorrect classes correspond to positive entries on the zero entries of yiy_i. Furthermore, all entries of sis_i must lie in [−1,1][-1, 1]. To enforce sparsity and incorporate sign constraints via over-parameterization, sis_i is parameterized with two vectors ui,vi∈[−1,1]Ku_i, v_i \in [-1, 1]^K:

    si=ui⊙ui⊙yi−vi⊙vi⊙(1−yi)s_i = u_i \odot u_i \odot y_i - v_i \odot v_i \odot (1 - y_i)

    where ⊙\odot denotes the entry-wise Hadamard product. The model parameters θ\theta and the per-sample over-parameterized noise variables {ui,vi}i=1N\{u_i, v_i\}_{i=1}^N are trained by solving:

    min⁡θ,{ui,vi}i=1N1N∑i=1Nℓ(f(xi;θ)+si,yi)s.t.ui,vi∈[−1,1]K\min_{\theta, \{u_i, v_i\}_{i=1}^N} \frac{1}{N} \sum_{i=1}^N \ell\left(f(x_i; \theta) + s_i, y_i\right) \quad \text{s.t.} \quad u_i, v_i \in [-1, 1]^K

    where ℓ(⋅,⋅)\ell(\cdot, \cdot) is a suitable classification loss function.

  2. Knowl 2 — Loss Function Selection and Split Gradient Updates in Sparse Over-Parameterization

    model/method

    When optimizing the Sparse Over-Parameterization (SOP) objective with si=ui⊙ui⊙yi−vi⊙vi⊙(1−yi)s_i = u_i \odot u_i \odot y_i - v_i \odot v_i \odot (1 - y_i), different loss functions are required for updating uiu_i and viv_i.

    For updating θ\theta and uiu_i, cross-entropy loss ℓCE\ell_{\mathrm{CE}} is applied after mapping the unnormalized score vector w=f(xi;θ)+siw = f(x_i; \theta) + s_i into a valid probability distribution using the projection ϕ(w)\phi(w):

    ϕ(w)=max⁡{w,ϵ1}∥max⁡{w,ϵ1∥1\phi(w) = \frac{\max\{w, \epsilon \mathbf{1}\}}{\|\max\{w, \epsilon \mathbf{1}\|_1}

    LCE(θ,ui,vi;xi,yi)=ℓCE(ϕ(f(xi;θ)+si),yi)L_{\mathrm{CE}}(\theta, u_i, v_i; x_i, y_i) = \ell_{\mathrm{CE}}\left(\phi(f(x_i; \theta) + s_i), y_i\right)

    where ϵ>0\epsilon > 0 is a small constant preventing zero division and 1\mathbf{1} is the all-ones vector.

    Cross-entropy loss cannot be used for updating viv_i because the gradient of LCEL_{\mathrm{CE}} with respect to viv_i satisfies:

    ∂LCE∂vi=2vi⊙(1−yi)1⊤(f(xi;θ)+si)\frac{\partial L_{\mathrm{CE}}}{\partial v_i} = \frac{2 v_i \odot (1 - y_i)}{\mathbf{1}^\top (f(x_i; \theta) + s_i)}

    which does not depend on the individual class-level model outputs f(xi;θ)f(x_i; \theta) other than through the global sum in the denominator. Consequently, viv_i cannot learn which specific false label entry was corrupted. To resolve this, viv_i is updated using the mean squared error (MSE) loss:

    LMSE(θ,ui,vi;xi,yi)=ℓMSE(f(xi;θ)+si,yi)L_{\mathrm{MSE}}(\theta, u_i, v_i; x_i, y_i) = \ell_{\mathrm{MSE}}\left(f(x_i; \theta) + s_i, y_i\right)

    whose gradient directly depends on individual entry discrepancies:

    ∂LMSE∂vi=4(f(xi;θ)+si−yi)⊙vi⊙(1−yi)\frac{\partial L_{\mathrm{MSE}}}{\partial v_i} = 4\left(f(x_i; \theta) + s_i - y_i\right) \odot v_i \odot (1 - y_i)

  3. Knowl 3 — Sparse Over-Parameterization Algorithm for Image Classification

    algorithm

    The Sparse Over-Parameterization (SOP) algorithm trains an over-parameterized neural network f(⋅;θ)f(\cdot; \theta) on noisy classification data (xi,yi)i=1N(x_i, y_i)_{i=1}^N while simultaneously recovering per-sample sparse noise parameters ui,vi∈[−1,1]Ku_i, v_i \in [-1, 1]^K.

    Input: Training dataset {(x_i, y_i)}_{i=1}^N, backbone network f(\cdot; \theta), number of epochs T, network learning rate \tau, noise learning rate ratios \alpha_u, \alpha_v, batch size \beta
    Initialization: Initialize \theta with standard schemes; initialize each entry of u_i, v_i i.i.d. from \mathcal{N}(0, 10^{-16}) for all i \in {1, \dots, N}
    for t = 1 to T do
        for b = 1 to N / \beta do
            Sample mini-batch B \subseteq {1, \dots, N} of size |B| = \beta
            \theta \leftarrow \theta - \tau \sum_{i \in B} \frac{\partial L_{CE}(\theta, u_i, v_i; x_i, y_i)}{\partial \theta}
        end for
        for i = 1 to N do
            u_i \leftarrow \mathcal{P}_{[-1, 1]}\left(u_i - \alpha_u \tau \frac{\partial L_{CE}(\theta, u_i, v_i; x_i, y_i)}{\partial u_i}\right)
            v_i \leftarrow \mathcal{P}_{[-1, 1]}\left(v_i - \alpha_v \tau \frac{\partial L_{MSE}(\theta, u_i, v_i; x_i, y_i)}{\partial v_i}\right)
        end for
    end for
    Output: Network parameters \theta and label noise variables {u_i, v_i}_{i=1}^N

    Here, P[−1,1](⋅)\mathcal{P}_{[-1, 1]}(\cdot) denotes entry-wise projection onto the interval [−1,1][-1, 1]. The network parameter θ\theta is optimized with SGD with momentum (e.g., 0.9) and weight decay (5×10−45 \times 10^{-4} or 10−310^{-3}), while uiu_i and viv_i are optimized with zero weight decay starting from an initial standard deviation of 10−810^{-8}.

  4. Knowl 4 — SOP+ with Consistency and Class-Balance Regularization

    model/method

    SOP+ augments the baseline Sparse Over-Parameterization objective with two regularization terms: a consistency regularizer LC(θ)\mathcal{L}_C(\theta) and a class-balance regularizer LB(θ)\mathcal{L}_B(\theta):

    LSOP+(θ,{ui,vi}i=1N)=L(θ,{ui,vi}i=1N)+λCLC(θ)+λBLB(θ)\mathcal{L}_{\mathrm{SOP+}}(\theta, \{u_i, v_i\}_{i=1}^N) = \mathcal{L}(\theta, \{u_i, v_i\}_{i=1}^N) + \lambda_C \mathcal{L}_C(\theta) + \lambda_B \mathcal{L}_B(\theta)

    where λC>0\lambda_C > 0 and λB>0\lambda_B > 0 are hyperparameter weights (e.g., λC=0.9,λB=0.1\lambda_C = 0.9, \lambda_B = 0.1).

    1. Consistency regularizer LC(θ)\mathcal{L}_C(\theta): Penalizes discrepancies between network softmax outputs on standard augmented images and heavily augmented images generated via Unsupervised Data Augmentation (UDA):

    LC(θ)=1N∑i=1NDKL(f(xi;θ) ∥ f(UDA(xi);θ))\mathcal{L}_C(\theta) = \frac{1}{N} \sum_{i=1}^N D_{\mathrm{KL}}\left(f(x_i; \theta) \,\|\, f(\mathrm{UDA}(x_i); \theta)\right)

    1. Class-balance regularizer LB(θ)\mathcal{L}_B(\theta): Prevents the network from collapsing into trivial constant predictions by minimizing the Kullback-Leibler divergence between the prior class distribution p=(p1,…,pK)p = (p_1, \dots, p_K) and the mean prediction across mini-batch BB:

    LB(θ)=∑k=1Kpklog⁡pkfˉk(x;θ)=−∑k=1Kpklog⁡fˉk(x;θ)\mathcal{L}_B(\theta) = \sum_{k=1}^K p_k \log \frac{p_k}{\bar{f}_k(x; \theta)} = -\sum_{k=1}^K p_k \log \bar{f}_k(x; \theta)

    where fˉ(x;θ)=1∣B∣∑x∈Bf(x;θ)\bar{f}(x; \theta) = \frac{1}{|B|} \sum_{x \in B} f(x; \theta) is the average prediction over the mini-batch.

  5. Knowl 5 — Implicit Sparse Regularization Induced by Hadamard Over-Parameterization

    theoretical result

    Consider the linearized model y=Jθ+sy = J\theta + s with measurement/Jacobian matrix J∈RN×pJ \in \mathbb{R}^{N \times p}, observations y∈RNy \in \mathbb{R}^N, and the non-convex over-parameterized objective:

    min⁡θ,u,vh(θ,u,v)=12∥Jθ+u⊙u−v⊙v−y∥22\min_{\theta, u, v} h(\theta, u, v) = \frac{1}{2} \|J\theta + u \odot u - v \odot v - y\|_2^2

    Let the gradient flow dynamics with learning rate scaling α\alpha on (u,v)(u, v) be defined by θ˙t=−J⊤rt\dot{\theta}_t = -J^\top r_t, u˙t=−2αut⊙rt\dot{u}_t = -2\alpha u_t \odot r_t, and v˙t=2αvt⊙rt\dot{v}_t = 2\alpha v_t \odot r_t, with residual rt=Jθt+ut⊙ut−vt⊙vt−yr_t = J\theta_t + u_t \odot u_t - v_t \odot v_t - y. Assume initialization θ0(γ,α)=0\theta_0(\gamma, \alpha) = 0, u0(γ,α)=γ1u_0(\gamma, \alpha) = \gamma \mathbf{1}, and v0(γ,α)=γ1v_0(\gamma, \alpha) = \gamma \mathbf{1} for small γ>0\gamma > 0.

    1. Global Convergence: For any fixed (γ,α)(\gamma, \alpha), if the limit (θ∞(γ,α),u∞(γ,α),v∞(γ,α))=lim⁡t→∞(θt,ut,vt)(\theta_\infty(\gamma, \alpha), u_\infty(\gamma, \alpha), v_\infty(\gamma, \alpha)) = \lim_{t \to \infty} (\theta_t, u_t, v_t) exists, it is a global minimizer of h(θ,u,v)h(\theta, u, v).
    2. Implicit ℓ1\ell_1-Regularization: Fix λ>0\lambda > 0 and couple the learning rate ratio α\alpha to initialization scale γ\gamma via:

    α(γ)=−log⁡γ2λ\alpha(\gamma) = -\frac{\log \gamma}{2\lambda}

    If the limit (θ^,u^,v^)=lim⁡γ→0(θ∞(γ,α(γ)),u∞(γ,α(γ)),v∞(γ,α(γ)))(\hat{\theta}, \hat{u}, \hat{v}) = \lim_{\gamma \to 0} (\theta_\infty(\gamma, \alpha(\gamma)), u_\infty(\gamma, \alpha(\gamma)), v_\infty(\gamma, \alpha(\gamma))) exists, then (θ^,s^)(\hat{\theta}, \hat{s}) with s^=u^⊙u^−v^⊙v^\hat{s} = \hat{u} \odot \hat{u} - \hat{v} \odot \hat{v} is an optimal solution to the regularized convex program:

    min⁡θ,s12∥θ∥22+λ∥s∥1s.t.y=Jθ+s\min_{\theta, s} \frac{1}{2} \|\theta\|_2^2 + \lambda \|s\|_1 \quad \text{s.t.} \quad y = J\theta + s

  6. Knowl 6 — Strict Saddle Property for Over-Parameterized Linear Robust Recovery

    theoretical result

    For the non-convex objective function:

    h(θ,u,v)=12∥Jθ+u⊙u−v⊙v−y∥22h(\theta, u, v) = \frac{1}{2} \|J\theta + u \odot u - v \odot v - y\|_2^2

    where J∈RN×pJ \in \mathbb{R}^{N \times p}, θ∈Rp\theta \in \mathbb{R}^p, u,v,y∈RNu, v, y \in \mathbb{R}^N, every critical point (θ,u,v)(\theta, u, v) is either:

    1. A global minimizer satisfying Jθ+u⊙u−v⊙v−y=0J\theta + u \odot u - v \odot v - y = 0; or
    2. A strict saddle point where the Hessian ∇2h(θ,u,v)\nabla^2 h(\theta, u, v) possesses at least one strictly negative eigenvalue.

    Specifically, for any critical point with non-zero residual entry ri≠0r_i \neq 0, it necessarily holds that ui=vi=0u_i = v_i = 0. Choosing a perturbation direction d=(dθ,du,dv)d = (d_\theta, d_u, d_v) along the ii-th coordinate yields a directional curvature of −2∣ri∣<0-2|r_i| < 0, guaranteeing that gradient descent with random initialization escapes saddle points almost surely.

  7. Knowl 7 — Matrix Coherence for Incoherent Robust Data Recovery

    definition

    Let J∈RN×pJ \in \mathbb{R}^{N \times p} be a matrix of rank rr with compact singular value decomposition J=UΣV⊤J = U \Sigma V^\top, where U∈RN×rU \in \mathbb{R}^{N \times r}, Σ∈Rr×r\Sigma \in \mathbb{R}^{r \times r}, and V∈Rp×rV \in \mathbb{R}^{p \times r}.

    The coherence of JJ with respect to the standard canonical basis {e1,…,eN}\{e_1, \dots, e_N\} of RN\mathbb{R}^N is defined as:

    μ(J)=Nrmax⁡1≤i≤N∥U⊤ei∥22\mu(J) = \frac{N}{r} \max_{1 \le i \le N} \|U^\top e_i\|_2^2

    The parameter μ(J)∈[1,N/r]\mu(J) \in [1, N/r] measures the degree to which the column space of JJ is spread out across coordinates. Small values of μ(J)\mu(J) indicate that the low-rank subspace is incoherent with the standard basis, preventing sparse corruptions from being indistinguishable from components of the underlying data representation.

  8. Knowl 8 — Exact Recovery of Sparse Corruptions and Low-Rank Linear Representations

    theoretical result

    Suppose observations y∈RNy \in \mathbb{R}^N are generated from an over-parameterized linear model y=Jθ∗+s∗y = J\theta_* + s_*, where J∈RN×pJ \in \mathbb{R}^{N \times p} is rank-rr with coherence μ(J)\mu(J), s∗∈RNs_* \in \mathbb{R}^N is a kk-sparse corruption vector, and θ∗=arg⁡min⁡θ∥θ∥22 s.t. y=Jθ+s∗\theta_* = \arg\min_\theta \|\theta\|_2^2 \text{ s.t. } y = J\theta + s_*.

    If the sparsity kk and rank rr satisfy:

    k2r<N4μ(J)k^2 r < \frac{N}{4\mu(J)}

    then there exists a threshold λ0=21+ρ1−ρ∥UΣ−1V⊤θ∗∥∞\lambda_0 = 2 \frac{1+\rho}{1-\rho} \|U\Sigma^{-1}V^\top \theta_*\|_\infty (where ρ∈(0,1)\rho \in (0, 1) satisfies (1+ρ)krNμ(J)≤ρ(1+\rho)k\sqrt{\frac{r}{N}\mu(J)} \le \rho) such that for all λ>λ0\lambda > \lambda_0, the unique minimizer of the convex program:

    min⁡θ,s12∥θ∥22+λ∥s∥1s.t.y=Jθ+s\min_{\theta, s} \frac{1}{2} \|\theta\|_2^2 + \lambda \|s\|_1 \quad \text{s.t.} \quad y = J\theta + s

    is exactly the ground-truth pair (θ∗,s∗)(\theta_*, s_*). Combined with the implicit regularization of gradient flow, gradient dynamics on the over-parameterized model with small initialization and learning rate ratio α<−log⁡γ2λ0\alpha < -\frac{\log \gamma}{2\lambda_0} exactly recover (θ∗,s∗)(\theta_*, s_*).

  9. Knowl 9 — Classification Accuracy under Synthetic Label Noise on CIFAR-10 and CIFAR-100

    data/table

    Under uniform symmetric label noise (20%,40%,60%,80%20\%, 40\%, 60\%, 80\%) and asymmetric class-pair flipping noise (40%40\%), Sparse Over-Parameterization (SOP and SOP+) was evaluated using ResNet-34 on CIFAR-10 and CIFAR-100 (mean ±\pm std over 5 runs). SOP avoids the generalization drop of standard Cross-Entropy (CE) training, consistently outperforming robust loss methods (Forward, GCE, SL, ELR) and semi-supervised ensemble methods (DivideMix, ELR+).

    CIFAR-10 (Symmetric) CIFAR-100 (Symmetric)
    Method 20% 40% 60% 80% 20% 40% 60% 80%
    CE 86.320.18 82.650.16 76.150.32 59.280.97 51.430.58 45.230.53 36.310.39 20.230.82
    FORWARD 87.990.36 83.250.38 74.960.65 54.640.44 39.192.61 31.051.44 19.121.95 8.990.58
    GCE 89.830.20 87.130.22 82.540.23 64.071.38 66.810.42 61.770.24 53.160.78 29.160.74
    SL 89.830.32 87.130.26 82.810.61 68.120.81 70.380.13 62.270.22 54.820.57 25.910.44
    ELR 91.160.08 89.150.17 86.120.49 73.860.61 74.210.22 68.280.31 59.280.67 29.780.56
    SOP (Ours) 93.180.57 90.090.27 86.760.22 68.320.77 74.670.30 70.120.57 60.260.41 30.200.63
    CIFAR-10 CIFAR-100
    Method Sym 20% Sym 50% Sym 80% Asym 40% Sym 20% Sym 50% Sym 80% Asym 40%
    DivideMix 96.1 94.6 93.2 93.4 77.1 74.6 60.2 72.1
    ELR+ 95.8 94.8 93.3 93.0 77.7 73.8 60.8 77.5
    SOP+ (Ours) 96.3 95.5 94.0 93.8 78.8 75.9 63.3 78.0
  10. Knowl 10 — Classification Accuracy under Real-World Label Noise Benchmarks

    data/table

    SOP and SOP+ achieve state-of-the-art performance across real-world noisy label datasets including Clothing1M (ResNet-50 pretrained), mini WebVision (InceptionResNetV2 evaluated on WebVision and ImageNet ILSVRC12 validation sets), and CIFAR-N (human worker label noise evaluated on CIFAR-10N and CIFAR-100N with PreActResNet-18).

    Method Clothing1M WebVision ILSVRC12
    CE 69.1 - -
    Forward 69.8 61.1 57.3
    Co-Teaching 69.2 63.6 61.5
    ELR 72.9 76.2 68.7
    CORES^2 73.2 - -
    SOP (Ours) 73.5 76.6 69.1
    CIFAR-10N CIFAR-100N
    Method Random 1 Random 2 Random 3 Aggregate Worst Noisy
    CE 85.020.65 86.461.79 85.160.61 87.770.38 77.691.55 55.500.66
    Forward 86.880.50 86.140.24 87.040.35 88.240.22 79.790.46 57.011.03
    Co-Teaching 90.330.13 90.300.17 90.150.18 91.200.13 83.830.13 60.370.27
    ELR+ 94.430.41 94.200.24 94.340.22 94.830.10 91.091.60 66.720.07
    CORES* 94.450.14 94.880.31 94.740.03 95.250.09 91.660.09 55.720.42
    SOP+ (Ours) 95.280.13 95.310.10 95.390.11 95.610.13 93.240.21 67.810.23

    In addition to higher accuracy, SOP and SOP+ offer training times of 1.0 hour and 2.1 hours respectively on CIFAR-10 (50% symmetric noise on a single V100 GPU), which is faster than DivideMix (5.4h), Co-teaching+ (4.4h), and ELR+ (2.3h).

Coverage note — None was omitted; all key theoretical formulations, landscapes, implicit regularizations, exact recovery theorems, optimization algorithms, and experimental results on synthetic and real label noise datasets are represented.

References

  1. 1.Algan, G. and Ulusoy, I. Image classification with deep learning in the presence of noisy labels: A survey. Knowledge-Based Systems, 215:106771, 2021.
  2. 2.Amid, E., Warmuth, M. K., Anil, R., and Koren, T. Robust bi-tempered logistic loss based on bregman divergences. Advances in Neural Information Processing Systems, 32: 15013–15022, 2019.
  3. 3.Arora, S., Cohen, N., Hu, W., and Luo, Y. Implicit regularization in deep matrix factorization. In Advances in Neural Information Processing Systems, pp. 7411–7422, 2019.
  4. 4.Belkin, M., Hsu, D., Ma, S., and Mandal, S. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019.
  5. 5.Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. A. Mixmatch: A holistic approach to semi-supervised learning. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  6. 6.Candes, E. J. and Tao, T. Decoding by linear programming. IEEE transactions on information theory, 51(12):4203–4215, 2005.
  7. 7.Candes, E. J., Li, X., Ma, Y., and Wright, J. Robust principal component analysis? Journal of the ACM (JACM), 58(3): 1–37, 2011.
  8. 8.Chang, H.-S., Learned-Miller, E., and McCallum, A. Active bias: Training more accurate neural networks by emphasizing high variance samples. Advances in Neural Information Processing Systems, 30:1002–1012, 2017.
  9. 9.Chen, T., Zhang, Z., Balachandra, S., Ma, H., Wang, Z., Wang, Z., et al. Sparsity winning twice: Better robust generalization from more efficient training. In International Conference on Learning Representations, 2021.
  10. 10.Chen, X. and Gupta, A. Webly supervised learning of convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1431–1439, 2015.
  11. 11.Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., and Liu, Y. Learning with instance-dependent label noise: A sample sieve approach. In International Conference on Learning Representations, 2021.
  12. 12.Chi, Y., Lu, Y. M., and Chen, Y. Nonconvex optimization meets low-rank matrix factorization: An overview. IEEE Transactions on Signal Processing, 67(20):5239–5269, 2019.
  13. 13.Chizat, L., Oyallon, E., and Bach, F. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, 32:2937–2947, 2019.
  14. 14.Chou, H.-H., Maly, J., and Rauhut, H. More is less: Inducing sparsity via overparameterization. arXiv preprint arXiv:2112.11027, 2021.
  15. 15.Claerbout, J. F. and Muir, F. Robust modeling with erratic data. Geophysics, 38(5):826–844, 1973.
  16. 16.Cohen, A., Dahmen, W., and DeVore, R. Compressed sensing and best k-term approximation. Journal of the American mathematical society, 22(1):211–231, 2009.
  17. 17.Davenport, M. A. and Romberg, J. An overview of low-rank matrix recovery from incomplete observations. IEEE Journal of Selected Topics in Signal Processing, 10(4): 608–622, 2016.
  18. 18.Diamond, S. and Boyd, S. CVXPY: A Python-embedded modeling language for convex optimization. Journal of Machine Learning Research, 17(83):1–5, 2016.
  19. 19.Ding, L., Jiang, L., Chen, Y., Qu, Q., and Zhu, Z. Rank over-specified robust matrix recovery: Subgradient method and exact recovery. Advances in Neural Information Processing Systems, 34, 2021.
  20. 20.Domahidi, A., Chu, E., and Boyd, S. ECOS: An SOCP solver for embedded systems. In European Control Conference (ECC), pp. 3071–3076, 2013.
  21. 21.Frenay, B. and Verleysen, M. Classification in the presence of label noise: a survey. IEEE transactions on neural networks and learning systems, 25(5):845–869, 2013.
  22. 22.Ge, R., Huang, F., Jin, C., and Yuan, Y. Escaping from saddle points—online stochastic gradient for tensor decomposition. In Conference on learning theory, pp. 797–842. PMLR, 2015.
  23. 23.Ghosh, A., Kumar, H., and Sastry, P. Robust loss functions under label noise for deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31, 2017.
  24. 24.Goldberger, J. and Ben-Reuven, E. Training deep neuralnetworks using a noise adaptation layer. 2017.
  25. 25.Gunasekar, S., Woodworth, B., Bhojanapalli, S., Neyshabur, B., and Srebro, N. Implicit regularization in matrix factorization. In 2018 Information Theory and Applications Workshop (ITA), pp. 1–10. IEEE, 2018.
  26. 26.Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  28. 28.Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K. Using trusted data to train deep networks on labels corrupted by severe noise. Advances in Neural Information Processing Systems, 31:10456–10465, 2018.
  29. 29.Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research, 22(241):1–124, 2021.
  30. 30.Hu, W., Li, Z., and Yu, D. Simple and effective regularization methods for training on noisily labeled data with generalization guarantee. In International Conference on Learning Representations, 2019.
  31. 31.Huang, L., Zhang, C., and Zhang, H. Self-adaptive training: beyond empirical risk minimization. arXiv preprint arXiv:2002.10319, 2020.
  32. 32.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. arXiv preprint arXiv:1806.07572, 2018.
  33. 33.Jacot, A., Ged, F., Gabriel, F., Simsek, B., and Hongler, C. Deep linear networks dynamics: Low-rank biases induced by initialization scale and l2 regularization. arXiv preprint arXiv:2106.15933, 2021.
  34. 34.Ji, Z., Dudık, M., Schapire, R. E., and Telgarsky, M. Gradient descent follows the regularization path for general losses. In Conference on Learning Theory, pp. 2109–2136. PMLR, 2020.
  35. 35.Jiang, L., Huang, D., Liu, M., and Yang, W. Beyond synthetic noise: Deep learning on controlled noisy labels. In International Conference on Machine Learning, pp. 4804–4815. PMLR, 2020.
  36. 36.Kalimeris, D., Kaplun, G., Nakkiran, P., Edelman, B., Yang, T., Barak, B., and Zhang, H. Sgd on neural networks learns functions of increasing complexity. Advances in Neural Information Processing Systems, 32:3496–3506, 2019.
  37. 37.Kim, T., Ko, J., Choi, J., Yun, S.-Y., et al. Fine samples for learning with noisy labels. Advances in Neural Information Processing Systems, 34, 2021.
  38. 38.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  39. 39.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25: 1097–1105, 2012.
  40. 40.Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. Gradient descent only converges to minimizers. In Conference on learning theory, pp. 1246–1257. PMLR, 2016.
  41. 41.Li, J., Wong, Y., Zhao, Q., and Kankanhalli, M. S. Learning to learn from noisy labeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5051–5059, 2019.
  42. 42.Li, J., Socher, R., and Hoi, S. C. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394, 2020a.
  43. 43.Li, J., Nguyen, T., Hegde, C., and Wong, K. W. Implicit sparse regularization: The impact of depth and early stopping. Advances in Neural Information Processing Systems, 34, 2021a.
  44. 44.Li, M., Soltanolkotabi, M., and Oymak, S. Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks. In International conference on artificial intelligence and statistics, pp. 4313–4324. PMLR, 2020b.
  45. 45.Li, W., Wang, L., Li, W., Agustsson, E., and Van Gool, L. Webvision database: Visual learning and understanding from web data. arXiv preprint arXiv:1708.02862, 2017.
  46. 46.Li, X., Liu, T., Han, B., Niu, G., and Sugiyama, M. Provably end-to-end label-noise learning without anchor points. arXiv preprint arXiv:2102.02400, 2021b.
  47. 47.Li, Y., Ma, T., and Zhang, H. Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations. In Conference On Learning Theory, pp. 2–47, 2018.
  48. 48.Li, Z., Luo, Y., and Lyu, K. Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning. In International Conference on Learning Representations, 2020c.
  49. 49.Lin, J. Z. and Bradic, J. Learning to combat noisy labels via classification margins. arXiv preprint arXiv:2102.00751, 2021.
  50. 50.Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C. Early-learning regularization prevents memorization of noisy labels. Advances in Neural Information Processing Systems, 33, 2020.
  51. 51.Liu, S., Liu, K., Zhu, W., Shen, Y., and Fernandez-Granda, C. Adaptive early-learning correction for segmentation from noisy annotations. ArXiv, abs/2110.03740, 2021a.
  52. 52.Liu, S., Yin, L., Mocanu, D. C., and Pechenizkiy, M. Do we actually need dense over-parameterization? in-time overparameterization in sparse training. In International Conference on Machine Learning, pp. 6989–7000. PMLR, 2021b.
  53. 53.Liu, T. and Tao, D. Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence, 38(3):447–461, 2015.
  54. 54.Loshchilov, I. and Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  55. 55.Lukasik, M., Bhojanapalli, S., Menon, A., and Kumar, S. Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp. 6448–6458. PMLR, 2020.
  56. 56.Ma, J. and Fattahi, S. Implicit regularization of sub-gradient method in robust matrix recovery: Don’t be afraid of outliers. arXiv preprint arXiv:2102.02969, 2021.
  57. 57.Ma, X., Wang, Y., Houle, M. E., Zhou, S., Erfani, S., Xia, S., Wijewickrema, S., and Bailey, J. Dimensionality-driven learning with noisy labels. In International Conference on Machine Learning, pp. 3355–3364. PMLR, 2018.
  58. 58.Ma, X., Huang, H., Wang, Y., Romano, S., Erfani, S., and Bailey, J. Normalized loss functions for deep learning with noisy labels. In International Conference on Machine Learning, pp. 6543–6553. PMLR, 2020.
  59. 59.Menon, A. K., Rawat, A. S., Reddi, S. J., and Kumar, S. Can gradient clipping mitigate label noise? In International Conference on Learning Representations, 2019.
  60. 60.Nacson, M. S., Srebro, N., and Soudry, D. Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 3051–3059. PMLR, 2019.
  61. 61.Neyshabur, B., Tomioka, R., and Srebro, N. In search of the real inductive bias: On the role of implicit regularization in deep learning. arXiv preprint arXiv:1412.6614, 2014.
  62. 62.Oymak, S. and Soltanolkotabi, M. Overparameterized nonlinear learning: Gradient descent takes the shortest path? In International Conference on Machine Learning, pp. 4951–4960, 2019.
  63. 63.Oymak, S., Fabian, Z., Li, M., and Soltanolkotabi, M. Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian. arXiv preprint arXiv:1906.05392, 2019.
  64. 64.Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1944–1952, 2017.
  65. 65.Razin, N. and Cohen, N. Implicit regularization in deep learning may not be explainable by norms. Advances in Neural Information Processing Systems, 33, 2020.
  66. 66.Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., and Rabinovich, A. Training deep neural networks on noisy labels with bootstrapping. arXiv preprint arXiv:1412.6596, 2014.
  67. 67.Song, H., Kim, M., and Lee, J.-G. Selfie: Refurbishing unclean samples for robust deep learning. In International Conference on Machine Learning, pp. 5907–5915. PMLR, 2019a.
  68. 68.Song, H., Kim, M., Park, D., and Lee, J.-G. Prestopping: How does early stopping help generalization against label noise? ArXiv, abs/1911.08059, 2019b.
  69. 69.Song, H., Kim, M., Park, D., Shin, Y., and Lee, J.-G. Learning from noisy labels with deep neural networks: A survey. arXiv preprint arXiv:2007.08199, 2020.
  70. 70.Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. The implicit bias of gradient descent on separable data. The Journal of Machine Learning Research, 19(1):2822–2878, 2018.
  71. 71.Stoger, D. and Soltanolkotabi, M. Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction. Advances in Neural Information Processing Systems, 34, 2021.
  72. 72.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818–2826, 2016.
  73. 73.Tanaka, D., Ikami, D., Yamasaki, T., and Aizawa, K. Joint optimization framework for learning with noisy labels. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5552–5560, 2018.
  74. 74.Tibshirani, R. J. Equivalences between sparse models and neural networks. 2021.
  75. 75.Vaskevicius, T., Kanade, V., and Rebeschini, P. Implicit regularization for optimal sparse recovery. In Advances in Neural Information Processing Systems, pp. 2968–2979, 2019.
  76. 76.Wang, B., Meng, Q., Zhang, H., Sun, R., Chen, W., and Ma, Z.-M. Momentum doesn’t change the implicit bias. arXiv preprint arXiv:2110.03891, 2021.
  77. 77.Wang, R., Liu, T., and Tao, D. Multiclass learning with partially corrupted labels. IEEE transactions on neural networks and learning systems, 29(6):2568–2580, 2017.
  78. 78.Wang, Y., Ma, X., Chen, Z., Luo, Y., Yi, J., and Bailey, J. Symmetric cross entropy for robust learning with noisy labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 322–330, 2019.
  79. 79.Wei, J. and Liu, Y. When optimizing f-divergence is robust with label noise. In International Conference on Learning Representations, 2021.
  80. 80.Wei, J., Liu, H., Liu, T., Niu, G., and Liu, Y. Understanding (generalized) label smoothing whenlearning with noisy labels. arXiv preprint arXiv:2106.04149, 2021a.
  81. 81.Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., and Liu, Y. Learning with noisy labels revisited: A study using real-world human annotations. arXiv preprint arXiv:2110.12088, 2021b.
  82. 82.Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N. Kernel and rich regimes in overparametrized models. In Conference on Learning Theory, pp. 3635–3673. PMLR, 2020.
  83. 83.Wright, J., Yang, A. Y., Ganesh, A., Sastry, S. S., and Ma, Y. Robust face recognition via sparse representation. IEEE transactions on pattern analysis and machine intelligence, 31(2):210–227, 2008.
  84. 84.Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., and Sugiyama, M. Are anchor points really indispensable in label-noise learning? Advances in Neural Information Processing Systems, 32:6838–6849, 2019.
  85. 85.Xia, X., Liu, T., Han, B., Gong, C., Wang, N., Ge, Z., and Chang, Y. Robust early-learning: Hindering the memorization of noisy labels. In International Conference on Learning Representations, 2020.
  86. 86.Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X. Learning from massive noisy labeled data for image classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2691–2699, 2015.
  87. 87.Xie, Q., Dai, Z., Hovy, E. H., Luong, M.-T., and Le, Q. V. Unsupervised data augmentation for consistency training. arXiv: Learning, 2020.
  88. 88.Yang, Z., Yu, Y., You, C., Steinhardt, J., and Ma, Y. Rethinking bias-variance trade-off for generalization of neural networks. In International Conference on Machine Learning, pp. 10767–10777. PMLR, 2020.
  89. 89.You, C., Zhu, Z., Qu, Q., and Ma, Y. Robust recovery via implicit bias of discrepant learning rates for double over-parameterization. Advances in Neural Information Processing Systems, 33:17733–17744, 2020.
  90. 90.Yu, Y., Chan, K. H. R., You, C., Song, C., and Ma, Y. Learning diverse and discriminative representations via the principle of maximal coding rate reduction. Advances in Neural Information Processing Systems, 33:9422–9434, 2020.
  91. 91.Zetterqvist, O., Jornsten, R., and Jonasson, J. Robust neural network classification via double regularization. arXiv preprint arXiv:2112.08102, 2021.
  92. 92.Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 64(3):107–115, 2021a.
  93. 93.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations, 2018.
  94. 94.Zhang, H., Xing, X., and Liu, L. Dualgraph: A graph-based method for reasoning about label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9654–9663, 2021b.
  95. 95.Zhang, Y., Niu, G., and Sugiyama, M. Learning noise transition matrix from only noisy labels via total variation regularization. arXiv preprint arXiv:2102.02414, 2021c.
  96. 96.Zhang, Z. and Sabuncu, M. R. Generalized cross entropy loss for training deep neural networks with noisy labels. In 32nd Conference on Neural Information Processing Systems (NeurIPS), 2018.
  97. 97.Zhao, P., Yang, Y., and He, Q.-C. Implicit regularization via hadamard product over-parametrization in high-dimensional linear regression. arXiv preprint arXiv:1903.09367, 2019.
  98. 98.Zheng, S., Wu, P., Goswami, A., Goswami, M., Metaxas, D., and Chen, C. Error-bounded correction of noisy labels. In International Conference on Machine Learning, pp. 11447–11457. PMLR, 2020.
  99. 99.Zhu, Z., Song, Y., and Liu, Y. Clusterability as an alternative to anchor points when learning with noisy labels. arXiv preprint arXiv:2102.05291, 2021.

Citation

MLA
Liu, S., et al. “Robust Training Under Label Noise by Over-parameterization”. International Conference on Machine Learning, vol. 162, 2022, pp. 14153–72, https://proceedings.mlr.press/v162/liu22w.html.
APA
Liu, S., Zhu, Z., Qu, Q., & You, C. (2022). Robust Training under Label Noise by Over-parameterization. International Conference on Machine Learning, 162, 14153–14172. https://proceedings.mlr.press/v162/liu22w.html
Chicago
Liu, S., Z. Zhu, Q. Qu, and C. You. 2022. “Robust Training Under Label Noise by Over-parameterization”. International Conference on Machine Learning 162: 14153–72. https://proceedings.mlr.press/v162/liu22w.html.
Harvard
Liu, S. et al. (2022) “Robust Training under Label Noise by Over-parameterization”, International Conference on Machine Learning. PMLR, pp. 14153–14172. Available at: https://proceedings.mlr.press/v162/liu22w.html.
Vancouver
1. Liu S, Zhu Z, Qu Q, You C (2022) Robust Training under Label Noise by Over-parameterization. In: International Conference on Machine Learning. PMLR, pp 14153–14172

BibTeX

@InProceedings{pmlr-v162-liu22w,
  title = 	 {Robust Training under Label Noise by Over-parameterization},
  author =       {Liu, Sheng and Zhu, Zhihui and Qu, Qing and You, Chong},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {14153--14172},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/liu22w/liu22w.pdf},
  url = 	 {https://proceedings.mlr.press/v162/liu22w.html},
  abstract = 	 {Recently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning. However, when the training data is corrupted, it has been well-known that over-parameterized networks tend to overfit and do not generalize. In this work, we propose a principled approach for robust training of over-parameterized deep networks in classification tasks where a proportion of training labels are corrupted. The main idea is yet very simple: label noise is sparse and incoherent with the network learned from clean data, so we model the noise and learn to separate it from the data. Specifically, we model the label noise via another sparse over-parameterization term, and exploit implicit algorithmic regularizations to recover and separate the underlying corruptions. Remarkably, when trained using such a simple method in practice, we demonstrate state-of-the-art test accuracy against label noise on a variety of real datasets. Furthermore, our experimental results are corroborated by theory on simplified linear models, showing that exact separation between sparse noise and low-rank data can be achieved under incoherent conditions. The work opens many interesting directions for improving over-parameterized models by using sparse over-parameterization and implicit regularization. Code is available at https://github.com/shengliu66/SOP.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/