DivideMix: Learning with Noisy Labels as Semi-supervised Learning

Junnan LiRichard SocherSteven C. H. Hoi

article2020ICLR1,431 citations

Proposes DivideMix, a framework that treats noisy label learning as a semi-supervised problem by dynamically partitioning data into clean and unlabeled sets and co-training two networks to prevent confirmation bias.

Listen

High-quality human data labeling for deep neural networks is expensive, time-consuming, and difficult to scale. Organizations frequently rely on automated web scraping or single-annotator pipelines to collect large datasets, but these practical alternatives inevitably introduce noisy, inaccurate labels that cause standard models to overfit and generalize poorly. The article introduces and evaluates a unified framework called DivideMix, which reformulates training on corrupted labels as a semi-supervised learning problem where questionable data points are stripped of their labels and repurposed to regularize model training.

The framework operates by training two neural networks simultaneously to avoid confirmation bias, where an individual model simply reinforces its own errors. In each training cycle, one network uses statistical loss modeling to partition the dataset into clean labeled samples and noisy unlabeled samples for the peer network. Both models are then trained on the combined data using data augmentation, cross-network label refinement for labeled data, and ensemble label guessing for unlabeled data. The article assesses this methodology across standard synthetic benchmarks, including CIFAR-10 and CIFAR-100 with noise levels ranging from 20% to 90%, as well as two large-scale real-world datasets, Clothing1M and WebVision.

The empirical findings demonstrate substantial performance gains over existing baseline methods. On synthetic benchmarks with extreme 80% to 90% label corruption, DivideMix achieves major accuracy improvements, outperforming existing techniques by up to approximately 10 to 12 percentage points on complex classification tasks. On real-world corrupted datasets, the method achieves state-of-the-art results, including over 12 percentage points higher top-1 accuracy on the WebVision benchmark and higher accuracy on Clothing1M. Ablation analyses confirm that dual-network co-training, data augmentation, and label refinement are critical to maintaining robustness and preventing performance degradation.

These results show that organizations do not need to discard imperfect or noisy data, nor do they need to invest heavily in exhaustive manual data cleaning. Repurposing suspect labels as unlabeled regularization data significantly lowers data curation costs, reduces operational risks tied to inaccurate training inputs, and maintains high classification performance. While training two networks concurrently requires more computing time than simple baseline training, the computational cost remains lower than multi-stage meta-learning alternatives.

Decision-makers and engineering teams working with noisy datasets should consider adopting dual-model semi-supervised training pipelines to salvage untrusted annotations and reduce data preparation overhead. Future initiatives should explore expanding this dual-network framework into domains outside computer vision, such as natural language processing, and test larger-scale deployment in production environments.

Cover for DivideMix: Learning with Noisy Labels as Semi-supervised Learning

Abstract

Deep neural networks are known to be annotation-hungry. Numerous efforts have been devoted to reducing the annotation cost when learning with deep networks. Two prominent directions include learning with noisy labels and semi-supervised learning by exploiting unlabeled data. In this work, we propose DivideMix, a novel framework for learning with noisy labels by leveraging semi-supervised learning techniques. In particular, DivideMix models the per-sample loss distribution with a mixture model to dynamically divide the training data into a labeled set with clean samples and an unlabeled set with noisy samples, and trains the model on both the labeled and unlabeled data in a semi-supervised manner. To avoid confirmation bias, we simultaneously train two diverged networks where each network uses the dataset division from the other network. During the semi-supervised training phase, we improve the MixMatch strategy by performing label co-refinement and label co-guessing on labeled and unlabeled samples, respectively. Experiments on multiple benchmark datasets demonstrate substantial improvements over state-of-the-art methods. Code is available at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Learning with Noisy Labels
  • 2.2 Semi-supervised Learning
  • 3 Method
  • 3.1 Co-Divide by Loss Modeling
  • 3.2 MixMatch with Label Co-Refinement and Co-Guessing
  • 4 Experiments
  • 4.1 Datasets and Implementation Details
  • 4.2 Comparison with State-of-the-art Methods
  • 4.3 Ablation Study
  • 5 Conclusion
  • References
  • A Additional Experiment Results
  • B Additional Training Details
  • C Additional Explanations for Ablation Study
  • D Training Time Analysis

Knowls

  1. Knowl 1 — DivideMix Algorithm for Learning with Noisy Labels

    algorithm

    DivideMix is a framework for robust deep neural network training in the presence of noisy labels by casting the problem into a semi-supervised learning (SSL) framework. Two networks, θ(1)\theta^{(1)} and θ(2)\theta^{(2)}, are trained simultaneously to prevent confirmation bias via co-divide, label co-refinement, and label co-guessing.

    Input: Model parameters θ(1)\theta^{(1)} and θ(2)\theta^{(2)}, dataset D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N, threshold τ\tau, augmentations MM, temperature TT, unsupervised loss weight λu\lambda_u, MixMatch parameter α\alpha, max epochs EmaxE_{\text{max}}.
    θ(1),θ(2)←WarmUp(D,θ(1),θ(2))\theta^{(1)}, \theta^{(2)} \leftarrow \text{WarmUp}(\mathcal{D}, \theta^{(1)}, \theta^{(2)})
    for e=1e = 1 to EmaxE_{\text{max}} do
        W(2)←FitGMM(D,θ(1))W^{(2)} \leftarrow \text{FitGMM}(\mathcal{D}, \theta^{(1)})
        W(1)←FitGMM(D,θ(2))W^{(1)} \leftarrow \text{FitGMM}(\mathcal{D}, \theta^{(2)})
        for k∈{1,2}k \in \{1, 2\} do
            X(k)←{(xi,yi,wi)∣wi≥τ,(xi,yi)∈D,wi∈W(k)}\mathcal{X}^{(k)} \leftarrow \{(x_i, y_i, w_i) \mid w_i \ge \tau, (x_i, y_i) \in \mathcal{D}, w_i \in W^{(k)}\}
            U(k)←{xi∣wi<τ,xi∈D,wi∈W(k)}\mathcal{U}^{(k)} \leftarrow \{x_i \mid w_i < \tau, x_i \in \mathcal{D}, w_i \in W^{(k)}\}
            for iter=1iter = 1 to NitersN_{\text{iters}} do
                Sample mini-batch {(xb,yb,wb)}b=1B\{(x_b, y_b, w_b)\}_{b=1}^B from X(k)\mathcal{X}^{(k)}
                Sample mini-batch {ub}b=1B\{u_b\}_{b=1}^B from U(k)\mathcal{U}^{(k)}
                for b=1b = 1 to BB do
                    for m=1m = 1 to MM do
                        x^b,m←Augment(xb)\hat{x}_{b,m} \leftarrow \text{Augment}(x_b)
                        u^b,m←Augment(ub)\hat{u}_{b,m} \leftarrow \text{Augment}(u_b)
                    end
                    pb←1M∑m=1Mpmodel(x^b,m;θ(k))p_b \leftarrow \frac{1}{M} \sum_{m=1}^M p_{\text{model}}(\hat{x}_{b,m}; \theta^{(k)})
                    yˉb←wbyb+(1−wb)pb\bar{y}_b \leftarrow w_b y_b + (1 - w_b) p_b
                    y^b←Sharpen(yˉb,T)\hat{y}_b \leftarrow \text{Sharpen}(\bar{y}_b, T)
                    qˉb←12M∑m=1M(pmodel(u^b,m;θ(1))+pmodel(u^b,m;θ(2)))\bar{q}_b \leftarrow \frac{1}{2M} \sum_{m=1}^M (p_{\text{model}}(\hat{u}_{b,m}; \theta^{(1)}) + p_{\text{model}}(\hat{u}_{b,m}; \theta^{(2)}))
                    qb←Sharpen(qˉb,T)q_b \leftarrow \text{Sharpen}(\bar{q}_b, T)
                end
                X^←{(x^b,m,y^b)∣b∈{1,…,B},m∈{1,…,M}}\hat{\mathcal{X}} \leftarrow \{(\hat{x}_{b,m}, \hat{y}_b) \mid b \in \{1,\dots,B\}, m \in \{1,\dots,M\}\}
                U^←{(u^b,m,qb)∣b∈{1,…,B},m∈{1,…,M}}\hat{\mathcal{U}} \leftarrow \{(\hat{u}_{b,m}, q_b) \mid b \in \{1,\dots,B\}, m \in \{1,\dots,M\}\}
                X′,U′←MixMatch(X^,U^,α)\mathcal{X}', \mathcal{U}' \leftarrow \text{MixMatch}(\hat{\mathcal{X}}, \hat{\mathcal{U}}, \alpha)
                L←LX(X′;θ(k))+λuLU(U′;θ(k))+λrLreg(X′,U′;θ(k))\mathcal{L} \leftarrow \mathcal{L}_{\mathcal{X}}(\mathcal{X}'; \theta^{(k)}) + \lambda_u \mathcal{L}_{\mathcal{U}}(\mathcal{U}'; \theta^{(k)}) + \lambda_r \mathcal{L}_{\text{reg}}(\mathcal{X}', \mathcal{U}'; \theta^{(k)})
                θ(k)←SGDUpdate(θ(k),∇θ(k)L)\theta^{(k)} \leftarrow \text{SGDUpdate}(\theta^{(k)}, \nabla_{\theta^{(k)}} \mathcal{L})
            end
        end
    end

    During inference, the ensemble prediction of both trained models 12(pmodel(x;θ(1))+pmodel(x;θ(2)))\frac{1}{2}(p_{\text{model}}(x; \theta^{(1)}) + p_{\text{model}}(x; \theta^{(2)})) is used to classify test samples.

  2. Knowl 2 — Co-Divide Dataset Partitioning via Gaussian Mixture Modeling

    model/method

    Because deep neural networks tend to learn clean patterns before fitting to label noise, clean samples typically exhibit lower cross-entropy loss than corrupted ones early in training. Given a dataset D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N where xix_i is an input sample and yi∈{0,1}Cy_i \in \{0, 1\}^C is its one-hot label over CC classes, the per-sample cross-entropy loss for network θ\theta is defined as:

    ℓi(θ)=−∑c=1Cyiclog⁡(pmodelc(xi;θ))\ell_i(\theta) = -\sum_{c=1}^C y_i^c \log(p_{\text{model}}^c(x_i; \theta))

    where pmodelc(xi;θ)p_{\text{model}}^c(x_i; \theta) denotes the softmax output probability for class cc.

    To model the distribution of clean versus noisy samples, a two-component Gaussian Mixture Model (GMM) is fitted to the max-normalized per-sample losses ℓ\ell using the Expectation-Maximization (EM) algorithm. For each sample xix_i, its probability of having a clean label is calculated as the posterior probability:

    wi=p(g∣ℓi)w_i = p(g \mid \ell_i)

    where gg represents the Gaussian component with the smaller mean loss.

    Samples are partitioned using a clean probability threshold τ\tau: samples with wi≥τw_i \ge \tau form a clean labeled dataset X\mathcal{X}, while samples with wi<τw_i < \tau have their labels discarded and form an unlabeled dataset U\mathcal{U}.

    To avoid self-training confirmation bias (where a single network repeatedly confirms its own misclassifications), DivideMix uses co-divide: two networks with parameters θ(1)\theta^{(1)} and θ(2)\theta^{(2)} are maintained with different random initializations and batch orderings. At each epoch, the GMM fitted to the loss distribution of network θ(1)\theta^{(1)} partitions the data used to train network θ(2)\theta^{(2)}, and vice versa.

  3. Knowl 3 — Confidence Penalty During Warm-Up for Asymmetric Noise

    model/method

    Before applying mixture modeling, the networks undergo a short warm-up period using standard cross-entropy loss. Under symmetric (uniform) label noise, warm-up naturally separates clean and noisy losses. However, under asymmetric (class-conditional) label noise, neural networks quickly overfit to corrupt labels and produce over-confident, low-entropy predictions. This causes per-sample losses across all samples to collapse near zero, preventing the Gaussian Mixture Model from separating clean from noisy samples.

    To counteract premature memorization of asymmetric noise, a confidence penalty based on negative entropy is added to the cross-entropy loss during the warm-up epochs. The prediction entropy for an input xx is given by:

    H=−∑c=1Cpmodelc(x;θ)log⁡(pmodelc(x;θ))H = -\sum_{c=1}^C p_{\text{model}}^c(x; \theta) \log(p_{\text{model}}^c(x; \theta))

    During warm-up, the network minimizes ℓCE−H\ell_{\text{CE}} - H, effectively penalizing confident output distributions and forcing the per-sample loss distribution to remain spread out and modelable by the GMM.

  4. Knowl 4 — Label Co-Refinement and Label Co-Guessing

    model/method

    In the semi-supervised phase of DivideMix, the two networks collaborate at the mini-batch level to refine labels of labeled samples and estimate labels for unlabeled samples.

    Label Co-Refinement: For a mini-batch of labeled samples {(xb,yb,wb)}b=1B\{(x_b, y_b, w_b)\}_{b=1}^B where wbw_b is the clean probability estimated by the other network, the ground-truth label yby_b is linearly combined with the current network's average prediction across MM stochastic augmentations x^b,m\hat{x}_{b,m}:

    pb=1M∑m=1Mpmodel(x^b,m;θ(k))p_b = \frac{1}{M} \sum_{m=1}^M p_{\text{model}}(\hat{x}_{b,m}; \theta^{(k)})

    yˉb=wbyb+(1−wb)pb\bar{y}_b = w_b y_b + (1 - w_b) p_b

    A temperature sharpening function with temperature parameter TT is applied to lower the label entropy:

    y^bc=Sharpen(yˉb,T)c=(yˉbc)1/T∑j=1C(yˉbj)1/T,c=1,…,C\hat{y}_b^c = \text{Sharpen}(\bar{y}_b, T)_c = \frac{(\bar{y}_b^c)^{1/T}}{\sum_{j=1}^C (\bar{y}_b^j)^{1/T}}, \quad c = 1, \dots, C

    Label Co-Guessing: For an unlabeled mini-batch {ub}b=1B\{u_b\}_{b=1}^B, pseudo-labels are guessed by ensembling the predictions from both networks across all MM augmentations:

    qˉb=12M∑m=1M(pmodel(u^b,m;θ(1))+pmodel(u^b,m;θ(2)))\bar{q}_b = \frac{1}{2M} \sum_{m=1}^M \left( p_{\text{model}}(\hat{u}_{b,m}; \theta^{(1)}) + p_{\text{model}}(\hat{u}_{b,m}; \theta^{(2)}) \right)

    qb=Sharpen(qˉb,T)q_b = \text{Sharpen}(\bar{q}_b, T)

  5. Knowl 5 — DivideMix Semi-Supervised Loss and Uniform Prior Regularization

    equation

    Given augmented labeled pairs X^={(x^b,m,y^b)}\hat{\mathcal{X}} = \{(\hat{x}_{b,m}, \hat{y}_b)\} and guessed unlabeled pairs U^={(u^b,m,qb)}\hat{\mathcal{U}} = \{(\hat{u}_{b,m}, q_b)\}, MixMatch transforms them via MixUp interpolation into mixed sets X′\mathcal{X}' and U′\mathcal{U}'. For a pair of samples (x1,x2)(x_1, x_2) with label vectors (p1,p2)(p_1, p_2),

    λ∼Beta(α,α),λ′=max⁡(λ,1−λ)\lambda \sim \text{Beta}(\alpha, \alpha), \quad \lambda' = \max(\lambda, 1 - \lambda) x′=λ′x1+(1−λ′)x2,p′=λ′p1+(1−λ′)p2x' = \lambda' x_1 + (1 - \lambda') x_2, \quad p' = \lambda' p_1 + (1 - \lambda') p_2

    The total optimization loss L\mathcal{L} for each network is:

    L=LX+λuLU+λrLreg\mathcal{L} = \mathcal{L}_{\mathcal{X}} + \lambda_u \mathcal{L}_{\mathcal{U}} + \lambda_r \mathcal{L}_{\text{reg}}

    where the supervised loss LX\mathcal{L}_{\mathcal{X}} is cross-entropy on X′\mathcal{X}', and the unsupervised loss LU\mathcal{L}_{\mathcal{U}} is the squared L2L_2 distance on U′\mathcal{U}':

    LX=−1∣X′∣∑x,p∈X′∑c=1Cpclog⁡(pmodelc(x;θ))\mathcal{L}_{\mathcal{X}} = -\frac{1}{|\mathcal{X}'|} \sum_{x, p \in \mathcal{X}'} \sum_{c=1}^C p^c \log(p_{\text{model}}^c(x; \theta))

    LU=1∣U′∣∑x,p∈U′∥p−pmodel(x;θ)∥22\mathcal{L}_{\mathcal{U}} = \frac{1}{|\mathcal{U}'|} \sum_{x, p \in \mathcal{U}'} \|p - p_{\text{model}}(x; \theta)\|_2^2

    To prevent the model from assigning all training instances to a single class under high noise ratios, Lreg\mathcal{L}_{\text{reg}} penalizes deviation from a uniform prior distribution π\pi (where πc=1/C\pi_c = 1/C):

    Lreg=∑c=1Cπclog⁡(πc1∣X′∣+∣U′∣∑x∈X′∪U′pmodelc(x;θ))\mathcal{L}_{\text{reg}} = \sum_{c=1}^C \pi_c \log \left( \frac{\pi_c}{\frac{1}{|\mathcal{X}'| + |\mathcal{U}'|} \sum_{x \in \mathcal{X}' \cup \mathcal{U}'} p_{\text{model}}^c(x; \theta)} \right)

    In implementations, λr=1\lambda_r = 1, and λu\lambda_u regulates the strength of the unlabeled data loss.

  6. Knowl 6 — CIFAR-10 and CIFAR-100 Benchmark Results Under Symmetric Noise

    data/table

    DivideMix was evaluated on CIFAR-10 and CIFAR-100 with symmetric label noise at corruption rates ranging from 20% to 90%, using an 18-layer PreAct ResNet architecture trained for 300 epochs. Both best test accuracy across all epochs and last test accuracy averaged over the final 10 epochs are reported.

    Dataset CIFAR-10 CIFAR-100
    Method / Noise Ratio 20% 50% 80% 90% 20% 50% 80% 90%
    Cross-Entropy Best 86.8 79.4 62.9 42.7 62.0 46.7 19.9 10.1
    Last 82.7 57.9 26.1 16.8 61.8 37.3 8.8 3.5
    Bootstrap Best 86.8 79.8 63.3 42.9 62.1 46.6 19.9 10.2
    Last 82.9 58.4 26.8 17.0 62.0 37.9 8.9 3.8
    F-correction Best 86.8 79.8 63.3 42.9 61.5 46.6 19.9 10.2
    Last 83.1 59.4 26.2 18.8 61.4 37.3 9.0 3.4
    Co-teaching+ Best 89.5 85.7 67.4 47.9 65.6 51.8 27.9 13.7
    Last 88.2 84.1 45.5 30.1 64.1 45.3 15.5 8.8
    Mixup Best 95.6 87.1 71.6 52.2 67.8 57.3 30.8 14.6
    Last 92.3 77.6 46.7 43.9 66.0 46.6 17.6 8.1
    P-correction Best 92.4 89.1 77.5 58.9 69.4 57.5 31.1 15.3
    Last 92.0 88.7 76.5 58.2 68.1 56.4 20.7 8.8
    Meta-Learning Best 92.9 89.3 77.4 58.7 68.5 59.2 42.4 19.5
    Last 92.0 88.8 76.1 58.3 67.7 58.0 40.1 14.3
    M-correction Best 94.0 92.0 86.8 69.1 73.9 66.1 48.2 24.3
    Last 93.8 91.9 86.6 68.7 73.4 65.4 47.6 20.5
    DivideMix Best 96.1 94.6 93.2 76.0 77.3 74.6 60.2 31.5
    Last 95.7 94.4 92.9 75.4 76.9 74.2 59.6 31.0

    DivideMix outperforms all previous methods across all noise ratios. On the most challenging CIFAR-100 setting with 80% and 90% noise, DivideMix exceeds prior state-of-the-art methods by approximately 12.0% and 7.2% in best accuracy, respectively.

  7. Knowl 7 — CIFAR-10 Benchmark Results Under Asymmetric Label Noise

    data/table

    DivideMix was evaluated on CIFAR-10 with 40% asymmetric (class-conditional) noise, designed to mimic real-world label mistakes by mapping semantically similar classes (e.g., deer →\to horse, dog ↔\leftrightarrow cat) to one another. All baselines were re-implemented using an 18-layer PreAct ResNet architecture under identical training conditions.

    Method Best (%) Last (%)
    Cross-Entropy 85.0 72.3
    F-correction 87.2 83.1
    M-correction 87.4 86.3
    Iterative-CV 88.6 88.0
    P-correction 88.5 88.1
    Joint-Optim 88.9 88.4
    Meta-Learning 89.2 88.6
    DivideMix 93.4 92.1

    DivideMix achieves 93.4% best test accuracy and 92.1% last test accuracy, outperforming the strongest baseline (Meta-Learning) by 4.2% in best accuracy and 3.5% in final averaged accuracy.

  8. Knowl 8 — Performance on Real-World Noisy Benchmarks: Clothing1M and WebVision

    data/table

    DivideMix was validated on two real-world noisy datasets: Clothing1M (1 million online shopping images with label noise from surrounding text, trained with ResNet-50 initialized from ImageNet) and mini WebVision 1.0 (the first 50 classes of the Google image subset, trained with Inception-ResNet-v2 from scratch and evaluated on both WebVision and ImageNet ILSVRC12 validation sets).

    Method Clothing1M Test Accuracy (%)
    Cross-Entropy 69.21
    F-correction 69.84
    M-correction 71.00
    Joint-Optim 72.16
    Meta-Cleaner 72.50
    Meta-Learning 73.47
    P-correction 73.49
    DivideMix 74.76
    Method WebVision Val (%) ILSVRC12 Val (%)
    Top-1 Top-5 Top-1 Top-5
    F-correction 61.12 82.68 57.36 82.36
    Decoupling 62.54 84.74 58.26 82.26
    D2L 62.68 84.00 57.80 81.36
    MentorNet 63.00 81.40 57.80 79.92
    Co-teaching 63.58 85.20 61.48 84.70
    Iterative-CV 65.24 85.34 61.60 84.98
    DivideMix 77.32 91.64 75.20 90.84

    On Clothing1M, DivideMix reaches 74.76% accuracy. On WebVision, DivideMix achieves 77.32% top-1 accuracy (an absolute improvement of 12.08% over the previous best result, Iterative-CV) and 75.20% top-1 accuracy when generalized to the ImageNet validation set (an absolute improvement of 13.60%).

  9. Knowl 9 — Ablation Study of DivideMix Components

    data/table

    An ablation study evaluated the contribution of individual algorithmic modules across symmetric (20%--90%) and asymmetric (40%) noise on CIFAR-10 and CIFAR-100.

    Dataset CIFAR-10 CIFAR-100
    Noise Type Symmetric Asymmetric Symmetric
    Methods / Noise Ratio 20% 50% 80% 90% 40% 20% 50% 80% 90%
    DivideMix Best 96.1 94.6 93.2 76.0 93.4 77.3 74.6 60.2 31.5
    Last 95.7 94.4 92.9 75.4 92.1 76.9 74.2 59.6 31.0
    DivideMix with θ(1)\theta^{(1)} test Best 95.2 94.2 93.0 75.5 92.7 75.2 72.8 58.3 29.9
    Last 95.0 93.7 92.4 74.2 91.4 74.8 72.1 57.6 29.2
    DivideMix w/o co-training Best 95.0 94.0 92.6 74.3 91.9 74.8 72.3 56.7 27.7
    Last 94.8 93.3 92.2 73.2 90.6 74.1 71.7 56.3 27.2
    DivideMix w/o label refinement Best 96.0 94.6 93.0 73.7 87.7 76.9 74.2 58.7 26.9
    Last 95.5 94.2 92.7 73.0 86.3 76.4 73.9 58.2 26.3
    DivideMix w/o augmentation Best 95.3 94.1 92.2 73.9 89.5 76.5 73.1 58.2 26.9
    Last 94.9 93.5 91.8 73.0 88.4 76.2 72.6 58.0 26.4
    Divide and MixMatch Best 94.1 92.8 89.7 70.1 86.5 73.7 70.5 55.3 25.0
    Last 93.5 92.3 89.1 68.6 85.2 72.4 69.7 53.9 23.7

    Key takeaways from the ablation are:

    1. Ensemble testing: Evaluating with the ensemble of both networks yields consistent gains over single-model inference (θ(1)\theta^{(1)} test).
    2. Co-training: Training a single network via self-divide causes performance degradation due to confirmation bias (especially under severe noise where CIFAR-100 90% drops from 31.0% to 27.2%).
    3. Label co-refinement: Crucial for high-noise and asymmetric scenarios (CIFAR-10 40% asymmetric drops from 92.1% to 86.3% without label refinement) because misdivided noisy samples in the labeled set are otherwise uncorrected.
    4. Augmentation and temperature sharpening: Input augmentation provides consistency regularization, and the full pipeline substantially outperforms naive dataset division combined with default MixMatch.
  10. Knowl 10 — Training Time Analysis and Computational Breakdown

    data/table

    The computational training time of DivideMix was evaluated on CIFAR-10 using a single NVIDIA V100 GPU and compared to other robust learning baselines.

    Co-teaching+ P-correction Meta-Learning DivideMix
    4.3 h 6.0 h 8.6 h 5.2 h

    Per-epoch operation time for DivideMix breaks down as follows:

    Co-Divide (GMM fitting) Data MixMatch (Augmentation/Interpolation) Forward-Backward Optimization
    17.2 s 16.0 s 12.5 s

    DivideMix requires 5.2 total hours on CIFAR-10, which is faster than P-correction (6.0 h) and Meta-Learning (8.6 h), while maintaining moderate per-epoch overhead across GMM fitting (17.2 s), MixMatch data preparation (16.0 s), and network backpropagation (12.5 s).

Coverage note — None was omitted; all key contributions including co-divide GMM loss modeling, confidence penalty warm-up, label co-refinement, label co-guessing, semi-supervised MixMatch loss formulation, and all empirical benchmarks/ablations are covered.

References

  1. 1.Eric Arazo, Diego Ortego, Paul Albert, Noel E. O’Connor, and Kevin McGuinness. Unsupervised label noise modeling and loss correction. In ICML, pp. 312–321, 2019.
  2. 2.Devansh Arpit, Stanislaw Jastrzkebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron C. Courville, Yoshua Bengio, and Simon Lacoste-Julien. A closer look at memorization in deep networks. In ICML, pp. 233–242, 2017.
  3. 3.David Berthelot, Nicholas Carlini, Ian J. Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. NeurIPS, 2019.
  4. 4.Pengfei Chen, Benben Liao, Guangyong Chen, and Shengyu Zhang. Understanding and utilizing deep neural networks trained with noisy labels. In ICML, pp. 1062–1070, 2019.
  5. 5.Yifan Ding, Liqiang Wang, Deliang Fan, and Boqing Gong. A semi-supervised two-stage approach to learning from noisy labels. In WACV, pp. 1215–1224, 2018.
  6. 6.Jacob Goldberger and Ehud Ben-Reuven. Training deep neural-networks using a noise adaptation layer. In ICLR, 2017.
  7. 7.Yves Grandvalet and Yoshua Bengio. Semi-supervised learning by entropy minimization. In NIPS, pp. 529–536, 2004.
  8. 8.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor W. Tsang, and Masashi Sugiyama. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In NeurIPS, pp. 8536–8546, 2018.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In ECCV, pp. 630–645, 2016.
  10. 10.Dan Hendrycks, Mantas Mazeika, Duncan Wilson, and Kevin Gimpel. Using trusted data to train deep networks on labels corrupted by severe noise. In NeurIPS, pp. 10477–10486, 2018.
  11. 11.Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In ICML, pp. 2309–2318, 2018.
  12. 12.Kyeongbo Kong, Junggi Lee, Youngchul Kwak, Minsung Kang, Seong Gyun Kim, and Woo-Jin Song. Recycling: Semi-supervised learning with noisy labels in deep neural networks. IEEE Access, 7:66998–67005, 2019.
  13. 13.Nikola Konstantinov and Christoph Lampert. Robust learning from untrusted sources. In ICML, pp. 3488–3498, 2019.
  14. 14.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Mater’s thesis, University of Toronto, 2009.
  15. 15.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. arXiv:1811.00982, 2018.
  16. 16.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. In ICLR, 2017.
  17. 17.Dong-Hyun Lee. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In ICML Workshop on Challenges in Representation Learning, volume 3, pp. 2, 2013.
  18. 18.Kuang-Huei Lee, Xiaodong He, Lei Zhang, and Linjun Yang. Cleannet: Transfer learning for scalable image classifier training with label noise. In CVPR, pp. 5447–5456, 2018.
  19. 19.Junnan Li, Yongkang Wong, Qi Zhao, and Mohan S. Kankanhalli. Learning to learn from noisy labeled data. In CVPR, pp. 5051–5059, 2019.
  20. 20.Wen Li, Limin Wang, Wei Li, Eirikur Agustsson, and Luc Van Gool. Webvision database: Visual learning and understanding from web data. arXiv:1708.02862, 2017a.
  21. 21.Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li. Learning from noisy labels with distillation. In ICCV, pp. 1928–1936, 2017b.
  22. 22.Xingjun Ma, Yisen Wang, Michael E. Houle, Shuo Zhou, Sarah M. Erfani, Shu-Tao Xia, Sudanthi N. R. Wijewickrema, and James Bailey. Dimensionality-driven learning with noisy labels. In ICML, pp. 3361–3370, 2018.
  23. 23.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9:2579–2605, 2008.
  24. 24.Dhruv Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, pp. 185–201, 2018.
  25. 25.Eran Malach and Shai Shalev-Shwartz. Decoupling “when to update” from “how to update”. In NIPS, pp. 960–970, 2017.
  26. 26.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: A regularization method for supervised and semi-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell., 41(8):1979–1993, 2019.
  27. 27.Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu. Making deep neural networks robust to label noise: A loss correction approach. In CVPR, pp. 2233–2241, 2017.
  28. 28.Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. Regularizing neural networks by penalizing confident output distributions. In ICLR Workshop, 2017.
  29. 29.Haim H. Permuter, Joseph M. Francos, and Ian Jermyn. A study of gaussian mixture models of color and texture features for image classification and segmentation. Pattern Recognition, 39(4): 695–706, 2006.
  30. 30.Scott E. Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich. Training deep neural networks on noisy labels with bootstrapping. In ICLR, 2015.
  31. 31.Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun. Learning to reweight examples for robust deep learning. In ICML, pp. 4331–4340, 2018.
  32. 32.Yanyao Shen and Sujay Sanghavi. Learning with bad training data via iterative trimmed loss minimization. In ICML, pp. 5739–5748, 2019.
  33. 33.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, pp. 4278–4284, 2017.
  34. 34.Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa. Joint optimization framework for learning with noisy labels. In CVPR, pp. 5552–5560, 2018.
  35. 35.Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C. Alexander, and Nathan Silberman. Learning from noisy labels by regularized estimation of annotator confusion. In CVPR, 2019.
  36. 36.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NIPS, pp. 1195–1204, 2017.
  37. 37.Sunil Thulasidasan, Tanmoy Bhattacharya, Jeff A. Bilmes, Gopinath Chennupati, and Jamal Mohd-Yusof. Combating label noise in deep learning using abstention. In ICML, pp. 6234–6243, 2019.
  38. 38.Arash Vahdat. Toward robustness against label noise in training deep discriminative neural networks. In NIPS, pp. 5601–5610, 2017.
  39. 39.Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta, and Serge J. Belongie. Learning from noisy large-scale datasets with minimal supervision. In CVPR, pp. 6575–6583, 2017.
  40. 40.Yisen Wang, Weiyang Liu, Xingjun Ma, James Bailey, Hongyuan Zha, Le Song, and Shu-Tao Xia. Iterative learning with open-set noisy labels. In CVPR, pp. 8688–8696, 2018.
  41. 41.Tong Xiao, Tian Xia, Yi Yang, Chang Huang, and Xiaogang Wang. Learning from massive noisy labeled data for image classification. In CVPR, pp. 2691–2699, 2015.
  42. 42.Kun Yi and Jianxin Wu. Probabilistic end-to-end noise correction for learning with noisy labels. In CVPR, 2019.
  43. 43.Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor W. Tsang, and Masashi Sugiyama. How does disagreement help generalization against label corruption? In ICML, pp. 7164–7173, 2019.
  44. 44.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In ICLR, 2017.
  45. 45.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  46. 46.Weihe Zhang, Yali Wang, and Yu Qiao. Metacleaner: Learning to hallucinate clean representations for noisy-labeled visual recognition. In CVPR, 2019.

Citation

MLA
Li, J., et al. “DivideMix: Learning with Noisy Labels as Semi-supervised Learning”. International Conference on Learning Representations, 2020, 2020, http://arxiv.org/abs/2002.07394v1.
APA
Li, J., Socher, R., & Hoi, S. C. H. (2020). DivideMix: Learning with Noisy Labels as Semi-supervised Learning. International Conference on Learning Representations, 2020. http://arxiv.org/abs/2002.07394v1
Chicago
Li, J., R. Socher, and S. C. H. Hoi. 2020. “DivideMix: Learning with Noisy Labels as Semi-supervised Learning”. International Conference on Learning Representations, 2020. http://arxiv.org/abs/2002.07394v1.
Harvard
Li, J., Socher, R. and Hoi, S.C.H. (2020) “DivideMix: Learning with Noisy Labels as Semi-supervised Learning”, International Conference on Learning Representations, 2020 [Preprint]. Available at: http://arxiv.org/abs/2002.07394v1.
Vancouver
1. Li J, Socher R, Hoi SCH (2020) DivideMix: Learning with Noisy Labels as Semi-supervised Learning. International Conference on Learning Representations, 2020

BibTeX

@article{li2020dividemix,
  title = {DivideMix: Learning with Noisy Labels as Semi-supervised Learning},
  author = {Li, Junnan and Socher, Richard and Hoi, Steven C. H.},
  year = {2020},
  journal = {International Conference on Learning Representations, 2020},
  url = {http://arxiv.org/abs/2002.07394v1},
  eprint = {2002.07394}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors