Selective-Supervised Contrastive Learning with Noisy Labels

Shikun LiXiaobo XiaShiming GeTongliang Liu

article2022CVPR252 citations

Proposes a selective-supervised contrastive learning framework that dynamically identifies reliable sample pairs without requiring prior knowledge of label noise rates, preventing corrupted annotations from degrading representation quality.

Listen

Deep learning models require massive datasets to achieve high accuracy, but obtaining clean, manually verified labels is costly and time-consuming. While web scraping and user tagging offer cheaper alternatives, they introduce noisy, incorrect labels that corrupt internal model representations and degrade real-world decision-making. Existing contrastive learning approaches attempt to build representations by comparing data pairs, but noisy labels inject incorrect pair relationships that undermine performance, especially when noise rates cannot be estimated in advance.

The article demonstrates a novel framework called selective-supervised contrastive learning to learn robust visual representations from noisily labeled data without prior knowledge of the exact noise rate. The primary objective is to evaluate whether dynamically filtering and selecting confident data pairs during training can shield deep neural networks from the harmful effects of mislabeled training data.

The researchers developed a two-stage approach that iteratively selects high-confidence examples and pairs during pre-training. First, confident individual examples are identified by checking agreement between learned representations and given labels across nearest neighbors. Next, confident pairs are constructed from these clean examples and augmented with additional pairs that share high representation similarity, even if their nominal class labels are wrong. The pre-trained model is then fine-tuned on confident examples using standard classification techniques. The authors benchmarked this framework across simulated noise conditions using standard image recognition datasets (CIFAR-10 and CIFAR-100) and evaluated real-world performance on the WebVision dataset.

The findings confirm that the proposed method consistently outperforms existing state-of-the-art baselines. Under heavy noise conditions of 80% symmetric noise on CIFAR-100, the method improved feature representation quality to 62.49% weighted accuracy, compared to 55.58% achieved by existing contrastive methods and 41.00% by standard supervised contrastive learning. Across simulated benchmark classifications, the framework delivered top accuracy, excelling particularly under asymmetric noise where mislabeling occurs between semantically related classes (achieving up to 74.2% accuracy under 40% asymmetric noise on CIFAR-100). On the real-world WebVision benchmark, the approach achieved the highest validation performance with a top-1 accuracy of 79.96% and top-5 accuracy of 92.64%, surpassing competing methods.

These results demonstrate that organizations can significantly lower data curation costs by effectively training reliable computer vision models directly on cheaper, web-scraped, or crowdsourced data. By focusing on representation similarity rather than relying on strict label correctness, the approach reduces the risk of model failure caused by systematic human labeling errors in real-world deployment pipelines.

Organizations training vision models on unverified datasets should consider adopting confident pair selection workflows to enhance robustness without paying for expensive re-labeling campaigns. Engineering teams can integrate the pre-training stage with existing fine-tuning pipelines to gain immediate accuracy improvements. Future development should focus on extending this framework to other operational domains, including object detection and text matching.

The main limitations are computational: contrastive pair-wise learning demands large batch sizes or extensive memory queues, and nearest-neighbor calculations increase processing overhead during training. Despite these computational costs, the empirical results provide high confidence in the method's ability to maintain robustness across varying noise types and levels.

Cover for Selective-Supervised Contrastive Learning with Noisy Labels

Abstract

Deep networks have strong capacities of embedding data into latent representations and finishing following tasks. However, the capacities largely come from high-quality annotated labels, which are expensive to collect. Noisy labels are more affordable, but result in corrupted representations, leading to poor generalization performance. To learn robust representations and handle noisy labels, we propose selective-supervised contrastive learning (Sel-CL) in this paper. Specifically, Sel-CL extend supervised contrastive learning (Sup-CL), which is powerful in representation learning, but is degraded when there are noisy labels. Sel-CL tackles the direct cause of the problem of Sup-CL. That is, as Sup-CL works in a pair-wise manner, noisy pairs built by noisy labels mislead representation learning. To alleviate the issue, we select confident pairs out of noisy ones for Sup-CL without knowing noise rates. In the selection process, by measuring the agreement between learned representations and given labels, we first identify confident examples that are exploited to build confident pairs. Then, the representation similarity distribution in the built confident pairs is exploited to identify more confident pairs out of noisy pairs. All obtained confident pairs are finally used for Sup-CL to enhance representations. Experiments on multiple noisy datasets demonstrate the robustness of the learned representations by our method, following the state-of-the-art performance. Source codes are available at https://github.com/ShikunLi/Sel-CL.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Learning with Noisy Labels
  • 2.2. Contrastive Learning
  • 3. Selective-Supervised Contrastive Learning
  • 3.1. Selecting Confident Examples
  • 3.2. Selecting Confident Pairs
  • 3.3. Representation Learning with Selected Pairs
  • 3.4. Classification Fine-tuning
  • 4. Experiments
  • 4.1. Datasets and Implementation Details
  • 4.2. Representation Learning Evaluations
  • 4.3. Results on Simulated Noisy Datasets
  • 4.4. Pair-wise Selection Analysis
  • 4.5. Results on the Real-World Noisy Dataset
  • 4.6. Ablation Study and Discussions
  • 5. Limitations
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Selective-Supervised Contrastive Learning (Sel-CL) Framework

    model/method

    Selective-Supervised Contrastive Learning (Sel-CL) is a two-stage representation learning framework designed to mitigate the detrimental impact of noisy labels in training datasets D~={(xi,y~i)}i=1n\tilde{\mathcal{D}} = \{(x_i, \tilde{y}_i)\}_{i=1}^n, where xix_i is an instance and y~i∈[C]={1,…,C}\tilde{y}_i \in [C] = \{1, \dots, C\} is its potentially corrupted label.

    The framework operates in two stages:

    1. Representation Pre-training: A neural network consisting of a backbone encoder ff (mapping xix_i to high-dimensional representation vi\mathbf{v}_i), a classification head (mapping vi\mathbf{v}_i to class probabilities p^(xi)\hat{\mathbf{p}}(x_i)), and a projection head (mapping vi\mathbf{v}_i to low-dimensional embedding zi\mathbf{z}_i) is pre-trained. After a brief warm-up with unsupervised contrastive learning, the network alternately identifies a set of confident examples T\mathcal{T} and selects a set of confident positive pairs G=G′∪G′′\mathcal{G} = \mathcal{G}' \cup \mathcal{G}'' at each epoch. The model parameters are optimized via a composite loss function combining supervised contrastive loss on selected pairs with Mixup, classification loss on confident examples, and a classifier-based similarity loss. This establishes a positive feedback loop: better confident pairs yield superior representations, which in turn enable more accurate confident pair selection.

    2. Classifier Fine-Tuning (Sel-CL+): The pre-trained encoder ff is frozen or fine-tuned with a new classification head f′f' trained exclusively on the identified confident examples T\mathcal{T} using a robust classification loss, or integrated with existing noisy-label learning frameworks.

  2. Knowl 2 — Confident Example Selection via K-Nearest Neighbor Pseudo-Labeling

    model/method

    To identify reliable training examples without prior knowledge of the dataset noise rate, Sel-CL measures the agreement between learned instance embeddings and given noisy labels using local neighborhood aggregation.

    Given the normalized representation zi\mathbf{z}_i of instance xix_i, cosine similarity between instances xix_i and xjx_j is computed as: d(zi,zj)=zizj⊤∥zi∥∥zj∥d(\mathbf{z}_i, \mathbf{z}_j) = \frac{\mathbf{z}_i \mathbf{z}_j^\top}{\|\mathbf{z}_i\| \|\mathbf{z}_j\|}

    For each example (xi,y~i)(x_i, \tilde{y}_i), its top-KK nearest neighbors Ni\mathcal{N}_i (K=250K = 250) in the embedding space are identified. A pseudo-label distribution q^(xi)=[q^1(xi),…,q^C(xi)]⊤\hat{\mathbf{q}}(x_i) = [\hat{q}_1(x_i), \dots, \hat{q}_C(x_i)]^\top approximating the clean posterior is constructed by voting across neighbor pseudo-labels: q^c(xi)=1K∑xk∈NiI[y^k=c],c∈[C]\hat{q}_c(x_i) = \frac{1}{K} \sum_{x_k \in \mathcal{N}_i} \mathbb{I}[\hat{y}_k = c], \quad c \in [C] where y^k\hat{y}_k is the dominant label in the neighborhood of xkx_k.

    Confident examples for class cc are identified by thresholding the cross-entropy loss ℓ\ell between q^(xi)\hat{\mathbf{q}}(x_i) and y~i\tilde{y}_i: Tc={(xi,y~i)∣ℓ(q^(xi),y~i)<γc,i∈[n]},c∈[C]\mathcal{T}_c = \{(x_i, \tilde{y}_i) \mid \ell(\hat{\mathbf{q}}(x_i), \tilde{y}_i) < \gamma_c, i \in [n]\}, \quad c \in [C] where γc\gamma_c is a dynamic per-class threshold set to the α\alpha-fractile of per-class label agreement counts ∑i=1nI[y^i=y~i][y~i=c]\sum_{i=1}^n \mathbb{I}[\hat{y}_i = \tilde{y}_i][\tilde{y}_i = c] across classes to ensure balanced selection. The full set of confident examples is T=⋃c=1CTc\mathcal{T} = \bigcup_{c=1}^C \mathcal{T}_c.

  3. Knowl 3 — Confident Pair Selection via Dynamic Representation Similarity Thresholding

    model/method

    In supervised contrastive learning, sample pairs sharing the same label are treated as positive pairs. When labels are noisy, pair selection in Sel-CL constructs a confident positive pair set G=G′∪G′′\mathcal{G} = \mathcal{G}' \cup \mathcal{G}'' through two complementary mechanisms:

    1. Example-Derived Confident Pairs (G′\mathcal{G}'): Pairs formed by instances that are both present in the confident example set T\mathcal{T} and share the same given label: G′={Pij∣y~i=y~j,(xi,y~i)∈T,(xj,y~j)∈T}\mathcal{G}' = \{P_{ij} \mid \tilde{y}_i = \tilde{y}_j, (x_i, \tilde{y}_i) \in \mathcal{T}, (x_j, \tilde{y}_j) \in \mathcal{T}\} where PijP_{ij} denotes the pair composed of (xi,y~i)(x_i, \tilde{y}_i) and (xj,y~j)(x_j, \tilde{y}_j).

    2. Similarity-Derived Confident Pairs (G′′\mathcal{G}''): Pairs that share the same nominal label (s~ij=I[y~i=y~j]=1\tilde{s}_{ij} = \mathbb{I}[\tilde{y}_i = \tilde{y}_j] = 1) and exhibit high latent representation similarity above a dynamic threshold γ\gamma: G′′={Pij∣s~ij=1,d(zi,zj)>γ}\mathcal{G}'' = \{P_{ij} \mid \tilde{s}_{ij} = 1, d(\mathbf{z}_i, \mathbf{z}_j) > \gamma\} where d(zi,zj)d(\mathbf{z}_i, \mathbf{z}_j) is the cosine similarity between low-dimensional embeddings zi\mathbf{z}_i and zj\mathbf{z}_j. The threshold γ\gamma is dynamically set as the β\beta-fractile of the representation similarity distribution computed over all pairs in G′\mathcal{G}'.

    This selection captures pairs where individual class labels might be incorrect, but both instances belong to the same true latent category.

  4. Knowl 4 — Multi-Objective Pre-training Loss Formulation in Sel-CL

    equation

    The overall pre-training objective of Sel-CL combines contrastive learning on selected pairs, classification on confident examples, and a pairwise similarity prediction loss: LALL=LMIX+λcLCLS+λsLSIM\mathcal{L}^{\text{ALL}} = \mathcal{L}^{\text{MIX}} + \lambda_c \mathcal{L}^{\text{CLS}} + \lambda_s \mathcal{L}^{\text{SIM}} where λc=1.0\lambda_c = 1.0 and λs=0.01\lambda_s = 0.01 are weighting hyperparameters.

    Given a mini-batch of NN samples with two random augmentations per sample yielding 2N2N views indexed by I=[2N]I = [2N], with A(i)=I∖{i}A(i) = I \setminus \{i\} and G(i)={g∈A(i)∣Pi′g′∈G}\mathcal{G}(i) = \{g \in A(i) \mid P_{i' g'} \in \mathcal{G}\} (where i′,g′i', g' denote the original dataset indices of augmented instances ii and gg):

    1. Contrastive Loss with Mixup (LMIX\mathcal{L}^{\text{MIX}}): For an individual sample embedding zi\mathbf{z}_i, the supervised contrastive loss over confident positive pairs is: Li(zi)=−1∣G(i)∣∑g∈G(i)log⁡exp⁡(zi⋅zg/τ)∑a∈A(i)exp⁡(zi⋅za/τ)\mathcal{L}_i(\mathbf{z}_i) = \frac{-1}{|\mathcal{G}(i)|} \sum_{g \in \mathcal{G}(i)} \log \frac{\exp(\mathbf{z}_i \cdot \mathbf{z}_g / \tau)}{\sum_{a \in A(i)} \exp(\mathbf{z}_i \cdot \mathbf{z}_a / \tau)} where τ>0\tau > 0 is a temperature hyperparameter (τ=0.1\tau = 0.1). If G(i)=∅\mathcal{G}(i) = \emptyset, unsupervised contrastive learning is applied. With Mixup interpolation xi=λxa+(1−λ)xbx_i = \lambda x_a + (1-\lambda)x_b (where λ∼Beta(αm,αm)\lambda \sim \text{Beta}(\alpha_m, \alpha_m) and αm=1\alpha_m = 1), the mixed contrastive loss is: LiMIX(zi)=λLa(zi)+(1−λ)Lb(zi),LMIX=∑i∈ILiMIX(zi)\mathcal{L}_i^{\text{MIX}}(\mathbf{z}_i) = \lambda \mathcal{L}_a(\mathbf{z}_i) + (1-\lambda) \mathcal{L}_b(\mathbf{z}_i), \quad \mathcal{L}^{\text{MIX}} = \sum_{i \in I} \mathcal{L}_i^{\text{MIX}}(\mathbf{z}_i)

    2. Classification Loss on Confident Examples (LCLS\mathcal{L}^{\text{CLS}}): LCLS=∑(xi,y~i)∈Tℓ(p^(xi),y~i)\mathcal{L}^{\text{CLS}} = \sum_{(x_i, \tilde{y}_i) \in \mathcal{T}} \ell(\hat{\mathbf{p}}(x_i), \tilde{y}_i) where ℓ\ell is the standard cross-entropy loss and p^(xi)\hat{\mathbf{p}}(x_i) is the classifier head output.

    3. Classifier Pairwise Similarity Loss (LSIM\mathcal{L}^{\text{SIM}}): LSIM=∑i∈I∑j∈A(i)ℓ(p^(xi)⊤p^(xj),I[Pi′j′∈G])\mathcal{L}^{\text{SIM}} = \sum_{i \in I} \sum_{j \in A(i)} \ell\left(\hat{\mathbf{p}}(x_i)^\top \hat{\mathbf{p}}(x_j), \mathbb{I}[P_{i' j'} \in \mathcal{G}]\right)

  5. Knowl 5 — Experimental Setup and Hyperparameter Configurations

    experimental setup

    The empirical evaluation of Sel-CL and Sel-CL+ uses the following configurations:

    • Datasets:
      • CIFAR-10 & CIFAR-100: 50k training and 10k test images (32×32×332 \times 32 \times 3). Evaluated under symmetric label noise (uniform random label flipping at 20%, 50%, 80%, 90%) and asymmetric label noise (class-conditional flipping at 10%, 20%, 30%, 40%: for CIFAR-10: truck →\rightarrow automobile, bird →\rightarrow airplane, deer →\rightarrow horse, cat ↔\leftrightarrow dog; for CIFAR-100: circular flips within super-classes).
      • WebVision-50: First 50 classes of Google subset of WebVision (2.4M web images total), tested on WebVision validation and ImageNet ILSVRC12 validation sets.
    • Architectures and Optimization:
      • CIFAR-10/100: PreAct ResNet-18 backbone, SGD optimizer (momentum 0.9, weight decay 10−410^{-4}, batch size 128). Pre-training runs for 250 epochs (1 epoch Uns-CL warm-up) with initial learning rate 0.1, decayed by 10×10\times at epochs 125 and 200. Fine-tuning runs for 70 epochs with learning rate 0.001. MoCo queue size is set to 30k.
      • WebVision-50: Standard ResNet-18, SGD optimizer (momentum 0.9, weight decay 10−410^{-4}, batch size 64). Pre-training runs for 130 epochs (5 epochs warm-up) with initial learning rate 0.1, decayed by 10×10\times at epochs 80 and 105. Fine-tuning runs for 50 epochs with learning rate 0.001. MoCo queue size is 60k.
    • General Hyperparameters: Mixup parameter αm=1.0\alpha_m = 1.0, temperature τ=0.1\tau = 0.1, loss weights λc=1.0,λs=0.01\lambda_c = 1.0, \lambda_s = 0.01.
    • Noise Detection Fractiles (α,β)(\alpha, \beta):
      • CIFAR-10 Symmetric: α=50%,β=25%\alpha = 50\%, \beta = 25\%
      • CIFAR-10 Asymmetric: α=50%,β=25%\alpha = 50\%, \beta = 25\%
      • CIFAR-100 Symmetric: α=75%,β=35%\alpha = 75\%, \beta = 35\%
      • CIFAR-100 Asymmetric: α=25%,β=0%\alpha = 25\%, \beta = 0\%
      • WebVision-50: α=40%,β=0%\alpha = 40\%, \beta = 0\%
  6. Knowl 6 — Classification Accuracy on Synthetic Noisy CIFAR-10 and CIFAR-100

    data/table

    Test classification accuracy (%) averaged over the last 10 training epochs on CIFAR-10 and CIFAR-100 under symmetric noise (20%, 50%, 80%, 90%) and asymmetric noise (10%, 20%, 30%, 40%) using PreAct ResNet-18 backbone:

    Dataset CIFAR-10 CIFAR-100
    Noise Type Symmetric Asymmetric Symmetric Asymmetric
    Noise Rate 20% 50% 80% 90% 10% 20% 30% 40% 20% 50% 80% 90% 10% 20% 30% 40%
    Cross-Entropy 82.7 57.9 26.1 16.8 88.8 86.1 81.7 76.0 61.8 37.3 8.8 3.5 68.1 63.6 53.3 44.5
    Mixup 92.3 77.6 46.7 43.9 93.3 88.0 83.3 77.7 66.0 46.6 17.6 8.1 72.4 65.1 57.6 48.1
    Forward 83.1 59.4 26.2 18.8 90.4 86.7 81.9 76.7 61.4 37.3 9.0 3.4 68.7 63.2 54.4 45.3
    GCE 86.6 81.9 54.6 21.2 89.5 85.6 80.6 76.0 59.2 47.8 15.8 7.2 68.0 58.6 51.4 42.9
    P-correction 92.0 88.7 76.5 58.2 93.1 92.9 92.6 91.6 68.1 56.4 20.7 8.8 76.1 68.9 59.3 48.3
    M-correction 93.8 91.9 86.6 68.7 89.6 91.8 92.2 91.2 73.4 65.4 47.6 20.5 67.1 64.5 58.6 47.4
    DivideMix 95.0 93.7 92.4 74.2 93.8 93.2 92.5 91.4 74.8 72.1 57.6 29.2 69.5 69.2 68.3 51.0
    ELR 93.8 92.6 88.0 63.3 94.4 93.3 91.5 85.3 74.5 70.2 45.2 20.5 75.8 74.8 73.6 70.0
    GCE (Uns-CL init.) 90.0 89.3 73.9 36.5 91.1 87.3 82.2 78.1 68.1 53.3 22.1 8.9 70.2 60.2 52.6 44.1
    ELR (Uns-CL init.) 94.4 93.0 88.3 86.2 95.0 94.7 94.4 93.3 76.2 71.9 57.9 40.8 77.2 75.5 74.3 70.4
    MOIT+ 94.1 91.8 81.1 74.7 94.2 94.3 94.3 93.3 75.9 70.6 47.6 41.8 77.4 76.4 75.1 74.0
    Sel-CL+ 95.5 93.9 89.2 81.9 95.6 95.2 94.5 93.4 76.5 72.4 59.6 48.8 78.7 77.5 76.4 74.2

    Sel-CL+ achieves state-of-the-art or competitive performance across all noise ratios, showing particularly pronounced gains under extreme symmetric noise (48.8% vs. 41.8% for MOIT+ and 29.2% for DivideMix on CIFAR-100 at 90% noise) and asymmetric noise regimes.

  7. Knowl 7 — Classification Accuracy on Real-World Noisy WebVision-50 and ILSVRC2012

    data/table

    Top-1 and top-5 accuracy (%) on the WebVision-50 validation set and the ImageNet ILSVRC2012 validation set (trained on WebVision-50 with standard ResNet-18):

    Methods WebVision ILSVRC12
    top-1 top-5 top-1 top-5
    Forward 61.12 82.68 57.36 82.36
    Decoupling 62.54 84.74 58.26 82.26
    D2L 62.68 84.00 57.80 81.36
    MentorNet 63.00 81.40 57.80 79.92
    Co-teaching 63.58 85.20 61.48 84.70
    Iterative-CV 65.24 85.34 61.60 84.98
    DivideMix 77.32 91.64 75.20 90.84
    ELR 76.26 91.26 68.71 87.84
    ELR+ 77.78 91.68 70.29 89.76
    ELR (Uns-CL init.) 79.93 92.00 71.23 88.23
    ProtoMix 76.3 91.5 73.3 91.2
    MoPro 77.59 – 76.31 –
    NGC 79.16 91.84 74.44 91.04
    Sel-CL+ 79.96 92.64 76.84 93.04

    Sel-CL+ outperforms all baselines on both the WebVision validation set (79.96% top-1 / 92.64% top-5) and the out-of-distribution ImageNet ILSVRC12 validation set (76.84% top-1 / 93.04% top-5), verifying robustness to real-world label noise.

  8. Knowl 8 — Representation Quality under Label Noise via Weighted KNN Evaluation

    data/table

    The representation quality of pre-trained encoders without fine-tuning is evaluated using a weighted KK-nearest neighbor (K=200K = 200) classifier on CIFAR-100:

    Methods Clean Symmetric Asymmetric
    0% 20% 80% 10% 40%
    Uns-CL 56.23 – – – –
    Sup-CL 72.66 58.32 41.00 71.11 68.00
    MOIT 77.48 67.42 55.58 74.86 72.60
    Sel-CL 77.94 75.36 62.49 76.77 72.71

    While standard Supervised Contrastive Learning (Sup-CL) suffers severe degradation under label noise (dropping from 72.66% clean to 41.00% under 80% symmetric noise, below Uns-CL's 56.23%), Sel-CL maintains high representation quality across all noise levels (e.g., 62.49% under 80% symmetric noise and 72.71% under 40% asymmetric noise).

  9. Knowl 9 — Ablation of Pair Selection Strategies and Architecture Components

    empirical result

    Ablation studies on CIFAR-10 and CIFAR-100 demonstrate the specific contributions of each module in Sel-CL:

    1. Effect of Selection Strategies (CIFAR-10 / CIFAR-100 under 40% asymmetric noise, evaluated via weighted KNN):

      • All examples and all pairs: 90.58% / 68.66%
      • Confident examples T\mathcal{T} and pairs G′\mathcal{G}': 90.64% / 70.25%
      • Confident examples T\mathcal{T} and pairs G′∪G′′\mathcal{G}' \cup \mathcal{G}'' (Sel-CL): 92.97% / 72.71%
      • Clean examples and associated pairs: 94.21% / 71.45%
      • Clean examples and all pairs: 94.76% / 73.43%
      • Clean examples and clean pairs (Oracle): 95.52% / 76.56% Adding similarity-derived confident pairs G′′\mathcal{G}'' provides significant improvements over using G′\mathcal{G}' alone by recovering true positive pairs among misclassified examples.
    2. Component Ablation on CIFAR-100 (reported as Test Acc / KNN Acc for Sel-CL, and Test Acc for Sel-CL+):

      • Sel-CL w/o Mixup Data Augmentation: 70.3% / 70.6% (Sym 20%), 64.2% / 66.2% (Asym 40%)
      • Sel-CL w/o MoCo Trick: 73.3% / 74.1% (Sym 20%), 69.2% / 71.5% (Asym 40%)
      • Sel-CL w/o Selection: 67.2% / 68.9% (Sym 20%), 49.9% / 68.7% (Asym 40%)
      • Sel-CL w/o Classifier Learning (LCLS\mathcal{L}^{\text{CLS}}): -- / 69.9% (Sym 20%), -- / 70.2% (Asym 40%)
      • Sel-CL w/o Similarity Loss (LSIM\mathcal{L}^{\text{SIM}}): 74.5% / 74.9% (Sym 20%), 71.8% / 72.5% (Asym 40%)
      • Full Sel-CL: 74.9% / 75.4% (Sym 20%), 72.0% / 72.7% (Asym 40%)
      • Full Sel-CL+: 76.5% (Sym 20%), 74.2% (Asym 40%)
  10. Knowl 10 — Compatibility with Warm-Up and Advanced Fine-Tuning Methods

    empirical result

    Sel-CL displays strong robustness to the warm-up scheme and enhances advanced downstream fine-tuning frameworks:

    1. Warm-up Method Invariance: Testing Sel-CL+ on CIFAR-10 and CIFAR-100 shows virtually identical final test accuracy whether warming up with unsupervised contrastive learning (Uns-CL) or supervised contrastive learning (Sup-CL):

      • CIFAR-10 Sym 20% / Sym 90% / Asym 40%: Uns-CL warmup achieves 95.5% / 81.9% / 93.4%; Sup-CL warmup achieves 95.5% / 81.6% / 93.4%.
      • CIFAR-100 Sym 20% / Sym 90% / Asym 40%: Uns-CL warmup achieves 76.5% / 48.8% / 74.2%; Sup-CL warmup achieves 76.8% / 51.4% / 74.5%.
    2. Integration with Advanced Fine-Tuning Frameworks: Pre-training with Sel-CL consistently boosts the performance of complex downstream fine-tuning methods compared to random initialization or Uns-CL initialization:

      • DivideMix (Sel-CL init): 96.3% (CIFAR-10 Sym 20%), 91.6% (CIFAR-10 Asym 40%), 78.7% (CIFAR-100 Sym 20%), 55.2% (CIFAR-100 Asym 40%), compared to standard DivideMix (95.7%, 92.1%, 76.9%, 53.8%) and DivideMix with Uns-CL init (96.2%, 90.8%, 78.3%, 52.9%).
      • ELR+ (Sel-CL init): 95.2% (CIFAR-10 Sym 20%), 94.6% (CIFAR-10 Asym 40%), 77.7% (CIFAR-100 Sym 20%), 72.9% (CIFAR-100 Asym 40%), compared to standard ELR+ (94.6%, 93.0%, 77.5%, 72.2%) and ELR+ with Uns-CL init (94.8%, 94.3%, 77.7%, 72.3%).
  11. Knowl 11 — Limitations of Selective-Supervised Contrastive Learning

    limitation

    The authors identify two primary limitations in the proposed Sel-CL framework:

    1. High Memory and Data Augmentation Requirements: As a contrastive learning method, representation quality depends heavily on adequate data augmentation and a large set of negative samples, necessitating memory banks or large queue sizes (e.g., 30k to 60k items via the MoCo mechanism), which imposes substantial storage and memory demands on hardware.
    2. Computational Overhead of K-Nearest Neighbors: Constructing pseudo-labels and evaluating neighborhood agreements requires querying KK-nearest neighbors (K=250K = 250) across the dataset embeddings at each epoch, introducing significant computational cost that requires fast approximate KNN algorithms when scaling to large datasets.

Coverage note — No substantial contributed material was omitted from the paper.

References

  1. 1.Eric Arazo, Diego Ortego, Paul Albert, Noel E. O'Connor, and Kevin McGuinness. Unsupervised label noise modeling and loss correction. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, ICML, pages 312–321, 2019. 6
  2. 2.Yingbin Bai and Tongliang Liu. Me-momentum: Extracting hard confident examples from noisily labeled data. In ICCV, 2021. 5
  3. 3.Haolan Chen, Fred X. Han, Di Niu, Dong Liu, Kunfeng Lai, Chenglin Wu, and Yu Xu. MIX: multi-channel information crossing for text matching. In KDD, pages 110–119, 2018. 1
  4. 4.Ling-Hao Chen, He Li, and Wenhao Yang. Anomman: Detect anomaly on multi-view attributed networks. arXiv preprint arXiv:2201.02822, 2022. 2
  5. 5.Pengfei Chen, Guangyong Chen, Junjie Ye, Pheng-Ann Heng, et al. Noise against noise: stochastic label noise helps combat inherent label noise. In ICLR, 2021. 2
  6. 6.Pengfei Chen, Benben Liao, Guangyong Chen, et al. Understanding and utilizing deep neural networks trained with noisy labels. In ICML, pages 1833–1841, 2019. 7
  7. 7.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, pages 1597–1607, 2020. 1, 2, 3, 4, 5, 6, 7, 8
  8. 8.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton. Big self-supervised models are strong semi-supervised learners. In NeurIPS, pages 22243–22255, 2020. 1, 2
  9. 9.De Cheng, Tongliang Liu, Yixiong Ning, Nannan Wang, Bo Han, Gang Niu, Xinbo Gao, and Masashi Sugiyama. Instance-dependent label-noise learning with manifoldregularized transition matrix estimation. In CVPR, 2022. 2
  10. 10.Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong, Xing Sun, and Yang Liu. Learning with instance-dependent label noise: A sample sieve approach. In ICLR, 2021. 2
  11. 11.Shiming Ge, Chunhui Zhang, Shikun Li, Dan Zeng, and Dacheng Tao. Cascaded correlation refinement for robust deep tracking. TNNLS, 32(3):1276–1288, 2021. 1
  12. 12.Aritra Ghosh, Himanshu Kumar, and P. S. Sastry. Robust loss functions under label noise for deep neural networks. In AAAI, pages 1919–1925, 2017. 2
  13. 13.Aritra Ghosh and Andrew Lan. Contrastive Learning Improves Model Robustness Under Label Noise. In CVPR Workshop, 2021. 1, 3, 6
  14. 14.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT press, 2016. 1
  15. 15.Bo Han, Gang Niu, Xingrui Yu, Quanming Yao, Miao Xu, Ivor Tsang, and Masashi Sugiyama. Sigua: Forgetting may make learning with noisy labels more robust. In ICML, pages 4006–4016, 2020. 1, 2
  16. 16.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. Coteaching: Robust training of deep neural networks with extremely noisy labels. In NeurIPS, pages 8536–8546, 2018. 2, 3, 7
  17. 17.Jiangfan Han, Ping Luo, and Xiaogang Wang. Deep selflearning from noisy labels. In ICCV, pages 5138–5147, 2019. 2
  18. 18.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, pages 9726–9735, 2020. 1, 2, 5
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 1
  20. 20.Dan Hendrycks, Mantas Mazeika, Duncan Wilson, and Kevin Gimpel. Using trusted data to train deep networks on labels corrupted by severe noise. In NeurIPS, pages 10477–10486, 2018. 2
  21. 21.Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty. In ICLR, 2020. 7
  22. 22.Yen-chang Hsu, Zhaoyang Lv, Joel Schlosser, Phillip Odom, and Zsolt Kira. Multi-class classification class labels without multi-class lables. In ICLR, 2019. 4
  23. 23.Wei Hu, Zhiyuan Li, and Dingli Yu. Simple and effective regularization methods for training on noisily labeled data with generalization guarantee. In ICLR, 2020. 2
  24. 24.Jiabo Huang, Qi Dong, Shaogang Gong, and Xiatian Zhu. Unsupervised deep learning by neighbourhood discovery. In ICML, pages 2849–2858, 2019. 6
  25. 25.Jinchi Huang, Lie Qu, Rongfei Jia, and Binqiang Zhao. O2unet: A simple noisy label detection approach for deep neural networks. In ICCV, pages 3326–3334, 2019. 2
  26. 26.Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In ICML, pages 2309–2318, 2018. 2, 7
  27. 27.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Dilip Krishnan, and Ce Liu. Supervised contrastive learning. In NeurIPS, pages 1–23, 2020. 2, 4, 5, 7, 8
  28. 28.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5
  29. 29.Kimin Lee, Sukmin Yun, Kibok Lee, Honglak Lee, Bo Li, et al. Robust inference via generative classifiers for handling noisy labels. In ICML, pages 3763–3772, 2019. 1
  30. 30.Junnan Li, Richard Socher, and Steven C. H. Hoi. DivideMix: learning with noisy labels as semi-supervised learning. In ICLR, 2020. 2, 5, 6, 7, 8
  31. 31.Junnan Li, Caiming Xiong, and Steven Hoi. Learning from noisy data with robust representation learning. In ICCV, 2021. 2, 7, 8
  32. 32.Junnan Li, Caiming Xiong, and Steven C.H. Hoi. MoPro: Webly supervised learning with momentum prototypes. In ICLR, 2021. 2, 3, 7
  33. 33.Shikun Li, Shiming Ge, Yingying Hua, Chunhui Zhang, Hao Wen, Tengfei Liu, and Weiqiang Wang. Coupled-view deep classifier learning from multiple noisy annotators. In AAAI, pages 4667–4674, 2020. 2
  34. 34.Shikun Li, Tongliang Liu, Jiyong Tan, Dan Zeng, and Shiming Ge. Trustable co-label learning from multiple noisy annotators. TMM, 2021. 2
  35. 35.Wen Li, Limin Wang, Wei Li, Eirikur Agustsson, and Luc Van Gool. Webvision database: Visual learning and understanding from web data. arXiv preprint arXiv:1708.02862, 2017. 5, 7
  36. 36.Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li. Learning from noisy labels with distillation. In ICCV, pages 1928–1936, 2017. 1
  37. 37.Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda. Early-learning regularization prevents memorization of noisy labels. In NeurIPS, 2020. 2, 5, 6, 7, 8
  38. 38.Tongliang Liu and Dacheng Tao. Classification with noisy labels by importance reweighting. TPAMI, 38(3):447–461, 2016. 2
  39. 39.Xingjun Ma, Yisen Wang, Michael E. Houle, Shuo Zhou, Sarah M. Erfani, Shu-Tao Xia, Sudanthi N. R. Wijewickrema, and James Bailey. Dimensionality-driven learning with noisy labels. In ICML, pages 3361–3370, 2018. 7
  40. 40.Eran Malach and Shai Shalev-Shwartz. Decoupling” when to update” from” how to update”. In NeurIPS, pages 960–970, 2017. 7
  41. 41.Diego Ortego, Eric Arazo, Paul Albert, Noel E. O'Connor, et al. Multi-objective interpolation training for robustness to label noise. In CVPR, 2021. 2, 3, 4, 5, 6, 8
  42. 42.Diego Ortego, Eric Arazo, Paul Albert, Noel E. O'Connor, and Kevin McGuinness. Towards Robust Learning with Different Label Noise Distributions. In ICPR, pages 7020–7027, 2021. 5
  43. 43.Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu. Making deep neural networks robust to label noise: A loss correction approach. In CVPR, pages 2233–2241, 2017. 6, 7
  44. 44.Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun. Learning to reweight examples for robust deep learning. In ICML, pages 4331–4340, 2018. 2
  45. 45.Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. Meta-weight-net: Learning an explicit mapping for sample weighting. In NeurIPS, pages 1917–1928, 2019. 2
  46. 46.Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa. Joint optimization framework for learning with noisy labels. In CVPR, pages 5552–5560, 2018. 2, 5
  47. 47.Ruijia Wang, Shuai Mou, Xiao Wang, Wanpeng Xiao, Qi Ju, Chuan Shi, and Xing Xie. Graph structure estimation neural networks. In WWW, pages 342–353, 2021. 2
  48. 48.Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip-driven referring image segmentation. arXiv preprint arXiv:2111.15174, 2021. 2
  49. 49.Tong Wei, Jiang-Xin Shi, Wei-Wei Tu, and Yu-Feng Li. Robust long-tailed learning under label noise. arXiv preprint arXiv:2108.11569, 2021. 2
  50. 50.Songhua Wu, Xiaobo Xia, Tongliang Liu, Bo Han, Mingming Gong, Nannan Wang, Haifeng Liu, and Gang Niu. Class2Simi: A New Perspective on Learning with Label Noise. In ICML, 2021. 4
  51. 51.Zhi-Fan Wu, Tong Wei, Jianwen Jiang, Chaojie Mao, Mingqian Tang, and Yu-Feng Li. NGC: A unified framework for learning with open-world noisy data. In ICCV, 2021. 2, 5, 7, 8
  52. 52.Xiaobo Xia, Tongliang Liu, Bo Han, Chen Gong, Nannan Wang, Zongyuan Ge, and Yi Chang. Robust early-learning: Hindering the memorization of noisy labels. In ICLR, 2021. 1
  53. 53.Xiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang, Mingming Gong, Haifeng Liu, Gang Niu, Dacheng Tao, and Masashi Sugiyama. Part-dependent label noise: Towards instance-dependent label noise. In NeurIPS, 2020. 2
  54. 54.Xiaobo Xia, Tongliang Liu, Nannan Wang, Bo Han, Chen Gong, Gang Niu, and Masashi Sugiyama. Are anchor points really indispensable in label-noise learning? In NeurIPS, pages 6835–6846, 2019. 2
  55. 55.Tong Xiao, Tian Xia, Yi Yang, Chang Huang, and Xiaogang Wang. Learning from massive noisy labeled data for image classification. In CVPR, pages 2691–2699, 2015. 1
  56. 56.Hansi Yang, Quanming Yao, Bo Han, Gang Niu, Hansi Yang, Bo Han, Gang Niu, and James Kwok. Searching to Exploit Memorization Effect in Learning from Corrupted Labels. In ICML, 2020. 2
  57. 57.Shuo Yang, Lu Liu, and Min Xu. Free lunch for few-shot learning: Distribution calibration. In ICLR, 2021. 1
  58. 58.Shuo Yang, Peize Sun, Yi Jiang, Xiaobo Xia, Ruiheng Zhang, Zehuan Yuan, Changhu Wang, Ping Luo, and Min Xu. Objects in semantic topology. In ICLR, 2022. 1
  59. 59.Kun Yi and Jianxin Wu. Probabilistic end-to-end noise correction for learning with noisy labels. In CVPR, pages 7017–7025, 2019. 6
  60. 60.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In ICLR, 2017. 1
  61. 61.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018. 2, 4, 6
  62. 62.Kangkai Zhang, Chunhui Zhang, Shikun Li, Dan Zeng, and Shiming Ge. Student network learning via evolutionary knowledge distillation. TCSVT, 2021. 1
  63. 63.Yikai Zhang, Songzhu Zheng, Pengxiang Wu, Mayank Goswami, and Chao Chen. Learning with feature-dependent label noise: A progressive approach. In ICLR, 2021. 2
  64. 64.Zhilu Zhang and Mert Sabuncu. Generalized cross entropy loss for training deep neural networks with noisy labels. In NeurIPS, pages 8778–8788, 2018. 2, 6
  65. 65.Evgenii Zheltonozhskii, Chaim Baskin, Avi Mendelson, Alex M. Bronstein, and Or Litany. Contrast to divide: Selfsupervised pre-training for learning with noisy labels. In WACV, 2022. 1, 3, 5, 6, 8
  66. 66.Songzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami, Dimitris Metaxas, and Chao Chen. Error-bounded correction of noisy labels. In ICML, pages 11447–11457, 2020. 2

Citation

MLA
Li, S., et al. “Selective-Supervised Contrastive Learning with Noisy Labels”. arXiv, 2022, http://arxiv.org/abs/2203.04181v1.
APA
Li, S., Xia, X., Ge, S., & Liu, T. (2022). Selective-Supervised Contrastive Learning with Noisy Labels. arXiv. http://arxiv.org/abs/2203.04181v1
Chicago
Li, S., X. Xia, S. Ge, and T. Liu. 2022. “Selective-Supervised Contrastive Learning with Noisy Labels”. arXiv. http://arxiv.org/abs/2203.04181v1.
Harvard
Li, S. et al. (2022) “Selective-Supervised Contrastive Learning with Noisy Labels”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.04181v1.
Vancouver
1. Li S, Xia X, Ge S, Liu T (2022) Selective-Supervised Contrastive Learning with Noisy Labels. arXiv

BibTeX

@article{li2022selective,
  title = {Selective-Supervised Contrastive Learning with Noisy Labels},
  author = {Li, Shikun and Xia, Xiaobo and Ge, Shiming and Liu, Tongliang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.04181v1},
  eprint = {2203.04181}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE