Selective-Supervised Contrastive Learning with Noisy Labels
Shikun LiXiaobo XiaShiming GeTongliang Liu
Proposes a selective-supervised contrastive learning framework that dynamically identifies reliable sample pairs without requiring prior knowledge of label noise rates, preventing corrupted annotations from degrading representation quality.
Deep learning models require massive datasets to achieve high accuracy, but obtaining clean, manually verified labels is costly and time-consuming. While web scraping and user tagging offer cheaper alternatives, they introduce noisy, incorrect labels that corrupt internal model representations and degrade real-world decision-making. Existing contrastive learning approaches attempt to build representations by comparing data pairs, but noisy labels inject incorrect pair relationships that undermine performance, especially when noise rates cannot be estimated in advance.
The article demonstrates a novel framework called selective-supervised contrastive learning to learn robust visual representations from noisily labeled data without prior knowledge of the exact noise rate. The primary objective is to evaluate whether dynamically filtering and selecting confident data pairs during training can shield deep neural networks from the harmful effects of mislabeled training data.
The researchers developed a two-stage approach that iteratively selects high-confidence examples and pairs during pre-training. First, confident individual examples are identified by checking agreement between learned representations and given labels across nearest neighbors. Next, confident pairs are constructed from these clean examples and augmented with additional pairs that share high representation similarity, even if their nominal class labels are wrong. The pre-trained model is then fine-tuned on confident examples using standard classification techniques. The authors benchmarked this framework across simulated noise conditions using standard image recognition datasets (CIFAR-10 and CIFAR-100) and evaluated real-world performance on the WebVision dataset.
The findings confirm that the proposed method consistently outperforms existing state-of-the-art baselines. Under heavy noise conditions of 80% symmetric noise on CIFAR-100, the method improved feature representation quality to 62.49% weighted accuracy, compared to 55.58% achieved by existing contrastive methods and 41.00% by standard supervised contrastive learning. Across simulated benchmark classifications, the framework delivered top accuracy, excelling particularly under asymmetric noise where mislabeling occurs between semantically related classes (achieving up to 74.2% accuracy under 40% asymmetric noise on CIFAR-100). On the real-world WebVision benchmark, the approach achieved the highest validation performance with a top-1 accuracy of 79.96% and top-5 accuracy of 92.64%, surpassing competing methods.
These results demonstrate that organizations can significantly lower data curation costs by effectively training reliable computer vision models directly on cheaper, web-scraped, or crowdsourced data. By focusing on representation similarity rather than relying on strict label correctness, the approach reduces the risk of model failure caused by systematic human labeling errors in real-world deployment pipelines.
Organizations training vision models on unverified datasets should consider adopting confident pair selection workflows to enhance robustness without paying for expensive re-labeling campaigns. Engineering teams can integrate the pre-training stage with existing fine-tuning pipelines to gain immediate accuracy improvements. Future development should focus on extending this framework to other operational domains, including object detection and text matching.
The main limitations are computational: contrastive pair-wise learning demands large batch sizes or extensive memory queues, and nearest-neighbor calculations increase processing overhead during training. Despite these computational costs, the empirical results provide high confidence in the method's ability to maintain robustness across varying noise types and levels.
- Paper: Supervised Contrastive Learning, Prannay Khosla et al. (2020). Introduces the supervised contrastive learning (SupCon) framework that Sel-CL explicitly adapts and modifies to handle noisy label pairs.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). Establishes modern sample-selection and semi-supervised paradigms for handling noisy labels that motivate Sel-CL's pair selection strategy.
- Paper: Co-teaching: Robust training of deep neural networks with extremely noisy labels, Bo Han et al. (2018). Provides the foundational sample-filtering and peer-network selection principles used across robust learning under heavy label noise.
- Paper: A Simple Framework for Contrastive Learning of Visual Representations, Ting Chen et al. (2020). Defines the core normalized temperature-scaled contrastive loss architecture and data augmentation framework underlying supervised contrastive methods.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). Presents foundational formulations for learning under label noise without ground-truth knowledge of noise rates.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). Offers a comprehensive survey of label-noise taxonomy, sample-selection strategies, and robust loss functions relevant to representation learning.
- Paper: Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations, Jiaheng Wei et al. (2022). Introduces real-world human annotation noise benchmarks (CIFAR-N) that systematically test the limits of robust learning methods like Sel-CL beyond synthetic noise.
- Paper: Robust Training under Label Noise by Over-parameterization, Sheng Liu et al. (2022). Explores an alternative over-parameterization paradigm for isolating label corruption during deep neural network optimization.
- Paper: Debiased Learning from Naturally Imbalanced Pseudo-Labels, Xudong Wang et al. (2022). Investigates how pseudo-labeling and sample-selection mechanisms suffer from class imbalances and introduces debiasing methods to rectify them.
