Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Jiaheng WeiZhaowei ZhuHao ChengTongliang LiuGang NiuYang Liu

article2022ICLR374 citations

Presents the CIFAR-10N and CIFAR-100N benchmarks with real-world human-annotated noise and verified ground truth, demonstrating that practical annotation errors are instance-dependent and establishing a standardized testbed for evaluating noise-resistant learning algorithms.

Listen

Data labeling in deep learning is expensive and prone to human error, resulting in noisy datasets that hinder model training. Machine learning research on learning with noisy labels has predominantly relied on synthetic label noise that assumes errors occur uniformly or strictly across categories, independent of image content. Existing real-world noisy benchmarks are often excessively large, lack ground-truth clean verification labels, or introduce artificial collection biases, preventing controlled and fair evaluations.

The article aims to introduce and analyze two accessible, verified real-world noisy label benchmark datasets—CIFAR-10N and CIFAR-100N (collectively termed CIFAR-N)—and systematically evaluate how real human annotation errors differ from theoretical synthetic noise models.

To conduct this evaluation, the researchers collected human annotations via Amazon Mechanical Turk for all 50,000 training images in both CIFAR-10 and CIFAR-100. For CIFAR-10N, three independent human annotations were collected per image to create aggregate, individual random, and worst-case noisy sets. For CIFAR-100N, workers provided both coarse and fine labels using an interactive hierarchical interface. The authors then conducted statistical hypothesis testing to evaluate feature dependency, analyzed neural network memorization patterns, and benchmarked a wide suite of noise-robust machine learning algorithms.

The analysis revealed several critical findings. First, human noise is statistically feature-dependent rather than purely class-dependent, rejecting the standard independent noise hypothesis at high significance levels (p-value of 1.8e-36 for CIFAR-10N). Second, error distributions are heavily skewed: annotators confuse visually similar categories (such as trucks with automobiles or worms with snakes) and introduce substantial class imbalances, with some categories receiving four times more annotations than others. Third, CIFAR-100N revealed an inherent multi-label phenomenon where human annotators chose prominent secondary subjects in images, such as labeling a person holding a flatfish as 'man' rather than 'flatfish'. Finally, deep neural networks consistently showed lower classification accuracy and faster memorization and overfitting on human noise compared to synthetic noise with identical transition matrices, exhibiting performance drops of up to 9% under high human noise.

These findings indicate that existing theoretical noise models and standard algorithmic solutions developed on synthetic data do not accurately capture real-world human annotation behaviors. Consequently, models developed under synthetic assumptions risk underperforming when deployed in practical, human-annotated machine learning pipelines. Furthermore, the presence of multiple legitimate objects within single-label datasets highlights a fundamental flaw in evaluating image classifiers against rigid single-label ground truths.

Researchers and machine learning practitioners should shift benchmark evaluations toward realistic, instance-dependent noise settings using open resources like CIFAR-N. Algorithmic development should prioritize methods combining semi-supervised techniques, data augmentation, and dual-network architectures (such as Divide-Mix and ELR+), which showed the highest resilience to human noise. Additionally, data collection workflows should incorporate multi-label handling to accommodate images with multiple valid subjects.

While the article provides high confidence through rigorous hypothesis testing and multi-run benchmarking across dozens of methods, limitations remain regarding image resolution, as the study focuses on 32x32 pixel images where low resolution inherently exacerbates human error. Further work across higher-resolution imagery and non-visual domains is needed to determine how these exact noise distributions generalize across broader machine learning settings.

No sufficiently relevant recommendations were found.

Cover for Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations

Abstract

Existing research on learning with noisy labels mainly focuses on synthetic label noise. Synthetic noise, though has clean structures which greatly enabled statistical analyses, often fails to model real-world noise patterns. The recent literature has observed several efforts to offer real-world noisy datasets, yet the existing efforts suffer from two caveats: (1) The lack of ground-truth verification makes it hard to theoretically study the property and treatment of real-world label noise; (2) These efforts are often of large scales, which may result in unfair comparisons of robust methods within reasonable and accessible computation power. To better understand real-world label noise, it is crucial to build controllable and moderate-sized real-world noisy datasets with both ground-truth and noisy labels. This work presents two new benchmark datasets CIFAR-10N, CIFAR-100N, equipping the training datasets of CIFAR-10, CIFAR-100 with human-annotated real-world noisy labels we collected from Amazon Mechanical Turk. We quantitatively and qualitatively show that real-world noisy labels follow an instance-dependent pattern rather than the classically assumed and adopted ones (e.g., class-dependent label noise). We then initiate an effort to benchmarking a subset of the existing solutions using CIFAR-10N and CIFAR-100N. We further proceed to study the memorization of correct and wrong predictions, which further illustrates the difference between human noise and class-dependent synthetic noise. We show indeed the real-world noise patterns impose new and outstanding challenges as compared to synthetic label noise. These observations require us to rethink the treatment of noisy labels, and we hope the availability of these two datasets would facilitate the development and evaluation of future learning with noisy label solutions. Datasets and leaderboards are available at this http URL.

Table of Contents

  • 1 Introduction
  • 1.1 Related works
  • 2 Synthetic label noise
  • 2.1 Class-dependent label noise
  • 2.2 Instance-dependent label noise
  • 3 Human annotated noisy labels on CIFAR-10, CIFAR-100
  • 3.1 CIFAR-10N real-world noisy label benchmark
  • 3.2 CIFAR-100N real-world noisy label benchmark
  • 4 Preliminary observations on CIFAR-10N, CIFAR-100N
  • 4.1 The noisy label distribution
  • 4.2 Human noisy labels v.s. synthetic noisy labels
  • 4.2.1 A qualitative aspect
  • 4.2.2 A quantitative aspect
  • 5 Learning with CIFAR-10N and CIFAR-100N
  • 5.1 Performance comparisons on CIFAR-10N and CIFAR-100N
  • 5.2 Memorization effects
  • 6 Conclusions
  • References
  • A CIFAR-10N real-world noisy label benchmark
  • A.1 Case studies before the formal collection
  • A.2 Dataset collection
  • A.3 The workers’ behaviors
  • A.4 More detailed dataset statistics
  • B CIFAR-100N real-world noisy label benchmark
  • B.1 Case studies before the formal collection
  • B.2 Dataset collection
  • B.3 More detailed dataset statistics
  • C Comparisons of the label collection procedure
  • C.1 Comparisons between CIFAR and CIFAR-N
  • C.2 Comparisons between CIFAR-10H and CIFAR-N
  • D Hypothesis testing of CIFAR-N
  • D.1 Hypothesis testing of CIFAR-10N
  • D.2 Hypothesis testing of CIFAR-100N
  • E Additional experiment results
  • E.1 Performance comparisons on CIFAR-N datasets
  • E.2 Performance comparisons on synthetic CIFAR datasets
  • E.3 Experiment details on CIFAR-10N and CIFAR-100N
  • E.4 Computing infrastructure

Knowls

  1. Knowl 1 — CIFAR-N Human-Annotated Noisy Label Benchmarks

    model/method

    CIFAR-N consists of two real-world noisy label benchmark datasets, CIFAR-10N and CIFAR-100N, created by collecting human annotations for the training sets of CIFAR-10 and CIFAR-100 via Amazon Mechanical Turk (MTurk) without altering the test sets.

    1. CIFAR-10N:

      • Collection: The 50,000 32×3232 \times 32 training images were split into 10 batches of 500 Human Intelligence Tasks (HITs), with 10 images per HIT. Each HIT was independently annotated by three distinct human workers. 747 workers contributed.
      • Label Sets:
        • Aggregate: Majority voting among the three human annotations per image (random tie-breaking); noise rate is 9.03%9.03\%.
        • Random 1, Random 2, Random 3: The first, second, and third submitted human annotations for each image, yielding noise rates of 17.23%17.23\%, 18.12%18.12\%, and 17.64%17.64\%, respectively.
        • Worst: Contains an incorrect label if any of the three annotators made an error (chosen randomly among wrong annotations), and the clean label only if all three agreed; noise rate is 40.21%40.21\%.
      • Consensus: 60.27%60.27\% of all training images received unanimous agreement across all three workers.
    2. CIFAR-100N:

      • Collection: The 50,000 training images were split into 10 batches of 1,000 HITs (5 images per HIT, scaled to 96×9696 \times 96 for labeling, one worker per HIT). Annotators followed a two-step hierarchy: first selecting one of 20 super-classes, then choosing among 4 to 6 fine classes within that super-class with reference images provided.
      • Noise Statistics: Provides both coarse (super-class) and fine noisy labels. Overall noise levels are 25.60%25.60\% for coarse labels and 40.20%40.20\% for fine labels. Among fine-label noise, 25%25\% of noisy labels fall outside the clean super-class, while 15%15\% remain inside the correct clean super-class.
  2. Knowl 2 — Statistical Hypothesis Testing Framework for Instance-Dependent Label Noise

    model/method

    To quantitatively determine whether human label noise is feature-dependent (instance-dependent) or conditionally independent of image features (class-dependent), a two-sample statistical hypothesis test is formulated using feature clustering under local noise homogeneity.

    1. Feature Clustering: For clean class i∈[K]i \in [K] where [K]:={1,2,…,K}[K] := \{1, 2, \dots, K\}, image feature representations extracted from a clean-trained deep network (the pre-classification layer of ResNet-34) are partitioned into 5 clusters using kk-means clustering. Let Ii,ν\mathcal{I}_{i,\nu} denote the instance index set in cluster ν∈{1,…,5}\nu \in \{1, \dots, 5\} for clean class ii.

    2. Transition Probability Vectors: The empirical noise transition vector pi,ν∈[0,1]Kp_{i,\nu} \in [0, 1]^K for cluster ν\nu of class ii is calculated by counting noisy label frequencies: pi,ν[j]=1∣Ii,ν∣∑n∈Ii,νI(y~n=j)p_{i,\nu}[j] = \frac{1}{|\mathcal{I}_{i,\nu}|} \sum_{n \in \mathcal{I}_{i,\nu}} \mathbb{I}(\tilde{y}_n = j) where y~n\tilde{y}_n is the noisy label assigned to instance xnx_n.

    3. Test Statistics:

      • Let pi,νhumanp^{\text{human}}_{i,\nu} be the transition vector from human noisy labels.
      • Let pi,νsyntheticp^{\text{synthetic}}_{i,\nu} be the transition vector obtained from synthetic class-dependent noise generated using the identical expected transition matrix T=E[T(X)]T = \mathbb{E}[T(X)].
      • Let pi,νsynthetic′p^{\text{synthetic}\prime}_{i,\nu} be the transition vector obtained from the same synthetic noise under different cluster assignments induced by random data augmentations.
      • Define distance samples across all classes i∈[K]i \in [K] and clusters ν∈{1,…,5}\nu \in \{1, \dots, 5\}: di,ν(1):=∥pi,νhuman−pi,νsynthetic∥22d^{(1)}_{i,\nu} := \|p^{\text{human}}_{i,\nu} - p^{\text{synthetic}}_{i,\nu}\|_2^2 di,ν(2):=∥pi,νsynthetic′−pi,νsynthetic∥22d^{(2)}_{i,\nu} := \|p^{\text{synthetic}\prime}_{i,\nu} - p^{\text{synthetic}}_{i,\nu}\|_2^2
    4. Hypothesis Formulation:

      • H0H_0: {di,ν(1)}i∈[K],ν∈[5]\{d^{(1)}_{i,\nu}\}_{i \in [K], \nu \in [5]} and {di,ν(2)}i∈[K],ν∈[5]\{d^{(2)}_{i,\nu}\}_{i \in [K], \nu \in [5]} come from the same distribution (label noise is feature-independent).
      • H1H_1: {di,ν(1)}\{d^{(1)}_{i,\nu}\} and {di,ν(2)}\{d^{(2)}_{i,\nu}\} come from different distributions (label noise is feature-dependent).

    A two-sided Student's tt-test across 10 repeated augmentations evaluates this null hypothesis at significance level α=0.05\alpha = 0.05.

  3. Knowl 3 — Feature-Dependency and Structural Non-Uniformity of Real-World Human Label Noise

    empirical result

    Statistical testing and qualitative analysis on CIFAR-10N and CIFAR-100N establish that real-world human label noise is fundamentally instance-dependent (feature-dependent) rather than class-dependent:

    1. Quantitative Hypothesis Testing:

      • On CIFAR-10N, the two-sided tt-test across all classes yields p=1.8×10−36p = 1.8 \times 10^{-36}, rejecting the null hypothesis of feature-independence. For 9 out of 10 individual classes (all except "automobile" with p=0.12971p = 0.12971), the null hypothesis of feature-independence is rejected (p<0.05p < 0.05).
      • On CIFAR-100N, testing across all fine classes yields p=5.2×10−16p = 5.2 \times 10^{-16}. At the individual class level, approximately 50 fine classes exhibit statistically significant feature dependency (p<0.05p < 0.05), while roughly 50 classes can be approximated as class-dependent.
    2. Qualitative Noise Characteristics:

      • Annotation Imbalance: While ground-truth training class counts are balanced (5,000 per class in CIFAR-10, 500 per class in CIFAR-100), human annotations exhibit strong marginal imbalances. In CIFAR-100N, "Man" is assigned ≥750\ge 750 times while "Streetcar" has only ≈200\approx 200 annotations. In CIFAR-10N, annotators systematically prefer "automobile" over "truck" and "horse" over "deer".
      • Feature-Similarity Flipping: Label noise concentrates among semantically and visually similar classes (e.g., ≈20%\approx 20\% mutual confusion between "snake" and "worm", and high mutual confusion between "cockroach" and "beetle", "fox" and "wolf", and among "boy"-"baby"-"girl"-"man").
      • Multi-Label Co-occurrence: A notable portion of "noisy" labels reflect secondary objects co-occurring in the image rather than perception mistakes (e.g., an image labeled clean "flatfish" showing a person holding the fish is frequently annotated by humans as "man").
  4. Knowl 4 — Accelerated Feature Memorization Under Human Label Noise

    empirical result

    Deep neural networks fit and memorize wrong annotations significantly faster and to a greater extent under human label noise than under synthetic class-dependent noise of the identical noise transition matrix TT and noise level.

    When training ResNet-34 with standard Cross-Entropy loss on CIFAR-10N noisy label sets (Aggregate, Random 1, and Worst) versus synthetic class-dependent noise matched to the exact same transition matrix TT:

    • For clean/correctly annotated samples, the trajectory of sample memorization (fraction of samples with prediction confidence max⁡iP(f(x)=i)>0.95\max_i \mathbb{P}(f(x) = i) > 0.95) is nearly identical across training epochs for both real human noise and synthetic noise.
    • For wrongly annotated samples, deep networks memorize real human label noise much earlier in training and reach substantially higher proportions of memorization than on synthetic noise.

    This discrepancy occurs because human labeling errors are non-randomly concentrated on visually ambiguous, low-resolution, or misleading instances that present salient, learnable feature patterns, inducing the network to over-fit erroneous supervision more aggressively.

  5. Knowl 5 — Benchmark Accuracy Comparison of Robust Learning Methods on CIFAR-10N and CIFAR-100N

    data/table

    Test classification accuracies (mean ±\pm standard deviation over 5 runs) of representative robust learning methods on CIFAR-10N and CIFAR-100N (fine labels) using ResNet-34 (SOP evaluated on Pre-activation ResNet-18) are summarized below:

    Method CIFAR-10N CIFAR-100N
    Clean Aggregate Random 1 Random 2 Random 3 Worst Clean Noisy
    CE (Standard) 92.92 ±\pm 0.11 87.77 ±\pm 0.38 85.02 ±\pm 0.65 86.46 ±\pm 1.79 85.16 ±\pm 0.61 77.69 ±\pm 1.55 76.70 ±\pm 0.74 55.50 ±\pm 0.66
    Forward T 93.02 ±\pm 0.12 88.24 ±\pm 0.22 86.88 ±\pm 0.50 86.14 ±\pm 0.24 87.04 ±\pm 0.35 79.79 ±\pm 0.46 76.18 ±\pm 0.37 57.01 ±\pm 1.03
    Backward T 93.10 ±\pm 0.05 88.13 ±\pm 0.29 87.14 ±\pm 0.34 86.28 ±\pm 0.80 86.86 ±\pm 0.41 77.61 ±\pm 1.05 76.79 ±\pm 0.60 57.14 ±\pm 0.92
    GCE 92.83 ±\pm 0.16 87.85 ±\pm 0.70 87.61 ±\pm 0.28 87.70 ±\pm 0.56 87.58 ±\pm 0.29 80.66 ±\pm 0.35 76.35 ±\pm 0.48 56.73 ±\pm 0.30
    Co-teaching 93.35 ±\pm 0.14 91.20 ±\pm 0.13 90.33 ±\pm 0.13 90.30 ±\pm 0.17 90.15 ±\pm 0.18 83.83 ±\pm 0.13 73.46 ±\pm 0.09 60.37 ±\pm 0.27
    Co-teaching+ 92.41 ±\pm 0.20 90.61 ±\pm 0.22 89.70 ±\pm 0.27 89.47 ±\pm 0.18 89.54 ±\pm 0.22 83.26 ±\pm 0.17 70.99 ±\pm 0.22 57.88 ±\pm 0.24
    T-Revision 93.35 ±\pm 0.23 88.52 ±\pm 0.17 88.33 ±\pm 0.32 87.71 ±\pm 1.02 87.79 ±\pm 0.67 80.48 ±\pm 1.20 72.83 ±\pm 0.21 51.55 ±\pm 0.31
    Peer Loss 93.99 ±\pm 0.13 90.75 ±\pm 0.25 89.06 ±\pm 0.11 88.76 ±\pm 0.19 88.57 ±\pm 0.09 82.00 ±\pm 0.60 74.67 ±\pm 0.36 57.59 ±\pm 0.61
    ELR 93.45 ±\pm 0.65 92.38 ±\pm 0.64 91.46 ±\pm 0.38 91.61 ±\pm 0.16 91.41 ±\pm 0.44 83.58 ±\pm 1.13 72.78 ±\pm 0.80 58.94 ±\pm 0.92
    ELR+ 95.39 ±\pm 0.05 94.83 ±\pm 0.10 94.43 ±\pm 0.41 94.20 ±\pm 0.24 94.34 ±\pm 0.22 91.09 ±\pm 1.60 78.57 ±\pm 0.12 66.72 ±\pm 0.07
    Positive-LS 94.77 ±\pm 0.17 91.57 ±\pm 0.07 89.80 ±\pm 0.28 89.35 ±\pm 0.33 89.82 ±\pm 0.14 82.76 ±\pm 0.53 76.25 ±\pm 0.35 55.84 ±\pm 0.48
    F-Div 94.88 ±\pm 0.12 91.64 ±\pm 0.34 89.70 ±\pm 0.40 89.79 ±\pm 0.12 89.55 ±\pm 0.49 82.53 ±\pm 0.52 76.14 ±\pm 0.36 57.10 ±\pm 0.65
    Divide-Mix 95.37 ±\pm 0.14 95.01 ±\pm 0.71 95.16 ±\pm 0.19 95.23 ±\pm 0.07 95.21 ±\pm 0.14 92.56 ±\pm 0.42 76.94 ±\pm 0.22 71.13 ±\pm 0.48
    Negative-LS 94.92 ±\pm 0.25 91.97 ±\pm 0.46 90.29 ±\pm 0.32 90.37 ±\pm 0.12 90.13 ±\pm 0.19 82.99 ±\pm 0.36 77.06 ±\pm 0.73 58.59 ±\pm 0.98
    JoCoR 93.40 ±\pm 0.24 91.44 ±\pm 0.05 90.30 ±\pm 0.20 90.21 ±\pm 0.19 90.11 ±\pm 0.21 83.37 ±\pm 0.30 74.07 ±\pm 0.33 59.97 ±\pm 0.24
    CORES2^2 93.43 ±\pm 0.24 91.23 ±\pm 0.11 89.66 ±\pm 0.32 89.91 ±\pm 0.45 89.79 ±\pm 0.50 83.60 ±\pm 0.53 75.56 ±\pm 0.53 61.15 ±\pm 0.73
    CORES∗^* 94.16 ±\pm 0.11 95.25 ±\pm 0.09 94.45 ±\pm 0.14 94.88 ±\pm 0.31 94.74 ±\pm 0.03 91.66 ±\pm 0.09 73.87 ±\pm 0.16 55.72 ±\pm 0.42
    VolMinNet 92.14 ±\pm 0.30 89.70 ±\pm 0.21 88.30 ±\pm 0.12 88.27 ±\pm 0.09 88.19 ±\pm 0.41 80.53 ±\pm 0.20 70.61 ±\pm 0.88 57.80 ±\pm 0.31
    CAL 94.50 ±\pm 0.31 91.97 ±\pm 0.32 90.93 ±\pm 0.31 90.75 ±\pm 0.30 90.74 ±\pm 0.24 85.36 ±\pm 0.16 75.67 ±\pm 0.25 61.73 ±\pm 0.42
    PES (Semi) 94.76 ±\pm 0.20 94.66 ±\pm 0.18 95.06 ±\pm 0.15 95.19 ±\pm 0.23 95.22 ±\pm 0.13 92.68 ±\pm 0.22 77.92 ±\pm 0.04 70.36 ±\pm 0.33
    SOP N/A 95.61 ±\pm 0.13 95.28 ±\pm 0.13 95.31 ±\pm 0.10 95.39 ±\pm 0.11 93.24 ±\pm 0.21 N/A 67.81 ±\pm 0.23

    Semi-supervised and sample-selection approaches incorporating advanced data augmentations (SOP, Divide-Mix, PES Semi, ELR+, and CORES∗\text{CORES}^*) attain the top test accuracies across all noise settings.

  6. Knowl 6 — Generalization Discrepancy Between Synthetic Class-Dependent and Human Label Noise

    empirical result

    When classifiers are evaluated on synthetic class-dependent noise versus real human noise generated with the exact same expected transition matrix T=E[T(X)]T = \mathbb{E}[T(X)], models trained on synthetic noise systematically achieve higher test accuracies on CIFAR-10, demonstrating that class-dependent synthetic noise underestimates the difficulty of real-world noise.

    The performance gap, calculated as Test Accuracysynthetic−Test Accuracyhuman\text{Test Accuracy}_{\text{synthetic}} - \text{Test Accuracy}_{\text{human}} across five runs on ResNet-34, shows consistent positive discrepancies:

    Method CIFAR-10 Gap (%) CIFAR-100 Gap (%)
    Aggregate Random 1 Random 2 Random 3 Worst Noisy
    CE (Standard) +4.35 +6.01 +4.50 +5.82 +9.00 +1.20
    Forward T +4.30 +4.82 +4.86 +4.27 +7.08 -0.14
    Co-teaching+ +0.89 +0.92 +0.86 +1.05 +2.63 -0.61
    Peer Loss +1.90 +2.42 +2.74 +1.95 +4.67 -0.85
    ELR -0.78 -0.81 -0.97 -0.51 -1.64 +1.05
    F-Div +0.72 +1.62 +1.33 +1.65 +4.14 +1.31
    Divide-Mix -0.01 +0.44 +0.42 +0.28 +0.29 +0.65
    Negative-LS +0.77 +1.31 +1.08 +1.36 +4.00 +1.26
    JoCoR +0.35 +0.78 +0.68 +1.01 +2.43 -0.48
    CORES2^2 +1.49 +1.69 +1.52 +1.66 +1.67 -0.72
    CAL +0.25 +0.04 +0.04 +0.09 +0.44 +0.47

    For standard Cross-Entropy, the performance gap widens as the noise rate increases, reaching +9.00%+9.00\% on CIFAR-10 Worst (40.21%40.21\% noise). Conversely, certain regularization-based methods such as ELR perform slightly better on real human noise than on synthetic noise (indicated by negative gap values).

  7. Knowl 7 — M-NN Noise Clusterability

    definition

    Let D={(xn,y~n)}n=1N\mathcal{D} = \{(x_n, \tilde{y}_n)\}_{n=1}^N be a training dataset where xn∈Xx_n \in \mathcal{X} is the feature vector and y~n∈[K]\tilde{y}_n \in [K] is the observed noisy label. Let T(x)∈[0,1]K×KT(x) \in [0, 1]^{K \times K} denote the instance-dependent noise transition matrix for input xx, whose entries are defined by: Ti,j(x):=P(Y~=j∣Y=i,X=x)T_{i,j}(x) := \mathbb{P}(\tilde{Y} = j \mid Y = i, X = x) where YY is the clean ground-truth label.

    The dataset D\mathcal{D} is said to satisfy MM-Nearest-Neighbor (MM-NN) noise clusterability if the MM nearest neighbors of any instance xnx_n share the identical noise transition matrix as xnx_n: T(xn)=T(xni),∀i∈[M]T(x_n) = T(x_n^i), \quad \forall i \in [M] where {xni}i=1M\{x_n^i\}_{i=1}^M denotes the MM nearest neighbors of instance xnx_n in the feature representation space.

  8. Knowl 8 — Memorized Feature in Deep Neural Classifiers

    definition

    In a KK-class classification task, given a trained classifier f:X→[0,1]Kf: \mathcal{X} \to [0, 1]^K, an input feature vector x∈Xx \in \mathcal{X}, and a scalar confidence threshold η∈(0,1)\eta \in (0, 1), the feature xx is defined to be memorized by ff if the model's predicted probability for at least one class exceeds η\eta: ∃i∈[K]such thatP(f(x)=i)>η\exists i \in [K] \quad \text{such that} \quad \mathbb{P}(f(x) = i) > \eta

  9. Knowl 9 — Standard Experimental Benchmark Protocol on CIFAR-N

    experimental setup

    The benchmark experiments on CIFAR-10N and CIFAR-100N employ the following standardized training configuration:

    • Backbone Architecture: ResNet-34 (Pre-activation ResNet-18 used for SOP).
    • Optimizer: Stochastic Gradient Descent (SGD) with momentum 0.90.9, initial learning rate 0.10.1, weight decay 0.00050.0005, and mini-batch size 128128.
    • Training Schedule: 100 epochs in total, with the learning rate multiplied by 0.10.1 at epoch 50.
    • Data Augmentation: Random crop with 4-pixel padding and random horizontal flips.
    • Method Settings: Algorithms with dedicated warm-up routines or dual networks (e.g., DivideMix, ELR+, CORES∗\text{CORES}^*) use their original configurations and mixup strategies. Pre-trained weights used for warm-up stages are kept identical across runs.

Coverage note — None was omitted; all key contributions including dataset creation, hypothesis testing framework, empirical discoveries, benchmark evaluation tables, and theoretical definitions were included.

References

  1. 1.Ehsan Amid, Manfred KK Warmuth, Rohan Anil, and Tomer Koren. Robust bi-tempered logistic loss based on bregman divergences. In Advances in Neural Information Processing Systems, pp. 14987–14996, 2019.
  2. 2.Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. A closer look at memorization in deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pp. 233–242. JMLR. org, 2017.
  3. 3.Yingbin Bai, Erkun Yang, Bo Han, Yanhua Yang, Jiatong Li, Yinian Mao, Gang Niu, and Tongliang Liu. Understanding and improving early stopping for learning with noisy labels. Advances in Neural Information Processing Systems, 34, 2021.
  4. 4.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. Advances in Neural Information Processing Systems, 32:5049–5059, 2019.
  5. 5.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European conference on computer vision, pp. 446–461. Springer, 2014.
  6. 6.Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong, Xing Sun, and Yang Liu. Learning with instancedependent label noise: A sample sieve approach. In International Conference on Learning Representations, 2021.
  7. 7.Aritra Ghosh, Himanshu Kumar, and PS Sastry. Robust loss functions under label noise for deep neural networks. In Thirty-First AAAI Conference on Artificial Intelligence, 2017.
  8. 8.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In Advances in neural information processing systems, pp. 8527–8537, 2018.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  10. 10.Dan Hendrycks, Mantas Mazeika, Duncan Wilson, and Kevin Gimpel. Using trusted data to train deep networks on labels corrupted by severe noise. Advances in Neural Information Processing Systems, 31:10456–10465, 2018.
  11. 11.Gang Hua, Chengjiang Long, Ming Yang, and Yan Gao. Collaborative active learning of a kernel machine ensemble for recognition. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1209–1216, 2013.
  12. 12.Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. Mentornet: Learning datadriven curriculum for very deep neural networks on corrupted labels. In International Conference on Machine Learning, pp. 2304–2313. PMLR, 2018.
  13. 13.Lu Jiang, Di Huang, Mason Liu, and Weilong Yang. Beyond synthetic noise: Deep learning on controlled noisy labels. In International Conference on Machine Learning, pp. 4804–4815. PMLR, 2020.
  14. 14.Zhimeng Jiang, Kaixiong Zhou, Zirui Liu, Li Li, Rui Chen, Soo-Hyun Choi, and Xia Hu. An information fusion approach to learning with instance-dependent label noise. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=ecH2FKaARUp.
  15. 15.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pp. 554–561, 2013.
  16. 16.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  17. 17.Kuang-Huei Lee, Xiaodong He, Lei Zhang, and Linjun Yang. Cleannet: Transfer learning for scalable image classifier training with label noise. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5447–5456, 2018.
  18. 18.Junnan Li, Richard Socher, and Steven C.H. Hoi. Dividemix: Learning with noisy labels as semisupervised learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HJgExaVtwr.
  19. 19.Wen Li, Limin Wang, Wei Li, Eirikur Agustsson, and Luc Van Gool. Webvision database: Visual learning and understanding from web data. arXiv preprint arXiv:1708.02862, 2017.
  20. 20.Xuefeng Li, Tongliang Liu, Bo Han, Gang Niu, and Masashi Sugiyama. Provably end-to-end label-noise learning without anchor points. arXiv preprint arXiv:2102.02400, 2021.
  21. 21.Yuan-Hong Liao, Amlan Kar, and Sanja Fidler. Towards good practices for efficiently annotating large-scale image classification datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4350–4359, 2021.
  22. 22.Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda. Early-learning regularization prevents memorization of noisy labels. Advances in Neural Information Processing Systems, 33, 2020.
  23. 23.Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You. Robust training under label noise by overparameterization. arXiv preprint arXiv:2202.14026, 2022.
  24. 24.Tongliang Liu and Dacheng Tao. Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence, 38(3):447–461, 2015.
  25. 25.Yang Liu and Hongyi Guo. Peer loss functions: Learning from noisy labels without knowing noise rates. In International Conference on Machine Learning, pp. 6226–6236. PMLR, 2020.
  26. 26.Chengjiang Long and Gang Hua. Multi-class multi-annotator active learning with robust gaussian process for visual recognition. In Proceedings of the IEEE international conference on computer vision, pp. 2839–2847, 2015.
  27. 27.Michal Lukasik, Srinadh Bhojanapalli, Aditya Menon, and Sanjiv Kumar. Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp. 6448–6458. PMLR, 2020.
  28. 28.Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari. Learning with noisy labels. In Advances in neural information processing systems, pp. 1196–1204, 2013.
  29. 29.Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu. Making deep neural networks robust to label noise: A loss correction approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1944–1952, 2017.
  30. 30.Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Russakovsky. Human uncertainty makes classification more robust. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9617–9626, 2019.
  31. 31.Hwanjun Song, Minseok Kim, and Jae-Gil Lee. Selfie: Refurbishing unclean samples for robust deep learning. In International Conference on Machine Learning, pp. 5907–5915, 2019.
  32. 32.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. Advances in neural information processing systems, 29:3630–3638, 2016.
  33. 33.Jialu Wang, Yang Liu, and Caleb Levy. Fair classification with group-dependent label noise. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 526–536, 2021a.
  34. 34.Jingkang Wang, Hongyi Guo, Zhaowei Zhu, and Yang Liu. Policy learning using weak supervision. Advances in Neural Information Processing Systems, 34, 2021b.
  35. 35.Yisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo, Jinfeng Yi, and James Bailey. Symmetric cross entropy for robust learning with noisy labels. In Proceedings of the IEEE International Conference on Computer Vision, pp. 322–330, 2019.
  36. 36.Hongxin Wei, Lei Feng, Xiangyu Chen, and Bo An. Combating noisy labels by agreement: A joint training method with co-regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13726–13735, 2020.
  37. 37.Jiaheng Wei and Yang Liu. When optimizing f-divergence is robust with label noise. In International Conference on Learning Representations, 2020.
  38. 38.Jiaheng Wei, Hangyu Liu, Tongliang Liu, Gang Niu, and Yang Liu. Understanding (generalized) label smoothing whenlearning with noisy labels. arXiv preprint arXiv:2106.04149, 2021.
  39. 39.Xiaobo Xia, Tongliang Liu, Nannan Wang, Bo Han, Chen Gong, Gang Niu, and Masashi Sugiyama. Are anchor points really indispensable in label-noise learning? In Advances in Neural Information Processing Systems, pp. 6838–6849, 2019.
  40. 40.Xiaobo Xia, Tongliang Liu, Bo Han, Chen Gong, Nannan Wang, Zongyuan Ge, and Yi Chang. Robust early-learning: Hindering the memorization of noisy labels. In International Conference on Learning Representations, 2020a.
  41. 41.Xiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang, Mingming Gong, Haifeng Liu, Gang Niu, Dacheng Tao, and Masashi Sugiyama. Part-dependent label noise: Towards instance-dependent label noise. In Advances in Neural Information Processing Systems, volume 33, pp. 7597–7610, 2020b.
  42. 42.Tong Xiao, Tian Xia, Yi Yang, Chang Huang, and Xiaogang Wang. Learning from massive noisy labeled data for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2691–2699, 2015.
  43. 43.Zeke Xie, Fengxiang He, Shaopeng Fu, Issei Sato, Dacheng Tao, and Masashi Sugiyama. Artificial neural variability for deep learning: on overfitting, noise memorization, and catastrophic forgetting. Neural computation, 33(8):2163–2192, 2021.
  44. 44.Yu Yao, Tongliang Liu, Bo Han, Mingming Gong, Jiankang Deng, Gang Niu, and Masashi Sugiyama. Dual T: Reducing estimation error for transition matrix in label-noise learning. In Advances in Neural Information Processing Systems, volume 33, pp. 7260–7271, 2020.
  45. 45.Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor W Tsang, and Masashi Sugiyama. How does disagreement help generalization against label corruption? arXiv preprint arXiv:1901.04215, 2019.
  46. 46.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 64(3):107–115, 2021a.
  47. 47.Yikai Zhang, Songzhu Zheng, Pengxiang Wu, Mayank Goswami, and Chao Chen. Learning with feature-dependent label noise: A progressive approach. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=ZPa2SyGcbwh.
  48. 48.Zhilu Zhang and Mert Sabuncu. Generalized cross entropy loss for training deep neural networks with noisy labels. In Advances in neural information processing systems, pp. 8778–8788, 2018.
  49. 49.Zhaowei Zhu, Tongliang Liu, and Yang Liu. A second-order approach to learning with instancedependent label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10113–10123, 2021a.
  50. 50.Zhaowei Zhu, Yiwen Song, and Yang Liu. Clusterability as an alternative to anchor points when learning with noisy labels. In International Conference on Machine Learning, pp. 12912–12923. PMLR, 2021b.
  51. 51.Zhaowei Zhu, Tianyi Luo, and Yang Liu. The rich get richer: Disparate impact of semi-supervised learning. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=DXPftn5kjQK.

Citation

MLA
Wei, J., et al. “Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations”. arXiv, 2021, http://arxiv.org/abs/2110.12088v2.
APA
Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., & Liu, Y. (2021). Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations. arXiv. http://arxiv.org/abs/2110.12088v2
Chicago
Wei, J., Z. Zhu, H. Cheng, T. Liu, G. Niu, and Y. Liu. 2021. “Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations”. arXiv. http://arxiv.org/abs/2110.12088v2.
Harvard
Wei, J. et al. (2021) “Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.12088v2.
Vancouver
1. Wei J, Zhu Z, Cheng H, Liu T, Niu G, Liu Y (2021) Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations. arXiv

BibTeX

@article{wei2021learning,
  title = {Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations},
  author = {Wei, Jiaheng and Zhu, Zhaowei and Cheng, Hao and Liu, Tongliang and Niu, Gang and Liu, Yang},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.12088v2},
  eprint = {2110.12088}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors