Debiased Learning from Naturally Imbalanced Pseudo-Labels

Xudong WangZhirong WuLong LianStella X. Yu

article2022CVPR166 citationsOral Presentation

Reveals that machine-generated pseudo-labels inherently suffer from severe class imbalance even on balanced datasets, and introduces a counterfactual debiasing framework with adaptive margins that substantially boosts accuracy in semi-supervised and zero-shot learning.

Listen

Modern computer vision models increasingly rely on pseudo-labels—model-generated predictions used to supervise training on massive unlabeled datasets—to cut annotation costs in semi-supervised and zero-shot learning. However, the article demonstrates that pseudo-labels naturally become highly imbalanced and biased due to inter-class similarities and model confusion, even when the underlying training and evaluation data are perfectly balanced. When models self-train on these biased predictions, they reinforce their own errors by creating false majority classes, trapping the systems in compounding performance deficits.

The article aims to introduce and validate a debiased learning framework, termed Debiased Pseudo-Labeling, that dynamically detects and removes these pseudo-label biases during training without requiring prior knowledge of the true class distributions.

To address this issue, the authors developed a two-part adaptive approach combining counterfactual reasoning to strip out response biases and an adaptive marginal loss to separate easily confused classes. They evaluated this framework across standard visual benchmarks—including ImageNet, CIFAR-10, and satellite imagery—testing its effectiveness across varying levels of supervision, severe class imbalances, and cross-domain zero-shot transfer.

The experimental findings show substantial performance gains. On ImageNet with only 0.2% labeled data, the proposed framework achieved an absolute top-1 accuracy gain of 26% over prior baselines, and delivered an 8.7% to 9% gain in zero-shot transfer learning. In zero-shot settings, a standard ResNet-50 model trained with this debiasing approach outperformed foundation models with up to fifteen times more parameters. Furthermore, the framework demonstrated universal utility across multiple learning architectures (such as FixMatch, MixMatch, and UDA) and improved accuracy by over 20% on difficult domain shifts, including an improvement of 25.7% on satellite land-cover classification.

These results demonstrate that machine learning teams can significantly lower human data annotation costs and compute requirements without sacrificing accuracy. Instead of relying on brute-force scale or manual data rebalancing, dynamically correcting algorithmic bias during training provides a robust, low-overhead safeguard against error propagation. Crucially, the method operates effectively even when the true distribution of the target data is unknown.

Organizations deploying semi-supervised and zero-shot pipelines should integrate adaptive debiasing modules into their existing training workflows as a plug-and-play enhancement. Before moving to full-scale deployment, teams should conduct internal pilot evaluations to calibrate debiasing hyperparameters, as excessively strong correction factors can impede model fitting while overly weak factors will fail to eliminate bias.

  • Paper: Contrastive Test-Time Adaptation, Dian Chen et al. (2022). Extends debiased online pseudo-labeling principles to test-time adaptation via contrastive learning and class diversification regularization.
Cover for Debiased Learning from Naturally Imbalanced Pseudo-Labels

Abstract

Pseudo-labels are confident predictions made on unlabeled target data by a classifier trained on labeled source data. They are widely used for adapting a model to unlabeled data, e.g., in a semi-supervised learning setting.

Our key insight is that pseudo-labels are naturally imbalanced due to intrinsic data similarity, even when a model is trained on balanced source data and evaluated on balanced target data. If we address this previously unknown imbalanced classification problem arising from pseudo-labels instead of ground-truth training labels, we could remove model biases towards false majorities created by pseudo-labels.

We propose a novel and effective debiased learning method with pseudo-labels, based on counterfactual reasoning and adaptive margins: the former removes the classifier response bias, whereas the latter adjusts the margin of each class according to the imbalance of pseudo-labels. Validated by extensive experimentation, our simple debiased learning delivers significant accuracy gains over the state-of-the-art on ImageNet-1K: 26% for semi-supervised learning with 0.2% annotations and 9% for zero-shot learning. Our code is available at: https://github.com/frank-xwang/debiased-pseudo-labeling.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Pseudo-Labels are Naturally Imbalanced
  • 3.1. Background
  • 3.2. Biases in Semi-supervised Learning
  • 3.3. Biases in Zero-Shot Learning
  • 3.4. Inter-Class Correlations
  • 4. Debiased Pseudo-Labeling
  • 4.1. Adaptive Debiasing
  • 4.2. Distinctions and Connections with Alternatives
  • 4.3. DebiasPL for T-ZSL and SSL
  • 5. Experiment
  • 5.1. Semi-supervised Learning
  • 5.2. Transductive Zero-Shot Learning
  • 6. Summary
  • References

Knowls

  1. Knowl 1 — Adaptive Debiasing via Counterfactual Reasoning for Pseudo-Labels

    model/method

    In pseudo-labeling, models generate biased predictions towards certain majority classes. To dynamically eliminate this model response bias without requiring prior knowledge of the true class distribution, DebiasPL formulates pseudo-labeling through causal inference using counterfactual reasoning.

    Let AiA_i denote the input instance, YY the prediction, DD the mediator, and MM the model bias. The pursuit of the direct causal effect along Ai→YA_i \to Y is formulated as the Controlled Direct Effect (CDE):

    CDE(Yi)=[Yi∣do(Ai),do(D)]−[Yi∣do(A^),do(D)]\text{CDE}(Y_i) = [Y_i \mid \text{do}(A_i), \text{do}(D)] - [Y_i \mid \text{do}(\hat{A}), \text{do}(D)]

    where A^={A1,…,An}\hat{A} = \{A_1, \dots, A_n\}. Because evaluating counterfactual outcomes over all training samples is computationally prohibitive, the Approximated Controlled Direct Effect (ACDE) approximates the counterfactual outcome using a momentum-updated prediction distribution p^\hat{p}. The debiased logit vector f~i∈RC\tilde{f}_i \in \mathbb{R}^C for a weakly augmented unlabeled sample α(xi)\alpha(x_i) across CC classes is computed as:

    f~i=f(α(xi))−λlog⁡p^\tilde{f}_i = f(\alpha(x_i)) - \lambda \log \hat{p}

    where f(α(xi))∈RCf(\alpha(x_i)) \in \mathbb{R}^C is the raw predicted logit vector, λ>0\lambda > 0 is the debias factor controlling the strength of the indirect effect removal, and p^∈[0,1]C\hat{p} \in [0, 1]^C is the running average class probability vector updated at each iteration as:

    p^←mp^+(1−m)1μB∑k=1μBpk\hat{p} \leftarrow m \hat{p} + (1 - m) \frac{1}{\mu B} \sum_{k=1}^{\mu B} p_k

    Here, m∈[0,1)m \in [0, 1) is the momentum coefficient, BB is the labeled batch size, μ\mu is the unlabeled-to-labeled sample ratio in a minibatch, and pk=softmax(f(α(xk)))∈[0,1]Cp_k = \text{softmax}(f(\alpha(x_k))) \in [0, 1]^C is the predicted probability distribution of unlabeled instance xkx_k. The logarithm log⁡p^\log \hat{p} rescales the probability vector to match the scale of the logits. The debiased logit f~i\tilde{f}_i replaces the raw model logit when assigning pseudo-labels.

  2. Knowl 2 — Adaptive Marginal Loss for Countering Inter-Class Confusion

    model/method

    To counteract inter-class confusion where minority or hard classes are misclassified into specific dominant classes, DebiasPL introduces an Adaptive Marginal Loss (AML). AML enforces a larger margin for classes that are frequently predicted (highly biased) relative to rarely predicted classes, preventing dominant categories from overwhelming model predictions during student update.

    Given the momentum-updated running average class probability distribution p^∈[0,1]C\hat{p} \in [0, 1]^C across CC classes, class-specific dynamic margins Δj\Delta_j are defined as:

    Δj=λlog⁡(1p^j)for j∈{1,…,C}\Delta_j = \lambda \log\left(\frac{1}{\hat{p}_j}\right) \quad \text{for } j \in \{1, \dots, C\}

    where λ>0\lambda > 0 is the debias factor. For a strongly augmented unlabeled instance β(xi)\beta(x_i) with model logits z=f(β(xi))∈RCz = f(\beta(x_i)) \in \mathbb{R}^C and an assigned pseudo-label y^i∈{1,…,C}\hat{y}_i \in \{1, \dots, C\}, the adaptive marginal loss LAML\mathcal{L}_{AML} is formulated as:

    LAML=−log⁡exp⁡(zy^i−Δy^i)exp⁡(zy^i−Δy^i)+∑k≠y^iCexp⁡(zk−Δk)\mathcal{L}_{AML} = -\log \frac{\exp(z_{\hat{y}_i} - \Delta_{\hat{y}_i})}{\exp(z_{\hat{y}_i} - \Delta_{\hat{y}_i}) + \sum_{k \neq \hat{y}_i}^C \exp(z_k - \Delta_k)}

    By dynamically enforcing larger margins against dominant classes based on online pseudo-label statistics p^\hat{p}, LAML\mathcal{L}_{AML} mitigates inter-class confounding without requiring fixed class count statistics or ground-truth class marginals.

  3. Knowl 3 — Debiased Pseudo-Labeling Framework for Semi-Supervised Learning

    model/method

    Debiased Pseudo-Labeling (DebiasPL) integrates adaptive debiasing and adaptive marginal loss into pseudo-labeling-based semi-supervised learning (SSL). Given a labeled dataset XL={(xi,yi)}i=1L\mathcal{X}_L = \{(x_i, y_i)\}_{i=1}^L with one-hot targets yi∈{0,1}Cy_i \in \{0, 1\}^C and an unlabeled dataset XU={xi}i=L+1L+U\mathcal{X}_U = \{x_i\}_{i=L+1}^{L+U}, the total objective is:

    L=Ls+λuLu\mathcal{L} = \mathcal{L}_s + \lambda_u \mathcal{L}_u

    where Ls\mathcal{L}_s is the standard supervised cross-entropy loss on weakly-augmented labeled instances α(xi)\alpha(x_i) with batch size BB:

    Ls=1B∑i=1BH(yi,softmax(f(α(xi))))\mathcal{L}_s = \frac{1}{B} \sum_{i=1}^B \mathcal{H}(y_i, \text{softmax}(f(\alpha(x_i))))

    and λu\lambda_u is a scalar balancing hyperparameter. The unsupervised loss Lu\mathcal{L}_u over μB\mu B unlabeled instances is formulated as:

    Lu=1μB∑i=1μB1[max⁡(softmax(f~i))≥τ]⋅LAML(f(β(xi)),y^i)\mathcal{L}_u = \frac{1}{\mu B} \sum_{i=1}^{\mu B} \mathbf{1}\left[\max(\text{softmax}(\tilde{f}_i)) \ge \tau\right] \cdot \mathcal{L}_{AML}(f(\beta(x_i)), \hat{y}_i)

    Here:

    • α(⋅)\alpha(\cdot) and β(⋅)\beta(\cdot) denote weak and strong data augmentations, respectively.
    • f~i=f(α(xi))−λlog⁡p^\tilde{f}_i = f(\alpha(x_i)) - \lambda \log \hat{p} is the counterfactually debiased logit vector, where p^\hat{p} is the momentum-updated running average probability across unlabeled batches.
    • y^i=arg⁡max⁡cf~ic\hat{y}_i = \arg\max_c \tilde{f}_i^c is the debiased pseudo-label.
    • τ∈(0,1)\tau \in (0, 1) is the confidence threshold.
    • LAML\mathcal{L}_{AML} is the adaptive marginal loss evaluated on strongly-augmented logits f(β(xi))f(\beta(x_i)) with margin offsets Δj=λlog⁡(1/p^j)\Delta_j = \lambda \log(1 / \hat{p}_j).

    Optionally, unlabeled instances with confidence below τ\tau can be regularized via cross-level instance-group discrimination loss (CLD) or pseudo-labeled using a vision-language model (CLIP) if CLIP confidence exceeds a threshold τclip\tau_{\text{clip}}.

  4. Knowl 4 — Transductive Zero-Shot Learning via CLIP and Debiased Pseudo-Labeling

    model/method

    Transductive Zero-Shot Learning (T-ZSL) adapts a visual recognition model to an unlabeled target dataset when candidate class names are known but no human-annotated labels are provided. While zero-shot CLIP predictions suffer from severe class imbalance that is amplified by standard self-training, DebiasPL resolves this by combining confidence-based pseudo-labeling with adaptive debiasing.

    1. Pseudo-Label Initialization from CLIP: Text prompts derived from target class names are embedded alongside target images using pre-trained CLIP encoders. The cosine similarities are normalized via softmax to produce class probability distributions. For each target sample whose maximum predicted class probability exceeds a high confidence threshold τclip\tau_{\text{clip}} (set to 0.950.95 by default), the argmax prediction is assigned as a one-hot pseudo-label, creating an initial pseudo-labeled set.
    2. Joint Optimization: The pseudo-labeled set is treated as the "labeled" data pool XL\mathcal{X}_L, and the full target dataset is treated as the unlabeled pool XU\mathcal{X}_U. The model (e.g., ResNet-50 initialized via MoCo v2 + EMAN) is trained using FixMatch augmented with Adaptive Debiasing (ACDE logit adjustment) to generate debiased pseudo-labels y^i=arg⁡max⁡(f(α(xi))−λlog⁡p^)\hat{y}_i = \arg\max (f(\alpha(x_i)) - \lambda \log \hat{p}) and Adaptive Marginal Loss (AML) on strongly-augmented samples f(β(xi))f(\beta(x_i)).

    Because CLIP predictions for the entire dataset can be computed once and cached, the computational overhead is negligible, while avoiding bias amplification during adaptation.

  5. Knowl 5 — Natural Imbalance and Inter-Class Confounding in Pseudo-Labels

    empirical result

    Even when deep neural networks are trained on curated, class-balanced source datasets and evaluated on class-balanced target datasets, the generated pseudo-labels are naturally class-imbalanced across both semi-supervised learning (SSL) and zero-shot learning (ZSL).

    Key empirical findings include:

    • SSL (FixMatch on CIFAR-10): When trained on 4 balanced labels per class, FixMatch generates strongly skewed pseudo-label distributions over balanced unlabeled CIFAR-10 data. This skew persists from early epochs (e.g., epoch 20) through late stages (epoch 100).
    • ZSL (CLIP on ImageNet-1K): Pre-trained CLIP produces heavily biased zero-shot predictions on balanced ImageNet-1K (e.g., class 0 is assigned over 3,500 predictions, >3 times its true count of ~1,300). High-frequency majority classes show high recall but substantially lower precision than medium/few-shot classes. Similar prediction skews occur on EuroSAT, MNIST, CIFAR-10, CIFAR-100, and Food101.
    • Root Cause (Inter-Class Correlation): Measuring cosine similarities between class image feature centroids reveals that low-frequency classes (the 10 least predicted classes) exhibit strong inter-class confusion with neighboring classes, whereas high-frequency classes have lower confounding. For instance, FixMatch systematically misclassifies instances of "ship" into "plane". Self-training without intervention reinforces these errors.
  6. Knowl 6 — Semi-Supervised Learning Performance on CIFAR-10 and CIFAR10-LT

    data/table

    DebiasPL evaluated on balanced CIFAR-10 and long-tailed CIFAR10-LT (with imbalance ratios γ∈{100,200}\gamma \in \{100, 200\}) using a Wide-ResNet-28-2 (WRN-28-2) backbone. Results represent top-1 classification accuracy (%) averaged over 5 folds.

    Method CIFAR10-LT: # of labels (percentage) CIFAR10: # of labels (percentage)
    γ=100\gamma=100 γ=200\gamma=200 40 80 250
    1244 (10%) 3726 (30%) 1125 (10%) 3365 (30%) (0.08%) (0.16%) (2%)
    UDA - - - - 71.0±6.071.0 \pm 6.0 - 91.2±1.191.2 \pm 1.1
    MixMatch 60.4±2.260.4 \pm 2.2 - 54.5±1.954.5 \pm 1.9 - 51.9±11.851.9 \pm 11.8 80.8±1.380.8 \pm 1.3 89.0±0.989.0 \pm 0.9
    CReST w/ DA 75.9±0.675.9 \pm 0.6 77.6±0.977.6 \pm 0.9 64.1±0.2264.1 \pm 0.22 67.7±0.867.7 \pm 0.8 - - -
    CReST+ w/ DA 78.1±0.878.1 \pm 0.8 79.2±0.279.2 \pm 0.2 67.7±1.467.7 \pm 1.4 70.5±0.670.5 \pm 0.6 - - -
    CoMatch w/ SimCLR - - - - 92.6±1.092.6 \pm 1.0 94.0±0.394.0 \pm 0.3 95.1±0.395.1 \pm 0.3
    FixMatch 67.3±1.267.3 \pm 1.2 73.1±0.673.1 \pm 0.6 59.7±0.659.7 \pm 0.6 67.7±0.867.7 \pm 0.8 86.1±3.586.1 \pm 3.5 92.1±0.992.1 \pm 0.9 94.9±0.794.9 \pm 0.7
    FixMatch w/ DA w/ LA 70.4±2.970.4 \pm 2.9 - 62.4±1.262.4 \pm 1.2 - - - -
    FixMatch w/ DA w/ SimCLR - - - - 89.7±4.689.7 \pm 4.6 93.3±0.593.3 \pm 0.5 94.9±0.794.9 \pm 0.7
    DebiasPL (w/ FixMatch) 79.2±1.0\mathbf{79.2 \pm 1.0} 80.6±0.5\mathbf{80.6 \pm 0.5} 71.4±2.0\mathbf{71.4 \pm 2.0} 74.1±0.6\mathbf{74.1 \pm 0.6} 94.6±1.3\mathbf{94.6 \pm 1.3} 95.2±0.1\mathbf{95.2 \pm 0.1} 95.4±0.1\mathbf{95.4 \pm 0.1}
    gains over best FixMatch +8.8 +7.5 +9.0 +6.4 +4.9 +1.9 +0.5

    DebiasPL uses a single unified set of hyperparameters across all benchmarks without requiring prior knowledge of the marginal class distribution, outperforming both methods designed for balanced data and methods specifically tailored to long-tailed distributions.

  7. Knowl 7 — ImageNet-1K Semi-Supervised Learning in Low-Shot Settings

    data/table

    Top-1 and top-5 classification accuracy (%) on ImageNet-1K semi-supervised learning under 1% and 0.2% labeled data fractions using a ResNet-50 backbone initialized with MoCo v2 + EMAN.

    Method B.S. #epochs Pre-train 1% labeled 0.2% labeled
    top-1 top-5 top-1 top-5
    FixMatch w/ DA 4096 400 No 53.4 74.4 - -
    FixMatch w/ DA 4096 400 Yes 59.9 79.8 - -
    FixMatch w/ EMAN 384 50 Yes 60.9 82.5 43.6 64.6
    SwAV 4096 50 Yes 53.9 78.5 - -
    SimCLRv2 (+ distillation) 4096 400 Yes 60.0 79.8 - -
    PAWS (multi-crops) 4096 50 Yes 66.5 - - -
    CoMatch (multi-views) 1440 400 Yes 67.1 87.1 - -
    CLIP (few-shot) 256 50 Yes 53.4 - 40.0 -
    DebiasPL w/ FixMatch 384 50 Yes 63.1 (+2.2) 83.6 (+1.1) 47.9 (+3.7) 69.6 (+5.0)
    DebiasPL (multi-views) 768 50 Yes 65.3 (+4.4) 85.2 (+2.7) 51.6 (+8.0) 73.3 (+8.7)
    DebiasPL (multi-views) 768 200 Yes 66.5 (+5.6) 85.6 (+3.1) 52.3 (+8.7) 73.5 (+8.9)
    DebiasPL (multi-views) 1536 300 Yes 67.1 (+6.2) 85.8 (+3.3) - -
    DebiasPL w/ CLIP 384 50 Yes 69.1 (+8.2) 89.1 (+6.6) 68.2 (+24.6) 88.2 (+23.6)
    DebiasPL w/ CLIP (multi-views) 768 50 Yes 70.9 (+10.0) 89.3 (+6.8) 69.6 (+26.0) 88.4 (+23.8)

    DebiasPL brings significant gains over the FixMatch w/ EMAN baseline, especially in the extreme 0.2% label regime (+26.0% top-1 with multi-views and CLIP integration).

  8. Knowl 8 — Transductive Zero-Shot Learning Results and Robustness to Domain Shift

    data/table

    Zero-shot classification performance on ImageNet-1K (Table 6) and robustness under domain shift across various benchmark datasets (Figure 9) using a ResNet-50 backbone.

    ImageNet-1K Zero-Shot Classification

    Method #param top-1 (%) top-5 (%)
    ConSE - 1.3 3.8
    DGP - 3.0 9.3
    ZSL-KG - 3.0 9.9
    Visual N-Grams - 11.5 -
    CLIP (prompt ensemble) 26M 59.6 -
    CLIP + FixMatch (vanilla) 26M 55.7 80.6
    CLIP + DebiasPL (ours) 26M 68.3 (+8.7) 88.9 (+8.3)
    CLIP (few-shot, ∼\sim1.5% labels) 26M 53.4 -
    CLIP + CoOp (few-shot, ∼\sim1.5% labels) 26M 60.9 -
    CLIP (ViT-B/32) 398M 63.2 -
    CLIP (ResNet50x4) 375M 65.8 -

    Applying vanilla FixMatch to CLIP pseudo-labels degrades performance from 59.6% to 55.7% due to bias amplification, whereas DebiasPL achieves 68.3% top-1, outperforming CLIP models with 15×15\times more parameters and models fine-tuned with human annotations.

    Robustness to Domain Shift Across Downstream Datasets

    • EuroSAT: CLIP zero-shot = 43.5%, DebiasPL = 69.8% (+25.7%)
    • CIFAR-100: CLIP zero-shot = 40.9%, DebiasPL = 63.4% (+22.5%)
    • MNIST: CLIP zero-shot = 61.7%, DebiasPL = 83.7% (+22.0%)
    • CIFAR-10: CLIP zero-shot = 72.3%, DebiasPL = 91.5% (+19.2%)
    • Food101: CLIP zero-shot = 80.5%, DebiasPL = 85.1% (+4.6%)

    DebiasPL achieves greater performance improvements on datasets exhibiting larger domain shifts relative to the pre-training distribution.

  9. Knowl 9 — DebiasPL as a Universal Add-On and Mismatched Distribution Robustness

    data/table

    DebiasPL acts as a modular add-on across different semi-supervised learning algorithms and remains robust when labeled and unlabeled data follow mismatched distributions.

    Universality Across SSL Methods (CIFAR-10, 4 labels/class, top-1 accuracy %)

    Algorithm FixMatch MixMatch UDA
    Baseline 89.7±4.689.7 \pm 4.6 47.5±11.547.5 \pm 11.5 29.1±5.929.1 \pm 5.9
    + DebiasPL 94.6±1.3\mathbf{94.6 \pm 1.3} 61.7±6.1\mathbf{61.7 \pm 6.1} 43.2±5.2\mathbf{43.2 \pm 5.2}
    Gain +4.9 +14.2 +14.1

    Robustness Under Mismatched Label/Unlabeled Distributions (CIFAR-10, 10% labeled data, γ=200\gamma=200)

    Method Labeled: Long-Tailed Labeled: Long-Tailed
    Unlabeled: Long-Tailed Unlabeled: Balanced
    FixMatch 62.3±1.662.3 \pm 1.6 72.1±2.372.1 \pm 2.3
    DebiasPL 71.4±2.0\mathbf{71.4 \pm 2.0} (+9.1) 83.5±2.4\mathbf{83.5 \pm 2.4} (+11.4)

    When unlabeled target data is class-balanced while the small labeled set is long-tailed, DebiasPL achieves an 11.4% gain over FixMatch, demonstrating adaptability to distribution mismatch without requiring distribution priors.

Coverage note — None was omitted; all key contributions including the causal formulation (ACDE), adaptive marginal loss, SSL/T-ZSL integration, empirical bias discovery, and complete experimental tables across CIFAR, ImageNet, and domain shift benchmarks are covered.

References

  1. 1.Andrew Arnold, Ramesh Nallapati, and William W Cohen. A comparative study of methods for transductive transfer learning. In Seventh IEEE international conference on data mining workshops (ICDMW 2007), pages 77–82. IEEE, 2007. 1
  2. 2.Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Armand Joulin, Nicolas Ballas, and Michael Rabbat. Semi-supervised learning of visual features by nonparametrically predicting view assignments with support samples. arXiv preprint arXiv:2104.13963, 2021. 2, 7
  3. 3.Elias Bareinboim and Judea Pearl. Controlling selection bias in causal inference. In Artificial Intelligence and Statistics, pages 100–108. PMLR, 2012. 4
  4. 4.David Berthelot, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring. In International Conference on Learning Representations, 2019. 1, 2, 5, 6, 7
  5. 5.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. Advances in Neural Information Processing Systems, 32:5049–5059, 2019. 1, 2, 7
  6. 6.Michel Besserve, Arash Mehrjou, R´emy Sun, and Bernhard Sch¨olkopf. Counterfactuals uncover the modular structure of deep generative models. In International Conference on Learning Representations, 2019. 4
  7. 7.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In European Conference on Computer Vision, 2014. 4, 8
  8. 8.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 2
  9. 9.Zhaowei Cai, Avinash Ravichandran, Subhransu Maji, Charless Fowlkes, Zhuowen Tu, and Stefano Soatto. Exponential moving average normalization for self-supervised and semisupervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 194–203, 2021. 7
  10. 10.Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. Learning imbalanced datasets with labeldistribution-aware margin loss. In Advances in Neural Information Processing Systems, pages 1567–1578, 2019. 1, 2, 6
  11. 11.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In Thirty-fourth Conference on Neural Information Processing Systems (NeurIPS), 2020. 7
  12. 12.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 7
  13. 13.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020. 2, 7
  14. 14.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020. 3
  15. 15.Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth. Describing objects by their attributes. In 2009 IEEE conference on computer vision and pattern recognition, pages 1778–1785. IEEE, 2009. 2
  16. 16.Andrea Frome, Greg Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. Devise: A deep visual-semantic embedding model. 2013. 2
  17. 17.Sander Greenland, James M Robins, and Judea Pearl. Confounding and collapsibility in causal inference. Statistical science, pages 29–46, 1999. 4, 5
  18. 18.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5356–5364, 2019. 1
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 7
  20. 20.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019. 4, 8
  21. 21.Paul W Holland. Statistics and causal inference. Journal of the American statistical Association, 81(396):945–960, 1986. 5
  22. 22.Sham M Kakade, Karthik Sridharan, and Ambuj Tewari. On the complexity of linear prediction: Risk bounds, margin bounds, and regularization. Advances in neural information processing systems, 21, 2008. 2
  23. 23.Michael Kampffmeyer, Yinbo Chen, Xiaodan Liang, Hao Wang, Yujia Zhang, and Eric P Xing. Rethinking knowledge graph propagation for zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11487–11496, 2019. 2, 8
  24. 24.Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis. Learning imbalanced datasets with label-distribution-aware margin loss. In International Conference on Learning Representations, pages 1567–1578, 2020. 1, 2
  25. 25.Guoliang Kang, Lu Jiang, Yi Yang, and Alexander G Hauptmann. Contrastive adaptation network for unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4893–4902, 2019. 1
  26. 26.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 4, 6, 8
  27. 27.Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. Attribute-based classification for zero-shot visual object categorization. IEEE transactions on pattern analysis and machine intelligence, 36(3):453–465, 2013. 2
  28. 28.Yann LeCun. The mnist database of handwritten digits. http://yann.lecun.com/exdb/mnist/, 1998. 4, 8
  29. 29.Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, 2013. 1, 2, 3
  30. 30.Ang Li, Allan Jabri, Armand Joulin, and Laurens van der Maaten. Learning visual n-grams from web data. In Proceedings of the IEEE International Conference on Computer Vision, pages 4183–4192, 2017. 8
  31. 31.Junnan Li, Caiming Xiong, and Steven Hoi. Comatch: Semisupervised learning with contrastive graph regularization. arXiv preprint arXiv:2011.11183, 2021. 2, 7
  32. 32.Bin Liu, Zhirong Wu, Han Hu, and Stephen Lin. Deep metric transfer for label propagation with limited annotated data. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 1
  33. 33.Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang. Large-margin softmax loss for convolutional neural networks. In ICML, volume 2, page 7, 2016. 2
  34. 34.Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54(6):1–35, 2021. 1
  35. 35.Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar. Long-tail learning via logit adjustment. In International Conference on Learning Representations, 2021. 2, 6, 7
  36. 36.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 41(8):1979–1993, 2018. 2
  37. 37.Jaemin Na, Heechul Jung, Hyung Jin Chang, and Wonjun Hwang. Fixbi: Bridging domain spaces for unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1094–1103, 2021. 1
  38. 38.Nihal V Nayak and Stephen H Bach. Zero-shot learning with common sense knowledge graphs. arXiv preprint arXiv:2006.10713, 2020. 2, 8
  39. 39.Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg Corrado, and Jeffrey Dean. Zero-shot learning by convex combination of semantic embeddings. In ICLR, 2014. 8
  40. 40.Judea Pearl. Causal inference in statistics: An overview. Statistics surveys, 3:96–146, 2009. 4, 5
  41. 41.Judea Pearl. Causality. Cambridge university press, 2009. 5
  42. 42.Judea Pearl. Direct and indirect effects. arXiv preprint arXiv:1301.2300, 2013. 4, 5
  43. 43.Judea Pearl and Dana Mackenzie. The book of why: the new science of cause and effect. Basic Books, 2018. 5
  44. 44.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543, 2014. 2
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021. 1, 2, 3, 6, 7, 8
  46. 46.Lorenzo Richiardi, Rino Bellocco, and Daniela Zugna. Mediation analysis in epidemiology: methods, interpretation and bias. International journal of epidemiology, 42(5):1511–1519, 2013. 5
  47. 47.Bernardino Romera-Paredes and Philip Torr. An embarrassingly simple approach to zero-shot learning. In International conference on machine learning, pages 2152–2161. PMLR, 2015. 2
  48. 48.Donald B Rubin. Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association, 100(469):322–331, 2005. 4
  49. 49.Donald B Rubin. Essential concepts of causal inference: a remarkable history and an intriguing future. Biostatistics & Epidemiology, 3(1):140–155, 2019. 4
  50. 50.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. 6, 8
  51. 51.Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. Regularization with stochastic transformations and perturbations for deep semi-supervised learning. arXiv preprint arXiv:1606.04586, 2016. 2
  52. 52.Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng. Zero-shot learning through cross-modal transfer. In Advances in Neural Information Processing Systems, pages 935–943, 2013. 2
  53. 53.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Advances in Neural Information Processing Systems, 33, 2020. 1, 2, 3, 7, 8
  54. 54.Kaihua Tang, Jianqiang Huang, and Hanwang Zhang. Longtailed classification by keeping the good and removing the bad momentum causal effect. Advances in Neural Information Processing Systems, 33, 2020. 2, 5, 6
  55. 55.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780, 2017. 2
  56. 56.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8769–8778, 2018. 1
  57. 57.Feng Wang, Weiyang Liu, Haijun Liu, and Jian Cheng. Additive margin softmax for face verification. arXiv preprint arXiv:1801.05599, 2018. 2
  58. 58.Wei Wang, Vincent W Zheng, Han Yu, and Chunyan Miao. A survey of zero-shot learning: Settings, methods, and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–37, 2019. 2
  59. 59.Xudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu, and Stella Yu. Long-tailed recognition by routing diverse distribution-aware experts. In International Conference on Learning Representations, 2021. 1, 2
  60. 60.Xudong Wang, Long Lian, and Stella X Yu. Data-centric semi-supervised learning. arXiv preprint arXiv:2110.03006, 2021. 2
  61. 61.Xudong Wang, Ziwei Liu, and Stella X Yu. Unsupervised feature learning by cross-level instance-group discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12586–12595, 2021. 5
  62. 62.Chen Wei, Kihyuk Sohn, Clayton Mellina, Alan Yuille, and Fan Yang. Crest: A class-rebalancing self-training framework for imbalanced semi-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10857–10866, 2021. 2, 7
  63. 63.Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. Unsupervised data augmentation for consistency training. arXiv preprint arXiv:1904.12848, 2019. 7
  64. 64.Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. Unsupervised data augmentation for consistency training. Advances in Neural Information Processing Systems, 33, 2020. 2
  65. 65.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10687–10698, 2020. 1, 2
  66. 66.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016. 7
  67. 67.Dong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua, and Qianru Sun. Causal intervention for weakly-supervised semantic segmentation. Advances in Neural Information Processing Systems, 33, 2020. 4
  68. 68.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. arXiv preprint arXiv:2109.01134, 2021. 7, 8

Citation

MLA
Wang, X., et al. “Debiased Learning from Naturally Imbalanced Pseudo-Labels”. arXiv, 2022, http://arxiv.org/abs/2201.01490v2.
APA
Wang, X., Wu, Z., Lian, L., & Yu, S. X. (2022). Debiased Learning from Naturally Imbalanced Pseudo-Labels. arXiv. http://arxiv.org/abs/2201.01490v2
Chicago
Wang, X., Z. Wu, L. Lian, and S. X. Yu. 2022. “Debiased Learning from Naturally Imbalanced Pseudo-Labels”. arXiv. http://arxiv.org/abs/2201.01490v2.
Harvard
Wang, X. et al. (2022) “Debiased Learning from Naturally Imbalanced Pseudo-Labels”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.01490v2.
Vancouver
1. Wang X, Wu Z, Lian L, Yu SX (2022) Debiased Learning from Naturally Imbalanced Pseudo-Labels. arXiv

BibTeX

@article{wang2022debiased,
  title = {Debiased Learning from Naturally Imbalanced Pseudo-Labels},
  author = {Wang, Xudong and Wu, Zhirong and Lian, Long and Yu, Stella X.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.01490v2},
  eprint = {2201.01490}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE