FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

Bowen ZhangYidong WangWenxin HouHao WuJindong WangManabu OkumuraTakahiro Shinozaki

article2021NeurIPS1,368 citations

Proposes Curriculum Pseudo Labeling and the resulting FlexMatch algorithm, which dynamically adjusts class-specific confidence thresholds during training to significantly improve semi-supervised learning accuracy and convergence speed without adding computational overhead.

Listen

Modern semi-supervised machine learning aims to build accurate predictive models while minimizing the high cost and labor required to manually label large volumes of data. Leading methods typically assign artificial labels to unlabeled data points only when model prediction confidence exceeds a rigid, uniform threshold. However, this one-size-fits-all approach ignores the fact that different categories vary in difficulty. As a result, standard methods underutilize data from difficult classes, especially in early training stages or under severe label scarcity.

The article evaluates whether dynamically adjusting confidence thresholds based on real-time learning status can enhance model accuracy and training efficiency. To achieve this, the authors introduce Curriculum Pseudo Labeling, which monitors the volume of confident predictions per class and lowers thresholds for harder, less-learned classes without requiring separate validation sets or additional computing overhead. Integrating this technique with the leading baseline model produces an improved algorithm termed FlexMatch.

Evaluation across multiple standard computer vision benchmarks demonstrates three primary findings. First, FlexMatch substantially cuts error rates when labeled data is extremely scarce: on CIFAR-100 with only 4 labeled examples per class, error fell from 46.42% to 39.94%, and on STL-10 with 40 total labels, error dropped from 35.97% to 29.15% (a 18.96% relative reduction). Second, the curriculum method accelerates training speeds by more than fivefold, allowing FlexMatch to surpass standard final model accuracy in less than 20% of the iterations. Third, the dynamic thresholding strategy consistently enhanced other popular semi-supervised algorithms, confirming broad generalizability.

These results demonstrate significant practical value by reducing the time, computing expenses, and data annotation costs necessary to achieve state-of-the-art model performance. For practitioners and decision-makers implementing machine learning in data-constrained settings, adopting dynamic curriculum pseudo-labeling offers an immediate performance boost at negligible computational cost. The authors have released an open-source codebase to facilitate adoption.

Confidence in these findings is high for balanced and complex classification tasks, but decision-makers should exercise caution in settings with heavy class imbalance. On the imbalanced digit recognition dataset evaluated in the article, dynamic thresholding underperformed fixed thresholds due to persistent threshold suppression in smaller classes. Future efforts and pilot deployments should focus on extending dynamic threshold adjustments to long-tailed, highly unbalanced operational environments.

arXiv: 2110.08263

No sufficiently relevant recommendations were found.

Cover for FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling

Abstract

The recently proposed FixMatch achieved state-of-the-art results on most semi-supervised learning (SSL) benchmarks. However, like other modern SSL algorithms, FixMatch uses a pre-defined constant threshold for all classes to select unlabeled data that contribute to the training, thus failing to consider different learning status and learning difficulties of different classes. To address this issue, we propose Curriculum Pseudo Labeling (CPL), a curriculum learning approach to leverage unlabeled data according to the model's learning status. The core of CPL is to flexibly adjust thresholds for different classes at each time step to let pass informative unlabeled data and their pseudo labels. CPL does not introduce additional parameters or computations (forward or backward propagation). We apply CPL to FixMatch and call our improved algorithm FlexMatch. FlexMatch achieves state-of-the-art performance on a variety of SSL benchmarks, with especially strong performances when the labeled data are extremely limited or when the task is challenging. For example, FlexMatch achieves 13.96% and 18.96% error rate reduction over FixMatch on CIFAR-100 and STL-10 datasets respectively, when there are only 4 labels per class. CPL also significantly boosts the convergence speed, e.g., FlexMatch can use only 1/5 training time of FixMatch to achieve even better performance. Furthermore, we show that CPL can be easily adapted to other SSL algorithms and remarkably improve their performances. We open-source our code at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 FlexMatch
  • 3.1 Curriculum Pseudo Labeling
  • 3.2 Threshold warm-up
  • 3.3 Non-linear mapping function
  • 4 Experiments
  • 4.1 Main results
  • 4.2 Results on ImageNet
  • 4.3 Convergence speed acceleration
  • 4.4 Ablation study
  • 5 Related Work
  • 6 Conclusion and Future Work
  • References
  • A Experimental Results
  • A.1 Hyperparameter setting
  • A.2 Class-wise accuracy improvement.
  • A.3 Median error rates
  • A.4 Detailed results
  • B TorchSSL: A PyTorch-based SSL Codebase
  • B.1 BatchNorm Controller
  • B.2 Benchmark results

Knowls

  1. Knowl 1 — Curriculum Pseudo Labeling (CPL)

    model/method

    Curriculum Pseudo Labeling (CPL) is a semi-supervised learning framework that dynamically adjusts pseudo-label selection thresholds on a per-class basis according to the model's learning status at each training time step.

    Standard threshold-based semi-supervised learning algorithms (such as FixMatch, Unsupervised Data Augmentation (UDA), and Pseudo-Labeling) use a single pre-defined, constant confidence threshold τ\tau for all classes. This treats all classes equally and ignores varying class learning difficulties, filtering out a large portion of unlabeled data for hard-to-learn classes especially early in training.

    CPL replaces the fixed scalar threshold with a dynamic class-dependent threshold Tt(c)T_t(c) for each class c∈{1,…,C}c \in \{1, \dots, C\} at training iteration tt. The threshold Tt(c)T_t(c) is computed from an internal estimate of how well class cc has been learned: classes that have fewer confident predictions receive lower thresholds, encouraging the model to incorporate more unlabeled training samples for difficult classes, while well-learned classes maintain thresholds near the maximum threshold τ\tau.

    CPL operates without requiring an external labeled validation set, extra network forward or backward passes, or additional learned parameters. The learning status metrics are maintained as running counts updated during the standard forward pass of consistency loss calculation.

  2. Knowl 2 — Formulation of Flexible Per-Class Thresholds in CPL

    equation

    Let NN denote the total number of unlabeled samples, τ∈(0,1)\tau \in (0, 1) denote a fixed base confidence threshold, and pm,t(y∣un)p_{m,t}(y \mid u_n) denote the predicted class probability distribution produced by model mm for unlabeled sample unu_n under weak augmentation at training step tt.

    The estimated learning status σt(c)\sigma_t(c) for class c∈{1,…,C}c \in \{1, \dots, C\} is the count of unlabeled samples currently assigned to class cc with prediction confidence exceeding τ\tau: σt(c)=∑n=1N1(max⁡ypm,t(y∣un)>τ)⋅1(arg⁡max⁡ypm,t(y∣un)=c)\sigma_t(c) = \sum_{n=1}^N \mathbf{1}\left(\max_y p_{m,t}(y \mid u_n) > \tau\right) \cdot \mathbf{1}\left(\arg\max_y p_{m,t}(y \mid u_n) = c\right)

    To account for initialization bias and prevent premature threshold escalation when few samples are confident, a warm-up normalization factor βt(c)∈[0,1]\beta_t(c) \in [0, 1] is computed by treating the unused unlabeled samples as an additional virtual class: βt(c)=σt(c)max⁡(max⁡c′σt(c′),  N−∑c′=1Cσt(c′))\beta_t(c) = \frac{\sigma_t(c)}{\max\left(\max_{c'} \sigma_t(c'),\; N - \sum_{c'=1}^C \sigma_t(c')\right)}

    The dynamic threshold Tt(c)T_t(c) for class cc at step tt is calculated via a non-linear mapping function M:[0,1]→[0,1]\mathcal{M}: [0, 1] \to [0, 1]: Tt(c)=M(βt(c))⋅τT_t(c) = \mathcal{M}(\beta_t(c)) \cdot \tau where M(x)=x2−x\mathcal{M}(x) = \frac{x}{2 - x} is a strictly increasing convex function that makes the threshold more sensitive as class learning approaches maturity.

  3. Knowl 3 — FlexMatch Optimization Objective

    equation

    FlexMatch integrates Curriculum Pseudo Labeling (CPL) into the FixMatch framework. For a labeled batch {(xb,yb)}b=1B\{(x_b, y_b)\}_{b=1}^B and an unlabeled batch {ub}b=1μB\{u_b\}_{b=1}^{\mu B} (where μ\mu is the unlabeled-to-labeled batch ratio), the overall loss at time step tt is: Lt=Ls+λLu,tL_t = L_s + \lambda L_{u,t} where λ\lambda is the unsupervised loss weight hyperparameter.

    The supervised loss LsL_s is the standard cross-entropy loss HH evaluated on weakly-augmented labeled data ω(xb)\omega(x_b): Ls=1B∑b=1BH(yb,pm(y∣ω(xb)))L_s = \frac{1}{B} \sum_{b=1}^B H\left(y_b, p_m(y \mid \omega(x_b))\right)

    The unsupervised consistency loss Lu,tL_{u,t} evaluates cross-entropy between the one-hot pseudo-label q^b=arg⁡max⁡(qb)\hat{q}_b = \arg\max(q_b) generated from the weakly-augmented sample qb=pm(y∣ω(ub))q_b = p_m(y \mid \omega(u_b)) and the prediction on the strongly-augmented sample pm(y∣Ω(ub))p_m(y \mid \Omega(u_b)), gated by the dynamic per-class threshold TtT_t: Lu,t=1μB∑b=1μB1(max⁡(qb)>Tt(arg⁡max⁡(qb)))H(q^b,pm(y∣Ω(ub)))L_{u,t} = \frac{1}{\mu B} \sum_{b=1}^{\mu B} \mathbf{1}\left(\max(q_b) > T_t(\arg\max(q_b))\right) H\left(\hat{q}_b, p_m(y \mid \Omega(u_b))\right) where ω(⋅)\omega(\cdot) is a weak data augmentation (e.g., random crop and horizontal flip) and Ω(⋅)\Omega(\cdot) is a strong data augmentation (e.g., RandAugment).

  4. Knowl 4 — FlexMatch Training Procedure

    algorithm

    The complete training algorithm of FlexMatch maintains a lookup vector of recent confident predictions for all unlabeled samples to dynamically compute per-class thresholds before calculating supervised and unsupervised losses.

    Input: Labeled set X={(xm,ym):m∈{1,…,M}}\mathcal{X} = \{(x_m, y_m) : m \in \{1, \dots, M\}\}, unlabeled set U={un:n∈{1,…,N}}\mathcal{U} = \{u_n : n \in \{1, \dots, N\}\}, base threshold τ\tau, mapping M(x)=x/(2−x)\mathcal{M}(x) = x / (2 - x), labeled batch size BB, unlabeled ratio μ\mu, loss weight λ\lambda, total iterations KK
    Output: Trained model parameters
    Initialize prediction memory for all unlabeled samples: u^n←−1\hat{u}_n \leftarrow -1 for all n∈{1,…,N}n \in \{1, \dots, N\}
    for step t=1t = 1 to KK do
        for c=1c = 1 to CC do
            σ(c)←∑n=1N1(u^n=c)\sigma(c) \leftarrow \sum_{n=1}^N \mathbf{1}(\hat{u}_n = c)
        end for
        Nunused←∑n=1N1(u^n=−1)N_{\text{unused}} \leftarrow \sum_{n=1}^N \mathbf{1}(\hat{u}_n = -1)
        for c=1c = 1 to CC do
            if max⁡c′σ(c′)<Nunused\max_{c'} \sigma(c') < N_{\text{unused}} then
                β(c)←σ(c)/Nunused\beta(c) \leftarrow \sigma(c) / N_{\text{unused}}
            else
                β(c)←σ(c)/max⁡c′σ(c′)\beta(c) \leftarrow \sigma(c) / \max_{c'} \sigma(c')
            end if
            T(c)←M(β(c))⋅τT(c) \leftarrow \mathcal{M}(\beta(c)) \cdot \tau
        end for
        Sample labeled batch {(xb,yb)}b=1B\{(x_b, y_b)\}_{b=1}^B from X\mathcal{X} and unlabeled batch {ub}b=1μB\{u_b\}_{b=1}^{\mu B} from U\mathcal{U}
        for b=1b = 1 to μB\mu B do
            qb←pm(y∣ω(ub))q_b \leftarrow p_m(y \mid \omega(u_b))
            if max⁡(qb)>τ\max(q_b) > \tau then
                u^b←arg⁡max⁡(qb)\hat{u}_b \leftarrow \arg\max(q_b)
            end if
        end for
        Compute LsL_s on labeled batch using weak augmentation
        Compute Lu,tL_{u,t} on unlabeled batch using threshold T(arg⁡max⁡(qb))T(\arg\max(q_b))
        Compute total loss Lt=Ls+λLu,tL_t = L_s + \lambda L_{u,t} and update model parameters via SGD
    end for
    return Model parameters
  5. Knowl 5 — Error Rates Across Standard Semi-Supervised Classification Benchmarks

    data/table

    The table below presents the best error rates (%) across three distinct random seeds on CIFAR-10, CIFAR-100, STL-10, and SVHN under varying numbers of labeled training samples. Base methods (Pseudo-Labeling [PL], UDA, and FixMatch) are evaluated alongside their CPL-augmented counterparts (Flex-PL, Flex-UDA, and FlexMatch).

    Dataset CIFAR-10 CIFAR-100 STL-10 SVHN
    Label Amount 40 250 4000 400 2500 10000 40 250 1000 40 1000
    PL 74.610.26 46.492.20 15.080.19 87.450.85 57.740.28 36.550.24 74.680.99 55.452.43 32.640.71 64.615.60 9.400.32
    Flex-PL 73.741.96 46.141.81 14.750.19 85.720.46 56.120.51 35.600.15 73.422.19 52.062.50 32.050.37 63.213.64 12.050.54
    UDA 10.623.75 5.160.06 4.290.07 46.391.59 27.730.21 22.490.23 37.428.44 9.721.15 6.640.17 5.124.27 1.890.01
    Flex-UDA 5.440.52 5.020.07 4.240.06 45.171.88 27.080.15 21.910.10 29.532.10 9.030.45 6.100.25 3.421.51 2.020.05
    FixMatch 7.470.28 4.860.05 4.210.08 46.420.82 28.030.16 22.200.12 35.974.14 9.811.04 6.250.33 3.811.18 1.960.03
    FlexMatch 4.970.06 4.980.09 4.190.01 39.941.62 26.490.20 21.900.15 29.154.16 8.230.39 5.770.18 8.193.20 6.720.30
    Fully-Supervised 4.620.05 19.300.09 - 2.130.02

    FlexMatch achieves state-of-the-art results across most configurations, demonstrating the largest performance gains under extremely limited label regimes (e.g., reducing error from 46.42% to 39.94% on CIFAR-100 with 400 labels, and from 35.97% to 29.15% on STL-10 with 40 labels). Adding CPL to UDA and Pseudo-Labeling also consistently improves their performance.

  6. Knowl 6 — Convergence Acceleration and Early Class-Wise Learning Dynamics

    empirical result

    FlexMatch significantly accelerates training convergence relative to FixMatch:

    1. Early convergence: On CIFAR-100 with 400 labels (4 labels per class), FlexMatch surpasses the final accuracy achieved by FixMatch after only 50K iterations (less than 1/5 of FixMatch's total training time to reach comparable performance).
    2. Early-stage class-wise accuracy: On CIFAR-10 with 40 labels at 200K iterations, FixMatch attains an overall accuracy of 56.35% because approximately half of the classes remain poorly learned due to their high fixed threshold. In contrast, FlexMatch achieves 94.29% accuracy at 200K iterations (surpassing FixMatch's final accuracy after 1M iterations) by lowering thresholds for difficult classes early on, allowing balanced class utilization.
  7. Knowl 7 — Ablation on Base Threshold, Threshold Mapping Function, and Warm-up

    empirical result

    Ablation experiments on CIFAR-10 (40 labels) and CIFAR-100 (400 labels) establish the following architectural properties for FlexMatch:

    1. Base threshold τ\tau: Testing τ∈{0.90,0.93,0.95,0.97,0.99}\tau \in \{0.90, 0.93, 0.95, 0.97, 0.99\} reveals an optimal error rate around τ=0.95\tau = 0.95. Tuning τ\tau affects both the upper threshold bound and the counting statistics σt(c)\sigma_t(c).
    2. Mapping function curvature M(x)\mathcal{M}(x):
      • Convex mapping M(x)=x2−x\mathcal{M}(x) = \frac{x}{2 - x} yields the lowest error rate (~4.67%).
      • Linear mapping M(x)=x\mathcal{M}(x) = x yields an intermediate error rate (~4.89%).
      • Concave mapping M(x)=ln⁡(x+1)ln⁡2\mathcal{M}(x) = \frac{\ln(x + 1)}{\ln 2} yields the worst error rate (~4.98%). The convex function keeps thresholds low while a class is difficult and raises thresholds more sharply only once learning becomes confident.
    3. Threshold warm-up: Incorporating the warm-up term N−∑cσt(c)N - \sum_c \sigma_t(c) into the normalization denominator provides an absolute error rate reduction of ~0.2% on CIFAR-10 (40 labels) and ~1.0% on CIFAR-100 (400 labels) by preventing unstable threshold spikes at initialization.
  8. Knowl 8 — Semi-Supervised Classification on ImageNet-1K

    data/table

    The table below shows Top-1 and Top-5 error rates (%) on ImageNet-1K using ResNet-50 trained for 2202^{20} iterations with 100K labeled examples (100 labels per class, representing less than 8% of the dataset) and unlabeled batch ratio μ=1\mu = 1, with base threshold τ=0.7\tau = 0.7.

    Method Top-1 Error Rate (%) Top-5 Error Rate (%)
    FixMatch 43.66 21.80
    FlexMatch 41.85 19.48

    FlexMatch reduces the Top-1 error rate by 1.81% and Top-5 error rate by 2.32% compared to FixMatch under identical hyperparameters, demonstrating the scalability of dynamic curriculum thresholding to large-scale, 1000-class datasets.

  9. Knowl 9 — Performance Degradation of CPL on Simple Class-Imbalanced Datasets

    limitation

    Curriculum Pseudo Labeling experiences performance degradation on datasets that are both inherently easy to classify and class-imbalanced, such as SVHN (where FlexMatch achieves 8.19% error with 40 labels and 6.72% with 1000 labels, underperforming FixMatch at 3.81% and 1.96%).

    Because the learning status σt(c)\sigma_t(c) is estimated via raw sample counts exceeding τ\tau, classes with intrinsically fewer examples in the unlabeled data yield small σt(c)\sigma_t(c) values. Consequently, βt(c)\beta_t(c) remains well below 1.0 even after the model has fully mastered the underrepresented class. The resulting persistently low flexible threshold Tt(c)T_t(c) allows noisy, incorrect pseudo-labels to pass through during training, causing fluctuations and suboptimal generalization.

  10. Knowl 10 — BatchNorm Controller for Semi-Supervised Model Training

    model/method

    When semi-supervised learning methods (such as Mean Teacher, Π\Pi-Model, and MixMatch) process labeled and unlabeled data in separate forward propagation passes, updating Batch Normalization (BatchNorm) running statistics on both labeled and unlabeled batches sequentially introduces significant training instability.

    The BatchNorm Controller solves this by updating BatchNorm statistics strictly during the forward pass of labeled data. Before performing forward propagation on unlabeled data, the controller records the current running mean and running variance of all BatchNorm layers and restores these statistics immediately after the unlabeled propagation pass completes, preventing unlabeled batch dynamics from distorting the normalization layers.

Coverage note — None was omitted; all key theoretical components, equations, algorithms, empirical benchmarks (CIFAR-10/100, STL-10, SVHN, ImageNet-1K), convergence dynamics, ablations, limitations, and codebase contributions (TorchSSL and BatchNorm Controller) are represented.

References

  1. 1.Philip Bachman, Ouais Alsharif, and Doina Precup. Learning with pseudo-ensembles. In NeurIPS, pages 3365–3373, 2014.
  2. 2.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. ICLR, 2017.
  3. 3.Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. Regularization with stochastic trans- formations and perturbations for deep semi-supervised learning. In NeurIPS, pages 1171–1179, 2016.
  4. 4.Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, 2013.
  5. 5.Geoffrey J McLachlan. Iterative reclassification procedure for constructing an asymptotically op- timal rule of allocation in discriminant analysis. Journal of the American Statistical Association, 70(350):365–369, 1975.
  6. 6.Chuck Rosenberg, Martial Hebert, and Henry Schneiderman. Semi-supervised self-training of object detection models. 2005.
  7. 7.Henry Scudder. Probability of error of some adaptive pattern-recognition machines. IEEE Transactions on Information Theory, 11(3):363–371, 1965.
  8. 8.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In CVPR, pages 10687–10698, 2020.
  9. 9.Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, and Tapani Raiko. Semi- supervised learning with ladder networks. In NeurIPS, pages 3546–3554, 2015.
  10. 10.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NeurIPS, pages 1195– 1204, 2017.
  11. 11.Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. Unsupervised data augmen- tation for consistency training. NeurIPS, 33, 2020.
  12. 12.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. NeurIPS, page 5050–5060, 2019.
  13. 13.David Berthelot, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring. In ICLR, 2019.
  14. 14.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raf- fel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fixmatch: Simplifying semi- supervised learning with consistency and confidence. NeurIPS, 33, 2020.
  15. 15.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In ICML, pages 41–48, 2009.
  16. 16.Yves Grandvalet, Yoshua Bengio, et al. Semi-supervised learning by entropy minimization. In CAP, pages 281–296, 2005.
  17. 17.Anastasia Pentina, Viktoriia Sharmanska, and Christoph H Lampert. Curriculum learning of multiple tasks. In CVPR, pages 5492–5500, 2015.
  18. 18.Lu Jiang, Deyu Meng, Qian Zhao, Shiguang Shan, and Alexander Hauptmann. Self-paced curriculum learning. In AAAI, volume 29, 2015.
  19. 19.Alex Graves, Marc G Bellemare, Jacob Menick, Remi Munos, and Koray Kavukcuoglu. Auto- mated curriculum learning for neural networks. In ICML, pages 1311–1320. PMLR, 2017.
  20. 20.Guy Hacohen and Daphna Weinshall. On the power of curriculum learning in training deep networks. In ICML, pages 2535–2544. PMLR, 2019.
  21. 21.Daphna Weinshall, Gad Cohen, and Dan Amir. Curriculum learning by transfer learning: Theory and experiments with deep networks. In ICML, pages 5238–5246. PMLR, 2018.
  22. 22.Eric Arazo, Diego Ortego, Paul Albert, Noel E O’Connor, and Kevin McGuinness. Pseudo- labeling and confirmation bias in deep semi-supervised learning. In IJCNN, pages 1–8, 2020.
  23. 23.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  24. 24.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  25. 25.Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsuper- vised feature learning. In AISTATS, pages 215–223, 2011.
  26. 26.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009.
  27. 27.Avital Oliver, Augustus Odena, Colin Raffel, Ekin D Cubuk, and Ian J Goodfellow. Realistic evaluation of deep semi-supervised learning algorithms. In NeurIPS, pages 3239–3250, 2018.
  28. 28.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 32:8026–8037, 2019.
  29. 29.Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. On the importance of initialization and momentum in deep learning. In ICML, pages 1139–1147. PMLR, 2013.
  30. 30.Boris T Polyak. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964.
  31. 31.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. ICLR, 2016.
  32. 32.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  33. 33.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  34. 34.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In BMVC, 2016.
  35. 35.Tianyi Zhou, Shengjie Wang, and Jeff Bilmes. Time-consistent self-supervision for semi- supervised learning. In ICML, pages 11523–11533. PMLR, 2020.
  36. 36.Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning. arXiv preprint arXiv:2101.06329, 2021.
  37. 37.Chen Gong, Dacheng Tao, Stephen J Maybank, Wei Liu, Guoliang Kang, and Jie Yang. Multi- modal curriculum learning for semi-supervised image classification. IEEE Transactions on Image Processing, 25(7):3249–3260, 2016.
  38. 38.Hoel Kervadec, Jose Dolz, Éric Granger, and Ismail Ben Ayed. Curriculum semi-supervised segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 568–576. Springer, 2019.
  39. 39.Qing Yu, Daiki Ikami, Go Irie, and Kiyoharu Aizawa. Multi-task curriculum framework for open-set semi-supervised learning. In ECCV, pages 438–454. Springer, 2020.
  40. 40.Yue Han, Yuhong Liu, and Zhigang Jin. Sentiment analysis via semi-supervised learning: a model based on dynamic threshold and multi-classifiers. Neural Computing and Applications, 32(9):5117–5129, 2020.
  41. 41.Zhedong Zheng and Yi Yang. Rectifying pseudo label learning via uncertainty estimation for domain adaptive semantic segmentation. IJCV, 129(4):1106–1120, 2021.
  42. 42.Paola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, and Vicente Ordonez. Curriculum labeling: Revisiting pseudo-labeling for semi-supervised learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 6912–6920, 2021.
  43. 43.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. fixmatch. https://github.com/google-research/fixmatch, 2020.
  44. 44.Lee Doyup and Cheon Yeongjae. Fixmatch-pytorch. https://github.com/LeeDoYup/FixMatch-pytorch, 2020.
  45. 45.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE TPAMI, 41(8):1979– 1993, 2018.
  46. 46.Hang Zhang, Kristin Dana, Jianping Shi, Zhongyue Zhang, Xiaogang Wang, Ambrish Tyagi, and Amit Agrawal. Context encoding for semantic segmentation. In CVPR, pages 7151–7160, 2018.

Citation

MLA
Zhang, B., et al. “FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling”. arXiv, 2021, http://arxiv.org/abs/2110.08263v3.
APA
Zhang, B., Wang, Y., Hou, W., Wu, H., Wang, J., Okumura, M., & Shinozaki, T. (2021). FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling. arXiv. http://arxiv.org/abs/2110.08263v3
Chicago
Zhang, B., Y. Wang, W. Hou, et al. 2021. “FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling”. arXiv. http://arxiv.org/abs/2110.08263v3.
Harvard
Zhang, B. et al. (2021) “FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.08263v3.
Vancouver
1. Zhang B, Wang Y, Hou W, Wu H, Wang J, Okumura M, Shinozaki T (2021) FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling. arXiv

BibTeX

@article{zhang2021flexmatch,
  title = {FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling},
  author = {Zhang, Bowen and Wang, Yidong and Hou, Wenxin and Wu, Hao and Wang, Jindong and Okumura, Manabu and Shinozaki, Takahiro},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.08263v3},
  eprint = {2110.08263}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/