Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data

Yuhao ChenXin TanBorui ZhaoZhaowei ChenRenjie SongJiajun LiangXuequan Lu

article2023CVPR64 citations

Presents FullMatch, a semi-supervised learning framework that fully exploits ambiguous unlabeled samples by suppressing non-target class competition and adaptively assigning negative pseudo-labels without introducing extra hyperparameters.

Listen

Modern computer vision models typically require large volumes of manually labeled data, which is expensive and time-consuming to create. Semi-supervised learning methods address this challenge by combining a small amount of labeled data with abundant unlabeled data. However, prevailing state-of-the-art frameworks, such as FixMatch, rely on rigid, high confidence thresholds to assign positive labels. As a result, they discard ambiguous or low-confidence examples, leaving a significant portion of unlabeled training data entirely unexploited during the early and middle stages of training.

The article introduces and evaluates FullMatch, a semi-supervised learning framework designed to utilize all unlabeled examples without adding new hyperparameters or computational overhead. The core objective is to improve model accuracy and data efficiency by extracting meaningful supervisory signals from both high-confidence and low-confidence unlabeled samples.

The researchers developed two complementary techniques and tested them across multiple standardized image benchmarks, including CIFAR-10, CIFAR-100, SVHN, STL-10, and ImageNet. The first technique, Entropy Meaning Loss, forces non-target classes to share remaining confidence uniformly, preventing competition with the target class and producing clearer decision boundaries. The second technique, Adaptive Negative Learning, identifies categories the model is confident an image does not belong to and assigns negative labels. This ranking cutoff is determined dynamically by comparing predictions across different data augmentations, avoiding the need for a separate validation set. The authors evaluated the approach across varying labeled data constraints and integrated it with existing baseline architectures.

The key findings indicate that FullMatch consistently outperforms baseline methods. In extremely data-scarce settings, such as four labeled examples per class, FullMatch improved average accuracy over FixMatch by more than one percentage point, including an approximate two percentage point increase on CIFAR-100. On the large-scale ImageNet benchmark, FullMatch improved top-1 accuracy by 1.1 percentage points over FixMatch. Furthermore, when combined with FlexMatch—an advanced framework utilizing curriculum pseudo-labeling—the resulting model established new state-of-the-art accuracy across nearly all evaluated datasets. Ablation experiments confirmed that both proposed techniques contribute complementary gains while adding negligible training time.

These findings demonstrate that ambiguous unlabeled data contains valuable information that can be safely harvested using negative pseudo-labeling and non-target class distribution constraints. For organizations deploying machine learning systems, this approach reduces the risk and cost associated with acquiring large labeled datasets while improving model performance. Because the framework does not introduce extra hyperparameters, it is practical to implement and easily integrates into existing semi-supervised training pipelines.

Organizations developing computer vision systems with limited annotations should adopt these non-target supervision and adaptive negative labeling techniques into their current workflows. Future work should evaluate this approach on additional visual tasks, such as object detection and segmentation, as well as complex real-world data containing severe class imbalance or out-of-distribution noise.

The reported conclusions carry high confidence based on consistent empirical gains across multiple standardized benchmarks and repeated experimental trials. However, readers should note that the evaluations were conducted on curated image classification benchmarks, meaning results may vary in non-vision domains or on less structured industrial datasets.

No sufficiently relevant recommendations were found.

Cover for Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data

Abstract

Semi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch) use a combination of consistency regularization and pseudo-labeling to achieve remarkable successes. However, these methods all suffer from the waste of complicated examples since all pseudo-labels have to be selected by a high threshold to filter out noisy ones. Hence, the examples with ambiguous predictions will not contribute to the training phase. For better leveraging all unlabeled examples, we propose two novel techniques: Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL). EML incorporates the prediction distribution of non-target classes into the optimization objective to avoid competition with target class, and thus generating more high-confidence predictions for selecting pseudo-label. ANL introduces the additional negative pseudo-label for all unlabeled data to leverage low-confidence examples. It adaptively allocates this label by dynamically evaluating the top-k performance of the model. EML and ANL do not introduce any additional parameter and hyperparameter. We integrate these techniques with FixMatch, and develop a simple yet powerful framework called FullMatch. Extensive experiments on several common SSL benchmarks (CIFAR-10/100, SVHN, STL-10 and ImageNet) demonstrate that FullMatch exceeds FixMatch by a large margin. Integrated with FlexMatch (an advanced FixMatch-based framework), we achieve state-of-the-art performance. Source code is available at https://github.com/megvii-research/FullMatch.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Entropy Meaning Loss
  • 3.3. Adaptive Negative Learning
  • 3.4. FullMatch
  • 4. Experiments
  • 4.1. Main Results
  • 4.2. Results on ImageNet
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Entropy Meaning Loss (EML) for Non-Target Classes

    model/method

    In pseudo-labeling semi-supervised learning methods (such as FixMatch), predictions on challenging unlabeled examples often fail to exceed the confidence threshold τ\tau due to competition between confusing candidate classes and the target class. Entropy Meaning Loss (EML) resolves this by supervising the non-target classes of samples that receive a positive pseudo-label, forcing the model to distribute the remaining prediction probability uniformly across all non-target classes.

    Let Q(i)=[q1(i),…,qC(i)]Q^{(i)} = [q_1^{(i)}, \dots, q_C^{(i)}] be the model prediction vector for the weakly-augmented version of unlabeled sample ii, where CC is the total number of classes. The target class indicator vector is S(i)=[s1(i),…,sC(i)]S^{(i)} = [s_1^{(i)}, \dots, s_C^{(i)}] where sc(i)=1(qc(i)≥τ)s_c^{(i)} = \mathbf{1}(q_c^{(i)} \ge \tau). The non-target class indicator vector U(i)=[u1(i),…,uC(i)]U^{(i)} = [u_1^{(i)}, \dots, u_C^{(i)}] identifies non-target categories for pseudo-labeled samples: uc(i)=1(max⁡(Q(i))≥τ)⋅1(sc(i)=0)u_c^{(i)} = \mathbf{1}\left(\max(Q^{(i)}) \ge \tau\right) \cdot \mathbf{1}\left(s_c^{(i)} = 0\right)

    Let P(i)=[p1(i),…,pC(i)]P^{(i)} = [p_1^{(i)}, \dots, p_C^{(i)}] denote the prediction confidence vector on the strongly-augmented version of sample ii. The target soft label yc(i)y_c^{(i)} assigned to non-target class cc is defined such that all non-target classes equally share the remaining confidence 1−ptc(i)1 - p_{tc}^{(i)} (where ptc(i)p_{tc}^{(i)} is the target class score): yc(i)=1−1(uc(i)=0)⋅pc(i)∑j=1C1(uj(i)=1)y_c^{(i)} = \frac{1 - \mathbf{1}(u_c^{(i)} = 0) \cdot p_c^{(i)}}{\sum_{j=1}^C \mathbf{1}(u_j^{(i)} = 1)}

    The Entropy Meaning Loss LemlL_{eml} is computed over batch size BB using binary cross-entropy (BCE): Leml=−1BC∑i=1B∑c=1Cuc(i)[yc(i)log⁡(pc(i))+(1−yc(i))log⁡(1−pc(i))]L_{eml} = -\frac{1}{B C} \sum_{i=1}^B \sum_{c=1}^C u_c^{(i)} \left[ y_c^{(i)} \log(p_c^{(i)}) + (1 - y_c^{(i)}) \log(1 - p_c^{(i)}) \right]

  2. Knowl 2 — Adaptive Negative Learning (ANL) for Ambiguous Predictions

    model/method

    Adaptive Negative Learning (ANL) assigns negative pseudo-labels to categories that an unlabeled sample confidently does not belong to, allowing low-confidence unlabeled samples (where max⁡(Q(i))<τ\max(Q^{(i)}) < \tau) to contribute to model training without requiring an external validation set or extra forward passes.

    ANL estimates model consistency across weak and strong augmentations. For a training batch at step tt, let Q^t=arg⁡max⁡(Q,t)\hat{Q}_t = \arg\max(Q, t) denote the temporary hard labels derived from weakly-augmented predictions QQ, and let PtP_t denote predictions on strongly-augmented views. ANL dynamically determines the smallest integer rank k∈[2,C]k \in [2, C] such that the top-kk predictions on PtP_t achieve 100% accuracy relative to pseudo-ground-truth Q^t\hat{Q}_t across the batch: k=arg⁡min⁡θ∈[2,C](Acc(Pt,Q^t,θ)=100%)k = \arg\min_{\theta \in [2, C]} \left( \text{Acc}(P_t, \hat{Q}_t, \theta) = 100\% \right) where Acc\text{Acc} computes the top-θ\theta classification accuracy across the batch.

    All categories ranked after the top-kk positions in the weakly-augmented prediction Q(i)Q^{(i)} (sorted in descending order of confidence via Rank(qc(i))\text{Rank}(q_c^{(i)})) are assigned negative pseudo-labels. The ANL loss LanlL_{anl} is formulated as: Lanl=−1B∑i=1B∑c=1C1[Rank(qc(i))>k]log⁡(1−pc(i))L_{anl} = -\frac{1}{B} \sum_{i=1}^B \sum_{c=1}^C \mathbf{1}\left[\text{Rank}(q_c^{(i)}) > k\right] \log(1 - p_c^{(i)}) where BB is the batch size, CC is the number of classes, and pc(i)p_c^{(i)} is the predicted probability for class cc under strong augmentation.

  3. Knowl 3 — FullMatch Semi-Supervised Learning Framework

    model/method

    FullMatch unifies consistency regularization, pseudo-labeling, Entropy Meaning Loss (EML), and Adaptive Negative Learning (ANL) into a single semi-supervised training framework.

    1. Negative Pseudo-Labeling via ANL: For all unlabeled samples in a batch, ANL dynamically evaluates batch-level weak-to-strong consistency to find the threshold rank kk, and penalizes low-rank classes via LanlL_{anl}. When an example receives a positive pseudo-label, the range of non-target classes constrained by EML is restricted to the top-kk classes excluding the target class, reducing the number of non-target classes from C−1C-1 to k−1k-1.
    2. Pseudo-Labeling Consistency: For samples whose maximum weakly-augmented confidence max⁡(Q(i))≥τ\max(Q^{(i)}) \ge \tau, standard cross-entropy consistency loss LusL_{us} is applied to the strongly-augmented prediction P(i)P^{(i)}, and EML loss LemlL_{eml} is applied to the non-target classes within the top-kk.

    The total training loss is: Lsum=Ls+Lus+α⋅Lanl+β⋅LemlL_{sum} = L_s + L_{us} + \alpha \cdot L_{anl} + \beta \cdot L_{eml} where:

    • Supervised loss on labeled batch BlB_l: Ls=1Bl∑i=1BlH(y(i),pm(y∣ω(x(i))))L_s = \frac{1}{B_l} \sum_{i=1}^{B_l} H(y^{(i)}, p_m(y|\omega(x^{(i)}))), with HH being cross-entropy, ω\omega denoting weak augmentation, and y(i)y^{(i)} being ground truth.
    • Unsupervised consistency loss on unlabeled batch BB: Lus=1B∑i=1B1(max⁡(Q(i))≥τ)H(Q^(i),P(i))L_{us} = \frac{1}{B} \sum_{i=1}^B \mathbf{1}(\max(Q^{(i)}) \ge \tau) H(\hat{Q}^{(i)}, P^{(i)}), with Q^(i)=arg⁡max⁡(Q(i))\hat{Q}^{(i)} = \arg\max(Q^{(i)}).
    • Loss weights are set to α=1.0\alpha = 1.0 and β=1.0\beta = 1.0 by default.

    When integrated with Curriculum Pseudo Labeling (FlexMatch), the resulting architecture is termed FullFlex.

  4. Knowl 4 — Gradient Alignment of Entropy Meaning Loss with Cross-Entropy Loss

    theoretical result

    For an unlabeled sample assigned a positive pseudo-label at target class index tc=arg⁡max⁡(Q(i))tc = \arg\max(Q^{(i)}), the Entropy Meaning Loss (EML) also produces a gradient with respect to the target class prediction ptc(i)p_{tc}^{(i)}, given by: gtc(i)=−1BC(C−1)log⁡(∏c=1,c≠tcC(1−pc(i))∏c=1,c≠tcCpc(i))g_{tc}^{(i)} = -\frac{1}{B C (C - 1)} \log\left( \frac{\prod_{c=1, c \neq tc}^C (1 - p_c^{(i)})}{\prod_{c=1, c \neq tc}^C p_c^{(i)}} \right) where BB is the batch size, CC is the number of classes, and pc(i)p_c^{(i)} is the predicted probability of class cc under strong augmentation.

    The gradient direction of gtc(i)g_{tc}^{(i)} is identical to the gradient direction of the standard unsupervised cross-entropy loss LusL_{us}. Consequently, EML cooperates with LusL_{us} to increase the confidence of the target class while simultaneously flattening the probability distribution across non-target categories.

  5. Knowl 5 — Top-1 Accuracy on Standard Semi-Supervised Benchmarks

    data/table

    The table compares top-1 accuracy (mean ±\pm standard deviation over 3 folds) of FullMatch and FullFlex against semi-supervised learning baselines on CIFAR-10, CIFAR-100, SVHN, and STL-10 under varying amounts of labeled training samples. Architectures used are Wide ResNet (WRN-28-2 and WRN-28-8). Training runs for 2202^{20} iterations with initial learning rate 0.03 and SGD optimizer (momentum 0.9, cosine decay).

    Dataset CIFAR-10 CIFAR-100 SVHN STL-10
    Label Amount 40 250 4000 400 2500 10000 40 1000 1000
    UDA 89.383.75 94.840.06 95.710.07 53.611.59 72.270.21 77.510.23 94.884.27 98.110.01 93.360.17
    RemixMatch 90.121.03 93.70.05 95.160.01 57.251.05 73.970.35 79.980.27 75.969.13 94.840.31 93.260.14
    Semco 92.130.22 94.880.27 96.200.08 55.891.18 68.070.01 75.550.12 - - 92.510.29
    Dash 86.783.75 95.440.13 95.920.06 55.240.96 72.820.21 78.030.14 96.971.59 97.970.06 92.740.40
    UPS 94.740.29 94.890.08 95.750.05 58.931.66 72.860.24 78.030.23 - - 93.980.28
    AlphaMatch 91.353.38 95.030.29 - 61.263.13 74.980.27 - 97.030.26 - 90.360.75
    CoMatch 93.120.92 95.100.35 95.940.03 59.981.11 72.990.21 78.170.23 - - 91.340.41
    SimMatch 94.401.37 95.160.39 96.040.01 62.192.21 74.930.32 79.420.11 - - -
    CR 94.310.9 94.960.3 95.840.13 50.770.79 72.420.37 78.970.23 96.331.84 97.610.06 93.040.42
    NP-Match 95.090.04 95.040.06 95.890.02 61.080.99 73.970.26 78.780.13 - - 94.410.24
    FixMatch 92.530.28 95.140.05 95.790.08 57.451.76 71.970.16 77.80.12 96.191.18 98.040.03 93.750.33
    FullMatch (ours) 94.111.01 95.360.12 96.250.08 59.421.40 73.060.40 78.560.10 97.650.10 98.010.03 94.260.09
    FlexMatch 95.030.06 95.020.09 95.810.01 60.061.62 73.510.2 78.10.15 96.081.24 97.370.06 94.230.18
    FullFlex (ours) 95.560.15 95.610.04 96.280.03 62.600.64 74.600.42 79.260.21 97.480.04 97.580.02 94.500.12

    FullMatch improves upon FixMatch across almost all settings, with the largest gains in the lowest-label regimes (e.g., +1.58% on CIFAR-10 with 40 labels, +1.97% on CIFAR-100 with 400 labels, and +1.46% on SVHN with 40 labels). Combining FullMatch with FlexMatch (FullFlex) consistently outperforms both baselines across all datasets.

  6. Knowl 6 — Semi-Supervised Image Classification Results on ImageNet

    data/table

    Evaluation on ImageNet using ResNet-50 trained for 2202^{20} iterations with 100 labeled samples per class (less than 8% of the total dataset):

    Method Top-1 (%) Top-5 (%)
    UPS 57.31 79.77
    NP-Match 58.22 80.67
    FixMatch 56.34 78.20
    FullMatch (ours) 57.44 (+1.10) 79.26 (+1.06)
    FlexMatch 58.15 80.52
    FullFlex (ours) 59.58 (+1.43) 81.38 (+0.86)

    FullMatch improves FixMatch by +1.10% Top-1 and +1.06% Top-5 accuracy, while FullFlex achieves 59.58% Top-1 accuracy, outperforming FlexMatch by +1.43% and establishing state-of-the-art semi-supervised classification performance on ImageNet.

  7. Knowl 7 — Ablation Study of EML and ANL Components

    data/table

    An ablation study on CIFAR-100 with 400 labeled samples demonstrates the individual and combined effects of Entropy Meaning Loss (EML) and Adaptive Negative Learning (ANL) over the baseline FixMatch (57.68% accuracy):

    Method CE BCE w PL w/o PL Accuracy (%) (%)
    FixMatch 57.68 -
    EML ✓ 58.35 +0.67
    EML ✓ 58.47 +0.79
    ANL ✓ 57.83 +0.15
    ANL ✓ 58.59 +0.91
    ANL ✓ ✓ 58.67 +0.99
    FullMatch ✓ ✓ ✓ 59.32 +1.64

    Key observations:

    • Implementing EML with binary cross-entropy (BCE) slightly outperforms cross-entropy (CE) (+0.79% vs +0.67%).
    • Applying ANL solely to samples without positive pseudo-labels ("w/o PL") yields a +0.91% gain, whereas applying it only to samples with positive pseudo-labels ("w PL") yields only +0.15%.
    • Applying ANL to all unlabeled data yields a +0.99% gain.
    • Combining BCE-based EML and full ANL in FullMatch achieves the highest overall accuracy (59.32%, a +1.64% improvement over FixMatch).
  8. Knowl 8 — Hyperparameter Robustness to Loss Weights alpha and beta

    data/table

    Evaluating FullMatch with different loss weight configurations for α\alpha (weight of ANL loss) and β\beta (weight of EML loss) on CIFAR-100 with 10,000 labeled samples shows that model accuracy is insensitive to changes in these hyperparameters:

    0.5 1.0 2.0
    0.5 1.0 2.0 0.5 1.0 2.0 0.5 1.0 2.0
    Accuracy (%) 78.43 78.36 78.50 78.38 78.46 78.48 78.49 78.31 78.47

    Across all combinations of α,β∈{0.5,1.0,2.0}\alpha, \beta \in \{0.5, 1.0, 2.0\}, accuracy remains virtually constant within a narrow band between 78.31% and 78.50%, justifying setting α=1.0\alpha = 1.0 and β=1.0\beta = 1.0 without dataset-specific tuning.

Coverage note — No substantial contributed material was omitted; qualitative visualizations (t-SNE and entropy distribution figures) were subsumed into the corresponding method and empirical knowls.

References

  1. 1.Philip Bachman, Ouais Alsharif, and Doina Precup. Learning with pseudo-ensembles. In NeurIPS, 2014.
  2. 2.David Berthelot, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation Anchoring. In ICLR, 2020.
  3. 3.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. MixMatch: a holistic approach to semi-supervised learning. In NeurIPS, 2019.
  4. 4.John Chen, Vatsal Shah, and Anastasios Kyrillidis. Negative sampling in semi-supervised learning. In ICML, 2020.
  5. 5.Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In AISTATS, 2011.
  6. 6.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In CVPR Workshops, 2020.
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  8. 8.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
  9. 9.Lee Doyup, Sungwoong Kim, Ildoo Kim, Yeongjae Cheon, Minsu Cho, and Wook-Shin Han. Contrastive Regularization for Semi-Supervised Learning. In CVPR, 2022.
  10. 10.Zhengyang Feng, Qianyu Zhou, Qiqi Gu, Xin Tan, Guangliang Cheng, Xuequan Lu, Jianping Shi, and Lizhuang Ma. Dmt: Dynamic mutual training for semi-supervised learning. PR, 2022.
  11. 11.Chengyue Gong, Dilin Wang, and Qiang Liu. Alphamatch: Improving consistency for semi-supervised learning with alpha-divergence. In CVPR, 2021.
  12. 12.Yves Grandvalet and Yoshua Bengio. Semi-supervised learning by entropy minimization. In NeurIPS, 2005.
  13. 13.Ethan Harris, Antonia Marcu, Matthew Painter, Mahesan Niranjan, Adam Prugel-Bennett, and Jonathon Hare. Fmix: Enhancing mixed sample data augmentation. arXiv preprint arXiv:2002.12047, 2020.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  15. 15.Byoungjip Kim, Jinho Choo, Yeong-Dae Kwon, Seongho Joe, Seungjai Min, and Youngjune Gwon. Selfmatch: Combining contrastive self-supervision and consistency for semi-supervised learning. arXiv preprint arXiv:2101.06480, 2021.
  16. 16.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Citeseer, 2009.
  17. 17.Dong-Hyun Lee. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In ICML Workshops, 2013.
  18. 18.Junnan Li, Caiming Xiong, and Steven CH Hoi. Comatch: Semi-supervised learning with contrastive graph regularization. In ICCV, 2021.
  19. 19.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  20. 20.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE TPAMI, 2018.
  21. 21.Islam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray Buntine, and Gholamreza Haffari. All labels are not created equal: Enhancing semi-supervision via label grouping and co-training. In CVPR, 2021.
  22. 22.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. In NeurIPS Workshop, 2011.
  23. 23.Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko. Semi-supervised learning with ladder networks. In NeurIPS, 2015.
  24. 24.Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah. In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning. In ICLR, 2021.
  25. 25.Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. Mutual exclusivity loss for semi-supervised deep learning. In ICIP, 2016.
  26. 26.Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence. In NeurIPS, 2020.
  27. 27.Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. On the importance of initialization and momentum in deep learning. In ICML, 2013.
  28. 28.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-SNE. JMLR, 2008.
  29. 29.Erik Wallin, Lennart Svensson, Fredrik Kahl, and Lars Hammarstrand. DoubleMatch: Improving Semi-Supervised Learning with Self-Supervision. arXiv preprint arXiv:2205.05575, 2022.
  30. 30.Jianfeng Wang, Thomas Lukasiewicz, Daniela Massiceti, Xiaolin Hu, Vladimir Pavlovic, and Alexandros Neophytou. NP-Match: When Neural Processes meet Semi-Supervised Learning. In ICML, 2022.
  31. 31.Yidong Wang, Hao Chen, Qiang Heng, Wenxin Hou, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Zhen Wu, and Jindong Wang. FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning. arXiv preprint arXiv:2205.07246, 2022.
  32. 32.Yuchao Wang, Haochen Wang, Yujun Shen, Jingjing Fei, Wei Li, Guoqiang Jin, Liwei Wu, Rui Zhao, and Xinyi Le. Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels. In CVPR, 2022.
  33. 33.Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. Unsupervised data augmentation for consistency training. In NeurIPS, 2020.
  34. 34.Yi Xu, Lei Shang, Jinxing Ye, Qi Qian, Yu-Feng Li, Baigui Sun, Hao Li, and Rong Jin. Dash: Semi-supervised learning with dynamic thresholding. In ICML, 2021.
  35. 35.Lihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi, and Yang Gao. St++: Make self-training work better for semi-supervised semantic segmentation. In CVPR, 2022.
  36. 36.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV, 2019.
  37. 37.Sergey Zagoruyko and Nikos Komodakis. Wide Residual Networks. In BMVC, 2016.
  38. 38.Bowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu, Jindong Wang, Manabu Okumura, and Takahiro Shinozaki. FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling. In NeurIPS, 2021.
  39. 39.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  40. 40.Wenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang, Andrew Makmur, Qingpeng Cai, and Beng Chin Ooi. Boostmis: Boosting medical image semi-supervised learning with adaptive pseudo labeling and informative active annotation. In CVPR, 2022.
  41. 41.Mingkai Zheng, Shan You, Lang Huang, Fei Wang, Chen Qian, and Chang Xu. SimMatch: Semi-supervised Learning with Similarity Matching. In CVPR, 2022.
  42. 42.Hongyu Zhou, Zheng Ge, Songtao Liu, Weixin Mao, Zeming Li, Haiyan Yu, and Jian Sun. Dense Teacher: Dense Pseudo-Labels for Semi-supervised Object Detection. In ECCV, 2022.
  43. 43.Tianyi Zhou, Shengjie Wang, and Jeff Bilmes. Time-consistent self-supervision for semi-supervised learning. In ICML, 2020.
  44. 44.Xiaojin Zhu and Andrew B Goldberg. Introduction to semi-supervised learning. Synthesis lectures on artificial intelligence and machine learning, 2009.

Citation

MLA
Chen, Y., et al. “Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data”. arXiv, 2023, http://arxiv.org/abs/2303.11066v1.
APA
Chen, Y., Tan, X., Zhao, B., Chen, Z., Song, R., Liang, J., & Lu, X. (2023). Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data. arXiv. http://arxiv.org/abs/2303.11066v1
Chicago
Chen, Y., X. Tan, B. Zhao, et al. 2023. “Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data”. arXiv. http://arxiv.org/abs/2303.11066v1.
Harvard
Chen, Y. et al. (2023) “Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.11066v1.
Vancouver
1. Chen Y, Tan X, Zhao B, Chen Z, Song R, Liang J, Lu X (2023) Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data. arXiv

BibTeX

@article{chen2023boosting,
  title = {Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data},
  author = {Chen, Yuhao and Tan, Xin and Zhao, Borui and Chen, Zhaowei and Song, Renjie and Liang, Jiajun and Lu, Xuequan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.11066v1},
  eprint = {2303.11066}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE