LaSSL: Label-Guided Self-Training for Semi-supervised Learning

Zhen ZhaoLuping ZhouLei WangYinghuan ShiYang Gao

article2022AAAI53 citations

Proposes a semi-supervised learning framework that iteratively refines pseudo-label quality and improves low-label image classification by coupling class-aware contrastive learning with mini-batch label propagation.

Listen

Building high-performing deep learning models typically requires massive volumes of manually labeled data, which is expensive, time-consuming, and often impractical in specialized domains like medicine. Semi-supervised learning offers a solution by training models on small sets of labeled data alongside large volumes of unlabeled data. However, prevailing methods rely on generating estimated labels, known as pseudo-labels, and filtering out low-confidence predictions. This practice discards useful unlabeled samples, underutilizes known label relationships, and risks compounding errors during model training.

The article develops and evaluates LaSSL, a label-guided self-training framework designed to maximize the utility of limited labeled data. The framework introduces two mutually reinforcing mechanisms: a class-aware contrastive loss that groups similar instances in feature space regardless of individual image variations, and a buffer-aided label propagation algorithm that spreads reliable label information across neighboring samples in small, efficient batches using a temporary memory buffer.

The approach was evaluated across four standard image classification benchmarks: CIFAR-10, CIFAR-100, SVHN, and Mini-ImageNet, under varying degrees of label scarcity. Key findings demonstrate that LaSSL outperforms existing state-of-the-art semi-supervised methods, particularly in extremely label-scarce environments. On CIFAR-10 with only 40 labeled samples (four per class), LaSSL reached 95.07% accuracy, matching performance levels that competing methods achieve only with 250 or more labels. On CIFAR-100 with four labels per class, it attained 62.33% accuracy, outperforming the leading baseline by approximately 7 percentage points. On the more complex Mini-ImageNet benchmark with 4,000 labeled images, LaSSL achieved 60.14% accuracy, exceeding the baseline by an absolute margin of 10.75 percentage points. Component analyses confirmed that contrastive loss rapidly scales the volume of confident predictions, while label propagation directly improves prediction accuracy.

These results indicate that organizations can significantly reduce data annotation costs and shorten deployment timelines without sacrificing predictive accuracy. By improving how models learn relationships across both labeled and unlabeled examples, high-accuracy vision systems become feasible in data-restricted operational environments.

Teams deploying machine learning under tight data-labeling budgets should consider incorporating iteration-level label propagation and class-aware feature grouping. When adopting this method, practitioners must tune key operational parameters, specifically the sample similarity thresholds and buffer sampling iterations, to balance computational overhead against label noise. Future work should focus on validating the framework on full-scale industry datasets, extending it beyond image classification to tasks such as object detection, and exploring automated hyperparameter tuning.

Cover for LaSSL: Label-Guided Self-Training for Semi-supervised Learning

Abstract

The key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches.

Table of Contents

  • Introduction
  • Related Work
  • Method
  • Overview
  • Buffer-aided Label Propagation Algorithm
  • Class-aware Contrastive Loss
  • Putting it all together
  • Experiments
  • CIFAR-10, CIFAR-100, and SVHN
  • Mini-ImageNet
  • Ablation Study
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — LaSSL Framework for Semi-Supervised Learning

    model/method

    Label-guided Self-training for Semi-supervised Learning (LaSSL) combines consistency-based self-training with representation learning and feature-level label propagation. Let X={(xb,pb)}b=1BX = \{(x_b, p_b)\}_{b=1}^B denote a mini-batch of BB labeled samples with ground-truth one-hot class distributions pb∈RNp_b \in \mathbb{R}^N over NN classes, and U={ub}b=1μBU = \{u_b\}_{b=1}^{\mu B} denote a mini-batch of μB\mu B unlabeled samples, where μ\mu is the unlabeled-to-labeled batch size ratio.

    The model architecture consists of:

    1. An encoder h(⋅)h(\cdot) extracting base representations from input images.
    2. A classification head f(⋅)f(\cdot) predicting class probability distributions F(x)=f(h(x))∈RNF(x) = f(h(x)) \in \mathbb{R}^N.
    3. A projection head g(⋅)g(\cdot) mapping representations to a low-dimensional embedding space G(x)=g(h(x))∈RdG(x) = g(h(x)) \in \mathbb{R}^d.

    Each training iteration comprises two phases:

    • Inference Phase: The network evaluates weakly augmented inputs a(ub)a(u_b) and a(xb)a(x_b) to produce projection features obu=G(a(ub))o_b^u = G(a(u_b)), obx=G(a(xb))o_b^x = G(a(x_b)), and initial predictions qbu=F(a(ub))q_b^u = F(a(u_b)). A distribution alignment (DA) step smooths predictions into qˉbu\bar{q}_b^u. Concurrently, a buffer-aided label propagation algorithm (BLPA) leverages historical embeddings and predictions from a queue QQ to generate label-propagated predictions q~bu\tilde{q}_b^u. The final pseudo-label is given by: q^bu=ηq~bu+(1−η)qˉbu\hat{q}_b^u = \eta \tilde{q}_b^u + (1 - \eta) \bar{q}_b^u where η∈[0,1]\eta \in [0, 1] is a balancing coefficient.

    • Training Phase: The network evaluates strongly augmented unlabeled images A(ub)A(u_b) and weakly augmented labeled images a(xb)a(x_b), producing predictions ybu=F(A(ub))y_b^u = F(A(u_b)), ybx=F(a(xb))y_b^x = F(a(x_b)), and projection embeddings zbu=G(A(ub))z_b^u = G(A(u_b)), zbx=G(a(xb))z_b^x = G(a(x_b)). The model is updated by minimizing a joint objective combining supervised cross-entropy, thresholded pseudo-label cross-entropy, and a class-aware contrastive loss (CACL). Model parameters are tracked using an exponential moving average (EMA) with a decay of 0.999.

  2. Knowl 2 — Buffer-aided Label Propagation Algorithm (BLPA)

    model/method

    Buffer-aided Label Propagation Algorithm (BLPA) performs label propagation at the mini-batch iteration level on the feature embedding space, utilizing the output queue Qi−1={(ob−1,qb−1)}Q_{i-1} = \{(o_{b-1}, q_{b-1})\} from the immediately preceding iteration i−1i-1.

    To mitigate noise from incorrect pseudo-labels, BLPA applies bootstrap aggregation (bagging) by drawing KK independent samples with replacement from Qi−1Q_{i-1}, yielding sampled sets {(ob−1(k),qb−1(k))}k=1K\{(o_{b-1}(k), q_{b-1}(k))\}_{k=1}^K. For each sample k∈{1,…,K}k \in \{1, \dots, K\}, the buffered data is partitioned based on a confidence threshold τ\tau:

    • High-confidence historical samples (treated as labeled): qb−1high(k)=1(max⁡(qb−1(k))≥τ)qb−1(k)q_{b-1}^{\text{high}}(k) = \mathbf{1}(\max(q_{b-1}(k)) \ge \tau) q_{b-1}(k) ob−1high(k)=1(max⁡(qb−1(k))≥τ)ob−1(k)o_{b-1}^{\text{high}}(k) = \mathbf{1}(\max(q_{b-1}(k)) \ge \tau) o_{b-1}(k)
    • Low-confidence historical samples (treated as unlabeled to explore manifold structure): ob−1low(k)=1(max⁡(qb−1(k))<τ)ob−1(k)o_{b-1}^{\text{low}}(k) = \mathbf{1}(\max(q_{b-1}(k)) < \tau) o_{b-1}(k)

    These subsets are merged with the current mini-batch labeled and unlabeled outputs (obx=G(a(xb))o_b^x = G(a(x_b)), obu=G(a(ub))o_b^u = G(a(u_b)), pbp_b) into compound labeled features os(k)=[obx,ob−1high(k)]o_s(k) = [o_b^x, o_{b-1}^{\text{high}}(k)], compound unlabeled features ot(k)=[obu,ob−1low(k)]o_t(k) = [o_b^u, o_{b-1}^{\text{low}}(k)], and compound label vectors qs(k)=[pb,qb−1high(k)]q_s(k) = [p_b, q_{b-1}^{\text{high}}(k)].

    An affinity matrix Ω(k)\Omega(k) is formed from pairwise feature cosine similarities between all compound samples (with zero diagonal). Its symmetrically normalized version is: Ω~(k)=D−1/2Ω(k)D−1/2\tilde{\Omega}(k) = D^{-1/2} \Omega(k) D^{-1/2} where DD is the diagonal degree matrix of Ω(k)\Omega(k) (Dii=∑jΩij(k)D_{ii} = \sum_j \Omega_{ij}(k)). Label propagation is solved via its closed-form solution: Φ∗(k)=(I−αΩ~(k))−1qs(k)\Phi^*(k) = (I - \alpha \tilde{\Omega}(k))^{-1} q_s(k) where α∈(0,1)\alpha \in (0, 1) governs the propagation fraction. Extracting the first μB\mu B entries corresponding to the current unlabeled batch gives ϕb(k)=Φ∗(k)[:μB]\phi_b(k) = \Phi^*(k)[:\mu B]. The ensemble prediction over all KK samplings is computed as: q~bu=1K∑k=1Kϕb(k)\tilde{q}_b^u = \frac{1}{K} \sum_{k=1}^K \phi_b(k)

  3. Knowl 3 — Class-Aware Contrastive Loss (CACL)

    model/method

    The Class-Aware Contrastive Loss (CACL) guides feature representation learning by leveraging both ground-truth labels and generated pseudo-labels across labeled and unlabeled samples in a mini-batch.

    Let y^=[pb,q^bu]∈R(B+μB)×N\hat{y} = [p_b, \hat{q}_b^u] \in \mathbb{R}^{(B + \mu B) \times N} denote the concatenated label distributions (one-hot ground truths pbp_b for labeled inputs and refined pseudo-labels q^bu\hat{q}_b^u for unlabeled inputs), and let z=[zbx,zbu]z = [z_b^x, z_b^u] denote the corresponding feature projection embeddings, where zbx=G(a(xb))z_b^x = G(a(x_b)) and zbu=G(A(ub))z_b^u = G(A(u_b)).

    The pairwise semantic similarity ωi,j\omega_{i,j} between instance ii and instance jj is computed as: ωi,j={1,if i=j0,if i≠j and y^i⋅y^j<εy^i⋅y^j,if i≠j and y^i⋅y^j≥ε\omega_{i,j} = \begin{cases} 1, & \text{if } i = j \\ 0, & \text{if } i \neq j \text{ and } \hat{y}_i \cdot \hat{y}_j < \varepsilon \\ \hat{y}_i \cdot \hat{y}_j, & \text{if } i \neq j \text{ and } \hat{y}_i \cdot \hat{y}_j \ge \varepsilon \end{cases} where ε∈[0,1]\varepsilon \in [0, 1] is a predefined similarity threshold, and y^i⋅y^j\hat{y}_i \cdot \hat{y}_j is the dot product between probability vectors.

    The class-aware contrastive loss is formulated as: Lbc=−∑i=1∣y^∣log⁡∑j=1∣y^∣ωi,jexp⁡(zi⋅zj/T)∑j=1,j≠i∣y^∣exp⁡(zi⋅zj/T)\mathcal{L}_b^c = - \sum_{i=1}^{|\hat{y}|} \log \frac{\sum_{j=1}^{|\hat{y}|} \omega_{i,j} \exp(z_i \cdot z_j / T)}{\sum_{j=1, j \neq i}^{|\hat{y}|} \exp(z_i \cdot z_j / T)} where T>0T > 0 is a temperature hyperparameter and zi⋅zjz_i \cdot z_j denotes the cosine similarity between normalized embedding vectors.

  4. Knowl 4 — Total Training Loss and Dynamic Ramp-Down Schedule

    equation

    In LaSSL, the total loss Lb\mathcal{L}_b optimized at each mini-batch iteration is: Lb=Lbx+λuLbu+λcLbc\mathcal{L}_b = \mathcal{L}_b^x + \lambda_u \mathcal{L}_b^u + \lambda_c \mathcal{L}_b^c

    The individual loss components are:

    1. Supervised cross-entropy loss: Lbx=H(pb,ybx)\mathcal{L}_b^x = H(p_b, y_b^x) where H(p,q)=−∑cp(c)log⁡q(c)H(p, q) = -\sum_c p(c) \log q(c), pbp_b are ground-truth labels, and ybx=F(a(xb))y_b^x = F(a(x_b)).
    2. Unsupervised consistency cross-entropy loss: Lbu=1(max⁡(q^bu)≥τ)H(q^bu,ybu)\mathcal{L}_b^u = \mathbf{1}(\max(\hat{q}_b^u) \ge \tau) H(\hat{q}_b^u, y_b^u) where 1(⋅)\mathbf{1}(\cdot) masks out predictions below confidence threshold τ\tau, q^bu\hat{q}_b^u is the refined pseudo-label, and ybu=F(A(ub))y_b^u = F(A(u_b)). The weight is set to λu=1.0\lambda_u = 1.0.
    3. Class-aware contrastive loss Lbc\mathcal{L}_b^c with a dynamic time-variant weight λc\lambda_c.

    To prioritize representation learning early and downstream classification tasks later, λc\lambda_c follows an exponential ramp-down schedule across training epochs t∈[1,Tt]t \in [1, T_t]: λc={λc0,if t≤Trλc0exp⁡(−(t−Tr)22(Tt−Tr)),otherwise\lambda_c = \begin{cases} \lambda_c^0, & \text{if } t \le T_r \\ \lambda_c^0 \exp\left(-\frac{(t - T_r)^2}{2(T_t - T_r)}\right), & \text{otherwise} \end{cases} where λc0\lambda_c^0 is the initial maximum weight, TtT_t is the total number of training epochs, and TrT_r is the ramp-down start epoch. When λc≤λ^c\lambda_c \le \hat{\lambda}_c (a minimum threshold), CACL and BLPA computations are deactivated for the remainder of training.

  5. Knowl 5 — LaSSL Mini-Batch Training Algorithm

    algorithm

    Algorithm for one training iteration of LaSSL:

    Input: Labeled batch (xb,pbx_b, p_b), unlabeled batch ubu_b, weight λc\lambda_c, threshold τ\tau, similarity threshold ε\varepsilon, combination ratio η\eta, bagging rounds KK, unsupervised weight λu\lambda_u.
    Output: Updated encoder hh, classifier ff, projector gg.
    // Phase I: Inference Phase
    Compute initial predictions on weakly augmented unlabeled data: qbu=f(h(a(ub)))q_b^u = f(h(a(u_b)))
    Apply distribution alignment with exponential moving average decay 0.99: qˉbu=DA(qbu)\bar{q}_b^u = \text{DA}(q_b^u)
    Compute projection embeddings: obx=g(h(a(xb)))o_b^x = g(h(a(x_b))) and obu=g(h(a(ub)))o_b^u = g(h(a(u_b)))
    Compute propagated pseudo-labels q~bu\tilde{q}_b^u via BLPA using obx,obu,pbo_b^x, o_b^u, p_b, and buffered queue Qi−1Q_{i-1}
    Compute final pseudo-labels: q^bu=ηq~bu+(1−η)qˉbu\hat{q}_b^u = \eta \tilde{q}_b^u + (1 - \eta) \bar{q}_b^u
    Update FIFO buffer queue: Qi={([obu,obx],[q^bu,pb])}Q_i = \{([o_b^u, o_b^x], [\hat{q}_b^u, p_b])\}
    // Phase II: Training Phase
    Compute predictions and projections on strongly augmented unlabeled data: ybu=f(h(A(ub)))y_b^u = f(h(A(u_b))), zbu=g(h(A(ub)))z_b^u = g(h(A(u_b)))
    Compute predictions and projections on weakly augmented labeled data: ybx=f(h(a(xb)))y_b^x = f(h(a(x_b))), zbx=g(h(a(xb)))z_b^x = g(h(a(x_b)))
    Compute supervised loss: Lbx=H(pb,ybx)\mathcal{L}_b^x = H(p_b, y_b^x)
    Compute unsupervised loss: Lbu=1(max⁡(q^bu)≥τ)H(q^bu,ybu)\mathcal{L}_b^u = \mathbf{1}(\max(\hat{q}_b^u) \ge \tau) H(\hat{q}_b^u, y_b^u)
    Compute class-aware contrastive loss Lbc\mathcal{L}_b^c using z=[zbx,zbu]z = [z_b^x, z_b^u] and y^=[pb,q^bu]\hat{y} = [p_b, \hat{q}_b^u] with threshold ε\varepsilon
    Compute total loss: Lb=Lbx+λuLbu+λcLbc\mathcal{L}_b = \mathcal{L}_b^x + \lambda_u \mathcal{L}_b^u + \lambda_c \mathcal{L}_b^c
    Back-propagate Lb\mathcal{L}_b and update parameters of h,f,gh, f, g
    Update exponential moving average (EMA) model parameters with decay 0.999
  6. Knowl 6 — Experimental Setup and Hyperparameters for LaSSL

    experimental setup

    LaSSL is evaluated on image classification benchmarks using standard semi-supervised evaluation protocols:

    • Architectures: Wide ResNet-28-2 (WRN-28-2) as encoder h(⋅)h(\cdot) for CIFAR-10 and SVHN; Wide ResNet-28-8 (WRN-28-8) for CIFAR-100; ResNet-18 for Mini-ImageNet. The predictor f(⋅)f(\cdot) is a single linear classification layer, and the projector g(⋅)g(\cdot) is a 2-layer MLP.
    • Optimization: SGD with momentum 0.90.9, weight decay 5×10−45 \times 10^{-4}, cosine decay learning rate scheduler, and exponential moving average (EMA) decay of 0.9990.999 for evaluation models.
    • Default Hyperparameters:
      • Labeled batch size B=64B = 64
      • Unlabeled-to-labeled batch size ratio μ=7\mu = 7
      • Number of bagging sampling rounds K=7K = 7
      • Label propagation factor α=0.8\alpha = 0.8
      • Prediction blending weight η=0.2\eta = 0.2
      • Pseudo-label confidence threshold τ=0.95\tau = 0.95 (set to 0.80.8 for Mini-ImageNet)
      • Semantic similarity threshold ε=0.7\varepsilon = 0.7
      • Total epochs Tt=512T_t = 512
      • Initial contrastive loss weight λc0=1.0\lambda_c^0 = 1.0 (set to 5.05.0 for Mini-ImageNet)
      • Contrastive termination threshold λ^c=0.1\hat{\lambda}_c = 0.1
      • Distribution alignment moving average decay =0.99= 0.99
  7. Knowl 7 — Classification Performance on CIFAR-10, CIFAR-100, and SVHN

    data/table

    LaSSL was benchmarked against state-of-the-art semi-supervised learning algorithms on CIFAR-10, CIFAR-100, and SVHN across 5 different random splits. Top-1 test accuracy (mean ±\pm standard deviation in %) is reported below:

    Methods CIFAR-10 CIFAR-100 SVHN
    40 labels 250 labels 400 labels 2500 labels 40 labels 250 labels
    Pseudo-label - 50.22 ±\pm 0.43 - 42.62 ±\pm 0.46 - 79.79 ±\pm 1.09
    Mean-Teacher - 67.68 ±\pm 2.30 - 46.09 ±\pm 0.57 - 96.43 ±\pm 0.11
    MixMatch 52.46 ±\pm 11.50 88.95 ±\pm 0.86 33.39 ±\pm 1.32 60.06 ±\pm 0.37 57.45 ±\pm 14.53 96.02 ±\pm 0.23
    UDA 70.95 ±\pm 5.93 91.18 ±\pm 1.08 40.72 ±\pm 0.88 66.87 ±\pm 0.22 47.37 ±\pm 20.51 94.31 ±\pm 2.76
    ReMixMatch 80.90 ±\pm 9.64 94.56 ±\pm 0.05 55.72 ±\pm 2.06 72.57 ±\pm 0.31 96.64 ±\pm 0.30 97.08 ±\pm 0.48
    FixMatch 86.19 ±\pm 3.37 94.93 ±\pm 0.65 51.15 ±\pm 1.75 71.71 ±\pm 0.11 96.04 ±\pm 2.17 97.52 ±\pm 0.38
    ACR 92.38 95.01 - - - -
    SelfMatch 93.19 ±\pm 1.08 95.13 ±\pm 0.26 - - 96.58 ±\pm 1.02 97.37 ±\pm 0.43
    CoMatch 93.09 ±\pm 1.39 95.09 ±\pm 0.33 - - - -
    Dash 86.78 ±\pm 3.75 95.44 ±\pm 0.13 55.24 ±\pm 0.96 72.82 ±\pm 0.21 96.97 ±\pm 1.59 97.83 ±\pm 0.10
    LaSSL 95.07 ±\pm 0.78 95.71 ±\pm 0.46 62.33 ±\pm 2.69 74.67 ±\pm 0.65 96.91 ±\pm 0.52 97.85 ±\pm 0.13

    LaSSL achieves significant improvements in extremely scarce label settings, reaching 95.07% on CIFAR-10 with 40 labels (4 labels per class), matching or exceeding competing methods trained with 250 labels. On CIFAR-100 with 400 labels, LaSSL achieves a 7.09% accuracy gain over the strongest baseline (Dash).

  8. Knowl 8 — Classification Performance on Mini-ImageNet

    empirical result

    On the Mini-ImageNet dataset (100 classes, 50,000 training images, 10,000 testing images center-cropped and resized to 84×8484 \times 84) using a ResNet-18 backbone with 4,000 labeled samples (40 labels per class):

    • The prior state-of-the-art method SimPLE achieved an average top-1 test accuracy of 49.39%49.39\%.
    • LaSSL achieved an average top-1 test accuracy of 60.14%±0.26%60.14\% \pm 0.26\%.

    This represents an absolute performance gain of 10.75%10.75\% over the previous state-of-the-art method under identical benchmark conditions.

  9. Knowl 9 — Ablation Study on CACL, BLPA, and Distribution Alignment Components

    data/table

    An ablation study on CIFAR-10 with 40 labeled samples (evaluated after 100 training epochs with fixed random seed 1) investigates the individual and combined effects of Class-Aware Contrastive Loss (CACL), Buffer-aided Label Propagation Algorithm (BLPA), and Distribution Alignment (DA).

    Metrics evaluated include:

    • Quantity (%): Ratio of high-confidence predictions (max⁡≥τ\\\max \ge \tau) to total unlabeled samples.
    • Quality (%): Ratio of high-confidence predictions that match ground-truth labels.
    • Accuracy (%): Top-1 classification accuracy.
    Method CACL BLPA DA Quant (%) Qual (%) Acc (%)
    Vanilla ×\times ×\times ×\times 83.91 81.98 75.54
    LaSSL-v1 ✓\checkmark ×\times ×\times 88.66 89.38 85.50
    LaSSL-v2 ✓\checkmark ✓\checkmark ×\times 89.08 94.31 90.24
    LaSSL-v3 ×\times ×\times ✓\checkmark 85.73 94.90 90.42
    LaSSL-v4 ✓\checkmark ×\times ✓\checkmark 87.46 94.89 91.11
    LaSSL-v5 (Full) ✓\checkmark ✓\checkmark ✓\checkmark 87.03 95.33 91.65

    Key observations:

    1. CACL rapidly boosts the quantity of high-confidence pseudo-labels (LaSSL-v1 achieves 88.66% quantity vs 83.91% for Vanilla).
    2. BLPA and DA substantially boost pseudo-label quality (increasing from 89.38% in v1 to 94.31% in v2, reaching 95.33% in the full model).
    3. Pseudo-label quality directly correlates with final testing accuracy.
  10. Knowl 10 — Sensitivity of Similarity Threshold and Bagging Sampling Rounds

    empirical result

    Ablation experiments on CIFAR-10 with 40 labeled samples evaluate the effect of the similarity threshold ε\varepsilon in CACL and the number of bagging sampling rounds KK in BLPA:

    1. Similarity Threshold ε\varepsilon in CACL (evaluated without BLPA and DA):

      • ε=0.6\varepsilon = 0.6: 87.64%87.64\% accuracy
      • ε=0.7\varepsilon = 0.7: 89.39%89.39\% accuracy (optimal)
      • ε=0.8\varepsilon = 0.8: 87.70%87.70\% accuracy
      • ε=0.9\varepsilon = 0.9: 87.36%87.36\% accuracy
      • ε=1.0\varepsilon = 1.0: 85.17%85.17\% accuracy (equivalent to standard instance discrimination where every sample is treated as a distinct class) The setting ε=0.7\varepsilon = 0.7 achieves the best balance: lower thresholds introduce false-positive pair associations, whereas higher thresholds fail to capture inter-instance semantic relationships.
    2. Bagging Sampling Rounds KK in BLPA (evaluated with default CACL and DA settings):

      • K=0K = 0 (no buffer-aided queue): 92.71%92.71\% accuracy
      • K=1K = 1 (direct buffer usage without bootstrap sampling): 92.10%92.10\% accuracy
      • K=3K = 3: 94.64%94.64\% accuracy
      • K=5K = 5: 93.43%93.43\% accuracy
      • K=7K = 7: 94.87%94.87\% accuracy Directly using all buffered data without bagging (K=1K=1) degrades accuracy due to unmitigated noise from false pseudo-labels. Bagging (K>1K > 1) filters noise and stabilizes label propagation.

Coverage note — None was omitted; all key methodology components (framework, BLPA, CACL, loss weighting schedule, iteration algorithm), benchmark results across four datasets, and ablation studies from the paper are fully covered.

References

  1. 1.Abuduweili, A.; Li, X.; Shi, H.; Xu, C.-Z.; and Dou, D. 2021. Adaptive Consistency Regularization for Semi-Supervised Transfer Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6923–6932.
  2. 2.Arazo, E.; Ortego, D.; Albert, P.; O’Connor, N. E.; and McGuinness, K. 2020. Pseudo-labeling and confirmation bias in deep semi-supervised learning. In 2020 International Joint Conference on Neural Networks (IJCNN), 1–8. IEEE.
  3. 3.Bengio, Y.; Louradour, J.; Collobert, R.; and Weston, J. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning (ICML), 41–48.
  4. 4.Berthelot, D.; Carlini, N.; Cubuk, E. D.; Kurakin, A.; Sohn, K.; Zhang, H.; and Raffel, C. 2020. Remixmatch: Semi-supervised learning with distribution alignment and augmentation anchoring. In 8th International Conference on Learning Representations (ICLR).
  5. 5.Berthelot, D.; Carlini, N.; Goodfellow, I.; Papernot, N.; Oliver, A.; and Raffel, C. A. 2019. MixMatch: A Holistic Approach to Semi-Supervised Learning. In Advances in Neural Information Processing Systems.
  6. 6.Blum, A.; and Mitchell, T. 1998. Combining labeled and unlabeled data with co-training. In Proceedings of the 17th annual conference on computational learning theory, 92–100.
  7. 7.Bridle, J. S.; Heading, A. J.; and MacKay, D. J. 1992. Unsupervised Classifiers, Mutual Information and 'Phantom Targets'. In Advances in Neural Information Processing Systems, NIPS Conference, Denver, Colorado, USA, December 2-5, 1991.
  8. 8.Cascante-Bonilla, P.; Tan, F.; Qi, Y.; and Ordonez, V. 2020. Curriculum Labeling: Revisiting Pseudo-Labeling for Semi-Supervised Learning. arXiv preprint arXiv:2001.06001.
  9. 9.Chapelle, O.; Scholkopf, B.; and Zien, A. 2009. Semi-supervised learning [book reviews]. IEEE Transactions on Neural Networks, 20(3): 542–542.
  10. 10.Chen, D.; Wang, W.; Gao, W.; and Zhou, Z. 2018. Tri-net for semi-supervised deep learning. In International Joint Conferences on Artificial Intelligence.
  11. 11.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020a. A simple framework for contrastive learning of visual representations. In International conference on machine learning (ICML), 1597–1607.
  12. 12.Chen, T.; Kornblith, S.; Swersky, K.; Norouzi, M.; and Hinton, G. 2020b. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029.
  13. 13.Goodfellow, I.; Bengio, Y.; Courville, A.; and Bengio, Y. 2016. Deep learning. MIT press Cambridge.
  14. 14.Grandvalet, Y.; and Bengio, Y. 2005. Semi-supervised learning by entropy minimization. In CAP, 281–296.
  15. 15.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9729–9738.
  16. 16.Hu, Z.; Yang, Z.; Hu, X.; and Nevatia, R. 2021. SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised Classification. arXiv preprint arXiv:2103.16725.
  17. 17.Iscen, A.; Tolias, G.; Avrithis, Y.; and Chum, O. 2019a. Label propagation for deep semi-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5070–5079.
  18. 18.Iscen, A.; Tolias, G.; Avrithis, Y.; and Chum, O. 2019b. Label propagation for deep semi-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5070–5079.
  19. 19.Jaiswal, A.; Babu, A. R.; Zadeh, M. Z.; Banerjee, D.; and Makedon, F. 2021. A survey on contrastive self-supervised learning. Technologies, 9(1): 2.
  20. 20.Kim, B.; Choo, J.; Kwon, Y.-D.; Joe, S.; Min, S.; and Gwon, Y. 2021. SelfMatch: Combining Contrastive Self-Supervision and Consistency for Semi-Supervised Learning. arXiv preprint arXiv:2101.06480.
  21. 21.Krizhevsky, A.; and Hinton, G. 2009. Learning multiple layers of features from tiny images. Handbook of Systemic Autoimmune Diseases, 1(4).
  22. 22.Laine, S.; and Aila, T. 2016. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242.
  23. 23.Lee, D.-H.; et al. 2013. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML.
  24. 24.Li, J.; Xiong, C.; and Hoi, S. 2020. CoMatch: Semi-supervised Learning with Contrastive Graph Regularization. arXiv preprint arXiv:2011.11183.
  25. 25.McLachlan, G. J. 1975. Iterative reclassification procedure for constructing an asymptotically optimal rule of allocation in discriminant analysis. Journal of the American Statistical Association, 70(350): 365–369.
  26. 26.Miyato, T.; Maeda, S.-i.; Koyama, M.; and Ishii, S. 2018. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 41(8): 1979–1993.
  27. 27.Netzer, Y.; and Wang, T. 2011. Reading Digits in Natural Images with Unsupervised Feature Learning. nips workshop on deep learning and unsupervised feature learning.
  28. 28.Oliver, A.; Odena, A.; Raffel, C.; Cubuk, E. D.; and Goodfellow, I. J. 2018. Realistic evaluation of deep semi-supervised learning algorithms. arXiv preprint arXiv:1804.09170.
  29. 29.Ouali, Y.; Hudelot, C.; and Tami, M. 2020. An Overview of Deep Semi-Supervised Learning. arXiv preprint arXiv:2006.05278.
  30. 30.Qiao, S.; Shen, W.; Zhang, Z.; Wang, B.; and Yuille, A. 2018. Deep co-training for semi-supervised image recognition. In Proceedings of the european conference on computer vision, 135–152.
  31. 31.Rasmus, A.; Valpola, H.; Honkala, M.; Berglund, M.; and Raiko, T. 2015. Semi-supervised learning with ladder networks. arXiv preprint arXiv:1507.02672.
  32. 32.Ravi, S.; and Larochelle, H. 2017. Optimization as a model for few-shot learning. In 5th International Conference on Learning Representations (ICLR).
  33. 33.Rizve, M. N.; Duarte, K.; Rawat, Y. S.; and Shah, M. 2021. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning. arXiv preprint arXiv:2101.06329.
  34. 34.Sohn, K.; Berthelot, D.; Li, C.-L.; Zhang, Z.; Carlini, N.; Cubuk, E. D.; Kurakin, A.; Zhang, H.; and Raffel, C. 2020. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685.
  35. 35.Tarvainen, A.; and Valpola, H. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780.
  36. 36.Verma, V.; Kawaguchi, K.; Lamb, A.; Kannala, J.; Bengio, Y.; and Lopez-Paz, D. 2019. Interpolation consistency training for semi-supervised learning. arXiv preprint arXiv:1903.03825.
  37. 37.Xie, Q.; Dai, Z.; Hovy, E.; Luong, M.-T.; and Le, Q. V. 2019. Unsupervised data augmentation for consistency training. arXiv preprint arXiv:1904.12848.
  38. 38.Xu, Y.; Shang, L.; Ye, J.; Qian, Q.; Li, Y.-F.; Sun, B.; Li, H.; and Jin, R. 2021. Dash: Semi-Supervised Learning with Dynamic Thresholding. In International Conference on Machine Learning (ICML), 11525–11536.
  39. 39.Yalniz, I. Z.; Jégou, H.; Chen, K.; Paluri, M.; and Mahajan, D. 2019. Billion-scale semi-supervised learning for image classification. arXiv preprint arXiv:1905.00546.
  40. 40.Zhai, X.; Oliver, A.; Kolesnikov, A.; and Beyer, L. 2019. S4l: Self-supervised semi-supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1476–1485.
  41. 41.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412.

Citation

MLA
Zhao, Z., et al. “LaSSL: Label-Guided Self-Training for Semi-supervised Learning”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 8, 2022, pp. 9208–16, https://doi.org/10.1609/AAAI.V36I8.20907.
APA
Zhao, Z., Zhou, L., Wang, L., Shi, Y., & Gao, Y. (2022). LaSSL: Label-Guided Self-Training for Semi-supervised Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 36(8), 9208–9216. https://doi.org/10.1609/AAAI.V36I8.20907
Chicago
Zhao, Z., L. Zhou, L. Wang, Y. Shi, and Y. Gao. 2022. “LaSSL: Label-Guided Self-Training for Semi-supervised Learning”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (8): 9208–16. https://doi.org/10.1609/AAAI.V36I8.20907.
Harvard
Zhao, Z. et al. (2022) “LaSSL: Label-Guided Self-Training for Semi-supervised Learning”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(8), pp. 9208–9216. Available at: https://doi.org/10.1609/AAAI.V36I8.20907.
Vancouver
1. Zhao Z, Zhou L, Wang L, Shi Y, Gao Y (2022) LaSSL: Label-Guided Self-Training for Semi-supervised Learning. Proceedings of the AAAI Conference on Artificial Intelligence 36:9208–9216

BibTeX

@article{Zhao_2022, title={LaSSL: Label-Guided Self-Training for Semi-supervised Learning}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I8.20907}, DOI={10.1609/aaai.v36i8.20907}, number={8}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhao, Zhen and Zhou, Luping and Wang, Lei and Shi, Yinghuan and Gao, Yang}, year={2022}, month=June, pages={9208–9216} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF