Class-Imbalanced Semi-Supervised Learning with Adaptive Thresholding
Lan-Zhe GuoYufeng Li
Proposes a class-dependent adaptive thresholding method with an efficient closed-form solution to improve pseudo-label selection and classification accuracy on minority classes in imbalanced semi-supervised learning.
Practical machine learning applications frequently face severe class imbalance alongside high data-labeling costs, leading to heavy reliance on semi-supervised learning techniques that use large volumes of unlabeled data. However, existing semi-supervised methods usually assume balanced categories and apply a single, fixed confidence threshold to assign pseudo-labels to unlabeled samples. This fixed-threshold approach severely biases model predictions toward majority categories, disproportionately discarding useful minority-class samples and causing substantial performance degradation on rare classes in real-world deployments.
The article evaluates and demonstrates a novel framework called Adaptive Thresholding (Adsh), which adjusts pseudo-label selection thresholds individually for each class during model training. The main objective is to establish an efficient, distribution-aware learning method that simultaneously minimizes prediction error and balances pseudo-label generation across imbalanced classes, without requiring prior knowledge of unlabeled data distributions.
To demonstrate this method, the authors developed a mathematical formulation that integrates class-specific selection biases into the learning objective, yielding an exact, efficient solution for setting adaptive thresholds. The framework was comprehensively evaluated using a standard deep neural network architecture across long-tailed variants of standard image benchmarks (CIFAR-10, SVHN, and STL-10). The experiments encompassed over twenty distinct imbalance ratios, varying volumes of labeled samples, and challenging scenarios where the class distributions of labeled and unlabeled sets differed significantly.
The primary findings show that Adsh consistently outperforms existing state-of-the-art semi-supervised and imbalanced-learning techniques across all evaluated settings. In heavily imbalanced scenarios on CIFAR-10, Adsh improved classification accuracy by up to 4 to 5 percentage points over leading baselines like FixMatch, DARP, and CReST. On SVHN and STL-10 datasets, Adsh achieved top accuracies of 92.13% and 79.25% respectively, demonstrating strong adaptability even when unlabeled class distributions were unknown. Furthermore, error analyses revealed that Adsh produced substantially less biased confusion matrices, and the framework readily combined with downstream re-balancing techniques to boost overall accuracy up to 86.21%.
These findings indicate that adapting confidence thresholds by class significantly enhances the robustness and performance of automated prediction systems in skewed environments, reducing operational risks associated with minority-class misclassification. Practitioners can implement Adsh with minimal computational overhead, as it operates via closed-form updates and avoids complex hyperparameter tuning. Organizations should consider Adsh as a drop-in enhancement for semi-supervised pipelines facing rare-event or class-skew challenges, starting with the recommended baseline configuration and fine-tuning as needed.
While empirical results strongly support the framework's practical efficacy across diverse vision datasets, the authors note that formal theoretical convergence guarantees remain an area for future research. Stakeholders should maintain high confidence in Adsh's empirical gains while piloting the approach across domain-specific data types to validate generalization beyond standard vision benchmarks.
- Paper: FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence, Kihyuk Sohn et al. (2020). FixMatch establishes the confidence-thresholded pseudo-labeling objective that this paper adapts to class imbalance.
- Paper: FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling, Bowen Zhang et al. (2021). FlexMatch develops class-specific adaptive confidence thresholds, providing the key prior SSL mechanism that this paper redirects toward balanced pseudo-label selection.
- Paper: A survey on semi-supervised learning, Jesper E. van Engelen et al. (2019). This survey organizes pseudo-labeling and other SSL approaches, helping situate the paper’s threshold-based method within the field’s established frameworks.
- Paper: Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data, Yuhao Chen et al. (2023). FullMatch extends confidence-filtered SSL by extracting supervisory signals from low-confidence examples as well as selected pseudo-labels, broadening the thresholding problem this paper addresses.
