ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image Classification
Fengbei LiuYu TianYuanhong ChenYuyuan LiuVasileios BelagiannisGustavo Carneiro
Proposes an anti-curriculum pseudo-labelling framework that prioritizes informative unlabeled samples and ensembles neural network predictions with nearest-neighbor classifiers to outperform state-of-the-art semi-supervised methods on class-imbalanced multi-label and multi-class medical diagnosis tasks.
Medical image analysis using deep learning often struggles with severe data constraints: while massive collections of unlabelled medical scans exist, obtaining high-quality expert annotations is costly, time-consuming, and scarce. Furthermore, clinical diagnostic tasks frequently involve multi-class or multi-label conditions alongside extreme class imbalances, where normal cases dominate and critical diseases represent only a small fraction of the data. Existing semi-supervised learning techniques that assign artificial labels to unlabelled samples typically select only high-confidence predictions. This practice reinforces majority-class bias, risks compounding classification errors, and struggles to generalise across both single- and multi-disease diagnostic settings.
The article introduces and evaluates anti-curriculum pseudo-labelling (ACPL), a semi-supervised learning method designed to handle imbalanced multi-class and multi-label medical image classification. The core objective is to demonstrate that selecting highly informative, low-density unlabelled samples—contrasting with standard easy-first curriculum learning—mitigates class imbalance and, when paired with an ensemble pseudo-labelling mechanism, improves overall diagnostic accuracy without relying on complex, task-specific data perturbations or computationally expensive self-supervised pre-training.
The approach was evaluated on two widely recognised public benchmarks: the multi-label Chest X-Ray14 dataset containing 112,120 chest radiographs across 14 disease classes, and the multi-class ISIC2018 dataset comprising 10,015 skin lesion images across seven conditions. The authors implemented a standard DenseNet-121 architecture and systematically compared ACPL against established consistency-based, pseudo-labelling, and self-supervised benchmarks across varying proportions of labelled training data ranging from 2% to 20%.
The findings confirm clear performance advantages across both diagnostic tasks. On the Chest X-Ray14 benchmark, ACPL achieved the top area under the ROC curve (AUC) across all evaluated labelled data splits, reaching an AUC of 74.82% with only 2% labelled data and 81.77% with 20% labelled data. This performance outperformed prior pseudo-labelling techniques by 3% to 20% and exceeded top consistency-based and self-supervised models. Class-level analysis on chest radiographs demonstrated superior accuracy in 10 of the 14 disease categories. On the ISIC2018 skin lesion dataset with 20% labelled data, the method set new top benchmarks with an AUC of 94.36%, sensitivity of 72.14%, and an F1 score of 62.23%, improving AUC by up to 3% over consistency methods and markedly outperforming standard self-training. Detailed component testing showed that prioritizing high-information samples raised the representation of rare minority diseases from under 10% to nearly 30% during training, directly resolving class imbalance while maintaining low variance across runs.
These results have significant operational and financial implications for deploying artificial intelligence in clinical environments. By achieving superior classification accuracy from minimal annotated data, healthcare organisations can cut data curation costs and accelerate model deployment timelines. Because ACPL relies on a standard pre-trained foundation rather than complex self-supervised pre-training, it lowers computational overhead while mitigating confirmation bias through an ensemble classifier combining deep network outputs with nearest-neighbour predictions.
Based on these findings, technical teams should consider adopting informative sample selection and ensemble-guided pseudo-labelling when training diagnostic models on imbalanced medical datasets. For organisations with limited labelling budgets, applying this method allows rapid bootstrapping of diagnostic classifiers using standard off-the-shelf architectures. Further validation should include piloting the algorithm on broader clinical imaging modalities and general computer vision tasks.
A primary limitation of this work is its assumption that unlabelled data originates entirely from within the same distribution as the labelled set. The authors note that model behaviour in the presence of out-of-distribution or corrupted clinical inputs remains untested. While confidence in the benchmarked performance is high, practitioners should exercise caution and conduct local validation before applying the method to operational clinical workflows containing out-of-distribution anomalies.
- Paper: FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling, Bowen Zhang et al. (2021). Introduces curriculum-based pseudo-labeling and adaptive thresholding for semi-supervised learning, providing the baseline curriculum paradigm that ACPL explicitly contrasts and inverts.
- Paper: Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss, Kaidi Cao et al. (2019). Establishes margin adjustments and optimization principles for handling severe class imbalance in deep networks, underpinning the class imbalance motivation in medical semi-supervised tasks.
- Paper: Deep Bayesian Active Learning with Image Data, Yarin Gal et al. (2017). Demonstrates how uncertainty estimation and informativeness criteria guide sample selection on medical image datasets, which motivates ACPL's informative sample selection.
- Paper: ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases, Xiaosong Wang et al. (2017). Presents the large-scale ChestX-ray dataset and multi-label benchmark environment that serves as one of the primary evaluation settings for ACPL.
- Paper: Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC), Noel Codella et al. (2019). Details the multi-class skin lesion benchmark (ISIC 2018) and evaluation metrics used directly to test ACPL under class-imbalanced conditions.
- Paper: A systematic study of the class imbalance problem in convolutional neural networks, Mateusz Buda et al. (2017). Provides a comprehensive study of the harmful effects of class imbalance on convolutional neural networks and evaluates classical mitigation strategies.
- Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). Presents key self-supervised and semi-supervised pre-training and distillation baselines against which ACPL benchmarks its compute-efficient pseudo-labeling.
- Paper: Debiased Learning from Naturally Imbalanced Pseudo-Labels, Xudong Wang et al. (2022). Extends the study of confirmation bias in self-training by introducing dynamic debiasing techniques to eliminate class-distribution distortion caused by pseudo-labels.
- Paper: LaSSL: Label-Guided Self-Training for Semi-supervised Learning, Zhen Zhao et al. (2022). Explores alternative pseudo-label refinement by coupling feature-space contrastive alignment with buffer-aided label propagation in label-scarce semi-supervised regimes.
- Paper: Uncertainty Estimation by Fisher Information-based Evidential Deep Learning, Danruo Deng et al. (2023). Generalizes evidential uncertainty and sample informativeness quantification using Fisher information to prevent overconfidence in multi-class learning.
