Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective
Yunrui ZhaoQianqian XuYangbangyan JiangPeisong WenQingming Huang
Proposes Dist-PU, a positive-unlabeled learning framework that aligns predicted and ground-truth label distributions while utilizing entropy minimization and Mixup regularization to eliminate negative-prediction bias in deep classifiers.
In many practical machine learning settings, such as medical diagnostics, recommendation engines, and cybersecurity, organizations only possess verified positive examples alongside a vast pool of unlabeled data. Traditional approaches to this problem, known as Positive-Unlabeled learning, typically minimize the penalty of misclassifying unlabeled data as negative cases. However, this common practice frequently forces flexible deep learning models to overfit and develop an unintended preference for negative predictions, degrading real-world performance.
The article introduces and evaluates Dist-PU, a framework designed to train accurate binary classifiers without negative labels by matching the overall predicted label distributions to their expected proportions. Rather than penalizing negative misclassifications directly, the article demonstrates how aligning distribution averages corrects model bias and produces superior classification accuracy.
The researchers developed an end-to-end framework combining three primary components: label distribution alignment, an entropy minimization technique to keep predictions sharp and decisive, and data-mixing regularization to prevent the model from becoming overly confident in early erroneous predictions. The method was evaluated across three distinct image classification benchmark datasets, ranging from standard vision benchmarks to medical imaging for Alzheimer's disease recognition, and benchmarked against nine established state-of-the-art algorithms using six evaluation metrics.
The findings confirm that Dist-PU consistently outperforms existing methods across all benchmarks. Overall, the approach improved performance metrics by roughly 1% to 3% over the second-best alternatives, achieving top scores such as 95.40% accuracy on Fashion-MNIST, 91.88% on CIFAR-10, and 71.57% on the Alzheimer dataset. Ablation tests revealed that label distribution alignment is the primary driver of performance, while data-mixing regularization significantly reduced the volume and overconfidence of incorrect predictions, resulting in a balanced trade-off between precision and recall.
These results indicate that organizations facing high labeling costs can achieve superior automated classification without exhaustively annotating negative data. By rectifying systemic negative-prediction bias, the proposed framework reduces operational labeling overhead while lowering the risk of false negatives in high-stakes operational environments. Practitioners can implement this objective directly into existing deep learning training pipelines with standard gradient optimization.
Organizations developing models under incomplete labeling conditions should consider deploying this distribution alignment framework in pilot settings, particularly where approximate positive class proportions can be estimated. Future initiatives should evaluate the framework's application across broader weakly supervised domains, such as text and tabular data, and explore automated tuning for regularization parameters to streamline deployment.
Confidence in these findings is supported by rigorous statistical significance testing across multiple benchmark tasks. Readers should note that the framework assumes the general proportion of positive cases in the unlabeled data is known or can be estimated accurately, which remains a key operational requirement for optimal performance.
- Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning, David Berthelot et al. (2019). Introduces the integration of Mixup and entropy minimization on unlabeled data in semi-supervised learning that Dist-PU directly adopts to regularize expectation alignment.
- Paper: The foundations of cost-sensitive learning, Charles Elkan (2001). Establishes the foundations of cost-sensitive risk minimization that traditional PU learning relies on and that Dist-PU identifies as prone to negative-prediction bias.
- Paper: A survey on semi-supervised learning, Jesper E. van Engelen et al. (2019). Provides a comprehensive taxonomy and theoretical grounding for learning from unlabeled data, contextualizing the transition from standard semi-supervised risk minimization to PU learning.
- Paper: Learning with Local and Global Consistency, Dengyong Zhou et al. (2003). Introduces the principle of global and local consistency regularization on unlabeled samples that motivates distribution-level alignment constraints.
- Paper: Debiased Learning from Naturally Imbalanced Pseudo-Labels, Xudong Wang et al. (2022). Extends the study of confirmation bias and distribution imbalances on unlabeled pseudo-labels by applying counterfactual reasoning to eliminate label bias.
- Paper: Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data, Yuhao Chen et al. (2023). Builds on regularizing unlabeled prediction distributions and mitigating confirmation bias by introducing entropy meaning losses and adaptive negative learning.
- Paper: Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness, Francesco Pinto et al. (2022). Investigates the specific regularizing mechanism of Mixup alongside standard objectives, directly continuing the analysis of Mixup regularization applied in Dist-PU.
- Paper: Generalized Category Discovery with Decoupled Prototypical Network, Wenbin An et al. (2023). Generalizes the problem of learning from partially labeled and unlabeled pools to open-world settings containing entirely novel, unannotated categories.
