Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective

Yunrui ZhaoQianqian XuYangbangyan JiangPeisong WenQingming Huang

article2022CVPR72 citations

Proposes Dist-PU, a positive-unlabeled learning framework that aligns predicted and ground-truth label distributions while utilizing entropy minimization and Mixup regularization to eliminate negative-prediction bias in deep classifiers.

Listen

In many practical machine learning settings, such as medical diagnostics, recommendation engines, and cybersecurity, organizations only possess verified positive examples alongside a vast pool of unlabeled data. Traditional approaches to this problem, known as Positive-Unlabeled learning, typically minimize the penalty of misclassifying unlabeled data as negative cases. However, this common practice frequently forces flexible deep learning models to overfit and develop an unintended preference for negative predictions, degrading real-world performance.

The article introduces and evaluates Dist-PU, a framework designed to train accurate binary classifiers without negative labels by matching the overall predicted label distributions to their expected proportions. Rather than penalizing negative misclassifications directly, the article demonstrates how aligning distribution averages corrects model bias and produces superior classification accuracy.

The researchers developed an end-to-end framework combining three primary components: label distribution alignment, an entropy minimization technique to keep predictions sharp and decisive, and data-mixing regularization to prevent the model from becoming overly confident in early erroneous predictions. The method was evaluated across three distinct image classification benchmark datasets, ranging from standard vision benchmarks to medical imaging for Alzheimer's disease recognition, and benchmarked against nine established state-of-the-art algorithms using six evaluation metrics.

The findings confirm that Dist-PU consistently outperforms existing methods across all benchmarks. Overall, the approach improved performance metrics by roughly 1% to 3% over the second-best alternatives, achieving top scores such as 95.40% accuracy on Fashion-MNIST, 91.88% on CIFAR-10, and 71.57% on the Alzheimer dataset. Ablation tests revealed that label distribution alignment is the primary driver of performance, while data-mixing regularization significantly reduced the volume and overconfidence of incorrect predictions, resulting in a balanced trade-off between precision and recall.

These results indicate that organizations facing high labeling costs can achieve superior automated classification without exhaustively annotating negative data. By rectifying systemic negative-prediction bias, the proposed framework reduces operational labeling overhead while lowering the risk of false negatives in high-stakes operational environments. Practitioners can implement this objective directly into existing deep learning training pipelines with standard gradient optimization.

Organizations developing models under incomplete labeling conditions should consider deploying this distribution alignment framework in pilot settings, particularly where approximate positive class proportions can be estimated. Future initiatives should evaluate the framework's application across broader weakly supervised domains, such as text and tabular data, and explore automated tuning for regularization parameters to streamline deployment.

Confidence in these findings is supported by rigorous statistical significance testing across multiple benchmark tasks. Readers should note that the framework assumes the general proportion of positive cases in the unlabeled data is known or can be estimated accurately, which remains a key operational requirement for optimal performance.

arXiv: 2212.02801
  • Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning, David Berthelot et al. (2019). Introduces the integration of Mixup and entropy minimization on unlabeled data in semi-supervised learning that Dist-PU directly adopts to regularize expectation alignment.
  • Paper: The foundations of cost-sensitive learning, Charles Elkan (2001). Establishes the foundations of cost-sensitive risk minimization that traditional PU learning relies on and that Dist-PU identifies as prone to negative-prediction bias.
  • Paper: A survey on semi-supervised learning, Jesper E. van Engelen et al. (2019). Provides a comprehensive taxonomy and theoretical grounding for learning from unlabeled data, contextualizing the transition from standard semi-supervised risk minimization to PU learning.
  • Paper: Learning with Local and Global Consistency, Dengyong Zhou et al. (2003). Introduces the principle of global and local consistency regularization on unlabeled samples that motivates distribution-level alignment constraints.
Cover for Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective

Abstract

Positive-Unlabeled (PU) learning tries to learn binary classifiers from a few labeled positive examples with many unlabeled ones. Compared with ordinary semi-supervised learning, this task is much more challenging due to the absence of any known negative labels. While existing cost-sensitive-based methods have achieved state-of-the-art performances, they explicitly minimize the risk of classifying unlabeled data as negative samples, which might result in a negative-prediction preference of the classifier. To alleviate this issue, we resort to a label distribution perspective for PU learning in this paper. Noticing that the label distribution of unlabeled data is fixed when the class prior is known, it can be naturally used as learning supervision for the model. Motivated by this, we propose to pursue the label distribution consistency between predicted and ground-truth label distributions, which is formulated by aligning their expectations. Moreover, we further adopt the entropy minimization and Mixup regularization to avoid the trivial solution of the label distribution consistency on unlabeled data and mitigate the consequent confirmation bias. Experiments on three benchmark datasets validate the effectiveness of the proposed method.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Problem setting
  • 3.2. Label distribution alignment
  • 3.3. Entropy minimization
  • 3.4. Confirmation bias
  • 4. Experiment
  • 4.1. Experimental settings
  • 4.2. Comparision with state-of-the-art methods
  • 4.3. Ablation studies
  • 4.4. Effectiveness of hyper-parameters
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Label Distribution Alignment Principle for Positive-Unlabeled Learning

    model/method

    In binary classification with input space X⊆Rd\mathcal{X} \subseteq \mathbb{R}^d and label space Y={0,1}\mathcal{Y} = \{0, 1\}, let πP=Pr⁡(y=1)\pi_P = \Pr(y=1) denote the positive class prior and πN=1−πP\pi_N = 1 - \pi_P the negative class prior. In positive-unlabeled (PU) learning under the Selected Completely at Random (SCAR) assumption, the training data consists of a labeled positive set XL={xi}i=1nL∼pP(x)X_L = \{x_i\}_{i=1}^{n_L} \sim p_P(x) and an unlabeled set XU={xj}j=1nU∼p(x)X_U = \{x_j\}_{j=1}^{n_U} \sim p(x), where p(x)=πPpP(x)+πNpN(x)p(x) = \pi_P p_P(x) + \pi_N p_N(x) is the marginal input distribution.

    Traditional cost-sensitive PU approaches optimize classification risk by treating unlabeled samples as negative instances with reweighting, which causes deep neural networks to overfit and produce a negative-prediction preference (where the predicted positive rate drops well below πP\pi_P). Label distribution alignment mitigates this by aligning the expected value of predicted labels with ground-truth label expectations.

    For a classifier f:X→Rf: \mathcal{X} \to \mathbb{R} with predicted label y^∈{0,1}\hat{y} \in \{0, 1\}, the expected zero-one risk decomposes into positive and negative risks:

    R=πP∣Ex∼pP(x)[y^]−1∣⏟RP+πN∣Ex∼pN(x)[y^]∣⏟RNR = \pi_P \underbrace{\left| \mathbb{E}_{x \sim p_P(x)}[\hat{y}] - 1 \right|}_{R_P} + \pi_N \underbrace{\left| \mathbb{E}_{x \sim p_N(x)}[\hat{y}] \right|}_{R_N}

    Using the identity Ex∼pU(x)[y^]−E(x,y)∼p(x,y)[y]=−πPRP+πNRN\mathbb{E}_{x \sim p_U(x)}[\hat{y}] - \mathbb{E}_{(x,y)\sim p(x,y)}[y] = -\pi_P R_P + \pi_N R_N, the negative risk term is expressed without negative labels as πNRN=πPRP+Ex∼pU(x)[y^]−πP\pi_N R_N = \pi_P R_P + \mathbb{E}_{x \sim p_U(x)}[\hat{y}] - \pi_P. To enforce that the expected prediction on unlabeled data matches the true class prior πP\pi_P stably, the distribution alignment risk incorporates an absolute deviation:

    Rlab=2πPRP+∣Ex∼pU(x)[y^]−πP∣R_{lab} = 2\pi_P R_P + \left| \mathbb{E}_{x \sim p_U(x)}[\hat{y}] - \pi_P \right|

  2. Knowl 2 — Differentiable Empirical Label Distribution Alignment Loss

    equation

    Let f:X→Rf: \mathcal{X} \to \mathbb{R} be a neural network parameterized classifier producing a logit for input xx. To make the expectation of predicted labels differentiable for gradient-based optimization, the hard expectation Ex[y^]\mathbb{E}_x[\hat{y}] is replaced with the expectation of the continuous sigmoid probability score:

    s=σ(f(x))=11+exp⁡(−f(x))s = \sigma(f(x)) = \frac{1}{1 + \exp(-f(x))}

    Given nLn_L labeled positive samples XLX_L and nUn_U unlabeled samples XUX_U, along with the known positive class prior πP∈(0,1)\pi_P \in (0, 1), the empirical label distribution alignment loss R^lab\hat{R}_{lab} is defined as:

    R^lab=2πPR^L+R^U=2πP∣1nL∑x∈XLs−1∣+∣1nU∑x∈XUs−πP∣\hat{R}_{lab} = 2\pi_P \hat{R}_L + \hat{R}_U = 2\pi_P \left| \frac{1}{n_L} \sum_{x \in X_L} s - 1 \right| + \left| \frac{1}{n_U} \sum_{x \in X_U} s - \pi_P \right|

    where R^L=∣1nL∑x∈XLs−1∣\hat{R}_L = \left| \frac{1}{n_L} \sum_{x \in X_L} s - 1 \right| measures the L1L_1-distance between the mean predicted score on labeled positives and 11, and R^U=∣1nU∑x∈XUs−πP∣\hat{R}_U = \left| \frac{1}{n_U} \sum_{x \in X_U} s - \pi_P \right| measures the L1L_1-distance between the mean predicted score on unlabeled instances and πP\pi_P.

  3. Knowl 3 — Generalization Bound for Label Distribution Alignment

    theoretical result

    Let F\mathcal{F} be a hypothesis class of bb-uniformly bounded functions with Vapnik–Chervonenkis (VC) dimension V\mathcal{V}. Let nLn_L be the number of labeled positive examples drawn identically and independently from pP(x)p_P(x), and nUn_U be the number of unlabeled examples drawn independently from the marginal distribution p(x)p(x). Let πP\pi_P denote the positive class prior probability and R^lab\hat{R}_{lab} the empirical label distribution alignment risk.

    For any δ∈(0,1)\delta \in (0, 1), with probability at least 1−δ1 - \delta, the true classification risk RR satisfies:

    R≤2R^lab+8πP⋅CVnL+12πPln⁡(4/δ)2nL+4CVnU+6ln⁡(4/δ)2nUR \le 2 \hat{R}_{lab} + 8\pi_P \cdot C \sqrt{\frac{\mathcal{V}}{n_L}} + 12\pi_P \sqrt{\frac{\ln(4/\delta)}{2 n_L}} + 4 C \sqrt{\frac{\mathcal{V}}{n_U}} + 6 \sqrt{\frac{\ln(4/\delta)}{2 n_U}}

    where CC is a universal constant. This bound establishes that minimizing the empirical label distribution alignment surrogate R^lab\hat{R}_{lab} minimizes an upper bound on the true zero-one classification risk RR, with the generalization gap vanishing as nL,nU→∞n_L, n_U \to \infty.

  4. Knowl 4 — Entropy Minimization for Unlabeled Data in Dist-PU

    model/method

    Optimizing the empirical unlabeled distribution loss R^U=∣1nU∑x∈XUs(x)−πP∣\hat{R}_U = \left| \frac{1}{n_U} \sum_{x \in X_U} s(x) - \pi_P \right| in isolation admits a degenerate trivial solution where the network outputs a constant score s(x)=πPs(x) = \pi_P for every unlabeled sample x∈XUx \in X_U, which makes R^U=0\hat{R}_U = 0 while yielding an uninformative classifier.

    To enforce the cluster assumption (that decision boundaries lie in low-density regions) and drive scores away from intermediate values toward 0 or 1, Dist-PU adds an entropy minimization loss over the unlabeled set XUX_U:

    Lent=−1nU∑x∈XU[(1−s)log⁡(1−s)+slog⁡s]L_{ent} = -\frac{1}{n_U} \sum_{x \in X_U} \left[ (1 - s) \log(1 - s) + s \log s \right]

    where s=σ(f(x))=1/(1+exp⁡(−f(x)))s = \sigma(f(x)) = 1 / (1 + \exp(-f(x))) is the predicted positive probability for unlabeled sample xx.

  5. Knowl 5 — Mixup Regularization and Mixed Entropy Loss for Mitigating Confirmation Bias

    model/method

    Entropy minimization can cause deep classifiers to become overconfident and reinforce early incorrect predictions on unlabeled data (confirmation bias). To mitigate this, Dist-PU incorporates Mixup regularization using predicted scores as soft label targets over all training samples XPU=XL∪XUX_{PU} = X_L \cup X_U.

    For any pair of samples x1,x2∈XPUx_1, x_2 \in X_{PU} with soft predictions s1=σ(f(x1))s_1 = \sigma(f(x_1)) and s2=σ(f(x2))s_2 = \sigma(f(x_2)), a mixing ratio is drawn from a symmetric Beta distribution:

    λ∼Beta(α,α),λ′=max⁡(λ,1−λ)\lambda \sim \text{Beta}(\alpha, \alpha), \quad \lambda' = \max(\lambda, 1 - \lambda)

    x′=λ′x1+(1−λ′)x2x' = \lambda' x_1 + (1 - \lambda') x_2

    Let s′=σ(f(x′))s' = \sigma(f(x')) and let lbce(t′,t)=−(1−t)log⁡(1−t′)−tlog⁡(t′)l_{bce}(t', t) = -(1-t)\log(1-t') - t\log(t') be the binary cross-entropy loss. The Mixup loss LmixL_{mix} over n=∣XPU∣n = |X_{PU}| pairs is:

    Lmix=1n∑x1,x2∈XPU[λ′lbce(s′,s1)+(1−λ′)lbce(s′,s2)]L_{mix} = \frac{1}{n} \sum_{x_1, x_2 \in X_{PU}} \left[ \lambda' l_{bce}(s', s_1) + (1 - \lambda') l_{bce}(s', s_2) \right]

    To ensure low entropy on virtual mixed samples as well, an entropy minimization penalty is also imposed on the mixed predictions:

    Lent′=−1n∑x1,x2∈XPUlbce(s′,s′)=−1n∑x1,x2∈XPU[(1−s′)log⁡(1−s′)+s′log⁡s′]L_{ent}' = -\frac{1}{n} \sum_{x_1, x_2 \in X_{PU}} l_{bce}(s', s') = -\frac{1}{n} \sum_{x_1, x_2 \in X_{PU}} \left[ (1 - s') \log(1 - s') + s' \log s' \right]

  6. Knowl 6 — Overall Dist-PU Training Objective and Optimization Scheme

    model/method

    The complete training objective for Dist-PU integrates label distribution alignment, unlabeled entropy minimization, Mixup consistency, and mixed sample entropy minimization:

    Ldist=R^lab+μLent+νLmix+γLent′L_{dist} = \hat{R}_{lab} + \mu L_{ent} + \nu L_{mix} + \gamma L_{ent}'

    where μ≥0\mu \ge 0 adjusts the importance of unlabeled entropy minimization LentL_{ent}, ν≥0\nu \ge 0 controls the strength of the Mixup consistency loss LmixL_{mix}, and γ≥0\gamma \ge 0 controls the entropy penalty Lent′L_{ent}' on mixed instances.

    Optimization is performed end-to-end via gradient descent using Adam with an initial learning rate of 5×10−45 \times 10^{-4} and weight decay of 5×10−35 \times 10^{-3}. Network logits are clamped to [−10,10][-10, 10] to prevent numerical errors in the entropy and log terms. Training follows a two-phase schedule: a warm-up stage of several epochs where ν=γ=0\nu = \gamma = 0 (without Mixup), followed by 60 epochs with Mixup enabled, where both the learning rate and the weight μ\mu are scheduled using cosine annealing.

  7. Knowl 7 — Experimental Setup and Benchmark Datasets for Dist-PU

    experimental setup

    Dist-PU was evaluated on three binary PU learning benchmarks with the following configurations:

    1. Fashion-MNIST (F-MNIST): 28 ×\times 28 grayscale images. nL=500n_L = 500 labeled positives, nU=60,000n_U = 60{,}000 unlabeled instances, 10,00010{,}000 test instances. Positive class prior πP=0.4\pi_P = 0.4. Positive classes are Top items (classes 0, 2, 4, 6); remaining classes are negative. Backbone is a 6-layer MLP; batch size is 256.

    2. CIFAR-10: 3 ×\times 32 ×\times 32 RGB images normalized with mean (0.485,0.456,0.406)(0.485, 0.456, 0.406) and standard deviation (0.229,0.224,0.225)(0.229, 0.224, 0.225). nL=1,000n_L = 1{,}000, nU=50,000n_U = 50{,}000, 10,00010{,}000 test instances. πP=0.4\pi_P = 0.4. Positive classes are Vehicles (classes 0, 1, 8, 9). Backbone is a 13-layer CNN; batch size is 256.

    3. Alzheimer Dataset: 3 ×\times 224 ×\times 224 MRI images. nL=769n_L = 769, nU=5,121n_U = 5{,}121, 1,2791{,}279 test instances. πP=0.5\pi_P = 0.5. Positive class corresponds to Alzheimer's Disease. Backbone is ResNet-50; batch size is 128.

    Hyperparameters μ\mu, ν\nu, γ\gamma, and α\alpha are searched within the ranges [0,0.1][0, 0.1], [0,10][0, 10], [0,0.3][0, 0.3], and [0.1,10][0.1, 10], respectively. All experiments are evaluated across five independent runs reporting mean and standard deviation for Accuracy (ACC), Precision (Prec.), Recall (Rec.), F1 score, Area Under ROC Curve (AUC), and Average Precision (AP).

  8. Knowl 8 — Comparative Performance on F-MNIST, CIFAR-10, and Alzheimer Benchmarks

    data/table

    Across all three benchmark datasets, Dist-PU outperformed baseline PU learning methods across accuracy, F1, AUC, and AP, maintaining balanced precision and recall.

    Dataset Method ACC (%) Prec. (%) Rec. (%) F1 (%) AUC (%) AP (%)
    F-MNIST naive 91.07 (0.93) 90.16 (2.25) 87.28 (2.08) 88.66 (1.14) 96.94 (0.67) 94.27 (1.40)
    uPU 94.02 (0.30) 92.50 (1.26) 92.59 (0.80) 92.53 (0.31) 97.34 (0.54) 96.60 (0.52)
    nnPU 94.44 (0.49) 91.69 (1.13) 94.69 (0.84) 93.16 (0.57) 97.53 (0.48) 96.39 (0.98)
    RP 92.37 (1.08) 88.58 (1.56) 92.94 (2.38) 90.69 (1.39) 97.14 (0.58) 94.39 (1.31)
    PUSB 94.50 (0.36) 93.12 (0.44) 93.12 (0.44) 93.12 (0.44) 97.31 (0.50) 96.28 (0.96)
    PUbN 94.82 (0.16) 92.92 (0.50) 94.24 (0.93) 93.57 (0.24) 94.72 (0.28) 89.87 (0.18)
    Self-PU 94.75 (0.25) 91.73 (0.80) 95.50 (0.61) 93.57 (0.28) 97.62 (0.31) 96.14 (0.70)
    aPU 94.71 (0.34) 92.71 (0.50) 94.20 (1.06) 93.44 (0.45) 97.67 (0.40) 96.64 (0.48)
    VPU 92.26 (1.11) 89.04 (2.00) 92.01 (2.00) 90.48 (1.35) 97.38 (0.44) 95.57 (0.62)
    ImbPU 94.54 (0.42) 92.81 (1.53) 93.66 (1.67) 93.21 (0.52) 97.67 (0.81) 96.73 (0.86)
    Dist-PU 95.40 (0.34) 94.18 (0.90) 94.34 (1.00) 94.25 (0.43) 98.57 (0.24) 97.90 (0.30)
    CIFAR-10 naive 84.92 (0.89) 83.59 (0.89) 77.57 (3.23) 80.43 (1.56) 92.29 (0.63) 88.23 (0.79)
    uPU 88.35 (0.45) 87.18 (2.39) 83.23 (2.68) 85.10 (0.56) 94.91 (0.62) 92.62 (1.11)
    nnPU 88.89 (0.45) 86.18 (1.15) 86.05 (1.42) 86.10 (0.58) 95.12 (0.52) 92.42 (1.38)
    RP 88.73 (0.15) 86.01 (1.01) 85.82 (1.51) 85.90 (0.32) 95.17 (0.23) 92.92 (0.56)
    PUSB 88.95 (0.41) 86.19 (0.51) 86.19 (0.51) 86.19 (0.51) 95.13 (0.52) 92.44 (1.34)
    PUbN 89.83 (0.30) 87.85 (0.98) 86.56 (1.87) 87.18 (0.54) 89.28 (0.54) 81.41 (0.37)
    Self-PU 89.28 (0.72) 86.16 (0.78) 87.21 (2.35) 86.67 (1.06) 95.47 (0.58) 93.28 (1.01)
    aPU 89.05 (0.52) 86.29 (1.30) 86.37 (0.79) 86.32 (0.56) 95.09 (0.42) 92.41 (1.23)
    VPU 87.99 (0.48) 86.72 (1.41) 82.71 (2.84) 84.63 (0.91) 94.51 (0.41) 92.00 (0.73)
    ImbPU 89.41 (0.46) 86.69 (0.87) 86.87 (0.82) 86.77 (0.56) 95.52 (0.27) 93.45 (0.45)
    Dist-PU 91.88 (0.52) 89.87 (1.09) 89.84 (0.81) 89.85 (0.62) 96.92 (0.45) 95.49 (0.72)
    Alzheimer naive 61.45 (3.81) 62.51 (5.87) 61.44 (12.5) 61.05 (4.65) 66.26 (6.39) 63.28 (5.98)
    uPU 68.48 (2.15) 69.65 (3.50) 66.13 (6.13) 67.62 (2.78) 73.75 (2.94) 69.53 (3.23)
    nnPU 68.33 (2.13) 68.01 (2.33) 69.48 (7.15) 68.55 (3.16) 72.90 (2.80) 69.45 (2.87)
    RP 61.61 (3.20) 61.89 (4.54) 64.60 (15.89) 62.10 (5.61) 66.13 (3.28) 63.82 (2.29)
    PUSB 69.21 (2.39) 69.16 (2.39) 69.26 (2.39) 69.21 (2.39) 74.43 (2.41) 70.00 (1.56)
    PUbN 69.98 (1.34) 69.42 (2.49) 72.02 (8.42) 70.38 (3.19) 69.98 (1.34) 63.84 (0.95)
    Self-PU 70.88 (0.72) 69.32 (2.54) 75.43 (5.07) 72.09 (1.09) 75.89 (1.79) 71.68 (3.84)
    aPU 68.52 (1.75) 66.21 (0.91) 75.71 (8.20) 70.46 (3.35) 73.80 (2.59) 70.71 (3.70)
    VPU 67.44 (0.65) 64.74 (1.12) 76.68 (3.60) 70.16 (1.08) 73.12 (0.85) 71.11 (0.75)
    ImbPU 68.18 (0.83) 67.54 (2.52) 70.64 (6.54) 68.83 (1.94) 73.81 (0.71) 70.46 (1.07)
    Dist-PU 71.57 (0.62) 68.48 (1.16) 80.09 (5.10) 73.74 (1.64) 77.13 (0.69) 73.33 (1.47)
  9. Knowl 9 — Ablation of Dist-PU Components on CIFAR-10

    data/table

    An ablation study evaluated the contribution of label distribution alignment (hatRlab\\hat{R}_{lab}), entropy minimization (Lent/Lent′L_{ent} / L_{ent}'), and Mixup regularization (LmixL_{mix}) on CIFAR-10.

    Variant R^lab\hat{R}_{lab} Lent/Lent′L_{ent} / L_{ent}' LmixL_{mix} ACC (%) Prec. (%) Rec. (%) F1 (%) AUC (%) AP (%)
    I 84.92 (0.89) 83.59 (0.89) 77.57 (3.23) 80.43 (1.56) 92.29 (0.63) 88.23 (0.79)
    II ✓ 89.30 (0.48) 86.96 (0.86) 86.19 (1.12) 86.57 (0.52) 95.40 (0.52) 93.22 (0.97)
    III ✓ 87.50 (0.24) 83.55 (0.17) 85.62 (0.78) 84.57 (0.36) 93.61 (0.33) 89.36 (0.98)
    IV ✓ 86.97 (0.23) 84.50 (4.35) 83.22 (6.47) 83.57 (1.11) 94.37 (0.22) 91.21 (0.73)
    V ✓ ✓ 89.41 (0.48) 86.40 (0.98) 87.30 (1.75) 86.83 (0.69) 95.61 (0.44) 93.51 (0.46)
    VI ✓ ✓ 91.62 (0.47) 90.96 (1.07) 87.78 (0.32) 89.34 (0.54) 96.80 (0.44) 95.54 (0.35)
    VII ✓ ✓ 87.67 (0.52) 85.01 (1.78) 84.07 (1.90) 84.51 (0.60) 94.32 (0.35) 90.65 (0.80)
    VIII ✓ ✓ ✓ 91.88 (0.52) 89.87 (1.09) 89.84 (0.81) 89.85 (0.62) 96.92 (0.45) 95.49 (0.72)

    The results establish three findings:

    1. Label distribution alignment (hatRlab\\hat{R}_{lab}) is primary, providing the single largest performance jump (Variant II vs. Variant I: ACC improves from 84.92%84.92\% to 89.30%89.30\%).
    2. Entropy minimization provides consistent modest gains when combined with other modules.
    3. Mixup provides large gains primarily when coupled with label distribution alignment (Variant VI vs. II improves ACC by 2.32%2.32\%, whereas Variant IV vs. I improves ACC by 2.05%2.05\% with much higher variance on precision and recall).
  10. Knowl 10 — Hyperparameter Sensitivity in Dist-PU

    empirical result

    Empirical sensitivity analysis of Dist-PU hyperparameters on CIFAR-10 revealed the following characteristics:

    1. Entropy weight μ\mu: When tested in isolation without Mixup, model performance across Accuracy, F1, AUC, and AP is stable for μ∈[0,0.1]\mu \in [0, 0.1]. Setting μ>0.5\mu > 0.5 causes severe degradation due to overly aggressive entropy penalization reinforcing misclassifications.

    2. Beta distribution parameter α\alpha: Larger values of α\alpha (e.g., α≥2.0\alpha \ge 2.0 up to 21.021.0) produce a narrower Beta distribution centered near λ=0.5\lambda = 0.5, which yields well-interpolated soft training targets between positive and unlabeled instances and improves overall accuracy.

    3. Mixup weights ν\nu and γ\gamma: Increasing the Mixup consistency weight ν\nu (up to ν≈9\nu \approx 9) directly weakens confirmation bias and yields higher accuracy. A moderate non-zero value for mixed-sample entropy weight γ\gamma (e.g., 0.1≤γ≤0.30.1 \le \gamma \le 0.3) provides additional positive regularizing effect on accuracy.

Coverage note — None was omitted; all core theoretical bounds, methodology components, experimental protocols, comparative tabular results, ablation data, and parameter sensitivity findings are fully included.

References

  1. 1.Eric Arazo, Diego Ortego, Paul Albert, Noel E. O’Connor, and Kevin McGuinness. Pseudo-labeling and confirmation bias in deep semi-supervised learning. In Proceedings of the IEEE International Joint Conference on Neural Networks, pages 1–8, 2020.
  2. 2.Jessa Bekker and Jesse Davis. Estimating the class prior in positive and unlabeled data through decision tree induction. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, pages 2712–2719, 2018.
  3. 3.Jessa Bekker and Jesse Davis. Learning from positive and unlabeled data: a survey. Machine Learning, 109(4):719–760, 2020.
  4. 4.Tristan Bepler, Andrew Morin, Micah Rapp, Julia Brasch, Lawrence Shapiro, Alex J Noble, and Bonnie Berger. Positive-unlabeled convolutional neural networks for particle picking in cryo-electron micrographs. Nature Methods, 16(11):1153–1160, 2019.
  5. 5.David Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring. In Proceedings of the 8th International Conference on Learning Representations, 2020.
  6. 6.David Berthelot, Nicholas Carlini, Ian J. Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems, pages 5050–5060, 2019.
  7. 7.Shizhen Chang, Bo Du, and Liangpei Zhang. Positive unlabeled learning with class-prior approximation. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, pages 2014–2021, 2020.
  8. 8.Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik. Vicinal risk minimization. In Advances in Neural Information Processing Systems, pages 416–422, 2000.
  9. 9.Sneha Chaudhari and Shirish K. Shevade. Learning from positive and unlabelled examples using maximum margin clustering. In Proceedings of the 19th International Conference on Neural Information Processing, volume 7665, pages 465–473, 2012.
  10. 10.Hui Chen, Fangqing Liu, Yin Wang, Liyue Zhao, and Hao Wu. A variational approach for learning from positive and unlabeled data. In Advances in Neural Information Processing Systems, 2020.
  11. 11.Xuxi Chen, Wuyang Chen, Tianlong Chen, Ye Yuan, Chen Gong, Kewei Chen, and Zhangyang Wang. Self-pu: Self boosted and calibrated positive-unlabeled training. In Proceedings of the 37th International Conference on Machine Learning, volume 119, pages 1510–1519, 2020.
  12. 12.Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Analysis of learning from positive and unlabeled data. In Advances in Neural Information Processing Systems, pages 703–711, 2014.
  13. 13.Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Class-prior estimation for learning from positive and unlabeled data. In Proceedings of The 7th Asian Conference on Machine Learning, volume 45, pages 221–236, 2015.
  14. 14.Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Convex formulation for learning from positive and unlabeled data. In Proceedings of the 32nd International Conference on Machine Learning, volume 37, pages 1386–1394, 2015.
  15. 15.Marthinus Christoffel du Plessis, Gang Niu, and Masashi Sugiyama. Class-prior estimation for learning from positive and unlabeled data. Machine Learning, 106(4):463–492, 2017.
  16. 16.Chen Gong, Hong Shi, Tongliang Liu, Chuang Zhang, Jian Yang, and Dacheng Tao. Loss decomposition and centroid estimation for positive and unlabeled learning. IEEE transactions on pattern analysis and machine intelligence, 43(3):918–932, 2019.
  17. 17.Chen Gong, Qizhou Wang, Tongliang Liu, Bo Han, Jane J You, Jian Yang, and Dacheng Tao. Instance-dependent positive and unlabeled learning with labeling bias estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  18. 18.Yves Grandvalet and Yoshua Bengio. Semi-supervised learning by entropy minimization. In Advances in Neural Information Processing Systems, pages 529–536, 2004.
  19. 19.Zayd Hammoudeh and Daniel Lowd. Learning from positive and unlabeled data with arbitrary positive shift. In Advances in Neural Information Processing Systems, 2020.
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  21. 21.Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory F. Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically. CoRR, abs/1712.00409, 2017.
  22. 22.Ming Hou, Brahim Chaib-Draa, Chao Li, and Qibin Zhao. Generative adversarial positive-unlabeled learning. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, pages 2255–2261, 2018.
  23. 23.Yu-Guan Hsieh, Gang Niu, and Masashi Sugiyama. Classification from positive, unlabeled and biased negative data. In Proceedings of the 36th International Conference on Machine Learning, volume 97, pages 2820–2829, 2019.
  24. 24.Wenpeng Hu, Ran Le, Bing Liu, Feng Ji, Jinwen Ma, Dongyan Zhao, and Rui Yan. Predictive adversarial learning from positive and unlabeled data. In Thirty-Fifth AAAI Conference on Artificial Intelligence, pages 7806–7814, 2021.
  25. 25.Masahiro Kato, Takeshi Teshima, and Junya Honda. Learning from positive and unlabeled data with a selection bias. In Preceedings of the 7th International Conference on Learning Representations, 2019.
  26. 26.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations, 2015.
  27. 27.Ryuichi Kiryo, Gang Niu, Marthinus Christoffel du Plessis, and Masashi Sugiyama. Positive-unlabeled learning with non-negative risk estimator. In Advances in Neural Information Processing Systems, pages 1675–1685, 2017.
  28. 28.Alex Krizhevsky et al. Learning multiple layers of features from tiny images. 2009.
  29. 29.Junnan Li, Richard Socher, and Steven C. H. Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. In Proceedings of the 8th International Conference on Learning Representations, 2020.
  30. 30.Bing Liu, Wee Sun Lee, Philip S. Yu, and Xiaoli Li. Partially supervised classification of text documents. In Proceedings of the Nineteenth International on Machine Learning, pages 387–394, 2002.
  31. 31.Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradient descent with warm restarts. In Proceedings of the 5th International Conference on Learning Representations, 2017.
  32. 32.Chuan Luo, Pu Zhao, Chen Chen, Bo Qiao, Chao Du, Hongyu Zhang, Wei Wu, Shaowei Cai, Bing He, Saravanakumar Rajmohan, and Qingwei Lin. PULNS: positive-unlabeled learning with effective negative sample selector. In Thirty-Fifth AAAI Conference on Artificial Intelligence, pages 8784–8792, 2021.
  33. 33.Dhruv Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. Exploring the limits of weakly supervised pretraining. In Proceedings of European Conference on Computer Vision, volume 11206, pages 185–201, 2018.
  34. 34.Curtis G. Northcutt, Tailin Wu, and Isaac L. Chuang. Learning with confident examples: Rank pruning for robust classification with noisy labels. In Proceedings of the Thirty-Third Conference on Uncertainty in Artificial Intelligence, 2017.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:140:1–140:67, 2020.
  36. 36.Harish G. Ramaswamy, Clayton Scott, and Ambuj Tewari. Mixture proportion estimation via kernel embeddings of distributions. In Proceedings of the 33nd International Conference on Machine Learning, volume 48, pages 2052–2060, 2016.
  37. 37.Patrice Y. Simard, Yann LeCun, John S. Denker, and Bernard Victorri. Transformation invariance in pattern recognition-tangent distance and tangent propagation. In Neural Networks: Tricks of the Trade, volume 1524, pages 239–27. Springer, 1996.
  38. 38.Guangxin Su, Weitong Chen, and Miao Xu. Positive-unlabeled learning from imbalanced data. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 2995–3001, 2021.
  39. 39.Vladimir Vapnik. Statistical learning theory. Wiley, 1998.
  40. 40.Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. CoRR, 2017.
  41. 41.Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, and Quoc Le. Unsupervised data augmentation for consistency training. In Advances in Neural Information Processing Systems, 2020.
  42. 42.Hwanjo Yu, Jiawei Han, and Kevin Chen-Chuan Chang. PEBL: web page classification without negative examples. IEEE Transactions on Knowledge and Data Engineering, 16(1):70–81, 2004.
  43. 43.Bangzuo Zhang and Wanli Zuo. Reliable negative extracting based on knn for learning from positive and unlabeled examples. Journal of Computers, 4(1):94–101, 2009.
  44. 44.Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In Proceedings of the 6th International Conference on Learning Representations, 2018.
  45. 45.Ya-Lin Zhang, Longfei Li, Jun Zhou, Xiaolong Li, Yujiang Liu, Yuanchao Zhang, and Zhi-Hua Zhou. POSTER: A PU learning based system for potential malicious URL detection. In Proceedings of the Conference on Computer and Communications Security, pages 2599–2601, 2017.
  46. 46.Dengyong Zhou, Olivier Bousquet, Thomas Navin Lal, Jason Weston, and Bernhard Schölkopf. Learning with local and global consistency. In Advances in Neural Information Processing Systems, pages 321–328, 2003.

Citation

MLA
Zhao, Y., et al. “Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective”. arXiv, 2022, http://arxiv.org/abs/2212.02801v1.
APA
Zhao, Y., Xu, Q., Jiang, Y., Wen, P., & Huang, Q. (2022). Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective. arXiv. http://arxiv.org/abs/2212.02801v1
Chicago
Zhao, Y., Q. Xu, Y. Jiang, P. Wen, and Q. Huang. 2022. “Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective”. arXiv. http://arxiv.org/abs/2212.02801v1.
Harvard
Zhao, Y. et al. (2022) “Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.02801v1.
Vancouver
1. Zhao Y, Xu Q, Jiang Y, Wen P, Huang Q (2022) Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective. arXiv

BibTeX

@article{zhao2022dist,
  title = {Dist-PU: Positive-Unlabeled Learning from a Label Distribution Perspective},
  author = {Zhao, Yunrui and Xu, Qianqian and Jiang, Yangbangyan and Wen, Peisong and Huang, Qingming},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.02801v1},
  eprint = {2212.02801}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE