Confidence Score for Source-Free Unsupervised Domain Adaptation
Jonghyun LeeDahuin JungJunho YimSungroh Yoon
Proposes a joint model-data structure confidence score and sample-weighted adaptation framework that mitigates noisy pseudo-labeling in source-free unsupervised domain adaptation by combining source model probabilities with target feature cluster distributions.
Deploying machine learning models to new target environments often leads to significant performance drops due to shifts in data distributions. While standard domain adaptation techniques mitigate this issue by using original source training data alongside unlabeled target data, strict privacy laws, security restrictions, and high computational costs frequently make original source datasets unavailable. Existing source-free methods rely only on a pre-trained model and attempt to assign pseudo-labels to target data based on clustering assumptions. However, treating every target sample with equal importance exposes these models to incorrect pseudo-labels and error accumulation, ultimately harming predictive accuracy.
The article develops and evaluates a sample-wise scoring method called the Joint Model-Data Structure (JMDS) score alongside an adaptation framework named Confidence score Weighting Adaptation using JMDS (CoWA-JMDS). The primary objective is to reliably estimate pseudo-label confidence by combining source model knowledge with target data structure, thereby improving adaptation performance without needing original source data.
The authors conducted extensive experimental evaluations using standard vision benchmarks: Office-31, Office-Home, and VisDA-2017. Their approach combines Gaussian Mixture Modeling in the target feature space to capture target domain distribution with pre-trained source model probabilities. The resulting framework applies these confidence scores as sample-specific training weights and incorporates a data augmentation strategy called weight Mixup to safely integrate lower-confidence samples.
The evaluation produced several key findings. First, the JMDS score consistently outperformed existing confidence metrics by achieving lower Area Under Risk-Coverage values across benchmarks, demonstrating superior error discrimination. Second, the CoWA-JMDS framework established new state-of-the-art results for closed-set adaptation, achieving average accuracies of 90.3% on Office-31, 72.5% on Office-Home, and 86.9% on VisDA-2017—surpassing existing source-free methods by 0.3% to 1.0% without requiring auxiliary generative networks. Third, in partial-set settings where target domains contain only a subset of source classes, the framework achieved 83.2% accuracy, beating prior techniques by nearly 4 percentage points. Finally, incorporating weight Mixup provided a 3.4% accuracy boost over standard Mixup on Office-31 by preventing low-confidence samples from introducing severe label noise.
These findings indicate that organizations can successfully adapt high-performing vision models to new operational environments while maintaining compliance with privacy and data-sharing constraints. By weighting training samples by combined confidence scores, teams can reduce the operational risk and manual verification costs associated with erroneous automated labels.
Practitioners facing privacy constraints or compute limitations should adopt sample-weighted adaptation frameworks like CoWA-JMDS rather than uniform pseudo-labeling. However, practitioners should note that weight Mixup degrades performance in open-set scenarios containing entirely unseen classes; in those situations, the article shows that using confidence weighting alone without sample mixing is preferred. While the reported empirical results show high confidence across several benchmark image sets, real-world deployment on complex, non-visual, or noisy domain shifts will require dedicated pilot testing.
- Paper: Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation, Jian Liang et al. (2020). It introduces the fundamental Source Hypothesis Transfer (SHOT) framework for source-free unsupervised domain adaptation, providing the essential baseline paradigm of adapting target features using pre-trained source hypotheses that the source paper seeks to improve via confidence weighting.
- Paper: MixMatch: A Holistic Approach to Semi-Supervised Learning, David Berthelot et al. (2019). It develops the MixMatch framework utilizing Mixup regularization on pseudo-labeled unlabeled batches, which provides the core mixing mechanics adapted by the source paper's proposed Weight Mixup.
- Paper: FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling, Bowen Zhang et al. (2021). It explores dynamic pseudo-label confidence evaluation to address the pitfalls of uniform sample importance across target classes, directly motivating sample-wise reliability metrics in adaptation.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). It establishes the foundational principles of dynamic instance reweighting to combat label noise and error accumulation during neural network optimization.
- Paper: Learning with Local and Global Consistency, Dengyong Zhou et al. (2003). It establishes the seminal mathematical principles of local and global data manifold consistency for label propagation on unlabeled data structures.
- Paper: Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection, Vibashan VS et al. (2023). It extends source-free unsupervised domain adaptation methodologies to complex structured prediction tasks by leveraging instance relation graphs for object detection.
- Paper: Feature Alignment and Uniformity for Test Time Adaptation, Shuai Wang et al. (2023). It advances beyond offline source-free domain adaptation by evaluating target feature alignment and representation uniformity under real-time, test-time adaptation settings.
- Paper: Improved Test-Time Adaptation for Domain Generalization, Liang Chen et al. (2023). It investigates test-time adaptation under domain generalization constraints, building on the principles of unsupervised online target adjustment without source data.
