Balancing Discriminability and Transferability for Source-Free Domain Adaptation

Jogendra Nath KunduAkshay R. KulkarniSuvaansh BhambriDeepesh MehtaShreyas Anand KulkarniVarun JampaniVenkatesh Babu Radhakrishnan

article2022ICML113 citations

Proposes an instance-level mixup strategy between original and generic domain representations to optimize the trade-off between discriminability and transferability in privacy-preserving, source-free domain adaptation across image classification and semantic segmentation tasks.

Listen

Machine learning models often experience substantial drops in performance when deployed in new target environments due to input distribution shifts. While conventional unsupervised domain adaptation aligns models across domains, it requires simultaneous access to labeled source data and unlabeled target data. In privacy-restricted, client-vendor settings, source data cannot be shared, necessitating source-free domain adaptation. Existing methods struggle with a fundamental trade-off: enhancing transferability across domains tends to degrade the model's discriminative ability to categorize tasks accurately, and attempting to map data to a generic representation causes critical domain-specific details to be lost.

The article demonstrates that creating an intermediate mixup domain—by linearly blending original domain samples with approximate generic domain representations—effectively balances transferability and discriminability. This approach provides a theoretically grounded method to tighten the target error bound and boost model performance under strict source-free constraints.

To evaluate this approach, the researchers applied two distinct mixup strategies: an input-space edge-mixup suited for spatial tasks like semantic segmentation, and a feature-space mixup utilizing augmented sub-domains suited for classification tasks. These modifications were integrated directly into established source-free and non-source-free frameworks and tested across multiple standard benchmarks, including Office-31, Office-Home, VisDA, DomainNet, and urban driving datasets (GTA5, SYNTHIA, Synscapes, and Cityscapes).

The evaluation yielded several key findings:

  1. Consistently outperformed prior methods: Adding the mixup framework to existing top-performing source-free models improved average accuracy by 0.7% to 1.6% on single-source benchmarks and by up to 3.6% on multi-source benchmarks (achieving 51.0% on DomainNet and 77.4% on Office-Home).
  2. Enhanced segmentation performance: In semantic segmentation, edge-mixup delivered an average 2.2% improvement across multi-source settings, surpassing even non-source-free approaches without requiring access to original source data.
  3. Accelerated model training: Incorporating the intermediate mixup representation narrowed the domain shift, leading to faster training convergence alongside higher final accuracy.
  4. Broad compatibility across paradigms: The mixup strategy generated substantial performance gains when integrated into conventional, non-source-free adaptation methods (improving baselines by up to 9.8%) and successfully transferred to speech and text benchmarks.

These findings indicate that organizations deploying machine learning models across diverse, privacy-sensitive client environments can achieve state-of-the-art accuracy without compromising data compliance. The mixup integration is lightweight, requires no architectural redesign, and reduces computational training cycles while eliminating the operational risks and liabilities associated with sharing centralized proprietary data.

Stakeholders adopting source-free adaptation should incorporate feature-space mixup for general classification tasks and input-space edge mixup for dense visual tasks, utilizing a small mixing ratio (such as 0.1). Technical teams should prioritize extending these generic domain realizations into automated, learnable pipelines, while decision-makers can proceed with high confidence given the consistent empirical validation across vision, language, and audio tasks.

Cover for Balancing Discriminability and Transferability for Source-Free Domain Adaptation

Abstract

Conventional domain adaptation (DA) techniques aim to improve domain transferability by learning domain-invariant representations; while concurrently preserving the task-discriminability knowledge gathered from the labeled source data. However, the requirement of simultaneous access to labeled source and unlabeled target renders them unsuitable for the challenging source-free DA setting. The trivial solution of realizing an effective original to generic domain mapping improves transferability but degrades task discriminability. Upon analyzing the hurdles from both theoretical and empirical standpoints, we derive novel insights to show that a mixup between original and corresponding translated generic samples enhances the discriminability-transferability trade-off while duly respecting the privacy-oriented source-free setting. A simple but effective realization of the proposed insights on top of the existing source-free DA approaches yields state-of-the-art performance with faster convergence. Beyond single-source, we also outperform multi-source prior-arts across both classification and semantic segmentation benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Theoretical insights
  • 3.1.1. Criteria to realize generic-domain mixup
  • 3.2. Training algorithm and realizable generic-domains
  • 3.2.1. Edge representation as a generic-domain
  • 3.2.2. Feature-space generic-domain representation
  • 4. Experiments
  • 4.1. Comparison with prior arts
  • 4.2. Analysis
  • 4.3. Compatibility with non-source-free DA
  • 5. Conclusion
  • References
  • Appendix
  • A. Notations
  • B. Discussions related to theory
  • B.1. Theorem 1 and Proof
  • B.2. Regarding result relating κ ( p s m , p t m ) and κ ( p s , p t )
  • C. Experiments
  • C.1. Implementation details
  • C.2. Experimental settings
  • C.3. Additional results
  • C.4. Extended comparisons

Knowls

  1. Knowl 1 — Instance-Level Generic-Domain Mixup for Source-Free Domain Adaptation

    model/method

    In source-free domain adaptation (SFDA) under the vendor-client paradigm, labeled source data Ds={(xs,ys)}\mathcal{D}_s = \{(x_s, y_s)\} drawn from distribution psp_s and unlabeled target data Dt={xt}\mathcal{D}_t = \{x_t\} drawn from distribution ptp_t are not accessible concurrently. Translating data entirely to a domain-generic representation (psgp_{s_g} or ptgp_{t_g}) maximizes domain transferability but severely harms task discriminability because domain-specific and task-discriminative features (such as color and texture) are non-linearly entangled.

    To balance transferability and discriminability, an instance-level convex mixup between an original sample and its corresponding generic-domain representation is performed independently on the vendor (source) side and client (target) side. In the input space, the mixup sample xmx_m is formed as:

    xm=λxg+(1−λ)xx_m = \lambda x_g + (1 - \lambda) x

    where x∈Xx \in \mathcal{X} is the original input image, xgx_g is its corresponding domain-generic translation, and λ∈[0,1]\lambda \in [0, 1] is a fixed mixup ratio. In the feature space of a feature extractor backbone h:X→Zh: \mathcal{X} \to \mathcal{Z}, the feature-mixup representation zmz_m is formed as:

    zm=λzg+(1−λ)zz_m = \lambda z_g + (1 - \lambda) z

    where z=h(x)z = h(x) and zgz_g is the generic feature representation for sample xx. Because the interpolation is performed strictly within each instance (between the instance and its own generic counterpart), the ground-truth semantic class label yy is preserved without requiring convex combinations of class labels.

  2. Knowl 2 — Reduction of H-Divergence via Generic-Domain Mixup

    theoretical result

    Let psp_s and ptp_t denote the marginal distributions of the source and target domains over input space X\mathcal{X}, and let psgp_{s_g} and ptgp_{t_g} denote the marginal distributions of their respective generic-domain representations. Let psmp_{s_m} and ptmp_{t_m} denote the marginal distributions of the mixup representations generated by xsm=λxsg+(1−λ)xsx_{s_m} = \lambda x_{s_g} + (1 - \lambda) x_s and xtm=λxtg+(1−λ)xtx_{t_m} = \lambda x_{t_g} + (1 - \lambda) x_t with mixup ratio λ∈[0,1]\lambda \in [0, 1].

    Assume the following preconditions:

    1. The original source psp_s and target ptp_t are easily separable by a linear domain classifier fd:Z→Rf_d: \mathcal{Z} \to \mathbb{R} (achieving perfect domain classification accuracy).
    2. The generic source distribution psgp_{s_g} and generic target distribution ptgp_{t_g} are inseparable, such that the classification accuracy of fdf_d equals that of a random classifier (1/21/2).
    3. The domain classifier fdf_d is linear, satisfying fd(ka+b)=kfd(a)+fd(b)f_d(k a + b) = k f_d(a) + f_d(b) for any scalar kk.

    Under these assumptions, the H\mathcal{H}-divergence between the mixup distributions is upper-bounded by the H\mathcal{H}-divergence between the original distributions:

    dH(psm,ptm)≤dH(ps,pt)d_{\mathcal{H}}(p_{s_m}, p_{t_m}) \le d_{\mathcal{H}}(p_s, p_t)

    where the H\mathcal{H}-divergence for hypothesis space H\mathcal{H} is defined as:

    dH(p,q)=sup⁡h′∈H∣Ex∼p[1(fd(h′(x))=1)]−Ex∼q[1(fd(h′(x))=1)]∣d_{\mathcal{H}}(p, q) = \sup_{h' \in \mathcal{H}} \left| \mathbb{E}_{x \sim p}[1(f_d(h'(x)) = 1)] - \mathbb{E}_{x \sim q}[1(f_d(h'(x)) = 1)] \right|

    This reduction in dHd_{\mathcal{H}} increases domain transferability while keeping joint optimal error κ(psm,ptm)\kappa(p_{s_m}, p_{t_m}) low, thereby providing a tighter upper bound on expected target risk ϵt(h)≤ϵs(h)+dH(psm,ptm)+κ(psm,ptm)\epsilon_t(h) \le \epsilon_s(h) + d_{\mathcal{H}}(p_{s_m}, p_{t_m}) + \kappa(p_{s_m}, p_{t_m}).

  3. Knowl 3 — Transferability and Discriminability Metrics in Domain Adaptation

    equation

    To evaluate the trade-off between transferability (invariance across domains) and discriminability (separation of task categories) for a backbone hypothesis space H\mathcal{H}, the transferability metric γT\gamma_T and discriminability metric γD\gamma_D are defined as:

    γT=1−dH(ps,pt)\gamma_T = 1 - d_{\mathcal{H}}(p_s, p_t)

    γD=1−12κ(ps,pt)\gamma_D = 1 - \frac{1}{2} \kappa(p_s, p_t)

    where dH(ps,pt)d_{\mathcal{H}}(p_s, p_t) is the H\mathcal{H}-divergence between source marginal psp_s and target marginal ptp_t given a domain classifier fd:Z→{0,1}f_d: \mathcal{Z} \to \{0, 1\}:

    dH(ps,pt)=sup⁡h′∈H∣Ex∼ps[1(fd(h′(x))=1)]−Ex∼pt[1(fd(h′(x))=1)]∣d_{\mathcal{H}}(p_s, p_t) = \sup_{h' \in \mathcal{H}} \left| \mathbb{E}_{x \sim p_s}[1(f_d(h'(x)) = 1)] - \mathbb{E}_{x \sim p_t}[1(f_d(h'(x)) = 1)] \right|

    and κ(ps,pt)\kappa(p_s, p_t) is the error of the joint ideal hypothesis on both domains:

    κ(ps,pt)=min⁡h′∈H(ϵs(h′)+ϵt(h′))\kappa(p_s, p_t) = \min_{h' \in \mathcal{H}} \left( \epsilon_s(h') + \epsilon_t(h') \right)

    with ϵs(h′)\epsilon_s(h') and ϵt(h′)\epsilon_t(h') denoting the expected source and target risks respectively. Since 0≤dH(ps,pt)≤10 \le d_{\mathcal{H}}(p_s, p_t) \le 1 and 0≤κ(ps,pt)≤20 \le \kappa(p_s, p_t) \le 2, both metrics satisfy γT,γD∈[0,1]\gamma_T, \gamma_D \in [0, 1].

  4. Knowl 4 — Feature-Space Generic Domain Realization via Augmented Sub-Domain Averaging

    model/method

    To realize a generic domain representation in deep feature space for non-dense classification tasks without discarding task-discriminative information, an instance is perturbed across KK sub-domains using domain-variant, task-preserving augmentations {A[i](⋅)}i=1K\{A^{[i]}(\cdot)\}_{i=1}^K. The augmentations chosen are:

    1. Fourier Domain Adaptation (FDA): swaps low-frequency FFT amplitude with reference images from an auxiliary style dataset while preserving phase/semantics.
    2. Adaptive Instance Normalization (AdaIN): alters feature statistics in instance normalization layers to stylize the image (stylization strength set to 0.30.3).
    3. Frost / Weather Augmentation: introduces synthetic weather artifacts at mild severity levels.
    4. Cartoonization: non-photorealistic artistic filter.

    For an input xx, the feature extractor backbone hh extracts features for each augmented view z[i]=h(A[i](x))z^{[i]} = h(A^{[i]}(x)). The generic feature representation zgz_g is the empirical mean over all KK sub-domain feature vectors:

    zg=1K∑i=1Kz[i]z_g = \frac{1}{K} \sum_{i=1}^K z^{[i]}

    The feature-mixup representation is computed as zm=λzg+(1−λ)h(x)z_m = \lambda z_g + (1 - \lambda) h(x). During backpropagation, gradients are stopped from flowing into zgz_g (only propagating through z=h(x)z = h(x)) to prevent conflicting gradients across the individual sub-domain transformations from destabilizing training. At deployment, the model is fine-tuned for a small number of iterations on the raw target dataset using the client adaptation algorithm, eliminating the need for feature-mixup computation during test inference.

  5. Knowl 5 — Edge-Mixup for Input-Space Generic Domain Realization

    model/method

    For dense prediction tasks such as semantic segmentation where object recognition depends heavily on spatial structure and object boundaries, edge maps serve as an approximate input-space generic domain. Edge representations preserve shape features while removing domain-variant textures and colors.

    Given an input image xx, a pre-trained CNN edge detector Ae(⋅)\mathcal{A}_e(\cdot) extracts edge map xg=Ae(x)x_g = \mathcal{A}_e(x). The edge-mixup sample xmx_m is computed directly in the input pixel space:

    xm=λAe(x)+(1−λ)xx_m = \lambda \mathcal{A}_e(x) + (1 - \lambda) x

    where λ∈[0,1]\lambda \in [0, 1] is the mixup coefficient (typically set to λ=0.1\lambda = 0.1). Labeled source dataset Ds\mathcal{D}_s and unlabeled target dataset Dt\mathcal{D}_t are converted to edge-mixup datasets Dsm={(xsm,ys)}\mathcal{D}_{s_m} = \{(x_{s_m}, y_s)\} and Dtm={xtm}\mathcal{D}_{t_m} = \{x_{t_m}\} prior to training.

  6. Knowl 6 — Generic-Domain Mixup Training Protocol for Source-Free Domain Adaptation

    algorithm

    The generic mixup framework integrates as a plug-and-play dataset wrapper into arbitrary source-free domain adaptation (SFDA) algorithms without modifying loss functions or network architectures.

    Input: Labeled source dataset DsD_s, Unlabeled target dataset DtD_t, Mixup type T∈{"edge-mixup","feature-mixup"}T \in \{\text{"edge-mixup"}, \text{"feature-mixup"}\}, Vendor-side algorithm VsAlgo\text{VsAlgo}, Client-side algorithm CsAlgo\text{CsAlgo}, Mixup ratio λ\lambda
    Output: Adapted client model (h,fc)(h, f_c)
    // Step 1: Create Mixup Datasets
    if T=="edge-mixup"T == \text{"edge-mixup"}:
        Dsm={(λAe(x)+(1−λ)x,y)∣(x,y)∈Ds}D_{s_m} = \{(\lambda A_e(x) + (1 - \lambda) x, y) \mid (x, y) \in D_s\}
        Dtm={λAe(x)+(1−λ)x∣x∈Dt}D_{t_m} = \{\lambda A_e(x) + (1 - \lambda) x \mid x \in D_t\}
    elif T=="feature-mixup"T == \text{"feature-mixup"}:
        Dsm={(λ1K∑i=1Kh(A[i](x))+(1−λ)h(x),y)∣(x,y)∈Ds}D_{s_m} = \{(\lambda \frac{1}{K}\sum_{i=1}^K h(A^{[i]}(x)) + (1 - \lambda) h(x), y) \mid (x, y) \in D_s\}
        Dtm={λ1K∑i=1Kh(A[i](x))+(1−λ)h(x)∣x∈Dt}D_{t_m} = \{\lambda \frac{1}{K}\sum_{i=1}^K h(A^{[i]}(x)) + (1 - \lambda) h(x) \mid x \in D_t\}
    // Step 2: Vendor-Side Source Training
    Initialize feature backbone hh and classifier fcf_c
    Train (h(s),fc(s))=arg⁡min⁡h,fcJ(VsAlgo(Dsm))(h^{(s)}, f_c^{(s)}) = \arg\min_{h, f_c} \mathcal{J}(\text{VsAlgo}(D_{s_m}))
    Vendor shares model weights (h(s),fc(s))(h^{(s)}, f_c^{(s)}) with Client (no source data shared)
    // Step 3: Client-Side Target Adaptation
    Initialize client model with (h(s),fc(s))(h^{(s)}, f_c^{(s)})
    Train (h(t),fc(t))=arg⁡min⁡h,fcJ(CsAlgo(Dtm))(h^{(t)}, f_c^{(t)}) = \arg\min_{h, f_c} \mathcal{J}(\text{CsAlgo}(D_{t_m}))
    // Step 4: Inference Finetuning
    Finetune (h(t),fc(t))(h^{(t)}, f_c^{(t)}) on original DtD_t using CsAlgo(Dt)\text{CsAlgo}(D_t) for a few iterations
    return (h(t),fc(t))(h^{(t)}, f_c^{(t)})
  7. Knowl 7 — Classification Performance on Single-Source and Multi-Source Domain Adaptation Benchmarks

    data/table

    Integrating edge-mixup and feature-mixup into leading source-free domain adaptation baselines (NRC and SHOT++) yields consistent performance improvements across single-source (SSDA) and multi-source (MSDA) classification benchmarks using ResNet-50 (Office-31, Office-Home, DomainNet) and ResNet-101 (VisDA).

    Method SF Office-31 Avg Office-Home Avg VisDA (S→\toR) DomainNet Avg
    FixBi (Non-SF SOTA) 91.4 72.7 87.2 -
    STEM (Non-SF SOTA) - - - 53.4
    NRC (Baseline SF SOTA) ✓ 89.4 72.2 85.9 47.4
    Ours (edge-mixup) + NRC ✓ 90.3 73.0 86.4 49.6
    Ours (feat-mixup) + NRC ✓ 90.5 73.8 87.3 51.0
    SHOT++ (Baseline SF SOTA) ✓ 89.2 73.0 87.3 -
    Ours (edge-mixup) + SHOT++ ✓ 90.2 73.7 87.5 -
    Ours (feat-mixup) + SHOT++ ✓ 90.7 74.5 87.8 -

    On Office-31 SSDA, feature-mixup improves NRC from 89.4%89.4\% to 90.5%90.5\% and SHOT++ from 89.2%89.2\% to 90.7%90.7\%. On Office-Home SSDA, feature-mixup improves NRC by +1.6%+1.6\% (73.8%73.8\%) and SHOT++ by +1.5%+1.5\% (74.5%74.5\%), exceeding prior non-source-free methods. On DomainNet MSDA (without domain labels), feature-mixup improves NRC by +3.6%+3.6\% (51.0%51.0\% vs. 47.4%47.4\%).

  8. Knowl 8 — Domain Adaptive Semantic Segmentation Performance under Edge-Mixup and Feature-Mixup

    data/table

    In semantic segmentation DA (evaluated on Cityscapes target validation set with standard DeepLabv2 with ResNet-101 backbone), edge-mixup outperforms feature-mixup. This occurs because edge detection directly preserves spatial shape boundaries critical for dense pixel-level prediction, whereas feature-mixup on spatial convolutional feature maps causes spatial boundary diffusion.

    Method SF SSDA (G) SSDA (Y) MSDA (G+S) MSDA (S+Y) MSDA (G+Y) MSDA (G+S+Y)
    FDA 50.5 52.5 - - - -
    ProDA 57.5 62.0 - - - -
    MSDA-CL - - 65.8 63.1 59.4 67.1
    SFDA ✓ 43.1 45.9 - - - -
    URMA ✓ 45.1 45.0 - - - -
    SFUDA ✓ 49.4 51.9 - - - -
    GtA (Baseline SF) ✓ 51.6 55.5 63.5 62.8 58.3 63.4
    Ours (feat-mixup) ✓ 51.9 55.6 63.6 63.2 61.4 64.3
    Ours (edge-mixup) ✓ 52.6 56.7 64.6 65.4 61.8 64.9

    Here G, Y, and S represent synthetic source domains GTA5 (19-class mIoU), SYNTHIA (13-class mIoU), and Synscapes (13-class mIoU), respectively. Edge-mixup yields consistent gains over the source-free SOTA (GtA) across all single-source (+1.0%+1.0\% on G, +1.2%+1.2\% on Y) and multi-source settings (+2.2%+2.2\% average gain), also surpassing the non-source-free SOTA MSDA-CL by an average of 0.4%0.4\%.

  9. Knowl 9 — Sensitivity to Mixup Ratio and Trade-Off Dynamics

    empirical result

    Empirical evaluation of the mixup ratio λ∈[0,1]\lambda \in [0, 1] on the Office-Home benchmark indicates distinct sensitivity behaviors for input-space edge-mixup versus feature-space mixup:

    1. For edge-mixup, performance increases from λ=0\lambda = 0 (baseline accuracy ∼72.2%\sim 72.2\%) up to an optimum at λ=0.1\lambda = 0.1 (73.0%73.0\%), but sharply degrades when λ>0.2\lambda > 0.2 (falling below 60%60\% at λ=1.0\lambda = 1.0). At high λ\lambda, the dominance of edge structures removes vital color and texture features, drastically reducing task discriminability γD\gamma_D.
    2. For feature-mixup, performance remains robust and superior to the baseline across a broad range of λ∈[0.1,0.8]\lambda \in [0.1, 0.8], peaking near λ=0.1\lambda = 0.1 (73.8%73.8\%). The feature mean across augmented sub-domains filters domain-specific noise while retaining semantic activations distributed across feature channels.
    3. Across all classification and segmentation benchmarks, λ=0.1\lambda = 0.1 provides the optimal empirical balance between transferability γT\gamma_T and discriminability γD\gamma_D, leading to faster optimization convergence.
  10. Knowl 10 — Compatibility of Feature-Mixup with Non-Source-Free Domain Adaptation Algorithms

    data/table

    Although designed for source-free DA, the feature-mixup framework is universally compatible with classical non-source-free domain adaptation methods that access source and target data concurrently. Training baseline models with feature-mixup representations on Office-Home yields substantial gains across single-source (SSDA) and multi-source (MSDA) settings.

    Method SSDA Avg Acc (%) MSDA Avg Acc (%)
    DANN 57.6 64.6
    DANN + feature-mixup 67.2 (+9.6) 72.9 (+8.3)
    CDAN+E 65.8 69.4
    CDAN+E + feature-mixup 72.2 (+6.4) 74.1 (+4.7)
    SRDC 71.3 73.1
    SRDC + feature-mixup 72.3 (+1.0) 75.1 (+2.0)

    Feature-mixup provides an absolute improvement of +9.6%+9.6\% and +8.3%+8.3\% on DANN for SSDA and MSDA respectively, +6.4%+6.4\% and +4.7%+4.7\% on CDAN+E, and +1.0%+1.0\% and +2.0%+2.0\% on SRDC, confirming that pre-aligning sub-domain distributions via instance mixup assists adversarial and clustering-based cross-domain alignment.

Coverage note — Multi-target domain adaptation (MTDA) results (Table 4), qualitative segmentation figures, and preliminary proof-of-concept experiments on audio (Libri-Adapt) and text (Amazon Reviews) in the appendix were omitted as secondary evaluations covered by the main classification and segmentation knowls.

References

  1. 1.Aggarwal, S., Kundu, J. N., Radhakrishnan, V. B., and Chakraborty, A. WAMDA: Weighted alignment of sources for multi-source domain adaptation. In BMVC, 2020.
  2. 2.Ahmed, S. M., Raychaudhuri, D. S., Paul, S., Oymak, S., and Roy-Chowdhury, A. K. Unsupervised multi-source domain adaptation without access to source data. In CVPR, 2021.
  3. 3.Ahmed, W., Morerio, P., and Murino, V. Cleaning noisy labels by negative ensemble learning for source-free unsupervised domain adaptation. In WACV, 2022.
  4. 4.Awais, M., Zhou, F., Xu, H., Hong, L., Luo, P., Bae, S.-H., and Li, Z. Adversarial robustness for unsupervised domain adaptation. In ICCV, 2021.
  5. 5.Ben-David, S., Blitzer, J., Crammer, K., and Pereira, F. Analysis of representations for domain adaptation. In NeurIPS, 2006.
  6. 6.Blitzer, J., Dredze, M., and Pereira, F. Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification. In ACL, 2007.
  7. 7.Chen, C., Zheng, Z., Ding, X., Huang, Y., and Dou, Q. Harmonizing transferability and discriminability for adapting object detectors. In CVPR, 2020.
  8. 8.Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40 (4):834–848, 2017.
  9. 9.Chen, X., Wang, S., Long, M., and Wang, J. Transferability vs. discriminability: Batch spectral penalization for adversarial domain adaptation. In ICML, 2019a.
  10. 10.Chen, Z., Zhuang, J., Liang, X., and Lin, L. Blending-target domain adaptation by adversarial meta-adaptation networks. In CVPR, 2019b.
  11. 11.Chou, H.-P., Chang, S.-C., Pan, J.-Y., Wei, W., and Juan, D.-C. Remix: Rebalanced mixup. In ECCV, 2020.
  12. 12.Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  13. 13.Dong, J., Fang, Z., Liu, A., Sun, G., and Liu, T. Confident anchor-induced multi-source free domain adaptation. In NeurIPS, 2021.
  14. 14.Fu, Y., Zhang, M., Xu, X., Cao, Z., Ma, C., Ji, Y., Zuo, K., and Lu, H. Partial feature selection and alignment for multi-source domain adaptation. In CVPR, 2021.
  15. 15.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V. Domain-adversarial training of neural networks. The Journal of Machine Learning Research, 17(1):2096–2030, 2016.
  16. 16.Gu, X., Sun, J., and Xu, Z. Spherical space domain adaptation with robust pseudo-label loss. In CVPR, 2020.
  17. 17.He, J., Jia, X., Chen, S., and Liu, J. Multi-source domain adaptation with collaborative learning for semantic segmentation. In CVPR, 2021.
  18. 18.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  19. 19.Hou, W., Wang, J., Tan, X., Qin, T., and Shinozaki, T. Cross-domain speech recognition with unsupervised character-level distribution matching. In Interspeech, 2021.
  20. 20.Huang, J., Guan, D., Xiao, A., and Lu, S. RDA: Robust domain adaptation via fourier adversarial attacking. In ICCV, 2021a.
  21. 21.Huang, J., Guan, D., Xiao, A., and Lu, S. Model adaptation: Historical contrastive learning for unsupervised domain adaptation without source data. In NeurIPS, 2021b.
  22. 22.Huang, X. and Belongie, S. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017.
  23. 23.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  24. 24.Jin, X., Lan, C., Zeng, W., and Chen, Z. Re-energizing domain discriminator with sample relabeling for adversarial domain adaptation. In ICCV, 2021.
  25. 25.Jung, A. B., Wada, K., Crall, J., Tanaka, S., Graving, J., Reinders, C., Yadav, S., Banerjee, J., Vecsei, G., Kraft, A., Rui, Z., Borovec, J., Vallentin, C., Zhydenko, S., Pfeiffer, K., Cook, B., Fernández, I., De Rainville, F.-M., Weng, C.-H., Ayala-Acevedo, A., Meudec, R., Laporte, M., et al. imgaug. https://github.com/aleju/imgaug, 2020. Online; accessed 01-Feb-2020.
  26. 26.Karouzos, C. F., Paraskevopoulos, G., and Potamianos, A. UDALM: Unsupervised domain adaptation through language modeling. In NAACL, 2021.
  27. 27.Kim, D., Saito, K., Oh, T.-H., Plummer, B. A., Sclaroff, S., and Saenko, K. CDS: Cross-domain self-supervised pre-training. In ICCV, 2021a.
  28. 28.Kim, J.-H., Choo, W., Jeong, H., and Song, H. O. Co-mixup: Saliency guided joint mixup with supermodular diversity. In ICLR, 2021b.
  29. 29.Kim, Y., Cho, D., Han, K., Panda, P., and Hong, S. Domain adaptation without source data. IEEE Transactions on Artificial Intelligence, 2(6):508–518, 2021c.
  30. 30.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  31. 31.Krueger, D., Caballero, E., Jacobsen, J.-H., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A. Out-of-distribution generalization via risk extrapolation (rex). In ICML, 2021.
  32. 32.Kundu, J. N., Venkat, N., and Babu, R. V. Universal source-free domain adaptation. In CVPR, 2020a.
  33. 33.Kundu, J. N., Venkat, N., Revanur, A., V, R. M., and Babu, R. V. Towards inheritable models for open-set domain adaptation. In CVPR, 2020b.
  34. 34.Kundu, J. N., Venkatesh, R. M., Venkat, N., Revanur, A., and Babu, R. V. Class-incremental domain adaptation. In ECCV, 2020c.
  35. 35.Kundu, J. N., Kulkarni, A., Singh, A., Jampani, V., and Babu, R. V. Generalize then adapt: Source-free domain adaptive semantic segmentation. In ICCV, 2021.
  36. 36.Kundu, J. N., Kulkarni, A., Bhambri, S., Jampani, V., and Babu, R. V. Amplitude spectrum transformation for open compound domain adaptive semantic segmentation. In AAAI, 2022.
  37. 37.Li, D., Yang, Y., Song, Y.-Z., and Hospedales, T. M. Deeper, broader and artier domain generalization. In ICCV, 2017.
  38. 38.Li, R., Jiao, Q., Cao, W., Wong, H.-S., and Wu, S. Model adaptation: Unsupervised domain adaptation without source data. In CVPR, 2020.
  39. 39.Li, S., Xie, M., Lv, F., Liu, C. H., Liang, J., Qin, C., and Li, W. Semantic concentration for domain adaptation. In ICCV, 2021a.
  40. 40.Li, Y., Yuan, L., Chen, Y., Wang, P., and Vasconcelos, N. Dynamic transfer for multi-source domain adaptation. In CVPR, 2021b.
  41. 41.Liang, J., Hu, D., and Feng, J. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In ICML, 2020.
  42. 42.Liang, J., Hu, D., Wang, Y., He, R., and Feng, J. Source data-absent unsupervised domain adaptation through hypothesis transfer and labeling transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  43. 43.Liu, Y., Zhang, W., and Wang, J. Source-free domain adaptation for semantic segmentation. In CVPR, 2021.
  44. 44.Long, M., Cao, Y., Wang, J., and Jordan, M. Learning transferable features with deep adaptation networks. In ICML, 2015.
  45. 45.Long, M., Zhu, H., Wang, J., and Jordan, M. I. Deep transfer learning with joint adaptation networks. In International conference on machine learning, pp. 2208–2217. PMLR, 2017.
  46. 46.Long, M., Cao, Z., Wang, J., and Jordan, M. I. Conditional adversarial domain adaptation. In NeurIPS, 2018.
  47. 47.Mathur, A., Kawsar, F., Berthouze, N., and Lane, N. D. Libri-adapt: a new speech dataset for unsupervised domain adaptation. In ICASSP, 2020.
  48. 48.Mitsuzumi, Y., Irie, G., Ikami, D., and Shibata, T. Generalized domain adaptation. In CVPR, 2021.
  49. 49.Morerio, P., Volpi, R., Ragonesi, R., and Murino, V. Generative pseudo-label refinement for unsupervised domain adaptation. In WACV, 2020.
  50. 50.Murez, Z., Kolouri, S., Kriegman, D., Ramamoorthi, R., and Kim, K. Image to image translation for domain adaptation. In CVPR, 2018.
  51. 51.Na, J., Jung, H., Chang, H. J., and Hwang, W. FixBi: Bridging domain spaces for unsupervised domain adaptation. In CVPR, 2021.
  52. 52.Nguyen, V.-A., Nguyen, T., Le, T., Tran, Q. H., and Phung, D. STEM: An approach to multi-source domain adaptation with guarantees. In ICCV, 2021.
  53. 53.Nguyen-Meidine, L. T., Belal, A., Kiran, M., Dolz, J., Blais-Morin, L.-A., and Granger, E. Unsupervised multi-target domain adaptation through knowledge distillation. In WACV, 2021.
  54. 54.Park, G. Y. and Lee, S. W. Information-theoretic regularization for multi-source domain adaptation. In ICCV, 2021.
  55. 55.Peng, X., Usman, B., Kaushik, N., Hoffman, J., Wang, D., and Saenko, K. VisDA: The visual domain adaptation challenge. In CVPRW, 2018.
  56. 56.Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B. Moment matching for multi-source domain adaptation. In ICCV, 2019.
  57. 57.Prabhu, V., Khare, S., Kartik, D., and Hoffman, J. SENTRY: Selective entropy optimization via committee consistency for unsupervised domain adaptation. In ICCV, 2021.
  58. 58.Qiu, Z., Zhang, Y., Lin, H., Niu, S., Liu, Y., Du, Q., and Tan, M. Source-free domain adaptation via avatar prototype generation and adaptation. In IJCAI, 2021.
  59. 59.Rangwani, H., Jain, A., Aithal, S. K., and Babu, R. V. S3VAADA: Submodular subset selection for virtual adversarial active domain adaptation. In ICCV, 2021.
  60. 60.Rangwani, H., Aithal, S. K., Mishra, M., Jain, A., and Babu, R. V. A closer look at smoothness in domain adversarial training. In ICML, 2022.
  61. 61.Richter, S. R., Vineet, V., Roth, S., and Koltun, V. Playing for data: Ground truth from computer games. In ECCV, 2016.
  62. 62.Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A. M. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In CVPR, 2016.
  63. 63.Roy, S., Krivosheev, E., Zhong, Z., Sebe, N., and Ricci, E. Curriculum graph co-teaching for multi-target domain adaptation. In CVPR, 2021.
  64. 64.Russo, P., Carlucci, F. M., Tommasi, T., and Caputo, B. From source to target and back: symmetric bi-directional adaptive GAN. In CVPR, 2018.
  65. 65.Saenko, K., Kulis, B., Fritz, M., and Darrell, T. Adapting visual category models to new domains. In ECCV, 2010.
  66. 66.Saito, K., Watanabe, K., Ushiku, Y., and Harada, T. Maximum classifier discrepancy for unsupervised domain adaptation. In CVPR, 2018.
  67. 67.Salimans, T. and Kingma, D. P. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In NeurIPS, 2016.
  68. 68.Scalbert, M., Vakalopoulou, M., and Couzini’e-Devy, F. Multi-source domain adaptation via supervised contrastive learning and confident consistency regularization. In BMVC, 2021.
  69. 69.Sivaprasad, P. T. and Fleuret, F. Uncertainty reduction for model adaptation in semantic segmentation. In CVPR, 2021.
  70. 70.Soria, X., Riba, E., and Sappa, A. Dense extreme inception network: Towards a robust cnn model for edge detection. In WACV, 2020.
  71. 71.Stephenson, C., Padhy, S., Ganesh, A., Hui, Y., Tang, H., and Chung, S. On the geometry of generalization and memorization in deep neural networks. In ICLR, 2021.
  72. 72.Tang, H., Chen, K., and Jia, K. Unsupervised domain adaptation via structurally regularized deep clustering. In CVPR, 2020.
  73. 73.Tian, J., Zhang, J., Li, W., and Xu, D. VDM-DA: Virtual domain modeling for source data-free domain adaptation. IEEE Transactions on Circuits and Systems for Video Technology, 2021.
  74. 74.Ulyanov, D., Vedaldi, A., and Lempitsky, V. Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis. In CVPR, 2017.
  75. 75.Venkat, N., Kundu, J. N., Singh, D. K., Revanur, A., and Babu, R. V. Your classifier can secretly suffice multi-source domain adaptation. In NeurIPS, 2020.
  76. 76.Venkateswara, H., Eusebio, J., Chakraborty, S., and Panchanathan, S. Deep hashing network for unsupervised domain adaptation. In CVPR, 2017.
  77. 77.Wei, G., Lan, C., Zeng, W., and Chen, Z. MetaAlign: Coordinating domain alignment and classification for unsupervised domain adaptation. In CVPR, 2021.
  78. 78.Wen, J., Greiner, R., and Schuurmans, D. Domain aggregation networks for multi-source domain adaptation. In ICML, 2020.
  79. 79.Wrenninge, M. and Unger, J. Synscapes: A photorealistic synthetic dataset for street scene parsing, 2018.
  80. 80.Xia, H., Zhao, H., and Ding, Z. Adaptive adversarial network for source-free domain adaptation. In ICCV, 2021.
  81. 81.Xu, R., Chen, Z., Zuo, W., Yan, J., and Lin, L. Deep cocktail network: Multi-source unsupervised domain adaptation with category shift. In CVPR, 2018.
  82. 82.Xu, Y., Kan, M., Shan, S., and Chen, X. Mutual learning of joint and separate domain alignments for multi-source domain adaptation. In WACV, 2022.
  83. 83.Yang, J., Zou, H., Zhou, Y., Zeng, Z., and Xie, L. Mind the discriminability: Asymmetric adversarial domain adaptation. In ECCV, 2020a.
  84. 84.Yang, S., van de Weijer, J., Herranz, L., Jui, S., et al. Exploiting the intrinsic neighborhood structure for source-free domain adaptation. In NeurIPS, 2021a.
  85. 85.Yang, S., Wang, Y., van de Weijer, J., Herranz, L., and Jui, S. Generalized source-free domain adaptation. In ICCV, 2021b.
  86. 86.Yang, X., Deng, C., Liu, T., and Tao, D. Heterogeneous graph attention network for unsupervised multiple-target domain adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020b.
  87. 87.Yang, Y. and Soatto, S. FDA: Fourier domain adaptation for semantic segmentation. In CVPR, 2020.
  88. 88.Ye, M., Zhang, J., Ouyang, J., and Yuan, D. Source data-free unsupervised domain adaptation for semantic segmentation. In ACMMM, 2021.
  89. 89.Yue, Z., Sun, Q., Hua, X.-S., and Zhang, H. Transporting causal mechanisms for unsupervised domain adaptation. In ICCV, 2021.
  90. 90.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  91. 91.Zhang, P., Zhang, B., Zhang, T., Chen, D., Wang, Y., and Wen, F. Prototypical pseudo label denoising and target structure learning for domain adaptive semantic segmentation. In CVPR, 2021.
  92. 92.Zhao, H., Zhang, S., Wu, G., Moura, J. M., Costeira, J. P., and Gordon, G. J. Adversarial multiple source domain adaptation. In NeurIPS, 2018.
  93. 93.Zhu, Y., Zhuang, F., and Wang, D. Aligning domain-specific distribution and classifier for cross-domain classification from multiple sources. In AAAI, 2019.

Citation

MLA
Kundu, J. N., et al. “Balancing Discriminability and Transferability for Source-Free Domain Adaptation”. International Conference on Machine Learning, vol. 162, 2022, pp. 11710–28, https://proceedings.mlr.press/v162/kundu22a.html.
APA
Kundu, J. N., Kulkarni, A. R., Bhambri, S., Mehta, D., Kulkarni, S. A., Jampani, V., & Radhakrishnan, V. B. (2022). Balancing Discriminability and Transferability for Source-Free Domain Adaptation. International Conference on Machine Learning, 162, 11710–11728. https://proceedings.mlr.press/v162/kundu22a.html
Chicago
Kundu, J. N., A. R. Kulkarni, S. Bhambri, et al. 2022. “Balancing Discriminability and Transferability for Source-Free Domain Adaptation”. International Conference on Machine Learning 162: 11710–28. https://proceedings.mlr.press/v162/kundu22a.html.
Harvard
Kundu, J.N. et al. (2022) “Balancing Discriminability and Transferability for Source-Free Domain Adaptation”, International Conference on Machine Learning. PMLR, pp. 11710–11728. Available at: https://proceedings.mlr.press/v162/kundu22a.html.
Vancouver
1. Kundu JN, Kulkarni AR, Bhambri S, Mehta D, Kulkarni SA, Jampani V, Radhakrishnan VB (2022) Balancing Discriminability and Transferability for Source-Free Domain Adaptation. In: International Conference on Machine Learning. PMLR, pp 11710–11728

BibTeX

@InProceedings{pmlr-v162-kundu22a,
  title = 	 {Balancing Discriminability and Transferability for Source-Free Domain Adaptation},
  author =       {Kundu, Jogendra Nath and Kulkarni, Akshay R and Bhambri, Suvaansh and Mehta, Deepesh and Kulkarni, Shreyas Anand and Jampani, Varun and Radhakrishnan, Venkatesh Babu},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {11710--11728},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/kundu22a/kundu22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/kundu22a.html},
  abstract = 	 {Conventional domain adaptation (DA) techniques aim to improve domain transferability by learning domain-invariant representations; while concurrently preserving the task-discriminability knowledge gathered from the labeled source data. However, the requirement of simultaneous access to labeled source and unlabeled target renders them unsuitable for the challenging source-free DA setting. The trivial solution of realizing an effective original to generic domain mapping improves transferability but degrades task discriminability. Upon analyzing the hurdles from both theoretical and empirical standpoints, we derive novel insights to show that a mixup between original and corresponding translated generic samples enhances the discriminability-transferability trade-off while duly respecting the privacy-oriented source-free setting. A simple but effective realization of the proposed insights on top of the existing source-free DA approaches yields state-of-the-art performance with faster convergence. Beyond single-source, we also outperform multi-source prior-arts across both classification and semantic segmentation benchmarks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/