Unsupervised Visual Domain Adaptation Using Subspace Alignment

Basura FernandoAmaury HabrardM. SebbanT. Tuytelaars

article2013ICCV1,381 citations

Proposes an unsupervised domain adaptation method that directly aligns source and target eigenvector subspaces through a fast, closed-form linear transformation without needing intermediate projections or parameter tuning.

Listen

Machine learning models in computer vision frequently underperform when deployed in real-world environments because the distribution of test data differs from the training data, a challenge known as dataset bias or domain shift. Traditional methods to fix this require expensive manual labeling of new data or rely on computationally complex algorithms that create intermediate representations. The article evaluates a novel, lightweight method called subspace alignment for unsupervised domain adaptation, which adapts models trained on labeled source images to unlabeled target images without needing target labels.

The evaluated approach uses principal component analysis to capture the core basis vectors of both source and target datasets, and then directly aligns the source coordinate system with the target coordinate system using a simple closed-form transformation matrix. The performance of this method was tested across benchmark visual datasets—including Office, Caltech, ImageNet, LabelMe, and PASCAL-VOC—using both nearest-neighbor and support vector machine classifiers, and was compared against leading domain adaptation methods and non-adapted baselines.

The findings show that subspace alignment consistently outperforms existing state-of-the-art techniques across multiple benchmarks. In classification tests across the Office and Caltech domains, the method achieved superior accuracy in 9 of 12 transfer tasks using a nearest-neighbor classifier and in 11 of 12 tasks using a support vector machine. On cross-dataset evaluations with ImageNet, LabelMe, and Caltech-256, it delivered an average nearest-neighbor accuracy of 45.0%, surpassing the closest alternative baseline of 37.9%. When training on ImageNet to classify PASCAL-VOC images, the method improved the mean average precision by approximately 27% relative to the leading geodesic flow kernel method and by 34% relative to no adaptation. In addition, the alignment effectively reduced theoretical domain divergence measures significantly more than competing approaches.

These results demonstrate that aligning source and target subspaces directly is faster, more robust, and more accurate than generating costly intermediate representations. Organizations deploying visual recognition models can achieve higher predictive accuracy across changing environments without the cost and delay of collecting new manual annotations. The underlying closed-form mathematical solution eliminates the need to tune complex regularization parameters, making implementation practical and scalable.

Decision-makers should consider adopting direct subspace alignment pipelines for visual recognition tasks where data collection conditions vary, such as transitioning between different camera types or web-scraped images. Further development should explore applying this method to large-scale image retrieval systems and real-time, on-the-fly model adaptation. Confidence in these findings is high for standard image classification settings, though performance remains bounded by the quality of initial feature extraction and the assumption of underlying linear relationships between the domain subspaces.

  • Paper: Adapting Visual Category Models to New Domains, Kate Saenko et al. (2010). Introduces the standard visual domain adaptation problem and the multi-domain Office benchmark that forms the empirical backbone of the subspace alignment evaluation.
  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). Provides the foundational learning-theoretic bounds and domain divergence definitions that motivate aligning distribution subspaces to reduce cross-domain target error.
  • Paper: Unbiased look at dataset bias, Antonio Torralba et al. (2011). Demonstrates the dataset bias effect across popular vision benchmarks like ImageNet and PASCAL VOC, framing the specific cross-dataset degradation addressed by subspace alignment.
  • Paper: Correcting Sample Selection Bias by Unlabeled Data, Jiayuan Huang et al. (2006). Establishes non-parametric sample and distribution matching using unlabeled target data, providing essential background on correcting covariate shifts without labels.
  • Paper: Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification, John Blitzer et al. (2007). Pioneers the strategy of mapping domain-specific representations into a shared, invariant subspace via unlabeled data.
  • Paper: Boosting for transfer learning, Wenyuan Dai et al. (2007). Presents early foundational instance-adaptation theory and methodology for transferring predictive power across differing source and target distributions.
Cover for Unsupervised Visual Domain Adaptation Using Subspace Alignment

Abstract

In this paper, we introduce a new domain adaptation (DA) algorithm where the source and target domains are represented by subspaces described by eigenvectors. In this context, our method seeks a domain adaptation solution by learning a mapping function which aligns the source subspace with the target one. We show that the solution of the corresponding optimization problem can be obtained in a simple closed form, leading to an extremely fast algorithm. We use a theoretical result to tune the unique hyperparameter corresponding to the size of the subspaces. We run our method on various datasets and show that, despite its intrinsic simplicity, it outperforms state of the art DA methods.

Table of Contents

  • Abstract
  • 1. Introduction
  • 2. Related work
  • 3. DA using unsupervised subspace alignment
  • 3.1. Subspace generation
  • 3.2. Domain adaptation with subspace alignment
  • 3.3. Consistency theorem on 𝑆𝑖𝑚 ( y S , y T )
  • 3.4. Divergence between source and target domains
  • 4. Experiments
  • 4.1. DA datasets and data preparation
  • 4.2. Experimental setup
  • 4.3. Selecting the optimal dimensionality
  • 4.4. Evaluating DA with divergence measures
  • 4.5. Classification Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Subspace Alignment Formulation and Closed-Form Solution

    model/method

    Let source data SS and target data TT be zz-normalized DD-dimensional feature vectors. Principal Component Analysis (PCA) is applied separately to SS and TT to obtain orthonormal basis matrices XS,XT∈RD×dX_S, X_T \in \mathbb{R}^{D \times d} consisting of the dd leading eigenvectors (such that XS′XS=IdX_S' X_S = I_d and XT′XT=IdX_T' X_T = I_d, where IdI_d is the d×dd \times d identity matrix and (⋅)′(\cdot)' denotes the matrix transpose).

    Subspace alignment optimizes a linear transformation matrix M∈Rd×dM \in \mathbb{R}^{d \times d} that maps the source subspace basis XSX_S directly to the target subspace basis XTX_T by minimizing the Frobenius norm Bregman divergence:

    M∗=arg⁡min⁡M∥XSM−XT∥F2M^* = \arg\min_M \|X_S M - X_T\|_F^2

    Because the Frobenius norm is invariant under orthonormal transformations and XS′XS=IdX_S' X_S = I_d, this objective simplifies as:

    ∥XSM−XT∥F2=∥XS′XSM−XS′XT∥F2=∥M−XS′XT∥F2\|X_S M - X_T\|_F^2 = \|X_S' X_S M - X_S' X_T\|_F^2 = \|M - X_S' X_T\|_F^2

    The optimal alignment transformation matrix has the exact closed-form solution:

    M∗=XS′XTM^* = X_S' X_T

    The target-aligned source coordinate system is given by:

    Xa=XSM∗=XSXS′XTX_a = X_S M^* = X_S X_S' X_T

    When source and target domains are identical, XS=XTX_S = X_T and M∗M^* reduces to the identity matrix IdI_d.

  2. Knowl 2 — Subspace Alignment Domain Adaptation Algorithm

    algorithm

    The Subspace Alignment algorithm adapts source domain data to target domain data by aligning their leading dd-dimensional PCA subspaces in closed form, then projects source data into the aligned subspace and target data into the target subspace before invoking a supervised classifier.

    Input: Source data S∈RnS×DS \in \mathbb{R}^{n_S \times D}, Target data T∈RnT×DT \in \mathbb{R}^{n_T \times D}, Source labels LSL_S, Subspace dimension dd
    Output: Predicted target labels LTL_T
    XS←PCA(S,d)X_S \leftarrow \text{PCA}(S, d)
    XT←PCA(T,d)X_T \leftarrow \text{PCA}(T, d)
    Xa←XSXS′XTX_a \leftarrow X_S X_S' X_T
    Sa←SXaS_a \leftarrow S X_a
    TT←TXTT_T \leftarrow T X_T
    LT←Classifier(Sa,TT,LS)L_T \leftarrow \text{Classifier}(S_a, T_T, L_S)
    return LTL_T

    Sa∈RnS×dS_a \in \mathbb{R}^{n_S \times d} represents the transformed source training instances and TT∈RnT×dT_T \in \mathbb{R}^{n_T \times d} represents the target test instances in the dd-dimensional subspace. The classification model (such as a Support Vector Machine or 1-Nearest Neighbor classifier) is trained on labeled pairs (Sa,LS)(S_a, L_S) and evaluated directly on TTT_T.

  3. Knowl 3 — Bilinear Cross-Domain Similarity Function

    equation

    Given a source vector yS∈R1×Dy_S \in \mathbb{R}^{1 \times D} and a target vector yT∈R1×Dy_T \in \mathbb{R}^{1 \times D}, their similarity is computed by projecting each vector into its respective domain subspace (XSX_S and XTX_T) and aligning them with the optimal transformation matrix M∗=XS′XTM^* = X_S' X_T:

    Sim(yS,yT)=(ySXSM∗)(yTXT)′=ySXSXS′XTXT′yT′=ySAyT′Sim(y_S, y_T) = (y_S X_S M^*) (y_T X_T)' = y_S X_S X_S' X_T X_T' y_T' = y_S A y_T'

    where the metric operator A∈RD×DA \in \mathbb{R}^{D \times D} is defined as:

    A=XSXS′XTXT′A = X_S X_S' X_T X_T'

    The matrix AA acts as a generalized dot-product metric (not necessarily positive semi-definite) that weights original feature dimensions according to the alignment between the source and target eigenvectors. Sim(yS,yT)Sim(y_S, y_T) is used directly for cross-domain nearest-neighbor classification.

  4. Knowl 4 — Concentration Bound for the Subspace Alignment Operator

    theoretical result

    Let XSdX_S^d and XTdX_T^d be the dd-dimensional projection operators of the true expected source and target covariance matrices with strictly ordered leading eigenvalues λ1S>⋯>λdS>λd+1S≥0\lambda_1^S > \dots > \lambda_d^S > \lambda_{d+1}^S \ge 0 and λ1T>⋯>λdT>λd+1T≥0\lambda_1^T > \dots > \lambda_d^T > \lambda_{d+1}^T \ge 0. Let XSndX_{S n}^d and XTndX_{T n}^d be the empirical projection operators constructed from sample sizes nSn_S and nTn_T, respectively, where all data vectors satisfy ∥x∥≤B\|x\| \le B. Let M=(XSd)′XTdM = (X_S^d)' X_T^d and Mn=(XSnd)′XTndM_n = (X_{S n}^d)' X_{T n}^d.

    For any confidence parameter δ∈(0,1)\delta \in (0, 1), with probability at least 1−δ1 - \delta:

    ∥XSdM(XTd)′−XSndMn(XTnd)′∥≤8d3/2B(1+ln⁡(2/δ)2)(1nS(λdS−λd+1S)+1nT(λdT−λd+1T))\|X_S^d M (X_T^d)' - X_{S n}^d M_n (X_{T n}^d)'\| \le 8 d^{3/2} B \left( 1 + \sqrt{\frac{\ln(2/\delta)}{2}} \right) \left( \frac{1}{\sqrt{n_S}(\lambda_d^S - \lambda_{d+1}^S)} + \frac{1}{\sqrt{n_T}(\lambda_d^T - \lambda_{d+1}^T)} \right)

  5. Knowl 5 — Subspace Dimensionality Selection via Eigenvalue Gap and Source Cross-Validation

    model/method

    To automatically select the subspace dimension dd while preventing overfitting, a theoretical bound on the projection operator deviation is used to define an upper limit dmaxd_{max}. Letting nmin=min⁡(nS,nT)n_{min} = \min(n_S, n_T) and (λdmin−λd+1min)=min⁡(λdS−λd+1S,λdT−λd+1T)(\lambda_{d}^{min} - \lambda_{d+1}^{min}) = \min(\lambda_d^S - \lambda_{d+1}^S, \lambda_d^T - \lambda_{d+1}^T), for a specified maximum deviation tolerance γ>0\gamma > 0 and confidence level δ>0\delta > 0, dmaxd_{max} is chosen as the largest dimension satisfying:

    (λdmaxmin−λdmax+1min)≥(1+ln⁡(2/δ)2)(16dmax3/2Bγnmin)(\lambda_{d_{max}}^{min} - \lambda_{d_{max}+1}^{min}) \ge \left( 1 + \sqrt{\frac{\ln(2/\delta)}{2}} \right) \left( \frac{16 d_{max}^{3/2} B}{\gamma \sqrt{n_{min}}} \right)

    For any d≤dmaxd \le d_{max}, the empirical alignment operator remains within deviation γ\gamma from its expected value. The optimal subspace dimension d∗∈{1,…,dmax}d^* \in \{1, \dots, d_{max}\} is subsequently selected by minimizing classification error on the labeled source dataset via 2-fold cross-validation.

  6. Knowl 6 — Target Density Around Source Discrepancy Metric

    definition

    Target Density Around Source (TDASTDAS) is an empirical divergence metric designed to measure local distribution alignment for nearest-neighbor classifiers under covariate shift and probabilistic Lipschitzness assumptions:

    TDAS=1nS∑yS∈S∣{yT∈T∣Sim(yS,yT)≥ϵ}∣TDAS = \frac{1}{n_S} \sum_{y_S \in S} |\{y_T \in T \mid Sim(y_S, y_T) \ge \epsilon\}|

    where nSn_S is the number of source instances, SS is the source sample set, TT is the target sample set, ϵ>0\epsilon > 0 is a neighborhood similarity threshold, and Sim(yS,yT)=ySAyT′Sim(y_S, y_T) = y_S A y_T' is the learned cross-domain similarity. TDASTDAS measures the average number of target instances situated in an ϵ\epsilon-neighborhood around each source point. Larger TDASTDAS values reflect closer distribution proximity and indicate superior expected performance for local classifiers.

  7. Knowl 7 — Unsupervised Domain Adaptation Accuracy on Office-Caltech10 Benchmark

    data/table

    Target classification accuracies (%) across all 12 domain transfer pairs of the Office and Caltech10 datasets (Amazon: A, Caltech10: C, DSLR: D, Webcam: W; 800-bin SURF bag-of-words features), averaged over 20 random train/test splits in the unsupervised setting.

    Method C→\rightarrowA D→\rightarrowA W→\rightarrowA A→\rightarrowC D→\rightarrowC W→\rightarrowC A→\rightarrowD C→\rightarrowD W→\rightarrowD A→\rightarrowW C→\rightarrowW D→\rightarrowW
    1-Nearest-Neighbor (NN) Classifier
    NA 21.5 26.9 20.8 22.8 24.8 16.4 22.4 21.7 40.5 23.3 20.0 53.0
    Baseline 1 38.0 29.8 35.5 30.9 29.6 31.3 34.6 37.4 71.8 35.1 33.5 74.0
    Baseline 2 40.5 33.0 38.0 33.3 31.2 31.9 34.7 36.4 72.9 36.8 34.4 78.4
    GFS 36.9 32.0 27.5 35.3 29.4 21.7 30.7 32.6 54.3 31.0 30.6 66.0
    GFK 36.9 32.5 31.1 35.6 29.8 27.2 35.2 35.2 70.6 34.4 33.7 74.9
    OUR (SA) 39.0 38.0 37.4 35.3 32.4 32.3 37.6 39.6 80.3 38.6 36.8 83.6
    Support Vector Machine (SVM) Classifier
    Baseline 1 44.3 36.8 32.9 36.8 29.6 24.9 36.1 38.9 73.6 42.5 34.6 75.4
    Baseline 2 44.5 38.6 34.2 37.3 31.6 28.4 32.5 35.3 73.6 37.3 34.2 80.5
    GFK 44.8 37.9 37.1 38.3 31.4 29.1 37.9 36.1 74.6 39.8 34.9 79.1
    OUR (SA) 46.1 42.0 39.3 39.9 35.0 31.8 38.8 39.4 77.9 39.6 38.9 82.3

    Subspace Alignment (OUR) outperforms Geodesic Flow Sampling (GFS), Geodesic Flow Kernel (GFK), PCA Source projection (Baseline 1), PCA Target projection (Baseline 2), and No Adaptation (NA) in 9 of 12 tasks with 1-NN and in 11 of 12 tasks with SVM.

  8. Knowl 8 — Unsupervised Domain Adaptation Accuracy on ImageNet, LabelMe, and Caltech-256

    data/table

    Target recognition accuracy (%) on 6 adaptation pairs across ImageNet (I), LabelMe (L), and Caltech-256 (C) using 5 shared classes (bird, car, chair, dog, person; 7719 total images; 2048-dimensional LLC spatial pyramid representations) over 20 random trials.

    Method L→\rightarrowC L→\rightarrowI C→\rightarrowL C→\rightarrowI I→\rightarrowL I→\rightarrowC AVG
    1-Nearest-Neighbor (NN) Classifier
    NA 46.0 38.4 29.5 31.3 36.9 45.5 37.9
    Baseline 1 24.2 27.2 46.9 41.8 35.7 33.8 34.9
    Baseline 2 24.6 27.4 47.0 42.0 35.6 33.8 35.0
    GFK 24.2 26.8 44.9 40.7 35.1 33.8 34.3
    OUR (SA) 49.1 41.2 47.0 39.1 39.4 54.5 45.0
    Support Vector Machine (SVM) Classifier
    NA 49.6 40.8 36.0 45.6 41.3 58.9 45.4
    Baseline 1 50.5 42.0 39.1 48.3 44.0 59.7 47.3
    Baseline 2 48.7 41.9 39.2 48.4 43.6 58.0 46.6
    GFK 52.3 43.5 39.6 49.0 45.3 61.8 48.6
    OUR (SA) 52.9 43.9 43.8 50.9 46.3 62.8 50.1

    Subspace Alignment achieves the highest average accuracy across all pairs for both 1-NN (45.0% vs. 34.3% for GFK) and SVM (50.1% vs. 48.6% for GFK), especially recovering accuracy on tasks with LabelMe as the source domain (L→CL \rightarrow C and L→IL \rightarrow I) where intermediate subspace methods degraded performance.

  9. Knowl 9 — Domain Discrepancy Reduction under Subspace Alignment

    data/table

    Average distribution discrepancy metrics across the 12 domain adaptation tasks from the Office dataset.

    Metric NA Baseline 1 Baseline 2 GFK OUR (SA)
    TDAS 1.25 3.34 2.74 2.84 4.26
    HΔHH\Delta H 98.1 99.0 99.0 74.3 53.2

    TDASTDAS measures target instance density around source points (higher values indicate closer alignment for local nearest-neighbor models). HΔHH\Delta H divergence is estimated via the error of a linear SVM trained to discriminate source from target instances (where 50.0% denotes indistinguishable distributions). Subspace Alignment achieves the highest TDASTDAS (4.26) and the lowest HΔHH\Delta H divergence (53.2%), demonstrating superior alignment in both local and global feature structures.

  10. Knowl 10 — Cross-Dataset Visual Adaptation from ImageNet to PASCAL-VOC-2007

    empirical result

    On cross-dataset adaptation from ImageNet to all 20 classes of the PASCAL-VOC-2007 test set using 2048-dimensional LLC spatial pyramid features with an SVM classifier:

    • In unsupervised domain adaptation, Subspace Alignment outperforms No Adaptation (NA) and Geodesic Flow Kernel (GFK) on every individual object category, achieving a 27% relative increase in mean average precision (mAP) over GFK (which itself improved mAP by 7% over NA).
    • In semi-supervised domain adaptation (with 50 labeled target samples per class added to training), Subspace Alignment improves mAP by 13% relative over GFK and 46% relative over No Adaptation, yielding an overall 10% relative mAP improvement over the unsupervised Subspace Alignment baseline.

Coverage note — Intermediate Lemma 1 is omitted as its result is fully subsumed by Theorem 2. Semi-supervised per-task detailed tables are in the supplementary material and are summarized in empirical results.

References

  1. 1.S. Ben-David, J. Blitzer, K. Crammer, and F. Pereira. Analysis of representations for domain adaptation. In NIPS. 2007.
  2. 2.S. Ben-David, S. Shalev-Shwartz, and R. Urner. Domain adaptation–can quantity compensate for quality? In International Symposium on Artificial Intelligence and Mathematics, 2012.
  3. 3.J. Blitzer, D. Foster, and S. Kakade. Domain adaptation with coupled subspaces. In Conference on Artificial Intelligence and Statistics, 2011.
  4. 4.J. Blitzer, R. McDonald, and F. Pereira. Domain adaptation with structural correspondence learning. In Conference on Empirical Methods in Natural Language Processing, 2006.
  5. 5.S.-F. Chang. Robust visual domain adaptation with low-rank reconstruction. In CVPR, 2012.
  6. 6.B. Chen, W. Lam, I. Tsang, and T.-L. Wong. Extracting discriminative concepts for domain adaptation in text mining. In ACM SIGKDD, 2009.
  7. 7.B. Gong, Y. Shi, F. Sha, and K. Grauman. Geodesic flow kernel for unsupervised domain adaptation. In CVPR, 2012.
  8. 8.R. Gopalan, R. Li, and R. Chellappa. Domain adaptation for object recognition: An unsupervised approach. In ICCV, 2011.
  9. 9.A. Khosla, T. Zhou, T. Malisiewicz, A. A. Efros, and A. Torralba. Undoing the damage of dataset bias. In ECCV, 2012.
  10. 10.B. Kulis, K. Saenko, and T. Darrell. What you saw is not what you get: Domain adaptation using asymmetric kernel transforms. In CVPR, 2011.
  11. 11.A. Margolis. A literature review of domain adaptation with unlabeled data. Technical report, University of Washington, 2011.
  12. 12.S. J. Pan, J. T. Kwok, and Q. Yang. Transfer learning via dimensionality reduction. In AAAI, 2008.
  13. 13.S. J. Pan, I. W. Tsang, J. T. Kwok, and Q. Yang. Domain adaptation via transfer component analysis. In IJCAI, 2009.
  14. 14.K. Saenko, B. Kulis, M. Fritz, and T. Darrell. Adapting visual category models to new domains. In ECCV, 2010.
  15. 15.A. Torralba and A. Efros. Unbiased look at dataset bias. In CVPR, 2011.
  16. 16.C. Wang and S. Mahadevan. Manifold alignment without correspondence. In IJCAI, 2009.
  17. 17.C. Wang and S. Mahadevan. Heterogeneous domain adaptation using manifold alignment. In IJCAI, 2011.
  18. 18.D. Zhai, B. Li, H. Chang, S. Shan, X. Chen, and W. Gao. Manifold alignment via corresponding projections. In BMVC, 2010.
  19. 19.L. Zwald and G. Blanchard. On the convergence of eigenspaces in kernel principal components analysis. In NIPS, 2005.

Citation

MLA
Fernando, B., et al. “Unsupervised Visual Domain Adaptation Using Subspace Alignment”. 2013 IEEE International Conference on Computer Vision, 2013, pp. 2960–67, https://doi.org/10.1109/ICCV.2013.368.
APA
Fernando, B., Habrard, A., Sebban, M., & Tuytelaars, T. (2013). Unsupervised Visual Domain Adaptation Using Subspace Alignment. 2013 IEEE International Conference on Computer Vision, 2960–2967. https://doi.org/10.1109/ICCV.2013.368
Chicago
Fernando, B., A. Habrard, M. Sebban, and T. Tuytelaars. 2013. “Unsupervised Visual Domain Adaptation Using Subspace Alignment”. 2013 IEEE International Conference on Computer Vision, 2960–67. https://doi.org/10.1109/ICCV.2013.368.
Harvard
Fernando, B. et al. (2013) “Unsupervised Visual Domain Adaptation Using Subspace Alignment”, 2013 IEEE International Conference on Computer Vision. IEEE, pp. 2960–2967. Available at: https://doi.org/10.1109/ICCV.2013.368.
Vancouver
1. Fernando B, Habrard A, Sebban M, Tuytelaars T (2013) Unsupervised Visual Domain Adaptation Using Subspace Alignment. In: 2013 IEEE International Conference on Computer Vision. IEEE, pp 2960–2967

BibTeX

@inproceedings{Fernando_2013, title={Unsupervised Visual Domain Adaptation Using Subspace Alignment}, url={http://dx.doi.org/10.1109/ICCV.2013.368}, DOI={10.1109/iccv.2013.368}, booktitle={2013 IEEE International Conference on Computer Vision}, publisher={IEEE}, author={Fernando, Basura and Habrard, Amaury and Sebban, Marc and Tuytelaars, Tinne}, year={2013}, month=Dec, pages={2960–2967} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE