Optimal Transport for Domain Adaptation

Nicolas CourtyRémi FlamaryDevis TuiaAlain Rakotomamonjy

article2014TPAMI1,383 citations

Proposes a regularized optimal transport framework that aligns probability distributions between distinct domains while preserving class structure, providing a principled geometric solution for visual domain adaptation tasks.

Listen

Modern data analytics and machine learning applications frequently encounter dataset drift, where predictive models trained on data from one acquisition source perform poorly when deployed on target data gathered under different conditions, such as altered lighting, background noise, or distinct sensor properties. This issue is especially acute in unsupervised domain adaptation, where target datasets lack label annotations entirely. The article evaluates a framework designed to align source and target data distributions by formulating domain adaptation as an optimal transport problem, which mathematically determines a minimal-effort mapping to transform labeled source data into the target space.

The evaluated approach introduces a regularized discrete optimal transport model that calculates non-linear, sample-specific transformations between empirical probability distributions. To prevent unrelated source samples from collapsing onto identical target points, the authors incorporate domain-specific regularizers: a convex group-lasso penalty that groups samples by class and a graph Laplacian regularizer that preserves local neighborhood structures. The framework also adapts to semi-supervised settings when limited target labels are present and scales to large datasets through a generalized conditional gradient optimization algorithm that leverages fast matrix scaling operations. The methodology was tested across synthetic non-linear benchmarks and standard real-world computer vision datasets, including digit recognition (USPS and MNIST), facial recognition across poses (PIE), and multi-domain object recognition (Caltech-Office).

The experimental results demonstrate that regularized optimal transport consistently outperforms existing domain adaptation baselines. On synthetic non-linear data, optimal transport maintained lower classification error rates across escalating transformation severities, achieving nearly flawless adaptation below forty-degree rotations where standard methods deteriorated. In real-world visual benchmarks, the proposed group-lasso and Laplacian regularizers improved average accuracy compared to unadapted baselines and standard subspace alignment algorithms, achieving up to 63.90% mean accuracy on digits and 47.70% on standard object recognition features. Furthermore, when applied to deep learning representations, optimal transport boosted accuracy by over 20 percentage points on specific domain shifts, such as between webcam and digital SLR categories. In semi-supervised evaluations with three target labels per class, embedding labels directly into the transport plan increased average accuracy to 55.6%, outperforming post-transport labeling approaches.

These findings indicate that aligning datasets via optimal transport allows organizations to transfer existing labeled assets to new domains without retraining entire deep architectures or manually re-annotating target datasets. Unlike conventional linear subspace alignments, this approach accommodates complex, non-linear physical distortions while preserving class integrity, directly mitigating the risks, operational costs, and development timelines associated with data re-labeling. Practitioners seeking to deploy models across shifting environments should consider implementing optimal transport alignment as an intermediate calibration layer, selecting group-lasso regularization when class coherence is paramount and incorporating any available target labels directly into the transportation cost matrix.

While the method provides exact recovery guarantees for affine transformations and robust empirical performance across complex shifts, users must note key boundary conditions. Performance degrades when extreme transformations cause severe ambiguity or when severe label proportion imbalances exist between source and target sets. Additionally, selecting optimization hyperparameters without target validation labels remains challenging in fully unsupervised environments. Overall confidence in the approach is high for visual adaptation tasks, and future work should focus on developing physics-informed regularizers and extending the formulation to multi-domain adaptation pipelines.

arXiv: 1507.00504
  • Paper: Analysis of Representations for Domain Adaptation, Shai Ben-David et al. (2006). This seminal work establishes the theoretical bounds on generalization error and distribution discrepancy across domains that underpin distribution-matching strategies in domain adaptation.
  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). It provides the foundational statistical learning theory and divergence measures bounding target error that justify learning aligned representations between source and target domains.
  • Paper: Transfer Feature Learning with Joint Distribution Adaptation, Mingsheng Long et al. (2013). It demonstrates how aligning both marginal and conditional distributions via learned representations enables effective unsupervised visual domain adaptation.
  • Paper: Adapting Visual Category Models to New Domains, Kate Saenko et al. (2010). It introduces the visual benchmark datasets and transformation-mapping formulation widely adopted to evaluate feature alignment under domain shifts.
  • Paper: A Survey on Transfer Learning, Sinno Jialin Pan et al. (2010). This survey formalizes the taxonomy of transfer learning and feature-representation alignment across divergent source and target data distributions.
Cover for Optimal Transport for Domain Adaptation

Abstract

Domain adaptation from one data space (or domain) to another is one of the most challenging tasks of modern data analytics. If the adaptation is done correctly, models built on a specific data space become more robust when confronted to data depicting the same semantic concepts (the classes), but observed by another observation system with its own specificities. Among the many strategies proposed to adapt a domain to another, finding a common representation has shown excellent properties: by finding a common representation for both domains, a single classifier can be effective in both and use labelled samples from the source domain to predict the unlabelled samples of the target domain. In this paper, we propose a regularized unsupervised optimal transportation model to perform the alignment of the representations in the source and target domains. We learn a transportation plan matching both PDFs, which constrains labelled samples in the source domain to remain close during transport. This way, we exploit at the same time the few labeled information in the source and the unlabelled distributions observed in both domains. Experiments in toy and challenging real visual adaptation examples show the interest of the method, that consistently outperforms state of the art approaches.

Table of Contents

  • I Introduction
  • I-A Related works
  • II Optimal transport and application to domain adaptation
  • II-A Problem and theoretical motivations
  • II-B Domain adaptation as a transportation problem
  • III Regularized discrete optimal transport
  • III-A Discrete optimal transport
  • III-B Regularized optimal transport
  • III-C OT-based mapping of the samples
  • III-D Discussing optimal transport for domain adaptation
  • IV Class-regularization for domain adaptation
  • IV-A Regularizing the transport with class labels
  • IV-A1 Regularization with group-sparsity
  • IV-A2 Laplacian regularization
  • IV-B Regularizing for semi-supervised domain adaptation
  • V Generalized conditional gradient for solving regularized OT problems
  • VI Numerical experiments
  • VI-A Two moons: simulated problem with controllable complexity
  • VI-B Visual adaptation datasets
  • VI-B1 Datasets
  • VI-B2 Experimental setup
  • VI-B3 Results on unsupervised domain adaptation
  • VI-B4 Semi-supervised domain adaptation
  • VII Conclusion
  • References

Knowls

  1. Knowl 1 — Optimal Transport Framework for Unsupervised Domain Adaptation

    model/method

    In unsupervised domain adaptation, a learner receives a labeled source dataset Xs={xis}i=1Ns⊂Ωs⊆RdX_s = \{\mathbf{x}_i^s\}_{i=1}^{N_s} \subset \Omega_s \subseteq \mathbb{R}^d with class labels Ys={yis}i=1Ns⊂CY_s = \{y_i^s\}_{i=1}^{N_s} \subset \mathcal{C} and an unlabeled target dataset Xt={xjt}j=1Nt⊂Ωt⊆RdX_t = \{\mathbf{x}_j^t\}_{j=1}^{N_t} \subset \Omega_t \subseteq \mathbb{R}^d. The domain discrepancy is modeled as an unknown, possibly non-linear input space mapping T:Ωs→ΩtT: \Omega_s \to \Omega_t that preserves conditional label distributions:

    Ps(y∣xs)=Pt(y∣T(xs)),implying ft(T(x))=fs(x)P_s(y \mid \mathbf{x}^s) = P_t(y \mid T(\mathbf{x}^s)), \quad \text{implying } f_t(T(\mathbf{x})) = f_s(\mathbf{x})

    where fsf_s and ftf_t are the Bayes decision functions in the source and target domains, respectively.

    The source and target empirical distributions are represented as discrete measures μs=∑i=1Nspisδxis\mu_s = \sum_{i=1}^{N_s} p_i^s \delta_{\mathbf{x}_i^s} and μt=∑j=1Ntpjtδxjt\mu_t = \sum_{j=1}^{N_t} p_j^t \delta_{\mathbf{x}_j^t}, where δx\delta_{\mathbf{x}} is the Dirac delta function, and ps∈R+Ns\mathbf{p}^s \in \mathbb{R}_+^{N_s}, pt∈R+Nt\mathbf{p}^t \in \mathbb{R}_+^{N_t} are probability vectors on the simplex (typically uniform weights pis=1/Nsp_i^s = 1/N_s and pjt=1/Ntp_j^t = 1/N_t). The discrete Kantorovich optimal transport formulation estimates a transportation plan γ0∈B\boldsymbol{\gamma}_0 \in \mathcal{B} by solving:

    γ0=arg⁡min⁡γ∈B⟨γ,C⟩F=arg⁡min⁡γ∈B∑i=1Ns∑j=1Ntγ(i,j)C(i,j)\boldsymbol{\gamma}_0 = \arg\min_{\boldsymbol{\gamma} \in \mathcal{B}} \langle \boldsymbol{\gamma}, \mathbf{C} \rangle_F = \arg\min_{\boldsymbol{\gamma} \in \mathcal{B}} \sum_{i=1}^{N_s} \sum_{j=1}^{N_t} \gamma(i, j) C(i, j)

    subject to the constraint set:

    B={γ∈(R+)Ns×Nt  |  γ1Nt=ps,  γ⊤1Ns=pt}\mathcal{B} = \left\{ \boldsymbol{\gamma} \in (\mathbb{R}_+)^{N_s \times N_t} \;\middle|\; \boldsymbol{\gamma} \mathbf{1}_{N_t} = \mathbf{p}^s, \; \boldsymbol{\gamma}^\top \mathbf{1}_{N_s} = \mathbf{p}^t \right\}

    where ⟨⋅,⋅⟩F\langle \cdot, \cdot \rangle_F is the Frobenius dot product, 1n\mathbf{1}_n is an nn-dimensional vector of ones, and C∈R+Ns×Nt\mathbf{C} \in \mathbb{R}_+^{N_s \times N_t} is the cost matrix with entries C(i,j)=c(xis,xjt)=∥xis−xjt∥22C(i, j) = c(\mathbf{x}_i^s, \mathbf{x}_j^t) = \|\mathbf{x}_i^s - \mathbf{x}_j^t\|_2^2 (squared Euclidean distance).

  2. Knowl 2 — Barycentric Mapping for Source Sample Transport

    model/method

    Once an optimal transport coupling γ0∈R+Ns×Nt\boldsymbol{\gamma}_0 \in \mathbb{R}_+^{N_s \times N_t} between source data Xs∈RNs×dX_s \in \mathbb{R}^{N_s \times d} and target data Xt∈RNt×dX_t \in \mathbb{R}^{N_t \times d} is obtained, each source sample xis\mathbf{x}_i^s is mapped into the target domain using a barycentric mapping x^is∈Rd\hat{\mathbf{x}}_i^s \in \mathbb{R}^d defined by the Fréchet mean:

    x^is=arg⁡min⁡x∈Rd∑j=1Ntγ0(i,j)c(x,xjt)\hat{\mathbf{x}}_i^s = \arg\min_{\mathbf{x} \in \mathbb{R}^d} \sum_{j=1}^{N_t} \gamma_0(i, j) c(\mathbf{x}, \mathbf{x}_j^t)

    For the squared ℓ2\ell_2 Euclidean cost c(x,x′)=∥x−x′∥22c(\mathbf{x}, \mathbf{x}') = \|\mathbf{x} - \mathbf{x}'\|_2^2, this barycenter is a weighted average residing in the convex hull of target points:

    x^is=1∑j=1Ntγ0(i,j)∑j=1Ntγ0(i,j)xjt\hat{\mathbf{x}}_i^s = \frac{1}{\sum_{j=1}^{N_t} \gamma_0(i, j)} \sum_{j=1}^{N_t} \gamma_0(i, j) \mathbf{x}_j^t

    In matrix notation, the transported source dataset X^s∈RNs×d\hat{X}_s \in \mathbb{R}^{N_s \times d} is given by:

    X^s=diag⁡(γ01Nt)−1γ0Xt\hat{X}_s = \operatorname{diag}(\boldsymbol{\gamma}_0 \mathbf{1}_{N_t})^{-1} \boldsymbol{\gamma}_0 X_t

    where diag⁡(v)\operatorname{diag}(\mathbf{v}) denotes the diagonal matrix formed from vector v\mathbf{v}, and 1Nt\mathbf{1}_{N_t} is an NtN_t-dimensional all-ones vector. When marginal distributions are uniform (pis=1/Nsp_i^s = 1/N_s and pjt=1/Ntp_j^t = 1/N_t), the forward mapping of source points and the inverse mapping of target points X^t∈RNt×d\hat{X}_t \in \mathbb{R}^{N_t \times d} simplify to linear transformations:

    X^s=Nsγ0XtandX^t=Ntγ0⊤Xs\hat{X}_s = N_s \boldsymbol{\gamma}_0 X_t \quad \text{and} \quad \hat{X}_t = N_t \boldsymbol{\gamma}_0^\top X_s

    A classifier trained on the transported labeled source samples (X^s,Ys)(\hat{X}_s, Y_s) is then applied directly to classify target domain points XtX_t.

  3. Knowl 3 — Exact Affine Transformation Recovery in Discrete Optimal Transport

    theoretical result

    Let μs=∑i=1n1nδxis\mu_s = \sum_{i=1}^n \frac{1}{n} \delta_{\mathbf{x}_i^s} and μt=∑j=1n1nδxjt\mu_t = \sum_{j=1}^n \frac{1}{n} \delta_{\mathbf{x}_j^t} be two discrete probability distributions in Rd\mathbb{R}^d with nn Dirac masses and uniform weights.

    Assume that:

    1. All source locations are distinct: xis≠xjs\mathbf{x}_i^s \neq \mathbf{x}_j^s for all i≠ji \neq j.
    2. The target samples are obtained via an affine transformation of source samples: xit=Axis+b\mathbf{x}_i^t = \mathbf{A} \mathbf{x}_i^s + \mathbf{b} for all i∈{1,…,n}i \in \{1, \dots, n\}.
    3. The translation vector satisfies b∈Rd\mathbf{b} \in \mathbb{R}^d, and A∈S+d\mathbf{A} \in \mathcal{S}_+^d is a strictly positive definite matrix.
    4. The ground cost is the squared Euclidean distance c(xs,xt)=∥xs−xt∥22c(\mathbf{x}^s, \mathbf{x}^t) = \|\mathbf{x}^s - \mathbf{x}^t\|_2^2.

    Under these conditions, the discrete Kantorovich optimal transport problem yields a solution T0T_0 such that:

    T0(xis)=Axis+b=xit,∀i∈{1,…,n}T_0(\mathbf{x}_i^s) = \mathbf{A}\mathbf{x}_i^s + \mathbf{b} = \mathbf{x}_i^t, \quad \forall i \in \{1, \dots, n\}

    This guarantees that optimal transport exactly recovers the underlying affine transformation on discrete samples and perfectly preserves source class labels during transport.

  4. Knowl 4 — Group-Sparse Regularization for Class-Preserving Optimal Transport

    model/method

    To prevent source samples with different class labels from transporting probability mass to the same target sample, class label information YsY_s from the source domain is incorporated into the optimal transport formulation via a convex group-lasso (ℓ1−ℓ2\ell_1-\ell_2) penalty on the columns of the coupling matrix γ\boldsymbol{\gamma}.

    The regularized optimal transport objective is defined as:

    min⁡γ∈B⟨γ,C⟩F+λΩs(γ)+ηΩc(γ)\min_{\boldsymbol{\gamma} \in \mathcal{B}} \langle \boldsymbol{\gamma}, \mathbf{C} \rangle_F + \lambda \Omega_s(\boldsymbol{\gamma}) + \eta \Omega_c(\boldsymbol{\gamma})

    where B={γ∈(R+)Ns×Nt∣γ1Nt=ps,γ⊤1Ns=pt}\mathcal{B} = \{ \boldsymbol{\gamma} \in (\mathbb{R}_+)^{N_s \times N_t} \mid \boldsymbol{\gamma} \mathbf{1}_{N_t} = \mathbf{p}^s, \boldsymbol{\gamma}^\top \mathbf{1}_{N_s} = \mathbf{p}^t \}, C\mathbf{C} is the cost matrix, λ≥0\lambda \ge 0 weights the entropic regularizer Ωs(γ)=∑i,jγ(i,j)log⁡γ(i,j)\Omega_s(\boldsymbol{\gamma}) = \sum_{i,j} \gamma(i,j) \log \gamma(i,j), and η≥0\eta \ge 0 weights the class regularizer Ωc(γ)\Omega_c(\boldsymbol{\gamma}).

    The group-sparse regularizer Ωc(γ)\Omega_c(\boldsymbol{\gamma}) is formulated as:

    Ωc(γ)=∑j=1Nt∑c∈C∥γ(Ic,j)∥2\Omega_c(\boldsymbol{\gamma}) = \sum_{j=1}^{N_t} \sum_{c \in \mathcal{C}} \|\boldsymbol{\gamma}(\mathcal{I}_c, j)\|_2

    where C\mathcal{C} is the set of source class labels, Ic⊂{1,…,Ns}\mathcal{I}_c \subset \{1, \dots, N_s\} denotes the row indices of source samples belonging to class cc, and γ(Ic,j)\boldsymbol{\gamma}(\mathcal{I}_c, j) is the vector of coupling coefficients from source samples of class cc to the jj-th target sample xjt\mathbf{x}_j^t. This penalty induces group sparsity over classes in each target sample's mass allocation.

  5. Knowl 5 — Graph Laplacian Regularization for Optimal Transport Domain Adaptation

    model/method

    To preserve local geometric structures and neighborhood relationships during transport, a graph Laplacian regularization term penalizes distortions between neighboring source samples after transportation.

    Given a symmetric similarity matrix Ss∈R+Ns×Ns\mathbf{S}_s \in \mathbb{R}_+^{N_s \times N_s} where Ss(i,j)≥0S_s(i, j) \ge 0 measures similarity between source samples xis\mathbf{x}_i^s and xjs\mathbf{x}_j^s with cross-class connections pruned (Ss(i,j)=0S_s(i, j) = 0 if yis≠yjsy_i^s \neq y_j^s), the Laplacian regularizer on the transported source samples x^is\hat{\mathbf{x}}_i^s is:

    Ωc(γ)=1Ns2∑i=1Ns∑j=1NsSs(i,j)∥x^is−x^js∥22\Omega_c(\boldsymbol{\gamma}) = \frac{1}{N_s^2} \sum_{i=1}^{N_s} \sum_{j=1}^{N_s} S_s(i, j) \|\hat{\mathbf{x}}_i^s - \hat{\mathbf{x}}_j^s\|_2^2

    When source marginals are uniform, this simplifies to the quadratic form:

    Ωc(γ)=Tr⁡(Xt⊤γ⊤LsγXt)\Omega_c(\boldsymbol{\gamma}) = \operatorname{Tr}(X_t^\top \boldsymbol{\gamma}^\top \mathbf{L}_s \boldsymbol{\gamma} X_t)

    where Ls=diag⁡(Ss1Ns)−Ss\mathbf{L}_s = \operatorname{diag}(\mathbf{S}_s \mathbf{1}_{N_s}) - \mathbf{S}_s is the graph Laplacian matrix of Ss\mathbf{S}_s, and Tr⁡(⋅)\operatorname{Tr}(\cdot) denotes the trace operator.

    When a target similarity matrix St∈R+Nt×Nt\mathbf{S}_t \in \mathbb{R}_+^{N_t \times N_t} with Laplacian Lt=diag⁡(St1Nt)−St\mathbf{L}_t = \operatorname{diag}(\mathbf{S}_t \mathbf{1}_{N_t}) - \mathbf{S}_t is also constructed (e.g., using an 8-nearest-neighbors graph), symmetric Laplacian regularization is defined with weighting parameter α∈[0,1]\alpha \in [0, 1]:

    Ωc(γ)=(1−α)Tr⁡(Xt⊤γ⊤LsγXt)+αTr⁡(Xs⊤γLtγ⊤Xs)\Omega_c(\boldsymbol{\gamma}) = (1 - \alpha) \operatorname{Tr}(X_t^\top \boldsymbol{\gamma}^\top \mathbf{L}_s \boldsymbol{\gamma} X_t) + \alpha \operatorname{Tr}(X_s^\top \boldsymbol{\gamma} \mathbf{L}_t \boldsymbol{\gamma}^\top X_s)

  6. Knowl 6 — Semi-Supervised Domain Adaptation via Optimal Transport Cost Masking

    model/method

    When a small number of labeled target samples are available alongside labeled source data, label consistency across domains is enforced without additional hyperparameters by incorporating an exact matching penalty Ωsemi(γ)\Omega_{\text{semi}}(\boldsymbol{\gamma}) into the optimal transport objective:

    Ωsemi(γ)=⟨γ,M⟩F=∑i=1Ns∑j=1Ntγ(i,j)M(i,j)\Omega_{\text{semi}}(\boldsymbol{\gamma}) = \langle \boldsymbol{\gamma}, \mathbf{M} \rangle_F = \sum_{i=1}^{N_s} \sum_{j=1}^{N_t} \gamma(i, j) M(i, j)

    where M∈(R∪{+∞})Ns×Nt\mathbf{M} \in (\mathbb{R} \cup \{+\infty\})^{N_s \times N_t} is a masking matrix defined by:

    M(i,j)={0if yis=yjt or if target sample j is unlabeled+∞if yis≠yjtM(i, j) = \begin{cases} 0 & \text{if } y_i^s = y_j^t \text{ or if target sample } j \text{ is unlabeled} \\ +\infty & \text{if } y_i^s \neq y_j^t \end{cases}

    Adding Ωsemi(γ)\Omega_{\text{semi}}(\boldsymbol{\gamma}) to the optimal transport objective is equivalent to modifying the original ground cost matrix C\mathbf{C} to C~=C+M\widetilde{\mathbf{C}} = \mathbf{C} + \mathbf{M}. This assigns infinite cost to transport between source and target points known to belong to different classes, guaranteeing that γ(i,j)=0\gamma(i, j) = 0 whenever yis≠yjty_i^s \neq y_j^t.

  7. Knowl 7 — Generalized Conditional Gradient Algorithm for Regularized Optimal Transport

    algorithm

    The regularized optimal transport problem min⁡γ∈B⟨γ,C⟩F+λΩs(γ)+ηΩc(γ)\min_{\boldsymbol{\gamma} \in \mathcal{B}} \langle \boldsymbol{\gamma}, \mathbf{C} \rangle_F + \lambda \Omega_s(\boldsymbol{\gamma}) + \eta \Omega_c(\boldsymbol{\gamma}) is solved by casting it as a composite optimization problem:

    min⁡γ∈Bf(γ)+g(γ)\min_{\boldsymbol{\gamma} \in \mathcal{B}} f(\boldsymbol{\gamma}) + g(\boldsymbol{\gamma})

    where f(γ)=⟨γ,C⟩F+ηΩc(γ)f(\boldsymbol{\gamma}) = \langle \boldsymbol{\gamma}, \mathbf{C} \rangle_F + \eta \Omega_c(\boldsymbol{\gamma}) is differentiable (for Laplacian regularization and for group lasso when initialized with positive entries), and g(γ)=λΩs(γ)=λ∑i,jγ(i,j)log⁡γ(i,j)g(\boldsymbol{\gamma}) = \lambda \Omega_s(\boldsymbol{\gamma}) = \lambda \sum_{i,j} \gamma(i,j) \log \gamma(i,j) is convex. The Generalized Conditional Gradient (GCG) algorithm linearizes f(γ)f(\boldsymbol{\gamma}) while keeping g(γ)g(\boldsymbol{\gamma}) intact, solving each subproblem using the Sinkhorn-Knopp scaling algorithm:

    Input: Cost matrix C∈R+Ns×Nt\mathbf{C} \in \mathbb{R}_+^{N_s \times N_t}, parameters λ>0\lambda > 0, η≥0\eta \ge 0, convergence tolerance ϵ\epsilon
    Output: Optimal transportation plan γ∗∈B\boldsymbol{\gamma}^* \in \mathcal{B}
    Initialize k←0k \leftarrow 0, γ0∈B\boldsymbol{\gamma}^0 \in \mathcal{B} with γ0(i,j)>0\gamma^0(i, j) > 0 for all i,ji, j
    repeat
        Compute gradient G←∇f(γk)=C+η∇Ωc(γk)\mathbf{G} \leftarrow \nabla f(\boldsymbol{\gamma}^k) = \mathbf{C} + \eta \nabla \Omega_c(\boldsymbol{\gamma}^k)
        Solve direction via Sinkhorn-Knopp: γ∗←arg⁡min⁡γ∈B⟨γ,G⟩F+λΩs(γ)\boldsymbol{\gamma}^* \leftarrow \arg\min_{\boldsymbol{\gamma} \in \mathcal{B}} \langle \boldsymbol{\gamma}, \mathbf{G} \rangle_F + \lambda \Omega_s(\boldsymbol{\gamma})
        Set search direction Δγ←γ∗−γk\Delta \boldsymbol{\gamma} \leftarrow \boldsymbol{\gamma}^* - \boldsymbol{\gamma}^k
        Find step size: αk←arg⁡min⁡0≤α≤1f(γk+αΔγ)+g(γk+αΔγ)\alpha^k \leftarrow \arg\min_{0 \le \alpha \le 1} f(\boldsymbol{\gamma}^k + \alpha \Delta \boldsymbol{\gamma}) + g(\boldsymbol{\gamma}^k + \alpha \Delta \boldsymbol{\gamma})
        Update iterate: γk+1←γk+αkΔγ\boldsymbol{\gamma}^{k+1} \leftarrow \boldsymbol{\gamma}^k + \alpha^k \Delta \boldsymbol{\gamma}
        k←k+1k \leftarrow k + 1
    until convergence of objective or ∥γk−γk−1∥<ϵ\|\boldsymbol{\gamma}^k - \boldsymbol{\gamma}^{k-1}\| < \epsilon
    return γk\boldsymbol{\gamma}^k
  8. Knowl 8 — Domain Adaptation Performance on the Two Moons Synthetic Benchmark

    empirical result

    The domain adaptation capability of optimal transport methods was evaluated on the two moons benchmark across increasing rotation angles from 10∘10^\circ to 90∘90^\circ. The source domain contains 300 samples (150 per moon class in R2\mathbb{R}^2), and target domains contain 300 samples generated by rotation. Generalization error was measured on 1,000 independent target test points using an SVM with a Gaussian kernel (5-fold cross-validated).

    Target rotation angle 10∘10^\circ 20∘20^\circ 30∘30^\circ 40∘40^\circ 50∘50^\circ 70∘70^\circ 90∘90^\circ
    SVM (no adaptation) 0.000 0.104 0.240 0.312 0.400 0.764 0.828
    DASVM 0.000 0.000 0.259 0.284 0.334 0.747 0.820
    PBDA 0.000 0.094 0.103 0.225 0.412 0.626 0.687
    OT-exact 0.000 0.028 0.065 0.109 0.206 0.394 0.507
    OT-IT 0.000 0.007 0.054 0.102 0.221 0.398 0.508
    OT-GL 0.000 0.000 0.000 0.013 0.196 0.378 0.508
    OT-Laplace 0.000 0.000 0.004 0.062 0.201 0.402 0.524

    The table lists the mean classification error rate over 10 independent realizations. Class-regularized optimal transport (OT-GL and OT-Laplace) maintained near-zero error rates up to 40∘40^\circ rotation, outperforming unregularized optimal transport (OT-exact), entropic optimal transport (OT-IT), and baselines (DASVM, PBDA). At 90∘90^\circ, all optimal transport variants achieved an error rate of approximately 0.508 due to symmetry (a −90∘-90^\circ rotation yields an identical geometric distribution with inverted class labels).

  9. Knowl 9 — Visual Domain Adaptation Benchmark on SURF Features

    empirical result

    Unsupervised domain adaptation methods were evaluated across visual recognition tasks using 800-bin SURF histogram descriptors: Digits (USPS denoted U, MNIST denoted M; 10 classes, 256-dimensional inputs), Face recognition (CMU PIE dataset with poses P1, P2, P3, P4; 68 classes, 1024-dimensional inputs), and Object recognition (Office-Caltech dataset across Caltech C, Amazon A, Webcam W, DSLR D; 10 classes, 800-dimensional SURF histograms). A 1-Nearest Neighbor (1NN) classifier was trained on the transported source points and tested on target points (averaged over 10 runs with 20 source samples per class, 8 for DSLR; target partitioned 50/50 for hyperparameter validation and testing).

    Domains 1NN PCA GFK TSL JDA OT-exact OT-IT OT-Laplace OT-LpL1 OT-GL
    U→\toM 39.00 37.83 44.16 40.66 54.52 50.67 53.66 57.42 60.15 57.85
    M→\toU 58.33 48.05 60.96 53.79 60.09 49.26 64.73 64.72 68.07 69.96
    mean 48.66 42.94 52.56 47.22 57.30 49.96 59.20 61.07 64.11 63.90
    PIE mean (12 pairs) 26.22 34.55 26.15 36.10 56.69 50.47 54.89 56.10 55.45 55.88
    C→\toA 20.54 35.17 35.29 45.25 40.73 30.54 37.75 38.96 48.21 44.17
    C→\toW 18.94 28.48 31.72 37.35 33.44 23.77 31.32 31.13 38.61 38.94
    C→\toD 19.62 33.75 35.62 39.25 39.75 26.62 34.50 36.88 39.62 44.50
    A→\toC 22.25 32.78 32.87 38.46 33.99 29.43 31.65 33.12 35.99 34.57
    A→\toW 23.51 29.34 32.05 35.70 36.03 25.56 30.40 30.33 35.63 37.02
    A→\toD 20.38 26.88 30.12 32.62 32.62 25.50 27.88 27.75 36.38 38.88
    W→\toC 19.29 26.95 27.75 29.02 31.81 25.87 31.63 31.37 33.44 35.98
    W→\toA 23.19 28.92 33.35 34.94 31.48 27.40 37.79 37.17 37.33 39.35
    W→\toD 53.62 79.75 79.25 80.50 84.25 76.50 80.00 80.62 81.38 84.00
    D→\toC 23.97 29.72 29.50 31.03 29.84 27.30 29.88 31.10 31.65 32.38
    D→\toA 27.10 30.67 32.98 36.67 32.85 29.08 32.77 33.06 37.06 37.17
    D→\toW 51.26 71.79 69.67 77.48 80.00 65.70 72.52 76.16 74.97 81.06
    mean 28.47 37.98 39.21 42.97 44.34 36.69 42.30 43.20 46.42 47.70

    The table displays classification accuracy in %. OT-GL achieved the highest mean accuracy on Digits (63.90%) and Office-Caltech Objects (47.70%), outperforming subspace methods (PCA, GFK, TSL, JDA) and unregularized OT (OT-exact: 36.69%). On PIE faces, JDA achieved the top mean score (56.69%) via its iterative EM pseudo-label refinement across 68 classes, followed closely by OT-Laplace (56.10%) and OT-GL (55.88%).

  10. Knowl 10 — Optimal Transport Domain Adaptation with Deep DeCAF Features

    empirical result

    Optimal transport domain adaptation was evaluated using 4,096-dimensional deep activation features extracted from the 6th (DeCAF-6) and 7th (DeCAF-7) fully connected layers of a convolutional network pre-trained on ImageNet and fine-tuned on the Office-Caltech dataset (10 classes; domains: Caltech C, Amazon A, Webcam W, DSLR D) with a 1NN classifier.

    Layer 6 (DeCAF-6) Layer 7 (DeCAF-7)
    Domains DeCAF JDA OT-IT OT-GL DeCAF JDA OT-IT OT-GL
    C→\toA 79.25 88.04 88.69 92.08 85.27 89.63 91.56 92.15
    C→\toW 48.61 79.60 75.17 84.17 65.23 79.80 82.19 83.84
    C→\toD 62.75 84.12 83.38 87.25 75.38 85.00 85.00 85.38
    A→\toC 64.66 81.28 81.65 85.51 72.80 82.59 84.22 87.16
    A→\toW 51.39 80.33 78.94 83.05 63.64 83.05 81.52 84.50
    A→\toD 60.38 86.25 85.88 85.00 75.25 85.50 86.62 85.25
    W→\toC 58.17 81.97 74.80 81.45 69.17 79.84 81.74 83.71
    W→\toA 61.15 90.19 80.96 90.62 72.96 90.94 88.31 91.98
    W→\toD 97.50 98.88 95.62 96.25 98.50 98.88 98.38 91.38
    D→\toC 52.13 81.13 77.71 84.11 65.23 81.21 82.02 84.93
    D→\toA 60.71 91.31 87.15 92.31 75.46 91.92 92.15 92.92
    D→\toW 85.70 97.48 93.77 96.29 92.25 97.02 96.62 94.17
    mean 65.20 86.72 83.64 88.18 75.93 87.11 87.53 88.11

    The table reports classification accuracy in %. OT-GL improved performance over raw DeCAF features by 22.98 percentage points on Layer 6 (from 65.20% to 88.18%) and by 12.18 percentage points on Layer 7 (from 75.93% to 88.11%), outperforming JDA (86.72% on Layer 6, 87.11% on Layer 7). Layer 7 features yielded comparable OT-GL performance to Layer 6 features (88.11% vs. 88.18%), demonstrating that optimal transport aligns representations effectively even without deeper network layer transformations.

  11. Knowl 11 — Semi-Supervised Visual Domain Adaptation Performance on SURF Features

    empirical result

    Semi-supervised domain adaptation was benchmarked on the Office-Caltech dataset using 800-bin SURF features with 3 labeled target samples per class. Two configurations were compared: "Unsupervised + labels", where optimal transport is learned without target labels and the 3 labeled target samples per class are only included during classifier training, and "Semi-supervised", where target labels constrain the optimal transportation plan via the infinite cost masking matrix M\mathbf{M}, compared against Max-Margin Domain Transfer (MMDT).

    Unsupervised + labels Semi-supervised
    Domains OT-IT OT-GL OT-IT OT-GL MMDT
    C→\toA 37.0±0.537.0 \pm 0.5 41.4±0.541.4 \pm 0.5 46.9±3.446.9 \pm 3.4 47.9±3.147.9 \pm 3.1 49.4±0.8\mathbf{49.4 \pm 0.8}
    C→\toW 28.5±0.728.5 \pm 0.7 37.4±1.137.4 \pm 1.1 64.8±3.064.8 \pm 3.0 65.0±3.1\mathbf{65.0 \pm 3.1} 63.8±1.163.8 \pm 1.1
    C→\toD 35.1±1.735.1 \pm 1.7 44.0±1.944.0 \pm 1.9 59.3±2.559.3 \pm 2.5 61.0±2.1\mathbf{61.0 \pm 2.1} 56.5±0.956.5 \pm 0.9
    A→\toC 32.3±0.132.3 \pm 0.1 36.7±0.236.7 \pm 0.2 36.0±1.336.0 \pm 1.3 37.1±1.1\mathbf{37.1 \pm 1.1} 36.4±0.836.4 \pm 0.8
    A→\toW 29.5±0.829.5 \pm 0.8 37.8±1.137.8 \pm 1.1 63.7±2.463.7 \pm 2.4 64.6±1.9\mathbf{64.6 \pm 1.9} 64.6±1.2\mathbf{64.6 \pm 1.2}
    A→\toD 36.9±1.536.9 \pm 1.5 46.2±2.046.2 \pm 2.0 57.6±2.557.6 \pm 2.5 59.1±2.3\mathbf{59.1 \pm 2.3} 56.7±1.356.7 \pm 1.3
    W→\toC 35.8±0.235.8 \pm 0.2 36.5±0.236.5 \pm 0.2 38.4±1.538.4 \pm 1.5 38.8±1.2\mathbf{38.8 \pm 1.2} 32.2±0.832.2 \pm 0.8
    W→\toA 39.6±0.339.6 \pm 0.3 41.9±0.441.9 \pm 0.4 47.2±2.547.2 \pm 2.5 47.3±2.547.3 \pm 2.5 47.7±0.9\mathbf{47.7 \pm 0.9}
    W→\toD 77.1±1.877.1 \pm 1.8 80.2±1.680.2 \pm 1.6 79.0±2.879.0 \pm 2.8 79.4±2.8\mathbf{79.4 \pm 2.8} 67.0±1.167.0 \pm 1.1
    D→\toC 32.7±0.332.7 \pm 0.3 34.7±0.334.7 \pm 0.3 35.5±2.135.5 \pm 2.1 36.8±1.5\mathbf{36.8 \pm 1.5} 34.1±1.534.1 \pm 1.5
    D→\toA 34.7±0.334.7 \pm 0.3 37.7±0.337.7 \pm 0.3 45.8±2.645.8 \pm 2.6 46.3±2.546.3 \pm 2.5 46.9±1.0\mathbf{46.9 \pm 1.0}
    D→\toW 81.9±0.681.9 \pm 0.6 84.5±0.484.5 \pm 0.4 83.9±1.483.9 \pm 1.4 84.0±1.5\mathbf{84.0 \pm 1.5} 74.1±0.874.1 \pm 0.8
    mean 41.8 46.6 54.8 55.6\mathbf{55.6} 52.5

    The table shows mean recognition accuracy ±\pm standard deviation in % across 10 realizations. Constraining the optimal transportation plan with target labels in the semi-supervised formulation increased mean accuracy from 46.6% to 55.6% for OT-GL, outperforming MMDT (52.5%).

Coverage note — The non-convex $\ell_p-\ell_1$ heuristic penalty from preliminary work (OT-LpL1) was omitted from standalone method knowls in favor of the primary convex $\ell_1-\ell_2$ group lasso formulation, though its experimental comparisons are fully included in the empirical tables.

References

  1. 1.R. K. Ahuja, T. L. Magnanti, and J. B. Orlin, Network Flows: Theory, Algorithms, and Applications. Upper Saddle River, NJ, USA: Prentice-Hall, Inc., 1993.
  2. 2.S. Ben-David, T. Luu, T. Lu, and D. Pál, “Impossibility theorems for domain adaptation.” in Artificial Intelligence and Statistics Conference (AISTATS), 2010, pp. 129–136.
  3. 3.J.-D. Benamou and Y. Brenier, “A computational fluid mechanics solution to the monge-kantorovich mass transfer problem,” Numerische Mathematik, vol. 84, no. 3, pp. 375–393, 2000.
  4. 4.D. P. Bertsekas, Nonlinear programming. Athena scientific Belmont, 1999.
  5. 5.N. Bonneel, J. Rabin, G. Peyré, and H. Pfister, “Sliced and radon Wasserstein barycenters of measures,” Journal of Mathematical Imaging and Vision, vol. 51, pp. 22–45, 2015.
  6. 6.N. Bonneel, M. van de Panne, S. Paris, and W. Heidrich, “Displacement interpolation using Lagrangian mass transport,” ACM Transaction on Graphics, vol. 30, no. 6, pp. 158:1–158:12, 2011.
  7. 7.K. Bredies, D. A. Lorenz, and P. Maass, “A generalized conditional gradient method and its connection to an iterative shrinkage method,” Computational Optimization and Applications, vol. 42, no. 2, pp. 173–193, 2009.
  8. 8.K. Bredies, D. Lorenz, and P. Maass, Equivalence of a generalized conditional gradient method and the method of surrogate functionals. Zentrum für Technomathematik, 2005.
  9. 9.L. Bruzzone and M. Marconcini, “Domain adaptation problems: A dasvm classification technique and a circular validation strategy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, no. 5, pp. 770–787, May 2010.
  10. 10.T. S. Caetano, T. Caelli, D. Schuurmans, and D. Barone, “Grapihcal models and point pattern matching,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, no. 10, pp. 1646–1663, 2006.
  11. 11.T. S. Caetano, J. J. McAuley, L. Cheng, Q. V. Le, and A. J. Smola, “Learning graph matching,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 31, no. 6, pp. 1048–1058, 2009.
  12. 12.G. Carlier, A. Oberman, and E. Oudet, “Numerical methods for matching for teams and Wasserstein barycenters,” Inria, Tech. Rep. hal-00987292, 2014.
  13. 13.M. Carreira-Perpinan and W. Wang, “LASS: A simple assignment model with laplacian smoothing,” in AAAI Conference on Artificial Intelligence, 2014.
  14. 14.N. Courty, R. Flamary, and D. Tuia, “Domain adaptation with regularized optimal transport,” in European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2014.
  15. 15.M. Cuturi, “Sinkhorn distances: Lightspeed computation of optimal transportation,” in Neural Information Processing Systems (NIPS), 2013, pp. 2292–2300.
  16. 16.M. Cuturi and D. Avis, “Ground metric learning,” Journal of Machine Learning Research, vol. 15, no. 1, pp. 533–564, Jan. 2014.
  17. 17.M. Cuturi and A. Doucet, “Fast computation of Wasserstein barycenters,” in International Conference on Machine Learning (ICML), 2014.
  18. 18.H. Daumé III, “Frustratingly easy domain adaptation,” in Ann. Meeting of the Assoc. Computational Linguistics, 2007.
  19. 19.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “DeCAF: a deep convolutional activation feature for generic visual recognition,” in International Conference on Machine Learning (ICML), 2014, pp. 647–655.
  20. 20.S. Ferradans, N. Papadakis, J. Rabin, G. Peyré, and J.-F. Aujol, “Regularized discrete optimal transport,” in Scale Space and Variational Methods in Computer Vision, SSVM, 2013, pp. 428–439.
  21. 21.W. Gangbo and R. J. McCann, “The geometry of optimal transportation,” Acta Mathematica, vol. 177, no. 2, pp. 113–161, 1996.
  22. 22.P. Germain, A. Habrard, F. Laviolette, and E. Morvant, “A PAC-Bayesian Approach for Domain Adaptation with Specialization to Linear Classifiers,” in International Conference on Machine Learning (ICML), Atlanta, USA, 2013, pp. 738–746.
  23. 23.B. Gong, Y. Shi, F. Sha, and K. Grauman, “Geodesic flow kernel for unsupervised domain adaptation.” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012, pp. 2066–2073.
  24. 24.R. Gopalan, R. Li, and R. Chellappa, “Domain adaptation for object recognition: An unsupervised approach,” in International Conference on Computer Vision (ICCV), 2011, pp. 999–1006.
  25. 25.G. Griffin, A. Holub, and P. Perona, “Caltech-256 Object Category Dataset,” California Institute of Technology, Tech. Rep. CNS-TR-2007-001, 2007.
  26. 26.J. Ham, D. Lee, and L. Saul, “Semisupervised alignment of manifolds,” in 10th International Workshop on Artificial Intelligence and Statistics, R. G. Cowell and Z. Ghahramani, Eds., 2005, pp. 120–127.
  27. 27.J. Hoffman, E. Rodner, J. Donahue, K. Saenko, and T. Darrell, “Efficient learning of domain invariant image representations,” in International Conference on Learning Representations (ICLR), 2013.
  28. 28.——, “Efficient learning of domain-invariant image representations,” in International Conference on Learning Representations (ICLR), 2013.
  29. 29.I.-H. Jhuo, D. Liu, D. T. Lee, and S.-F. Chang, “Robust visual domain adaptation with low-rank reconstruction,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012, pp. 2168–2175.
  30. 30.L. Kantorovich, “On the translocation of masses,” C.R. (Doklady) Acad. Sci. URSS (N.S.), vol. 37, pp. 199–201, 1942.
  31. 31.P. Knight, “The sinkhorn-knopp algorithm: Convergence and applications,” SIAM Journal on Matrix Analysis and Applications, vol. 30, no. 1, pp. 261–275, 2008.
  32. 32.B. Kulis, K. Saenko, and T. Darrell, “What you saw is not what you get: domain adaptation using asymmetric kernel transforms,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, 2011.
  33. 33.A. Kumar, H. Daumé III, and D. Jacobs, “Generalized multiview analysis: A discriminative latent space,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
  34. 34.M. Long, J. Wang, G. Ding, J. Sun, and P. Yu, “Transfer feature learning with joint distribution adaptation,” in International Conference on Computer Vision (ICCV), Dec 2013, pp. 2200–2207.
  35. 35.B. Luo and R. Hancock, “Structural graph matching using the em algorithm and singular value decomposition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 23, no. 10, pp. 1120–1136, 2001.
  36. 36.Y. Mansour, M. Mohri, and A. Rostamizadeh, “Domain adaptation: Learning bounds and algorithms,” in Conference on Learning Theory (COLT), 2009, pp. 19–30.
  37. 37.S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on Knowledge and Data Engineering, vol. 22, no. 10, pp. 1345–1359, 2010.
  38. 38.——, “Domain adaptation via transfer component analysis,” IEEE Transactions on Neural Networks, vol. 22, pp. 199–210, 2011.
  39. 39.V. M. Patel, R. Gopalan, R. Li, and R. Chellappa, “Visual domain adaptation: an overview of recent advances,” IEEE Signal Processing Magazine, vol. 32, no. 3, 2015.
  40. 40.J. Rabin, G. Peyré, J. Delon, and M. Bernot, “Wasserstein barycenter and its application to texture mixing,” in Scale Space and Variational Methods in Computer Vision, ser. Lecture Notes in Computer Science, 2012, vol. 6667, pp. 435–446.
  41. 41.Y. Rubner, C. Tomasi, and L. Guibas, “A metric for distributions with applications to image databases,” in International Conference on Computer Vision (ICCV), 1998, pp. 59–66.
  42. 42.K. Saenko, B. Kulis, M. Fritz, and T. Darrell, “Adapting visual category models to new domains,” in European Conference on Computer Vision (ECCV), ser. LNCS, 2010, pp. 213–226.
  43. 43.F. Santambrogio, “Optimal transport for applied mathematicians,” Birkäuser, NY, 2015.
  44. 44.S. Si, D. Tao, and B. Geng, “Bregman divergence-based regularization for transfer subspace learning,” IEEE Transactions on Knowledge and Data Engineering, vol. 22, no. 7, pp. 929–942, July 2010.
  45. 45.J. Solomon, R. Rustamov, G. Leonidas, and A. Butscher, “Wasserstein propagation for semi-supervised learning,” in International Conference on Machine Learning (ICML), 2014, pp. 306–314.
  46. 46.M. Sugiyama, S. Nakajima, H. Kashima, P. Buenau, and M. Kawanabe, “Direct importance estimation with model selection and its application to covariate shift adaptation,” in Neural Information Processing Systems (NIPS), 2008.
  47. 47.D. Tuia and G. Camps-Valls, “Kernel manifold alignment for domain adaptation,” PLoS One, vol. 11, no. 2, p. e0148655, 2016.
  48. 48.D. Tuia, R. Flamary, A. Rakotomamonjy, and N. Courty, “Multitemporal classification without new labels: a solution with optimal transport,” in 8th International Workshop on the Analysis of Multitemporal Remote Sensing Images, 2015.
  49. 49.C. Villani, Optimal transport: old and new, ser. Grundlehren der mathematischen Wissenschaften. Springer, 2009.
  50. 50.C. Wang, P. Krafft, and S. Mahadevan, “Manifold alignment,” in Manifold Learning: Theory and Applications, Y. Ma and Y. Fu, Eds. CRC Press, 2011.
  51. 51.C. Wang and S. Mahadevan, “Manifold alignment without correspondence,” in International Joint Conference on Artificial Intelligence (IJCAI), Pasadena, CA, 2009.
  52. 52.——, “Heterogeneous domain adaptation using manifold alignment,” in International Joint Conference on Artificial Intelligence (IJCAI). AAAI Press, 2011, pp. 1541–1546.
  53. 53.K. Zhang, V. W. Zheng, Q. Wang, J. T. Kwok, Q. Yang, and I. Marsic, “Covariate shift in Hilbert space: A solution via surrogate kernels,” in International Conference on Machine Learning (ICML), 2013.
  54. 54.J. Zheng, M.-Y. Liu, R. Chellappa, and P. Phillips, “A Grassmann manifold-based domain adaptation approach,” in International Conference on Pattern Recognition (ICPR), Nov 2012, pp. 2095–2099.

Citation

MLA
Courty, N., et al. “Optimal Transport for Domain Adaptation”. arXiv, 2015, http://arxiv.org/abs/1507.00504v2.
APA
Courty, N., Flamary, R., Tuia, D., & Rakotomamonjy, A. (2015). Optimal Transport for Domain Adaptation. arXiv. http://arxiv.org/abs/1507.00504v2
Chicago
Courty, N., R. Flamary, D. Tuia, and A. Rakotomamonjy. 2015. “Optimal Transport for Domain Adaptation”. arXiv. http://arxiv.org/abs/1507.00504v2.
Harvard
Courty, N. et al. (2015) “Optimal Transport for Domain Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1507.00504v2.
Vancouver
1. Courty N, Flamary R, Tuia D, Rakotomamonjy A (2015) Optimal Transport for Domain Adaptation. arXiv

BibTeX

@article{courty2015optimal,
  title = {Optimal Transport for Domain Adaptation},
  author = {Courty, Nicolas and Flamary, Rémi and Tuia, Devis and Rakotomamonjy, Alain},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1507.00504v2},
  eprint = {1507.00504}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF