Simultaneous Deep Transfer Across Domains and Tasks

Eric TzengJudy HoffmanTrevor DarrellKate Saenko

article2015ICCV1,417 citations

Introduces a deep learning architecture that simultaneously bridges domain and task gaps by pairing domain-invariance optimization with soft label distribution matching, enabling effective visual transfer with minimal labeled target data.

Listen

Deploying machine vision models into real-world operational environments often leads to steep performance drops caused by dataset bias and domain shift, such as differences in lighting, camera angles, or backgrounds. While traditional deep learning models require extensive labeled training data in every new environment to maintain accuracy, collecting hundreds of annotations per category in practical settings is costly and unrealistic. The article addresses this challenge by evaluating a new framework designed to adapt visual recognition models across operational domains and recognition tasks when target data is mostly unlabeled and labeled examples are scarce or missing entirely for several categories.

The main objective of the article is to demonstrate a deep convolutional neural network architecture that jointly optimizes for domain invariance and semantic task transfer, thereby enabling accurate image classification in a new target environment with minimal human supervision. To accomplish this, the authors introduced an architecture that combines standard classification loss with two novel mechanisms: a domain confusion loss that renders internal image representations statistically indistinguishable between source and target environments, and a soft label distribution loss that distills category relationships—such as the visual similarity between laptops and monitors—from the source domain to guide unannotated target classes. The approach was experimentally evaluated across benchmark image datasets, specifically the Office benchmark across three domains and a cross-dataset shift between ImageNet and Caltech-256.

The evaluation yielded several key findings. In semi-supervised settings where target domain annotations were available for only 15 out of 31 categories, the proposed joint method achieved an average classification accuracy of 66.4% on the 16 completely unannotated categories, outperforming the source-only baseline of 62.0% and achieving an approximate 13% relative improvement over prior domain adaptation methods on the most difficult domain shifts. In fully supervised settings with three labeled examples per class, the joint method achieved an average accuracy of 82.22%, surpassing standard joint fine-tuning (81.50%) and alternative domain adaptation baselines. On large-scale cross-dataset transfers, the method demonstrated substantial accuracy gains when very few target labels were present (1 to 5 examples per class), whereas standard domain classification models degraded to near-random performance (56% accuracy versus 99% on unadapted baselines), confirming that the learned internal representations achieved high domain invariance.

These findings indicate that organizations can deploy computer vision systems into new operating environments at substantially lower data collection and labeling costs without sacrificing recognition performance. The approach mitigates operational risks by ensuring that unannotated categories still benefit from structural knowledge learned in prior domains. For technical leaders and engineering teams, the article recommends adopting soft-label matching and domain confusion optimization as standard fine-tuning strategies when target training data is limited or partially annotated. When substantial labeled target data is already accessible, standard joint fine-tuning remains a viable alternative, though domain adaptation provides the greatest performance advantage under severe data constraints. While the results demonstrate robust gains on standard visual classification benchmarks using standard network backbones, organizations should validate the architecture on their specific domain distributions and complex operational imagery before full deployment.

arXiv: 1510.02192
  • Paper: Deep Domain Confusion: Maximizing for Domain Invariance, Eric Tzeng et al. (2014). This paper establishes the deep domain confusion framework using maximum mean discrepancy to optimize domain invariance, which directly precedes and motivates simultaneous domain and task transfer.
  • Paper: DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition, Jeff Donahue et al. (2013). This work demonstrates that deep convolutional features transfer effectively to downstream visual tasks but retain dataset bias, framing the core challenge the source paper aims to solve.
  • Paper: Adapting Visual Category Models to New Domains, Kate Saenko et al. (2010). This foundational paper introduces the standard Office domain adaptation benchmark and metric learning formulations that ground visual domain transfer.
  • Paper: Transfer Feature Learning with Joint Distribution Adaptation, Mingsheng Long et al. (2013). This study introduces joint distribution adaptation to align both marginal and conditional distributions across domains, providing key conceptual foundations for simultaneous transfer.
  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). This foundational theory provides formal generalization bounds and divergence measures between domains that justify minimizing distribution discrepancy during transfer.
  • Paper: Analysis of Representations for Domain Adaptation, Shai Ben-David et al. (2006). This paper develops the foundational theoretical bounds demonstrating that effective domain adaptation requires learning representations that minimize cross-domain divergence while preserving low source error.
Cover for Simultaneous Deep Transfer Across Domains and Tasks

Abstract

Recent reports suggest that a generic supervised deep CNN model trained on a large-scale dataset reduces, but does not remove, dataset bias. Fine-tuning deep models in a new domain can require a significant amount of labeled data, which for many applications is simply not available. We propose a new CNN architecture to exploit unlabeled and sparsely labeled target domain data. Our approach simultaneously optimizes for domain invariance to facilitate domain transfer and uses a soft label distribution matching loss to transfer information between tasks. Our proposed adaptation method offers empirical performance which exceeds previously published results on two standard benchmark visual domain adaptation tasks, evaluated across supervised and semi-supervised adaptation settings.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Joint CNN architecture for domain and task transfer
  • 3.1 Aligning domains via domain confusion
  • 3.2 Aligning source and target classes via soft labels
  • 4 Evaluation
  • 4.1 Adaptation on the Office dataset
  • 4.2 Adaptation between diverse domains
  • 5 Analysis
  • 5.1 Domain confusion enforces domain invariance
  • 5.2 Soft labels for task transfer
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Joint Objective for Simultaneous Deep Transfer Across Domains and Tasks

    model/method

    The joint deep transfer framework simultaneously aligns source and target domain feature representations and transfers semantic task structure from a fully labeled source domain to a sparsely labeled target domain. The model operates on labeled source data {xS,yS}\{x_S, y_S\} and target data {xT,yT}\{x_T, y_T\} (where labels yTy_T are provided for only a subset of instances and/or classes).

    The feature representation is parameterized by θrepr\theta_{\text{repr}} (corresponding to layers conv1–fc7 in an AlexNet/CaffeNet architecture) to produce features f(x;θrepr)f(x; \theta_{\text{repr}}), category classification layer fc8 parameterized by θC\theta_C, and an auxiliary domain classifier layer fcD parameterized by θD\theta_D.

    The joint objective function to be minimized is: L(xS,yS,xT,yT,θD;θrepr,θC)=LC(xS,yS,xT,yT;θrepr,θC)+λLconf(xS,xT,θD;θrepr)+νLsoft(xT,yT;θrepr,θC)\mathcal{L}(x_S, y_S, x_T, y_T, \theta_D; \theta_{\text{repr}}, \theta_C) = \mathcal{L}_C(x_S, y_S, x_T, y_T; \theta_{\text{repr}}, \theta_C) + \lambda \mathcal{L}_{\text{conf}}(x_S, x_T, \theta_D; \theta_{\text{repr}}) + \nu \mathcal{L}_{\text{soft}}(x_T, y_T; \theta_{\text{repr}}, \theta_C) where:

    • LC\mathcal{L}_C is the standard softmax cross-entropy classification loss on all available hard-labeled data across KK classes: LC(x,y;θrepr,θC)=−∑k=1KI[y=k]log⁡pk\mathcal{L}_C(x, y; \theta_{\text{repr}}, \theta_C) = -\sum_{k=1}^K \mathbb{I}[y = k] \log p_k, with p=softmax(θCTf(x;θrepr))p = \text{softmax}(\theta_C^T f(x; \theta_{\text{repr}})).
    • Lconf\mathcal{L}_{\text{conf}} is the domain confusion loss that penalizes representations from which domain identity can be reliably inferred.
    • Lsoft\mathcal{L}_{\text{soft}} is the soft label matching loss that transfers inter-class correlations from source to target.
    • λ≥0\lambda \ge 0 and ν≥0\nu \ge 0 are hyperparameters balancing the domain confusion and task transfer objectives (set empirically to λ=0.01\lambda = 0.01 and ν=0.1\nu = 0.1).
  2. Knowl 2 — Domain Invariance via Domain Classifier Confusion

    model/method

    To learn a domain-invariant representation f(x;θrepr)f(x; \theta_{\text{repr}}), an auxiliary domain classification layer θD\theta_D is introduced to predict the domain index d∈{1,…,D}d \in \{1, \dots, D\} (with D=2D=2 for source vs. target). The domain classifier is trained to minimize the standard multi-class cross-entropy loss: LD(xS,xT,θrepr;θD)=−∑d=1DI[yD=d]log⁡qd\mathcal{L}_D(x_S, x_T, \theta_{\text{repr}}; \theta_D) = -\sum_{d=1}^D \mathbb{I}[y_D = d] \log q_d where q=softmax(θDTf(x;θrepr))q = \text{softmax}(\theta_D^T f(x; \theta_{\text{repr}})) is the predicted probability distribution over domains.

    To enforce domain invariance in the representation, θrepr\theta_{\text{repr}} is optimized to maximize confusion between domains by minimizing the cross-entropy between qq and a uniform distribution over domain labels: Lconf(xS,xT,θD;θrepr)=−∑d=1D1Dlog⁡qd\mathcal{L}_{\text{conf}}(x_S, x_T, \theta_D; \theta_{\text{repr}}) = -\sum_{d=1}^D \frac{1}{D} \log q_d

    Because minimizing LD\mathcal{L}_D and minimizing Lconf\mathcal{L}_{\text{conf}} represent opposing objectives, parameter optimization alternates between training the domain classifier θD\theta_D to separate domains and training the representation θrepr\theta_{\text{repr}} to confound the classifier: min⁡θDLD(xS,xT,θrepr;θD)\min_{\theta_D} \mathcal{L}_D(x_S, x_T, \theta_{\text{repr}}; \theta_D) min⁡θreprLconf(xS,xT,θD;θrepr)\min_{\theta_{\text{repr}}} \mathcal{L}_{\text{conf}}(x_S, x_T, \theta_D; \theta_{\text{repr}})

  3. Knowl 3 — Inter-Class Semantic Transfer via High-Temperature Soft Label Matching

    model/method

    To transfer category relationships learned in the source domain to categories in the target domain (including target categories with no labeled instances), the network matches target predictions against category-level "soft labels" extracted from a source-trained CNN.

    For each class k∈{1,…,K}k \in \{1, \dots, K\}, the source soft label vector l(k)∈RKl^{(k)} \in \mathbb{R}^K is defined as the mean softmax activation over all source training examples of class kk, computed with temperature scaling τ>1\tau > 1: l(k)=1Nk∑x∈XSksoftmax(zS(x)τ)l^{(k)} = \frac{1}{N_k} \sum_{x \in X_S^k} \text{softmax}\left(\frac{z_S(x)}{\tau}\right) where zS(x)z_S(x) is the logit vector produced by the source CNN for input xx, and NkN_k is the number of source training samples in category kk. The high temperature parameter τ\tau prevents the distribution from collapsing onto a one-hot representation, preserving non-zero probability mass over visually and semantically related classes.

    For a target training image xTx_T with known class label yTy_T, the soft label loss minimizes the cross-entropy between the target model's temperature-scaled output probability p=softmax(θCTf(xT;θrepr)/τ)p = \text{softmax}(\theta_C^T f(x_T; \theta_{\text{repr}}) / \tau) and the precomputed source soft label l(yT)l^{(y_T)}: Lsoft(xT,yT;θrepr,θC)=−∑i=1Kli(yT)log⁡pi\mathcal{L}_{\text{soft}}(x_T, y_T; \theta_{\text{repr}}, \theta_C) = -\sum_{i=1}^K l^{(y_T)}_i \log p_i This objective updates category parameters even for classes lacking target labels by propagating probability distributions from related classes.

  4. Knowl 4 — Alternating Optimization Algorithm for Simultaneous Domain and Task Transfer

    algorithm

    The simultaneous transfer architecture is trained using an alternating gradient descent procedure that updates the domain classifier and feature representation iteratively.

    Input: Source dataset DS={(xS,yS)}D_S = \{(x_S, y_S)\}, target dataset DT={(xT,yT)}D_T = \{(x_T, y_T)\} with sparse labels, source-pretrained CNN weights θrepr,θC\theta_{\text{repr}}, \theta_C, precomputed per-class soft labels {l(k)}k=1K\{l^{(k)}\}_{k=1}^K, temperature τ\tau, loss weights λ=0.01\lambda = 0.01, ν=0.1\nu = 0.1, learning rate η=0.001\eta = 0.001
    Output: Adapted representation θrepr\theta_{\text{repr}} and target classifier θC\theta_C
    Initialize domain classifier weights θD\theta_D
    repeat
        Sample mini-batch of source instances BS⊂DSB_S \subset D_S and target instances BT⊂DTB_T \subset D_T
        
        // Step 1: Update domain classifier to discriminate domains
        Compute domain predictions q(x)=softmax(θDTf(x;θrepr))q(x) = \text{softmax}(\theta_D^T f(x; \theta_{\text{repr}})) for x∈BS∪BTx \in B_S \cup B_T
        Evaluate domain classification loss LD=−∑x∈BSlog⁡q1(x)−∑x∈BTlog⁡q2(x)L_D = -\sum_{x \in B_S} \log q_1(x) - \sum_{x \in B_T} \log q_2(x)
        Update domain classifier: θD←θD−η∇θDLD\theta_D \leftarrow \theta_D - \eta \nabla_{\theta_D} L_D
        
        // Step 2: Update representation and classification parameters
        Compute classification loss LCL_C on available labeled samples in BS∪BTB_S \cup B_T
        Compute domain confusion loss Lconf=−∑x∈BS∪BT∑d=1212log⁡qd(x)L_{\text{conf}} = -\sum_{x \in B_S \cup B_T} \sum_{d=1}^2 \frac{1}{2} \log q_d(x)
        Compute soft label loss Lsoft=−∑(xT,yT)∈BT∑i=1Kli(yT)log⁡pi(xT)L_{\text{soft}} = -\sum_{(x_T, y_T) \in B_T} \sum_{i=1}^K l^{(y_T)}_i \log p_i(x_T) where p(xT)=softmax(θCTf(xT;θrepr)/τ)p(x_T) = \text{softmax}(\theta_C^T f(x_T; \theta_{\text{repr}}) / \tau)
        
        Compute total joint loss L=LC+λLconf+νLsoftL = L_C + \lambda L_{\text{conf}} + \nu L_{\text{soft}}
        Update representation and classifier: θrepr←θrepr−η∇θreprL\theta_{\text{repr}} \leftarrow \theta_{\text{repr}} - \eta \nabla_{\theta_{\text{repr}}} L, θC←θC−η∇θCL\theta_C \leftarrow \theta_C - \eta \nabla_{\theta_C} L
    until convergence
    return θrepr,θC\theta_{\text{repr}}, \theta_C
  5. Knowl 5 — Office Benchmark Supervised Adaptation Performance

    data/table

    The supervised domain adaptation performance was evaluated on the 31-class Office dataset across six domain shifts between Amazon (A), Webcam (W), and DSLR (D). Source training sets used 20 examples per category for Amazon and 8 per category for DSLR/Webcam. The target training set contained 3 labeled examples per category, with evaluation performed on the remaining unlabeled target images. Multi-class accuracy was computed by averaging per-class accuracies across all 31 categories over 5 random splits.

    Method A →\to W A →\to D W →\to A W →\to D D →\to A D →\to W Average
    DLID 51.9 – – 89.9 – 78.2 –
    DeCAF6_6 S+T 80.7 ±\pm 2.3 – – – – 94.8 ±\pm 1.2 –
    DaNN 53.6 ±\pm 0.2 – – 83.5 ±\pm 0.0 – 71.2 ±\pm 0.0 –
    Source CNN 56.5 ±\pm 0.3 64.6 ±\pm 0.4 42.7 ±\pm 0.1 93.6 ±\pm 0.2 47.6 ±\pm 0.1 92.4 ±\pm 0.3 66.22
    Target CNN 80.5 ±\pm 0.5 81.8 ±\pm 1.0 59.9 ±\pm 0.3 81.8 ±\pm 1.0 59.9 ±\pm 0.3 80.5 ±\pm 0.5 74.05
    Source+Target CNN 82.5 ±\pm 0.9 85.2 ±\pm 1.1 65.2 ±\pm 0.7 96.3 ±\pm 0.5 65.8 ±\pm 0.5 93.9 ±\pm 0.5 81.50
    Ours: dom confusion only 82.8 ±\pm 0.9 85.9 ±\pm 1.1 64.9 ±\pm 0.5 97.5 ±\pm 0.2 66.2 ±\pm 0.4 95.6 ±\pm 0.4 82.13
    Ours: soft labels only 82.7 ±\pm 0.7 84.9 ±\pm 1.2 65.2 ±\pm 0.6 98.3 ±\pm 0.3 66.0 ±\pm 0.5 95.9 ±\pm 0.6 82.17
    Ours: dom confusion+soft labels 82.7 ±\pm 0.8 86.1 ±\pm 1.2 65.0 ±\pm 0.5 97.6 ±\pm 0.2 66.2 ±\pm 0.3 95.7 ±\pm 0.5 82.22

    The combined domain confusion and soft label model achieves an overall average accuracy of 82.22%82.22\%, outperforming shallow methods (DLID, DaNN) and standard source/target fine-tuning baselines (81.50%81.50\%).

  6. Knowl 6 — Office Benchmark Semi-Supervised Task Adaptation Performance

    data/table

    In the semi-supervised task adaptation setting on the Office dataset, the source domain contains labeled examples for all 31 classes, but target domain training labels (10 examples per category) are available for only 15 auxiliary categories. Models are evaluated exclusively on the remaining 16 held-out categories that received zero labeled target training instances. Accuracies are multi-class averages across the 16 held-out categories over 5 train/test splits.

    Method A →\to W A →\to D W →\to A W →\to D D →\to A D →\to W Average
    MMDT – 44.6 ±\pm 0.3 – 58.3 ±\pm 0.5 – – –
    Source CNN 54.2 ±\pm 0.6 63.2 ±\pm 0.4 34.7 ±\pm 0.1 94.5 ±\pm 0.2 36.4 ±\pm 0.1 89.3 ±\pm 0.5 62.0
    Ours: dom confusion only 55.2 ±\pm 0.6 63.7 ±\pm 0.9 41.1 ±\pm 0.0 96.5 ±\pm 0.1 41.2 ±\pm 0.1 91.3 ±\pm 0.4 64.8
    Ours: soft labels only 56.8 ±\pm 0.4 65.2 ±\pm 0.9 38.8 ±\pm 0.4 96.5 ±\pm 0.2 41.7 ±\pm 0.3 89.6 ±\pm 0.1 64.8
    Ours: dom confusion+soft labels 59.3 ±\pm 0.6 68.0 ±\pm 0.5 40.5 ±\pm 0.2 97.5 ±\pm 0.1 43.1 ±\pm 0.2 90.0 ±\pm 0.2 66.4

    Both domain confusion alone and soft label matching alone improve classification accuracy on unseen classes by 2.8% on average over the source CNN baseline (64.8% vs. 62.0%). Combining both losses achieves 66.4% average accuracy, providing a 13% relative improvement over the baselines on the four most difficult domain shifts (A→WA \to W, A→DA \to D, W→AW \to A, D→AD \to A).

  7. Knowl 7 — Cross-Dataset Supervised Transfer on ImageNet to Caltech-256

    empirical result

    On the dense cross-dataset testbed comprising 40 overlapping classes between ImageNet (source domain, 5,534 images) and Caltech-256 (target domain, 4,366 images) across 5 train/test splits, adaptation was evaluated with varying target labeled sample sizes (0,1,3,50, 1, 3, 5 labeled images per class):

    • With 0 target labeled examples, the domain confusion model outperforms the source-only CNN baseline (≈72.5% \approx 72.5\% accuracy).
    • For 1, 3, and 5 labeled target examples per class, the proposed method combining domain confusion and soft label matching achieves the highest multi-class accuracy, improving over both Source CNN and Source+Target CNN joint fine-tuning.
    • In contrast, fine-tuning on the target labeled data alone achieves only 36.6±0.6%36.6 \pm 0.6\% (1 example), 60.9±0.5%60.9 \pm 0.5\% (3 examples), and 67.7±0.5%67.7 \pm 0.5\% (5 examples), demonstrating that target-only fine-tuning fails in low-data regimes.
    • All deep transfer variants substantially exceed the prior state-of-the-art result of 24.8%24.8\% obtained with SURF Bag-of-Words features.
  8. Knowl 8 — Empirical Verification of Domain Invariance via Domain Classifier Degradation

    empirical result

    Domain invariance was quantified by evaluating whether an optimal domain discriminator could distinguish between Amazon and Webcam feature representations:

    • Support Vector Machines (SVMs) were trained on 160 images (80 from Amazon, 80 from Webcam) using either standard CaffeNet fc7 features or fc7 features learned with the domain confusion loss, and tested on the remaining images from both domains.
    • An SVM trained on the standard CaffeNet fc7 representation separated the two domains with 99%99\% test accuracy, indicating strong domain bias.
    • An SVM trained on the representation learned with the domain confusion loss achieved only 56%56\% test accuracy (approaching the 50%50\% random guessing baseline), demonstrating that the domain confusion loss effectively eliminates domain-identifying features from the latent space.
  9. Knowl 9 — Inter-Class Knowledge Transfer to Unlabeled Target Categories via Soft Labels

    empirical result

    Qualitative and categorical distribution analyses demonstrate how the soft label loss propagates supervision to classes with no labeled target data:

    • In the semi-supervised Amazon →\to Webcam setting where the class monitor has zero labeled target training instances, the auxiliary class laptop computer is present during training with target labels.
    • The precomputed source soft label for laptop computer places significant probability mass on monitor due to their visual and semantic similarity in the source domain.
    • During target training, matching laptop computer target samples to this source soft label forces the network to update parameters for the monitor class.
    • At test time, this mechanism enables the network to correctly classify target monitor images (which the source-only baseline misclassifies as ring binder), demonstrating cross-task transfer via empirical class correlations.

Coverage note — No substantial contributed material was omitted from the paper.

References

  1. 1.L. T. Alessandro Bergamo. Exploiting weakly-labeled web images to improve object classification: a domain adaptation approach. In Neural Information Processing Systems (NIPS), Dec. 2010.
  2. 2.Y. Aytar and A. Zisserman. Tabula rasa: Model transfer for object category detection. In Proc. ICCV, 2011.
  3. 3.J. Ba and R. Caruana. Do deep nets really need to be deep? In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger, editors, Advances in Neural Information Processing Systems 27, pages 2654–2662. Curran Associates, Inc., 2014.
  4. 4.A. Berg, J. Deng, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge 2012. 2012.
  5. 5.K. M. Borgwardt, A. Gretton, M. J. Rasch, H.-P. Kriegel, B. Scholkopf, and A. J. Smola. Integrating structured biological data by kernel maximum mean discrepancy. In Bioinformatics, 2006.
  6. 6.J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Sackinger, and R. Shah. Signature verification using a siamese time delay neural network. International Journal of Pattern Recognition and Artificial Intelligence, 7(04):669–688, 1993.
  7. 7.S. Chopra, S. Balakrishnan, and R. Gopalan. DLID: Deep learning for domain adaptation by interpolating between domains. In ICML Workshop on Challenges in Representation Learning, 2013.
  8. 8.S. Chopra, R. Hadsell, and Y. LeCun. Learning a similarity metric discriminatively, with application to face verification. In Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on, volume 1, pages 539–546. IEEE, 2005.
  9. 9.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition. In Proc. ICML, 2014.
  10. 10.L. Duan, D. Xu, and I. W. Tsang. Learning with augmented features for heterogeneous domain adaptation. In Proc. ICML, 2012.
  11. 11.B. Fernando, A. Habrard, M. Sebban, and T. Tuytelaars. Unsupervised visual domain adaptation using subspace alignment. In Proc. ICCV, 2013.
  12. 12.Y. Ganin and V. Lempitsky. Unsupervised Domain Adaptation by Backpropagation. ArXiv e-prints, Sept. 2014.
  13. 13.M. Ghifary, W. B. Kleijn, and M. Zhang. Domain adaptive neural networks for object recognition. CoRR, abs/1409.6041, 2014.
  14. 14.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. arXiv e-prints, 2013.
  15. 15.B. Gong, Y. Shi, F. Sha, and K. Grauman. Geodesic flow kernel for unsupervised domain adaptation. In Proc. CVPR, 2012.
  16. 16.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2014.
  17. 17.J. Hoffman, S. Guadarrama, E. Tzeng, R. Hu, J. Donahue, R. Girshick, T. Darrell, and K. Saenko. LSDA: Large scale detection through adaptation. In Neural Information Processing Systems (NIPS), 2014.
  18. 18.J. Hoffman, E. Rodner, J. Donahue, K. Saenko, and T. Darrell. Efficient learning of domain-invariant image representations. In Proc. ICLR, 2013.
  19. 19.J. Hoffman, E. Tzeng, J. Donahue, , Y. Jia, K. Saenko, and T. Darrell. One-shot learning of supervised deep convolutional models. In arXiv 1312.6204; presented at ICLR Workshop, 2014.
  20. 20.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  21. 21.D. Kifer, S. Ben-David, and J. Gehrke. Detecting change in data streams. In Proc. VLDB, 2004.
  22. 22.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In Proc. NIPS, 2012.
  23. 23.B. Kulis, K. Saenko, and T. Darrell. What you saw is not what you get: Domain adaptation using asymmetric kernel transforms. In Proc. CVPR, 2011.
  24. 24.M. Long and J. Wang. Learning transferable features with deep adaptation networks. CoRR, abs/1502.02791, 2015.
  25. 25.Y. Mansour, M. Mohri, and A. Rostamizadeh. Domain adaptation: Learning bounds and algorithms. In COLT, 2009.
  26. 26.J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning (ICML-11), pages 689–696, 2011.
  27. 27.S. J. Pan, I. W. Tsang, J. T. Kwok, and Q. Yang. Domain adaptation via transfer component analysis. In IJCA, 2009.
  28. 28.K. Saenko, B. Kulis, M. Fritz, and T. Darrell. Adapting visual category models to new domains. In Proc. ECCV, 2010.
  29. 29.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. CoRR, abs/1312.6229, 2013.
  30. 30.T. Tommasi, T. Tuytelaars, and B. Caputo. A testbed for cross-dataset analysis. In TASK-CV Workshop, ECCV, 2014.
  31. 31.A. Torralba and A. Efros. Unbiased look at dataset bias. In Proc. CVPR, 2011.
  32. 32.J. Yang, R. Yan, and A. Hauptmann. Adapting SVM classifiers to data with shifted distributions. In ICDM Workshops, 2007.

Citation

MLA
Tzeng, E., et al. “Simultaneous Deep Transfer Across Domains and Tasks”. arXiv, 2015, http://arxiv.org/abs/1510.02192v1.
APA
Tzeng, E., Hoffman, J., Darrell, T., & Saenko, K. (2015). Simultaneous Deep Transfer Across Domains and Tasks. arXiv. http://arxiv.org/abs/1510.02192v1
Chicago
Tzeng, E., J. Hoffman, T. Darrell, and K. Saenko. 2015. “Simultaneous Deep Transfer Across Domains and Tasks”. arXiv. http://arxiv.org/abs/1510.02192v1.
Harvard
Tzeng, E. et al. (2015) “Simultaneous Deep Transfer Across Domains and Tasks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1510.02192v1.
Vancouver
1. Tzeng E, Hoffman J, Darrell T, Saenko K (2015) Simultaneous Deep Transfer Across Domains and Tasks. arXiv

BibTeX

@article{tzeng2015simultaneous,
  title = {Simultaneous Deep Transfer Across Domains and Tasks},
  author = {Tzeng, Eric and Hoffman, Judy and Darrell, Trevor and Saenko, Kate},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1510.02192v1},
  eprint = {1510.02192}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE