Taskonomy: Disentangling Task Transfer Learning

Amir ZamirAlexander SaxWilliam ShenLeonidas GuibasJitendra MalikSilvio Savarese

article2018CVPR1,431 citationsBest Paper Award

Establishes a computational taxonomy of transfer learning dependencies across twenty-six common visual tasks, revealing how to reduce labeled training data requirements by roughly two-thirds across multi-task vision systems.

Listen

Developing comprehensive artificial intelligence and computer vision systems typically requires training separate neural networks for every specific capability, such as detecting edges, recognizing objects, or estimating depth. This isolated approach demands massive volumes of expensive labeled data, requires heavy computational infrastructure, and fails to leverage natural relationships between related perceptual functions. The article addresses this operational inefficiency by demonstrating a principled, computational method to identify how visual tasks relate to one another and map how knowledge can transfer between them to minimize supervisory requirements.

The investigation evaluated a dictionary of 26 common two-dimensional, three-dimensional, and semantic vision tasks across a standardized dataset of 4 million indoor images from roughly 600 buildings. The researchers trained individual models on identical images, evaluated approximately 3,000 transfer learning pathways using low-capacity readout networks, and normalized performance across varying task domains. They then formulated an optimization model that automatically selects the most efficient transfer strategy given any specified budget of fully trained source models.

The findings show that visual tasks exhibit strong, quantifiable transfer learning relationships that often run counter to human intuition. By implementing the optimized transfer structure, the total number of labeled data points needed to solve a suite of 10 tasks was reduced by roughly two-thirds compared to training models independently, while maintaining comparable performance. Furthermore, providing autonomous agents with representations from this optimized structure substantially improved sample efficiency and navigation performance in unseen environments compared to training from raw visual inputs.

These results provide a practical framework for organizations to cut data annotation costs, accelerate model deployment timelines, and design scalable perception systems for robotics and computer vision. Leadership teams building multi-task or autonomous systems should move away from training isolated perception models and instead adopt a shared transfer policy that prioritizes high-leverage source tasks. While the findings are highly reliable within the tested domain, users should exercise appropriate caution before directly deploying the specific mappings to outdoor environments or unconventional sensor modalities without localized validation.

arXiv: 1804.08328
  • Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). This work directly addresses negative transfer and conflicting gradient interference that arise when training multi-task models derived from task relationship taxonomies like Taskonomy.
  • Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Unified-IO scales the unification of diverse 2D, 3D, semantic, and multimodal visual tasks explored in Taskonomy into a single sequence-to-sequence foundation model.
  • Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). Contrastive Multiview Coding leverages cross-modal and multi-task visual correlations—such as depth and surface normals analyzed in Taskonomy—to learn powerful self-supervised representations.
  • Paper: A Comprehensive Survey on Transfer Learning, Fuzhen Zhuang et al. (2019). This comprehensive survey categorizes modern transfer learning advances, contextualizing task-space mapping approaches like Taskonomy within the broader transfer landscape.
Cover for Taskonomy: Disentangling Task Transfer Learning

Abstract

Do visual tasks have a relationship, or are they unrelated? For instance, could having surface normals simplify estimating the depth of an image? Intuition answers these questions positively, implying existence of a structure among visual tasks. Knowing this structure has notable values; it is the concept underlying transfer learning and provides a principled way for identifying redundancies across tasks, e.g., to seamlessly reuse supervision among related tasks or solve many tasks in one system without piling up the complexity.

We proposes a fully computational approach for modeling the structure of space of visual tasks. This is done via finding (first and higher-order) transfer learning dependencies across a dictionary of twenty six 2D, 2.5D, 3D, and semantic tasks in a latent space. The product is a computational taxonomic map for task transfer learning. We study the consequences of this structure, e.g. nontrivial emerged relationships, and exploit them to reduce the demand for labeled data. For example, we show that the total number of labeled datapoints needed for solving a set of 10 tasks can be reduced by roughly 2/3 (compared to training independently) while keeping the performance nearly the same. We provide a set of tools for computing and probing this taxonomical structure including a solver that users can employ to devise efficient supervision policies for their use cases.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Step I: Task-Specific Modeling
  • 3.2 Step II: Transfer Modeling
  • 3.3 Step III: Ordinal Normalization using Analytic Hierarchy Process (AHP)
  • 3.4 Step IV: Computing the Global Taxonomy
  • 4 Experiments
  • 4.1 Evaluation of Computed Taxonomies
  • 4.2 Generalization to Novel Tasks
  • 5 Significance Test of the Structure
  • 5.1 Evaluation on MIT Places & ImageNet
  • 5.2 Universality of the Structure
  • 5.3 Task Similarity Tree
  • 6 Limitations and Discussion
  • References

Knowls

  1. Knowl 1 — Four-Step Framework for Computational Task Taxonomy (Taskonomy)

    model/method

    Taskonomy computes the transfer learning structure among visual tasks by modeling task relationships as a directed hypergraph through four sequential steps:

    1. Task-Specific Modeling: A set of fully supervised feedforward convolutional neural networks is trained independently from scratch for each source task s∈Ss \in \mathcal{S}, using a standardized, homogeneous architecture across all tasks to eliminate architectural bias.
    2. Transfer Modeling: For every target task t∈Tt \in \mathcal{T} and source task subset {s1,…,sk}⊆S\{s_1, \dots, s_k\} \subseteq \mathcal{S}, a small, low-capacity readout network (transfer function) is trained on a small amount of target task data. The transfer network takes as input the frozen latent representations extracted by the encoders of the source task networks to predict the output of target task tt.
    3. Ordinal Normalization: Because loss metrics across distinct visual domains (e.g., depth regression vs. semantic segmentation classification) have non-commensurable numerical ranges and scales, the test performances of all transfer functions for a given target are converted into scale-free, normalized task affinities using the Analytic Hierarchy Process (AHP).
    4. Global Taxonomy Extraction: A global transfer policy is computed using Binary Integer Programming (BIP), which extracts an optimal subgraph selecting the combination of source tasks and transfer hyperedges that maximizes aggregate target task performance subject to a user-defined supervision budget γ\gamma (maximum allowable source tasks trained from scratch).
  2. Knowl 2 — Global Taxonomy Subgraph Selection via Binary Integer Programming

    model/method

    Let V=T∪S\mathcal{V} = \mathcal{T} \cup \mathcal{S} denote a dictionary of visual tasks, where T={t1,…,tn}\mathcal{T} = \{t_1, \dots, t_n\} is the set of target tasks to be solved, and S\mathcal{S} is the set of source tasks that can be trained from scratch. Tasks in T∖(T∩S)\mathcal{T} \setminus (\mathcal{T} \cap \mathcal{S}) are target-only tasks, tasks in T∩S\mathcal{T} \cap \mathcal{S} can act as both source and target, and tasks in S∖(T∩S)\mathcal{S} \setminus (\mathcal{T} \cap \mathcal{S}) are auxiliary source-only tasks.

    Finding the optimal transfer policy is formulated as a constraint satisfaction subgraph selection problem on a directed hypergraph and solved via Binary Integer Programming (BIP):

    max⁡{xe},{ys}∑t∈T∑e∈Etwexe\max_{\{x_e\}, \{y_s\}} \sum_{t \in \mathcal{T}} \sum_{e \in \mathcal{E}_t} w_e x_e

    subject to: ∑s∈Sys≤γ\sum_{s \in \mathcal{S}} y_s \le \gamma ∑e∈Etxe=1∀t∈T\sum_{e \in \mathcal{E}_t} x_e = 1 \quad \forall t \in \mathcal{T} xe≤ys∀e∈E,  ∀s∈sources(e)x_e \le y_s \quad \forall e \in \mathcal{E}, \; \forall s \in \text{sources}(e) xe∈{0,1},ys∈{0,1}x_e \in \{0, 1\}, \quad y_s \in \{0, 1\}

    where γ∈Z+\gamma \in \mathbb{Z}^+ is the total supervision budget (the maximum number of source tasks allowed to be trained from scratch), ysy_s is a binary decision variable indicating whether source task s∈Ss \in \mathcal{S} is trained from scratch, Et\mathcal{E}_t is the set of candidate transfer hyperedges into target task tt, xex_e is a binary decision variable indicating whether hyperedge ee is selected, sources(e)\text{sources}(e) is the subset of source tasks providing representations for edge ee, and wew_e is the normalized transfer affinity weight of hyperedge ee obtained via Analytic Hierarchy Process normalization.

  3. Knowl 3 — Low-Capacity Transfer Modeling and Information Accessibility

    model/method

    To evaluate whether a source task ss transfers effectively to a target task tt, the source representation must be informative for tt and this information must be easily accessible (computationally read out without retraining a deep network from scratch). If an excessively expressive transfer model were used, high target performance could stem from the transfer network learning the task itself rather than extracting features from the source representation.

    To ensure that transfer affinity measures true representation utility:

    • The encoder weights of the task-specific source network are kept frozen.
    • The transfer function is restricted to a low-capacity neural network with few parameters.
    • The transfer function is trained on a small amount of target training data.

    Under these constraints, high validation performance on the target task indicates that the latent feature space of the source encoder exposes the necessary statistical properties in an easily extractable form.

  4. Knowl 4 — Ordinal Normalization of Transfer Affinities via Analytic Hierarchy Process

    model/method

    Visual tasks employ heterogeneous evaluation metrics and loss functions with incompatible numerical distributions (for example, pixel-wise angular error for surface normal estimation, ℓ1\ell_1 loss for metric depth, and cross-entropy for multi-class classification). Consequently, raw transfer loss values cannot be directly compared or summed across different target tasks.

    To normalize affinities into a common scale:

    1. For each target task tt, the validation performances of all candidate source transfers (single source and higher-order combinations) are computed.
    2. An ordinal ranking of sources is created for each target task based on relative performance.
    3. The Analytic Hierarchy Process (AHP) is used to map pairwise ordinal comparisons and ratios into a normalized, ratio-scale affinity vector for each target task, such that transfer weights for a target sum to 11.

    This produces a calibrated directed affinity matrix that allows the Binary Integer Program to balance trade-offs across disparate visual tasks during global taxonomy optimization.

  5. Knowl 5 — Taskonomy Standardized Multi-Task Benchmark and Dictionary

    experimental setup

    To isolate intrinsic task relationships from domain shift or dataset-specific artifacts, the Taskonomy framework evaluates transfer learning across a controlled multi-task setup:

    • Dataset: 4 million indoor scene images sampled across approximately 600 buildings. All camera views are aligned with building-wide 3D triangular meshes, enabling programmatic derivation of ground-truth annotations across multiple modalities for identical pixel inputs.
    • Task Dictionary: 26 tasks spanning 2D, 2.5D, 3D, and semantic perceptual levels. The dictionary contains 22 primary tasks (including surface normals, Euclidean distance, z-depth, 2D edges, 3D edges, 2D keypoints, 3D keypoints, 2D segmentation, 2.5D segmentation, semantic segmentation, 3D curvature, reshading, denoising, room layout, vanishing points, camera pose estimation, object classification, and scene classification) and 4 source-only tasks (colorization, inpainting, jigsaw puzzle solving, and random projection).
    • Transfer Function Training: For 1st-order transfers (k=1k=1), 22×25=55022 \times 25 = 550 transfer networks are evaluated. Across 1st and higher kk-th order transfers (22×(25k)22 \times \binom{25}{k} possibilities), roughly 3,000 transfer networks were sampled and trained, totaling 47,886 cloud GPU hours.
  6. Knowl 6 — Labeled Data Reduction via Taxonomy Transfer Policies

    empirical result

    By selecting transfer policies computed via the Taskonomy global optimization graph, the total number of labeled training samples required to solve a collection of 10 visual target tasks is reduced by approximately two-thirds (≈67%\approx 67\%) compared to training 10 independent fully supervised models from scratch, while preserving near-equivalent overall performance.

    The quality of the transfer policy is evaluated using two primary metrics:

    • Gain: The win rate of the transfer-learned model against a baseline network trained on the same limited target data without transfer learning.
    • Quality: The win rate of the transfer-learned model against a fully supervised gold-standard network trained on the complete dataset from scratch.

    As the supervision budget γ\gamma increases, both Gain and Quality monotonically increase across target tasks until plateauing near the fully supervised upper bound.

  7. Knowl 7 — Computational Directional Asymmetry in Neural Task Transfer

    empirical result

    Empirically measured transfer learning affinities between visual tasks frequently display directional asymmetry that contradicts analytical mathematical intuitions:

    • Analytically, surface normals can be computed as the derivative of depth, suggesting that depth estimation representations should readily transfer to surface normal estimation. However, empirical transfer modeling reveals that transferring from a surface normal encoder to depth estimation is computationally much more effective than the reverse direction (normals→depth>depth→normalsnormals \to depth > depth \to normals).
    • Multi-source transfers reveal complementary synergies: combining low-level geometric and appearance representations (such as 2D edges and reshading) provides higher transferability to spatial orientation tasks (such as camera pose estimation) than transferring from either source task alone.

    These results establish that transferability in deep neural representations is governed by empirical accessibility and optimization properties within neural network function classes rather than analytical or closed-form mathematical derivability.

  8. Knowl 8 — Mid-Level Visual Representations as Generic Priors for Visuomotor Control

    empirical result

    Pre-trained representations from the Taskonomy task dictionary can serve as frozen mid-level vision modules for reinforcement learning (RL) agents solving active visuomotor tasks (such as robotic indoor navigation):

    • Feeding policies with frozen mid-level features (e.g., surface normals, 2D edges, occlusion boundaries, room layout) instead of raw pixels yields substantially higher sample efficiency and improves policy generalization to unseen indoor environments where end-to-end learning from scratch fails.
    • The performance gain is sensitive to the chosen set of visual features. Selecting a maximal-coverage subset of mid-level tasks using the Taskonomy transferability structure provides a compact, general visual prior that outperforms ad-hoc feature selections for downstream robotic tasks.

Coverage note — Specific implementation details regarding individual network layer counts, exact training loss equations for each of the 26 tasks, and full hyperparameter tables from the full CVPR 2018 paper were omitted because this IJCAI-19 summary paper references them to the original publication and presents them in condensed form.

References

  1. 1.Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017.
  2. 2.Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence, 35(8):1798–1828, 2013.
  3. 3.Zhiyuan Chen and Bing Liu. Lifelong Machine Learning. Morgan & Claypool Publishers, 2016.
  4. 4.Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In International Conference on Computer Vision, 2015.
  5. 5.Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. Why does unsupervised pre-training help deep learning? Journal of Machine Learning Research, 2010.
  6. 6.Rong Ge. Provable algorithms for machine learning problems. PhD thesis, Princeton University, 2013.
  7. 7.Stuart Geman, Daniel F Potter, and Zhiyi Chi. Composition systems. Quarterly of Applied Mathematics, 60(4):707–736, 2002.
  8. 8.Alison Gopnik, Andrew N Meltzoff, and Patricia K Kuhl. The scientist in the crib: Minds, brains, and how children learn. William Morrow & Co, 1999.
  9. 9.Inc. Gurobi Optimization. Gurobi optimizer reference manual, 2016.
  10. 10.Kevin Henry. The theory and applications of homomorphic cryptography. 2008.
  11. 11.Yedid Hoshen and Shmuel Peleg. Visual learning of arithmetic operations. CoRR, 2015.
  12. 12.Iasonas Kokkinos. Ubernet: Training a universal’convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. arXiv, 2016.
  13. 13.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and Brain Sciences, pages 1–101, 2016.
  14. 14.Michael Mccloskey and Neil J. Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. The Psychology of Learning and Motivation, 24, 1989.
  15. 15.Anastasia Pentina and Christoph H Lampert. Multi-task learning with labeled and unlabeled tasks. stat, 1050:1, 2017.
  16. 16.Jean Piaget and Margaret Cook. The origins of intelligence in children, volume 8. International Universities Press New York, 1952.
  17. 17.Lorien Y Pratt. Discriminability-based transfer between neural networks. In Advances in neural information processing systems, pages 204–211, 1993.
  18. 18.Olga Russakovsky et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 2015.
  19. 19.R. W. Saaty. The analytic hierarchy process – what it is and how it is used. Mathematical Modeling, 1987.
  20. 20.Ruslan Salakhutdinov, Joshua Tenenbaum, and Antonio Torralba. One-shot learning with a hierarchical nonparametric bayesian model. In Proceedings of ICML Workshop on Unsupervised and Transfer Learning, 2012.
  21. 21.Alexander Sax, Bradley Emi, Amir Zamir, Leonidas Guibas, Silvio Savarese, and Jitendra Malik. Mid-level visual representations improve generalization and sample efficiency for learning visuomotor policies. arXiv preprint, 2018.
  22. 22.Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In Conference on computer vision and pattern recognition workshops, 2014.
  23. 23.Daniel L. Silver, Qiang Yang, and Lianghao Li. Lifelong machine learning systems: Beyond learning algorithms. In in AAAI Spring Symposium Series, 2013.
  24. 24.Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng. Zero-shot learning through cross-modal transfer. In NIPS, pages 935–943, 2013.
  25. 25.Trevor Standley, Amir Zamir, Dawn Chen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Which tasks should be learned together in multi-task learning? arXiv preprint arXiv:1905.07553, 2019.
  26. 26.Sebastian Thrun and Lorien Pratt. Learning to learn. Springer Science & Business Media, 2012.
  27. 27.Alan M Turing. Computing machinery and intelligence. Mind, 59(236):433–460, 1950.
  28. 28.Amir Zamir, Alexander Sax, William Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  29. 29.Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. Learning deep features for scene recognition using places database. In Advances in neural information processing systems, pages 487–495, 2014.

Citation

MLA
Zamir, A., et al. “Taskonomy: Disentangling Task Transfer Learning”. arXiv, 2018, http://arxiv.org/abs/1804.08328v1.
APA
Zamir, A., Sax, A., Shen, W., Guibas, L., Malik, J., & Savarese, S. (2018). Taskonomy: Disentangling Task Transfer Learning. arXiv. http://arxiv.org/abs/1804.08328v1
Chicago
Zamir, A., A. Sax, W. Shen, L. Guibas, J. Malik, and S. Savarese. 2018. “Taskonomy: Disentangling Task Transfer Learning”. arXiv. http://arxiv.org/abs/1804.08328v1.
Harvard
Zamir, A. et al. (2018) “Taskonomy: Disentangling Task Transfer Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1804.08328v1.
Vancouver
1. Zamir A, Sax A, Shen W, Guibas L, Malik J, Savarese S (2018) Taskonomy: Disentangling Task Transfer Learning. arXiv

BibTeX

@article{zamir2018taskonomy,
  title = {Taskonomy: Disentangling Task Transfer Learning},
  author = {Zamir, Amir and Sax, Alexander and Shen, William and Guibas, Leonidas and Malik, Jitendra and Savarese, Silvio},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1804.08328v1},
  eprint = {1804.08328}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE