Taskonomy: Disentangling Task Transfer Learning
Amir ZamirAlexander SaxWilliam ShenLeonidas GuibasJitendra MalikSilvio Savarese
Establishes a computational taxonomy of transfer learning dependencies across twenty-six common visual tasks, revealing how to reduce labeled training data requirements by roughly two-thirds across multi-task vision systems.
Developing comprehensive artificial intelligence and computer vision systems typically requires training separate neural networks for every specific capability, such as detecting edges, recognizing objects, or estimating depth. This isolated approach demands massive volumes of expensive labeled data, requires heavy computational infrastructure, and fails to leverage natural relationships between related perceptual functions. The article addresses this operational inefficiency by demonstrating a principled, computational method to identify how visual tasks relate to one another and map how knowledge can transfer between them to minimize supervisory requirements.
The investigation evaluated a dictionary of 26 common two-dimensional, three-dimensional, and semantic vision tasks across a standardized dataset of 4 million indoor images from roughly 600 buildings. The researchers trained individual models on identical images, evaluated approximately 3,000 transfer learning pathways using low-capacity readout networks, and normalized performance across varying task domains. They then formulated an optimization model that automatically selects the most efficient transfer strategy given any specified budget of fully trained source models.
The findings show that visual tasks exhibit strong, quantifiable transfer learning relationships that often run counter to human intuition. By implementing the optimized transfer structure, the total number of labeled data points needed to solve a suite of 10 tasks was reduced by roughly two-thirds compared to training models independently, while maintaining comparable performance. Furthermore, providing autonomous agents with representations from this optimized structure substantially improved sample efficiency and navigation performance in unseen environments compared to training from raw visual inputs.
These results provide a practical framework for organizations to cut data annotation costs, accelerate model deployment timelines, and design scalable perception systems for robotics and computer vision. Leadership teams building multi-task or autonomous systems should move away from training isolated perception models and instead adopt a shared transfer policy that prioritizes high-leverage source tasks. While the findings are highly reliable within the tested domain, users should exercise appropriate caution before directly deploying the specific mappings to outdoor environments or unconventional sensor modalities without localized validation.
- Paper: Multitask Learning, RICH CARUANA (1997). Caruana’s seminal work defines the principles of learning multiple tasks simultaneously to exploit shared inductive bias, providing the foundational conceptual underpinning for Taskonomy’s transfer learning taxonomy.
- Paper: Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-scale Convolutional Architecture, David Eigen et al. (2014). This paper demonstrates joint prediction of depth, surface normals, and semantic labels within a unified convolutional architecture, establishing the 2D, 2.5D, and 3D visual task settings analyzed in Taskonomy.
- Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). This work explores learning dynamic feature-sharing structures across paired visual tasks like surface normals and semantic segmentation, directly motivating Taskonomy's formalization of cross-task dependencies.
- Paper: Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, Alex Kendall et al. (2017). Kendall et al. develop multi-task loss weighting across geometry and semantics, addressing multi-task optimization trade-offs that Taskonomy maps into an empirical transfer space.
- Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). This paper quantifies the layer-wise transferability and specificity of deep representations across tasks, motivating Taskonomy's latent space transfer evaluation.
- Paper: A Survey on Multi-Task Learning, Yu Zhang et al. (2017). Zhang and Yang provide a comprehensive taxonomy of multi-task learning paradigms and task relationship modeling that frames Taskonomy’s computational transfer dictionary.
- Paper: A Survey on Transfer Learning, Sinno Jialin Pan et al. (2010). This survey formalizes the fundamental theoretical definitions and paradigms of transfer learning that Taskonomy adapts to the space of 26 visual tasks.
- Paper: DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition, Jeff Donahue et al. (2013). Donahue et al. establish how deep representations pretrained on one task can be effectively repurposed as generic visual features for others.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). This work directly addresses negative transfer and conflicting gradient interference that arise when training multi-task models derived from task relationship taxonomies like Taskonomy.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Unified-IO scales the unification of diverse 2D, 3D, semantic, and multimodal visual tasks explored in Taskonomy into a single sequence-to-sequence foundation model.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). Contrastive Multiview Coding leverages cross-modal and multi-task visual correlations—such as depth and surface normals analyzed in Taskonomy—to learn powerful self-supervised representations.
- Paper: A Comprehensive Survey on Transfer Learning, Fuzhen Zhuang et al. (2019). This comprehensive survey categorizes modern transfer learning advances, contextualizing task-space mapping approaches like Taskonomy within the broader transfer landscape.
