End-To-End Multi-Task Learning With Attention
Shikun LiuEdward JohnsAndrew J. Davison
Introduces the Multi-Task Attention Network (MTAN), an architecture that applies task-specific soft-attention masks to a shared feature extractor, achieving state-of-the-art visual performance while reducing sensitivity to loss weighting schemes.
Modern computer vision systems often require models capable of performing several tasks at once, such as object recognition, depth perception, and surface boundary detection. Deploying independent, dedicated networks for each task creates high computational costs, excessive memory consumption, and slow operational speeds. Conversely, conventional multi-task architectures struggle to balance general feature sharing with specialized task requirements, frequently requiring manual and labor-intensive tuning of loss functions to prevent easier tasks from dominating the training process.
The article demonstrates and evaluates the Multi-Task Attention Network (MTAN), an end-to-end framework designed to efficiently share a compact pool of global features while learning task-specific features via soft attention mechanisms. The primary objective is to prove that MTAN achieves high predictive accuracy across diverse tasks while remaining parameter-efficient and structurally robust to variations in loss-weight balancing schemes.
The researchers evaluated the proposed approach through comparative benchmark experiments across both dense pixel-level prediction and multi-domain image classification. For dense visual tasks—such as semantic segmentation, depth estimation, and surface normal prediction—MTAN was integrated with a standard encoder-decoder architecture and tested on the urban outdoor CityScapes dataset and the complex indoor NYUv2 dataset. For classification, the approach was tested across ten distinct domains within the Visual Decathlon Challenge using a Wide Residual Network backbone. The architecture was benchmarked against single-task networks and multiple competitive multi-task baselines across multiple loss-balancing strategies, including an introduced loss-balancing technique called Dynamic Weight Average (DWA).
The experimental findings show that MTAN achieves state-of-the-art or highly competitive performance while scaling significantly better than existing approaches. On the challenging NYUv2 indoor benchmark, MTAN outperformed all baseline multi-task methods across every task and loss-weighting strategy. Unlike architectures that duplicate network capacity linearly as tasks are added, MTAN requires only about a 10% parameter increase per additional task. Furthermore, the architecture proved substantially less sensitive to the choice of training loss weights, maintaining consistent learning curves across equal, uncertainty-based, and dynamic weighting schemes. The performance advantage over single-task models expanded notably as task complexity increased, with learned attention masks effectively filtering shared representations for task-specific needs.
These results indicate that organizations can deploy unified vision systems that substantially reduce memory footprints and inference latency without sacrificing accuracy. Because MTAN exhibits natural resilience to loss weighting, engineering teams can minimize time-consuming manual hyperparameter tuning during deployment pipelines. This makes the architecture particularly suitable for resource-constrained production settings such as autonomous driving and robotics where multiple concurrent visual perception tasks must operate reliably in real time.
Decision-makers should consider adopting soft attention-based multi-task designs over disjoint model ensembles or heavily duplicated networks when deploying multi-functional computer vision systems. Before full production rollout, engineering teams should conduct pilot implementations on their target hardware platforms to validate operational latency and determine whether simple dynamic weighting methods like DWA are sufficient for their domain. Confidence in these conclusions is high given the consistent empirical performance across varied network backbones and standard datasets, though users should evaluate performance on domain-specific datasets if edge cases fall outside the standard benchmarks evaluated in the article.
- Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). Cross-Stitch Networks establish foundational methods for learning shared and task-specific visual representations end-to-end, directly preceding the design of task-specific attention mechanisms on shared backbones.
- Paper: GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks, Zhao Chen et al. (2017). GradNorm introduces dynamic loss balancing in multi-task networks, providing the baseline optimization framework and motivation for architectural alternatives like MTAN that are less sensitive to loss weighting.
- Paper: Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, Alex Kendall et al. (2017). This paper establishes the standard benchmark suite of multi-task computer vision (semantic segmentation and depth estimation) and homoscedastic uncertainty weighting evaluated in MTAN.
- Paper: Residual Attention Network for Image Classification, Fei Wang et al. (2017). This work demonstrates how soft-attention residual modules can be integrated into feed-forward convolutional backbones to modulate feature pools.
- Paper: An Overview of Multi-Task Learning in Deep Neural Networks, Sebastian Ruder (2017). Provides a comprehensive overview of hard versus soft parameter sharing in deep multi-task architectures, clarifying the architectural space MTAN improves upon.
- Paper: Multitask Learning, RICH CARUANA (1997). Rich Caruana's foundational text provides the core principles of multi-task learning through shared representations that modern deep architectures like MTAN build upon.
- Paper: Multi-Task Learning as Multi-Objective Optimization, Ozan Sener et al. (2018). Formulates multi-task learning as multi-objective optimization (MGDA) on shared representations, providing an optimization-centric counterpart to MTAN's architectural solution for task conflicts.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Introduces gradient projection (PCGrad) to mitigate conflicting gradient directions across tasks, complementing feature-level multi-task attention architectures.
- Paper: Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts, Jiaqi Ma et al. (2018). Explores task relationships in multi-task settings through task-specific gating over shared sub-networks (Mixture-of-Experts), offering an alternative routing mechanism to soft attention masks.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Expands multi-task representation sharing into a unified, large-scale multi-modal transformer architecture handling diverse vision and language tasks simultaneously.
