VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding
Yi XinJunlong DuQiang WangZhiwen LinKe Yan
Proposes an efficient multi-task adapter framework that adapts pre-trained vision models to multiple dense scene understanding tasks simultaneously with constant-time efficiency and only one percent of trainable parameters.
Large vision models deliver strong performance across diverse visual tasks, but adapting them by updating all model weights requires massive compute and storage budgets. While parameter-efficient fine-tuning has emerged to update only tiny submodules, existing approaches primarily focus on single tasks. When applied to multi-task visual problems—such as simultaneously identifying object categories, segmenting human parts, estimating surface normals, and detecting salient areas—current methods either scale up parameter counts linearly or require running images through the backbone model multiple times, causing severe latency and memory bottlenecks.
The article evaluates whether a unified, multi-task adapter framework can efficiently transfer large pre-trained vision transformers to several dense scene understanding tasks at once without multiplying training or inference costs. To demonstrate this, the authors develop the Vision Multi-Task Adapter (VMT-Adapter) and a lightweight variant (VMT-Adapter-Lite), testing them on the standard PASCAL-Context benchmark comprising four distinct dense prediction tasks.
The framework pairs shared projection layers that capture cross-task interactions with minimal task-specific modules that extract distinct features via simple scaling and shifting operations. This unified design allows images to pass through the frozen visual encoder only once, regardless of the number of tasks. The authors benchmarked their method against full fine-tuning and several existing parameter-efficient techniques across different encoder scales (Swin-Tiny and Swin-Base) and decoder architectures.
The experimental findings show significant performance and efficiency gains. First, VMT-Adapter outperformed single-task full fine-tuning by an average of 3.96% across all four dense tasks while updating only about 1% (1.13 million) of the encoder parameters. Second, the lightweight variant, VMT-Adapter-Lite, achieved a 1.34% gain over full fine-tuning using just 0.36% (0.40 million) of the parameters. Third, unlike prior multi-task adapters that process inputs repeatedly per task, the proposed design achieves constant O(1) training and inference throughput relative to task count. Fourth, the architecture scales effectively with larger vision backbones, widening its relative improvement over full fine-tuning to 7.10% on the larger Swin-Base model.
These results demonstrate that organizations can deploy complex, multi-functional vision systems at a fraction of the computational and storage expenses normally required. By mitigating negative gradient interference between simultaneous tasks, the shared-plus-specific structure boosts overall accuracy while removing the deployment cost barriers of storing multiple heavy task-specific models.
For practical implementation, teams should adopt VMT-Adapter when top task performance is essential, or VMT-Adapter-Lite with a matrix parameter dimension of m=3 when operating under stringent storage or memory constraints. Looking ahead, practitioners should conduct additional pilots on broader datasets and monitor scaling behavior, as the cumulative size of task-specific modules may increase if scaled to hundreds or thousands of simultaneous tasks.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). Introduces the foundational bottleneck adapter architecture for parameter-efficient transfer learning that VMT-Adapter adapts and builds upon for multi-task dense vision models.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Establishes a unified structural taxonomy for parameter-efficient fine-tuning mechanisms, providing the design principles leveraged by modern visual adapter frameworks.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Demonstrates parameter-efficient fine-tuning directly on vision transformer backbones, establishing key benchmarks that VMT-Adapter extends to multi-task visual settings.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). Introduces the hierarchical Swin Transformer architecture that serves as the core pre-trained visual backbone adapted in VMT-Adapter's experiments.
- Paper: End-To-End Multi-Task Learning With Attention, Shikun Liu et al. (2018). Pioneers shared-versus-task-specific attention mechanisms for multi-task dense visual prediction, establishing foundational concepts for multi-task visual learning.
- Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Analyzes gradient conflicts in simultaneous multi-task training, motivating VMT-Adapter's shared-plus-specific modular design to mitigate negative interference.
- Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). Provides foundational strategies for sharing and splitting visual representations across concurrent tasks to avoid brute-force separate model training.
- Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). Defines multi-task visual parsing across diverse dense scene understanding tasks, formulating the multi-task evaluation paradigm used in dense scene benchmarks.
No sufficiently relevant recommendations were found.
