VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding

Yi XinJunlong DuQiang WangZhiwen LinKe Yan

article2024AAAI93 citations

Proposes an efficient multi-task adapter framework that adapts pre-trained vision models to multiple dense scene understanding tasks simultaneously with constant-time efficiency and only one percent of trainable parameters.

Listen

Large vision models deliver strong performance across diverse visual tasks, but adapting them by updating all model weights requires massive compute and storage budgets. While parameter-efficient fine-tuning has emerged to update only tiny submodules, existing approaches primarily focus on single tasks. When applied to multi-task visual problems—such as simultaneously identifying object categories, segmenting human parts, estimating surface normals, and detecting salient areas—current methods either scale up parameter counts linearly or require running images through the backbone model multiple times, causing severe latency and memory bottlenecks.

The article evaluates whether a unified, multi-task adapter framework can efficiently transfer large pre-trained vision transformers to several dense scene understanding tasks at once without multiplying training or inference costs. To demonstrate this, the authors develop the Vision Multi-Task Adapter (VMT-Adapter) and a lightweight variant (VMT-Adapter-Lite), testing them on the standard PASCAL-Context benchmark comprising four distinct dense prediction tasks.

The framework pairs shared projection layers that capture cross-task interactions with minimal task-specific modules that extract distinct features via simple scaling and shifting operations. This unified design allows images to pass through the frozen visual encoder only once, regardless of the number of tasks. The authors benchmarked their method against full fine-tuning and several existing parameter-efficient techniques across different encoder scales (Swin-Tiny and Swin-Base) and decoder architectures.

The experimental findings show significant performance and efficiency gains. First, VMT-Adapter outperformed single-task full fine-tuning by an average of 3.96% across all four dense tasks while updating only about 1% (1.13 million) of the encoder parameters. Second, the lightweight variant, VMT-Adapter-Lite, achieved a 1.34% gain over full fine-tuning using just 0.36% (0.40 million) of the parameters. Third, unlike prior multi-task adapters that process inputs repeatedly per task, the proposed design achieves constant O(1) training and inference throughput relative to task count. Fourth, the architecture scales effectively with larger vision backbones, widening its relative improvement over full fine-tuning to 7.10% on the larger Swin-Base model.

These results demonstrate that organizations can deploy complex, multi-functional vision systems at a fraction of the computational and storage expenses normally required. By mitigating negative gradient interference between simultaneous tasks, the shared-plus-specific structure boosts overall accuracy while removing the deployment cost barriers of storing multiple heavy task-specific models.

For practical implementation, teams should adopt VMT-Adapter when top task performance is essential, or VMT-Adapter-Lite with a matrix parameter dimension of m=3 when operating under stringent storage or memory constraints. Looking ahead, practitioners should conduct additional pilots on broader datasets and monitor scaling behavior, as the cumulative size of task-specific modules may increase if scaled to hundreds or thousands of simultaneous tasks.

  • Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). Introduces the foundational bottleneck adapter architecture for parameter-efficient transfer learning that VMT-Adapter adapts and builds upon for multi-task dense vision models.
  • Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). Establishes a unified structural taxonomy for parameter-efficient fine-tuning mechanisms, providing the design principles leveraged by modern visual adapter frameworks.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Demonstrates parameter-efficient fine-tuning directly on vision transformer backbones, establishing key benchmarks that VMT-Adapter extends to multi-task visual settings.
  • Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). Introduces the hierarchical Swin Transformer architecture that serves as the core pre-trained visual backbone adapted in VMT-Adapter's experiments.
  • Paper: End-To-End Multi-Task Learning With Attention, Shikun Liu et al. (2018). Pioneers shared-versus-task-specific attention mechanisms for multi-task dense visual prediction, establishing foundational concepts for multi-task visual learning.
  • Paper: Gradient Surgery for Multi-Task Learning, Tianhe Yu et al. (2020). Analyzes gradient conflicts in simultaneous multi-task training, motivating VMT-Adapter's shared-plus-specific modular design to mitigate negative interference.
  • Paper: Cross-Stitch Networks for Multi-task Learning, Ishan Misra et al. (2016). Provides foundational strategies for sharing and splitting visual representations across concurrent tasks to avoid brute-force separate model training.
  • Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). Defines multi-task visual parsing across diverse dense scene understanding tasks, formulating the multi-task evaluation paradigm used in dense scene benchmarks.

No sufficiently relevant recommendations were found.

Cover for VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding

Abstract

Large-scale pre-trained models have achieved remarkable success in various computer vision tasks. A standard approach to leverage these models is to fine-tune all model parameters for downstream tasks, which poses challenges in terms of computational and storage costs. Recently, inspired by Natural Language Processing (NLP), parameter-efficient transfer learning has been successfully applied to vision tasks. However, most existing techniques primarily focus on single-task adaptation, and despite limited research on multi-task adaptation, these methods often exhibit suboptimal training and inference efficiency. In this paper, we first propose an once-for-all Vision Multi-Task Adapter (VMT-Adapter), which strikes approximately O(1) training and inference efficiency w.r.t task number. Concretely, VMT-Adapter shares the knowledge from multiple tasks to enhance cross-task interaction while preserves task-specific knowledge via independent knowledge extraction modules. Notably, since task-specific modules require few parameters, VMT-Adapter can handle an arbitrary number of tasks with a negligible increase of trainable parameters. We also propose VMT-Adapter-Lite, which further reduces the trainable parameters by learning shared parameters between down- and up-projections. Extensive experiments on four dense scene understanding tasks demonstrate the superiority of VMT-Adapter(-Lite), achieving a 3.96% (1.34%) relative improvement compared to single-task full fine-tuning, while utilizing merely ~ 1% (0.36%) trainable parameters of the pre-trained model.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background
  • 4 Method
  • 4.1 VMT-Adapter
  • 4.2 A More Lightweight VMT-Adapter
  • 4.3 Discussion
  • 5 Experiment
  • 5.1 Experimental Settings
  • 5.2 Baselines
  • 5.3 Performance Comparisons
  • 5.4 Ablation Studies
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — VMT-Adapter jointly models task-generic and task-specific knowledge

    model/method

    VMT-Adapter is a multi-task adapter inserted into transformer layers while the pre-trained transformer parameters remain frozen. For an input feature xℓ∈Rdx_\ell\in\mathbb{R}^{d} from layer ℓ\ell, all tasks share a down-projection Wdown∈Rd×kW_{\mathrm{down}}\in\mathbb{R}^{d\times k}, a ReLU activation, and an up-projection Wup∈Rk×dW_{\mathrm{up}}\in\mathbb{R}^{k\times d}. The shared projections produce a cross-task representation:

    F=ReLU⁡(xℓWdown)Wup.F=\operatorname{ReLU}(x_\ell W_{\mathrm{down}})W_{\mathrm{up}}.

    A scalar gating ratio ss partitions FF into a task-generic component and a task-specific component. For task i∈{1,…,T}i\in\{1,\ldots,T\}, an independent knowledge-extraction module applies a learned scale vector αi∈Rd\alpha_i\in\mathbb{R}^{d} and shift vector γi∈Rd\gamma_i\in\mathbb{R}^{d} to the task-specific portion:

    fti=αi⊙(sF)+γi,f_{t_i}=\alpha_i\odot(sF)+\gamma_i,

    where ⊙\odot denotes elementwise scaling. The shared projections enable cross-task interaction, while the independent (αi,γi)(\alpha_i,\gamma_i) pairs preserve task-specific information. Each task-specific module contributes only 2d2d parameters.

  2. Knowl 2 — VMT-Adapter-Lite factorizes the shared projections with Kronecker products

    model/method

    VMT-Adapter-Lite reduces the parameters in the two shared projections by representing them as sums of Kronecker products. Let mm be a positive integer that divides dd and kk, let Ai∈Rm×mA_i\in\mathbb{R}^{m\times m} be shared between the down- and up-projections, and let Bidown∈Rd/m×k/mB_i^{\mathrm{down}}\in\mathbb{R}^{d/m\times k/m} and Biup∈Rk/m×d/mB_i^{\mathrm{up}}\in\mathbb{R}^{k/m\times d/m} be projection-specific matrices. The projections are constructed as

    Wdown=∑i=1mAi⊗Bidown,Wup=∑i=1mAi⊗Biup.W_{\mathrm{down}}=\sum_{i=1}^{m}A_i\otimes B_i^{\mathrm{down}},\qquad W_{\mathrm{up}}=\sum_{i=1}^{m}A_i\otimes B_i^{\mathrm{up}}.

    The shared matrices contribute m3m^3 parameters, and each projection has dk/mdk/m parameters instead of dkdk. Thus, for a transformer with LL layers and TT tasks, the adapter-related parameter count is m3+2Ldk/m+2Tdm^3+2Ldk/m+2Td. The projection parameters are reduced by approximately a factor of mm relative to VMT-Adapter, although the additional m3m^3 shared parameters remain.

  3. Knowl 3 — VMT-Adapter supplies multi-scale generic and task-specific features to every decoder

    model/method

    In the dense prediction architecture, VMT-Adapter is placed in parallel with the MLP inside each Swin Transformer layer. For each of the four hierarchical transformer blocks, the task-generic representation is added back to the encoder stream and propagated to later blocks, whereas the task-specific representation for each task is sent directly to that task's decoder. The decoder for task ii receives four multi-scale features:

    Fims=[(fg1+fti1), (fg2+fti2), (fg3+fti3), (fg4+fti4)],F_i^{\mathrm{ms}}=\left[(f_g^1+f_{t_i}^1),\,(f_g^2+f_{t_i}^2),\,(f_g^3+f_{t_i}^3),\,(f_g^4+f_{t_i}^4)\right],

    where fgjf_g^j and ftijf_{t_i}^j are respectively the task-generic and task-ii-specific representations from transformer block jj. The architecture diagram on page 4 depicts this separation: the shared generic pathway traverses the encoder once, while task-specific branches feed the separate task decoders.

  4. Knowl 4 — VMT-Adapter has task-number-independent encoder computation

    theoretical result

    Let TT be the number of tasks, LL the number of transformer layers, dd the feature dimension, k=d/ρk=d/\rho the adapter bottleneck dimension, and ρ\rho the down-projection ratio. The adapter-related trainable parameter counts are

    Multiple Adapter=2TLd2ρ,Shared Adapter=2Ld2ρ,VMT-Adapter=2Ld2ρ+2Td,VMT-Adapter-Lite=m3+2Ld2mρ+2Td.\begin{aligned} \text{Multiple Adapter}&=\frac{2TLd^2}{\rho},\\ \text{Shared Adapter}&=\frac{2Ld^2}{\rho},\\ \text{VMT-Adapter}&=\frac{2Ld^2}{\rho}+2Td,\\ \text{VMT-Adapter-Lite}&=m^3+\frac{2Ld^2}{m\rho}+2Td. \end{aligned}

    Multiple Adapter evaluates separate encoder paths for all TT tasks, giving training and inference efficiency O(T)O(T). VMT-Adapter and VMT-Adapter-Lite propagate only task-generic features through the encoder and compute task-specific representations in parallel, giving approximately O(1)O(1) training and inference efficiency with respect to task number. The task-specific term 2Td2Td is negligible relative to the shared projections when T≪kT\ll k, but it can become dominant for extremely large task collections.

  5. Knowl 5 — VMT-Adapter improves multi-task transfer with roughly one percent of encoder parameters

    data/table

    On PASCAL-Context with a SwinTransformer-Tiny encoder, the paper compares multi-task parameter-efficient methods on 21-class semantic segmentation, 7-class human-part segmentation, saliency detection, and surface-normal estimation. Segmentation, human-part segmentation, and saliency use mIoU, for which higher is better; surface normals use mean error (mErr), for which lower is better. The final column is relative improvement over Single-task Full Fine-tuning. Trainable-parameter counts are reported for the encoder alone and for the complete model, in millions. The results table on page 6 shows that VMT-Adapter gives the best average improvement while using far fewer trainable parameters than the strongest competing methods.

    Could not parse LaTeX table

    VMT-Adapter improves the average result by 3.96%3.96\% over Single-task Full Fine-tuning using 1.131.13 million encoder parameters, approximately 1%1\% of the pre-trained encoder. VMT-Adapter-Lite uses only 0.400.40 million encoder parameters, approximately 0.36%0.36\% of the encoder, while retaining a 1.34%1.34\% average improvement. Hyperformer obtains a 2.64%2.64\% improvement but uses 72.7772.77 million trainable encoder parameters, and Polyhistor obtains 2.34%2.34\% with 6.416.41 million.

  6. Knowl 6 — Shared VMT-Adapter gradients exhibit fewer task conflicts

    theoretical result

    For task ii, let Li(θsh,θspi)L_i(\theta_{\mathrm{sh}},\theta_{\mathrm{sp}}^i) be its loss, where θsh\theta_{\mathrm{sh}} are the projections shared by all tasks and θspi\theta_{\mathrm{sp}}^i are task-ii-specific parameters. Define the shared-parameter gradient gi=∇θshLi(θsh,θspi)g_i=\nabla_{\theta_{\mathrm{sh}}}L_i(\theta_{\mathrm{sh}},\theta_{\mathrm{sp}}^i). After a sufficiently small gradient step of size η\eta using task ii, the first-order change in task jj's loss is

    ΔLj=−ηgi⋅gj+o(η).\Delta L_j=-\eta g_i\cdot g_j+o(\eta).

    Therefore, gi⋅gj>0g_i\cdot g_j>0 means that updating the shared parameters for task ii has a positive first-order effect on task jj, while a negative dot product indicates interference. The paper measures the relationship using

    cos⁡ϕij=gi⋅gj∥gi∥ ∥gj∥.\cos\phi_{ij}=\frac{g_i\cdot g_j}{\lVert g_i\rVert\,\lVert g_j\rVert}.

    Across training iterations for the four PASCAL-Context tasks, VMT-Adapter has a higher probability of cos⁡ϕij>0\cos\phi_{ij}>0 and a lower probability of cos⁡ϕij<0\cos\phi_{ij}<0 than Shared Adapter. This indicates increased cross-task synergy and reduced gradient conflict in the shared projections.

  7. Knowl 7 — PASCAL-Context benchmark and common training protocol

    experimental setup

    The evaluation uses PASCAL-Context, with 4,998 training images and 5,105 test images. The four simultaneous dense prediction tasks are 21-class semantic segmentation, 7-class human-part segmentation, surface-normal estimation, and saliency detection. mIoU is used for the two segmentation tasks and saliency, while mean error is used for surface normals.

    All methods use an ImageNet-pre-trained Swin Transformer encoder and separate task decoders based on the SegFormer All-MLP decoder. The encoder and decoder structure is held fixed across comparisons. Experiments are run in PyTorch on four NVIDIA V100 GPUs with batch size 12 for 60 epochs. Optimization uses Adam with learning rate 10−410^{-4} and weight decay 10−410^{-4}; the learning rate is linearly decreased over training. The baselines include full fine-tuning, decoder-only fine-tuning, Multiple and Shared variants of BitFit, Relative Bias, LoRA, Adapter, and Low-rank Adapter, as well as Hyperformer and Polyhistor.

  8. Knowl 8 — The down-projection ratio controls the VMT-Adapter efficiency–accuracy trade-off

    data/table

    The paper varies the down-projection ratio ρ=d/k\rho=d/k for VMT-Adapter and the Kronecker-sharing dimension mm for VMT-Adapter-Lite using SwinTransformer-Tiny. The table on page 7 reports encoder/all trainable parameters in millions, the four task metrics, and relative improvement over Single-task Full Fine-tuning.

    Could not parse LaTeX table

    For VMT-Adapter, the tested ratios yield improvements from 3.02%3.02\% to 4.27%4.27\% while using 0.5%0.5\% to 3.9%3.9\% of the trainable parameters of Single-task Full Fine-tuning. For VMT-Adapter-Lite, increasing mm decreases parameters but also decreases performance; m=3m=3 provides the reported balance between accuracy and parameter count.

  9. Knowl 9 — VMT-Adapter generalizes across encoder and decoder architectures

    empirical result

    The paper evaluates VMT-Adapter with a larger ImageNet-pre-trained SwinTransformer-Base encoder and with an alternative HRNet-V2 high-resolution decoder. Relative improvements are measured against Single-task Full Fine-tuning on the same four PASCAL-Context tasks.

    With the SwinTransformer-Tiny encoder and SegFormer All-MLP decoder, VMT-Adapter achieves a 3.96%3.96\% average improvement. Replacing the encoder with SwinTransformer-Base increases the average improvement to 7.10%7.10\%, indicating stronger gains with the larger backbone. With SwinTransformer-Tiny retained, replacing the All-MLP decoder by the HRNet-V2 high-resolution decoder yields a 2.41%2.41\% average improvement. The comparison charts on page 7 therefore support applicability to different encoder scales and decoder designs rather than only to the original architecture pairing.

  10. Knowl 10 — VMT-Adapter has a stated scalability limitation for extremely many tasks

    limitation

    VMT-Adapter's task-specific knowledge-extraction modules add 2Td2Td parameters for TT tasks and feature dimension dd. Although this term is negligible for the task counts evaluated in the paper, the authors note that when the number of tasks reaches thousands, the task-specific parameters can exceed the shared projection parameters. Thus, the claimed parameter efficiency is strongest for moderate task counts and is not guaranteed to remain dominant in extremely large-task regimes.

Coverage note — No substantial contributed material was omitted; the adapter designs, efficiency analysis, gradient analysis, benchmark, main comparisons, ablations, generalization tests, and stated limitation are represented.

References

  1. 1.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  2. 2.Bruggemann, D.; Kanakis, M.; Obukhov, A.; Georgoulis, S.; and Van Gool, L. 2021. Exploring relational context for multi-task dense prediction. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  3. 3.Chen, S.; Ge, C.; Tong, Z.; Wang, J.; Song, Y.; Wang, J.; and Luo, P. 2022. Adaptformer: Adapting vision transformers for scalable visual recognition. In Advances in Neural Information Processing Systems (NeurIPS).
  4. 4.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).
  5. 5.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR).
  6. 6.Gao, Y.; Ma, J.; Zhao, M.; Liu, W.; and Yuille, A. L. 2019. Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  7. 7.He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022a. Towards a Unified View of Parameter-Efficient Transfer Learning. In Proceedings of the International Conference on Learning Representations (ICLR).
  8. 8.He, K.; Chen, X.; Xie, S.; Li, Y.; Dollar, P.; and Girshick, R. 2022b. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  9. 9.He, K.; Girshick, R.; and Dollar, P. 2019. Rethinking imagenet pre-training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  10. 10.He, X.; Li, C.; Zhang, P.; Yang, J.; and Wang, X. E. 2022c. Parameter-efficient fine-tuning for vision transformers. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  11. 11.Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning (ICML).
  12. 12.Hu, E.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, L.; and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR).
  13. 13.Karimi Mahabadi, R.; Ruder, S.; Dehghani, M.; and Henderson, J. 2021. Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  14. 14.Li, W.-H.; Liu, X.; and Bilen, H. 2022. Learning multiple dense prediction tasks from partially annotated data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  15. 15.Liu, S.; Johns, E.; and Davison, A. J. 2019. End-to-end multi-task learning with attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  16. 16.Liu, Y.-C.; Ma, C.-Y.; Tian, J.; He, Z.; and Kira, Z. 2022. Polyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision Tasks. In Advances in Neural Information Processing Systems (NeurIPS).
  17. 17.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  18. 18.mahabadi, R. K.; Henderson, J.; and Ruder, S. 2021. Compacter: Efficient Low-Rank Hypercomplex Adapter Layers. In Advances in Neural Information Processing Systems (NeurIPS).
  19. 19.Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015. Imagenet large scale visual recognition challenge. In International Journal of Computer Vision (IJCV).
  20. 20.Sun, K.; Zhao, Y.; Jiang, B.; Cheng, T.; Xiao, B.; Liu, D.; Mu, Y.; Wang, X.; Liu, W.; and Wang, J. 2019. High-resolution representations for labeling pixels and regions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  21. 21.Vandenhende, S.; Georgoulis, S.; Gansbeke, W. V.; Proesmans, M.; Dai, D.; and Gool, L. V. 2021. Multi-Task Learning for Dense Prediction Tasks: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
  22. 22.Vandenhende, S.; Georgoulis, S.; and Van Gool, L. 2020. MTI-Net: Multi-Scale Task Interaction Networks for Multi-Task Learning. In Proceedings of the European Conference on Computer Vision (ECCV).
  23. 23.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS).
  24. 24.Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J. M.; and Luo, P. 2021. SegFormer: Simple and efficient design for semantic segmentation with transformers. In Advances in Neural Information Processing Systems (NeurIPS).
  25. 25.Xin, Y.; Du, J.; Wang, Q.; Yan, K.; and Ding, S. 2023. MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning. In arXiv preprint arXiv:2312.08636.
  26. 26.Xu, Y.; Yang, Y.; and Zhang, L. 2023. DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense Prediction. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  27. 27.Yi-Lin Sung, M. B., Jaemin Cho. 2022. VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  28. 28.Yu, B. X.; Chang, J.; Liu, L.; Tian, Q.; and Chen, C. W. 2022. Towards a Unified View on Visual Parameter-Efficient Transfer Learning. In Proceedings of the International Conference on Learning Representations (ICLR).
  29. 29.Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022. BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  30. 30.Zhang, X.; Zhou, L.; Li, Y.; Cui, Z.; Xie, J.; and Yang, J. 2021. Transfer vision patterns for multi-task pixel learning. In Proceedings of the ACM Conference on Multimedia (MM).

Citation

MLA
Xin, Y., et al. “VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding”. arXiv, 2023, http://arxiv.org/abs/2312.08733v2.
APA
Xin, Y., Du, J., Wang, Q., Lin, Z., & Yan, K. (2023). VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding. arXiv. http://arxiv.org/abs/2312.08733v2
Chicago
Xin, Y., J. Du, Q. Wang, Z. Lin, and K. Yan. 2023. “VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding”. arXiv. http://arxiv.org/abs/2312.08733v2.
Harvard
Xin, Y. et al. (2023) “VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.08733v2.
Vancouver
1. Xin Y, Du J, Wang Q, Lin Z, Yan K (2023) VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding. arXiv

BibTeX

@article{xin2023vmt,
  title = {VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding},
  author = {Xin, Yi and Du, Junlong and Wang, Qiang and Lin, Zhiwen and Yan, Ke},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.08733v2},
  eprint = {2312.08733}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF