MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning

Yi XinJunlong DuQiang WangKe YanShouhong Ding

article2024AAAI95 citations

Presents a parameter-efficient multi-task prompt learning framework for CLIP that synchronizes vision and text prompts using shared source representations and gradient-driven task grouping, surpassing full fine-tuning with only 0.09% trainable parameters.

Listen

Deploying computer vision models across multiple related tasks typically requires substantial computational resources. Conventional multi-task learning models rely on complex, task-specific decoders that scale linearly in cost with every added task. While large pre-trained vision-language foundation models like CLIP provide strong zero-shot recognition capabilities without dedicated decoders, fine-tuning their entire architecture (roughly 150 million parameters) is computationally expensive and prone to overfitting when downstream training data is scarce. Existing parameter-efficient tuning techniques alleviate this burden but often focus only on single modalities or single tasks, disrupting the pre-trained alignment between image and text features.

The article introduces and evaluates the Multi-modal Alignment Prompt (MmAP), a parameter-efficient fine-tuning framework designed to adapt pre-trained vision-language models to cross-domain multi-task image recognition. The core objective is to achieve high multi-task accuracy while updating only a tiny fraction of total model parameters.

To evaluate this approach, the authors designed a unified benchmarking setup using CLIP with a Vision Transformer backbone on two standard cross-domain multi-task datasets: Office-Home (four tasks, 65 categories) and MiniDomainNet (four tasks, 126 categories). They evaluated performance under limited data scenarios ranging from 1% to 20% of available data (equivalent to 3 to 12 training examples per class). The proposed method generates coupled text and visual prompts simultaneously from a single shared source prompt using Kronecker matrix operations, keeping the base foundation model frozen. It clusters tasks based on gradient cosine similarity to share prompts across complementary tasks while assigning distinct prompts to preserve task-specific nuances.

The experiments show that MmAP matches or exceeds full model fine-tuning accuracy while updating only 0.13 million parameters—approximately 0.09% of the full 149.62 million parameters. On Office-Home, the framework achieved an average accuracy of 86.5% with 10% data and 87.8% with 20% data, outperforming existing single-modality and multi-modal prompt baselines. On the more challenging MiniDomainNet dataset, full model fine-tuning suffered from overfitting and fell behind parameter-efficient approaches, whereas the proposed method attained the highest overall accuracy across both 1% and 2% splits (84.9% and 86.1%, respectively). Furthermore, ablation experiments confirmed that gradient-based task grouping outperformed both random grouping and monolithic single-group training by roughly 0.4% to 0.85%.

These findings indicate that maintaining direct alignment between language and visual modalities is superior to tuning one modality in isolation or directly altering pre-trained weights. For enterprise deployments, this provides an efficient route to scale vision-language systems across multiple operational domains, cutting down model storage, training time, and compute overhead without sacrificing accuracy.

Organizations adapting foundational vision-language models to multi-task visual recognition should prioritize multi-modal prompt tuning frameworks over full fine-tuning or bias-only modifications. Implementation pipelines should incorporate gradient-driven task similarity checks before training to prevent negative transfer between conflicting tasks. A recommended operational threshold is to provide at least three training examples per class, as performance drops below zero-shot baselines under extreme single-example (1-shot) conditions.

The evaluated scope is limited to cross-domain image classification tasks using the ViT-B/16 architecture and does not cover dense visual tasks like segmentation or object detection. Confidence in the reported results is high for few-shot image classification, but stakeholders should validate performance on their specific domain workflows and task distributions before wide-scale deployment.

Cover for MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning

Abstract

Multi-Task Learning (MTL) is designed to train multiple correlated tasks simultaneously, thereby enhancing the performance of individual tasks. Typically, a multi-task network structure consists of a shared backbone and task-specific decoders. However, the complexity of the decoders increases with the number of tasks. To tackle this challenge, we integrate the decoder-free vision-language model CLIP, which exhibits robust zero-shot generalization capability. Recently, parameter-efficient transfer learning methods have been extensively explored with CLIP for adapting to downstream tasks, where prompt tuning showcases strong potential. Nevertheless, these methods solely fine-tune a single modality (text or visual), disrupting the modality structure of CLIP. In this paper, we first propose Multi-modal Alignment Prompt (MmAP) for CLIP, which aligns text and visual modalities during fine-tuning process. Building upon MmAP, we develop an innovative multi-task prompt learning framework. On the one hand, to maximize the complementarity of tasks with high similarity, we utilize a gradient-driven task grouping method that partitions tasks into several disjoint groups and assign a group-shared MmAP to each group. On the other hand, to preserve the unique characteristics of each task, we assign an task-specific MmAP to each task. Comprehensive experiments on two large multi-task learning datasets demonstrate that our method achieves significant performance improvements compared to full fine-tuning while only utilizing approximately ~ 0.09% of trainable parameters.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Contrastive Language-Image Pre-training
  • 3.2 Multi-modal Alignment Prompt
  • 3.3 Multi-Task Prompt Learning Framework
  • 4 Experiment
  • 4.1 Benchmark Setting
  • 4.2 Experiment Results
  • 4.3 Ablation Study
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Multi-Modal Alignment Prompt Generation via Kronecker Product

    model/method

    Multi-Modal Alignment Prompt (MmAP) aligns visual and textual modalities when adapting a frozen vision-language model (specifically CLIP) to downstream tasks. Let bb denote prompt length, dld_l denote the dimension of text tokens, and dvd_v denote the dimension of image patch tokens. MmAP initializes a compact source prompt matrix Ps∈Rm×nP_s \in \mathbb{R}^{m \times n} and two modality-specific scaling matrices Ml∈Rbm×dlnM_l \in \mathbb{R}^{\frac{b}{m} \times \frac{d_l}{n}} and Mv∈Rbm×dvnM_v \in \mathbb{R}^{\frac{b}{m} \times \frac{d_v}{n}}.

    The text prompt Pl∈Rb×dlP_l \in \mathbb{R}^{b \times d_l} and visual prompt Pv∈Rb×dvP_v \in \mathbb{R}^{b \times d_v} are generated simultaneously via the Kronecker product (⊗\otimes):

    Pl=Ml⊗Ps=(Ml11Ps⋯Ml1nPs⋮⋱⋮Mlm1Ps⋯MlmnPs)P_l = M_l \otimes P_s = \begin{pmatrix} M_{l_{11}} P_s & \cdots & M_{l_{1n}} P_s \\ \vdots & \ddots & \vdots \\ M_{l_{m1}} P_s & \cdots & M_{l_{mn}} P_s \end{pmatrix}

    Pv=Mv⊗Ps=(Mv11Ps⋯Mv1nPs⋮⋱⋮Mvm1Ps⋯MvmnPs)P_v = M_v \otimes P_s = \begin{pmatrix} M_{v_{11}} P_s & \cdots & M_{v_{1n}} P_s \\ \vdots & \ddots & \vdots \\ M_{v_{m1}} P_s & \cdots & M_{v_{mn}} P_s \end{pmatrix}

    Using the Kronecker product preserves the underlying information in PsP_s across both modalities to enforce cross-modal alignment while reducing the number of trainable prompt parameters per layer from K(dl+dv)K(d_l + d_v) to mn+K(dl+dv)mnmn + \frac{K(d_l + d_v)}{mn}, where KK is the number of transformer layers.

  2. Knowl 2 — Gradient-Driven Task Grouping for Vision Multi-Task Learning

    model/method

    In cross-domain multi-task prompt learning, joint training of similar tasks yields mutual benefits, whereas joint training of conflicting tasks degrades performance. To determine which tasks should share prompt representations, task similarity is evaluated by measuring the alignment of gradients computed on a temporary global shared prompt PglbP_{glb} while keeping the underlying vision-language model parameters frozen.

    For any two downstream tasks TiT_i and TjT_j, their pairwise similarity is defined by the dot product of their task-specific loss gradients with respect to PglbP_{glb}:

    sim(Ti,Tj)=∇PglbLTi(Pglb)⋅∇PglbLTj(Pglb)\text{sim}(T_i, T_j) = \nabla_{P_{glb}} \mathcal{L}_{T_i}(P_{glb}) \cdot \nabla_{P_{glb}} \mathcal{L}_{T_j}(P_{glb})

    where LT\mathcal{L}_T represents the cross-entropy loss for task TT. A condition where sim(Ti,Tj)>0\text{sim}(T_i, T_j) > 0 indicates positive transfer (mutual gain). To ensure estimation stability, the gradient similarities are averaged across multiple training snapshots during the training of PglbP_{glb}. The set of all NN tasks {Ti}i=1N\{T_i\}_{i=1}^N is subsequently partitioned into disjoint groups GG that maximize overall intra-group task similarity.

  3. Knowl 3 — Hierarchical Prompt Injection and Multi-Task Optimization

    model/method

    The multi-task prompt learning framework structures prompt parameters into two distinct hierarchies for each transformer layer k∈{1,…,K}k \in \{1, \dots, K\}:

    1. A group-shared prompt set for each task group GG, parameterized by source prompts PGkP_G^k, text scaling matrices MGlkM_{Gl}^k, and visual scaling matrices MGvkM_{Gv}^k.
    2. A task-specific prompt set for each individual task T∈GT \in G, parameterized by source prompts PTkP_T^k, text scaling matrices MTlkM_{Tl}^k, and visual scaling matrices MTvkM_{Tv}^k.

    For each transformer layer kk, prompts are generated via Kronecker products and prepended alongside input tokens to the language transformer layer Lk\mathcal{L}_k and vision transformer layer Vk\mathcal{V}_k:

    […,Wk]=Lk([PGk⊗MGlk,PTk⊗MTlk,Wk−1])[\dots, W_k] = \mathcal{L}_k([P_G^k \otimes M_{Gl}^k, P_T^k \otimes M_{Tl}^k, W_{k-1}])

    [ck,…,Ek]=Vk([ck−1,PGk⊗MGvk,PTk⊗MTvk,Ek−1])[c_k, \dots, E_k] = \mathcal{V}_k([c_{k-1}, P_G^k \otimes M_{Gv}^k, P_T^k \otimes M_{Tv}^k, E_{k-1}])

    where Wk−1W_{k-1} represents intermediate text token embeddings, Ek−1E_{k-1} represents intermediate image patch embeddings, ck−1c_{k-1} is the image class token, and [⋅,⋅][\cdot, \cdot] denotes sequence concatenation.

    The group-shared prompts are updated collectively across all tasks within group GG using the accumulated loss:

    L(PG,MGl,MGv)=∑T∈GLT(PG,MGl,MGv)\mathcal{L}(P_G, M_{Gl}, M_{Gv}) = \sum_{T \in G} \mathcal{L}_T(P_G, M_{Gl}, M_{Gv})

    while task-specific prompts are updated individually for each task TT via L(PT,MTl,MTv)=LT(PT,MTl,MTv)\mathcal{L}(P_T, M_{Tl}, M_{Tv}) = \mathcal{L}_T(P_T, M_{Tl}, M_{Tv}), where LT\mathcal{L}_T is the task cross-entropy loss.

  4. Knowl 4 — Evaluation on Office-Home Multi-Task Benchmark

    data/table

    The Multi-modal Alignment Prompt for Multi-Task Learning (MmAP-MT) was benchmarked on the Office-Home cross-domain dataset (containing Art, Clipart, Product, and Real World tasks across 65 classes) against zero-shot CLIP, single-task learning (STL) baselines, multi-task learning (MTL) full fine-tuning, and parameter-efficient transfer learning methods (BitFit, CLIP-Adapter, CoOp-MT, VPT-MT, MaPLe-MT) using CLIP ViT-B/16.

    Single Task Learning Multi Task Learning
    Method Zero Full FT CoOp VPT MaPLe Full FT Adapter BitFit CoOp VPT MaPLe Ours
    10% split
    Art 82.9 84.9 84.2 83.7 84.4 85.8 82.3 79.1 84.3 84.0 84.8 85.7
    Clipart 68.3 75.4 72.6 70.5 72.8 76.3 71.7 67.8 73.0 72.4 73.3 76.3
    Product 89.3 91.6 92.4 90.9 92.2 92.1 90.8 86.7 92.7 91.7 92.7 92.9
    Real 90.1 89.8 90.5 89.2 90.4 90.2 89.2 85.9 90.7 90.6 90.8 90.9
    Avg. 82.6 85.4 84.9 83.6 85.0 86.1 83.5 79.9 85.2 84.7 85.4 86.5
    20% split
    Art 84.6 87.1 85.6 85.4 85.9 87.4 83.2 81.7 86.0 85.9 86.3 88.2
    Clipart 68.2 77.9 74.5 71.4 74.2 78.8 75.4 69.6 73.9 72.3 74.2 77.1
    Product 89.5 91.9 93.0 91.5 92.8 93.0 91.7 87.2 92.9 92.1 92.9 93.5
    Real 90.7 89.8 91.8 90.9 91.8 91.9 90.6 86.7 92.0 91.7 92.0 92.4
    Avg. 83.3 86.7 86.2 84.8 86.2 87.8 85.2 85.5 86.3 85.5 86.4 87.8
    #Params - 598.48M 0.04M 0.68M 19.2M 149.62M 0.53M 0.17M 0.01M 0.17M 4.8M 0.13M

    MmAP-MT attains an average accuracy of 86.5%86.5\% on the 10% data split and 87.8%87.8\% on the 20% data split, matching or exceeding multi-task full fine-tuning (86.1%86.1\% and 87.8%87.8\%) while training only 0.13M parameters (~0.09% of the full model's 149.62M parameters).

  5. Knowl 5 — Evaluation on MiniDomainNet Multi-Task Benchmark

    data/table

    Experiments on MiniDomainNet (140,000 images across 126 classes with 4 tasks: Clipart, Painting, Real, Sketch) compare parameter-efficient multi-task methods against full fine-tuning under low-data regimes (1% and 2% splits, corresponding to 3-shot and 6-shot per class).

    Single Task Learning Multi Task Learning
    Method Zero Full FT CoOp VPT MaPLe Full FT Adapter BitFit CoOp VPT MaPLe Ours
    1% split
    Clipart 82.6 82.1 82.7 82.3 82.9 82.8 82.6 78.9 83.4 83.0 83.4 83.9
    Paint 82.3 81.8 81.8 81.7 82.0 81.5 80.4 74.7 82.3 81.9 82.5 83.5
    Real 91.2 89.1 91.9 91.6 92.0 89.1 90.9 84.2 91.3 90.1 91.4 92.2
    Sketch 79.9 77.0 77.1 78.5 78.5 77.2 78.3 72.4 79.2 78.6 79.1 79.8
    Avg. 84.0 82.5 83.4 83.5 83.9 82.7 83.0 77.6 84.0 83.4 84.1 84.9
    2% split
    Clipart 82.6 82.2 83.8 83.5 83.8 82.8 83.1 81.5 84.7 83.8 84.5 85.7
    Paint 82.3 82.1 82.5 82.4 82.7 82.1 81.5 76.8 83.2 82.2 83.6 85.0
    Real 91.2 89.2 91.9 91.5 91.6 89.3 90.6 85.9 91.7 90.5 91.9 92.3
    Sketch 80.0 77.4 79.0 79.6 79.9 77.7 78.7 74.8 80.1 79.0 80.5 81.5
    Avg. 84.0 82.7 84.3 84.2 84.5 83.0 83.4 79.8 84.9 83.9 85.1 86.1
    #Params - 598.48M 0.04M 0.68M 19.2M 149.62M 0.53M 0.17M 0.01M 0.17M 4.8M 0.13M

    Full fine-tuning suffers severe overfitting in this low-data setting (82.7%82.7\% on 1% and 83.0%83.0\% on 2%, lower than zero-shot 84.0%84.0\%). In contrast, MmAP achieves 84.9%84.9\% and 86.1%86.1\%, outperforming full fine-tuning, BitFit (77.6%,79.8%77.6\%, 79.8\%), and multi-task prompt baselines while requiring only 0.13M parameters.

  6. Knowl 6 — Ablation of Multi-Task Prompt Learning Components

    data/table

    An ablation study evaluated the relative importance of task-specific MmAP, group-shared MmAP, and the gradient-based task grouping strategy on the Office-Home dataset (10% and 20% training splits).

    Task Specific Group Shared Office-Home
    Task Group Random 10% 20%
    ✓ ×\times ×\times 85.76 86.97
    ×\times ✓ ×\times 86.05 87.29
    ✓ ×\times ✓ 85.80 86.92
    ✓ ✓ ×\times 86.48 87.77
    ✓ All in one group 86.09 87.36

    Combining task-specific prompts with gradient-grouped shared prompts yields the highest accuracy (86.48%86.48\% and 87.77%87.77\%). Replacing gradient-driven task grouping with random task grouping degrades performance by 0.68%0.68\% (at 10%) and 0.85%0.85\% (at 20%). Merging all tasks into a single group without gradient clustering reduces performance by 0.39%0.39\% and 0.41%0.41\%.

  7. Knowl 7 — Ablation on Multi-Modal Alignment Architectures: MmAP vs. MLP vs. Joint Training

    empirical result

    Ablation comparing MmAP's Kronecker product formulation to alternative multi-modal prompt tuning architectures was conducted on Office-Home (10% split) and MiniDomainNet (2% split):

    1. Joint Train: Independent visual prompts PvP_v and textual prompts PlP_l trained simultaneously without cross-modal generation constraints. Achieves 85.48%85.48\% accuracy on Office-Home (10%) and 85.51%85.51\% on MiniDomainNet (2%) with 0.24M parameters.
    2. MLP-based Generation: Textual and visual prompts generated via two multilayer perceptron (MLP) branches from a shared representation. Achieves 85.89%85.89\% on Office-Home (10%) and 85.63%85.63\% on MiniDomainNet (2%) with 3.96M parameters.
    3. MmAP (Kronecker Product): Generates PlP_l and PvP_v from PsP_s via Kronecker products. Achieves 86.43%86.43\% on Office-Home (10%) and 86.13%86.13\% on MiniDomainNet (2%) with 0.13M parameters.

    MmAP improves accuracy over the MLP approach by +0.54%+0.54\% on Office-Home while using ~30x fewer trainable parameters (0.13M vs. 3.96M), demonstrating that the Kronecker product better preserves source prompt semantics without adding parameter overhead.

  8. Knowl 8 — Low-Shot Prompt Tuning Performance and 1-Shot Limitation

    empirical result

    Evaluation of MmAP across kk-shot training budgets (k∈{1,3,6,12}k \in \{1, 3, 6, 12\} samples per class) on the four Office-Home tasks (Art, Clipart, Product, RealWorld) demonstrates a clear scaling profile:

    • Under the 1-shot setting (k=1k=1), all tested prompt tuning methods (CoOp-MT, VPT-MT, MaPLe-MT, MmAP) underperform zero-shot CLIP on Art and RealWorld tasks because a single sample is overly specific and causes the prompt to overfit rather than capturing a general task representation.
    • When provided with 3 or more shots per class (k≥3k \ge 3), MmAP consistently outperforms all baseline prompt methods across all four domains, widening the performance gap over zero-shot CLIP monotonically as kk increases.

Coverage note — None was omitted. All primary contributions—including the Kronecker-based MmAP formulation, the gradient-driven grouping mechanism, full multi-task framework formulation, and all experimental results/ablations—are represented.

References

  1. 1.Chen, S.; Ge, C.; Tong, Z.; Wang, J.; Song, Y.; Wang, J.; and Luo, P. 2022. Adaptformer: Adapting vision transformers for scalable visual recognition. In Advances in Neural Information Processing Systems (NeurIPS).
  2. 2.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR).
  3. 3.Fifty, C.; Amid, E.; Zhao, Z.; Yu, T.; Anil, R.; and Finn, C. 2021. Efficiently identifying task groupings for multi-task learning. In Advances in Neural Information Processing Systems (NeurIPS).
  4. 4.Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; and Qiao, Y. 2021. CLIP-Adapter: Better Vision-Language Models with Feature Adapters. In arXiv preprint arXiv:2110.04544.
  5. 5.Gao, Y.; Ma, J.; Zhao, M.; Liu, W.; and Yuille, A. L. 2019. Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  6. 6.He, Y.; Zheng, S.; Tay, Y.; Gupta, J.; Du, Y.; Aribandi, V.; Zhao, Z.; Li, Y.; Chen, Z.; Metzler, D.; et al. 2022. Hyperprompt: Prompt-based task-conditioning of transformers. In Proceedings of the International Conference on Machine Learning (ICML).
  7. 7.Hu, E. J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR).
  8. 8.Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning (ICML).
  9. 9.Jia, M.; Tang, L.; Chen, B.; Cardie, C.; Belongie, S. J.; Hariharan, B.; and Lim, S. 2022a. Visual prompt tuning. In Proceedings of the European Conference on Computer Vision (ECCV).
  10. 10.Jia, M.; Tang, L.; Chen, B.-C.; Cardie, C.; Belongie, S.; Hariharan, B.; and Lim, S.-N. 2022b. Visual Prompt Tuning. In Proceedings of the European Conference on Computer Vision (ECCV).
  11. 11.Khattak, M. U.; Rasheed, H. A.; Maaz, M.; Khan, S.; and Khan, F. S. 2023. Maple: Multi-modal prompt learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  12. 12.Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. 2022. Grounded language-image pre-training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  13. 13.Liang, X.; Niu, M.; Han, J.; Xu, H.; Xu, C.; and Liang, X. 2023. Visual Exemplar Driven Task-Prompting for Unified Perception in Autonomous Driving. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  14. 14.Liu, Y.; Lu, Y.; Liu, H.; An, Y.; Xu, Z.; Yao, Z.; Zhang, B.; Xiong, Z.; and Gui, C. 2023. Hierarchical Prompt Learning for Multi-Task Learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  15. 15.Liu, Y.-C.; Ma, C.-Y.; Tian, J.; He, Z.; and Kira, Z. 2022. Polyhistor: Parameter-Efficient Multi-Task Adaptation for Dense Vision Tasks. In Advances in Neural Information Processing Systems (NeurIPS).
  16. 16.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  17. 17.Long, M.; Cao, Z.; Wang, J.; and Yu, P. S. 2017. Learning multiple tasks with multilinear relationship networks. In Advances in Neural Information Processing Systems (NeurIPS).
  18. 18.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning (ICML).
  19. 19.Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2022. Denseclip: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  20. 20.Shen, J.; Zhen, X.; Worring, M.; and Shao, L. 2021. Variational multi-task learning with gumbel-softmax priors. In Advances in Neural Information Processing Systems (NeurIPS).
  21. 21.Standley, T.; Zamir, A.; Chen, D.; Guibas, L.; Malik, J.; and Savarese, S. 2020. Which tasks should be learned together in multi-task learning? In Proceedings of the International Conference on Machine Learning (ICML).
  22. 22.Sung, Y.; Cho, J.; and Bansal, M. 2022. VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  23. 23.Venkateswara, H.; Eusebio, J.; Chakraborty, S.; and Panchanathan, S. 2017. Deep hashing network for unsupervised domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  24. 24.Wang, Q.; Du, J.; Yan, K.; and Ding, S. 2023. Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts Learning. In Proceedings of the ACM Conference on Multimedia (MM).
  25. 25.Wortsman, M.; Ilharco, G.; Kim, J. W.; Li, M.; Kornblith, S.; Roelofs, R.; Lopes, R. G.; Hajishirzi, H.; Farhadi, A.; Namkoong, H.; et al. 2022. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  26. 26.Xin, Y.; Du, J.; Wang, Q.; Lin, Z.; and Yan, K. 2023. VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding. In arXiv preprint arXiv:2312.08733.
  27. 27.Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; and Wang, X. 2022. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  28. 28.Xu, Y.; Yang, Y.; and Zhang, L. 2023. DeMT: Deformable mixer transformer for multi-task learning of dense prediction. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  29. 29.Ye, H.; and Xu, D. 2023. Taskprompter: Spatial-channel multi-task prompting for dense scene understanding. In Proceedings of the International Conference on Learning Representations (ICLR).
  30. 30.Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  31. 31.Zhang, X.; Zhou, L.; Li, Y.; Cui, Z.; Xie, J.; and Yang, J. 2021. Transfer vision patterns for multi-task pixel learning. In Proceedings of the ACM Conference on Multimedia (MM).
  32. 32.Zhong, Y.; Yang, J.; Zhang, P.; Li, C.; Codella, N.; Li, L. H.; Zhou, L.; Dai, X.; Yuan, L.; Li, Y.; et al. 2022. Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  33. 33.Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022. Learning to prompt for vision-language models. In International Journal of Computer Vision (IJCV).
  34. 34.Zhou, K.; Yang, Y.; Qiao, Y.; and Xiang, T. 2021. Domain adaptive ensemble learning. In IEEE Transactions on Image Processing (TIP).

Citation

MLA
Xin, Y., et al. “MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning”. arXiv, 2023, http://arxiv.org/abs/2312.08636v1.
APA
Xin, Y., Du, J., Wang, Q., Yan, K., & Ding, S. (2023). MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning. arXiv. http://arxiv.org/abs/2312.08636v1
Chicago
Xin, Y., J. Du, Q. Wang, K. Yan, and S. Ding. 2023. “MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning”. arXiv. http://arxiv.org/abs/2312.08636v1.
Harvard
Xin, Y. et al. (2023) “MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.08636v1.
Vancouver
1. Xin Y, Du J, Wang Q, Yan K, Ding S (2023) MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning. arXiv

BibTeX

@article{xin2023mmap,
  title = {MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning},
  author = {Xin, Yi and Du, Junlong and Wang, Qiang and Yan, Ke and Ding, Shouhong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.08636v1},
  eprint = {2312.08636}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF