MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning
Yi XinJunlong DuQiang WangKe YanShouhong Ding
Presents a parameter-efficient multi-task prompt learning framework for CLIP that synchronizes vision and text prompts using shared source representations and gradient-driven task grouping, surpassing full fine-tuning with only 0.09% trainable parameters.
Deploying computer vision models across multiple related tasks typically requires substantial computational resources. Conventional multi-task learning models rely on complex, task-specific decoders that scale linearly in cost with every added task. While large pre-trained vision-language foundation models like CLIP provide strong zero-shot recognition capabilities without dedicated decoders, fine-tuning their entire architecture (roughly 150 million parameters) is computationally expensive and prone to overfitting when downstream training data is scarce. Existing parameter-efficient tuning techniques alleviate this burden but often focus only on single modalities or single tasks, disrupting the pre-trained alignment between image and text features.
The article introduces and evaluates the Multi-modal Alignment Prompt (MmAP), a parameter-efficient fine-tuning framework designed to adapt pre-trained vision-language models to cross-domain multi-task image recognition. The core objective is to achieve high multi-task accuracy while updating only a tiny fraction of total model parameters.
To evaluate this approach, the authors designed a unified benchmarking setup using CLIP with a Vision Transformer backbone on two standard cross-domain multi-task datasets: Office-Home (four tasks, 65 categories) and MiniDomainNet (four tasks, 126 categories). They evaluated performance under limited data scenarios ranging from 1% to 20% of available data (equivalent to 3 to 12 training examples per class). The proposed method generates coupled text and visual prompts simultaneously from a single shared source prompt using Kronecker matrix operations, keeping the base foundation model frozen. It clusters tasks based on gradient cosine similarity to share prompts across complementary tasks while assigning distinct prompts to preserve task-specific nuances.
The experiments show that MmAP matches or exceeds full model fine-tuning accuracy while updating only 0.13 million parameters—approximately 0.09% of the full 149.62 million parameters. On Office-Home, the framework achieved an average accuracy of 86.5% with 10% data and 87.8% with 20% data, outperforming existing single-modality and multi-modal prompt baselines. On the more challenging MiniDomainNet dataset, full model fine-tuning suffered from overfitting and fell behind parameter-efficient approaches, whereas the proposed method attained the highest overall accuracy across both 1% and 2% splits (84.9% and 86.1%, respectively). Furthermore, ablation experiments confirmed that gradient-based task grouping outperformed both random grouping and monolithic single-group training by roughly 0.4% to 0.85%.
These findings indicate that maintaining direct alignment between language and visual modalities is superior to tuning one modality in isolation or directly altering pre-trained weights. For enterprise deployments, this provides an efficient route to scale vision-language systems across multiple operational domains, cutting down model storage, training time, and compute overhead without sacrificing accuracy.
Organizations adapting foundational vision-language models to multi-task visual recognition should prioritize multi-modal prompt tuning frameworks over full fine-tuning or bias-only modifications. Implementation pipelines should incorporate gradient-driven task similarity checks before training to prevent negative transfer between conflicting tasks. A recommended operational threshold is to provide at least three training examples per class, as performance drops below zero-shot baselines under extreme single-example (1-shot) conditions.
The evaluated scope is limited to cross-domain image classification tasks using the ViT-B/16 architecture and does not cover dense visual tasks like segmentation or object detection. Confidence in the reported results is high for few-shot image classification, but stakeholders should validate performance on their specific domain workflows and task distributions before wide-scale deployment.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. CLIP provides the foundational contrastive vision-language representation and zero-shot architecture that MmAP parameter-efficiently adapts for cross-domain multi-task recognition.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Context Optimization (CoOp) establishes the baseline framework for continuous prompt learning in frozen vision-language models like CLIP.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Visual Prompt Tuning introduces prompt token optimization directly into Vision Transformer architectures, forming the visual counterpart to textual prompt learning.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). CoCoOp extends static prompt tuning with dynamic visual conditioning, illustrating the core challenge of overfitting and cross-modal alignment that MmAP addresses across multi-task domains.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This work establishes the structural and parameter-efficient transfer learning taxonomy that underlies frozen backbone adaptation strategies like multi-modal prompting.
- Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). Prompt-based Multimodal Fusion explores interactive two-way prompting across modalities, offering key concepts for jointly coordinating vision and text prompts.
- Paper: SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer, Tu Vu et al. (2022). SPoT provides foundational methods for transferring and sharing soft prompt parameters across distinct tasks in multi-task prompt learning.
- Paper: Diversity-Aware Meta Visual Prompting, Qidong Huang et al. (2023). DAM-VP introduces cluster-based prompt partitioning for high-diversity datasets, directly relating to MmAP's task-grouping mechanisms.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). VMT-Adapter expands multi-task parameter-efficient adaptation from classification models like MmAP to multi-task dense scene understanding tasks in frozen vision transformers.
- Paper: Localizing Task Information for Improved Model Merging and Compression, Ke Wang et al. (2024). This paper presents a model merging and task-isolation strategy to mitigate multi-task interference without retraining, offering an alternative paradigm to prompt grouping.
- Paper: EMR-Merging: Tuning-Free High-Performance Model Merging, Chenyu Huang et al. (2024). EMR-Merging extends multi-task transfer learning by merging fine-tuned parameters into unified representations without additional training or hyperparameter tuning.
- Paper: Tuning LayerNorm in Attention: Towards Efficient Multi-Modal LLM Finetuning, Bingchen Zhao et al. (2024). This work investigates alternative parameter-efficient multi-modal fine-tuning by modifying LayerNorm parameters in attention blocks instead of prepending prompts.
- Paper: MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI, Kaining Ying et al. (2024). MMT-Bench provides a large-scale multi-task evaluation framework to assess and benchmark multi-modal adaptation and task relationships across complex vision-language domains.
