π-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation
Chengyue WuTeng WangYixiao GeZeyu LuRuisong ZhouYing ShanPing Luo
Proposes -Tuning, a parameter-efficient transfer framework that maps cross-modal task similarities using Fisher information embeddings and interpolates lightweight task experts to improve downstream vision, language, and multimodal performance.
Modern foundation models handle diverse vision, language, and multimodal tasks using unified architectures. However, adapting these massive models to specialized downstream tasks remains costly and inefficient. Traditional full fine-tuning requires substantial computational resources and separate model copies, while standard parameter-efficient transfer learning methods typically optimize for a single target task in isolation, ignoring valuable cross-task and cross-modal synergies that could enhance performance.
The article introduces and evaluates Predict-Interpolate Tuning (π-Tuning), a universal parameter-efficient transfer learning framework designed to transfer multimodal foundation models by identifying similar tasks and interpolating lightweight, task-specific expert modules.
The researchers evaluated the framework using pre-trained foundation models across 14 unimodal and 6 multimodal datasets spanning computer vision, natural language processing, and vision-language domains. The method first computes task embeddings using a diagonal approximation of the Fisher Information Matrix to build a scalable similarity graph across modalities. It then selects the top-ranked auxiliary task experts and combines their parameters with the target task expert through learned interpolation weights. Experiments evaluated performance across standard parameter-efficient techniques—including Adapters, Prompt Tuning, and LoRA—in full-data, few-shot, zero-shot, and out-of-distribution transfer settings.
The evaluation produced four key findings. First, π-Tuning consistently outperformed existing parameter-efficient methods and matched or exceeded full fine-tuning across multimodal and unimodal benchmarks while updating only a fraction of total parameters. Second, the performance benefits were especially pronounced in data-scarce settings: in 16-shot image classification, π-Adapter improved accuracy by an average of 4.55 percentage points over standard adapters, with individual dataset gains reaching up to 16.39 percentage points. Third, cross-modal transfer proved highly effective; vision-language tasks like image captioning served as strong auxiliary experts for pure computer vision tasks, and visual entailment boosted natural language inference. Finally, combining about two highly similar auxiliary experts provided the optimal performance gain, whereas adding dissimilar tasks introduced negative interference.
These findings demonstrate that lightweight experts trained on related tasks reside in the same optimization basin, allowing direct parameter interpolation without adding complex fusion architectures or runtime inference latency. For engineering and deployment teams, this approach substantially reduces storage footprint and training overhead while maintaining the execution throughput of standard adapter modules. Furthermore, the resulting models exhibit greater robustness against distribution shifts compared to models fine-tuned solely on a single domain.
Organizations deploying large multimodal foundation models should consider adopting task-similarity-guided expert interpolation over isolated task-level tuning, particularly when developing applications in low-data regimes. Practitioners should focus interpolation on a small set of top-ranked similar tasks (typically the top two) rather than combining broad, uncurated pools of experts.
The study's limitations include reliance on a diagonal Fisher Information Matrix approximation for task similarity, the manual tuning of the number of auxiliary experts, and empirical validation focused primarily on base and large variants of OFA and T5 backbones. While confidence in the reported efficiency gains and accuracy improvements across evaluated benchmarks is high, further validation is warranted before applying the framework to substantially larger foundation models or generative diffusion pipelines.
- Paper: ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft Prompts, Akari Asai et al. (2022). ATTEMPT establishes the multi-task prompt-mixture idea that makes π-Tuning’s transfer and interpolation of task-specific experts easier to follow.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). This unified account of adapters, prompts, and LoRA supplies the parameter-efficient tuning vocabulary and method distinctions π-Tuning builds on.
- Paper: SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer, Tu Vu et al. (2022). SPoT shows how transferring learned prompts from source tasks can improve a target task, a precursor to π-Tuning’s cross-task expert transfer.
- Paper: Parameter-Efficient Transfer Learning for NLP, Neil Houlsby et al. (2019). The adapter framework provides a foundational example of frozen-backbone, task-specific modules—the kind of lightweight experts π-Tuning interpolates.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). CLIP-Adapter grounds the multimodal setting in lightweight vision-language adaptation, helping clarify the adapter experts π-Tuning transfers across tasks.
No sufficiently relevant recommendations were found.
