π-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation

Chengyue WuTeng WangYixiao GeZeyu LuRuisong ZhouYing ShanPing Luo

article2023ICML50 citations

Proposes π\pi-Tuning, a parameter-efficient transfer framework that maps cross-modal task similarities using Fisher information embeddings and interpolates lightweight task experts to improve downstream vision, language, and multimodal performance.

Listen

Modern foundation models handle diverse vision, language, and multimodal tasks using unified architectures. However, adapting these massive models to specialized downstream tasks remains costly and inefficient. Traditional full fine-tuning requires substantial computational resources and separate model copies, while standard parameter-efficient transfer learning methods typically optimize for a single target task in isolation, ignoring valuable cross-task and cross-modal synergies that could enhance performance.

The article introduces and evaluates Predict-Interpolate Tuning (π-Tuning), a universal parameter-efficient transfer learning framework designed to transfer multimodal foundation models by identifying similar tasks and interpolating lightweight, task-specific expert modules.

The researchers evaluated the framework using pre-trained foundation models across 14 unimodal and 6 multimodal datasets spanning computer vision, natural language processing, and vision-language domains. The method first computes task embeddings using a diagonal approximation of the Fisher Information Matrix to build a scalable similarity graph across modalities. It then selects the top-ranked auxiliary task experts and combines their parameters with the target task expert through learned interpolation weights. Experiments evaluated performance across standard parameter-efficient techniques—including Adapters, Prompt Tuning, and LoRA—in full-data, few-shot, zero-shot, and out-of-distribution transfer settings.

The evaluation produced four key findings. First, π-Tuning consistently outperformed existing parameter-efficient methods and matched or exceeded full fine-tuning across multimodal and unimodal benchmarks while updating only a fraction of total parameters. Second, the performance benefits were especially pronounced in data-scarce settings: in 16-shot image classification, π-Adapter improved accuracy by an average of 4.55 percentage points over standard adapters, with individual dataset gains reaching up to 16.39 percentage points. Third, cross-modal transfer proved highly effective; vision-language tasks like image captioning served as strong auxiliary experts for pure computer vision tasks, and visual entailment boosted natural language inference. Finally, combining about two highly similar auxiliary experts provided the optimal performance gain, whereas adding dissimilar tasks introduced negative interference.

These findings demonstrate that lightweight experts trained on related tasks reside in the same optimization basin, allowing direct parameter interpolation without adding complex fusion architectures or runtime inference latency. For engineering and deployment teams, this approach substantially reduces storage footprint and training overhead while maintaining the execution throughput of standard adapter modules. Furthermore, the resulting models exhibit greater robustness against distribution shifts compared to models fine-tuned solely on a single domain.

Organizations deploying large multimodal foundation models should consider adopting task-similarity-guided expert interpolation over isolated task-level tuning, particularly when developing applications in low-data regimes. Practitioners should focus interpolation on a small set of top-ranked similar tasks (typically the top two) rather than combining broad, uncurated pools of experts.

The study's limitations include reliance on a diagonal Fisher Information Matrix approximation for task similarity, the manual tuning of the number of auxiliary experts, and empirical validation focused primarily on base and large variants of OFA and T5 backbones. While confidence in the reported efficiency gains and accuracy improvements across evaluated benchmarks is high, further validation is warranted before applying the framework to substantially larger foundation models or generative diffusion pipelines.

No sufficiently relevant recommendations were found.

Cover for π-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation

Abstract

Foundation models have achieved great advances in multi-task learning with a unified interface of unimodal and multimodal tasks. However, the potential of such multi-task learners has not been exploited during transfer learning. In this work, we present a universal parameter-efficient transfer learning method, termed Predict-Interpolate Tuning (π-Tuning), for vision, language, and vision-language tasks. It aggregates the parameters of lightweight task-specific experts learned from similar tasks to aid the target downstream task. The task similarities are predicted in a unified modality-independent space, yielding a scalable graph to demonstrate task relationships. π-Tuning has several appealing benefits. First, it flexibly explores both intra- and inter-modal transferability between similar tasks to improve the accuracy and robustness of transfer learning, especially in data-scarce scenarios. Second, it offers a systematical solution for transfer learning with multi-task prediction-and-then-interpolation, compatible with diverse types of parameter-efficient experts, such as prompt and adapter. Third, an extensive study of task-level mutual benefits on 14 unimodal and 6 multimodal datasets shows that π-Tuning surpasses fine-tuning and other parameter-efficient transfer learning methods both in full-shot and low-shot regimes. The task graph also enables an in-depth interpretable analysis of task transferability across modalities. The code will be available at https://github.com/TencentARC/pi-Tuning.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Preliminary
  • 2.2. π-Tuning
  • 3. Experiments
  • 3.1. Experimental Settings
  • 3.2. Comparison with PETL Methods
  • 3.3. Task Relationships
  • 3.4. Cross-Task Transferability
  • 3.5. Ablation Study
  • 4. Scope and Limitations
  • 5. Related Work
  • 6. Conclusion
  • References
  • A. Experimental Setups
  • A.1. Experimental Setting for VL Tasks
  • A.2. Experimental Setting for Vision Tasks
  • A.3. Experimental Setting for Language Tasks
  • B. Analysis For Relationship Between FIM and Landscape
  • B.1. Notation and preliminaries
  • B.2. Restatement of Assumption
  • B.3. Proof of Theorem 2.1

Knowls

  1. Knowl 1 — π-Tuning interpolates target and similar-task parameter-efficient experts

    model/method

    Let θ0\theta_0 be a frozen pretrained sequence-to-sequence model, τt\tau_t a target task, ϕt\phi_t its task-specific lightweight expert, and ϕ1sim,…,ϕksim\phi^{\mathrm{sim}}_1,\ldots,\phi^{\mathrm{sim}}_k experts trained on the kk most similar source tasks. π-Tuning forms one expert by a softmax-weighted interpolation:

    ϕˉ=softmax⁡(α)0ϕt+∑j=1ksoftmax⁡(α)jϕjsim,\bar\phi=\operatorname{softmax}(\alpha)_0\phi_t+\sum_{j=1}^{k}\operatorname{softmax}(\alpha)_j\phi^{\mathrm{sim}}_j,

    where α∈Rk+1\alpha\in\mathbb{R}^{k+1} contains trainable interpolation logits and index 00 denotes the target expert. The experts are first trained individually; after selecting the source experts, the interpolated expert is tuned on τt\tau_t. The paper’s default setting tunes both α\alpha and the selected experts’ parameters. The resulting expert has the same inference-time parameter dimensions as the original parameter-efficient expert, so interpolation adds no inference latency. On RefCOCO, the reported π-Adapter training cost was 17 A100 GPU-hours, with 18.60 samples/second throughput, 2.60% deployment parameters, and 7.42% training parameters; the Adapter figures were 344 GPU-hours, 18.60 samples/second, 2.60%, and 2.60%, respectively.

  2. Knowl 2 — Diagonal empirical Fisher embeddings provide cross-modal task similarities

    equation

    For task τi\tau_i with sis_i labeled examples (xij,yij)(x_i^j,y_i^j), π-Tuning represents the task using the diagonal of the empirical Fisher information matrix of its parameter-efficient expert ϕi\phi_i. The per-example gradient is taken with respect to expert parameters while the pretrained model remains frozen:

    Fi=diag⁡ ⁣(1si∑j=1sigij(gij)⊤),gij=∇ϕilog⁡Pϕi(yij∣xij).F_i=\operatorname{diag}\!\left(\frac{1}{s_i}\sum_{j=1}^{s_i}g_i^j(g_i^j)^\top\right),\qquad g_i^j=\nabla_{\phi_i}\log P_{\phi_i}(y_i^j\mid x_i^j).

    Here Pϕi(y∣x)P_{\phi_i}(y\mid x) is the model’s conditional output distribution, and FiF_i is used as a vector embedding; retaining only diagonal entries avoids the cost of the full matrix. The method assumes correlations between different expert parameters are not essential for this representation. It computes cosine similarity between task embeddings, ranks candidate source tasks, and selects the top kk. Because each task is embedded in the same parameter space, the procedure can compare vision, language, and vision-language tasks and add new tasks to the candidate pool.

  3. Knowl 3 — Similar Fisher information can bound the distance between task optima

    theoretical result

    Consider two cross-entropy task losses L1L_1 and L2L_2 over the same n0n_0-dimensional parameter-efficient expert ϕ\phi. Let ϕ0\phi_0 be their common initialization, and let ϕ1\phi_1 and ϕ2\phi_2 be the respective local minima reached from it. Write FϕiF^i_\phi for task ii’s Fisher information matrix at ϕ\phi, treated as the negative loss Hessian in the paper’s analysis. The paper’s bound applies when: (1) the initialization gradients satisfy ∥∇L1(ϕ0)−∇L2(ϕ0)∥2≤C1\|\nabla L_1(\phi_0)-\nabla L_2(\phi_0)\|_2\le C_1; (2) along the segment from ϕ0\phi_0 to ϕ1\phi_1, ∥Fϕ1−Fϕ2∥2≤C2\|F^1_\phi-F^2_\phi\|_2\le C_2; and (3) along the segment from ϕ1\phi_1 to ϕ2\phi_2, task 2’s Fisher matrix remains sufficiently close to its value at ϕ1\phi_1 to keep the integrated matrix nonsingular. More precisely, for some 0<C3<10<C_3<1 the third condition is ∥Fϕ2−Fϕ12∥1≤(1−C3)/(n0∥(Fϕ12)−1∥2)\|F^2_\phi-F^2_{\phi_1}\|_1\le(1-C_3)/(n_0\|(F^2_{\phi_1})^{-1}\|_2). Under these conditions,

    ∥ϕ1−ϕ2∥2≤cC3(C1+C2R0),c=∥(Fϕ12)−1∥2,R0=∥ϕ1−ϕ0∥2.\|\phi_1-\phi_2\|_2\le\frac{c}{C_3}(C_1+C_2R_0),\qquad c=\|(F^2_{\phi_1})^{-1}\|_2,\quad R_0=\|\phi_1-\phi_0\|_2.

    Thus, under the stated gradient and Fisher-stability conditions, similar Fisher information along the relevant parameter paths constrains how far apart the two task-specific local minima can be.

  4. Knowl 4 — π-Adapter is competitive across the evaluated vision-language benchmarks

    empirical result

    On OFA-Large, π-Adapter generally improved on the single-task Adapter across referring-expression comprehension, visual entailment, image captioning, and visual question answering. For RefCOCO, π-Adapter versus Adapter scored 90.49 versus 90.05 on val, 92.93 versus 92.42 on testA, and 85.91 versus 84.83 on testB. For RefCOCO+, the corresponding scores were 84.92 versus 84.50, 90.03 versus 89.66, and 77.91 versus 77.26; for RefCOCOg they were 86.60 versus 85.48 on val-u and 86.92 versus 85.88 on test-u. On SNLI-VE, dev/test accuracy was 90.16/90.01 versus 90.04/89.59. On COCO captioning, the reported BLEU@4, METEOR, CIDEr, and SPICE scores were 41.70, 31.40, 140.7, and 24.50, compared with 41.80, 31.30, 140.6, and 24.50 for Adapter. On VQA, test-dev/test-std scores were 78.78/78.82 versus 78.55/78.62. The captioning results are mixed across metrics; overall, the measurements show gains or comparable performance on most reported splits and metrics, without an inference-throughput penalty.

  5. Knowl 5 — π-Tuning improves full-data vision and language task scores

    empirical result

    Across eight vision classification tasks on OFA-Base, the mean score rose from 79.05 with Adapter to 81.47 with π-Adapter using the full training data. The largest reported gains were on Aircraft, from 45.24 to 52.42 (+7.18), and Cars, from 53.71 to 61.30 (+7.59). On six GLUE language-understanding tasks, OFA-Large’s mean score increased from 87.12 to 88.00; on RTE it rose from 72.20 to 76.53 (+4.33). With T5-Base, the corresponding six-task mean increased from 87.93 to 89.68, including an RTE increase from 73.57 to 82.86 (+9.29). These results show that the interpolation method also benefits unimodal task transfer and is not restricted to OFA’s multimodal backbone.

  6. Knowl 6 — Transfer gains are larger in the reported low-data evaluations

    empirical result

    In 16-shot image classification on OFA-Base, π-Adapter increased the mean score across eight vision tasks from 61.46 for Adapter to 66.01, a reported gain of 4.55 points. In the zero-shot GLUE evaluation on OFA-Large, where no target-task expert was available, the method used the expert from the target’s nearest-neighbor task; the mean score increased from 36.53 for OFA-Large to 42.58 for this π-Adapter transfer, a gain of 6.05 points. In both settings the paper reports greater average improvement than in the corresponding full-data comparisons, supporting the usefulness of source-task experts when target labels are limited or absent.

  7. Knowl 7 — The task graph indicates vision-language proximity to vision and a central role for captioning

    empirical result

    The Fisher-embedding task graph, constructed across vision, language, and vision-language datasets, places vision-language tasks closer to vision tasks than to language tasks. Image captioning is close to many vision and vision-language tasks, making it a central task in the graph. The paper also evaluates expert interpolation on RefCOCO, RefCOCO+, and RefCOCOg: RefCOCOg’s target expert is interpolated with other task experts using coefficients from 0 to 1 in increments of 0.05. Higher-similarity experts show more favorable interpolation behavior, and the accuracy of direct transfer and the maximum interpolation accuracy both track task similarity; the most similar experts’ best results require greater interpolation weights. The reported loss-landscape visualization likewise places similar-task checkpoints in a shared basin, where interpolation can reach lower test error than the individual checkpoints.

  8. Knowl 8 — Cross-task interpolation improves reported multi-task results and distribution-shift robustness

    empirical result

    On referring-expression comprehension, π-Adapter scored higher than Adapter on every reported split for both OFA-Base and OFA-Large. For OFA-Base, for example, RefCOCO+ val/testA/testB scores increased from 74.60/83.20/69.79 to 79.46/84.63/71.43; for OFA-Large, they increased from 83.81/89.08/76.54 to 84.77/90.31/77.68. The paper characterizes π-Tuning as outperforming direct multi-task optimization across the reported task splits, including exceeding task-specific performance in this comparison. In experiments assessing shifted test distributions on referring-expression datasets, π-Adapter had a smaller relative performance drop than domain-specific Adapter and full fine-tuning. The authors attribute this qualitative robustness to retaining knowledge from experts trained on other domains while adapting to the target.

  9. Knowl 9 — Ablations support trained source experts, a small source set, and PETL-method flexibility

    empirical result

    On RefCOCO, RefCOCO+, and RefCOCOg, randomly initialized auxiliary experts (“w/o init.”) performed worse than the original Adapter, indicating that interpolation alone was not sufficient to obtain the reported transfer gains. Optimizing only the interpolation weights (“only scale”) performed better than Adapter but slightly worse than π-Adapter, which also tunes expert parameters. The number-of-experts experiment found that performance first increased and then declined as more auxiliary experts were added; the best value was around k=2k=2, the paper’s default, with later-ranked and less-similar experts hypothesized to introduce negative transfer. π-Tuning also transferred beyond adapters: on the three referring-expression datasets, π-LoRA improved mean accuracy by 4.20 points over LoRA, and π-Prompt improved it by 1.94 points over Prompt Tuning.

  10. Knowl 10 — The evaluation leaves backbone scale and adaptive expert selection open

    limitation

    The paper evaluates π-Tuning on OFA-Base, OFA-Large, and T5-Base, and states that computational-resource limits prevented broader evaluation on larger backbones and more fields. The number of auxiliary experts is manually chosen in the reported method; although the ablation favors about two experts in its tested setting, the paper identifies adaptive selection of the number of source tasks as an open direction. The similarity measure also uses a diagonal Fisher approximation; the authors note that other computationally manageable task embeddings could be considered.

Coverage note — The appendix’s task-specific hyperparameter settings and the remaining per-split benchmark entries are omitted as replication detail; the main method, theoretical result, cross-modal analyses, principal performance findings, ablations, and stated limitations are represented.

References

  1. 1.Achille, A., Lam, M., Tewari, R., Ravichandran, A., Maji, S., Fowlkes, C. C., Soatto, S., and Perona, P. Task2vec: Task embedding for meta-learning. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6430–6439, 2019.
  2. 2.Aghajanyan, A., Gupta, A., Shrivastava, A., Chen, X., Zettlemoyer, L., and Gupta, S. Muppet: Massive multi-task representations with pre-finetuning. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 5799–5811. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.468. URL https://doi.org/10.18653/v1/2021.emnlp-main.468.
  3. 3.Alonso, H. M. and Plank, B. When is multitask learning effective? semantic sequence prediction under varying data conditions. In Lapata, M., Blunsom, P., and Koller, A. (eds.), Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, Valencia, Spain, April 3-7, 2017, Volume 1: Long Papers, pp. 44–53. Association for Computational Linguistics, 2017. doi: 10.18653/v1/e17-1005. URL https://doi.org/10.18653/v1/e17-1005.
  4. 4.Amari, S.-I. Natural gradient works efficiently in learning. Neural computation, 10(2):251–276, 1998.
  5. 5.Bingel, J. and Søgaard, A. Identifying beneficial task relations for multi-task learning in deep neural networks. In Lapata, M., Blunsom, P., and Koller, A. (eds.), Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, Valencia, Spain, April 3-7, 2017, Volume 2: Short Papers, pp. 164–169. Association for Computational Linguistics, 2017. doi: 10.18653/v1/e17-2026. URL https://doi.org/10.18653/v1/e17-2026.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  7. 7.Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  8. 8.Chen, Y., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. UNITER: universal image-text representation learning. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J. (eds.), Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXX, volume 12375 of Lecture Notes in Computer Science, pp. 104–120. Springer, 2020. doi: 10.1007/978-3-030-58577-8_7. URL https://doi.org/10.1007/978-3-030-58577-8_7.
  9. 9.Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. Uniter: Learning universal image-text representations. 2019.
  10. 10.Cho, J., Lei, J., Tan, H., and Bansal, M. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pp. 1931–1942. PMLR, 2021.
  11. 11.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  12. 12.Dietterich, T. G. Ensemble methods in machine learning. In Kittler, J. and Roli, F. (eds.), Multiple Classifier Systems, First International Workshop, MCS 2000, Cagliari, Italy, June 21-23, 2000, Proceedings, volume 1857 of Lecture Notes in Computer Science, pp. 1–15. Springer, 2000. doi: 10.1007/3-540-45014-9_1. URL https://doi.org/10.1007/3-540-45014-9_1.
  13. 13.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  14. 14.Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. Essentially no barriers in neural network energy landscape. In International conference on machine learning, pp. 1309–1318. PMLR, 2018.
  15. 15.Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning, pp. 3259–3269. PMLR, 2020.
  16. 16.Gan, Z., Chen, Y.-C., Li, L., Zhu, C., Cheng, Y., and Liu, J. Large-scale adversarial training for vision-and-language representation learning. Advances in Neural Information Processing Systems, 33:6616–6628, 2020.
  17. 17.Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. Loss surfaces, mode connectivity, and fast ensembling of dnns. Advances in neural information processing systems, 31, 2018.
  18. 18.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6904–6913, 2017.
  19. 19.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pp. 2790–2799. PMLR, 2019.
  20. 20.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  21. 21.Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., and Carion, N. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1780–1790, 2021.
  22. 22.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 6402–6413, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/9ef2ed4b7fd2c810847ffa5fa85bce38-Abstract.html.
  23. 23.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 3045–3059. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.243. URL https://doi.org/10.18653/v1/2021.emnlp-main.243.
  24. 24.Li, L. H., Yatskar, M., Yin, D., Hsieh, C., and Chang, K. Visualbert: A simple and performant baseline for vision and language. CoRR, abs/1908.03557, 2019. URL http://arxiv.org/abs/1908.03557.
  25. 25.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  26. 26.Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. GPT understands, too. CoRR, abs/2103.10385, 2021. URL https://arxiv.org/abs/2103.10385.
  27. 27.Lu, J., Goswami, V., Rohrbach, M., Parikh, D., and Lee, S. 12-in-1: Multi-task vision and language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10437–10446, 2020.
  28. 28.Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022.
  29. 29.Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 11–20, 2016.
  30. 30.Mou, C., Wang, X., Xie, L., Zhang, J., Qi, Z., Shan, Y., and Qie, X. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023.
  31. 31.Neyshabur, B., Sedghi, H., and Zhang, C. What is being transferred in transfer learning? Advances in neural information processing systems, 33:512–523, 2020.
  32. 32.Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I. Adapterfusion: Non-destructive task composition for transfer learning. arXiv preprint arXiv:2005.00247, 2020.
  33. 33.Poth, C., Pfeiffer, J., Rücklé, A., and Gurevych, I. What to pre-train on? efficient intermediate task selection. arXiv preprint arXiv:2104.08247, 2021.
  34. 34.Pruksachatkun, Y., Phang, J., Liu, H., Htut, P. M., Zhang, X., Pang, R. Y., Vania, C., Kann, K., and Bowman, S. R. Intermediate-task transfer learning with pretrained models for natural language understanding: When and why does it work? arXiv preprint arXiv:2005.00628, 2020.
  35. 35.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021.
  36. 36.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  37. 37.Schick, T. and Schütze, H. It’s not just size that matters: Small language models are also few-shot learners. arXiv preprint arXiv:2009.07118, 2020.
  38. 38.Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. FLAVA: A foundational language and vision alignment model. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 15617–15629. IEEE, 2022. doi: 10.1109/CVPR52688.2022.01519. URL https://doi.org/10.1109/CVPR52688.2022.01519.
  39. 39.Vu, T., Wang, T., Munkhdalai, T., Sordoni, A., Trischler, A., Mattarella-Micke, A., Maji, S., and Iyyer, M. Exploring and predicting transferability across nlp tasks. arXiv preprint arXiv:2005.00770, 2020.
  40. 40.Vu, T., Lester, B., Constant, N., Al-Rfou, R., and Cer, D. Spot: Better frozen model adaptation through soft prompt transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5039–5059, 2022.
  41. 41.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  42. 42.Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pp. 23318–23340. PMLR, 2022.
  43. 43.Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In International Conference on Machine Learning, pp. 23965–23998. PMLR, 2022.
  44. 44.Xie, N., Lai, F., Doran, D., and Kadav, A. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:1901.06706, 2019.
  45. 45.Yang, H., Lin, J., Yang, A., Wang, P., Zhou, C., and Yang, H. Prompt tuning for generative multimodal pretrained models. arXiv preprint arXiv:2208.02532, 2022.
  46. 46.Yang, Z., Gan, Z., Wang, J., Hu, X., Ahmed, F., Liu, Z., Lu, Y., and Wang, L. Crossing the format boundary of text and boxes: Towards unified vision-language modeling. arXiv preprint arXiv:2111.12085, 2021.
  47. 47.Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. Modeling context in referring expressions. In European Conference on Computer Vision, pp. 69–85. Springer, 2016.
  48. 48.Zaken, E. B., Ravfogel, S., and Goldberg, Y. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv preprint arXiv:2106.10199, 2021.
  49. 49.Zamir, A. R., Sax, A., Shen, W. B., Guibas, L. J., Malik, J., and Savarese, S. Taskonomy: Disentangling task transfer learning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 3712–3722. Computer Vision Foundation / IEEE Computer Society, 2018. doi: 10.1109/CVPR.2018.00391. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Zamir_Taskonomy_Disentangling_Task_CVPR_2018_paper.html.
  50. 50.Zhou, K., Yang, J., Loy, C. C., and Liu, Z. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022a.
  51. 51.Zhou, W., Xu, C., and McAuley, J. Efficiently tuned parameters are task embeddings. arXiv preprint arXiv:2210.11705, 2022b.
  52. 52.Zhu, X., Zhu, J., Li, H., Wu, X., Li, H., Wang, X., and Dai, J. Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 16783–16794. IEEE, 2022. doi: 10.1109/CVPR52688.2022.01630. URL https://doi.org/10.1109/CVPR52688.2022.01630.

Citation

MLA
Wu, C., et al. “$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation”. International Conference on Machine Learning, vol. 202, 2023, pp. 37713–27, https://proceedings.mlr.press/v202/wu23t.html.
APA
Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., & Luo, P. (2023). $\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation. International Conference on Machine Learning, 202, 37713–37727. https://proceedings.mlr.press/v202/wu23t.html
Chicago
Wu, C., T. Wang, Y. Ge, et al. 2023. “$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation”. International Conference on Machine Learning 202: 37713–27. https://proceedings.mlr.press/v202/wu23t.html.
Harvard
Wu, C. et al. (2023) “$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation”, International Conference on Machine Learning. PMLR, pp. 37713–37727. Available at: https://proceedings.mlr.press/v202/wu23t.html.
Vancouver
1. Wu C, Wang T, Ge Y, Lu Z, Zhou R, Shan Y, Luo P (2023) $\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation. In: International Conference on Machine Learning. PMLR, pp 37713–37727

BibTeX

@InProceedings{pmlr-v202-wu23t,
  title = 	 {$\pi$-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation},
  author =       {Wu, Chengyue and Wang, Teng and Ge, Yixiao and Lu, Zeyu and Zhou, Ruisong and Shan, Ying and Luo, Ping},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {37713--37727},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/wu23t/wu23t.pdf},
  url = 	 {https://proceedings.mlr.press/v202/wu23t.html},
  abstract = 	 {Foundation models have achieved great advances in multi-task learning with a unified interface of unimodal and multimodal tasks. However, the potential of such multi-task learners has not been exploited during transfer learning. In this work, we present a universal parameter-efficient transfer learning method, termed Predict-Interpolate Tuning ($\pi$-Tuning), for vision, language, and vision-language tasks. It aggregates the parameters of lightweight task-specific experts learned from similar tasks to aid the target downstream task. The task similarities are predicted in a unified modality-independent space, yielding a scalable graph to demonstrate task relationships. $\pi$-Tuning has several appealing benefits. First, it flexibly explores both intra- and inter-modal transferability between similar tasks to improve the accuracy and robustness of transfer learning, especially in data-scarce scenarios. Second, it offers a systematical solution for transfer learning with multi-task prediction-and-then-interpolation, compatible with diverse types of parameter-efficient experts, such as prompt and adapter. Third, an extensive study of task-level mutual benefits on 14 unimodal and 6 multimodal datasets shows that $\pi$-Tuning surpasses fine-tuning and other parameter-efficient transfer learning methods both in full-shot and low-shot regimes. The task graph also enables an in-depth interpretable analysis of task transferability across modalities. The code will be available at https://github.com/TencentARC/pi-Tuning.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/