Efficient Multimodal Fusion via Interactive Prompting

Yaowei LiRuijie QuanLinchao ZhuYi Yang

article2023CVPR68 citations

Proposes a parameter- and memory-efficient multimodal fusion framework that uses interactive deep-layer prompts across frozen unimodal transformers, matching full fine-tuning performance while cutting training memory usage by up to 66% and updating fewer than 3% of parameters.

Listen

Modern artificial intelligence increasingly relies on combining diverse data sources, such as text and images, to solve complex real-world tasks. While large multimodal models deliver strong predictive performance, updating their full set of parameters for specialized downstream applications requires massive graphics processing unit memory and computational infrastructure. Existing parameter-efficient tuning methods reduce the number of adjusted weights but still require passing gradient calculations through every layer of the network, offering minimal relief for real-world memory bottlenecks.

The article evaluates a memory-efficient framework called Prompt-based Multimodal Fusion (PMF), designed to integrate independently pretrained unimodal vision and language models for multimodal tasks. The objective is to demonstrate that targeted, two-way interactive prompting can drastically reduce memory usage and trainable parameter counts while matching the accuracy of fully finetuned models.

To evaluate this approach, the researchers conducted extensive empirical testing across three established multimodal benchmark datasets spanning recipe categorization, movie genre classification, and visual entailment. The modular architecture freezes the underlying image and text transformers and establishes a parallel, two-way communication pathway between them. Instead of inserting continuous prompt vectors across the entire model, the framework introduces three specialized prompt types—query prompts, query context prompts, and fusion context prompts—only into the deepest layers of the networks. Lightweight mapping functions then translate queried representations from one modality to the other.

The experimental findings show that PMF significantly improves computational efficiency without sacrificing predictive accuracy. First, the method reduces training memory usage by up to 66% compared to standard finetuning baselines and by up to 55% compared to existing prompt-based fusion techniques. Second, it updates less than 3% of the total model parameters—and under 2.5% in base configurations—while achieving performance comparable to full-model finetuning across all evaluated datasets. Third, the framework outperforms prior prompt-based multimodal approaches across every benchmark. Finally, applying the strategy to larger underlying transformer backbones further improves accuracy with minimal added memory overhead.

These results demonstrate that organizations can deploy high-performing multimodal AI systems on standard, lower-memory hardware, substantially lowering computing infrastructure costs and broadening deployment feasibility. Furthermore, because the framework pairs independently trained vision and language models, organizations do not need expensive, paired multimodal datasets for initial pretraining. Decoupling prompts into dedicated querying and fusion stages overcomes the historical performance gap seen in earlier parameter-efficient fusion methods.

For practical implementation, engineering teams should consider adopting deep-layer prompt fusion when deploying multimodal applications under strict compute constraints. The source also demonstrates that applying automated architecture search can optimize prompt lengths and layer choices for specific tasks. For future exploration, researchers should evaluate PMF across additional complex multimodal domains, such as visual question answering, and investigate new prompting designs to fully close the remaining small performance gap with full finetuning.

Decision-makers should note that while PMF performs competitively, its accuracy remains slightly behind full finetuning on certain base configurations, and introducing distinct prompt types requires tuning task-specific hyperparameters. Nonetheless, the reported findings provide strong empirical confidence that deep interactive prompting is a highly effective, cost-efficient strategy for multimodal model integration.

arXiv: 2304.06306
Cover for Efficient Multimodal Fusion via Interactive Prompting

Abstract

Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multimodal learning models constantly increases, leading to an urgent need to reduce the massive computational cost of finetuning these models for downstream tasks. In this paper, we propose an efficient and flexible multimodal fusion method, namely PMF, tailored for fusing unimodally pretrained transformers. Specifically, we first present a modular multimodal fusion framework that exhibits high flexibility and facilitates mutual interactions among different modalities. In addition, we disentangle vanilla prompts into three types in order to learn different optimizing objectives for multimodal learning. It is also worth noting that we propose to add prompt vectors only on the deep layers of the unimodal transformers, thus significantly reducing the training memory usage. Experiment results show that our proposed method achieves comparable performance to several other multimodal finetuning methods with less than 3% trainable parameters and up to 66% saving of training memory usage.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Prompt-based Multimodal Fusion
  • 3.1. Unimodal Transformers
  • 3.2. Unimodal Base feature Extraction
  • 3.3. Multimodal Fusion Layer
  • 4. Experiments
  • 4.1. Datasets and Metrics
  • 4.2. Existing Methods and Baselines
  • 4.3. Implementation Details
  • 4.4. Main Results
  • 4.5. Ablation Study
  • 4.6. Modularity and Flexibility
  • 4.7. PMF with NAS
  • 5. Limitations and Future Works
  • 6. Conclusion
  • 7. Acknowledgement
  • References

Knowls

  1. Knowl 1 — PMF’s parallel modular fusion architecture

    model/method

    Prompt-based Multimodal Fusion (PMF) combines two independently pretrained transformers—a vision transformer for an image and a language transformer for text—in parallel rather than placing one modality encoder after the other. Each input is converted into a sequence beginning with its modality-specific CLS token, and each sequence is processed by its own LL-layer transformer. If zl\mathbf z^l and z′l\mathbf z^{\prime l} are the two modality sequences at layer ll, PMF first performs unimodal feature extraction for layers before the fusion point LfL_f:

    zl+1=TransLayer⁡l(zl;θ),l<Lf,\mathbf z^{l+1}=\operatorname{TransLayer}^{l}(\mathbf z^l;\theta),\qquad l<L_f,

    where θ\theta denotes frozen pretrained parameters. From layer LfL_f onward, the two streams exchange information through PMF fusion layers, allowing bidirectional rather than one-way cross-modal interaction. The modular design permits either unimodal backbone to be replaced independently and can accommodate backbones with different hidden dimensions through learned mappings.

    After the final layer LL, PMF applies separate linear classifiers to the two CLS representations, zCLSL\mathbf z^L_{\mathrm{CLS}} and zCLS′L\mathbf z^{\prime L}_{\mathrm{CLS}}, and averages their pre-softmax logits for the final classification prediction. The pretrained transformer parameters remain frozen; trainable parameters consist of prompt vectors, cross-modal nonlinear mappings, and the final classifiers.

  2. Knowl 2 — Three-stage interactive prompting for bidirectional fusion

    model/method

    Each PMF multimodal fusion layer contains a querying stage followed by a fusion stage. For each modality, PMF learns three distinct prompt types: a query prompt (QP), a query-context prompt (QCP), and a fusion-context prompt (FCP). QP extracts information to send to the other modality, QCP supplies context while the query is formed, and FCP supplies context when the extracted information is incorporated into the other modality.

    For an unprimed modality with input sequence zl\mathbf z^l, QCP zqcpl\mathbf z^l_{qcp}, and QP zqpl\mathbf z^l_{qp}, the querying stage concatenates the prompts to the sequence and applies the frozen unimodal transformer layer:

    [z^l ∣∣ z^qcpl ∣∣ z^qpl]=TransLayer⁡l([zl ∣∣ zqcpl ∣∣ zqpl];θ).[\hat{\mathbf z}^{l}\,||\,\hat{\mathbf z}^{l}_{qcp}\,||\,\hat{\mathbf z}^{l}_{qp}] =\operatorname{TransLayer}^{l}([\mathbf z^{l}\,||\,\mathbf z^{l}_{qcp}\,||\,\mathbf z^{l}_{qp}];\theta).

    Only the output corresponding to QP, z^qpl\hat{\mathbf z}^{l}_{qp}, is retained as the information to transmit; the QCP output is discarded after providing query context. A learned nonlinear mapping flf^l converts this information into the representation space of the other modality:

    yqpl=fl(z^qpl).\mathbf y^l_{qp}=f^l(\hat{\mathbf z}^{l}_{qp}).

    The mapping consists of two linear layers with a bottleneck, with a ReLU activation after only the first linear layer. The other modality then performs fusion by concatenating its original sequence, its FCP, and the mapped query output, followed by its frozen transformer layer:

    [z′l+1 ∣∣ z^fcp′l ∣∣ y^qpl]=TransLayer⁡l([z′l ∣∣ zfcp′l ∣∣ yqpl];θ′).[\mathbf z^{\prime l+1}\,||\,\hat{\mathbf z}^{\prime l}_{fcp}\,||\,\hat{\mathbf y}^{l}_{qp}] =\operatorname{TransLayer}^{l}([\mathbf z^{\prime l}\,||\,\mathbf z^{\prime l}_{fcp}\,||\,\mathbf y^{l}_{qp}];\theta^{\prime}).

    The same operations are performed in the reverse direction using the primed modality’s prompts and mapping f′lf^{\prime l}, so each modality both queries and receives information at every fusion layer. The resulting two modality sequences become the input to the next fusion layer.

  3. Knowl 3 — Deep-layer prompting reduces training memory

    model/method

    PMF reduces training memory by inserting trainable prompts only in the deep layers of the frozen unimodal transformers, beginning at the fusion layer LfL_f, rather than inserting prompts throughout all LL layers. Layers before LfL_f are used only to compute base features and do not need to retain intermediate activations for gradient computation because their parameters are frozen and they contain no trainable prompts.

    Consequently, backpropagation traverses only the suffix of the two transformer stacks containing the multimodal fusion layers, the prompt vectors, the nonlinear mappings, and the classifiers. Increasing LfL_f shortens this trainable suffix and lowers memory usage, while potentially reducing the amount of unimodal processing available for cross-modal fusion. The trainable nonlinear mappings account for more than 95% of PMF’s trainable parameters, while the full pretrained backbones remain unchanged.

  4. Knowl 4 — Main multimodal classification results and memory comparison

    data/table

    PMF was evaluated against unimodal fine-tuning, full multimodal fine-tuning, and prompt-based fusion methods on SNLI-VE, UPMC-Food-101, and MM-IMDB. SNLI-VE and UPMC-Food-101 use accuracy; MM-IMDB uses F1-Macro/F1-Micro. The reported values are means over three random seeds. Training and inference memory are reported as maximum GPU memory in gigabytes, and the parameter column gives updated parameters in millions. A dash denotes fewer than 0.10.1 million trainable parameters.

    Could not parse LaTeX table

    With the base backbones, PMF reaches an average score of 75.02 while updating only 2.5 million parameters and using 12.84 GB of training memory. This is comparable to LateConcat’s 75.18 average despite using far fewer trainable parameters and about 66% less training memory than LateConcat. PMF also outperforms all listed prompt-based methods. With BERT-large and ViT-large, PMF-large reaches a 75.99 average and exceeds LateConcat’s 75.18 average while using 4.5 million trainable parameters and 18.44 GB of training memory.

  5. Knowl 5 — Datasets, backbones, and training protocol

    experimental setup

    PMF was tested on three vision-language classification datasets. UPMC-Food-101 contains 90,840 image-recipe pairs from 101 food classes; because it provides only training and testing splits, 5,000 training examples were held out for validation. MM-IMDB contains 25,956 movie-poster/movie-plot pairs and requires multilabel genre prediction over 23 long-tailed classes. SNLI-VE contains 565,286 image-premise/text-hypothesis pairs, with labels entailment, contradiction, or neutrality; PMF uses only the image premise and text hypothesis as input.

    Accuracy is reported for UPMC-Food-101 and SNLI-VE, while macro-F1 and micro-F1 are reported for MM-IMDB. Unless otherwise specified, the image encoder is an ImageNet-21k-pretrained ViT-base and the language encoder is BERT-base-uncased, both with 12 hidden layers. Prompt vectors are initialized from a Gaussian distribution with mean 00 and standard deviation 0.020.02. Training uses SGD with momentum 0.90.9, weight decay 10−410^{-4}, batch size 64 for SNLI-VE, and batch size 32 for UPMC-Food-101 and MM-IMDB. Cross-entropy loss is used, with inverse-frequency class weighting for UPMC-Food-101 and MM-IMDB.

  6. Knowl 6 — Every PMF component contributes to fusion quality

    empirical result

    An ablation on MM-IMDB evaluated the four trainable PMF components—QP, the nonlinear mapping ff, QCP, and FCP—with Lf=10L_f=10. A single check-marked prompt has length 4; the doubled QP setting has length 8. The results are F1-Macro/F1-Micro.

    Could not parse LaTeX table

    The no-component configuration is equivalent to a linear classifier on the unimodal CLS features. Adding QP or FCP alone reduces performance, showing that isolated prompting of the frozen top layers does not achieve useful fusion. Adding the mapping together with QP produces the largest single improvement, but the mapping cannot operate without QP-derived information. Adding QCP and FCP separately gives further gains, and the complete three-prompt design performs best. Extending QP to length 8 does not replace QCP: the QP-plus-mapping configuration with length 8 reaches 57.98/63.69, whereas adding QCP with length-4 QP reaches 58.30/64.07.

  7. Knowl 7 — Fusion depth and prompt length control the memory–accuracy trade-off

    empirical result

    On MM-IMDB, PMF was evaluated with fusion starting at layers Lf∈{0,2,4,6,8,10,12}L_f\in\{0,2,4,6,8,10,12\} while using prompt length M=4M=4 for each prompt type. Training memory decreases continuously as fusion begins later because fewer transformer layers participate in backpropagation. Performance remains relatively stable for Lf≤10L_f\leq 10, but starting fusion too late harms performance; the authors therefore identify deep-layer prompting, approximately 10<l<L10<l<L, as the best empirical trade-off for the base model.

    A separate experiment fixed Lf=10L_f=10 and used equal lengths MM for QP, QCP, and FCP, testing M∈{1,4,8,16,32}M\in\{1,4,8,16,32\}. Performance improves as MM grows through 16 and declines when the prompts become too long at M=32M=32. Increasing MM from 1 to 16 raises training memory by only about 1 GB, indicating that fusion depth is a substantially more important memory control than prompt length.

  8. Knowl 8 — PMF scales to interchangeable unimodal backbones

    data/table

    PMF’s modular design allows the language and vision transformers to be replaced independently. When the two encoders have different depths, fusion starts at Lf-img=Limg−2L_{f\text{-img}}=L_{\text{img}}-2 for the image encoder and Lf-txt=Ltxt−2L_{f\text{-txt}}=L_{\text{txt}}-2 for the text encoder, leaving two unimodal layers before fusion in each stream. Differences in hidden dimension are handled by the learned nonlinear mappings. With prompt length M=4M=4, the following results were obtained on MM-IMDB; memory is reported as training/inference GPU memory in GB and scores are F1-Macro/F1-Micro.

    Could not parse LaTeX table

    Replacing either base encoder with a large encoder improves MM-IMDB performance, and using both large encoders increases the score from 58.77/64.51 to 61.66/66.72. The associated training memory rises from 12.84 GB to 18.44 GB, supporting the paper’s claim that PMF can exploit larger unimodal models with a limited memory increase.

  9. Knowl 9 — Automatic fusion-structure search improves PMF with less manual tuning

    algorithm

    PMF has two main fusion-structure hyperparameters: the starting fusion layer LfL_f and the prompt length MM. The authors apply AutoFormer-based neural architecture search to automatically select a fusion structure rather than manually choosing these values. The search evaluates candidate PMF structures over this fusion design space and returns a task-specific configuration; the supplied paper reports the outcome but does not provide the detailed search space or evolutionary-search steps.

    Using the same vision and language encoder families as regular PMF, the searched configuration obtains the following results:

    Could not parse LaTeX table

    The NAS configuration improves the average score from 75.02 for the regular base PMF configuration to 75.66, at the cost of higher training memory, and reduces the manual effort needed to find a suitable fusion structure.

  10. Knowl 10 — Stated limitations of PMF

    limitation

    PMF does not consistently match full fine-tuning when the same pretrained backbones are used: although its performance is comparable to several fine-tuning baselines with far fewer trainable parameters, the paper states that it remains behind fine-tuning on the evaluated datasets. This indicates that the prompting mechanism does not yet extract all of the useful knowledge stored in the frozen pretrained models.

    The second limitation is hyperparameter complexity. Separating prompts into QP, QCP, and FCP gives them different fusion roles, but it also creates more choices for prompt lengths and fusion locations, increasing the tuning effort required to obtain the best task-specific configuration. The authors identify broader evaluation on multimodal tasks such as visual question answering and with additional model architectures as future work.

Coverage note — The supplementary AutoFormer search-space and evolutionary-search details are not reconstructed because they are not included in the supplied paper text; no other substantial contributed material is deliberately omitted.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022. 2
  2. 2.John Arevalo, Thamar Solorio, Manuel Montes-y Gomez, and Fabio A Gonzalez. Gated multimodal units for information fusion. arXiv preprint arXiv:1702.01992, 2017. 2, 5
  3. 3.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Visual prompting: Modifying pixel space to adapt pre-trained models. arXiv preprint arXiv:2203.17274, 2022. 3
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 3
  5. 5.Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling. Autoformer: Searching transformers for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12270–12280, 2021. 8
  6. 6.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758, 2021. 1
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 6
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1, 5
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3, 5
  10. 10.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 1
  11. 11.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021. 2, 3
  12. 12.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. arXiv preprint arXiv:2203.12119, 2022. 3, 5
  13. 13.Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. Prompting visual-language models for efficient video understanding. arXiv preprint arXiv:2112.04478, 2021. 1, 3
  14. 14.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. arXiv preprint arXiv:2210.03117, 2022. 1, 3
  15. 15.Douwe Kiela, Suvrat Bhooshan, Hamed Firooz, Ethan Perez, and Davide Testuggine. Supervised multimodal bitransformers for classifying images and text. arXiv preprint arXiv:1909.02950, 2019. 2, 3, 5, 6
  16. 16.Ryan Kiros, Ruslan Salakhutdinov, and Rich Zemel. Multimodal neural language models. In International conference on machine learning, pages 595–603. PMLR, 2014. 2
  17. 17.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021. 1, 3
  18. 18.Guangrui Li, Guoliang Kang, Yi Zhu, Yunchao Wei, and Yi Yang. Domain consensus clustering for universal domain adaptation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  19. 19.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021. 1, 3
  20. 20.Sheng Liang, Mengjie Zhao, and Hinrich Schutze. Modular and parameter-efficient multimodal fusion with prompting. arXiv preprint arXiv:2203.08055, 2022. 1, 2, 3, 6
  21. 21.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021. 1, 3
  22. 22.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too. arXiv preprint arXiv:2103.10385, 2021. 1, 3
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 1
  24. 24.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019. 2
  25. 25.Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian. Prompt distribution learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5206–5215, 2022. 3
  26. 26.Oscar Manas, Pau Rodriguez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, and Aishwarya Agrawal. Mapl: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting. arXiv preprint arXiv:2210.07179, 2022. 2
  27. 27.Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. Attention bottlenecks for multimodal fusion. Advances in Neural Information Processing Systems, 34:14200–14213, 2021. 2, 6
  28. 28.Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. Multimodal deep learning. In ICML, 2011. 2
  29. 29.Guanghui Qin and Jason Eisner. Learning how to ask: Querying lms with mixtures of soft prompts. arXiv preprint arXiv:2104.06599, 2021. 1, 3
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 2, 3
  31. 31.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 1
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020. 1
  33. 33.Nitish Srivastava and Russ R Salakhutdinov. Multimodal learning with deep boltzmann machines. Advances in neural information processing systems, 25, 2012. 2
  34. 34.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems, 34:200–212, 2021. 2, 3
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3
  36. 36.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pages 23318–23340. PMLR, 2022. 2
  37. 37.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14733–14743, 2022. 2
  38. 38.Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Frederic Precioso. Recipe recognition with large multimodal food dataset. In 2015 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), pages 1–6. IEEE, 2015. 2, 5
  39. 39.Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu, and Yu-Gang Jiang. Semi-supervised vision transformers. In ECCV, 2022. 1
  40. 40.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45, 2020. 6
  41. 41.Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. Visual entailment: A novel task for fine-grained image understanding. arXiv preprint arXiv:1901.06706, 2019. 2, 5
  42. 42.Ning Xiong and Per Svensson. Multi-sensor management for information fusion: issues and approaches. Information fusion, 3(2):163–186, 2002. 2
  43. 43.Hao Yang, Junyang Lin, An Yang, Peng Wang, Chang Zhou, and Hongxia Yang. Prompt tuning for generative multimodal pretrained models. arXiv preprint arXiv:2208.02532, 2022. 1, 3, 5
  44. 44.Yi Yang, Yueting Zhuang, and Yunhe Pan. Multiple knowledge representation for big data artificial intelligence: framework, applications, and case studies. Frontiers of Information Technology & Electronic Engineering, 22(12):1551–1558, 2021. 2
  45. 45.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 2
  46. 46.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825, 2022. 3
  47. 47.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022. 1, 3
  48. 48.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8746–8755, 2020. 1

Citation

MLA
Li, Y., et al. “Efficient Multimodal Fusion via Interactive Prompting”. arXiv, 2023, http://arxiv.org/abs/2304.06306v2.
APA
Li, Y., Quan, R., Zhu, L., & Yang, Y. (2023). Efficient Multimodal Fusion via Interactive Prompting. arXiv. http://arxiv.org/abs/2304.06306v2
Chicago
Li, Y., R. Quan, L. Zhu, and Y. Yang. 2023. “Efficient Multimodal Fusion via Interactive Prompting”. arXiv. http://arxiv.org/abs/2304.06306v2.
Harvard
Li, Y. et al. (2023) “Efficient Multimodal Fusion via Interactive Prompting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.06306v2.
Vancouver
1. Li Y, Quan R, Zhu L, Yang Y (2023) Efficient Multimodal Fusion via Interactive Prompting. arXiv

BibTeX

@article{li2023efficient,
  title = {Efficient Multimodal Fusion via Interactive Prompting},
  author = {Li, Yaowei and Quan, Ruijie and Zhu, Linchao and Yang, Yi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.06306v2},
  eprint = {2304.06306}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE