Multimodal Prompting with Missing Modalities for Visual Recognition

Yi-Lun LeeYi-Hsuan TsaiWei-Chen ChiuChen-Yu Lee

article2023CVPR226 citations

Proposes a parameter-efficient prompt learning framework that adapts frozen multimodal transformers to arbitrary missing-modality scenarios in training or testing by tuning less than 1% of the model's parameters.

Listen

Modern artificial intelligence applications increasingly rely on multimodal models that combine different data types, such as text and images. In real-world deployments, however, complete data is rarely guaranteed due to privacy constraints, hardware failures, or network issues, resulting in missing information during either system training or active deployment. At the same time, adapting large pretrained transformer models to handle these edge cases typically requires full-model retraining, which demands massive computational budgets and millions of GPU hours. The article evaluates whether lightweight prompt learning can make multimodal models robust to arbitrary missing data while bypassing the extreme computational cost of full-scale fine-tuning.

The researchers propose a missing-aware prompt learning framework that keeps the underlying transformer model entirely frozen, introducing small learnable parameters called prompts conditioned on which modality is absent. The study evaluates two primary prompt insertion strategies—input-level and attention-level prompting—across three benchmark datasets representing multi-label classification, noisy image-text categorization, and hate speech detection under simulated data loss rates of up to 70% to 90%.

The findings demonstrate that missing-aware prompts substantially improve robustness across diverse missing-modality conditions. By training only 221,000 prompt parameters—less than 0.2% of the full 113-million-parameter backbone—the system achieves performance competitive with full-model fine-tuning (reaching a 42.66 F1-Macro score on movie genre classification compared to 46.45 for full retraining, while outperforming frozen baseline models by around 6 to 10 points across tasks). Input-level prompting generally yields the highest overall accuracy, whereas attention-level prompting exhibits greater stability across varying input sequence lengths. Furthermore, the analysis reveals that attaching prompts to the earliest transformer layers is far more critical to accuracy than prompt length or attaching prompts to later layers, as early intervention guides multimodal fusion before modality-specific traits are merged.

These results show that organizations can deploy resilient multimodal AI systems at a fraction of standard training, memory, and infrastructure costs. The approach eliminates the need to maintain separate, expensive models for every potential permutation of incomplete data. Organizations operating with limited computing budgets or large foundation models can prioritize lightweight prompt tuning over full retraining. When implementing this architecture, engineering teams should place prompt tokens within early transformer layers and choose attention-level prompting if target data lengths vary widely, or input-level prompting for maximum performance.

Confidence in these findings is high for vision-language classification benchmarks, but practical limitations remain. The evaluations focus on two-modality vision-text scenarios on a single base transformer model, leaving performance on more complex architectures, generative tasks, or three-way combinations (such as audio-video-text) unverified. Future work should validate the framework on larger billion-parameter foundations and operational production pipelines with unpredictable missing-data patterns.

arXiv: 2303.03369
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Introduces Visual Prompt Tuning (VPT) for parameter-efficient adaptation of vision transformers, providing the foundational technique that the source adapts for missing-modality scenarios.
  • Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Pioneers continuous prompt tuning for pre-trained vision-language architectures, establishing the core parameter-efficient paradigm upon which multimodal prompt learning builds.
  • Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). Extends prompt learning with dynamic conditional mechanisms, providing key background on input-dependent prompt generation used in adaptive multimodal models.
  • Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Establishes the foundational mechanics and scaling benefits of parameter-efficient continuous prompt tuning in frozen transformer models.
  • Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). Offers classical foundations on learning multimodal deep representations that remain robust when specific modalities are missing during inference.
  • Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). Describes fundamental alignment and fusion architectures for multimodal transformers, establishing the backbone design principles utilized in prompt-based multimodal learning.
Cover for Multimodal Prompting with Missing Modalities for Visual Recognition

Abstract

In this paper, we tackle two challenges in multimodal learning for visual recognition: 1) when missing-modality occurs either during training or testing in real-world situations; and 2) when the computation resources are not available to finetune on heavy transformer models. To this end, we propose to utilize prompt learning and mitigate the above two challenges together. Specifically, our modality-missing-aware prompts can be plugged into multimodal transformers to handle general missing-modality cases, while only requiring less than 1% learnable parameters compared to training the entire model. We further explore the effect of different prompt configurations and analyze the robustness to missing modality. Extensive experiments are conducted to show the effectiveness of our prompt learning framework that improves the performance under various missing-modality cases, while alleviating the requirement of heavy model re-training. Code is available.1

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Method
  • 3.1. Overall Framework
  • 3.2. Prompt Learning for Missing Modalities
  • 3.3. Prompt Design
  • 4. Experimental Results
  • 4.1. Implementation Details
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — General modality-incomplete learning framework

    model/method

    The paper studies multimodal recognition in which different samples may have different missing modalities, and the missing-modality pattern may differ between training and testing. For two modalities m1m_1 and m2m_2, the dataset is decomposed as D={Dc,Dm1,Dm2}D=\{D^c,D^{m_1},D^{m_2}\}, where Dc={(xim1,xim2,yi)}D^c=\{(x_i^{m_1},x_i^{m_2},y_i)\} contains complete pairs, Dm1={(xjm1,yj)}D^{m_1}=\{(x_j^{m_1},y_j)\} contains samples missing modality m2m_2, and Dm2={(xkm2,yk)}D^{m_2}=\{(x_k^{m_2},y_k)\} contains samples missing modality m1m_1. Missing inputs are replaced with fixed dummy inputs, producing a uniform pair of modality inputs for every sample.

    For each missing-modality case m∈{c,m1,m2}m\in\{c,m_1,m_2\}, the method selects a corresponding learnable missing-aware prompt PmP_m and inserts it into multiple layers of a frozen pretrained multimodal transformer. The transformer output associated with its text task token is sent to a trainable pooler and task-specific fully connected classifier. The workflow diagram on page 3 depicts this process: case-dependent prompt selection, dummy-input completion, prompt insertion into multimodal self-attention layers, and classification using only the trainable prompt and task-head components.

  2. Knowl 2 — Input-level missing-aware prompting

    model/method

    Let a pretrained multimodal transformer have NN multi-head self-attention layers. At layer ii, let hi∈RL×dh^i\in\mathbb{R}^{L\times d} be the sequence of LL input embeddings with embedding dimension dd, and let pmi∈RLp×dp_m^i\in\mathbb{R}^{L_p\times d} be the learnable prompt of length LpL_p for missing-modality case mm. Input-level prompting prepends the case-specific prompt to the layer input:

    fpromptinput(pmi,hi)=[pmi;hi],f_{\mathrm{prompt}}^{\mathrm{input}}(p_m^i,h^i)=[p_m^i;h^i],

    where [⋅;⋅][\cdot;\cdot] denotes concatenation along the sequence dimension. If prompts are inserted into NpN_p layers, the sequence can grow to NpLp+LN_pL_p+L at the deepest prompted layer because prompts introduced in earlier layers are retained and can interact with later prompts and modality tokens. This design gives each prompted layer access to inherited prompt information, but makes the method sensitive to the relative lengths of prompts and multimodal inputs.

  3. Knowl 3 — Attention-level missing-aware prompting

    model/method

    Attention-level prompting injects the missing-aware prompt into the key and value streams rather than prepending prompt tokens to the query input. For layer ii, let hi∈RL×dh^i\in\mathbb{R}^{L\times d} be the input sequence, let WQi,WKi,WVi∈Rd×dW_Q^i,W_K^i,W_V^i\in\mathbb{R}^{d\times d} be the frozen query, key, and value projection matrices, and define

    Qi=hiWQi,Ki=hiWKi,Vi=hiWVi.Q^i=h^iW_Q^i,\qquad K^i=h^iW_K^i,\qquad V^i=h^iW_V^i.

    The prompt pmi∈RLp×dp_m^i\in\mathbb{R}^{L_p\times d} is split, for even LpL_p, into key and value sub-prompts pki,pvi∈R(Lp/2)×dp_k^i,p_v^i\in\mathbb{R}^{(L_p/2)\times d}. The prompted attention operation is

    fpromptattn(pmi,hi)=softmax⁡(Qi[pki;Ki]⊤d)[pvi;Vi].f_{\mathrm{prompt}}^{\mathrm{attn}}(p_m^i,h^i)=\operatorname{softmax}\left(\frac{Q^i[p_k^i;K^i]^\top}{\sqrt d}\right)[p_v^i;V^i].

    Because the prompt is not added to the query sequence, the output sequence retains length LL. Each prompt affects the attention computation of its own layer without accumulating prompt tokens in deeper layers.

  4. Knowl 4 — Multi-layer prompt placement

    model/method

    The method assigns a separate prompt to each selected transformer layer. For contiguous layer indices from ss through ee, the prompts for missing-modality case mm are represented by

    Pm={pmi}i=se∈RNp×Lp×d,Np=e−s+1,P_m=\{p_m^i\}_{i=s}^{e}\in\mathbb{R}^{N_p\times L_p\times d},\qquad N_p=e-s+1,

    where pmip_m^i is attached either to the layer input in input-level prompting or to the key and value streams in attention-level prompting. Experiments show that increasing the number of prompted layers generally improves recognition, but the starting layer is more important than the total count: earlier layers are more effective because modality-specific information is still distinct before deeper transformer layers fuse the modalities. The paper consequently uses the early half of the transformer, corresponding to layers 00 through 55 of the 12-layer ViLT backbone.

  5. Knowl 5 — Frozen-backbone training objective

    model/method

    The pretrained multimodal transformer fθf_\theta is frozen. The trainable parameters are the missing-aware prompt parameters θp\theta_p and the downstream task parameters θt\theta_t, consisting of the pooler and fully connected classifier. For a modality pair (xnm1,xnm2)(x_n^{m_1},x_n^{m_2}) in the dummy-completed training set and label yny_n, optimization minimizes the task-specific loss:

    L=Ltask(xnm1,xnm2,yn;θt,θp).\mathcal{L}=\mathcal{L}_{\mathrm{task}}(x_n^{m_1},x_n^{m_2},y_n;\theta_t,\theta_p).

    The loss can be, for example, binary cross-entropy for multi-label movie-genre classification. Thus, the approach adapts a large multimodal transformer without updating its pretrained parameters; only the case-dependent prompts and the task head learn from the downstream data.

  6. Knowl 6 — Experimental protocol for missing modalities

    experimental setup

    Experiments use the frozen ViLT vision-language transformer on MM-IMDb, UPMC Food-101, and Hateful Memes. The evaluation metrics are F1-Macro, classification accuracy, and AUROC, respectively. Text is tokenized with the bert-base-uncased tokenizer, with maximum lengths 1024 for MM-IMDb, 512 for UPMC Food-101, and 128 for Hateful Memes. Missing text is replaced by an empty string. Images are resized so that the shorter side is 384 pixels and the longer side is at most 640 pixels, then divided into 32×3232\times32 patches; missing images are replaced by an image whose pixel values are all one.

    The default prompt length is Lp=16L_p=16, and prompts are inserted into ViLT layers 00 through 55. All experiments use AdamW with learning rate 10−210^{-2}, weight decay 2×10−22\times10^{-2}, a warm-up over 10% of training steps, and linear decay to zero afterward.

    The missing rate η\eta is the fraction of modality-incomplete samples. For missing-text experiments, η\eta of samples are image-only and 1−η1-\eta are complete; for missing-image experiments, η\eta are text-only and 1−η1-\eta are complete. For missing-both experiments, η/2\eta/2 are text-only, η/2\eta/2 are image-only, and 1−η1-\eta are complete. The default is η=70%\eta=70\%, and training and testing can use different missing-modality cases and rates.

  7. Knowl 7 — Performance across missing-modality cases

    data/table

    The quantitative comparison below evaluates a frozen ViLT baseline that trains only the pooler and classifier against attention-level and input-level missing-aware prompts. The entries in the training and testing columns are the percentages of available image and text modalities, respectively; all experiments use a 70% missing rate. Scores are F1-Macro for MM-IMDb, accuracy for UPMC Food-101, and AUROC for Hateful Memes.

    Could not parse LaTeX table

    Attention-level prompts improve the frozen baseline in all nine conditions. Input-level prompts give the best score in eight of nine conditions and provide the largest gains on MM-IMDb and UPMC Food-101; the exception is the missing-text Hateful Memes condition, where attention-level prompting reaches 62.17 AUROC and input-level prompting reaches 59.11.

  8. Knowl 8 — Robustness to mismatched training and testing missing rates

    empirical result

    The missing-rate curves reported on pages 6–7 show that missing-aware prompts remain useful when the missing-modality rate at testing differs from the rate used during training. When models are trained with missing-both data at a 70% rate and tested with missing-text data across rates, both prompt designs outperform the frozen baseline. Attention-level prompting is stronger when the testing missing rate is below 30%, whereas input-level prompting becomes more effective as the testing missing rate increases beyond 30%.

    For input-level prompting trained on missing-both data, training at a low missing rate of 10% gives higher scores when the test data are mostly complete, while training at a high missing rate of 90% remains competitive when test data contain many incomplete samples. When training starts from modality-complete data but randomly assigns complete, text-only, or image-only inputs across training epochs, both prompt designs still consistently improve the baseline across testing missing rates.

  9. Knowl 9 — Prompt configuration and length trade-offs

    limitation

    Input-level prompting generally achieves the highest recognition scores, but its accumulated prompt tokens make it sensitive to the ratio between prompt length and input length. This effect is especially visible for Hateful Memes, whose maximum text length is only 128 tokens: with the default prompt length of 16, input-level prompting can make task-specific feature learning less stable and performs worse than attention-level prompting in the missing-text condition.

    Attention-level prompting is more stable across datasets because prompts are inserted into each layer's key and value streams without extending the query sequence, so each prompt primarily controls its corresponding layer. In prompt-length ablations, performance improves as the length grows initially but both designs reach their best performance at Lp=16L_p=16; even Lp=1L_p=1 remains competitive. Adding an equivalent number of ordinary classifier parameters does not provide a comparable improvement, indicating that the location and function of the prompt parameters matter.

  10. Knowl 10 — Parameter efficiency compared with full fine-tuning

    empirical result

    On MM-IMDb under the missing-both condition, fully fine-tuning the 113-million-parameter ViLT obtains 46.45 F1-Macro. The proposed approach freezes ViLT and trains 221 thousand missing-aware prompt parameters, corresponding to only 0.2% of the backbone parameters, while obtaining 42.66 F1-Macro with input-level prompting. The parameter ratio excludes the task classifier because a downstream classifier is required in either setting. This demonstrates that the method preserves much of the full-fine-tuning performance while avoiding the computational cost of updating the entire multimodal transformer.

Coverage note — The paper's supplementary qualitative examples and additional plots are omitted because the supplied pages only state that they exist without reporting enough detailed data to form reproducible standalone knowls.

References

  1. 1.John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González. Gated multimodal units for information fusion. In International Conference on Learning Representations (ICLR) Workshops, 2017. 5, 6
  2. 2.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Visual prompting: Modifying pixel space to adapt pre-trained models. ArXiv:2203.17274, 2022. 2, 3
  3. 3.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision (ECCV), 2014. 5
  4. 4.Adam Botach, Evgenii Zheltonozhskii, and Chaim Baskin. End-to-end referring video object segmentation with multimodal transformers. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 2020. 1, 2, 3
  6. 6.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021. 4, 5
  7. 7.Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. Multi-modal transformer for video retrieval. In European Conference on Computer Vision (ECCV), 2020. 2
  8. 8.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In European Conference on Computer Vision (ECCV), 2022. 2, 3
  9. 9.Haotian Ju, Dongyue Li, and Hongyang R Zhang. Robust fine-tuning of deep neural networks with hessian-based generalization guarantees. ArXiv:2206.02659, 2022. 1
  10. 10.Mayank Kejriwal and Ke Shen. Do fine-tuned commonsense language models really generalize? ArXiv:2011.09159, 2020. 1
  11. 11.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. ArXiv:2210.03117, 2022. 3
  12. 12.Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in Neural Information Processing Systems (NeurIPS), 2020. 5, 6
  13. 13.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning (ICML), 2021. 1, 2, 4, 5, 6
  14. 14.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 2017. 6
  15. 15.Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister. Formnet: Structural encoding beyond sequential modeling in form document information extraction. In ACL, 2022. 1
  16. 16.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. ArXiv:2104.08691, 2021. 2, 3
  17. 17.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems (NeurIPS), 2021. 1, 2
  18. 18.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. ArXiv:2101.00190, 2021. 2, 3
  19. 19.Sheng Liang, Mengjie Zhao, and Hinrich Schütze. Modular and parameter-efficient multimodal fusion with prompting. In Findings of the Association for Computational Linguistics: ACL 2022, 2022. 3
  20. 20.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, 2014. 6
  21. 21.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2018. 6
  22. 22.Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, and Xi Peng. Are multimodal transformers robust to missing modality? In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 2, 3, 4, 5
  23. 23.Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng. Smil: Multimodal learning with severely missing modality. In AAAI Conference on Artificial Intelligence (AAAI), 2021. 2
  24. 24.Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines. In International Conference on Learning Representations (ICLR), 2021. 1
  25. 25.Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabas Póczos. Found in translation: Learning robust joint representations by cyclic translations between modalities. In AAAI Conference on Artificial Intelligence (AAAI), 2019. 1
  26. 26.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. ArXiv:2112.11446, 2021. 1
  27. 27.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 2020. 1
  28. 28.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems (NeurIPS), 2021. 5
  29. 29.Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister. Prefix conditioning unifies language and label supervision. ArXiv:2206.01125, 2022. 2
  30. 30.Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister. Pic2word: Mapping pictures to words for zero-shot composed image retrieval. In CVPR, 2023. 1
  31. 31.Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. What language model to train if you have one million gpu hours? ArXiv:2210.15424, 2022. 1
  32. 32.Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems (NeurIPS), 2021. 2, 3
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017. 5
  34. 34.Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Frederic Precioso. Recipe recognition with large multimodal food dataset. In IEEE International Conference on Multimedia & Expo (ICME) Workshops, 2015. 5, 6
  35. 35.Zilong Wang, Zhaohong Wan, and Xiaojun Wan. Transmodality: An end2end fusion method with transformer for multimodal sentiment analysis. In ACM Web Conference (WWW), 2020. 1
  36. 36.Zifeng Wang, Zizhao Zhang, Sayna Ebrahimi, Ruoxi Sun, Han Zhang, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, et al. Dualprompt: Complementary prompting for rehearsal-free continual learning. In European Conference on Computer Vision (ECCV), 2022. 2, 3, 4
  37. 37.Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister. Learning to prompt for continual learning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  38. 38.Jinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis, and Jingkuan Song. Prompting for multi-modal tracking. In ACM Conference on Multimedia (MM), 2022. 3
  39. 39.Jiandian Zeng, Tianyi Liu, and Jiantao Zhou. Tag-assisted multimodal sentiment analysis under uncertain missing modalities. ArXiv:2204.13707, 2022. 2
  40. 40.Jinming Zhao, Ruichen Li, and Qin Jin. Missing modality imagination network for emotion recognition with uncertain missing modalities. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), 2021. 2
  41. 41.Jinming Zhao, Ruichen Li, Qin Jin, Xinchao Wang, and Haizhou Li. Memobert: Pre-training model with prompt-based learning for multimodal emotion recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022. 3
  42. 42.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision (IJCV), 2022. 2, 3

Citation

MLA
Lee, Y.-L., et al. “Multimodal Prompting with Missing Modalities for Visual Recognition”. arXiv, 2023, http://arxiv.org/abs/2303.03369v2.
APA
Lee, Y.-L., Tsai, Y.-H., Chiu, W.-C., & Lee, C.-Y. (2023). Multimodal Prompting with Missing Modalities for Visual Recognition. arXiv. http://arxiv.org/abs/2303.03369v2
Chicago
Lee, Y.-L., Y.-H. Tsai, W.-C. Chiu, and C.-Y. Lee. 2023. “Multimodal Prompting with Missing Modalities for Visual Recognition”. arXiv. http://arxiv.org/abs/2303.03369v2.
Harvard
Lee, Y.-L. et al. (2023) “Multimodal Prompting with Missing Modalities for Visual Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.03369v2.
Vancouver
1. Lee Y-L, Tsai Y-H, Chiu W-C, Lee C-Y (2023) Multimodal Prompting with Missing Modalities for Visual Recognition. arXiv

BibTeX

@article{lee2023multimodal,
  title = {Multimodal Prompting with Missing Modalities for Visual Recognition},
  author = {Lee, Yi-Lun and Tsai, Yi-Hsuan and Chiu, Wei-Chen and Lee, Chen-Yu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.03369v2},
  eprint = {2303.03369}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE