Multimodal Prompting with Missing Modalities for Visual Recognition
Yi-Lun LeeYi-Hsuan TsaiWei-Chen ChiuChen-Yu Lee
Proposes a parameter-efficient prompt learning framework that adapts frozen multimodal transformers to arbitrary missing-modality scenarios in training or testing by tuning less than 1% of the model's parameters.
Modern artificial intelligence applications increasingly rely on multimodal models that combine different data types, such as text and images. In real-world deployments, however, complete data is rarely guaranteed due to privacy constraints, hardware failures, or network issues, resulting in missing information during either system training or active deployment. At the same time, adapting large pretrained transformer models to handle these edge cases typically requires full-model retraining, which demands massive computational budgets and millions of GPU hours. The article evaluates whether lightweight prompt learning can make multimodal models robust to arbitrary missing data while bypassing the extreme computational cost of full-scale fine-tuning.
The researchers propose a missing-aware prompt learning framework that keeps the underlying transformer model entirely frozen, introducing small learnable parameters called prompts conditioned on which modality is absent. The study evaluates two primary prompt insertion strategies—input-level and attention-level prompting—across three benchmark datasets representing multi-label classification, noisy image-text categorization, and hate speech detection under simulated data loss rates of up to 70% to 90%.
The findings demonstrate that missing-aware prompts substantially improve robustness across diverse missing-modality conditions. By training only 221,000 prompt parameters—less than 0.2% of the full 113-million-parameter backbone—the system achieves performance competitive with full-model fine-tuning (reaching a 42.66 F1-Macro score on movie genre classification compared to 46.45 for full retraining, while outperforming frozen baseline models by around 6 to 10 points across tasks). Input-level prompting generally yields the highest overall accuracy, whereas attention-level prompting exhibits greater stability across varying input sequence lengths. Furthermore, the analysis reveals that attaching prompts to the earliest transformer layers is far more critical to accuracy than prompt length or attaching prompts to later layers, as early intervention guides multimodal fusion before modality-specific traits are merged.
These results show that organizations can deploy resilient multimodal AI systems at a fraction of standard training, memory, and infrastructure costs. The approach eliminates the need to maintain separate, expensive models for every potential permutation of incomplete data. Organizations operating with limited computing budgets or large foundation models can prioritize lightweight prompt tuning over full retraining. When implementing this architecture, engineering teams should place prompt tokens within early transformer layers and choose attention-level prompting if target data lengths vary widely, or input-level prompting for maximum performance.
Confidence in these findings is high for vision-language classification benchmarks, but practical limitations remain. The evaluations focus on two-modality vision-text scenarios on a single base transformer model, leaving performance on more complex architectures, generative tasks, or three-way combinations (such as audio-video-text) unverified. Future work should validate the framework on larger billion-parameter foundations and operational production pipelines with unpredictable missing-data patterns.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Introduces Visual Prompt Tuning (VPT) for parameter-efficient adaptation of vision transformers, providing the foundational technique that the source adapts for missing-modality scenarios.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Pioneers continuous prompt tuning for pre-trained vision-language architectures, establishing the core parameter-efficient paradigm upon which multimodal prompt learning builds.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). Extends prompt learning with dynamic conditional mechanisms, providing key background on input-dependent prompt generation used in adaptive multimodal models.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Establishes the foundational mechanics and scaling benefits of parameter-efficient continuous prompt tuning in frozen transformer models.
- Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). Offers classical foundations on learning multimodal deep representations that remain robust when specific modalities are missing during inference.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). Describes fundamental alignment and fusion architectures for multimodal transformers, establishing the backbone design principles utilized in prompt-based multimodal learning.
- Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). Builds on parameter-efficient multimodal prompting by introducing interactive two-way deep prompt fusion between frozen unimodal vision and language transformers.
- Paper: MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning, Yi Xin et al. (2024). Extends multimodal prompt learning to multi-task and cross-domain adaptation while explicitly preserving cross-modal alignment across tasks.
- Paper: Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization, Jameel Abdul Samadh et al. (2023). Applies prompt learning at test time to resolve distribution shifts in multimodal foundation models without updating pre-trained weights.
- Paper: MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer, Jianjian Cao et al. (2024). Explores a complementary direction for efficient multimodal transformers by dynamically pruning aligned tokens across multimodal branches.
