MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
Yatai JiJunjie WangYuan GongLin ZhangYanru ZhuHongfa WangJiaxing ZhangTetsuya SakaiYujiu Yang
Proposes a vision-language pre-training framework that models multimodal features as Gaussian distributions instead of deterministic points to capture inter- and intra-modal semantic uncertainty across downstream tasks like visual reasoning and image-text retrieval.
Artificial intelligence systems for multimodal understanding face significant challenges with ambiguity and noise across visual and textual data. In real-world scenarios, an image region often contains multiple objects, and words can hold multiple meanings or synonyms, creating uncertainty both within individual modalities and between them. Most conventional vision-language models treat concepts as single fixed points in representation space, which fails to capture complex conceptual hierarchies and restricts the diversity of model predictions.
The main objective of the article is to develop and evaluate a vision-language pre-training framework that explicitly models semantic uncertainty by representing multimodal features as probability distributions rather than fixed points.
To achieve this, the article introduces the Multimodal Uncertainty-Aware Vision-Language Pre-training (MAP) model, centered around a newly designed Probability Distribution Encoder. This module models text tokens and image patches as multivariate Gaussian distributions by incorporating both sequence-level and feature-level interactions. The authors establish three distribution-based pre-training tasks to handle cross-modal alignment on large-scale unlabeled datasets: contrastive learning, masked language modeling, and image-text matching. Pre-trained on standard multimodal benchmarks—including MSCOCO, Visual Genome, SBU, and Conceptual Captions—the framework was subsequently fine-tuned and tested across multiple downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment.
The evaluations yielded several key findings. First, MAP achieved state-of-the-art performance across downstream tasks, notably outperforming comparable base-size models on visual question answering (78.03 on VQA2.0 test-dev) and visual reasoning (83.30 on NLVR2 dev). Second, in image retrieval on MSCOCO, MAP surpassed competitive baselines including ALBEF, even exceeding variants trained on significantly larger datasets of over 10 million images. Third, ablation analyses confirmed that distribution representations consistently outperformed deterministic point representations across all tasks, with masked language modeling proving to be the most vital pre-training objective. Finally, qualitative and toy analyses revealed that the learned distributions capture semantic overlap effectively and allow models to generate multiple diverse, plausible answers through sampling rather than being restricted to a single rigid prediction.
These results demonstrate that capturing semantic uncertainty substantially improves cross-modal understanding, representation robustness, and predictive flexibility. In practice, adopting distribution-based modeling allows organizations to achieve superior task accuracy without relying strictly on massive parameter scaling or exorbitantly large pre-training datasets, offering efficiency gains in data utilization. Furthermore, the capacity to produce diverse, valid outputs is particularly valuable for applications requiring nuanced reasoning or open-ended interactions.
Based on these findings, teams developing vision-language systems should consider adopting probabilistic encoders to handle real-world ambiguity in multimodal data. The article recommends utilizing Softmax-based sequence interactions within distribution encoders to best capture relational dependencies. Before broad commercial deployment, organizations should conduct pilot studies and further research exploring additional distribution subspaces and testing performance on larger, more diverse datasets. While confidence in the reported experimental improvements is high due to comprehensive baseline comparisons and statistical validation, current evaluations remain bounded by the specific pre-training corpora and distribution assumptions studied in the article.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). ALBEF established the standard paradigm of aligning image-text features before multimodal fusion alongside contrastive and masked modeling tasks that MAP directly probabilisticizes.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). UNITER formalized universal multimodal pre-training objectives (MLM, ITM, and contrastive alignment) that serve as the direct deterministic foundations adapted into MAP's distribution-based objectives.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). VisualBERT introduces the unified transformer baseline for vision-and-language tasks that MAP augments with uncertainty-aware probabilistic representations.
- Paper: What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?, Alex Kendall et al. (2017). This paper lays the conceptual groundwork for modeling aleatoric and epistemic uncertainty in deep vision architectures.
- Paper: Robust Cross-Modal Representation Learning with Progressive Self-Distillation, Alex Andonian et al. (2022). This work demonstrates how to handle noisy many-to-many cross-modal correspondences using soft alignment, directly motivating MAP's distribution-based cross-modal modeling.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This survey provides the foundational taxonomy for multimodal representation, alignment, and fusion challenges that MAP's probability distribution encoder aims to address.
- Paper: SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger, Yuting Gao et al. (2024). SoftCLIP extends the mitigation of rigid one-to-one multimodal alignment constraints by leveraging softened intra-modal targets for cross-modal contrastive representation learning.
- Paper: Provable Dynamic Fusion for Low-Quality Multimodal Data, Qingyang Zhang et al. (2023). This work advances dynamic multimodal fusion under low-quality, uncertain data by providing rigorous theoretical justifications for robust cross-modal interaction.
- Paper: Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization, Jameel Abdul Samadh et al. (2023). PromptAlign applies distribution-level alignment at test time to vision-language representations to resolve uncertainty and performance degradation caused by domain shifts.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). PreFLMR scales up fine-grained cross-modal retrieval by deploying late-interaction token alignments across diverse downstream multimodal tasks.
