Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal uncertainty-aware vision-language pre-training

Multimodal uncertainty-aware vision-language pre-training is a machine learning framework that trains computational models on paired visual and textual data while explicitly accounting for semantic ambiguity and multiple plausible interpretations within and between modalities. Rather than mapping images and text to static, deterministic vector embeddings, this approach represents inputs as probabilistic distributions across a shared semantic space to model the inherent uncertainty of cross-modal correspondence. By incorporating distribution-based pre-training objectives such as probabilistic contrastive learning, masked language modeling, and image-text alignment, the framework captures complex, non-deterministic relationships across visual and linguistic elements. This enables models to develop more expressive representations that improve robustness and generalization when adapted to downstream tasks such as image-text retrieval, visual question answering, and visual reasoning.

1 item

MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model

MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model

Yatai Ji, Junjie Wang, Yuan Gong, Lin Zhang, Yanru Zhu, Hongfa Wang, Jiaxing Zhang, Tetsuya Sakai, Yujiu Yang

OrganizationsInternational Digital Economy AcademyTencentTsinghua UniversityWaseda University

Why you should read this

Proposes a vision-language pre-training framework that models multimodal features as Gaussian distributions instead of deterministic points to capture inter- and intra-modal semantic uncertainty across downstream tasks like visual reasoning and image-text retrieval.

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pre-training on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pre-training frameworks and propose suitable pre-training tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM). The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results.

Added

2026-09-26