keyword
multimodal uncertainty-aware vision-language pre-training
Multimodal uncertainty-aware vision-language pre-training is a machine learning framework that trains computational models on paired visual and textual data while explicitly accounting for semantic ambiguity and multiple plausible interpretations within and between modalities. Rather than mapping images and text to static, deterministic vector embeddings, this approach represents inputs as probabilistic distributions across a shared semantic space to model the inherent uncertainty of cross-modal correspondence. By incorporating distribution-based pre-training objectives such as probabilistic contrastive learning, masked language modeling, and image-text alignment, the framework captures complex, non-deterministic relationships across visual and linguistic elements. This enables models to develop more expressive representations that improve robustness and generalization when adapted to downstream tasks such as image-text retrieval, visual question answering, and visual reasoning.
1 item

