keyword
Multimodal contrastive loss
Multimodal contrastive loss is an objective function in machine learning designed to align representations from multiple distinct data modalities, such as text, images, or audio, within a shared embedding space. It operates by maximizing the similarity between representations of semantically corresponding cross-modal pairs while minimizing the similarity between unrelated or non-matching pairs. By encouraging the geometric alignment of paired inputs across different sensory channels, this loss function enables neural network encoders to map heterogeneous data into a unified latent space that preserves shared semantic relationships, supporting downstream applications such as cross-modal retrieval, classification, and robust inference under missing modalities.
1 item

