keyword
vision-language contrastive learning
Vision-language contrastive learning is a multimodal machine learning technique that trains computational models to understand the relationships between visual data, such as images or video frames, and textual descriptions. In this approach, separate neural network encoders map visual and language inputs into a shared embedding space using a contrastive objective, which pulls corresponding image-text pairs closer together while pushing non-matching pairs farther apart. By aligning cross-modal semantic representations at scale without requiring labor-intensive manual labels, this pre-training method enables models to learn rich, generalized representations that transfer effectively to downstream tasks such as cross-modal retrieval, zero-shot image classification, visual question answering, and visual reasoning.
1 item

