keyword
retrieval augmented contrastive language-image pre-training
Retrieval augmented contrastive language-image pre-training is a multimodal artificial intelligence training approach that enhances vision-language models by integrating an external retrieval mechanism into the contrastive learning process. Instead of forcing a neural network to memorize all visual concepts and semantic relationships entirely within its internal parameters, the framework queries a reference dataset of image-text pairs during training to fetch contextually relevant examples. These retrieved references are used to enrich the feature representations of input data, allowing the model to leverage explicit cross-modal correspondences rather than purely parameterized memory. By operating like an open-book reference system, this method improves data efficiency and enhances performance on downstream visual tasks such as zero-shot image classification and object recognition.
1 item

