Built independently by an author, for readers. Read the story and support ChapterPal

keyword

retrieval augmented contrastive language-image pre-training

Retrieval augmented contrastive language-image pre-training is a multimodal artificial intelligence training approach that enhances vision-language models by integrating an external retrieval mechanism into the contrastive learning process. Instead of forcing a neural network to memorize all visual concepts and semantic relationships entirely within its internal parameters, the framework queries a reference dataset of image-text pairs during training to fetch contextually relevant examples. These retrieved references are used to enrich the feature representations of input data, allowing the model to leverage explicit cross-modal correspondences rather than purely parameterized memory. By operating like an open-book reference system, this method improves data efficiency and enhances performance on downstream visual tasks such as zero-shot image classification and object recognition.

1 item

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

Chen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng, Deli Zhao, Jingren Zhou

OrganizationsAlibaba Group

Why you should read this

Proposes an open-book contrastive learning framework that uses online image-text retrieval from a reference set to augment visual embeddings, boosting zero-shot classification performance without requiring models to memorize vast training concepts.

Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we propose a novel and efficient framework: Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) to augment embeddings by online retrieval. Specifically, we sample part of image-text data as a hold-out reference set. Given an input image, relevant image-text pairs are retrieved from the reference set to enrich the representation of input image. This process can be considered as an open-book exam: with the reference set as a cheat sheet, the proposed method doesn't need to memorize all visual concepts in the training data. It explores how to recognize visual concepts by exploiting correspondence between images and texts in the cheat sheet. The proposed RA-CLIP implements this idea and comprehensive experiments are conducted to show how RA-CLIP works. Performances on 10 image classification datasets and 2 object detection datasets show that RA-CLIP outperforms vanilla CLIP baseline by a large margin on zero-shot image classification task (+12.7%), linear probe image classification task (+6.9%) and zero-shot ROI classification task (+2.8%).

Added

2026-09-26