Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RA-CLIP

RA-CLIP, short for Retrieval Augmented Contrastive Language-Image Pre-Training, is a multimodal artificial intelligence framework that enhances vision-language models by retrieving relevant image-text data to augment representations during learning and inference. Unlike standard contrastive models that rely purely on internal parameters to memorize vast amounts of visual concepts from scratch, RA-CLIP utilizes an external reference repository of paired visual and textual examples to help align cross-modal features. By leveraging this external knowledge base to provide context for input images, the framework reduces the memorization burden of the network while improving accuracy and sample efficiency on downstream visual recognition tasks such as zero-shot image classification.

1 item

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

Chen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng, Deli Zhao, Jingren Zhou

OrganizationsAlibaba Group

Why you should read this

Proposes an open-book contrastive learning framework that uses online image-text retrieval from a reference set to augment visual embeddings, boosting zero-shot classification performance without requiring models to memorize vast training concepts.

Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we propose a novel and efficient framework: Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) to augment embeddings by online retrieval. Specifically, we sample part of image-text data as a hold-out reference set. Given an input image, relevant image-text pairs are retrieved from the reference set to enrich the representation of input image. This process can be considered as an open-book exam: with the reference set as a cheat sheet, the proposed method doesn't need to memorize all visual concepts in the training data. It explores how to recognize visual concepts by exploiting correspondence between images and texts in the cheat sheet. The proposed RA-CLIP implements this idea and comprehensive experiments are conducted to show how RA-CLIP works. Performances on 10 image classification datasets and 2 object detection datasets show that RA-CLIP outperforms vanilla CLIP baseline by a large margin on zero-shot image classification task (+12.7%), linear probe image classification task (+6.9%) and zero-shot ROI classification task (+2.8%).

Added

2026-09-26