RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

Chen-Wei XieSiyang SunXiong XiongYun ZhengDeli ZhaoJingren Zhou

article2023CVPR60 citations

Proposes an open-book contrastive learning framework that uses online image-text retrieval from a reference set to augment visual embeddings, boosting zero-shot classification performance without requiring models to memorize vast training concepts.

Listen

Modern vision-language systems such as Contrastive Language-Image Pre-training (CLIP) have transformed visual recognition by enabling models to recognize new visual concepts without task-specific training data. However, standard models require massive datasets of tens to hundreds of millions of image-text pairs to memorize concepts directly within their model parameters. This excessive data requirement makes training prohibitively expensive for most organizations and research teams.

The article demonstrates and evaluates Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP), a framework designed to make vision-language training substantially more data-efficient. The core objective is to show that providing an external reference dataset during evaluation allows a visual model to look up relevant descriptions rather than memorizing every concept internally.

The researchers implemented this approach by splitting image-text data into a primary training set and a hold-out reference set consisting of 1.6 million image-text pairs, roughly one-tenth of the total training volume. When processing an input image, the system uses an efficient image retriever to search the reference set for the top matching image-text pairs. A multi-head cross-attention module then integrates textual and visual information from these retrieved pairs directly into the input image representation. The system was pre-trained using 15 million image-text pairs from a standard public dataset and evaluated across ten image classification benchmarks and two object detection benchmarks.

The findings show that the proposed framework consistently outperforms standard methods across diverse recognition tasks. First, the retrieval-augmented framework achieved an average zero-shot image classification accuracy of 52.0% across ten benchmarks, outperforming standard baseline models by an absolute margin of 12.7% and beating previous state-of-the-art methods such as DeCLIP and SLIP. Second, on the standard ImageNet benchmark, zero-shot accuracy rose from 37.7% to 53.5% under the same training budget, and scaling the reference pool from 1,000 to 10 million pairs produced steady accuracy gains. Third, the framework improved linear probe classification accuracy by an average of 6.9% over standard baselines (reaching 75.5%) and improved zero-shot region-of-interest detection performance on standard object detection benchmarks, particularly for small and medium objects.

These results demonstrate that treating concept recognition as an open-book lookup rather than internal memorization significantly improves data and computational efficiency. Organizations can deploy higher-performing vision-language models at reduced pre-training data scale, lowering infrastructure costs and development timelines. Furthermore, the findings show that external reference sets from different standard image-text collections perform robustly without requiring dataset-specific tuning.

Organizations developing or deploying visual AI systems should adopt retrieval-augmented architectures to improve performance when training data budgets are constrained. Teams should use standard pre-trained uni-modal encoders to extract reference embeddings offline, minimizing runtime latency and computational overhead. When implementing the retrieval module, practitioners should focus retrieval augmentation primarily on the image representations rather than text representations, as the article found that text-to-text augmentation introduces noise.

Decision-makers should note certain operational boundaries: the retrieval mechanism adds minor system complexity and search latency during inference, and performance gains begin to show diminishing returns as the reference dataset expands past several million pairs. Confidence in the empirical results is high given rigorous benchmarking across twelve diverse datasets, but organizations should conduct pilot testing to optimize retrieval indexing and latency for real-time edge applications.

Cover for RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

Abstract

Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we propose a novel and efficient framework: Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) to augment embeddings by online retrieval. Specifically, we sample part of image-text data as a hold-out reference set. Given an input image, relevant image-text pairs are retrieved from the reference set to enrich the representation of input image. This process can be considered as an open-book exam: with the reference set as a cheat sheet, the proposed method doesn't need to memorize all visual concepts in the training data. It explores how to recognize visual concepts by exploiting correspondence between images and texts in the cheat sheet. The proposed RA-CLIP implements this idea and comprehensive experiments are conducted to show how RA-CLIP works. Performances on 10 image classification datasets and 2 object detection datasets show that RA-CLIP outperforms vanilla CLIP baseline by a large margin on zero-shot image classification task (+12.7%), linear probe image classification task (+6.9%) and zero-shot ROI classification task (+2.8%).

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Contrastive Language Image Pre-training
  • 2.2. Knowledge-Enhanced Models
  • 2.3. Zero-shot Visual Recognition
  • 3. Method
  • 3.1. Dual-Encoder Architecture
  • 3.2. Overview of RA-CLIP
  • 3.3. Loss Function.
  • 4. Experiment
  • 4.1. Implementation Details
  • 4.2. Ablation Study
  • 4.2.1 Effectiveness of the Retrieval Augmented CLIP
  • 4.2.2 Different image-text data as reference set
  • 4.2.3 Ablation of pre-trained SentenceT and DINO
  • 4.2.4 Different hyper-parameters and design choices
  • 4.3. Main Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) Framework

    model/method

    Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) is an open-vocabulary vision-language representation learning framework that augments target image representations via online multi-modal retrieval. Given an initial collection of NN image-text pairs {(Ii,Ti)}i=1N\{(I_i, T_i)\}_{i=1}^N, the dataset is partitioned into a training set T\mathcal{T} and a disjoint hold-out reference set R\mathcal{R}.

    For an input image IiI_i, an initial normalized visual embedding vi∈Rdv_i \in \mathbb{R}^d is generated by a trainable Vision Transformer image encoder. Concurrently, a frozen, self-supervised uni-modal image retriever ϕ\phi (e.g., DINO-S/8) extracts an offline index of image embeddings for all pairs in R\mathcal{R}. Using the ϕ\phi-embedding of query IiI_i, the top-KK nearest reference images {rkI}k=1K\{r_k^I\}_{k=1}^K and their paired captions {rkT}k=1K\{r_k^T\}_{k=1}^K are retrieved from R\mathcal{R}.

    A Retrieval Augmented Module (RAM) fuses the query embedding viv_i with the retrieved image embeddings ekI=ϕ(rkI)e_k^I = \phi(r_k^I) and frozen text embeddings ekT=ψ(rkT)e_k^T = \psi(r_k^T) (extracted by a uni-modal text encoder ψ\psi) via multi-head cross-attention, yielding an augmented visual representation vi′v'_i. The augmented image embedding vi′v'_i is aligned with the corresponding normalized text embedding ti∈Rdt_i \in \mathbb{R}^d (produced by a trainable Transformer text encoder) using symmetric contrastive InfoNCE loss. This decouples visual concept recognition from pure parametric memorization.

  2. Knowl 2 — Retrieval Augmented Module (RAM) Cross-Attention Formulations

    equation

    The Retrieval Augmented Module (RAM) constructs an augmented image representation vi′v'_i from the initial image representation vi∈Rdv_i \in \mathbb{R}^d and the retrieved reference pairs {(rkI,rkT)}k=1K\{(r_k^I, r_k^T)\}_{k=1}^K. Reference visual embeddings {ekI}k=1K\{e_k^I\}_{k=1}^K and textual embeddings {ekT}k=1K\{e_k^T\}_{k=1}^K are extracted via frozen uni-modal encoders ϕ\phi and ψ\psi:

    ekI=ϕ(rkI),ekT=ψ(rkT)e_k^I = \phi(r_k^I), \quad e_k^T = \psi(r_k^T)

    RAM calculates a text-augmented context vector aiTa_i^T by querying the retrieved reference text features using the visual query viv_i conditioned on reference image keys:

    aiT=MultiheadAttn(vi,{ekI}k=1K,{ekT}k=1K)a_i^T = \text{MultiheadAttn}\left(v_i, \{e_k^I\}_{k=1}^K, \{e_k^T\}_{k=1}^K\right)

    where viv_i serves as the query (QQ), {ekI}k=1K\{e_k^I\}_{k=1}^K serves as the key (KK), and {ekT}k=1K\{e_k^T\}_{k=1}^K serves as the value (VV).

    Symmetrically, RAM computes an image-augmented context vector aiIa_i^I using the reference text keys and reference image values:

    aiI=MultiheadAttn(vi,{ekT}k=1K,{ekI}k=1K)a_i^I = \text{MultiheadAttn}\left(v_i, \{e_k^T\}_{k=1}^K, \{e_k^I\}_{k=1}^K\right)

    where viv_i serves as the query (QQ), {ekT}k=1K\{e_k^T\}_{k=1}^K serves as the key (KK), and {ekI}k=1K\{e_k^I\}_{k=1}^K serves as the value (VV).

    The final augmented image embedding vi′v'_i is formed by a residual fusion across all three representations:

    vi′=vi+aiT+aiIv'_i = v_i + a_i^T + a_i^I

  3. Knowl 3 — Contrastive Training Objective for RA-CLIP

    equation

    Given a mini-batch of NN image-text pairs producing augmented image representations {vi′}i=1N\{v'_i\}_{i=1}^N and text embeddings {ti}i=1N\{t_i\}_{i=1}^N, RA-CLIP minimizes the bidirectional InfoNCE contrastive loss L=Lv2t+Lt2v\mathcal{L} = \mathcal{L}_{v2t} + \mathcal{L}_{t2v}. The image-to-text loss Lv2t\mathcal{L}_{v2t} and text-to-image loss Lt2v\mathcal{L}_{t2v} are defined as:

    Lv2t=−1N∑i=1Nlog⁡(exp⁡(σ(ti,vi′)/τ)∑j=1Nexp⁡(σ(ti,vj′)/τ))\mathcal{L}_{v2t} = - \frac{1}{N} \sum_{i=1}^N \log \left( \frac{\exp(\sigma(t_i, v'_i) / \tau)}{\sum_{j=1}^N \exp(\sigma(t_i, v'_j) / \tau)} \right)

    Lt2v=−1N∑i=1Nlog⁡(exp⁡(σ(vi′,ti)/τ)∑j=1Nexp⁡(σ(vi′,tj)/τ))\mathcal{L}_{t2v} = - \frac{1}{N} \sum_{i=1}^N \log \left( \frac{\exp(\sigma(v'_i, t_i) / \tau)}{\sum_{j=1}^N \exp(\sigma(v'_i, t_j) / \tau)} \right)

    where σ(x,y)=x⊤y∥x∥2∥y∥2\sigma(x, y) = \frac{x^\top y}{\|x\|_2 \|y\|_2} denotes cosine similarity, and τ∈R+\tau \in \mathbb{R}^+ is a learnable temperature parameter initialized to 0.070.07 and optimized end-to-end.

  4. Knowl 4 — Zero-Shot and Linear Probe Image Classification Performance on 10 Benchmarks

    data/table

    The table below compares RA-CLIP against vanilla CLIP and state-of-the-art vision-language pre-training methods (MS-CLIP, SLIP, DeCLIP, K-Lite) across 10 image classification datasets. Models are pre-trained on YFCC15M (with 1.6M reference pairs for RA-CLIP) using a ViT-B/32 image encoder. RA-CLIP* denotes the variant training a custom text encoder (extwidth=512 ext{width} = 512, extheads=8 ext{heads} = 8) from scratch for matched architecture comparison against DeCLIP.

    Method ImageNet ImageNetV2 Pets CIFAR10 CIFAR100 SUN397 Food101 Caltech101 DTD Dogs Avg.
    Zero-Shot Accuracy (%)
    K-Lite 45.3 – – – – – – – – – –
    CLIP 37.7 32.8 16.1 76.0 48.6 50.8 21.8 69.8 28.3 11.1 39.3
    MS-CLIP 36.7 30.2 – – – – – – – 5.6 –
    SLIP 38.3 33.3 28.3 72.2 45.3 45.1 44.7 65.9 21.8 11.8 40.7
    DeCLIP 43.2 36.1 30.2 72.1 39.7 51.6 46.9 70.1 24.2 11.7 42.6
    RA-CLIP* 51.2 45.4 50.5 89.4 61.8 45.7 43.9 76.1 24.6 22.0 51.1
    RA-CLIP 53.5 47.2 49.0 89.4 62.3 46.5 43.8 76.9 25.6 26.1 52.0
    Linear Probe Accuracy (%)
    CLIP 63.5 51.3 69.8 91.7 74.1 64.7 69.1 84.9 66.5 50.5 68.6
    MS-CLIP 68.1 49.8 62.1 87.2 66.7 71.7 76.0 81.6 69.4 46.1 67.9
    SLIP 68.1 52.1 75.4 90.5 75.3 73.5 77.1 87.2 71.1 52.6 72.3
    DeCLIP 69.2 53.1 76.5 88.6 71.6 75.9 79.3 88.0 69.1 49.9 72.1
    RA-CLIP* 73.3 62.3 88.2 94.9 78.4 60.7 65.5 86.8 65.5 75.5 75.1
    RA-CLIP 72.9 61.9 88.2 95.2 78.9 61.1 66.5 87.2 66.6 76.0 75.5

    RA-CLIP improves zero-shot classification over vanilla CLIP by +12.7%+12.7\% on average and outperforms DeCLIP by +9.4%+9.4\%. For linear probe evaluation, RA-CLIP achieves an average accuracy of 75.5%75.5\%, outperforming vanilla CLIP by +6.9%+6.9\% and prior state-of-the-art representations by +3.2%+3.2\%.

  5. Knowl 5 — Ablations on Reference Datasets and Uni-Modal Encoder Modules

    data/table

    The table below presents ablation experiments on ImageNet zero-shot top-1 accuracy to isolate the contributions of reference dataset sources, encoder initializations, and uni-modal encoders (ϕ\phi and ψ\psi).

    ID Method Init. Img Enc Init. Txt Enc Pretrain Set Ref Set ϕ\phi ψ\psi ImageNet Top-1 (%)
    1 CLIP ViT rand. BERT YFCC None None None 37.7
    2 CLIP DINO-S SentenceT YFCC None None None 21.0
    3 CLIP ViT IN1K BERT YFCC None None None 46.1
    4 CLIP ViT rand. BERT YFCC+CC None None None 42.1
    5 RA-CLIP ViT rand. BERT YFCC YFCC SentenceT DINO-S 53.5
    6 RA-CLIP ViT rand. BERT YFCC CC SentenceT DINO-S 54.5
    7 RA-CLIP ViT rand. BERT YFCC LAION SentenceT DINO-S 54.2
    8 RA-CLIP ViT rand. BERT YFCC CC Text Encoder DINO-S 54.4

    Key observations include:

    1. Freezing uni-modal representations (DINO-S and SentenceT) directly in standard CLIP without cross-attention augmentation yields poor multi-modal alignment (21.0%21.0\% Top-1).
    2. Standard CLIP initialized with supervised ImageNet-1K pre-trained weights reaches 46.1%46.1\%, while RA-CLIP reaches 53.5%53.5\% without ImageNet pre-training supervision.
    3. Using external reference sets (CC12M or LAION at 1.6M pairs) provides consistent gains (54.5%54.5\% and 54.2%54.2\%) comparable to in-domain YFCC reference pairs (53.5%53.5\%).
    4. Substituting SentenceT with RA-CLIP's own trainable text encoder yields comparable accuracy (54.4%54.4\% vs. 54.5%54.5\%), confirming SentenceT is not strictly necessary but allows pre-computing offline reference embeddings.
  6. Knowl 6 — Zero-Shot Region-of-Interest (ROI) Classification on COCO and LVIS

    data/table

    Zero-shot object detection region classification evaluated on the validation splits of MS COCO and LVIS benchmarks following the RegionCLIP protocol using ground-truth region proposal bounding boxes.

    LVIS COCO
    Method AP APs\text{AP}_s APm\text{AP}_m APl\text{AP}_l AP APs\text{AP}_s APm\text{AP}_m APl\text{AP}_l
    Region CLIP 21.6 8.7 31.0 45.7 44.4 21.9 51.0 61.8
    Region RA-CLIP 23.2 10.9 34.2 44.9 48.4 29.3 57.9 61.9

    Region RA-CLIP achieves a +1.6%+1.6\% AP gain on LVIS and +4.0%+4.0\% AP gain on COCO over Region CLIP. The largest improvements occur on small objects (APs\text{AP}_s: +2.2%+2.2\% on LVIS, +7.4%+7.4\% on COCO) and medium objects (APm\text{AP}_m: +3.2%+3.2\% on LVIS, +6.9%+6.9\% on COCO).

  7. Knowl 7 — Ablation of Retrieval Hyperparameters and RAM Fusion Designs

    empirical result

    Ablation experiments on ImageNet zero-shot classification (trained on YFCC15M with 1.6M CC12M reference pairs) evaluate RAM feature fusion components, retrieval size KK, and reference set scale:

    1. Feature Fusion Type: Using only retrieved text contexts aiTa_i^T yields 52.1%52.1\% top-1 accuracy; combining both retrieved contexts aiT+aiIa_i^T + a_i^I yields 51.8%51.8\%; and full residual fusion aiT+aiI+via_i^T + a_i^I + v_i achieves the highest accuracy of 54.5%54.5\%.
    2. Retrieval Number KK: For residual fusion aiT+aiI+via_i^T + a_i^I + v_i, setting K=16K=16 yields 54.3%54.3\%, K=64K=64 yields 54.5%54.5\%, and K=128K=128 yields 53.9%53.9\%, showing robustness across retrieval set sizes with optimal results at K=64K=64.
    3. Reference Set Scale: Evaluated across reference dataset sizes from CC12M, ImageNet top-1 zero-shot accuracy scales logarithmically with the number of reference pairs: 1K→46.2%1\text{K} \to 46.2\%, 10K→47.7%10\text{K} \to 47.7\%, 100K→50.1%100\text{K} \to 50.1\%, 1.6M→54.5%1.6\text{M} \to 54.5\%, and 10M→56.7%10\text{M} \to 56.7\%.
  8. Knowl 8 — Symmetric Text Representation Augmentation Degradation

    empirical result

    Applying retrieval augmentation symmetrically to both image and text branches degrades zero-shot ImageNet top-1 accuracy compared to visual-only augmentation.

    When text query sentences retrieve top-KK reference image-text pairs via Sentence Transformer and augment text token embeddings using cross-attention analogously to the visual branch, performance drops from 54.5%54.5\% to 53.1%53.1\%. This performance degradation occurs because text prompts are compact and semantically less rich than images; text-based retrieval yields higher semantic variance and noisy reference pairs that dilute the specific category semantics.

  9. Knowl 9 — RA-CLIP Pre-training Implementation Setup and Architecture Configuration

    experimental setup

    The standard implementation of RA-CLIP consists of the following components and optimization parameters:

    • Image Encoder: Vision Transformer ViT-B/32 (12 layers, 12 attention heads, hidden dimension 768), randomly initialized; input resolution is 224×224224 \times 224 pixels.
    • Text Encoder: BERT-base (12 layers, 12 attention heads, hidden dimension 768) initialized with BERT weights, or trained from scratch with 512 width and 8 heads (for RA-CLIP*); input tokens are truncated to a maximum length of 77.
    • Projection Layers: Linear projections map image and text representations to a joint d=384d = 384-dimensional embedding space.
    • Offline Encoders: DINO-S/8 pre-trained self-supervised on ImageNet-1K serves as the uni-modal image retriever and encoder ϕ\phi. A 6-layer Sentence Transformer (SentenceT) serves as the uni-modal text encoder ψ\psi.
    • Retrieval Augmented Module (RAM): 6 layers of multi-head cross-attention blocks.
    • Training Details: Trained for 32 epochs using the LAMB optimizer with an initial learning rate of 2.5×10−32.5 \times 10^{-3}, cosine decay schedule, 5 epochs of linear warmup, weight decay of 0.20.2, and batch size of 4096 across 8 NVIDIA Tesla A100 GPUs.
  10. Knowl 10 — Retrieval Augmentation on OpenAI Pre-trained Foundation CLIP

    empirical result

    RA-CLIP can be applied to large-scale pre-trained foundation models without full end-to-end retraining. Initializing the image encoder, text encoder, ϕ\phi, and ψ\psi with OpenAI's CLIP-B/32 (originally pre-trained on 400 million image-text pairs), freezing these backbones, and training only the Retrieval Augmented Module (RAM) and projection layers on YFCC15M improves ImageNet zero-shot top-1 accuracy from 63.3%63.3\% (vanilla OpenAI CLIP-B/32) to 68.2%68.2\% (a +4.9%+4.9\% absolute gain).

Coverage note — Qualitative visual retrieval case study examples (Figure 5) were excluded as standalone knowls since their underlying concepts and mechanism are fully characterized in the RAM architecture and quantitative evaluation knowls.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. CoRR, 2022.
  2. 2.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. Improving language models by retrieving from trillions of tokens. CoRR, 2022.
  3. 3.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Eur. Conf. Comput. Vis., pages 446–461, 2014.
  4. 4.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision Transformers. In Int. Conf. Comput. Vis., pages 9650–9660, 2021.
  5. 5.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3558–3568, 2021.
  6. 6.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In Int. Conf. Mach. Learn., pages 1597–1607, 2020.
  7. 7.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO captions: Data collection and evaluation server. CoRR, 2015.
  8. 8.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3606–3613, 2014.
  9. 9.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In IEEE Conf. Comput. Vis. Pattern Recog. Worksh., pages 702–703, 2020.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In IEEE Conf. Comput. Vis. Pattern Recog., pages 248–255, 2009.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional Transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4171–4186, 2019.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Int. Conf. Learn. Represent., 2020.
  13. 13.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In IEEE Conf. Comput. Vis. Pattern Recog. Worksh., pages 178–178, 2004.
  14. 14.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., volume 00, pages 580–587, 2014.
  15. 15.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 5356–5364, 2019.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In IEEE Conf. Comput. Vis. Pattern Recog., pages 770–778, 2016.
  17. 17.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. Few-shot learning with retrieval augmented language models. CoRR, 2022.
  18. 18.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Int. Conf. Mach. Learn., pages 4904–4916. PMLR, 2021.
  19. 19.Jeff Johnson, Matthijs Douze, and Herve Jegou. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 2017.
  20. 20.Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Fei-Fei Li. Novel dataset for fine-grained image categorization: Stanford dogs. In Proc. CVPR workshop on fine-grained visual categorization (FGVC), volume 2, 2011.
  21. 21.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  22. 22.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet classification with deep convolutional neural networks. In Adv. Neural Inform. Process. Syst., 2012.
  23. 23.Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin, Chenguang Zhu, Michael Zeng, Heng Ji, and Shih-Fu Chang. CLIP-Event: Connecting text and images with event structures. In IEEE Conf. Comput. Vis. Pattern Recog., pages 16399–16408, 2022.
  24. 24.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In Int. Conf. Learn. Represent., 2022.
  25. 25.Norman Mu, Alexander Kirillov, David A. Wagner, and Saining Xie. SLIP: self-supervision meets language-image pre-training. In Eur. Conf. Comput. Vis., volume 13686, pages 529–544, 2022.
  26. 26.Antonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella, Emanuele Rodola, and Francesco Locatello. ASIF: Coupled data turns unimodal models to multimodal without training. arXiv preprint arXiv:2210.01738, 2022.
  27. 27.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3498–3505, 2012.
  28. 28.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In Adv. Neural Inform. Process. Syst., pages 8024–8035, 2019.
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Int. Conf. Mach. Learn., pages 8748–8763. PMLR, 2021.
  30. 30.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In Int. Conf. Mach. Learn., volume 97, pages 5389–5400, 2019.
  31. 31.Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2019.
  32. 32.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs. CoRR, 2021.
  33. 33.Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Anna Rohrbach, Zhe Gan, Lijuan Wang, Lu Yuan, Ce Liu, Kurt Keutzer, Trevor Darrell, and Jianfeng Gao. K-LITE: Learning transferable visual models with external knowledge. Adv. Neural Inform. Process. Syst., 2022.
  34. 34.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Int. Conf. Learn. Represent., 2014.
  35. 35.Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. YFCC100M: the new data in multimedia research. Commun. ACM, 59(2):64–73, 2016.
  36. 36.Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. CoRR, 2018.
  37. 37.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Adv. Neural Inform. Process. Syst., 2017.
  38. 38.Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva. SUN database: Exploring a large collection of scene categories. Int. J. Comput. Vis., 119(1):3–22, 2016.
  39. 39.Haoxuan You, Luowei Zhou, Bin Xiao, Noel Codella, Yu Cheng, Ruochen Xu, Shih-Fu Chang, and Lu Yuan. Learning visual representation from modality-shared contrastive language-image pre-training. In Eur. Conf. Comput. Vis., pages 69–87. Springer, 2022.
  40. 40.Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training BERT in 76 minutes. In Int. Conf. Learn. Represent., 2020.
  41. 41.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. CoCa: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research, 2022.
  42. 42.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang. Florence: A new foundation model for computer vision. CoRR, 2021.
  43. 43.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. RegionCLIP: Region-based language-image pretraining. In IEEE Conf. Comput. Vis. Pattern Recog., pages 16793–16803, 2022.

Citation

MLA
Xie, C.-W., et al. “RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 19265–74, https://doi.org/10.1109/CVPR52729.2023.01846.
APA
Xie, C.-W., Sun, S., Xiong, X., Zheng, Y., Zhao, D., & Zhou, J. (2023). RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19265–19274. https://doi.org/10.1109/CVPR52729.2023.01846
Chicago
Xie, C.-W., S. Sun, X. Xiong, Y. Zheng, D. Zhao, and J. Zhou. 2023. “RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19265–74. https://doi.org/10.1109/CVPR52729.2023.01846.
Harvard
Xie, C.-W. et al. (2023) “RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 19265–19274. Available at: https://doi.org/10.1109/CVPR52729.2023.01846.
Vancouver
1. Xie C-W, Sun S, Xiong X, Zheng Y, Zhao D, Zhou J (2023) RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 19265–19274

BibTeX

@inproceedings{Xie_2023, title={RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01846}, DOI={10.1109/cvpr52729.2023.01846}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Xie, Chen-Wei and Sun, Siyang and Xiong, Xiong and Zheng, Yun and Zhao, Deli and Zhou, Jingren}, year={2023}, month=June, pages={19265–19274} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE