Retrieval-Augmented Multimodal Language Modeling

Michihiro YasunagaArmen AghajanyanWeijia ShiRichard JamesJure LeskovecPercy LiangMike LewisLuke ZettlemoyerWen-Tau Yih

article2023ICML172 citations

Proposes RA-CM3, a retrieval-augmented multimodal architecture that fetches relevant text-image documents to generate both modalities, significantly cutting training compute while outperforming models like DALL-E and enabling multimodal in-context learning.

Listen

Modern multimodal artificial intelligence systems, such as models that generate images from text or descriptions from images, traditionally store all their world knowledge directly within their neural network parameters. As a result, expanding their knowledge base or updating information requires exponentially larger models and massive training datasets, which dramatically increases computing costs. The article evaluates a modular approach called retrieval augmentation, demonstrating that connecting a base multimodal generative model to an external document memory allows it to reference external text and images dynamically, rather than relying solely on memorized parameters.

To demonstrate this capability, the authors developed Retrieval-Augmented CM3 (RA-CM3), the first model capable of both retrieving and generating mixed combinations of text and images. The system pairs an off-the-shelf multimodal retriever based on CLIP with a 2.7-billion-parameter CM3 Transformer generator. The model was trained from scratch on 150 million web-collected text-image pairs from the LAION dataset using 256 GPUs over five days, using the exact same dataset as both training data and external memory to ensure a controlled evaluation against non-retrieval baselines.

The investigation produced several key findings regarding model efficiency and performance. First, RA-CM3 significantly outperformed baseline models without retrieval on standard benchmarks, improving image generation quality by approximately 13 to 14 FID points (reducing the score from 29.5 to 15.7, where lower is better) and caption quality by over 17 CIDEr points (increasing from 71.9 to 89.1). Second, RA-CM3 achieved these superior results while requiring less than 30% of the training compute and parameters utilized by comparable autoregressive systems like DALL-E (12B). Third, the architecture proved substantially more capable at generating rare or knowledge-intensive concepts, such as correctly rendering historical artifacts and composite scenes where standard generative models hallucinate or substitute common defaults. Finally, the training process enabled zero-shot and few-shot in-context learning, allowing users to control image styling directly via image prompts and improving few-shot classification accuracy from 0.78 at one-shot to 0.90 at eight-shot.

These findings indicate that retrieval augmentation offers a path to lower training expenses, faster development cycles, and higher factual reliability. Rather than expanding model scale to memorize rare entities, organizations can utilize external memory stores that are easily updated without retraining the entire neural network. Furthermore, retrieving source documents inherently provides provenance, improving interpretability and reducing unintended hallucinations in generative outputs.

Moving forward, technical teams should consider retrieval-augmented architectures when designing multimodal systems that require frequent knowledge updates, rare entity processing, or budget-constrained compute budgets. Recommended subsequent steps include investigating retriever fine-tuning rather than relying on frozen retrievers, expanding the framework to additional modalities beyond text and image pairs, and utilizing ensembling techniques to incorporate larger numbers of retrieved reference examples.

Decision-makers should note certain limitations: the model is a research prototype trained on filtered web data, meaning risks of biased or unsafe outputs persist. In addition, context length constraints in current sequence architectures practically limit direct retrieval to one or two documents per pass before needing ensemble mechanisms. Nevertheless, the reported improvements in training efficiency and generation fidelity are supported by consistent scaling behavior across model sizes, providing high confidence in the fundamental advantages of retrieval-augmented multimodal modeling.

arXiv: 2211.12561
Cover for Retrieval-Augmented Multimodal Language Modeling

Abstract

Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all their knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). Specifically, for the retriever, we use a pretrained CLIP, and for the generator, we train a CM3 Transformer on the LAION dataset. Our resulting model, named Retrieval-Augmented CM3 (RA-CM3), is the first multimodal model that can retrieve and generate both text and images. We show that RA-CM3 significantly outperforms baseline multimodal models such as DALL-E and CM3 on both image and caption generation tasks (12 FID and 17 CIDEr improvements on MS-COCO), while requiring much less compute for training (<30% of DALL-E). Moreover, we show that RA-CM3 exhibits novel capabilities, such as faithful image generation and multimodal in-context learning (e.g., image generation from demonstrations).

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Approach
  • 3.1. Preliminaries
  • 3.2. Multimodal retrieval
  • 3.3. Multimodal generator
  • 3.4. Training and inference
  • 4. Experiments
  • 4.1. Training setup
  • 4.2. Evaluation setup
  • 4.3. Main results
  • 4.4. Analysis
  • 5. Qualitative results
  • 5.1. Knowledge-intensive multimodal generation
  • 5.2. Image infilling and editing
  • 5.3. Controlled image generation
  • 5.4. One-shot and few-shot image classification
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Ethics and societal impact
  • B. Related work
  • C. Additional results
  • C.1. Intrinsic evaluation of CLIP-based retriever
  • C.2. Scaling laws of RA-CM3
  • C.3. Analysis of RA-CM3 designs
  • D. Additional discussions
  • D.1. Fair comparison of the retrieval-augmented model and non-retrieval-augmented model
  • D.2. Taking an existing model (e.g. vanilla CM3) and finetune it with retrieval-augmentation, instead of training the retrieval-augmented model (RA-CM3) from scratch
  • D.3. How the number of retrieved documents used for the generator (K) was set

Knowls

  1. Knowl 1 — Retrieval-Augmented Multimodal Modeling Architecture (RA-CM3)

    model/method

    The Retrieval-Augmented Causal Masked Multimodal (RA-CM3) architecture generalizes retrieval-augmented language modeling to arbitrary sequences of interleaved text and visual data. The framework consists of two core components:

    1. Multimodal Dense Retriever (RR): Given a multimodal input query xx (e.g., text prompt, image, or text-image pair) and an external memory M\mathcal{M} of multimodal documents, the retriever scores candidates using a bi-encoder and retrieves the top KK relevant multimodal documents M=(m1,…,mK)⊆MM = (m_1, \dots, m_K) \subseteq \mathcal{M}.
    2. Retrieval-Augmented Generator (GG): A Transformer decoder based on the CM3 architecture. Multimodal documents are represented as structured sequences (e.g., HTML strings formatted as "<img alt=[text] src=[image]>"), where text is tokenized into standard discrete text tokens and each image is encoded into 1024 discrete image tokens using a VQGAN tokenizer. The retrieved documents M=(m1,…,mK)M = (m_1, \dots, m_K) are prepended as in-context demonstrations to the input sequence xx, forming the concatenated input (m1,…,mK,x)(m_1, \dots, m_K, x) fed directly into the Transformer.

    The model generates outputs autoregressively, supporting text-to-image synthesis, image-to-text captioning, multimodal infilling, and contextual editing within a single unified model.

  2. Knowl 2 — CLIP-Based Mixed-Modal Dense Retriever

    model/method

    The dense multimodal retriever computes a relevance score r(q,m)r(q, m) between a multimodal query document qq and a candidate memory document m∈Mm \in \mathcal{M} using a dual-encoder inner-product scoring function:

    r(q,m)=EQ(q)⊤EM(m)r(q, m) = E_Q(q)^\top E_M(m)

    where EQE_Q is the query encoder and EME_M is the memory document encoder. Both encoders use frozen weights from a pretrained CLIP (ViT-L/14) model.

    To construct a unified representation for a multimodal document dd containing both a text component dtextd_{\text{text}} and an image component dimgd_{\text{img}}:

    1. The text sequence is encoded via the frozen CLIP text encoder into vector vtext=CLIPtext(dtext)v_{\text{text}} = \text{CLIP}_{\text{text}}(d_{\text{text}}).
    2. The image is encoded via the frozen CLIP image encoder into vector vimg=CLIPimg(dimg)v_{\text{img}} = \text{CLIP}_{\text{img}}(d_{\text{img}}).
    3. The document vector representation is the L2L_2-normalized average of the text and image embeddings:

    E(d)=vtext+vimg∥vtext+vimg∥2E(d) = \frac{v_{\text{text}} + v_{\text{img}}}{\|v_{\text{text}} + v_{\text{img}}\|_2}

    Candidate search across the memory index M\mathcal{M} is conducted via Maximum Inner Product Search (MIPS) using FAISS.

  3. Knowl 3 — Joint Generator Training Objective with Retrieved Document Loss Weighting

    equation

    The retrieval-augmented generator is trained using teacher forcing on the concatenated sequence of retrieved context documents (m1,…,mK)(m_1, \dots, m_K) and the target document xx. The training objective minimizes a weighted joint causal language modeling cross-entropy loss:

    L=Lmain+αLretr=−log⁡p(x∣m1,…,mK)−αlog⁡p(m1,…,mK)L = L_{\text{main}} + \alpha L_{\text{retr}} = -\log p(x \mid m_1, \dots, m_K) - \alpha \log p(m_1, \dots, m_K)

    where:

    • xx is the primary target multimodal document (tokenized into text and VQGAN image tokens).
    • (m1,…,mK)(m_1, \dots, m_K) is the sequence of KK retrieved multimodal documents prepended to xx in context.
    • Lmain=−log⁡p(x∣m1,…,mK)L_{\text{main}} = -\log p(x \mid m_1, \dots, m_K) is the autoregressive next-token prediction cross-entropy loss evaluated on the tokens of xx.
    • Lretr=−log⁡p(m1,…,mK)L_{\text{retr}} = -\log p(m_1, \dots, m_K) is the autoregressive token prediction cross-entropy loss evaluated across the tokens of the retrieved documents.
    • α≥0\alpha \ge 0 is a scalar hyperparameter that weights the retrieved context loss. Setting α=0.1\alpha = 0.1 optimizes training efficiency and generator perplexity by leveraging token logits already computed over the retrieved image tokens (1024 tokens per image) without discarding computation.
  4. Knowl 4 — Diverse Multimodal Document Retrieval Strategy

    algorithm

    The retrieval procedure fetches diverse multimodal documents by applying similarity thresholding to avoid redundancy and query token dropout for training regularization:

    Input: Query document qq, external memory index M\mathcal{M}, number of documents to retrieve KK, similarity threshold τ=0.9\tau = 0.9, token dropout rate pdrop=0.2p_{\text{drop}} = 0.2, training mode boolean is_training\text{is\_training}
    Output: Retrieved multimodal document list M=(m1,…,mK)M = (m_1, \dots, m_K)
    if is_training\text{is\_training}:
        qquery←randomly mask/drop pdrop fraction of tokens from qq_{\text{query}} \leftarrow \text{randomly mask/drop } p_{\text{drop}} \text{ fraction of tokens from } q
    else:
        qquery←qq_{\text{query}} \leftarrow q
    S←retrieve candidate documents from M ranked by descending r(qquery,m)S \leftarrow \text{retrieve candidate documents from } \mathcal{M} \text{ ranked by descending } r(q_{\text{query}}, m)
    M←[]M \leftarrow []
    for candidate document m∈Sm \in S:
        if length(MM) == KK:
            break
        if r(qquery,m)>τr(q_{\text{query}}, m) > \tau:
            continue
        if any(r(m,m′)>τ for m′∈Mr(m, m') > \tau \text{ for } m' \in M):
            continue
        append mm to MM
    return MM

    During training, qq is set to either only the text portion or only the image portion of xx to mirror inference queries and avoid trivial retrieval matching.

  5. Knowl 5 — MS-COCO Zero-Shot Text-to-Image and Image-to-Text Benchmark Performance

    data/table

    Zero-shot performance comparison on the MS-COCO validation set for text-to-image generation (measured by Fréchet Inception Distance, FID, where lower is better) and image-to-text generation (measured by CIDEr, where higher is better). RA-CM3 (2.7B parameters) was pretrained on 150M LAION text-image pairs and evaluated without task-specific fine-tuning.

    Approach Model Type MS-COCO FID (↓\downarrow) CIDEr (↑\uparrow)
    Retrieval Baseline - 17.97 84.1
    KNN-Diffusion Diffusion 16.66 -
    Stable Diffusion Diffusion 12.63 -
    GLIDE Diffusion 12.24 -
    DALL-E 2 Diffusion 10.39 -
    Imagen Diffusion 7.27 -
    Re-Imagen Diffusion 6.88 -
    DALL-E (12B) Autoregressive ∼\sim28 20.2 (Small)
    CogView (4B) Autoregressive 27.1 -
    CogView2 (6B) Autoregressive 24.0 -
    Parti (20B) Autoregressive 7.23 83.9
    Flamingo (3B; 4-shot) Autoregressive - 85.0
    Flamingo (80B; 4-shot) Autoregressive - 103.0
    Vanilla CM3 (2.7B) Autoregressive 29.5 71.9
    RA-CM3 (2.7B, Ours) Autoregressive 15.7 89.1

    RA-CM3 improves by 13.8 FID points over the compute-matched vanilla CM3 baseline in text-to-image generation and achieves 89.1 CIDEr in caption generation (outperforming Parti 20B and Flamingo-3B 4-shot), while utilizing under 30% of the training compute required by DALL-E.

  6. Knowl 6 — Ablation of Retrieval Relevance, Modality, Diversity, and Loss Weighting

    data/table

    Ablation study evaluating key architectural and training design choices for a 2.7B parameter RA-CM3 model trained for one day on 256 GPUs. Generative performance is evaluated via image perplexity and text perplexity on the MS-COCO validation set (lower is better across both metrics).

    Method Design Choice Image PPL (↓\downarrow) Text PPL (↓\downarrow)
    Retrieval Relevance Random at train infer 246 23
    Retrieve at train, random at infer 246 24
    Random at train, retrieve at infer 243 18
    Retrieve at train infer (final) 227 13
    Retrieval Modality Only image or only text 234 15
    Multimodal document (final) 227 13
    Retrieval Diversity Simply take top KK 244 17
    Avoid redundancy 235 15
    Avoid redundancy + Query dropout (final) 227 13
    Generator Training α=0\alpha = 0 239 17
    α=1\alpha = 1 240 17
    α=0.3\alpha = 0.3 231 14
    α=0.1\alpha = 0.1 (final) 227 13

    The ablation confirms that:

    1. Training with retrieved relevant documents is essential; training with random context and retrieving only at test time degrades text perplexity from 13 to 18 and image perplexity from 227 to 243.
    2. Multimodal retrieval (retaining both text and images) outperforms retrieving unimodal documents.
    3. Redundancy filtering and query dropout jointly reduce image perplexity from 244 to 227.
    4. Joint loss optimization with α=0.1\alpha = 0.1 outperforms unweighted optimization (α=0\alpha = 0) and overweighted optimization (α=1\alpha = 1). image tokens naturally have higher perplexity numbers than text tokens.
  7. Knowl 7 — Multimodal In-Context Demonstration for Controllable Generation and Image Editing

    model/method

    Training the RA-CM3 generator on retrieved multimodal sequences induces in-context multimodal reasoning, enabling user control over image generation and editing via demonstration examples without fine-tuning:

    1. Controllable Style and Visual Synthesis: By manually inserting demonstration image-caption pairs into the generator prompt before the target query caption (e.g., prepending an image of a triangular wooden house and an image of autumn foliage before the prompt "Photo of a house taken on an autumn day"), RA-CM3 synthesizes output images adhering to the visual style and entity characteristics of the demonstrations.
    2. Controllable Image Infilling and Editing: Image infilling is framed as sequence completion with the format "[unmasked patch sequence] <mask > [unmasked patch sequence]", where the model predicts the token sequence inside <mask >. By inserting demonstration images into the context (e.g., placing an image of a person in a red jacket into context), the model alters target features (e.g., editing a black jacket to red) or accurately infills masked body regions and equipment (e.g., synthesizing both legs and skis rather than legs alone).
  8. Knowl 8 — Few-Shot Image Classification with Non-Semantic Labels via Context Ensembling

    data/table

    Few-shot in-context classification evaluated on a binary image classification task constructed from ImageNet using abstract, non-semantic labels ("animal X" vs "animal Y") to eliminate prior semantic memorization. For 1-shot evaluation, the prompt consists of ([image X],"animal X",[image Y],"animal Y")([\text{image } X], \text{"animal X"}, [\text{image } Y], \text{"animal Y"}) followed by ([test image],"animal ")([\text{test image}], \text{"animal "}), predicting the next token probability of "X" vs "Y". For kk-shot evaluation (k∈{1,2,4,8}k \in \{1, 2, 4, 8\}), kk demonstration pairs are evaluated across kk separate passes, and predicted token probabilities are averaged via an ensemble.

    Model kk-shot Accuracy
    k=1k=1 k=2k=2 k=4k=4 k=8k=8
    Baseline CM3 0.53 0.50 0.56 0.56
    RA-CM3 (Ours) 0.78 0.79 0.86 0.90

    RA-CM3 achieves 0.78 one-shot accuracy on non-semantic labels compared to 0.53 for vanilla CM3, and scales monotonically to 0.90 accuracy at k=8k=8 via demonstration ensembling.

  9. Knowl 9 — Perplexity Scaling Laws for Retrieval-Augmented Multimodal Language Models

    empirical result

    Across model parameter scales of 125M, 350M, 1.3B, and 2.7B parameters trained with identical compute budgets (two days on 256 A100 GPUs), RA-CM3 achieves consistent validation perplexity improvements over compute-matched vanilla CM3 on the MS-COCO dataset:

    • Validation Image Perplexity: RA-CM3 decreases from ∼240\sim 240 (125M) to ∼220\sim 220 (2.7B), maintaining a lower perplexity than vanilla CM3 (which scales from ∼250\sim 250 at 125M to ∼230\sim 230 at 2.7B).
    • Validation Text Perplexity: RA-CM3 decreases from ∼17\sim 17 (125M) to ∼11\sim 11 (2.7B), outperforming vanilla CM3 (which scales from ∼21\sim 21 at 125M to ∼15\sim 15 at 2.7B).

    Across both modalities, a 1.3B-parameter RA-CM3 achieves lower validation perplexity than a 2.7B-parameter vanilla CM3, demonstrating that retrieval augmentation reduces the parameter requirements needed to reach a given generative performance level without displaying diminishing returns up to 2.7B parameters.

  10. Knowl 10 — Effect of the Number of Retrieved Multimodal Documents on Generation Perplexity

    data/table

    Evaluation of caption-to-image generation perplexity on the MS-COCO validation set as a function of the number of retrieved multimodal context documents KK prepended to the generator:

    Model Image Perplexity (↓\downarrow)
    K=1K=1 K=2K=2 K=4K=4 K=8K=8
    RA-CM3 228 227 228 232

    Image perplexity reaches an optimal value of 227 at K=2K = 2. Setting K=2K = 2 allows the total sequence (comprising up to three multimodal documents of ∼1024\sim 1024 image tokens plus text tokens each) to fit within a 4096-token Transformer context window without exceeding GPU memory constraints.

Coverage note — Intrinsic recall evaluation of the frozen CLIP bi-encoder retriever on MS-COCO (Appendix C.1) was omitted as it represents a standard validation of off-the-shelf CLIP representations rather than a novel model contribution.

References

  1. 1.Agarwal, O., Ge, H., Shakeri, S., and Al-Rfou, R. Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training. In North American Chapter of the Association for Computational Linguistics (NAACL), 2021.
  2. 2.Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., and Zettlemoyer, L. CM3: A causal masked multimodal model of the internet. arXiv preprint arXiv:2201.07520, 2022.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  4. 4.Ashual, O., Sheynin, S., Polyak, A., Singer, U., Gafni, O., Nachmani, E., and Taigman, Y. Knn-diffusion: Image generation via large-scale retrieval. arXiv preprint arXiv:2204.02849, 2022.
  5. 5.Birhane, A., Prabhu, V. U., and Kahembwe, E. Multimodal datasets: misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963, 2021a.
  6. 6.Birhane, A., Prabhu, V. U., and Kahembwe, E. Ethical considerations of generative AI. AI for Content Creation Workshop, CVPR, 2021b.
  7. 7.Blattmann, A., Rombach, R., Oktay, K., and Ommer, B. Retrieval-augmented diffusion models. arXiv preprint arXiv:2204.11824, 2022.
  8. 8.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al. Improving language models by retrieving from trillions of tokens. In International Conference on Machine Learning (ICML), 2022.
  9. 9.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  10. 10.Chen, W., Hu, H., Chen, X., Verga, P., and Cohen, W. W. Murag: Multimodal retrieval-augmented generator for open question answering over images and text. In Empirical Methods in Natural Language Processing (EMNLP), 2022a.
  11. 11.Chen, W., Hu, H., Saharia, C., and Cohen, W. W. Re-imagen: Retrieval-augmented text-to-image generator. arXiv preprint arXiv:2209.14491, 2022b.
  12. 12.Cho, J., Lu, J., Schwenk, D., Hajishirzi, H., and Kembhavi, A. X-lxmert: Paint, caption and answer questions with multi-modal transformers. arXiv preprint arXiv:2009.11278, 2020.
  13. 13.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2009.
  14. 14.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al. Cogview: Mastering text-to-image generation via transformers. In Neural Information Processing Systems (NeurIPS), 2021.
  15. 15.Ding, M., Zheng, W., Hong, W., and Tang, J. Cogview2: Faster and better text-to-image generation via hierarchical transformers. arXiv preprint arXiv:2204.14217, 2022.
  16. 16.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  17. 17.Forever, A. rudall-e. https://github.com/ai-forever/ru-dalle, 2021.
  18. 18.Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, W.-t., Zettlemoyer, L., and Lewis, M. Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999, 2022.
  19. 19.Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y. Make-a-scene: Scene-based text-to-image generation with human priors. arXiv preprint arXiv:2203.13131, 2022.
  20. 20.Gur, S., Neverova, N., Stauffer, C., Lim, S.-N., Kiela, D., and Reiter, A. Cross-modal retrieval augmentation for multi-modal classification. arXiv preprint arXiv:2104.08108, 2021.
  21. 21.Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W. Realm: Retrieval-augmented language model pre-training. In International Conference on Machine Learning (ICML), 2020.
  22. 22.Hao, Y., Song, H., Dong, L., Huang, S., Chi, Z., Wang, W., Ma, S., and Wei, F. Language models are general-purpose interfaces. arXiv preprint arXiv:2206.06336, 2022.
  23. 23.Hashimoto, T. B., Guu, K., Oren, Y., and Liang, P. S. A retrieve-and-edit framework for predicting structured outputs. In Neural Information Processing Systems (NeurIPS), 2018.
  24. 24.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Neural Information Processing Systems (NeurIPS), 2017.
  25. 25.Johnson, J., Douze, M., and Jegou, H. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 2019.
  26. 26.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. Dense passage retrieval for open-domain question answering. In Empirical Methods in Natural Language Processing (EMNLP), 2020.
  27. 27.Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858, 2019.
  28. 28.Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. Generalization through memorization: Nearest neighbor language models. In International Conference on Learning Representations (ICLR), 2019.
  29. 29.Kim, S., Cho, S., Kim, C., Lee, D., and Baek, W. mindall-e on conceptual captions. https://github.com/kakaobrain/minDALL-E, 2021.
  30. 30.Kingma, D. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  31. 31.Lewis, M., Ghazvininejad, M., Ghosh, G., Aghajanyan, A., Wang, S., and Zettlemoyer, L. Pre-training via paraphrasing. In Neural Information Processing Systems (NeurIPS), 2020a.
  32. 32.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kuttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems (NeurIPS), 2020b.
  33. 33.Li, B., Qi, X., Lukasiewicz, T., and Torr, P. Controllable text-to-image generation. Neural Information Processing Systems (NeurIPS), 2019.
  34. 34.Li, Y., Pan, Y., Yao, T., and Mei, T. Comprehending and ordering semantics for image captioning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  35. 35.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European conference on computer vision, 2014.
  36. 36.Metzler, D., Tay, Y., Bahri, D., and Najork, M. Rethinking search: making domain experts out of dilettantes. In ACM SIGIR Forum, 2021.
  37. 37.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  38. 38.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Neural Information Processing Systems (NeurIPS), 2019.
  39. 39.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  40. 40.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 2021.
  41. 41.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  42. 42.Ramos, R., Martins, B., Elliott, D., and Kementchedjhieva, Y. Smallcap: Lightweight image captioning prompted with retrieval augmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  43. 43.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  44. 44.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  45. 45.Sarto, S., Cornia, M., Baraldi, L., and Cucchiara, R. Retrieval-augmented transformer for image captioning. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing, 2022.
  46. 46.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  47. 47.Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652, 2023.
  48. 48.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems (NeurIPS), 2017.
  49. 49.Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  50. 50.Wang, P. Dall-e in pytorch. https://github.com/lucidrains/DALLE-pytorch, 2021.
  51. 51.Wang, Y., Yasunaga, M., Ren, H., Wada, S., and Leskovec, J. Vqa-gnn: Reasoning with multimodal semantic graph for visual question answering. arXiv preprint arXiv:2205.11501, 2022a.
  52. 52.Wang, Z., Yu, J., Yu, A. W., Dai, Z., Tsvetkov, Y., and Cao, Y. Simvlm: Simple visual language model pretraining with weak supervision. In International Conference on Learning Representations (ICLR), 2022b.
  53. 53.Xie, T., Wu, C. H., Shi, P., Zhong, R., Scholak, T., Yasunaga, M., Wu, C.-S., Zhong, M., Yin, P., Wang, S. I., Zhong, V., Wang, B., Li, C., Boyle, C., Ni, A., Yao, Z., Radev, D., Xiong, C., Kong, L., Zhang, R., Smith, N. A., Zettlemoyer, L., and Yu, T. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. In Empirical Methods in Natural Language Processing (EMNLP), 2022.
  54. 54.Yasunaga, M., Ren, H., Bosselut, A., Liang, P., and Leskovec, J. QA-GNN: Reasoning with language models and knowledge graphs for question answering. In North American Chapter of the Association for Computational Linguistics (NAACL), 2021.
  55. 55.Yasunaga, M., Bosselut, A., Ren, H., Zhang, X., Manning, C. D., Liang, P., and Leskovec, J. Deep bidirectional language-knowledge graph pretraining. In Neural Information Processing Systems (NeurIPS), 2022a.
  56. 56.Yasunaga, M., Leskovec, J., and Liang, P. LinkBERT: Pre-training language models with document links. In Association for Computational Linguistics (ACL), 2022b.
  57. 57.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022.
  58. 58.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  59. 59.Zhang, Z., Han, X., Liu, Z., Jiang, X., Sun, M., and Liu, Q. Ernie: Enhanced language representation with informative entities. In Association for Computational Linguistics (ACL), 2019.

Citation

MLA
Yasunaga, M., et al. “Retrieval-Augmented Multimodal Language Modeling”. International Conference on Machine Learning, vol. 202, 2023, pp. 39755–69, https://proceedings.mlr.press/v202/yasunaga23a.html.
APA
Yasunaga, M., Aghajanyan, A., Shi, W., James, R., Leskovec, J., Liang, P., Lewis, M., Zettlemoyer, L., & Yih, W.-T. (2023). Retrieval-Augmented Multimodal Language Modeling. International Conference on Machine Learning, 202, 39755–39769. https://proceedings.mlr.press/v202/yasunaga23a.html
Chicago
Yasunaga, M., A. Aghajanyan, W. Shi, et al. 2023. “Retrieval-Augmented Multimodal Language Modeling”. International Conference on Machine Learning 202: 39755–69. https://proceedings.mlr.press/v202/yasunaga23a.html.
Harvard
Yasunaga, M. et al. (2023) “Retrieval-Augmented Multimodal Language Modeling”, International Conference on Machine Learning. PMLR, pp. 39755–39769. Available at: https://proceedings.mlr.press/v202/yasunaga23a.html.
Vancouver
1. Yasunaga M, Aghajanyan A, Shi W, James R, Leskovec J, Liang P, Lewis M, Zettlemoyer L, Yih W-T (2023) Retrieval-Augmented Multimodal Language Modeling. In: International Conference on Machine Learning. PMLR, pp 39755–39769

BibTeX

@InProceedings{pmlr-v202-yasunaga23a,
  title = 	 {Retrieval-Augmented Multimodal Language Modeling},
  author =       {Yasunaga, Michihiro and Aghajanyan, Armen and Shi, Weijia and James, Richard and Leskovec, Jure and Liang, Percy and Lewis, Mike and Zettlemoyer, Luke and Yih, Wen-Tau},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {39755--39769},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/yasunaga23a/yasunaga23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/yasunaga23a.html},
  abstract = 	 {Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all their knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). Specifically, for the retriever, we use a pretrained CLIP, and for the generator, we train a CM3 Transformer on the LAION dataset. Our resulting model, named Retrieval-Augmented CM3 (RA-CM3), is the first multimodal model that can retrieve and generate both text and images. We show that RA-CM3 significantly outperforms baseline multimodal models such as DALL-E and CM3 on both image and caption generation tasks (12 FID and 17 CIDEr improvements on MS-COCO), while requiring much less compute for training ($<$30% of DALL-E). Moreover, we show that RA-CM3 exhibits novel capabilities such as faithful image generation and multimodal in-context learning (e.g., image generation from demonstrations).}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/