Retrieval-Augmented Multimodal Language Modeling
Michihiro YasunagaArmen AghajanyanWeijia ShiRichard JamesJure LeskovecPercy LiangMike LewisLuke ZettlemoyerWen-Tau Yih
Proposes RA-CM3, a retrieval-augmented multimodal architecture that fetches relevant text-image documents to generate both modalities, significantly cutting training compute while outperforming models like DALL-E and enabling multimodal in-context learning.
Modern multimodal artificial intelligence systems, such as models that generate images from text or descriptions from images, traditionally store all their world knowledge directly within their neural network parameters. As a result, expanding their knowledge base or updating information requires exponentially larger models and massive training datasets, which dramatically increases computing costs. The article evaluates a modular approach called retrieval augmentation, demonstrating that connecting a base multimodal generative model to an external document memory allows it to reference external text and images dynamically, rather than relying solely on memorized parameters.
To demonstrate this capability, the authors developed Retrieval-Augmented CM3 (RA-CM3), the first model capable of both retrieving and generating mixed combinations of text and images. The system pairs an off-the-shelf multimodal retriever based on CLIP with a 2.7-billion-parameter CM3 Transformer generator. The model was trained from scratch on 150 million web-collected text-image pairs from the LAION dataset using 256 GPUs over five days, using the exact same dataset as both training data and external memory to ensure a controlled evaluation against non-retrieval baselines.
The investigation produced several key findings regarding model efficiency and performance. First, RA-CM3 significantly outperformed baseline models without retrieval on standard benchmarks, improving image generation quality by approximately 13 to 14 FID points (reducing the score from 29.5 to 15.7, where lower is better) and caption quality by over 17 CIDEr points (increasing from 71.9 to 89.1). Second, RA-CM3 achieved these superior results while requiring less than 30% of the training compute and parameters utilized by comparable autoregressive systems like DALL-E (12B). Third, the architecture proved substantially more capable at generating rare or knowledge-intensive concepts, such as correctly rendering historical artifacts and composite scenes where standard generative models hallucinate or substitute common defaults. Finally, the training process enabled zero-shot and few-shot in-context learning, allowing users to control image styling directly via image prompts and improving few-shot classification accuracy from 0.78 at one-shot to 0.90 at eight-shot.
These findings indicate that retrieval augmentation offers a path to lower training expenses, faster development cycles, and higher factual reliability. Rather than expanding model scale to memorize rare entities, organizations can utilize external memory stores that are easily updated without retraining the entire neural network. Furthermore, retrieving source documents inherently provides provenance, improving interpretability and reducing unintended hallucinations in generative outputs.
Moving forward, technical teams should consider retrieval-augmented architectures when designing multimodal systems that require frequent knowledge updates, rare entity processing, or budget-constrained compute budgets. Recommended subsequent steps include investigating retriever fine-tuning rather than relying on frozen retrievers, expanding the framework to additional modalities beyond text and image pairs, and utilizing ensembling techniques to incorporate larger numbers of retrieved reference examples.
Decision-makers should note certain limitations: the model is a research prototype trained on filtered web data, meaning risks of biased or unsafe outputs persist. In addition, context length constraints in current sequence architectures practically limit direct retrieval to one or two documents per pass before needing ensemble mechanisms. Nevertheless, the reported improvements in training efficiency and generation fidelity are supported by consistent scaling behavior across model sizes, providing high confidence in the fundamental advantages of retrieval-augmented multimodal modeling.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). Read this landmark autoregressive text-to-image work first to understand the generation paradigm RA-CM3 extends with external multimodal retrieval.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Its foundational RAG design clarifies how a generator can condition on retrieved external evidence rather than relying only on parametric memory.
- Paper: Hierarchical Text-Conditional Image Generation with CLIP Latents, Aditya Ramesh et al. (2022). Its CLIP-latent generation pipeline provides useful context for how aligned image-text representations can support image generation.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). PreFLMR advances multimodal retrieval toward knowledge-intensive visual question answering, extending the retrieval component into a broader set of retrieval tasks.
- Paper: Kosmos-G: Generating Images in Context with Multimodal Large Language Models, Xichen Pan et al. (2024). KOSMOS-G carries multimodal in-context generation forward by using interleaved text and image references to create images grounded in visual demonstrations.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o continues the effort to unify multimodal understanding and generation in a single model, broadening the generation-centered design explored here.
