Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study
Boxin WangWei PingPeng XuLawrence McAfeeZihan LiuMohammad ShoeybiYi DongOleksii KuchaievBo LiChaowei Xiao
Demonstrates through scalable reproduction up to 9.5B parameters that pretraining autoregressive language models with retrieval substantially improves factual accuracy and downstream knowledge-intensive task performance over standard GPT baselines while introducing RETRO++ to boost open-domain question answering.
Standard autoregressive language models require immense parameter counts to store factual knowledge, making them expensive to deploy, difficult to update with new information, and prone to factual hallucinations. Although retrieval mechanisms have been added to models during fine-tuning or inference, whether large generative language models should be pretrained with retrieval capabilities from scratch has remained an open question. The article evaluates this question by systematically comparing standard generative models with a retrieval-augmented architecture across text generation quality, factual accuracy, downstream task performance, and question answering.
The researchers reproduced and trained a scalable retrieval-augmented architecture, known as RETRO, across configurations ranging from 148 million to 9.5 billion parameters on a 330-billion-token pretraining dataset. The model incorporates a retrieval database containing 5.3 billion text chunks indexed using dense similarity search, enabling it to retrieve relevant external text during both pretraining and generation. Crucially, the experiments evaluated these models against standard architectures under identical data schedules and model sizes to isolate the precise impact of pretraining with retrieval.
The findings show that retrieval-augmented pretraining consistently outperforms standard generative models across several critical metrics. First, in open-ended text generation, the retrieval architecture reduced repetitive text generation by an average of 21% while maintaining equal fluency and coherence. Second, the model demonstrated moderately higher factual accuracy and lower hallucination rates across established factuality benchmarks. Third, in zero-shot evaluations across nine standard tasks, the architecture significantly outperformed standard models on knowledge-intensive benchmarks while remaining competitive on reasoning tasks. Fourth, an improved variant introduced in the study, RETRO++, substantially elevated open-domain question answering performance, boosting the exact match accuracy on the Natural Questions benchmark to 54.1% compared to 40.9% for the baseline retrieval model.
These results demonstrate that pretraining models with retrieval is a highly effective design pattern for foundation models. Incorporating retrieval allows models to access external knowledge dynamically rather than relying solely on memorized parameters, facilitating easier factual updates and reducing the risk of generating incorrect or outdated information. Because the approach requires only modest computational overhead—adding less than 25% to pretraining GPU hours—it presents a favorable cost-to-performance trade-off for organizations deploying large language models.
Engineering and research teams should consider pretraining autoregressive models with retrieval when developing systems for knowledge-intensive domains and factual question answering. When implementing these architectures, practitioners should adopt enhanced evidence-routing designs such as RETRO++ and implement flexible retrieval intervals to balance latency against output accuracy. Furthermore, organizations must carefully curate the external retrieval database, as using toxic or low-quality source texts directly degrades generation safety and accuracy.
The study notes certain limitations, particularly that model evaluations were capped at 9.5 billion parameters rather than the tens or hundreds of billions used in the largest commercial systems. Additionally, system outputs remain strongly dependent on the quality, neutrality, and freshness of the underlying datastore. Nevertheless, the consistent performance gains across all evaluated model sizes provide strong confidence that pretraining with retrieval is a robust, scalable architecture for future generative systems.
- Paper: Improving language models by retrieving from trillions of tokens, Sebastian Borgeaud et al. (2022). Read the original RETRO study first to understand the retrieval-augmented architecture and experimental baseline this comprehensive evaluation reproduces and extends.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). REALM establishes the earlier precedent for integrating retrieval into language-model pretraining, clarifying the design lineage behind retrieval-augmented pretraining.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). The foundational RAG paper introduces retrieval-grounded generation and its factuality motivation, useful context for the source’s comparison of retrieval approaches.
- Paper: InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining, Boxin Wang et al. (2024). Building on retrieval-augmented pretraining, InstructRetro scales the approach and adds instruction tuning to improve task following and zero-shot performance.
