InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
Boxin WangWei PingLawrence McAfeePeng XuBo LiMohammad ShoeybiBryan Catanzaro
Demonstrates that scaling retrieval-augmented pretraining to a 48-billion-parameter model yields substantial zero-shot performance gains after instruction tuning, while revealing that the retrieval encoder can be removed post-training without degrading downstream task accuracy.
Large language models often struggle with factual accuracy, high computational training costs, and hallucination when relying solely on internal parameters. Integrating external database retrieval during the initial training phase can mitigate these issues, but prior retrieval-augmented models have remained relatively small (around 7 to 11 billion parameters), limiting their ability to follow complex human instructions and perform well on new tasks without task-specific training.
The article demonstrates the scaling of retrieval-augmented language models up to 48 billion parameters—creating Retro 48B and its instruction-tuned variant, InstructRetro—and evaluates whether retrieval during pretraining produces a foundation model that substantially outperforms traditional generative models on zero-shot tasks.
To achieve this, the researchers applied a continued pretraining technique on an existing 43-billion-parameter foundation model using an additional 100 billion tokens while dynamically retrieving information from an external index of 1.2 trillion tokens (19 billion chunks). Crucially, unlike prior methods that freeze model weights, all parameters were unfrozen during training. The resulting foundation model was then fine-tuned using a high-quality blend of 128,000 conversational instruction examples across various domains.
The evaluation revealed several key findings:
- Efficiency: Continued pretraining with retrieval added only 2.58% in overall computing hours compared to standard training, while achieving language modeling accuracy comparable to standard models four times larger.
- Downstream Performance: After instruction tuning, InstructRetro outperformed its standard instruction-tuned counterpart across all benchmarks, showing an average relative gain of 7% on eight short-form question-answering tasks, 10% on four long-form question-answering tasks, and 16% on three document summarization benchmarks.
- Architectural Simplification: Researchers discovered that the dedicated retrieval encoder can be completely deactivated during downstream use. Relying solely on the main decoder backbone (InstructRetro 43B) yielded comparable or even slightly superior results to the 48-billion-parameter configuration with active cross-attention.
These findings indicate that retrieval-augmented pretraining fundamentally conditions the core language model to better utilize contextual evidence presented in prompts, even when operating as a standard decoder-only system. For organizations deploying generative artificial intelligence, this approach delivers the performance of much larger models at substantially lower computational, hardware, and operational costs.
Decision-makers should consider adopting retrieval-augmented continued pretraining pipelines when training proprietary foundation models to maximize task accuracy per compute dollar. Future efforts should focus on exploring retrieval-augmented instruction tuning datasets to assess whether keeping retrieval encoders active during fine-tuning can unlock additional accuracy gains.
Confidence in these findings is high across standard English benchmarks and enterprise documentation tasks. However, stakeholders should note that the base pretraining dataset excluded specialized coding and roleplay data, where performance showed little to no advantage over baseline models. Deployment in domains requiring proprietary or privacy-sensitive data still requires standard governance to address residual data leakage and bias risks.
- Paper: Improving language models by retrieving from trillions of tokens, Sebastian Borgeaud et al. (2022). Introduces the foundational RETRO architecture and chunked cross-attention pretraining mechanism that InstructRetro directly scales up and instruction-tunes.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Establishes the core retrieval-augmented generation framework combining neural retrieval with generative decoders to mitigate hallucination on knowledge-intensive tasks.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). Pioneers the paradigm of retrieval-augmented language model pretraining by incorporating an external knowledge retriever directly into the pretraining objective.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Provides the foundational instruction-tuning methodology that InstructRetro applies post-pretraining to unlock zero-shot generalization across diverse benchmarks.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). Demonstrates that conditioning language models on retrieved context enhances effective parameter scale, motivating InstructRetro's exploration of decoder-only contextual utilization.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). Extends instruction tuning for retrieval-augmented models by training a single language model backbone to unify context reranking directly with answer generation.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Advances beyond static retrieval architectures by instruction-tuning models to selectively retrieve, self-evaluate, and critique evidence on demand during generation.
- Paper: DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models, Weihang Su et al. (2024). Builds on instruction-tuned retrieval systems by dynamically evaluating internal model uncertainty and attention distributions to determine exact retrieval timing and query formulation.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). Investigates and strengthens the resilience of retrieval-augmented language models against irrelevant or counterfactual retrieved context through adaptive adversarial training.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Progresses from supervised instruction-tuned retrieval backbones toward autonomous, multi-turn search and self-verification learned via reinforcement learning.
