InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining

Boxin WangWei PingLawrence McAfeePeng XuBo LiMohammad ShoeybiBryan Catanzaro

article2024ICML79 citations

Demonstrates that scaling retrieval-augmented pretraining to a 48-billion-parameter model yields substantial zero-shot performance gains after instruction tuning, while revealing that the retrieval encoder can be removed post-training without degrading downstream task accuracy.

Listen

Large language models often struggle with factual accuracy, high computational training costs, and hallucination when relying solely on internal parameters. Integrating external database retrieval during the initial training phase can mitigate these issues, but prior retrieval-augmented models have remained relatively small (around 7 to 11 billion parameters), limiting their ability to follow complex human instructions and perform well on new tasks without task-specific training.

The article demonstrates the scaling of retrieval-augmented language models up to 48 billion parameters—creating Retro 48B and its instruction-tuned variant, InstructRetro—and evaluates whether retrieval during pretraining produces a foundation model that substantially outperforms traditional generative models on zero-shot tasks.

To achieve this, the researchers applied a continued pretraining technique on an existing 43-billion-parameter foundation model using an additional 100 billion tokens while dynamically retrieving information from an external index of 1.2 trillion tokens (19 billion chunks). Crucially, unlike prior methods that freeze model weights, all parameters were unfrozen during training. The resulting foundation model was then fine-tuned using a high-quality blend of 128,000 conversational instruction examples across various domains.

The evaluation revealed several key findings:

  1. Efficiency: Continued pretraining with retrieval added only 2.58% in overall computing hours compared to standard training, while achieving language modeling accuracy comparable to standard models four times larger.
  2. Downstream Performance: After instruction tuning, InstructRetro outperformed its standard instruction-tuned counterpart across all benchmarks, showing an average relative gain of 7% on eight short-form question-answering tasks, 10% on four long-form question-answering tasks, and 16% on three document summarization benchmarks.
  3. Architectural Simplification: Researchers discovered that the dedicated retrieval encoder can be completely deactivated during downstream use. Relying solely on the main decoder backbone (InstructRetro 43B) yielded comparable or even slightly superior results to the 48-billion-parameter configuration with active cross-attention.

These findings indicate that retrieval-augmented pretraining fundamentally conditions the core language model to better utilize contextual evidence presented in prompts, even when operating as a standard decoder-only system. For organizations deploying generative artificial intelligence, this approach delivers the performance of much larger models at substantially lower computational, hardware, and operational costs.

Decision-makers should consider adopting retrieval-augmented continued pretraining pipelines when training proprietary foundation models to maximize task accuracy per compute dollar. Future efforts should focus on exploring retrieval-augmented instruction tuning datasets to assess whether keeping retrieval encoders active during fine-tuning can unlock additional accuracy gains.

Confidence in these findings is high across standard English benchmarks and enterprise documentation tasks. However, stakeholders should note that the base pretraining dataset excluded specialized coding and roleplay data, where performance showed little to no advantage over baseline models. Deployment in domains requiring proprietary or privacy-sensitive data still requires standard governance to address residual data leakage and bias risks.

Cover for InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining

Abstract

Pretraining auto-regressive large language models (LLMs) with retrieval demonstrates better perplexity and factual accuracy by leveraging external databases. However, the size of existing pretrained retrieval-augmented LLM is still limited (e.g., Retro has 7.5B parameters), which limits the effectiveness of instruction tuning and zero-shot generalization. In this work, we introduce Retro 48B, the largest LLM pretrained with retrieval. Specifically, we continue to pretrain a 43B GPT model on additional 100 billion tokens using the Retro augmentation method by retrieving from 1.2 trillion tokens. Notably, the obtained foundation model, Retro 48B, largely outperforms the counterpart GPT 43B trained on 1.2T tokens in terms of perplexity with only 2.58% additional GPU hours, demonstrating the significant scaling potential of the method. After instruction tuning on Retro, InstructRetro demonstrates significant improvement over the instruction tuned GPT on a wide range of zero-shot tasks. Specifically, the average improvement of InstructRetro is 7% over its GPT counterpart across 8 short-form QA and reading comprehension tasks, 10% over GPT across 4 challenging long-form QA tasks, and 16% over GPT across 3 summarization tasks. Surprisingly, we find that one can ablate the encoder from InstructRetro architecture and directly use its decoder backbone, while achieving comparable results. Our results highlight the promising direction to obtain a better GPT decoder through continued pretraining with retrieval before instruction tuning. Our code and checkpoints are publicly available at: https://huggingface.co/nvidia/retro-48b-instruct-4k.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Continued Pretraining of GPT with Retrieval
  • 3.1. Preliminaries of Retro
  • 3.2. Retro-fitting: continued pretraining with retrieval
  • 4. Instruction Tuning
  • 4.1. Datasets Blending
  • 4.2. Training details
  • 5. Experiments
  • 5.1. Experimental setup
  • 5.2. Zero-shot evaluation on QA tasks
  • 5.3. Zero-shot evaluation on summarization tasks
  • 5.4. Ablation stuides
  • 5.5. Evaluation on MT-Bench
  • 6. Conclusion
  • Impact Statement
  • References
  • A. Details of Pretraining
  • A.1. Pretraining corpus
  • A.2. Continued pretraining schedules
  • A.3. Computational cost for continued pretraining
  • B. Details of retrieval database
  • B.1. Faiss index configuration
  • B.2. Computational cost on building retrieval database
  • B.3. Ablation studies on Faiss index confirations
  • C. Qualitative examples
  • C.1. An example From the instruction tuning data
  • C.2. An example From the downstream QA dataset: SQuAD 1.1
  • D. Experimental results on MT Bench
  • E. Potential Negative Social Impacts

Knowls

  1. Knowl 1 — InstructRetro Architecture and Gated Cross-Attention Mechanism

    model/method

    InstructRetro is a framework that combines retrieval-augmented pretraining with instruction tuning across a three-stage pipeline:

    1. Stage I (GPT Backbone Pretraining): A standard auto-regressive decoder-only GPT model (e.g., 43B parameters) is pretrained from scratch on 1.1 trillion tokens.

    2. Stage II (Continued Pretraining with Retrieval / Retro-fitting): The model is augmented with a shallow 2-layer bidirectional Retro encoder and chunk-wise cross-attention layers, adding approximately 10% additional parameters (yielding Retro 48B). The entire model (both decoder backbone and encoder/cross-attention modules) is continued in pretraining on an additional 100 billion tokens while retrieving nearest neighbor chunks from a 1.2-trillion-token retrieval database.

    3. Stage III (Instruction Tuning with Encoder Gating): The model is fine-tuned on a blend of 128,000 high-quality conversational instruction samples. Because instruction tuning datasets generally lack high-quality pretraining-style neighbor chunks, the Retro encoder is bypassed via a binary gate g∈{0,1}g \in \{0, 1\} situated between the cross-attention output and the self-attention residual connection:

    hout=FFN(hself+g⋅hcross)h_{\text{out}} = \text{FFN}(h_{\text{self}} + g \cdot h_{\text{cross}})

    During Stage II pretraining, the gate is active (g=1g=1). During Stage III instruction tuning and downstream inference, the gate is deactivated (g=0g=0). This effectively freezes the cross-attention and encoder parameters and updates only the weights of the GPT decoder backbone, turning the model into a pure decoder (InstructRetro 43B) that directly ingests in-context retrieved evidence in a Retrieval-Augmented Generation (RAG) format.

  2. Knowl 2 — Unfrozen-Decoder Retro-fitting Pretraining

    model/method

    In retrieval-augmented pretraining (Retro), an input token sequence X=(x1,…,xn)X = (x_1, \dots, x_n) of maximum length n=4096n=4096 is partitioned into ll contiguous chunks (C1,…,Cl)(C_1, \dots, C_l) each of size m=64m=64.

    To generate tokens in chunk CiC_i, the dense BERT embedding of the preceding chunk Ci−1C_{i-1} is used to retrieve its top-k=2k=2 nearest neighbor chunks N(Ci−1)\mathcal{N}(C_{i-1}) from the 1.2T-token database. Using Ci−1C_{i-1} rather than CiC_i ensures causality in autoregressive language modeling. The neighbor chunks are encoded by a 2-layer bidirectional Transformer (Retro encoder) sharing the hidden dimension of the decoder, and integrated into the decoder via chunk-wise cross-attention layers.

    Unlike standard Retro-fitting which freezes the decoder parameters when adding retrieval modules to a pretrained language model, this method unfreezes all decoder parameters during continued pretraining on 100B tokens (a 9% expansion of the initial 1.1T pretraining budget). Unfreezing the decoder leads to faster loss convergence, lower validation perplexity across all model sizes (823M, 2.25B, 8.5B, 22B, and 48B), and imbues the decoder backbone with an enhanced intrinsic ability to utilize retrieved context.

  3. Knowl 3 — Downstream Transferability of Retrieval-Pretrained Decoders Without Retrieval Encoders

    empirical result

    Ablating the Retro encoder from InstructRetro 48B at downstream inference time—by disabling the cross-attention gate (g=0g=0) and providing retrieved documents directly in the prompt context (standard Retrieval-Augmented Generation format)—yields a 43B parameter decoder-only model (InstructRetro 43B) that achieves downstream performance matching or exceeding the full 48B model with an active encoder.

    On zero-shot downstream benchmarks:

    • Across 8 short-form QA and reading comprehension tasks, InstructRetro 43B achieves a 7% average relative improvement over InstructGPTRAG_{\text{RAG}} 43B, while InstructRetro 48B achieves a 6% relative improvement.
    • Across 4 long-form QA tasks, InstructRetro 43B achieves a 10% average relative improvement over InstructGPTRAG_{\text{RAG}} 43B (matching InstructRetro 48B's 10% relative improvement).
    • Across 3 long-document summarization tasks, InstructRetro 43B achieves a 16% average relative improvement over InstructGPTRAG_{\text{RAG}} 43B.

    This confirms that continued pretraining with retrieval and an unfrozen decoder directly enhances the decoder backbone's reasoning and in-context evidence integration capabilities, rendering the dedicated retrieval encoder unnecessary during downstream task inference.

  4. Knowl 4 — Zero-Shot Evaluation on Short-Form QA and Reading Comprehension Benchmarks

    data/table

    Zero-shot performance across eight short-form question answering and reading comprehension benchmarks demonstrates that InstructRetro 43B outperforms its instruction-tuned baseline InstructGPTRAG_{\text{RAG}} 43B (by 7% on average) and surpasses larger models such as Llama 2 70B with RAG:

    Task NQ TriviaQA NewsQA SQuAD 2.0 SQuAD 1.1 Quoref NarrativeQA DROP
    Metric EM EM F1 F1 / EM F1 / EM F1 F1 F1
    GPT-3 175B 14.6 64.3 - 59.5 / 52.6 - - - 23.6
    PaLM 2-L 37.5 - - - - - - -
    GLaM 64B 24.7 71.3 - 71.1 / 64.7 - - - 57.3
    FLAN-LaMDA 137B 20.7 68.1 - 44.2 / - 80.1 / - - - 22.7
    Llama 2 RAG 70B 37.7 65.6 53.4 71.4 / 64.1 73.4 / 66.2 69.7 52.7 57.2
    Retro 7.5B 8.9 36.0 - - - - - -
    Retro++ 9B 25.8 48.3 - - - - - -
    Atlas 11B 26.7 56.9 - - - - - -
    Raven 11B 29.6 65.7 - - - - - -
    RA-DIT 65B 35.2 75.4 - - - - - -
    InstructGPTRAG_{\text{RAG}} 43B 37.0 78.1 52.4 70.7 / 64.3 72.4 / 65.8 71.5 53.9 51.8
    InstructRetro 43B (w/o encoder) 38.9 78.3 57.4 75.6 / 69.3 77.1 / 70.4 76.2 60.0 54.8
    InstructRetro 48B (w/ encoder) 38.6 77.8 57.0 74.8 / 67.7 76.4 / 69.0 76.1 59.8 54.6

    For open-domain benchmarks (NQ, TriviaQA), the DRAGON+ dense retriever was used to fetch top-5 context passages. For SQuAD, NewsQA, Quoref, NarrativeQA, and DROP, gold contexts provided in the datasets were supplied in the prompt.

  5. Knowl 5 — Zero-Shot Performance on Long-Form QA and Document Summarization Benchmarks

    data/table

    On four open-ended long-form QA benchmarks (doc2dial, two proprietary car manual QA datasets, and an IT documentation dataset) evaluated using F1 score, and three long-document summarization benchmarks (GovReport, SummScreenFD, QMSum) evaluated using the geometric mean of ROUGE-1, ROUGE-2, and ROUGE-L, InstructRetro demonstrates significant gains over InstructGPTRAG_{\text{RAG}} 43B and Llama 2 RAG 70B:

    Model doc2dial Car #1 Car #2 IT Doc
    Llama 2 RAG 70B 32.33 49.63 45.89 25.70
    InstructGPTRAG_{\text{RAG}} 43B 32.87 58.18 50.88 31.40
    InstructRetro 43B (w/o encoder) 35.74 (+8.73%) 63.52 (+9.18%) 57.49 (+12.99%) 34.08 (+8.54%)
    InstructRetro 48B (w/ encoder) 35.95 (+9.37%) 63.16 (+8.56%) 56.82 (+11.67%) 34.07 (+8.50%)
    Model GovReport SummScreenFD QMSum
    Llama 2 RAG 70B 16.98 10.02 14.50
    InstructGPTRAG_{\text{RAG}} 43B 12.59 10.43 15.06
    InstructRetro 43B (w/o encoder) 17.46 (+38.68%) 10.93 (+4.79%) 15.61 (+3.65%)

    InstructRetro 43B achieves an average relative gain of 10% on long-form QA and 16% on summarization tasks compared to InstructGPTRAG_{\text{RAG}} 43B.

  6. Knowl 6 — Synergy Between Retrieval-Augmented Pretraining and Instruction Tuning

    empirical result

    Before instruction tuning, zero-shot Exact Match (EM) accuracy on Natural Questions reveals that scaling parameter size leads to performance saturation: base Retro outperforms base GPT at smaller parameter counts (e.g., 2.25B), but as model size reaches 43B/48B, both base models saturate (at approximately 33% EM for GPT 43B and 34% EM for Retro 48B) due to a shared instruction-following bottleneck.

    Applying conversational instruction tuning removes this instruction-following bottleneck for both models. After instruction tuning, the benefit of retrieval pretraining becomes significantly magnified: InstructRetro 48B reaches 38.9% EM compared to 37.0% EM for InstructGPTRAG_{\text{RAG}} 43B. This indicates that retrieval pretraining and instruction tuning are mutually complementary: instruction tuning is required to unlock the contextual reasoning capabilities acquired during retrieval pretraining.

  7. Knowl 7 — Computational Cost and Scaling Efficiency of Retro-fitting

    theoretical result

    Continued pretraining with retrieval (Retro-fitting) on 100B tokens adds between 27% and 37% computational overhead in GPU hours compared to standard GPT pretraining on 100B tokens across model scales:

    GPT Model GPT GPU Hours (100B tokens) Retro Model Retro GPU Hours (100B tokens) Overhead (100B tokens)
    823M 1,408 878M 1,920 36%
    2.25B 3,226 2.5B 4,096 27%
    8.5B 12,698 9.5B 17,325 37%
    22B 37,888 24B 52,152 37%
    43B 53,329 48B 69,995 31%

    When evaluating overall pretraining budget starting from a 43B GPT model pretrained on 1.1T tokens and continued for 0.1T tokens with retrieval to obtain Retro 48B, the total compute relative to training a 43B GPT on 1.2T tokens is:

    Relative Pretraining Compute=1.1T×1.0+0.1T×(1+0.31)1.2T×1.0=102.58%\text{Relative Pretraining Compute} = \frac{1.1\text{T} \times 1.0 + 0.1\text{T} \times (1 + 0.31)}{1.2\text{T} \times 1.0} = 102.58\%

    Thus, Retro 48B requires only 2.58% additional overall GPU hours while achieving validation perplexity comparable to a standard GPT model with 4×\times larger parameter scale.

  8. Knowl 8 — Large-Scale Faiss Dense Retrieval Index Configuration for Trillion-Token Corpora

    experimental setup

    To support retrieval from a 1.2-trillion-token pretraining corpus divided into 19 billion 64-token chunks, a dense approximate nearest neighbor retrieval index is constructed using Faiss:

    • Embedding: Every chunk is encoded with BERT-large-cased at a throughput of 6.22M chunks per A100 GPU hour, totaling 3,054 GPU hours.
    • Dimension Reduction: Optimized Product Quantization (OPQ64_128) applies a rotation to improve quantization accuracy.
    • Clustering and Graph Search: Inverted File Index with 222=4,194,3042^{22} = 4,194,304 centroids accelerated by Hierarchical Navigable Small World graphs (IVF4194304_HNSW32), trained on 600M chunks in under 4 hours on a DGX-A100 node. Adding the full 19B chunks requires 192 CPU hours.
    • Encoding: Product Quantization compressing dense embeddings into 64 bits (PQ64), keeping total memory under 1 TB.
    • Query Parameters: Setting runtime parameters nprobe=4096\text{nprobe}=4096 and efSearch=32\text{efSearch}=32 achieves an average query time of 4 ms per chunk on a DGX-A100 node with a recall accuracy of 0.93 (at top-K=2000K=2000). Total query time for 100B pretraining tokens is approximately 1,736 CPU hours.
  9. Knowl 9 — Multi-Turn Dialogue Evaluation on MT-Bench

    data/table

    When aligned on the Tulu-v2 instruction dataset and evaluated on the multi-turn MT-Bench benchmark across 8 conversation categories, InstructRetro-Tulu-v2-43B achieves a higher average score (6.70) than InstructGPT-Tulu-v2-43B (6.44):

    MT-Bench Category InstructRetro-Tulu-v2-43B InstructGPT-Tulu-v2-43B
    Writing 8.85 8.15
    Roleplay 7.75 7.80
    Reasoning 5.40 4.75
    Math 3.15 2.35
    Coding 3.40 4.10
    Extraction 6.80 6.75
    STEM 8.58 8.53
    Humanities 9.68 9.10
    Turn 1 6.89 6.67
    Turn 2 6.51 6.21
    Average 6.70 6.44

    InstructRetro exhibits notable improvements in Writing (+0.70), Reasoning (+0.65), and Math (+0.80), with comparable performance on Extraction and STEM, while slightly underperforming in Coding and Roleplay due to the lack of coding and roleplay text in the base GPT pretraining corpus.

Coverage note — None was omitted; all key contributions, architecture designs, scaling results, downstream benchmark evaluations, and index ablation details are covered.

References

  1. 1.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al. Improving language models by retrieving from trillions of tokens. In ICML, 2022.
  2. 2.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. NeurIPS, 2020a.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020b.
  4. 4.Chen, M., Chu, Z., Wiseman, S., and Gimpel, K. Summscreen: A dataset for abstractive screenplay summarization. ArXiv, abs/2104.07091, 2021. URL https://api.semanticscholar.org/CorpusID:233240744.
  5. 5.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  6. 6.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models. arXiv preprint arXiv: 2210.11416, 2022. URL https://arxiv.org/abs/2210.11416v5.
  7. 7.Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R. Free dolly: Introducing the world’s first truly open instruction-tuned llm. databricks, 2023.
  8. 8.Dasigi, P., Liu, N. F., Marasovic, A., Smith, N. A., and Gardner, M. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. Conference on Empirical Methods in Natural Language Processing, 2019. doi: 10.18653/v1/D19-1606.
  9. 9.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C. Glam: Efficient scaling of language models with mixture-of-experts. International Conference on Machine Learning, 2021. URL https://arxiv.org/abs/2112.06905v2.
  11. 11.Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. North American Chapter of the Association for Computational Linguistics, 2019. doi: 10.18653/v1/N19-1246.
  12. 12.Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M. Eli5: Long form question answering. Annual Meeting of the Association for Computational Linguistics, 2019. doi: 10.18653/v1/P19-1346.
  13. 13.Feng, S., Wan, H., Gunasekara, R. C., Patel, S., Joshi, S., and Lastras, L. doc2dial: A goal-oriented document-grounded dialogue dataset. Conference on Empirical Methods in Natural Language Processing, 2020. doi: 10.18653/v1/2020.emnlp-main.652.
  14. 14.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  15. 15.Ge, T., He, K., Ke, Q., and Sun, J. Optimized product quantization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(4):744–755, 2014. doi: 10.1109/TPAMI.2013.240.
  16. 16.Gray, R. and Neuhoff, D. Quantization. IEEE Transactions on Information Theory, 44(6):2325–2383, 1998. doi: 10.1109/18.720541.
  17. 17.Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. REALM: Retrieval augmented language model pre-training. In ICML, 2020.
  18. 18.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  19. 19.Honovich, O., Scialom, T., Levy, O., and Schick, T. Unnatural instructions: Tuning language models with (almost) no human labor. Annual Meeting of the Association for Computational Linguistics, 2022. doi: 10.48550/arXiv.2212.09689.
  20. 20.Huang, J., Ping, W., Xu, P., Shoeybi, M., Chang, K. C.-C., and Catanzaro, B. Raven: In-context learning with retrieval augmented encoder-decoder language models. arXiv preprint arXiv:2308.07922, 2023.
  21. 21.Huang, L. R., Cao, S., Parulian, N. N., Ji, H., and Wang, L. Efficient attentions for long document summarization. ArXiv, abs/2104.02112, 2021. URL https://api.semanticscholar.org/CorpusID:233033613.
  22. 22.Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E. Atlas: Few-shot learning with retrieval augmented language models. arXiv preprint arXiv: 2208.03299, 2022a.
  23. 23.Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299, 2022b.
  24. 24.Johnson, J., Douze, M., and Jégou, H. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7 (3):535–547, 2019.
  25. 25.Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. Annual Meeting of the Association for Computational Linguistics, 2017. doi: 10.18653/v1/P17-1147.
  26. 26.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. Dense passage retrieval for open-domain question answering. In EMNLP, 2020.
  27. 27.Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. Generalization through memorization: Nearest neighbor language models. 2020.
  28. 28.Kim, H., Hessel, J., Jiang, L., West, P., Lu, X., Yu, Y., Zhou, P., Bras, R. L., Alikhani, M., Kim, G., Sap, M., and Choi, Y. Soda: Million-scale dialogue distillation with social commonsense contextualization. arXiv preprint arXiv: 2212.10465, 2022. URL https://arxiv.org/abs/2212.10465v2.
  29. 29.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2014.
  30. 30.Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6:317–328, 2018.
  31. 31.Kudo, T. and Richardson, J. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. Conference on Empirical Methods in Natural Language Processing, 2018. doi: 10.18653/v1/D18-2012.
  32. 32.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466, 2019. doi: 10.1162/tacl_a_00276. URL https://aclanthology.org/Q19-1026.
  33. 33.Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A. Openassistant conversations - democratizing large language model alignment. arXiv preprint arXiv: 2304.07327, 2023.
  34. 34.Laurençon, H., Saulnier, L., Wang, T., Akiki, C., Villanova del Moral, A., Le Scao, T., Von Werra, L., Mou, C., González Ponferrada, E., Nguyen, H., et al. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. Advances in Neural Information Processing Systems, 35:31809–31826, 2022.
  35. 35.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS, 2020.
  36. 36.Lin, S.-C., Asai, A., Li, M., Oguz, B., Lin, J., Mehdad, Y., tau Yih, W., and Chen, X. How to train your dragon: Diverse augmentation towards generalizable dense retrieval. arXiv preprint arXiv: 2302.07452, 2023.
  37. 37.Lin, X. V., Chen, X., Chen, M., Shi, W., Lomeli, M., James, R., Rodriguez, P., Kahn, J., Szilvasy, G., Lewis, M., Zettlemoyer, L., and tau Yih, W. RA-DIT: Retrieval-augmented dual instruction tuning. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=22OTbutug9.
  38. 38.Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., and Roberts, A. The flan collection: Designing data and methods for effective instruction tuning. International Conference on Machine Learning, 2023. doi: 10.48550/arXiv.2301.13688.
  39. 39.Malkov, Y. A. and Yashunin, D. A. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence, 42(4):824–836, 2018.
  40. 40.Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. Cross-task generalization via natural language crowdsourcing instructions. In ACL, 2022.
  41. 41.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  42. 42.OpenAI. ChatGPT. https://chat.openai.com, 2022.
  43. 43.OpenAI. GPT-4 technical report. arXiv, 2023.
  44. 44.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. NeurIPS, 2022.
  45. 45.Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., Plachouras, V., Rocktäschel, T., and Riedel, S. KILT: a benchmark for knowledge intensive language tasks. In NAACL, 2021.
  46. 46.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  47. 47.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020.
  48. 48.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. Conference on Empirical Methods in Natural Language Processing, 2016. doi: 10.18653/v1/D16-1264.
  49. 49.Rajpurkar, P., Jia, R., and Liang, P. Know what you don’t know: Unanswerable questions for squad. Annual Meeting of the Association for Computational Linguistics, 2018. doi: 10.18653/v1/P18-2124.
  50. 50.Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Scao, T. L., Biderman, S., Gao, L., Wolf, T., and Rush, A. M. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022a. URL https://openreview.net/forum?id=9Vrb9D0WI4.
  51. 51.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. Multitask prompted training enables zero-shot task generalization. In ICLR, 2022b.
  52. 52.Shi, W., Min, S., Lomeli, M., Zhou, C., Li, M., Lin, V., Smith, N. A., Zettlemoyer, L., Yih, S., and Lewis, M. In-context pretraining: Language modeling beyond document boundaries. arXiv preprint arXiv:2310.10638, 2023a.
  53. 53.Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652, 2023b.
  54. 54.Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., Zhang, E., Child, R., Aminabadi, R. Y., Bernauer, J., Song, X., Shoeybi, M., He, Y., Houston, M., Tiwary, S., and Catanzaro, B. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv, 2022.
  55. 55.Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  56. 56.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models. ARXIV, 2023a. URL https://arxiv.org/abs/2302.13971v1.
  57. 57.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023b. URL https://arxiv.org/abs/2307.09288v2.
  58. 58.Trischler, A., Wang, T., Yuan, X., Harris, J., Sordoni, A., Bachman, P., and Suleman, K. Newsqa: A machine comprehension dataset. REP4NLP@ACL, 2016. doi: 10.18653/v1/W17-2623.
  59. 59.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In NIPS, 2017.
  60. 60.Wang, B., Ping, W., Xu, P., McAfee, L., Liu, Z., Shoeybi, M., Dong, Y., Kuchaiev, O., Li, B., Xiao, C., et al. Shall we pretrain autoregressive language models with retrieval? a comprehensive study. In EMNLP, 2023a.
  61. 61.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022a.
  62. 62.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language models with self-generated instructions. Annual Meeting of the Association for Computational Linguistics, 2022b. doi: 10.48550/arXiv.2212.10560.
  63. 63.Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., and Hajishirzi, H. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv: 2306.04751, 2023b. URL https://arxiv.org/abs/2306.04751v1.
  64. 64.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In ICLR, 2022a.
  65. 65.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022b.
  66. 66.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022c.
  67. 67.Yogatama, D., de Masson d’Autume, C., and Kong, L. Adaptive semiparametric language models. Transactions of the Association for Computational Linguistics, 2021.
  68. 68.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.
  69. 69.Zhong, M., Yin, D., Yu, T., Zaidi, A., Mutuma, M., Jha, R., Awadallah, A. H., Celikyilmaz, A., Liu, Y., Qiu, X., et al. Qmsum: A new benchmark for query-based multi-domain meeting summarization. arXiv preprint arXiv:2104.05938, 2021.

Citation

MLA
Wang, B., et al. “InstructRetro: Instruction Tuning Post Retrieval-Augmented Pretraining”. arXiv, 2023, http://arxiv.org/abs/2310.07713v3.
APA
Wang, B., Ping, W., McAfee, L., Xu, P., Li, B., Shoeybi, M., & Catanzaro, B. (2023). InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining. arXiv. http://arxiv.org/abs/2310.07713v3
Chicago
Wang, B., W. Ping, L. McAfee, et al. 2023. “InstructRetro: Instruction Tuning Post Retrieval-Augmented Pretraining”. arXiv. http://arxiv.org/abs/2310.07713v3.
Harvard
Wang, B. et al. (2023) “InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.07713v3.
Vancouver
1. Wang B, Ping W, McAfee L, Xu P, Li B, Shoeybi M, Catanzaro B (2023) InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining. arXiv

BibTeX

@article{wang2023instructretro,
  title = {InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining},
  author = {Wang, Boxin and Ping, Wei and McAfee, Lawrence and Xu, Peng and Li, Bo and Shoeybi, Mohammad and Catanzaro, Bryan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.07713v3},
  eprint = {2310.07713}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/