Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study

Boxin WangWei PingPeng XuLawrence McAfeeZihan LiuMohammad ShoeybiYi DongOleksii KuchaievBo LiChaowei Xiao

article2023EMNLP85 citations

Demonstrates through scalable reproduction up to 9.5B parameters that pretraining autoregressive language models with retrieval substantially improves factual accuracy and downstream knowledge-intensive task performance over standard GPT baselines while introducing RETRO++ to boost open-domain question answering.

Listen

Standard autoregressive language models require immense parameter counts to store factual knowledge, making them expensive to deploy, difficult to update with new information, and prone to factual hallucinations. Although retrieval mechanisms have been added to models during fine-tuning or inference, whether large generative language models should be pretrained with retrieval capabilities from scratch has remained an open question. The article evaluates this question by systematically comparing standard generative models with a retrieval-augmented architecture across text generation quality, factual accuracy, downstream task performance, and question answering.

The researchers reproduced and trained a scalable retrieval-augmented architecture, known as RETRO, across configurations ranging from 148 million to 9.5 billion parameters on a 330-billion-token pretraining dataset. The model incorporates a retrieval database containing 5.3 billion text chunks indexed using dense similarity search, enabling it to retrieve relevant external text during both pretraining and generation. Crucially, the experiments evaluated these models against standard architectures under identical data schedules and model sizes to isolate the precise impact of pretraining with retrieval.

The findings show that retrieval-augmented pretraining consistently outperforms standard generative models across several critical metrics. First, in open-ended text generation, the retrieval architecture reduced repetitive text generation by an average of 21% while maintaining equal fluency and coherence. Second, the model demonstrated moderately higher factual accuracy and lower hallucination rates across established factuality benchmarks. Third, in zero-shot evaluations across nine standard tasks, the architecture significantly outperformed standard models on knowledge-intensive benchmarks while remaining competitive on reasoning tasks. Fourth, an improved variant introduced in the study, RETRO++, substantially elevated open-domain question answering performance, boosting the exact match accuracy on the Natural Questions benchmark to 54.1% compared to 40.9% for the baseline retrieval model.

These results demonstrate that pretraining models with retrieval is a highly effective design pattern for foundation models. Incorporating retrieval allows models to access external knowledge dynamically rather than relying solely on memorized parameters, facilitating easier factual updates and reducing the risk of generating incorrect or outdated information. Because the approach requires only modest computational overhead—adding less than 25% to pretraining GPU hours—it presents a favorable cost-to-performance trade-off for organizations deploying large language models.

Engineering and research teams should consider pretraining autoregressive models with retrieval when developing systems for knowledge-intensive domains and factual question answering. When implementing these architectures, practitioners should adopt enhanced evidence-routing designs such as RETRO++ and implement flexible retrieval intervals to balance latency against output accuracy. Furthermore, organizations must carefully curate the external retrieval database, as using toxic or low-quality source texts directly degrades generation safety and accuracy.

The study notes certain limitations, particularly that model evaluations were capped at 9.5 billion parameters rather than the tens or hundreds of billions used in the largest commercial systems. Additionally, system outputs remain strongly dependent on the quality, neutrality, and freshness of the underlying datastore. Nevertheless, the consistent performance gains across all evaluated model sizes provide strong confidence that pretraining with retrieval is a robust, scalable architecture for future generative systems.

Wang et al (2023).pdf
Cover for Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study

Table of Contents

  • 1 Introduction
  • 2 Key Findings
  • 2.1 Text Generation
  • 2.2 LM Evaluation Harness Benchmark
  • 2.3 Open-domain QA
  • 3 Related Work
  • 4 Model and Implementation
  • 4.1 Preliminaries of RETRO
  • 4.2 Implementation
  • 4.2.1 Retrieval Database
  • 4.2.2 Pretraining RETRO Models
  • 4.2.3 Retrieval-augmented Generation
  • 4.2.4 Batched Training for Downstream Tasks
  • 5 Open-ended Text Generation
  • 5.1 Text Quality
  • 5.1.1 Automatic Evaluation
  • 5.1.2 Human Evaluation
  • 5.2 Factuality
  • 5.3 Toxicity
  • 6 LM Evaluation Harness Benchmark
  • 7 Open-domain Question Answering
  • 7.1 Experimental Setup
  • 7.2 Results and Analysis
  • 7.3 Zero-shot evaluation with and without instruction tuning
  • 8 Conclusion
  • Limitations
  • References
  • Appendix
  • A Details of Retrieval Index
  • B Details of Pre-trained LMs
  • C Implementation Details of Retrieval-Augmented Generation
  • C.1 “Left Padding” Rule
  • C.2 Frequency of Retrieval in Text Generation
  • D Details of Evaluation for Text Generation Quality
  • D.1 Details of Automatic Evaluation for Text Generation Quality
  • D.2 Details of Human Evaluation for Text Generation Quality
  • E Details of Factuality Evaluation
  • E.1 Experimental Setup
  • E.2 Evaluation Metrics
  • F Details of Toxicity Evaluation
  • F.1 Experimental Setup
  • F.2 Evaluation Metrics
  • G Details of LM Evaluation Harness Benchmark
  • G.1 Task Details
  • G.2 Evaluation Protocol
  • G.3 Fine-tuning Performance.
  • G.4 Put Retrieval Evidence in Context for GPT in zero-shot evaluation
  • H Details of Open-domain QA
  • H.1 Experimental Setup
  • H.2 Training Details
  • H.3 Qualitative Study on NQ

Knowls

  1. Knowl 1 — RETRO retrieves chunk neighbors to condition autoregressive generation

    model/method

    RETRO is a decoder-only autoregressive language model that supplements its parametric state with retrieved text. It divides the input and retrieval corpus into contiguous chunks; in this study the chunk size is 64 tokens. To generate a chunk, RETRO retrieves nearest-neighbor chunks using the immediately preceding input chunk, then combines the retrieved text with preceding-context information through chunk-wise cross-attention. Retrieval is based on the previous chunk rather than the chunk being generated, so the model does not use future tokens to retrieve evidence for its current prediction. The retrieval database stores text chunks as values and corresponding BERT embeddings as search keys.

  2. Knowl 2 — A 330-billion-token reproduction improves validation perplexity at moderate added training cost

    empirical result

    The authors pretrained GPT and RETRO from scratch using the same 330B-token English corpus, comparable Transformer configurations, and matched training schedules. The corpus yielded 5.3B retrieval chunks of 64 tokens. Its Faiss dense-retrieval index used an IVF index with 2222^{22} centroids, HNSW acceleration, and 64-bit product-quantization codes; batched queries averaged 4 ms per chunk on a DGX-2H node, with memory use below 400 GB. RETRO models ranged from 148M to 9.5B parameters; the matched GPT models ranged from 126M to 8.3B. Validation perplexity was lower for RETRO at all four reported sizes:

    Model size Small Medium XL XXL
    GPT 17.76 13.18 10.18 7.86
    RETRO (k=2k=2 neighbors) 12.99 10.06 8.10 6.72

    For training on 330B tokens, GPT versus RETRO required 1,240 versus 1,560 GPU-hours at Small size, 3,600 versus 4,480 at Medium, and 12,000 versus 13,440 at XL, corresponding to reported RETRO overheads of 25.80%, 24.44%, and 12.00%, respectively. These costs were measured on DGX-2H nodes with eight A100 GPUs.

  3. Knowl 3 — Left padding and adjustable retrieval frequency enable practical RETRO generation

    model/method

    RETRO's chunk alignment can prevent retrieval from helping when a short prompt and the next-token prediction fall within the first chunk: there is no preceding chunk from which to retrieve. The authors address this with a left-padding rule. Padding tokens are added before the context so that the context forms a preceding chunk and the next-token prediction begins in the following chunk; for contexts longer than one chunk, this also makes the rightmost context available for retrieval. In downstream batched training and inference, question or context chunks are left-padded and answer chunks are right-padded, keeping questions separate from candidate answers and preserving autoregressive causality.

    Generation can update retrieval after every token (retrieval step 1), use the same retrieved neighbors for a full chunk (step 64 for 64-token chunks), or use an intermediate step. Step 1 gives the most frequently updated evidence but incurs more retrieval work; step 64 reduces that work but can leave neighbors stale as the generated context changes. The authors recommend shorter steps for short-answer tasks and longer steps for efficient long-form generation.

  4. Knowl 4 — RETRO reduces repetitive generations while preserving fluency and diversity

    empirical result

    On 5,000 prompts, GPT and RETRO generated up to 200 tokens using nucleus sampling with p=0.9p=0.9, retrieval step 64, and two retrieved neighbors. RETRO had a lower repetition rate at every model size; the authors report an average reduction of 21% relative to GPT. Self-BLEU and Zipf coefficients were similar, indicating no major measured change in diversity or vocabulary use. In human evaluation of the XXL models, RETRO's mean relevance score was 3.726 versus GPT's 3.715, and its mean fluency score was 3.826 versus GPT's 3.818, on five-point scales.

    Small Medium XL XXL
    Metric GPT RETRO GPT RETRO GPT RETRO GPT RETRO
    Repetition (%) 2.86 2.26 1.70 1.50 1.44 0.96 1.40 1.12
    Self-BLEU 0.29 0.30 0.29 0.30 0.29 0.29 0.31 0.31
    Zipf coefficient 0.98 0.98 0.96 0.98 0.97 0.98 0.96 0.96

    The repetition metric is the percentage of generations containing a phrase of at least two tokens repeated at least three times at the end. Lower Self-BLEU indicates greater diversity; the reported coefficients for Zipf vocabulary distributions are shown as measured.

  5. Knowl 5 — Factuality gains depend on the benchmark and decoding strategy

    empirical result

    For XL models, RETRO performed better than GPT on the FACTUALITYPROMPTS measures of hallucinated named entities and entailment under both top-pp sampling and greedy decoding. Hallucinated named-entity error (NEER) is the percentage of generations containing an entity absent from the ground-truth source; EntailR is the percentage judged entailed by that source. Greedy decoding produced lower NEER and higher EntailR than top-pp sampling for both models.

    Decoding Prompt type GPT NEER (%) RETRO NEER (%) GPT EntailR (%) RETRO EntailR (%)
    top-p=0.9p=0.9 Factual 52.42 52.14 2.93 3.11
    top-p=0.9p=0.9 Nonfactual 56.82 56.75 2.04 2.06
    Greedy Factual 39.87 37.42 12.91 16.66
    Greedy Nonfactual 45.02 42.45 8.75 10.88

    On TruthfulQA, RETRO with its pretraining retrieval database also scored above GPT on both MC1 (single-choice accuracy) and MC2 (multi-choice probability): 0.239 versus 0.222 and 0.382 versus 0.377 in the question-answer prompt format, and 0.248 versus 0.234 and 0.439 versus 0.435 in the null format. In the null format, RETRO with Wikipedia and DPR retrieval databases scored 0.242/0.437 and 0.245/0.439 on MC1/MC2, respectively. The paper notes that this improvement was less pronounced on TruthfulQA than on FACTUALITYPROMPTS; retrieval did not resolve some adversarial logical or indexical questions.

  6. Knowl 6 — Toxicity outcomes change with the retrieval database

    empirical result

    On REALTOXICITYPROMPTS, the retrieval source affected RETRO's toxicity. With the pretraining database, RETRO (XL) was more toxic than GPT (XL) on the full prompt set; with Wikipedia retrieval, its reported toxicity measures were lower than GPT's. The evaluation used 10,000 prompts and 25 sampled continuations of up to 20 tokens per prompt. Expected Maximum Toxicity averages, across prompts, each prompt's maximum toxicity among the 25 generations. Toxicity Probability is the percentage of prompts for which at least one of the 25 continuations had toxicity at least 0.5. Values are shown for all prompts and the toxic and nontoxic prompt subsets.

    Model Database Expected Maximum Toxicity Toxicity Probability (%)
    Full Toxic Nontoxic Full Toxic Nontoxic
    GPT – 0.44 0.64 0.39 37 74 27
    RETRO, top-N=2N=2, top-K=2K=2 Pretraining 0.46 0.66 0.40 40 76 30
    RETRO, top-N=2N=2, top-K=2K=2 Wikipedia 0.43 0.64 0.38 35 73 25

    Here, top-NN is the number of nearest neighbors considered and top-KK is the number returned after toxicity filtering. The paper reports similar Wikipedia results for top-N=5N=5 and 10, and finds that increasing NN and filtering for nontoxic neighbors did not provide a further clear benefit. The authors attribute the direction of the database effect to the toxicity of retrieved evidence: the pretraining evidence had higher measured toxicity than the generations, whereas Wikipedia evidence had lower measured toxicity.

  7. Knowl 7 — Zero-shot RETRO gains are concentrated in some knowledge-intensive tasks

    empirical result

    On nine LM Evaluation Harness classification tasks, RETRO's average zero-shot accuracy exceeded GPT's at each of four model sizes, although task-level results were mixed. The largest advantages appeared on knowledge-intensive tasks, particularly HellaSwag and BoolQ, while the seven other tasks did not show a consistent RETRO advantage. Scores are percentages; each entry below gives GPT followed by RETRO.

    Task Small Medium XL XXL
    HellaSwag 31.3 / 36.2 43.2 / 46.2 56.7 / 59.0 72.3 / 70.6
    BoolQ 59.3 / 61.8 57.4 / 57.2 62.2 / 62.7 67.3 / 70.7
    Lambada 41.7 / 41.4 54.1 / 55.0 63.9 / 64.0 73.9 / 72.7
    RACE 34.6 / 32.5 37.3 / 37.3 40.8 / 39.9 44.3 / 43.2
    PiQA 64.3 / 64.8 70.2 / 68.7 73.7 / 74.1 78.5 / 77.4
    WinoGrande 52.4 / 52.0 53.8 / 55.2 59.0 / 60.1 68.5 / 65.8
    ANLI-R2 35.1 / 36.2 33.5 / 33.3 34.3 / 35.3 32.2 / 35.5
    HANS 51.5 / 51.4 50.5 / 50.5 50.1 / 50.0 50.8 / 56.5
    WiC 50.0 / 50.0 50.2 / 50.0 47.8 / 49.8 52.4 / 52.4
    Average accuracy 46.7 / 47.4 50.0 / 50.4 54.3 / 55.0 60.0 / 60.5

    For example, RETRO's HellaSwag advantage was 4.9 percentage points at Small, 3.0 at Medium, and 2.3 at XL, but it was 1.7 points lower than GPT at XXL. Prepending the top retrieved passage directly to GPT in a separate zero-shot HellaSwag ablation reduced accuracy to 24.5%, 25.2%, 24.2%, and 24.1% across the four sizes, compared with GPT's 31.3%, 43.2%, 56.7%, and 72.3%. Thus, simply supplying retrieved text at inference did not reproduce RETRO's results.

  8. Knowl 8 — RETRO++ routes the strongest retrieved evidence to the decoder for open-domain QA

    model/method

    For open-domain question answering, RETRO++ modifies the original RETRO evidence pathway: it places the single most relevant retrieved passage at the beginning of the decoder input as context, while the encoder processes the remaining retrieved evidence. The study used DPR-retrieved Wikipedia passages, greedy decoding, and exact match (EM) after normalizing predictions and answers. This design differs from original RETRO's use of retrieved passages through the encoder.

    The largest reported configurations produced the following EM scores on Natural Questions (NQ) and TriviaQA. A dash means that the paper did not report a score.

    Method NQ EM TriviaQA EM
    GPT, close-book 36.1 45.1
    REALM 40.4 –
    DPR 41.5 56.8
    RAG-BART 44.5 56.1
    RAG-GPT 50.9 60.9
    FiD-Large 51.4 67.6
    RETRO, this study 40.9 59.9
    RETRO, original report 45.5 –
    RETRO++ 54.1 66.7

    RETRO++ exceeded the original RETRO report by 8.6 NQ EM points and exceeded this study's RETRO by 13.2 points; the authors identify putting the top DPR passage in the decoder context as key to the latter improvement. RETRO++ also outperformed RAG-GPT, which adds retrieval during fine-tuning, on both reported datasets. The authors note that their RETRO pretraining used 330B tokens, versus 600B in the original RETRO report.

  9. Knowl 9 — Instruction tuning improves zero-shot NQ accuracy for both retrieval-augmented models

    empirical result

    The authors instruction-tuned XXL GPT and RETRO using 128K instruction examples, then evaluated zero-shot NQ exact match. Fine-tuning used a batch size of 128, learning rate 5×10−65\times10^{-6}, weight decay 0.01, Adam with β1=0.9\beta_1=0.9 and β2=0.98\beta_2=0.98, and 1,000 steps. Evaluation used RAG-GPT-style prompting for GPT and RETRO++-style prompting for RETRO. Instruction tuning raised the scores of both methods, and RETRO++ remained higher in both conditions:

    Condition RAG-GPT EM RETRO++ EM
    Without instruction tuning 24.43 25.93
    With instruction tuning 29.75 31.16

    The instruction mixture combined social dialogue, long-form QA, generated instructions, FLAN and chain-of-thought data, and human-written conversation datasets. These results concern zero-shot NQ evaluation and are separate from the fine-tuned open-domain QA results.

  10. Knowl 10 — Retrieval database quality and training scale remain limitations

    limitation

    The paper identifies retrieval-database quality as a limitation: factual accuracy and toxicity outcomes depend on the contents and coverage of the database, and inaccurate, biased, or outdated material can compromise model outputs. Its experiments pretrain models up to 9.5B parameters on a 330B-token corpus, at substantial computational cost; the authors leave scaling retrieval-augmented pretraining further as future work. Consequently, the reported comparisons do not establish performance at the much larger model and corpus scales needed to settle how retrieval pretraining behaves at foundation-model scale.

Coverage note — The paper's fine-tuning results on Lambada and HellaSwag and its qualitative NQ examples are omitted because their findings are supplementary to the included zero-shot benchmark and open-domain QA results.

References

  1. 1.Matthew W Bilotti, Paul Ogilvie, Jamie Callan, and Eric Nyberg. 2007. Structured retrieval for question answering. In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval.
  2. 2.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020. Piqa: Reasoning about physical commonsense in natural language. In AAAI.
  3. 3.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In ICML.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv: 2210.11416.
  7. 7.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. In NAACL.
  8. 8.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  10. 10.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, J. Weston, and Michael Auli. 2019. Eli5: Long form question answering. Annual Meeting of the Association for Computational Linguistics.
  11. 11.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  12. 12.Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun. 2014. Optimized product quantization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(4):744–755.
  13. 13.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings in EMNLP.
  14. 14.R.M. Gray and D.L. Neuhoff. 1998. Quantization. IEEE Transactions on Information Theory, 44(6):2325–2383.
  15. 15.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. REALM: Retrieval augmented language model pre-training. In ICML.
  16. 16.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. International Conference On Learning Representations.
  17. 17.Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022. Unnatural instructions: Tuning language models with (almost) no human labor. Annual Meeting of the Association for Computational Linguistics.
  18. 18.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research.
  19. 19.Gautier Izacard and Édouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880.
  20. 20.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  21. 21.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547.
  22. 22.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In EMNLP.
  23. 23.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through memorization: Nearest neighbor language models.
  24. 24.Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, and Yejin Choi. 2022. Soda: Million-scale dialogue distillation with social commonsense contextualization. arXiv preprint arXiv: 2212.10465.
  25. 25.Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization.
  26. 26.Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021. Internet-augmented dialogue generation. arXiv preprint arXiv:2107.07566.
  27. 27.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. arXiv preprint arXiv: 2304.07327.
  28. 28.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. In EMNLP.
  29. 29.Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Factuality enhanced language models for open-ended text generation. NeurIPS.
  30. 30.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020a. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL.
  31. 31.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020b. Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS.
  32. 32.Stephanie C. Lin, Jacob Hilton, and Owain Evans. 2021. TruthfulQA: Measuring how models mimic human falsehoods. ACL.
  33. 33.S. Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023. The flan collection: Designing data and methods for effective instruction tuning. International Conference on Machine Learning.
  34. 34.Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence, 42(4):824–836.
  35. 35.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual knowledge in GPT. In NeurIPS.
  36. 36.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  37. 37.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial nli: A new benchmark for natural language understanding. In ACL.
  38. 38.OpenAI. 2022. ChatGPT. https://chat.openai.com.
  39. 39.OpenAI. 2023. GPT-4 technical report. arXiv.
  40. 40.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context. In NAACL.
  41. 41.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? In EMNLP.
  42. 42.Steven T. Piantadosi. 2014. Zipf’s word frequency law in natural language: A critical review and future directions. Psychonomic Bulletin & Review, 21:1112–1130.
  43. 43.Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. Wic: the word-in-context dataset for evaluating context-sensitive meaning representations. In NAACL.
  44. 44.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  45. 45.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  46. 46.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research.
  47. 47.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Winogrande: An adversarial winograd schema challenge at scale. In AAAI.
  48. 48.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567.
  49. 49.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv.
  50. 50.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS.
  52. 52.Boxin Wang, Wei Ping, Lawrence McAfee, Peng Xu, Bo Li, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Instructretro: Instruction tuning post retrieval-augmented pretraining. arXiv preprint arXiv: 2310.07713.
  53. 53.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language models with self-generated instructions. Annual Meeting of the Association for Computational Linguistics.
  54. 54.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In ICLR.
  55. 55.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  56. 56.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in detoxifying language models. Findings of EMNLP.
  57. 57.Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021. Adaptive semiparametric language models. Transactions of the Association for Computational Linguistics.
  58. 58.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In ACL.
  59. 59.Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018. Guiding neural machine translation with retrieved translation pieces. In NAACL.
  60. 60.Yangqiaoyu Zhou and Chenhao Tan. 2021. Investigating the effect of natural language explanations on out-of-distribution generalization in few-shot NLI. In Proceedings of the Second Workshop on Insights from Negative Results in NLP, pages 117–124, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  61. 61.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, page 1097–1100, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Wang, B., et al. “Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 7763–86, https://doi.org/10.18653/v1/2023.emnlp-main.482.
APA
Wang, B., Ping, W., Xu, P., McAfee, L., Liu, Z., Shoeybi, M., Dong, Y., Kuchaiev, O., Li, B., Xiao, C., Anandkumar, A., & Catanzaro, B. (2023). Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7763–7786. https://doi.org/10.18653/v1/2023.emnlp-main.482
Chicago
Wang, B., W. Ping, P. Xu, et al. 2023. “Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7763–86. https://doi.org/10.18653/v1/2023.emnlp-main.482.
Harvard
Wang, B. et al. (2023) “Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7763–7786. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.482.
Vancouver
1. Wang B, Ping W, Xu P, et al (2023) Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7763–7786

BibTeX

@inproceedings{wang-etal-2023-shall,
    title = "Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study",
    author = "Wang, Boxin  and
      Ping, Wei  and
      Xu, Peng  and
      McAfee, Lawrence  and
      Liu, Zihan  and
      Shoeybi, Mohammad  and
      Dong, Yi  and
      Kuchaiev, Oleksii  and
      Li, Bo  and
      Xiao, Chaowei  and
      Anandkumar, Anima  and
      Catanzaro, Bryan",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.482/",
    doi = "10.18653/v1/2023.emnlp-main.482",
    pages = "7763--7786"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/