Retrieval-Augmented Generation for Large Language Models: A Survey

Yunfan GaoYun XiongXinyu GaoKang-Xiang JiaJin PanYuxi BiYi DaiJiawei SunQian-Yu GuoMeng Wang

article2023arXiv4,092 citations

Synthesizes the architectural progression of retrieval-augmented generation across naive, advanced, and modular paradigms, providing a clear taxonomy of retrieval, generation, and evaluation methods designed to mitigate hallucinations in large language models.

Listen

Retrieval-Augmented Generation (RAG) has emerged as a practical way to strengthen large language models by pulling relevant external information at inference time, thereby reducing hallucinations, supplying up-to-date facts, and supporting knowledge-intensive tasks where the modelsparametric memory alone falls short.

The survey maps the rapid evolution of RAG research since the arrival of ChatGPT, organizes more than one hundred studies into three successive paradigmsNaive, Advanced, and Modular RAGand examines the core technical components of retrieval, generation, and augmentation, together with evaluation practices and open challenges.

Across these studies the authors show that targeted pre- and post-retrieval optimizations, modular pipeline designs, and tighter integration with fine-tuning consistently raise retrieval precision and answer quality; that RAG outperforms unsupervised fine-tuning on both familiar and novel factual questions; and that adaptive or iterative retrieval patterns further improve performance on multi-step reasoning tasks while lowering unnecessary context length.

These gains matter because they allow organizations to deploy LLMs in domains that require current or proprietary knowledge without retraining the entire model, while preserving traceability to source documents and lowering the risk of unsupported output.

The survey recommends combining RAG with selective fine-tuning or lightweight adapters, extending the approach to multimodal data, and developing more robust benchmarks that test noise tolerance, negative rejection, and counterfactual robustness. It also calls for production-oriented work on latency, data security, and evaluation tooling.

The analysis rests on a broad but still rapidly changing literature; quantitative comparisons across paradigms remain limited, and many reported gains depend on specific datasets or model scales whose generalizability is not yet fully established.

Cover for Retrieval-Augmented Generation for Large Language Models: A Survey

Abstract

Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising solution by incorporating knowledge from external databases. This enhances the accuracy and credibility of the generation, particularly for knowledge-intensive tasks, and allows for continuous knowledge updates and integration of domain-specific information. RAG synergistically merges LLMs' intrinsic knowledge with the vast, dynamic repositories of external databases. This comprehensive review paper offers a detailed examination of the progression of RAG paradigms, encompassing the Naive RAG, the Advanced RAG, and the Modular RAG. It meticulously scrutinizes the tripartite foundation of RAG frameworks, which includes the retrieval, the generation and the augmentation techniques. The paper highlights the state-of-the-art technologies embedded in each of these critical components, providing a profound understanding of the advancements in RAG systems. Furthermore, this paper introduces up-to-date evaluation framework and benchmark. At the end, this article delineates the challenges currently faced and points out prospective avenues for research and development.

Table of Contents

  • I Introduction
  • II Overview of RAG
  • II-A Naive RAG
  • II-B Advanced RAG
  • II-C Modular RAG
  • II-C1 New Modules
  • II-C2 New Patterns
  • II-D RAG vs Fine-tuning
  • III Retrieval
  • III-A Retrieval Source
  • III-A1 Data Structure
  • III-A2 Retrieval Granularity
  • III-B Indexing Optimization
  • III-B1 Chunking Strategy
  • III-B2 Metadata Attachments
  • III-B3 Structural Index
  • III-C Query Optimization
  • III-C1 Query Expansion
  • III-C2 Query Transformation
  • III-C3 Query Routing
  • III-D Embedding
  • III-D1 Mix/hybrid Retrieval
  • III-D2 Fine-tuning Embedding Model
  • III-E Adapter
  • IV Generation
  • IV-A Context Curation
  • IV-A1 Reranking
  • IV-A2 Context Selection/Compression
  • IV-B LLM Fine-tuning
  • V Augmentation process in RAG
  • V-A Iterative Retrieval
  • V-B Recursive Retrieval
  • V-C Adaptive Retrieval
  • VI Task and Evaluation
  • VI-A Downstream Task
  • VI-B Evaluation Target
  • VI-C Evaluation Aspects
  • VI-C1 Quality Scores
  • VI-C2 Required Abilities
  • VI-D Evaluation Benchmarks and Tools
  • VII Discussion and Future Prospects
  • VII-A RAG vs Long Context
  • VII-B RAG Robustness
  • VII-C Hybrid Approaches
  • VII-D Scaling laws of RAG
  • VII-E Production-Ready RAG
  • VII-F Multi-modal RAG
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Retrieval-Augmented Generation Paradigms: Naive, Advanced, and Modular RAG

    model/method

    Retrieval-Augmented Generation (RAG) integrates external knowledge retrieval into large language model (LLM) workflows. RAG architectures are categorized into three developmental paradigms:

    1. Naive RAG: Follows a standard sequential Retrieve-Read pipeline consisting of three stages:

      • Indexing: Documents are parsed, split into plain text chunks, encoded into vector representations with an embedding model, and indexed in a vector database.
      • Retrieval: Given an input query, the same embedding model encodes the query to search for the top-KK most semantically similar chunks via cosine or dot-product similarity.
      • Generation: The retrieved chunks and the original user query are concatenated into an augmented prompt and passed to the LLM to generate the final response. Limitations: Low retrieval precision/recall, context fragmentation, generation hallucinations unsupported by retrieved context, and redundancy.
    2. Advanced RAG: Augments the Naive pipeline with specialized pre-retrieval and post-retrieval strategies:

      • Pre-retrieval: Focuses on optimizing the data index (e.g., sliding windows, metadata tagging, hierarchical indexing) and refining the user query (e.g., rewriting, expansion, multi-query routing).
      • Post-retrieval: Processes retrieved chunks prior to LLM ingestion via reranking (relocating key passages to prompt boundaries to mitigate the lost-in-the-middle phenomenon) and context selection/compression.
    3. Modular RAG: Transcends fixed sequential flows by decomposing RAG into reconfigurable functional modules and adaptable interaction patterns:

      • New Modules: Specialized modules such as Search (engine/database query generation), Memory (dynamic unbounded memory pools), Routing (semantic or metadata pathing), Predict (LLM-generated context), and Task Adapters (zero-shot prompt/retriever specialization).
      • Flexible Patterns: Non-linear workflows such as iterative retrieval (alternating retrieve and generate), recursive retrieval (multi-hop problem decomposition), and adaptive retrieval (using internal model confidence or special control tokens to trigger retrieval on demand).
  2. Knowl 2 — Comparison Framework: RAG vs. Fine-Tuning vs. Prompt Engineering

    model/method

    Optimization methods for large language models (LLMs) can be analyzed across two dimensions: External Knowledge Required and Model Adaptation Required:

    • Prompt Engineering: Characterized by low external knowledge requirements and low model adaptation requirements. It leverages the model's frozen, pre-existing parametric capabilities with minimal intervention.
    • Naive and Advanced RAG: Characterized by high external knowledge requirements and low model adaptation requirements. It supplies external facts in real time without retraining model parameters, offering high interpretability, traceable provenance, and low computational update overhead, though incurring additional retrieval latency.
    • Fine-Tuning (FT): Characterized by low external factual updates and high model adaptation requirements. It updates model weights to internalize specific styles, linguistic structures, or output formats. However, unsupervised fine-tuning is inefficient at instilling brand-new factual knowledge, requires substantial computation, and risks hallucinating on unseen data distributions.
    • Modular RAG / Hybrid Integration: Bridges the high external knowledge and high model adaptation regimes by combining external document retrieval with joint generator/retriever fine-tuning, adapter alignment, and reinforcement learning.
  3. Knowl 3 — Pre-Retrieval Optimization Strategies: Indexing and Query Transformation

    model/method

    Pre-retrieval optimizations aim to improve the quality of indexed representations and align ambiguous user queries with document search spaces:

    1. Indexing Optimization:

      • Chunking Granularity: Balances context length and noise. Strategies include Small2Big (indexing small units such as sentences or propositions for high retrieval precision, while returning surrounding large paragraphs as context to the LLM) and sliding window chunking with token overlap.
      • Metadata Enrichment: Augmenting chunks with document timestamps, page numbers, authors, section headers, or synthetic hypothetical questions generated by an LLM (Reverse HyDE) to bridge query-chunk semantic mismatch.
      • Structural Indices: Organizing documents into hierarchical parent-child trees (storing summaries at parent nodes for fast tree-traversal) or Knowledge Graph (KG) indices that capture structural and semantic entity-relation links.
    2. Query Optimization:

      • Query Expansion: Expanding a single query into multiple perspectives via multi-query generation, decomposing complex queries into sub-questions using least-to-most prompting, or validating expanded sub-queries through Chain-of-Verification (CoVe).
      • Query Transformation: Modifying the search string via query rewriting (using LLMs or specialized small language models like RRR and BEQUE), generating hypothetical answers as search targets (HyDE), or abstracting the query into a high-level conceptual problem (Step-back Prompting).
      • Query Routing: Directing queries dynamically to specific knowledge bases, search engines, or KG endpoints using metadata filters or semantic vector routing.
  4. Knowl 4 — Post-Retrieval Optimization Strategies: Reranking and Context Compression

    model/method

    Feeding all retrieved documents directly into an LLM causes context dilution, noise accumulation, and degradation due to the lost-in-the-middle effect (where LLMs attend primarily to the beginning and end of long prompts). Post-retrieval optimizes the retrieved content through:

    1. Reranking:

      • Reorders candidate passages to place the most relevant information at prompt edges and discards low-scoring passages.
      • Implemented via rule-based metrics (e.g., Diversity, MRR), specialized cross-encoder models (e.g., SpanBERT, BGE-Reranker, Cohere Rerank), or general LLM-based prompting.
    2. Context Compression and Selection:

      • Token-Level Elimination: Small language models (SLMs such as GPT-2 Small or LLaMA-7B in LLMLingua / LongLLMLingua) identify and strip non-essential tokens, maintaining language structure while compressing prompt length.
      • Information Extraction and Condensation: Encoders trained via contrastive loss (e.g., RECOMP, PRCA) compress retrieved multi-document context into compact summaries or selective extracts conditioned on the query.
      • Filter-Reranker Paradigm: SLMs act as coarse filters to discard irrelevant context, followed by LLMs performing fine-grained reordering and relevance critique prior to generation.
  5. Knowl 5 — Retrieval Augmentation Processes: Iterative, Recursive, and Adaptive Retrieval

    model/method

    Beyond single-turn (once) retrieval, advanced RAG frameworks employ three dynamic augmentation workflows:

    1. Iterative Retrieval:

      • Alternates cyclically between document retrieval and partial response generation (e.g., ITER-RETGEN).
      • The LLM's intermediate generated output serves as augmented context to formulate subsequent search queries, iteratively collecting complementary context until a stopping condition or iteration limit is reached.
    2. Recursive Retrieval:

      • Gradually refines queries by breaking down complex multi-hop problems into hierarchical or sequential sub-problems (e.g., IRCoT, Tree of Clarifications).
      • In structured corpora, it starts with document-level summaries to identify pertinent sections, followed by secondary fine-grained retrievals within those specific sections.
    3. Adaptive Retrieval:

      • Enables the RAG system or LLM to dynamically determine when retrieval is required and what content to retrieve.
      • Approaches include confidence-based thresholds (e.g., FLARE triggers retrieval when token generation probabilities drop below a predefined threshold) and reflection tokens (e.g., Self-RAG utilizes special [Retrieve] and [Critic] tokens to autonomously evaluate retrieval necessity and grade retrieved document relevance).
  6. Knowl 6 — Evaluation System for RAG: Targets, Quality Dimensions, and Key Capabilities

    definition

    The evaluation of RAG systems is structured around two evaluation targets, three primary quality scores, and four core robustness abilities:

    1. Evaluation Targets:

      • Retrieval Quality: Evaluates the precision and recall of the context retrieval module using information retrieval metrics such as Hit Rate, Mean Reciprocal Rank (MRR), and Normalized Discounted Cumulative Gain (NDCG).
      • Generation Quality: Evaluates the coherence, relevance, factual accuracy, and safety of the final generated output.
    2. Primary Quality Scores (The RAG Triad):

      • Context Relevance: Assesses whether the retrieved context contains only pertinent information without excessive noise.
      • Answer Faithfulness: Measures whether the generated answer is strictly grounded in and supported by the retrieved context, penalizing hallucinations.
      • Answer Relevance: Measures whether the generated response directly and accurately addresses the user's input query.
    3. Core Required Abilities:

      • Noise Robustness: Ability to generate accurate answers when retrieved context contains distracting or irrelevant passages.
      • Negative Rejection: Ability of the model to abstain from answering when the retrieved documents do not contain sufficient information.
      • Information Integration: Ability to synthesize answers requiring multi-hop reasoning across multiple retrieved documents.
      • Counterfactual Robustness: Ability to detect and disregard known factual errors or contradictions injected into retrieved documents.
  7. Knowl 7 — Summary of RAG Evaluation Frameworks and Automated Benchmarks

    data/table

    RAG evaluation employs standardized benchmarks to assess specific model capabilities, alongside automated LLM-as-a-judge frameworks to compute quality scores across retrieval and generation:

    Framework Evaluation Targets Evaluation Aspects Quantitative Metrics
    RGB Retrieval Generation Quality Noise Robustness, Negative Rejection, Information Integration, Counterfactual Robustness Accuracy, Exact Match (EM)
    RECALL Generation Quality Counterfactual Robustness Reappearance Rate (R-Rate)
    RAGAS Retrieval Generation Quality Context Relevance, Faithfulness, Answer Relevance LLM-judged custom scores, Cosine Similarity
    ARES Retrieval Generation Quality Context Relevance, Faithfulness, Answer Relevance Accuracy (using fine-tuned LLM judges)
    TruLens Retrieval Generation Quality Context Relevance, Faithfulness, Answer Relevance LLM-judged custom triad scores
    CRUD Retrieval Generation Quality Creative Generation, Knowledge-intensive QA, Error Correction, Summarization BLEU, ROUGE-L, BERTScore, RAGQuestEval

    Benchmarks (RGB, RECALL, CRUD) test model robustness under specific stress scenarios (e.g., injecting noisy, counterfactual, or multi-document context), while evaluation toolkits (RAGAS, ARES, TruLens) leverage LLM-based prompting or fine-tuned evaluators to score context relevance, faithfulness, and answer relevance without requiring extensive human annotations.

  8. Knowl 8 — Retrieval Source Modalities and Granularity Trade-offs in RAG

    model/method

    RAG systems operate across diverse data structures and retrieval unit granularities, each presenting distinct trade-offs:

    1. Retrieval Source Types:

      • Unstructured Text: Standard corpora and encyclopedic dumps (e.g., Wikipedia, PubMed, domain-specific text).
      • Semi-Structured Data: Tables and PDFs. Addressed via LLM code generation (Text-to-SQL, TableGPT) or linearizing tables into text format.
      • Structured Knowledge Graphs (KGs): Verified entities, triples, and subgraphs (e.g., KnowledGPT, G-Retriever via Prize-Collecting Steiner Tree optimization) providing high precision and multi-hop reasoning pathways.
      • LLM-Generated Content: Using model-generated rationales or memory pools (e.g., GenRead, Self-Mem) as internal knowledge retrieval sources.
      • Multimodal Modalities: Images (RA-CM3, BLIP-2), Audio/Speech (GSS, UEOP), Video (Vid2Seq), and Code repositories (RBPS).
    2. Retrieval Granularity Spectrum:

      • Granularities range from fine to coarse: Token \rightarrow Phrase \rightarrow Sentence \rightarrow Proposition \rightarrow Chunk \rightarrow Document (or Entity \rightarrow Triplet \rightarrow Subgraph in KGs).
      • Coarse-grained units (documents, large chunks) capture rich contextual context but introduce noise and distract dense retrievers.
      • Fine-grained units (sentences, propositions) maximize semantic precision and relevance but risk lacking surrounding context, necessitating techniques like Small2Big retrieval.
  9. Knowl 9 — Retriever-Generator Alignment and Cooperative Fine-Tuning in RAG

    model/method

    To overcome the disparity between independent retrievers and generators, RAG incorporates joint alignment and tuning strategies:

    1. LM-Supervised Retrieval (LSR): Uses generator feedback to supervise the retriever. For instance, REPLUG computes the Kullback-Leibler (KL) divergence between the retriever's document similarity distribution PR(dq)P_R(d|q) and the generator's document likelihood distribution PG(yq,d)P_G(y|q, d) to optimize retriever embeddings without requiring joint cross-attention.
    2. Dual Instruction Tuning (e.g., RA-DIT): Updates both retriever and LLM parameters simultaneously, aligning retriever scoring functions with LLM generation probabilities via KL divergence, followed by supervised fine-tuning of the generator on task prompts.
    3. Pluggable Contextual Adapters: When LLM parameters are inaccessible (black-box APIs) or fine-tuning is computationally constrained, intermediate adapters (e.g., AAR, PRCA, BGM) are trained between the retriever and generator to reorder, filter, or reformat retrieved passages to maximize LLM generation performance.
  10. Knowl 10 — Technical Challenges and Future Research Horizons in RAG

    limitation

    Several fundamental challenges define ongoing and future RAG research:

    1. RAG vs. Super-Long Context LLMs: As LLMs expand context windows beyond 200k tokens, direct context insertion becomes feasible. However, RAG remains essential due to lower inference latency, reduced computational cost, and explicit, verifiable source attribution. Future work focuses on hybrid systems where RAG selects information for long-context reasoning.
    2. Robustness to Misinformation and Adversarial Noise: Unfiltered or counterfactual retrieved documents can severely degrade generation quality. Interestingly, some empirical studies indicate that including specific irrelevant documents can occasionally improve generation accuracy, underscoring the need for robust noise-filtering mechanisms.
    3. Scaling Laws in RAG: While parameter scaling laws are established for standalone LLMs, their behavior in joint retrieval-augmented pre-training (e.g., RETRO++) and the existence of Inverse Scaling Laws (where smaller RAG models might outperform larger standalone LLMs) remain under active investigation.
    4. Production Readiness and Enterprise Deployment: Key engineering hurdles include low-latency retrieval over billion-scale vectors, multi-tenant data access control, preventing sensitive metadata/source leakage through LLM prompts, and transitioning from loose module pipelines to robust enterprise platforms.

Coverage note — No substantial contributed material was omitted. The knowls comprehensively cover the three RAG paradigms, pre/post-retrieval optimizations, modular extensions and workflows, RAG vs fine-tuning taxonomy, evaluation criteria and benchmarks, multi-source/granularity retrieval, joint alignment methods, and key future research challenges.

References

  1. 1.N. Kandpal, H. Deng, A. Roberts, E. Wallace, and C. Raffel, “Large language models struggle to learn long-tail knowledge,” in International Conference on Machine Learning. PMLR, 2023, pp. 15 696–15 707.
  2. 2.Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen et al., “Siren’s song in the ai ocean: A survey on hallucination in large language models,” arXiv preprint arXiv:2309.01219, 2023.
  3. 3.D. Arora, A. Kini, S. R. Chowdhury, N. Natarajan, G. Sinha, and A. Sharma, “Gar-meets-rag paradigm for zero-shot information retrieval,” arXiv preprint arXiv:2310.20158, 2023.
  4. 4.P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W.-t. Yih, T. Rockt ¨ aschel ¨ et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020.
  5. 5.S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark et al., “Improving language models by retrieving from trillions of tokens,” in International conference on machine learning. PMLR, 2022, pp. 2206–2240.
  6. 6.L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” Advances in neural information processing systems, vol. 35, pp. 27 730–27 744, 2022.
  7. 7.X. Ma, Y. Gong, P. He, H. Zhao, and N. Duan, “Query rewriting for retrieval-augmented large language models,” arXiv preprint arXiv:2305.14283, 2023.
  8. 8.I. ILIN, “Advanced rag techniques: an illustrated overview,” https://pub.towardsai.net/advanced-rag-techniques-an-illustrated-overview-04d193d8fec6, 2023.
  9. 9.W. Peng, G. Li, Y. Jiang, Z. Wang, D. Ou, X. Zeng, E. Chen et al., “Large language model based long-tail query rewriting in taobao search,” arXiv preprint arXiv:2311.03758, 2023.
  10. 10.H. S. Zheng, S. Mishra, X. Chen, H.-T. Cheng, E. H. Chi, Q. V. Le, and D. Zhou, “Take a step back: Evoking reasoning via abstraction in large language models,” arXiv preprint arXiv:2310.06117, 2023.
  11. 11.L. Gao, X. Ma, J. Lin, and J. Callan, “Precise zero-shot dense retrieval without relevance labels,” arXiv preprint arXiv:2212.10496, 2022.
  12. 12.V. Blagojevi, “Enhancing rag pipelines in haystack: Introducing diversityranker and lostinthemiddleranker,” https://towardsdatascience.com/enhancing-rag-pipelines-in-haystack-45f14e2bc9f5, 2023.
  13. 13.W. Yu, D. Iter, S. Wang, Y. Xu, M. Ju, S. Sanyal, C. Zhu, M. Zeng, and M. Jiang, “Generate rather than retrieve: Large language models are strong context generators,” arXiv preprint arXiv:2209.10063, 2022.
  14. 14.Z. Shao, Y. Gong, Y. Shen, M. Huang, N. Duan, and W. Chen, “Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy,” arXiv preprint arXiv:2305.15294, 2023.
  15. 15.X. Wang, Q. Yang, Y. Qiu, J. Liang, Q. He, Z. Gu, Y. Xiao, and W. Wang, “Knowledgpt: Enhancing large language models with retrieval and storage access on knowledge bases,” arXiv preprint arXiv:2308.11761, 2023.
  16. 16.A. H. Raudaschl, “Forget rag, the future is rag-fusion,” https://towardsdatascience.com/forget-rag-the-future-is-rag-fusion-1147298d8ad1, 2023.
  17. 17.X. Cheng, D. Luo, X. Chen, L. Liu, D. Zhao, and R. Yan, “Lift yourself up: Retrieval-augmented text generation with self memory,” arXiv preprint arXiv:2305.02437, 2023.
  18. 18.S. Wang, Y. Xu, Y. Fang, Y. Liu, S. Sun, R. Xu, C. Zhu, and M. Zeng, “Training data is more valuable than you think: A simple and effective method by retrieving from training data,” arXiv preprint arXiv:2203.08773, 2022.
  19. 19.X. Li, E. Nie, and S. Liang, “From classification to generation: Insights into crosslingual retrieval augmented icl,” arXiv preprint arXiv:2311.06595, 2023.
  20. 20.D. Cheng, S. Huang, J. Bi, Y. Zhan, J. Liu, Y. Wang, H. Sun, F. Wei, D. Deng, and Q. Zhang, “Uprise: Universal prompt retrieval for improving zero-shot evaluation,” arXiv preprint arXiv:2303.08518, 2023.
  21. 21.Z. Dai, V. Y. Zhao, J. Ma, Y. Luan, J. Ni, J. Lu, A. Bakalov, K. Guu, K. B. Hall, and M.-W. Chang, “Promptagator: Few-shot dense retrieval from 8 examples,” arXiv preprint arXiv:2209.11755, 2022.
  22. 22.Z. Sun, X. Wang, Y. Tay, Y. Yang, and D. Zhou, “Recitation-augmented language models,” arXiv preprint arXiv:2210.01296, 2022.
  23. 23.O. Khattab, K. Santhanam, X. L. Li, D. Hall, P. Liang, C. Potts, and M. Zaharia, “Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp,” arXiv preprint arXiv:2212.14024, 2022.
  24. 24.Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig, “Active retrieval augmented generation,” arXiv preprint arXiv:2305.06983, 2023.
  25. 25.A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-rag: Learning to retrieve, generate, and critique through self-reflection,” arXiv preprint arXiv:2310.11511, 2023.
  26. 26.Z. Ke, W. Kong, C. Li, M. Zhang, Q. Mei, and M. Bendersky, “Bridging the preference gap between retrievers and llms,” arXiv preprint arXiv:2401.06954, 2024.
  27. 27.X. V. Lin, X. Chen, M. Chen, W. Shi, M. Lomeli, R. James, P. Rodriguez, J. Kahn, G. Szilvasy, M. Lewis et al., “Ra-dit: Retrieval-augmented dual instruction tuning,” arXiv preprint arXiv:2310.01352, 2023.
  28. 28.O. Ovadia, M. Brief, M. Mishaeli, and O. Elisha, “Fine-tuning or retrieval? comparing knowledge injection in llms,” arXiv preprint arXiv:2312.05934, 2023.
  29. 29.T. Lan, D. Cai, Y. Wang, H. Huang, and X.-L. Mao, “Copy is all you need,” in The Eleventh International Conference on Learning Representations, 2022.
  30. 30.T. Chen, H. Wang, S. Chen, W. Yu, K. Ma, X. Zhao, D. Yu, and H. Zhang, “Dense x retrieval: What retrieval granularity should we use?” arXiv preprint arXiv:2312.06648, 2023.
  31. 31.F. Luo and M. Surdeanu, “Divide & conquer for entailment-aware multi-hop evidence retrieval,” arXiv preprint arXiv:2311.02616, 2023.
  32. 32.Q. Gou, Z. Xia, B. Yu, H. Yu, F. Huang, Y. Li, and N. Cam-Tu, “Diversify question generation with retrieval-augmented style transfer,” arXiv preprint arXiv:2310.14503, 2023.
  33. 33.Z. Guo, S. Cheng, Y. Wang, P. Li, and Y. Liu, “Prompt-guided retrieval augmentation for non-knowledge-intensive tasks,” arXiv preprint arXiv:2305.17653, 2023.
  34. 34.Z. Wang, J. Araki, Z. Jiang, M. R. Parvez, and G. Neubig, “Learning to filter context for retrieval-augmented generation,” arXiv preprint arXiv:2311.08377, 2023.
  35. 35.M. Seo, J. Baek, J. Thorne, and S. J. Hwang, “Retrieval-augmented data augmentation for low-resource domain tasks,” arXiv preprint arXiv:2402.13482, 2024.
  36. 36.Y. Ma, Y. Cao, Y. Hong, and A. Sun, “Large language model is not a good few-shot information extractor, but a good reranker for hard samples!” arXiv preprint arXiv:2303.08559, 2023.
  37. 37.X. Du and H. Ji, “Retrieval-augmented generative question answering for event argument extraction,” arXiv preprint arXiv:2211.07067, 2022.
  38. 38.L. Wang, N. Yang, and F. Wei, “Learning to retrieve in-context examples for large language models,” arXiv preprint arXiv:2307.07164, 2023.
  39. 39.S. Rajput, N. Mehta, A. Singh, R. H. Keshavan, T. Vu, L. Heldt, L. Hong, Y. Tay, V. Q. Tran, J. Samost et al., “Recommender systems with generative retrieval,” arXiv preprint arXiv:2305.05065, 2023.
  40. 40.B. Jin, H. Zeng, G. Wang, X. Chen, T. Wei, R. Li, Z. Wang, Z. Li, Y. Li, H. Lu et al., “Language models as semantic indexers,” arXiv preprint arXiv:2310.07815, 2023.
  41. 41.R. Anantha, T. Bethi, D. Vodianik, and S. Chappidi, “Context tuning for retrieval augmented generation,” arXiv preprint arXiv:2312.05708, 2023.
  42. 42.G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave, “Few-shot learning with retrieval augmented language models,” arXiv preprint arXiv:2208.03299, 2022.
  43. 43.J. Huang, W. Ping, P. Xu, M. Shoeybi, K. C.-C. Chang, and B. Catanzaro, “Raven: In-context learning with retrieval augmented encoder-decoder language models,” arXiv preprint arXiv:2308.07922, 2023.
  44. 44.B. Wang, W. Ping, P. Xu, L. McAfee, Z. Liu, M. Shoeybi, Y. Dong, O. Kuchaiev, B. Li, C. Xiao et al., “Shall we pretrain autoregressive language models with retrieval? a comprehensive study,” arXiv preprint arXiv:2304.06762, 2023.
  45. 45.B. Wang, W. Ping, L. McAfee, P. Xu, B. Li, M. Shoeybi, and B. Catanzaro, “Instructretro: Instruction tuning post retrieval-augmented pre-training,” arXiv preprint arXiv:2310.07713, 2023.
  46. 46.S. Siriwardhana, R. Weerasekera, E. Wen, T. Kaluarachchi, R. Rana, and S. Nanayakkara, “Improving the domain adaptation of retrieval augmented generation (rag) models for open domain question answering,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 1–17, 2023.
  47. 47.Z. Yu, C. Xiong, S. Yu, and Z. Liu, “Augmentation-adapted retriever improves generalization of language models as generic plug-in,” arXiv preprint arXiv:2305.17331, 2023.
  48. 48.O. Yoran, T. Wolfson, O. Ram, and J. Berant, “Making retrieval-augmented language models robust to irrelevant context,” arXiv preprint arXiv:2310.01558, 2023.
  49. 49.H.-T. Chen, F. Xu, S. A. Arora, and E. Choi, “Understanding retrieval augmentation for long-form question answering,” arXiv preprint arXiv:2310.12150, 2023.
  50. 50.W. Yu, H. Zhang, X. Pan, K. Ma, H. Wang, and D. Yu, “Chain-of-note: Enhancing robustness in retrieval-augmented language models,” arXiv preprint arXiv:2311.09210, 2023.
  51. 51.S. Xu, L. Pang, H. Shen, X. Cheng, and T.-S. Chua, “Search-in-the-chain: Towards accurate, credible and traceable large language models for knowledgeintensive tasks,” CoRR, vol. abs/2304.14732, 2023.
  52. 52.M. Berchansky, P. Izsak, A. Caciularu, I. Dagan, and M. Wasserblat, “Optimizing retrieval-augmented reader models via token elimination,” arXiv preprint arXiv:2310.13682, 2023.
  53. 53.J. Lala, O. O’Donoghue, A. Shtedritski, S. Cox, S. G. Rodriques, ´ and A. D. White, “Paperqa: Retrieval-augmented generative agent for scientific research,” arXiv preprint arXiv:2312.07559, 2023.
  54. 54.F. Cuconasu, G. Trappolini, F. Siciliano, S. Filice, C. Campagnano, Y. Maarek, N. Tonellotto, and F. Silvestri, “The power of noise: Redefining retrieval for rag systems,” arXiv preprint arXiv:2401.14887, 2024.
  55. 55.Z. Zhang, X. Zhang, Y. Ren, S. Shi, M. Han, Y. Wu, R. Lai, and Z. Cao, “Iag: Induction-augmented generation framework for answering reasoning questions,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1–14.
  56. 56.N. Thakur, L. Bonifacio, X. Zhang, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, B. Chen, M. Rezagholizadeh et al., “Nomiracl: Knowing when you don’t know for robust multilingual retrieval-augmented generation,” arXiv preprint arXiv:2312.11361, 2023.
  57. 57.G. Kim, S. Kim, B. Jeon, J. Park, and J. Kang, “Tree of clarifications: Answering ambiguous questions with retrieval-augmented large language models,” arXiv preprint arXiv:2310.14696, 2023.
  58. 58.Y. Wang, P. Li, M. Sun, and Y. Liu, “Self-knowledge guided retrieval augmentation for large language models,” arXiv preprint arXiv:2310.05002, 2023.
  59. 59.Z. Feng, X. Feng, D. Zhao, M. Yang, and B. Qin, “Retrieval-generation synergy augmented large language models,” arXiv preprint arXiv:2310.05149, 2023.
  60. 60.P. Xu, W. Ping, X. Wu, L. McAfee, C. Zhu, Z. Liu, S. Subramanian, E. Bakhturina, M. Shoeybi, and B. Catanzaro, “Retrieval meets long context large language models,” arXiv preprint arXiv:2310.03025, 2023.
  61. 61.H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,” arXiv preprint arXiv:2212.10509, 2022.
  62. 62.R. Ren, Y. Wang, Y. Qu, W. X. Zhao, J. Liu, H. Tian, H. Wu, J.-R. Wen, and H. Wang, “Investigating the factual knowledge boundary of large language models with retrieval augmentation,” arXiv preprint arXiv:2307.11019, 2023.
  63. 63.P. Sarthi, S. Abdullah, A. Tuli, S. Khanna, A. Goldie, and C. D. Manning, “Raptor: Recursive abstractive processing for tree-organized retrieval,” arXiv preprint arXiv:2401.18059, 2024.
  64. 64.O. Ram, Y. Levine, I. Dalmedigos, D. Muhlgay, A. Shashua, K. Leyton-Brown, and Y. Shoham, “In-context retrieval-augmented language models,” arXiv preprint arXiv:2302.00083, 2023.
  65. 65.Y. Ren, Y. Cao, P. Guo, F. Fang, W. Ma, and Z. Lin, “Retrieve-and-sample: Document-level event argument extraction via hybrid retrieval augmentation,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 293–306.
  66. 66.Z. Wang, X. Pan, D. Yu, D. Yu, J. Chen, and H. Ji, “Zemi: Learning zero-shot semi-parametric language models from multiple tasks,” arXiv preprint arXiv:2210.00185, 2022.
  67. 67.S.-Q. Yan, J.-C. Gu, Y. Zhu, and Z.-H. Ling, “Corrective retrieval augmented generation,” arXiv preprint arXiv:2401.15884, 2024.
  68. 68.P. Jain, L. B. Soares, and T. Kwiatkowski, “1-pager: One pass answer generation and evidence retrieval,” arXiv preprint arXiv:2310.16568, 2023.
  69. 69.H. Yang, Z. Li, Y. Zhang, J. Wang, N. Cheng, M. Li, and J. Xiao, “Prca: Fitting black-box large language models for retrieval question answering via pluggable reward-driven contextual adapter,” arXiv preprint arXiv:2310.18347, 2023.
  70. 70.S. Zhuang, B. Liu, B. Koopman, and G. Zuccon, “Open-source large language models are strong zero-shot query likelihood models for document ranking,” arXiv preprint arXiv:2310.13243, 2023.
  71. 71.F. Xu, W. Shi, and E. Choi, “Recomp: Improving retrieval-augmented lms with compression and selective augmentation,” arXiv preprint arXiv:2310.04408, 2023.
  72. 72.W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W.-t. Yih, “Replug: Retrieval-augmented black-box language models,” arXiv preprint arXiv:2301.12652, 2023.
  73. 73.E. Melz, “Enhancing llm intelligence with arm-rag: Auxiliary rationale memory for retrieval augmented generation,” arXiv preprint arXiv:2311.04177, 2023.
  74. 74.H. Wang, W. Huang, Y. Deng, R. Wang, Z. Wang, Y. Wang, F. Mi, J. Z. Pan, and K.-F. Wong, “Unims-rag: A unified multi-source retrieval-augmented generation for personalized dialogue systems,” arXiv preprint arXiv:2401.13256, 2024.
  75. 75.Z. Luo, C. Xu, P. Zhao, X. Geng, C. Tao, J. Ma, Q. Lin, and D. Jiang, “Augmented large language models with parametric knowledge guiding,” arXiv preprint arXiv:2305.04757, 2023.
  76. 76.X. Li, Z. Liu, C. Xiong, S. Yu, Y. Gu, Z. Liu, and G. Yu, “Structure-aware language model pretraining improves dense retrieval on structured data,” arXiv preprint arXiv:2305.19912, 2023.
  77. 77.M. Kang, J. M. Kwak, J. Baek, and S. J. Hwang, “Knowledge graph-augmented language models for knowledge-grounded dialogue generation,” arXiv preprint arXiv:2305.18846, 2023.
  78. 78.W. Shen, Y. Gao, C. Huang, F. Wan, X. Quan, and W. Bi, “Retrieval-generation alignment for end-to-end task-oriented dialogue system,” arXiv preprint arXiv:2310.08877, 2023.
  79. 79.T. Shi, L. Li, Z. Lin, T. Yang, X. Quan, and Q. Wang, “Dual-feedback knowledge retrieval for task-oriented dialogue systems,” arXiv preprint arXiv:2310.14528, 2023.
  80. 80.P. Ranade and A. Joshi, “Fabula: Intelligence report generation using retrieval-augmented narrative construction,” arXiv preprint arXiv:2310.13848, 2023.
  81. 81.X. Jiang, R. Zhang, Y. Xu, R. Qiu, Y. Fang, Z. Wang, J. Tang, H. Ding, X. Chu, J. Zhao et al., “Think and retrieval: A hypothesis knowledge graph enhanced medical large language models,” arXiv preprint arXiv:2312.15883, 2023.
  82. 82.J. Baek, S. Jeong, M. Kang, J. C. Park, and S. J. Hwang, “Knowledge-augmented language model verification,” arXiv preprint arXiv:2310.12836, 2023.
  83. 83.L. Luo, Y.-F. Li, G. Haffari, and S. Pan, “Reasoning on graphs: Faithful and interpretable large language model reasoning,” arXiv preprint arXiv:2310.01061, 2023.
  84. 84.X. He, Y. Tian, Y. Sun, N. V. Chawla, T. Laurent, Y. LeCun, X. Bresson, and B. Hooi, “G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,” arXiv preprint arXiv:2402.07630, 2024.
  85. 85.L. Zha, J. Zhou, L. Li, R. Wang, Q. Huang, S. Yang, J. Yuan, C. Su, X. Li, A. Su et al., “Tablegpt: Towards unifying tables, nature language and commands into one gpt,” arXiv preprint arXiv:2307.08674, 2023.
  86. 86.M. Gaur, K. Gunaratna, V. Srinivasan, and H. Jin, “Iseeq: Information seeking question generation using dynamic meta-information retrieval and knowledge graphs,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 10, 2022, pp. 10 672–10 680.
  87. 87.F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. H. Chi, N. Scharli, ¨ and D. Zhou, “Large language models can be easily distracted by irrelevant context,” in International Conference on Machine Learning. PMLR, 2023, pp. 31 210–31 227.
  88. 88.R. Teja, “Evaluating the ideal chunk size for a rag system using llamaindex,” https://www.llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5, 2023.
  89. 89.Langchain, “Recursively split by character,” https://python.langchain.com/docs/modules/data_connection/document_transformers/recursive_text_splitter, 2023.
  90. 90.S. Yang, “Advanced rag 01: Small-to-big retrieval,” https://towardsdatascience.com/advanced-rag-01-small-to-big-retrieval-172181b396d4, 2023.
  91. 91.Y. Wang, N. Lipka, R. A. Rossi, A. Siu, R. Zhang, and T. Derr, “Knowledge graph prompting for multi-document question answering,” arXiv preprint arXiv:2308.11730, 2023.
  92. 92.D. Zhou, N. Scharli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schu- ¨ urmans, C. Cui, O. Bousquet, Q. Le et al., “Least-to-most prompting enables complex reasoning in large language models,” arXiv preprint arXiv:2205.10625, 2022.
  93. 93.S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. Celikyilmaz, and J. Weston, “Chain-of-verification reduces hallucination in large language models,” arXiv preprint arXiv:2309.11495, 2023.
  94. 94.X. Li and J. Li, “Angle-optimized text embeddings,” arXiv preprint arXiv:2309.12871, 2023.
  95. 95.VoyageAI, “Voyage’s embedding models,” https://docs.voyageai.com/embeddings/, 2023.
  96. 96.BAAI, “Flagembedding,” https://github.com/FlagOpen/FlagEmbedding, 2023.
  97. 97.P. Zhang, S. Xiao, Z. Liu, Z. Dou, and J.-Y. Nie, “Retrieve anything to augment large language models,” arXiv preprint arXiv:2310.07554, 2023.
  98. 98.N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” arXiv preprint arXiv:2307.03172, 2023.
  99. 99.Y. Gao, T. Sheng, Y. Xiang, Y. Xiong, H. Wang, and J. Zhang, “Chat-rec: Towards interactive and explainable llms-augmented recommender system,” arXiv preprint arXiv:2303.14524, 2023.
  100. 100.N. Anderson, C. Wilson, and S. D. Richardson, “Lingua: Addressing scenarios for live interpretation and automatic dubbing,” in Proceedings of the 15th Biennial Conference of the Association for Machine Translation in the Americas (Volume 2: Users and Providers Track and Government Track), J. Campbell, S. Larocca, J. Marciano, K. Savenkov, and A. Yanishevsky, Eds. Orlando, USA: Association for Machine Translation in the Americas, Sep. 2022, pp. 202–209. [Online]. Available: https://aclanthology.org/2022.amta-upg.14
  101. 101.H. Jiang, Q. Wu, X. Luo, D. Li, C.-Y. Lin, Y. Yang, and L. Qiu, “Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression,” arXiv preprint arXiv:2310.06839, 2023.
  102. 102.V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, ˘ and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” arXiv preprint arXiv:2004.04906, 2020.
  103. 103.Y. Ma, Y. Cao, Y. Hong, and A. Sun, “Large language model is not a good few-shot information extractor, but a good reranker for hard samples!” ArXiv, vol. abs/2303.08559, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:257532405
  104. 104.J. Cui, Z. Li, Y. Yan, B. Chen, and L. Yuan, “Chatlaw: Open-source legal large language model with integrated external knowledge bases,” arXiv preprint arXiv:2306.16092, 2023.
  105. 105.O. Yoran, T. Wolfson, O. Ram, and J. Berant, “Making retrieval-augmented language models robust to irrelevant context,” arXiv preprint arXiv:2310.01558, 2023.
  106. 106.X. Li, R. Zhao, Y. K. Chia, B. Ding, L. Bing, S. Joty, and S. Poria, “Chain of knowledge: A framework for grounding large language models with structured knowledge bases,” arXiv preprint arXiv:2305.13269, 2023.
  107. 107.H. Yang, S. Yue, and Y. He, “Auto-gpt for online decision making: Benchmarks and additional opinions,” arXiv preprint arXiv:2306.02224, 2023.
  108. 108.T. Schick, J. Dwivedi-Yu, R. Dess`ı, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” arXiv preprint arXiv:2302.04761, 2023.
  109. 109.J. Zhang, “Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt,” arXiv preprint arXiv:2304.11116, 2023.
  110. 110.R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders et al., “Webgpt: Browser-assisted question-answering with human feedback,” arXiv preprint arXiv:2112.09332, 2021.
  111. 111.T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee et al., “Natural questions: a benchmark for question answering research,” Transactions of the Association for Computational Linguistics, vol. 7, pp. 453–466, 2019.
  112. 112.Y. Liu, S. Yavuz, R. Meng, M. Moorthy, S. Joty, C. Xiong, and Y. Zhou, “Exploring the integration strategies of retriever and large language models,” arXiv preprint arXiv:2308.12574, 2023.
  113. 113.M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, “Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,” arXiv preprint arXiv:1705.03551, 2017.
  114. 114.P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” arXiv preprint arXiv:1606.05250, 2016.
  115. 115.J. Berant, A. Chou, R. Frostig, and P. Liang, “Semantic parsing on freebase from question-answer pairs,” in Proceedings of the 2013 conference on empirical methods in natural language processing, 2013, pp. 1533–1544.
  116. 116.A. Mallen, A. Asai, V. Zhong, R. Das, H. Hajishirzi, and D. Khashabi, “When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories,” arXiv preprint arXiv:2212.10511, 2022.
  117. 117.T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng, “Ms marco: A human-generated machine reading comprehension dataset,” 2016.
  118. 118.Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning, “Hotpotqa: A dataset for diverse, explainable multi-hop question answering,” arXiv preprint arXiv:1809.09600, 2018.
  119. 119.X. Ho, A.-K. D. Nguyen, S. Sugawara, and A. Aizawa, “Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps,” arXiv preprint arXiv:2011.01060, 2020.
  120. 120.H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “Musique: Multihop questions via single-hop question composition,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 539–554, 2022.
  121. 121.A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli, “Eli5: Long form question answering,” arXiv preprint arXiv:1907.09190, 2019.
  122. 122.T. Kocisk ˇ y, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, ` and E. Grefenstette, “The narrativeqa reading comprehension challenge,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 317–328, 2018.
  123. 123.K.-H. Lee, X. Chen, H. Furuta, J. Canny, and I. Fischer, “A human-inspired reading agent with gist memory of very long contexts,” arXiv preprint arXiv:2402.09727, 2024.
  124. 124.I. Stelmakh, Y. Luan, B. Dhingra, and M.-W. Chang, “Asqa: Factoid questions meet long-form answers,” arXiv preprint arXiv:2204.06092, 2022.
  125. 125.M. Zhong, D. Yin, T. Yu, A. Zaidi, M. Mutuma, R. Jha, A. H. Awadallah, A. Celikyilmaz, Y. Liu, X. Qiu et al., “Qmsum: A new benchmark for query-based multi-domain meeting summarization,” arXiv preprint arXiv:2104.05938, 2021.
  126. 126.P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner, “A dataset of information-seeking questions and answers anchored in research papers,” arXiv preprint arXiv:2105.03011, 2021.
  127. 127.T. Moller, A. Reina, R. Jayakumar, and M. Pietsch, “Covid-qa: A ¨ question answering dataset for covid-19,” in ACL 2020 Workshop on Natural Language Processing for COVID-19 (NLP-COVID), 2020.
  128. 128.X. Wang, G. H. Chen, D. Song, Z. Zhang, Z. Chen, Q. Xiao, F. Jiang, J. Li, X. Wan, B. Wang et al., “Cmb: A comprehensive medical benchmark in chinese,” arXiv preprint arXiv:2308.08833, 2023.
  129. 129.H. Zeng, “Measuring massive multitask chinese understanding,” arXiv preprint arXiv:2304.12986, 2023.
  130. 130.R. Y. Pang, A. Parrish, N. Joshi, N. Nangia, J. Phang, A. Chen, V. Padmakumar, J. Ma, J. Thompson, H. He et al., “Quality: Question answering with long input texts, yes!” arXiv preprint arXiv:2112.08608, 2021.
  131. 131.P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try arc, the ai2 reasoning challenge,” arXiv preprint arXiv:1803.05457, 2018.
  132. 132.A. Talmor, J. Herzig, N. Lourie, and J. Berant, “Commonsenseqa: A question answering challenge targeting commonsense knowledge,” arXiv preprint arXiv:1811.00937, 2018.
  133. 133.E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston, “Wizard of wikipedia: Knowledge-powered conversational agents,” arXiv preprint arXiv:1811.01241, 2018.
  134. 134.H. Wang, M. Hu, Y. Deng, R. Wang, F. Mi, W. Wang, Y. Wang, W.-C. Kwan, I. King, and K.-F. Wong, “Large language models as source planner for personalized knowledge-grounded dialogue,” arXiv preprint arXiv:2310.08840, 2023.
  135. 135.——, “Large language models as source planner for personalized knowledge-grounded dialogue,” arXiv preprint arXiv:2310.08840, 2023.
  136. 136.X. Xu, Z. Gou, W. Wu, Z.-Y. Niu, H. Wu, H. Wang, and S. Wang, “Long time no see! open-domain conversation with long-term persona memory,” arXiv preprint arXiv:2203.05797, 2022.
  137. 137.T.-H. Wen, M. Gasic, N. Mrksic, L. M. Rojas-Barahona, P.-H. Su, S. Ultes, D. Vandyke, and S. Young, “Conditional generation and snapshot learning in neural dialogue systems,” arXiv preprint arXiv:1606.03352, 2016.
  138. 138.R. He and J. McAuley, “Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering,” in proceedings of the 25th international conference on world wide web, 2016, pp. 507–517.
  139. 139.S. Li, H. Ji, and J. Han, “Document-level event argument extraction by conditional generation,” arXiv preprint arXiv:2104.05919, 2021.
  140. 140.S. Ebner, P. Xia, R. Culkin, K. Rawlins, and B. Van Durme, “Multi-sentence argument linking,” arXiv preprint arXiv:1911.03766, 2019.
  141. 141.H. Elsahar, P. Vougiouklis, A. Remaci, C. Gravier, J. Hare, F. Laforest, and E. Simperl, “T-rex: A large scale alignment of natural language with knowledge base triples,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), 2018.
  142. 142.O. Levy, M. Seo, E. Choi, and L. Zettlemoyer, “Zero-shot relation extraction via reading comprehension,” arXiv preprint arXiv:1706.04115, 2017.
  143. 143.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” arXiv preprint arXiv:1905.07830, 2019.
  144. 144.S. Kim, S. J. Joo, D. Kim, J. Jang, S. Ye, J. Shin, and M. Seo, “The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning,” arXiv preprint arXiv:2305.14045, 2023.
  145. 145.A. Saha, V. Pahuja, M. Khapra, K. Sankaranarayanan, and S. Chandar, “Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018.
  146. 146.D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300, 2020.
  147. 147.S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” arXiv preprint arXiv:1609.07843, 2016.
  148. 148.M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant, “Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,” Transactions of the Association for Computational Linguistics, vol. 9, pp. 346–361, 2021.
  149. 149.J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, “Fever: a large-scale dataset for fact extraction and verification,” arXiv preprint arXiv:1803.05355, 2018.
  150. 150.N. Kotonya and F. Toni, “Explainable automated fact-checking for public health claims,” arXiv preprint arXiv:2010.09926, 2020.
  151. 151.R. Lebret, D. Grangier, and M. Auli, “Neural text generation from structured data with application to the biography domain,” arXiv preprint arXiv:1603.07771, 2016.
  152. 152.H. Hayashi, P. Budania, P. Wang, C. Ackerson, R. Neervannan, and G. Neubig, “Wikiasp: A dataset for multi-domain aspect-based summarization,” Transactions of the Association for Computational Linguistics, vol. 9, pp. 211–225, 2021.
  153. 153.S. Narayan, S. B. Cohen, and M. Lapata, “Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization,” arXiv preprint arXiv:1808.08745, 2018.
  154. 154.S. Saha, J. A. Junaed, M. Saleki, A. S. Sharma, M. R. Rifat, M. Rahouti, S. I. Ahmed, N. Mohammed, and M. R. Amin, “Vio-lens: A novel dataset of annotated social network posts leading to different forms of communal violence and its evaluation,” in Proceedings of the First Workshop on Bangla Language Processing (BLP-2023), 2023, pp. 72–84.
  155. 155.X. Li and D. Roth, “Learning question classifiers,” in COLING 2002: The 19th International Conference on Computational Linguistics, 2002.
  156. 156.R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proceedings of the 2013 conference on empirical methods in natural language processing, 2013, pp. 1631–1642.
  157. 157.H. Husain, H.-H. Wu, T. Gazit, M. Allamanis, and M. Brockschmidt, “Codesearchnet challenge: Evaluating the state of semantic code search,” arXiv preprint arXiv:1909.09436, 2019.
  158. 158.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano et al., “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021.
  159. 159.R. Steinberger, B. Pouliquen, A. Widiger, C. Ignat, T. Erjavec, D. Tufis, and D. Varga, “The jrc-acquis: A multilingual aligned parallel corpus with 20+ languages,” arXiv preprint cs/0609058, 2006.
  160. 160.Y. Hoshi, D. Miyashita, Y. Ng, K. Tatsuno, Y. Morioka, O. Torii, and J. Deguchi, “Ralle: A framework for developing and evaluating retrieval-augmented large language models,” arXiv preprint arXiv:2308.10633, 2023.
  161. 161.J. Liu, “Building production-ready rag applications,” https://www.ai.engineer/summit/schedule/building-production-ready-rag-applications, 2023.
  162. 162.I. Nguyen, “Evaluating rag part i: How to evaluate document retrieval,” https://www.deepset.ai/blog/rag-evaluation-retrieval, 2023.
  163. 163.Q. Leng, K. Uhlenhuth, and A. Polyzotis, “Best practices for llm evaluation of rag applications,” https://www.databricks.com/blog/LLM-auto-eval-best-practices-RAG, 2023.
  164. 164.S. Es, J. James, L. Espinosa-Anke, and S. Schockaert, “Ragas: Automated evaluation of retrieval augmented generation,” arXiv preprint arXiv:2309.15217, 2023.
  165. 165.J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, “Ares: An automated evaluation framework for retrieval-augmented generation systems,” arXiv preprint arXiv:2311.09476, 2023.
  166. 166.C. Jarvis and J. Allard, “A survey of techniques for maximizing llm performance,” https://community.openai.com/t/openai-dev-day-2023-breakout-sessions/505213#a-survey-of-techniques-for-maximizing-llm-performance-2, 2023.
  167. 167.J. Chen, H. Lin, X. Han, and L. Sun, “Benchmarking large language models in retrieval-augmented generation,” arXiv preprint arXiv:2309.01431, 2023.
  168. 168.Y. Liu, L. Huang, S. Li, S. Chen, H. Zhou, F. Meng, J. Zhou, and X. Sun, “Recall: A benchmark for llms robustness against external counterfactual knowledge,” arXiv preprint arXiv:2311.08147, 2023.
  169. 169.Y. Lyu, Z. Li, S. Niu, F. Xiong, B. Tang, W. Wang, H. Wu, H. Liu, T. Xu, and E. Chen, “Crud-rag: A comprehensive chinese benchmark for retrieval-augmented generation of large language models,” arXiv preprint arXiv:2401.17043, 2024.
  170. 170.P. Xu, W. Ping, X. Wu, L. McAfee, C. Zhu, Z. Liu, S. Subramanian, E. Bakhturina, M. Shoeybi, and B. Catanzaro, “Retrieval meets long context large language models,” arXiv preprint arXiv:2310.03025, 2023.
  171. 171.C. Packer, V. Fang, S. G. Patil, K. Lin, S. Wooders, and J. E. Gonzalez, “Memgpt: Towards llms as operating systems,” arXiv preprint arXiv:2310.08560, 2023.
  172. 172.G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” arXiv preprint arXiv:2309.17453, 2023.
  173. 173.T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, and J. E. Gonzalez, “Raft: Adapting language model to domain specific rag,” arXiv preprint arXiv:2403.10131, 2024.
  174. 174.J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361, 2020.
  175. 175.U. Alon, F. Xu, J. He, S. Sengupta, D. Roth, and G. Neubig, “Neurosymbolic language modeling with automaton-augmented retrieval,” in International Conference on Machine Learning. PMLR, 2022, pp. 468–485.
  176. 176.M. Yasunaga, A. Aghajanyan, W. Shi, R. James, J. Leskovec, P. Liang, M. Lewis, L. Zettlemoyer, and W.-t. Yih, “Retrieval-augmented multimodal language modeling,” arXiv preprint arXiv:2211.12561, 2022.
  177. 177.J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping languageimage pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597, 2023.
  178. 178.W. Zhu, A. Yan, Y. Lu, W. Xu, X. E. Wang, M. Eckstein, and W. Y. Wang, “Visualize before you write: Imagination-guided open-ended text generation,” arXiv preprint arXiv:2210.03765, 2022.
  179. 179.J. Zhao, G. Haffar, and E. Shareghi, “Generating synthetic speech from spokenvocab for speech translation,” arXiv preprint arXiv:2210.08174, 2022.
  180. 180.D. M. Chan, S. Ghosh, A. Rastrow, and B. Hoffmeister, “Using external off-policy speech-to-text mappings in contextual end-to-end automated speech recognition,” arXiv preprint arXiv:2301.02736, 2023.
  181. 181.A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, “Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 714–10 726.
  182. 182.N. Nashid, M. Sintaha, and A. Mesbah, “Retrieval-based prompt selection for code-related few-shot learning,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 2450–2462.

Citation

MLA
Gao, Y., et al. “Retrieval-Augmented Generation for Large Language Models: A Survey”. arXiv, 2023, http://arxiv.org/abs/2312.10997v5.
APA
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv. http://arxiv.org/abs/2312.10997v5
Chicago
Gao, Y., Y. Xiong, X. Gao, et al. 2023. “Retrieval-Augmented Generation for Large Language Models: A Survey”. arXiv. http://arxiv.org/abs/2312.10997v5.
Harvard
Gao, Y. et al. (2023) “Retrieval-Augmented Generation for Large Language Models: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10997v5.
Vancouver
1. Gao Y, Xiong Y, Gao X, Jia K, Pan J, Bi Y, Dai Y, Sun J, Wang M, Wang H (2023) Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv

BibTeX

@article{gao2023retrieval,
  title = {Retrieval-Augmented Generation for Large Language Models: A Survey},
  author = {Gao, Yunfan and Xiong, Yun and Gao, Xinyu and Jia, Kangxiang and Pan, Jinliu and Bi, Yuxi and Dai, Yi and Sun, Jiawei and Wang, Meng and Wang, Haofen},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10997v5},
  eprint = {2312.10997}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors