DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models

Weihang SuYichen TangQingyao AiZhijing WuYiqun Liu

article2024ACL100 citations

Proposes a training-free dynamic retrieval-augmented generation framework that evaluates token uncertainty and self-attention weights to dynamically decide when to retrieve external knowledge and how to formulate queries across the full context.

Listen

Large language models often generate plausible yet factually incorrect text, commonly referred to as hallucination. To address this risk in knowledge-intensive and multi-step tasks, retrieval-augmented generation connects models to external data sources. However, conventional dynamic retrieval systems rely on rigid rules, such as retrieving text at fixed token intervals or after every sentence. These approaches either retrieve unnecessarily—introducing distracting noise and inflating computational costs—or restrict search queries to only the most recent tokens, missing broader contextual needs.

The article introduces and evaluates DRAGIN (Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models), a lightweight framework designed to dynamically determine both when to retrieve external data and what to search for during text generation without requiring model retraining or prompt engineering.

To decide the optimal timing for retrieval, DRAGIN uses a mechanism that evaluates token uncertainty, semantic importance, and the influence of a given token on subsequent words via attention scores. To determine search content, the framework extracts key tokens across the full preceding context based on internal self-attention distributions. The authors evaluated DRAGIN across four knowledge-intensive benchmarks (2WikiMultihopQA, HotpotQA, StrategyQA, and IIRC) spanning multi-hop question answering, reading comprehension, and commonsense reasoning, using three open-source models: LLaMA-2-Chat-7B, LLaMA-2-Chat-13B, and Vicuna-13B-v1.5.

The evaluation revealed several key findings:

  1. DRAGIN consistently achieved state-of-the-art performance across all four benchmarks, outperforming standard generation and existing dynamic retrieval methods.
  2. Performance gains were largest in complex multi-step reasoning tasks, where exact match accuracy and answer quality improved markedly over baseline systems.
  3. DRAGIN maintained superior retrieval efficiency, triggering external searches significantly less often than fixed-interval or fixed-sentence baselines while preserving higher accuracy.
  4. Standard lexical search (BM25) consistently outperformed dense neural retrieval methods when paired with the framework, offering higher accuracy with lower operational complexity.

These findings indicate that aligning external data retrieval directly with a model's real-time internal information needs significantly enhances factual accuracy while controlling computational overhead and query latency. Moreover, the strong performance of simpler lexical search engines suggests organizations can achieve high retrieval-augmented performance without investing in complex dense retrieval infrastructure.

For practical implementation, technical teams deploying open-source models should consider dynamic, attention-based retrieval triggers to improve generation accuracy and mitigate hallucination risk. Organizations should tune the activation threshold to balance execution speed and accuracy according to specific application requirements.

A primary limitation of DRAGIN is its dependency on internal self-attention scores from transformer architectures. As a result, the framework is currently compatible only with open-source or locally hosted models, and cannot be applied directly to commercial black-box application programming interfaces that conceal internal attention states. Additionally, for smaller models that struggle with long-context comprehension, incorporating retrieved documents can occasionally cause distraction, indicating that future work should focus on extended context management.

arXiv: 2403.10081
Cover for DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models

Abstract

Dynamic retrieval augmented generation (RAG) paradigm actively decides when and what to retrieve during the text generation process of Large Language Models (LLMs). There are two key elements of this paradigm: identifying the optimal moment to activate the retrieval module (deciding when to retrieve) and crafting the appropriate query once retrieval is triggered (determining what to retrieve). However, current dynamic RAG methods fall short in both aspects. Firstly, the strategies for deciding when to retrieve often rely on static rules. Moreover, the strategies for deciding what to retrieve typically limit themselves to the LLM’s most recent sentence or the last few tokens, while the LLM’s information needs may span across the entire context. To overcome these limitations, we introduce a new framework, DRAGIN, i.e., Dynamic Retrieval Augmented Generation based on the Information Needs of LLMs. Our framework is specifically designed to make decisions on when and what to retrieve based on the LLM’s information needs during the text generation process. We evaluate DRAGIN along with existing methods comprehensively over 4 knowledge-intensive generation datasets. Experimental results show that DRAGIN achieves superior performance on all tasks, demonstrating the effectiveness of our method1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Single-round Retrieval-augmented LLM
  • 2.2 Multi-round Retrieval-augmented LLM
  • 3 Methodology
  • 3.1 Real-time Information Need Detection
  • 3.2 Query Formulation based on Self-attention
  • 3.3 Continue Generation after Retrieval
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Settings for each Dataset
  • 4.3 Baselines
  • 4.4 Selected LLMs
  • 4.5 Implementation Details
  • 5 Experimental Results
  • 5.1 Overall Results of DRAGIN and Baselines
  • 5.2 Efficiency
  • 5.3 Timing of Retrieval
  • 5.4 Query Formulation
  • 5.5 Impact of Retriever
  • 6 Conclusions and Future Works
  • 7 Limitations
  • 8 Ethics Statement
  • Acknowledgments
  • References
  • A Datasets and Settings
  • B Evaluation Details
  • C Case Study
  • D Error Analysis
  • E Hyperparameters
  • F Prompt Template

Knowls

  1. Knowl 1 — DRAGIN dynamically aligns retrieval with the language model’s information needs

    model/method

    DRAGIN is a dynamic retrieval-augmented generation framework that makes both retrieval decisions during generation: when to retrieve and what to retrieve. It combines two components:

    • Real-time Information Need Detection (RIND) evaluates generated tokens using uncertainty, attention-based influence on subsequent context, and a semantic stopword filter. Retrieval is activated only when a token-level information-need score exceeds a threshold.
    • Query Formulation based on Self-attention (QFS) uses the triggering token’s self-attention over the entire preceding context to select query terms, rather than using only the latest sentence or a fixed window of tokens.

    After retrieval, the language model continues generation using the retrieved passages as external knowledge. DRAGIN requires no additional model training or fine-tuning and is designed for Transformer-based language models whose self-attention scores are accessible.

  2. Knowl 2 — RIND scores retrieval need from uncertainty, contextual influence, and semantic content

    equation

    For a generated token sequence T={t1,t2,…,tn}T=\{t_1,t_2,\ldots,t_n\}, let VV be the language model’s vocabulary and let pi(v)p_i(v) be the probability of generating vocabulary token v∈Vv\in V at position ii. RIND measures token uncertainty with the entropy

    Hi=−∑v∈Vpi(v)log⁡pi(v).H_i=-\sum_{v\in V}p_i(v)\log p_i(v).

    Let Ai,jA_{i,j} be the last-layer Transformer self-attention value between positions ii and jj, with QiQ_i the query vector at position ii, KjK_j the key vector at position jj, and dkd_k the key-vector dimensionality. For i<ji<j, the attention value is computed as

    Ai,j=softmax⁡j(QiKjTdk),amax⁡(i)=max⁡j>iAi,j.A_{i,j}=\operatorname{softmax}_{j}\left(\frac{Q_iK_j^{\mathsf T}}{\sqrt{d_k}}\right), \qquad a_{\max}(i)=\max_{j>i}A_{i,j}.

    Let SS be the stopword set. The semantic indicator is

    si={0,ti∈S,1,ti∉S.s_i=\begin{cases} 0,&t_i\in S,\\ 1,&t_i\notin S. \end{cases}

    RIND combines the three signals into the token score

    SRIND(ti)=Hi amax⁡(i) si.S_{\mathrm{RIND}}(t_i)=H_i\,a_{\max}(i)\,s_i.

    Given a retrieval threshold θ\theta, the retrieval module is activated when at least one already-generated token satisfies SRIND(ti)>θS_{\mathrm{RIND}}(t_i)>\theta. Stopwords therefore cannot independently trigger retrieval, while uncertain, semantically meaningful tokens with strong attention influence are more likely to do so.

  3. Knowl 3 — QFS constructs queries from the most attended tokens in the full context

    algorithm

    QFS formulates a retrieval query when RIND identifies a trigger position ii. It uses the last Transformer layer’s attention distribution from the triggered position to the preceding tokens {t1,…,ti−1}\{t_1,\ldots,t_{i-1}\}.

    Input: Generated token sequence T, trigger position i, and number n of query tokens
    Output: Retrieval query Q_i
    Extract the attention scores A_i = {a_i,1, ..., a_i,i-1} from the last Transformer layer.
    Rank the preceding token positions by descending attention score.
    Select the n positions with the largest attention scores.
    Map the selected token pieces to their corresponding words.
    Restore the selected words to their original left-to-right order in T.
    Concatenate the restored words to form Q_i.
    Return Q_i.

    The resulting query is sparse but context-wide: it preserves the words that the language model considered most important for generating the trigger position, even when those words occur far earlier than the latest sentence. The value of nn is a dataset- and language-model-specific hyperparameter.

  4. Knowl 4 — Retrieved passages replace the interrupted generation context and support continued decoding

    model/method

    When RIND identifies a trigger token tit_i, QFS produces a query and an external retriever returns passages Di1,Di2,Di3D_{i1},D_{i2},D_{i3}. The language-model output is truncated immediately before the trigger position, producing a prefix T′=truncate⁡(T,ti)T'=\operatorname{truncate}(T,t_i). DRAGIN then prompts the language model with the retrieved passages, the original question, and the truncated answer prefix in the following structure:

    Below are the external knowledge references:

    [1] Di1D_{i1}
    [2] Di2D_{i2}
    [3] Di3D_{i3}

    Please answer the question based on the external knowledge:
    Question: [original question]
    Answer: T′T'

    The language model resumes generation from the truncation point using the retrieved knowledge. If RIND later detects another information need at position jj, QFS creates a new query, retrieves Dj1,Dj2,Dj3D_{j1},D_{j2},D_{j3}, replaces the previous passages, and repeats the same continuation procedure.

  5. Knowl 5 — Experimental evaluation compares DRAGIN with fixed, single-round, and uncertainty-triggered retrieval

    experimental setup

    DRAGIN was evaluated on four knowledge-intensive benchmarks: 2WikiMultihopQA and HotpotQA for multi-hop question answering, IIRC for incomplete-information reading comprehension, and StrategyQA for commonsense reasoning. The datasets contain 1,000, 1,000, 954, and 1,000 examples, respectively. Exact match, token-level F1, and precision were used for the first three datasets; StrategyQA was evaluated with accuracy.

    Experiments used Llama-2-Chat-7B, Llama-2-Chat-13B, and Vicuna-13B-v1.5. Wikipedia was segmented into 100-token passages, BM25 retrieved the top 3 passages, and generation used greedy decoding. The prompts used 6 demonstrations for 2WikiMultihopQA and 8 demonstrations for each of HotpotQA, IIRC, and StrategyQA.

    The baselines were: direct generation without retrieval (wo-RAG); single-round retrieval from the initial question (SR-RAG); retrieval every fixed number of tokens using the preceding token window (FL-RAG); retrieval after every generated sentence using the preceding sentence (FS-RAG); and FLARE, which retrieves when a generated token falls below a probability threshold and queries with the latest sentence after removing low-probability tokens.

    The main DRAGIN hyperparameters were:

    LLM Hyperparameter 2WikiMultihopQA HotpotQA IIRC StrategyQA
    Llama2-13b-chat generate length 64 100 128 100
    θ\theta 0.6 1.2 1.25 1.0
    top nn tokens 25 35 25 25
    Llama2-7b-chat generate length 64 100 128 100
    θ\theta 1.0 1.3 1.3 0.75
    top nn tokens 25 35 35 35
    Vicuna-13b-v1.5 generate length 64 100 128 100
    θ\theta 1.2 1.2 1.3 1.5
    top nn tokens 25 35 35 25
  6. Knowl 6 — DRAGIN improves knowledge-intensive generation across models and benchmarks

    data/table

    The main comparison measures exact match (EM), F1, and StrategyQA accuracy. DRAGIN is the strongest retrieval-augmented method on nearly every model–dataset combination and improves especially strongly on the multi-hop datasets. It does not surpass direct generation on StrategyQA with Llama2-7B, where direct generation reaches accuracy 0.6590.659 and DRAGIN reaches 0.6410.641.

    LLM RAG method 2Wiki EM 2Wiki F1 Hotpot EM Hotpot F1 Strategy Acc. IIRC EM / F1
    Llama2-13b-chat wo-RAG 0.187 0.2721 0.223 0.3097 0.650 0.168 / 0.2039
    SR-RAG 0.245 0.3364 0.263 0.3706 0.654 0.196 / 0.2303
    FL-RAG 0.217 0.3054 0.177 0.2682 0.648 0.155 / 0.1875
    FS-RAG 0.270 0.3610 0.267 0.3715 0.655 0.171 / 0.2061
    FLARE 0.224 0.3076 0.180 0.2756 0.655 0.138 / 0.1667
    DRAGIN 0.304 0.3931 0.314 0.4238 0.689 0.185 / 0.2221
    Llama2-7b-chat wo-RAG 0.146 0.2232 0.184 0.2745 0.659 0.139 / 0.1731
    SR-RAG 0.169 0.2549 0.164 0.2499 0.645 0.187 / 0.2258
    FL-RAG 0.112 0.1922 0.146 0.2107 0.635 0.172 / 0.2023
    FS-RAG 0.189 0.2652 0.214 0.3035 0.629 0.178 / 0.2157
    FLARE 0.143 0.2134 0.149 0.2208 0.627 0.136 / 0.1644
    DRAGIN 0.220 0.2926 0.232 0.3344 0.641 0.192 / 0.2336
    Vicuna-13b-v1.5 wo-RAG 0.146 0.2232 0.228 0.3256 0.682 0.175 / 0.2149
    SR-RAG 0.170 0.2564 0.254 0.3531 0.686 0.217 / 0.2564
    FL-RAG 0.135 0.2133 0.187 0.3039 0.645 0.0985 / 0.1285
    FS-RAG 0.188 0.2625 0.185 0.3216 0.622 0.1027 / 0.1344
    FLARE 0.157 0.2257 0.092 0.1808 0.599 0.1174 / 0.1469
    DRAGIN 0.252 0.3516 0.288 0.4164 0.687 0.2233 / 0.2652

    For example, with Llama2-13B, DRAGIN raises HotpotQA EM from 0.2230.223 without retrieval and 0.2630.263 with single-round retrieval to 0.3140.314, while raising StrategyQA accuracy from 0.6500.650 to 0.6890.689. The results show that retrieval timing and query construction jointly matter; adding retrieval according to fixed schedules can be worse than single-round retrieval.

  7. Knowl 7 — Controlled ablations show that both RIND timing and QFS query construction are necessary

    data/table

    Two controlled experiments isolate the two major DRAGIN decisions. For the timing experiment, every method uses the same query—the last complete generated sentence—and only the retrieval trigger differs. On IIRC, DRAGIN outperforms FLARE, FL-RAG, and FS-RAG for both Llama2-13B and Vicuna-13B.

    LLM Timing method EM F1 Prec.
    Llama2-13B FLARE 0.128 0.1599 0.1677
    FL-RAG 0.155 0.1875 0.1986
    FS-RAG 0.171 0.2061 0.2185
    DRAGIN 0.187 0.2242 0.2319
    Vicuna-13B FLARE 0.097 0.1277 0.1324
    FL-RAG 0.099 0.1285 0.1324
    FS-RAG 0.103 0.1344 0.1358
    DRAGIN 0.196 0.2367 0.2476

    For the query experiment, all methods use RIND as the trigger and differ only in query formulation. QFS is compared with FLARE’s sentence with low-probability tokens removed, FS-RAG’s preceding sentence, FL-RAG’s nearest 25 tokens, and the full context. On HotpotQA, QFS is best for both language models.

    LLM Query method EM F1 Prec.
    Llama2-13B FLARE 0.262 0.3674 0.3792
    Full Context 0.252 0.3584 0.3711
    FS-RAG 0.255 0.3574 0.3685
    FL-RAG 0.241 0.3394 0.3495
    DRAGIN/QFS 0.314 0.4238 0.4401
    Vicuna-13B FLARE 0.225 0.3366 0.3420
    Full Context 0.221 0.3402 0.3457
    FS-RAG 0.216 0.3432 0.3507
    FL-RAG 0.214 0.3268 0.3264
    DRAGIN/QFS 0.288 0.4164 0.4226

    Using the entire context indiscriminately is inferior to QFS, supporting the claim that selecting attention-important terms removes redundant context while preserving the model’s immediate information need.

  8. Knowl 8 — DRAGIN uses fewer retrieval calls than fixed-schedule baselines while remaining stable across thresholds

    data/table

    Retrieval-call counts were measured as averages over the four datasets. FLARE makes the fewest calls, while DRAGIN generally makes fewer calls than FS-RAG and FL-RAG and still achieves higher answer quality in the main evaluation.

    LLM Method 2WikiMultihopQA HotpotQA StrategyQA IIRC
    L13B FL-RAG 3.770 3.194 3.626 3.426
    FS-RAG 3.131 4.583 4.885 4.305
    FLARE 1.592 3.378 0.625 5.521
    DRAGIN 2.631 3.505 4.786 2.829
    L7B FL-RAG 3.342 3.809 3.757 2.839
    FS-RAG 3.833 4.152 4.546 4.210
    FLARE 0.941 1.064 1.271 1.095
    DRAGIN 2.836 3.013 4.629 2.927
    V13B FL-RAG 4.199 3.564 3.591 3.189
    FS-RAG 3.720 5.701 6.820 6.032
    FLARE 1.093 1.078 1.118 0.335
    DRAGIN 2.542 3.184 3.744 3.120

    For Llama2-13B on HotpotQA, varying the RIND threshold from 0.30.3 to 1.01.0 changes EM only from 0.2950.295 to 0.2930.293 and keeps F1 between 0.38560.3856 and 0.39440.3944. The reported threshold sweep is:

    θ\theta EM F1 Prec.
    0.3 0.295 0.3856 0.3873
    0.4 0.297 0.387 0.389
    0.5 0.299 0.3897 0.3915
    0.6 0.304 0.3931 0.3946
    0.7 0.304 0.3927 0.3937
    0.8 0.301 0.392 0.3927
    0.9 0.301 0.3944 0.3947
    1 0.293 0.3869 0.3875

    Increasing θ\theta reduces activation frequency, providing a practical efficiency–accuracy control, while the answer metrics remain comparatively insensitive over the tested range.

  9. Knowl 9 — BM25 outperforms SGPT as the retriever within DRAGIN’s generation loop

    data/table

    The paper compares lexical BM25 with the dense retriever SGPT while holding the Llama2-13B DRAGIN configuration fixed. BM25 performs better on every evaluated dataset, suggesting that lexical matching is a strong retrieval choice for the dynamically generated QFS queries.

    Dataset Retriever EM F1 Prec.
    2WikiMultihopQA BM25 0.304 0.393 0.395
    SGPT 0.273 0.356 0.357
    HotpotQA BM25 0.314 0.424 0.437
    SGPT 0.264 0.371 0.388
    IIRC BM25 0.185 0.222 0.235
    SGPT 0.169 0.201 0.207

    The largest difference occurs on HotpotQA, where BM25 reaches EM 0.3140.314 and F1 0.4240.424, compared with SGPT’s EM 0.2640.264 and F1 0.3710.371.

  10. Knowl 10 — DRAGIN depends on accessible self-attention and can be weakened by long retrieved contexts

    limitation

    Both RIND and QFS require access to Transformer self-attention scores. Consequently, DRAGIN cannot be directly applied to language-model APIs that expose generated text but do not expose internal attention values.

    The paper also reports a long-context failure mode. RIND can retrieve multiple passages that extend the language model’s context, and models that handle long contexts poorly may confuse or mix information from those passages. In an analyzed case, Llama2-7B received a relevant passage among three retrieved passages but still failed to produce the correct answer. The authors identify improved handling of extended retrieved contexts and attention-free alternatives as future work.

Coverage note — The illustrative Einstein and arena case studies, prompt-template exemplars, and full error-analysis example were omitted because they demonstrate the method qualitatively but do not add independent methodological or quantitative contributions.

References

  1. 1.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206–2240. PMLR.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Jia Chen, Haitao Li, Weihang Su, Qingyao Ai, and Yiqun Liu. 2023. Thuir at wsdm cup 2023 task 1: Unbiased learning to rank. arXiv preprint arXiv:2304.12650.
  4. 4.Xuesong Chen, Ziyi Ye, Xiaohui Xie, Yiqun Liu, Xiaorong Gao, Weihang Su, Shuqi Zhu, Yike Sun, Min Zhang, and Shaoping Ma. 2022. Web search via an efficient and effective brain-machine interface. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, pages 1569–1572.
  5. 5.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna.lmsys.org (accessed 14 April 2023).
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Yan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao, Weihang Su, Jia Chen, and Yiqun Liu. 2024. Scaling laws for dense retrieval. arXiv preprint arXiv:2403.18684.
  8. 8.James Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot, and Pradeep Dasigi. 2020. Iirc: A dataset of incomplete information reading comprehension questions. arXiv preprint arXiv:2011.07127.
  9. 9.Luyu Gao and Jamie Callan. 2021. Condenser: a pre-training architecture for dense retrieval. arXiv preprint arXiv:2104.08253.
  10. 10.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  11. 11.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  12. 12.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. arXiv preprint arXiv:2011.01060.
  13. 13.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282.
  14. 14.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  15. 15.Zhengbao Jiang, Luyu Gao, Jun Araki, Haibo Ding, Zhiruo Wang, Jamie Callan, and Graham Neubig. 2022. Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer. arXiv preprint arXiv:2212.02027.
  16. 16.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. arXiv preprint arXiv:2305.06983.
  17. 17.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019. Generalization through memorization: Nearest neighbor language models. arXiv preprint arXiv:1911.00172.
  18. 18.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  19. 19.Haitao Li, Jia Chen, Weihang Su, Qingyao Ai, and Yiqun Liu. 2023a. Towards better web search performance: Pre-training, fine-tuning and learning to rank. arXiv preprint arXiv:2303.04710.
  20. 20.Haitao Li, Weihang Su, Changyue Wang, Yueyue Wu, Qingyao Ai, and Yiqun Liu. 2023b. Thuir@ coliee 2023: Incorporating structural knowledge into pre-trained language models for legal case retrieval. arXiv preprint arXiv:2305.06812.
  21. 21.Haitao Li, Changyue Wang, Weihang Su, Yueyue Wu, Qingyao Ai, and Yiqun Liu. 2023c. Thuir@ coliee 2023: More parameters and legal knowledge for legal case entailment. arXiv preprint arXiv:2305.06817.
  22. 22.Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jingyuan Wang, Jian-Yun Nie, and Ji-Rong Wen. 2023d. The web can be your oyster for improving large language models. arXiv preprint arXiv:2305.10998.
  23. 23.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2021. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv preprint arXiv:2104.08704.
  24. 24.Yixiao Ma, Yueyue Wu, Weihang Su, Qingyao Ai, and Yiqun Liu. 2023. Caseencoder: A knowledge-enhanced pre-trained model for legal case encoding. arXiv preprint arXiv:2305.05393.
  25. 25.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. arXiv preprint arXiv:2005.00661.
  26. 26.Niklas Muennighoff. 2022. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904.
  27. 27.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. arXiv preprint arXiv:2302.00083.
  28. 28.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  29. 29.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman ´ Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  30. 30.Hemlata Shelar, Gagandeep Kaur, Neha Heda, and Poorva Agrawal. 2020. Named entity recognition approaches and their comparison for custom ner model. Science & Technology Libraries, 39(3):324–337.
  31. 31.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. Replug: Retrieval-augmented black-box language models. arXiv preprint arXiv:2301.12652.
  32. 32.Weihang Su, Qingyao Ai, Xiangsheng Li, Jia Chen, Yiqun Liu, Xiaolong Wu, and Shengluan Hou. 2023a. Wikiformer: Pre-training with structured information of wikipedia for ad-hoc retrieval. arXiv preprint arXiv:2312.10661.
  33. 33.Weihang Su, Qingyao Ai, Yueyue Wu, Yixiao Ma, Haitao Li, and Yiqun Liu. 2023b. Caseformer: Pre-training for legal case retrieval. arXiv preprint arXiv:2311.00333.
  34. 34.Weihang Su, Xiangsheng Li, Yiqun Liu, Min Zhang, and Shaoping Ma. 2023c. Thuir2 at ntcir-16 session search (ss) task. arXiv preprint arXiv:2307.00250.
  35. 35.Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, and Yiqun Liu. 2024. Unsupervised real-time hallucination detection based on the internal states of large language models. arXiv preprint arXiv:2403.06448.
  36. 36.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  37. 37.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  38. 38.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509.
  39. 39.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  40. 40.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  42. 42.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600.
  43. 43.Ziyi Ye, Xiaohui Xie, Qingyao Ai, Yiqun Liu, Zhihong Wang, Weihang Su, and Min Zhang. 2024. Relevance feedback with brain signals. ACM Transactions on Information Systems, 42(4):1–37.
  44. 44.ChengXiang Zhai. 2008. Statistical language models for information retrieval. Synthesis lectures on human language technologies, 1(1):1–141.
  45. 45.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  46. 46.Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Paco Guzman, Luke Zettlemoyer, and Marjan Ghazvininejad. 2020. Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593.

Citation

MLA
Su, W., et al. “DRAGIN: Dynamic Retrieval Augmented Generation Based on the Real-time Information Needs of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 12991–3013, https://doi.org/10.18653/v1/2024.acl-long.702.
APA
Su, W., Tang, Y., Ai, Q., Wu, Z., & Liu, Y. (2024). DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12991–13013. https://doi.org/10.18653/v1/2024.acl-long.702
Chicago
Su, W., Y. Tang, Q. Ai, Z. Wu, and Y. Liu. 2024. “DRAGIN: Dynamic Retrieval Augmented Generation Based on the Real-time Information Needs of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12991–13013. https://doi.org/10.18653/v1/2024.acl-long.702.
Harvard
Su, W. et al. (2024) “DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12991–13013. Available at: https://doi.org/10.18653/v1/2024.acl-long.702.
Vancouver
1. Su W, Tang Y, Ai Q, Wu Z, Liu Y (2024) DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12991–13013

BibTeX

@inproceedings{su-etal-2024-dragin,
    title = "{DRAGIN}: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models",
    author = "Su, Weihang  and
      Tang, Yichen  and
      Ai, Qingyao  and
      Wu, Zhijing  and
      Liu, Yiqun",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.702/",
    doi = "10.18653/v1/2024.acl-long.702",
    pages = "12991--13013"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/