Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity

Soyeong JeongJinheon BaekSukmin ChoSung Ju HwangJong Park

article2024NAACL704 citations

Proposes a dynamic framework that uses a lightweight classifier to route queries across no-retrieval, single-step, and multi-step retrieval strategies based on question complexity, significantly reducing computational overhead while improving question-answering accuracy.

Listen

Large Language Models often struggle with factual errors and outdated information. While Retrieval-Augmented Generation addresses this by pulling external knowledge from document repositories, current implementations use a rigid strategy: either a fast single retrieval step that fails on complex, multi-hop questions, or a resource-intensive iterative retrieval process that wastes computational power on simple questions. Because real-world user queries vary significantly in difficulty, existing one-size-fits-all and simplistic binary retrieval systems create substantial computational overhead and suboptimal response accuracy.

The article introduces and evaluates Adaptive-RAG, an adaptive framework designed to dynamically route incoming questions to the most suitable processing strategy based on predicted question complexity. The objective is to balance accuracy and operational efficiency across simple, moderate, and highly complex questions without altering the underlying language model architectures.

The framework relies on a small, dedicated classifier model trained to categorize incoming questions into three complexity tiers: straightforward questions answered using internal model memory alone (no retrieval), moderate questions answered via a single document retrieval step, and complex questions handled through iterative, multi-step retrieval and reasoning. The classifier is trained without manual human annotation by automatically labeling data based on baseline model success rates and inherent structural biases in benchmark datasets. The researchers evaluated the system across six standard open-domain question answering benchmarks (comprising both single-hop and multi-hop datasets) using multiple large language models, including GPT-3.5 and the FLAN-T5 series.

The evaluation yielded three primary findings. First, Adaptive-RAG achieved superior overall accuracy compared to standard single-step and existing adaptive retrieval baselines, reaching an F1 score of 50.91 and an exact match rate of 37.97 on GPT-3.5 across all benchmarks. Second, the system substantially improved operational efficiency compared to fully iterative multi-step systems; on GPT-3.5, it reduced the average retrieval-and-generation steps from 2.81 to 1.03 per query and cut processing time by more than 50%. Third, tests with varying classifier sizes demonstrated that even small classifier models (around 60 million parameters) deliver performance comparable to larger classifiers, minimizing system overhead.

These findings indicate that routing queries adaptively offers a practical path to scaling enterprise question answering systems. By reserving expensive multi-step reasoning for genuinely complex inquiries and answering straightforward queries with minimal resources, organizations can significantly lower API expenses and server compute costs while maintaining high response reliability.

Decision-makers should consider adopting dynamic routing frameworks like Adaptive-RAG when deploying retrieval-augmented systems in production. To operationalize this approach, engineering teams should implement automated classifier training pipelines using historical query performance and conduct pilots to validate latency and cost savings on production traffic. Organizations must also pair these systems with content moderation layers to filter potentially harmful user queries and retrieved text.

The main limitation is that classifier accuracy remains below an ideal theoretical ceiling, showing moderate misclassification between adjacent complexity tiers (such as classifying multi-step queries as single-step queries about 31% of the time). Because training labels rely on automated heuristics rather than expert human annotations, there is residual uncertainty in classifier precision. Nonetheless, confidence in the demonstrated trade-off benefits is high across diverse model scales and question types.

Cover for Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity

Abstract

Retrieval-Augmented Large Language Models (LLMs), which incorporate the non-parametric knowledge from external knowledge bases into LLMs, have emerged as a promising approach to enhancing response accuracy in several tasks, such as Question-Answering (QA). However, even though there are various approaches dealing with queries of different complexities, they either handle simple queries with unnecessary computational overhead or fail to adequately address complex multi-step queries; yet, not all user requests fall into only one of the simple or complex categories. In this work, we propose a novel adaptive QA framework that can dynamically select the most suitable strategy for (retrieval-augmented) LLMs from the simplest to the most sophisticated ones based on the query complexity. Also, this selection process is operationalized with a classifier, which is a smaller LM trained to predict the complexity level of incoming queries with automatically collected labels, obtained from actual predicted outcomes of models and inherent inductive biases in datasets. This approach offers a balanced strategy, seamlessly adapting between the iterative and single-step retrieval-augmented LLMs, as well as the no-retrieval methods, in response to a range of query complexities. We validate our model on a set of open-domain QA datasets, covering multiple query complexities, and show that ours enhances the overall efficiency and accuracy of QA systems, compared to relevant baselines including the adaptive retrieval approaches. Code is available at: https://github.com/starsuzi/Adaptive-RAG.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminaries
  • 3.2 Adaptive-RAG: Adaptive Retrieval-Augmented Generation
  • 4 Experimental Setups
  • 4.1 Datasets
  • 4.2 Models
  • 4.3 Evaluation Metrics
  • 4.4 Implementation Details
  • 5 Experimental Results and Analyses
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Additional Experimental Setups
  • A.1 Datasets
  • A.2 Models
  • A.3 Implementation Details
  • B Additional Experimental Results

Knowls

  1. Knowl 1 — Adaptive Retrieval-Augmented Generation Framework (Adaptive-RAG)

    model/method

    Adaptive Retrieval-Augmented Generation (Adaptive-RAG) is a question-answering framework that dynamically adjusts its operational strategy according to the intrinsic complexity of each input query qq. Instead of applying a uniform retrieval strategy to all queries, the framework routes incoming queries across three execution tiers:

    1. Non-Retrieval Strategy (Class AA): Intended for straightforward queries answerable directly via the internal parametric memory of the large language model (LLM): aˉ=LLM(q)\bar{a} = \text{LLM}(q) where qq represents the user input query tokens and aˉ\bar{a} is the generated answer sequence.

    2. Single-Step Retrieval Strategy (Class BB): Intended for queries of moderate complexity requiring factual grounding from an external knowledge corpus D\mathcal{D}. A document d∈Dd \in \mathcal{D} is retrieved once and passed along with qq to the LLM: d=Retriever(q;D)d = \text{Retriever}(q; \mathcal{D}) aˉ=LLM(q,d)\bar{a} = \text{LLM}(q, d)

    3. Multi-Step Retrieval Strategy (Class CC): Intended for complex, multi-hop queries that require synthesizing facts from multiple documents through iterative reasoning. At reasoning step ii, new documents di∈Dd_i \in \mathcal{D} are retrieved based on the query and preceding context ci=(d1,…,di−1,aˉ1,…,aˉi−1)c_i = (d_1, \dots, d_{i-1}, \bar{a}_1, \dots, \bar{a}_{i-1}), and intermediate reasoning outcomes aˉi\bar{a}_i are produced: di=Retriever(q,ci;D)d_i = \text{Retriever}(q, c_i; \mathcal{D}) aˉi=LLM(q,di,ci)\bar{a}_i = \text{LLM}(q, d_i, c_i) This process iterates until a termination criterion is met or a maximum step limit is reached.

    The routing decision is made prior to generation by a lightweight classifier LM, o=Classifier(q)∈{A,B,C}o = \text{Classifier}(q) \in \{A, B, C\}, enabling the framework to balance computational cost and answer correctness across heterogeneous query distributions without modifying internal LLM parameters.

  2. Knowl 2 — Automatic Training Data Construction for Query Complexity Classification

    model/method

    Because manual labels for query complexity levels {A,B,C}\{A, B, C\} are unavailable, the query complexity classifier is trained using a two-stage automatic labeling strategy that combines model performance outcomes with dataset inductive biases:

    1. Silver Data Generation via Empirical Execution: Each training query qq is independently processed by the non-retrieval model LLM(q)\text{LLM}(q), the single-step model LLM(q,d)\text{LLM}(q, d), and the multi-step model LLM(q,d,c)\text{LLM}(q, d, c).
    • If the non-retrieval model correctly answers qq, it is assigned label AA.
    • If non-retrieval fails but either single-step or multi-step retrieval succeeds, label BB is assigned by applying a tie-breaking rule that favors the simpler model.
    1. Dataset Inductive Bias Fallback: For queries where all three execution approaches fail to generate the ground-truth answer, labels are assigned according to the architectural design of the originating dataset:
    • Queries from single-hop QA benchmarks (e.g., SQuAD, Natural Questions, TriviaQA) receive label BB.
    • Queries from multi-hop QA benchmarks requiring compositional reasoning (e.g., MuSiQue, HotpotQA, 2WikiMultiHopQA) receive label CC.

    The classifier is trained on the resulting dataset of pairs (q,o)(q, o) using standard cross-entropy loss.

  3. Knowl 3 — Adaptive Query Routing and Answering Procedure

    algorithm

    The Adaptive-RAG inference procedure classifies an incoming query into a complexity tier and executes the corresponding retrieval-augmented generation workflow.

    Input: User query qq, Knowledge Base D\mathcal{D}, Classifier Classifier\text{Classifier}, Base LLM\text{LLM}, Retriever Retriever\text{Retriever}, Multi-Step RAG Module IRCoT\text{IRCoT}
    Output: Predicted answer aˉ\bar{a}
    o←Classifier(q)o \leftarrow \text{Classifier}(q)
    if o=′A′o = 'A' then
        aˉ←LLM(q)\bar{a} \leftarrow \text{LLM}(q)
    else if o=′B′o = 'B' then
        d←Retriever(q;D)d \leftarrow \text{Retriever}(q; \mathcal{D})
        aˉ←LLM(q,d)\bar{a} \leftarrow \text{LLM}(q, d)
    else if o=′C′o = 'C' then
        aˉ←IRCoT(q,LLM,Retriever,D)\bar{a} \leftarrow \text{IRCoT}(q, \text{LLM}, \text{Retriever}, \mathcal{D})
    end if
    return aˉ\bar{a}

    The classifier Classifier\text{Classifier} runs as a lightweight T5 encoder-decoder model prior to any document retrieval. When o=′C′o = 'C', iterative retrieval-and-reasoning (such as Interleaved Retrieval with Chain-of-Thought, IRCoT) is invoked; otherwise, either direct decoding (AA) or single-pass BM25 retrieval followed by generation (BB) is performed.

  4. Knowl 4 — Question-Answering Accuracy and Efficiency Across LLM Architectures

    data/table

    Adaptive-RAG evaluated across FLAN-T5-XL (3B), FLAN-T5-XXL (11B), and GPT-3.5-Turbo-Instruct demonstrates superior accuracy compared to static single-step and prior adaptive baselines, while requiring substantially fewer retrieval/generation steps and lower runtime than uniform multi-step retrieval.

    FLAN-T5-XL (3B) FLAN-T5-XXL (11B) GPT-3.5 (Turbo)
    Types Methods EM F1 Acc Step Time EM F1 Acc Step Time EM F1 Acc Step Time
    Simple No Retrieval 14.87 21.12 15.97 0.00 0.11 17.83 25.14 19.33 0.00 0.08 35.77 48.56 44.27 0.00 0.71
    Single-step Approach 34.83 44.31 38.87 1.00 1.00 37.87 47.63 41.90 1.00 1.00 34.73 46.99 45.27 1.00 1.00
    Adaptive Adaptive Retrieval 23.87 32.24 26.73 0.50 0.56 26.93 35.67 29.73 0.50 0.54 35.90 48.20 45.30 0.50 0.86
    Self-RAG 9.90 20.79 31.57 0.72 0.43 10.87 22.98 34.13 0.74 0.23 10.87 22.98 34.13 0.74 1.50
    Adaptive-RAG (Ours) 37.17 46.94 42.10 2.17 3.60 38.90 48.62 43.77 1.35 2.00 37.97 50.91 48.97 1.03 1.46
    Complex Multi-step Approach 39.00 48.85 43.70 4.69 8.81 40.13 50.09 45.20 2.13 3.80 38.13 50.87 49.70 2.81 3.33
    Oracle Adaptive-RAG w/ Oracle 45.00 56.28 49.90 1.28 2.11 47.17 58.60 52.20 0.84 1.10 47.70 62.80 58.57 0.50 1.03

    Metrics include Exact Match (EM), token F1 score, answer Accuracy (Acc, ground truth containment), average retrieval/generation steps (Step), and relative time per query normalized to the single-step approach (Time = 1.00). On GPT-3.5, Adaptive-RAG achieves 50.91 F1 with a relative time of 1.46, matching the Multi-step approach (50.87 F1) while reducing execution time by more than half (relative time 3.33 vs 1.46).

  5. Knowl 5 — Performance Variations Across Single-Hop and Multi-Hop QA Benchmarks

    empirical result

    Across individual evaluation datasets using FLAN-T5-XL (3B), Adaptive-RAG adjusts its resource allocation according to query difficulty:

    1. Single-Hop Datasets (SQuAD v1.1, Natural Questions, TriviaQA):
    • On SQuAD: Adaptive-RAG achieves 38.30 F1 (1.37 steps, 2.02 relative time) compared to Single-step 39.30 F1 (1.00 step, 1.00 time) and Multi-step 35.60 F1 (4.52 steps, 9.03 time).
    • On Natural Questions: Adaptive-RAG achieves 47.30 F1 (1.00 step, 1.00 relative time) matching Single-step (47.30 F1) while avoiding the 10.18 relative time overhead of Multi-step (47.80 F1).
    • On TriviaQA: Adaptive-RAG achieves 60.70 F1 (1.23 steps, 1.54 time) vs Single-step 62.40 F1 (1.00 step, 1.00 time) and Multi-step 62.40 F1 (5.28 steps, 9.22 time).
    1. Multi-Hop Datasets (MuSiQue, HotpotQA, 2WikiMultiHopQA):
    • On MuSiQue: Adaptive-RAG achieves 31.80 F1 (3.22 steps, 6.61 time), approaching Multi-step 31.90 F1 (3.60 steps, 7.58 time) and outperforming Single-step 22.80 F1.
    • On HotpotQA: Adaptive-RAG reaches 53.82 F1 (3.55 steps, 5.99 time) vs Single-step 46.15 F1 (1.00 step, 1.00 time) and Multi-step 56.54 F1 (5.53 steps, 9.38 time).
    • On 2WikiMultiHopQA: Adaptive-RAG achieves 49.75 F1 (2.63 steps, 4.68 time) vs Single-step 47.90 F1 (1.00 step, 1.00 time) and Multi-step 58.85 F1 (4.17 steps, 7.37 time).
  6. Knowl 6 — Ablation Analysis of Classifier Training Data Annotation Strategies

    data/table

    Ablation experiments comparing classifier data annotation sources demonstrate that combining silver labels from model predictions with dataset inductive bias is necessary for balancing efficiency and multi-hop accuracy.

    QA Classifier (Accuracy %)
    Training Strategies F1 Step All No (A) One (B) Multi (C)
    Adaptive-RAG (Ours) 46.94 1084 54.52 30.52 66.28 65.45
    w/o Binary 43.43 640 60.30 62.19 65.70 39.55
    w/o Silver 48.79 1464 40.00 0.00 53.98 75.91
    • w/o Binary (Only Silver Data): When inductive dataset biases are omitted, the classifier struggles to identify complex queries, dropping Multi-step classification accuracy from 65.45%65.45\% to 39.55%39.55\%, which lowers downstream QA F1 to 43.4343.43.
    • w/o Silver (Only Dataset Inductive Bias): When silver labels from model predictions are omitted, the classifier cannot identify zero-retrieval opportunities (No-retrieval classification accuracy is 0.00%0.00\%), inflating total inference steps to 14641464.
    • Full Strategy: Achieves balanced class accuracies (30.52%30.52\% on A, 66.28%66.28\% on B, 65.45%65.45\% on C) and maintains an overall F1 of 46.9446.94 with 10841084 steps.
  7. Knowl 7 — Effect of Complexity Classifier Parameter Scale on Performance

    data/table

    Varying the parameter size of the T5 classifier used for query complexity prediction demonstrates that compact language models achieve performance comparable to larger variants.

    QA Classifier (Accuracy %)
    Sizes F1 Step All No (A) One (B) Multi (C)
    Small (60M) 45.83 964 53.48 26.65 70.62 53.18
    Base (223M) 45.97 983 53.41 26.42 69.46 56.82
    Large (770M) 46.94 1084 54.52 30.52 66.28 65.45

    T5-Small (60M parameters) achieves a downstream QA F1 of 45.8345.83 with 964964 total steps, closely matching T5-Large (770M parameters, 46.9446.94 F1 with 10841084 steps). This indicates that query complexity assessment can be deployed in resource-constrained environments using lightweight models without significant performance degradation.

  8. Knowl 8 — Query Complexity Classification Latency and Label Distribution

    data/table

    Empirical runtime per query increases non-linearly with complexity class, confirming that routing queries to the minimum necessary tier yields substantial latency reductions.

    Complexity Label Time/Query (Sec.) Percentage (%)
    No Retrieval (A) 0.35 8.60
    Single-step Retrieval (B) 3.08 53.33
    Multi-step Retrieval (C) 27.18 38.07

    Class AA queries require 0.350.35 seconds, Class BB queries require 3.083.08 seconds, and Class CC queries require 27.1827.18 seconds. By routing 61.93%61.93\% of test queries to Class AA (8.60%8.60\%) or Class BB (53.33%53.33\%), the system avoids the 27.1827.18-second latency of Class CC on the majority of queries.

    Analysis of the classifier confusion matrix reveals the following distribution of predictions across ground-truth classes:

    • True AA instances: 31%31\% predicted as AA, 47%47\% as BB, 22%22\% as CC.
    • True BB instances: 10%10\% predicted as AA, 66%66\% as BB, 23%23\% as CC.
    • True CC instances: 3%3\% predicted as AA, 31%31\% as BB, 65%65\% as CC.
  9. Knowl 9 — Benchmark Datasets and Retrieval Evaluation Protocol

    experimental setup

    The experimental evaluation integrates single-hop and multi-hop open-domain question answering datasets into a unified benchmark:

    1. Single-Hop Datasets: SQuAD v1.1 (reading comprehension converted to open-domain), Natural Questions (Google search queries), and TriviaQA (quiz website trivia). The knowledge corpus is the Wikipedia dump preprocessed by Karpukhin et al. (2020).
    2. Multi-Hop Datasets: MuSiQue (2-4 hop compositional queries), HotpotQA (multi-article reasoning queries), and 2WikiMultiHopQA (2-hop Wikipedia knowledge graph queries). The knowledge corpus is the Wikipedia dump preprocessed by Trivedi et al. (2023).
    3. Retriever: Term-based sparse retrieval using Okapi BM25.
    4. Target Generator LLMs: FLAN-T5-XL (3B parameters), FLAN-T5-XXL (11B parameters), and GPT-3.5-Turbo-Instruct.
    5. Classifier Model & Training: T5-Large (770M parameters) fine-tuned using the AdamW optimizer with a learning rate of 3×10−53 \times 10^{-5} for up to 100 iterations, selecting the checkpoint with the highest validation accuracy. Training queries are sampled at 400 instances per dataset (non-overlapping with test samples), annotated using silver model outcomes and dataset inductive biases.
  10. Knowl 10 — Limitations and Theoretical Performance Upper Bound

    limitation

    The Adaptive-RAG framework exhibits three primary limitations:

    1. Heuristic Noise in Classifier Training Data: The automated dataset construction combines empirical model predictions and dataset inductive biases without human verification. Incorrect answers from base models or misaligned dataset assumptions introduce label noise during classifier training.
    2. Coarse-Grained Complexity Taxonomy: The framework models query complexity as a discrete 3-tier space {A,B,C}\{A, B, C\}, which does not differentiate finer levels of multi-hop reasoning (such as 2-hop versus 4-hop queries).
    3. Performance Gap Relative to Oracle Routing: Equipping Adaptive-RAG with an oracle classifier yields substantial improvements over the learned classifier (e.g., on GPT-3.5, Oracle achieves an F1 of 62.8062.80 in 1.031.03 relative time versus 50.9150.91 F1 in 1.461.46 relative time; on FLAN-T5-XXL, Oracle achieves 58.6058.60 F1 in 1.101.10 relative time versus 48.6248.62 F1 in 2.002.00 relative time), highlighting that classification errors are the main bottleneck constraining overall system performance.

Coverage note — None was omitted; all contributed models, automatic labeling methods, experimental datasets, comparative results, ablations, classifier analyses, and stated limitations are fully represented.

References

  1. 1.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  2. 2.Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations.
  3. 3.Jinheon Baek, Soyeong Jeong, Minki Kang, Jong Park, and Sung Ju Hwang. 2023. Knowledge-augmented language model verification. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 1720–1736. Association for Computational Linguistics.
  4. 4.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2022. Improving language models by retrieving from trillions of tokens. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 2206–2240. PMLR.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer opendomain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1870–1879. Association for Computational Linguistics.
  7. 7.Sukmin Cho, Jeongyeon Seo, Soyeong Jeong, and Jong C. Park. 2023. Improving zero-shot reader by reducing distractions from irrelevant documents in open-domain question answering. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 3145–3157. Association for Computational Linguistics.
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  9. 9.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 6609–6625. International Committee on Computational Linguistics.
  10. 10.Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021, pages 874–880. Association for Computational Linguistics.
  11. 11.Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. J. Mach. Learn. Res., 24:251:1–251:43.
  12. 12.Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park. 2023. Test-time self-adaptive small language models for question answering. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 15459–15469. Association for Computational Linguistics.
  13. 13.Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In EMNLP 2023.
  14. 14.Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1601–1611. Association for Computational Linguistics.
  15. 15.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, November 16-20, 2020. Association for Computational Linguistics.
  16. 16.Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir R. Radev, Noah A. Smith, Yejin Choi, and Kentaro Inui. 2022. Realtime QA: what’s the answer right now? arXiv preprint arXiv:2207.13332.
  17. 17.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-searchpredict: Composing retrieval and language models for knowledge-intensive NLP. arXiv preprint arXiv.2212.14024, abs/2212.14024.
  18. 18.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  19. 19.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  20. 20.Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022. Internetaugmented language models through few-shot prompting for open-domain question answering. arXiv preprint arXiv:2203.05115.
  21. 21.Belinda Z. Li, Sewon Min, Srinivasan Iyer, Yashar Mehdad, and Wen-tau Yih. 2020. Efficient one-pass end-to-end entity linking for questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6433–6441. Association for Computational Linguistics.
  22. 22.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  23. 23.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 9802–9822. Association for Computational Linguistics.
  24. 24.OpenAI. 2023. GPT-4 technical report. arXiv preprint arXiv:2303.08774.
  25. 25.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, pages 8024–8035.
  26. 26.Jayr Alencar Pereira, Robson do Nascimento Fidalgo, Roberto de Alencar Lotufo, and Rodrigo Frassetto Nogueira. 2023. Visconde: Multi-document QA with GPT-3 and neural reranking. In Advances in Information Retrieval - 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, April 2-6, 2023, Proceedings, Part II, volume 13981 of Lecture Notes in Computer Science, pages 534–543. Springer.
  27. 27.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023.
  28. 28.Peng Qi, Haejun Lee, Tg Sido, and Christopher D. Manning. 2021. Answering open-domain questions of varying reasoning steps from text. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 3599–3614. Association for Computational Linguistics.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  30. 30.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392. The Association for Computational Linguistics.
  31. 31.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented language models. Transactions of the Association for Computational Linguistics.
  32. 32.Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994. Okapi at TREC-3. In Proceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2-4, 1994, volume 500-225 of NIST Special Publication, pages 109–126. National Institute of Standards and Technology (NIST).
  33. 33.Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2023. REPLUG: retrievalaugmented black-box language models. arXiv preprint arXiv:2301.12652.
  34. 34.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and finetuned chat models. arXiv preprint arXiv:2307.09288.
  35. 35.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022a. Musique: Multihop questions via single-hop question composition. Trans. Assoc. Comput. Linguistics, 10:539–554.
  36. 36.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022b. ♪ MuSiQue: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554.
  37. 37.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving retrieval with chain-of-thought reasoning for knowledgeintensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 10014–10037. Association for Computational Linguistics.
  38. 38.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. Trans. Mach. Learn. Res., 2022.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022b. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  40. 40.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 - Demos, pages 38–45. Association for Computational Linguistics.
  41. 41.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  42. 42.Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin. 2019. End-to-end open-domain question answering with bertserini. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Demonstrations, pages 72–77. Association for Computational Linguistics.
  43. 43.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  44. 44.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  45. 45.Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, and Jianshu Chen. 2023. Thrust: Adaptively propels large language models with external knowledge. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  46. 46.Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv preprint arXiv:2101.00774.

Citation

MLA
Jeong, S., et al. “Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models Through Question Complexity”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7036–50, https://doi.org/10.18653/v1/2024.naacl-long.389.
APA
Jeong, S., Baek, J., Cho, S., Hwang, S. J., & Park, J. C. (2024). Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7036–7050. https://doi.org/10.18653/v1/2024.naacl-long.389
Chicago
Jeong, S., J. Baek, S. Cho, S. J. Hwang, and J. C. Park. 2024. “Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models Through Question Complexity”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7036–50. https://doi.org/10.18653/v1/2024.naacl-long.389.
Harvard
Jeong, S. et al. (2024) “Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7036–7050. Available at: https://doi.org/10.18653/v1/2024.naacl-long.389.
Vancouver
1. Jeong S, Baek J, Cho S, Hwang SJ, Park JC (2024) Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7036–7050

BibTeX

@inproceedings{jeong-etal-2024-adaptive,
    title = "Adaptive-{RAG}: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity",
    author = "Jeong, Soyeong  and
      Baek, Jinheon  and
      Cho, Sukmin  and
      Hwang, Sung Ju  and
      Park, Jong",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.389/",
    doi = "10.18653/v1/2024.naacl-long.389",
    pages = "7036--7050"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/