Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Harsh TrivediNiranjan BalasubramanianTushar KhotAshish Sabharwal

article2023ACL922 citations

Proposes IRCoT, a framework that interleaves chain-of-thought reasoning with step-by-step external retrieval to reduce model hallucinations and improve multi-step open-domain question answering across both large and small language models without extra training.

Listen

Large language models frequently generate convincing yet incorrect answers—known as hallucinations—when addressing complex, multi-step questions that require missing or up-to-date knowledge. Standard retrieval solutions typically rely on a one-step query based solely on the original question. This setup breaks down during multi-step reasoning because identifying which information to fetch next depends on intermediate conclusions that are not yet evident from the initial prompt.

The article demonstrates and evaluates an interleaved framework called IRCoT (Interleaved Retrieval guided by Chain-of-Thought Reasoning). The primary objective is to test whether alternating step-by-step reasoning sentences with iterative document retrieval can improve both information retrieval quality and question-answering accuracy without requiring additional model training.

The authors evaluated the framework across four established multi-step question-answering datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC. The methodology coupled standard BM25 document search with various language models, ranging from OpenAI’s 175-billion-parameter GPT-3 to smaller Flan-T5 models ranging from 0.2 to 11 billion parameters. The process alternates between generating a single step of reasoning and using that step as a search query to gather relevant paragraphs, repeating until reaching a final answer or a step limit.

The evaluation yielded several key findings. First, IRCoT substantially enhanced retrieval quality, boosting recall by 11 to 21 percentage points when paired with GPT-3 compared to standard one-step retrieval. Second, this improved retrieval translated into significant downstream question-answering gains, increasing accuracy by up to 15 F1 points and reducing factual reasoning errors by up to 50%. Third, smaller open-source models achieved strong results; a 3-billion-parameter Flan-T5 model using IRCoT outperformed the 58-times larger GPT-3 model running on standard one-step retrieval. Finally, the approach generalized well in out-of-distribution tests where demonstration prompts from one dataset were applied to entirely different datasets.

These results demonstrate that multi-step knowledge gathering can overcome the memory and knowledge limits of language models while dramatically reducing hallucination risks. For organizations deploying generative systems on knowledge-intensive workflows, the findings suggest that architectural orchestration between retrieval and reasoning can deliver higher accuracy than simply adopting larger, more expensive models. This offers substantial performance and cost advantages.

Decision-makers should consider adopting iterative, step-guided retrieval architectures for complex knowledge management and search applications. Because each generated reasoning step triggers a separate model and search query, engineering teams must weigh the trade-off between higher computational latency and increased factual accuracy. Organizations should run targeted pilots to evaluate whether their specific workflows warrant dynamic stopping mechanisms to control call volumes.

The findings carry high confidence for structured multi-hop question-answering, but readers should note specific limitations. IRCoT requires models with sufficient context windows and baseline step-by-step reasoning capabilities, and the increased number of queries introduces additional latency and computational expense that may require filtering or caching optimizations in high-throughput production environments.

arXiv: 2212.10509stonybrooknlp/ircot
Cover for Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Abstract

Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA). They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters. While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA. Here, what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved. To address this, we propose IRCoT, a new approach for multi-step QA that interleaves retrieval with steps (sentences) in a CoT, guiding the retrieval with CoT and in turn using retrieved results to improve CoT. Using IRCoT with GPT3 substantially improves retrieval (up to 21 points) as well as downstream QA (up to 15 points) on four datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC. We observe similar substantial gains in out-of-distribution (OOD) settings as well as with much smaller models such as Flan-T5-large without additional training. IRCoT reduces model hallucination, resulting in factually more accurate CoT reasoning.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Chain-of-Thought-Guided Retrieval and Open-Domain QA
  • 3.1 Interleaving Retrieval with Chain-of-Thought Reasoning
  • 3.2 Question Answering Reader
  • 4 Experimental Setup
  • 4.1 Models
  • 5 Results
  • 6 Conclusions
  • References
  • A Constructing Retrieval Corpora
  • B Special Handling of Models for IIRC
  • C Comparison with Previous Systems for ODQA with LLMs
  • D Additional CoT Generation Examples
  • E Direct vs CoT Prompting Readers
  • F Separate Reader in IRCoT QA
  • G Prompts

Knowls

  1. Knowl 1 — IRCoT alternates reasoning and retrieval

    algorithm

    IRCoT answers a knowledge-intensive multi-step question by alternating language-model reasoning with external retrieval. It requires a question, a paragraph retriever, a language model that can generate chain-of-thought (CoT) text, and a corpus. In the paper’s experimental implementation, the retriever is BM25 and the CoT generator is GPT-3 code-davinci-002 or a Flan-T5 model.

    1. Retrieve an initial set of paragraphs using the question as the query.
    2. Ask the language model to generate the next CoT sentence using the question, all paragraphs collected so far, and all CoT sentences generated so far. If it generates multiple sentences, retain only the first.
    3. Use that sentence as a query to retrieve more paragraphs, and add them to the collection.
    4. Repeat the reasoning and retrieval steps until a generated sentence contains “answer is:” or the maximum of 8 reasoning steps is reached. Return all collected paragraphs, with the total collection capped at 15 paragraphs in the experiments.

    The reasoner’s demonstrations contain complete CoTs, gold supporting paragraphs, and shuffled distractor paragraphs. At test time, the reasoner receives all paragraphs collected to that point. A separate QA reader then answers using the returned paragraphs; IRCoT’s internally generated CoT is not itself used as the final reader output in the main experiments.

  2. Knowl 2 — IRCoT improves retrieval recall over one-step question retrieval

    empirical result

    On HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC, IRCoT retrieved more gold supporting paragraphs than one-step BM25 retrieval using only the question. The comparison used a maximum budget of 15 paragraphs; the reported recall is the best recall obtained after selecting retrieval hyperparameters on development data. Since the systems return an unranked and variable number of paragraphs, the authors use this fixed-budget optimal recall rather than a rank-based or fixed-kk metric.

    Recall scores for Flan-T5-XXL, in dataset order HotpotQA / 2WikiMultihopQA / MuSiQue / IIRC, were 61.5 / 68.1 / 44.6 / 38.4 for one-step retrieval and 69.4 / 82.4 / 48.1 / 48.6 for IRCoT. For GPT-3, the corresponding scores were 61.5 / 68.1 / 44.6 / 36.0 and 72.8 / 90.7 / 57.1 / 57.2. Thus IRCoT’s gains were 7.9, 14.3, 3.5, and 10.2 points for Flan-T5-XXL, and 11.3, 22.6, 12.5, and 21.2 points for GPT-3, respectively.

  3. Knowl 3 — IRCoT improves answer F1 with selected reader prompts

    empirical result

    In open-domain QA, IRCoT retrieval generally produced higher answer F1 than either one-step retrieval or answering from the language model without retrieved paragraphs. The reported values are means and standard deviations over three demonstration-set runs. Flan-T5-XXL used direct-answer prompting, while GPT-3 used CoT prompting, the reader strategies selected by the authors for their main comparisons.

    For HotpotQA / 2WikiMultihopQA / MuSiQue / IIRC, Flan-T5-XXL F1 was 25.3±0.3 / 32.7±0.3 / 13.7±0.3 / 28.9±0.3 without retrieval; 49.7±0.5 / 51.2±0.3 / 25.8±0.6 / 40.0±1.3 with one-step retrieval; and 59.1±0.9 / 66.5±1.4 / 30.8±0.2 / 42.5±2.1 with IRCoT. GPT-3 F1 was 47.5±0.4 / 41.2±1.0 / 25.2±1.2 / 52.1±0.1 without retrieval; 53.6±0.7 / 54.8±2.1 / 29.4±0.8 / 49.8±2.3 with one-step retrieval; and 60.7±1.1 / 68.0±1.5 / 36.5±1.2 / 49.9±1.1 with IRCoT. IRCoT improved over one-step retrieval in all Flan-T5-XXL conditions and in GPT-3 conditions except IIRC, where the F1 difference was only 0.1 point despite IRCoT’s higher retrieval recall.

  4. Knowl 4 — IRCoT reduces factual errors in generated reasoning

    empirical result

    The authors manually assessed GPT-3-generated CoTs for 40 randomly sampled questions from each of HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC. A CoT counted as having a factual error if at least one fact in a sentence before the final answer sentence was false. The numbers of questions with at least one such error, ordered by dataset and then by method (no retrieval / one-step retrieval / IRCoT), were: HotpotQA 15 / 10 / 5; 2WikiMultihopQA 29 / 23 / 14; MuSiQue 28 / 27 / 23; and IIRC 15 / 14 / 11. IRCoT therefore had the fewest factual-error cases in each dataset; relative to one-step retrieval, the reduction was 50% on HotpotQA and approximately 40% on 2WikiMultihopQA. The paper notes that a factual error need not imply an incorrect answer, and an incorrect answer need not imply a factual error in the reasoning.

  5. Knowl 5 — IRCoT generalizes across datasets in OOD evaluation

    empirical result

    The authors tested out-of-distribution generalization by using CoT demonstrations from one dataset and evaluating on a different dataset, using the evaluation dataset’s corpus for retrieval. Across all directed dataset pairs among HotpotQA, 2WikiMultihopQA, and MuSiQue, both Flan-T5-XXL and GPT-3 showed the same qualitative pattern as in in-distribution evaluation: IRCoT retrieval had higher recall than one-step retrieval, and IRCoT QA had higher answer F1 than both one-step retrieval QA and retrieval-free QA. IIRC was excluded from this cross-dataset experiment because its question format and retrieval procedure require special handling.

  6. Knowl 6 — IRCoT benefits retrieval across language-model sizes

    empirical result

    On HotpotQA, 2WikiMultihopQA, and MuSiQue, IRCoT retrieval outperformed one-step retrieval with each tested CoT generator: Flan-T5-base (0.2B parameters), Flan-T5-large (0.7B), Flan-T5-XL (3B), Flan-T5-XXL (11B), and GPT-3 code-davinci-002 (175B). One-step retrieval does not use a language model, so its recall is fixed across model sizes. In the order HotpotQA / 2WikiMultihopQA / MuSiQue, one-step recall was 61.5 / 68.1 / 44.6. IRCoT recall for Flan-T5 sizes 0.2B, 0.7B, 3B, 11B, and GPT-3 175B was, respectively: 64.9 / 74.4 / 46.8; 68.6 / 82.6 / 48.9; 71.5 / 83.8 / 52.1; 69.4 / 82.4 / 48.1; and 72.8 / 90.7 / 57.1.

    For answer F1, IRCoT QA outperformed one-step retrieval QA at all tested model sizes except the smallest 0.2B model. In particular, IRCoT QA with the 3B Flan-T5-XL outperformed both one-step retrieval QA and retrieval-free QA using the 175B GPT-3 on all three datasets. IIRC was excluded from the size comparison because the smaller models were not reliable at identifying Wikipedia page titles, which the IIRC retrieval procedure requires.

  7. Knowl 7 — Benchmark construction and experimental protocol

    experimental setup

    The experiments cover HotpotQA, 2WikiMultihopQA, the answerable subset of MuSiQue, and the answerable subset of IIRC. HotpotQA uses its supplied Wikipedia corpus. For 2WikiMultihopQA and MuSiQue, the open-domain corpus combines supporting and nonsupporting paragraphs associated with questions across the datasets’ train, development, and test splits. For IIRC, the corpus contains paragraphs from Wikipedia pages represented in the dataset. The resulting corpus sizes were 5,233,329 paragraphs for HotpotQA, 430,225 for 2WikiMultihopQA, 139,416 for MuSiQue, and 1,882,415 for IIRC.

    For each dataset, the authors sampled 100 development questions for hyperparameter tuning and 500 separate questions for testing. They wrote CoT annotations for 20 questions per dataset, then formed three demonstration sets by sampling 15 questions per set. Hyperparameters were selected using the first demonstration set; results were evaluated with each of the three sets and reported as their mean and standard deviation. BM25, implemented in Elasticsearch, was the base retriever. One-step retrieval selected its paragraph count from 5, 7, 9, 11, 13, or 15; IRCoT selected the number per retrieval step from 2, 4, 6, or 8, with the total collection capped at 15. The number of distractor paragraphs in demonstrations was selected from 1, 2, or 3. For IIRC, the main passage was always supplied; the model generated three candidate Wikipedia page titles at test time, those titles were mapped to corpus titles using BM25, and subsequent paragraph retrieval was restricted to those pages.

  8. Knowl 8 — IRCoT’s published-system comparison is strong but not head-to-head

    empirical result

    The authors compared GPT-3 IRCoT QA with previously published LLM-based open-domain QA scores, while cautioning that the systems used different models, retrieval sources, APIs, and test subsets, so the comparison is not a controlled head-to-head evaluation. IRCoT’s EM / F1 scores were 45.8 / 58.5 on HotpotQA bridge questions, 49.3 / 60.7 on full HotpotQA, 57.7 / 68.0 on 2WikiMultihopQA, 34.2 / 43.8 on 2-hop MuSiQue, and 26.5 / 36.5 on full MuSiQue. In the updated comparison, DSP scored 51.4 / 62.9 on HotpotQA, and DecomP scored 53.5 / 70.8 on 2WikiMultihopQA; these are higher than IRCoT on those datasets’ reported F1 scores. The authors report that IRCoT remained the strongest listed system on MuSiQue and was close to the best reported scores on HotpotQA and 2WikiMultihopQA.

  9. Knowl 9 — IRCoT QA uses a separate reader after retrieval

    empirical result

    Although IRCoT generates a CoT while retrieving, the main QA system uses a separate reader to produce the final answer. An ablation compared this design with extracting the answer directly from the retrieval-stage CoT. For Flan-T5-XXL, using a separate reader yielded higher F1 on all four datasets: 59.1 versus 52.6 on HotpotQA, 66.5 versus 60.9 on 2WikiMultihopQA, 30.8 versus 24.9 on MuSiQue, and 42.5 versus 40.3 on IIRC. For GPT-3, separate-reader versus no-reader F1 was 60.7 versus 61.0, 68.0 versus 70.4, 36.5 versus 31.5, and 49.9 versus 48.4, respectively. Thus the separate reader was not uniformly better for GPT-3, but was better or close on three datasets and substantially better on MuSiQue; the authors retained the separate-reader design for their experiments.

  10. Knowl 10 — IRCoT depends on CoT capability, long context, and extra inference calls

    limitation

    IRCoT requires a base language model capable of zero- or few-shot CoT generation; the authors note that this capability is less common in models below 20B parameters, limiting applicability to smaller models. It also requires the model to accept the collected paragraphs together with demonstrations, creating a long-context requirement. Each reasoning sentence triggers a separate language-model call, so the retrieval and QA gains incur additional computation. Finally, some experiments used OpenAI’s code-davinci-002 API, which was deprecated after submission and makes those runs difficult to reproduce; the authors state that they believe the qualitative trends would persist, and note that their Flan-T5 experiments remain reproducible using publicly available weights.

Coverage note — The qualitative example CoTs and the direct-versus-CoT reader-prompting comparison are omitted because they illustrate or refine the reported factuality and QA findings rather than adding a separate central result.

References

  1. 1.Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, and Caiming Xiong. 2020. Learning to retrieve reasoning paths over wikipedia graph for question answering. In International Conference on Learning Representations.
  2. 2.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack Rae, Erich Elsen, and Laurent Sifre. 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 2206–2240. PMLR.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  6. 6.Rajarshi Das, Shehzaad Dhuliawala, Manzil Zaheer, and Andrew McCallum. 2019. Multi-step retriever-reader interaction for scalable open-domain question answering. In International Conference on Learning Representations.
  7. 7.Yair Feldman and Ran El-Yaniv. 2019. Multi-hop paragraph retrieval for open-domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2296–2309, Florence, Italy. Association for Computational Linguistics.
  8. 8.James Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot, and Pradeep Dasigi. 2020. IIRC: A dataset of incomplete information reading comprehension questions. In EMNLP.
  9. 9.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3929–3938. PMLR.
  10. 10.Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022. Large language models are reasoning teachers. arXiv preprint arXiv:2212.10071.
  11. 11.Xanh Ho, A. Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. In COLING.
  12. 12.Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Atlas: Few-shot learning with retrieval augmented language models. arXiv preprint arXiv:2208.03299.
  13. 13.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  14. 14.Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, and Kentaro Inui. 2022. RealTime QA: What’s the answer right now? arXiv preprint arXiv:2207.13332.
  15. 15.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021. Baleen: Robust multi-hop reasoning at scale via condensed retrieval. In Advances in Neural Information Processing Systems.
  16. 16.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2023. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP.
  17. 17.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks.
  18. 18.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  19. 19.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In ICML 2022 Workshop on Knowledge Retrieval and Language Models.
  20. 20.Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022. Internet-augmented language models through few-shot prompting for open-domain question answering. arXiv preprint arXiv:2203.05115.
  21. 21.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474. Curran Associates, Inc.
  22. 22.Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2022. Teaching small language models to reason. arXiv preprint arXiv:2212.08410.
  23. 23.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  24. 24.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems.
  25. 25.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  26. 26.Peng Qi, Xiaowen Lin, Leo Mehr, Zijian Wang, and Christopher D. Manning. 2019. Answering complex open-domain questions through iterative query generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2590–2602, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  28. 28.Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models. arXiv preprint arXiv:2210.01296.
  29. 29.Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. UL2: Unifying language learning paradigms. In The Eleventh International Conference on Learning Representations.
  30. 30.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multi-hop questions via single-hop question composition. TACL, 10:539–554.
  31. 31.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  32. 32.Wenhan Xiong, Xiang Li, Srini Iyer, Jingfei Du, Patrick Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel, Douwe Kiela, and Barlas Oguz. 2021. Answering complex open-domain questions with multi-hop dense retrieval. In International Conference on Learning Representations.
  33. 33.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In EMNLP.
  34. 34.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  35. 35.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023. Generate rather than retrieve: Large language models are strong context generators. In The Eleventh International Conference on Learning Representations.
  36. 36.Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. arXiv preprint arXiv:2101.00774.

Citation

MLA
Trivedi, H., et al. “Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 10014–37, https://doi.org/10.18653/v1/2023.acl-long.557.
APA
Trivedi, H., Balasubramanian, N., Khot, T., & Sabharwal, A. (2023). Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10014–10037. https://doi.org/10.18653/v1/2023.acl-long.557
Chicago
Trivedi, H., N. Balasubramanian, T. Khot, and A. Sabharwal. 2023. “Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10014–37. https://doi.org/10.18653/v1/2023.acl-long.557.
Harvard
Trivedi, H. et al. (2023) “Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10014–10037. Available at: https://doi.org/10.18653/v1/2023.acl-long.557.
Vancouver
1. Trivedi H, Balasubramanian N, Khot T, Sabharwal A (2023) Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10014–10037

BibTeX

@inproceedings{trivedi-etal-2023-interleaving,
    title = "Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions",
    author = "Trivedi, Harsh  and
      Balasubramanian, Niranjan  and
      Khot, Tushar  and
      Sabharwal, Ashish",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.557/",
    doi = "10.18653/v1/2023.acl-long.557",
    pages = "10014--10037"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/