Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Harsh TrivediNiranjan BalasubramanianTushar KhotAshish Sabharwal
Proposes IRCoT, a framework that interleaves chain-of-thought reasoning with step-by-step external retrieval to reduce model hallucinations and improve multi-step open-domain question answering across both large and small language models without extra training.
Large language models frequently generate convincing yet incorrect answers—known as hallucinations—when addressing complex, multi-step questions that require missing or up-to-date knowledge. Standard retrieval solutions typically rely on a one-step query based solely on the original question. This setup breaks down during multi-step reasoning because identifying which information to fetch next depends on intermediate conclusions that are not yet evident from the initial prompt.
The article demonstrates and evaluates an interleaved framework called IRCoT (Interleaved Retrieval guided by Chain-of-Thought Reasoning). The primary objective is to test whether alternating step-by-step reasoning sentences with iterative document retrieval can improve both information retrieval quality and question-answering accuracy without requiring additional model training.
The authors evaluated the framework across four established multi-step question-answering datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC. The methodology coupled standard BM25 document search with various language models, ranging from OpenAI’s 175-billion-parameter GPT-3 to smaller Flan-T5 models ranging from 0.2 to 11 billion parameters. The process alternates between generating a single step of reasoning and using that step as a search query to gather relevant paragraphs, repeating until reaching a final answer or a step limit.
The evaluation yielded several key findings. First, IRCoT substantially enhanced retrieval quality, boosting recall by 11 to 21 percentage points when paired with GPT-3 compared to standard one-step retrieval. Second, this improved retrieval translated into significant downstream question-answering gains, increasing accuracy by up to 15 F1 points and reducing factual reasoning errors by up to 50%. Third, smaller open-source models achieved strong results; a 3-billion-parameter Flan-T5 model using IRCoT outperformed the 58-times larger GPT-3 model running on standard one-step retrieval. Finally, the approach generalized well in out-of-distribution tests where demonstration prompts from one dataset were applied to entirely different datasets.
These results demonstrate that multi-step knowledge gathering can overcome the memory and knowledge limits of language models while dramatically reducing hallucination risks. For organizations deploying generative systems on knowledge-intensive workflows, the findings suggest that architectural orchestration between retrieval and reasoning can deliver higher accuracy than simply adopting larger, more expensive models. This offers substantial performance and cost advantages.
Decision-makers should consider adopting iterative, step-guided retrieval architectures for complex knowledge management and search applications. Because each generated reasoning step triggers a separate model and search query, engineering teams must weigh the trade-off between higher computational latency and increased factual accuracy. Organizations should run targeted pilots to evaluate whether their specific workflows warrant dynamic stopping mechanisms to control call volumes.
The findings carry high confidence for structured multi-hop question-answering, but readers should note specific limitations. IRCoT requires models with sufficient context windows and baseline step-by-step reasoning capabilities, and the increased number of queries introduces additional latency and computational expense that may require filtering or caching optimizations in high-throughput production environments.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Read this foundational account of chain-of-thought prompting first: IRCoT depends on generating intermediate reasoning steps to guide what information to retrieve next.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This introduces retrieval-augmented generation’s core retriever-and-generator setup, the foundation IRCoT adapts by making retrieval iterative and reasoning-guided.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Its self-ask method decomposes multi-hop questions into successive subquestions, preparing readers for IRCoT’s use of intermediate reasoning to formulate retrieval queries.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This establishes why parametric memory can fail on factual questions and motivates the external retrieval that IRCoT interleaves with reasoning.
- Paper: Successive Prompting for Decomposing Complex Questions, Dheeru Dua et al. (2022). Its iterative question-decomposition approach provides useful groundwork for understanding how IRCoT turns each emerging reasoning step into a next action.
- Paper: DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models, Weihang Su et al. (2024). DRAGIN extends dynamic retrieval by deciding both when a model needs evidence and what to search for as generation unfolds.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Search-R1 carries the reasoning-and-search loop into training, teaching models to generate queries, use evidence, and verify answers autonomously.
- Paper: Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity, Soyeong Jeong et al. (2024). Adaptive-RAG builds on iterative retrieval by routing questions of different complexity to no-retrieval, one-shot, or multi-step strategies.
- Paper: End-to-End Beam Retrieval for Multi-Hop Question Answering, Jiahao Zhang et al. (2024). Beam Retrieval extends multi-hop question answering with an end-to-end search over alternative evidence chains and dynamically terminates retrieval.
- Paper: LLatrieval: LLM-Verified Retrieval for Verifiable Generation, Xiaonan Li et al. (2024). LLatrieval continues iterative retrieval by checking whether gathered evidence is sufficient and searching specifically for missing information.
