End-to-End Beam Retrieval for Multi-Hop Question Answering
Jiahao ZhangHaiyang ZhangDongmei ZhangYong LiuShen Huang
Presents Beam Retrieval, an end-to-end framework that tracks multiple passage hypotheses across variable reasoning hops to overcome early-stage search errors and boost downstream multi-hop question answering accuracy.
Complex question-answering systems often face significant challenges when an answer requires gathering and connecting evidence across multiple distinct passages, a process known as multi-step or multi-hop reasoning. Existing passage retrieval methods generally focus on simple two-step queries and optimize individual steps in isolation. This fragmented design leads to brittle performance on more realistic, multi-step scenarios, because an initial retrieval error tends to derail the entire sequence and feed noisy or incomplete context to the answering model.
The article evaluates a unified framework called Beam Retrieval, which treats multi-step passage selection as an end-to-end decoding process across a variable number of reasoning steps. The objective is to demonstrate that simultaneously training an underlying language model encoder alongside two dedicated classification heads—one for the initial hop and another for subsequent hops—significantly improves retrieval precision and boosts downstream answer accuracy.
The approach was validated through extensive experiments on established multi-hop benchmark datasets, primarily MuSiQue-Ans (which features challenging 2- to 4-hop queries), HotpotQA, 2WikiMultihopQA, and the Incomplete Information Reading Comprehension (IIRC) dataset. Beam Retrieval uses a beam search mechanism that tracks multiple candidate chains of evidence simultaneously, terminating the search dynamically when confidence falls below a set threshold. The retrieved passages were then supplied either to dedicated supervised reading models or to few-shot large language models to evaluate final question-answering accuracy.
The findings show that Beam Retrieval establishes a new state of the art across all evaluated benchmarks. On the challenging MuSiQue-Ans dataset, it improved exact-match retrieval accuracy by nearly 50% relative to prior baselines (rising from 53.50% to 79.31%), and it achieved 99.9% precision on 2WikiMultihopQA. In downstream question answering, the framework enabled a supervised reader to reach a 91.4% supporting-passage score on MuSiQue-Ans (approaching the human benchmark of 93.9%) and drove substantial accuracy gains when pairing large language models with retrieved passages compared to providing raw, unfiltered candidates. Additionally, ablation results confirmed that maintaining identical search beam sizes between training and inference, as well as retaining two distinct classification heads, is critical to optimal performance.
These results demonstrate that joint optimization and multi-hypothesis tracking effectively insulate retrieval pipelines from early-stage compounding errors. For operational question-answering and search architectures, filtering inputs with a robust multi-hop retriever sharply reduces extraneous text, which can cut computation overhead for downstream large language models while minimizing hallucinations.
For practical implementation, the article suggests adopting a beam size of 1 for standard applications, as it matches the low latency and resource consumption of existing methods while still outperforming them; a beam size of 2 is recommended when maximum retrieval accuracy is essential on highly complex tasks. Future development should focus on extending the framework to fully open-domain web environments and engineering optimizations to mitigate the increased memory footprint encountered during training with larger beam sizes.
Confidence in these findings is high for bounded reading comprehension settings with moderate candidate pools (10 to 25 passages per query). However, caution is warranted when extrapolating directly to vast, open-domain web corpora without an initial candidate generation stage, as the framework was tested primarily as a reranker rather than a standalone open-web search engine.
- Paper: Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps, Xanh Ho et al. (2020). It introduces 2WikiMultihopQA, one of the primary benchmark datasets used by the source paper to evaluate multi-hop passage retrieval and reasoning.
- Paper: Dense Passage Retrieval for Open-Domain Question Answering, Vladimir Karpukhin et al. (2020). It establishes Dense Passage Retrieval (DPR) for question answering, providing the foundational neural retrieval paradigm that multi-step retrieval methods build upon.
- Paper: Passage Re-ranking with BERT, Rodrigo Nogueira et al. (2019). It details how cross-encoder neural language models can be fine-tuned to score and re-rank candidate passages, a foundational mechanism behind beam retrieval's classification heads.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). It formulates the core Retrieval-Augmented Generation (RAG) framework connecting passage retrieval to downstream generator and reader models.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). It explores dynamic, multi-step active retrieval during reasoning, offering essential background on the limitations of single-step retrieval for multi-hop QA.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). It formalizes the compositionality gap and demonstrates how multi-hop reasoning breaks down across sub-facts in multi-step QA benchmarks.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It introduces the conceptual foundation of end-to-end multi-hop memory retrieval across sequential reasoning steps.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). It advances multi-step retrieval from supervised beam search over candidate passages to an autonomous, reinforcement-learned interactive search engine agent.
- Paper: Search-o1: Agentic Search-Enhanced Large Reasoning Models, Xiaoxi Li et al. (2025). It extends multi-step retrieval and reasoning to large reasoning models by integrating dynamic agentic search queries and knowledge refinement directly into chain-of-thought problem solving.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). It builds on multi-passage retrieval filtering by unifying context reranking directly into the instruction-tuned generator model itself.
- Paper: Self-Improving Language Models with Bidirectional Evolutionary Search, Guowei Xu et al. (2026). It generalizes multi-hop reasoning search strategies by combining forward expansion with backward goal decomposition on multi-hop benchmarks like MuSiQue.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). It extends multi-hop exploration strategies to structured knowledge graphs by introducing adaptive path planning and dynamic self-correcting memory.
