End-to-End Beam Retrieval for Multi-Hop Question Answering

Jiahao ZhangHaiyang ZhangDongmei ZhangYong LiuShen Huang

article2024NAACL64 citations

Presents Beam Retrieval, an end-to-end framework that tracks multiple passage hypotheses across variable reasoning hops to overcome early-stage search errors and boost downstream multi-hop question answering accuracy.

Listen

Complex question-answering systems often face significant challenges when an answer requires gathering and connecting evidence across multiple distinct passages, a process known as multi-step or multi-hop reasoning. Existing passage retrieval methods generally focus on simple two-step queries and optimize individual steps in isolation. This fragmented design leads to brittle performance on more realistic, multi-step scenarios, because an initial retrieval error tends to derail the entire sequence and feed noisy or incomplete context to the answering model.

The article evaluates a unified framework called Beam Retrieval, which treats multi-step passage selection as an end-to-end decoding process across a variable number of reasoning steps. The objective is to demonstrate that simultaneously training an underlying language model encoder alongside two dedicated classification heads—one for the initial hop and another for subsequent hops—significantly improves retrieval precision and boosts downstream answer accuracy.

The approach was validated through extensive experiments on established multi-hop benchmark datasets, primarily MuSiQue-Ans (which features challenging 2- to 4-hop queries), HotpotQA, 2WikiMultihopQA, and the Incomplete Information Reading Comprehension (IIRC) dataset. Beam Retrieval uses a beam search mechanism that tracks multiple candidate chains of evidence simultaneously, terminating the search dynamically when confidence falls below a set threshold. The retrieved passages were then supplied either to dedicated supervised reading models or to few-shot large language models to evaluate final question-answering accuracy.

The findings show that Beam Retrieval establishes a new state of the art across all evaluated benchmarks. On the challenging MuSiQue-Ans dataset, it improved exact-match retrieval accuracy by nearly 50% relative to prior baselines (rising from 53.50% to 79.31%), and it achieved 99.9% precision on 2WikiMultihopQA. In downstream question answering, the framework enabled a supervised reader to reach a 91.4% supporting-passage score on MuSiQue-Ans (approaching the human benchmark of 93.9%) and drove substantial accuracy gains when pairing large language models with retrieved passages compared to providing raw, unfiltered candidates. Additionally, ablation results confirmed that maintaining identical search beam sizes between training and inference, as well as retaining two distinct classification heads, is critical to optimal performance.

These results demonstrate that joint optimization and multi-hypothesis tracking effectively insulate retrieval pipelines from early-stage compounding errors. For operational question-answering and search architectures, filtering inputs with a robust multi-hop retriever sharply reduces extraneous text, which can cut computation overhead for downstream large language models while minimizing hallucinations.

For practical implementation, the article suggests adopting a beam size of 1 for standard applications, as it matches the low latency and resource consumption of existing methods while still outperforming them; a beam size of 2 is recommended when maximum retrieval accuracy is essential on highly complex tasks. Future development should focus on extending the framework to fully open-domain web environments and engineering optimizations to mitigate the increased memory footprint encountered during training with larger beam sizes.

Confidence in these findings is high for bounded reading comprehension settings with moderate candidate pools (10 to 25 passages per query). However, caution is warranted when extrapolating directly to vast, open-domain web corpora without an initial candidate generation stage, as the framework was tested primarily as a reranker rather than a standalone open-web search engine.

arXiv: 2308.08973
Cover for End-to-End Beam Retrieval for Multi-Hop Question Answering

Abstract

Multi-hop question answering (QA) involves finding multiple relevant passages and step-by-step reasoning to answer complex questions, indicating a retrieve-and-read paradigm. However, previous retrievers were customized for two-hop questions, and most of them were trained separately across different hops, resulting in a lack of supervision over the entire multi-hop retrieval process and leading to poor performance in complicated scenarios beyond two hops. In this work, we introduce Beam Retrieval, an end-to-end beam retrieval framework for multi-hop QA. This approach models the multi-hop retrieval process in an end-to-end manner by jointly optimizing an encoder and two classification heads across all hops. Moreover, Beam Retrieval maintains multiple partial hypotheses of relevant passages at each step, expanding the search space and reducing the risk of missing relevant passages. To establish a complete QA system, we incorporate a supervised reader or a large language model (LLM). Experimental results demonstrate that Beam Retrieval achieves a nearly 50% improvement compared with baselines on challenging MuSiQue-Ans, and it also surpasses all previous retrievers on HotpotQA and achieves 99.9% precision on 2WikiMultiHopQA. Providing high-quality context, Beam Retrieval helps our supervised reader achieve new state-of-the-art performance and substantially improves the few-shot QA performance of LLMs1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Beam Retrieval
  • 3.1 Problem Formulation
  • 3.2 Scoring
  • 3.3 End-to-End Inference
  • 3.4 Joint Optimization
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Models
  • 4.2.1 Beam Retrieval
  • 4.2.2 Downstream Reader
  • 4.3 Evaluation Metrics
  • 5 Results
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Multi-Task Supervised Reader
  • B Few-Shot Templates
  • B.1 without Beam Retrieval
  • B.2 with Beam Retrieval
  • C Analysis of Beam Search Algorithm in Beam Retrieval
  • D Beam Retrieval Performance on Dataset IIRC
  • E Ablation Study on Shuffle Operation

Knowls

  1. Knowl 1 — End-to-end variable-hop retrieval formulation

    model/method

    Beam Retrieval treats multi-hop passage selection as sequence decoding. Given a question QQ, a candidate set of nn passages D={p1,…,pn}\mathcal{D}=\{p_1,\ldots,p_n\}, and an unknown or known maximum of kk hops, the retriever outputs a chain of distinct passages (p^1,…,p^k)(\hat p_1,\ldots,\hat p_k). At hop tt, the next passage is selected conditionally on QQ and the passages already selected, written as p^<t=(p^1,…,p^t−1)\hat p_{<t}=(\hat p_1,\ldots,\hat p_{t-1}).

    The method maps candidate passages to possible decoding tokens and maps the question to a start token. For every unused candidate p∈D∖p^<tp\in\mathcal{D}\setminus\hat p_{<t}, it assigns a score S(p∣Q,p^<t)S(p\mid Q,\hat p_{<t}). The conditional next-passage distribution is defined by normalizing these scores over unused candidates:

    P(p^t=p∣Q,p^<t)=S(p∣Q,p^<t)∑q∈D∖p^<tS(q∣Q,p^<t).P(\hat p_t=p\mid Q,\hat p_{<t})=\frac{S(p\mid Q,\hat p_{<t})}{\sum_{q\in\mathcal{D}\setminus\hat p_{<t}}S(q\mid Q,\hat p_{<t})}.

    Passage uniqueness is enforced because a ground-truth relevant chain contains no duplicate passage. Unlike one-step or two-step sequence-labeling retrievers, this formulation represents the complete multi-hop retrieval process and can terminate at a data-dependent hop.

  2. Knowl 2 — Two-head encoder scoring architecture

    model/method

    Beam Retrieval uses one shared auto-encoder language model and two classification heads. At the first hop, each candidate passage pip_i is encoded with the sequence [CLS]∥Q∥pi∥[SEP][\mathrm{CLS}]\mathbin{\Vert}Q\mathbin{\Vert}p_i\mathbin{\Vert}[\mathrm{SEP}], where ∥\mathbin{\Vert} denotes concatenation. If the encoded sequence has length LiL_i and the encoder hidden size is hh, its representation is Hi∈RLi×hH_i\in\mathbb{R}^{L_i\times h}. The first classification head maps the final [CLS][\mathrm{CLS}] representation to two outputs, corresponding to irrelevant and relevant, and the relevant output is the first-hop score S(pi∣Q)S(p_i\mid Q).

    At every later hop t>1t>1, each partial passage chain Pt−1=(p1,…,pt−1)P_{t-1}=(p_1,\ldots,p_{t-1}) is concatenated with an unused candidate pp as [CLS]∥Q∥Pt−1∥p∥[SEP][\mathrm{CLS}]\mathbin{\Vert}Q\mathbin{\Vert}P_{t-1}\mathbin{\Vert}p\mathbin{\Vert}[\mathrm{SEP}]. The same encoder and a second two-classification-output head produce S(p∣Q,Pt−1)S(p\mid Q,P_{t-1}). The first head handles the fixed set of nn first-hop candidates, whereas the second head handles the variable-sized set of candidates generated by beam expansion. The two heads are jointly trained with the shared encoder across all hops.

  3. Knowl 3 — Beam-search inference procedure

    algorithm

    Beam Retrieval maintains BB partial passage chains rather than committing to one chain after each hop. The paper uses a stopping threshold τ=−1\tau=-1 and experiments primarily with beam sizes B=1B=1 and B=2B=2.

    Input: question Q, candidate passages D, beam size B, stopping threshold tau
    Output: one retrieved chain of relevant passages
    Score every singleton chain (p) for p in D with classifier1.
    Keep the B highest-scoring singleton chains as the current beam.
    Set the previous beam to the current beam.
    for hop t = 2, 3, ...:
        Expand every chain in the current beam with every unused passage.
        Score all expanded chains with classifier2.
        If the highest expanded-chain score is below tau:
            return the highest-scoring chain in the previous beam
        Set the previous beam to the current beam.
        Keep the B highest-scoring expanded chains as the current beam.
    return the highest-scoring chain in the current beam

    At each expansion, a chain cannot reuse a passage already contained in that chain. If the current best score falls below τ\tau, the method returns the highest-scoring chain from the preceding hop; otherwise, it continues expanding. For a kk-hop chain, the implementation calls the encoder kk times, the first classifier once, and the second classifier k−1k-1 times. Beam search enlarges the explored hypothesis set and reduces the chance that an early retrieval error eliminates the correct chain.

  4. Knowl 4 — Joint loss over all retrieval hops

    equation

    Beam Retrieval jointly optimizes the shared encoder, the first-hop classifier, and the later-hop classifier. Let (p1,…,pk)(p_1,\ldots,p_k) be the ground-truth relevant passages, let BB be the training beam size, let Pt−1bP_{t-1}^b be the bbth partial hypothesis at hop t−1t-1, and let S(p∣Q,P)S(p\mid Q,P) be the relevant-class probability for candidate passage pp. The first-hop binary cross-entropy loss is

    L1=−∑p∈D[l1,plog⁡S(p∣Q)+(1−l1,p)log⁡(1−S(p∣Q))].\mathcal{L}_1=-\sum_{p\in\mathcal{D}}\left[l_{1,p}\log S(p\mid Q)+(1-l_{1,p})\log\bigl(1-S(p\mid Q)\bigr)\right].

    For each later hop t∈{2,…,k}t\in\{2,\ldots,k\}, the loss over all beam hypotheses is

    Lt=−∑b=1B∑p∈D∖Pt−1b[lt,plog⁡S(p∣Q,Pt−1b)+(1−lt,p)log⁡(1−S(p∣Q,Pt−1b))].\mathcal{L}_t=-\sum_{b=1}^{B}\sum_{p\in\mathcal{D}\setminus P_{t-1}^{b}}\left[l_{t,p}\log S(p\mid Q,P_{t-1}^{b})+(1-l_{t,p})\log\bigl(1-S(p\mid Q,P_{t-1}^{b})\bigr)\right].

    The total retrieval loss is the sum across hops:

    L=∑t=1kLt.\mathcal{L}=\sum_{t=1}^{k}\mathcal{L}_t.

    When the dataset specifies an ordering, lt,p=1l_{t,p}=1 only for the ground-truth passage ptp_t at hop tt, and it is 00 for every other candidate. When only an unordered set of relevant passages is available, lt,p=1l_{t,p}=1 for every passage in the ground-truth relevant set and 00 otherwise. Increasing BB supplies more irrelevant passage sequences as negative training examples, helping the model learn when to stop, but also increases the imbalance between negative and positive examples.

  5. Knowl 5 — Benchmark and implementation protocol

    experimental setup

    The primary experiments evaluate passage retrieval in the reading-comprehension setting on MuSiQue-Ans, distractor-setting HotpotQA, and 2WikiMultihopQA. Each question has respectively 20, 10, and 10 candidate passages. MuSiQue-Ans contains 2-, 3-, and 4-hop questions; HotpotQA contains 2-hop questions; and 2WikiMultihopQA contains 2- and 4-hop questions. The entity-relation annotations supplied by 2WikiMultihopQA are not used.

    Beam Retrieval uses base or large DeBERTa encoders, 16 training epochs, batch size 1 at the example level, AdamW with learning rate 2×10−52\times10^{-5}, maximum input length 512, and gradient checkpointing. The retrieval threshold is τ=−1\tau=-1. When concatenated passages exceed the length limit, passages are truncated using their average length, and passage order is shuffled during training.

    For supervised downstream QA, the retrieved context is concatenated with the question. DeBERTa-large is used for MuSiQue-Ans and 2WikiMultihopQA, while DeBERTa-xxlarge is used for HotpotQA. These readers are trained for 12 epochs with batch size 4, AdamW learning rate 5×10−65\times10^{-6}, and maximum input length 1024. Retrieval is evaluated with passage-level exact match (EM) and F1, ignoring the order among relevant passages. Downstream QA is evaluated with answer and supporting-fact EM/F1; MuSiQue-Ans reports answer F1 and supporting-passage F1.

  6. Knowl 6 — Retrieval performance across three benchmarks

    data/table

    On the development sets, Beam Retrieval substantially outperforms the reported prior retrievers, especially on the more complex MuSiQue-Ans benchmark. The values are retrieval EM/F1, where higher is better.

    • MuSiQue-Ans: EE: 21.47/67.61; SA: 30.37/72.30; Ex(EE): 48.78/77.79; Ex(SA): 53.50/79.24; Beam Retrieval with beam size 1: 77.37/89.77; Beam Retrieval with beam size 2: 79.31/90.51.
    • HotpotQA: SAE: 91.98/95.76; SA Selector: 93.06/96.43; S2G: 95.77/97.82; FE2H: 96.32/98.02; Smoothing R3: 96.85/98.32; Beam Retrieval with beam size 1: 97.29/98.55; Beam Retrieval with beam size 2: 97.52/98.68.
    • 2WikiMultihopQA: SA Selector: 98.25/99.13; Beam Retrieval with beam size 1: 99.93/99.96.

    Thus, relative to the strongest listed MuSiQue-Ans baseline, beam size 1 raises retrieval EM from 53.50 to 77.37, while beam size 2 raises it to 79.31. On 2WikiMultihopQA, beam size 1 achieves 99.93 EM and 99.96 F1; the paper describes this as 99.9% retrieval precision.

  7. Knowl 7 — Supervised reader gains from retrieved context

    data/table

    Replacing prior retrieval systems with Beam Retrieval's context improves the supervised multi-hop QA reader. The following are test-set results; all entries are reported as EM/F1, with separate answer and supporting-fact metrics where available.

    • MuSiQue-Ans answer F1/support-passage F1: EE: 40.7/69.4; SA: 52.3/75.2; Ex(EE): 46.4/78.1; Ex(SA): 49.0/80.6; RoHTmix: 63.6/0; Beam Retrieval with beam size 1: 66.9/90.0; Beam Retrieval with beam size 2: 69.2/91.4.
    • HotpotQA answer EM/F1 and supporting-fact EM/F1: HGN: 69.22/82.19 and 62.76/88.47; SAE: 66.92/79.62 and 61.53/86.86; S2G: 70.72/83.53 and 64.30/88.72; FE2H: 71.89/84.44 and 64.98/89.14; Smoothing R3: 72.07/84.34 and 65.44/89.55; Beam Retrieval with beam size 2: 72.69/85.04 and 66.25/90.09.
    • 2WikiMultihopQA answer EM/F1 and supporting-fact EM/F1: CRERC: 69.58/72.33 and 82.86/90.68; NA-Reviewer: 76.73/81.91 and 89.61/94.31; BigBird-base: 74.05/79.68 and 77.14/92.13; Beam Retrieval with beam size 1: 88.47/90.87 and 95.87/98.15.

    The reported Beam Retrieval systems establish the best result among the compared systems on all three benchmarks. On MuSiQue-Ans, the supporting-passage F1 of 91.4 is also reported as close to the human score of 93.9.

  8. Knowl 8 — Beam size, inference cost, and configuration ablations

    empirical result

    Beam size 2 provides the best primary trade-off on MuSiQue-Ans, while larger beams add training cost and eventually reduce retrieval quality. With the base encoder on MuSiQue-Ans, beam sizes 1, 2, 3, and 4 produce respectively EM/F1 values of 74.18/87.46, 75.47/88.27, 74.56/87.84, and 74.43/87.65. Relative training memory is respectively 100%, 119%, 150%, and 194%, and relative training speed is 100%, 58%, 42%, and 36%.

    On HotpotQA development data, inference times and EM are: FE2H, 96.82 ms and 96.35; Smoothing R3, 127.75 ms and 96.85; Beam Retrieval with beam size 1, 124.64 ms and 97.29; and Beam Retrieval with beam size 2, 196.35 ms and 97.52. This motivates beam size 1 for resource-constrained practical use.

    The model is also sensitive to consistency between training and inference beam sizes. Training and reasoning with beam sizes (1,1)(1,1), (2,2)(2,2), and (3,3)(3,3) gives EM/F1 of 74.18/87.46, 75.47/88.27, and 74.56/87.84. Mismatched settings (3,2)(3,2), (3,1)(3,1), and (2,1)(2,1) reduce performance to 74.31/87.84, 74.06/87.67, and 75.13/88.17. Using four separate heads, one per possible hop, gives 72.16/87.04, and using one shared head gives 73.11/87.32, both below the two-head configuration. The paper attributes the advantage of two heads to the different candidate-sequence structures at the first and later hops.

  9. Knowl 9 — Few-shot LLM and transfer-setting improvements

    empirical result

    Beam Retrieval is also evaluated as a context-selection plugin for few-shot language models. On 500 sampled questions per dataset, the paper compares using all candidate passages with using only Beam Retrieval's passages. Answer F1 for GPT-3.5-turbo-16k without/with Beam Retrieval is 32.3/60.9 on MuSiQue-Ans, 65.0/79.0 on HotpotQA, and 51.2/69.2 on 2WikiMultihopQA. For LongChat-13B-16K, the corresponding values are 15.6/40.9, 46.5/66.2, and 31.8/50.6. The retrieved context therefore substantially improves few-shot QA, in some cases approaching supervised-reader performance.

    The method also transfers to the IIRC knowledge-intensive reading-comprehension dataset. After dividing linked Wikipedia pages into ten-sentence passages, the constructed data contains 7,566 training and 954 test examples, with 1--6 relevant passages and 10--25 irrelevant passages per question. Retrieval EM is 57.35 for a one-step retriever, 85.01 for Beam Retrieval with beam size 1, 86.90 with beam size 2, and 86.37 with beam size 3.

    In a full-wiki HotpotQA reranking experiment, MDR alone obtains retrieval EM 65.9 when its top-two chain is evaluated; MDR reranking obtains 81.2, Beam Retrieval reranking obtains 82.2, and reranking with gold chain information obtains 85.6. These results support using Beam Retrieval as a reranker over open-domain candidate chains, although not as a fully independent open-domain retriever.

  10. Knowl 10 — Stated limitations

    limitation

    The paper identifies two main limitations. First, training resource consumption increases rapidly as the beam size grows because each beam creates additional passage hypotheses and encoder computations. Second, Beam Retrieval is difficult to apply independently in open-domain retrieval, where the candidate space is much larger than the fixed candidate sets used in the primary experiments. The authors leave reducing training cost and enabling variable-hop open-domain retrieval as future work.

Coverage note — The appendix's detailed supervised-reader loss formulas and shuffle-operation ablation were omitted because they support downstream implementation and robustness analysis rather than constituting separate load-bearing contributions of the retrieval framework.

References

  1. 1.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  2. 2.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879, Vancouver, Canada. Association for Computational Linguistics.
  3. 3.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Yuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai, Shuo-hang Wang, and Jingjing Liu. 2020. Hierarchical graph network for multi-hop question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8823–8838, Online. Association for Computational Linguistics.
  5. 5.James Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot, and Pradeep Dasigi. 2020. IIRC: A dataset of incomplete information reading comprehension questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1137–1147, Online. Association for Computational Linguistics.
  6. 6.Ruiliu Fu, Han Wang, Xuejun Zhang, Jun Zhou, and Yonghong Yan. 2021. Decomposing complex questions makes multi-hop QA easier and more interpretable. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 169–180, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  7. 7.Ruiliu Fu, Han Wang, Jun Zhou, and Xuejun Zhang. 2022. Na-reviewer: Reviewing the context to improve the error accumulation issue for multi-hop qa. Electronics Letters, 58(6):237–239.
  8. 8.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  9. 9.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2023. Analyzing the effectiveness of the underlying reasoning tasks in multi-hop question answering. In Findings of the Association for Computational Linguistics: EACL 2023, pages 1163–1180, Dubrovnik, Croatia. Association for Computational Linguistics.
  10. 10.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps. CoRR, abs/2011.01060.
  11. 11.Xin-Yi Li, Wei-Jun Lei, and Yu-Bin Yang. 2023. From easy to hard: Two-stage selector and reader for multi-hop question answering. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.
  12. 12.Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  13. 13.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  14. 14.OpenAI. 2023. Gpt-4 technical report.
  15. 15.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc.
  16. 16.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multi-hop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554.
  17. 17.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10014–10037, Toronto, Canada. Association for Computational Linguistics.
  18. 18.Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bowen Zhou. 2020. Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 9073–9080. AAAI Press.
  19. 19.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  20. 20.Bohong Wu, Zhuosheng Zhang, and Hai Zhao. 2021. Graph-free multi-hop reading comprehension: A select-to-guide strategy. CoRR, abs/2107.11823.
  21. 21.Wenhan Xiong, Xiang Li, Srini Iyer, Jingfei Du, Patrick Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel, Douwe Kiela, and Barlas Oguz. 2021. Answering complex open-domain questions with multi-hop dense retrieval. In International Conference on Learning Representations.
  22. 22.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  23. 23.Jiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv, Juanzi Li, Lei Hou, Jiaxin Shi, and Qi Tian. 2023. Reasoning over hierarchical question decomposition tree for explainable question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14556–14570, Toronto, Canada. Association for Computational Linguistics.
  24. 24.Yin Zhangyue, Wang Yuxin, Hu Xiannian, Wu Yiguang, Yan Hang, Zhang Xinyu, Cao Zhao, Huang Xuanjing, and Qiu Xipeng. 2023. Rethinking label smoothing on multi-hop question answering. In Proceedings of the 22nd Chinese National Conference on Computational Linguistics, pages 611–623, Harbin, China. Chinese Information Processing Society of China.
  25. 25.Chen Zhao, Chenyan Xiong, Jordan Boyd-Graber, and Hal Daumé III. 2021. Multi-step reasoning over unstructured text with beam dense retrieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4635–4641, Online. Association for Computational Linguistics.
  26. 26.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.
  27. 27.Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-domain question answering. CoRR, abs/2101.00774.

Citation

MLA
Zhang, J., et al. “End-to-End Beam Retrieval for Multi-Hop Question Answering”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 1718–31, https://doi.org/10.18653/v1/2024.naacl-long.96.
APA
Zhang, J., Zhang, H., Zhang, D., Yong, L., & Huang, S. (2024). End-to-End Beam Retrieval for Multi-Hop Question Answering. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1718–1731. https://doi.org/10.18653/v1/2024.naacl-long.96
Chicago
Zhang, J., H. Zhang, D. Zhang, L. Yong, and S. Huang. 2024. “End-to-End Beam Retrieval for Multi-Hop Question Answering”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1718–31. https://doi.org/10.18653/v1/2024.naacl-long.96.
Harvard
Zhang, J. et al. (2024) “End-to-End Beam Retrieval for Multi-Hop Question Answering”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1718–1731. Available at: https://doi.org/10.18653/v1/2024.naacl-long.96.
Vancouver
1. Zhang J, Zhang H, Zhang D, Yong L, Huang S (2024) End-to-End Beam Retrieval for Multi-Hop Question Answering. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 1718–1731

BibTeX

@inproceedings{zhang-etal-2024-end,
    title = "End-to-End Beam Retrieval for Multi-Hop Question Answering",
    author = "Zhang, Jiahao  and
      Zhang, Haiyang  and
      Zhang, Dongmei  and
      Yong, Liu  and
      Huang, Shen",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.96/",
    doi = "10.18653/v1/2024.naacl-long.96",
    pages = "1718--1731"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/