Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Zhengliang ShiShuo ZhangWeiwei SunShen GaoPengjie RenZhumin ChenZhaochun Ren

article2024ACL92 citations

Proposes GenGround, a framework that counters noisy retrieval in multi-hop question answering by having large language models generate intermediate answers first and then revise them against retrieved evidence, supplemented by a distillation technique that transfers this capability to smaller models.

Listen

Modern artificial intelligence systems increasingly struggle with complex questions that require multi-step reasoning and information from multiple sources. Standard approaches typically follow a retrieve-then-read model, searching external databases first and feeding retrieved texts into language models to generate an answer. However, this workflow often fails when search tools miss key evidence or when retrieved documents contain irrelevant and misleading text that causes the model to generate incorrect conclusions.

The main objective of the article is to demonstrate and evaluate a new framework called GenGround (generate-then-ground), which integrates a model's internal knowledge with external reference documents to solve multi-hop reasoning tasks. The authors also assess a distillation technique designed to transfer these capabilities into smaller, more cost-effective open-source language models.

The authors evaluated the framework across four standard benchmark question-answering datasets—HotpotQA, MuSiQue, 2WikimultihopQA, and StrategyQA—using both proprietary systems (such as ChatGPT) and smaller open-source models (Mistral-7B). The approach alternates between two phases: first, the model deduces a simplified sub-question and generates an immediate answer from memory; second, it grounds and revises that answer by citing evidence from retrieved documents in manageable batches. The distillation method trained a smaller student model on roughly 45,700 synthesized question-and-revision trajectories derived from larger models.

The findings show consistent performance advantages over existing approaches. First, the framework achieved the highest accuracy across all four benchmarks, outperforming established retrieval-augmented baselines such as DSPy and SearChain by 4 to 6 percentage points in accuracy. Second, a fine-grained analysis revealed that the model answered 28.7% of questions correctly using internal knowledge alone and successfully revised another 24.5% using retrieved documents, while maintaining a very low revision error rate of 5.6%. Third, the distillation process significantly improved the performance of smaller models, yielding a 9.8% relative accuracy gain on HotpotQA and a 26.4% gain on MuSiQue over standard prompting. Finally, the framework reduced computational inference costs by processing fewer tokens than competing iterative baselines (averaging approximately 3,542 tokens compared to 7,806 and 8,918 for baselines).

These results indicate that generating an initial hypothesis before grounding it in external documents produces higher accuracy, reduces hallucination risks, and lowers operational computing costs. For organizations deploying conversational AI and automated research tools, this approach provides a more reliable method to handle complex multi-step inquiries without relying exclusively on expensive large models or flawless retrieval pipelines.

Decision-makers should consider adopting a generate-then-ground structure for knowledge-intensive question answering, particularly when dealing with noisy document repositories. Teams facing budget or latency constraints can deploy distilled open-source models, which offer performance comparable to proprietary baselines. Further testing in organization-specific domains is recommended before full production deployment.

The framework's performance remains contingent on two key assumptions: complex inquiries must be successfully decomposed into simpler steps, and external sources must contain valid information to correct initial errors. If the model fails at early question decomposition or if reference databases contain significant misinformation, performance may degrade.

Shi et al (2024).pdf
Cover for Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Abstract

Multi-Hop Question Answering (MHQA) tasks present a significant challenge for large language models (LLMs) due to the intensive knowledge required. Current solutions, like Retrieval-Augmented Generation, typically retrieve potential documents from an external corpus to read an answer. However, the performance of this retrieve-then-read paradigm is constrained by the retriever and the inevitable noise in the retrieved documents. To mitigate these challenges, we introduce a novel generate-then-ground (GenGround) framework, synergizing the parametric knowledge of LLMs and external documents to solve a multi-hop question. GenGround empowers LLMs to alternate two phases until the final answer is derived: (1) formulate a simpler, single-hop question and directly generate the answer; (2) ground the question-answer pair in retrieved documents, amending any wrong predictions in the answer. We also propose an instructional grounding distillation method to generalize our method into smaller models. Extensive experiments conducted on four datasets illustrate the superiority of our method.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Multi-hop Question Answering
  • 2.2 Retrieval-Augmented Generation
  • 3 Generate-then-Ground with LLMs
  • 3.1 Answer Deduction
  • 3.2 Instructional Knowledge Grounding
  • 3.3 Batch Knowledge Ground
  • 4 Generalization with Grounding Distillation
  • 4.1 Synthesize the Training Dataset
  • 4.2 Training Objective
  • 5 Experimental Setup
  • 5.1 Datasets
  • 5.2 Baselines
  • 5.3 Evaluation Metrics
  • 5.4 Implementation Details
  • 6 Experimental Results
  • 6.1 Experimental Results
  • 6.2 Ablation Study
  • 6.3 Case Study
  • 7 Analysis and Discussion
  • 7.1 Result Consistency and Stability
  • 7.2 Knowledge Incorporation
  • 7.3 Qualitative Analysis for Efficiency
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • Appendices
  • Appendix I. Evaluation Metrics Details
  • Appendix II. Case Study
  • Appendix III. Training Dataset for Grounding Distillation
  • Appendix IV. Human Evaluation

Knowls

  1. Knowl 1 — GenGround alternates parametric answer generation with evidence-based revision

    model/method

    Generate-then-Ground (GenGround) answers a multi-hop question by repeatedly using an LLM’s parametric knowledge to propose an answer and then using retrieved documents to check and, when warranted, revise it. Let QQ be the input question and HiH_i the context before iteration ii, containing earlier sub-questions and their revised answers. The LLM uses QQ and HiH_i to formulate a single-hop sub-question qiq_i and directly generate an immediate answer aia_i. A retriever uses qiq_i to find documents; the LLM then grounds the pair (qi,ai)(q_i,a_i) in the documents, citing relevant evidence and producing a revised answer. The revised answer and sub-question are added to the context for the next iteration, so later questions can use earlier corrected information. The process continues until the LLM derives a final answer. If grounding finds no relevant evidence, the immediate answer is retained rather than changed.

  2. Knowl 2 — Batch grounding searches retrieved documents in successive groups

    algorithm

    GenGround limits the effect of long, noisy document lists by grounding an answer against successive batches rather than the entire list at once. Let DiD_i be the ranked documents retrieved for sub-question qiq_i, and let bb be the batch size. For each iteration, process the first bb documents together and ask the LLM to cite evidence and revise the immediate answer. If the LLM finds evidence, accept that revision and end grounding for this iteration. If it signals that no evidence was found, process the next bb documents and repeat. Stop when evidence is found or the retrieved list is exhausted; if all batches produce no evidence, keep the immediate answer unchanged. The experiments used b=3b=3 and retrieved up to 10 documents per question.

  3. Knowl 3 — Instructional grounding distillation trains smaller models to cite and revise

    model/method

    Instructional grounding distillation (IGD) transfers the grounding behavior of ChatGPT to a smaller student model. The authors sampled 50,000 single-hop questions from Natural Questions, paired each with a ground-truth document and noise documents, and generated an immediate answer with a smaller model such as Mistral-7B. ChatGPT then produced a grounding trajectory containing cited evidence and a revised answer. The authors filtered trajectories that lacked evidence, lacked a revised answer, or had a revision misaligned with the ground-truth evidence; the resulting synthetic dataset contained 45,710 examples. For each example, the student is trained to generate the reference-and-revision trajectory given the grounding instruction, question, and document list. If rr is the target trajectory, IGI_G the grounding instruction, qq the question, d∗d^* its ground-truth document, DD the noise-document set, and heta heta the student parameters, the language-model training loss is L_G(\theta)=-\sum_{t=1}^{|r|}\log P_\theta(r_t\mid r_{<t},I_G,q,{d^*}\cup D),$$ where rtr_t is token tt of the target trajectory and r<tr_{<t} is its preceding prefix. The reported Mistral-7B training used DeepSpeed ZeRO, learning rate 5×10−55\times10^{-5}, weight decay 0.010.01, and three NVIDIA A100-PCIE-80GB GPUs; training took 18 hours.

  4. Knowl 4 — GenGround leads the reported multi-hop QA benchmark comparisons

    empirical result

    With GPT-3.5-turbo as the LLM and ColBERTv2 retrieval, GenGround achieved the best reported results across HotpotQA, MuSiQue, 2WikiMultihopQA, and StrategyQA. Scores below are F1/accuracy/semantic accuracy, except StrategyQA, for which only accuracy is reported. GenGround scored 52.26/47.27/55.73 on HotpotQA, 27.36/20.24/24.77 on MuSiQue, and 50.21/45.61/48.58 on 2WikiMultihopQA; its StrategyQA accuracy was 77.12. For comparison, DSPy scored 47.80/42.43/50.07, 20.11/13.40/17.40, and 44.77/43.43/45.43 on the first three datasets, and 71.78 accuracy on StrategyQA. SearChain, which was evaluated only with accuracy and semantic accuracy because it generated long-form answers, scored 46.76/48.12, 17.07/20.45, and 42.14/46.27 on the first three datasets, and 76.95 accuracy on StrategyQA. Accuracy measures whether the reference answer is contained in the model answer; semantic accuracy was judged by GPT-3.5-turbo-instruct.

  5. Knowl 5 — Benchmark ablations show that deduction, grounding, and batching each contribute

    empirical result

    Ablations with GPT-3.5-turbo on HotpotQA and StrategyQA show lower performance when any of GenGround’s three stages is removed. Each entry gives HotpotQA F1/accuracy/semantic accuracy followed by StrategyQA accuracy; the parenthetical values are the reported absolute drops from full GenGround. Removing answer deduction yielded 42.65 (↓9\downarrow 9)/41.08 (↓6\downarrow 6)/43.14 (↓12\downarrow 12), and 66.51 (↓10\downarrow 10). Removing instructional grounding yielded 45.14 (↓7\downarrow 7)/41.35 (↓4\downarrow 4)/43.23 (↓5\downarrow 5), and 72.34 (↓5\downarrow 5). Removing batch grounding yielded 47.27 (↓5\downarrow 5)/45.03 (↓2\downarrow 2)/51.19 (↓4\downarrow 4), and 71.72 (↓5\downarrow 5). The deduction ablation answered the original multi-hop question directly before retrieval and revision; the grounding ablation used retrieved documents to generate answers without the evidence-citing revision phase; the batch ablation presented the long retrieved list without successive grouping.

  6. Knowl 6 — Grounding distillation improves Mistral-7B multi-hop QA accuracy

    empirical result

    With Mistral-7B and ColBERTv2 retrieval, prompting the model to use GenGround without distillation achieved accuracies of 38.31 on HotpotQA, 11.34 on MuSiQue, and 29.45 on 2WikiMultihopQA. After IGD training, the respective accuracies rose to 42.08, 14.37, and 32.69, an average absolute gain of 3.35 points across the three datasets. The distilled model also exceeded the strongest listed baseline on each dataset: DSPy scored 36.41 on HotpotQA and 28.31 on 2WikiMultihopQA, while GRG with decomposition scored 9.34 on MuSiQue. The paper reports relative improvements over the prompted model of 9.84% on HotpotQA and 26.4% on MuSiQue.

  7. Knowl 7 — GenGround retains its lead with BM25 and Google Search retrieval

    empirical result

    In experiments using GPT-3.5-turbo, replacing ColBERTv2 with either BM25 or Google Search did not remove GenGround’s accuracy advantage on HotpotQA, MuSiQue, or 2WikiMultihopQA. With BM25, GenGround scored 42.21, 18.32, and 40.32, respectively; the strongest competing scores for those datasets were 41.31, 15.62, and 38.84, all from GRG with decomposition. With Google Search, GenGround scored 48.95, 21.54, and 46.87; the strongest competing scores were 46.86 on HotpotQA and 20.71 on MuSiQue from DSPy, and 43.21 on 2WikiMultihopQA from GRG with decomposition. These comparisons use accuracy and indicate that GenGround’s performance advantage held across the tested retrievers.

  8. Knowl 8 — Multi-hop QA evaluation uses four benchmarks and controlled retrieval conditions

    experimental setup

    The evaluation covered HotpotQA, MuSiQue, 2WikiMultihopQA, and StrategyQA. The paper reports randomly sampling 1,400 questions for the benchmark evaluation following prior setups. GPT-3.5-turbo was the backbone for GenGround and the baselines, with decoding temperature 0; Mistral-7B was used for open-model comparisons. The main retriever was ColBERTv2, which returned the top 10 documents, and the batch-grounding size was 3. The retrieval corpus was Wikipedia 2017 for HotpotQA and a large-scale passage collection built on Wikipedia 2018 for the other open-domain QA benchmarks; BM25 and Google Search were also tested. Evaluation used accuracy, token-level F1, and semantic accuracy judged by GPT-3.5-turbo-instruct; StrategyQA results were reported with accuracy.

  9. Knowl 9 — Trajectory analysis finds both direct answering and correction contribute to success

    empirical result

    Three annotators examined 100 randomly sampled HotpotQA cases to assess how GenGround’s answer deduction and grounding contributed to correctness. They classified a case as a success if the model answered correctly immediately or corrected an initially incorrect answer; as a failure if it remained wrong; and as an error if grounding changed a correct immediate answer into an incorrect one. The overall success rate was 53.2%: 28.7% of cases were answered correctly before grounding and 24.5% were initially wrong but corrected with external evidence. The failure rate was 41.2%, while the error rate was 5.6%. In a separate human evaluation of 120 randomly selected cases from the four benchmarks, three annotators assessed correctness on a three-level scale, with at least two rating each case and a third resolving disagreements. GenGround scored 52.75, compared with 49.71 for GRG, 51.24 for DSPy, and 46.30 for SearChain; inter-annotator agreement was Cohen’s κ=0.73\kappa=0.73.

  10. Knowl 10 — GenGround depends on useful decomposition and corrective evidence

    limitation

    The authors identify three conditions that can limit GenGround. First, the method depends on the LLM producing a meaningful immediate answer, which may vary with the task and restrict applicability to domains where the initial answer is not useful. Second, it assumes complex questions can be decomposed into simpler questions, although decomposition itself is difficult and is not fully addressed by the framework. Third, grounding can correct an initially non-factual answer only when external documents contain the needed information and are not themselves misleading; missing evidence or misinformation can therefore compromise the method.

Coverage note — The token-consumption comparison and repeated-run stability analysis are omitted as supplementary efficiency and consistency diagnostics; the central framework, distillation, benchmark findings, ablations, trajectory analysis, and stated limitations are included.

References

  1. 1.Abdelrahman Abdallah and Adam Jatowt. 2023. Generator-retriever-generator: A novel approach to open-domain question answering. arXiv preprint arXiv:2307.11278.
  2. 2.Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024. Evaluating correctness and faithfulness of instruction-following models for question answering. In Transactions of the Association for Computational Linguistics: TACL.
  3. 3.Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, et al. 2023. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence: AAAI.
  4. 4.Liang Chen, Yang Deng, Yatao Bian, Zeyu Qin, Bingzhe Wu, Tat-Seng Chua, and Kam-Fai Wong. 2023. Beyond factuality: A comprehensive evaluation of large language models as knowledge generators. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP.
  5. 5.Zhangyin Feng, Xiaocheng Feng, Dezhi Zhao, Maojin Yang, and Bing Qin. 2023. Retrieval-generation synergy augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP.
  6. 6.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al. 2023a. Rarr: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: ACL.
  7. 7.Shen Gao, Zhengliang Shi, Minghang Zhu, Bowen Fang, Xin Xin, Pengjie Ren, Zhumin Chen, and Jun Ma. 2024. Confucius: Iterative tool learning from introspection feedback by easy-to-difficult curriculum. In Proceedings of the AAAI Conference on Artificial Intelligence.
  8. 8.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023b. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
  9. 9.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies. Transactions of the Association for Computational Linguistics (TACL).
  10. 10.Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics.
  11. 11.Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023a. Mistral 7b. ArXiv, abs/2310.06825.
  12. 12.Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023b. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP.
  13. 13.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP. arXiv preprint arXiv:2212.14024.
  14. 14.Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714.
  15. 15.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  16. 16.Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems.
  17. 17.Ruosen Li and Xinya Du. 2023. Leveraging structured information for explainable multi-hop question answering and reasoning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP.
  18. 18.Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting for retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP.
  19. 19.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: ACL, pages 9802–9822.
  20. 20.Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. 2022. A survey on multi-hop question answering and generation. arXiv preprint arXiv:2204.09140.
  21. 21.Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, et al. 2023. Webcpm: Interactive web search for chinese long-form question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: ACL.
  22. 22.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In SIGKDD, page 3505–3506.
  23. 23.Ruiyang Ren, Yuhao Wang, Yingqi Qu, Wayne Xin Zhao, Jing Liu, Hao Tian, Hua Wu, Ji-Rong Wen, and Haifeng Wang. 2023. Investigating the factual knowledge boundary of large language models with retrieval augmentation. arXiv preprint arXiv:2307.11019.
  24. 24.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  25. 25.Keshav Santhanam, O. Khattab, Jon Saad-Falcon, Christopher Potts, and Matei A. Zaharia. 2021. Colbertv2: Effective and efficient retrieval via lightweight late interaction. In North American Chapter of the Association for Computational Linguistics.
  26. 26.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  27. 27.Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy. In Findings of the Association for Computational Linguistics: EMNLP 2023.
  28. 28.Zhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang, Pengjie Ren, and Zhaochun Ren. 2023. Rade: Reference-assisted dialogue evaluation for open-domain dialogue. ArXiv, abs/2309.08156.
  29. 29.Weiwei Sun, Zhengliang Shi, Shen Gao, Pengjie Ren, Maarten de Rijke, and Zhaochun Ren. 2023a. Contrastive learning reduces hallucination in conversations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 13618–13626.
  30. 30.Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023b. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agent. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.
  31. 31.Yixuan Tang, Hwee Tou Ng, and Anthony Tung. 2021. Do multi-hop question answering systems know how to answer the single-hop sub-questions? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: ACL, pages 3244–3249. Association for Computational Linguistics.
  32. 32.Nandan Thakur, Luiz Bonifacio, Xinyu Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, et al. 2023a. Nomiracl: Knowing when you don’t know for robust multilingual retrieval-augmented generation. arXiv preprint arXiv:2312.11361.
  33. 33.Nandan Thakur, Luiz Bonifacio, Xinyu Crystina Zhang, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Boxing Chen, Mehdi Rezagholizadeh, and Jimmy J. Lin. 2023b. Nomiracl: Knowing when you don’t know for robust multilingual retrieval-augmented generation. ArXiv, abs/2312.11361.
  34. 34.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop questions via single-hop question composition. In Transactions of the Association for Computational Linguistics: TACL.
  35. 35.Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query expansion with large language models. corr abs/2303.07678 (2023). In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.
  36. 36.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022a. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  37. 37.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Huai hsin Chi, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. ArXiv, abs/2203.11171.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Huai hsin Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903.
  39. 39.Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2023. Adaptive chameleon or stubborn sloth: Unraveling the behavior of large language models in knowledge conflicts. arXiv preprint arXiv:2305.13300.
  40. 40.Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2024. Search-in-the-chain: Towards accurate, credible and traceable large language models for knowledge-intensive tasks. In WWW.
  41. 41.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  42. 42.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.
  43. 43.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  44. 44.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2022. Generate rather than retrieve: Large language models are strong context generators. arXiv preprint arXiv:2209.10063.
  45. 45.Wenhao Yu, Hongming Zhang, Xiaoman Pan, Kaixin Ma, Hongwei Wang, and Dong Yu. 2023. Chain-of-note: Enhancing robustness in retrieval-augmented language models. arXiv preprint arXiv:2311.09210.
  46. 46.Jiahao Zhang, Haiyang Zhang, Dongmei Zhang, Yong Liu, and Shen Huang. 2024. End-to-end beam retrieval for multi-hop question answering. In 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics.
  47. 47.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219.
  48. 48.Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023. Verify-and-edit: A knowledge-enhanced chain-of-thought framework. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada. Association for Computational Linguistics.
  49. 49.Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey. arXiv preprint arXiv:2308.07107.

Citation

MLA
Shi, Z., et al. “Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7339–53, https://doi.org/10.18653/v1/2024.acl-long.397.
APA
Shi, Z., Zhang, S., Sun, W., Gao, S., Ren, P., Chen, Z., & Ren, Z. (2024). Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7339–7353. https://doi.org/10.18653/v1/2024.acl-long.397
Chicago
Shi, Z., S. Zhang, W. Sun, et al. 2024. “Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7339–53. https://doi.org/10.18653/v1/2024.acl-long.397.
Harvard
Shi, Z. et al. (2024) “Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7339–7353. Available at: https://doi.org/10.18653/v1/2024.acl-long.397.
Vancouver
1. Shi Z, Zhang S, Sun W, Gao S, Ren P, Chen Z, Ren Z (2024) Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7339–7353

BibTeX

@inproceedings{shi-etal-2024-generate,
    title = "Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering",
    author = "Shi, Zhengliang  and
      Zhang, Shuo  and
      Sun, Weiwei  and
      Gao, Shen  and
      Ren, Pengjie  and
      Chen, Zhumin  and
      Ren, Zhaochun",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.397/",
    doi = "10.18653/v1/2024.acl-long.397",
    pages = "7339--7353"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/