Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
Zhengliang ShiShuo ZhangWeiwei SunShen GaoPengjie RenZhumin ChenZhaochun Ren
Proposes GenGround, a framework that counters noisy retrieval in multi-hop question answering by having large language models generate intermediate answers first and then revise them against retrieved evidence, supplemented by a distillation technique that transfers this capability to smaller models.
Modern artificial intelligence systems increasingly struggle with complex questions that require multi-step reasoning and information from multiple sources. Standard approaches typically follow a retrieve-then-read model, searching external databases first and feeding retrieved texts into language models to generate an answer. However, this workflow often fails when search tools miss key evidence or when retrieved documents contain irrelevant and misleading text that causes the model to generate incorrect conclusions.
The main objective of the article is to demonstrate and evaluate a new framework called GenGround (generate-then-ground), which integrates a model's internal knowledge with external reference documents to solve multi-hop reasoning tasks. The authors also assess a distillation technique designed to transfer these capabilities into smaller, more cost-effective open-source language models.
The authors evaluated the framework across four standard benchmark question-answering datasets—HotpotQA, MuSiQue, 2WikimultihopQA, and StrategyQA—using both proprietary systems (such as ChatGPT) and smaller open-source models (Mistral-7B). The approach alternates between two phases: first, the model deduces a simplified sub-question and generates an immediate answer from memory; second, it grounds and revises that answer by citing evidence from retrieved documents in manageable batches. The distillation method trained a smaller student model on roughly 45,700 synthesized question-and-revision trajectories derived from larger models.
The findings show consistent performance advantages over existing approaches. First, the framework achieved the highest accuracy across all four benchmarks, outperforming established retrieval-augmented baselines such as DSPy and SearChain by 4 to 6 percentage points in accuracy. Second, a fine-grained analysis revealed that the model answered 28.7% of questions correctly using internal knowledge alone and successfully revised another 24.5% using retrieved documents, while maintaining a very low revision error rate of 5.6%. Third, the distillation process significantly improved the performance of smaller models, yielding a 9.8% relative accuracy gain on HotpotQA and a 26.4% gain on MuSiQue over standard prompting. Finally, the framework reduced computational inference costs by processing fewer tokens than competing iterative baselines (averaging approximately 3,542 tokens compared to 7,806 and 8,918 for baselines).
These results indicate that generating an initial hypothesis before grounding it in external documents produces higher accuracy, reduces hallucination risks, and lowers operational computing costs. For organizations deploying conversational AI and automated research tools, this approach provides a more reliable method to handle complex multi-step inquiries without relying exclusively on expensive large models or flawless retrieval pipelines.
Decision-makers should consider adopting a generate-then-ground structure for knowledge-intensive question answering, particularly when dealing with noisy document repositories. Teams facing budget or latency constraints can deploy distilled open-source models, which offer performance comparable to proprietary baselines. Further testing in organization-specific domains is recommended before full production deployment.
The framework's performance remains contingent on two key assumptions: complex inquiries must be successfully decomposed into simpler steps, and external sources must contain valid information to correct initial errors. If the model fails at early question decomposition or if reference databases contain significant misinformation, performance may degrade.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational RAG paper establishes the retrieve-and-generate architecture that GenGround revises by generating an initial answer before grounding it in retrieved evidence.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). FLARE develops retrieval triggered by a model’s uncertainty during generation, providing a key precedent for GenGround’s adaptive alternation between internal answers and external evidence.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Its analysis of failures in composing known facts, and its self-ask decomposition method, motivate GenGround’s use of simpler sub-questions for multi-hop reasoning.
- Paper: Successive Prompting for Decomposing Complex Questions, Dheeru Dua et al. (2022). Successive Prompting shows how decomposing a complex question into sequential question-answer steps supports the sub-question reasoning GenGround relies on.
- Paper: Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks, Minki Kang et al. (2023). KARD establishes how retrieved evidence can support reasoning distillation into smaller models, a direct precursor to GenGround’s distillation of grounded QA trajectories.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). Search-R1 extends multi-hop QA beyond GenGround’s generate-then-ground workflow by training models to interleave reasoning and search through reinforcement learning.
