Merging Generated and Retrieved Knowledge for Open-Domain QA
Yunxiang ZhangMuhammad KhalifaLajanugen LogeswaranMoontae LeeHonglak LeeLu Wang
Proposes a compatibility-oriented framework that pairs LLM-generated texts with retrieved documents to resolve knowledge conflicts and improve open-domain question answering accuracy.
Modern open-domain question answering systems increasingly combine factual information retrieved from external text sources with relevant knowledge generated by large language models. While combining these two sources can broaden background coverage, language models frequently generate fabricated statements—known as hallucinations—that contradict factual retrieved contexts. Standard question answering architectures tend to favor generated text even when it contains errors, causing systems to produce incorrect answers when these knowledge conflicts arise.
The article introduces and evaluates a compatibility-oriented knowledge merging framework called COMBO. The primary objective is to demonstrate that pairing and prioritizing mutually supportive passages from retrieved documents and language models allows reading systems to extract correct answers more reliably while mitigating the adverse effects of knowledge conflicts.
To evaluate this framework, the authors conducted supervised experiments across four benchmark question answering datasets representing single-step and complex multi-step reasoning tasks. The approach automatically generates pseudo-labels—termed silver labels—by measuring how a reader model's accuracy changes when specific text passages are added or removed, eliminating the need for manual human annotation. These labels train two neural classifiers to score passage evidentiality and consistency. A bipartite matching algorithm then pairs generated and retrieved passages to maximize overall mutual support before passing the ranked pairs into a sequence-to-sequence reader model.
The findings show that COMBO consistently improves accuracy over baseline approaches on single-step question benchmarks. On the tested single-hop datasets, COMBO outperformed direct passage merging by up to 1.9 exact match points and achieved an average improvement of 1.3 points. The performance advantage became significantly wider in scenarios with severe knowledge conflicts, where baseline methods suffered substantial accuracy drops. However, on the complex multi-step benchmark, the method achieved only marginal gains on bridge questions and showed no overall improvement, as multi-step comparison tasks frequently arrived at correct answers despite underlying hallucinations.
These results indicate that resolving conflicting evidence before answer generation improves accuracy with modest operational overhead. The framework requires approximately 23% more memory and 27% more training time than direct merging, making it a cost-effective upgrade for existing retrieval-augmented systems. By guiding the reader to attend to mutually compatible evidence, the method lowers the operational risk of deploying hallucination-prone language models in automated answering systems.
Organizations implementing retrieval-augmented language models should adopt compatibility scoring and structured passage pairing to insulate systems against conflicting generated data. For future development, technical teams should explore using in-context prompt scoring to eliminate weakly supervised classifier training entirely, and develop fine-grained compatibility mechanisms specifically tailored to complex multi-step reasoning chains.
Confidence in these findings is high for standard single-step questions, supported by human verification confirming 78% classifier accuracy. Readers should exercise caution when applying the method to complex multi-step reasoning or high-stakes domains such as healthcare and finance, where complete elimination of model hallucinations remains unguaranteed.
- Paper: Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks, Akari Asai et al. (2022). Its silver-label method uses leave-one-out passage removal to identify evidence, a direct precursor to COMBO’s reader-impact-based evidentiality labels.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational RAG work establishes the retriever–generator architecture that COMBO modifies by resolving conflicts among retrieved and generated passages.
- Paper: Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models, Fei Wang et al. (2025). ASTUTE RAG continues the knowledge-conflict problem by testing a training-free way to reconcile retrieved evidence with a model’s internal knowledge.
- Paper: Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering, Zhengliang Shi et al. (2024). GenGround takes the evidence-integration challenge into multi-hop QA, generating answers first and then grounding and revising them against retrieved documents.
- Paper: Bridging the Preference Gap between Retrievers and LLMs, Zixuan Ke et al. (2024). BGM extends passage selection for RAG by learning a bridge that chooses and reorders retrieved evidence to better suit the generator.
- Paper: Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models, Wenhao Yu et al. (2024). Chain-of-Note extends retrieval robustness by having models assess each document’s relevance and credibility before answering.
