Merging Generated and Retrieved Knowledge for Open-Domain QA

Yunxiang ZhangMuhammad KhalifaLajanugen LogeswaranMoontae LeeHonglak LeeLu Wang

article2023EMNLP59 citations

Proposes a compatibility-oriented framework that pairs LLM-generated texts with retrieved documents to resolve knowledge conflicts and improve open-domain question answering accuracy.

Listen

Modern open-domain question answering systems increasingly combine factual information retrieved from external text sources with relevant knowledge generated by large language models. While combining these two sources can broaden background coverage, language models frequently generate fabricated statements—known as hallucinations—that contradict factual retrieved contexts. Standard question answering architectures tend to favor generated text even when it contains errors, causing systems to produce incorrect answers when these knowledge conflicts arise.

The article introduces and evaluates a compatibility-oriented knowledge merging framework called COMBO. The primary objective is to demonstrate that pairing and prioritizing mutually supportive passages from retrieved documents and language models allows reading systems to extract correct answers more reliably while mitigating the adverse effects of knowledge conflicts.

To evaluate this framework, the authors conducted supervised experiments across four benchmark question answering datasets representing single-step and complex multi-step reasoning tasks. The approach automatically generates pseudo-labels—termed silver labels—by measuring how a reader model's accuracy changes when specific text passages are added or removed, eliminating the need for manual human annotation. These labels train two neural classifiers to score passage evidentiality and consistency. A bipartite matching algorithm then pairs generated and retrieved passages to maximize overall mutual support before passing the ranked pairs into a sequence-to-sequence reader model.

The findings show that COMBO consistently improves accuracy over baseline approaches on single-step question benchmarks. On the tested single-hop datasets, COMBO outperformed direct passage merging by up to 1.9 exact match points and achieved an average improvement of 1.3 points. The performance advantage became significantly wider in scenarios with severe knowledge conflicts, where baseline methods suffered substantial accuracy drops. However, on the complex multi-step benchmark, the method achieved only marginal gains on bridge questions and showed no overall improvement, as multi-step comparison tasks frequently arrived at correct answers despite underlying hallucinations.

These results indicate that resolving conflicting evidence before answer generation improves accuracy with modest operational overhead. The framework requires approximately 23% more memory and 27% more training time than direct merging, making it a cost-effective upgrade for existing retrieval-augmented systems. By guiding the reader to attend to mutually compatible evidence, the method lowers the operational risk of deploying hallucination-prone language models in automated answering systems.

Organizations implementing retrieval-augmented language models should adopt compatibility scoring and structured passage pairing to insulate systems against conflicting generated data. For future development, technical teams should explore using in-context prompt scoring to eliminate weakly supervised classifier training entirely, and develop fine-grained compatibility mechanisms specifically tailored to complex multi-step reasoning chains.

Confidence in these findings is high for standard single-step questions, supported by human verification confirming 78% classifier accuracy. Readers should exercise caution when applying the method to complex multi-step reasoning or high-stakes domains such as healthcare and finance, where complete elimination of model hallucinations remains unguaranteed.

arXiv: 2310.14393yunx-z/COMBO
Cover for Merging Generated and Retrieved Knowledge for Open-Domain QA

Abstract

Open-domain question answering (QA) systems are often built with retrieval modules. However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage. Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance. Yet, LLMs tend to “hallucinate” content that conflicts with the retrieved knowledge. Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented Knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information. Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels. Then a Fusion-in-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer. Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks. Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts.¹

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • Parametric and Retrieved Knowledge for QA.
  • 3 Method
  • 3.1 Defining Compatibility
  • 3.2 Mining Silver Labels
  • 3.3 Matching Retrieved and LLM-generated Passages into Pairs
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Results and Analysis
  • 5 Conclusion and Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Preliminaries
  • A.1 Knowledge Conflicts Prediction
  • A.2 Retrieve/Generate-then-read Open-domain QA Framework
  • A.3 Extension of Compatibility Definition from Single-hop to Multi-hop QA
  • A.4 Validating Assumptions behind the Compatibility Definition
  • B Dataset Details
  • C Implementation Details
  • C.1 Silver Evidentiality Label Mining
  • D Additional Experimental Results and Analysis
  • D.1 More Results
  • D.2 Comparison Questions from HotpotQA are Less Sensitive to Hallucinations in LLM-generated Passages
  • D.3 Distribution of Pairwise Relationships
  • D.4 Analyzing Attention Distributions Over Different Types of Passage Pairs

Knowls

  1. Knowl 1 — Compatibility means shared answer-supporting evidence

    definition

    For a question QQ, let lil_i be an LLM-generated passage and rjr_j a passage retrieved from an external corpus. COMBO assumes retrieved passages are factually reliable but may be irrelevant to QQ, while LLM-generated passages usually provide plausible, question-relevant evidence but may be factually wrong. Under these assumptions, (li,rj)(l_i,r_j) is compatible when both passages contain correct evidence supporting the answer to QQ. It is conflicting when the retrieved passage contains correct evidence but the generated passage contains incorrect evidence; it is non-evidential when the retrieved passage lacks correct evidence.

  2. Knowl 2 — Compatibility decomposes into retrieved evidence and conditional agreement

    equation

    For a question QQ, let ELE_L denote the event that an LLM-generated passage lil_i contains correct evidence for QQ, and let ERE_R denote the event that a retrieved passage rjr_j contains correct evidence. COMBO defines the compatibility score as the probability that both events hold and factors it into retrieved-passage evidentiality and generated-passage evidence conditional on retrieved evidence:

    cijQ=P(EL,ER)=P(ER)P(EL∣ER).c_{ij}^{Q}=P(E_L,E_R)=P(E_R)P(E_L\mid E_R).

    Both factors, and therefore the compatibility score, are probabilities in [0,1][0,1]. An evidentiality discriminator DED_E estimates P(ER)P(E_R) from (Q,rj)(Q,r_j); a consistency discriminator DCD_C estimates P(EL∣ER)P(E_L\mid E_R) from (Q,li,rj)(Q,l_i,r_j). The factorization reflects the assumption that generated passages are already likely to offer plausible evidence, so their factuality is judged relative to evidence in the retrieved passage.

  3. Knowl 3 — Silver labels train the evidentiality and consistency discriminators

    model/method

    COMBO mines training labels from a fine-tuned Fusion-in-Decoder (FiD) reader rather than requiring manual labels. For evidentiality, the reader receives the top 10 retrieved passages for a question. A retrieved passage is labeled evidential if the reader answers correctly with all passages but incorrectly after that passage is removed. It is labeled non-evidential if the full input yields an incorrect answer but removing that passage makes the answer correct.

    For a candidate generated–retrieved pair (li,rj)(l_i,r_j), the reader is evaluated on four inputs: (I) all retrieved passages; (II) all retrieved passages except rjr_j; (III) all retrieved passages plus lil_i; and (IV) all retrieved passages except rjr_j, plus lil_i. The pair is labeled consistent if the answer is wrong on II and correct on I, III, and IV. It is labeled conflicting if the answer is correct on I and wrong on II, III, and IV. The method needs predictions on III and IV only when I is correct and II is incorrect. The resulting silver labels supervise the binary discriminators.

    In a manual check on 150 sampled pairs from NaturalQuestions, the discriminator predictions over compatible, conflicting, and non-evidential categories achieved 78% overall accuracy.

  4. Knowl 4 — Compatibility-guided matching pairs passages by maximum total score

    algorithm

    For a question QQ, let MM be the number of generated passages and NN the number of retrieved passages. COMBO scores every possible pair using the consistency discriminator, but sets the score to zero unless the evidentiality discriminator assigns the retrieved passage probability greater than 0.50.5:

    cijQ=P(EL∣ER) 1[P(ER)>0.5].c_{ij}^{Q}=P(E_L\mid E_R)\,\mathbf{1}[P(E_R)>0.5].

    Here ELE_L and ERE_R mean that the generated and retrieved passages, respectively, contain correct evidence for QQ; P(EL∣ER)P(E_L\mid E_R) is estimated by DCD_C, P(ER)P(E_R) by DED_E, and 1[⋅]\mathbf{1}[\cdot] is 1 when its condition holds and 0 otherwise. This evidentiality cutoff prevents non-evidential retrieved passages from receiving positive pair scores.

    Treat passages as nodes in a complete bipartite graph, with generated passages on one side and retrieved passages on the other, and use cijQc_{ij}^{Q} as each edge weight. Select a maximum-weight perfect matching, thereby pairing all passages while maximizing the sum of pair scores; the Hungarian algorithm solves this assignment problem in polynomial time. COMBO matches rather than discards pairs. The experiments use 10 passages from each source per question; in the passage-count scaling experiments, generated passages are duplicated when there are fewer of them than retrieved passages so the sets can be matched.

  5. Knowl 5 — FiD reads matched passages as ordered pairs ranked by compatibility

    model/method

    COMBO adapts a Fusion-in-Decoder reader so each matched generated–retrieved pair is encoded as one input unit, independently of the other pairs. The question and the two passages are marked with the prefixes “question:”, “generated passage:”, and “retrieved passage:”. Within every pair, the generated passage comes first and the retrieved passage second, enabling encoder self-attention across the two sources while keeping factual retrieved evidence in a consistent position. The reader receives the pairs sorted from highest to lowest compatibility score and generates the answer from their encoded representations. Thus the reader can jointly use evidence within a pair and is given a ranking that prioritizes pairs the discriminators judge more compatible.

  6. Knowl 6 — Benchmark design combines retrieved and LLM-generated contexts

    experimental setup

    COMBO was evaluated in a fully supervised setting on NaturalQuestions Open, TriviaQA unfiltered, and WebQuestions for single-hop QA, and HotpotQA full wiki for multi-hop QA. DPR retrieved passages for the single-hop datasets and MDR retrieved passages for HotpotQA; the retrievers were not further fine-tuned and no gold passages were supplied. Single-hop generated passages came from InstructGPT generations released by prior work; HotpotQA passage chains were generated by querying ChatGPT (gpt-3.5-turbo). The reader was FiD-large (770M parameters), and the standard two-source comparisons used 10 retrieved and 10 generated passages per question. The discriminators used RoBERTa-large for single-hop QA and DeBERTa-large for multi-hop QA, with separate discriminator models trained for each dataset; WebQuestions discriminator training was warmed up using silver-label data from the other two single-hop datasets. Exact Match (EM) was the principal metric. Two-source experiments were run with three random seeds and report their average.

  7. Knowl 7 — COMBO improves single-hop EM but shows little overall gain on HotpotQA

    data/table

    The results compare retrieval-only and generation-only readers with direct merging, random passage matching, and COMBO. EM is reported on the NaturalQuestions, TriviaQA, and WebQuestions test sets, and on the HotpotQA development set for all questions and for bridge questions. All methods use the same FiD-large reader; two-source results are averages over three random seeds.

    Method NQ test TQA test WebQ test HQA all Q dev HQA bridge Q dev
    Retrieved passage only 46.7 61.9 48.1 59.9 55.4
    LLM passage only 40.3 67.8 51.5 42.6 35.9
    Direct merging 52.7 74.2 51.1 61.6 57.8
    Random matching 53.3 74.2 51.6 61.5 57.7
    COMBO 54.2 74.6 53.0 61.6 58.0

    COMBO exceeds direct merging by 1.5 EM on NaturalQuestions, 0.4 on TriviaQA, and 1.9 on WebQuestions, an average gain of about 1.3 points across the three single-hop test sets. On HotpotQA, COMBO matches direct merging at 61.6 EM for all questions and is 0.2 points higher on bridge questions; the paper therefore reports no overall HotpotQA improvement and characterizes the bridge-question gain as minor.

  8. Knowl 8 — Ablations support the main COMBO design choices

    empirical result

    On the NaturalQuestions and TriviaQA development sets, COMBO scores 52.3 and 73.9 EM, respectively. Removing individual components lowers performance: without the evidentiality discriminator, scores are 51.8 and 73.5; without pairwise input, 51.5 and 73.4; without sorting pairs by compatibility, 51.7 and 73.7; without fixing generated-before-retrieved order within each pair, 50.8 and 73.3; without the evidentiality cutoff, 51.7 and 73.6; and without maximum-weight matching, 51.6 and 73.6.

    The pairwise-input ablation linearizes matched passages instead of encoding each pair together; the ordering ablation shuffles which source appears first; the sorting ablation shuffles the order of pairs; removing the cutoff uses the product of the two discriminator probabilities; and removing optimal matching replaces it with greedy matching. These development-set comparisons support the usefulness of each tested component, with the largest reported drop occurring when the fixed within-pair order is removed.

  9. Knowl 9 — COMBO's advantage over direct merging tends to grow with answer-based conflict rate

    data/table

    On the NaturalQuestions development set, the paper defines conflict rate as the proportion of all generated–retrieved passage pairs in which the retrieved passage contains the gold answer string but the generated passage does not. Let NN be the number of retrieved passages, MM the number of generated passages, NAN_A the number of retrieved passages containing the gold answer string AA, and MAM_A the number of generated passages containing AA. Then the measured rate is NA(M−MA)/(NM)N_A(M-M_A)/(NM). This is an answer-string-based proxy for conflict, not a direct annotation of semantic contradiction.

    Across conflict-rate bins, the subset percentage and EM scores for retrieval-only, direct merging, and COMBO are: 0–0.1, 56.2%, 41.6, 47.7, and 48.5 (+0.8 over direct merging); 0.1–0.2, 22.7%, 45.1, 53.1, and 53.3 (+0.2); 0.2–0.3, 15.5%, 52.5, 59.0, and 59.9 (+0.9); 0.3–0.4, 1.8%, 62.5, 65.0, and 67.5 (+2.5); 0.4–0.5, 2.2%, 57.4, 64.0, and 66.0 (+2.0); and 0.5–1.0, 1.6%, 61.8, 56.6, and 61.0 (+4.4). The table shows a generally larger COMBO advantage over direct merging in higher-conflict bins, although in the highest bin COMBO remains below retrieval-only EM (61.0 versus 61.8).

  10. Knowl 10 — Scope and operational limitations of COMBO

    limitation

    The experiments evaluate COMBO only for open-domain QA, using one generator and one retriever per dataset and a FiD-based generative reader; the results do not establish performance with other task types or model families. The paper reports no overall gain over direct merging on HotpotQA and notes that compatibility modeling may need to represent relations between evidence hops more finely for multi-hop questions. The amount of mined silver data can also be inadequate for dataset-specific discriminator training on small datasets. Finally, on NaturalQuestions, training the FiD reader with COMBO required approximately 23% more GPU memory and 27% more training time than direct merging. The authors also note that if all generated passages are free of hallucinations or all retrieved passages lack correct evidence, the compatibility scores become uniform and matching reduces to random pairing.

Coverage note — The supplementary same-answer oracle-matching comparison and passage-count and attention-distribution analyses are omitted because they are secondary diagnostics; the main benchmark, ablations, conflict-stratified results, and limitations are retained.

References

  1. 1.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Ábrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, and et al. 2023. Palm 2 technical report. CoRR, abs/2305.10403.
  2. 2.Akari Asai, Matt Gardner, and Hannaneh Hajishirzi. 2022. Evidentiality-guided generation for knowledge-intensive NLP tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 2226–2243. Association for Computational Linguistics.
  3. 3.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, pages 1533–1544. ACL.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  5. 5.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1870–1879. Association for Computational Linguistics.
  6. 6.Hung-Ting Chen, Michael J. Q. Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 2292–2307. Association for Computational Linguistics.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  8. 8.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. CoRR, abs/2208.14271.
  9. 9.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  10. 10.Martin Fajcik, Martin Docekal, Karel Ondrej, and Pavel Smrz. 2021. R2-D2: A modular baseline for open-domain question answering. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 854–870. Association for Computational Linguistics.
  11. 11.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR.
  12. 12.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  13. 13.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Towards unsupervised dense information retrieval with contrastive learning. CoRR, abs/2112.09118.
  14. 14.Gautier Izacard and Edouard Grave. 2021a. Distilling knowledge from reader to retriever for question answering. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  15. 15.Gautier Izacard and Edouard Grave. 2021b. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021, pages 874–880. Association for Computational Linguistics.
  16. 16.Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022. Few-shot learning with retrieval augmented language models. CoRR, abs/2208.03299.
  17. 17.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12).
  18. 18.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know When language models know? on the calibration of language models for question answering. Trans. Assoc. Comput. Linguistics, 9:962–977.
  19. 19.Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1601–1611. Association for Computational Linguistics.
  20. 20.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Association for Computational Linguistics.
  21. 21.Harold W Kuhn. 1955. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97.
  22. 22.Harold W Kuhn. 1956. Variants of the hungarian method for assignment problems. Naval research logistics quarterly, 3(4):253–258.
  23. 23.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: a benchmark for question answering research. Trans. Assoc. Comput. Linguistics, 7:452–466.
  24. 24.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 6086–6096. Association for Computational Linguistics.
  25. 25.Kyungjae Lee, Seung-won Hwang, Sang-eun Han, and Dohyeon Lee. 2021. Robustifying multi-hop QA through pseudo-evidentiality training. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 6110–6119. Association for Computational Linguistics.
  26. 26.Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2022. Factuality enhanced language models for open-ended text generation. CoRR, abs/2206.04624.
  27. 27.Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix X. Yu, and Sanjiv Kumar. 2022a. Large language models with controllable working memory. CoRR, abs/2211.05110.
  28. 28.Junlong Li, Zhuosheng Zhang, and Hai Zhao. 2022b. Self-prompting large language models for open-domain QA. CoRR, abs/2212.08635.
  29. 29.Jiacheng Liu, Skyler Hallinan, Ximing Lu, Pengfei He, Sean Welleck, Hannaneh Hajishirzi, and Yejin Choi. 2022a. Rainier: Reinforced knowledge introspector for commonsense question answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 8938–8958. Association for Computational Linguistics.
  30. 30.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2022b. Generated knowledge prompting for commonsense reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3154–3169. Association for Computational Linguistics.
  31. 31.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  32. 32.Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 7052–7063. Association for Computational Linguistics.
  33. 33.Kaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2023. Chain-of-skills: A configurable model for open-domain question answering. CoRR, abs/2305.03130.
  34. 34.Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adámek, Eric Malmi, and Aliaksei Severyn. 2022. Teaching small language models to reason. CoRR, abs/2212.08410.
  35. 35.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. CoRR, abs/2212.10511.
  36. 36.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 9802–9822. Association for Computational Linguistics.
  37. 37.Ella Neeman, Roee Aharoni, Or Honovich, Leshem Choshen, Idan Szpektor, and Omri Abend. 2022. Disentqa: Disentangling parametric and contextual knowledge with counterfactual question answering. CoRR, abs/2211.05655.
  38. 38.Barlas Oguz, Xilun Chen, Vladimir Karpukhin, Stan Peshterliev, Dmytro Okhonko, Michael Sejr Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Scott Yih. 2022. Unik-qa: Unified representations of structured and unstructured knowledge for open-domain question answering. In Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 1535–1546. Association for Computational Linguistics.
  39. 39.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  40. 40.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. CoRR, abs/2203.02155.
  41. 41.Liangming Pan, Wenhu Chen, Min-Yen Kan, and William Yang Wang. 2021. Contraqa: Question answering under contradicting contexts. CoRR, abs/2110.07803.
  42. 42.Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. CoRR, abs/2302.12813.
  43. 43.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick S. H. Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 2523–2544. Association for Computational Linguistics.
  44. 44.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  45. 45.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100, 000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016, pages 2383–2392. The Association for Computational Linguistics.
  46. 46.Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4615–4629. Association for Computational Linguistics.
  47. 47.Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models. CoRR, abs/2210.01296.
  48. 48.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and verification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 809–819. Association for Computational Linguistics.
  49. 49.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. CoRR, abs/2212.10509.
  50. 50.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  51. 51.Wenya Wang, Vivek Srikumar, Hanna Hajishirzi, and Noah A. Smith. 2022. Elaboration-generating commonsense question answering at scale. CoRR, abs/2209.01232.
  52. 52.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  53. 53.Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022. Symbolic knowledge distillation: from general language models to commonsense models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 4602–4625. Association for Computational Linguistics.
  54. 54.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771.
  55. 55.Wenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel, Douwe Kiela, and Barlas Oguz. 2021. Answering complex open-domain questions with multi-hop dense retrieval. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  56. 56.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 2369–2380. Association for Computational Linguistics.
  57. 57.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023. Generate rather than retrieve: Large language models are strong context generators. In 11th International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  58. 58.Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2022. A survey of knowledge-enhanced text generation. ACM Comput. Surv., 54(11s):227:1–227:38.
  59. 59.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. CoRR, abs/2210.03493.
  60. 60.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.

Citation

MLA
Zhang, Y., et al. “Merging Generated and Retrieved Knowledge for Open-Domain QA”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4710–28, https://doi.org/10.18653/v1/2023.emnlp-main.286.
APA
Zhang, Y., Khalifa, M., Logeswaran, L., Lee, M., Lee, H., & Wang, L. (2023). Merging Generated and Retrieved Knowledge for Open-Domain QA. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4710–4728. https://doi.org/10.18653/v1/2023.emnlp-main.286
Chicago
Zhang, Y., M. Khalifa, L. Logeswaran, M. Lee, H. Lee, and L. Wang. 2023. “Merging Generated and Retrieved Knowledge for Open-Domain QA”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4710–28. https://doi.org/10.18653/v1/2023.emnlp-main.286.
Harvard
Zhang, Y. et al. (2023) “Merging Generated and Retrieved Knowledge for Open-Domain QA”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4710–4728. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.286.
Vancouver
1. Zhang Y, Khalifa M, Logeswaran L, Lee M, Lee H, Wang L (2023) Merging Generated and Retrieved Knowledge for Open-Domain QA. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4710–4728

BibTeX

@inproceedings{zhang-etal-2023-merging,
    title = "Merging Generated and Retrieved Knowledge for Open-Domain {QA}",
    author = "Zhang, Yunxiang  and
      Khalifa, Muhammad  and
      Logeswaran, Lajanugen  and
      Lee, Moontae  and
      Lee, Honglak  and
      Wang, Lu",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.286/",
    doi = "10.18653/v1/2023.emnlp-main.286",
    pages = "4710--4728"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/