LLatrieval: LLM-Verified Retrieval for Verifiable Generation

Xiaonan LiChangtai ZhuLinyang LiZhangyue YinTianxiang SunXipeng Qiu

article2024NAACL72 citations

Proposes an iterative framework where large language models verify and update retrieved documents to overcome retrieval bottlenecks and improve citation accuracy in verifiable text generation.

Listen

Large language models frequently generate factually incorrect content, commonly known as hallucinations. To make these systems trustworthy for high-stakes applications such as medical diagnosis and technical reporting, verifiable generation requires models to cite supporting source documents for their answers. However, standard systems rely on smaller, separate search tools that act as a performance bottleneck. These conventional search tools often fail to find the correct evidence, and because language models typically receive these results passively without providing feedback, the overall accuracy and verifiability of the final answers suffer.

The article demonstrates that actively involving the language model in evaluating and refining retrieved evidence substantially improves answer correctness and citation quality. It introduces LLatrieval, a framework that establishes an iterative verification and update loop between the retrieval system and the language model to ensure all generated claims are backed by solid evidence.

To evaluate this approach, the researchers conducted extensive experiments across three standard long-form question answering benchmarks (ASQA, QAMPARI, and ELI5). The framework operates by first having the language model verify whether the initial search results can fully answer the user query. If the evidence falls short, the system updates the retrieval set through two coordinated mechanisms: progressive selection, which filters out irrelevant or redundant items from candidate lists, and missing-information querying, which prompts the language model to retrieve specific missing facts. This verify-update loop repeats until the evidence meets quality standards or reaches an iteration limit. The framework was evaluated across multiple leading language models, including GPT-3.5 and 70-billion-parameter open-source models.

The evaluation produced several key findings. First, the proposed framework established new state-of-the-art results across all evaluated benchmarks, improving answer correctness by an average of 3.4 points and citation quality by 5.9 points over baseline search methods. Second, it outperformed other language model enhancement techniques—such as RankGPT and query rewriting—while using fewer document evaluations because the system dynamically halts once sufficient evidence is found. Third, the internal verification performed by the language model matched the accuracy of feedback derived from human gold-standard answers. Finally, retrieval performance scaled positively with model size, demonstrating stronger results when powered by more capable language models.

These findings indicate that addressing the evidence retrieval bottleneck is critical for deploying reliable generative AI in enterprise settings. Rather than relying on static search pipelines, systems that integrate iterative model verification can significantly reduce factual hallucinations and compliance risks. Furthermore, because the verification mechanism allows the process to stop as soon as adequate evidence is gathered, organizations can better manage computational costs and operational efficiency compared to fixed-step alternatives.

Organizations implementing verifiable question-answering systems should adopt iterative verification frameworks to ensure output reliability. Decision-makers should leverage the system's adjustable verification thresholds to balance accuracy requirements against computing budgets based on their specific operational needs. For high-stakes deployments, teams should conduct pilot programs to establish domain-specific thresholds and assess system behavior when relevant reference data is missing entirely.

Confidence in these findings is supported by consistent improvements across diverse datasets and model architectures. However, decision-makers should note that the iterative process requires real-time model inference, which may introduce latency constraints in environments requiring ultra-fast response times. Performance gains also remain bounded by whether the necessary information actually exists within the underlying document repositories.

No sufficiently relevant recommendations were found.

Cover for LLatrieval: LLM-Verified Retrieval for Verifiable Generation

Table of Contents

  • 1 Introduction
  • 2 Background: Verifiable Generation
  • 2.1 Pipeline
  • 3.1 Retrieval Verification
  • 3.2 Retrieval Update
  • 3.3 Verify-Update Iteration
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Main Results
  • 4.3 Analyses
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • A Implementation Details
  • A.1 Method Details
  • A.2 Instructions
  • A.3 Average Document Candidates
  • B The Correlation between Retrieval Quality and Generation Quality
  • C Performance across Various Hyper-Parameters

Knowls

  1. Knowl 1 — LLatrieval uses LLM feedback to refine supporting evidence

    model/method

    LLatrieval is a retrieval framework for verifiable generation in which a large language model (LLM) does not simply accept an initial retrieval result. Given a question, it checks whether the current documents support answering that question; if not, it updates the documents and checks again. The update combines LLM selection among retrieved candidates with retrieval queries aimed at information missing from the current documents. The process ends when the LLM accepts the evidence or the iteration limit is reached, after which the retained documents are supplied to the answer-generating LLM. This makes retrieval quality subject to iterative LLM feedback rather than a one-time ranking decision.

  2. Knowl 2 — Verify-update procedure for LLatrieval

    algorithm

    For a question qq, LLatrieval returns a set DD of supporting documents. The document pool is the corpus searched by retriever RR; TT is the maximum number of iterations, NN is the number of candidate documents requested per retrieval, and kk is the target number of retained documents. A sliding window divides candidates into manageable groups. At each iteration, the LLM selects up to kk documents from the current set and a candidate window; after all windows have been processed, the LLM verifies the resulting set. The procedure stops on a positive verification or after TT iterations.

    Input: Question qq, corpus, retriever RR, LLM, maximum iterations TT, candidates per retrieval NN, retained-document count kk
    Output: Supporting-document set DD
    Set query QQ to qq
    Set DD to the empty set
    For iteration ii from 1 to TT:
        If DD is not empty:
            Ask the LLM to generate a query for information missing from DD for answering qq
            Set QQ to that query
        Retrieve NN candidate documents from the corpus using RR and query QQ
        For each sliding window WW of the retrieved candidates:
            Ask the LLM to select up to kk documents from DD and WW
            Replace DD with the selected documents
        Ask the LLM whether DD can support answering qq
        If verification is positive:
            Stop iterating
    Return DD
  3. Knowl 3 — LLM-based verification of retrieved evidence

    model/method

    LLatrieval implements retrieval verification with two instruction-prompted alternatives. In classification mode, an LLM receives the question and current documents and returns a binary judgment of whether the documents can support a direct, accurate answer. In score-and-filter mode, the LLM assigns the documents a support score from 0 to 10; verification passes when the score is at least a chosen threshold τ\tau, and fails otherwise. A larger τ\tau imposes a stricter evidence criterion and generally requires more retrieval-update iterations. Classification was used for the main benchmark comparison because the score-and-filter results varied with the selected threshold.

  4. Knowl 4 — Progressive selection of document candidates

    model/method

    Progressive Selection updates the retained evidence by asking an LLM to choose a combination of documents, rather than merely reranking candidates individually. The retriever's candidate list is processed in sliding windows. For each window, the LLM selects up to kk documents from the union of that window and the current retained set; the selected set replaces the previous set and remains the same target size. This allows the LLM to discard irrelevant or redundant documents while preserving useful evidence already found and incorporating useful candidates from later windows. In the experiments, the current documents were placed before new candidates in the LLM input, document summaries were used to reduce input length, and the standard settings used a window size of 20 and k=5k=5.

  5. Knowl 5 — Missing-information queries broaden retrieval

    model/method

    When the current documents do not support answering a question, LLatrieval asks an LLM to identify what information is missing and generate a query for retrieving it. The paper evaluates two query forms: a new question about the missing information, and a pseudo-passage expressing that information. The retriever uses the generated query to obtain additional candidates, which are then considered alongside the already retained documents by Progressive Selection. The two query forms are complementary ways of expressing the information gap; the paper does not establish that either form is uniformly better.

  6. Knowl 6 — Benchmark and implementation conditions

    experimental setup

    LLatrieval was evaluated in the retrieval-read setting on the ALCE benchmark's ASQA, QAMPARI, and ELI5 datasets. Answer correctness was measured with exact-match recall (EM-R) for ASQA, answer F1 for QAMPARI, and the Claim metric for ELI5. Verifiability was measured with citation recall, precision, and F1. The main implementation used OpenAI gpt-3.5-turbo-0301 for retrieval decisions and answer generation, with temperature 0; the retrieval procedure used 50 candidates per query, a sliding-window size of 20, five retained supporting documents, and at most four iterations. Wikipedia was the document corpus for ASQA and QAMPARI, with BGE-large as retriever; ELI5 used the Sphere web corpus and BM25. The ELI5 first-iteration query was generated using HYDE.

  7. Knowl 7 — LLatrieval outperforms retrieval and reranking baselines on ALCE

    empirical result

    On ALCE, both missing-information query styles achieved higher reported correctness and citation quality than the retrieval and reranking baselines under the same retrieval-read evaluation. With passage-style queries, LLatrieval scored 57.7 EM-R on ASQA, with citation recall/precision/F1 of 60.9/61.3/61.1; 24.3 answer F1 on QAMPARI, with citation scores of 32.4/36.6/34.4; and 16.5 Claim on ELI5, with citation scores of 75.8/67.4/71.3. Its reported overall correctness and citation F1 were 32.9 and 55.6. With question-style queries, the corresponding ASQA scores were 57.3 and 60.5/60.6/60.6; QAMPARI scores were 23.6 and 30.9/34.8/32.7; ELI5 scores were 16.7 and 75.4/68.0/71.5; and overall scores were 32.5 and 54.9. For comparison, RankGPT with 100 candidates—the strongest reranking baseline by the reported overall scores—achieved overall correctness 30.8 and citation F1 50.1. The reported scores use the benchmark's 0–100 scale.

  8. Knowl 8 — Ablations show that selection, verification, and updating contribute

    empirical result

    On ASQA and QAMPARI, Progressive Selection improved the initial BGE-large retrieval results: ASQA EM-R/citation F1 rose from 55.9/57.5 to 56.8/60.2, and QAMPARI answer F1/citation F1 rose from 17.3/24.0 to 23.0/32.3. The verification split identified substantially weaker cases: among ASQA examples that failed verification, EM-R/citation F1 were 41.0/40.0, versus 60.3/64.5 for examples that passed; for QAMPARI, the corresponding scores were 10.1/9.1 versus 25.8/36.6. Updating the failed examples improved their scores to 44.3/44.3 on ASQA and 17.7/21.1 on QAMPARI. The counts were 948 ASQA and 1,000 QAMPARI examples before verification; 779 and 827 passed, while 169 and 173 failed and were updated. These results show that verification distinguishes low-support cases and that updating those cases improves their subsequent generation results.

  9. Knowl 9 — Stricter verification and more iterations trade additional retrieval cost for quality

    empirical result

    In score-and-filter experiments on ASQA and QAMPARI, examples retained under higher verification thresholds had better generation quality on average. Raising the threshold generally improved benchmark performance while increasing the number of candidates processed, which the paper used as a proxy for LLM input cost because candidate documents dominated the input tokens. The trade-off was not strictly monotonic: a threshold of 8 did not improve ASQA correctness, which the authors attributed as possible causes to missing relevant information in the corpus or an LLM failure to follow the missing-information-query instruction. Experiments that increased the maximum number of iterations also showed improving performance as retrieval was refined. With passage-style queries, the numbers of examples reaching iterations 1–4 were 948, 169, 105, and 84 for ASQA; 1,000, 173, 70, and 46 for QAMPARI; and 1,000, 234, 144, and 111 for ELI5. The method therefore allows retrieval effort to vary with the evidence needed for individual questions.

  10. Knowl 10 — LLatrieval improves results with multiple LLMs

    empirical result

    Experiments with several LLMs showed improvements over the original retriever on both QAMPARI and ELI5, though the size of the gains varied by model. The original retriever scored 17.3 QAMPARI answer F1 and 24.0 citation F1, and 15.2 ELI5 Claim and 66.7 citation F1. With GPT-3.5-turbo-0301, LLatrieval scored 24.3/34.4 on QAMPARI and 16.5/71.3 on ELI5; with GPT-3.5-turbo-0613, 24.8/32.9 and 17.1/71.7; and with GPT-3.5-turbo-1106, 23.9/33.6 and 17.3/71.7. Results with open-source models were also above the original retriever: Llama-2-70B-Chat scored 21.0/28.9 and 15.8/61.0; Xwin-LM-70B-V0.1 scored 22.5/30.8 and 16.1/70.8; and Tulu-2-DPO-70B scored 22.7/32.0 and 16.9/71.8. The paper interprets the stronger results with GPT-3.5 relative to Llama-2-70B as evidence of potential to improve retrieval by using stronger LLMs, not as a demonstrated scaling law.

  11. Knowl 11 — LLM verification compares favorably with gold-answer feedback

    empirical result

    The paper compared LLatrieval's internal LLM verification with an external gold-answer criterion that passed retrieval only when gold answers appeared in the retrieved documents. On ASQA, internal feedback used an average of 68.9 candidates and achieved correctness 57.7 and citation score 61.4; gold feedback used 136.9 candidates and achieved 58.1 and 60.8. On QAMPARI, internal feedback used 64.5 candidates and achieved correctness 24.3 and citation score 34.4; gold feedback used 101.0 candidates and achieved 25.2 and 33.6. The authors characterize the generation results as comparable, with internal feedback requiring fewer candidates. The table labels the citation measure as Cite.

  12. Knowl 12 — LLatrieval has latency and bias limitations

    limitation

    LLatrieval requires real-time LLM inference for selection, missing-information query generation, and verification. The paper cautions that this may make it unsuitable for applications requiring low latency or high throughput, even though adaptive stopping can reduce retrieval effort when evidence is sufficient. The method also relies on LLM judgments and generated queries, so LLM biases may affect which documents are retained. The authors suggest that prompting choices and advances in inference efficiency may mitigate these issues, but do not report an evaluation that resolves them.

Coverage note — The appendix's full hyperparameter and retriever sensitivity grid and its retrieval-quality/generation-quality correlation analysis are omitted as supporting robustness evidence rather than distinct core contributions.

References

  1. 1.Samuel Joseph Amouyal, Ohad Rubin, Ori Yoran, Tomer Wolfson, Jonathan Herzig, and Jonathan Berant. 2022. QAMPARI: : An open-domain question answering benchmark for questions with many answers from multiple paragraphs. CoRR, abs/2205.12665.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In FAccT ’21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021, pages 610–623. ACM.
  3. 3.Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, and Kellie Webster. 2022. Attributed question answering: Evaluation and modeling for attributed large language models. CoRR, abs/2212.08037.
  4. 4.Luiz Henrique Bonifacio, Hugo Queiroz Abonizio, Marzieh Fadaee, and Rodrigo Frassetto Nogueira. 2022. Inpars: Data augmentation for information retrieval using large language models. CoRR, abs/2202.05144.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, and Dong Yu. 2023. Sub-sentence encoder: Contrastive learning of propositional semantic representations. CoRR, abs/2311.04335.
  7. 7.I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023. Factool: Factuality detection in generative AI - A tool augmented framework for multi-task and multi-domain scenarios. CoRR, abs/2307.13528.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2023. Palm: Scaling language modeling with pathways. J. Mach. Learn. Res., 24:240:1–240:113.
  9. 9.Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot dense retrieval from 8 examples. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  10. 10.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: long form question answering. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 3558–3567. Association for Computational Linguistics.
  11. 11.Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023a. RARR: researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 16477–16508. Association for Computational Linguistics.
  12. 12.Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023b. Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 1762–1777. Association for Computational Linguistics.
  13. 13.Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023c. Enabling large language models to generate text with citations. CoRR, abs/2305.14627.
  14. 14.Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022. TRUE: re-evaluating factual consistency evaluation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 3905–3920. Association for Computational Linguistics.
  15. 15.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2023a. Large language models cannot self-correct reasoning yet. CoRR, abs/2310.01798.
  16. 16.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023b. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. CoRR, abs/2311.05232.
  17. 17.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2.
  18. 18.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised dense information retrieval with contrastive learning. Trans. Mach. Learn. Res., 2022.
  19. 19.Vitor Jeronymo, Luiz Henrique Bonifacio, Hugo Queiroz Abonizio, Marzieh Fadaee, Roberto de Alencar Lotufo, Jakub Zavrel, and Rodrigo Frassetto Nogueira. 2023. Inpars-v2: Large language models as efficient dataset generators for information retrieval. CoRR, abs/2301.01820.
  20. 20.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Association for Computational Linguistics.
  21. 21.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  22. 22.Dongfang Li, Zetian Sun, Xinshuo Hu, Zhenyu Liu, Ziyang Chen, Baotian Hu, Aiguo Wu, and Min Zhang. 2023a. A survey of large language models attribution. CoRR, abs/2311.03731.
  23. 23.Xinze Li, Yixin Cao, Liangming Pan, Yubo Ma, and Aixin Sun. 2023b. Towards verifiable generation: A benchmark for knowledge-aware language model attribution. CoRR, abs/2310.05634.
  24. 24.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023a. Lost in the middle: How language models use long contexts. CoRR, abs/2307.03172.
  25. 25.Nelson F. Liu, Tianyi Zhang, and Percy Liang. 2023b. Evaluating verifiability in generative search engines. CoRR, abs/2304.09848.
  26. 26.Xiao Liu, Hanyu Lai, Hao Yu, Yifan Xu, Aohan Zeng, Zhengxiao Du, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023c. Webglm: Towards an efficient web-enhanced question answering system with human preferences. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2023, Long Beach, CA, USA, August 6-10, 2023, pages 4549–4560. ACM.
  27. 27.Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting for retrieval-augmented large language models. CoRR, abs/2305.14283.
  28. 28.Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2023. Expertqa: Expert-curated questions and attributed answers. CoRR, abs/2309.07852.
  29. 29.Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, H. Francis Song, Martin J. Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese. 2022. Teaching language models to support answers with verified quotes. CoRR, abs/2203.11147.
  30. 30.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. CoRR, abs/2305.14251.
  31. 31.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. Webgpt: Browser-assisted question-answering with human feedback. CoRR, abs/2112.09332.
  32. 32.Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large dual encoders are generalizable retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9844–9855, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  33. 33.Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 708–718, Online. Association for Computational Linguistics.
  34. 34.Rodrigo Frassetto Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. CoRR, abs/1901.04085.
  35. 35.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  36. 36.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  37. 37.Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick S. H. Lewis, Barlas Oguz, Edouard Grave, Wen-tau Yih, and Sebastian Riedel. 2021. The web is your oyster - knowledge-intensive NLP against a very large web corpus. CoRR, abs/2112.09924.
  38. 38.Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2023a. Webcpm: Interactive web search for chinese long-form question answering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 8968–8988. Association for Computational Linguistics.
  39. 39.Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Bendersky. 2023b. Large language models are effective text rankers with pairwise ranking prompting. CoRR, abs/2306.17563.
  40. 40.Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021. Measuring attribution in natural language generation models. CoRR, abs/2112.12870.
  41. 41.Vipula Rawte, Amit P. Sheth, and Amitava Das. 2023. A survey of hallucination in large foundation models. CoRR, abs/2309.05922.
  42. 42.Revanth Gangi Reddy, Yi R. Fung, Qi Zeng, Manling Li, Ziqi Wang, Paul Sullivan, and Heng Ji. 2023. Smartbook: Ai-assisted situation report generation. CoRR, abs/2303.14337.
  43. 43.Stephen E. Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Found. Trends Inf. Retr., 3(4):333–389.
  44. 44.Devendra Singh Sachan, Mike Lewis, Dani Yogatama, Luke Zettlemoyer, Joelle Pineau, and Manzil Zaheer. 2022. Questions are all you need to train a dense passage retriever. CoRR, abs/2206.10658.
  45. 45.M. Salvagno, F.S. Taccone, A.G. Gerli, and ChatGPT. 2023. Can artificial intelligence help for scientific writing? Critical Care, 27(1). Publisher: BioMed Central Ltd.
  46. 46.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin c! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 624–643. Association for Computational Linguistics.
  47. 47.Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing retrieval-augmented large language models with iterative retrieval-generation synergy. CoRR, abs/2305.15294.
  48. 48.Tao Shen, Guodong Long, Xiubo Geng, Chongyang Tao, Tianyi Zhou, and Daxin Jiang. 2023. Large language models are strong zero-shot retriever. CoRR, abs/2304.14233.
  49. 49.Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati. 2023. GPT-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems. CoRR, abs/2310.12397.
  50. 50.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: factoid questions meet long-form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 8273–8288. Association for Computational Linguistics.
  51. 51.Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. One embedder, any task: Instruction-finetuned text embeddings. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pages 1102–1121. Association for Computational Linguistics.
  52. 52.Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023. Is chatgpt good at search? investigating large language models as re-ranking agent. CoRR, abs/2304.09542.
  53. 53.Raphael Tang, Xinyu Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture. 2023. Found in the middle: Permutation self-consistency improves listwise ranking in large language models. CoRR, abs/2310.07712.
  54. 54.Xwin-LM Team. 2023. Xwin-lm.
  55. 55.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  56. 56.Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023. Can large language models really improve by self-critiquing their own plans? CoRR, abs/2310.08118.
  57. 57.Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Jiayang Cheng, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang. 2023a. Survey on factuality in large language models: Knowledge, retrieval and domain-specificity. CoRR, abs/2310.07521.
  58. 58.Liang Wang, Nan Yang, and Furu Wei. 2023b. Query2doc: Query expansion with large language models. CoRR, abs/2303.07678.
  59. 59.Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md. Rizwan Parvez, and Graham Neubig. 2023c. Learning to filter context for retrieval-augmented generation. CoRR, abs/2311.08377.
  60. 60.Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof. 2023. C-pack: Packaged resources to advance general chinese embedding. CoRR, abs/2309.07597.
  61. 61.Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2023. Making retrieval-augmented language models robust to irrelevant context. CoRR, abs/2310.01558.
  62. 62.Hang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2022a. Adversarial retriever-ranker for dense text retrieval. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  63. 63.Shunyu Zhang, Yaobo Liang, Ming Gong, Daxin Jiang, and Nan Duan. 2022b. Multi-view document representation learning for open-domain dense retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 5990–6000. Association for Computational Linguistics.
  64. 64.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023. Siren’s song in the AI ocean: A survey on hallucination in large language models. CoRR, abs/2309.01219.
  65. 65.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.

Citation

MLA
Li, X., et al. “LLatrieval: LLM-Verified Retrieval for Verifiable Generation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5453–71, https://doi.org/10.18653/V1/2024.NAACL-LONG.305.
APA
Li, X., Zhu, C., Li, L., Yin, Z., Sun, T., & Qiu, X. (2024). LLatrieval: LLM-Verified Retrieval for Verifiable Generation. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5453–5471. https://doi.org/10.18653/V1/2024.NAACL-LONG.305
Chicago
Li, X., C. Zhu, L. Li, Z. Yin, T. Sun, and X. Qiu. 2024. “LLatrieval: LLM-Verified Retrieval for Verifiable Generation”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5453–71. https://doi.org/10.18653/V1/2024.NAACL-LONG.305.
Harvard
Li, X. et al. (2024) “LLatrieval: LLM-Verified Retrieval for Verifiable Generation”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5453–5471. Available at: https://doi.org/10.18653/V1/2024.NAACL-LONG.305.
Vancouver
1. Li X, Zhu C, Li L, Yin Z, Sun T, Qiu X (2024) LLatrieval: LLM-Verified Retrieval for Verifiable Generation. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5453–5471

BibTeX

@inproceedings{Li_2024, title={LLatrieval: LLM-Verified Retrieval for Verifiable Generation}, url={http://dx.doi.org/10.18653/V1/2024.NAACL-LONG.305}, DOI={10.18653/v1/2024.naacl-long.305}, booktitle={Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Li, Xiaonan and Zhu, Changtai and Li, Linyang and Yin, Zhangyue and Sun, Tianxiang and Qiu, Xipeng}, year={2024}, pages={5453–5471} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/