LLatrieval: LLM-Verified Retrieval for Verifiable Generation
Xiaonan LiChangtai ZhuLinyang LiZhangyue YinTianxiang SunXipeng Qiu
Proposes an iterative framework where large language models verify and update retrieved documents to overcome retrieval bottlenecks and improve citation accuracy in verifiable text generation.
Large language models frequently generate factually incorrect content, commonly known as hallucinations. To make these systems trustworthy for high-stakes applications such as medical diagnosis and technical reporting, verifiable generation requires models to cite supporting source documents for their answers. However, standard systems rely on smaller, separate search tools that act as a performance bottleneck. These conventional search tools often fail to find the correct evidence, and because language models typically receive these results passively without providing feedback, the overall accuracy and verifiability of the final answers suffer.
The article demonstrates that actively involving the language model in evaluating and refining retrieved evidence substantially improves answer correctness and citation quality. It introduces LLatrieval, a framework that establishes an iterative verification and update loop between the retrieval system and the language model to ensure all generated claims are backed by solid evidence.
To evaluate this approach, the researchers conducted extensive experiments across three standard long-form question answering benchmarks (ASQA, QAMPARI, and ELI5). The framework operates by first having the language model verify whether the initial search results can fully answer the user query. If the evidence falls short, the system updates the retrieval set through two coordinated mechanisms: progressive selection, which filters out irrelevant or redundant items from candidate lists, and missing-information querying, which prompts the language model to retrieve specific missing facts. This verify-update loop repeats until the evidence meets quality standards or reaches an iteration limit. The framework was evaluated across multiple leading language models, including GPT-3.5 and 70-billion-parameter open-source models.
The evaluation produced several key findings. First, the proposed framework established new state-of-the-art results across all evaluated benchmarks, improving answer correctness by an average of 3.4 points and citation quality by 5.9 points over baseline search methods. Second, it outperformed other language model enhancement techniques—such as RankGPT and query rewriting—while using fewer document evaluations because the system dynamically halts once sufficient evidence is found. Third, the internal verification performed by the language model matched the accuracy of feedback derived from human gold-standard answers. Finally, retrieval performance scaled positively with model size, demonstrating stronger results when powered by more capable language models.
These findings indicate that addressing the evidence retrieval bottleneck is critical for deploying reliable generative AI in enterprise settings. Rather than relying on static search pipelines, systems that integrate iterative model verification can significantly reduce factual hallucinations and compliance risks. Furthermore, because the verification mechanism allows the process to stop as soon as adequate evidence is gathered, organizations can better manage computational costs and operational efficiency compared to fixed-step alternatives.
Organizations implementing verifiable question-answering systems should adopt iterative verification frameworks to ensure output reliability. Decision-makers should leverage the system's adjustable verification thresholds to balance accuracy requirements against computing budgets based on their specific operational needs. For high-stakes deployments, teams should conduct pilot programs to establish domain-specific thresholds and assess system behavior when relevant reference data is missing entirely.
Confidence in these findings is supported by consistent improvements across diverse datasets and model architectures. However, decision-makers should note that the iterative process requires real-time model inference, which may introduce latency constraints in environments requiring ultra-fast response times. Performance gains also remain bounded by whether the necessary information actually exists within the underlying document repositories.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). Read this foundational account of retrieval-augmented generation first to understand the retrieve-and-generate pipeline LLatrieval strengthens with iterative evidence verification.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). Its ALCE benchmark establishes the citation-quality measures and ASQA, QAMPARI, and ELI5 evaluation setting that LLatrieval uses to assess its improvements.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). FLARE introduces inference-time, need-driven retrieval for long-form answers, preparing you for LLatrieval’s iterative retrieval updates when evidence is incomplete.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). Self-RAG’s retrieve-and-critique framework provides a useful precursor for understanding LLatrieval’s model-led assessment of retrieved evidence.
No sufficiently relevant recommendations were found.
