Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Satyapriya KrishnaKalpesh KrishnaAnhad MohananeySteven SchwarczAdam StamblerShyam UpadhyayManaal Faruqui

article2025NAACL200 citations

Introduces FRAMES, a benchmark of multi-hop questions requiring information synthesis across multiple documents to evaluate retrieval-augmented generation systems simultaneously on factuality, retrieval, and complex reasoning.

Listen

Modern artificial intelligence applications increasingly rely on retrieval-augmented generation, a technique that pairs large language models with search systems to fetch external documents and synthesize accurate, up-to-date answers. Despite widespread deployment, existing benchmarks typically evaluate factual correctness, search retrieval, and multi-step reasoning in isolation rather than together. This fragmented testing fails to capture how language models perform in realistic scenarios where they must locate multiple disparate facts and correctly reason over them to answer complex questions.

The main objective of the article is to introduce a unified evaluation benchmark, named FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), and to evaluate how effectively state-of-the-art language models perform end-to-end multi-document retrieval and complex reasoning tasks.

To establish this benchmark, human experts created a dataset of 824 challenging, multi-hop questions grounded in Wikipedia articles. Each question requires synthesizing facts across 2 to 15 different articles and tests skills such as numerical calculations, timeline tracking, and constraint matching. The article evaluated leading models, including Gemini Pro 1.5, Gemini Flash 1.5, Gemma 2, Llama 3.2, and Qwen 2.5, across single-step answering, standard document retrieval, and multi-step iterative search pipelines.

The key findings reveal significant performance gaps in current systems. First, leading language models struggle severely when answering complex multi-source questions in a single step without external search, with top models achieving an accuracy of only about 41%. Second, providing models with perfect ground-truth context establishes an upper performance bound of roughly 73% accuracy; approximately 80% of the remaining errors stem from failures in numerical, tabular, and post-processing reasoning rather than missing facts. Third, implementing a multi-step retrieval and search-planning pipeline dramatically improves accuracy to 66%—a more than 50% relative improvement over standard single-step prompting. Fourth, unguided iterative search often traps models in repetitive, incorrect query loops, whereas search planning prompts that encourage diverse queries allow models to recover and approach upper-bound accuracy.

These results demonstrate that simply giving language models search tools is insufficient for solving complex information tasks. System reliability depends heavily on structured, multi-step search planning and specialized numerical and tabular reasoning. Relying on single-step model generation for multi-source knowledge tasks introduces high error rates and operational risk, whereas iterative retrieval pipelines offer a viable path to high accuracy.

To build robust systems, organizations should adopt iterative search architectures with explicit planning and anti-repetition instructions rather than relying on standard single-pass question answering. Future technical efforts should focus on training specialized dense retrieval models for multi-hop contexts, implementing step-by-step verification methods to improve reasoning accuracy, and optimizing search pipelines to reduce the computational cost of multiple query cycles.

Readers should interpret these findings within certain limitations. The benchmark is restricted to Wikipedia-based data, which may not reflect all domain-specific enterprise settings, and there is a potential risk that models encountered parts of this public information during pre-training. Nevertheless, the study provides high confidence that current models face genuine reasoning bottlenecks when integrating facts across multiple sources, highlighting the necessity of iterative retrieval planning.

arXiv: 2409.12941

No sufficiently relevant recommendations were found.

Cover for Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Abstract

Large Language Models (LLMs) have shown significant improvements across cognitive tasks, with an emerging application in enhancing retrieval-augmented generation (RAG) capabilities. These systems require LLMs to understand queries, retrieve relevant information, and synthesize accurate responses. Given their increasing real-world deployment, comprehensive evaluation is crucial. We propose FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), a high-quality dataset designed to test LLMs’ factual responses, retrieval capabilities, and reasoning in generating final answers. Unlike previous work evaluating these abilities in isolation, FRAMES offers a unified framework for assessing LLM performance in end-to-end RAG scenarios. Our dataset comprises challenging multi-hop questions requiring integration of information from multiple sources. Baseline results show that even state-of-the-art LLMs struggle, achieving 0.408 accuracy without retrieval. However, our proposed multi-step retrieval pipeline significantly improves accuracy to 0.66 (>50% improvement). We aim to bridge evaluation gaps and assist in developing more robust RAG systems.

Table of Contents

  • 1 Introduction
  • 2 FRAMES
  • 3 Empirical Analysis
  • 3.1 Single-Step Evaluations
  • 3.2 Multi-Step Evaluations
  • 4 Related Works
  • 5 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Future Work
  • B Additional Results
  • Synthetic Data Generation

Knowls

  1. Knowl 1 — FRAMES evaluates retrieval-augmented generation end to end

    definition

    FRAMES (Factuality, Retrieval, And reasoning MEasurement Set) is an evaluation dataset of 824 questions intended to assess retrieval-augmented generation systems across factual grounding, retrieval of information, and reasoning that combines retrieved facts into an answer. Its questions are designed to require information from multiple Wikipedia articles, so evaluation tests these capabilities together rather than in isolation.

  2. Knowl 2 — Human-authored questions combine facts from multiple Wikipedia articles

    experimental setup

    Expert human annotators created FRAMES questions using information from Wikipedia articles. Each question has a gold answer and a list of articles needed to answer it; questions require information from 2–15 articles. Annotators were asked to create standalone factoid questions with a single unambiguous answer and to label the reasoning abilities they require. The resulting dataset spans topics including history, sports, science, animals, and health. About 36% of questions require two articles, about 35% require three, and about 16% require four.

  3. Knowl 3 — Multi-step retrieval repeatedly expands context before answering

    algorithm

    The FRAMES multi-step evaluation pipeline takes a question, a Wikipedia article index, a number of retrieval rounds nn, a number of search queries per round kk, and a number of documents ndocsn_{\text{docs}} to retrieve per query. It initializes the context with the question. At each round, the language model generates kk search queries based on the question and accumulated context; BM25 retrieves the top ndocsn_{\text{docs}} articles for each query, and articles not already in context are added. After nn rounds, the model produces its final answer using the expanded context. Experiments compare a vanilla query-generation prompt with a search-planning prompt that provides examples of useful query sequences and instructs the model not to repeat queries and to reason step by step. With n=5n=5, the pipeline requires five sequential retrieval-generation calls and one final answering call.

  4. Knowl 4 — Search planning raises Gemini Pro accuracy to 0.66

    empirical result

    On FRAMES, Gemini-Pro-1.5-0514 achieved accuracy 0.66 with multi-step retrieval and search-planning instructions using five rounds, five queries per round, and ten retrieved documents per query. This approached the 0.729 accuracy achieved when the model received all gold Wikipedia articles. In the vanilla multi-step setting with five rounds, five queries per round, and two documents per query, accuracy rose from about 0.45 to about 0.52 as retrieval was iterated. The paper reports that planning improved performance across reasoning types, with numerical-reasoning accuracy exceeding the oracle result in the plotted comparison.

  5. Knowl 5 — Single-step accuracy improves with retrieved and oracle articles

    data/table

    The following accuracies compare single-step prompting conditions on FRAMES. Naive prompting supplies no retrieved articles; BM25-R adds the top two or four Wikipedia articles retrieved for the question; the oracle condition supplies all gold articles identified by annotators. Dashes indicate results not reported because of context-length constraints. The table shows that adding retrieved articles helped both Gemini models, while supplying all gold articles produced higher accuracy still. Responses were scored against free-form gold answers by an LLM autorater; comparison with human evaluation of Gemini-Pro-1.5-0514 autoratings yielded accuracy 0.96 and Cohen’s κ=0.889\kappa=0.889.

    Condition G-Pro-1.5 G-Flash-1.5 Gemma2-27b Llama3.2-3B-I Qwen2.5-3B-I
    Naive prompt 0.408 0.263 0.308 0.115 0.095
    BM25-R (ndoc=2n_{\text{doc}}=2) 0.452 0.288 – – –
    BM25-R (ndoc=4n_{\text{doc}}=4) 0.474 0.315 – – –
    Oracle prompt 0.729 0.665 – – –

    G-Pro-1.5 and G-Flash-1.5 are Gemini-Pro-1.5-0514 and Gemini-Flash-1.5-0514, respectively.

  6. Knowl 6 — FRAMES distinguishes five kinds of reasoning demand

    definition

    FRAMES assigns questions to one or more of five reasoning types. Numerical reasoning involves counting, comparison, or calculation. Tabular reasoning requires extracting or analyzing statistics in tables or infoboxes. Multiple-constraint reasoning finds an answer satisfying several constraints jointly. Temporal reasoning requires reasoning over dates or timelines. Post-processing requires transforming an answer after retrieving the necessary facts, such as converting a calculated year into Roman numerals. Multiple labels may apply to the same question. Multiple-constraint questions are the largest category at about 36% of the dataset, and numerical-reasoning questions account for about 20%.

  7. Knowl 7 — Dataset checks target answer validity, freshness, and guessability

    experimental setup

    FRAMES underwent several human quality checks. Annotators rechecked whether each answer was correct and supported by its associated Wikipedia pages three months after initial collection; 5.5% of samples were removed because their answers were no longer true. Questions with answers that could change over time were given date context to disambiguate them. Yes/no questions were excluded to avoid making random guessing a viable route to substantial accuracy. Restricting evidence to Wikipedia was intended to improve interpretability and reliability.

  8. Knowl 8 — Oracle-context errors cluster in numerical, tabular, and post-processing tasks

    empirical result

    Even when Gemini-Pro-1.5-0514 received all gold Wikipedia articles, it answered 27% of FRAMES questions incorrectly. About 80% of those oracle-condition errors were in numerical reasoning, tabular reasoning, or post-processing, indicating that access to the relevant facts did not eliminate reasoning failures. The paper also reports that BM25 retrieval chiefly improved multiple-constraint questions (about 8%) and post-processing questions (about 10%) relative to naive prompting.

  9. Knowl 9 — Synthetic question generation required substantial human correction

    empirical result

    The authors tested generating multi-article questions with a state-of-the-art language model before using human annotation for the final FRAMES dataset. More than 30% of generated questions and answers were hallucinated, and the model struggled to produce questions that strictly required information from more than four articles. After manually removing hallucinated examples, the same model achieved about 32% accuracy on the remaining generated questions. The authors therefore used the synthetic task instructions as guidance for human annotation rather than relying on synthetic examples as the final benchmark.

  10. Knowl 10 — Potential pretraining contamination and limited coverage remain concerns

    limitation

    The authors identify possible overlap between FRAMES’ Wikipedia evidence and language models’ pretraining data as a limitation that could inflate measured performance and weaken conclusions about generalization. They do not quantify the extent of contamination. They also note that, despite efforts to include diverse questions, the dataset may not represent the full range of real-world queries, limiting its applicability to some domains and use cases.

Coverage note — Background comparisons with prior benchmarks, ethical considerations, and proposed future research directions are omitted because they do not add independent results or methods needed to reconstruct the paper’s main contribution.

References

  1. 1.Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Wang. 2020. Hybridqa: A dataset of multi-hop question answering over tabular and textual data. arXiv preprint arXiv:2004.07347.
  2. 2.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  3. 3.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. Eli5: Long form question answering. arXiv preprint arXiv:1907.09190.
  4. 4.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821.
  5. 5.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
  6. 6.Team Gemma, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118.
  7. 7.Google. 2024a. Gemini 1.5 flash. https://deepmind.google/technologies/gemini/flash/.
  8. 8.Google. 2024b. Gemini 1.5 pro. https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/.
  9. 9.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929–3938. PMLR.
  10. 10.Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551.
  11. 11.Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Akari Asai, Xinyan Yu, Dragomir Radev, Noah A Smith, Yejin Choi, Kentaro Inui, et al. 2024. Realtime qa: what’s the answer right now? Advances in Neural Information Processing Systems, 36.
  12. 12.Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714.
  13. 13.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, pages 39–48.
  14. 14.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213.
  15. 15.Satyapriya Krishna, Chirag Agarwal, and Himabindu Lakkaraju. 2024. Understanding the effects of iterative prompting on truthfulness. CoRR, abs/2402.06625.
  16. 16.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453–466.
  17. 17.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  18. 18.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step. arXiv preprint arXiv:2305.20050.
  19. 19.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
  20. 20.Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3214–3252. Association for Computational Linguistics.
  21. 21.Tie-Yan Liu et al. 2009. Learning to rank for information retrieval. Foundations and Trends® in Information Retrieval, 3(3):225–331.
  22. 22.Meta. 2024. Llama-3.2. https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct.
  23. 23.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP.
  24. 24.Qwen. 2024. Qwen-2.5. https://huggingface.co/Qwen/Qwen2.5-3B-Instruct.
  25. 25.Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at trec-3. Nist Special Publication Sp, 109:109.
  26. 26.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36.
  27. 27.Julian Schnitzler, Xanh Ho, Jiahao Huang, Florian Boudin, Saku Sugawara, and Akiko Aizawa. 2024. Morehopqa: More than multi-hop reasoning. arXiv preprint arXiv:2406.13397.
  28. 28.Yixuan Tang and Yi Yang. 2024. Multihop-rag: Benchmarking retrieval-augmented generation for multi-hop queries. arXiv preprint arXiv:2401.15391.
  29. 29.Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Musique: Multihop questions via single-hop question composition. Trans. Assoc. Comput. Linguistics, 10:539–554.
  30. 30.Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al. 2023. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214.
  31. 31.Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2017. Constructing datasets for multi-hop reading comprehension across documents. CoRR, abs/1710.06481.
  32. 32.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018a. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600.
  33. 33.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018b. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 2369–2380. Association for Computational Linguistics.
  34. 34.Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024a. Evaluation of retrieval-augmented generation: A survey. ArXiv, abs/2405.07437.
  35. 35.Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024b. Evaluation of retrieval-augmented generation: A survey. arXiv preprint arXiv:2405.07437.
  36. 36.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223.

Citation

MLA
Krishna, S., et al. “Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 4745–59, https://doi.org/10.18653/v1/2025.naacl-long.243.
APA
Krishna, S., Krishna, K., Mohananey, A., Schwarcz, S., Stambler, A., Upadhyay, S., & Faruqui, M. (2025). Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 4745–4759. https://doi.org/10.18653/v1/2025.naacl-long.243
Chicago
Krishna, S., K. Krishna, A. Mohananey, et al. 2025. “Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 4745–59. https://doi.org/10.18653/v1/2025.naacl-long.243.
Harvard
Krishna, S. et al. (2025) “Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4745–4759. Available at: https://doi.org/10.18653/v1/2025.naacl-long.243.
Vancouver
1. Krishna S, Krishna K, Mohananey A, Schwarcz S, Stambler A, Upadhyay S, Faruqui M (2025) Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 4745–4759

BibTeX

@inproceedings{krishna-etal-2025-fact,
    title = "Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation",
    author = "Krishna, Satyapriya  and
      Krishna, Kalpesh  and
      Mohananey, Anhad  and
      Schwarcz, Steven  and
      Stambler, Adam  and
      Upadhyay, Shyam  and
      Faruqui, Manaal",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.243/",
    doi = "10.18653/v1/2025.naacl-long.243",
    pages = "4745--4759",
    ISBN = "979-8-89176-189-6"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/