Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Satyapriya KrishnaKalpesh KrishnaAnhad MohananeySteven SchwarczAdam StamblerShyam UpadhyayManaal Faruqui
Introduces FRAMES, a benchmark of multi-hop questions requiring information synthesis across multiple documents to evaluate retrieval-augmented generation systems simultaneously on factuality, retrieval, and complex reasoning.
Modern artificial intelligence applications increasingly rely on retrieval-augmented generation, a technique that pairs large language models with search systems to fetch external documents and synthesize accurate, up-to-date answers. Despite widespread deployment, existing benchmarks typically evaluate factual correctness, search retrieval, and multi-step reasoning in isolation rather than together. This fragmented testing fails to capture how language models perform in realistic scenarios where they must locate multiple disparate facts and correctly reason over them to answer complex questions.
The main objective of the article is to introduce a unified evaluation benchmark, named FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), and to evaluate how effectively state-of-the-art language models perform end-to-end multi-document retrieval and complex reasoning tasks.
To establish this benchmark, human experts created a dataset of 824 challenging, multi-hop questions grounded in Wikipedia articles. Each question requires synthesizing facts across 2 to 15 different articles and tests skills such as numerical calculations, timeline tracking, and constraint matching. The article evaluated leading models, including Gemini Pro 1.5, Gemini Flash 1.5, Gemma 2, Llama 3.2, and Qwen 2.5, across single-step answering, standard document retrieval, and multi-step iterative search pipelines.
The key findings reveal significant performance gaps in current systems. First, leading language models struggle severely when answering complex multi-source questions in a single step without external search, with top models achieving an accuracy of only about 41%. Second, providing models with perfect ground-truth context establishes an upper performance bound of roughly 73% accuracy; approximately 80% of the remaining errors stem from failures in numerical, tabular, and post-processing reasoning rather than missing facts. Third, implementing a multi-step retrieval and search-planning pipeline dramatically improves accuracy to 66%—a more than 50% relative improvement over standard single-step prompting. Fourth, unguided iterative search often traps models in repetitive, incorrect query loops, whereas search planning prompts that encourage diverse queries allow models to recover and approach upper-bound accuracy.
These results demonstrate that simply giving language models search tools is insufficient for solving complex information tasks. System reliability depends heavily on structured, multi-step search planning and specialized numerical and tabular reasoning. Relying on single-step model generation for multi-source knowledge tasks introduces high error rates and operational risk, whereas iterative retrieval pipelines offer a viable path to high accuracy.
To build robust systems, organizations should adopt iterative search architectures with explicit planning and anti-repetition instructions rather than relying on standard single-pass question answering. Future technical efforts should focus on training specialized dense retrieval models for multi-hop contexts, implementing step-by-step verification methods to improve reasoning accuracy, and optimizing search pipelines to reduce the computational cost of multiple query cycles.
Readers should interpret these findings within certain limitations. The benchmark is restricted to Wikipedia-based data, which may not reflect all domain-specific enterprise settings, and there is a potential risk that models encountered parts of this public information during pre-training. Nevertheless, the study provides high confidence that current models face genuine reasoning bottlenecks when integrating facts across multiple sources, highlighting the necessity of iterative retrieval planning.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational RAG paper establishes the retrieval-and-generation setup that FRAMES evaluates end to end.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). Its survey maps RAG architectures and evaluation gaps, clarifying the context for FRAMES’s unified benchmark.
- Paper: ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, Jon Saad-Falcon et al. (2024). ARES evaluates context relevance, faithfulness, and answer relevance, supplying a direct antecedent for FRAMES’s broader RAG evaluation.
- Paper: Ragas: Automated Evaluation of Retrieval Augmented Generation, Shahul Es et al. (2024). Ragas operationalizes reference-free evaluation of retrieval context and generated answers, preparing readers for FRAMES’s combined measures.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). ALCE establishes end-to-end evaluation of retrieval-grounded answers and their evidence, a useful precursor to FRAMES’s integrated assessment.
No sufficiently relevant recommendations were found.
