Built independently by an author, for readers. Read the story and support ChapterPal

keyword

retrieval-augmented generation

Retrieval-augmented generation is a method in which a language model retrieves relevant information from an external source and uses it as context to generate a response. By grounding generation in retrieved material, it can draw on information beyond what is encoded in the model’s parameters.

55 items

Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval

Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval

Peter Baile Chen, Yi Zhang, Dan Roth

OrganizationsAmazon Web ServicesMassachusetts Institute of TechnologyUniversity of Pennsylvania

Why you should read this

Proposes a join-aware table retrieval framework using mixed-integer programming to re-rank candidate tables by jointly evaluating query-table relevance and table joinability, improving end-to-end question-answering accuracy over multi-table databases.

Retrieving relevant tables containing the necessary information to accurately answer a given question over tables is critical to open-domain question-answering (QA) systems. Previous methods assume the answer to such a question can be found either in a single table or multiple tables identified through question decomposition or rewriting. However, neither of these approaches is sufficient, as many questions require retrieving multiple tables and joining them through a join plan that cannot be discerned from the user query itself. If the join plan is not considered in the retrieval stage, the subsequent steps of reasoning and answering based on those retrieved tables are likely to be incorrect. To address this problem, we introduce a method that uncovers useful join relations for any query and database during table retrieval. We use a novel re-ranking method formulated as a mixed-integer program that considers not only table-query relevance but also table-table relevance that requires inferring join relationships. Our method outperforms the state-of-the-art approaches for table retrieval by up to 9.3% in F1 score and for end-to-end QA by up to 5.4% in accuracy.

Added

2026-10-05

Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together

Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together

Dilara Soylu, Christopher Potts, Omar Khattab

OrganizationsStanford University

Why you should read this

Demonstrates that alternating between prompt optimization and model weight fine-tuning enables modular language model pipelines to teach themselves, substantially outperforming either approach used alone across complex reasoning and retrieval tasks.

Natural Language Processing (NLP) systems are increasingly taking the form of sophisticated modular pipelines, e.g., Retrieval Augmented Generation (RAG), where each module may involve a distinct Language Model (LM) and an associated prompt template. These compound systems often lack intermediate labels or gradient flow to optimize each module, making their end-to-end optimization challenging. Here we seek strategies to optimize both the module-level LM weights and the associated prompt templates of such systems to maximize a downstream task metric. We propose for the first time combining the weight and prompt optimization strategies to optimize a modular LM pipeline by alternating between the two to get the same LM to teach itself. In experiments with multi-hop QA, mathematical reasoning, and feature-based classification using mistral-7b, llama-2-7b, and llama-3-8b, these BetterTogether strategies optimizing the weights and prompts of a pipeline together outperform directly optimizing weights alone and prompts alone by up to 60% and 6%, respectively, on average across LMs and tasks. Our BetterTogether optimizer is released in DSPy at http://dspy.ai.

Added

2026-10-05

Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Yuhan Chen, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang, Yongbin Li, Rui Yan

OrganizationsAlibaba GroupRenmin University of China

Why you should read this

Proposes Attention Buckets, a training-free inference method that eliminates blind spots in LLM context retrieval caused by rotary position embedding attention waveforms by ensembling parallel processes with complementary angle bases, boosting 7B models to GPT-4-level tool-use accuracy.

In this paper, we demonstrate that an inherent waveform pattern in the attention allocation of large language models (LLMs) significantly affects their performance in tasks demanding a high degree of context awareness, such as utilizing LLMs for tool-use. Specifically, the crucial information in the context will be potentially overlooked by model when it is positioned in the trough zone of the attention waveform, leading to decreased performance. To address this issue, we propose a novel inference method named Attention Buckets. It allows LLMs to process their input through multiple parallel processes. Each process utilizes a distinct base angle for the rotary position embedding, thereby creating a unique attention waveform. By compensating an attention trough of a particular process with an attention peak of another process, our approach enhances LLM’s awareness to various contextual positions, thus mitigating their risk of overlooking crucial information. In the largest tool-use benchmark, our method elevates a 7B model to achieve state-of-the-art performance comparable to that of GPT-4. On other benchmarks and some RAG tasks, which also demand a thorough understanding of contextual content, Attention Buckets also exhibited notable enhancements in performance.

Added

2026-10-05

Training Language Models to Generate Text with Citations via Fine-grained Rewards

Training Language Models to Generate Text with Citations via Fine-grained Rewards

Chengyu Huang, Zeqiu Wu, Yushi Hu, Wenya Wang

OrganizationsNanyang Technological UniversityNational University of SingaporeUniversity of Washington

Why you should read this

Proposes a reinforcement learning and rejection sampling framework using sentence-level fine-grained rewards for citation precision, recall, and answer correctness, enabling smaller open-source models like LLaMA-2-7B to outperform GPT-3.5-turbo in generating factually accurate, well-cited text.

While recent Large Language Models (LLMs) have proven useful in answering user queries, they are prone to hallucination, and their responses often lack credibility due to missing references to reliable sources. An intuitive solution to these issues would be to include in-text citations referring to external documents as evidence. While previous works have directly prompted LLMs to generate in-text citations, their performances are far from satisfactory, especially when it comes to smaller LLMs. In this work, we propose an effective training framework using fine-grained rewards to teach LLMs to generate highly supportive and relevant citations, while ensuring the correctness of their responses. We also conduct a systematic analysis of applying these fine-grained rewards to common LLM training strategies, demonstrating its advantage over conventional practices. We conduct extensive experiments on Question Answering (QA) datasets taken from the ALCE benchmark and validate the model's generalizability using EXPERTQA. On LLaMA-2-7B, the incorporation of fine-grained rewards achieves the best performance among the baselines, even surpassing that of GPT-3.5-turbo.1

Added

2026-10-05

Why language models hallucinate

Why language models hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang

OrganizationsGeorgia Institute of TechnologyOpenAI

Why you should read this

Explains how standard training objectives and benchmark scoring inherently reward language models for guessing rather than expressing uncertainty, framing hallucinations as predictable statistical classification errors that require reformed evaluation metrics to fix.

Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty. Such "hallucinations" persist even in state-of-the-art systems and undermine trust. We argue that language models hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty, and we analyze the statistical causes of hallucinations in the modern training pipeline. Hallucinations need not be mysterious -- they originate simply as errors in binary classification. If incorrect statements cannot be distinguished from facts, then hallucinations in pretrained language models will arise through natural statistical pressures. We then argue that hallucinations persist due to the way most evaluations are graded -- language models are optimized to be good test-takers, and guessing when uncertain improves test performance. This "epidemic" of penalizing uncertain responses can only be addressed through a socio-technical mitigation: modifying the scoring of existing benchmarks that are misaligned but dominate leaderboards, rather than introducing additional hallucination evaluations. This change may steer the field toward more trustworthy AI systems.

Added

2026-10-05

Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare

Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare

Ravi Shankar, Sheng Wong, Lin Li, Magdalena Bachmann, Alex Silverthorne, Beth Albert, Gabriel Jones

OrganizationsDepartment of Computer ScienceUniversity of Oxford

Why you should read this

Proposes an energy-based modeling approach that significantly outperforms calibrated softmax confidence for selective abstention in healthcare retrieval-augmented generation, substantially reducing false positive rates on hard near-distribution queries.

Reliable abstention is critical for retrieval-augmented generation (RAG) systems, particularly in safety-critical domains such as women's health, where incorrect answers can lead to harm. We present an energy-based model (EBM) that learns a smooth energy landscape over a dense semantic corpus of 2.6M guideline-derived questions, enabling the system to decide when to generate or abstain. We benchmark the EBM against a calibrated softmax baseline and a k-nearest neighbour (kNN) density heuristic across both easy and hard abstention splits, where hard cases are semantically challenging near-distribution queries. The EBM achieves superior abstention performance abstention on semantically hard cases, reaching AUROC 0.961 versus 0.950 for softmax, while also reducing FPR@95 (0.235 vs 0.331). On easy negatives, performance is comparable across methods, but the EBM's advantage becomes most pronounced in safety-critical hard distributions. A comprehensive ablation with controlled negative sampling and fair data exposure shows that robustness stems primarily from the energy scoring head, while the inclusion or exclusion of specific negative types (hard, easy, mixed) sharpens decision boundaries but is not essential for generalisation to hard cases. These results demonstrate that energy-based abstention scoring offers a more reliable confidence signal than probability-based softmax confidence, providing a scalable and interpretable foundation for safe RAG systems.

Added

2026-10-04

Training Transformers for KV Cache Compressibility

Training Transformers for KV Cache Compressibility

Yoav Gelberg, Yam Eitan, Michael Bronstein, Yarin Gal, Haggai Maron

OrganizationsAITHYRANVIDIATechnion – Israel Institute of TechnologyUniversity of Oxford

Why you should read this

Introduces KV-Compression Aware Training (KV-CAT), a pretraining method that masks key-value slots during training to produce representations that substantially improve the performance of downstream KV cache compression algorithms on long-context tasks.

Long-context language modeling is increasingly constrained by the Key-Value (KV) cache, whose memory and decode-time access costs scale linearly with the prefix length. This bottleneck has motivated a range of context-compression methods, from token-level summarization to recent optimization-based KV compression methods. These post-hoc methods operate on the KV cache of a fixed pretrained model, so their effectiveness is fundamentally limited by how well the model's internal representations can be compressed. In this work, we formalize the notion of KV compressibility and show that it is a property of the learned representations, rather than of the context alone. We prove that almost any sequence-to-vector function admits both highly compressible and inherently non-compressible transformer implementations, highlighting the need to guide transformers toward compressible representations during training. Motivated by this, we propose KV-Compression Aware Training (KV-CAT), a continued pretraining procedure that incentivizes the emergence of compressible representations. We introduce a train-time KV sparsification policy that masks KV slots during training. This forces the model to use fewer KV slots and encourages it to learn representations amenable to post-hoc compression. Empirically, we show that KV-CAT improves the quality-budget tradeoff of downstream compression methods across retrieval, long-context question answering, and perplexity-based evaluation of compressed-prefix continuation.

Added

2026-10-04

Context-DPO: Aligning Language Models for Context-Faithfulness

Context-DPO: Aligning Language Models for Context-Faithfulness

Baolong Bi, Shaohan Huang, Yiwei Wang, Tianchi Yang, Zihan Zhang, Haizhen Huang, Lingrui Mei, Junfeng Fang, Zehao Li, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, Shenghua Liu

OrganizationsMicrosoftNational University of SingaporeUniversity of California, MercedUniversity of Chinese Academy of Sciences

Why you should read this

Proposes Context-DPO, a direct preference optimization method that resolves knowledge conflicts in retrieval-augmented generation and increases language model context-faithfulness by up to 280% without degrading general generative performance.

Reliable responses from large language models (LLMs) require adherence to user instructions and retrieved information. While alignment techniques help LLMs align with human intentions and values, improving context-faithfulness through alignment remains underexplored. To address this, we propose Context-DPO\textbf{Context-DPO}, the first alignment method specifically designed to enhance LLMs' context-faithfulness. We introduce ConFiQA\textbf{ConFiQA}, a benchmark that simulates Retrieval-Augmented Generation (RAG) scenarios with knowledge conflicts to evaluate context-faithfulness. By leveraging faithful and stubborn responses to questions with provided context from ConFiQA, our Context-DPO aligns LLMs through direct preference optimization. Extensive experiments demonstrate that our Context-DPO significantly improves context-faithfulness, achieving 35% to 280% improvements on popular open-source models. Further analysis demonstrates that Context-DPO preserves LLMs' generative capabilities while providing interpretable insights into context utilization. Our code and data are released at this https URL

Added

2026-10-04

Creative Commons License
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang

OrganizationsMThreads, Inc.

Why you should read this

Proposes a hybrid offline-online framework that precomputes chunk-level key-value caches and stitches them at inference using independent attention and reordered rotary position embeddings, accelerating time-to-first-token in retrieval-augmented generation by up to 9.4x without sacrificing accuracy or requiring architectural modifications.

Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of online computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we introduce TurboRAG, a hybrid offline-online paradigm that (i) pre-computes chunk-level key-value (KV) caches, (ii) stitches them together at inference time using independent-attention and reordered-RoPE techniques, and (iii) preserves answer quality without changing the model architecture. Our approach is applicable to most existing large language models and their applications without any requirement in modification of models and inference systems. Experimental results across a suite of RAG benchmarks demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG systems (on an average of 8.6x), but reserving comparable performance to the standard RAG systems.

Added

2026-10-04

CompAct: Compressing Retrieved Documents Actively for Question Answering

CompAct: Compressing Retrieved Documents Actively for Question Answering

Chanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong, Jaewoo Kang

OrganizationsAIGEN SciencesKorea UniversityUpstage AI

Why you should read this

Presents CompAct, an active context compression framework that dynamically integrates multi-hop evidence across retrieved documents and applies early termination, achieving up to 47x compression while boosting reader accuracy on complex question-answering benchmarks.

Retrieval-augmented generation supports language models to strengthen their factual groundings by providing external contexts. However, language models often face challenges when given extensive information, diminishing their effectiveness in solving questions. Context compression tackles this issue by filtering out irrelevant information, but current methods still struggle in realistic scenarios where crucial information cannot be captured with a single-step approach. To overcome this limitation, we introduce CompAct, a novel framework that employs an active strategy to condense extensive documents without losing key information. Our experiments demonstrate that CompAct brings significant improvements in both performance and compression rate on multi-hop question-answering benchmarks. CompAct flexibly operates as a cost-efficient plug-in module with various off-the-shelf retrievers or readers, achieving exceptionally high compression rates (47x).

Added

2026-10-03

MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation

MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation

Chia-Yuan Chang, Zhimeng Jiang, Vineeth Rakesh, Menghai Pan, Chin-Chia Michael Yeh, Guanchu Wang, Mingzhi Hu, Zhichao Xu, Yan Zheng, Mahashweta Das, Na Zou

OrganizationsRice UniversityTexas A&M UniversityUniversity of HoustonUniversity of UtahVisa ResearchWorcester Polytechnic Institute

Why you should read this

Proposes a training-free multi-agent framework that uses dynamic score thresholds to filter out noisy retrieved documents, boosting question-answering accuracy by up to 11% without model fine-tuning.

Large Language Models (LLMs) are becoming essential tools for various natural language processing tasks but often suffer from generating outdated or incorrect information. Retrieval-Augmented Generation (RAG) addresses this issue by incorporating external, real-time information retrieval to ground LLM responses. However, the existing RAG systems frequently struggle with the quality of retrieval documents, as irrelevant or noisy documents degrade performance, increase computational overhead, and undermine response reliability. To tackle this problem, we propose Multi-Agent Filtering Retrieval-Augmented Generation (MAIN-RAG), a training-free RAG framework that leverages multiple LLM agents to collaboratively filter and score retrieved documents. Specifically, MAIN-RAG introduces an adaptive filtering mechanism that dynamically adjusts the relevance filtering threshold based on score distributions, effectively minimizing noise while maintaining high recall of relevant documents. The proposed approach leverages inter-agent consensus to ensure robust document selection without requiring additional training data or fine-tuning. Experimental results across four QA benchmarks demonstrate that MAIN-RAG consistently outperforms traditional RAG approaches, achieving a 2–11% improvement in answer accuracy while reducing the number of irrelevant retrieved documents. Quantitative analysis further reveals that our approach achieves superior response consistency and answer accuracy over baseline methods, offering a competitive and practical alternative to training-based solutions.

Added

2026-10-03

Instruction-tuned Language Models are Better Knowledge Learners

Instruction-tuned Language Models are Better Knowledge Learners

Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodríguez, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Wen-tau Yih, Srini Iyer

OrganizationsCarnegie Mellon UniversityMetaUniversity of Washington

Why you should read this

Proposes pre-instruction-tuning to train language models on question-answer pairs before continued pre-training on new documents, significantly improving their ability to retain and accurately answer questions about newly acquired factual knowledge.

In order for large language model (LLM)-based assistants to effectively adapt to evolving information needs, it must be possible to update their factual knowledge through continued training on new data. The standard recipe for doing so involves continued pre-training on new documents followed by instruction-tuning on question-answer (QA) pairs. However, we find that LLMs trained with this recipe struggle to answer questions, even though the perplexity of documents is minimized. We found that QA pairs are generally straightforward, while documents are more complex, weaving many factual statements together in an intricate manner. Therefore, we hypothesize that it is beneficial to expose LLMs to QA pairs before continued pre-training on documents so that the process of encoding knowledge from complex documents takes into account how this knowledge is accessed through questions. Based on this, we propose pre-instruction-tuning (PIT), a method that instruction-tunes on questions prior to training on documents. This contrasts with standard instruction-tuning, which learns how to extract knowledge after training on documents. Extensive experiments and ablation studies demonstrate that PIT significantly enhances the ability of LLMs to absorb knowledge from new documents, outperforming standard instruction-tuning by 17.8%.

Added

2026-10-02

Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Zhengliang Shi, Shuo Zhang, Weiwei Sun, Shen Gao, Pengjie Ren, Zhumin Chen, Zhaochun Ren

Why you should read this

Proposes GenGround, a framework that counters noisy retrieval in multi-hop question answering by having large language models generate intermediate answers first and then revise them against retrieved evidence, supplemented by a distillation technique that transfers this capability to smaller models.

Multi-Hop Question Answering (MHQA) tasks present a significant challenge for large language models (LLMs) due to the intensive knowledge required. Current solutions, like Retrieval-Augmented Generation, typically retrieve potential documents from an external corpus to read an answer. However, the performance of this retrieve-then-read paradigm is constrained by the retriever and the inevitable noise in the retrieved documents. To mitigate these challenges, we introduce a novel generate-then-ground (GenGround) framework, synergizing the parametric knowledge of LLMs and external documents to solve a multi-hop question. GenGround empowers LLMs to alternate two phases until the final answer is derived: (1) formulate a simpler, single-hop question and directly generate the answer; (2) ground the question-answer pair in retrieved documents, amending any wrong predictions in the answer. We also propose an instructional grounding distillation method to generalize our method into smaller models. Extensive experiments conducted on four datasets illustrate the superiority of our method.

Added

2026-10-02

Bridging the Preference Gap between Retrievers and LLMs

Bridging the Preference Gap between Retrievers and LLMs

Zixuan Ke, Weize Kong, Cheng Li, Mingyang Zhang, Qiaozhu Mei, Michael Bendersky

OrganizationsGoogleUniversity of Illinois ChicagoUniversity of Michigan

Why you should read this

Proposes a sequence-to-sequence bridge framework that connects frozen retrievers and large language models, training on supervised and reinforcement learning to select and reorder passages according to the model's actual context preferences rather than human ranking assumptions.

Large Language Models (LLMs) have demonstrated superior results across a wide range of tasks, and Retrieval-augmented Generation (RAG) is an effective way to enhance the performance by locating relevant information and placing it into the context window of the LLM. However, the relationship between retrievers and LLM in a RAG is still under-investigated. Most existing work treats the retriever and the LLM as independent components and leaves a gap between retrieving human-“friendly” information and assembling a LLM-“friendly” context. In this work, we examine a novel bridge mechanism. We validate the ranking and selection assumptions of retrievers in the context of RAG and propose a framework that chains together supervised and reinforcement learning to train a bridge model that optimizes the connection between the retriever and the LLM. Empirical results demonstrate the effectiveness of our method in both question-answering and personalized generation tasks.

Added

2026-10-02

DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, Jun Wang

OrganizationsJilin UniversityShanghai Jiao Tong UniversityUniversity College London

Why you should read this

Proposes DS-Agent, a framework integrating case-based reasoning with large language models to iteratively retrieve, adapt, and refine expert Kaggle solutions for automated machine learning pipeline development at minimal computational cost.

In this work, we investigate the potential of large language models (LLMs) based agents to automate data science tasks, with the goal of comprehending task requirements, then building and training the best-fit machine learning models. Despite their widespread success, existing LLM agents are hindered by generating unreasonable experiment plans within this scenario. To this end, we present DS-Agent, a novel automatic framework that harnesses LLM agent and case-based reasoning (CBR). In the development stage, DS-Agent follows the CBR framework to structure an automatic iteration pipeline, which can flexibly capitalize on the expert knowledge from Kaggle, and facilitate consistent performance improvement through the feedback mechanism. Moreover, DS-Agent implements a low-resource deployment stage with a simplified CBR paradigm to adapt past successful solutions from the development stage for direct code generation, significantly reducing the demand on foundational capabilities of LLMs. Empirically, DS-Agent with GPT-4 achieves 100% success rate in the development stage, while attaining 36% improvement on average one pass rate across alternative LLMs in the deployment stage. In both stages, DS-Agent achieves the best rank in performance, costing 1.60and1.60 and0.13 per run with GPT-4, respectively. Our data and code are open-sourced at https://github.com/guosyjlu/DS-Agent.

Added

2026-10-01

From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

Bernal Jimnez Gutirrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su

OrganizationsThe Ohio State UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes HippoRAG 2, a non-parametric continual learning framework that integrates knowledge graphs with Personalized PageRank and online language model reasoning to outperform standard retrieval-augmented generation across factual, sense-making, and associative memory tasks.

Our ability to continuously acquire, organize, and leverage knowledge is a key feature of human intelligence that AI systems must approximate to unlock their full potential. Given the challenges in continual learning with large language models (LLMs), retrieval-augmented generation (RAG) has become the dominant way to introduce new information. However, its reliance on vector retrieval hinders its ability to mimic the dynamic and interconnected nature of human long-term memory. Recent RAG approaches augment vector embeddings with various structures like knowledge graphs to address some of these gaps, namely sense-making and associativity. However, their performance on more basic factual memory tasks drops considerably below standard RAG. We address this unintended deterioration and propose HippoRAG 2, a framework that outperforms standard RAG comprehensively on factual, sense-making, and associative memory tasks. HippoRAG 2 builds upon the Personalized PageRank algorithm used in HippoRAG and enhances it with deeper passage integration and more effective online use of an LLM. This combination pushes this RAG system closer to the effectiveness of human long-term memory, achieving a 7% improvement in associative memory tasks over the state-of-the-art embedding model while also exhibiting superior factual knowledge and sense-making memory capabilities. This work paves the way for non-parametric continual learning for LLMs. Code and data are available at https://github.com/OSU-NLP-Group/HippoRAG.

Added

2026-10-01

Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui

OrganizationsGoogleHarvard UniversityMeta

Why you should read this

Introduces FRAMES, a benchmark of multi-hop questions requiring information synthesis across multiple documents to evaluate retrieval-augmented generation systems simultaneously on factuality, retrieval, and complex reasoning.

Large Language Models (LLMs) have shown significant improvements across cognitive tasks, with an emerging application in enhancing retrieval-augmented generation (RAG) capabilities. These systems require LLMs to understand queries, retrieve relevant information, and synthesize accurate responses. Given their increasing real-world deployment, comprehensive evaluation is crucial. We propose FRAMES (Factuality, Retrieval, And reasoning MEasurement Set), a high-quality dataset designed to test LLMs’ factual responses, retrieval capabilities, and reasoning in generating final answers. Unlike previous work evaluating these abilities in isolation, FRAMES offers a unified framework for assessing LLM performance in end-to-end RAG scenarios. Our dataset comprises challenging multi-hop questions requiring integration of information from multiple sources. Baseline results show that even state-of-the-art LLMs struggle, achieving 0.408 accuracy without retrieval. However, our proposed multi-step retrieval pipeline significantly improves accuracy to 0.66 (>50% improvement). We aim to bridge evaluation gaps and assist in developing more robust RAG systems.

Added

2026-10-01