Built independently by an author, for readers. Read the story and support ChapterPal

keyword

parametric knowledge

Parametric knowledge refers to the factual information, patterns, and concepts that an artificial intelligence model, such as a large language model, encodes and stores directly within its internal weights and parameters during training. Unlike non-parametric knowledge, which is dynamically retrieved from external documents, databases, or provided contexts at inference time, parametric knowledge acts as the model's internal memory and remains fixed unless the model is updated or retrained. This stored knowledge enables the system to generate answers, complete text, and perform reasoning tasks autonomously, though it remains inherently bounded by the data seen during training and can become outdated or prone to factual errors when addressing less common, rapidly changing, or missing information.

8 items

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

OrganizationsAllen Institute for AIStony Brook University

Why you should read this

Proposes IRCoT, a framework that interleaves chain-of-thought reasoning with step-by-step external retrieval to reduce model hallucinations and improve multi-step open-domain question answering across both large and small language models without extra training.

Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA). They struggle, however, when the necessary knowledge is either unavailable to the LLM or not up-to-date within its parameters. While using the question to retrieve relevant text from an external knowledge source helps LLMs, we observe that this one-step retrieve-and-read approach is insufficient for multi-step QA. Here, what to retrieve depends on what has already been derived, which in turn may depend on what was previously retrieved. To address this, we propose IRCoT, a new approach for multi-step QA that interleaves retrieval with steps (sentences) in a CoT, guiding the retrieval with CoT and in turn using retrieved results to improve CoT. Using IRCoT with GPT3 substantially improves retrieval (up to 21 points) as well as downstream QA (up to 15 points) on four datasets: HotpotQA, 2WikiMultihopQA, MuSiQue, and IIRC. We observe similar substantial gains in out-of-distribution (OOD) settings as well as with much smaller models such as Flan-T5-large without additional training. IRCoT reduces model hallucination, resulting in factually more accurate CoT reasoning.¹

Added

2026-10-05

Merging Generated and Retrieved Knowledge for Open-Domain QA

Merging Generated and Retrieved Knowledge for Open-Domain QA

Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang

OrganizationsLG AI ResearchUniversity of Illinois ChicagoUniversity of Michigan

Why you should read this

Proposes a compatibility-oriented framework that pairs LLM-generated texts with retrieved documents to resolve knowledge conflicts and improve open-domain question answering accuracy.

Open-domain question answering (QA) systems are often built with retrieval modules. However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage. Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance. Yet, LLMs tend to “hallucinate” content that conflicts with the retrieved knowledge. Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented Knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information. Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels. Then a Fusion-in-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer. Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks. Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts.¹

Added

2026-10-03

Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering

Zhengliang Shi, Shuo Zhang, Weiwei Sun, Shen Gao, Pengjie Ren, Zhumin Chen, Zhaochun Ren

Why you should read this

Proposes GenGround, a framework that counters noisy retrieval in multi-hop question answering by having large language models generate intermediate answers first and then revise them against retrieved evidence, supplemented by a distillation technique that transfers this capability to smaller models.

Multi-Hop Question Answering (MHQA) tasks present a significant challenge for large language models (LLMs) due to the intensive knowledge required. Current solutions, like Retrieval-Augmented Generation, typically retrieve potential documents from an external corpus to read an answer. However, the performance of this retrieve-then-read paradigm is constrained by the retriever and the inevitable noise in the retrieved documents. To mitigate these challenges, we introduce a novel generate-then-ground (GenGround) framework, synergizing the parametric knowledge of LLMs and external documents to solve a multi-hop question. GenGround empowers LLMs to alternate two phases until the final answer is derived: (1) formulate a simpler, single-hop question and directly generate the answer; (2) ground the question-answer pair in retrieved documents, amending any wrong predictions in the answer. We also propose an instructional grounding distillation method to generalize our method into smaller models. Extensive experiments conducted on four datasets illustrate the superiority of our method.

Added

2026-10-02

From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

From RAG to Memory: Non-Parametric Continual Learning for Large Language Models

Bernal Jimnez Gutirrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su

OrganizationsThe Ohio State UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes HippoRAG 2, a non-parametric continual learning framework that integrates knowledge graphs with Personalized PageRank and online language model reasoning to outperform standard retrieval-augmented generation across factual, sense-making, and associative memory tasks.

Our ability to continuously acquire, organize, and leverage knowledge is a key feature of human intelligence that AI systems must approximate to unlock their full potential. Given the challenges in continual learning with large language models (LLMs), retrieval-augmented generation (RAG) has become the dominant way to introduce new information. However, its reliance on vector retrieval hinders its ability to mimic the dynamic and interconnected nature of human long-term memory. Recent RAG approaches augment vector embeddings with various structures like knowledge graphs to address some of these gaps, namely sense-making and associativity. However, their performance on more basic factual memory tasks drops considerably below standard RAG. We address this unintended deterioration and propose HippoRAG 2, a framework that outperforms standard RAG comprehensively on factual, sense-making, and associative memory tasks. HippoRAG 2 builds upon the Personalized PageRank algorithm used in HippoRAG and enhances it with deeper passage integration and more effective online use of an LLM. This combination pushes this RAG system closer to the effectiveness of human long-term memory, achieving a 7% improvement in associative memory tasks over the state-of-the-art embedding model while also exhibiting superior factual knowledge and sense-making memory capabilities. This work paves the way for non-parametric continual learning for LLMs. Code and data are available at https://github.com/OSU-NLP-Group/HippoRAG.

Added

2026-10-01

Factuality of Large Language Models: A Survey

Factuality of Large Language Models: A Survey

Yuxia Wang, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, Preslav Nakov

OrganizationsGoogleMohamed bin Zayed University of Artificial IntelligenceMonash UniversitySofia University “St. Kliment Ohridski”

Why you should read this

Synthesizes recent advances in large language model factuality across text and vision modalities by categorizing evaluation benchmarks, clarifying distinctions between factuality and hallucination, and analyzing error-mitigation techniques alongside calibration strategies.

Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.

Added

2026-09-26

Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering

Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, Pasquale Minervini

OrganizationsMiniml.AISapienza University of RomeThe Chinese University of Hong KongUniversity College LondonUniversity of Edinburgh

Why you should read this

Introduces SPARE, a training-free representation engineering method that leverages sparse auto-encoders to detect mid-layer conflict signals and steer whether large language models rely on parametric memory or contextual evidence during question answering.

Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context—this phenomenon, known as context-memory knowledge conflicts 1, can lead to undesirable model behaviour, such as reliance on outdated or incorrect information. Analysing the internal activations of LLMs, we find that they can internally register the signals of knowledge conflict at mid-layers. Such signals allow us to detect whether a knowledge conflict occurs and use inference-time intervention strategies to resolve it. In this work, we propose SPARE, a training-free representation engineering method that uses pre-trained sparse auto-encoders (SAEs) to control the knowledge selection behaviour of LLMs. SPARE identifies the functional features that control the knowledge selection behaviours and applies them to edit the internal activations of LLMs at inference time. Our experimental results show that SPARE can effectively control the usage of either knowledge source to resolve knowledge conflict in open-domain question-answering tasks, surpassing existing representation engineering methods (+10%) as well as contrastive decoding methods (+15%).

Added

2026-09-26

Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models

Mosh Levy, Alon Jacoby, Yoav Goldberg

OrganizationsAllen Institute for AIBar-Ilan University

Why you should read this

Demonstrates through a controlled question-answering framework that large language models experience sharp declines in multi-step reasoning performance as context length increases, failing well before reaching their technical context limits regardless of padding type or fact placement.

This paper explores the impact of extending input lengths on the capabilities of Large Language Models (LLMs). Despite LLMs advancements in recent times, their performance consistency across different input lengths is not well understood. We investigate this aspect by introducing a novel QA reasoning framework, specifically designed to assess the impact of input length. We isolate the effect of input length using multiple versions of the same sample, each being extended with padding of different lengths, types and locations. Our findings show a notable degradation in LLMs’ reasoning performance at much shorter input lengths than their technical maximum. We show that the degradation trend appears in every version of our dataset, although at different intensities. Additionally, our study reveals that the traditional metric of next word prediction correlates negatively with performance of LLMs’ on our reasoning dataset. We analyse our results and identify failure modes that can serve as useful guides for future research, potentially informing strategies to address the limitations observed in LLMs.

Added

2026-09-26

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Alex Troy Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, Daniel Khashabi

OrganizationsAllen Institute for AIJohns Hopkins UniversityUniversity of Washington

Why you should read this

Reveals that scaling language models fails to resolve factual errors on long-tail knowledge and introduces an adaptive retrieval strategy on the PopQA benchmark that queries external memory only when needed, significantly cutting inference costs while improving factual accuracy.

Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

Added

2026-09-25