Built independently by an author, for readers. Read the story and support ChapterPal

keyword

retrieval-augmented LMs

Retrieval-augmented language models are artificial intelligence systems that integrate a standard language model with an external search or retrieval mechanism to fetch relevant documents from a knowledge base during the text generation process. Rather than relying entirely on the static factual information stored in their neural network parameters during training, these models dynamically query non-parametric data sources to ground their outputs in accurate and contextually relevant evidence. By combining internal generative capabilities with real-time access to external information, retrieval-augmented language models reduce factual errors and hallucinations, enhance performance on knowledge-intensive and long-tail domain queries, and allow knowledge bases to be updated continuously without requiring expensive model retraining.

1 item

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Alex Troy Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, Daniel Khashabi

OrganizationsAllen Institute for AIJohns Hopkins UniversityUniversity of Washington

Why you should read this

Reveals that scaling language models fails to resolve factual errors on long-tail knowledge and introduces an adaptive retrieval strategy on the PopQA benchmark that queries external memory only when needed, significantly cutting inference costs while improving factual accuracy.

Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

Added

2026-09-25