When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Alex Troy MallenAkari AsaiVictor ZhongRajarshi DasHannaneh HajishirziDaniel Khashabi
Reveals that scaling language models fails to resolve factual errors on long-tail knowledge and introduces an adaptive retrieval strategy on the PopQA benchmark that queries external memory only when needed, significantly cutting inference costs while improving factual accuracy.
Large language models often struggle to reliably answer questions that require fine-grained, real-world factual knowledge, leading to hallucinations and factual inaccuracies. The article investigates the limits of language models in memorizing factual knowledge, evaluates how scaling model size compares against augmenting models with external search retrieval, and introduces a method to selectively deploy retrieval only when needed.
The authors conducted a large-scale evaluation using ten language models across three model families—ranging from 1.3 billion parameters to large-scale GPT-3 models—tested on open-domain question answering. To evaluate factual knowledge across different frequency levels, the authors introduced PopQA, a new dataset of 14,000 entity-centric questions mapped against Wikipedia monthly page views as a proxy for subject popularity, and cross-evaluated findings on the EntityQuestions dataset. The analysis evaluated both unassisted model memory and semi-parametric configurations augmented with search retrievers such as BM25 and Contriever.
The investigation produced four primary findings. First, a language model's factual accuracy strongly correlates with subject entity popularity; models reliably recall well-known facts but struggle heavily with less frequent, long-tail entities. Second, simply scaling up model parameters fails to resolve long-tail gaps: for the 4,000 least popular questions in PopQA, increasing model size from 6 billion parameters to GPT-3 only improved accuracy from 16% to 19%. Third, external retrieval dramatically improves accuracy on rare entities—allowing a 2.7-billion-parameter model to outperform unassisted GPT-3 on tail knowledge—yet retrieval actually degrades GPT-3's accuracy by 10% on popular entities due to irrelevant or misleading search passages. Fourth, an "Adaptive Retrieval" strategy, which consults external documents only when entity popularity falls below a specific threshold, improved GPT-3 accuracy by up to 5.3% over fixed retrieval pipelines while cutting GPT-3 inference API costs by roughly half.
These findings indicate that relying solely on increasing parameter size is an inefficient and ineffective approach for mastering long-tail world knowledge. Organizations deploying language models should not assume that frontier-scale models are factual across all domains, nor should they blindly attach retrieval pipelines to every prompt. Instead, engineering teams should implement adaptive retrieval systems that evaluate query characteristics beforehand to reduce infrastructure latency, lower operating expenses, and prevent misleading search contexts from confusing the underlying model.
Decision-makers should note that these experiments relied on synthetic, template-generated questions and used monthly Wikipedia page views as a static proxy for entity popularity. While confidence in the core trade-off between model scale, popularity, and retrieval utility is high, further validation on complex, real-world enterprise queries and alternative domain-specific popularity metrics is recommended before final production deployment.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This paper establishes the foundational LAMA probing benchmark that measures how much factual knowledge is stored implicitly within language model parameters.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This work introduces the hybrid retrieval-augmented generation framework combining parametric and non-parametric memory that serves as the baseline architecture investigated in the source.
- Paper: REALM: Retrieval-Augmented Language Model Pre-Training, Kelvin Guu et al. (2020). This study demonstrates how augmenting language models with neural retrieval over external corpora mitigates the limitations of purely parametric knowledge storage.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). This paper establishes the Fusion-in-Decoder architecture for open-domain question answering, providing key empirical evidence on how generative models utilize retrieved text passages.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). This research evaluates the efficacy and limitations of probing pre-trained language models to extract their internal factual associations.
- Paper: Generalization through Memorization: Nearest Neighbor Language Models, Urvashi Khandelwal et al. (2020). This work introduces nearest-neighbor language modeling, formalizing the distinction between implicit parametric memorization and explicit non-parametric lookup.
- Paper: Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen et al. (2017). This seminal work establishes open-domain question answering by reading Wikipedia, defining the retrieve-and-read paradigm for factual information access.
- Paper: WebGPT: Browser-assisted question-answering with human feedback, Reiichiro Nakano et al. (2021). This paper investigates integrating search-engine retrieval with large language models to overcome factual inaccuracies and hallucinations in parametric generation.
- Paper: Active Retrieval Augmented Generation, Zhengbao Jiang et al. (2023). This work extends the concept of selective retrieval by introducing an active mechanism (FLARE) that dynamically retrieves external knowledge whenever model generation uncertainty is detected.
- Paper: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection, Akari Asai et al. (2024). This research builds upon adaptive retrieval principles by training language models with reflection tokens to evaluate when retrieval is required and critique retrieved content.
- Paper: Retrieval-Augmented Generation for Large Language Models: A Survey, Yunfan Gao et al. (2023). This survey provides a comprehensive synthesis of modern retrieval-augmented generation paradigms, categorizing adaptive and modular extensions developed after initial factual probing studies.
- Paper: REPLUG: Retrieval-Augmented Black-Box Language Models, Weijia Shi et al. (2024). This paper generalizes non-parametric memory integration by designing an ensemble retrieval-augmentation method specifically tailored for frozen, black-box language models.
- Paper: Atlas: Few-shot Learning with Retrieval Augmented Language Models, Gautier Izacard et al. (2023). This study advances retrieval-augmented architectures by pre-training unified dense retrieval and sequence-to-sequence readers to achieve state-of-the-art few-shot knowledge performance.
- Paper: Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning, Bowen Jin et al. (2025). This work scales dynamic knowledge retrieval by using reinforcement learning to train models to autonomously interleave multi-turn search interactions with reasoning.
- Paper: Search-o1: Agentic Search-Enhanced Large Reasoning Models, Xiaoxi Li et al. (2025). This paper extends on-demand retrieval strategies to complex reasoning models, enabling agentic web querying when internal parametric confidence is low during multi-step problem solving.
- Paper: In-Context Retrieval-Augmented Language Models, Ori Ram et al. (2023). This work investigates lightweight non-parametric augmentation by prepending retrieved documents into the context window of frozen language models without architectural modifications.
