How Language Model Hallucinations Can Snowball
Muru ZhangOfir PressWilliam MerrillAlisa LiuNoah A. Smith
Reveals that large language models frequently invent false justifications to maintain consistency with their own earlier mistakes, generating secondary errors that they are otherwise capable of correctly identifying as false in isolation.
As large language models are increasingly deployed in real-world information retrieval and decision-making systems, their tendency to produce plausible but false statements presents a major operational and safety risk. While conventional industry wisdom assumes these errors stem from gaps in training knowledge, the article demonstrates that language models frequently invent incorrect claims that they can separately recognize as false in isolation. This behavior, termed hallucination snowballing, occurs when a model commits to an incorrect initial response and subsequently manufactures false supporting details to maintain internal consistency.
The article evaluates why and how frequently this phenomenon occurs across leading commercial and open-source models, specifically testing GPT-3.5, GPT-4, and LLaMA2-70B-chat. The authors designed three targeted question-answering datasets of 500 questions each spanning distinct domains: primality testing, biographical senator verification, and multi-step flight connectivity. In a two-stage evaluation, models were first evaluated on whether they answered the queries correctly under standard zero-shot prompting. When an incorrect answer was produced, the supporting claims were extracted and fed back into the same model in separate sessions to verify whether the model possessed the correct underlying knowledge.
The findings reveal that models struggle heavily with single-step commitments on sequential reasoning tasks, exhibiting average error rates between 60% and 83%. More crucially, when models generated incorrect explanations to support their false answers, they were capable of identifying their own supporting claims as false 67.37% of the time for GPT-3.5, 87.03% for GPT-4, and 93.67% for LLaMA2-70B-chat. Standard decoding adjustments, including higher sampling temperatures and beam search, failed to alleviate this behavior. While zero-shot chain-of-thought prompting ("Let's think step by step") lowered error rates on simple benchmarks for GPT-3.5 and GPT-4, it proved ineffective on composite, multi-part reasoning questions, where both models still failed on more than half of the questions and continued to experience snowballed errors in over 65% of those failures.
These results indicate that popular deployment mitigations, such as knowledge retrieval augmentation, are insufficient to solve hallucinations because the root problem is structural rather than purely informational. Models prioritize output consistency over factuality once an initial token commitment is made. For organizations utilizing language models, this creates hidden compliance and performance risks, as confident justifications cannot be treated as evidence of factual reasoning.
To address this systemic risk, the article suggests re-evaluating model training and generation workflows. System developers should implement architectures that force explicit reasoning chains prior to answering, and explore fine-tuning techniques that teach models how to backtrack and revise prior statements rather than forcing consistency. In the interim, decision-makers should recognize that model self-confidence is fragile, and critical applications should avoid relying on single-pass model outputs without independent, modular verification pipelines.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This paper establishes that language models possess internal calibration to separate truth from falsehood in their own outputs, providing the foundational premise for investigating why models can recognize errors yet still hallucinate in multi-step generation.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This work demonstrates how step-by-step reasoning can generate unfaithful explanations that rationalize incorrect choices, directly motivating the study of error cascading during chain-of-thought generation.
- Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models, Potsawee Manakul et al. (2023). This paper develops sampling-based self-checking methods to evaluate factuality without external knowledge, establishing key mechanisms used to probe model self-awareness of hallucinations.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). This study demonstrates that multi-turn inquiry can expose internal factual inconsistencies in language models, setting up the empirical basis for interrogating claims that models can self-detect as false.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). This research proves that models maintain latent factual representations separate from their generated surface text, directly informing how models can simultaneously produce and detect hallucinations.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). This survey establishes the standard taxonomy of factuality and faithfulness hallucinations in large language models necessary for contextualizing downstream generation errors.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). This paper provides formal analysis showing that language models make locally valid deduction steps while failing global reasoning paths, underpinning the dynamics of multi-step snowballing errors.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This foundational benchmark demonstrates how language models mimic human falsehoods and fail truthfulness, providing baseline problem formulations for hallucination analysis.
- Paper: Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation, Xiaoying Zhang et al. (2024). This paper leverages the fact that language models possess internal knowledge to recognize their own errors by using self-evaluation preference data to align models against hallucinations.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). This work extends the finding that models implicitly recognize their hallucinations by building lightweight internal probe circuits to predict generation failures in real time.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). This study develops an automated Direct Preference Optimization framework that exploits internal model confidence to train models out of multi-sentence factual hallucination cascades.
- Paper: Confabulation: The Surprising Value of Large Language Model Hallucinations, Peiqi Sui et al. (2024). This article analyzes the narrative and communicative dynamics of snowballing inaccuracies, reframing cascading hallucinations as coherent narrative confabulation.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This paper proposes an autonomous retrofitting framework that detects and corrects cascading intermediate errors in multi-step reasoning using structured knowledge graphs.
- Paper: Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?, Gal Yona et al. (2024). This research builds on self-knowledge discrepancies by evaluating whether models can faithfully communicate their internal uncertainty in natural language.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). This work introduces representation engineering using sparse auto-encoders to steer whether a model relies on internal memory versus external context during knowledge conflicts.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). This survey synthesizes modern findings on factual errors, evaluation frameworks, and post-generation mitigation strategies across text and multimodal foundation models.
