Confabulation: The Surprising Value of Large Language Model Hallucinations
Peiqi SuiEamon DuedeSophie WuRichard Jean So
Demonstrates that large language model hallucinations exhibit higher narrativity and semantic coherence than factual outputs, reframing these errors as confabulations that drive coherent story generation rather than purely harmful flaws.
As artificial intelligence systems become integral across high-stakes sectors such as law, medicine, finance, and science, large language model hallucinations—the generation of factually inaccurate or fabricated text—are widely treated as a major safety and reliability flaw. The standard industry response focuses on eliminating these errors entirely. However, emerging theoretical work suggests that hallucinations are statistically unavoidable and tied to creative text generation. The article addresses this core tension by investigating whether these inaccurate outputs serve an unacknowledged communicative purpose rather than functioning purely as system defects.
The main objective of the article is to evaluate the semantic properties of inaccurate language model outputs and demonstrate that what is commonly termed hallucination is better understood as confabulation: an impulse to generate structured, coherent narratives to fill information gaps, closely mirroring human sense-making.
To evaluate this relationship, the authors conducted a statistical analysis of tens of thousands of dialogue samples across three major benchmark datasets: FaithDial, BEGIN, and HaluEval. The study measured narrative content using a specialized classifier trained for story detection and evaluated conversational coherence across more than 65,000 samples using an automated dialogue evaluation metric. The researchers used logistic and beta regression models to examine the relationships among factual accuracy, storytelling characteristics, and conversational coherence.
The analysis produced three primary findings. First, across all three benchmarks, inaccurate model outputs consistently exhibited higher average narrativity scores than their truthful or human-edited counterparts. Second, statistical modeling showed that higher narrativity is a significant positive predictor of whether an output contains factual errors. Third, higher narrativity strongly correlated with increased conversational coherence across the entire dialogue dataset. Rather than producing disjointed or random falsehoods, models generate structured, self-consistent narratives when information is missing.
These findings suggest that attempts to uniformly eliminate inaccuracies may inadvertently degrade a model's ability to generate coherent, fluent, and persuasive narrative text. Because human communication relies heavily on storytelling to establish context and build common understanding, an absolute suppression of confabulation poses risks to model performance in applications that depend on creative synthesis, hypothesis generation, or high user engagement. While strictly factual domains such as legal compliance or medical record retrieval still require strict accuracy controls, other applications may benefit from balancing factuality against narrative richness.
Organizations should reconsider one-size-fits-all strategies that treat all factual inaccuracies as categorical failures. Decision-makers should tailor model parameters and mitigation pipelines according to the intended use case, distinguishing truth-critical tasks from exploratory, creative, or communication-centric applications. Further research and user studies with human participants are needed to confirm whether human end-users actually realize tangible experience and comprehension benefits from engaging with high-narrative confabulations.
Confidence in the statistical correlation between inaccuracy, narrativity, and coherence across the studied dialogue benchmarks is high. However, readers should remain cautious regarding causal claims: the current evidence demonstrates strong association rather than direct causation, and the findings are bounded by the automated metrics and conversational datasets evaluated in the article.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Provides a comprehensive taxonomy and foundational principles of hallucination in large language models that contextualize the source's reinterpretation of factual errors as communicative confabulation.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). Establishes core definitions and categorizations of intrinsic versus extrinsic hallucinations in natural language generation, laying the conceptual groundwork for analyzing model generation trade-offs.
- Paper: LaMDA: Language Models for Dialog Applications, Romal Thoppilan et al. (2022). Explores the fundamental tension between factual grounding and conversational quality metrics like sensibleness and interestingness in dialogue systems.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Introduces benchmark methodologies for measuring how language models generate plausible falsehoods and mimic human-like errors rather than adhering strictly to truth.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). Demonstrates early empirical evidence that neural models frequently fabricate fluent, extrinsic details to bridge information gaps in generated text.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Critiques conventional evaluation practices in natural language generation, highlighting why surface fluency often masks underlying factual inaccuracies.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). Offers an extensive subsequent survey on LLM factuality, expanding on the tensions between sentence probability, conversational fluency, and objective truth.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). Applies targeted preference optimization pipelines to strictly mitigate the factual errors that the source identifies as deeply entangled with narrative coherence.
- Paper: Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?, Gal Yona et al. (2024). Investigates how linguistic phrasing and verbalized confidence relate to actual model certainty, complementing the source's study on narrative framing in inaccurate text.
- Paper: Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation, Dongjin Kang et al. (2024). Examines conversational and emotional engagement strategies in dialogue models, providing a practical setting where narrative coherence often clashes with factual strictness.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Extends the evaluation of narrative alignment and conversational consistency across very long-term conversational memory settings.
