Confabulation: The Surprising Value of Large Language Model Hallucinations

Peiqi SuiEamon DuedeSophie WuRichard Jean So

article2024ACL52 citations

Demonstrates that large language model hallucinations exhibit higher narrativity and semantic coherence than factual outputs, reframing these errors as confabulations that drive coherent story generation rather than purely harmful flaws.

Listen

As artificial intelligence systems become integral across high-stakes sectors such as law, medicine, finance, and science, large language model hallucinations—the generation of factually inaccurate or fabricated text—are widely treated as a major safety and reliability flaw. The standard industry response focuses on eliminating these errors entirely. However, emerging theoretical work suggests that hallucinations are statistically unavoidable and tied to creative text generation. The article addresses this core tension by investigating whether these inaccurate outputs serve an unacknowledged communicative purpose rather than functioning purely as system defects.

The main objective of the article is to evaluate the semantic properties of inaccurate language model outputs and demonstrate that what is commonly termed hallucination is better understood as confabulation: an impulse to generate structured, coherent narratives to fill information gaps, closely mirroring human sense-making.

To evaluate this relationship, the authors conducted a statistical analysis of tens of thousands of dialogue samples across three major benchmark datasets: FaithDial, BEGIN, and HaluEval. The study measured narrative content using a specialized classifier trained for story detection and evaluated conversational coherence across more than 65,000 samples using an automated dialogue evaluation metric. The researchers used logistic and beta regression models to examine the relationships among factual accuracy, storytelling characteristics, and conversational coherence.

The analysis produced three primary findings. First, across all three benchmarks, inaccurate model outputs consistently exhibited higher average narrativity scores than their truthful or human-edited counterparts. Second, statistical modeling showed that higher narrativity is a significant positive predictor of whether an output contains factual errors. Third, higher narrativity strongly correlated with increased conversational coherence across the entire dialogue dataset. Rather than producing disjointed or random falsehoods, models generate structured, self-consistent narratives when information is missing.

These findings suggest that attempts to uniformly eliminate inaccuracies may inadvertently degrade a model's ability to generate coherent, fluent, and persuasive narrative text. Because human communication relies heavily on storytelling to establish context and build common understanding, an absolute suppression of confabulation poses risks to model performance in applications that depend on creative synthesis, hypothesis generation, or high user engagement. While strictly factual domains such as legal compliance or medical record retrieval still require strict accuracy controls, other applications may benefit from balancing factuality against narrative richness.

Organizations should reconsider one-size-fits-all strategies that treat all factual inaccuracies as categorical failures. Decision-makers should tailor model parameters and mitigation pipelines according to the intended use case, distinguishing truth-critical tasks from exploratory, creative, or communication-centric applications. Further research and user studies with human participants are needed to confirm whether human end-users actually realize tangible experience and comprehension benefits from engaging with high-narrative confabulations.

Confidence in the statistical correlation between inaccuracy, narrativity, and coherence across the studied dialogue benchmarks is high. However, readers should remain cautious regarding causal claims: the current evidence demonstrates strong association rather than direct causation, and the findings are bounded by the automated metrics and conversational datasets evaluated in the article.

arXiv: 2406.04175
Cover for Confabulation: The Surprising Value of Large Language Model Hallucinations

Abstract

This paper presents a systematic defense of large language model (LLM) hallucinations or ‘confabulations’ as a potential resource instead of a categorically negative pitfall. The standard view is that confabulations are inherently problematic and AI research should eliminate this flaw. In this paper, we argue and empirically demonstrate that measurable semantic characteristics of LLM confabulations mirror a human propensity to utilize increased narrativity as a cognitive resource for sense-making and communication. In other words, it has potential value. Specifically, we analyze popular hallucination benchmarks and reveal that hallucinated outputs display increased levels of narrativity and semantic coherence relative to veridical outputs. This finding reveals a tension in our usually dismissive understandings of confabulation. It suggests, counter-intuitively, that the tendency for LLMs to confabulate may be intimately associated with a positive capacity for coherent narrative-text generation.

Table of Contents

  • 1 Background
  • Definition of Confabulation
  • 2 Related Work
  • 2.1 Confabulation vs Hallucination
  • 2.2 Towards a Narrative-Centered Definition of Confabulation
  • 3 Data, Methods, and Results
  • 3.1 Empirical Results for Higher Narrativity in Hallucinations
  • 4 In defense of confabulation
  • 4.0.1 Empirical Support for Association of Narrativity and Coherence in Confabulated Texts
  • 4.1 Narrative, Discourse, and Coherence
  • 4.1.1 Narratives Help Us Articulate and Understand Complex Arguments
  • 4.1.2 Narratives Maintain the Consistency of Our Own Internal World Models
  • 4.1.3 Narratives Enable Patients to Negotiate the Coherence of Their Experiences
  • 5 Limitations and Directions for Future Research
  • 6 Acknowledgement
  • References

Knowls

  1. Knowl 1 — Narrative-Centered Definition of LLM Confabulation

    definition

    In natural language processing, large language model (LLM) confabulation is defined from a narrative-centric perspective as a latent narrative impulse to schematize information into self-consistent stories, particularly when available context or explicit details are insufficient. Under this definition, rather than being treated strictly as an erroneous or pathological failure mode, confabulation reflects compensatory storytelling analogous to human cognitive sense-making: an agent generates plausible, fictional, or unverified narrative details to maintain contextual coherence and semantic verisimilitude. The extent of this compensatory storytelling is operationalized as a continuous scalar construct termed narrativity.

  2. Knowl 2 — Softmax Probability Classifier for Quantifying Text Narrativity

    model/method

    To quantify the degree of narrativity in generated dialogue responses, a text classification approach is employed. An ELECTRA-large encoder model is fine-tuned on an expert-annotated narrative detection dataset sourced from online discourse (Reddit). For any input dialogue turn or text sequence xx, the narrativity score is operationalized as the scalar softmax probability output associated with the story class:

    Narrativity(x)=P(class=story∣x)∈[0,1]\text{Narrativity}(x) = P(\text{class} = \text{story} \mid x) \in [0, 1]

    Encoder models trained under this formulation achieve area under the ROC curve (AUC) performance between 0.830.83 and 0.850.85 on bootstrapped narrative test sets and between 0.810.81 and 0.840.84 on independent narrative evaluation benchmarks.

  3. Knowl 3 — Summary Statistics of Narrativity Scores Across Dialogue Hallucination Benchmarks

    data/table

    Across multiple dialogue benchmarks, text responses labeled as hallucinations consistently exhibit higher mean and median narrativity scores than partial hallucinations and truthful ground truth counterparts.

    FaithDial HaluEval BEGIN
    Metric Hallucinated Partial Truth Hallucinated Truth Hallucinated Partial Truth
    Count 4485 14108 2852 10000 10000 1019 239 1139
    Mean 0.620 0.606 0.518 0.655 0.638 0.658 0.612 0.561
    Std 0.178 0.178 0.186 0.138 0.168 0.183 0.188 0.187
    Min 0.041 0.052 0.061 0.167 0.029 0.064 0.073 0.059
    25% 0.500 0.486 0.379 0.563 0.528 0.541 0.477 0.433
    50% 0.640 0.624 0.524 0.669 0.664 0.695 0.642 0.567
    75% 0.756 0.741 0.664 0.759 0.768 0.799 0.761 0.698
    Max 0.975 0.974 0.959 0.952 0.985 0.965 0.956 0.965

    In FaithDial, mean narrativity drops from 0.6200.620 in hallucinated outputs to 0.6060.606 in partial hallucinations and 0.5180.518 in truth. In HaluEval, hallucinated responses average 0.6550.655 compared to 0.6380.638 for ground truth. In BEGIN, hallucinated responses average 0.6580.658, partial hallucinations average 0.6120.612, and truthful responses average 0.5610.561.

  4. Knowl 4 — Logistic Regression of Hallucination Classification on Narrativity

    empirical result

    The predictive relationship between a response's narrativity score xi∈[0,1]x_i \in [0, 1] and its binary classification as a hallucination yi∈{0,1}y_i \in \{0, 1\} (yi=1y_i = 1 for hallucinated or partially hallucinated responses, yi=0y_i = 0 for non-hallucinated ground truth) is modeled via binomial logistic regression estimated by maximum likelihood estimation (MLE):

    P(yi=1∣xi)=exp⁡(β0+β1xi)1+exp⁡(β0+β1xi)P(y_i = 1 \mid x_i) = \frac{\exp(\beta_0 + \beta_1 x_i)}{1 + \exp(\beta_0 + \beta_1 x_i)}

    Fitted over N=43,842N = 43{,}842 dialogue responses from the FaithDial and BEGIN benchmarks (log-likelihood: −27,397.6-27{,}397.6):

    • Intercept β0=0.368\beta_0 = 0.368 (standard error: 0.0380.038, p<0.01p < 0.01)
    • Narrativity coefficient β1=0.631\beta_1 = 0.631 (standard error: 0.0590.059, p<0.01p < 0.01)

    The positive, statistically significant coefficient indicates that higher narrativity is a significant predictor of a response containing hallucinatory content.

  5. Knowl 5 — Beta Regression Analysis of Dialogue Coherence and Narrativity

    empirical result

    The association between narrativity and dialogic semantic coherence is evaluated using beta regression. Coherence scores are computed via the DEAM metric (a RoBERTa-large model fine-tuned on conversation-level semantic perturbations to assess discourse coherence), producing continuous values bounded in (0,1)(0, 1) with epsilon smoothing applied to handle boundary extremes.

    The regression model is specified as:

    Beta(Coherencei)=α+β1⋅Narrativityi\text{Beta}(\text{Coherence}_i) = \alpha + \beta_1 \cdot \text{Narrativity}_i

    Across N=65,287N = 65{,}287 observations from the FaithDial, BEGIN, and HaluEval benchmarks (log-likelihood: 109,935.2109{,}935.2, pseudo-R2=0.004R^2 = 0.004):

    • Intercept α=0.433\alpha = 0.433 (standard error: 0.0180.018, p<0.01p < 0.01)
    • Narrativity coefficient β1=0.372\beta_1 = 0.372 (standard error: 0.0290.029, p<0.01p < 0.01)

    The positive, statistically significant coefficient (p<0.01p < 0.01) demonstrates that higher degrees of narrativity in dialogue outputs are systematically associated with higher conversational coherence.

  6. Knowl 6 — Benchmark Datasets and Response Aggregation for Dialogue Hallucination Evaluation

    experimental setup

    Empirical evaluation of narrativity and hallucination is conducted across three dialogue datasets:

    1. FaithDial: A benchmark of 21,445 information-seeking dialogue turns derived from Wizard of Wikipedia, labeled by crowdsourced annotators as either hallucinated or truthful (subdivided into entailment, uncooperative, and generic classes).
    2. BEGIN: An expert-curated dialogue set containing information-seeking queries alongside responses generated by humans and language models (GPT-2, DoHA, and CTRL), labeled as hallucinated, partially hallucinated, or truthful.
    3. HaluEval: A large-scale benchmark containing 10,000 dialogue turns comparing ChatGPT-generated plausible hallucinations directly against ground truth responses.

    To standardize labels across datasets, instances without hallucination labels are treated as 'truth', instances possessing solely hallucination labels are treated as 'hallucinated', and instances in FaithDial and BEGIN containing both truthful and hallucinated properties are categorized as 'partial'.

  7. Knowl 7 — Theoretical Framing of Confabulation via the Narrative Paradigm and Cognitive Narratology

    theoretical result

    Grounding large language model generation in the communication framework of the Narrative Paradigm and Cognitive Narratology establishes that narrative generation serves as a primary cognitive mechanism for discourse-level sense-making. In the Narrative Paradigm, communication effectiveness is determined by narrative rationality, which comprises narrative coherence (internal contextual consistency) and narrative fidelity (resonance with existing mental models). Cognitive narratology demonstrates that constructing narrative storyworlds and counterfactual simulations maintains the self-consistency and adaptability of internal world representations. Consequently, LLM confabulations—which synthesize plausible, narrative details when factual groundings are absent—mirror human communicative strategies for establishing discourse coherence, indicating that the wholesale elimination of confabulation risks degrading an LLM's capacity for coherent narrative sense-making.

  8. Knowl 8 — Limitations in Causal Attribution and User Experience Generalization

    limitation

    The empirical and theoretical defense of confabulation has two main limitations:

    1. Non-Causal Relationship: While empirical analyses demonstrate statistically significant correlations between increased narrativity, hallucination labels, and dialogue coherence, the findings do not establish that narrativity causally drives coherence.
    2. Unvalidated User Experience in Human-AI Interaction: The hypothesized benefits of narratively rich confabulations are derived from human-to-human communication paradigms and automated metrics (an ELECTRA-large narrativity classifier and the DEAM coherence evaluator). Direct behavioral evaluations with human participants are required to determine whether users actually experience improved utility or satisfaction when interacting with high-narrative confabulations compared to low-narrative factual outputs.

Coverage note — No substantial contributed material was omitted; all primary theoretical definitions, experimental setups, tabular data, statistical regression models, and stated limitations are captured.

References

  1. 1.Hussam Alkaissi and Samy I McFarlane. 2023. Artificial hallucinations in chatgpt: implications in scientific writing. Cureus, 15(2).
  2. 2.Ivan Anishchenko, Samuel J Pellock, Tamuka M Chidyausiku, Theresa A Ramelot, Sergey Ovchinnikov, Jingzhou Hao, Khushboo Bafna, Christoffer Norn, Alex Kang, Asim K Bera, et al. 2021. De novo protein design by deep network hallucination. Nature, 600(7889):547–552.
  3. 3.Maria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, and Andrew Piper. 2023. Where do people tell stories online? story detection across online communities. arXiv preprint arXiv:2311.09675.
  4. 4.Anjulyn M Ballard, Ashlee Davis, and Cynthia A Hoffner. 2021. The impact of health narratives on persuasion in african american women: A systematic review and meta-analysis. Health Communication, 36(5):560–571.
  5. 5.Brian Boyd. 2009. On the origin of stories: Evolution, cognition, and fiction. Harvard University Press.
  6. 6.Teva D Brender. 2023. Chatbot confabulations are not hallucinations—reply. JAMA Internal Medicine, 183(10):1177–1178.
  7. 7.Gillian Brown and George Yule. 1983. Discourse analysis. Cambridge university press.
  8. 8.Jerome Bruner. 1987. Life as narrative. Social research, pages 11–32.
  9. 9.Jerome Bruner. 1991. The narrative construction of reality. Critical inquiry, 18(1):1–21.
  10. 10.Meng Cao, Yue Dong, and Jackie Cheung. 2022. Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3340–3354, Dublin, Ireland. Association for Computational Linguistics.
  11. 11.Rita Charon. 2001. Narrative medicine: a model for empathy, reflection, profession, and trust. Jama, 286(15):1897–1902.
  12. 12.Meng Chen, Robert A Bell, and Laramie D Taylor. 2016. Narrator point of view and persuasion in health narratives: the role of protagonist–reader similarity, identification, and self-referencing. Journal of health communication, 21(8):908–918.
  13. 13.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555.
  14. 14.Heather Coats, Nadia Shive, Ardith Z Doorenbos, and Sarah J Schmiege. 2020. Integration of person-centered narratives into the electronic health record: study protocol. Nursing research, 69(6):483.
  15. 15.Shawn Curran, Sam Lansley, and Oliver Bethell. 2023. Hallucination is the last thing you need. arXiv preprint arXiv:2306.11520.
  16. 16.Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024. Large legal fictions: Profiling legal hallucinations in large language models. arXiv preprint arXiv:2401.01301.
  17. 17.Michael F Dahlstrom. 2010. The role of causality in information acceptance in narratives: An example from science communication. Communication Research, 37(6):857–875.
  18. 18.Hilary P Dannenberg. 2008. Coincidence and counterfactuality: Plotting time and space in narrative fiction. U of Nebraska Press.
  19. 19.Daniel C Dennett. 1988. Why everyone is a novelist. TLS-THE TIMES LITERARY SUPPLEMENT, 4459(1016).
  20. 20.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018. Wizard of wikipedia: Knowledge-powered conversational agents. arXiv preprint arXiv:1811.01241.
  21. 21.Lubomír Doležel. 1998. Heterocosmica. Johns Hopkins University Press.
  22. 22.Eamon Duede. 2022. Instruments, agents, and artificial intelligence: novel epistemic categories of reliability. Synthese, 200(6):491.
  23. 23.Eamon Duede. 2023. Deep learning opacity in scientific discovery. Philosophy of Science, 90(5):1089–1099.
  24. 24.Eamon Duede, William Dolan, André Bauer, Ian Foster, and Karim Lakhani. 2024. Oil & water? diffusion of ai within and across scientific fields. arXiv preprint arXiv:2405.15828.
  25. 25.Eamon Duede and Richard Jean So. Forthcoming. The humanistic case for AI optimism. Poetics Today.
  26. 26.Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar Zaiane, Mo Yu, Edoardo M. Ponti, and Siva Reddy. 2022a. FaithDial: A faithful benchmark for information-seeking dialogue. Transactions of the Association for Computational Linguistics, 10:1473–1490.
  27. 27.Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy. 2022b. On the origin of hallucinations in conversational models: Is it the datasets or the models? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5271–5285, Seattle, United States. Association for Computational Linguistics.
  28. 28.Nancy Easterlin. 2012. A biocultural approach to literary theory and interpretation. JHU Press.
  29. 29.Robin Emsley. 2023. Chatgpt: these are not hallucinations–they’re fabrications and falsifications. Schizophrenia, 9(1):52.
  30. 30.Hao Fei, Qian Liu, Meishan Zhang, Min Zhang, and Tat-Seng Chua. 2023. Scene graph as pivoting: Inference-time image-free unsupervised multimodal machine translation with visual scene hallucination. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5980–5994, Toronto, Canada. Association for Computational Linguistics.
  31. 31.Silvia Ferrari and Francisco Cribari-Neto. 2004. Beta regression for modelling rates and proportions. Journal of applied statistics, 31(7):799–815.
  32. 32.Walter R Fisher. 1984. Narration as a human communication paradigm: The case of public moral argument. Communications Monographs, 51(1):1–22.
  33. 33.Walter R Fisher. 2021. Human communication as narration: Toward a philosophy of reason, value, and action. Univ of South Carolina Press.
  34. 34.Arthur W Frank. 2013. The wounded storyteller: Body, illness & ethics. University of Chicago Press.
  35. 35.Lauren French, Maryanne Garry, and Elizabeth Loftus. 2009. False memories: A kind of confabulation in non-clinical. Confabulation: Views from neuroscience, psychiatry, psychology, and philosophy.
  36. 36.Richard Gerrig. 2018. Experiencing narrative worlds. Routledge.
  37. 37.Sarik Ghazarian, Nuan Wen, Aram Galstyan, and Nanyun Peng. 2022. DEAM: Dialogue coherence evaluation using AMR-based semantic manipulations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 771–785, Dublin, Ireland. Association for Computational Linguistics.
  38. 38.Google. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  39. 39.Herbert P Grice. 1975. Logic and conversation. In Speech acts, pages 41–58. Brill.
  40. 40.Rami Hatem, Brianna Simmons, and Joseph E Thornton. 2023. Chatbot confabulations are not hallucinations. JAMA Internal Medicine, 183(10):1177–1177.
  41. 41.David Herman. 2009. Basic elements of narrative. John Wiley & Sons.
  42. 42.David Herman. 2013. Storytelling and the Sciences of Mind. MIT press.
  43. 43.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  44. 44.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  45. 45.Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. arXiv preprint arXiv:2307.10169.
  46. 46.Adam Tauman Kalai and Santosh S Vempala. 2023. Calibrated language models must hallucinate. arXiv preprint arXiv:2311.14648.
  47. 47.Haoqiang Kang and Xiao-Yang Liu. 2023. Deficiency of large language models in finance: An empirical examination of hallucination. arXiv preprint arXiv:2311.15548.
  48. 48.Saul Kripke. 1963. Semantical considerations on modal logic. Acta Philosophica Fennica, 16:83–94.
  49. 49.Minhyeok Lee. 2023. A mathematical investigation of hallucination and creativity in gpt models. Mathematics, 11(10):2320.
  50. 50.David Lewis. 1979. Counterfactual dependence and time’s arrow. Noûs, pages 455–476.
  51. 51.Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. HaluEval: A large-scale hallucination evaluation benchmark for large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6449–6464, Singapore. Association for Computational Linguistics.
  52. 52.Zihao Li. 2023. The dark side of chatgpt: Legal and ethical challenges from stochastic parrots and hallucination. arXiv preprint arXiv:2304.14347.
  53. 53.Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023a. Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness? arXiv preprint arXiv:2312.03729.
  54. 54.Zhenzhen Liu, Chao Wan, Varsha Kishore, Jin Peng Zhou, Minmin Chen, and Kilian Q Weinberger. 2023b. Correction with backtracking reduces hallucination in summarization. arXiv preprint arXiv:2310.16176.
  55. 55.Junliang Luo, Tianyu Li, Di Wu, Michael Jenkin, Steve Liu, and Gregory Dudek. 2024. Hallucination detection and hallucination mitigation: An investigation. arXiv preprint arXiv:2401.08358.
  56. 56.Jean Matter Mandler. 2014. Stories, scripts, and scenes: Aspects of schema theory. Psychology Press.
  57. 57.Daniel McKee, Bing Shuai, Andrew Berneshawi, Manchen Wang, Davide Modolo, Svetlana Lazebnik, and Joseph Tighe. 2021. Multi-object tracking with hallucinated and unlabeled videos. arXiv preprint arXiv:2108.08836.
  58. 58.Hugo Mercier and Dan Sperber. 2017. The enigma of reason. Harvard University Press.
  59. 59.Anirban Mukherjee and Hannah Chang. 2023. The creative frontier of generative ai: Managing the novelty-usefulness tradeoff. arXiv preprint arXiv:2306.03601.
  60. 60.Mohamed Nejjar, Luca Zacharias, Fabian Stiehle, and Ingo Weber. 2023. Llms for science: Usage for code generation and data analysis. arXiv preprint arXiv:2311.16733.
  61. 61.OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  62. 62.Corinna Oschatz and Caroline Marker. 2020. Long-term persuasive effects in narrative communication research: A meta-analysis. Journal of Communication, 70(4):473–496.
  63. 63.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  64. 64.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2023. Med-HALT: Medical domain hallucination test for large language models. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 314–334, Singapore. Association for Computational Linguistics.
  65. 65.Andrew Piper and Sunyam Bagga. 2022. Toward a data-driven theory of narrativity. New Literary History, 54(1):879–901.
  66. 66.Andrew Piper, Richard Jean So, and David Bamman. 2021. Narrative theory for computational narrative understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 298–311.
  67. 67.Paul Riesthuis, Henry Otgaar, Glynis Bogaard, and Ivan Mangiulli. 2023. Factors affecting the forced confabulation effect: a meta-analysis of laboratory studies. Memory, 31(5):635–651.
  68. 68.Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. 2023. Mathematical discoveries from program search with large language models. Nature, pages 1–3.
  69. 69.FRANCISCO Rosário, ISABEL CORREIA, JORGE OLIVEIRA, and JOAO RAPOSO. 2018. Narrative medicine in diabetes—a randomized trial. Diabetes, 67(Supplement_1).
  70. 70.Marie-Laure Ryan. 1991. Possible worlds, artificial intelligence, and narrative theory. Indiana University Press.
  71. 71.Marie-Laure Ryan. 2001. Narrative as virtual reality. Immersion and Interactivity in Literature, pages 357–359.
  72. 72.Oliver Sacks. 1985. The man who mistook his wife for a hat and other clinical tales.
  73. 73.Roger C Schank and Robert P Abelson. 2013. Scripts, plans, goals, and understanding: An inquiry into human knowledge structures. Psychology Press.
  74. 74.Constanze Schreiner, Markus Appel, Maj-Britt Isberner, and Tobias Richter. 2018. Argument strength and the persuasiveness of stories. Discourse Processes, 55(4):371–386.
  75. 75.Anshul Shah, Aniket Roy, Ketul Shah, Shlok Mishra, David Jacobs, Anoop Cherian, and Rama Chellappa. 2023. Halp: Hallucinating latent positives for skeleton-based self-supervised learning of actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18846–18856.
  76. 76.Ritwik Sinha, Zhao Song, and Tianyi Zhou. 2023. A mathematical abstraction for balancing the trade-off between creativity and reality in large language models. arXiv preprint arXiv:2306.02295.
  77. 77.Andrew L Smith, Felix Greaves, and Trishan Panch. 2023. Hallucination or confabulation? neuroanatomy as metaphor in large language models. PLOS Digital Health, 2(11):e0000388.
  78. 78.Michael Stubbs. 1983. Discourse analysis: The sociolinguistic analysis of natural language. Basil Blackwell.
  79. 79.Wannapon Suraworachet, Jennifer Seon, and Mutlu Cukurova. 2024. Predicting challenge moments from students’ discourse: A comparison of gpt-4 to two traditional natural language processing approaches. arXiv preprint arXiv:2401.01692.
  80. 80.Charles Taylor. 1992. Sources of the Self. Harvard University Press.
  81. 81.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  82. 82.Mark Turner. 1996. The literary mind: The origins of thought and language. Oxford University Press.
  83. 83.Skyler Wang, Ned Cooper, Margaret Eby, and Eun Seo Jo. 2023. From human-centered to social-centered artificial intelligence: Assessing chatgpt’s impact through disruptive events. arXiv preprint arXiv:2306.00227.
  84. 84.Amanda Wiggins and Jessica L. Bunin. 2023. Confabulation. StatPearls Publishing, Treasure Island (FL).
  85. 85.Jing Wu, Jennifer Hobbs, and Naira Hovakimyan. 2023. Hallucination improves the performance of unsupervised visual representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 16132–16143.
  86. 86.Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817.
  87. 87.Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, and Li Yuan. 2023. Llm lies: Hallucinations are not bugs, but features as adversarial examples. arXiv preprint arXiv:2310.01469.
  88. 88.Xiao Zhan, Yifan Xu, and Stefan Sarkadi. 2023. Deceptive ai ecosystems: The case of chatgpt. In Proceedings of the 5th International Conference on Conversational User Interfaces, CUI ’23, New York, NY, USA. Association for Computing Machinery.
  89. 89.Chen Zhang. 2023. User-controlled knowledge fusion in large language models: Balancing creativity and hallucination. arXiv preprint arXiv:2307.16139.
  90. 90.Kyrie Zhixuan Zhou, Zachary Kilhoffer, Madelyn Rose Sanfilippo, Ted Underwood, Ece Gumusel, Mengyi Wei, Abhinav Choudhry, and Jinjun Xiong. 2024. " the teachers are confused as well": A multiple-stakeholder ethics discussion on large language models in computing education. arXiv e-prints, pages arXiv–2401.
  91. 91.Lisa Zunshine. 2006. Why we read fiction: Theory of mind and the novel. Ohio State University Press.

Citation

MLA
Sui, P., et al. “Confabulation: The Surprising Value of Large Language Model Hallucinations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 14274–84, https://doi.org/10.18653/v1/2024.acl-long.770.
APA
Sui, P., Duede, E., Wu, S., & So, R. (2024). Confabulation: The Surprising Value of Large Language Model Hallucinations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14274–14284. https://doi.org/10.18653/v1/2024.acl-long.770
Chicago
Sui, P., E. Duede, S. Wu, and R. So. 2024. “Confabulation: The Surprising Value of Large Language Model Hallucinations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14274–84. https://doi.org/10.18653/v1/2024.acl-long.770.
Harvard
Sui, P. et al. (2024) “Confabulation: The Surprising Value of Large Language Model Hallucinations”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14274–14284. Available at: https://doi.org/10.18653/v1/2024.acl-long.770.
Vancouver
1. Sui P, Duede E, Wu S, So R (2024) Confabulation: The Surprising Value of Large Language Model Hallucinations. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14274–14284

BibTeX

@inproceedings{sui-etal-2024-confabulation,
    title = "Confabulation: The Surprising Value of Large Language Model Hallucinations",
    author = "Sui, Peiqi  and
      Duede, Eamon  and
      Wu, Sophie  and
      So, Richard",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.770/",
    doi = "10.18653/v1/2024.acl-long.770",
    pages = "14274--14284"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/