On Large Language Models' Hallucination with Regard to Known Facts

Che JiangBiqing QiXiangyu HongDayuan FuYang ChengFandong MengMo YuBowen ZhouJie Zhou

article2024NAACL78 citations

Reveals distinct token probability trajectories across model layers during failed knowledge recall, enabling an 88% accurate classifier for detecting when large language models hallucinate facts they actually know.

Listen

Large language models frequently generate confident yet incorrect statements, known as factual hallucinations. While some errors stem from missing information, models often fail even when they have stored the correct facts in memory. This inconsistency undermines the reliability of artificial intelligence systems in high-stakes operational environments, where organizations must depend on truthful and verifiable outputs.

The article investigates the internal mechanisms that cause language models to fail at recalling facts they already know, and demonstrates an automated method to detect these hallucinations during generation.

To evaluate this behavior, the authors constructed a dataset of over 30,000 factual queries across multiple relationship categories. By asking for the same factual fact using different phrasing, they isolated instances where the model answered correctly under one prompt but hallucinated under another. Using internal mapping and tracking techniques on open-source foundation models—primarily Llama-2-7B-chat alongside OPT and Pythia models—the researchers observed layer-by-layer internal state changes across the network's depth during text generation.

The analysis revealed four key findings regarding how hallucinations unfold internally:

  • Known-fact hallucination stems from failed internal retrieval rather than missing knowledge: when the model produces an incorrect output, the correct answer appears as the top candidate in intermediate layers only about 31% of the time, compared to nearly 78% during correct generations.
  • Correct outputs display distinct layer dynamics: during accurate generation, the correct answer's probability surges steeply in the middle-to-late processing layers (around layer 20 in 32-layer models), whereas incorrect outputs speculate early in shallow layers without a definitive extraction point.
  • Feed-forward neural network components contribute more to errors than attention components: these feed-forward modules actively suppress the correct answer and promote erroneous choices in the final decoding stages.
  • Tracking internal layer dynamics enables accurate automated hallucination detection: training a standard support vector machine classifier solely on internal token probability trajectories distinguished correct answers from hallucinations with up to an 88% success rate across multiple model architectures.

These findings indicate that hallucination on known knowledge is a systematic failure during internal semantic parsing and retrieval rather than random noise or entity obscurity. Because erroneous responses follow distinct internal trajectories, organizations deploying open-source models can detect hallucinations in real time by inspecting hidden states without needing pre-existing answer keys or external fact-checking databases, substantially lowering verification costs and operational risks.

Organizations should consider implementing white-box monitoring probes on internal generation trajectories for automated quality assurance in critical workflows. When hallucinations occur, prompt rephrasing can often trigger successful recall without requiring expensive model retraining. However, further validation is necessary before applying this approach to complex multi-step reasoning, unstructured generation tasks, or proprietary models where internal hidden states cannot be inspected.

Cover for On Large Language Models' Hallucination with Regard to Known Facts

Abstract

Large language models are successful in answering factoid questions but are also prone to hallucination. We investigate the phenomenon of LLMs possessing correct answer knowledge yet still hallucinating from the perspective of inference dynamics, an area not previously covered in studies on hallucinations. We are able to conduct this analysis via two key ideas. First, we identify the factual questions that query the same triplet knowledge but result in different answers. The difference between the model behaviors on the correct and incorrect outputs hence suggests the patterns when hallucinations happen. Second, to measure the pattern, we utilize mappings from the residual streams to vocabulary space. We reveal the different dynamics of the output token probabilities along the depths of layers between the correct and hallucinated cases. In hallucinated cases, the output token’s information rarely demonstrates abrupt increases and consistent superiority in the later stages of the model. Leveraging the dynamic curve as a feature, we build a classifier capable of accurately detecting hallucinatory predictions with an 88% success rate. Our study shed light on understanding the reasons for LLMs’ hallucinations on their known facts, and more importantly, on accurately predicting when they are hallucinating.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Experimental Setup
  • 3.1 Dataset
  • 3.2 Model
  • 3.3 Observation methods
  • 4 Results
  • 4.1 Accuracy Statistics
  • 4.2 Lens Observation
  • 4.3 Module contributions
  • 4.4 Logit evolution pattern
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgements
  • References
  • A Dataset Statistics

Knowls

  1. Knowl 1 — Correct and hallucinated answers follow different layerwise inference dynamics

    empirical result

    For factual triplet completion, the probability assigned to the first token of a correct answer typically rises sharply in the middle-to-late layers, after being near zero in early layers. Hallucinated output tokens generally lack a comparable abrupt increase and can be represented before the query has been fully parsed. With the Tuned Lens, correct-answer information tends to emerge around layer 20 in the 32-layer Llama2-7B-chat model, whereas information associated with erroneous outputs is often apparent earlier. With the Logit Lens, the erroneous token tends to become decodable only in the last two layers, while the correctly recalled token becomes decodable earlier and remains so. The authors interpret this contrast as suggesting that successful recall involves an intermediate extraction point after which the answer is retained for decoding, while hallucinated outputs can bypass that process.

  2. Knowl 2 — Failed recall rarely promotes the correct answer to the top logit rank

    empirical result

    The study compared three first-token cases for the same factual triplet: the correct answer token under a prompt answered correctly (Suc.), that same correct token under a different prompt answered incorrectly (Fail.), and the incorrect output token under that latter prompt (Hal.). Across relation types, the average frequency with which each token reached the top ranks of Logit Lens vocabulary predictions was:

    • Top 1: Suc. 77.57%, Fail. 31.28%, Hal. 68.04%.
    • Top 5: Suc. 93.21%, Fail. 56.71%, Hal. 92.70%.

    The correct token under failed-recall prompts thus reached top rank much less often than under successful prompts. The authors conclude that in most observed hallucinations the target knowledge was not extracted into a strong intermediate prediction; in some cases, however, the correct token briefly became competitive and was then suppressed.

  3. Knowl 3 — MLP activity is associated with suppressing correct answers during failed recall

    empirical result

    The authors compared attention and MLP contributions to the last input token’s residual stream and ablated module outputs at selected positions. In successful recalls, both attention and MLP modules contributed to answer-related information, particularly around layer 20. In failed recalls, some answer information could still be extracted around that layer, but MLP contributions had a stronger inhibitory effect on the correct token toward the final layers and contributed to decoding erroneous outputs. Ablations over more than 200 pairs of successful and failed cases also showed that, in the early half of the model, subject-position semantic parsing had less impact on the final output in failed recalls than in successful ones; later output processing was concentrated at the last input-token position. These observations support, but do not by themselves establish as a complete causal account, the proposal that weak or biased parsing can impair later extraction and that MLP activity can help erroneous information prevail.

  4. Knowl 4 — Layerwise output-token curves support hallucination classification

    empirical result

    The authors trained a linear support vector classifier using the output-token probability curve across layers, extracted with the Logit Lens, the Tuned Lens, or both curves concatenated. They used scikit-learn’s default SVC hyperparameters, shuffled the curves, and assigned 20% to testing and 80% to training. The reported classification accuracies were:

    • Llama-7B-chat: Logit 0.839, Tuned 0.854, Both 0.879.
    • Llama-13B-chat: Logit 0.849, Tuned 0.840, Both 0.878.
    • OPT-6.7B: Logit 0.856, Tuned 0.858, Both 0.865.
    • Pythia-6.9B: Logit 0.824, Tuned 0.764, Both 0.822.

    Thus, output-token dynamics classified hallucinated versus non-hallucinated predictions at above 80% accuracy for all four models; concatenating the two lens curves gave the best accuracy for the three models where it improved on either single curve, reaching 0.879 for Llama-7B-chat. The classifier does not require the correct answer as an input, but it does require access to the model’s internal states.

  5. Knowl 5 — Operational definition of a known-fact hallucination

    definition

    A known-fact hallucination occurs when a language model receives different prompts that query the same factual triplet (s,r,o)(s,r,o)—where ss is a subject, rr a relation, and oo the correct object—and produces both a correct answer and an incorrect response across those prompts. Incorrect responses include wrong entities, irrelevant information, and uncertain or uninformative answers. In this setup, producing the correct answer for at least one prompt is evidence that the model can recall the fact under some query formulation; an incorrect answer to another formulation is the hallucination being studied.

  6. Knowl 6 — Triplet-query dataset pairs prompt variation with model response variation

    model/method

    The authors adapted COUNTERFACT triplets by creating more than 30,000 statement or question prompts that end immediately before the object to be predicted. They devised multiple query formulations for the same subject–relation pair and manually clarified ambiguous templates so that the expected answer belonged to the intended semantic category. This enabled comparisons of different prompts for the same fact, including cases where one prompt elicited the correct object and another elicited an incorrect response. The dataset covered relation types such as capitals, occupations, native languages, and places of birth; the model’s first generated token was used to track the answer’s development.

  7. Knowl 7 — Residual-stream lens analysis tracks three answer-token roles

    model/method

    For a factual triplet (s,r,o)(s,r,o), the analysis follows the hidden state at the final input-token position across model layers, since that position predicts the next token. The Logit Lens maps each layer’s residual state directly into vocabulary space using the model’s unembedding; the Tuned Lens uses layer-specific learned transformations before mapping to vocabulary space. For a pair of prompts querying the same triplet, the authors track: (1) Suc., the correct answer’s first token under the prompt answered correctly; (2) Fail., that same correct token under the prompt answered incorrectly; and (3) Hal., the incorrect response’s first token under the latter prompt. These curves distinguish whether the correct answer is represented, whether it is suppressed, and when the erroneous output becomes decodable.

  8. Knowl 8 — Model and decoding protocol for the primary analysis

    experimental setup

    The primary inference-dynamics analysis used Llama2-7B-chat, a 32-layer Transformer with hidden dimension 4,096 and vocabulary size 32,000. At layer ll, the hidden state xilx_i^{l} for token position ii is updated as xil=xil−1+ail+milx_i^{l}=x_i^{l-1}+a_i^{l}+m_i^{l}, where aila_i^{l} and milm_i^{l} are that layer’s attention and MLP outputs. Decoding was fixed to greedy selection to avoid variation from decoding strategies. A response counted as correct when its first ten generated tokens contained the ground-truth answer; correct samples containing negation terms or words resembling multiple-choice answers were filtered out.

  9. Knowl 9 — Subject popularity showed little relationship with error categories

    empirical result

    The authors manually categorized errors in roughly 200 sampled cases as incorrect entities, uncertain responses, or irrelevant information, and grouped subjects by Wikipedia-page browsing counts for the preceding year. The category counts by popularity group were: fewer than 10410^4 views, 28 incorrect, 12 uncertain, and 11 irrelevant; 10410^4–10510^5, 26, 8, and 16; 10510^5–10610^6, 27, 9, and 16; and more than 10610^6, 28, 9, and 14. They found no significant association between popularity and these error categories. In a broader analysis of facts with four query variants, lower access frequency was weakly associated with more extraction errors; the authors therefore do not claim that popularity has no relationship to recall performance.

  10. Knowl 10 — Scope and interpretation are limited

    limitation

    The investigation focuses on factual knowledge expressed as subject–relation–object triplets, so its findings may not transfer to other knowledge structures or domains. The detailed inference-dynamics analysis centers on Llama2-7B-chat; classification experiments on Llama2-13B-chat, OPT-6.7B, and Pythia-6.9B provide additional evidence for the detector but do not establish that the proposed internal mechanism generalizes across models. The authors also characterize their account of how internal-state changes produce outputs as preliminary, leaving the detailed causal influence of those changes unresolved.

Coverage note — The extensive relation-by-relation prompt-template inventory and per-relation visualizations are omitted because they operationalize the dataset and illustrate the reported aggregate findings rather than forming separate standalone contributions.

References

  1. 1.Guillaume Alain and Yoshua Bengio. 2016. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644.
  2. 2.Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. 2022. A review on language models as knowledge bases. arXiv preprint arXiv:2204.06031.
  3. 3.Amos Azaria and Tom Mitchell. 2023. The internal state of an llm knows when its lying. arXiv preprint arXiv:2304.13734.
  4. 4.Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor V. Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens. ArXiv, abs/2303.08112.
  5. 5.Stella Biderman, Hailey Schoelkopf, Quentin G. Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023a. Pythia: A suite for analyzing large language models across training and scaling. ArXiv, abs/2304.01373.
  6. 6.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023b. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397–2430. PMLR.
  7. 7.Shiyu Chang, Yang Zhang, Mo Yu, and Tommi Jaakkola. 2020. Invariant rationalization. In International Conference on Machine Learning, pages 1448–1458. PMLR.
  8. 8.Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. Dola: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883.
  9. 9.Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. 2023. Crawling the internal knowledge-base of language models. arXiv preprint arXiv:2301.12810.
  10. 10.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023. Analyzing transformers in embedding space. In Annual Meeting of the Association for Computational Linguistics.
  11. 11.Javier Ferrando, Gerard I Gállego, Ioannis Tsiamas, and Marta R Costa-jussà. 2023. Explaining how transformers use context to build predictions. arXiv preprint arXiv:2305.12535.
  12. 12.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767.
  13. 13.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45.
  14. 14.Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt. 2023. Overthinking the truth: Understanding how language models process false demonstrations. ArXiv, abs/2307.09476.
  15. 15.Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023. How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. ArXiv, abs/2305.00586.
  16. 16.Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2022. Understanding transformer memorization recall through idioms. arXiv preprint arXiv:2210.03588.
  17. 17.Evan Hernandez, Arnab Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2023. Linearity of relation decoding in transformer language models. ArXiv, abs/2308.09124.
  18. 18.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  19. 19.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  20. 20.Nikhil Kandpal, H. Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning.
  21. 21.Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023a. Trustworthy ai: From principles to practices. ACM Comput. Surv., 55(9).
  22. 22.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023b. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341.
  23. 23.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372.
  24. 24.Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229.
  25. 25.Neel Nanda, Andrew Lee, and Martin Wattenberg. 2023. Emergent linear representations in world models of self-supervised sequence models. ArXiv, abs/2309.00941.
  26. 26.nostalgebraist. 2020. interpreting gpt: the logit lens. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens.
  27. 27.Adam Pearce, Asma Ghandeharioun, Nada Hussein, Nithum Thain, Martin Wattenberg, and Lucas Dixon. 2023. Do machine learning models memorize or generalize? https://pair.withgoogle.com/explorables/grokking/.
  28. 28.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? arXiv preprint arXiv:1909.01066.
  29. 29.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. In Conference on Empirical Methods in Natural Language Processing.
  30. 30.Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023. The curious case of hallucinatory (un) answerability: Finding truths in the hidden states of over-confident large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3607–3625.
  31. 31.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288.
  32. 32.Miles Turpin, Julian Michael, Ethan Perez, and Sam Bowman. 2023. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. ArXiv, abs/2305.04388.
  33. 33.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don’t know? In Findings of the Association for Computational Linguistics: ACL 2023, pages 8653–8665, Toronto, Canada. Association for Computational Linguistics.
  34. 34.Paul Youssef, Osman Alperen Koracs, Meijie Li, Jorg Schlotterer, and Christin Seifert. 2023. Give me the facts! a survey on factual knowledge probing in pre-trained language models. In Conference on Empirical Methods in Natural Language Processing.
  35. 35.Mo Yu, Shiyu Chang, Yang Zhang, and Tommi S Jaakkola. 2019. Rethinking cooperative rationalization: Introspective extraction and complement control. In Conference on Empirical Methods in Natural Language Processing and International Joint Conference on Computational Linguistics. Association for Computational Linguistics.
  36. 36.Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi. 2024. Attention satisfies: A constraint-satisfaction lens on factual errors of language models. In The Twelfth International Conference on Learning Representations.
  37. 37.Fred Zhang and Neel Nanda. 2023. Towards best practices of activation patching in language models: Metrics and methods. arXiv preprint arXiv:2309.16042.
  38. 38.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  39. 39.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023. Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv preprint arXiv:2309.01219.
  40. 40.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [mask]: Learning vs. learning to recall. In North American Chapter of the Association for Computational Linguistics.

Citation

MLA
Jiang, C., et al. “On Large Language Models’ Hallucination with Regard to Known Facts”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 1041–53, https://doi.org/10.18653/v1/2024.naacl-long.60.
APA
Jiang, C., Qi, B., Hong, X., Fu, D., Cheng, Y., Meng, F., Yu, M., Zhou, B., & Zhou, J. (2024). On Large Language Models’ Hallucination with Regard to Known Facts. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1041–1053. https://doi.org/10.18653/v1/2024.naacl-long.60
Chicago
Jiang, C., B. Qi, X. Hong, et al. 2024. “On Large Language Models’ Hallucination with Regard to Known Facts”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 1041–53. https://doi.org/10.18653/v1/2024.naacl-long.60.
Harvard
Jiang, C. et al. (2024) “On Large Language Models’ Hallucination with Regard to Known Facts”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1041–1053. Available at: https://doi.org/10.18653/v1/2024.naacl-long.60.
Vancouver
1. Jiang C, Qi B, Hong X, Fu D, Cheng Y, Meng F, Yu M, Zhou B, Zhou J (2024) On Large Language Models’ Hallucination with Regard to Known Facts. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 1041–1053

BibTeX

@inproceedings{jiang-etal-2024-large,
    title = "On Large Language Models' Hallucination with Regard to Known Facts",
    author = "Jiang, Che  and
      Qi, Biqing  and
      Hong, Xiangyu  and
      Fu, Dayuan  and
      Cheng, Yang  and
      Meng, Fandong  and
      Yu, Mo  and
      Zhou, Bowen  and
      Zhou, Jie",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.60/",
    doi = "10.18653/v1/2024.naacl-long.60",
    pages = "1041--1053"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/