On Large Language Models' Hallucination with Regard to Known Facts
Che JiangBiqing QiXiangyu HongDayuan FuYang ChengFandong MengMo YuBowen ZhouJie Zhou
Reveals distinct token probability trajectories across model layers during failed knowledge recall, enabling an 88% accurate classifier for detecting when large language models hallucinate facts they actually know.
Large language models frequently generate confident yet incorrect statements, known as factual hallucinations. While some errors stem from missing information, models often fail even when they have stored the correct facts in memory. This inconsistency undermines the reliability of artificial intelligence systems in high-stakes operational environments, where organizations must depend on truthful and verifiable outputs.
The article investigates the internal mechanisms that cause language models to fail at recalling facts they already know, and demonstrates an automated method to detect these hallucinations during generation.
To evaluate this behavior, the authors constructed a dataset of over 30,000 factual queries across multiple relationship categories. By asking for the same factual fact using different phrasing, they isolated instances where the model answered correctly under one prompt but hallucinated under another. Using internal mapping and tracking techniques on open-source foundation models—primarily Llama-2-7B-chat alongside OPT and Pythia models—the researchers observed layer-by-layer internal state changes across the network's depth during text generation.
The analysis revealed four key findings regarding how hallucinations unfold internally:
- Known-fact hallucination stems from failed internal retrieval rather than missing knowledge: when the model produces an incorrect output, the correct answer appears as the top candidate in intermediate layers only about 31% of the time, compared to nearly 78% during correct generations.
- Correct outputs display distinct layer dynamics: during accurate generation, the correct answer's probability surges steeply in the middle-to-late processing layers (around layer 20 in 32-layer models), whereas incorrect outputs speculate early in shallow layers without a definitive extraction point.
- Feed-forward neural network components contribute more to errors than attention components: these feed-forward modules actively suppress the correct answer and promote erroneous choices in the final decoding stages.
- Tracking internal layer dynamics enables accurate automated hallucination detection: training a standard support vector machine classifier solely on internal token probability trajectories distinguished correct answers from hallucinations with up to an 88% success rate across multiple model architectures.
These findings indicate that hallucination on known knowledge is a systematic failure during internal semantic parsing and retrieval rather than random noise or entity obscurity. Because erroneous responses follow distinct internal trajectories, organizations deploying open-source models can detect hallucinations in real time by inspecting hidden states without needing pre-existing answer keys or external fact-checking databases, substantially lowering verification costs and operational risks.
Organizations should consider implementing white-box monitoring probes on internal generation trajectories for automated quality assurance in critical workflows. When hallucinations occur, prompt rephrasing can often trigger successful recall without requiring expensive model retraining. However, further validation is necessary before applying this approach to complex multi-step reasoning, unstructured generation tasks, or proprietary models where internal hidden states cannot be inspected.
- Paper: Dissecting Recall of Factual Associations in Auto-Regressive Language Models, Mor Geva et al. (2023). Its account of how attention and feed-forward layers retrieve factual associations gives a mechanistic foundation for tracing where correct answers emerge—or fail to emerge—during generation.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its tests of models’ confidence in their own answers establish why internal signals can reveal knowledge that an incorrect output conceals.
- Paper: Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?, Kevin Liu et al. (2023). Its distinction between probe-detectable truth and erroneous generated answers prepares readers to interpret the source’s analysis of internal knowledge and output failures.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). Building on internal traces as signals of correctness, Gnosis turns them into a compact self-checking mechanism that predicts failures across tasks.
- Paper: Inside-Out: Hidden Factual Knowledge in LLMs, Jonathan Herzig et al. (2025). It extends the distinction between correct internal knowledge and faulty output by systematically measuring factual knowledge hidden in intermediate computations.
