LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Hadas OrgadMichael TokerZorik GekhmanRoi ReichartIdan SzpektorHadas KotekYonatan Belinkov
Demonstrates that large language models often internally represent the correct answer while generating hallucinations, and identifies token-level truthfulness signals that predict specific error types despite failing to transfer across datasets.
Large language models frequently generate inaccurate or biased information, widely known as hallucinations. While most conventional methods evaluate these errors based on external user perception or simple output confidence, this human-centric perspective fails to capture how models internally encode truthfulness. Understanding how models internally represent errors is critical for building trustworthy systems and implementing precise safety mechanisms. The article evaluates how truthfulness is represented across the internal layers and tokens of language models, demonstrating how these internal signals can be used to detect errors, classify error types, and select correct answers.
To conduct this evaluation, the researchers tested four open-source language models across ten diverse datasets covering factual retrieval, commonsense reasoning, natural language inference, sentiment analysis, and mathematics. They allowed models to generate unrestricted, long-form answers to mirror practical use. The analysis involved training lightweight linear classifiers, called probing classifiers, on intermediate model activations across different layers and token positions. The study compared these probing methods against common baseline techniques, such as output probability aggregation and direct self-evaluation prompting, and evaluated their performance across tasks, error distributions, and multiple generated samples.
Key findings reveal that internal truthfulness signals are highly concentrated in the exact answer tokens rather than the prompt or final generated tokens. Exploiting activations at these exact answer tokens substantially improved error detection accuracy, outperforming all logit- and probability-based baselines across all datasets. Second, the article demonstrates that truthfulness encoding is skill-specific rather than universal. Probes trained on one dataset failed to generalize to fundamentally different tasks beyond what simple output probabilities could already predict, refuting prior assumptions of a single, universal internal truth mechanism. Third, internal representations successfully predicted fine-grained error types, such as whether a model would consistently repeat the same error or generate widespread guesses. Finally, the analysis revealed a severe discrepancy between internal knowledge and external output: even when models consistently generated incorrect responses, their internal representations often identified the correct answer among resampled candidates, allowing probe-based selection to improve task accuracy by 30 to 40 percentage points on difficult error categories.
These findings imply that relying on a single, general-purpose truth detector introduces operational risks, as models handle different reasoning skills through distinct internal mechanisms. However, organizations can build highly reliable, task-specific error detectors for high-stakes domains such as medicine and law by targeting exact answer tokens. The results also show that models often possess the required factual knowledge internally but fail to generate it due to decoding mechanisms that prioritize text likelihood over truthfulness, indicating significant potential to recover accurate answers without retraining entire models.
Organizations deploying language models should avoid universal truth-detecting filters and instead implement task-specific probing architectures focused on exact answer tokens within specialized workflows. For complex domains, engineering teams can combine internal probing with answer resampling or targeted retrieval interventions based on the predicted error type. Further development should focus on building lightweight extraction modules for real-time answer token identification and testing internal decoding interventions on open-ended generation tasks.
The conclusions are limited to open-source models, as the methodology requires direct access to intermediate neural network activations, making it inapplicable to closed, black-box systems. Additionally, the benchmarks focused on structured tasks with verifiable reference labels rather than open-ended creative text. Confidence in these findings is high across standard question-answering and reasoning benchmarks, though caution is warranted when attempting to extrapolate probe behavior to unseen task domains without task-specific validation.
- Paper: Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?, Kevin Liu et al. (2023). Read this earlier analysis of probe–output disagreements to understand the competing explanations and calibration issues that motivate examining where truthfulness is encoded internally.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). Its method for extracting latent factual knowledge from hidden activations provides a key precedent for the source’s use of probes to recover information that generation fails to express.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). This earlier study establishes how models can estimate whether their answers are correct, preparing readers for the source’s finer-grained investigation of internal truthfulness signals.
- Paper: Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators, Matéo Mahaut et al. (2024). Its comparison of internal probes with confidence estimators clarifies the baseline approaches the source tests against and the value of inspecting hidden representations.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). This earlier work shows how probes can locate internal truth-related representations, a prerequisite for understanding the source’s layer- and token-level probing analysis.
- Paper: On Large Language Models' Hallucination with Regard to Known Facts, Che Jiang et al. (2024). Its account of known-fact hallucinations as failures of retrieval provides context for the source’s finding that correct answers can remain internally available despite erroneous outputs.
- Paper: The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models, Aviv Slobodkin et al. (2023). Its hidden-state analysis of answerability introduces the use of internal signals to detect when generation is unreliable, which the source extends to truthfulness and error types.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). This benchmark establishes the TruthfulQA evaluation context used in earlier work on internal truthfulness and helps frame the factual-error detection problem.
- Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits, Amirhosein Ghasemabadi et al. (2025). Building on internal signals for correctness detection, Gnosis extends failure prediction to compressed traces of hidden states and attention across broader reasoning tasks.
