Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?
Kevin LiuStephen CasperDylan Hadfield-MenellJacob Andreas
Explains why internal probes often outperform direct language model outputs by categorizing query–probe disagreements into confabulation, deception, and heterogeneity, showing that superior probe accuracy stems primarily from better uncertainty calibration rather than intentional deception.
Artificial intelligence language models frequently generate factual errors, which creates risks for organizations relying on them as automated knowledge sources. Prior studies revealed that examining a model's internal representations via linear classifiers—known as probes—often yields more accurate factual judgments than asking the model directly through queries. This disparity has led to claims that language models exhibit deception by internally knowing the truth while outputting false statements.
The article investigates why model outputs disagree with their internal representations of truthfulness. It evaluates whether this mismatch stems from deliberate deception-like dynamics or alternative mechanisms, including differences in uncertainty calibration and performance across distinct subsets of data.
The researchers conducted an empirical evaluation using two autoregressive language models, GPT2-XL and GPT-J, across three question-answering benchmarks: BoolQ, SciQ, and CREAK. They compared direct query outputs against linear probes trained on the models' final internal hidden states. The study established a taxonomy classifying disagreements into three main types: model confabulation (where queries answer confidently despite probe uncertainty), deception (where both query and probe are confident but reach opposite conclusions), and heterogeneity (where probes and queries succeed on different subsets of inputs).
The analysis produced three primary findings. First, apparent deception is rare; high-confidence contradictions between probes and queries occurred infrequently across the benchmarks, showing up meaningfully only on the CREAK dataset. Second, the superior performance of probes is largely driven by better statistical calibration on uncertain answers rather than a greater volume of confident, correct facts. Third, queries and probes exhibit complementary strengths on different subsets of data: by combining both approaches into a weighted ensemble, the researchers increased overall accuracy beyond either individual method in four out of six experimental settings, such as raising GPT-J accuracy on SciQ from 84.1% via query and 88.6% via probe to 90.5% via the ensemble.
These findings indicate that internal-versus-external discrepancies stem from differing computational prediction pathways rather than an underlying intent to mislead. For organizational leadership, this clarifies that anthropomorphic framing around model lying is largely inaccurate. However, it also highlights that internal probing cannot fully resolve language model unreliability, as models still make significant factual errors across both pathways.
Organizations developing or deploying factual language systems should consider ensembling output probabilities with internal probe signals to achieve incremental accuracy gains and better uncertainty estimates. When high accuracy is critical, decision-makers should fine-tune underlying models directly on domain data rather than relying solely on probing pretrained models, as fine-tuning consistently outperformed zero-shot queries and probes on datasets like SciQ and CREAK.
Confidence in these findings is supported by consistent error distributions across different model sizes and probe regularizations. Nonetheless, decision-makers should exercise caution: the study was limited to binary question answering, evaluated a single prompt template per task, and tested models that continue to generate substantial factual errors, making them unsuitable for fully autonomous deployment in mission-critical settings without human oversight.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). Its Contrast-Consistent Search method establishes how hidden-state probes can recover factual knowledge that ordinary language-model answers fail to express, the central premise this study tests.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its analysis of models’ self-evaluation and calibration provides essential context for this study’s finding that probe advantages often reflect better handling of uncertainty.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Its TruthfulQA benchmark frames the problem of fluent, confident falsehoods that motivates studying gaps between internal truth signals and generated answers.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Its methods for eliciting factual knowledge from language models establish the earlier problem of separating what models know from what their prompts reveal.
- Paper: Inside-Out: Hidden Factual Knowledge in LLMs, Jonathan Herzig et al. (2025). It extends the internal-versus-external knowledge comparison with a formal probe-based measurement, finding that hidden factual knowledge exceeds what standard generation and observable probabilities reveal.
- Paper: TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space, Shaolei Zhang et al. (2024). It turns the gap between internal truthfulness and generated output into an intervention, editing internal representations to improve factual answers.
- Paper: How Language Model Hallucinations Can Snowball, Muru Zhang et al. (2024). It carries the knowledge-expression mismatch into multi-turn interactions, showing how an initial error can lead a model to produce further false claims despite recognizing them separately.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). It extends internal probing into selective control, steering whether a model uses its stored knowledge or conflicting retrieved context.
