Built independently by an author, for readers. Read the story and support ChapterPal

keyword

cognitive dissonance

In artificial intelligence and natural language processing, cognitive dissonance refers to the phenomenon where a language model generates textual outputs that contradict or misalign with the factual truthfulness captured within its own internal representations. This discrepancy occurs when probing the internal hidden layers of a model indicates an accurate representation of a statement or truth value, while the model probability queries or generated texts produce an inaccurate, confabulated, or inconsistent response. Such disagreements highlight a divergence between external generation behaviors and internal latent knowledge, often driven by differing calibration under uncertainty, varying performance across input distributions, or misalignment between surface text generation objectives and internal representation spaces.

1 item

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, Jacob Andreas

OrganizationsMassachusetts Institute of Technology

Why you should read this

Explains why internal probes often outperform direct language model outputs by categorizing query–probe disagreements into confabulation, deception, and heterogeneity, showing that superior probe accuracy stems primarily from better uncertainty calibration rather than intentional deception.

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these two procedures sometimes disagree, and that probes tend to be more accurate than LM outputs. This has led some researchers to conclude that LMs “lie” or otherwise encode non-cooperative communicative intents. Is this an accurate description of today’s LMs, or can query–probe disagreement arise in other ways? We identify three different classes of disagreement, which we term confabulation, deception, and heterogeneity. In many cases, the superiority of probes is simply attributable to better calibration on uncertain answers rather than a greater fraction of correct, high-confidence answers. In some cases, queries and probes perform better on different subsets of inputs, and accuracy can further be improved by ensembling the two.¹

Added

2026-10-03