Built independently by an author, for readers. Read the story and support ChapterPal

keyword

high-confidence answers

High-confidence answers are outputs or predictions generated by a computational model, such as an artificial intelligence or question-answering system, for which the system assigns a high probability, score, or estimated certainty of correctness. In machine learning and computational evaluation, these responses are characterized by strong internal representation signals, high conditional probability scores, or explicit certainty metrics indicating that the system treats the output as definitive rather than speculative. They are commonly distinguished from uncertain answers to evaluate calibration, which measures how reliably a system confidence level matches its empirical accuracy. While well-calibrated systems demonstrate a strong alignment between high confidence and factual truth, models can also generate high-confidence errors when internal probability estimates fail to correspond with real-world correctness.

1 item

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, Jacob Andreas

OrganizationsMassachusetts Institute of Technology

Why you should read this

Explains why internal probes often outperform direct language model outputs by categorizing query–probe disagreements into confabulation, deception, and heterogeneity, showing that superior probe accuracy stems primarily from better uncertainty calibration rather than intentional deception.

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these two procedures sometimes disagree, and that probes tend to be more accurate than LM outputs. This has led some researchers to conclude that LMs “lie” or otherwise encode non-cooperative communicative intents. Is this an accurate description of today’s LMs, or can query–probe disagreement arise in other ways? We identify three different classes of disagreement, which we term confabulation, deception, and heterogeneity. In many cases, the superiority of probes is simply attributable to better calibration on uncertain answers rather than a greater fraction of correct, high-confidence answers. In some cases, queries and probes perform better on different subsets of inputs, and accuracy can further be improved by ensembling the two.¹

Added

2026-10-03