Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
Matéo MahautLaura AinaPaula CzarnowskaMomchil HardalovThomas MüllerLluís Màrquez
Evaluates five major families of factual confidence estimators across multiple large language models and tasks, revealing that internal hidden-state probes achieve superior reliability while exposing how easily model confidence fluctuates across semantically equivalent prompts.
Large Language Models often generate inaccurate statements or assert false claims with high confidence, leading to hallucinations, misinformation, and degraded user trust. Accurately estimating how confident a model is in a given fact is vital for deciding when a system should answer or abstain, yet existing confidence estimators have lacked systematic comparison across tasks and model families.
The article systematically evaluates the reliability and robustness of five factual confidence estimation techniques across two complementary settings: verifying whether a statement is true and estimating whether a model knows the answer to a question. The authors developed a standardized experimental framework testing eight open-weight models, ranging from 7 billion to 46.7 billion parameters, across multiple benchmark datasets, semantic paraphrases, and multilingual translations.
The investigation produced several key findings. First, supervised probes trained directly on the model's internal hidden layers significantly outperformed all black-box and prompting-based estimators, exceeding sequence-probability baselines by an average of 0.3 in the area under the precision-recall curve. Second, while trained probes generalized well across different datasets and translated languages, non-trained methods performed poorly, frequently hovering near chance levels in question-answering settings. Third, instruction-tuned and larger models exhibited better confidence calibration when using prompting methods than their non-tuned counterparts. Finally, models demonstrated widespread instability under semantic variations: altering the wording of a question or statement frequently caused substantial fluctuations in estimated confidence, showing that models often fail to abstract facts independently of specific phrasing.
These findings highlight significant practical implications for operational risk, cost, and safety. Relying on simple prompting or output token probabilities to detect hallucinations in high-stakes environments creates a false sense of security due to poor reliability. Conversely, extracting confidence via internal probes provides superior risk mitigation but introduces architectural constraints, requiring white-box access to model weights and labeled training data.
Organizations deploying open-weight models should prioritize hidden-state probing to monitor output factuality and gate critical workflows. Where only black-box instruction-tuned models are available, verbalized confidence or surrogate token probabilities can serve as secondary alternatives, though decision-makers must account for their lower reliability. Teams should also implement consistency-checking pipelines that test multiple semantically equivalent prompts to identify unstable factual knowledge before acting on critical model responses.
Decision-makers should view these results in light of specific limitations. The empirical analysis focused primarily on atomic, single-entity facts in academic benchmarks and did not evaluate multi-step reasoning, complex document synthesis, or proprietary black-box models. Further pilot testing and analysis are necessary to establish whether these probing advantages fully generalize to complex, enterprise-scale reasoning tasks.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its early tests of whether language models can recognize their own correct and incorrect answers establish the core self-knowledge problem that this study compares confidence estimators against.
- Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models, Potsawee Manakul et al. (2023). SelfCheckGPT introduces response-consistency sampling as a black-box hallucination detector, making it useful context for one of the source’s main estimator families.
- Paper: Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?, Kevin Liu et al. (2023). Its comparison of internal hidden-state probes with direct model answers clarifies the probe-based confidence approach that the source finds strongest.
- Paper: The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models, Aviv Slobodkin et al. (2023). Its use of hidden-state classifiers to detect when a model cannot answer provides a concrete precursor to the source’s supervised probing method.
- Paper: Inducing Artificial Uncertainty in Language Models, Sophia Hager et al. (2026). It takes the source’s finding that hidden-state probes estimate correctness well further by showing how to train them when difficult calibration examples are unavailable.
- Paper: Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning, Lorenzo Jaime Flores et al. (2026). It extends the comparison of confidence metrics by testing whether their calibration survives supervised fine-tuning across models and tasks.
- Paper: LUQ: Long-text Uncertainty Quantification for LLMs, Caiqi Zhang et al. (2024). It carries uncertainty estimation beyond the source’s atomic factual answers to sentence-level confidence in long-form generation.
