Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao XiongZhiyuan HuXinyang LuYifei LiJie FuJunxian HeBryan Hooi
Presents a systematic framework for evaluating black-box uncertainty estimation in large language models, showing that verbalized prompting and response consistency can effectively reduce overconfidence to rival white-box calibration methods.
As large language models become central to commercial applications, accurately assessing their uncertainty is critical for risk assessment, safety, and reducing factual errors. Most commercial models operate as black boxes accessible only via text prompts, making traditional white-box methods—such as inspecting internal probabilities or performing expensive fine-tuning—infeasible. Consequently, decision-makers face major reliability risks when relying on a model's stated confidence.
The article systematically evaluates how effectively large language models express their own uncertainty using purely black-box techniques. It introduces a modular framework combining human-inspired prompting, response sampling, and consistency aggregation to measure and improve confidence calibration and failure prediction across diverse tasks.
The researchers benchmarked five major language models (including GPT-4, GPT-3.5, and LLaMA 2) across eight datasets spanning arithmetic, commonsense, symbolic, ethical, and professional knowledge domains. They tested five prompt designs (such as asking for step-by-step reasoning or multiple guesses), three sampling methods (random temperature variations, paraphrasing, and misleading cues), and several mathematical aggregation rules to compute uncertainty from multiple outputs.
The investigation produced five key findings. First, when asked directly, models display severe overconfidence, routinely assigning 80% to 100% confidence even to incorrect answers and mimicking human conversational rounding by expressing values in multiples of five. Second, while larger models like GPT-4 show improved calibration compared to older models like GPT-3, their ability to flag their own incorrect answers remains poor, with failure prediction metrics often hovering near random guessing. Third, human-inspired prompting strategies (such as eliciting multiple ranked guesses) improve calibration mainly by boosting task accuracy rather than refining the model’s true self-awareness of error. Fourth, sampling multiple responses significantly improves error detection—for instance, boosting failure prediction accuracy from near-random to over 92% on arithmetic tasks. Fifth, hybrid aggregation methods combining response consistency with stated verbal confidence consistently outperform methods relying solely on answer agreement.
These results demonstrate that a single query cannot reliably convey model uncertainty, posing operational and compliance risks if deployed blindly in high-stakes environments. While black-box methods narrow the gap with traditional internal-probability techniques, models still struggle substantially on complex tasks requiring specialized expertise, such as professional law.
For practical deployment, the article recommends a balanced, stable approach: prompt the model to generate multiple top guesses, sample several random outputs (around four to five responses to balance cost and accuracy), and aggregate them using hybrid confidence-weighting strategies. Decision-makers should apply caution, as the evaluations were limited to question-answering tasks with single ground-truth answers rather than open-ended text generation. Stated confidence scores should not be treated as objective probabilities without ensemble validation.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Its earlier tests of whether language models can recognize correct answers establish the self-evaluation problem that this paper reexamines under black-box confidence elicitation.
- Paper: Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness, Jiuhai Chen et al. (2024). It extends black-box confidence elicitation by combining sampled-response consistency with self-reflection to detect errors across answer-generation tasks.
- Paper: LUQ: Long-text Uncertainty Quantification for LLMs, Caiqi Zhang et al. (2024). It carries the paper’s sampling-and-consistency approach into long-form generation, estimating uncertainty sentence by sentence and supporting abstention.
- Paper: SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales, Tianyang Xu et al. (2024). It continues the confidence-calibration problem by training models to produce calibrated scores and uncertainty rationales in a single pass.
- Paper: Language Models with Conformal Factuality Guarantees, Christopher Mohri et al. (2024). It turns uncertainty estimates into conformal factuality guarantees by removing claims whose estimated reliability falls below a calibrated threshold.
