Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Katherine TianEric MitchellAllan ZhouArchit SharmaRafael RafailovHuaxiu YaoChelsea FinnChristopher D. Manning
Demonstrates that directly prompting RLHF-tuned language models like GPT-4 and Claude to verbalize their confidence scores achieves significantly better calibration than extracting their conditional probabilities, reducing expected calibration error by up to 50% across standard benchmarks.
Deploying artificial intelligence systems in high-stakes environments requires trustworthy uncertainty estimates, enabling systems to flag or defer low-confidence predictions to human experts. While base language models generally output well-calibrated internal probabilities, fine-tuning them with reinforcement learning from human feedback to better follow user instructions often degrades this calibration, causing models to assert incorrect information with unwarranted confidence.
The article evaluates practical methods for extracting calibrated confidence scores from human-feedback fine-tuned language models. It demonstrates that directly asking models to state their confidence in natural language or numbers provides significantly more reliable uncertainty estimates than using their underlying token probabilities.
To evaluate uncertainty extraction, the researchers tested five widely used models—including ChatGPT, GPT-4, Claude 1, Claude 2, and Llama-2-70B-Chat—across three factual question-answering benchmarks comprising roughly 2,800 total questions. The study compared the models' internal conditional probabilities (estimated via sampling) against direct verbalization techniques. These techniques prompted models to output numerical probabilities or descriptive likelihood phrases (such as "likely" or "almost certain") in either single-stage or two-stage dialogues, alongside prompt variants that required generating multiple answer choices or step-by-step reasoning.
The investigation yielded four central findings. First, verbalized confidence scores were consistently better-calibrated than raw internal model probabilities across leading proprietary systems, frequently reducing the expected calibration error by roughly 50%. Second, prompting models to generate and evaluate multiple candidate answers before stating a confidence score notably improved calibration, mirroring human psychological strategies for reducing overconfidence. Third, language models verbalized numerical probabilities with equal or greater accuracy than descriptive text phrases. Fourth, incorporating chain-of-thought reasoning into prompts failed to improve confidence calibration.
These findings have direct operational and cost implications for organizations deploying language models. Because proprietary application programming interfaces (APIs) rarely expose underlying token log probabilities and sampling multiple responses is expensive, prompting models to state numerical confidences in a single response provides a simpler, cheaper, and more reliable safeguard against automated errors and hallucinations. Notably, this ability appears naturally in instruction-following models without requiring extra calibration training.
Organizations seeking to measure model reliability should adopt prompt strategies that ask models to provide their top candidate answers alongside explicit numerical confidence ratings in a single query. When using closed commercial models, teams can apply post-hoc temperature scaling on a small validation dataset to further enhance calibration. Developers should avoid adding chain-of-thought reasoning solely for calibration purposes, as it increases latency and token costs without measurable gains.
These conclusions carry certain limitations. The evaluation focused primarily on short-form factual question-answering rather than complex mathematical reasoning, multi-step logic, or long-form generation. Furthermore, the open-source Llama-2 model exhibited less consistent calibration improvements than proprietary systems, and the opacity of closed-source model training limits deeper analysis. Nonetheless, for short-form factual tasks in advanced commercial systems, there is high confidence that verbalized confidence scoring significantly improves prediction reliability.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Read this earlier study of language models’ self-evaluation first to understand the confidence-elicitation foundation that the source tests after human-feedback fine-tuning.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Its account of calibration, expected calibration error, and temperature scaling supplies the core measurement concepts the source uses to compare confidence estimates.
- Paper: Linguistic Calibration of Long-Form Generations, Neil Band et al. (2024). This study carries verbal confidence beyond short answers, developing training methods for calibrated uncertainty in long-form generations.
- Paper: SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales, Tianyang Xu et al. (2024). SaySelf advances confidence prompting into a trained, single-pass approach that pairs calibrated scores with explanations of knowledge gaps.
- Paper: Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?, Gal Yona et al. (2024). This follow-up tests whether verbal hedging actually reflects models’ intrinsic uncertainty, probing a key limitation of confidence stated in words.
- Paper: Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning, Lorenzo Jaime Flores et al. (2026). By testing how supervised fine-tuning changes confidence-metric calibration, this later study extends the source’s concern with calibration after model tuning.
