keyword
selective prediction
Selective prediction is a machine learning formulation in which a model is allowed to abstain from making a prediction on an input when its confidence is low or uncertainty is high. Rather than forcing a model to generate an output for every query, this framework evaluates internal confidence or uncertainty metrics to selectively produce answers only for instances meeting a reliability threshold. The selection mechanism commonly relies on uncertainty quantification techniques such as probability estimates, verbalized confidence, or consistency across sampled outputs. By trading off the proportion of answered queries against output accuracy, selective prediction enhances the trustworthiness of automated systems and helps prevent costly errors or incorrect outputs in high-stakes settings.
3 items

Selectively Answering Ambiguous Questions
Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein
Why you should read this
Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.
Trustworthy language models should abstain from answering questions when they do not know the answer. However, the answer to a question can be unknown for a variety of reasons. Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context. We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous. In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work. We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning. Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.
Added
2026-10-03

Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
Lorenzo Jaime Flores, Cesare Spinoso di-Piano, Jackie Cheung
Why you should read this
Demonstrates that supervised fine-tuning unpredictably alters language model confidence calibration across text generation tasks, significantly undermining the effectiveness of standard uncertainty metrics for hallucination detection and selective prediction.
Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection and selective prediction. While prior work has developed various confidence metrics and demonstrated their calibration for classification tasks or using verbalized confidence, the robustness of probability-based and self-consistency-based UQ metrics for natural language generation remains underexplored particularly under model adaptation. Since practitioners routinely apply supervised fine-tuning to adapt models to new tasks, a key question arises: do confidence metrics maintain their calibration when models are fine-tuned? We investigate this question across NLG tasks including translation, question answering, and mathematical reasoning. We find that calibration shifts substantially after SFT: across 216 configurations, it degrades in 112 cases and improves in 104, with confidence scores shifting due to factors beyond output quality, such as proximity to the training distribution. Degradation is therefore neither universal nor rare, and its direction cannot be anticipated from the pre-SFT model. Through a downstream task evaluation, we show that this miscalibration substantially reduces the practical utility of confidence scores for identifying correct answers. Our findings reveal that existing confidence metrics for NLG cannot be reliably deployed off-the-shelf after fine-tuning, highlighting the need for calibration-robust UQ methods under model adaptation.
Added
2026-09-29

Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
Gal Yona, Roee Aharoni, Mor Geva
Why you should read this
Reveals that leading large language models consistently fail to faithfully communicate their intrinsic confidence in natural language, exposing critical limitations in how models express certainty and hedge their answers on knowledge-intensive tasks.
We posit that large language models (LLMs) should be capable of expressing their intrinsic uncertainty in natural language. For example, if the LLM is equally likely to output two contradicting answers to the same question, then its generated response should reflect this uncertainty by hedging its answer (e.g., “I’m not sure, but I think…”). We formalize faithful response uncertainty based on the gap between the model’s intrinsic confidence in the assertions it makes and the decisiveness by which they are conveyed. This example-level metric reliably indicates whether the model reflects its uncertainty, as it penalizes both excessive and insufficient hedging. We evaluate a variety of aligned LLMs at faithfully communicating uncertainty on several knowledge-intensive question answering tasks. Our results provide strong evidence that modern LLMs are poor at faithfully conveying their uncertainty, and that better alignment is necessary to improve their trustworthiness.
Added
2026-09-26
