Built independently by an author, for readers. Read the story and support ChapterPal

keyword

uncertainty estimation

Uncertainty estimation is the process of quantifying the degree of confidence or doubt associated with a machine learning model's predictions and outputs. Rather than providing only a single definitive answer, systems equipped with uncertainty estimation measure the reliability of their decisions, helping to identify when outputs are likely to be incorrect, ambiguous, or based on data that differs from their training distribution. This capability is critical for safety-sensitive decision-making, misclassification detection, and human-AI collaboration, as it enables models to express calibrated confidence and helps developers distinguish between inherent noise in the input data and limitations in the model's own knowledge. Common techniques for estimating predictive uncertainty include Bayesian neural network approximations, ensemble methods, evidential learning frameworks, and consistency- or probability-based scoring methods.

14 items

Referential Uncertainty in Human-AI Collaboration

Referential Uncertainty in Human-AI Collaboration

Christian Poelitz, Finale Doshi-Velez, Siân Lindley

OrganizationsHarvard UniversityMicrosoft

Why you should read this

Demonstrates that frontier vision-language models routinely conceal referential ambiguity from human partners, and proves that accurately targeted uncertainty hedges cut human acceptance of AI errors from 78% to 36%.

Effective human-AI collaboration requires partners to establish references through interaction, which becomes fragile when descriptions are ambiguous, similar referents compete, or partners see different things. We study referential uncertainty - uncertainty over which candidate object a description refers to - in a collaborative puzzle task where a human Helper instructs an AI Worker to place pieces. The Worker must identify and communicate its uncertainty, and the Helper must recognize and act on it. We show that a separately elicited belief distribution over candidate pieces is better calibrated (ECE 0.15) and better discriminates correct from incorrect placements (AUROC 0.65) than raw action-token probabilities, which are severely overconfident (0.97 mean confidence, ECE 0.44). Across three frontier vision-language models (GPT-4.1, GPT-5, GPT-5.5), this elicited uncertainty rises predictably with instruction vagueness, but not with competing referents in context, even when those increase errors. The models seldom externalize it, asking for clarification on only 3.5-16.7% of turns. In a controlled human study (N=210), participants given only the Worker's default message accept 78% of wrong placements and cannot tell right from wrong (AUC 0.50). Precise descriptions and, especially, well-targeted hedges cut wrong-move acceptance to 36% while largely preserving correct-move acceptance, compensating for missing shared awareness such as not seeing the Worker's action. But this benefit depends on targeting: a deployable hedge derived from the model's own belief entropy inherits that signal's weakness and can do more harm than good. Externalized uncertainty helps a human partner only when it is accurately targeted.

Added

2026-10-05

Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, Bryan Hooi

OrganizationsÉcole Polytechnique Fédérale de LausanneNational University of SingaporeThe Hong Kong University of Science and Technology

Why you should read this

Presents a systematic framework for evaluating black-box uncertainty estimation in large language models, showing that verbalized prompting and response consistency can effectively reduce overconfidence to rival white-box calibration methods.

Empowering large language models to accurately express confidence in their answers is essential for trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on white-box access to internal model information or model fine-tuning, have become less suitable for LLMs, especially closed-source commercial APIs. This leads to a growing need to explore the untapped area of black-box approaches for LLM uncertainty estimation. To better break down the problem, we define a systematic framework with three components: prompting strategies for eliciting verbalized confidence, sampling methods for generating multiple responses, and aggregation techniques for computing consistency. We then benchmark these methods on two key tasks-confidence calibration and failure prediction-across five types of datasets (e.g., commonsense and arithmetic reasoning) and five widely-used LLMs including GPT-4 and LLaMA 2 Chat. Our analysis uncovers several key insights: 1) LLMs, when verbalizing their confidence, tend to be overconfident, potentially imitating human patterns of expressing confidence. 2) As model capability scales up, both calibration and failure prediction performance improve. 3) Employing our proposed strategies, such as human-inspired prompts, consistency among multiple responses, and better aggregation strategies can help mitigate this overconfidence from various perspectives. 4) Comparisons with white-box methods indicate that while white-box methods perform better, the gap is narrow, e.g., 0.522 to 0.605 in AUROC. Despite these advancements, none of these techniques consistently outperform others, and all investigated methods struggle in challenging tasks, such as those requiring professional knowledge, indicating significant scope for improvement. We believe this study can serve as a strong baseline and provide insights for eliciting confidence in black-box LLMs.

Added

2026-10-04

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr

OrganizationsAmazonUniversity of Southern California

Why you should read this

Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.

Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.

Added

2026-10-04

Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Kaitlyn Zhou, Dan Jurafsky, Tatsunori Hashimoto

OrganizationsStanford University

Why you should read this

Reveals that prompt expressions of high confidence paradoxically degrade language model accuracy by up to seven percent compared to expressions of uncertainty, exposing how pretraining data patterns cause models to mimic surface linguistic habits rather than genuinely track factual truth.

The increased deployment of LMs for real-world tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs. Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures. We develop a typology of epistemic markers and inject 50 markers into prompts for question answering. We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%. Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance. Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions. These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.

Added

2026-10-01

Inducing Artificial Uncertainty in Language Models

Inducing Artificial Uncertainty in Language Models

Sophia Hager, Simon Zeng, Nicholas Andrews

OrganizationsJohns Hopkins UniversityMicrosoft

Why you should read this

Demonstrates that training uncertainty probes on artificially induced uncertainty in language models significantly improves confidence calibration on difficult tasks where naturally challenging training data is scarce.

In safety-critical applications, language models should be able to characterize their uncertainty with meaningful probabilities. Many uncertainty quantification approaches require supervised data; however, finding suitable unseen challenging data is increasingly difficult for large language models trained on vast amounts of scraped data. If the model is consistently (and correctly) confident in its predictions, the uncertainty quantification method may consistently overestimate confidence on new and unfamiliar data. Finding data which exhibits enough uncertainty to train supervised uncertainty quantification methods for high-performance models may therefore be challenging, and will increase in difficulty as LLMs saturate datasets. To address this issue, we first introduce the problem of inducing artificial uncertainty in language models, then investigate methods of inducing artificial uncertainty on trivially easy data in the absence of challenging data at training time. We use probes trained to recognize artificial uncertainty on the original model, and find that these probes trained on artificial uncertainty outperform probes trained without artificial uncertainty in recognizing real uncertainty, achieving notably higher calibration on hard data with minimal loss of performance on easy data.

Added

2026-09-29

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Danruo Deng, Guangyong Chen, Yang Yu, Furui Liu, Pheng-Ann Heng

OrganizationsInstitute of Medical Intelligence and XRThe Chinese University of Hong KongZhejiang Lab

Why you should read this

Proposes a Fisher Information-based evidential deep learning framework that dynamically reweights loss terms to prevent over-penalizing ambiguous training samples, significantly improving uncertainty quantification and few-shot classification reliability.

Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network’s outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning (I-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings.

Added

2026-09-26

Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Gal Yona, Roee Aharoni, Mor Geva

OrganizationsGoogleTel Aviv University

Why you should read this

Reveals that leading large language models consistently fail to faithfully communicate their intrinsic confidence in natural language, exposing critical limitations in how models express certainty and hedge their answers on knowledge-intensive tasks.

We posit that large language models (LLMs) should be capable of expressing their intrinsic uncertainty in natural language. For example, if the LLM is equally likely to output two contradicting answers to the same question, then its generated response should reflect this uncertainty by hedging its answer (e.g., “I’m not sure, but I think…”). We formalize faithful response uncertainty based on the gap between the model’s intrinsic confidence in the assertions it makes and the decisiveness by which they are conveyed. This example-level metric reliably indicates whether the model reflects its uncertainty, as it penalizes both excessive and insufficient hedging. We evaluate a variety of aligned LLMs at faithfully communicating uncertainty on several knowledge-intensive question answering tasks. Our results provide strong evidence that modern LLMs are poor at faithfully conveying their uncertainty, and that better alignment is necessary to improve their trustworthiness.

Added

2026-09-26

Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness

Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness

Francesco Pinto, Harry Yang, Ser Nam Lim, Philip H. S. Torr, Puneet K. Dokania

OrganizationsFive AIMetaUniversity of Oxford

Why you should read this

Proposes RegMixup, a simple modification that applies Mixup as an auxiliary regularizer alongside standard cross-entropy loss rather than as a standalone objective, substantially improving classification accuracy and out-of-distribution detection without requiring architectural changes or costly ensembles.

We show that the effectiveness of the well celebrated Mixup [Zhang et al., 2018] can be further improved if instead of using it as the sole learning objective, it is utilized as an additional regularizer to the standard cross-entropy loss. This simple change not only improves accuracy but also significantly improves the quality of the predictive uncertainty estimation of Mixup in most cases under various forms of covariate shifts and out-of-distribution detection experiments. In fact, we observe that Mixup otherwise yields much degraded performance on detecting out-of-distribution samples possibly, as we show empirically, due to its tendency to learn models exhibiting high-entropy throughout; making it difficult to differentiate in-distribution samples from out-of-distribution ones. To show the efficacy of our approach (RegMixup²), we provide thorough analyses and experiments on vision datasets (ImageNet & CIFAR-10/100) and compare it with a suite of recent approaches for reliable uncertainty estimation.

Added

2026-09-26

A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges

A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges

Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Acharya, Vladimir Makarenkov, Saeid Nahavandi

OrganizationsChinese Academy of SciencesDeakin UniversityDibrugarh UniversityGoogleNgee Ann PolytechnicShenzhen UniversitySwinburne University of TechnologyUniversité du Québec à MontréalUniversity of OuluUniversity of Waterloo

Why you should read this

Surveys Bayesian approximations, deep ensembles, and non-Bayesian uncertainty quantification techniques across deep learning and reinforcement learning, comparing their practical implementations in computer vision, medical diagnostics, and natural language processing while outlining unresolved research challenges.

Abstract—Uncertainty quantification (UQ) plays a pivotal role in the reduction of uncertainties during both optimization and decision making, applied to solve a variety of real-world applications in science and engineering. Bayesian approximation and ensemble learning techniques are two of the most widely-used UQ methods in the literature. In this regard, researchers have proposed different UQ methods and examined their performance in a variety of applications such as computer vision (e.g., self-driving cars and object detection), image processing (e.g., image restoration), medical image analysis (e.g., medical image classification and segmentation), natural language processing (e.g., text classification, social media texts and recidivism risk-scoring), bioinformatics, etc. This study reviews recent advances in UQ methods used in deep learning, investigates the application of these methods in reinforcement learning, and highlight the fundamental research challenges and directions associated with the UQ field.

Added

2026-09-14

Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning

Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning

Yarin Gal, Zoubin Ghahramani

OrganizationsUniversity of Cambridge

Why you should read this

Establishes a novel theoretical framework that transforms dropout training into approximate Bayesian inference within deep Gaussian processes, enabling deep neural networks to quantify model uncertainty without compromising computational efficiency or test accuracy.

Deep learning tools have gained tremendous attention in applied machine learning. However such tools for regression and classification do not capture model uncertainty. In comparison, Bayesian models offer a mathematically grounded framework to reason about model uncertainty, but usually come with a prohibitive computational cost. In this paper we develop a new theoretical framework casting dropout training in deep neural networks (NNs) as approximate Bayesian inference in deep Gaussian processes. A direct result of this theory gives us tools to model uncertainty with dropout NNs -- extracting information from existing models that has been thrown away so far. This mitigates the problem of representing uncertainty in deep learning without sacrificing either computational complexity or test accuracy. We perform an extensive study of the properties of dropout's uncertainty. Various network architectures and non-linearities are assessed on tasks of regression and classification, using MNIST as an example. We show a considerable improvement in predictive log-likelihood and RMSE compared to existing state-of-the-art methods, and finish by using dropout's uncertainty in deep reinforcement learning.

Added

2026-04-27

Creative Commons License