Built independently by an author, for readers. Read the story and support ChapterPal

keyword

epistemic uncertainty

Epistemic uncertainty refers to uncertainty that arises from a lack of knowledge, incomplete information, or insufficient data about the true underlying system or model. Often described as model uncertainty or reducible uncertainty, it reflects the degree of doubt a predictive model or learning system has regarding its own parameters, structure, or outputs when exposed to unfamiliar or sparsely sampled regions of the problem domain. Unlike aleatoric uncertainty, which stems from the inherent, irreducible randomness of a data-generating process, epistemic uncertainty can theoretically be reduced or eliminated by gathering more training data, incorporating additional domain knowledge, or refining the model. Consequently, estimating epistemic uncertainty is essential for tasks such as active learning, out-of-distribution detection, and reliable decision-making in safety-critical applications.

11 items

Distinguishing the Knowable from the Unknowable with Language Models

Distinguishing the Knowable from the Unknowable with Language Models

Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, Benjamin L. Edelman

OrganizationsHarvard University

Why you should read this

Demonstrates that language models internally separate reducible epistemic uncertainty from inherent aleatoric entropy, enabling simple probes and unsupervised methods to accurately detect when an uncertain prediction is caused by a lack of knowledge rather than inherent ambiguity.

We study the feasibility of identifying epistemic uncertainty (reflecting a lack of knowledge), as opposed to aleatoric uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over free-form text. In the absence of ground-truth probabilities, we explore a setting where, in order to (approximately) disentangle a given LLM’s uncertainty, a significantly larger model stands in as a proxy for the ground truth. We show that small linear probes trained on the embeddings of frozen, pretrained models accurately predict when larger models will be more confident at the token level and that probes trained on one text domain generalize to others. Going further, we propose a fully unsupervised method that achieves non-trivial accuracy on the same task. Taken together, we interpret these results as evidence that LLMs naturally contain internal representations of different types of uncertainty that could potentially be leveraged to devise more informative indicators of model confidence in diverse practical settings. Code can be found at: https://github.com/KempnerInstitute/llm_uncertainty

Added

2026-10-05

Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

Fabian Falck, Ziyu Wang, Christopher C. Holmes

OrganizationsUniversity of California, BerkeleyUniversity of Oxford

Why you should read this

Falsifies the widely held belief that in-context learning is Bayesian by deriving martingale-based diagnostic tests and showing that large language models violate the fundamental predictive invariance and uncertainty scaling expected of exchangeable Bayesian systems.

In-context learning (ICL) has emerged as a particularly remarkable characteristic of Large Language Models (LLM): given a pretrained LLM and an observed dataset, LLMs can make predictions for new data points from the same distribution without fine-tuning. Numerous works have postulated ICL as approximately Bayesian inference, rendering this a natural hypothesis. In this work, we analyse this hypothesis from a new angle through the martingale property, a fundamental requirement of a Bayesian learning system for exchangeable data. We show that the martingale property is a necessary condition for unambiguous predictions in such scenarios, and enables a principled, decomposed notion of uncertainty vital in trustworthy, safety-critical systems. We derive actionable checks with corresponding theory and test statistics which must hold if the martingale property is satisfied. We also examine if uncertainty in LLMs decreases as expected in Bayesian learning when more data is observed. In three experiments, we provide evidence for violations of the martingale property, and deviations from a Bayesian scaling behaviour of uncertainty, falsifying the hypothesis that ICL is Bayesian.

Added

2026-10-05

Activation-Space Uncertainty Quantification for Pretrained Networks

Activation-Space Uncertainty Quantification for Pretrained Networks

Richard Bergna, Stefan Depeweg, Sergio Calvo Ordoñez, Jonathan Plenk, Álvaro Cartea, Jose Miguel Hernández-Lobato

OrganizationsDepartment of EngineeringMathematical InstituteSiemens AGUniversity of CambridgeUniversity of Oxford

Why you should read this

Introduces Gaussian Process Activations, a post-hoc method that enables single-pass, closed-form uncertainty quantification for pretrained vision and language models while strictly preserving their original predictions without retraining or sampling.

Reliable uncertainty estimates are crucial for deploying pretrained models; yet, many strong methods for quantifying uncertainty require retraining, Monte Carlo sampling, or expensive second-order computations and may alter a frozen backbone's predictions. To address this, we introduce Gaussian Process Activations (GAPA), a post-hoc method that shifts Bayesian modeling from weights to activations. GAPA replaces standard nonlinearities with Gaussian-process activations whose posterior mean exactly matches the original activation, preserving the backbone's point predictions by construction while providing closed-form epistemic variances in activation space. To scale to modern architectures, we use a sparse variational inducing-point approximation over cached training activations, combined with local k-nearest-neighbor subset conditioning, enabling deterministic single-pass uncertainty propagation without sampling, backpropagation, or second-order information. Across regression, classification, image segmentation, and language modeling, GAPA matches or outperforms strong post-hoc baselines in calibration and out-of-distribution detection while remaining efficient at test time.

Added

2026-10-04

Uncertainty Quantification for In-Context Learning of Large Language Models

Uncertainty Quantification for In-Context Learning of Large Language Models

Chen Ling, Xujiang Zhao, Xuchao Zhang, Wei Cheng, Yanchi Liu, Yiyou Sun, Mika Oishi, Takao Osaki, Katsushi Matsuda, Jie Ji, Guangji Bai, Liang Zhao, Haifeng Chen

OrganizationsEmory UniversityMicrosoftNEC CorporationNEC Laboratories America, Inc.

Why you should read this

Presents a Bayesian framework that decomposes predictive uncertainty in large language model in-context learning into prompt-induced aleatoric and model-induced epistemic components, enabling unsupervised diagnostic evaluation of output reliability across both white-box and black-box settings.

In-context learning has emerged as a ground-breaking ability of Large Language Models (LLMs) and revolutionized various fields by providing a few task-relevant demonstrations in the prompt. However, trustworthy issues with LLM’s response, such as hallucination, have also been actively discussed. Existing works have been devoted to quantifying the uncertainty in LLM’s response, but they often overlook the complex nature of LLMs and the uniqueness of in-context learning. In this work, we delve into the predictive uncertainty of LLMs associated with in-context learning, highlighting that such uncertainties may stem from both the provided demonstrations (aleatoric uncertainty) and ambiguities tied to the model’s configurations (epistemic uncertainty). We propose a novel formulation and corresponding estimation method to quantify both types of uncertainties. The proposed method offers an unsupervised way to understand the prediction of in-context learning in a plug-and-play fashion. Extensive experiments are conducted to demonstrate the effectiveness of the decomposition. The code and data are available at: https://github.com/lingchen0331/UQ_ICL.

Added

2026-10-03

Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Model-Bellman Inconsistency for Model-based Offline Reinforcement Learning

Yihao Sun, Jiaji Zhang, Chengxing Jia, Haoxin Lin, Junyin Ye, Yang Yu

Why you should read this

Proposes MOBILE, a model-based offline reinforcement learning algorithm that measures uncertainty using Bellman estimation inconsistency across a dynamics ensemble to approximate true Bellman errors and achieve state-of-the-art performance on D4RL and NeoRL benchmarks.

For offline reinforcement learning (RL), model-based methods are expected to be data-efficient as they incorporate dynamics models to generate more data. However, due to inevitable model errors, straightforwardly learning a policy in the model typically fails in the offline setting. Previous studies have incorporated conservatism to prevent out-of-distribution exploration. For example, MOPO penalizes rewards through uncertainty measures from predicting the next states, which we have discovered are loose bounds of the ideal uncertainty, i.e., the Bellman error. In this work, we propose MOdel-Bellman Inconsistency penalized OffLinE Policy Optimization (MOBILE), a novel uncertainty-driven offline RL algorithm. MOBILE conducts uncertainty quantification through the inconsistency of Bellman estimations under an ensemble of learned dynamics models, which can be a better approximator to the true Bellman error, and penalizes the Bellman estimation based on this uncertainty. Empirically we have verified that our proposed uncertainty quantification can be significantly closer to the true Bellman error than the compared methods. Consequently, MOBILE outperforms prior offline RL approaches on most tasks of D4RL and NeoRL benchmarks.

Added

2026-10-03

Selectively Answering Ambiguous Questions

Selectively Answering Ambiguous Questions

Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein

OrganizationsDuke UniversityGoogleUniversity of Texas at Austin

Why you should read this

Demonstrates that measuring answer consistency across repeatedly sampled outputs provides a much more reliable confidence score than model likelihoods or self-verification prompts for deciding when language models should abstain from answering ambiguous questions.

Trustworthy language models should abstain from answering questions when they do not know the answer. However, the answer to a question can be unknown for a variety of reasons. Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context. We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous. In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work. We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning. Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions.

Added

2026-10-03

Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Decomposing Uncertainty for Large Language Models through Input Clarification Ensembling

Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, Yang Zhang

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes input clarification ensembling, a practical framework that separates large language model uncertainty into input ambiguity and model knowledge deficits without modifying model parameters or training procedures.

Uncertainty decomposition refers to the task of decomposing the total uncertainty of a predictive model into aleatoric (data) uncertainty, resulting from inherent randomness in the data-generating process, and epistemic (model) uncertainty, resulting from missing information in the model’s training data. In large language models (LLMs) specifically, identifying sources of uncertainty is an important step toward improving reliability, trustworthiness, and interpretability, but remains an important open research question. In this paper, we introduce an uncertainty decomposition framework for LLMs, called input clarification ensembling, which can be applied to any pre-trained LLM. Our approach generates a set of clarifications for the input, feeds them into an LLM, and ensembles the corresponding predictions. We show that, when aleatoric uncertainty arises from ambiguity or under-specification in LLM inputs, this approach makes it possible to factor an (un-clarified) LLM’s predictions into separate aleatoric and epistemic terms, using a decomposition similar to the one employed by Bayesian neural networks. Empirical evaluations demonstrate that input clarification ensembling provides accurate and reliable uncertainty quantification on several language processing tasks. Code and data are available at https://github.com/UCSB-NLP-Chang/llm_uncertainty.

Added

2026-10-01

Inducing Artificial Uncertainty in Language Models

Inducing Artificial Uncertainty in Language Models

Sophia Hager, Simon Zeng, Nicholas Andrews

OrganizationsJohns Hopkins UniversityMicrosoft

Why you should read this

Demonstrates that training uncertainty probes on artificially induced uncertainty in language models significantly improves confidence calibration on difficult tasks where naturally challenging training data is scarce.

In safety-critical applications, language models should be able to characterize their uncertainty with meaningful probabilities. Many uncertainty quantification approaches require supervised data; however, finding suitable unseen challenging data is increasingly difficult for large language models trained on vast amounts of scraped data. If the model is consistently (and correctly) confident in its predictions, the uncertainty quantification method may consistently overestimate confidence on new and unfamiliar data. Finding data which exhibits enough uncertainty to train supervised uncertainty quantification methods for high-performance models may therefore be challenging, and will increase in difficulty as LLMs saturate datasets. To address this issue, we first introduce the problem of inducing artificial uncertainty in language models, then investigate methods of inducing artificial uncertainty on trivially easy data in the absence of challenging data at training time. We use probes trained to recognize artificial uncertainty on the original model, and find that these probes trained on artificial uncertainty outperform probes trained without artificial uncertainty in recognizing real uncertainty, achieving notably higher calibration on hard data with minimal loss of performance on easy data.

Added

2026-09-29

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Danruo Deng, Guangyong Chen, Yang Yu, Furui Liu, Pheng-Ann Heng

OrganizationsInstitute of Medical Intelligence and XRThe Chinese University of Hong KongZhejiang Lab

Why you should read this

Proposes a Fisher Information-based evidential deep learning framework that dynamically reweights loss terms to prevent over-penalizing ambiguous training samples, significantly improving uncertainty quantification and few-shot classification reliability.

Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network’s outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning (I-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings.

Added

2026-09-26

Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods

Eyke Hüllermeier, Willem Waegeman

OrganizationsGhent UniversityPaderborn University

Why you should read this

Clarifies the critical distinction between irreducible data randomness and reducible model ignorance, providing a comprehensive framework for quantifying both aleatoric and epistemic uncertainty to build safer, more reliable machine learning systems.

The notion of uncertainty is of major importance in machine learning and constitutes a key element of machine learning methodology. In line with the statistical tradition, uncertainty has long been perceived as almost synonymous with standard probability and probabilistic predictions. Yet, due to the steadily increasing relevance of machine learning for practical applications and related issues such as safety requirements, new problems and challenges have recently been identified by machine learning scholars, and these problems may call for new methodological developments. In particular, this includes the importance of distinguishing between (at least) two different types of uncertainty, often referred to as aleatoric and epistemic. In this paper, we provide an introduction to the topic of uncertainty in machine learning as well as an overview of attempts so far at handling uncertainty in general and formalizing this distinction in particular.

Added

2026-09-16

Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift

Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift

Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, Jasper Snoek

OrganizationsGoogle

Why you should read this

Establishes a critical large-scale benchmark for evaluating state-of-the-art predictive uncertainty quantification methods in machine learning, specifically under challenging dataset shift conditions, revealing that methods marginalizing over models significantly outperform traditional approaches.

Modern machine learning methods including deep learning have achieved great success in predictive accuracy for supervised learning tasks, but may still fall short in giving useful estimates of their predictive {\em uncertainty}. Quantifying uncertainty is especially critical in real-world settings, which often involve input distributions that are shifted from the training distribution due to a variety of factors including sample bias and non-stationarity. In such settings, well calibrated uncertainty estimates convey information about when a model's output should (or should not) be trusted. Many probabilistic deep learning methods, including Bayesian-and non-Bayesian methods, have been proposed in the literature for quantifying predictive uncertainty, but to our knowledge there has not previously been a rigorous large-scale empirical comparison of these methods under dataset shift. We present a large-scale benchmark of existing state-of-the-art methods on classification problems and investigate the effect of dataset shift on accuracy and calibration. We find that traditional post-hoc calibration does indeed fall short, as do several other previous methods. However, some methods that marginalize over models give surprisingly strong results across a broad spectrum of tasks.

Added

2026-04-27

License

Published with permission