Built independently by an author, for readers. Read the story and support ChapterPal

keyword

prior networks

Prior networks are neural network models designed for predictive uncertainty estimation that parameterize a prior probability distribution directly over the model predictive distributions. Rather than placing probability distributions over the network weights as in Bayesian neural networks, or producing a single point estimate for predictions, a prior network outputs the parameters of a higher-order conjugate distribution, such as a Dirichlet distribution for classification tasks or a Normal-Wishart distribution for regression tasks. This structure enables a single deterministic neural network to emulate ensemble behavior in a single forward pass and to explicitly separate data uncertainty, which arises from inherent class overlap or measurement noise, from distributional uncertainty, which occurs when test inputs diverge from the training data distribution.

3 items

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Danruo Deng, Guangyong Chen, Yang Yu, Furui Liu, Pheng-Ann Heng

OrganizationsInstitute of Medical Intelligence and XRThe Chinese University of Hong KongZhejiang Lab

Why you should read this

Proposes a Fisher Information-based evidential deep learning framework that dynamically reweights loss terms to prevent over-penalizing ambiguous training samples, significantly improving uncertainty quantification and few-shot classification reliability.

Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network’s outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning (I-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings.

Added

2026-09-26

Deep Ensembles Work, But Are They Necessary?

Deep Ensembles Work, But Are They Necessary?

Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel, John P. Cunningham

OrganizationsColumbia University

Why you should read this

Demonstrates that the uncertainty quantification and distribution-shift advantages typically credited to deep ensembles can be fully matched by properly scaled single neural networks.

Ensembling neural networks is an effective way to increase accuracy, and can often match the performance of individual larger models. This observation poses a natural question: given the choice between a deep ensemble and a single neural network with similar accuracy, is one preferable over the other? Recent work suggests that deep ensembles may offer distinct benefits beyond predictive power: namely, uncertainty quantification and robustness to dataset shift. In this work, we demonstrate limitations to these purported benefits, and show that a single (but larger) neural network can replicate these qualities. First, we show that ensemble diversity, by any metric, does not meaningfully contribute to an ensemble’s un-certainty quantification on out-of-distribution (OOD) data, but is instead highly correlated with the relative improvement of a single larger model. Second, we show that the OOD performance afforded by ensembles is strongly determined by their in-distribution (InD) performance, and—in this sense—is not indicative of any “effective robustness.” While deep ensembles are a practical way to achieve improvements to predictive power, uncertainty quantification, and robustness, our results show that these improvements can be replicated by a (larger) single model.

Added

2026-09-26