Uncertainty Estimation by Fisher Information-based Evidential Deep Learning
Danruo DengGuangyong ChenYang YuFurui LiuPheng-Ann Heng
Proposes a Fisher Information-based evidential deep learning framework that dynamically reweights loss terms to prevent over-penalizing ambiguous training samples, significantly improving uncertainty quantification and few-shot classification reliability.
Reliable uncertainty estimation is essential for deploying deep learning models in safety-critical domains such as medical diagnosis and autonomous systems, where models must signal when they are uncertain or encountering unfamiliar data. Existing evidential deep learning frameworks address this by predicting parameters for a Dirichlet probability distribution over classes, but they suffer from a key limitation: when training on ambiguous or noisy samples labeled with rigid one-hot targets, the framework over-penalizes and suppresses evidence for alternative, plausible classes, leading to underestimated data uncertainty and model overfitting.
The article introduces and evaluates Fisher Information-based Evidential Deep Learning, a framework designed to improve classification accuracy and the reliability of uncertainty quantification without requiring out-of-distribution training data. The primary objective is to demonstrate that measuring class informativeness via the Fisher Information Matrix allows the network to dynamically adapt its loss function, preserving evidence for uncertain classes and avoiding overconfidence.
To accomplish this, the authors model label generation through an anisotropic multivariate Gaussian distribution whose variance is governed by the inverse of the Fisher Information Matrix. They also integrate a PAC-Bayesian bound to theoretically ground generalization performance. The approach was evaluated across standard image benchmarks (MNIST and CIFAR10) and complex few-shot classification setups on mini-ImageNet and tiered-ImageNet across 10,000 evaluation episodes, testing confidence calibration, out-of-distribution detection, and noisy sample identification against standard baseline architectures.
The findings show consistent and substantial performance gains over existing evidential and Bayesian baselines. On standard image classification, the proposed method raised CIFAR10 accuracy from 83.55% to 89.20% and improved confidence estimation. In out-of-distribution detection across four standard benchmark pairs, the method improved precision-recall metrics by 0.5% to 3.8% over runner-up techniques without exposure to out-of-distribution training data. In few-shot settings, classification accuracy gains ranged from 1.62% to 9.31% (reaching 78.60% under 10-way 20-shot conditions), while out-of-distribution detection performance rose by up to 9.36%. Furthermore, the framework improved noisy sample detection by more than 11% compared to competing methods.
These results demonstrate that incorporating information-theoretic weighting reduces operational risks by making uncertainty estimates far more separable between familiar and unfamiliar inputs. For practitioners, this translates to improved model reliability and safety without introducing costly computational overhead or separate calibration steps during inference. The findings challenge the conventional practice of penalizing non-target class evidence uniformly during evidential model training.
Organizations developing safety-critical classification pipelines should consider adopting information-weighted evidential training to enhance failure detection. In implementation, practitioners must carefully balance the weighting hyperparameters, as optimal settings for out-of-distribution detection slightly diverge from those maximizing pure classification accuracy.
The primary limitation of the current work is its mathematical reliance on the Dirichlet distribution, restricting its direct application to discrete classification tasks rather than continuous regression problems. Nevertheless, given the extensive empirical evaluations and narrow confidence intervals reported across thousands of test episodes, decision-makers can place high confidence in these results for image classification and anomaly detection workflows.
- Paper: A survey of uncertainty in deep neural networks, Jakob Gawlikowski et al. (2021). Provides a comprehensive taxonomy and evaluation of deep uncertainty estimation methods, establishing the foundational concepts of single-pass evidential models and Dirichlet prior networks that this paper builds upon.
- Paper: Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods, Eyke Hüllermeier et al. (2019). Clarifies the theoretical and conceptual distinctions between aleatoric and epistemic uncertainties, which are essential for understanding the representation learning dynamics in evidential deep learning.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). Introduces dynamic example re-weighting strategies to combat label noise and imbalance, motivating the source paper's Fisher Information-based objective loss reweighting mechanism.
- Paper: A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges, Moloud Abdar et al. (2020). Offers a broad survey of uncertainty quantification techniques and applications across deep neural architectures, providing key context for the evidential deep learning problem setting.
- Paper: On Calibration of Modern Neural Networks, Chuan Guo et al. (2017). Establishes standard metrics and insights into why modern neural networks produce miscalibrated and overconfident predictive confidences.
No sufficiently relevant recommendations were found.
