Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Danruo DengGuangyong ChenYang YuFurui LiuPheng-Ann Heng

article2023ICML79 citations

Proposes a Fisher Information-based evidential deep learning framework that dynamically reweights loss terms to prevent over-penalizing ambiguous training samples, significantly improving uncertainty quantification and few-shot classification reliability.

Listen

Reliable uncertainty estimation is essential for deploying deep learning models in safety-critical domains such as medical diagnosis and autonomous systems, where models must signal when they are uncertain or encountering unfamiliar data. Existing evidential deep learning frameworks address this by predicting parameters for a Dirichlet probability distribution over classes, but they suffer from a key limitation: when training on ambiguous or noisy samples labeled with rigid one-hot targets, the framework over-penalizes and suppresses evidence for alternative, plausible classes, leading to underestimated data uncertainty and model overfitting.

The article introduces and evaluates Fisher Information-based Evidential Deep Learning, a framework designed to improve classification accuracy and the reliability of uncertainty quantification without requiring out-of-distribution training data. The primary objective is to demonstrate that measuring class informativeness via the Fisher Information Matrix allows the network to dynamically adapt its loss function, preserving evidence for uncertain classes and avoiding overconfidence.

To accomplish this, the authors model label generation through an anisotropic multivariate Gaussian distribution whose variance is governed by the inverse of the Fisher Information Matrix. They also integrate a PAC-Bayesian bound to theoretically ground generalization performance. The approach was evaluated across standard image benchmarks (MNIST and CIFAR10) and complex few-shot classification setups on mini-ImageNet and tiered-ImageNet across 10,000 evaluation episodes, testing confidence calibration, out-of-distribution detection, and noisy sample identification against standard baseline architectures.

The findings show consistent and substantial performance gains over existing evidential and Bayesian baselines. On standard image classification, the proposed method raised CIFAR10 accuracy from 83.55% to 89.20% and improved confidence estimation. In out-of-distribution detection across four standard benchmark pairs, the method improved precision-recall metrics by 0.5% to 3.8% over runner-up techniques without exposure to out-of-distribution training data. In few-shot settings, classification accuracy gains ranged from 1.62% to 9.31% (reaching 78.60% under 10-way 20-shot conditions), while out-of-distribution detection performance rose by up to 9.36%. Furthermore, the framework improved noisy sample detection by more than 11% compared to competing methods.

These results demonstrate that incorporating information-theoretic weighting reduces operational risks by making uncertainty estimates far more separable between familiar and unfamiliar inputs. For practitioners, this translates to improved model reliability and safety without introducing costly computational overhead or separate calibration steps during inference. The findings challenge the conventional practice of penalizing non-target class evidence uniformly during evidential model training.

Organizations developing safety-critical classification pipelines should consider adopting information-weighted evidential training to enhance failure detection. In implementation, practitioners must carefully balance the weighting hyperparameters, as optimal settings for out-of-distribution detection slightly diverge from those maximizing pure classification accuracy.

The primary limitation of the current work is its mathematical reliance on the Dirichlet distribution, restricting its direct application to discrete classification tasks rather than continuous regression problems. Nevertheless, given the extensive empirical evaluations and narrow confidence intervals reported across thousands of test episodes, decision-makers can place high confidence in these results for image classification and anomaly detection workflows.

No sufficiently relevant recommendations were found.

Cover for Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

Abstract

Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network’s outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning (I-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings.

Table of Contents

  • 1. Introduction
  • 2. Preliminary
  • 3. Method
  • 3.1. Generative Model of Evidential Network
  • 3.2. Fisher Information-based Evidential Network
  • 3.3. Learning with PAC-Bayesian Bound
  • 4. Related Work
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Confidence Evaluation
  • 5.3. OOD detection
  • 5.4. Few-shot Learning
  • 5.5. Noisy data detection
  • 5.6. Ablation Study
  • 5.7. Analysis of Uncertainty estimation
  • 6. Conclusion
  • Acknowledgments
  • References
  • A. Derivation and Proof
  • A.1. FIM Derivation for Dirichlet Distribution
  • A.2. Derivation of the objective function Eq. 5
  • A.3. Proof of Theorem 3.1
  • B. Derivations for Uncertainty Measures and Energy Distance
  • B.1. Expected Entropy of Dirichlet-based Uncertainty Models
  • B.2. Mutual Information of Dirichlet-based Uncertainty Models
  • B.3. Differential Entropy of Dirichlet-based Uncertainty Models
  • B.4. Energy Distance
  • C. Experimental Details and Additional Results
  • C.1. Datasets
  • C.2. Implementation details
  • C.3. Additional Experimental Results on OOD detection
  • C.4. Additional Experimental Results on Few-shot Learning
  • C.5. Additional Ablation Study
  • C.6. Additional Analysis of Uncertainty estimation

Knowls

  1. Knowl 1 — Generative Model of Fisher Information-based Evidential Deep Learning

    model/method

    In Fisher Information-based Evidential Deep Learning (I\mathcal{I}-EDL), an evidential neural network fθ:Rd→R+Kf_\theta: \mathbb{R}^d \to \mathbb{R}_+^K predicts the concentration parameters α=fθ(x)+1∈R+K\alpha = f_\theta(x) + \mathbf{1} \in \mathbb{R}_+^K of a Dirichlet distribution Dir(p∣α)\mathrm{Dir}(p|\alpha) over the class probability simplex p∈ΔK−1p \in \Delta^{K-1}.

    Unlike classical Evidential Deep Learning (EDL), which assumes an isotropic Gaussian generative model y∼N(p,σ2I)y \sim \mathcal{N}(p, \sigma^2 I) for one-hot target vectors y∈RKy \in \mathbb{R}^K, I\mathcal{I}-EDL models the target variable yy as generated from a multivariate Gaussian distribution whose covariance matrix is proportional to the inverse Fisher Information Matrix (FIM) I(α)−1\mathcal{I}(\alpha)^{-1} of the Dirichlet distribution:

    y∼N(p,σ2I(α)−1),p∼Dir(α)y \sim \mathcal{N}(p, \sigma^2 \mathcal{I}(\alpha)^{-1}), \quad p \sim \mathrm{Dir}(\alpha)

    where σ2>0\sigma^2 > 0 is a scalar covariance scaling parameter. Because classes with higher predicted evidence possess smaller Fisher information (and thus larger variance in I(α)−1\mathcal{I}(\alpha)^{-1}), this formulation prevents over-penalizing mislabeled classes in high data-uncertainty samples while maximizing the likelihood of observed labels.

  2. Knowl 2 — Closed-Form Fisher Information Matrix and Log-Determinant for Dirichlet Distributions

    equation

    For a Dirichlet distribution Dir(p∣α)\mathrm{Dir}(p|\alpha) over KK categories parameterized by concentration parameters α=[α1,…,αK]T∈R+K\alpha = [\alpha_1, \dots, \alpha_K]^T \in \mathbb{R}_+^K with Dirichlet precision α0=∑k=1Kαk\alpha_0 = \sum_{k=1}^K \alpha_k, the Fisher Information Matrix (FIM) I(α)∈RK×K\mathcal{I}(\alpha) \in \mathbb{R}^{K \times K} with respect to α\alpha is defined by I(α)=EDir(p∣α)[−∂2log⁡Dir(p∣α)∂α∂αT]\mathcal{I}(\alpha) = \mathbb{E}_{\mathrm{Dir}(p|\alpha)} \left[ -\frac{\partial^2 \log \mathrm{Dir}(p|\alpha)}{\partial \alpha \partial \alpha^T} \right].

    The entries of the FIM are given in closed form by:

    [I(α)]ij={ψ(1)(αi)−ψ(1)(α0),if i=j−ψ(1)(α0),if i≠j[\mathcal{I}(\alpha)]_{ij} = \begin{cases} \psi^{(1)}(\alpha_i) - \psi^{(1)}(\alpha_0), & \text{if } i = j \\ -\psi^{(1)}(\alpha_0), & \text{if } i \neq j \end{cases}

    where ψ(1)(x)=d2dx2ln⁡Γ(x)\psi^{(1)}(x) = \frac{d^2}{dx^2} \ln \Gamma(x) is the trigamma function. In matrix form:

    I(α)=diag([ψ(1)(α1),…,ψ(1)(αK)])−ψ(1)(α0)11T\mathcal{I}(\alpha) = \mathrm{diag}\left([\psi^{(1)}(\alpha_1), \dots, \psi^{(1)}(\alpha_K)]\right) - \psi^{(1)}(\alpha_0)\mathbf{1}\mathbf{1}^T

    where 1=[1,…,1]T∈RK\mathbf{1} = [1, \dots, 1]^T \in \mathbb{R}^K.

    By the Matrix-Determinant Lemma, the log-determinant of the FIM is:

    log⁡∣I(α)∣=∑i=1Klog⁡ψ(1)(αi)+log⁡(1−∑i=1Kψ(1)(α0)ψ(1)(αi))\log |\mathcal{I}(\alpha)| = \sum_{i=1}^K \log \psi^{(1)}(\alpha_i) + \log\left( 1 - \sum_{i=1}^K \frac{\psi^{(1)}(\alpha_0)}{\psi^{(1)}(\alpha_i)} \right)

  3. Knowl 3 — Objective Function of $\mathcal{I}$-EDL

    equation

    For a dataset D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N of NN samples with inputs xi∈Rdx_i \in \mathbb{R}^d and one-hot ground-truth labels yi∈{0,1}Ky_i \in \{0, 1\}^K, let αi=fθ(xi)+1∈R+K\alpha_i = f_\theta(x_i) + \mathbf{1} \in \mathbb{R}_+^K denote the predicted Dirichlet parameters and αi0=∑j=1Kαij\alpha_{i0} = \sum_{j=1}^K \alpha_{ij} the Dirichlet precision. The optimization objective of I\mathcal{I}-EDL is:

    min⁡θ1N∑i=1N(LiI-MSE−λ1Li∣I∣+λ2LiKL)\min_\theta \frac{1}{N} \sum_{i=1}^N \left( \mathcal{L}_i^{\mathcal{I}\text{-MSE}} - \lambda_1 \mathcal{L}_i^{|\mathcal{I}|} + \lambda_2 \mathcal{L}_i^{\text{KL}} \right)

    where λ1≥0,λ2≥0\lambda_1 \ge 0, \lambda_2 \ge 0 are weighting hyperparameters, and the individual components are:

    1. FIM-Weighted Mean Squared Error: LiI-MSE=∑j=1K((yij−αijαi0)2+αij(αi0−αij)αi02(αi0+1))ψ(1)(αij)\mathcal{L}_i^{\mathcal{I}\text{-MSE}} = \sum_{j=1}^K \left( \left(y_{ij} - \frac{\alpha_{ij}}{\alpha_{i0}}\right)^2 + \frac{\alpha_{ij}(\alpha_{i0} - \alpha_{ij})}{\alpha_{i0}^2(\alpha_{i0} + 1)} \right) \psi^{(1)}(\alpha_{ij}) where ψ(1)\psi^{(1)} is the trigamma function.

    2. Negative Log-Determinant of FIM Penalty: Li∣I∣=∑j=1Klog⁡ψ(1)(αij)+log⁡(1−∑j=1Kψ(1)(αi0)ψ(1)(αij))\mathcal{L}_i^{|\mathcal{I}|} = \sum_{j=1}^K \log \psi^{(1)}(\alpha_{ij}) + \log \left( 1 - \sum_{j=1}^K \frac{\psi^{(1)}(\alpha_{i0})}{\psi^{(1)}(\alpha_{ij})} \right)

    3. PAC-Bayesian Kullback-Leibler Regularizer: LiKL=log⁡Γ(∑j=1Kα^ij)−log⁡Γ(K)−∑j=1Klog⁡Γ(α^ij)+∑j=1K(α^ij−1)[ψ(α^ij)−ψ(∑k=1Kα^ik)]\mathcal{L}_i^{\text{KL}} = \log \Gamma\left(\sum_{j=1}^K \hat{\alpha}_{ij}\right) - \log \Gamma(K) - \sum_{j=1}^K \log \Gamma(\hat{\alpha}_{ij}) + \sum_{j=1}^K (\hat{\alpha}_{ij} - 1) \left[ \psi(\hat{\alpha}_{ij}) - \psi\left(\sum_{k=1}^K \hat{\alpha}_{ik}\right) \right] where α^i=αi⊙(1−yi)+yi\hat{\alpha}_i = \alpha_i \odot (1 - y_i) + y_i suppresses the predicted concentration of the true class, Γ(⋅)\Gamma(\cdot) is the gamma function, and ψ(⋅)\psi(\cdot) is the digamma function.

  4. Knowl 4 — PAC-Bayesian Generalization Bound for Evidential Neural Networks

    theoretical result

    Let P\mathcal{P} denote a data distribution over X×Y\mathcal{X} \times \mathcal{Y}, Θ\Theta the hypothesis parameter space, and π\pi a prior distribution over Θ\Theta. For any δ∈(0,1]\delta \in (0, 1] and λ>0\lambda > 0, with probability at least 1−δ1 - \delta over samples D∼Pn\mathcal{D} \sim \mathcal{P}^n, any posterior distribution ρ\rho satisfies:

    Eρ(θ)[L(heta)]≤Eρ(θ)[L^D(θ)]+1λ[DKL(ρ∥π)+log⁡1δ+ΨP,π(λ,n)]\mathbb{E}_{\rho(\theta)}[\mathcal{L}( heta)] \le \mathbb{E}_{\rho(\theta)}[\hat{\mathcal{L}}_{\mathcal{D}}(\theta)] + \frac{1}{\lambda} \left[ D_{\mathrm{KL}}(\rho \parallel \pi) + \log \frac{1}{\delta} + \Psi_{\mathcal{P}, \pi}(\lambda, n) \right]

    where ΨP,π(λ,n)=log⁡Eπ(θ)ED∼Pn[eλ(L(θ)−L^D(θ))]\Psi_{\mathcal{P}, \pi}(\lambda, n) = \log \mathbb{E}_{\pi(\theta)} \mathbb{E}_{\mathcal{D} \sim \mathcal{P}^n} \left[ e^{\lambda(\mathcal{L}(\theta) - \hat{\mathcal{L}}_{\mathcal{D}}(\theta))} \right].

    In evidential neural networks, setting the posterior as the Dirichlet distribution Dir(p∣α)\mathrm{Dir}(p|\alpha) and the prior as Dir(p∣μ)\mathrm{Dir}(p|\mu) (where μk=β≫1\mu_k = \beta \gg 1 for the ground-truth class and 11 for all other classes) justifies minimizing the empirical FIM loss together with the Dirichlet KL-divergence term DKL(Dir(p∣α^)∥Dir(p∣1))D_{\mathrm{KL}}(\mathrm{Dir}(p|\hat{\alpha}) \parallel \mathrm{Dir}(p|\mathbf{1})), where α^=α⊙(1−y)+y\hat{\alpha} = \alpha \odot (1 - y) + y.

  5. Knowl 5 — $\mathcal{I}$-Evidential Deep Learning Training Procedure

    algorithm

    The training procedure for Fisher Information-based Evidential Deep Learning (I\mathcal{I}-EDL) optimizes network parameters θ\theta using mini-batch gradient descent on the composite FIM loss with linear warm-up on the PAC-Bayes KL term.

    Input: Training set D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N, batch size bb, learning rate βlr\beta_{\mathrm{lr}}, total epochs TT, hyperparameter λ\lambda
    Output: Trained model parameters θ\theta
    Initialize network parameters θ\theta
    for t=0,1,…,Tt = 0, 1, \dots, T do
        λt←min⁡(1.0,t/T)\lambda_t \leftarrow \min(1.0, t / T)
        for mini-batch Db∼D\mathcal{D}_b \sim \mathcal{D} do
            for each sample (xi,yi)∈Db(x_i, y_i) \in \mathcal{D}_b do
                αi←fθ(xi)+1\alpha_i \leftarrow f_\theta(x_i) + 1
                α^i←αi⊙(1−yi)+yi\hat{\alpha}_i \leftarrow \alpha_i \odot (1 - y_i) + y_i
                LiI-MSE←∑j=1K((yij−αijαi0)2+αij(αi0−αij)αi02(αi0+1))ψ(1)(αij)\mathcal{L}_i^{\mathcal{I}\text{-MSE}} \leftarrow \sum_{j=1}^K \left( (y_{ij} - \frac{\alpha_{ij}}{\alpha_{i0}})^2 + \frac{\alpha_{ij}(\alpha_{i0} - \alpha_{ij})}{\alpha_{i0}^2(\alpha_{i0} + 1)} \right) \psi^{(1)}(\alpha_{ij})
                Li∣I∣←∑j=1Klog⁡ψ(1)(αij)+log⁡(1−∑j=1Kψ(1)(αi0)ψ(1)(αij))\mathcal{L}_i^{|\mathcal{I}|} \leftarrow \sum_{j=1}^K \log \psi^{(1)}(\alpha_{ij}) + \log \left( 1 - \sum_{j=1}^K \frac{\psi^{(1)}(\alpha_{i0})}{\psi^{(1)}(\alpha_{ij})} \right)
                LiKL←log⁡Γ(∑j=1Kα^ij)−log⁡Γ(K)−∑j=1Klog⁡Γ(α^ij)+∑j=1K(α^ij−1)[ψ(α^ij)−ψ(∑k=1Kα^ik)]\mathcal{L}_i^{\text{KL}} \leftarrow \log \Gamma(\sum_{j=1}^K \hat{\alpha}_{ij}) - \log \Gamma(K) - \sum_{j=1}^K \log \Gamma(\hat{\alpha}_{ij}) + \sum_{j=1}^K (\hat{\alpha}_{ij} - 1) [\psi(\hat{\alpha}_{ij}) - \psi(\sum_{k=1}^K \hat{\alpha}_{ik})]
                Li←LiI-MSE−λLi∣I∣+λtLiKL\mathcal{L}_i \leftarrow \mathcal{L}_i^{\mathcal{I}\text{-MSE}} - \lambda \mathcal{L}_i^{|\mathcal{I}|} + \lambda_t \mathcal{L}_i^{\text{KL}}
            end for
            L←1b∑i=1bLi\mathcal{L} \leftarrow \frac{1}{b} \sum_{i=1}^b \mathcal{L}_i
            θ←θ−βlr∇θL\theta \leftarrow \theta - \beta_{\mathrm{lr}} \nabla_\theta \mathcal{L}
        end for
    end for
    return θ\theta
  6. Knowl 6 — Misclassification Detection and Classification Accuracy on CIFAR-10

    data/table

    Evaluation of CIFAR-10 classification using a VGG-16 architecture compares I\mathcal{I}-EDL with MC Dropout and Dirichlet-based uncertainty models: Prior Networks trained with KL (KL-PN) and reverse KL (RKL-PN), Posterior Network (PostN), and standard Evidential Deep Learning (EDL). Confidence evaluation is measured via Area Under the Precision-Recall Curve (AUPR, %) for detecting correct predictions using maximum probability (Max.P=max⁡cpc\mathrm{Max}.P = \max_c p_c) or maximum Dirichlet parameter (Max.α=max⁡cαc\mathrm{Max}.\alpha = \max_c \alpha_c).

    Method Max.P AUPR (%) Max.α\alpha AUPR (%) Accuracy (%)
    MC Dropout 97.15±0.097.15 \pm 0.0 - 82.84±0.182.84 \pm 0.1
    KL-PN 50.61±4.050.61 \pm 4.0 52.49±4.252.49 \pm 4.2 27.46±1.727.46 \pm 1.7
    RKL-PN 86.11±0.486.11 \pm 0.4 85.59±0.385.59 \pm 0.3 64.76±0.364.76 \pm 0.3
    PostN 97.76±0.097.76 \pm 0.0 97.25±0.097.25 \pm 0.0 84.85±0.084.85 \pm 0.0
    EDL 97.86±0.297.86 \pm 0.2 97.86±0.297.86 \pm 0.2 83.55±0.683.55 \pm 0.6
    I\mathcal{I}-EDL 98.72±0.1\mathbf{98.72 \pm 0.1} 98.63±0.1\mathbf{98.63 \pm 0.1} 89.20±0.3\mathbf{89.20 \pm 0.3}

    I\mathcal{I}-EDL outperforms all baselines in top-1 accuracy (89.20%89.20\%, +5.65%+5.65\% over EDL) and achieves state-of-the-art confidence estimation AUPR (98.72%98.72\% for Max.P and 98.63%98.63\% for Max.α\alpha).

  7. Knowl 7 — Out-of-Distribution Detection Benchmark Performance

    data/table

    Out-of-Distribution (OOD) detection performance is measured by AUPR (%) where in-distribution (ID) data receives label 1 and OOD data receives label 0. Models are trained on ID data without exposure to real OOD samples. Evaluated uncertainty metrics are Max.P=max⁡cpc\mathrm{Max}.P = \max_c p_c and Dirichlet precision α0=∑c=1Kαc\alpha_0 = \sum_{c=1}^K \alpha_c, averaged over 5 runs.

    MNIST →\rightarrow KMNIST MNIST →\rightarrow FMNIST CIFAR10 →\rightarrow SVHN CIFAR10 →\rightarrow CIFAR100
    Method Max.P α0\alpha_0 Max.P α0\alpha_0 Max.P α0\alpha_0 Max.P α0\alpha_0
    MC Dropout 94.00±0.194.00 \pm 0.1 - 96.56±0.296.56 \pm 0.2 - 51.39±0.151.39 \pm 0.1 - 45.57±1.045.57 \pm 1.0 -
    KL-PN 92.97±1.292.97 \pm 1.2 93.39±1.093.39 \pm 1.0 98.44±0.198.44 \pm 0.1 98.16±0.098.16 \pm 0.0 43.96±1.943.96 \pm 1.9 43.23±2.343.23 \pm 2.3 61.41±2.861.41 \pm 2.8 61.53±3.461.53 \pm 3.4
    RKL-PN 60.76±2.960.76 \pm 2.9 53.76±3.453.76 \pm 3.4 78.45±3.178.45 \pm 3.1 72.18±3.672.18 \pm 3.6 53.61±1.153.61 \pm 1.1 49.37±0.849.37 \pm 0.8 55.42±2.655.42 \pm 2.6 54.74±2.854.74 \pm 2.8
    PostN 95.75±0.295.75 \pm 0.2 94.59±0.394.59 \pm 0.3 97.78±0.297.78 \pm 0.2 97.24±0.397.24 \pm 0.3 80.21±0.280.21 \pm 0.2 77.71±0.377.71 \pm 0.3 81.96±0.881.96 \pm 0.8 82.06±0.882.06 \pm 0.8
    EDL 97.02±0.897.02 \pm 0.8 96.31±2.096.31 \pm 2.0 98.10±0.498.10 \pm 0.4 98.08±0.498.08 \pm 0.4 78.87±3.578.87 \pm 3.5 79.12±3.779.12 \pm 3.7 84.30±0.784.30 \pm 0.7 84.18±0.784.18 \pm 0.7
    I\mathcal{I}-EDL 98.34±0.2\mathbf{98.34 \pm 0.2} 98.33±0.2\mathbf{98.33 \pm 0.2} 98.89±0.3\mathbf{98.89 \pm 0.3} 98.86±0.3\mathbf{98.86 \pm 0.3} 83.26±2.4\mathbf{83.26 \pm 2.4} 82.96±2.2\mathbf{82.96 \pm 2.2} 85.35±0.7\mathbf{85.35 \pm 0.7} 84.84±0.6\mathbf{84.84 \pm 0.6}

    I\mathcal{I}-EDL consistently outperforms standard EDL and all other DBU and BNN baselines across all four OOD detection benchmarks in both Max.P\mathrm{Max}.P and α0\alpha_0 scores.

  8. Knowl 8 — Few-Shot Classification, Confidence Evaluation, and OOD Detection on mini-ImageNet

    data/table

    Under NN-way KK-shot episodic meta-learning on mini-ImageNet with a pre-trained WideResNet-28-10 feature extractor and a 1-layer evidential classifier, I\mathcal{I}-EDL is compared against classical EDL across 10,000 episodes. OOD detection is evaluated using query samples from the Caltech-UCSD Birds (CUB) dataset. Top-1 accuracy (%), confidence evaluation AUPR (%) via Max.α\mathrm{Max}.\alpha, and OOD detection AUPR (%) via α0\alpha_0 are reported with 95% confidence intervals.

    5-Way 1-Shot 5-Way 5-Shot 5-Way 20-Shot
    Method Acc. Conf. OOD Acc. Conf. OOD Acc. Conf. OOD
    EDL 61.0061.00 80.5980.59 65.4065.40 80.3880.38 93.9293.92 76.5376.53 85.5485.54 97.5197.51 79.7879.78
    I\mathcal{I}-EDL 63.82\mathbf{63.82} 82.00\mathbf{82.00} 74.76\mathbf{74.76} 82.00\mathbf{82.00} 94.09\mathbf{94.09} 82.48\mathbf{82.48} 88.12\mathbf{88.12} 97.54\mathbf{97.54} 85.40\mathbf{85.40}
    Δ\Delta +2.82+2.82 +1.41+1.41 +9.36+9.36 +1.62+1.62 +0.17+0.17 +5.95+5.95 +2.58+2.58 +0.04+0.04 +5.62+5.62
    10-Way 1-Shot 10-Way 5-Shot 10-Way 20-Shot
    Method Acc. Conf. OOD Acc. Conf. OOD Acc. Conf. OOD
    EDL 44.5544.55 65.9765.97 67.8367.83 62.5262.52 86.8186.81 76.3476.34 69.2969.29 94.2194.21 76.8876.88
    I\mathcal{I}-EDL 49.37\mathbf{49.37} 68.29\mathbf{68.29} 71.95\mathbf{71.95} 67.89\mathbf{67.89} 87.45\mathbf{87.45} 82.29\mathbf{82.29} 78.60\mathbf{78.60} 94.40\mathbf{94.40} 82.52\mathbf{82.52}
    Δ\Delta +4.82+4.82 +2.32+2.32 +4.12+4.12 +5.37+5.37 +0.64+0.64 +5.95+5.95 +9.31+9.31 +0.19+0.19 +5.64+5.64

    I\mathcal{I}-EDL yields consistent gains over EDL across all few-shot configurations, improving classification accuracy by up to +9.31%+9.31\% (10-way 20-shot), confidence evaluation by up to +2.32%+2.32\% (10-way 1-shot), and OOD detection by up to +9.36%+9.36\% (5-way 1-shot).

  9. Knowl 9 — Ablation of FIM-Weighted MSE and FIM Log-Determinant Penalty

    data/table

    An ablation study on mini-ImageNet under 5-way 5-shot classification evaluates the separate and combined contributions of the FIM-weighted mean squared error (I\mathcal{I}-MSE) and the negative log-determinant penalty (∣I∣|\mathcal{I}|). Standard EDL corresponds to the absence of both terms.

    I\mathcal{I}-MSE ∣I∣|\mathcal{I}| Accuracy (%) Conf. Max.α\alpha AUPR (%) OOD α0\alpha_0 AUPR (%)
    80.38±0.1580.38 \pm 0.15 93.92±0.0993.92 \pm 0.09 76.53±0.2776.53 \pm 0.27
    ✓ 81.82±0.1481.82 \pm 0.14 93.97±0.0993.97 \pm 0.09 79.68±0.2579.68 \pm 0.25
    ✓ 81.27±0.1481.27 \pm 0.14 94.42±0.0894.42 \pm 0.08 81.75±0.2281.75 \pm 0.22
    ✓ ✓ 82.00±0.14\mathbf{82.00 \pm 0.14} 94.09±0.09\mathbf{94.09 \pm 0.09} 82.48±0.20\mathbf{82.48 \pm 0.20}

    Individually, I\mathcal{I}-MSE improves accuracy by +1.44%+1.44\% and OOD AUPR by +3.15%+3.15\%, while ∣I∣|\mathcal{I}| improves accuracy by +0.89%+0.89\% and OOD AUPR by +5.22%+5.22\%. Combining both objectives produces the best performance across accuracy (82.00%82.00\%) and OOD detection (82.48%82.48\%).

  10. Knowl 10 — Inapplicability of Dirichlet-Based $\mathcal{I}$-EDL to Continuous Regression

    limitation

    Because I\mathcal{I}-EDL is formulated using the Dirichlet distribution over the categorical probability simplex ΔK−1\Delta^{K-1}, it is intrinsically designed for discrete classification tasks and cannot be directly applied to continuous regression problems. While continuous regression targets can be discretized into ordinal bins to fit a Dirichlet distribution, such discretization loses metric information, ordering, and continuous distance relationships among real-valued regression labels.

Coverage note — OOD detection curves and tiered-ImageNet plots (Figures 3, 5, 6, 7, 8) and extended 50-shot tables (Table 10) were omitted as they reinforce the exact numeric results and conclusions already captured in the included tables.

References

  1. 1.Ablain, M., Meyssignac, B., Zawadzki, L., Jugier, R., Ribes, A., Spada, G., Benveniste, J., Cazenave, A., and Picot, N. Uncertainty in satellite estimates of global mean sea-level changes, trend and acceleration. Earth System Science Data, 11(3):1189–1202, 2019.
  2. 2.Alquier, P., Ridgway, J., and Chopin, N. On the properties of variational approximations of gibbs posteriors. The Journal of Machine Learning Research, 17(1):8374–8414, 2016.
  3. 3.Amini, A., Schwarting, W., Soleimany, A., and Rus, D. Deep evidential regression. Advances in Neural Information Processing Systems, 33:14927–14937, 2020.
  4. 4.Bao, W., Yu, Q., and Kong, Y. Evidential deep learning for open set action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13349–13358, 2021.
  5. 5.Bao, W., Yu, Q., and Kong, Y. Opental: Towards open set temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2979–2989, 2022.
  6. 6.Bengs, V., Hüllermeier, E., and Waegeman, W. Pitfalls of epistemic uncertainty quantification through loss minimisation. In Advances in Neural Information Processing Systems, 2022.
  7. 7.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight uncertainty in neural network. In International conference on machine learning, pp. 1613–1622. PMLR, 2015.
  8. 8.Charpentier, B., Zügner, D., and Günnemann, S. Posterior network: Uncertainty estimation without ood samples via density-based pseudo-counts. Advances in Neural Information Processing Systems, 33:1356–1367, 2020.
  9. 9.Choi, J., Chun, D., Kim, H., and Lee, H.-J. Gaussian yolov3: An accurate and fast object detector using localization uncertainty for autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 502–511, 2019.
  10. 10.Clanuwat, T., Bober-Irizar, M., Kitamoto, A., Lamb, A., Yamamoto, K., and Ha, D. Deep learning for classical japanese literature. arXiv preprint arXiv:1812.01718, 2018.
  11. 11.Collier, M., Mustafa, B., Kokiopoulou, E., Jenatton, R., and Berent, J. Correlated input-dependent label noise in large-scale image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1551–1560, 2021.
  12. 12.Cui, P., Yue, Y., Deng, Z., and Zhu, J. Confidence-based reliable learning under dual noises. Advances in Neural Information Processing Systems, 35:35116–35129, 2022.
  13. 13.Dempster, A. P. A generalization of bayesian inference. Journal of the Royal Statistical Society: Series B (Methodological), 30(2):205–232, 1968.
  14. 14.Feng, D., Rosenbaum, L., and Dietmayer, K. Towards safe autonomous driving: Capture uncertainty in the deep neural network for lidar 3d vehicle detection. In 2018 21st international conference on intelligent transportation systems (ITSC), pp. 3266–3273. IEEE, 2018.
  15. 15.Gal, Y. Uncertainty in Deep Learning. PhD thesis, University of Cambridge, 2016.
  16. 16.Gal, Y. and Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning, pp. 1050–1059. PMLR, 2016.
  17. 17.Gal, Y., Islam, R., and Ghahramani, Z. Deep bayesian active learning with image data. In International Conference on Machine Learning, pp. 1183–1192. PMLR, 2017.
  18. 18.Germain, P., Lacasse, A., Laviolette, F., and Marchand, M. Pac-bayesian learning of linear classifiers. In Proceedings of the 26th Annual International Conference on Machine Learning, pp. 353–360, 2009.
  19. 19.Ghaffari, S., Saleh, E., Forsyth, D., and Wang, Y.-X. On the importance of firth bias reduction in few-shot classification. arXiv preprint arXiv:2110.02529, 2021.
  20. 20.Graves, A. Practical variational inference for neural networks. Advances in neural information processing systems, 24, 2011.
  21. 21.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International conference on machine learning, pp. 1321–1330. PMLR, 2017.
  22. 22.Hemmer, P., Kühl, N., and Schöffer, J. Deal: deep evidential active learning for image classification. In Deep Learning Applications, Volume 3, pp. 171–192. Springer, 2022.
  23. 23.Jøsang, A. Artificial reasoning with subjective logic. In Proceedings of the second Australian workshop on commonsense reasoning, volume 48, pp. 34. Citeseer, 1997.
  24. 24.Jøsang, A. Subjective logic, volume 3. Springer, 2016.
  25. 25.Karandikar, A., Cain, N., Tran, D., Lakshminarayanan, B., Shlens, J., Mozer, M. C., and Roelofs, B. Soft calibration objectives for neural networks. Advances in Neural Information Processing Systems, 34:29768–29779, 2021.
  26. 26.Kristiadi, A., Hein, M., and Hennig, P. Learnable uncertainty under laplace approximations. In Uncertainty in Artificial Intelligence, pp. 344–353. PMLR, 2021.
  27. 27.Kristiadi, A., Hein, M., and Hennig, P. Being a bit frequentist improves bayesian neural networks. In International Conference on Artificial Intelligence and Statistics, pp. 529–545. PMLR, 2022.
  28. 28.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. cs.toronto.edu, 2009.
  29. 29.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems, 30, 2017.
  30. 30.LeCun, Y. The mnist database of handwritten digits. http://yann. lecun. com/exdb/mnist/, 1998.
  31. 31.Lehmann, E. L. and Casella, G. Theory of point estimation. Springer Science & Business Media, 2006.
  32. 32.Liang, S., Li, Y., and Srikant, R. Enhancing the reliability of out-of-distribution image detection in neural networks. In International Conference on Learning Representations, 2018.
  33. 33.Ma, Y.-A., Chen, T., and Fox, E. A complete recipe for stochastic gradient mcmc. Advances in neural information processing systems, 28, 2015.
  34. 34.Malinin, A. and Gales, M. Predictive uncertainty estimation via prior networks. Advances in neural information processing systems, 31, 2018.
  35. 35.Malinin, A. and Gales, M. Reverse kl-divergence training of prior networks: Improved uncertainty and adversarial robustness. Advances in Neural Information Processing Systems, 32, 2019.
  36. 36.Masegosa, A. Learning under model misspecification: Applications to variational and ensemble methods. Advances in Neural Information Processing Systems, 33: 5479–5491, 2020.
  37. 37.McAllester, D. A. Some pac-bayesian theorems. Machine Learning, 37(3):355–363, 1999.
  38. 38.Meinert, N., Gawlikowski, J., and Lavin, A. The unreasonable effectiveness of deep evidential regression. arXiv preprint arXiv:2205.10060, 2022.
  39. 39.Nair, T., Precup, D., Arnold, D. L., and Arbel, T. Exploring uncertainty measures in deep networks for multiple sclerosis lesion detection and segmentation. Medical image analysis, 59:101557, 2020.
  40. 40.Nandy, J., Hsu, W., and Lee, M. L. Towards maximizing the representation gap between in-domain & out-of-distribution examples. Advances in Neural Information Processing Systems, 33:9239–9250, 2020.
  41. 41.Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. The street view house numbers (svhn) dataset. Technical report, Technical report, Accessed 2016-08-01.[Online], 2018.
  42. 42.Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., and Snoek, J. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in neural information processing systems, 32, 2019.
  43. 43.Pandey, D. S. and Yu, Q. Multidimensional belief quantification for label-efficient meta-learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14391–14400, 2022.
  44. 44.Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G. Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548, 2017.
  45. 45.Quinonero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D. Dataset shift in machine learning. Mit Press, 2008.
  46. 46.Ren, M., Triantafillou, E., Ravi, S., Snell, J., Swersky, K., Tenenbaum, J. B., Larochelle, H., and Zemel, R. S. Meta-learning for semi-supervised few-shot classification. arXiv preprint arXiv:1803.00676, 2018.
  47. 47.Ritter, H., Botev, A., and Barber, D. A scalable laplace approximation for neural networks. In 6th International Conference on Learning Representations, ICLR 2018-Conference Track Proceedings, volume 6. International Conference on Representation Learning, 2018.
  48. 48.Roelofs, R., Cain, N., Shlens, J., and Mozer, M. C. Mitigating bias in calibration error estimation. In International Conference on Artificial Intelligence and Statistics, pp. 4036–4054. PMLR, 2022.
  49. 49.Sastry, C. S. and Oore, S. Detecting out-of-distribution examples with gram matrices. In International Conference on Machine Learning, pp. 8491–8501. PMLR, 2020.
  50. 50.Seeböck, P., Orlando, J. I., Schlegl, T., Waldstein, S. M., Bogunović, H., Klimscha, S., Langs, G., and Schmidt-Erfurth, U. Exploiting epistemic uncertainty of anatomy segmentation for anomaly detection in retinal oct. IEEE transactions on medical imaging, 39(1):87–98, 2019.
  51. 51.Sensoy, M., Kaplan, L., and Kandemir, M. Evidential deep learning to quantify classification uncertainty. Advances in Neural Information Processing Systems, 31, 2018.
  52. 52.Sentz, K. and Ferson, S. Combination of evidence in dempster-shafer theory. US Department of Energy, 2002.
  53. 53.Shafer, G. A mathematical theory of evidence, volume 42. Princeton university press, 1976.
  54. 54.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  55. 55.Soleimany, A. P., Amini, A., Goldman, S., Rus, D., Bhatia, S. N., and Coley, C. W. Evidential deep learning for guided molecular property prediction and discovery. ACS central science, 7(8):1356–1367, 2021.
  56. 56.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818–2826, 2016.
  57. 57.Székely, G. J. and Rizzo, M. L. Energy statistics: A class of statistics based on distances. Journal of statistical planning and inference, 143(8):1249–1272, 2013.
  58. 58.Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al. Matching networks for one shot learning. Advances in neural information processing systems, 29, 2016.
  59. 59.Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. The caltech-ucsd birds-200-2011 dataset. caltech.edu, 2011.
  60. 60.Welling, M. and Teh, Y. W. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th international conference on machine learning (ICML-11), pp. 681–688. Citeseer, 2011.
  61. 61.Wilson, A. G. and Izmailov, P. Bayesian deep learning and a probabilistic perspective of generalization. Advances in neural information processing systems, 33:4697–4708, 2020.
  62. 62.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  63. 63.Yang, S., Liu, L., and Xu, M. Free lunch for few-shot learning: Distribution calibration. arXiv preprint arXiv:2101.06395, 2021.
  64. 64.Zaidi, S., Zela, A., Elsken, T., Holmes, C. C., Hutter, F., and Teh, Y. Neural ensemble search for uncertainty estimation and dataset shift. Advances in Neural Information Processing Systems, 34:7898–7911, 2021.
  65. 65.Zhao, X., Chen, F., Hu, S., and Cho, J.-H. Uncertainty aware semi-supervised learning on graph data. Advances in Neural Information Processing Systems, 33:12827–12836, 2020.

Citation

MLA
Deng, D., et al. “Uncertainty Estimation by Fisher Information-based Evidential Deep Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 7596–616, https://proceedings.mlr.press/v202/deng23b.html.
APA
Deng, D., Chen, G., Yu, Y., Liu, F., & Heng, P.-A. (2023). Uncertainty Estimation by Fisher Information-based Evidential Deep Learning. International Conference on Machine Learning, 202, 7596–7616. https://proceedings.mlr.press/v202/deng23b.html
Chicago
Deng, D., G. Chen, Y. Yu, F. Liu, and P.-A. Heng. 2023. “Uncertainty Estimation by Fisher Information-based Evidential Deep Learning”. International Conference on Machine Learning 202: 7596–7616. https://proceedings.mlr.press/v202/deng23b.html.
Harvard
Deng, D. et al. (2023) “Uncertainty Estimation by Fisher Information-based Evidential Deep Learning”, International Conference on Machine Learning. PMLR, pp. 7596–7616. Available at: https://proceedings.mlr.press/v202/deng23b.html.
Vancouver
1. Deng D, Chen G, Yu Y, Liu F, Heng P-A (2023) Uncertainty Estimation by Fisher Information-based Evidential Deep Learning. In: International Conference on Machine Learning. PMLR, pp 7596–7616

BibTeX

@InProceedings{pmlr-v202-deng23b,
  title = 	 {Uncertainty Estimation by {F}isher Information-based Evidential Deep Learning},
  author =       {Deng, Danruo and Chen, Guangyong and Yu, Yang and Liu, Furui and Heng, Pheng-Ann},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {7596--7616},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/deng23b/deng23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/deng23b.html},
  abstract = 	 {Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network’s outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning ($\mathcal{I}$-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/