A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness

Jeremiah Zhe LiuShreyas PadhyJie RenZi LinYeming WenGhassen JerfelZachary NadoJasper SnoekDustin TranBalaji Lakshminarayanan

article2023JMLR92 citations

Proposes Spectral-normalized Neural Gaussian Process (SNGP), a method that combines spectral normalization with a Gaussian process output layer to achieve distance-aware, high-quality uncertainty estimation in a single deterministic network without the computational overhead of deep ensembles.

Listen

Deep neural networks are increasingly deployed in safety-critical applications such as healthcare, genomics, and autonomous systems. However, standard networks frequently exhibit overconfidence when making incorrect predictions or encountering inputs outside their training data. Existing solutions, such as Bayesian neural networks and deep ensembles, combine predictions across multiple parameter sets or models. While effective, these multi-model techniques require substantial memory and multiple computational passes, making them impractical for real-time, resource-constrained environments.

The article develops and evaluates a principled approach to improve uncertainty quantification in a single, deterministic deep neural network using distance awareness—the capability to measure how far an unseen test input lies from the training data manifold.

To establish a theoretical foundation, the authors formulate uncertainty estimation as a minimax decision problem, proving that distance awareness is a mathematically necessary condition for high-quality uncertainty estimates. To implement this property, the article introduces the Spectral-normalized Neural Gaussian Process (SNGP). This framework modifies modern residual networks via two straightforward adjustments: applying spectral normalization to the hidden layers to ensure smooth representations that avoid representation collapse, and replacing the final classification layer with a distance-aware Gaussian process layer approximated via random features and Laplace approximation. The authors evaluate SNGP across synthetic benchmarks and complex real-world datasets across three modalities: computer vision (CIFAR-10, CIFAR-100, and ImageNet using Wide-ResNet and ResNet-50 architectures), natural language intent detection (CLINC out-of-scope dataset with BERT), and genomics (a 1D convolutional network identifying bacterial sequences).

The evaluation reveals several key findings. First, SNGP significantly improves calibration and out-of-distribution detection over baseline neural networks and competing single-model methods without sacrificing standard predictive accuracy. For instance, on CIFAR-100, SNGP cuts expected calibration error on corrupted data from approximately 0.258 down to 0.060—an improvement of over 75%—while raising out-of-distribution detection area under the curve from roughly 0.799 to 0.846. Second, SNGP scales smoothly to large tasks like ImageNet, where single-model alternatives struggle, reducing corrupted calibration error from 0.103 to 0.045 while maintaining 76.1% clean classification accuracy. Third, SNGP demonstrates strong cross-domain generalization, achieving top single-model out-of-scope detection on BERT (0.969 AUROC) and genomics sequence identification. Finally, SNGP acts as an orthogonal building block that stacks effectively with ensemble methods and data augmentation pipelines (such as AugMix), with an ensemble of SNGP models achieving the best overall accuracy and uncertainty scores across benchmarks.

These findings indicate that organizations can deploy reliable, safety-oriented deep learning models in latency-critical and edge-computing environments without the multi-fold compute and memory overhead of ensemble models. By replacing standard output layers and normalizing residual weights, systems gain the ability to recognize unfamiliar scenarios and avoid confident, high-risk failures.

For practical implementation, engineering teams can adopt SNGP as a drop-in replacement for standard classification architectures where fast single-pass inference is required. In unconstrained computing environments where maximum reliability is necessary, practitioners should combine SNGP base models with ensembling and domain-specific data augmentation to compound predictive and uncertainty gains.

The approach relies on approximations to retain computational efficiency, including random feature expansions and Laplace posterior estimates, and its representation guarantees depend on residual network architectures. Nevertheless, extensive empirical validation across multiple domains provides high confidence in SNGP as an efficient and practical standard for single-model uncertainty estimation.

Liu et al (2023).pdf
Cover for A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness

Abstract

Uncertainty quantification is a key component of modern machine learning systems. A popular approach for quantifying predictive uncertainty on neural network models is deep ensembles. However, training multiple independently initialized networks leads to prohibitive computational costs at both train time and test time. Alternatively, single-model approaches such as MC-dropout and its variants have been proposed but typically underperform relative to their ensemble-based counterparts. In this work we propose a simple method which brings together distance-awareness and spectral normalization for improving robustness to distributional shift while retaining model calibration. We demonstrate significant gains over baseline methods on benchmark tasks including out-of-distribution detection.

Table of Contents

  • 1. Introduction
  • 2. Theoretical Motivation for Distance Awareness
  • 2.1 Uncertainty Estimation as a Minimax Learning Problem
  • 2.2 Distance Awareness as a Necessary Condition
  • 3. Our Proposed Method: Spectral-normalized Neural Gaussian Process (SNGP)
  • 3.1 Distance-aware Output Layer via Laplace-approximated Neural Gaussian Process
  • 3.2 Approximately Distance-preserving Hidden Mapping via Spectral Normalization
  • 4. Where does SNGP fit in the current landscape of methods?
  • 5. Related Work
  • 6. Benchmarking Experiments
  • 6.1 Low-dimensional Experiments
  • 6.1.1 REGRESSION
  • 6.1.2 CLASSIFICATION
  • 6.2 Image Classification
  • 6.2.1 CIFAR-10 AND CIFAR-100
  • 6.2.2 IMAGENET
  • 6.3 Generalization to other data modalities
  • 6.3.1 CONVERSATIONAL LANGUAGE UNDERSTANDING
  • 6.3.2 BACTERIA GENOMICS SEQUENCE IDENTIFICATION
  • 7. SNGP as a Building Block for Probabilistic Deep Learning
  • 7.1 Ensembling Approaches for Hidden Representation Uncertainty
  • 7.2 Data Augmentation for Improved Distance Awareness
  • 8. Conclusions and Discussion
  • 8.1 Further Discussions
  • Acknowledgments
  • References
  • Appendix A. Method Summary
  • A.1 Extension to Regression and Multi-class Classification
  • A.2 Hyperparameter Configuration
  • Appendix B. Formal Statements
  • Appendix C. Experiment Details
  • C.1 Model Configuration
  • C.2 Evaluation
  • C.3 Theoretical Convergence to Optimal Behaviour
  • Appendix D. An Example Formalization of 'Semantic Distance'.
  • Appendix E. Proof
  • E.1 Proof of Proposition 3
  • E.2 Proof of Lemma 7

Knowls

  1. Knowl 1 — Minimax uncertainty requires an in-domain prediction and a conservative OOD prediction

    theoretical result

    Consider classification with KK classes, an in-domain region XIND\mathcal{X}_{\mathrm{IND}}, and an out-of-domain region whose label distribution is otherwise unconstrained. Suppose the model has learned an in-domain predictive distribution q(y∣x,x∈XIND)q(y\mid x, x\in\mathcal{X}_{\mathrm{IND}}) and the true probability of an input being in-domain is known. Under minimax risk for the paper’s strictly proper separable Bregman scoring rules—including the Brier and logarithmic scores—the unique optimal predictive distribution combines the learned in-domain prediction with a uniform distribution over classes for OOD inputs: p(y∣x)=q(y∣x,x∈XIND)p∗(x∈XIND)+puniform(y∣x,x∉XIND)p∗(x∉XIND)p(y\mid x)=q(y\mid x,x\in\mathcal{X}_{\mathrm{IND}})p^*(x\in\mathcal{X}_{\mathrm{IND}})+p_{\mathrm{uniform}}(y\mid x,x\notin\mathcal{X}_{\mathrm{IND}})p^*(x\notin\mathcal{X}_{\mathrm{IND}}), where puniform(y=k)=1/Kp_{\mathrm{uniform}}(y=k)=1/K. The uniform OOD prediction is optimal under the stated worst-case assumption, in which nature may choose any OOD label distribution; the result does not assert that uniform predictions are optimal when OOD classes are semantically related to in-domain classes.

  2. Knowl 2 — Distance awareness is the uncertainty property needed to identify unfamiliar inputs

    definition

    A predictive distribution p(y∣x)p(y\mid x) trained on an in-domain set XIND\mathcal{X}_{\mathrm{IND}} is input-distance-aware if it has an uncertainty summary u(x)u(x)—such as predictive entropy or variance—that is a monotonic function of the distance from xx to the training domain: u(x)=v(d(x,XIND))u(x)=v(d(x,\mathcal{X}_{\mathrm{IND}})), where vv is monotonic and d(x,XIND)=Ex′∼XIND[dX(x,x′)]d(x,\mathcal{X}_{\mathrm{IND}})=\mathbb{E}_{x'\sim\mathcal{X}_{\mathrm{IND}}}[d_{\mathcal{X}}(x,x')] for a suitable metric dXd_{\mathcal{X}} on the data manifold. The minimax result makes this property consequential: a model must estimate whether an input belongs to the training domain to combine its learned in-domain prediction with the conservative OOD prediction. In a deep classifier g∘hg\circ h, the paper identifies two design requirements: the output layer gg must express uncertainty based on hidden-space distance, and the representation hh must preserve meaningful input-space distances.

  3. Knowl 3 — Lipschitz-bounded residual blocks preserve distances through a network

    theoretical result

    Let h=hB∘⋯∘h1h=h_B\circ\cdots\circ h_1 be a composition of BB equal-dimensional residual blocks, with hb(x)=x+gb(x)h_b(x)=x+g_b(x). If every residual mapping gbg_b is α\alpha-Lipschitz for a common 0<α<10<\alpha<1, then for every pair of inputs x,x′x,x' the representation satisfies (1−α)BdX(x,x′)≤∥h(x)−h(x′)∥H≤(1+α)BdX(x,x′)(1-\alpha)^B d_{\mathcal{X}}(x,x')\leq\|h(x)-h(x')\|_{\mathcal{H}}\leq(1+\alpha)^B d_{\mathcal{X}}(x,x'). Thus the representation is bi-Lipschitz: it neither collapses distinct inputs together nor expands distances without bound. For a residual branch of the form gb(x)=a(Wbx+cb)g_b(x)=a(W_bx+c_b) with a 1-Lipschitz activation, bounding the spectral norm of WbW_b bounds the branch’s Lipschitz constant. The strict two-sided guarantee requires the stated residual-branch condition; spectral normalization is the practical mechanism used to control the upper bound.

  4. Knowl 4 — SNGP combines spectrally normalized features with a Laplace-approximated GP output

    model/method

    Spectral-normalized Neural Gaussian Process (SNGP) modifies a neural classifier in two ways: it spectrally normalizes hidden-layer weights, and replaces the dense output layer with a Gaussian-process layer implemented using random Fourier features and a Laplace approximation. For an input xx, let h(x)∈Rdhh(x)\in\mathbb{R}^{d_h} be the penultimate representation. Construct a DD-dimensional random feature vector ϕ(x)=2σ2/D cos⁡(Wh(x)+b)\phi(x)=\sqrt{2\sigma^2/D}\,\cos(Wh(x)+b), where WW and bb are fixed random weights and biases sampled for the RBF-kernel approximation, and σ\sigma controls kernel amplitude. Train output weights β\beta and the ordinary hidden weights by minibatch SGD on the task negative log likelihood plus a quadratic prior penalty. Estimate each hidden weight matrix’s spectral norm with one power iteration and rescale it when necessary to enforce a chosen bound cc. In the final training epoch, accumulate the output-weight posterior precision from training features; for binary classification this is Σ^−1=τI+∑i=1Np^i(1−p^i)ϕiϕi⊤\hat\Sigma^{-1}=\tau I+\sum_{i=1}^{N}\hat p_i(1-\hat p_i)\phi_i\phi_i^\top, where NN is the number of training examples, ϕi=ϕ(xi)\phi_i=\phi(x_i), p^i\hat p_i is the fitted positive-class probability, II is the D×DD\times D identity, and τ\tau is a positive ridge term. The Laplace covariance is Σ^=(Σ^−1)−1\hat\Sigma=(\hat\Sigma^{-1})^{-1}. At prediction time, compute the logit mean m(x)=ϕ(x)⊤βm(x)=\phi(x)^\top\beta and variance v(x)=ϕ(x)⊤Σ^ϕ(x)v(x)=\phi(x)^\top\hat\Sigma\phi(x); for multiclass tasks the paper uses a shared covariance approximation. The mean costs O(D)O(D) and the variance costs O(D2)O(D^2) per input. Experiments used D=1024D=1024 for most models and D=2048D=2048 for BERT, one power iteration, and an RBF length scale of 2.0; the spectral bound was task-dependent, including c=6c=6 for Wide ResNets and c=0.95c=0.95 for the BERT pooler.

  5. Knowl 5 — SNGP approaches a uniform predictive distribution far from training data

    theoretical result

    Assume the feature mapping is distance-preserving and the SNGP output uses an RBF-kernel random-feature approximation with a Laplace posterior. As a test input moves arbitrarily far from the training manifold, its kernel similarities to training examples tend to zero. Consequently, the SNGP predictive logit mean tends to zero while its predictive variance tends to a data-independent prior value. Under the paper’s mean-field approximation to the Gaussian-softmax integral, the resulting class probabilities converge to the uniform distribution over the KK classes; the Monte Carlo softmax approximation likewise approaches the uniform prediction. This is the model’s intended far-OOD behavior under the paper’s assumptions, rather than a claim that every input outside the training set must receive a uniform prediction.

  6. Knowl 6 — Controlled toy tasks show that both SNGP components matter for distance-aware uncertainty

    empirical result

    The paper compared exact Gaussian processes, random-feature Gaussian processes, variational Gaussian processes, and neural classifiers on a one-dimensional bimodal regression task and on two-dimensional two-ovals and two-moons classification tasks. On regression, the random-feature GP reproduced the exact GP’s near-zero mean and high variance far from training data, with periodic artifacts from the Fourier approximation; the tested variational GP’s predictions away from data depended on inducing-point locations. For the classification tasks, the neural models used a 12-layer residual feedforward network with 128 hidden units, trained on 500 examples from each in-domain class; an additional OOD class was held out. Deep ensembles and MC Dropout (10 models or samples, respectively) mainly expressed uncertainty near the decision boundary and could remain confident far from the data. A GP output without spectral normalization also failed to be reliably distance-aware because its hidden representation could collapse distinct inputs. Combining the GP output with spectral normalization produced uncertainty surfaces resembling those of the exact GP, supporting the paper’s claim that a distance-aware output layer alone is insufficient when the representation discards input-distance information.

  7. Knowl 7 — Image benchmarks show improved calibration and OOD detection without sacrificing accuracy

    empirical result

    On CIFAR-10 and CIFAR-100, the authors trained Wide ResNet-28-10 models and evaluated clean and corrupted test accuracy, ECE, NLL, and maximum-softmax-probability OOD AUROC; reported results average 10 seeds. For CIFAR-10, SNGP achieved clean/corrupted accuracy 95.7±0.140/79.3±0.34095.7\pm0.140/79.3\pm0.340, clean/corrupted ECE 0.017±0.003/0.099±0.0080.017\pm0.003/0.099\pm0.008, clean/corrupted NLL 0.149±0.005/0.745±0.0260.149\pm0.005/0.745\pm0.026, and AUROC 0.960±0.0040.960\pm0.004 on SVHN and 0.902±0.0030.902\pm0.003 on CIFAR-100. The deterministic network’s corresponding values were 95.8±0.190/79.0±0.35095.8\pm0.190/79.0\pm0.350 accuracy, 0.028±0.002/0.153±0.0050.028\pm0.002/0.153\pm0.005 ECE, 0.183±0.007/1.042±0.0380.183\pm0.007/1.042\pm0.038 NLL, and 0.946±0.005/0.893±0.0010.946\pm0.005/0.893\pm0.001 AUROC. For CIFAR-100, SNGP achieved accuracy 80.3±0.230/55.3±0.19080.3\pm0.230/55.3\pm0.190, ECE 0.030±0.004/0.060±0.0040.030\pm0.004/0.060\pm0.004, NLL 0.761±0.007/1.919±0.0130.761\pm0.007/1.919\pm0.013, and AUROC 0.846±0.0190.846\pm0.019 on SVHN and 0.798±0.0010.798\pm0.001 on CIFAR-10; the deterministic network achieved 80.4±0.290/55.0±0.18080.4\pm0.290/55.0\pm0.180 accuracy, 0.107±0.004/0.258±0.0040.107\pm0.004/0.258\pm0.004 ECE, 0.941±0.016/2.663±0.0450.941\pm0.016/2.663\pm0.045 NLL, and 0.799±0.020/0.795±0.0010.799\pm0.020/0.795\pm0.001 AUROC. On ImageNet with ResNet-50, SNGP achieved clean/corrupted accuracy 76.1±0.01/41.1±0.0176.1\pm0.01/41.1\pm0.01, ECE 0.013±0.001/0.045±0.0120.013\pm0.001/0.045\pm0.012, and NLL 0.93±0.01/3.03±0.010.93\pm0.01/3.03\pm0.01; the deterministic model achieved 76.2±0.01/40.5±0.0176.2\pm0.01/40.5\pm0.01 accuracy, 0.032±0.002/0.103±0.0110.032\pm0.002/0.103\pm0.011 ECE, and 0.939±0.01/3.21±0.020.939\pm0.01/3.21\pm0.02 NLL. These results show that the GP layer and spectral normalization improved uncertainty metrics while keeping predictive accuracy close to the deterministic baseline, including at ImageNet scale.

  8. Knowl 8 — Language and genomics experiments extend SNGP beyond image classification

    empirical result

    For CLINC out-of-scope intent detection, the authors fine-tuned BERTBase on 150 in-domain intents, with 150 training sentences per intent, and evaluated on in-domain and 1,500 natural OOD utterances. SNGP obtained accuracy 96.6±0.0596.6\pm0.05, ECE 0.014±0.0050.014\pm0.005, NLL 1.218±0.031.218\pm0.03, OOD AUROC 0.969±0.010.969\pm0.01, and AUPR 0.880±0.010.880\pm0.01. The deterministic BERT baseline scored 96.5±0.1196.5\pm0.11, 0.024±0.0020.024\pm0.002, 3.559±0.113.559\pm0.11, 0.897±0.010.897\pm0.01, and 0.757±0.020.757\pm0.02, respectively; the Deep Ensemble had accuracy 97.5±0.0397.5\pm0.03, ECE 0.013±0.0020.013\pm0.002, NLL 1.062±0.021.062\pm0.02, AUROC 0.964±0.010.964\pm0.01, and AUPR 0.862±0.010.862\pm0.01. For bacterial genomic-sequence classification, 10 species were in-domain and 60 species were OOD. SNGP achieved accuracy 85.71±0.10085.71\pm0.100, ECE 0.019±0.0040.019\pm0.004, NLL 0.417±0.0040.417\pm0.004, AUROC 0.672±0.0110.672\pm0.011, and AUPR 0.637±0.0090.637\pm0.009, compared with the deterministic model’s 84.40±0.39084.40\pm0.390, 0.049±0.0070.049\pm0.007, 0.487±0.0070.487\pm0.007, 0.640±0.0050.640\pm0.005, and 0.609±0.0050.609\pm0.005. Thus, in both modalities SNGP improved on the single deterministic baseline in OOD detection and uncertainty metrics; its relative performance against ensembles differed by task.

  9. Knowl 9 — SNGP complements ensembles and data augmentation

    empirical result

    The paper treats representation diversity from ensembling and representation improvements from data augmentation as complementary to SNGP’s distance-awareness. On CIFAR-10, an ensemble of SNGP models achieved clean/corrupted accuracy 96.4±0.040/81.1±0.13096.4\pm0.040/81.1\pm0.130, ECE 0.009±0.001/0.042±0.0020.009\pm0.001/0.042\pm0.002, and AUROC 0.967±0.0020.967\pm0.002 for SVHN and 0.920±0.0010.920\pm0.001 for CIFAR-100; a conventional Deep Ensemble scored 96.4±0.090/80.4±0.15096.4\pm0.090/80.4\pm0.150, 0.011±0.001/0.092±0.0030.011\pm0.001/0.092\pm0.003, and 0.947±0.002/0.914±0.0000.947\pm0.002/0.914\pm0.000 on those metrics. AugMix also improved SNGP’s corrupted-data and OOD performance: on CIFAR-10, SNGP plus AugMix achieved corrupted accuracy 87.1±0.33087.1\pm0.330, corrupted ECE 0.026±0.0050.026\pm0.005, and AUROC 0.976±0.0050.976\pm0.005 for SVHN and 0.925±0.0030.925\pm0.003 for CIFAR-100, versus SNGP alone at 79.3±0.34079.3\pm0.340, 0.099±0.0080.099\pm0.008, and 0.960±0.004/0.902±0.0030.960\pm0.004/0.902\pm0.003. On CIFAR-100, SNGP plus AugMix achieved corrupted accuracy 66.4±0.19066.4\pm0.190 and SVHN AUROC 0.870±0.0240.870\pm0.024, versus 55.3±0.19055.3\pm0.190 and 0.846±0.0190.846\pm0.019 for SNGP alone; corrupted ECE was 0.064±0.0020.064\pm0.002 with AugMix versus 0.060±0.0040.060\pm0.004 without it, so not every metric improved. These experiments support combining distance-aware base models with other uncertainty or representation-learning techniques, while showing that gains depend on task and metric.

  10. Knowl 10 — SNGP’s guarantees and performance depend on approximations and tuning

    limitation

    The distance-preservation argument assumes a suitable input metric and, for the residual-block guarantee, equal-dimensional representations with residual branches strictly below a Lipschitz bound of one. Practical networks may reduce dimension, so exact bi-Lipschitz preservation is generally unavailable; the spectral bound must balance representation smoothness against expressiveness and predictive accuracy. The method also approximates an exact GP using random Fourier features, a Laplace posterior, and—in the reported experiments—a mean-field approximation to the Gaussian-softmax predictive distribution. The authors report that more accurate feature and softmax approximations did not meaningfully improve their preliminary results, but do not rule out gains from a more accurate GP posterior. Kernel amplitude affects calibration and is tuned on held-out in-domain data; the spectral bound is also task-dependent. The paper therefore establishes a practical method and conditional theoretical motivation, not an exact semantic-distance guarantee or universal superiority over more computationally intensive uncertainty methods.

Coverage note — Detailed comparisons among alternative OOD scores, the appendix’s illustrative formalization of semantic distance, and exhaustive hyperparameter and augmentation ablations are omitted because they are secondary analyses rather than load-bearing parts of the method, theory, or principal benchmark findings.

References

  1. 1.Ittai Abraham, Yair Bartal, and Ofer Neiman. Advances in metric embedding theory. Advances in Mathematics, 228(6):3026–3126, 2011.
  2. 2.Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mane.́ Concrete Problems in AI Safety. arXiv:1606.06565 [cs], June 2016.
  3. 3.Senjian An, Farid Boussaid, and Mohammed Bennamoun. How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances? In International Conference on Machine Learning, pages 514–523, June 2015. ISSN: 1938-7228 Section: Machine Learning.
  4. 4.Cem Anil, James Lucas, and Roger Grosse. Sorting Out Lipschitz Function Approximation. In International Conference on Machine Learning, pages 291–301, May 2019. ISSN: 1938-7228 Section: Machine Learning.
  5. 5.Francis Bach. Breaking the Curse of Dimensionality with Convex Neural Networks. Journal of Machine Learning Research, 18(19):1–53, 2017. ISSN 1533-7928.
  6. 6.Peter Bartlett, Steven Evans, and Phil Long. Representing smooth functions as compositions of near-identity functions with implications for deep network optimization. arXiv, 2018.
  7. 7.Jens Behrmann, Will Grathwohl, Ricky T. Q. Chen, David Duvenaud, and Joern-Henrik Jacobsen. Invertible Residual Networks. In International Conference on Machine Learning, pages 573–582, May 2019. ISSN: 1938-7228 Section: Machine Learning.
  8. 8.Jens Behrmann, Paul Vicol, Kuan-Chieh Wang, Roger Grosse, and Joern-Henrik Jacobsen. Understanding and mitigating exploding inverses in invertible neural networks. In International Conference on Artificial Intelligence and Statistics, pages 1792–1800. PMLR, 2021.
  9. 9.Abhijit Bendale and Terrance E. Boult. Towards Open Set Deep Networks. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  10. 10.James O. Berger. Statistical Decision Theory and Bayesian Analysis. Springer Series in Statistics. Springer-Verlag, New York, 2 edition, 1985. ISBN 978-0-387-96098-2.
  11. 11.Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, New York, April 2011. ISBN 978-0-387-31073-2.
  12. 12.Avrim Blum. Random Projection, Margins, Kernels, and Feature-Selection. In Subspace, Latent Structure and Feature Selection, Lecture Notes in Computer Science, pages 52–68, Berlin, Heidelberg, 2006. Springer. ISBN 978-3-540-34138-3. doi: 10.1007/11752790 3.
  13. 13.Charles Blundell, Yee Whye Teh, and Katherine A Heller. Bayesian rose trees. In Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence, pages 65–72, 2010.
  14. 14.John Bradshaw, Alexander G. de G. Matthews, and Zoubin Ghahramani. Adversarial Examples, Uncertainty, and Transfer Testing Robustness in Gaussian Process Hybrid Deep Networks. arXiv:1707.02476 [stat], July 2017. arXiv: 1707.02476.
  15. 15.Jochen Brocker. Reliability, sufficiency, and the decomposition of proper scores. ¨ Quarterly Journal of the Royal Meteorological Society: A journal of the atmospheric sciences, applied meteorology and physical oceanography, 135(643):1512–1519, 2009.
  16. 16.Roberto Calandra, Jan Peters, Carl E. Rasmussen, and Marc Peter Deisenroth. Manifold Gaussian Processes for regression. 2016 International Joint Conference on Neural Networks (IJCNN), 2016. doi: 10.1109/IJCNN.2016.7727626.
  17. 17.Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, Vancouver, Canada, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/S17-2001.
  18. 18.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. Universal sentence encoder for english. In Proceedings of the 2018 conference on empirical methods in natural language processing: system demonstrations, pages 169–174, 2018.
  19. 19.Dhivya Chandrasekaran and Vijay Mago. Evolution of semantic similarity—a survey. ACM Computing Surveys (CSUR), 54(2):1–37, 2021.
  20. 20.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020a.
  21. 21.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020b.
  22. 22.Leena Chennuru Vankadara and Ulrike von Luxburg. Measures of distortion for machine learning. Advances in Neural Information Processing Systems, 31, 2018.
  23. 23.Artem Chernodub and Dimitri Nowicki. Norm-preserving Orthogonal Permutation Linear Unit Activation Functions (OPLU). arXiv:1604.02313 [cs], January 2017. arXiv: 1604.02313.
  24. 24.Krzysztof Choromanski, Mark Rowland, Tamas Sarlos, Vikas Sindhwani, Richard Turner, and Adrian Weller. The Geometry of Random Features. In International Conference on Artificial Intelligence and Statistics, pages 1–9, March 2018.
  25. 25.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  26. 26.Peng Cui, Zhijie Deng, Wenbo Hu, and Jun Zhu. Accurate and reliable forecasting using stochastic differential equations. arXiv preprint arXiv:2103.15041, 2021.
  27. 27.Andreas Damianou and Neil Lawrence. Deep Gaussian Processes. In Artificial Intelligence and Statistics, pages 207–215, April 2013.
  28. 28.Jean Daunizeau. Semi-analytical approximations to statistical moments of sigmoid and softmax mappings of normal variables. arXiv preprint arXiv:1703.00091, 2017.
  29. 29.Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig. Laplace redux-effortless Bayesian deep learning. Advances in Neural Information Processing Systems, 34, 2021.
  30. 30.Morris H DeGroot and Mark J Schervish. Probability and statistics. Pearson Education, 2012.
  31. 31.Guillaume P. Dehaene. A deterministic and computable Bernstein-von Mises theorem. ArXiv, 2019.
  32. 32.John S. Denker and Yann LeCun. Transforming Neural-Net Output Levels to Probability Distributions. In Advances in Neural Information Processing Systems 3, pages 853–859. Morgan-Kaufmann, 1991.
  33. 33.Thomas Deselaers and Vittorio Ferrari. Visual and semantic similarity in imagenet. In CVPR 2011, pages 1777–1784. IEEE, 2011.
  34. 34.Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: Non-linear Independent Components Estimation. arXiv:1410.8516 [cs], October 2014. arXiv: 1410.8516.
  35. 35.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. arXiv:1605.08803 [cs, stat], May 2016. arXiv: 1605.08803.
  36. 36.Mike Dusenberry, Ghassen Jerfel, Yeming Wen, Yian Ma, Jasper Snoek, Katherine Heller, Balaji Lakshminarayanan, and Dustin Tran. Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors. Proceedings of the International Conference on Machine Learning, 1, 2020.
  37. 37.Vincent Dutordoir, Nicolas Durrande, and James Hensman. Sparse gaussian processes with spherical harmonic features. In International Conference on Machine Learning, pages 2793–2802. PMLR, 2020.
  38. 38.Runa Eschenhagen, Erik Daxberger, Philipp Hennig, and Agustinus Kristiadi. Mixtures of laplace approximations for improved post-hoc uncertainty in deep learning. CoRR, abs/2111.03577, 2021.
  39. 39.Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. Deep Ensembles: A Loss Landscape Perspective. arXiv:1912.02757 [cs, stat], December 2019. arXiv: 1912.02757.
  40. 40.David Freedman. Wald Lecture: On the Bernstein-von Mises theorem with infinite-dimensional parameters. The Annals of Statistics, 27(4):1119–1141, August 1999. ISSN 0090-5364, 2168-8966. doi: 10.1214/aos/1017938917.
  41. 41.Yarin Gal and Zoubin Ghahramani. Dropout As a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML’16, pages 1050–1059, New York, NY, USA, 2016. JMLR.org.
  42. 42.Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. Bayesian Data Analysis. Chapman and Hall/CRC, Boca Raton, 3 edition edition, November 2013. ISBN 978-1-4398-4095-5.
  43. 43.Tilmann Gneiting and Adrian E Raftery. Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association, 102(477):359–378, March 2007. ISSN 0162-1459. doi: 10.1198/016214506000001437.
  44. 44.Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E. Raftery. Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69(2):243–268, April 2007. ISSN 1467-9868. doi: 10.1111/j.1467-9868.2007.00587.x.
  45. 45.Henry Gouk, Eibe Frank, Bernhard Pfahringer, and Michael J Cree. Regularisation of neural networks by enforcing lipschitz continuity. Machine Learning, 110(2):393–416, 2021.
  46. 46.Peter D. Grunwald and A. Philip Dawid. Game theory, maximum entropy, minimum discrepancy ¨ and robust Bayesian decision theory. Annals of Statistics, 32(4):1367–1433, August 2004. ISSN 0090-5364, 2168-8966. doi: 10.1214/009053604000000553. Publisher: Institute of Mathematical Statistics.
  47. 47.Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. Improved training of wasserstein GANs. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS’17, pages 5769–5779, Long Beach, California, USA, December 2017. Curran Associates Inc. ISBN 978-1-5108-6096-4.
  48. 48.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On Calibration of Modern Neural Networks. In International Conference on Machine Learning, pages 1321–1330, July 2017. ISSN: 1938-7228 Section: Machine Learning.
  49. 49.Abhishek Gupta, Alagan Anpalagan, Ling Guan, and Ahmed Shaharyar Khwaja. Deep learning for object detection and scene perception in self-driving cars: Survey, challenges, and open issues. Array, 10:100057, 2021. ISSN 2590-0056. doi: https://doi.org/10.1016/j.array.2021.100057.
  50. 50.Danijar Hafner, Dustin Tran, Timothy Lillicrap, Alex Irpan, and James Davidson. Noise contrastive priors for functional uncertainty. In Uncertainty in Artificial Intelligence, pages 905–914. PMLR, 2020.
  51. 51.Kehang Han, Balaji Lakshminarayanan, and Jeremiah Zhe Liu. Reliable graph neural networks for drug discovery under distributional shift. In NeurIPS 2021 Workshop on Distribution Shifts: Connecting Methods and Applications, 2021.
  52. 52.Richard Harang and Ethan M Rudd. Towards principled uncertainty estimation for deep neural networks. arXiv preprint arXiv:1810.12278, 2018.
  53. 53.Tatsunori B Hashimoto, David Alvarez-Melis, and Tommi S Jaakkola. Word embeddings as metric recovery in semantic spaces. Transactions of the Association for Computational Linguistics, 4:273–286, 2016.
  54. 54.Michael Hauser and Asok Ray. Principles of Riemannian geometry in neural networks. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017a.
  55. 55.Michael Hauser and Asok Ray. Principles of Riemannian Geometry in Neural Networks. In Advances in Neural Information Processing Systems 30, pages 2807–2816. Curran Associates, Inc., 2017b.
  56. 56.Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf. Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 41–50, 2019a.
  57. 57.Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf. Why ReLU Networks Yield High-Confidence Predictions Far Away From the Training Data and How to Mitigate the Problem. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 41–50, 2019b.
  58. 58.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2018.
  59. 59.Dan Hendrycks and Kevin Gimpel. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In International Conference on Learning Representations, Toulon, France, April 2017.
  60. 60.Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. In International Conference on Learning Representations, 2018.
  61. 61.Dan Hendrycks, Kimin Lee, and Mantas Mazeika. Using Pre-Training Can Improve Model Robustness and Uncertainty. In International Conference on Machine Learning, pages 2712–2721, May 2019. ISSN: 1938-7228 Section: Machine Learning.
  62. 62.Dan Hendrycks*, Norman Mu*, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A Simple Method to Improve Robustness and Uncertainty under Data Shift. In International Conference on Learning Representations, 2020.
  63. 63.James Hensman, Alexander Matthews, and Zoubin Ghahramani. Scalable variational Gaussian process classification. In Artificial Intelligence and Statistics, pages 351–360. PMLR, 2015.
  64. 64.Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner. Towards a definition of disentangled representations. arXiv preprint arXiv:1812.02230, 2018.
  65. 65.Marius Hobbhahn, Agustinus Kristiadi, and Philipp Hennig. Fast predictive uncertainty for classification with bayesian deep networks. In Uncertainty in Artificial Intelligence, pages 822–832. PMLR, 2022.
  66. 66.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015.
  67. 67.Joern-Henrik Jacobsen, Jens Behrmann, Richard Zemel, and Matthias Bethge. Excessive invariance causes adversarial vulnerability. In International Conference on Learning Representations, 2019a.
  68. 68.Jorn-Henrik Jacobsen, Jens Behrmannn, Nicholas Carlini, Florian Tramer, and Nicolas Papernot. ¨ Exploiting excessive invariance caused by norm-bounded adversarial robustness. arXiv preprint arXiv:1903.10484, 2019b.
  69. 69.Jorn-Henrik Jacobsen, Arnold W.M. Smeulders, and Edouard Oyallon. i-RevNet: Deep invertible ¨ networks. In International Conference on Learning Representations, 2018.
  70. 70.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zˇ´ıdek, Anna Potapenko, et al. Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873):583–589, 2021.
  71. 71.Mohammad Emtiyaz E Khan, Alexander Immer, Ehsan Abedi, and Maciej Korzepa. Approximate Inference Turns Deep Networks into Gaussian Processes. In Advances in Neural Information Processing Systems 32, pages 3094–3104. Curran Associates, Inc., 2019.
  72. 72.Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational autoencoders and nonlinear ica: A unifying framework. In International Conference on Artificial Intelligence and Statistics, pages 2207–2217. PMLR, 2020.
  73. 73.Ian Kivlichan, Zi Lin, Jeremiah Liu, and Lucy Vasserman. Measuring and improving model-moderator collaboration using uncertainty estimation. In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021), pages 36–53, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.woah-1.5.
  74. 74.Agustinus Kristiadi, Matthias Hein, and Philipp Hennig. Being bayesian, even just a bit, fixes overconfidence in relu networks. In International conference on machine learning, pages 5436–5446. PMLR, 2020.
  75. 75.Agustinus Kristiadi, Matthias Hein, and Philipp Hennig. Learnable uncertainty under Laplace approximations. In Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, volume 161 of Proceedings of Machine Learning Research, pages 344–353. PMLR, 27–30 Jul 2021.
  76. 76.Agustinus Kristiadi, Matthias Hein, and Philipp Hennig. Being a bit frequentist improves bayesian neural networks. In International Conference on Artificial Intelligence and Statistics, pages 529–545. PMLR, 2022.
  77. 77.Frederik Kunstner, Philipp Hennig, and Lukas Balles. Limitations of the empirical fisher approximation for natural gradient descent. Advances in neural information processing systems, 32, 2019.
  78. 78.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. In Advances in Neural Information Processing Systems 30, pages 6402–6413. Curran Associates, Inc., 2017.
  79. 79.Jurgen Landes. Probabilism, entropies and strictly proper scoring rules. International Journal of Approximate Reasoning, 63:1–21, August 2015. ISSN 0888-613X.
  80. 80.Stefan Larson, Anish Mahendran, Joseph J Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K Kummerfeld, Kevin Leach, Michael A Laurenzano, Lingjia Tang, et al. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1311–1316, 2019.
  81. 81.Neil D Lawrence and Joaquin Quinonero-Candela. Local distance preservation in the GP-LVM through back constraints. In Proceedings of the 23rd international conference on Machine learning, pages 513–520, 2006.
  82. 82.L. LeCam. Convergence of Estimates Under Dimensionality Restrictions. The Annals of Statistics, 1(1):38–53, January 1973. ISSN 0090-5364, 2168-8966. doi: 10.1214/aos/1193342380.
  83. 83.Kimin Lee, Honglak Lee, Kibok Lee, and Jinwoo Shin. Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples. In International Conference on Learning Representations, 2018a.
  84. 84.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. Advances in neural information processing systems, 31, 2018b.
  85. 85.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks. In Advances in Neural Information Processing Systems 31, pages 7167–7177. Curran Associates, Inc., 2018c.
  86. 86.Shiyu Liang, Yixuan Li, and R Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. In International Conference on Learning Representations, 2018.
  87. 87.Fanghui Liu, Xiaolin Huang, Yudong Chen, and Johan AK Suykens. Random features for kernel approximation: A survey on algorithms, theory, and beyond. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):7128–7148, 2021.
  88. 88.Jeremiah Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax Weiss, and Balaji Lakshminarayanan. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. Advances in Neural Information Processing Systems, 33:7498–7512, 2020a.
  89. 89.Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. Advances in Neural Information Processing Systems, 33:21464–21475, 2020b.
  90. 90.Zhiyun Lu, Eugene Ie, and Fei Sha. Mean-field approximation to Gaussian-softmax integral with application to uncertainty estimation. arXiv preprint arXiv:2006.07584, 2020.
  91. 91.David Macedo and Teresa Ludermir. Enhanced isotropy maximization loss: Seamless and high-performance out-of-distribution detection simply replacing the softmax loss. arXiv preprint arXiv:2105.14399, 2021.
  92. 92.David Macedo, Tsang Ing Ren, Cleber Zanchettin, Adriano L. I. Oliveira, Alain Tapp, and Teresa Ludermir. Isotropic Maximization Loss and Entropic Score: Fast, Accurate, Scalable, Unexposed, Turnkey, and Native Neural Networks Out-of-Distribution Detection. arXiv:1908.05569 [cs, stat], February 2020. arXiv: 1908.05569.
  93. 93.David Macedo, Cleber Zanchettin, and Teresa Ludermir. Distinction maximization loss: Efficiently improving classification accuracy, uncertainty estimation, and out-of-distribution detection simply replacing the loss and calibrating. arXiv preprint arXiv:2205.05874, 2022.
  94. 94.David J. C. MacKay. A practical Bayesian framework for backpropagation networks. Neural Computation, 4(3):448–472, May 1992. ISSN 0899-7667. Number: 3 Publisher: MIT Press.
  95. 95.Andrey Malinin and Mark Gales. Predictive Uncertainty Estimation via Prior Networks. In Advances in Neural Information Processing Systems 31, pages 7047–7058. Curran Associates, Inc., 2018a.
  96. 96.Andrey Malinin and Mark Gales. Prior Networks for Detection of Adversarial Attacks. arXiv:1812.02575 [cs, stat], December 2018b. arXiv: 1812.02575.
  97. 97.Jirı Matousek. Lecture notes on metric embeddings. Technical report, Technical report, ETH Z ˇ urich, ¨ 2013.
  98. 98.Alexander Meinke and Matthias Hein. Towards neural networks that provably know when they don’t know. In International Conference on Learning Representations, 2020.
  99. 99.Thomas P. Minka. A family of algorithms for approximate Bayesian inference. phd, Massachusetts Institute of Technology, USA, 2001. AAI0803033.
  100. 100.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral Normalization for Generative Adversarial Networks. In International Conference on Learning Representations, 2018.
  101. 101.Saif Mohammad and Graeme Hirst. Distributional measures as proxies for semantic distance: A survey. Computational Linguistics, 1(1), 2006.
  102. 102.Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. Deterministic neural networks with appropriate inductive biases capture epistemic and aleatoric uncertainty. arXiv preprint arXiv:2102.11582, 2021.
  103. 103.Zachary Nado, Neil Band, Mark Collier, Josip Djolonga, Michael W Dusenberry, Sebastian Farquhar, Angelos Filos, Marton Havasi, Rodolphe Jenatton, Ghassen Jerfel, et al. Uncertainty baselines: Benchmarks for uncertainty & robustness in deep learning. arXiv preprint arXiv:2106.04015, 2021.
  104. 104.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng. Reading Digits in Natural Images with Unsupervised Feature Learning. In NeurIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011.
  105. 105.Anh Nguyen, Jason Yosinski, and Jeff Clune. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 427–436, 2015.
  106. 106.Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. Measuring calibration in deep learning. In CVPR Workshop, 2019.
  107. 107.Kazuki Osawa, Siddharth Swaroop, Mohammad Emtiyaz E Khan, Anirudh Jain, Runa Eschenhagen, Richard E Turner, and Rio Yokota. Practical Deep Learning with Bayesian Principles. In Advances in Neural Information Processing Systems 32, pages 4287–4299. Curran Associates, Inc., 2019.
  108. 108.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’ s uncertainty? evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  109. 109.Shreyas Padhy, Zachary Nado, Jie Ren, Jeremiah Liu, Jasper Snoek, and Balaji Lakshminarayanan. Revisiting one-vs-all classifiers for predictive uncertainty and out-of-distribution detection in neural networks. arXiv preprint arXiv:2007.05134, 2020.
  110. 110.Maxim Panov and Vladimir Spokoiny. Finite Sample Bernstein von Mises Theorem for Semiparametric Problems. Bayesian Analysis, 10(3):665–710, September 2015. ISSN 1936-0975, 1931-6690. doi: 10.1214/14-BA926.
  111. 111.Vardan Papyan. Traces of class/cross-class structure pervade deep learning spectra. Journal of Machine Learning Research, 21(252):1–64, 2020.
  112. 112.Matthew Parry, A. Philip Dawid, and Steffen Lauritzen. Proper local scoring rules. Annals of Statistics, 40(1):561–592, February 2012. ISSN 0090-5364, 2168-8966. doi: 10.1214/12-AOS971. Publisher: Institute of Mathematical Statistics.
  113. 113.Dominique Perrault-Joncas and Marina Meila. Metric learning and manifolds: Preserving the intrinsic geometry. Preprint Department of Statistics, University of Washington, 2012.
  114. 114.Ryan Poplin, Pi-Chuan Chang, David Alexander, Scott Schwartz, Thomas Colthurst, Alexander Ku, Dan Newburger, Jojo Dijamco, Nam Nguyen, Pegah T Afshar, et al. A universal SNP and small-indel variant caller using deep neural networks. Nature biotechnology, 36(10):983–987, 2018.
  115. 115.Ali Rahimi and Benjamin Recht. Uniform approximation of functions with random bases. In 2008 46th Annual Allerton Conference on Communication, Control, and Computing, pages 555–561. IEEE, 2008a.
  116. 116.Ali Rahimi and Benjamin Recht. Random Features for Large-Scale Kernel Machines. In Advances in Neural Information Processing Systems 20, pages 1177–1184. Curran Associates, Inc., 2008b.
  117. 117.Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. University Press Group Limited, January 2006. ISBN 978-0-262-18253-9. Google-Books-ID: vWtwQgAACAAJ.
  118. 118.Jamie Reilly, Bonnie Zuckerman, Ann Marie Finley, Celia Paula Litovsky, and Yoed Kenett. What is semantic distance? a review and proposed method for modeling conceptual transitions in natural language. 2022.
  119. 119.Jie Ren, Peter J Liu, Emily Fertig, Jasper Snoek, Ryan Poplin, Mark Depristo, Joshua Dillon, and Balaji Lakshminarayanan. Likelihood ratios for out-of-distribution detection. Advances in neural information processing systems, 32, 2019.
  120. 120.Jie Ren, Stanislav Fort, Jeremiah Liu, Abhijit Guha Roy, Shreyas Padhy, and Balaji Lakshminarayanan. A simple fix to Mahalanobis distance for improving near-OOD detection. arXiv preprint arXiv:2106.09022, 2021.
  121. 121.Carlos Riquelme, George Tucker, and Jasper Snoek. Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling. In International Conference on Learning Representations, 2018.
  122. 122.Hippolyt Ritter, Aleksandar Botev, and David Barber. A Scalable Laplace Approximation for Neural Networks. In International Conference on Learning Representations, 2018.
  123. 123.Francois Rousseau, Lucas Drumetz, and Ronan Fablet. Residual Networks as Flows of Diffeomorphisms. Journal of Mathematical Imaging and Vision, 62(3):365–375, April 2020. ISSN 1573-7683. doi: 10.1007/s10851-019-00890-3.
  124. 124.Abhijit Guha Roy, Jie Ren, Shekoofeh Azizi, Aaron Loh, Vivek Natarajan, Basil Mustafa, Nick Pawlowski, Jan Freyberg, Yuan Liu, Zach Beaver, et al. Does your dermatology classifier know what it doesn’t know? detecting the long-tail of unseen conditions. Medical Image Analysis, 75:102274, 2022.
  125. 125.Wenjie Ruan, Xiaowei Huang, and Marta Kwiatkowska. Reachability analysis of deep neural networks with provable guarantees. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, IJCAI’18, pages 2651–2659, Stockholm, Sweden, July 2018. AAAI Press. ISBN 978-0-9992411-2-7.
  126. 126.Walter Rudin et al. Principles of mathematical analysis, volume 3. McGraw-hill New York, 1976.
  127. 127.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3):211–252, December 2015. ISSN 1573-1405. doi: 10.1007/s11263-015-0816-y.
  128. 128.Ruslan Salakhutdinov and Andriy Mnih. Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo. In Proceedings of the 25th International Conference on Machine Learning, ICML ’08, pages 880–887, New York, NY, USA, 2008. ACM. ISBN 978-1-60558-205-4. doi: 10.1145/1390156.1390267.
  129. 129.Russ R Salakhutdinov and Geoffrey E Hinton. Using deep belief nets to learn covariance kernels for gaussian processes. In Advances in Neural Information Processing Systems, 2007.
  130. 130.Hugh Salimbeni and Marc Deisenroth. Doubly stochastic variational inference for deep Gaussian processes. Advances in neural information processing systems, 30, 2017.
  131. 131.Walter J. Scheirer, Lalit P. Jain, and Terrance E. Boult. Probability Models for Open Set Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(11):2317–2324, November 2014. ISSN 1939-3539. doi: 10.1109/TPAMI.2014.2321392. Conference Name: IEEE Transactions on Pattern Analysis and Machine Intelligence.
  132. 132.Micheal O. Searcod. Metric Spaces. Springer London, London, 2007 edition edition, August 2006. ISBN 978-1-84628-369-7.
  133. 133.Murat Sensoy, Lance Kaplan, and Melih Kandemir. Evidential deep learning to quantify classification uncertainty. Advances in Neural Information Processing Systems, 31, 2018a.
  134. 134.Murat Sensoy, Lance Kaplan, and Melih Kandemir. Evidential Deep Learning to Quantify Classification Uncertainty. In Advances in Neural Information Processing Systems 31, pages 3179–3189. Curran Associates, Inc., 2018b.
  135. 135.Lei Shu, Hu Xu, and Bing Liu. Doc: Deep open classification of text documents. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2911–2916, 2017.
  136. 136.Nicki Skafte, Martin Jørgensen, and Søren Hauberg. Reliable training and estimation of variance networks. Advances in Neural Information Processing Systems, 32, 2019.
  137. 137.Lewis Smith, Joost van Amersfoort, Haiwen Huang, Stephen Roberts, and Yarin Gal. Can convolutional resnets approximately preserve input distances? a frequency analysis perspective. arXiv preprint arXiv:2106.02469, 2021.
  138. 138.Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams. Scalable bayesian optimization using deep neural networks. In International conference on machine learning, pages 2171–2180. PMLR, 2015.
  139. 139.Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues. Robust Large Margin Deep Neural Networks. IEEE Transactions on Signal Processing, 2017. doi: 10.1109/TSP.2017.2708039.
  140. 140.Natasa Tagasovska and David Lopez-Paz. Single-Model Uncertainties for Deep Learning. In Advances in Neural Information Processing Systems 32, pages 6417–6428. Curran Associates, Inc., 2019.
  141. 141.Joshua B Tenenbaum, Vin de Silva, and John C Langford. A global geometric framework for nonlinear dimensionality reduction. Science, 290(5500):2319–2323, 2000.
  142. 142.Sunil Thulasidasan, Gopinath Chennupati, Jeff A Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. On Mixup Training: Improved Calibration and Predictive Uncertainty for Deep Neural Networks. In Advances in Neural Information Processing Systems 32, pages 13888–13899. Curran Associates, Inc., 2019.
  143. 143.Luke Tierney, Robert E. Kass, and Joseph B. Kadane. Approximate Marginal Densities of Nonlinear Functions. Biometrika, 76(3):425–433, 1989. ISSN 0006-3444. doi: 10.2307/2336109. Publisher: [Oxford University Press, Biometrika Trust].
  144. 144.Michalis Titsias. Variational Learning of Inducing Variables in Sparse Gaussian Processes. In Artificial Intelligence and Statistics, pages 567–574, April 2009.
  145. 145.Gia-Lac Tran, Edwin V. Bonilla, John Cunningham, Pietro Michiardi, and Maurizio Filippone. Calibrating Deep Convolutional Gaussian Processes. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 1554–1563, April 2019. ISSN: 1938-7228 Section: Machine Learning.
  146. 146.Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. Lipschitz-Margin Training: Scalable Certification of Perturbation Invariance for Deep Neural Networks. In Advances in Neural Information Processing Systems 31, pages 6541–6550. Curran Associates, Inc., 2018.
  147. 147.Joost Van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. Uncertainty estimation using a single deep deterministic neural network. In International conference on machine learning, pages 9690–9700. PMLR, 2020.
  148. 148.Joost van Amersfoort, Lewis Smith, Andrew Jesson, Oscar Key, and Yarin Gal. On feature collapse and deep kernel learning for single forward pass uncertainty. arXiv preprint arXiv:2102.11409, 2021.
  149. 149.Nikhita Vedula, Nedim Lipka, Pranav Maneriker, and Srinivasan Parthasarathy. Towards Open Intent Discovery for Conversational Text. arXiv:1904.08524 [cs], April 2019. arXiv: 1904.08524.
  150. 150.Grace Wahba. Spline Models for Observational Data. SIAM, September 1990. ISBN 978-0-89871-244-5.
  151. 151.Yeming Wen, Dustin Tran, and Jimmy Ba. BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong Learning. In International Conference on Learning Representations, 2020.
  152. 152.Tsui-Wei Weng, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Dong Su, Yupeng Gao, Cho-Jui Hsieh, and Luca Daniel. Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach. In International Conference on Learning Representations, 2018.
  153. 153.Florian Wenzel, Kevin Roth, Bastiaan Veeling, Jakub Swiatkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin. How good is the Bayes posterior in deep neural networks really? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 10248–10259. PMLR, 13–18 Jul 2020.
  154. 154.Andrew G Wilson and Pavel Izmailov. Bayesian deep learning and a probabilistic perspective of generalization. Advances in neural information processing systems, 33:4697–4708, 2020.
  155. 155.Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P Xing. Deep kernel learning. In Artificial intelligence and statistics, pages 370–378. PMLR, 2016a.
  156. 156.Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P. Xing. Stochastic Variational Deep Kernel Learning. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NeurIPS’16, pages 2594–2602, USA, 2016b. Curran Associates Inc. ISBN 978-1-5108-3881-9.
  157. 157.Mohammad-Ali Yaghoub-Zadeh-Fard, Boualem Benatallah, Fabio Casati, Moshe Chai Barukh, and Shayan Zamanirad. User Utterance Acquisition for Training Task-Oriented Bots: A Review of Challenges, Techniques and Opportunities. IEEE Internet Computing, pages 1–1, 2020. ISSN 1941-0131. doi: 10.1109/MIC.2020.2978157. Conference Name: IEEE Internet Computing.
  158. 158.Felix Xinnan X Yu, Ananda Theertha Suresh, Krzysztof M Choromanski, Daniel N Holtmann-Rice, and Sanjiv Kumar. Orthogonal Random Features. In Advances in Neural Information Processing Systems 29, pages 1975–1983. Curran Associates, Inc., 2016.
  159. 159.Sergey Zagoruyko and Nikos Komodakis. Wide Residual Networks. arXiv:1605.07146 [cs], June 2017. arXiv: 1605.07146.
  160. 160.Yinhe Zheng, Guanyi Chen, and Minlie Huang. Out-of-Domain Detection for Natural Language Understanding in Dialog Systems. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:1198–1209, 2020. ISSN 2329-9304. doi: 10.1109/TASLP.2020.2983593. Conference Name: IEEE/ACM Transactions on Audio, Speech, and Language Processing.

Citation

MLA
Liu, J. Z., et al. “A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness”. Journal of Machine Learning Research, vol. 24, no. 42, 2023, pp. 1–3, https://www.jmlr.org/papers/v24/22-0479.html.
APA
Liu, J. Z., Padhy, S., Ren, J., Lin, Z., Wen, Y., Jerfel, G., Nado, Z., Snoek, J., Tran, D., & Lakshminarayanan, B. (2023). A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness. Journal of Machine Learning Research, 24(42), 1–63. https://www.jmlr.org/papers/v24/22-0479.html
Chicago
Liu, J. Z., S. Padhy, J. Ren, et al. 2023. “A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness”. Journal of Machine Learning Research 24 (42): 1–63. https://www.jmlr.org/papers/v24/22-0479.html.
Harvard
Liu, J.Z. et al. (2023) “A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness”, Journal of Machine Learning Research, 24(42), pp. 1–63. Available at: https://www.jmlr.org/papers/v24/22-0479.html.
Vancouver
1. Liu JZ, Padhy S, Ren J, Lin Z, Wen Y, Jerfel G, Nado Z, Snoek J, Tran D, Lakshminarayanan B (2023) A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness. Journal of Machine Learning Research 24:1–63

BibTeX

@article{JMLR:v24:22-0479,
  author  = {Jeremiah Zhe Liu and Shreyas Padhy and Jie Ren and Zi Lin and Yeming Wen and Ghassen Jerfel and Zachary Nado and Jasper Snoek and Dustin Tran and Balaji Lakshminarayanan},
  title   = {A Simple Approach to Improve Single-Model Deep Uncertainty via Distance-Awareness},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {42},
  pages   = {1--63},
  url     = {http://jmlr.org/papers/v24/22-0479.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/