All You Need is a Good Functional Prior for Bayesian Deep Learning

Ba-Hien TranSimone RossiDimitrios MiliosMaurizio Filippone

article2022JMLR84 citations

Proposes a method to align Bayesian neural network weight priors with interpretable Gaussian process functional priors by minimizing their Wasserstein distance, significantly boosting predictive performance across regression and classification tasks.

Listen

Deep neural networks are central to modern machine learning, yet standard deployments struggle to quantify predictive uncertainty reliably. While Bayesian approaches offer a principled way to represent uncertainty by placing prior distributions over parameters, modern networks contain thousands to millions of weights. In practice, standard independent Gaussian priors placed on these parameters induce pathological and uncontrollable behaviors in function space, such as collapsing to flat, uninformative functions in deep architectures or assigning extreme, overconfident probabilities to single classes. This inability to establish interpretable and well-behaved prior beliefs has severely limited the reliability and adoption of Bayesian deep learning in high-stakes settings.

The article introduces and evaluates a practical framework to tune the parameter priors of Bayesian neural networks so that their induced function-level distributions match interpretable target functional priors, specifically Gaussian processes. By optimizing this alignment directly in function space before observing training data, the authors demonstrate that Bayesian neural networks can achieve superior predictive accuracy, robust uncertainty calibration, and resilience to data corruption.

The approach formulates prior selection as a sample-based distance minimization problem using the Wasserstein distance, optimized via an alternating gradient-based algorithm. The authors explore parameterizations of increasing flexibility, including layer-wise Gaussian distributions, hierarchical Inverse-Gamma distributions, and normalizing flows. The target functional properties are specified via standard and hierarchical Gaussian processes. Empirical credibility is established across a comprehensive benchmark suite: standard regression and classification datasets from the UCI repository, active learning tasks, and deep vision architectures (including LeNet-5, VGG-16, and PreResNet-20) evaluated on MNIST, CIFAR-10, and corrupted CIFAR-10C datasets. Posterior inference is primarily conducted using scale-adapted stochastic gradient Hamiltonian Monte Carlo sampling.

The investigation yields several key findings. First, optimizing parameter priors via Wasserstein distance consistently aligns neural network outputs with desired functional behaviors, avoiding the optimization instabilities seen in Kullback-Leibler divergence approaches. Second, Gaussian process-induced priors deliver systematically superior predictive performance across benchmarks; on CIFAR-10 image classification, the hierarchical induced prior achieved top accuracies of 76.51% on LeNet-5, 87.03% on VGG-16, and 88.20% on PreResNet-20, outperforming standard priors, posterior temperature scaling, and popular non-Bayesian deep ensembles. Third, models trained with these functional priors exhibit superior calibration and robustness under severe covariate shift and out-of-distribution inputs, maintaining high predictive entropy on unfamiliar data rather than making overconfident erroneous predictions. Fourth, in active learning scenarios, the proposed priors guided faster, more sample-efficient data acquisition than fixed prior alternatives.

These findings indicate that prior specification—rather than approximate Bayesian inference itself—has been the critical missing link in Bayesian deep learning. For decision-makers, adopting functional priors significantly reduces operational risks associated with model overconfidence and data corruption without requiring costly ad-hoc interventions like data-driven cross-validation or heuristic posterior tempering. Furthermore, this prior optimization is performed completely prior to training, preserving Bayesian integrity while reducing computational exploration time compared to massive hyperparameter grid searches.

Organizations deploying deep neural networks in safety-critical, active learning, or distributionally shifting environments should integrate functional prior optimization into their Bayesian modeling pipelines. When selecting prior parameterizations, practitioners face a trade-off: layer-wise hierarchical priors scale easily to deep convolutional models and deliver the strongest overall performance, whereas normalizing flow priors offer maximum expressiveness but scale linearly with parameter count and are currently better suited for smaller architectures. Future efforts should focus on sparsifying normalizing flow priors for deep models, extending the framework to unsupervised and latent variable architectures, and reducing the computational cost of the underlying distance optimization.

The authors express high confidence in their findings, supported by rigorous convergence diagnostics and validation across multiple sampling schemes, including full-batch Hamiltonian Monte Carlo. However, readers should note that computational overhead remains cubic with respect to the number of measurement points used in Gaussian process sampling, and the alignment relies on smooth network activations and finite measurement grids. Despite these boundary conditions, the evidence confirms that functional prior alignment is an effective, scalable strategy for robust deep learning.

Tran et al (2022).pdf
  • Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Bayes by Backprop establishes how distributions over neural-network weights yield Bayesian predictions, the parameter-space starting point for understanding why this paper instead tunes priors by their induced functions.
  • Paper: Deep Gaussian Processes, Andreas C. Damianou et al. (2012). Deep Gaussian Processes develops the function-space Bayesian models used here as interpretable target priors, making its account of deep probabilistic functions a direct conceptual prerequisite.
  • Paper: Practical Variational Inference for Neural Networks, Alex Graves (2011). Practical Variational Inference for Neural Networks introduces distributions over network weights and their approximate Bayesian treatment, clarifying the parameter-prior setting this paper seeks to improve.
  • Paper: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, Yarin Gal et al. (2016). Dropout as a Bayesian Approximation connects neural networks to deep Gaussian processes, preparing readers to understand the paper’s comparison between parameterized networks and function-space priors.

No sufficiently relevant recommendations were found.

Cover for All You Need is a Good Functional Prior for Bayesian Deep Learning

Table of Contents

  • 1. Introduction
  • Contributions
  • 2. Related Work
  • 3. Preliminaries
  • 3.1 Bayesian Neural Networks
  • 3.2 Gaussian Process Priors
  • 3.3 Wasserstein Distance
  • 4. Imposing Gaussian Process Priors on Bayesian Neural Networks
  • 4.1 Wasserstein Distance Optimization
  • 4.2 Prior Parameterization for Neural Networks
  • 4.3 Algorithm and Complexity
  • 5. Examples and Practical Considerations
  • 5.1 Visualization on a 1D regression synthetic dataset
  • 5.2 The effects of the GP prior on the BNN posterior
  • 5.3 Wasserstein distance vs KL divergence
  • 6. Experimental Evaluation
  • 6.1 Baselines
  • 6.2 UCI regression benchmark
  • 6.3 UCI classification benchmark
  • 6.4 Bayesian convolutional neural networks for image classification
  • 6.5 Optimizing priors with data: cross-validation and empirical Bayes
  • 6.6 Active learning
  • 6.7 Maximum-a-posteriori (MAP) estimation with GP-induced prior
  • 7. Conclusions
  • Acknowledgments
  • Appendix A. Implementation and experimental details
  • A.1 Deep Ensemble
  • A.2 Likelihoods for BNNs
  • A.3 Sampling from the posterior using scale-adapted SGHMC
  • A.4 Tempered posterior
  • A.5 Details on the sampling scheme for BNN hierarchical priors
  • A.6 MAP estimation with Gaussian prior
  • A.7 Network architectures
  • A.8 Measuring similarity between GPs and BNNs using maximum mean discrepancy
  • A.9 Details on the experiments with functional BNNs and empirical Bayes
  • Appendix B. Additional results
  • B.1 Additional results on MAP estimation with GP-induced priors
  • B.2 Tabular results on the UCI benchmarks
  • B.3 Convergence of Wasserstein optimization
  • B.4 Additional comparisons with the empirical Bayes approach
  • B.5 Additional results with full-batch Hamiltonian Monte Carlo
  • B.6 Additional discussion on the optimization of Wasserstein distance
  • Appendix C. A primer on Wasserstein Distance
  • Step 1: Kantorovich duality
  • Step 2: D-Transforms
  • Step 2.1: Euclidean distance
  • Step 3. Putting everything together
  • References

Knowls

  1. Knowl 1 — Match a neural network’s functional prior to a target process using Wasserstein distance

    model/method

    A prior over neural-network weights induces a distribution over functions, but its effect on those functions is difficult to specify directly. The proposed method instead tunes weight-prior parameters to match the neural network’s induced function distribution to a target stochastic process, such as a Gaussian process (GP), using the 1-Wasserstein distance. The comparison is made at a finite set of inputs: draw MM measurement inputs XM=(x1,…,xM)X_M=(x_1,\ldots,x_M) from a chosen input distribution qq, and let fMf_M denote the vector of function values at those inputs. If pNN(fM;ψ)p_{\mathrm{NN}}(f_M;\psi) is the neural-network prior induced by weight-prior parameters ψ\psi, and pGP(fM)p_{\mathrm{GP}}(f_M) is the target process, optimize

    min⁡ψmax⁡θ  EXM∼q[EfM∼pGPϕθ(fM)−EfM∼pNN(⋅;ψ)ϕθ(fM)].\min_{\psi}\max_{\theta}\;\mathbb{E}_{X_M\sim q}\left[\mathbb{E}_{f_M\sim p_{\mathrm{GP}}}\phi_\theta(f_M)-\mathbb{E}_{f_M\sim p_{\mathrm{NN}}(\cdot;\psi)}\phi_\theta(f_M)\right].

    Here ϕθ\phi_\theta is a 1-Lipschitz scalar-valued critic with parameters θ\theta. The objective needs samples from the two processes, not a closed-form density for the neural network’s induced function prior; accordingly, the target process can be any sampleable stochastic process.

  2. Knowl 2 — Use interpretable GP kernels, optionally with hyperpriors, to specify target functions

    definition

    The target functional prior is specified by a GP with mean zero and a covariance kernel. For inputs x,x′∈RDx,x'\in\mathbb{R}^D, the radial basis function (RBF) kernel used in the method is

    κα,l(x,x′)=α2exp⁡(−∥x−x′∥22l2),\kappa_{\alpha,l}(x,x')=\alpha^2\exp\left(-\frac{\|x-x'\|_2^2}{l^2}\right),

    where α>0\alpha>0 controls the marginal standard deviation and l>0l>0 controls the length scale of variation. To avoid fixing these kernel hyperparameters, the paper also uses a hierarchical target: sample α\alpha and ll independently from log-normal distributions with user-chosen log-space mean mm and variance s2s^2, then sample f∼GP(0,κα,l)f\sim\mathcal{GP}(0,\kappa_{\alpha,l}). This construction allows the neural-network prior to match a distribution over GP functions, rather than only one GP with fixed hyperparameters.

  3. Knowl 3 — Estimate the distance with alternating critic and prior updates

    algorithm

    The optimization procedure takes as input a sampleable target process, a sampleable neural-network prior controlled by ψ\psi, an input distribution qq, a number NsN_s of samples per process, and a finite measurement set of size NMN_M. It returns optimized neural-network prior parameters ψ\psi. At each outer iteration, it samples measurement inputs from qq, draws target-process and neural-network functions at those inputs, trains a critic to distinguish the two sample distributions, and then updates ψ\psi to reduce the critic’s estimated Wasserstein discrepancy. The critic is trained for multiple steps before each prior update; the paper’s implementation uses a multilayer perceptron with two hidden layers of 200 units and softplus activations. The critic is constrained with a gradient penalty: for interpolated functions f^=ϵfNN+(1−ϵ)fGP\hat f=\epsilon f_{\mathrm{NN}}+(1-\epsilon)f_{\mathrm{GP}}, where ϵ∼Uniform[0,1]\epsilon\sim\mathrm{Uniform}[0,1], penalize λ(∥∇f^ϕθ(f^)∥2−1)2\lambda(\|\nabla_{\hat f}\phi_\theta(\hat f)\|_2-1)^2. The paper uses λ=10\lambda=10, nLipschitz=200n_{\mathrm{Lipschitz}}=200 critic updates per prior update, Adagrad with learning rate 0.020.02 for the critic, and RMSprop for the prior parameters with learning rate 0.050.05 on UCI and banana experiments and 0.010.01 on other experiments. Unless a particular experiment specifies otherwise, measurement inputs combine observed training inputs with uniformly sampled inputs from the domain. The outer optimization stops when ψ\psi is judged to have converged. The stated per-outer-step cost for a hierarchical GP target is O(nLipschitzNs2NM3)O(n_{\mathrm{Lipschitz}}N_s^2N_M^3), dominated by repeated GP sampling and matrix factorizations; this prior-fitting cost is incurred before posterior sampling.

  4. Knowl 4 — Fit Gaussian, hierarchical, or flow-based neural-network weight priors

    model/method

    The method considers three parameterizations of the neural-network weight prior. In the Gaussian version (GPI-G), weights and biases in each layer have separate, zero-mean Gaussian distributions, each sharing one variance within that layer; the optimized parameters are the layerwise weight and bias variances. Positivity is enforced by optimizing an unconstrained parameter ρ\rho and setting the standard deviation to σ=log⁡(1+eρ)\sigma=\log(1+e^\rho); samples use the reparameterization w=σεw=\sigma\varepsilon, with ε∼N(0,1)\varepsilon\sim\mathcal{N}(0,1). In the hierarchical version (GPI-H), each layer’s weight and bias variances have separate inverse-gamma priors, and parameters are conditionally Gaussian given those variances; the inverse-gamma shape and rate parameters are optimized. Posterior sampling alternates SGHMC updates of the weights with Gibbs updates of the variances. In the normalizing-flow version (GPI-NF), each layer’s prior starts with a factorized Gaussian and applies four planar transformations. For a layer’s vector z∈Rdz\in\mathbb{R}^{d}, a transformation is tk(z)=z+uktanh⁡(ak⊤z+bk)t_k(z)=z+u_k\tanh(a_k^\top z+b_k), with trainable vectors uk,ak∈Rdu_k,a_k\in\mathbb{R}^{d} and scalar bkb_k. The log density accounts for the Jacobian determinant of each transformation. This more flexible parameterization can improve the match to a target GP, but its number of prior parameters grows as O(KNbnn)O(KN_{\mathrm{bnn}}) for KK transformations and NbnnN_{\mathrm{bnn}} network parameters; the paper applies it only to fully connected networks.

  5. Knowl 5 — Wasserstein matching converges more consistently than the tested KL-based alternatives

    empirical result

    On six UCI regression datasets, the paper compared Wasserstein-based prior fitting with KL-based fitting using early stopping and with KL-based fitting augmented by a spectral Stein gradient estimator for the entropy gradient. Similarity between the fitted neural-network prior and target GP was assessed using squared maximum mean discrepancy (MMD), alongside test root mean squared error (RMSE) and negative log-likelihood (NLL). The KL-based variants initially improved the approximation but subsequently degraded it, including the variant with the entropy-gradient estimator. Wasserstein optimization instead consistently improved the GP match over iterations and did not require estimating the neural-network prior’s entropy. These results support the proposed optimization choice for the tested settings; they do not establish that Wasserstein matching is universally preferable.

  6. Knowl 6 — GP-induced priors improve results on UCI regression and classification benchmarks

    empirical result

    The UCI regression experiments used randomly selected 90%/10% training/test splits, repeated 10 times (5 times for protein), and two-layer tanh multilayer perceptrons with 100 hidden units for smaller datasets and 200 for protein. The classification experiments used seven datasets, two-layer tanh networks, and 10 random splits. For regression prior fitting, the target was a hierarchical ARD RBF GP, with 128 function samples and measurement sets of 100 inputs comprising 70% training examples and 30% uniformly sampled domain points. The GP-induced Gaussian, hierarchical, and—on some regression datasets—flow priors generally improved predictive performance over fixed Gaussian and fixed hierarchical priors; they also outperformed the tested deep ensembles on most regression datasets. On classification, GPI-H achieved the best reported NLL across the seven datasets and the best or near-best accuracy; GPI-G also generally improved on a fixed Gaussian prior. These comparisons used test RMSE for regression and test accuracy and NLL for classification.

  7. Knowl 7 — GPI-H improves CIFAR-10 prediction across three convolutional architectures

    data/table

    The CIFAR-10 comparison evaluated posterior predictions from SGHMC on LeNet5, PreResNet20, and VGG16 without data augmentation or adversarial examples. The hierarchical target GP used log-normal hyperpriors on variance and length scale, and the neural-network priors were fitted using 128 process samples and 32 measurement inputs. After a 10,000-iteration burn-in, 200 posterior samples were collected with 10,000 simulation iterations between samples. The table reports test accuracy and NLL as mean ± one standard error over four runs. GPI-H achieved the best accuracy and NLL among the listed methods for all three architectures; in particular, it improved on the corresponding fixed hierarchical prior in each case.

    Could not parse LaTeX table
  8. Knowl 8 — GP-induced priors improve uncertainty behavior under shift and out-of-distribution inputs

    empirical result

    Several image experiments tested uncertainty as well as in-distribution accuracy. For PreResNet20 trained on CIFAR-10 and evaluated on CIFAR-10-C, which contains 16 corruption types at five severity levels, GPI-H outperformed the deep ensemble at every reported corruption level in accuracy and NLL; GPI-G also improved on the fixed Gaussian prior. When training subsets of different sizes were used, GPI-H performed best across the tested proportions, and GPI-G outperformed the fixed Gaussian prior. In a separate LeNet5 experiment trained on MNIST, predictive entropies were compared on MNIST and NOT-MNIST. The GP-induced priors produced better out-of-distribution uncertainty than the fixed priors in the reported entropy-CDF comparison, while the deep ensemble tended to be overconfident on both in-distribution and out-of-distribution examples. These are empirical findings for the architectures and datasets tested.

  9. Knowl 9 — GPI priors support more effective uncertainty-driven regression active learning

    empirical result

    The active-learning experiment used the UCI regression settings with data split into 20% initial training, 60% unlabeled pool, and 20% test sets. At each of 10 rounds, the model selected the 5% of the initial pool size with the largest posterior predictive variance, added those points to training, and was evaluated by test RMSE. Results are mean test RMSE at the final round, with reported standard errors. GPI-H had the lowest or tied-lowest RMSE on most datasets and improved on the fixed hierarchical prior on several; GPI-G generally improved on the fixed Gaussian prior. The values below are reported without additional rounding.

    Could not parse LaTeX table
  10. Knowl 10 — The method is heuristic and has finite-domain and computational limitations

    limitation

    The paper treats GP-based functional priors as a useful, interpretable heuristic, not as a theoretically guaranteed superior prior; other priors may be better suited to particular data or architectures. Matching is performed on finite measurement sets sampled from a chosen input distribution, so the optimized prior is directly constrained only at those sampled locations. Prior fitting adds computational cost, particularly for GP targets requiring repeated matrix factorizations. The flow-based prior is especially costly in large networks because its parameter count grows linearly with the number of network parameters, and the paper uses it only for fully connected networks. The authors identify reducing Wasserstein-optimization complexity and improving flow-prior scalability as future work.

Coverage note — The MAP-estimation demonstration and detailed SGHMC/HMC implementation comparisons are omitted because they are secondary to the paper’s central functional-prior method and its main predictive evaluations.

References

  1. 1.R. Amit and R. Meir. Meta-Learning by Adjusting Priors Based on Extended PAC-Bayes Theory. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 205–214. PMLR, 2018.
  2. 2.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein Generative Adversarial Networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 214–223. PMLR, 2017.
  3. 3.A. Ashukha, A. Lyzhov, D. Molchanov, and D. Vetrov. Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning. In International Conference on Learning Representations, 2020.
  4. 4.A. Atanov, A. Ashukha, K. Struminsky, D. Vetrov, and M. Welling. The Deep Weight Prior. In International Conference on Learning Representations, 2019.
  5. 5.C. M. Bishop. Pattern recognition and machine learning. Springer, 1st ed. 2006. corr. 2nd printing 2011 edition, Aug. 2006.
  6. 6.C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra. Weight Uncertainty in Neural Network. In International Conference on Machine Learning, pages 1613–1622. PMLR, 2015.
  7. 7.F.-X. Briol, C. J. Oates, M. Girolami, M. A. Osborne, and D. Sejdinovic. Probabilistic Integration: A Role in Statistical Computation? Statistical Science, 34(1):1–22, 02 2019.
  8. 8.T. Chen, E. Fox, and C. Guestrin. Stochastic Gradient Hamiltonian Monte Carlo. In Proceedings of the 31st International Conference on Machine Learning, Proceedings of Machine Learning Research, pages 1683–1691. PMLR, 2014.
  9. 9.X. Chen, D. P. Kingma, T. Salimans, Y. Duan, P. Dhariwal, J. Schulman, I. Sutskever, and P. Abbeel. Variational Lossy Autoencoder. In International Conference on Learning Representations, 2017.
  10. 10.J. Cockayne, C. J. Oates, I. C. Ipsen, and M. Girolami. A Bayesian Conjugate Gradient Method (with Discussion). Bayesian Analysis, 14(3):937–1012, 09 2019.
  11. 11.E. A. Daxberger, A. Kristiadi, A. Immer, R. Eschenhagen, M. Bauer, and P. Hennig. Laplace Redux – Effortless Bayesian Deep Learning. In Advances in Neural Information Processing Systems, volume 34, pages 20089–20103, 2021.
  12. 12.S. Delattre and N. Fournier. On the Kozachenko–Leonenko entropy estimator. Journal of Statistical Planning and Inference, 185:69–93, 2017.
  13. 13.D. Dua and C. Graff. UCI machine learning repository. University of California, Irvine, School of Information and Computer Sciences, 2017. URL http://archive.ics.uci.edu/ml.
  14. 14.S. Duane, A. Kennedy, B. J. Pendleton, and D. Roweth. Hybrid Monte Carlo. Physics Letters B, 195(2):216 – 222, 1987.
  15. 15.J. Duchi, E. Hazan, and Y. Singer. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. Journal of Machine Learning Research, 12(61):2121–2159, 2011.
  16. 16.D. Duvenaud, O. Rippel, R. Adams, and Z. Ghahramani. Avoiding Pathologies in Very Deep Networks. In Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, volume 33 of Proceedings of Machine Learning Research, pages 202–210. PMLR, 2014.
  17. 17.D. Flam-Shepherd, J. Requeima, and D. Duvenaud. Mapping Gaussian Process Priors to Bayesian Neural Networks. In NeurIPS workshop on Bayesian Deep Learning, 2017.
  18. 18.D. Flam-Shepherd, J. Requeima, and D. Duvenaud. Characterizing and Warping the Function space of Bayesian Neural Networks. In NeurIPS workshop on Bayesian Deep Learning, 2018.
  19. 19.Y. Gal and Z. Ghahramani. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In Proceedings of the 33nd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1050–1059. JMLR, June 19-24 2016.
  20. 20.A. Gelman and D. B. Rubin. Inference from Iterative Simulation using Multiple Sequences. Statistical Science, 7(4):457–472, 1992.
  21. 21.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, volume 27, pages 2672–2680. Curran Associates, Inc., 2014.
  22. 22.A. Graves. Practical Variational Inference for Neural Networks. In Advances in Neural Information Processing Systems, volume 24, pages 2348–2356. Curran Associates, Inc., 2011.
  23. 23.A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. J. Smola. A Kernel Two-Sample Test. Journal of Machine Learning Research, 13:723–773, 2012.
  24. 24.A. Grover, M. Dhar, and S. Ermon. Flow-GAN: Combining Maximum Likelihood and Adversarial Learning in Generative Models. In Proceedings of the 32nd Conference on Artificial Intelligence, pages 3069–3076. AAAI Press, 2018.
  25. 25.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville. Improved Training of Wasserstein GANs. In Advances in Neural Information Processing Systems, volume 30, pages 5767–5777. Curran Associates, Inc., 2017.
  26. 26.D. Ha, A. M. Dai, and Q. V. Le. Hypernetworks. In International Conference on Learning Representations, 2017.
  27. 27.D. Hafner, D. Tran, T. P. Lillicrap, A. Irpan, and J. Davidson. Noise Contrastive Priors for Functional Uncertainty. In Proceedings of the 35h Conference on Uncertainty in Artificial Intelligence, page 332. AUAI Press, 2019.
  28. 28.K. He, X. Zhang, S. Ren, and J. Sun. Identity Mappings in Deep Residual Networks. In Proceeding of the 14th European Conference on Computer Vision, volume 9908 (Part IV) of Lecture Notes in Computer Science, pages 630–645. Springer, 2016.
  29. 29.J. Heek and N. Kalchbrenner. Bayesian Inference for Large Scale Image Classification. arXiv:1908.03491, 2019.
  30. 30.D. Hendrycks and T. Dietterich. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In International Conference on Learning Representations, 2019.
  31. 31.M. D. Hoffman and A. Gelman. The No-U-turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo. Journal of Machine Learning Research, 15(1):1593–1623, 2014.
  32. 32.N. Houlsby, F. Huszar, Z. Ghahramani, and J. Hernández-lobato. Collaborative Gaussian Processes for Preference Learning. In Advances in Neural Information Processing Systems, volume 25, pages 2096–2104. Curran Associates, Inc., 2012.
  33. 33.A. Immer, M. Bauer, V. Fortuin, G. Rätsch, and M. E. Khan. Scalable Marginal Likelihood Estimation for Model Selection in Deep Learning. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 4563–4573. PMLR, 2021a.
  34. 34.A. Immer, M. Korzepa, and M. Bauer. Improving Predictions of Bayesian Neural Nets via Local Linearization. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pages 703–711. PMLR, 2021b.
  35. 35.A. Jacot, F. Gabriel, and C. Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. In Advances in Neural Information Processing Systems, volume 31, pages 8571–8580. Curran Associates, Inc., 2018.
  36. 36.M. Jankowiak and F. Obermeyer. Pathwise Derivatives Beyond the Reparameterization Trick. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2240–2249. PMLR, 2018.
  37. 37.L. V. Kantorovich. On the transfer of masses. Doklady Akademii Nauk SSSR, 37:227–229, 1942.
  38. 38.L. V. Kantorovich. On a problem of Monge. Uspekhi Matematicheskikh Nauk, 3:225–226, 1948.
  39. 39.T. Karaletsos and T. D. Bui. Gaussian Process Meta-Representations For Hierarchical Neural Network Weight Priors. In 2nd Symposium on Advances in Approximate Bayesian Inference, 2019.
  40. 40.T. Karaletsos and T. D. Bui. Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights. In Advances in Neural Information Processing Systems, volume 33, 2020.
  41. 41.A. Kendall and Y. Gal. What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision? In Advances in Neural Information Processing Systems, volume 30, pages 5574–5584. Curran Associates, Inc., 2017.
  42. 42.M. E. Khan, A. Immer, E. Abedi, and M. Korzepa. Approximate Inference Turns Deep Networks into Gaussian Processes. In Advances in Neural Information Processing Systems, pages 3088–3098, 2019.
  43. 43.D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations, 2015.
  44. 44.D. P. Kingma and M. Welling. Auto-Encoding Variational Bayes. In International Conference on Learning Representations, 2014.
  45. 45.D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling. Improved Variational Inference with Inverse Autoregressive Flow. In Advances in Neural Information Processing Systems, volume 29, pages 4743–4751. Curran Associates, Inc., 2016.
  46. 46.A. Krizhevsky and G. Hinton. Learning Multiple Layers of Features from Tiny Images. Master’s thesis, Department of Computer Science, University of Toronto, 2009.
  47. 47.B. Lakshminarayanan, A. Pritzel, and C. Blundell. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. In Advances in Neural Information Processing Systems, volume 30, pages 6402–6413. Curran Associates, Inc., 2017.
  48. 48.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  49. 49.Y. LeCun, Y. Bengio, and G. Hinton. Deep learning. Nature, 521(7553):436–444, May 2015.
  50. 50.J. Lee, S. S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein. Finite Versus Infinite Neural Networks: an Empirical Study. In Advances in Neural Information Processing Systems, volume 33, 2020.
  51. 51.H. Liu, Y. S. Ong, X. Shen, and J. Cai. When Gaussian Process Meets Big Data: A Review of Scalable GPs. IEEE Transactions on Neural Networks and Learning Systems, 31(11):4405–4423, 2020.
  52. 52.Q. Liu and D. Wang. Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm. In Advances in Neural Information Processing Systems, volume 29, pages 2378–2386. Curran Associates, Inc., 2016.
  53. 53.C. Louizos and M. Welling. Multiplicative Normalizing Flows for Variational Bayesian Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 2218–2227. PMLR, 2017.
  54. 54.C. Ma, Y. Li, and J. M. Hernández-Lobato. Variational Implicit Processes. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 4222–4233. PMLR, 2019.
  55. 55.D. J. MacKay. Information-based objective functions for active data selection. Neural computation, 4(4):590–604, 1992.
  56. 56.D. J. MacKay. Probable Networks and Plausible Predictions - a Review of Practical Bayesian Methods for Supervised Neural Networks. Network: Computation in Neural Systems, 6 (3):469–505, 1995.
  57. 57.D. J. MacKay. Bayesian non-linear modeling for the prediction competition. In Maximum Entropy and Bayesian Methods, pages 221–234. Springer, 1996.
  58. 58.D. J. C. Mackay. Information Theory, Inference and Learning Algorithms. Cambridge University Press, 1st edition, 2003.
  59. 59.T. Matsubara, C. J. Oates, and F. Briol. The Ridgelet Prior: A Covariance Function Approach to Prior Specification for Bayesian Neural Networks. Journal of Machine Learning Research, 22:1–57, 2021.
  60. 60.A. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani. Gaussian Process Behaviour in Wide Deep Neural Networks. In International Conference on Learning Representations, 2018.
  61. 61.J. Močkus. On Bayesian Methods for Seeking the Extremum. In Optimization Techniques IFIP Technical Conference Novosibirsk, pages 400–404. Springer Berlin Heidelberg, 1975.
  62. 62.A. Müller. Integral Probability Metrics and Their Generating Classes of Functions. Advances in Applied Probability, 29(2):429–443, 1997.
  63. 63.E. T. Nalisnick, J. Gordon, and J. M. Hernández-Lobato. Predictive Complexity Priors. In Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pages 694–702. PMLR, 2021.
  64. 64.R. M. Neal. Bayesian Learning for Neural Networks (Lecture Notes in Statistics). Springer, 1st edition, Aug. 1996.
  65. 65.F. Nogueira. Bayesian Optimization: Open source constrained global optimization tool for Python, 2014. URL https://github.com/fmfn/BayesianOptimization.
  66. 66.A. O’Hagan. Bayes–Hermite quadrature. Journal of Statistical Planning and Inference, 29 (3):245 – 260, 1991.
  67. 67.K. Osawa, S. Swaroop, M. E. E. Khan, A. Jain, R. Eschenhagen, R. E. Turner, and R. Yokota. Practical Deep Learning with Bayesian Principles. In Advances in Neural Information Processing Systems, volume 32, pages 4287–4299. Curran Associates, Inc., 2019.
  68. 68.Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek. Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift. In Advances in Neural Information Processing Systems, volume 32, pages 13991–14002. Curran Associates, Inc., 2019.
  69. 69.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems, volume 32, pages 8026–8037. Curran Associates, Inc., 2019.
  70. 70.T. Pearce, R. Tsuchida, M. Zaki, A. Brintrup, and A. Neely. Expressive Priors in Bayesian Neural Networks: Kernel Combinations and Periodic Functions. In Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence, page 25. AUAI Press, 2019.
  71. 71.D. Phan, N. Pradhan, and M. Jankowiak. Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro. arXiv:1912.11554, 2019.
  72. 72.C. E. Rasmussen and Z. Ghahramani. Bayesian Monte Carlo. In Advances in Neural Information Processing Systems, volume 15, pages 489–496. MIT Press, 2002.
  73. 73.C. E. Rasmussen and C. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006.
  74. 74.D. Rezende and S. Mohamed. Variational Inference with Normalizing Flows. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1530–1538, Lille, France, 07–09 Jul 2015. PMLR.
  75. 75.D. J. Rezende, S. Mohamed, and D. Wierstra. Stochastic Backpropagation and Approximate Inference in Deep Generative Models. In Proceedings of the 31th International Conference on Machine Learning, volume 32 of Proceeding of Machine Learning Research, pages 1278–1286, Beijing, China, 21-26 June 2014. PMLR.
  76. 76.S. Rossi, P. Michiardi, and M. Filippone. Good Initializations of Variational Bayes for Deep Models. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 5487–5497, Long Beach, California, USA, 09–15 Jun 2019. PMLR.
  77. 77.S. Rossi, S. Marmin, and M. Filippone. Walsh-Hadamard Variational Inference for Bayesian Deep Learning. In Advances in Neural Information Processing Systems, volume 33, 2020.
  78. 78.B. Settles. Active Learning Literature Survey. Technical report, University of Wisconsin-Madison Department of Computer Sciences, 2009.
  79. 79.J. Shi, S. Sun, and J. Zhu. A Spectral Approach to Gradient Estimation for Implicit Distributions. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 4651–4660. PMLR, 2018.
  80. 80.J. Shi, M. E. Khan, and J. Zhu. Scalable Training of Inference Networks for Gaussian-Process Models. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 5758–5768. PMLR, 2019.
  81. 81.K. Simonyan and A. Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations, 2015.
  82. 82.N. Skafte, M. Jorgensen, and S. Hauberg. Reliable Training and Estimation of Variance Networks. In Advances in Neural Information Processing Systems, volume 32, pages 6326–6336. Curran Associates, Inc., 2019.
  83. 83.J. Snoek, H. Larochelle, and R. P. Adams. Practical Bayesian Optimization of Machine Learning Algorithms. In Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012.
  84. 84.J. T. Springenberg, A. Klein, S. Falkner, and F. Hutter. Bayesian Optimization with Robust Bayesian Neural Networks. In Advances in Neural Information Processing Systems, volume 29, pages 4134–4142. Curran Associates, Inc., 2016.
  85. 85.N. Srinivas, A. Krause, S. M. Kakade, and M. W. Seeger. Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. In Proceedings of the 27th International Conference on Machine Learning, pages 1015–1022. Omnipress, 2010.
  86. 86.S. Sun, G. Zhang, J. Shi, and R. Grosse. Functional Variational Bayesian Neural Networks. In International Conference on Learning Representations, 2019.
  87. 87.T. Tieleman and G. Hinton. Lecture 6.5—RmsProp: Divide the Gradient by a Running Average of Its Recent Magnitude. COURSERA: Neural Networks for Machine Learning, 2012.
  88. 88.Tishby, Levin, and Solla. Consistent inference of probabilities in layered networks: predictions and generalizations. In International 1989 Joint Conference on Neural Networks, pages 403–409 vol.2, 1989.
  89. 89.B.-H. Tran, S. Rossi, D. Milios, P. Michiardi, E. V. Bonilla, and M. Filippone. Model Selection for Bayesian Autoencoders. In Advances in Neural Information Processing Systems, volume 34, pages 19730–19742. Curran Associates, Inc., 2021.
  90. 90.C. Villani. Topics in Optimal Transportation. Graduate studies in mathematics. American Mathematical Society, 2003.
  91. 91.F. Wenzel, K. Roth, B. S. Veeling, J. Swi¸atkowski, L. Tran, S. Mandt, J. Snoek, T. Salimans, R. Jenatton, and S. Nowozin. How Good is the Bayes Posterior in Deep Neural Networks Really? In Proceeding of the 37th International Conference on Machine Learning, 2020.
  92. 92.W. Yang, L. Lorch, M. A. Graule, S. Srinivasan, A. Suresh, J. Yao, M. F. Pradier, and F. Doshi-velez. Output-Constrained Bayesian Neural Networks. In ICML workshop on Uncertainty & Robustness in Deep Learning, 2019.
  93. 93.Y. Yao, L. Rosasco, and A. Caponnetto. On Early Stopping in Gradient Descent Learning. Constructive Approximation, 26(2):289–315, 2007.
  94. 94.R. Zhang, C. Li, J. Zhang, C. Chen, and A. G. Wilson. Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning. In International Conference on Learning Representations, 2020.

Citation

MLA
Tran, B.-H., et al. “All You Need Is a Good Functional Prior for Bayesian Deep Learning”. Journal of Machine Learning Research, vol. 23, no. 74, 2022, pp. 1–6, https://www.jmlr.org/papers/v23/20-1340.html.
APA
Tran, B.-H., Rossi, S., Milios, D., & Filippone, M. (2022). All You Need is a Good Functional Prior for Bayesian Deep Learning. Journal of Machine Learning Research, 23(74), 1–56. https://www.jmlr.org/papers/v23/20-1340.html
Chicago
Tran, B.-H., S. Rossi, D. Milios, and M. Filippone. 2022. “All You Need Is a Good Functional Prior for Bayesian Deep Learning”. Journal of Machine Learning Research 23 (74): 1–56. https://www.jmlr.org/papers/v23/20-1340.html.
Harvard
Tran, B.-H. et al. (2022) “All You Need is a Good Functional Prior for Bayesian Deep Learning”, Journal of Machine Learning Research, 23(74), pp. 1–56. Available at: https://www.jmlr.org/papers/v23/20-1340.html.
Vancouver
1. Tran B-H, Rossi S, Milios D, Filippone M (2022) All You Need is a Good Functional Prior for Bayesian Deep Learning. Journal of Machine Learning Research 23:1–56

BibTeX

@article{JMLR:v23:20-1340,
  author  = {Ba-Hien Tran and Simone Rossi and Dimitrios Milios and Maurizio Filippone},
  title   = {All You Need is a Good Functional Prior for Bayesian Deep Learning},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {74},
  pages   = {1--56},
  url     = {http://jmlr.org/papers/v23/20-1340.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/