Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI

Theodore PapamarkouMaria SkoularidouKonstantina PallaLaurence AitchisonJulyan ArbelDavid B. DunsonMaurizio FilipponeVincent FortuinPhilipp HennigJosé Miguel Hernández-Lobato

article2024ICML79 citations

Argues that integrating Bayesian deep learning into massive foundation models is essential for mitigating overconfident hallucinations, improving uncertainty quantification in safety-critical decision-making, and enabling data-efficient learning.

Listen

Modern artificial intelligence relies heavily on massive, overparameterized neural networks that achieve strong predictive accuracy on standard supervised benchmarks. However, these systems frequently generate incorrect outputs with dangerously high confidence, fail to adapt cleanly to evolving data, and struggle when high-quality training examples are scarce. In safety-critical fields such as healthcare, autonomous driving, and scientific discovery, uncalibrated overconfidence introduces severe operational risks. Bayesian deep learning addresses these fundamental flaws by treating network parameters and model architectures probabilistically rather than relying on brittle point estimates.

The article systematically demonstrates why Bayesian principles are essential for large-scale AI and outlines a strategic technical roadmap to make probabilistic deep learning practical, scalable, and computationally efficient for foundation models and mission-critical workflows.

To establish this framework, the article provides a comprehensive evaluation of existing Bayesian paradigms against standard deep learning methods. It reviews structural approximations such as Laplace methods and variational inference, stochastic sampling techniques including Markov chain Monte Carlo, and non-parametric formulations such as deep kernel processes. The synthesis incorporates theoretical insights, algorithmic trade-offs, and empirical findings across diverse computational domains to identify the primary performance and engineering bottlenecks currently hindering widespread adoption.

The findings establish that probabilistic integration provides four major capabilities where traditional deep learning falls short. First, it enables reliable uncertainty quantification, allowing models to recognize out-of-distribution inputs, mitigate hallucinations in large language models, and safely route uncertain decisions to human experts. Second, it yields superior data efficiency by embedding domain-informed priors, preventing overfitting on sparse datasets, and improving active learning by strategically selecting the most informative training points. Third, it enhances continuous adaptation and knowledge transfer across changing domains without catastrophically forgetting previous patterns. Fourth, contrary to common assumptions about excessive overhead, modern implementations such as post-hoc Laplace approximations and training trajectory recycling require minimal compute beyond standard optimization—with some structured Bayesian network approaches achieving up to four times greater parameter memory efficiency than deterministic baselines.

These insights demonstrate that incorporating Bayesian inference directly reduces the asymmetric costs and safety hazards of automated errors. In high-stakes environments, deploying models that reliably quantify confidence allows organizations to avoid catastrophic failures, lower manual validation costs, and ensure compliance with emerging safety standards. While deep ensembling is frequently used to capture predictive variability, it becomes economically and computationally unsustainable as base models scale to billions of parameters, reinforcing the necessity of principled, lightweight Bayesian approximations.

Decision-makers and research teams should pursue targeted hybrid implementations rather than attempting full, brute-force Bayesian inference. Promising initial pathways include applying Bayesian techniques selectively to high-impact subcomponents, such as the final layer of large networks or via parameter-efficient fine-tuning methods like low-rank adaptation. Organizations should also invest in scalable posterior sampling algorithms, integrate hardware-aware low-precision numerics, and leverage domain-specific functional priors. Further development of standardized sequential decision benchmarks and user-friendly probabilistic software is necessary to lower adoption barriers for practitioners.

Significant technical challenges remain, notably the difficulty of accurately mapping complex, high-dimensional posterior landscapes and the absence of universal diagnostic metrics for probabilistic deep models. While high confidence supports the core value of Bayesian uncertainty calibration and data efficiency, engineering teams must recognize that fully sampling billion-parameter distributions remains computationally prohibitive, necessitating pragmatic, structured approximations in near-term production environments.

Papamarkou et al (2024).pdf
Cover for Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI

Table of Contents

  • 1. Introduction
  • 2. Why Bayesian Deep Learning Matters
  • 2.1. Uncertainty Quantification
  • 2.2. Data Efficiency
  • 2.3. Adaptability to New and Evolving Domains
  • 2.4. Model Misspecification and Interpretability
  • 3. Current Challenges
  • 3.1. Laplace and Variational Approximations
  • 3.2. Ensembles
  • 3.3. Posterior Sampling Algorithms
  • 3.4. Prior Specification
  • 3.5. Scalability
  • 3.6. Foundation Models
  • 4. Proposed Future Directions
  • 4.1. Posterior Sampling Algorithms
  • 4.2. Hybrid Bayesian Approaches
  • 4.3. Deep Kernel Processes and Machines
  • 4.4. Semi-Supervised and Self-Supervised Learning
  • 4.5. Mixed Precision and Tensor Computations
  • 4.6. Compression Strategies
  • 4.7. Other Future Directions
  • 5. Final Remarks
  • Acknowledgements
  • Impact Statement
  • References
  • A. Background
  • A.1. Laplace Approximations
  • A.2. Variational Inference
  • A.3. Ensembles
  • A.4. Posterior Sampling Algorithms
  • A.5. Prior Specification
  • A.6. Deep Kernel Processes
  • A.7. Deep Kernel Machines
  • B. Diagnostics, Metrics and Benchmarks
  • C. Software Usability
  • D. Topical Developments

Knowls

  1. Knowl 1 — Bayesian deep learning’s central role and bottleneck

    model/method

    Bayesian deep learning (BDL) applies Bayesian inference to deep learning by representing uncertainty over neural-network parameters and using that distribution to form predictive distributions. The paper’s position is that this framework can address needs that point-estimate deep learning does not adequately address, especially reliable uncertainty assessment and the incorporation of prior knowledge in safety-critical or data-limited settings. It argues that the principal barrier to broader use is the computational cost of approximate posterior inference—not that uncertainty has become irrelevant in large-scale AI—and calls for BDL methods whose time, memory, and predictive utility are practical at contemporary scales.

  2. Knowl 2 — Predictive uncertainty supports safer decisions

    model/method

    The paper argues that a predictive distribution from BDL can make uncertainty operational in decision-making: when a prediction is ambiguous or out of distribution, a system can flag uncertainty, defer a decision to a human, or wait for more information rather than act on an unreliable point prediction. This is especially relevant when errors have asymmetric costs and when confident but incorrect outputs, such as those produced by language models, can cause harm. The paper also distinguishes BDL from conformal prediction by noting that BDL does not require exchangeability and can represent dependence across spatiotemporal data through suitable latent variables.

  3. Knowl 3 — Approximate inference trades posterior coverage against computational cost

    limitation

    The paper reviews a central trade-off among common BDL inference approaches. Laplace and Gaussian variational approximations provide analytically tractable, often computationally practical uncertainty estimates, but are generally local approximations that capture one posterior mode; the paper also notes concerns about parameterization dependence and Gaussian mismatch to local posterior geometry. SWAG constructs a Gaussian approximation from an SGD trajectory with relatively little extra computation. Deep ensembles instead train from multiple initializations and can capture prediction differences across modes, but require multiple model trainings and executions. Sampling methods such as Hamiltonian Monte Carlo can produce strong posterior approximations, but at substantial additional computational cost. These approaches therefore cannot be ranked as uniformly superior: their usefulness depends on the posterior, data, and computational budget.

  4. Knowl 4 — Scalability requires balancing time, memory, and model utility

    limitation

    The paper identifies simultaneous time efficiency, memory efficiency, predictive performance, and uncertainty calibration as the central scalability challenge for BDL. It cautions against treating Bayesian neural networks as inherently slower or more memory-intensive than deterministic networks: it points to reported results of up to fourfold memory efficiency in parameter count for some Bayesian networks, as well as approximate-posterior methods that reuse standard training trajectories with negligible additional computation. These examples do not remove the challenge; the paper’s target is uncertainty quantification whose computational cost is competitive with point estimation, including when GPU resources are constrained.

  5. Knowl 5 — Prior specification should target useful function behavior

    model/method

    In BDL, a prior over neural-network weights induces a prior over functions, and the paper argues that the function-level prior is what matters most for generalization. A default independent, zero-centered Gaussian weight prior can be unsatisfactory: its variance controls regularization and can affect whether network units behave more linearly or nonlinearly, while a misspecified prior can contribute to the cold-posterior effect. The paper calls for computationally efficient, informative priors that encode desired properties such as sparsity, equivariance, fairness, reliable uncertainty, or robustness to covariate shift. Function-space priors offer a direct way to express beliefs about functions, but can create support mismatches, problematic variational objectives, or costly Gaussian-process computations.

  6. Knowl 6 — Bayesian updating supports data efficiency and adaptation

    model/method

    The paper argues that BDL can make effective use of limited or costly data by combining prior information with observations, with Bayesian regularization and uncertainty modeling potentially reducing overfitting and sensitivity to outliers. The same updating of beliefs offers a route to transfer, continual, and lifelong learning: information from earlier tasks can influence downstream solutions, while posterior updates can incorporate new evidence as domains evolve. Predictive uncertainty can also guide active learning by selecting informative examples for labeling. The paper presents these as advantages and research opportunities, not as a guarantee that every BDL method will outperform deterministic learning in small-data settings.

  7. Knowl 7 — Bayesian model averaging represents structural uncertainty

    model/method

    Bayesian model averaging (BMA) accounts for uncertainty over model structures by assigning prior and posterior probabilities to multiple candidate models and averaging their predictions, rather than committing to one architecture. In BDL, this can represent uncertainty over neural-network architectures or patterns of included weights. The paper argues that averaging across structures can reduce the effect of model misspecification and complement parameter uncertainty. It also identifies reproducible, interpretable Bayesian inferences from deep networks as an ongoing need, particularly in scientific applications where black-box prediction is not the sole objective.

  8. Knowl 8 — Foundation models create both a need and an opportunity for BDL

    model/method

    The paper identifies large language models and other foundation models as important but relatively underexplored settings for BDL. Their scale makes repeated training and conventional posterior sampling especially costly, while their deployment can involve out-of-distribution inputs, overconfident errors, personalization, causal questions, or fine-tuning with sparse data—all settings in which uncertainty estimates and prior information may be useful. The paper notes early directions including Bayesian low-rank adaptation, Bayesian optimization, and Bayesian reward modeling, but emphasizes that meaningful applications and scalable methods are still needed to establish when Bayesian inference tools add value to foundation models.

  9. Knowl 9 — Posterior sampling needs to explore modes or target functionals

    model/method

    The paper calls for stochastic-gradient Markov chain Monte Carlo (SG-MCMC) methods that explore high-dimensional, multimodal neural-network posteriors at a speed closer to stochastic gradient descent while preserving useful posterior accuracy. Proposed research directions include tempered sampling and methods based on transport, diffusion, or flow models to move between isolated modes; running inference in linear, sparse, or other structured parameter subspaces; and targeting identifiable, lower-dimensional functionals of interest rather than attempting to represent the full parameter posterior. The paper also suggests hybrid samplers that combine structured variational inference with MCMC to trade computation against accuracy.

  10. Knowl 10 — Hybrid inference can concentrate Bayesian computation where it matters

    model/method

    The paper proposes hybrid BDL as a practical strategy: use Bayesian inference for parts of a model where uncertainty is especially useful and affordable, and point estimation for the remaining parts. A last-layer Laplace approximation is given as an example. Related directions include posterior inference in parameter subspaces or targeted subnetworks, and models combining neural networks with Gaussian processes. The underlying proposal is to avoid requiring full Bayesian treatment of every parameter when a restricted Bayesian component can provide useful uncertainty at lower cost.

  11. Knowl 11 — Deep kernel processes and machines avoid weight-space symmetries

    model/method

    The paper presents deep kernel processes (DKPs) as a function- or kernel-centered alternative to placing priors directly on neural-network weights or intermediate features. DKPs treat kernels, represented through layerwise Gram matrices, as random variables; this representation avoids permutation and rotation symmetries of feature or weight parameterizations that can make posteriors highly multimodal. Deep kernel machines (DKMs) take an infinite-width limit of this approach and temper the likelihood so that representation learning is retained. In a DKM, learned layerwise Gram matrices describe the representations, while the likelihood fits the data and regularization terms draw the model toward a neural-network Gaussian-process prior. The paper identifies these methods as promising for understanding function-space inference and representation learning, while noting that scalable training remains important.

  12. Knowl 12 — BDL needs evaluation beyond point-prediction accuracy

    limitation

    The paper identifies a lack of consensus on diagnostics, metrics, datasets, and benchmarks tailored to BDL. In particular, convergence of high-dimensional, multimodal posterior samplers is difficult to establish; predictive evaluations often emphasize the mean while giving insufficient attention to predictive variance; and systematic measures of epistemic and aleatoric uncertainty are lacking beyond established classification measures such as AUROC and AUPRC. The paper recommends stronger evaluation under distribution shift and misspecification, including calibration and probabilistic performance guarantees, and argues for sequential tasks—such as experimental design, Bayesian optimization, active learning, and bandits—where uncertainty can be evaluated through its effect on decisions.

  13. Knowl 13 — Computational co-design could make BDL more practical

    model/method

    The paper proposes addressing BDL’s computational burden through systems and numerical methods as well as inference algorithms. Mixed-precision computation can introduce uncertainty that Bayesian methods could propagate to predictions, while Bayesian decision-making could help select numerical precision for different computations. Compression strategies include sparsity-inducing priors, entropy-based weight coding, and Bayesian quantization, potentially allowing compute–precision trade-offs at prediction time. The paper also highlights probabilistic numerics, specialized hardware, and easier-to-use software as directions for reducing computational or practitioner overhead.

  14. Knowl 14 — Self-supervised objectives can be given a Bayesian interpretation

    model/method

    The paper observes that semi-supervised and self-supervised learning objectives can appear ad hoc from a Bayesian perspective when they do not correspond to an explicit likelihood. It notes that improvements from semi-supervised learning may depend on patterns introduced by data curation, which limits straightforward extrapolation to uncurated data. For self-supervised learning, the paper points to recognition-parameterized probabilistic models as a way to formulate a rigorous likelihood for objectives based on related augmented views. This perspective may clarify how self-supervised learning works and how it could be used to construct informative Bayesian priors.

Coverage note — The paper’s domain-specific application examples and specialized directions in human–AI interaction, decentralized learning, reinforcement learning, and computer vision are omitted as illustrative extensions rather than distinct core contributions; their broader uncertainty, evaluation, and scalability themes are represented above.

References

  1. 1.Abdar, M., Samami, M., Mahmoodabad, S. D., Doan, T., Mazoure, B., Hashemifesharaki, R., Liu, L., Khosravi, A., Acharya, U. R., Makarenkov, V., et al. Uncertainty quantification in skin cancer classification using three-way decision-based Bayesian deep learning. Computers in Biology and Medicine, 135:104418, 2021.
  2. 2.Abdullah, A. A., Hassan, M. M., and Mustafa, Y. T. A review on Bayesian deep learning in healthcare: Applications and challenges. IEEE Access, 10:36538–36562, 2022.
  3. 3.Agrawal, D., Papamarkou, T., and Hinkle, J. Wide neural networks with bottlenecks are deep Gaussian processes. Journal of Machine Learning Research, 21(175):1–66, 2020.
  4. 4.Aitchison, L. A statistical theory of cold posteriors in deep neural networks. International Conference on Learning Representations, 2021.
  5. 5.Aitchison, L. and Ganev, S. InfoNCE is variational inference in a recognition parameterised model. arXiv preprint arXiv:2107.02495, 2023.
  6. 6.Aitchison, L., Yang, A., and Ober, S. W. Deep kernel processes. In International Conference on Machine Learning, 2021.
  7. 7.Alemi, A. A. and Poole, B. Variational prediction. In Fifth Symposium on Advances in Approximate Bayesian Inference, 2023.
  8. 8.Alexos, A., Boyd, A. J., and Mandt, S. Structured stochastic gradient MCMC. In International Conference on Machine Learning, 2022.
  9. 9.Andriushchenko, M. Adversarial attacks on GPT-4 via simple random search. Preprint, 2023.
  10. 10.Antoran, J., Bhatt, U., Adel, T., Weller, A., and Hernandez-Lobato, J. M. Getting a CLUE: A method for explaining uncertainty estimates. In International Conference on Learning Representations, 2021.
  11. 11.Antoran, J., Allingham, J. U., Janz, D., Daxberger, E., Nalisnick, E., and Hernandez-Lobato, J. M. Linearised Laplace inference in networks with normalisation layers and the neural g-prior. In Fourth Symposium on Advances in Approximate Bayesian Inference, 2022.
  12. 12.Antoran, J., Padhy, S., Barbano, R., Nalisnick, E., Janz, D., and Hernandez-Lobato, J. M. Sampling-based inference for large linear models, with application to linearised Laplace. In International Conference on Learning Representations, 2023.
  13. 13.Arbel, J., Pitas, K., Vladimirova, M., and Fortuin, V. A primer on Bayesian neural networks: review and debates. arXiv preprint arXiv:2309.16314, 2023.
  14. 14.Atzeni, M., Sachan, M., and Loukas, A. Infusing lattice symmetry priors in attention mechanisms for sample-efficient abstract geometric reasoning. arXiv preprint arXiv:2306.03175, 2023.
  15. 15.Bamler, R., Salehi, F., and Mandt, S. Augmenting and tuning knowledge graph embeddings. In Conference on Uncertainty in Artificial Intelligence, 2020.
  16. 16.Band, N., Rudner, T. G. J., Feng, Q., Filos, A., Nado, Z., Dusenberry, M. W., Jerfel, G., Tran, D., and Gal, Y. Benchmarking Bayesian deep learning on diabetic retinopathy detection tasks. In Advances in Neural Information Processing Systems, 2021.
  17. 17.Bayes, T. An essay towards solving a problem in the doctrine of chances. Philosophical transactions of the Royal Society of London, 53:370–418, 1763. By the late Rev. Mr. Bayes, FRS communicated by Mr. Price, in a letter to John Canton, AMFRS.
  18. 18.Bhatt, U., Antoran, J., Zhang, Y., Liao, Q. V., Sattigeri, P., Fogliato, R., Melançon, G., Krishnan, R., Stanley, J., Tickoo, O., Nachman, L., Chunara, R., Srikumar, M., Weller, A., and Xiang, A. Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty. In Conference on AI, Ethics, and Society, 2021.
  19. 19.Bingham, E., Chen, J. P., Jankowiak, M., Obermeyer, F., Pradhan, N., Karaletsos, T., Singh, R., Szerlip, P., Horsfall, P., and Goodman, N. D. Pyro: Deep universal probabilistic programming. Journal of Machine Learning Research, 20(28):1–6, 2019.
  20. 20.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight uncertainty in neural networks. In International Conference on Machine Learning, 2015a.
  21. 21.Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. Weight uncertainty in neural networks. In International Conference on Machine Learning, 2015b.
  22. 22.Bouchiat, K., Immer, A., Yeche, H., Ratsch, G., and Fortuin, V. Laplace-approximated neural additive models: Improving interpretability with Bayesian inference. arXiv preprint arXiv:2305.16905, 2023.
  23. 23.Brooks, S., Gelman, A., Jones, G., and Meng, X.-L. Handbook of Markov chain Monte Carlo. CRC Press, 2011.
  24. 24.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Advances in Neural Information Processing Systems, 2020.
  25. 25.Burt, D. R., Ober, S. W., Garriga-Alonso, A., and van der Wilk, M. Understanding variational inference in function-space. In Third Symposium on Advances in Approximate Bayesian Inference, 2020.
  26. 26.Carvalho, C. M., Polson, N. G., and Scott, J. G. Handling sparsity via the horseshoe. In Artificial Intelligence and Statistics, 2009.
  27. 27.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al. Open problems and fundamental limitations of reinforcement learning from human feedback. Transactions on Machine Learning Research, 2023.
  28. 28.Chang, P. G., Murphy, K. P., and Jones, M. On diagonal approximations to the extended Kalman filter for online training of Bayesian neural networks. In Continual Lifelong Learning Workshop at ACML 2022, 2022.
  29. 29.Chen, T., Fox, E., and Guestrin, C. Stochastic gradient Hamiltonian Monte Carlo. In International Conference on Machine Learning, 2014.
  30. 30.Ciosek, K., Fortuin, V., Tomioka, R., Hofmann, K., and Turner, R. Conservative uncertainty estimation by fitting prior networks. In International Conference on Learning Representations, 2020.
  31. 31.Cohen, S. Bayesian analysis in natural language processing. Springer Nature, 2022.
  32. 32.Corani, G. and Mignatti, A. Credal model averaging for classification: representing prior ignorance and expert opinions. International Journal of Approximate Reasoning, 56:264–277, 2015.
  33. 33.Cranmer, M., Tamayo, D., Rein, H., Battaglia, P., Hadden, S., Armitage, P. J., Ho, S., and Spergel, D. N. A Bayesian neural network predicts the dissolution of compact planetary systems. Proceedings of the National Academy of Sciences, 118(40):e2026053118, 2021.
  34. 34.Damianou, A. and Lawrence, N. D. Deep Gaussian processes. In International Conference on Artificial Intelligence and Statistics, 2013.
  35. 35.D’Angelo, F. and Fortuin, V. Repulsive deep ensembles are Bayesian. Advances in Neural Information Processing Systems, 2021.
  36. 36.D’Angelo, F., Fortuin, V., and Wenzel, F. On Stein variational neural network ensembles. arXiv preprint arXiv:2106.10760, 2021.
  37. 37.Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P. Laplace redux - effortless Bayesian deep learning. In Advances in Neural Information Processing Systems, 2021a.
  38. 38.Daxberger, E., Nalisnick, E., Allingham, J. U., Antoran, J., and Hernandez-Lobato, J. M. Bayesian deep learning via subnetwork inference. In International Conference on Machine Learning, 2021b.
  39. 39.de G. Matthews, A. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z. Gaussian process behaviour in wide deep neural networks. International Conference on Learning Representations, 2018.
  40. 40.Detommaso, G., Gasparin, A., Donini, M., Seeger, M., Wilson, A. G., and Archambeau, C. Fortuna: A library for uncertainty quantification in deep learning. arXiv preprint arXiv:2302.04019, 2023.
  41. 41.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. QLoRA: Efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems, 2023.
  42. 42.Dhahri, R., Immer, A., Charpentier, B., Gunnemann, S., and Fortuin, V. Shaving weights with Occam’s razor: Bayesian sparsification for neural networks using the marginal likelihood. arXiv preprint arXiv:2402.15978, 2024.
  43. 43.Dold, D., Rugamer, D., Sick, B., and Durr, O. Semi-structured subspace inference. In International Conference on Artificial Intelligence and Statistics, 2024.
  44. 44.Ferreira, L., Conselice, C. J., Duncan, K., Cheng, T.-Y., Griffiths, A., and Whitney, A. Galaxy merger rates up to z ∼ 3 using a Bayesian deep learning model: A major-merger classifier using illustrisTNG simulation data. The Astrophysical Journal, 895(2):115, 2020.
  45. 45.Finzi, M., Benton, G., and Wilson, A. G. Residual pathway priors for soft equivariance constraints. In Advances in Neural Information Processing Systems, 2021.
  46. 46.Fortuin, V. Priors in Bayesian deep learning: a review. International Statistical Review, 90(3):563–591, 2022.
  47. 47.Fortuin, V., Garriga-Alonso, A., van der Wilk, M., and Aitchison, L. BNNpriors: A library for Bayesian neural network inference with different prior distributions. Software Impacts, 9:100079, 2021.
  48. 48.Fortuin, V., Garriga-Alonso, A., Ober, S. W., Wenzel, F., Ratsch, G., Turner, R. E., van der Wilk, M., and Aitchison, L. Bayesian neural network priors revisited. In International Conference on Learning Representations, 2022.
  49. 49.Gal, Y. and Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, 2016.
  50. 50.Gal, Y., Islam, R., and Ghahramani, Z. Deep Bayesian active learning with image data. In International Conference on Machine Learning, 2017.
  51. 51.Ganev, S. K. and Aitchison, L. Semi-supervised learning with a principled likelihood from a generative model of data curation. In International Conference on Learning Representations, 2023.
  52. 52.Garnelo, M., Rosenbaum, D., Maddison, C., Ramalho, T., Saxton, D., Shanahan, M., Teh, Y. W., Rezende, D., and Eslami, S. A. Conditional neural processes. In International Conference on Machine Learning, 2018.
  53. 53.Garriga-Alonso, A. and Fortuin, V. Exact Langevin dynamics with stochastic gradients. In Third Symposium on Advances in Approximate Bayesian Inference, 2021.
  54. 54.Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. Bayesian data analysis. Chapman and Hall/CRC, 3rd edition, 2013.
  55. 55.Gelman, A., Vehtari, A., Simpson, D., Margossian, C. C., Carpenter, B., Yao, Y., Kennedy, L., Gabry, J., Burkner, P.-C., and Modrak, M. Bayesian workflow. arXiv preprint arXiv:2011.01808, 2020.
  56. 56.George, E. I. Dilution priors: compensating for model space redundancy. In Borrowing Strength: Theory Powering Applications–A Festschrift for Lawrence D. Brown, volume 6, pp. 158–166. Institute of Mathematical Statistics, 2010.
  57. 57.Ghosh, S., Yao, J., and Doshi-Velez, F. Structured variational learning of bayesian neural networks with horseshoe priors. In International Conference on Machine Learning, 2018.
  58. 58.Goli, L., Reading, C., Sellan, S., Jacobson, A., and Tagliasacchi, A. Bayes’ rays: uncertainty quantification in neural radiance fields. Conference on Computer Vision and Pattern Recognition, 2024.
  59. 59.Gordon, J., Bruinsma, W. P., Foong, A. Y., Requeima, J., Dubois, Y., and Turner, R. E. Convolutional conditional neural processes. In International Conference on Learning Representations, 2019.
  60. 60.Graves, A. Practical variational inference for neural networks. In Advances in Neural Information Processing Systems, 2011.
  61. 61.Griffiths, R.-R., Klarner, L., Moss, H. B., Ravuri, A., Truong, S., Stanton, S., Tom, G., Rankovic, B., Du, Y., Jamasb, A., et al. GAUCHE: A library for Gaussian processes in chemistry. In Advances in Neural Information Processing Systems, 2023.
  62. 62.Grun, B. and Hofmarcher, P. Identifying groups of determinants in Bayesian model averaging using Dirichlet process clustering. Scandinavian Journal of Statistics, 48(3):1018–1045, 2021.
  63. 63.Grunwald, P. and Van Ommen, T. Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis, 12(4):1069–1103, 2017.
  64. 64.Gruver, N., Stanton, S., Kirichenko, P., Finzi, M., Maffettone, P., Myers, V., Delaney, E., Greenside, P., and Wilson, A. G. Effective surrogate models for protein design with Bayesian optimization. In ICML Workshop on Computational Biology, 2021.
  65. 65.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. arXiv preprint arXiv:2310.07820, 2023a.
  66. 66.Gruver, N., Stanton, S., Frey, N., Rudner, T. G., Hotzel, I., Lafrance-Vanasse, J., Rajpal, A., Cho, K., and Wilson, A. G. Protein design with guided discrete diffusion. Advances in Neural Information Processing Systems, 2023b.
  67. 67.Gu, Y. and Dunson, D. B. Bayesian pyramids: identifiable multilayer discrete latent structure models for discrete data. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(2):399–426, 2023.
  68. 68.Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. In International Conference on Machine Learning, 2017.
  69. 69.Gustafsson, F. K., Danelljan, M., and Schon, T. B. Evaluating scalable Bayesian deep learning methods for robust computer vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020.
  70. 70.Havasi, M., Peharz, R., and Hernandez-Lobato, J. M. Minimal random code learning: Getting bits back from compressed model parameters. In International Conference on Learning Representations, 2019.
  71. 71.He, B., Lakshminarayanan, B., and Teh, Y. W. Bayesian deep ensembles via the neural tangent kernel. In Advances in Neural Information Processing Systems, 2020.
  72. 72.Hennig, P., Osborne, M. A., and Kersting, H. P. Probabilistic Numerics: Computation as Machine Learning. Cambridge University Press, 2022.
  73. 73.Hernandez, S. and Lopez, J. L. Uncertainty quantification for plant disease detection using Bayesian deep learning. Applied Soft Computing, 96:106597, 2020.
  74. 74.Hobbhahn, M., Kristiadi, A., and Hennig, P. Fast predictive uncertainty for classification with Bayesian deep networks. In Conference on Uncertainty in Artificial Intelligence, 2022.
  75. 75.Hoeting, J. A., Madigan, D., Raftery, A. E., and Volinsky, C. T. Bayesian model averaging. In AAAI Workshop on Integrating Multiple Learned Models, 1998.
  76. 76.Hoeting, J. A., Madigan, D., Raftery, A. E., and Volinsky, C. T. Bayesian model averaging: a tutorial. Statistical Science, 14(4):382–417, 1999.
  77. 77.Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J. Stochastic variational inference. Journal of Machine Learning Research, 14(1):1303–1347, May 2013.
  78. 78.Hubin, A. and Storvik, G. Combining model and parameter uncertainty in Bayesian neural networks. arXiv preprint arXiv:1903.07594, 2019.
  79. 79.Hubin, A. and Storvik, G. Sparse Bayesian neural networks: bridging model and parameter uncertainty through scalable variational inference. Mathematics, 12(6):788, 2024.
  80. 80.Hubin, A., Storvik, G., and Frommlet, F. Flexible Bayesian nonlinear model configuration. Journal of Artificial Intelligence Research, 72:901–942, 2021.
  81. 81.Ibrahim, J. G. and Laud, P. W. On Bayesian analysis of generalized linear models using Jeffreys’s prior. Journal of the American Statistical Association, 86(416):981–986, 1991.
  82. 82.Immer, A., Bauer, M., Fortuin, V., Ratsch, G., and Khan, M. E. Scalable marginal likelihood estimation for model selection in deep learning. In International Conference on Machine Learning, 2021a.
  83. 83.Immer, A., Korzepa, M., and Bauer, M. Improving predictions of Bayesian neural nets via local linearization. In International Conference on Artificial Intelligence and Statistics, 2021b.
  84. 84.Immer, A., Hennigen, L. T., Fortuin, V., and Cotterell, R. Probing as quantifying inductive bias. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, pp. 1839–1851, 2022a.
  85. 85.Immer, A., van der Ouderaa, T., Ratsch, G., Fortuin, V., and van der Wilk, M. Invariance learning in deep neural networks with differentiable Laplace approximations. Advances in Neural Information Processing Systems, 2022b.
  86. 86.Immer, A., Van Der Ouderaa, T. F., Van Der Wilk, M., Ratsch, G., and Scholkopf, B. Stochastic marginal likelihood gradients using neural tangent kernels. In International Conference on Machine Learning, 2023.
  87. 87.Izmailov, P., Maddox, W. J., Kirichenko, P., Garipov, T., Vetrov, D., and Wilson, A. G. Subspace inference for Bayesian deep learning. In Conference on Uncertainty in Artificial Intelligence, 2020.
  88. 88.Izmailov, P., Nicholson, P., Lotfi, S., and Wilson, A. G. Dangers of Bayesian model averaging under covariate shift. Advances in Neural Information Processing Systems, 2021a.
  89. 89.Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. G. What are Bayesian neural network posteriors really like? In International Conference on Machine Learning, 2021b.
  90. 90.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, 2018.
  91. 91.Janz, D., Hron, J., Mazur, P., Hofmann, K., Hernandez-Lobato, J. M., and Tschiatschek, S. Successor uncertainties: exploration and uncertainty in temporal difference learning. Advances in Neural Information Processing Systems, 32, 2019.
  92. 92.Jaynes, E. T. Probability theory: The logic of science. Cambridge University Press, 2003.
  93. 93.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 2023.
  94. 94.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  95. 95.Kapoor, S., Maddox, W. J., Izmailov, P., and Wilson, A. G. On uncertainty, tempering, and data augmentation in Bayesian classification. Advances in Neural Information Processing Systems, 2022.
  96. 96.Karaletsos, T. and Bui, T. D. Hierarchical Gaussian process priors for Bayesian neural network weights. In Advances in Neural Information Processing Systems, 2020.
  97. 97.Karaletsos, T., Dayan, P., and Ghahramani, Z. Probabilistic meta-representations of neural networks. arXiv preprint arXiv:1810.00555, 2018.
  98. 98.Khan, M. E. and Rue, H. The Bayesian learning rule. Journal of Machine Learning Research, 24(281):1–46, 2023.
  99. 99.Khan, M. E., Immer, A., Abedi, E., and Korzepa, M. Approximate inference turns deep networks into Gaussian processes. In Advances in Neural Information Processing Systems, 2019.
  100. 100.Kirichenko, P., Izmailov, P., and Wilson, A. G. Last layer re-training is sufficient for robustness to spurious correlations. In International Conference on Learning Representations, 2023.
  101. 101.Klarner, L., Rudner, T. G. J., Reutlinger, M., Schindler, T., Morris, G. M., Deane, C., and Teh, Y. W. Drug Discovery under Covariate Shift with Domain-Informed Prior Distributions over Functions. In International Conference on Machine Learning, 2023.
  102. 102.Kolb, C., Muller, C. L., Bischl, B., and Rugamer, D. Smoothing the edges: A general framework for smooth optimization in sparse regularization using Hadamard over-parametrization. arXiv preprint arXiv:2307.03571, 2023.
  103. 103.Kou, S., Gan, L., Wang, D., Li, C., and Deng, Z. BayesDiff: estimating pixel-wise uncertainty in diffusion via Bayesian inference. In International Conference on Learning Representations, 2024.
  104. 104.Kristiadi, A., Hein, M., and Hennig, P. Being Bayesian, even just a bit, fixes overconfidence in ReLU networks. In International Conference on Machine Learning, 2020.
  105. 105.Kristiadi, A., Hein, M., and Hennig, P. An infinite-feature extension for Bayesian ReLU nets that fixes their asymptotic overconfidence. In Advances in Neural Information Processing Systems, 2021a.
  106. 106.Kristiadi, A., Hein, M., and Hennig, P. Learnable uncertainty under Laplace approximations. In Conference on Uncertainty in Artificial Intelligence, 2021b.
  107. 107.Kristiadi, A., Hein, M., and Hennig, P. Being a bit frequentist improves Bayesian neural networks. In International Conference on Artificial Intelligence and Statistics, 2022.
  108. 108.Kristiadi, A., Dangel, F., and Hennig, P. The geometry of neural nets’ parameter spaces under reparametrization. In Advances in Neural Information Processing Systems, 2023.
  109. 109.Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., and Pleiss, G. A sober look at LLMs for material discovery: are they actually good for Bayesian optimization over molecules? arXiv preprint arXiv:2402.05015, 2024.
  110. 110.Krueger, D., Huang, C.-W., Islam, R., Turner, R., Lacoste, A., and Courville, A. Bayesian hypernetworks. arXiv preprint arXiv:1710.04759, 2017.
  111. 111.Lakshminarayanan, B., Pritzel, A., and Blundell, C. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in Neural Information Processing Systems, 2017.
  112. 112.Lampinen, J. and Vehtari, A. Bayesian approach for neural networks—review and case studies. Neural networks, 14(3):257–274, 2001.
  113. 113.Langford, J. and Shawe-Taylor, J. PAC-Bayes & margins. In Advances in Neural Information Processing Systems, 2002.
  114. 114.Lawrence, N. D. Variational inference in probabilistic models. PhD thesis, University of Cambridge, 2001.
  115. 115.Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J. Deep neural networks as Gaussian processes. In International Conference on Learning Representations, 2017.
  116. 116.Leitherer, A., Ziletti, A., and Ghiringhelli, L. M. Robust recognition and exploratory analysis of crystal structures via Bayesian deep learning. Nature Communications, 12(1):6234, 2021.
  117. 117.Li, A., Boyd, A., Smyth, P., and Mandt, S. Detecting and adapting to irregular distribution shifts in Bayesian online learning. Advances in Neural Information Processing Systems, 2021.
  118. 118.Li, J., Miao, Z., Qiu, Q., and Zhang, R. Training Bayesian neural networks with sparse subspace variational inference. In International Conference on Learning Representations, 2024.
  119. 119.Li, Y. and Clyde, M. A. Mixtures of g-priors in generalized linear models. Journal of the American Statistical Association, 113(524):1828–1845, 2018.
  120. 120.Li, Y. L., Rudner, T. G., and Wilson, A. G. A study of Bayesian neural network surrogates for Bayesian optimization. arXiv preprint arXiv:2305.20028, 2023.
  121. 121.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In International Conference on Learning Representations, 2022.
  122. 122.Liu, Q. and Wang, D. Stein variational gradient descent: A general purpose Bayesian inference algorithm. Advances in Neural Information Processing Systems, 2016.
  123. 123.Lopez, J. L., Rudner, T. G. J., and Shamout, F. Informative priors improve the reliability of multimodal clinical data classification. In Machine Learning for Health Symposium Findings, 2023.
  124. 124.Lotfi, S., Izmailov, P., Benton, G., Goldblum, M., and Wilson, A. G. Bayesian model selection, the marginal likelihood, and generalization. In International Conference on Machine Learning, pp. 14223–14247, 2022.
  125. 125.Louizos, C., Ullrich, K., and Welling, M. Bayesian compression for deep learning. Advances in Neural Information Processing Systems, 2017.
  126. 126.Luo, X., Nadiga, B. T., Park, J. H., Ren, Y., Xu, W., and Yoo, S. A Bayesian deep learning approach to near-term climate prediction. Journal of Advances in Modeling Earth Systems, 14(10):e2022MS003058, 2022.
  127. 127.MacKay, D. J. Bayesian interpolation. Neural Computation, 4(3):415–447, 1992.
  128. 128.MacKay, D. J. Bayesian neural networks and density networks. Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 354(1):73–80, 1995.
  129. 129.MacKay, D. J. Choice of basis for Laplace approximation. Machine Learning, 33:77–86, 1998.
  130. 130.Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G. A simple baseline for Bayesian uncertainty in deep learning. Advances in Neural Information Processing Systems, 32, 2019.
  131. 131.Mandt, S., Hoffman, M. D., and Blei, D. M. Stochastic gradient descent as approximate Bayesian inference. Journal of Machine Learning Research, 18(134):1–35, 2017.
  132. 132.Manogaran, G., Shakeel, P. M., Fouad, H., Nam, Y., Baskar, S., Chilamkurti, N., and Sundarasekar, R. Wearable IoT smart-log patch: An edge computing-based Bayesian deep learning network system for multi access physical monitoring system. Sensors, 19(13):3030, 2019.
  133. 133.Mansinghka, V. K. Natively probabilistic computation. PhD thesis, Massachusetts Institute of Technology, Department of Brain and Cognitive Sciences, 2009.
  134. 134.Margatina, K., Barrault, L., and Aletras, N. On the importance of effectively adapting pretrained language models for active learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022.
  135. 135.Margatina, K., Schick, T., Aletras, N., and Dwivedi-Yu, J. Active learning principles for in-context learning with large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, 2023.
  136. 136.Martens, J. Deep learning via Hessian-free optimization. In International Conference on Machine Learning, 2010.
  137. 137.McAllister, R., Gal, Y., Kendall, A., Van Der Wilk, M., Shah, A., Cipolla, R., and Weller, A. Concrete problems for autonomous vehicle safety: Advantages of Bayesian deep learning. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 2017.
  138. 138.Meister, J. A., Nguyen, K. A., Kapetanakis, S., and Luo, Z. A novel deep learning approach for one-step conformal prediction approximation. Annals of Mathematics and Artificial Intelligence, pp. 1–28, 2023.
  139. 139.Milsom, E., Anson, B., and Aitchison, L. Convolutional deep kernel machines. arXiv preprint arXiv:2309.09814, 2023.
  140. 140.Molchanov, D., Ashukha, A., and Vetrov, D. Variational dropout sparsifies deep neural networks. In International Conference on Machine Learning, 2017.
  141. 141.Moor, M., Banerjee, O., Abad, Z. S. H., Krumholz, H. M., Leskovec, J., Topol, E. J., and Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265, 2023.
  142. 142.Muller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F. Transformers can do Bayesian inference. In International Conference on Learning Representations, 2021.
  143. 143.Mur-Labadia, L., Martinez-Cantin, R., and Guerrero, J. J. Bayesian deep learning for affordance segmentation in images. arXiv preprint arXiv:2303.00871, 2023.
  144. 144.Murphy, K. P. Probabilistic machine learning: Advanced topics. MIT Press, 2023.
  145. 145.Nabarro, S., Ganev, S., Garriga-Alonso, A., Fortuin, V., van der Wilk, M., and Aitchison, L. Data augmentation in Bayesian neural networks and the cold posterior effect. In Uncertainty in Artificial Intelligence, 2022.
  146. 146.Neal, R. M. Priors for infinite networks. Bayesian learning for neural networks, pp. 29–53, 1996.
  147. 147.Neklyudov, K., Molchanov, D., Ashukha, A., and Vetrov, D. P. Structured Bayesian pruning via log-normal multiplicative noise. In Advances in Neural Information Processing Systems, 2017.
  148. 148.Neklyudov, K., Molchanov, D., Ashukha, A., and Vetrov, D. Variance networks: when expectation does not meet your expectations. In International Conference on Learning Representations, 2018.
  149. 149.Nemeth, C. and Fearnhead, P. Stochastic gradient Markov chain Monte Carlo. Journal of the American Statistical Association, 116(533):433–450, 2021.
  150. 150.Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E. Variational continual learning. In International Conference on Learning Representations, 2018.
  151. 151.Oates, C. J. and Sullivan, T. J. A modern retrospective on probabilistic numerics. Statistics and Computing, 29(6):1335–1351, 2019.
  152. 152.Ober, S. and Aitchison, L. A variational approximate posterior for the deep Wishart process. Advances in Neural Information Processing Systems, 2021a.
  153. 153.Ober, S. W. and Aitchison, L. Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes. In International Conference on Machine Learning, 2021b.
  154. 154.Ober, S. W., Anson, B., Milsom, E., and Aitchison, L. An improved variational approximate posterior for the deep Wishart process. In Conference on Uncertainty in Artificial Intelligence, 2023.
  155. 155.Onal, E., Floge, K., Caldwell, E., Sheverdin, A., and Fortuin, V. Gaussian stochastic weight averaging for Bayesian low-rank adaptation of large language models. arXiv preprint arXiv:2405.03425, 2024.
  156. 156.Osband, I., Wen, Z., Asghari, S. M., Dwaracherla, V., Lu, X., Ibrahimi, M., Lawson, D., Hao, B., O’Donoghue, B., and Van Roy, B. The neural testbed: Evaluating joint predictions. Advances in Neural Information Processing Systems, 2022.
  157. 157.Osband, I., Wen, Z., Asghari, S. M., Dwaracherla, V., Ibrahimi, M., Lu, X., and Van Roy, B. Approximate Thompson sampling via epistemic neural networks. In Conference on Uncertainty in Artificial Intelligence, 2023.
  158. 158.Papadopoulos, H., Vovk, V., and Gammerman, A. Conformal prediction with neural networks. In International Conference on Tools with Artificial Intelligence, 2007.
  159. 159.Papamarkou, T. Approximate blocked Gibbs sampling for Bayesian neural networks. Statistics and Computing, 33, 2023.
  160. 160.Papamarkou, T., Hinkle, J., Young, M. T., and Womble, D. Challenges in Markov chain Monte Carlo for Bayesian neural networks. Statistical Science, 37(3):425–442, 2022.
  161. 161.Parrado-Hernandez, E., Ambroladze, A., Shawe-Taylor, J., and Sun, S. PAC-Bayes bounds with data dependent priors. Journal of Machine Learning Research, 13(1):3507–3531, 2012.
  162. 162.Patacchiola, M., Turner, J., Crowley, E. J., O' Boyle, M., and Storkey, A. J. Bayesian meta-learning for the few-shot setting via deep kernels. In Advances in Neural Information Processing Systems, 2020.
  163. 163.Pearce, T., Leibfried, F., and Brintrup, A. Uncertainty in neural networks: Approximately Bayesian ensembling. In International Conference on Artificial Intelligence and Statistics, pp. 234–244. PMLR, 2020.
  164. 164.Peng, W., Ye, Z.-S., and Chen, N. Bayesian deep-learning-based health prognostics toward prognostics uncertainty. IEEE Transactions on Industrial Electronics, 67(3):2283–2293, 2019.
  165. 165.Pielok, T., Bischl, B., and Rugamer, D. Approximate Bayesian inference with Stein functional variational gradient descent. In International Conference on Learning Representations, 2022.
  166. 166.Polson, N. G. and Rockova, V. Posterior concentration for sparse deep learning. Advances in Neural Information Processing Systems, 2018.
  167. 167.Qiu, S., Rudner, T. G. J., Kapoor, S., and Wilson, A. G. Should we learn most likely functions or parameters? In Advances in Neural Information Processing Systems, 2023.
  168. 168.Rainforth, T., Foster, A., Ivanova, D. R., and Bickford Smith, F. Modern bayesian experimental design. Statistical Science, 39(1):100–114, 2024.
  169. 169.Ritter, H. and Karaletsos, T. TyXe: Pyro-based Bayesian neural nets for Pytorch. In Proceedings of Machine Learning and Systems, 2022.
  170. 170.Ritter, H., Botev, A., and Barber, D. A scalable Laplace approximation for neural networks. In International Conference on Learning Representations, 2018.
  171. 171.Ritter, H., Kukla, M., Zhang, C., and Li, Y. Sparse uncertainty representation in deep learning with inducing weights. In Advances in Neural Information Processing Systems, 2021.
  172. 172.Robbins, H. E. A stochastic approximation method. Annals of Mathematical Statistics, 22:400–407, 1951.
  173. 173.Rothfuss, J., Fortuin, V., Josifoski, M., and Krause, A. PACOH: Bayes-optimal meta-learning with PAC-guarantees. In International Conference on Machine Learning, 2021.
  174. 174.Rothfuss, J., Josifoski, M., Fortuin, V., and Krause, A. PAC-Bayesian meta-learning: From theory to practice. arXiv preprint arXiv:2211.07206, 2022.
  175. 175.Rudner, T. G. J., Chen, Z., Teh, Y. W., and Gal, Y. Tractable function-space variational inference in Bayesian neural networks. In Advances in Neural Information Processing Systems, 2022a.
  176. 176.Rudner, T. G. J., Smith, F. B., Feng, Q., Teh, Y. W., and Gal, Y. Continual Learning via Sequential Function-Space Variational Inference. In International Conference on Machine Learning, 2022b.
  177. 177.Rudner, T. G. J., Kapoor, S., Qiu, S., and Wilson, A. G. Function-Space Regularization in Neural Networks: A Probabilistic Perspective. In International Conference on Machine Learning, 2023.
  178. 178.Rudner, T. G. J., Pan, X., Li, Y. L., Shwartz-Ziv, R., and Wilson, A. G. Uncertainty-aware priors for finetuning pretrained models. In Preprint, 2024a.
  179. 179.Rudner, T. G. J., Zhang, Y. S., Wilson, A. G., and Kempe, J. Mind the GAP: Improving robustness to subpopulation shifts with group-aware priors. In International Conference on Artificial Intelligence and Statistics, 2024b.
  180. 180.Rugamer, D. A new PHO-rmula for improved performance of semi-structured networks. In International Conference on Machine Learning, 2023.
  181. 181.Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., and Wen, Z. A tutorial on Thompson sampling. Foundations and Trends® in Machine Learning, 11(1):1–96, 2018.
  182. 182.Sam, D., Pukdee, R., Jeong, D. P., Byun, Y., and Kolter, J. Z. Bayesian neural networks with domain knowledge priors. arXiv preprint arXiv:2402.13410, 2024.
  183. 183.Schraudolph, N. N., Yu, J., and Gunter, S. A stochastic quasi-Newton method for online convex optimization. In International Conference on Artificial Intelligence and Statistics, 2007.
  184. 184.Schwobel, P., Jørgensen, M., Ober, S. W., and Van Der Wilk, M. Last layer marginal likelihood for invariance learning. In International Conference on Artificial Intelligence and Statistics, 2022.
  185. 185.Sen, D., Papamarkou, T., and Dunson, D. Bayesian neural networks and dimensionality reduction. In Handbook of Bayesian, fiducial, and frequentist inference. Chapmann and Hall/CRC Press, 2024.
  186. 186.Shah, A., Wilson, A., and Ghahramani, Z. Student-t processes as alternatives to Gaussian processes. In International Conference on Artificial Intelligence and Statistics, 2014.
  187. 187.Sharma, M., Rainforth, T., Teh, Y. W., and Fortuin, V. Incorporating unlabelled data into Bayesian neural networks. arXiv preprint arXiv:2304.01762, 2023.
  188. 188.Shen, Y., Daheim, N., Cong, B., Nickl, P., Marconi, G. M., Bazan, C., Yokota, R., Gurevych, I., Cremers, D., and Khan, M. E. Variational learning is effective for large deep networks. In International Conference on Machine Learning, 2024.
  189. 189.Shi, L., Copot, C., and Vanlanduit, S. A Bayesian deep neural network for safe visual servoing in human–robot interaction. Frontiers in Robotics and AI, 8:687031, 2021.
  190. 190.Shwartz-Ziv, R., Goldblum, M., Souri, H., Kapoor, S., Zhu, C., LeCun, Y., and Wilson, A. G. Pre-train your loss: Easy Bayesian transfer learning with informative priors. In Advances in Neural Information Processing Systems, 2022.
  191. 191.Skaaret-Lund, L., Storvik, G., and Hubin, A. Sparsifying Bayesian neural networks with latent binary variables and normalizing flows. arXiv preprint arXiv:2305.03395, 2023.
  192. 192.Soboczenski, F., Himes, M. D., O’Beirne, M. D., Zorzan, S., Baydin, A. G., Cobb, A. D., Gal, Y., Angerhausen, D., Mascaro, M., Arney, G. N., et al. Bayesian deep learning for exoplanet atmospheric retrieval. arXiv preprint arXiv:1811.03390, 2018.
  193. 193.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2020.
  194. 194.Stanton, S., Maddox, W., Gruver, N., Maffettone, P., Delaney, E., Greenside, P., and Wilson, A. G. Accelerating Bayesian optimization for biological sequence design with denoising autoencoders. In International Conference on Machine Learning, 2022.
  195. 195.Strieth-Kalthoff, F., Hao, H., Rathore, V., Derasp, J., Gaudin, T., Angello, N. H., Seifrid, M., Trushina, E., Guy, M., Liu, J., and et al. Delocalized, asynchronous, closed-loop discovery of organic laser emitters. ChemRxiv, 2023.
  196. 196.Sun, S., Zhang, G., Shi, J., and Grosse, R. Functional variational Bayesian neural networks. In International Conference on Learning Representations, 2019.
  197. 197.Sunde, L.-M. S. Spherical priors for Bayesian deep learning. Master’s thesis, University of Oslo, 2023.
  198. 198.Tatzel, L., Wenger, J., Schneider, F., and Hennig, P. Accelerating generalized linear models by trading off computation for uncertainty. arXiv preprint arXiv:2310.20285, 2023.
  199. 199.Tom, G., Hickman, R. J., Zinzuwadia, A., Mohajeri, A., Sanchez-Lengeling, B., and Aspuru-Guzik, A. Calibration and generalizability of probabilistic models on low-data chemical datasets with DIONYSUS. Digital Discovery, 2023.
  200. 200.Tran, B.-H., Rossi, S., Milios, D., and Filippone, M. All you need is a good functional prior for Bayesian deep learning. Journal of Machine Learning Research, 23(1), 2022a.
  201. 201.Tran, D., Liu, J., Dusenberry, M. W., Phan, D., Collier, M., Ren, J., Han, K., Wang, Z., Mariet, Z., Hu, H., Band, N., Rudner, T. G. J., Singhal, K., Nado, Z., van Amersfoort, J., Kirsch, A., Jenatton, R., Thain, N., Yuan, H., Buchanan, K., Murphy, K., Sculley, D., Gal, Y., Ghahramani, Z., Snoek, J., and Lakshminarayanan, B. Plex: Towards Reliability Using Pretrained Large Model Extensions. In ICML 2022 Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward, 2022b.
  202. 202.Tran, M.-N., Nguyen, N., Nott, D., and Kohn, R. Bayesian deep net GLM and GLMM. Journal of Computational and Graphical Statistics, 29(1):97–113, 2020.
  203. 203.Tump, A. N., Wolf, M., Romanczuk, P., and Kurvers, R. H. J. M. Avoiding costly mistakes in groups: the evolution of error management in collective decision making. PLOS Computational Biology, 18(8):1–21, 2022.
  204. 204.van der Ouderaa, T. F., Immer, A., and van der Wilk, M. Learning layer-wise equivariances automatically using gradients. In Advances in Neural Information Processing Systems, 2023.
  205. 205.Vandal, T., Kodra, E., Dy, J., Ganguly, S., Nemani, R., and Ganguly, A. R. Quantifying uncertainty in discrete-continuous and skewed data with Bayesian deep learning. In International Conference on Knowledge Discovery & Data Mining, 2018.
  206. 206.Vehtari, A., Gelman, A., Simpson, D., Carpenter, B., and Burkner, P.-C. Rank-normalization, folding, and localization: An improved Rb for assessing convergence of MCMC. Bayesian Analysis, 16(2):667–718, 2021.
  207. 207.Villani, C. Topics in optimal transportation, volume 58. American Mathematical Soc., 2021.
  208. 208.Vladimirova, M., Verbeek, J., Mesejo, P., and Arbel, J. Understanding priors in Bayesian neural networks at the unit level. In International Conference on Machine Learning, 2019.
  209. 209.Vladimirova, M., Arbel, J., and Girard, S. Bayesian neural network unit priors and generalized Weibull-tail property. In Asian Conference on Machine Learning, 2021.
  210. 210.Vovk, V., Gammerman, A., and Shafer, G. Algorithmic learning in a random world, volume 29. Springer, 2005.
  211. 211.Wainwright, M. J. and Jordan, M. I. Graphical Models, Exponential Families, and Variational Inference. Now Publishers Inc., Hanover, MA, USA, 2008.
  212. 212.Wang, Y., Rudner, T. G. J., and Wilson, A. G. Visual explanations of image-text representations via multi-modal information bottleneck attribution. In Advances in Neural Information Processing Systems, 2023a.
  213. 213.Wang, Z., Chen, Y., Song, Q., and Zhang, R. Enhancing low-precision sampling via stochastic gradient Hamiltonian Monte Carlo. arXiv preprint arXiv:2310.16320, 2023b.
  214. 214.Wasserman, L. Bayesian model selection and model averaging. Journal of Mathematical Psychology, 44(1):92–107, 2000.
  215. 215.Watanabe, S. Algebraic Geometry and Statistical Learning Theory. Cambridge University Press, USA, 2009.
  216. 216.Way, G. P. and Greene, C. S. Bayesian deep learning for single-cell analysis. Nature Methods, 15(12):1009–1010, 2018.
  217. 217.Wei, S. and Lau, E. Variational Bayesian neural networks via resolution of singularities. arXiv preprint arXiv:2302.06035, 2023.
  218. 218.Welling, M. and Teh, Y. W. Bayesian learning via stochastic gradient Langevin dynamics. In International Conference on Machine Learning, 2011.
  219. 219.Wen, Z., Osband, I., Qin, C., Lu, X., Ibrahimi, M., Dwaracherla, V., Asghari, M., and Van Roy, B. From predictions to decisions: The importance of joint predictive distributions. arXiv preprint arXiv:2107.09224, 2021.
  220. 220.Wenger, J., Pleiss, G., Pfortner, M., Hennig, P., and Cunningham, J. P. Posterior and computational uncertainty in Gaussian processes. In Advances in Neural Information Processing Systems, 2022.
  221. 221.Wenzel, F., Roth, K., Veeling, B., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S. How good is the Bayes posterior in deep neural networks really? In International Conference on Machine Learning, 2020.
  222. 222.Wiese, J. G., Wimmer, L., Papamarkou, T., Bischl, B., Gunnemann, S., and Rugamer, D. Towards efficient MCMC sampling in Bayesian neural networks by exploiting symmetry. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2023.
  223. 223.Wild, V. D., Ghalebikesabi, S., Sejdinovic, D., and Knoblauch, J. A rigorous link between deep ensembles and (variational) Bayesian methods. In Conference on Neural Information Processing Systems, 2023.
  224. 224.Williams, P. M. Bayesian regularization and pruning using a Laplace prior. Neural Computation, 7(1):117–143, 1995.
  225. 225.Wilson, A. G. and Izmailov, P. Bayesian deep learning and a probabilistic perspective of generalization. Advances in Neural Information Processing Systems, 2020.
  226. 226.Wilson, A. G., Knowles, D. A., and Ghahramani, Z. Gaussian process regression networks. In International Conference on Machine Learning, 2012.
  227. 227.Wilson, A. G., Hu, Z., Salakhutdinov, R., and Xing, E. P. Deep kernel learning. In International Conference on Artificial Intelligence and Statistics, 2016.
  228. 228.Wilson, A. G., Izmailov, P., Hoffman, M. D., Gal, Y., Li, Y., Pradier, M. F., Vikram, S., Foong, A., Lotfi, S., and Farquhar, S. Evaluating approximate inference in bayesian deep learning. In NeurIPS 2021 Competitions and Demonstrations Track, 2022.
  229. 229.Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. An explanation of in-context learning as implicit Bayesian inference. In International Conference on Learning Representations, 2021.
  230. 230.Yang, A. X., Robeyns, M., Milsom, E., Anson, B., Schoots, N., and Aitchison, L. A theory of representation learning gives a deep generalisation of kernel methods. In International Conference on Machine Learning, 2023a.
  231. 231.Yang, A. X., Robeyns, M., Coste, T., Wang, J., Bou-Ammar, H., and Aitchison, L. Bayesian reward models for LLM alignment. arXiv preprint arXiv:2402.13210, 2024a.
  232. 232.Yang, A. X., Robeyns, M., Wang, X., and Aitchison, L. Bayesian low-rank adaptation for large language models. International Conference on Learning Representations, 2024b.
  233. 233.Yang, Y., Li, W., Gulliver, T. A., and Li, S. Bayesian deep learning-based probabilistic load forecasting in smart grids. IEEE Transactions on Industrial Informatics, 16(7):4703–4713, 2019.
  234. 234.Yang, Y., Bamler, R., and Mandt, S. Variational Bayesian quantization. In International Conference on Machine Learning, 2020.
  235. 235.Yang, Y., Hui, B., Yuan, H., Gong, N., and Cao, Y. SneakyPrompt: Evaluating robustness of text-to-image generative models’ safety filters. In IEEE Symposium on Security and Privacy, 2023b.
  236. 236.Yang, Y., Mandt, S., and Theis, L. An introduction to neural data compression. Foundations and Trends® in Computer Graphics and Vision, 15(2):113–200, 2023c.
  237. 237.Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S. Bayesian model-agnostic meta-learning. In Advances in Neural Information Processing Systems, 2018.
  238. 238.Zellner, A. On assessing prior distributions and Bayesian regression analysis with g-prior distributions. Bayesian Inference and Decision Techniques, 1986.
  239. 239.Zhang, J., Jennings, J., Zhang, C., and Ma, C. Towards causal foundation model: on duality between causal inference and attention. arXiv preprint arXiv:2310.00809, 2023.
  240. 240.Zhang, R., Cooper, A. F., and De Sa, C. AMAGOLD: amortized Metropolis adjustment for efficient stochastic gradient MCMC. In International Conference on Artificial Intelligence and Statistics, 2020a.
  241. 241.Zhang, R., Li, C., Zhang, J., Chen, C., and Wilson, A. G. Cyclical stochastic gradient MCMC for Bayesian deep learning. In International Conference on Learning Representations, 2020b.
  242. 242.Zhang, R., Wilson, A. G., and De Sa, C. Low-precision stochastic gradient Langevin dynamics. In International Conference on Machine Learning, 2022a.
  243. 243.Zhang, X., Li, Y., Li, W., Guo, K., and Shao, Y. Personalized federated learning via variational Bayesian inference. In International Conference on Machine Learning, 2022b.
  244. 244.Zhao, S., Kim, M. P., Sahoo, R., Ma, T., and Ermon, S. Calibrating predictions to decisions: a novel approach to multi-class calibration. In Advances in Neural Information Processing Systems, 2021.
  245. 245.Zhou, Z., Yu, H., and Shi, H. Human activity recognition based on improved Bayesian convolution network to analyze health care data using wearable iot device. IEEE Access, 8:86411–86418, 2020.

Citation

MLA
Papamarkou, T., et al. “Position: Bayesian Deep Learning Is Needed in the Age of Large-Scale AI”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.00809.
APA
Papamarkou, T., Skoularidou, M., Palla, K., Aitchison, L., Arbel, J., Dunson, D., Filippone, M., Fortuin, V., Hennig, P., Hernández-Lobato, J. M., Hubin, A., Immer, A., Karaletsos, T., Khan, M. E., Kristiadi, A., Li, Y., Mandt, S., Nemeth, C., Osborne, M. A., … Zhang, R. (2024). Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI. arXiv. https://doi.org/10.48550/arxiv.2402.00809
Chicago
Papamarkou, T., M. Skoularidou, K. Palla, et al. 2024. “Position: Bayesian Deep Learning Is Needed in the Age of Large-Scale AI”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.00809.
Harvard
Papamarkou, T. et al. (2024) “Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.00809.
Vancouver
1. Papamarkou T, Skoularidou M, Palla K, et al (2024) Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI. https://doi.org/10.48550/arxiv.2402.00809

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.00809,
  doi = {10.48550/ARXIV.2402.00809},
  url = {https://arxiv.org/abs/2402.00809},
  author = {Papamarkou, Theodore and Skoularidou, Maria and Palla, Konstantina and Aitchison, Laurence and Arbel, Julyan and Dunson, David and Filippone, Maurizio and Fortuin, Vincent and Hennig, Philipp and Hernández-Lobato, José Miguel and Hubin, Aliaksandr and Immer, Alexander and Karaletsos, Theofanis and Khan, Mohammad Emtiyaz and Kristiadi, Agustinus and Li, Yingzhen and Mandt, Stephan and Nemeth, Christopher and Osborne, Michael A. and Rudner, Tim G. J. and Rügamer, David and Teh, Yee Whye and Welling, Max and Wilson, Andrew Gordon and Zhang, Ruqi},
  keywords = {Machine Learning (cs.LG), Machine Learning (stat.ML), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution Non Commercial No Derivatives 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/