All You Need is a Good Functional Prior for Bayesian Deep Learning
Ba-Hien TranSimone RossiDimitrios MiliosMaurizio Filippone
Proposes a method to align Bayesian neural network weight priors with interpretable Gaussian process functional priors by minimizing their Wasserstein distance, significantly boosting predictive performance across regression and classification tasks.
Deep neural networks are central to modern machine learning, yet standard deployments struggle to quantify predictive uncertainty reliably. While Bayesian approaches offer a principled way to represent uncertainty by placing prior distributions over parameters, modern networks contain thousands to millions of weights. In practice, standard independent Gaussian priors placed on these parameters induce pathological and uncontrollable behaviors in function space, such as collapsing to flat, uninformative functions in deep architectures or assigning extreme, overconfident probabilities to single classes. This inability to establish interpretable and well-behaved prior beliefs has severely limited the reliability and adoption of Bayesian deep learning in high-stakes settings.
The article introduces and evaluates a practical framework to tune the parameter priors of Bayesian neural networks so that their induced function-level distributions match interpretable target functional priors, specifically Gaussian processes. By optimizing this alignment directly in function space before observing training data, the authors demonstrate that Bayesian neural networks can achieve superior predictive accuracy, robust uncertainty calibration, and resilience to data corruption.
The approach formulates prior selection as a sample-based distance minimization problem using the Wasserstein distance, optimized via an alternating gradient-based algorithm. The authors explore parameterizations of increasing flexibility, including layer-wise Gaussian distributions, hierarchical Inverse-Gamma distributions, and normalizing flows. The target functional properties are specified via standard and hierarchical Gaussian processes. Empirical credibility is established across a comprehensive benchmark suite: standard regression and classification datasets from the UCI repository, active learning tasks, and deep vision architectures (including LeNet-5, VGG-16, and PreResNet-20) evaluated on MNIST, CIFAR-10, and corrupted CIFAR-10C datasets. Posterior inference is primarily conducted using scale-adapted stochastic gradient Hamiltonian Monte Carlo sampling.
The investigation yields several key findings. First, optimizing parameter priors via Wasserstein distance consistently aligns neural network outputs with desired functional behaviors, avoiding the optimization instabilities seen in Kullback-Leibler divergence approaches. Second, Gaussian process-induced priors deliver systematically superior predictive performance across benchmarks; on CIFAR-10 image classification, the hierarchical induced prior achieved top accuracies of 76.51% on LeNet-5, 87.03% on VGG-16, and 88.20% on PreResNet-20, outperforming standard priors, posterior temperature scaling, and popular non-Bayesian deep ensembles. Third, models trained with these functional priors exhibit superior calibration and robustness under severe covariate shift and out-of-distribution inputs, maintaining high predictive entropy on unfamiliar data rather than making overconfident erroneous predictions. Fourth, in active learning scenarios, the proposed priors guided faster, more sample-efficient data acquisition than fixed prior alternatives.
These findings indicate that prior specification—rather than approximate Bayesian inference itself—has been the critical missing link in Bayesian deep learning. For decision-makers, adopting functional priors significantly reduces operational risks associated with model overconfidence and data corruption without requiring costly ad-hoc interventions like data-driven cross-validation or heuristic posterior tempering. Furthermore, this prior optimization is performed completely prior to training, preserving Bayesian integrity while reducing computational exploration time compared to massive hyperparameter grid searches.
Organizations deploying deep neural networks in safety-critical, active learning, or distributionally shifting environments should integrate functional prior optimization into their Bayesian modeling pipelines. When selecting prior parameterizations, practitioners face a trade-off: layer-wise hierarchical priors scale easily to deep convolutional models and deliver the strongest overall performance, whereas normalizing flow priors offer maximum expressiveness but scale linearly with parameter count and are currently better suited for smaller architectures. Future efforts should focus on sparsifying normalizing flow priors for deep models, extending the framework to unsupervised and latent variable architectures, and reducing the computational cost of the underlying distance optimization.
The authors express high confidence in their findings, supported by rigorous convergence diagnostics and validation across multiple sampling schemes, including full-batch Hamiltonian Monte Carlo. However, readers should note that computational overhead remains cubic with respect to the number of measurement points used in Gaussian process sampling, and the alignment relies on smooth network activations and finite measurement grids. Despite these boundary conditions, the evidence confirms that functional prior alignment is an effective, scalable strategy for robust deep learning.
- Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Bayes by Backprop establishes how distributions over neural-network weights yield Bayesian predictions, the parameter-space starting point for understanding why this paper instead tunes priors by their induced functions.
- Paper: Deep Gaussian Processes, Andreas C. Damianou et al. (2012). Deep Gaussian Processes develops the function-space Bayesian models used here as interpretable target priors, making its account of deep probabilistic functions a direct conceptual prerequisite.
- Paper: Practical Variational Inference for Neural Networks, Alex Graves (2011). Practical Variational Inference for Neural Networks introduces distributions over network weights and their approximate Bayesian treatment, clarifying the parameter-prior setting this paper seeks to improve.
- Paper: Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, Yarin Gal et al. (2016). Dropout as a Bayesian Approximation connects neural networks to deep Gaussian processes, preparing readers to understand the paper’s comparison between parameterized networks and function-space priors.
No sufficiently relevant recommendations were found.
