An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference
Jeremias KnoblauchJack JewsonTheodoros Damoulas
Presents Generalized Variational Inference, a modular optimization framework that extends standard Bayesian updating to handle misspecified priors, misspecified likelihoods, and computational constraints in deep probabilistic models.
Modern machine learning and large-scale data analytics frequently apply Bayesian statistical methods to quantify uncertainty and improve predictive performance. However, standard Bayesian inference relies on three core assumptions: correctly specified prior distributions, accurately specified data-generating models (likelihoods), and infinite computational resources. In complex, high-dimensional applications—such as Bayesian neural networks and deep Gaussian processes—these assumptions are routinely broken. Ad-hoc priors, model misspecification, outliers, and heavy computational constraints often cause standard Bayesian and variational methods to produce brittle, overconfident, or distorted predictions.
The article demonstrates that Bayesian inference can be generalized into a unified, optimization-centric framework that systematically overcomes these limitations. Its main objective is to establish an axiomatic formulation of Bayesian updating as an optimization problem and to introduce Generalized Variational Inference, a flexible, scalable methodology that accommodates imperfect priors, misspecified models, and finite computational budgets.
To develop this framework, the authors formulate an axiomatic foundation based on information regularizers and risk minimization, proving that standard Bayesian updating and existing variational methods are constrained instances of a single optimization structure termed the Rule of Three. The authors establish theoretical guarantees, including Frequentist consistency and bounds linking this approach to approximate evidence lower bounds. They then develop practical quasi-conjugate and black-box variational optimization algorithms and evaluate them empirically on benchmark machine learning tasks, comparing predictive performance (root mean square error and negative log likelihood) against standard variational inference and discrepancy-based alternatives on regression data sets.
The article presents four primary findings. First, any exact or variational Bayesian posterior can be represented as an optimization problem characterized by three modular components: an empirical loss function, a prior divergence regularizer, and a constrained family of feasible distributions. Second, this modularity guarantees that modifying the loss tackles model misspecification and outliers without altering uncertainty quantification, while modifying the divergence corrects for poor priors and adjusts posterior variances without warping parameter estimation. Third, in Bayesian neural network experiments, using robust divergence regularizers—specifically Rényi's alpha-divergence with alpha greater than one—significantly improved predictive accuracy and out-of-sample likelihood over standard variational inference, outperforming alternative discrepancy-based methods that accidentally collapsed predictive uncertainty during hyperparameter optimization. Fourth, robust scoring functions derived from beta- and gamma-divergences successfully mitigated data contamination and outliers in both changepoint detection and deep Gaussian process regression while retaining closed-form computational efficiency.
These findings imply that statistical machine learning systems do not need to rely on the unrealistic assumption that models or priors are perfect descriptions of reality. By treating posterior inference as a direct, modular optimization problem rather than an inflexible probability update, practitioners can engineer algorithms with greater robustness against anomalous data, misinformed initial assumptions, and restrictive computational limits. This structure reduces the operational risk of model failure and prevents costly predictive errors caused by outlier contamination and under-estimated uncertainty.
For practitioners and engineering teams, the source supports adopting Generalized Variational Inference as a drop-in replacement for standard evidence lower bound optimizations in high-dimensional probabilistic models. Specifically, teams should use additive, robust divergence-based losses (such as beta- or gamma-losses with hyperparameter tuning between 0.01 and 0.1 on standardized data) when data streams contain noise or outliers, and adopt Rényi's alpha-divergence when factorized default priors cause overconcentration. Standard variational software architectures can integrate these updates with minimal code modification using black-box gradient routines.
The primary limitations of the proposed approach involve hyperparameter selection and computational trade-offs. Selecting optimal divergence parameters often requires tuning on standardized data, and non-additive robust losses can scale poorly with sample size. Furthermore, while theoretical consistency guarantees hold under mild regularity conditions, closed-form gradient evaluations are mainly restricted to exponential family distributions. Readers can place high confidence in the foundational theory and empirical performance improvements reported for benchmark deep models, but pilots and validation splits remain necessary to calibrate tuning parameters for specific industrial datasets.
- Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). This review establishes the ELBO, variational-family, and KL-minimization framework that the source generalizes into its optimization-centric account of Bayesian inference.
- Paper: Deep Gaussian Processes, Andreas C. Damianou et al. (2012). Its variational treatment of deep Gaussian processes provides a concrete inference setup that the source later revisits with generalized variational posteriors.
- Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Bayes by Backprop supplies the Bayesian-neural-network setting in which the source explores the robustness and posterior-marginal effects of generalized variational inference.
No sufficiently relevant recommendations were found.
