An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference

Jeremias KnoblauchJack JewsonTheodoros Damoulas

article2022JMLR170 citations

Presents Generalized Variational Inference, a modular optimization framework that extends standard Bayesian updating to handle misspecified priors, misspecified likelihoods, and computational constraints in deep probabilistic models.

Listen

Modern machine learning and large-scale data analytics frequently apply Bayesian statistical methods to quantify uncertainty and improve predictive performance. However, standard Bayesian inference relies on three core assumptions: correctly specified prior distributions, accurately specified data-generating models (likelihoods), and infinite computational resources. In complex, high-dimensional applications—such as Bayesian neural networks and deep Gaussian processes—these assumptions are routinely broken. Ad-hoc priors, model misspecification, outliers, and heavy computational constraints often cause standard Bayesian and variational methods to produce brittle, overconfident, or distorted predictions.

The article demonstrates that Bayesian inference can be generalized into a unified, optimization-centric framework that systematically overcomes these limitations. Its main objective is to establish an axiomatic formulation of Bayesian updating as an optimization problem and to introduce Generalized Variational Inference, a flexible, scalable methodology that accommodates imperfect priors, misspecified models, and finite computational budgets.

To develop this framework, the authors formulate an axiomatic foundation based on information regularizers and risk minimization, proving that standard Bayesian updating and existing variational methods are constrained instances of a single optimization structure termed the Rule of Three. The authors establish theoretical guarantees, including Frequentist consistency and bounds linking this approach to approximate evidence lower bounds. They then develop practical quasi-conjugate and black-box variational optimization algorithms and evaluate them empirically on benchmark machine learning tasks, comparing predictive performance (root mean square error and negative log likelihood) against standard variational inference and discrepancy-based alternatives on regression data sets.

The article presents four primary findings. First, any exact or variational Bayesian posterior can be represented as an optimization problem characterized by three modular components: an empirical loss function, a prior divergence regularizer, and a constrained family of feasible distributions. Second, this modularity guarantees that modifying the loss tackles model misspecification and outliers without altering uncertainty quantification, while modifying the divergence corrects for poor priors and adjusts posterior variances without warping parameter estimation. Third, in Bayesian neural network experiments, using robust divergence regularizers—specifically Rényi's alpha-divergence with alpha greater than one—significantly improved predictive accuracy and out-of-sample likelihood over standard variational inference, outperforming alternative discrepancy-based methods that accidentally collapsed predictive uncertainty during hyperparameter optimization. Fourth, robust scoring functions derived from beta- and gamma-divergences successfully mitigated data contamination and outliers in both changepoint detection and deep Gaussian process regression while retaining closed-form computational efficiency.

These findings imply that statistical machine learning systems do not need to rely on the unrealistic assumption that models or priors are perfect descriptions of reality. By treating posterior inference as a direct, modular optimization problem rather than an inflexible probability update, practitioners can engineer algorithms with greater robustness against anomalous data, misinformed initial assumptions, and restrictive computational limits. This structure reduces the operational risk of model failure and prevents costly predictive errors caused by outlier contamination and under-estimated uncertainty.

For practitioners and engineering teams, the source supports adopting Generalized Variational Inference as a drop-in replacement for standard evidence lower bound optimizations in high-dimensional probabilistic models. Specifically, teams should use additive, robust divergence-based losses (such as beta- or gamma-losses with hyperparameter tuning between 0.01 and 0.1 on standardized data) when data streams contain noise or outliers, and adopt Rényi's alpha-divergence when factorized default priors cause overconcentration. Standard variational software architectures can integrate these updates with minimal code modification using black-box gradient routines.

The primary limitations of the proposed approach involve hyperparameter selection and computational trade-offs. Selecting optimal divergence parameters often requires tuning on standardized data, and non-additive robust losses can scale poorly with sample size. Furthermore, while theoretical consistency guarantees hold under mild regularity conditions, closed-form gradient evaluations are mainly restricted to exponential family distributions. Readers can place high confidence in the foundational theory and empirical performance improvements reported for benchmark deep models, but pilots and validation splits remain necessary to calibrate tuning parameters for specific industrial datasets.

  • Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). This review establishes the ELBO, variational-family, and KL-minimization framework that the source generalizes into its optimization-centric account of Bayesian inference.
  • Paper: Deep Gaussian Processes, Andreas C. Damianou et al. (2012). Its variational treatment of deep Gaussian processes provides a concrete inference setup that the source later revisits with generalized variational posteriors.
  • Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Bayes by Backprop supplies the Bayesian-neural-network setting in which the source explores the robustness and posterior-marginal effects of generalized variational inference.

No sufficiently relevant recommendations were found.

Cover for An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference

Abstract

We advocate an optimization-centric view of Bayesian inference. Our inspiration is the representation of Bayes’ rule as infinite-dimensional optimization (Csiszár, 1975; Donsker and Varadhan, 1975; Zellner, 1988). Equipped with this perspective, we study Bayesian inference when one does not have access to (1) well-specified priors, (2) well-specified likelihoods, (3) infinite computing power. While these three assumptions underlie the standard Bayesian paradigm, they are typically inappropriate for modern Machine Learning applications. We propose addressing this through an optimization-centric generalization of Bayesian posteriors that we call the Rule of Three (RoT). The RoT can be justified axiomatically and recovers Bayesian, PAC-Bayesian and VI posteriors as special cases. While the RoT is primarily a conceptual and theoretical device, it also encompasses a novel sub-class of tractable posteriors which we call Generalized Variational Inference (GVI) posteriors. Just as the RoT, GVI posteriors are specified by three arguments: a loss, a divergence and a variational family. They also possess a number of desirable properties, including modularity, Frequentist consistency and an interpretation as approximate ELBO. We explore applications of GVI posteriors, and show that they can be used to improve robustness and posterior marginals on Bayesian Neural Networks and Deep Gaussian Processes.

Table of Contents

  • 1. Introduction
  • 2. An optimization-centric view on Bayesian inference
  • 2.1 Preliminaries
  • 2.2 Bayesian inference as infinite-dimensional optimization
  • 2.3 Optimality of standard Variational Inference
  • 2.3.1 VI as log evidence bound
  • 2.3.3 VI as constrained optimization
  • 2.4 Reconciling (sub)optimality with empirical evidence
  • 3. A reality check: Re-examining the traditional Bayesian paradigm
  • 3.1 The traditional Bayesian paradigm
  • 3.2 Machine Learning: Challenging the traditional Bayesian paradigm
  • 3.3 Prior misspecification
  • 3.4 Likelihood Misspecification
  • 3.5 Computation mismatch
  • 4. The Rule of Three: Optimization-Centric Bayesian Inference
  • 4.1 An axiomatic foundation for Bayesian inference
  • 4.2 The Rule of Three
  • 4.3 Modularity of the Rule of Three
  • 4.4 Connecting the Rule of Three to existing methods
  • 4.4.1 Standard Variational Inference (VI)
  • 4.4.2 Coherence and the RoT
  • 4.4.3 PAC-Bayes
  • 4.4.4 Latent Variable Models & Variational Autoencoders
  • 4.4.5 Links with Information Theory
  • 5. Generalized Variational Inference (GVI)
  • 5.1 Operationalizing the Optimization-Centric View on Bayesian Inference
  • 5.2 Choosing ℓ and D: Robustness, better marginals, and beyond
  • 5.2.1 Robustness to prior misspecification via D
  • 5.2.2 Robustness to model misspecification via ℓ
  • 5.2.3 Beyond robustness
  • 5.3 Theoretical properties of GVI
  • 5.3.1 Frequentist consistency
  • 5.3.2 GVI as a posterior approximation
  • 5.4 Inference with Generalized Variational Inference (GVI)
  • 5.4.1 Quasi-conjugate inference
  • 5.4.2 Additional details on Black-Box GVI (BBGVI)
  • 6. Experiments
  • 6.1 Bayesian Neural Network Regression
  • 6.1.1 Typical patterns (Figure 11)
  • 6.1.2 The surprising benefits of modularity (Figure 12)
  • 6.2 Deep Gaussian Processes
  • 6.2.1 Preliminaries for DGPs
  • 6.2.2 Doubly stochastic VI in DGPs
  • 6.2.3 Adaption to GVI
  • 6.2.4 Results
  • 7. Discussion & Conclusion
  • Acknowledgments
  • Appendix A. Definitions for robust divergences
  • Appendix B. Comparing robust divergences as prior regularizer
  • B.1 A cautionary tale: The boundedness of the α-divergence (D^(α)_A)
  • B.2 Larger divergences produce larger marginal variances
  • B.3 Robustness to the prior
  • B.3.1 Weighted KLD (1/w KLD)
  • B.3.2 Rényi's α-divergence (D^(α)_AR)
  • B.3.3 β-divergence (D^(β)_B)
  • B.3.4 γ-divergence (D^(γ)_G)
  • Appendix C. Proof of Theorem 10
  • Appendix D. Link to the Predictive Information Bottleneck
  • Appendix E. Experimental Details for Figure 4
  • Appendix F. Proof of Theorem 14 and additional lower bounds
  • F.1 Proof for D^(α)_AR (Theorem 14)
  • F.2 The D^(β)_B prior regulariser
  • F.3 The D^(γ)_G prior regulariser
  • F.3.1 Interpretation
  • Appendix G. Proof of Proposition 16
  • Appendix H. Black Box GVI (BBGVI)
  • H.1 Preliminaries and assumptions
  • H.2 Standard black box VI with (A2) and (A3)
  • H.3 BBGVI under (A2)
  • H.3.1 Gradients if (D1) holds, not using (A2)
  • H.3.2 Gradients if (D2) holds, not using (A2)
  • H.3.3 Rao-Blackwellization for variance reduction, using (A2)
  • H.4 BBGVI if neither (A2) nor (A3) hold
  • H.5 Generically applicable variance reduction
  • Appendix I. Closed forms for divergences & proof of Proposition 17
  • I.1 High-level overview of results and preliminaries
  • I.2 Results, proofs & examples
  • I.2.1 Master result for D^(α,β,r)_G
  • I.2.3 Corollary: The special cases of D^(β)_B, D^(γ)_G
  • Appendix J. Experiments
  • J.1 Bayesian Neural Networks (BNNs)
  • J.1.1 First set of additional experiments (Figure 21)
  • J.1.2 Second set of additional experiments (Figure 22)
  • J.2 Deep Gaussian Processes (DGPs)
  • J.2.1 Proof of Corollary 18
  • J.2.2 Proof of Proposition 19
  • J.2.3 Additional experiments varying D (Figure 23)
  • References

Knowls

  1. Knowl 1 — Axiomatic characterization of the Rule of Three

    theoretical result

    Let Θ\Theta be a parameter space, P(Θ)\mathcal{P}(\Theta) its probability measures, π\pi a prior, and x1:n=(x1,…,xn)x_{1:n}=(x_1,\ldots,x_n) observed data. Suppose a posterior q∗q^* minimizes an objective over some feasible set Π⊆P(Θ)\Pi\subseteq\mathcal{P}(\Theta), and the objective increases with both the expected cumulative loss and a statistical divergence from the prior. If the objective's combining function is independent of the prior, data, loss, divergence, and feasible set, and the construction recovers the standard Bayesian posterior when the divergence is Kullback–Leibler (KL) and Π=P(Θ)\Pi=\mathcal{P}(\Theta), the resulting objective is additive:

    q∗=P(ℓ,D,Π)=arg⁡min⁡q∈Π{Eq[∑i=1nℓ(θ,xi)]+D(q∥π)}.q^*=\mathcal{P}(\ell,D,\Pi) =\arg\min_{q\in\Pi}\left\{\mathbb{E}_{q}\left[\sum_{i=1}^n\ell(\theta,x_i)\right]+D(q\|\pi)\right\}.

    Here ℓ(θ,xi)\ell(\theta,x_i) is the loss for parameter θ\theta and observation xix_i, and D(q∥π)≥0D(q\|\pi)\geq 0 vanishes only when qq and π\pi agree almost everywhere. The paper calls this three-part specification—the loss, prior divergence, and feasible posterior space—the Rule of Three (RoT).

  2. Knowl 2 — Standard variational inference is optimal within its chosen family

    theoretical result

    For a fixed variational family Q⊆P(Θ)\mathcal{Q}\subseteq\mathcal{P}(\Theta), standard variational inference (VI) minimizes the same expected-loss-plus-KL objective as unrestricted Bayesian inference, but restricts the optimization to Q\mathcal{Q}:

    qVI∗=arg⁡min⁡q∈Q{Eq[∑i=1nℓ(θ,xi)]+KL(q∥π)}.q^*_{\mathrm{VI}}=\arg\min_{q\in\mathcal{Q}}\left\{\mathbb{E}_{q}\left[\sum_{i=1}^n\ell(\theta,x_i)\right]+\mathrm{KL}(q\|\pi)\right\}.

    Thus, for the objective defining the Bayesian posterior and any fixed finite-dimensional Q\mathcal{Q}, standard VI gives the optimal posterior in that family. This is an objective-relative result: it does not imply that VI will predict better in practice if the objective is misspecified or the family cannot meaningfully approximate the unrestricted posterior.

  3. Knowl 3 — Generalized variational inference as a tractable RoT posterior

    model/method

    Generalized Variational Inference (GVI) applies the Rule of Three with a parameterized variational family Q={q(θ∣κ):κ∈K}\mathcal{Q}=\{q(\theta\mid\kappa):\kappa\in\mathcal{K}\}, where κ\kappa is a finite-dimensional variational parameter:

    qGVI∗=arg⁡min⁡q∈Q{Eq[∑i=1nℓ(θ,xi)]+D(q∥π)}.q^*_{\mathrm{GVI}}=\arg\min_{q\in\mathcal{Q}}\left\{\mathbb{E}_{q}\left[\sum_{i=1}^n\ell(\theta,x_i)\right]+D(q\|\pi)\right\}.

    The loss ℓ\ell, divergence DD, and family Q\mathcal{Q} can be selected separately. Unlike VI methods motivated solely as approximations to a standard Bayesian posterior, GVI defines a posterior directly through the chosen objective; it is intended to be better suited to the inference task when the standard prior, likelihood, or computational assumptions are unsuitable.

  4. Knowl 4 — Modularity separates target fit, prior regularization, and computation

    theoretical result

    For a Rule-of-Three posterior P(ℓ,D,Π)\mathcal{P}(\ell,D,\Pi), hold the data, prior, and feasible set Π\Pi fixed. The paper's modularity result assigns distinct roles to the objective components: changing the loss ℓ\ell changes how parameter values are assessed and can address model misspecification; changing the divergence DD changes how the posterior is regularized toward the prior and can address prior misspecification; and changing DD can also alter uncertainty quantification without changing the loss-defined parameter target. The feasible set Π\Pi determines which posteriors can be computed, so restricting it to a tractable family addresses computational limits. These are modular design roles, not guarantees that every choice of loss or divergence will improve inference.

  5. Knowl 5 — Robust likelihood scoring rules for model misspecification

    model/method

    The paper proposes replacing the negative log likelihood with additive scoring rules derived from the β\beta- and γ\gamma-divergences. For likelihood density p(y∣θ)p(y\mid\theta), observation yy, and c>0c>0, define Ip,c(θ)=∫p(z∣θ)c dzI_{p,c}(\theta)=\int p(z\mid\theta)^c\,dz. The scores are

    ℓp(β)(θ,y)=−p(y∣θ)β−1β−1+Ip,β(θ)β,\ell_p^{(\beta)}(\theta,y)=-\frac{p(y\mid\theta)^{\beta-1}}{\beta-1}+\frac{I_{p,\beta}(\theta)}{\beta}, ℓp(γ)(θ,y)=−γγ−1 p(y∣θ)γ−1Ip,γ(θ)−(γ−1)/γ.\ell_p^{(\gamma)}(\theta,y)=-\frac{\gamma}{\gamma-1}\,p(y\mid\theta)^{\gamma-1}I_{p,\gamma}(\theta)^{-(\gamma-1)/\gamma}.

    For β>1\beta>1 or γ>1\gamma>1, respectively, these scores reduce the influence of observations that fit the model poorly; as the corresponding parameter approaches 11, the score approaches the negative log likelihood. The paper reports that values just above 11 can balance robustness and efficiency, and suggests 1+ε1+\varepsilon with ε∈[0.01,0.1]\varepsilon\in[0.01,0.1] for standardized data as a useful empirical range, not a universal optimum.

  6. Knowl 6 — Rényi divergence can improve robustness to prior misspecification

    empirical result

    For posterior density q(θ)q(\theta) and prior density π(θ)\pi(\theta), the paper considers the Rényi α\alpha-divergence

    DAR(α)(q∥π)=1α(α−1)log⁡∫q(θ)απ(θ)1−α dθ,D_{\mathrm{AR}}^{(\alpha)}(q\|\pi)=\frac{1}{\alpha(\alpha-1)}\log\int q(\theta)^\alpha\pi(\theta)^{1-\alpha}\,d\theta,

    which recovers KL divergence as α→1\alpha\to1. In the paper's finite-sample analyses, using DAR(α)D_{\mathrm{AR}}^{(\alpha)} with 0<α<10<\alpha<1 can make GVI less sensitive to a badly specified prior while producing wider marginal variances than KL-regularized VI. The reported comparisons found this combination more useful than simply rescaling the KL penalty, which trades prior sensitivity against variance in an unfavorable way. These are findings for the divergences and settings examined, rather than a general guarantee for every model or variational family.

  7. Knowl 7 — Black-box GVI with stochastic gradients

    algorithm

    Black-box GVI optimizes a parameterized posterior using minibatches and Monte Carlo estimates, and can therefore be applied when the objective has no closed form. Inputs are observations x1:nx_{1:n}, prior π\pi, loss ℓ\ell, divergence DD, family q(θ∣κ)q(\theta\mid\kappa), initial parameter κ0\kappa_0, minibatch size KK, number of posterior samples SS, learning-rate schedule, stopping rule, and optionally a control variate. At iteration tt, sample KK observations without replacement and draw SS independent values θ(s)∼q(θ∣κt)\theta^{(s)}\sim q(\theta\mid\kappa_t). Estimate the loss gradient with the score-function terms and scale the minibatch sum by n/Kn/K. Add the divergence gradient when it is available in closed form; when D(q∥π)=Eq[rκ,π(θ)]D(q\|\pi)=\mathbb{E}_{q}[r_{\kappa,\pi}(\theta)], estimate its gradient with the corresponding score-function and direct-gradient terms. Average across the SS draws, optionally apply variance reduction, and use the selected optimizer to minimize the objective. Stop when the supplied stopping rule is met and return q(θ∣κ^)q(\theta\mid\widehat{\kappa}). The paper does not prescribe universal values for KK, SS, or the learning-rate schedule; these are algorithm inputs. Its stated gradient estimator for a nonlinear transformation of an expectation is a plug-in estimate and may be biased.

    Input: Data x1:nx_{1:n}, prior π\pi, loss ℓ\ell, divergence DD, family q(θ∣κ)q(\theta\mid\kappa), initial κ0\kappa_0, minibatch size KK, sample count SS, learning-rate schedule, stopping rule, optional control variate hh
    Set t=0t=0
    Repeat
        Sample KK indices without replacement from {1,…,n}\{1,\ldots,n\}
        Draw SS independent samples θ(s)∼q(θ∣κt)\theta^{(s)}\sim q(\theta\mid\kappa_t)
        For each draw, estimate the loss gradient using nK∑i in minibatchℓ(θ(s),xi)∇κlog⁡q(θ(s)∣κt)\frac{n}{K}\sum_{i\text{ in minibatch}}\ell(\theta^{(s)},x_i)\nabla_\kappa\log q(\theta^{(s)}\mid\kappa_t)
        Add the divergence gradient, using its closed form or its applicable Monte Carlo estimator
        Average the gradient estimates over the SS draws
        If hh is supplied, apply the control-variate correction
        Update κt\kappa_t with the chosen optimizer to reduce the objective
        Set t←t+1t\leftarrow t+1
    Until the stopping rule is satisfied
    Return q(θ∣κ^)q(\theta\mid\widehat{\kappa})
  8. Knowl 8 — Frequentist consistency of GVI

    theoretical result

    For independent, identically distributed observations from a true distribution PXP_X, let the population-optimal parameter be the unique minimizer θ∗=arg⁡min⁡θEPX[ℓ(θ,X)]\theta^*=\arg\min_{\theta}\mathbb{E}_{P_X}[\ell(\theta,X)]. The paper establishes that GVI posteriors concentrate at this parameter under regularity conditions ensuring the optimization problems are well defined and expected loss is finite. One stated version uses a mean-field normal variational family, a divergence DD that is lower-semicontinuous in its posterior argument and finite for every member of the family, and a prior satisfying Eπ[EPX[ℓ(θ,X)]]<∞\mathbb{E}_{\pi}[\mathbb{E}_{P_X}[\ell(\theta,X)]]<\infty. Under these conditions and the paper's additional regularity assumptions, the posterior based on nn observations converges weakly to a point mass:

    qGVI,n∗⇒δθ∗as n→∞.q^*_{\mathrm{GVI},n}\Rightarrow\delta_{\theta^*}\qquad\text{as }n\to\infty.

    The stated result does not require DD to be KL divergence; the limiting target is determined by the population risk of the selected loss.

  9. Knowl 9 — Certain GVI objectives bound a power-likelihood ELBO

    theoretical result

    With the Rényi divergence DAR(α)D_{\mathrm{AR}}^{(\alpha)}, the GVI objective has an evidence-lower-bound interpretation. Let Ln(θ)=∑i=1nℓ(θ,xi)L_n(\theta)=\sum_{i=1}^n\ell(\theta,x_i), take α>0\alpha>0 with α≠1\alpha\ne1, set w(α)=max⁡{1,α}w(\alpha)=\max\{1,\alpha\} and c(α)=min⁡{1,1/α}c(\alpha)=\min\{1,1/\alpha\}, and define the power-loss posterior qBwℓ(θ)∝π(θ)exp⁡[−w(α)Ln(θ)]q_B^{w\ell}(\theta)\propto\pi(\theta)\exp[-w(\alpha)L_n(\theta)]. For any admissible variational density qq, let ELBOwℓ(q)=Eq[−w(α)Ln(θ)+log⁡π(θ)−log⁡q(θ)]\mathrm{ELBO}_{w\ell}(q)=\mathbb{E}_q[-w(\alpha)L_n(\theta)+\log\pi(\theta)-\log q(\theta)]. Then

    Eq[Ln(θ)]+DAR(α)(q∥π)≥−c(α) ELBOwℓ(q)+S1(α,q,π),\mathbb{E}_q[L_n(\theta)]+D_{\mathrm{AR}}^{(\alpha)}(q\|\pi)\geq -c(\alpha)\,\mathrm{ELBO}_{w\ell}(q)+S_1(\alpha,q,\pi),

    where S1(α,q,π)=DAR(α)(q∥π)−KL(q∥π)S_1(\alpha,q,\pi)=D_{\mathrm{AR}}^{(\alpha)}(q\|\pi)-\mathrm{KL}(q\|\pi) for 0<α<10<\alpha<1, and S1=0S_1=0 for α>1\alpha>1. Thus the connection is to a scaled ELBO for a generalized Bayesian posterior using a power loss, with a nonzero slack term in the range 0<α<10<\alpha<1.

  10. Knowl 10 — BNN regression: Rényi-regularized GVI improves prediction for selected settings

    empirical result

    The Bayesian neural network (BNN) regression experiments compared standard VI, discrepancy-VI (DVI), and GVI using Rényi-divergence regularization. Each model used one hidden layer with 50 ReLU units, factorized normal prior and variational distributions, probabilistic backpropagation, and Adam for 500 epochs with batch size 32. On UCI datasets, performance was assessed over 50 random 90%/10% train/test splits using average negative log likelihood (NLL) and root mean square error (RMSE); each test prediction used 100 posterior samples. In the reported datasets, GVI with α>1\alpha>1 improved both NLL and RMSE relative to standard VI, whereas 0<α<10<\alpha<1 generally worsened predictive performance. Performance varied non-monotonically with α\alpha: reducing uncertainty initially helped, but excessive concentration harmed prediction. DVI sometimes improved NLL relative to standard VI but showed less consistent gains for RMSE. The reported comparison therefore supports tuning posterior concentration to the task, rather than treating wider variational uncertainty as uniformly beneficial.

  11. Knowl 11 — DGP regression: robust scoring improves reported predictive performance

    empirical result

    For deep Gaussian process (DGP) regression, the paper modified the likelihood loss rather than the prior divergence, replacing the log score with the robust γ\gamma-divergence score for γ∈{1.01,1.05}\gamma\in\{1.01,1.05\}. The experiments used the doubly stochastic DGP implementation with an RBF kernel and dimension-wise lengthscales, 100 inducing points, minibatches of size min⁡(1000,n)\min(1000,n), and 20,000 Adam iterations at learning rate 0.01. The latent width was min⁡(Dx,30)\min(D_x,30), where DxD_x is the input dimension. Predictive NLL and RMSE were evaluated on 50 random 90%/10% train/test splits of UCI datasets. Across the reported comparisons, GVI with the robust score outperformed standard VI on both metrics; γ=1.01\gamma=1.01 generally did better than 1.051.05, though the two were comparable on some datasets. The authors suggest that the robust score helps by down-weighting non-informative regions of the latent spaces; this is their explanation for the observed gains, not a separately established causal result.

Coverage note — The auxiliary connections to PAC-Bayes, VAEs, and the Predictive Information Bottleneck, along with detailed closed-form divergence derivations and secondary experiments, are omitted because they extend or instantiate the RoT/GVI framework rather than changing its principal results.

References

  1. 1.Ryan Prescott Adams and David J. C. MacKay. Bayesian online changepoint detection. arXiv preprint arXiv:0710.3742, 2007.
  2. 2.James Aitchison. Goodness of prediction fit. Biometrika, 62(3):547–554, 1975.
  3. 3.Alexander A. Alemi. Variational predictive information bottleneck. In Workshop on Information Theory, Advances in Neural Information Processing Systems, 2019.
  4. 4.Pierre Alquier. Non-exponentially weighted aggregation: regret bounds for unbounded loss functions. arXiv preprint arXiv:2009.03017, 2020.
  5. 5.Pierre Alquier and Benjamin Guedj. Simpler PAC-Bayesian bounds for hostile data. Machine Learning, 107(5):887–902, 2018.
  6. 6.Pierre Alquier, James Ridgway, and Nicolas Chopin. On the properties of variational approximations of Gibbs posteriors. The Journal of Machine Learning Research, 17(1):8374–8414, 2016.
  7. 7.Shun-ichi Amari. Differential-geometrical methods in statistics, volume 28. Springer Science & Business Media, 2012.
  8. 8.Maximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton, Benjamin Letham, Andrew Gordon Wilson, and Eytan Bakshy. BoTorch: Programmable Bayesian optimization in PyTorch. arXiv preprint arXiv:1910.06403, 2019.
  9. 9.A. Barp, F.-X. Briol, A. B. Duncan, M. Girolami, and L. Mackey. Minimum Stein discrepancy estimators. In Neural Information Processing Systems, pages 12964–12976, 2019.
  10. 10.Ayanendranath Basu, Ian R. Harris, Nils L. Hjort, and M. C. Jones. Robust and efficient estimation by minimising a density power divergence. Biometrika, 85(3):549–559, 1998.
  11. 11.Thomas Bayes. An essay towards solving a problem in the doctrine of chances. Philosophical Transactions of the Royal Society of London, 53:370–418, 1763.
  12. 12.Matthew James Beal. Variational algorithms for approximate Bayesian inference. University College London, 2003.
  13. 13.Luc B´egin, Pascal Germain, Fran¸cois Laviolette, and Jean-Francis Roy. PAC-Bayesian bounds based on the R´enyi divergence. In Artificial Intelligence and Statistics, pages 435–444, 2016.
  14. 14.Rudolf Beran et al. Minimum hellinger distance estimates for parametric models. The annals of Statistics, 5(3):445–463, 1977.
  15. 15.James O. Berger. The case for objective Bayesian analysis. Bayesian analysis, 1(3):385–402, 2006.
  16. 16.James O Berger and Jos´e M Bernardo. On the development of the reference prior method. Bayesian statistics, 4(4):35–60, 1992.
  17. 17.James O. Berger, Elias Moreno, Luis Raul Pericchi, M. Jesus Bayarri, Jose M. Bernardo, Juan A. Cano, Julian De la Horra, Jacinto Martin, David Rios-Insua, and Bruno Betro. An overview of robust Bayesian analysis. Test, 3(1):5–124, 1994.
  18. 18.Jose M. Bernardo. Reference posterior distributions for Bayesian inference. Journal of the Royal Statistical Society: Series B (Methodological), 41(2):113–128, 1979.
  19. 19.Jos´e M. Bernardo. Bayesian theory. Wiley Series in Probability and Statistics. 23 cm. 586 p., 2000.
  20. 20.Alexandros Beskos, Natesh Pillai, Gareth Roberts, Jesus-Maria Sanz-Serna, and Andrew Stuart. Optimal tuning of the hybrid Monte Carlo algorithm. Bernoulli, 19(5A):1501–1534, 2013.
  21. 21.William Bialek, Ilya Nemenman, and Naftali Tishby. Predictability, complexity, and learning. Neural computation, 13(11):2409–2463, 2001.
  22. 22.Pier Giovanni Bissiri, Chris Holmes, and Stephen Walker. A general framework for updating belief distributions. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 78(5):1103–1130, 2016.
  23. 23.Edwin V. Bonilla, Karl Krauth, and Amir Dezfouli. Generic inference in latent Gaussian process models. Journal of Machine Learning Research, 20(117):1–63, 2019.
  24. 24.George E. P. Box. Sampling and Bayes’ inference in scientific modelling and robustness. Journal of the Royal Statistical Society. Series A (General), pages 383–430, 1980.
  25. 25.F-X. Briol, A. Barp, A. B. Duncan, and M. Girolami. Statistical inference for generative models with maximum mean discrepancy. arXiv:1906.05944, 2019.
  26. 26.Thang Bui, Daniel Hern´andez-Lobato, Jose Hernandez-Lobato, Yingzhen Li, and Richard Turner. Deep Gaussian processes for regression using approximate expectation propagation. In International Conference on Machine Learning, pages 1472–1481, 2016.
  27. 27.Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov. Importance weighted autoencoders. In International Conference on Learning Representations, 2016.
  28. 28.Fran¸cois Caron, Arnaud Doucet, and Raphael Gottardo. On-line changepoint detection and parameter estimation with application to genomic data. Statistics and Computing, 22(2):579–595, 2012.
  29. 29.Olivier Catoni. Pac-bayesian supervised classification: the thermodynamics of statistical learning. arXiv preprint arXiv:0712.0248, 2007.
  30. 30.Liqun Chen, Chenyang Tao, Ruiyi Zhang, Ricardo Henao, and Lawrence Carin Duke. Variational inference and model selection with generalized evidence bounds. In International Conference on Machine Learning, pages 892–901, 2018.
  31. 31.Badr-Eddine Ch´erief-Abdellatif and Pierre Alquier. MMD-Bayes: Robust Bayesian estimation via maximum mean discrepancy. arXiv preprint arXiv:1909.13339, 2019a.
  32. 32.Badr-Eddine Ch´erief-Abdellatif and Pierre Alquier. Finite sample properties of parametric mmd estimation: robustness to misspecification and dependence. arXiv preprint arXiv:1912.05737, 2019b.
  33. 33.Herman Chernoff. A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations. The Annals of Mathematical Statistics, 23(4):493–507, 1952.
  34. 34.Anna Choromanska, Mikael Henaff, Michael Mathieu, G´erard Ben Arous, and Yann LeCun. The loss surfaces of multilayer networks. In Artificial Intelligence and Statistics, pages 192–204, 2015.
  35. 35.Andrzej Cichocki and Shun-ichi Amari. Families of alpha-beta-and gamma-divergences: Flexible and robust measures of similarities. Entropy, 12(6):1532–1568, 2010.
  36. 36.Imre Csisz´ar. I-divergence geometry of probability distributions and minimization problems. The Annals of Probability, pages 146–158, 1975.
  37. 37.Kurt Cutajar, Edwin V Bonilla, Pietro Michiardi, and Maurizio Filippone. Random feature expansions for deep Gaussian processes. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 884–893. JMLR, 2017a.
  38. 38.Kurt Cutajar, Edwin V Bonilla, Pietro Michiardi, and Maurizio Filippone. Random feature expansions for deep Gaussian processes. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 884–893. JMLR. org, 2017b.
  39. 39.Z. Dai, A. Damianou, J. Gonzalez, and N. Lawrence. Variational auto-encoded deep Gaussian processes. In International Conference on Learning Representations, 2016.
  40. 40.Andreas Damianou and Neil Lawrence. Deep Gaussian processes. In Artificial Intelligence and Statistics, pages 207–215, 2013.
  41. 41.G Darmois. Sur les lois de probabilit´e `a estimation exhaustive. Comptes Rendus de l’Acad´emie des Sciences, 200:1265–1266, 1935.
  42. 42.Herbert Aron David. First (?) occurrence of common terms in probability and statistics—a second list, with corrections. The American Statistician, 52(1):36–40, 1998.
  43. 43.Pierre-Simon De Laplace. M´emoire sur la probabilit´e des causes par les ´ev´enements. M´em. de math. et phys. pr´esent´es `a l’Acad. roy. des sci, 6:621–656, 1774.
  44. 44.Arthur P Dempster, Nan M Laird, and Donald B Rubin. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977.
  45. 45.Luc Devroye and G´abor Lugosi. Combinatorial methods in density estimation. Springer Science & Business Media, New York, 2012.
  46. 46.Adji Bousso Dieng, Dustin Tran, Rajesh Ranganath, John Paisley, and David Blei. Variational inference via χ upper bound minimization. In Advances in Neural Information Processing Systems, pages 2732–2741, 2017.
  47. 47.Justin Domke and Daniel R. Sheldon. Importance weighting and variational inference. In Advances in neural information processing systems, pages 4470–4479, 2018.
  48. 48.Monroe D Donsker and SR Srinivasa Varadhan. Asymptotic evaluation of certain Markov process expectations for large time, i. Communications on Pure and Applied Mathematics, 28(1):1–47, 1975.
  49. 49.Paul Fearnhead and Zhen Liu. On-line inference for multiple changepoint problems. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69(4):589–605, 2007.
  50. 50.Stephen E Fienberg. When did Bayesian inference become ”Bayesian”? Bayesian analysis, 1(1):1–40, 2006.
  51. 51.Ronald Aylmer Fisher. Contributions to mathematical statistics. 1950.
  52. 52.Edwin Fong and Chris Holmes. On the marginal likelihood and cross-validation. arXiv preprint arXiv:1905.08737, 2019.
  53. 53.Hironori Fujisawa and Shinto Eguchi. Robust parameter estimation with a small bias against heavy contamination. Journal of Multivariate Analysis, 99(9):2053–2081, 2008.
  54. 54.Futoshi Futami, Issei Sato, and Masashi Sugiyama. Variational inference based on robust divergences. In Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics, volume 84 of Proceedings of Machine Learning Research, pages 813–822. PMLR, 2018.
  55. 55.Kuzman Ganchev, Jennifer Gillenwater, and Ben Taskar. Posterior regularization for structured latent variable models. Journal of Machine Learning Research, 11:2001–2049, 2010.
  56. 56.Jacob Gardner, Geoff Pleiss, Kilian Q. Weinberger, David Bindel, and Andrew G. Wilson. Gpytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration. In Advances in Neural Information Processing Systems, pages 7576–7586, 2018.
  57. 57.Andrew Gelman, Daniel Simpson, and Michael Betancourt. The prior can often only be understood in the context of the likelihood. Entropy, 19(10):555, 2017.
  58. 58.Pascal Germain, Francis Bach, Alexandre Lacoste, and Simon Lacoste-Julien. PAC-Bayesian theory meets Bayesian inference. In Advances in Neural Information Processing Systems, pages 1884–1892, 2016.
  59. 59.Subhashis Ghosal. A review of consistency and convergence rates of posterior distributions. In Proc. Varanasi Symp. on Bayesian Inference, 1998.
  60. 60.Subhashis Ghosal, Jayanta K Ghosh, and Aad W Van Der Vaart. Convergence rates of posterior distributions. Annals of Statistics, 28(2):500–531, 2000.
  61. 61.Abhik Ghosh and Ayanendranath Basu. Robust Bayes estimation using the density power divergence. Annals of the Institute of Statistical Mathematics, 68(2):413–437, 2016.
  62. 62.Manuel Gil. On R´enyi divergence measures for continuous alphabet sources. PhD thesis, 2011.
  63. 63.Manuel Gil, Fady Alajaji, and Tamas Linder. R´enyi divergence measures for commonly used univariate continuous distributions. Information Sciences, 249:124–131, 2013.
  64. 64.M Goldstein. Influence and belief adjustment. Influence Diagrams, Belief Nets and Decision Analysis, pages 143–174, 1990.
  65. 65.Michael Goldstein. Subjective Bayesian analysis: principles and practice. Bayesian Analysis, 1(3):403–420, 2006.
  66. 66.Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, and David Duvenaud. Backpropagation through the void: Optimizing control variates for black-box gradient estimation. arXiv preprint arXiv:1711.00123, 2017.
  67. 67.Peter Gr¨unwald. Safe learning: bridging the gap between Bayes, MDL and statistical learning theory via empirical convexity. In Proceedings of the 24th Annual Conference on Learning Theory, pages 397–420, 2011.
  68. 68.Peter Gr¨unwald. The safe Bayesian. In International Conference on Algorithmic Learning Theory, pages 169–183. Springer, 2012.
  69. 69.Peter Gr¨unwald and Thijs Van Ommen. Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis, 12(4):1069–1103, 2017.
  70. 70.Benjamin Guedj. A primer on PAC-Bayesian learning. arXiv preprint arXiv:1901.05353, 2019.
  71. 71.Oliver Hamelijnck, Theodoros Damoulas, Kangrui Wang, and Mark Girolami. Multi-resolution multi-task gaussian processes. In Advances in Neural Information Processing Systems, 2019.
  72. 72.Frank R Hampel, Elvezio M Ronchetti, Peter J Rousseeuw, and Werner A Stahel. Robust statistics: the approach based on influence functions, volume 196. John Wiley & Sons, 2011.
  73. 73.Pashupati Hegde, Markus Heinonen, Harri L¨ahdesm¨aki, and Samuel Kaski. Deep learning with differential gaussian process flows. In The 22nd International Conference on Artificial Intelligence and Statistics, pages 1812–1821, 2019.
  74. 74.James Hensman and Neil D. Lawrence. Nested variational compression in deep Gaussian processes. stat, 1050:3, 2014.
  75. 75.Jos´e Miguel Hern´andez-Lobato and Ryan Adams. Probabilistic backpropagation for scalable learning of Bayesian neural networks. In International Conference on Machine Learning, pages 1861–1869, 2015.
  76. 76.Jos´e Miguel Hern´andez-Lobato, Yingzhen Li, Mark Rowland, Daniel Hern´andez-Lobato, Thang D Bui, and Richard E Turner. Black-box α-divergence minimization. In Proceedings of the 33rd International Conference on International Conference on Machine Learning-Volume 48, pages 1511–1520, 2016.
  77. 77.Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-vae: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, volume 3, 2017.
  78. 78.Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley. Stochastic variational inference. The Journal of Machine Learning Research, 14(1):1303–1347, 2013.
  79. 79.Chris Holmes and Stephen Walker. Assigning a value to a power likelihood in a general bayesian model. Biometrika, 104(2):497–503, 2017.
  80. 80.Giles Hooker and Anand N Vidyashankar. Bayesian model robustness via disparities. Test, 23(3):556–584, 2014.
  81. 81.Chin-Wei Huang, Shawn Tan, Alexandre Lacoste, and Aaron C. Courville. Improving explorability in variational inference with annealed variational objectives. In Advances in Neural Information Processing Systems, pages 9724–9734, 2018.
  82. 82.Hung Hung, Zhi-Yu Jou, and Su-Yun Huang. Robust mislabel logistic regression without modeling mislabel probabilities. Biometrics, 74(1):145–154, 2018.
  83. 83.Aapo Hyv¨arinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 6:695–708, 2005.
  84. 84.Martin Jankowiak, Geoff Pleiss, and Jacob R Gardner. Sparse gaussian process regression beyond variational inference. arXiv preprint arXiv:1910.07123, 2019.
  85. 85.Edwin T. Jaynes. Probability theory: The logic of science. Cambridge university press, 2003.
  86. 86.H. Jeffreys. Theory of probability: Oxford Univ. Press (earlier editions 1939, 1948), 1961.
  87. 87.Jack Jewson, Jim Smith, and Chris Holmes. Principles of Bayesian inference using general divergence criteria. Entropy, 20(6):442, 2018.
  88. 88.Michael I Jordan, Zoubin Ghahramani, Tommi S Jaakkola, and Lawrence K Saul. An introduction to variational methods for graphical models. Machine learning, 37(2):183–233, 1999.
  89. 89.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2014.
  90. 90.Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. In International Conference on Learning Representations, 2013.
  91. 91.Jeremias Knoblauch. Frequentist consistency of generalized variational inference. arXiv preprint arXiv:1912.04946, 2019a.
  92. 92.Jeremias Knoblauch. Robust deep Gaussian processes. arXiv preprint arXiv:1904.02303, 2019b.
  93. 93.Jeremias Knoblauch and Theodoros Damoulas. Spatio-temporal Bayesian on-line changepoint detection with model selection. In Proceedings of the 27th International Conference on Machine Learning (ICML), 2018.
  94. 94.Jeremias Knoblauch and Lara Vomfell. Robust bayesian inference for discrete outcomes with the total variation distance. ArXiv, abs/2010.13456, 2020.
  95. 95.Jeremias Knoblauch, Jack Jewson, and Theodoros Damoulas. Doubly robust Bayesian inference for non-stationary streaming data using β-divergences. In Advances in Neural Information Processing Systems (NeurIPS), pages 64–75, 2018.
  96. 96.Jeremias Knoblauch, Jack Jewson, and Theodoros Damoulas. Generalized variational inference. arXiv preprint arXiv:1904.02063, 2019.
  97. 97.Bernard Osgood Koopman. On distributions admitting a sufficient statistic. Transactions of the American Mathematical society, 39(3):399–409, 1936.
  98. 98.Solomon Kullback and Richard A Leibler. On information and sufficiency. The annals of mathematical statistics, 22(1):79–86, 1951.
  99. 99.Sebastian Kurtek and Karthik Bharath. Bayesian sensitivity analysis with the Fisher–Rao metric. Biometrika, 102(3):601–616, 2015.
  100. 100.Tomasz Ku´smierczyk, Joseph Sakaya, and Arto Klami. Variational Bayesian decision-making for continuous utilities. In Advances in Neural Information Processing Systems, 2019.
  101. 101.Simon Lacoste-Julien, Ferenc Husz´ar, and Zoubin Ghahramani. Approximate inference for the loss-calibrated Bayesian. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, pages 416–424, 2011.
  102. 102.Ga¨el Letarte, Pascal Germain, Benjamin Guedj, and Fran¸cois Laviolette. Dichotomize and generalize: Pac-bayesian binary activated deep neural networks. In Advances in Neural Information Processing Systems, 2019.
  103. 103.Yingzhen Li and Richard E Turner. R´enyi divergence variational inference. In Advances in Neural Information Processing Systems, pages 1073–1081, 2016.
  104. 104.Moshe Lichman. UCI machine learning repository, 2013.
  105. 105.F Liese and I Vajda. Convex statistical distances, volume 95 of teubner texts in mathematics. BSB BG Teubner Verlagsgesellschaft, Leipzig, 1987.
  106. 106.Bruce G Lindsay et al. Efficiency versus robustness: the case for minimum hellinger distance and related methods. The annals of statistics, 22(2):1081–1114, 1994.
  107. 107.Gabriel Loaiza-Ganem and John P. Cunningham. The continuous Bernoulli: fixing a pervasive error in variational autoencoders. In Advances in Neural Information Processing Systems, 2019.
  108. 108.Chao Ma, Yingzhen Li, and Jos´e Miguel Hern´andez-Lobato. Variational implicit processes. In International Conference on Machine Learning, pages 4222–4233. PMLR, 2019.
  109. 109.David J. C. MacKay. Bayesian methods for backpropagation networks. In Models of neural networks III, pages 211–254. Springer, 1996.
  110. 110.David J. C. MacKay. Choice of basis for Laplace approximation. Machine learning, 33(1):77–86, 1998.
  111. 111.Takuo Matsubara, Chris J Oates, and Fran¸cois-Xavier Briol. The ridgelet prior: A covariance function approach to prior specification for bayesian neural networks. arXiv preprint arXiv:2010.08488, 2020.
  112. 112.Alexander G. de G. Matthews, James Hensman, Richard Turner, and Zoubin Ghahramani. On sparse variational methods and the Kullback-Leibler divergence between stochastic processes. Journal of Machine Learning Research, 51:231–239, 2016.
  113. 113.Alexander G. de G. Matthews, Mark Van Der Wilk, Tom Nickson, Keisuke Fujii, Alexis Boukouvalas, Pablo Le´on-Villagr´a, Zoubin Ghahramani, and James Hensman. Gpflow: A Gaussian process library using tensorflow. The Journal of Machine Learning Research, 18(1):1299–1304, 2017.
  114. 114.Conor Mayo-Wilson and Aditya Saraf. Qualitative robust bayesianism and the likelihood principle. arXiv preprint arXiv:2009.03879, 2020.
  115. 115.David A. McAllester. Some PAC-Bayesian theorems. Machine Learning, 37(3):355–363, 1999a.
  116. 116.David A. McAllester. PAC-Bayesian model averaging. In Proceedings of the twelfth annual conference on Computational learning theory, pages 164–170. ACM, 1999b.
  117. 117.Minami Mihoko and Shinto Eguchi. Robust blind source separation by beta divergence. Neural computation, 14(8):1859–1886, 2002.
  118. 118.Jeffrey W. Miller and David B. Dunson. Robust Bayesian inference via coarsening. Journal of the American Statistical Association, 114(527):1113–1125, 2019.
  119. 119.Thomas Minka. Divergence measures and message passing. Technical report, Technical report, Microsoft Research, 2005.
  120. 120.Thomas P Minka. Expectation propagation for approximate Bayesian inference. In Proceedings of the Seventeenth conference on Uncertainty in artificial intelligence, pages 362–369. Morgan Kaufmann Publishers Inc., 2001.
  121. 121.Tomoyuki Nakagawa and Shintaro Hashimoto. Robust Bayesian inference via γ-divergence. Communications in Statistics-Theory and Methods, pages 1–18, 2019.
  122. 122.Eric Nalisnick, Jonathan Gordon, and Jos´e Miguel Hern´andez-Lobato. Predictive complexity priors. arXiv preprint arXiv:2006.10801, 2020.
  123. 123.Radford M Neal. Bayesian learning for neural networks, volume 118. Springer Science & Business Media, 2012.
  124. 124.Radford M Neal and Geoffrey E Hinton. A view of the EM algorithm that justifies incremental, sparse, and other variants. In Learning in graphical models, pages 355–368. Springer, 1998.
  125. 125.Anthony O’Hagan and Jeremy E Oakley. Probability is perfect, but we can’t elicit it perfectly. Reliability Engineering & System Safety, 85(1):239–248, 2004.
  126. 126.Y. Ohnishi and J. Honorio. Novel change of measure inequalities with applications to pac-bayesian bounds and monte carlo estimation. arXiv: Learning, 2020.
  127. 127.Manfred Opper and Ole Winther. Gaussian processes for classification: Mean-field algorithms. Neural computation, 12(11):2655–2684, 2000.
  128. 128.Joseph J. K. O’Ruanaidh. Numerical Bayesian methods applied to signal processing. PhD thesis, University of Cambridge, 1994.
  129. 129.John Paisley, David M. Blei, and Michael I. Jordan. Variational Bayesian inference with stochastic search. In Proceedings of the 29th International Coference on International Conference on Machine Learning, pages 1363–1370, 2012.
  130. 130.Francesco Pauli, Walter Racugno, and Laura Ventura. Bayesian composite marginal likelihoods. Statistica Sinica, pages 149–164, 2011.
  131. 131.Fengchun Peng and Dipak K Dey. Bayesian analysis of outlier problems using divergence measures. Canadian Journal of Statistics, 23(2):199–213, 1995.
  132. 132.E.J.G. Pitman. Sufficient statistics and intrinsic accuracy. Proceedings of the Cambridge Philosophical Society, 32, 1936.
  133. 133.Joaquin Qui˜nonero-Candela and Carl Edward Rasmussen. A unifying view of sparse approximate Gaussian process regression. Journal of Machine Learning Research, 6(Dec):1939–1959, 2005.
  134. 134.Rajesh Ranganath, Sean Gerrish, and David Blei. Black box variational inference. In Artificial Intelligence and Statistics, pages 814–822, 2014.
  135. 135.Rajesh Ranganath, Dustin Tran, Jaan Altosaar, and David Blei. Operator variational inference. In Advances in Neural Information Processing Systems, pages 496–504, 2016.
  136. 136.Jean-Baptiste Regli and Ricardo Silva. Alpha-beta divergence for variational inference. arXiv preprint arXiv:1805.01045, 2018.
  137. 137.Alfr´ed R´enyi. On measures of entropy and information. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics. The Regents of the University of California, 1961.
  138. 138.Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on International Conference on Machine Learning-Volume 37, pages 1530–1538, 2015.
  139. 139.Mathieu Ribatet, Daniel Cooley, and Anthony C Davison. Bayesian inference from composite likelihoods, with an application to spatial extremes. Statistica Sinica, pages 813–845, 2012.
  140. 140.Gareth O Roberts and Jeffrey S Rosenthal. Optimal scaling of discrete approximations to Langevin diffusions. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 60(1):255–268, 1998.
  141. 141.Gareth O Roberts, Andrew Gelman, and Walter R Gilks. Weak convergence and optimal scaling of random walk Metropolis algorithms. The annals of applied probability, 7(1):110–120, 1997.
  142. 142.Simone Rossi, Sebastien Marmin, and Maurizio Filippone. Walsh-Hadamard variational inference for Bayesian deep learning. arXiv preprint arXiv:1905.11248, 2019a.
  143. 143.Simone Rossi, Pietro Michiardi, and Maurizio Filippone. Good initializations of variational Bayes for deep models. In Proceedings of the 36th International Conference on Machine Learning, volume 97, pages 5487–5497, 2019b.
  144. 144.H˚avard Rue, Sara Martino, and Nicolas Chopin. Approximate Bayesian inference for latent Gaussian models by using integrated nested Laplace approximations. Journal of the royal statistical society: Series b (statistical methodology), 71(2):319–392, 2009.
  145. 145.Yunus Saat¸ci, Ryan D. Turner, and Carl E. Rasmussen. Gaussian process change point models. In Proceedings of the 27th International Conference on Machine Learning, pages 927–934, 2010.
  146. 146.Abhijoy Saha, Karthik Bharath, and Sebastian Kurtek. A geometric variational approach to bayesian inference. Journal of the American Statistical Association, pages 1–25, 2019.
  147. 147.Tim Salimans and David A Knowles. On using control variates with stochastic approximation for variational Bayes and its connection to stochastic linear regression. arXiv preprint arXiv:1401.1022, 2014.
  148. 148.Hugh Salimbeni and Marc Deisenroth. Doubly stochastic variational inference for deep Gaussian processes. In Advances in Neural Information Processing Systems, pages 4588–4599, 2017.
  149. 149.John Shawe-Taylor and Robert C Williamson. A PAC analysis of a Bayesian estimator. In Annual Workshop on Computational Learning Theory: Proceedings of the tenth annual conference on Computational learning theory, volume 6, pages 2–9, 1997.
  150. 150.Xiaotong Shen and Larry Wasserman. Rates of convergence of posterior distributions. The Annals of Statistics, 29(3):687–714, 2001.
  151. 151.Jiaxin Shi, Shengyang Sun, and Jun Zhu. Kernel implicit variational inference. In International Conference on Learning Representations, 2018.
  152. 152.Zhenming Shun and Peter McCullagh. Laplace approximation of high dimensional integrals. Journal of the Royal Statistical Society: Series B (Methodological), 57(4):749–760, 1995.
  153. 153.Douglas G Simpson. Minimum hellinger distance estimation for the analysis of count data. Journal of the American statistical Association, 82(399):802–807, 1987.
  154. 154.Edward Snelson and Zoubin Ghahramani. Sparse Gaussian processes using pseudo-inputs. In Advances in neural information processing systems, pages 1257–1264, 2006.
  155. 155.Elliott Sober. Evidence and evolution: The logic behind the science. Cambridge University Press, 2008.
  156. 156.Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther. Ladder variational autoencoders. In Advances in neural information processing systems, pages 3738–3746, 2016.
  157. 157.Nicholas Syring and Ryan Martin. Calibrating general posterior credible regions. Biometrika, 106(2):479–486, 2019.
  158. 158.Roy N Tamura and Dennis D Boos. Minimum hellinger distance estimation for multivariate location and covariance. Journal of the American Statistical Association, 81(393):223–229, 1986.
  159. 159.Louis C Tiao, Edwin V Bonilla, and Fabio Ramos. Cycle-consistent adversarial learning as approximate bayesian inference. arXiv preprint arXiv:1806.01771, 2018.
  160. 160.Luke Tierney and Joseph B Kadane. Accurate approximations for posterior moments and marginal densities. Journal of the american statistical association, 81(393):82–86, 1986.
  161. 161.Naftali Tishby, Fernando C Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000.
  162. 162.Michalis Titsias. Variational learning of inducing variables in sparse Gaussian processes. In Artificial Intelligence and Statistics, pages 567–574, 2009.
  163. 163.Michalis Titsias and Miguel L´azaro-Gredilla. Doubly stochastic variational Bayes for non-conjugate inference. In International Conference on Machine Learning, pages 1971–1979, 2014.
  164. 164.Udo v Toussaint, Silvio Gori, and Volker Dose. Invariance priors for bayesian feed-forward neural networks. Neural Networks, 19(10):1550–1557, 2006.
  165. 165.Dustin Tran, Rajesh Ranganath, and David M Blei. The variational gaussian process. In 4th International Conference on Learning Representations, ICLR 2016, 2016.
  166. 166.Dustin Tran, Rajesh Ranganath, and David Blei. Hierarchical implicit models and likelihood-free variational inference. In Advances in Neural Information Processing Systems, pages 5523–5533, 2017.
  167. 167.John W Tukey. A survey of sampling from contaminated distributions. Contributions to probability and statistics, pages 448–485, 1960.
  168. 168.R. E. Turner and M. Sahani. Two problems with variational expectation maximisation for time-series models. In Bayesian time series models. Cambridge University Press, 2011.
  169. 169.Ryan D. Turner, Steven Bottone, and Clay J. Stanek. Online variational approximations to non-exponential family change point models: with application to radar tracking. In Advances in Neural Information Processing Systems, pages 306–314, 2013.
  170. 170.Keyon Vafa. Training deep Gaussian processes with sampling. In NIPS 2016 Workshop on Advances in Approximate Bayesian Inference, 2016.
  171. 171.Tim Van Erven and Peter Harremos. R´enyi divergence and Kullback-Leibler divergence. IEEE Transactions on Information Theory, 60(7):3797–3820, 2014.
  172. 172.Cristiano Varin, Nancy Reid, and David Firth. An overview of composite likelihood methods. Statistica Sinica, pages 5–42, 2011.
  173. 173.Stephen Walker. New approaches to Bayesian consistency. The Annals of Statistics, 32(5):2028–2043, 2004.
  174. 174.Dilin Wang, Hao Liu, and Qiang Liu. Variational inference with tail-adaptive f-divergence. In Advances in Neural Information Processing Systems, pages 5742–5752, 2018.
  175. 175.Ke Alexander Wang, Geoff Pleiss, Jacob R Gardner, Stephen Tyree, Kilian Q Weinberger, and Andrew Gordon Wilson. Exact Gaussian processes on a million data points. arXiv preprint arXiv:1903.08114, 2019.
  176. 176.Yali Wang, Marcus Brubaker, Brahim Chaib-Draa, and Raquel Urtasun. Sequential inference for deep Gaussian process. In Artificial Intelligence and Statistics, pages 694–703, 2016.
  177. 177.Sumio Watanabe. Mathematical Theory of Bayesian Statistics. CRC Press, 2018.
  178. 178.Christopher KI Williams and Matthias Seeger. Using the Nystr¨om method to speed up kernel machines. In Advances in neural information processing systems, pages 682–688, 2001.
  179. 179.Robert C Wilson, Matthew R Nassar, and Joshua I Gold. Bayesian online learning of the hazard rate in change-point problems. Neural computation, 22(9):2452–2476, 2010.
  180. 180.Mike Wu, Noah Goodman, and Stefano Ermon. Differentiable antithetic sampling for variance reduction in stochastic variational inference. In Proceedings of Machine Learning Research, volume 89, pages 2877–2886, 2019.
  181. 181.Yue Yang, Ryan Martin, and Howard Bondell. Variational approximations using Fisher divergence. arXiv preprint arXiv:1905.05284, 2019.
  182. 182.Yun Yang, Debdeep Pati, and Anirban Bhattacharya. α-variational inference with statistical guarantees. arXiv preprint arXiv:1710.03266, 2017.
  183. 183.Yannis G Yatracos. Rates of convergence of minimum distance estimators and Kolmogorov’s entropy. The Annals of Statistics, pages 768–774, 1985.
  184. 184.Arnold Zellner. Maximal data information prior distributions. New developments in the applications of Bayesian methods, pages 211–232, 1977.
  185. 185.Arnold Zellner. Optimal information processing and Bayes’s theorem. The American Statistician, 42(4):278–280, 1988.
  186. 186.Guodong Zhang, Shengyang Sun, David Duvenaud, and Roger Grosse. Noisy natural gradient as variational inference. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80, pages 5852–5861, 2018.
  187. 187.Tong Zhang. From ϵ-entropy to kl-entropy: Analysis of minimum information complexity density estimation. The Annals of Statistics, 34(5):2180–2210, 2006.
  188. 188.Jun Zhu, Ning Chen, and Eric P Xing. Bayesian inference with posterior regularization and applications to infinite latent svms. The Journal of Machine Learning Research, 15(1):1799–1847, 2014.

Citation

MLA
Knoblauch, J., et al. “An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference”. Journal of Machine Learning Research, vol. 23, no. 132, 2022, pp. 1–9, https://www.jmlr.org/papers/v23/19-1047.html.
APA
Knoblauch, J., Jewson, J., & Damoulas, T. (2022). An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference. Journal of Machine Learning Research, 23(132), 1–109. https://www.jmlr.org/papers/v23/19-1047.html
Chicago
Knoblauch, J., J. Jewson, and T. Damoulas. 2022. “An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference”. Journal of Machine Learning Research 23 (132): 1–109. https://www.jmlr.org/papers/v23/19-1047.html.
Harvard
Knoblauch, J., Jewson, J. and Damoulas, T. (2022) “An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference”, Journal of Machine Learning Research, 23(132), pp. 1–109. Available at: https://www.jmlr.org/papers/v23/19-1047.html.
Vancouver
1. Knoblauch J, Jewson J, Damoulas T (2022) An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference. Journal of Machine Learning Research 23:1–109

BibTeX

@article{JMLR:v23:19-1047,
  author  = {Jeremias Knoblauch and Jack Jewson and Theodoros Damoulas},
  title   = {An Optimization-centric View on Bayes' Rule: Reviewing and Generalizing Variational Inference},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {132},
  pages   = {1--109},
  url     = {http://jmlr.org/papers/v23/19-1047.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/