Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm

Qiang LiuDilin Wang

article2016NeurIPS1,320 citations

Introduces Stein Variational Gradient Descent, a deterministic particle-based inference algorithm that unifies optimization and Bayesian computation by iteratively transporting particles to match target distributions via functional gradient descent on the KL divergence.

Listen

Modern data-driven decision-making increasingly relies on Bayesian inference to reason under uncertainty in complex models. However, exact calculation of posterior distributions is computationally intractable. Traditional sampling techniques, such as Markov Chain Monte Carlo, struggle to scale to large datasets and assess convergence reliably, while conventional variational methods require restrictive distributional assumptions and complex, model-specific derivations. There is a strong need for a fast, scalable, and general-purpose algorithm that automates Bayesian inference without requiring specialized manual derivations for every new model.

The article develops and evaluates Stein Variational Gradient Descent, a general-purpose deterministic algorithm that transports a set of particles to approximate complex probability distributions. The algorithm bridges the gap between fast optimization and full Bayesian inference, reducing to standard gradient ascent for point estimation when using a single particle while capturing full posterior uncertainty as the number of particles increases.

The researchers established a theoretical connection showing that the derivative of the Kullback-Leibler divergence under smooth transformations directly yields an optimal descent direction using the kernelized Stein discrepancy. To evaluate practical performance, the authors conducted empirical experiments across varied benchmarks, including a multimodal synthetic distribution, large-scale Bayesian logistic regression on the Covertype dataset with over 580,000 data points, and Bayesian neural networks across ten regression benchmarks.

The key findings demonstrate that Stein Variational Gradient Descent consistently matches or outperforms leading alternative methods in both accuracy and computational efficiency. First, on multimodal synthetic targets, the deterministic repulsive force successfully pushes particles to discover distant probability modes and achieves mean squared error comparable to or lower than exact Monte Carlo sampling. Second, on large-scale logistic regression, the proposed algorithm achieved higher test classification accuracy and faster convergence than standard stochastic sampling and variational baselines. Third, on Bayesian neural network benchmarks, the method achieved lower prediction error and higher test log-likelihood across nine of ten datasets while delivering dramatic speed improvements—reducing training runtime by up to 90% compared to specialized probabilistic backpropagation baselines.

These results show that organizations can achieve high-quality uncertainty quantification at the speed and simplicity of standard gradient-based optimization. Because the algorithm avoids complex matrix inversions, Jacobian determinant calculations, and weight degeneracy issues, it lowers the computational cost and implementation barrier for deploying Bayesian models in large-scale machine learning workflows.

Technical leaders and practitioners should consider adopting this particle-based framework as a drop-in counterpart to gradient descent for tasks requiring robust uncertainty estimation. When scaling to massive particle sets, teams should utilize data mini-batching, parallelized particle updates, and kernel approximation techniques. Future work should focus on establishing formal theoretical convergence rates and expanding empirical validation to modern deep learning architectures.

  • Paper: Variational Inference: A Review for Statisticians, David M. Blei et al. (2016). Provides a comprehensive foundation in variational inference and KL divergence minimization that SVGD directly generalizes through functional particle optimization.
  • Paper: Bayesian Learning via Stochastic Gradient Langevin Dynamics, Max Welling et al. (2011). Introduces scalable gradient-based sampling via Langevin dynamics, serving as the core particle-based sampling baseline that SVGD aims to improve upon with deterministic repulsive dynamics.
  • Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). Explores continuous transformations and flows for flexible variational approximations, establishing the conceptual bridge between density transformations and particle transport.
  • Paper: Stochastic variational inference, Matt Hoffman et al. (2012). Presents stochastic gradient optimization schemes for variational inference, which SVGD builds upon for general non-parametric inference.
  • Paper: An Introduction to Variational Methods for Graphical Models, MICHAEL I. JORDAN et al. (1999). Outlines foundational principles of variational approximation methods in graphical models that motivate non-parametric particle-based extensions.
  • Paper: Computational Optimal Transport, Gabriel Peyré et al. (2018). Examines optimal transport metrics and Wasserstein spaces, providing rigorous theoretical and computational machinery to analyze SVGD as a gradient flow of the KL divergence.
  • Paper: Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, Xingchao Liu et al. (2023). Extends particle transport and continuous-flow generative modeling by constructing straight ODE paths between distributions, building on velocity field transport principles.
  • Paper: Demystifying MMD GANs, Mikołaj Bińkowski et al. (2018). Investigates kernel-based discrepancy metrics and gradient estimators in distribution matching, directly related to the kernelized Stein discrepancy updates in SVGD.
  • Paper: Stochastic Gradient Descent over P2, Maria Oprea et al. (2026). Develops formal optimization theory for stochastic gradient descent directly over probability distributions in Wasserstein space, generalizing SVGD-style particle systems.
  • Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). Unifies transport equations and velocity fields connecting probability distributions via ODE and SDE flows, expanding on continuous particle transport concepts.
Cover for Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm

Abstract

We propose a general purpose variational inference algorithm that forms a natural counterpart of gradient descent for optimization. Our method iteratively transports a set of particles to match the target distribution, by applying a form of functional gradient descent that minimizes the KL divergence. Empirical studies are performed on various real world models and datasets, on which our method is competitive with existing state-of-the-art methods. The derivation of our method is based on a new theoretical result that connects the derivative of KL divergence under smooth transforms with Stein's identity and a recently proposed kernelized Stein discrepancy, which is of independent interest.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Variational Inference Using Smooth Transforms
  • 3.1 Stein Operator as the Derivative of KL Divergence
  • 3.2 Stein Variational Gradient Descent
  • 4 Related Works
  • 5 Experiments
  • 6 Conclusion
  • References
  • A Proof of Theorem
  • B Proof of Theorem
  • C Connection with de Bruijn’s identity and Fisher Divergence
  • D Additional Experiments
  • D.1 Bayesian Logistic Regression on Small Datasets

Knowls

  1. Knowl 1 — Stein Variational Gradient Descent Algorithm

    algorithm

    Stein Variational Gradient Descent (SVGD) is a deterministic particle-based variational inference algorithm. It transports a set of nn particles {xi}i=1n⊂Rd\{x_i\}_{i=1}^n \subset \mathbb{R}^d to approximate a continuous target probability distribution p(x)=pˉ(x)/Zp(x) = \bar{p}(x)/Z, where the unnormalized posterior pˉ(x)=p0(x)∏k=1Np(Dk∣x)\bar{p}(x) = p_0(x) \prod_{k=1}^N p(D_k \mid x) is differentiable, without requiring the normalization constant ZZ.

    At each iteration ℓ\ell, every particle xiℓx_i^\ell is updated along a direction ϕ^∗(xiℓ)\hat{\phi}^*(x_i^\ell) determined by the empirical expectation over all particles using a positive definite kernel k(x,x′)k(x, x'):

    Input: Target unnormalized score function ∇xlog⁡p(x)\nabla_x \log p(x), positive definite kernel k(x,x′)k(x, x'), step size schedule {ϵℓ}\{\epsilon_\ell\}, and initial particles {xi0}i=1n\{x_i^0\}_{i=1}^n.
    Output: Final particle set {xi}i=1n\{x_i\}_{i=1}^n approximating p(x)p(x).
    for iteration ℓ=0,1,2,…\ell = 0, 1, 2, \dots do
        for each particle i∈{1,…,n}i \in \{1, \dots, n\} do
            ϕ^∗(xiℓ)←1n∑j=1n[k(xjℓ,xiℓ)∇xjℓlog⁡p(xjℓ)+∇xjℓk(xjℓ,xiℓ)]\hat{\phi}^*(x_i^\ell) \leftarrow \frac{1}{n} \sum_{j=1}^n [k(x_j^\ell, x_i^\ell) \nabla_{x_j^\ell} \log p(x_j^\ell) + \nabla_{x_j^\ell} k(x_j^\ell, x_i^\ell)]
            xiℓ+1←xiℓ+ϵℓϕ^∗(xiℓ)x_i^{\ell+1} \leftarrow x_i^\ell + \epsilon_\ell \hat{\phi}^*(x_i^\ell)
        end for
    end for

    For large datasets where NN is large, the exact target score ∇xlog⁡p(x)\nabla_x \log p(x) is approximated using a mini-batch Ω⊂{1,…,N}\Omega \subset \{1, \dots, N\}: ∇xlog⁡p(x)≈∇xlog⁡p0(x)+N∣Ω∣∑k∈Ω∇xlog⁡p(Dk∣x)\nabla_x \log p(x) \approx \nabla_x \log p_0(x) + \frac{N}{|\Omega|} \sum_{k \in \Omega} \nabla_x \log p(D_k \mid x)

    The computational complexity per iteration is O(n2)O(n^2) kernel evaluations and nn target score evaluations, which can be parallelized across the nn particles.

  2. Knowl 2 — First-Order Variation of KL Divergence under Smooth Transformations

    theoretical result

    Let p(x)p(x) and q(x)q(x) be continuously differentiable probability densities supported on X⊆Rd\mathcal{X} \subseteq \mathbb{R}^d. Consider an incremental transformation T(x)=x+ϵϕ(x)T(x) = x + \epsilon \phi(x), where ϕ:X→Rd\phi: \mathcal{X} \to \mathbb{R}^d is a smooth perturbation direction vector and ϵ∈R\epsilon \in \mathbb{R} is a perturbation magnitude. Let q[T](z)q_{[T]}(z) denote the probability density of the pushforward variable z=T(x)z = T(x) when x∼q(x)x \sim q(x).

    The derivative of the Kullback-Leibler (KL) divergence KL(q[T]∥p)\text{KL}(q_{[T]} \parallel p) with respect to ϵ\epsilon evaluated at ϵ=0\epsilon = 0 is given by: ∇ϵKL(q[T]∥p)∣ϵ=0=−Ex∼q[trace(Apϕ(x))]\left. \nabla_\epsilon \text{KL}(q_{[T]} \parallel p) \right|_{\epsilon=0} = -\mathbb{E}_{x \sim q}[\text{trace}(\mathcal{A}_p \phi(x))] where Ap\mathcal{A}_p is the Stein operator acting on the vector function ϕ(x)=[ϕ1(x),…,ϕd(x)]⊤\phi(x) = [\phi_1(x), \dots, \phi_d(x)]^\top, defined as: Apϕ(x)=ϕ(x)∇xlog⁡p(x)⊤+∇xϕ(x)\mathcal{A}_p \phi(x) = \phi(x) \nabla_x \log p(x)^\top + \nabla_x \phi(x)

    When ϕ\phi belongs to the Stein class of pp (satisfying the boundary condition p(x)ϕ(x)=0p(x)\phi(x) = 0 for all x∈∂Xx \in \partial \mathcal{X} if X\mathcal{X} is compact, or lim⁡∥x∥→∞ϕ(x)p(x)=0\lim_{\|x\| \to \infty} \phi(x)p(x) = 0 when X=Rd\mathcal{X} = \mathbb{R}^d), Ex∼p[trace(Apϕ(x))]=0\mathbb{E}_{x \sim p}[\text{trace}(\mathcal{A}_p \phi(x))] = 0 by Stein's identity.

  3. Knowl 3 — Optimal Perturbation Direction as RKHS Functional Gradient

    theoretical result

    Let Hd=H×⋯×H\mathcal{H}^d = \mathcal{H} \times \dots \times \mathcal{H} be the vector-valued reproducing kernel Hilbert space (RKHS) associated with a positive definite kernel k(x,x′)k(x, x'). For a smooth density q(x)q(x) and target density p(x)p(x) on X⊆Rd\mathcal{X} \subseteq \mathbb{R}^d, consider perturbations of the identity map T(x)=x+ϵϕ(x)T(x) = x + \epsilon \phi(x) with ϕ∈Hd\phi \in \mathcal{H}^d.

    The direction of steepest descent that maximizes the negative derivative −∇ϵKL(q[T]∥p)∣ϵ=0-\left. \nabla_\epsilon \text{KL}(q_{[T]} \parallel p) \right|_{\epsilon=0} within the RKHS ball B={ϕ∈Hd:∥ϕ∥Hd2≤S(q,p)}\mathcal{B} = \{\phi \in \mathcal{H}^d : \|\phi\|_{\mathcal{H}^d}^2 \le \mathbb{S}(q, p)\} is: ϕq,p∗(⋅)=Ex∼q[k(x,⋅)∇xlog⁡p(x)+∇xk(x,⋅)]\phi^*_{q,p}(\cdot) = \mathbb{E}_{x \sim q}[k(x, \cdot) \nabla_x \log p(x) + \nabla_x k(x, \cdot)]

    Evaluating the derivative along this optimal direction yields the negative Kernelized Stein Discrepancy (KSD): ∇ϵKL(q[x+ϵϕq,p∗]∥p)∣ϵ=0=−S(q,p)=−∥ϕq,p∗∥Hd2\left. \nabla_\epsilon \text{KL}(q_{[x + \epsilon \phi^*_{q,p}]} \parallel p) \right|_{\epsilon=0} = -\mathbb{S}(q, p) = -\|\phi^*_{q,p}\|_{\mathcal{H}^d}^2 where S(q,p)=max⁡∥ϕ∥Hd≤1(Ex∼q[trace(Apϕ(x))])2\mathbb{S}(q, p) = \max_{\|\phi\|_{\mathcal{H}^d} \le 1} \left( \mathbb{E}_{x \sim q}[\text{trace}(\mathcal{A}_p \phi(x))] \right)^2.

    Equivalently, considering the functional F[f]=KL(q[x+f(x)]∥p)F[f] = \text{KL}(q_{[x + f(x)]} \parallel p) for f∈Hdf \in \mathcal{H}^d, its functional gradient at zero perturbation f=0f = 0 is: ∇fKL(q[x+f(x)]∥p)∣f=0=−ϕq,p∗(x)\left. \nabla_f \text{KL}(q_{[x + f(x)]} \parallel p) \right|_{f=0} = -\phi^*_{q,p}(x)

  4. Knowl 4 — Two-Force Interaction Dynamics in SVGD Updates

    model/method

    The empirical velocity field that transports particles {xi}i=1n\{x_i\}_{i=1}^n in Stein Variational Gradient Descent, ϕ^∗(x)=1n∑j=1n[k(xj,x)∇xjlog⁡p(xj)+∇xjk(xj,x)]\hat{\phi}^*(x) = \frac{1}{n} \sum_{j=1}^n \left[ k(x_j, x) \nabla_{x_j} \log p(x_j) + \nabla_{x_j} k(x_j, x) \right] decomposes into two distinct interacting forces:

    1. Smoothed drive force: The first term, 1n∑j=1nk(xj,x)∇xjlog⁡p(xj)\frac{1}{n} \sum_{j=1}^n k(x_j, x) \nabla_{x_j} \log p(x_j), is a kernel-weighted average of the target score functions evaluated across all particles. It drives particles towards high-probability regions and modes of p(x)p(x).

    2. Repulsive force: The second term, 1n∑j=1n∇xjk(xj,x)\frac{1}{n} \sum_{j=1}^n \nabla_{x_j} k(x_j, x), prevents particles from collapsing onto local modes. For an RBF kernel k(x,x′)=exp⁡(−1h∥x−x′∥22)k(x, x') = \exp(-\frac{1}{h}\|x - x'\|_2^2), this term equals 1n∑j=1n2h(x−xj)k(xj,x)\frac{1}{n} \sum_{j=1}^n \frac{2}{h}(x - x_j) k(x_j, x), which pushes particle xx away from neighboring particles xjx_j that have high kernel similarity k(xj,x)k(x_j, x).

    Special Cases and Limiting Behaviors:

    • Single Particle (n=1n=1): When n=1n=1, if the kernel satisfies ∇xk(x,x)=0\nabla_x k(x, x) = 0 (which holds for RBF kernels), the repulsive force vanishes and the algorithm reduces exactly to standard gradient ascent for Maximum A Posteriori (MAP) estimation: xℓ+1=xℓ+ϵℓ∇xlog⁡p(xℓ)x^{\ell+1} = x^\ell + \epsilon_\ell \nabla_x \log p(x^\ell).
    • Zero Bandwidth (h→0h \to 0): The repulsive term vanishes and all particles decouple into nn independent gradient ascent chains converging to local MAP modes.
  5. Knowl 5 — Adaptive Bandwidth Selection via the Median Heuristic in SVGD

    model/method

    When using the radial basis function (RBF) kernel k(x,x′)=exp⁡(−1h∥x−x′∥22)k(x, x') = \exp(-\frac{1}{h}\|x - x'\|_2^2) with nn particles {xi}i=1n⊂Rd\{x_i\}_{i=1}^n \subset \mathbb{R}^d, the kernel bandwidth hh is updated adaptively at each iteration according to: h=med2log⁡nh = \frac{\text{med}^2}{\log n} where med\text{med} is the median of all pairwise Euclidean distances ∥xi−xj∥2\|x_i - x_j\|_2 across the current particle set {xi}i=1n\{x_i\}_{i=1}^n.

    This choice ensures that for each particle xix_i, the sum of kernel weights from all particles satisfies: ∑j=1nk(xi,xj)≈nexp⁡(−med2h)=nexp⁡(−log⁡n)=1\sum_{j=1}^n k(x_i, x_j) \approx n \exp\left( -\frac{\text{med}^2}{h} \right) = n \exp(-\log n) = 1 This dynamically balances the particle's own gradient contribution (k(xi,xi)=1k(x_i, x_i) = 1) against the aggregate repulsive and smoothing influence exerted by the remaining n−1n-1 particles.

  6. Knowl 6 — Deterministic de Bruijn Identity via Fisher Divergence

    theoretical result

    Let p(x)p(x) and q(x)q(x) be smooth densities on Rd\mathbb{R}^d. For the incremental transform T(x)=x+ϵϕ(x)T(x) = x + \epsilon \phi(x), choosing the perturbation direction to be the score difference function, ϕq,p(x)=∇xlog⁡p(x)−∇xlog⁡q(x)\phi_{q,p}(x) = \nabla_x \log p(x) - \nabla_x \log q(x) causes the first variation of the KL divergence to evaluate to the negative Fisher divergence F(q,p)\mathcal{F}(q, p): ∇ϵKL(q[T]∥p)∣ϵ=0=−F(q,p)=−Ex∼q[∥∇xlog⁡p(x)−∇xlog⁡q(x)∥22]\left. \nabla_\epsilon \text{KL}(q_{[T]} \parallel p) \right|_{\epsilon=0} = -\mathcal{F}(q, p) = -\mathbb{E}_{x \sim q}\left[ \|\nabla_x \log p(x) - \nabla_x \log q(x)\|_2^2 \right]

    This result provides a deterministic counterpart to the classical de Bruijn identity, which establishes the same relationship between the derivative of KL divergence and Fisher divergence using a stochastic Gaussian perturbation T(x)=x+ϵξT(x) = x + \sqrt{\epsilon} \xi with ξ∼N(0,I)\xi \sim \mathcal{N}(0, I).

  7. Knowl 7 — Benchmark Comparison of SVGD and Probabilistic Backpropagation on Bayesian Neural Networks

    data/table

    Bayesian neural network regression benchmarks comparing Probabilistic Backpropagation (PBP) and Stein Variational Gradient Descent (SVGD, Our Method) on single-hidden-layer neural networks (50 hidden units for all datasets except 100 for Protein and Year; ReLU activation max⁡(0,x)\max(0, x); Gamma(1,0.1)\text{Gamma}(1, 0.1) prior on inverse covariances). SVGD uses n=20n=20 particles, AdaGrad with momentum, and mini-batch size 100 (1000 for Year). Datasets use 90% train / 10% test splits averaged over 20 random trials (5 for Protein, 1 for Year).

    Avg. Test RMSE Avg. Test LL Avg. Time (Secs)
    Dataset PBP Our Method PBP Our Method PBP Ours
    Boston 2.977±0.0932.977 \pm 0.093 2.957±0.0992.957 \pm 0.099 −2.579±0.052-2.579 \pm 0.052 −2.504±0.029-2.504 \pm 0.029 18 16
    Concrete 5.506±0.1035.506 \pm 0.103 5.324±0.1045.324 \pm 0.104 −3.137±0.021-3.137 \pm 0.021 −3.082±0.018-3.082 \pm 0.018 33 24
    Energy 1.734±0.0511.734 \pm 0.051 1.374±0.0451.374 \pm 0.045 −1.981±0.028-1.981 \pm 0.028 −1.767±0.024-1.767 \pm 0.024 25 21
    Kin8nm 0.098±0.0010.098 \pm 0.001 0.090±0.0010.090 \pm 0.001 0.901±0.0100.901 \pm 0.010 0.984±0.0080.984 \pm 0.008 118 41
    Naval 0.006±0.0000.006 \pm 0.000 0.004±0.0000.004 \pm 0.000 3.735±0.0043.735 \pm 0.004 4.089±0.0124.089 \pm 0.012 173 49
    Combined 4.052±0.0314.052 \pm 0.031 4.033±0.0334.033 \pm 0.033 −2.819±0.008-2.819 \pm 0.008 −2.815±0.008-2.815 \pm 0.008 136 51
    Protein 4.623±0.0094.623 \pm 0.009 4.606±0.0134.606 \pm 0.013 −2.950±0.002-2.950 \pm 0.002 −2.947±0.003-2.947 \pm 0.003 682 68
    Wine 0.614±0.0080.614 \pm 0.008 0.609±0.0100.609 \pm 0.010 −0.931±0.014-0.931 \pm 0.014 −0.925±0.014-0.925 \pm 0.014 26 22
    Yacht 0.778±0.0420.778 \pm 0.042 0.864±0.0520.864 \pm 0.052 −1.211±0.044-1.211 \pm 0.044 −1.225±0.042-1.225 \pm 0.042 25 25
    Year 8.733±NA8.733 \pm \text{NA} 8.684±NA8.684 \pm \text{NA} −3.586±NA-3.586 \pm \text{NA} −3.580±NA-3.580 \pm \text{NA} 7777 684

    SVGD achieves lower test RMSE and higher test log-likelihood (LL) across 9 of the 10 benchmark datasets compared to PBP, while requiring substantially less computation time (e.g., Year runtime is reduced from 7777s to 684s, and Naval runtime from 173s to 49s).

  8. Knowl 8 — Scalable Bayesian Logistic Regression on Covertype Dataset

    empirical result

    In a Bayesian logistic regression experiment on the binary Covertype dataset (N=581,012N = 581,012 data points, d=54d = 54 features) with prior p0(w∣α)=N(w;0,α−1I)p_0(w \mid \alpha) = \mathcal{N}(w; 0, \alpha^{-1} I) and p0(α)=Gamma(α;1,0.01)p_0(\alpha) = \text{Gamma}(\alpha; 1, 0.01), SVGD was evaluated against parallel stochastic gradient Langevin dynamics (Parallel SGLD, n=100n=100 parallel chains), sequential SGLD, particle mirror descent (PMD), and doubly stochastic variational inference (DSVI).

    Experimental setup: Data partitioned into 80% training and 20% testing (averaged over 50 random trials), mini-batch size ∣Ω∣=50|\Omega| = 50, step size ϵt=a/(t+1)0.55\epsilon_t = a / (t + 1)^{0.55}, and n=100n=100 particles for SVGD, parallel SGLD, and PMD. PMD used h=0.002×med2h = 0.002 \times \text{med}^2 and required initialized weights to be scaled down by 10 to avoid numerical instability.

    Results:

    • SVGD achieved higher testing accuracy faster than all baselines, reaching ≈75%\approx 75\% accuracy within 1 epoch.
    • When varying particle count n∈{1,10,50,250}n \in \{1, 10, 50, 250\} at iteration 3000 (≈0.32\approx 0.32 epochs), SVGD accuracy increased monotonically with nn, outperforming parallel SGLD across all particle sizes.
    • Matrix parallelization in MATLAB allowed parallel SGLD and SVGD (n=100n=100) to run only ≈3×\approx 3\times slower per iteration than single-chain sequential SGLD despite performing 100 times more likelihood evaluations.
  9. Knowl 9 — Mode Capture and Expectation Estimation in 1D Gaussian Mixture

    empirical result

    In a 1D synthetic experiment, SVGD was tested on a bimodal target distribution p(x)=13N(x;−2,1)+23N(x;2,1)p(x) = \frac{1}{3}\mathcal{N}(x; -2, 1) + \frac{2}{3}\mathcal{N}(x; 2, 1) initialized with particles drawn from q0(x)=N(x;−10,1)q_0(x) = \mathcal{N}(x; -10, 1), creating an initial distribution with virtually zero probability overlap with p(x)p(x).

    Results:

    • Despite initialization far from the target mass, SVGD with n=100n=100 particles successfully drove particles to escape the first local mode at x=−2x=-2 and recover both modes with the correct 1:21:2 probability proportions by 500 iterations, whereas Particle Mirror Descent suffered from weight degeneracy.
    • When estimating test expectations Ep[h(x)]\mathbb{E}_p[h(x)] for h(x)=xh(x) = x, h(x)=x2h(x) = x^2, and h(x)=cos⁡(ωx+b)h(x) = \cos(\omega x + b) (with ω∼N(0,1)\omega \sim \mathcal{N}(0, 1) and b∼Uniform([0,2π])b \sim \text{Uniform}([0, 2\pi])), SVGD achieved comparable or lower mean squared error (MSE) across sample sizes n∈{10,50,250}n \in \{10, 50, 250\} compared to exact i.i.d. Monte Carlo sampling, because the deterministic repulsive force distributes particles more evenly across the support than random samples.

Coverage note — Omitted the toy 2D Bayesian logistic regression visualization and the non-parametric variational inference (NPV) / NUTS baseline comparison across the 8 small UCI benchmarks (Appendix D.1), as results across methods on these small datasets were virtually identical and added no distinct algorithmic insights beyond the Covertype benchmark.

References

  1. 1.M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley. Stochastic variational inference. JMLR, 2013.
  2. 2.M. Welling and Y. W. Teh. Bayesian learning via stochastic gradient Langevin dynamics. In ICML, 2011.
  3. 3.D. Maclaurin and R. P. Adams. Firefly Monte Carlo: Exact MCMC with subsets of data. In UAI, 2014.
  4. 4.R. Ranganath, S. Gerrish, and D. M. Blei. Black box variational inference. In AISTATS, 2014.
  5. 5.S. Gershman, M. Hoffman, and D. Blei. Nonparametric variational inference. In ICML, 2012.
  6. 6.A. Kucukelbir, R. Ranganath, A. Gelman, and D. Blei. Automatic variational inference in STAN. In NIPS, 2015.
  7. 7.B. Dai, N. He, H. Dai, and L. Song. Provable Bayesian inference via particle mirror descent. In AISTATS, 2016.
  8. 8.Q. Liu, J. D. Lee, and M. I. Jordan. A kernelized Stein discrepancy for goodness-of-fit tests and model evaluation. ICML, 2016.
  9. 9.C. J. Oates, M. Girolami, and N. Chopin. Control functionals for Monte Carlo integration. Journal of the Royal Statistical Society, Series B, 2016.
  10. 10.K. Chwialkowski, H. Strathmann, and A. Gretton. A kernel test of goodness-of-fit. ICML, 2016.
  11. 11.J. Gorham and L. Mackey. Measuring sample quality with Stein’s method. In NIPS, pages 226–234, 2015.
  12. 12.C. Villani. Optimal transport: old and new, volume 338. Springer Science & Business Media, 2008.
  13. 13.D. J. Rezende and S. Mohamed. Variational inference with normalizing flows. In ICML, 2015.
  14. 14.Y. Marzouk, T. Moselhy, M. Parno, and A. Spantini. An introduction to sampling via measure transport. arXiv preprint arXiv:1602.05023, 2016.
  15. 15.P. Del Moral. Mean field simulation for Monte Carlo integration. CRC Press, 2013.
  16. 16.M. Kac. Probability and related topics in physical sciences, volume 1. American Mathematical Soc., 1959.
  17. 17.A. Rahimi and B. Recht. Random features for large-scale kernel machines. In NIPS, pages 1177–1184, 2007.
  18. 18.D. Tran, R. Ranganath, and D. M. Blei. Variational Gaussian process. In ICLR, 2016.
  19. 19.M. Titsias and M. Lázaro-Gredilla. Doubly stochastic variational Bayes for non-conjugate inference. In ICML, pages 1971–1979, 2014.
  20. 20.E. Challis and D. Barber. Affine independent variational inference. In NIPS, 2012.
  21. 21.S. Han, X. Liao, D. B. Dunson, and L. Carin. Variational Gaussian copula inference. In AISTATS, 2016.
  22. 22.D. Tran, D. M. Blei, and E. M. Airoldi. Copula variational inference. In NIPS, 2015.
  23. 23.C. M. B. N. Lawrence and T. J. M. I. Jordan. Approximating posterior distributions in belief networks using mixtures. In NIPS, 1998.
  24. 24.T. S. Jaakkola and M. I. Jordon. Improving the mean field approximation via the use of mixture distributions. In Learning in graphical models, pages 163–173. MIT Press, 1999.
  25. 25.N. D. Lawrence. Variational inference in probabilistic models. PhD thesis, University of Cambridge, 2001.
  26. 26.T. D. Kulkarni, A. Saeedi, and S. Gershman. Variational particle approximations. arXiv preprint arXiv:1402.5715, 2014.
  27. 27.C. Robert and G. Casella. Monte Carlo statistical methods. Springer Science & Business Media, 2013.
  28. 28.A. Smith, A. Doucet, N. de Freitas, and N. Gordon. Sequential Monte Carlo methods in practice. Springer Science & Business Media, 2013.
  29. 29.M. D. Hoffman and A. Gelman. The No-U-Turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo. The Journal of Machine Learning Research, 15(1):1593–1623, 2014.
  30. 30.J. M. Hernández-Lobato and R. P. Adams. Probabilistic backpropagation for scalable learning of Bayesian neural networks. In ICML, 2015.
  31. 31.C. Stein, P. Diaconis, S. Holmes, G. Reinert, et al. Use of exchangeable pairs in the analysis of simulations. In Stein’s Method, pages 1–25. Institute of Mathematical Statistics, 2004.
  32. 32.Y. Li, J. M. Hernández-Lobato, and R. E. Turner. Stochastic expectation propagation. In NIPS, 2015.
  33. 33.Y. Li and R. E. Turner. Variational inference with Renyi divergence. arXiv preprint arXiv:1602.02311, 2016.
  34. 34.Y. Gal and Z. Ghahramani. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. arXiv preprint arXiv:1506.02142, 2015.
  35. 35.T. M. Cover and J. A. Thomas. Elements of information theory. John Wiley & Sons, 2012.
  36. 36.S. Lyu. Interpretation and generalization of score matching. In UAI, pages 359–366, 2009.

Citation

MLA
Liu, Q., and D. Wang. “Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm”. arXiv, 2016, http://arxiv.org/abs/1608.04471v3.
APA
Liu, Q., & Wang, D. (2016). Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm. arXiv. http://arxiv.org/abs/1608.04471v3
Chicago
Liu, Q., and D. Wang. 2016. “Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm”. arXiv. http://arxiv.org/abs/1608.04471v3.
Harvard
Liu, Q. and Wang, D. (2016) “Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1608.04471v3.
Vancouver
1. Liu Q, Wang D (2016) Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm. arXiv

BibTeX

@article{liu2016stein,
  title = {Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm},
  author = {Liu, Qiang and Wang, Dilin},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1608.04471v3},
  eprint = {1608.04471}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors