keyword
stochastic variational inference
Stochastic variational inference is an algorithm for approximating complex posterior probability distributions in Bayesian machine learning models when working with massive datasets. Standard variational inference frames Bayesian posterior approximation as an optimization problem, fitting a tractable family of distributions to the true posterior, but it typically requires processing the complete dataset at every iteration. Stochastic variational inference overcomes this computational bottleneck by utilizing stochastic optimization techniques, subsampling small mini-batches of data at each step to compute noisy yet unbiased estimates of the variational objective gradient. This mini-batch approach dramatically reduces per-iteration computational overhead and memory requirements, enabling scalable learning and uncertainty quantification across large-scale probabilistic models, including topic models, Gaussian processes, and deep probabilistic programs.
5 items

Deep Gaussian Processes
Andreas C. Damianou, Neil D. Lawrence
Why you should read this
Introduces deep Gaussian processes and a variational inference framework that enables fully Bayesian hierarchical modeling and automated architecture selection on small datasets without overfitting.
In this paper we introduce deep Gaussian process (GP) models. Deep GPs are a deep belief network based on Gaussian process mappings. The data is modeled as the output of a multivariate GP. The inputs to that Gaussian process are then governed by another GP. A single layer model is equivalent to a standard GP or the GP latent variable model (GP-LVM). We perform inference in the model by approximate variational marginalization. This results in a strict lower bound on the marginal likelihood of the model which we use for model selection (number of layers and nodes per layer). Deep belief networks are typically applied to relatively large data sets using stochastic gradient descent for optimization. Our fully Bayesian treatment allows for the application of deep models even when data is scarce. Model selection by our variational bound shows that a five layer hierarchy is justified even when modelling a digit data set containing only 150 examples.
Added
2026-09-25

Pyro: Deep Universal Probabilistic Programming
Eli Bingham, Jonathan P. Chen, Martin Jankowiak, Fritz Obermeyer, Neeraj Pradhan, Theofanis Karaletsos, Rohit Singh, Paul Szerlip, Paul Horsfall, Noah D. Goodman
Why you should read this
Presents Pyro, a universal probabilistic programming language built on PyTorch that couples deep neural networks with stochastic variational inference to scale expressive Bayesian models to high-dimensional datasets.
Pyro is a probabilistic programming language built on Python as a platform for developing advanced probabilistic models in AI research. To scale to large datasets and high-dimensional models, Pyro uses stochastic variational inference algorithms and probability distributions built on top of PyTorch, a modern GPU-accelerated deep learning framework. To accommodate complex or model-specific algorithmic behavior, Pyro leverages Poutine, a library of composable building blocks for modifying the behavior of probabilistic programs.
Added
2026-09-25

Gaussian Processes for Big Data
James Hensman, Nicolo Fusi, Neil D. Lawrence
Why you should read this
Develops a stochastic variational inference framework for Gaussian processes that overcomes cubic computational constraints, enabling mini-batch training on datasets with millions of data points.
We introduce stochastic variational inference for Gaussian process models. This enables the application of Gaussian process (GP) models to data sets containing millions of data points. We show how GPs can be vari- ationally decomposed to depend on a set of globally relevant inducing variables which factorize the model in the necessary manner to perform variational inference. Our ap- proach is readily extended to models with non-Gaussian likelihoods and latent variable models based around Gaussian processes. We demonstrate the approach on a simple toy problem and two real world data sets.
Added
2026-09-25

Stochastic variational inference
Matt Hoffman, David M. Blei, Chong Wang, John Paisley
Why you should read this
Develops stochastic variational inference, a scalable algorithm that applies stochastic optimization to variational bounds, enabling complex Bayesian models to perform posterior inference on massive datasets containing millions of documents.
We develop stochastic variational inference, a scalable algorithm for approximating posterior distributions. We develop this technique for a large class of probabilistic models and we demonstrate it with two probabilistic topic models, latent Dirichlet allocation and the hierarchical Dirichlet process topic model. Using stochastic variational inference, we analyze several large collections of documents: 300K articles from Nature, 1.8M articles from The New York Times, and 3.8M articles from Wikipedia. Stochastic inference can easily handle data sets of this size and outperforms traditional variational inference, which can only handle a smaller subset. (We also show that the Bayesian nonparametric topic model outperforms its parametric counterpart.) Stochastic variational inference lets us apply complex Bayesian models to massive data sets.
Added
2026-09-14

Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, Jasper Snoek
Why you should read this
Establishes a critical large-scale benchmark for evaluating state-of-the-art predictive uncertainty quantification methods in machine learning, specifically under challenging dataset shift conditions, revealing that methods marginalizing over models significantly outperform traditional approaches.
Modern machine learning methods including deep learning have achieved great success in predictive accuracy for supervised learning tasks, but may still fall short in giving useful estimates of their predictive {\em uncertainty}. Quantifying uncertainty is especially critical in real-world settings, which often involve input distributions that are shifted from the training distribution due to a variety of factors including sample bias and non-stationarity. In such settings, well calibrated uncertainty estimates convey information about when a model's output should (or should not) be trusted. Many probabilistic deep learning methods, including Bayesian-and non-Bayesian methods, have been proposed in the literature for quantifying predictive uncertainty, but to our knowledge there has not previously been a rigorous large-scale empirical comparison of these methods under dataset shift. We present a large-scale benchmark of existing state-of-the-art methods on classification problems and investigate the effect of dataset shift on accuracy and calibration. We find that traditional post-hoc calibration does indeed fall short, as do several other previous methods. However, some methods that marginalize over models give surprisingly strong results across a broad spectrum of tasks.
Added
2026-04-27
License
Published with permission
