keyword
approximate inference
Approximate inference is a set of computational techniques used in statistics and machine learning to estimate posterior probability distributions, marginal likelihoods, or expected values when exact calculations are mathematically intractable or computationally prohibitive. In complex probabilistic models and Bayesian networks, exact inference frequently requires calculating high-dimensional integrals or summing over an exponential combination of latent states. Approximate inference resolves this limitation by trading absolute precision for computational feasibility, primarily through deterministic optimization strategies such as variational inference and loopy belief propagation, or through stochastic sampling methods such as Markov chain Monte Carlo and Gibbs sampling. These methods allow scalable parameter estimation, uncertainty quantification, and predictive reasoning in complex, high-dimensional data settings.
14 items

Deep Gaussian Processes
Andreas C. Damianou, Neil D. Lawrence
Why you should read this
Introduces deep Gaussian processes and a variational inference framework that enables fully Bayesian hierarchical modeling and automated architecture selection on small datasets without overfitting.
In this paper we introduce deep Gaussian process (GP) models. Deep GPs are a deep belief network based on Gaussian process mappings. The data is modeled as the output of a multivariate GP. The inputs to that Gaussian process are then governed by another GP. A single layer model is equivalent to a standard GP or the GP latent variable model (GP-LVM). We perform inference in the model by approximate variational marginalization. This results in a strict lower bound on the marginal likelihood of the model which we use for model selection (number of layers and nodes per layer). Deep belief networks are typically applied to relatively large data sets using stochastic gradient descent for optimization. Our fully Bayesian treatment allows for the application of deep models even when data is scarce. Model selection by our variational bound shows that a five layer hierarchy is justified even when modelling a digit data set containing only 150 examples.
Added
2026-09-25

Pyro: Deep Universal Probabilistic Programming
Eli Bingham, Jonathan P. Chen, Martin Jankowiak, Fritz Obermeyer, Neeraj Pradhan, Theofanis Karaletsos, Rohit Singh, Paul Szerlip, Paul Horsfall, Noah D. Goodman
Why you should read this
Presents Pyro, a universal probabilistic programming language built on PyTorch that couples deep neural networks with stochastic variational inference to scale expressive Bayesian models to high-dimensional datasets.
Pyro is a probabilistic programming language built on Python as a platform for developing advanced probabilistic models in AI research. To scale to large datasets and high-dimensional models, Pyro uses stochastic variational inference algorithms and probability distributions built on top of PyTorch, a modern GPU-accelerated deep learning framework. To accommodate complex or model-specific algorithmic behavior, Pyro leverages Poutine, a library of composable building blocks for modifying the behavior of probabilistic programs.
Added
2026-09-25

Bayesian probabilistic matrix factorization using Markov chain Monte Carlo
R. Salakhutdinov, A. Mnih
Why you should read this
Presents a fully Bayesian treatment of probabilistic matrix factorization using Markov chain Monte Carlo methods to automatically control model complexity and significantly improve recommendation accuracy on large-scale collaborative filtering datasets like Netflix.
Low-rank matrix approximation methods provide one of the simplest and most effective approaches to collaborative filtering. Such models are usually fitted to data by finding a MAP estimate of the model parameters, a procedure that can be performed efficiently even on very large datasets. However, unless the regularization parameters are tuned carefully, this approach is prone to overfitting because it finds a single point estimate of the parameters. In this paper we present a fully Bayesian treatment of the Probabilistic Matrix Factorization (PMF) model in which model capacity is controlled automatically by integrating over all model parameters and hyperparameters. We show that Bayesian PMF models can be efficiently trained using Markov chain Monte Carlo methods by applying them to the Netflix dataset, which consists of over 100 million movie ratings. The resulting models achieve significantly higher prediction accuracy than PMF models trained using MAP estimation.
Added
2026-09-24

Loopy Belief Propagation for Approximate Inference: An Empirical Study
Kevin P. Murphy, Yair Weiss, Michael I. Jordan
Why you should read this
Evaluates loopy belief propagation across diverse Bayesian networks, revealing that Pearl's polytree algorithm often converges to accurate marginal approximations in loopy graphs while diagnosing the causes of oscillatory failure in complex real-world models.
Recently, researchers have demonstrated that loopy belief propagation - the use of Pearls polytree algorithm IN a Bayesian network WITH loops OF error- correcting this http URL most dramatic instance OF this IS the near Shannon - limit performance OF Turbo Codes codes whose decoding algorithm IS equivalent TO loopy belief propagation IN a chain - structured Bayesian network. IN this paper we ask : IS there something special about the error - correcting code context, OR does loopy propagation WORK AS an approximate inference schemeIN a more general setting? We compare the marginals computed using loopy propagation TO the exact ones IN four Bayesian network architectures, including two real - world networks : ALARM AND this http URL find that the loopy beliefs often converge AND WHEN they do, they give a good approximation TO the correct this http URL,ON the QMR network, the loopy beliefs oscillated AND had no obvious relationship TO the correct posteriors. We present SOME initial investigations INTO the cause OF these oscillations, AND show that SOME simple methods OF preventing them lead TO the wrong results.
Added
2026-09-18

Deep Bayesian Active Learning with Image Data
Yarin Gal, Riashat Islam, Zoubin Ghahramani
Why you should read this
Develops a Bayesian active learning framework for high-dimensional image data that uses uncertainty estimation in deep convolutional networks to substantially reduce the amount of labeled training data needed for vision tasks.
Even though active learning forms an important pillar of machine learning, deep learning tools are not prevalent within it. Deep learning poses several difficulties when used in an active learning setting. First, active learning (AL) methods generally rely on being able to learn and update models from small amounts of data. Recent advances in deep learning, on the other hand, are notorious for their dependence on large amounts of data. Second, many AL acquisition functions rely on model uncertainty, yet deep learning methods rarely represent such model uncertainty. In this paper we combine recent advances in Bayesian deep learning into the active learning framework in a practical way. We develop an active learning framework for high dimensional data, a task which has been extremely challenging so far, with very sparse existing literature. Taking advantage of specialised models such as Bayesian convolutional neural networks, we demonstrate our active learning techniques with image data, obtaining a significant improvement on existing active learning approaches. We demonstrate this on both the MNIST dataset, as well as for skin cancer diagnosis from lesion images (ISIC2016 task).
Added
2026-09-17

A Unifying View of Sparse Approximate Gaussian Process Regression
Joaquin Quiñonero-Candela, Carl Edward Rasmussen
We provide a new unifying view, including all existing proper probabilistic sparse approximations for Gaussian process regression. Our approach relies on expressing the effective prior which the methods are using. This allows new insights to be gained, and highlights the relationship between existing methods. It also allows for a clear theoretically justified ranking of the closeness of the known approximations to the corresponding full GPs. Finally we point directly to designs of new better sparse approximations, combining the best of the existing strategies, within attractive computational constraints.
Added
2026-09-16

Stochastic variational inference
Matt Hoffman, David M. Blei, Chong Wang, John Paisley
Why you should read this
Develops stochastic variational inference, a scalable algorithm that applies stochastic optimization to variational bounds, enabling complex Bayesian models to perform posterior inference on massive datasets containing millions of documents.
We develop stochastic variational inference, a scalable algorithm for approximating posterior distributions. We develop this technique for a large class of probabilistic models and we demonstrate it with two probabilistic topic models, latent Dirichlet allocation and the hierarchical Dirichlet process topic model. Using stochastic variational inference, we analyze several large collections of documents: 300K articles from Nature, 1.8M articles from The New York Times, and 3.8M articles from Wikipedia. Stochastic inference can easily handle data sets of this size and outperforms traditional variational inference, which can only handle a smaller subset. (We also show that the Bayesian nonparametric topic model outperforms its parametric counterpart.) Stochastic variational inference lets us apply complex Bayesian models to massive data sets.
Added
2026-09-14

Incorporating Non-local Information into Information Extraction Systems by Gibbs Sampling
Jenny Rose Finkel, Trond Grenager, Christopher D. Manning
Why you should read this
Proposes using Gibbs sampling with simulated annealing during inference to incorporate long-range consistency constraints into conditional random fields, boosting information extraction accuracy on standard benchmarks without requiring structural retraining.
Most current statistical natural language processing models use only local features so as to permit dynamic programming in inference, but this makes them unable to fully account for the long distance structure that is prevalent in language use. We show how to solve this dilemma with Gibbs sampling, a simple Monte Carlo method used to perform approximate inference in factored probabilistic models. By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference. We use this technique to augment an existing CRF-based information extraction system with long-distance dependency models, enforcing label consistency and extraction template consistency constraints. This technique results in an error reduction of up to 9% over state-of-the-art systems on two established information extraction tasks.
Added
2026-09-11

An Introduction to Variational Methods for Graphical Models
MICHAEL I. JORDAN, ZOUBIN GHAHRAMANI, TOMMI S. JAAKKOLA, LAWRENCE K. SAUL
Why you should read this
Establishes a foundational framework for approximate inference and learning in graphical models by using convex duality to convert intractable probabilistic calculations into tractable optimization problems.
This paper presents a tutorial introduction to the use of variational methods for inference and learning in graphical models (Bayesian networks and Markov random fields). We present a number of examples of graphical models, including the QMR-DT database, the sigmoid belief network, the Boltzmann machine, and several variants of hidden Markov models, in which it is infeasible to run exact inference algorithms. We then introduce variational methods, which exploit laws of large numbers to transform the original graphical model into a simplified graphical model in which inference is efficient. Inference in the simplified model provides bounds on probabilities of interest in the original model. We describe a general framework for generating variational transformations based on convex duality. Finally we return to the examples and demonstrate how variational algorithms can be formulated in each case.
Added
2026-09-10

Decision-Making with Auto-Encoding Variational Bayes
Romain Lopez, Pierre Boyeau, N. Yosef, Michael I. Jordan, J. Regier
Why you should read this
Proposes using multiple importance sampling over varied approximate posteriors to correct decision-making biases in auto-encoding variational Bayes, outperforming existing methods on complex single-cell RNA sequencing benchmarks.
To make decisions based on a model fit with auto-encoding variational Bayes (AEVB), practitioners often let the variational distribution serve as a surrogate for the posterior distribution. This approach yields biased estimates of the expected risk, and therefore leads to poor decisions for two reasons. First, the model fit with AEVB may not equal the underlying data distribution. Second, the variational distribution may not equal the posterior distribution under the fitted model. We explore how fitting the variational distribution based on several objective functions other than the ELBO, while continuing to fit the generative model based on the ELBO, affects the quality of downstream decisions. For the probabilistic principal component analysis model, we investigate how importance sampling error, as well as the bias of the model parameter estimates, varies across several approximate posteriors when used as proposal distributions. Our theoretical results suggest that a posterior approximation distinct from the variational distribution should be used for making decisions. Motivated by these theoretical results, we propose learning several approximate proposals for the best model and combining them using multiple importance sampling for decision-making. In addition to toy examples, we present a full-fledged case study of single-cell RNA sequencing. In this challenging instance of multiple hypothesis testing, our proposed approach surpasses the current state of the art.
Added
2026-09-05


HoloClean: Holistic Data Repairs with Probabilistic Inference
Theodoros Rekatsinas, Xu Chu, Ihab F. Ilyas, Christopher Ré
Why you should read this
Introduces HoloClean, a holistic data repairing framework that unifies qualitative and quantitative signals through probabilistic inference to achieve significant improvements in accuracy and scalability over state-of-the-art methods.
We introduce HoloClean, a framework for holistic data repairing driven by probabilistic inference. HoloClean unifies qualitative data repairing, which relies on integrity constraints or external data sources, with quantitative data repairing methods, which leverage statistical properties of the input data. Given an inconsistent dataset as input, HoloClean automatically generates a probabilistic program that performs data repairing. Inspired by recent theoretical advances in probabilistic inference, we introduce a series of optimizations which ensure that inference over HoloClean’s probabilistic model scales to instances with millions of tuples. We show that HoloClean finds data repairs with an average precision of ~ 90% and an average recall of above ~ 76% across a diverse array of datasets exhibiting different types of errors. This yields an average F1 improvement of more than 2x against state-of-the-art methods.
Added
2026-04-18

Adversarially Learned Inference
Vincent Dumoulin, Ishmael Belghazi, Ben Poole, Olivier Mastropietro, Alex Lamb, Martin Arjovsky, Aaron Courville
Why you should read this
Proposes a framework to simultaneously learn a generator and an inference network (encoder) within the adversarial game.
We introduce the adversarially learned inference (ALI) model, which jointly learns a generation network and an inference network using an adversarial process. The generation network maps samples from stochastic latent variables to the data space while the inference network maps training examples in data space to the space of latent variables. An adversarial game is cast between these two networks and a discriminative network is trained to distinguish between joint latent/data-space samples from the generative network and joint samples from the inference network. We illustrate the ability of the model to learn mutually coherent inference and generation networks through the inspections of model samples and reconstructions and confirm the usefulness of the learned representations by obtaining a performance competitive with state-of-the-art on the semi-supervised SVHN and CIFAR10 tasks.
Added
2026-03-07

Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Danilo Jimenez Rezende, Shakir Mohamed, Daan Wierstra
Why you should read this
Develops a new algorithm for scalable inference and learning in deep generative models, enabling the creation of realistic samples, accurate data imputation, and high-dimensional visualization through a novel stochastic backpropagation technique.
We marry ideas from deep neural networks and approximate Bayesian inference to derive a generalised class of deep, directed generative models, endowed with a new algorithm for scalable inference and learning. Our algorithm introduces a recognition model to represent approximate posterior distributions, and that acts as a stochastic encoder of the data. We develop stochastic back-propagation -- rules for back-propagation through stochastic variables -- and use this to develop an algorithm that allows for joint optimisation of the parameters of both the generative and recognition model. We demonstrate on several real-world data sets that the model generates realistic samples, provides accurate imputations of missing data and is a useful tool for high-dimensional data visualisation.
Added
2026-03-01


Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio
Why you should read this
Introduces a novel adversarial framework that enables the creation of highly realistic generative models without complex inference or Markov chains, revolutionizing synthetic data generation.
We propose a new framework for estimating generative models via an adversarial process, in which we simultaneously train two models: a generative model G that captures the data distribution, and a discriminative model D that estimates the probability that a sample came from the training data rather than G. The training procedure for G is to maximize the probability of D making a mistake. This framework corresponds to a minimax two-player game. In the space of arbitrary functions G and D, a unique solution exists, with G recovering the training data distribution and D equal to 1/2 everywhere. In the case where G and D are defined by multilayer perceptrons, the entire system can be trained with backpropagation. There is no need for any Markov chains or unrolled approximate inference networks during either training or generation of samples. Experiments demonstrate the potential of the framework through qualitative and quantitative evaluation of the generated samples.
Added
2026-02-14
License
Published with permission
