keyword
Markov chain Monte Carlo
Markov chain Monte Carlo is a class of computational algorithms used to generate representative samples from complex, often high-dimensional probability distributions that are difficult or impossible to sample directly. The approach works by constructing a Markov chain, which is a sequence of random steps where each new state depends solely on the current state, structured such that the long-term stationary distribution of the chain converges to the target distribution. Once the chain has run long enough to reach equilibrium, the accumulated states can be treated as samples from the target distribution to approximate integrals, calculate expectations, and estimate posterior distributions in Bayesian statistics, machine learning, and physical sciences. Common implementations include the Metropolis-Hastings algorithm, Gibbs sampling, and Hamiltonian Monte Carlo, which employ different strategies for proposing and accepting state transitions across the parameter space.
8 items

Compositional Score Modeling for Simulation-Based Inference
Tomas Geffner, George Papamakarios, Andriy Mnih
Why you should read this
Proposes a sample-efficient simulation-based inference framework using conditional score modeling that naturally aggregates an arbitrary number of observations at test time without relying on MCMC or variational approximations.
Neural Posterior Estimation methods for simulation-based inference can be ill-suited for dealing with posterior distributions obtained by conditioning on multiple observations, as they tend to require a large number of simulator calls to learn accurate approximations. In contrast, Neural Likelihood Estimation methods can handle multiple observations at inference time after learning from individual observations, but they rely on standard inference methods, such as MCMC or variational inference, which come with certain performance drawbacks. We introduce a new method based on conditional score modeling that enjoys the benefits of both approaches. We model the scores of the (diffused) posterior distributions induced by individual observations, and introduce a way of combining the learned scores to approximately sample from the target posterior distribution. Our approach is sample-efficient, can naturally aggregate multiple observations at inference time, and avoids the drawbacks of standard inference methods.
Added
2026-10-04

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching
Aaron J. Havens, Benjamin Kurt Miller, Bing Yan, Carles Domingo-Enrich, Anuroop Sriram, Daniel S. Levine, Brandon M. Wood, Bin Hu, Brandon Amos, Brian Karrer, Xiang Fu, Guan-Horng Liu, Ricky T. Q. Chen
Why you should read this
Introduces a scalable stochastic optimal control framework that trains diffusion samplers from unnormalized densities with dramatically fewer expensive energy evaluations, enabling efficient amortized molecular conformer generation.
We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry. Code & and Benchmarks provided at github.com/facebookresearch/adjoint_sampling.
Added
2026-10-03

Sampling-Based Accuracy Testing of Posterior Estimators for General Inference
Pablo Lemos, Adam Coogan, Yashar Hezaveh, Laurence Perreault Levasseur
Why you should read this
Introduces Tests of Accuracy with Random Points (TARP), a sample-only coverage testing framework that provides necessary and sufficient guarantees for validating high-dimensional generative posterior estimators without requiring explicit density evaluations.
Parameter inference, i.e. inferring the posterior distribution of the parameters of a statistical model given some data, is a central problem to many scientific disciplines. Generative models can be used as an alternative to Markov Chain Monte Carlo methods for conducting posterior inference, both in likelihood-based and simulation-based problems. However, assessing the accuracy of posteriors encoded in generative models is not straightforward. In this paper, we introduce ‘Tests of Accuracy with Random Points’ (TARP) coverage testing as a method to estimate coverage probabilities of generative posterior estimators. Our method differs from previously-existing coverage-based methods, which require posterior evaluations. We prove that our approach is necessary and sufficient to show that a posterior estimator is accurate. We demonstrate the method on a variety of synthetic examples, and show that TARP can be used to test the results of posterior inference analyses in high-dimensional spaces. We also show that our method can detect inaccurate inferences in cases where existing methods fail.
Added
2026-09-26

Bayesian probabilistic matrix factorization using Markov chain Monte Carlo
R. Salakhutdinov, A. Mnih
Why you should read this
Presents a fully Bayesian treatment of probabilistic matrix factorization using Markov chain Monte Carlo methods to automatically control model complexity and significantly improve recommendation accuracy on large-scale collaborative filtering datasets like Netflix.
Low-rank matrix approximation methods provide one of the simplest and most effective approaches to collaborative filtering. Such models are usually fitted to data by finding a MAP estimate of the model parameters, a procedure that can be performed efficiently even on very large datasets. However, unless the regularization parameters are tuned carefully, this approach is prone to overfitting because it finds a single point estimate of the parameters. In this paper we present a fully Bayesian treatment of the Probabilistic Matrix Factorization (PMF) model in which model capacity is controlled automatically by integrating over all model parameters and hyperparameters. We show that Bayesian PMF models can be efficiently trained using Markov chain Monte Carlo methods by applying them to the Netflix dataset, which consists of over 100 million movie ratings. The resulting models achieve significantly higher prediction accuracy than PMF models trained using MAP estimation.
Added
2026-09-24

Bayesian Learning via Stochastic Gradient Langevin Dynamics
Max Welling, Yee Whye Teh
Why you should read this
Introduces Stochastic Gradient Langevin Dynamics, a scalable framework that enables Bayesian posterior sampling on large datasets by injecting balanced Gaussian noise into mini-batch gradient updates without requiring full-dataset Metropolis-Hastings acceptance steps.
In this paper we propose a new framework for learning from large scale datasets based on iterative learning from small mini-batches. By adding the right amount of noise to a standard stochastic gradient optimization algorithm we show that the iterates will converge to samples from the true posterior distribution as we anneal the stepsize. This seamless transition between optimization and Bayesian posterior sampling provides an in-built protection against overfitting. We also propose a practical method for Monte Carlo estimates of posterior statistics which monitors a “sampling threshold” and collects samples after it has been surpassed. We apply the method to three models: a mixture of Gaussians, logistic regression and ICA with natural gradients.
Added
2026-09-12

Density estimation using Real NVP
Laurent Dinh, Jascha Narain Sohl-Dickstein, Samy Bengio
Why you should read this
Introduces Real NVP, a normalizing flow architecture that achieves exact log-likelihood computation, fast sampling, and tractable latent inference for high-dimensional generative modeling through invertible affine coupling layers.
Unsupervised learning of probabilistic models is a central yet challenging problem in machine learning. Specifically, designing models with tractable learning, sampling, inference and evaluation is crucial in solving this task. We extend the space of such models using real-valued non-volume preserving (real NVP) transformations, a set of powerful invertible and learnable transformations, resulting in an unsupervised learning algorithm with exact log-likelihood computation, exact sampling, exact inference of latent variables, and an interpretable latent space. We demonstrate its ability to model natural images on four datasets through sampling, log-likelihood evaluation and latent variable manipulations.
Added
2026-09-10
License
Published with permission

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew D. Hoffman, Andrew Gelman
Why you should read this
Introduces the No-U-Turn Sampler (NUTS), an algorithm that automates path length selection and step-size adaptation in Hamiltonian Monte Carlo to enable efficient, hands-free sampling for high-dimensional Bayesian models.
Hamiltonian Monte Carlo (HMC) is a Markov chain Monte Carlo (MCMC) algorithm that avoids the random walk behavior and sensitivity to correlated parameters that plague many MCMC methods by taking a series of steps informed by first-order gradient information. These features allow it to converge to high-dimensional target distributions much more quickly than simpler methods such as random walk Metropolis or Gibbs sampling. However, HMC's performance is highly sensitive to two user-specified parameters: a step size {\epsilon} and a desired number of steps L. In particular, if L is too small then the algorithm exhibits undesirable random walk behavior, while if L is too large the algorithm wastes computation. We introduce the No-U-Turn Sampler (NUTS), an extension to HMC that eliminates the need to set a number of steps L. NUTS uses a recursive algorithm to build a set of likely candidate points that spans a wide swath of the target distribution, stopping automatically when it starts to double back and retrace its steps. Empirically, NUTS perform at least as efficiently as and sometimes more efficiently than a well tuned standard HMC method, without requiring user intervention or costly tuning runs. We also derive a method for adapting the step size parameter {\epsilon} on the fly based on primal-dual averaging. NUTS can thus be used with no hand-tuning at all. NUTS is also suitable for applications such as BUGS-style automatic inference engines that require efficient "turnkey" sampling algorithms.
Added
2026-09-09


An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning
Maximilian Dax, Theo Heimel, Gilles Louppe
Why you should read this
Unifies Bayesian and frequentist foundations of neural simulation-based inference, explaining how methods like neural posterior and likelihood estimation solve scientific inverse problems alongside concrete validation strategies.
Simulation-based inference (SBI) with machine learning is an increasingly important tool for solving inverse problems in science and engineering, including parameter inference and the inversion of detector effects. We provide an overview of the Bayesian and frequentist statistical frameworks, describe how machine-learning-based SBI methods, such as neural posterior estimation and neural likelihood estimation, can be used for parameter estimation within these frameworks, and show that the same methods can also be applied to Empirical Bayes or unfolding tasks. We also discuss how to validate inference results and the limitations of SBI with machine learning.
Added
2026-08-24

