Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Markov chain Monte Carlo

Markov chain Monte Carlo is a class of computational algorithms used to generate representative samples from complex, often high-dimensional probability distributions that are difficult or impossible to sample directly. The approach works by constructing a Markov chain, which is a sequence of random steps where each new state depends solely on the current state, structured such that the long-term stationary distribution of the chain converges to the target distribution. Once the chain has run long enough to reach equilibrium, the accumulated states can be treated as samples from the target distribution to approximate integrals, calculate expectations, and estimate posterior distributions in Bayesian statistics, machine learning, and physical sciences. Common implementations include the Metropolis-Hastings algorithm, Gibbs sampling, and Hamiltonian Monte Carlo, which employ different strategies for proposing and accepting state transitions across the parameter space.

8 items

Compositional Score Modeling for Simulation-Based Inference

Compositional Score Modeling for Simulation-Based Inference

Tomas Geffner, George Papamakarios, Andriy Mnih

OrganizationsGoogleUniversity of Massachusetts Amherst

Why you should read this

Proposes a sample-efficient simulation-based inference framework using conditional score modeling that naturally aggregates an arbitrary number of observations at test time without relying on MCMC or variational approximations.

Neural Posterior Estimation methods for simulation-based inference can be ill-suited for dealing with posterior distributions obtained by conditioning on multiple observations, as they tend to require a large number of simulator calls to learn accurate approximations. In contrast, Neural Likelihood Estimation methods can handle multiple observations at inference time after learning from individual observations, but they rely on standard inference methods, such as MCMC or variational inference, which come with certain performance drawbacks. We introduce a new method based on conditional score modeling that enjoys the benefits of both approaches. We model the scores of the (diffused) posterior distributions induced by individual observations, and introduce a way of combining the learned scores to approximately sample from the target posterior distribution. Our approach is sample-efficient, can naturally aggregate multiple observations at inference time, and avoids the drawbacks of standard inference methods.

Added

2026-10-04

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Aaron J. Havens, Benjamin Kurt Miller, Bing Yan, Carles Domingo-Enrich, Anuroop Sriram, Daniel S. Levine, Brandon M. Wood, Bin Hu, Brandon Amos, Brian Karrer, Xiang Fu, Guan-Horng Liu, Ricky T. Q. Chen

OrganizationsMetaMicrosoftNew York UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces a scalable stochastic optimal control framework that trains diffusion samplers from unnormalized densities with dramatically fewer expensive energy evaluations, enabling efficient amortized molecular conformer generation.

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry. Code & and Benchmarks provided at github.com/facebookresearch/adjoint_sampling.

Added

2026-10-03

Sampling-Based Accuracy Testing of Posterior Estimators for General Inference

Sampling-Based Accuracy Testing of Posterior Estimators for General Inference

Pablo Lemos, Adam Coogan, Yashar Hezaveh, Laurence Perreault Levasseur

OrganizationsCIELA InstituteFlatiron InstituteMila – Québec Artificial Intelligence InstituteUniversité de Montréal

Why you should read this

Introduces Tests of Accuracy with Random Points (TARP), a sample-only coverage testing framework that provides necessary and sufficient guarantees for validating high-dimensional generative posterior estimators without requiring explicit density evaluations.

Parameter inference, i.e. inferring the posterior distribution of the parameters of a statistical model given some data, is a central problem to many scientific disciplines. Generative models can be used as an alternative to Markov Chain Monte Carlo methods for conducting posterior inference, both in likelihood-based and simulation-based problems. However, assessing the accuracy of posteriors encoded in generative models is not straightforward. In this paper, we introduce ‘Tests of Accuracy with Random Points’ (TARP) coverage testing as a method to estimate coverage probabilities of generative posterior estimators. Our method differs from previously-existing coverage-based methods, which require posterior evaluations. We prove that our approach is necessary and sufficient to show that a posterior estimator is accurate. We demonstrate the method on a variety of synthetic examples, and show that TARP can be used to test the results of posterior inference analyses in high-dimensional spaces. We also show that our method can detect inaccurate inferences in cases where existing methods fail.

Added

2026-09-26

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo

The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo

Matthew D. Hoffman, Andrew Gelman

OrganizationsColumbia University

Why you should read this

Introduces the No-U-Turn Sampler (NUTS), an algorithm that automates path length selection and step-size adaptation in Hamiltonian Monte Carlo to enable efficient, hands-free sampling for high-dimensional Bayesian models.

Hamiltonian Monte Carlo (HMC) is a Markov chain Monte Carlo (MCMC) algorithm that avoids the random walk behavior and sensitivity to correlated parameters that plague many MCMC methods by taking a series of steps informed by first-order gradient information. These features allow it to converge to high-dimensional target distributions much more quickly than simpler methods such as random walk Metropolis or Gibbs sampling. However, HMC's performance is highly sensitive to two user-specified parameters: a step size {\epsilon} and a desired number of steps L. In particular, if L is too small then the algorithm exhibits undesirable random walk behavior, while if L is too large the algorithm wastes computation. We introduce the No-U-Turn Sampler (NUTS), an extension to HMC that eliminates the need to set a number of steps L. NUTS uses a recursive algorithm to build a set of likely candidate points that spans a wide swath of the target distribution, stopping automatically when it starts to double back and retrace its steps. Empirically, NUTS perform at least as efficiently as and sometimes more efficiently than a well tuned standard HMC method, without requiring user intervention or costly tuning runs. We also derive a method for adapting the step size parameter {\epsilon} on the fly based on primal-dual averaging. NUTS can thus be used with no hand-tuning at all. NUTS is also suitable for applications such as BUGS-style automatic inference engines that require efficient "turnkey" sampling algorithms.

Added

2026-09-09

Creative Commons License