Pathfinder: Parallel quasi-Newton variational inference
Lu ZhangBob CarpenterAndrew GelmanAki Vehtari
Introduces Pathfinder, a parallel variational inference algorithm that constructs normal approximations along quasi-Newton optimization trajectories to produce high-quality posterior draws with one to two orders of magnitude fewer gradient evaluations than standard methods.
Bayesian statistical computation is widely used across scientific and industrial applications, but modern complex models often require substantial computational time and resources. Standard Markov chain Monte Carlo (MCMC) sampling methods provide asymptotically exact solutions but are computationally intensive and frequently suffer from slow initial warmup phases or become trapped in local minor modes. Faster approximate methods, such as automatic differentiation variational inference (ADVI), often struggle to converge due to the high variance of stochastic gradient estimates and inherently sequential optimization loops.
The article introduces Pathfinder, a fast variational inference algorithm designed to sample approximately from differentiable probability distributions. Its primary objective is to evaluate whether constructing local normal approximations along quasi-Newton optimization trajectories can produce accurate, robust posterior approximations and rapid initializations for MCMC while drastically reducing computation time.
The authors evaluated the algorithm on a benchmark suite of 20 diverse Bayesian models spanning generalized linear models, Gaussian processes, differential equations, and time series, supplemented by high-dimensional case studies. Pathfinder uses a quasi-Newton optimization trajectory (specifically L-BFGS) to move from random initial points toward high-probability regions. Along this path, it constructs local normal approximations using inverse Hessian estimates for local curvature, evaluates the evidence lower bound in parallel to pick the optimal approximation, and draws samples. A multi-path extension runs several trajectories in parallel and applies Pareto-smoothed importance resampling to handle non-normal distributions and filter out inferior local modes.
Key findings show that Pathfinder produces approximate draws that range from slightly worse to significantly better than mean-field and dense ADVI, and comparable to short chains of Hamiltonian Monte Carlo, as measured by 1-Wasserstein distance. Computationally, Pathfinder achieved these results requiring one to two orders of magnitude fewer probability density and gradient evaluations—evaluating approximately 30 to 50 times fewer operations than baseline methods on average. In an applied Gaussian process case study, initializing MCMC chains with Pathfinder successfully avoided minor modes and cut single-chain warmup runtimes by roughly two-thirds.
These results demonstrate that incorporating curvature information via quasi-Newton methods dramatically accelerates approximate Bayesian computation and enhances workflow reliability. Using Pathfinder as a standalone variational approximation or as a drop-in replacement for the initial phase of MCMC warmup reduces computing costs, lowers developer wait times, and diminishes the risk of chains stalling in minor modes.
The authors recommend adopting multi-path Pathfinder with Pareto-smoothed importance resampling as standard practice for fast exploratory analysis and MCMC initialization. Future extensions could incorporate stochastic gradient subsampling to handle massive datasets and integrate Bayesian optimization to fine-tune point selection along optimization paths.
The article notes that confidence should be tempered in high-dimensional distributions with severe multimodality, highly non-normal geometries, or weak parameter identifiability, where approximations can become overly concentrated. For such complex models, the authors advise using Pathfinder primarily as an initialization tool followed by full stochastic MCMC sampling.
- Paper: The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo, Matthew D. Hoffman et al. (2011). It introduces the No-U-Turn Sampler (dynamic Hamiltonian Monte Carlo), which serves as a primary benchmark and initialization target for the fast posterior approximations evaluated in Pathfinder.
- Paper: Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm, Qiang Liu et al. (2016). It introduces deterministic particle optimization for approximate Bayesian inference, providing foundational principles for gradient-based posterior approximation without conventional MCMC.
- Paper: Bayesian Learning via Stochastic Gradient Langevin Dynamics, Max Welling et al. (2011). It establishes core concepts of bridging continuous optimization with posterior sampling, forming key background for trajectory-based approximate Bayesian computation.
- Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). It develops flexible non-Gaussian variational distributions through smooth transformations, motivating the need for expressive yet computationally efficient variational approximations.
- Paper: Identifying and attacking the saddle point problem in high-dimensional non-convex optimization, Yann Dauphin et al. (2014). It examines saddle-point landscapes in high-dimensional non-convex optimization, explaining the pathology that Pathfinder's multi-path quasi-Newton search and importance resampling are designed to overcome.
- Paper: Continual Learning via Sequential Function-Space Variational Inference, Tim G. J. Rudner et al. (2022). It extends variational inference principles into sequential and continual learning settings by formulating Bayesian function-space updates.
- Paper: On the geometry of Stein variational gradient descent, Andrew B. Duncan et al. (2023). It provides a rigorous geometric analysis of particle-based variational sampling dynamics, offering deeper theoretical insight into optimization-driven posterior approximation.
- Paper: Score-based Data Assimilation, François Rozet et al. (2023). It applies advanced score-based trajectory sampling to solve high-dimensional Bayesian inverse problems and data assimilation.
- Paper: An Introduction to Bayesian and Frequentist Simulation-Based Inference with Machine Learning, Maximilian Dax et al. (2026). It offers an extensive synthesis of machine learning-driven approximate inference strategies for complex Bayesian posterior estimation where direct likelihood evaluation is intractable.
