Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling
Denis BlessingXiaogang JiaJohannes EsslingerFrancisco VargasGerhard Neumann
Establishes a standardized benchmarking framework and new evaluation metrics like evidence upper bounds to rigorously measure mode collapse and compare variational sampling methods across diverse tasks.
Modern computational applications in machine learning, statistics, and the physical sciences depend heavily on sampling from complex, high-dimensional probability distributions and calculating their normalizing constants. Although recent developments have merged Monte Carlo techniques with variational inference to handle challenging distributions, the field has lacked a standardized evaluation framework. Existing studies rely on fragmented metrics, such as the evidence lower bound, which fail to detect when algorithms miss entire regions of high probability—a failure mode known as mode collapse. The article's main objective is to establish a standardized benchmark suite that evaluates modern variational sampling methods across diverse tasks and introduces dedicated metrics to rigorously quantify mode coverage.
To conduct this evaluation, the article examines over a dozen leading sampling algorithms across three major classes: tractable density models, sequential importance sampling methods, and continuous-time diffusion-based methods. The benchmarking suite tests these algorithms against twelve synthetic and real-world target distributions, ranging from low-dimensional mixture models and complex funnel geometries to Bayesian logistic regressions and high-dimensional spatial statistics scaling up to 1,600 dimensions. Alongside traditional performance measures and probability transport distances, the article introduces entropic mode coverage, an entropy-based metric designed to quantify how evenly a model covers all modes of a target distribution.
The findings reveal several critical insights into algorithm performance. First, no single algorithm outperforms the others in all settings. Gaussian mixture models and annealed flow bootstrap methods achieve the highest evidence lower bounds on many real-world tasks and show exceptional sample efficiency, requiring orders of magnitude fewer function queries to converge. Second, mode collapse worsens sharply in higher dimensions; sequential importance sampling methods that capture modes effectively in low dimensions collapse almost entirely in high dimensions due to particle resampling. Third, diffusion-based methods maintain strong resilience against mode collapse in high-dimensional multimodal distributions, but they suffer from poor sample efficiency and high computational runtime due to evaluating gradients at every discretization step. Finally, traditional evaluation metrics prove misleading: algorithms suffering from complete mode collapse can still achieve deceptively high evidence lower bounds and reverse estimates.
These results demonstrate that practitioners cannot rely on traditional evidence lower bounds alone to assess model quality in risk-sensitive applications. Using standard optimization heuristics, such as pre-training or learning the initial proposal distribution end-to-end, inadvertently triggers mode collapse by prioritizing local optimization over broad exploration. Furthermore, the choice of sampling method introduces direct trade-offs between computational cost, wall-clock time, and mode discovery.
Based on these findings, decision-makers should tailor algorithm selection to their problem constraints. For scenarios where target evaluations are computationally expensive and dimensions are moderate, Gaussian mixture models offer the best balance of efficiency and accuracy. When target distributions exhibit high-dimensional multimodality and preventing mode collapse is critical, diffusion-based samplers are recommended despite their higher compute budget. For sequential Monte Carlo systems, adopting Hamiltonian dynamics over standard random-walk steps substantially improves robustness. Future work should focus on developing scalable methods that combine the sample efficiency of mixture models with the mode-preserving properties of diffusion frameworks. Although ground-truth mode detection remains difficult to verify on unconstrained real-world targets without known normalizers, the comparative trade-offs established in the article provide high confidence for selecting appropriate sampling architectures.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Its continuous-time SDE framework establishes the score-based diffusion samplers that the source evaluates against other sampling methods.
- Paper: Variational Inference with Normalizing Flows, Danilo Jimenez Rezende et al. (2015). Its normalizing-flow foundations clarify the tractable-density variational methods and ELBO-based evaluation central to the benchmark.
- Paper: Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm, Qiang Liu et al. (2016). Its particle-transport view of variational Bayesian inference provides the conceptual groundwork for understanding particle-based sampling methods compared in the source.
No sufficiently relevant recommendations were found.
