Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling

Denis BlessingXiaogang JiaJohannes EsslingerFrancisco VargasGerhard Neumann

article2024ICML54 citations

Establishes a standardized benchmarking framework and new evaluation metrics like evidence upper bounds to rigorously measure mode collapse and compare variational sampling methods across diverse tasks.

Listen

Modern computational applications in machine learning, statistics, and the physical sciences depend heavily on sampling from complex, high-dimensional probability distributions and calculating their normalizing constants. Although recent developments have merged Monte Carlo techniques with variational inference to handle challenging distributions, the field has lacked a standardized evaluation framework. Existing studies rely on fragmented metrics, such as the evidence lower bound, which fail to detect when algorithms miss entire regions of high probability—a failure mode known as mode collapse. The article's main objective is to establish a standardized benchmark suite that evaluates modern variational sampling methods across diverse tasks and introduces dedicated metrics to rigorously quantify mode coverage.

To conduct this evaluation, the article examines over a dozen leading sampling algorithms across three major classes: tractable density models, sequential importance sampling methods, and continuous-time diffusion-based methods. The benchmarking suite tests these algorithms against twelve synthetic and real-world target distributions, ranging from low-dimensional mixture models and complex funnel geometries to Bayesian logistic regressions and high-dimensional spatial statistics scaling up to 1,600 dimensions. Alongside traditional performance measures and probability transport distances, the article introduces entropic mode coverage, an entropy-based metric designed to quantify how evenly a model covers all modes of a target distribution.

The findings reveal several critical insights into algorithm performance. First, no single algorithm outperforms the others in all settings. Gaussian mixture models and annealed flow bootstrap methods achieve the highest evidence lower bounds on many real-world tasks and show exceptional sample efficiency, requiring orders of magnitude fewer function queries to converge. Second, mode collapse worsens sharply in higher dimensions; sequential importance sampling methods that capture modes effectively in low dimensions collapse almost entirely in high dimensions due to particle resampling. Third, diffusion-based methods maintain strong resilience against mode collapse in high-dimensional multimodal distributions, but they suffer from poor sample efficiency and high computational runtime due to evaluating gradients at every discretization step. Finally, traditional evaluation metrics prove misleading: algorithms suffering from complete mode collapse can still achieve deceptively high evidence lower bounds and reverse estimates.

These results demonstrate that practitioners cannot rely on traditional evidence lower bounds alone to assess model quality in risk-sensitive applications. Using standard optimization heuristics, such as pre-training or learning the initial proposal distribution end-to-end, inadvertently triggers mode collapse by prioritizing local optimization over broad exploration. Furthermore, the choice of sampling method introduces direct trade-offs between computational cost, wall-clock time, and mode discovery.

Based on these findings, decision-makers should tailor algorithm selection to their problem constraints. For scenarios where target evaluations are computationally expensive and dimensions are moderate, Gaussian mixture models offer the best balance of efficiency and accuracy. When target distributions exhibit high-dimensional multimodality and preventing mode collapse is critical, diffusion-based samplers are recommended despite their higher compute budget. For sequential Monte Carlo systems, adopting Hamiltonian dynamics over standard random-walk steps substantially improves robustness. Future work should focus on developing scalable methods that combine the sample efficiency of mixture models with the mode-preserving properties of diffusion frameworks. Although ground-truth mode detection remains difficult to verify on unconstrained real-world targets without known normalizers, the comparative trade-offs established in the article provide high confidence for selecting appropriate sampling architectures.

No sufficiently relevant recommendations were found.

Cover for Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling

Abstract

Monte Carlo methods, Variational Inference, and their combinations play a pivotal role in sampling from intractable probability distributions. However, current studies lack a unified evaluation framework, relying on disparate performance measures and limited method comparisons across diverse tasks, complicating the assessment of progress and hindering the decision-making of practitioners. In response to these challenges, our work introduces a benchmark that evaluates sampling methods using a standardized task suite and a broad range of performance criteria. Moreover, we study existing metrics for quantifying mode collapse and introduce novel metrics for this purpose. Our findings provide insights into strengths and weaknesses of existing sampling methods, serving as a valuable reference for future developments. The code is publicly available here.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Quantifying Mode-Collapse
  • 4 Benchmarking Methods
  • 5 Benchmarking Target Densities
  • 6 Hyperparameters and Tuning
  • 7 Experiments
  • 7.1 Evaluation on Synthetic Target Densities
  • 7.2 Evaluation on Real Target Densities
  • 8 Discussion and Conclusion
  • 9 Conclusion
  • References
  • A Performance Criteria Details
  • A.1 Density-Ratio-Based Criteria
  • A.2 Integral Probability Metrics
  • A.3 Extending the Entropic Mode Coverage
  • B Details on Unnormalized Importance Weights / Density Ratios
  • C Benchmark Target Details
  • C.1 Bayesian Logistic Regression
  • C.2 Random Effect Regression
  • C.3 Time Series Models
  • C.4 Spatial Statistics
  • C.5 Synthetic Targets
  • D Algorithms and Parameter Choices
  • E Further Experimental results
  • F Ablation Studies
  • F.1 Ablation Study: Batchsize and Number of Particles
  • F.2 Ablation Study: Number of Temperatures / Timesteps T
  • F.3 Ablation Study: Sequential Monte Carlo Design Choices
  • F.4 Ablation Study: Initial Model Support
  • F.5 Ablation Study: Langevin Methods
  • F.6 Ablation Study: Transport Flow Type
  • F.7 Ablation Study: Gradient Guidance
  • F.8 Ablation Study: Pre-training the Proposal/Base-Distribution π0\pi_{0}

Knowls

  1. Knowl 1 — Entropy-based mode coverage for balanced and unbalanced modes

    definition

    Entropic mode coverage (EMC) is a heuristic for measuring whether samples from a model distribution qθq_\theta represent the mode descriptors of a target. Let {ξi}i=1M\{\xi_i\}_{i=1}^M be disjoint descriptors of the target’s modes, and let ai(x)=p(x∈ξi)a_i(x)=p(x\in\xi_i) be an auxiliary probability that assigns sample xx to descriptor ii. EMC is the model-sample expectation of the entropy of this assignment distribution, normalized by the maximum entropy: EMC=Ex∼qθ[−∑i=1Mai(x)log⁡Mai(x)]\mathrm{EMC}=\mathbb{E}_{x\sim q_\theta}\left[-\sum_{i=1}^M a_i(x)\log_M a_i(x)\right]. Its Monte Carlo estimate averages this quantity over model samples. The normalization gives values in [0,1][0,1]: zero means samples are consistently assigned to one descriptor, while one means assignments are uniform across descriptors. EMC requires known mode descriptors and is not an appropriate optimum when the target’s mode probabilities are unequal. For that case, the paper proposes expected Jensen–Shannon divergence, EJS=Ex∼qθ[DJS(a(x)∥r)]\mathrm{EJS}=\mathbb{E}_{x\sim q_\theta}[D_{\mathrm{JS}}(a(x)\|r)], where rr is the known target mode-probability vector; with base-2 logarithms, EJS lies in [0,1][0,1], equals zero when the distributions match, and equals one when they have disjoint probability mass.

  2. Knowl 2 — Forward evidence upper bound as a mode-collapse diagnostic

    equation

    Let π(x)=γ(x)/Z\pi(x)=\gamma(x)/Z be a normalized target density on Rd\mathbb{R}^d, with evaluable unnormalized density γ\gamma, normalizer ZZ, and model density qθq_\theta. The importance ratio is w(x)=γ(x)/qθ(x)w(x)=\gamma(x)/q_\theta(x). The evidence upper bound is EUBO=Ex∼π[log⁡w(x)]\mathrm{EUBO}=\mathbb{E}_{x\sim\pi}[\log w(x)], so DKL(π∥qθ)=EUBO−log⁡ZD_{\mathrm{KL}}(\pi\|q_\theta)=\mathrm{EUBO}-\log Z and EUBO≥log⁡Z\mathrm{EUBO}\geq\log Z, with equality only when qθ=πq_\theta=\pi. Because it averages under the target, the EUBO penalizes model density that is too small in target-supported regions and can reveal missing modes. By contrast, the ELBO, Ex∼qθ[log⁡w(x)]\mathbb{E}_{x\sim q_\theta}[\log w(x)], and other model-sample criteria can look favorable even when the model omits target modes. For models with latent variables, the extended EUBO is an upper bound on the marginal EUBO and can be looser, making direct comparisons across model types difficult.

  3. Knowl 3 — Benchmark composition and evaluation protocol

    experimental setup

    The benchmark compares 12 sampling methods: mean-field Gaussian variational inference (MFVI), Gaussian-mixture variational inference (GMMVI), sequential Monte Carlo (SMC), annealed flow transport (AFT), continual repeated AFT (CRAFT), flow annealed importance sampling bootstrap (FAB), Monte Carlo diffusion (MCD), Langevin diffusion variational inference (LDVI), path integral sampler (PIS), time-reversed diffusion sampler (DIS), denoising diffusion sampler (DDS), and general bridge sampler (GBS). Targets span Bayesian logistic regression (Credit, Cancer, Ionosphere, Sonar; dimensions 25, 31, 35, and 61), random-effects regression (Seeds; 26), a Brownian time-series model (32), a log Gaussian Cox process (LGCP; 1600), a Funnel target (10), image-density targets (Digits; 196, and Fashion; 784), and Gaussian and Student-t mixtures (MoG and MoS; dimensions 2, 50, and 200). Availability of exact normalizers, target samples, and mode descriptors varies by target. Where the required target information was available, the benchmark reports 2-Wasserstein distance, maximum mean discrepancy, forward and reverse normalizer-estimation errors, ELBO and EUBO, effective sample sizes, and EMC. Evaluation criteria were computed 100 times during training and smoothed with a length-5 running average; four random seeds were used, and the best result from each run was averaged. Criteria used 2000 samples. EMC was reported at the training point with the highest ELBO. Main training settings included 100,000 gradient steps for MFVI, 2000 particles and 128 annealing steps for most sequential methods, and 128 time steps with batches of 2000 and 40,000 gradient steps for diffusion-based methods.

  4. Knowl 4 — Mode collapse worsens on high-dimensional mixtures

    empirical result

    On the MoG and MoS targets, EMC showed a marked dimension-dependent change in mode coverage. At dimension d=2d=2, every evaluated method except MFVI was reported to sample from all modes, with EMC approximately one. At d=50d=50 and d=200d=200, methods outside the diffusion-based group generally exhibited mode collapse. This result demonstrates that success on low-dimensional multimodal targets did not reliably predict mode coverage at higher dimensions.

  5. Knowl 5 — Flow-based methods best captured the Funnel target

    empirical result

    On the Funnel target, most evaluated samplers represented its overall funnel shape but had difficulty generating samples at the narrow neck and broad opening. FAB and GMMVI were the exceptions reported to capture those regions. They also achieved the strongest reverse and forward estimates of the normalizer and the best evidence-bound results among the compared methods on this target.

  6. Knowl 6 — Image targets expose different coverage and density-estimation failures

    empirical result

    On the 14×14 Digits target, most methods found the majority of modes and produced visually plausible samples, consistent with generally high EMC. Many methods, particularly diffusion-based methods, nevertheless had poor normalizer estimates and weak ELBO and EUBO results. On the 28×28 Fashion target, methods either showed mode collapse or produced low-quality samples. Notably, methods suffering from mode collapse could still have the smallest forward and reverse normalizer-estimation errors, illustrating that those errors alone do not establish good mode coverage.

  7. Knowl 7 — Real-target comparisons favor GMMVI and FAB, with a scalability caveat

    empirical result

    For Credit, Seeds, Cancer, Brownian, Ionosphere, Sonar, and LGCP, the benchmark lacked ground-truth normalizers or target samples, so it reported ELBO rather than sample-based distances or forward criteria. GMMVI performed well across the tasks and often outperformed more complex variational Monte Carlo methods, but encountered a memory scalability problem on the 1600-dimensional LGCP target. FAB performed well on a majority of the real-target tasks. These comparisons assess ELBO performance; without target samples or normalizers, they do not establish which methods best avoid mode collapse.

  8. Knowl 8 — SMC resampling and transition-kernel choice affect coverage

    empirical result

    In an SMC ablation on MoG targets, not using resampling prevented the observed mode collapse across dimensions, with EMC approximately one; the authors also report that this choice could worsen ELBO. Hamiltonian Monte Carlo (HMC) outperformed Metropolis–Hastings (MH) on both ELBO and EMC across the tested dimensions when the kernels were compared using the same number of target-function evaluations and hand-tuned step sizes with rejection rates around 0.65. Thus, resampling and transition-kernel choice materially affected both sample quality and mode coverage.

  9. Knowl 9 — Initial support trades evidence tightness against mode coverage

    empirical result

    On multimodal MoG targets, varying the scale of the Gaussian initial proposal or base distribution π0=N(0,σ02I)\pi_0=\mathcal{N}(0,\sigma_0^2 I) revealed an exploration–exploitation trade-off for sequential and diffusion-based methods. Smaller initial support often gave tighter ELBO values but could confine samples to few modes, producing EMC near zero; larger support generally improved coverage but loosened the ELBO, particularly in higher dimensions. Learning proposal parameters end-to-end by maximizing the extended ELBO, or pretraining a Gaussian base distribution using MFVI, could yield higher ELBO values while inducing mode collapse. The study therefore found that a better ELBO did not necessarily imply broader coverage.

  10. Knowl 10 — More annealing or diffusion steps tighten evidence bounds

    empirical result

    An ablation on MoG targets varied the number TT of annealing temperatures for sequential methods and discretization steps for SDE-based methods. Increasing TT generally improved ELBO and EUBO values across the tested dimensions. The improvement came with longer runtimes and greater memory demands; some configurations with T=256T=256 failed from out-of-memory errors.

Coverage note — The benchmark’s detailed per-target metric tables, target-specific hyperparameter grid, and supporting ablations on batch size, flow architecture, proposal pretraining, and diffusion gradient guidance are omitted as narrower supporting results rather than core benchmark conclusions.

References

  1. 1.Agakov, F. V. and Barber, D. An auxiliary variational method. In Advances in Neural Information Processing Systems, 2004.
  2. 2.Akhound-Sadegh, T., Rector-Brooks, J., Bose, A. J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., et al. Iterated denoising energy matching for sampling from boltzmann densities. arXiv preprint arXiv:2402.06121, 2024.
  3. 3.Anderson, B. D. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
  4. 4.Arbel, M., Matthews, A., and Doucet, A. Annealed flow transport monte carlo. In International Conference on Machine Learning, pp. 318–330. PMLR, 2021.
  5. 5.Arenz, O., Neumann, G., and Zhong, M. Efficient gradient-free variational inference using policy search. In International conference on machine learning, pp. 234–243. PMLR, 2018.
  6. 6.Arenz, O., Dahlinger, P., Ye, Z., Volpp, M., and Neumann, G. A unified perspective on natural gradient variational inference with gaussian mixture models. arXiv preprint arXiv:2209.11533, 2022.
  7. 7.Aronszajn, N. Theory of reproducing kernels. Transactions of the American mathematical society, 68(3):337–404, 1950.
  8. 8.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. arXiv preprint arXiv:2211.01364, 2022.
  9. 9.Bishop, C. Pattern recognition and machine learning. Springer google schola, 2:531–537, 2006.
  10. 10.Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. Variational inference: A review for statisticians. Journal of the American statistical Association, 112(518):859–877, 2017.
  11. 11.Chib, S. and Greenberg, E. Understanding the metropolis-hastings algorithm. The american statistician, 49(4):327–335, 1995.
  12. 12.Cover, T. M. Elements of information theory. John Wiley & Sons, 1999.
  13. 13.Crowder, M. J. Beta-binomial anova for proportions. Applied statistics, pp. 34–37, 1978.
  14. 14.Dai Pra, P. A stochastic control approach to reciprocal diffusion processes. Applied mathematics and Optimization, 23(1):313–329, 1991.
  15. 15.Del Moral, P., Doucet, A., and Jasra, A. Sequential Monte Carlo samplers. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 68(3):411–436, 2006.
  16. 16.Dinh, L., Krueger, D., and Bengio, Y. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014.
  17. 17.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real nvp. arXiv preprint arXiv:1605.08803, 2016.
  18. 18.Douc, R. and Cappé, O. Comparison of resampling schemes for particle filtering. In ISPA 2005. Proceedings of the 4th International Symposium on Image and Signal Processing and Analysis, 2005., pp. 64–69. Ieee, 2005.
  19. 19.Doucet, A., Grathwohl, W., Matthews, A. G., and Strathmann, H. Score-based diffusion meets annealed importance sampling. Advances in Neural Information Processing Systems, 35:21482–21494, 2022a.
  20. 20.Doucet, A., Grathwohl, W., Matthews, A. G. d. G., and Strathmann, H. Score-based diffusion meets annealed importance sampling. In Advances in Neural Information Processing Systems, 2022b.
  21. 21.Duane, S., Kennedy, A. D., Pendleton, B. J., and Roweth, D. Hybrid monte carlo. Physics letters B, 195(2):216–222, 1987.
  22. 22.Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G. Neural spline flows. Advances in neural information processing systems, 32, 2019.
  23. 23.Frenkel, D. and Smit, B. Understanding molecular simulation: from algorithms to applications. Elsevier, 2023.
  24. 24.Geffner, T. and Domke, J. MCMC variational inference via uncorrected Hamiltonian annealing. In Advances in Neural Information Processing Systems, 2021.
  25. 25.Geffner, T. and Domke, J. Langevin diffusion variational inference. arXiv preprint arXiv:2208.07743, 2022.
  26. 26.Gorman, R. P. and Sejnowski, T. J. Analysis of hidden units in a layered network trained to classify sonar targets. Neural networks, 1(1):75–89, 1988.
  27. 27.Gretton, A., Borgwardt, K. M., Rasch, M. J., Scholkopf, B., and Smola, A. A kernel two-sample test. The Journal of Machine Learning Research, 13(1):723–773, 2012.
  28. 28.Hammersley, J. Monte carlo methods. Springer Science & Business Media, 2013.
  29. 29.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  30. 30.Jankowiak, M. and Phan, D. Surrogate likelihoods for variational annealed importance sampling. In International Conference on Machine Learning, pp. 9881–9901. PMLR, 2022.
  31. 31.Ji, C. and Shen, H. Stochastic variational inference via upper bound. arXiv preprint arXiv:1912.00650, 2019.
  32. 32.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  33. 33.Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. Improved variational inference with inverse autoregressive flow. Advances in neural information processing systems, 29, 2016.
  34. 34.Lahlou, S., Deleu, T., Lemos, P., Zhang, D., Volokhova, A., Hernandez-García, A., Ezzine, L. N., Bengio, Y., and Malkin, N. A theory of continuous generative flow networks. In International Conference on Machine Learning, pp. 18269–18300. PMLR, 2023.
  35. 35.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, pp. 2278–2324, 1998. doi: 10.1109/5.726791.
  36. 36.Leonard, C. A survey of the schr ¨\” odinger problem and some of its connections with optimal transport. arXiv preprint arXiv:1308.0215, 2013.
  37. 37.Li, Y. and Turner, R. E. Renyi divergence variational inference. Advances in neural information processing systems, 29, 2016.
  38. 38.Liu, J. S. and Liu, J. S. Monte Carlo strategies in scientific computing, volume 75. Springer, 2001.
  39. 39.Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. Trajectory balance: Improved credit assignment in gflownets. Advances in Neural Information Processing Systems, 35:5955–5967, 2022a.
  40. 40.Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. Gflownets and variational inference. arXiv preprint arXiv:2210.00580, 2022b.
  41. 41.Matthews, A., Arbel, M., Rezende, D. J., and Doucet, A. Continual repeated annealed flow transport monte carlo. In International Conference on Machine Learning, pp. 15196–15219. PMLR, 2022.
  42. 42.Midgley, L. I., Stimper, V., Simm, G. N., Scholkopf, B., and Hernandez-Lobato, J. M. Flow annealed importance sampling bootstrap. arXiv preprint arXiv:2208.01893, 2022.
  43. 43.Midgley, L. I., Stimper, V., Antoran, J., Mathieu, E., Scholkopf, B., and Hernández-Lobato, J. M. Se (3) equivariant augmented coupling flows. arXiv preprint arXiv:2308.10364, 2023.
  44. 44.Mittal, S., Bracher, N. L., Lajoie, G., Jaini, P., and Brubaker, M. A. Exploring exchangeable dataset amortization for bayesian posterior inference. In ICML 2023 Workshop on Structured Probabilistic Inference {&} Generative Modeling, 2023.
  45. 45.Møller, J., Syversveen, A. R., and Waagepetersen, R. P. Log gaussian cox processes. Scandinavian journal of statistics, 25(3):451–482, 1998.
  46. 46.Neal, R. M. Annealed importance sampling. Statistics and Computing, 11(2):125–139, 2001.
  47. 47.Neal, R. M. Slice sampling. The annals of statistics, 31(3):705–767, 2003.
  48. 48.Nishihara, R., Murray, I., and Adams, R. P. Parallel mcmc with generalized elliptical slice sampling. The Journal of Machine Learning Research, 15(1):2087–2112, 2014.
  49. 49.Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. Normalizing flows for probabilistic modeling and inference. The Journal of Machine Learning Research, 22(1):2617–2680, 2021.
  50. 50.Peyre, G., Cuturi, M., et al. Computational optimal transport: With applications to data science. Foundations and Trends® in Machine Learning, 11(5-6):355–607, 2019.
  51. 51.Rezende, D. and Mohamed, S. Variational inference with normalizing flows. In International conference on machine learning, pp. 1530–1538. PMLR, 2015.
  52. 52.Richter, L., Boustati, A., Nusken, N., Ruiz, F., and Akyildiz, O. D. Vargrad: a low-variance gradient estimator for variational inference. Advances in Neural Information Processing Systems, 33:13481–13492, 2020.
  53. 53.Richter, L., Berner, J., and Liu, G.-H. Improved sampling via learned diffusions. arXiv preprint arXiv:2307.01198, 2023.
  54. 54.Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. Improved techniques for training gans. Advances in neural information processing systems, 29, 2016.
  55. 55.Sarkk ¨ a, S. and Solin, A. ¨ Applied stochastic differential equations, volume 10. Cambridge University Press, 2019.
  56. 56.Sendera, M., Kim, M., Mittal, S., Lemos, P., Scimeca, L., Rector-Brooks, J., Adam, A., Bengio, Y., and Malkin, N. On diffusion models for amortized inference: Benchmarking and improving stochastic control and sampling. arXiv preprint arXiv:2402.05098, 2024.
  57. 57.Shapiro, A. Monte carlo sampling methods. Handbooks in operations research and management science, 10:353–425, 2003.
  58. 58.Sigillito, V. G., Wing, S. P., Hutton, L. V., and Baker, K. B. Classification of radar returns from the ionosphere using neural networks. Johns Hopkins APL Technical Digest, 10(3):262–266, 1989.
  59. 59.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  60. 60.Sountsov, P., Radul, A., and contributors. Inference gym, 2020. URL https://pypi.org/project/inference_gym.
  61. 61.Stoltz, G., Rousset, M., et al. Free energy computations: A mathematical perspective. World Scientific, 2010.
  62. 62.Thin, A., Kotelevskii, N., Durmus, A., Moulines, E., Panov, M., and Doucet, A. Monte Carlo variational auto-encoders. In International Conference on Machine Learning, 2021.
  63. 63.Tzen, B. and Raginsky, M. Neural stochastic differential equations: Deep latent gaussian models in the diffusion limit. arXiv preprint arXiv:1905.09883, 2019.
  64. 64.Vargas, F., Grathwohl, W., and Doucet, A. Denoising diffusion samplers. arXiv preprint arXiv:2302.13834, 2023a.
  65. 65.Vargas, F., Ovsianas, A., Fernandes, D., Girolami, M., Lawrence, N. D., and Nusken, N. Bayesian learning via neural schrodinger–föllmer flows. Statistics and Computing, 33(1):3, 2023b.
  66. 66.Vargas, F., Padhy, S., Blessing, D., and Nusken, N. Transport meets variational inference: Controlled monte carlo diffusions. In The Twelfth International Conference on Learning Representations, 2024.
  67. 67.Wainwright, M. J. and Jordan, M. I. Graphical Models, Exponential Families, and Variational Inference. Foundations and Trends® in Machine Learning, 1(1–2):1–305, November 2008.
  68. 68.Wan, N., Li, D., and Hovakimyan, N. F-divergence variational inference. Advances in neural information processing systems, 33:17370–17379, 2020.
  69. 69.Wu, H., Kohler, J., and No ¨ e, F. Stochastic normalizing flows. Advances in Neural Information Processing Systems, 33:5933–5944, 2020a.
  70. 70.Wu, H., Kohler, J., and No ¨ e, F. Stochastic normalizing flows. In Advances in Neural Information Processing Systems, 2020b.
  71. 71.Zhang, D., Chen, R. T. Q., Liu, C.-H., Courville, A., and Bengio, Y. Diffusion generative flow samplers: Improving learning signals through partial trajectory optimization. arXiv preprint arXiv:2310.02679, 2023.
  72. 72.Zhang, G., Hsu, K., Li, J., Finn, C., and Grosse, R. Differentiable annealed importance sampling and the perils of gradient noise. In Advances in Neural Information Processing Systems, 2021.
  73. 73.Zhang, Q. and Chen, Y. Path integral sampler: a stochastic control approach for sampling. arXiv preprint arXiv:2111.15141, 2021.

Citation

MLA
Blessing, D., et al. “Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling”. arXiv, 2024, http://arxiv.org/abs/2406.07423v1.
APA
Blessing, D., Jia, X., Esslinger, J., Vargas, F., & Neumann, G. (2024). Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling. arXiv. http://arxiv.org/abs/2406.07423v1
Chicago
Blessing, D., X. Jia, J. Esslinger, F. Vargas, and G. Neumann. 2024. “Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling”. arXiv. http://arxiv.org/abs/2406.07423v1.
Harvard
Blessing, D. et al. (2024) “Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.07423v1.
Vancouver
1. Blessing D, Jia X, Esslinger J, Vargas F, Neumann G (2024) Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling. arXiv

BibTeX

@article{blessing2024beyond,
  title = {Beyond ELBOs: A Large-Scale Evaluation of Variational Methods for Sampling},
  author = {Blessing, Denis and Jia, Xiaogang and Esslinger, Johannes and Vargas, Francisco and Neumann, Gerhard},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.07423v1},
  eprint = {2406.07423}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/