Iterated Denoising Energy Matching for Sampling from Boltzmann Densities

Tara Akhound-SadeghJarrid Rector-BrooksAvishek Joey BoseSarthak MittalPablo LemosCheng-Hao LiuMarcin SenderaSiamak RavanbakhshGauthier GidelYoshua Bengio

article2024ICML130 citations

Proposes an iterative, simulation-free score matching algorithm that trains diffusion samplers directly from unnormalized energy functions without data or MCMC, achieving up to five times faster training and successfully scaling to the challenging 55-particle Lennard-Jones system.

Listen

Sampling from unnormalized probability distributions, known as Boltzmann distributions, is a foundational challenge in computational chemistry, physics, and materials discovery. In these physical systems, determining the equilibrium states of molecular and particle systems requires exploring complex, high-dimensional energy landscapes. Traditional numerical techniques such as Markov Chain Monte Carlo and Molecular Dynamics are computationally expensive and scale poorly to high dimensions. While recent deep learning generative models offer an alternative, existing methods either require pre-existing data samples, rely on restrictive model architectures, or require computationally demanding trajectory simulations during training that prevent scaling to larger systems.

The article introduces and evaluates Iterated Denoising Energy Matching (iDEM), a neural sampling framework designed to learn directly from a system's energy function and its gradients without requiring any prior training data. The primary objective is to demonstrate that iDEM can efficiently scale to high-dimensional physical systems while outperforming existing neural and flow-based samplers in both sample quality and computational efficiency.

The approach operates via a bi-level iterative scheme that combines a diffusion-based model with an off-policy replay buffer. In the inner loop, a neural network is trained using a simulation-free score matching objective that estimates target score directions through local Gaussian perturbations directly on the energy function. In the outer loop, the updated diffusion model generates candidate states to populate a replay buffer without computing costly backpropagation gradients through the simulation. This iterative feedback loop smooths the rugged energy landscape and allows the model to explore multiple isolated modes effectively. The framework was evaluated across synthetic multimodal distributions and physical particle benchmark systems ranging up to 165 dimensions, incorporating physical geometric symmetries.

The evaluation yielded several key findings regarding speed, scalability, and sample accuracy. First, iDEM achieved convergence between two to five times faster than leading alternative frameworks, reducing training times by approximately a factor of four on high-dimensional benchmarks. Second, iDEM demonstrated state-of-the-art sample quality across standard metrics, including superior negative log-likelihood, mode coverage, and effective sample size. Third, iDEM was the first energy-trained neural method capable of scaling to the challenging 55-particle Lennard-Jones system, whereas competing diffusion-based neural samplers failed to converge or experienced catastrophic divergence. Finally, analysis confirmed that the error in the internal score estimator primarily impacts the magnitude rather than the direction of the score vectors, which can be stabilized in practice through simple gradient clipping.

These results demonstrate that simulation-free score matching can substantially reduce computational overhead while mitigating mode-collapse issues in scientific machine learning. By eliminating the need for expensive trajectory integration during parameter updates, organizations can dramatically lower compute costs and shorten experimental timelines for molecular design, protein modeling, and materials simulation.

Organizations developing computational sampling pipelines should consider adopting bi-level, simulation-free energy matching frameworks when scaling up molecular and atomic simulations. Future development should focus on integrating adaptive variance-reduction techniques into the score estimator and exploring faster numerical solvers for the outer-loop generation step to further optimize training throughput.

While iDEM demonstrates robust performance, key limitations include the statistical bias inherent in finite Monte Carlo score estimation within sparse or low-density regions of the energy landscape, as well as the requirement of having access to analytically computable energy gradients. Nevertheless, the empirical stability across high-dimensional tasks supports strong confidence in the framework's effectiveness for unnormalized density sampling.

arXiv: 2402.06121jarridrb/DEM

No sufficiently relevant recommendations were found.

Cover for Iterated Denoising Energy Matching for Sampling from Boltzmann Densities

Abstract

Efficiently generating statistically independent samples from an unnormalized probability distribution, such as equilibrium samples of many-body systems, is a foundational problem in science. In this paper, we propose Iterated Denoising Energy Matching (iDEM), an iterative algorithm that uses a novel stochastic score matching objective leveraging solely the energy function and its gradient—and no data samples—to train a diffusion-based sampler. Specifically, iDEM alternates between (I) sampling regions of high model density from a diffusion-based sampler and (II) using these samples in our stochastic matching objective to further improve the sampler. iDEM is scalable to high dimensions as the inner matching objective, is simulation-free, and requires no MCMC samples. Moreover, by leveraging the fast mode mixing behavior of diffusion, iDEM smooths out the energy landscape enabling efficient exploration and learning of an amortized sampler. We evaluate iDEM on a suite of tasks ranging from standard synthetic energy functions to invariant n-body particle systems. We show that the proposed approach achieves state-of-the-art performance on all metrics and trains 2 − 5× faster, which allows it to be the first method to train using energy on the challenging 55-particle Lennard-Jones system.

Table of Contents

  • 1. Introduction
  • 2. Background and preliminaries
  • 2.1. Classical sampling methods
  • 2.2. Denoising diffusion
  • 3. ITERATED DENOISING ENERGY MATCHING
  • 3.1. Denoising diffusion with a Boltzmann target (C1)
  • 3.2. Amortized sampling with a diffusion sampler (C2)
  • 3.3. Incorporating symmetries in iDEM
  • 4. Experimental results
  • 4.1. Main results
  • 4.2. Ablation experiments
  • 5. Related work
  • 6. Conclusion
  • Impact statement
  • Contribution statement
  • Acknowledgments
  • References
  • A. Proofs of propositions
  • B. iDEM for non-VE noising processes
  • C. iDEM and flow matching
  • D. Sampling with MCMC
  • D.1. Metropolis-Hastings
  • D.2. Hamiltonian Monte Carlo
  • D.3. Metropolis-Adjusted Langevin Algorithm
  • D.4. Annealed importance sampling
  • D.5. Boltzmann generators
  • D.6. Sequential Monte Carlo
  • D.7. Nested Sampling
  • E. Related simulation-based and simulation-free diffusion-like samplers
  • F. Additional Details on the Experiments
  • F.1. Experimental Setup
  • F.2. Metrics reported in Table 2 and Table 5
  • F.3. Timing experiment setup
  • F.4. Task details
  • F.4.1. GAUSSIAN MIXTURE MODEL
  • F.4.3. LENNARD-JONES POTENTIAL
  • G. Additional results
  • G.1. Additional metrics
  • G.2. Interatomic distances
  • G.3. Further ablations

Knowls

  1. Knowl 1 — Denoising Energy Matching estimates noisy target scores without target samples

    model/method

    Let the target density on Rd\mathbb{R}^d be p0(x)=exp⁡[−E(x)]/Zp_0(x)=\exp[-E(x)]/Z, where the energy EE and its gradient are available but the normalizing constant ZZ and samples from p0p_0 are not. Under variance-exploding Gaussian noising, pt=p0∗N(0,σt2Id)p_t=p_0*\mathcal{N}(0,\sigma_t^2 I_d), where t∈[0,1]t\in[0,1] and σt2\sigma_t^2 is the noise variance. For a chosen noisy point xt∈Rdx_t\in\mathbb{R}^d, draw independent ϵi∼N(0,σt2Id)\epsilon_i\sim\mathcal{N}(0,\sigma_t^2 I_d) and define xi=xt+ϵix_i=x_t+\epsilon_i for i=1,…,Ki=1,\ldots,K. The DEM score target is

    SK(xt,t)=∇xtlog⁡∑i=1Kexp⁡[−E(xt+ϵi)]=−∑i=1Kwi∇E(xi),wi=exp⁡[−E(xi)]∑j=1Kexp⁡[−E(xj)].S_K(x_t,t)=\nabla_{x_t}\log\sum_{i=1}^K\exp[-E(x_t+\epsilon_i)]=-\sum_{i=1}^K w_i\nabla E(x_i),\qquad w_i=\frac{\exp[-E(x_i)]}{\sum_{j=1}^K\exp[-E(x_j)]}.

    The weights normalize away ZZ, so evaluating this target requires energy and gradient evaluations plus Gaussian noise samples, not samples from the target or a simulated trajectory. A neural score model sθ(xt,t)s_\theta(x_t,t) is trained with ∥SK(xt,t)−sθ(xt,t)∥2\|S_K(x_t,t)-s_\theta(x_t,t)\|^2. The Monte Carlo target is biased at finite KK but consistent as KK increases. Because it can be evaluated at arbitrary xtx_t, the regression can use off-policy points, including samples retained from earlier training iterations.

  2. Knowl 2 — iDEM alternates score fitting with diffusion-based replay-buffer exploration

    algorithm

    Iterated Denoising Energy Matching (iDEM) trains a diffusion score model sθs_\theta using two alternating loops. Its inputs are a tractable prior p1p_1, a noise schedule σt2\sigma_t^2, a reverse-SDE integrator with LL steps, a replay buffer BB, batch size bb, Monte Carlo sample count KK, and an optimizer. The buffer may optionally be initialized with existing samples; no target samples are required.

    In each outer iteration, sample bb initial states from p1p_1 and integrate the reverse SDE defined by the current sθs_\theta to obtain bb candidate target-space points. Keep θ\theta fixed during this integration, add the resulting points to BB, and enforce the chosen maximum buffer size. In each inner update, draw a buffer point x0x_0 uniformly, sample t∼U(0,1)t\sim\mathcal{U}(0,1) and xt∼N(x0,σt2Id)x_t\sim\mathcal{N}(x_0,\sigma_t^2 I_d), form the KK-sample DEM score target at (xt,t)(x_t,t), and update θ\theta to reduce the squared score-regression loss. Repeat outer and inner updates to return the trained sampler. Reverse-SDE simulation is used to populate the buffer, but gradients are not backpropagated through that simulation; the DEM inner loop itself requires no SDE integration. The authors characterize this combination of model-generated points and off-policy buffer reuse as a hybrid between on-policy and off-policy training.

  3. Knowl 3 — Finite-sample DEM score error has a concentration bound

    theoretical result

    For a fixed noisy point xt∈Rdx_t\in\mathbb{R}^d and time tt, suppose that, under X∼N(xt,σt2Id)X\sim\mathcal{N}(x_t,\sigma_t^2 I_d), both exp⁡[−E(X)]\exp[-E(X)] and ∥∇exp⁡[−E(X)]∥\|\nabla\exp[-E(X)]\| are sub-Gaussian. Then the paper's Monte Carlo score estimator SK(xt,t)S_K(x_t,t), based on KK independent draws from this Gaussian, satisfies, for a constant c(xt)c(x_t) and 0<δ<10<\delta<1,

    Pr⁡ ⁣(∥SK(xt,t)−∇log⁡pt(xt)∥≤c(xt)log⁡(1/δ)K)≥1−δ.\Pr\!\left(\|S_K(x_t,t)-\nabla\log p_t(x_t)\|\leq \frac{c(x_t)\log(1/\delta)}{\sqrt{K}}\right)\geq 1-\delta.

    Thus the stated bound gives an O(K−1/2)O(K^{-1/2}) error rate, with a point-dependent constant. It also implies that the estimator can require more Monte Carlo samples where the Gaussian-averaged unnormalized density is small.

  4. Knowl 4 — DEM can preserve particle-system symmetries

    model/method

    For an nn-particle system in three dimensions whose energy is invariant to rigid motions and particle permutations, the target density is invariant under G=SE(3)×SnG=SE(3)\times S_n. The paper establishes that the Monte Carlo DEM score estimator is GG-equivariant when its Gaussian sampling distribution transforms invariantly with the noisy input. In practice, rotational and permutation symmetry are compatible with standard isotropic Gaussian noise; translation invariance is handled by using Gaussian noise with zero center of mass, which confines perturbations to the translation-invariant subspace. An equivariant diffusion score model can then be used in iDEM so that both the score estimation and sampling respect the system's symmetries.

  5. Knowl 5 — Sampler-quality comparison across four energy benchmarks

    data/table

    The comparison covers a 40-component two-dimensional Gaussian mixture (GMM), a four-particle double-well system (DW-4, d=8d=8), and three-dimensional Lennard-Jones systems with 13 and 55 particles (LJ-13, d=39d=39; LJ-55, d=165d=165). Entries are mean ±\pm standard deviation over three seeds. NLL is estimated using the same OT conditional-flow-matching evaluation procedure for samples from each method; lower NLL and W2W_2 are better, while higher normalized ESS is better. An asterisk marks divergent training. The results show that iDEM is competitive on the lower-dimensional tasks and is the strongest successful method on LJ-55 across these reported metrics; PIS, DDS, and pDEM diverge on LJ-55.

    GMM (d=2d=2) DW-4 (d=8d=8) LJ-13 (d=39d=39) LJ-55 (d=165d=165)
    Method NLL ESS W2W_2 NLL ESS W2W_2 NLL ESS W2W_2 NLL ESS W2W_2
    FAB 7.14±0.017.14\pm0.01 0.653±0.0170.653\pm0.017 12.0±5.7312.0\pm5.73 7.16±0.017.16\pm0.01 0.947±0.0070.947\pm0.007 2.15±0.022.15\pm0.02 17.52±0.1717.52\pm0.17 0.101±0.0590.101\pm0.059 4.35±0.014.35\pm0.01 200.32±62.3200.32\pm62.3 0.063±0.0010.063\pm0.001 18.03±1.2118.03\pm1.21
    PIS 7.72±0.037.72\pm0.03 0.295±0.0180.295\pm0.018 7.64±0.927.64\pm0.92 7.19±0.017.19\pm0.01 0.901±0.0030.901\pm0.003 2.13±0.022.13\pm0.02 47.05±12.4647.05\pm12.46 0.004±0.0020.004\pm0.002 4.67±0.114.67\pm0.11 ∗* ∗* ∗*
    DDS 7.43±0.467.43\pm0.46 0.687±0.2080.687\pm0.208 9.31±0.829.31\pm0.82 11.27±1.2411.27\pm1.24 0.408±0.0010.408\pm0.001 2.15±0.042.15\pm0.04 ∗* ∗* ∗* ∗* ∗* ∗*
    pDEM 7.10±0.027.10\pm0.02 0.634±0.0840.634\pm0.084 12.20±0.1412.20\pm0.14 7.44±0.057.44\pm0.05 0.547±0.0100.547\pm0.010 2.11±0.032.11\pm0.03 18.80±0.4818.80\pm0.48 0.044±0.0130.044\pm0.013 4.21±0.064.21\pm0.06 ∗* ∗* ∗*
    iDEM 6.96±0.076.96\pm0.07 0.734±0.0920.734\pm0.092 7.42±3.447.42\pm3.44 7.17±0.007.17\pm0.00 0.825±0.0020.825\pm0.002 2.13±0.042.13\pm0.04 17.68±0.1417.68\pm0.14 0.231±0.0050.231\pm0.005 4.26±0.034.26\pm0.03 125.86±18.03125.86\pm18.03 0.106±0.0220.106\pm0.022 16.128±0.07116.128\pm0.071

    On GMM, iDEM has the best NLL and ESS among the listed methods and the lowest W2W_2. On LJ-13 it has the highest ESS, while FAB has slightly lower NLL and pDEM slightly lower W2W_2. On DW-4 its NLL is close to FAB's, and its ESS is below FAB's. On LJ-55, iDEM is the only neural sampler in this comparison to train successfully, and it also outperforms FAB on all three listed metrics.

  6. Knowl 6 — Additional total-variation and partition-function estimates

    data/table

    This comparison reports total variation (TV; lower is better) and estimated log⁡Z\log Z for the same four benchmarks. The log⁡Z\log Z estimates are lower bounds, so larger values are preferred; for LJ-55, the FAB dagger indicates that only one of its three runs converged. The results provide additional evidence of distributional fit: iDEM's TV is tied with the best reported value on GMM and LJ-13 and is lower than FAB's on LJ-55, while its log⁡Z\log Z estimate is the largest reported for LJ-13 and LJ-55. It is not best on every metric, including DW-4.

    GMM (d=2d=2) DW-4 (d=8d=8) LJ-13 (d=39d=39) LJ-55 (d=165d=165)
    Method TV log⁡Z\log Z TV log⁡Z\log Z TV log⁡Z\log Z TV log⁡Z\log Z
    FAB 0.88±0.020.88\pm0.02 −1.165±0.164-1.165\pm0.164 0.09±0.000.09\pm0.00 29.602±0.01929.602\pm0.019 0.04±0.000.04\pm0.00 4.35±0.014.35\pm0.01 0.24±0.090.24\pm0.09 32.809†32.809^{\dagger}
    PIS 0.92±0.010.92\pm0.01 −2.243±0.070-2.243\pm0.070 0.09±0.000.09\pm0.00 29.599±0.00929.599\pm0.009 0.25±0.010.25\pm0.01 46.685±1.47146.685\pm1.471 ∗* ∗*
    DDS 0.82±0.020.82\pm0.02 −0.358±0.209-0.358\pm0.209 0.16±0.010.16\pm0.01 28.382±0.15828.382\pm0.158 ∗* ∗* ∗* ∗*
    pDEM 0.82±0.020.82\pm0.02 −0.370±0.005-0.370\pm0.005 0.13±0.000.13\pm0.00 29.191±0.03629.191\pm0.036 0.06±0.020.06\pm0.02 32.450±3.19132.450\pm3.191 ∗* ∗*
    iDEM 0.82±0.010.82\pm0.01 −0.340±0.075-0.340\pm0.075 0.10±0.010.10\pm0.01 29.567±0.01429.567\pm0.014 0.04±0.010.04\pm0.01 49.969±2.78449.969\pm2.784 0.09±0.010.09\pm0.01 273.167±22.226273.167\pm22.226

    The reported TV for the GMM is computed over a two-dimensional histogram; for the particle systems, TV is computed from interatomic-distance distributions. The log⁡Z\log Z values are estimated using an OT-CFM proposal, and are lower bounds on the true log partition function. An asterisk denotes divergence.

  7. Knowl 7 — iDEM reduces training time, especially on high-dimensional particle systems

    data/table

    Training time to convergence is reported in hours, excluding evaluation. Across the four tasks, iDEM is faster than FAB and the simulation-based diffusion samplers PIS and DDS wherever those methods converge. Its advantage over FAB grows on LJ-13 and LJ-55, consistent with DEM avoiding trajectory integration and backpropagation in its inner training loop. pDEM is faster still, but its LJ-55 runs are unstable. An asterisk denotes divergent training.

    Method GMM DW-4 LJ-13 LJ-55
    FAB 1.71 6.87 21.78 40.35
    PIS 4.11 11.29 17.36 ∗*
    DDS 1.81 5.65 ∗* ∗*
    pDEM 0.36 1.40 1.79 ∗*
    iDEM 0.87 4.30 6.55 7.75

    The authors report iDEM as approximately 1.8×1.8\times faster than FAB on GMM and DW-4 and approximately 4×4\times faster on LJ-13 and LJ-55.

  8. Knowl 8 — Estimator ablations reveal the effects of sample count and noise level

    empirical result

    On the GMM benchmark, increasing the Monte Carlo count KK decreases the bias and mean squared error (MSE) of the DEM score estimator; a regression to the measured bias gives an empirical asymptotic rate of O(1/K)O(1/K), faster than the paper's stated theoretical O(1/K)O(1/\sqrt{K}) bound. For fixed KK, estimator MSE increases at larger diffusion times, where samples are farther from high-density regions. On DW-4, increasing KK improves the training-time TV result. These findings support the need for more score-estimation samples in noisier or less informative regions.

    The authors also compared score-target constructions. Their log-sum-exp estimator was the well-behaved option in these tests: a direct ratio estimator using exponentiated energies frequently produced non-finite values below 500 Monte Carlo samples, while a Jensen-based estimate retained substantial bias as sample count increased. The finite-KK bias and sensitivity to estimator variance are identified as limitations of DEM.

  9. Knowl 9 — A direct DEM score estimator was tested on a 10,000-dimensional mixture

    empirical result

    The authors extended the 40-component GMM from two dimensions to 10,000 dimensions and used the DEM score estimator directly to generate samples, rather than first fitting a neural network. Across the dimensions and noise levels they examined, estimator bias and MSE increased with dimension and noise, while the cosine similarity between the mean estimated score and the true score remained close to one; the observed error was therefore mainly in score magnitude rather than direction. With score-norm clipping, generated samples visually matched the true distribution when projected onto the first two principal components for clipping values from 10410^4 to 10810^8. The paper hypothesizes that magnitude overestimation at high noise may be tolerable if score directions remain informative and estimates become more accurate nearer the target.

Coverage note — The appendix's non-variance-exploding-noise reparameterization and flow-matching discussion are omitted because they are extensions or connections rather than central contributions; auxiliary per-task hyperparameter grids and interatomic-distance plots are omitted as supporting experimental detail.

References

  1. 1.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. International Conference on Learning Representations (ICLR), 2023.
  2. 2.Albergo, M. S., Kanwar, G., and Shanahan, P. E. Flow-based generative models for markov chain monte carlo in lattice field theory. Physical Review D, 100(3):034515, 2019.
  3. 3.Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., and Baraniuk, R. G. Self-consuming generative models go mad. International Conference on Learning Representations (ICLR), 2023.
  4. 4.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. arXiv preprint arXiv:2211.01364, 2022.
  5. 5.Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M., and Gidel, G. On the stability of iterative retraining of generative models on their own data. International Conference on Learning Representations (ICLR), 2023.
  6. 6.Bose, A. J., Brubaker, M., and Kobyzev, I. Equivariant finite normalizing flows. arXiv preprint arXiv:2110.08649, 2021.
  7. 7.Bose, A. J., Akhound-Sadegh, T., Fatras, K., Huguet, G., Rector-Brooks, J., Liu, C.-H., Nica, A. C., Korablyov, M., Bronstein, M., and Tong, A. SE(3)-stochastic flow matching for protein backbone generation. International Conference on Learning Representations (ICLR), 2024.
  8. 8.Brehmer, J., Bose, J., De Haan, P., and Cohen, T. EDGI: Equivariant diffusion for planning with embodied agents. Neural Information Processing Systems (NeurIPS), 2023a.
  9. 9.Brehmer, J., De Haan, P., Behrends, S., and Cohen, T. Geometric algebra transformer. Neural Information Processing Systems (NeurIPS), 2023b.
  10. 10.Bugallo, M. F., Elvira, V., Martino, L., Luengo, D., Miguez, J., and Djuric, P. M. Adaptive importance sampling: The past, the present, and the future. IEEE Signal Processing Magazine, 34(4):60–79, 2017.
  11. 11.Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. Neural Information Processing Systems (NIPS), 2018.
  12. 12.Dai Pra, P. A stochastic control approach to reciprocal diffusion processes. Applied mathematics and Optimization, 23:313–329, 1991.
  13. 13.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion schrodinger bridge with applications to score-based generative modeling. Neural Information Processing Systems (NeurIPS), 2021.
  14. 14.Del Moral, P., Doucet, A., and Jasra, A. Sequential Monte Carlo samplers. Journal of the Royal Statistical Society Series B: Statistical Methodology, 68(3):411–436, 2006.
  15. 15.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using Real NVP. International Conference on Learning Representations (ICLR), 2017.
  16. 16.Doucet, A., Grathwohl, W., Matthews, A. G., and Strathmann, H. Score-based diffusion meets annealed importance sampling. Neural Information Processing Systems (NeurIPS), 2022.
  17. 17.Feroz, F., Hobson, M. P., Cameron, E., and Pettitt, A. N. Importance nested sampling and the MultiNest algorithm. arXiv preprint arXiv:1306.2144, 2013.
  18. 18.Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Bois-bunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., Gautheron, L., Gayraud, N. T., Janati, H., Rakotomamonjy, A., Redko, I., Rolet, A., Schutz, A., Seguy, V., Sutherland, D. J., Tavenard, R., Tong, A., and Vayer, T. Pot: Python optimal transport. Journal of Machine Learning Research, 22(78):1–8, 2021. URL http://jmlr.org/papers/v22/20-451.html.
  19. 19.Garcia Satorras, V., Hoogeboom, E., Fuchs, F., Posner, I., and Welling, M. E(n) equivariant normalizing flows. Neural Information Processing Systems (NeurIPS), 2021.
  20. 20.Geffner, T. and Domke, J. MCMC variational inference via uncorrected Hamiltonian annealing. Neural Information Processing Systems (NeurIPS), 2021.
  21. 21.Geffner, T. and Domke, J. Langevin diffusion variational inference. Artificial Intelligence and Statistics (AISTATS), 2023.
  22. 22.Grenander, U. and Miller, M. I. Representations of knowledge in complex systems. Journal of the Royal Statistical Society: Series B (Methodological), 56(4):549–581, 1994.
  23. 23.Grenioux, L., Durmus, A., Moulines, E., and Gabrié, M. On sampling with approximate transport maps. arXiv preprint arXiv:2302.04763, 2023.
  24. 24.Handley, W., Hobson, M., and Lasenby, A. Polychord: nested sampling for cosmology. Monthly Notices of the Royal Astronomical Society: Letters, 450(1):L61–L65, 2015.
  25. 25.Hastings, W. K. Monte carlo sampling methods using markov chains and their applications. 1970.
  26. 26.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Neural Information Processing Systems (NeurIPS), 2020.
  27. 27.Hoffman, M. D., Gelman, A., et al. The no-u-turn sampler: adaptively setting path lengths in hamiltonian monte carlo. Journal of Machine Learning Research, 15(1):1593–1623, 2014.
  28. 28.Hoogeboom, E., Satorras, V. G., Vignac, C., and Welling, M. Equivariant diffusion for molecule generation in 3d. International Conference on Machine Learning (ICML), 2022.
  29. 29.Huang, X., Dong, H., Hao, Y., Ma, Y.-A., and Zhang, T. Reverse diffusion Monte Carlo. International Conference on Learning Representations (ICLR), 2024.
  30. 30.Igashov, I., Stark, H., Vignac, C., Satorras, V. G., Frossard, P., Welling, M., Bronstein, M., and Correia, B. Equivariant 3d-conditional diffusion models for molecular linker design. International Conference on Learning Representations (ICLR), 2022.
  31. 31.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  32. 32.Kirkpatrick, S., Gelatt Jr, C. D., and Vecchi, M. P. Optimization by simulated annealing. science, 220(4598):671–680, 1983.
  33. 33.Klein, L., Foong, A. Y., Fjelde, T. E., Mlodozeniec, B., Brockschmidt, M., Nowozin, S., Noe, F., and Tomioka, R. Timewarp: Transferable acceleration of molecular dynamics by learning time-coarsened dynamics. Neural Information Processing Systems (NeurIPS), 2023a.
  34. 34.Klein, L., Kramer, A., and Noé, F. Equivariant flow matching. Neural Information Processing Systems (NeurIPS), 2023b.
  35. 35.Köhler, J., Klein, L., and Noé, F. Equivariant flows: exact likelihood generative learning for symmetric densities. International Conference on Machine Learning (ICML), 2020.
  36. 36.Köhler, J., Invernizzi, M., De Haan, P., and Noé, F. Rigid body flows for sampling molecular crystal structures. International Conference on Machine Learning (ICML), 2023.
  37. 37.Lahlou, S., Deleu, T., Lemos, P., Zhang, D., Volokhova, A., Hernandez-García, A., Ezzine, L. N., Bengio, Y., and Malkin, N. A theory of continuous generative flow networks. International Conference on Machine Learning (ICML), 2023.
  38. 38.Leimkuhler, B. and Matthews, C. Rational construction of stochastic numerical methods for molecular sampling. Applied Mathematics Research eXpress, 2013(1):34–56, 2013.
  39. 39.Lemos, P., Malkin, N., Handley, W., Bengio, Y., Hezaveh, Y., and Perreault-Levasseur, L. Improving gradient-guided nested sampling for posterior inference. arXiv preprint arXiv:2312.03911, 2023.
  40. 40.Li, S.-H. and Wang, L. Neural network renormalization group. Physical review letters, 121(26):260601, 2018.
  41. 41.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. International Conference on Learning Representations (ICLR), 2023.
  42. 42.Liu, Q. Rectified flow: A marginal preserving approach to optimal transport. arXiv preprint arXiv:2209.14577, 2022.
  43. 43.Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. GFlowNets and variational inference. International Conference on Learning Representations (ICLR), 2023.
  44. 44.Matthews, A., Arbel, M., Rezende, D. J., and Doucet, A. Continual repeated annealed flow transport monte carlo. International Conference on Machine Learning (ICML), 2022.
  45. 45.Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., and Teller, E. Equation of state calculations by fast computing machines. The journal of chemical physics, 21(6):1087–1092, 1953.
  46. 46.Midgley, L. I., Stimper, V., Antoran, J., Mathieu, E., Scholkopf, B., and Hernández-Lobato, J. M. SE(3) equivariant augmented coupling flows. Neural Information Processing Systems (NeurIPS), 2023a.
  47. 47.Midgley, L. I., Stimper, V., Simm, G. N., Scholkopf, B., and Hernández-Lobato, J. M. Flow annealed importance sampling bootstrap. International Conference on Learning Representations (ICLR), 2023b.
  48. 48.Neal, R. M. Annealed importance sampling. Statistics and computing, 11:125–139, 2001.
  49. 49.Neal, R. M. Slice sampling. The annals of statistics, 31(3):705–767, 2003.
  50. 50.Neal, R. M. et al. MCMC using Hamiltonian dynamics. Handbook of Markov chain Monte Carlo, 2(11):2, 2011.
  51. 51.Nicoli, K. A., Nakajima, S., Strodthoff, N., Samek, W., Müller, K.-R., and Kessel, P. Asymptotically unbiased estimation of physical observables with neural samplers. Physical Review E, 101(2):023304, 2020.
  52. 52.Noé, F., Olsson, S., Köhler, J., and Wu, H. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365(6457):eaaw1147, 2019.
  53. 53.Owen, A. B. Monte Carlo theory, methods and examples. https://artowen.su.domains/mc/, 2013.
  54. 54.Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. Normalizing flows for probabilistic modeling and inference. Journal of Machine Learning Research, 22(1):2617–2680, 2021.
  55. 55.Pavon, M. Stochastic control and nonequilibrium thermodynamical systems. Applied Mathematics and Optimization, 19:187–202, 1989.
  56. 56.Rezende, D. and Mohamed, S. Variational inference with normalizing flows. International Conference on Machine Learning (ICML), 2015.
  57. 57.Richter, L., Berner, J., and Liu, G.-H. Improved sampling via learned diffusions. International Conference on Learning Representations (ICLR), 2024.
  58. 58.Robert, C. P., Casella, G., and Casella, G. Monte Carlo statistical methods, volume 2. Springer, 1999.
  59. 59.Roberts, G. O. and Rosenthal, J. S. Optimal scaling of discrete approximations to langevin diffusions. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 60(1):255–268, 1998.
  60. 60.Roberts, G. O. and Tweedie, R. L. Exponential convergence of langevin distributions and their discrete approximations. Bernoulli, pp. 341–363, 1996.
  61. 61.Satorras, V. G., Hoogeboom, E., and Welling, M. E (n) equivariant graph neural networks. International Conference on Machine Learning (ICML), 2021.
  62. 62.Skilling, J. Nested sampling for general Bayesian computation. 2006.
  63. 63.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. International Conference on Machine Learning (ICML), 2015.
  64. 64.Song, J., Zhang, Q., Yin, H., Mardani, M., Liu, M.-Y., Kautz, J., Chen, Y., and Vahdat, A. Loss-guided diffusion models for plug-and-play controllable generation. International Conference on Machine Learning (ICML), 2023.
  65. 65.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. International Conference on Learning Representations (ICLR), 2021.
  66. 66.Speagle, J. S. DYNESTY: a dynamic nested sampling package for estimating Bayesian posteriors and evidences. Monthly Notices of the Royal Astronomical Society, 493(3):3132–3158, 2020.
  67. 67.Thin, A., Kotelevskii, N., Doucet, A., Durmus, A., Moulines, E., and Panov, M. Monte Carlo variational auto-encoders. International Conference on Machine Learning (ICML), 2021.
  68. 68.Tieleman, T. Training restricted boltzmann machines using approximations to the likelihood gradient. International Conference on Machine Learning (ICML), 2008.
  69. 69.Tong, A., Malkin, N., Huguet, G., Zhang, Y., Rector-Brooks, J., Fatras, K., Wolf, G., and Bengio, Y. Improving and generalizing flow-based generative models with mini-batch optimal transport. arXiv preprint arXiv:2302.00482, 2023.
  70. 70.Tzen, B. and Raginsky, M. Neural stochastic differential equations: Deep latent Gaussian models in the diffusion limit. arXiv preprint arXiv:1905.09883, 2019a.
  71. 71.Tzen, B. and Raginsky, M. Theoretical guarantees for sampling and inference in generative models with latent diffusions. Conference on Learning Theory (CoLT), 2019b.
  72. 72.Vargas, F., Grathwohl, W., and Doucet, A. Denoising diffusion samplers. International Conference on Learning Representations (ICLR), 2023.
  73. 73.Vargas, F., Padhy, S., Blessing, D., and Nüsken, N. Transport meets variational inference: Controlled Monte Carlo diffusions. International Conference on Learning Representations (ICLR), 2024.
  74. 74.Vershynin, R. High-dimensional probability: An introduction with applications in data science. Cambridge university press, 2018.
  75. 75.Wainwright, M. J., Jordan, M. I., et al. Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning, 1(1–2):1–305, 2008.
  76. 76.Wu, H., Köhler, J., and Noé, F. Stochastic normalizing flows. Neural Information Processing Systems (NeurIPS), 2020.
  77. 77.Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022.
  78. 78.Yim, J., Campbell, A., Foong, A. Y. K., Gastegger, M., Jimenez-Luna, J., Lewis, S., Satorras, V. G., Veeling, B. S., Barzilay, R., Jaakkola, T., and Noé, F. Fast protein backbone generation with SE(3) flow matching. arXiv preprint arXiv:2310.05297, 2023a.
  79. 79.Yim, J., Trippe, B. L., De Bortoli, V., Mathieu, E., Doucet, A., Barzilay, R., and Jaakkola, T. SE(3) diffusion model with application to protein backbone generation. International Conference on Machine Learning (ICML), 2023b.
  80. 80.Zhang, D., Chen, R. T. Q., Malkin, N., and Bengio, Y. Unifying generative models with GFlowNets and beyond. arXiv preprint arXiv:2209.02606v2, 2023.
  81. 81.Zhang, Q. and Chen, Y. Path integral sampler: a stochastic control approach for sampling. International Conference on Learning Representations (ICLR), 2022.

Citation

MLA
Akhound-Sadegh, T., et al. “Iterated Denoising Energy Matching for Sampling from Boltzmann Densities”. arXiv, 2024, http://arxiv.org/abs/2402.06121v2.
APA
Akhound-Sadegh, T., Rector-Brooks, J., Bose, A. J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., Malkin, N., & Tong, A. (2024). Iterated Denoising Energy Matching for Sampling from Boltzmann Densities. arXiv. http://arxiv.org/abs/2402.06121v2
Chicago
Akhound-Sadegh, T., J. Rector-Brooks, A. J. Bose, et al. 2024. “Iterated Denoising Energy Matching for Sampling from Boltzmann Densities”. arXiv. http://arxiv.org/abs/2402.06121v2.
Harvard
Akhound-Sadegh, T. et al. (2024) “Iterated Denoising Energy Matching for Sampling from Boltzmann Densities”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.06121v2.
Vancouver
1. Akhound-Sadegh T, Rector-Brooks J, Bose AJ, et al (2024) Iterated Denoising Energy Matching for Sampling from Boltzmann Densities. arXiv

BibTeX

@article{akhoundsadegh2024iterated,
  title = {Iterated Denoising Energy Matching for Sampling from Boltzmann Densities},
  author = {Akhound-Sadegh, Tara and Rector-Brooks, Jarrid and Bose, Avishek Joey and Mittal, Sarthak and Lemos, Pablo and Liu, Cheng-Hao and Sendera, Marcin and Ravanbakhsh, Siamak and Gidel, Gauthier and Bengio, Yoshua and Malkin, Nikolay and Tong, Alexander},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.06121v2},
  eprint = {2402.06121}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/