NETS: A Non-equilibrium Transport Sampler

Michael Samuel AlbergoEric Vanden-Eijnden

article2025ICML76 citations

Introduces an unbiased sampling framework for unnormalized target distributions that augments non-equilibrium Langevin dynamics with a learned drift field trained without backpropagation to minimize importance weight variance.

Listen

Sampling from complex, unnormalized probability distributions is fundamental to modern machine learning, Bayesian inference, and computational physical sciences. Standard Markov Chain Monte Carlo methods often struggle with multimodal distributions because particles become trapped in local energy wells, resulting in prohibitively slow convergence. While non-equilibrium methods like annealed importance sampling transport a simple initial distribution toward a complex target in finite time, they suffer from high-variance corrective weights whenever the sampling trajectory lags behind the evolving distribution. The article introduces the Non-Equilibrium Transport Sampler (NETS), a framework that resolves this challenge by augmenting annealed Langevin dynamics with a learned transport drift to maintain particle alignment with the target distribution.

The authors develop and evaluate this approach through an optimize-then-discretize mathematical framework and extensive numerical simulations. A neural network learns the optimal velocity field using off-policy objectives based on physics-informed neural networks or action matching. Crucially, these training objectives do not require backpropagating through the simulation differential equations, and they provably bound the divergence between the generated and target distributions. The article validates the sampler across standard challenging benchmarks, including multi-modal Gaussian mixture models ranging from 2 to 200 dimensions, a 10-dimensional funnel distribution, a 50-dimensional mixture of Student-t distributions, and a 400-dimensional lattice field theory model near its critical phase transition.

The experimental findings show substantial improvements in sampling efficiency and statistical accuracy. On a standard 40-mode Gaussian mixture benchmark, the proposed method achieves an effective sample size between 97.9% and 99.3%, whereas standard annealed importance sampling collapses to under 1%. In high-dimensional scaling tests up to 200 dimensions, learned transport alone achieves an effective sample size of approximately 60%, whereas unaugmented annealing fails completely. Furthermore, in lattice field theory simulations, the framework proves nearly two orders of magnitude more statistically efficient than conventional annealed sampling, accurately capturing physical phase transitions and producing unbiased magnetization estimates matching gold-standard Hybrid Monte Carlo baselines.

These results demonstrate that combining learned deterministic transport with stochastic diffusion allows practitioners to generate high-quality, unbiased samples from difficult distributions at significantly lower computational variance. The method allows post-training tuning of the diffusion coefficient and integration step size, offering a practical trade-off between computation time and sample fidelity. Organizations conducting complex Bayesian modeling or physical simulations should consider integrating this learned-transport framework into their workflows, pairing it with particle resampling when effective sample sizes decline. However, practitioners should be aware that resolving dynamics near sharp physical phase transitions still requires finer numerical discretization, demanding 1,500 to 2,000 integration steps. Future work should focus on developing adaptive time-stepping schemes and testing the architecture on large-scale molecular dynamics and industrial posterior inference problems.

arXiv: 2410.02711

No sufficiently relevant recommendations were found.

Cover for NETS: A Non-equilibrium Transport Sampler

Abstract

We introduce the Non-Equilibrium Transport Sampler (NETS), an algorithm for sampling from unnormalized probability distributions. NETS builds on non-equilibrium sampling strategies that transport a simple base distribution into the target distribution in finite time, as pioneered in Neal’s annealed importance sampling (AIS). In the continuous-time setting, this transport is accomplished by evolving walkers using Langevin dynamics with a time-dependent potential, while simultaneously evolving importance weights to debias their solutions following Jarzynski’s equality. The key innovation of NETS is to add to the dynamics a learned drift term that offsets the need for these corrective weights by minimizing their variance through an objective that can be estimated without backpropagation and provably bounds the Kullback-Leibler divergence between the estimated and target distributions. NETS provides unbiased samples and features a tunable diffusion coefficient that can be adjusted after training to maximize the effective sample size. In experiments on standard benchmarks, high-dimensional Gaussian mixtures, and statistical lattice field theory models, NETS shows compelling performances.

Table of Contents

  • 1. Introduction
  • 1.1. Related work
  • 2. Methods
  • 2.1. Setup and Notations
  • 2.2. Non-equilibrium Sampling with Importance Weights
  • 2.3. Non-equilibrium Sampling with Perfect Transport
  • 2.4. Non-Equilibrium Transport Sampler
  • 2.5. Estimating the Drift b t ( x ) via a PINN Objective
  • 2.6. Control of the Kullback-Leibler Divergence
  • 2.7. Implementation
  • 2.8. Learning the U t of Stochastic Interpolants
  • 3. Numerical Experiments
  • 3.1. 40-Mode Gaussian Mixture
  • 3.2. Funnel and Student-t Mixture
  • 3.3. Scaling on High-dimensional GMMs
  • 3.4. Lattice φ 4 Theory
  • Impact Statement
  • References
  • A. Proofs of Section 2
  • B. Time-discretized Version of Proposition 2.4
  • C. Solving for the Optimal Drift via Feynman-Kac Formula
  • D. Extensions and Generalizations
  • D.1. Inertial NETS
  • D.2. Multimarginal NETS
  • E. Estimating the Drift b t ( x ) = ∇ ϕ t ( x ) via Action Matching (AM)
  • F. Link with CMCD (Vargas et al., 2024)
  • G. Hutchinson's Trace Estimator for the Evaluation of ∇· ˆ b t ( x )
  • H. Details on Numerical Experiments
  • H.1. Performance metrics
  • H.2. 40-mode GMM
  • H.3. Neal's 10d funnel
  • H.4. 50d Mixture of Student-T distributions
  • I. Details of the φ 4 model
  • I.1. Free theory λ = 0
  • I.2. φ 4 numerical details

Knowls

  1. Knowl 1 — NETS gives an exact weighted sampling identity for learned transport

    model/method

    Let Ut:Rd→RU_t:\mathbb{R}^d\to\mathbb{R}, t∈[0,1]t\in[0,1], be a twice-differentiable path of normalizable potentials, with densities ρt(x)=Zt−1e−Ut(x)\rho_t(x)=Z_t^{-1}e^{-U_t(x)} and free energies Ft=−log⁡ZtF_t=-\log Z_t. Assume the initial density ρ0\rho_0 can be sampled and its normalizer Z0Z_0 is known. For a chosen time-dependent diffusion coefficient ϵt≥0\epsilon_t\geq 0 and a sufficiently regular learned drift b^t:Rd→Rd\hat b_t:\mathbb{R}^d\to\mathbb{R}^d, evolve

    dXt=(−ϵt∇Ut(Xt)+b^t(Xt)) dt+2ϵt dWt,dAt=(∇⋅b^t(Xt)−∇Ut(Xt)⋅b^t(Xt)−∂tUt(Xt)) dt,dX_t=(-\epsilon_t\nabla U_t(X_t)+\hat b_t(X_t))\,dt+\sqrt{2\epsilon_t}\,dW_t, \qquad dA_t=(\nabla\cdot\hat b_t(X_t)-\nabla U_t(X_t)\cdot\hat b_t(X_t)-\partial_tU_t(X_t))\,dt,

    with X0∼ρ0X_0\sim\rho_0 and A0=0A_0=0, where WtW_t is standard dd-dimensional Brownian motion. For any integrable test function hh, the law of these paths satisfies

    Eρt[h(X)]=E[eAth(Xt)]E[eAt],ZtZ0=E[eAt].\mathbb{E}_{\rho_t}[h(X)]=\frac{\mathbb{E}[e^{A_t}h(X_t)]}{\mathbb{E}[e^{A_t}]}, \qquad \frac{Z_t}{Z_0}=\mathbb{E}[e^{A_t}].

    Thus even an imperfect transport drift preserves an exact importance-weighted identity for target expectations and the partition-function ratio. The weights are generally variable; unweighted endpoints are exact target samples only when the added drift is the perfect transport.

  2. Knowl 2 — The perfect transport drift removes weight variance

    theoretical result

    For a prescribed density path ρt(x)=Zt−1e−Ut(x)\rho_t(x)=Z_t^{-1}e^{-U_t(x)}, an added velocity field bt(x)b_t(x) is a perfect transport if it solves the continuity equation

    ∇⋅(btρt)=−∂tρt.\nabla\cdot(b_t\rho_t)=-\partial_t\rho_t.

    If XtX_t follows dXt=(−ϵt∇Ut(Xt)+bt(Xt))dt+2ϵt dWtdX_t=(-\epsilon_t\nabla U_t(X_t)+b_t(X_t))dt+\sqrt{2\epsilon_t}\,dW_t, starting from X0∼ρ0X_0\sim\rho_0, then its marginal density is ρt\rho_t at every t∈[0,1]t\in[0,1]. In the NETS weight equation, setting b^t=bt\hat b_t=b_t makes the log weight deterministic: At=−Ft+F0A_t=-F_t+F_0. Consequently, the dynamics sample the prescribed density without importance-weight variance, and the same weight gives the exact normalizer ratio Zt/Z0Z_t/Z_0.

  3. Knowl 3 — PINN loss identifies the perfect drift and free energy

    theoretical result

    Let b^t(x)\hat b_t(x) be a candidate drift, F^t\widehat F_t a candidate free energy, and ρ^t(x)>0\widehat\rho_t(x)>0 any probability density used to weight locations. Define the residual and loss, for T∈(0,1]T\in(0,1], by

    qt(x)=∇⋅b^t(x)−∇Ut(x)⋅b^t(x)−∂tUt(x)+∂tF^t,LPINNT[b^,F^]=∫0TEx∼ρ^t[∣qt(x)∣2]dt.q_t(x)=\nabla\cdot\hat b_t(x)-\nabla U_t(x)\cdot\hat b_t(x)-\partial_tU_t(x)+\partial_t\widehat F_t, \qquad L_{\mathrm{PINN}}^T[\hat b,\widehat F]=\int_0^T\mathbb{E}_{x\sim\widehat\rho_t}[|q_t(x)|^2]dt.

    Under the paper’s regularity and normalizability assumptions, the minimum loss is zero, and every zero-loss minimizer has F^t=Ft\widehat F_t=F_t (with the initial free-energy value fixed by Z0Z_0) and a drift satisfying ∇⋅(b^tρt)=−∂tρt\nabla\cdot(\hat b_t\rho_t)=-\partial_t\rho_t throughout [0,T][0,T]. The sampling density ρ^t\widehat\rho_t may be chosen off-policy; target samples are not required to define the objective. Using ρ^t=ρt\widehat\rho_t=\rho_t emphasizes the regions carrying target probability mass.

  4. Knowl 4 — PINN residual controls endpoint KL divergence

    theoretical result

    Let ρ^t\hat\rho_t solve

    ∂tρ^t+∇⋅(b^tρ^t)=ϵt∇⋅(∇Utρ^t+∇ρ^t),ρ^0=ρ0,\partial_t\hat\rho_t+\nabla\cdot(\hat b_t\hat\rho_t) =\epsilon_t\nabla\cdot(\nabla U_t\hat\rho_t+\nabla\hat\rho_t), \qquad \hat\rho_0=\rho_0,

    where ϵt≥0\epsilon_t\geq0, ρt=Zt−1e−Ut\rho_t=Z_t^{-1}e^{-U_t} is the target density path, and Ft=−log⁡ZtF_t=-\log Z_t. With the PINN loss evaluated on the target path, the endpoint discrepancy obeys

    DKL(ρ^1∥ρ1)≤LPINN1(b^,F).D_{\mathrm{KL}}(\hat\rho_1\|\rho_1)\leq\sqrt{L_{\mathrm{PINN}}^1(\hat b,F)}.

    If the free-energy estimate F^t\widehat F_t instead satisfies ∫01∣∂tF^t−∂tFt∣2dt≤δ\int_0^1|\partial_t\widehat F_t-\partial_tF_t|^2dt\leq\delta for some δ≥0\delta\geq0, then

    DKL(ρ^1∥ρ1)≤2LPINN1(b^,F^)+2δ.D_{\mathrm{KL}}(\hat\rho_1\|\rho_1)\leq\sqrt{2L_{\mathrm{PINN}}^1(\hat b,\widehat F)+2\delta}.

    These bounds quantify how the residual loss, and error in the estimated free-energy derivative, constrain the sampler’s endpoint distribution.

  5. Knowl 5 — PINN training estimates its gradient without differentiating through paths

    algorithm

    NETS can train the PINN drift and free-energy models using walker paths while stopping gradients through the simulated paths. For on-policy training, the paths and log weights under the current drift provide importance-weighted estimates of target-path expectations; for off-policy training, locations may instead be sampled from another positive density. The training horizon TT can be increased gradually from a small value to 11 to limit early weight variance. Resampling may optionally be triggered when effective sample size falls below a chosen threshold.

    Input: Potential path U_t, initial sampler rho_0, n walkers, time horizon T, time grid with K steps, diffusion coefficient epsilon_t, models b_hat_t and F_hat_t, learning rate eta
    Initialize x_i from rho_0 and A_i = 0 for each walker i
    Repeat until the training loss converges:
        Draw and sort time points in [0,T]; initialize or continue the walker paths
        For each consecutive time interval [t_k,t_{k+1}] and each walker i:
            Update x_i by an Euler-Maruyama step for drift b_hat_t - epsilon_t grad U_t and noise sqrt(2 epsilon_t)
            Update A_i using div b_hat_t - grad U_t dot b_hat_t - partial_t U_t
        Estimate the on-policy PINN loss by time-averaging
            sum_i exp(A_i) abs(q_t(x_i))^2 / sum_i exp(A_i)
            where q_t = div b_hat_t - grad U_t dot b_hat_t - partial_t U_t + partial_t F_hat_t
        Treat simulated locations and weights as detached data; take a gradient step on b_hat_t and F_hat_t
        Optionally resample walkers if their effective sample size is below the selected threshold
        Optionally increase T toward 1 as training proceeds
    Output: Trained drift b_hat_t and free-energy model F_hat_t

    The paper leaves the walker count, grid resolution, learning rate, and resampling threshold to the experiment; it does not prescribe universal values. In the reported high-dimensional Gaussian-mixture training, the networks used width 512 and depth 4 and were trained for 4,000 iterations.

  6. Knowl 6 — Action Matching provides a second objective for gradient-form transport

    model/method

    A transport drift can be restricted to gradient form bt(x)=∇ϕt(x)b_t(x)=\nabla\phi_t(x), where ϕt:Rd→R\phi_t:\mathbb{R}^d\to\mathbb{R} is a scalar potential. For the prescribed density path ρt=Zt−1e−Ut\rho_t=Z_t^{-1}e^{-U_t}, NETS’s Action Matching objective for a candidate ϕ^t\widehat\phi_t, over T∈(0,1]T\in(0,1], is

    LAMT[ϕ^]=∫0T∫Rd(12∣∇ϕ^t(x)∣2+∂tϕ^t(x))ρt(x) dx dt+∫Rd(ϕ^0(x)ρ0(x)−ϕ^T(x)ρT(x))dx.L_{\mathrm{AM}}^T[\widehat\phi] =\int_0^T\int_{\mathbb{R}^d}\left(\frac12|\nabla\widehat\phi_t(x)|^2+\partial_t\widehat\phi_t(x)\right)\rho_t(x)\,dx\,dt +\int_{\mathbb{R}^d}\left(\widehat\phi_0(x)\rho_0(x)-\widehat\phi_T(x)\rho_T(x)\right)dx.

    Its minimizer is unique up to an additive spatial constant, and bt=∇ϕtb_t=\nabla\phi_t satisfies ∇⋅(btρt)=−∂tρt\nabla\cdot(b_t\rho_t)=-\partial_t\rho_t. Unlike the PINN objective, Action Matching requires the correct ρt\rho_t and therefore is not off-policy. The paper also gives a weight representation for gradient-form drifts that avoids explicitly evaluating the Laplacian Δϕ^t\Delta\widehat\phi_t when ϵt>0\epsilon_t>0.

  7. Knowl 7 — Stochastic-interpolant parameterization learns both the potential path and drift

    model/method

    A target density path can be parameterized through the stochastic interpolant It=αtx0+βtx1I_t=\alpha_t x_0+\beta_t x_1, where x0∼ρ0x_0\sim\rho_0, x1∼ρ1x_1\sim\rho_1, α0=β1=1\alpha_0=\beta_1=1, α1=β0=0\alpha_1=\beta_0=0, α˙t<0\dot\alpha_t<0, and β˙t>0\dot\beta_t>0. Its density evolves under the velocity bt(x)=E[α˙tx0+β˙tx1∣It=x]b_t(x)=\mathbb{E}[\dot\alpha_t x_0+\dot\beta_t x_1\mid I_t=x]. For Gaussian ρ0=N(0,Id)\rho_0=\mathcal N(0,I_d), the associated potential obeys ∇Ut(x)=αt−1E[x0∣It=x]\nabla U_t(x)=\alpha_t^{-1}\mathbb{E}[x_0\mid I_t=x].

    The paper uses αt=1−t2\alpha_t=\sqrt{1-t^2} and βt=t\beta_t=t, for which these relations imply tbt(x)=x−∇Ut(x)t b_t(x)=x-\nabla U_t(x). It parameterizes the potential by

    U^t(x)=12(1−t)∣x∣2+tU1(x)+t(1−t)f^t(x),b^t(x)=x−∇U1(x)−(1−t)∇f^t(x),\widehat U_t(x)=\frac12(1-t)|x|^2+tU_1(x)+t(1-t)\widehat f_t(x), \qquad \widehat b_t(x)=x-\nabla U_1(x)-(1-t)\nabla\widehat f_t(x),

    where U1U_1 is the target potential and f^t\widehat f_t is a neural network. Substituting these parameterizations into the PINN residual allows joint learning of the interpolating potential and transport without target samples. At zero residual, the resulting density matches the interpolant’s density path.

  8. Knowl 8 — NETS improves ESS and Wasserstein distance on the 40-mode Gaussian mixture

    data/table

    The paper compares NETS variants with published samplers on the 40-mode, two-dimensional Gaussian mixture whose means are interpolated from a standard Gaussian base with scale σ=2\sigma=2. Each method is evaluated using effective sample size (ESS) and the 2-Wasserstein distance W2W_2, estimated from 2,000 generated samples; reported values are means and standard deviations. The strongest NETS PINN and resampling variants have higher ESS and lower W2W_2 than the listed prior methods. The comparison is reproduced below.

    Method ESS ↑\uparrow W2↓W_2\downarrow
    FAB 0.653±0.0170.653\pm0.017 12.0±5.7312.0\pm5.73
    PIS 0.295±0.0180.295\pm0.018 7.64±0.927.64\pm0.92
    DDS 0.687±0.2080.687\pm0.208 9.31±0.829.31\pm0.82
    pDEM 0.634±0.0840.634\pm0.084 12.20±0.1412.20\pm0.14
    iDEM 0.734±0.0920.734\pm0.092 7.42±3.447.42\pm3.44
    CMCD-KL 0.268±0.0690.268\pm0.069 9.32±0.719.32\pm0.71
    CMCD-LV 0.655±0.0230.655\pm0.023 4.01±0.254.01\pm0.25
    NETS-AM, ϵt=5\epsilon_t=5 0.808±0.0310.808\pm0.031 3.89±0.223.89\pm0.22
    NETS-PINN, ϵt=0\epsilon_t=0 0.954±0.0030.954\pm0.003 3.55±0.573.55\pm0.57
    NETS-PINN, ϵt=4\epsilon_t=4 0.979±0.0020.979\pm0.002 3.14±0.463.14\pm0.46
    NETS-PINN with resampling 0.993±0.0040.993\pm0.004 3.27±0.313.27\pm0.31

    NETS used 100 sampling steps for these results. The resampling run used one resampling step when ESS fell below 98%; its ESS reached 0.993. The table’s NETS results demonstrate that learned transport can maintain high-weight efficiency and improve distributional fit on a multimodal target.

  9. Knowl 9 — NETS performs strongly on Neal’s Funnel and a 50-dimensional Student-t mixture

    data/table

    The paper evaluates NETS on Neal’s 10-dimensional Funnel and a 50-dimensional mixture of Student-t distributions (MoS), comparing maximum mean discrepancy (MMD) and 2-Wasserstein distance W2W_2 from the target. Scores are based on 2,000 model samples and 2,000 target samples; values are reported as means and standard deviations. NETS with resampling has the best listed MMD on Funnel, while on MoS NETS with resampling has the best listed MMD and W2W_2. For Funnel, NETS is competitive in MMD but has larger W2W_2 than the strongest listed baseline. Dashes indicate unreported results.

    Method Funnel MMD ↓\downarrow Funnel W2↓W_2\downarrow MoS MMD ↓\downarrow MoS W2↓W_2\downarrow
    FAB 0.032±0.0000.032\pm0.000 153.894±3.916153.894\pm3.916 0.093±0.0140.093\pm0.014 1204.160±147.71204.160\pm147.7
    GMMVI 0.031±0.0000.031\pm0.000 105.620±3.472105.620\pm3.472 0.135±0.0170.135\pm0.017 1255.216±296.91255.216\pm296.9
    PIS −- −- 0.218±0.0070.218\pm0.007 2113.172±31.172113.172\pm31.17
    DDS 0.172±0.0310.172\pm0.031 142.890±9.552142.890\pm9.552 0.131±0.0010.131\pm0.001 2154.884±3.8612154.884\pm3.861
    AFT 0.159±0.0100.159\pm0.010 145.138±6.061145.138\pm6.061 0.395±0.0820.395\pm0.082 2648.410±301.32648.410\pm301.3
    CRAFT 0.115±0.0030.115\pm0.003 134.335±0.663134.335\pm0.663 0.257±0.0240.257\pm0.024 1893.926±117.31893.926\pm117.3
    CMCD-KL 0.095±0.0030.095\pm0.003 513.339±192.4513.339\pm192.4 −- −-
    NETS-AM 0.041±0.0010.041\pm0.001 435.793±96.17435.793\pm96.17 0.0396±0.0010.0396\pm0.001 407.827±69.64407.827\pm69.64
    NETS-PINN 0.033±0.0020.033\pm0.002 388.91±141.5388.91\pm141.5 0.032±0.0010.032\pm0.001 482.393±174.6482.393\pm174.6
    NETS-PINN with resampling 0.027±0.0030.027\pm0.003 343.78±65.25343.78\pm65.25 0.030±0.0000.030\pm0.000 400.076±59.31400.076\pm59.31

    NETS used 100 sampling steps. For NETS-AM, ϵt=5\epsilon_t=5 on Funnel and 44 on MoS; for NETS-PINN, ϵt=5\epsilon_t=5 on both. Resampling was triggered when ESS fell below 70%.

  10. Knowl 10 — Transport enables high-dimensional Gaussian-mixture sampling, while diffusion improves it

    empirical result

    The paper trains a PINN NETS drift on 8-mode Gaussian mixtures in dimensions d=36,64,128,200d=36,64,128,200, using a feed-forward network of width 512 and depth 4 for 4,000 training iterations. In the high-dimensional ESS-versus-diffusion plot, transport alone (ϵt=0\epsilon_t=0) achieves about 60% ESS at d=200d=200, whereas annealed Langevin dynamics without learned transport yields ESS near zero across the tested dimensions. Increasing the diffusion coefficient improves performance when learned transport is present; at large diffusion the methods approach nearly independent sampling and the gaps across dimensions shrink. This requires finer time discretization: the experiments use between 100 steps at ϵt=0\epsilon_t=0 and 2,000 steps at ϵt=80\epsilon_t=80.

    A separate 50-dimensional MoS ablation reports decreasing W2W_2 as post-training diffusion is increased toward the ϵ→∞\epsilon\to\infty regime. The paper emphasizes that this limit requires increased integration resolution. Together, these experiments show that the diffusion coefficient can be adjusted after drift training to improve sampling, rather than being fixed by the training procedure.

  11. Knowl 11 — NETS improves sampling of lattice $\phi^4$ fields near and beyond a phase transition

    empirical result

    The paper tests NETS on two-dimensional lattice ϕ4\phi^4 theory, where a configuration is a real-valued field on an L×LL\times L lattice. The interpolating energy is Ut(ϕ)=∑x∼y∣ϕx−ϕy∣2+∑x(mt2ϕx2+λtϕx4)U_t(\phi)=\sum_{x\sim y}|\phi_x-\phi_y|^2+\sum_x(m_t^2\phi_x^2+\lambda_t\phi_x^4), with mt2m_t^2 and λt\lambda_t interpolated from a free-theory base with λ0=0\lambda_0=0. NETS is trained with the Action Matching objective.

    For the 400-dimensional L=20L=20 lattice, the target parameters are m12=−1.0m_1^2=-1.0, λ1=0.9\lambda_1=0.9, at the identified transition. NETS maintains nearly two orders of magnitude greater statistical efficiency than AIS without transport, as measured by ESS over the integration path. Its magnetization estimates are consistent with Hybrid Monte Carlo (HMC) reference samples, while AIS weights are too variable for a useful reweighted histogram. For the 256-dimensional L=16L=16 lattice, parameters m12=−1.0m_1^2=-1.0, λ1=0.8\lambda_1=0.8 place the system in the ordered phase; NETS samples both positive and negative ordered configurations, whereas AIS fails to sample the correct distribution and has unusably variable weights. The field-theory experiments require 1,500–2,000 integration steps to resolve dynamics near criticality.

Coverage note — The paper’s inertial and multimarginal NETS extensions, exact time-discretized weight correction, Hutchinson divergence estimator, and detailed lattice-action derivation are omitted because they are ancillary extensions or implementation details rather than the central sampler, guarantees, and experiments.

References

  1. 1.Akhound-Sadegh, T., Rector-Brooks, J., Bose, J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., Malkin, N., and Tong, A. Iterated denoising energy matching for sampling from boltzmann densities. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 760–786. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/akhound-sadegh24a.html.
  2. 2.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations, 2022.
  3. 3.Albergo, M. S., Kanwar, G., and Shanahan, P. E. Flow-based generative models for markov chain monte carlo in lattice field theory. Phys. Rev. D, 100:034515, Aug 2019. doi: 10.1103/PhysRevD.100.034515. URL https://link.aps.org/doi/10.1103/PhysRevD.100.034515.
  4. 4.Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797, 2023.
  5. 5.Arbel, M., Matthews, A. G. D. G., and Doucet, A. Annealed flow transport monte carlo. In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, 18–24 Jul 2021.
  6. 6.Arenz, O., Dahlinger, P., Ye, Z., Volpp, M., and Neumann, G. A unified perspective on natural gradient variational inference with gaussian mixture models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=tLBjsX4tjs.
  7. 7.Behjoo, H. and Chertkov, M. Harmonic path integral diffusion, 2024. URL https://arxiv.org/abs/2409.15166.
  8. 8.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling, 2024. URL https://arxiv.org/abs/2211.01364.
  9. 9.Blessing, D., Jia, X., Esslinger, J., Vargas, F., and Neumann, G. Beyond ELBOs: A large-scale evaluation of variational methods for sampling. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 4205–4229. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/blessing24a.html.
  10. 10.Bolic, M., Djurić, P. M., and Hong, S. Resampling algorithms for particle filters: A computational complexity perspective. EURASIP Journal on Advances in Signal Processing, 2004(15):403686, 2004. doi: 10.1155/S1110865704405149. URL https://doi.org/10.1155/S1110865704405149.
  11. 11.Bonanno, C., Nada, A., and Vadacchino, D. Mitigating topological freezing using out-of-equilibrium simulations. Journal of High Energy Physics, 2024(4):126, 2024. doi: 10.1007/JHEP04(2024)126. URL https://doi.org/10.1007/JHEP04(2024)126.
  12. 12.Bruna, J. and Han, J. Posterior sampling with denoising oracles via tilted transport, 2024. URL https://arxiv.org/abs/2407.00745.
  13. 13.Caselle, M., Cellini, E., Nada, A., and Panero, M. Stochastic normalizing flows as non-equilibrium transformations. Journal of High Energy Physics, 2022(7):15, 2022. doi: 10.1007/JHEP07(2022)015. URL https://doi.org/10.1007/JHEP07(2022)015.
  14. 14.Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/69386f6bb1dfed68692a24c8686939b9-Paper.pdf.
  15. 15.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion schrodinger bridge with applications to score-based generative modeling. In Advances in Neural Information Processing Systems, volume 34, pp. 17695–17709, 2021.
  16. 16.Del Moral, P. Nonlinear filtering: Interacting particle resolution. Comptes Rendus de l’Academie des Sciences - Series I - Mathematics, 325(6):653–658, 1997. ISSN 0764-4442. doi: https://doi.org/10.1016/S0764-4442(97)84778-7. URL https://www.sciencedirect.com/science/article/pii/S0764444297847787.
  17. 17.Doucet, A., de Freitas, N., and Gordon, N. J. (eds.). Sequential Monte Carlo Methods in Practice. Statistics for Engineering and Information Science. Springer, 2001. ISBN 978-1-4419-2887-0. doi: 10.1007/978-1-4757-3437-9. URL https://doi.org/10.1007/978-1-4757-3437-9.
  18. 18.Fan, M., Zhou, R., Tian, C., and Qian, X. Path-guided particle-based sampling. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=Kt4fwiuKqf.
  19. 19.Faulkner, M. F. and Livingstone, S. Sampling algorithms in statistical physics: a guide for statistics and machine learning, 2023. URL https://arxiv.org/abs/2208.04751.
  20. 20.Gabrié, M., Rotskoff, G. M., and Vanden-Eijnden, E. Adaptive monte carlo augmented with normalizing flows. Proceedings of the National Academy of Sciences, 119(10):e2109420119, 2022. doi: 10.1073/pnas.2109420119. URL https://www.pnas.org/doi/abs/10.1073/pnas.2109420119.
  21. 21.Grathwohl, W., Chen, R. T. Q., Bettencourt, J., and Duvenaud, D. Scalable reversible generative models with free-form continuous dynamics. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJxgknCcK7.
  22. 22.Hartmann, C., Schutte, C., and Zhang, W. Jarzynski’s equality, fluctuation theorems, and variance reduction: Mathematical analysis and numerical algorithms. Journal of Statistical Physics, 175:1214–1261, 2018. URL https://api.semanticscholar.org/CorpusID:59394118.
  23. 23.Hastings, W. K. Monte carlo sampling methods using markov chains and their applications. Biometrika, 57(1):97–109, 1970. ISSN 00063444, 14643510. URL http://www.jstor.org/stable/2334940.
  24. 24.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in neural information processing systems, volume 33, pp. 6840–6851, 2020.
  25. 25.Hua, M., Laurière, M., and Vanden-Eijnden, E. An efficient on-policy deep learning framework for stochastic optimal control. arXiv preprint arXiv:2410.05163, 2024.
  26. 26.Henin, J., Lelievre, T., Shirts, M. R., Valsson, O., and Delemotte, L. Enhanced sampling methods for molecular dynamics simulations [article v1.0]. Living Journal of Computational Molecular Science, 4(1):1583, Dec. 2022. doi: 10.33011/livecoms.4.1.1583. URL https://livecomsjournal.org/index.php/livecoms/article/view/v4i1e1583.
  27. 27.Jarzynski, C. Nonequilibrium equality for free energy differences. Phys. Rev. Lett., 78:2690–2693, Apr 1997. doi: 10.1103/PhysRevLett.78.2690. URL https://link.aps.org/doi/10.1103/PhysRevLett.78.2690.
  28. 28.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2022.
  29. 29.Liu, J. S. Metropolized independent sampling with comparisons to rejection sampling and importance sampling. Statistics and Computing, 6(2):113–119, 1996. doi: 10.1007/BF00162521. URL https://doi.org/10.1007/BF00162521.
  30. 30.Liu, X., Gong, C., and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2022.
  31. 31.Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. Gflownets and variational inference, 2023. URL https://arxiv.org/abs/2210.00580.
  32. 32.Matthews, A. G. D. G., Arbel, M., Rezende, D. J., and Doucet, A. Continual repeated annealed flow transport monte carlo. In Proceedings of the 39th International Conference on Machine Learning, Proceedings of Machine Learning Research, Jul 2022.
  33. 33.Midgley, L. I., Stimper, V., Simm, G. N. C., Scholkopf, B., and Hernandez-Lobato, J. M. Flow annealed importance sampling bootstrap. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=XCTVFJwS9LJ.
  34. 34.Maté, B. and Fleuret, F. Learning interpolations between boltzmann densities, 2023. URL https://arxiv.org/abs/2301.07388.
  35. 35.Neal, R. M. Probabilistic inference using markov chain monte carlo methods. Technical Report CRG-TR-93-1, Department of Computer Science, University of Toronto, September 1993.
  36. 36.Neal, R. M. Annealed importance sampling. Statistics and Computing, 11(2):125–139, 2001. doi: 10.1023/A:1008923215028. URL https://doi.org/10.1023/A:1008923215028.
  37. 37.Neklyudov, K., Brekelmans, R., Severo, D., and Makhzani, A. Action matching: Learning stochastic dynamics from samples. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 25858–25889. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/neklyudov23a.html.
  38. 38.Noe, F., Olsson, S., Köhler, J., and Wu, H. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365(6457):eaaw1147, 2019. doi: 10.1126/science.aaw1147. URL https://www.science.org/doi/abs/10.1126/science.aaw1147.
  39. 39.Parno, M. D. and Marzouk, Y. M. Transport map accelerated markov chain monte carlo. SIAM/ASA Journal on Uncertainty Quantification, 6(2):645–682, 2018. doi: 10.1137/17M1134640. URL https://doi.org/10.1137/17M1134640.
  40. 40.Sendera, M., Kim, M., Mittal, S., Lemos, P., Scimeca, L., Rector-Brooks, J., Adam, A., Bengio, Y., and Malkin, N. Improved off-policy training of diffusion samplers, 2024. URL https://arxiv.org/abs/2402.05098.
  41. 41.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  42. 42.Tian, Y., Panda, N., and Lin, Y. T. Liouville flow importance sampler. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 48186–48210. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/tian24c.html.
  43. 43.Vaikuntanathan, S. and Jarzynski, C. Escorted free energy simulations: Improving convergence by reducing dissipation. Phys. Rev. Lett., 100:190601, May 2008. doi: 10.1103/PhysRevLett.100.190601. URL https://link.aps.org/doi/10.1103/PhysRevLett.100.190601.
  44. 44.Vargas, F., Grathwohl, W. S., and Doucet, A. Denoising diffusion samplers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=8pvnfTAbu1f.
  45. 45.Vargas, F., Padhy, S., Blessing, D., and Nüsken, N. Transport meets variational inference: Controlled monte carlo diffusions. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=PP1rudnxiW.
  46. 46.Vierhaus, I. Simulation of ϕ⁴ theory in the strong coupling expansion beyond the Ising Limit. PhD thesis, Humboldt University of Berlin, 07 2010.
  47. 47.Wilson, K. G. Confinement of quarks. Phys. Rev. D, 10:2445–2459, Oct 1974. doi: 10.1103/PhysRevD.10.2445. URL https://link.aps.org/doi/10.1103/PhysRevD.10.2445.
  48. 48.Zhang, Q. and Chen, Y. Path integral sampler: A stochastic control approach for sampling. In International Conference on Learning Representations, 2021.
  49. 49.Zhang, Q. and Chen, Y. Path integral sampler: A stochastic control approach for sampling. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=_uCb2ynRu7Y.

Citation

MLA
Albergo, M. S., and E. Vanden-Eijnden. “NETS: A Non-Equilibrium Transport Sampler”. arXiv, 2024, http://arxiv.org/abs/2410.02711v3.
APA
Albergo, M. S., & Vanden-Eijnden, E. (2024). NETS: A Non-Equilibrium Transport Sampler. arXiv. http://arxiv.org/abs/2410.02711v3
Chicago
Albergo, M. S., and E. Vanden-Eijnden. 2024. “NETS: A Non-Equilibrium Transport Sampler”. arXiv. http://arxiv.org/abs/2410.02711v3.
Harvard
Albergo, M.S. and Vanden-Eijnden, E. (2024) “NETS: A Non-Equilibrium Transport Sampler”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.02711v3.
Vancouver
1. Albergo MS, Vanden-Eijnden E (2024) NETS: A Non-Equilibrium Transport Sampler. arXiv

BibTeX

@article{albergo2024nets,
  title = {NETS: A Non-Equilibrium Transport Sampler},
  author = {Albergo, Michael S. and Vanden-Eijnden, Eric},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.02711v3},
  eprint = {2410.02711}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/