Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Aaron J. HavensBenjamin Kurt MillerBing YanCarles Domingo-EnrichAnuroop SriramDaniel S. LevineBrandon M. WoodBin HuBrandon AmosBrian Karrer

article2025ICML87 citations

Introduces a scalable stochastic optimal control framework that trains diffusion samplers from unnormalized densities with dramatically fewer expensive energy evaluations, enabling efficient amortized molecular conformer generation.

Listen

Sampling from complex physical distributions using unnormalized energy functions is a fundamental challenge across computational chemistry, physics-based inference, and molecular modeling. Existing techniques—such as Markov Chain Monte Carlo, flow models, and diffusion samplers—struggle with high dimensions or require extensive evaluations of expensive physics-based energy calculations at every gradient step. The article introduces Adjoint Sampling, an on-policy variational inference framework based on stochastic optimal control that trains continuous diffusion samplers using unnormalized energy functions without requiring pre-existing ground truth data.

The main objective of the article is to demonstrate that Adjoint Sampling enables scalable, highly efficient sampling by decoupling model optimization from costly energy evaluations and trajectory simulations. The approach leverages a modified training objective called Reciprocal Adjoint Matching combined with an experience replay buffer. Instead of simulating full trajectories and computing energy gradients during every parameter update, the algorithm stores final states and their energy gradients, reconstructing intermediate states via exact analytical formulas. The methodology incorporates physical symmetries and periodic boundary conditions using equivariant neural networks, evaluating performance on synthetic physical potentials (such as 55-particle Lennard-Jones systems) and large-scale molecular conformer generation across thousands of organic molecules from the SPICE and GEOM-DRUGS benchmarks.

The findings show that Adjoint Sampling significantly reduces computational overhead while maintaining superior sample accuracy. On synthetic multi-particle systems, the method achieves competitive geometric distance metrics while reducing energy function evaluations per gradient update by several orders of magnitude compared to prior score-matching approaches. In molecular conformer generation benchmarks, Adjoint Sampling consistently outperforms standard industry baselines such as RDKit in coverage recall across varied structural complexities. When combined with standard post-generation relaxation, pretrained Cartesian Adjoint Sampling attains 96.65% recall coverage on SPICE and 87.01% on GEOM-DRUGS, substantially exceeding classical rule-based methods.

These results imply that high-fidelity molecular conformer generation and physical simulation can scale effectively without ground-truth structural training sets, reducing computational costs and turnaround times in drug discovery and material design workflows. Practitioners should consider adopting Adjoint Sampling for large-scale molecular conformation pipelines, utilizing torsional models when fast unrelaxed sampling is required and pretrained Cartesian models when highest recall after structural relaxation is prioritized. However, users should note that the framework targets unnormalized energy approximations whose reliability depends on the underlying energy potential, requiring standard relaxation and validation steps before making downstream experimental decisions.

arXiv: 2504.11713facebookresearch/adjoint_sampling

No sufficiently relevant recommendations were found.

Cover for Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

Abstract

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry. Code & and Benchmarks provided at github.com/facebookresearch/adjoint_sampling.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Optimizing Diffusion Processes for Sampling
  • 3. Adjoint Sampling
  • 3.1. Adjoint Matching
  • 3.2. Reciprocal Adjoint Matching
  • 3.3. Adjoint Sampling Theory
  • 3.4. Geometric Extensions
  • 4. Related Work
  • 5. Experiments
  • 5.1. Synthetic Energy Functions
  • 5.2. Sampling Conformers from an Energy Function
  • 6. Conclusion
  • Impact Statement
  • References
  • A. Additional preliminaries
  • A.1. Stochastic optimal control
  • A.2. Adjoint Matching
  • B. Base process derivations
  • B.1. Euclidean space
  • B.2. Flat tori
  • C. Theory of Reciprocal Adjoint Matching
  • C.1. Reciprocal projections decrease the SOC objective
  • C.2. Adjoint Sampling Preserves Critical Points of Adjoint Matching
  • D. Geometric Extensions
  • D.1. Graph-Conditioned SE(3)-Invariant Sampling
  • D.2. Periodic boundary conditions
  • E. Bridge Matching and Pretraining
  • F. Synthetic Energy Experiment Details
  • F.1. Double Well Potential (DW-4)
  • F.2. Lennard-Jones Potential (LJ-13, LJ-55)
  • F.3. Architectures and Hyper parameters
  • F.4. Reported Metrics
  • G. Sampling Molecular Conformers
  • G.1. Sampling the target molecule in Cartesian coordinates
  • G.2. Torsion Angles and Torques
  • G.3. Coverage Recall and Precision Metrics
  • G.4. Hyperparameters and Architecture Details for SPICE and GEOM-DRUGS
  • H. Data Preparation for the Conformation Benchmark
  • H.1. SPICE Dataset
  • H.2. GEOM-DRUGS Dataset
  • I. Additional Experimental Results
  • I.1. Energy Histograms for Synthetic Energy Experiments
  • I.2. Runtime Analysis of Molecular Conformer Generation
  • I.3. Choice of energy function
  • I.4. Ablating Reciprocal Projection
  • I.5. Threshold Ablation for Molecular Conformation Generation

Knowls

  1. Knowl 1 — Sampling a Boltzmann target as a minimum-energy control problem

    model/method

    Let x∈Rdx\in\mathbb{R}^d be a configuration with energy E(x)E(x), temperature τ>0\tau>0, and finite normalizer Z=∫Rdexp⁡[−E(x)/τ]dxZ=\int_{\mathbb{R}^d}\exp[-E(x)/\tau]dx. The target density is μ(x)=Z−1exp⁡[−E(x)/τ]\mu(x)=Z^{-1}\exp[-E(x)/\tau]. Adjoint Sampling represents a sampler as the controlled diffusion dXt=σ(t)u(Xt,t)dt+σ(t)dBtdX_t=\sigma(t)u(X_t,t)dt+\sigma(t)dB_t, X0=0X_0=0, for t∈[0,1]t\in[0,1], where BtB_t is standard dd-dimensional Brownian motion, σ(t)\sigma(t) is a scalar noise schedule, and uu is a learnable drift control. Write p1basep^{\mathrm{base}}_1 for the time-1 density of the same diffusion with u=0u=0. The control is learned by minimizing J(u)=EX∼pu[∫0112∥u(Xt,t)∥2dt+g(X1)]J(u)=\mathbb{E}_{X\sim p^u}[\int_0^1\tfrac12\|u(X_t,t)\|^2dt+g(X_1)], with terminal cost g(x)=log⁡p1base(x)+E(x)/τg(x)=\log p^{\mathrm{base}}_1(x)+E(x)/\tau. The omitted additive constant in gg is log⁡Z\log Z and does not affect the minimizing control. Under the paper's stated mild regularity conditions, the optimal control produces terminal density p1u∗=μp^{u^*}_1=\mu; equivalently, it realizes the minimum-path-KL Schrödinger bridge from the point mass at the origin to μ\mu. This reformulation lets training use the energy function and its gradient without requiring target-distribution samples.

  2. Knowl 2 — Adjoint Matching simplifies to regression on the terminal energy gradient

    equation

    For the zero-base-drift, zero-running-cost diffusion dXt=σ(t)u(Xt,t)dt+σ(t)dBtdX_t=\sigma(t)u(X_t,t)dt+\sigma(t)dB_t with terminal cost gg, the lean adjoint state is constant along a trajectory: it equals ∇g(X1)\nabla g(X_1). Thus the Adjoint Matching objective can be written as LAM(u)=EX∼puˉ[∫0112∥u(Xt,t)+σ(t)∇g(X1)∥2dt]\mathcal{L}_{\mathrm{AM}}(u)=\mathbb{E}_{X\sim p^{\bar u}}[\int_0^1\tfrac12\|u(X_t,t)+\sigma(t)\nabla g(X_1)\|^2dt]. Here X=(Xt)t∈[0,1]X=(X_t)_{t\in[0,1]} is a trajectory, puˉp^{\bar u} is generated by the current control with its sampling path stopped from receiving gradients, and uˉ=stopgrad⁡(u)\bar u=\operatorname{stopgrad}(u). The control is trained to regress toward −σ(t)∇g(X1)-\sigma(t)\nabla g(X_1) at intermediate states. In the sampling problem, g(x)=log⁡p1base(x)+E(x)/τg(x)=\log p^{\mathrm{base}}_1(x)+E(x)/\tau, so this target includes both the terminal base-density gradient and the temperature-scaled energy gradient. This simplification avoids a separate backward simulation of the lean adjoint.

  3. Knowl 3 — Reciprocal Adjoint Matching replaces trajectory pairs with tractable bridge samples

    model/method

    Adjoint Sampling replaces the trajectory expectation in Adjoint Matching with an expectation over terminal samples from the current controlled diffusion and intermediate states drawn from the zero-control base diffusion conditioned on those terminal samples. Its Reciprocal Adjoint Matching objective is LRAM(u)=∫01EX1∼p1uˉ, Xt∼pt∣1base(⋅∣X1)[λ(t)2∥u(Xt,t)+σ(t)∇g(X1)∥2]dt\mathcal{L}_{\mathrm{RAM}}(u)=\int_0^1\mathbb{E}_{X_1\sim p^{\bar u}_1,\,X_t\sim p^{\mathrm{base}}_{t|1}(\cdot|X_1)}[\tfrac{\lambda(t)}{2}\|u(X_t,t)+\sigma(t)\nabla g(X_1)\|^2]dt, with λ(t)=1/σ(t)2\lambda(t)=1/\sigma(t)^2. Here p1uˉp^{\bar u}_1 is the time-1 distribution of the current stopped control, and pt∣1basep^{\mathrm{base}}_{t|1} is the conditional distribution of the zero-drift base process at time tt given its time-1 state. The base conditional is available in closed form for the schedules used in the paper, making it possible to create many intermediate-state training examples from each costly terminal sample. The conditional draws at different times need not come from one simulated trajectory. The time weighting changes numerical scaling without changing the stated optimal solution.

  4. Knowl 4 — Alternating replay-buffer training algorithm

    algorithm

    The procedure alternates costly on-policy terminal sampling and energy-gradient evaluation with repeated inexpensive RAM updates. Inputs are a terminal cost gg, a base process with tractable pt∣1basep^{\mathrm{base}}_{t|1}, a parameterized control uθu_\theta, outer-loop batch size nn, inner-loop batch size mm, a replay buffer, and optimizer settings. It returns the learned diffusion drift. The paper uses Euler–Maruyama without gradient tracking to generate terminal samples. It does not specify a single universal stopping budget or buffer capacity; those are selected for each experiment.

    Input: Terminal cost g; base posterior p_base(t|1); control u_theta; outer batch size n; inner batch size m; replay buffer B; optimizer
    Initialize B as empty
    while training budget remains:
        Sample n terminal states X1 from the controlled diffusion with stopgrad(u_theta)
        Evaluate grad g(X1) once for each new terminal state
        Add each pair (X1, grad g(X1)) to B, retaining the configured buffer contents
        Repeat the configured number of inner updates:
            Sample m pairs (X1, grad g(X1)) uniformly from B
            Sample m times t independently and uniformly from [0, 1]
            For each pair, draw Xt from p_base(t|1)(.|X1)
            Compute the batch mean of lambda(t)/2 * ||u_theta(Xt,t) + sigma(t) grad g(X1)||^2
            Update theta by an optimizer step on this loss
    Output: The trained control u_theta defining the diffusion sampler
  5. Knowl 5 — Reciprocal projection improves the control objective and preserves the optimal fixed point

    theoretical result

    For a control uu, define its reciprocal projection P(u)\mathcal{P}(u) as the control whose path distribution minimizes path KL divergence to the zero-drift base process conditioned on the time-1 marginal of uu: P(u)=arg⁡min⁡vDKL(pv(X)∥pbase(X∣X1)p1u(X1))\mathcal{P}(u)=\arg\min_v D_{\mathrm{KL}}(p^v(X)\|p^{\mathrm{base}}(X|X_1)p^u_1(X_1)). Define the stochastic-control objective J(u)=EX∼pu[∫0112∥u(Xt,t)∥2dt+g(X1)]J(u)=\mathbb{E}_{X\sim p^u}[\int_0^1\tfrac12\|u(X_t,t)\|^2dt+g(X_1)]. The paper establishes that LRAM(P(u))=LAM(P(u))\mathcal{L}_{\mathrm{RAM}}(\mathcal{P}(u))=\mathcal{L}_{\mathrm{AM}}(\mathcal{P}(u)) and J(P(u))≤J(u)J(\mathcal{P}(u))\leq J(u). In the idealized alternating procedure, where terminal samples are drawn from the current control and the RAM regression is fully minimized before collecting new samples, the update satisfies ui+1=P(ui)−δLAMδu(P(ui))u_{i+1}=\mathcal{P}(u_i)-\frac{\delta\mathcal{L}_{\mathrm{AM}}}{\delta u}(\mathcal{P}(u_i)). A fixed point satisfying u=P(u)u=\mathcal{P}(u) and u=u−δLAMδu(u)u=u-\frac{\delta\mathcal{L}_{\mathrm{AM}}}{\delta u}(u) is the optimal control for the stated stochastic-control problem. These guarantees concern the idealized update; practical training uses a finite replay buffer and partial optimization between sample-collection steps.

  6. Knowl 6 — Equivariant Cartesian sampling enforces molecular symmetries and removes translation

    model/method

    For a graph of kk particles with spatial dimension dd, the Cartesian sampler conditions its drift on graph features and makes it equivariant to graph automorphisms and spatial rotations. Translation invariance is enforced by restricting positions to the zero-center-of-mass subspace. For unconstrained coordinates y∈Rkdy\in\mathbb{R}^{kd}, set x=Ayx=Ay, where A=(Ik−1k1k1k⊤)⊗IdA=(I_k-\frac{1}{k}\mathbf{1}_k\mathbf{1}_k^\top)\otimes I_d; IkI_k and IdI_d are identity matrices and 1k\mathbf{1}_k is the all-ones vector. The controlled process uses projected drift and noise, dXt=σ(t)Auθ(Xt,t;G)dt+σ(t)AdBtdX_t=\sigma(t)A u_\theta(X_t,t;G)dt+\sigma(t)A dB_t, with graph GG. Its base marginal is a possibly singular Gaussian with covariance νtAA⊤\nu_tAA^\top, where νt=∫0tσ(s)2ds\nu_t=\int_0^t\sigma(s)^2ds. The corresponding RAM regression is projected through AA, so the learned update remains in the centered subspace. Under the paper's invariance assumptions on the energy and base density, the construction yields distributions invariant to graph-preserving permutations and rotations. The Cartesian implementation uses an equivariant graph neural network.

  7. Knowl 7 — Torsional sampling models molecular conformations on a flat torus

    model/method

    The torsional variant represents a molecule by its drotd_{\mathrm{rot}} torsion angles ϕ∈[−π,π)drot\phi\in[-\pi,\pi)^{d_{\mathrm{rot}}}, with opposite endpoints identified, so the state space is a flat torus. A torsion is associated with four atoms around a central bond; changing it rotates atoms on one side of that bond while preserving bond lengths and bond angles. The initial state is chosen with all torsions set to zero, using a cis-like convention in which the outer atoms of each torsion are as close as possible. The diffusion and RAM posterior are defined on the torus, using wrapped base-process distributions to handle periodicity. The energy-gradient regression target is transformed from Cartesian coordinates to torsion coordinates through the Jacobian of the map from torsion angles to atomic positions; the resulting torsion gradient is clipped during training. The model predicts torque quantities rather than unconstrained Cartesian updates. Because only torsions change, this representation cannot break molecular bonds and does not need the Cartesian bond-structure energy regularizer.

  8. Knowl 8 — Synthetic energy benchmarks show competitive sample quality with far fewer energy evaluations

    empirical result

    The paper evaluates sampling on the 2D four-particle double-well system DW-4 (d=8d=8), and 3D Lennard–Jones systems LJ-13 (d=39d=39) and LJ-55 (d=165d=165), comparing samples with long-run MCMC references. Geometric W2W_2 accounts for point-cloud symmetries; E(⋅)W2E(\cdot)W_2 is the one-dimensional Wasserstein-2 distance between generated and reference energy distributions; path-ESS is a normalized path effective sample size, omitted for iDEM because its objective does not define the relevant path weights. Lower distances and higher path-ESS are better. The table reports selected methods and all three metrics where reported. The evaluation counts are per gradient update under the LJ-55 experiment hyperparameters. Adjoint Sampling is comparable to iDEM on geometric W2W_2 while substantially improving energy-distribution agreement for LJ-13 and LJ-55; it also uses only 0.002 energy evaluations and 3 control-network evaluations per update, compared with 512 and 3 for full iDEM and 1 and 1000 for PIS. The one-Monte-Carlo-sample iDEM variant has an unstable LJ-55 energy metric.

    DW-4 (d=8d=8) LJ-13 (d=39d=39) LJ-55 (d=165d=165) Evaluations/update
    Method path-ESS W2W_2 E(⋅)W2E(\cdot)W_2 path-ESS W2W_2 E(⋅)W2E(\cdot)W_2 path-ESS W2W_2 E(⋅)W2E(\cdot)W_2 Energy uθu_\theta
    PIS 0.462±\pm0.081 0.68±\pm0.23 0.65±\pm0.25 0.012±\pm0.011 1.93±\pm0.07 18.02±\pm1.12 0.001±\pm0.000 4.79±\pm0.45 228.70±\pm131.27 1 1000
    iDEM – 0.70±\pm0.06 0.55±\pm0.14 – 1.61±\pm0.01 30.78±\pm24.46 – 4.69±\pm1.52 93.53±\pm16.31 512 3
    iDEM (1 MC sample) – 1.21±\pm0.02 2.70±\pm0.18 – 2.03±\pm0.02 22.41±\pm0.18 – 5.79±\pm1.60 1e32±\pm1e32 1 3
    Adjoint Sampling without reciprocal projection 0.448±\pm0.110 0.63±\pm0.11 1.03±\pm0.23 0.159±\pm0.068 1.68±\pm0.01 2.91±\pm1.39 0.094±\pm0.025 4.50±\pm0.10 94.48±\pm76.12 0.002 3
    Adjoint Sampling 0.627±\pm0.037 0.62±\pm0.06 0.55±\pm0.12 0.220±\pm0.041 1.67±\pm0.01 2.40±\pm1.25 0.066±\pm0.037 4.50±\pm0.05 58.04±\pm20.98 0.002 3
  9. Knowl 9 — Amortized conformer-generation benchmark and evaluation protocol

    experimental setup

    The molecular task trains graph-conditioned samplers to generate conformers from molecular topology and an energy model, without using reference atomic configurations during training. Training uses 24,477 SPICE molecular graphs; evaluation uses 80 held-out SPICE molecules and 80 GEOM-DRUGS molecules, each test set spanning molecules with 3–10 rotatable bonds. Reference conformers are obtained from molecular-chemistry pipelines for SPICE and from GEOM-DRUGS for the external test set. The energy model is eSEN, pretrained on SPICE-MACE-OFF; Cartesian sampling also adds a flat-bottom bond-structure regularizer, whereas torsional sampling changes only torsion angles. An optional Bridge Matching pretraining phase uses RDKit conformers to initialize exploration. Results are evaluated both before and after energy-based L-BFGS relaxation. Coverage recall measures the fraction of reference conformers with a generated match below the RMSD threshold; coverage precision measures the fraction of generated conformers matching a reference. AMR is the average minimum RMSD in the corresponding direction. The reported table uses a 1.25 Å threshold. The released reference conformers are intended as representative held-out low-energy structures, not as an exhaustive enumeration or samples from a fixed-temperature Boltzmann distribution.

  10. Knowl 10 — Conformer sampling improves recall over RDKit, especially with relaxation and pretraining

    empirical result

    The results below compare RDKit ETKDG with torsional and Cartesian Adjoint Sampling on SPICE and GEOM-DRUGS. Coverage is reported in percent at an RMSD threshold of 1.25 Å; AMR is in Å, and lower AMR is better. Each cell is mean ±\pm standard deviation across molecules. Without relaxation, both Adjoint Sampling variants generally improve recall over RDKit on SPICE; on GEOM-DRUGS, they also improve recall, while precision is mixed and Cartesian precision is lower. Relaxation raises coverage for all methods and particularly benefits Cartesian Adjoint Sampling with pretraining. The best recall in each dataset's relaxed comparison is 96.65% on SPICE and 87.01% on GEOM-DRUGS for Cartesian Adjoint Sampling with pretraining. On the external GEOM-DRUGS set, its relaxed precision coverage remains below RDKit's, illustrating the recall–precision trade-off.

    SPICE GEOM-DRUGS
    Recall Precision Recall Precision
    Method Cov. AMR Cov. AMR Cov. AMR Cov. AMR
    Without relaxation
    RDKit ETKDG 72.74±\pm33.18 1.04±\pm0.52 69.68±\pm37.11 1.14±\pm0.64 63.51±\pm34.74 1.15±\pm0.61 69.77±\pm38.23 1.09±\pm0.66
    Torsional Adjoint Sampling 85.06±\pm24.61 0.86±\pm0.30 70.42±\pm34.54 1.06±\pm0.54 72.91±\pm31.17 0.98±\pm0.40 67.85±\pm36.01 1.09±\pm0.55
    Cartesian Adjoint Sampling 82.22±\pm25.72 0.96±\pm0.26 49.13±\pm33.01 1.26±\pm0.38 60.93±\pm35.15 1.20±\pm0.43 28.44±\pm27.77 1.86±\pm0.64
    Cartesian Adjoint Sampling (+pretrain) 89.42±\pm17.48 0.84±\pm0.24 65.93±\pm29.53 1.12±\pm0.34 72.98±\pm30.82 1.02±\pm0.34 45.14±\pm31.46 1.47±\pm0.52
    With relaxation
    RDKit ETKDG 81.61±\pm27.58 0.79±\pm0.44 74.51±\pm35.07 0.97±\pm0.64 71.72±\pm29.73 0.93±\pm0.53 75.37±\pm32.76 0.89±\pm0.60
    Torsional Adjoint Sampling 88.25±\pm21.17 0.72±\pm0.33 74.66±\pm32.96 0.94±\pm0.59 76.62±\pm27.73 0.87±\pm0.40 71.85±\pm32.34 0.97±\pm0.54
    Cartesian Adjoint Sampling 94.10±\pm15.67 0.68±\pm0.28 62.80±\pm29.63 1.02±\pm0.37 79.08±\pm29.44 0.89±\pm0.45 47.88±\pm30.92 1.40±\pm0.61
    Cartesian Adjoint Sampling (+pretrain) 96.65±\pm7.51 0.60±\pm0.23 77.20±\pm25.76 0.92±\pm0.35 87.01±\pm22.79 0.76±\pm0.34 61.96±\pm31.55 1.10±\pm0.46

Coverage note — Detailed architecture settings, the separate energy-model comparison, and the full reference-conformer construction pipeline are omitted as supporting experimental detail; the main method, theoretical guarantee, benchmark design, caveat, and principal quantitative results are included.

References

  1. 1.Akhound-Sadegh, T., Rector-Brooks, J., Bose, J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., et al. Iterated denoising energy matching for sampling from boltzmann densities. In Forty-first International Conference on Machine Learning, 2024.
  2. 2.Albergo, M. S. and Vanden-Eijnden, E. Nets: A non-equilibrium transport sampler. arXiv preprint arXiv:2410.02711, 2024.
  3. 3.Albergo, M. S., Kanwar, G., and Shanahan, P. E. Flow-based generative models for markov chain monte carlo in lattice field theory. Physical Review D, 100(3):034515, 2019.
  4. 4.Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797, 2023.
  5. 5.Arbel, M., Matthews, A., and Doucet, A. Annealed flow transport monte carlo. In International Conference on Machine Learning, pp. 318–330. PMLR, 2021.
  6. 6.Axelrod, S. and Gomez-Bombarelli, R. Geom, energy-annotated molecular conformations for property prediction and molecular generation. Scientific Data, 9(1):185, 2022.
  7. 7.Batatia, I., Kovacs, D. P., Simm, G., Ortner, C., and Csanyi, G. Mace: Higher order equivariant message passing neural networks for fast and accurate force fields. Advances in neural information processing systems, 35:11423–11436, 2022.
  8. 8.Bellman, R. Dynamic programming. Princeton Landmarks in Mathematics. Princeton University Press, Princeton, NJ, 2010., 1957.
  9. 9.Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E. Gflownet foundations. Journal of Machine Learning Research, 24(210):1–55, 2023.
  10. 10.Berner, J., Richter, L., and Ullrich, K. An optimal control perspective on diffusion-based generative modeling. arXiv preprint arXiv:2211.01364, 2023.
  11. 11.Bryson, A. E. and Ho, Y. C. Applied Optimal Control. Blaisdell, 1969.
  12. 12.Chen, J., Richter, L., Berner, J., Blessing, D., Neumann, G., and Anandkumar, A. Sequential controlled langevin diffusions. arXiv preprint arXiv:2412.07081, 2024a.
  13. 13.Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.
  14. 14.Chen, Y., Georgiou, T. T., and Pavon, M. On the relation between optimal transport and schrodinger bridges: A stochastic control viewpoint. Journal of Optimization Theory and Applications, 169:671–691, 2016.
  15. 15.Chen, Y., Goldstein, M., Hua, M., Albergo, M. S., Boffi, N. M., and Vanden-Eijnden, E. Probabilistic forecasting with stochastic interpolants and follmer processes. arXiv preprint arXiv:2403.13724, 2024b.
  16. 16.Dai Pra, P. A stochastic control approach to reciprocal diffusion processes. Applied mathematics and Optimization, 23(1):313–329, 1991.
  17. 17.De Bortoli, V., Mathieu, E., Hutchinson, M., Thornton, J., Teh, Y. W., and Doucet, A. Riemannian score-based generative modelling. Advances in Neural Information Processing Systems, 35:2406–2422, 2022.
  18. 18.De Bortoli, V., Hutchinson, M., Wirnsberger, P., and Doucet, A. Target score matching. arXiv preprint arXiv:2402.08667, 2024.
  19. 19.Del Moral, P., Doucet, A., and Jasra, A. Sequential monte carlo samplers. Journal of the Royal Statistical Society Series B: Statistical Methodology, 68(3):411–436, 2006.
  20. 20.Diez, J. V., Atance, S. R., Engkvist, O., and Olsson, S. Generation of conformational ensembles of small molecules via surrogate model-assisted molecular dynamics. Machine Learning: Science and Technology, 5(2):025010, 2024.
  21. 21.Domingo-Enrich, C., Drozdzal, M., Karrer, B., and Chen, R. T. Q. Adjoint matching: Fine-tuning flow and diffusion generative models with memoryless stochastic optimal control, 2024. URL https://arxiv.org/abs/2409.08861.
  22. 22.Eastman, P., Behara, P. K., Dotson, D. L., Galvelis, R., Herr, J. E., Horton, J. T., Mao, Y., Chodera, J. D., Pritchard, B. P., Wang, Y., et al. Spice, a dataset of drug-like molecules and peptides for training machine learning potentials. Scientific Data, 10(1):11, 2023.
  23. 23.Fleming, W. and Rishel, R. Deterministic and Stochastic Optimal Control. Stochastic Modelling and Applied Probability. Springer New York, 2012.
  24. 24.Follmer, H. An entropy approach to the time reversal of diffusion processes. In Stochastic Differential Systems Filtering and Control: Proceedings of the IFIP-WG 7/1 Working Conference Marseille-Luminy, France, March 12–17, 1984, pp. 156–163. Springer, 2005.
  25. 25.Friede, M., Holzer, C., Ehlert, S., and Grimme, S. dxtb—an efficient and fully differentiable framework for extended tight-binding. The Journal of Chemical Physics, 161(6), 2024.
  26. 26.Fu, X., Wood, B. M., Barroso-Luque, L., Levine, D. S., Gao, M., Dzamba, M., and Zitnick, C. L. Learning smooth and expressive interatomic potentials for physical property prediction. arXiv preprint arXiv:2502.12147, 2025.
  27. 27.Gabrie , M., Rotskoff, G. M., and Vanden-Eijnden, E. Adaptive monte carlo augmented with normalizing flows. Proceedings of the National Academy of Sciences, 119(10), mar 2022.
  28. 28.Ganea, O., Pattanaik, L., Coley, C., Barzilay, R., Jensen, K., Green, W., and Jaakkola, T. Geomol: Torsional geometric generation of molecular 3d conformer ensembles. Advances in Neural Information Processing Systems, 34:13757–13769, 2021.
  29. 29.Geiger, M. and Smidt, T. e3nn: Euclidean neural networks, 2022. URL https://arxiv.org/abs/2207.09453.
  30. 30.Geiger, M., Smidt, T., M., A., Miller, B. K., Boomsma, W., Dice, B., Lapchevskyi, K., Weiler, M., Tyszkiewicz, M., Batzner, S., Madisetti, D., Uhrin, M., Frellsen, J., Jung, N., Sanborn, S., Wen, M., Rackers, J., Rød, M., and Bailey, M. Euclidean neural networks: e3nn, April 2022. URL https://doi.org/10.5281/zenodo.6459381.
  31. 31.Grimme, S., Bannwarth, C., and Shushkov, P. A robust and accurate tight-binding quantum chemical method for structures, vibrational frequencies, and noncovalent interactions of large molecular systems parametrized for all spd-block elements (z= 1–86). Journal of chemical theory and computation, 13(5):1989–2009, 2017.
  32. 32.Hassan, M., Shenoy, N., Lee, J., Stark, H., Thaler, S., and Beaini, D. Et-flow: Equivariant flow-matching for molecular conformer generation. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  33. 33.He, J., Du, Y., Vargas, F., Zhang, D., Padhy, S., OuYang, R., Gomes, C., and Hernandez-Lobato, J. M. No trick, no treat: Pursuits and challenges towards simulation-free training of neural samplers. arXiv preprint arXiv:2502.06685, 2025.
  34. 34.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33. Curran Associates, Inc., 2020.
  35. 35.Holderrieth, P., Albergo, M. S., and Jaakkola, T. LEAPS: A discrete neural sampler via locally equivariant networks. arXiv preprint arXiv:2502.10843, 2025.
  36. 36.Hua, M., Lauriere, M., and Vanden-Eijnden, E. A simulation-free deep learning approach to stochastic optimal control. arXiv preprint arXiv:2410.05163, 2024.
  37. 37.Huang, C.-W., Aghajohari, M., Bose, J., Panangaden, P., and Courville, A. C. Riemannian diffusion models. Advances in Neural Information Processing Systems, 35:2750–2761, 2022.
  38. 38.Jing, B., Corso, G., Chang, J., Barzilay, R., and Jaakkola, T. Torsional diffusion for molecular conformer generation. Advances in Neural Information Processing Systems, 35:24240–24253, 2022.
  39. 39.Kappen, H. J. Path integrals and symmetry breaking for optimal control theory. Journal of Statistical Mechanics: Theory and Experiment, 2005(11), nov 2005.
  40. 40.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. Advances in neural information processing systems, 35:26565–26577, 2022.
  41. 41.Klein, L., Foong, A., Fjelde, T., Mlodozeniec, B., Brockschmidt, M., Nowozin, S., Noe, F., and Tomioka, R. Timewarp: Transferable acceleration of molecular dynamics by learning time-coarsened dynamics. Advances in Neural Information Processing Systems, 36, 2024a.
  42. 42.Klein, L., Kramer, A., and Noe, F. Equivariant flow matching. Advances in Neural Information Processing Systems, 36, 2024b.
  43. 43.Kohler, J., Klein, L., and Noe, F. Equivariant flows: exact likelihood generative learning for symmetric densities. In International conference on machine learning, pp. 5361–5370. PMLR, 2020.
  44. 44.Kohler, J., Invernizzi, M., De Haan, P., and Noe, F. Rigid body flows for sampling molecular crystal structures. In International Conference on Machine Learning, pp. 17301–17326. PMLR, 2023.
  45. 45.Kondor, R., Lin, Z., and Trivedi, S. Clebsch-gordan nets: a fully fourier space spherical convolutional neural network, 2018. URL https://arxiv.org/abs/1806.09231.
  46. 46.Kovacs, D. P., Moore, J. H., Browning, N. J., Batatia, I., Horton, J. T., Kapil, V., Witt, W. C., Magdau, I.-B., Cole, D. J., and Csanyi, G. Mace-off23: Transferable machine learning force fields for organic molecules. arXiv preprint arXiv:2312.15211, 2023.
  47. 47.Lahlou, S., Deleu, T., Lemos, P., Zhang, D., Volokhova, A., Hernandez-Garcıa, A., Ezzine, L. N., Bengio, Y., and Malkin, N. A theory of continuous generative flow networks. In International Conference on Machine Learning, pp. 18269–18300. PMLR, 2023.
  48. 48.Landrum, G. Rdkit documentation. Release, 1(1-79):4, 2013.
  49. 49.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023.
  50. 50.Liu, G.-H., Vahdat, A., Huang, D.-A., Theodorou, E. A., Nie, W., and Anandkumar, A. I2sb: Image-to-image schr\” odinger bridge. arXiv preprint arXiv:2302.05872, 2023.
  51. 51.Liu, G.-H., Lipman, Y., Nickel, M., Karrer, B., Theodorou, E., and Chen, R. T. Q. Generalized schrodinger bridge matching. In The Twelfth International Conference on Learning Representations, 2024.
  52. 52.Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In European Conference on Computer Vision, pp. 23–40. Springer, 2024.
  53. 53.Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. Trajectory balance: Improved credit assignment in gflownets. arXiv preprint arXiv:2201.13259, 2023.
  54. 54.Matthews, A., Arbel, M., Rezende, D. J., and Doucet, A. Continual repeated annealed flow transport monte carlo. In International Conference on Machine Learning, pp. 15196–15219. PMLR, 2022.
  55. 55.Midgley, L., Stimper, V., Antoran, J., Mathieu, E., Scholkopf, B., and Hernandez-Lobato, J. M. SE(3) equivariant augmented coupling flows. Advances in Neural Information Processing Systems, 36, 2024.
  56. 56.Midgley, L. I., Stimper, V., Simm, G. N., Scholkopf, B., and Hernandez-Lobato, J. M. Flow annealed importance sampling bootstrap. In The Twelfth International Conference on Learning Representations: ICLR 2024, 2023.
  57. 57.Miller, B. K., Geiger, M., Smidt, T. E., and Noe, F. Relevance of rotationally equivariant convolutions for predicting molecular properties. arXiv preprint arXiv:2008.08461, 2020.
  58. 58.Neal, R. M. Annealed importance sampling. Statistics and computing, 11:125–139, 2001.
  59. 59.Neal, R. M. et al. MCMC using Hamiltonian dynamics. Handbook of Markov Chain Monte Carlo, 2(11):2, 2011.
  60. 60.Neese, F. The orca program system. WIRES Comput. Molec. Sci., 2(1):73–78, 2012. doi: 10.1002/wcms.81.
  61. 61.Noe, F., Olsson, S., Kohler, J., and Wu, H. Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning. Science, 365(6457):eaaw1147, 2019.
  62. 62.Pavon, M. Stochastic control and nonequilibrium thermodynamical systems. Applied Mathematics and Optimization, 19:187–202, 1989.
  63. 63.Peluchetti, S. Diffusion bridge mixture transports, schrodinger bridge problems and generative modeling. Journal of Machine Learning Research, 24(374):1–51, 2023a.
  64. 64.Peluchetti, S. Non-denoising forward-time diffusions. arXiv preprint arXiv:2312.14589, 2023b.
  65. 65.Phillips, A., Dau, H.-D., Hutchinson, M. J., De Bortoli, V., Deligiannidis, G., and Doucet, A. Particle denoising diffusion sampler. arXiv preprint arXiv:2402.06320, 2024.
  66. 66.Pracht, P., Bohle, F., and Grimme, S. Automated exploration of the low-energy chemical space with fast quantum chemical methods. Physical Chemistry Chemical Physics, 22(14):7169–7192, 2020.
  67. 67.Pracht, P., Grimme, S., Bannwarth, C., Bohle, F., Ehlert, S., Feldmann, G., Gorges, J., Muller, M., Neudecker, T., Plett, C., et al. Crest—a program for the exploration of low-energy molecular chemical space. The Journal of Chemical Physics, 160(11), 2024.
  68. 68.Protter, P. E. and Protter, P. E. Stochastic differential equations. Springer, 2005.
  69. 69.Rezende, D. and Mohamed, S. Variational inference with normalizing flows. In Proceedings of the 32nd International Conference on Machine Learning, 2015.
  70. 70.Richter, L. and Berner, J. Improved sampling via learned diffusions. arXiv preprint arXiv:2307.01198, 2023.
  71. 71.Richter, L. and Berner, J. Improved sampling via learned diffusions. In The Twelfth International Conference on Learning Representations, 2024.
  72. 72.Riniker, S. and Landrum, G. A. Better informed distance geometry: using what we know to improve conformation generation. Journal of chemical information and modeling, 55(12):2562–2574, 2015.
  73. 73.Satorras, V. G., Hoogeboom, E., and Welling, M. E (n) equivariant graph neural networks. In International conference on machine learning, pp. 9323–9332. PMLR, 2021.
  74. 74.Sethi, S. Optimal Control Theory: Applications to Management Science and Economics. Springer International Publishing, 2018.
  75. 75.Shi, Y., De Bortoli, V., Campbell, A., and Doucet, A. Diffusion schrodinger bridge matching. Advances in Neural Information Processing Systems, 36, 2023.
  76. 76.Somnath, V. R., Pariset, M., Hsieh, Y.-P., Martinez, M. R., Krause, A., and Bunne, C. Aligned diffusion schrodinger bridges. In Uncertainty in Artificial Intelligence, pp. 1985–1995. PMLR, 2023.
  77. 77.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. arXiv preprint arXiv:1907.05600, 2019.
  78. 78.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR 2021), 2021.
  79. 79.Thomas, N., Smidt, T., Kearnes, S., Yang, L., Li, L., Kohlhoff, K., and Riley, P. Tensor field networks: Rotation- and translation-equivariant neural networks for 3d point clouds, 2018. URL https://arxiv.org/abs/1802.08219.
  80. 80.Thornton, J., Hutchinson, M., Mathieu, E., De Bortoli, V., Teh, Y. W., and Doucet, A. Riemannian diffusion schr\” odinger bridge. arXiv preprint arXiv:2207.03024, 2022.
  81. 81.Tzen, B. and Raginsky, M. Neural stochastic differential equations: Deep latent Gaussian models in the diffusion limit. arXiv:1905.09883, 2019a.
  82. 82.Tzen, B. and Raginsky, M. Theoretical guarantees for sampling and inference in generative models with latent diffusions. arXiv:1903.01608, 2019b.
  83. 83.Vargas, F., Grathwohl, W. S., and Doucet, A. Denoising diffusion samplers. In The Eleventh International Conference on Learning Representations, 2023.
  84. 84.Vargas, F., Padhy, S., Blessing, D., and Nusken, N. Transport meets variational inference: Controlled monte carlo diffusions. In The Twelfth International Conference on Learning Representations: ICLR 2024, 2024.
  85. 85.Veber, D. F., Johnson, S. R., Cheng, H.-Y., Smith, B. R., Ward, K. W., and Kopple, K. D. Molecular properties that influence the oral bioavailability of drug candidates. Journal of medicinal chemistry, 45(12):2615–2623, 2002.
  86. 86.Weiler, M., Geiger, M., Welling, M., Boomsma, W., and Cohen, T. 3d steerable cnns: Learning rotationally equivariant features in volumetric data, 2018. URL https://arxiv.org/abs/1807.02547.
  87. 87.Wu, H., Kohler, J., and Noe, F. Stochastic normalizing flows. Advances in Neural Information Processing Systems, 33:5933–5944, 2020.
  88. 88.Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022.
  89. 89.Zhang, D., Chen, R. T. Q., Liu, C.-H., Courville, A., and Bengio, Y. Diffusion generative flow samplers: Improving learning signals through partial trajectory optimization. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=OIsahq1UYC.
  90. 90.Zhang, Q. and Chen, Y. Path integral sampler: A stochastic control approach for sampling. In International Conference on Learning Representations, 2022.

Citation

MLA
Havens, A., et al. “Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching”. arXiv, 2025, https://doi.org/10.48550/arxiv.2504.11713.
APA
Havens, A., Miller, B. K., Yan, B., Domingo-Enrich, C., Sriram, A., Wood, B., Levine, D., Hu, B., Amos, B., Karrer, B., Fu, X., Liu, G.-H., & Chen, R. T. Q. (2025). Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching. arXiv. https://doi.org/10.48550/arxiv.2504.11713
Chicago
Havens, A., B. K. Miller, B. Yan, et al. 2025. “Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2504.11713.
Harvard
Havens, A. et al. (2025) “Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching”. arXiv. Available at: https://doi.org/10.48550/arxiv.2504.11713.
Vancouver
1. Havens A, Miller BK, Yan B, et al (2025) Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching. https://doi.org/10.48550/arxiv.2504.11713

BibTeX

@misc{https://doi.org/10.48550/arxiv.2504.11713,
  doi = {10.48550/ARXIV.2504.11713},
  url = {https://arxiv.org/abs/2504.11713},
  author = {Havens, Aaron and Miller, Benjamin Kurt and Yan, Bing and Domingo-Enrich, Carles and Sriram, Anuroop and Wood, Brandon and Levine, Daniel and Hu, Bin and Amos, Brandon and Karrer, Brian and Fu, Xiang and Liu, Guan-Horng and Chen, Ricky T. Q.},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/