Diffusion-based Molecule Generation with Informative Prior Bridges

Lemeng WuChengyue GongXingchao LiuMao YeQiang Liu

article2022NeurIPS151 citations

Develops physically informed diffusion bridges via Lyapunov functions to embed physical and statistical priors into generative models, producing significantly more stable 3D molecular structures and uniform point clouds.

Listen

Generating realistic 3D structures through artificial intelligence is critical for fields like drug discovery, antibody design, and computer vision. While diffusion-based generative models have achieved notable success, standard methods corrupt and reconstruct data using uninformative noise trajectories that ignore underlying physical laws and domain geometry. Historically, researchers have attempted to enforce physical realism by constraining neural network architectures, but this often restricts model flexibility and fails to prevent irregular structural outputs. Addressing these limitations is essential to produce chemically viable molecules and regularly distributed 3D point clouds without prohibitive computational overhead.

The article develops and evaluates a framework that injects physical and statistical prior information directly into the training trajectory of diffusion models via informative diffusion bridges. Rather than redesigning the neural architecture, the approach steers training using physically guided stochastic processes that are mathematically guaranteed to reach the target data point at a fixed end time.

To construct these valid endpoints, the authors establish a mathematical criterion based on Lyapunov functions, allowing flexible physical forces to be combined with Brownian bridge dynamics without violating end-state guarantees. For molecule generation, the framework incorporates either molecular force-field potentials inspired by AMBER or data-driven statistical energies measuring bond lengths and angles across nearest neighbors. For 3D point cloud generation, the authors apply uniformity-promoting forces, such as Riesz and nearest-neighbor distance energies, to ensure generated points distribute smoothly across object surfaces.

The key findings demonstrate significant performance gains across multiple domains. On small-molecule generation benchmarks using the QM9 dataset, the proposed bridge model combined with statistical forces increased molecular stability from 82.0% to 84.6% and atom stability from 98.7% to 98.8% compared to standard equivariant diffusion models, while also improving chemical novelty from 65.7% to 68.8%. On the larger GEOM-DRUG dataset, atom stability rose from 81.3% to 82.4%, confirming effectiveness on larger molecular systems. In efficiency evaluations, the model retained robust performance when the number of diffusion sampling steps was substantially reduced—achieving 69.2% molecular stability at 50 steps and 83.7% at 500 steps compared to baseline scores of 66.4% and 81.2%. For 3D point cloud synthesis on the ShapeNet airplane and chair benchmarks, incorporating nearest-neighbor statistical priors generated noticeably more uniform surface shapes and achieved competitive 100-step performance quality in as few as 10 diffusion steps, while adding only an 8% training and 3% inference computational overhead.

These results imply that guiding the diffusion training trajectory directly with domain physics is more effective and versatile than relying exclusively on specialized network architectures. By producing higher-quality and more stable structures with fewer sampling steps, this strategy lowers computational costs, shortens generation timelines, and reduces downstream failure rates in applications like molecular screening and 3D surface meshing.

Organizations developing molecular generative pipelines or 3D geometry applications should consider integrating prior bridge drift terms into existing diffusion training routines. Future technical work should focus on extending the framework to very large macromolecular systems like proteins, incorporating complex torsional angle dynamics that are currently omitted due to dynamic bonding verification limits, and resolving training bottlenecks that appear when scaling to large batch sizes.

arXiv: 2209.00865
  • Paper: Equivariant Diffusion for Molecule Generation in 3D, Emiel Hoogeboom et al. (2022). This foundational work establishes equivariant 3D molecular generation via continuous-discrete diffusion, providing the primary baseline and geometric representation framework modified by informative prior bridges.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Understanding the fundamental formulation and denoising objectives of Denoising Diffusion Probabilistic Models (DDPM) is essential to following how standard forward-reverse dynamics are generalized to diffusion bridges.
  • Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). This paper establishes the core metrics and generative principles for 3D point cloud synthesis, which serves as a primary benchmark domain for evaluating uniformity-promoted prior bridges.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Reading this paper provides necessary context on non-Markovian forward and reverse trajectories that underpin modern fast sampling and bridge-like formulations in diffusion models.
  • Paper: SchNet: A continuous-filter convolutional neural network for modeling quantum interactions, Kristof Schütt et al. (2017). This work introduces continuous-filter convolutions that enforce physical invariants on continuous 3D atomic coordinates, supplying essential background for physically grounded molecular modeling.
Cover for Diffusion-based Molecule Generation with Informative Prior Bridges

Abstract

AI-based molecule generation provides a promising approach to a large area of biomedical sciences and engineering, such as antibody design, hydrolase engineering, or vaccine development. Because the molecules are governed by physical laws, a key challenge is to incorporate prior information into the training procedure to generate high-quality and realistic molecules. We propose a simple and novel approach to steer the training of diffusion-based generative models with physical and statistics prior information. This is achieved by constructing physically informed diffusion bridges, stochastic processes that guarantee to yield a given observation at the fixed terminal time. We develop a Lyapunov function based method to construct and determine bridges, and propose a number of proposals of informative prior bridges for both high-quality molecule generation and uniformity-promoted 3D point cloud generation. With comprehensive experiments, we show that our method provides a powerful approach to the 3D generation task, yielding molecule structures with better quality and stability scores and more uniformly distributed point clouds of high qualities.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Method
  • 3.1 Learning Diffusion Generative Models with Prior Bridges
  • 3.2 Designing Informative Prior Bridges
  • 4 Molecule and 3D Generation with Informative Prior Bridges
  • 4.1 Prior Bridges for Molecule Generation
  • 4.2 Prior Bridges for Point Cloud Generation
  • 5 Experiment
  • 5.1 Force Guided Molecule Generation
  • 5.2 Force Guided Point Cloud Generation
  • 6 Conclusion and Limitations
  • References
  • A Proofs
  • B Model Details
  • B.1 Model Architecture for Molecule Generation.
  • B.2 Model Architecture for Point Cloud Generation.
  • C More Visualization for Point Cloud Generation

Knowls

  1. Knowl 1 — Prior-Guided Diffusion Generative Model Training via Diffusion Bridges

    model/method

    A generative modeling framework on Rd\mathbb{R}^d over time t∈[0,1]t \in [0, 1] trains a parameterized forward diffusion model PθP^\theta:

    dZt=stθ(Zt)dt+σt(Zt)dWt,t∈[0,1],Z0∼μ0dZ_t = s_t^\theta(Z_t)dt + \sigma_t(Z_t)dW_t, \quad t \in [0, 1], \quad Z_0 \sim \mu_0

    to match the target data distribution Π∗\Pi^* at terminal time t=1t = 1 (i.e., P1θ=Π∗P_1^\theta = \Pi^*), where WtW_t is standard Brownian motion, σt:Rd→Rd×d\sigma_t: \mathbb{R}^d \to \mathbb{R}^{d \times d} is a positive definite diffusion covariance coefficient, and μ0\mu_0 is an initial distribution.

    To guide the trajectory generation with physical or problem-dependent priors, an imputation bridge process QxQ^x pinned at terminal data point xx (Qx(Z1=x)=1Q^x(Z_1 = x) = 1) is defined for each x∈Rdx \in \mathbb{R}^d:

    Qx:dZt=bt(Zt∣x)dt+σt(Zt)dWt,Z0∼μ0Q^x: dZ_t = b_t(Z_t \mid x)dt + \sigma_t(Z_t)dW_t, \quad Z_0 \sim \mu_0

    where bt(Zt∣x)b_t(Z_t \mid x) incorporates domain prior forces. The mixture over data distribution Π∗\Pi^* is QΠ∗:=∫Qx(⋅)Π∗(dx)Q^{\Pi^*} := \int Q^x(\cdot)\Pi^*(dx), ensuring Q1Π∗=Π∗Q_1^{\Pi^*} = \Pi^*.

    Training minimizes the Kullback-Leibler divergence min⁡θKL(QΠ∗∥Pθ)\min_\theta \mathrm{KL}(Q^{\Pi^*} \parallel P^\theta), which via Girsanov theorem reduces to the denoised score matching loss:

    L(θ)=EZ∼QΠ∗[12∫01∥σt(Zt)−1(stθ(Zt)−bt(Zt∣Z1))∥22dt]+const\mathcal{L}(\theta) = \mathbb{E}_{Z \sim Q^{\Pi^*}} \left[ \frac{1}{2} \int_0^1 \left\| \sigma_t(Z_t)^{-1} \left(s_t^\theta(Z_t) - b_t(Z_t \mid Z_1)\right) \right\|_2^2 dt \right] + \text{const}

    The global optimum θ∗\theta^* satisfies stθ∗(z)=EZ∼QΠ∗[bt(z∣Z1)∣Zt=z]s_t^{\theta^*}(z) = \mathbb{E}_{Z \sim Q^{\Pi^*}}[b_t(z \mid Z_1) \mid Z_t = z]. To incorporate a domain force field ft(z)=−∇E(z)f_t(z) = -\nabla E(z) derived from an energy function E(z)E(z), the learned drift is parameterized as stθ(z)=αft(z)+s~tθ(z)s_t^\theta(z) = \alpha f_t(z) + \tilde{s}_t^\theta(z), where s~tθ\tilde{s}_t^\theta is parameterized by a neural network and α\alpha is a learnable or preset scalar weighting parameter.

  2. Knowl 2 — Endpoint Preservation of Perturbed Brownian Bridges

    theoretical result

    Let Qx,bbQ^{x,\mathrm{bb}} denote a time-scaled Brownian bridge pinned at endpoint x∈Rdx \in \mathbb{R}^d on t∈[0,1]t \in [0, 1]:

    dZt=σt2x−Ztβ1−βtdt+σtdWt,Z0∼μ0dZ_t = \sigma_t^2 \frac{x - Z_t}{\beta_1 - \beta_t} dt + \sigma_t dW_t, \quad Z_0 \sim \mu_0

    where σt>0\sigma_t > 0, WtW_t is standard Brownian motion, and βt=∫0tσs2ds\beta_t = \int_0^t \sigma_s^2 ds. Let Qx,bb,fQ^{x,\mathrm{bb},f} be a perturbed diffusion process with an additional prior force drift ftf_t:

    dZt=(σtft(Zt)+σt2x−Ztβ1−βt)dt+σtdWt,Z0∼μ0dZ_t = \left( \sigma_t f_t(Z_t) + \sigma_t^2 \frac{x - Z_t}{\beta_1 - \beta_t} \right) dt + \sigma_t dW_t, \quad Z_0 \sim \mu_0

    If the prior force satisfies EQx,bb,f[∥ft(Zt)∥22]<+∞\mathbb{E}_{Q^{x,\mathrm{bb},f}}[\|f_t(Z_t)\|_2^2] < +\infty for all t∈[0,1]t \in [0, 1] and σt>0\sigma_t > 0, then Qx,bb,fQ^{x,\mathrm{bb},f} and Qx,bbQ^{x,\mathrm{bb}} are mutually absolutely continuous with finite Kullback-Leibler divergence:

    KL(Qx,bb∥Qx,bb,f)=12EQx,bb[∫01∥ft(Zt)∥22dt]<+∞\mathrm{KL}(Q^{x,\mathrm{bb}} \parallel Q^{x,\mathrm{bb},f}) = \frac{1}{2} \mathbb{E}_{Q^{x,\mathrm{bb}}} \left[ \int_0^1 \|f_t(Z_t)\|_2^2 dt \right] < +\infty

    Consequently, Qx,bb,fQ^{x,\mathrm{bb},f} shares the exact support of Qx,bbQ^{x,\mathrm{bb}}, guaranteeing the bridge endpoint condition Qx,bb,f(Z1=x)=1Q^{x,\mathrm{bb},f}(Z_1 = x) = 1 almost surely.

  3. Knowl 3 — Lyapunov Characterization for Diffusion Bridge Construction

    theoretical result

    Let A⊂RdA \subset \mathbb{R}^d be a measurable set. A function Ut(z)=U(z,t)U_t(z) = U(z, t) with U(⋅,t)∈C2(Rd)U(\cdot, t) \in C^2(\mathbb{R}^d) and U(z,⋅)∈C1([0,1])U(z, \cdot) \in C^1([0, 1]) is a Lyapunov function for AA at t=1t = 1 if U1(z)≥0U_1(z) \ge 0 for all z∈Rdz \in \mathbb{R}^d and U1(z)=0  ⟺  z∈AU_1(z) = 0 \iff z \in A.

    Consider the diffusion process QAQ^A:

    dZt=(−αt∇zUt(Zt)+νt(Zt))dt+σt(Zt)dWt,t∈[0,1],Z0∼μ0dZ_t = \left(-\alpha_t \nabla_z U_t(Z_t) + \nu_t(Z_t)\right) dt + \sigma_t(Z_t) dW_t, \quad t \in [0, 1], \quad Z_0 \sim \mu_0

    where αt>0\alpha_t > 0 is a time-varying step size, νt\nu_t is a perturbation drift, and σt\sigma_t is the diffusion coefficient. QAQ^A is a valid bridge to AA (i.e., QA(Z1∈A)=1Q^A(Z_1 \in A) = 1) if the following conditions hold:

    1. UU satisfies an expected Polyak-Łojasiewicz condition: EQA[Ut(Zt)]−EQA[∥∇zUt(Zt)∥22]≤0\mathbb{E}_{Q^A}[U_t(Z_t)] - \mathbb{E}_{Q^A}[\|\nabla_z U_t(Z_t)\|_2^2] \le 0 for all tt.

    2. Defining βt=EQA[∇zUt(Zt)⊤νt(Zt)]\beta_t = \mathbb{E}_{Q^A}[\nabla_z U_t(Z_t)^\top \nu_t(Z_t)], γt=EQA[∂tUt(Zt)+12tr(∇z2Ut(Zt)σt2(Zt))]\gamma_t = \mathbb{E}_{Q^A}[\partial_t U_t(Z_t) + \frac{1}{2}\mathrm{tr}(\nabla_z^2 U_t(Z_t) \sigma_t^2(Z_t))], and integrating factor ζt=exp⁡(∫0tαsds)\zeta_t = \exp\left(\int_0^t \alpha_s ds\right), the coefficients satisfy:

    lim⁡t↑1ζt=+∞andlim⁡t↑1ζt∫0tζs(βs+γs)ds=+∞\lim_{t \uparrow 1} \zeta_t = +\infty \quad \text{and} \quad \lim_{t \uparrow 1} \frac{\zeta_t}{\int_0^t \zeta_s (\beta_s + \gamma_s) ds} = +\infty

  4. Knowl 4 — Statistical Geometric Prior Energy for Molecule Generation

    model/method

    For a molecule represented by 3D coordinates xr=[x1r,…,xmr]∈Rm×3x^r = [x_1^r, \dots, x_m^r] \in \mathbb{R}^{m \times 3} and continuous type vectors xh=[x1h,…,xmh]∈Rm×kx^h = [x_1^h, \dots, x_m^h] \in \mathbb{R}^{m \times k} where x^ih=argmax⁡(xih)∈{1,…,k}\hat{x}_i^h = \operatorname{argmax}(x_i^h) \in \{1, \dots, k\} is the assigned discrete atom category, empirical bond length and bond angle distributions learned from a training set define the statistical prior energy Estat(x)E_{\text{stat}}(x):

    Estat(x)=∑ij∈knn(x)1σ^x^ihx^jh2(∥xir−xjr∥2−μ^x^ihx^jh)2+∑ij,jk∈knn(x)1σ^x^ihx^jhx^kh2(Ang(xijkr)−μ^x^ihx^jhx^kh)2E_{\text{stat}}(x) = \sum_{ij \in \mathrm{knn}(x)} \frac{1}{\hat{\sigma}_{\hat{x}_i^h \hat{x}_j^h}^2} \left( \|x_i^r - x_j^r\|_2 - \hat{\mu}_{\hat{x}_i^h \hat{x}_j^h} \right)^2 + \sum_{ij, jk \in \mathrm{knn}(x)} \frac{1}{\hat{\sigma}_{\hat{x}_i^h \hat{x}_j^h \hat{x}_k^h}^2} \left( \mathrm{Ang}(x_{ijk}^r) - \hat{\mu}_{\hat{x}_i^h \hat{x}_j^h \hat{x}_k^h} \right)^2

    where knn(x)\mathrm{knn}(x) denotes the KK-nearest neighbor graph over Euclidean distances in xrx^r; Ang(xijkr)\mathrm{Ang}(x_{ijk}^r) is the angle formed between vectors xir−xjrx_i^r - x_j^r and xkr−xjrx_k^r - x_j^r; μ^rc\hat{\mu}_{rc} and σ^rc2\hat{\sigma}_{rc}^2 are empirical means and variances of bond lengths between atom types rr and cc; and μ^rcr′\hat{\mu}_{rcr'} and σ^rcr′2\hat{\sigma}_{rcr'}^2 are empirical means and variances of angles for triplet atom types r,c,r′r, c, r'.

    The resulting prior force injected into the diffusion bridge drift is ft(x)=−∇Estat(x)f_t(x) = -\nabla E_{\text{stat}}(x).

  5. Knowl 5 — AMBER-Inspired Physical Energy for Molecule Generation

    model/method

    To inject physical force fields into diffusion bridge processes for 3D molecules with atom coordinates xr∈Rm×3x^r \in \mathbb{R}^{m \times 3} and discrete atom types x^h∈{1,…,k}m\hat{x}^h \in \{1, \dots, k\}^m, a composite potential energy E(x)E(x) is defined:

    E(x)=Ebond(x)+Eangle(x)+ELJ(x)+ECoulomb(x)E(x) = E_{\text{bond}}(x) + E_{\text{angle}}(x) + E_{\text{LJ}}(x) + E_{\text{Coulomb}}(x)

    1. Bond energy: Ebond(x)=∑ij∈bond(x)(∥xir−xjr∥2−ℓ0(x^ih,x^jh))2E_{\text{bond}}(x) = \sum_{ij \in \mathrm{bond}(x)} (\|x_i^r - x_j^r\|_2 - \ell^0(\hat{x}_i^h, \hat{x}_j^h))^2, where bond(x)\mathrm{bond}(x) includes pairs with distance <1.15< 1.15 times the covalent radius, and ℓ0(r,c)\ell^0(r, c) is the dataset expected bond length between types rr and cc.

    2. Angle energy: Eangle(x)=∑ijk∈angle(x)(Ang(xijkr)−ω0(x^ih,x^jh,x^kh))2E_{\text{angle}}(x) = \sum_{ijk \in \mathrm{angle}(x)} (\mathrm{Ang}(x_{ijk}^r) - \omega^0(\hat{x}_i^h, \hat{x}_j^h, \hat{x}_k^h))^2, where angle(x)\mathrm{angle}(x) denotes adjacent bond pairs, Ang(xijkr)\mathrm{Ang}(x_{ijk}^r) is the angle formed by vectors xir−xjrx_i^r - x_j^r and xkr−xjrx_k^r - x_j^r, and ω0\omega^0 is the dataset expected bond angle.

    3. Lennard-Jones (LJ) energy: ELJ(x)=∑i≠je(∥xir−xjr∥2)E_{\text{LJ}}(x) = \sum_{i \ne j} e(\|x_i^r - x_j^r\|_2) with e(ℓ)=(σ/ℓ)12−2(σ/ℓ)6e(\ell) = (\sigma / \ell)^{12} - 2(\sigma / \ell)^6, where σ\sigma represents average nucleus distance.

    4. Coulomb potential energy: ECoulomb(x)=κ∑i≠jq(x^ih)q(x^jh)∥xir−xjr∥2E_{\text{Coulomb}}(x) = \kappa \sum_{i \ne j} \frac{q(\hat{x}_i^h) q(\hat{x}_j^h)}{\|x_i^r - x_j^r\|_2}, where κ\kappa is Coulomb's constant and q(r)q(r) is the atomic point charge of type rr.

  6. Knowl 6 — Uniformity-Promoting Prior Energies for 3D Point Cloud Diffusion

    model/method

    For 3D point cloud generation where each point cloud is an un-typed coordinate set xr=[x1r,…,xmr]∈Rm×3x^r = [x_1^r, \dots, x_m^r] \in \mathbb{R}^{m \times 3}, spatial uniformity across the surface is encouraged during diffusion bridge training via two potential energy formulations:

    1. Riesz Energy: Introduces pairwise repulsive forces:

    ERiesz(x)=12∑i=1m∑j≠i∥xir−xjr∥2−2E_{\text{Riesz}}(x) = \frac{1}{2} \sum_{i=1}^m \sum_{j \ne i} \|x_i^r - x_j^r\|_2^{-2}

    1. KNN Distance Energy: Penalizes deviations from the empirical dataset mean distance to the KK-nearest neighbors:

    Eknn(x)=∑i=1m(knn-disti(xr)−μknn)2E_{\text{knn}}(x) = \sum_{i=1}^m \left( \mathrm{knn\text{-}dist}_i(x^r) - \mu_{\text{knn}} \right)^2

    where knn-disti(xr)=1K∑j∈NK(xir;xr)∥xir−xjr∥2\mathrm{knn\text{-}dist}_i(x^r) = \frac{1}{K} \sum_{j \in N_K(x_i^r; x^r)} \|x_i^r - x_j^r\|_2, NK(xir;xr)N_K(x_i^r; x^r) is the index set of the KK-nearest neighbors of xirx_i^r, and μknn\mu_{\text{knn}} is the empirical average KK-NN distance computed across the training dataset (K=4K = 4 for 2D surface meshes).

    The prior force added to the drift of the bridge process is ft(x)=−∇E(x)f_t(x) = -\nabla E(x).

  7. Knowl 7 — Prior-Guided Diffusion Bridge Training and Sampling Algorithm

    algorithm

    The algorithm trains a generative diffusion model PθP^\theta with drift stθ(z)=αft(z)+s~tθ(z)s_t^\theta(z) = \alpha f_t(z) + \tilde{s}_t^\theta(z) using informative bridge processes QxQ^x, and simulates PθP^\theta to draw new samples.

    Input: Training dataset D={x(k)}k=1nD = \{x^{(k)}\}_{k=1}^n, bridge noise schedule σt\sigma_t, prior force ft(z)=−∇E(z)f_t(z) = -\nabla E(z), number of discretization time steps TT, neural network drift model s~tθ\tilde{s}_t^\theta.
    Training:
      repeat
        Sample data point x∼Dx \sim D
        Sample random time step t∈[0,1]t \in [0, 1]
        Simulate bridge state Zt∼Qx,bb,fZ_t \sim Q^{x,\mathrm{bb},f} initialized from Z0∼μ0Z_0 \sim \mu_0 via Euler-Maruyama
        Compute target bridge drift bt(Zt∣x)=σtft(Zt)+σt2x−Ztβ1−βtb_t(Z_t \mid x) = \sigma_t f_t(Z_t) + \sigma_t^2 \frac{x - Z_t}{\beta_1 - \beta_t}
        Compute loss L(θ)=12∥σt−1(stθ(Zt)−bt(Zt∣x))∥22\mathcal{L}(\theta) = \frac{1}{2} \|\sigma_t^{-1} (s_t^\theta(Z_t) - b_t(Z_t \mid x))\|_2^2
        Update parameters θ\theta (and optionally weighting factor α\alpha) via stochastic gradient descent
      until convergence
    Sampling:
      Sample Z0∼μ0Z_0 \sim \mu_0
      for time step i=0i = 0 to T−1T - 1 do
        Let t=i/Tt = i / T and Δt=1/T\Delta t = 1 / T
        Sample ϵt∼N(0,I)\epsilon_t \sim \mathcal{N}(0, I)
        Compute drift stθ(Zt)=αft(Zt)+s~tθ(Zt)s_t^\theta(Z_t) = \alpha f_t(Z_t) + \tilde{s}_t^\theta(Z_t)
        Update Zt+Δt=Zt+stθ(Zt)Δt+σtΔt ϵtZ_{t + \Delta t} = Z_t + s_t^\theta(Z_t) \Delta t + \sigma_t \sqrt{\Delta t} \, \epsilon_t
      end for
      Round categorical features in Z1Z_1 to discrete types
    Output: Generated sample Z1Z_1
  8. Knowl 8 — Molecule Generation Quality on QM9 and GEOM-DRUG Benchmarks

    data/table

    Molecules generated by the proposed Bridge with Statistical Force (K=5K=5) were evaluated against baseline methods on QM9 (small molecules with up to 9 heavy atoms) and GEOM-DRUG (larger drug conformations with 44 atoms on average). Metrics evaluate atom stability (percentage of atoms with valid valencies), molecular stability (percentage of molecules with all atoms stable), RDKit-based novelty (percentage of valid molecules not in training set), and valid + unique rate.

    QM9 GEOM-DRUG
    Method Atom Sta (%) ↑\uparrow Mol Sta (%) ↑\uparrow Novelty (%) ↑\uparrow Valid + Unique ↑\uparrow Atom Sta (%) ↑\uparrow Mol Sta (%) ↑\uparrow
    EN-Flow 85.0 4.9 81.4 0.349 75.0 0.0
    GDM 97.0 63.2 74.6 - 75.0 0.0
    E-GDM 98.7 ±\pm 0.1 82.0 ±\pm 0.4 65.7 ±\pm 0.2 0.902 81.3 0.0
    Bridge 98.7 ±\pm 0.1 81.8 ±\pm 0.2 66.0 ±\pm 0.2 0.902 81.0 ±\pm 0.7 0.0
    Bridge + Force 98.8 ±\pm 0.1 84.6 ±\pm 0.3 68.8 ±\pm 0.2 0.907 82.4 ±\pm 0.8 0.0

    Adding the statistical prior force to the bridge improves QM9 molecular stability from 82.0% to 84.6% and novelty from 65.7% to 68.8% over E-GDM, while increasing training and inference computational cost by 8% and 3%, respectively.

  9. Knowl 9 — Point Cloud Generation Quality on ShapeNet

    data/table

    3D point clouds (2048 points sampled per shape) were generated for Chair and Airplane categories from ShapeNet and evaluated under 10 and 100 diffusion time steps. Metrics include Minimum Matching Distance (MMD, lower is better; Chamfer Distance [CD] ×103\times 10^3, Earth Mover's Distance [EMD] ×10\times 10) and Coverage (COV, higher is better; %\%).

    10 Steps 100 Steps
    MMD ↓\downarrow COV ↑\uparrow MMD ↓\downarrow COV ↑\uparrow
    Category Method CD EMD CD EMD CD EMD CD EMD
    Chair Diffusion 14.01 3.23 32.72 29.36 12.32 1.79 47.41 47.59
    Bridge 13.04 2.14 46.01 42.59 12.47 1.85 47.83 47.13
    + Riesz 12.84 1.95 47.21 44.31 12.31 1.82 48.14 47.42
    + Statistic 12.65 1.84 47.58 45.23 12.25 1.78 48.39 47.56
    Airplane Diffusion 3.71 1.31 43.12 39.94 3.28 1.04 48.74 46.38
    Bridge 3.44 1.24 46.90 43.46 3.37 1.08 47.11 46.17
    + Riesz 3.39 1.20 47.11 43.12 3.24 1.09 48.62 46.23
    + Statistic 3.30 1.12 47.02 44.67 3.24 1.06 48.53 46.73

    At 10 sampling steps, adding statistical or Riesz forces to the bridge produces sample quality and coverage comparable to 100-step standard diffusion models. Statistical KNN energy outperforms Riesz energy, which occasionally causes outlier artifacts due to pure repulsion.

  10. Knowl 10 — Sampling Step Efficiency and Energy Ablations on QM9

    data/table

    Evaluating generation fidelity with reduced discretization steps and different physical energy ablations on the QM9 dataset demonstrates model robustness and individual component contributions.

    Sampling Discretization Steps:

    50 Steps 100 Steps 500 Steps
    Method Atom Sta (%) Mol Sta (%) Atom Sta (%) Mol Sta (%) Atom Sta (%) Mol Sta (%)
    EGM 97.0 ±\pm 0.1 66.4 ±\pm 0.2 97.3 ±\pm 0.1 69.8 ±\pm 0.2 98.5 ±\pm 0.1 81.2 ±\pm 0.1
    Bridge + Force 97.3 ±\pm 0.1 69.2 ±\pm 0.2 97.9 ±\pm 0.1 72.3 ±\pm 0.2 98.7 ±\pm 0.1 83.7 ±\pm 0.1

    Energy Function Variations (1000 steps):

    Method Atom Sta (%) Mol Sta (%) Method Atom Sta (%) Mol Sta (%)
    Force (Stat), k=7k=7 98.8 ±\pm 0.1 84.5 ±\pm 0.2 Force (AMBER) 98.7 ±\pm 0.1 83.1 ±\pm 0.2
    Force (Stat), k=5k=5 98.8 ±\pm 0.1 84.6 ±\pm 0.3 Force (AMBER) w/o bond 98.7 ±\pm 0.1 82.5 ±\pm 0.1
    Force (Stat), k=3k=3 98.8 ±\pm 0.1 83.9 ±\pm 0.3 Force (AMBER) w/o angle 98.7 ±\pm 0.1 82.4 ±\pm 0.2
    Force (Stat), k=1k=1 98.8 ±\pm 0.1 82.7 ±\pm 0.3 Force (AMBER) w/o LJ/Coulomb 98.7 ±\pm 0.1 82.7 ±\pm 0.2

    Bridge + Force at 500 steps outperforms 1000-step EGM (83.7% vs. 82.0% Mol Sta). For the statistical force, performance saturates at k=5k=5, while removing any component of the AMBER potential degrades stability.

  11. Knowl 11 — Limitations in Torsional Potential Modeling and Large-Batch Training

    limitation

    The proposed prior bridge framework exhibits two identified limitations:

    1. Omission of Torsional Potentials: Neither the physical (AMBER-inspired) nor statistical energy formulations include dihedral/torsional angle potentials because dynamically verifying covalent bond connectivity among 4-atom sequences during intermediate noisy diffusion steps is computationally intractable and ambiguous.

    2. Training Scalability Bottlenecks: Similar to score-based diffusion baselines, training deep diffusion bridge networks requires long wall-clock times (e.g., approximately 10 days on a Tesla V100 GPU for QM9 and GEOM-DRUG). Attempts to accelerate training via larger batch sizes (such as 512 or 1024) caused empirical performance degradation.

Coverage note — None was omitted; all major theoretical results, methodological definitions, algorithms, empirical benchmarks, and stated limitations are included.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning repre-sentations and generative models for 3d point clouds. In International conference on machine learning, pages 40–49. PMLR, 2018.
  2. 2.Miguel Alcalde, Manuel Ferrer, Francisco J Plou, and Antonio Ballesteros. Environmental bio-catalysis: from remediation with enzymes to novel green processes. TRENDS in Biotechnology, 24(6):281–287, 2006.
  3. 3.Namrata Anand, Raphael Eguchi, Irimpan I Mathews, Carla P Perez, Alexander Derry, Russ B Altman, and Po-Ssu Huang. Protein sequence design with a learned potential. Nature communi-cations, 13(1):1–11, 2022.
  4. 4.Simon Axelrod and Rafael Gómez-Bombarelli. Geom, energy-annotated molecular conforma-tions for property prediction and molecular generation. Scientific Data, 9(1):185, 2022.
  5. 5.Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge Belongie, Noah Snavely, and Bharath Hariharan. Learning gradient fields for shape generation. In European Conference on Computer Vision, pages 364–381. Springer, 2020.
  6. 6.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015.
  7. 7.Tianrong Chen, Guan-Horng Liu, and Evangelos A Theodorou. Likelihood training of schr" odinger bridge using forward-backward sdes theory. arXiv preprint arXiv:2110.11291, 2021.
  8. 8.Wendy D Cornell, Piotr Cieplak, Christopher I Bayly, Ian R Gould, Kenneth M Merz, David M Ferguson, David C Spellmeyer, Thomas Fox, James W Caldwell, and Peter A Kollman. A second generation force field for the simulation of proteins, nucleic acids, and organic molecules. Journal of the American Chemical Society, 117(19):5179–5197, 1995.
  9. 9.Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet. Diffusion schrödinger bridge with applications to score-based generative modeling. Advances in Neural Information Processing Systems, 34, 2021.
  10. 10.Joseph L Doob and JI Doob. Classical potential theory and its probabilistic counterpart, volume 549. Springer, 1984.
  11. 11.Yuanqi Du, Tianfan Fu, Jimeng Sun, and Shengchao Liu. Molgensurvey: A systematic survey in machine learning models for molecule design. arXiv preprint arXiv:2203.14500, 2022.
  12. 12.Yong Duan, Chun Wu, Shibasish Chowdhury, Mathew C Lee, Guoming Xiong, Wei Zhang, Rong Yang, Piotr Cieplak, Ray Luo, Taisung Lee, et al. A point-charge force field for molecular mechanics simulations of proteins based on condensed-phase quantum mechanical calculations. Journal of computational chemistry, 24(16):1999–2012, 2003.
  13. 13.Niklas Gebauer, Michael Gastegger, and Kristof Schütt. Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 7566–7578. Curran Associates, Inc., 2019.
  14. 14.Dwaraknath Gnaneshwar, Bharath Ramsundar, Dhairya Gandhi, Rachel Kurchin, and Venkata-subramanian Viswanathan. Score-based generative models for molecule generation. arXiv preprint arXiv:2203.04698, 2022.
  15. 15.Chengyue Gong, Lemeng Wu, and Qiang Liu. How to fill the optimum set? population gradient descent with harmless diversity. arXiv preprint arXiv:2202.08376, 2022.
  16. 16.Mario Götz. On the riesz energy of measures. Journal of Approximation Theory, 122(1):62–78, 2003.
  17. 17.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  18. 18.Emiel Hoogeboom, Victor Garcia Satorras, Clément Vignac, and Max Welling. Equivariant diffusion for molecule generation in 3d. arXiv preprint arXiv:2203.17003, 2022.
  19. 19.Bowen Jing, Gabriele Corso, Regina Barzilay, and Tommi S Jaakkola. Torsional diffusion for molecular conformer generation. In ICLR2022 Machine Learning for Drug Discovery, 2022.
  20. 20.William L Jorgensen, David S Maxwell, and Julian Tirado-Rives. Development and testing of the opls all-atom force field on conformational energetics and properties of organic liquids. Journal of the American Chemical Society, 118(45):11225–11236, 1996.
  21. 21.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron-neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Zdek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  22. 22.Greg Landrum. Rdkit documentation. Release, 1(1-79):4, 2013.
  23. 23.Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. Pre-training molecular graph representation with 3d geometry. arXiv preprint arXiv:2110.07728, 2021.
  24. 24.Xingchao Liu, Lemeng Wu, Mao Ye, and Qiang Liu. Let us build bridges: Understanding and extending diffusion generative models, 2022.
  25. 25.Hongyuan Lu, Daniel J Diaz, Natalie J Czarnecki, Congzhi Zhu, Wantae Kim, Raghav Shroff, Daniel J Acosta, Bradley R Alexander, Hannah O Cole, Yan Zhang, et al. Machine learning-aided engineering of hydrolases for pet depolymerization. Nature, 604(7907):662–667, 2022.
  26. 26.Shitong Luo, Jiaqi Guan, Jianzhu Ma, and Jian Peng. A 3d molecule generative model for structure-based drug design. arXiv preprint arXiv:2203.10446, 2022.
  27. 27.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021.
  28. 28.Shitong Luo, Jiahan Li, Jiaqi Guan, Yufeng Su, Chaoran Cheng, Jian Peng, and Jianzhu Ma. Equivariant point cloud analysis via learning orientations for message passing. arXiv preprint arXiv:2203.14486, 2022.
  29. 29.Shitong Luo, Chence Shi, Minkai Xu, and Jian Tang. Predicting molecular conformation via dynamic graph score matching. Advances in Neural Information Processing Systems, 34, 2021.
  30. 30.M Riad Manaa, Laurence E Fried, Carl F Melius, Marcus Elstner, and Th Frauenheim. De-composition of hmx at extreme conditions: A molecular dynamics simulation. The Journal of Physical Chemistry A, 106(39):9024–9029, 2002.
  31. 31.Elman Mansimov, Omar Mahmood, Seokho Kang, and Kyunghyun Cho. Molecular geometry prediction using a deep generative graph neural network. Scientific reports, 9(1):1–13, 2019.
  32. 32.Xuerong Mao. Stochastic differential equations and applications. Elsevier, 2007.
  33. 33.B. Oksendal. Stochastic differential equations: an introduction with applications. Springer Science & Business Media, 6 edition, 2013.
  34. 34.Stefano Peluchetti. Non-denoising forward-time diffusions, 2022.
  35. 35.Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole Von Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific data, 1(1):1–7, 2014.
  36. 36.Victor Garcia Satorras, Emiel Hoogeboom, Fabian B Fuchs, Ingmar Posner, and Max Welling. E (n) equivariant normalizing flows. arXiv preprint arXiv:2105.09016, 2021.
  37. 37.Chence Shi, Shitong Luo, Minkai Xu, and Jian Tang. Learning gradient fields for molecular conformation generation. In International Conference on Machine Learning, 2021.
  38. 38.Gregor NC Simm and José Miguel Hernández-Lobato. A generative model for molecular distance geometry. arXiv preprint arXiv:1909.11459, 2019.
  39. 39.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsuper-vised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  40. 40.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  41. 41.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34, 2021.
  42. 42.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  43. 43.Francisco Vargas, Pierre Thodoroff, Austen Lamacraft, and Neil Lawrence. Solving schrödinger bridges via maximum likelihood. Entropy, 23(9):1134, 2021.
  44. 44.Gefei Wang, Yuling Jiao, Qian Xu, Yang Wang, and Can Yang. Deep generative learning via schrödinger bridge. In International Conference on Machine Learning, pages 10794–10804. PMLR, 2021.
  45. 45.Sheng Wang, Siqi Sun, Zhen Li, Renyu Zhang, and Jinbo Xu. Accurate de novo prediction of protein contact map by ultra-deep learning model. PLoS computational biology, 13(1):e1005324, 2017.
  46. 46.Fang Wu, Qiang Zhang, Xurui Jin, Yinghui Jiang, and Stan Z Li. A score-based geometric model for molecular dynamics simulations. arXiv preprint arXiv:2204.08672, 2022.
  47. 47.Minkai Xu, Shitong Luo, Yoshua Bengio, Jian Peng, and Jian Tang. Learning neural generative dynamics for molecular conformation generation. arXiv preprint arXiv:2102.10240, 2021.
  48. 48.Minkai Xu, Wujie Wang, Shitong Luo, Chence Shi, Yoshua Bengio, Rafael Gomez-Bombarelli, and Jian Tang. An end-to-end framework for molecular conformation generation via bilevel programming. In International Conference on Machine Learning, 2021.
  49. 49.Minkai Xu, Lantao Yu, Yang Song, Chence Shi, Stefano Ermon, and Jian Tang. Geodiff: A geometric diffusion model for molecular conformation generation. In International Conference on Learning Representations, 2022.
  50. 50.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4541–4550, 2019.
  51. 51.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5826–5835, 2021.

Citation

MLA
Wu, L., et al. “Diffusion-based Molecule Generation with Informative Prior Bridges”. arXiv, 2022, http://arxiv.org/abs/2209.00865v1.
APA
Wu, L., Gong, C., Liu, X., Ye, M., & Liu, Q. (2022). Diffusion-based Molecule Generation with Informative Prior Bridges. arXiv. http://arxiv.org/abs/2209.00865v1
Chicago
Wu, L., C. Gong, X. Liu, M. Ye, and Q. Liu. 2022. “Diffusion-based Molecule Generation with Informative Prior Bridges”. arXiv. http://arxiv.org/abs/2209.00865v1.
Harvard
Wu, L. et al. (2022) “Diffusion-based Molecule Generation with Informative Prior Bridges”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.00865v1.
Vancouver
1. Wu L, Gong C, Liu X, Ye M, Liu Q (2022) Diffusion-based Molecule Generation with Informative Prior Bridges. arXiv

BibTeX

@article{wu2022diffusion,
  title = {Diffusion-based Molecule Generation with Informative Prior Bridges},
  author = {Wu, Lemeng and Gong, Chengyue and Liu, Xingchao and Ye, Mao and Liu, Qiang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.00865v1},
  eprint = {2209.00865}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors