MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space

Yanru QuKeyue QiuYuxuan SongJingjing GongJiawei HanMingyue ZhengHao ZhouWei-Ying Ma

article2024ICML56 citations

Proposes MolCRAFT, a structure-based drug design model operating in continuous parameter space with noise-reduced sampling to generate structurally realistic 3D molecular conformations while achieving reference-level binding affinities without mode collapse.

Listen

Structure-based drug design uses three-dimensional biological target structures to generate therapeutic molecule candidates computationally. While modern generative methods propose high-affinity binders, they frequently generate physically unrealistic structures—a problem known as false positives. Sequential autoregressive models suffer from mode collapse by repeatedly producing narrow sets of simple rings, while diffusion models struggle with the mathematical gap between continuous atomic positions and discrete atom types, leading to high-strain, distorted geometries and severe steric clashes.

The article demonstrates that shifting the generative process into a unified, continuous parameter space resolves these conformational and structural issues. The primary objective is to evaluate a new model, MolCRAFT, which integrates Bayesian Flow Networks with three-dimensional symmetry preservation and a noise-reduced sampling strategy, against leading autoregressive and diffusion frameworks on standard benchmark targets.

The approach was evaluated on the CrossDocked benchmark dataset using 100,000 training protein-ligand pairs and 100 test proteins, sampling 100 molecules per test target under controlled molecule sizes. MolCRAFT models both continuous atomic coordinates and discrete atom types as continuous probability distributions, updating parameters smoothly during generation and bypassing intermediate discrete sampling noise.

The key findings demonstrate major improvements across binding quality, physical stability, and computational speed. First, MolCRAFT achieved a reference-level binding score of -6.59 kcal/mol on generated 3D poses without requiring artificial optimization, outperforming competing baselines by up to -0.84 kcal/mol. Second, it lowered median molecular strain energy to 195 kcal/mol—reducing physical distortion by roughly an order of magnitude compared to diffusion baselines, which exhibited median strain energies between 421 and 1,243 kcal/mol. Third, 41.8% of its generated complexes remained within 2 Å of docking positions upon verification, outperforming all baselines and closely matching natural binding consistency. Finally, MolCRAFT generated complete and valid molecules with a 96.7% success rate in 141 seconds per 100 samples, representing an approximate 24-fold to 44-fold acceleration over leading diffusion models.

These results show that earlier artificial high affinity scores were largely distorted artifacts of redocking software rearranging poorly formed poses. By capturing realistic interatomic interactions directly in three dimensions, MolCRAFT provides a more dependable virtual drug screening pipeline. This reduces the risk of wasting expensive laboratory synthesis resources on structurally unfeasible computational false positives, while drastically cutting computational runtimes.

Organizations developing computational discovery pipelines should consider transitioning from hybrid continuous-discrete diffusion methods to unified continuous parameter frameworks. Benchmarking protocols must also enforce strict controls on candidate molecular sizes, as unconstrained size inflation creates misleadingly high affinity scores. Before committing candidates to physical synthesis, teams should conduct wet-lab experimental validations to confirm actual binding performance.

Confidence in these computational benchmarks is high given the extensive comparative metrics across 100 test proteins. However, the evaluation relies on in silico docking approximations and simulated datasets, which can introduce distribution shifts. Practical adoption must account for the boundary conditions of computational modeling until validated by wet-lab synthesis.

Cover for MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space

Abstract

Generative models for structure-based drug design (SBDD) have shown promising results in recent years. Existing works mainly focus on how to generate molecules with higher binding affinity, ignoring the feasibility prerequisites for generated 3D poses and resulting in false positives. We conduct thorough studies on key factors of ill-conformational problems when applying autoregressive methods and diffusion to SBDD, including mode collapse and hybrid continuous-discrete space. In this paper, we introduce MolCRAFT, the first SBDD model that operates in the continuous parameter space, together with a novel noise reduced sampling strategy. Empirical results show that our model consistently achieves superior performance in binding affinity with more stable 3D structure, demonstrating our ability to accurately model interatomic interactions. To our best knowledge, MolCRAFT is the first to achieve reference-level Vina Scores (-6.59 kcal/mol) with comparable molecular size, outperforming other strong baselines by a wide margin (-0.84 kcal/mol). Code is available at https://github.com/AlgoMole/MolCRAFT.

Table of Contents

  • 1. Introduction
  • 2. Challenges of Generative Models in SBDD
  • 2.1. Failure Modes of Generated Molecules
  • 2.2. Molecular Mode Collapse
  • 2.3. Hybrid Continuous-Discrete Space
  • 3. Preliminary
  • 3.1. Problem Definition
  • 3.2. Molecular Generation in Parameter Space
  • 4. Methodology
  • 4.1. Resolving Different Modalities in Parameter Space
  • 4.2. Noise Reduced Sampling in Parameter Space
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Main Results
  • 5.3. Ablation Study of Sampling Strategy
  • 6. Related Work
  • 7. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Detailed Formulation of BFN
  • B. Proof of SE-(3) Invariant Objective and SE-(3) Equivariant Sampling Process
  • C. Implementation Details
  • C.1. Parameterization with SE-(3) Equivariant Network
  • C.2. Featurization
  • C.3. Model Hyperparameters
  • D. Additional Experimental Results
  • D.1. Full Evaluation Results
  • D.2. Ablation Studies

Knowls

  1. Knowl 1 — MolCRAFT Unified Parameter Space for Continuous Coordinates and Discrete Atom Types

    model/method

    MolCRAFT formulates structure-based drug design (SBDD) within a Bayesian Flow Network (BFN) operating entirely in a unified, continuous parameter space θ=[θx,θv]\theta = [\theta^x, \theta^v] to eliminate the continuous-discrete discrepancy common in hybrid molecular diffusion models.

    Given a protein binding pocket P={(xP(i),vP(i))}i=1NPP = \{(x_P^{(i)}, v_P^{(i)})\}_{i=1}^{N_P} with NPN_P atoms (xP(i)∈R3x_P^{(i)} \in \mathbb{R}^3, vP(i)∈RDPv_P^{(i)} \in \mathbb{R}^{D_P}), the objective is to generate a ligand molecule M={(xM(i),vM(i))}i=1NMM = \{(x_M^{(i)}, v_M^{(i)})\}_{i=1}^{N_M} (xM(i)∈R3x_M^{(i)} \in \mathbb{R}^3, vM(i)∈RDMv_M^{(i)} \in \mathbb{R}^{D_M}), denoted compactly as p=[xP,vP]p = [x_P, v_P] and m=[xM,vM]m = [x_M, v_M].

    1. Continuous Atom Coordinates θx\theta^x: Coordinates xx are parameterized by a Gaussian distribution N(x∣μ,ρ−1I)\mathcal{N}(x \mid \mu, \rho^{-1} I) with parameter θx={μ,ρ}\theta^x = \{\mu, \rho\}, where μ∈RNM×3\mu \in \mathbb{R}^{N_M \times 3} is learned and ρ∈R\rho \in \mathbb{R} is accumulated precision. For coordinate noise factor αi\alpha_i, the Bayesian update function {μi,ρi}←h({μi−1,ρi−1},yix,αi)\{\mu_i, \rho_i\} \leftarrow h(\{\mu_{i-1}, \rho_{i-1}\}, y^x_i, \alpha_i) is: ρi=ρi−1+αi\rho_i = \rho_{i-1} + \alpha_i μi=μi−1ρi−1+yixαiρi\mu_i = \frac{\mu_{i-1}\rho_{i-1} + y^x_i \alpha_i}{\rho_i} Continuous coordinate noise defines the sender distribution: pS(yx∣xM;α)=N(yx∣xM,α−1I)p_S(y^x \mid x_M; \alpha) = \mathcal{N}(y^x \mid x_M, \alpha^{-1} I)

    2. Discrete Atom Types θv\theta^v: Atom types vv are parameterized by categorical probabilities θv∈RNM×K\theta^v \in \mathbb{R}^{N_M \times K}, where KK is the number of atom types. For noise parameter αi′\alpha'_i, the Bayesian update is: h(θi−1v,yiv,αi′)k=exp⁡(yi,kv)(θi−1v)k∑j=1Kexp⁡(yi,jv)(θi−1v)jh(\theta^v_{i-1}, y^v_i, \alpha'_i)_k = \frac{\exp(y^v_{i,k}) (\theta^v_{i-1})_k}{\sum_{j=1}^K \exp(y^v_{i,j}) (\theta^v_{i-1})_j} Continuous Gaussian noise is applied to discrete one-hot representations evM∈RNM×Ke_{v_M} \in \mathbb{R}^{N_M \times K} via the sender distribution: pS(yv∣vM;α′)=N(yv∣α′(KevM−1),α′KI)p_S(y^v \mid v_M; \alpha') = \mathcal{N}(y^v \mid \alpha'(K e_{v_M} - 1), \alpha' K I)

    3. Receiver Distributions: The neural network Φ(θi−1,p,ti)\Phi(\theta_{i-1}, p, t_i) reconstructs the clean molecule estimate m^=[x^,v^]∼pO(m∣θi−1,p;ti)\hat{m} = [\hat{x}, \hat{v}] \sim p_O(m \mid \theta_{i-1}, p; t_i), giving receiver distributions: pR(yx∣θx,p;t)=N(yx∣x^(θx,p,t),α−1I)p_R(y^x \mid \theta^x, p; t) = \mathcal{N}(y^x \mid \hat{x}(\theta^x, p, t), \alpha^{-1} I) pR((yv)(d)∣θv,p;t)=∑k=1KpOv(k∣θ;t) pSv((yv)(d)∣k;α′)p_R((y^v)^{(d)} \mid \theta^v, p; t) = \sum_{k=1}^K p_O^v(k \mid \theta; t) \, p_S^v((y^v)^{(d)} \mid k; \alpha') for each atom d∈{1,…,NM}d \in \{1, \dots, N_M\}.

  2. Knowl 2 — MolCRAFT Noise-Reduced Sampling in Continuous Parameter Space

    algorithm

    Standard Bayesian Flow Network sampling draws noisy sample-space latents yiy_i at each step ({θi−1→Φm^→pSyi→pUθi}\{ \theta_{i-1} \xrightarrow{\Phi} \hat{m} \xrightarrow{p_S} y_i \xrightarrow{p_U} \theta_i \}), which reintroduces high variance, especially for discrete atom types. MolCRAFT performs noise-reduced sampling directly in the continuous parameter space by taking the continuous expectation m^=[x^,v^]\hat{m} = [\hat{x}, \hat{v}] from the output network Φ\Phi and evaluating the Bayesian flow distribution pF(θi∣m^,p;t)p_F(\theta_i \mid \hat{m}, p; t) in closed form:

    γ(t)=β(t)1−β(t)\gamma(t) = \frac{\beta(t)}{1 - \beta(t)} pF(μ∣x^,p;t)=N(μ∣γ(t)x^,γ(t)(1−γ(t))I)p_F(\mu \mid \hat{x}, p; t) = \mathcal{N}(\mu \mid \gamma(t)\hat{x}, \gamma(t)(1-\gamma(t))I) pF(θv∣v^,p;t)=Eyv∼N(yv∣β′(t)(Kev^−1),β′(t)KI)[δ(θv−softmax(yv))]p_F(\theta^v \mid \hat{v}, p; t) = \mathbb{E}_{y^v \sim \mathcal{N}(y^v \mid \beta'(t)(K e_{\hat{v}} - 1), \beta'(t)KI)} \left[ \delta(\theta^v - \text{softmax}(y^v)) \right]

    function update(x_hat, v_hat, beta_t, beta_prime_t, t)
        gamma = beta_t / (1.0 - beta_t)
        sample mu ~ N(gamma * x_hat, gamma * (1.0 - gamma) * I)
        sample y_v ~ N(beta_prime_t * (K * e_v_hat - 1.0), beta_prime_t * K * I)
        theta_v = softmax(y_v along atom type dimension)
        return mu, theta_v
    end function
    Input: Network Phi, pocket features p, sampling steps N, atom count n_M, type count K, hyperparams sigma_1, beta_1
    Output: Generated molecule [x_hat, v_hat]
    mu = 0
    rho = 1
    theta_v = matrix of shape (n_M, K) filled with 1.0 / K
    for i = 1 to N do
        t = (i - 1) / N
        x_hat, v_hat = Phi(mu, theta_v, p, t)
        beta_t = sigma_1^(-2 * t) - 1.0
        beta_prime_t = (t^2) * beta_1
        mu, theta_v = update(x_hat, v_hat, beta_t, beta_prime_t, t)
    end for
    x_hat, p_v_O = Phi(mu, theta_v, p, 1.0)
    sample v_hat ~ p_v_O
    return [x_hat, v_hat]

    Sampling runs for N=100N = 100 steps with noise schedules β1=1.5\beta_1 = 1.5 and σ1=0.03\sigma_1 = 0.03.

  3. Knowl 3 — MolCRAFT Discrete-Time Training Objective and Loss Formulation

    algorithm

    MolCRAFT is trained using an nn-step discrete-time variational lower bound objective minimizing the Kullback-Leibler (KL) divergence between the sender distribution pSp_S and receiver distribution pRp_R over nn discretization steps:

    Ln(m,p)=Ei∼U(1,n)Eyi∼pS,θi−1∼pF[DKL(pS∥pR)]=Lxn+LvnL^n(m, p) = \mathbb{E}_{i \sim \mathcal{U}(1, n)} \mathbb{E}_{y_i \sim p_S, \theta_{i-1} \sim p_F} \left[ D_{\text{KL}}(p_S \parallel p_R) \right] = L_x^n + L_v^n

    1. Coordinate Loss (analytic Gaussian KL divergence): Lxn=DKL(N(xM,αi−1I)∥N(x^(θi−1,p,t),αi−1I))=αi2∥xM−x^(θi−1,p,t)∥2=1−σ12/n2σ12i/n∥xM−x^∥2L_x^n = D_{\text{KL}}\left(\mathcal{N}(x_M, \alpha_i^{-1} I) \parallel \mathcal{N}(\hat{x}(\theta_{i-1}, p, t), \alpha_i^{-1} I)\right) = \frac{\alpha_i}{2} \|x_M - \hat{x}(\theta_{i-1}, p, t)\|^2 = \frac{1 - \sigma_1^{2/n}}{2 \sigma_1^{2i/n}} \|x_M - \hat{x}\|^2

    2. Atom Type Loss: Lvn=ln⁡pS(yv∣vM;α)−∑d=1NMln⁡(∑k=1KpO(k∣θ;t) pS((yv)(d)∣k;α))L_v^n = \ln p_S(y^v \mid v_M; \alpha) - \sum_{d=1}^{N_M} \ln \left( \sum_{k=1}^K p_O(k \mid \theta; t) \, p_S((y^v)^{(d)} \mid k; \alpha) \right)

    Input: Ligand coordinates x_M, ligand types v_M, pocket p, hyperparams sigma_1, beta_1, training steps n
    Output: Scalar loss L_n
    i ~ UniformInteger(1, n)
    t = (i - 1) / n
    beta_t = sigma_1^(-2 * t) - 1.0
    beta_prime_t = (t^2) * beta_1
    sample mu ~ p_F^x(mu | x_M, p; t, beta_t)
    sample theta_v ~ p_F^v(theta_v | v_M, p; t, beta_prime_t)
    x_hat, v_hat = p_O(mu, theta_v, p, t)
    L_x = ((1.0 - sigma_1^(2.0 / n)) / (2.0 * sigma_1^(2.0 * i / n))) * ||x_M - x_hat||^2
    alpha = beta_1 * (2 * i - 1) / (n^2)
    sample y_v ~ p_S^v(y_v | v_M, alpha)
    L_v = ln p_S^v(y_v | v_M; alpha) - ln p_R^v(y_v | v_hat; alpha, t)
    return L_x + L_v

    The model is trained with n=1000n = 1000 training steps, Adam optimizer (learning rate 0.0050.005, batch size 88), exponential moving average of parameters with decay factor 0.9990.999, converging within 15 epochs.

  4. Knowl 4 — SE(3) Invariance of Likelihood and SE(3) Equivariant Coordinate Sampling

    theoretical result

    Let Tg(x)=Rx+bT_g(x) = R x + b denote an arbitrary transformation in the Special Euclidean group SE(3)\mathrm{SE}(3), where R∈SO(3)R \in \mathrm{SO}(3) is a 3D rotation matrix and b∈R3b \in \mathbb{R}^3 is a 3D translation vector.

    Preconditions:

    1. The center of mass (CoM) of the protein atom coordinates is translated to zero: x~P=QxP\tilde{x}_P = Q x_P where Q=I3⊗(INP−1NP1NP1NP⊤)Q = I_3 \otimes (I_{N_P} - \frac{1}{N_P} \mathbf{1}_{N_P} \mathbf{1}_{N_P}^\top). Ligand atom coordinates xMx_M and parameter mean μ\mu are translated by the identical protein CoM vector, defining zero-CoM representations p~,m~,μ~\tilde{p}, \tilde{m}, \tilde{\mu}.
    2. The neural network backbone Φ(μ,xP)\Phi(\mu, x_P) is an SE(3)\mathrm{SE}(3)-equivariant graph neural network satisfying Φ(Rμ,RxP)=RΦ(μ,xP)\Phi(R \mu, R x_P) = R \Phi(\mu, x_P).

    Result:

    1. The loss objective is SE(3)\mathrm{SE}(3)-invariant: ∥Tg(xM)−Tg(x^)∥2=∥RxM+b−Rx^−b∥2=∥R(xM−x^)∥2=∥xM−x^∥2\|T_g(x_M) - T_g(\hat{x})\|^2 = \|R x_M + b - R \hat{x} - b\|^2 = \|R(x_M - \hat{x})\|^2 = \|x_M - \hat{x}\|^2 Consequently, the likelihood is invariant to joint SE(3)\mathrm{SE}(3) transformations on the protein-ligand complex: pϕ(Tg(m∣p))=pϕ(m∣p)p_\phi(T_g(m \mid p)) = p_\phi(m \mid p).
    2. Sampling generation in continuous parameter space produces coordinate outputs xNx_N that are SE(3)\mathrm{SE}(3)-equivariant with respect to the pocket coordinates xPx_P.
  5. Knowl 5 — MolCRAFT SE(3)-Equivariant PosNet3D Backbone and Graph Featurization

    model/method

    MolCRAFT parameterizes the denoising network Φ(θ,p,t)\Phi(\theta, p, t) using PosNet3D on dynamically constructed dual graphs: a kk-nearest neighbors (k=32k=32) graph between ligand and protein atoms, and a fully connected graph between ligand atoms.

    For each layer l∈{0,…,L−1}l \in \{0, \dots, L-1\} (L=9L=9, hidden dimension 128128, 1616 attention heads, ReLU activations with Layer Normalization): hil+1=hil+∑j∈NG(i)ϕh(dijl,hil,hjl,eij,t)h_i^{l+1} = h_i^l + \sum_{j \in \mathcal{N}_G(i)} \phi_h(d_{ij}^l, h_i^l, h_j^l, e_{ij}, t) Δxi=∑j∈NG(i)(xjl−xil) ϕx(dijl,hil+1,hjl+1,eij,t)\Delta x_i = \sum_{j \in \mathcal{N}_G(i)} (x_j^l - x_i^l) \, \phi_x(d_{ij}^l, h_i^{l+1}, h_j^{l+1}, e_{ij}, t) xil+1=xil+Δxi⋅1molx_i^{l+1} = x_i^l + \Delta x_i \cdot \mathbf{1}_{\text{mol}} where NG(i)\mathcal{N}_G(i) is the neighborhood of atom ii, dijl=∥xil−xjl∥2d_{ij}^l = \|x_i^l - x_j^l\|_2, eije_{ij} is a 4-dimensional one-hot edge type (protein-protein, ligand-ligand, protein-ligand, ligand-protein), and 1mol\mathbf{1}_{\text{mol}} restricts coordinate updates to ligand atoms only. ϕh\phi_h and ϕx\phi_x are multi-head attention blocks using hilh_i^l as query and [hil,hjl,eij][h_i^l, h_j^l, e_{ij}] as keys and values. Distance features are expanded using radial basis functions.

    Initializations and Featurization:

    • Layer 0 inputs: x0=[μ,xP]x^0 = [\mu, x_P] and h0=linear(θv,vP,t)h^0 = \text{linear}(\theta^v, v_P, t).
    • Protein features vPv_P: one-hot element (H, C, N, O, S, Se), 20-dim one-hot amino acid type, 1-dim backbone indicator, 1-hot arm/scaffold pocket region indicator.
    • Ligand features: 1-hot element (C, N, O, F, P, S, Cl), 1-hot aromatic vs. non-aromatic indicator.
    • Output layer: directly yields coordinate estimate x^=Φx\hat{x} = \Phi^x and categorical distributions v^(d)=softmax((Φv)(d))\hat{v}^{(d)} = \text{softmax}((\Phi^v)^{(d)}).
  6. Knowl 6 — Benchmark Evaluation of MolCRAFT on CrossDocked2020 Dataset

    data/table

    Evaluation of generated molecules on 100 test protein pockets from CrossDocked2020 (100 samples per pocket) under molecular size control. MolCRAFT achieves reference-level direct Vina Scores (-6.59 kcal/mol vs -6.36 kcal/mol for reference) without requiring redocking-based pose rearrangements, while maintaining the highest Binding Feasibility (33.8%) and lowest median Strain Energy among diffusion baselines (195 kcal/mol vs 1243 kcal/mol for TargetDiff).

    Methods Vina Score () Vina Min () Vina Dock () Strain Energy () Clash () RMSD () SA () QED () BF () Size
    Avg. Med. Avg. Med. Avg. Med. 25% 50% 75% Avg. % < 2Ä Avg. Avg. (%) Avg.
    Reference -6.36 -6.46 -6.71 -6.49 -7.45 -7.26 34 107 196 5.51 34.0 0.73 0.48 29.0 22.8
    AR -5.75 -5.64 -6.18 -5.88 -6.75 -6.62 259 595 2286 4.49 31.1 0.64 0.51 17.3 17.7
    Pocket2Mol -5.14 -4.70 -6.42 -5.82 -7.15 -6.79 102 189 374 6.24 30.8 0.76 0.57 24.6 17.7
    Ours-small -5.96 -5.89 -6.33 -6.04 -6.98 -6.63 44 103 274 4.88 38.6 0.74 0.52 33.3 17.8
    FLAG 16.48 4.53 1.21 -4.04 -5.63 -6.61 143 396 1164 40.83 8.2 0.70 0.49 3.8 21.5
    TargetDiff -5.47 -6.30 -6.64 -6.83 -7.80 -7.91 369 1243 13871 10.84 29.4 0.58 0.48 14.4 24.2
    Decomp-R -5.19 -5.27 -6.03 -6.00 -7.03 -7.16 115 421 1424 8.16 22.7 0.66 0.51 14.6 21.2
    Ours -6.59 -7.04 -7.27 -7.26 -7.92 -8.01 83 195 510 7.09 41.8 0.69 0.50 33.8 22.7
    Decomp-O -5.67 -6.04 -7.04 -7.09 -8.39 -8.43 379 983 4133 14.63 23.9 0.61 0.45 11.1 29.4
    Ours-large -6.61 -8.14 -8.14 -8.42 -9.25 -9.20 171 333 1110 10.73 42.7 0.62 0.46 31.1 29.4

    Definitions: BF (Binding Feasibility) is the percentage of generated molecules simultaneously satisfying Vina Score <−2.49< -2.49 kcal/mol, Strain Energy <836< 836 kcal/mol, and symmetry-corrected ligand-redocking RMSD <2A˚< 2\text{\AA} (95th percentile reference thresholds). Clash counts steric overlap interactions in the protein-ligand complex.

  7. Knowl 7 — Substructural Bond and Angle Geometry Distribution Fidelity

    empirical result

    MolCRAFT captures substructural geometric distributions with higher fidelity than existing autoregressive and diffusion-based SBDD models, as measured by the Jensen-Shannon Divergence (JSD) between generated and reference distributions across bond lengths, bond angles, and torsion angles on the CrossDocked test set.

    1. Average Substructural JSD Summary:
    • Bond Length JSD: MolCRAFT achieves 0.3190.319, outperforming Decomp-R (0.3480.348), Decomp-O (0.3590.359), TargetDiff (0.3820.382), Pocket2Mol (0.4850.485), and AR (0.5540.554).
    • Bond Angle JSD: MolCRAFT achieves 0.3790.379, outperforming FLAG (0.4060.406), Decomp-R (0.4120.412), Decomp-O (0.4140.414), TargetDiff (0.4350.435), and Pocket2Mol (0.4820.482).
    • Torsion Angle JSD: MolCRAFT achieves 0.3220.322, superior to TargetDiff (0.4100.410), Pocket2Mol (0.4640.464), and AR (0.5470.547), trailing only FLAG (0.2700.270, which is explicitly fragment-based but produces 0.2%0.2\% unparseable outlier molecules).
    1. Multi-Modal Bond Length Distribution Capture: While autoregressive baselines (AR, Pocket2Mol) collapse to unimodal bond lengths across diverse bond types, MolCRAFT is the only tested model that successfully recovers the true multi-modal bond length distributions characteristic of C-C, C-N, and C-O chemical bonds.
  8. Knowl 8 — Mode Collapse in Autoregressive SBDD Models

    empirical result

    Atom-based autoregressive SBDD models exhibit severe mode collapse, generating repetitive substructures due to unnatural sequential ordering constraints during generation. In contrast, MolCRAFT achieves high uniqueness and balanced ring distributions matching natural reference molecules on the CrossDocked benchmark:

    1. Sample Uniqueness (percentage of unique generated molecules per pocket):
    • AR: 36.2%36.2\%
    • Pocket2Mol: 73.7%73.7\%
    • DecompDiff-R: 50.3%50.3\%
    • DecompDiff-O: 61.6%61.6\%
    • TargetDiff: 99.6%99.6\%
    • FLAG: 99.7%99.7\%
    • MolCRAFT: 97.7%97.7\%
    1. Ring Substructure Proportions (relative to ring-containing molecules):
    • Fused Rings: Reference =30.0%= 30.0\%, Train =21.6%= 21.6\%. Pocket2Mol collapses heavily toward fused rings (52.0%52.0\%), whereas MolCRAFT matches the reference closely (30.9%30.9\%).
    • 3-Membered Rings: Reference =4.0%= 4.0\%, Train =3.8%= 3.8\%. AR exhibits severe pathological mode preference (50.8%50.8\% 3-membered rings), whereas MolCRAFT produces 0.0%0.0\%.
    • 6-Membered Rings: Reference =84.0%= 84.0\%, Train =90.9%= 90.9\%. AR yields 71.9%71.9\%, Pocket2Mol 88.6%88.6\%, and MolCRAFT 85.1%85.1\%.
  9. Knowl 9 — Ablation of Parameter-Space Sampling Strategy and Computational Efficiency

    empirical result

    Ablation experiments comparing parameter-space sampling against data-space sampling across step counts on 100 test proteins (10 samples per pocket) demonstrate significant gains in both binding quality and sampling speed:

    1. Binding Affinity Boost: Switching from data-space sampling to parameter-space noise-reduced sampling improves the mean Vina Score / Vina Min from −5.42-5.42 / −6.30-6.30 kcal/mol to −6.51-6.51 / −7.13-7.13 kcal/mol.

    2. Sampling Step Convergence: Unlike data-space sampling (which requires 1000 steps and degrades due to over-smoothing at fine noise scales), parameter-space sampling reaches optimal molecular property scores (QED, Synthetic Accessibility, Completeness) within 100100 sampling steps.

    3. Runtime and Generation Success: MolCRAFT generates 100 samples in 141141 seconds with 96.7%96.7\% Generation Success (valid and complete molecules). This represents a 24.3×24.3\times speedup over TargetDiff (34283428 seconds for 100 samples) and a 43.9×43.9\times speedup over DecompDiff (61896189 seconds for 100 samples).

  10. Knowl 10 — Protein-Ligand Interaction Counts Under Molecular Size Stratification

    data/table

    Evaluation of key non-covalent protein-ligand interactions (Hydrogen Bond Donors, Hydrogen Bond Acceptors, van der Waals contacts, and Hydrophobic interactions) using PoseCheck across size-stratified model groups on the CrossDocked benchmark. MolCRAFT forms the highest number of hydrogen bonds across all size tiers, confirming accurate physical interaction modeling rather than surface volume exploitation.

    Size Tier Methods HB Donors (Avg.) HB Acceptors (Avg.) vdWs (Avg.) Hydrophobic (Avg.) Size (Avg.)
    Reference Reference 0.87 1.42 6.61 5.06 22.8
    Smaller-size AR 0.51 0.90 5.54 3.78 17.7
    Pocket2Mol 0.32 0.63 5.25 4.53 17.7
    FLAG 0.28 0.30 5.85 3.76 16.7
    Ours-small 0.62 1.09 6.24 4.42 17.8
    Reference-size TargetDiff 0.63 0.98 7.92 5.43 24.2
    Decomp-R 0.56 0.99 6.70 4.37 21.2
    Ours 0.71 1.25 7.38 5.07 22.7
    Larger-size Decomp-O 0.52 0.87 9.14 6.84 29.4
    Ours-large 0.75 1.38 8.64 6.07 29.4

Coverage note — Deliberately omitted individual per-bond-type and per-angle-type JSD breakdown tables (Tables 5-7) and baseline-specific prior preparation details (such as AlphaSpace2 pocket prior extraction for Decomp-O), retaining instead the aggregate substructural JSD benchmark and key failure analyses.

References

  1. 1.Alhossary, A., Handoko, S. D., Mu, Y., and Kwoh, C.- K. Fast, accurate, and reliable molecular docking with quickvina 2. Bioinformatics, 31(13):2214–2216, 2015.
  2. 2.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  3. 3.Bjerrum, E. J. and Threlfall, R. Molecular generation with recurrent neural networks (rnns). arXiv preprint arXiv:1705.04612, 2017.
  4. 4.Francoeur, P. G., Masuda, T., Sunseri, J., Jia, A., Iovanisci, R. B., Snyder, I., and Koes, D. R. Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design. Journal of Chemical Information and Modeling, 60(9):4200–4215, 2020a. doi: 10.1021/acs.jcim.0c00411. URL https://doi.org/10.1021/acs.jcim.0c00411. PMID: 32865404.
  5. 5.Francoeur, P. G., Masuda, T., Sunseri, J., Jia, A., Iovanisci, R. B., Snyder, I., and Koes, D. R. Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design. Journal of chemical information and modeling, 60(9):4200–4215, 2020b.
  6. 6.Gómez-Bombarelli, R., Wei, J. N., Duvenaud, D., Hernández-Lobato, J. M., Sánchez-Lengeling, B., Sheberla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic chemical design using a data-driven continuous representation of molecules. ACS central science, 4(2):268–276, 2018.
  7. 7.Graves, A., Srivastava, R. K., Atkinson, T., and Gomez, F. Bayesian flow networks. arXiv preprint arXiv:2308.07037, 2023.
  8. 8.Guan, J., Qian, W. W., Ma, W.-Y., Ma, J., and Peng, J. Energy-inspired molecular conformation optimization. In international conference on learning representations, 2021.
  9. 9.Guan, J., Qian, W. W., Peng, X., Su, Y., Peng, J., and Ma, J. 3d equivariant diffusion for target-aware molecule generation and affinity prediction. In The Eleventh International Conference on Learning Representations, 2022.
  10. 10.Guan, J., Zhou, X., Yang, Y., Bao, Y., Peng, J., Ma, J., Liu, Q., Wang, L., and Gu, Q. DecompDiff: Diffusion models with decomposed priors for structure-based drug design. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 11827–11846. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/guan23a.html.
  11. 11.Harris, C., Didi, K., Jamasb, A. R., Joshi, C. K., Mathis, S. V., Lio, P., and Blundell, T. Benchmarking generated poses: How rational is structure-based drug design with generative models? arXiv preprint arXiv:2308.07413, 2023.
  12. 12.Hassan, N. M., Alhossary, A. A., Mu, Y., and Kwoh, C.-K. Protein-ligand blind docking using quickvina-w with inter-process spatio-temporal integration. Scientific reports, 7(1):15451, 2017.
  13. 13.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  14. 14.Hoogeboom, E., Satorras, V. G., Vignac, C., and Welling, M. Equivariant diffusion for molecule generation in 3d. In International conference on machine learning, pp. 8867–8887. PMLR, 2022.
  15. 15.Isert, C., Atz, K., and Schneider, G. Structure-based drug design with geometric deep learning. Current Opinion in Structural Biology, 79:102548, April 2023. ISSN 0959440X. doi: 10.1016/j.sbi.2023. 102548. URL https://linkinghub.elsevier.com/retrieve/pii/S0959440X23000222.
  16. 16.Jiang, Y., Zhang, G., You, J., Zhang, H., Yao, R., Xie, H., Zhang, L., Xia, Z., Dai, M., Wu, Y., et al. Pocketflow is a data-and-knowledge-driven structure-based molecular generative model. Nature Machine Intelligence, pp. 1–12, 2024.
  17. 17.Katigbak, J., Li, H., Rooklin, D., and Zhang, Y. Alphaspace 2.0: Representing concave biomolecular surfaces using beta-clusters. Journal of Chemical Information and Modeling, 60(3):1494–1508, 2020. doi: 10.1021/acs.jcim.9b00652. URL https://doi.org/10.1021/acs.jcim.9b00652. PMID: 31995373.
  18. 18.Liu, M., Luo, Y., Uchino, K., Maruhashi, K., and Ji, S. Generating 3D Molecules for Target Protein Binding, May 2022. URL http://arxiv.org/abs/2204.09410. arXiv:2204.09410 [cs, q-bio].
  19. 19.Long, S., Zhou, Y., Dai, X., and Zhou, H. Zero-shot 3d drug design by sketching and generating. In NeurIPS, 2022.
  20. 20.Luo, S., Guan, J., Ma, J., and Peng, J. A 3D Generative Model for Structure-Based Drug Design. Advances in Neural Information Processing Systems, 34: 6229–6239, 2021. URL http://arxiv.org/abs/2203.10446.
  21. 21.Masuda, T., Ragoza, M., and Koes, D. R. Generating 3D Molecular Structures Conditional on a Receptor Binding Site with Deep Generative Models, November 2020. URL http://arxiv.org/abs/2010.14442. arXiv:2010.14442 [physics, q-bio].
  22. 22.McNutt, A. T., Francoeur, P., Aggarwal, R., Masuda, T., Meli, R., Ragoza, M., Sunseri, J., and Koes, D. R. Gnina 1.0: molecular docking with deep learning. Journal of cheminformatics, 13(1):1–20, 2021.
  23. 23.Peng, X., Luo, S., Guan, J., Xie, Q., Peng, J., and Ma, J. Pocket2Mol: Efficient molecular sampling based on 3D protein pockets. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 17644–17655. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/peng22b.html.
  24. 24.Peng, X., Guan, J., Liu, Q., and Ma, J. Moldiff: Addressing the atom-bond inconsistency problem in 3d molecule diffusion generation. In International Conference on Machine Learning, pp. 27611–27629. PMLR, 2023.
  25. 25.Powers, A. S., Yu, H. H., Suriana, P. A., and Dror, R. O. Fragment-based ligand generation guided by geometric deep learning on protein-ligand structures. In ICLR2022 Machine Learning for Drug Discovery, 2022. URL https://openreview.net/forum?id=192L9cr-8HU.
  26. 26.Satorras, V. G., Hoogeboom, E., and Welling, M. E (n) equivariant graph neural networks. In International conference on machine learning, pp. 9323–9332. PMLR, 2021.
  27. 27.Schneuing, A., Du, Y., Harris, C., Jamasb, A., Igashov, I., Du, W., Blundell, T., Lio, P., Gomes, C., Welling, M., Bronstein, M., and Correia, B. Structure-based Drug Design with Equivariant Diffusion Models, October 2022. URL http://arxiv.org/abs/2210.13695. arXiv:2210.13695 [cs, q-bio].
  28. 28.Segler, M. H., Kogej, T., Tyrchan, C., and Waller, M. P. Generating focused molecule libraries for drug discovery with recurrent neural networks. ACS central science, 4 (1):120–131, 2018.
  29. 29.Song, Y., Gong, J., Qu, Y., Zhou, H., Zheng, M., Liu, J., and Ma, W.-Y. Unified generative modeling of 3d molecules via bayesian flow networks. arXiv preprint arXiv:2403.15441, 2024a.
  30. 30.Song, Y., Gong, J., Xu, M., Cao, Z., Lan, Y., Ermon, S., Zhou, H., and Ma, W.-Y. Equivariant flow matching with hybrid probability transport for 3d molecule generation. Advances in Neural Information Processing Systems, 36, 2024b.
  31. 31.Tian, S., Wang, J., Li, Y., Li, D., Xu, L., and Hou, T. The application of in silico drug-likeness predictions in pharmaceutical research. Advanced drug delivery reviews, 86: 2–10, 2015.
  32. 32.Trott, O. and Olson, A. J. Autodock vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of Computational Chemistry, 31 (2):455–461, 2010. doi: https://doi.org/10.1002/jcc.21334. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/jcc.21334.
  33. 33.Ursu, O., Rayan, A., Goldblum, A., and Oprea, T. I. Understanding drug-likeness. Wiley Interdisciplinary Reviews: Computational Molecular Science, 1(5):760–781, 2011.
  34. 34.Walters, W. P. Virtual chemical libraries. Journal of Medicinal Chemistry, 62(3):1116–1124, 2019. doi: 10.1021/acs.jmedchem.8b01048. URL https://doi.org/10.1021/acs.jmedchem.8b01048. PMID: 30148631.
  35. 35.Wang, M., Wang, Z., Sun, H., Wang, J., Shen, C., Weng, G., Chai, X., Li, H., Cao, D., and Hou, T. Deep learning approaches for de novo drug design: An overview. Current Opinion in Structural Biology, 72:135–144, February 2022. ISSN 0959440X. doi: 10.1016/j.sbi.2021.10.001. URL https://linkinghub.elsevier.com/retrieve/pii/S0959440X21001433.
  36. 36.Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J. Geodiff: A geometric diffusion model for molecular conformation generation. In International Conference on Learning Representations, 2021.
  37. 37.Zhang, Z. and Liu, Q. Learning Subpocket Prototypes for Generalizable Structure-based Drug Design, May 2023. URL http://arxiv.org/abs/2305.13997. arXiv:2305.13997 [cs, q-bio].
  38. 38.Zhang, Z., Min, Y., Zheng, S., and Liu, Q. Molecule Generation For Target Protein Binding with Structural Motifs. 2023.

Citation

MLA
Qu, Y., et al. “MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space”. arXiv, 2024, http://arxiv.org/abs/2404.12141v4.
APA
Qu, Y., Qiu, K., Song, Y., Gong, J., Han, J., Zheng, M., Zhou, H., & Ma, W.-Y. (2024). MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space. arXiv. http://arxiv.org/abs/2404.12141v4
Chicago
Qu, Y., K. Qiu, Y. Song, et al. 2024. “MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space”. arXiv. http://arxiv.org/abs/2404.12141v4.
Harvard
Qu, Y. et al. (2024) “MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.12141v4.
Vancouver
1. Qu Y, Qiu K, Song Y, Gong J, Han J, Zheng M, Zhou H, Ma W-Y (2024) MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space. arXiv

BibTeX

@article{qu2024molcraft,
  title = {MolCRAFT: Structure-Based Drug Design in Continuous Parameter Space},
  author = {Qu, Yanru and Qiu, Keyue and Song, Yuxuan and Gong, Jingjing and Han, Jiawei and Zheng, Mingyue and Zhou, Hao and Ma, Wei-Ying},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.12141v4},
  eprint = {2404.12141}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/