Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs

Rui JiaoJiaqi HanWenbing HuangYu RongYang Liu

article2023AAAI65 citations

Proposes an energy-motivated equivariant pretraining framework that pairs E(3)-invariant force prediction via Riemann-Gaussian denoising with graph-level noise scale estimation to improve 3D molecular representation learning.

Listen

Accurately modeling three-dimensional molecular structures without relying on scarce labeled data is critical for advancing drug discovery, materials design, and molecular simulations. Conventional machine learning pretraining primarily focuses on two-dimensional graph representations, failing to capture essential spatial geometries and physical symmetries, such as invariance to rotations and translations in space.

The article introduces and evaluates a self-supervised pretraining framework termed 3D Equivariant Molecular Graph Pretraining (3D-EMGP). The main objective is to establish whether training an energy-motivated, symmetry-preserving model on unlabeled 3D molecular structures improves downstream molecular property and force predictions compared to existing pretraining approaches.

The authors designed a physics-inspired approach using an equivariant neural network backbone that relates atomic forces to potential energy gradients. The framework combines two self-supervised objectives: a node-level force prediction task framed as position denoising using a rotation- and translation-invariant Riemann-Gaussian distribution, and a graph-level classification task that identifies the magnitude of noise applied to a molecule. The model was pretrained on 100,000 unlabeled molecular conformations from the GEOM-QM9 dataset and subsequently evaluated on two standard benchmarks: MD17 for force and energy simulation and QM9 for molecular property prediction.

The evaluation produced several key findings. First, 3D-EMGP significantly outperformed baseline and existing pretraining methods on force prediction, cutting the average mean absolute error on MD17 to 0.0968—an improvement of more than 22% over the next best approach and a greater than 50% reduction relative to training without pretraining. Second, the framework achieved state-of-the-art results across most quantum-chemical properties in QM9, such as internal energies and orbital gaps. Third, several traditional 2D pretraining techniques exhibited negative transfer on 3D tasks, producing worse results than no pretraining at all. Fourth, ablation studies verified that both the node-level force prediction and graph-level noise detection contributed meaningfully, while the Riemann-Gaussian formulation prevented performance degradation observed under standard Gaussian assumptions. Finally, testing on alternative model backbones demonstrated average error reductions of 6.9% to 36.1%, confirming the versatility of the method.

These results indicate that embedding physical symmetries and 3D geometric tasks into unsupervised representation learning substantially enhances predictive performance and model stability. For research and development organizations, this strategy reduces the costly experimental or computational overhead required to generate labeled quantum-mechanical data, accelerates molecular screening, and provides more physically reliable energy landscapes.

Organizations developing computational chemistry pipelines should integrate symmetry-aware 3D pretraining workflows rather than relying on legacy 2D graph methods when spatial geometry governs target properties. Future efforts should evaluate the framework on larger macro-molecules, such as proteins, and expand pretraining across broader, more diverse conformational datasets. While the experimental evidence strongly supports the method's effectiveness on small organic molecules, stakeholders should exercise caution when applying the model to properties governed purely by global electronic spatial extents, where transfer gains remain limited.

Cover for Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs

Abstract

Pretraining molecular representation models without labels is fundamental to various applications. Conventional methods mainly process 2D molecular graphs and focus solely on 2D tasks, making their pretrained models incapable of characterizing 3D geometry and thus defective for downstream 3D tasks. In this work, we tackle 3D molecular pretraining in a complete and novel sense. In particular, we first propose to adopt an equivariant energy-based model as the backbone for pretraining, which enjoys the merits of fulfilling the symmetry of 3D space. Then we develop a node-level pretraining loss for force prediction, where we further exploit the Riemann-Gaussian distribution to ensure the loss to be E(3)-invariant, enabling more robustness. Moreover, a graph-level noise scale prediction task is also leveraged to further promote the eventual performance. We evaluate our model pretrained from a large-scale 3D dataset GEOM-QM9 on two challenging 3D benchmarks: MD17 and QM9. Experimental results demonstrate the efficacy of our method against current state-of-the-art pretraining approaches, and verify the validity of our design for each proposed component. Code is available at https://github.com/jiaor17/3D-EMGP.

Table of Contents

  • Introduction
  • Related Works
  • Method Energy-based Molecular Modeling
  • Node-Level: Equivariant Force Prediction
  • Graph-Level: Invariant Noise-scale Prediction
  • Experiments
  • Experimental Setup
  • Main Results
  • Ablation Studies
  • Visualization
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Energy-Based Equivariant Molecular Representation Modeling

    model/method

    To model molecular conformations while preserving 3D physical symmetries, a molecule of NN atoms is represented by coordinates X=[x1,…,xN]∈R3×NX = [x_1, \dots, x_N] \in \mathbb{R}^{3 \times N} and node features H=[h1,…,hN]∈Rm×NH = [h_1, \dots, h_N] \in \mathbb{R}^{m \times N}. An E(3)\mathrm{E}(3)-equivariant graph neural network φEGN\varphi_{\text{EGN}} produces invariant latent node representations:

    H′=φEGN(X,H,E)H' = \varphi_{\text{EGN}}(X, H, \mathcal{E})

    where H′=[h1′,…,hN′]∈Rk×NH' = [h'_1, \dots, h'_N] \in \mathbb{R}^{k \times N} and E\mathcal{E} denotes molecular graph connectivity. The scalar molecular potential energy E^(X)∈R\hat{E}(X) \in \mathbb{R} is obtained via node pooling and a multi-layer perceptron projection head φProj:Rk→R\varphi_{\text{Proj}}: \mathbb{R}^k \to \mathbb{R}:

    E^(X)=φProj(∑i=1Nhi′)\hat{E}(X) = \varphi_{\text{Proj}}\left(\sum_{i=1}^N h'_i\right)

    Because atomic forces point in the direction of steepest potential energy decrease, the predicted 3D force matrix F^(X)∈R3×N\hat{F}(X) \in \mathbb{R}^{3 \times N} acting on all atoms is derived as the negative spatial gradient of the predicted energy:

    F^(X)=−∇XE^(X)\hat{F}(X) = -\nabla_X \hat{E}(X)

    Under Euclidean group transformations g∈E(3)g \in \mathrm{E}(3) (comprising 3D orthogonal transformations O∈R3×3O \in \mathbb{R}^{3 \times 3} and translations t∈R3t \in \mathbb{R}^3), the predicted energy E^(X)\hat{E}(X) is E(3)\mathrm{E}(3)-invariant (E^(g⋅X)=E^(X)\hat{E}(g \cdot X) = \hat{E}(X)), while the derived force field F^(X)\hat{F}(X) is O(3)\mathrm{O}(3)-equivariant and translation-invariant (F^(OX+t)=OF^(X)\hat{F}(OX + t) = O\hat{F}(X)).

  2. Knowl 2 — Doubly E(3)-Invariant Riemann-Gaussian Denoising Formulation

    model/method

    Under a Boltzmann energy distribution p(X)=1Zexp⁡(−E(X)kT)p(X) = \frac{1}{Z} \exp\left(-\frac{E(X)}{kT}\right), the score function corresponds directly to molecular force: ∇Xlog⁡p(X)∝−∇XE(X)=F\nabla_X \log p(X) \propto -\nabla_X E(X) = F. Score matching enables learning this force without ground-truth labels by minimizing denoising error against perturbed coordinates X~∼p(X~∣X)\tilde{X} \sim p(\tilde{X}|X):

    LEFP-DN=EG∼G,X~∼p(X~∣X)[∥F^(X~)−∇X~log⁡p(X~∣X)∥F2]\mathcal{L}_{\text{EFP-DN}} = \mathbb{E}_{G \sim \mathcal{G}, \tilde{X} \sim p(\tilde{X}|X)} \left[ \|\hat{F}(\tilde{X}) - \nabla_{\tilde{X}} \log p(\tilde{X}|X)\|_F^2 \right]

    Standard Gaussian noise violates the requirement that geometric perturbation probabilities must be independent of spatial orientation and position (p(g1⋅X~∣g2⋅X)=p(X~∣X)p(g_1 \cdot \tilde{X} \mid g_2 \cdot X) = p(\tilde{X} \mid X) for all g1,g2∈E(3)g_1, g_2 \in \mathrm{E}(3)). To satisfy double E(3)\mathrm{E}(3)-invariance, the perturbation is modeled via a Riemann-Gaussian distribution with noise scale σ\sigma:

    pσ(X~∣X)=Rieσ(X~∣X)=1Z(σ)exp⁡(−d2(X~,X)4σ2)p_\sigma(\tilde{X}|X) = \text{Rie}_\sigma(\tilde{X}|X) = \frac{1}{Z(\sigma)} \exp\left(-\frac{d^2(\tilde{X}, X)}{4\sigma^2}\right)

    where the distance metric is defined over zero-mean centered coordinates Y=X−μ(X)Y = X - \mu(X) and Y~=X~−μ(X~)\tilde{Y} = \tilde{X} - \mu(\tilde{X}) (with μ(X)∈R3\mu(X) \in \mathbb{R}^3 being the mean atom position):

    d(X1,X2)=∥Y1⊤Y1−Y2⊤Y2∥Fd(X_1, X_2) = \|Y_1^\top Y_1 - Y_2^\top Y_2\|_F

    The analytical score function of this Riemann-Gaussian distribution with respect to X~\tilde{X} is:

    ∇X~log⁡pσ(X~∣X)=−1σ2[(Y~Y~⊤)Y~−(YY~⊤)Y]\nabla_{\tilde{X}} \log p_\sigma(\tilde{X}|X) = -\frac{1}{\sigma^2}\left[(\tilde{Y}\tilde{Y}^\top)\tilde{Y} - (Y\tilde{Y}^\top)Y\right]

    which has computational complexity O(N)\mathcal{O}(N) for NN atoms.

  3. Knowl 3 — Double E(3)-Invariance of the Riemann-Gaussian Conformation Distribution

    theoretical result

    Let X,X~∈R3×NX, \tilde{X} \in \mathbb{R}^{3 \times N} denote molecular atomic coordinate matrices. Let Y=X−μ(X)Y = X - \mu(X) and Y~=X~−μ(X~)\tilde{Y} = \tilde{X} - \mu(\tilde{X}) denote coordinates translated to zero mean, where μ(X)=1N∑i=1Nxi\mu(X) = \frac{1}{N} \sum_{i=1}^N x_i. Define the conformation distance metric as:

    d(X~,X)=∥Y~⊤Y~−Y⊤Y∥Fd(\tilde{X}, X) = \|\tilde{Y}^\top \tilde{Y} - Y^\top Y\|_F

    and the corresponding Riemann-Gaussian conditional probability density with scale parameter σ\sigma as:

    pσ(X~∣X)=1Z(σ)exp⁡(−d2(X~,X)4σ2)p_\sigma(\tilde{X}|X) = \frac{1}{Z(\sigma)} \exp\left(-\frac{d^2(\tilde{X}, X)}{4\sigma^2}\right)

    For any rigid Euclidean transformations g1,g2∈E(3)g_1, g_2 \in \mathrm{E}(3) where gk⋅X=OkX+tkg_k \cdot X = O_k X + t_k (Ok∈O(3)O_k \in \mathrm{O}(3), tk∈R3t_k \in \mathbb{R}^3), the distribution satisfies double E(3)\mathrm{E}(3)-invariance:

    pσ(g1⋅X~∣g2⋅X)=pσ(X~∣X)p_\sigma(g_1 \cdot \tilde{X} \mid g_2 \cdot X) = p_\sigma(\tilde{X} \mid X)

    Additionally, the distance d(X~,X)d(\tilde{X}, X) and density pσ(X~∣X)p_\sigma(\tilde{X}|X) are permutation invariant with respect to the ordering of columns (atoms) in X~\tilde{X} and XX.

  4. Knowl 4 — Multi-Scale Equivariant Force Prediction (EFP) Loss

    equation

    The node-level Equivariant Force Prediction (EFP) training loss across multiple noise scales {σl}l=1L\{\sigma_l\}_{l=1}^L is defined as:

    LEFP-Final=EG∼G,l∼U(1,L),X~∼pσl(X~∣X)[σl2∥1σlF^(X~)−1α∇X~log⁡pσl(X~∣X)∥F2]\mathcal{L}_{\text{EFP-Final}} = \mathbb{E}_{G \sim \mathcal{G}, l \sim \mathcal{U}(1,L), \tilde{X} \sim p_{\sigma_l}(\tilde{X}|X)} \left[ \sigma_l^2 \left\| \frac{1}{\sigma_l} \hat{F}(\tilde{X}) - \frac{1}{\alpha} \nabla_{\tilde{X}} \log p_{\sigma_l}(\tilde{X}|X) \right\|_F^2 \right]

    where G\mathcal{G} is the pretraining dataset of molecular conformers, U(1,L)\mathcal{U}(1, L) is a discrete uniform distribution over LL noise levels, F^(X~)=−∇X~E^(X~)\hat{F}(\tilde{X}) = -\nabla_{\tilde{X}} \hat{E}(\tilde{X}) is the predicted force vector matrix, ∇X~log⁡pσl(X~∣X)=−1σl2[(Y~Y~⊤)Y~−(YY~⊤)Y]\nabla_{\tilde{X}} \log p_{\sigma_l}(\tilde{X}|X) = -\frac{1}{\sigma_l^2}[(\tilde{Y}\tilde{Y}^\top)\tilde{Y} - (Y\tilde{Y}^\top)Y] is the Riemann-Gaussian score, and α\alpha is a doubly E(3)\mathrm{E}(3)-invariant normalization coefficient given by:

    α=12(∥Y~Y~⊤∥F+∥YY~⊤∥F)\alpha = \frac{1}{2}\left(\|\tilde{Y}\tilde{Y}^\top\|_F + \|Y\tilde{Y}^\top\|_F\right)

  5. Knowl 5 — Graph-Level Invariant Noise-Scale Prediction (INP) Loss

    model/method

    To complement node-level force prediction with global structural awareness, a graph-level classification task is defined to identify the magnitude of perturbation applied to an input molecule. Given original coordinates XX and perturbed coordinates X~∼pσl(X~∣X)\tilde{X} \sim p_{\sigma_l}(\tilde{X}|X) at noise level index l∈{1,…,L}l \in \{1, \dots, L\}, the equivariant graph neural network backbone φEGN\varphi_{\text{EGN}} generates invariant graph-level representations u=∑i=1Nhi′u = \sum_{i=1}^N h'_i and u~=∑i=1Nh~i′\tilde{u} = \sum_{i=1}^N \tilde{h}'_i.

    A classification head φScale\varphi_{\text{Scale}} operates on the concatenated graph embeddings to predict logits over all LL noise levels:

    p=φScale(u∥u~)∈RLp = \varphi_{\text{Scale}}(u \mathbin{\Vert} \tilde{u}) \in \mathbb{R}^L

    The invariant noise-scale prediction loss is the multi-class cross-entropy between the predicted logits and the one-hot target I[l]\mathbb{I}[l]:

    LINP=EG∼G,l∼U(1,L),X~∼pσl(X~∣X)[LCE(I[l],p)]\mathcal{L}_{\text{INP}} = \mathbb{E}_{G \sim \mathcal{G}, l \sim \mathcal{U}(1,L), \tilde{X} \sim p_{\sigma_l}(\tilde{X}|X)} \left[ \mathcal{L}_{\text{CE}}(\mathbb{I}[l], p) \right]

  6. Knowl 6 — 3D Equivariant Molecular Graph Pretraining (3D-EMGP) Framework

    model/method

    The 3D-EMGP framework unifies local physical force estimation and global noise perception into a joint pretraining objective:

    L=λ1LEFP-Final+λ2LINP\mathcal{L} = \lambda_1 \mathcal{L}_{\text{EFP-Final}} + \lambda_2 \mathcal{L}_{\text{INP}}

    where λ1\lambda_1 and λ2\lambda_2 are balancing hyperparameter coefficients, LEFP-Final\mathcal{L}_{\text{EFP-Final}} is the node-level multi-scale equivariant force prediction loss, and LINP\mathcal{L}_{\text{INP}} is the graph-level invariant noise scale prediction loss.

    Pretraining proceeds on unlabeled 3D molecular conformations sampled from large datasets (e.g., GEOM-QM9). Conformations X~\tilde{X} are sampled from the Riemann-Gaussian distribution using Langevin dynamics. Once pretrained, the backbone parameters φEGN\varphi_{\text{EGN}} are transferred to downstream 3D molecular property prediction or force field simulation benchmarks (such as QM9 and MD17) and fine-tuned with task-specific projection heads.

  7. Knowl 7 — MD17 Benchmark Force Prediction Performance

    data/table

    The table below reports the Mean Absolute Error (MAE) for atomic force prediction across 8 organic molecules in the MD17 dataset. All evaluated methods utilize the identical EGNN 3D backbone architecture.

    Method Aspirin Benzene Ethanol Malon. Naph. Salicylic Toluene Uracil Average
    Base 0.3885 0.1861 0.0599 0.1464 0.3310 0.2683 0.1563 0.1323 0.2086
    AttrMask 0.3643 0.2277 0.0567 0.1456 0.1773 0.3890 0.1093 0.1560 0.2032
    EdgePred 0.4707 0.2036 0.0743 0.1268 0.2310 0.3400 0.1854 0.1933 0.2281
    GPT-GNN 0.4278 0.2492 0.0703 0.1484 0.2080 0.3609 0.1541 0.2219 0.2301
    InfoGraph 0.6578 0.2743 0.1257 0.2647 0.2860 0.5793 0.3821 0.4238 0.3742
    GCC 0.3996 0.2346 0.0662 0.1484 0.2798 0.4263 0.3378 0.2369 0.2662
    GraphCL 0.2333 0.1845 0.0503 0.0852 0.0966 0.1587 0.0725 0.1167 0.1247
    JOAO 0.3646 0.2331 0.0642 0.1029 0.2017 0.3020 0.1322 0.1683 0.1961
    JOAOv2 0.3447 0.2198 0.0568 0.0981 0.1889 0.2753 0.1001 0.1850 0.1836
    GraphMVP 0.3198 0.2800 0.0629 0.0788 0.2350 0.2641 0.0903 0.1339 0.1831
    3D Infomax 0.4592 0.1914 0.0705 0.1263 0.2642 0.3401 0.2032 0.1836 0.2298
    GEM 0.3994 0.2105 0.0871 0.1161 0.1489 0.2344 0.1193 0.1827 0.1873
    PosPred 0.3050 0.2023 0.0519 0.0937 0.0971 0.2481 0.0945 0.1270 0.1525
    3D-EMGP 0.1560 0.1648 0.0389 0.0737 0.0829 0.1187 0.0619 0.0773 0.0968

    3D-EMGP achieves the lowest force MAE across every individual molecule and an overall average MAE of 0.0968, representing a substantial improvement over the second-best baseline GraphCL (0.1247) and the un-pretrained Base model (0.2086).

  8. Knowl 8 — QM9 Chemical Property Prediction Benchmark Performance

    data/table

    The table below presents the Mean Absolute Error (MAE) across 12 target properties on the QM9 dataset, evaluated with an EGNN backbone across all pretraining strategies.

    Method α\alpha Δϵ\Delta\epsilon ϵHOMO\epsilon_{\text{HOMO}} ϵLUMO\epsilon_{\text{LUMO}} μ\mu CvC_v GG HH R2R^2 UU U0U_0 ZPVE
    Base 0.070 49.9 28.0 24.3 0.031 0.031 10.1 10.9 0.067 9.7 9.3 1.51
    AttrMask 0.072 50.0 31.3 37.8 0.020 0.062 11.2 11.4 0.423 10.8 10.7 1.90
    EdgePred 0.086 58.2 37.4 31.9 0.039 0.038 14.5 14.8 0.112 14.2 14.7 1.81
    GPT-GNN 0.103 54.1 35.7 28.8 0.039 0.032 12.2 14.8 0.158 24.8 12.0 1.75
    InfoGraph 0.099 72.2 48.1 38.1 0.041 0.030 16.5 14.5 0.114 14.9 16.4 1.69
    GCC 0.085 57.7 37.7 32.3 0.041 0.034 12.8 14.5 0.104 13.2 13.1 1.66
    GraphCL 0.066 45.5 26.8 22.9 0.027 0.028 10.2 9.6 0.095 9.7 9.6 1.42
    JOAO 0.068 46.0 28.2 22.8 0.028 0.030 10.5 10.0 0.076 9.9 10.1 1.48
    JOAOv2 0.066 45.0 27.8 22.2 0.027 0.028 9.9 9.2 0.087 9.8 9.5 1.43
    GraphMVP 0.070 46.9 28.5 26.3 0.031 0.033 11.2 10.4 0.082 10.3 10.2 1.63
    3D Infomax 0.075 48.8 29.8 25.7 0.034 0.033 13.0 12.4 0.122 12.5 12.7 1.67
    GEM 0.081 52.1 33.8 27.7 0.034 0.035 13.2 13.3 0.089 12.6 13.4 1.73
    PosPred 0.067 40.6 25.1 20.9 0.024 0.035 10.9 10.2 0.115 10.3 10.2 1.46
    3D-EMGP 0.057 37.1 21.3 18.2 0.020 0.026 9.3 8.7 0.092 8.6 8.6 1.38

    3D-EMGP achieves the best performance across 10 of the 12 property targets (such as α=0.057\alpha = 0.057, Δϵ=37.1\Delta\epsilon = 37.1, ϵHOMO=21.3\epsilon_{\text{HOMO}} = 21.3, ϵLUMO=18.2\epsilon_{\text{LUMO}} = 18.2, μ=0.020\mu = 0.020, Cv=0.026C_v = 0.026, G=9.3G = 9.3, H=8.7H = 8.7, U=8.6U = 8.6, U0=8.6U_0 = 8.6, and ZPVE=1.38\text{ZPVE} = 1.38), exhibiting consistent advantages over 2D and 3D pretraining baselines.

  9. Knowl 9 — Ablation Study of 3D-EMGP Pretraining Components

    data/table

    The ablation study below evaluates the impact of each design component in 3D-EMGP on MD17 average energy and force prediction MAE:

    Configuration EFP INP RG Energy Head Energy MAE Force MAE
    Base 0.1191 0.2086
    Full 3D-EMGP ✓ ✓ ✓ ✓ 0.0876 0.0968
    INP only ✓ ✓ ✓ 0.0974 0.1350
    EFP only ✓ ✓ ✓ 0.0905 0.1193
    Gaussian noise ✓ ✓ ✓ 0.0912 0.1060
    Distance denoising ✓ ✓ 0.0931 0.1292
    Direct force ✓ ✓ ✓ 0.0914 0.1267

    The results show that:

    1. Removing either the node-level force prediction (EFP only vs. INP only) degrades performance relative to joint training (Full 3D-EMGP).
    2. Replacing the Riemann-Gaussian (RG) noise with standard Euclidean Gaussian noise increases force MAE from 0.0968 to 0.1060 due to broken double E(3)\mathrm{E}(3)-invariance.
    3. Predicting forces directly via equivariant vector output (Direct force) rather than via the negative gradient of a pooled energy head (Energy Head) worsens force MAE from 0.0968 to 0.1267.
  10. Knowl 10 — 3D-EMGP Generalization Across Geometric GNN Backbones

    empirical result

    When applied to different 3D graph neural network architectures, 3D-EMGP pretraining consistently improves downstream force prediction on MD17:

    1. For SchNet, 3D-EMGP reduces the average MD17 force prediction MAE by 36.1% compared to training SchNet from scratch.
    2. For TorchMD-ET (Equivariant Transformer), 3D-EMGP reduces the average MD17 force prediction MAE by 6.9% compared to training TorchMD-ET from scratch.

    These gains indicate that the benefits of 3D-EMGP pretraining are architecture-agnostic and transfer across diverse equivariant and invariant 3D GNN designs.

Coverage note — None was omitted; all key contributions including the energy-based equivariant model, Riemann-Gaussian formulation, EFP and INP pretraining losses, benchmark results (MD17 and QM9), ablations, and backbone generalizations are covered.

References

  1. 1.Anderson, B.; Hy, T. S.; and Kondor, R. 2019. Cormorant: Covariant molecular neural networks. Advances in neural information processing systems, 32.
  2. 2.Axelrod, S.; and Gomez-Bombarelli, R. 2022. GEOM, energy-annotated molecular conformations for property prediction and molecular generation. Scientific Data, 9(1): 1–14.
  3. 3.Boltzmann, L. 1868. Studien uber das Gleichgewicht der lebenden Kraft. Wissenschafiliche Abhandlungen, 1: 49–96.
  4. 4.Chmiela, S.; Tkatchenko, A.; Sauceda, H. E.; Poltavsky, I.; Schutt, K. T.; and M¨ uller, K.-R. 2017. Machine learning of accurate energy-conserving molecular force fields. Science advances, 3(5): e1603015.
  5. 5.Fang, X.; Liu, L.; Lei, J.; He, D.; Zhang, S.; Zhou, J.; Wang, F.; Wu, H.; and Wang, H. 2022. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 1–8.
  6. 6.Finzi, M.; Stanton, S.; Izmailov, P.; and Wilson, A. G. 2020. Generalizing Convolutional Neural Networks for Equivariance to Lie Groups on Arbitrary Continuous Data. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, 3165–3176. PMLR.
  7. 7.Fuchs, F.; Worrall, D. E.; Fischer, V.; and Welling, M. 2020. SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  8. 8.Gilmer, J.; Schoenholz, S. S.; Riley, P. F.; Vinyals, O.; and Dahl, G. E. 2017. Neural message passing for quantum chemistry. In International conference on machine learning, 1263–1272. PMLR.
  9. 9.Hamilton, W.; Ying, Z.; and Leskovec, J. 2017. Inductive representation learning on large graphs. Advances in neural information processing systems, 30.
  10. 10.Han, J.; Rong, Y.; Xu, T.; and Huang, W. 2022. Geometrically Equivariant Graph Neural Networks: A Survey. arXiv preprint arXiv:2202.07230.
  11. 11.Hu, W.; Liu, B.; Gomes, J.; Zitnik, M.; Liang, P.; Pande, V. S.; and Leskovec, J. 2020a. Strategies for Pre-training Graph Neural Networks. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  12. 12.Hu, Z.; Dong, Y.; Wang, K.; Chang, K.; and Sun, Y. 2020b. GPT-GNN: Generative Pre-Training of Graph Neural Networks. In Gupta, R.; Liu, Y.; Tang, J.; and Prakash, B. A., eds., KDD '20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, August 23-27, 2020, 1857–1867. ACM.
  13. 13.Hutchinson, M. J.; Le Lan, C.; Zaidi, S.; Dupont, E.; Teh, Y. W.; and Kim, H. 2021. Lietransformer: Equivariant self-attention for lie groups. In International Conference on Machine Learning, 4533–4543. PMLR.
  14. 14.Kearnes, S.; McCloskey, K.; Berndl, M.; Pande, V.; and Riley, P. 2016. Molecular graph convolutions: moving beyond fingerprints. Journal of computer-aided molecular design, 30(8): 595–608.
  15. 15.Kipf, T. N.; and Welling, M. 2016. Variational Graph Auto-Encoders. NIPS Workshop on Bayesian Deep Learning.
  16. 16.Liu, S.; Wang, H.; Liu, W.; Lasenby, J.; Guo, H.; and Tang, J. 2021. Pre-training Molecular Graph Representation with 3D Geometry. In International Conference on Learning Representations.
  17. 17.Liu, Y.; Wang, L.; Liu, M.; Lin, Y.; Zhang, X.; Oztekin, B.; and Ji, S. 2022. Spherical Message Passing for 3D Molecular Graphs. In International Conference on Learning Representations.
  18. 18.Luo, S.; and Hu, W. 2020. Differentiable manifold reconstruction for point cloud denoising. In Proceedings of the 28th ACM international conference on multimedia, 1330–1338.
  19. 19.Luo, S.; Shi, C.; Xu, M.; and Tang, J. 2021. Predicting Molecular Conformation via Dynamic Graph Score Matching. Advances in Neural Information Processing Systems, 34.
  20. 20.Qiu, J.; Chen, Q.; Dong, Y.; Zhang, J.; Yang, H.; Ding, M.; Wang, K.; and Tang, J. 2020. Gcc: Graph contrastive coding for graph neural network pre-training. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1150–1160.
  21. 21.Ramakrishnan, R.; Dral, P. O.; Rupp, M.; and Von Lilienfeld, O. A. 2014. Quantum chemistry structures and properties of 134 kilo molecules. Scientific data, 1(1): 1–7.
  22. 22.Rong, Y.; Bian, Y.; Xu, T.; Xie, W.; Wei, Y.; Huang, W.; and Huang, J. 2020. Self-Supervised Graph Transformer on Large-Scale Molecular Data. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  23. 23.Rosenstein, M. T.; Marx, Z.; Kaelbling, L. P.; and Dietterich, T. G. 2005. To transfer or not to transfer. In In NIPS'05 Workshop, Inductive Transfer: 10 Years Later.
  24. 24.Said, S.; Bombrun, L.; Berthoumieu, Y.; and Manton, J. H. 2017. Riemannian Gaussian distributions on the space of symmetric positive definite matrices. IEEE Transactions on Information Theory, 63(4): 2153–2170.
  25. 25.Satorras, V. G.; Hoogeboom, E.; and Welling, M. 2021. E (n) equivariant graph neural networks. In International Conference on Machine Learning, 9323–9332. PMLR.
  26. 26.Schlick, T. 2010. Molecular modeling and simulation: an interdisciplinary guide, volume 2. Springer.
  27. 27.Schütt, K.; Kindermans, P.-J.; Sauceda Felix, H. E.; Chmiela, S.; Tkatchenko, A.; and Müller, K.-R. 2017. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. Advances in neural information processing systems, 30.
  28. 28.Shi, C.; Luo, S.; Xu, M.; and Tang, J. 2021. Learning gradient fields for molecular conformation generation. In International Conference on Machine Learning, 9558–9568. PMLR.
  29. 29.Song, Y.; and Ermon, S. 2020. Improved techniques for training score-based generative models. Advances in neural information processing systems, 33: 12438–12448.
  30. 30.Stärk, H.; Beaini, D.; Corso, G.; Tossou, P.; Dallago, C.; Günnemann, S.; and Liò, P. 2022. 3d infomax improves gnns for molecular property prediction. In International Conference on Machine Learning, 20479–20502. PMLR.
  31. 31.Sun, F.; Hoffmann, J.; Verma, V.; and Tang, J. 2020. InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  32. 32.Sun, M.; Xing, J.; Wang, H.; Chen, B.; and Zhou, J. 2021. MoCL: Data-Driven Molecular Fingerprint via Knowledge-Aware Contrastive Learning from Molecular Graph. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, KDD '21, 3585–3594. New York, NY, USA: Association for Computing Machinery. ISBN 9781450383325.
  33. 33.Thölke, P.; and De Fabritiis, G. 2021. Equivariant transformers for neural network based molecular potentials. In International Conference on Learning Representations.
  34. 34.Thomas, N.; Smidt, T.; Kearnes, S.; Yang, L.; Li, L.; Kohlhoff, K.; and Riley, P. 2018. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds. arXiv preprint arXiv:1802.08219.
  35. 35.Velickovic, P.; Fedus, W.; Hamilton, W. L.; Lio, P.; Bengio, Y.; and Hjelm, R. D. 2019. Deep Graph Infomax. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  36. 36.Vincent, P. 2011. A connection between score matching and denoising autoencoders. Neural computation, 23(7): 1661–1674.
  37. 37.Wallach, I.; Dzamba, M.; and Heifets, A. 2015. AtomNet: a deep convolutional neural network for bioactivity prediction in structure-based drug discovery. arXiv preprint arXiv:1510.02855.
  38. 38.Xu, K.; Hu, W.; Leskovec, J.; and Jegelka, S. 2019. How Powerful are Graph Neural Networks? In International Conference on Learning Representations.
  39. 39.You, Y.; Chen, T.; Shen, Y.; and Wang, Z. 2021a. Graph Contrastive Learning Automated. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, 12121–12132. PMLR.
  40. 40.You, Y.; Chen, T.; Shen, Y.; and Wang, Z. 2021b. Graph contrastive learning automated. In International Conference on Machine Learning, 12121–12132. PMLR.
  41. 41.You, Y.; Chen, T.; Sui, Y.; Chen, T.; Wang, Z.; and Shen, Y. 2020. Graph Contrastive Learning with Augmentations. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  42. 42.Zheng, L.; Fan, J.; and Mu, Y. 2019. Onionnet: a multiple-layer intermolecular-contact-based convolutional neural network for protein–ligand binding affinity prediction. ACS omega, 4(14): 15956–15965.
  43. 43.Zhou, G.; Gao, Z.; Ding, Q.; Zheng, H.; Xu, H.; Wei, Z.; Zhang, L.; and Ke, G. 2023. Uni-Mol: A Universal 3D Molecular Representation Learning Framework. In The Eleventh International Conference on Learning Representations.
  44. 44.Zhu, J.; Xia, Y.; Wu, L.; Xie, S.; Qin, T.; Zhou, W.; Li, H.; and Liu, T.-Y. 2022. Unified 2d and 3d pre-training of molecular representations. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2626–2636.

Citation

MLA
Jiao, R., et al. “Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs”. arXiv, 2022, http://arxiv.org/abs/2207.08824v4.
APA
Jiao, R., Han, J., Huang, W., Rong, Y., & Liu, Y. (2022). Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs. arXiv. http://arxiv.org/abs/2207.08824v4
Chicago
Jiao, R., J. Han, W. Huang, Y. Rong, and Y. Liu. 2022. “Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs”. arXiv. http://arxiv.org/abs/2207.08824v4.
Harvard
Jiao, R. et al. (2022) “Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.08824v4.
Vancouver
1. Jiao R, Han J, Huang W, Rong Y, Liu Y (2022) Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs. arXiv

BibTeX

@article{jiao2022energy,
  title = {Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs},
  author = {Jiao, Rui and Han, Jiaqi and Huang, Wenbing and Rong, Yu and Liu, Yang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.08824v4},
  eprint = {2207.08824}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/