Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules

Jun XiaChengshuai ZhaoBozhen HuZhangyang GaoCheng TanYue LiuSiyuan LiStan Z. Li

article2023ICLR234 citations
Listen

Computational modeling of molecular graphs using graph neural networks is essential for accelerating drug discovery and chemical property prediction. To make use of vast amounts of unlabeled molecular data, standard pre-training methods typically mask and predict raw atom types. However, this common setup frequently leads to negative transfer, where pre-trained models actually perform worse on downstream tasks than models trained entirely from scratch. This failure happens because raw chemical elements form a very small and heavily imbalanced vocabulary—dominated by carbon—making the training task overly simplistic and prone to bias.

The article introduces and evaluates Mole-BERT, an end-to-end, fully data-driven pre-training framework designed to resolve these representation bottlenecks. Mole-BERT replaces raw atom identities with a context-aware tokenizer based on a grouped vector-quantized variational autoencoder, which maps atoms and their chemical neighborhoods into a richer, balanced discrete codebook. It couples this node-level pre-training task, called Masked Atoms Modeling, with a graph-level pre-training task called Triplet Masked Contrastive Learning, which uses multiple masking ratios to capture fine-grained chemical similarity relationships.

The researchers pre-trained the system on two million unlabeled molecules from the ZINC database and benchmarked it across eight standard molecular classification datasets, four property regression datasets, and two drug-target affinity prediction benchmarks. In extensive comparative evaluations, the key findings demonstrate that:

  • Mole-BERT achieved an average classification score of 74.04%, outperforming training from scratch (67.15%) by 6.89 percentage points and surpassing top competing methods, including those relying on complex three-dimensional geometric data.
  • The context-aware node-level task completely eliminated the negative transfer failures seen in standard masking methods on sensitive benchmarks such as HIV and blood-brain barrier penetration datasets.
  • When evaluated across multiple model backbones, Mole-BERT delivered consistent performance improvements, generating relative gains of 6.5% to 10.3% across various graph neural network architectures.
  • In molecule retrieval and drug-target binding affinity tasks, the framework produced lower error rates and retrieved reference molecules that aligned substantially better with established chemical similarity metrics.

These findings indicate that machine learning models in drug discovery do not necessarily require expensive manual annotations or complex spatial geometry datasets to achieve state-of-the-art accuracy. By properly framing pre-training around context-aware tokenization and graded similarity learning, organizations can improve predictive reliability while lowering data preparation overhead and reducing the risk of model failure during lead optimization.

For practical implementation, the authors recommend adopting the context-aware tokenizer as a standard, off-the-shelf component in molecular graph pipelines and combining masked atom modeling with graph-level contrastive tasks. Future work should pilot this framework on related biological problems, such as protein sequence and structure modeling, where vocabulary imbalances also present significant modeling challenges. Users should note that optimal performance depends on tuning vocabulary sizes and masking ratios, and empirical testing remains necessary when transferring the model to entirely new chemical domains.

Cover for Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules

Table of Contents

  • 1 INTRODUCTION
  • 2 RELATED WORK
  • 2.1 PRE-TRAINING ON MOLECULES
  • 2.2 MLM-STYLE PRE-TRAINING STRATEGIES
  • 3 PRELIMINARY
  • 4 PROPOSED PRE-TRAINING FRAMEWORK: MOLE-BERT
  • 4.1 MASKED ATOMS MODELING (MAM)
  • 4.2 GRAPH-LEVEL TASK: TRIPLET MASKED CONTRASTIVE LEARNING (TMCL)
  • 5 EXPERIMENTS
  • 5.1 DATASETS
  • 5.2 EXPERIMENTS CONFIGURATION
  • 5.3 RESULTS AND ANALYSIS
  • 6 CONCLUSIONS AND FUTURE WORKS
  • 7 ACKNOWLEDGEMENTS
  • REFERENCES
  • A DATASETS
  • B MORE EXPERIMENTAL DETAILS
  • C THE INFLUENCE OF THE MASKING RATIO
  • D MORE RESULTS OF MOLECULE RETRIEVAL
  • E IMPLEMENTATION DETAILS OF BASELINES
  • F DISCRETE VAE V.S. GROUP VQ-VAE
  • G TRAINING AND VALIDATION CURVES
  • H ABLATIONS FOR GROUP VQ-VAE
  • I DETAILED INFORMATION OF THE STRUCTURAL ABBREVIATIONS

Knowls

  1. Knowl 1 — Mole-BERT Pre-training Framework

    model/method

    Mole-BERT is a self-supervised pre-training framework for molecular Graph Neural Networks (GNNs) that combines node-level Masked Atoms Modeling (MAM) and graph-level Triplet Masked Contrastive Learning (TMCL). Given an unlabeled molecular dataset D\mathcal{D}, the total pre-training objective LMole-BERT\mathcal{L}_{\text{Mole-BERT}} is formulated as:

    LMole-BERT=LMAM+LTMCL\mathcal{L}_{\text{Mole-BERT}} = \mathcal{L}_{\text{MAM}} + \mathcal{L}_{\text{TMCL}}

    where LMAM\mathcal{L}_{\text{MAM}} trains the GNN encoder to predict context-aware discrete atom codes generated by a group Vector Quantized Variational Autoencoder (group VQ-VAE), and LTMCL\mathcal{L}_{\text{TMCL}} regularizes the graph embeddings via contrastive and triplet ranking objectives across differently masked graph views. Both objectives share the same underlying GNN backbone.

  2. Knowl 2 — Context-Aware Group VQ-VAE Molecular Tokenizer

    model/method

    To represent atoms as discrete tokens capturing local chemical context rather than static element identities, Mole-BERT employs a group VQ-VAE tokenizer. Given a molecular graph G=(V,E)G = (\mathcal{V}, \mathcal{E}) with nn atoms, a 5-layer Graph Isomorphism Network (GIN) encoder computes continuous atom embeddings {h1,h2,…,hn}\{h_1, h_2, \dots, h_n\}.

    A codebook A={e1,e2,…,e∣A∣}\mathcal{A} = \{e_1, e_2, \dots, e_{|\mathcal{A}|}\} of size ∣A∣=512|\mathcal{A}| = 512 is partitioned into four element-specific index groups to prevent code collision between distinct chemical elements:

    • Carbon (C): restricted to code indices [1,128][1, 128]
    • Nitrogen (N): restricted to code indices [129,256][129, 256]
    • Oxygen (O): restricted to code indices [257,384][257, 384]
    • Rare atoms (e.g., P, S, F, Cl, Br, I): restricted to code indices [385,512][385, 512]

    For each atom ii, the quantized discrete token ziz_i is determined by nearest-neighbor lookup within its assigned element group:

    zi=arg⁡min⁡j∥hi−ej∥2z_i = \arg\min_j \|h_i - e_j\|_2

    The selected codebook embeddings {ez1,ez2,…,ezn}\{e_{z_1}, e_{z_2}, \dots, e_{z_n}\} are then passed to a 5-layer GIN decoder to reconstruct the original input node attributes.

  3. Knowl 3 — Group VQ-VAE Tokenizer Loss Formulation

    equation

    The group VQ-VAE tokenizer is trained on molecular graphs using a three-component loss function LVQ\mathcal{L}_{\text{VQ}}:

    LVQ=1n∑i=1n(1−viTv^i∥vi∥∥v^i∥)γ+1n∑i=1n∥sg[hi]−ezi∥22+βn∑i=1n∥sg[ezi]−hi∥22\mathcal{L}_{\text{VQ}} = \frac{1}{n} \sum_{i=1}^{n} \left( 1 - \frac{v_i^T \hat{v}_i}{\|v_i\| \|\hat{v}_i\|} \right)^\gamma + \frac{1}{n} \sum_{i=1}^{n} \|\text{sg}[h_i] - e_{z_i}\|_2^2 + \frac{\beta}{n} \sum_{i=1}^{n} \|\text{sg}[e_{z_i}] - h_i\|_2^2

    where:

    • viv_i is the ground-truth attribute vector of atom ii, and v^i\hat{v}_i is the reconstructed attribute vector from the decoder.
    • The first term is the scaled cosine error reconstruction loss with scaling parameter γ≥1\gamma \ge 1.
    • The second term is the vector quantization loss that optimizes the codebook vectors ezie_{z_i} toward the encoder representations hih_i.
    • The third term is the commitment loss weighted by hyperparameter β=0.25\beta = 0.25, which prevents encoder representations from growing arbitrarily.
    • sg[⋅]\text{sg}[\cdot] denotes the stop-gradient operator, and gradients are propagated from decoder to encoder using the straight-through estimator.
  4. Knowl 4 — Masked Atoms Modeling (MAM)

    model/method

    Masked Atoms Modeling (MAM) is a node-level pre-training pretext task. Given an input molecular graph GG, a random subset M\mathcal{M} containing 15% of the atoms is masked, resulting in a corrupted graph GMG^\mathcal{M}. The GNN encoder processes GMG^\mathcal{M}, and a softmax classification head predicts the discrete token values zi∈Az_i \in \mathcal{A} generated by the group VQ-VAE tokenizer over the vocabulary size ∣A∣=512|\mathcal{A}| = 512.

    The MAM pre-training objective is defined as the negative log-likelihood:

    LMAM=−∑G∈D∑i∈Mlog⁡p(zi∣GM)\mathcal{L}_{\text{MAM}} = - \sum_{G \in \mathcal{D}} \sum_{i \in \mathcal{M}} \log p(z_i \mid G^\mathcal{M})

    where D\mathcal{D} is the pre-training molecule dataset.

  5. Knowl 5 — Triplet Masked Contrastive Learning (TMCL)

    model/method

    Triplet Masked Contrastive Learning (TMCL) is a graph-level self-supervised objective designed to reflect graded semantic similarities between molecular structures rather than treating all non-identical graphs as equally dissimilar.

    Given an anchor graph GG, two corrupted versions GM1G^{M_1} and GM2G^{M_2} are generated by masking atom subsets M1M_1 and M2M_2 with masking ratios of 15% and 30%, respectively, establishing the prior that GM1G^{M_1} is semantically closer to GG than GM2G^{M_2}. Using graph-level representations hGh_G, hM1h_{M_1}, and hM2h_{M_2} obtained via mean pooling readout, the TMCL loss is defined as:

    LTMCL=Lcon+μLtri\mathcal{L}_{\text{TMCL}} = \mathcal{L}_{\text{con}} + \mu \mathcal{L}_{\text{tri}}

    Lcon=−∑G∈Dlog⁡exp⁡(sim(hM1,hM2)/τ)∑G′∈Bexp⁡(sim(hM1,hG′)/τ)\mathcal{L}_{\text{con}} = - \sum_{G \in \mathcal{D}} \log \frac{\exp(\text{sim}(h_{M_1}, h_{M_2}) / \tau)}{\sum_{G' \in \mathcal{B}} \exp(\text{sim}(h_{M_1}, h_{G'}) / \tau)}

    Ltri=∑G∈Dmax⁡(sim(hG,hM2)−sim(hG,hM1),0)\mathcal{L}_{\text{tri}} = \sum_{G \in \mathcal{D}} \max\left(\text{sim}(h_G, h_{M_2}) - \text{sim}(h_G, h_{M_1}), 0\right)

    where sim(⋅,⋅)\text{sim}(\cdot, \cdot) is cosine similarity, B\mathcal{B} is the sampled mini-batch, τ=0.1\tau = 0.1 is the temperature hyperparameter, and μ∈{0.1,0.3,0.5}\mu \in \{0.1, 0.3, 0.5\} is a trade-off coefficient tuned on validation data.

  6. Knowl 6 — Molecular Property Prediction Benchmarking on MoleculeNet

    data/table

    Mole-BERT and various molecular pre-training baselines were evaluated on 8 binary classification datasets from MoleculeNet using 10 random seeds with scaffold splitting (80% train, 10% validation, 10% test). Downstream fine-tuning used a 5-layer GIN backbone (hidden dimension 300) trained for 100 epochs, reporting test ROC-AUC at the epoch with the highest validation performance.

    Method Tox21 ToxCast Sider ClinTox MUV HIV BBBP Bace Average
    No pretrain 74.6 (0.4) 61.7 (0.5) 58.2 (1.7) 58.4 (6.4) 70.7 (1.8) 75.5 (0.8) 65.7 (3.3) 72.4 (3.8) 67.15
    InfoGraph 73.3 (0.6) 61.8 (0.4) 58.7 (0.6) 75.4 (4.3) 74.4 (1.8) 74.2 (0.9) 68.7 (0.6) 74.3 (2.6) 70.10
    GPT-GNN 74.9 (0.3) 62.5 (0.4) 58.1 (0.3) 58.3 (5.2) 75.9 (2.3) 65.2 (2.1) 64.5 (1.4) 77.9 (3.2) 68.45
    EdgePred 76.0 (0.6) 64.1 (0.6) 60.4 (0.7) 64.1 (3.7) 75.1 (1.2) 76.3 (1.0) 67.3 (2.4) 77.3 (3.5) 70.08
    ContextPred 73.6 (0.3) 62.6 (0.6) 59.7 (1.8) 74.0 (3.4) 72.5 (1.5) 75.6 (1.0) 70.6 (1.5) 78.8 (1.2) 70.93
    GraphLoG 75.0 (0.6) 63.4 (0.6) 59.6 (1.9) 75.7 (2.4) 75.5 (1.6) 76.1 (0.8) 68.7 (1.6) 78.6 (1.0) 71.56
    GraphCL 75.1 (0.7) 63.0 (0.4) 59.8 (1.3) 77.5 (3.8) 76.4 (0.4) 75.1 (0.7) 67.8 (2.4) 74.6 (2.1) 71.16
    GraphMAE 75.2 (0.9) 63.6 (0.3) 60.5 (1.2) 76.5 (3.0) 76.4 (2.0) 76.8 (0.6) 71.2 (1.0) 78.2 (1.5) 72.30
    3D InfoMax 74.5 (0.7) 63.5 (0.8) 56.8 (2.1) 62.7 (3.3) 76.2 (1.4) 76.1 (1.3) 69.1 (1.2) 78.6 (1.9) 69.69
    GraphMVP 74.9 (0.8) 63.1 (0.2) 60.2 (1.1) 79.1 (2.8) 77.7 (0.6) 76.0 (0.1) 70.8 (0.5) 79.3 (1.5) 72.64
    MGSSL 75.2 (0.6) 63.3 (0.5) 61.6 (1.0) 77.1 (4.5) 77.6 (0.4) 75.8 (0.4) 68.8 (0.6) 78.8 (0.9) 72.28
    AttrMask 75.1 (0.9) 63.3 (0.6) 60.5 (0.9) 73.5 (4.3) 75.8 (1.0) 75.3 (1.5) 65.2 (1.4) 77.8 (1.8) 70.81
    MAM 76.2 (0.5) 63.9 (0.3) 61.4 (1.9) 75.1 (3.0) 77.4 (2.1) 77.5 (1.0) 66.8 (1.5) 78.9 (1.1) 72.16
    TMCL 74.9 (0.7) 63.2 (0.7) 59.6 (1.4) 77.0 (4.2) 77.2 (0.3) 75.3 (1.1) 67.6 (1.3) 75.1 (1.2) 71.24
    Mole-BERT 76.8 (0.5) 64.3 (0.2) 62.8 (1.1) 78.9 (3.0) 78.6 (1.8) 78.2 (0.8) 71.9 (1.6) 80.8 (1.4) 74.04

    Mole-BERT attains the highest average ROC-AUC of 74.04%, improving over training from scratch by 6.89% and outperforming 3D-geometry-based pre-training methods (GraphMVP: 72.64%) in a purely 2D data-driven manner.

  7. Knowl 7 — Molecular Property Regression and Drug-Target Affinity (DTA) Performance

    data/table

    Mole-BERT was evaluated on four molecular property regression datasets (scaffold split, reporting RMSE across 3 seeds; lower is better) and two Drug-Target Affinity (DTA) benchmarks (random split, reporting MSE across 3 seeds; lower is better).

    Molecular Property Prediction (RMSE ↓\downarrow) Drug-Target Affinity (MSE ↓\downarrow)
    Dataset ESOL Lipo Malaria CEP Davis KIBA
    No Pre-train 1.178 (0.044) 0.744 (0.007) 1.127 (0.003) 1.254 (0.030) 0.286 (0.006) 0.206 (0.004)
    ContextPred 1.196 (0.037) 0.702 (0.020) 1.101 (0.015) 1.243 (0.025) 0.279 (0.002) 0.198 (0.004)
    JOAO 1.120 (0.019) 0.708 (0.007) 1.145 (0.010) 1.293 (0.003) 0.281 (0.004) 0.196 (0.005)
    GraphMVP 1.064 (0.045) 0.691 (0.013) 1.106 (0.013) 1.228 (0.001) 0.274 (0.002) 0.175 (0.001)
    AttrMask 1.112 (0.048) 0.730 (0.004) 1.119 (0.014) 1.256 (0.000) 0.291 (0.007) 0.203 (0.003)
    MAM 1.098 (0.025) 0.711 (0.010) 1.107 (0.009) 1.240 (0.006) 0.278 (0.005) 0.188 (0.007)
    TMCL 1.116 (0.042) 0.704 (0.014) 1.123 (0.017) 1.262 (0.011) 0.282 (0.005) 0.194 (0.002)
    Mole-BERT 1.015 (0.030) 0.676 (0.017) 1.074 (0.009) 1.232 (0.009) 0.266 (0.004) 0.157 (0.001)

    Mole-BERT outperforms all baselines across all 6 regression benchmarks, demonstrating versatility on both pure graph property regression and heterogeneous molecule-protein interaction modeling.

  8. Knowl 8 — Enhancement of Multi-Task and Supervised Graph Pre-training via MAM

    data/table

    Masked Atoms Modeling (MAM) can serve as a drop-in replacement for the standard node-attribute masking sub-task (AttrMask) in multi-task and supervised pre-training pipelines. Performance is measured using mean ROC-AUC (standard deviation across 10 seeds) on 8 MoleculeNet benchmarks.

    Pre-training Scheme Tox21 ToxCast Sider ClinTox MUV HIV BBBP Bace Average
    MGSSL (AttrMask) 75.2 (0.6) 63.3 (0.5) 61.6 (1.0) 77.1 (4.5) 77.6 (0.4) 75.8 (0.4) 68.8 (0.6) 78.8 (0.9) 72.28
    MGSSL (MAM) 76.6 (0.7) 64.5 (0.9) 62.1 (0.8) 78.2 (3.8) 78.7 (0.5) 76.9 (0.7) 70.5 (1.1) 80.2 (1.5) 73.46
    Supervised 76.8 (0.8) 65.2 (0.5) 61.7 (0.8) 57.0 (2.8) 79.8 (1.6) 74.3 (1.5) 67.9 (0.9) 77.7 (0.8) 70.05
    Supervised + AttrMask 77.8 (0.6) 65.3 (0.8) 63.2 (0.8) 73.8 (3.6) 80.9 (1.6) 77.5 (1.3) 66.8 (1.4) 80.7 (1.3) 73.25
    Supervised + MAM 78.6 (0.5) 66.9 (0.4) 64.0 (1.0) 75.4 (2.9) 81.8 (1.6) 78.8 (1.0) 69.1 (1.7) 82.3 (1.2) 74.61

    Replacing AttrMask with MAM yields consistent gains: MGSSL average ROC-AUC increases from 72.28% to 73.46% (+1.18%), and Supervised pre-training (on ChEMBL) improves from 73.25% with AttrMask to 74.61% with MAM (+1.36%).

  9. Knowl 9 — Diagnosis of Negative Transfer in Graph Attribute Masking (AttrMask)

    empirical result

    The standard AttrMask pretext task directly masks raw atom elements and predicts the 118 elemental types. This setup incurs negative transfer on downstream tasks (e.g., scoring lower than training from scratch on HIV: 75.3% vs 75.5%, and BBBP: 65.2% vs 65.7%) due to two primary causes:

    1. Small vocabulary and task triviality: 118-way elemental classification is too simple, causing pre-training training accuracy to saturate quickly at ∼96%\sim 96\% within 20 epochs (compared to text MLM in BERT with ∼30k\sim 30\text{k} vocabulary reaching only ∼70%\sim 70\% accuracy).
    2. Extreme class imbalance: The element distribution in pre-training data is severely skewed—Carbon accounts for 73.93%, Nitrogen 10.99%, Oxygen 10.82%, Sulfur 1.87%, Fluorine 1.15%, Iodine 0.04%, Phosphorus 0.007%, and Boron 0.003%. This skews network predictions toward dominant elements and limits transferable representation learning.
  10. Knowl 10 — GNN Backbone Agnosticism and Codebook Vocabulary Sensitivity

    empirical result

    Mole-BERT was evaluated across various GNN backbones and tokenizer vocabulary sizes on 8 MoleculeNet datasets (reporting average ROC-AUC):

    1. Backbone Agnosticism: Mole-BERT consistently improves over non-pretrained baselines regardless of the GNN architecture:

      • GIN: 67.15% →\rightarrow 74.04% (+10.26% relative gain)
      • GraphSAGE: 68.46% →\rightarrow 73.74% (+7.71% relative gain)
      • R-GCN: 68.32% →\rightarrow 73.51% (+7.60% relative gain)
      • GCN: 68.77% →\rightarrow 73.22% (+6.47% relative gain)
    2. Vocabulary Size Sensitivity: Varying codebook size ∣A∣|\mathcal{A}| from 128 to 2048 demonstrates that MAM outperforms AttrMask (70.81%) even at size 128 (71.42% for MAM, 73.23% for Mole-BERT). Performance peaks around ∣A∣=512|\mathcal{A}| = 512 (72.16% for MAM, 74.04% for Mole-BERT) and ∣A∣=1024|\mathcal{A}| = 1024 (72.21% for MAM, 74.02% for Mole-BERT), validating 512 as an effective trade-off between capacity and computation.

Coverage note — Omitted qualitative t-SNE embedding plots and the specific chemical structure abbreviations table for functional groups (Table 10), as they represent supplementary visualization and reference data fully encapsulated by the core method and empirical results.

References

  1. 1.Dávid Bajusz, Anita Rácz, and Károly Héberger. Why is tanimoto index an appropriate choice for fingerprint-based similarity calculations? Journal of cheminformatics, 7(1):1–13, 2015.
  2. 2.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
  3. 3.Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  4. 4.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In International conference on machine learning, pp. 1691–1703. PMLR, 2020.
  5. 5.Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: Large-scale self-supervised pretraining for molecular property prediction. CoRR, abs/2010.09885, 2020.
  6. 6.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. ELECTRA: Pre-training text encoders as discriminators rather than generators. In ICLR, 2020. URL https://openreview.net/pdf?id=r1xMH1BtvB.
  7. 7.Jacob Devlin, Ming-Wei Chang, and others. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  9. 9.Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4(2):127–134, 2022a.
  10. 10.Yin Fang, Qiang Zhang, Haihong Yang, Xiang Zhuang, Shumin Deng, Wen Zhang, Ming Qin, Zhuo Chen, Xiaohui Fan, and Huajun Chen. Molecular contrastive learning with chemical element knowledge graph. In Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence (AAAI), 2022b.
  11. 11.Anna Gaulton, Louisa J. Bellis, A. Patrícia Bento, Jon Chambers, Mark Davies, Anne Hersey, Yvonne Light, Shaun McGlinchey, David Michalovich, Bissan Al-Lazikani, and John P. Overington. Chembl: a large-scale bioactivity database for drug discovery. Nucleic Acids Research, 2012.
  12. 12.Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive representation learning on large graphs. Advances in neural information processing systems, 30, 2017.
  13. 13.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000–16009, 2022.
  14. 14.Zhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong, Hongxia Yang, Chunjie Wang, and Jie Tang. Graphmae: Self-supervised masked graph autoencoders. arXiv e-prints, pp. arXiv–2205, 2022.
  15. 15.Bozhen Hu, Jun Xia, Jiangbin Zheng, Cheng Tan, Yufei Huang, Yongjie Xu, and Stan Z Li. Protein language models and structure prediction: Connection and progression. arXiv preprint arXiv:2211.16742, 2022.
  16. 16.Weihua Hu, Bowen Liu, and others. Strategies for pre-training graph neural networks. ICLR, 2020.
  17. 17.Ziniu Hu and others. Gpt-gnn: Generative pre-training of graph neural networks. KDD, 2020.
  18. 18.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77, 2020.
  19. 19.N. Thomas Kipf. and M. Welling. Semi-supervised classification with graph convolutional networks. ICLR, 2017.
  20. 20.Thomas N Kipf and Max Welling. Variational graph auto-encoders. arXiv preprint arXiv:1611.07308, 2016.
  21. 21.Matej Kosec, Sheng Fu, and Mario Michael Krell. Packing: Towards 2x nlp bert acceleration. arXiv preprint arXiv:2107.02027, 2021.
  22. 22.Greg Landrum et al. Rdkit: A software suite for cheminformatics, computational chemistry, and predictive modeling, 2013.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019.
  24. 24.Pengyong Li, Jun Wang, and others. An effective self-supervised framework for learning expressive molecular global representations to drug discovery. BIB, 2021a.
  25. 25.Pengyong Li, Jun Wang, Yixuan Qiao, Hao Chen, Yihuan Yu, Xiaojun Yao, Peng Gao, Guotong Xie, and Sen Song. An effective self-supervised framework for learning expressive molecular global representations to drug discovery. Briefings in Bioinformatics, 22(6):bbab109, 2021b.
  26. 26.Zhaowen Li, Zhiyang Chen, Fan Yang, Wei Li, Yousong Zhu, Chaoyang Zhao, Rui Deng, Liwei Wu, Rui Zhao, Ming Tang, et al. Mst: Masked self-supervised transformer for visual representation. Advances in Neural Information Processing Systems, 34:13165–13176, 2021c.
  27. 27.Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. Pre-training molecular graph representation with 3d geometry. In International Conference on Learning Representations, 2022a. URL https://openreview.net/forum?id=xQUe1pOKPam.
  28. 28.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  29. 29.Yue Liu, Wenxuan Tu, Sihang Zhou, Xinwang Liu, Linxuan Song, Xihong Yang, and En Zhu. Deep graph clustering via dual correlation reduction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 7603–7611, 2022b.
  30. 30.Yue Liu, Jun Xia, Sihang Zhou, Siwei Wang, Xifeng Guo, Xihong Yang, Ke Liang, Wenxuan Tu, Z. Stan Li, and Xinwang Liu. A survey of deep graph clustering: Taxonomy, challenge, and application. arXiv preprint arXiv:2211.12875, 2022c.
  31. 31.Yue Liu, Xihong Yang, Sihang Zhou, Xinwang Liu, Zhen Wang, Ke Liang, Wenxuan Tu, Liang Li, Jingcan Duan, and Cancan Chen. Hard sample aware network for contrastive deep graph clustering. In Proc. of AAAI, 2023.
  32. 32.Diego Mesquita, Amauri Souza, and Samuel Kaski. Rethinking pooling in graph neural networks. Advances in Neural Information Processing Systems, 33:2220–2231, 2020.
  33. 33.Thin Nguyen, Hang Le, Thomas P Quinn, Tri Nguyen, Thuc Duy Le, and Svetha Venkatesh. Graphdta: Predicting drug–target binding affinity with graph neural networks. Bioinformatics, 37 (8):1140–1147, 2021.
  34. 34.Jiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang, Hongxia Yang, Ming Ding, Kuansan Wang, and Jie Tang. Gcc: Graph contrastive coding for graph neural network pre-training. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 1150–1160, 2020a.
  35. 35.Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. Pre-trained models for natural language processing: A survey. Science China Technological Sciences, 63(10): 1872–1897, 2020b.
  36. 36.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  37. 37.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pp. 8821–8831. PMLR, 2021.
  38. 38.Bharath Ramsundar, Peter Eastman, Patrick Walters, and Vijay Pande. Deep learning for the life sciences: applying deep learning to genomics, microscopy, drug discovery, and more. O’Reilly Media, 2019.
  39. 39.Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. Contrastive learning with hard negative samples. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=CR1XOQ0UTh-.
  40. 40.David Rogers and Mathew Hahn. Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754, 2010.
  41. 41.Yu Rong, Yatao Bian, and others. Self-supervised graph transformer on large-scale molecular data. NIPS, 2020a.
  42. 42.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33:12559–12571, 2020b.
  43. 43.Lars Ruddigkeit, Ruud Van Deursen, Lorenz C Blum, and Jean-Louis Reymond. Enumeration of 166 billion organic small molecules in the chemical universe database gdb-17. Journal of chemical information and modeling, 52(11):2864–2875, 2012.
  44. 44.Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In European semantic web conference, pp. 593–607. Springer, 2018.
  45. 45.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909, 2015.
  46. 46.Hannes Stark, Dominique Beaini, Gabriele Corso, Prudencio Tossou, Christian Dallago, Stephan Günnemann, and Pietro Liò. 3d infomax improves gnns for molecular property prediction. In International Conference on Machine Learning, pp. 20479–20502. PMLR, 2022.
  47. 47.T. Sterling and John J. Irwin. Zinc 15 – ligand discovery for everyone. Journal of Chemical Information and Modeling, 55:2324 – 2337, 2015.
  48. 48.Hannes Stark, Dominique Beaini, Gabriele Corso, Prudencio Tossou, Christian Dallago, Stephan Günnemann, and Pietro Liò. 3d infomax improves gnns for molecular property prediction. arXiv preprint arXiv:2110.04126, 2021.
  49. 49.Fan-Yun Sun, Jordan Hoffman, Vikas Verma, and Jian Tang. Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. In International Conference on Learning Representations, 2020a. URL https://openreview.net/forum?id=r1lfF2NYvH.
  50. 50.Fan-Yun Sun, Jordan Hoffmann, and others. Infograph: Unsupervised and semi-supervised graph-level representation learning via mutual information maximization. ICLR, 2020b.
  51. 51.Mengying Sun, Jing Xing, and others. Mocl: Contrastive learning on molecular graphs with multi-level domain knowledge. KDD, 2021.
  52. 52.Susheel Suresh, Pan Li, Cong Hao, and Jennifer Neville. Adversarial graph augmentation to improve graph contrastive learning. Advances in Neural Information Processing Systems, 34, 2021.
  53. 53.Cheng Tan, Jun Xia, Lirong Wu, and Stan Z Li. Co-learning: Learning from noisy labels with self-supervision. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 1405–1413, 2021.
  54. 54.Cheng Tan, Zhangyang Gao, Jun Xia, and Stan Z Li. Generative de novo protein design with global context. arXiv preprint arXiv:2204.10673, 2022.
  55. 55.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017.
  56. 56.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In ICLR (Poster), 2018.
  57. 57.Petar Velickovic, William Fedus, William L Hamilton, Pietro Liò, Yoshua Bengio, and R Devon Hjelm. Deep graph infomax. In ICLR (Poster), 2019.
  58. 58.Sheng Wang, Yuzhi Guo, Yuhong Wang, Hongmao Sun, and Junzhou Huang. SMILES-BERT: large scale unsupervised pre-training for molecular property prediction. In BCB, pp. 429–436. ACM, 2019.
  59. 59.Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning, pp. 9929–9939. PMLR, 2020.
  60. 60.Yifei Wang, Shiyang Chen, Guobin Chen, Ethan Shurberg, Hang Liu, and Pengyu Hong. Motif-based graph representation learning with application to chemical molecules, 2023. URL https://openreview.net/forum?id=70_umOqc6_-.
  61. 61.Yuyang Wang, Jianren Wang, Zhonglin Cao, and Amir Barati Farimani. Molecular contrastive learning of representations via graph neural networks. Nature Machine Intelligence, 4(3):279–287, 2022.
  62. 62.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14668–14678, 2022.
  63. 63.David Weininger, Arthur Weininger, and L. Joseph Weininger. Smiles. 2. algorithm for generation of unique smiles notation. JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES, 1989.
  64. 64.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
  65. 65.Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018.
  66. 66.Jun Xia, Haitao Lin, Yongjie Xu, Lirong Wu, Zhangyang Gao, Siyuan Li, and Stan Z. Li. Towards robust graph neural networks against label noise, 2021. URL https://openreview.net/forum?id=H38f_9b90BO.
  67. 67.Jun Xia, Cheng Tan, Lirong Wu, Yongjie Xu, and Stan Z Li. Ot cleaner: Label correction as optimal transport. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3953–3957. IEEE, 2022a.
  68. 68.Jun Xia, Lirong Wu, , Jintao Chen, Bozhen Hu, and Stan Z. Li. SimGRACE: A Simple Framework for Graph Contrastive Learning without Data Augmentation. In Proceedings of The Web Conference 2022. Association for Computing Machinery, 2022b.
  69. 69.Jun Xia, Lirong Wu, Ge Wang, Jintao Chen, and Stan Z Li. Progcl: Rethinking hard negative mining in graph contrastive learning. In International Conference on Machine Learning, pp. 24332–24346. PMLR, 2022c.
  70. 70.Jun Xia, Jiangbin Zheng, Cheng Tan, Ge Wang, and Stan Z Li. Towards effective and generalizable fine-tuning for pre-trained molecular graph models. bioRxiv, 2022d.
  71. 71.Jun Xia, Yanqiao Zhu, Yuanqi Du, and Stan Z. Li. Pre-training graph neural networks for molecular representations: Retrospect and prospect. In ICML 2022 2nd AI for Science Workshop, 2022e. URL https://openreview.net/forum?id=dhXLkrY2Nj3.
  72. 72.Jun Xia, Yanqiao Zhu, Yuanqi Du, Yue Liu, and Stan Z Li. A systematic survey of molecular pre-trained models. arXiv preprint arXiv:2210.16484, 2022f.
  73. 73.Keyulu Xu, Weihua Hu, and others. How powerful are graph neural networks? In ICLR, 2019.
  74. 74.Minghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo, and Jian Tang. Self-supervised graph-level representation learning with local and global structure. In International Conference on Machine Learning, pp. 11548–11558. PMLR, 2021a.
  75. 75.Minghao Xu, Hang Wang, and others. Self-supervised graph-level representation learning with local and global structure. ICML, 2021b.
  76. 76.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  77. 77.Y. You, T. Chen, and others. Graph contrastive learning with augmentations. In NeurIPS, 2020.
  78. 78.Yuning You, Tianlong Chen, and others. Graph contrastive learning automated. ICML, 2021.
  79. 79.Zaixi Zhang, Qi Liu, and others. Motif-based graph self-supervised learning for molecular property prediction. NeurIPS, 2021.
  80. 80.Jiangbin Zheng, Yile Wang, Ge Wang, Jun Xia, Yufei Huang, Guojiang Zhao, Yue Zhang, and Stan Li. Using context-to-vector with graph retrofitting to improve word embeddings. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8154–8163, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.561. URL https://aclanthology.org/2022.acl-long.561.
  81. 81.Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, and Stan Z Li. Cvt-slr: Contrastive visual-textual transformation for sign language recognition with variational alignment. arXiv preprint arXiv:2303.05725, 2023.
  82. 82.Jinhua Zhu, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. Unified 2d and 3d pre-training of molecular representations. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2626–2636, 2022.
  83. 83.Yanqiao Zhu, Yichen Xu, Qiang Liu, and Shu Wu. An empirical study of graph contrastive learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021a. URL https://openreview.net/forum?id=UuUbIYnHKO.
  84. 84.Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang. Graph contrastive learning with adaptive augmentation. In Proceedings of the Web Conference 2021, pp. 2069–2080, 2021b.

Citation

MLA
Xia, J., et al. “Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules”. [], American Chemical Society (ACS), 2023, https://doi.org/10.26434/chemrxiv-2023-dngg4.
APA
Xia, J., Zhao, C., Hu, B., Gao, Z., Tan, C., Liu, Y., Li, S., & Li, S. Z. (2023). Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules. In []. American Chemical Society (ACS). https://doi.org/10.26434/chemrxiv-2023-dngg4
Chicago
Xia, J., C. Zhao, B. Hu, et al. 2023. “Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules”. In []. American Chemical Society (ACS). https://doi.org/10.26434/chemrxiv-2023-dngg4.
Harvard
Xia, J. et al. (2023) “Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules”, []. American Chemical Society (ACS). Available at: https://doi.org/10.26434/chemrxiv-2023-dngg4.
Vancouver
1. Xia J, Zhao C, Hu B, Gao Z, Tan C, Liu Y, Li S, Li SZ (2023) Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules. []. https://doi.org/10.26434/chemrxiv-2023-dngg4

BibTeX

@article{Xia_2023, title={Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules}, url={http://dx.doi.org/10.26434/chemrxiv-2023-dngg4}, DOI={10.26434/chemrxiv-2023-dngg4}, publisher={American Chemical Society (ACS)}, author={Xia, Jun and Zhao, Chengshuai and Hu, Bozhen and Gao, Zhangyang and Tan, Cheng and Liu, Yue and Li, Siyuan and Li, Stan Z.}, year={2023}, month=Apr }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors