Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules
Jun XiaChengshuai ZhaoBozhen HuZhangyang GaoCheng TanYue LiuSiyuan LiStan Z. Li
Computational modeling of molecular graphs using graph neural networks is essential for accelerating drug discovery and chemical property prediction. To make use of vast amounts of unlabeled molecular data, standard pre-training methods typically mask and predict raw atom types. However, this common setup frequently leads to negative transfer, where pre-trained models actually perform worse on downstream tasks than models trained entirely from scratch. This failure happens because raw chemical elements form a very small and heavily imbalanced vocabulary—dominated by carbon—making the training task overly simplistic and prone to bias.
The article introduces and evaluates Mole-BERT, an end-to-end, fully data-driven pre-training framework designed to resolve these representation bottlenecks. Mole-BERT replaces raw atom identities with a context-aware tokenizer based on a grouped vector-quantized variational autoencoder, which maps atoms and their chemical neighborhoods into a richer, balanced discrete codebook. It couples this node-level pre-training task, called Masked Atoms Modeling, with a graph-level pre-training task called Triplet Masked Contrastive Learning, which uses multiple masking ratios to capture fine-grained chemical similarity relationships.
The researchers pre-trained the system on two million unlabeled molecules from the ZINC database and benchmarked it across eight standard molecular classification datasets, four property regression datasets, and two drug-target affinity prediction benchmarks. In extensive comparative evaluations, the key findings demonstrate that:
- Mole-BERT achieved an average classification score of 74.04%, outperforming training from scratch (67.15%) by 6.89 percentage points and surpassing top competing methods, including those relying on complex three-dimensional geometric data.
- The context-aware node-level task completely eliminated the negative transfer failures seen in standard masking methods on sensitive benchmarks such as HIV and blood-brain barrier penetration datasets.
- When evaluated across multiple model backbones, Mole-BERT delivered consistent performance improvements, generating relative gains of 6.5% to 10.3% across various graph neural network architectures.
- In molecule retrieval and drug-target binding affinity tasks, the framework produced lower error rates and retrieved reference molecules that aligned substantially better with established chemical similarity metrics.
These findings indicate that machine learning models in drug discovery do not necessarily require expensive manual annotations or complex spatial geometry datasets to achieve state-of-the-art accuracy. By properly framing pre-training around context-aware tokenization and graded similarity learning, organizations can improve predictive reliability while lowering data preparation overhead and reducing the risk of model failure during lead optimization.
For practical implementation, the authors recommend adopting the context-aware tokenizer as a standard, off-the-shelf component in molecular graph pipelines and combining masked atom modeling with graph-level contrastive tasks. Future work should pilot this framework on related biological problems, such as protein sequence and structure modeling, where vocabulary imbalances also present significant modeling challenges. Users should note that optimal performance depends on tuning vocabulary sizes and masking ratios, and empirical testing remains necessary when transferring the model to entirely new chemical domains.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). It introduces standard masked node attribute pre-training on molecular graphs and analyzes negative transfer, which Mole-BERT directly addresses and improves upon with context-aware tokenization.
- Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). It establishes graph contrastive learning through augmentations and masking, providing the foundational contrastive principles utilized in Mole-BERT's graph-level pre-training task.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). It demonstrates discrete visual tokenization via vector quantization for masked transformer modeling, inspiring Mole-BERT's context-aware VQ-VAE tokenizer for molecular graphs.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). It defines the foundational masked language modeling paradigm that Mole-BERT adapts into masked atom modeling for molecular representation learning.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). It provides the canonical MoleculeNet benchmark suite and splitting protocols used to evaluate Mole-BERT across downstream molecular property tasks.
- Paper: Do Transformers Really Perform Badly for Graph Representation?, Chengxuan Ying et al. (2021). It presents Graphormer and essential structural encodings for graph transformers, which serve as foundational backbones benchmarked in Mole-BERT.
- Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). It establishes the standard neural message passing framework for learning molecular properties on 2D chemical graphs.
- Paper: DeepDTA: deep drug–target binding affinity prediction, Hakime Öztürk et al. (2018). It provides the standard formulation and benchmark datasets (Davis and KIBA) used by Mole-BERT for evaluating drug-target binding affinity prediction.
- Paper: Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs, Rui Jiao et al. (2023). Extends molecular graph pretraining from 2D topological graphs into 3D geometric structures and interatomic forces using equivariant energy-based modeling.
- Paper: DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations, Yuanfeng Ji et al. (2023). Builds upon molecular representation learning and affinity prediction models like Mole-BERT to establish realistic out-of-distribution evaluation benchmarks with noisy annotations.
- Paper: A Generalization of ViT/MLP-Mixer to Graphs, Xiaoxin He et al. (2023). Generalizes transformer and mixer architectures to graph domains using sub-graph patches, providing an alternative scalable backbone for molecular graph modeling.
- Paper: Protein Representation Learning by Geometric Structure Pretraining, Zuobai Zhang et al. (2023). Explores self-supervised geometric representation learning on biomolecular structures, extending graph pre-training principles from small molecules to full protein targets.
