Analyzing Learned Molecular Representations for Property Prediction
Kevin YangKyle SwansonWengong JinConnor ColeyPhilipp EidenHua GaoAngel Guzman-PerezTimothy HopperBrian KelleyMiriam Mathea
Presents a graph convolutional architecture for molecular property prediction and demonstrates through evaluation across 19 public and 16 proprietary industrial datasets that learned representations consistently outperform traditional fixed molecular descriptors.
Molecular property prediction is a vital component of modern drug discovery and chemical development. In recent years, computational methods have expanded from traditional approaches relying on human-engineered chemical descriptors to advanced deep learning architectures that learn representations directly from molecular graphs. However, prior research has shown conflicting findings regarding whether learned graph representations outperform fixed descriptors, and these models have rarely undergone rigorous evaluation across realistic industrial pipelines.
This article evaluates the predictive accuracy and generalizability of learned molecular representations compared to traditional fingerprint and descriptor methods across both public benchmarks and proprietary industrial datasets. It also demonstrates a tailored graph convolutional neural network architecture designed to improve property prediction across diverse chemical tasks.
To conduct this evaluation, the authors performed over 850 experiments across 19 public datasets and 16 proprietary industry datasets provided by Amgen, BASF, and Novartis. The investigated model, termed the Directed Message Passing Neural Network (D-MPNN), passes informational messages across directed chemical bonds rather than individual atoms, avoiding unnecessary feedback loops during graph encoding. To address the limitations of small datasets and local representations, the architecture optionally integrates 200 computed global chemical descriptors. Models were rigorously assessed using scaffold-based and chronological data splits to measure how well they generalize to entirely new chemical structures, followed by hyperparameter tuning via Bayesian optimization and model ensembling.
The analysis produced several key findings. First, the D-MPNN architecture matched or outperformed baseline models on 11 of the 19 public datasets and 15 of the 16 proprietary industry datasets, providing consistently superior or competitive performance across varied chemical endpoints. Second, data splitting methodology fundamentally dictates model evaluation: scaffold-based splits serve as an effective, realistic proxy for real-world chronological splits, whereas standard random splits severely overestimate generalization performance. Third, combining learned graph representations with computed global features, Bayesian hyperparameter optimization, and ensembling provided systematic performance gains, yielding dramatic improvements of up to 37% on certain physical datasets. Finally, while the optimized D-MPNN outperforms existing in-house industrial models, all evaluated computational models still fall substantially short of the upper performance bounds set by experimental assay reproducibility.
These findings indicate that learned molecular graph models are robust, highly practical, and ready for integration into industrial discovery pipelines. Relying on scaffold or chronological splits will reduce risk by preventing false confidence in virtual screening performance before committing to costly chemical synthesis. The success of the hybrid feature approach also indicates that domain-specific chemical descriptors still offer valuable regularization, especially when models face limited training data.
Organizations should adopt bond-level graph neural network architectures as strong starting baselines for molecular property prediction and mandate scaffold-based or chronological splitting during model validation. To further enhance predictive accuracy, development teams should routinely apply Bayesian hyperparameter optimization, model ensembling, and relevant global descriptors. However, decision-makers should recognize that model accuracy cannot yet replace laboratory screening, and performance degrades when applied to extremely small datasets (under 1,000 compounds), severe class imbalances (such as datasets with under 1% positive hits), or tasks that strictly depend on three-dimensional molecular geometry. Future work should focus on integrating three-dimensional spatial coordinates, developing effective pretraining strategies on large chemical repositories, and improving model stability under severe data imbalance.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). It establishes the standardized MoleculeNet benchmark datasets, splits, and baseline evaluation protocols directly analyzed and expanded upon by the source paper.
- Paper: Neural Message Passing for Quantum Chemistry, Justin Gilmer et al. (2017). It formulates Message Passing Neural Networks (MPNNs) for molecular property prediction, providing the foundational architecture that the source study modifies and benchmarks.
- Paper: Convolutional Networks on Graphs for Learning Molecular Fingerprints, David Duvenaud et al. (2015). It introduces end-to-end differentiable neural graph fingerprints, providing the primary conceptual precursor to modern graph convolutional neural networks in chemoinformatics.
- Paper: How Powerful are Graph Neural Networks?, Keyulu Xu et al. (2019). It provides the theoretical framework connecting message-passing graph neural networks with the Weisfeiler-Lehman test, establishing expressiveness baselines evaluated in learned molecular representations.
- Paper: Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks, Christopher Morris et al. (2019). It establishes theoretical equivalence between standard graph neural networks and the 1-WL isomorphism test on molecular benchmarks like QM9.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). It outlines the relational inductive biases and unified graph network formulation foundational to modern molecular message-passing architectures.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). It formalizes first-order graph convolutional networks, which serve as a primary architecture compared against fixed molecular descriptors.
- Paper: Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules, R. Gómez-Bombarelli et al. (2016). It pioneers data-driven continuous molecular representations for property prediction and generation, motivating the study of learned versus hand-crafted molecular features.
- Paper: Strategies for Pre-training Graph Neural Networks, Weihua Hu et al. (2020). It develops self-supervised node- and graph-level pre-training strategies to improve out-of-distribution generalization for molecular property prediction.
- Paper: Open Graph Benchmark: Datasets for Machine Learning on Graphs, Weihua Hu et al. (2020). It expands upon molecular property benchmarking by establishing standardized large-scale datasets and realistic scaffold splits across machine learning on graphs.
- Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). It designs graph contrastive learning strategies that enhance representation transfer and performance on downstream biochemical property prediction tasks.
- Paper: Do Transformers Really Perform Bad for Graph Representation?, Chengxuan Ying et al. (2021). It extends molecular representation learning beyond traditional GNNs by adapting Transformer architectures with structural encodings for large-scale property prediction.
- Paper: GNNExplainer: Generating Explanations for Graph Neural Networks, Rex Ying et al. (2019). It introduces an explainability method that identifies salient subgraphs and features driving graph neural network predictions on molecular datasets.
- Paper: E(n) Equivariant Graph Neural Networks, Victor Garcia Satorras et al. (2021). It advances molecular modeling by incorporating 3D Euclidean equivariance directly into graph neural network message-passing for physical and chemical property prediction.
