Analyzing Learned Molecular Representations for Property Prediction

Kevin YangKyle SwansonWengong JinConnor ColeyPhilipp EidenHua GaoAngel Guzman-PerezTimothy HopperBrian KelleyMiriam Mathea

article2019Journal of Chemical Information and Modeling1,924 citations

Presents a graph convolutional architecture for molecular property prediction and demonstrates through evaluation across 19 public and 16 proprietary industrial datasets that learned representations consistently outperform traditional fixed molecular descriptors.

Listen

Molecular property prediction is a vital component of modern drug discovery and chemical development. In recent years, computational methods have expanded from traditional approaches relying on human-engineered chemical descriptors to advanced deep learning architectures that learn representations directly from molecular graphs. However, prior research has shown conflicting findings regarding whether learned graph representations outperform fixed descriptors, and these models have rarely undergone rigorous evaluation across realistic industrial pipelines.

This article evaluates the predictive accuracy and generalizability of learned molecular representations compared to traditional fingerprint and descriptor methods across both public benchmarks and proprietary industrial datasets. It also demonstrates a tailored graph convolutional neural network architecture designed to improve property prediction across diverse chemical tasks.

To conduct this evaluation, the authors performed over 850 experiments across 19 public datasets and 16 proprietary industry datasets provided by Amgen, BASF, and Novartis. The investigated model, termed the Directed Message Passing Neural Network (D-MPNN), passes informational messages across directed chemical bonds rather than individual atoms, avoiding unnecessary feedback loops during graph encoding. To address the limitations of small datasets and local representations, the architecture optionally integrates 200 computed global chemical descriptors. Models were rigorously assessed using scaffold-based and chronological data splits to measure how well they generalize to entirely new chemical structures, followed by hyperparameter tuning via Bayesian optimization and model ensembling.

The analysis produced several key findings. First, the D-MPNN architecture matched or outperformed baseline models on 11 of the 19 public datasets and 15 of the 16 proprietary industry datasets, providing consistently superior or competitive performance across varied chemical endpoints. Second, data splitting methodology fundamentally dictates model evaluation: scaffold-based splits serve as an effective, realistic proxy for real-world chronological splits, whereas standard random splits severely overestimate generalization performance. Third, combining learned graph representations with computed global features, Bayesian hyperparameter optimization, and ensembling provided systematic performance gains, yielding dramatic improvements of up to 37% on certain physical datasets. Finally, while the optimized D-MPNN outperforms existing in-house industrial models, all evaluated computational models still fall substantially short of the upper performance bounds set by experimental assay reproducibility.

These findings indicate that learned molecular graph models are robust, highly practical, and ready for integration into industrial discovery pipelines. Relying on scaffold or chronological splits will reduce risk by preventing false confidence in virtual screening performance before committing to costly chemical synthesis. The success of the hybrid feature approach also indicates that domain-specific chemical descriptors still offer valuable regularization, especially when models face limited training data.

Organizations should adopt bond-level graph neural network architectures as strong starting baselines for molecular property prediction and mandate scaffold-based or chronological splitting during model validation. To further enhance predictive accuracy, development teams should routinely apply Bayesian hyperparameter optimization, model ensembling, and relevant global descriptors. However, decision-makers should recognize that model accuracy cannot yet replace laboratory screening, and performance degrades when applied to extremely small datasets (under 1,000 compounds), severe class imbalances (such as datasets with under 1% positive hits), or tasks that strictly depend on three-dimensional molecular geometry. Future work should focus on integrating three-dimensional spatial coordinates, developing effective pretraining strategies on large chemical repositories, and improving model stability under severe data imbalance.

Cover for Analyzing Learned Molecular Representations for Property Prediction

Abstract

Advancements in neural machinery have led to a wide range of algorithmic solutions for molecular property prediction. Two classes of models in particular have yielded promising results: neural networks applied to computed molecular fingerprints or expert-crafted descriptors, and graph convolutional neural networks that construct a learned molecular representation by operating on the graph structure of the molecule. However, recent literature has yet to clearly determine which of these two methods is superior when generalizing to new chemical space. Furthermore, prior research has rarely examined these new models in industry research settings in comparison to existing employed models. In this paper, we benchmark models extensively on 19 public and 16 proprietary industrial datasets spanning a wide variety of chemical endpoints. In addition, we introduce a graph convolutional model that consistently matches or outperforms models using fixed molecular descriptors as well as previous graph neural architectures on both public and proprietary datasets. Our empirical findings indicate that while approaches based on these representations have yet to reach the level of experimental reproducibility, our proposed model nevertheless offers significant improvements over models currently used in industrial workflows.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Methods
  • 3.1 Message Passing Neural Networks
  • 3.2 Directed MPNN
  • 3.3 Initial Featurization
  • 3.4 D-MPNN with Features
  • 3.5 Hyperparameter Optimization
  • 3.6 Ensembling
  • 3.7 Implementation
  • 4 Experiments
  • 4.1 Data
  • 4.2 Experimental Procedure
  • 5 Results and Discussion
  • 5.1 Comparison to Baselines
  • 5.1.1 Comparison to MoleculeNet
  • 5.1.2 Comparison to Mayr et al.Mayr et al. (2018)
  • 5.1.3 Out-of-the-Box Comparison of D-MPNN to Other Baselines
  • 5.2 Proprietary Datasets
  • 5.2.1 Amgen
  • 5.2.2 BASF
  • 5.2.3 Novartis
  • 5.3 Experimental Error
  • 5.4 Analysis of Split Type
  • 5.5 Ablations
  • 5.5.1 Message Type
  • 5.5.2 RDKit Features
  • 5.5.3 Hyperparameter Optimization
  • 5.5.4 Ensembling
  • 5.5.5 Effect of Data Size
  • 6 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Directed Message Passing Neural Network (D-MPNN) Architecture

    model/method

    The Directed Message Passing Neural Network (D-MPNN) operates on an undirected molecular graph G=(V,E)G = (V, E) with atom feature vectors xv∈Rdvx_v \in \mathbb{R}^{d_v} for nodes v∈Vv \in V and bond feature vectors evw∈Rdee_{vw} \in \mathbb{R}^{d_e} for edges (v,w)∈E(v, w) \in E.

    Unlike standard MPNNs that pass messages between vertices, D-MPNN passes messages along directed edges (bonds) to prevent backtracking ("totters"), ensuring that a message along v→wv \to w is not immediately passed back along w→vw \to v in the next step.

    Edge hidden states are initialized at step t=0t=0 by: hvw0=τ(Wi[xv ∥ evw])h^0_{vw} = \tau(W_i [x_v \,\|\, e_{vw}]) where [xv ∥ evw]∈Rdv+de[x_v \,\|\, e_{vw}] \in \mathbb{R}^{d_v + d_e} denotes vector concatenation, Wi∈Rh×(dv+de)W_i \in \mathbb{R}^{h \times (d_v + d_e)} is a learned parameter matrix, hh is the hidden size, and τ\tau is the ReLU activation function.

    For message passing iterations t∈{0,…,T−1}t \in \{0, \dots, T-1\}, messages mvwt+1m^{t+1}_{vw} and updated edge states hvwt+1h^{t+1}_{vw} are computed as: mvwt+1=∑k∈N(v)∖{w}hkvtm^{t+1}_{vw} = \sum_{k \in N(v) \setminus \{w\}} h^t_{kv} hvwt+1=τ(hvw0+Wmmvwt+1)h^{t+1}_{vw} = \tau(h^0_{vw} + W_m m^{t+1}_{vw}) where N(v)N(v) denotes the set of atom neighbors of vv, Wm∈Rh×hW_m \in \mathbb{R}^{h \times h} is a learned weight matrix shared across all steps, and hvw0h^0_{vw} provides a skip connection to the initial edge embedding.

    After TT steps, edge representations are mapped to atom-level representations by aggregating incoming bond features: mv=∑w∈N(v)hwvTm_v = \sum_{w \in N(v)} h^T_{wv} hv=τ(Wa[xv ∥ mv])h_v = \tau(W_a [x_v \,\|\, m_v]) where Wa∈Rh×(dv+h)W_a \in \mathbb{R}^{h \times (d_v + h)} is a learned matrix.

    The global molecular representation vector h∈Rhh \in \mathbb{R}^h is computed by summing atom representations across the graph: h=∑v∈Vhvh = \sum_{v \in V} h_v

    Finally, property predictions y^\hat{y} are generated via a feed-forward neural network ff: y^=f(h)\hat{y} = f(h).

  2. Knowl 2 — Hybrid Molecular Representation Combining Graph Convolutions with CDF-Normalized Descriptors

    model/method

    To improve property prediction on small datasets (where graph neural networks risk overfitting) and to incorporate global chemical context beyond the local message-passing diameter T<diam(G)T < \text{diam}(G), a hybrid molecular representation combines the D-MPNN learned graph embedding with 200 calculated global 2D molecular descriptors computed using RDKit.

    To prevent features with large numeric ranges from dominating or suffering from outlier bias during training, all 200 raw continuous and count descriptors are normalized by mapping them through empirical Cumulative Density Functions (CDFs) fitted to a large chemical reference library (100,000 compounds from an internal screening catalog). The CDF transform bounds each feature to [0,1][0, 1], where each value represents the population percentile observed below that raw feature value.

    Let h∈Rhh \in \mathbb{R}^h be the graph-level embedding produced by the D-MPNN readout phase, and let hf∈[0,1]200h_f \in [0, 1]^{200} be the vector of CDF-normalized global molecular descriptors. The readout network computes property predictions via: y^=f([h ∥ hf])\hat{y} = f([h \,\|\, h_f]) where [h ∥ hf][h \,\|\, h_f] is the concatenation of the learned and engineered representations, and ff is a feed-forward neural network.

  3. Knowl 3 — Initial Atom and Bond Featurization for D-MPNN

    definition

    The D-MPNN featurizes molecular graphs using fixed one-hot and numerical descriptors computed via RDKit:

    1. Atom Features (xvx_v, total size 127):
    • Atom type: One-hot vector over 100 chemical elements by atomic number (size 100).
    • Number of bonds: One-hot vector for degree {0,1,2,3,4,5}\{0, 1, 2, 3, 4, 5\} (size 6).
    • Formal charge: Integer electronic charge assigned to the atom (size 5).
    • Chirality: One-hot vector for unspecified, tetrahedral clockwise (CW), tetrahedral counter-clockwise (CCW), or other (size 4).
    • Number of bonded hydrogens: One-hot vector for count {0,1,2,3,4}\{0, 1, 2, 3, 4\} (size 5).
    • Hybridization: One-hot vector for spsp, sp2sp^2, sp3sp^3, sp3dsp^3d, or sp3d2sp^3d^2 (size 5).
    • Aromaticity: Binary indicator of aromatic system membership (size 1).
    • Atomic mass: Mass of the atom divided by 100 as a scaled real number (size 1).
    1. Bond Features (evwe_{vw}, total size 12):
    • Bond type: One-hot vector for single, double, triple, or aromatic (size 4).
    • Conjugation: Binary indicator for conjugation (size 1).
    • Ring membership: Binary indicator of whether the bond is in a ring (size 1).
    • Stereochemistry: One-hot vector for None, any, E/Z, or cis/trans (size 6).
  4. Knowl 4 — Scaffold Splitting as an Evaluation Proxy for Prospective Chronological Generalization

    empirical result

    In drug discovery, machine learning models are applied prospectively to predict properties of newly synthesized chemical series, making chronological (temporal) splits the ideal evaluation setup. However, when synthesis dates are unavailable, standard random splits fail as an evaluation protocol because train and test sets share 70% to 80% of the same molecular scaffolds, causing models that memorize scaffolds to appear artificially accurate.

    Scaffold-based splitting groups molecules by their Bemis-Murcko scaffolds. Scaffolds with sizes exceeding half the desired test set size are assigned to the training set to preserve test scaffold diversity, while the remaining scaffolds are randomly assigned to train, validation, and test subsets (e.g., 80:10:10 ratio) ensuring 0% scaffold overlap between train and test sets.

    Empirical evaluations across industrial datasets (Amgen ADME and physical chemistry endpoints) and public datasets (such as PDBbind) demonstrate that scaffold splits provide a substantially better proxy for prospective chronological performance than random splits. Both scaffold and chronological splits result in significantly higher test errors than random splits due to the distribution shift across chemical space. Furthermore, on certain industrial datasets, chronological splits prove even more challenging than scaffold splits, showing that temporal splits remain the preferred benchmark when time annotations are available.

  5. Knowl 5 — Comparative Performance of D-MPNN Across Public Benchmark Datasets

    data/table

    Across 19 public benchmarks spanning quantum mechanics, physical chemistry, biophysics, and physiology evaluated under scaffold splits, D-MPNN consistently matches or outperforms baseline models using fixed fingerprints or prior graph neural network architectures.

    Baseline D-MPNN is better D-MPNN is same D-MPNN is worse # Datasets
    MoleculeNet best 5 3 2 10
    Mayr et al. FFN 8 10 1 19
    Random Forest (Morgan) 9 1 4 15
    FFN (Morgan binary) 14 5 0 19
    FFN (Morgan counts) 15 4 0 19
    FFN (RDKit descriptors) 8 5 4 19

    Statistical significance is evaluated at p<0.05p < 0.05 using paired one-sided Wilcoxon signed-rank tests across test molecules/subsets and Welch's t-tests across splits. D-MPNN matches or outperforms all baselines on 11 of 19 public datasets: QM7, QM8, QM9, ESOL, FreeSolv, Lipophilicity, BBBP, PDBbind-F, PCBA, Tox21, and ClinTox. On the remaining datasets, no single baseline consistently outperforms D-MPNN. Baseline models only outperform D-MPNN when models exploit 3D structural coordinates (e.g., PDBbind-C) or on datasets characterized by extreme label sparsity (e.g., MUV, with 0.2% active labels).

  6. Knowl 6 — Transferability and Performance of D-MPNN on Proprietary Industrial Datasets

    empirical result

    D-MPNN evaluated on 16 proprietary industrial datasets across three pharmaceutical and chemical companies demonstrates that learned graph representations transfer effectively to real-world industrial discovery pipelines, outperforming baseline models on 15 of the 16 datasets:

    1. Amgen Datasets (5 tasks evaluated on chronological splits): On rat plasma protein binding free fraction (rPPB), intrinsic clearance in rat liver microsomes (RLM), solubility (Sol HCL, Sol PBS, Sol SIF), and human pregnane X receptor activation (hPXR regression and classification), D-MPNN outperforms or matches Random Forest, feed-forward networks on Morgan fingerprints/descriptors, and the Mayr et al. baseline on 4 of the 5 datasets.

    2. BASF Datasets (10 quantum mechanical datasets with 13 property targets each, across different solvents on scaffold splits): D-MPNN and its ensemble achieve higher R2R^2 correlation across all 10 solvent datasets than feed-forward networks trained on Morgan binary fingerprints, Morgan count fingerprints, or RDKit descriptors.

    3. Novartis Dataset (logP regression on a chronological split of 20,294 compounds): D-MPNN achieves lower RMSE than all fingerprint-based and descriptor-based baselines.

  7. Knowl 7 — Experimental Assay Reproducibility Ceiling in Molecular Property Prediction

    empirical result

    Benchmarking deep learning models against the agreement between repeated runs of the same physical assay (experimental error ceiling) on industrial ADME and physical chemistry endpoints (Amgen rPPB, RLM, solubility, and hPXR datasets evaluated on chronological splits) reveals a substantial gap between machine learning performance and assay reproducibility.

    While D-MPNN consistently outperforms industry-standard models using expert-crafted descriptors, both machine learning approaches achieve R2R^2 values ranging from 0.250.25 to 0.700.70, falling well below the experimental reproducibility bound of R2≈0.60R^2 \approx 0.60 to 0.850.85 observed for replicate wet-lab assays. This demonstrates that significant predictive headroom remains for molecular property prediction architectures.

  8. Knowl 8 — Ablation of Message Passing Centering: Directed Bonds vs Undirected Bonds vs Atoms

    empirical result

    Comparison of three message passing schemes within the MPNN framework across public regression and classification datasets reveals the impact of edge-directed representations:

    1. Directed bond message passing (D-MPNN): Updates messages along directed bonds v→wv \to w, excluding the reverse incoming bond w→vw \to v from the aggregation step, which prevents immediate two-step cyclic message looping (totters).
    2. Undirected bond message passing: Exchanges messages over undirected edges without directional distinction.
    3. Atom-centered message passing: Updates messages on nodes by aggregating information across all adjacent vertices.

    Directed bond message passing achieves lower average errors on regression tasks and higher average AUC on classification tasks compared to both undirected bond and atom-centered message passing. However, when tested individually per dataset, the performance differences between the three message passing formulations are largely not statistically significant (p≥0.05p \ge 0.05).

  9. Knowl 9 — Sensitivity of Learned vs Fixed Molecular Representations to Dataset Size

    empirical result

    The relative superiority of learned molecular representations (graph convolutions) versus fixed representations (expert-engineered fingerprints and descriptors) depends heavily on the available training sample size:

    1. Data-sparse regimes (N≤1,000N \le 1,000 molecules or tasks with fewer than 300 positive labels): Fixed fingerprints (such as Morgan fingerprints) combined with feed-forward networks or random forests match or exceed the performance of pure graph convolutional networks, as graph neural networks overfit due to insufficient training signals to learn molecular features from scratch.

    2. Data-rich regimes (N≥5,000N \ge 5,000 to 10,00010,000 molecules): In multi-target datasets such as ChEMBL filtered at increasing molecule count thresholds per target (300, 1,000, 5,000, and 10,000 molecules), D-MPNN's ROC-AUC systematically surpasses fixed-descriptor feed-forward networks (e.g., the Mayr et al. baseline), confirming that graph convolutional representations scale more effectively with dataset volume.

  10. Knowl 10 — Bayesian Hyperparameter Optimization and Ensembling for D-MPNN

    model/method

    The predictive performance of D-MPNN is optimized via two complementary procedures:

    1. Bayesian Hyperparameter Optimization: The message passing depth TT (number of message steps), hidden state dimension hh, number of feed-forward readout layers, and dropout probability are optimized using 20 iterations of Bayesian optimization via Hyperopt over cross-validation splits. Hyperparameter tuning improves performance across nearly all datasets, yielding a 2% to 5% typical performance gain and up to a 37% error reduction on quantum chemistry benchmarks (such as QM9).

    2. Model Ensembling: Combining five independently initialized D-MPNN models (trained with different random weight initializations) via unweighted prediction averaging yields an additional 1% to 5% improvement in accuracy and error metrics across both regression and classification datasets.

Coverage note — Omitted minor implementation details regarding the specific web demo and individual per-dataset hyperparameter configurations from the supplementary material, as they are engineering artifacts rather than core scientific contributions.

References

  1. 1.Duvenaud, D. K.; Maclaurin, D.; Iparraguirre, J.; Bombarell, R.; Hirzel, T.; Aspuru-Guzik, A.; Adams, R. P. Convolutional Networks on Graphs for Learning Molecular Fingerprints. Advances in Neural Information Processing Systems 2015, 2224–2232.
  2. 2.Wu, Z.; Ramsundar, B.; Feinberg, E.; Gomes, J.; Geniesse, C.; Pappu, A. S.; Leswing, K.; Pande, V. MoleculeNet: A Benchmark for Molecular Machine Learning. Chem. Sci. 2018, 9, 513530.
  3. 3.Kearnes, S.; McCloskey, K.; Berndl, M.; Pande, V.; Riley, P. Molecular Graph Convolutions: Moving Beyond Fingerprints. J. Comput.-Aided Mol. Des. 2016, 30, 595–608.
  4. 4.Gilmer, J.; Schoenholz, S. S.; Riley, P. F.; Vinyals, O.; Dahl, G. E. Neural Message Passing for Quantum Chemistry. Proceedings of the 34th International Conference on Machine Learning 2017,
  5. 5.Li, Y.; Tarlow, D.; Brockschmidt, M.; Zemel, R. Gated Graph Sequence Neural Networks. arXiv preprint arXiv:1511.05493 2015,
  6. 6.Kipf, T. N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. arXiv preprint arXiv:1609.02907 2016,
  7. 7.Defferrard, M.; Bresson, X.; Vandergheynst, P. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. Advances in Neural Information Processing Systems 2016, 3844–3852.
  8. 8.Bruna, J.; Zaremba, W.; Szlam, A.; LeCun, Y. Spectral Networks and Locally Connected Networks on Graphs. arXiv preprint arXiv:1312.6203 2013,
  9. 9.Coley, C. W.; Barzilay, R.; Green, W. H.; Jaakkola, T. S.; Jensen, K. F. Convolutional Embedding of Attributed Molecular Graphs for Physical Property Prediction. J. Chem. Inf. Model. 57, 1757–1772.
  10. 10.Schütt, K. T.; Arbabzadah, F.; Chmiela, S.; Müller, K. R.; Tkatchenko, A. Quantum-Chemical Insights from Deep Tensor Neural Networks. Nat. Commun. 2017, 8, 13890.
  11. 11.Battaglia, P.; Pascanu, R.; Lai, M.; Rezende, D. J. Interaction Networks for Learning about Objects, Relations and Physics. Advances in Neural Information Processing Systems 2016, 4502–4510.
  12. 12.Mayr, A.; Klambauer, G.; Unterthiner, T.; Steijaert, M.; Wegner, J. K.; Ceulemans, H.; Clevert, D.-A.; Hochreiter, S. Large-Scale Comparison of Machine Learning Methods for Drug Target Prediction on ChEMBL. Chem. Sci. 2018, 9, 5441–5451.
  13. 13.Sheridan, R. P. Time-Split Cross-Validation as a Method for Estimating the Goodness of Prospective Prediction. J. Chem. Inf. Model. 2013, 53, 783–790.
  14. 14.Shahriari, B.; Swersky, K.; Wang, Z.; Adams, R. P.; de Freitas, N. Taking the Human Out of the Loop: A Review of Bayesian Optimization. Proceedings of the IEEE 2016, 104, 148–175.
  15. 15.Dietterich, T. G. Ensemble Methods in Machine Learning. International Workshop on Multiple Classifier Systems 2000, 1–15.
  16. 16.Hamilton, W.; Ying, Z.; Leskovec, J. Inductive Representation Learning on Large Graphs. Advances in Neural Information Processing Systems 2017, 1024–1034.
  17. 17.Scarselli, F.; Gori, M.; Tsoi, A. C.; Hagenbuchner, M.; Monfardini, G. The Graph Neural Network Model. IEEE Transactions on Neural Networks 2009, 20, 61–80.
  18. 18.Henaff, M.; Bruna, J.; LeCun, Y. Deep Convolutional Networks on Graph-Structured Data. arXiv preprint arXiv:1506.05163 2015,
  19. 19.Dai, H.; Dai, B.; Song, L. Discriminative Embeddings of Latent Variable Models for Structured Data. International Conference on Machine Learning 2016, 2702–2711.
  20. 20.Lei, T.; Jin, W.; Barzilay, R.; Jaakkola, T. Deriving Neural Architectures from Sequence and Graph Kernels. arXiv preprint arXiv:1705.09037 2017,
  21. 21.Kusner, M. J.; Paige, B.; Hernández-Lobato, J. M. Grammar Variational Autoencoder. arXiv preprint arXiv:1703.01925 2017,
  22. 22.Gómez-Bombarelli, R.; Wei, J. N.; Duvenaud, D.; Hernández-Lobato, J. M.; Sánchez-Lengeling, B.; Sheberla, D.; Aguilera-Iparraguirre, J.; Hirzel, T. D.; Adams, R. P.; Aspuru-Guzik, A. Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Cent. Sci. 2018, 4, 268–276.
  23. 23.Jin, W.; Barzilay, R.; Jaakkola, T. Junction Tree Variational Autoencoder for Molecular Graph Generation. arXiv preprint arXiv:1802.04364 2018,
  24. 24.Jin, W.; Yang, K.; Barzilay, R.; Jaakkola, T. Learning Multimodal Graph-to-Graph Translation for Molecular Optimization. arXiv preprint arXiv:1812.01070 2018,
  25. 25.Cortes, C.; Vapnik, V. Support Vector Machine. Machine Learning 1995, 20, 273–297.
  26. 26.Breiman, L. Random Forests. Machine Learning 2001, 45, 5–32.
  27. 27.Mauri, A.; Consonni, V.; Pavan, M.; Todeschini, R. Dragon Software: An Easy Approach to Molecular Descriptor Calculations. Match 2006, 56, 237–248.
  28. 28.Rogers, D.; Hahn, M. Extended-Connectivity Fingerprints. J. Chem. Inf. Model. 2010, 50, 742–754.
  29. 29.Swamidass, S. J.; Chen, J.; Bruand, J.; Phung, P.; Ralaivola, L.; Baldi, P. Kernels for Small Molecules and the Prediction of Mutagenicity, Toxicity and Anti-Cancer Activity. Bioinformatics 2005, 21, i359–i368.
  30. 30.Cao, D.-S.; Xu, Q.-S.; Hu, Q.-N.; Liang, Y.-Z. ChemoPy: Freely Available Python Package for Computational Biology and Chemoinformatics. Bioinformatics 2013, 29, 1092–1094.
  31. 31.Durant, J. L.; Leland, B. A.; Henry, D. R.; Nourse, J. G. Reoptimization of MDL Keys for Use in Drug Discovery. J. Chem. Inf. Model. 2002, 42, 1273–1280.
  32. 32.Moriwaki, H.; Tian, Y.-S.; Kawashita, N.; Takagi, T. Mordred: A Molecular Descriptor Calculator. J. Cheminf. 2018, 10, 4.
  33. 33.Schütt, K.; Kindermans, P.-J.; Felix, H. E. S.; Chmiela, S.; Tkatchenko, A.; Müller, K.-R. SchNet: A Continuous-Filter Convolutional Neural Network for Modeling Quantum Interactions. Advances in Neural Information Processing Systems 2017, 991–1001.
  34. 34.Kondor, R.; Son, H. T.; Pan, H.; Anderson, B.; Trivedi, S. Covariant Compositional Networks for Learning Graphs. arXiv preprint arXiv:1801.02144 2018,
  35. 35.Faber, F. A.; Hutchison, L.; Huang, B.; Gilmer, J.; Schoenholz, S. S.; Dahl, G. E.; Vinyals, O.; Kearnes, S.; Riley, P. F.; von Lilienfeld, O. A. Machine Learning Prediction Errors Better than DFT Accuracy. arXiv preprint arXiv:1702.05532 2017,
  36. 36.Feinberg, E. N.; Sur, D.; Wu, Z.; Husic, B. E.; Mai, H.; Li, Y.; Sun, S.; Yang, J.; Ramsundar, B.; Pande, V. S. PotentialNet for Molecular Property Prediction. ACS Cent. Sci. 2018, 4, 1520–1530.
  37. 37.Lee, A. A.; Yang, Q.; Bassyouni, A.; Butler, C. R.; Hou, X.; Jenkinson, S.; Price, D. A. Ligand Biological Activity Predicted by Cleaning Positive and Negative Chemical Correlations. Proc. Natl. Acad. Sci. U. S. A. 2019,
  38. 38.Weininger, D. SMILES, A Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules. J. Chem. Inf. Model. 1988, 28, 31–36.
  39. 39.Ishiguro, K.; Maeda, S.-i.; Koyama, M. Graph Warp Module: An Auxiliary Module for Boosting the Power of Graph Neural Networks. arXiv preprint arXiv:1902.01020 2019,
  40. 40.Liu, K.; Sun, X.; Jia, L.; Ma, J.; Xing, H.; Wu, J.; Gao, H.; Sun, Y.; Boulnois, F.; Fan, J. Chemi-net: A Graph Convolutional Network for Accurate Drug Property Prediction. arXiv preprint arXiv:1803.06236 2018,
  41. 41.Mahé, P.; Ueda, N.; Akutsu, T.; Perret, J.-L.; Vert, J.-P. Extensions of Marginalized Graph Kernels. Proceedings of the Twenty-First International Conference on Machine learning 2004, 70.
  42. 42.Koller, D.; Friedman, N.; Bach, F. Probabilistic Graphical Models: Principles and Techniques; MIT Press 2009,
  43. 43.Nair, V.; Hinton, G. E. Rectified Linear Units Improve Restricted Boltzmann Machines. Proceedings of the 27th International Conference on Machine Learning 2010,
  44. 44.Landrum, G. RDKit: Open-Source Cheminformatics. 2006, https://rdkit.org/docs/index.html, last visited 2019-05-24.
  45. 45.Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 2011, 12, 2825–2830.
  46. 46.Irwin, J. J.; Shoichet, B. K. ZINC–A Free Database of Commercially Available Compounds for Virtual Screening. J. Chem. Inf. Model. 2005, 45, 177–182.
  47. 47.Kim, S.; Thiessen, P. A.; Bolton, E. E.; Chen, J.; Fu, G.; Gindulyte, A.; Han, L.; He, J.; He, S.; Shoemaker, B. A. PubChem Substance and Compound Databases. Nucleic Acids Res. 2015, 44, D1202–D1213.
  48. 48.Sartor, M. A.; Leikauf, G. D.; Medvedovic, M. LRpath: A Logistic Regression Approach for Identifying Enriched Biological Groups in Gene Expression Data. Bioinformatics 2008, 25, 211–217.
  49. 49.Distributed Asynchronous Hyperparameter Optimization in Python. https://github.com/hyperopt/hyperopt, last visited 2019-05-24.
  50. 50.Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; Lerer, A. Automatic differentiation in PyTorch. 31st Conference on Neural Information Processing Systems 2017,
  51. 51.Message Passing Neural Networks for Molecule Property Prediction. https://github.com/swansonk14/chemprop, last visited 2019-05-24.
  52. 52.Descriptor computation(chemistry) and (optional) storage for machine learning. https://github.com/bp-kelley/descriptastorus, last visited 2019-05-24.
  53. 53.Chemprop Machine Learning for Molecular Property Prediction. http://chemprop.csail.mit.edu, last visited 2019-05-24.
  54. 54.Friedman, J. H. Greedy Function Approximation: A Gradient Boosting Machine. Annals of Statistics 2001, 1189–1232.
  55. 55.Cramer, J. S. Logit Models from Economics and Other Fields; 2003,
  56. 56.Lusci, A.; Pollastri, G.; Baldi, P. Deep Architectures and Deep Learning in Chemoinformatics: The Prediction of Aqueous Solubility for Drug-like Molecules. J. Chem. Inf. Model. 2013, 53, 1563–1575.
  57. 57.Cortes, C.; Vapnik, V. Support-Vector Networks. Machine Learning 1995, 273–297.
  58. 58.Ma, J.; Sheridan, R. P.; Liaw, A.; Dahl, G. E.; Svetnik, V. Deep Neural Nets as a Method for Quantitative Structure–Activity Relationships. J. Chem. Inf. Model. 2015, 55, 263–274.
  59. 59.Ramsundar, B.; Liu, B.; Wu, Z.; Verras, A.; Tudor, M.; Sheridan, R. P.; Pande, V. Is Multitask Deep Learning Practical for Pharma? J. Chem. Inf. Model. 2017, 57, 2068–2076.
  60. 60.Swamidass, S. J.; Azencott, C.-A.; Lin, T.-W.; Gramajo, H.; Tsai, S.-C.; Baldi, P. Influence Relevance Voting: An Accurate and Interpretable Virtual High Throughput Screening Method. J. Chem. Inf. Model. 2009, 49, 756–766.
  61. 61.Smith, J. S.; Isayev, O.; Roitberg, A. E. ANI-1: An Extensible Neural Network Potential with DFT Accuracy at Force Field Computational Cost. Chem. Sci. 2017, 8, 3192–3203.
  62. 62.Ramsundar, B.; Eastman, P.; Leswing, K.; Walters, P.; Pande, V. Deep Learning for the Life Sciences; O’Reilly Media 2019,
  63. 63.Scripts for running lsc model on other datasets. https://github.com/yangkevin2/lsc_experiments, last visited 2019-05-24.
  64. 64.Navarin, N.; Tran, D. V.; Sperduti, A. Pre-training Graph Neural Networks with Kernels. arXiv preprint arXiv:1811.06930 2018,
  65. 65.Goh, G. B.; Siegel, C.; Vishnu, A.; Hodas, N. Using Rule-Based Labels for Weak Supervised Learning: A ChemNet for Transferable Chemical Property Prediction. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining 2018, 302–310.

Citation

MLA
Yang, K., et al. “Analyzing Learned Molecular Representations for Property Prediction”. Journal of Chemical Information and Modeling 59.8 (2019): 3370-3388, 2019, http://arxiv.org/abs/1904.01561v5.
APA
Yang, K., Swanson, K., Jin, W., Coley, C., Eiden, P., Gao, H., Guzman-Perez, A., Hopper, T., Kelley, B., Mathea, M., Palmer, A., Settels, V., Jaakkola, T., Jensen, K., & Barzilay, R. (2019). Analyzing Learned Molecular Representations for Property Prediction. Journal of Chemical Information and Modeling 59.8 (2019): 3370-3388. http://arxiv.org/abs/1904.01561v5
Chicago
Yang, K., K. Swanson, W. Jin, et al. 2019. “Analyzing Learned Molecular Representations for Property Prediction”. Journal of Chemical Information and Modeling 59.8 (2019): 3370-3388. http://arxiv.org/abs/1904.01561v5.
Harvard
Yang, K. et al. (2019) “Analyzing Learned Molecular Representations for Property Prediction”, Journal of chemical information and modeling 59.8 (2019): 3370-3388 [Preprint]. Available at: http://arxiv.org/abs/1904.01561v5.
Vancouver
1. Yang K, Swanson K, Jin W, et al (2019) Analyzing Learned Molecular Representations for Property Prediction. Journal of chemical information and modeling 59.8 (2019): 3370-3388

BibTeX

@article{yang2019analyzing,
  title = {Analyzing Learned Molecular Representations for Property Prediction},
  author = {Yang, Kevin and Swanson, Kyle and Jin, Wengong and Coley, Connor and Eiden, Philipp and Gao, Hua and Guzman-Perez, Angel and Hopper, Timothy and Kelley, Brian and Mathea, Miriam and Palmer, Andrew and Settels, Volker and Jaakkola, Tommi and Jensen, Klavs and Barzilay, Regina},
  year = {2019},
  journal = {Journal of chemical information and modeling 59.8 (2019): 3370-3388},
  url = {http://arxiv.org/abs/1904.01561v5},
  eprint = {1904.01561}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF