Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization

Wenhao GaoTianfan FuJimeng SunConnor W. Coley

article2022NeurIPS178 citations

Establishes a standardized open-source benchmark evaluating 25 molecular design algorithms across 23 tasks under realistic oracle budgets, revealing that many state-of-the-art methods fail to outperform simpler predecessors when sample efficiency is strictly constrained.

Listen

Designing new molecules computationally is critical for accelerating drug discovery and materials design. While many artificial intelligence methods have emerged to generate candidate structures, real-world discovery relies on expensive physical experiments or high-accuracy simulations to test each candidate. Consequently, computational algorithms must be sample-efficient, discovering high-performing molecules with as few evaluation queries as possible. Prior literature frequently overlooked this constraint, reporting performance over unconstrained budgets, using trivial benchmarks, or neglecting the substantial run-to-run variation inherent to non-deterministic algorithms.

The article introduces the Practical Molecular Optimization benchmark to evaluate how effectively and efficiently diverse molecular optimization methods perform under a realistic evaluation budget. Specifically, the study assesses the optimization capability, sample efficiency, robustness, and generalizability of 25 molecular design algorithms across 23 standardized target tasks relevant to therapeutics.

To establish a rigorous and fair comparison, the authors limited each algorithm to an evaluation budget of 10,000 oracle queries. Performance was measured by calculating the area under the curve of the top-10 average objective score over time, a metric that rewards algorithms that discover top candidates earlier. The benchmark encompassed string-based, graph-based, and chemical synthesis-based molecular assembly strategies paired with diverse optimization techniques, including genetic algorithms, reinforcement learning, and Bayesian optimization. All algorithms were tuned on standardized tasks and evaluated over five independent trials to account for non-deterministic behavior.

The findings show that no existing algorithm is sample-efficient enough to optimize complex, de novo molecular targets within a realistic experimental budget of only hundreds of evaluations. Furthermore, established older methods such as REINVENT and Graph GA consistently outperformed newer deep learning approaches across the benchmark. Robust representations like SELFIES did not provide a general performance advantage over standard SMILES strings for modern language models, except in specific genetic algorithm implementations. In addition, model-based methods using surrogate predictors only improved sample efficiency when the surrogate model was carefully calibrated, sometimes underperforming simpler model-free counterparts if the surrogate misled the search. Finally, optimization performance depended heavily on the problem landscape, with string-based evolutionary algorithms performing best on atomic composition targets and other methods excelling on structural similarity tasks.

These findings indicate that the molecular design field risks misallocating resources toward increasingly complex architectures that do not improve practical discovery performance. For organizations investing in drug discovery workflows, using unvetted state-of-the-art algorithms without budget constraints risks generating unstable or low-quality candidates while exhausting experimental resources. Simpler, established algorithms currently offer better reliability and efficiency when tuned correctly.

For future molecular design research and deployment, decision-makers and researchers should enforce strict evaluation budgets, evaluate algorithms across multiple diverse target landscapes, perform extensive task-specific hyperparameter tuning, and report performance distributions from multiple independent trials rather than single best runs.

The authors note limitations in the study, including the inability to exhaustively test every algorithmic variant or hyperparameter combination, a potential evaluation bias toward similarity-based target functions, and the exclusion of multi-objective synthesizability constraints and physical docking simulations. Nevertheless, the benchmark provides a high level of confidence in its comparative conclusions by enforcing reproducible, standardized constraints across all evaluated methods.

Cover for Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization

Abstract

Molecular optimization is a fundamental goal in the chemical sciences and is of central interest to drug and material design. In recent years, significant progress has been made in solving challenging problems across various aspects of computational molecular optimizations, emphasizing high validity, diversity, and, most recently, synthesizability. Despite this progress, many papers report results on trivial or self-designed tasks, bringing additional challenges to directly assessing the performance of new methods. Moreover, the sample efficiency of the optimization—the number of molecules evaluated by the oracle—is rarely discussed, despite being an essential consideration for realistic discovery applications.

To fill this gap, we have created an open-source benchmark for practical molecular optimization, PMO, to facilitate the transparent and reproducible evaluation of algorithmic advances in molecular optimization. This paper thoroughly investigates the performance of 25 molecular design algorithms on 23 single-objective (scalar) optimization tasks with a particular focus on sample efficiency. Our results show that most “state-of-the-art” methods fail to outperform their predecessors under a limited oracle budget allowing 10K queries and that no existing algorithm can efficiently solve certain molecular optimization problems in this setting. We analyze the influence of the optimization algorithm choices, molecular assembly strategies, and oracle landscapes on the optimization performance to inform future algorithm development and benchmarking. PMO provides a standardized experimental setup to comprehensively evaluate and compare new molecule optimization methods with existing ones. All code can be found at https://github.com/wenhao-gao/mol_opt.

Table of Contents

  • 1 Introduction
  • 2 Algorithms
  • 2.1 Preliminaries
  • 2.2 Molecular assembly strategies
  • 2.3 Optimization algorithms
  • 3 Experiments
  • 3.1 Benchmark setup
  • 3.2 Results & Analysis
  • 4 Conclusions
  • Acknowledgments and Disclosure of Funding
  • Reproducibility Statement
  • References
  • Checklist

Knowls

  1. Knowl 1 — PMO standardized benchmark for practical molecular optimization

    model/method

    The paper introduces PMO, an open-source benchmark designed to compare molecular optimization algorithms under a controlled oracle budget. PMO evaluates 25 representative molecular design methods on 23 single-objective scalar molecular-property tasks, tunes methods under a common protocol, runs each method with five independent random seeds, and measures both optimization quality and the number of oracle evaluations required. The benchmark caps each run at 10,000 oracle calls and is intended to expose whether an algorithm is useful when property evaluations are expensive. Code, parameters, and releasable results are publicly released at https://github.com/wenhao-gao/mol_opt.

  2. Knowl 2 — Black-box molecular optimization objective and sample-efficiency metric

    equation

    For a molecular structure mm in a chemical space M\mathcal{M} and a scalar black-box oracle O:M→RO:\mathcal{M}\rightarrow\mathbb{R}, the benchmark formulates molecular design as

    m⋆=arg⁡max⁡m∈MO(m).m^\star=\arg\max_{m\in\mathcal{M}}O(m).

    The oracle returns the ground-truth property value for a queried molecule; its analytic form and derivatives are unavailable. PMO evaluates a run using the area under the curve of the average oracle value of the best K=10K=10 molecules found so far versus the number of oracle calls. This AUC Top-10 metric rewards both high final molecular quality and reaching that quality with fewer queries. All reported AUC values are min–max scaled to [0,1][0,1], and every method is evaluated within a maximum budget of 10,000 oracle calls.

  3. Knowl 3 — Coverage of molecular representations and optimization algorithms

    model/method

    PMO separates a molecular optimization method into a molecular assembly strategy and an optimization algorithm. Assembly strategies represent or construct molecules as SMILES strings, SELFIES strings, atom-level molecular graphs, fragment-level molecular graphs, or synthesis pathways. The ranked 25-method benchmark includes Screening, MolPAL, REINVENT, SELFIES-REINVENT, MolDQN, Graph MCTS, SMILES GA, STONED, Graph GA, SynNet, GP BO, SMILES-LSTM-HC, SELFIES-LSTM-HC, MIMOSA, DoG-Gen, Pasithea, DST, SMILES-VAE, SELFIES-VAE, JT-VAE, DoG-AE, MARS, GFlowNet, GFlowNet-AL, and GA+D. These methods cover random and model-based screening, genetic algorithms, Monte Carlo tree search, Bayesian optimization, variational autoencoders, score-based models, hill climbing, reinforcement learning, and gradient-ascent approaches. BOSS and ChemBO were also attempted as Bayesian-optimization baselines but were early-stopped after failing to produce meaningful results, partly because of poor scaling of the string-subsequence kernel.

  4. Knowl 4 — Oracle suite and controlled experimental protocol

    experimental setup

    The benchmark uses 23 scalar oracles: qed, drd2, gsk3b, jnk3, albuterol_similarity, amlodipine_mpo, celecoxib_rediscovery, deco_hop, fexofenadine_mpo, isomers_c7h8n2o2, isomers_c9h10n2o2pf2cl, median1, median2, mestranol_similarity, osimertinib_mpo, perindopril_mpo, ranolazine_mpo, scaffold_hop, sitagliptin_mpo, thiothixene_rediscovery, troglitazone_rediscovery, valsartan_smarts, and zaleplon_mpo. QED is a drug-likeness heuristic; DRD2, GSK3β, and JNK3 are machine-learning bioactivity predictors; and the remaining tasks are GuacaMol objectives involving similarity, molecular properties, isomer-specific atomic contributions, or substructure constraints. Oracle scores are normalized so that 1 is optimal. Whenever a database is required, methods are restricted to the approximately 250,000-molecule ZINC 250K dataset; generative models are pretrained on this dataset and required fragments are extracted from it. Hyperparameters are tuned using three independent runs on zaleplon_mpo and perindopril_mpo, while final results use five independent runs. Docking-based oracles are excluded because they are more costly while still providing only coarse binding-affinity estimates.

  5. Knowl 5 — Overall benchmark ranking favors REINVENT and Graph GA

    data/table

    The ten highest-ranked methods are determined by the sum of their task-wise mean AUC Top-10 scores over all 23 tasks. The scores and ranks are:

    • REINVENT: 14.196, rank 1
    • Graph GA: 13.751, rank 2
    • SELFIES-REINVENT: 13.471, rank 3
    • GP BO: 13.156, rank 4
    • STONED: 13.024, rank 5
    • SMILES-LSTM-HC: 12.223, rank 6
    • SMILES GA: 12.054, rank 7
    • SynNet: 11.498, rank 8
    • DoG-Gen: 11.456, rank 9
    • DST: 10.989, rank 10

    Across six aggregate metrics—AUC Top-1, AUC Top-10, AUC Top-100, Top-1, Top-10, and Top-100—the corresponding mean ranks for these ten methods are 1.00, 2.33, 3.16, 4.83, 5.66, 4.66, 8.00, 10.00, 7.50, and 9.66, respectively, when ordered as REINVENT, Graph GA, SELFIES-REINVENT, GP BO, STONED, SMILES-LSTM-HC, SMILES GA, SynNet, DoG-Gen, and DST. The results show that older methods, especially REINVENT and Graph GA, outperform many newer methods under the benchmark’s limited-query setting.

  6. Knowl 6 — Most methods remain inefficient at realistic oracle budgets

    empirical result

    Under the PMO protocol, none of the evaluated methods reliably optimizes even the benchmark’s relatively simple objectives within only a few hundred oracle calls, except for near-trivial oracles such as QED, DRD2, and osimertinib_mpo. Some tasks remain difficult even with the full 10,000-query budget. AUC Top-10 changes the apparent comparison between methods because it penalizes slow improvement: SMILES-LSTM-HC can eventually reach performance comparable to Graph GA but requires more oracle evaluations, whereas REINVENT reaches high performance substantially earlier. Methods that construct molecules token-by-token or atom-by-atom from a single starting point, including GA+D, MolDQN, and Graph MCTS, are especially data-inefficient because many generated candidates are undesirable, unstable, or unsynthesizable and therefore consume oracle queries without useful progress.

  7. Knowl 7 — SELFIES generally does not improve optimization over SMILES

    empirical result

    Head-to-head comparisons of corresponding string-based methods show that most SELFIES variants do not outperform their SMILES counterparts in optimization quality or sample efficiency. SELFIES exceeded the SMILES variant on only 4 of 23 tasks for the LSTM hill-climbing pair, 6 of 23 tasks for the VAE-based pair, and 3 of 23 tasks for the REINVENT pair. The advantage of SELFIES in guaranteeing syntactic validity therefore does not translate into a general optimization advantage for modern language-model-based methods. SELFIES-based genetic algorithms did perform better than the evaluated SMILES-based genetic algorithm, but this was not a strictly representation-only comparison because genetic-algorithm mutation and crossover rules differed. The authors further observed that many SELFIES token combinations can collapse to a small set of valid molecules, so syntactic validity alone does not guarantee broad chemical-space exploration.

  8. Knowl 8 — Model-based optimization helps only when the predictive model and search loops are well designed

    empirical result

    The benchmark supports a qualified advantage for model-based search. MolPAL, which trains a predictive model to prioritize a library, outperformed random Screening on 22 of the 23 tasks. However, adding a predictive model to a generative optimizer was not consistently beneficial: GP BO outperformed Graph GA on 12 tasks, yet Graph GA had the higher aggregate score, and GFlowNet outperformed its active-learning variant GFlowNet-AL on almost every task. These comparisons indicate that model-based optimization can improve sample efficiency, but errors in the predictive model or poorly designed inner-loop and outer-loop optimization can direct the search away from promising molecules, especially early in a run when little oracle data are available.

  9. Knowl 9 — Oracle landscape determines which optimization strategy works best

    empirical result

    Clustering oracle tasks by relative AUC Top-10 reveals distinct landscape classes with different method preferences. String-based genetic algorithms such as SMILES GA and STONED perform particularly well on isomer-type objectives: isomers_c7h8n2o2, isomers_c9h10n2o2pf2cl, sitagliptin_mpo, and zaleplon_mpo. These objectives are based on sums of atomic contributions, whereas most other multi-property objectives are dominated by fingerprint similarity. Similarity-based tasks involving logP and topological polar surface area, such as fexofenadine_mpo and osimertinib_mpo, form a related cluster, while rediscovery and median objectives form another; machine-learning bioactivity oracles also tend to align with the similarity-based cluster. QED is so easy that nearly all methods obtain similar values, while substructure-oriented tasks such as deco_hop, valsartan_smarts, and scaffold_hop show more varied method performance. Thus, no single optimizer is uniformly best, and the benchmark does not establish which oracle landscape most closely represents a true pharmaceutical design objective.

  10. Knowl 10 — Hyperparameter tuning, independent replicates, and benchmark limitations

    limitation

    PMO shows that default hyperparameters from an algorithm’s original publication are not reliably optimal in a limited-oracle-budget setting. For example, the best value found for a key REINVENT hyperparameter was much larger than the value recommended in its original study. Randomness is also substantial: Graph GA and GP BO exhibit high run-to-run variance because of their exploratory behavior. Consequently, molecular-optimization studies should retune hyperparameters when the task or evaluation environment changes, run multiple independent replicates, and report the distribution of outcomes; for costly experimental oracles, worst-case performance may be more relevant than the mean. The benchmark is limited by incomplete coverage of methods and hyperparameter configurations, the possibility that its representative implementations are not best-in-class, potential bias toward similarity-based oracles, and limited analysis of synthesizability and diversity. Its oracle budget is counted from scratch, including data used to train surrogate models, so conclusions may differ when a previously collected dataset is available for surrogate pretraining.

Coverage note — Detailed per-method architectures, hardware settings, appendix-only analyses, and the complete 23-by-25 score matrix were omitted because they support reproducibility but do not add separate load-bearing contributions beyond the benchmark protocol, rankings, comparative analyses, and stated limitations.

References

  1. 1.Jan H Jensen. A graph-based genetic algorithm and generative model/monte carlo tree search for the exploration of chemical space. Chemical science, 10(12):3567–3572, 2019.
  2. 2.Yutong Xie, Chence Shi, Hao Zhou, Yuwei Yang, Weinan Zhang, Yong Yu, and Lei Li. MARS: Markov molecular sampling for multi-objective drug discovery. In ICLR, 2021.
  3. 3.David E Graff, Eugene I Shakhnovich, and Connor W Coley. Accelerating high-throughput virtual screening through molecular pool-based active learning. Chemical science, 12(22):7866–7881, 2021.
  4. 4.Francesco Gentile, Jean Charle Yaacoub, James Gleave, Michael Fernandez, Anh-Tien Ton, Fuqiang Ban, Abraham Stern, and Artem Cherkasov. Artificial intelligence–enabled virtual screening of ultra-large chemical libraries with deep docking. Nature Protocols, pages 1–26, 2022.
  5. 5.Marcus Olivecrona, Thomas Blaschke, Ola Engkvist, and Hongming Chen. Molecular de-novo design through deep reinforcement learning. Journal of cheminformatics, 9(1):1–14, 2017.
  6. 6.Rafael Gómez-Bombarelli, Jennifer N Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D Hirzel, Ryan P Adams, and Alán Aspuru-Guzik. Automatic chemical design using a data-driven continuous representation of molecules. ACS central science, 2018.
  7. 7.Matt J Kusner, Brooks Paige, and José Miguel Hernández-Lobato. Grammar variational autoencoder. In International Conference on Machine Learning, pages 1945–1954. PMLR, 2017.
  8. 8.Wengong Jin, Regina Barzilay, and Tommi Jaakkola. Junction tree variational autoencoder for molecular graph generation. ICML, 2018.
  9. 9.Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024, 2020.
  10. 10.Yoshua Bengio, Tristan Deleu, Edward J. Hu, Salem Lahlou, Mo Tiwari, and Emmanuel Bengio. GFlowNet foundations. CoRR, abs/2111.09266, 2021.
  11. 11.John Bradshaw, Brooks Paige, Matt J Kusner, Marwin Segler, and José Miguel Hernández-Lobato. Barking up the right tree: an approach to search over molecule synthesis dags. Advances in Neural Information Processing Systems, 33:6852–6866, 2020.
  12. 12.Wenhao Gao, Rocío Mercado, and Connor W Coley. Amortized tree generation for bottom-up synthesis planning and synthesizable molecular design. International Conference on Learning Representations, 2022.
  13. 13.Nathan Brown, Marco Fiscato, Marwin HS Segler, and Alain C Vaucher. GuacaMol: benchmarking models for de novo molecular design. Journal of chemical information and modeling, 59(3):1096–1108, 2019.
  14. 14.Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. Therapeutics data commons: Machine learning datasets and tasks for therapeutics. NeurIPS Track Datasets and Benchmarks, 2021.
  15. 15.Austin Tripp, Gregor NC Simm, and José Miguel Hernández-Lobato. A fresh look at de novo molecular design benchmarks. In NeurIPS 2021 AI for Science Workshop, 2021.
  16. 16.Zhenpeng Zhou, Steven Kearnes, Li Li, Richard N Zare, and Patrick Riley. Optimization of molecules via deep reinforcement learning. Scientific reports, 9(1):1–10, 2019.
  17. 17.AkshatKumar Nigam, Pascal Friederich, Mario Krenn, and Alán Aspuru-Guzik. Augmenting genetic algorithms with deep neural networks for exploring the chemical space. In ICLR, 2020.
  18. 18.Sai Krishna Gottipati, Boris Sattarov, Sufeng Niu, Yashaswi Pathak, Haoran Wei, Shengchao Liu, Simon Blackburn, Karam Thomas, Connor Coley, Jian Tang, et al. Learning to navigate the synthetically accessible chemical space using reinforcement learning. In International Conference on Machine Learning, pages 3668–3679. PMLR, 2020.
  19. 19.Ksenia Korovina, Sailun Xu, Kirthevasan Kandasamy, Willie Neiswanger, Barnabas Poczos, Jeff Schneider, and Eric Xing. ChemBO: Bayesian optimization of small organic molecules with synthesizable recommendations. In International Conference on Artificial Intelligence and Statistics, pages 3393–3403. PMLR, 2020.
  20. 20.Tianfan Fu, Wenhao Gao, Cao Xiao, Jacob Yasonik, Connor W Coley, and Jimeng Sun. Differentiable scaffolding tree for molecular optimization. International Conference on Learning Representations, 2022.
  21. 21.Emmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup, and Yoshua Bengio. Flow network based generative models for non-iterative diverse candidate generation. Advances in Neural Information Processing Systems, 34, 2021.
  22. 22.Natalie Maus, Haydn T Jones, Juston S Moore, Matt J Kusner, John Bradshaw, and Jacob R Gardner. Local latent space bayesian optimization over structured inputs. arXiv preprint arXiv:2201.11872, 2022.
  23. 23.Antoine Grosnit, Rasul Tutunov, Alexandre Max Maraval, Ryan-Rhys Griffiths, Alexander I Cowen-Rivers, Lin Yang, Lin Zhu, Wenlong Lyu, Zhitang Chen, Jun Wang, et al. High-dimensional Bayesian optimisation with variational autoencoders and deep metric learning. arXiv preprint arXiv:2106.03609, 2021.
  24. 24.G Richard Bickerton, Gaia V Paolini, Jérémy Besnard, Sorel Muresan, and Andrew L Hopkins. Quantifying the chemical beauty of drugs. Nature chemistry, 4(2):90, 2012.
  25. 25.Jiaxuan You, Bowen Liu, Zhitao Ying, Vijay Pande, and Jure Leskovec. Graph convolutional policy network for goal-directed molecular graph generation. Advances in neural information processing systems, 31, 2018.
  26. 26.Regine S Bohacek, Colin McMartin, and Wayne C Guida. The art and practice of structure-based drug design: a molecular modeling perspective. Medicinal research reviews, 16(1):3–50, 1996.
  27. 27.Brian Goldman, Steven Kearnes, Trevor Kramer, Patrick Riley, and W Patrick Walters. Defining levels of automated chemical design. Journal of Medicinal Chemistry, 2022.
  28. 28.Wenhao Gao, Priyanka Raghavan, and Connor W Coley. Autonomous platforms for data-driven organic synthesis. Nature Communications, 13(1):1–4, 2022.
  29. 29.AkshatKumar Nigam, Robert Pollice, Mario Krenn, Gabriel dos Passos Gomes, and Alan Aspuru-Guzik. Beyond generative models: superfast traversal, optimization, novelty, exploration and discovery (STONED) algorithm for molecules using SELFIES. Chemical science, 12(20):7079–7090, 2021.
  30. 30.Henry Moss, David Leslie, Daniel Beck, Javier Gonzalez, and Paul Rayson. BOSS: Bayesian optimization over string spaces. Advances in neural information processing systems, 33:15476–15486, 2020.
  31. 31.Benjamin Sanchez-Lengeling, Carlos Outeiral, Gabriel L Guimaraes, and Alan Aspuru-Guzik. Optimizing distributions over molecular space. an objective-reinforced generative adversarial network for inverse-design chemistry (ORGANIC). 2017.
  32. 32.Nicola De Cao and Thomas Kipf. MolGAN: An implicit generative model for small molecular graphs. arXiv preprint arXiv:1805.11973, 2018.
  33. 33.Tianfan Fu, Cao Xiao, Xinhao Li, Lucas M Glass, and Jimeng Sun. MIMOSA: Multi-constraint molecule sampling for molecule optimization. AAAI, 2021.
  34. 34.Wengong Jin, Regina Barzilay, and Tommi Jaakkola. Multi-objective molecule generation using interpretable substructures. In International Conference on Machine Learning, pages 4849–4859. PMLR, 2020.
  35. 35.Soojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu, and Sung Ju Hwang. Hit and lead discovery with explorative RL and fragment-based molecule generation. Advances in Neural Information Processing Systems, 34, 2021.
  36. 36.Julien Horwood and Emmanuel Noutahi. Molecular design in synthetically accessible chemical space via deep reinforcement learning. ACS omega, 5(51):32984–32994, 2020.
  37. 37.Cynthia Shen, Mario Krenn, Sagi Eppel, and Alan Aspuru-Guzik. Deep molecular dreaming: Inverse machine learning for de-novo molecular design and interpretability with surjective representations. Machine Learning: Science and Technology, 2021.
  38. 38.David Weininger. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36, 1988.
  39. 39.Jakob Lykke Andersen, Christoph Flamm, Daniel Merkle, and Peter F Stadler. Chemical graph transformation with stereo-information. In International Conference on Graph Transformation, pages 54–69. Springer, 2017.
  40. 40.Lagnajit Pattanaik, Octavian-Eugen Ganea, Ian Coley, Klavs F Jensen, William H Green, and Connor W Coley. Message passing networks for molecules with tetrahedral chirality. arXiv preprint arXiv:2012.00094, 2020.
  41. 41.Keir Adams, Lagnajit Pattanaik, and Connor W Coley. Learning 3D representations of molecular chirality with invariance to bond rotations. International Conference on Learning Representations, 2022.
  42. 42.Teague Sterling and John J Irwin. ZINC 15–ligand discovery for everyone. Journal of chemical information and modeling, 55(11):2324–2337, 2015.
  43. 43.Fredrik Svensson, Ulf Norinder, and Andreas Bender. Improving screening efficiency through iterative screening using docking and conformal prediction. Journal of chemical information and modeling, 57(3):439–444, 2017.
  44. 44.José Miguel Hernández-Lobato, James Requeima, Edward O Pyzer-Knapp, and Alán Aspuru-Guzik. Parallel and distributed Thompson sampling for large-scale accelerated exploration of chemical space. In International conference on machine learning, pages 1470–1479. PMLR, 2017.
  45. 45.Laeeq Ahmed, Valentin Georgiev, Marco Capuccini, Salman Toor, Wesley Schaal, Erwin Laure, and Ola Spjuth. Efficient iterative virtual screening with Apache Spark and conformal prediction. Journal of cheminformatics, 10(1):1–8, 2018.
  46. 46.Francesco Gentile, Vibudh Agrawal, Michael Hsing, Anh-Tien Ton, Fuqiang Ban, Ulf Norinder, Martin E Gleave, and Artem Cherkasov. Deep docking: A deep learning platform for augmentation of structure based drug discovery. ACS central science, 6(6):939–949, 2020.
  47. 47.David E Graff, Matteo Aldeghi, Joseph A Morrone, Kirk E Jordan, Edward O Pyzer-Knapp, and Connor W Coley. Self-focusing virtual screening with active design space pruning. arXiv preprint arXiv:2205.01753, 2022.
  48. 48.Naruki Yoshikawa, Kei Terayama, Masato Sumita, Teruki Homma, Kenta Oono, and Koji Tsuda. Population-based de novo molecule generation, using grammatical evolution. Chemistry Letters, 47(11):1431–1434, 2018.
  49. 49.Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas. Taking the human out of the loop: A review of Bayesian optimization. Proceedings of the IEEE, 104(1):148–175, 2015.
  50. 50.Marc Deisenroth and Jun Wei Ng. Distributed Gaussian processes. In International Conference on Machine Learning, pages 1481–1490. PMLR, 2015.
  51. 51.Diederik P Kingma and Max Welling. Auto-encoding variational Bayes. International Conference on Learning Representations (ICLR), 2014.
  52. 52.Maximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton, Benjamin Letham, Andrew Gordon Wilson, and Eytan Bakshy. BoTorch: A Framework for Efficient Monte-Carlo Bayesian Optimization. Advances in neural information processing systems, 33, 2020.
  53. 53.Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, et al. Molecular sets (MOSES): A benchmarking platform for molecular generation models. Frontiers in pharmacology, 2020.
  54. 54.Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. A tutorial on the cross-entropy method. Annals of operations research, 134(1):19–67, 2005.
  55. 55.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229–256, 1992.
  56. 56.Austin Tripp, Wenlin Chen, and José Miguel Hernández-Lobato. An evaluation framework for the objective functions of de novo drug design benchmarks. In ICLR2022 Machine Learning for Drug Discovery, 2022.
  57. 57.Yibo Li, Liangren Zhang, and Zhenming Liu. Multi-objective de novo drug design with conditional graph generative model. Journal of cheminformatics, 10(1):1–24, 2018.
  58. 58.Tobiasz Cieplinski, Tomasz Danel, Sabina Podlewska, and Stanislaw Jastrzebski. We should at least be able to design molecules that dock well. arXiv preprint arXiv:2006.16955, 2020.
  59. 59.Miguel García-Ortegón, Gregor NC Simm, Austin J Tripp, José Miguel Hernández-Lobato, Andreas Bender, and Sergio Bacallado. DOCKSTRING: Easy molecular docking yields better benchmarks for ligand design. Journal of Chemical Information and Modeling, 2021.
  60. 60.Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba. Benchmarking model-based reinforcement learning. arXiv preprint arXiv:1907.02057, 2019.
  61. 61.Wenhao Gao and Connor W Coley. The synthesizability of molecules proposed by generative models. Journal of chemical information and modeling, 60(12):5714–5723, 2020.
  62. 62.Michalis Titsias. Variational learning of inducing variables in sparse Gaussian processes. In Artificial intelligence and statistics, pages 567–574. PMLR, 2009.
  63. 63.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  64. 64.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representations, 2014.
  65. 65.Aditya R Thawani, Ryan-Rhys Griffiths, Arian Jamasb, Anthony Bourached, Penelope Jones, William McCorkindale, Alexander A Aldrick, and Alpha A Lee. The photoswitch dataset: a molecular machine learning benchmark for the advancement of synthetic chemistry. arXiv preprint arXiv:2008.03226, 2020.
  66. 66.Alice Capecchi, Daniel Probst, and Jean-Louis Reymond. One molecular fingerprint to rule them all: drugs, biomolecules, and the metabolome. Journal of cheminformatics, 12(1):1–15, 2020.
  67. 67.Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A Hunter, Costas Bekas, and Alpha A Lee. Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction. ACS central science, 5(9):1572–1583, 2019.
  68. 68.Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. Gated graph sequence neural networks. arXiv preprint arXiv:1511.05493, 2015.
  69. 69.Gabriel Lima Guimaraes, Benjamin Sanchez-Lengeling, Carlos Outeiral, Pedro Luis Cunha Farias, and Alán Aspuru-Guzik. Objective-reinforced generative adversarial networks (ORGAN) for sequence generation models. arXiv preprint arXiv:1705.10843, 2017.
  70. 70.Jasper Snoek, Hugo Larochelle, and Ryan P Adams. Practical Bayesian optimization of machine learning algorithms. Advances in neural information processing systems, 25, 2012.
  71. 71.Choon Hui Teo and S.V.N. Vishwanathan. Fast and space efficient string kernels using suffix arrays. In Proceedings of the 23rd international conference on Machine learning, pages 929–936, 2006.
  72. 72.Lukas Biewald. Experiment tracking with weights and biases, 2020. Software available from wandb.com.

Citation

MLA
Gao, W., et al. “Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization”. arXiv, 2022, http://arxiv.org/abs/2206.12411v2.
APA
Gao, W., Fu, T., Sun, J., & Coley, C. W. (2022). Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization. arXiv. http://arxiv.org/abs/2206.12411v2
Chicago
Gao, W., T. Fu, J. Sun, and C. W. Coley. 2022. “Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization”. arXiv. http://arxiv.org/abs/2206.12411v2.
Harvard
Gao, W. et al. (2022) “Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.12411v2.
Vancouver
1. Gao W, Fu T, Sun J, Coley CW (2022) Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization. arXiv

BibTeX

@article{gao2022sample,
  title = {Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization},
  author = {Gao, Wenhao and Fu, Tianfan and Sun, Jimeng and Coley, Connor W.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.12411v2},
  eprint = {2206.12411}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors