Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization
Wenhao GaoTianfan FuJimeng SunConnor W. Coley
Establishes a standardized open-source benchmark evaluating 25 molecular design algorithms across 23 tasks under realistic oracle budgets, revealing that many state-of-the-art methods fail to outperform simpler predecessors when sample efficiency is strictly constrained.
Designing new molecules computationally is critical for accelerating drug discovery and materials design. While many artificial intelligence methods have emerged to generate candidate structures, real-world discovery relies on expensive physical experiments or high-accuracy simulations to test each candidate. Consequently, computational algorithms must be sample-efficient, discovering high-performing molecules with as few evaluation queries as possible. Prior literature frequently overlooked this constraint, reporting performance over unconstrained budgets, using trivial benchmarks, or neglecting the substantial run-to-run variation inherent to non-deterministic algorithms.
The article introduces the Practical Molecular Optimization benchmark to evaluate how effectively and efficiently diverse molecular optimization methods perform under a realistic evaluation budget. Specifically, the study assesses the optimization capability, sample efficiency, robustness, and generalizability of 25 molecular design algorithms across 23 standardized target tasks relevant to therapeutics.
To establish a rigorous and fair comparison, the authors limited each algorithm to an evaluation budget of 10,000 oracle queries. Performance was measured by calculating the area under the curve of the top-10 average objective score over time, a metric that rewards algorithms that discover top candidates earlier. The benchmark encompassed string-based, graph-based, and chemical synthesis-based molecular assembly strategies paired with diverse optimization techniques, including genetic algorithms, reinforcement learning, and Bayesian optimization. All algorithms were tuned on standardized tasks and evaluated over five independent trials to account for non-deterministic behavior.
The findings show that no existing algorithm is sample-efficient enough to optimize complex, de novo molecular targets within a realistic experimental budget of only hundreds of evaluations. Furthermore, established older methods such as REINVENT and Graph GA consistently outperformed newer deep learning approaches across the benchmark. Robust representations like SELFIES did not provide a general performance advantage over standard SMILES strings for modern language models, except in specific genetic algorithm implementations. In addition, model-based methods using surrogate predictors only improved sample efficiency when the surrogate model was carefully calibrated, sometimes underperforming simpler model-free counterparts if the surrogate misled the search. Finally, optimization performance depended heavily on the problem landscape, with string-based evolutionary algorithms performing best on atomic composition targets and other methods excelling on structural similarity tasks.
These findings indicate that the molecular design field risks misallocating resources toward increasingly complex architectures that do not improve practical discovery performance. For organizations investing in drug discovery workflows, using unvetted state-of-the-art algorithms without budget constraints risks generating unstable or low-quality candidates while exhausting experimental resources. Simpler, established algorithms currently offer better reliability and efficiency when tuned correctly.
For future molecular design research and deployment, decision-makers and researchers should enforce strict evaluation budgets, evaluate algorithms across multiple diverse target landscapes, perform extensive task-specific hyperparameter tuning, and report performance distributions from multiple independent trials rather than single best runs.
The authors note limitations in the study, including the inability to exhaustively test every algorithmic variant or hyperparameter combination, a potential evaluation bias toward similarity-based target functions, and the exclusion of multi-objective synthesizability constraints and physical docking simulations. Nevertheless, the benchmark provides a high level of confidence in its comparative conclusions by enforcing reproducible, standardized constraints across all evaluated methods.
- Paper: Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules, Rafael Gómez-Bombarelli et al. (2016). This seminal paper introduced continuous latent-space variational autoencoders for molecular generation and property optimization, providing the conceptual foundation for modern computational molecular design algorithms evaluated in the benchmark.
- Paper: Junction Tree Variational Autoencoder for Molecular Graph Generation, Wengong Jin et al. (2018). This paper establishes the Junction Tree VAE architecture for molecular graph generation and property optimization, representing a key graph-based baseline tested under sample-efficiency constraints.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). This foundational work introduced standardized benchmarking datasets and evaluation protocols for molecular machine learning that motivate rigorous molecular optimization benchmarks.
- Paper: A Tutorial on Bayesian Optimization, Peter I. Frazier (2018). This tutorial explains the principles of Bayesian optimization and surrogate modeling under low evaluation budgets, which are critical for understanding sample-efficient black-box molecular optimization.
- Paper: Analyzing Learned Molecular Representations for Property Prediction, Kevin Yang et al. (2019). This paper systematically analyzes learned graph representations versus traditional descriptors for molecular property prediction, providing essential context for evaluating surrogate model accuracy in molecular design.
- Paper: Practical Bayesian Optimization of Machine Learning Algorithms, Jasper Snoek et al. (2012). This work establishes practical Bayesian optimization methodology and Gaussian process acquisition strategies under limited evaluation budgets.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). This work develops a bilevel framework combining large language models and physical simulations for molecular structure and parameter discovery, offering an advanced approach to budget-constrained scientific optimization.
- Paper: Empirical Design in Reinforcement Learning, Andrew Patterson et al. (2024). This guide details robust statistical methodologies, multi-trial reporting, and experimental design principles for reinforcement learning that directly address the reproducibility and variance issues identified in molecular optimization benchmarks.
- Paper: Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules, Jun Xia et al. (2023). This paper advances molecular graph pre-training with balanced discrete codebooks, addressing representation bottlenecks relevant to improving surrogate and generative models in molecular design.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). This work establishes empirical scaling laws for proxy reward overoptimization, providing theoretical and empirical insight into how surrogate models mislead optimization searches.
