LIMO: Latent Inceptionism for Targeted Molecule Generation

Peter EckmannKunyang SunBo ZhaoMudong FengMichael K. GilsonRose Yu

article2022ICML59 citations

Presents a gradient-based reverse-optimization framework that pairs a variational autoencoder with sequential property predictors to rapidly generate drug-like molecules with high target binding affinities at a fraction of the computational cost of reinforcement learning methods.

Listen

Early-stage drug discovery is a multi-billion-dollar process that often takes years, largely because researchers must screen enormous chemical libraries to find drug-like molecules that bind tightly to disease-related proteins. Existing machine learning methods, such as reinforcement learning and sampling-based models, struggle to efficiently optimize computationally expensive properties like physics-based binding affinity. The article presents Latent Inceptionism on Molecules (LIMO), a generative modeling framework designed to rapidly generate novel, drug-like compounds with high binding affinities to target proteins.

LIMO works by mapping molecules into a continuous representation using a variational autoencoder and predicting properties from the decoded molecular output rather than directly from the latent space. It then uses gradient-based optimization to backpropagate directly into the latent space to generate optimized molecules, followed by filtering and fine-tuning steps to ensure drug-likeness. The framework was evaluated across standard benchmarks and applied to de novo molecule design targeting two human proteins: human estrogen receptor (ESR1), which has known drug binders, and acetyl-CoA acyl transferase 1 (ACAA1), an enzyme with no known binders.

Key findings show that LIMO operates 6 to 8 times faster than leading reinforcement learning approaches and 12 times faster than sampling methods, completing optimization tasks in roughly one hour. In docking-based binding affinity tests, LIMO achieved superior target binding compared to all baselines, reaching the nanomolar range for both proteins. Furthermore, rigorous molecular dynamics simulations revealed a generated ESR1 candidate with an exceptionally strong predicted dissociation constant of 6·10⁻¹⁴ M, well beyond the affinity of existing cancer drugs such as tamoxifen and raloxifene. LIMO also effectively maintained high structural diversity and demonstrated the unique capability to optimize molecular properties while keeping specific chemical substructures fixed.

These findings suggest that LIMO can significantly compress early drug discovery timelines and lower computational costs by rapidly producing diverse, high-affinity leads for novel and validated disease targets alike. The primary limitation of the study is that binding affinities are derived from computational simulations rather than physical laboratory experiments, meaning false positives can occur despite promising physics-based validation. Moving forward, researchers and drug discovery teams should integrate automated validation tools into the design pipeline and synthesize top-performing candidates to confirm their biological activity through physical laboratory testing.

arXiv: 2206.09010
Cover for LIMO: Latent Inceptionism for Targeted Molecule Generation

Abstract

Generation of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily employ reinforcement learning, Markov sampling, or deep generative models guided by Gaussian processes, which can be prohibitively slow when generating molecules with high binding affinity calculated by computationally-expensive physics-based methods. We present Latent Inceptionism on Molecules (LIMO), which significantly accelerates molecule generation with an inceptionism-like technique. LIMO employs a variational autoencoder-generated latent space and property prediction by two neural networks in sequence to enable faster gradient-based reverse-optimization of molecular properties. Comprehensive experiments show that LIMO performs competitively on benchmark tasks and markedly outperforms state-of-the-art techniques on the novel task of generating drug-like compounds with high binding affinity, reaching nanomolar range against two protein targets. We corroborate these docking-based results with more accurate molecular dynamics-based calculations of absolute binding free energy and show that one of our generated drug-like compounds has a predicted KD (a measure of binding affinity) of 6 · 10−14 M against the human estrogen receptor, well beyond the affinities of typical early-stage drug candidates and most FDA-approved drugs to their respective targets. Code is available at https://github.com/Rose-STL-Lab/LIMO.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Variational Autoencoder
  • 3.2. Property Predictor
  • 3.3. Reverse Optimization
  • 3.4. Refinement
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. QED and Penalized logP Maximization
  • 4.3. logP Targeting
  • 4.4. Similarity-constrained Penalized logP Maximization
  • 4.5. Substructure-constrained logP Extremization
  • 4.6. Single-objective Binding Affinity Maximization
  • 4.7. Multi-objective Binding Affinity Maximization
  • 5. Discussion and Conclusions
  • Acknowledgements
  • References
  • A. Experiment description and baselines
  • A.1. Tasks
  • A.2. Molecule metrics
  • A.3. Fine-tuning algorithm
  • A.4. Baselines
  • A.5. Experimental details
  • A.6. Autodock-GPU
  • A.7. Absolute binding free energy
  • B. Additional experiments
  • B.1. Random generation of molecules
  • B.2. Justification of multi-objective optimization

Knowls

  1. Knowl 1 — LIMO’s stacked latent-space generation architecture

    model/method

    Latent Inceptionism on Molecules (LIMO) generates molecules by combining a variational autoencoder (VAE) with a separate molecular-property predictor. The VAE encoder maps a SELFIES molecular string xx to a continuous latent vector z∈Rmz\in\mathbb{R}^{m}, and the VAE decoder maps zz to a differentiable, real-valued molecular representation. A property predictor then consumes the decoder output rather than zz directly. After the predictor is trained, LIMO treats zz as an optimizable variable and backpropagates the desired property objective through the decoder and predictor. This stacked decoder–predictor design is intended to make property prediction easier than learning a direct latent-space-to-property mapping while retaining efficient gradient-based optimization.

  2. Knowl 2 — SELFIES VAE representation and valid molecular decoding

    model/method

    A molecule is represented as a length-nn SELFIES string x=(x1,…,xn)x=(x_1,\ldots,x_n) whose symbols belong to a vocabulary S={s1,…,sd}S=\{s_1,\ldots,s_d\}. The VAE decoder produces, for every position ii, a probability vector yi∈[0,1]dy_i\in[0,1]^d over the dd SELFIES symbols. The discrete output is obtained by selecting the most probable symbol at each position:

    x^i=sdi∗,di∗=arg⁡max⁡1≤j≤dyi,j.\hat{x}_i=s_{d_i^*},\qquad d_i^*=\arg\max_{1\leq j\leq d}y_{i,j}.

    Because SELFIES is designed so that every SELFIES string corresponds to a chemically valid molecule, decoding the continuous representation into a SELFIES string preserves chemical validity. The VAE learns a latent space in which molecules can be sampled and optimized; random molecules are generated by sampling z∼N(0m,Im)z\sim\mathcal{N}(0_m,I_m) and decoding it.

  3. Knowl 3 — Property prediction from decoded molecular representations

    model/method

    LIMO freezes the trained VAE before fitting a property predictor gθg_\theta. Let fdecf_{\mathrm{dec}} be the VAE decoder and let π\pi be a ground-truth molecular-property function, such as logP, QED, synthetic accessibility, or docking-based binding affinity. Training examples are obtained by sampling latent vectors z∼N(0m,Im)z\sim\mathcal{N}(0_m,I_m), decoding them, and evaluating the resulting molecules with π\pi. The predictor is trained with mean-squared error:

    ℓ0(θ)=∥gθ ⁣(fdec(z))−π ⁣(fdec(z))∥22.\ell_0(\theta)=\left\|g_\theta\!\left(f_{\mathrm{dec}}(z)\right)-\pi\!\left(f_{\mathrm{dec}}(z)\right)\right\|_2^2.

    Training the predictor after the VAE has been completed avoids requiring expensive property labels for the full generative-model dataset, permits generated molecules to provide diverse property-training examples, and allows new properties to be added without retraining the VAE. The predictor receives the decoder’s intermediate molecular representation, not the latent vector itself.

  4. Knowl 4 — Gradient-based reverse optimization with optional substructure preservation

    algorithm

    After training, LIMO freezes the decoder fdecf_{\mathrm{dec}} and property predictors g1,…,gkg_1,\ldots,g_k, then initializes a trainable latent vector from z∼N(0m,Im)z\sim\mathcal{N}(0_m,I_m). For properties to be maximized, it minimizes the weighted objective

    ℓ1(z)=−∑i=1kwi gi ⁣(fdec(z)),\ell_1(z)=-\sum_{i=1}^{k}w_i\,g_i\!\left(f_{\mathrm{dec}}(z)\right),

    where wiw_i are property weights selected by random grid search. The Adam optimizer updates zz through the differentiable decoder–predictor composition. Property-minimization or targeting objectives are implemented by changing the objective direction or weights.

    For substructure-constrained optimization, a binary mask M∈{0,1}n×dM\in\{0,1\}^{n\times d} marks SELFIES positions and symbols that must remain fixed. An input molecule containing the desired substructure is reconstructed as x^start=fdec(fenc(xstart))\hat{x}_{\mathrm{start}}=f_{\mathrm{dec}}(f_{\mathrm{enc}}(x_{\mathrm{start}})). LIMO adds the penalty

    ℓ2(z)=λ∑i=1n∑j=1d[Mi,j(fdec(z)i,j−x^start,i,j)]2,\ell_2(z)=\lambda\sum_{i=1}^{n}\sum_{j=1}^{d}\left[M_{i,j}\left(f_{\mathrm{dec}}(z)_{i,j}-\hat{x}_{\mathrm{start},i,j}\right)\right]^2,

    where λ=1000\lambda=1000 in the experiments. The combined objective preserves the masked molecular substructure while allowing the remaining molecular positions to change.

  5. Knowl 5 — Drug-likeness filtering after multi-objective optimization

    model/method

    After reverse optimization, LIMO removes molecules that fail simple drug-likeness and synthesizability criteria. A molecule is retained only if its quantitative estimate of drug-likeness satisfies QED>0.4\mathrm{QED}>0.4, its synthetic accessibility score satisfies SA<5.5\mathrm{SA}<5.5, and it contains no ring with fewer than 5 or more than 6 atoms. The ring restriction supplements SA because very small and very large rings can be difficult to synthesize without being adequately penalized by the SA score. The authors also used an auxiliary greedy refinement procedure for some tasks: repeatedly test replacing eligible carbon atoms with N, O, Cl, or F, retain the best valid improving replacement, and stop when a complete pass yields no improvement.

  6. Knowl 6 — Experimental configuration for molecular optimization

    experimental setup

    Optimization experiments use the approximately 250,000 purchasable drug-like molecules in ZINC250k; random-generation experiments use the approximately 2-million-molecule ZINC-derived MOSES dataset. The shared VAE latent dimension is 1024. Its 64-dimensional symbol embedding feeds four batch-normalized fully connected layers with ReLU activations, using width 2,000 in the first layer and 1,000 thereafter; the decoder has the same layer dimensions and uses position-wise softmax outputs. The VAE is trained for 18 epochs with Adam at learning rate 10−410^{-4}, using an ELBO whose reconstruction and KL terms are weighted 0.9 and 0.1, respectively. Each property predictor has three 1,000-unit ReLU fully connected layers; it is trained for 5 epochs, followed by 10 reverse-optimization epochs with latent-space learning rate 0.1. There are 100,000 predictor-training examples for ordinary properties and 10,000 for binding affinity because docking is expensive.

    The optimized properties are logP, penalized logP, QED, SA, and docking-based binding affinity. Binding-affinity experiments target human estrogen receptor ESR1 using crystal structure PDB 1ERR and human peroxisomal acetyl-CoA acyl transferase 1 (ACAA1) using PDB 2IIK. LIMO receives only the protein structure and binding-site information, not known binders. AutoDock-GPU supplies docking affinities, RDKit supplies ordinary molecular properties, and comparisons use JT-VAE, GCPN, MolDQN, MARS, and GraphDF. Experiments run on two GTX 1080 Ti GPUs, four CPU cores, and 32 GB of memory.

  7. Knowl 7 — Fast QED and penalized-logP maximization

    data/table

    On 100,000 generated molecules, LIMO matches or approaches the strongest benchmark scores while requiring substantially less total generation-and-evaluation time. The reported top-three scores are:

    • JT-VAE: penalized logP 5.30,4.93,4.495.30,4.93,4.49; QED 0.925,0.911,0.9100.925,0.911,0.910; 24 hours; no output-length limit.
    • GCPN: penalized logP 7.98,7.85,7.807.98,7.85,7.80; QED 0.948,0.947,0.9460.948,0.947,0.946; 8 hours; length limit.
    • MolDQN: penalized logP 11.8,11.8,11.811.8,11.8,11.8; QED 0.948,0.943,0.9430.948,0.943,0.943; 24 hours; length limit.
    • MARS: penalized logP 45.0,44.3,43.845.0,44.3,43.8; QED 0.948,0.948,0.9480.948,0.948,0.948; 12 hours; no length limit.
    • GraphDF: penalized logP 13.7,13.2,13.213.7,13.2,13.2; QED 0.948,0.948,0.9480.948,0.948,0.948; 8 hours; no length limit.
    • LIMO with prediction directly on zz: penalized logP 6.52,6.38,5.596.52,6.38,5.59; QED 0.910,0.909,0.8920.910,0.909,0.892; 1 hour; length limit.
    • LIMO with prediction on the decoded representation: penalized logP 10.5,9.69,9.6010.5,9.69,9.60; QED 0.947,0.946,0.9450.947,0.946,0.945; 1 hour; length limit.

    On an unseen set of 1,000 generated molecules, the coefficient of determination between predicted and actual properties is R2=0.04R^2=0.04 for direct prediction from zz and R2=0.38R^2=0.38 for the stacked decoder–predictor architecture. Replacing the fully connected encoder and decoder with eight-layer, 512-hidden-unit LSTMs reduced performance; the maximum QED observed in that variant was 0.3.

  8. Knowl 8 — LogP targeting, similarity-constrained optimization, and fixed-substructure editing

    empirical result

    For targeting the range −2.5<logP<−2.0-2.5<\mathrm{logP}<-2.0, LIMO places 10.4% of generated molecules in the target range and achieves diversity 0.914, where diversity is one minus mean pairwise Morgan-fingerprint Tanimoto similarity. The corresponding success/diversity values are JT-VAE 11.3%/0.846, GCPN 85.5%/0.392, MolDQN 9.66%/0.854, and GraphDF 0% with diversity unavailable. LIMO generates 33 molecules per second within the target range.

    For similarity-constrained penalized-logP maximization, the experiment starts from the 800 ZINC250k molecules with the lowest penalized logP, performs 1,000 latent gradient-ascent steps per molecule, and selects the best candidate satisfying a minimum Morgan-fingerprint Tanimoto similarity δ\delta. LIMO’s mean improvement and success rate are 10.1±2.310.1\pm2.3 and 100% at δ=0.0\delta=0.0; 5.8±2.65.8\pm2.6 and 99.0% at δ=0.2\delta=0.2; 3.6±2.33.6\pm2.3 and 93.7% at δ=0.4\delta=0.4; and 1.8±2.01.8\pm2.0 and 85.5% at δ=0.6\delta=0.6. LIMO has the largest improvement at the two least restrictive similarity levels and remains competitive at higher similarity constraints.

    In a separate fixed-substructure experiment, LIMO successfully increased or decreased logP from two ZINC250k starting molecules while preserving a specified substructure. This demonstrates optimization around a fixed scaffold rather than merely enforcing whole-molecule similarity.

  9. Knowl 9 — Single-objective docking-affinity generation

    data/table

    LIMO was evaluated by generating 10,000 molecules per method for ESR1 and ACAA1 and reporting the three lowest predicted dissociation constants KDK_D in nanomolar; lower KDK_D means higher computed affinity. LIMO produced the strongest docking scores with the shortest runtime:

    • GCPN: ESR1 6.4,6.6,8.56.4,6.6,8.5 nM; ACAA1 75,83,8475,83,84 nM; 6 hours.
    • MolDQN: ESR1 373,588,1062373,588,1062 nM; ACAA1 240,337,608240,337,608 nM; 6 hours.
    • GraphDF: ESR1 25,47,5125,47,51 nM; ACAA1 370,520,590370,520,590 nM; 12 hours.
    • MARS: ESR1 17,64,6917,64,69 nM; ACAA1 163,203,236163,203,236 nM; 6 hours.
    • LIMO: ESR1 0.72,0.89,1.40.72,0.89,1.4 nM; ACAA1 37,37,4137,37,41 nM; 1 hour.

    Although single-objective optimization generated very favorable docking scores, the resulting molecules often contained problematic polyenes or rings with at least 8 atoms. The authors regarded these structures as concerns for reactivity, toxicity, or synthesis, motivating simultaneous optimization of affinity, QED, and SA.

  10. Knowl 10 — Multi-objective affinity optimization and molecular-dynamics corroboration

    empirical result

    For each protein target, LIMO generated 100,000 molecules while simultaneously optimizing docking affinity, QED, and SA, then applied the drug-likeness and ring filters. The two best retained ligands for each target had the following reported properties; KDK_D values are in nM, QED is higher-is-better, SA is lower-is-better, and ABFE values are more rigorous than docking values.

    • ESR1 LIMO molecule 1: docking KD=4.6K_D=4.6, QED 0.43, SA 4.8, ABFE 6×10−56\times10^{-5} nM, Lipinski pass, 0 PAINS alerts, Fsp3=0.16Fsp^3=0.16, MCE-18 90.
    • ESR1 LIMO molecule 2: docking KD=2.8K_D=2.8, QED 0.64, SA 4.9, ABFE 1000 nM, Lipinski pass, 0 PAINS alerts, Fsp3=0.52Fsp^3=0.52, MCE-18 76.
    • ESR1 GCPN molecules: docking KD=810K_D=810 and 2.7×1042.7\times10^4 nM, with QED 0.43 and 0.80 and SA 4.2 and 3.7; their ABFE values were not reported.
    • For comparison, tamoxifen had docking KD=87K_D=87 nM and experimental ABFE-derived KD=1.5K_D=1.5 nM; raloxifene had docking KD=7.9×106K_D=7.9\times10^6 nM and experimental KD=0.030K_D=0.030 nM.
    • ACAA1 LIMO molecule 1: docking KD=28K_D=28, QED 0.57, SA 5.5, ABFE 4×1044\times10^4 nM, Lipinski pass, 0 PAINS alerts, Fsp3=0.52Fsp^3=0.52, MCE-18 52.
    • ACAA1 LIMO molecule 2: docking KD=31K_D=31, QED 0.44, SA 4.9, no ABFE binding detected, Lipinski pass, 0 PAINS alerts, Fsp3=0.81Fsp^3=0.81, MCE-18 45.
    • ACAA1 GCPN molecules both had docking KD=8500K_D=8500 nM, with QED 0.69 and 0.54 and SA 4.2 and 4.3.

    The authors inspected docked three-dimensional poses and found that representative LIMO ligands fit the target pockets with favorable interactions. They then used molecular-dynamics-based absolute binding free-energy calculations on promising ligands. For each ligand, five docking poses were evaluated, their free energies were combined into an overall binding free energy, and the reported value was averaged over two independent runs. One of two ESR1 candidates was assigned KD=6×10−5K_D=6\times10^{-5} nM, equivalent to 6×10−146\times10^{-14} M, by ABFE. This result is computational and was not experimentally confirmed.

Coverage note — The detailed greedy carbon-to-heteroatom fine-tuning pseudocode and the auxiliary MOSES random-generation benchmark were not expanded into separate knowls because they are secondary refinement/diversity analyses rather than load-bearing components of LIMO’s targeted optimization contribution.

References

  1. 1.Abdelraheem, E. M. M., Kurpiewska, K., Kalinowska-Tłuscik, J., and Döm​​ling, A. Artificial macrocycles by Ugi reaction and Passerini ring closure. The Journal of Organic Chemistry, 81(19):8789–8795, September 2016.
  2. 2.Allen, W. J., Fochtman, B. C., Balius, T. E., and Rizzo, R. C. Customizable de novo design strategies for DOCK: Application to HIVgp41 and other therapeutic targets. Journal of Computational Chemistry, 38(30):2641–2663, September 2017.
  3. 3.Baell, J. B. and Holloway, G. A. New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. Journal of Medicinal Chemistry, 53(7):2719–2740, April 2010.
  4. 4.Bickerton, G. R., Paolini, G. V., Besnard, J., Muresan, S., and Hopkins, A. L. Quantifying the chemical beauty of drugs. Nature chemistry, 4:90–98, 2012.
  5. 5.Birch, M. and Sibley, G. Antifungal chemistry review. In Comprehensive Medicinal Chemistry III, pp. 703–716. Elsevier, 2017.
  6. 6.Boitreaud, J., Mallet, V., Oliver, C., and Waldispühl, J. Optimol: Optimization of binding affinities in chemical space for drug discovery. Journal of Chemical Information and Modeling, 60(12):5658–5666, September 2020.
  7. 7.Carracedo-Reboredo, P., Linares-Blanco, J., Rodríguez-Fernandez, N., Cedrón, F., Novoa, F. J., Carballal, A., Maojo, V., Pazos, A., and Fernandez-Lozano, C. A review on machine learning approaches and trends in drug discovery. Computational and Structural Biotechnology Journal, 19:4538–4558, 2021.
  8. 8.Cournia, Z., Allen, B. K., Beuming, T., Pearlman, D. A., Radak, B. K., and Sherman, W. Rigorous Free Energy Simulations in Virtual Screening. Journal of Chemical Information and Modeling, 60(9):4153–4169, September 2020. ISSN 1549-9596, 1549-960X.
  9. 9.Dai, H., Tian, Y., Dai, B., Skiena, S., and Song, L. Syntax-directed variational autoencoder for structured data. In International Conference on Learning Representations, 2018.
  10. 10.De Cao, N. and Kipf, T. MolGAN: An implicit generative model for small molecular graphs. arXiv preprint arXiv:1805.11973, 2018.
  11. 11.Ertl, P. and Schuffenhauer, A. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of Cheminformatics, 1(1):8, Jun 2009.
  12. 12.Friesner, R. A., Banks, J. L., Murphy, R. B., Halgren, T. A., Klicic, J. J., Mainz, D. T., Repasky, M. P., Knoll, E. H., Shelley, M., Perry, J. K., Shaw, D. E., Francis, P., and Shenkin, P. S. Glide: a new approach for rapid, accurate docking and scoring. 1. method and assessment of docking accuracy. Journal of Medicinal Chemistry, 47 (7):1739–1749, February 2004.
  13. 13.Fu, T., Gao, W., Xiao, C., Yasonik, J., Coley, C. W., and Sun, J. Differentiable scaffolding tree for molecular optimization. In International Conference on Learning Representations, 2022.
  14. 14.Gilson, M., Given, J., Bush, B., and McCammon, J. The statistical-thermodynamic basis for computation of binding affinities: a critical review. Biophysical Journal, 72 (3):1047–1069, March 1997.
  15. 15.Gilson, M. K. and Zhou, H.-X. Calculation of protein-ligand binding affinities. Annual Review of Biophysics and Biomolecular Structure, 36(1):21–42, 2007.
  16. 16.Gomez-Bombarelli, R., N.Wei, J., Duvenaud, D., Hernandez-Lobato, J. M., Sánchez-Lengeling, B., Sheberla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic chemical design using a data-driven continuous representation of molecules. In ACS Cent. Sci., volume 4, pp. 268–276, 2018.
  17. 17.Guilloux, V. L., Schmidtke, P., and Tuffery, P. Fpocket: An open source platform for ligand pocket detection. BMC Bioinformatics, 10(1), June 2009.
  18. 18.Guimaraes, G. L., Sanchez-Lengeling, B., Outeiral, C., Farias, P. L. C., and Aspuru-Guzik, A. Objective-reinforced generative adversarial networks (ORGAN) for sequence generation models. arXiv preprint arXiv:1705.10843, 2017.
  19. 19.Hataya, R., Nakayama, H., and Yoshizoe, K. Graph energy-based model for substructure preserving molecular design. arXiv preprint arXiv:2102.04600, 2021.
  20. 20.Heinzelmann, G. and Gilson, M. K. Automation of absolute protein-ligand binding free energy calculations for docking refinement and compound evaluation. Scientific Reports, 11(1):1116, December 2021. ISSN 2045-2322.
  21. 21.Hughes, J., Rees, S., Kalindjian, S., and Philpott, K. Principles of early drug discovery. British Journal of Pharmacology, 162(6):1239–1249, February 2011.
  22. 22.Hussain, A., Yousuf, S. K., and Mukherjee, D. Importance and synthesis of benzannulated medium-sized and macrocyclic rings (BMRs). RSC Adv., 4(81):43241–43257, 2014.
  23. 23.Irwin, J. J., Sterling, T., Mysinger, M. M., Bolstad, E. S., and Coleman, R. G. ZINC: A free tool to discover chemistry for biology. J. Chem. Inf. Model., 52, 2012.
  24. 24.Ivanenkov, Y. A., Zagribelnyy, B. A., and Aladinskiy, V. A. Are we opening the door to a new era of medicinal chemistry or being collapsed to a chemical singularity? Journal of Medicinal Chemistry, 62(22):10026–10043, June 2019.
  25. 25.Jeon, W. and Kim, D. Autonomous molecule generation using reinforcement learning and docking to develop potential novel inhibitors. Scientific Reports, 10(1), December 2020.
  26. 26.Jin, W., Barzilay, R., and Jaakkola, T. Junction tree variational autoencoder for molecular graph generation. In Proceedings of the 35th International Conference on Machine Learning, 2018.
  27. 27.Jin, W., Yang, K., Barzilay, R., and Jaakkola, T. Learning multimodal graph-to-graph translation for molecular optimization. In International Conference on Learning Representations, 2019.
  28. 28.Jin, W., Barzilay, D., and Jaakkola, T. Hierarchical generation of molecular graphs using structural motifs. In III, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 4839–4848. PMLR, 13–18 Jul 2020a.
  29. 29.Jin, W., Barzilay, R., and Jaakkola, T. Multi-objective molecule generation using interpretable substructures. In Proceedings of the 37th International Conference on Machine Learning, 2020b.
  30. 30.Kadurin, A., Nikolenko, S., Khrabrov, K., Aliper, A., and Zhavoronkov, A. druGAN: An advanced generative adversarial autoencoder model for de novo generation of new molecules with desired molecular properties in silico. Mol. Pharmaceutics, 14(9):3098–3104, 2017.
  31. 31.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  32. 32.Krenn, M., Hase, F., Nigam, A., Friederich, P., and Aspuru-Guzik, A. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024, October 2020.
  33. 33.Kusner, M. J., Paige, B., and Hernandez-Lobato, J. M. Grammar variational autoencoder. Proceedings of the 34th International Conference on Machine Learning, 2017.
  34. 34.Li, Y., Zhang, L., and Liu, Z. Multi-objective de novo drug design with conditional graph generative model. Journal of cheminformatics, 10(1):1–24, 2018.
  35. 35.Lim, J., Hwang, S.-Y., Moon, S., Kim, S., and Kim, W. Y. Scaffold-based molecular design with a graph generative model. Chemical Science, 11(4):1153–1164, 2020.
  36. 36.Lipinski, C. A., Lombardo, F., Dominy, B. W., and Feeney, P. J. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Advanced Drug Delivery Reviews, 46(1-3):3–26, March 2001.
  37. 37.Liu, T., Lin, Y., Wen, X., Jorissen, R. N., and Gilson, M. K. BindingDB: a web-accessible database of experimentally determined protein-ligand binding affinities. Nucleic Acids Research, 35(Database):D198–D201, January 2007.
  38. 38.Luo, S., Guan, J., Ma, J., and Peng, J. A 3d generative model for structure-based drug design. In Advances in Neural Information Processing Systems, volume 34, pp. 6229–6239. Curran Associates, Inc., 2021a.
  39. 39.Luo, Y., Yan, K., and Ji, S. GraphDF: A discrete flow model for molecular graph generation. In International Conference on Machine Learning. PMLR 139, 2021b.
  40. 40.Ma, T., Chen, J., and Xiao, C. Constrained generation of semantically valid graphs via regularizing variational autoencoders. Advances in Neural Information Processing Systems, 31:7113–7124, 2018.
  41. 41.Maziarz, K., Jackson-Flux, H., Cameron, P., Sirockin, F., Schneider, N., Stiefl, N., Segler, M., and Brockschmidt, M. Learning to extend molecular scaffolds with structural motifs. In International Conference on Learning Representations. arXiv, 2021.
  42. 42.Mordvintsev, A., Olah, C., and Tyka, M. Inceptionism: Going deeper into neural networks, 2015. URL https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html.
  43. 43.Nigam, A., Friederich, P., Krenn, M., and Aspuru-Guzik, A. Augmenting genetic algorithms with deep neural networks for exploring the chemical space. arXiv preprint arXiv:1909.11655, 2020.
  44. 44.Notin, P., Hernandez-Lobato, J. M., and Gal, Y. Improving black-box optimization in vae latent space using decoder uncertainty. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 802–814. Curran Associates, Inc., 2021.
  45. 45.O'Boyle, N. M., Banck, M., James, C. A., Morley, C., Vandermeersch, T., and Hutchison, G. R. Open babel: An open chemical toolbox. Journal of Cheminformatics, 3 (1), October 2011.
  46. 46.Olivecrona, M., Blaschke, T., Engkvist, O., and Chen, H. Molecular de-novo design through deep reinforcement learning. Journal of Cheminformatics, 9(48), 2017.
  47. 47.Paul, S. M., Mytelka, D. S., Dunwiddie, C. T., Persinger, C. C., Munos, B. H., Lindborg, S. R., and Schacht, A. L. How to improve r&d productivity: the pharmaceutical industry's grand challenge. Nature Reviews Drug Discovery, 9(3):203–214, February 2010.
  48. 48.Polishchuk, P. G., Madzhidov, T. I., and Varnek, A. Estimation of the size of drug-like chemical space based on GDB-17 data. Journal of Computer-Aided Molecular Design, 27(8):675–679, August 2013.
  49. 49.Polykovskiy, D., Zhebrak, A., Sanchez-Lengeling, B., Golovanov, S., Tatanov, O., Belyaev, S., Kurbanov, R., Artamonov, A., Aladinskiy, V., Veselov, M., Kadurin, A., Johansson, S., Chen, H., Nikolenko, S., Aspuru-Guzik, A., and Zhavoronkov, A. Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models. Frontiers in Pharmacology, 2020.
  50. 50.Popova, M., Shvets, M., Oliva, J., and Isayev, O. MolecularRNN: Generating realistic molecular graphs with optimized properties. In arXiv preprint arXiv:1905.13372, 2019.
  51. 51.Rogers, D. and Hahn, M. Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5):742–754, April 2010.
  52. 52.Santos-Martins, D., Solis-Vasquez, L., Tillack, A. F., Sanner, M. F., Koch, A., and Forli, S. Accelerating AutoDock4 with GPUs and gradient-based local search. Journal of Chemical Theory and Computation, 17(2):1060–1073, January 2021.
  53. 53.Shen, C., Krenn, M., Eppel, S., and Aspuru-Guzik, A. Deep molecular dreaming: inverse machine learning for de-novo molecular design and interpretability with surjective representations. Machine Learning: Science and Technology, 2(3), 2021.
  54. 54.Shi, C., Xu, M., Zhu, Z., Zhang, W., Zhang, M., and Tang, J. GraphAF: A flow-based autoregressive model for molecular graph generation. In International Conference on Machine Learning, 2020.
  55. 55.Simonovsky, M. and Komodakis, N. GraphVAE: Towards generation of small graphs using variational autoencoders. International Conference on Artificial Neural Networks, 2018.
  56. 56.Spiegel, J. O. and Durrant, J. D. Autogrow4: an open-source genetic algorithm for de novo drug design and lead optimization. Journal of Cheminformatics, 12(25), 2020.
  57. 57.Tripp, A., Daxberger, E., and Hernandez-Lobato, J. M. Sample-efficient optimization in the latent space of deep generative models via weighted retraining. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 11259–11272. Curran Associates, Inc., 2020.
  58. 58.Wei, W., Cherukupalli, S., Jing, L., Liu, X., and Zhan, P. Fsp3: A new parameter for drug-likeness. Drug Discovery Today, 25(10):1839–1845, October 2020.
  59. 59.Weininger, D. SMILES, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Modeling, 28(1):31–36, February 1988.
  60. 60.Xie, Y., Shi, C., Zhou, H., Yang, Y., Zhang, W., Yu, Y., and LI, L. MARS: Markov molecular sampling for multi-objective drug discovery. In International Conference on Learning Representations, 2021.
  61. 61.Xiong, G., Wu, Z., Yi, J., Fu, L., Yang, Z., Hsieh, C., Yin, M., Zeng, X., Wu, C., Lu, A., Chen, X., Hou, T., and Cao, D. ADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties. Nucleic Acids Research, 49(W1):W5–W14, April 2021.
  62. 62.You, J., Liu, B., Ying, R., Pande, V., and Leskovec, J. Graph convolutional policy network for goal-directed molecular graph generation. In 32nd Conference on Neural Information Processing Systems, 2018.
  63. 63.Zang, C. and Wang, F. MoFlow: An invertible flow model for generating molecular graphs. In In Proceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2020.
  64. 64.Zhavoronkov, A., Ivanenkov, Y. A., Aliper, A., Veselov, M. S., Aladinskiy, V. A., Aladinskaya, A. V., Terentiev, V. A., Polykovskiy, D. A., Kuznetsov, M. D., Asadulaev, A., Volkov, Y., Zholus, A., Shayakhmetov, R. R., Zhebrak, A., Minaeva, L. I., Zagribelnyy, B. A., Lee, L. H., Soll, R., Madge, D., Xing, L., Guo, T., and Aspuru-Guzik, A. Deep learning enables rapid identification of potent ddr1 kinase inhibitors. Nature Biotechnology, 37:1038–1040, 2019.
  65. 65.Zhou, Z., Kearnes, S., Li, L., Zare, R. N., and patrick Riley. Optimization of molecules via deep reinforcement learning. Scientific Reports, 9, 2019.

Citation

MLA
Eckmann, P., et al. “LIMO: Latent Inceptionism for Targeted Molecule Generation”. International Conference on Machine Learning, vol. 162, 2022, pp. 5777–92, https://proceedings.mlr.press/v162/eckmann22a.html.
APA
Eckmann, P., Sun, K., Zhao, B., Feng, M., Gilson, M., & Yu, R. (2022). LIMO: Latent Inceptionism for Targeted Molecule Generation. International Conference on Machine Learning, 162, 5777–5792. https://proceedings.mlr.press/v162/eckmann22a.html
Chicago
Eckmann, P., K. Sun, B. Zhao, M. Feng, M. Gilson, and R. Yu. 2022. “LIMO: Latent Inceptionism for Targeted Molecule Generation”. International Conference on Machine Learning 162: 5777–92. https://proceedings.mlr.press/v162/eckmann22a.html.
Harvard
Eckmann, P. et al. (2022) “LIMO: Latent Inceptionism for Targeted Molecule Generation”, International Conference on Machine Learning. PMLR, pp. 5777–5792. Available at: https://proceedings.mlr.press/v162/eckmann22a.html.
Vancouver
1. Eckmann P, Sun K, Zhao B, Feng M, Gilson M, Yu R (2022) LIMO: Latent Inceptionism for Targeted Molecule Generation. In: International Conference on Machine Learning. PMLR, pp 5777–5792

BibTeX

@InProceedings{pmlr-v162-eckmann22a,
  title = 	 {{LIMO}: Latent Inceptionism for Targeted Molecule Generation},
  author =       {Eckmann, Peter and Sun, Kunyang and Zhao, Bo and Feng, Mudong and Gilson, Michael and Yu, Rose},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {5777--5792},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/eckmann22a/eckmann22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/eckmann22a.html},
  abstract = 	 {Generation of drug-like molecules with high binding affinity to target proteins remains a difficult and resource-intensive task in drug discovery. Existing approaches primarily employ reinforcement learning, Markov sampling, or deep generative models guided by Gaussian processes, which can be prohibitively slow when generating molecules with high binding affinity calculated by computationally-expensive physics-based methods. We present Latent Inceptionism on Molecules (LIMO), which significantly accelerates molecule generation with an inceptionism-like technique. LIMO employs a variational autoencoder-generated latent space and property prediction by two neural networks in sequence to enable faster gradient-based reverse-optimization of molecular properties. Comprehensive experiments show that LIMO performs competitively on benchmark tasks and markedly outperforms state-of-the-art techniques on the novel task of generating drug-like compounds with high binding affinity, reaching nanomolar range against two protein targets. We corroborate these docking-based results with more accurate molecular dynamics-based calculations of absolute binding free energy and show that one of our generated drug-like compounds has a predicted $K_D$ (a measure of binding affinity) of $6 \cdot 10^{-14}$ M against the human estrogen receptor, well beyond the affinities of typical early-stage drug candidates and most FDA-approved drugs to their respective targets. Code is available at https://github.com/Rose-STL-Lab/LIMO.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/