A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?

Agustinus KristiadiFelix Strieth-KalthoffMarta SkretaPascal PoupartAlán Aspuru-GuzikGeoff Pleiss

article2024ICML69 citations

Demonstrates that large language models improve principled Bayesian optimization for molecular material discovery only when pretrained or finetuned on domain-specific chemistry data rather than applied out-of-the-box.

Listen

Accelerating the discovery of novel materials and therapeutics is critical to addressing urgent global challenges across healthcare, clean energy, and manufacturing. Bayesian optimization has emerged as an essential tool to automate this search by balancing the exploration of untested chemical candidates with the exploitation of known promising structures. While general-purpose large language models have recently drawn significant attention for their apparent scientific reasoning, prior attempts to apply them in material discovery have largely relied on heuristic, non-Bayesian prompting techniques. These methods fail to provide calibrated uncertainty estimates and often demand prohibitive computational or financial budgets.

The article evaluates whether large language models can effectively accelerate principled Bayesian optimization across molecular search spaces. Specifically, it investigates whether models function effectively as fixed feature extractors or as adaptive surrogate models when paired with parameter-efficient fine-tuning and rigorous Bayesian uncertainty estimation.

To conduct this evaluation, the authors benchmarked eight distinct model configurations across eight real-world chemistry discovery tasks spanning drug candidate binding, battery electrolyte stability, photovoltaics, and optical materials. The study compared general-purpose language models against chemistry-specific transformers and traditional algorithmic molecular fingerprints. The methodological framework tested two primary setups: using language models as static feature extractors feeding into Gaussian process and Laplace-approximated neural network surrogates, and fine-tuning models dynamically with low-rank adaptation while performing Bayesian inference across the newly added parameters.

The analysis yielded several key findings. First, general-purpose large language models underperform simple molecular fingerprints when used as out-of-the-box feature extractors, demonstrating that surface-level text generation does not translate into informative chemistry representations. Second, domain-specific models trained on chemistry data consistently outperformed general-purpose models across single- and multi-objective optimization benchmarks. Third, parameter-efficient fine-tuning combined with Laplace approximations significantly improved optimization efficiency across most tasks. Fourth, principled Bayesian surrogates using lightweight chemistry models decisively outperformed proprietary in-context prompting methods such as GPT-4 in both optimization speed and cost, avoiding the thousands of dollars in query expenses associated with conversational prompting.

These results demonstrate that the raw pretraining data domain matters far more than the general language scale or chat fluency of a model. For decision-makers and research organizations, relying on large, commercial conversational models for molecular discovery introduces unnecessary operating costs and sub-optimal search trajectories. Instead, deploying compact, chemistry-specialized models running locally on standard computing hardware yields superior experimental performance at a fraction of the operational budget and risk.

Organizations advancing automated chemical discovery should prioritize small, domain-specific foundation models integrated with principled Bayesian uncertainty frameworks over general-purpose chat interfaces. Practitioners should maintain prompting templates aligned closely with pretraining representations—such as standard chemical line notations—and utilize parameter-efficient adaptation to fine-tune surrogates as experimental data accumulates. Furthermore, engineering workflows should optimize candidate forward-pass screening, as model inference across large libraries forms the primary operational bottleneck rather than surrogate retraining.

These conclusions are supported by thorough empirical benchmarks across simulated discovery tasks, though the current scope is limited to discrete candidate libraries within chemical domains. While confidence in the comparative superiority of domain-tailored models is high, further validation in continuous molecular generative spaces and wet-lab physical robotic platforms will be necessary to fully map out physical experimental constraints.

No sufficiently relevant recommendations were found.

Cover for A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?

Abstract

Automation is one of the cornerstones of contemporary material discovery. Bayesian optimization (BO) is an essential part of such workflows, enabling scientists to leverage prior domain knowledge into efficient exploration of a large molecular space. While such prior knowledge can take many forms, there has been significant fanfare around the ancillary scientific knowledge encapsulated in large language models (LLMs). However, existing work thus far has only explored LLMs for heuristic materials searches. Indeed, recent work obtains the uncertainty estimate—an integral part of BO—from point-estimated, non-Bayesian LLMs. In this work, we study the question of whether LLMs are actually useful to accelerate principled Bayesian optimization in the molecular space. We take a sober, dispassionate stance in answering this question. This is done by carefully (i) viewing LLMs as fixed feature extractors for standard but principled BO surrogate models and by (ii) leveraging parameter-efficient finetuning methods and Bayesian neural networks to obtain the posterior of the LLM surrogate. Our extensive experiments with real-world chemistry problems show that LLMs can be useful for BO over molecules, but only if they have been pretrained or finetuned with domain-specific data.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Bayesian optimization
  • 2.1.1. BO in Chemical Space
  • 2.2. Bayesian neural networks
  • 2.2.1. Laplace Approximations
  • 2.3. Large language models
  • 2.3.1. Parameter-Efficient Fine-Tuning
  • 3. Experiment Setup
  • 4. How Informative are Pretrained LLMs?
  • 4.1. General or domain-specific LLMs?
  • 4.2. Multiobjective optimization
  • 4.3. Effects of prompting
  • 4.4. The case of in-context learning
  • 5. How Useful are Finetuned LLMs?
  • 5.1. Are finetuned LLM surrogates preferable?
  • 6. Related Work
  • 7. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Additional Details
  • A.1. Pseudocodes
  • A.2. Datasets
  • A.3. Training
  • A.3.1. Fixed-Feature Surrogates
  • A.3.2. Finetuned Surrogates
  • A.4. Prompting
  • A.5. In-context-learning baselines
  • B. Additional Results
  • B.1. Laplace vs. GP
  • B.2. Acquisition functions
  • B.3. Prompts
  • B.4. Textual representations
  • B.5. Computational costs

Knowls

  1. Knowl 1 — Bayesian inference over parameter-efficiently adapted LLMs

    model/method

    For a molecule xx, the adaptive surrogate is a regression head with weights ww composed with a pretrained transformer of frozen weights W∗W^* and a parameter-efficient fine-tuning (PEFT) component with weights ω\omega. The inferred parameters are θ=(w,ω)\theta=(w,\omega); the original transformer weights are conditioned on, rather than included in, the posterior over θ\theta. For observed molecule–property data DD, the model defines

    p(θ∣D;W=W∗)∝p(θ;W=W∗) p(D∣θ;W=W∗).p(\theta\mid D;W=W^*)\propto p(\theta;W=W^*)\,p(D\mid\theta;W=W^*).

    The authors find a maximum-a-posteriori estimate θ∗\theta^* and apply a linearized Laplace approximation (LLA) to obtain a posterior predictive distribution. For an input molecule xx, this predictive distribution is Gaussian with mean gθ∗(x)g_{\theta^*}(x) and covariance J∗(x)Σ∗J∗(x)⊤J^*(x)\Sigma^*J^*(x)^\top, where gθ∗g_{\theta^*} is the regression output, J∗(x)J^*(x) is its Jacobian with respect to the inferred weights at θ∗\theta^*, and Σ∗\Sigma^* is the Laplace covariance over those weights. This gives uncertainty over both the regression head and PEFT weights while keeping the large pretrained transformer fixed. The method applies to PEFT methods generally; the experiments use LoRA.

  2. Knowl 2 — Domain-specific molecular representations are more useful than general-purpose LLM features

    empirical result

    In discrete-pool Bayesian optimization (BO) across the tested chemistry tasks, features from general-purpose T5, GPT-2-Medium, and Llama-2-7B generally performed worse than 1024-bit Morgan fingerprints. Chemistry-focused transformer representations, from T5-Chem and MolFormer, were generally more useful than general-purpose LLM representations and were competitive with or better than fingerprints. MolFormer performed slightly better on average than T5-Chem, despite having 44 million rather than 220 million parameters; the paper notes that MolFormer was pretrained on more chemistry data (100 million versus 33 million examples). The authors interpret this pattern as evidence that chemistry-specific pretraining data may matter more here than general natural-language capability. Across the tested tasks, Laplace-approximated neural-network surrogates also generally performed better than Gaussian-process surrogates, although both showed a similar ordering among feature representations.

  3. Knowl 3 — Fine-tuning usually improves BO, with one reported regression

    empirical result

    The authors compared fixed-feature surrogates with adaptive surrogates that fine-tune T5 or chemistry-specific T5-Chem using LoRA and then place a Laplace approximation over the LoRA and regression-head weights. At each BO round, they reinitialized and trained the LoRA weights using the observations collected so far, then used the resulting posterior predictive uncertainty for acquisition. Fine-tuning improved BO performance on most tested problems for both T5 and T5-Chem. It did not provide a substantial gain on every task, and fine-tuning T5-Chem reduced performance on the Photovoltaics task. The authors suggest that using the same learning-rate, weight-decay, and Laplace hyperparameter settings across all problems may explain that exception; the results therefore do not establish that fine-tuning always helps.

  4. Knowl 4 — Fixed-feature BO uses frozen molecular embeddings with a Bayesian surrogate

    model/method

    The fixed-feature approach converts each molecule xx into a text context c(x)c(x), feeds that context to a pretrained transformer with frozen weights W∗W^*, and uses its last-layer embedding as the molecule representation. In the experiments, token embeddings were averaged across the sequence while padding and end-of-sequence tokens were excluded. A separate Bayesian surrogate is fit to the embeddings and observed molecular-property values; the paper uses Gaussian processes or Laplace-approximated neural networks. At each BO round, the acquisition function is evaluated over the remaining candidate molecules, the selected molecule is evaluated, and the new observation is added to the training data. The transformer itself is not adapted, so uncertainty comes from the Bayesian surrogate rather than from the pretrained LLM.

  5. Knowl 5 — Benchmark tasks, representations, and evaluation

    experimental setup

    The study evaluates discrete-pool BO on eight chemistry problems with simulator-derived property labels: Redoxmer (1,407 molecules; minimize redox potential), Solvation (the same 1,407 molecules; minimize solvation energy), Kinase (10,449; minimize docking score), Laser (10,000 sampled from 182,858; maximize fluorescence oscillator strength), Photovoltaics (10,000 sampled from 2,320,648; maximize power conversion efficiency), Photoswitches (392; maximize transition wavelength), Multi-Redox (1,407; jointly optimize redox potential and solvation energy), and Multi-Laser (10,000; jointly optimize fluorescence oscillator strength and electronic gap). The tested representations include 1024-bit Morgan fingerprints, pretrained MolFormer features, general-purpose T5-Base, GPT-2-Medium and Llama-2-7B, and chemistry-specific T5-Chem. Fingerprint features were paired with a Tanimoto-kernel GP, and transformer features with a Matérn-kernel GP; Laplace-approximated neural-network surrogates were also evaluated. Thompson sampling was the principal acquisition strategy. Single-objective outcomes were compared using task-specific optima and normalized GAP; multiobjective outcomes were evaluated by hypervolume.

  6. Knowl 6 — Small chemistry-specific Bayesian surrogates outperform the tested in-context optimizer

    empirical result

    On a cost-limited Redoxmer comparison, the authors used a candidate pool of 200 molecules, an initial set of 5 labeled molecules, and 15 BO rounds. BO-LIFT, an in-context-learning optimizer whose uncertainty is based on variability in LLM completions, performed poorly with Llama-2-7B but substantially better with GPT-4. Even so, a fixed-feature BO surrogate using the chemistry-specific T5-Chem model performed better and was much cheaper. A GPT-4 BO-LIFT run cost approximately 12–12–18, and five random-seed runs cost $75.81. The comparison supports the paper’s claim for this setup: a small domain-specific LLM paired with a Bayesian surrogate can be preferable to a much larger prompted model, both in optimization performance and monetary cost.

  7. Knowl 7 — Molecular prompts and string representations affect optimization performance

    empirical result

    The study compared four ways to present a molecule to an LLM: the bare SMILES string, a sentence ending in a property-value completion, a direct natural-language prediction request, and a request for a numerical-only answer. Prompt choice affected BO performance across the tested datasets, but effects were model-dependent rather than uniform. T5-Chem generally performed best with the bare SMILES string, matching its SMILES-based pretraining, and it performed well across prompt choices without requiring prompt engineering. T5 and Llama-2-7B showed no consistent improvement from a particular prompt; Llama-2-7B was largely insensitive to prompt changes. Comparing SMILES with IUPAC names also changed outcomes for T5 and T5-Chem, whereas Llama-2-7B was relatively insensitive to the representation. The results indicate that input formats close to a model’s pretraining format are a useful choice, but that prompt and string-format effects depend on the model.

  8. Knowl 8 — T5-Chem leads the tested multiobjective BO experiments

    empirical result

    The multiobjective experiments combined redox potential with solvation energy for Multi-Redox, and added electronic gap as an objective to the Laser problem for Multi-Laser. The surrogate predicted all objectives jointly, and the acquisition strategy used scalarized Thompson sampling with fixed uniform weights. Performance was measured by hypervolume. T5-Chem had the best overall performance across the two problems. MolFormer performed better than T5 and T5-Chem early in optimization, but fell behind both LLM-based representations at later rounds. This pattern is consistent with the single-objective findings that chemistry-specific representations are useful, while also showing that the relative ranking can change over the course of optimization.

  9. Knowl 9 — Candidate evaluation, not fine-tuning, dominates adaptive-surrogate runtime

    empirical result

    In timing experiments on candidate pools of 392, 1,407, and 10,000 molecules, the per-round prediction cost of fine-tuned surrogates scaled approximately linearly with the number of candidates. The authors attribute most of the runtime to forwarding the LLM over the remaining candidate pool, whose size can greatly exceed the number of observations used to train the surrogate. Training the PEFT weights and fitting the Laplace approximation were not the main bottlenecks in these experiments. Computing the LLA Jacobian does add the cost of several backward passes, and the experiments used a minibatch size of 16 because of GPU memory limits. By contrast, fixed-feature BO can cache molecule embeddings and reuse them across rounds.

  10. Knowl 10 — Scope is limited to discrete molecular pools and chemistry

    limitation

    The study concerns BO over a predetermined, finite set of molecules, as in workflows where candidate libraries are specified in advance. It does not evaluate continuous-space BO with LLM-based Bayesian surrogates, which the authors leave for future work, and its experiments focus on chemistry rather than other scientific domains. The molecular property labels are obtained from physically founded simulations, so the reported results are not a demonstration of performance in closed-loop experimental discovery.

Coverage note — The supplementary Thompson-sampling versus expected-improvement comparison is not a separate knowl because its aggregate differences were reported as insignificant and task-level differences went in both directions; implementation details of the released software library are omitted because the paper does not evaluate them as a distinct methodological contribution.

References

  1. 1.Agarwal, G., Doan, H. A., Robertson, L. A., Zhang, L., and Assary, R. S. Discovery of energy storage molecular materials using quantum chemistry-guided multiobjective bayesian optimization. Chemistry of Materials, 33(20), 2021.
  2. 2.Angello, N., Friday, D., Hwang, C., Yi, S., Cheng, A., Torres-Flores, T., Jira, E., Wang, W., Aspuru-Guzik, A., Burke, M., and et al. Closed-loop transfer enables AI to yield chemical knowledge. ChemRxiv, 2023.
  3. 3.Auer, P., Cesa-Bianchi, N., and Fischer, P. Finite-time analysis of the multi-armed bandit problem. Machine Learning, 47, 2002.
  4. 4.Balandat, M., Karrer, B., Jiang, D., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. BoTorch: A framework for efficient Monte-Carlo Bayesian optimization. In NeurIPS, 2020.
  5. 5.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In ACM Conference on Fairness, Accountability, and Transparency, 2021.
  6. 6.Bergamin, F., Moreno-Munoz, P., Hauberg, S., and Arvanitidis, G. Riemannian Laplace approximations for Bayesian neural networks. In NeurIPS, 2023.
  7. 7.Botev, A., Ritter, H., and Barber, D. Practical Gauss-Newton optimisation for deep learning. In ICML, 2017.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In NeurIPS, 2020.
  9. 9.Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., and Androutsopoulos, I. LEGAL-BERT: The muppets straight out of law school. In EMNLP, 2020.
  10. 10.Chithrananda, S., Grand, G., and Ramsundar, B. ChemBERTa: Large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885, 2020.
  11. 11.Christofidellis, D., Giannone, G., Born, J., Winther, O., Laino, T., and Manica, M. Unifying molecular and textual representations via multi-task language modelling. In ICML, 2023.
  12. 12.Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P. Laplace redux – effortless Bayesian deep learning. In NeurIPS, 2021.
  13. 13.de Regt, H. W. Understanding, values and the aims of science. Philosophy of Science, 87(5), 2020.
  14. 14.Enamine Ltd. The Enamine REAL database, 2023. URL https://enamine.net/compound-collections/real-compounds/real-database. Accessed: 2023-11-29.
  15. 15.Fortuin, V., Garriga-Alonso, A., Ober, S. W., Wenzel, F., Ratsch, G., Turner, R. E., van der Wilk, M., and Aitchison, L. Bayesian neural network priors revisited. In ICLR, 2022.
  16. 16.Garnett, R. Bayesian optimization. Cambridge University Press, 2023.
  17. 17.Gomez-Bombarelli, R., Wei, J. N., Duvenaud, D., Hernandez-Lobato, J. M., Sanchez-Lengeling, B., Sheberla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic chemical design using a data-driven continuous representation of molecules. ACS central science, 4(2), 2018.
  18. 18.Gorgulla, C., Nigam, A., Koop, M., C¸ ınaroglu, S. S., Secker, C., Haddadnia, M., Kumar, A., Malets, Y., Hasson, A., Li, M., Tang, M., Levin-Konigsberg, R., Radchenko, D., Kumar, A., Gehev, M., Aquilanti, P.-Y., Gabb, H., Alhossary, A., Wagner, G., Aspuru-Guzik, A., Moroz, Y. S., Fackeldey, K., and Arthanari, H. VirtualFlow 2.0 - the next generation drug discovery platform enabling adaptive screens of 69 billion molecules. bioRxiv, 2023.
  19. 19.Graff, D. E., Shakhnovich, E. I., and Coley, C. W. Accelerating high-throughput virtual screening through molecular pool-based active learning. Chemical Science, 12(22), 2021.
  20. 20.Greenaway, R. L., Jelfs, K. E., Spivey, A. C., and Yaliraki, S. N. From alchemist to AI chemist. Nature Reviews Chemistry, 7(8), 2023.
  21. 21.Griffiths, R.-R., Greenfield, J. L., Thawani, A. R., Jamasb, A. R., Moss, H. B., Bourached, A., Jones, P., McCorkindale, W., Aldrick, A. A., Fuchter, M. J., and Lee, A. A. Data-driven discovery of molecular photoswitches with multioutput Gaussian processes. Chemical Science, 13(45), 2022.
  22. 22.Griffiths, R.-R., Klarner, L., Moss, H. B., Ravuri, A., Truong, S., Stanton, S., Tom, G., Rankovic, B., Du, Y., Jamasb, A., Deshwal, A., Schwartz, J., Tripp, A., Kell, G., Frieder, S., Bourached, A., Chan, A., Moss, J., Guo, C., Durholt, J., Chaurasia, S., Strieth-Kalthoff, F., Lee, A. A., Cheng, B., Aspuru-Guzik, A., Schwaller, P., and Tang, J. GAUCHE: A library for Gaussian processes in chemistry. In NeurIPS, 2023.
  23. 23.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. In NeurIPS, 2023.
  24. 24.Guo, T., Guo, K., Nan, B., Liang, Z., Guo, Z., Chawla, N. V., Wiest, O., and Zhang, X. What can large language models do in chemistry? A comprehensive benchmark on eight tasks. arXiv preprint arXiv:2305.18365, 2023.
  25. 25.Han, C., Wang, Z., Zhao, H., and Ji, H. In-context learning of large language models explained as kernel regression. arXiv preprint arXiv:2305.12766, 2023.
  26. 26.Hernandez-Lobato, J. M., Requeima, J., Pyzer-Knapp, E. O., and Aspuru-Guzik, A. Parallel and distributed Thompson sampling for large-scale accelerated exploration of chemical space. In ICML, 2017.
  27. 27.Hickman, R., Parakh, P., Cheng, A., Ai, Q., Schrier, J., Aldeghi, M., and Aspuru-Guzik, A. Olympus, enhanced: Benchmarking mixed-parameter and multi-objective optimization in chemistry and materials science. ChemRxiv, 2023.
  28. 28.Hickman, R. J., Aldeghi, M., Hase, F., and Aspuru-Guzik, A. Bayesian optimization with known experimental and design constraints for chemistry applications. Digital Discovery, 1(5), 2022.
  29. 29.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for NLP. In ICML, 2019.
  30. 30.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In ICLR, 2022.
  31. 31.Hase, F., Aldeghi, M., Hickman, R. J., Roch, L. M., and Aspuru-Guzik, A. Gryffin: An algorithm for Bayesian optimization of categorical variables informed by expert knowledge. Applied Physics Reviews, 8(3), 2021.
  32. 32.Immer, A., Korzepa, M., and Bauer, M. Improving predictions of Bayesian neural nets via local linearization. In AISTATS, 2021.
  33. 33.Jablonka, K. M., Ai, Q., Al-Feghali, A., Badhwar, S., Bran, J. D., Bringuier, S., Brinson, L. C., Choudhary, K., Circi, D., Cox, S., et al. 14 examples of how LLMs can transform materials science and chemistry: A reflection on a large language model hackathon. arXiv preprint arXiv:2306.06283, 2023a.
  34. 34.Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A., and Smit, B. Leveraging large language models for predictive chemistry. ChemRxiv, 2023b.
  35. 35.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In NIPS, 2018.
  36. 36.Jiang, S., Chai, H., Gonzalez, J., and Garnett, R. BINOCULARS for efficient, nonmyopic sequential experimental design. In ICML, 2020.
  37. 37.Jones, D. R., Schonlau, M., and Welch, W. J. Efficient global optimization of expensive black-box functions. Journal of global optimization, 13, 1998.
  38. 38.Kasneci, E., Seßler, K., Kuchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Gunnemann, S., Hüllermeier, E., et al. ChatGPT for good? on opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 2023.
  39. 39.Kim, S., Lu, P. Y., Loh, C., Smith, J., Snoek, J., and Soljacic, M. Deep learning for Bayesian optimization of scientific problems with high-dimensional structure. TMLR, 2022.
  40. 40.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR, 2015.
  41. 41.Korovina, K., Xu, S., Kandasamy, K., Neiswanger, W., Poczos, B., Schneider, J., and Xing, E. P. ChemBO: Bayesian optimization of small organic molecules with synthesizable recommendations. In AISTATS, 2020.
  42. 42.Kristiadi, A., Eschenhagen, R., and Hennig, P. Posterior refinement improves sample efficiency in Bayesian neural networks. In NeurIPS, 2022.
  43. 43.Kristiadi, A., Immer, A., Eschenhagen, R., and Fortuin, V. Promises and pitfalls of the linearized Laplace in Bayesian optimization. In Advances in Approximate Bayesian Inference, 2023.
  44. 44.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In ACL, 2021.
  45. 45.Li, Y. L., Rudner, T. G., and Wilson, A. G. A study of Bayesian neural network surrogates for Bayesian optimization. In ICLR, 2024.
  46. 46.Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In NeurIPS, 2022.
  47. 47.Liu, T., Astorga, N., Seedat, N., and van der Schaar, M. Large language models to enhance Bayesian optimization. In ICLR, 2024.
  48. 48.Lopez, S. A., Pyzer-Knapp, E. O., Simm, G. N., Lutzow, T., Li, K., Seress, L. R., Hachmann, J., and Aspuru-Guzik, A. The harvard organic photovoltaic dataset. Scientific Data, 3(160086), 2016.
  49. 49.Loshchilov, I. and Hutter, F. SGDR: Stochastic gradient descent with warm restarts. In ICLR, 2017.
  50. 50.Lyu, J., Wang, S., Balius, T. E., Singh, I., Levit, A., Moroz, Y. S., O’Meara, M. J., Che, T., Algaa, E., Tolmachova, K., Tolmachev, A. A., Shoichet, B. K., Roth, B. L., and Irwin, J. J. Ultra-large library docking for discovering new chemotypes. Nature, 566(7743), 2019.
  51. 51.MacKay, D. J. The evidence framework applied to classification networks. Neural Computation, 4(5), 1992a.
  52. 52.MacKay, D. J. A practical Bayesian framework for backpropagation networks. Neural Computation, 4(3), 1992b.
  53. 53.Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B. PEFT: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  54. 54.Maus, N., Jones, H., Moore, J., Kusner, M. J., Bradshaw, J., and Gardner, J. Local latent space Bayesian optimization over structured inputs. In NeurIPS, 2022.
  55. 55.Microsoft Research AI4Science and Microsoft Azure Quantum. The impact of large language models on scientific discovery: a preliminary study using GPT-4. arXiv preprint arXiv:2311.07361, 2023.
  56. 56.Mockus, J. On Bayesian methods for seeking the extremum. In Optimization Techniques IFIP Technical Conference, 1975.
  57. 57.Morgan, H. L. The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. Journal of Chemical Documentation, 5(2), 1965.
  58. 58.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  59. 59.Paria, B., Kandasamy, K., and Poczos, B. A flexible framework for multi-objective Bayesian optimization using random scalarizations. In UAI, 2020.
  60. 60.Pyzer-Knapp, E. O., Suh, C., Gomez-Bombarelli, R., Aguilera-Iparraguirre, J., and Aspuru-Guzik, A. What is high-throughput virtual screening? A perspective from organic materials discovery. Annual Review of Materials Research, 45(1), 2015.
  61. 61.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI Blog, 2019.
  62. 62.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(1), 2020.
  63. 63.Ramos, M. C., Michtavy, S. S., Porosoff, M. D., and White, A. D. Bayesian optimization of catalysts with in-context learning. arXiv preprint arXiv:2304.05341, 2023.
  64. 64.Rankovic, B. and Schwaller, P. BoChemian: Large language model embeddings for Bayesian optimization of chemical reactions. In NeurIPS 2023 Workshop on Adaptive Experimental Design and Active Learning in the Real World, 2023.
  65. 65.Rasmussen, C. E. and Williams, C. K. I. Gaussian processes in machine learning. The MIT Press, 2006.
  66. 66.Restrepo, G. Chemical space: limits, evolution and modelling of an object bigger than our universal library. Digital Discovery, 1(5), 2022.
  67. 67.Ritter, H., Botev, A., and Barber, D. A scalable Laplace approximation for neural networks. In ICLR, 2018.
  68. 68.Ross, J., Belgodere, B., Chenthamarakshan, V., Padhi, I., Mroueh, Y., and Das, P. Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4(12), 2022.
  69. 69.Schneider, G. Virtual screening: An endless staircase? Nature Reviews Drug Discovery, 9(4), 2010.
  70. 70.Sharir, O., Peleg, B., and Shoham, Y. The cost of training NLP models: A concise overview. arXiv preprint arXiv:2004.08900, 2020.
  71. 71.Shoichet, B. K. Virtual screening of chemical libraries. Nature, 432(7019), 2004.
  72. 72.Shwartz-Ziv, R., Goldblum, M., Souri, H., Kapoor, S., Zhu, C., LeCun, Y., and Wilson, A. G. Pre-train your loss: Easy Bayesian transfer learning with informative priors. In NeurIPS, 2022.
  73. 73.Snoek, J., Larochelle, H., and Adams, R. P. Practical Bayesian optimization of machine learning algorithms. In NIPS, 2012.
  74. 74.Stanton, S., Maddox, W., Gruver, N., Maffettone, P., Delaney, E., Greenside, P., and Wilson, A. G. Accelerating Bayesian optimization for biological sequence design with denoising autoencoders. In ICML, 2022.
  75. 75.Strieth-Kalthoff, F., Hao, H., Rathore, V., Derasp, J., Gaudin, T., Angello, N. H., Seifrid, M., Trushina, E., Guy, M., Liu, J., and et al. Delocalized, asynchronous, closed-loop discovery of organic laser emitters. Science, 384(6697), 2024.
  76. 76.Sun, C., Huang, L., and Qiu, X. Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence. In NAACL, 2019.
  77. 77.Thompson, W. R. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3-4), 1933.
  78. 78.Tom, G., Schmid, S. P., Baird, S. G., Cao, Y., Darvish, K., Hao, H., Lo, S., Pablo-Garcia, S., Rajaonson, E. M., Skreta, M., Yoshikawa, N., Corapi, S., Akkoc, G. D., Strieth-Kalthoff, F., Seifrid, M., and Aspuru-Guzik, A. Self-driving laboratories for chemistry and materials sciences. ChemRxiv, 2024.
  79. 79.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  80. 80.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  81. 81.Tripp, A., Daxberger, E., and Hernandez-Lobato, J. M. Sample-efficient optimization in the latent space of deep generative models via weighted retraining. In NeurIPS, 2020.
  82. 82.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In NIPS, 2017.
  83. 83.Vig, J., Madani, A., Varshney, L. R., Xiong, C., Socher, R., and Rajani, N. F. BERTology meets biology: Interpreting attention in protein language models. In ICLR, 2021.
  84. 84.Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Katwyk, P. V., Deac, A., Anandkumar, A., Bergen, K., Gomes, C. P., Ho, S., Kohli, P., Lasenby, J., Leskovec, J., Liu, T.-Y., Manrai, A., Marks, D., Ramsundar, B., Song, L., Sun, J., Tang, J., Velickovic, P., Welling, M., Zhang, L., Coley, C. W., Bengio, Y., and Zitnik, M. Scientific discovery in the age of artificial intelligence. Nature, 620(7972), 2023.
  85. 85.Weininger, D. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1), 1988.
  86. 86.Wilson, A. G., Hu, Z., Salakhutdinov, R., and Xing, E. P. Deep kernel learning. In AISTATS, 2016.
  87. 87.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  88. 88.Yang, A. X., Robeyns, M., Wang, X., and Aitchison, L. Bayesian low-rank adaptation for large language models. arXiv preprint arXiv:2308.13111, 2023.
  89. 89.Zhang, Y. and Lee, A. A. Bayesian semi-supervised learning for uncertainty-calibrated prediction of molecular properties and active learning. Chemical Science, 10(35), 2019.
  90. 90.Zitzler, E. Evolutionary algorithms for multiobjective optimization: Methods and applications, volume 63. Shaker Ithaca, 1999.

Citation

MLA
Kristiadi, A., et al. “A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?”. arXiv, 2024, http://arxiv.org/abs/2402.05015v2.
APA
Kristiadi, A., Strieth-Kalthoff, F., Skreta, M., Poupart, P., Aspuru-Guzik, A., & Pleiss, G. (2024). A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?. arXiv. http://arxiv.org/abs/2402.05015v2
Chicago
Kristiadi, A., F. Strieth-Kalthoff, M. Skreta, P. Poupart, A. Aspuru-Guzik, and G. Pleiss. 2024. “A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?”. arXiv. http://arxiv.org/abs/2402.05015v2.
Harvard
Kristiadi, A. et al. (2024) “A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05015v2.
Vancouver
1. Kristiadi A, Strieth-Kalthoff F, Skreta M, Poupart P, Aspuru-Guzik A, Pleiss G (2024) A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?. arXiv

BibTeX

@article{kristiadi2024sober,
  title = {A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?},
  author = {Kristiadi, Agustinus and Strieth-Kalthoff, Felix and Skreta, Marta and Poupart, Pascal and Aspuru-Guzik, Alán and Pleiss, Geoff},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05015v2},
  eprint = {2402.05015}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/