LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery

Pingchuan MaTsun-Hsuan WangMinghao GuoZhiqing SunJoshua B. TenenbaumDaniela RusChuang GanWojciech Matusik

article2024ICML73 citations

Proposes Scientific Generative Agent, a bilevel optimization framework that couples language models for discrete hypothesis generation with differentiable simulations for continuous parameter optimization to discover physical laws and design molecules.

Listen

Accelerating scientific discovery requires systems that can formulate valid hypotheses while grounding them in rigorous physical laws. Large language models possess extensive cross-disciplinary knowledge and reasoning capabilities, yet they struggle with numerical precision and simulating observational feedback on their own. Conversely, numerical simulations accurately model continuous physical behaviors but lack the creative reasoning needed to propose symbolic equations or novel molecular structures.

The article demonstrates the Scientific Generative Agent, a unified framework that combines large language models with differentiable physical simulations. The objective is to automate the discovery of discrete scientific structures—such as material equations and chemical formulas—alongside the optimization of their continuous numerical parameters.

To evaluate this framework, the authors implemented a bilevel optimization process tested on two complex physical problems: discovering constitutive material laws from motion trajectories and designing molecular structures to match target quantum mechanical properties. The outer level employs a large language model to propose discrete symbolic expressions (such as mathematical code or molecular string sequences) and define the continuous parameter search space. The inner level utilizes differentiable simulation engines and gradient-based algorithms to calibrate continuous parameters against physical observations and feed loss curves back to the language model. An evolutionary exploration-exploitation strategy regulates model temperature to balance conservative refinements with novel hypotheses across multiple iterations.

The primary findings demonstrate that this bilevel agent significantly outperforms existing automated discovery and symbolic regression methods. First, across eight benchmark tasks, the framework achieved errors orders of magnitude lower than standard prompt-based baselines. Second, in ablation tests on difficult tasks, removing the bilevel simulation feedback or the exploit-and-explore mechanism caused performance degradation of over 50%, highlighting the necessity of combining both levels. Third, in synthetic tests featuring a non-existent, imaginary material law, the model successfully recovered the hybrid behavior without relying on memorized training data. Fourth, extending optimization from 5 to 20 iterations yielded massive accuracy gains, improving performance by up to 52,400% on the most challenging plasticity tasks.

These findings imply that artificial intelligence can discover valid, unconventional scientific solutions that deviate from traditional human heuristics while remaining physically sound. By replacing rigid, domain-specific symbolic regression tools with code-generating models, the approach reduces the labor and time required to characterize complex materials and molecular candidates. GPT-4 served as the top-performing backbone model, though open-source alternatives also demonstrated viable baseline utility.

For practitioners and research organizations seeking to adopt this methodology, the authors recommend tuning the number of optimization iterations as the primary control lever for difficult discovery tasks and maintaining an approximate 1:3 balance between conservative and exploratory proposals. Prior to deploying such systems autonomously, organizations should implement code execution sandboxes, incorporate domain-specific constraints or human-in-the-loop feedback, and utilize key-value cache reuse to manage computational and financial costs, which averaged approximately $10 per task under commercial model pricing.

Confidence in these findings is strong across the tested simulation benchmarks in mechanics and molecular design. However, readers should note limitations: the system was evaluated within simulated environments rather than physical laboratory experiments, the interpretability of generated code remains challenging to guarantee, and the approach depends heavily on the availability of differentiable simulators.

Cover for LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery

Abstract

Large Language Models have recently gained significant attention in scientific discovery for their extensive knowledge and advanced reasoning capabilities. However, they encounter challenges in effectively simulating observational feedback and grounding it with language to propel advancements in physical scientific discovery. Conversely, human scientists undertake scientific discovery by formulating hypotheses, conducting experiments, and revising theories through observational analysis. Inspired by this, we propose to enhance the knowledge-driven, abstract reasoning abilities of LLMs with the computational strength of simulations. We introduce Scientific Generative Agent (SGA), a bilevel optimization framework: LLMs act as knowledgeable and versatile thinkers, proposing scientific hypotheses and reason about discrete components, such as physics equations or molecule structures; meanwhile, simulations function as experimental platforms, providing observational feedback and optimizing via differentiability for continuous parts, such as physical parameters. We conduct extensive experiments to demonstrate our framework’s efficacy in constitutive law discovery and molecular design, unveiling novel solutions that differ from conventional human expectations yet remain coherent upon analysis.

Table of Contents

  • 1. Introduction
  • 2. Scientific Generative Agent
  • 2.1. Bilevel Optimization Pipeline
  • 2.2. LLM-Driven Outer-Level Search
  • 2.3. Differentiable Inner-Level Optimization
  • 3. Experiments
  • 3.1. Problem Definitions
  • 3.2. Experiment Setup
  • 3.3. Physical Scientific Discovery
  • 3.4. Ablation Study
  • 3.5. Case Study
  • 4. Related Work
  • 4.1. Automated Scientific Discovery
  • 4.2. Large Language Models and Agents
  • 4.3. Bilevel Optimization
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Full Prompts
  • B. More Explanations
  • B.1. Data Workflow
  • B.2. Differences to Symbolic Regression Task
  • C. More Experiments
  • C.1. Symbolic Regression
  • C.2. Longer Iteration
  • D. More Results
  • D.1. Constitutive Law Discovery (a)
  • D.2. Constitutive Law Discovery (b)
  • D.3. Constitutive Law Discovery (c)
  • D.4. Constitutive Law Discovery (d)
  • D.5. Molecule Design (e)
  • D.6. Molecule Design (f)
  • D.7. Molecule Design (g)
  • D.8. Molecule Design (h)
  • D.9. Imaginary Constitutive Law

Knowls

  1. Knowl 1 — Bilevel Optimization Formulation of Scientific Generative Agent

    model/method

    Scientific Generative Agent (SGA) frames automated physical scientific discovery as a bilevel optimization problem where an outer loop driven by a Large Language Model (LLM) searches over discrete hypotheses and continuous parameterizations, while an inner loop powered by differentiable physical simulation optimizes continuous parameters.

    Let Φ\Phi denote a simulator that takes as inputs a discrete scientific expression E\mathcal{E} (e.g., symbolic constitutive equations or molecular topology) and a vector of continuous physical parameters θ∈Θ\theta \in \Theta (e.g., material stiffness parameters or atomic coordinates). The simulator produces simulated physical phenomena yy and additional observational feedback zz (such as particle trajectories): y,z=Φ(θ;E)y, z = \Phi(\theta; \mathcal{E})

    Given an objective metric L(y)\mathcal{L}(y) measuring error against ground-truth phenomena (such as kinematic trajectory reconstruction error or deviation from target quantum mechanical properties), the bilevel optimization problem is defined as: min⁡E,ΘL(y(E,Θ,θ^;Φ))\min_{\mathcal{E}, \Theta} \mathcal{L}\left(y\left(\mathcal{E}, \Theta, \hat{\theta}; \Phi\right)\right) s.t.G(E,Θ;Φ)≤0\text{s.t.} \quad G(\mathcal{E}, \Theta; \Phi) \le 0 θ^∈arg⁡min⁡θ∈ΘL(y(θ;Φ,E))\hat{\theta} \in \arg\min_{\theta \in \Theta} \mathcal{L}(y(\theta; \Phi, \mathcal{E})) where:

    • E\mathcal{E} is the discrete symbolic scientific hypothesis proposed by the LLM.
    • Θ\Theta is the continuous parameterization space specified by the LLM (defining which decision variables the inner continuous optimizer will update).
    • G(E,Θ;Φ)≤0G(\mathcal{E}, \Theta; \Phi) \le 0 enforces simulation validity constraints (i.e., whether E\mathcal{E} is simulatable, well-formed, and executable in Φ\Phi).
    • θ^\hat{\theta} is the optimal continuous parameter vector obtained by gradient descent through the differentiable simulation.

    The LLM generates candidate updates (E,Θ)(\mathcal{E}, \Theta) conditioned on a prompt PP and a history of top-KK previous proposals and experimental outcomes: E,Θ=LLM({L(yk),zk,ok,Ek,Θk}k∈[K];P)\mathcal{E}, \Theta = \text{LLM}\left(\{\mathcal{L}(y_k), z_k, o_k, \mathcal{E}_k, \Theta_k\}_{k \in [K]}; P\right) where oko_k summarizes intermediate inner-optimization results (such as loss curves and optimization trajectories), zkz_k represents physical observations, and k∈[K]k \in [K] indexes the retained historical solutions.

  2. Knowl 2 — Scientific Generative Agent (SGA) Search Algorithm

    algorithm

    Scientific Generative Agent (SGA) executes an iterative evolutionary search combining an LLM outer proposer with a gradient-based inner optimizer. The outer loop maintains a top-KK priority heap HH of ranked (expression, parameter) tuples. In each iteration, it queries the LLM to generate MlM_l exploiting candidate solutions at a lower sampling temperature TlT_l and MhM_h exploring candidate solutions at a higher sampling temperature ThT_h. Differentiable simulation then optimizes continuous parameters θ^∈Θ\hat{\theta} \in \Theta for each proposed candidate before appending them to the heap.

    Input: Initial discrete expression and continuous parameters (E,θ∈Θ\mathcal{E}, \theta \in \Theta), number of exploitation candidates MlM_l, number of exploration candidates MhM_h, exploitation temperature TlT_l, exploration temperature ThT_h, number of iterations NN, top-KK history size KK, differentiable simulator Φ\Phi.
    Output: Optimal solution (E∗,θ^∗)(\mathcal{E}^*, \hat{\theta}^*).
    H←heap()H \leftarrow \text{heap}()
    θ^←optim(E,θ;Φ)\hat{\theta} \leftarrow \text{optim}(\mathcal{E}, \theta; \Phi)
    H.append((E,θ^))H.\text{append}((\mathcal{E}, \hat{\theta}))
    for i=1,…,Ni = 1, \dots, N do
        (E,Θ)[:Ml]←LLM(H.topk(K),Tl)(\mathcal{E}, \Theta)[:M_l] \leftarrow \text{LLM}(H.\text{topk}(K), T_l)
        (E,Θ)[Ml:Ml+Mh]←LLM(H.topk(K),Th)(\mathcal{E}, \Theta)[M_l : M_l + M_h] \leftarrow \text{LLM}(H.\text{topk}(K), T_h)
        for m=1,…,Ml+Mhm = 1, \dots, M_l + M_h do
            θ^←optim(Em,θ∈Θm;Φ)\hat{\theta} \leftarrow \text{optim}(\mathcal{E}_m, \theta \in \Theta_m; \Phi)
            H.append((Em,θ^))H.\text{append}((\mathcal{E}_m, \hat{\theta}))
        end for
    end for
    return H.topk(1)H.\text{topk}(1)

    In standard experimental settings:

    • Number of outer iterations: N=5N = 5 (or N=20N = 20 in prolonged experiments).
    • Heap history capacity: K=5K = 5.
    • Candidate offspring split: Ml=4M_l = 4 exploitation candidates, Mh=12M_h = 12 exploration candidates (total M=16M = 16).
    • Sampling temperatures: Tl=0.5T_l = 0.5 (exploit) and Th=1.0T_h = 1.0 (explore).
    • Inner optimizer: Adam optimizer using Mean Squared Error (MSE) loss through differentiable simulation.
  3. Knowl 3 — Exploration and Exploitation Strategy via Temperature Tuning

    model/method

    To balance conservative refinement of high-performing hypotheses with aggressive discovery of novel structures, SGA partitions the generation of candidate offspring in each iteration into two distinct groups by adjusting the LLM decoding temperature:

    1. Exploitation group (MexploitM_{\text{exploit}}): Generated with a low decoding temperature (Tl=0.5T_l = 0.5). These candidates make cautious, incremental modifications to the best solutions retrieved from the top-KK history heap, following observed loss gradients and refining existing parameter structures.
    2. Exploration group (MexploreM_{\text{explore}}): Generated with a higher decoding temperature (Th=1.0T_h = 1.0). These candidates propose daring structural modifications or entirely new functional forms to escape local minima in complex, non-linear search spaces.

    Empirically, generating purely exploitative offspring leads to repetitive, redundant proposals across iterations, while generating purely exploratory offspring leads to noisy, uninformative, or invalid code expressions that violate simulation constraints. A ratio of 1:31:3 between MexploitM_{\text{exploit}} and MexploreM_{\text{explore}} (specifically 4 exploiting candidates and 12 exploring candidates per iteration) achieves a reliable balance between valid simulation execution rates and global optimization effectiveness.

  4. Knowl 4 — Problem Formulations for Constitutive Law Discovery and Molecular Design

    experimental setup

    The SGA framework is evaluated on two physical scientific discovery domains:

    1. Constitutive Law Discovery: The objective is to identify both the discrete symbolic material model φ(⋅)\varphi(\cdot) and continuous material parameters θ\theta from observed particle trajectories X^t∈[1,…,T]\hat{X}_{t \in [1,\dots,T]} across TT time steps generated by a Material Point Method (MPM) simulator implemented in Warp. The models consider:

    • Elastic constitutive law: φE(F;θE)↦τ\varphi_E(F; \theta_E) \mapsto \tau, where F∈R3×3F \in \mathbb{R}^{3\times 3} is the deformation gradient, τ∈R3×3\tau \in \mathbb{R}^{3\times 3} is the Kirchhoff stress tensor, and θE\theta_E are continuous parameters (such as Young's modulus and Poisson's ratio).
    • Plastic constitutive law: φP(F;θP)↦Fcorrected\varphi_P(F; \theta_P) \mapsto F_{\text{corrected}}, where Fcorrected∈R3×3F_{\text{corrected}} \in \mathbb{R}^{3\times 3} is the plastic return-mapping corrected deformation gradient and θP\theta_P are plastic parameters. The output trajectory Xt∈[1,…,T]=sim(φ(⋅;θ))X_{t \in [1,\dots,T]} = \text{sim}(\varphi(\cdot; \theta)) is optimized by minimizing MSE loss against X^t∈[1,…,T]\hat{X}_{t \in [1,\dots,T]}. Four tasks are evaluated: (a) non-linear elastic material starting from linear elastic, (b) von Mises plastic material starting from purely elastic, (c) granular material starting from purely elastic, and (d) weakly compressible fluid starting from purely elastic.

    2. Molecular Design: The goal is to discover molecular structures (represented as SMILES strings) and 3D atomic coordinates matching target quantum mechanical properties normalized on the QM9 dataset. In the outer loop, the LLM proposes SMILES strings and initial 3D coordinate guesses. In the inner loop, 3D conformations are generated via the RDKit ETKGD algorithm followed by MMFF force field energy minimization, and atom coordinates are optimized using the UniMol transformer model to match target property values. Four tasks are evaluated: (e) Highest Occupied Molecular Orbital (HOMO) set to 00, (f) Lowest Unoccupied Molecular Orbital (LUMO) set to 00, (g) HOMO-LUMO energy gap set to 00, and (h) HOMO-LUMO energy gap set to −2-2.

  5. Knowl 5 — Benchmark Results on Constitutive Law Search and Molecular Design

    data/table

    SGA was evaluated across 8 tasks: 4 constitutive law discovery tasks ((a)-(d)) and 4 molecular design tasks ((e)-(h)). It was compared against baseline LLM-based prompting and search methods (Chain-of-Thought [CoT], FunSearch, Eureka, OPRO) and two ablation variants of SGA: one without bilevel optimization (Ours (no bilevel)) and one without exploitation temperature sampling (Ours (no exploit)).

    Method #Iter. #Hist. Bilevel Constitutive Law Search Molecule Design
    (a) ↓\downarrow (b) ↓\downarrow (c) ↓\downarrow (d) ↓\downarrow (e) ↓\downarrow (f) ↓\downarrow (g) ↓\downarrow (h) ↓\downarrow
    CoT 1 5 N/A X 298.5 1462.3 150.0 384.1 3.0 32.1 18.6 6.0
    FunSearch 20 2 0 / 4 X 210.3 872.2 82.8 139.5 1.1 7.1 8.3 1.1
    Eureka 5 1 0 / 16 X 128.0 531.0 101.7 150.1 4.3 9.8 3.3 9.7e-1
    OPRO 5 5 0 / 16 X 136.2 508.3 99.2 128.8 2.4 9.4 3.1 1.3
    Ours (no bilevel) 5 5 4 / 12 X 90.2 517.0 83.6 68.4 8.6e-1 9.1 1.8 1.4
    Ours (no exploit) 5 5 0 / 16 ✓ 3.0e-3 3.9e-1 6.6e-2 1.4e-12 4.0e-4 1.5e-1 6.1e-1 2.8e-5
    Ours 5 5 4 / 12 ✓ 5.2e-5 2.1e-1 6.0e-2 1.4e-12 1.3e-4 1.1e-1 5.4e-1 3.6e-5

    Key findings:

    1. SGA achieves lower MSE loss by multiple orders of magnitude compared to non-bilevel baselines across all 8 tasks (e.g., 5.2×10−55.2 \times 10^{-5} vs 128.0128.0 on task (a), and 1.3×10−41.3 \times 10^{-4} vs 1.11.1 on task (e)).
    2. Removing bilevel optimization causes drastic performance degradation across all tasks, demonstrating the necessity of differentiable simulation for continuous parameter tuning.
    3. Incorporating the exploit-and-explore strategy (4 exploit / 12 explore) outperforms purely exploratory optimization (0 exploit / 16 explore) by over 50% on challenging tasks such as von Mises plasticity discovery (b) and LUMO target matching (f).
  6. Knowl 6 — Generalization Verification via Imaginary Constitutive Law Invention

    empirical result

    To test whether SGA discovers physical laws via systematic exploration and simulation feedback rather than memorization of existing training data, an imaginary constitutive law was created by synthetically blending three distinct material behaviors: von Mises plasticity (50%50\%), granular material (30%30\%), and weakly compressible fluid (20%20\%).

    Under identical optimization settings, SGA was evaluated against baseline LLM optimizers on recovering this synthetic non-physical material from observed motion trajectories:

    Method FunSearch Eureka OPRO Ours
    Loss ↓\downarrow 105.0 89.1 98.0 1.3e-3

    SGA achieved an MSE loss of 1.3×10−31.3 \times 10^{-3}, outperforming FunSearch (105.0105.0), Eureka (89.189.1), and OPRO (98.098.0) by several orders of magnitude. The discovered model produced a motion trajectory visually indistinguishable from the ground truth, confirming that the bilevel optimization framework generalizes to completely novel, unrecorded physical equations outside the LLM's pretraining corpus.

  7. Knowl 7 — Benchmark Comparison Against Symbolic Regression and Molecule Design Baselines

    data/table

    To evaluate symbolic equation discovery against specialized regression baselines, task (a) (non-linear constitutive law discovery) was reformulated as a direct symbolic regression problem where stress tensors served as direct target outputs and the 9 output dimensions were split into 9 independent problems without back-propagation through time (BPTT). SGA (operating under its full physical simulation setting) was compared against 14 SRBench baselines and 3 pretrained symbolic regression models:

    Method R2 ↑\uparrow MSE ↓\downarrow MAE ↓\downarrow Symbolic
    AIFeynman 0.05105 22814675.8 2520.0 ✓
    DSR 0.57527 10966411.0 2045.0 ✓
    BSR 0.66526 8642965.0 1938.6 ✓
    AdaBoost 0.75058 6439962.9 1777.7 X
    GP-GOMEA 0.77734 5749076.4 1580.1 ✓
    SBP-GP 0.81773 4706077.0 1367.5 ✓
    LightGBM 0.83368 4294433.7 1129.9 X
    XGBoost 0.87775 3156500.5 1109.2 X
    MRGP 0.91074 2304682.5 950.5 ✓
    EPLEX 0.91851 2104070.1 122.2 ✓
    FFX 0.93124 1775263.7 801.7 ✓
    MLP 0.98240 454461.5 366.3 X
    FEAT 0.98761 319800.6 336.1 ✓
    DSO 0.99642 92374.9 168.6 ✓
    Operon 0.99684 81577.9 92.4 ✓
    SymbolicGPT 0.52333 6862154.7 1680.7 ✓
    NeSymReS N/A to >3>3 variables ✓
    T-JSL N/A to >2>2 variables ✓
    Ours 0.99901 17424.6 86.4 ✓

    SGA achieved the highest R2R^2 (0.999010.99901), lowest MSE (1.74×1041.74 \times 10^4), and lowest MAE (86.486.4) among all evaluated methods. Pretrained neural symbolic regression models (NeSymReS and T-JSL) failed due to input dimension constraints (>3>3 and >2>2 variables respectively), whereas SGA handled multi-dimensional tensor operations natively through general code representation.

    Furthermore, when compared on molecular design tasks ((e)-(h)) against ChemGE (a population-based molecular generation baseline), SGA achieved lower target error: on task (e) 1.3×10−41.3 \times 10^{-4} vs 4.8×10−34.8 \times 10^{-3}; on task (f) 1.1×10−11.1 \times 10^{-1} vs 1.81.8; on task (g) 5.4×10−15.4 \times 10^{-1} vs 1.51.5; and on task (h) 3.6×10−53.6 \times 10^{-5} vs 9.8×10−59.8 \times 10^{-5}.

  8. Knowl 8 — Effect of Extended Optimization Iterations on SGA Performance

    empirical result

    Extending the number of outer question-answering cycles NN in SGA from 5 to 20 iterations yields substantial improvements on complex and constrained discovery problems, while easier tasks converge early:

    #Iterations (a) ↓\downarrow (b) ↓\downarrow (c) ↓\downarrow (d) ↓\downarrow (e) ↓\downarrow (f) ↓\downarrow (g) ↓\downarrow (h) ↓\downarrow
    5 5.2e-5 2.1e-1 6.0e-2 1.4e-12 1.3e-4 1.1e-1 5.4e-1 3.6e-5
    20 4.2e-6 4.0e-4 2.5e-3 1.4e-12 1.3e-4 6.5e-2 1.2e-1 5.6e-6
    Improvement +1138.1% +52400.0% +2300.0% 0.0% 0.0% +69.2% +350.0% +542.9%

    The largest performance gains occur on the most challenging physical discovery tasks: von Mises plasticity discovery (b) improves by +52400.0%+52400.0\% (loss dropping from 0.210.21 to 4.0×10−44.0 \times 10^{-4}), granular material discovery (c) improves by +2300.0%+2300.0\%, and non-linear elastic discovery (a) improves by +1138.1%+1138.1\%. Tasks that quickly saturate or reach machine precision (such as fluid discovery (d) at 1.4×10−121.4 \times 10^{-12} and HOMO matching (e) at 1.3×10−41.3 \times 10^{-4}) show 0.0%0.0\% additional change.

  9. Knowl 9 — Comparative Performance of LLM Backbones in SGA

    empirical result

    SGA was evaluated across all 8 benchmark tasks using four distinct LLM backbones: gpt-4-turbo-preview (GPT-4), gpt-3.5-turbo (GPT-3.5), Claude-3-Sonnet, and Mixtral-8x7B.

    Key findings include:

    1. GPT-4 achieves the best overall performance across all tasks, achieving the highest aggregate rank across the entire benchmark suite.
    2. Claude-3-Sonnet consistently achieves the second-highest performance on constitutive law search tasks ((a)-(d)).
    3. Mixtral-8x7B (an open-source mixture-of-experts model) achieves competitive results and tops the ranking on two molecular design tasks ((e) and (f)), demonstrating that SGA functions effectively with open-weights models.
    4. GPT-3.5 consistently ranks lowest across both domain categories due to weaker complex physical reasoning and higher syntax error rates in generating valid simulation code.
  10. Knowl 10 — Limitations of Scientific Generative Agent

    limitation

    The Scientific Generative Agent framework has several identified limitations:

    1. Interpretability: Although prompts instruct the LLM to generate step-by-step reasoning, plans, and code comments, the resulting symbolic expressions and non-linear mathematical constructs are not guaranteed to be human-interpretable.
    2. AI Safety and Code Execution: Generated code snippets are directly executed in the simulation environment without automated sandboxing or verification filtering, creating potential security and execution hazards on host systems.
    3. Absence of Domain-Specific Invariance Constraints: SGA relies entirely on the internal knowledge priors of the LLM rather than enforcing physical principles (e.g., thermodynamic consistency, frame indifference, symmetry constraints) via formal hard-coded rules.
    4. Dependence on Simulation Differentiability: The inner-level continuous optimization requires differentiable simulators to compute parameter gradients ∇θΦ(θ;E)\nabla_\theta \Phi(\theta; \mathcal{E}). While zero-order optimization methods could be applied, their scalability to high-dimensional parameter spaces remains unproven in this framework.
    5. Inference Cost: Operating the outer loop with commercial LLMs is computationally expensive; completing a single 5-iteration discovery task with GPT-4 costs approximately $10, scaling linearly with the number of iterations.
    6. Redundant Context Processing: Neighboring iterations submit largely overlapping prompts containing the top-KK heap history, creating repetitive key-value (KV) cache computation that requires specialized KV caching optimizations to reduce latency.

Coverage note — No substantial contributed material was omitted; full experimental details, prompts, algorithms, and comparative results are comprehensively covered.

References

  1. 1.AI4Science, M. R. and Quantum, M. A. The impact of large language models on scientific discovery: a preliminary study using gpt-4. arXiv preprint arXiv:2311.07361, 2023.
  2. 2.Anthropic. Introducing the next generation of claude, 2024. URL https://www.anthropic.com/news/claude-3-family.
  3. 3.Arnaldo, I., Krawiec, K., and O’Reilly, U.-M. Multiple regression genetic programming. In Proceedings of the 2014 Annual Conference on Genetic and Evolutionary Computation, pp. 879–886, 2014.
  4. 4.Bender, G., Kindermans, P.-J., Zoph, B., Vasudevan, V., and Le, Q. Understanding and simplifying one-shot architecture search. In International conference on machine learning, pp. 550–559. PMLR, 2018.
  5. 5.Biggio, L., Bendinelli, T., Neitz, A., Lucchi, A., and Parascandolo, G. Neural symbolic regression that scales. In International Conference on Machine Learning, pp. 936–945. Pmlr, 2021.
  6. 6.Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  8. 8.Cai, H., Zhu, L., and Han, S. ProxylessNAS: Direct neural architecture search on target task and hardware. In International Conference on Learning Representations, 2019.
  9. 9.Cai, T., Wang, X., Ma, T., Chen, X., and Zhou, D. Large language models as tool makers. In International Conference on Learning Representations, 2024.
  10. 10.Cava, W. L., Singh, T. R., Taggart, J., Suri, S., and Moore, J. Learning concise representations for regression by evolving networks of trees. In International Conference on Learning Representations, 2019.
  11. 11.Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pp. 785–794, 2016.
  12. 12.Chithrananda, S., Grand, G., and Ramsundar, B. Chemberta: large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885, 2020.
  13. 13.Colson, B., Marcotte, P., and Savard, G. An overview of bilevel optimization. Annals of operations research, 153: 235–256, 2007.
  14. 14.Du, T., Wu, K., Ma, P., Wah, S., Spielberg, A., Rus, D., and Matusik, W. Diffpd: Differentiable projective dynamics. ACM Transactions on Graphics (TOG), 41(2):1–21, 2021.
  15. 15.Fang, X., Liu, L., Lei, J., He, D., Zhang, S., Zhou, J., Wang, F., Wu, H., and Wang, H. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4(2):127–134, 2022.
  16. 16.Fortunato, S., Bergstrom, C. T., Börner, K., Evans, J. A., Helbing, D., Milojević, S., Petersen, A. M., Radicchi, F., Sinatra, R., Uzzi, B., et al. Science of science. Science, 359(6379):eaao0185, 2018.
  17. 17.Halgren, T. A. Merck molecular force field. i. basis, form, scope, parameterization, and performance of mmff94. Journal of computational chemistry, 17(5-6):490–519, 1996.
  18. 18.Hansen, N. The cma evolution strategy: a comparing review. Towards a new evolutionary computation: Advances in the estimation of distribution algorithms, pp. 75–102, 2006.
  19. 19.Huang, Q., Vora, J., Liang, P., and Leskovec, J. Benchmarking large language models as ai research agents. arXiv preprint arXiv:2310.03302, 2023.
  20. 20.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  21. 21.Jiang, C., Schroeder, C., Teran, J., Stomakhin, A., and Selle, A. The material point method for simulating continuum materials. In Acm siggraph 2016 courses, pp. 1–52. 2016.
  22. 22.Jin, W., Barzilay, R., and Jaakkola, T. Junction tree variational autoencoder for molecular graph generation. In International conference on machine learning, pp. 2323–2332. PMLR, 2018.
  23. 23.Jin, Y., Fu, W., Kang, J., Guo, J., and Guo, J. Bayesian symbolic regression. arXiv preprint arXiv:1910.08892, 2019.
  24. 24.Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems, 30, 2017.
  25. 25.Kingma, D. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations, San Diega, CA, USA, 2015.
  26. 26.Kommenda, M., Burlacu, B., Kronberger, G., and Affenzeller, M. Parameter identification for symbolic regression using nonlinear least squares. Genetic Programming and Evolvable Machines, 21(3):471–501, 2020.
  27. 27.Kramer, S., Cerrato, M., Džeroski, S., and King, R. Automated scientific discovery: From equation discovery to autonomous discovery systems. arXiv preprint arXiv:2305.02251, 2023.
  28. 28.La Cava, W., Helmuth, T., Spector, L., and Moore, J. H. A probabilistic and multi-objective analysis of lexicase selection and ε-lexicase selection. Evolutionary Computation, 27(3):377–402, 2019.
  29. 29.La Cava, W., Orzechowski, P., Burlacu, B., de Franca, F., Virgolin, M., Jin, Y., Kommenda, M., and Moore, J. Contemporary symbolic regression methods and their relative performance. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021.
  30. 30.Landrum, G. et al. Rdkit: A software suite for cheminformatics, computational chemistry, and predictive modeling. Greg Landrum, 8:31, 2013.
  31. 31.Li, J., Liu, Y., Fan, W., Wei, X.-Y., Liu, H., Tang, J., and Li, Q. Empowering molecule discovery for moleculecaption translation with large language models: A chatgpt perspective. arXiv preprint arXiv:2306.06615, 2023.
  32. 32.Li, W., Li, W., Sun, L., Wu, M., Yu, L., Liu, J., Li, Y., and Tian, S. Transformer-based model for symbolic regression via joint supervised learning. In The Eleventh International Conference on Learning Representations, 2022.
  33. 33.Liu, H., Simonyan, K., and Yang, Y. DARTS: Differentiable architecture search. In International Conference on Learning Representations, 2019.
  34. 34.Liu, R., Mu, P., Yuan, X., Zeng, S., and Zhang, J. A general descent aggregation framework for gradient-based bi-level optimization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):38–57, 2022.
  35. 35.Liu, Z., Roberts, R. A., Lal-Nag, M., Chen, X., Huang, R., and Tong, W. Ai-based language models powering drug discovery and development. Drug Discovery Today, 26 (11):2593–2607, 2021.
  36. 36.Ma, P., Du, T., Tenenbaum, J. B., Matusik, W., and Gan, C. Risp: Rendering-invariant state predictor with differentiable simulation and rendering for cross-domain parameter estimation. In International Conference on Learning Representations, 2021.
  37. 37.Ma, P., Chen, P. Y., Deng, B., Tenenbaum, J. B., Du, T., Gan, C., and Matusik, W. Learning neural constitutive laws from motion observations for generalizable pde dynamics. In International Conference on Machine Learning. PMLR, 2023.
  38. 38.Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A. Eureka: Human-level reward design via coding large language models. In International Conference on Learning Representations, 2024.
  39. 39.Macklin, M. Warp: A high-performance python framework for gpu simulation and graphics, March 2022. NVIDIA GPU Technology Conference.
  40. 40.McConaghy, T. Ffx: Fast, scalable, deterministic symbolic regression technology. Genetic Programming Theory and Practice IX, pp. 235–260, 2011.
  41. 41.Mundhenk, T. N., Landajuela, M., Glatt, R., Santiago, C. P., Faissol, D. M., and Petersen, B. K. Symbolic regression via neural-guided genetic programming population seeding. In Advances in Neural Information Processing Systems, 2021.
  42. 42.OpenAI. OpenAI: Introducing ChatGPT, 2022. URL https://openai.com/blog/chatgpt.
  43. 43.OpenAI. OpenAI: GPT-4, 2023. URL https://openai.com/research/gpt-4.
  44. 44.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
  45. 45.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  46. 46.Petersen, B. K., Larma, M. L., Mundhenk, T. N., Santiago, C. P., Kim, S. K., and Kim, J. T. Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients. In International Conference on Learning Representations, 2020.
  47. 47.Popper, K. The logic of scientific discovery. Routledge, 2005.
  48. 48.Ramakrishnan, R., Dral, P. O., Rupp, M., and Von Lilienfeld, O. A. Quantum chemistry structures and properties of 134 kilo molecules. Scientific data, 1(1):1–7, 2014.
  49. 49.Riniker, S. and Landrum, G. A. Better informed distance geometry: using what we know to improve conformation generation. Journal of chemical information and modeling, 55(12):2562–2574, 2015.
  50. 50.Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al. Mathematical discoveries from program search with large language models. Nature, pp. 1–3, 2023.
  51. 51.Rosenberg, A. and McIntyre, L. Philosophy of science: A contemporary introduction. Routledge, 2019.
  52. 52.Schapire, R. E. The boosting approach to machine learning: An overview. Nonlinear estimation and classification, pp. 149–171, 2003.
  53. 53.Schneider, G. Automating drug discovery. Nature reviews drug discovery, 17(2):97–113, 2018.
  54. 54.Sharma, G. and Thakur, A. Chatgpt in drug discovery. 2023.
  55. 55.Sinha, A., Malo, P., and Deb, K. Evolutionary algorithm for bilevel optimization using approximations of the lower level optimal solution mapping. European Journal of Operational Research, 257(2):395–411, 2017a.
  56. 56.Sinha, A., Malo, P., and Deb, K. A review on bilevel optimization: From classical to evolutionary approaches and applications. IEEE Transactions on Evolutionary Computation, 22(2):276–295, 2017b.
  57. 57.Sulsky, D., Zhou, S.-J., and Schreyer, H. L. Application of a particle-in-cell method to solid mechanics. Computer physics communications, 87(1-2):236–252, 1995.
  58. 58.Sumers, T., Yao, S., Narasimhan, K., and Griffiths, T. Cognitive architectures for language agents. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. Survey Certification.
  59. 59.Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476–482, 2024.
  60. 60.Udrescu, S.-M., Tan, A., Feng, J., Neto, O., Wu, T., and Tegmark, M. Ai feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularity. Advances in Neural Information Processing Systems, 33:4860–4871, 2020.
  61. 61.Valipour, M., You, B., Panju, M., and Ghodsi, A. Symbolicgpt: A generative transformer model for symbolic regression. arXiv preprint arXiv:2106.14131, 2021.
  62. 62.Virgolin, M., Alderliesten, T., and Bosman, P. A. Linear scaling with and within semantic backpropagation-based genetic programming for symbolic regression. In Proceedings of the genetic and evolutionary computation conference, pp. 1084–1092, 2019.
  63. 63.Virgolin, M., Alderliesten, T., Witteveen, C., and Bosman, P. A. Improving model-based genetic programming for symbolic regression of small expressions. Evolutionary computation, 29(2):211–237, 2021.
  64. 64.Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Van Katwyk, P., Deac, A., et al. Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60, 2023.
  65. 65.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022.
  66. 66.Weininger, D. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36, 1988.
  67. 67.Wuestman, M., Hoekman, J., and Frenken, K. A typology of scientific breakthroughs. Quantitative Science Studies, 1(3):1203–1222, 2020.
  68. 68.Xue, C., Wang, X., Yan, J., Hu, Y., Yang, X., and Sun, K. Rethinking bi-level optimization in neural architecture search: A gibbs sampling perspective. In AAAI Conference on Artificial Intelligence, volume 35, pp. 10551–10559, 2021.
  69. 69.Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. Large language models as optimizers. In International Conference on Learning Representations, 2024.
  70. 70.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R. Tree of thoughts: Deliberate problem solving with large language models. In Conference on Neural Information Processing Systems, 2023a.
  71. 71.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations, 2023b.
  72. 72.Yoshikawa, N., Terayama, K., Sumita, M., Homma, T., Oono, K., and Tsuda, K. Population-based de novo molecule generation, using grammatical evolution. Chemistry Letters, 47(11):1431–1434, 2018.
  73. 73.Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al. Efficiently programming large language models using sglang. arXiv preprint arXiv:2312.07104, 2023.
  74. 74.Zhou, G., Gao, Z., Ding, Q., Zheng, H., Xu, H., Wei, Z., Zhang, L., and Ke, G. Uni-mol: A universal 3d molecular representation learning framework. In International Conference on Learning Representations, 2023.
  75. 75.Zhou, Z., Kearnes, S., Li, L., Zare, R. N., and Riley, P. Optimization of molecules via deep reinforcement learning. Scientific reports, 9(1):10752, 2019.

Citation

MLA
Ma, P., et al. “LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery”. arXiv, 2024, http://arxiv.org/abs/2405.09783v1.
APA
Ma, P., Wang, T.-H., Guo, M., Sun, Z., Tenenbaum, J. B., Rus, D., Gan, C., & Matusik, W. (2024). LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery. arXiv. http://arxiv.org/abs/2405.09783v1
Chicago
Ma, P., T.-H. Wang, M. Guo, et al. 2024. “LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery”. arXiv. http://arxiv.org/abs/2405.09783v1.
Harvard
Ma, P. et al. (2024) “LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.09783v1.
Vancouver
1. Ma P, Wang T-H, Guo M, Sun Z, Tenenbaum JB, Rus D, Gan C, Matusik W (2024) LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery. arXiv

BibTeX

@article{ma2024llm,
  title = {LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery},
  author = {Ma, Pingchuan and Wang, Tsun-Hsuan and Guo, Minghao and Sun, Zhiqing and Tenenbaum, Joshua B. and Rus, Daniela and Gan, Chuang and Matusik, Wojciech},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.09783v1},
  eprint = {2405.09783}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/