LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
Pingchuan MaTsun-Hsuan WangMinghao GuoZhiqing SunJoshua B. TenenbaumDaniela RusChuang GanWojciech Matusik
Proposes Scientific Generative Agent, a bilevel optimization framework that couples language models for discrete hypothesis generation with differentiable simulations for continuous parameter optimization to discover physical laws and design molecules.
Accelerating scientific discovery requires systems that can formulate valid hypotheses while grounding them in rigorous physical laws. Large language models possess extensive cross-disciplinary knowledge and reasoning capabilities, yet they struggle with numerical precision and simulating observational feedback on their own. Conversely, numerical simulations accurately model continuous physical behaviors but lack the creative reasoning needed to propose symbolic equations or novel molecular structures.
The article demonstrates the Scientific Generative Agent, a unified framework that combines large language models with differentiable physical simulations. The objective is to automate the discovery of discrete scientific structures—such as material equations and chemical formulas—alongside the optimization of their continuous numerical parameters.
To evaluate this framework, the authors implemented a bilevel optimization process tested on two complex physical problems: discovering constitutive material laws from motion trajectories and designing molecular structures to match target quantum mechanical properties. The outer level employs a large language model to propose discrete symbolic expressions (such as mathematical code or molecular string sequences) and define the continuous parameter search space. The inner level utilizes differentiable simulation engines and gradient-based algorithms to calibrate continuous parameters against physical observations and feed loss curves back to the language model. An evolutionary exploration-exploitation strategy regulates model temperature to balance conservative refinements with novel hypotheses across multiple iterations.
The primary findings demonstrate that this bilevel agent significantly outperforms existing automated discovery and symbolic regression methods. First, across eight benchmark tasks, the framework achieved errors orders of magnitude lower than standard prompt-based baselines. Second, in ablation tests on difficult tasks, removing the bilevel simulation feedback or the exploit-and-explore mechanism caused performance degradation of over 50%, highlighting the necessity of combining both levels. Third, in synthetic tests featuring a non-existent, imaginary material law, the model successfully recovered the hybrid behavior without relying on memorized training data. Fourth, extending optimization from 5 to 20 iterations yielded massive accuracy gains, improving performance by up to 52,400% on the most challenging plasticity tasks.
These findings imply that artificial intelligence can discover valid, unconventional scientific solutions that deviate from traditional human heuristics while remaining physically sound. By replacing rigid, domain-specific symbolic regression tools with code-generating models, the approach reduces the labor and time required to characterize complex materials and molecular candidates. GPT-4 served as the top-performing backbone model, though open-source alternatives also demonstrated viable baseline utility.
For practitioners and research organizations seeking to adopt this methodology, the authors recommend tuning the number of optimization iterations as the primary control lever for difficult discovery tasks and maintaining an approximate 1:3 balance between conservative and exploratory proposals. Prior to deploying such systems autonomously, organizations should implement code execution sandboxes, incorporate domain-specific constraints or human-in-the-loop feedback, and utilize key-value cache reuse to manage computational and financial costs, which averaged approximately $10 per task under commercial model pricing.
Confidence in these findings is strong across the tested simulation benchmarks in mechanics and molecular design. However, readers should note limitations: the system was evaluated within simulated environments rather than physical laboratory experiments, the interpretability of generated code remains challenging to guarantee, and the approach depends heavily on the availability of differentiable simulators.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Provides a comprehensive architectural foundation for LLM-based autonomous agents, clarifying the planning, memory, and tool-use modules that SGA builds upon.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Introduces the generative agent framework for synthesizing higher-level insights and maintaining persistent reflections, directly informing the agentic hypothesis-generation paradigm in SGA.
- Paper: Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data, Anuj Karpatne et al. (2016). Establishes the theory-guided data science paradigm that couples physical scientific principles with computational models, motivating the bilevel integration of simulations and machine learning.
- Paper: AI Feynman: A physics-inspired method for symbolic regression, Silviu-Marian Udrescu et al. (2019). Demonstrates physics-inspired symbolic discovery of governing equations, providing foundational context for SGA's discrete hypothesis reasoning over physics formulas.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Pioneers closed-loop reasoning and replanning using environmental feedback, which SGA adapts into simulation-driven observational feedback loops for scientific discovery.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). Surveys the integration of LLM reasoning cores with external environments and tools, framing how language models interact with simulation backends.
- Paper: The frontier of simulation-based inference, Kyle Cranmer et al. (2019). Surveys methods for inferring physical parameters and statistical feedback through simulation execution paths, underpinning the continuous optimization tier in SGA.
- Paper: Junction Tree Variational Autoencoder for Molecular Graph Generation, Wengong Jin et al. (2018). Introduces structured molecular graph representations for optimization, which provides relevant background for SGA's molecular design experiments.
- Paper: Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models, Kevin Murphy (2026). Extends LLM-driven scientific hypothesis generation by coupling language model proposals with Bayesian active experiment design to discover mechanistic world models efficiently.
- Paper: Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, Fiona Y. Wang et al. (2026). Advances agentic scientific discovery beyond fixed parameter and hypothesis spaces by formalizing category-theoretic self-revising frameworks for mechanics and materials workflows.
- Paper: Accelerating Scientific Research with Gemini in the Real-World, Samuel Schmidgall et al. (2026). Builds upon computational and simulated discovery paradigms by expanding multi-agent LLM systems to interface with real-world laboratory equipment and physical experimental execution.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). Expands LLM scientific discovery architectures into the multimodal regime to reason over complex experimental artifacts, electron micrographs, and chemical structures.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). Generalizes the bilevel search across code and agentic workflows by using meta-agents to programmatically design and optimize end-to-end agent architectures.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Provides a comprehensive post-2024 synthesis of agentic reasoning, planning, and interactive tool-use frameworks applicable to complex scientific domains.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). Applies agent-driven scientific research workflows to systematically structure and generate complete research narratives from abstract conceptual ideas.
