Symbolic Regression with a Learned Concept Library
Arya GrayeliAtharva SehgalOmar Costilla-ReyesMiles D. CranmerSwarat Chaudhuri
Proposes LASR, a symbolic regression framework that guides genetic algorithms with large language models to extract reusable conceptual abstractions, achieving state-of-the-art accuracy on the Feynman benchmark and discovering a novel scaling law for language models.
Automated scientific discovery relies heavily on symbolic regression to uncover compact, interpretable mathematical formulas from raw data. However, searching through the exponentially large space of possible mathematical expressions poses a severe computational bottleneck. Traditional genetic algorithms explore this space primarily through random mutation and crossover operations, unlike human scientists who naturally synthesize observations into higher-level concepts and use domain intuition to guide their investigations.
The article introduces and evaluates LASR, a novel symbolic regression method designed to accelerate scientific formula discovery. The primary objective is to demonstrate that integrating an evolving library of abstract, natural-language concepts learned via large language models into genetic programming substantially improves the discovery of precise mathematical hypotheses.
The approach combines standard evolutionary search from the PySR framework with zero-shot queries to large language models across three alternating phases. In the first phase, genetic expression search is periodically guided by language-model operations conditioned on current concepts. In the second phase, Pareto-optimal and low-performing expressions are summarized by the language model into new qualitative textual concepts. In the third phase, these concepts are further evolved and generalized. The authors evaluated the system on 100 physics equations from the Feynman Lectures under noisy conditions, a benchmark of 41 synthetic equations designed to prevent data memorization, and a real-world machine learning task using 53,812 evaluations from the BIG-Bench benchmark.
The findings show that LASR achieves an exact match solve rate of 72% on the Feynman benchmark, outperforming the previous leading baseline of 59% and deep learning alternatives ranging between 20% and 40%. Even with a small open-source model and minimal language guidance, the method solved 67% of the equations. On the synthetic benchmark, LASR achieved a high predictive accuracy of 0.913 compared to 0.070 for the standard genetic baseline, confirming that performance gains stem from concept-guided reasoning rather than data memorization. Furthermore, when applied to language model scaling behaviors, LASR successfully discovered an empirical scaling law that matches established formulations like Chinchilla while requiring only three free parameters instead of five.
These results indicate that concept libraries introduce meaningful semantic biases that help algorithms escape local minima and filter out unnecessary mathematical terms, producing cleaner, more generalizable formulas. This allows researchers to discover complex empirical relationships without manually specifying rigid functional forms beforehand, reducing discovery timelines and engineering overhead.
Organizations exploring automated scientific discovery and empirical modeling should consider augmenting genetic search workflows with language-guided concept extraction. For practitioners adopting this method, incorporating high-level human hints can further accelerate convergence. Future development should focus on testing the approach with larger compute budgets, exploring model fine-tuning, and adapting the concept induction mechanism to broader scientific optimization tasks.
Confidence in the system's search capabilities is high across the evaluated benchmarks, but practical users must exercise caution. The system does not guarantee that induced concepts are scientifically accurate or mutually consistent, meaning qualitative concepts should be treated as heuristic search drivers rather than proven domain principles until formally verified.
- Paper: AI Feynman: A physics-inspired method for symbolic regression, Silviu-Marian Udrescu et al. (2019). This paper establishes the standard physics benchmark dataset (the Feynman equations) and foundational neural-symbolic methodology that LASR builds upon and directly evaluates against.
- Paper: Gene Expression Programming: A New Adaptive Algorithm for Solving Problems, Cândida Ferreira (2001). It provides foundational principles for evolutionary program representation and genetic algorithms in symbolic regression that LASR integrates with LLM-guided exploration.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). It demonstrates how large language models can generate functional programs from specifications, establishing the foundation for LLM-driven programmatic hypothesis search.
- Paper: Competition-level code generation with AlphaCode, Yujia Li et al. (2022). It establishes large-scale program search, sampling, and filtering strategies with language models that inform LASR's generation of candidate programmatic expressions.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). This work extends LLM-driven symbolic hypothesis generation into a bilevel framework coupled directly with differentiable physical simulations for discovering physical laws.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). This survey provides a comprehensive synthesis of how LLMs and foundation models are applied across scientific domains, placing symbolic regression and scientific discovery frameworks like LASR into a broader ecosystem.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). This work introduces Kolmogorov-Arnold Networks as an interpretable neural alternative for discovering mathematical formulas and physical laws from numerical data.
- Paper: Automated Design of Agentic Systems, Shengran Hu et al. (2025). This paper broadens the paradigm of evolutionary search over code representations using LLM meta-agents to design entire agentic architectures.
- Paper: ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models, Jinheon Baek et al. (2025). It scales LLM-driven scientific concept generation and iterative refinement from individual equations to full end-to-end research proposals.
