Symbolic Regression with a Learned Concept Library

Arya GrayeliAtharva SehgalOmar Costilla-ReyesMiles D. CranmerSwarat Chaudhuri

article2024NeurIPS103 citations

Proposes LASR, a symbolic regression framework that guides genetic algorithms with large language models to extract reusable conceptual abstractions, achieving state-of-the-art accuracy on the Feynman benchmark and discovering a novel scaling law for language models.

Listen

Automated scientific discovery relies heavily on symbolic regression to uncover compact, interpretable mathematical formulas from raw data. However, searching through the exponentially large space of possible mathematical expressions poses a severe computational bottleneck. Traditional genetic algorithms explore this space primarily through random mutation and crossover operations, unlike human scientists who naturally synthesize observations into higher-level concepts and use domain intuition to guide their investigations.

The article introduces and evaluates LASR, a novel symbolic regression method designed to accelerate scientific formula discovery. The primary objective is to demonstrate that integrating an evolving library of abstract, natural-language concepts learned via large language models into genetic programming substantially improves the discovery of precise mathematical hypotheses.

The approach combines standard evolutionary search from the PySR framework with zero-shot queries to large language models across three alternating phases. In the first phase, genetic expression search is periodically guided by language-model operations conditioned on current concepts. In the second phase, Pareto-optimal and low-performing expressions are summarized by the language model into new qualitative textual concepts. In the third phase, these concepts are further evolved and generalized. The authors evaluated the system on 100 physics equations from the Feynman Lectures under noisy conditions, a benchmark of 41 synthetic equations designed to prevent data memorization, and a real-world machine learning task using 53,812 evaluations from the BIG-Bench benchmark.

The findings show that LASR achieves an exact match solve rate of 72% on the Feynman benchmark, outperforming the previous leading baseline of 59% and deep learning alternatives ranging between 20% and 40%. Even with a small open-source model and minimal language guidance, the method solved 67% of the equations. On the synthetic benchmark, LASR achieved a high predictive accuracy of 0.913 compared to 0.070 for the standard genetic baseline, confirming that performance gains stem from concept-guided reasoning rather than data memorization. Furthermore, when applied to language model scaling behaviors, LASR successfully discovered an empirical scaling law that matches established formulations like Chinchilla while requiring only three free parameters instead of five.

These results indicate that concept libraries introduce meaningful semantic biases that help algorithms escape local minima and filter out unnecessary mathematical terms, producing cleaner, more generalizable formulas. This allows researchers to discover complex empirical relationships without manually specifying rigid functional forms beforehand, reducing discovery timelines and engineering overhead.

Organizations exploring automated scientific discovery and empirical modeling should consider augmenting genetic search workflows with language-guided concept extraction. For practitioners adopting this method, incorporating high-level human hints can further accelerate convergence. Future development should focus on testing the approach with larger compute budgets, exploring model fine-tuning, and adapting the concept induction mechanism to broader scientific optimization tasks.

Confidence in the system's search capabilities is high across the evaluated benchmarks, but practical users must exercise caution. The system does not guarantee that induced concepts are scientifically accurate or mutually consistent, meaning qualitative concepts should be treated as heuristic search drivers rather than proven domain principles until formally verified.

arXiv: 2409.09359
Cover for Symbolic Regression with a Learned Concept Library

Abstract

We present a novel method for symbolic regression (SR), the task of searching for compact programmatic hypotheses that best explain a dataset. The problem is commonly solved using genetic algorithms; we show that we can enhance such methods by inducing a library of abstract textual concepts. Our algorithm, called LASR, uses zero-shot queries to a large language model (LLM) to discover and evolve abstract concepts occurring in known high-performing hypotheses. We discover new hypotheses using a mix of standard evolutionary steps and LLM-guided steps (obtained through zero-shot LLM queries) conditioned on discovered concepts. Once discovered, hypotheses are used in a new round of concept abstraction and evolution. We validate LASR on the Feynman equations, a popular SR benchmark, as well as a set of synthetic tasks. On these benchmarks, LASR substantially outperforms a variety of state-of-the-art SR approaches based on deep learning and evolutionary algorithms. Moreover, we show that LASR can be used to discover a new and powerful scaling law for LLMs.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 3 Method
  • 4 Experiments
  • 4.1 Comparison against baselines in the Feynman Equation Dataset
  • 4.2 Cascading Experiments
  • 4.3 Ablation Experiments
  • 4.4 Qualitative Analysis and User Hints
  • 4.5 Data Leakage Validation
  • 4.6 Using LASR to discover LLM Scaling Laws
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Broader Societal Impacts
  • A.2 LLMPrompts
  • A.3 Implementation Details
  • A.3.1 Compute Usage
  • A.3.2 Concept Sampling
  • A.3.3 Hyperparameters
  • A.4 Dataset Details
  • A.4.1 Feynman Equations
  • A.4.2 Synthetic Dataset
  • A.5 Additional Experiments
  • A.5.1 Asymmetric comparison with PySR
  • A.5.2 Subset of equations discovered by LaSR on the Feynman Dataset
  • A.5.3 Qualitative comparison of Synthetic Dataset equations
  • A.5.4 Stochasticity of LASR and PySR
  • A.6 Using LASR to find an LLM Scaling Law
  • A.7 Metrics for Cascading experiment
  • A.8 Further Qualitative Analysis
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Hierarchical Bayesian Formulation of Symbolic Regression with Latent Concepts

    model/method

    Classical symbolic regression formulates the search for a mathematical expression π∈L\pi \in \mathcal{L} that explains an empirical dataset D={(xi,yi)}i=1N\mathcal{D} = \{(x_i, y_i)\}_{i=1}^N as a maximum a posteriori (MAP) estimation problem:

    π⋆=arg⁡max⁡πpL(π∣D)=arg⁡max⁡πpL(D∣π) pL(π)\pi^\star = \arg\max_{\pi} p_{\mathcal{L}}(\pi | \mathcal{D}) = \arg\max_{\pi} p_{\mathcal{L}}(\mathcal{D} | \pi) \, p_{\mathcal{L}}(\pi)

    where pL(D∣π)p_{\mathcal{L}}(\mathcal{D} | \pi) is the data execution likelihood (fitness) and pL(π)p_{\mathcal{L}}(\pi) is a prior distribution penalizing syntactic expression complexity.

    To integrate natural-language empirical intuitions into the search, symbolic regression with a latent concept library introduces an open-ended collection of textual concepts C\mathcal{C}. The joint distribution over programmatic hypotheses π\pi and concept library C\mathcal{C} given dataset D\mathcal{D} is formulated as the hierarchical optimization problem:

    arg⁡max⁡π,Cp(π,C∣D)=arg⁡max⁡π,Cp(D∣π)⋅p(π∣C)⋅p(C)\arg\max_{\pi, \mathcal{C}} p(\pi, \mathcal{C} | \mathcal{D}) = \arg\max_{\pi, \mathcal{C}} p(\mathcal{D} | \pi) \cdot p(\pi | \mathcal{C}) \cdot p(\mathcal{C})

    where:

    • p(D∣π)p(\mathcal{D} | \pi) is the exact likelihood obtained via program execution on dataset D\mathcal{D}.
    • p(π∣C)p(\pi | \mathcal{C}) is the conditional likelihood of candidate expressions conforming to concept set C\mathcal{C}, approximated by querying a Large Language Model (LLM) zero-shot with the concepts and task grammar.
    • p(C)p(\mathcal{C}) is the prior distribution over natural-language scientific concepts, also approximated via LLM generative sampling.
  2. Knowl 2 — The LASR Algorithm for Concept-Guided Symbolic Regression

    algorithm

    LASR (Language-Augmented Symbolic Regression) is an alternating optimization algorithm that simultaneously evolves a pool of mathematical hypotheses across parallel populations and a library of natural-language concepts.

    Input: Optional user hints C0C_0, dataset D={(xi,yi)}i=1ND = \{(x_i, y_i)\}_{i=1}^N, iterations II, populations KK, concept evolution steps MM, LLM operator probability pp
    Output: Evolved concept library CC, best programmatic hypothesis π⋆\pi^\star
    C = InitializeConceptLibrary(C_0)
    {Pi_1, ..., Pi_K} = InitializePopulations(C, K)
    for iter = 1 to I do
        for i = 1 to K do
            Pi_i = SRCycle(Pi_i, D, C, p)
        end for
        F = ExtractParetoFrontier({Pi_1, ..., Pi_K}, D)
        C = C union ConceptAbstraction(F, C)
        for m = 1 to M do
            C = ConceptEvolution(C)
        end for
    end for
    pi_star = BestExpression(F)
    return C, pi_star

    Execution Stages:

    1. Initialization (InitializePopulations): Initializes KK populations of expression trees, optionally using LLM-generated expressions seeded from user hints C0C_0 or LLM concept priors.
    2. Hypothesis Search (SRCycle): Evolves populations in parallel over multiple generations. In each step, symbolic mutation and crossover operators are substituted with probability pp by LLM-guided operators (LLMMutate, LLMCrossover) conditioned on concepts sampled from CC. Expressions are simplified, constants are optimized via continuous optimization, and elite expressions migrate across populations.
    3. Concept Abstraction (ConceptAbstraction): Extracts the Pareto frontier F\mathcal{F} of expressions (trading off accuracy against syntactic complexity) alongside the lowest-performing expressions. A zero-shot LLM prompt synthesizes natural-language concepts capturing positive structural motifs and penalizing negative motifs.
    4. Concept Evolution (ConceptEvolution): Recombines existing concepts over MM iterations via zero-shot LLM prompts to synthesize more abstract and general domain concepts.
  3. Knowl 3 — Concept-Conditioned LLM Genetic Operators

    model/method

    LASR injects semantic language priors into evolutionary search by defining three zero-shot LLM-augmented genetic operators that replace standard genetic programming (GP) operations with probability pp (and execute standard symbolic operations with probability 1−p1-p):

    • LLMInit: Generates initial mathematical expression trees by prompting an LLM with available variable names, allowed arithmetic/trigonometric operators, and concepts sampled from the concept library C\mathcal{C} (or initial user hints C0\mathcal{C}_0).
    • LLMMutate: Samples a subset of ll concepts from C\mathcal{C} and prompts the LLM to mutate a given expression πi\pi_i into an updated expression πj\pi_j that implements the sampled conceptual patterns while preserving valid grammar and variable constraints.
    • LLMCrossover: Samples ll concepts from C\mathcal{C} alongside two parent expressions πi\pi_i and πj\pi_j, prompting the LLM to construct an offspring expression πk\pi_k that recombines sub-expression trees from both parents in accordance with the specified concepts.

    Stochastically interleaving these LLM operators at probability pp enables language-guided macro-steps in program space without bottlenecking fine-grained local search.

  4. Knowl 4 — Concept Abstraction and Evolution via LLM In-Context Reasoning

    model/method

    LASR expands and refines its latent concept library C\mathcal{C} through two zero-shot LLM prompting mechanisms:

    • Concept Abstraction: Collects the Pareto-optimal expressions across all populations (optimizing MSE loss against syntactic complexity) alongside the highest-loss expressions from the current iteration: F={π1+,…,πa+,π1−,…,πb−}\mathcal{F} = \{\pi^+_1, \dots, \pi^+_a, \pi^-_1, \dots, \pi^-_b\}. A zero-shot prompt asks the LLM to hypothesize underlying structural, algebraic, and physical assumptions that differentiate the high-performing expressions from the poor ones. The synthesized textual descriptions are appended to C\mathcal{C}.
    • Concept Evolution: Samples multiple concepts from C\mathcal{C} and prompts the LLM to merge and extrapolate these ideas into novel, diverse scientific hypotheses (e.g., connecting exponential decay to temperature-dependent dynamical equilibria). Because concept fitness cannot be evaluated prior to program execution, all evolved concepts are preserved in C\mathcal{C} to maximize exploration diversity.
  5. Knowl 5 — Recency-Based Concept Sampling Strategy for Evolution and Diversity

    model/method

    To balance exploitation of high-quality concepts with exploration of novel conceptual spaces, LASR applies opposing sampling filters on the concept library C\mathcal{C} depending on the target operation:

    • Hypothesis Guidance Sampling: When sampling concepts for LLMMutate and LLMCrossover, LASR uniformly samples from the top-KK most recently generated concepts in C\mathcal{C} (with K=20K = 20). This conditions program search on the most up-to-date and refined conceptual abstractions.
    • Concept Evolution Sampling: When sampling concepts for ConceptEvolution, the top-KK most recent concepts are explicitly excluded. The LLM is prompted with older concepts from the library, preventing the concept generation process from collapsing into a narrow cluster of recent ideas and preserving diverse conceptual lineages.
  6. Knowl 6 — Exact Solve Rates on the Noisy Feynman Physics Benchmark

    data/table

    On the 100 equations of the Feynman Lectures on Physics dataset (evaluated with target Gaussian noise of 0.0010.001 on the dependent variable and additional random noise variables to test feature selection), LASR outperforms deep learning and genetic programming baselines under identical iteration budgets (I=40I=40).

    Method Exact Match Solve Rate
    GPlearn 20 / 100
    AFP 24 / 100
    AFP-FE 26 / 100
    DSR 23 / 100
    uDSR 40 / 100
    AIFeynman 38 / 100
    PySR 59 / 100
    LASR (GPT-3.5-turbo, p=1%p=1\%) 72 / 100

    PySR represents the direct baseline equivalent to LASR without LLM operators and concept induction. When PySR is executed uninterrupted for 10 hours per equation up to 10610^6 iterations, it reaches an exact solve rate of 62/10062/100 (59+359 + 3), remaining well below LASR's 40-iteration rate of 72/10072/100.

  7. Knowl 7 — Performance of LASR Under LLM Backbone and Guidance Probability Cascades

    data/table

    A model cascade across LLM backbones (Llama-3-8B and GPT-3.5-turbo) and concept guidance probabilities p∈{1%,5%,10%}p \in \{1\%, 5\%, 10\%\} on the 100 Feynman equations demonstrates monotonic improvements in equation discovery as LLM capability and query frequency increase.

    LASR (Llama3-8B) LASR (GPT-3.5)
    Solve Category PySR p=1%p = 1\% p=5%p = 5\% p=10%p = 10\% p=1%p = 1\%
    Exact Solve 59/100 67/100 69/100 71/100 72/100
    Almost Solve 7/100 5/100 6/100 2/100 3/100
    Close 16/100 9/100 12/100 12/100 10/100
    Not Close 18/100 19/100 13/100 16/100 15/100

    Definitions of qualitative solve tiers:

    • Exact Solve: Discovered expression symbolically simplifies to the ground truth equation.
    • Almost Solve: Low MSE loss, but the expression contains one extraneous term or lacks one term.
    • Close: Captures the general nested functional structure of the solution without symbolic equivalence.
    • Not Close: Fails to capture the structure of the ground truth equation.
  8. Knowl 8 — Ablation Analysis of Concept Library, Variable Names, and User Hints

    empirical result

    Ablation experiments on the Feynman physics dataset measuring convergence under an MSE threshold of <10−11< 10^{-11} over 40 iterations demonstrate the role of each component of LASR:

    1. Semantic Variable Names: Stripping domain variable names (e.g., replacing θ,r,ϵ\theta, r, \epsilon with generic symbols) leads to a substantial performance drop. LLM operators rely on variable semantics to infer relevant mathematical operations (e.g., associating angular variable θ\theta with trigonometric functions).
    2. Concept Library and Concept Evolution: Removing concept library abstraction and evolution slows down convergence, widening the performance gap against standard genetic search as concept guidance probability pp increases. Ablating concept evolution while retaining concept abstraction yields intermediate performance.
    3. User Hints: Seeding the initial concept library C0\mathcal{C}_0 with brief chapter titles from the Feynman lectures significantly accelerates equation discovery speed and final solve count.
  9. Knowl 9 — Validation Against Memorization and Data Leakage Using Synthetic Equations

    empirical result

    To confirm that LASR's performance does not stem from pretraining memorization or dataset leakage:

    1. Non-Physical Synthetic Benchmark: A benchmark of 41 procedurally generated equations was created with arbitrary variable names and non-standard mathematical compositions (e.g., y=0.782x3+0.536x2ex1(log⁡x2−x2ecos⁡x1)y = \frac{0.782x_3 + 0.536}{x_2 e^{x_1}(\log x_2 - x_2 e^{\cos x_1})}) designed such that PySR achieves poor fit (MSE >1> 1 after 400 iterations). On this dataset, LASR (using Llama3-8B at p=0.1%p=0.1\%) achieves a test set R2R^2 of 0.9130.913 after 20 iterations, compared to 0.0700.070 for PySR.
    2. Stochastic Variation Across Random Seeds: Evaluating LASR across different random seeds on Equation 53 of the Feynman dataset (point source heat flow, h=P4πR2h = \frac{P}{4\pi R^2}) yields syntactically distinct discovered expressions that each simplify to the true inverse-square relation:
    h=2.88411330.000219−r(−36.240696−(sin⁡(r)⋅0.002057)P)r(loss 2×10−6,complexity 16)h = \frac{2.8841133}{0.000219 - r\left(\frac{-36.240696 - (\sin(r)\cdot 0.002057)}{P}\right)r} \quad (\text{loss } 2\times 10^{-6}, \text{complexity } 16)

    and

    h=Pr(r⋅17.252695−0.001495)⋅1.372847(loss 2×10−6,complexity 11)h = \frac{P}{r(r\cdot 17.252695 - 0.001495)} \cdot 1.372847 \quad (\text{loss } 2\times 10^{-6}, \text{complexity } 11)

    This structural divergence confirms that LASR conducts stochastic search guided by abstracted concepts rather than reproducing memorized token sequences.

  10. Knowl 10 — In-Context Scaling Law for Large Language Models Discovered by LASR

    empirical result

    When applied without an initial equation template to 53,812 evaluation points from BIG-Bench multiple-choice grade tasks (across 55 language models), LASR discovered the following empirical scaling law connecting training steps and in-context shot count:

    score=A(train_stepsB)#shots+E\text{score} = \frac{A}{\left(\frac{\text{train\_steps}}{B}\right)^{\text{\#shots}}} + E

    with fitted numerical parameters:

    score=−0.0248235(train_steps116050.999)#shots+0.360124\text{score} = \frac{-0.0248235}{\left(\frac{\text{train\_steps}}{116050.999}\right)^{\text{\#shots}}} + 0.360124

    where train_steps\text{train\_steps} is the model training step count, #shots\text{\#shots} is the number of in-context inference examples, and E≈0.360E \approx 0.360 is the baseline score.

    Scaling Law Skeleton Validation MSE Loss Free Parameters
    LASR Discovered Equation 0.03598±0.002650.03598 \pm 0.00265 3
    Chinchilla 0.03610±0.002680.03610 \pm 0.00268 5
    Modified Chinchilla 0.03598±0.002640.03598 \pm 0.00264 5
    Residual Term Only (score=E\text{score} = E) 0.09324±0.019920.09324 \pm 0.01992 1

    The discovered 3-parameter formula matches the predictive accuracy of the 5-parameter Chinchilla formulation on a held-out set of 10,763 validation samples, capturing the phenomenon that in-context prompting yields exponential gains for low-compute models while exhibiting diminishing returns as pretraining compute grows.

  11. Knowl 11 — Limitations of Concept-Guided Symbolic Regression

    limitation

    LASR has three primary limitations:

    1. Unverified Concept Correctness: Textual concepts induced by the LLM are derived from empirical correlations and model heuristics; they can capture spurious dataset artifacts or model biases, potentially misleading human domain experts.
    2. Lack of Mutual Consistency Checking: The algorithm lacks a formal verification mechanism to enforce that newly evolved concepts are logically or mathematically consistent with existing concepts in the library C\mathcal{C}.
    3. Compute and Token Overhead: The number of LLM invocations scales with iterations and guidance probability (reaching up to 60,000 calls and ≈25M\approx 25\text{M} tokens for 40 iterations at p=1%p=1\%), making Candidate generation substantially slower than purely symbolic operations.

Coverage note — None was omitted; all primary contributions, including the hierarchical Bayesian formulation, the LASR algorithm, LLM genetic operators, concept sampling heuristics, benchmark experiments, synthetic data validation, discovered scaling law, and limitations, are fully covered.

References

  1. 1.Ibrahim Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. Revisiting neural scaling laws in language and vision. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022.
  2. 2.Rohit Batra, Le Song, and Rampi Ramprasad. Emerging materials intelligence ecosystems propelled by machine learning. Nature Reviews Materials, 6(8):655–678, 2021.
  3. 3.Matthew Bowers, Theo X Olausson, Lionel Wong, Gabriel Grand, Joshua B Tenenbaum, Kevin Ellis, and Armando Solar-Lezama. Top-down synthesis for library learning. Proceedings of the ACM on Programming Languages, 7(POPL):1182–1213, 2023.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  5. 5.Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger. Broken neural scaling laws. In The Eleventh International Conference on Learning Representations, 2023.
  6. 6.Swarat Chaudhuri, Kevin Ellis, Oleksandr Polozov, Rishabh Singh, Armando Solar-Lezama, Yisong Yue, et al. Neurosymbolic programming. Foundations and Trends® in Programming Languages, 7(3):158–243, 2021.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  8. 8.Xinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton, Hanjun Dai, Max Lin, and Denny Zhou. Spreadsheetcoder: Formula prediction from semi-structured context. In International Conference on Machine Learning, pages 1661–1672. PMLR, 2021.
  9. 9.Mia Chiquier, Utkarsh Mall, and Carl Vondrick. Evolving interpretable visual classifiers with large language models, 2024.
  10. 10.Miles Cranmer. Interpretable machine learning for science with pysr and symbolicregression. jl. arXiv preprint arXiv:2305.01582, 2023.
  11. 11.Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho. Discovering symbolic models from deep learning with inductive biases. In Neural Information Processing Systems, 2020.
  12. 12.Benjamin L Davis and Zehao Jin. Discovery of a planar black hole mass scaling relation for spiral galaxies. The Astrophysical Journal Letters, 956(1):L22, 2023.
  13. 13.Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, and Pushmeet Kohli. Robustfill: Neural program learning under noisy I/O. In ICML, 2017.
  14. 14.Kevin Ellis, Lucas Morales, Mathias Sablé-Meyer, Armando Solar-Lezama, and Josh Tenenbaum. Learning libraries of subroutines for neurally–guided Bayesian program induction. In Advances in Neural Information Processing Systems, pages 7805–7815, 2018.
  15. 15.Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sable-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, and Joshua B Tenenbaum. Dreamcoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning. arXiv preprint arXiv:2006.08381, 2020.
  16. 16.Aarohi Srivastava et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023.
  17. 17.Abhimanyu Dubey et al. The llama 3 herd of models, 2024.
  18. 18.Donald Gerwin. Information processing, data inferences, and scientific generalization. Behavioral Science, 19(5):314–325, 1974.
  19. 19.Gabriel Grand, Lionel Wong, Matthew Bowers, Theo X Olausson, Muxin Liu, Joshua B Tenenbaum, and Jacob Andreas. Lilo: Learning interpretable libraries by compressing and documenting code. arXiv preprint arXiv:2310.19791, 2023.
  20. 20.Arthur Grundner, Tom Beucler, Pierre Gentine, and Veronika Eyring. Data-driven equation discovery of a cloud cover parameterization. Journal of Advances in Modeling Earth Systems, 16(3):e2023MS003763, 2024.
  21. 21.Tanmay Gupta and Aniruddha Kembhavi. Visual programming: Compositional visual reasoning without training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962, 2023.
  22. 22.Alberto Hernandez, Adarsh Balasubramanian, Fenglin Yuan, Simon AM Mason, and Tim Mueller. Fast, accurate, and transferable many-body interatomic potentials by symbolic regression. npj Computational Materials, 5(1):112, 2019.
  23. 23.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  24. 24.John E. Hopcroft, Rajeev Motwani, and Jeffrey D. Ullman. Introduction to automata theory, languages, and computation, 3rd Edition. Pearson international edition. Addison-Wesley, 2007.
  25. 25.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  26. 26.William La Cava, Patryk Orzechowski, Bogdan Burlacu, Fabrício Olivetti de França, Marco Virgolin, Ying Jin, Michael Kommenda, and Jason H Moore. Contemporary symbolic regression methods and their relative performance. arXiv preprint arXiv:2107.14351, 2021.
  27. 27.Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. Human-level concept learning through probabilistic program induction. Science, 350(6266):1332–1338, 2015.
  28. 28.Mikel Landajuela, Chak Shing Lee, Jiachen Yang, Ruben Glatt, Claudio P Santiago, Ignacio Aravena, Terrell Mundhenk, Garrett Mulcahy, and Brenden K Petersen. A unified framework for deep symbolic regression. Advances in Neural Information Processing Systems, 35:33985–33998, 2022.
  29. 29.Pat Langley. Bacon: A production system that discovers empirical laws. In International Joint Conference on Artificial Intelligence, 1977.
  30. 30.Pablo Lemos, Niall Jeffrey, Miles Cranmer, Shirley Ho, and Peter Battaglia. Rediscovering orbital mechanics with machine learning. Machine Learning: Science and Technology, 4(4):045002, 2023.
  31. 31.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023.
  32. 32.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  33. 33.Nour Makke and Sanjay Chawla. Interpretable scientific discovery with symbolic regression: a review. Artificial Intelligence Review, 57(1):2, 2024.
  34. 34.Matteo Merler, Nicola Dainese, and Katsiaryna Haitsiukevich. In-context symbolic regression: Leveraging language models for function discovery. arXiv preprint arXiv:2404.19094, 2024.
  35. 35.Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K. Hoover, and Joel Lehman. Language model crossover: Variation through few-shot prompting, 2024.
  36. 36.Vijayaraghavan Murali, Letao Qi, Swarat Chaudhuri, and Chris Jermaine. Neural sketch learning for conditional program generation. ICLR, 2018.
  37. 37.Deep Symbolic Optimization Organization. Srbench symbolic solution. https://github.com/dso-org/deep-symbolic-optimization/blob/master/images/srbench_symbolic-solution.png. Accessed: 2024-05-22.
  38. 38.Brenden K Petersen, Mikel Landajuela, T Nathan Mundhenk, Claudio P Santiago, Soo K Kim, and Joanne T Kim. Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients. arXiv preprint arXiv:1912.04871, 2019.
  39. 39.Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. Regularized evolution for image classifier architecture search. In Proceedings of the aaai conference on artificial intelligence, volume 33, pages 4780–4789, 2019.
  40. 40.Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024.
  41. 41.Michael Schmidt and Hod Lipson. Distilling free-form natural laws from experimental data. Science, 324(5923):81–85, 2009.
  42. 42.Michael D. Schmidt and Hod Lipson. Age-fitness pareto optimization. In Annual Conference on Genetic and Evolutionary Computation, 2010.
  43. 43.Ameesh Shah, Eric Zhan, Jennifer J Sun, Abhinav Verma, Yisong Yue, and Swarat Chaudhuri. Learning differentiable programs with admissible neural heuristics. In Advances in Neural Information Processing Systems, 2020.
  44. 44.Richard Shin, Miltiadis Allamanis, Marc Brockschmidt, and Oleksandr Polozov. Program synthesis and semantic parsing with learned code idioms. In Advances in Neural Information Processing Systems, pages 10825–10835, 2019.
  45. 45.Parshin Shojaee, Kazem Meidani, Shashank Gupta, Amir Barati Farimani, and Chandan K Reddy. Llm-sr: Scientific equation discovery via programming with large language models. arXiv preprint arXiv:2404.18400, 2024.
  46. 46.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.
  47. 47.Trevor Stephens. gplearn: Genetic programming in python, with a scikit-learn inspired api, 2024. Accessed: 2024-05-22.
  48. 48.Dídac Surís, Sachit Menon, and Carl Vondrick. Vipergpt: Visual inference via python execution for reasoning. arXiv preprint arXiv:2303.08128, 2023.
  49. 49.Silviu-Marian Udrescu and Max Tegmark. Ai feynman: A physics-inspired method for symbolic regression. Science Advances, 6(16):eaay2631, 2020.
  50. 50.Sergiy Verstyuk and Michael R Douglas. Machine learning the gravity equation for international trade. Available at SSRN 4053795, 2022.
  51. 51.Marco Virgolin, Ziyuan Wang, Tanja Alderliesten, and Peter AN Bosman. Machine learning for the prediction of pseudorealistic pediatric abdominal phantoms for radiation dose reconstruction. Journal of Medical Imaging, 7(4):046501–046501, 2020.
  52. 52.Eugene P Wigner. The unreasonable effectiveness of mathematics in the natural sciences. In Mathematics and science, pages 291–306. World Scientific, 1990.
  53. 53.Catherine Wong, Kevin M Ellis, Joshua Tenenbaum, and Jacob Andreas. Leveraging language to learn program abstractions and search heuristics. In International conference on machine learning, pages 11193–11204. PMLR, 2021.
  54. 54.Eric Zelikman, Qian Huang, Gabriel Poesia, Noah Goodman, and Nick Haber. Parsel: Algorithmic reasoning with language models by composing decompositions. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Citation

MLA
Grayeli, A., et al. “Symbolic Regression with a Learned Concept Library”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 44678–709, https://proceedings.neurips.cc/paper_files/paper/2024/file/4ec3ddc465c6d650c9c419fb91f1c00a-Paper-Conference.pdf.
APA
Grayeli, A., Sehgal, A., Costilla-Reyes, O., Cranmer, M., & Chaudhuri, S. (2024). Symbolic Regression with a Learned Concept Library. Advances in Neural Information Processing Systems, 37, 44678–44709. https://proceedings.neurips.cc/paper_files/paper/2024/file/4ec3ddc465c6d650c9c419fb91f1c00a-Paper-Conference.pdf
Chicago
Grayeli, A., A. Sehgal, O. Costilla-Reyes, M. Cranmer, and S. Chaudhuri. 2024. “Symbolic Regression with a Learned Concept Library”. Advances in Neural Information Processing Systems 37: 44678–709. https://proceedings.neurips.cc/paper_files/paper/2024/file/4ec3ddc465c6d650c9c419fb91f1c00a-Paper-Conference.pdf.
Harvard
Grayeli, A. et al. (2024) “Symbolic Regression with a Learned Concept Library”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 44678–44709. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/4ec3ddc465c6d650c9c419fb91f1c00a-Paper-Conference.pdf.
Vancouver
1. Grayeli A, Sehgal A, Costilla-Reyes O, Cranmer M, Chaudhuri S (2024) Symbolic Regression with a Learned Concept Library. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 44678–44709

BibTeX

@inproceedings{grayeli2024symbolic,
  title = {Symbolic Regression with a Learned Concept Library},
  author = {Grayeli, Arya and Sehgal, Atharva and Costilla-Reyes, Omar and Cranmer, Miles and Chaudhuri, Swarat},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {44678-44709},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/4ec3ddc465c6d650c9c419fb91f1c00a-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors