AI Feynman: A physics-inspired method for symbolic regression
Silviu-Marian UdrescuMax Tegmark
Develops AI Feynman, a recursive symbolic regression algorithm that combines neural network fitting with physical properties such as symmetry and separability to extract governing equations from data, raising discovery rates on difficult benchmarks from 15% to 90%.
Modern data-driven science faces a fundamental bottleneck in symbolic regression: automatically discovering exact, human-interpretable mathematical formulas from numerical data. Standard brute-force search over mathematical symbols is computationally impossible for complex expressions because search spaces grow exponentially. Meanwhile, popular genetic algorithm tools often get trapped in suboptimal approximations when formulas contain nested or non-additive structures.
The article develops and evaluates AI Feynman, a recursive symbolic regression algorithm that combines standard fitting techniques with neural networks to discover underlying physical properties—such as symmetries and separability—to break complex multi-variable problems into simpler, solvable sub-problems.
To evaluate the algorithm, the authors constructed a benchmark database of 100 physics equations from the Feynman Lectures on Physics and a separate, more challenging test set of 20 advanced physics formulas from standard graduate-level texts. Synthetic datasets with up to 100,000 points per equation were generated and evaluated under varying levels of data volume and injected noise. AI Feynman systematically applies dimensional analysis, polynomial fitting, and brute-force searches alongside neural network interpolators that test for translational symmetry, scaling, and additive or multiplicative separability.
The analysis produced several key findings. First, AI Feynman discovered 100% (100 of 100) of the basic Feynman benchmark equations, significantly outperforming the commercial genetic algorithm software Eureqa, which solved 71%. Second, on the 20 advanced test equations, AI Feynman improved the state-of-the-art discovery rate from 15% (3 of 20 solved by Eureqa) to 90% (18 of 20). Third, even when physical units were omitted and dimensional analysis was disabled, AI Feynman maintained a 93% success rate on the benchmark by relying directly on neural network property detection. Finally, the algorithm exhibited strong data efficiency and noise resilience: most simpler equations required as few as 10 to 100 data points and tolerated up to 1% injected noise, whereas complex multi-variable equations required larger sample sizes (up to 1,000,000 points) and lower noise thresholds to train the neural network effectively.
These results indicate that automated scientific discovery does not require intractable brute-force computation if algorithms exploit structural simplifications inherent in natural laws. By reducing multi-variable dependencies into modular sub-tasks, AI Feynman provides a deterministic path toward exact analytic solutions, reducing the risk of false convergence common in genetic programming. This capability accelerates scientific modeling and allows automated distillation of governing physical laws directly from experimental data.
For future development, the article recommends hybridizing AI Feynman with genetic algorithms to generate Pareto-optimal candidate expressions under noisy conditions. It also proposes extending the architecture to automatically discover differential equations by estimating derivatives, incorporating broader functional operators such as arbitrary exponentiation, and adopting optimized neural network architectures to lower fitting error.
The findings are supported with high confidence on clean, well-sampled synthetic data with known physical properties. However, readers should exercise caution when evaluating highly arbitrary mathematical formulas lacking physical symmetries or applications involving noisy experimental data with noise in the independent variables, where neural network training and property thresholding become more challenging.
- Paper: Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data, Anuj Karpatne et al. (2016). It provides the foundational conceptual framework for theory-guided data science and embedding physical properties into machine learning, which directly motivates the physics-inspired regression paradigm of AI Feynman.
- Paper: Gene Expression Programming: A New Adaptive Algorithm for Solving Problems, Cândida Ferreira (2001). It introduces foundational evolutionary and expression-tree representations for symbolic regression, establishing the traditional baseline methods that AI Feynman seeks to dramatically improve upon.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). It extends AI Feynman's goal of interpretable, physics-inspired scientific discovery by introducing Kolmogorov-Arnold Networks as an interpretable alternative to MLPs for discovering symbolic and mathematical formulas.
- Paper: Scientific Machine Learning Through Physics–Informed Neural Networks: Where we are and What’s Next, Salvatore Cuomo et al. (2022). It surveys the subsequent landscape of scientific machine learning and physics-informed models, providing broader context on how physics-guided neural architectures have evolved for scientific discovery.
