Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data
Anuj KarpatneGowtham AtluriJames FaghmousMichael SteinbachArindam BanerjeeAuroop GangulyShashi ShekharNagiza SamatovaVipin Kumar
Establishes theory-guided data science as a foundational paradigm that integrates physical principles into machine learning to produce scientifically consistent, interpretable, and generalizable models across disciplines.
Modern data science has driven remarkable breakthroughs in commercial settings, but purely data-driven models frequently stumble when applied to complex physical and scientific problems. In critical scientific applications, standard machine learning methods often fail to generalize because training data is sparse, noisy, and non-stationary, while the underlying physical processes involve numerous interacting variables. Black-box models can easily identify deceptive, non-causal patterns that perform well on benchmark tests but fail completely in real-world deployment, as demonstrated when Google Flu Trends overestimated flu visits by more than a factor of two. Conversely, traditional scientific models based purely on first principles often rely on oversimplified assumptions when dealing with complex dynamical systems. The article addresses this fundamental divide by establishing the conceptual framework of Theory-guided Data Science (TGDS), a paradigm designed to systematically merge established scientific knowledge with advanced data science techniques to produce physically consistent, generalizable, and interpretable models.
To demonstrate this paradigm, the article conducts an extensive conceptual evaluation and literature synthesis across multiple domains, including climate science, hydrology, material discovery, quantum chemistry, and aerospace engineering. The framework introduces physical consistency as a core performance metric alongside training accuracy and model simplicity, showing that domain knowledge can eliminate physically impossible solutions and reduce model variance without sacrificing accuracy. The article structures TGDS into five core approaches: guiding model design through tailored response functions and physically informed neural network architectures; guiding learning algorithms using informed parameter initializations, probabilistic priors, domain-driven constraints, and structured regularization; refining model outputs using explicit physical equations or inferred implicit orderings; constructing hybrid systems where machine learning resolves deficiencies in theoretical equations; and augmenting theory-based systems via observational data assimilation and automated parameter calibration.
The findings across various scientific case studies confirm that blending domain principles with data-driven methods delivers substantial performance gains. In materials science, integrating probabilistic models with theoretical calculations accelerated compound discovery, identifying approximately one hundred previously unknown ternary oxide compounds. In Earth observation, incorporating implicit elevation-based physical constraints enabled the accurate tracking and global mapping of dynamic surface water bodies despite noisy satellite data. In aerospace engineering, embedding machine learning corrections directly into Reynolds-averaged Navier-Stokes equations significantly closed performance gaps in complex turbulence simulations. Furthermore, in macroecology, initializing matrix completion models with species-level biological averages improved plant trait prediction accuracy significantly compared to standard, randomly initialized machine learning algorithms.
These results demonstrate that theory-guided data science mitigates the severe risks and costs of deploying unconstrained black-box models in high-stakes fields such as healthcare, aerospace, and environmental management. Anchoring models in established physical laws ensures that algorithms reflect causal mechanisms rather than spurious mathematical correlations. Consequently, organizations can safely extract value from large observational datasets without sacrificing interpretability, regulatory compliance, or operational safety. Decision-makers should prioritize TGDS methodologies over purely black-box machine learning when addressing complex physical systems, and research teams should develop hybrid workflows that explicitly embed physical laws into their model training objectives and evaluation pipelines. Future research must expand these methodologies beyond supervised regression to unsupervised pattern mining, automated workflow design, and comprehensive uncertainty quantification, while establishing standardized techniques to handle scientific knowledge that is uncertain, incomplete, or represented across differing physical scales.
No sufficiently relevant recommendations were found.
- Paper: DeepXDE: A Deep Learning Library for Solving Differential Equations, Lu Lu et al. (2019). DeepXDE puts theory-guided data science into software practice by embedding governing differential equations and physical laws directly into deep learning loss functions.
- Paper: When and why PINNs fail to train: A neural tangent kernel perspective, Sifan Wang et al. (2020). This paper investigates the theoretical failure modes and optimization dynamics of physics-informed neural networks, advancing the core TGDS methodology of embedding physical constraints into machine learning.
- Paper: Learning to Simulate Complex Physics with Graph Networks, Alvaro Sanchez-Gonzalez et al. (2020). This work demonstrates how relational inductive biases and graph networks can learn to simulate complex physical dynamics directly from data while maintaining physical consistency.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). It formalizes relational inductive biases in deep neural networks and graph architectures, providing a key mechanism for incorporating domain structure into machine learning models.
- Paper: Analyzing Learned Molecular Representations for Property Prediction, Kevin Yang et al. (2019). This study applies domain-guided graph neural representations and physical descriptors to molecular property prediction in chemical science.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). Kolmogorov-Arnold Networks offer an alternative interpretable architecture tailored for discovering exact physical equations and solving scientific machine learning tasks.
- Paper: Towards A Rigorous Science of Interpretable Machine Learning, Finale Doshi-Velez et al. (2017). It provides a rigorous framework for evaluating interpretability in machine learning, addressing the primary TGDS requirement of producing scientifically meaningful models.
- Paper: Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models, Kevin Murphy (2026). This work operationalizes automated scientific discovery by combining Bayesian experiment design with language models to uncover mechanistic physical and chemical laws.
