Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data

Anuj KarpatneGowtham AtluriJames FaghmousMichael SteinbachArindam BanerjeeAuroop GangulyShashi ShekharNagiza SamatovaVipin Kumar

article2016TKDE1,419 citations

Establishes theory-guided data science as a foundational paradigm that integrates physical principles into machine learning to produce scientifically consistent, interpretable, and generalizable models across disciplines.

Listen

Modern data science has driven remarkable breakthroughs in commercial settings, but purely data-driven models frequently stumble when applied to complex physical and scientific problems. In critical scientific applications, standard machine learning methods often fail to generalize because training data is sparse, noisy, and non-stationary, while the underlying physical processes involve numerous interacting variables. Black-box models can easily identify deceptive, non-causal patterns that perform well on benchmark tests but fail completely in real-world deployment, as demonstrated when Google Flu Trends overestimated flu visits by more than a factor of two. Conversely, traditional scientific models based purely on first principles often rely on oversimplified assumptions when dealing with complex dynamical systems. The article addresses this fundamental divide by establishing the conceptual framework of Theory-guided Data Science (TGDS), a paradigm designed to systematically merge established scientific knowledge with advanced data science techniques to produce physically consistent, generalizable, and interpretable models.

To demonstrate this paradigm, the article conducts an extensive conceptual evaluation and literature synthesis across multiple domains, including climate science, hydrology, material discovery, quantum chemistry, and aerospace engineering. The framework introduces physical consistency as a core performance metric alongside training accuracy and model simplicity, showing that domain knowledge can eliminate physically impossible solutions and reduce model variance without sacrificing accuracy. The article structures TGDS into five core approaches: guiding model design through tailored response functions and physically informed neural network architectures; guiding learning algorithms using informed parameter initializations, probabilistic priors, domain-driven constraints, and structured regularization; refining model outputs using explicit physical equations or inferred implicit orderings; constructing hybrid systems where machine learning resolves deficiencies in theoretical equations; and augmenting theory-based systems via observational data assimilation and automated parameter calibration.

The findings across various scientific case studies confirm that blending domain principles with data-driven methods delivers substantial performance gains. In materials science, integrating probabilistic models with theoretical calculations accelerated compound discovery, identifying approximately one hundred previously unknown ternary oxide compounds. In Earth observation, incorporating implicit elevation-based physical constraints enabled the accurate tracking and global mapping of dynamic surface water bodies despite noisy satellite data. In aerospace engineering, embedding machine learning corrections directly into Reynolds-averaged Navier-Stokes equations significantly closed performance gaps in complex turbulence simulations. Furthermore, in macroecology, initializing matrix completion models with species-level biological averages improved plant trait prediction accuracy significantly compared to standard, randomly initialized machine learning algorithms.

These results demonstrate that theory-guided data science mitigates the severe risks and costs of deploying unconstrained black-box models in high-stakes fields such as healthcare, aerospace, and environmental management. Anchoring models in established physical laws ensures that algorithms reflect causal mechanisms rather than spurious mathematical correlations. Consequently, organizations can safely extract value from large observational datasets without sacrificing interpretability, regulatory compliance, or operational safety. Decision-makers should prioritize TGDS methodologies over purely black-box machine learning when addressing complex physical systems, and research teams should develop hybrid workflows that explicitly embed physical laws into their model training objectives and evaluation pipelines. Future research must expand these methodologies beyond supervised regression to unsupervised pattern mining, automated workflow design, and comprehensive uncertainty quantification, while establishing standardized techniques to handle scientific knowledge that is uncertain, incomplete, or represented across differing physical scales.

arXiv: 1612.08544

No sufficiently relevant recommendations were found.

Cover for Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data

Abstract

Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of data science models in enabling scientific discovery. The overarching vision of TGDS is to introduce scientific consistency as an essential component for learning generalizable models. Further, by producing scientifically interpretable models, TGDS aims to advance our scientific understanding by discovering novel domain insights. Indeed, the paradigm of TGDS has started to gain prominence in a number of scientific disciplines such as turbulence modeling, material discovery, quantum chemistry, bio-medical science, bio-marker discovery, climate science, and hydrology. In this paper, we formally conceptualize the paradigm of TGDS and present a taxonomy of research themes in TGDS. We describe several approaches for integrating domain knowledge in different research themes using illustrative examples from different disciplines. We also highlight some of the promising avenues of novel research for realizing the full potential of theory-guided data science.

Table of Contents

  • I Introduction
  • II Theory-Guided Data Science
  • III Theory-guided Design of Data Science Models
  • III-A Theory-guided Specification of Response
  • III-B Theory-guided Design of Model Architecture
  • IV Theory-guided Learning of Data Science Models
  • IV-A Theory-guided Initialization
  • IV-B Theory-guided Probabilistic Models
  • IV-C Theory-guided Constrained Optimization
  • IV-D Theory-guided Regularization
  • V Theory-guided Refinement of Data Science Outputs
  • V-A Using Explicit Domain Knowledge
  • V-B Using Implicit Domain Knowledge
  • VI Learning Hybrid Models of Theory and Data Science
  • VII Augmenting Theory-based Models using Data Science
  • VII-A Data Assimilation in Theory-based Models
  • VII-B Calibrating Theory-based Models using Data
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Theory-Guided Data Science Paradigm and Objective

    definition

    Theory-guided data science (TGDS) is a paradigm for scientific discovery that systematically integrates domain-specific scientific principles and physical knowledge into data science and machine learning models. Traditional data science models optimize performance primarily as a balance of empirical training accuracy and statistical model complexity:

    Performance∝Accuracy+Simplicity\text{Performance} \propto \text{Accuracy} + \text{Simplicity}

    In scientific domains characterized by sample paucity, high dimensionality, non-stationarity, and physical constraints, purely data-driven models frequently learn spurious, non-generalizable correlations. TGDS reformulates the learning objective to explicitly include scientific and physical consistency:

    Performance∝Accuracy+Simplicity+Consistency\text{Performance} \propto \text{Accuracy} + \text{Simplicity} + \text{Consistency}

    Under this paradigm, candidate models must simultaneously fit observational data, maintain low model complexity, and adhere to established first principles, physical laws, or domain domain theories.

  2. Knowl 2 — Variance Reduction via Physical Consistency Constraints

    theoretical result

    In statistical machine learning, high-capacity model families M3\mathcal{M}_3 achieve low approximation bias on average but suffer from high estimation variance when trained on limited or non-representative empirical datasets. Conversely, highly constrained model families M1\mathcal{M}_1 exhibit low variance but suffer from high bias due to oversimplified assumptions.

    Scientific knowledge provides an independent constraint that prunes candidates within high-capacity model spaces that violate physical laws (such as mass and energy conservation, thermodynamic admissibility, or monotonic physical relationships). Restricting the optimization or parameter search to the physically consistent sub-manifold reduces the effective estimation variance of complex models without degrading their approximation bias relative to the true underlying physical process.

  3. Knowl 3 — Theory-Guided Design of Model Architecture and Response

    model/method

    Theory-guided model design integrates domain knowledge into the structural specification of predictive models prior to learning:

    1. Theory-Guided Specification of Response: In generalized linear models (GLMs), the expected mean μ=E[y∣x]\mu = \mathbb{E}[y \mid \mathbf{x}] of response variable yy given inputs x\mathbf{x} is defined via a link function g(⋅)g(\cdot), model parameters w\mathbf{w}, and intercept bb as:
    g(μ)=wTx+b  ⟺  μ=g−1(wTx+b)g(\mu) = \mathbf{w}^T \mathbf{x} + b \iff \mu = g^{-1}(\mathbf{w}^T \mathbf{x} + b)

    Domain physics dictates the appropriate probability distribution P(y∣x)P(y \mid \mathbf{x}) and link function g(⋅)g(\cdot). For physical phenomena involving extreme events (such as severe floods or droughts), standard Gaussian assumptions fail and are replaced with distributions from extreme value theory, such as the Gumbel distribution.

    1. Domain-Modular and Structured Neural Architectures: Deep neural networks can be structured by decomposing a physical system into interconnected sub-networks corresponding to distinct physical sub-processes (e.g., separate coupled modules for precipitation/evaporation, surface runoff, and subsurface seepage in hydrological systems). Furthermore, layer connectivity can reflect physical dynamics, such as using Long Short-Term Memory (LSTM) recurrent networks with domain-specific lag connections to model processes with differing latency characteristics (e.g., immediate surface runoff versus delayed groundwater seepage).
  4. Knowl 4 — Theory-Guided Parameter Initialization

    model/method

    Iterative optimization algorithms (such as gradient descent in artificial neural networks or non-convex matrix factorization) are sensitive to parameter initialization, often converging to poor local minima, flat saddle points, or physically meaningless solutions when observational data is sparse. Theory-guided initialization leverages physical domain estimates and simulation models to set initial parameter values:

    1. Domain-Mean Initialization: In sparse matrix completion tasks (such as imputing incomplete plant-trait matrices where rows represent organisms across environments and columns represent physiological traits), missing entries are initialized using species-level domain averages rather than arbitrary constants or random values, establishing a physically robust baseline before iterative matrix completion.

    2. Simulation-Based Pre-training: In deep neural network learning for complex physical systems (such as fluid flow), network weights are pre-trained on synthetic datasets generated by computationally inexpensive, approximate numerical simulations of first-principle physics models. The pre-trained model is subsequently fine-tuned using limited, high-fidelity empirical observations or ground-truth experimental data.

  5. Knowl 5 — Theory-Guided Priors in Probabilistic Models

    model/method

    In high-dimensional scientific inverse problems where the number of unknown physical state variables far exceeds the available sensor measurements, unconstrained probabilistic models are prone to overfitting and spurious correlations. Theory-guided probabilistic modeling incorporates dynamical physical differential equations as informative prior distributions over model states and parameters.

    For example, in non-invasive cardiac electrophysiological imaging, electrical activations across thousands of myocardial locations must be inferred from a sparse torso electrocardiogram (ECG) array. Physical differential equations governing electrical signal propagation along myocardial fibers are used to construct conditional spatial prior distributions P(st∣st−1)P(\mathbf{s}_t \mid \mathbf{s}_{t-1}), specifying the expected spatial distribution of cardiac potentials st\mathbf{s}_t at time tt given the state at t−1t-1. Incorporating these domain-derived priors within a hierarchical Bayesian inference framework regularizes the under-determined inverse problem and restricts parameter estimation to physically valid trajectories.

  6. Knowl 6 — Theory-Guided Constrained Optimization

    model/method

    Domain knowledge can be incorporated into machine learning models by formulating physical laws as explicit mathematical constraints within the optimization problem:

    1. Self-Consistent Density Functional Constraints: In computational quantum chemistry, machine learning models (such as kernel ridge regression) trained to predict the kinetic energy functional T[n]T[n] as a function of electron density n(r)n(\mathbf{r}) must satisfy the Euler-Lagrange variational condition for the ground-state density n0(r)n_0(\mathbf{r}):
    δT^[n0]δn0(r)=μ−v(r)\frac{\delta \hat{T}[n_0]}{\delta n_0(\mathbf{r})} = \mu - v(\mathbf{r})

    where v(r)v(\mathbf{r}) is the external potential, μ\mu is a Lagrange multiplier (chemical potential), and δδn\frac{\delta}{\delta n} denotes the functional derivative. Constrained optimization restricts the functional derivatives of the learned estimator T^[n]\hat{T}[n] to lie on the density manifold of the physical system, preventing unconstrained estimators from arriving at ill-conditioned solutions.

    1. Physical Ordering Constraints: In spatial classification problems, physical topography provides ordering constraints. In surface water body extent mapping from multi-spectral remote sensing, the concave bathymetric structure of water basins dictates that lower-elevation locations must fill before higher-elevation locations. Incorporating elevation monotonicity constraints enforces that if a location at elevation z1z_1 is classified as dry land, all connected locations at elevation z2>z1z_2 > z_1 must also be classified as dry land.
  7. Knowl 7 — Theory-Guided Regularization and Structured Sparsity

    model/method

    Standard regularizers (such as L1L_1 Lasso or L2L_2 ridge penalties) penalize model complexity based solely on mathematical norms, which can eliminate physically causal attributes in favor of correlated non-causal proxies. Theory-guided regularization designs penalty terms based on physical structures:

    1. Structured and Group Sparsity: Group Lasso and Sparse Group Lasso penalties enforce parameter selection over predefined, physically meaningful feature groupings. In climate science, spatially contiguous variables are grouped to identify coherent spatial regions and teleconnection pathways.

    2. Linkage and Difference Penalties: In bioinformatics (e.g., genome-wide association studies), linkage disequilibrium dictates that adjacent genomic markers travel together across generations. This physical structure is encoded via a difference regularizer that penalizes discrepancies between adjacent regression coefficients βj\beta_j and βj+1\beta_{j+1}, such as ∑j(βj−βj+1)2\sum_j (\beta_j - \beta_{j+1})^2 or smoothed minimax concave penalties.

    3. Multi-Task Learning over Heterogeneous Sub-populations: When physical relationships vary across distinct sub-domains (e.g., differing vegetation types in remote sensing), multi-task learning treats each sub-domain as a separate task, regularizing parameters by sharing information across physically related tasks inferred via domain-guided clustering of contextual variables.

  8. Knowl 8 — Theory-Guided Post-Processing and Output Refinement

    model/method

    Theory-guided output refinement enforces physical consistency on the predictions generated by data science models during post-processing:

    1. Explicit Theory Verification: Candidate outputs generated by computationally efficient probabilistic models (e.g., predicted stable crystal structures in high-throughput materials discovery) are screened and filtered using rigorous, first-principle ab initio calculations (such as density functional theory), eliminating physically invalid predictions.

    2. Implicit Constraint Estimation and Iterative Refinement: When exact physical parameters are unobserved (such as fine-scale lake bathymetry in global remote sensing), physical constraints can be inferred from long-term prediction statistics. Under the physical rule that lower-elevation points are submerged more frequently than higher-elevation points, historical time-series classifications of water presence are used to infer a latent elevation ordering across spatial pixels. The model outputs and latent physical ordering are then iteratively solved and refined to ensure full elevation-monotonic consistency across all spatial classification maps.

  9. Knowl 9 — Hybrid Theory-Data Science Modeling Architectures

    model/method

    Hybrid modeling combines theory-based dynamical models and data-driven learning components into unified frameworks:

    1. Two-Component Cascading and Downscaling: Coarse-resolution simulations from numerical dynamical models (such as general circulation models in climate science) serve as input features to statistical machine learning models that learn downscaling functions to predict physical variables at fine spatial and temporal scales.

    2. Direct ML Augmentation of Governing Equations (Field Inversion and Machine Learning - FIML): In fluid dynamics governed by the Navier-Stokes equations, direct numerical simulations are computationally prohibitive, leading to approximations via Reynolds-averaged Navier-Stokes (RANS) equations with unclosed Reynolds stress τ\tau. Machine learning models can predict an additive discrepancy term:

    τ=τRANS+ΔτML\tau = \tau_{\text{RANS}} + \Delta \tau_{\text{ML}}

    or embed directly into the transport equations for turbulent eddy viscosity ν\nu:

    −τij=2ρνSij∗−23ρKδij-\tau_{ij} = 2 \rho \nu S^*_{ij} - \frac{2}{3} \rho K \delta_{ij} DνDt=βML×P−D+T\frac{D\nu}{Dt} = \beta_{\text{ML}} \times P - D + T

    where ρ\rho is fluid density, Sij∗S^*_{ij} is the traceless rate of strain tensor, KK is turbulent kinetic energy, δij\delta_{ij} is the Kronecker delta, DDt\frac{D}{Dt} is the material derivative, PP, DD, and TT denote physical production, destruction, and transport terms, and βML\beta_{\text{ML}} is a correction factor learned via an artificial neural network.

  10. Knowl 10 — Augmenting Theory-Based Models via Data Assimilation and Calibration

    model/method

    Data science approaches enhance theory-based numerical models through state estimation and parameter calibration:

    1. Data Assimilation: In numerical dynamical systems (e.g., numerical weather prediction and hydrology), data assimilation infers the optimal sequence of physical system states by constraining model states to depend simultaneously on physics-governed transition models from preceding states and incoming observational measurements at each time step (e.g., via ensemble Kalman filtering or variational assimilation schemes).

    2. Bayesian and Multi-Armed Bandit Parameter Calibration: Theory-based physical models contain high-dimensional, empirically unmeasurable parameter vectors. Calibration methods like Generalized Likelihood Uncertainty Estimation (GLUE) update probability distributions over the parameter space using Monte Carlo sampling and Bayesian updates against observed data. Machine learning multi-armed bandit formulations (including continuum-armed bandits) optimize the exploration-exploitation trade-off to identify parameter configurations that maximize data likelihood under finite computational evaluations.

Coverage note — All primary conceptual foundations, taxonomy themes, and domain-specific formulation frameworks of theory-guided data science introduced in the paper are covered; domain background reviews and author biographical notes were omitted.

References

  1. 1.G. Bell, T. Hey, and A. Szalay, ‘‘Beyond the data deluge,’’ Science, vol. 323, no. 5919, pp. 1297–1298, 2009.
  2. 2.Economist, ‘‘The data deluge,’’ Special Supplement, 2010.
  3. 3.M. James, C. Michael, B. Brad, B. Jacques, D. Richard, R. Charles, and H. Angela, ‘‘Big data: The next frontier for innovation, com- petition, and productivity,’’ The McKinsey Global Institute, 2011.
  4. 4.A. Halevy, P. Norvig, and F. Pereira, ‘‘The unreasonable effective- ness of data,’’ Intelligent Systems, IEEE, vol. 24, no. 2, pp. 8–12, 2009.
  5. 5.B. P. Roe, H.-J. Yang, J. Zhu, Y. Liu, I. Stancu, and G. McGregor, ‘‘Boosted decision trees as an alternative to artificial neural net- works for particle identification,’’ Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, vol. 543, no. 2, pp. 577–584, 2005.
  6. 6.D. Castelvecchi et al., ‘‘Artificial intelligence called in to tackle lhc data deluge,’’ Nature, vol. 528, no. 7580, pp. 18–19, 2015.
  7. 7.P. Baldi and S. Brunak, Bioinformatics: the machine learning approach. MIT press, 2001.
  8. 8.J. H. Faghmous and V. Kumar, ‘‘A Big Data Guide to Understand- ing Climate Change: The Case for Theory-Guided Data Science,’’ Big Data, vol. 3, 2014.
  9. 9.J. H. Faghmous, V. Kumar, and S. Shekhar, ‘‘Computing and climate,’’ Computing in Science & Engineering, vol. 17, no. 6, pp. 6–8, 2015.
  10. 10.D. Graham-Rowe, D. Goldston, C. Doctorow, M. Waldrop, C. Lynch, F. Frankel, R. Reid, S. Nelson, D. Howe, S. Rhee et al., ‘‘Big data: science in the petabyte era,’’ Nature, vol. 455, no. 7209, pp. 8–9, 2008.
  11. 11.T. Jonathan, A. Gerald et al., ‘‘Special issue: dealing with data,’’ Science, vol. 331, no. 6018, pp. 639–806, 2011.
  12. 12.T. J. Sejnowski, P. S. Churchland, and J. A. Movshon, ‘‘Putting big data to good use in neuroscience,’’ Nature neuroscience, vol. 17, no. 11, pp. 1440–1441, 2014.
  13. 13.C. Anderson, ‘‘The End of Theory: The Data Deluge Makes the Scientific Method Obsolete,’’ Wired Magazine, 2008.
  14. 14.P. M. Caldwell, C. S. Bretherton, M. D. Zelinka, S. A. Klein, B. D. Santer, and B. M. Sanderson, ‘‘Statistical significance of climate sensitivity predictors obtained by data mining,’’ Geophysical Re- search Letters, vol. 41, no. 5, pp. 1803–1808, 2014.
  15. 15.D. Lazer, R. Kennedy, G. King, and A. Vespignani, ‘‘The Parable of Google Flu: Traps in Big Data Analysis,’’ Science (New York, N.Y.), vol. 343, no. 6176, pp. 1203–5, Mar. 2014. [Online]. Available: http://www.ncbi.nlm.nih.gov/pubmed/24626916
  16. 16.G. Marcus and E. Davis, ‘‘Eight (no, nine!) problems with big data,’’ The New York Times, vol. 6, no. 04, p. 2014, 2014.
  17. 17.J. Ginsberg, M. H. Mohebbi, R. S. Patel, L. Brammer, M. S. Smolin- ski, and L. Brilliant, ‘‘Detecting influenza epidemics using search engine query data,’’ Nature, vol. 457, no. 7232, pp. 1012–1014, 2009.
  18. 18.J. Kawale, S. Liess, A. Kumar, M. Steinbach, P. Snyder, V. Kumar, A. R. Ganguly, N. F. Samatova, and F. Semazzi, ‘‘A graph-based approach to find teleconnections in climate data,’’ Statistical Anal- ysis and Data Mining, vol. 6, no. 3, pp. 158–179, 2013.
  19. 19.J. H. Faghmous, I. Frenger, Y. Yao, R. Warmka, A. Lindell, and V. Kumar, ‘‘A daily global mesoscale ocean eddy dataset from satellite altimetry,’’ Scientific data, vol. 2, 2015.
  20. 20.A. P. Singh, S. Medida, and K. Duraisamy, ‘‘Machine learning- augmented predictive modeling of turbulent separated flows over airfoils,’’ arXiv preprint arXiv:1608.03990, 2016.
  21. 21.J.-X. Wang, J.-L. Wu, and H. Xiao, ‘‘Physics-informed ma- chine learning for predictive turbulence modeling: Using data to improve rans modeled reynolds stresses,’’ arXiv preprint arXiv:1606.07987, 2016.
  22. 22.G. Hautier, C. C. Fischer, A. Jain, T. Mueller, and G. Ceder, ‘‘Find- ing natures missing ternary oxide compounds using machine learning and density functional theory,’’ Chemistry of Materials, vol. 22, no. 12, pp. 3762–3767, 2010.
  23. 23.C. C. Fischer, K. J. Tibbetts, D. Morgan, and G. Ceder, ‘‘Predicting crystal structure by merging data mining with quantum mechan- ics,’’ Nature materials, vol. 5, no. 8, pp. 641–646, 2006.
  24. 24.S. Curtarolo, G. L. Hart, M. B. Nardelli, N. Mingo, S. Sanvito, and O. Levy, ‘‘The high-throughput highway to computational materials design,’’ Nature materials, vol. 12, no. 3, pp. 191–201, 2013.
  25. 25.L. Li, J. C. Snyder, I. M. Pelaschier, J. Huang, U.-N. Niranjan, P. Duncan, M. Rupp, K.-R. Müller, and K. Burke, ‘‘Understand- ing machine-learned density functionals,’’ International Journal of Quantum Chemistry, 2015.
  26. 26.K. C. Wong, L. Wang, and P. Shi, ‘‘Active model with orthotropic hyperelastic material for cardiac image analysis,’’ in Functional Imaging and Modeling of the Heart. Springer, 2009, pp. 229–238.
  27. 27.J. Xu, J. L. Sapp, A. R. Dehaghani, F. Gao, M. Horacek, and L. Wang, ‘‘Robust transmural electrophysiological imaging: Inte- grating sparse and dynamic physiological models into ecg-based inference,’’ in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015. Springer, 2015, pp. 519–527.
  28. 28.J. Liu, K. Wang, S. Ma, and J. Huang, ‘‘Accounting for linkage disequilibrium in genome-wide association studies: a penalized regression method,’’ Statistics and its interface, vol. 6, no. 1, p. 99, 2013.
  29. 29.A. Khandelwal, V. Mithal, and V. Kumar, ‘‘Post classification label refinement using implicit ordering constraint among data instances,’’ in Data Mining (ICDM), 2015 IEEE International Confer- ence on. IEEE, 2015, pp. 799–804.
  30. 30.A. Khandelwal, A. Karpatne, M. Marlier, J. Kim, D. Lettenmaier, and V. Kumar, ‘‘An approach for global monitoring of surface water extent variations using modis data,’’ in Remote Sensing of Environment (in review), 2017.
  31. 31.N. Wagner and J. M. Rondinelli, ‘‘Theory-guided machine learning in materials science,’’ Frontiers in Materials, vol. 3, p. 28, 2016.
  32. 32.J. Faghmous, A. Banerjee, S. Shekhar, M. Steinbach, V. Kumar, A. R. Ganguly, and N. Samatova, ‘‘Theory-guided data science for climate change,’’ Computer, vol. 47, no. 11, pp. 74–78, 2014.
  33. 33.A. R. Ganguly, E. A. Kodra, A. Agrawal, A. Banerjee, S. Bo- riah, S. Chatterjee, S. Chatterjee, A. Choudhary, D. Das, J. Fagh- mous, P. Ganguli, S. Ghosh, K. Hayhoe, C. Hays, W. Hen- drix, Q. Fu, J. Kawale, D. Kumar, V. Kumar, W. Liao, S. Liess, R. Mawalagedara, V. Mithal, R. Oglesby, K. Salvi, P. K. Snyder, K. Steinhaeuser, D. Wang, and D. Wuebbles, ‘‘Toward enhanced understanding and projections of climate extremes using physics- guided data mining techniques,’’ Nonlinear Processes in Geophysics, vol. 21, no. 4, pp. 777–795, 2014.
  34. 34.Physics Informed Machine Learning Conference, Santa Fe, New Mex- ico, 2016.
  35. 35.‘‘Physical analytics, ibm research,’’ http://researcher.watson.ibm. com/researcher/view_group.php?id=6566, accessed: 2016-10-20.
  36. 36.M. Ghasemizade and M. Schirmer, ‘‘Subsurface flow contribution in the hydrological cycle: lessons learned and challenges ahead–a review,’’ Environmental earth sciences, vol. 69, no. 2, pp. 707–718, 2013.
  37. 37.C. Paniconi and M. Putti, ‘‘Physically based modeling in catch- ment hydrology at 50: Survey and outlook,’’ Water Resources Research, vol. 51, no. 9, pp. 7090–7129, 2015.
  38. 38.M. F. Bierkens, ‘‘Global hydrology 2015: State, trends, and direc- tions,’’ Water Resources Research, vol. 51, no. 7, pp. 4923–4947, 2015.
  39. 39.J. Friedman, T. Hastie, and R. Tibshirani, The elements of statistical learning. Springer series in statistics Springer, Berlin, 2001, vol. 1.
  40. 40.P.-N. Tan, M. Steinbach, and V. Kumar, Intorduction to Data Mining. Addison-Wesley, 2005.
  41. 41.V. N. Vapnik and V. Vapnik, Statistical learning theory. Wiley New York, 1998, vol. 1.
  42. 42.M. D. Twa, S. Parthasarathy, C. Roberts, A. M. Mahmoud, T. W. Raasch, and M. A. Bullimore, ‘‘Automated decision tree classi- fication of corneal shape,’’ Optometry and vision science: official publication of the American Academy of Optometry, vol. 82, no. 12, p. 1038, 2005.
  43. 43.J. Z. Leibo, Q. Liao, F. Anselmi, W. A. Freiwald, and T. Pog- gio, ‘‘View-tolerant face recognition and hebbian learning imply mirror-symmetric neural tuning to head orientation,’’ Current Bi- ology, 2016.
  44. 44.T. Mikolov, M. Karafiát, L. Burget, J. Černocký, and S. Khudanpur, ‘‘Recurrent neural network based language model.’’ in Interspeech, vol. 2, 2010, p. 3.
  45. 45.H. Sak, A. W. Senior, and F. Beaufays, ‘‘Long short-term memory recurrent neural network architectures for large scale acoustic modeling.’’ in INTERSPEECH, 2014, pp. 338–342.
  46. 46.A. Krizhevsky, I. Sutskever, and G. E. Hinton, ‘‘Imagenet classifi- cation with deep convolutional neural networks,’’ in Advances in neural information processing systems, 2012, pp. 1097–1105.
  47. 47.F. Schrodt, J. Kattge, H. Shan, F. Fazayeli, J. Joswig, A. Banerjee, M. Reichstein, G. Bönisch, S. Díaz, J. Dickie et al., ‘‘Bhpmf–a hierarchical bayesian approach to gap-filling and trait prediction for macroecology and functional biogeography,’’ Global Ecology and Biogeography, vol. 24, no. 12, pp. 1510–1521, 2015.
  48. 48.J. Kattge, S. Diaz, S. Lavorel, I. Prentice, P. Leadley, G. Bönisch, E. Garnier, M. Westoby, P. B. Reich, I. Wright et al., ‘‘Try–a global database of plant traits,’’ Global change biology, vol. 17, no. 9, pp. 2905–2935, 2011.
  49. 49.P. Melville and V. Sindhwani, ‘‘Recommender systems,’’ in Ency- clopedia of machine learning. Springer, 2011, pp. 829–838.
  50. 50.J. Friedman, T. Hastie, and R. Tibshirani, ‘‘Sparse inverse covari- ance estimation with the graphical lasso,’’ Biostatistics, vol. 9, no. 3, pp. 432–441, 2008.
  51. 51.H. Denli, N. Subrahmanya et al., ‘‘Multi-scale graphical models for spatio-temporal processes,’’ in Advances in Neural Information Processing Systems, 2014, pp. 316–324.
  52. 52.J.-F. Boulicaut and B. Jeudy, ‘‘Constraint-based data mining,’’ in Data Mining and Knowledge Discovery Handbook. Springer, 2005, pp. 399–416.
  53. 53.J. Pei and J. Han, ‘‘Constrained frequent pattern mining: a pattern- growth view,’’ ACM SIGKDD Explorations Newsletter, vol. 4, no. 1, pp. 31–39, 2002.
  54. 54.S. Basu, I. Davidson, and K. Wagstaff, Constrained clustering: Ad- vances in algorithms, theory, and applications. CRC Press, 2008.
  55. 55.A. J. Majda and J. Harlim, ‘‘Physics constrained nonlinear regres- sion models for time series,’’ Nonlinearity, vol. 26, no. 1, p. 201, 2012.
  56. 56.A. J. Majda and Y. Yuan, ‘‘Fundamental limitations of ad hoc linear and quadratic multi-level regression models for physical systems,’’ Discrete and Continuous Dynamical Systems B, vol. 17, no. 4, pp. 1333–1363, 2012.
  57. 57.P. Hohenberg and W. Kohn, ‘‘Inhomogeneous electron gas,’’ Phys- ical review, vol. 136, no. 3B, p. B864, 1964.
  58. 58.A. Karpatne, A. Khandelwal, X. Chen, V. Mithal, J. Faghmous, and V. Kumar, ‘‘Global monitoring of inland water dynamics: state-of-the-art, challenges, and opportunities,’’ in Computational Sustainability. Springer, 2016, pp. 121–147.
  59. 59.A. Karpatne, Z. Jiang, R. R. Vatsavai, S. Shekhar, and V. Kumar, ‘‘Monitoring land-cover changes: A machine-learning perspec- tive,’’ IEEE Geoscience and Remote Sensing Magazine, vol. 4, no. 2, pp. 8–21, 2016.
  60. 60.G. M. James and P. Radchenko, ‘‘A generalized dantzig selector with shrinkage tuning,’’ Biometrika, vol. 96, no. 2, pp. 323–337, 2009.
  61. 61.S. Chatterjee, S. Chen, and A. Banerjee, ‘‘Generalized dantzig selector: Application to the k-support norm,’’ in Advances in Neural Information Processing Systems, 2014, pp. 1934–1942.
  62. 62.M. Yuan and Y. Lin, ‘‘Model selection and estimation in regression with grouped variables,’’ Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 68, no. 1, pp. 49–67, 2006.
  63. 63.L. Jacob, G. Obozinski, and J.-P. Vert, ‘‘Group lasso with overlap and graph lasso,’’ in Proceedings of the 26th annual international conference on machine learning. ACM, 2009, pp. 433–440.
  64. 64.S. Kim and E. P. Xing, ‘‘Tree-guided group lasso for multi-task regression with structured sparsity,’’ in Proceedings of the 27th International Conference on Machine Learning (ICML-10), 2010, pp. 543–550.
  65. 65.J. Friedman, T. Hastie, and R. Tibshirani, ‘‘A note on the group lasso and a sparse group lasso,’’ arXiv preprint arXiv:1001.0736, 2010.
  66. 66.S. Chatterjee, K. Steinhaeuser, A. Banerjee, S. Chatterjee, and A. R. Ganguly, ‘‘Sparse group lasso: Consistency and climate applica- tions.’’ in SDM. SIAM, 2012, pp. 47–58.
  67. 67.A. Karpatne, A. Khandelwal, S. Boriah, and V. Kumar, ‘‘Predictive learning in the presence of heterogeneity and limited training data.’’ in SDM. SIAM, 2014, pp. 253–261.
  68. 68.V. Mithal, A. Khandelwal, S. Boriah, K. Steinhaeuser, and V. Ku- mar, ‘‘Change detection from temporal sequences of class labels: application to land cover change mapping,’’ in SIAM International Conference on Data mining, Austin, TX, USA. Citeseer, 2013, pp. 2–4.
  69. 69.X. Jia, A. Khandelwal, J. Gerber, K. Carlson, P. West, L. Samberg, and V. Kumar, ‘‘Automated plantation mapping in southeast asia using remote sensing data,’’ Department of Computer Science, University of Minnesota, Twin Cities, Tech. Rep. 16-029, 2016.
  70. 70.X. Jia, A. Khandelwal, N. Guru, J. Gerber, K. Carlson, P. West, and V. Kumar, ‘‘Predict land covers with transition modeling and incremental learning,’’ in SIAM International Conference on Data Mining (accepted), 2017.
  71. 71.R. L. Wilby, T. Wigley, D. Conway, P. Jones, B. Hewitson, J. Main, and D. Wilks, ‘‘Statistical downscaling of general circulation model output: a comparison of methods,’’ Water resources research, vol. 34, no. 11, pp. 2995–3008, 1998.
  72. 72.P. Sadowski, D. Fooshee, N. Subrahmanya, and P. Baldi, ‘‘Syn- ergies between quantum mechanics and machine learning in reaction prediction,’’ Journal of Chemical Information and Modeling, vol. 56, no. 11, pp. 2125–2128, 2016.
  73. 73.E. J. Parish and K. Duraisamy, ‘‘A paradigm for data-driven predictive modeling using field inversion and machine learning,’’ Journal of Computational Physics, vol. 305, pp. 758–774, 2016.
  74. 74.G. Evensen, Data assimilation: the ensemble Kalman filter. Springer Science & Business Media, 2009.
  75. 75.K. Beven and A. Binley, ‘‘The future of distributed models: model calibration and uncertainty prediction,’’ Hydrological processes, vol. 6, no. 3, pp. 279–298, 1992.
  76. 76.O. Chapelle and L. Li, ‘‘An empirical evaluation of thompson sampling,’’ in Advances in neural information processing systems, 2011, pp. 2249–2257.
  77. 77.L. Li, W. Chu, J. Langford, and R. E. Schapire, ‘‘A contextual- bandit approach to personalized news article recommendation,’’ in Proceedings of the 19th international conference on World wide web. ACM, 2010, pp. 661–670.
  78. 78.M. Jordan and T. Mitchell, ‘‘Machine learning: Trends, perspec- tives, and prospects,’’ Science, vol. 349, no. 6245, pp. 255–260, 2015.
  79. 79.R. Agrawal, ‘‘The continuum-armed bandit problem,’’ SIAM jour- nal on control and optimization, vol. 33, no. 6, pp. 1926–1951, 1995.
  80. 80.R. D. Kleinberg, ‘‘Nearly tight bounds for the continuum-armed bandit problem,’’ in Advances in Neural Information Processing Sys- tems, 2004, pp. 697–704.
  81. 81.J. H. Faghmous, L. Styles, V. Mithal, S. Boriah, S. Liess, F. Vikebo, M. d. S. Mesquita, and V. Kumar, ‘‘Eddyscan: A physically con- sistent ocean eddy monitoring application,’’ in Intelligent Data Understanding (CIDU), 2012 Conference on, oct. 2012, pp. 96 –103.
  82. 82.J. H. Faghmous, H. Nguyen, M. Le, and V. Kumar, ‘‘Spatio- temporal consistency as a means to identify unlabeled objects in a continuous data field,’’ in AAAI, 2014, pp. 410–416.
  83. 83.V. G. Honavar, ‘‘The promise and potential of big data: A case for discovery informatics,’’ Review of Policy Research, vol. 31, no. 4, pp. 326–330, 2014.
  84. 84.Y. Gil, E. Deelman, M. Ellisman, T. Fahringer, G. Fox, D. Gannon, C. Goble, M. Livny, L. Moreau, and J. Myers, ‘‘Examining the challenges of scientific workflows,’’ Ieee computer, vol. 40, no. 12, pp. 26–34, 2007.
  85. 85.M. Paganini, L. d. Oliveira, and B. Nachman, ‘‘Calogan: Simulating 3d high energy particle showers in multi-layer electromagnetic calorimeters with generative adversarial networks,’’ arXiv preprint arXiv:1705.02355, 2017.
  86. 86.T. Hey, S. Tansley, K. M. Tolle et al., The fourth paradigm: data- intensive scientific discovery. Microsoft research Redmond, WA, 2009, vol. 1.

Citation

MLA
Karpatne, A., et al. “Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data”. IEEE Transactions on Knowledge and Data Engineering, vol. 29, no. 10, 2017, pp. 2318–31, https://doi.org/10.1109/TKDE.2017.2720168.
APA
Karpatne, A., Atluri, G., Faghmous, J. H., Steinbach, M., Banerjee, A., Ganguly, A., Shekhar, S., Samatova, N., & Kumar, V. (2017). Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data. IEEE Transactions on Knowledge and Data Engineering, 29(10), 2318–2331. https://doi.org/10.1109/TKDE.2017.2720168
Chicago
Karpatne, A., G. Atluri, J. H. Faghmous, et al. 2017. “Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data”. IEEE Transactions on Knowledge and Data Engineering 29 (10): 2318–31. https://doi.org/10.1109/TKDE.2017.2720168.
Harvard
Karpatne, A. et al. (2017) “Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data”, IEEE Transactions on Knowledge and Data Engineering, 29(10), pp. 2318–2331. Available at: https://doi.org/10.1109/TKDE.2017.2720168.
Vancouver
1. Karpatne A, Atluri G, Faghmous JH, Steinbach M, Banerjee A, Ganguly A, Shekhar S, Samatova N, Kumar V (2017) Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data. IEEE Transactions on Knowledge and Data Engineering 29:2318–2331

BibTeX

@article{Karpatne_2017, title={Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data}, volume={29}, ISSN={1041-4347}, url={http://dx.doi.org/10.1109/TKDE.2017.2720168}, DOI={10.1109/tkde.2017.2720168}, number={10}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Karpatne, Anuj and Atluri, Gowtham and Faghmous, James H. and Steinbach, Michael and Banerjee, Arindam and Ganguly, Auroop and Shekhar, Shashi and Samatova, Nagiza and Kumar, Vipin}, year={2017}, month=Oct, pages={2318–2331} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF