Characterizing possible failure modes in physics-informed neural networks

Aditi S. KrishnapriyanAmir GholamiShandian ZheRobert M. KirbyMichael W. Mahoney

article2021NeurIPS1,380 citations

Demonstrates that physics-informed neural networks fail on complex differential equations due to severe optimization bottlenecks rather than limited model capacity, proposing curriculum regularization and sequence-to-sequence training to reduce prediction error by up to two orders of magnitude.

Listen

Scientific machine learning increasingly relies on physics-informed neural networks to solve complex differential equations across engineering and scientific disciplines. These frameworks integrate known physical laws directly into the neural network training objective as soft constraints or penalty terms. However, severe and often overlooked failure modes arise when applying these models to even moderately complex physical scenarios, presenting substantial operational risks for organizations deploying them in high-stakes simulations.

The article systematically evaluates why standard physics-informed neural network formulations fail on foundational differential equations and demonstrates effective algorithmic interventions to resolve these failures. The authors investigate benchmark systems covering transport, chemical reaction, and reaction-diffusion processes using fully connected networks optimized with quasi-Newton methods across varying parameter regimes, comparing predictive accuracy against known analytical solutions.

The investigation reveals three critical findings. First, standard networks succeed only under simple conditions with small physical coefficients; once convection, reaction, or diffusion coefficients increase moderately, relative prediction errors surge to between 50% and nearly 100%. Second, network architecture capacity is not the root cause. Instead, incorporating differential equations as soft constraints creates highly rugged, non-convex optimization landscapes that trap standard training algorithms in poor local minima. Third, the authors propose and evaluate two targeted solutions: curriculum regularization—which gradually scales problem complexity during training—and sequence-to-sequence time-marching, which solves the domain across sequential time intervals rather than predicting entire space-time trajectories simultaneously. Both approaches smooth the optimization landscape and reduce prediction errors by one to two orders of magnitude.

These findings indicate that relying on out-of-the-box physics-informed neural networks without specialized training workflows introduces significant simulation error and performance failure. Organizations should immediately discontinue global space-time training for complex physical systems, adopting curriculum regularization or sequential time-stepping to ensure robust model convergence and lower error variance. While the analysis is currently limited to canonical one-dimensional benchmarks, confidence in the diagnostic findings remains high, underscoring the necessity of evaluating advanced multi-dimensional implementations before large-scale deployment.

Cover for Characterizing possible failure modes in physics-informed neural networks

Abstract

Recent work in scientific machine learning has developed so-called physics-informed neural network (PINN) models. The typical approach is to incorporate physical domain knowledge as soft constraints on an empirical loss function and use existing machine learning methodologies to train the model. We demonstrate that, while existing PINN methodologies can learn good models for relatively trivial problems, they can easily fail to learn relevant physical phenomena for even slightly more complex problems. In particular, we analyze several distinct situations of widespread physical interest, including learning differential equations with convection, reaction, and diffusion operators. We provide evidence that the soft regularization in PINNs, which involves PDE-based differential operators, can introduce a number of subtle problems, including making the problem more ill-conditioned. Importantly, we show that these possible failure modes are not due to the lack of expressivity in the NN architecture, but that the PINN's setup makes the loss landscape very hard to optimize. We then describe two promising solutions to address these failure modes. The first approach is to use curriculum regularization, where the PINN's loss term starts from a simple PDE regularization, and becomes progressively more complex as the NN gets trained. The second approach is to pose the problem as a sequence-to-sequence learning task, rather than learning to predict the entire space-time at once. Extensive testing shows that we can achieve up to 1-2 orders of magnitude lower error with these methods as compared to regular PINN training.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Possible failure modes for physics-informed neural networks
  • 3.1 Learning convection
  • 3.2 Learning reaction-diffusion
  • 4 Diagnosing possible failure modes for physics-informed NNs
  • 4.1 Soft PDE regularization and optimization difficulties
  • 5 Expressivity versus optimization difficulty
  • 5.1 Curriculum PINN Regularization
  • 5.2 Sequence-to-sequence learning vs learning the entire space-time solution
  • 6 Conclusions
  • 7 Acknowledgements.
  • References
  • A Learning reaction
  • B A PDE perspective on ill-conditioned regularization
  • B.1 Approximate condition number scaling for PDE regularization in PINNs
  • C Learning convection
  • D Learning reaction-diffusion
  • E Extra Results
  • E.1 Extra results for loss landscapes when varying the λ\lambda parameter
  • E.2 Extra curriculum regularization results
  • E.3 Extra sequence-to-sequence learning results

Knowls

  1. Knowl 1 — Condition Number Scaling of Differential Regularization Operators in PINNs

    theoretical result

    Incorporating partial differential equation (PDE) constraints as soft regularization terms in physics-informed neural network (PINN) loss functions introduces differential operators that can be severely ill-conditioned. For a uniformly discretized grid of spatial size NN with spacing h=1/Nh = 1/N and time step δt\delta t:

    1. Convection operator F(u)=∂u∂t+β∂u∂x\mathcal{F}(u) = \frac{\partial u}{\partial t} + \beta \frac{\partial u}{\partial x}: For a perturbation δu\delta u, the perturbation in the operator residual scales as O(δt−1)+O(βh−1)\mathcal{O}(\delta t^{-1}) + \mathcal{O}(\beta h^{-1}). When δt\delta t is not vanishingly small, the condition number of the operator scales as O(βN)\mathcal{O}(\beta N). Under an L2L_2 squared penalty, this ill-conditioning scales quadratically as O(βN)2\mathcal{O}(\beta N)^2.

    2. Diffusion operator F(u)=∂u∂t−ν∂2u∂x2\mathcal{F}(u) = \frac{\partial u}{\partial t} - \nu \frac{\partial^2 u}{\partial x^2}: For a perturbation δu\delta u, the perturbation in the operator residual scales as O(δt−1)+O(νh−2)\mathcal{O}(\delta t^{-1}) + \mathcal{O}(\nu h^{-2}). The condition number scales as O(νN2)\mathcal{O}(\nu N^2), and with an L2L_2 squared penalty, it scales quadratically as O(νN2)2\mathcal{O}(\nu N^2)^2.

    Because the diffusion operator's condition number scales with N2N^2 rather than NN, soft regularization for diffusion-dominated systems induces significantly faster growth in numerical instability and optimization difficulty than convection-dominated systems as spatial resolution or physical coefficients increase.

  2. Knowl 2 — Optimization Pathology Versus Expressivity in PINN Failure Modes

    empirical result

    The failure of standard physics-informed neural networks (PINNs) to learn solutions for non-trivial PDE parameter regimes (such as high convection velocity β\beta, reaction rate ρ\rho, or diffusion coefficient ν\nu) is caused by optimization failure rather than an architectural expressivity bottleneck.

    Visualizing the empirical loss landscape of trained PINN models along the two dominant Hessian eigenvectors shows:

    • In easy regimes (e.g., convection coefficient β=1\beta = 1), the loss landscape is smooth and convex around the solution.
    • As physical coefficients increase (e.g., β≥20\beta \ge 20), the loss landscape becomes highly non-convex, non-symmetric, and rugged with numerous local minima, causing optimizers like L-BFGS to become trapped at poor, high-loss configurations.
    • Scaling the soft-regularization weight λF\lambda_F does not resolve this pathology; higher λF\lambda_F increases the roughness and magnitude of the loss landscape, while lower λF\lambda_F leads to solutions that fail to satisfy the governing equations.
    • When trained with optimization warm-starting (curriculum regularization) or domain decomposition (sequence-to-sequence time marching), identical neural network architectures achieve low errors, verifying that the models possess sufficient representation capacity.
  3. Knowl 3 — Curriculum Regularization for PINN Training

    algorithm

    Curriculum regularization is an optimization strategy for physics-informed neural networks (PINNs) that prevents the optimizer from becoming trapped in poor local minima by warm-starting network parameters on easier versions of the target PDE problem before progressively increasing the difficulty of the PDE operator.

    Input: Initial PDE parameter α0\alpha_0, target PDE parameter αtarget\alpha_{\text{target}}, step size Δα\Delta \alpha, initial condition h(x)h(x), spatial domain Ω\Omega, time domain [0,T][0, T]
    Output: Trained neural network parameters θ\theta
    Initialize neural network parameters θ\theta randomly
    α←α0\alpha \leftarrow \alpha_0
    while α≤αtarget\alpha \le \alpha_{\text{target}} do
        Train network NN(θ,x,t)\text{NN}(\theta, x, t) by minimizing L(θ;α)=Lu0+Lub+λFLF(θ;α)\mathcal{L}(\theta; \alpha) = \mathcal{L}_{u_0} + \mathcal{L}_{u_b} + \lambda_F \mathcal{L}_{F}(\theta; \alpha) via L-BFGS
        θ←θ∗\theta \leftarrow \theta^*
        α←α+Δα\alpha \leftarrow \alpha + \Delta \alpha
    return θ\theta

    Here, α\alpha represents a physical parameter such as the convection coefficient β\beta or reaction coefficient ρ\rho. The network parameters θ∗\theta^* obtained from training at a lower parameter value serve as the weight initialization for the subsequent, more difficult parameter value. This progressive continuation smooths the loss landscape, substantially lowers error variance across random seeds, and yields 1 to 2 orders of magnitude lower relative error compared to training directly on the target parameter.

  4. Knowl 4 — Sequence-to-Sequence Time-Marching PINN Learning

    algorithm

    Sequence-to-sequence (seq2seq) time-marching PINN learning decomposes a global space-time domain Ω×[0,T]\Omega \times [0, T] into consecutive discrete time intervals of length Δt\Delta t, training the neural network over one time segment at a time instead of predicting the entire space-time volume simultaneously.

    Input: Total time horizon TT, segment duration Δt\Delta t, initial state condition h(x)h(x), spatial domain Ω\Omega, collocation point budget per segment NfN_f
    Output: Predicted state trajectory u^(x,t)\hat{u}(x, t) for all t∈[0,T]t \in [0, T]
    tstart←0t_{\text{start}} \leftarrow 0
    uinit(x)←h(x)u_{\text{init}}(x) \leftarrow h(x)
    while tstart<Tt_{\text{start}} < T do
        tend←tstart+Δtt_{\text{end}} \leftarrow t_{\text{start}} + \Delta t
        Initialize neural network NN(θ,x,t)\text{NN}(\theta, x, t) on domain Ω×[tstart,tend]\Omega \times [t_{\text{start}}, t_{\text{end}}]
        Train NN(θ,x,t)\text{NN}(\theta, x, t) by minimizing data/boundary misfit and PDE residual over [tstart,tend][t_{\text{start}}, t_{\text{end}}], using uinit(x)u_{\text{init}}(x) as the initial condition at t=tstartt = t_{\text{start}}
        uinit(x)←NN(θ,x,tend)u_{\text{init}}(x) \leftarrow \text{NN}(\theta, x, t_{\text{end}})
        tstart←tendt_{\text{start}} \leftarrow t_{\text{end}}
    return u^\hat{u}

    To ensure fair comparison with standard PINNs, the total number of interior collocation points across all segments is kept identical to the total points used in the single space-time model (e.g., Nf=100N_f = 100 per segment for 10 segments of Δt=0.1\Delta t = 0.1, matching 1000 total points for T=[0,1]T = [0, 1]). Marching in time restricts the ill-conditioned operator to a smaller horizon, allowing the network to capture steep gradients and reducing relative error by up to 2 orders of magnitude.

  5. Knowl 5 — Soft-Constrained Physics-Informed Neural Network Formulation

    model/method

    A physics-informed neural network (PINN) approximates the solution u(x,t)u(x, t) to a partial differential equation F(u(x,t))=0\mathcal{F}(u(x, t)) = 0 defined on x∈Ω⊂Rdx \in \Omega \subset \mathbb{R}^d and t∈[0,T]t \in [0, T] by training a neural network u^(x,t)=NN(θ,x,t)\hat{u}(x, t) = \text{NN}(\theta, x, t) with parameters θ\theta. Physical equations and boundary constraints are enforced through soft regularization in an unconstrained empirical objective:

    min⁡θL(θ)=Lu0(θ)+Lub(θ)+λFLF(θ)\min_\theta \mathcal{L}(\theta) = \mathcal{L}_{u_0}(\theta) + \mathcal{L}_{u_b}(\theta) + \lambda_F \mathcal{L}_F(\theta)

    where λF≥0\lambda_F \ge 0 is a regularization hyperparameter and the individual loss terms are defined over sampled collocation points as:

    Lu0(θ)=1Nu∑i=1Nu(u^(xi,0)−u0(xi))2\mathcal{L}_{u_0}(\theta) = \frac{1}{N_u} \sum_{i=1}^{N_u} \left( \hat{u}(x_i, 0) - u_0(x_i) \right)^2

    Lub(θ)=1Nb∑i=1Nb(u^(xileft,ti)−u^(xiright,ti))2\mathcal{L}_{u_b}(\theta) = \frac{1}{N_b} \sum_{i=1}^{N_b} \left( \hat{u}(x_i^{\text{left}}, t_i) - \hat{u}(x_i^{\text{right}}, t_i) \right)^2

    LF(θ)=1Nf∑i=1Nf(F(u^(xi,ti)))2\mathcal{L}_F(\theta) = \frac{1}{N_f} \sum_{i=1}^{N_f} \left( \mathcal{F}(\hat{u}(x_i, t_i)) \right)^2

    where NuN_u is the number of initial condition points, NbN_b is the number of boundary condition points (expressed here for periodic boundary conditions), and NfN_f is the number of residual evaluation points in the interior domain.

  6. Knowl 6 — Empirical Failure Modes of Standard PINNs Across Physical Regimes

    empirical result

    When evaluated on canonical 1D differential equations with analytical solutions, standard PINNs trained via L-BFGS across learning rates from 10−410^{-4} to 2.02.0 accurately learn solutions only in trivial parameter regimes (low coefficients) and fail completely in moderately higher regimes:

    1. 1D Convection (∂u∂t+β∂u∂x=0\frac{\partial u}{\partial t} + \beta \frac{\partial u}{\partial x} = 0, u(x,0)=sin⁡(x)u(x, 0) = \sin(x), periodic boundary conditions on [0,2π)[0, 2\pi)): For β=1\beta = 1, the relative L2L_2 error is 7.84×10−37.84 \times 10^{-3}. For β=20,30,40\beta = 20, 30, 40, the relative L2L_2 error surges to 75.0%75.0\%, 89.7%89.7\%, and 96.1%96.1\% respectively. The network satisfies boundary conditions but fails to propagate the wave past the earliest time steps.

    2. 1D Reaction (∂u∂t−ρu(1−u)=0\frac{\partial u}{\partial t} - \rho u(1 - u) = 0, Gaussian initial condition): For ρ≥5\rho \ge 5, the relative L2L_2 error reaches 97.9%97.9\% to 99.6%99.6\%, with the network predicting an almost homogeneous near-zero solution everywhere.

    3. 1D Reaction-Diffusion (∂u∂t−ν∂2u∂x2−ρu(1−u)=0\frac{\partial u}{\partial t} - \nu \frac{\partial^2 u}{\partial x^2} - \rho u(1 - u) = 0 with ρ=5\rho = 5): For diffusion coefficients ν=2,3,4,5,6\nu = 2, 3, 4, 5, 6, regular PINNs yield relative L2L_2 errors of 50.7%50.7\%, 79.8%79.8\%, 88.4%88.4\%, 93.5%93.5\%, and 96.0%96.0\% respectively, failing to capture both sharp reaction fronts and diffusive transitions.

  7. Knowl 7 — Performance of Curriculum Regularization on 1D Convection

    data/table

    The table compares relative and absolute L2L_2 prediction errors between regular PINN training and curriculum regularization on the 1D linear convection equation ∂u∂t+β∂u∂x=0\frac{\partial u}{\partial t} + \beta \frac{\partial u}{\partial x} = 0 with initial condition u(x,0)=sin⁡(x)u(x, 0) = \sin(x) and periodic boundaries on [0,2π)[0, 2\pi) over t∈[0,1]t \in [0, 1]. Both methods use a 4-layer fully connected network with 50 neurons per layer, tanh activations, and L-BFGS optimization.

    Problem Metric Regular PINN Curriculum Training
    1D convection: β=20\beta = 20 Relative error 7.50×10−17.50 \times 10^{-1} 9.84×10−39.84 \times 10^{-3}
    Absolute error 4.32×10−14.32 \times 10^{-1} 5.42×10−35.42 \times 10^{-3}
    1D convection: β=30\beta = 30 Relative error 8.97×10−18.97 \times 10^{-1} 2.02×10−22.02 \times 10^{-2}
    Absolute error 5.42×10−15.42 \times 10^{-1} 1.10×10−21.10 \times 10^{-2}
    1D convection: β=40\beta = 40 Relative error 9.61×10−19.61 \times 10^{-1} 5.33×10−25.33 \times 10^{-2}
    Absolute error 5.82×10−15.82 \times 10^{-1} 2.69×10−22.69 \times 10^{-2}

    While regular PINN training fails with near 100%100\% relative error at high convection speeds (β=20,30,40\beta = 20, 30, 40), curriculum training reduces relative and absolute error by nearly two orders of magnitude by incrementally warm-starting network weights from lower β\beta values.

  8. Knowl 8 — Performance of Sequence-to-Sequence PINNs on 1D Reaction-Diffusion

    data/table

    The table compares the performance of predicting the entire space-time domain simultaneously versus time-marching sequence-to-sequence (seq2seq) learning with step sizes Δt=0.05\Delta t = 0.05 and Δt=0.1\Delta t = 0.1 for the 1D reaction-diffusion equation ∂u∂t−ν∂2u∂x2−ρu(1−u)=0\frac{\partial u}{\partial t} - \nu \frac{\partial^2 u}{\partial x^2} - \rho u(1 - u) = 0 with fixed reaction coefficient ρ=5\rho = 5 across various diffusion coefficients ν\nu.

    Parameters Metric Entire state space Δt=0.05\Delta t = 0.05 Δt=0.1\Delta t = 0.1
    ν=2,ρ=5\nu = 2, \rho = 5 Relative error 5.07×10−15.07 \times 10^{-1} 2.04×10−22.04 \times 10^{-2} 1.18×10−21.18 \times 10^{-2}
    Absolute error 2.70×10−12.70 \times 10^{-1} 1.06×10−21.06 \times 10^{-2} 6.41×10−36.41 \times 10^{-3}
    ν=3,ρ=5\nu = 3, \rho = 5 Relative error 7.98×10−17.98 \times 10^{-1} 1.92×10−21.92 \times 10^{-2} 1.56×10−21.56 \times 10^{-2}
    Absolute error 4.79×10−14.79 \times 10^{-1} 1.01×10−21.01 \times 10^{-2} 8.17×10−38.17 \times 10^{-3}
    ν=4,ρ=5\nu = 4, \rho = 5 Relative error 8.84×10−18.84 \times 10^{-1} 2.37×10−22.37 \times 10^{-2} 1.59×10−21.59 \times 10^{-2}
    Absolute error 5.74×10−15.74 \times 10^{-1} 1.15×10−21.15 \times 10^{-2} 8.01×10−38.01 \times 10^{-3}
    ν=5,ρ=5\nu = 5, \rho = 5 Relative error 9.35×10−19.35 \times 10^{-1} 2.36×10−22.36 \times 10^{-2} 2.39×10−22.39 \times 10^{-2}
    Absolute error 6.46×10−16.46 \times 10^{-1} 1.09×10−21.09 \times 10^{-2} 1.15×10−21.15 \times 10^{-2}
    ν=6,ρ=5\nu = 6, \rho = 5 Relative error 9.60×10−19.60 \times 10^{-1} 2.81×10−22.81 \times 10^{-2} 2.69×10−22.69 \times 10^{-2}
    Absolute error 6.84×10−16.84 \times 10^{-1} 1.17×10−21.17 \times 10^{-2} 1.28×10−21.28 \times 10^{-2}

    The seq2seq time-marching formulation consistently achieves 1 to 2 orders of magnitude lower relative and absolute errors compared to predicting the full space-time at once, holding the total number of collocation points constant.

  9. Knowl 9 — Performance of Sequence-to-Sequence PINNs on 1D Reaction

    data/table

    The table compares relative and absolute L2L_2 prediction errors between predicting the entire space-time domain simultaneously and sequence-to-sequence (seq2seq) time marching with time step sizes Δt=0.05\Delta t = 0.05 and Δt=0.1\Delta t = 0.1 for the 1D reaction equation ∂u∂t−ρu(1−u)=0\frac{\partial u}{\partial t} - \rho u(1 - u) = 0 across reaction parameters ρ∈{5,6,7,8,9,10}\rho \in \{5, 6, 7, 8, 9, 10\} on t∈[0,1]t \in [0, 1].

    Parameter Metric Entire state space Δt=0.05\Delta t = 0.05 Δt=0.1\Delta t = 0.1
    ρ=5\rho = 5 Relative error 9.79×10−19.79 \times 10^{-1} 7.06×10−27.06 \times 10^{-2} 7.09×10−27.09 \times 10^{-2}
    Absolute error 5.40×10−15.40 \times 10^{-1} 2.52×10−22.52 \times 10^{-2} 2.39×10−22.39 \times 10^{-2}
    ρ=6\rho = 6 Relative error 9.88×10−19.88 \times 10^{-1} 8.25×10−28.25 \times 10^{-2} 7.78×10−27.78 \times 10^{-2}
    Absolute error 5.88×10−15.88 \times 10^{-1} 3.02×10−23.02 \times 10^{-2} 2.65×10−22.65 \times 10^{-2}
    ρ=7\rho = 7 Relative error 9.92×10−19.92 \times 10^{-1} 8.16×10−28.16 \times 10^{-2} 7.56×10−27.56 \times 10^{-2}
    Absolute error 6.31×10−16.31 \times 10^{-1} 3.03×10−23.03 \times 10^{-2} 2.69×10−22.69 \times 10^{-2}
    ρ=8\rho = 8 Relative error 9.94×10−19.94 \times 10^{-1} 8.19×10−28.19 \times 10^{-2} 7.44×10−27.44 \times 10^{-2}
    Absolute error 6.69×10−16.69 \times 10^{-1} 3.10×10−23.10 \times 10^{-2} 2.73×10−22.73 \times 10^{-2}
    ρ=9\rho = 9 Relative error 9.95×10−19.95 \times 10^{-1} 7.02×10−27.02 \times 10^{-2} 8.63×10−28.63 \times 10^{-2}
    Absolute error 7.02×10−17.02 \times 10^{-1} 2.83×10−22.83 \times 10^{-2} 3.21×10−23.21 \times 10^{-2}
    ρ=10\rho = 10 Relative error 9.96×10−19.96 \times 10^{-1} 6.88×10−26.88 \times 10^{-2} 7.47×10−27.47 \times 10^{-2}
    Absolute error 7.31×10−17.31 \times 10^{-1} 2.85×10−22.85 \times 10^{-2} 2.85×10−22.85 \times 10^{-2}

    Standard space-time PINNs fail uniformly across all ρ≥5\rho \ge 5 with >97%>97\% relative error. In contrast, seq2seq learning maintains relative errors under 9%9\% and absolute errors around 0.0250.025--0.0320.032, effectively recovering sharp reaction behavior.

Coverage note — None omitted; all core contributions, theoretical condition number scalings, loss landscape Hessian analyses, mitigation algorithms (curriculum regularization and sequence-to-sequence time marching), and benchmark empirical results are completely covered.

References

  1. 1.Y. Bengio, J. Louradour, R. Collobert, and J. Weston. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48, 2009.
  2. 2.S. L. Brunton, B. R. Noack, and P. Koumoutsakos. Machine learning for fluid mechanics. Annual Review of Fluid Mechanics, 52:477–508, 2020.
  3. 3.Y. Chen, L. Lu, G. E. Karniadakis, and L. Dal Negro. Physics-informed neural networks for inverse problems in nano-optics and metamaterials. Optics express, 28(8):11618–11633, 2020.
  4. 4.M. Cranmer, S. Greydanus, S. Hoyer, P. Battaglia, D. Spergel, and S. Ho. Lagrangian neural networks. arXiv preprint arXiv:2003.04630, 2020.
  5. 5.M. Dissanayake and N. Phan-Thien. Neural-network-based approximations for solving partial differential equations. communications in Numerical Methods in Engineering, 10(3):195–201, 1994.
  6. 6.P. L. Donti, D. Rolnick, and J. Z. Kolter. Dc3: A learning method for optimization with hard constraints. arXiv preprint arXiv:2104.12225, 2021.
  7. 7.V. Dwivedi, N. Parashar, and B. Srinivasan. Distributed learning machines for solving forward and inverse problems in partial differential equations. Neurocomputing, 420:299–316, 2021.
  8. 8.K. Eriksson, D. Estep, P. Hansbo, and C. Johnson. Computational differential equations. Cambridge University Press, 1996.
  9. 9.B. Fornberg. A practical guide to pseudospectral methods. Cambridge university press, 1998.
  10. 10.O. Fuks and H. A. Tchelepi. Limitations of physics informed machine learning for nonlinear two-phase transport in porous media. Journal of Machine Learning for Modeling and Computing, 1 (1), 2020.
  11. 11.N. Geneva and N. Zabaras. Modeling the dynamics of pde systems with physics-constrained deep auto-regressive networks. Journal of Computational Physics, 403:109056, 2020.
  12. 12.S. Greydanus, M. Dzamba, and J. Yosinski. Hamiltonian neural networks. Advances in Neural Information Processing Systems, 32:15379–15389, 2019.
  13. 13.J. Han, A. Jentzen, and E. Weinan. Solving high-dimensional partial differential equations using deep learning. Proceedings of the National Academy of Sciences, 115(34):8505–8510, 2018.
  14. 14.O. Hennigh, S. Narasimhan, M. A. Nabian, A. Subramaniam, K. Tangsali, Z. Fang, M. Rietmann, W. Byeon, and S. Choudhry. Nvidia simnet™: An ai-accelerated multi-physics simulation framework. In International Conference on Computational Science, pages 447–461. Springer, 2021.
  15. 15.W. Ji, W. Qiu, Z. Shi, S. Pan, and S. Deng. Stiff-pinn: Physics-informed neural network for stiff chemical kinetics. arXiv preprint arXiv:2011.04520, 2020.
  16. 16.X. Jin, S. Cai, H. Li, and G. E. Karniadakis. Nsfnets (navier-stokes flow nets): Physics-informed neural networks for the incompressible navier-stokes equations. Journal of Computational Physics, 426:109951, 2021.
  17. 17.G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang. Physics-informed machine learning. Nature Reviews Physics, 3(6):422–440, 2021.
  18. 18.I. E. Lagaris, A. Likas, and D. I. Fotiadis. Artificial neural networks for solving ordinary and partial differential equations. IEEE transactions on neural networks, 9(5):987–1000, 1998.
  19. 19.Z. Long, Y. Lu, X. Ma, and B. Dong. Pde-net: Learning pdes from data. In International Conference on Machine Learning, pages 3208–3216. PMLR, 2018.
  20. 20.L. Lu, X. Meng, Z. Mao, and G. E. Karniadakis. Deepxde: A deep learning library for solving differential equations. SIAM Review, 63(1):208–228, 2021.
  21. 21.L. Lu, R. Pestourie, W. Yao, Z. Wang, F. Verdugo, and S. G. Johnson. Physics-informed neural networks with hard constraints for inverse design. arXiv preprint arXiv:2102.04626, 2021.
  22. 22.P. Márquez-Neila, M. Salzmann, and P. Fua. Imposing hard constraints on deep networks: Promises and limitations. arXiv preprint arXiv:1706.02025, 2017.
  23. 23.P. Moin. Fundamentals of engineering numerical analysis. Cambridge University Press, 2010.
  24. 24.Y. Nandwani, A. Pathak, P. Singla, et al. A primal dual formulation for deep learning with constraints. Advances in Neural Information Processing Systems, 2019.
  25. 25.D. R. Parisi, M. C. Mariani, and M. A. Laborde. Solving differential equations with unsupervised neural networks. Chemical Engineering and Processing: Process Intensification, 42(8-9):715–721, 2003.
  26. 26.C. possible failure modes in physics-informed neural networks. https://github.com/a1k12/characterizing-pinns-failure-modes, 2021.
  27. 27.C. Rackauckas, Y. Ma, J. Martensen, C. Warner, K. Zubov, R. Supekar, D. Skinner, A. Ramadhan, and A. Edelman. Universal differential equations for scientific machine learning. arXiv preprint arXiv:2001.04385, 2020.
  28. 28.M. Raissi, P. Perdikaris, and G. E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019.
  29. 29.M. Raissi, A. Yazdani, and G. E. Karniadakis. Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations. Science, 367(6481):1026–1030, 2020.
  30. 30.F. Sahli Costabal, Y. Yang, P. Perdikaris, D. E. Hurtado, and E. Kuhl. Physics-informed neural networks for cardiac activation mapping. Frontiers in Physics, 8:42, 2020.
  31. 31.J. Sirignano and K. Spiliopoulos. Dgm: A deep learning algorithm for solving partial differential equations. Journal of computational physics, 375:1339–1364, 2018.
  32. 32.B. P. van Milligen, V. Tribaldos, and J. Jiménez. Neural network differential equation and plasma equilibrium solver. Physical review letters, 75(20):3594, 1995.
  33. 33.L. von Rueden, S. Mayer, K. Beckh, B. Georgiev, S. Giesselbach, R. Heese, B. Kirsch, J. Pfrommer, A. Pick, R. Ramamurthy, et al. Informed machine learning–a taxonomy and survey of integrating knowledge into learning systems. arXiv preprint arXiv:1903.12394, 2019.
  34. 34.B. Wang, W. Zhang, and W. Cai. Multi-scale deep neural network (mscalednn) methods for oscillatory stokes flows in complex domains. arXiv preprint arXiv:2009.12729, 2020.
  35. 35.S. Wang, Y. Teng, and P. Perdikaris. Understanding and mitigating gradient pathologies in physics-informed neural networks. arXiv preprint arXiv:2001.04536, 2020.
  36. 36.S. Wang, H. Wang, and P. Perdikaris. On the eigenvector bias of fourier feature networks: From regression to solving multi-scale pdes with physics-informed neural networks. arXiv preprint arXiv:2012.10047, 2020.
  37. 37.S. Wang, X. Yu, and P. Perdikaris. When and why pinns fail to train: A neural tangent kernel perspective. arXiv preprint arXiv:2007.14527, 2020.
  38. 38.E. Weinan, J. Han, and A. Jentzen. Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations. Communications in Mathematics and Statistics, 5(4):349–380, 2017.
  39. 39.J. Willard, X. Jia, S. Xu, M. Steinbach, and V. Kumar. Integrating physics-based modeling with machine learning: A survey. arXiv preprint arXiv:2003.04919, 2020.
  40. 40.K. Xu and E. Darve. Physics constrained learning for data-driven inverse modeling from sparse observations. arXiv preprint arXiv:2002.10521, 2020.
  41. 41.Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Mahoney. Hessian-based analysis of large batch training and robustness to adversaries. Advances in Neural Information Processing Systems, 2018.
  42. 42.Z. Yao, A. Gholami, K. Keutzer, and M. W. Mahoney. Pyhessian: Neural networks through the lens of the hessian. In 2020 IEEE International Conference on Big Data (Big Data), pages 581–590. IEEE, 2020.
  43. 43.Y. Zhu, N. Zabaras, P.-S. Koutsourelakis, and P. Perdikaris. Physics-constrained deep learning for high-dimensional surrogate modeling and uncertainty quantification without labeled data. Journal of Computational Physics, 394:56–81, 2019.
  44. 44.O. C. Zienkiewicz, R. L. Taylor, P. Nithiarasu, and J. Zhu. The finite element method, volume 3. McGraw-hill London, 1977.

Citation

MLA
Krishnapriyan, A., et al. “Characterizing Possible Failure Modes in Physics-informed Neural Networks”. Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 26548–60, https://proceedings.neurips.cc/paper_files/paper/2021/file/df438e5206f31600e6ae4af72f2725f1-Paper.pdf.
APA
Krishnapriyan, A., Gholami, A., Zhe, S., Kirby, R., & Mahoney, M. (2021). Characterizing possible failure modes in physics-informed neural networks. Advances in Neural Information Processing Systems, 34, 26548–26560. https://proceedings.neurips.cc/paper_files/paper/2021/file/df438e5206f31600e6ae4af72f2725f1-Paper.pdf
Chicago
Krishnapriyan, A., A. Gholami, S. Zhe, R. Kirby, and M. Mahoney. 2021. “Characterizing Possible Failure Modes in Physics-informed Neural Networks”. Advances in Neural Information Processing Systems 34: 26548–60. https://proceedings.neurips.cc/paper_files/paper/2021/file/df438e5206f31600e6ae4af72f2725f1-Paper.pdf.
Harvard
Krishnapriyan, A. et al. (2021) “Characterizing possible failure modes in physics-informed neural networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 26548–26560. Available at: https://proceedings.neurips.cc/paper_files/paper/2021/file/df438e5206f31600e6ae4af72f2725f1-Paper.pdf.
Vancouver
1. Krishnapriyan A, Gholami A, Zhe S, Kirby R, Mahoney M (2021) Characterizing possible failure modes in physics-informed neural networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 26548–26560

BibTeX

@inproceedings{krishnapriyan2021characterizing,
  title = {Characterizing possible failure modes in physics-informed neural networks},
  author = {Krishnapriyan, Aditi and Gholami, Amir and Zhe, Shandian and Kirby, Robert and Mahoney, Michael},
  year = {2021},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {34},
  pages = {26548-26560},
  url = {https://proceedings.neurips.cc/paper_files/paper/2021/file/df438e5206f31600e6ae4af72f2725f1-Paper.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors