When and why PINNs fail to train: A neural tangent kernel perspective
Sifan WangXinling YuParis Perdikaris
Explains why physics-informed neural networks fail to train using neural tangent kernel analysis and develops an adaptive weighting method based on kernel eigenvalues to resolve convergence discrepancies across loss terms.
Physics-informed neural networks offer a flexible approach for modeling complex engineering and physical systems described by differential equations. Despite widespread adoption, standard fully-connected networks frequently fail during training, especially when simulating systems with multi-scale behaviors or high-frequency variations. Understanding the root causes of these training failures is critical for building trustworthy, robust machine learning tools for scientific and industrial computing.
This article investigates the training dynamics of physics-informed neural networks through neural tangent kernel theory to determine why standard training fails and how to systematically resolve these failures.
The authors analyze the theoretical behavior of neural networks in the infinite-width limit during gradient descent optimization. They derive the specific mathematical kernel that governs physics-informed networks and evaluate training dynamics across canonical benchmark problems, including a one-dimensional Poisson equation and a one-dimensional wave equation using single- and multi-layer architectures.
The analysis reveals three key findings regarding network behavior. First, as the network becomes very wide, the governing kernel converges to a fixed, deterministic matrix that remains practically constant throughout training. Second, standard models exhibit severe spectral bias, causing gradient descent to learn low-frequency features rapidly while learning high-frequency components extremely slowly. Third, a fundamental imbalance emerges during optimization because the training error rates of the physics residual and the boundary conditions are governed by vastly different kernel eigenvalues. Typically, the physics residual dominates, causing the model to satisfy the internal differential equation while failing to fit the boundary and initial conditions.
These findings explain why standard physics-informed networks often yield inaccurate results or fail completely: the optimization process prioritizes one loss component at the expense of others. To correct this pathology, the authors introduce an adaptive training algorithm that uses the trace of the kernel matrices to balance the convergence rates across all loss terms. In numerical benchmarks, this adaptive weighting scheme improved predictive accuracy by approximately two orders of magnitude, reducing relative error from over 40% down to less than 0.2% on a challenging wave equation problem without requiring costly trial-and-error hyperparameter tuning.
Organizations developing machine learning for physical modeling should incorporate eigenvalue- or trace-based adaptive weighting schemes into their training pipelines to stabilize convergence and eliminate manual weight calibration. While the formal mathematical proofs are primarily established for single-layer networks solving linear problems under infinitesimal learning rates, numerical experiments demonstrate high empirical stability across deeper architectures and modern optimizers. Future work should focus on extending formal theoretical guarantees to nonlinear equations, advanced network architectures, and inverse problem settings.
- Paper: Neural Tangent Kernel: Convergence and Generalization in Neural Networks, Arthur Jacot et al. (2018). This foundational work introduces the Neural Tangent Kernel framework and establishes how infinite-width neural network training dynamics under gradient descent reduce to linear differential equations governed by a deterministic kernel.
- Paper: DeepXDE: A Deep Learning Library for Solving Differential Equations, Lu Lu et al. (2019). This paper establishes the core physics-informed neural network formulation and software implementation for enforcing partial differential equations and boundary conditions via composite loss functions.
- Paper: DGM: A deep learning algorithm for solving partial differential equations, Justin Sirignano et al. (2017). This work introduces deep learning methods for high-dimensional partial differential equations trained on batches of random collocation points, serving as essential context for understanding neural PDE training mechanics.
- Paper: The Deep Ritz Method: A Deep Learning-Based Numerical Algorithm for Solving Variational Problems, Weinan E et al. (2017). This paper presents variational deep learning formulations for differential equations, providing foundational background on optimization challenges in neural network-based PDE solvers.
- Paper: Scientific Machine Learning Through Physics–Informed Neural Networks: Where we are and What’s Next, Salvatore Cuomo et al. (2022). This comprehensive review synthesizes developments across physics-informed neural networks, surveying training failure modes and modern optimization algorithms like NTK-based weighting.
- Paper: Physics-informed neural networks (PINNs) for fluid mechanics: a review, Shengze Cai et al. (2021). This survey examines physics-informed neural networks applied specifically to fluid dynamics, discussing practical training difficulties and algorithmic enhancements for complex Navier-Stokes problems.
- Paper: Fourier Neural Operator for Parametric Partial Differential Equations, Zongyi Li et al. (2020). This paper introduces Fourier Neural Operators as an alternative operator-learning paradigm that circumvents point-wise PINN training dynamics by mapping directly between function spaces.
- Paper: Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Lu Lu et al. (2021). This work formulates DeepONet for learning general nonlinear operators, advancing beyond single-instance PINN training by building parameterized operator surrogates.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). This text explores Kolmogorov-Arnold Networks as an architectural alternative to multilayer perceptrons for scientific machine learning and PDE solving, addressing traditional MLP optimization barriers.
- Paper: Neural means and kernel corrections for operator learning, Yitzchak Shmalo (2026). This book bridges neural scientific surrogates with exact kernel methods, expanding on the interplay between deep networks and kernel dynamics in scientific machine learning.
