When and why PINNs fail to train: A neural tangent kernel perspective

Sifan WangXinling YuParis Perdikaris

article2020Journal of Computational Physics1,794 citations

Explains why physics-informed neural networks fail to train using neural tangent kernel analysis and develops an adaptive weighting method based on kernel eigenvalues to resolve convergence discrepancies across loss terms.

Listen

Physics-informed neural networks offer a flexible approach for modeling complex engineering and physical systems described by differential equations. Despite widespread adoption, standard fully-connected networks frequently fail during training, especially when simulating systems with multi-scale behaviors or high-frequency variations. Understanding the root causes of these training failures is critical for building trustworthy, robust machine learning tools for scientific and industrial computing.

This article investigates the training dynamics of physics-informed neural networks through neural tangent kernel theory to determine why standard training fails and how to systematically resolve these failures.

The authors analyze the theoretical behavior of neural networks in the infinite-width limit during gradient descent optimization. They derive the specific mathematical kernel that governs physics-informed networks and evaluate training dynamics across canonical benchmark problems, including a one-dimensional Poisson equation and a one-dimensional wave equation using single- and multi-layer architectures.

The analysis reveals three key findings regarding network behavior. First, as the network becomes very wide, the governing kernel converges to a fixed, deterministic matrix that remains practically constant throughout training. Second, standard models exhibit severe spectral bias, causing gradient descent to learn low-frequency features rapidly while learning high-frequency components extremely slowly. Third, a fundamental imbalance emerges during optimization because the training error rates of the physics residual and the boundary conditions are governed by vastly different kernel eigenvalues. Typically, the physics residual dominates, causing the model to satisfy the internal differential equation while failing to fit the boundary and initial conditions.

These findings explain why standard physics-informed networks often yield inaccurate results or fail completely: the optimization process prioritizes one loss component at the expense of others. To correct this pathology, the authors introduce an adaptive training algorithm that uses the trace of the kernel matrices to balance the convergence rates across all loss terms. In numerical benchmarks, this adaptive weighting scheme improved predictive accuracy by approximately two orders of magnitude, reducing relative error from over 40% down to less than 0.2% on a challenging wave equation problem without requiring costly trial-and-error hyperparameter tuning.

Organizations developing machine learning for physical modeling should incorporate eigenvalue- or trace-based adaptive weighting schemes into their training pipelines to stabilize convergence and eliminate manual weight calibration. While the formal mathematical proofs are primarily established for single-layer networks solving linear problems under infinitesimal learning rates, numerical experiments demonstrate high empirical stability across deeper architectures and modern optimizers. Future work should focus on extending formal theoretical guarantees to nonlinear equations, advanced network architectures, and inverse problem settings.

Cover for When and why PINNs fail to train: A neural tangent kernel perspective

Abstract

Physics-informed neural networks (PINNs) have lately received great attention thanks to their flexibility in tackling a wide range of forward and inverse problems involving partial differential equations. However, despite their noticeable empirical success, little is known about how such constrained neural networks behave during their training via gradient descent. More importantly, even less is known about why such models sometimes fail to train at all. In this work, we aim to investigate these questions through the lens of the Neural Tangent Kernel (NTK); a kernel that captures the behavior of fully-connected neural networks in the infinite width limit during training via gradient descent. Specifically, we derive the NTK of PINNs and prove that, under appropriate conditions, it converges to a deterministic kernel that stays constant during training in the infinite-width limit. This allows us to analyze the training dynamics of PINNs through the lens of their limiting NTK and find a remarkable discrepancy in the convergence rate of the different loss components contributing to the total training error. To address this fundamental pathology, we propose a novel gradient descent algorithm that utilizes the eigenvalues of the NTK to adaptively calibrate the convergence rate of the total training error. Finally, we perform a series of numerical experiments to verify the correctness of our theory and the practical effectiveness of the proposed algorithms. The data and code accompanying this manuscript are publicly available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Infinitely Wide Neural Networks
  • 3 Physics-informed Neural Networks (PINNs)
  • 3.1 Neural tangent kernel theory for PINNs
  • 4 Analyzing the training dynamics of PINNs through the lens of their NTK
  • 5 Spectral bias in physics-informed neural networks
  • 6 Practical insights
  • 7 Numerical Experiments
  • 7.1 Convergence of the NTK of PINNs
  • 7.2 Adaptive training for PINNs
  • 7.3 One-dimensional wave equation
  • 8 Discussion
  • References
  • A Proof of Lemma
  • B Proof of Theorem
  • C Proof of Theorem
  • D Proof of Theorem

Knowls

  1. Knowl 1 — Adaptive NTK-Based Weighting Algorithm for PINNs

    algorithm

    To alleviate the convergence discrepancy between different loss components in physics-informed neural networks (PINNs), the loss term weights are adaptively calibrated using the traces of their corresponding Neural Tangent Kernel (NTK) sub-matrices.

    Consider a PDE L[u](x)=f(x)\mathcal{L}[u](x) = f(x) for x∈Ωx \in \Omega with boundary condition u(x)=g(x)u(x) = g(x) for x∈∂Ωx \in \partial \Omega, approximated by a neural network u(x,θ)u(x, \theta) with parameter vector θ\theta. The composite loss function is: L(θ)=λbLb(θ)+λrLr(θ)\mathcal{L}(\theta) = \lambda_b \mathcal{L}_b(\theta) + \lambda_r \mathcal{L}_r(\theta) where Lb(θ)=12Nb∑i=1Nb∣u(xbi,θ)−g(xbi)∣2\mathcal{L}_b(\theta) = \frac{1}{2 N_b} \sum_{i=1}^{N_b} |u(x_b^i, \theta) - g(x_b^i)|^2 and Lr(θ)=12Nr∑i=1Nr∣Lu(xri,θ)−f(xri)∣2\mathcal{L}_r(\theta) = \frac{1}{2 N_r} \sum_{i=1}^{N_r} |\mathcal{L}u(x_r^i, \theta) - f(x_r^i)|^2.

    Input: Boundary training points {xbi,g(xbi)}i=1Nb\{x_b^i, g(x_b^i)\}_{i=1}^{N_b}, residual points {xri,f(xri)}i=1Nr\{x_r^i, f(x_r^i)\}_{i=1}^{N_r}, learning rate η\eta, total steps SS, update frequency FF.
    Output: Optimized network parameters θS\theta_S.
    Initialize network parameters θ0\theta_0, set λb=1\lambda_b = 1, λr=1\lambda_r = 1.
    for n=0,1,…,S−1n = 0, 1, \dots, S - 1 do
        if n(modF)==0n \pmod F == 0 then
            Compute Jacobian matrices Ju(n)=∇θu(xb,θn)J_u(n) = \nabla_\theta u(x_b, \theta_n) and Jr(n)=∇θLu(xr,θn)J_r(n) = \nabla_\theta \mathcal{L}u(x_r, \theta_n).
            Compute NTK sub-matrices Kuu(n)=Ju(n)Ju(n)TK_{uu}(n) = J_u(n) J_u(n)^T and Krr(n)=Jr(n)Jr(n)TK_{rr}(n) = J_r(n) J_r(n)^T.
            Compute total NTK trace Tr(K(n))=Tr(Kuu(n))+Tr(Krr(n))\text{Tr}(K(n)) = \text{Tr}(K_{uu}(n)) + \text{Tr}(K_{rr}(n)).
            Update weights:
            λb=Tr(K(n))Tr(Kuu(n))\lambda_b = \frac{\text{Tr}(K(n))}{\text{Tr}(K_{uu}(n))}
            λr=Tr(K(n))Tr(Krr(n))\lambda_r = \frac{\text{Tr}(K(n))}{\text{Tr}(K_{rr}(n))}
        end if
        Compute loss gradient ∇θL(θn)=λb∇θLb(θn)+λr∇θLr(θn)\nabla_\theta \mathcal{L}(\theta_n) = \lambda_b \nabla_\theta \mathcal{L}_b(\theta_n) + \lambda_r \nabla_\theta \mathcal{L}_r(\theta_n).
        Update parameters: θn+1=θn−η∇θL(θn)\theta_{n+1} = \theta_n - \eta \nabla_\theta \mathcal{L}(\theta_n).
    end for
    return θS\theta_S

    For problems with MM separate loss components {Lm}m=1M\{\mathcal{L}_m\}_{m=1}^M (such as distinct initial and boundary conditions), the update generalizes to λk=∑m=1MTr(Km(n))Tr(Kk(n))\lambda_k = \frac{\sum_{m=1}^M \text{Tr}(K_m(n))}{\text{Tr}(K_k(n))} for each component k∈{1,…,M}k \in \{1, \dots, M\}.

  2. Knowl 2 — Continuous Gradient Flow Dynamics and Neural Tangent Kernel of PINNs

    theoretical result

    Consider a partial differential equation L[u](x)=f(x)\mathcal{L}[u](x) = f(x) on Ω⊂Rd\Omega \subset \mathbb{R}^d with boundary condition u(x)=g(x)u(x) = g(x) on ∂Ω\partial \Omega, parameterized by a neural network u(x,θ)u(x, \theta) with parameter vector θ\theta. Given boundary data {xbi,g(xbi)}i=1Nb\{x_b^i, g(x_b^i)\}_{i=1}^{N_b} and collocation points {xri,f(xri)}i=1Nr\{x_r^i, f(x_r^i)\}_{i=1}^{N_r}, the composite loss is: L(θ)=12∑i=1Nb∣u(xbi,θ)−g(xbi)∣2+12∑i=1Nr∣Lu(xri,θ)−f(xri)∣2\mathcal{L}(\theta) = \frac{1}{2}\sum_{i=1}^{N_b} |u(x_b^i, \theta) - g(x_b^i)|^2 + \frac{1}{2}\sum_{i=1}^{N_r} |\mathcal{L}u(x_r^i, \theta) - f(x_r^i)|^2 Under continuous-time gradient descent flow dθdt=−∇θL(θ)\frac{d\theta}{dt} = -\nabla_\theta \mathcal{L}(\theta), the time evolution of the network outputs u(xb,θ(t))u(x_b, \theta(t)) and Lu(xr,θ(t))\mathcal{L}u(x_r, \theta(t)) satisfies: [du(xb,θ(t))dtdLu(xr,θ(t))dt]=−K(t)[u(xb,θ(t))−g(xb)Lu(xr,θ(t))−f(xr)]\begin{bmatrix} \frac{du(x_b, \theta(t))}{dt} \\[6pt] \frac{d\mathcal{L}u(x_r, \theta(t))}{dt} \end{bmatrix} = -\mathbf{K}(t) \begin{bmatrix} u(x_b, \theta(t)) - g(x_b) \\[6pt] \mathcal{L}u(x_r, \theta(t)) - f(x_r) \end{bmatrix} where the PINN Neural Tangent Kernel K(t)\mathbf{K}(t) is defined blockwise by: K(t)=[Kuu(t)Kur(t)Kru(t)Krr(t)]=[Ju(t)Jr(t)][Ju(t)TJr(t)T]\mathbf{K}(t) = \begin{bmatrix} \mathbf{K}_{uu}(t) & \mathbf{K}_{ur}(t) \\[4pt] \mathbf{K}_{ru}(t) & \mathbf{K}_{rr}(t) \end{bmatrix} = \begin{bmatrix} J_u(t) \\[4pt] J_r(t) \end{bmatrix} \begin{bmatrix} J_u(t)^T & J_r(t)^T \end{bmatrix} with entries given by the parameter inner products: (Kuu)ij(t)=⟨∂u(xbi,θ(t))∂θ,∂u(xbj,θ(t))∂θ⟩(\mathbf{K}_{uu})_{ij}(t) = \left\langle \frac{\partial u(x_b^i, \theta(t))}{\partial \theta}, \frac{\partial u(x_b^j, \theta(t))}{\partial \theta} \right\rangle (Kur)ij(t)=⟨∂u(xbi,θ(t))∂θ,∂Lu(xrj,θ(t))∂θ⟩(\mathbf{K}_{ur})_{ij}(t) = \left\langle \frac{\partial u(x_b^i, \theta(t))}{\partial \theta}, \frac{\partial \mathcal{L}u(x_r^j, \theta(t))}{\partial \theta} \right\rangle (Krr)ij(t)=⟨∂Lu(xri,θ(t))∂θ,∂Lu(xrj,θ(t))∂θ⟩(\mathbf{K}_{rr})_{ij}(t) = \left\langle \frac{\partial \mathcal{L}u(x_r^i, \theta(t))}{\partial \theta}, \frac{\partial \mathcal{L}u(x_r^j, \theta(t))}{\partial \theta} \right\rangle and Kru(t)=Kur(t)T\mathbf{K}_{ru}(t) = \mathbf{K}_{ur}(t)^T. The matrix K(t)\mathbf{K}(t) is symmetric and positive semi-definite.

  3. Knowl 3 — Constancy of PINN Neural Tangent Kernel During Training

    theoretical result

    For a one-dimensional Poisson problem uxx(x)=f(x)u_{xx}(x) = f(x) on Ω\Omega with u(x)=g(x)u(x) = g(x) on ∂Ω\partial \Omega, approximated by a fully-connected single-hidden-layer neural network u(x,θ)=1NW(1)σ(W(0)x+b(0))+b(1)u(x, \theta) = \frac{1}{\sqrt{N}} W^{(1)} \sigma(W^{(0)} x + b^{(0)}) + b^{(1)} trained via continuous gradient descent flow with infinitesimal learning rate, suppose for any fixed time horizon T>0T > 0:

    1. Parameters remain uniformly bounded: sup⁡t∈[0,T]∥θ(t)∥∞≤C\sup_{t \in [0, T]} \|\theta(t)\|_\infty \le C for a constant CC independent of width NN.
    2. Cumulative training errors remain bounded: ∫0T∣∑i=1Nb(u(xbi,θ(τ))−g(xbi))∣dτ≤C,∫0T∣∑i=1Nr(uxx(xri,θ(τ))−f(xri))∣dτ≤C\int_0^T \left| \sum_{i=1}^{N_b} (u(x_b^i, \theta(\tau)) - g(x_b^i)) \right| d\tau \le C, \quad \int_0^T \left| \sum_{i=1}^{N_r} (u_{xx}(x_r^i, \theta(\tau)) - f(x_r^i)) \right| d\tau \le C
    3. The activation function σ\sigma is smooth and satisfies ∣σ(k)∣≤C|\sigma^{(k)}| \le C for all derivatives k∈{0,1,2,3,4}k \in \{0, 1, 2, 3, 4\}.

    Then, in the infinite-width limit, the PINN Neural Tangent Kernel K(t)\mathbf{K}(t) remains constant over the entire training trajectory: lim⁡N→∞sup⁡t∈[0,T]∥K(t)−K(0)∥2=0\lim_{N \to \infty} \sup_{t \in [0, T]} \|\mathbf{K}(t) - \mathbf{K}(0)\|_2 = 0

  4. Knowl 4 — Deterministic Limiting Neural Tangent Kernel of PINNs at Initialization

    theoretical result

    For a single-hidden-layer fully-connected network u(x,θ)=1NW(1)σ(W(0)x+b(0))+b(1)u(x, \theta) = \frac{1}{\sqrt{N}} W^{(1)} \sigma(W^{(0)} x + b^{(0)}) + b^{(1)} with independent standard normal parameter initialization N(0,1)\mathcal{N}(0,1), the PINN Neural Tangent Kernel at initialization K(0)\mathbf{K}(0) converges in probability as the width N→∞N \to \infty to a deterministic limiting kernel matrix: K(0)=[Kuu(0)Kur(0)Kru(0)Krr(0)]→P[Θuu(1)Θur(1)Θru(1)Θrr(1)]=:K∗\mathbf{K}(0) = \begin{bmatrix} \mathbf{K}_{uu}(0) & \mathbf{K}_{ur}(0) \\[4pt] \mathbf{K}_{ru}(0) & \mathbf{K}_{rr}(0) \end{bmatrix} \xrightarrow{\mathcal{P}} \begin{bmatrix} \Theta_{uu}^{(1)} & \Theta_{ur}^{(1)} \\[4pt] \Theta_{ru}^{(1)} & \Theta_{rr}^{(1)} \end{bmatrix} =: \mathbf{K}^* where for inputs x,x′x, x': Θuu(1)(x,x′)=Σ˙(1)(x,x′)(xx′)+Σ(1)(x,x′)+Σ˙(1)(x,x′)+1\Theta_{uu}^{(1)}(x, x') = \dot{\Sigma}^{(1)}(x, x')(x x') + \Sigma^{(1)}(x, x') + \dot{\Sigma}^{(1)}(x, x') + 1 Θrr(1)(x,x′)=Arr(x,x′)+Brr(x,x′)+Crr(x,x′)\Theta_{rr}^{(1)}(x, x') = A_{rr}(x, x') + B_{rr}(x, x') + C_{rr}(x, x') Θur(1)(x,x′)=Aur(x,x′)+Bur(x,x′)+Cur(x,x′)\Theta_{ur}^{(1)}(x, x') = A_{ur}(x, x') + B_{ur}(x, x') + C_{ur}(x, x') with covariance Σ(1)(x,x′)=E(u,v)∼N(0,Λ(1))[σ(u)σ(v)]+1\Sigma^{(1)}(x, x') = \mathbb{E}_{(u,v) \sim \mathcal{N}(0, \Lambda^{(1)})}[\sigma(u)\sigma(v)] + 1, derivative covariance Σ˙(1)(x,x′)=E(u,v)∼N(0,Λ(1))[σ˙(u)σ˙(v)]\dot{\Sigma}^{(1)}(x, x') = \mathbb{E}_{(u,v) \sim \mathcal{N}(0, \Lambda^{(1)})}[\dot{\sigma}(u)\dot{\sigma}(v)], and components defined via expectations over standard Gaussians Wk(0)∼N(0,1)W_k^{(0)} \sim \mathcal{N}(0,1): Arr=E[(Wk(0))4σ...(Wk(0)x+bk(0))σ...(Wk(0)x′+bk(0))]xx′+2E[(Wk(0))3σ...(Wk(0)x+bk(0))σ¨(Wk(0)x′+bk(0))]x+2E[(Wk(0))3σ...(Wk(0)x′+bk(0))σ¨(Wk(0)x+bk(0))]x′+4E[(Wk(0))2σ¨(Wk(0)x+bk(0))σ¨(Wk(0)x′+bk(0))]A_{rr} = \mathbb{E}[(W_k^{(0)})^4 \dddot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \dddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] x x' + 2\mathbb{E}[(W_k^{(0)})^3 \dddot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] x + 2\mathbb{E}[(W_k^{(0)})^3 \dddot{\sigma}(W_k^{(0)}x'+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x+b_k^{(0)})] x' + 4\mathbb{E}[(W_k^{(0)})^2 \ddot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] Brr=E[(Wk(0))4σ¨(Wk(0)x+bk(0))σ¨(Wk(0)x′+bk(0))]B_{rr} = \mathbb{E}[(W_k^{(0)})^4 \ddot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] Crr=E[(Wk(0))4σ...(Wk(0)x+bk(0))σ...(Wk(0)x′+bk(0))]C_{rr} = \mathbb{E}[(W_k^{(0)})^4 \dddot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \dddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] Aur=E[(Wk(0))2σ˙(Wk(0)x+bk(0))σ...(Wk(0)x′+bk(0))]xx′+2E[Wk(0)σ˙(Wk(0)x+bk(0))σ¨(Wk(0)x′+bk(0))]xA_{ur} = \mathbb{E}[(W_k^{(0)})^2 \dot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \dddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] x x' + 2\mathbb{E}[W_k^{(0)} \dot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] x Bur=E[(Wk(0))2σ(Wk(0)x+bk(0))σ¨(Wk(0)x′+bk(0))]B_{ur} = \mathbb{E}[(W_k^{(0)})^2 \sigma(W_k^{(0)}x+b_k^{(0)}) \ddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})] Cur=E[(Wk(0))2σ˙(Wk(0)x+bk(0))σ...(Wk(0)x′+bk(0))]C_{ur} = \mathbb{E}[(W_k^{(0)})^2 \dot{\sigma}(W_k^{(0)}x+b_k^{(0)}) \dddot{\sigma}(W_k^{(0)}x'+b_k^{(0)})]

  5. Knowl 5 — Gaussian Process Correspondence for Infinitely Wide PINNs on Linear PDEs

    theoretical result

    For a single-hidden-layer fully-connected network u(x,θ)=1NW(1)σ(W(0)x+b(0))+b(1)u(x, \theta) = \frac{1}{\sqrt{N}} W^{(1)} \sigma(W^{(0)} x + b^{(0)}) + b^{(1)} with independent standard normal parameter initialization N(0,1)\mathcal{N}(0, 1) and a smooth activation function σ\sigma with bounded second derivative σ¨\ddot{\sigma}, the network output u(x,θ)u(x, \theta) and its second derivative uxx(x,θ)u_{xx}(x, \theta) asymptotically converge in distribution as width N→∞N \to \infty to centered Gaussian processes: u(x,θ)→DGP(0,Σ(1)(x,x′))u(x, \theta) \xrightarrow{\mathcal{D}} \mathcal{GP}\left(0, \Sigma^{(1)}(x, x')\right) uxx(x,θ)→DGP(0,Σxx(1)(x,x′))u_{xx}(x, \theta) \xrightarrow{\mathcal{D}} \mathcal{GP}\left(0, \Sigma_{xx}^{(1)}(x, x')\right) where: Σ(1)(x,x′)=E(u,v)∼N(0,Λ(1))[σ(u)σ(v)]+1,Λ(1)(x,x′)=[xTx+1xTx′+1x′Tx+1x′Tx′+1]\Sigma^{(1)}(x, x') = \mathbb{E}_{(u,v) \sim \mathcal{N}(0, \Lambda^{(1)})}[\sigma(u)\sigma(v)] + 1, \quad \Lambda^{(1)}(x, x') = \begin{bmatrix} x^T x + 1 & x^T x' + 1 \\[4pt] x'^T x + 1 & x'^T x' + 1 \end{bmatrix} Σxx(1)(x,x′)=Eu,v∼N(0,1)[u4σ¨(ux+v)σ¨(ux′+v)]\Sigma_{xx}^{(1)}(x, x') = \mathbb{E}_{u,v \sim \mathcal{N}(0,1)}\left[ u^4 \ddot{\sigma}(ux + v) \ddot{\sigma}(ux' + v) \right] Because linear combinations and differential operations on Gaussian processes preserve Gaussianity, this PINN-GP correspondence generalizes to any linear partial differential operator L\mathcal{L} under suitable regularity conditions.

  6. Knowl 6 — Convergence Rate Discrepancy and Spectral Bias in PINNs

    theoretical result

    Given the positive semi-definite PINN NTK K(0)=QTΛQ\mathbf{K}(0) = \mathbf{Q}^T \mathbf{\Lambda} \mathbf{Q} with orthonormal eigenvectors Q\mathbf{Q} and non-negative eigenvalues Λ=diag(λ1,…,λn)\mathbf{\Lambda} = \text{diag}(\lambda_1, \dots, \lambda_n), the training error under gradient flow evolves as: Q([u(xb,θ(t))uxx(xr,θ(t))]−[g(xb)f(xr)])≈−e−ΛtQ[g(xb)f(xr)]\mathbf{Q} \left( \begin{bmatrix} u(x_b, \theta(t)) \\[4pt] u_{xx}(x_r, \theta(t)) \end{bmatrix} - \begin{bmatrix} g(x_b) \\[4pt] f(x_r) \end{bmatrix} \right) \approx -e^{-\mathbf{\Lambda} t} \mathbf{Q} \begin{bmatrix} g(x_b) \\[4pt] f(x_r) \end{bmatrix} Each error component along eigenvector ii decays at exponential rate e−λite^{-\lambda_i t}. Higher eigenvalues correspond to lower-frequency components, causing rapid decay for low frequencies and severe slowing for high frequencies (spectral bias).

    The average convergence rate of an n×nn \times n kernel matrix K\mathbf{K} is defined as: c=∑i=1nλin=Tr(K)nc = \frac{\sum_{i=1}^n \lambda_i}{n} = \frac{\text{Tr}(\mathbf{K})}{n} Kernel K1\mathbf{K}_1 dominates K2\mathbf{K}_2 if c1≫c2c_1 \gg c_2. In standard PINNs, the residual kernel Krr\mathbf{K}_{rr} typically dominates the boundary kernel Kuu\mathbf{K}_{uu} (crr≫cuuc_{rr} \gg c_{uu}). As a consequence, the PDE residual loss converges significantly faster than the boundary condition loss, preventing the network from learning the correct boundary-constrained PDE solution.

  7. Knowl 7 — Gradient Flow Dynamics and Forward Euler Stability Limit for Weighted PINN Losses

    theoretical result

    For a weighted PINN loss function with component weights λb\lambda_b and λr\lambda_r: L(θ)=λb2Nb∑i=1Nb∣u(xbi,θ)−g(xbi)∣2+λr2Nr∑i=1Nr∣Lu(xri,θ)−f(xri)∣2\mathcal{L}(\theta) = \frac{\lambda_b}{2 N_b} \sum_{i=1}^{N_b} |u(x_b^i, \theta) - g(x_b^i)|^2 + \frac{\lambda_r}{2 N_r} \sum_{i=1}^{N_r} |\mathcal{L}u(x_r^i, \theta) - f(x_r^i)|^2 the continuous gradient descent dynamics of the predictions follow: [du(xb,θ(t))dtdLu(xr,θ(t))dt]=−K~(t)[u(xb,θ(t))−g(xb)Lu(xr,θ(t))−f(xr)]\begin{bmatrix} \frac{du(x_b, \theta(t))}{dt} \\[6pt] \frac{d\mathcal{L}u(x_r, \theta(t))}{dt} \end{bmatrix} = -\tilde{\mathbf{K}}(t) \begin{bmatrix} u(x_b, \theta(t)) - g(x_b) \\[6pt] \mathcal{L}u(x_r, \theta(t)) - f(x_r) \end{bmatrix} where the effective weighted NTK is: K~(t)=[λbNbKuu(t)λrNrKur(t)λbNbKru(t)λrNrKrr(t)]\tilde{\mathbf{K}}(t) = \begin{bmatrix} \frac{\lambda_b}{N_b}\mathbf{K}_{uu}(t) & \frac{\lambda_r}{N_r}\mathbf{K}_{ur}(t) \\[6pt] \frac{\lambda_b}{N_b}\mathbf{K}_{ru}(t) & \frac{\lambda_r}{N_r}\mathbf{K}_{rr}(t) \end{bmatrix} When discretized with forward Euler (gradient descent with learning rate η\eta), numerical stability requires: η≤2λmax⁡(K~(t))\eta \le \frac{2}{\lambda_{\max}(\tilde{\mathbf{K}}(t))} In terms of convergence rate effects, scaling the weight λb\lambda_b or λr\lambda_r is mathematically equivalent to inverse-scaling the corresponding batch size NbN_b or NrN_r.

  8. Knowl 8 — Empirical Validation of Adaptive Weighting on 1D Poisson Equation

    empirical result

    The adaptive NTK weighting method was evaluated on a 1D Poisson boundary value problem: uxx(x)=−16π2sin⁡(4πx),x∈[0,1],u(0)=u(1)=0u_{xx}(x) = -16\pi^2 \sin(4\pi x), \quad x \in [0, 1], \quad u(0) = u(1) = 0 with exact solution u(x)=sin⁡(4πx)u(x) = \sin(4\pi x). A fully-connected neural network with 1 hidden layer of width 100 and tanh⁡\tanh activations was trained using full-batch gradient descent (batch sizes Nb=Nr=100N_b = N_r = 100) with learning rate 10−510^{-5} for 40,000 iterations.

    • Standard PINN (unweighted baseline with λb=λr=1\lambda_b = \lambda_r = 1) achieved a relative L2L^2 error of 2.40×10−12.40 \times 10^{-1}.
    • NTK-calibrated PINN (fixing λr=1\lambda_r = 1 and setting λb=Tr(Krr(0))Tr(Kuu(0))≈100\lambda_b = \frac{\text{Tr}(\mathbf{K}_{rr}(0))}{\text{Tr}(\mathbf{K}_{uu}(0))} \approx 100) achieved a relative L2L^2 error of 1.63×10−31.63 \times 10^{-3}, improving predictive accuracy by over two orders of magnitude.
    • Manual hyper-parameter sweeps across λb∈[1,500]\lambda_b \in [1, 500] confirmed that the relative L2L^2 error reaches its global minimum near λb≈100\lambda_b \approx 100, verifying that the NTK trace ratio directly identifies the optimal weighting.
  9. Knowl 9 — Resolution of Training Failure on 1D Wave Equation via Adaptive NTK Weighting

    empirical result

    The adaptive NTK weighting method was tested on a 1D wave equation exhibiting high stiffness: utt(x,t)−4uxx(x,t)=0,(x,t)∈(0,1)×(0,1)u_{tt}(x, t) - 4u_{xx}(x, t) = 0, \quad (x,t) \in (0, 1) \times (0, 1) with boundary conditions u(0,t)=u(1,t)=0u(0, t) = u(1, t) = 0, initial position u(x,0)=sin⁡(πx)+0.5sin⁡(4πx)u(x, 0) = \sin(\pi x) + 0.5\sin(4\pi x), and initial velocity ut(x,0)=0u_t(x, 0) = 0. The exact analytical solution is u(x,t)=sin⁡(πx)cos⁡(2πt)+0.5sin⁡(4πx)cos⁡(8πt)u(x, t) = \sin(\pi x)\cos(2\pi t) + 0.5\sin(4\pi x)\cos(8\pi t).

    A 5-layer fully-connected network with 500 neurons per layer was trained using the Adam optimizer for 80,000 steps with mini-batch sizes Nu=Nut=Nr=300N_u = N_{u_t} = N_r = 300.

    • Standard PINN (equal loss weights λu=λut=λr=1\lambda_u = \lambda_{u_t} = \lambda_r = 1) completely failed to learn the wave solution, yielding a relative L2L^2 error of 4.518×10−14.518 \times 10^{-1} (over 45%45\% error), caused by Kr\mathbf{K}_r and Kut\mathbf{K}_{u_t} dominating Ku\mathbf{K}_u.
    • Adaptive NTK-weighted PINN (updating weights λu,λut,λr\lambda_u, \lambda_{u_t}, \lambda_r every 1,000 steps via trace ratios λk=Tr(Ku)+Tr(Kut)+Tr(Kr)Tr(Kk)\lambda_k = \frac{\text{Tr}(\mathbf{K}_u) + \text{Tr}(\mathbf{K}_{u_t}) + \text{Tr}(\mathbf{K}_r)}{\text{Tr}(\mathbf{K}_k)}) resolved the eigenvalue discrepancy and achieved a relative L2L^2 error of 1.728×10−31.728 \times 10^{-3}.
  10. Knowl 10 — Methodological and Theoretical Limitations of NTK Analysis for PINNs

    limitation

    The NTK formulation for PINNs possesses three primary limitations:

    1. Loss weighting adjusts relative convergence rates across tasks but does not alter the underlying eigenvalue spectrum within individual kernel blocks; consequently, adaptive loss weighting cannot resolve the fundamental spectral bias of neural networks against high frequencies.
    2. The weighted continuous dynamics matrix K~\tilde{\mathbf{K}} is generally asymmetric and indefinite, which can introduce complex eigenvalues and lead to numerical instability during training if weights λb\lambda_b or λr\lambda_r become excessively large.
    3. The theoretical proofs of NTK constancy during training and Gaussian process convergence are rigorously derived only for single-hidden-layer networks applied to linear Poisson equations under continuous gradient flow. Extensions to deep multi-layer architectures, non-linear partial differential equations, and momentum-based optimizers (such as Adam) remain empirical observations and conjectures.

Coverage note — None omitted; all primary theoretical derivations, theorems, algorithms, empirical validations, and limitations contributed by the paper have been extracted as self-contained knowls.

References

  1. 1.Maziar Raissi, Alireza Yazdani, and George Em Karniadakis. Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations. Science, 367(6481):1026–1030, 2020.
  2. 2.Luning Sun, Han Gao, Shaowu Pan, and Jian-Xun Wang. Surrogate modeling for fluid flows based on physics-constrained deep learning without simulation data. Computer Methods in Applied Mechanics and Engineering, 361:112732, 2020.
  3. 3.Maziar Raissi, Hessam Babaee, and Peyman Givi. Deep learning of turbulent scalar mixing. Physical Review Fluids, 4(12):124501, 2019.
  4. 4.Maziar Raissi, Zhicheng Wang, Michael S Triantafyllou, and George Em Karniadakis. Deep learning of vortex-induced vibrations. Journal of Fluid Mechanics, 861:119–137, 2019.
  5. 5.Xiaowei Jin, Shengze Cai, Hui Li, and George Em Karniadakis. Nsfnets (navier-stokes flow nets): Physics-informed neural networks for the incompressible navier-stokes equations. arXiv preprint arXiv:2003.06496, 2020.
  6. 6.Francisco Sahli Costabal, Yibo Yang, Paris Perdikaris, Daniel E Hurtado, and Ellen Kuhl. Physics-informed neural networks for cardiac activation mapping. Frontiers in Physics, 8:42, 2020.
  7. 7.Georgios Kissas, Yibo Yang, Eileen Hwuang, Walter R Witschey, John A Detre, and Paris Perdikaris. Machine learning in cardiovascular flows modeling: Predicting arterial blood pressure from non-invasive 4d flow mri data using physics-informed neural networks. Computer Methods in Applied Mechanics and Engineering, 358:112623, 2020.
  8. 8.Zhiwei Fang and Justin Zhan. Deep physical informed neural networks for metamaterial design. IEEE Access, 8:24506–24513, 2019.
  9. 9.Dehao Liu and Yan Wang. Multi-fidelity physics-constrained neural network and its application in materials modeling. Journal of Mechanical Design, 141(12), 2019.
  10. 10.Yuyao Chen, Lu Lu, George Em Karniadakis, and Luca Dal Negro. Physics-informed neural networks for inverse problems in nano-optics and metamaterials. Optics Express, 28(8):11618–11633, 2020.
  11. 11.Sifan Wang and Paris Perdikaris. Deep learning of free boundary and stefan problems. arXiv preprint arXiv:2006.05311, 2020.
  12. 12.Yibo Yang and Paris Perdikaris. Adversarial uncertainty quantification in physics-informed neural networks. Journal of Computational Physics, 394:136–152, 2019.
  13. 13.Yinhao Zhu, Nicholas Zabaras, Phaedon-Stelios Koutsourelakis, and Paris Perdikaris. Physics-constrained deep learning for high-dimensional surrogate modeling and uncertainty quantification without labeled data. Journal of Computational Physics, 394:56–81, 2019.
  14. 14.Yibo Yang and Paris Perdikaris. Physics-informed deep generative models. arXiv preprint arXiv:1812.03511, 2018.
  15. 15.Luning Sun and Jian-Xun Wang. Physics-constrained bayesian neural network for fluid flow reconstruction with sparse and noisy data. arXiv preprint arXiv:2001.05542, 2020.
  16. 16.Liu Yang, Xuhui Meng, and George Em Karniadakis. B-pinns: Bayesian physics-informed neural networks for forward and inverse pde problems with noisy data. arXiv preprint arXiv:2003.06097, 2020.
  17. 17.Justin Sirignano and Konstantinos Spiliopoulos. Dgm: A deep learning algorithm for solving partial differential equations. Journal of computational physics, 375:1339–1364, 2018.
  18. 18.Jiequn Han, Arnulf Jentzen, and E Weinan. Solving high-dimensional partial differential equations using deep learning. Proceedings of the National Academy of Sciences, 115(34):8505–8510, 2018.
  19. 19.Dongkun Zhang, Ling Guo, and George Em Karniadakis. Learning in modal space: Solving time-dependent stochastic PDEs using physics-informed neural networks. SIAM Journal on Scientific Computing, 42(2):A639–A665, 2020.
  20. 20.Guofei Pang, Lu Lu, and George Em Karniadakis. fpinns: Fractional physics-informed neural networks. SIAM Journal on Scientific Computing, 41(4):A2603–A2626, 2019.
  21. 21.Guofei Pang, Marta D’Elia, Michael Parks, and George E Karniadakis. npinns: nonlocal physics-informed neural networks for a parametrized nonlocal universal laplacian operator. algorithms and applications. arXiv preprint arXiv:2004.04276, 2020.
  22. 22.Alexandre M Tartakovsky, Carlos Ortiz Marrero, Paris Perdikaris, Guzel D Tartakovsky, and David Barajas-Solano. Learning parameters and constitutive relationships with physics informed deep neural networks. arXiv preprint arXiv:1808.03398, 2018.
  23. 23.Lu Lu, Pengzhan Jin, and George Em Karniadakis. Deeponet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators. arXiv preprint arXiv:1910.03193, 2019.
  24. 24.AM Tartakovsky, C Ortiz Marrero, Paris Perdikaris, GD Tartakovsky, and D Barajas-Solano. Physics-informed deep neural networks for learning parameters and constitutive relationships in subsurface flow problems. Water Resources Research, 56(5):e2019WR026731, 2020.
  25. 25.Yeonjong Shin, Jerome Darbon, and George Em Karniadakis. On the convergence and generalization of physics informed neural networks. arXiv preprint arXiv:2004.01806, 2020.
  26. 26.Hamdi A Tchelepi and Olga Fuks. Limitations of physics informed machine learning for nonlinear two-phase transport in porous media. Journal of Machine Learning for Modeling and Computing, 1(1), 2020.
  27. 27.Maziar Raissi. Deep hidden physics models: Deep learning of nonlinear partial differential equations. The Journal of Machine Learning Research, 19(1):932–955, 2018.
  28. 28.Sifan Wang, Yujun Teng, and Paris Perdikaris. Understanding and mitigating gradient pathologies in physics-informed neural networks. arXiv preprint arXiv:2001.04536, 2020.
  29. 29.Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. In International Conference on Machine Learning, pages 5301–5310, 2019.
  30. 30.Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu. Towards understanding the spectral bias of deep learning. arXiv preprint arXiv:1912.01198, 2019.
  31. 31.Matthew Tancik, Pratul P Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. arXiv preprint arXiv:2006.10739, 2020.
  32. 32.Ronen Basri, Meirav Galun, Amnon Geifman, David Jacobs, Yoni Kasten, and Shira Kritchman. Frequency bias in neural networks for input of non-uniform density. arXiv preprint arXiv:2003.04560, 2020.
  33. 33.Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in neural information processing systems, pages 8571–8580, 2018.
  34. 34.Greg Yang. Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation. arXiv preprint arXiv:1902.04760, 2019.
  35. 35.Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani. Gaussian process behaviour in wide deep neural networks. arXiv preprint arXiv:1804.11271, 2018.
  36. 36.Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. Deep neural networks as gaussian processes. arXiv preprint arXiv:1711.00165, 2017.
  37. 37.David JC MacKay. Introduction to gaussian processes. NATO ASI Series F Computer and Systems Sciences, 168:133–166, 1998.
  38. 38.Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang. On exact computation with an infinitely wide neural net. In Advances in Neural Information Processing Systems, pages 8141–8150, 2019.
  39. 39.Maziar Raissi, Paris Perdikaris, and George E Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019.
  40. 40.Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. In Advances in neural information processing systems, pages 8572–8583, 2019.
  41. 41.Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma. Frequency principle: Fourier analysis sheds light on deep neural networks. arXiv preprint arXiv:1901.06523, 2019.
  42. 42.Basri Ronen, David Jacobs, Yoni Kasten, and Shira Kritchman. The convergence rate of neural networks for learned functions of different frequencies. In Advances in Neural Information Processing Systems, pages 4761–4771, 2019.
  43. 43.Lu Lu, Xuhui Meng, Zhiping Mao, and George E Karniadakis. Deepxde: A deep learning library for solving differential equations. arXiv preprint arXiv:1907.04502, 2019.
  44. 44.Parviz Moin. Fundamentals of engineering numerical analysis. Cambridge University Press, 2010.
  45. 45.L.C. Evans and American Mathematical Society. Partial Differential Equations. Graduate studies in mathematics. American Mathematical Society, 1998.
  46. 46.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256, 2010.
  47. 47.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.

Citation

MLA
Wang, S., et al. “When and Why PINNs Fail to Train: A Neural Tangent Kernel Perspective”. arXiv, 2020, http://arxiv.org/abs/2007.14527v1.
APA
Wang, S., Yu, X., & Perdikaris, P. (2020). When and why PINNs fail to train: A neural tangent kernel perspective. arXiv. http://arxiv.org/abs/2007.14527v1
Chicago
Wang, S., X. Yu, and P. Perdikaris. 2020. “When and Why PINNs Fail to Train: A Neural Tangent Kernel Perspective”. arXiv. http://arxiv.org/abs/2007.14527v1.
Harvard
Wang, S., Yu, X. and Perdikaris, P. (2020) “When and why PINNs fail to train: A neural tangent kernel perspective”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2007.14527v1.
Vancouver
1. Wang S, Yu X, Perdikaris P (2020) When and why PINNs fail to train: A neural tangent kernel perspective. arXiv

BibTeX

@article{wang2020when,
  title = {When and why PINNs fail to train: A neural tangent kernel perspective},
  author = {Wang, Sifan and Yu, Xinling and Perdikaris, Paris},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2007.14527v1},
  eprint = {2007.14527}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF