PDEBench: An Extensive Benchmark for Scientific Machine Learning

Makoto TakamotoTimothy PraditiaRaphael LeiteritzDaniel MacKinlayFrancesco AlesianiDirk PflügerMathias Niepert

article2022NeurIPS383 citationsSimTech Best Paper Award 2023

Presents PDEBench, an extensive benchmark suite featuring diverse time-dependent physical simulations, large ready-to-use datasets, and standardized baselines to systematically evaluate and compare scientific machine learning models against classical numerical methods.

Listen

Data-driven scientific machine learning holds substantial promise for modeling complex physical systems governed by partial differential equations, such as fluid dynamics, environmental contaminant transport, and wave propagation. Traditional numerical solvers are computationally expensive and struggle with non-differentiable formulations, which complicates optimization, sensitivity analysis, and the resolution of inverse problems. However, progress in this field has been severely limited by the lack of standardized, large-scale, and challenging benchmark datasets that evaluate whether neural surrogate models accurately capture real-world physics.

The article introduces and evaluates PDEBENCH, an open-source, extensible benchmark suite designed to assess machine learning models against classical numerical simulations across diverse scientific modeling tasks. The primary objective is to provide a rigorous, standardized platform comprising extensive datasets, reproducible baseline models, and physics-informed evaluation metrics to identify key challenges and guide algorithm development in scientific machine learning.

The benchmark encompasses 35 datasets across 11 distinct partial differential equation systems spanning one-, two-, and three-dimensional domains. These include stylized baseline models, such as advection and reaction-diffusion, as well as complex scenarios featuring compressible and incompressible fluid flows, shallow-water waves, and groundwater contaminant transport. Datasets incorporate realistic boundary conditions and parameter variations, such as varying viscosities and Mach numbers. Using standard deep learning frameworks, the authors evaluated prominent machine learning architectures—specifically the Fourier Neural Operator, U-Net, and Physics-Informed Neural Networks—on both forward time-stepping emulation and gradient-based inverse inference for estimating unknown initial conditions.

The experimental findings demonstrate critical insights into current model performance. First, the Fourier Neural Operator consistently outperformed other architectures across most forward and inverse tasks, maintaining stable predictions across the frequency spectrum. Second, standard machine learning error metrics proved inadequate; global root-mean-squared error masked critical local physical errors in turbulent regimes and across discontinuities. Third, machine learning models delivered massive computational speedups during deployment—generating predictions up to three orders of magnitude faster than classical solvers—while bypassing classical time-stepping stability restrictions. Finally, significant performance degradations occurred during temporal extrapolation beyond trained time horizons, as well as in regimes governed by high Reynolds numbers or sharp shock waves.

These results demonstrate that machine learning emulators can drastically reduce operational compute costs for physical simulations, making real-time design and optimization viable. However, the findings also show that current models struggle with generalization over longer horizons and non-smooth physical regimes. Relying on global aggregate errors introduces operational risk because a model may appear statistically accurate while severely violating physical conservation laws or boundary dynamics.

Organizations developing or deploying scientific machine learning tools should adopt multi-faceted, physics-informed metrics rather than relying solely on standard error metrics. Decision-makers should prioritize architectures that operate effectively across frequency spaces, such as neural operators, for surrogate modeling workflows. Additionally, future research must focus on stabilizing autoregressive predictions over extended time horizons, improving performance on high-frequency shock dynamics, and expanding benchmarks to multi-phase flows and complex geometric domains.

Cover for PDEBench: An Extensive Benchmark for Scientific Machine Learning

Abstract

Machine learning-based modeling of physical systems has experienced increased interest in recent years. Despite some impressive progress, there is still a lack of benchmarks for Scientific ML that are easy to use but still challenging and representative of a wide range of problems. We introduce PDEBench, a benchmark suite of time-dependent simulation tasks based on Partial Differential Equations (PDEs). PDEBench comprises both code and data to benchmark the performance of novel machine learning models against both classical numerical simulations and machine learning baselines. Our proposed set of benchmark problems contribute the following unique features: (1) A much wider range of PDEs compared to existing benchmarks, ranging from relatively common examples to more realistic and difficult problems; (2) much larger ready-to-use datasets compared to prior work, comprising multiple simulation runs across a larger number of initial and boundary conditions and PDE parameters; (3) more extensible source codes with user-friendly APIs for data generation and baseline results with popular machine learning models (FNO, U-Net, PINN, Gradient-Based Inverse Method). PDEBench allows researchers to extend the benchmark freely for their own purposes using a standardized API and to compare the performance of new models to existing baseline methods. We also propose new evaluation metrics with the aim to provide a more holistic understanding of learning methods in the context of Scientific ML. With those metrics we identify tasks which are challenging for recent ML methods and propose these tasks as future challenges for the community. The code is available at this https URL.

Table of Contents

  • 1 Motivation
  • 2 Related Work
  • 3 PDEBench: A Benchmark for Scientific Machine Learning
  • 3.1 General Problem Definition
  • 3.2 Overview of Datasets and PDEs
  • 3.3 Overview of Metrics
  • 3.4 Existing Baseline Surrogate Models
  • 3.5 Data Format, Benchmark Access, Maintenance, and Extensibility
  • 4 A Selection of Experiments
  • 4.1 Baseline Setups
  • 4.2 Baseline Performance
  • 4.3 Temporal Error Analysis
  • 4.4 Inference Time Comparison
  • 5 Conclusions and Limitations
  • References
  • A Continuation of Related Work
  • B Detailed metrics description
  • B.1 Inverse Problem Metrics
  • C Training Protocol and Hyperparameters
  • C.1 Inverse problem
  • D Detailed Problem Description
  • D.1 1D Advection Equation
  • D.2 1D Diffusion-Reaction Equation
  • D.3 Burgers equation
  • D.4 Darcy Flow
  • D.5 Compressible Navier-Stokes equation
  • D.6 Inhomogenous, incompressible Navier-Stokes
  • D.7 2D Shallow-Water Equations
  • D.8 Diffusion-Sorption Equation
  • D.9 2D Diffusion-Reaction Equation
  • D.10 Gradient-Based Inverse Method
  • E Detailed Baseline Score
  • F Detailed Runtime Comparison
  • G Resolution Sensitivity of Inference Time
  • H Error Comparison with PDE Solver
  • I Visualization of Model Predictions
  • J Visualization of Initial Conditions
  • K PDEBench’s Data Sheet
  • K.1 Motivation
  • K.2 Composition
  • K.3 Collection Process
  • K.4 Preprocessing/cleaning/labeling
  • K.5 Uses
  • K.6 Distribution
  • K.7 Maintenance
  • K.8 Reproducibility of the baseline score
  • K.9 Reading and using the dataset
  • K.10 Data Format

Knowls

  1. Knowl 1 — PDEBENCH Benchmark Suite Composition

    data/table

    PDEBENCH is a standardized benchmark suite comprising 35 datasets across 11 distinct partial differential equation (PDE) systems. The suite spans 1D, 2D, and 3D spatial domains (Nd∈{1,2,3}N_d \in \{1, 2, 3\}) and includes both time-dependent and steady-state systems across varying physical parameters (e.g., advection velocity, diffusion coefficients, viscosity, Mach numbers), diverse initial condition distributions (superposition of sinusoidal modes, random Gaussian fields, Riemann shock-tube problems), and diverse boundary conditions (periodic, Dirichlet, Neumann, Cauchy/derivative flux, and outgoing absorbing boundaries).

    PDE System NdN_d Time-dependent Spatial Resolution (NsN_s) Temporal Steps (NtN_t) Sample Count
    Advection 1 Yes 1024 200 10000
    Burgers' 1 Yes 1024 200 10000
    Diffusion-Reaction 1 Yes 1024 200 10000
    Diffusion-Reaction 2 Yes 128×128128 \times 128 100 1000
    Diffusion-Sorption 1 Yes 1024 100 10000
    Compressible Navier-Stokes 1 Yes 1024 100 10000
    Compressible Navier-Stokes 2 Yes 512×512512 \times 512 21 1000
    Compressible Navier-Stokes 3 Yes 128×128×128128 \times 128 \times 128 21 100
    Incompressible Navier-Stokes 2 Yes 256×256256 \times 256 1000 1000
    Darcy Flow 2 No 128×128128 \times 128 – 10000
    Shallow-Water 2 Yes 128×128128 \times 128 100 1000

    The datasets are formatted in HDF5 files structured as (b×t×x1×⋯×xd×v)(b \times t \times x_1 \times \dots \times x_d \times v), where bb denotes the batch sample dimension, tt the temporal index, x1,…,xdx_1, \dots, x_d the spatial dimensions, and vv the physical state variable channels (e.g., density, velocity components, pressure).

  2. Knowl 2 — Forward Propagator Surrogate Learning Formulation

    model/method

    Let vθ:T×S×Θ→Rd\mathbf{v}_\theta: \mathcal{T} \times \mathcal{S} \times \Theta \to \mathbb{R}^d denote the vector-valued solution to a partial differential equation on a spatial domain S\mathcal{S} with temporal index T\mathcal{T} and parameter space Θ\Theta. The forward propagator operator Fθ:vθ(t,⋅)↦vθ(t+1,⋅)\mathcal{F}_\theta: \mathbf{v}_\theta(t, \cdot) \mapsto \mathbf{v}_\theta(t+1, \cdot) advances the system state by one discrete time step.

    To accommodate temporal derivative approximations via finite differences, the discretized forward propagator F˚θ\mathring{\mathcal{F}}_\theta operates on a history of ℓ≥1\ell \ge 1 consecutive timesteps:

    F˚θ:vθ([t−ℓ:t−1],⋅)↦vθ(t,⋅)\mathring{\mathcal{F}}_\theta: \mathbf{v}_\theta([t-\ell:t-1], \cdot) \mapsto \mathbf{v}_\theta(t, \cdot)

    where vθ([t−ℓ:t−1],⋅)≡(vθ(t−ℓ,⋅),…,vθ(t−1,⋅))\mathbf{v}_\theta([t-\ell:t-1], \cdot) \equiv (\mathbf{v}_\theta(t-\ell, \cdot), \ldots, \mathbf{v}_\theta(t-1, \cdot)).

    Given a simulation dataset D={vθk(k)([0:tmax⁡],⋅)∣k=1,…,K}\mathcal{D} = \{\mathbf{v}_{\theta_k}^{(k)}([0:t_{\max}], \cdot) \mid k = 1, \ldots, K\} of discretized trajectory samples generated by high-precision numerical solvers, a parametric machine learning surrogate Fθ,ϕ\mathcal{F}_{\theta, \phi} with trainable parameters ϕ\phi is trained by minimizing an empirical loss functional L\mathcal{L}:

    ϕ^=arg⁡min⁡ϕ∑t=1tmax⁡∑k=1KL(Fθk,ϕ{vθk(k)([t−ℓ:t−1],⋅)},vθk(k)(t,⋅))\hat{\phi} = \arg\min_\phi \sum_{t=1}^{t_{\max}} \sum_{k=1}^K \mathcal{L}\left(\mathcal{F}_{\theta_k, \phi}\left\{\mathbf{v}_{\theta_k}^{(k)}([t-\ell:t-1], \cdot)\right\}, \mathbf{v}_{\theta_k}^{(k)}(t, \cdot)\right)

  3. Knowl 3 — Surrogate Gradient-Based PDE Inversion Framework

    model/method

    In inverse PDE problems, an unobserved initial condition u0=vθ(0,⋅)\mathbf{u}_0 = \mathbf{v}_\theta(0, \cdot) or latent parameter vector θ\theta is inferred from observed outputs vθ([t:t+ℓ],⋅)\mathbf{v}_\theta([t:t+\ell], \cdot) at future time steps. Assuming an observational model with mean-zero additive noise ϵ\boldsymbol{\epsilon}:

    vθ(t,⋅)=Fθ,ϕ{vθ([t−ℓ:t−1],⋅)}+ϵ\mathbf{v}_\theta(t, \cdot) = \mathcal{F}_{\theta, \phi}\{\mathbf{v}_\theta([t-\ell:t-1], \cdot)\} + \boldsymbol{\epsilon}

    Because the forward neural surrogate model Fθ,ϕ\mathcal{F}_{\theta, \phi} is continuously differentiable with respect to its inputs, inverse estimation is performed via gradient-based optimization directly through the surrogate network. For an unknown initial condition u0\mathbf{u}_0, a parameterization u^0=pθ(u0)\hat{\mathbf{u}}_0 = p_\theta(\mathbf{u}_0) is constructed using a low-dimensional latent grid (e.g., 64 hidden nodes) upsampled via bilinear interpolation. The estimated initial condition is obtained by gradient descent minimizing the prediction loss at a target evaluation horizon t=Tt = T:

    Linv=L(u(t=T,x∣u0),u(t=T,x∣u^0))\mathcal{L}_{\text{inv}} = \mathcal{L}\left(\mathbf{u}(t = T, \mathbf{x} \mid \mathbf{u}_0), \mathbf{u}(t = T, \mathbf{x} \mid \hat{\mathbf{u}}_0)\right)

  4. Knowl 4 — Evaluation Metrics for Scientific Machine Learning Surrogates

    definition

    To evaluate physical consistency, scale invariance, and multi-scale frequency accuracy of scientific ML surrogates beyond global root-mean-squared error (RMSE), the following physics-informed metrics are defined for true field utrue\mathbf{u}_{\text{true}} and predicted field upred\mathbf{u}_{\text{pred}} over NN spatial discretization points:

    1. Normalized RMSE (nRMSE) for scale-invariant error assessment: nRMSE=∥upred−utrue∥2∥utrue∥2\text{nRMSE} = \frac{\|\mathbf{u}_{\text{pred}} - \mathbf{u}_{\text{true}}\|_2}{\|\mathbf{u}_{\text{true}}\|_2}

    2. Maximum Error (max error) measuring local worst-case deviations and stability: max error=max⁡x∣upred(x)−utrue(x)∣\text{max error} = \max_{\mathbf{x}} |\mathbf{u}_{\text{pred}}(\mathbf{x}) - \mathbf{u}_{\text{true}}(\mathbf{x})|

    3. Conserved Quantity RMSE (cRMSE) measuring violations of physical conservation laws: cRMSE=1N∥∑xupred(x)−∑xutrue(x)∥2\text{cRMSE} = \frac{1}{N} \left\| \sum_{\mathbf{x}} \mathbf{u}_{\text{pred}}(\mathbf{x}) - \sum_{\mathbf{x}} \mathbf{u}_{\text{true}}(\mathbf{x}) \right\|_2

    4. Boundary RMSE (bRMSE) assessing boundary condition preservation over domain boundary ∂S\partial \mathcal{S}: bRMSE=1∣∂S∣∑x∈∂S∣upred(x)−utrue(x)∣2\text{bRMSE} = \sqrt{\frac{1}{|\partial \mathcal{S}|} \sum_{\mathbf{x} \in \partial \mathcal{S}} |\mathbf{u}_{\text{pred}}(\mathbf{x}) - \mathbf{u}_{\text{true}}(\mathbf{x})|^2}

    5. Fourier Space Band-Filtered RMSE (fRMSE) measuring spectral error across wavenumber bands [kmin⁡,kmax⁡][k_{\min}, k_{\max}]: fRMSE(kmin⁡,kmax⁡)=∑k=kmin⁡kmax⁡∣F(upred)(k)−F(utrue)(k)∣2kmax⁡−kmin⁡+1\text{fRMSE}(k_{\min}, k_{\max}) = \sqrt{\frac{\sum_{k=k_{\min}}^{k_{\max}} |\mathcal{F}(\mathbf{u}_{\text{pred}})(k) - \mathcal{F}(\mathbf{u}_{\text{true}})(k)|^2}{k_{\max} - k_{\min} + 1}} where F\mathcal{F} is the discrete Fourier transform. Frequency bands are partitioned into Low (kmin⁡=0,kmax⁡=4k_{\min}=0, k_{\max}=4), Middle (kmin⁡=5,kmax⁡=12k_{\min}=5, k_{\max}=12), and High (kmin⁡=13,kmax⁡=∞k_{\min}=13, k_{\max}=\infty).

  5. Knowl 5 — Compressible Navier-Stokes Benchmark Formulation

    equation

    The compressible Navier-Stokes equations model non-linear compressible fluid dynamics across subsonic and supersonic regimes. The governing equations in conservative form are:

    ∂tρ+∇⋅(ρv)=0\partial_t \rho + \nabla \cdot (\rho \mathbf{v}) = 0

    ρ(∂tv+v⋅∇v)=−∇p+ηΔv+(ζ+η3)∇(∇⋅v)\rho (\partial_t \mathbf{v} + \mathbf{v} \cdot \nabla \mathbf{v}) = -\nabla p + \eta \Delta \mathbf{v} + \left(\zeta + \frac{\eta}{3}\right) \nabla (\nabla \cdot \mathbf{v})

    ∂t(ϵ+12ρv2)+∇⋅[(p+ϵ+12ρv2)v−v⋅σ′]=0\partial_t \left(\epsilon + \frac{1}{2}\rho \mathbf{v}^2\right) + \nabla \cdot \left[\left(p + \epsilon + \frac{1}{2}\rho \mathbf{v}^2\right)\mathbf{v} - \mathbf{v} \cdot \boldsymbol{\sigma}'\right] = 0

    where ρ\rho is the mass density, v\mathbf{v} is the fluid velocity vector, pp is the gas pressure, η\eta and ζ\zeta are shear and bulk viscosity coefficients, σ′\boldsymbol{\sigma}' is the viscous stress tensor, and ϵ=p/(Γ−1)\epsilon = p / (\Gamma - 1) is the internal energy with adiabatic index Γ=5/3\Gamma = 5/3.

    The benchmark implements configurations across 1D, 2D, and 3D with:

    1. Mach numbers M=∣v∣/cs∈{0.1,1.0}M = |\mathbf{v}|/c_s \in \{0.1, 1.0\}, where sound speed cs=Γp/ρc_s = \sqrt{\Gamma p / \rho}.
    2. Viscosities ranging from viscous regimes (η=ζ=0.1,0.01\eta = \zeta = 0.1, 0.01) to inviscid limits (η=ζ=10−8\eta = \zeta = 10^{-8}).
    3. Initial fields including random superposition fields, turbulent divergence-free velocity fields via Helmholtz decomposition in Fourier space, and Riemann shock-tube problems with discontinuous initial states (ρL,vL,pL)(\rho_L, \mathbf{v}_L, p_L) and (ρR,vR,pR)(\rho_R, \mathbf{v}_R, p_R) generating shock and rarefaction waves.
  6. Knowl 6 — 2D Shallow-Water Hyperbolic Flow Formulation

    equation

    The 2D shallow-water equations represent free-surface hyperbolic flow derived by depth-integrating the Navier-Stokes equations under hydrostatic pressure assumptions:

    ∂th+∂x(hu)+∂y(hv)=0\partial_t h + \partial_x (h u) + \partial_y (h v) = 0

    ∂t(hu)+∂x(hu2+12grh2)+∂y(huv)=−grh∂xb\partial_t (h u) + \partial_x \left(h u^2 + \frac{1}{2} g_r h^2\right) + \partial_y (h u v) = -g_r h \partial_x b

    ∂t(hv)+∂y(hv2+12grh2)+∂x(huv)=−grh∂yb\partial_t (h v) + \partial_y \left(h v^2 + \frac{1}{2} g_r h^2\right) + \partial_x (h u v) = -g_r h \partial_y b

    where h(t,x,y)h(t, x, y) is the total water depth, u=(u,v)T\mathbf{u} = (u, v)^T is the depth-averaged horizontal velocity vector, b(x,y)b(x, y) is the spatially varying bathymetry, and grg_r is gravitational acceleration. The quantities huh u and hvh v represent directional momentum.

    The benchmark implements a 2D radial dam break problem on a square domain Ω=[−2.5,2.5]2\Omega = [-2.5, 2.5]^2 with circular discontinuity initial condition:

    h(0,x,y)={2.0if x2+y2<r1.0if x2+y2≥rh(0, x, y) = \begin{cases} 2.0 & \text{if } \sqrt{x^2 + y^2} < r \\ 1.0 & \text{if } \sqrt{x^2 + y^2} \ge r \end{cases}

    where radius r∼U(0.3,0.7)r \sim \mathcal{U}(0.3, 0.7) is uniformly distributed, and Neumann no-flow conditions are enforced on the boundaries.

  7. Knowl 7 — 1D Diffusion-Sorption Equation with Freundlich Isotherm

    equation

    The 1D diffusion-sorption equation models non-linear retarded solute transport in porous media:

    ∂tu(t,x)=DR(u)∂xxu(t,x),x∈(0,1),  t∈(0,500]\partial_t u(t, x) = \frac{D}{R(u)} \partial_{xx} u(t, x), \quad x \in (0, 1), \; t \in (0, 500]

    where D=5×10−4D = 5 \times 10^{-4} is the effective diffusion coefficient and R(u)R(u) is the non-linear retardation factor governed by the Freundlich sorption isotherm:

    R(u)=1+1−ϕϕρsknfunf−1R(u) = 1 + \frac{1 - \phi}{\phi} \rho_s k n_f u^{n_f - 1}

    with porous medium porosity ϕ=0.29\phi = 0.29, bulk density ρs=2880\rho_s = 2880, Freundlich distribution coefficient k=3.5×10−4k = 3.5 \times 10^{-4}, and exponent nf=0.874n_f = 0.874. The model exhibits a mathematical singularity as u→0u \to 0.

    The boundary conditions are non-periodic and include a derivative Neumann-type flux boundary:

    u(t,0)=1.0,u(t,1)=D∂xu(t,1)u(t, 0) = 1.0, \quad u(t, 1) = D \partial_x u(t, 1)

    with initial condition sampled uniformly as u(0,x)∼U(0,0.2)u(0, x) \sim \mathcal{U}(0, 0.2).

  8. Knowl 8 — Baseline ML Surrogate Architectures and Training Configurations

    model/method

    PDEBENCH evaluates three primary baseline surrogate paradigms:

    1. Fourier Neural Operator (FNO): Parameterizes integral kernel operators directly in the Fourier frequency domain using discrete Fourier transforms, enabling mesh- and resolution-invariant forward propagation. Trained for 500 epochs using Adam with initial learning rate 10−310^{-3}, decayed by 0.50.5 every 100 epochs.

    2. U-Net with Pushforward Trick: Multi-scale encoder-decoder convolutional network with skip connections, extended to 1D, 2D, and 3D spatial dimensions. Fully open-loop autoregressive training accumulates catastrophic rollout error; thus, U-Net is trained using a pushforward unrolling trick over multi-step temporal horizons to stabilize autoregressive evaluation.

    3. Physics-Informed Neural Networks (PINNs): Fully connected deep networks (depth 6, width 40 per layer) trained using DeepXDE for 15,000 epochs (Adam, learning rate 10−310^{-3}) to minimize PDE residuals, initial condition loss, and boundary condition loss. Because PINNs optimize for a single fixed instance, evaluations are averaged across 10 distinct initial condition samples per dataset.

  9. Knowl 9 — Comparative Performance and Failure Modes of Baseline Surrogates

    empirical result

    Empirical evaluations across the 11 PDE benchmarks reveal distinct model behaviors and failure modes:

    1. FNO Superiority in Smooth Regimes: FNO achieves the lowest overall RMSE, cRMSE (conserved quantities), and bRMSE (boundary tracking) across most 1D, 2D, and 3D datasets, maintaining a consistent spectral error of approximately 4×10−44 \times 10^{-4} across low- and mid-frequency bands.

    2. Gibbs Phenomenon in Discontinuous/Inviscid Regimes: In low-diffusion regimes (e.g., Burgers' equation as ν→0.001\nu \to 0.001 and inviscid Compressible Navier-Stokes), FNO's high-frequency Fourier error (extfRMSEhigh ext{fRMSE}_{\text{high}}) increases by over two orders of magnitude due to frequency truncation and Gibbs oscillations near sharp shock fronts.

    3. Scale Sensitivity of Normalized RMSE: In Darcy flow benchmarks, nRMSE increases sharply as the forcing scale β\beta decreases (reaching nRMSE>1.0\text{nRMSE} > 1.0 at β=0.01\beta = 0.01, where mean field magnitude is ≈0.01\approx 0.01), demonstrating that global RMSE is inadequate for small-amplitude solutions.

    4. PINN Generalization Bottleneck: PINNs demonstrate lower high-frequency error on specific individual samples but cannot infer over varying parameter or initial condition distributions without complete retraining.

  10. Knowl 10 — Inference Efficiency and Diffusion CFL Stability Bypass

    empirical result

    Neural surrogate inference provides 2 to 3 orders of magnitude acceleration relative to classical numerical simulation solvers for forward evaluation.

    PDE Scenario Resolution Classical Solver Time (s) FNO Inference Time (s) FNO Training Time (s)
    Diffusion-sorption (1D) 102411024^1 59.83 0.32 48760
    2D CFD (η=ζ=0.1\eta = \zeta = 0.1) 5122512^2 582.61 0.14 107567
    3D CFD (inviscid) 64364^3 60.06 0.14 12387
    2D Diffusion-reaction 1282128^2 2.21 0.40 54140
    2D Shallow-water 1282128^2 0.62 0.37 52580

    Explicit numerical solvers are strictly constrained by the Courant-Friedrichs-Lewy (CFL) parabolic stability condition, which dictates time step size Δt∝Δx2/η\Delta t \propto \Delta x^2 / \eta, causing numerical runtime to scale quadratically with mesh refinement and inversely with physical viscosity/diffusion η\eta. In contrast, neural operator (FNO and U-Net) inference time is independent of η\eta, yielding dramatic computational advantages in strongly diffusive or viscous regimes.

  11. Knowl 11 — Temporal Extrapolation Breakdown in Autoregressive Surrogates

    empirical result

    When neural operator models (FNO) are trained on trajectory intervals up to time ttrain=1.0t_{\text{train}} = 1.0 and evaluated on future horizons up to t=2.0t = 2.0 (tested on 1D Advection and 1D Burgers' equations), prediction error increases monotonically and steeply immediately for t>ttraint > t_{\text{train}}. Surrogates fail to capture long-term dynamic time invariance, demonstrating that out-of-distribution temporal extrapolation remains an open challenge for Scientific ML models.

  12. Knowl 12 — PDEBENCH Scope and Structural Limitations

    limitation

    PDEBENCH has several defined domain and structural limitations:

    1. Geometry and Discretization: All simulations are restricted to rectangular/Cartesian computational domains ([0,1]d[0, 1]^d or [−2.5,2.5]2[-2.5, 2.5]^2); irregular boundary geometries, non-Euclidean meshes, and unstructured grids are not included.
    2. Physical Scope: The benchmark focuses on hydromechanical and transport field equations (compressible/incompressible flow, advection, reaction-diffusion, shallow-water); multi-phase flows, phase transitions, and quantum mechanical systems are excluded.
    3. 3D Sample Size Limitations: Due to disk storage and computational simulation constraints, 3D Compressible Navier-Stokes datasets are limited to 100 samples at 1283128^3 (or 64364^3) resolution, which may constrain deep neural network convergence.

Coverage note — None was omitted; all 11 benchmark equations, forward/inverse formulations, novel evaluation metrics, baseline surrogate architectures, empirical findings, and stated limitations were converted into self-contained knowls.

References

  1. 1.K. R. Allen, T. Lopez-Guevara, K. Stachenfeld, A. Sanchez-Gonzalez, P. Battaglia, J. Hamrick, and T. Pfaff. Physical design using differentiable learned simulators. arXiv preprint arXiv:2202.00728, 2022.
  2. 2.F. d. A. Belbute-Peres, Y.-f. Chen, and F. Sha. Hyperpinn: Learning parameterized differential equations with physics-informed hypernetworks. In International Conference on Learning Representations, 2021. URL https://openreview.net/pdf?id=LxUuRDUhRjM.
  3. 3.D. A. Bezgin, A. B. Buhendwa, and N. A. Adams. JAX-FLUIDS: A fully-differentiable high-order computational fluid dynamics solver for compressible two-phase flows, 2022. URL https://arxiv.org/abs/2203.13760.
  4. 4.J. Brandstetter, D. Worrall, and M. Welling. Message passing neural pde solvers. In The Tenth International Conference on Learning Representations, 2022. URL https://openreview.net/pdf?id=vSix3HPYKSU.
  5. 5.G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba. OpenAI Gym, 2016.
  6. 6.S. L. Brunton and J. N. Kutz. Data-Driven Science and Engineering: Machine Learning, Dynamical Systems, and Control. Cambridge University Press, 2019. ISBN 978-1-108-42209-3. URL https://databookuw.com.
  7. 7.K. Cao. Inverse Problems for the Heat Equation Using Conjugate Gradient Methods. PhD thesis, University of Leeds, 2018.
  8. 8.T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud. Neural ordinary differential equations. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 6572–6583. Curran Associates, Inc., 2018.
  9. 9.M. Chu, N. Thuerey, H.-P. Seidel, C. Theobalt, and R. Zayer. Learning meaningful controls for fluids. ACM Transactions on Graphics, 40(4):1–13, Aug. 2021. doi: 10.1145/3476576.3476661.
  10. 10.M. Crosas. The dataverse network: An open-source application for sharing, discovering and preserving data. D-Lib Magazine, Volume 17, 2011. URL http://www.dlib.org/dlib/january11/crosas/01crosas.html.
  11. 11.H. Eivazi, M. Tahani, P. Schlatter, and R. Vinuesa. Physics-informed neural networks for solving reynolds-averaged navier-stokes equations. arXiv:2107.10711 [physics], Jul 2021. URL http://arxiv.org/abs/2107.10711. arXiv: 2107.10711.
  12. 12.R. P. Feynman. Feynman lectures on physics - Volume 1. 1963.
  13. 13.C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem. Brax–a differentiable physics engine for large scale rigid body simulation. arXiv preprint arXiv:2106.13281, 2021. URL https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/file/3def184ad8f4755ff269862ea77393dd-Paper-round1.pdf.
  14. 14.D. W. Gladish, D. E. Pagendam, L. J. M. Peeters, P. M. Kuhnert, and J. Vaze. Emulation Engines: Choice and Quantification of Uncertainty for Complex Hydrological Models. Journal of Agricultural, Biological and Environmental Statistics, 23(1):39–62, Mar. 2018. doi: 10.1007/s13253-017-0308-3.
  15. 15.T. H. Group. An overview of the hdf5 technology suite and its applications, 2022. URL https://portal.hdfgroup.org/display/HDF5/HDF5.
  16. 16.M. Guo and J. S. Hesthaven. Data-driven reduced order modeling for time-dependent problems. Computer Methods in Applied Mechanics and Engineering, 345:75–99, 2019. ISSN 0045-7825. doi: https://doi.org/10.1016/j.cma.2018.10.029. URL https://www.sciencedirect.com/science/article/pii/S0045782518305334.
  17. 17.O. Hennigh, S. Narasimhan, M. A. Nabian, A. Subramaniam, K. Tangsali, M. Rietmann, J. d. A. Ferrandis, W. Byeon, Z. Fang, and S. Choudhry. NVIDIA SimNet{TM}: An AI-accelerated multi-physics simulation framework. arXiv:2012.07938 [physics], Dec. 2020.
  18. 18.P. Holl and V. Koltun. Phiflow: A differentiable PDE solving framework for deep learning via physical simulations. page 5, 2020.
  19. 19.Z. Huang, T. Schneider, M. Li, C. Jiang, D. Zorin, and D. Panozzo. A large-scale benchmark for the incompressible Navier-Stokes equations, 2021. URL https://arxiv.org/abs/2112.05309.
  20. 20.T. Inoue, S.-i. Inutsuka, and H. Koyama. The Role of Ambipolar Diffusion in the Formation Process of Moderately Magnetized Diffuse Clouds. The Astrophysical Journal, 658(2):L99–L102, Apr. 2007. doi: 10.1086/514816.
  21. 21.M. Karlbauer, T. Praditia, S. Otte, S. Oladyshkin, W. Nowak, and M. V. Butz. Composing partial differential equations with physics-aware neural networks. In Proceedings of the 39th International Conference on Machine Learning, Proceedings of Machine Learning Research, Baltimore, USA, 16–23 Jul 2022.
  22. 22.K. Kashinath, M. Mustafa, A. Albert, J.-L. Wu, C. Jiang, S. Esmaeilzadeh, K. Azizzadenesheli, R. Wang, A. Chattopadhyay, A. Singh, A. Manepalli, D. Chirila, R. Yu, R. Walters, B. White, H. Xiao, H. A. Tchelepi, P. Marcus, A. Anandkumar, P. Hassanzadeh, and n. Prabhat. Physics-informed machine learning: Case studies for weather and climate modelling. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2194):20200093, Apr. 2021. doi: 10.1098/rsta.2020.0093.
  23. 23.D. I. Ketcheson, K. T. Mandli, A. J. Ahmadia, A. Alghamdi, M. Quezada de Luna, M. Parsani, M. G. Knepley, and M. Emmett. PyClaw: Accessible, Extensible, Scalable Tools for Wave Propagation Problems. SIAM Journal on Scientific Computing, 34(4):C210–C231, Nov. 2012.
  24. 24.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  25. 25.G. Klaasen and W. Troy. Stationary wave solutions of a system of reaction-diffusion equations derived from the fitzhugh–nagumo equations. SIAM Journal on Applied Mathematics, 44(1):96–110, 1984. doi: 10.1137/0144008.
  26. 26.N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, and A. Anandkumar. Neural operator: Learning Maps Between Function Spaces. arXiv:2108.08481 [cs, math], Sept. 2021.
  27. 27.A. Krishnapriyan, A. Gholami, S. Zhe, R. Kirby, and M. W. Mahoney. Characterizing possible failure modes in physics-informed neural networks. Advances in Neural Information Processing Systems, 34, 2021.
  28. 28.A. Lavin, H. Zenil, B. Paige, D. Krakauer, J. Gottschlich, T. Mattson, A. Anandkumar, S. Choudry, K. Rocki, A. G. Baydin, C. Prunkl, B. Paige, O. Isayev, E. Peterson, P. L. McMahon, J. Macke, K. Cranmer, J. Zhang, H. Wainwright, A. Hanuka, M. Veloso, S. Assefa, S. Zheng, and A. Pfeffer. Simulation Intelligence: Towards a New Generation of Scientific Methods. arXiv:2112.03235 [cs], Dec. 2021.
  29. 29.R. Leiteritz, M. Hurler, and D. Pflüger. Learning free-surface flow with physics-informed neural networks. In 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA), pages 1668–1673, 2021. doi: 10.1109/ICMLA52953.2021.00266.
  30. 30.R. Leiteritz, P. Buchfink, B. Haasdonk, and D. Pflüger. Surrogate-data-enriched physics-aware neural networks. Proceedings of the Northern Lights Deep Learning Workshop, 2, 04 2022. doi: 10.7557/18.6268.
  31. 31.H. Lewy, K. Friedrichs, and R. Courant. Über die partiellen differenzengleichungen der mathematischen physik. Mathematische annalen, 100:32–74, 1928.
  32. 32.Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Fourier neural operator for parametric partial differential equations. International Conference on Learning Representations (ICLR), 2021.
  33. 33.G. Limousin, J.-P. Gaudet, L. Charlet, S. Szenknect, V. Barthès, and M. Krimissa. Sorption isotherms: A review on physical bases, modeling and measurement. Applied Geochemistry, 22(2):249–275, 2007. ISSN 0883-2927. doi: https://doi.org/10.1016/j.apgeochem.2006.09.010. URL https://www.sciencedirect.com/science/article/pii/S0883292706002629.
  34. 34.L. Lu, X. Meng, Z. Mao, and G. E. Karniadakis. DeepXDE: A deep learning library for solving differential equations. SIAM Review, 63(1):208–228, 2021. doi: 10.1137/19M1274067.
  35. 35.D. MacKinlay, D. Pagendam, P. M. Kuhnert, T. Cui, D. Robertson, and S. Janardhanan. Model Inversion for Spatio-temporal Processes using the Fourier Neural Operator. In Neurips Workshop on Machine Learning for the Physical Sciences, page 7, 2021.
  36. 36.S. Makridakis, E. Spiliotis, and V. Assimakopoulos. The M4 Competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting, 36(1):54–74, Jan. 2020. doi: 10.1016/j.ijforecast.2019.04.014.
  37. 37.S. K. Mitusch, S. W. Funke, and J. S. Dokken. Dolfin-Adjoint 2018.1: Automated adjoints for FEniCS and Firedrake. Journal of Open Source Software, 4(38):1292, June 2019. doi: 10.21105/joss.01292.
  38. 38.J. Močkus. On Bayesian Methods for Seeking the Extremum. In G. I. Marchuk, editor, Optimization Techniques IFIP Technical Conference: Novosibirsk, July 1–7, 1974, Lecture Notes in Computer Science, pages 400–404, Berlin, Heidelberg, 1975. Springer. ISBN 978-3-662-38527-2. doi: 10.1007/978-3-662-38527-2_55.
  39. 39.F. Moukalled, L. Mangani, and M. Darwish. The Finite Volume Method in Computational Fluid Dynamics. Springer, 1 edition, 2016. doi: 10.1007/978-3-319-16874-6.
  40. 40.J. Nocedal and S. J. Wright. Numerical optimization. Springer, 1999.
  41. 41.W. Nowak and A. Guthke. Entropy-based experimental design for optimal model discrimination in the geosciences. Entropy, 18(11), 2016. doi: 10.3390/e18110409.
  42. 42.A. O’Hagan. Curve Fitting and Optimal Design for Prediction. Journal of the Royal Statistical Society: Series B (Methodological), 40(1):1–24, 1978. doi: 10.1111/j.2517-6161.1978.tb01643.x.
  43. 43.R. S. Olson, W. La Cava, P. Orzechowski, R. J. Urbanowicz, and J. H. Moore. PMLB: A large benchmark suite for machine learning evaluation and comparison. BioData Mining, 10(1):36, Dec. 2017. doi: 10.1186/s13040-017-0154-4.
  44. 44.K. Otness, A. Gjoka, J. Bruna, D. Panozzo, B. Peherstorfer, T. Schneider, and D. Zorin. An extensible benchmark suite for learning to simulate physical systems. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021. URL https://openreview.net/forum?id=pY9MHwmrymR.
  45. 45.A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc., 2019. URL http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf.
  46. 46.C. Rackauckas. The essential tools of scientific machine learning (Scientific ML). The Winnower, Aug. 2019. doi: 10.15200/winn.156631.13064.
  47. 47.M. Raissi, P. Perdikaris, and G. E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, Feb. 2019. doi: 10.1016/j.jcp.2018.10.045.
  48. 48.O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation, May 2015.
  49. 49.L. Ruthotto and E. Haber. Deep Neural Networks motivated by Partial Differential Equations. arXiv:1804.04272 [cs, math, stat], Apr. 2018.
  50. 50.J. Snoek, H. Larochelle, and R. P. Adams. Practical bayesian optimization of machine learning algorithms. In Advances in Neural Information Processing Systems, pages 2951–2959. Curran Associates, Inc., 2012.
  51. 51.K. Stachenfeld, D. B. Fielding, D. Kochkov, M. Cranmer, T. Pfaff, J. Godwin, C. Cui, S. Ho, P. Battaglia, and A. Sanchez-Gonzalez. Learned coarse models for efficient turbulence simulation, 2022. URL https://arxiv.org/abs/2112.15275.
  52. 52.J. M. Stone and M. L. Norman. ZEUS-2D: A Radiation Magnetohydrodynamics Code for Astrophysical Flows in Two Space Dimensions. I. The Hydrodynamic Algorithms and Tests. The Astrophysical Journal Supplement, 80:753, June 1992. doi: 10.1086/191680.
  53. 53.A. M. Stuart. Inverse problems: A Bayesian perspective. Acta Numerica, 19:451–559, 2010. doi: 10.1017/S0962492910000061.
  54. 54.D. J. Tait and T. Damoulas. Variational Autoencoding of PDE Inverse Problems. arXiv:2006.15641 [cs, stat], June 2020.
  55. 55.M. Takamoto, T. Pradita, R. Leiteritz, D. MacKinlay, F. Alesiani, D. Pflüger, and M. Niepert. PDEBench: A diverse and comprehensive benchmark for scientific machine learning, 2022. URL https://darus.uni-stuttgart.de/privateurl.xhtml?token=1be27526-348a-40ed-9fd0-c62f588efc01.
  56. 56.A. Tarantola. Inverse Problem Theory and Methods for Model Parameter Estimation. SIAM, Jan. 2005. ISBN 978-0-89871-792-1.
  57. 57.E. F. Toro, M. Spruce, and W. Speares. Restoration of the contact surface in the HLL-Riemann solver. Shock Waves, 4(1):25–34, July 1994. doi: 10.1007/BF01414629.
  58. 58.A. Turing. The chemical basis of morphogenesis. Philosophical Transactions of the Royal Society B, 237:37–72, 1952.
  59. 59.B. van Leer. Towards the Ultimate Conservative Difference Scheme. V. A Second-Order Sequel to Godunov’s Method. Journal of Computational Physics, 32(1):101–136, July 1979. doi: 10.1016/0021-9991(79)90145-1.
  60. 60.P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, İ. Polat, Y. Feng, E. W. Moore, J. VanderPlas, D. Laxalde, J. Perktold, R. Cimrman, I. Henriksen, E. A. Quintero, C. R. Harris, A. M. Archibald, A. H. Ribeiro, F. Pedregosa, P. van Mulbregt, and SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nature Methods, 17:261–272, 2020. doi: 10.1038/s41592-019-0686-2.
  61. 61.R. Wang, K. Kashinath, M. Mustafa, A. Albert, and R. Yu. Towards Physics-informed Deep Learning for Turbulent Flow Prediction. arXiv:1911.08655 [physics, stat], June 2020.
  62. 62.S. Wang, X. Yu, and P. Perdikaris. When and why pinns fail to train: A neural tangent kernel perspective. Journal of Computational Physics, 449:110768, 2022. ISSN 0021-9991. doi: https://doi.org/10.1016/j.jcp.2021.110768. URL https://www.sciencedirect.com/science/article/pii/S002199912100663X.
  63. 63.M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al. The fair guiding principles for scientific data management and stewardship. Scientific data, 3(1):1–9, 2016.
  64. 64.O. Yadan. Hydra - a framework for elegantly configuring complex applications. Github, 2019. URL https://github.com/facebookresearch/hydra.

Citation

MLA
Takamoto, M., et al. “PDEBENCH: An Extensive Benchmark for Scientific Machine Learning”. arXiv, 2022, http://arxiv.org/abs/2210.07182v7.
APA
Takamoto, M., Praditia, T., Leiteritz, R., MacKinlay, D., Alesiani, F., Pflüger, D., & Niepert, M. (2022). PDEBENCH: An Extensive Benchmark for Scientific Machine Learning. arXiv. http://arxiv.org/abs/2210.07182v7
Chicago
Takamoto, M., T. Praditia, R. Leiteritz, et al. 2022. “PDEBENCH: An Extensive Benchmark for Scientific Machine Learning”. arXiv. http://arxiv.org/abs/2210.07182v7.
Harvard
Takamoto, M. et al. (2022) “PDEBENCH: An Extensive Benchmark for Scientific Machine Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.07182v7.
Vancouver
1. Takamoto M, Praditia T, Leiteritz R, MacKinlay D, Alesiani F, Pflüger D, Niepert M (2022) PDEBENCH: An Extensive Benchmark for Scientific Machine Learning. arXiv

BibTeX

@article{takamoto2022pdebench,
  title = {PDEBENCH: An Extensive Benchmark for Scientific Machine Learning},
  author = {Takamoto, Makoto and Praditia, Timothy and Leiteritz, Raphael and MacKinlay, Dan and Alesiani, Francesco and Pflüger, Dirk and Niepert, Mathias},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.07182v7},
  eprint = {2210.07182}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors