PDEBench: An Extensive Benchmark for Scientific Machine Learning
Makoto TakamotoTimothy PraditiaRaphael LeiteritzDaniel MacKinlayFrancesco AlesianiDirk PflügerMathias Niepert
Presents PDEBench, an extensive benchmark suite featuring diverse time-dependent physical simulations, large ready-to-use datasets, and standardized baselines to systematically evaluate and compare scientific machine learning models against classical numerical methods.
Data-driven scientific machine learning holds substantial promise for modeling complex physical systems governed by partial differential equations, such as fluid dynamics, environmental contaminant transport, and wave propagation. Traditional numerical solvers are computationally expensive and struggle with non-differentiable formulations, which complicates optimization, sensitivity analysis, and the resolution of inverse problems. However, progress in this field has been severely limited by the lack of standardized, large-scale, and challenging benchmark datasets that evaluate whether neural surrogate models accurately capture real-world physics.
The article introduces and evaluates PDEBENCH, an open-source, extensible benchmark suite designed to assess machine learning models against classical numerical simulations across diverse scientific modeling tasks. The primary objective is to provide a rigorous, standardized platform comprising extensive datasets, reproducible baseline models, and physics-informed evaluation metrics to identify key challenges and guide algorithm development in scientific machine learning.
The benchmark encompasses 35 datasets across 11 distinct partial differential equation systems spanning one-, two-, and three-dimensional domains. These include stylized baseline models, such as advection and reaction-diffusion, as well as complex scenarios featuring compressible and incompressible fluid flows, shallow-water waves, and groundwater contaminant transport. Datasets incorporate realistic boundary conditions and parameter variations, such as varying viscosities and Mach numbers. Using standard deep learning frameworks, the authors evaluated prominent machine learning architectures—specifically the Fourier Neural Operator, U-Net, and Physics-Informed Neural Networks—on both forward time-stepping emulation and gradient-based inverse inference for estimating unknown initial conditions.
The experimental findings demonstrate critical insights into current model performance. First, the Fourier Neural Operator consistently outperformed other architectures across most forward and inverse tasks, maintaining stable predictions across the frequency spectrum. Second, standard machine learning error metrics proved inadequate; global root-mean-squared error masked critical local physical errors in turbulent regimes and across discontinuities. Third, machine learning models delivered massive computational speedups during deployment—generating predictions up to three orders of magnitude faster than classical solvers—while bypassing classical time-stepping stability restrictions. Finally, significant performance degradations occurred during temporal extrapolation beyond trained time horizons, as well as in regimes governed by high Reynolds numbers or sharp shock waves.
These results demonstrate that machine learning emulators can drastically reduce operational compute costs for physical simulations, making real-time design and optimization viable. However, the findings also show that current models struggle with generalization over longer horizons and non-smooth physical regimes. Relying on global aggregate errors introduces operational risk because a model may appear statistically accurate while severely violating physical conservation laws or boundary dynamics.
Organizations developing or deploying scientific machine learning tools should adopt multi-faceted, physics-informed metrics rather than relying solely on standard error metrics. Decision-makers should prioritize architectures that operate effectively across frequency spaces, such as neural operators, for surrogate modeling workflows. Additionally, future research must focus on stabilizing autoregressive predictions over extended time horizons, improving performance on high-frequency shock dynamics, and expanding benchmarks to multi-phase flows and complex geometric domains.
- Paper: Fourier Neural Operator for Parametric Partial Differential Equations, Zongyi Li et al. (2020). Introduces the Fourier Neural Operator architecture, which serves as a primary baseline and the top-performing model evaluated across PDEBench.
- Paper: DeepXDE: A Deep Learning Library for Solving Differential Equations, Lu Lu et al. (2019). Establishes standard methodologies and software implementations for physics-informed neural network baselines used throughout PDEBench for forward and inverse modeling.
- Paper: Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Lu Lu et al. (2021). Formulates neural operator learning between infinite-dimensional function spaces, providing the theoretical foundation for surrogate modeling frameworks benchmarked in PDEBench.
- Paper: Characterizing possible failure modes in physics-informed neural networks, Aditi S. Krishnapriyan et al. (2021). Identifies foundational failure modes and spectral optimization limitations in physics-informed neural networks that PDEBench's multi-regime datasets explicitly evaluate.
- Paper: When and why PINNs fail to train: A neural tangent kernel perspective, Sifan Wang et al. (2020). Provides a theoretical explanation for the spectral bias and training difficulties of physics-informed models when handling the high-frequency dynamics benchmarked in the study.
- Paper: Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data, Anuj Karpatne et al. (2016). Defines the core principles of theory-guided data science and physics-informed metrics that motivate PDEBench's holistic evaluation suite.
- Paper: Convolutional Neural Operators for robust and accurate learning of PDEs, Bogdan Raonic et al. (2023). Develops Convolutional Neural Operators to address the aliasing, out-of-distribution, and high-frequency shock degradation issues exposed by PDEBench's baseline evaluations.
- Paper: Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs, Nikola Kovachki et al. (2023). Synthesizes the broader mathematical foundations and comprehensive empirical properties of neural operators across diverse partial differential equation systems.
- Paper: Learning Operators with Coupled Attention, Georgios Kissas et al. (2022). Introduces an attention-coupled kernel operator architecture designed to overcome the data-efficiency and extrapolation limits identified in standard operator benchmarks.
- Paper: AirfRANS: High Fidelity Computational Fluid Dynamics Dataset for Approximating Reynolds-Averaged Navier-Stokes Solutions, Florent Bonnet et al. (2022). Extends standardized benchmarking to high-fidelity aerodynamic fluid flows on irregular computational meshes using Reynolds-averaged Navier-Stokes equations.
- Paper: Neural Stochastic PDEs: Resolution-Invariant Learning of Continuous Spatiotemporal Dynamics, Cristopher Salvi et al. (2022). Generalizes resolution-invariant operator learning to continuous spatiotemporal systems driven by external stochastic disturbances.
- Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). Combines neural operators with generative diffusion models to enable resolution-invariant probabilistic sampling and inverse problem solving in infinite-dimensional function spaces.
- Paper: Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors, Alexander Scheinker (2026). Proposes a bidirectional self-supervised error estimation technique to address the compounding rollout error and temporal extrapolation degradation reported in PDE surrogate models.
