Convolutional Neural Operators for robust and accurate learning of PDEs
Bogdan RaonicRoberto MolinaroTim De RyckTobias RohnerFrancesca BartolucciRima AlaifariSiddhartha MishraEmmanuel de Bézenac
Develops continuous-discrete equivalent convolutional neural operators that overcome aliasing errors in traditional CNNs, proving their universal approximation capability for partial differential equations and outperforming existing operator learning baselines across diverse multiscale benchmarks.
Simulating complex physical systems described by partial differential equations is critical across engineering and the physical sciences, yet traditional numerical methods often incur prohibitive computational costs in multi-query settings like optimization and uncertainty quantification. While machine learning surrogates have emerged to address this bottleneck, conventional convolutional neural networks suffer from severe aliasing errors across varying grid resolutions, causing existing models to rely heavily on Fourier-based or fully connected architectures that face trade-offs in computational cost, expressivity, and spatial localization.
The article develops and evaluates Convolutional Neural Operators (CNOs), a new continuous-discrete equivalent framework that adapts convolutional networks to accurately, robustly, and efficiently learn mapping operators for partial differential equations.
The authors design CNOs as an operator adaptation of the U-Net architecture that maps between spaces of bandlimited functions, utilizing continuous-discrete equivalent filters alongside upsampled and downsampled activation layers to eliminate aliasing. To validate the method, they prove a universal approximation theorem and evaluate CNO against leading benchmarks—including Fourier Neural Operators (FNO), DeepONet, Galerkin Transformers, standard U-Nets, and Residual Networks—across a newly designed Representative PDE Benchmark suite spanning seven diverse linear and nonlinear two-dimensional systems (such as Poisson, Wave, Transport, Navier-Stokes, Darcy Flow, and Compressible Euler equations).
The evaluation yields several key findings: First, CNO achieved the lowest in-distribution error on six of the seven benchmarks, outperforming FNO by nearly a factor of 20 on the Poisson problem (0.21% vs 4.98% relative median error) and maintaining superior performance on complex fluid flows. Second, CNO demonstrated exceptional zero-shot out-of-distribution generalization, maintaining test errors below 1.2% across most test cases, whereas baselines like Galerkin Transformers and standard networks suffered catastrophic error spikes (exceeding 100%) on spatial transport tasks due to lacking translation equivariance. Third, CNO demonstrated strict resolution invariance, keeping test errors steady across various grid resolutions while FNO errors varied by up to 25% and standard U-Net errors increased by up to a factor of three. Finally, CNO exhibited superior computational and data efficiency, training nearly twice as fast as the best-performing FNO model and requiring roughly four times fewer data samples (about 14,300 versus 60,100 samples) to achieve a 1% error target on Navier-Stokes simulations.
These findings show that CNO bridges the longstanding gap between standard vision architectures and functional operator learning, delivering a localized, alias-free surrogate model that significantly reduces the computational overhead and data requirements for physics-based modeling. For decision-makers, adopting CNO can lower surrogate training budgets, accelerate engineering design loops, and mitigate the operational risk of model degradation across different grid resolutions.
Organizations developing scientific machine learning workflows should consider piloting CNO architectures as primary surrogate models for complex two-dimensional partial differential equations, particularly when spatial translation equivariance and multi-scale resolution handling are critical. Future technical roadmaps should focus on extending CNO frameworks to three-dimensional domains, irregular geometries via coordinate transformations, and time-dependent trajectories using autoregressive techniques.
Confidence in these findings is high for two-dimensional Cartesian domains given the rigorous mathematical proofs and comprehensive empirical benchmarking against leading baselines. However, users should exercise caution when extending the architecture directly to complex irregular geometries or three-dimensional settings without the necessary domain transformations and hardware memory considerations.
- Paper: Fourier Neural Operator for Parametric Partial Differential Equations, Zongyi Li et al. (2020). Introduces Fourier Neural Operators (FNO) for learning discretization-invariant solution mappings of PDEs, establishing the primary baseline and foundational operator learning framework that CNOs seek to improve upon.
- Paper: Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Lu Lu et al. (2021). Establishes the universal operator approximation theorem and DeepONet architecture for mapping between infinite-dimensional function spaces, providing core theoretical grounding for neural operator design.
- Paper: A guide to convolution arithmetic for deep learning, Vincent Dumoulin et al. (2016). Provides a comprehensive mathematical formulation of discrete convolution arithmetic, strides, and padding essential for designing continuous and resolution-consistent convolutional architectures.
- Paper: Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs, Nikola Kovachki et al. (2023). Provides a broader, comprehensive unifying framework and mathematical theory for learning maps between infinite-dimensional function spaces across diverse neural operator parameterizations.
- Paper: Neural means and kernel corrections for operator learning, Yitzchak Shmalo (2026). Extends deep neural operator frameworks by pairing neural mean estimators with exact kernel ridge regression corrections to enhance accuracy and uncertainty quantification.
- Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). Presents Kolmogorov-Arnold Networks as an alternative paradigm for function approximation and PDE solving, providing an interpretable architecture that moves beyond standard convolutional and multi-layer operator formulations.
