Convolutional Neural Operators for robust and accurate learning of PDEs

Bogdan RaonicRoberto MolinaroTim De RyckTobias RohnerFrancesca BartolucciRima AlaifariSiddhartha MishraEmmanuel de Bézenac

article2023NeurIPS192 citations

Develops continuous-discrete equivalent convolutional neural operators that overcome aliasing errors in traditional CNNs, proving their universal approximation capability for partial differential equations and outperforming existing operator learning baselines across diverse multiscale benchmarks.

Listen

Simulating complex physical systems described by partial differential equations is critical across engineering and the physical sciences, yet traditional numerical methods often incur prohibitive computational costs in multi-query settings like optimization and uncertainty quantification. While machine learning surrogates have emerged to address this bottleneck, conventional convolutional neural networks suffer from severe aliasing errors across varying grid resolutions, causing existing models to rely heavily on Fourier-based or fully connected architectures that face trade-offs in computational cost, expressivity, and spatial localization.

The article develops and evaluates Convolutional Neural Operators (CNOs), a new continuous-discrete equivalent framework that adapts convolutional networks to accurately, robustly, and efficiently learn mapping operators for partial differential equations.

The authors design CNOs as an operator adaptation of the U-Net architecture that maps between spaces of bandlimited functions, utilizing continuous-discrete equivalent filters alongside upsampled and downsampled activation layers to eliminate aliasing. To validate the method, they prove a universal approximation theorem and evaluate CNO against leading benchmarks—including Fourier Neural Operators (FNO), DeepONet, Galerkin Transformers, standard U-Nets, and Residual Networks—across a newly designed Representative PDE Benchmark suite spanning seven diverse linear and nonlinear two-dimensional systems (such as Poisson, Wave, Transport, Navier-Stokes, Darcy Flow, and Compressible Euler equations).

The evaluation yields several key findings: First, CNO achieved the lowest in-distribution error on six of the seven benchmarks, outperforming FNO by nearly a factor of 20 on the Poisson problem (0.21% vs 4.98% relative median error) and maintaining superior performance on complex fluid flows. Second, CNO demonstrated exceptional zero-shot out-of-distribution generalization, maintaining test errors below 1.2% across most test cases, whereas baselines like Galerkin Transformers and standard networks suffered catastrophic error spikes (exceeding 100%) on spatial transport tasks due to lacking translation equivariance. Third, CNO demonstrated strict resolution invariance, keeping test errors steady across various grid resolutions while FNO errors varied by up to 25% and standard U-Net errors increased by up to a factor of three. Finally, CNO exhibited superior computational and data efficiency, training nearly twice as fast as the best-performing FNO model and requiring roughly four times fewer data samples (about 14,300 versus 60,100 samples) to achieve a 1% error target on Navier-Stokes simulations.

These findings show that CNO bridges the longstanding gap between standard vision architectures and functional operator learning, delivering a localized, alias-free surrogate model that significantly reduces the computational overhead and data requirements for physics-based modeling. For decision-makers, adopting CNO can lower surrogate training budgets, accelerate engineering design loops, and mitigate the operational risk of model degradation across different grid resolutions.

Organizations developing scientific machine learning workflows should consider piloting CNO architectures as primary surrogate models for complex two-dimensional partial differential equations, particularly when spatial translation equivariance and multi-scale resolution handling are critical. Future technical roadmaps should focus on extending CNO frameworks to three-dimensional domains, irregular geometries via coordinate transformations, and time-dependent trajectories using autoregressive techniques.

Confidence in these findings is high for two-dimensional Cartesian domains given the rigorous mathematical proofs and comprehensive empirical benchmarking against leading baselines. However, users should exercise caution when extending the architecture directly to complex irregular geometries or three-dimensional settings without the necessary domain transformations and hardware memory considerations.

  • Paper: Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs, Nikola Kovachki et al. (2023). Provides a broader, comprehensive unifying framework and mathematical theory for learning maps between infinite-dimensional function spaces across diverse neural operator parameterizations.
  • Paper: Neural means and kernel corrections for operator learning, Yitzchak Shmalo (2026). Extends deep neural operator frameworks by pairing neural mean estimators with exact kernel ridge regression corrections to enhance accuracy and uncertainty quantification.
  • Paper: KAN: Kolmogorov-Arnold Networks, Ziming Liu et al. (2025). Presents Kolmogorov-Arnold Networks as an alternative paradigm for function approximation and PDE solving, providing an interpretable architecture that moves beyond standard convolutional and multi-layer operator formulations.
Cover for Convolutional Neural Operators for robust and accurate learning of PDEs

Abstract

Although very successfully used in conventional machine learning, convolution based neural network architectures – believed to be inconsistent in function space – have been largely ignored in the context of learning solution operators of PDEs. Here, we present novel adaptations for convolutional neural networks to demonstrate that they are indeed able to process functions as inputs and outputs. The resulting architecture, termed as convolutional neural operators (CNOs), is designed specifically to preserve its underlying continuous nature, even when implemented in a discretized form on a computer. We prove a universality theorem to show that CNOs can approximate operators arising in PDEs to desired accuracy. CNOs are tested on a novel suite of benchmarks, encompassing a diverse set of PDEs with possibly multi-scale solutions and are observed to significantly outperform baselines, paving the way for an alternative framework for robust and accurate operator learning.

Table of Contents

  • 1 Introduction.
  • 2 Convolutional Neural Operators.
  • 3 Universal Approximation by CNOs.
  • 4 Experiments.
  • 5 Discussion.
  • References

Knowls

  1. Knowl 1 — Convolutional Neural Operator Architecture

    model/method

    The Convolutional Neural Operator (CNO) is an operator learning framework designed to map between function spaces of bandlimited functions Bw(D,RdX)→Bw(D,RdY)B_w(D, \mathbb{R}^{d_X}) \to B_w(D, \mathbb{R}^{d_Y}) on a domain D=T2D = \mathbb{T}^2 (the 2D torus) with bandlimit w>0w > 0, where

    Bw(D)={f∈L2(D):supp(f^)⊆[−w,w]2}.B_w(D) = \left\{f \in L^2(D) : \text{supp}(\hat{f}) \subseteq [-w, w]^2\right\}.

    The CNO mapping G:Bw(D,RdX)→Bw(D,RdY)\mathcal{G} : B_w(D, \mathbb{R}^{d_X}) \to B_w(D, \mathbb{R}^{d_Y}) is defined as an operator composition:

    G:u↦P(u)=v0↦v1↦⋯↦vL↦Q(vL)=uˉ,\mathcal{G} : u \mapsto \mathcal{P}(u) = v_0 \mapsto v_1 \mapsto \dots \mapsto v_L \mapsto \mathcal{Q}(v_L) = \bar{u},

    where intermediate function representations evolve as

    vl+1=Pl∘Σl∘Kl(vl),1≤l≤L−1.v_{l+1} = \mathcal{P}_l \circ \Sigma_l \circ \mathcal{K}_l(v_l), \quad 1 \le l \le L-1.

    Here, the input function uu is first mapped to a latent bandlimited space of d0>dXd_0 > d_X channels via a local convolution lifting operator P:Bw(D,RdX)→Bw(D,Rd0)\mathcal{P} : B_w(D, \mathbb{R}^{d_X}) \to B_w(D, \mathbb{R}^{d_0}). Each intermediate layer applies a spatial convolution operator Kl\mathcal{K}_l, a modulated bandlimited activation operator Σl\Sigma_l, and a sampling operator Pl\mathcal{P}_l (which performs upsampling U\mathcal{U} or downsampling D\mathcal{D}). The output function vL∈Bw(D,RdL)v_L \in B_w(D, \mathbb{R}^{d_L}) is mapped to the target output space by a projection convolution operator Q:Bw(D,RdL)→Bw(D,RdY)\mathcal{Q} : B_w(D, \mathbb{R}^{d_L}) \to B_w(D, \mathbb{R}^{d_Y}).

    The full architecture is instantiated as an Operator U-Net comprising:

    • Encoder blocks where functions are downsampled in spatial bandlimit while channel width is increased;
    • Decoder blocks where functions are upsampled in spatial bandlimit while channel width is reduced;
    • ResNet blocks Rw,wˉ(v)=v+Kw∘Σw,wˉ∘Kw(v)\mathcal{R}_{w, \bar{w}}(v) = v + \mathcal{K}_w \circ \Sigma_{w, \bar{w}} \circ \mathcal{K}_w(v) and Invariant blocks Iw,wˉ(v)=Σw,wˉ∘Kw(v)\mathcal{I}_{w, \bar{w}}(v) = \Sigma_{w, \bar{w}} \circ \mathcal{K}_w(v) connecting corresponding encoder and decoder levels;
    • Skip connections operating via channel-concatenation without altering spatial resolution.
  2. Knowl 2 — Elementary Bandlimited Operators for Continuous-Discrete Equivalence

    model/method

    To preserve continuous-discrete equivalence (CDE) and prevent aliasing errors when operating on bandlimited function spaces Bw(D)={f∈L2(D):supp(f^)⊆[−w,w]2}B_w(D) = \{f \in L^2(D) : \text{supp}(\hat{f}) \subseteq [-w, w]^2\} on D=T2D = \mathbb{T}^2, CNO implements the following elementary operators:

    1. Physical-Space Convolution Operator: For a kernel size k∈Nk \in \mathbb{N}, discrete weights kij∈Rk_{ij} \in \mathbb{R}, and uniform grid points zijz_{ij} with spacing ≤1/(2w)\le 1/(2w) satisfying the Whittaker-Shannon-Kotelnikov sampling condition, the continuous convolution Kw:Bw(D)→Bw(D)\mathcal{K}_w : B_w(D) \to B_w(D) is defined by:

    Kwf(x)=(Kw⋆f)(x)=∫DKw(x−y)f(y) dy=∑i,j=1kkijf(x−zij),∀x∈D.\mathcal{K}_w f(x) = (K_w \star f)(x) = \int_D K_w(x - y)f(y)\,dy = \sum_{i,j=1}^k k_{ij} f(x - z_{ij}), \quad \forall x \in D.

    1. Upsampling Operator: For target bandlimit wˉ>w\bar{w} > w, the operator Uw,wˉ:Bw(D)→Bwˉ(D)\mathcal{U}_{w, \bar{w}} : B_w(D) \to B_{\bar{w}}(D) embeds the function into a higher bandlimit space:

    Uw,wˉf(x)=f(x),∀x∈D.\mathcal{U}_{w, \bar{w}} f(x) = f(x), \quad \forall x \in D.

    1. Downsampling Operator: For target bandlimit w‾<w\underline{w} < w, the operator Dw,w‾:Bw(D)→Bw‾(D)\mathcal{D}_{w, \underline{w}} : B_w(D) \to B_{\underline{w}}(D) is defined via convolution with a 2D sinc interpolation filter hw‾(x0,x1)=sinc(2w‾x0)⋅sinc(2w‾x1)h_{\underline{w}}(x_0, x_1) = \text{sinc}(2\underline{w}x_0) \cdot \text{sinc}(2\underline{w}x_1):

    Dw,w‾f(x)=(w‾w)2(hw‾⋆f)(x)=(w‾w)2∫Dhw‾(x−y)f(y) dy,∀x∈D.\mathcal{D}_{w, \underline{w}} f(x) = \left(\frac{\underline{w}}{w}\right)^2 (h_{\underline{w}} \star f)(x) = \left(\frac{\underline{w}}{w}\right)^2 \int_D h_{\underline{w}}(x - y)f(y)\,dy, \quad \forall x \in D.

    1. Alias-Free Activation Layer: Pointwise evaluation of a non-linear activation σ\sigma generates frequencies beyond bandlimit ww. To enforce exact bandlimits, the activation operator Σw,wˉ:Bw(D)→Bw(D)\Sigma_{w, \bar{w}} : B_w(D) \to B_w(D) upsamples ff to a sufficiently large intermediate bandlimit wˉ>w\bar{w} > w such that σ(Bw)⊂Bwˉ\sigma(B_w) \subset B_{\bar{w}}, applies σ\sigma, and downsamples back to ww:

    Σw,wˉf(x)=Dwˉ,w(σ∘Uw,wˉf)(x),∀x∈D.\Sigma_{w, \bar{w}} f(x) = \mathcal{D}_{\bar{w}, w}\left(\sigma \circ \mathcal{U}_{w, \bar{w}} f\right)(x), \quad \forall x \in D.

  3. Knowl 3 — Representation Equivalence of Convolutional Neural Operators

    theoretical result

    Let D=T2D = \mathbb{T}^2, w>0w > 0, and let G:Bw(D,RdX)→Bw(D,RdY)\mathcal{G} : B_w(D, \mathbb{R}^{d_X}) \to B_w(D, \mathbb{R}^{d_Y}) be a Convolutional Neural Operator constructed from the elementary bandlimited operations (convolution Kw\mathcal{K}_w, upsampling Uw,wˉ\mathcal{U}_{w, \bar{w}}, downsampling Dw,w‾\mathcal{D}_{w, \underline{w}}, and modulated activation Σw,wˉ\Sigma_{w, \bar{w}}).

    Then G\mathcal{G} is a Representation Equivalent Neural Operator (ReNO).

    As a consequence of representation equivalence, the continuous operator mapping between spaces of bandlimited functions commutes exactly with discrete representations obtained by sampling on uniform grids of spacing ≤1/(2w)\le 1/(2w) and reconstructing via sinc interpolation. This continuous-discrete equivalence guarantees that the discrete execution of CNO on a computer is free from aliasing errors and evaluates consistently across arbitrary grid resolutions.

  4. Knowl 4 — Universal Approximation of PDE Solution Operators by CNOs

    theoretical result

    Consider the abstract boundary value problem on the domain D=T2D = \mathbb{T}^2:

    L(u)=0,B(u)=0,\mathcal{L}(u) = 0, \quad \mathcal{B}(u) = 0,

    where L\mathcal{L} is a differential operator depending on spatial coordinate xx through a coefficient function a∈Hr(D)a \in H^r(D), and B\mathcal{B} is a boundary operator. Let G†:X∗⊂Hr(D)→Hr(D)\mathcal{G}^\dagger : \mathcal{X}^* \subset H^r(D) \to H^r(D) denote the continuous solution operator a↦ua \mapsto u.

    Assume there exists a modulus of continuity for G†\mathcal{G}^\dagger such that

    ∥G†(a)−G†(a′)∥Lp(D)≤ω(∥a−a′∥Hσ(D)),∀a,a′∈X∗,\|\mathcal{G}^\dagger(a) - \mathcal{G}^\dagger(a')\|_{L^p(D)} \le \omega\left(\|a - a'\|_{H^\sigma(D)}\right), \quad \forall a, a' \in \mathcal{X}^*,

    where p∈{2,∞}p \in \{2, \infty\}, σ∈N0\sigma \in \mathbb{N}_0 with 0≤σ≤r−10 \le \sigma \le r - 1, and ω:[0,∞)→[0,∞)\omega : [0, \infty) \to [0, \infty) is a monotonically increasing function satisfying lim⁡y→0ω(y)=0\lim_{y \to 0} \omega(y) = 0.

    Let r>max⁡{σ,2/p}r > \max\{\sigma, 2/p\} and B>0B > 0. For every tolerance ε>0\varepsilon > 0, there exists a Convolutional Neural Operator G:Bw(D)→Bw(D)\mathcal{G} : B_w(D) \to B_w(D) such that for every coefficient a∈X∗a \in \mathcal{X}^* satisfying ∥a∥Hr(D)≤B\|a\|_{H^r(D)} \le B, the approximation error satisfies:

    ∥G†(a)−G(a)∥Lp(D)<ε.\|\mathcal{G}^\dagger(a) - \mathcal{G}(a)\|_{L^p(D)} < \varepsilon.

  5. Knowl 5 — Benchmark Performance Comparison Across Representative PDE Benchmarks

    data/table

    The relative median L1L^1 test errors for both in-distribution (In) and out-of-distribution (Out) evaluation across the Representative PDE Benchmarks (RPB) for Feedforward Neural Networks with residual connections (FFNN), Galerkin Transformer (GT), U-Net, ResNet, DeepONet (DON), Fourier Neural Operator (FNO), and Convolutional Neural Operator (CNO):

    Benchmark In/Out FFNN GT UNet ResNet DON FNO CNO
    Poisson Equation In 5.74% 2.77% 0.71% 0.43% 12.92% 4.98% 0.21%
    Out 5.35% 2.84% 1.27% 1.10% 9.15% 7.05% 0.27%
    Wave Equation In 2.51% 1.44% 1.51% 0.79% 2.26% 1.02% 0.63%
    Out 3.01% 1.79% 2.03% 1.36% 2.83% 1.77% 1.17%
    Smooth Transport In 7.09% 0.98% 0.49% 0.39% 1.14% 0.28% 0.24%
    Out 650.6% 875.4% 1.28% 0.96% 157.2% 3.90% 0.46%
    Discontinuous Transport In 13.0% 1.55% 1.31% 1.01% 5.78% 1.15% 1.01%
    Out 257.3% 22691.1% 1.35% 1.16% 117.1% 2.89% 1.09%
    Allen-Cahn Equation In 18.27% 0.77% 0.82% 1.40% 13.63% 0.28% 0.54%
    Out 46.93% 2.90% 2.18% 3.74% 19.86% 1.10% 2.23%
    Navier-Stokes Equations In 8.05% 4.14% 3.54% 3.69% 11.64% 3.57% 2.76%
    Out 16.12% 11.09% 10.93% 9.68% 15.05% 9.58% 7.04%
    Darcy Flow In 2.14% 0.86% 0.54% 0.42% 1.13% 0.80% 0.38%
    Out 2.23% 1.17% 0.64% 0.60% 1.61% 1.11% 0.50%
    Compressible Euler In 0.78% 2.09% 0.38% 1.70% 1.93% 0.44% 0.35%
    Out 1.34% 2.94% 0.76% 2.06% 2.88% 0.69% 0.59%

    CNO outperforms all baselines on 7 of the 8 benchmarks in both in-distribution and out-of-distribution regimes, trailing only FNO on the Allen-Cahn equation. On the Poisson benchmark, CNO outperforms FNO by a factor of over 20 in-distribution (0.21% vs 4.98%) and over 26 out-of-distribution (0.27% vs 7.05%). On transport tasks, architectures lacking translation equivariance (FFNN, GT, DON) experience extreme out-of-distribution errors (up to 22691.1% for GT on discontinuous transport), whereas CNO maintains errors of 0.46% and 1.09%.

  6. Knowl 6 — Representative PDE Benchmarks Specification

    experimental setup

    The Representative PDE Benchmarks (RPB) consist of eight 2D PDE problems on Cartesian domains designed to evaluate multiscale resolution, generalization, and computational fidelity:

    1. Poisson Equation: −Δu=f-\Delta u = f in D=(0,1)2D = (0, 1)^2 with u∣∂D=0u|_{\partial D} = 0 and source term f(x,y)=πK2∑i,j=1Kaij(i2+j2)0.5sin⁡(πix)sin⁡(πjy)f(x, y) = \frac{\pi}{K^2} \sum_{i,j=1}^K a_{ij} (i^2 + j^2)^{0.5} \sin(\pi i x) \sin(\pi j y) where aij∼U[−1,1]a_{ij} \sim \mathcal{U}[-1, 1]. K=16K=16 for training and in-distribution testing; K=20K=20 for out-of-distribution (OOD) testing to evaluate high-frequency extrapolation.

    2. Wave Equation: utt−c2Δu=0u_{tt} - c^2 \Delta u = 0 (c=0.1c = 0.1) on D×(0,T)D \times (0, T) with initial standing wave f(x,y)f(x, y) decaying as (i2+j2)−r(i^2 + j^2)^{-r}. In-distribution setup uses T=5,K=24,r=1T=5, K=24, r=1; OOD setup uses K=32,r=0.85K=32, r=0.85.

    3. Smooth and Discontinuous Transport Equations: ut+v⋅∇u=0u_t + v \cdot \nabla u = 0 with constant velocity v=(0.2,0.2)v = (0.2, 0.2) evaluated at T=1T=1. In-distribution training uses Gaussian blobs (smooth) and disk indicators (discontinuous) with centers in (0.2,0.4)2(0.2, 0.4)^2. OOD testing shifts initial centers to (0.4,0.6)2(0.4, 0.6)^2 to test translation equivariance.

    4. Allen-Cahn Equation: ut=Δu−ε2u(u2−1)u_t = \Delta u - \varepsilon^2 u(u^2 - 1) with reaction rate ε=220\varepsilon = 220 at T=0.0002T = 0.0002. Training uses K=24,r=1K=24, r=1; OOD testing sets K=16K=16 with r∼U[0.85,1.15]r \sim \mathcal{U}[0.85, 1.15].

    5. Navier-Stokes Equations: Incompressible 2D fluid flow ut+(u⋅∇)u+∇p=νΔuu_t + (u \cdot \nabla)u + \nabla p = \nu \Delta u, div u=0\text{div } u = 0 on T2\mathbb{T}^2 with high-Reynolds spectral viscosity ν=4×10−4\nu = 4 \times 10^{-4} applied to modes ≥12\ge 12. Initial condition is a thin shear layer with thickness ρ=0.1\rho = 0.1 and 10 perturbation modes; OOD testing reduces thickness to ρ=0.09\rho = 0.09 and shifts layers vertically.

    6. Darcy Flow: −∇⋅(a∇u)=1-\nabla \cdot (a \nabla u) = 1 with u∣∂D=0u|_{\partial D} = 0. Permeability aa is drawn from a thresholded Gaussian process with length scale l=0.1l = 0.1 (in-distribution) and l=0.05l = 0.05 (OOD).

    7. Flow Past Airfoils (Compressible Euler): Steady-state transonic density fields past RAE2822 airfoils perturbed with 20 Hicks-Henne bump functions (in-distribution) and 30 bump functions (OOD).

  7. Knowl 7 — Resolution Invariance and Aliasing Error Suppression

    empirical result

    Evaluation of CNO, FNO, and standard U-Net on the 2D Navier-Stokes thin shear layer problem across grid resolutions ranging from 32232^2 to 1282128^2 (trained at 64264^2) establishes the following:

    1. Fourier Spectrum Fidelity: The exact solution spectrum exhibits a rich multi-scale vortex distribution. CNO matches the true decay rate across high and low frequency modes in Fourier space. In contrast, FNO artificially amplifies frequencies along the horizontal spectrum axis due to aliasing errors, while standard U-Net exhibits high-frequency amplification.

    2. Resolution Invariance: CNO test error remains constant and flat when evaluated at sub-resolutions (32232^2) and super-resolutions (962,128296^2, 128^2). In contrast, FNO test error increases by up to 25% on lower resolutions and by 10% on higher resolutions, demonstrating that FNO is not resolution invariant for multiscale solutions. Standard U-Net exhibits an error increase of up to a factor of 3 across tested resolutions.

  8. Knowl 8 — Data and Computational Scaling Laws for CNO

    empirical result

    On the 2D Navier-Stokes benchmark, test error EE scales with the number of training samples NN according to the power law

    E=(N0N)−r,E = \left(\frac{N_0}{N}\right)^{-r},

    where rr is the convergence rate and N0N_0 is the normalized sample count required to achieve a 1% test error:

    • CNO achieves a convergence rate of r=0.37r = 0.37 and requires N0≈14.3KN_0 \approx 14.3\text{K} training samples to reach 1% error.
    • FNO achieves a rate of r=0.28r = 0.28 and requires N0≈60.1KN_0 \approx 60.1\text{K} samples (over 4×4\times more data than CNO).
    • Galerkin Transformer (GT) achieves r=0.27r = 0.27 and requires N0≈133.2KN_0 \approx 133.2\text{K} samples (nearly 10×10\times more data than CNO).
    • Model capacity acts as a scaling bottleneck: a smaller CNO model (0.82M parameters) scales with a shallower exponent r=0.30r = 0.30 compared to the 7.8M parameter CNO model (r=0.37r = 0.37).

    In computational efficiency, for a given model parameter count, CNO achieves lower validation error than FNO. The best-performing CNO model trains nearly twice as fast per epoch as the best-performing FNO model while attaining lower test error.

  9. Knowl 9 — Dimensionality and Domain Geometry Limitations of CNO

    limitation

    The Convolutional Neural Operator architecture and its empirical validations are formulated for two-dimensional Cartesian domains (D=T2D = \mathbb{T}^2 or rectangular subsets). Extending CNO to three spatial dimensions is computationally and memory demanding due to 3D continuous sinc filtering and multi-channel convolutions. Furthermore, handling irregular or non-Cartesian geometries is not natively supported by the spatial convolution operations and requires mapping the irregular domain onto a Cartesian reference domain via learned coordinate transformations or diffeomorphisms.

Coverage note — Omitted material includes intermediate technical lemmas in the supplementary material (SM A, SM B), detailed multi-channel index expansions, and individual baseline hyperparameter search ranges in SM C.

References

  1. 1.I. Ayed, E. de Bézenac, A. Pajot, J. Brajard, and P. Gallinari. Learning dynamical systems from partial observations. CoRR, abs/1902.11136, 2019.
  2. 2.F. Bartolucci, E. de Bézenac, B. Raonic, R. Molinaro, S. Mishra, and R. Alaifari. Are neural operators really neural operators ? frame theory meets operator learning. Technical Report 2023-21, Seminar for Applied Mathematics, ETH Zürich, Switzerland, 2023.
  3. 3.J. Bell, P. Collela, and H. M. Glaz. A second-order projection method for the incompressible Navier-Stokes equations. J. Comput. Phys., 85:257–283, 1989.
  4. 4.M. Bertero, P. Bocacci, and C. De Mol. Introduction to inverse problems in imaging. CRC press, 2021.
  5. 5.K. Bhattacharya, B. Hosseini, N. B. Kovachki, and A. M. Stuart. Model Reduction And Neural Networks For Parametric PDEs. The SMAI journal of computational mathematics, 7:121–157, 2021.
  6. 6.N. Boullé, Y. Nakatsukasa, and A. Townsend. Rational neural networks. Advances in Neural Information Processing Systems, 33:14243–14253, 2020.
  7. 7.S. Cai, Z. Wang, L. Lu, T. A. Zaki, and G. E. Karniadakis. DeepM&Mnet: Inferring the electroconvection multiphysics fields based on operator approximation by neural networks. Journal of Computational Physics, 436:110296, 2021.
  8. 8.S. Cao. Choose a transformer: Fourier or galerkin. In 35th conference on neural information processing systems, 2021.
  9. 9.T. Chen and H. Chen. Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems. IEEE Transactions on Neural Networks, 6(4):911–917, 1995.
  10. 10.T. De Ryck, S. Lanthaler, and S. Mishra. On the approximation of functions by tanh neural networks. Neural Networks, 2021.
  11. 11.T. De Ryck and S. Mishra. Generic bounds on the approximation error for physics-informed (and) operator learning. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  12. 12.Q. Delfosse, P. Schramowski, M. Mundt, A. Molina, and K. Kersting. Adaptive rational activations to boost deep reinforcement learning. arXiv preprint arXiv:2102.09407, 2021.
  13. 13.L. C. Evans. Partial differential equations, volume 19. American Mathematical Soc., 2010.
  14. 14.V. Fanaskov and I. Oseledets. Spectral neural operators. arXiv preprint arXiv:2205.10573v1, 2022.
  15. 15.I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio. Deep learning, volume 1. MIT Press, 2016.
  16. 16.J. K. Gupta and J. Brandstetter. Towards multi-spatiotemporal-scale generalized pde modeling, 2022.
  17. 17.E. Haber and L. Ruthotto. Stable architectures for deep neural networks. Inverse problems, 34, 2018.
  18. 18.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv: 1512.03385, 2015.
  19. 19.J. S. Hesthaven, S. Gottlieb, and D. Gottlieb. Spectral methods for time-dependent problems, volume 21. Cambridge University Press, 2007.
  20. 20.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  21. 21.G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang. Physics informed machine learning. Nature Reviews Physics, pages 1–19, may 2021.
  22. 22.T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila. Alias-free generative adversarial networks. Advances in Neural Information Processing Systems, 34:852–863, 2021.
  23. 23.G. Kissas, J. H. Seidman, L. F. Guilhoto, V. M. Preciado, G. J. Pappas, and P. Perdikaris. Learning operators with coupled attention. Journal of Machine Learning Research, 23(215):1–63, 2022.
  24. 24.N. Kovachki, S. Lanthaler, and S. Mishra. On universal approximation and error bounds for fourier neural operators. Journal of Machine Learning Research, 22:Art–No, 2021.
  25. 25.N. Kovachki, Z. Li, B. Liu, K. Azizzadensheli, K. Bhattacharya, A. Stuart, and A. Anandkumar. Neural operator: Learning maps between function spaces. arXiv preprint arXiv:2108.08481v3, 2021.
  26. 26.S. Lanthaler, S. Mishra, and G. E. Karniadakis. Error estimates for DeepONets: A deep learning framework in infinite dimensions. Transactions of Mathematics and Its Applications, 6(1):tnac001, 2022.
  27. 27.S. Lanthaler, S. Mishra, and C. Parés-Pulido. Statistical solutions of the incompressible euler equations. Mathematical Models and Methods in Applied Sciences, 31(02):223–292, Feb 2021.
  28. 28.S. Lanthaler, R. Molinaro, P. Hadorn, and S. Mishra. Nonlinear reconstruction for operator learning of pdes with discontinuities. In International Conference on Learning Representations, 2023.
  29. 29.Y. LeCun, Y. Bengio, and G. Hinton. Deep learning. Nature, 521(7553):436–444, 2015.
  30. 30.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  31. 31.M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural networks, 6(6):861–867, 1993.
  32. 32.Z. Li, D. Z. Huang, B. Liu, and A. Anandkumar. Fourier neural operator with learned deformations for pdes on general geometries, 2022.
  33. 33.Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, 2021.
  34. 34.Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. M. Stuart, and A. Anandkumar. Neural operator: Graph kernel network for partial differential equations. CoRR, abs/2003.03485, 2020.
  35. 35.Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. Liu, A. M. Stuart, K. Bhattacharya, and A. Anandkumar. Multipole graph neural operator for parametric partial differential equations. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems (NeurIPS), volume 33, pages 6755–6766. Curran Associates, Inc., 2020.
  36. 36.Z. Li, H. Zheng, N. Kovachki, D. Jin, H. Chen, B. Liu, K. Azizzadenesheli, and A. Anandkumar. Physics-informed neural operator for learning partial differential equations. arXiv preprint arXiv:2111.03794, 2021.
  37. 37.Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie. A convnet for the 2020s. In Proceedings - 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 11966–11976. IEEE Computer Society, 2022.
  38. 38.Z. Long, Y. Lu, X. Ma, and B. Dong. Pde-net: Learning pdes from data. In J. G. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 3214–3222. PMLR, 2018.
  39. 39.L. Lu, P. Jin, G. Pang, Z. Zhang, and G. E. Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, 2021.
  40. 40.K. O. Lye, S. Mishra, D. Ray, and P. Chandrashekar. Iterative surrogate model optimization (ISMO): An active learning algorithm for PDE constrained optimization with deep neural networks. Computer Methods in Applied Mechanics and Engineering, 374:113575, 2021.
  41. 41.Z. Mao, L. Lu, O. Marxen, T. Zaki, and G. E. Karniadakis. DeepMandMnet for hypersonics: Predicting the coupled flow and finite-rate chemistry behind a normal shock using neural-network approximation of operators. Preprint, available from arXiv:2011.03349v1, 2020.
  42. 42.D. A. Masters, N. J. Taylor, T. Rendall, C. B. Allen, and D. J. Poole. Geometric comparison of aerofoil shape parameterization methods. AIAA Journal, pages 1575–1589, 2017.
  43. 43.A. Molina, P. Schramowski, and K. Kersting. Pad\’e activation units: End-to-end learning of flexible activation functions in deep networks. arXiv preprint arXiv:1907.06732, 2019.
  44. 44.R. Molinaro, y. Yang, E. Engquist, and S. Mishra. Neural inverse operators for solving pde inverse problems. arXiv:2301.11167, 2023.
  45. 45.J. Pathak, S. Subramanian, P. Harrington, S. Raja, A. Chattopadhyay, M. Mardani, T. Kurth, D. Hall, Z. Li, K. Azizzadenesheli, p. Hassanzadeh, K. Kashinath, and A. Anandkumar. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators. arXiv preprint arXiv:2202.11214, 2022.
  46. 46.P. Petersen and F. Voigtlaender. Equivalence of approximation by convolutional neural networks and fully-connected networks. Proceedings of the American Mathematical Society, 148(4):1567–1581, 2020.
  47. 47.M. Prasthofer, T. De Ryck, and S. Mishra. Variable input deep operator networks. arXiv preprint arXiv:2205.11404, 2022.
  48. 48.A. Quarteroni and A. Valli. Numerical approximation of Partial differential equations, volume 23. Springer, 1994.
  49. 49.M. Raissi, P. Perdikaris, and G. E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019.
  50. 50.O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015.
  51. 51.T. Schanze. Sinc interpolation of discrete periodic signals. IEEE Transactions on Signal Processing, 43(6):1502–1503, 1995.
  52. 52.J. H. Seidman, G. Kissas, P. Perdikaris, and G. J. Pappas. NOMAD: Nonlinear manifold decoders for operator learning. arXiv preprint arXiv:2206.03551, 2022.
  53. 53.M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng. Fourier features let networks learn high frequency functions in low dimensional domains. Advances in Neural Information Processing Systems, 33:7537–7547, 2020.
  54. 54.M. Telgarsky. Neural networks and rational functions. In International Conference on Machine Learning, pages 3387–3393. PMLR, 2017.
  55. 55.A. Tran, A. Mathews, L. Xie, and C. S. Ong. Factorized fourier neural operators. In The Eleventh International Conference on Learning Representations, 2023.
  56. 56.M. Unser. Sampling-50 years after shannon. Proceedings of the IEEE, 88(4):569–587, 2000.
  57. 57.M. Vetterli, J. Kovacevic, and V. Goyal. Foundations of Signal Processing. Cambridge University Press, 2014.
  58. 58.S. Wang, S. Suo, W. Ma, A. Pokrovsky, and R. Urtasun. Deep parametric continuous convolutional neural networks. CoRR, abs/2101.06742, 2021.
  59. 59.S. Wang, H. Wang, and P. Perdikaris. Learning the solution operator of parametric partial differential equations with physics-informed DeepOnets. arXiv preprint arXiv:2103.10974, 2021.
  60. 60.S. E. Wei. Aliasing-free nonlinear signal processing using implicitly defined functions. IEEE Access, 10:76281–76295, 2022.
  61. 61.R. Wightman, H. Touvron, and H. Jégou. Resnet strikes back: An improved training procedure in timm. CoRR, abs/2110.00476, 2021.
  62. 62.J. Yang, Q. Du, and W. Zhang. Uniform l p-bound of the allen-cahn equation and its numerical discretization. International Journal of Numerical Analysis & Modeling, 15, 2018.
  63. 63.Y. Zhu and N. Zabaras. Bayesian deep convolutional encoder–decoder networks for surrogate modeling and uncertainty quantification. Journal of Computational Physics, 336:415–447, 2018.

Citation

MLA
Raonić, B., et al. “Convolutional Neural Operators for Robust and Accurate Learning of PDEs”. arXiv, 2023, http://arxiv.org/abs/2302.01178v3.
APA
Raonić, B., Molinaro, R., Ryck, T. D., Rohner, T., Bartolucci, F., Alaifari, R., Mishra, S., & Bézenac, E. de . (2023). Convolutional Neural Operators for robust and accurate learning of PDEs. arXiv. http://arxiv.org/abs/2302.01178v3
Chicago
Raonić, B., R. Molinaro, T. D. Ryck, et al. 2023. “Convolutional Neural Operators for Robust and Accurate Learning of PDEs”. arXiv. http://arxiv.org/abs/2302.01178v3.
Harvard
Raonić, B. et al. (2023) “Convolutional Neural Operators for robust and accurate learning of PDEs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.01178v3.
Vancouver
1. Raonić B, Molinaro R, Ryck TD, Rohner T, Bartolucci F, Alaifari R, Mishra S, Bézenac E de (2023) Convolutional Neural Operators for robust and accurate learning of PDEs. arXiv

BibTeX

@article{raonic2023convolutional,
  title = {Convolutional Neural Operators for robust and accurate learning of PDEs},
  author = {Raonić, Bogdan and Molinaro, Roberto and Ryck, Tim De and Rohner, Tobias and Bartolucci, Francesca and Alaifari, Rima and Mishra, Siddhartha and Bézenac, Emmanuel de},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.01178v3},
  eprint = {2302.01178}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors