The Deep Ritz Method: A Deep Learning-Based Numerical Algorithm for Solving Variational Problems

Weinan EBing Yu

article2017Communications in Mathematics and Statistics1,955 citations

Proposes the Deep Ritz method, a numerical framework that solves variational problems and high-dimensional partial differential equations by training neural networks to minimize energy formulations using stochastic gradient descent.

Listen

Traditional numerical methods for solving partial differential equations often struggle with complex geometries, singularities, and high-dimensional spaces. Because conventional grid-based techniques suffer from severe computational scaling issues when problem dimensions increase, finding flexible, scalable alternatives is critical for computational science and engineering applications.

The article introduces and evaluates the Deep Ritz Method, a deep learning-based framework designed to solve variational problems and partial differential equations. The objective is to demonstrate how deep neural networks can approximate trial functions within a Ritz formulation and scale effectively across different dimensions.

The approach integrates deep residual networks to represent trial functions and uses stochastic gradient descent with random mini-batch spatial sampling for numerical integration. The authors evaluate this technique across several benchmark problems, including two-dimensional Poisson equations with corner singularities, high-dimensional Poisson problems up to 100 dimensions, Neumann boundary condition setups, and quantum eigenvalue problems such as the infinite potential well and harmonic oscillator.

The experiments yield several key findings. First, in low-dimensional singular problems, the Deep Ritz Method achieved higher accuracy than the standard finite difference method while requiring significantly fewer parameters (for example, achieving a 0.0072 relative error with 811 parameters compared to 0.0125 error with 625 parameters in finite difference). Second, the method successfully scaled to high dimensions, solving a 10-dimensional Poisson problem to roughly 0.4% relative error and a 100-dimensional problem to about 2.2% error. Third, transfer learning accelerated early-stage training when system forcing terms changed. Finally, for eigenvalue estimation, the framework maintained accuracy in low to moderate dimensions (0.11% to 1.6% error in 5 dimensions), though errors degraded in 10-dimensional eigenvalue benchmarks (up to 12.6% error).

These findings indicate that neural network representations provide a naturally adaptive, nonlinear framework capable of circumventing the curse of dimensionality for many partial differential equations. This has significant implications for reducing modeling complexity and expanding computational feasibility in fields like quantum mechanics and high-dimensional physics without requiring tailored spatial meshing.

Organizations evaluating this approach should consider pilot implementations on high-dimensional or singularity-prone problems where standard finite element or finite difference methods are intractable. For operational deployment, teams should combine the method with transfer learning pipelines to minimize initial training time when evaluating recurring problem variants.

Despite these strengths, confidence in the method should be balanced against key limitations. Transforming the variational formulation into a neural network parameter optimization introduces non-convex loss landscapes with local minima, and strict mathematical convergence rates remain unproven. Furthermore, handling essential boundary conditions requires penalty terms, and performance degrades in high-dimensional eigenvalue settings, indicating that further refinements to architectures and optimization strategies are necessary before broad production use.

Cover for The Deep Ritz Method: A Deep Learning-Based Numerical Algorithm for Solving Variational Problems

Abstract

We propose a deep learning based method, the Deep Ritz Method, for numerically solving variational problems, particularly the ones that arise from partial differential equations. The Deep Ritz method is naturally nonlinear, naturally adaptive and has the potential to work in rather high dimensions. The framework is quite simple and fits well with the stochastic gradient descent method used in deep learning. We illustrate the method on several problems including some eigenvalue problems.

Table of Contents

  • 1 Introduction
  • 2 The Deep Ritz Method
  • 2.1 Building trial functions
  • 2.2 The stochastic gradient descent algorithm and the quadrature rule
  • 3 Numerical Results
  • 3.1 The Poisson equation in two dimension
  • 3.2 Poisson equation in high dimension
  • 3.3 An example with the Neumann boundary condition
  • 3.4 Transfer learning
  • 3.5 Eigenvalue problems
  • 4 Discussion
  • References

Knowls

  1. Knowl 1 — The Deep Ritz Method for Variational Problems

    model/method

    The Deep Ritz method numerically solves partial differential equations by reformulating them as variational energy minimization problems and parameterizing the trial function using a deep neural network u(x;θ)u(x; \theta) with parameter vector heta heta.

    For a domain Ω⊂Rd\Omega \subset \mathbb{R}^d, a source term f(x)f(x), and Dirichlet boundary conditions u(x)=g(x)u(x) = g(x) on ∂Ω\partial\Omega, the continuous variational problem is formulated with a boundary penalty term:

    I(u)=∫Ω(12∣∇xu(x)∣2−f(x)u(x))dx+β∫∂Ω∣u(x)−g(x)∣2dsI(u) = \int_\Omega \left( \frac{1}{2} |\nabla_x u(x)|^2 - f(x)u(x) \right) dx + \beta \int_{\partial\Omega} |u(x) - g(x)|^2 ds

    where β>0\beta > 0 is a penalty parameter enforcing the essential boundary condition.

    Rather than discretizing the domain with a fixed grid or mesh, the spatial integrals are approximated via Monte Carlo numerical quadrature using randomly sampled mini-batches of points in Ω\Omega and on ∂Ω\partial\Omega at each iteration of stochastic gradient descent (SGD or Adam). This allows the method to remain mesh-free and naturally adaptive across different spatial regions.

  2. Knowl 2 — Residual Network Architecture with Smooth Activation for Deep Ritz

    model/method

    The trial function u(x;θ)u(x; \theta) in the Deep Ritz method is constructed using a deep residual neural network. Given an input coordinate x∈Rdx \in \mathbb{R}^d, it is mapped to a hidden representation s∈Rms \in \mathbb{R}^m (using zero-padding if d<md < m or a linear projection if d>md > m).

    The network comprises a stack of residual blocks fi:Rm→Rmf_i: \mathbb{R}^m \to \mathbb{R}^m, where each block consists of two linear transformations, two activations, and a skip connection:

    fi(s)=ϕ(Wi,2⋅ϕ(Wi,1s+bi,1)+bi,2)+sf_i(s) = \phi\left(W_{i,2} \cdot \phi(W_{i,1} s + b_{i,1}) + b_{i,2}\right) + s

    with weight matrices Wi,1,Wi,2∈Rm×mW_{i,1}, W_{i,2} \in \mathbb{R}^{m \times m}, bias vectors bi,1,bi,2∈Rmb_{i,1}, b_{i,2} \in \mathbb{R}^m, and element-wise activation function ϕ\phi.

    To ensure C2C^2 smoothness necessary for computing spatial derivatives ∇xu(x;θ)\nabla_x u(x; \theta) while maintaining gradient flow, the smooth cubic rectified activation function is used:

    ϕ(x)=max⁡{x3,0}\phi(x) = \max\{x^3, 0\}

    After passing through nn sequential blocks zθ(x)=(fn∘⋯∘f1)(x)z_\theta(x) = (f_n \circ \dots \circ f_1)(x), the scalar trial solution is computed by a linear output layer:

    u(x;θ)=a⋅zθ(x)+bu(x; \theta) = a \cdot z_\theta(x) + b

    where a∈Rma \in \mathbb{R}^m and b∈Rb \in \mathbb{R}, and θ\theta denotes the complete set of parameters {Wi,j,bi,j,a,b}\{W_{i,j}, b_{i,j}, a, b\}.

  3. Knowl 3 — Deep Ritz Optimization Algorithm

    algorithm

    The Deep Ritz algorithm trains a neural network parameterizing a trial function u(x;θ)u(x; \theta) by minimizing the variational loss using stochastic quadrature.

    Input: Domain Ω⊂Rd\Omega \subset \mathbb{R}^d, boundary ∂Ω\partial\Omega, source function f(x)f(x), boundary function g(x)g(x), penalty parameter β\beta, batch sizes NΩN_\Omega and N∂ΩN_{\partial\Omega}, learning rate η\eta, total steps KK
    Output: Optimized network parameters θ\theta
    Initialize network parameters θ(0)\theta^{(0)}
    for k=0,1,…,K−1k = 0, 1, \dots, K-1 do
        Sample NΩN_\Omega points {xj}j=1NΩ\{x_{j}\}_{j=1}^{N_\Omega} uniformly and independently at random from Ω\Omega
        Sample N∂ΩN_{\partial\Omega} points {yj}j=1N∂Ω\{y_{j}\}_{j=1}^{N_{\partial\Omega}} uniformly and independently at random from ∂Ω\partial\Omega
        Compute domain loss approximation:
            LΩ(θ(k))=Vol(Ω)NΩ∑j=1NΩ(12∣∇xu(xj;θ(k))∣2−f(xj)u(xj;θ(k)))L_\Omega(\theta^{(k)}) = \frac{\mathrm{Vol}(\Omega)}{N_\Omega} \sum_{j=1}^{N_\Omega} \left( \frac{1}{2} |\nabla_x u(x_j; \theta^{(k)})|^2 - f(x_j)u(x_j; \theta^{(k)}) \right)
        Compute boundary loss approximation:
            L∂Ω(θ(k))=Area(∂Ω)N∂Ω∑j=1N∂Ω∣u(yj;θ(k))−g(yj)∣2L_{\partial\Omega}(\theta^{(k)}) = \frac{\mathrm{Area}(\partial\Omega)}{N_{\partial\Omega}} \sum_{j=1}^{N_{\partial\Omega}} |u(y_j; \theta^{(k)}) - g(y_j)|^2
        Evaluate total loss: L(θ(k))=LΩ(θ(k))+βL∂Ω(θ(k))L(\theta^{(k)}) = L_\Omega(\theta^{(k)}) + \beta L_{\partial\Omega}(\theta^{(k)})
        Compute parameter gradient: G(k)=∇θL(θ(k))G^{(k)} = \nabla_\theta L(\theta^{(k)})
        Update parameters: θ(k+1)=OptimizerUpdate(θ(k),G(k),η)\theta^{(k+1)} = \mathrm{OptimizerUpdate}(\theta^{(k)}, G^{(k)}, \eta)
    end for
    return θ(K)\theta^{(K)}

    In standard implementations, OptimizerUpdate\mathrm{OptimizerUpdate} is performed using the Adam optimizer.

  4. Knowl 4 — Deep Ritz Method for Quantum Eigenvalue Problems

    model/method

    To compute the ground-state eigenvalue λ0\lambda_0 of the Schrödinger-type eigenvalue problem −Δu(x)+v(x)u(x)=λu(x)-\Delta u(x) + v(x)u(x) = \lambda u(x) on Ω\Omega with u∣∂Ω=0u|_{\partial\Omega} = 0 (where v(x)v(x) is a potential function), the Deep Ritz method minimizes the Rayleigh quotient subject to normalization and boundary constraints.

    The unconstrained objective functional is formulated with penalty terms:

    L(u(⋅;θ))=∫Ω∣∇u∣2dx+∫Ωvu2dx∫Ωu2dx+β∫∂Ωu(x)2ds+γ(∫Ωu(x)2dx−1)2L(u(\cdot; \theta)) = \frac{\int_\Omega |\nabla u|^2 dx + \int_\Omega v u^2 dx}{\int_\Omega u^2 dx} + \beta \int_{\partial\Omega} u(x)^2 ds + \gamma \left( \int_\Omega u(x)^2 dx - 1 \right)^2

    where β>0\beta > 0 penalizes boundary non-zero values, and γ>0\gamma > 0 stabilizes the L2L_2-norm normalization. Keeping the Rayleigh quotient denominator ∫Ωu2dx\int_\Omega u^2 dx alongside the penalty parameter γ\gamma accelerates training and avoids requiring excessively large penalty weights.

    For eigenvalue problems, a DenseNet-style architecture is used with pairwise cross-layer skip connections:

    yi=ϕ(Wi−1xi−1+bi−1),xi=[xi−1;yi]y_i = \phi(W_{i-1}x_{i-1} + b_{i-1}), \quad x_i = [x_{i-1}; y_i]

    with quadratic rectified activation ϕ(x)=max⁡(0,x)2\phi(x) = \max(0, x)^2 (ReLU2\mathrm{ReLU}^2), which prevents gradient explosion during optimization.

  5. Knowl 5 — Accuracy of Deep Ritz vs Finite Difference on 2D Corner Singularity Poisson Problem

    data/table

    The Deep Ritz Method (DRM) was compared with the standard Finite Difference Method (FDM) on the 2D domain Ω=(−1,1)2∖[0,1)×{0}\Omega = (-1, 1)^2 \setminus [0, 1) \times \{0\} for the Poisson equation Δu(x)=0\Delta u(x) = 0 in Ω\Omega with exact solution u(r,θ)=r1/2sin⁡(θ/2)u(r, \theta) = r^{1/2} \sin(\theta/2) on ∂Ω\partial\Omega. The solution exhibits a corner singularity at the origin where derivative magnitudes diverge.

    Method Number of Blocks Parameters Relative L2L_2 Error
    DRM 3 591 0.0079
    DRM 4 811 0.0072
    DRM 5 1031 0.00647
    DRM 6 1251 0.0057
    FDM - 625 0.0125
    FDM - 2401 0.0063

    The results demonstrate that DRM achieves a smaller relative L2L_2 error than FDM with fewer parameters (e.g., DRM with 591 parameters yields 0.0079 error vs FDM with 625 parameters yielding 0.0125). This improvement in accuracy with fewer degrees of freedom is attributed to the natural nonlinear adaptivity of the neural network representation near the singularity.

  6. Knowl 6 — High-Dimensional Poisson Equation Solving via Deep Ritz Method

    empirical result

    The Deep Ritz method was evaluated on high-dimensional Poisson problems on unit hypercubes Ω=(0,1)d\Omega = (0, 1)^d:

    1. Ten-dimensional Dirichlet problem (d=10d=10): −Δu=0in (0,1)10,u(x)=∑k=15x2k−1x2kon ∂(0,1)10-\Delta u = 0 \quad \text{in } (0, 1)^{10}, \quad u(x) = \sum_{k=1}^5 x_{2k-1} x_{2k} \quad \text{on } \partial(0, 1)^{10} Using a network with 6 fully-connected layers, 3 skip connections (671 total parameters), β=103\beta = 10^3, and sampling 1,000 domain points plus 100 points per bounding hyperplane per step, the relative L2L_2 error was reduced to approximately 0.4%0.4\% after 50,000 SGD/Adam iterations.

    2. One-hundred-dimensional Poisson problem (d=100d=100): −Δu=−200in (0,1)100,u(x)=∑k=1100xk2on ∂(0,1)100-\Delta u = -200 \quad \text{in } (0, 1)^{100}, \quad u(x) = \sum_{k=1}^{100} x_k^2 \quad \text{on } \partial(0, 1)^{100} Using 3 residual blocks of width m=100m=100, the relative L2L_2 error was reduced to approximately 2.2%2.2\% after 50,000 iterations.

    These results confirm that the Deep Ritz method can solve high-dimensional PDEs without suffering from the grid-based curse of dimensionality.

  7. Knowl 7 — Deep Ritz Method with Natural Neumann Boundary Conditions

    empirical result

    The Deep Ritz method naturally handles homogeneous Neumann boundary conditions without requiring penalty terms on the boundary.

    For the dd-dimensional PDE system:

    −Δu(x)+π2u(x)=2π2∑k=1dcos⁡(πxk)for x∈[0,1]d-\Delta u(x) + \pi^2 u(x) = 2\pi^2 \sum_{k=1}^d \cos(\pi x_k) \quad \text{for } x \in [0, 1]^d ∂u∂n∣∂[0,1]d=0\frac{\partial u}{\partial n}\Big|_{\partial[0, 1]^d} = 0

    with exact solution u(x)=∑k=1dcos⁡(πxk)u(x) = \sum_{k=1}^d \cos(\pi x_k), the energy functional is minimized directly over the domain:

    I(u)=∫Ω(12(∣∇u(x)∣2+π2u(x)2)−f(x)u(x))dxI(u) = \int_\Omega \left( \frac{1}{2} \left( |\nabla u(x)|^2 + \pi^2 u(x)^2 \right) - f(x)u(x) \right) dx

    Because the zero-flux condition is a natural boundary condition for this variational form, no boundary penalty integral is added. Using a residual neural network architecture, the trained model achieved a relative L2L_2 error of 1.3%1.3\% for d=5d = 5 and 1.9%1.9\% for d=10d = 10 after 50,000 iterations.

  8. Knowl 8 — Transfer Learning in the Deep Ritz Method Across Forcing Functions

    empirical result

    Transferring pre-trained weights significantly accelerates the convergence of the Deep Ritz method when solving a PDE with a modified forcing function on the same domain.

    When solving −Δu(x)=6(1+x1)(1−x1)x2+2(1+x2)(1−x2)x2-\Delta u(x) = 6(1+x_1)(1-x_1)x_2 + 2(1+x_2)(1-x_2)x_2 on the slit domain Ω=(−1,1)2∖[0,1)×{0}\Omega = (-1, 1)^2 \setminus [0, 1) \times \{0\}, network parameters were initialized using the pre-trained weights from the homogeneous Laplace equation −Δu(x)=0-\Delta u(x) = 0 on the same domain.

    Compared to random weight initialization, weight transfer resulted in:

    1. A much steeper decrease in solution error during early training iterations.
    2. Substantially smaller parameter change norms ln⁡∥ΔW∥22\ln \|\Delta W\|_2^2 across training iterations.

    This indicates that transferring weights provides an effective warm-start strategy for parametric or sequentially varying PDEs.

  9. Knowl 9 — Ground-State Eigenvalue Estimation Across Dimensions

    data/table

    The Deep Ritz method was tested for ground-state eigenvalue computation across dimensions d∈{1,5,10}d \in \{1, 5, 10\} for an infinite potential well (v(x)=0v(x) = 0 on [0,1]d[0,1]^d, with exact eigenvalue λ0=dπ2\lambda_0 = d\pi^2) and a truncated harmonic oscillator (v(x)=∣x∣2v(x) = |x|^2 on [−3,3]d[-3, 3]^d, with exact eigenvalue λ0=d\lambda_0 = d).

    Problem Dimension dd Exact λ0\lambda_0 Approximate λ0\lambda_0 Relative Error
    Infinite Potential Well 1 9.87 9.85 0.20%
    Infinite Potential Well 5 49.35 49.29 0.11%
    Infinite Potential Well 10 98.70 92.35 6.43%
    Harmonic Oscillator 1 1.00 1.0016 0.16%
    Harmonic Oscillator 5 5.00 5.0814 1.60%
    Harmonic Oscillator 10 10.00 11.26 12.60%

    For lower dimensions (d=1,5d=1, 5), the relative errors remain below 2%2\%. However, as dimension increases to d=10d=10, accuracy degrades substantially (e.g., reaching 6.43%6.43\% error for the potential well and 12.6%12.6\% for the harmonic oscillator), highlighting optimization and domain truncation challenges in higher-dimensional eigenvalue problems.

  10. Knowl 10 — Limitations of the Deep Ritz Method

    limitation

    The Deep Ritz method exhibits several core limitations:

    1. Non-convex parameter optimization: Even when the underlying continuous variational problem I(u)I(u) is convex, composing the objective with a nonlinear deep neural network parameterization yields a non-convex optimization objective L(θ)L(\theta), exposing training to local minima and saddle points.
    2. Absence of convergence rates: There is no established theoretical convergence rate for the combination of neural network variational approximation and stochastic gradient optimization.
    3. Essential boundary enforcement: Essential (Dirichlet) boundary conditions must be enforced via penalty terms, introducing sensitive penalty hyperparameters (e.g., β\beta) and potential optimization stiffness.
    4. Scalability degradation on eigenvalue problems: In higher dimensions (e.g., d=10d=10), eigenvalue estimation accuracy degrades significantly compared to low-dimensional cases.

Coverage note — No substantial contributed material was omitted. All formulations, network architectures, algorithms, benchmark experiments, transfer learning investigations, eigenvalue results, and stated limitations are covered.

References

  1. 1.I. Goodfellow, Y. Bengio and A. Courville, Deep Learning. MIT Press, 2016.
  2. 2.W. E, “A proposal for machine learning via dynamical systems”, Communications in Mathematics and Statistics, March 2017, Volume 5, Issue 1, pp 1-11.
  3. 3.J. Q. Han, A. Jentzen and W. E, “Overcoming the curse of dimensionality: Solving high-dimensional partial differential equations using deep learning”, submitted, arXiv:1707.02568.
  4. 4.W. E, J. Q. Han and A. Jentzen, “Deep learning-based numerical methods for highdimensional parabolic partial differential equations and backward stochastic differential equations”, submitted, arXiv:1706.04702.
  5. 5.C. Beck, W. E and Arnulf Jentzen, “Machine learning approximation algorithms for high-dimensional fully nonlinear partial differential equations and second-order backward stochastic differential equations”, submitted. arXiv:1709.05963.
  6. 6.J. Q. Han, L. Zhang, R. Car and W. E, “Deep Potential: A general and “first-principle” representation of the potential energy”, submitted, arXiv:1707.01478.
  7. 7.L. Zhang, J.Q. Han, H. Wang, R. Car and W.E, “Deep Potential Molecular Dynamics: A scalable model with the accuracy of quantum mechanics”, submitted, arXiv:1707.09571.
  8. 8.L. C. Evans, Partial Differential Equations, 2nd ed. American Mathematical Society, 2010.
  9. 9.K. M. He, X. Y. Zhang, S. Q. Ren, J. Sun, “Deep residual learning for image recognition”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), vol. 00, no. , pp. 770-778, 2016, doi:10.1109/CVPR.2016.90
  10. 10.D. P. Kingma, and J. Ba. “Adam: A method for stochastic optimization.” arXiv preprint arXiv:1412.6980, 2014.
  11. 11.G. Strang and G. Fix, An Analysis of the Finite Element Method. Prentice-Hall, 1973.
  12. 12.G. Huang, Z. Liu, K. Q. Weinberger, V. D. M. Laurens, “Densely connected convolutional networks.”, arXiv preprint arXiv:1608.06993, 2016.

Citation

MLA
E, W., and B. Yu. “The Deep Ritz Method: A Deep Learning-based Numerical Algorithm for Solving Variational Problems”. arXiv, 2017, http://arxiv.org/abs/1710.00211v1.
APA
E, W., & Yu, B. (2017). The Deep Ritz method: A deep learning-based numerical algorithm for solving variational problems. arXiv. http://arxiv.org/abs/1710.00211v1
Chicago
E, W., and B. Yu. 2017. “The Deep Ritz Method: A Deep Learning-based Numerical Algorithm for Solving Variational Problems”. arXiv. http://arxiv.org/abs/1710.00211v1.
Harvard
E, W. and Yu, B. (2017) “The Deep Ritz method: A deep learning-based numerical algorithm for solving variational problems”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1710.00211v1.
Vancouver
1. E W, Yu B (2017) The Deep Ritz method: A deep learning-based numerical algorithm for solving variational problems. arXiv

BibTeX

@article{e2017the,
  title = {The Deep Ritz method: A deep learning-based numerical algorithm for solving variational problems},
  author = {E, Weinan and Yu, Bing},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1710.00211v1},
  eprint = {1710.00211}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF