The Deep Ritz Method: A Deep Learning-Based Numerical Algorithm for Solving Variational Problems
Weinan EBing Yu
Proposes the Deep Ritz method, a numerical framework that solves variational problems and high-dimensional partial differential equations by training neural networks to minimize energy formulations using stochastic gradient descent.
Traditional numerical methods for solving partial differential equations often struggle with complex geometries, singularities, and high-dimensional spaces. Because conventional grid-based techniques suffer from severe computational scaling issues when problem dimensions increase, finding flexible, scalable alternatives is critical for computational science and engineering applications.
The article introduces and evaluates the Deep Ritz Method, a deep learning-based framework designed to solve variational problems and partial differential equations. The objective is to demonstrate how deep neural networks can approximate trial functions within a Ritz formulation and scale effectively across different dimensions.
The approach integrates deep residual networks to represent trial functions and uses stochastic gradient descent with random mini-batch spatial sampling for numerical integration. The authors evaluate this technique across several benchmark problems, including two-dimensional Poisson equations with corner singularities, high-dimensional Poisson problems up to 100 dimensions, Neumann boundary condition setups, and quantum eigenvalue problems such as the infinite potential well and harmonic oscillator.
The experiments yield several key findings. First, in low-dimensional singular problems, the Deep Ritz Method achieved higher accuracy than the standard finite difference method while requiring significantly fewer parameters (for example, achieving a 0.0072 relative error with 811 parameters compared to 0.0125 error with 625 parameters in finite difference). Second, the method successfully scaled to high dimensions, solving a 10-dimensional Poisson problem to roughly 0.4% relative error and a 100-dimensional problem to about 2.2% error. Third, transfer learning accelerated early-stage training when system forcing terms changed. Finally, for eigenvalue estimation, the framework maintained accuracy in low to moderate dimensions (0.11% to 1.6% error in 5 dimensions), though errors degraded in 10-dimensional eigenvalue benchmarks (up to 12.6% error).
These findings indicate that neural network representations provide a naturally adaptive, nonlinear framework capable of circumventing the curse of dimensionality for many partial differential equations. This has significant implications for reducing modeling complexity and expanding computational feasibility in fields like quantum mechanics and high-dimensional physics without requiring tailored spatial meshing.
Organizations evaluating this approach should consider pilot implementations on high-dimensional or singularity-prone problems where standard finite element or finite difference methods are intractable. For operational deployment, teams should combine the method with transfer learning pipelines to minimize initial training time when evaluating recurring problem variants.
Despite these strengths, confidence in the method should be balanced against key limitations. Transforming the variational formulation into a neural network parameter optimization introduces non-convex loss landscapes with local minima, and strict mathematical convergence rates remain unproven. Furthermore, handling essential boundary conditions requires penalty terms, and performance degrades in high-dimensional eigenvalue settings, indicating that further refinements to architectures and optimization strategies are necessary before broad production use.
- Book: Convex Optimization: Algorithms and Complexity, Sébastien Bubeck (2015). Provides foundational principles and complexity guarantees for continuous optimization algorithms, which underpin the minimization of variational energies in deep learning solvers.
- Paper: An overview of gradient descent optimization algorithms, Sebastian Ruder (2016). Provides a comprehensive foundation on stochastic gradient descent and adaptive optimization methods used to minimize the Deep Ritz variational loss.
- Paper: Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming, Saeed Ghadimi et al. (2013). Establishes convergence properties of stochastic first-order optimization methods for nonconvex objectives that are essential for training neural networks on variational problem formulations.
- Paper: An Energy Approach to the Solution of Partial Differential Equations in Computational Mechanics via Machine Learning: Concepts, Implementation and Applications, Esteban Samaniego et al. (2019). Builds directly on the Deep Ritz variational philosophy to formulate the Deep Energy Method for structural and continuum mechanics PDEs.
- Paper: DGM: A deep learning algorithm for solving partial differential equations, Justin Sirignano et al. (2017). Explores the complementary strong-form collocation approach for solving high-dimensional partial differential equations via deep neural networks.
- Paper: DeepXDE: A Deep Learning Library for Solving Differential Equations, Lu Lu et al. (2019). Implements an open-source software platform uniting physics-informed and variational deep learning solvers across forward and inverse differential equation benchmarks.
- Paper: Scientific Machine Learning Through Physics–Informed Neural Networks: Where we are and What’s Next, Salvatore Cuomo et al. (2022). Synthesizes the development and applications of deep learning methods for solving PDEs, surveying the landscape established by foundational techniques like the Deep Ritz method.
- Paper: Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators, Lu Lu et al. (2021). Generalizes deep learning solvers from finding individual PDE solutions to learning nonlinear solution operators across infinite-dimensional function spaces.
- Paper: Fourier Neural Operator for Parametric Partial Differential Equations, Zongyi Li et al. (2020). Advances deep scientific computing by learning mesh-independent operator mappings for parametric PDEs using Fourier neural representations.
