Transolver: A Fast Transformer Solver for PDEs on General Geometries
Haixu WuHuakun LuoHaowen WangJianmin WangMingsheng Long
Develops Transolver, a linear-complexity Transformer architecture for solving partial differential equations on complex geometries by adaptively grouping mesh points into physics-aware slices, achieving a 22% relative performance gain across standard benchmarks and scaling to large-scale industrial aerodynamic simulations.
Simulating complex physical systems—such as aerodynamic vehicle design, fluid flows, and structural mechanics—relies heavily on solving partial differential equations (PDEs). Traditional numerical simulation methods can take hours or days to evaluate complex designs, creating severe bottlenecks in engineering workflows. While deep learning models, particularly Transformers, offer the promise of near-instant surrogate simulations, existing architectures struggle when scaling to large, irregular 2D and 3D simulation meshes due to excessive computational costs and difficulty learning intricate physical relationships directly from millions of discrete points.
The article demonstrates and evaluates Transolver, a physics-inspired neural solver designed to solve PDEs rapidly and accurately across diverse, unstructured geometries. The main objective is to establish an efficient Transformer architecture that models high-level physical states rather than calculating attention directly over massive collections of raw mesh points.
To achieve this, the authors introduced "Physics-Attention," an approach that adaptively partitions discretized geometric meshes into a compact set of learnable slices sharing similar physical properties, aggregates them into physics-aware tokens, performs attention among these tokens, and projects the results back to the mesh. The evaluation benchmarked Transolver against more than 20 competing neural operators and geometric models across six standard academic datasets (spanning regular grids, point clouds, and structured meshes) and two large-scale industrial tasks: 3D vehicle surface pressure and airflow from the ShapeNet dataset, and aerodynamic airfoil design from the AirfRANS dataset.
The findings establish that Transolver consistently sets new state-of-the-art results. Across the six standard benchmarks, Transolver achieved an average relative error reduction of about 22% compared to the strongest prior models, including error reductions of roughly 25% to 30% in solid mechanics and pipe flow tasks. In large-scale vehicle simulations involving over 32,000 mesh points, Transolver significantly outperformed baselines in predicting velocity and pressure fields while achieving a top Spearman’s rank correlation of 0.9935 for drag force ranking. In testing on unseen aerodynamic flow conditions, the architecture maintained rank correlations of nearly 99%, and computational benchmarks demonstrated that Transolver scales linearly with mesh size, executing nearly five times faster with roughly 77% less memory consumption than competing Transformer baselines on large grids.
These results demonstrate that focusing machine learning attention on underlying physical states rather than raw geometric points resolves key trade-offs between speed, memory, and physical accuracy. For industrial and engineering organizations, this approach enables accurate surrogate simulations that shorten product design cycles from days to seconds, reduces high-performance computing infrastructure costs, and reliably ranks design alternatives during aerodynamic and structural optimization.
Based on these findings, teams should consider piloting Transolver for high-throughput design screening workflows, particularly where geometry varies significantly across iterations. Organizations exploring foundation models for engineering simulations should also assess multiscale slicing strategies to capture hierarchical flow features efficiently. Before broad deployment in safety-critical manufacturing, further validation is recommended on highly turbulent or extreme out-of-distribution operating conditions with expanded physical datasets, as model performance was established on specific benchmark regimes and showed slight sensitivity to slice count tuning.
- Paper: Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs, Nikola Kovachki et al. (2023). Establishes the foundational neural operator framework for learning mappings between infinite-dimensional function spaces independently of mesh discretization.
- Paper: Fourier Neural Operator for Parametric Partial Differential Equations, Zongyi Li et al. (2020). Introduces Fourier Neural Operators for parameterizing continuous PDE solution operators, providing the essential baseline paradigm that Transolver aims to generalize to complex geometries.
- Paper: Learning Mesh-Based Simulation with Graph Networks, Tobias Pfaff et al. (2020). Demonstrates mesh-based physical simulations on irregular geometries using graph neural networks, illustrating the mesh-point modeling challenges that Transolver addresses via physics-attention.
- Paper: Learning Operators with Coupled Attention, Georgios Kissas et al. (2022). Explores operator learning coupled with attention mechanisms across continuous domains, directly informing attention-based neural PDE solver design.
- Paper: Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer, Wang Zeng et al. (2022). Presents dynamic token clustering to group flexible non-grid regions into learned tokens, motivating Transolver's concept of adaptively slicing discretized physical states.
- Paper: Convolutional Neural Operators for robust and accurate learning of PDEs, Bogdan Raonic et al. (2023). Analyzes aliasing and continuous-discrete mappings in neural PDE operators, setting key context for discretization-invariant PDE modeling.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). Surveys linear-complexity and efficient attention mechanisms that underpin the design of scalable Transformers for large-scale scientific data.
- Paper: Neural Operators with Localized Integral and Differential Kernels, Miguel Liu-Schiaffini et al. (2024). Extends continuous neural operator architectures to unstructured and irregular geometries by incorporating localized differential and integral kernels.
