Geometric and Physical Quantities improve E(3) Equivariant Message Passing

Johannes BrandstetterRob HesselinkElise van der PolErik J. BekkersMax Welling

article2022ICLR334 citations

Develops Steerable E(3) Equivariant Graph Neural Networks (SEGNNs), which incorporate steerable geometric vectors and tensors directly into message passing to improve physical and chemical property prediction over invariant scalar architectures.

Listen

Simulating complex physical systems and molecular interactions is essential for accelerating materials design, drug discovery, and clean energy technologies. However, standard machine learning methods often struggle with 3D physical data because they fail to naturally respect fundamental geometric symmetries, such as 3D rotations, translations, and reflections, or they discard crucial directional information such as velocity and forces by only using scalar values.

The article demonstrates a new machine learning architecture, Steerable E(3) Equivariant Graph Neural Networks (SEGNNs). The primary objective is to evaluate whether allowing graph neural networks to process rich geometric and physical vectors directly within both message passing and node update stages improves predictive accuracy and physical consistency across complex scientific tasks.

The researchers evaluated SEGNNs through controlled experiments and ablation studies across three primary benchmarks: a charged and gravitational multi-body physical dynamic system, the QM9 molecular property benchmark consisting of small organic molecules, and the large-scale Open Catalyst 2020 (OC20) benchmark containing over 450,000 catalyst-adsorbate combinations. The architecture was tested against existing linear point convolutions and invariant message passing baselines using identical parameter budgets to isolate the exact impact of non-linear steerable operations.

The analysis produced several key findings: First, SEGNN achieved state-of-the-art performance on the OC20 energy prediction benchmark, outperforming existing competitive architectures across both in-distribution and out-of-distribution splits (achieving an in-distribution mean absolute error of 0.5327 eV). Second, incorporating physical quantities such as velocity directly into the node updates reduced prediction error by approximately 39% compared to standard invariant baselines in multi-body physical simulations. Third, ablation studies established that non-linear message passing systematically outperforms traditional linear point convolutions. Finally, on molecular property benchmarks, the model maintained high predictive accuracy even when graph connectivity was restricted to a narrow 2 Å cutoff radius, reducing the total message volume by roughly six-fold.

These findings demonstrate that preserving full directional equivariance without collapsing vectors into invariant scalars significantly improves model generalization and sample efficiency. For engineering and scientific workflows, this allows surrogate machine learning models to provide fast, highly accurate shortcuts for computationally prohibitive quantum chemistry and physical dynamics simulations, directly lowering computational costs and accelerating research timelines.

Organizations developing machine learning for molecular modeling or physical simulations should adopt steerable, non-linear message passing architectures when directional physical quantities are present. Practitioners should select moderate harmonic orders (such as order-1 features and attributes), as higher orders yielded minimal accuracy gains while significantly increasing computational overhead. Before broad deployment, teams should pilot these models on their domain-specific datasets and profile runtime constraints, as computing Clebsch-Gordan tensor products introduces higher computational latency than standard scalar networks.

Cover for Geometric and Physical Quantities improve E(3) Equivariant Message Passing

Abstract

Including covariant information, such as position, force, velocity or spin is important in many tasks in computational physics and chemistry. We introduce Steerable E(3) Equivariant Graph Neural Networks (SEGNNs) that generalise equivariant graph networks, such that node and edge attributes are not restricted to invariant scalars, but can contain covariant information, such as vectors or tensors. This model, composed of steerable MLPs, is able to incorporate geometric and physical information in both the message and update functions. Through the definition of steerable node attributes, the MLPs provide a new class of activation functions for general use with steerable feature fields. We discuss ours and related work through the lens of equivariant non-linear convolutions, which further allows us to pin-point the successful components of SEGNNs: non-linear message aggregation improves upon classic linear (steerable) point convolutions; steerable messages improve upon recent equivariant graph networks that send invariant messages. We demonstrate the effectiveness of our method on several tasks in computational physics and chemistry and provide extensive ablation studies.

Table of Contents

  • 1 Introduction
  • 2 Generalised E(33) equivariant steerable message passing
  • 2.1 Steerable E(33) Equivariant Graph Neural Networks
  • 3 Message Passing as Convolution, Related Work
  • 4 Experiments
  • 5 Conclusion
  • 6 Reproducibility Statement
  • 7 Ethical Statement
  • References
  • A Mathematical background
  • A.1 Group definition and the groups E(33) and O(33)
  • A.2 Invariance, equivariance and representations
  • A.3 Steerable vectors, Wigner-D matrices and irreducible representations
  • A.4 Spherical harmonics
  • A.5 Clebsch-Gordan product and steerable MLPs
  • B Steerable group convolutions
  • C Experiments
  • C.1 Pseudocode of SEGNN and ablated architectures
  • C.2 Experimental details
  • D Licenses

Knowls

  1. Knowl 1 — Steerable E(3) Equivariant Graph Neural Network Layer

    model/method

    A Steerable E(3)E(3) Equivariant Graph Neural Network (SEGNN) layer updates steerable node features f~i∈VL\tilde{\mathbf{f}}_i \in \mathcal{V}_L at node viv_i in a graph G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E}) with 3D spatial coordinates xi∈R3\mathbf{x}_i \in \mathbb{R}^3. The layer operates via message computation, aggregation, and node feature updating:

    m~ij=ϕm(f~i,f~j,∥xj−xi∥2,a~ij)\tilde{\mathbf{m}}_{ij} = \phi_m\left(\tilde{\mathbf{f}}_i, \tilde{\mathbf{f}}_j, \|\mathbf{x}_j - \mathbf{x}_i\|^2, \tilde{\mathbf{a}}_{ij}\right)

    f~i′=f~i+ϕf(f~i,∑j∈N(i)m~ij,a~i)\tilde{\mathbf{f}}_i' = \tilde{\mathbf{f}}_i + \phi_f\left(\tilde{\mathbf{f}}_i, \sum_{j \in \mathcal{N}(i)} \tilde{\mathbf{m}}_{ij}, \tilde{\mathbf{a}}_i\right)

    where:

    • VL=⨁l=0lfVl\mathcal{V}_L = \bigoplus_{l=0}^{l_f} \mathcal{V}_l is a direct sum of type-ll steerable vector spaces up to maximum degree lfl_f.
    • a~ij∈Vla\tilde{\mathbf{a}}_{ij} \in \mathcal{V}_{l_a} is a steerable edge attribute encoding geometric or physical relations (such as spherical harmonic embeddings of relative positions xj−xi\mathbf{x}_j - \mathbf{x}_i or relative velocities).
    • a~i∈Vla\tilde{\mathbf{a}}_i \in \mathcal{V}_{l_a} is a steerable node attribute (such as the neighborhood average of edge embeddings combined with node-specific vectors like velocity, force, or spin).
    • ϕm\phi_m is an O(3)O(3)-steerable MLP that computes messages steered by a~ij\tilde{\mathbf{a}}_{ij}.
    • ϕf\phi_f is an O(3)O(3)-steerable MLP that computes node updates steered by a~i\tilde{\mathbf{a}}_i.
    • N(i)\mathcal{N}(i) denotes the set of neighbors of node viv_i.

    Equivariance to the Euclidean group E(3)E(3) is guaranteed because translations enter exclusively through relative coordinate differences xj−xi\mathbf{x}_j - \mathbf{x}_i, and all transformations within ϕm\phi_m and ϕf\phi_f are equivariant with respect to orthogonal transformations (rotations and reflections) in O(3)O(3).

  2. Knowl 2 — Steerable Multi-Layer Perceptrons and Gated Activations

    model/method

    A steerable Multi-Layer Perceptron (steerable MLP) maps between steerable vector spaces Vin\mathcal{V}_{\text{in}} and Vout\mathcal{V}_{\text{out}} by interleaving conditioned steerable linear mappings with equivariant non-linearities.

    A steerable linear layer parameterized by learnable weights W\mathbf{W} and conditioned on a steerable geometric or physical attribute vector a~\tilde{\mathbf{a}} is defined via the Clebsch-Gordan (CG) tensor product ⊗cgW\otimes_{\text{cg}}^{\mathbf{W}}:

    Wa~h~:=h~⊗cgWa~\mathbf{W}_{\tilde{\mathbf{a}}} \tilde{\mathbf{h}} := \tilde{\mathbf{h}} \otimes_{\text{cg}}^{\mathbf{W}} \tilde{\mathbf{a}}

    For a layer receiving input h~k−1\tilde{\mathbf{h}}^{k-1} and conditioning attribute a~\tilde{\mathbf{a}}, the layer update is:

    h~k=σ(Wa~kh~k−1)\tilde{\mathbf{h}}^k = \sigma\left(\mathbf{W}_{\tilde{\mathbf{a}}}^k \tilde{\mathbf{h}}^{k-1}\right)

    The non-linearity σ\sigma is implemented using gated activations: scalar (type-0) irreducible representations (irreps) are passed through standard scalar activation functions (such as Swish), while higher-order irreps (l>0l > 0) are element-wise multiplied by scalar gating values generated by the linear layer and passed through scalar activations:

    Gate(h~,g)=h~(0)⊕⨁l>0σscalar(gl)h~(l)\text{Gate}(\tilde{\mathbf{h}}, \mathbf{g}) = \tilde{\mathbf{h}}^{(0)} \oplus \bigoplus_{l > 0} \sigma_{\text{scalar}}(g_l) \tilde{\mathbf{h}}^{(l)}

    Steerable MLPs satisfy the equivariance property for every group element g∈O(3)g \in O(3) with representation matrices D(g)\mathbf{D}(g) and D′(g)\mathbf{D}'(g):

    MLPD(g)a~(D(g)h~0)=D′(g)MLPa~(h~0)\text{MLP}_{\mathbf{D}(g)\tilde{\mathbf{a}}}(\mathbf{D}(g)\tilde{\mathbf{h}}_0) = \mathbf{D}'(g)\text{MLP}_{\tilde{\mathbf{a}}}(\tilde{\mathbf{h}}_0)

  3. Knowl 3 — Clebsch-Gordan Tensor Product for Steerable Vector Spaces

    equation

    Let Vl=R2l+1\mathcal{V}_l = \mathbb{R}^{2l+1} denote a type-ll steerable vector space transforming under O(3)O(3) via the (2l+1)×(2l+1)(2l+1) \times (2l+1) Wigner-D matrix D(l)(g)\mathbf{D}^{(l)}(g). The Clebsch-Gordan (CG) tensor product ⊗cgw:Vl1×Vl2→Vl\otimes_{\text{cg}}^w: \mathcal{V}_{l_1} \times \mathcal{V}_{l_2} \to \mathcal{V}_l between a type-l1l_1 vector h~(l1)\tilde{\mathbf{h}}^{(l_1)} and a type-l2l_2 vector h~(l2)\tilde{\mathbf{h}}^{(l_2)} parameterized by a learnable scalar weight w∈Rw \in \mathbb{R} produces a type-ll vector whose mm-th component (m∈{−l,−l+1,…,l}m \in \{-l, -l+1, \dots, l\}) is:

    (h~(l1)⊗cgwh~(l2))m(l)=w∑m1=−l1l1∑m2=−l2l2C(l1,m1)(l2,m2)(l,m)hm1(l1)hm2(l2)\left(\tilde{\mathbf{h}}^{(l_1)} \otimes_{\text{cg}}^w \tilde{\mathbf{h}}^{(l_2)}\right)_m^{(l)} = w \sum_{m_1=-l_1}^{l_1} \sum_{m_2=-l_2}^{l_2} C_{(l_1, m_1)(l_2, m_2)}^{(l, m)} h_{m_1}^{(l_1)} h_{m_2}^{(l_2)}

    where C(l1,m1)(l2,m2)(l,m)C_{(l_1, m_1)(l_2, m_2)}^{(l, m)} are the Clebsch-Gordan coefficients.

    The Clebsch-Gordan coefficients satisfy selection rules and vanish identically unless:

    ∣l1−l2∣≤l≤l1+l2|l_1 - l_2| \le l \le l_1 + l_2

    For mixed-type steerable spaces V=⨁lmlVl\mathcal{V} = \bigoplus_{l} m_l \mathcal{V}_l with multiplicities mlm_l, the weighted tensor product ⊗cgW\otimes_{\text{cg}}^{\mathbf{W}} maintains separate learnable weights for every allowable channel path across valid (l1,l2,l)(l_1, l_2, l) triplets.

  4. Knowl 4 — Spherical Harmonic Embeddings for Geometric and Physical Quantities

    model/method

    A three-dimensional vector x∈R3∖{0}\mathbf{x} \in \mathbb{R}^3 \setminus \{\mathbf{0}\} is converted into a type-ll steerable vector a~(l)∈Vl=R2l+1\tilde{\mathbf{a}}^{(l)} \in \mathcal{V}_l = \mathbb{R}^{2l+1} by evaluating real spherical harmonics Ym(l):S2→RY_m^{(l)}: S^2 \to \mathbb{R} on the normalized direction:

    a~(l)(x)=(Ym(l)(x∥x∥))m=−l,−l+1,…,lT\tilde{\mathbf{a}}^{(l)}(\mathbf{x}) = \left( Y_m^{(l)}\left( \frac{\mathbf{x}}{\|\mathbf{x}\|} \right) \right)_{m=-l, -l+1, \dots, l}^T

    Under orthogonal transformations R∈O(3)R \in O(3), the embedding satisfies:

    a~(l)(Rx)=D(l)(R)a~(l)(x)\tilde{\mathbf{a}}^{(l)}(R\mathbf{x}) = \mathbf{D}^{(l)}(R)\tilde{\mathbf{a}}^{(l)}(\mathbf{x})

    where D(l)(R)\mathbf{D}^{(l)}(R) is the order-ll Wigner-D representation matrix.

    In steerable message passing:

    • Steerable edge attributes a~ij\tilde{\mathbf{a}}_{ij} embed relative positions xj−xi\mathbf{x}_j - \mathbf{x}_i, relative velocities vj−vi\mathbf{v}_j - \mathbf{v}_i, or relative forces.
    • Steerable node attributes a~i\tilde{\mathbf{a}}_i are constructed by combining the average neighboring edge attributes with node-specific physical quantities vik\mathbf{v}_i^k (such as velocity, acceleration, force, or spin):

    a~i=1∣N(i)∣∑j∈N(i)a~ij+∑ka~(lk)(vik)\tilde{\mathbf{a}}_i = \frac{1}{|\mathcal{N}(i)|} \sum_{j \in \mathcal{N}(i)} \tilde{\mathbf{a}}_{ij} + \sum_k \tilde{\mathbf{a}}^{(l_k)}(\mathbf{v}_i^k)

  5. Knowl 5 — Equivariant Message Passing as Steerable Non-Linear Convolutions

    theoretical result

    Equivariant graph neural network layers can be systematically categorized through the lens of group convolutions on point clouds:

    1. Linear Steerable Point Convolutions: Conventional steerable point convolutions compute node updates as: f~i′=∑j∈N(i)Wa~ij(∥xj−xi∥)f~j\tilde{\mathbf{f}}_i' = \sum_{j \in \mathcal{N}(i)} \mathbf{W}_{\tilde{\mathbf{a}}_{ij}}(\|\mathbf{x}_j - \mathbf{x}_i\|) \tilde{\mathbf{f}}_j where Wa~ij\mathbf{W}_{\tilde{\mathbf{a}}_{ij}} acts as a linear kernel transformation of neighboring features f~j\tilde{\mathbf{f}}_j, and non-linearities enter only after neighbor aggregation via gated point-wise activations.

    2. Isotropic Non-Linear Message Passing (EGNN): Invariant message passing frameworks compute non-linear messages: mij=MLP(fi,fj,∥xj−xi∥2)\mathbf{m}_{ij} = \text{MLP}(\mathbf{f}_i, \mathbf{f}_j, \|\mathbf{x}_j - \mathbf{x}_i\|^2) These represent non-linear convolutions restricted to rotationally invariant (isotropic) scalar interactions.

    3. Steerable Non-Linear Message Passing (SEGNN): SEGNN unifies both paradigms by computing non-linear messages conditioned on steerable directional attributes a~ij\tilde{\mathbf{a}}_{ij}: m~ij=MLPa~ij(f~i,f~j,∥xj−xi∥2)\tilde{\mathbf{m}}_{ij} = \text{MLP}_{\tilde{\mathbf{a}}_{ij}}(\tilde{\mathbf{f}}_i, \tilde{\mathbf{f}}_j, \|\mathbf{x}_j - \mathbf{x}_i\|^2) and updating nodes via steerable MLPs conditioned on steerable node attributes a~i\tilde{\mathbf{a}}_i: f~i′=f~i+MLPa~i(f~i,∑j∈N(i)m~ij)\tilde{\mathbf{f}}_i' = \tilde{\mathbf{f}}_i + \text{MLP}_{\tilde{\mathbf{a}}_i}\left(\tilde{\mathbf{f}}_i, \sum_{j \in \mathcal{N}(i)} \tilde{\mathbf{m}}_{ij}\right) This incorporates non-linear feature propagation into steerable group convolutions while expanding invariant message passing to higher-order covariant geometric and physical quantities.

  6. Knowl 6 — SEGNN Message Passing Layer Algorithm

    algorithm

    The following algorithm details the execution of a single SEGNN message passing layer for a node viv_i receiving steerable node features f~i\tilde{\mathbf{f}}_i, relative position vectors xij=xj−xi\mathbf{x}_{ij} = \mathbf{x}_j - \mathbf{x}_i, and optional node physical vectors vi1,vi2\mathbf{v}_i^1, \mathbf{v}_i^2.

    Input: Node features f~i\tilde{\mathbf{f}}_i, relative positions xij\mathbf{x}_{ij}, physical vectors vi1,vi2\mathbf{v}_i^1, \mathbf{v}_i^2
    Output: Updated steerable node features f~i′\tilde{\mathbf{f}}_i'
    function O3_TENSOR_PRODUCT(input1, input2)
        output = CGTensorProduct(input1, input2)
        output = output + bias_on_type_0_irreps
        return output
    end function
    function O3_TENSOR_PRODUCT_SWISH_GATE(input1, input2)
        output, g = O3_TENSOR_PRODUCT(input1, input2)
        output_gated = Gate(output, Swish(g))
        return output_gated
    end function
    Compute edge attributes a~ij=SphericalHarmonicEmbedding(xij)\tilde{\mathbf{a}}_{ij} = \text{SphericalHarmonicEmbedding}(\mathbf{x}_{ij})
    Compute node attribute embeddings v~i1=SphericalHarmonicEmbedding(vi1)\tilde{\mathbf{v}}_i^1 = \text{SphericalHarmonicEmbedding}(\mathbf{v}_i^1), v~i2=SphericalHarmonicEmbedding(vi2)\tilde{\mathbf{v}}_i^2 = \text{SphericalHarmonicEmbedding}(\mathbf{v}_i^2)
    Aggregate node attributes a~i=∑j∈N(i)a~ij+v~i1+v~i2\tilde{\mathbf{a}}_i = \sum_{j \in \mathcal{N}(i)} \tilde{\mathbf{a}}_{ij} + \tilde{\mathbf{v}}_i^1 + \tilde{\mathbf{v}}_i^2
    for each neighbor j∈N(i)j \in \mathcal{N}(i) do
        h~ij=f~i⊕f~j⊕∥xij∥2\tilde{\mathbf{h}}_{ij} = \tilde{\mathbf{f}}_i \oplus \tilde{\mathbf{f}}_j \oplus \|\mathbf{x}_{ij}\|^2
        \tilde{\mathbf{m}}_{ij} = \text{O3_TENSOR_PRODUCT_SWISH_GATE}(\tilde{\mathbf{h}}_{ij}, \tilde{\mathbf{a}}_{ij})
        \tilde{\mathbf{m}}_{ij} = \text{O3_TENSOR_PRODUCT_SWISH_GATE}(\tilde{\mathbf{m}}_{ij}, \tilde{\mathbf{a}}_{ij})
    end for
    Aggregate messages m~i=∑j∈N(i)m~ij\tilde{\mathbf{m}}_i = \sum_{j \in \mathcal{N}(i)} \tilde{\mathbf{m}}_{ij}
    \tilde{\mathbf{f}}_i^{(1)} = \text{O3_TENSOR_PRODUCT_SWISH_GATE}(\tilde{\mathbf{f}}_i \oplus \tilde{\mathbf{m}}_i, \tilde{\mathbf{a}}_i)
    \tilde{\mathbf{f}}_i' = \tilde{\mathbf{f}}_i + \text{O3_TENSOR_PRODUCT}(\tilde{\mathbf{f}}_i^{(1)}, \tilde{\mathbf{a}}_i)
    return f~i′\tilde{\mathbf{f}}_i'
  7. Knowl 7 — Performance on Open Catalyst Project OC20 IS2RE Benchmark

    data/table

    The Initial Structure to Relaxed Energy (IS2RE) task in the Open Catalyst 2020 (OC20) dataset requires predicting the relaxed ground state energy from initial catalyst-adsorbate coordinates. Models are evaluated using Mean Absolute Error (MAE in eV) and Energy within Threshold (EwT, percentage of predictions within ϵ=0.02 eV\epsilon = 0.02\text{ eV} of ground truth) across four test splits: In-Distribution (ID), Out-of-Distribution Adsorbates (OOD Ads), Out-of-Distribution Catalysts (OOD Cat), and Out-of-Distribution Both (OOD Both).

    Energy MAE [eV] ↓\downarrow EwT [%] ↑\uparrow
    Model ID OOD Ads OOD Cat OOD Both ID OOD Ads OOD Cat OOD Both
    Median baseline 1.7499 1.8793 1.7090 1.6636 0.71% 0.72% 0.89% 0.74%
    CGCNN 0.6149 0.9155 0.6219 0.8511 3.40% 1.93% 3.10% 2.00%
    SchNet 0.6387 0.7342 0.6616 0.7037 2.96% 2.33% 2.94% 2.21%
    EdgeUpdateNet 0.5839 0.7252 0.6016 0.6862 3.48% 2.35% 3.30% 2.57%
    EnergyNet 0.6366 0.7170 0.6387 0.6626 3.30% 2.20% 3.07% 2.34%
    DimeNet++ 0.5620 0.7252 0.5756 0.6613 4.25% 2.07% 4.10% 2.41%
    SphereNet 0.5630 0.7030 0.5710 0.6380 4.47% 2.29% 4.09% 2.41%
    SEGNN (Ours) 0.5327 0.6921 0.5369 0.6790 5.37% 2.46% 4.91% 2.63%

    SEGNN achieves superior performance on all four test splits on the IS2RE task among methods trained on the IS2RE training partition, outperforming second-order directional architectures (SphereNet, DimeNet++) without requiring explicit 3-body or 4-body angular message updates.

  8. Knowl 8 — Evaluation on Charged N-Body Particle System Dynamics

    data/table

    The charged N-body system benchmark estimates the 3D positions of 5 interacting charged particles after 1,000 timesteps. Performance is measured by Mean Squared Error (MSE) on predicted particle positions, and computational cost is evaluated by forward pass execution time in seconds for a batch size of 100 samples on a GeForce RTX 2080 Ti GPU.

    Method MSE Time [s]
    SE(3)-Transformer .0244 .0742
    TFN .0155 .0182
    NMP .0107 .0017
    Radial Field .0104 .0019
    EGNN .0070 ±\pm .00022 .0029
    SElinear\text{SE}_{\text{linear}} (lf=2,la=2l_f=2, l_a=2) .0116 ±\pm .00021 .0640
    SEnon-linear\text{SE}_{\text{non-linear}} (lf=1,la=1l_f=1, l_a=1) .0060 ±\pm .00019 .0310
    SEGNNG\text{SEGNN}_{\text{G}} (lf=1,la=1l_f=1, l_a=1) .0056 ±\pm .00025 .0250
    SEGNNG+P\text{SEGNN}_{\text{G+P}} (lf=1,la=1l_f=1, l_a=1) .0043 ±\pm .00015 .0260

    Key takeaways from the comparison:

    1. Transitioning from linear steerable point convolutions (SElinear\text{SE}_{\text{linear}}) to non-linear steerable messages (SEnon-linear\text{SE}_{\text{non-linear}}) reduces MSE from .0116 to .0060.
    2. Incorporating geometric node attributes (neighbor-averaged orientations in SEGNNG\text{SEGNN}_{\text{G}}) further improves MSE to .0056.
    3. Adding physical velocity embeddings to node attributes (SEGNNG+P\text{SEGNN}_{\text{G+P}}) reduces MSE to .0043, establishing state-of-the-art performance.
  9. Knowl 9 — Molecular Property Prediction on the QM9 Benchmark

    data/table

    Performance of SEGNN and baseline models on the QM9 dataset across 12 quantum chemical properties, reported as Mean Absolute Error (MAE). Target properties include dipole moment (μ\mu), isotropic polarizability (α\alpha), highest occupied molecular orbital energy (ϵHOMO\epsilon_{\text{HOMO}}), lowest unoccupied molecular orbital energy (ϵLUMO\epsilon_{\text{LUMO}}), HOMO-LUMO gap (Δϵ\Delta\epsilon), spatial extent (R2R^2), zero-point vibrational energy (ZPVE\text{ZPVE}), heat capacity at constant volume (CvC_v), and thermodynamic energies (U0,U,H,GU_0, U, H, G).

    Task α\alpha Δϵ\Delta\epsilon ϵHOMO\epsilon_{\text{HOMO}} ϵLUMO\epsilon_{\text{LUMO}} μ\mu CvC_v GG HH R2R^2 UU U0U_0 ZPVE
    Units bohr3\text{bohr}^3 meV meV meV D calmol K\frac{\text{cal}}{\text{mol K}} meV meV bohr3\text{bohr}^3 meV meV meV
    NMP .092 69 43 38 .030 .040 19 17 .180 20 20 1.50
    SchNet .235 63 41 34 .033 .033 14 14 .073 19 14 1.70
    Cormorant .085 61 34 38 .038 .026 20 21 .961 21 22 2.02
    L1Net .088 68 46 35 .043 .031 14 14 .354 14 13 1.56
    LieConv .084 49 30 25 .032 .038 22 24 .800 19 19 2.28
    TFN .223 58 40 38 .064 .101 - - - - - -
    SE(3)-Tr. .142 53 35 33 .051 .054 - - - - - -
    EGNN .071 48 29 25 .029 .031 12 12 .106 12 12 1.55
    SEGNN (Ours) .060 42 24 21 .023 .031 15 16 .660 13 15 1.62

    Cutoff Radius Ablation on QM9: By utilizing steerable representations with lf=2,la=3l_f=2, l_a=3, SEGNN maintains high accuracy at a tight cutoff radius of 2 A˚2\text{ \AA} (sending ∼6×\sim 6\times fewer messages per layer than a 5 A˚5\text{ \AA} cutoff). In contrast, an invariant EGNN (lf=0,la=0l_f=0, l_a=0) degrades significantly at 2 A˚2\text{ \AA} (e.g., Δϵ\Delta\epsilon error increases from 53 meV53\text{ meV} on a fully connected graph to 98 meV98\text{ meV} at 2 A˚2\text{ \AA}, whereas SEGNN achieves 42 meV42\text{ meV} at 2 A˚2\text{ \AA}).

  10. Knowl 10 — Position and Force Prediction on 100-Body Gravitational Dynamics

    data/table

    Evaluation of position and force prediction on a 100-body gravitational particle simulation after 1,000 steps without boundary conditions. Models predict either the next position vector or the instantaneous force vector from current coordinates and velocities. Performance is measured in Mean Squared Error (MSE) alongside inference time per batch of 20 samples on an NVIDIA RTX 3090 GPU for varying neighbor connectivity (5, 20, 50 neighbors).

    5 neighbors 20 neighbors 50 neighbors
    Method pos force Time [s] pos force Time [s] pos force Time [s]
    MPNN .297 .299 .0012 .277 .273 .0014 .262 .268 .0029
    EGNN .301 unstable .0024 .256 unstable .0025 .239 unstable .0047
    SEGNN (lf=0,la=0l_f=0, l_a=0) .292 .296 .0085 .266 .276 .0088 .251 .265 .0100
    SEGNN (lf=1,la=1l_f=1, l_a=1) .265 .273 .0208 .237 .244 .0212 .212 .223 .0416

    Standard message passing networks (MPNN) fail to preserve equivariance for vector inputs and outputs, while EGNN is numerically unstable for direct force regression. SEGNN with lf=1,la=1l_f=1, l_a=1 preserves exact E(3)E(3) equivariance for both position and vector-force outputs, consistently outperforming invariant baselines across all neighborhood graph densities.

  11. Knowl 11 — Computational Overhead of Higher-Order Irreducible Representations in SEGNN

    limitation

    The computational cost of Clebsch-Gordan tensor products scales significantly with the maximum irreducible representation (irrep) order lfl_f and attribute order lal_a. For any l>0l > 0, computing the tensor product requires multiple tensor contractions across Clebsch-Gordan paths rather than a single matrix-vector multiply.

    On the charged N-body system benchmark, forward pass latency increases by more than an order of magnitude when moving from scalar to steerable representations:

    • Scalar baseline (lf=0,la=0l_f=0, l_a=0): 0.004 s
    • Order-1 steerable model (lf=1,la=1l_f=1, l_a=1): 0.026 s
    • Order-2 steerable model (lf=2,la=3l_f=2, l_a=3): 0.071 s

    Furthermore, increasing lfl_f and lal_a beyond 1 provides diminishing returns on several physical benchmarks (such as N-body dynamics and OC20 IS2RE) where first-order geometric and physical vectors (e.g., velocity and relative displacements) capture the primary physical symmetries.

Coverage note — None was omitted; all key architectural components, mathematical formulations, algorithms, benchmarks (N-body, QM9, OC20, Gravitational 100-body), and computational ablations are covered.

References

  1. 1.Sanchez-Gonzalez Alvaro, Godwin Jonathan, Pfaff Tobias, Ying Rex, Leskovec Jure, and Battaglia Peter. Learning to simulate complex physics with graph networks. In Proceedings of the 37th International Conference on Machine Learning, volume 119, pp. 8459–8468, 2020.
  2. 2.Brandon Anderson, Truong Son Hy, and Risi Kondor. Cormorant: Covariant molecular neural networks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  3. 3.Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu. Relational inductive biases, deep learning, and graph networks. arXiv preprint arXiv:1806.01261, 2018.
  4. 4.Simon Batzner, Tess E Smidt, Lixin Sun, Jonathan P Mailoa, Mordechai Kornbluth, Nicola Molinari, and Boris Kozinsky. Se (3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials. arXiv preprint arXiv:2101.03164, 2021.
  5. 5.Erik J Bekkers. B-spline cnns on lie groups. In International Conference on Learning Representations, 2019.
  6. 6.Erik J Bekkers, Maxime W Lafarge, Mitko Veta, Koen AJ Eppenhof, Josien PW Pluim, and Remco Duits. Roto-translation covariant convolutional networks for medical image analysis. In International conference on medical image computing and computer-assisted intervention, pp. 440–448. Springer, 2018.
  7. 7.Lukas Biewald. Experiment tracking with weights and biases, 2020. URL https://www.wandb.com/. Software available from wandb.com.
  8. 8.Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras-Domingo, Caleb Ho, Weihua Hu, Aini Palizhati, Anuroop Sriram, Brandon Wood, Junwoong Yoon, Devi Parikh, C. Lawrence Zitnick, and Zachary Ulissi. Open catalyst 2020 (oc20) dataset and community challenges. ACS Catalysis, 0(0):6059–6072, 2020. doi: 10.1021/acscatal.0c04525. URL https://doi.org/10.1021/acscatal.0c04525.
  9. 9.Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras-Domingo, Caleb Ho, Weihua Hu, Aini Palizhati, Anuroop Sriram, Brandon Wood, Junwoong Yoon, Devi Parikh, C. Lawrence Zitnick, and Zachary Ulissi. The open catalyst 2020 (oc20) dataset and community challenges, 2021.
  10. 10.Taco Cohen and Max Welling. Group equivariant convolutional networks. In International conference on machine learning, pp. 2990–2999. PMLR, 2016.
  11. 11.Taco S. Cohen and Max Welling. Steerable cnns. In International Conference on Learning Representations (ICLR), 2017.
  12. 12.Taco S. Cohen, Mario Geiger, Jonas Koehler, and Max Welling. Spherical cnns. In International Conference on Learning Representations (ICLR), 2018.
  13. 13.Michael Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016.
  14. 14.Walter Dehnen and Justin I Read. N-body simulations of gravitational dynamics. The European Physical Journal Plus, 126(5):1–28, 2011.
  15. 15.Congyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard, Andrea Tagliasacchi, and Leonidas Guibas. Vector neurons: A general framework for so(3)-equivariant networks. arXiv preprint arXiv:2104.12229, 2021.
  16. 16.Matthias Fey and Jan E. Lenssen. Fast graph representation learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019.
  17. 17.Marc Finzi, Samuel Stanton, Pavel Izmailov, and Andrew Gordon Wilson. Generalizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data. In International Conference on Machine Learning, pp. 3165–3176. PMLR, 2020.
  18. 18.Marc Finzi, Max Welling, and Andrew Gordon Wilson. A practical method for constructing equivariant multilayer perceptrons for arbitrary matrix groups. arXiv preprint arXiv:2104.09459, 2021.
  19. 19.William T Freeman, Edward H Adelson, et al. The design and use of steerable filters. IEEE Transactions on Pattern analysis and machine intelligence, 13(9):891–906, 1991.
  20. 20.Fabian B Fuchs, Daniel E Worrall, Volker Fischer, and Max Welling. Se (3)-transformers: 3d roto-translation equivariant attention networks. arXiv preprint arXiv:2006.10503, 2020.
  21. 21.Mario Geiger, Tess Smidt, Alby M., Benjamin Kurt Miller, Wouter Boomsma, Bradley Dice, Kostiantyn Lapchevskyi, Maurice Weiler, Michał Tyszkiewicz, Simon Batzner, Jes Frellsen, Nuri Jung, Sophia Sanborn, Josh Rackers, and Michael Bailey. e3nn/e3nn: 2021-04-21, April 2021a. URL https://doi.org/10.5281/zenodo.4708275.
  22. 22.Mario Geiger, Tess Smidt, Alby M., Benjamin Kurt Miller, Wouter Boomsma, Bradley Dice, Kostiantyn Lapchevskyi, Maurice Weiler, Michał Tyszkiewicz, Simon Batzner, Jes Frellsen, Nuri Jung, Sophia Sanborn, Josh Rackers, and Michael Bailey. e3nn/e3nn: 2021-05-10, May 2021b. URL https://doi.org/10.5281/zenodo.4745784.
  23. 23.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. In International Conference on Machine Learning, pp. 1263–1272. PMLR, 2017.
  24. 24.Yacov. Hel-Or and Patrick Teo. Canonical decomposition of steerable functions. In Proceedings CVPR IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 809–816, 1996.
  25. 25.Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael JL Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. arXiv preprint arXiv:2009.01411, 2020.
  26. 26.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel. Neural relational inference for interacting systems. In International Conference on Machine Learning, pp. 2688–2697. PMLR, 2018.
  28. 28.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations (ICLR), 2017.
  29. 29.Johannes Klicpera, Janek Groß, and Stephan Günnemann. Directional message passing for molecular graphs. In International Conference on Learning Representations, 2019.
  30. 30.Johannes Klicpera, Shankari Giri, Johannes T Margraf, and Stephan Günnemann. Fast and uncertainty-aware directional message passing for non-equilibrium molecules. arXiv preprint arXiv:2011.14115, 2020.
  31. 31.Johannes Klicpera, Florian Becker, and Stephan Günnemann. Gemnet: Universal directional graph neural networks for molecules. arXiv preprint arXiv:2106.08903, 2021.
  32. 32.Jonas Köhler, Leon Klein, and Frank Noé. Equivariant flows: sampling configurations for multi-body systems with symmetric energies. arXiv preprint arXiv:1910.00753, 2019.
  33. 33.Risi Kondor and Shubhendu Trivedi. On the generalization of equivariance and convolution in neural networks to the action of compact groups. In International Conference on Machine Learning, pp. 2747–2755. PMLR, 2018.
  34. 34.Risi Kondor, Zhen Lin, and Shubhendu Trivedi. Clebsch–gordan nets: a fully fourier space spherical convolutional neural network. Advances in Neural Information Processing Systems, 31:10117–10126, 2018.
  35. 35.Schütt Kristof, Kindermans Pieter-Jan, Sauceda Huziel, Chmiela Stefan, Tkatchenko Alexandre, and Klaus-Robert Müller. Schnet: a continuous-filter convolutional neural network for modeling quantum interactions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 992–1002, 2017.
  36. 36.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097–1105, 2012.
  37. 37.Leon Lang and Maurice Weiler. A wigner-eckart theorem for group equivariant convolution kernels. In International Conference on Learning Representations, 2020.
  38. 38.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  39. 39.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436–444, 2015.
  40. 40.Yi Liu, Limei Wang, Meng Liu, Xuan Zhang, Bora Oztekin, and Shuiwang Ji. Spherical message passing for 3d graph networks. arXiv preprint arXiv:2102.05013, 2021.
  41. 41.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  42. 42.Andreas Mayr, Sebastian Lehner, Arno Mayrhofer, Christoph Kloss, Sepp Hochreiter, and Johannes Brandstetter. Boundary graph neural networks for 3d simulations. arXiv preprint arXiv:2106.11299, 2021.
  43. 43.Benjamin Kurt Miller, Mario Geiger, Tess E Smidt, and Frank Noé. Relevance of rotationally equivariant convolutions for predicting molecular properties. arXiv preprint arXiv:2008.08461, 2020.
  44. 44.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. arXiv preprint arXiv:1912.01703, 2019.
  45. 45.Prajit Ramachandran, Barret Zoph, and Quoc V. Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017.
  46. 46.Raghunathan Ramakrishnan, Pavlo Dral, Matthias Rupp, and Anatole von Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific Data, 1, 08 2014.
  47. 47.David Romero, Erik Bekkers, Jakub Tomczak, and Mark Hoogendoorn. Attentive group equivariant convolutional networks. In Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, pp. 8188–8199, 2020.
  48. 48.Lars Ruddigkeit, Ruud Van Deursen, Lorenz C Blum, and Jean-Louis Reymond. Enumeration of 166 billion organic small molecules in the chemical universe database gdb-17. Journal of chemical information and modeling, 52(11):2864–2875, 2012.
  49. 49.Jun J. Sakurai and Jim Napolitano. Modern Quantum Mechanics. Cambridge University Press, 2 edition, 2017.
  50. 50.Victor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E (n) equivariant graph neural networks. arXiv preprint arXiv:2102.09844, 2021.
  51. 51.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE Transactions on Neural Networks, 20(1):61–80, 2009.
  52. 52.Jürgen Schmidhuber. Deep learning in neural networks: An overview. Neural networks, 61:85–117, 2015.
  53. 53.Kristof T Schütt, Oliver T Unke, and Michael Gastegger. Equivariant message passing for the prediction of tensorial properties and molecular spectra. arXiv preprint arXiv:2102.03150, 2021.
  54. 54.Nathaniel Thomas, Tess Smidt, Steven Kearnes, Lusann Yang, Li Li, Kai Kohlhoff, and Patrick Riley. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds. arXiv preprint arXiv:1802.08219, 2018.
  55. 55.Pfaff Tobias, Fortunato Meire, Sanchez-Gonzalez Alvaro, and Battaglia Peter W. Learning mesh-based simulation with graph networks, 2020.
  56. 56.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016.
  57. 57.Maurice Weiler and Gabriele Cesa. General e (2)-equivariant steerable cnns. Advances in Neural Information Processing Systems, 32:14334–14345, 2019.
  58. 58.Maurice Weiler, Mario Geiger, Max Welling, Wouter Boomsma, and Taco S Cohen. 3d steerable cnns: Learning rotationally equivariant features in volumetric data. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.
  59. 59.Maurice Weiler, Patrick Forré, Erik Verlinde, and Max Welling. Coordinate independent convolutional networks–isometry and gauge equivariant convolutions on riemannian manifolds. arXiv preprint arXiv:2106.06020, 2021.
  60. 60.Marysia Winkels and Taco S Cohen. 3d g-cnns for pulmonary nodule detection. In International Conference on Medi- cal Imaging with Deep Learning (MIDL), 2018.
  61. 61.Daniel Worrall and Gabriel Brostow. Cubenet: Equivariance to 3d rotation and translation. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 567–584, 2018.
  62. 62.Daniel E Worrall, Stephan J Garbin, Daniyar Turmukhambetov, and Gabriel J Brostow. Harmonic networks: Deep translation and rotation equivariance. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5028–5037, 2017.
  63. 63.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9621–9630, 2019.
  64. 64.Tian Xie and Jeffrey C Grossman. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Physical review letters, 120(14):145301, 2018.
  65. 65.C. Lawrence Zitnick, Lowik Chanussot, Abhishek Das, Siddharth Goyal, Javier Heras-Domingo, Caleb Ho, Weihua Hu, Thibaut Lavril, Aini Palizhati, Morgane Riviere, Muhammed Shuaibi, Anuroop Sriram, Kevin Tran, Brandon Wood, Junwoong Yoon, Devi Parikh, and Zachary Ulissi. An introduction to electrocatalyst design using machine learning for renewable energy storage, 2020.

Citation

MLA
Brandstetter, J., et al. “Geometric and Physical Quantities Improve E(3) Equivariant Message Passing”. arXiv, 2021, http://arxiv.org/abs/2110.02905v3.
APA
Brandstetter, J., Hesselink, R., Pol, E. van . der ., Bekkers, E. J., & Welling, M. (2021). Geometric and Physical Quantities Improve E(3) Equivariant Message Passing. arXiv. http://arxiv.org/abs/2110.02905v3
Chicago
Brandstetter, J., R. Hesselink, E. van . der . Pol, E. J. Bekkers, and M. Welling. 2021. “Geometric and Physical Quantities Improve E(3) Equivariant Message Passing”. arXiv. http://arxiv.org/abs/2110.02905v3.
Harvard
Brandstetter, J. et al. (2021) “Geometric and Physical Quantities Improve E(3) Equivariant Message Passing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.02905v3.
Vancouver
1. Brandstetter J, Hesselink R, Pol E van der, Bekkers EJ, Welling M (2021) Geometric and Physical Quantities Improve E(3) Equivariant Message Passing. arXiv

BibTeX

@article{brandstetter2021geometric,
  title = {Geometric and Physical Quantities Improve E(3) Equivariant Message Passing},
  author = {Brandstetter, Johannes and Hesselink, Rob and Pol, Elise van der and Bekkers, Erik J and Welling, Max},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.02905v3},
  eprint = {2110.02905}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors