EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction Reasoning

Chenxin XuRobby T. TanYuhong TanSiheng ChenYu Guang WangXinchao WangYanfeng Wang

article2023CVPR206 citations

Proposes an efficient multi-agent motion prediction framework that guarantees Euclidean equivariance and interaction invariance, achieving substantial error reductions across particle dynamics, molecular modeling, human skeleton motion, and pedestrian trajectory benchmarks.

Listen

Predicting the future movement of interacting agents is critical across high-impact technologies such as autonomous driving, robotics, physics modeling, and molecular simulation. Conventional prediction models transform past motion into abstract feature spaces that lose geometric orientation. As a result, they fail to preserve natural physical symmetries: predictions change inconsistently when the input coordinate frame is translated, rotated, or reflected. Addressing this requires models that satisfy two mathematical properties: motion equivariance (where transforming the input motion transforms the predicted future by the exact same shift or rotation) and interaction invariance (where relational behaviors between agents remain unchanged regardless of the coordinate system).

The article demonstrates that embedding mathematical equivariance directly into a sequence-to-sequence neural network significantly boosts prediction accuracy, model generalization, and computational efficiency across multiple diverse domains.

The authors developed EqMotion, an end-to-end framework that couples equivariant geometric feature learning with invariant pattern learning and relational reasoning. Unlike existing equivariant models limited to single timestamp transitions, EqMotion handles multi-step sequences. The network operates by first deriving translation-sensitive geometric features alongside rotation-invariant pattern features (such as speed and turning angle). It infers an invariant interaction graph to categorize relationships between agents and uses attention, spatial aggregation, and custom equivariant non-linear operations to update representations. The authors evaluated EqMotion across four distinct benchmark domains: particle physics simulations (springs and charged systems), molecular dynamics (MD17 dataset across four molecules), 3D human skeleton motion forecasting (Human 3.6M dataset), and pedestrian trajectory prediction (ETH-UCY benchmark).

Across all four evaluated domains, EqMotion established new state-of-the-art benchmarks, reducing prediction error by 24.0% in particle dynamics, 30.1% in molecular dynamics, 8.6% in 3D human skeleton motion, and 9.2% in pedestrian trajectory prediction. In relational reasoning tests, the model achieved perfect 100% consistency under arbitrary spatial transformations and reached 97.6% interaction category accuracy in spring simulations. In human motion benchmarks, EqMotion outperformed dedicated domain-specific baseline models while maintaining a lightweight footprint, requiring less than 30% of the parameter size of competing architectures. Furthermore, when trained on only 5% of available training data, EqMotion achieved performance comparable to or better than competing models trained on 100% of the dataset.

These findings prove that explicitly enforcing geometric physical laws in model architecture prevents networks from having to memorize arbitrary spatial rotations and translations. This fundamental architectural guarantee dramatically cuts data collection and labeling costs, reduces model parameter size, lowers inference overhead, and improves safety and reliability in physical automation and robotic control systems.

Organizations developing motion forecasting systems should prioritize embedding geometric symmetries into their network architectures rather than relying solely on post-processing or brute-force data augmentation. Implementing EqMotion architectures can yield immediate reductions in required training data and compute footprints. Future work should evaluate the framework in complex open environments that integrate rich static context, such as high-definition road maps in autonomous driving.

No sufficiently relevant recommendations were found.

Cover for EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction Reasoning

Abstract

Learning to predict agent motions with relationship reasoning is important for many applications. In motion prediction tasks, maintaining motion equivariance under Euclidean geometric transformations and invariance of agent interaction is a critical and fundamental principle. However, such equivariance and invariance properties are overlooked by most existing methods. To fill this gap, we propose EqMotion, an efficient equivariant motion prediction model with invariant interaction reasoning. To achieve motion equivariance, we propose an equivariant geometric feature learning module to learn a Euclidean transformable feature through dedicated designs of equivariant operations. To reason agent’s interactions, we propose an invariant interaction reasoning module to achieve a more stable interaction modeling. To further promote more comprehensive motion features, we propose an invariant pattern feature learning module to learn an invariant pattern feature, which cooperates with the equivariant geometric feature to enhance network expressiveness. We conduct experiments for the proposed model on four distinct scenarios: particle dynamics, molecule dynamics, human skeleton motion prediction and pedestrian trajectory prediction. Experimental results show that our method is not only generally applicable, but also achieves state-of-the-art prediction performances on all the four tasks, improving by 24.0/30.1/8.6/9.2%. Code is available at https://github.com/MediaBrain-SJTU/EqMotion.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background and Problem Formulation
  • 3.1. Motion Prediction
  • 3.2. Equivariance and Invariance
  • 4. Methodology
  • 4.1. Feature Initialization
  • 4.2. Invariant Reasoning Module
  • 4.3. Equivariant Geometric Feature Learning
  • 4.4. Invariant Pattern Feature Learning
  • 4.5. Equivariant Output Layer
  • 4.6. Theoretical Analysis
  • 5. Experiment
  • 5.1. Scenario 1: Particle Dynamic Prediction
  • 5.2. Scenario 2: Molecule Dynamic Prediction
  • 5.3. Scenario 3: Human Skeleton Motion Prediction
  • 5.4. Scenario 4: Pedestrian Trajectory Prediction
  • 5.5. Ablation Studies
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Euclidean Equivariance and Invariance Definitions in Multi-Agent Motion Prediction

    definition

    In multi-agent motion prediction, a system consists of MM interacting agents observed in an nn-dimensional Euclidean space (where n∈{2,3}n \in \{2, 3\}). The observed past trajectory of agent ii over TpT_p timestamps is Xi=[xi1,xi2,…,xiTp]∈RTp×n\mathbf{X}_i = [\mathbf{x}_i^1, \mathbf{x}_i^2, \dots, \mathbf{x}_i^{T_p}] \in \mathbb{R}^{T_p \times n}, and the ground-truth future trajectory over TfT_f timestamps is Yi=[yi1,yi2,…,yiTf]∈RTf×n\mathbf{Y}_i = [\mathbf{y}_i^1, \mathbf{y}_i^2, \dots, \mathbf{y}_i^{T_f}] \in \mathbb{R}^{T_f \times n}. The aggregate historical observations and future targets for all agents are denoted by X=[X1,X2,…,XM]∈RM×Tp×n\mathbb{X} = [\mathbf{X}_1, \mathbf{X}_2, \dots, \mathbf{X}_M] \in \mathbb{R}^{M \times T_p \times n} and Y=[Y1,Y2,…,YM]∈RM×Tf×n\mathbb{Y} = [\mathbf{Y}_1, \mathbf{Y}_2, \dots, \mathbf{Y}_M] \in \mathbb{R}^{M \times T_f \times n}, respectively.

    A spatial Euclidean transformation is parameterized by a translation vector t∈Rn\mathbf{t} \in \mathbb{R}^n and an orthogonal rotation or reflection matrix R∈SO(n)\mathbf{R} \in \mathrm{SO}(n) (or O(n)\mathrm{O}(n)). Equivariance and invariance for an operator F(⋅)\mathcal{F}(\cdot) with input X\mathbb{X} and output Z=F(X)\mathbb{Z} = \mathcal{F}(\mathbb{X}) are defined as follows:

    1. Euclidean Equivariance: An operation F(⋅)\mathcal{F}(\cdot) is equivariant under Euclidean transformations if transforming the input coordinates transforms the output coordinates by the exact same transformation:

    F(XR+t)=ZR+t∀R∈SO(n),  ∀t∈Rn\mathcal{F}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \mathbb{Z}\mathbf{R} + \mathbf{t} \quad \forall \mathbf{R} \in \mathrm{SO}(n), \; \forall \mathbf{t} \in \mathbb{R}^n

    A motion prediction model Fpred(⋅)\mathcal{F}_{\text{pred}}(\cdot) is motion equivariant if Fpred(XR+t)=Y^R+t\mathcal{F}_{\text{pred}}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \widehat{\mathbb{Y}}\mathbf{R} + \mathbf{t} for all R∈SO(n)\mathbf{R} \in \mathrm{SO}(n) and t∈Rn\mathbf{t} \in \mathbb{R}^n, where Y^=Fpred(X)\widehat{\mathbb{Y}} = \mathcal{F}_{\text{pred}}(\mathbb{X}).

    1. Euclidean Invariance: An operation F(⋅)\mathcal{F}(\cdot) is invariant under Euclidean transformations if transforming the input coordinates leaves the output unchanged:

    F(XR+t)=Z∀R∈SO(n),  ∀t∈Rn\mathcal{F}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \mathbb{Z} \quad \forall \mathbf{R} \in \mathrm{SO}(n), \; \forall \mathbf{t} \in \mathbb{R}^n

    An interaction reasoning model Freason(⋅)\mathcal{F}_{\text{reason}}(\cdot) is interaction invariant if the inferred relationships between agents satisfy Freason(XR+t)=Freason(X)\mathcal{F}_{\text{reason}}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \mathcal{F}_{\text{reason}}(\mathbb{X}).

  2. Knowl 2 — EqMotion Model Architecture and Prediction Pipeline

    model/method

    EqMotion is a sequence-to-sequence multi-agent motion prediction network that guarantees exact Euclidean equivariance of predicted trajectories and exact invariance of inferred agent interactions. The network maintains two coupled feature representations across LL layers: equivariant geometric features G(ℓ)=[G1(ℓ),…,GM(ℓ)]∈RM×C×n\mathbb{G}^{(\ell)} = [\mathbf{G}_1^{(\ell)}, \dots, \mathbf{G}_M^{(\ell)}] \in \mathbb{R}^{M \times C \times n} (preserving CC spatial coordinates per agent in Rn\mathbb{R}^n) and invariant pattern features H(ℓ)=[h1(ℓ),…,hM(ℓ)]∈RM×D\mathbf{H}^{(\ell)} = [\mathbf{h}_1^{(\ell)}, \dots, \mathbf{h}_M^{(\ell)}] \in \mathbb{R}^{M \times D} (capturing DD-dimensional scalar motion attributes).

    Given the input motion X∈RM×Tp×n\mathbb{X} \in \mathbb{R}^{M \times T_p \times n}, the five-stage feedforward pipeline is:

    1. Feature Initialization: Initial geometric and pattern features are extracted by FIL(⋅)\mathcal{F}_{\text{IL}}(\cdot):

    G(0),H(0)=FIL(X)\mathbb{G}^{(0)}, \mathbf{H}^{(0)} = \mathcal{F}_{\text{IL}}(\mathbb{X})

    1. Invariant Interaction Reasoning: When interaction relationships are not given a priori, an invariant reasoning module FIRM(⋅)\mathcal{F}_{\text{IRM}}(\cdot) infers an interaction graph with categorical edge distributions {cij}\{\mathbf{c}_{ij}\}:

    {cij}=FIRM(G(0),H(0))\{\mathbf{c}_{ij}\} = \mathcal{F}_{\text{IRM}}(\mathbb{G}^{(0)}, \mathbf{H}^{(0)})

    1. Equivariant Geometric Feature Learning (EGFL): In each layer ℓ∈{0,1,…,L−1}\ell \in \{0, 1, \dots, L-1\}, geometric features are updated via equivariant temporal attention, spatial aggregation weighted by {cij}\{\mathbf{c}_{ij}\}, and non-linear coordinate projections:

    G(ℓ+1)=FEGFL(ℓ)(G(ℓ),H(ℓ),{cij})\mathbb{G}^{(\ell+1)} = \mathcal{F}_{\text{EGFL}}^{(\ell)}(\mathbb{G}^{(\ell)}, \mathbf{H}^{(\ell)}, \{\mathbf{c}_{ij}\})

    1. Invariant Pattern Feature Learning (IPFL): In parallel, invariant pattern features are updated via invariant message passing informed by Euclidean pairwise distances:

    H(ℓ+1)=FIPFL(ℓ)(G(ℓ),H(ℓ))\mathbf{H}^{(\ell+1)} = \mathcal{F}_{\text{IPFL}}^{(\ell)}(\mathbb{G}^{(\ell)}, \mathbf{H}^{(\ell)})

    1. Equivariant Output Layer (EOL): The final geometric features G(L)\mathbb{G}^{(L)} are mapped to future trajectory predictions Y^=[Y^1,…,Y^M]∈RM×Tf×n\widehat{\mathbb{Y}} = [\widehat{\mathbf{Y}}_1, \dots, \widehat{\mathbf{Y}}_M] \in \mathbb{R}^{M \times T_f \times n} using a linear coordinate projection:

    Y^=FEOL(G(L)),Y^i=Wout(Gi(L)−G‾(L))+G‾(L)\widehat{\mathbb{Y}} = \mathcal{F}_{\text{EOL}}(\mathbb{G}^{(L)}), \quad \widehat{\mathbf{Y}}_i = \mathbf{W}_{\text{out}}\left(\mathbf{G}_i^{(L)} - \overline{\mathbb{G}}^{(L)}\right) + \overline{\mathbb{G}}^{(L)}

    where Wout∈RTf×C\mathbf{W}_{\text{out}} \in \mathbb{R}^{T_f \times C} is a learnable matrix and G‾(L)=1M⋅C∑i=1M∑c=1Cgi,c(L)∈R1×n\overline{\mathbb{G}}^{(L)} = \frac{1}{M \cdot C} \sum_{i=1}^M \sum_{c=1}^C \mathbf{g}_{i,c}^{(L)} \in \mathbb{R}^{1 \times n} is the global centroid.

  3. Knowl 3 — Equivariant Geometric Feature Learning Layer

    model/method

    The Equivariant Geometric Feature Learning layer FEGFL(ℓ)(⋅)\mathcal{F}_{\text{EGFL}}^{(\ell)}(\cdot) updates the geometric coordinates Gi(ℓ)∈RC×n\mathbf{G}_i^{(\ell)} \in \mathbb{R}^{C \times n} of each agent ii while preserving Euclidean equivariance through three sequential sub-operations:

    1. Equivariant Inner-Agent Attention: To model temporal dependencies across coordinate channels CC without disturbing geometric orientations, scalar attention weights are learned from the invariant pattern features hi(ℓ)∈RD\mathbf{h}_i^{(\ell)} \in \mathbb{R}^D and multiplied coordinate-wise relative to the global coordinate centroid G‾(ℓ)∈R1×n\overline{\mathbb{G}}^{(\ell)} \in \mathbb{R}^{1 \times n}:

    Gi(ℓ)←ϕatt(ℓ)(hi(ℓ))⊙(Gi(ℓ)−G‾(ℓ))+G‾(ℓ)\mathbf{G}_i^{(\ell)} \leftarrow \phi_{\text{att}}^{(\ell)}(\mathbf{h}_i^{(\ell)}) \odot \left(\mathbf{G}_i^{(\ell)} - \overline{\mathbb{G}}^{(\ell)}\right) + \overline{\mathbb{G}}^{(\ell)}

    where ϕatt(ℓ):RD→RC\phi_{\text{att}}^{(\ell)}: \mathbb{R}^D \to \mathbb{R}^C is an MLP and ⊙\odot is channel-wise broadcasting multiplication.

    1. Equivariant Inter-Agent Aggregation: To capture physical and social interactions, pairwise influences between agent ii and neighbors j∈Nij \in \mathcal{N}_i are propagated along their relative directional vectors Gi(ℓ)−Gj(ℓ)\mathbf{G}_i^{(\ell)} - \mathbf{G}_j^{(\ell)}:

    eij(ℓ)=∑k=1Kcij,k ϕe,k(ℓ)([hi(ℓ);hj(ℓ);∥Gi(ℓ)−Gj(ℓ)∥2,col])\mathbf{e}_{ij}^{(\ell)} = \sum_{k=1}^K c_{ij,k} \, \phi_{e,k}^{(\ell)}\left([\mathbf{h}_i^{(\ell)}; \mathbf{h}_j^{(\ell)}; \|\mathbf{G}_i^{(\ell)} - \mathbf{G}_j^{(\ell)}\|_{2,\text{col}} ]\right)

    Gi(ℓ)←Gi(ℓ)+∑j∈Nieij(ℓ)⊙(Gi(ℓ)−Gj(ℓ))\mathbf{G}_i^{(\ell)} \leftarrow \mathbf{G}_i^{(\ell)} + \sum_{j \in \mathcal{N}_i} \mathbf{e}_{ij}^{(\ell)} \odot \left(\mathbf{G}_i^{(\ell)} - \mathbf{G}_j^{(\ell)}\right)

    where cij=[cij,1,…,cij,K]⊤∈[0,1]K\mathbf{c}_{ij} = [c_{ij,1}, \dots, c_{ij,K}]^{\top} \in [0, 1]^K is the interaction category probability vector, ∥Gi(ℓ)−Gj(ℓ)∥2,col∈RC\|\mathbf{G}_i^{(\ell)} - \mathbf{G}_j^{(\ell)}\|_{2,\text{col}} \in \mathbb{R}^C denotes column-wise Euclidean distances across coordinate channels, and each ϕe,k(ℓ)(⋅)\phi_{e,k}^{(\ell)}(\cdot) is an MLP outputting channel weights in RC\mathbb{R}^C.

    1. Equivariant Non-Linear Coordinate Activation: An invariant condition splits coordinates to apply query-key directional projections, producing the updated geometric feature Gi(ℓ+1)\mathbf{G}_i^{(\ell+1)}.
  4. Knowl 4 — Equivariant Non-Linear Activation for Geometric Coordinates

    equation

    To introduce non-linear expressiveness into coordinate transformations while maintaining strict rotational and translational equivariance, EqMotion defines an equivariant activation based on invariant inner products and directional vector clipping.

    Let Gi(ℓ)∈RC×n\mathbf{G}_i^{(\ell)} \in \mathbb{R}^{C \times n} be the geometric coordinate feature of agent ii at layer ℓ\ell, and let G‾(ℓ)=1M⋅C∑i=1M∑c=1Cgi,c(ℓ)∈R1×n\overline{\mathbb{G}}^{(\ell)} = \frac{1}{M \cdot C} \sum_{i=1}^M \sum_{c=1}^C \mathbf{g}_{i,c}^{(\ell)} \in \mathbb{R}^{1 \times n} be the global spatial centroid across all MM agents and CC coordinate channels. Centered query coordinates Qi(ℓ)∈RC×n\mathbf{Q}_i^{(\ell)} \in \mathbb{R}^{C \times n} and key coordinates Ki(ℓ)∈RC×n\mathbf{K}_i^{(\ell)} \in \mathbb{R}^{C \times n} are computed using learnable channel-mixing weight matrices WQ(ℓ),WK(ℓ)∈RC×C\mathbf{W}_{\text{Q}}^{(\ell)}, \mathbf{W}_{\text{K}}^{(\ell)} \in \mathbb{R}^{C \times C}:

    Qi(ℓ)=WQ(ℓ)(Gi(ℓ)−G‾(ℓ)),Ki(ℓ)=WK(ℓ)(Gi(ℓ)−G‾(ℓ))\mathbf{Q}_i^{(\ell)} = \mathbf{W}_{\text{Q}}^{(\ell)}\left(\mathbf{G}_i^{(\ell)} - \overline{\mathbb{G}}^{(\ell)}\right), \quad \mathbf{K}_i^{(\ell)} = \mathbf{W}_{\text{K}}^{(\ell)}\left(\mathbf{G}_i^{(\ell)} - \overline{\mathbb{G}}^{(\ell)}\right)

    For each coordinate channel c∈{1,…,C}c \in \{1, \dots, C\}, let qi,c(ℓ),ki,c(ℓ)∈Rn\mathbf{q}_{i,c}^{(\ell)}, \mathbf{k}_{i,c}^{(\ell)} \in \mathbb{R}^n denote the cc-th row vectors of Qi(ℓ)\mathbf{Q}_i^{(\ell)} and Ki(ℓ)\mathbf{K}_i^{(\ell)}. The updated coordinate vector gi,c(ℓ+1)∈Rn\mathbf{g}_{i,c}^{(\ell+1)} \in \mathbb{R}^n is defined by:

    gi,c(ℓ+1)={qi,c(ℓ)+G‾(ℓ),if ⟨qi,c(ℓ),ki,c(ℓ)⟩≥0,qi,c(ℓ)−⟨qi,c(ℓ),ki,c(ℓ)∥ki,c(ℓ)∥2⟩ki,c(ℓ)∥ki,c(ℓ)∥2+G‾(ℓ),if ⟨qi,c(ℓ),ki,c(ℓ)⟩<0\mathbf{g}_{i,c}^{(\ell+1)} = \begin{cases} \mathbf{q}_{i,c}^{(\ell)} + \overline{\mathbb{G}}^{(\ell)}, & \text{if } \langle \mathbf{q}_{i,c}^{(\ell)}, \mathbf{k}_{i,c}^{(\ell)} \rangle \ge 0, \\[2mm] \mathbf{q}_{i,c}^{(\ell)} - \left\langle \mathbf{q}_{i,c}^{(\ell)}, \dfrac{\mathbf{k}_{i,c}^{(\ell)}}{\|\mathbf{k}_{i,c}^{(\ell)}\|_2} \right\rangle \dfrac{\mathbf{k}_{i,c}^{(\ell)}}{\|\mathbf{k}_{i,c}^{(\ell)}\|_2} + \overline{\mathbb{G}}^{(\ell)}, & \text{if } \langle \mathbf{q}_{i,c}^{(\ell)}, \mathbf{k}_{i,c}^{(\ell)} \rangle < 0 \end{cases}

    where ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle is the standard vector inner product in Rn\mathbb{R}^n and ∥⋅∥2\|\cdot\|_2 is the Euclidean norm. Because ⟨qi,c(ℓ)R,ki,c(ℓ)R⟩=⟨qi,c(ℓ),ki,c(ℓ)⟩\langle \mathbf{q}_{i,c}^{(\ell)}\mathbf{R}, \mathbf{k}_{i,c}^{(\ell)}\mathbf{R} \rangle = \langle \mathbf{q}_{i,c}^{(\ell)}, \mathbf{k}_{i,c}^{(\ell)} \rangle for any R∈SO(n)\mathbf{R} \in \mathrm{SO}(n), the branching condition is invariant under Euclidean transformations, and both branch assignments transform equivariantly.

  5. Knowl 5 — Invariant Interaction Reasoning Module (IRM)

    model/method

    The Invariant Interaction Reasoning Module FIRM(⋅)\mathcal{F}_{\text{IRM}}(\cdot) infers an interaction graph whose edge weights categorize spatial relationships into KK interaction types without requiring external supervision. The module operates strictly on invariant scalar quantities to ensure that the inferred edge categories are invariant under Euclidean transformations of agent coordinates.

    Given initial pattern features hi(0)∈RD\mathbf{h}_i^{(0)} \in \mathbb{R}^D and initial geometric features Gi(0)∈RC×n\mathbf{G}_i^{(0)} \in \mathbb{R}^{C \times n}, message passing is performed across agent neighbors Ni\mathcal{N}_i:

    mij′=ϕrm([hi(0);hj(0);∥Gi(0)−Gj(0)∥2,col])\mathbf{m}'_{ij} = \phi_{\text{rm}}\left([\mathbf{h}_i^{(0)}; \mathbf{h}_j^{(0)}; \|\mathbf{G}_i^{(0)} - \mathbf{G}_j^{(0)}\|_{2,\text{col}} ]\right)

    pi′=∑j∈Nimij′,hi′=ϕrh([pi′;hi(0)])\mathbf{p}'_i = \sum_{j \in \mathcal{N}_i} \mathbf{m}'_{ij}, \quad \mathbf{h}'_i = \phi_{\text{rh}}\left([\mathbf{p}'_i; \mathbf{h}_i^{(0)}]\right)

    cij=softmax⁡(ϕrc([hi′;hj′;∥Gi(0)−Gj(0)∥2,col])τ)\mathbf{c}_{ij} = \operatorname{softmax}\left( \dfrac{\phi_{\text{rc}}\left([\mathbf{h}'_i; \mathbf{h}'_j; \|\mathbf{G}_i^{(0)} - \mathbf{G}_j^{(0)}\|_{2,\text{col}} ]\right)}{\tau} \right)

    where ∥Gi(0)−Gj(0)∥2,col∈RC\|\mathbf{G}_i^{(0)} - \mathbf{G}_j^{(0)}\|_{2,\text{col}} \in \mathbb{R}^C is the vector of column-wise Euclidean distances across all CC coordinate channels, [⋅;⋅][\cdot; \cdot] denotes vector concatenation, ϕrm,ϕrh,ϕrc\phi_{\text{rm}}, \phi_{\text{rh}}, \phi_{\text{rc}} are learnable MLPs, τ>0\tau > 0 is a temperature hyperparameter controlling categorical smoothness, and cij∈[0,1]K\mathbf{c}_{ij} \in [0, 1]^K is the resulting categorical probability vector for the edge between agents ii and jj. The reasoning module is trained end-to-end alongside the full prediction network.

  6. Knowl 6 — Feature Initialization and Invariant Pattern Feature Learning

    model/method

    EqMotion initializes geometric and pattern representations and iteratively refines pattern features via invariant message passing:

    1. Feature Initialization (FIL\mathcal{F}_{\text{IL}}): Given historical trajectory Xi∈RTp×n\mathbf{X}_i \in \mathbb{R}^{T_p \times n} for each agent ii and global mean coordinate X‾=1M⋅Tp∑j=1M∑t=1Tpxjt∈R1×n\overline{\mathbb{X}} = \frac{1}{M \cdot T_p} \sum_{j=1}^M \sum_{t=1}^{T_p} \mathbf{x}_j^t \in \mathbb{R}^{1 \times n}:

      • Geometric feature: Gi(0)=Winit_g(Xi−X‾)+X‾∈RC×n\mathbf{G}_i^{(0)} = \mathbf{W}_{\text{init\_g}}(\mathbf{X}_i - \overline{\mathbb{X}}) + \overline{\mathbb{X}} \in \mathbb{R}^{C \times n}, where Winit_g∈RC×Tp\mathbf{W}_{\text{init\_g}} \in \mathbb{R}^{C \times T_p} is a learnable matrix performing temporal linear combination.
      • Kinematic invariants: Velocity Vi=ΔXi∈RTp×n\mathbf{V}_i = \Delta \mathbf{X}_i \in \mathbb{R}^{T_p \times n}, velocity magnitudes ρit=∥vit∥2\rho_i^t = \|\mathbf{v}_i^t\|_2, and turning angles θit=angle⁡(vit,vit−1)\theta_i^t = \operatorname{angle}(\mathbf{v}_i^t, \mathbf{v}_i^{t-1}).
      • Pattern feature: hi(0)=ϕinit_h([hoi;hetai])∈RD\mathbf{h}_i^{(0)} = \phi_{\text{init\_h}}([\boldsymbol{ ho}_i; \boldsymbol{ heta}_i]) \in \mathbb{R}^D, parameterized by an MLP or LSTM.
    2. Invariant Pattern Feature Learning Layer (FIPFL(ℓ)\mathcal{F}_{\text{IPFL}}^{(\ell)}): Pattern features are updated by aggregating invariant pairwise geometric differences and neighboring pattern states:

    mij(ℓ)=ϕm(ℓ)([hi(ℓ);hj(ℓ);∥Gi(ℓ)−Gj(ℓ)∥2,col])\mathbf{m}_{ij}^{(\ell)} = \phi_m^{(\ell)}\left([\mathbf{h}_i^{(\ell)}; \mathbf{h}_j^{(\ell)}; \|\mathbf{G}_i^{(\ell)} - \mathbf{G}_j^{(\ell)}\|_{2,\text{col}} ]\right)

    pi(ℓ)=∑j∈Nimij(ℓ),hi(ℓ+1)=ϕh(ℓ)([hi(ℓ);pi(ℓ)])\mathbf{p}_i^{(\ell)} = \sum_{j \in \mathcal{N}_i} \mathbf{m}_{ij}^{(\ell)}, \quad \mathbf{h}_i^{(\ell+1)} = \phi_h^{(\ell)}\left([\mathbf{h}_i^{(\ell)}; \mathbf{p}_i^{(\ell)}]\right)

    where ϕm(ℓ)\phi_m^{(\ell)} and ϕh(ℓ)\phi_h^{(\ell)} are MLPs, mij(ℓ)\mathbf{m}_{ij}^{(\ell)} represents edge messages, and pi(ℓ)\mathbf{p}_i^{(\ell)} is the aggregated neighborhood representation for agent ii.

  7. Knowl 7 — Theoretical Equivariance and Invariance Guarantees of EqMotion

    theoretical result

    EqMotion provides theoretical guarantees that all geometric representations and final predicted trajectories are equivariant under Euclidean transformations, while pattern features and inferred interaction graphs are invariant.

    Theorem 1: For any translation vector t∈Rn\mathbf{t} \in \mathbb{R}^n and orthogonal transformation matrix R∈SO(n)\mathbf{R} \in \mathrm{SO}(n) applied to the input trajectories X\mathbb{X}, the constituent modules of EqMotion satisfy:

    1. FIL(XR+t)=(G(0)R+t,  H(0))\mathcal{F}_{\text{IL}}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \left(\mathbb{G}^{(0)}\mathbf{R} + \mathbf{t}, \; \mathbf{H}^{(0)}\right)
    2. FIRM(G(0)R+t,  H(0))={cij}\mathcal{F}_{\text{IRM}}(\mathbb{G}^{(0)}\mathbf{R} + \mathbf{t}, \; \mathbf{H}^{(0)}) = \{\mathbf{c}_{ij}\}
    3. FEGFL(ℓ)(G(ℓ)R+t,  H(ℓ),  {cij})=G(ℓ+1)R+t\mathcal{F}_{\text{EGFL}}^{(\ell)}(\mathbb{G}^{(\ell)}\mathbf{R} + \mathbf{t}, \; \mathbf{H}^{(\ell)}, \; \{\mathbf{c}_{ij}\}) = \mathbb{G}^{(\ell+1)}\mathbf{R} + \mathbf{t}
    4. FIPFL(ℓ)(G(ℓ)R+t,  H(ℓ))=H(ℓ+1)\mathcal{F}_{\text{IPFL}}^{(\ell)}(\mathbb{G}^{(\ell)}\mathbf{R} + \mathbf{t}, \; \mathbf{H}^{(\ell)}) = \mathbf{H}^{(\ell+1)}
    5. FEOL(G(L)R+t)=Y^R+t\mathcal{F}_{\text{EOL}}(\mathbb{G}^{(L)}\mathbf{R} + \mathbf{t}) = \widehat{\mathbb{Y}}\mathbf{R} + \mathbf{t}

    Corollary 1: Combining the initialization, invariant interaction reasoning, LL alternating EGFL/IPFL layers, and the equivariant output layer, the end-to-end motion prediction network Fpred(⋅)\mathcal{F}_{\text{pred}}(\cdot) satisfies sequence-to-sequence Euclidean equivariance:

    Fpred(XR+t)=Y^R+t∀R∈SO(n),  ∀t∈Rn\mathcal{F}_{\text{pred}}(\mathbb{X}\mathbf{R} + \mathbf{t}) = \widehat{\mathbb{Y}}\mathbf{R} + \mathbf{t} \quad \forall \mathbf{R} \in \mathrm{SO}(n), \; \forall \mathbf{t} \in \mathbb{R}^n

  8. Knowl 8 — 3D Human Skeleton Motion Prediction Results on Human3.6M

    data/table

    Human skeleton motion prediction is evaluated on Human3.6M (H3.6M) across 15 actions using Mean Per Joint Position Error (MPJPE in millimeters, measuring average ℓ2\ell_2 distance between predicted and ground-truth 3D joints). Models are trained on 6 subjects and evaluated on Subject 5 for short-term (up to 400 ms) and long-term (up to 1000 ms) horizons.

    Method Short-Term Average MPJPE (ms) Long-Term Average MPJPE (ms)
    80ms 160ms 320ms 400ms 560ms 1000ms
    Res-sup. (CVPR'17) 34.7 62.0 101.1 115.5 129.2 165.0
    Traj-GCN (ICCV'19) 12.7 26.1 52.3 63.5 81.6 114.3
    DMGNN (CVPR'20) 17.0 33.6 65.9 79.7 93.6 127.6
    MSRGCN (ICCV'21) 12.1 25.6 51.6 62.9 81.1 114.2
    PGBIG (CVPR'22) 10.3 22.7 47.4 58.5 76.9 110.3
    SPGSN (ECCV'22) 10.4 22.3 47.1 58.3 77.4 109.6
    EqMotion (Ours) 9.1 20.1 43.7 55.0 73.4 106.9

    EqMotion outperforms all prior methods across all prediction intervals, reducing the average short-term MPJPE by 8.6% relative to the strongest baselines (PGBIG and SPGSN) and achieving a 3.8% improvement on long-term predictions (560–1000 ms).

  9. Knowl 9 — Molecular and Physical Dynamics Prediction and Interaction Reasoning Results

    data/table

    EqMotion was evaluated on two physical dynamic simulation benchmarks: physical NN-body systems (Springs and Charged particles) and molecular dynamics trajectories from the MD17 dataset (Aspirin, Benzene, Ethanol, Malonaldehyde).

    1. Interaction Category Recognition on NN-body Dynamics: Interaction reasoning accuracy (%) and consistency (% agreement across 20 random Euclidean transformations, mean ±\pm std over 5 independent runs):
    Model Springs Charged
    Accuracy (%) Consistency (%) Accuracy (%) Consistency (%)
    Corr.(path) 58.1 ±\pm 0.0 99.8 ±\pm 0.1 57.5 ±\pm 0.1 87.9 ±\pm 0.1
    Corr.(LSTM) 53.5 ±\pm 0.5 92.4 ±\pm 2.1 57.2 ±\pm 0.4 91.7 ±\pm 1.1
    EGNN 61.0 ±\pm 1.3 100.0 ±\pm 0.0 58.2 ±\pm 1.4 100.0 ±\pm 0.0
    NRI 93.0 ±\pm 1.1 93.7 ±\pm 1.2 70.0 ±\pm 0.6 88.5 ±\pm 1.3
    dNRI 93.3 ±\pm 2.0 89.6 ±\pm 2.0 70.4 ±\pm 1.7 83.6 ±\pm 1.8
    EqMotion (Ours) 97.6 ±\pm 1.1 100.0 ±\pm 0.0 80.9 ±\pm 3.4 100.0 ±\pm 0.0
    Supervised (Upper bound) 98.7 ±\pm 0.2 100.0 ±\pm 0.0 97.4 ±\pm 0.2 100.0 ±\pm 0.0
    1. Molecular Trajectory Prediction on MD17 (ADE / FDE ×10−2\times 10^{-2}):
    Method Aspirin Benzene Ethanol Malonaldehyde
    Radial Field 17.98 / 26.20 7.73 / 12.47 8.10 / 10.61 16.53 / 25.10
    TFN 15.02 / 21.35 7.55 / 12.30 8.05 / 10.57 15.21 / 24.32
    SE(3)-Trans 15.70 / 22.39 7.62 / 12.50 8.05 / 10.86 15.44 / 24.47
    EGNN 14.61 / 20.65 7.50 / 12.16 8.01 / 10.22 15.21 / 24.00
    LSTM 17.59 / 24.79 6.06 / 9.46 7.73 / 9.88 15.14 / 22.90
    S-LSTM 13.12 / 18.14 3.06 / 3.52 7.23 / 9.85 11.93 / 18.43
    NRI 12.60 / 18.50 1.89 / 2.58 6.69 / 8.78 12.79 / 19.86
    NMMP 10.41 / 14.67 2.21 / 3.33 6.17 / 7.86 9.50 / 14.89
    GroupNet 10.62 / 14.00 2.02 / 2.95 6.00 / 7.88 7.99 / 12.49
    EqMotion (Ours) 5.95 / 8.38 1.18 / 1.73 5.05 / 7.02 5.85 / 9.02

    EqMotion achieves state-of-the-art accuracy across all molecules, decreasing ADE and FDE by 34.2% and 30.1% on average compared to prior methods, and attains perfect 100% invariance consistency in interaction reasoning.

  10. Knowl 10 — Pedestrian Trajectory Prediction Results on ETH-UCY

    data/table

    Pedestrian trajectory prediction performance evaluated on the ETH-UCY benchmark (ETH, HOTEL, UNIV, ZARA1, ZARA2) using 8 past timestamps (3.2 s) to predict 12 future timestamps (4.8 s). Evaluation metrics are Average Displacement Error (ADE in meters) and Final Displacement Error (FDE in meters) under both deterministic (single output) and multi-prediction (best of 20 samples) settings.

    Deterministic Method ETH Hotel Univ Zara1 Zara2 Average
    S-LSTM 1.09 / 2.35 0.79 / 1.76 0.67 / 1.40 0.47 / 1.00 0.56 / 1.17 0.72 / 1.54
    SGAN-ind 1.13 / 2.21 1.01 / 2.18 0.60 / 1.28 0.42 / 0.91 0.52 / 1.11 0.74 / 1.54
    Trajectron++ 1.02 / 2.00 0.33 / 0.62 0.53 / 1.19 0.44 / 0.99 0.32 / 0.73 0.53 / 1.11
    TransF 1.03 / 2.10 0.36 / 0.71 0.53 / 1.32 0.44 / 1.00 0.34 / 0.76 0.54 / 1.17
    MemoNet 1.00 / 2.08 0.35 / 0.67 0.55 / 1.19 0.46 / 1.00 0.37 / 0.82 0.55 / 1.15
    EqMotion (Ours) 0.96 / 1.92 0.30 / 0.58 0.50 / 1.10 0.39 / 0.86 0.30 / 0.68 0.49 / 1.03
    Multi-prediction (K=20K=20) ETH Hotel Univ Zara1 Zara2 Average
    SGAN 0.87 / 1.62 0.67 / 1.37 0.76 / 0.52 0.35 / 0.68 0.42 / 0.84 0.61 / 1.21
    NMMP 0.61 / 1.08 0.33 / 0.63 0.52 / 1.11 0.32 / 0.66 0.43 / 0.85 0.41 / 0.82
    Trajectron++ 0.61 / 1.02 0.19 / 0.28 0.30 / 0.54 0.24 / 0.42 0.18 / 0.31 0.30 / 0.51
    PECNet 0.54 / 0.87 0.18 / 0.24 0.35 / 0.60 0.22 / 0.39 0.17 / 0.30 0.29 / 0.48
    AgentFormer 0.45 / 0.75 0.14 / 0.22 0.25 / 0.45 0.18 / 0.30 0.14 / 0.24 0.23 / 0.39
    GroupNet 0.46 / 0.73 0.15 / 0.25 0.26 / 0.49 0.21 / 0.39 0.17 / 0.33 0.25 / 0.44
    MID 0.39 / 0.66 0.13 / 0.22 0.22 / 0.45 0.17 / 0.30 0.13 / 0.27 0.21 / 0.38
    GP-Graph 0.43 / 0.63 0.18 / 0.30 0.24 / 0.42 0.17 / 0.31 0.15 / 0.29 0.23 / 0.39
    EqMotion (Ours) 0.40 / 0.61 0.12 / 0.18 0.23 / 0.43 0.18 / 0.32 0.13 / 0.23 0.21 / 0.35

    EqMotion achieves state-of-the-art results, reducing average FDE by 10.4% under the deterministic setting and by 7.9% under the multi-prediction setting compared to previous best methods.

  11. Knowl 11 — Ablation Study on EqMotion Modules and Equivariant Operations

    data/table

    Ablation experiments conducted on the Human3.6M skeleton motion prediction benchmark quantify the contributions of individual modules and specific equivariant operations. Performance is reported in MPJPE (mm) across short-term timestamps:

    1. Ablation of Key Network Modules:
    EGFL IPFL IRM 80ms 160ms 320ms 400ms Average
    12.9 31.9 68.2 82.4 48.9
    ✓ 10.1 22.6 48.7 60.7 35.5
    ✓ ✓ 9.2 20.8 45.4 57.0 33.1
    ✓ ✓ ✓ 9.1 20.1 43.7 55.0 32.0
    1. Ablation of EGFL Internal Operations:
    Configuration 80ms 160ms 320ms 400ms Average
    Without Inner-Agent Attention 9.2 20.5 44.3 55.7 32.4
    Without Inter-Agent Aggregation 9.7 22.0 47.2 58.9 34.5
    Without Non-Linear Activation 9.4 21.3 46.7 58.6 34.0
    Full EqMotion 9.1 20.1 43.7 55.0 32.0

    The equivariant geometric feature learning module (EGFL) provides the largest single gain (reducing error from 48.9 to 35.5 mm). Combining all three modules (EGFL, IPFL, IRM) and all three equivariant geometric operations achieves the lowest error (32.0 mm).

  12. Knowl 12 — Data Efficiency and Parameter Compactness of EqMotion

    empirical result

    EqMotion demonstrates high sample efficiency and parameter compactness on the Human3.6M dataset due to its built-in geometric equivariance:

    1. Sample Efficiency: When trained on only 5% of the Human3.6M training data, EqMotion achieves prediction accuracy comparable to or better than competitive baseline models (such as Traj-GCN and MSR-GCN) trained on 100% of the data. Across all training data fractions (5%, 10%, 30%, 50%, 70%, 100%), EqMotion consistently attains lower MPJPE than Traj-GCN, MSR-GCN, and SPGSN.

    2. Parameter Compactness: EqMotion requires a model parameter size that is less than 30% of the model sizes of prominent 3D human motion prediction baselines (such as DMGNN, HisRep, MSGGCN, and SPGSN), while simultaneously attaining lower MPJPE. The exact Euclidean equivariance relieves the network from needing extra parameter capacity to learn invariance/equivariance across rotational and translational transformations.

Coverage note — None was omitted; all primary architectural components, mathematical equations, theoretical results, empirical benchmark tables across four domains, ablations, and data efficiency findings are represented.

References

  1. 1.Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social lstm: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 961–971, 2016.
  2. 2.Inhwan Bae, Jin-Hwi Park, and Hae-Gon Jeon. Learning pedestrian group representations for multi-modal trajectory prediction. In European Conference on Computer Vision, pages 270–289. Springer, 2022.
  3. 3.Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. Advances in neural information processing systems, 29, 2016.
  4. 4.Yujun Cai, Lin Huang, Yiwei Wang, Tat-Jen Cham, Jianfei Cai, Junsong Yuan, Jun Liu, Xu Yang, Yiheng Zhu, Xiaohui Shen, et al. Learning progressive joint propagation for human motion prediction. In European Conference on Computer Vision, pages 226–242. Springer, 2020.
  5. 5.Sergio Casas, Wenjie Luo, and Raquel Urtasun. Intentnet: Learning to predict intention from raw sensor data. In Conference on Robot Learning, pages 947–956. PMLR, 2018.
  6. 6.Yuning Chai, Benjamin Sapp, Mayank Bansal, and Dragomir Anguelov. Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. arXiv preprint arXiv:1910.05449, 2019.
  7. 7.Stefan Chmiela, Alexandre Tkatchenko, Huziel E Sauceda, Igor Poltavsky, Kristof T Schütt, and Klaus-Robert Müller. Machine learning of accurate energy-conserving molecular force fields. Science advances, 3(5):e1603015, 2017.
  8. 8.Taco Cohen and Max Welling. Group equivariant convolutional networks. In International conference on machine learning, pages 2990–2999. PMLR, 2016.
  9. 9.Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang, and Guiqing Li. Msr-gcn: Multi-scale residual graph convolution networks for human motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11467–11476, 2021.
  10. 10.Congyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard, Andrea Tagliasacchi, and Leonidas J Guibas. Vector neurons: A general framework for so (3)-equivariant networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12200–12209, 2021.
  11. 11.Carlos Esteves, Christine Allen-Blanchette, Xiaowei Zhou, and Kostas Daniilidis. Polar transformer networks. arXiv preprint arXiv:1709.01889, 2017.
  12. 12.Marc Finzi, Samuel Stanton, Pavel Izmailov, and Andrew Gordon Wilson. Generalizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data. In International Conference on Machine Learning, pages 3165–3176. PMLR, 2020.
  13. 13.Katerina Fragkiadaki, Sergey Levine, Panna Felsen, and Jitendra Malik. Recurrent network models for human dynamics. In Proceedings of the IEEE international conference on computer vision, pages 4346–4354, 2015.
  14. 14.Fabian Fuchs, Daniel Worrall, Volker Fischer, and Max Welling. Se (3)-transformers: 3d roto-translation equivariant attention networks. Advances in Neural Information Processing Systems, 33:1970–1981, 2020.
  15. 15.Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid. Vectornet: Encoding hd maps and agent dynamics from vectorized representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11525–11533, 2020.
  16. 16.Francesco Giuliari, Irtiza Hasan, Marco Cristani, and Fabio Galasso. Transformer networks for trajectory forecasting. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 10335–10342. IEEE, 2021.
  17. 17.Colin Graber and Alexander G Schwing. Dynamic neural relational inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8513–8522, 2020.
  18. 18.Tianpei Gu, Guangyi Chen, Junlong Li, Chunze Lin, Yongming Rao, Jie Zhou, and Jiwen Lu. Stochastic trajectory prediction via motion indeterminacy diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17113–17122, 2022.
  19. 19.Xiao Guo and Jongmoo Choi. Human motion prediction via learning local structure representations and temporal dependencies. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 2580–2587, 2019.
  20. 20.Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social gan: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2255–2264, 2018.
  21. 21.Dirk Helbing and Peter Molnar. Social force model for pedestrian dynamics. Physical review E, 51(5):4282, 1995.
  22. 22.Yue Hu, Siheng Chen, Ya Zhang, and Xiao Gu. Collaborative motion prediction via neural motion message passing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6319–6328, 2020.
  23. 23.Wenbing Huang, Jiaqi Han, Yu Rong, Tingyang Xu, Fuchun Sun, and Junzhou Huang. Equivariant graph mechanics networks with constraints. arXiv preprint arXiv:2203.06442, 2022.
  24. 24.Michael J Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont, Yee Whye Teh, and Hyunjik Kim. Lietransformer: Equivariant self-attention for lie groups. In International Conference on Machine Learning, pages 4533–4543. PMLR, 2021.
  25. 25.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, 36(7):1325–1339, 2013.
  26. 26.Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena. Structural-rnn: Deep learning on spatio-temporal graphs. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 5308–5317, 2016.
  27. 27.Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael JL Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. arXiv preprint arXiv:2009.01411, 2020.
  28. 28.Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel. Neural relational inference for interacting systems. In International Conference on Machine Learning, pages 2688–2697. PMLR, 2018.
  29. 29.Kris M Kitani, Brian D Ziebart, James Andrew Bagnell, and Martial Hebert. Activity forecasting. In European conference on computer vision, pages 201–214. Springer, 2012.
  30. 30.Miltiadis Kofinas, Naveen Nagaraja, and Efstratios Gavves. Roto-translated local coordinate frames for interacting dynamical systems. Advances in Neural Information Processing Systems, 34:6417–6429, 2021.
  31. 31.Jonas Köhler, Leon Klein, and Frank Noé. Equivariant flows: sampling configurations for multi-body systems with symmetric energies. arXiv preprint arXiv:1910.00753, 2019.
  32. 32.Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker. Desire: Distant future prediction in dynamic scenes with interacting agents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 336–345, 2017.
  33. 33.Andreas M Lehrmann, Peter V Gehler, and Sebastian Nowozin. Efficient nonlinear markov models for human motion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1314–1321, 2014.
  34. 34.Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. Crowds by example. In Computer graphics forum, volume 26, pages 655–664. Wiley Online Library, 2007.
  35. 35.Jesse Levinson, Jake Askeland, Jan Becker, Jennifer Dolson, David Held, Soeren Kammel, J Zico Kolter, Dirk Langer, Oliver Pink, Vaughan Pratt, et al. Towards fully autonomous driving: Systems and algorithms. In 2011 IEEE intelligent vehicles symposium (IV), pages 163–168. IEEE, 2011.
  36. 36.Chen Li, Zhen Zhang, Wee Sun Lee, and Gim Hee Lee. Convolutional sequence to sequence model for human dynamics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5226–5234, 2018.
  37. 37.Jiachen Li, Fan Yang, Masayoshi Tomizuka, and Chiho Choi. Evolvegraph: Multi-agent trajectory prediction with dynamic relational reasoning. Proceedings of the Neural Information Processing Systems (NeurIPS), 2020.
  38. 38.Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3595–3603, 2019.
  39. 39.Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. Symbiotic graph neural networks for 3d skeleton-based human action recognition and motion prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6):3316–3333, 2021.
  40. 40.Maosen Li, Siheng Chen, Zijing Zhang, Lingxi Xie, Qi Tian, and Ya Zhang. Skeleton-parted graph scattering networks for 3d human motion prediction. In European Conference on Computer Vision, 2022.
  41. 41.Maosen Li, Siheng Chen, Yangheng Zhao, Ya Zhang, Yanfeng Wang, and Qi Tian. Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 214–223, 2020.
  42. 42.Junwei Liang, Lu Jiang, Kevin Murphy, Ting Yu, and Alexander Hauptmann. The garden of forking paths: Towards multi-future trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10508–10518, 2020.
  43. 43.Ming Liang, Bin Yang, Rui Hu, Yun Chen, Renjie Liao, Song Feng, and Raquel Urtasun. Learning lane graph representations for motion forecasting. In European Conference on Computer Vision, pages 541–556. Springer, 2020.
  44. 44.Tiezheng Ma, Yongwei Nie, Chengjiang Long, Qing Zhang, and Guiqing Li. Progressively generating better initial guesses towards next stages for high-quality human motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6437–6446, 2022.
  45. 45.Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan-Hui Lee, Ehsan Adeli, Jitendra Malik, and Adrien Gaidon. It is not the journey but the destination: Endpoint conditioned trajectory prediction. In European Conference on Computer Vision, pages 759–776. Springer, 2020.
  46. 46.Wei Mao, Miaomiao Liu, and Mathieu Salzmann. History repeats itself: Human motion prediction via motion attention. In European Conference on Computer Vision, pages 474–489. Springer, 2020.
  47. 47.Wei Mao, Miaomiao Liu, Mathieu Salzmann, and Hongdong Li. Learning trajectory dependencies for human motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9489–9497, 2019.
  48. 48.Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, and Alberto Del Bimbo. Mantra: Memory augmented networks for multiple trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7143–7152, 2020.
  49. 49.Diego Marcos, Michele Volpi, Nikos Komodakis, and Devis Tuia. Rotation equivariant vector field networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 5048–5057, 2017.
  50. 50.Julieta Martinez, Michael J Black, and Javier Romero. On human motion prediction using recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2891–2900, 2017.
  51. 51.Ramin Mehran, Alexis Oyama, and Mubarak Shah. Abnormal crowd behavior detection using social force model. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 935–942. IEEE, 2009.
  52. 52.Jeremy Morton, Tim A Wheeler, and Mykel J Kochenderfer. Analysis of recurrent neural networks for probabilistic modeling of driver behavior. IEEE Transactions on Intelligent Transportation Systems, 18(5):1289–1298, 2016.
  53. 53.Damian Mrowca, Chengxu Zhuang, Elias Wang, Nick Haber, Li F Fei-Fei, Josh Tenenbaum, and Daniel L Yamins. Flexible neural representation for physics prediction. Advances in neural information processing systems, 31, 2018.
  54. 54.Stefano Pellegrini, Andreas Ess, Konrad Schindler, and Luc Van Gool. You’ll never walk alone: Modeling social behavior for multi-target tracking. In 2009 IEEE 12th International Conference on Computer Vision, pages 261–268. IEEE, 2009.
  55. 55.Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Multi-agent generative trajectory forecasting with heterogeneous data for control. 2020.
  56. 56.Alvaro Sanchez-Gonzalez, Victor Bapst, Kyle Cranmer, and Peter Battaglia. Hamiltonian graph networks with ode integrators. arXiv preprint arXiv:1909.12790, 2019.
  57. 57.Alvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying, Jure Leskovec, and Peter Battaglia. Learning to simulate complex physics with graph networks. In International Conference on Machine Learning, pages 8459–8468. PMLR, 2020.
  58. 58.Vıctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E (n) equivariant graph neural networks. In International conference on machine learning, pages 9323–9332. PMLR, 2021.
  59. 59.Bohan Tang, Yiqi Zhong, Ulrich Neumann, Gang Wang, Ya Zhang, and Siheng Chen. Collaborative uncertainty in multi-agent trajectory forecasting. Advances in Neural Information Processing Systems, 34, 2021.
  60. 60.Graham W Taylor and Geoffrey E Hinton. Factored conditional restricted boltzmann machines for modeling motion style. In Proceedings of the 26th annual international conference on machine learning, pages 1025–1032, 2009.
  61. 61.Nathaniel Thomas, Tess Smidt, Steven Kearnes, Lusann Yang, Li Li, Kai Kohlhoff, and Patrick Riley. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds. arXiv preprint arXiv:1802.08219, 2018.
  62. 62.Benjamin Ummenhofer, Lukas Prantl, Nils Thuerey, and Vladlen Koltun. Lagrangian fluid simulation with continuous convolutions. In International Conference on Learning Representations, 2019.
  63. 63.Anirudh Vemula, Katharina Muelling, and Jean Oh. Social attention: Modeling attention in human crowds. In 2018 IEEE international Conference on Robotics and Automation (ICRA), pages 4601–4607. IEEE, 2018.
  64. 64.Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert. The pose knows: Video forecasting by generating pose futures. In Proceedings of the IEEE international conference on computer vision, pages 3332–3341, 2017.
  65. 65.Jack M Wang, David J Fleet, and Aaron Hertzmann. Gaussian process dynamical models for human motion. IEEE transactions on pattern analysis and machine intelligence, 30(2):283–298, 2007.
  66. 66.Maurice Weiler, Fred A Hamprecht, and Martin Storath. Learning steerable filters for rotation equivariant cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 849–858, 2018.
  67. 67.Daniel E Worrall, Stephan J Garbin, Daniyar Turmukhambetov, and Gabriel J Brostow. Harmonic networks: Deep translation and rotation equivariance. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5028–5037, 2017.
  68. 68.Chenxin Xu, Siheng Chen, Maosen Li, and Ya Zhang. Invariant teacher and equivariant student for unsupervised 3d human pose estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3013–3021, 2021.
  69. 69.Chenxin Xu, Maosen Li, Zhenyang Ni, Ya Zhang, and Siheng Chen. Groupnet: Multiscale hypergraph neural networks for trajectory prediction with relational reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6498–6507, 2022.
  70. 70.Chenxin Xu, Weibo Mao, Wenjun Zhang, and Siheng Chen. Remember intentions: Retrospective-memory-based trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6488–6497, 2022.
  71. 71.Chenxin Xu, Yuxi Wei, Bohan Tang, Sheng Yin, Ya Zhang, and Siheng Chen. Dynamic-group-aware networks for multi-agent trajectory prediction with relational reasoning. arXiv preprint arXiv:2206.13114, 2022.
  72. 72.Xingyi Yang, Jingwen Ye, and Xinchao Wang. Factorizing knowledge in neural networks. In European Conference on Computer Vision, 2022.
  73. 73.Xingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye, and Xinchao Wang. Deep model reassembly. In Conference on Neural Information Processing Systems, 2022.
  74. 74.Yiding Yang, Zunlei Feng, Mingli Song, and Xinchao Wang. Factorizable graph convolutional networks. In Conference on Neural Information Processing Systems, 2020.
  75. 75.Yiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao, and Xinchao Wang. Distilling knowledge from graph convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  76. 76.Ye Yuan, Xinshuo Weng, Yanglan Ou, and Kris M Kitani. Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9813–9823, 2021.
  77. 77.Yiqi Zhong, Zhenyang Ni, Siheng Chen, and Ulrich Neumann. Aware of the history: Trajectory forecasting with the local behavior data. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, pages 393–409. Springer, 2022.

Citation

MLA
Xu, C., et al. “EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning”. arXiv, 2023, http://arxiv.org/abs/2303.10876v2.
APA
Xu, C., Tan, R. T., Tan, Y., Chen, S., Wang, Y. G., Wang, X., & Wang, Y. (2023). EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning. arXiv. http://arxiv.org/abs/2303.10876v2
Chicago
Xu, C., R. T. Tan, Y. Tan, et al. 2023. “EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning”. arXiv. http://arxiv.org/abs/2303.10876v2.
Harvard
Xu, C. et al. (2023) “EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.10876v2.
Vancouver
1. Xu C, Tan RT, Tan Y, Chen S, Wang YG, Wang X, Wang Y (2023) EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning. arXiv

BibTeX

@article{xu2023eqmotion,
  title = {EqMotion: Equivariant Multi-agent Motion Prediction with Invariant Interaction Reasoning},
  author = {Xu, Chenxin and Tan, Robby T. and Tan, Yuhong and Chen, Siheng and Wang, Yu Guang and Wang, Xinchao and Wang, Yanfeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.10876v2},
  eprint = {2303.10876}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/