Adaptive Trajectory Prediction via Transferable GNN

Yi XuLichen WangYizhou WangYun Fu

article2022CVPR105 citations

Proposes a transferable graph neural network framework that aligns feature distributions across disparate environments to prevent performance drops in cross-domain pedestrian trajectory forecasting.

Listen

Accurately predicting pedestrian movement is essential for safety, navigation, and automated planning in autonomous driving and robotics. Current prediction systems typically assume that pedestrian movement patterns observed during training will match those encountered during real-world deployment. In practice, however, pedestrian density, walking speeds, and acceleration vary dramatically across different physical environments, such as open streets versus crowded university plazas. This environmental mismatch creates a domain shift problem that severely degrades the accuracy of standard predictive models.

The article demonstrates this vulnerability across standard benchmarks and introduces a new predictive framework called the Transferable Graph Neural Network to overcome it. The primary objective is to evaluate whether integrating unsupervised domain adaptation directly into a network can bridge distribution gaps and generate reliable pedestrian trajectory predictions in new, unseen target environments.

To evaluate this framework, the authors designed a transfer-learning benchmark across five real-world pedestrian datasets: ETH, Hotel, University, Zara1, and Zara2. This benchmark generated 20 cross-scene prediction tasks where models trained on one source environment were required to adapt to a distinct target environment. The framework uses graph neural network layers to capture spatial relationships and pedestrian interactions, combines them with an attention mechanism to extract fine-grained, individual-level movement features, and aligns source and target feature distributions using a distance-minimizing alignment loss. Crucially, the model accesses only the past observed trajectory history of the new target scene without requiring future trajectory labels, reflecting realistic deployment constraints.

The experimental findings show substantial improvements over existing techniques. Across all 20 cross-domain tasks, the proposed framework achieved an average prediction error of 0.96 meters and a final endpoint error of 1.82 meters, consistently outperforming five leading baseline models. Compared to top-performing baselines like Social-STGCNN, SGCN, and PECNet, the framework reduced average trajectory error by approximately 21.3% and final displacement error by roughly 20.5%. Furthermore, ablation tests confirmed that incorporating an individual-level attention module and explicit domain alignment significantly outperforms simply pooling features or training models on mixed datasets without alignment.

These findings have direct operational and safety implications for autonomous vehicles and mobile robotics. High prediction error in unfamiliar surroundings increases collision risks, delays vehicle decision-making, and limits operational scalability. The results show that deploying models naively across disparate environments without domain alignment introduces severe bias, whereas an adaptive framework allows systems to transfer learned motion dynamics to new operating domains without costly, manual retraining or ground-truth data collection.

Organizations developing autonomous navigation systems should adopt domain-invariant learning frameworks to make predictive models resilient to environmental shifts. Development teams should prioritize fine-grained individual feature alignment over standard sample-level domain adaptation techniques. Before deploying these models into safety-critical operations, organizations should conduct pilot trials in edge-case environments characterized by extreme crowd densities or non-standard movement dynamics to validate operational safety thresholds.

While the evaluation provides high confidence across standard cross-scene benchmarks, the scope of the evaluation relies on five public surveillance-style datasets with specific observation-to-prediction ratios. Readers should account for uncertainties in unseen edge environments that may feature significantly more complex occlusions, high vehicle-pedestrian mixed traffic, or sensor noise not present in these standard datasets.

arXiv: 2203.05046
Cover for Adaptive Trajectory Prediction via Transferable GNN

Abstract

Pedestrian trajectory prediction is an essential component in a wide range of AI applications such as autonomous driving and robotics. Existing methods usually assume the training and testing motions follow the same pattern while ignoring the potential distribution differences (e.g., shopping mall and street). This issue results in inevitable performance decrease. To address this issue, we propose a novel Transferable Graph Neural Network (T-GNN) framework, which jointly conducts trajectory prediction as well as domain alignment in a unified framework. Specifically, a domain-invariant GNN is proposed to explore the structural motion knowledge where the domain-specific knowledge is reduced. Moreover, an attention-based adaptive knowledge learning module is further proposed to explore fine-grained individual-level feature representations for knowledge transfer. By this way, disparities across different trajectory domains will be better alleviated. More challenging while practical trajectory prediction experiments are designed, and the experimental results verify the superior performance of our proposed model. To the best of our knowledge, our work is the pioneer which fills the gap in benchmarks and techniques for practical pedestrian trajectory prediction across different domains.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Forecasting Pedestrian Trajectory
  • 2.2. Graph-Involved Forecasting Models
  • 2.3. Domain Adaptation
  • 3. Our Method
  • 3.1. Problem Definition
  • 3.2. Spatial-Temporal Feature Representations
  • 3.3. Attention-Based Adaptive Learning
  • 3.4. Temporal Prediction Module
  • 3.5. Objective Function
  • 4. Experiments
  • 4.1. Quantitative Analysis
  • 4.2. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Transferable Graph Neural Network Framework for Cross-Domain Trajectory Prediction

    model/method

    The Transferable Graph Neural Network (T-GNN) framework performs pedestrian trajectory forecasting under domain shift between a source domain (with NsN_s pedestrians) and an unlabelled target domain (with NtN_t pedestrians). Given observed pedestrian trajectories Γ={Γ1,…,ΓN}\Gamma = \{\Gamma^1, \dots, \Gamma^N\} from time step t=1t = 1 to TobsT_{obs}, where each pedestrian's trajectory is Γi={o1i,…,oobsi}\Gamma^i = \{o_1^i, \dots, o_{obs}^i\} with coordinates oti=(xti,yti)∈R2o_t^i = (x_t^i, y_t^i) \in \mathbb{R}^2, the model forecasts future trajectories Γˉ={Γˉ1,…,ΓˉN}\bar{\Gamma} = \{\bar{\Gamma}^1, \dots, \bar{\Gamma}^N\} from t=Tobs+1t = T_{obs}+1 to TpredT_{pred}, where Γˉi={oobs+1i,…,opredi}\bar{\Gamma}^i = \{o_{obs+1}^i, \dots, o_{pred}^i\}.

    T-GNN unifies trajectory forecasting and unsupervised domain adaptation through three interacting components:

    1. A weight-shared Spatial-Temporal Graph Neural Network (GCN) that constructs dynamic graphs for source and target scenes, extracting spatial-temporal feature matrices F(s)∈RNs×Df×LobsF_{(s)} \in \mathbb{R}^{N_s \times D_f \times L_{obs}} and F(t)∈RNt×Df×LobsF_{(t)} \in \mathbb{R}^{N_t \times D_f \times L_{obs}}, where DfD_f is feature dimension and Lobs=TobsL_{obs} = T_{obs} is the observation length.
    2. An Attention-Based Adaptive Knowledge Learning Module that collapses individual pedestrian representations into domain-level context vectors c(s),c(t)∈RDvc_{(s)}, c_{(t)} \in \mathbb{R}^{D_v} (with Dv=Df×LobsD_v = D_f \times L_{obs}) using attention weights, allowing computation of a distribution alignment loss Lalign\mathcal{L}_{align} despite variable pedestrian counts (Ns≠NtN_s \neq N_t).
    3. A Temporal Prediction Module based on Temporal Convolutional Networks (TCN) that maps source features F(s)F_{(s)} to future trajectory distributions parameterized as bivariate Gaussians, optimized via a prediction loss Lpre\mathcal{L}_{pre}.

    The entire network is trained end-to-end by jointly minimizing the weighted multi-task objective L=Lpre+λLalign\mathcal{L} = \mathcal{L}_{pre} + \lambda \mathcal{L}_{align}.

  2. Knowl 2 — Spatial-Temporal Feature Representation with Dynamic Graph Attention and GCN

    model/method

    To extract spatial-temporal interaction features across pedestrians while eliminating scene size biases, T-GNN employs coordinate decentralization, dynamic graph attention, and cascaded graph convolutions.

    1. Coordinate Decentralization: Coordinates of all NN pedestrians in a scene are normalized relative to the mean position at the final observed frame TobsT_{obs}: o′ti=oti−1N∑i=1Noobsi{o'}_t^i = o_t^i - \frac{1}{N}\sum_{i=1}^N o_{obs}^i where o′ti=(x′ti,y′ti){o'}_t^i = ({x'}_t^i, {y'}_t^i) are relative coordinates.

    2. Graph Construction and Feature Embedding: At time step tt, the graph Gt=(Vt,Et,Ft)G_t = (V_t, E_t, F_t) consists of vertices Vt={vt;i∣i=1,…,N}V_t = \{v_{t; i} \mid i = 1, \dots, N\}, feature matrix Ft={ft;i∣i=1,…,N}∈RN×DfF_t = \{f_{t; i} \mid i = 1, \dots, N\} \in \mathbb{R}^{N \times D_f} with embeddings ft;i=σ((x′ti,y′ti)Wo)f_{t; i} = \sigma(({x'}_t^i, {y'}_t^i) W_o) (where Wo∈R2×DfW_o \in \mathbb{R}^{2 \times D_f} is a learnable projection matrix and σ(⋅)\sigma(\cdot) is the ReLU activation function), and adjacency matrix At={at;i,j}∈RN×NA_t = \{a_{t; i,j}\} \in \mathbb{R}^{N \times N} initialized with pairwise Euclidean distances: at;i,j=∥o′ti−o′tj∥2a_{t; i,j} = \|{o'}_t^i - {o'}_t^j\|_2

    3. Dynamic Graph Attention Layer (GAL): Dynamic spatial relations are updated via graph attention coefficients αt;i,j\alpha_{t; i, j}: αt;i,j=exp⁡(ϕ(Wl[at;i⊕at;j]))∑j=1Nexp⁡(ϕ(Wl[at;i⊕at;j]))\alpha_{t; i,j} = \frac{\exp\left(\phi\left(W_l [a_{t; i} \oplus a_{t; j}]\right)\right)}{\sum_{j=1}^N \exp\left(\phi\left(W_l [a_{t; i} \oplus a_{t; j}]\right)\right)} where at;i∈RN×1a_{t; i} \in \mathbb{R}^{N \times 1} is the ii-th column of AtA_t, ⊕\oplus is row concatenation, Wl∈R1×2NW_l \in \mathbb{R}^{1 \times 2N} is a learnable parameter vector, and ϕ(⋅)\phi(\cdot) is LeakyReLU with negative slope θ=0.2\theta = 0.2. The updated adjacency column is pt;i=σ(∑j=1Nαt;i,jat;j)p_{t; i} = \sigma\left(\sum_{j=1}^N \alpha_{t; i, j} a_{t; j}\right), yielding updated adjacency matrix At′∈RN×NA'_t \in \mathbb{R}^{N \times N}.

    4. Spatial-Temporal Graph Convolution: Self-loops are added via A^t=At′+I\hat{A}_t = A'_t + I. Stacking across observation steps t=1,…,Tobst = 1, \dots, T_{obs} forms A^={A^1,…,A^obs}∈RN×N×Lobs\hat{A} = \{\hat{A}_1, \dots, \hat{A}_{obs}\} \in \mathbb{R}^{N \times N \times L_{obs}} and degree matrices D^={D^1,…,D^obs}\hat{D} = \{\hat{D}_1, \dots, \hat{D}_{obs}\}. The ll-th GCN layer output is computed as: F(l+1)=σ(D^−12A^D^−12F(l)W(l))F^{(l+1)} = \sigma\left(\hat{D}^{-\frac{1}{2}} \hat{A} \hat{D}^{-\frac{1}{2}} F^{(l)} W^{(l)}\right) where W(l)W^{(l)} are trainable parameters. Three cascaded GCN layers (l=3l = 3) extract final spatial-temporal representation matrices F(s)∈RNs×Df×LobsF_{(s)} \in \mathbb{R}^{N_s \times D_f \times L_{obs}} and F(t)∈RNt×Df×LobsF_{(t)} \in \mathbb{R}^{N_t \times D_f \times L_{obs}} for source and target scenes.

  3. Knowl 3 — Attention-Based Individual-Level Domain Alignment Module

    model/method

    In cross-domain trajectory prediction, standard domain adaptation methods cannot align global scene representations directly because the number of pedestrians NsN_s in the source domain and NtN_t in the target domain vary per frame. The attention-based adaptive knowledge learning module overcomes this misalignment by computing individual-level attention weights over pedestrian trajectories to form domain context vectors.

    Feature maps F(s)∈RNs×Df×LobsF_{(s)} \in \mathbb{R}^{N_s \times D_f \times L_{obs}} and F(t)∈RNt×Df×LobsF_{(t)} \in \mathbb{R}^{N_t \times D_f \times L_{obs}} are reformatted into collections of pedestrian vectors: F(s)=[f1(s),f2(s),…,fNs(s)],fi(s)∈RDvF_{(s)} = [f_1^{(s)}, f_2^{(s)}, \dots, f_{N_s}^{(s)}], \quad f_i^{(s)} \in \mathbb{R}^{D_v} F(t)=[f1(t),f2(t),…,fNt(t)],fi(t)∈RDvF_{(t)} = [f_1^{(t)}, f_2^{(t)}, \dots, f_{N_t}^{(t)}], \quad f_i^{(t)} \in \mathbb{R}^{D_v} where Dv=Df×LobsD_v = D_f \times L_{obs}, DfD_f is feature dimensionality, and LobsL_{obs} is observation length.

    Individual representativeness scores βi(s)\beta_i^{(s)} and βi(t)\beta_i^{(t)} within their respective domains are computed using shared parameters h∈RDvh \in \mathbb{R}^{D_v} and Wf∈RDv×DvW_f \in \mathbb{R}^{D_v \times D_v}: βi(s)=exp⁡(h⊤tanh⁡(Wffi(s)))∑j=1Nsexp⁡(h⊤tanh⁡(Wffj(s)))\beta_i^{(s)} = \frac{\exp\left(h^\top \tanh\left(W_f f_i^{(s)}\right)\right)}{\sum_{j=1}^{N_s} \exp\left(h^\top \tanh\left(W_f f_j^{(s)}\right)\right)} βi(t)=exp⁡(h⊤tanh⁡(Wffi(t)))∑j=1Ntexp⁡(h⊤tanh⁡(Wffj(t)))\beta_i^{(t)} = \frac{\exp\left(h^\top \tanh\left(W_f f_i^{(t)}\right)\right)}{\sum_{j=1}^{N_t} \exp\left(h^\top \tanh\left(W_f f_j^{(t)}\right)\right)}

    The refined domain-level context vectors c(s),c(t)∈RDvc_{(s)}, c_{(t)} \in \mathbb{R}^{D_v} are weighted sums over pedestrians: c(s)=∑i=1Nsβi(s)fi(s),c(t)=∑i=1Ntβi(t)fi(t)c_{(s)} = \sum_{i=1}^{N_s} \beta_i^{(s)} f_i^{(s)}, \quad c_{(t)} = \sum_{i=1}^{N_t} \beta_i^{(t)} f_i^{(t)}

    The distribution alignment loss Lalign\mathcal{L}_{align} is calculated as the normalized squared L2L_2 distance between context vectors: Lalign=1Df∥c(s)−c(t)∥22\mathcal{L}_{align} = \frac{1}{D_f} \|c_{(s)} - c_{(t)}\|_2^2

  4. Knowl 4 — Temporal Prediction Module and Bi-variate Gaussian Training Objective

    model/method

    The temporal prediction module generates future trajectories from source domain feature representations F(s)∈RNs×Df×LobsF_{(s)} \in \mathbb{R}^{N_s \times D_f \times L_{obs}} using Temporal Convolutional Networks (TCN) along the temporal dimension, avoiding sequential RNN error accumulation and gradient vanishing.

    The layer-wise operation across three cascaded TCN layers (l=3l = 3) is: F(s)(l+1)=TCN(F(s)(l);Wt(l))F_{(s)}^{(l+1)} = \text{TCN}\left(F_{(s)}^{(l)}; W_t^{(l)}\right) where Wt(l)W_t^{(l)} denotes trainable parameters, producing prediction features F(s),pred∈RNs×Df×LpredF_{(s), pred} \in \mathbb{R}^{N_s \times D_f \times L_{pred}}, with Lpred=Tpred−TobsL_{pred} = T_{pred} - T_{obs}.

    A linear layer parameterizes a bivariate Gaussian distribution for each pedestrian ii at future time step t∈{Tobs+1,…,Tpred}t \in \{T_{obs}+1, \dots, T_{pred}\}: (μ^ti,σ^ti,ρ^ti)=Linear(F(s),pred;Wp)(\hat{\mu}_t^i, \hat{\sigma}_t^i, \hat{\rho}_t^i) = \text{Linear}\left(F_{(s), pred}; W_p\right) where μ^ti=(μ^x,μ^y)ti\hat{\mu}_t^i = (\hat{\mu}_x, \hat{\mu}_y)_t^i is the predicted mean position, σ^ti=(σ^x,σ^y)ti\hat{\sigma}_t^i = (\hat{\sigma}_x, \hat{\sigma}_y)_t^i represents standard deviations, and ρ^ti\hat{\rho}_t^i is the cross-coordinate correlation coefficient. Future pedestrian positions are modeled as (xti,yti)∼N(μ^ti,σ^ti,ρ^ti)(x_t^i, y_t^i) \sim \mathcal{N}(\hat{\mu}_t^i, \hat{\sigma}_t^i, \hat{\rho}_t^i).

    The trajectory prediction loss Lpre\mathcal{L}_{pre} is the negative log-likelihood of ground-truth future positions: Lpre=−∑t=Tobs+1Tpredlog⁡P((xti,yti)∣μ^ti,σ^ti,ρ^ti)\mathcal{L}_{pre} = -\sum_{t=T_{obs}+1}^{T_{pred}} \log P\left((x_t^i, y_t^i) \mid \hat{\mu}_t^i, \hat{\sigma}_t^i, \hat{\rho}_t^i\right)

    The total training loss jointly optimizes prediction accuracy on the source domain and feature alignment between source and target domains: L=Lpre+λLalign\mathcal{L} = \mathcal{L}_{pre} + \lambda \mathcal{L}_{align} where λ\lambda balances prediction and alignment losses (set empirically to λ=1\lambda = 1).

  5. Knowl 5 — Cross-Domain Trajectory Prediction Benchmark and Evaluation Protocol

    experimental setup

    The cross-domain trajectory prediction benchmark evaluates models under distribution shifts across five real-world scenes from the ETH dataset (ETH, HOTEL) and UCY dataset (UNIV, ZARA1, ZARA2), denoted as domains A, B, C, D, and E.

    Experimental Tasks: Each scene serves as a distinct domain, creating 20 directional cross-domain transfer tasks: A→{B,C,D,E}A \to \{B, C, D, E\}, B→{A,C,D,E}B \to \{A, C, D, E\}, C→{A,B,D,E}C \to \{A, B, D, E\}, D→{A,B,C,E}D \to \{A, B, C, E\}, and E→{A,B,C,D}E \to \{A, B, C, D\}.

    Evaluation Protocol: For a task S→TS \to T (source SS, target TT):

    • Standard baselines are trained on the training set of SS concatenated with the validation set of TT.
    • Unsupervised domain adaptation models (including T-GNN) are trained on the training set of SS (with full trajectory labels) and only the observed trajectory history (T1T_1 to TobsT_{obs}) of the validation set of TT (without access to target future trajectories during training).
    • All models are evaluated exclusively on the independent, non-overlapping testing set of TT.

    Configuration and Hyperparameters:

    • Observation horizon: Lobs=8L_{obs} = 8 frames; prediction horizon: Lpred=12L_{pred} = 12 frames.
    • Network structure: 3 GCN layers, 3 TCN layers, feature dimension Df=64D_f = 64, batch size 16, alignment trade-off parameter λ=1\lambda = 1.
    • Training: Adam optimizer for 200 epochs with initial learning rate 0.001, decayed to 0.0005 at epoch 100.
    • Inference: 20 trajectory samples are generated per pedestrian; the best-of-20 prediction is evaluated.

    Evaluation Metrics:

    • Average Displacement Error (ADE): ADE=∑i=1Nt∑t=Tobs+1Tpred∥o^ti−oti∥2Nt(Tpred−Tobs)\text{ADE} = \frac{\sum_{i=1}^{N_t} \sum_{t=T_{obs}+1}^{T_{pred}} \|\hat{o}_t^i - o_t^i\|_2}{N_t(T_{pred} - T_{obs})}
    • Final Displacement Error (FDE): FDE=∑i=1Nt∥o^predi−opredi∥2Nt\text{FDE} = \frac{\sum_{i=1}^{N_t} \|\hat{o}_{pred}^i - o_{pred}^i\|_2}{N_t} where NtN_t is total pedestrians in the target domain, o^ti\hat{o}_t^i is predicted position, and otio_t^i is ground truth.
  6. Knowl 6 — Accuracy Comparison Across 20 Cross-Domain Trajectory Prediction Tasks

    data/table

    Performance of T-GNN compared to five baseline models (Social-STGCNN, PECNet, RSBG, Tra2Tra, SGCN) across all 20 cross-domain transfer tasks evaluated on the ETH/UCY datasets (A: ETH, B: HOTEL, C: UNIV, D: ZARA1, E: ZARA2) using Average Displacement Error (ADE) and Final Displacement Error (FDE) in meters.

    Method A2B A2C A2D A2E B2A B2C B2D B2E C2A C2B C2D C2E D2A D2B D2C D2E E2A E2B E2C E2D Ave
    Average Displacement Error (ADE, meters)
    Social-STGCNN 1.83 1.58 1.30 1.31 3.02 1.38 2.63 1.58 1.16 0.70 0.82 0.54 1.04 1.05 0.73 0.47 0.98 1.09 0.74 0.50 1.22
    PECNet 1.97 1.68 1.24 1.35 3.11 1.35 2.69 1.62 1.39 0.82 0.93 0.57 1.10 1.17 0.92 0.52 1.01 1.25 0.83 0.61 1.31
    RSBG 2.21 1.59 1.48 1.42 3.18 1.49 2.72 1.73 1.23 0.87 1.04 0.60 1.19 1.21 0.80 0.49 1.09 1.37 1.03 0.78 1.38
    Tra2Tra 1.72 1.58 1.27 1.37 3.32 1.36 2.67 1.58 1.16 0.70 0.85 0.60 1.09 1.07 0.81 0.52 1.03 1.10 0.75 0.52 1.25
    SGCN 1.68 1.54 1.26 1.28 3.22 1.38 2.62 1.58 1.14 0.70 0.82 0.52 1.05 0.97 0.80 0.48 0.97 1.08 0.75 0.51 1.22
    T-GNN (Ours) 1.13 1.25 0.94 1.03 2.54 1.08 2.25 1.41 0.97 0.54 0.61 0.23 0.88 0.78 0.59 0.32 0.87 0.72 0.65 0.34 0.96
    Final Displacement Error (FDE, meters)
    Social-STGCNN 3.24 2.86 2.53 2.43 5.16 2.51 4.86 2.88 2.30 1.34 1.74 1.10 2.21 1.99 1.41 0.88 2.10 2.05 1.47 1.01 2.30
    PECNet 3.33 2.83 2.53 2.45 5.23 2.48 4.90 2.86 2.22 1.32 1.68 1.12 2.20 2.05 1.52 0.88 2.10 1.84 1.45 0.98 2.29
    RSBG 3.42 2.96 2.75 2.50 5.28 2.59 5.19 3.10 2.36 1.55 1.99 1.37 2.28 2.22 1.77 0.97 2.19 2.29 1.81 1.34 2.50
    Tra2Tra 3.29 2.88 2.66 2.45 5.22 2.50 4.89 2.90 2.29 1.33 1.78 1.09 2.26 2.12 1.63 0.92 2.18 2.06 1.52 1.17 2.34
    SGCN 3.22 2.81 2.52 2.40 5.18 2.47 4.83 2.85 2.24 1.32 1.71 1.03 2.23 1.90 1.48 0.97 2.10 1.95 1.52 0.99 2.29
    T-GNN (Ours) 2.18 2.25 1.78 1.84 4.15 1.82 4.04 2.53 1.91 1.12 1.30 0.87 1.92 1.46 1.25 0.65 1.86 1.45 1.28 0.72 1.82

    T-GNN consistently outperforms all competing baselines across every transfer task. Overall, T-GNN achieves an average ADE of 0.96 m (a 21.31% error reduction compared to Social-STGCNN and SGCN at 1.22 m) and an average FDE of 1.82 m (a 20.52% error reduction compared to PECNet and SGCN at 2.29 m).

  7. Knowl 7 — Domain Discrepancy Statistics Across Trajectory Datasets

    data/table

    Statistical analysis across five pedestrian trajectory scenes demonstrates substantial domain gaps in crowd density, velocity, and acceleration patterns.

    Metric ETH HOTEL UNIV ZARA1 ZARA2 E-D S-D
    NoS 70 301 947 602 921 877 383.63
    NoP 181 1053 24334 2253 5833 24153 10073.07
    AN 2.586 3.498 25.696 3.743 6.333 23.11 9.78
    AV (m/s) 0.437 0.178 0.205 0.369 0.206 0.259 0.11
    AA (m/s2^2) 0.131 0.060 0.035 0.039 0.026 0.105 0.04

    Metric definitions:

    • NoS: number of trajectory sequences to be predicted.
    • NoP: total number of unique pedestrians across the scene.
    • AN: average number of pedestrians present in each sequence.
    • AV: average velocity of pedestrians per sequence in meters per second (m/s).
    • AA: average acceleration of pedestrians per sequence in meters per second squared (m/s2^2).
    • E-D: extreme deviation (maximum minus minimum value across scenes).
    • S-D: standard deviation across scenes.

    These statistics reveal significant domain disparities: pedestrian density (AN) in UNIV (25.696) is nearly 10×10\times that in ETH (2.586); mean velocity (AV) in ETH (0.437 m/s) is roughly 2.5×2.5\times higher than in HOTEL (0.178 m/s); and mean acceleration (AA) in ETH (0.131 m/s2^2) is 5×5\times higher than in ZARA2 (0.026 m/s2^2).

  8. Knowl 8 — Comparison of Domain Adaptation Distance Measures for Trajectory Transfer

    data/table

    Average trajectory prediction performance (ADE / FDE in meters) across all 20 transfer tasks when varying the domain discrepancy alignment loss Lalign\mathcal{L}_{align} within the T-GNN architecture.

    Method Average ADE / FDE (m)
    T-GNN + MMD 1.11 / 2.11
    T-GNN + CORAL 1.07 / 2.01
    T-GNN + GFK 1.15 / 2.08
    T-GNN + UDA 1.07 / 2.09
    T-GNN (Ours, L2L_2 context distance) 0.96 / 1.82

    Compared DA strategies:

    • T-GNN + MMD: Multi-Kernel Maximum Mean Discrepancies loss.
    • T-GNN + CORAL: Deep Correlation Alignment (Deep CORAL) loss.
    • T-GNN + GFK: Geodesic Flow Kernel domain adaptation strategy.
    • T-GNN + UDA: Unsupervised domain adaptive graph convolutional network using an adversarial domain discriminator loss.
    • T-GNN (L2L_2): Normalized squared L2L_2 distance between attention-weighted context vectors c(s)c_{(s)} and c(t)c_{(t)}.

    The attention-weighted L2L_2 distance achieves superior performance compared to distribution-matching and adversarial losses, indicating that individual-weighted feature representations better preserve the underlying spatial characteristics needed for trajectory forecasting.

  9. Knowl 9 — Ablation Study on Architecture Components and Loss Balancing Weight

    data/table

    Ablation experiments evaluate the contribution of individual T-GNN components on five transfer tasks (A2B, B2C, C2D, D2E, E2A; A: ETH, B: HOTEL, C: UNIV, D: ZARA1, E: ZARA2) alongside the sensitivity of performance to the trade-off parameter λ\lambda.

    Component Variants:

    • T-GNN w/o GAL: Dynamic graph attention layer removed; adjacency attention αt;i,j\alpha_{t; i, j} is static.
    • T-GNN w/o AAL w/ AP: Attention-based adaptive learning module replaced with average pooling across pedestrians.
    • T-GNN w/o AAL w/ LL: Attention-based adaptive learning module replaced with a single trainable linear layer.
    • Social-STGCNN-V1 / SGCN-V1 / T-GNN-V1: Trained purely on source data without target validation data or adaptation loss.
    • Social-STGCNN / SGCN / T-GNN-V2: Trained directly on mixed source and target validation data without domain adaptation.
    • T-GNN (Ours): Complete model with dynamic graph attention and attention-based context alignment.
    Variants ID A2B B2C C2D D2E E2A
    T-GNN w/o GAL 1 1.51 / 2.34 1.17 / 1.90 0.69 / 1.42 0.39 / 0.71 0.90 / 1.98
    T-GNN w/o AAL w/ AP 2 1.78 / 2.85 1.23 / 2.02 0.77 / 1.53 0.42 / 0.79 0.96 / 2.03
    T-GNN w/o AAL w/ LL 3 1.81 / 2.91 1.25 / 2.03 0.76 / 1.48 0.43 / 0.79 0.94 / 2.01
    Social-STGCNN-V1 4 2.18 / 3.68 2.30 / 3.21 1.59 / 2.54 1.23 / 1.72 1.73 / 2.98
    SGCN-V1 5 2.03 / 3.53 2.35 / 3.22 1.68 / 2.71 1.12 / 1.59 1.81 / 3.02
    T-GNN-V1 6 2.12 / 3.58 2.28 / 3.21 1.73 / 2.76 1.19 / 1.58 1.74 / 2.95
    Social-STGCNN 7 1.83 / 3.24 1.38 / 2.51 0.82 / 1.74 0.47 / 0.88 0.98 / 2.10
    SGCN 8 1.68 / 3.24 1.38 / 2.47 0.82 / 1.71 0.48 / 0.97 0.97 / 2.10
    T-GNN-V2 9 1.89 / 3.25 1.35 / 2.48 0.88 / 1.93 0.53 / 0.97 0.98 / 2.16
    T-GNN (Ours) 10 1.13 / 2.18 1.08 / 1.82 0.61 / 1.30 0.32 / 0.65 0.87 / 1.86

    Sensitivity to Alignment Weight λ\lambda (Average across 20 tasks):

    Metric λ=0.01\lambda = 0.01 λ=0.1\lambda = 0.1 λ=1.0\lambda = 1.0 λ=5.0\lambda = 5.0 λ=10.0\lambda = 10.0
    ADE (m) 1.19 1.05 0.96 1.16 1.31
    FDE (m) 2.16 2.02 1.82 2.07 2.45

    Removing graph attention (Variant 1) or substituting the adaptive attention learning module with average pooling (Variant 2) or a linear layer (Variant 3) causes substantial degradation. Setting λ=1.0\lambda = 1.0 yields optimal balance between trajectory prediction loss and domain alignment loss.

Coverage note — No substantial contributed material was omitted; all key architectural components, objectives, domain shift benchmarks, quantitative results, loss comparisons, and ablation studies are included.

References

  1. 1.Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social LSTM: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 961–971, 2016. 1, 3
  2. 2.Javad Amirian, Jean-Bernard Hayet, and Julien Pettre. Social ways: Learning multi-modal distributions of pedestrian trajectories with GANs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 2964–2972, 2019. 3
  3. 3.Haoyu Bai, Shaojun Cai, Nan Ye, David Hsu, and Wee Sun Lee. Intention-aware online POMDP planning for autonomous driving in a crowd. In Proceedings of the IEEE International Conference on Robotics and Automation, pages 454–460, 2015. 1
  4. 4.Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv:1803.01271, 2018. 5
  5. 5.Niccolo Bisagno, Bo Zhang, and Nicola Conci. Group LSTM: Group trajectory prediction in crowded scenarios. In Proceedings of the European Conference on Computer Vision, pages 213–225, 2018. 3
  6. 6.Ruichu Cai, Fengzhu Wu, Zijian Li, Pengfei Wei, Lingling Yi, and Kun Zhang. Graph domain adaptation: A generative view. arXiv preprint arXiv:2106.07482, 2021. 3
  7. 7.Guangyi Chen, Junlong Li, Jiwen Lu, and Jie Zhou. Human trajectory prediction via counterfactual analysis. In Proceedings of the IEEE International Conference on Computer Vision, pages 9824–9833, 2021. 1, 3
  8. 8.Guangyi Chen, Junlong Li, Nuoxing Zhou, Liangliang Ren, and Jiwen Lu. Personalized trajectory prediction via distribution discrimination. In Proceedings of the IEEE International Conference on Computer Vision, pages 15580–15589, 2021. 1
  9. 9.Hao Cheng, Wentong Liao, Xuejiao Tang, Michael Ying Yang, Monika Sester, and Bodo Rosenhahn. Exploring dynamic context for multi-path trajectory prediction. arXiv preprint arXiv:2010.16267, 2020. 1, 3
  10. 10.Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014. 5
  11. 11.Patrick Dendorfer, Sven Elflein, and Laura Leal-Taixe. Mg-gan: A multi-generator model preventing out-of-distribution samples in pedestrian trajectory prediction. In Proceedings of the IEEE International Conference on Computer Vision, pages 13158–13167, 2021. 1, 3
  12. 12.Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 3
  13. 13.P. Kingma Diederik and Ba Jimmy. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations, 2015. 7
  14. 14.Zhengming Ding, Sheng Li, Ming Shao, and Yun Fu. Graph adaptive knowledge transfer for unsupervised domain adaptation. In Proceedings of the European Conference on Computer Vision, pages 37–52, 2018. 3
  15. 15.David Ellis, Eric Sommerlade, and Ian Reid. Modelling pedestrian trajectory patterns with gaussian processes. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 1229–1234, 2009. 3
  16. 16.Tharindu Fernando, Simon Denman, Sridha Sridharan, and Clinton Fookes. GD-GAN: Generative adversarial networks for trajectory prediction and group detection in crowds. In Proceedings of the Asian Conference on Computer Vision, pages 314–330, 2018. 1, 3
  17. 17.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In Proceedings of the International Conference on Machine Learning, pages 1180–1189, 2015. 2, 3, 5
  18. 18.Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid. VectorNet: Encoding HD maps and agent dynamics from vectorized representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 11522–11530, 2020. 3
  19. 19.Francesco Giuliari, Irtiza Hasan, Marco Cristani, and Fabio Galasso. Transformer networks for trajectory forecasting. In Proceedings of the IEEE International Conference on Pattern Recognition, pages 10335–10342, 2020. 3
  20. 20.Boqing Gong, Yuan Shi, Fei Sha, and K. Grauman. Geodesic flow kernel for unsupervised domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2066–2073, 2015. 3, 6, 7
  21. 21.Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social GAN: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2255–2264, 2018. 1
  22. 22.Gewen He, Xiaofeng Liu, Fangfang Fan, and Jane You. Classification-aware semi-supervised domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 4147–4156, 2020. 3
  23. 23.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997. 5
  24. 24.Yue Hu, Siheng Chen, Ya Zhang, and Xiao Gu. Collaborative motion prediction via neural motion message passing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6319–6328, 2020. 1, 3
  25. 25.Hal Daume III, Abhishek Kumar, and Avishek Saha. Co-regularization based semi-supervised domain adaptation. In Proceedings of the Advances in Neural Information Processing Systems, 2010. 3
  26. 26.Boris Ivanovic and Marco Pavone. The trajectron: Probabilistic multi-agent trajectory modeling with dynamic spatiotemporal graphs. In Proceedings of the IEEE International Conference on Computer Vision, pages 2375–2384, 2019. 2, 3
  27. 27.Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena. Structural-RNN: Deep learning on spatio-temporal graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5308–5317, 2016. 3
  28. 28.Magdiel Jimenez-Guarneros and Pilar Gomez-Gil. A study of the effects of negative transfer on deep unsupervised domain adaptation methods. Expert Systems with Applications, 167:114088, 2020. 3
  29. 29.Christopher Keat and Christian Laugier. Modelling smooth paths using gaussian processes. In Proceedings of the International Conference on Field and Service Robotics, pages 381–390, 2007. 3
  30. 30.Thomas N Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations, 2017. 5
  31. 31.Kris M. Kitani, Brian D. Ziebart, James Andrew Bagnell, and Martial Hebert. Activity forecasting. In Proceedings of the European Conference on Computer Vision, pages 201–214, 2012. 3
  32. 32.Vineet Kosaraju, Amir Sadeghian, Roberto Martın-Martın, Ian Reid, Hamid Rezatofighi, and Silvio Savarese. Socialbigat: Multimodal trajectory forecasting using bicycle-gan and graph attention networks. In Proceedings of the Advances in Neural Information Processing Systems, pages 137–146, 2019. 3
  33. 33.Parth Kothari, Brian Sifringer, and Alexandre Alahi. Interpretable social anchors for human trajectory forecasting in crowds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 15556–15566, 2021. 3
  34. 34.Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. Crowds by example. Computer Graphics Forum, 26(3):655–664, 2010. 6
  35. 35.Jiachen Li, Hengbo Ma, and Masayoshi Tomizuka. Conditional generative neural system for probabilistic trajectory prediction. arXiv preprint arXiv:1905.01631, 2019. 1, 3
  36. 36.Shijie Li, Yanying Zhou, Jinhui Yi, and Juergen Gall. Spatial-temporal consistency network for low-latency trajectory forecasting. In Proceedings of the IEEE International Conference on Computer Vision, pages 1940–1949, 2021. 2
  37. 37.Junwei Liang, Lu Jiang, and Alexander Hauptmann. Simaug: Learning robust representations from simulation for trajectory prediction. In Proceedings of the European Conference on Computer Vision, pages 275–292, 2020. 3
  38. 38.Junwei Liang, Lu Jiang, Juan Carlos Niebles, Alexander G Hauptmann, and Li Fei-Fei. Peeking into the future: Predicting future person activities and locations in videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5725–5734, 2019. 1, 3
  39. 39.Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. Learning transferable features with deep adaptation networks. In Proceedings of the International Conference on Machine Learning, pages 97–105, 2015. 2, 3, 5, 6, 7
  40. 40.Matthias Luber, Johannes A Stork, Gian Diego Tipaldi, and Kai O Arras. People tracking with human motion predictions from social forces. In Proceedings of the IEEE International Conference on Robotics and Automation, pages 464–469, 2010. 1
  41. 41.Zimeng Luo, Jiani Hu, Weihong Deng, and Haifeng Shen. Deep unsupervised domain adaptation for face recognition. In Proceedings of the IEEE International Conference on Automatic Face & Gesture Recognition, pages 453–457, 2018. 3
  42. 42.Qianqian Ma, Yang-Yu Liu, and Alex Olshevsky. Optimal lockdown for pandemic control. arXiv preprint arXiv:2010.12923, 2020. 1
  43. 43.Qianqian Ma and Alex Olshevsky. Adversarial crowdsourcing through robust rank-one matrix completion. In Proceedings of the Advances in Neural Information Processing Systems, pages 21841–21852, 2020. 3
  44. 44.Osama Makansi, Ozgun Cicek, Yassine Marrakchi, and Thomas Brox. On exposing the challenging long tail in future prediction of traffic actors. In Proceedings of the IEEE International Conference on Computer Vision, pages 13127–13137, 2021. 3
  45. 45.Dimitrios Makris and Tim Ellis. Spatial and probabilistic modelling of pedestrian behaviour. In Proceedings of the British Machine Vision Conference, pages 54.1–54.10, 2002. 3
  46. 46.Karttikeya Mangalam, Yang An, Harshayu Girase, and Jitendra Malik. From goals, waypoints & paths to long term human trajectory forecasting. In Proceedings of the IEEE International Conference on Computer Vision, pages 15233–15242, 2021. 3
  47. 47.Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan Hui Lee, Ehsan Adeli, Jitendra Malik, and Adrien Gaidon. It is not the journey but the destination: Endpoint conditioned trajectory prediction. In Proceedings of the European Conference on Computer Vision, pages 759–776, 2020. 1, 3, 6, 7
  48. 48.Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-STGCNN: A social spatiotemporal graph convolutional neural network for human trajectory prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 14424–14432, 2020. 2, 3, 5, 6, 7, 8
  49. 49.Mehdi Moussaıd, Niriaska Perozo, Simon Garnier, Dirk Helbing, and Guy Theraulaz. The walking behaviour of pedestrian social groups and its impact on crowd dynamics. PloS one, 5(4):e10047, 2010. 1
  50. 50.Basam Musleh, Fernando Garcıa, Javier Otamendi, Jose Ma Armingol, and De La Escalera Arturo. Identifying and tracking pedestrians based on sensor fusion and motion stability predictions. Sensors, 10(9):8028–8053, 2010. 1
  51. 51.Jie Ni, Qiang Qiu, and Rama Chellappa. Subspace interpolation via dictionary learning for unsupervised domain adaptation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 692–699, 2013. 2, 3, 5
  52. 52.Takenori Obo and Yuto Nakamura. Intelligent robot navigation based on human emotional model in human-aware environment. In Proceedings of the International Conference on Machine Learning and Cybernetics, pages 1–6, 2020. 1
  53. 53.Bo Pang, Tianyang Zhao, Xu Xie, and Ying Nian Wu. Trajectory prediction with latent belief energy-based model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 11814–11824, 2021. 3
  54. 54.Stefano Pellegrini, Andreas Ess, Konrad Schindler, and Luc J. Van Gool. You’ll never walk alone: Modeling social behavior for multi-target tracking. In Proceedings of the IEEE International Conference on Computer Vision, pages 261–268, 2009. 6
  55. 55.Christoph Rosmann, Malte Oeljeklaus, Frank Hoffmann, and Torsten Bertram. Online trajectory prediction and planning for social robot navigation. In Proceedings of the IEEE International Conference on Advanced Intelligent Mechatronics, pages 1255–1260, 2017. 3
  56. 56.Amir Sadeghian, Vineet Kosaraju, Ali Sadeghian, Noriaki Hirose, Hamid Rezatofighi, and Silvio Savarese. SoPhie: An attentive GAN for predicting paths compliant to social and physical constraints. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1349–1358, 2019. 1, 3
  57. 57.Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell. Adapting visual category models to new domains. In Proceedings of the European Conference on Computer Vision, pages 213–226, 2010. 3
  58. 58.Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data. In Proceedings of the European Conference on Computer Vision, pages 683–700, 2020. 1
  59. 59.Nasim Shafiee, Taskin Padir, and Ehsan Elhamifar. Introvert: Human trajectory prediction via conditional 3d attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 16815–16825, 2021. 1
  60. 60.Liushuai Shi, Le Wang, Chengjiang Long, Sanping Zhou, Mo Zhou, Zhenxing Niu, and Gang Hua. Sgcn: Sparse graph convolution network for pedestrian trajectory prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8994–9003, 2021. 2, 3, 5, 6, 7, 8
  61. 61.Baochen Sun and Kate Saenko. Deep CORAL: Correlation alignment for deep domain adaptation. In Proceedings of the European Conference on Computer Vision, pages 443–450. Springer, 2016. 2, 3, 5, 6
  62. 62.Jianhua Sun, Qinhong Jiang, and Cewu Lu. Recursive social behavior graph for trajectory prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 660–669, 2020. 2, 3, 6, 7
  63. 63.Jianhua Sun, Yuxuan Li, Hao-Shu Fang, and Cewu Lu. Three steps to multimodal trajectory prediction: Modality clustering, classification and synthesis. In Proceedings of the IEEE International Conference on Computer Vision, 2021. 3
  64. 64.Ben Talbot, Feras Dayoub, Peter Corke, and Gordon Wyeth. Robot navigation in unseen spaces using an abstract map. IEEE Transactions on Cognitive and Developmental Systems, 13(4):791–805, 2021. 1
  65. 65.Hung Tran, Vuong Le, and Truyen Tran. Goal-driven longterm trajectory prediction. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, pages 796–805, 2021. 3
  66. 66.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, pages 5998–6008, 2017. 3
  67. 67.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017. 3, 4
  68. 68.Anirudh Vemula, Katharina Muelling, and Jean Oh. Social attention: Modeling attention in human crowds. In Proceedings of the IEEE International Conference on Robotics and Automation, pages 1–7, 2018. 3
  69. 69.Chengxin Wang, Shaofeng Cai, and Gary Tan. Graphtcn: Spatio-temporal interaction modeling for human trajectory prediction. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, pages 3450–3459, 2021. 2
  70. 70.Lichen Wang, Bo Zong, Qianqian Ma, Wei Cheng, Jingchao Ni, Wenchao Yu, Yanchi Liu, Dongjin Song, Haifeng Chen, and Yun Fu. Inductive and unsupervised representation learning on graph structured objects. In Proceedings of the International Conference on Learning Representations, 2019. 3
  71. 71.Man Wu, Shirui Pan, Chuan Zhou, Xiaojun Chang, and Xingquan Zhu. Unsupervised domain adaptive graph convolutional networks. In Proceedings of the Web Conference, pages 1457–1467, 2020. 2, 3, 5, 6, 7
  72. 72.Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S Yu Philip. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems, 32(1):4–24, 2021. 3
  73. 73.Yi Xu, Dongchun Ren, Mingxia Li, Yuehai Chen, Mingyu Fan, and Huaxia Xia. Robust trajectory prediction of multiple interacting pedestrians via incremental active learning. In Proceedings of the International Conference on Neural Information Processing, pages 141–150, 2021. 3
  74. 74.Yi Xu, Dongchun Ren, Mingxia Li, Yuehai Chen, Mingyu Fan, and Huaxia Xia. Tra2tra: Trajectory-to-trajectory prediction with a global social spatial-temporal attentive neural network. IEEE Robotics and Automation Letters, 6(2):1574–1581, 2021. 1, 2, 4, 6, 7
  75. 75.Yi Xu, Jing Yang, and Shaoyi Du. CF-LSTM: Cascaded feature-based long short-term networks for predicting pedestrian trajectory. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 12541–12548, 2020. 1
  76. 76.Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial temporal graph convolutional networks for skeleton-based action recognition. arXiv preprint arXiv:1801.07455, 2018. 3
  77. 77.Baoyao Yang, Andy J. Ma, and Pong C. Yuen. Learning domain-shared group-sparse representation for unsupervised domain adaptation. Pattern Recognition, 81:615–632, 2018. 3
  78. 78.M. Yasuno, N. Yasuda, and M. Aoki. Pedestrian detection and tracking in far infrared images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 125–125, 2004. 1
  79. 79.Cunjun Yu, Xiao Ma, Jiawei Ren, Haiyu Zhao, and Shuai Yi. Spatio-temporal graph transformer networks for pedestrian trajectory prediction. arXiv preprint arXiv:2005.08514, 2020. 3
  80. 80.Ye Yuan, Xinshuo Weng, Yanglan Ou, and Kris Kitani. Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting. In Proceedings of the IEEE International Conference on Computer Vision, pages 9793–9803, 2021. 3
  81. 81.Luo Yuanfu, Cai Panpan, Bera Aniket, Hsu David, Lee Wee Sun, and Manocha Dinesh. PORCA: Modeling and planning for autonomous driving among many pedestrians. IEEE Robotics and Automation Letters, 3:3418–3425, 2018. 1
  82. 82.Pu Zhang, Wanli Ouyang, Pengfei Zhang, Jianru Xue, and Nanning Zheng. SR-LSTM: State refinement for LSTM towards pedestrian trajectory prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12085–12094, 2019. 1, 3
  83. 83.He Zhao and Richard P Wildes. Where are you heading? Dynamic trajectory prediction with expert goal examples. In Proceedings of the IEEE International Conference on Computer Vision, pages 7629–7638, 2021. 3
  84. 84.Fang Zheng, Le Wang, Sanping Zhou, Wei Tang, Zhenxing Niu, Nanning Zheng, and Gang Hua. Unlimited neighborhood interaction for heterogeneous trajectory prediction. In Proceedings of the IEEE International Conference on Computer Vision, pages 13168–13177, 2021. 3
  85. 85.Yanliang Zhu, Deheng Qian, Dongchun Ren, and Huaxia Xia. Starnet: Pedestrian trajectory prediction using deep neural network in star topology. In Proceedings of the IEEE International Conference on Intelligent Robots and Systems, pages 8075–8080, 2019. 3, 4
  86. 86.Yanliang Zhu, Dongchun Ren, Yi Xu, Deheng Qian, Mingyu Fan, Xin Li, and Huaxia Xia. Simultaneous past and current social interaction-aware trajectory prediction for multiple intelligent agents in dynamic scenes. ACM Transactions on Intelligent Systems and Technology, 13:1–16, 2021. 1
  87. 87.Junbao Zhuo, Shuhui Wang, and Weigang Zhang. Deep unsupervised convolutional domain adaptation. In Proceedings of the ACM International Conference on Multimedia, pages 261–269, 2017. 2, 3, 5, 7

Citation

MLA
Xu, Y., et al. “Adaptive Trajectory Prediction via Transferable GNN”. arXiv, 2022, http://arxiv.org/abs/2203.05046v2.
APA
Xu, Y., Wang, L., Wang, Y., & Fu, Y. (2022). Adaptive Trajectory Prediction via Transferable GNN. arXiv. http://arxiv.org/abs/2203.05046v2
Chicago
Xu, Y., L. Wang, Y. Wang, and Y. Fu. 2022. “Adaptive Trajectory Prediction via Transferable GNN”. arXiv. http://arxiv.org/abs/2203.05046v2.
Harvard
Xu, Y. et al. (2022) “Adaptive Trajectory Prediction via Transferable GNN”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.05046v2.
Vancouver
1. Xu Y, Wang L, Wang Y, Fu Y (2022) Adaptive Trajectory Prediction via Transferable GNN. arXiv

BibTeX

@article{xu2022adaptive,
  title = {Adaptive Trajectory Prediction via Transferable GNN},
  author = {Xu, Yi and Wang, Lichen and Wang, Yizhou and Fu, Yun},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.05046v2},
  eprint = {2203.05046}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE