Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data

Tim SalzmannBoris IvanovicPunarjay ChakravartyMarco Pavone

article2020ECCV1,404 citations

Introduces Trajectron++, a graph-structured trajectory forecasting framework that incorporates kinematic constraints, semantic map data, and planned ego-agent motions to generate physically feasible multi-agent predictions for robotic control systems.

Listen

Autonomous systems operating in human-populated environments—such as self-driving vehicles—must accurately anticipate the future movements of pedestrians, vehicles, and other dynamic agents to ensure safe and socially compliant navigation. However, existing trajectory forecasting methods frequently produce unrealistic predictions because they ignore physical dynamic constraints (such as a vehicle’s inability to slide sideways) and fail to incorporate rich environmental context (such as high-definition maps) or the autonomous system's own planned actions.

The article develops and evaluates Trajectron++, a graph-structured, deep generative model designed to forecast multi-agent trajectories while explicitly enforcing dynamic feasibility, ingesting heterogeneous environmental data, and optionally conditioning predictions on the ego-agent's planned path.

The authors evaluated the framework across standard real-world benchmark datasets, including the ETH and UCY pedestrian datasets (encompassing five data subsets and 1,536 unique pedestrians) and the large-scale nuScenes autonomous driving dataset (featuring multi-class agents and 11-layer high-definition semantic maps). The model represents dynamic scenes as directed spatiotemporal graphs, predicts probability distributions over control actions rather than raw coordinates, and integrates physical vehicle and pedestrian motion dynamics to output physically feasible position trajectories.

The key findings demonstrate that Trajectron++ substantially outperforms current state-of-the-art predictive methods. First, the model achieves 55% to 60% lower average displacement errors compared to leading generative baselines across pedestrian benchmarks and improves final displacement error by 33% over deterministic regressors. Second, integrating system dynamics serves as the single most critical driver of performance, directly eliminating physically impossible trajectories while substantially reducing negative log-likelihood across all datasets. Third, incorporating semantic map data reduces collision and obstacle-violation rates among pedestrians from 4.6% to 1.0% overall, and from 21.5% down to 4.9% for agents situated close to obstacles. Finally, conditioning predictions on the ego-vehicle's planned future motions yields marked reductions in error and decreases road-boundary violations from 7.6% to 4.2% on the nuScenes dataset, all while executing well within real-time robotics constraints (under 1.2 seconds per scene).

These findings indicate that integrating physical dynamic models and semantic map contexts directly into learning-based forecasting pipelines dramatically reduces safety risks, improves probability calibration, and minimizes false predictions in complex traffic scenarios. Crucially, the capability to evaluate human and agent responses conditional on the robot's future actions enables downstream motion planners to test multiple candidate trajectories safely before execution.

Decision-makers and engineering teams developing autonomous navigation stacks should consider adopting control-space prediction pipelines that explicitly integrate kinematic models and semantic map layers rather than relying on unconstrained coordinate regressors. Integration efforts should prioritize linking the forecasting framework directly with ego-vehicle path planning to leverage interaction-aware conditioning.

The reported results carry high confidence across standard public benchmarks; however, limitations remain regarding the use of simplified unicycle dynamics rather than full physical vehicle models (such as complete bicycle models with online parameter estimation). Readers should also exercise caution when evaluating downstream real-world deployments, as the study assumed clean object detections and ground-truth tracking inputs rather than accounting for sensor noise and detection latency in end-to-end perception stacks.

arXiv: 2001.03093
  • Paper: Planning-oriented Autonomous Driving, Yi Hu et al. (2022). UniAD extends modular trajectory forecasting and dynamic agent interaction concepts into a fully unified, end-to-end perception, prediction, and planning autonomous driving pipeline.
  • Paper: Deep Reinforcement Learning for Autonomous Driving: A Survey, Bangalore Ravi Kiran et al. (2020). This survey explores how dynamic multi-agent trajectory forecasting models are integrated into downstream reinforcement learning and motion planning frameworks for autonomous vehicles.
Cover for Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data

Abstract

Reasoning about human motion is an important prerequisite to safe and socially-aware robotic navigation. As a result, multi-agent behavior prediction has become a core component of modern human-robot interactive systems, such as self-driving cars. While there exist many methods for trajectory forecasting, most do not enforce dynamic constraints and do not account for environmental information (e.g., maps). Towards this end, we present Trajectron++, a modular, graph-structured recurrent model that forecasts the trajectories of a general number of diverse agents while incorporating agent dynamics and heterogeneous data (e.g., semantic maps). Trajectron++ is designed to be tightly integrated with robotic planning and control frameworks; for example, it can produce predictions that are optionally conditioned on ego-agent motion plans. We demonstrate its performance on several challenging real-world trajectory forecasting datasets, outperforming a wide array of state-of-the-art deterministic and generative methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Formulation
  • 4 Trajectron++
  • 5 Experiments
  • 5.1 ETH and UCY Datasets
  • 5.2 nuScenes Dataset
  • 6 Conclusion
  • References
  • 0.A Single Integrator Distribution Integration
  • 0.A.1 Mean Derivation22 2 These equations are also found in the Kalman Filter prediction step [24].
  • 0.A.2 Covariance Derivation22footnotemark: 2
  • 0.B Dynamically-Extended Unicycle Distribution Integration
  • 0.B.1 Mean Derivation44 4 These equations are also found in the Extended Kalman Filter prediction step [46].
  • 0.B.2 Covariance Derivation33footnotemark: 3
  • 0.C Average and Final Displacement Error Evaluation
  • 0.D Additional Training Information
  • 0.D.1 Choosing α,β\alpha,\beta in
  • 0.D.2 Separate Map Encoder Learning Rate
  • 0.D.3 Data Augmentation
  • 0.E ETH Pedestrians Map Encoding
  • 0.F Online Runtime

Knowls

  1. Knowl 1 — Trajectron++ Spatiotemporal Graph and Heterogeneous Encoder Architecture

    model/method

    Trajectron++ represents a dynamic scene with a time-varying number N(t)N(t) of interacting agents A1,…,AN(t)A_1, \dots, A_{N(t)} as a directed spatiotemporal graph G=(V,E)G = (V, E). Each node corresponds to an agent with a semantic class SiS_i (e.g., Pedestrian, Car, Bus). A directed edge (Ai,Aj)∈E(A_i, A_j) \in E exists if agent AiA_i influences agent AjA_j, determined by Euclidean distance: ∥pi−pj∥2≤dSj\|p_i - p_j\|_2 \le d_{S_j}, where pi,pj∈R2p_i, p_j \in \mathbb{R}^2 are 2D world positions and dSjd_{S_j} represents the perception range of class SjS_j.

    The encoder generates a representation vector exe_x for each modeled agent by fusing several heterogeneous input representations:

    1. Agent History: The current and past HH states xi=si(t−H:t)∈R(H+1)×Dx_i = s_i^{(t-H:t)} \in \mathbb{R}^{(H+1) \times D} are encoded using a Long Short-Term Memory (LSTM) network with 32 hidden dimensions.
    2. Agent Interactions: For a target agent, incoming edges from neighboring agents of the same semantic class are aggregated via an element-wise sum and passed to a class-specific edge LSTM (8 hidden dimensions). The outputs of all edge LSTMs connecting to the node are combined using an additive attention mechanism to yield an edge influence vector.
    3. Heterogeneous Environmental Data (Map): A local geometric semantic map Mi(t)∈RHM×WM×LM_i^{(t)} \in \mathbb{R}^{H_M \times W_M \times L} around agent AiA_i, rotated to align with the agent's current heading, is processed by a 4-layer Convolutional Neural Network (CNN) with filter sizes {5,5,5,3}\{5, 5, 5, 3\}, strides {2,2,1,1}\{2, 2, 1, 1\}, and leaky ReLU activations (alpha=0.2\\alpha = 0.2), followed by a 32-dimensional dense layer.
    4. Ego-Agent Future Motion Plans: When evaluating potential robot trajectories, the ego-agent's planned trajectory over the future TT timesteps yR=sR(t+1:t+T)y_R = s_R^{(t+1:t+T)} is encoded via a 32-dimensional bi-directional LSTM.

    The node history representation, edge influence representation, map feature vector, and ego future representation are concatenated to produce the unified conditioning representation vector exe_x.

  2. Knowl 2 — Dynamically-Feasible Trajectory Generation via Control Integration

    model/method

    Rather than directly predicting agent positions, Trajectron++ predicts distributions over low-level control actions u(t)u^{(t)} and integrates them forward through explicit differential equations governing agent dynamics.

    A Gated Recurrent Unit (GRU) decoder with 128 hidden dimensions receives the concatenated encoder representation exe_x, a discrete latent mode zz, and the previous state/action to output the parameters of a bivariate Gaussian distribution over control actions N(μu(t),Σu(t))\mathcal{N}(\mu_u^{(t)}, \Sigma_u^{(t)}) at each time step tt.

    For agents modeled with linear dynamics (e.g., pedestrians modeled as single integrators with control u(t)=p˙(t)u^{(t)} = \dot{p}^{(t)}), system dynamics are linear Gaussian, enabling exact closed-form propagation of the position mean and covariance.

    For agents governed by nonlinear kinematics (e.g., wheeled vehicles modeled as dynamically-extended unicycles with control u(t)=[ω(t),a(t)]Tu^{(t)} = [\omega^{(t)}, a^{(t)}]^T), the system dynamics are linearized about the current state mean and control mean at each step to propagate the mean and covariance in an Extended Kalman Filter (EKF) style.

    Supervision during training is applied directly to the integrated position distribution pψ(y∣x,z)p_\psi(y \mid x, z), allowing the network to optimize positional forecasting error while strictly guaranteeing dynamic feasibility of the generated trajectories.

  3. Knowl 3 — Conditional Latent Variable Formulation and Discrete InfoVAE Objective

    model/method

    Trajectron++ captures multimodal future behaviors using a Conditional Variational Autoencoder (CVAE) framework with a discrete categorical latent variable z∈Zz \in Z, where ∣Z∣=25|Z| = 25. The future trajectory distribution is expressed as the mixture:

    p(y∣x)=∑z∈Zpψ(y∣x,z)pθ(z∣x)p(y \mid x) = \sum_{z \in Z} p_\psi(y \mid x, z) p_\theta(z \mid x)

    where pθ(z∣x)p_\theta(z \mid x) is the categorical prior parameterized by the encoder representation exe_x, and pψ(y∣x,z)p_\psi(y \mid x, z) is the decoder output distribution after dynamical integration.

    During training, ground truth future trajectories yy are encoded by a 32-dimensional bi-directional LSTM to produce the variational posterior qϕ(z∣x,y)q_\phi(z \mid x, y). The network parameters (ϕ,θ,ψ)(\phi, \theta, \psi) are optimized using the InfoVAE objective adapted for discrete conditional variables:

    max⁡ϕ,θ,ψ∑i=1NEz∼qϕ(⋅∣xi,yi)[log⁡pψ(yi∣xi,z)]−βDKL(qϕ(z∣xi,yi)∥pθ(z∣xi))+αIq(x;z)\max_{\phi, \theta, \psi} \sum_{i=1}^N \mathbb{E}_{z \sim q_\phi(\cdot \mid x_i, y_i)} [\log p_\psi(y_i \mid x_i, z)] - \beta D_{\mathrm{KL}}(q_\phi(z \mid x_i, y_i) \parallel p_\theta(z \mid x_i)) + \alpha I_q(x; z)

    where Iq(x;z)I_q(x; z) is the mutual information between xx and zz under qϕ(x,z)q_\phi(x, z), computed by approximating qϕ(z∣xi,yi)q_\phi(z \mid x_i, y_i) with pθ(z∣xi)p_\theta(z \mid x_i) and marginalizing over the training batch with weighting α=1.0\alpha = 1.0. Because ∣Z∣=25|Z| = 25 is finite and discrete, the expectation Ez∼qϕ\mathbb{E}_{z \sim q_\phi} is computed analytically by summing over all 25 elements, avoiding gradient estimators such as the Gumbel-Softmax reparameterization. The hyperparameter β\beta is annealed during training using an increasing sigmoid schedule.

  4. Knowl 4 — Dynamically-Extended Unicycle Dynamics and Uncertainty Propagation

    equation

    For wheeled vehicles, Trajectron++ models motion via the dynamically-extended unicycle model with state s=[x,y,ϕ,v]Ts = [x, y, \phi, v]^T (representing 2D position, heading angle, and speed) and control inputs u=[ω,a]Tu = [\omega, a]^T (representing heading rate and acceleration). The continuous dynamics are given by:

    [x˙y˙ϕ˙v˙]=[vcos⁡(ϕ)vsin⁡(ϕ)ωa]\begin{bmatrix} \dot{x} \\ \dot{y} \\ \dot{\phi} \\ \dot{v} \end{bmatrix} = \begin{bmatrix} v \cos(\phi) \\ v \sin(\phi) \\ \omega \\ a \end{bmatrix}

    Assuming a zero-order hold over sampling interval Δt\Delta t, the discrete-time state update s(t+1)=f(s(t),u(t))s^{(t+1)} = f(s^{(t)}, u^{(t)}) for ∣ω∣>10−3|\omega| > 10^{-3} is:

    [x(t+1)y(t+1)ϕ(t+1)v(t+1)]=[x(t)y(t)ϕ(t)v(t)]+[v(t)DS(t)+a(t)ω(t)sin⁡(ϕ(t)+ω(t)Δt)Δt+a(t)ω(t)DC(t)−v(t)DC(t)−a(t)ω(t)cos⁡(ϕ(t)+ω(t)Δt)Δt+a(t)ω(t)DS(t)ω(t)Δta(t)Δt]\begin{bmatrix} x^{(t+1)} \\ y^{(t+1)} \\ \phi^{(t+1)} \\ v^{(t+1)} \end{bmatrix} = \begin{bmatrix} x^{(t)} \\ y^{(t)} \\ \phi^{(t)} \\ v^{(t)} \end{bmatrix} + \begin{bmatrix} v^{(t)} D_S^{(t)} + \frac{a^{(t)}}{\omega^{(t)}} \sin(\phi^{(t)} + \omega^{(t)}\Delta t)\Delta t + \frac{a^{(t)}}{\omega^{(t)}} D_C^{(t)} \\ -v^{(t)} D_C^{(t)} - \frac{a^{(t)}}{\omega^{(t)}} \cos(\phi^{(t)} + \omega^{(t)}\Delta t)\Delta t + \frac{a^{(t)}}{\omega^{(t)}} D_S^{(t)} \\ \omega^{(t)} \Delta t \\ a^{(t)} \Delta t \end{bmatrix}

    where DS(t)=sin⁡(ϕ(t)+ω(t)Δt)−sin⁡(ϕ(t))ω(t)D_S^{(t)} = \frac{\sin(\phi^{(t)} + \omega^{(t)}\Delta t) - \sin(\phi^{(t)})}{\omega^{(t)}} and DC(t)=cos⁡(ϕ(t)+ω(t)Δt)−cos⁡(ϕ(t))ω(t)D_C^{(t)} = \frac{\cos(\phi^{(t)} + \omega^{(t)}\Delta t) - \cos(\phi^{(t)})}{\omega^{(t)}}. For ∣ω∣≤10−3|\omega| \le 10^{-3}, the limit dynamics as ω→0\omega \to 0 are used:

    [x(t+1)y(t+1)ϕ(t+1)v(t+1)]=[x(t)y(t)ϕ(t)v(t)]+[v(t)cos⁡(ϕ(t))Δt+0.5a(t)cos⁡(ϕ(t))(Δt)2v(t)sin⁡(ϕ(t))Δt+0.5a(t)sin⁡(ϕ(t))(Δt)20a(t)Δt]\begin{bmatrix} x^{(t+1)} \\ y^{(t+1)} \\ \phi^{(t+1)} \\ v^{(t+1)} \end{bmatrix} = \begin{bmatrix} x^{(t)} \\ y^{(t)} \\ \phi^{(t)} \\ v^{(t)} \end{bmatrix} + \begin{bmatrix} v^{(t)} \cos(\phi^{(t)})\Delta t + 0.5 a^{(t)} \cos(\phi^{(t)})(\Delta t)^2 \\ v^{(t)} \sin(\phi^{(t)})\Delta t + 0.5 a^{(t)} \sin(\phi^{(t)})(\Delta t)^2 \\ 0 \\ a^{(t)} \Delta t \end{bmatrix}

    Given predicted control distribution u(t)∼N(μu(t),Σu(t))u^{(t)} \sim \mathcal{N}(\mu_u^{(t)}, \Sigma_u^{(t)}) with Σu=[σω2ρωaσωσaρωaσωσaσa2]\Sigma_u = \begin{bmatrix} \sigma_\omega^2 & \rho_{\omega a}\sigma_\omega\sigma_a \\ \rho_{\omega a}\sigma_\omega\sigma_a & \sigma_a^2 \end{bmatrix}, the state covariance Σp,ϕ,v\Sigma_{p,\phi,v} is propagated via first-order Taylor expansion:

    Σp,ϕ,v(t+1)=F(t)Σp,ϕ,v(t)(F(t))T+G(t)Σu(t)(G(t))T\Sigma_{p,\phi,v}^{(t+1)} = F^{(t)} \Sigma_{p,\phi,v}^{(t)} (F^{(t)})^T + G^{(t)} \Sigma_u^{(t)} (G^{(t)})^T

    where F(t)=∂f∂s∣(μs(t),μu(t))F^{(t)} = \left.\frac{\partial f}{\partial s}\right|_{(\mu_s^{(t)}, \mu_u^{(t)})} and G(t)=∂f∂u∣(μs(t),μu(t))G^{(t)} = \left.\frac{\partial f}{\partial u}\right|_{(\mu_s^{(t)}, \mu_u^{(t)})}.

  5. Knowl 5 — Single Integrator Dynamics and Linear Gaussian Uncertainty Propagation

    equation

    For pedestrian agents, the state is defined as position s=p=[x,y]T∈R2s = p = [x, y]^T \in \mathbb{R}^2 and the control action is velocity u=p˙=[x˙,y˙]T∈R2u = \dot{p} = [\dot{x}, \dot{y}]^T \in \mathbb{R}^2. The discrete linear system dynamics over time step Δt\Delta t are:

    p(t+1)=I2×2p(t)+ΔtI2×2p˙(t)p^{(t+1)} = I_{2 \times 2} p^{(t)} + \Delta t I_{2 \times 2} \dot{p}^{(t)}

    The decoder outputs a bivariate Gaussian distribution N(μu(t),Σu(t))\mathcal{N}(\mu_u^{(t)}, \Sigma_u^{(t)}) over velocity:

    μu=[μx˙μy˙],Σu=[σx˙2ρx˙y˙σx˙σy˙ρx˙y˙σx˙σy˙σy˙2]\mu_u = \begin{bmatrix} \mu_{\dot{x}} \\ \mu_{\dot{y}} \end{bmatrix}, \quad \Sigma_u = \begin{bmatrix} \sigma_{\dot{x}}^2 & \rho_{\dot{x}\dot{y}}\sigma_{\dot{x}}\sigma_{\dot{y}} \\ \rho_{\dot{x}\dot{y}}\sigma_{\dot{x}}\sigma_{\dot{y}} & \sigma_{\dot{y}}^2 \end{bmatrix}

    Because the dynamics are linear and Σu(t)\Sigma_u^{(t)} is the sole source of uncertainty, position mean μp\mu_p and covariance Σp\Sigma_p are propagated analytically according to linear Gaussian sum rules:

    μp(t+1)=μp(t)+Δtμu(t)\mu_p^{(t+1)} = \mu_p^{(t)} + \Delta t \mu_u^{(t)}

    Σp(t+1)=Σp(t)+(Δt)2Σu(t)\Sigma_p^{(t+1)} = \Sigma_p^{(t)} + (\Delta t)^2 \Sigma_u^{(t)}

  6. Knowl 6 — Trajectron++ Output Configurations

    definition

    Depending on downstream robotic task requirements, Trajectron++ supports four inference configurations:

    1. Most Likely (ML): Deterministic mode prediction where both the latent behavior mode and trajectory are the distribution modes: zmode=arg⁡max⁡zpθ(z∣x),y=arg⁡max⁡ypψ(y∣x,zmode)z_{\mathrm{mode}} = \arg\max_z p_\theta(z \mid x), \quad y = \arg\max_y p_\psi(y \mid x, z_{\mathrm{mode}})
    2. zmodez_{\mathrm{mode}} Mode Sampling: Trajectories sampled exclusively from the most likely high-level behavior mode: zmode=arg⁡max⁡zpθ(z∣x),y∼pψ(y∣x,zmode)z_{\mathrm{mode}} = \arg\max_z p_\theta(z \mid x), \quad y \sim p_\psi(y \mid x, z_{\mathrm{mode}})
    3. Full Sampling: Unconditional sampling where latent discrete mode zz and trajectory yy are sampled sequentially: z∼pθ(z∣x),y∼pψ(y∣x,z)z \sim p_\theta(z \mid x), \quad y \sim p_\psi(y \mid x, z)
    4. Analytic Distribution: Closed-form mixture evaluation over the full discrete latent space: p(y∣x)=∑z∈Zpψ(y∣x,z)pθ(z∣x)p(y \mid x) = \sum_{z \in Z} p_\psi(y \mid x, z) p_\theta(z \mid x)
  7. Knowl 7 — Pedestrian Trajectory Forecasting Performance on ETH and UCY Benchmarks

    data/table

    Trajectron++ was evaluated on the ETH and UCY pedestrian benchmark datasets (comprising ETH, Hotel, Univ, Zara 1, and Zara 2) using a leave-one-out cross-validation protocol (8 timesteps / 3.2s observed, 12 timesteps / 4.8s predicted).

    Dataset Deterministic ADE / FDE (m)
    Linear LSTM S-LSTM S-ATTN Ours (ML) Ours+∫\int (ML)
    ETH 1.33 / 2.94 1.09 / 2.41 1.09 / 2.35 0.39 / 3.74 0.71 / 1.66 0.71 / 1.68
    Hotel 0.39 / 0.72 0.86 / 1.91 0.79 / 1.76 0.29 / 2.64 0.22 / 0.46 0.22 / 0.46
    Univ 0.82 / 1.59 0.61 / 1.31 0.67 / 1.40 0.33 / 3.92 0.44 / 1.17 0.41 / 1.07
    Zara 1 0.62 / 1.21 0.41 / 0.88 0.47 / 1.00 0.20 / 0.52 0.30 / 0.79 0.30 / 0.77
    Zara 2 0.77 / 1.48 0.52 / 1.11 0.56 / 1.17 0.30 / 2.13 0.23 / 0.59 0.23 / 0.59
    Average 0.79 / 1.59 0.70 / 1.52 0.72 / 1.54 0.30 / 2.59 0.38 / 0.93 0.37 / 0.91
    Dataset Probabilistic Best-of-20 ADE / FDE (m)
    S-GAN SoPhie Trajectron MATF Ours (Full) Ours+∫\int (Full)
    ETH 0.81 / 1.52 0.70 / 1.43 0.59 / 1.14 1.01 / 1.75 0.39 / 0.83 0.43 / 0.86
    Hotel 0.72 / 1.61 0.76 / 1.67 0.35 / 0.66 0.43 / 0.80 0.12 / 0.21 0.12 / 0.19
    Univ 0.60 / 1.26 0.54 / 1.24 0.54 / 1.13 0.44 / 0.91 0.20 / 0.44 0.22 / 0.43
    Zara 1 0.34 / 0.69 0.30 / 0.63 0.43 / 0.83 0.26 / 0.45 0.15 / 0.33 0.17 / 0.32
    Zara 2 0.42 / 0.84 0.38 / 0.78 0.43 / 0.85 0.26 / 0.57 0.11 / 0.25 0.12 / 0.25
    Average 0.58 / 1.18 0.54 / 1.15 0.47 / 0.92 0.48 / 0.90 0.19 / 0.41 0.21 / 0.41
    Dataset Kernel Density Estimate NLL (2000 samples)
    S-GAN Trajectron Ours (Full) Ours+∫\int (Full)
    ETH 15.70 2.99 1.80 1.31
    Hotel 8.10 2.26 -1.29 -1.94
    Univ 2.88 1.05 -0.89 -1.13
    Zara 1 1.36 1.86 -1.13 -1.41
    Zara 2 0.96 0.81 -2.19 -2.53
    Average 5.80 1.79 -0.74 -1.14

    These results demonstrate that Trajectron++'s deterministic ML mode reduces average FDE by 33% relative to deterministic baselines. Under probabilistic evaluation (Best-of-20), Trajectron++ achieves 55% to 60% lower average error than previous generative models (MATF and Trajectron). Under KDE Negative Log Likelihood, the dynamics-integrated model (Ours+$\int$) achieves the best log-likelihood across all datasets (-1.14 mean NLL).

  8. Knowl 8 — Trajectory Forecasting Performance and Horizon Generalization on nuScenes

    data/table

    Trajectory prediction performance evaluated on the nuScenes dataset for vehicles and pedestrians across forecasting horizons of 1s, 2s, 3s, and 4s (trained only on a 3s horizon). Reported baseline vehicle metrics have detection/tracking error (22–24 cm) subtracted for fair comparison.

    Method Vehicle FDE (m)
    @1s @2s @3s @4s
    Constant Velocity 0.32 0.89 1.70 2.73
    Social LSTM 0.47 - 1.61 -
    Convolutional Social Pooling 0.46 - 1.50 -
    CAR-Net 0.38 - 1.35 -
    SpAGNN 0.36 - 1.23 -
    Ours (ML) 0.18 0.57 1.25 2.24
    Ours+∫\int, M (ML) 0.07 0.45 1.14 2.20
    Pedestrian Model Pedestrian KDE NLL / FDE (m)
    @1s @2s @3s @4s
    Ours (ML) -2.69 / 0.03 -2.46 / 0.17 -1.76 / 0.37 -1.09 / 0.60
    Ours+∫\int, M (ML) -5.58 / 0.01 -3.96 / 0.17 -2.77 / 0.37 -1.89 / 0.62

    Trajectron++ (Ours+$\int$, M) outperforms all vehicle baselines at all horizons and maintains low error at 4s despite being trained only up to 3s, demonstrating temporal generalizability beyond its training horizon.

  9. Knowl 9 — Ablation of Dynamics, Semantic Maps, and Ego-Vehicle Future Conditioning

    data/table

    Ablation study on the nuScenes dataset evaluating the impact of dynamics integration (∫\int), HD semantic map encoding (MM), and ego-robot future motion plan conditioning (yRy_R) on vehicles.

    Configuration KDE NLL FDE ML (m) Boundary Viol. (%)
    ∫\int MM yRy_R @1s @2s @3s @4s @1s @2s @3s @4s @1s @2s @3s @4s
    (a) Including Ego-Vehicle in Evaluation
    - - - 0.81 0.05 0.37 0.87 0.18 0.57 1.25 2.24 0.2 0.6 2.8 6.9
    ✓ - - -4.28 -2.82 -1.67 -0.76 0.07 0.45 1.13 2.17 0.2 0.7 3.2 8.1
    ✓ ✓ - -4.17 -2.74 -1.62 -0.70 0.07 0.45 1.14 2.20 0.3 0.6 2.8 7.6
    (b) Excluding Ego-Vehicle (Conditioning on Ego-Vehicle Future)
    ✓ ✓ - -4.26 -2.86 -1.76 -0.87 0.07 0.44 1.09 2.09 0.3 0.6 2.8 7.6
    ✓ ✓ ✓ -3.90 -2.76 -1.75 -0.93 0.08 0.34 0.81 1.50 0.3 0.5 1.6 4.2

    The ablation indicates:

    1. Dynamics integration (∫\int) is the single largest contributor to performance, drastically reducing KDE NLL (from +0.81 to -4.28 @ 1s) and FDE (from 0.18 m to 0.07 m @ 1s).
    2. Semantic map encoding (MM) curbs the slight increase in road boundary violations caused by training in position space (reducing violations from 8.1% to 7.6% @ 4s).
    3. Ego-robot future conditioning (yRy_R) provides substantial gains, lowering 4s FDE from 2.09 m to 1.50 m and reducing 4s road boundary violations from 7.6% to 4.2%.
  10. Knowl 10 — Impact of Semantic Map Encoding on Obstacle and Boundary Violations

    empirical result

    Incorporating environmental map representations through the CNN map encoder significantly reduces the incidence of physically infeasible and colliding trajectory predictions across datasets:

    1. ETH University Dataset (Binary Obstacle Maps): Across 2000 sampled trajectories using the Full output mode, Trajectron++ without map encoding generates obstacle-colliding predictions 4.6% of the time, which drops to 1.0% with map encoding. In high-risk situations where pedestrians are close to obstacles (defined as agents with at least one colliding trajectory sample), the collision rate drops from 21.5% without map encoding to 4.9% with map encoding.
    2. nuScenes Dataset (HD Semantic Maps): For vehicle trajectory predictions at a 4-second horizon, introducing semantic HD map encoding reduces the road boundary violation rate from 8.1% to 7.6%, and further down to 4.2% when combined with conditioning on the ego-vehicle's future motion plan.

Coverage note — None was omitted; all key contributions—including the spatiotemporal graph architecture, discrete CVAE and InfoVAE objective, dynamical integration models (single integrator and dynamically-extended unicycle), covariance linearization, output modes, comprehensive empirical benchmarks across ETH/UCY and nuScenes, map-based collision reduction, and ablation studies—are covered.

References

  1. 1.Alahi, A., Goel, K., Ramanathan, V., Robicquet, A., Fei-Fei, L., Savarese, S.: Social LSTM: Human trajectory prediction in crowded spaces. In: IEEE Conf. on Computer Vision and Pattern Recognition (2016) 3, 5, 9, 10, 11, 12
  2. 2.Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: Int. Conf. on Learning Representations (2015) 6
  3. 3.Battaglia, P.W., Pascanu, R., Lai, M., Rezende, D., Kavukcuoglu, K.: Interaction networks for learning about objects, relations and physics. In: Conf. on Neural Information Processing Systems (2016) 6
  4. 4.Bowman, S.R., Vilnis, L., Vinyals, O., Dai, A.M., Jozefowicz, R., Bengio, S.: Generating sentences from a continuous space. In: Proc. Annual Meeting of the Association for Computational Linguistics (2015) 22
  5. 5.Britz, D., Goldie, A., Luong, M.T., Le, Q.V.: Massive exploration of neural machine translation architectures. In: Proc. of Conf. on Empirical Methods in Natural Language Processing. pp. 1442–1451 (2017) 7
  6. 6.Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuScenes: A multimodal dataset for autonomous driving (2019) 4, 8, 12
  7. 7.Casas, S., Gulino, C., Liao, R., Urtasun, R.: SpAGNN: Spatially-aware graph neural networks for relational behavior forecasting from sensor data (2019) 3, 10, 12, 13
  8. 8.Casas, S., Luo, W., Urtasun, R.: IntentNet: Learning to predict intention from raw sensor data. In: Conf. on Robot Learning. pp. 947–956 (2018) 3
  9. 9.Chang, M.F., Lambert, J., Sangkloy, P., Singh, J., Bak, S., Hartnett, A., Wang, D., Carr, P., Lucey, S., Ramanan, D., Hays, J.: Argoverse: 3d tracking and forecasting with rich maps. In: IEEE Conf. on Computer Vision and Pattern Recognition (2019) 4
  10. 10.Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y.: Learning phrase representations using rnn encoder-decoder for statistical machine translation. In: Proc. of Conf. on Empirical Methods in Natural Language Processing. pp. 1724–1734 (2014) 7
  11. 11.Deo, N., Trivedi, M.M.: Multi-modal trajectory prediction of surrounding vehicles with maneuver based lstms. In: IEEE Intelligent Vehicles Symposium (2018) 4, 9, 12
  12. 12.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: Conf. on Neural Information Processing Systems (2014) 3, 4
  13. 13.Gupta, A., Johnson, J., Li, F., Savarese, S., Alahi, A.: Social GAN: Socially acceptable trajectories with generative adversarial networks. In: IEEE Conf. on Computer Vision and Pattern Recognition (2018) 4, 5, 9, 10, 11
  14. 14.Gweon, H., Saxe, R.: Developmental cognitive neuroscience of theory of mind. In: Neural Circuit Development and Function in the Brain, chap. 20, pp. 367–377. Academic Press (2013). https://doi.org/https://doi.org/10.1016/B978-0-12-397267-5.00057-1, http://www.sciencedirect.com/science/article/pii/B9780123972675000571 1
  15. 15.Hallac, D., Leskovec, J., Boyd, S.: Network lasso: Clustering and optimization in large graphs. In: ACM Int. Conf. on Knowledge Discovery and Data Mining (2015) 23
  16. 16.Helbing, D., Molnar, P.: Social force model for pedestrian dynamics. Physical Review E 51(5), 4282–4286 (1995) 3
  17. 17.Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., Lerchner, A.: beta-VAE: Learning basic visual concepts with a constrained variational framework. In: Int. Conf. on Learning Representations (2017) 21
  18. 18.Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural Computation (1997) 6
  19. 19.Ivanovic, B., Pavone, M.: The Trajectron: Probabilistic multi-agent trajectory modeling with dynamic spatiotemporal graphs. In: IEEE Int. Conf. on Computer Vision (2019) 2, 3, 4, 5, 9, 10, 11, 21
  20. 20.Ivanovic, B., Schmerling, E., Leung, K., Pavone, M.: Generative modeling of multimodal multi-human behavior. In: IEEE/RSJ Int. Conf. on Intelligent Robots & Systems (2018) 4, 5, 6
  21. 21.Jain, A., Zamir, A.R., Savarese, S., Saxena, A.: Structural-RNN: Deep learning on spatio-temporal graphs. In: IEEE Conf. on Computer Vision and Pattern Recognition (2016) 5, 6
  22. 22.Jain, A., Casas, S., Liao, R., Xiong, Y., Feng, S., Segal, S., Urtasun, R.: Discrete residual flow for probabilistic pedestrian behavior prediction. In: Conf. on Robot Learning (2019) 3
  23. 23.Jang, E., Gu, S., Poole, B.: Categorial reparameterization with gumbel-softmax. In: Int. Conf. on Learning Representations (2017) 8
  24. 24.Kalman, R.E.: A new approach to linear filtering and prediction problems. ASME Journal of Basic Engineering 82, 35–45 (1960) 7, 18
  25. 25.Kesten, R., Usman, M., Houston, J., Pandya, T., Nadhamuni, K., Ferreira, A., Yuan, M., Low, B., Jain, A., Ondruska, P., Omari, S., Shah, S., Kulkarni, A., Kazakova, A., Tao, C., Platinsky, L., Jiang, W., Shet, V.: Lyft Level 5 AV Dataset 2019. https://level5.lyft.com/dataset/ (2019) 4
  26. 26.Kong, J., Pfeifer, M., Schildbach, G., Borrelli, F.: Kinematic and dynamic vehicle models for autonomous driving control design. In: IEEE Intelligent Vehicles Symposium (2015) 2, 6
  27. 27.Kosaraju, V., Sadeghian, A., Martın-Martın, R., Reid, I., Rezatofighi, S.H., Savarese, S.: Social-BiGAT: Multimodal trajectory forecasting using bicycle-GAN and graph attention networks. In: Conf. on Neural Information Processing Systems (2019) 3, 4, 9, 10
  28. 28.LaValle, S.M.: Better unicycle models. In: Planning Algorithms, pp. 743–743. Cambridge Univ. Press (2006) 6, 19
  29. 29.LaValle, S.M.: A simple unicycle. In: Planning Algorithms, pp. 729–730. Cambridge Univ. Press (2006) 19
  30. 30.Lee, N., Choi, W., Vernaza, P., Choy, C.B., Torr, P.H.S., Chandraker, M.: DESIRE: distant future prediction in dynamic scenes with interacting agents. In: IEEE Conf. on Computer Vision and Pattern Recognition (2017) 3, 4
  31. 31.Lee, N., Kitani, K.M.: Predicting wide receiver trajectories in American football. In: IEEE Winter Conf. on Applications of Computer Vision (2016) 3
  32. 32.Lerner, A., Chrysanthou, Y., Lischinski, D.: Crowds by example. Computer Graphics Forum 26(3), 655–664 (2007) 4, 8, 10
  33. 33.Morton, J., Wheeler, T.A., Kochenderfer, M.J.: Analysis of recurrent neural networks for probabilistic modeling of driver behavior. IEEE Transactions on Pattern Analysis & Machine Intelligence 18(5), 1289–1298 (2017) 3
  34. 34.Paden, B., Cap, M., Yong, S.Z., Yershov, D., Frazzoli, E.: A survey of motion ˇ planning and control techniques for self-driving urban vehicles. IEEE Transactions on Intelligent Vehicles 1(1), 33–55 (2016) 2, 6, 19
  35. 35.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A.: Automatic differentiation in PyTorch. In: Conf. on Neural Information Processing Systems - Autodiff Workshop (2017) 9
  36. 36.Pellegrini, S., Ess, A., Schindler, K., Gool, L.v.: You’ll never walk alone: Modeling social behavior for multi-target tracking. In: IEEE Int. Conf. on Computer Vision (2009) 4, 8, 10
  37. 37.Rasmussen, C.E., Williams, C.K.I.: Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning). MIT Press, first edn. (2006) 3
  38. 38.Rhinehart, N., McAllister, R., Kitani, K., Levine, S.: PRECOG: Prediction conditioned on goals in visual multi-agent settings. In: IEEE Int. Conf. on Computer Vision (2019) 3, 4
  39. 39.Rudenko, A., Palmieri, L., Herman, M., Kitani, K.M., Gavrila, D.M., Arras, K.O.: Human motion trajectory prediction: A survey. Int. Journal of Robotics Research 39(8), 895–935 (2020) 3
  40. 40.Sadeghian, A., Kosaraju, V., Sadeghian, A., Hirose, N., Rezatofighi, S.H., Savarese, S.: SoPhie: An attentive GAN for predicting paths compliant to social and physical constraints. In: IEEE Conf. on Computer Vision and Pattern Recognition (2019) 4, 9, 10, 12
  41. 41.Sadeghian, A., Legros, F., Voisin, M., Vesel, R., Alahi, A., Savarese, S.: CAR-Net: Clairvoyant attentive recurrent network. In: European Conf. on Computer Vision (2018) 4, 10, 12
  42. 42.Scholler, C., Aravantinos, V., Lay, F., Knoll, A.: What the constant velocity model can teach us about pedestrian motion prediction. IEEE Robotics and Automation Letters (2020) 22
  43. 43.Sohn, K., Lee, H., Yan, X.: Learning structured output representation using deep conditional generative models. In: Conf. on Neural Information Processing Systems (2015) 3, 4, 7
  44. 44.Tang, Y.C., Salakhutdinov, R.: Multiple futures prediction. In: Conf. on Neural Information Processing Systems (2019) 3
  45. 45.Thiede, L.A., Brahma, P.P.: Analyzing the variety loss in the context of probabilistic trajectory prediction. In: IEEE Int. Conf. on Computer Vision (2019) 9, 11
  46. 46.Thrun, S., Burgard, W., Fox, D.: The extended Kalman filter. In: Probabilistic Robotics, pp. 54–64. MIT Press (2005) 7, 20, 21
  47. 47.Vemula, A., Muelling, K., Oh, J.: Social attention: Modeling attention in human crowds. In: Proc. IEEE Conf. on Robotics and Automation (2018) 3, 5, 9, 10, 11
  48. 48.Wang, J.M., Fleet, D.J., Hertzmann, A.: Gaussian process dynamical models for human motion. IEEE Transactions on Pattern Analysis & Machine Intelligence 30(2), 283–298 (2008) 3
  49. 49.Waymo: Safety report (2018), Available at https://waymo.com/safety/. Retrieved on November 9, 2019 2
  50. 50.Waymo: Waymo Open Dataset: An autonomous driving dataset. https://waymo.com/open/ (2019) 4
  51. 51.Zeng, W., Luo, W., Suo, S., Sadat, A., Yang, B., Casas, S., Urtasun, R.: End-toend interpretable neural motion planner. In: IEEE Conf. on Computer Vision and Pattern Recognition (2019) 3
  52. 52.Zhao, S., Song, J., Ermon, S.: InfoVAE: Balancing learning and inference in variational autoencoders. In: Proc. AAAI Conf. on Artificial Intelligence (2019) 8
  53. 53.Zhao, T., Xu, Y., Monfort, M., Choi, W., Baker, C., Zhao, Y., Wang, Y., Wu, Y.N.: Multi-agent tensor fusion for contextual trajectory prediction. In: IEEE Conf. on Computer Vision and Pattern Recognition (2019) 3, 4, 9, 10, 12

Citation

MLA
Salzmann, T., et al. “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data”. arXiv, 2020, http://arxiv.org/abs/2001.03093v5.
APA
Salzmann, T., Ivanovic, B., Chakravarty, P., & Pavone, M. (2020). Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data. arXiv. http://arxiv.org/abs/2001.03093v5
Chicago
Salzmann, T., B. Ivanovic, P. Chakravarty, and M. Pavone. 2020. “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data”. arXiv. http://arxiv.org/abs/2001.03093v5.
Harvard
Salzmann, T. et al. (2020) “Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2001.03093v5.
Vancouver
1. Salzmann T, Ivanovic B, Chakravarty P, Pavone M (2020) Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data. arXiv

BibTeX

@article{salzmann2020trajectron,
  title = {Trajectron++: Dynamically-Feasible Trajectory Forecasting With Heterogeneous Data},
  author = {Salzmann, Tim and Ivanovic, Boris and Chakravarty, Punarjay and Pavone, Marco},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2001.03093v5},
  eprint = {2001.03093}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF