Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction

Daehee ParkHobin RyuYunseo YangJegyeong ChoJiwon KimKuk-Jin Yoon

article2023ICLR80 citations

Proposes a vehicle trajectory prediction framework that infers stochastic future interactions from lane-level occupancy and adjacency dynamics, achieving state-of-the-art long-term forecasting accuracy on the nuScenes and Argoverse benchmarks.

Listen

Safe autonomous driving requires vehicles to accurately forecast how surrounding traffic will move in dynamic, complex road environments. Existing trajectory forecasting systems typically infer multi-vehicle interactions based solely on past movements, often relying on deterministic assumptions that fail to capture the wide variety of possible human driving behaviors. The article introduces and evaluates a novel vehicle trajectory prediction framework designed to explicitly model future relationships between agents by incorporating road map structures and probabilistic reasoning.

The proposed method addresses this challenge by first estimating the coarse future path of each vehicle as lane-level waypoint occupancy probabilities over time. It then evaluates the likelihood that pairs of vehicles will traverse adjacent lane segments to determine their potential for future interaction. By combining these geometric relationship probabilities with a probabilistic Gaussian Mixture distribution, the system generates diverse, socially compliant future paths within a conditional generative framework. The approach was evaluated on two widely used real-world autonomous driving datasets: nuScenes, covering six-second future horizons, and Argoverse, covering three-second future horizons.

The results demonstrate substantial improvements over existing methods. On the nuScenes benchmark, the framework achieved state-of-the-art performance, outperforming previous top models across all evaluation metrics. When forecasting ten potential trajectory samples, the model reduced the average prediction error by 5.3% and the miss rate by 8.8% relative to the prior leading approach. Ablation experiments showed that explicit future relationship modeling delivers greater value over longer forecasting horizons: the method achieved over a 10% error reduction on six-second predictions, whereas the performance benefit was halved to roughly 5.6% on three-second prediction tasks. On the Argoverse benchmark, the model delivered significant improvements over baseline models, remaining competitive with top-performing methods.

These findings indicate that integrating lane-level road geometry into probabilistic interaction modeling produces safer, more realistic trajectory forecasts without averaging out complex maneuvers such as yielding or passing. The performance disparity between short- and long-horizon tasks demonstrates that future interaction reasoning becomes critical as autonomous vehicles plan further ahead in time. Engineering and research teams developing autonomous perception and planning systems should consider integrating probabilistic, map-aware interaction modules into long-range motion prediction pipelines.

While the model establishes leading performance in long-range forecasting, its impact is naturally more modest over shorter forecasting horizons where past momentum dominates vehicle behavior. Decision-makers should note that the system relies on high-definition map availability and graph-based lane representations. Further development should explore integrating this future-relationship reasoning module with advanced training strategies and newer goal-conditioned baseline architectures to maximize predictive accuracy across diverse driving conditions.

arXiv: 2305.14715
Cover for Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction

Abstract

Understanding the interaction between multiple agents is crucial for realistic vehicle trajectory prediction. Existing methods have attempted to infer the interaction from the observed past trajectories of agents using pooling, attention, or graph-based methods, which rely on a deterministic approach. However, these methods can fail under complex road structures, as they cannot predict various interactions that may occur in the future. In this paper, we propose a novel approach that uses lane information to predict a stochastic future relationship among agents. To obtain a coarse future motion of agents, our method first predicts the probability of lane-level waypoint occupancy of vehicles. We then utilize the temporal probability of passing adjacent lanes for each agent pair, assuming that agents passing adjacent lanes will highly interact. We also model the interaction using a probabilistic distribution, which allows for multiple possible future interactions. The distribution is learned from the posterior distribution of interaction obtained from ground truth future trajectories. We validate our method on popular trajectory prediction datasets: nuScenes and Argoverse. The results show that the proposed method brings remarkable performance gain in prediction accuracy, and achieves state-of-the-art performance in long-term prediction benchmark dataset.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Goal-conditioned trajectory prediction
  • 2.2 Interaction modeling
  • 2.3 Multi-modal trajectory prediction
  • 3 Formulation
  • 4 Method
  • 4.1 Waypoint Occupancy
  • 4.2 Future Relationship Module (FRM)
  • 4.2.1 Inter-agent Proximity
  • 4.2.2 Prior of the Interaction
  • 4.2.3 CVAE posterior
  • 4.3 Decoder
  • 4.4 Training
  • 5 Experiments
  • 5.1 Quantitative result
  • 5.2 Qualitative result
  • 5.3 Ablation study
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Future Relationship Trajectory Prediction Architecture

    model/method

    The Future Relationship reasoning framework models multi-agent vehicle trajectory prediction as a Conditional Variational Autoencoder (CVAE) problem where interactions among vehicles are conditioned on coarse, lane-based future motion representations rather than solely on past motion histories.

    Given NN agents and a High-Definition (HD) lane graph G=(ℓ,e)\mathcal{G} = (\ell, e) comprising MM lane polyline segments, the architecture proceeds through four main components:

    1. Past and Lane Encoding: Past trajectories x−∈RN×(tp+1)×2\mathbf{x}^- \in \mathbb{R}^{N \times (t_p + 1) \times 2} over time steps −tp:0-t_p:0 and the lane graph G\mathcal{G} are encoded into past motion features hx−\mathbf{h}_x^- and lane features hℓ\mathbf{h}_\ell.
    2. Waypoint Occupancy Prediction: An intermediate representation τ−∈RN×M×tf\tau^- \in \mathbb{R}^{N \times M \times t_f} is predicted, specifying the probability of each agent occupying each lane segment at each future step 1:tf1:t_f.
    3. Future Relationship Module (FRM): FRM calculates smoothed inter-agent proximity (PRPR) across intermediate steps 1:tf−11:t_f-1 using a Graph Convolutional Network (GCN) over lane connectivity. It infers stochastic, directed interaction edge latents ze∈RN×N×d\mathbf{z}_e \in \mathbb{R}^{N \times N \times d} via a Gaussian Mixture (GM) prior distribution pθ(ze∣X,τ)p_\theta(\mathbf{z}_e \mid \mathbf{X}, \tau) during inference, and a single Gaussian posterior distribution qϕ(ze∣X,τ+)q_\phi(\mathbf{z}_e \mid \mathbf{X}, \tau^+) during training from ground-truth future trajectories x+\mathbf{x}^+.
    4. Message Passing and Trajectory Decoding: Sampled interaction edges ze\mathbf{z}_e undergo message passing across agent pairs to form interaction features hR\mathbf{h}_R. The trajectory decoder concatenates hR\mathbf{h}_R with intention features hI\mathbf{h}_I (derived from past motion and goal features sampled from waypoint occupancy at the final horizon τtf\tau_{t_f}) to decode FF multimodal future trajectories Y∈RN×F×tf×2\mathbf{Y} \in \mathbb{R}^{N \times F \times t_f \times 2}.
  2. Knowl 2 — Waypoint Occupancy Formulation

    equation

    Waypoint occupancy represents the temporal lane-level probability distribution of vehicle positions over future time horizons. For NN vehicles, MM lane polyline segments, and a prediction horizon of tft_f steps, the predicted waypoint occupancy tensor τ1:tf−∈RN×M×tf\boldsymbol{\tau}^-_{1:t_f} \in \mathbb{R}^{N \times M \times t_f} is computed from concatenated past motion features hx−\mathbf{h}_x^- and lane features hℓ\mathbf{h}_\ell:

    τ1:tf−=softmax(MLP([hx−,hℓ]))\boldsymbol{\tau}^-_{1:t_f} = \text{softmax}\left(\text{MLP}\left([\mathbf{h}_x^-, \mathbf{h}_\ell]\right)\right)

    where [⋅,⋅][\cdot, \cdot] denotes feature concatenation along the embedding dimension, and the softmax function is normalized across the MM lane segments such that at each future time step t∈{1,…,tf}t \in \{1, \dots, t_f\} and for each agent i∈{1,…,N}i \in \{1, \dots, N\}:

    ∑m=1Mτti,m=1\sum_{m=1}^M \tau_{t}^{i, m} = 1

    During training, ground truth waypoint occupancy τ1:tf+\boldsymbol{\tau}^+_{1:t_f} is derived directly from ground truth future positions and headings.

  3. Knowl 3 — Inter-Agent Proximity via Multi-Relational Lane Graph Smoothing

    equation

    To model interactions where vehicles passing adjacent lanes influence each other, waypoint occupancy τ1:tf−1∈RN×M×(tf−1)\boldsymbol{\tau}_{1:t_f-1} \in \mathbb{R}^{N \times M \times (t_f-1)} over intermediate horizons is smoothed using a 2-hop multi-relational Graph Convolutional Network (GCN) over lane connectivity types E={successor,predecessor,right neighbor,left neighbor,in-same-intersection}\mathcal{E} = \{\text{successor}, \text{predecessor}, \text{right neighbor}, \text{left neighbor}, \text{in-same-intersection}\}:

    τ~1:tf−1=∑e∈Eσ(De−1Aeτ1:tf−1We)∈RN×M×(tf−1)\tilde{\boldsymbol{\tau}}_{1:t_f-1} = \sum_{e \in \mathcal{E}} \sigma\left( D_e^{-1} A_e \boldsymbol{\tau}_{1:t_f-1} W_e \right) \in \mathbb{R}^{N \times M \times (t_f-1)}

    where Ae∈RM×MA_e \in \mathbb{R}^{M \times M} is the adjacency matrix for relationship type ee, DeD_e is the corresponding diagonal degree matrix, WeW_e is a learnable layer weight matrix, and σ(⋅)\sigma(\cdot) denotes a softmax layer followed by a ReLU activation.

    The resulting smoothed occupancy τ~1:tf−1\tilde{\boldsymbol{\tau}}_{1:t_f-1} is contracted across the lane dimension MM via a dot product to compute the temporal inter-agent proximity tensor PR∈RN×N×(tf−1)PR \in \mathbb{R}^{N \times N \times (t_f-1)}:

    PR=τ~1:tf−1⋅(τ~1:tf−1)⊤PR = \tilde{\boldsymbol{\tau}}_{1:t_f-1} \cdot \left(\tilde{\boldsymbol{\tau}}_{1:t_f-1}\right)^\top

  4. Knowl 4 — Gaussian Mixture Prior and Gaussian Posterior for Interaction Edges

    model/method

    Future interaction between vehicle pairs is modeled as a stochastic latent edge zeij∈Rd\mathbf{z}_e^{ij} \in \mathbb{R}^d connecting agent ii to agent jj.

    Prior Distribution (Inference and Training)

    The prior is parameterized as a Gaussian Mixture Model (GMM) with KK discrete interaction modes to capture multimodal social relations from past motion features hx−\mathbf{h}_x^- and inter-agent proximity prij−=PR−[i,j,:]∈Rtf−1pr_{ij}^- = PR^-[i, j, :] \in \mathbb{R}^{t_f - 1}:

    μKij,σKij,πKij=Fθ([prij−,hx−,i,hx−,j])\boldsymbol{\mu}_K^{ij}, \boldsymbol{\sigma}_K^{ij}, \boldsymbol{\pi}_K^{ij} = F_\theta\left([pr_{ij}^-, \mathbf{h}_x^{-, i}, \mathbf{h}_x^{-, j}]\right)

    where FθF_\theta comprises 1D convolutional layers and MLPs, outputting mean μKij∈RK×d\boldsymbol{\mu}_K^{ij} \in \mathbb{R}^{K \times d}, standard deviation σKij∈RK×d\boldsymbol{\sigma}_K^{ij} \in \mathbb{R}^{K \times d}, and mixture logits πKij∈RK\boldsymbol{\pi}_K^{ij} \in \mathbb{R}^K. Discrete mode selection kk is sampled using the Gumbel-Max trick, and the latent prior edge ze−,ij\mathbf{z}_e^{-, ij} is sampled using reparameterized Gaussian noise ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}):

    k=arg⁡max⁡k(πK,kij+gk),gk∼Gumbel(0,1)k = \arg\max_k \left( \pi_{K, k}^{ij} + g_k \right), \quad g_k \sim \text{Gumbel}(0, 1)

    ze−,ij=μkij+σkij⊙ϵ\mathbf{z}_e^{-, ij} = \boldsymbol{\mu}_k^{ij} + \boldsymbol{\sigma}_k^{ij} \odot \boldsymbol{\epsilon}

    Posterior Distribution (Training Only)

    The approximate posterior is parameterized as a single Gaussian conditioned on ground truth future motion features hx+\mathbf{h}_x^+ and ground truth proximity prij+pr_{ij}^+:

    μij,σij=Fϕ([prij+,hx+,i,hx+,j])∈Rd,Rd\boldsymbol{\mu}^{ij}, \boldsymbol{\sigma}^{ij} = F_\phi\left([pr_{ij}^+, \mathbf{h}_x^{+, i}, \mathbf{h}_x^{+, j}]\right) \in \mathbb{R}^d, \mathbb{R}^d

    ze+,ij=μij+σij⊙ϵ,ϵ∼N(0,I)\mathbf{z}_e^{+, ij} = \boldsymbol{\mu}^{ij} + \boldsymbol{\sigma}^{ij} \odot \boldsymbol{\epsilon}, \quad \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})

  5. Knowl 5 — Interaction Feature Aggregation via Directed Message Passing

    equation

    Given latent interaction edge embeddings zeij∈Rd\mathbf{z}_e^{ij} \in \mathbb{R}^d for each pair of vehicles i,j∈{1,…,N}i, j \in \{1, \dots, N\} and agent motion features hxj∈Rdx\mathbf{h}_x^j \in \mathbb{R}^{d_x}, the interaction feature hRi∈Rd\mathbf{h}_R^i \in \mathbb{R}^d for agent ii is aggregated through an asymmetric message passing function:

    hRi=σ′(1N−1∑j≠iNzeij⊗Fp(hxj))\mathbf{h}_R^i = \sigma'\left( \frac{1}{N - 1} \sum_{j \neq i}^N \mathbf{z}_e^{ij} \otimes F_p\left(\mathbf{h}_x^j\right) \right)

    where ⊗\otimes represents the element-wise (Hadamard) product, Fp:Rdx→RdF_p: \mathbb{R}^{d_x} \to \mathbb{R}^d is a learnable projection function mapping agent motion embeddings into the interaction edge feature space, and σ′(⋅)\sigma'(\cdot) is a non-linear activation function. The scaling factor 1N−1\frac{1}{N-1} averages messages across all surrounding agents.

  6. Knowl 6 — Training Objective and NRI Degeneracy Mitigation

    equation

    The overall model is trained end-to-end using a joint objective function consisting of waypoint occupancy negative log-likelihood Lnll\mathcal{L}_{nll}, closed-form Gaussian mixture Kullback–Leibler divergence LKL\mathcal{L}_{KL}, and a goal-restricted reconstruction loss Lrecon\mathcal{L}_{recon}:

    Lall=Lnll+LKL+Lrecon\mathcal{L}_{all} = \mathcal{L}_{nll} + \mathcal{L}_{KL} + \mathcal{L}_{recon}

    1. Waypoint Occupancy Loss: Lnll=−τ+log⁡(τ−)\mathcal{L}_{nll} = -\boldsymbol{\tau}^+ \log\left(\boldsymbol{\tau}^-\right)

    2. GMM Kullback–Leibler Divergence Loss: LKL=−KL[qϕ(ze∣X,τ+) ∥ pθ(ze∣X,τ−)]≈log⁡∑k=1Kπkexp⁡(−KL[qϕ ∥ pθ,k])\mathcal{L}_{KL} = -\text{KL}\left[q_\phi(\mathbf{z}_e \mid \mathbf{X}, \boldsymbol{\tau}^+) \,\Vert\, p_\theta(\mathbf{z}_e \mid \mathbf{X}, \boldsymbol{\tau}^-)\right] \approx \log \sum_{k=1}^K \pi_k \exp\left(-\text{KL}\left[q_\phi \,\Vert\, p_{\theta, k}\right]\right)

    3. Degeneracy-Mitigating Reconstruction Loss: To prevent Neural Relational Inference (NRI) degeneracy—where decoders learn to ignore relation edges during training—the GT trajectory is explicitly conditioned on the GT goal feature. This restricts the latent interaction edge ze\mathbf{z}_e to governing momentary interactive motions: Lrecon=min⁡ze{E[log⁡pθ(Y∣X,ze,τ+)]}\mathcal{L}_{recon} = \min_{\mathbf{z}_e} \left\{ \mathbb{E}\left[ \log p_\theta\left(\mathbf{Y} \mid \mathbf{X}, \mathbf{z}_e, \boldsymbol{\tau}^+\right) \right] \right\}

  7. Knowl 7 — Evaluation on the nuScenes Trajectory Prediction Benchmark

    data/table

    Performance of the proposed method on the nuScenes test set benchmark for 6-second future trajectory forecasting (using 2 seconds of history). Lower values are better for all metrics: Minimum Average Displacement Error (mADEKmADE_K in meters), Miss Rate (MRKMR_K within 2 meters), and Minimum Final Displacement Error (mFDE1mFDE_1 in meters).

    Paper mADE5\text{mADE}_5 mADE10\text{mADE}_{10} MR5\text{MR}_5 MR10\text{MR}_{10} mFDE1\text{mFDE}_1
    Trajectron++ (Salzmann et al., 2020) 1.88 1.51 0.70 0.57 9.52
    P2T (Deo Trivedi, 2020) 1.45 1.16 0.64 0.46 10.5
    AgentFormer (Yuan et al., 2021) 1.86 1.45 - - -
    LaPred (Kim et al., 2021) 1.47 1.12 0.53 0.46 8.37
    MultiPath (Chai et al., 2020) 1.44 1.14 - - 7.69
    GOHOME (Gilles et al., 2022a) 1.42 1.15 0.57 0.47 6.99
    Autobot (Girgis et al., 2021) 1.37 1.03 0.62 0.44 8.19
    THOMAS (Gilles et al., 2022b) 1.33 1.04 0.55 0.42 6.71
    PGP (Deo et al., 2022) 1.27 0.94 0.52 0.34 7.17
    Ours 1.18 0.88 0.48 0.30 6.59

    The proposed method outperforms all baseline and state-of-the-art models across every benchmark metric on the nuScenes test set, achieving relative reductions over the runner-up method (PGP) of 7.1% on mADE5\text{mADE}_5, 6.4% on mADE10\text{mADE}_{10}, 7.7% on MR5\text{MR}_5, 11.8% on MR10\text{MR}_{10}, and outperforming THOMAS on mFDE1\text{mFDE}_1 without post-processing recombination heuristics.

  8. Knowl 8 — Evaluation on the Argoverse Motion Forecasting Benchmark

    data/table

    Performance of the proposed method on the Argoverse validation and test sets (predicting 3 seconds of future from 2 seconds of past). Metrics are Minimum Average Displacement Error (mADE6\text{mADE}_6 in meters) and Minimum Final Displacement Error (mFDE6\text{mFDE}_6 in meters) using K=6K=6 trajectory samples.

    Paper Val set Test set
    mADE6\text{mADE}_6 mFDE6\text{mFDE}_6 mADE6\text{mADE}_6 mFDE6\text{mFDE}_6
    TNT (Zhao et al., 2021) 0.73 1.29 0.94 1.54
    LaneRCNN (Zeng et al., 2021) 0.77 1.19 0.90 1.45
    TPCN (Ye et al., 2021) 0.73 1.15 0.87 1.38
    Autobot (Girgis et al., 2021) 0.73 1.10 0.89 1.41
    mmTransformer (Liu et al., 2021) 0.72 1.21 0.84 1.34
    SceneTransformer (Varadarajan et al., 2022) - - 0.80 1.23
    Multipath++ (Varadarajan et al., 2022) - - 0.79 1.21
    HiVT (Zhou et al., 2022) 0.66 0.96 0.77 1.17
    Baseline 0.71 1.03 0.86 1.30
    Ours 0.68 0.99 0.82 1.27

    The proposed method improves over its baseline on the Argoverse validation set (reducing mADE6\text{mADE}_6 from 0.71 to 0.68 and mFDE6\text{mFDE}_6 from 1.03 to 0.99) and test set (reducing mADE6\text{mADE}_6 from 0.86 to 0.82 and mFDE6\text{mFDE}_6 from 1.30 to 1.27).

  9. Knowl 9 — Impact of Prediction Horizon on Future Relationship Modeling Gain

    data/table

    Comparison of trajectory prediction performance gains across differing future prediction horizons (6 seconds vs. 3 seconds on nuScenes, and 3 seconds on Argoverse). Metrics are mADE1\text{mADE}_1 and mADE6\text{mADE}_6 in meters.

    Dataset Horizon Baseline Ours Relative Improvement (%)
    nuScenes (6 sec) 3.23 / 1.17 2.89 / 1.10 10.5% / 6.0%
    nuScenes (3 sec) 1.26 / 0.50 1.19 / 0.48 5.6% / 4.0%
    Argoverse (3 sec) 1.41 / 0.71 1.33 / 0.68 5.7% / 4.2%

    The relative improvement gained from Future Relationship modeling over the baseline on nuScenes decreases by approximately half (from 10.5% to 5.6% on mADE1\text{mADE}_1) when the prediction horizon is truncated from 6 seconds to 3 seconds. The 3-second gain on nuScenes closely matches the 5.7% gain observed on Argoverse (which natively uses a 3-second horizon), demonstrating that future relationship reasoning delivers substantially greater utility over longer forecasting horizons.

  10. Knowl 10 — Ablation Analysis of Future Relationship Architecture Components

    data/table

    Ablation experiments conducted on the nuScenes dataset evaluating architectural choices and stochastic modeling schemes for sample counts F=1F=1 and F=5F=5. Metrics are mADE/mFDE\text{mADE} / \text{mFDE} in meters.

    Configuration F=1F=1 (mADE / mFDE) F=5F=5 (mADE / mFDE)
    Impact of model design
    Baseline 3.23 / 7.60 1.26 / 2.49
    Ours w/o FR 3.21 / 7.59 1.26 / 2.50
    Ours w/o GCN 3.04 / 6.94 1.22 / 2.41
    Ours w/ Sym 2.99 / 6.78 1.22 / 2.35
    Importance of multimodal stochastic interaction
    Ours w/ GP 2.98 / 6.78 1.20 / 2.33
    Ours w/ Deterministic 2.96 / 6.80 1.28 / 2.52
    Ours (Full) 2.89 / 6.61 1.19 / 2.30

    Key findings from the ablation study include:

    • Future Relationship (FR): Removing FR (Ours w/o FR), which infers relations strictly from past trajectories similar to standard NRI, degrades performance to baseline levels.
    • GCN Smoothing: Omitting multi-relational GCN smoothing (Ours w/o GCN) degrades mADE1\text{mADE}_1 from 2.89 to 3.04 due to noisy binary ground truth occupancy proximity.
    • Asymmetry: Forcing symmetric interaction edges (Ours w/ Sym) underperforms the asymmetric formulation.
    • Stochastic vs. Deterministic: Using a single Gaussian prior (Ours w/ GP) drops performance due to single-modality. Predicting deterministic edge means (Ours w/ Deterministic) yields substantial degradation at F=5F=5 (mADE5\text{mADE}_5 increases from 1.19 to 1.28), demonstrating that stochastic latent edge sampling is essential for diverse multi-sample prediction.

Coverage note — No substantial contributed material was omitted; qualitative trajectory visualizations and standard implementation hyperparameter details from the supplementary material were summarized within relevant method and evaluation knowls.

References

  1. 1.Inhwan Bae, Jin-Hwi Park, and Hae-Gon Jeon. Non-probability sampling network for stochastic human trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6477–6487, 2022.
  2. 2.Alexander Barth and Uwe Franke. Where will the oncoming vehicle be the next second? In 2008 IEEE Intelligent Vehicles Symposium, pp. 1068–1073. IEEE, 2008.
  3. 3.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11618–11628. IEEE Computer Society, 2020.
  4. 4.Defu Cao, Jiachen Li, Hengbo Ma, and Masayoshi Tomizuka. Spectral temporal graph neural network for trajectory prediction. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 1839–1845. IEEE, 2021.
  5. 5.Sandra Carrasco, D Fernández Llorca, and MA Sotelo. Scout: Socially-consistent and understandable graph attention network for trajectory prediction of vehicles and vrus. In 2021 IEEE Intelligent Vehicles Symposium (IV), pp. 1501–1508. IEEE, 2021.
  6. 6.Sergio Casas, Cole Gulino, Simon Suo, Katie Luo, Renjie Liao, and Raquel Urtasun. Implicit latent variable model for scene-consistent motion forecasting. In 2020 European Conference on Computer Vision (ECCV). Springer, 2020.
  7. 7.Yuning Chai, Benjamin Sapp, Mayank Bansal, and Dragomir Anguelov. Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conference on Robot Learning, pp. 86–99. PMLR, 2020.
  8. 8.Rohan Chandra, Tianrui Guan, Srujan Panuganti, Trisha Mittal, Uttaran Bhattacharya, Aniket Bera, and Dinesh Manocha. Forecasting trajectory and behavior of road-agents using spectral clustering in graph-lstms. IEEE Robotics and Automation Letters, 5(3):4882–4890, 2020.
  9. 9.Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. Argoverse: 3d tracking and forecasting with rich maps. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8740–8749. IEEE Computer Society, 2019.
  10. 10.Siyuan Chen, Jiahai Wang, and Guoqing Li. Neural relational inference with efficient message passing mechanisms. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 7055–7063, 2021.
  11. 11.Nachiket Deo and Mohan M Trivedi. Convolutional social pooling for vehicle trajectory prediction. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1549–15498. IEEE, 2018.
  12. 12.Nachiket Deo and Mohan M Trivedi. Trajectory forecasts in unknown environments conditioned on grid-based plans. arXiv preprint arXiv:2001.00735, 2020.
  13. 13.Nachiket Deo, Eric Wolff, and Oscar Beijbom. Multimodal trajectory prediction conditioned on lane-graph traversals. In Conference on Robot Learning, pp. 203–212. PMLR, 2022.
  14. 14.Nat Dilokthanakul, Pedro AM Mediano, Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan. Deep unsupervised clustering with gaussian mixture variational autoencoders. arXiv preprint arXiv:1611.02648, 2016.
  15. 15.Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid. Vectornet: Encoding hd maps and agent dynamics from vectorized representation. In 2020 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  16. 16.Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, and Fabien Moutarde. Gohome: Graph-oriented heatmap output for future motion estimation. In 2022 International Conference on Robotics and Automation (ICRA), pp. 9107–9114. IEEE, 2022a.
  17. 17.Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, and Fabien Moutarde. Thomas: Trajectory heatmap output with learned multi-agent sampling. In International Conference on Learning Representations, 2022b.
  18. 18.Roger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss, Jim Aldon D’Souza, Samira Ebrahimi Kahou, Felix Heide, and Christopher Pal. Latent variable sequential set transformers for joint multi-agent motion prediction. In International Conference on Learning Representations, 2021.
  19. 19.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In 2021 Advances in neural information processing systems (NeurIPS), 2014.
  20. 20.Junru Gu, Chen Sun, and Hang Zhao. Densetnt: End-to-end trajectory prediction from dense goal sets. In 2021 IEEE/CVF Conference on Computer Vision (ICCV), pp. 15283–15292. IEEE, 2021.
  21. 21.Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social gan: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2255–2264, 2018.
  22. 22.Boris Ivanovic and Marco Pavone. The trajectron: Probabilistic multi-agent trajectory modeling with dynamic spatiotemporal graphs. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2375–2384, 2019.
  23. 23.ByeoungDo Kim, Seokhwan Lee, Seong Hyeon Park, Elbek Khoshimjonov, Dongsuk Kum, Junsoo Kim, and Jun Won Choi. Lapred: Lane-aware prediction of multi-modal future trajectories of dynamic agents. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, pp. 14636–14645. IEEE Computer Vision and Pattern Recognition, 2021.
  24. 24.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  25. 25.Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel. Neural relational inference for interacting systems. In International Conference on Machine Learning, pp. 2688–2697. PMLR, 2018.
  26. 26.Vineet Kosaraju, Amir Sadeghian, Roberto Martín-Martín, Ian Reid, Hamid Rezatofighi, and Silvio Savarese. Social-bigat: Multimodal trajectory forecasting using bicycle-gan and graph attention networks. Advances in Neural Information Processing Systems, 32, 2019.
  27. 27.Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker. Desire: Distant future prediction in dynamic scenes with interacting agents. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 336–345, 2017.
  28. 28.Jiachen Li, Fan Yang, Masayoshi Tomizuka, and Chiho Choi. Evolvegraph: Multi-agent trajectory prediction with dynamic relational reasoning. Advances in neural information processing systems, 33:19783–19794, 2020.
  29. 29.Longyuan Li, Jian Yao, Li Wenliang, Tong He, Tianjun Xiao, Junchi Yan, David Wipf, and Zheng Zhang. Grin: Generative relation and intention network for multi-agent trajectory prediction. Advances in Neural Information Processing Systems, 34:27107–27118, 2021a.
  30. 30.Xiao Li, Guy Rosman, Igor Gilitschenski, Cristian-Ioan Vasile, Jonathan A DeCastro, Sertac Karaman, and Daniela Rus. Vehicle trajectory prediction using generative adversarial network with temporal logic syntax tree features. IEEE Robotics and Automation Letters, 6(2):3459–3466, 2021b.
  31. 31.Yaguang Li, Chuizheng Meng, Cyrus Shahabi, and Yan Liu. Structure-informed graph auto-encoder for relational inference and simulation. In ICML Workshop on Learning and Reasoning with Graph-Structured Data, volume 8, pp. 2, 2019.
  32. 32.Ming Liang, Bin Yang, Rui Hu, Yun Chen, Renjie Liao, Song Feng, and Raquel Urtasun. Learning lane graph representations for motion forecasting. In 2020 European Conference on Computer Vision (ECCV). Springer, 2020.
  33. 33.Chiu-Feng Lin, A Galip Ulsoy, and David J LeBlanc. Vehicle dynamics and external disturbance estimation for vehicle path prediction. IEEE Transactions on Control Systems Technology, 8(3): 508–518, 2000.
  34. 34.Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, and Bolei Zhou. Multimodal motion prediction with stacked transformers. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7573–7582. IEEE Computer Society, 2021.
  35. 35.Yecheng Jason Ma, Jeevana Priya Inala, Dinesh Jayaraman, and Osbert Bastani. Likelihood-based diverse sampling for trajectory forecasting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 13279–13288, 2021.
  36. 36.Jean Mercat, Thomas Gilles, Nicole El Zoghby, Guillaume Sandou, Dominique Beauvois, and Guillermo Pita Gil. Multi-head attention for multi-modal joint vehicle motion forecasting. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 9638–9644. IEEE, 2020.
  37. 37.Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, David J Weiss, Ben Sapp, Zhifeng Chen, and Jonathon Shlens. Scene transformer: A unified architecture for predicting future trajectories of multiple agents. In 2022 International Conference on Learning Representations (ICLR), 2022. URL https://openreview.net/forum?id=Wm3EA5OlHsG.
  38. 38.Tung Phan-Minh, Elena Corina Grigore, Freddy A Boulton, Oscar Beijbom, and Eric M Wolff. Covernet: Multimodal behavior prediction using trajectory sets. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14062–14071. IEEE, 2020.
  39. 39.Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data. In European Conference on Computer Vision, pp. 683–700. Springer, 2020.
  40. 40.Charlie Tang and Russ R Salakhutdinov. Multiple futures prediction. Advances in Neural Information Processing Systems, 32, 2019.
  41. 41.Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava, Khaled S Refaat, Nigamaa Nayakanti, Andre Cornman, Kan Chen, Bertrand Douillard, Chi Pang Lam, Dragomir Anguelov, et al. Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction. In 2022 International Conference on Robotics and Automation (ICRA), pp. 7814–7821. IEEE, 2022.
  42. 42.Anirudh Vemula, Katharina Muelling, and Jean Oh. Social attention: Modeling attention in human crowds. In 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018.
  43. 43.Mingkun Wang, Xinge Zhu, Changqian Yu, Wei Li, Yuexin Ma, Ruochun Jin, Xiaoguang Ren, Dongchun Ren, Mingxu Wang, and Wenjing Yang. Ganet: Goal area network for motion forecasting. arXiv preprint arXiv:2209.09723, 2022.
  44. 44.Max Welling and Thomas N Kipf. Semi-supervised classification with graph convolutional networks. In J. International Conference on Learning Representations (ICLR 2017), 2016.
  45. 45.Maosheng Ye, Tongyi Cao, and Qifeng Chen. Tpcn: Temporal point cloud networks for motion forecasting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11318–11327, 2021.
  46. 46.Maosheng Ye, Jiamiao Xu, Xunnong Xu, Tongyi Cao, and Qifeng Chen. Dcms: Motion forecasting with dual consistency and multi-pseudo-target supervision. arXiv preprint arXiv:2204.05859, 2022.
  47. 47.Ye Yuan, Xinshuo Weng, Yanglan Ou, and Kris M Kitani. Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9813–9823, 2021.
  48. 48.Wenyuan Zeng, Ming Liang, Renjie Liao, and Raquel Urtasun. Lanercnn: Distributed representations for graph-centric motion forecasting. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 532–539. IEEE, 2021.
  49. 49.Lingyao Zhang, Po-Hsun Su, Jerrick Hoang, Galen Clark Haynes, and Micol Marchetti-Bowick. Map-adaptive goal-based trajectory prediction. In Conference on Robot Learning, pp. 1371–1383. PMLR, 2021.
  50. 50.Hang Zhao, Jiyang Gao, Tian Lan, Chen Sun, Ben Sapp, Balakrishnan Varadarajan, Yue Shen, Yi Shen, Yuning Chai, Cordelia Schmid, et al. Tnt: Target-driven trajectory prediction. In Conference on Robot Learning, pp. 895–904. PMLR, 2021.
  51. 51.Zikang Zhou, Luyao Ye, Jianping Wang, Kui Wu, and Kejie Lu. Hivt: Hierarchical vector transformer for multi-agent motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8823–8833, 2022.

Citation

MLA
Park, D., et al. “Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction”. arXiv, 2023, http://arxiv.org/abs/2305.14715v1.
APA
Park, D., Ryu, H., Yang, Y., Cho, J., Kim, J., & Yoon, K.-J. (2023). Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction. arXiv. http://arxiv.org/abs/2305.14715v1
Chicago
Park, D., H. Ryu, Y. Yang, J. Cho, J. Kim, and K.-J. Yoon. 2023. “Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction”. arXiv. http://arxiv.org/abs/2305.14715v1.
Harvard
Park, D. et al. (2023) “Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.14715v1.
Vancouver
1. Park D, Ryu H, Yang Y, Cho J, Kim J, Yoon K-J (2023) Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction. arXiv

BibTeX

@article{park2023leveraging,
  title = {Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction},
  author = {Park, Daehee and Ryu, Hobin and Yang, Yunseo and Cho, Jegyeong and Kim, Jiwon and Yoon, Kuk-Jin},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.14715v1},
  eprint = {2305.14715}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors