M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction

Qiao SunXin HuangJunru GuBrian C. WilliamsHang Zhao

article2022CVPR160 citations

Proposes a tractable framework for multi-agent interactive trajectory prediction that factorizes joint distributions into influencer-reactor pairs, avoiding exponential search spaces while achieving state-of-the-art accuracy on the Waymo Open Motion Dataset.

Listen

Autonomous driving systems rely heavily on accurately predicting the future movements of surrounding road participants to avoid collisions and navigate safely. While modern artificial intelligence models effectively forecast individual agent paths in isolation, jointly forecasting realistic, coordinated paths across multiple interacting road users remains a major technical bottleneck. Directly modeling all agents at once causes computational complexity to explode exponentially, whereas predicting paths independently frequently produces unrealistic collisions and ignores how road users react to one another.

The article demonstrates and evaluates M2I, a new framework designed to generate realistic, multi-agent future trajectory predictions efficiently by breaking the complex joint prediction task into structured marginal and conditional forecasting steps.

To achieve this, the approach first classifies pairs of interacting agents into specific roles—either as an influencer that moves independently or a reactor that yields—achieving 90.09% classification accuracy. The system then uses a standard model to predict independent paths for the influencer and subsequently generates reactor trajectories directly conditioned on the influencer’s forecasted paths. The framework was trained and evaluated on the Waymo Open Motion Dataset, analyzing 8-second future trajectories across more than 200,000 real-world driving scenarios using combined vector and image-based scene representations.

Evaluation on the benchmark interactive test set demonstrated several core findings. First, the article's approach achieved state-of-the-art performance in mean average precision, the primary benchmark metric assessing trajectory quality and confidence calibration, reaching 0.08 overall and 0.16 for vehicles, noticeably outperforming established baselines and prior challenge-winning models. Second, the framework significantly reduced overlapping, physically impossible trajectories between agents, lowering overlap rates from 0.42 to 0.20 when evaluated on alternative predictor architectures. Third, while the model slightly trailed specialized architectures in displacement error metrics at final time steps, it produced substantially fewer false-positive predictions. Finally, ablation experiments confirmed that the framework is modular and generalizable across different underlying neural network architectures.

These results indicate that factoring complex driving interactions into directional influencer-reactor relationships is a computationally efficient and scalable path forward for autonomous vehicle safety. By producing scene-compliant, collision-free forecasts with well-calibrated confidence scores, autonomous planning systems can better assess collision risks and execute smoother, safer driving decisions without suffering prohibitive computational delays. Furthermore, the conditional framework enables valuable counterfactual simulation capabilities by allowing planners to simulate how other drivers would react to hypothetical trajectory changes.

Organizations developing autonomous driving software should consider adopting factored conditional architectures like M2I to improve multi-agent interaction forecasting. Immediate technical efforts should focus on pairing the relation and conditional framework with advanced transformer-based context encoders to reduce final displacement errors. Practitioners should also integrate the conditional predictor into simulation pipelines to evaluate interactive safety scenarios.

Key limitations center on data distribution and behavioral assumptions. The framework achieved substantial improvements on vehicles due to ample training data, but gains were negligible on pedestrians and cyclists where interactive examples were scarce. Additionally, the model assumes unidirectional influence and does not explicitly account for mutual, simultaneous negotiation between agents. High confidence is warranted for vehicle-to-vehicle interaction forecasting, while cautious validation and expanded data collection are advised before relying on the system for vulnerable road users and highly complex multi-party interactions.

Cover for M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction

Abstract

Predicting future motions of road participants is an important task for driving autonomously in urban scenes. Existing models excel at predicting marginal trajectories for single agents, yet it remains an open question to jointly predict scene compliant trajectories over multiple agents. The challenge is due to exponentially increasing prediction space as a function of the number of agents. In this work, we exploit the underlying relations between interacting agents and decouple the joint prediction problem into marginal prediction problems. Our proposed approach M2I first classifies interacting agents as pairs of influencers and reactors, and then leverages a marginal prediction model and a conditional prediction model to predict trajectories for the influencers and reactors, respectively. The predictions from interacting agents are combined and selected according to their joint likelihoods. Experiments show that our simple but effective approach achieves state-of-the-art performance on the Waymo Open Motion Dataset interactive prediction benchmark.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Interactive Trajectory Prediction
  • 2.2. Conditional Trajectory Prediction
  • 3. Approach
  • 3.1. Problem Formulation
  • 3.2. Model Overview
  • 3.3. Relation Predictor
  • 3.4. Marginal Trajectory Predictor
  • 3.5. Conditional Trajectory Predictor
  • 3.6. Sample Selector
  • 3.7. Inference
  • 4. Experiments
  • 4.1. Dataset
  • 4.2. Metrics
  • 4.3. Model Details
  • 4.3.1 Context Encoder
  • 4.3.2 Relation Prediction Head
  • 4.3.3 Trajectory Prediction Head
  • 4.3.4 Training Details
  • 4.4. Quantitative Results
  • 4.4.1 Validation Set
  • 4.4.2 Testing Set
  • 4.5. Ablation Study
  • 4.5.1 Relation Prediction
  • 4.5.2 Conditional Prediction
  • 4.5.3 Generalizing to Other Predictors
  • 4.6. Qualitative Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Factorized influencer–reactor formulation

    model/method

    M2I predicts the joint future trajectories of interacting road agents by assigning one agent as an influencer and the other as a reactor. Given observed scene context X=(M,S)X=(M,S), where MM is the map state and SS contains the observed states of all agents, and future trajectories Y=(YI,YR)Y=(Y_I,Y_R) over a finite horizon, M2I approximates the joint distribution as

    P(Y∣X)=P(YI,YR∣X)≈P(YI∣X)P(YR∣X,YI).P(Y\mid X)=P(Y_I,Y_R\mid X)\approx P(Y_I\mid X)P(Y_R\mid X,Y_I).

    Here, YIY_I is the influencer trajectory predicted marginally without conditioning on the reactor, and YRY_R is the reactor trajectory predicted conditionally on the influencer trajectory. This replaces direct learning in the joint trajectory space with one marginal and one conditional prediction problem.

    For a pair of agents with no interaction, M2I instead assumes conditional independence:

    P(Y∣X)≈P(YI∣X)P(YR∣X).P(Y\mid X)\approx P(Y_I\mid X)P(Y_R\mid X).

    The proposed extension to N>2N>2 interacting agents chains the distributions according to the directed influencer relations:

    PN>2(Y∣X)≈∏i=1NP(Yi∣X,Yiinf),P_{N>2}(Y\mid X)\approx \prod_{i=1}^{N}P\bigl(Y_i\mid X,\mathbf{Y}^{\mathrm{inf}}_i\bigr),

    where YiY_i is the future trajectory of agent ii, and Yiinf\mathbf{Y}^{\mathrm{inf}}_i is the set of trajectories of agents predicted to influence agent ii. This extension assumes that the directed dependency graph contains no influence loops.

  2. Knowl 2 — Heuristic relation labeling and relation prediction

    model/method

    M2I represents the interaction between two agents with one of three relation types: PASS, YIELD, or NONE. Training labels are generated from the agents’ ground-truth future positions y1τ,y2τ∈R2y_1^\tau,y_2^\tau\in\mathbb{R}^2 at time steps τ∈{1,…,T}\tau\in\{1,\ldots,T\}. The closest distance between the two future trajectories is

    dI=min⁡τ1∈{1,…,T}min⁡τ2∈{1,…,T}∥y1τ1−y2τ2∥2,d_I=\min_{\tau_1\in\{1,\ldots,T\}}\min_{\tau_2\in\{1,\ldots,T\}}\left\|y_1^{\tau_1}-y_2^{\tau_2}\right\|_2,

    where dId_I is measured in spatial units and ϵd\epsilon_d is a dynamic distance threshold determined by the agents’ sizes. If dI>ϵdd_I>\epsilon_d, the relation is labeled NONE. Otherwise, each agent’s time of reaching the closest interaction point is computed as

    t1=arg⁡min⁡τ1∈{1,…,T}min⁡τ2∈{1,…,T}∥y1τ1−y2τ2∥2, t_1=\mathop{\arg\min}_{\tau_1\in\{1,\ldots,T\}}\min_{\tau_2\in\{1,\ldots,T\}}\left\|y_1^{\tau_1}-y_2^{\tau_2}\right\|_2, t2=arg⁡min⁡τ2∈{1,…,T}min⁡τ1∈{1,…,T}∥y1τ1−y2τ2∥2. t_2=\mathop{\arg\min}_{\tau_2\in\{1,\ldots,T\}}\min_{\tau_1\in\{1,\ldots,T\}}\left\|y_1^{\tau_1}-y_2^{\tau_2}\right\|_2.

    If t1>t2t_1>t_2, agent 1 is labeled as yielding to agent 2; otherwise, agent 1 is labeled as passing agent 2. A YIELD relation assigns the yielding agent as reactor and the other agent as influencer, while a PASS relation reverses these assignments. NONE assigns both agents independent marginal predictors.

    The learned relation predictor uses a context encoder over observed agent states, nearby agents, and map coordinates, followed by a one-layer MLP relation head that outputs probabilities over PASS, YIELD, and NONE. It is trained with cross-entropy loss between the predicted relation distribution and the heuristic relation label. At inference, the relation with maximum predicted probability determines the influencer–reactor assignment.

  3. Knowl 3 — Joint-sample generation and likelihood selection

    algorithm

    M2I generates scene-compliant joint samples using a marginal predictor for the influencer, a conditional predictor for the reactor, and a likelihood-based selector.

    Inputs: observed scene context XX; a trained relation predictor; marginal and conditional trajectory predictors; the number NN of samples produced by each predictor; and the required output count KK.

    Output: KK joint trajectory samples for the interacting agents, with confidence scores.

    1. Apply the relation predictor to XX and select the highest-probability relation.

    2. If the relation is PASS or YIELD, assign the influencer and reactor according to the predicted relation.

    3. Use the marginal predictor to produce NN influencer trajectories with probabilities pI(i)p_I^{(i)}, i∈{1,…,N}i\in\{1,\ldots,N\}.

    4. For every influencer trajectory ii, augment the scene context with that future trajectory and use the conditional predictor to produce NN reactor trajectories with conditional probabilities pR∣I(j∣i)p_{R\mid I}^{(j\mid i)}, j∈{1,…,N}j\in\{1,\ldots,N\}.

    5. Form all N2N^2 joint candidates. The joint probability of candidate (i,j)(i,j) is the product pI(i)pR∣I(j∣i)p_I^{(i)}p_{R\mid I}^{(j\mid i)}.

    6. Return the KK candidates with the highest joint probabilities.

    7. If the predicted relation is NONE, independently generate NN marginal trajectories for each agent, assign each pair the product of its two marginal probabilities, and retain the KK highest-probability pairs.

    The paper uses K=6K=6 for benchmark evaluation. Selecting from N2N^2 candidates preserves multiple possible influencer futures while avoiding the much larger joint goal space required by a fully joint predictor.

  4. Knowl 4 — Modular marginal and conditional predictors

    model/method

    M2I contains three learned modules: a relation predictor, a marginal trajectory predictor, and a conditional trajectory predictor. The modules share the same encoder–decoder design and scene-context encoder, but are trained separately. The architecture diagram on page 3 shows the relation predictor first producing an influencer–reactor direction, the marginal predictor generating influencer samples, and the conditional predictor branching from each influencer sample to generate reactor samples.

    The marginal predictor produces multimodal trajectories and confidence scores for the influencer using only the observed scene context. The conditional predictor uses the same trajectory prediction head but receives an augmented context containing a candidate future trajectory of the influencer. This allows the reactor distribution to depend on the specific predicted behavior of the influencer rather than only on its past and the static scene.

    At training time, the conditional predictor uses teacher forcing: the ground-truth future trajectory of the influencer is supplied as the conditioning trajectory. At inference time, each predicted influencer trajectory is used instead. The framework is predictor-agnostic in principle, so different context encoders and trajectory heads can replace the particular implementations used in the experiments.

  5. Knowl 5 — Concrete M2I architecture and training configuration

    experimental setup

    The implemented M2I model combines vectorized and rasterized scene encoding. The vector encoder follows VectorNet: observed agent trajectories and map lane segments are represented as polylines; an MLP encodes vectors within each polyline, a graph neural network models dependencies, and max pooling summarizes each polyline. Cross-attention over agent and map polylines produces agent features. In parallel, the raster encoder converts the scene into a 60-channel 224×224224\times224 image with a spatial resolution of 1 m×1 m1\,\mathrm{m}\times1\,\mathrm{m} per pixel and processes it with a pretrained VGG16 network. The vector and raster features are concatenated.

    For conditional prediction, the influencer’s future trajectory is added as an extra vector polyline and is also rasterized into 80 additional channels, one for each future time step over the 8-second horizon. The relation head is an MLP with hidden size 128, layer normalization, and ReLU activation, producing logits for the three relation types.

    The trajectory head is DenseTNT. It first predicts a heatmap distribution over agent goals using lane scoring, goal–lane attention, and probability estimation, then regresses the complete trajectory conditioned on the selected goal.

    The relation, marginal, and conditional models are trained separately for 30 epochs with batch size 64 on eight NVIDIA RTX 3080 GPUs. Training uses Adam with initial learning rate 10−310^{-3}; the learning rate is multiplied by 0.70.7 every five epochs. The default hidden size is 128.

  6. Knowl 6 — Waymo interactive-prediction evaluation protocol

    experimental setup

    M2I is evaluated on the Waymo Open Motion Dataset interactive-prediction task. Each example contains map state and up to 1.1 seconds of observed agent states represented by 11 time steps, and the model predicts the joint future trajectories of two interacting agents over the next 8 seconds represented by 80 time steps. The training set contains 204,166 scenarios and the validation set contains 43,479 examples. The dataset identifies likely interacting agent pairs but does not provide their PASS, YIELD, or NONE direction, so the proposed heuristic is used to create relation labels during training.

    The evaluation uses minADE, the average displacement error of the closest of K=6K=6 joint samples; minFDE, the final displacement error of the closest joint sample at the prediction horizon; miss rate (MR), the fraction of examples for which none of the six samples lies within the benchmark’s velocity-dependent lateral and longitudinal thresholds; overlap rate (OR), the percentage of predicted trajectories overlapping another predicted agent trajectory, measured using the most likely joint sample; and mean average precision (mAP), the area under the precision–recall curve of the confidence-scored prediction samples. Lower displacement, miss, and overlap metrics are better, whereas higher mAP is better. The paper uses a modified OR that measures overlap only among the predicted agents, rather than among all objects in the scene.

  7. Knowl 7 — Interactive benchmark performance

    empirical result

    On the Waymo interactive benchmark, M2I generally improves confidence quality and scene-consistent joint prediction, especially in mAP. The following values report mFDE, MR, and mAP for vehicles, pedestrians, and cyclists, followed by mAP over all agent types; lower mFDE and MR are better, and higher mAP is better.

    Validation set: Waymo LSTM Baseline: vehicle (−,0.88,0.01)(-,0.88,0.01), pedestrian (−,0.93,0.02)(-,0.93,0.02), cyclist (−,0.98,0.00)(-,0.98,0.00), all mAP 0.010.01; Waymo Full Baseline: (6.07,0.66,0.08)(6.07,0.66,0.08), (4.20,1.00,0.00)(4.20,1.00,0.00), (6.46,0.83,0.01)(6.46,0.83,0.01), all mAP 0.030.03; SceneTransformer: (3.99,0.49,0.11)(3.99,0.49,0.11), (3.15,0.62,0.06)(3.15,0.62,0.06), (4.69,0.71,0.04)(4.69,0.71,0.04), all mAP 0.070.07; Baseline Marginal: (6.26,0.60,0.16)(6.26,0.60,0.16), (3.59,0.63,0.04)(3.59,0.63,0.04), (6.47,0.76,0.03)(6.47,0.76,0.03), all mAP 0.070.07; Baseline Joint: (11.31,0.64,0.14)(11.31,0.64,0.14), (3.44,0.93,0.01)(3.44,0.93,0.01), (7.16,0.82,0.01)(7.16,0.82,0.01), all mAP 0.050.05; M2I: (5.49,0.55,0.18)(5.49,0.55,0.18), (3.61,0.60,0.06)(3.61,0.60,0.06), (6.26,0.73,0.04)(6.26,0.73,0.04), all mAP 0.090.09.

    Test set: Waymo LSTM Baseline: vehicle (12.40,0.87,0.01)(12.40,0.87,0.01), pedestrian (6.85,0.92,0.00)(6.85,0.92,0.00), cyclist (10.84,0.97,0.00)(10.84,0.97,0.00), all mAP 0.000.00; HeatIRm4: (7.20,0.80,0.07)(7.20,0.80,0.07), (4.06,0.80,0.05)(4.06,0.80,0.05), (6.69,0.85,0.01)(6.69,0.85,0.01), all mAP 0.040.04; AIR2: (5.00,0.64,0.10)(5.00,0.64,0.10), (3.68,0.71,0.04)(3.68,0.71,0.04), (5.47,0.81,0.04)(5.47,0.81,0.04), all mAP 0.050.05; SceneTransformer: (4.08,0.50,0.10)(4.08,0.50,0.10), (3.19,0.62,0.05)(3.19,0.62,0.05), (4.65,0.70,0.04)(4.65,0.70,0.04), all mAP 0.060.06; M2I: (5.65,0.57,0.16)(5.65,0.57,0.16), (3.73,0.60,0.06)(3.73,0.60,0.06), (6.16,0.74,0.03)(6.16,0.74,0.03), all mAP 0.080.08.

    M2I exceeds both Waymo baselines on the reported validation metrics and improves validation all-agent mAP from 0.070.07 for SceneTransformer and the marginal baseline to 0.090.09. On the test set, M2I obtains all-agent mAP 0.080.08, compared with 0.060.06 for SceneTransformer. The improvement is not uniform across displacement metrics: SceneTransformer has lower vehicle and all-agent final displacement errors, while M2I provides stronger confidence calibration and fewer false-positive predictions according to mAP.

  8. Knowl 8 — Relation and conditional-prediction ablations

    empirical result

    The learned relation predictor reaches 90.09% accuracy on the Waymo validation set. Replacing predicted relations with ground-truth relations changes vehicle prediction mAP at 8 seconds by 3.05 percentage points, showing that relation classification quality materially affects the joint predictor.

    For vehicle reactors at 8 seconds, the marginal predictor and conditional predictor produce the following results:

    Could not parse LaTeX table

    Here, M2I Conditional GT conditions on the ground-truth influencer trajectory, whereas M2I Conditional P1 conditions only on the single best predicted influencer trajectory. Ground-truth conditioning improves every metric relative to the marginal predictor, supporting the modeled dependence between influencer and reactor futures. Conditioning on only one predicted influencer trajectory performs worse because errors in that trajectory propagate to the reactor. M2I recovers the benefit by conditioning on multiple influencer samples and selecting the most likely joint samples.

  9. Knowl 9 — Generalization to a different trajectory predictor

    empirical result

    M2I remains effective when its context encoder and trajectory head are replaced by VectorNet and TNT, respectively. On the vehicle validation set at 8 seconds, the comparison is:

    Could not parse LaTeX table

    TNT M2I improves minADE, minFDE, overlap rate, and mAP relative to both TNT Marginal and TNT Joint. The largest gains are in scene compliance and confidence quality: OR decreases from 0.420.42 to 0.200.20 relative to TNT Marginal, while mAP increases from 0.100.10 to 0.140.14. This supports the claim that the influencer–reactor factorization is not tied to DenseTNT.

  10. Knowl 10 — Scope limitations of the factorized interaction assumption

    limitation

    M2I has three stated limitations. First, its final-displacement performance remains behind the state-of-the-art, particularly compared with SceneTransformer; the authors identify replacing the current context encoder with SceneTransformer as a possible future improvement. Second, performance depends strongly on the amount of interactive training data, especially for relation and conditional predictors. The largest mAP gains occur for vehicles, which have more interaction examples, while gains for pedestrians and cyclists are smaller because their interactive training data are limited. Third, the factorization assumes one-way influence and therefore excludes mutual influence between two agents and loopy influence graphs involving more than two agents. The authors report that their heuristic finds an obvious influencer in almost all WOMD interactive cases, but more complicated mutual-influence scenarios remain outside the demonstrated method.

Coverage note — No substantial contributed material was omitted; the page-8 qualitative example was not made a separate knowl because it illustrates the already quantified interaction and scene-compliance benefits without adding a distinct method or aggregate result.

References

  1. 1.Waymo open motion dataset interaction prediction. https://waymo.com/open/challenges/2021/interaction-prediction/, 2021. [Online; Accessed November 16th 2021].
  2. 2.Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social LSTM: Human trajectory prediction in crowded spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 961–971, 2016.
  3. 3.Sergio Casas, Cole Gulino, Renjie Liao, and Raquel Urtasun. SpAGNN: Spatially-aware graph neural networks for relational behavior forecasting from sensor data. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 9491–9497. IEEE, 2020.
  4. 4.Sergio Casas, Cole Gulino, Simon Suo, Katie Luo, Renjie Liao, and Raquel Urtasun. Implicit latent variable model for scene-consistent motion forecasting. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2020.
  5. 5.Yuning Chai, Benjamin Sapp, Mayank Bansal, and Dragomir Anguelov. MultiPath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conference on Robot Learning (CoRL), 2019.
  6. 6.Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. Argoverse: 3d tracking and forecasting with rich maps. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8748–8757, 2019.
  7. 7.Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou, Tsung-Han Lin, Thi Nguyen, Tzu-Kuo Huang, Jeff Schneider, and Nemanja Djuric. Multimodal trajectory predictions for autonomous driving using deep convolutional networks. In 2019 International Conference on Robotics and Automation (ICRA), pages 2090–2096. IEEE, 2019.
  8. 8.Shengzhe Dai, Zhiheng Li, Li Li, Nanning Zheng, and Shuofeng Wang. A flexible and explainable vehicle motion prediction and inference framework combining semi-supervised aog and st-lstm. IEEE Transactions on Intelligent Transportation Systems, 2020.
  9. 9.Nachiket Deo and Mohan M Trivedi. Multi-modal trajectory prediction of surrounding vehicles with maneuver based lstms. In 2018 IEEE Intelligent Vehicles Symposium (IV), pages 1179–1184. IEEE, 2018.
  10. 10.Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R Qi, Yin Zhou, et al. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9710–9719, 2021.
  11. 11.Liangji Fang, Qinhong Jiang, Jianping Shi, and Bolei Zhou. TPNet: Trajectory proposal network for motion prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6797–6806, 2020.
  12. 12.Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid. VectorNet: Encoding hd maps and agent dynamics from vectorized representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11525–11533, 2020.
  13. 13.Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, and Fabien Moutarde. GOHOME: Graph-oriented heatmap output for future motion estimation. arXiv preprint arXiv:2109.01827, 2021.
  14. 14.Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, and Fabien Moutarde. HOME: Heatmap output for future motion estimation. In IEEE International Intelligent Transportation Systems Conference (ITSC), pages 500–507. IEEE, 2021.
  15. 15.Junru Gu, Chen Sun, and Hang Zhao. DenseTNT: End-to-end trajectory prediction from dense goal sets. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15303–15312, 2021.
  16. 16.Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social GAN: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2255–2264, 2018.
  17. 17.Dirk Helbing and Peter Molnar. Social force model for pedestrian dynamics. Physical review E, 51(5):4282, 1995.
  18. 18.Xin Huang, Stephen G McGill, Jonathan A DeCastro, Luke Fletcher, John J Leonard, Brian C Williams, and Guy Rosman. DiversityGAN: Diversity-aware vehicle motion prediction via latent semantic sampling. IEEE Robotics and Automation Letters, 5(4):5089–5096, 2020.
  19. 19.Xin Huang, Guy Rosman, Igor Gilitschenski, Ashkan Jasour, Stephen G McGill, John J Leonard, and Brian C Williams. HYPER: Learned hybrid trajectory prediction via factored inference and adaptive sampling. In International Conference on Robotics and Automation (ICRA), 2022.
  20. 20.Siddhesh Khandelwal, William Qi, Jagjeet Singh, Andrew Hartnett, and Deva Ramanan. What-if motion prediction for autonomous driving. arXiv preprint arXiv:2008.10587, 2020.
  21. 21.ByeoungDo Kim, Seong Hyeon Park, Seokhwan Lee, Elbek Khoshimjonov, Dongsuk Kum, Junsoo Kim, Jeong Soo Kim, and Jun Won Choi. LaPred: Lane-aware prediction of multimodal future trajectories of dynamic agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14636–14645, 2021.
  22. 22.Vineet Kosaraju, Amir Sadeghian, Roberto Mart´ın-Mart´ın, Ian Reid, Hamid Rezatofighi, and Silvio Savarese. Social-BiGAT: Multimodal trajectory forecasting using bicycle-gan and graph attention networks. Advances in Neural Information Processing Systems, 32, 2019.
  23. 23.Sumit Kumar, Yiming Gu, Jerrick Hoang, Galen Clark Haynes, and Micol Marchetti-Bowick. Interaction-based trajectory prediction over a hybrid traffic graph. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5530–5535. IEEE, 2020.
  24. 24.Yen-Ling Kuo, Xin Huang, Andrei Barbu, Stephen G McGill, Boris Katz, John J Leonard, and Guy Rosman. Trajectory prediction with linguistic representations. In International Conference on Robotics and Automation (ICRA), 2022.
  25. 25.Donsuk Lee, Yiming Gu, Jerrick Hoang, and Micol Marchetti-Bowick. Joint interaction and trajectory prediction for autonomous driving using graph neural networks. arXiv preprint arXiv:1912.07882, 2019.
  26. 26.Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker. DESIRE: Distant future prediction in dynamic scenes with interacting agents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 336–345, 2017.
  27. 27.Lingyun Luke Li, Bin Yang, Ming Liang, Wenyuan Zeng, Mengye Ren, Sean Segal, and Raquel Urtasun. End-to-end contextual perception and prediction with interaction transformer. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5784–5791. IEEE, 2020.
  28. 28.Ming Liang, Bin Yang, Rui Hu, Yun Chen, Renjie Liao, Song Feng, and Raquel Urtasun. Learning lane graph representations for motion forecasting. In European Conference on Computer Vision, pages 541–556. Springer, 2020.
  29. 29.Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan-Hui Lee, Ehsan Adeli, Jitendra Malik, and Adrien Gaidon. It is not the journey but the destination: Endpoint conditioned trajectory prediction. In Proceedings of the European Conference on Computer Vision (ECCV), August 2020.
  30. 30.Xiaoyu Mo, Zhiyu Huang, and Chen Lv. Multi-modal interactive agent trajectory prediction using heterogeneous edge-enhanced graph attention network. Workshop on Autonomous Driving, CVPR, 2021.
  31. 31.Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-STGCNN: A social spatio-temporal graph convolutional neural network for human trajectory prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14424–14432, 2020.
  32. 32.Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan, Zhengdong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, et al. Scene Transformer: A unified architecture for predicting multiple agent trajectories. In International Conference on Learning Representations (ICLR), 2022.
  33. 33.Kamra Nitin, Zhu Hao, Trivedi Dweep, Zhang Ming, and Liu Yan. Multi-agent trajectory prediction with fuzzy query attention. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  34. 34.Nicholas Rhinehart, Rowan McAllister, Kris Kitani, and Sergey Levine. PRECOG: Prediction conditioned on goals in visual multi-agent settings. In Proceedings of the IEEE International Conference on Computer Vision, pages 2821–2830, 2019.
  35. 35.Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Multi-agent generative trajectory forecasting with heterogeneous data for control. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, 2020.
  36. 36.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  37. 37.Haoran Song, Di Luan, Wenchao Ding, Michael Y Wang, and Qifeng Chen. Learning to predict vehicle trajectories with model-based planning. In Conference on Robot Learning (CoRL), 2021.
  38. 38.Charlie Tang and Russ R Salakhutdinov. Multiple futures prediction. In Advances in Neural Information Processing Systems, pages 15424–15434, 2019.
  39. 39.Ekaterina Tolstaya, Reza Mahjourian, Carlton Downey, Balakrishnan Vadarajan, Benjamin Sapp, and Dragomir Anguelov. Identifying driver interactions via conditional behavior prediction. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 3473–3479. IEEE, 2021.
  40. 40.T. van der Heiden, N. S. Nagaraja, C. Weiss, and E. Gavves. SafeCritic: Collision-aware trajectory prediction. In British Machine Vision Conference Workshop, 2019.
  41. 41.Allen Wang, Xin Huang, Ashkan Jasour, and Brian Williams. Fast risk assessment for autonomous vehicles using learned models of agent futures. In Robotics: Science and Systems, 2020.
  42. 42.David Wu and Yunan Wu. Air2 for interaction prediction. Workshop on Autonomous Driving, CVPR, 2021.
  43. 43.Kota Yamaguchi, Alexander C Berg, Luis E Ortiz, and Tamara L Berg. Who are you with and where are you going? In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1345–1352, 2011.
  44. 44.Ye Yuan and Kris Kitani. Diverse trajectory forecasting with determinantal point processes. In International Conference on Learning Representations (ICLR), 2020.
  45. 45.Wei Zhan, Liting Sun, Di Wang, Haojie Shi, Aubrey Clausse, Maximilian Naumann, Julius Kummerle, Hendrik Konigshof, Christoph Stiller, Arnaud de La Fortelle, et al. Interaction dataset: An international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps. arXiv preprint arXiv:1910.03088, 2019.
  46. 46.Hang Zhao, Jiyang Gao, Tian Lan, Chen Sun, Benjamin Sapp, Balakrishnan Varadarajan, Yue Shen, Yi Shen, Yuning Chai, Cordelia Schmid, et al. TNT: Target-driven trajectory prediction. In Conference on Robot Learning (CoRL), 2020.
  47. 47.Tianyang Zhao, Yifei Xu, Mathew Monfort, Wongun Choi, Chris Baker, Yibiao Zhao, Yizhou Wang, and Ying Nian Wu. Multi-agent tensor fusion for contextual trajectory prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12126–12134, 2019.

Citation

MLA
Sun, Q., et al. “M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 6533–42, https://doi.org/10.1109/CVPR52688.2022.00643.
APA
Sun, Q., Huang, X., Gu, J., Williams, B. C., & Zhao, H. (2022). M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6533–6542. https://doi.org/10.1109/CVPR52688.2022.00643
Chicago
Sun, Q., X. Huang, J. Gu, B. C. Williams, and H. Zhao. 2022. “M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6533–42. https://doi.org/10.1109/CVPR52688.2022.00643.
Harvard
Sun, Q. et al. (2022) “M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 6533–6542. Available at: https://doi.org/10.1109/CVPR52688.2022.00643.
Vancouver
1. Sun Q, Huang X, Gu J, Williams BC, Zhao H (2022) M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 6533–6542

BibTeX

@inproceedings{Sun_2022, title={M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00643}, DOI={10.1109/cvpr52688.2022.00643}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Sun, Qiao and Huang, Xin and Gu, Junru and Williams, Brian C. and Zhao, Hang}, year={2022}, month=June, pages={6533–6542} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE