M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction
Qiao SunXin HuangJunru GuBrian C. WilliamsHang Zhao
Proposes a tractable framework for multi-agent interactive trajectory prediction that factorizes joint distributions into influencer-reactor pairs, avoiding exponential search spaces while achieving state-of-the-art accuracy on the Waymo Open Motion Dataset.
Autonomous driving systems rely heavily on accurately predicting the future movements of surrounding road participants to avoid collisions and navigate safely. While modern artificial intelligence models effectively forecast individual agent paths in isolation, jointly forecasting realistic, coordinated paths across multiple interacting road users remains a major technical bottleneck. Directly modeling all agents at once causes computational complexity to explode exponentially, whereas predicting paths independently frequently produces unrealistic collisions and ignores how road users react to one another.
The article demonstrates and evaluates M2I, a new framework designed to generate realistic, multi-agent future trajectory predictions efficiently by breaking the complex joint prediction task into structured marginal and conditional forecasting steps.
To achieve this, the approach first classifies pairs of interacting agents into specific roles—either as an influencer that moves independently or a reactor that yields—achieving 90.09% classification accuracy. The system then uses a standard model to predict independent paths for the influencer and subsequently generates reactor trajectories directly conditioned on the influencer’s forecasted paths. The framework was trained and evaluated on the Waymo Open Motion Dataset, analyzing 8-second future trajectories across more than 200,000 real-world driving scenarios using combined vector and image-based scene representations.
Evaluation on the benchmark interactive test set demonstrated several core findings. First, the article's approach achieved state-of-the-art performance in mean average precision, the primary benchmark metric assessing trajectory quality and confidence calibration, reaching 0.08 overall and 0.16 for vehicles, noticeably outperforming established baselines and prior challenge-winning models. Second, the framework significantly reduced overlapping, physically impossible trajectories between agents, lowering overlap rates from 0.42 to 0.20 when evaluated on alternative predictor architectures. Third, while the model slightly trailed specialized architectures in displacement error metrics at final time steps, it produced substantially fewer false-positive predictions. Finally, ablation experiments confirmed that the framework is modular and generalizable across different underlying neural network architectures.
These results indicate that factoring complex driving interactions into directional influencer-reactor relationships is a computationally efficient and scalable path forward for autonomous vehicle safety. By producing scene-compliant, collision-free forecasts with well-calibrated confidence scores, autonomous planning systems can better assess collision risks and execute smoother, safer driving decisions without suffering prohibitive computational delays. Furthermore, the conditional framework enables valuable counterfactual simulation capabilities by allowing planners to simulate how other drivers would react to hypothetical trajectory changes.
Organizations developing autonomous driving software should consider adopting factored conditional architectures like M2I to improve multi-agent interaction forecasting. Immediate technical efforts should focus on pairing the relation and conditional framework with advanced transformer-based context encoders to reduce final displacement errors. Practitioners should also integrate the conditional predictor into simulation pipelines to evaluate interactive safety scenarios.
Key limitations center on data distribution and behavioral assumptions. The framework achieved substantial improvements on vehicles due to ample training data, but gains were negligible on pedestrians and cyclists where interactive examples were scarce. Additionally, the model assumes unidirectional influence and does not explicitly account for mutual, simultaneous negotiation between agents. High confidence is warranted for vehicle-to-vehicle interaction forecasting, while cautious validation and expanded data collection are advised before relying on the system for vulnerable road users and highly complex multi-party interactions.
- Paper: Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data, Tim Salzmann et al. (2020). Trajectron++ provides foundational methodology for graph-structured multi-agent trajectory forecasting and conditional motion prediction under kinematic constraints.
- Paper: Interaction Networks for Learning about Objects, Relations and Physics, Peter W. Battaglia et al. (2016). This paper establishes the core relational reasoning framework for factorizing dynamic systems into distinct pairwise object interactions.
- Paper: Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks, Agrim Gupta et al. (2018). Social GAN introduces key concepts in multi-agent pooling and diverse multi-future forecasting that underpin interactive trajectory prediction.
- Paper: Social LSTM: Human Trajectory Prediction in Crowded Spaces, Alexandre Alahi et al. (2016). Social LSTM provides the classic foundational baseline for data-driven social pooling across interacting agents in shared environments.
- Paper: Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction, Daehee Park et al. (2023). This work extends interactive vehicle prediction by explicitly modeling stochastic future relational dependencies and lane-level occupancy beyond deterministic historical interaction.
- Paper: EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction Reasoning, Chenxin Xu et al. (2023). EqMotion builds upon relational interaction reasoning by enforcing mathematical motion equivariance and geometric invariance in multi-agent forecasting.
- Paper: Planning-oriented Autonomous Driving, Yi Hu et al. (2022). UniAD incorporates multi-agent motion forecasting into an integrated, planning-oriented autonomous driving framework.
- Paper: Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?, Zhiqi Li et al. (2024). This paper critically analyzes how end-to-end planning models utilize perception and agent status versus ego information in open-loop driving benchmarks.
