TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty
Zhengming ZhangRenran TianZhengming Ding
Proposes a compact transformer-based evidential prediction model that captures temporal dynamics from motion features and quantifies pedestrian crossing intention uncertainty to align AI confidence with human annotator disagreements.
As automated driving systems advance toward higher levels of autonomy, safely and smoothly interacting with pedestrians in complex urban environments remains a primary obstacle. Traditional collision-avoidance systems rely heavily on pedestrian trajectory prediction, which typically forecasts only one to two seconds into the future. Human drivers and automated systems require a longer planning horizon of at least three seconds to negotiate interactions comfortably and avoid sudden disruptions. While predicting pedestrian crossing intent offers a solution to extend this horizon, existing models produce simple binary probabilities that overlook the inherent ambiguity of dynamic street scenes and the disagreements common among human observers.
The article develops and evaluates a new algorithm named Transformer-Based Evidential Prediction to address these challenges. The objective is to accurately predict whether a pedestrian intends to cross the street while simultaneously quantifying the model's confidence through an explicit uncertainty metric. Rather than processing heavy raw video streams, the approach relies entirely on compact tabular data, including pedestrian bounding box geometry and vehicle motion details. It uses self-attention mechanisms to capture temporal patterns across video frames and applies evidential deep learning based on Dirichlet probability distributions to estimate uncertainty directly from data.
The algorithm was evaluated on three major benchmark datasets: JAAD, PIE, and PSI. Across all three benchmarks, the method outperformed existing state-of-the-art models. On the JAAD dataset, it raised the area under the receiver operating characteristic curve by approximately nine percent. On the PSI dataset, it boosted classification accuracy by seven percent (reaching 83%) and the balanced F1 score by twelve percent (reaching 0.88). The analysis also revealed a strong inverse relationship between prediction uncertainty and accuracy: filtering out high-uncertainty cases further improved performance, with accuracy climbing to 91% on JAAD and 93% on PIE when rejecting the most uncertain predictions. Furthermore, the model’s predicted uncertainty showed a strong positive correlation (0.60) with human annotator disagreement on frame-by-frame labeled data.
These findings have direct operational and safety implications for automated vehicle design. Quantifying uncertainty provides automated driving systems with a principled mechanism to recognize when a pedestrian scenario is ambiguous. In practice, vehicles can use this confidence score to trigger cautious behaviors, such as slowing down earlier or initiating safer transitions back to manual human control well before an emergency occurs. Because the architecture uses lightweight tabular features rather than full video feeds, it also reduces onboard computational overhead without sacrificing predictive power.
For future development, the article suggests incorporating dynamic, frame-by-frame intention labels into training to better mirror human judgment during ambiguous interactions. System developers should establish calibrated uncertainty thresholds that balance automated decision-making against precautionary interventions. While the results demonstrate robust performance across standard benchmarks, readers should note limitations regarding training diversity, as the model showed higher uncertainty when encountering rare scenarios, such as groups boarding transit or children negotiating crossings. Expanding datasets to cover a broader variety of pedestrian demographics and complex corner cases will be necessary before broad real-world deployment.
- Paper: Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data, Tim Salzmann et al. (2020). Trajectron++ establishes standard deep generative benchmarks and dynamics modeling for multi-agent pedestrian trajectory prediction, framing the foundational forecasting paradigm that TrEP extends into intention and uncertainty estimation.
- Paper: Social LSTM: Human Trajectory Prediction in Crowded Spaces, Alexandre Alahi et al. (2016). Social LSTM introduces foundational deep sequence modeling and interaction pooling for predicting pedestrian motion in shared spaces.
- Paper: Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks, Agrim Gupta et al. (2018). Social GAN provides foundational techniques for generating multi-modal, socially plausible pedestrian trajectories across standard benchmarks.
- Paper: You'll never walk alone: Modeling social behavior for multi-target tracking, S. Pellegrini et al. (2009). This paper establishes the core principles of modeling goal-oriented pedestrian intent and collision avoidance dynamics in crowded scenes.
- Paper: Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models, Kurtland Chua et al. (2018). PETS demonstrates principled aleatoric and epistemic uncertainty quantification in deep trajectory dynamics, providing essential context for evidential uncertainty modeling.
- Paper: Adaptive Trajectory Prediction via Transferable GNN, Yi Xu et al. (2022). This work explores domain adaptation and generalization across diverse pedestrian benchmarks, addressing distribution shifts relevant to multi-dataset intent prediction.
- Paper: M2I: From Factored Marginal Trajectory Prediction to Interactive Prediction, Qiao Sun et al. (2022). M2I provides key methods for structuring influencer-reactor relationships in multi-agent trajectory prediction, informing how vehicle-pedestrian interactions are framed.
- Paper: EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction Reasoning, Chenxin Xu et al. (2023). EqMotion extends sequence-based motion prediction by enforcing geometric equivariance and relational invariance across interacting agents.
- Paper: Leveraging Future Relationship Reasoning for Vehicle Trajectory Prediction, Daehee Park et al. (2023). This work advances multi-agent forecasting by modeling future relational reasoning and probabilistic occupancy distributions over longer planning horizons.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). iTransformer generalizes transformer-based temporal forecasting by inverting attention mechanisms across variates to handle multivariate time series more effectively.
- Paper: Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?, Zhiqi Li et al. (2024). This study critically investigates the dependency on ego-vehicle status versus perceptual features in end-to-end motion planning pipelines.
