TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty

Zhengming ZhangRenran TianZhengming Ding

article2023AAAI83 citations

Proposes a compact transformer-based evidential prediction model that captures temporal dynamics from motion features and quantifies pedestrian crossing intention uncertainty to align AI confidence with human annotator disagreements.

Listen

As automated driving systems advance toward higher levels of autonomy, safely and smoothly interacting with pedestrians in complex urban environments remains a primary obstacle. Traditional collision-avoidance systems rely heavily on pedestrian trajectory prediction, which typically forecasts only one to two seconds into the future. Human drivers and automated systems require a longer planning horizon of at least three seconds to negotiate interactions comfortably and avoid sudden disruptions. While predicting pedestrian crossing intent offers a solution to extend this horizon, existing models produce simple binary probabilities that overlook the inherent ambiguity of dynamic street scenes and the disagreements common among human observers.

The article develops and evaluates a new algorithm named Transformer-Based Evidential Prediction to address these challenges. The objective is to accurately predict whether a pedestrian intends to cross the street while simultaneously quantifying the model's confidence through an explicit uncertainty metric. Rather than processing heavy raw video streams, the approach relies entirely on compact tabular data, including pedestrian bounding box geometry and vehicle motion details. It uses self-attention mechanisms to capture temporal patterns across video frames and applies evidential deep learning based on Dirichlet probability distributions to estimate uncertainty directly from data.

The algorithm was evaluated on three major benchmark datasets: JAAD, PIE, and PSI. Across all three benchmarks, the method outperformed existing state-of-the-art models. On the JAAD dataset, it raised the area under the receiver operating characteristic curve by approximately nine percent. On the PSI dataset, it boosted classification accuracy by seven percent (reaching 83%) and the balanced F1 score by twelve percent (reaching 0.88). The analysis also revealed a strong inverse relationship between prediction uncertainty and accuracy: filtering out high-uncertainty cases further improved performance, with accuracy climbing to 91% on JAAD and 93% on PIE when rejecting the most uncertain predictions. Furthermore, the model’s predicted uncertainty showed a strong positive correlation (0.60) with human annotator disagreement on frame-by-frame labeled data.

These findings have direct operational and safety implications for automated vehicle design. Quantifying uncertainty provides automated driving systems with a principled mechanism to recognize when a pedestrian scenario is ambiguous. In practice, vehicles can use this confidence score to trigger cautious behaviors, such as slowing down earlier or initiating safer transitions back to manual human control well before an emergency occurs. Because the architecture uses lightweight tabular features rather than full video feeds, it also reduces onboard computational overhead without sacrificing predictive power.

For future development, the article suggests incorporating dynamic, frame-by-frame intention labels into training to better mirror human judgment during ambiguous interactions. System developers should establish calibrated uncertainty thresholds that balance automated decision-making against precautionary interventions. While the results demonstrate robust performance across standard benchmarks, readers should note limitations regarding training diversity, as the model showed higher uncertainty when encountering rare scenarios, such as groups boarding transit or children negotiating crossings. Expanding datasets to cover a broader variety of pedestrian demographics and complex corner cases will be necessary before broad real-world deployment.

Cover for TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty

Abstract

With rapid development in hardware (sensors and processors) and AI algorithms, automated driving techniques have entered the public's daily life and achieved great success in supporting human driving performance. However, due to the high contextual variations and temporal dynamics in pedestrian behaviors, the interaction between autonomous-driving cars and pedestrians remains challenging, impeding the development of fully autonomous driving systems. This paper focuses on predicting pedestrian intention with a novel transformer-based evidential prediction (TrEP) algorithm. We develop a transformer module towards the temporal correlations among the input features within pedestrian video sequences and a deep evidential learning model to capture the AI uncertainty under scene complexities. Experimental results on three popular pedestrian intent benchmarks have verified the effectiveness of our proposed model over the state-of-the-art. The algorithm performance can be further boosted by controlling the uncertainty level. We systematically compare human disagreements with AI uncertainty to further evaluate AI performance in confusing scenes. The code is released at https://github.com/zzmonlyyou/TrEP.git.

Table of Contents

  • Introduction
  • Related Works
  • Our Proposed Method
  • Preliminary & Motivation
  • Framework Overview
  • Experiment
  • Dataset
  • Evaluation and Metrics
  • Implementation Details
  • Comparison Results
  • Ablation Study
  • Uncertainty Analysis
  • Disagreement Analysis
  • Case Study
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Transformer-Based Evidential Prediction Architecture for Pedestrian Intention

    model/method

    The Transformer-Based Evidential Prediction (TrEP) architecture predicts pedestrian crossing intentions and quantifies decision uncertainty from a temporal sequence of bounding box tracking coordinates and vehicle states, operating entirely without visual image tokens. For an observed sequence of ll frames, the tabular feature vector at frame ii is defined as xi=[bi,ci,ai,ri,actioni]x_i = [b_i, c_i, a_i, r_i, \text{action}_i], where:

    • bi=(x1,y1,x2,y2)b_i = (x_1, y_1, x_2, y_2) is the quaternion representing the upper-left and bottom-right 2D bounding box coordinates of the pedestrian.
    • ci=(cx,cy)c_i = (c_x, c_y) denotes the 2D bounding box center coordinates.
    • aia_i is the area of the pedestrian bounding box.
    • rir_i is the bounding box aspect ratio (length-to-width ratio).
    • actioni\text{action}_i represents ego-vehicle dynamics (such as vehicle speed or discrete driver behavior annotations).

    The feature vectors are processed through the following stages:

    1. Linear feature projection: A feed-forward linear layer projects each xi∈Rfdx_i \in \mathbb{R}^{f_d} to an 8-dimensional representation.
    2. Sinusoidal positional encoding: Temporal order information gig_i is generated via sinusoidal functions of varying frequencies and summed with the projected features: ki=Linear(xi)+gik_i = \text{Linear}(x_i) + g_i.
    3. Transformer encoder: A multi-head self-attention module (1 layer with 2 heads for PIE and PSI datasets; 2 layers with 2 heads for JAAD; feed-forward dimension expanded to 16; dropout rate 0.1) explicitly captures temporal dependencies across the sequence.
    4. Embedding flattening: The sequence embeddings across all ll frames are flattened into a global feature vector f=Flatten(Transformer(k1,k2,…,kl))f = \text{Flatten}(\text{Transformer}(k_1, k_2, \ldots, k_l)).
    5. Evidential output layer: A linear layer followed by a Rectified Linear Unit (ReLU\text{ReLU}) maps ff to a non-negative evidence vector e=[e1,…,eK]⊤∈R≥0Ke = [e_1, \ldots, e_K]^\top \in \mathbb{R}_{\ge 0}^K (where K=2K=2 for binary crossing classification), which parameterizes a Dirichlet distribution over class probabilities.
  2. Knowl 2 — Dirichlet Evidential Formulation and Uncertainty Quantification

    equation

    In TrEP, evidential deep learning replaces point-estimate softmax probabilities with a Dirichlet probability distribution over class assignments for binary pedestrian crossing intention (K=2K=2, where class 1 represents crossing and class 0 represents non-crossing).

    Given the non-negative evidence vector ei=[ei1,…,eiK]⊤e_i = [e_{i1}, \ldots, e_{iK}]^\top produced by the ReLU\text{ReLU} output layer for sample ii, the Dirichlet distribution parameters αi=[αi1,…,αiK]⊤\alpha_i = [\alpha_{i1}, \ldots, \alpha_{iK}]^\top are defined as: αij=eij+1\alpha_{ij} = e_{ij} + 1 The Dirichlet strength (total evidence metric) SiS_i is computed as: Si=∑j=1Kαij=∑j=1K(eij+1)S_i = \sum_{j=1}^K \alpha_{ij} = \sum_{j=1}^K (e_{ij} + 1) The expected class probability E[pij]\mathbb{E}[p_{ij}] and variance Var(pij)\text{Var}(p_{ij}) for class jj under the Dirichlet distribution Dir(pi∣αi)\text{Dir}(p_i \mid \alpha_i) are given by: E[pij]=αijSi\mathbb{E}[p_{ij}] = \frac{\alpha_{ij}}{S_i} Var(pij)=E[pij](1−E[pij])Si+1\text{Var}(p_{ij}) = \frac{\mathbb{E}[p_{ij}](1 - \mathbb{E}[p_{ij}])}{S_i + 1} The overall model predictive uncertainty ui∈(0,1]u_i \in (0, 1] is defined as the ratio of the number of classes KK to the Dirichlet strength SiS_i: ui=KSiu_i = \frac{K}{S_i} When zero evidence is predicted (eij=0e_{ij} = 0 for all jj), Si=KS_i = K, yielding maximal uncertainty ui=1u_i = 1 and a uniform probability expectation E[pij]=1/K\mathbb{E}[p_{ij}] = 1/K. As total evidence grows (Si→∞S_i \to \infty), uncertainty approaches zero (ui→0u_i \to 0).

  3. Knowl 3 — Evidential Loss Function with Dirichlet Regularization

    equation

    The TrEP model parameters Θ\Theta are trained by minimizing an objective function combining an evidential mean squared error with a Kullback-Leibler (KL) divergence regularizer, replacing cross-entropy loss: L(Θ)=∑i=1N(Li(Θ)+λ KL[D(pi∣αi)∥D(pi∣1)])\mathcal{L}(\Theta) = \sum_{i=1}^N \left( \mathcal{L}_i(\Theta) + \lambda \,\text{KL}\left[ D(p_i \mid \alpha_i) \parallel D(p_i \mid \mathbf{1}) \right] \right) where NN is the total number of training samples, λ>0\lambda > 0 is a regularization trade-off parameter (set to λ=10\lambda = 10), yi=[yi1,…,yiK]⊤y_i = [y_{i1}, \ldots, y_{iK}]^\top is the ground-truth one-hot label vector for sample ii, and D(pi∣1)D(p_i \mid \mathbf{1}) represents a uniform Dirichlet prior parameterizing complete uncertainty (the "I do not know" state).

    The sample loss Li(Θ)\mathcal{L}_i(\Theta) minimizes expected prediction error while penalizing the variance of the predicted Dirichlet distribution across all KK classes: Li(Θ)=∑j=1K((yij−E[pij])2+Var(pij))\mathcal{L}_i(\Theta) = \sum_{j=1}^K \left( (y_{ij} - \mathbb{E}[p_{ij}])^2 + \text{Var}(p_{ij}) \right) where E[pij]=αijSi\mathbb{E}[p_{ij}] = \frac{\alpha_{ij}}{S_i} and Var(pij)=E[pij](1−E[pij])Si+1\text{Var}(p_{ij}) = \frac{\mathbb{E}[p_{ij}](1 - \mathbb{E}[p_{ij}])}{S_i + 1}.

    The KL divergence term regularizes predictions by penalizing misleading evidence allocations that deviate from the uniform prior without improving sample fit: KL[D(pi∣αi)∥D(pi∣1)]=ln⁡(Γ(∑j=1Kαij)Γ(K)∏j=1KΓ(αij))+∑j=1K(αij−1)[ψ(αij)−ψ(∑m=1Kαim)]\text{KL}\left[ D(p_i \mid \alpha_i) \parallel D(p_i \mid \mathbf{1}) \right] = \ln\left( \frac{\Gamma\left(\sum_{j=1}^K \alpha_{ij}\right)}{\Gamma(K) \prod_{j=1}^K \Gamma(\alpha_{ij})} \right) + \sum_{j=1}^K (\alpha_{ij} - 1)\left[ \psi(\alpha_{ij}) - \psi\left(\sum_{m=1}^K \alpha_{im}\right) \right] where Γ(⋅)\Gamma(\cdot) is the gamma function and ψ(⋅)\psi(\cdot) is the digamma function.

  4. Knowl 4 — Pedestrian Crossing Intention Prediction Performance on PIE and JAAD Benchmarks

    data/table

    The performance of the proposed Transformer-Based Evidential Prediction (TrEP) model is compared against baseline models on the PIE and JAAD datasets. Baselines include ATGC, I3D, MM-LSTM, SF-GRU, PCPA, MMHA, and BiPed. Evaluations for the proposed model include the Base Model (standard cross-entropy), the Evidential Model with all samples (u≤1u \le 1), and the Evidential Model with an uncertainty rejection threshold of u≤0.6u \le 0.6.

    Model PIE JAAD
    Accuracy AUC F1 Precision Accuracy AUC F1 Precision
    ATGC 0.59 0.55 0.36 0.35 0.64 0.60 0.53 0.50
    I3D 0.79 0.75 0.64 0.61 0.82 0.75 0.55 0.49
    MM-LSTM 0.84 0.84 0.75 0.68 0.80 0.77 0.58 0.51
    SF-GRU 0.86 0.83 0.75 0.73 0.83 0.77 0.58 0.51
    PCPA 0.86 0.84 0.76 0.73 0.83 0.77 0.57 0.50
    MMHA 0.89 0.88 0.81 0.77 0.84 0.80 0.62 0.54
    BiPed 0.91 0.90 0.85 0.82 0.83 0.79 0.60 0.52
    Ours (Base) 0.91 0.93 0.85 0.84 0.87 0.88 0.63 0.63
    Ours (u≤1u \le 1) 0.92 0.94 0.85 0.88 0.88 0.86 0.61 0.70
    Ours (u≤0.6u \le 0.6) 0.93 0.94 0.87 0.89 0.91 0.86 0.69 0.71

    When evaluating all samples (u≤1u \le 1), the proposed model achieves 0.92 Accuracy and 0.94 AUC on PIE, and 0.88 Accuracy and 0.86 AUC on JAAD. Setting an uncertainty threshold u≤0.6u \le 0.6 retains 96% of PIE and 89% of JAAD samples, boosting PIE Accuracy to 0.93 and JAAD Accuracy to 0.91 while increasing JAAD F1 score from 0.61 to 0.69.

  5. Knowl 5 — Pedestrian Crossing Intention Prediction Performance on PSI Benchmark

    data/table

    The performance of the proposed model is evaluated on the PSI pedestrian intention dataset against VR-GCN, PIE-Intention, and PSI-Intention. Performance metrics include Accuracy, Balanced Accuracy, and F1 score for the Base Model, the Evidential Model retaining all samples (u≤1u \le 1), and the Evidential Model filtering uncertain samples with threshold u≤0.6u \le 0.6.

    Model Accuracy Balanced Accuracy F1
    VR-GCN 0.74 0.61 0.64
    PIE-Intention 0.69 0.58 0.79
    PSI-Intention 0.76 0.67 0.66
    Ours (Base) 0.83 0.75 0.88
    Ours (u≤1u \le 1) 0.82 0.75 0.87
    Ours (u≤0.6u \le 0.6, 75% included) 0.85 0.77 0.90

    On the PSI benchmark, the proposed evidential model (u≤1u \le 1) achieves 0.82 Accuracy, 0.75 Balanced Accuracy, and 0.87 F1 score, outperforming PSI-Intention by 6 percentage points in accuracy and 21 percentage points in F1 score. Restricting predictions to confident cases (u≤0.6u \le 0.6, retaining 75% of the data) further elevates Accuracy to 0.85 and F1 score to 0.90.

  6. Knowl 6 — Feature and Positional Encoding Ablation in TrEP Base Model

    data/table

    An ablation study demonstrates the impact of individual tabular bounding box features and sinusoidal positional encoding on the base Transformer architecture across PIE, JAAD, and PSI benchmarks.

    Model Configuration PIE JAAD PSI
    Accuracy F1 Accuracy F1 Accuracy F1
    Bbox+Action 0.80 0.72 0.79 0.58 0.72 0.69
    Bbox+Action+Center 0.91 0.85 0.87 0.63 0.80 0.85
    Bbox+Action+Center+Ratio 0.89 0.81 0.86 0.65 0.83 0.88
    No Pos. Encoder 0.90 0.83 0.85 0.61 0.81 0.87

    The ablation results reveal:

    1. Adding bounding box center coordinates (+Center) to the baseline coordinates and ego-vehicle action provides the largest single performance gain, increasing prediction accuracy by at least 8% across all datasets (0.80 to 0.91 on PIE, 0.79 to 0.87 on JAAD, and 0.72 to 0.80 on PSI).
    2. Incorporating bounding box area and aspect ratio (+Ratio) yields minor decreases in accuracy on PIE (0.91 to 0.89) and JAAD (0.87 to 0.86) while improving accuracy on PSI (0.80 to 0.83).
    3. Eliminating the positional encoder (No Pos. Encoder) leads to a consistent drop in Accuracy and F1 score across all three benchmarks, confirming the utility of explicit temporal ordering.
  7. Knowl 7 — Negative Correlation Between Evidential Uncertainty and Model Prediction Accuracy

    empirical result

    Empirical evaluation of the TrEP evidential model across JAAD, PIE, and PSI shows a strong negative relationship between predicted uncertainty uu and model performance metrics. When grouping test instances into uncertainty bins u∈(0,0.1],(0.1,0.2],…,(0.9,1.0]u \in (0, 0.1], (0.1, 0.2], \ldots, (0.9, 1.0]:

    1. Precision, F1 score, Accuracy, and AUC decline systematically as predicted uncertainty increases.
    2. In cases where uncertainty reaches u=1.0u = 1.0, the model outputs near-zero evidence (e≈0e \approx 0), defaulting toward crossing classifications that produce sharp drops in F1 score and balanced accuracy.
    3. Setting an uncertainty rejection threshold at u≤0.6u \le 0.6 allows the model to reliably filter out high-error cases, increasing Accuracy from 0.92 to 0.93 on PIE (96% retention), 0.88 to 0.91 on JAAD (89% retention), and 0.82 to 0.85 on PSI (75% retention).
  8. Knowl 8 — Alignment Between Evidential AI Uncertainty and Human Annotator Disagreement

    empirical result

    To analyze whether evidential uncertainty reflects human perceptual ambiguity, TrEP model uncertainty uu was evaluated against human annotator disagreement measured by decision entropy H(p)=−∑cpclog⁡pcH(p) = -\sum_{c} p_c \log p_c from crowd-sourced annotations on the PIE and PSI datasets. Across both datasets, model accuracy drops as human disagreement increases.

    However, the linear correlation between predicted evidential uncertainty uu and annotator disagreement entropy differs substantially depending on dataset annotation scheme:

    • On the PSI dataset, where annotators label crossing intentions dynamically at every frame (intention segmentation), model uncertainty exhibits a strong positive correlation with human disagreement (r=0.60,p<0.001r = 0.60, p < 0.001).
    • On the PIE dataset, where pedestrians receive a single fixed intention label across the entire multi-second sequence, model uncertainty shows a weak negative correlation with human disagreement (r=−0.17,p<0.001r = -0.17, p < 0.001).

    This demonstrates that frame-level dynamic intention annotations enable evidential models to learn human-aligned cognitive uncertainty patterns, whereas static sequence-level labeling obscures temporal negotiation dynamics.

  9. Knowl 9 — Pedestrian Intention Prediction Experimental Protocols and Dataset Specifications

    experimental setup

    The experimental validation of pedestrian crossing intention prediction uses three benchmark datasets with distinct sampling protocols and feature inputs:

    1. PIE Dataset: Dashcam video from a 4-hour drive in Toronto downtown. Sequences are sampled at least 1.0 second prior to crossing action onset with an overlap ratio of 0.5, yielding 3,980 training sequences (995 crossing). Vehicle dynamics are annotated via vehicle speed.
    2. JAAD Dataset: Dashcam video clips sampled at least 1.0 second prior to crossing onset with an overlap ratio of 0.5, yielding 3,955 training sequences (805 crossing). Vehicle dynamics are provided as discrete driver behavior actions.
    3. PSI Dataset: Dashcam video clips sampled across entire pedestrian tracks with an overlap ratio of 0.8. The task takes 15 historical frames as input to predict the crossing intention label at the 16th frame, yielding 6,262 training sequences (3,927 crossing). PSI provides frame-by-frame intention segmentation labels and contains no ego-vehicle action annotations.

    Model Training Setup:

    • Input feature dimension fdf_d is linearly projected to 8, and the transformer fully connected layers project to 16.
    • Multi-head self-attention: 1 layer with 2 heads for PIE and PSI; 2 layers with 2 heads for JAAD; dropout rate of 0.1.
    • Training: Adam optimizer with learning rate 5×10−35 \times 10^{-3}, batch size 64, trained for 2,000 epochs with regularization weight λ=10\lambda = 10.
  10. Knowl 10 — Limitations in Pedestrian Intention Modeling from Annotation Granularity and Scene Rarity

    limitation

    The effectiveness and uncertainty calibration of the TrEP intention prediction model face two primary limitations:

    1. Static Labeling Granularity: Datasets that assign a single static intention label across long video spans (such as PIE) force the model to overlook localized intention dynamics and temporal negotiation, leading to an inverse correlation between evidential uncertainty and human disagreement (r=−0.17r = -0.17).
    2. Scarcity of Atypical Social Interactions: In complex or atypical scenarios—such as pedestrians walking in front of a vehicle specifically to board a bus, or children crossing with adults while actively negotiating right-of-way—the lack of representative training samples causes the model to predict high uncertainty (u=1.0u = 1.0) or misclassify crossing intentions despite realistic lateral trajectories.

Coverage note — None was omitted; all key architectural components, evidential mathematical formulations, empirical benchmark comparisons, ablation studies, uncertainty analyses, human disagreement correlations, and stated limitations are fully covered.

References

  1. 1.Aliakbarian, M. S.; Saleh, F. S.; Salzmann, M.; Fernando, B.; Petersson, L.; and Andersson, L. 2018. VIENA2: A Driving Anticipation Dataset. In Asian Conference on Computer Vision, 449–466. Springer.
  2. 2.Amini, A.; Schwarting, W.; Soleimany, A.; and Rus, D. 2020. Deep evidential regression. Advances in Neural Information Processing Systems, 33: 14927–14937.
  3. 3.Bao, W.; Yu, Q.; and Kong, Y. 2021. Evidential deep learning for open set action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 13349–13358.
  4. 4.Bewley, A.; Ge, Z.; Ott, L.; Ramos, F.; and Upcroft, B. 2016. Simple online and realtime tracking. In 2016 IEEE International Conference on Image Processing (ICIP), 3464–3468.
  5. 5.Carreira, J.; and Zisserman, A. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6299–6308.
  6. 6.Chen, T.; and Tian, R. 2021. A survey on deep-learning methods for pedestrian behavior prediction from the egocentric view. In 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), 1898–1905. IEEE.
  7. 7.Chen, T.; Tian, R.; Chen, Y.; Domeyer, J.; Toyoda, H.; Sherony, R.; Jing, T.; and Ding, Z. 2021. PSI: A Pedestrian Behavior Dataset for Socially Intelligent Autonomous Car. arXiv preprint arXiv:2112.02604.
  8. 8.Chen, T.; Tian, R.; and Ding, Z. 2021. Visual reasoning using graph convolutional networks for predicting pedestrian crossing intention. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 3103–3109.
  9. 9.Cui, Y.; Cao, Z.; Xie, Y.; Jiang, X.; Tao, F.; Chen, Y. V.; Li, L.; and Liu, D. 2022. Dg-labeler and dgl-mots dataset: Boost the autonomous driving perception. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 58–67.
  10. 10.Ding, Y.; Harirchi, F.; Yong, S. Z.; Jacobsen, E.; and Ozay, N. 2018. Optimal input design for affine model discrimination with applications in intention-aware vehicles. In 2018 ACM/IEEE 9th International Conference on Cyber-Physical Systems (ICCPS), 297–307. IEEE.
  11. 11.Domeyer, J. E.; Lee, J. D.; and Toyoda, H. 2020. Vehicle automation–Other road user communication and coordination: Theory and mechanisms. IEEE Access, 8: 19860–19872.
  12. 12.Eriksson, A.; and Stanton, N. A. 2017. Takeover time in highly automated vehicles: noncritical transitions to and from manual control. Human factors, 59(4): 689–705.
  13. 13.Fang, Z.; and Lopez, A. M. 2018. Is the pedestrian going to cross? answering by 2d pose estimation. In 2018 IEEE intelligent vehicles symposium (IV), 1271–1276. IEEE.
  14. 14.Gujjar, P.; and Vaughan, R. 2019. Classifying pedestrian actions in advance using predicted video of urban driving scenes. In 2019 International Conference on Robotics and Automation (ICRA), 2097–2103. IEEE.
  15. 15.Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; and Wang, Y. 2021. Transformer in transformer. Advances in Neural Information Processing Systems, 34.
  16. 16.Herman, M.; Wagner, J.; Prabhakaran, V.; Moser, N.; Ziesche, H.; Ahmed, W.; Burkle, L.; Kloppenburg, E.; and Glaser, C. 2021. Pedestrian Behavior Prediction for Automated Driving: Requirements, Metrics, and Relevant Features. IEEE Transactions on Intelligent Transportation Systems.
  17. 17.Ji, W.; Yu, S.; Wu, J.; Ma, K.; Bian, C.; Bi, Q.; Li, J.; Liu, H.; Cheng, L.; and Zheng, Y. 2021. Learning calibrated medical image segmentation via multi-rater agreement modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12341–12351.
  18. 18.Jing, T.; Xia, H.; Tian, R.; Ding, H.; Luo, X.; Domeyer, J.; Sherony, R.; and Ding, Z. 2022. Inaction: Interpretable action decision making for autonomous driving. In European Conference on Computer Vision, 370–387. Springer.
  19. 19.Kotseruba, I.; Rasouli, A.; and Tsotsos, J. K. 2020. Do they want to cross? understanding pedestrian intention for behavior prediction. In 2020 IEEE Intelligent Vehicles Symposium (IV), 1688–1693. IEEE.
  20. 20.Kotseruba, I.; Rasouli, A.; and Tsotsos, J. K. 2021. Benchmark for evaluating pedestrian action prediction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 1258–1268.
  21. 21.Litman, T. 2017. Autonomous vehicle implementation predictions. Victoria Transport Policy Institute Victoria, BC, Canada.
  22. 22.Liu, B.; Adeli, E.; Cao, Z.; Lee, K.-H.; Shenoi, A.; Gaidon, A.; and Niebles, J. C. 2020a. Spatiotemporal relationship reasoning for pedestrian intent prediction. IEEE Robotics and Automation Letters, 5(2): 3485–3492.
  23. 23.Liu, D.; Cui, Y.; Chen, Y.; Zhang, J.; and Fan, B. 2020b. Video object detection for autonomous driving: Motion-aid feature calibration. Neurocomputing, 409: 1–11.
  24. 24.Liu, D.; Cui, Y.; Guo, X.; Ding, W.; Yang, B.; and Chen, Y. 2021. Visual localization for autonomous driving: Mapping the accurate location in the city maze. In 2020 25th International Conference on Pattern Recognition (ICPR), 3170–3177. IEEE.
  25. 25.Liu, X.; Masoud, N.; Zhu, Q.; and Khojandi, A. 2022. A markov decision process framework to incorporate network-level data in motion planning for connected and automated vehicles. Transportation Research Part C: Emerging Technologies, 136: 103550.
  26. 26.Liu, X.; Zhao, G.; Masoud, N.; and Zhu, Q. 2020c. Trajectory planning for connected and automated vehicles: Cruising, lane changing, and platooning. arXiv preprint arXiv:2001.08620.
  27. 27.Ma, X.; Karimpour, A.; and Wu, Y.-J. 2020. Statistical evaluation of data requirement for ramp metering performance assessment. Transportation Research Part A: Policy and Practice, 141: 248–261.
  28. 28.Merat, N.; Jamson, A. H.; Lai, F. C.; Daly, M.; and Carsten, O. M. 2014. Transition to manual: Driver behaviour when resuming control from a highly automated vehicle. Transportation research part F: traffic psychology and behaviour, 27: 274–282.
  29. 29.Pang, Y.; Guo, Z.; and Zhuang, B. 2022. ProspectNet: Weighted Conditional Attention for Future Interaction Modeling in Behavior Prediction. arXiv preprint arXiv:2208.13848.
  30. 30.Qu, X.; Mei, Q.; Liu, P.; and Hickey, T. 2020. Using EEG to distinguish between writing and typing for the same cognitive task. In Brain Function Assessment in Learning: Second International Conference, BFAL 2020, Heraklion, Crete, Greece, October 9–11, 2020, Proceedings 2, 66–74. Springer.
  31. 31.Rasouli, A.; Kotseruba, I.; Kunic, T.; and Tsotsos, J. K. 2019. PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory Prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  32. 32.Rasouli, A.; Kotseruba, I.; and Tsotsos, J. K. 2017. Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior. In Proceedings of the IEEE International Conference on Computer Vision Workshops, 206–213.
  33. 33.Rasouli, A.; Kotseruba, I.; and Tsotsos, J. K. 2020. Pedestrian action anticipation using contextual feature fusion in stacked rnns. arXiv preprint arXiv:2005.06582.
  34. 34.Rasouli, A.; Rohani, M.; and Luo, J. 2021. Bifold and Semantic Reasoning for Pedestrian Behavior Prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 15600–15610.
  35. 35.Rasouli, A.; Yau, T.; Rohani, M.; and Luo, J. 2022. Multi-Modal Hybrid Architecture for Pedestrian Action Prediction. In 2022 IEEE Intelligent Vehicles Symposium (IV), 91–97.
  36. 36.Sensoy, M.; Kaplan, L.; Cerutti, F.; and Saleki, M. 2020. Uncertainty-aware deep classifiers using generative models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 5620–5627.
  37. 37.Sensoy, M.; Kaplan, L.; and Kandemir, M. 2018. Evidential deep learning to quantify classification uncertainty. Advances in neural information processing systems, 31.
  38. 38.Shi, L.; Wang, L.; Long, C.; Zhou, S.; Zhou, M.; Niu, Z.; and Hua, G. 2021. SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8994–9003.
  39. 39.Simonyan, K.; and Zisserman, A. 2014. Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems, 27.
  40. 40.Tang, Y.; Song, S.; Gui, S.; Chao, W.; Cheng, C.; and Qin, R. 2023. Active and Low-Cost Hyperspectral Imaging for the Spectral Analysis of a Low-Light Environment. Sensors, 23(3): 1437.
  41. 41.Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, 4489–4497.
  42. 42.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017a. Attention is all you need. Advances in neural information processing systems, 30.
  43. 43.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017b. Attention is All you Need. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  44. 44.Wang, C.; Wang, Y.; Xu, M.; and Crandall, D. J. 2022. Stepwise goal-driven networks for trajectory prediction. IEEE Robotics and Automation Letters, 7(2): 2716–2723.
  45. 45.Wu, J.; Fang, H.; Shang, F.; Wang, Z.; Yang, D.; Zhou, W.; Yang, Y.; and Xu, Y. 2022. Learning self-calibrated optic disc and cup segmentation from multi-rater annotations. arXiv preprint arXiv:2206.05092.
  46. 46.Wu, J.; Fu, R.; Fang, H.; Zhang, Y.; and Xu, Y. 2023. MedSegDiffV2: Diffusion based Medical Image Segmentation with Transformer. arXiv preprint arXiv:2301.11798.
  47. 47.Xu, T.; Chen, W.; Pichao, W.; Wang, F.; Li, H.; and Jin, R. 2021. CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation. In International Conference on Learning Representations.
  48. 48.Xu, Y.; Piao, Z.; and Gao, S. 2018. Encoding Crowd Interaction With Deep Neural Network for Pedestrian Trajectory Prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  49. 49.Yagi, T.; Mangalam, K.; Yonetani, R.; and Sato, Y. 2018. Future Person Localization in First-Person Videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  50. 50.Yang, D.; Zhang, H.; Yurtsever, E.; Redmill, K.; and Ozguner, U. 2022. Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention. IEEE Transactions on Intelligent Vehicles.
  51. 51.Yao, Y.; Atkins, E.; Johnson-Roberson, M.; Vasudevan, R.; and Du, X. 2021. BiTraP: Bi-Directional Pedestrian Trajectory Prediction With Multi-Modal Goal Estimation. IEEE Robotics and Automation Letters, 6(2): 1463–1470.
  52. 52.Yi, L.; and Qu, X. 2022. Attention-Based CNN Capturing EEG Recording’s Average Voltage and Local Change. In Artificial Intelligence in HCI: 3rd International Conference, AI-HCI 2022, Held as Part of the 24th HCI International Conference, HCII 2022, Virtual Event, June 26–July 1, 2022, Proceedings, 448–459. Springer.
  53. 53.Zeng, Z.; Zhao, W.; Qian, P.; Zhou, Y.; Zhao, Z.; Chen, C.; and Guan, C. 2021. Robust Traffic Prediction From Spatial–Temporal Data Based on Conditional Distribution Learning. IEEE Transactions on Cybernetics, 52(12): 13458–13471.
  54. 54.Zhang, P.; Ouyang, W.; Zhang, P.; Xue, J.; and Zheng, N. 2019. SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  55. 55.Zhang, S.; Abdel-Aty, M.; Wu, Y.; and Zheng, O. 2021a. Pedestrian crossing intention prediction at red-light using pose estimation. IEEE Transactions on Intelligent Transportation Systems, 23(3): 2331–2339.
  56. 56.Zhang, Z.; Shen, D.; Tian, R.; Li, L.; Chen, Y.; Sturdevant, J.; and Cox, E. 2021b. Implementation and Performance Evaluation of Invehicle Highway Back-of-Queue Alerting System Using the Driving Simulator. In 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), 1753–1759.
  57. 57.Zhang, Z.; Tian, R.; and Duffy, V. G. 2023. Trust in Automated Vehicle: A Meta-Analysis, 221–234. Cham: Springer International Publishing. ISBN 978-3-031-10784-9.
  58. 58.Zhang, Z.; Tian, R.; Duffy, V. G.; and Li, L. 2022a. The Comfort of the Soft-Safety Driver Alerts: Measurements and Evaluation. International Journal of Human–Computer Interaction, 0(0): 1–11.
  59. 59.Zhang, Z.; Tian, R.; Elahi, F. M.; Luo, X.; Domeyer, J.; and Sherony, R. 2022b. Modeling Pedestrian Situated Intent in Dynamic Driving Scenes from the Driver’s Perspective. Available at SSRN 4281923.
  60. 60.Zhang, Z.; Tian, R.; Sherony, R.; Domeyer, J.; and Ding, Z. 2022c. Attention-Based Interrelation Modeling for Explainable Automated Driving. IEEE Transactions on Intelligent Vehicles.

Citation

MLA
Zhang, Z., et al. “TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, 2023, pp. 3534–42, https://doi.org/10.1609/AAAI.V37I3.25463.
APA
Zhang, Z., Tian, R., & Ding, Z. (2023). TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty. Proceedings of the AAAI Conference on Artificial Intelligence, 37(3), 3534–3542. https://doi.org/10.1609/AAAI.V37I3.25463
Chicago
Zhang, Z., R. Tian, and Z. Ding. 2023. “TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (3): 3534–42. https://doi.org/10.1609/AAAI.V37I3.25463.
Harvard
Zhang, Z., Tian, R. and Ding, Z. (2023) “TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(3), pp. 3534–3542. Available at: https://doi.org/10.1609/AAAI.V37I3.25463.
Vancouver
1. Zhang Z, Tian R, Ding Z (2023) TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty. Proceedings of the AAAI Conference on Artificial Intelligence 37:3534–3542

BibTeX

@article{Zhang_2023, title={TrEP: Transformer-Based Evidential Prediction for Pedestrian Intention with Uncertainty}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I3.25463}, DOI={10.1609/aaai.v37i3.25463}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhang, Zhengming and Tian, Renran and Ding, Zhengming}, year={2023}, month=June, pages={3534–3542} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF