Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?

Zhiqi LiZhiding YuShiyi LanJiahan LiJan KautzTong LuJosé M. Álvarez

article2024CVPR288 citations

Reveals that existing open-loop autonomous driving benchmarks over-rely on ego vehicle status rather than visual perception, introducing a road-adherence evaluation metric and a perception-free baseline to reassess model planning quality.

Listen

The article investigates a critical flaw in current open-loop benchmarks for end-to-end autonomous driving systems, which directly map sensory inputs to vehicle trajectory planning. While recent models report leading performance on the standard nuScenes dataset, questions remain about whether these scores reflect genuine driving intelligence or reliance on superficial shortcuts.

The main objective of the article is to determine whether ego status—the vehicle's current velocity, acceleration, yaw rate, and driving commands—dominates trajectory planning at the expense of visual perception, and to evaluate whether current benchmark datasets and evaluation metrics accurately assess driving safety and quality.

The authors conducted a comprehensive experimental analysis on the nuScenes dataset by testing established state-of-the-art frameworks alongside custom baselines. They introduced Ego-MLP, an extremely simple multi-layer perceptron network that uses only the vehicle's physical status without any camera or sensor inputs. They also designed BEV-Planner, a straightforward baseline that predicts paths directly from sensory features without relying on intermediate human-labeled tasks like 3D object detection or high-definition maps. To systematically assess reliance on ego status versus perception, the researchers introduced perturbations—including weather noise, completely blanked images, and altered speed inputs—and introduced a new metric called Curb Collision Rate to track how frequently planned paths intersect road boundaries.

The analysis revealed several crucial findings. First, roughly 73.9% of the nuScenes dataset consists of simple straight-line driving, allowing models that rely purely on vehicle status to achieve competitive planning scores without processing environmental surroundings. Second, Ego-MLP matched or exceeded complex perception-based models on standard displacement error and obstacle collision metrics. Third, visual input perturbations showed that blanking camera feeds entirely caused almost no degradation in planning metrics, whereas artificially changing the input velocity severely degraded planned paths. Fourth, standard metrics failed to capture off-road hazards, but evaluating models using the new Curb Collision Rate revealed that status-only models and naive straight-driving baselines frequently drove off the road. Finally, incorporating map perception improved boundary compliance during complex turning maneuvers, which account for only 13% of the data, but worsened average scores across the dataset due to noise introduced in the dominant straight-driving scenarios.

These findings imply that current benchmark scores provide a false sense of safety and progress. Open-loop driving models are suffering from an ego status shortcut, essentially ignoring environmental vision because the dataset is dominated by straight paths. This poses severe safety risks, as a model over-reliant on current speed rather than visual cues cannot safely navigate dynamic real-world environments. Furthermore, existing optimization techniques that minimize obstacle collision rates can inadvertently push vehicles off road boundaries.

Senior leaders and developers should avoid relying exclusively on displacement error and obstacle collision rates on the nuScenes dataset to gauge autonomous driving readiness. The research community must prioritize the development of more diverse, representative datasets with balanced driving maneuvers, alongside comprehensive evaluation suites that incorporate road boundary adherence and closed-loop validation before deploying end-to-end planning models.

The primary limitation noted in the article is that open-loop evaluations cannot capture how surrounding traffic dynamically reacts to an autonomous vehicle's actions. Additionally, certain annotated road boundaries in real-world data may be technically drivable, requiring careful metric calibration. Nevertheless, the experimental evidence offers strong confidence that existing benchmark practices overstate model planning capabilities due to ego status dominance.

Cover for Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?

Abstract

End-to-end autonomous driving recently emerged as a promising research direction to target autonomy from a full-stack perspective. Along this line, many of the latest works follow an open-loop evaluation setting on nuScenes to study the planning behavior. In this paper, we delve deeper into the problem by conducting thorough analyses and demystifying more devils in the details. We initially observed that the nuScenes dataset, characterized by relatively simple driving scenarios, leads to an under-utilization of perception information in end-to-end models incorporating ego status, such as the ego vehicle's velocity. These models tend to rely predominantly on the ego vehicle's status for future path planning. Beyond the limitations of the dataset, we also note that current metrics do not comprehensively assess the planning quality, leading to potentially biased conclusions drawn from existing benchmarks. To address this issue, we introduce a new metric to evaluate whether the predicted trajectories adhere to the road. We further propose a simple baseline able to achieve competitive results without relying on perception annotations. Given the current limitations on the benchmark and metrics, we suggest the community reassess relevant prevailing research and be cautious about whether the continued pursuit of state-of-the-art would yield convincing and universal conclusions. Code and models are available at https://github.com/NVlabs/BEV-Planner.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. BEV perception
  • 2.2. End-to-end autonomous driving
  • 3. Proposed BEV-Planner
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Metrics
  • 4.3. Discussion
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Curb Collision Rate Metric for Autonomous Path Planning

    definition

    The Curb Collision Rate (CCR) measures the safety and road adherence of autonomous vehicle trajectory predictions by calculating the proportion of predicted trajectories that intersect with road boundaries (curbs).

    To compute CCR, road boundaries from map annotations are rasterized into a grid representation at a spatial resolution of 0.1 m0.1\,\text{m}. A collision event is registered whenever the ego vehicle's planned footprint along its future trajectory overlaps with a rasterized curb grid cell. CCR serves to detect trajectories that stray outside the drivable area, revealing failure modes such as driving off the road that traditional L2L_2 distance and obstacle-only collision metrics fail to capture.

  2. Knowl 2 — Ego-MLP: Open-Loop Trajectory Prediction from Ego Status Alone

    model/method

    Ego-MLP is an open-loop motion planning baseline designed to isolate the predictive capability of ego-vehicle kinematics without using any visual or spatial perception features.

    Unlike predecessor methods (such as AD-MLP) that inadvertently introduced future label leakage by feeding ground-truth historical trajectory waypoints, Ego-MLP restricts its input strictly to instantaneous vehicle dynamic parameters: ego velocity vv, longitudinal acceleration aa, yaw angle/turning rate ω\omega, and categorical high-level driving commands cc (e.g., go straight, turn left, turn right). A multi-layer perceptron (MLP) directly maps these kinematic variables to future trajectory waypoints τ={(xt,yt)}t=1T\tau = \{(x_t, y_t)\}_{t=1}^T across horizons T∈{1 s,2 s,3 s}T \in \{1\,\text{s}, 2\,\text{s}, 3\,\text{s}\}.

  3. Knowl 3 — BEV-Planner: End-to-End Planning Baseline from BEV Features

    model/method

    BEV-Planner is a minimalist end-to-end autonomous driving architecture that generates future vehicle trajectory waypoints directly from multi-camera bird's-eye-view (BEV) features without requiring intermediate perception supervision (such as 3D bounding boxes, tracking identifiers, or HD map annotations).

    Given multi-view camera inputs, a backbone and view transformation network extract spatial BEV feature maps Bt∈RH×W×CB_t \in \mathbb{R}^{H \times W \times C} with a resolution of 128×128128 \times 128 over a 50 m50\,\text{m} perceptual range. Temporal information is incorporated by concatenating BEV features from the preceding 4 timesteps along the channel dimension without explicit geometric alignment, followed by a BEV convolutional encoder that compresses channel depth to 256 dimensions to yield the temporal feature map BB. Future trajectories τ\tau are generated by cross-attention between a learnable ego query QQ and BB, decoded via MLPs: τ=MLP(CrossAttention(q=Q,k=B,v=B))\tau = \text{MLP}(\text{CrossAttention}(q=Q, k=B, v=B)) The entire network is trained end-to-end using only an L1L_1 loss between predicted waypoints τ\tau and human driver trajectory ground truths.

  4. Knowl 4 — Open-Loop Planning Performance Benchmarking on nuScenes

    data/table

    Comparative open-loop evaluation of end-to-end planning models on the nuScenes dataset across displacement error (L2L_2), obstacle collision rate (Colli.), and Curb Collision Rate (CCR) demonstrates that kinematic shortcuts match or exceed complex full-stack perception pipelines under traditional metrics.

    Method Ego Status L2L_2 (m) ↓\downarrow Colli. (%) ↓\downarrow CCR (%) ↓\downarrow
    in BEV in Plan 1s 2s 3s 1s 2s 3s 1s 2s 3s
    ST-P3 1.59 2.64 3.73 0.69 3.62 8.39 2.53 8.17 14.4
    UniAD 0.59 1.01 1.48 0.16 0.51 1.64 0.35 1.46 3.99
    UniAD (Official) ✓ 0.35 0.63 0.99 0.16 0.43 1.27 0.21 1.32 3.63
    UniAD ✓ ✓ 0.20 0.42 0.75 0.02 0.25 0.84 0.20 1.33 3.24
    VAD-Base 0.69 1.22 1.83 0.06 0.68 2.52 1.02 3.44 7.00
    VAD-Base (Official) ✓ 0.41 0.70 1.06 0.04 0.43 1.15 0.60 2.38 5.18
    VAD-Base (Official) ✓ ✓ 0.17 0.34 0.60 0.04 0.27 0.67 0.21 2.13 5.06
    GoStraight – ✓ 0.38 0.79 1.33 0.15 0.60 2.50 2.07 8.09 15.7
    Ego-MLP – ✓ 0.15 0.32 0.59 0.00 0.27 0.85 0.27 2.52 6.60
    BEV-Planner* 0.27 0.54 0.90 0.04 0.35 1.80 0.63 3.38 7.93
    BEV-Planner 0.30 0.52 0.83 0.10 0.37 1.30 0.78 3.79 8.22
    BEV-Planner+ ✓ 0.28 0.42 0.68 0.04 0.37 1.07 0.70 3.77 8.15
    BEV-Planner++ ✓ ✓ 0.16 0.32 0.57 0.00 0.29 0.73 0.35 2.62 6.51

    Ego-MLP achieves an average L2L_2 error of 0.35 m0.35\,\text{m} and collision rate of 0.37%0.37\%, performing on par with full-stack models (UniAD at 0.46 m0.46\,\text{m} / 0.37%0.37\% and VAD-Base at 0.37 m0.37\,\text{m} / 0.33%0.33\%). However, CCR reveals that vision-free and naive linear extrapolation methods leave the drivable area at substantially higher rates (6.60%6.60\% for Ego-MLP and 15.7%15.7\% at 3 s3\,\text{s} for GoStraight vs. 3.24%3.24\% at 3 s3\,\text{s} for UniAD).

  5. Knowl 5 — Ego Status Sensitivity vs. Perceptual Robustness in Open-Loop Planning

    data/table

    Evaluating VAD-Base under systematic corruptions applied independently to sensor images and ego velocity inputs reveals that the planner's decisions are dominated by ego status while largely ignoring visual perception.

    Corruption Mode Noise / Scale L2L_2 Avg (m) ↓\downarrow Colli. Avg (%) ↓\downarrow CCR Avg (%) ↓\downarrow Det. (NDS) ↑\uparrow Map (mAP) ↑\uparrow
    Clean None 0.37 0.33 2.47 45.5 47.0
    Snow Image 0.45 0.32 2.82 36.1 29.4
    Fog Image 0.45 0.30 2.78 34.3 29.4
    Glare Image 0.44 0.26 2.63 41.7 38.3
    Rain Image 0.45 0.29 2.89 29.1 13.0
    Blank Image 0.46 0.54 3.71 0.0 0.0
    Velocity Perturbation v×0.0v \times 0.0 6.16 7.98 1.23 45.5 47.0
    Velocity Perturbation v×0.5v \times 0.5 3.19 1.71 2.83 45.5 47.0
    Velocity Perturbation v×1.5v \times 1.5 3.20 3.08 7.01 45.5 47.0
    Velocity Perturbation v=100 m/sv = 100\,\text{m/s} 208.0 9.38 27.0 45.5 47.0

    When complete image degradation (Blank) causes 3D detection (NDS) and map segmentation (mAP) to collapse to 0.00.0, the average L2L_2 planning error increases by only 0.09 m0.09\,\text{m} (0.37 m→0.46 m0.37\,\text{m} \to 0.46\,\text{m}). Conversely, perturbing ego velocity without altering the camera inputs causes catastrophic planning failure (e.g., v×0.0v \times 0.0 degrades L2L_2 to 6.16 m6.16\,\text{m} and collision rate to 7.98%7.98\%; v=100 m/sv = 100\,\text{m/s} produces an average L2L_2 error of 208.0 m208.0\,\text{m}).

  6. Knowl 6 — Trade-off of Map Perception in Straight Versus Turning Maneuvers

    data/table

    Integrating HD map perception into BEV-Planner produces conflicting effects: it degrades overall average L2L_2 distance and collision metrics due to the overwhelming dominance of straight-driving data, but significantly improves safety during complex turning maneuvers.

    Method L2L_2-ST Avg (m) ↓\downarrow L2L_2-LR Avg (m) ↓\downarrow Colli-ST Avg (%) ↓\downarrow Colli-LR Avg (%) ↓\downarrow CCR Avg (%) ↓\downarrow
    BEV-Planner 0.48 0.81 0.40 2.25 4.26
    BEV-Planner+Map 0.97 0.89 0.91 0.78 2.60

    Here, -ST denotes evaluation restricted to 'go straight' commands, while -LR denotes evaluation restricted to left/right turning commands. Incorporating map perception (BEV-Planner+Map) reduces the turning collision rate from 2.25%2.25\% to 0.78%0.78\% and improves overall Curb Collision Rate (CCR) from 4.26%4.26\% to 2.60%2.60\%. However, in straight driving scenarios (which constitute 73.9%73.9\% to 87%87\% of the benchmark), the map learning objective introduces interference, increasing straight L2L_2 error (0.48 m→0.97 m0.48\,\text{m} \to 0.97\,\text{m}) and straight collision rate (0.40%→0.91%0.40\% \to 0.91\%).

  7. Knowl 7 — Inflation of Road Departure Risks by Post-Processing Collision Optimization

    data/table

    Trajectory post-processing optimization modules designed for obstacle avoidance can severely compromise road boundary compliance.

    Method Post-Processing (P.P.) Avg. L2L_2 (m) ↓\downarrow Avg. Colli. (%) ↓\downarrow Avg. CCR (%) ↓\downarrow
    UniAD ✓ 0.77 0.51 7.83
    UniAD 0.66 0.62 1.72

    UniAD employs a non-linear optimization post-processing stage to nudge predicted trajectory points away from occupancy grids, successfully lowering the obstacle collision rate from 0.62%0.62\% to 0.51%0.51\%. However, because this optimization lacks map prior constraints such as curb locations, it inadvertently forces trajectories across road boundaries, increasing the Curb Collision Rate (CCR) by more than 4.5×4.5\times (1.72%→7.83%1.72\% \to 7.83\%).

  8. Knowl 8 — Feature Attention Degeneration in Planning-Guided BEV Encoders

    empirical result

    When ego status (velocity, acceleration, yaw rate) is explicitly concatenated into the ego query within the planning module (as in BEV-Planner++), the model converges significantly faster during training on trajectory loss, but its spatial feature representation suffers severe degeneration.

    Visualization of cross-attention spatial activations over the BEV plane shows that BEV-Planner (without ego status) focuses attention predominantly on the forward roadway ahead of the ego vehicle. In contrast, BEV-Planner++ concentrates its attention in the immediate vicinity or behind the ego vehicle. The presence of explicit kinematic inputs removes the gradient incentive for the query to attend to forward visual representations, causing the network to shortcut trajectory regression via ego status alone.

  9. Knowl 9 — Driving Trajectory Distribution Imbalance in nuScenes

    empirical result

    Analysis of the nuScenes autonomous driving dataset indicates that 73.9%73.9\% of all driving trajectories consist of straightforward driving scenarios, with turning maneuvers (left and right turns) accounting for only 13%13\% of the dataset.

    Because the vast majority of scenarios require only maintaining current velocity and heading, open-loop planning evaluations are biased toward models that exploit kinematic continuity. This distribution imbalance enables simple kinematic extrapolation to achieve near-state-of-the-art results on aggregate L2L_2 and collision metrics while penalizing models that learn multi-task perceptual representations.

Coverage note — None was omitted. All principal contributions—including the empirical diagnosis of ego status shortcutting, the CCR metric, Ego-MLP and BEV-Planner architectures, corruption benchmarks, straight vs. turning analyses, and post-processing evaluation—are represented.

References

  1. 1.Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al. End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316, 2016. 1
  2. 2.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020. 2
  3. 3.Sergio Casas, Abbas Sadat, and Raquel Urtasun. Mp3: A unified model to map, perceive, predict and plan. In CVPR, 2021. 3
  4. 4.Li Chen, Chonghao Sima, Yang Li, Zehan Zheng, Jiajie Xu, Xiangwei Geng, Hongyang Li, Conghui He, Jianping Shi, Yu Qiao, et al. Persformer: 3d lane detection via perspective transformer and the openlane benchmark. In European Conference on Computer Vision, pages 550–567. Springer, 2022. 3
  5. 5.Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li. End-to-end autonomous driving: Challenges and frontiers. arXiv preprint arXiv:2306.16927, 2023. 1
  6. 6.Laurene Claussmann, Marc Revilloud, Dominique Gruyer, and Sébastien Glaser. A review of motion planning for highway autonomous driving. IEEE Transactions on Intelligent Transportation Systems, 21(5):1826–1848, 2019. 1
  7. 7.Dmitri Dolgov, Sebastian Thrun, Michael Montemerlo, and James Diebel. Practical search techniques in path planning for autonomous driving. Ann Arbor, 1001(48105):18–80, 2008. 1
  8. 8.Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. Carla: An open urban driving simulator. 2017. 1, 3
  9. 9.Junru Gu, Chenxu Hu, Tianyuan Zhang, Xuanyao Chen, Yilun Wang, Yue Wang, and Hang Zhao. Vip3d: End-to-end visual trajectory prediction via 3d agent queries. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5496–5506, 2023. 3
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 3
  11. 11.Peiyun Hu, Aaron Huang, John Dolan, David Held, and Deva Ramanan. Safe local motion planning with selfsupervised freespace forecasting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12732–12741, 2021. 2
  12. 12.Shengchao Hu, Li Chen, Penghao Wu, Hongyang Li, Junchi Yan, and Dacheng Tao. St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning. In European Conference on Computer Vision, pages 533–549. Springer, 2022. 2, 3
  13. 13.Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al. Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17853–17862, 2023. 1, 2, 3, 5, 8
  14. 14.Junjie Huang and Guan Huang. BEVDet4D: Exploit temporal cues in multi-camera 3d object detection. arXiv preprint arXiv:2203.17054, 2022. 3, 4
  15. 15.Linyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang, Jingdong Wang, Yu Qiao, and Hongyang Li. Leveraging vision-centric multi-modal expertise for 3d object detection. arXiv preprint arXiv:2310.15670, 2023. 3
  16. 16.Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Vad: Vectorized scene representation for efficient autonomous driving. arXiv preprint arXiv:2303.12077, 2023. 2, 3, 6
  17. 17.Tarasha Khurana, Peiyun Hu, Achal Dave, Jason Ziglar, David Held, and Deva Ramanan. Differentiable raycasting for self-supervised occupancy forecasting. In European Conference on Computer Vision, pages 353–369. Springer, 2022. 2
  18. 18.Hongyang Li, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Enze Xie, Zhiqi Li, Hanming Deng, Hao Tian, et al. Delving into the devils of bird’s-eye-view perception: A review, evaluation and recipe. arXiv preprint arXiv:2209.05324, 2022. 3
  19. 19.Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning. IEEE transactions on pattern analysis and machine intelligence, 45(3):3461–3475, 2022. 1, 3
  20. 20.Qi Li, Yue Wang, Yilun Wang, and Hang Zhao. Hdmapnet: An online hd map construction and evaluation framework. 2022. 3
  21. 21.Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran Wang, Yukang Shi, Jianjian Sun, and Zeming Li. BEVDepth: Acquisition of reliable depth for multi-view 3d object detection. arXiv preprint arXiv:2206.10092, 2022. 3
  22. 22.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv preprint arXiv:2203.17270, 2022. 3
  23. 23.Zhiqi Li, Zhiding Yu, David Austin, Mingsheng Fang, Shiyi Lan, Jan Kautz, and Jose M Alvarez. Fb-occ: 3d occupancy prediction based on forward-backward view transformation. arXiv preprint arXiv:2307.01492, 2023. 3
  24. 24.Zhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu, and Jose M Alvarez. Fb-bev: Bev representation from forward-backward view transformations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6919–6928, 2023. 3
  25. 25.Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang, Wenyu Liu, and Chang Huang. Maptr: Structured modeling and learning for online vectorized hd map construction. arXiv preprint arXiv:2208.14437, 2022. 3
  26. 26.Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and Zhizhong Su. Sparse4d v2: Recurrent temporal fusion with sparse model. arXiv preprint arXiv:2305.14018, 2023. 3
  27. 27.Haisong Liu, Yao Teng, Tao Lu, Haiguang Wang, and Limin Wang. Sparsebev: High-performance sparse 3d object detection from multi-camera videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18580–18590, 2023.
  28. 28.Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun. PETRv2: A unified framework for 3d perception from multi-camera images. arXiv preprint arXiv:2206.01256, 2022. 3
  29. 29.Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang, and Hang Zhao. Vectormapnet: End-to-end vectorized hd map learning. In International Conference on Machine Learning, pages 22352–22369. PMLR, 2023. 3
  30. 30.Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. Time will tell: New outlooks and a baseline for temporal multi-view 3d object detection. arXiv preprint arXiv:2210.02443, 2022. 3
  31. 31.Jonah Philion and Sanja Fidler. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In ECCV, 2020. 3
  32. 32.Aditya Prakash, Kashyap Chitta, and Andreas Geiger. Multimodal fusion transformer for end-to-end autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7077–7087, 2021. 1
  33. 33.Abbas Sadat, Sergio Casas, Mengye Ren, Xinyu Wu, Pranaab Dhawan, and Raquel Urtasun. Perceive, predict, and plan: Safe motion planning through interpretable semantic representations. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pages 414–430. Springer, 2020. 3
  34. 34.Hao Shao, Letian Wang, Ruobing Chen, Hongsheng Li, and Yu Liu. Safety-enhanced autonomous driving using interpretable sensor fusion transformer. In Conference on Robot Learning, pages 726–737. PMLR, 2023. 3
  35. 35.Ardi Tampuu, Tambet Matiisen, Maksym Semikin, Dmytro Fishman, and Naveed Muhammad. A survey of end-to-end driving: Architectures and training methods. IEEE Transactions on Neural Networks and Learning Systems, 33(4): 1364–1384, 2020. 1
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 3
  37. 37.Li Wang, Xinyu Zhang, Baowei Xv, Jinzhao Zhang, Rong Fu, Xiaoyu Wang, Lei Zhu, Haibing Ren, Pingping Lu, Jun Li, et al. Interfusion: Interaction-based 4d radar and lidar fusion for 3d object detection. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 12247–12253. IEEE, 2022. 3
  38. 38.Ruihao Wang, Jian Qin, Kaiying Li, Yaochen Li, Dong Cao, and Jintao Xu. Bev-lanedet: An efficient 3d lane detection based on virtual camera via key-points. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1002–1011, 2023. 3
  39. 39.Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xiangyu Zhang. Exploring object-centric temporal modeling for efficient multi-view 3d object detection. arXiv preprint arXiv:2303.11926, 2023. 3
  40. 40.Penghao Wu, Xiaosong Jia, Li Chen, Junchi Yan, Hongyang Li, and Yu Qiao. Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline. Advances in Neural Information Processing Systems, 35:6119–6132, 2022. 3
  41. 41.Enze Xie, Zhiding Yu, Daquan Zhou, Jonah Philion, Anima Anandkumar, Sanja Fidler, Ping Luo, and Jose M Alvarez. M2BEV: Multi-camera joint 3d detection and segmentation with unified birds-eye view representation. arXiv preprint arXiv:2204.05088, 2022. 3
  42. 42.Chenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang, Gao Huang, Hongyang Li, Yu Qiao, Lewei Lu, et al. Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17830–17839, 2023. 3
  43. 43.Tengju Ye, Wei Jing, Chunyong Hu, Shikun Huang, Lingping Gao, Fangzhen Li, Jingke Wang, Ke Guo, Wencong Xiao, Weibo Mao, et al. Fusionad: Multi-modality fusion for prediction and planning tasks of autonomous driving. arXiv preprint arXiv:2308.01006, 2023. 2, 3, 6
  44. 44.Tianyuan Yuan, Yicheng Liu, Yue Wang, Yilun Wang, and Hang Zhao. Streammapnet: Streaming mapping network for vectorized online hd map construction. arXiv preprint arXiv:2308.12570, 2023. 3
  45. 45.Jiang-Tian Zhai, Ze Feng, Jinhao Du, Yongqiang Mao, Jiang-Jiang Liu, Zichang Tan, Yifu Zhang, Xiaoqing Ye, and Jingdong Wang. Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes. arXiv preprint arXiv:2305.10430, 2023. 2, 3, 4

Citation

MLA
Li, Z., et al. “Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?”. arXiv, 2023, http://arxiv.org/abs/2312.03031v2.
APA
Li, Z., Yu, Z., Lan, S., Li, J., Kautz, J., Lu, T., & Alvarez, J. M. (2023). Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?. arXiv. http://arxiv.org/abs/2312.03031v2
Chicago
Li, Z., Z. Yu, S. Lan, et al. 2023. “Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?”. arXiv. http://arxiv.org/abs/2312.03031v2.
Harvard
Li, Z. et al. (2023) “Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.03031v2.
Vancouver
1. Li Z, Yu Z, Lan S, Li J, Kautz J, Lu T, Alvarez JM (2023) Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?. arXiv

BibTeX

@article{li2023ego,
  title = {Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?},
  author = {Li, Zhiqi and Yu, Zhiding and Lan, Shiyi and Li, Jiahan and Kautz, Jan and Lu, Tong and Alvarez, Jose M.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.03031v2},
  eprint = {2312.03031}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE