Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
Zhiqi LiZhiding YuShiyi LanJiahan LiJan KautzTong LuJosé M. Álvarez
Reveals that existing open-loop autonomous driving benchmarks over-rely on ego vehicle status rather than visual perception, introducing a road-adherence evaluation metric and a perception-free baseline to reassess model planning quality.
The article investigates a critical flaw in current open-loop benchmarks for end-to-end autonomous driving systems, which directly map sensory inputs to vehicle trajectory planning. While recent models report leading performance on the standard nuScenes dataset, questions remain about whether these scores reflect genuine driving intelligence or reliance on superficial shortcuts.
The main objective of the article is to determine whether ego status—the vehicle's current velocity, acceleration, yaw rate, and driving commands—dominates trajectory planning at the expense of visual perception, and to evaluate whether current benchmark datasets and evaluation metrics accurately assess driving safety and quality.
The authors conducted a comprehensive experimental analysis on the nuScenes dataset by testing established state-of-the-art frameworks alongside custom baselines. They introduced Ego-MLP, an extremely simple multi-layer perceptron network that uses only the vehicle's physical status without any camera or sensor inputs. They also designed BEV-Planner, a straightforward baseline that predicts paths directly from sensory features without relying on intermediate human-labeled tasks like 3D object detection or high-definition maps. To systematically assess reliance on ego status versus perception, the researchers introduced perturbations—including weather noise, completely blanked images, and altered speed inputs—and introduced a new metric called Curb Collision Rate to track how frequently planned paths intersect road boundaries.
The analysis revealed several crucial findings. First, roughly 73.9% of the nuScenes dataset consists of simple straight-line driving, allowing models that rely purely on vehicle status to achieve competitive planning scores without processing environmental surroundings. Second, Ego-MLP matched or exceeded complex perception-based models on standard displacement error and obstacle collision metrics. Third, visual input perturbations showed that blanking camera feeds entirely caused almost no degradation in planning metrics, whereas artificially changing the input velocity severely degraded planned paths. Fourth, standard metrics failed to capture off-road hazards, but evaluating models using the new Curb Collision Rate revealed that status-only models and naive straight-driving baselines frequently drove off the road. Finally, incorporating map perception improved boundary compliance during complex turning maneuvers, which account for only 13% of the data, but worsened average scores across the dataset due to noise introduced in the dominant straight-driving scenarios.
These findings imply that current benchmark scores provide a false sense of safety and progress. Open-loop driving models are suffering from an ego status shortcut, essentially ignoring environmental vision because the dataset is dominated by straight paths. This poses severe safety risks, as a model over-reliant on current speed rather than visual cues cannot safely navigate dynamic real-world environments. Furthermore, existing optimization techniques that minimize obstacle collision rates can inadvertently push vehicles off road boundaries.
Senior leaders and developers should avoid relying exclusively on displacement error and obstacle collision rates on the nuScenes dataset to gauge autonomous driving readiness. The research community must prioritize the development of more diverse, representative datasets with balanced driving maneuvers, alongside comprehensive evaluation suites that incorporate road boundary adherence and closed-loop validation before deploying end-to-end planning models.
The primary limitation noted in the article is that open-loop evaluations cannot capture how surrounding traffic dynamically reacts to an autonomous vehicle's actions. Additionally, certain annotated road boundaries in real-world data may be technically drivable, requiring careful metric calibration. Nevertheless, the experimental evidence offers strong confidence that existing benchmark practices overstate model planning capabilities due to ego status dominance.
- Paper: Planning-oriented Autonomous Driving, Yi Hu et al. (2022). UniAD establishes the foundational paradigm and nuScenes benchmark framework for multi-task, end-to-end autonomous driving planning that the source study directly audits and critiques.
- Paper: Shortcut learning in deep neural networks, Robert Geirhos et al. (2020). This paper formulates the theoretical framework of shortcut learning in neural networks, which the source directly applies to diagnose ego status dominance in autonomous driving models.
- Paper: Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline, Penghao Wu et al. (2022). Reading this work introduces baseline methods for trajectory planning and control prediction in end-to-end driving models evaluated under open- and closed-loop settings.
- Paper: End-to-End Driving Via Conditional Imitation Learning, Felipe Codevilla et al. (2017). This seminal paper introduces conditional imitation learning using high-level ego commands, providing foundational context for how motion intents and status condition downstream planners.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). BEVFormer establishes the bird's-eye-view multi-camera perception architecture widely used by the state-of-the-art end-to-end planners evaluated in the source paper.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). This foundational paper presents the Lift-Splat-Shoot architecture that underpins bird's-eye-view representations and trajectory cost-map evaluation in vision-based driving pipelines.
- Paper: A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others, Zhiheng Li et al. (2023). This study analyzes how mitigating one shortcut can amplify others in neural networks, providing important conceptual groundwork for the source's exploration of benchmark metrics and shortcut dependencies.
- Paper: DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving, Bencheng Liao et al. (2025). DiffusionDrive advances beyond the deterministic shortcuts analyzed in the source by using truncated diffusion over trajectory anchors to generate diverse multimodal driving paths evaluated across both open-loop and closed-loop metrics.
- Paper: SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment, Katrin Renz et al. (2025). SimLingo addresses the source paper's warning about open-loop evaluation artifacts by designing an aligned vision-language framework evaluated directly in reactive closed-loop simulation environments.
- Paper: DeepAccident: A Motion and Accident Prediction Benchmark for V2X Autonomous Driving, Tianqi Wang et al. (2024). DeepAccident responds to the lack of safety-critical edge cases in standard datasets like nuScenes by establishing a multi-agent benchmark specifically tailored to motion forecasting and collision prediction.
- Paper: Evaluating the World Model Implicit in a Generative Model, Keyon Vafa et al. (2024). This work generalizes the source's critique of deceptive benchmark metrics by establishing formal criteria to test whether generative sequence models possess genuine world models or merely mimic local next-step shortcuts.
