DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
Bencheng LiaoShaoyu ChenHaoran YinBo JiangCheng WangSixu YanXinbang ZhangXiangyu LiYing ZhangQian Zhang
Proposes a truncated diffusion policy anchored on prior driving patterns that reduces denoising to two steps, enabling diverse multi-mode trajectory generation at 45 FPS and setting a new performance record on the NAVSIM benchmark.
End-to-end autonomous driving systems aim to generate vehicle control and motion plans directly from raw sensor inputs, offering a scalable alternative to traditional hand-crafted rules. However, existing approaches face a fundamental trade-off. Single-trajectory methods fail to capture multiple valid driving decisions in complex traffic, while systems using large fixed libraries of thousands of trajectories struggle with unmodeled situations and computational bottlenecks. While diffusion models—a class of generative models that iteratively remove noise from data—excel at producing diverse actions in robotics, directly applying them to vehicles leads to overlapping, repetitive trajectories and slow multi-step denoising that cannot operate in real time.
The article introduces and evaluates DiffusionDrive, a motion-planning framework that adapts diffusion models for real-time autonomous driving by combining prior driving patterns with a streamlined denoising schedule.
The approach introduces a truncated diffusion policy alongside an efficient cascaded transformer decoder. Instead of starting from completely random noise across 20 iterative steps, the model initializes generation around a small set of 20 learned trajectory anchors with bounded noise, reducing the required denoising process to just two steps. The decoder uses sparse spatial cross-attention to interact directly with Bird's Eye View map representations and surrounding obstacles. The framework was evaluated on the benchmark NAVSIM driving dataset against closed-loop planning metrics and tested for open-loop precision on the nuScenes dataset using standard computing hardware.
The findings show that DiffusionDrive achieved a record planning score of 88.1 on NAVSIM, outperforming leading single-mode and multi-mode systems, including challenge-winning models that relied on thousands of anchors and complex post-processing. It reduced the required diffusion steps tenfold from 20 to 2, achieving a processing speed of 45 frames per second on a single high-end processor. The framework also improved trajectory diversity by 64% over standard diffusion baselines, successfully generating high-quality alternate actions such as lane changes and collision-avoidance maneuvers. Additionally, on the nuScenes benchmark, the framework delivered a 20.8% lower position error and a 63.6% lower collision rate compared to previous vectorized baselines while running 1.8 times faster.
These results demonstrate that autonomous planners can generate diverse, high-quality driving decisions without relying on oversized trajectory libraries or sacrificing inference speed. By slashing computational latency to millisecond levels while improving safety and comfort metrics, the framework makes generative planning commercially viable for real-time onboard vehicle deployment, lowering hardware compute requirements and expanding operational safety margins in dynamic environments.
Organizations developing autonomous driving stacks should consider adopting truncated diffusion policies to replace rigid anchor tables and deterministic planning heads. Engineering teams should investigate integrating this cascaded decoder with existing upstream vision and map perception modules. Because evaluation primarily relied on standardized simulation benchmarks and public datasets, closed-loop physical road testing and validation in extreme edge cases remain necessary next steps before production deployment.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). Establishes the foundational formulation of representing sequential action and trajectory generation as a conditional denoising diffusion policy, which DiffusionDrive adapts and accelerates for autonomous vehicle motion planning.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Introduces trajectory-level diffusion modeling for decision-making and control, providing the conceptual basis for applying diffusion processes directly to multi-step planning.
- Paper: Planning-oriented Autonomous Driving, Yi Hu et al. (2022). Presents the UniAD paradigm for end-to-end planning-oriented autonomous driving using bird's-eye-view representations and transformer queries that modern end-to-end planners build upon.
- Paper: Efficient Diffusion Policies For Offline Reinforcement Learning, Bingyi Kang et al. (2023). Addresses the computational latency bottleneck of diffusion-based policies by demonstrating fast sampling and single-pass action approximations essential for real-time control.
- Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation, Zhijian Liu et al. (2022). Details the unified bird's-eye-view (BEV) multi-sensor representation pipeline that feeds into the spatial cross-attention mechanisms used by downstream planners like DiffusionDrive.
- Paper: Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline, Penghao Wu et al. (2022). Analyzes the trade-offs between trajectory planning and control prediction in end-to-end driving architectures, contextualizing the need for high-quality multi-modal trajectory generation.
- Paper: Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?, Zhiqi Li et al. (2024). Critically examines open-loop planning shortcuts and perception utilization on the nuScenes benchmark, establishing the motivation for robust closed-loop evaluation on NAVSIM.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Introduces the standard denoising diffusion probabilistic model framework whose multi-step denoising schedule DiffusionDrive truncates for rapid inference.
- Paper: Generative Modeling via Drifting, Mingyang Deng et al. (2026). Extends the push toward ultra-fast generative policies by introducing a drifting field paradigm that eliminates multi-step sampling altogether to enable one-step inference in robotics and vision.
