DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Bencheng LiaoShaoyu ChenHaoran YinBo JiangCheng WangSixu YanXinbang ZhangXiangyu LiYing ZhangQian Zhang

article2025CVPR415 citations

Proposes a truncated diffusion policy anchored on prior driving patterns that reduces denoising to two steps, enabling diverse multi-mode trajectory generation at 45 FPS and setting a new performance record on the NAVSIM benchmark.

Listen

End-to-end autonomous driving systems aim to generate vehicle control and motion plans directly from raw sensor inputs, offering a scalable alternative to traditional hand-crafted rules. However, existing approaches face a fundamental trade-off. Single-trajectory methods fail to capture multiple valid driving decisions in complex traffic, while systems using large fixed libraries of thousands of trajectories struggle with unmodeled situations and computational bottlenecks. While diffusion models—a class of generative models that iteratively remove noise from data—excel at producing diverse actions in robotics, directly applying them to vehicles leads to overlapping, repetitive trajectories and slow multi-step denoising that cannot operate in real time.

The article introduces and evaluates DiffusionDrive, a motion-planning framework that adapts diffusion models for real-time autonomous driving by combining prior driving patterns with a streamlined denoising schedule.

The approach introduces a truncated diffusion policy alongside an efficient cascaded transformer decoder. Instead of starting from completely random noise across 20 iterative steps, the model initializes generation around a small set of 20 learned trajectory anchors with bounded noise, reducing the required denoising process to just two steps. The decoder uses sparse spatial cross-attention to interact directly with Bird's Eye View map representations and surrounding obstacles. The framework was evaluated on the benchmark NAVSIM driving dataset against closed-loop planning metrics and tested for open-loop precision on the nuScenes dataset using standard computing hardware.

The findings show that DiffusionDrive achieved a record planning score of 88.1 on NAVSIM, outperforming leading single-mode and multi-mode systems, including challenge-winning models that relied on thousands of anchors and complex post-processing. It reduced the required diffusion steps tenfold from 20 to 2, achieving a processing speed of 45 frames per second on a single high-end processor. The framework also improved trajectory diversity by 64% over standard diffusion baselines, successfully generating high-quality alternate actions such as lane changes and collision-avoidance maneuvers. Additionally, on the nuScenes benchmark, the framework delivered a 20.8% lower position error and a 63.6% lower collision rate compared to previous vectorized baselines while running 1.8 times faster.

These results demonstrate that autonomous planners can generate diverse, high-quality driving decisions without relying on oversized trajectory libraries or sacrificing inference speed. By slashing computational latency to millisecond levels while improving safety and comfort metrics, the framework makes generative planning commercially viable for real-time onboard vehicle deployment, lowering hardware compute requirements and expanding operational safety margins in dynamic environments.

Organizations developing autonomous driving stacks should consider adopting truncated diffusion policies to replace rigid anchor tables and deterministic planning heads. Engineering teams should investigate integrating this cascaded decoder with existing upstream vision and map perception modules. Because evaluation primarily relied on standardized simulation benchmarks and public datasets, closed-loop physical road testing and validation in extreme edge cases remain necessary next steps before production deployment.

  • Paper: Generative Modeling via Drifting, Mingyang Deng et al. (2026). Extends the push toward ultra-fast generative policies by introducing a drifting field paradigm that eliminates multi-step sampling altogether to enable one-step inference in robotics and vision.
Cover for DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Abstract

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising direction. However, the numerous denoising steps in the robotic diffusion policy and the more dynamic, open-world nature of traffic scenes pose substantial challenges for generating diverse driving actions at a real-time speed. To address these challenges, we propose a novel truncated diffusion policy that incorporates prior multi-mode anchors and truncates the diffusion schedule, enabling the model to learn denoising from anchored Gaussian distribution to the multi-mode driving action distribution. Additionally, we design an efficient cascade diffusion decoder for enhanced interaction with conditional scene context. The proposed model, DiffusionDrive, demonstrates 10× reduction in denoising steps compared to vanilla diffusion policy, delivering superior diversity and quality in just 2 steps. On the planning-oriented NAVSIM dataset, with aligned ResNet-34 backbone, DiffusionDrive achieves 88.1 PDMS without bells and whistles, setting a new record, while running at a real-time speed of 45 FPS on an NVIDIA 4090. Qualitative results on challenging scenarios further confirm that DiffusionDrive can robustly generate diverse plausible driving actions.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminary
  • 3.2. Investigation
  • 3.3. Truncated Diffusion
  • 3.4. Architecture
  • 4. Experiment
  • 4.1. Dataset
  • 4.2. Implementation Detail
  • 4.3. Quantitative Comparison
  • 4.4. Roadmap
  • 4.5. Ablation Study
  • 4.6. Qualitative Comparison
  • 4.7. Quantitative Comparison on nuScenes dataset
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Truncated Diffusion Policy for Trajectory Planning

    model/method

    In end-to-end autonomous driving, policy models take onboard sensor data (such as camera images and LiDAR point clouds) to predict an ego-vehicle future trajectory τ={(xt,yt)}t=1Tf\tau = \{(x_t, y_t)\}_{t=1}^{T_f} over a planning horizon TfT_f.

    While vanilla robotic diffusion policies denoise from standard Gaussian noise N(0,I)\mathcal{N}(0, I) over many denoising steps (e.g., 20 steps) and frequently suffer from mode collapse where distinct noise samples converge to redundant paths, the truncated diffusion policy initiates generation from an anchored Gaussian distribution. A set of NanchorN_{\text{anchor}} trajectory anchors {ak}k=1Nanchor\{a_k\}_{k=1}^{N_{\text{anchor}}} is derived by applying K-Means clustering to ground-truth trajectories in the training dataset, where each anchor is ak={(xt,yt)}t=1Tfa_k = \{(x_t, y_t)\}_{t=1}^{T_f}.

    The forward diffusion schedule is truncated so that only a fractional portion of Gaussian noise is added to the anchors rather than diffusing them to pure isotropic noise:

    τki=αˉiak+1−αˉiϵ,ϵ∼N(0,I),i∈[1,Ttrunc]\tau_k^i = \sqrt{\bar{\alpha}_i} a_k + \sqrt{1 - \bar{\alpha}_i} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I), \quad i \in [1, T_{\text{trunc}}]

    where ii denotes the diffusion timestep, αˉi=∏s=1i(1−βs)\bar{\alpha}_i = \prod_{s=1}^i (1 - \beta_s) is determined by the noise schedule βs\beta_s, and Ttrunc≪TT_{\text{trunc}} \ll T represents the truncated horizon (e.g., Ttrunc=50T_{\text{trunc}} = 50 out of T=1000T = 1000 total diffusion steps). By diffusing within sub-Gaussian distributions localized around empirical driving anchors, the policy models multi-mode action distributions with as few as 2 denoising steps and avoids the need for massive fixed vocabularies.

  2. Knowl 2 — Training Objective for Truncated Diffusion Trajectory Planning

    equation

    During training of the truncated diffusion policy, the diffusion decoder fθf_\theta receives NanchorN_{\text{anchor}} noisy trajectories {τki}k=1Nanchor\{\tau_k^i\}_{k=1}^{N_{\text{anchor}}} along with conditional scene context zz, and outputs trajectory coordinate predictions {τ^k}k=1Nanchor\{\hat{\tau}_k\}_{k=1}^{N_{\text{anchor}}} together with classification confidence scores {s^k}k=1Nanchor\{\hat{s}_k\}_{k=1}^{N_{\text{anchor}}}:

    {s^k,τ^k}k=1Nanchor=fθ({τki}k=1Nanchor,z)\{\hat{s}_k, \hat{\tau}_k\}_{k=1}^{N_{\text{anchor}}} = f_\theta(\{\tau_k^i\}_{k=1}^{N_{\text{anchor}}}, z)

    Positive and negative sample labels are assigned based on spatial proximity to the ground-truth demonstration τgt\tau_{\text{gt}}: the noisy sample corresponding to the anchor closest to τgt\tau_{\text{gt}} is designated as the positive sample (yk=1y_k = 1), while all other anchor samples are designated as negative samples (yk=0y_k = 0).

    The overall model is trained by minimizing the composite loss:

    L=∑k=1Nanchor[ykLrec(τ^k,τgt)+λ BCE(s^k,yk)]\mathcal{L} = \sum_{k=1}^{N_{\text{anchor}}} \left[ y_k \mathcal{L}_{\text{rec}}(\hat{\tau}_k, \tau_{\text{gt}}) + \lambda \, \text{BCE}(\hat{s}_k, y_k) \right]

    where Lrec(τ^k,τgt)\mathcal{L}_{\text{rec}}(\hat{\tau}_k, \tau_{\text{gt}}) is the L1L_1 coordinate reconstruction loss between the predicted waypoints and ground-truth waypoints, BCE(s^k,yk)\text{BCE}(\hat{s}_k, y_k) is the binary cross-entropy classification loss, and λ\lambda is a balancing hyperparameter.

  3. Knowl 3 — Cascade Transformer Diffusion Decoder Architecture

    model/method

    The diffusion decoder iteratively refines sampled noisy trajectories through structured interactions with multimodal scene context. For an input set of noisy trajectories {τki}\{\tau_k^i\}, each decoder layer executes the following sequence:

    1. Spatial Deformable Cross-Attention: Samples feature representations from Bird's Eye View (BEV) or Perspective View (PV) feature maps at locations dynamically determined by the trajectory waypoint coordinates.
    2. Perception Query Cross-Attention: Trajectory tokens perform cross-attention with structured agent and map queries produced by an upstream perception backbone to incorporate high-level dynamic obstacle and road topology context.
    3. Feed-Forward Network (FFN): Standard dense projection and activation layer.
    4. Timestep Modulation: Modulates trajectory features with sinusoidal embeddings of the diffusion timestep ii.
    5. Prediction Head: A multi-layer perceptron (MLP) predicts a classification confidence score s^k\hat{s}_k and offset coordinates relative to the input noisy coordinates to yield updated coordinates τ^k\hat{\tau}_k.

    Multiple decoder layers (e.g., 2 cascade stages) are stacked sequentially, with the refined trajectory outputs of one stage serving as the input anchors for the subsequent stage. The parameters of this cascade decoder are shared across all reverse denoising timesteps.

  4. Knowl 4 — Inference and Dynamic Sampling in Truncated Diffusion Planning

    algorithm

    The truncated diffusion planner allows an arbitrary number of trajectory candidates NinferN_{\text{infer}} to be evaluated at test time, independent of the training anchor set size NanchorN_{\text{anchor}}. The denoising process begins from anchored Gaussian distributions and applies a small number of Denoising Diffusion Implicit Models (DDIM) steps before selecting the trajectory with the highest predicted score.

    Input: Prior anchor trajectories {a_k} for k in 1 to N_infer, conditional scene context z, truncated noise schedule parameters alpha_bar, number of denoising steps S (e.g., S = 2)
    Output: Executable ego-vehicle trajectory tau_star
    for k = 1 to N_infer do
        Sample epsilon_k ~ Normal(0, I)
        tau_k_current = sqrt(alpha_bar_Ttrunc) * a_k + sqrt(1 - alpha_bar_Ttrunc) * epsilon_k
    end for
    for step = S down to 1 do
        timestep_i = map_denoising_step_to_timestep(step)
        for stage = 1 to num_cascade_stages do
            {s_hat_k, tau_hat_k} = DecoderStage({tau_k_current}, z, timestep_i)
            tau_k_current = tau_hat_k
        end for
        if step > 1 then
            timestep_next = map_denoising_step_to_timestep(step - 1)
            for k = 1 to N_infer do
                tau_k_current = DDIM_Update(tau_k_current, tau_hat_k[k], timestep_i, timestep_next)
            end for
        end if
    end for
    best_idx = argmax(s_hat_k over all k in 1 to N_infer)
    tau_star = tau_hat_k[best_idx]
    return tau_star
  5. Knowl 5 — Mode Diversity Metric for Planned Trajectory Ensembles

    equation

    To quantitatively evaluate the diversity of generated multi-mode trajectory distributions and detect mode collapse in generative driving policies, the mode diversity score DD is computed as one minus the mean spatial Intersection over Union (mIoU) between each predicted trajectory and the union of all predicted trajectories:

    D=1−1N∑i=1NArea(τi∩⋃j=1Nτj)Area(τi∪⋃j=1Nτj)D = 1 - \frac{1}{N} \sum_{i=1}^N \frac{\text{Area}\left(\tau_i \cap \bigcup_{j=1}^N \tau_j\right)}{\text{Area}\left(\tau_i \cup \bigcup_{j=1}^N \tau_j\right)}

    where τi\tau_i represents the ii-th denoised trajectory, NN is the total number of sampled trajectories in the ensemble, and ⋃j=1Nτj\bigcup_{j=1}^N \tau_j is the geometric union of all trajectories in the ego coordinate plane. A score of D=0D = 0 indicates complete mode collapse (all samples perfectly overlap), whereas higher values of DD indicate greater spatial dispersion and diversity among distinct driving hypotheses.

  6. Knowl 6 — Progression and Computational Efficiency from Vanilla Diffusion to Truncated DiffusionDrive

    data/table

    The table traces the architectural roadmap from the baseline Transfuser model to DiffusionDrive on the NAVSIM navtest split, evaluating predictive driving model score (PDMS), trajectory diversity score (DD), parameter counts, and latency per step on an NVIDIA RTX 4090 GPU.

    Method NC ↑\uparrow DAC ↑\uparrow TTC ↑\uparrow Comf. ↑\uparrow EP ↑\uparrow PDMS ↑\uparrow Arch. Step Time ↓\downarrow Steps ↓\downarrow Total Time ↓\downarrow D↑D \uparrow Para. ↓\downarrow FPS ↑\uparrow
    Transfuser 97.7 92.8 92.8 100 79.2 84.0 MLP 0.2ms 1 0.2ms 0% 56M 60
    TransfuserDP 97.5 93.7 92.7 100 79.4 84.6 UNet 6.5ms 20 130.0ms 11% 101M 7
    TransfuserTD 97.9 94.2 93.9 100 80.2 85.7 UNet 6.9ms 2 13.8ms 70% 102M 27
    DiffusionDrive 98.2 96.2 94.7 100 82.2 88.1 Dec. 3.8ms 2 7.6ms 74% 60M 45

    Directly converting Transfuser into a vanilla DDIM diffusion policy (TransfuserDP) increases PDMS by +0.6 and diversity DD to 11%, but introduces a 650×\times inference overhead (reducing FPS from 60 to 7). Applying the truncated diffusion policy (TransfuserTD) reduces denoising steps from 20 to 2, increasing diversity to 70% and boosting FPS to 27. Integrating the cascade diffusion decoder in DiffusionDrive achieves 88.1 PDMS, 74% diversity, and 45 FPS with only 60M parameters.

  7. Knowl 7 — Closed-Loop Planning Performance on NAVSIM Benchmark

    data/table

    The table compares DiffusionDrive against state-of-the-art end-to-end autonomous driving models on the NAVSIM navtest split under closed-loop evaluation metrics using an aligned ResNet-34 image backbone.

    Method Input Anchor NC ↑\uparrow DAC ↑\uparrow TTC ↑\uparrow Comf. ↑\uparrow EP ↑\uparrow PDMS ↑\uparrow
    UniAD Camera 0 97.8 91.9 92.9 100 78.8 83.4
    PARA-Drive Camera 0 97.9 92.4 93.0 99.8 79.3 84.0
    LTF Camera 0 97.4 92.8 92.4 100 79.0 83.8
    Transfuser C L 0 97.7 92.8 92.8 100 79.2 84.0
    DRAMA C L 0 98.0 93.1 94.8 100 80.1 85.5
    VADv2-V8192 C L 8192 97.2 89.1 91.6 100 76.0 80.9
    Hydra-MDP-V8192 C L 8192 97.9 91.7 92.9 100 77.6 83.0
    Hydra-MDP-V8192-W-EP C L 8192 98.3 96.0 94.6 100 78.7 86.5
    DiffusionDrive C L 20 98.2 96.2 94.7 100 82.2 88.1

    Here, 'C & L' denotes camera and LiDAR sensor inputs. Sub-metrics comprise No At-Fault Collisions (NC), Drivable Area Compliance (DAC), Time-to-Collision (TTC), Comfort (Comf.), Ego Progress (EP), and Predictive Driver Model Score (PDMS). DiffusionDrive achieves 88.1 PDMS with 20 anchors, surpassing VADv2 by 7.2 PDMS while reducing anchor count by 400×\times, and exceeding Hydra-MDP-V8192-W-EP by 1.6 PDMS without requiring rule-based distillation or confidence score post-processing.

  8. Knowl 8 — Open-Loop Planning Evaluation on nuScenes Dataset

    data/table

    The table compares the open-loop motion planning performance and inference runtime of DiffusionDrive (built on SparseDrive with a ResNet-50 backbone and 18 anchors) against prior methods on the nuScenes dataset across 1s, 2s, 3s horizons, and their average.

    Method L2 Error (m) ↓\downarrow Collision Rate (%) ↓\downarrow FPS ↑\uparrow
    1s 2s 3s Avg. 1s 2s 3s Avg.
    ST-P3 1.33 2.11 2.90 2.11 0.23 0.62 1.27 0.71 1.6
    UniAD 0.45 0.70 1.04 0.73 0.62 0.58 0.63 0.61 1.8
    OccNet 1.29 2.13 2.99 2.14 0.21 0.59 1.37 0.72 2.6
    VAD 0.41 0.70 1.05 0.72 0.07 0.17 0.41 0.22 4.5
    SparseDrive 0.29 0.58 0.96 0.61 0.01 0.05 0.18 0.08 9.0
    DiffusionDrive 0.27 0.54 0.90 0.57 0.03 0.05 0.16 0.08 8.2

    All FPS measurements are conducted on a single NVIDIA RTX 4090 GPU. DiffusionDrive achieves the lowest average L2 error of 0.57 m and an average collision rate of 0.08%, outperforming VAD by 20.8% lower L2 error and 63.6% lower collision rate while running 1.8×\times faster.

  9. Knowl 9 — Architectural Ablation of Diffusion Decoder Components

    data/table

    The table presents an ablation study on the NAVSIM navtest split isolating the contribution of each architectural component inside the diffusion decoder.

    UNet Dec. Ego Query Spatial Cross-attn Agent/Map Cross-attn Cascade Dec. Param. ↓\downarrow NC ↑\uparrow DAC ↑\uparrow TTC ↑\uparrow Comf. ↑\uparrow EP ↑\uparrow PDMS ↑\uparrow
    ✓ ✓ × × × 102M 97.9 94.2 93.9 100 80.2 85.7
    × ✓ × × × 57M 88.7 83.2 80.0 84.8 43.3 55.1
    × ✓ ✓ × × 58M 98.2 95.4 94.4 100 81.3 87.1
    × ✓ × ✓ × 58M 97.9 93.5 93.8 100 79.8 85.1
    × ✓ ✓ ✓ × 59M 98.0 95.8 94.4 100 81.7 87.4
    × ✓ ✓ ✓ ✓ 60M 98.2 96.2 94.7 100 82.2 88.1

    Key empirical findings from the ablation:

    • Relying solely on the ego query without spatial scene attention degrades PDMS severely to 55.1.
    • Spatial deformable cross-attention with BEV/PV maps provides the largest single gain, boosting PDMS by +32.0 points to 87.1.
    • Cross-attention with agent/map perception tokens provides additional semantic context, improving PDMS from 87.1 to 87.4.
    • Stacking 2 cascade decoder layers improves PDMS further to 88.1 while reducing model parameters from 102M (UNet) to 60M.
  10. Knowl 10 — Sensitivity to Denoising Steps, Cascade Stages, and Candidate Trajectory Count

    empirical result

    Ablation experiments on NAVSIM navtest demonstrate the impact of inference denoising steps, cascade decoder stages, and candidate noise sample counts NinferN_{\text{infer}} on planning quality:

    1. Denoising Steps: Starting from the anchored Gaussian distribution, a single denoising step (S=1S = 1) achieves 87.9 PDMS. Increasing to S=2S = 2 steps reaches the maximum performance of 88.1 PDMS (NC: 98.2, DAC: 96.2, TTC: 94.7, EP: 82.2), while S=3S = 3 steps maintains 88.1 PDMS.
    2. Cascade Decoder Stages: A single cascade stage (1 layer, 59M parameters) achieves 87.4 PDMS. Stacking 2 cascade stages (60M parameters) increases performance to 88.1 PDMS. Increasing to 4 stages (65M parameters) yields 88.2 PDMS, indicating that 2 stages optimal balance quality and parameter efficiency.
    3. Inference Sample Count (NinferN_{\text{infer}}): Evaluating Ninfer=10N_{\text{infer}} = 10 noisy candidate samples yields 84.9 PDMS. Increasing candidate sampling to Ninfer=20N_{\text{infer}} = 20 raises PDMS to 88.1, and Ninfer=40N_{\text{infer}} = 40 reaches 88.2 PDMS, demonstrating that 20 candidate trajectories provide near-complete coverage of the driving action space.

Coverage note — None was omitted; all key contributions, mathematical formulations, decoder architecture details, roadmap benchmarks, main comparative results, and ablation studies have been captured.

References

  1. 1.Anurag Ajay, Yilun Du, Abhi Gupta, Joshua B. Tenenbaum, Tommi S. Jaakkola, and Pulkit Agrawal. Is conditional generative modeling all you need for decision making? In ICLR, 2023.
  2. 2.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020.
  3. 3.Shaoyu Chen, Tianheng Cheng, Xinggang Wang, Wenming Meng, Qian Zhang, and Wenyu Liu. Efficient and robust 2d-to-bev representation learning via geometry-guided kernel transformer. CoRR, abs/2206.04584, 2022.
  4. 4.Shaoyu Chen, Bo Jiang, Hao Gao, Bencheng Liao, Qing Xu, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang. Vadv2: End-to-end vectorized autonomous driving via probabilistic planning. arXiv preprint arXiv:2402.13243, 2024.
  5. 5.Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang, Chang Huang, and Wenyu Liu. Polardetr: Polar parametrization for vision-based surround-view 3d detection. Image and Vision Computing, 156:105438, 2025.
  6. 6.Zhili Chen, Maosheng Ye, Shuangjie Xu, Tongyi Cao, and Qifeng Chen. Ppad: Iterative interactions of prediction and planning for end-to-end autonomous driving. In ECCV, 2024.
  7. 7.Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In RSS, 2023.
  8. 8.Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger. Transfuser: Imitation with transformer-based sensor fusion for autonomous driving. TPAMI, 2022.
  9. 9.Younwoo Choi, Ray Coden Mercurius, Soheil Mohamad Alizadeh Shabestary, and Amir Rasouli. Dice: Diverse diffusion model with scoring for trajectory prediction. In IV, 2024.
  10. 10.OpenScene Contributors. Openscene: The largest up-to-date 3d occupancy prediction benchmark in autonomous driving. https://github.com/OpenDriveLab/OpenScene, 2023.
  11. 11.Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, and Kashyap Chitta. Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking. In NeurIPS, 2024.
  12. 12.Xunjiang Gu, Guanyu Song, Igor Gilitschenski, Marco Pavone, and Boris Ivanovic. Producing and leveraging online map uncertainty in trajectory prediction. In CVPR, 2024.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  15. 15.Shengchao Hu, Li Chen, Penghao Wu, Hongyang Li, Junchi Yan, and Dacheng Tao. St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning. In ECCV, 2022.
  16. 16.Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al. Planning-oriented autonomous driving. In CVPR, 2023.
  17. 17.Junjie Huang, Guan Huang, Zheng Zhu, Yun Ye, and Dalong Du. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv preprint arXiv:2112.11790, 2021.
  18. 18.Zhiyu Huang, Zixu Zhang, Ameya Vaidya, Yuxiao Chen, Chen Lv, and Jaime Fernandez Fisac. Versatile scene-consistent traffic scenario generation as optimization with diffusion. arXiv preprint arXiv:2404.02524, 2024.
  19. 19.Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine. Planning with diffusion for flexible behavior synthesis. In ICLR, 2022.
  20. 20.Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Vad: Vectorized scene representation for efficient autonomous driving. In ICCV, 2023.
  21. 21.Chiyu Jiang, Andre Cornman, Cheolho Park, Benjamin Sapp, Yin Zhou, Dragomir Anguelov, et al. Motiondiffuser: Controllable multi-agent motion prediction using diffusion. In CVPR, 2023.
  22. 22.Napat Karnchanachari, Dimitris Geromichalos, Kok Seang Tan, Nanxiang Li, Christopher Eriksen, Shakiba Yaghoubi, Noushin Mehdipour, Gianmarco Bernasconi, Whye Kit Fong, Yiluan Guo, et al. Towards learning-based planning: The nuplan benchmark for real-world autonomous driving. In ICRA, 2024.
  23. 23.Yingyan Li, Lue Fan, Jiawei He, Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang, and Tieniu Tan. Enhancing end-to-end autonomous driving with latent world model. arXiv preprint arXiv:2406.08481, 2024.
  24. 24.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In ECCV, 2022.
  25. 25.Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, et al. Hydra-mdp: End-to-end multimodal planning with multi-target hydra-distillation. arXiv preprint arXiv:2406.06978, 2024.
  26. 26.Zhiqi Li, Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu, and Jose M Alvarez. Is ego status all you need for open-loop end-to-end autonomous driving? In CVPR, 2024.
  27. 27.Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang, Wenyu Liu, and Chang Huang. MapTR: Structured modeling and learning for online vectorized HD map construction. In ICLR, 2023.
  28. 28.Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. Maptrv2: An end-to-end framework for online vectorized hd map construction. IJCV, 2024.
  29. 29.Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and Zhizhong Su. Sparse4d: Multi-view 3d object detection with sparse spatial-temporal fusion. arXiv preprint arXiv:2211.10581, 2022.
  30. 30.Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023.
  31. 31.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
  32. 32.Yicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang, and Hang Zhao. Vectormapnet: End-to-end vectorized hd map learning. In ICML, 2023.
  33. 33.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  34. 34.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  35. 35.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021.
  36. 36.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  37. 37.Ajay Sridhar, Dhruv Shah, Catherine Glossop, and Sergey Levine. Nomad: Goal masked diffusion policies for navigation and exploration. In ICRA, 2024.
  38. 38.Maria Stamatopoulou, Jianwei Liu, and Dimitrios Kanoulas. Dippest: Diffusion-based path planner for synthesizing trajectories applied on quadruped robots. arXiv preprint arXiv:2405.19232, 2024.
  39. 39.Wenchao Sun, Xuewu Lin, Yining Shi, Chuang Zhang, Haoran Wu, and Sifa Zheng. Sparsedrive: End-to-end autonomous driving via sparse scene representation. arXiv preprint arXiv:2405.19620, 2024.
  40. 40.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019.
  41. 41.Wenwen Tong, Chonghao Sima, Tai Wang, Li Chen, Silei Wu, Hanming Deng, Yi Gu, Lewei Lu, Ping Luo, Dahua Lin, et al. Scene as occupancy. In ICCV, 2023.
  42. 42.Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In CoRL, 2022.
  43. 43.Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. In CVPR, 2024.
  44. 44.Yixiao Wang, Chen Tang, Lingfeng Sun, Simone Rossi, Yichen Xie, Chensheng Peng, Thomas Hannagan, Stefano Sabatini, Nicola Poerio, Masayoshi Tomizuka, et al. Optimizing diffusion models for joint trajectory prediction and controllable generation. In ECCV, 2024.
  45. 45.Xinshuo Weng, Boris Ivanovic, Yan Wang, Yue Wang, and Marco Pavone. Para-drive: Parallelized architecture for real-time autonomous driving. In CVPR, 2024.
  46. 46.Zehang Weng, Haofei Lu, Danica Kragic, and Jens Lundell. Dexdiffuser: Generating dexterous grasps with diffusion models. arXiv preprint arXiv:2402.02989, 2024.
  47. 47.Sixu Yan, Zeyu Zhang, Muzhi Han, Zaijin Wang, Qi Xie, Zhitian Li, Zhehan Li, Hangxin Liu, Xinggang Wang, and Song-Chun Zhu. M2diffuser: Diffusion-based trajectory optimization for mobile manipulation in 3d scenes. arXiv preprint arXiv:2410.11402, 2024.
  48. 48.Brian Yang, Huangyuan Su, Nikolaos Gkanatsios, Tsung-Wei Ke, Ayush Jain, Jeff Schneider, and Katerina Fragkiadaki. Diffusion-es: Gradient-free planning with diffusion for autonomous driving and zero-shot instruction following. In CVPR, 2024.
  49. 49.Jingfeng Yao and Xinggang Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. arXiv preprint arXiv:2501.01423, 2025.
  50. 50.Jingfeng Yao, Cheng Wang, Wenyu Liu, and Xinggang Wang. Fasterdit: Towards faster diffusion transformers training without architecture modification. arXiv preprint arXiv:2410.10356, 2024.
  51. 51.Wenhao Yu, Jie Peng, Huanyu Yang, Junrui Zhang, Yifan Duan, Jianmin Ji, and Yanyong Zhang. Ldp: A local diffusion planner for efficient robot navigation and collision avoidance. arXiv preprint arXiv:2407.01950, 2024.
  52. 52.Chengran Yuan, Zhanqi Zhang, Jiawei Sun, Shuo Sun, Zefan Huang, Christina Dao Wen Lee, Dongen Li, Yuhang Han, Anthony Wong, Keng Peng Tee, et al. Drama: An efficient end-to-end motion planner for autonomous driving with mamba. arXiv preprint arXiv:2408.03601, 2024.
  53. 53.Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 3d diffusion policy. In RSS, 2024.
  54. 54.Fangao Zeng, Bin Dong, Yuang Zhang, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. Motr: End-to-end multiple-object tracking with transformer. In ECCV, 2022.
  55. 55.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and re-identification in multiple object tracking. IJCV, 2021.
  56. 56.Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. Bytetrack: Multi-object tracking by associating every detection box. In ECCV, 2022.
  57. 57.Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. Truncated diffusion probabilistic models and diffusion-based adversarial auto-encoders. In The Eleventh International Conference on Learning Representations, 2023.
  58. 58.Wenzhao Zheng, Ruiqi Song, Xianda Guo, Chenming Zhang, and Long Chen. Genad: Generative end-to-end autonomous driving. In ECCV, 2024.
  59. 59.Ziyuan Zhong, Davis Rempe, Yuxiao Chen, Boris Ivanovic, Yulong Cao, Danfei Xu, Marco Pavone, and Baishakhi Ray. Language-guided traffic simulation via scene-level diffusion. In CoRL, 2023.
  60. 60.Ziyuan Zhong, Davis Rempe, Danfei Xu, Yuxiao Chen, Sushant Veer, Tong Che, Baishakhi Ray, and Marco Pavone. Guided conditional diffusion for controllable traffic simulation. In ICRA, 2023.
  61. 61.Lianghui Zhu, Zilong Huang, Bencheng Liao, Jun Hao Liew, Hanshu Yan, Jiashi Feng, and Xinggang Wang. Dig: Scalable and efficient diffusion models with gated linear attention. arXiv preprint arXiv:2405.18428, 2024.
  62. 62.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR, 2021.

Citation

MLA
Liao, B., et al. “DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 12037–47, https://doi.org/10.1109/CVPR52734.2025.01124.
APA
Liao, B., Chen, S., Yin, H., Jiang, B., Wang, C., Yan, S., Zhang, X., Li, X., Zhang, Y., Zhang, Q., & Wang, X. (2025). DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12037–12047. https://doi.org/10.1109/CVPR52734.2025.01124
Chicago
Liao, B., S. Chen, H. Yin, et al. 2025. “DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12037–47. https://doi.org/10.1109/CVPR52734.2025.01124.
Harvard
Liao, B. et al. (2025) “DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 12037–12047. Available at: https://doi.org/10.1109/CVPR52734.2025.01124.
Vancouver
1. Liao B, Chen S, Yin H, et al (2025) DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 12037–12047

BibTeX

@inproceedings{Liao_2025, title={DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01124}, DOI={10.1109/cvpr52734.2025.01124}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Liao, Bencheng and Chen, Shaoyu and Yin, Haoran and Jiang, Bo and Wang, Cheng and Yan, Sixu and Zhang, Xinbang and Li, Xiangyu and Zhang, Ying and Zhang, Qian and Wang, Xinggang}, year={2025}, month=June, pages={12037–12047} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE