DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
Xiaoyu ZhouZhiwei LinXiaojun ShanYongtao WangDeqing SunMing-Hsuan Yang
Proposes a composite 3D Gaussian splatting framework that couples incremental background modeling with dynamic object graphs and LiDAR priors to achieve high-fidelity surround-view rendering for complex autonomous driving scenarios.
High-fidelity three-dimensional reconstruction and simulation of dynamic driving environments are vital for developing and testing autonomous vehicles, especially for generating safety-critical edge cases. However, existing methods struggle to model large-scale, 360-degree surrounding scenes when vehicles and surrounding objects move at high speeds. Previous neural reconstruction techniques are computationally demanding, suffer from visual blurring, and fail to maintain multi-camera consistency across outward-facing views with minimal overlap.
The article demonstrates and evaluates DrivingGaussian, a novel framework designed to reconstruct complex, large-scale dynamic driving scenes and synthesize photorealistic surrounding views. The framework decomposes scenes into static backgrounds and moving objects using sequential multi-sensor data, such as camera images and laser-based depth sensors (LiDAR).
The approach introduces two core components: an incremental model that builds the expansive static background sequentially across time, and a dynamic graph model that tracks, reconstructs, and positions moving objects individually. These representations are combined into a global model that correctly captures real-world occlusions. Furthermore, the method integrates LiDAR data to initialize geometric priors and improve spatial consistency across multi-camera setups. The authors evaluated this approach against leading neural radiance and Gaussian-based methods on established autonomous driving benchmarks, specifically nuScenes (multi-camera) and KITTI-360 (monocular).
The results show that the proposed method outperforms all existing state-of-the-art baselines. On the nuScenes benchmark, DrivingGaussian achieved superior image quality metrics, recording a peak signal-to-noise ratio (PSNR) of 28.74 dB—outperforming the leading competitor EmerNeRF (26.75 dB) and baseline 3D Gaussian Splatting (26.08 dB)—while reducing perceptual error (LPIPS) to 0.237 compared to over 0.298–0.311 for competing models. The method also proved effective in monocular settings on KITTI-360, leading the field with a PSNR of 25.62 dB. Ablation studies revealed that dynamic object modeling and static background decomposition are critical to performance, while LiDAR priors notably improve fine structural detail without requiring overly dense point clouds.
These findings mean that autonomous vehicle programs can generate realistic, multi-view dynamic simulations and test safety-critical edge cases—such as sudden pedestrian hazards—without high physical testing costs. By eliminating the need to estimate complex motion flows and addressing previous rendering artifacts, the framework improves simulation fidelity and developer safety validation pipelines.
Organizations developing autonomous systems should consider integrating composite Gaussian modeling pipelines into their simulation workflows for sensor validation and corner-case testing. However, the evaluation assumes the availability of bounding box detections (or tracking models) and calibrated sensor arrays. While the system demonstrates high resilience and functions well even with traditional point initialization instead of LiDAR, users should test performance in extreme optical conditions or unannotated environments before deploying it as a core component of production validation pipelines.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). This foundational work introduces 3D Gaussian Splatting and its real-time differentiable rendering pipeline, which DrivingGaussian adopts and adapts for dynamic driving scene reconstruction.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). This work establishes dynamic 4D Gaussian Splatting formulations for moving scenes, providing essential context for representing time-varying Gaussian primitives.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). This paper presents foundational scene contraction and rendering techniques for unbounded 360-degree environments that directly motivate large-scale surrounding scene modeling in driving setups.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). This paper introduces methods for decoupling dynamic foreground objects from static background scenes in neural rendering, directly preceding DrivingGaussian's composite decomposition approach.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). This work formulates compositional neural feature fields to separate foreground objects from background environments, establishing the conceptual foundation for composite scene modeling.
- Paper: VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction, Jiaqi Lin et al. (2024). This work scales 3D Gaussian Splatting to vast, expansive environments via spatial partitioning and decoupled appearance modeling, extending the large-scale reconstruction capabilities highlighted in DrivingGaussian.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). This paper advances Gaussian splatting by introducing 2D planar disk primitives to achieve geometrically accurate surface reconstruction and multi-view perspective consistency.
- Paper: GaussianPro: 3D Gaussian Splatting with Progressive Propagation, Kai Cheng et al. (2024). This work develops progressive propagation for Gaussian splatting in large-scale and texture-less scenes, addressing point cloud initialization limitations encountered in complex environments.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). This study introduces a massive 10K-scene real-world benchmark to systematically evaluate and compare advanced novel view synthesis methods like 3D Gaussian Splatting across complex, unbounded spaces.
