3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis
Zhicheng LuXiang GuoLe HuiTianrui ChenMin YangXiao TangFeng ZhuYuchao Dai
Proposes a dynamic novel view synthesis framework that voxelizes 3D Gaussians and applies sparse 3D convolutions alongside continuous 6D rotation representations to model geometrically coherent non-rigid scene deformations.
Synthesizing photo-realistic views of moving scenes from single-camera video is a critical capability for immersive technologies like virtual and augmented reality. However, existing methods struggle to accurately reconstruct dynamic 3D environments because they learn motion deformations without adequate awareness of local 3D geometric structure. This lack of structural coherence leads to visual artifacts, unnatural distortions, and poor reconstruction quality during motion.
The article demonstrates a novel 3D geometry-aware deformable framework that models dynamic scenes using 3D Gaussian representations. The main objective is to establish whether explicitly incorporating 3D geometric features and continuous rotation parameters into the deformation process can significantly enhance dynamic view synthesis and structural consistency.
To achieve this, the approach represents the scene via a static canonical field paired with a deformation network. The system converts 3D Gaussian distributions into point clouds, extracts local structural features using efficient 3D sparse convolutions, and combines them with point-level features. A deformation model then predicts position, scale, and continuous six-dimensional rotation adjustments across time. The authors validated this methodology through extensive quantitative benchmarks and visual evaluations on standard synthetic datasets (the eight dynamic scenes of D-NeRF) and real-world video datasets (HyperNeRF) using a single modern GPU.
The findings show that the proposed framework consistently outperforms existing state-of-the-art methods across all standard image quality and perceptual error metrics. On synthetic benchmarks, the method achieved an overall peak signal-to-noise ratio of 38.01, surpassing the vanilla model baseline of 35.23 and outperforming alternative point-cloud and planar feature techniques. Ablation experiments confirmed that the sparse convolution geometry branch was the single most impactful component for performance gains, followed by continuous six-dimensional rotation modeling and time-specific density control. Furthermore, training completed in approximately two hours on average, with rendering speeds reaching 12 frames per second and compact total storage requirements of under 50 megabytes.
These results indicate that embedding explicit 3D geometric constraints into motion modeling yields smoother, physically plausible nonrigid deformations while maintaining computational efficiency. By overcoming the geometric inconsistencies inherent in earlier radiance field methods, this technique reduces the engineering risks and computational overhead of deploying high-quality dynamic 3D content in interactive applications.
Organizations developing immersive media, simulation, and spatial computing pipelines should consider adopting explicit geometry-aware Gaussian frameworks to model dynamic scenes. For next steps, technical teams should explore motion segmentation cues that separate moving foregrounds from static backgrounds to further refine motion modeling and accelerate rendering speeds.
While the method delivers high-confidence reconstructions across evaluated benchmarks, performance on real-world datasets remains constrained by narrow camera viewing ranges and inherent motion-shape ambiguities. Stakeholders should account for these real-world data limitations when deploying the pipeline in unconstrained capture environments.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). It introduces 3D Gaussian Splatting for real-time radiance field rendering, providing the foundational explicit primitive representation that the source deforms and enhances.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). It establishes deformation fields coupled with 3D Gaussian splatting for dynamic scene rendering, setting up the exact baseline framework that the source improves with 3D geometry awareness.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). It introduces canonical space deformation fields for dynamic view synthesis and establishes the benchmark dynamic scene dataset evaluated by the source.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). It introduces continuous deformable fields and geometric regularization for non-rigid dynamic scene synthesis from monocular video.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). It details the sparse 3D/4D convolution operations utilized by the source to efficiently extract local geometric features from point-converted Gaussian primitives.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). It lays the groundwork for coordinate-based neural view synthesis and volume rendering upon which dynamic radiance fields and Gaussian methods are built.
- Paper: DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes, Xiaoyu Zhou et al. (2024). It extends Gaussian splatting to dynamic, multi-camera autonomous driving environments by explicitly decoupling moving objects from static backgrounds.
- Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). It builds on dynamic 3D Gaussian modeling to capture high-fidelity non-rigid facial motions and realistic continuous relighting.
- Paper: VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction, Jiaqi Lin et al. (2024). It explores scalable spatial partitioning techniques for Gaussian representations to reconstruct vast, unconstrained real-world environments.
