3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis

Zhicheng LuXiang GuoLe HuiTianrui ChenMin YangXiao TangFeng ZhuYuchao Dai

article2024CVPR117 citations

Proposes a dynamic novel view synthesis framework that voxelizes 3D Gaussians and applies sparse 3D convolutions alongside continuous 6D rotation representations to model geometrically coherent non-rigid scene deformations.

Listen

Synthesizing photo-realistic views of moving scenes from single-camera video is a critical capability for immersive technologies like virtual and augmented reality. However, existing methods struggle to accurately reconstruct dynamic 3D environments because they learn motion deformations without adequate awareness of local 3D geometric structure. This lack of structural coherence leads to visual artifacts, unnatural distortions, and poor reconstruction quality during motion.

The article demonstrates a novel 3D geometry-aware deformable framework that models dynamic scenes using 3D Gaussian representations. The main objective is to establish whether explicitly incorporating 3D geometric features and continuous rotation parameters into the deformation process can significantly enhance dynamic view synthesis and structural consistency.

To achieve this, the approach represents the scene via a static canonical field paired with a deformation network. The system converts 3D Gaussian distributions into point clouds, extracts local structural features using efficient 3D sparse convolutions, and combines them with point-level features. A deformation model then predicts position, scale, and continuous six-dimensional rotation adjustments across time. The authors validated this methodology through extensive quantitative benchmarks and visual evaluations on standard synthetic datasets (the eight dynamic scenes of D-NeRF) and real-world video datasets (HyperNeRF) using a single modern GPU.

The findings show that the proposed framework consistently outperforms existing state-of-the-art methods across all standard image quality and perceptual error metrics. On synthetic benchmarks, the method achieved an overall peak signal-to-noise ratio of 38.01, surpassing the vanilla model baseline of 35.23 and outperforming alternative point-cloud and planar feature techniques. Ablation experiments confirmed that the sparse convolution geometry branch was the single most impactful component for performance gains, followed by continuous six-dimensional rotation modeling and time-specific density control. Furthermore, training completed in approximately two hours on average, with rendering speeds reaching 12 frames per second and compact total storage requirements of under 50 megabytes.

These results indicate that embedding explicit 3D geometric constraints into motion modeling yields smoother, physically plausible nonrigid deformations while maintaining computational efficiency. By overcoming the geometric inconsistencies inherent in earlier radiance field methods, this technique reduces the engineering risks and computational overhead of deploying high-quality dynamic 3D content in interactive applications.

Organizations developing immersive media, simulation, and spatial computing pipelines should consider adopting explicit geometry-aware Gaussian frameworks to model dynamic scenes. For next steps, technical teams should explore motion segmentation cues that separate moving foregrounds from static backgrounds to further refine motion modeling and accelerate rendering speeds.

While the method delivers high-confidence reconstructions across evaluated benchmarks, performance on real-world datasets remains constrained by narrow camera viewing ranges and inherent motion-shape ambiguities. Stakeholders should account for these real-world data limitations when deploying the pipeline in unconstrained capture environments.

arXiv: 2404.06270
Cover for 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Novel View Synthesis
  • 2.2. Dynamic View Synthesis
  • 3. Method
  • 3.1. Preliminary
  • 3.2. Gaussian Canonical Field
  • 3.3. Deformation Field
  • 3.4. Rasterization
  • 3.5. Optimization
  • 4. Experiments
  • 4.1. Dataset
  • 4.2. Implementation Details
  • 4.3. Quantitative Results
  • 4.4. Visualization Results
  • 4.5. Ablation Study
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Canonical Geometry Feature Learning via Sparse 3D Convolution

    model/method

    To capture local 3D geometric structures in dynamic scenes, the canonical space of a scene is represented using 3D Gaussian distributions treated as a 3D point cloud P∈RN×3\mathbf{P} \in \mathbb{R}^{N \times 3}, where NN is the number of Gaussians and only their 3D coordinates are used. A two-branch network extracts geometry-aware features:

    1. Geometric Branch: Converts the point cloud into sparse voxels by spatial discretization with a fixed grid size ss: V=floor⁡(P/s),\mathbf{V} = \operatorname{floor}(\mathbf{P} / s), where V∈ZM×3\mathbf{V} \in \mathbb{Z}^{M \times 3} and MM is the number of non-empty voxels. A sparse 3D U-Net composed of stacked sparse convolutions with skip connections processes V\mathbf{V} to produce aggregated local voxel features Fv∈RM×C\mathbf{F}_v \in \mathbb{R}^{M \times C}, where CC is the feature dimension. The voxel features are de-voxelized back to the point level by assigning each voxel's feature to the constituent points, yielding point-level geometric features Fp′∈RN×C\mathbf{F}'_p \in \mathbb{R}^{N \times C}.

    2. Identity Branch: A Multi-Layer Perceptron (MLP) maps individual 3D point coordinates directly to embedding features Fp∈RN×C\mathbf{F}_p \in \mathbb{R}^{N \times C} to preserve high-resolution, point-specific identity.

    3. Feature Fusion: The outputs of both branches are concatenated and projected by an MLP to form fused point-level geometry-aware features Ffuse∈RN×C\mathbf{F}_{\text{fuse}} \in \mathbb{R}^{N \times C}: Ffuse=MLP⁡(Concat⁡(Fp′,Fp)).\mathbf{F}_{\text{fuse}} = \operatorname{MLP}(\operatorname{Concat}(\mathbf{F}'_p, \mathbf{F}_p)).

  2. Knowl 2 — Continuous 6D Rotation Representation for 3D Gaussians

    model/method

    To avoid the discontinuities in parameter space associated with quaternion representations during neural network optimization, the orientation of each 3D Gaussian in canonical space and its temporal rotation deformation are parameterized using a continuous 6D rotation representation.

    Each Gaussian's rotation is represented by a pair of 3D column vectors [a1,a2]∈R3×2[\mathbf{a}_1, \mathbf{a}_2] \in \mathbb{R}^{3 \times 2} (initialized to [1,0,0]⊤[1, 0, 0]^\top and [0,1,0]⊤[0, 1, 0]^\top for identity rotation). The mapping fV2M:R6→SO(3)f_{\text{V2M}}: \mathbb{R}^6 \to SO(3) from the 6D vector to an orthogonal 3×33 \times 3 rotation matrix [b1,b2,b3]∈SO(3)[\mathbf{b}_1, \mathbf{b}_2, \mathbf{b}_3] \in SO(3) is defined via Gram-Schmidt orthogonalization: fV2M([∣∣a1a2∣∣])=[∣∣∣b1b2b3∣∣∣],f_{\text{V2M}}\left(\begin{bmatrix} | & | \\ \mathbf{a}_1 & \mathbf{a}_2 \\ | & | \end{bmatrix}\right) = \begin{bmatrix} | & | & | \\ \mathbf{b}_1 & \mathbf{b}_2 & \mathbf{b}_3 \\ | & | & | \end{bmatrix}, b1=N(a1),\mathbf{b}_1 = \mathcal{N}(\mathbf{a}_1), b2=N(a2−(b1⋅a2)b1),\mathbf{b}_2 = \mathcal{N}(\mathbf{a}_2 - (\mathbf{b}_1 \cdot \mathbf{a}_2)\mathbf{b}_1), b3=b1×b2,\mathbf{b}_3 = \mathbf{b}_1 \times \mathbf{b}_2, where N(v)=v/∥v∥2\mathcal{N}(\mathbf{v}) = \mathbf{v} / \|\mathbf{v}\|_2 denotes the vector normalization operator, ⋅\cdot denotes the inner product, and ×\times denotes the cross product.

  3. Knowl 3 — Gaussian Deformation Field and Temporal State Transformation

    model/method

    Given a canonical 3D Gaussian with position x∈R3\mathbf{x} \in \mathbb{R}^3, 6D rotation vector r∈R6\mathbf{r} \in \mathbb{R}^6, scaling vector s∈R3\mathbf{s} \in \mathbb{R}^3, spherical harmonics color coefficients c\mathbf{c}, opacity o∈Ro \in \mathbb{R}, and fused geometric feature Ffuse\mathbf{F}_{\text{fuse}}, a deformation decoder MLP GΦ\mathcal{G}_\Phi predicts time-dependent transformations at timestamp tt: (Δxt,Δrt,Δst)=GΦ(Ffuse,γ(x),γ(t)),(\Delta \mathbf{x}_t, \Delta \mathbf{r}_t, \Delta \mathbf{s}_t) = \mathcal{G}_\Phi(\mathbf{F}_{\text{fuse}}, \gamma(\mathbf{x}), \gamma(t)), where Δxt∈R3\Delta \mathbf{x}_t \in \mathbb{R}^3 is position deformation, Δrt∈R6\Delta \mathbf{r}_t \in \mathbb{R}^6 is continuous rotation deformation, Δst∈R3\Delta \mathbf{s}_t \in \mathbb{R}^3 is scale deformation, and γ(⋅)\gamma(\cdot) is a multiscale sinusoidal positional encoding: γ(p)=(sin⁡(20πp),cos⁡(20πp),…,sin⁡(2L−1πp),cos⁡(2L−1πp)).\gamma(p) = (\sin(2^0 \pi p), \cos(2^0 \pi p), \dots, \sin(2^{L-1} \pi p), \cos(2^{L-1} \pi p)).

    The deformed attributes of the Gaussian at time tt are computed as: xt=x+Δxt,\mathbf{x}_t = \mathbf{x} + \Delta \mathbf{x}_t, st=s+Δst,\mathbf{s}_t = \mathbf{s} + \Delta \mathbf{s}_t, rt=fV2M(Δrt)×fV2M(r),\mathbf{r}_t = f_{\text{V2M}}(\Delta \mathbf{r}_t) \times f_{\text{V2M}}(\mathbf{r}), where fV2Mf_{\text{V2M}} maps 6D vectors to SO(3)SO(3) rotation matrices, and matrix multiplication combines the base and delta rotations.

    Color parameters c\mathbf{c} and opacity oo are kept invariant over time to enforce physical consistency of individual Gaussian primitives across the timeline. The transformed Gaussians (xt,c,rt,st,o)(\mathbf{x}_t, \mathbf{c}, \mathbf{r}_t, \mathbf{s}_t, o) are subsequently rendered via differentiable tile rasterization using camera intrinsic K\mathbf{K} and extrinsic [R∣T][\mathbf{R}|\mathbf{T}] parameters.

  4. Knowl 4 — Dynamic Adaptive Density Control with Inverse Transformation

    model/method

    In dynamic scenes, Gaussian densification (cloning under-reconstructed Gaussians or splitting oversized Gaussians) cannot rely solely on canonical parameters because regions of high variance or under-reconstruction change across time.

    Adaptive density control is executed based on the deformed attributes at timestamp tt:

    1. Gaussians in under-reconstructed regions are identified and duplicated.
    2. Oversized Gaussians with high positional gradient variance are split into two smaller Gaussians, scaling their sizes down by a factor ϕ=1.6\phi = 1.6 and initializing their positions by sampling from the probability density function (PDF) of the original 3D Gaussian.
    3. Because Gaussian parameters are maintained in canonical space, an inverse deformation transformation ft−1(Δxt,Δrt,Δst)f_t^{-1}(\Delta \mathbf{x}_t, \Delta \mathbf{r}_t, \Delta \mathbf{s}_t) maps the newly cloned or split Gaussians from timestamp space tt back into canonical space.
  5. Knowl 5 — Optimization Objective for Deformable Gaussian Splatting

    equation

    The overall training objective combines a photometric reconstruction loss and a motion regularization term: L=Lphoto+ωLmotion,L = L_{\text{photo}} + \omega L_{\text{motion}}, where ω\omega is a trade-off hyperparameter.

    The photometric loss LphotoL_{\text{photo}} balances pixel-level L1L_1 loss and structural similarity loss LD-SSIML_{\text{D-SSIM}} between the rendered image C^t\hat{\mathbf{C}}_t and ground-truth image Ct\mathbf{C}_t: Lphoto=(1−λ)L1+λLD-SSIM,L_{\text{photo}} = (1 - \lambda) L_1 + \lambda L_{\text{D-SSIM}}, where λ\lambda balances the two metrics.

    The motion regularization term LmotionL_{\text{motion}} enforces scene rigidity under the assumption that the majority of scene points are static and motion amplitudes are small: Lmotion=∥Δxt∥1,L_{\text{motion}} = \|\Delta \mathbf{x}_t\|_1, where Δxt∈R3\Delta \mathbf{x}_t \in \mathbb{R}^3 is the predicted positional displacement of each Gaussian at timestamp tt.

  6. Knowl 6 — Implementation and Optimization Specifications

    experimental setup

    The model is optimized over 40,000 iterations. During the first 3,000 iterations, only the static canonical field is trained; the deformation field is introduced thereafter for dynamic optimization.

    • Optimizer & Learning Rate: Adam optimizer with an exponentially decaying learning rate from 8×10−48 \times 10^{-4} to 1.6×10−61.6 \times 10^{-6}.
    • Network Architectures:
      • Front point feature extractor: 2-layer MLP with a hidden width of 64.
      • Back point feature fusion: 3-layer MLP with a hidden width of 64.
      • Deformation decoder: 5-layer MLP with a width of 256 and skip connections.
    • Positional Encoding: L=10L = 10 frequency bands for spatial coordinates x\mathbf{x} and L=6L = 6 frequency bands for timestamp tt.
    • Initialization: For datasets without point clouds (e.g., D-NeRF), 150,000 points are randomly initialized; for datasets with provided point clouds (e.g., HyperNeRF), provided points are used as initial positions.
    • Hardware & Efficiency: Experiments execute on a single NVIDIA RTX 4090 GPU. Training on the D-NeRF dataset takes approximately 2 hours on average, achieves a rendering frame rate of 12 FPS at fixed viewpoints, and uses 48 MB total storage (34 MB point cloud + 14 MB network weights).
  7. Knowl 7 — Quantitative Evaluation on the D-NeRF Synthetic Benchmark

    data/table

    The performance of 3D geometry-aware deformable Gaussian Splatting was evaluated against prior state-of-the-art methods across 8 synthetic scenes from the D-NeRF dataset using PSNR (higher is better), SSIM (higher is better), and LPIPS (lower is better).

    Method Hell Warrior Mutant Hook
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    3D-GS 15.3924 0.8776 0.1300 21.7554 0.9359 0.0575 18.6933 0.8733 0.1144
    D-NeRF 25.0293 0.9506 0.0691 31.2900 0.9739 0.0268 29.2567 0.9650 0.1174
    TiNeuVox-B 28.2058 0.9661 0.0631 33.9029 0.9771 0.0301 31.7929 0.9718 0.0436
    NDVG 26.4933 0.9600 0.0670 34.4131 0.9801 0.0270 30.0009 0.9626 0.0463
    FDNeRF 27.7120 0.9665 0.0508 34.9727 0.9810 0.0312 32.2867 0.9756 0.0388
    4D-GS 28.1196 0.9730 0.0276 38.3411 0.9936 0.0062 33.1560 0.9810 0.0168
    Ours 32.2712 0.9835 0.0164 41.4284 0.9969 0.0029 36.9647 0.9916 0.0076
    Method Bouncing Balls Lego T-Rex
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    3D-GS 22.5575 0.9485 0.0647 23.0991 0.9329 0.0567 25.7496 0.9567 0.0474
    D-NeRF 38.9300 0.9900 0.1031 21.6427 0.8394 0.1654 31.7568 0.9767 0.0396
    TiNeuVox-B 40.8536 0.9913 0.0401 25.1748 0.9217 0.0689 32.7750 0.9783 0.0307
    NDVG 37.5157 0.9874 0.0751 25.0416 0.9395 0.0534 32.6229 0.9781 0.0330
    FDNeRF 40.0191 0.9912 0.0395 25.2700 0.9390 0.0460 30.7068 0.9731 0.0368
    4D-GS 40.7418 0.9941 0.0105 25.4024 0.9434 0.0377 33.3912 0.9869 0.0130
    Ours 43.5929 0.9960 0.0061 25.4411 0.9474 0.0329 39.0285 0.9952 0.0052
    Method Stand Up Jumping Jacks
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    3D-GS 19.3779 0.9200 0.0909 20.7163 0.9227 0.0980
    D-NeRF 32.7992 0.9818 0.0215 32.8031 0.9810 0.0373
    TiNeuVox-B 36.2031 0.9859 0.0199 34.7390 0.9823 0.0328
    NDVG 33.2158 0.9793 0.0302 31.2530 0.9737 0.0398
    FDNeRF 36.9107 0.9878 0.0188 33.5521 0.9812 0.0329
    4D-GS 38.2610 0.9923 0.0071 35.6656 0.9882 0.0159
    Ours 42.2101 0.9966 0.0028 37.9604 0.9928 0.0088

    The method outperforms existing implicit NeRF and Gaussian splatting baselines across all 8 scenes on all three metrics, achieving an average PSNR gain of over 3 dB on several complex dynamic sequences (e.g., Mutant, Stand Up, T-Rex).

  8. Knowl 8 — Quantitative Evaluation on the HyperNeRF Real Scene Benchmark

    data/table

    The method was benchmarked against existing dynamic view synthesis techniques on the real-world dataset proposed by HyperNeRF across multiple capture scenes.

    Individual scene results (PSNR and MS-SSIM):

    Method Chicken 3D Printer Broom Peel Banana
    PSNR↑\uparrow MS-SSIM↑\uparrow PSNR↑\uparrow MS-SSIM↑\uparrow PSNR↑\uparrow MS-SSIM↑\uparrow PSNR↑\uparrow MS-SSIM↑\uparrow
    TiNeuVox 28.2861 0.9474 22.7514 0.8392 21.2682 0.6832 24.5136 0.8743
    NDVG 27.0536 0.9390 22.4196 0.8389 21.4658 0.7028 22.8204 0.8279
    FDNeRF 27.9627 0.9438 22.8027 0.8453 21.9091 0.7154 24.2515 0.8645
    3D-GS 20.8915 0.7426 18.3991 0.6114 20.3953 0.6598 20.5654 0.8094
    Ours 28.5342 0.9331 22.0403 0.8098 20.8994 0.5241 25.5785 0.9067

    Average performance over 4 scenes (Cut Lemon, Chicken, 3D Printer, and Split Cookie):

    Method PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    TiNeuVox-B 27.16 0.76 0.40
    3D-GS 21.26 0.69 0.40
    4D-GS 26.98 0.78 0.31
    Ours 27.52 0.80 0.25

    The results confirm that the geometry-aware 3D Gaussian formulation maintains state-of-the-art novel view synthesis fidelity in real-world scenarios characterized by narrow camera baselines and motion ambiguities.

  9. Knowl 9 — Ablation Study on Architecture and Deformation Components

    data/table

    An ablation study conducted on the synthetic D-NeRF dataset (800×800800 \times 800 resolution) analyzed the contribution of each component to novel view synthesis quality:

    Method PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow
    w/o geo. feat. 37.5757 0.9841 0.0173
    w/o 6D rotation 37.8750 0.9851 0.0154
    canonical DC 37.8026 0.9847 0.0166
    vanilla 35.2307 0.9793 0.0242
    PointNet feat. 36.7353 0.9826 0.0184
    Plane feat. 35.9054 0.9811 0.0212
    ours full 38.0134 0.9853 0.0153

    Key takeaways:

    1. Geometric features: Removing the 3D sparse convolution branch (w/o geo. feat.) causes a 0.44 dB drop in PSNR. Using PointNet-style features (PointNet feat.) or 2D multi-plane projection features (Plane feat.) leads to drops of 1.28 dB and 2.11 dB in PSNR, respectively, demonstrating that sparse 3D convolution in canonical space extracts significantly better local geometric context than 2D projection or point MLPs.
    2. 6D Rotation representation: Replacing the continuous 6D rotation vector with unit quaternions (w/o 6D rotation) reduces PSNR from 38.01 dB to 37.88 dB, confirming that avoiding quaternion discontinuities helps optimization.
    3. Dynamic Density Control: Controlling densification purely in canonical space (canonical DC) degrades PSNR by 0.21 dB because it fails to capture temporal variances at specific timestamps.
    4. Baseline comparison: The full proposed model outperforms the vanilla MLP deformation model (vanilla) by 2.78 dB in PSNR.

Coverage note — None was omitted; all key methodological formulations (canonical 3D sparse feature learning, continuous 6D rotation parameterization, deformation field, temporal density control), experimental settings, comparative tables on synthetic/real benchmarks, and ablation studies are covered.

References

  1. 1.Kara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lempitsky. Neural point-based graphics. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  2. 2.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-NeRF: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  3. 3.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Zip-NeRF: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 2
  4. 4.Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Barron, Ce Liu, and Hendrik Lensch. NeRD: Neural reflectance decomposition from image collections. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  5. 5.Chris Buehler, Michael Bosse, Leonard McMillan, Steven Gortler, and Michael Cohen. Unstructured lumigraph rendering. In Proceedings of the Conference on Computer Graphics and Interactive Techniques, 2001. 2
  6. 6.Ang Cao and Justin Johnson. HexPlane: A fast representation for dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  7. 7.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  8. 8.Shenchang Eric Chen and Lance Williams. View interpolation for image synthesis. In Proceedings of the Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 1993. 2
  9. 9.Inchang Choi, Orazio Gallo, Alejandro Troccoli, Min H Kim, and Jan Kautz. Extreme view synthesis. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019. 2
  10. 10.Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B Tenenbaum, and Jiajun Wu. Neural radiance flow for 4d view synthesis and video processing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 3
  11. 11.Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. Fast dynamic radiance fields with time-aware neural voxels. In Proceedings of the Conference on Computer Graphics and Interactive Techniques in Asia (SIGGRAPH ASIA), 2022. 1, 2, 3, 6, 7
  12. 12.John Flynn, Ivan Neulander, James Philbin, and Noah Snavely. DeepStereo: Learning to predict new views from the world's imagery. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2
  13. 13.Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-Planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3
  14. 14.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 3
  15. 15.Ned Greene. Environment mapping and other applications of world projections. IEEE Computer Graphics and Applications, 1986. 2
  16. 16.Xiang Guo, Guanying Chen, Yuchao Dai, Xiaoqing Ye, Jiadai Sun, Xiao Tan, and Errui Ding. Neural deformable voxel grid for fast optimization of dynamic view synthesis. In Proceedings of the Asian Conference on Computer Vision (ACCV), 2022. 1, 2, 3, 6
  17. 17.Xiang Guo, Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan, Errui Ding, Yumeng Zhang, and Jingdong Wang. Forward flow for novel view synthesis of dynamic scenes. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 1, 2, 3, 6
  18. 18.Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering. ACM Transactions on Graphics (TOG), 2018. 2
  19. 19.Wenbo Hu, Yuling Wang, Lin Ma, Bangbang Yang, Lin Gao, Xiao Liu, and Yuewen Ma. Tri-MipRF: Tri-mip representation for efficient anti-aliasing neural radiance fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 2
  20. 20.Nima Khademi Kalantari, Ting-Chun Wang, and Ravi Ramamoorthi. Learning-based view synthesis for light field cameras. ACM Transactions on Graphics (TOG), 2016. 2
  21. 21.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3D Gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (TOG), 2023. 1, 2, 3, 4, 5, 6, 7
  22. 22.Marc Levoy and Pat Hanrahan. Light field rendering. In Proceedings of the Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 1996. 2
  23. 23.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  24. 24.David B Lindell, Julien NP Martel, and Gordon Wetzstein. AutoInt: Automatic integration for fast neural volume rendering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  25. 25.Baoyuan Liu, Min Wang, Hassan Foroosh, Marshall Tappen, and Marianna Pensky. Sparse convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 4
  26. 26.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020. 2
  27. 27.Yu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim, Yung-Yu Chuang, Johannes Kopf, and Jia-Bin Huang. Robust dynamic radiance fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2, 3
  28. 28.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics (TOG), 2021. 2
  29. 29.Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  30. 30.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 1, 2, 4, 5
  31. 31.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (TOG), 2022. 2
  32. 32.Thomas Neff, Pascal Stadlbauer, Mathias Parger, Andreas Kurz, Joerg H. Mueller, Chakravarty R. Alla Chaitanya, Anton S. Kaplanyan, and Markus Steinberger. DONeRF: Towards real-time rendering of compact neural radiance fields using depth oracle networks. Computer Graphics Forum (CGF), 2021. 2
  33. 33.Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 1, 3
  34. 34.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. HyperNeRF: A higher-dimensional representation for topologically varying neural radiance fields. ACM Transactions on Graphics (TOG), 2021. 1, 3, 6, 7
  35. 35.Eric Penner and Li Zhang. Soft 3D reconstruction for view synthesis. ACM Transactions on Graphics (TOG), 2017. 2
  36. 36.Martin Piala and Ronald Clark. TermiNeRF: Ray termination prediction for efficient neural rendering. In Proceedings of the International Conference on 3D Vision (3DV), 2021. 2
  37. 37.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 2, 3, 6
  38. 38.Daniel Rebain, Wei Jiang, Soroosh Yazdani, Ke Li, Kwang Moo Yi, and Andrea Tagliasacchi. DeRF: Decomposed radiance fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  39. 39.Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. KiloNeRF: Speeding up neural radiance fields with thousands of tiny MLPs. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  40. 40.Konstantinos Rematas, Andrew Liu, Pratul P. Srinivasan, Jonathan T. Barron, Andrea Tagliasacchi, Thomas Funkhouser, and Vittorio Ferrari. Urban radiance fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  41. 41.Gernot Riegler and Vladlen Koltun. Free view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  42. 42.Gernot Riegler and Vladlen Koltun. Stable view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  43. 43.Harry Shum and Sing Bing Kang. Review of image-based rendering techniques. In Visual Communications and Image Processing (VCIP), 2000. 2
  44. 44.Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. NeRF-Player: A streamable dynamic scene representation with decomposed neural radiance fields. IEEE Transactions on Visualization and Computer Graphics (TVCG), 2023. 3
  45. 45.Pratul P Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T Barron. NeRV: Neural reflectance and visibility fields for relighting and view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  46. 46.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct Voxel Grid Optimization: Super-fast convergence for radiance fields reconstruction. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  47. 47.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-NeRF: Scalable large scene neural view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  48. 48.Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nießner, et al. State of the art on neural rendering. In Computer Graphics Forum (CGF), 2020. 2
  49. 49.Ayush Tewari, O Fried, J Thies, V Sitzmann, S Lombardi, Z Xu, T Simon, M Nießner, E Tretschk, L Liu, et al. Advances in neural rendering. In Proceedings of the Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 2021. 2
  50. 50.Justus Thies, Michael Zollhöfer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 2019. 2
  51. 51.Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer, Christoph Lassner, and Christian Theobalt. Non-rigid neural radiance fields: Reconstruction and novel view synthesis of a dynamic scene from monocular video. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 1, 3
  52. 52.Alex Trevithick and Bo Yang. GRF: Learning a general radiance field for 3D representation and rendering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  53. 53.Haithem Turki, Jason Y Zhang, Francesco Ferroni, and Deva Ramanan. SUDS: Scalable urban dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1, 2, 3
  54. 54.Chaoyang Wang, Ben Eckart, Simon Lucey, and Orazio Gallo. Neural trajectory fields for dynamic novel view synthesis. arXiv preprint arXiv:2105.05994, 2021. 3
  55. 55.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srinivasan, Howard Zhou, Jonathan T Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. IBRNet: Learning multi-view image-based rendering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  56. 56.Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4D Gaussian splatting for real-time dynamic scene rendering. arXiv preprint arXiv:2310.08528, 2023. 3, 6, 7
  57. 57.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  58. 58.Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Anyi Rao, Christian Theobalt, Bo Dai, and Dahua Lin. BungeeNeRF: Progressive neural radiance field for extreme multi-scale scene rendering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  59. 59.Zexiang Xu, Sai Bi, Kalyan Sunkavalli, Sunil Hadap, Hao Su, and Ravi Ramamoorthi. Deep view synthesis from sparse photometric images. ACM Transactions on Graphics (TOG), 2019. 2
  60. 60.Wenqi Yang, Guanying Chen, Chaofeng Chen, Zhenfang Chen, and Kwan-Yee K Wong. S3-NeRF: Neural reflectance field from shading and shadow under a single viewpoint. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2022. 2
  61. 61.Jae Shin Yoon, Kihwan Kim, Orazio Gallo, Hyun Soo Park, and Jan Kautz. Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  62. 62.Alex Yu, Sara Fridovich-Keil, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  63. 63.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021. 2
  64. 64.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  65. 65.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. NeRF++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 2
  66. 66.Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 6
  67. 67.Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul Debevec, William T Freeman, and Jonathan T Barron. NeRFactor: Neural factorization of shape and reflectance under an unknown illumination. ACM Transactions on Graphics (TOG), 2021. 2
  68. 68.Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2, 4
  69. 69.M. Zwicker, H. Pfister, J. van Baar, and M. Gross. Ewa volume splatting. In Proceedings of IEEE Visualization (VIS), 2001. 4

Citation

MLA
Lu, Z., et al. “3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis”. arXiv, 2024, http://arxiv.org/abs/2404.06270v2.
APA
Lu, Z., Guo, X., Hui, L., Chen, T., Yang, M., Tang, X., Zhu, F., & Dai, Y. (2024). 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis. arXiv. http://arxiv.org/abs/2404.06270v2
Chicago
Lu, Z., X. Guo, L. Hui, et al. 2024. “3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis”. arXiv. http://arxiv.org/abs/2404.06270v2.
Harvard
Lu, Z. et al. (2024) “3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.06270v2.
Vancouver
1. Lu Z, Guo X, Hui L, Chen T, Yang M, Tang X, Zhu F, Dai Y (2024) 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis. arXiv

BibTeX

@article{lu2024geometry,
  title = {3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis},
  author = {Lu, Zhicheng and Guo, Xiang and Hui, Le and Chen, Tianrui and Yang, Min and Tang, Xiao and Zhu, Feng and Dai, Yuchao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.06270v2},
  eprint = {2404.06270}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE