D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video
Tianhao WuFangcheng ZhongAndrea TagliasacchiForrester ColeCengiz Öztireli
Presents a self-supervised neural radiance field framework that reconstructs static 3D scenes from monocular video by decoupling moving objects and dynamic shadows using separate field networks and a skewed-entropy loss.
Reconstructing clean three-dimensional environments from casual, handheld monocular video is a core challenge for applications in robotics, autonomous navigation, and augmented reality. In everyday footage, moving objects and their associated shadows frequently occlude the static environment. Conventional two-dimensional image segmentation and inpainting techniques often struggle because they lack a multi-view 3D understanding of the scene. Meanwhile, existing neural 3D modeling methods typically rely on pre-trained supervised segmentation tools or struggle when applied to complex, non-rigid motion and time-varying shadows captured from a single moving camera.
The article demonstrates a self-supervised approach, named Decoupled Dynamic Neural Radiance Field (D2NeRF), that simultaneously separates dynamic moving objects and their dynamic shadows while reconstructing a complete, static 3D model of the background environment from monocular video without human supervision or pre-trained object detectors.
The method decomposes the scene into two separate neural radiance fields: a static field representing the background and a dynamic field representing moving elements over time. Because dynamic networks have high expressive capacity and can mistakenly represent static background elements, the researchers introduced a skewed-entropy loss along with ray and static regularizers to cleanly divide space between static and dynamic components. In addition, they implemented a dedicated shadow field that dynamically modulates surface brightness to capture cast shadows without altering the underlying static geometry and color. The authors evaluated the approach on synthetic benchmarks featuring moving objects and ground-truth shadows, as well as ten real-world video sequences captured on handheld smartphones.
Key experimental findings demonstrate that D2NeRF significantly outperforms existing state-of-the-art baselines across core benchmarks. In scene decoupling and static background reconstruction, D2NeRF achieved an average peak signal-to-noise ratio of 31.18, compared to 26.27 for NeuralDiff and 24.28 for NeRF-W. On video object segmentation, the method achieved a mean score of 0.717 across tested scenes, outperforming Motion Grouping (0.424) and NeuralDiff (0.372). Ablation studies verified that the combination of the skewed-entropy regularizer and ray constraints is critical, reducing perceptual image errors from 0.215 to 0.080, while the shadow field successfully removed large, view-correlated dynamic shadows that prior methods failed to resolve.
These findings indicate that high-quality 3D digital twins and clean background models can be extracted directly from casual consumer video captures. By removing the need for labor-intensive manual masking, pre-trained object detectors, or multi-camera setups, the method reduces operational costs and workflow complexity for 3D asset generation and environment mapping. Unlike prior frameworks limited to single rigid objects, the architecture reliably handles multiple non-rigid and topologically varying objects.
Organizations evaluating this technology should conduct pilot implementations on target operational video pipelines, paying attention to computational demands. Training takes approximately two hours across four specialized high-memory graphics processors per video sequence. For optimal performance, practitioners should perform hyperparameter tuning based on camera speed and motion levels in their specific use cases.
Decision-makers should note certain operational boundaries. The system relies on accurate camera calibration and constant scene illumination. Performance degrades on highly reflective surfaces, where view-dependent highlights may be misinterpreted as object motion, as well as on texture-less objects that move minimally within a narrow area over the duration of the capture. Nevertheless, the reported quantitative results and robust synthetic benchmarks provide high confidence in the method's effectiveness under standard lighting and well-calibrated camera trajectories.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Provides the foundational volumetric neural radiance field representation and volume rendering pipeline that D²NeRF decomposes into dynamic and static components.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). Introduces continuous dynamic neural radiance fields parameterized over time, establishing the baseline dynamic formulation that D²NeRF builds upon and regularizes to prevent dynamic overfitting.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Establishes deformation-based dynamic neural radiance fields for casual monocular video, providing essential background for representing time-varying volumetric scenes.
- Paper: NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, Ricardo Martin-Brualla et al. (2021). Introduces the concept of decomposing scenes into separate static and transient volumetric fields, directly informing D²NeRF's dual-field architecture for scene decoupling.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Pioneers compositional neural scene representations by explicitly separating distinct foreground objects from static backgrounds in 3D feature fields.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). Extends dynamic 3D scene representation from implicit neural volume rendering to explicit 4D Gaussian primitives for real-time rendering of moving scenes.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Advances beyond volumetric ray-marching radiance fields by introducing real-time 3D Gaussian splatting for radiance field rendering.
