D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

Tianhao WuFangcheng ZhongAndrea TagliasacchiForrester ColeCengiz Öztireli

article2022NeurIPS198 citations

Presents a self-supervised neural radiance field framework that reconstructs static 3D scenes from monocular video by decoupling moving objects and dynamic shadows using separate field networks and a skewed-entropy loss.

Listen

Reconstructing clean three-dimensional environments from casual, handheld monocular video is a core challenge for applications in robotics, autonomous navigation, and augmented reality. In everyday footage, moving objects and their associated shadows frequently occlude the static environment. Conventional two-dimensional image segmentation and inpainting techniques often struggle because they lack a multi-view 3D understanding of the scene. Meanwhile, existing neural 3D modeling methods typically rely on pre-trained supervised segmentation tools or struggle when applied to complex, non-rigid motion and time-varying shadows captured from a single moving camera.

The article demonstrates a self-supervised approach, named Decoupled Dynamic Neural Radiance Field (D2NeRF), that simultaneously separates dynamic moving objects and their dynamic shadows while reconstructing a complete, static 3D model of the background environment from monocular video without human supervision or pre-trained object detectors.

The method decomposes the scene into two separate neural radiance fields: a static field representing the background and a dynamic field representing moving elements over time. Because dynamic networks have high expressive capacity and can mistakenly represent static background elements, the researchers introduced a skewed-entropy loss along with ray and static regularizers to cleanly divide space between static and dynamic components. In addition, they implemented a dedicated shadow field that dynamically modulates surface brightness to capture cast shadows without altering the underlying static geometry and color. The authors evaluated the approach on synthetic benchmarks featuring moving objects and ground-truth shadows, as well as ten real-world video sequences captured on handheld smartphones.

Key experimental findings demonstrate that D2NeRF significantly outperforms existing state-of-the-art baselines across core benchmarks. In scene decoupling and static background reconstruction, D2NeRF achieved an average peak signal-to-noise ratio of 31.18, compared to 26.27 for NeuralDiff and 24.28 for NeRF-W. On video object segmentation, the method achieved a mean score of 0.717 across tested scenes, outperforming Motion Grouping (0.424) and NeuralDiff (0.372). Ablation studies verified that the combination of the skewed-entropy regularizer and ray constraints is critical, reducing perceptual image errors from 0.215 to 0.080, while the shadow field successfully removed large, view-correlated dynamic shadows that prior methods failed to resolve.

These findings indicate that high-quality 3D digital twins and clean background models can be extracted directly from casual consumer video captures. By removing the need for labor-intensive manual masking, pre-trained object detectors, or multi-camera setups, the method reduces operational costs and workflow complexity for 3D asset generation and environment mapping. Unlike prior frameworks limited to single rigid objects, the architecture reliably handles multiple non-rigid and topologically varying objects.

Organizations evaluating this technology should conduct pilot implementations on target operational video pipelines, paying attention to computational demands. Training takes approximately two hours across four specialized high-memory graphics processors per video sequence. For optimal performance, practitioners should perform hyperparameter tuning based on camera speed and motion levels in their specific use cases.

Decision-makers should note certain operational boundaries. The system relies on accurate camera calibration and constant scene illumination. Performance degrades on highly reflective surfaces, where view-dependent highlights may be misinterpreted as object motion, as well as on texture-less objects that move minimally within a narrow area over the duration of the capture. Nevertheless, the reported quantitative results and robust synthetic benchmarks provide high confidence in the method's effectiveness under standard lighting and well-calibrated camera trajectories.

Cover for D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

Abstract

Given a monocular video, segmenting and decoupling dynamic objects while recovering the static environment is a widely studied problem in machine intelligence. Existing solutions usually approach this problem in the image domain, limiting their performance and understanding of the environment. We introduce Decoupled Dynamic Neural Radiance Field (D²NeRF), a self-supervised approach that takes a monocular video and learns a 3D scene representation which decouples moving objects, including their shadows, from the static background. Our method represents the moving objects and the static background by two separate neural radiance fields with only one allowing for temporal changes. A naive implementation of this approach leads to the dynamic component taking over the static one as the representation of the former is inherently more general and prone to overfitting. To this end, we propose a novel loss to promote correct separation of phenomena. We further propose a shadow field network to detect and decouple dynamically moving shadows. We introduce a new dataset containing various dynamic objects and shadows and demonstrate that our method can achieve better performance than state-of-the-art approaches in decoupling dynamic and static 3D objects, occlusion and shadow removal, and image segmentation for moving objects. Project page: d2nerf.github.io

Citation

MLA
Wu, T., et al. “D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 32653–66, https://proceedings.neurips.cc/paper_files/paper/2022/file/d2cc447db9e56c13b993c11b45956281-Paper-Conference.pdf.
APA
Wu, T., Zhong, F., Tagliasacchi, A., Cole, F., & Oztireli, C. (2022). D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video. Advances in Neural Information Processing Systems, 35, 32653–32666. https://proceedings.neurips.cc/paper_files/paper/2022/file/d2cc447db9e56c13b993c11b45956281-Paper-Conference.pdf
Chicago
Wu, T., F. Zhong, A. Tagliasacchi, F. Cole, and C. Oztireli. 2022. “D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video”. Advances in Neural Information Processing Systems 35: 32653–66. https://proceedings.neurips.cc/paper_files/paper/2022/file/d2cc447db9e56c13b993c11b45956281-Paper-Conference.pdf.
Harvard
Wu, T. et al. (2022) “D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32653–32666. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d2cc447db9e56c13b993c11b45956281-Paper-Conference.pdf.
Vancouver
1. Wu T, Zhong F, Tagliasacchi A, Cole F, Oztireli C (2022) D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32653–32666

BibTeX

@inproceedings{wu20222nerf,
  title = {D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video},
  author = {Wu, Tianhao and Zhong, Fangcheng and Tagliasacchi, Andrea and Cole, Forrester and Oztireli, Cengiz},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {32653-32666},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d2cc447db9e56c13b993c11b45956281-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission