Decomposing NeRF for Editing via Feature Field Distillation

Sosuke KobayashiEiichi MatsumotoVincent Sitzmann

article2022NeurIPS438 citations

Proposes distilling 2D foundation model representations into 3D feature fields alongside neural radiance fields, enabling query-based 3D semantic segmentation and localized scene editing without retraining.

Listen

Neural radiance fields (NeRF) have emerged as a powerful technique for constructing high-fidelity 3D scene representations and generating novel viewpoints from 2D imagery. However, standard NeRF models implicitly encode entire scenes into interconnected network weights, making selective manipulation or object-centric editing difficult. Prior segmentation methods typically rely on expensive manual annotations or restricted closed-set categories, which limits practical interactive editing. The article addresses this operational bottleneck by introducing distilled feature fields (DFF), a method that transfers semantic understanding into 3D scenes without requiring manual 3D annotations.

The main objective of the article is to demonstrate that knowledge from pre-trained 2D vision and language models can be distilled into 3D feature fields to enable query-based zero-shot semantic decomposition and local scene editing. The authors formulate a framework where a 3D feature field is optimized alongside standard radiance parameters via differentiable volume rendering, supervised directly by pre-trained 2D teacher models such as LSeg for text queries and DINO for visual image patch queries. To evaluate the framework, the authors conducted quantitative segmentation benchmarks on indoor scenes from the Replica dataset and performed qualitative appearance, geometry, and optimization-based editing experiments across complex real-world captures.

The evaluation revealed several critical findings. First, DFF achieved superior 3D semantic segmentation performance on the Replica benchmark compared to a supervised 3D convolutional baseline trained on the ScanNet dataset, improving mean intersection-over-union from 0.475 to 0.589 and overall accuracy from 75.8% to 85.5%. Second, adding a feature branch to the neural field introduced minimal computational overhead and preserved visual reconstruction fidelity, showing essentially identical novel view synthesis quality (32.85 dB peak signal-to-noise ratio versus 32.87 dB for baseline NeRF). Third, the approach enabled precise, view-consistent local edits—including color alterations, deletions, and geometric translations—driven by text prompts, image patches, or point selections. Fourth, when integrated with text-driven optimization frameworks like CLIPNeRF, DFF prevented unintended global style alterations by successfully confining modifications strictly to designated target regions.

These findings indicate that 3D neural graphics can leverage off-the-shelf 2D foundational models to achieve open-vocabulary, interactive scene editing without the cost of collecting specialized 3D training datasets. By eliminating the need to retrain models for new semantic categories, this method significantly reduces workflow friction and computation costs for downstream graphics, spatial computing, and digital asset creation. Furthermore, it demonstrates that multi-view volume rendering inherently denoises 2D features, producing robust 3D segmentation fields that outperform discrete point-cloud baselines.

Organizations developing neural rendering pipelines should adopt feature field distillation to enable flexible scene manipulation and asset decomposition. Practitioners should utilize coarse sampling or remove high-frequency positional encodings in feature branches to enforce spatial smoothness and reduce boundary artifacts. For broader integration, lightweight independent feature networks can be deployed to retrofit existing, off-the-shelf NeRF models or modern hash-grid representations without retraining underlying geometry. Future research should prioritize enhancing background reconstruction behind removed objects and developing view-dependent referring mechanisms to support spatial language queries.

Confidence in the reported semantic segmentation and local editing capabilities is high, backed by both quantitative benchmarks and diverse real-world visual demonstrations. Nonetheless, key operational constraints remain. The student representation is ultimately bounded by the resolution and vocabulary limits of its 2D teacher models, and underlying geometric reconstruction artifacts in the radiance field can introduce noise into the feature supervision. Additionally, deleting prominent foreground objects can occasionally expose blurred or incomplete background geometry where camera views were occluded.

arXiv: 2205.15585
Cover for Decomposing NeRF for Editing via Feature Field Distillation

Abstract

Emerging neural radiance fields (NeRF) are a promising scene representation for computer graphics, enabling high-quality 3D reconstruction and novel view synthesis from image observations. However, editing a scene represented by a NeRF is challenging, as the underlying connectionist representations such as MLPs or voxel grids are not object-centric or compositional. In particular, it has been difficult to selectively edit specific regions or objects. In this work, we tackle the problem of semantic scene decomposition of NeRFs to enable query-based local editing of the represented 3D scenes. We propose to distill the knowledge of off-the-shelf, self-supervised 2D image feature extractors such as CLIP-LSeg or DINO into a 3D feature field optimized in parallel to the radiance field. Given a user-specified query of various modalities such as text, an image patch, or a point-and-click selection, 3D feature fields semantically decompose 3D space without the need for re-training and enable us to semantically select and edit regions in the radiance field. Our experiments validate that the distilled feature fields (DFFs) can transfer recent progress in 2D vision and language foundation models to 3D scene representations, enabling convincing 3D segmentation and selective editing of emerging neural graphics representations.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Neural Radiance Fields (NeRF)
  • 3.2 Pre-trained Models and Zero-shot Segmentation of Images
  • 4 Distilled Feature Fields
  • 4.1 Distilling Foundation Modules into 3D Feature Fields via Volume Rendering
  • 4.2 Query-based Decomposition and Editing
  • 5 Experiments
  • 5.1 3D Semantic Segmentation
  • 5.2 Editable Novel View Synthesis
  • 6 Discussion, Limitations, and Conclusions
  • References
  • A Training and Model Architectures
  • B Editing Procedure
  • C Feature Encoders
  • D Replica Dataset Experiment
  • E Ablation Experiments of Variants
  • F Implementation of CLIPNeRF Experiment
  • G CLIP-inspired Segmentation Models

Citation

MLA
Kobayashi, S., et al. “Decomposing NeRF for Editing via Feature Field Distillation”. arXiv, 2022, http://arxiv.org/abs/2205.15585v2.
APA
Kobayashi, S., Matsumoto, E., & Sitzmann, V. (2022). Decomposing NeRF for Editing via Feature Field Distillation. arXiv. http://arxiv.org/abs/2205.15585v2
Chicago
Kobayashi, S., E. Matsumoto, and V. Sitzmann. 2022. “Decomposing NeRF for Editing via Feature Field Distillation”. arXiv. http://arxiv.org/abs/2205.15585v2.
Harvard
Kobayashi, S., Matsumoto, E. and Sitzmann, V. (2022) “Decomposing NeRF for Editing via Feature Field Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.15585v2.
Vancouver
1. Kobayashi S, Matsumoto E, Sitzmann V (2022) Decomposing NeRF for Editing via Feature Field Distillation. arXiv

BibTeX

@article{kobayashi2022decomposing,
  title = {Decomposing NeRF for Editing via Feature Field Distillation},
  author = {Kobayashi, Sosuke and Matsumoto, Eiichi and Sitzmann, Vincent},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.15585v2},
  eprint = {2205.15585}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission