Decomposing NeRF for Editing via Feature Field Distillation
Sosuke KobayashiEiichi MatsumotoVincent Sitzmann
Proposes distilling 2D foundation model representations into 3D feature fields alongside neural radiance fields, enabling query-based 3D semantic segmentation and localized scene editing without retraining.
Neural radiance fields (NeRF) have emerged as a powerful technique for constructing high-fidelity 3D scene representations and generating novel viewpoints from 2D imagery. However, standard NeRF models implicitly encode entire scenes into interconnected network weights, making selective manipulation or object-centric editing difficult. Prior segmentation methods typically rely on expensive manual annotations or restricted closed-set categories, which limits practical interactive editing. The article addresses this operational bottleneck by introducing distilled feature fields (DFF), a method that transfers semantic understanding into 3D scenes without requiring manual 3D annotations.
The main objective of the article is to demonstrate that knowledge from pre-trained 2D vision and language models can be distilled into 3D feature fields to enable query-based zero-shot semantic decomposition and local scene editing. The authors formulate a framework where a 3D feature field is optimized alongside standard radiance parameters via differentiable volume rendering, supervised directly by pre-trained 2D teacher models such as LSeg for text queries and DINO for visual image patch queries. To evaluate the framework, the authors conducted quantitative segmentation benchmarks on indoor scenes from the Replica dataset and performed qualitative appearance, geometry, and optimization-based editing experiments across complex real-world captures.
The evaluation revealed several critical findings. First, DFF achieved superior 3D semantic segmentation performance on the Replica benchmark compared to a supervised 3D convolutional baseline trained on the ScanNet dataset, improving mean intersection-over-union from 0.475 to 0.589 and overall accuracy from 75.8% to 85.5%. Second, adding a feature branch to the neural field introduced minimal computational overhead and preserved visual reconstruction fidelity, showing essentially identical novel view synthesis quality (32.85 dB peak signal-to-noise ratio versus 32.87 dB for baseline NeRF). Third, the approach enabled precise, view-consistent local edits—including color alterations, deletions, and geometric translations—driven by text prompts, image patches, or point selections. Fourth, when integrated with text-driven optimization frameworks like CLIPNeRF, DFF prevented unintended global style alterations by successfully confining modifications strictly to designated target regions.
These findings indicate that 3D neural graphics can leverage off-the-shelf 2D foundational models to achieve open-vocabulary, interactive scene editing without the cost of collecting specialized 3D training datasets. By eliminating the need to retrain models for new semantic categories, this method significantly reduces workflow friction and computation costs for downstream graphics, spatial computing, and digital asset creation. Furthermore, it demonstrates that multi-view volume rendering inherently denoises 2D features, producing robust 3D segmentation fields that outperform discrete point-cloud baselines.
Organizations developing neural rendering pipelines should adopt feature field distillation to enable flexible scene manipulation and asset decomposition. Practitioners should utilize coarse sampling or remove high-frequency positional encodings in feature branches to enforce spatial smoothness and reduce boundary artifacts. For broader integration, lightweight independent feature networks can be deployed to retrofit existing, off-the-shelf NeRF models or modern hash-grid representations without retraining underlying geometry. Future research should prioritize enhancing background reconstruction behind removed objects and developing view-dependent referring mechanisms to support spatial language queries.
Confidence in the reported semantic segmentation and local editing capabilities is high, backed by both quantitative benchmarks and diverse real-world visual demonstrations. Nonetheless, key operational constraints remain. The student representation is ultimately bounded by the resolution and vocabulary limits of its 2D teacher models, and underlying geometric reconstruction artifacts in the radiance field can introduce noise into the feature supervision. Additionally, deleting prominent foreground objects can occasionally expose blurred or incomplete background geometry where camera views were occluded.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces the foundational Neural Radiance Fields (NeRF) formulation and differentiable volume rendering framework upon which the feature field distillation and scene representation are built.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Pioneers the concept of neural feature fields within volume rendering to achieve compositional and controllable 3D scene representations.
- Paper: Instant neural graphics primitives with a multiresolution hash encoding, Thomas Müller et al. (2022). Establishes multiresolution hash encodings for fast neural graphics primitives, which provide the underlying acceleration structure used to optimize neural fields efficiently.
- Paper: Unsupervised Semantic Segmentation by Distilling Feature Correspondences, Mark Hamilton et al. (2022). Demonstrates how distilling self-supervised feature correspondences enables semantic segmentation without manual supervision, providing key conceptual foundations for zero-shot 2D feature transfer.
- Paper: Neural Sparse Voxel Fields, Lingjie Liu et al. (2020). Presents sparse voxel-bounded neural fields that enable localized evaluation and discrete object manipulation in neural implicit representations.
- Paper: NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, Ricardo Martin-Brualla et al. (2021). Introduces the decomposition of neural scene fields into separate components to handle transient and static scene elements.
- Paper: Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields, Shijie Zhou et al. (2024). Adapts the concept of 2D foundation model feature field distillation from implicit NeRFs to explicit 3D Gaussian Splatting for real-time semantic segmentation and editing.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). Extends zero-shot feature field distillation in NeRF by introducing active novel-view synthesis to resolve cross-view semantic ambiguities in uncertain regions.
- Paper: GARField: Group Anything with Radiance Fields, Chung Min Kim et al. (2024). Builds on semantic field decomposition by learning hierarchical 3D grouping conditioned on physical scale using 2D foundation model masks.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). Advances semantic lifting in neural fields by resolving cross-view 2D instance inconsistencies to construct unified 3D panoptic representations.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Explores an alternative instruction-driven 3D scene editing paradigm using iterative 2D diffusion image editing rather than explicit feature field decomposition.
- Paper: GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting, Yiwen Chen et al. (2024). Applies localized 2D mask projections and semantic tracing to enable rapid object-centric editing directly within 3D Gaussian Splatting representations.
- Paper: DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis, Yinghao Xu et al. (2023). Extends compositional scene editing into a generative framework by using spatial bounding-box priors to disentangle multi-object radiance fields.
