Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
Shijie ZhouHaoran ChangSicheng JiangZhiwen FanZehao ZhuDejia XuPradyumna ChariSuya YouZhangyang WangAchuta Kadambi
Presents an efficient framework for distilling 2D foundation model features into 3D Gaussian Splatting representations, enabling real-time novel-view semantic segmentation, interactive point prompting, and language-guided 3D scene editing.
Modern 3D computer vision and graphics have increasingly shifted toward creating interactive representations of physical scenes. While early methods based on neural radiance fields excelled at rendering novel camera views, extending them to understand semantics—such as recognizing, segmenting, or editing specific objects—has been severely hindered by slow rendering speeds, heavy computational costs, and visual artifacts. Although explicit representations like 3D Gaussian Splatting have delivered fast, high-quality visual rendering, they natively lack the semantic feature embeddings necessary for downstream scene understanding and automated editing.
The article demonstrates an explicit 3D scene representation framework, termed Feature 3DGS, that combines 3D Gaussian Splatting with feature field distillation from large 2D vision foundation models. The primary objective is to enable fast, high-fidelity novel view synthesis alongside promptable semantic segmentation and language-guided 3D scene editing.
To achieve this, the authors extend each 3D Gaussian point to store arbitrary-dimensional semantic features alongside spatial, opacity, and color parameters. The framework distills knowledge from pre-trained 2D vision models, specifically Segment Anything (SAM) and LSeg, using a parallel N-dimensional rasterizer that renders color and semantic feature maps jointly. To avoid the computational bottleneck of rasterizing high-dimensional features directly, the pipeline learns compact, low-dimensional feature vectors that are subsequently upsampled using an optional, lightweight convolutional speed-up module.
The evaluation revealed several key findings across standard synthetic and real-world benchmark datasets. First, the proposed approach achieved a 23% improvement in mean intersection-over-union for semantic segmentation compared to neural radiance field baselines, while improving overall visual rendering quality. Second, the framework operated up to 2.7 times faster in distillation and rendering than prior neural feature distillation methods, with the speed-up module more than doubling rendering frame rates to 14.55 frames per second on test benchmarks. Third, leveraging distilled 2D features allowed direct novel-view instance segmentation via prompt points and bounding boxes at speeds up to 1.7 times faster than processing full 2D images. Finally, the framework successfully executed text-driven 3D modifications, including clean object removal, object extraction across occluded views, and localized appearance recoloring without distorting background structures.
These findings indicate that explicit point-based 3D representations can effectively retain both photorealistic visual fidelity and rich semantic meaning without the trade-offs typical of implicit neural networks. By decoupling feature extraction from 2D image pipelines and enabling real-time rendering, this method significantly reduces compute latency and operational costs for interactive applications in augmented reality, virtual reality, robotics, and digital content creation.
Stakeholders and practitioners deploying interactive 3D applications should consider adopting explicit feature-distilled splatting architectures over traditional implicit neural radiance fields when real-time interaction is required. Organizations should pilot the optional convolutional speed-up module to balance feature fidelity with frame-rate performance for resource-constrained edge systems. Future development should focus on refining the base splatting pipeline to mitigate background noise artifacts and improving the quality of teacher foundation models to reduce downstream labeling errors.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Introduces 3D Gaussian Splatting for real-time radiance field rendering, providing the foundational representation and differentiable rendering pipeline that Feature 3DGS equips with distilled feature fields.
- Paper: Decomposing NeRF for Editing via Feature Field Distillation, Sosuke Kobayashi et al. (2022). Establishes the concept of distilling 2D foundation model features into 3D neural fields for semantic scene decomposition and editing, setting up the paradigm that Feature 3DGS translates to explicit Gaussians.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Provides the seminal neural radiance field and volume rendering formulation that motivated subsequent 3D feature distillation and faster explicit rendering alternatives.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Pioneers the use of continuous 3D neural feature fields for rendering and scene decomposition beyond pure RGB color fields.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Demonstrates instruction-based 3D scene editing driven by 2D image models, contextualizing the need for real-time semantic manipulation enabled by distilled feature fields.
- Paper: GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting, Yiwen Chen et al. (2024). Applies Gaussian Splatting and feature-guided editing to enable swift, controllable 3D scene manipulation and semantic object modifications.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). Extends Gaussian Splatting toward generative text-to-3D synthesis by distilling 2D diffusion guidance directly into 3D Gaussian representations.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). Advances open-set 3D semantic understanding and zero-shot novel view distillation, continuing the study of embedding 2D vision-language features into 3D spatial fields.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). Replaces 3D volumetric Gaussians with planar 2D oriented splats to substantially improve the underlying geometric fidelity and surface reconstructions of splatting architectures.
- Paper: Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering, Tao Lu et al. (2024). Enhances the efficiency and robustness of Gaussian Splatting via anchor-based spatial structures that alleviate geometry fitting artifacts across varied views.
- Paper: GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction, Yuanhui Huang 0002 et al. (2025). Applies semantic 3D Gaussian representations to downstream autonomous driving tasks through probabilistic superposition for 3D occupancy prediction.
