Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields

Shijie ZhouHaoran ChangSicheng JiangZhiwen FanZehao ZhuDejia XuPradyumna ChariSuya YouZhangyang WangAchuta Kadambi

article2024CVPR516 citationsCVPR 2024 Highlight

Presents an efficient framework for distilling 2D foundation model features into 3D Gaussian Splatting representations, enabling real-time novel-view semantic segmentation, interactive point prompting, and language-guided 3D scene editing.

Listen

Modern 3D computer vision and graphics have increasingly shifted toward creating interactive representations of physical scenes. While early methods based on neural radiance fields excelled at rendering novel camera views, extending them to understand semantics—such as recognizing, segmenting, or editing specific objects—has been severely hindered by slow rendering speeds, heavy computational costs, and visual artifacts. Although explicit representations like 3D Gaussian Splatting have delivered fast, high-quality visual rendering, they natively lack the semantic feature embeddings necessary for downstream scene understanding and automated editing.

The article demonstrates an explicit 3D scene representation framework, termed Feature 3DGS, that combines 3D Gaussian Splatting with feature field distillation from large 2D vision foundation models. The primary objective is to enable fast, high-fidelity novel view synthesis alongside promptable semantic segmentation and language-guided 3D scene editing.

To achieve this, the authors extend each 3D Gaussian point to store arbitrary-dimensional semantic features alongside spatial, opacity, and color parameters. The framework distills knowledge from pre-trained 2D vision models, specifically Segment Anything (SAM) and LSeg, using a parallel N-dimensional rasterizer that renders color and semantic feature maps jointly. To avoid the computational bottleneck of rasterizing high-dimensional features directly, the pipeline learns compact, low-dimensional feature vectors that are subsequently upsampled using an optional, lightweight convolutional speed-up module.

The evaluation revealed several key findings across standard synthetic and real-world benchmark datasets. First, the proposed approach achieved a 23% improvement in mean intersection-over-union for semantic segmentation compared to neural radiance field baselines, while improving overall visual rendering quality. Second, the framework operated up to 2.7 times faster in distillation and rendering than prior neural feature distillation methods, with the speed-up module more than doubling rendering frame rates to 14.55 frames per second on test benchmarks. Third, leveraging distilled 2D features allowed direct novel-view instance segmentation via prompt points and bounding boxes at speeds up to 1.7 times faster than processing full 2D images. Finally, the framework successfully executed text-driven 3D modifications, including clean object removal, object extraction across occluded views, and localized appearance recoloring without distorting background structures.

These findings indicate that explicit point-based 3D representations can effectively retain both photorealistic visual fidelity and rich semantic meaning without the trade-offs typical of implicit neural networks. By decoupling feature extraction from 2D image pipelines and enabling real-time rendering, this method significantly reduces compute latency and operational costs for interactive applications in augmented reality, virtual reality, robotics, and digital content creation.

Stakeholders and practitioners deploying interactive 3D applications should consider adopting explicit feature-distilled splatting architectures over traditional implicit neural radiance fields when real-time interaction is required. Organizations should pilot the optional convolutional speed-up module to balance feature fidelity with frame-rate performance for resource-constrained edge systems. Future development should focus on refining the base splatting pipeline to mitigate background noise artifacts and improving the quality of teacher foundation models to reduce downstream labeling errors.

arXiv: 2312.03203
Cover for Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields

Abstract

3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times, some work has emerged that aims to extend the functionality of NeRF beyond view synthesis, for semantically aware tasks such as editing and segmentation using 3D feature field distillation from 2D foundation models. However, these methods have two major limitations: (a) they are limited by the rendering speed of NeRF pipelines, and (b) implicitly represented feature fields suffer from continuity artifacts reducing feature quality. Recently, 3D Gaussian Splatting has shown state-of-the-art performance on real-time radiance field rendering. In this work, we go one step further: in addition to radiance field rendering, we enable 3D Gaussian splatting on arbitrary-dimension semantic features via 2D foundation model distillation. This translation is not straightforward: naively in-

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Implicit Radiance Field Representations
  • 2.2. Explicit Radiance Field Representations
  • 2.3. Feature Field Distillation
  • 3. Method
  • 3.1. High-dimensional Semantic Feature Rendering
  • 3.2. Optimization and Speed-up
  • 3.3. Promptable Explicit Scene Representation
  • 4. Experiments
  • 4.1. Novel view semantic segmentation
  • 4.2. Segment Anything from Any View
  • 4.3. Language-guided Editing
  • 5. Discussion and Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Parameterization of Explicit 3D Gaussians for Distilled Feature Fields

    model/method

    In Feature 3DGS, the 3D scene is represented by an explicit collection of anisotropic 3D Gaussians initialized from sparse Structure from Motion (SfM) point clouds. Each individual 3D Gaussian ii is parameterized by a tuple of optimizable attributes:

    Θi={xi,qi,si,αi,ci,fi}\Theta_i = \{x_i, q_i, s_i, \alpha_i, c_i, f_i\}

    where:

    • xi∈R3x_i \in \mathbb{R}^3 denotes the 3D spatial center position of the Gaussian.
    • qi∈R4q_i \in \mathbb{R}^4 is a unit quaternion representing the 3D rotation matrix RR.
    • si∈R3s_i \in \mathbb{R}^3 is a scaling vector defining the diagonal scaling matrix SS.
    • αi∈R\alpha_i \in \mathbb{R} represents the opacity of the Gaussian.
    • ci∈R3c_i \in \mathbb{R}^3 is the diffuse color representation (or up to 4 bands of spherical harmonics coefficients optimized incrementally).
    • fi∈RNf_i \in \mathbb{R}^N is an arbitrary NN-dimensional semantic feature embedding vector that captures latent foundation model representations.

    The 3D spatial covariance matrix Σ\Sigma is parameterized to remain positive semi-definite via the decomposition Σ=RSSTRT\Sigma = R S S^T R^T. When projected into the 2D image plane under a world-to-camera matrix WW and the Jacobian JJ of the projective transformation, the 2D covariance matrix is given by:

    Σ′=JWΣWTJT\Sigma' = J W \Sigma W^T J^T

  2. Knowl 2 — Joint Volumetric Alpha-Blending for Color and Semantic Feature Fields

    equation

    To render color and high-dimensional semantic features simultaneously, Feature 3DGS applies point-based α\alpha-blending in front-to-back depth order across the sorted set N\mathcal{N} of 3D Gaussians overlapping a target pixel:

    C=∑i∈NciαiTi,Fs=∑i∈NfiαiTiC = \sum_{i \in \mathcal{N}} c_i \alpha_i T_i, \quad F_s = \sum_{i \in \mathcal{N}} f_i \alpha_i T_i

    where:

    • C∈R3C \in \mathbb{R}^3 is the rendered pixel color.
    • Fs∈RNF_s \in \mathbb{R}^N is the rendered student semantic feature vector at the pixel.
    • ci∈R3c_i \in \mathbb{R}^3 and fi∈RNf_i \in \mathbb{R}^N are the color and semantic feature vectors stored at the ii-th Gaussian.
    • αi∈[0,1]\alpha_i \in [0, 1] is the learned opacity of the ii-th Gaussian.
    • Ti=∏j=1i−1(1−αj)T_i = \prod_{j=1}^{i-1} (1 - \alpha_j) is the cumulative transmittance representing the probability that the ray reaches Gaussian ii without prior occlusion.

    Both color and feature maps share the same tile-based parallel rasterization procedure (e.g., 16×1616 \times 16 pixel tiles) to preserve per-pixel spatial fidelity and geometric alignment between appearance and semantics.

  3. Knowl 3 — Joint Photometric and Feature Distillation Loss

    equation

    The optimization objective in Feature 3DGS combines a photometric radiance loss with a semantic feature distillation loss:

    L=Lrgb+γLf\mathcal{L} = \mathcal{L}_{\text{rgb}} + \gamma \mathcal{L}_f

    where the individual loss terms are defined as:

    Lrgb=(1−λ)L1(I,I^)+λLD-SSIM(I,I^)\mathcal{L}_{\text{rgb}} = (1 - \lambda) \mathcal{L}_1(I, \hat{I}) + \lambda \mathcal{L}_{\text{D-SSIM}}(I, \hat{I})

    Lf=∥Ft(I)−Fs(I^)∥1\mathcal{L}_f = \| F_t(I) - F_s(\hat{I}) \|_1

    Here:

    • I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3} is the ground truth RGB image, and I^\hat{I} is the rendered image.
    • Ft(I)∈RH×W×MF_t(I) \in \mathbb{R}^{H \times W \times M} is the teacher feature embedding produced by the 2D foundation model (e.g., LSeg or SAM) on ground truth image II.
    • Fs(I^)F_s(\hat{I}) is the rendered student feature map resized to match the spatial dimensions H×WH \times W via bilinear interpolation.
    • λ∈[0,1]\lambda \in [0, 1] balances the L1\mathcal{L}_1 pixel difference and structural dissimilarity LD-SSIM\mathcal{L}_{\text{D-SSIM}} (set to λ=0.2\lambda = 0.2).
    • γ\gamma weights the feature loss against the RGB loss (set to γ=1.0\gamma = 1.0). Unlike implicit NeRF representations which require tiny γ\gamma values to avoid interference between radiance and semantic branches, explicit 3DGS allows equal weighting (γ=1.0\gamma = 1.0) without degrading novel view visual quality.
  4. Knowl 4 — Convolutional Speed-Up Module for High-Dimensional Feature Rasterization

    model/method

    Directly rendering high-dimensional semantic feature maps where MM is large (e.g., M=512M=512 for LSeg, M=256M=256 for SAM) significantly reduces training and rendering frame rates. To mitigate this computational bottleneck, Feature 3DGS introduces an optional speed-up module:

    1. Each 3D Gaussian is initialized with a lower-dimensional feature vector fi∈RNf_i \in \mathbb{R}^N where N≪MN \ll M (e.g., N=128N=128).
    2. The parallel rasterizer renders an intermediate low-dimensional feature map Fslow∈RHfeat×Wfeat×NF_s^{\text{low}} \in \mathbb{R}^{H_{\text{feat}} \times W_{\text{feat}} \times N}.
    3. A lightweight, learnable 1×11 \times 1 convolutional decoder upsamples the feature channel dimension from NN to the teacher foundation model dimension MM, producing Fs(I^)∈RH×W×MF_s(\hat{I}) \in \mathbb{R}^{H \times W \times M}.

    Because the 1×11 \times 1 convolution operates on the rasterized 2D feature map rather than per-Gaussian 3D points, it introduces minimal computational overhead while enabling learnable cross-channel communication, yielding up to 2.7×2.7\times faster training and rendering without sacrificing downstream segmentation accuracy.

  5. Knowl 5 — Promptable 3D Gaussian Filtering and Scene Editing

    algorithm

    Language- and point-prompted manipulation is performed directly in the explicit 3D Gaussian domain using semantic query matching:

    Input: Set of 3D Gaussians with parameters G = {(x_i, q_i, s_i, alpha_i, c_i, f_i)}_{i=1}^N, query prompt tau (text or point), prompt candidate set T
    Output: Edited set of 3D Gaussians G_edited
    Extract feature query embedding q(tau) using foundation text/point encoder
    for each 3D Gaussian x in G do
        Compute cosine similarity:
        s(x, tau) = (f(x) . q(tau)) / (||f(x)|| * ||q(tau)||)
        for each label j in T do
            Compute s(x, j) = (f(x) . q(j)) / (||f(x)|| * ||q(j)||)
        end for
        Compute prompt probability via softmax:
        p(tau | x) = exp(s(x, tau)) / sum_{j in T} exp(s(x, j))
    end for
    Define 3D Gaussian mask M_G based on probability scores:
    if Hard Selection then
        M_G(x) = 1 if p(tau | x) == max_{j in T} p(j | x) else 0
    else if Soft Selection then
        M_G(x) = 1 if p(tau | x) >= threshold else 0
    end if
    for each 3D Gaussian x in G do
        if M_G(x) == 1 then
            Apply target operation: Extract, Delete (set alpha(x) = 0), or Recolor (update c(x))
        end if
    end for
    return Edited Gaussians G_edited
  6. Knowl 6 — Novel View Promptable Segmentation via 3D Distilled Foundation Decoders

    model/method

    Traditional foundation model application to novel view 3D scenes requires a two-step inference process: first synthesizing an RGB image at the target camera viewpoint, and then executing the complete heavy encoder-decoder pipeline of the 2D foundation model (such as the Segment Anything Model, SAM).

    Feature 3DGS eliminates the heavy 2D image encoder during inference by rendering the distilled SAM feature map FsF_s directly from the 3D Gaussian field for any specified camera pose. Point or bounding-box prompt coordinates are passed directly into the lightweight SAM mask decoder alongside the rendered 2D feature map. This bypasses the foundation model image encoder entirely, achieving up to a 1.7×1.7\times speedup in total rendering and segmentation latency while preserving segmentation mask fidelity.

  7. Knowl 7 — Semantic Segmentation and Radiance Quality on the Replica Benchmark

    data/table

    Novel view synthesis and semantic segmentation performance on the Replica indoor dataset after 5,000 training iterations, comparing standard 3D Gaussian Splatting (Base 3DGS), Distilled Feature Fields on NeRF (NeRF-DFF), and Feature 3DGS (with and without the 1×11 \times 1 convolutional speed-up module where feature dimension N=128N=128):

    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow mIoU ↑\uparrow Accuracy ↑\uparrow FPS ↑\uparrow
    Base 3DGS 36.133 0.965 0.033 - - -
    NeRF-DFF - - - 0.636 0.864 5.38
    Feature 3DGS (Full) 36.915 0.970 0.024 0.787 0.943 6.84
    Feature 3DGS (w/ speed-up) 37.012 0.971 0.023 0.782 0.943 14.55

    These results demonstrate that distilling semantic feature fields into 3D Gaussians improves novel view radiance quality (PSNR increases from 36.133 to 37.012 dB), while achieving a 23.7% relative improvement in mIoU (0.787 vs. 0.636) and over 2.7×2.7\times higher rendering speed (14.55 FPS vs. 5.38 FPS) compared to NeRF-DFF.

  8. Knowl 8 — Language-Guided 3D Object Extraction, Deletion, and Appearance Modification

    empirical result

    Distilling LSeg features into 3D Gaussians enables zero-shot language-guided editing directly in 3D explicit space using text prompts encoded with CLIP (ViT-B/32):

    • Object Extraction: Text queries such as "extract the banana" isolate full 3D object geometry, including occluded regions (e.g., reconstructing the entire banana geometry even when partially occluded by an apple from specific viewpoints) with significantly fewer background floater artifacts than NeRF-DFF.
    • Object Deletion: Text queries such as "delete the car" set the opacity α\alpha of corresponding Gaussians to zero, cleanly removing target foreground geometry while preserving occluded background geometry (such as vegetation behind the car).
    • Appearance Modification: Queries targeting semantic categories (e.g., "change sidewalk color", "change leaves color") selectively modify diffuse color parameters cc of Gaussians matching the target semantic score without modifying adjacent semantic objects (such as retaining the red stop sign color).
  9. Knowl 9 — Limitations of Feature 3DGS

    limitation

    The Feature 3DGS framework exhibits two primary limitations:

    1. Teacher Model Dependency: The quality of the learned 3D student feature field is strictly bounded by the accuracy, resolution limits, and semantic noise of the pre-trained 2D teacher foundation models (such as CLIP-LSeg or SAM).
    2. Floater Artifacts from Explicit Splatting: The underlying 3D Gaussian Splatting rasterization pipeline inherently generates isolated, out-of-surface Gaussian floaters. These floaters can acquire noisy semantic features and introduce visual or segmentation artifacts during novel view rendering and language-based Gaussian filtering.

Coverage note — None was omitted. All key contributions, mathematical formulations (Gaussian feature parameterization, joint rasterization, distillation loss), architectural components (speed-up module, promptable explicit representations), experimental benchmark tables, and stated limitations have been fully captured.

References

  1. 1.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. ICCV, 2021. 2
  2. 2.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. ICCV, 2023. 2
  3. 3.Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Perez. Zero-shot semantic segmentation. Advances in Neural Information Processing Systems, 32, 2019. 4
  4. 4.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 3
  5. 5.Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al. Efficient geometry-aware 3d generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16123–16133, 2022. 3
  6. 6.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14124–14133, 2021. 3
  7. 7.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision, pages 333–350. Springer, 2022. 3
  8. 8.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111:98–136, 2015. 5
  9. 9.Zhiwen Fan, Peihao Wang, Yifan Jiang, Xinyu Gong, Dejia Xu, and Zhangyang Wang. Nerf-sos: Any-view self-supervised object segmentation on complex scenes. arXiv preprint arXiv:2209.08776, 2022. 2, 3
  10. 10.Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12479–12488, 2023. 3
  11. 11.Rahul Goel, Dhawal Sirikonda, Saurabh Saini, and PJ Narayanan. Interactive segmentation of radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4201–4211, 2023. 3
  12. 12.Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang. Context-aware feature generation for zero-shot semantic segmentation. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1921–1929, 2020. 4
  13. 13.Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering. ACM Transactions on Graphics (ToG), 37(6):1–15, 2018. 1, 7
  14. 14.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 4
  15. 15.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023. 2, 4, 6
  16. 16.Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embedded radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19729–19739, 2023. 2, 3, 8
  17. 17.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 4, 5
  18. 18.Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann. Decomposing nerf for editing via feature field distillation. In NeurIPS, 2022. 2, 3, 5, 6, 7, 8
  19. 19.Georgios Kopanas, Julien Philip, Thomas Leimkuhler, and George Drettakis. Point-based neural rendering with per-view optimization. In Computer Graphics Forum, pages 29–43. Wiley Online Library, 2021. 4
  20. 20.Boyi Li, Kilian Q. Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation, 2022. 4, 5, 6
  21. 21.Xiang Li, Tianhan Wei, Yau Pun Chen, Yu-Wing Tai, and Chi-Keung Tang. Fss-1000: A 1000-class dataset for few-shot segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2869–2878, 2020. 5
  22. 22.Stefan Lionar, Xiangyu Xu, Min Lin, and Gim Hee Lee. Nu-mcc: Multiview compressive coding with neighborhood decoder and repulsive udf. arXiv preprint arXiv:2307.09112, 2023. 3
  23. 23.Kunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu, Yingchen Yu, Abdulmotaleb El Saddik, Christian Theobalt, Eric Xing, and Shijian Lu. Weakly supervised 3d open-vocabulary segmentation. Advances in Neural Information Processing Systems, 36, 2024. 3
  24. 24.Kirill Mazur, Edgar Sucar, and Andrew J Davison. Feature-realistic neural fusion for real-time, open set scene understanding. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 8201–8207. IEEE, 2023. 3
  25. 25.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4460–4470, 2019. 2
  26. 26.Ben Mildenhall, Pratul P Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 38(4):1–14, 2019. 6
  27. 27.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 2, 3
  28. 28.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 2, 3
  29. 29.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 165–174, 2019. 2
  30. 30.Fabian Pedregosa, Gaèl Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830, 2011. 7
  31. 31.Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16, pages 523–540. Springer, 2020. 2
  32. 32.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 5
  33. 33.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12179–12188, 2021. 5
  34. 34.Zhongzheng Ren, Aseem Agarwala†, Bryan Russell†, Alexander G. Schwing†, and Oliver Wang†. Neural volumetric object selection. In CVPR, 2022. († alphabetic ordering). 3
  35. 35.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016. 4
  36. 36.Nur Muhammad Mahi Shafiullah, Chris Paxton, Lerrel Pinto, Soumith Chintala, and Arthur Szlam. Clip-fields: Weakly supervised semantic fields for robotic memory. arXiv preprint arXiv:2210.05663, 2022. 3
  37. 37.William Shen, Ge Yang, Alan Yu, Jansen Wong, Leslie Pack Kaelbling, and Phillip Isola. Distilled feature fields enable few-shot language-guided manipulation. In 7th Annual Conference on Robot Learning, 2023. 3
  38. 38.Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Norman Müller, Matthias Nießner, Angela Dai, and Peter Kontschieder. Panoptic lifting for 3d scene understanding with neural fields. arXiv preprint arXiv:2212.09802, 2022. 3
  39. 39.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. The replica dataset: A digital replica of indoor spaces, 2019. 6
  40. 40.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8248–8258, 2022. 3
  41. 41.Nikolaos Tsagkas, Oisin Mac Aodha, and Chris Xiaoxuan Lu. Vl-fields: Towards language-grounded neural implicit spatial representations. In 2023 IEEE International Conference on Robotics and Automation. IEEE, 2023. 3
  42. 42.Vadim Tschernezki, Iro Laina, Diane Larlus, and Andrea Vedaldi. Neural Feature Fusion Fields: 3D distillation of self-supervised 2D image representations. In 3DV, 2022. 2, 3
  43. 43.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In CVPR, 2021. 3
  44. 44.Zhen Wang, Shijie Zhou, Jeong Joon Park, Despoina Paschalidou, Suya You, Gordon Wetzstein, Leonidas Guibas, and Achuta Kadambi. Alto: Alternating latent topologies for implicit 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 259–270, 2023. 2
  45. 45.Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5438–5448, 2022. 3
  46. 46.Jianglong Ye, Naiyan Wang, and Xiaolong Wang. Featurenerf: Learning generalizable nerfs by distilling foundation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8962–8973, 2023. 3
  47. 47.Brent Yi, Weijia Zeng, Sam Buchanan, and Yi Ma. Canonical factors for hybrid neural fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3414–3426, 2023. 3
  48. 48.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In CVPR, 2021. 3
  49. 49.Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and Andrew Davison. In-place scene labelling and understanding with implicit scene representation. In ICCV, 2021. 3
  50. 50.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision, 127:302–321, 2019. 5
  51. 51.Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Ewa volume splatting. In Proceedings Visualization, 2001. VIS’01., pages 29–538. IEEE, 2001. 4

Citation

MLA
Zhou, S., et al. “Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields”. arXiv, 2023, http://arxiv.org/abs/2312.03203v3.
APA
Zhou, S., Chang, H., Jiang, S., Fan, Z., Zhu, Z., Xu, D., Chari, P., You, S., Wang, Z., & Kadambi, A. (2023). Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields. arXiv. http://arxiv.org/abs/2312.03203v3
Chicago
Zhou, S., H. Chang, S. Jiang, et al. 2023. “Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields”. arXiv. http://arxiv.org/abs/2312.03203v3.
Harvard
Zhou, S. et al. (2023) “Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.03203v3.
Vancouver
1. Zhou S, Chang H, Jiang S, Fan Z, Zhu Z, Xu D, Chari P, You S, Wang Z, Kadambi A (2023) Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields. arXiv

BibTeX

@article{zhou2023feature,
  title = {Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields},
  author = {Zhou, Shijie and Chang, Haoran and Jiang, Sicheng and Fan, Zhiwen and Zhu, Zehao and Xu, Dejia and Chari, Pradyumna and You, Suya and Wang, Zhangyang and Kadambi, Achuta},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.03203v3},
  eprint = {2312.03203}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE