OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

Francis EngelmannFabian ManhardtMichael NiemeyerKeisuke TatenoFederico Tombari

article2024ICLR82 citations

Presents an open-vocabulary 3D scene segmentation framework that encodes pixel-wise vision-language features directly into neural radiance fields and uses novel-view synthesis to classify unobserved regions, achieving superior accuracy over existing models like LERF and OpenScene.

Listen

Modern autonomous systems, including augmented reality devices and service robots, require a detailed understanding of their three-dimensional environments. Traditional computer vision methods rely on supervised learning with fixed, predefined categories, limiting an agent's ability to identify novel objects or adapt to changing environments without costly manual retraining. While recent vision-language models allow open-vocabulary recognition in two-dimensional images, lifting this capability into complex 3D scenes remains challenging due to the resolution limits of standard 3D meshes and difficulties in aligning 2D text-image features with 3D space.

The article demonstrates OpenNeRF, an implicit neural radiance field framework designed for open-set 3D semantic segmentation. The objective is to evaluate whether directly embedding dense, pixel-aligned vision-language features into a continuous neural volumetric representation—augmented by synthesizing views of ambiguous areas—outperforms existing explicit mesh-based and implicit methods on zero-shot 3D understanding.

To achieve this, the approach distills 2D pixel-level vision-language embeddings from OpenSeg into a neural radiance field without requiring complex multi-scale sampling or auxiliary regularizations. When feature predictions from different camera angles disagree, the system calculates per-point uncertainty across the scene. It then generates targeted, novel camera viewpoints focused on these uncertain regions, renders synthetic images, extracts additional feature maps, and updates the neural representation. The authors evaluated this system on eight standardized indoor scenes from the Replica dataset across 51 semantic categories divided into head, common, and tail frequency distributions, while also testing on real-world smartphone scans.

The evaluation yielded several key findings. First, OpenNeRF achieved an overall 3D segmentation score of 20.4 mean intersection over union, outperforming the leading baseline OpenScene by 4.5 points and LERF by 9.9 points. Second, the performance advantage was most pronounced on rare and smaller tail categories, where OpenNeRF reached 5.8 mIoU compared to OpenScene's 1.5 mIoU, successfully segmenting small functional items such as wall plugs, clocks, and tissue boxes that baseline systems missed entirely. Third, the uncertainty-driven view synthesis mechanism directly enhanced accuracy, improving overall segmentation performance by 1.0 mIoU over fixed-view setups, whereas rendering from naive random viewpoints severely degraded performance to 15.4 mIoU.

These results demonstrate that continuous neural representations combined with targeted novel-view synthesis offer a more flexible, scalable alternative to explicit 3D mesh pipelines. By eliminating the need for expensive 3D semantic pre-training data and heavy architectural regularization, OpenNeRF reduces development complexity while enhancing zero-shot scene querying across objects, abstract physical properties, and material textures.

Stakeholders developing spatial computing or robotic platforms should consider adopting implicit neural representations with pixel-aligned feature distillation to improve interaction with novel items. Future work should focus on scaling the framework to larger outdoor environments, accelerating training and inference speeds for real-time deployment, and evaluating performance in dynamic settings.

Confidence in these findings is supported by consistent benchmark improvements and systematic ablation tests across multiple indoor rooms. However, readers should note that overall segmentation accuracy on rare, long-tail categories remains relatively low across all tested models, indicating that open-vocabulary 3D segmentation in complex environments is still an evolving field that warrants cautious deployment in safety-critical tasks.

arXiv: 2404.03650
Cover for OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

Abstract

Large visual-language models (VLMs), like CLIP, enable open-set image segmentation to segment arbitrary concepts from an image in a zero-shot manner. This goes beyond the traditional closed-set assumption, i.e., where models can only segment classes from a pre-defined training set. More recently, first works on open-set segmentation in 3D scenes have appeared in the literature. These methods are heavily influenced by closed-set 3D convolutional approaches that process point clouds or polygon meshes. However, these 3D scene representations do not align well with the image-based nature of the visual-language models. Indeed, point cloud and 3D meshes typically have a lower resolution than images and the reconstructed 3D scene geometry might not project well to the underlying 2D image sequences used to compute pixel-aligned CLIP features. To address these challenges, we propose OpenNeRF which naturally operates on posed images and directly encodes the VLM features within the NeRF. This is similar in spirit to LERF, however our work shows that using pixel-wise VLM features (instead of global CLIP features) results in an overall less complex architecture without the need for additional DINO regularization. Our OpenNeRF further leverages NeRF's ability to render novel views and extract open-set VLM features from areas that are not well observed in the initial posed images. For 3D point cloud segmentation on the Replica dataset, OpenNeRF outperforms recent open-vocabulary methods such as LERF and OpenScene by at least +4.9 mIoU.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Open-Set Radiance Fields
  • 3.2 Training Objectives
  • 3.3 Rendering Novel Views
  • 3.4 Implementation and Training Details
  • 4 Experiments
  • 4.1 Methods in comparison.
  • 4.2 Results on 3D Semantic Segmentation
  • 4.3 Analysis Experiments
  • 4.4 Qualitative Results for 3D Scene Segmentation and Open-Set Applications
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — OpenNeRF Scene Representation and Distillation Architecture

    model/method

    OpenNeRF represents a 3D scene as a continuous neural radiance field that predicts volume density σ\sigma, RGB color c\mathbf{c}, and an open-set semantic feature vector o\mathbf{o} for any given 3D position x∈R3\mathbf{x} \in \mathbb{R}^3 and viewing direction d∈S2\mathbf{d} \in \mathbb{S}^2:

    fθ(x,d)↦(σ,c,o)f_\theta(\mathbf{x}, \mathbf{d}) \mapsto (\sigma, \mathbf{c}, \mathbf{o})

    where θ\theta denotes the trainable network parameters, σ∈[0,∞)\sigma \in [0, \infty), c∈[0,1]3\mathbf{c} \in [0, 1]^3, and o∈RD\mathbf{o} \in \mathbb{R}^D is a DD-dimensional semantic descriptor (D=768D = 768). The backbone is parameterized following Mip-NeRF with positional encodings for high-frequency spatial reasoning, while an additional multi-layer perceptron (MLP) head branches off to output the open-set semantic feature vector o\mathbf{o}.

    Volume rendering calculates the color c^r\hat{\mathbf{c}}_r, expected depth d^r\hat{d}_r, and open-set feature vector o^r\hat{\mathbf{o}}_r along a camera ray r(t)=oc+td\mathbf{r}(t) = \mathbf{o}_c + t\mathbf{d} by numerically integrating samples along the ray according to volumetric transmission and density weights. The open-set branch is directly supervised using 2D pixel-aligned feature maps extracted from pre-trained OpenSeg models. By utilizing dense pixel-aligned features, OpenNeRF eliminates the need for multi-resolution patch sampling during training and scale-pyramid rendering during inference, and avoids auxiliary regularization terms such as self-supervised DINO feature matching.

  2. Knowl 2 — OpenNeRF Training Objectives and Optimization Constraints

    model/method

    OpenNeRF is optimized end-to-end on a set of sampled camera rays R\mathcal{R} using a composite loss function:

    L=LRGB+λopenLopen+λdepthLdepth\mathcal{L} = \mathcal{L}_{\text{RGB}} + \lambda_{\text{open}} \mathcal{L}_{\text{open}} + \lambda_{\text{depth}} \mathcal{L}_{\text{depth}}

    where λopen\lambda_{\text{open}} and λdepth\lambda_{\text{depth}} are loss weighting coefficients. The loss components are defined as follows:

    1. Appearance Loss: Evaluates the Euclidean reconstruction error between rendered color c^r\hat{\mathbf{c}}_r and ground-truth sensor color cr\mathbf{c}_r:

    LRGB=1∣R∣∑r∈R∥cr−c^r∥22\mathcal{L}_{\text{RGB}} = \frac{1}{|\mathcal{R}|} \sum_{r \in \mathcal{R}} \|\mathbf{c}_r - \hat{\mathbf{c}}_r\|_2^2

    1. Depth Loss: Minimizes the smooth L1L_1 (Huber) loss between the rendered expected ray depth d^r\hat{d}_r and observed depth sensor measurement drd_r where depth observations exist:

    Ldepth=1∣R∣∑r∈RHuber(dr−d^r)\mathcal{L}_{\text{depth}} = \frac{1}{|\mathcal{R}|} \sum_{r \in \mathcal{R}} \text{Huber}(d_r - \hat{d}_r)

    1. Open-Set Cosine Loss: Maximizes the cosine similarity between the normalized rendered feature vector o^r\hat{\mathbf{o}}_r and the normalized OpenSeg ground-truth feature vector or\mathbf{o}_r:

    Lopen=1∣R∣∑r∈R(−or∥or∥2⋅o^r∥o^r∥2)\mathcal{L}_{\text{open}} = \frac{1}{|\mathcal{R}|} \sum_{r \in \mathcal{R}} \left( - \frac{\mathbf{o}_r}{\|\mathbf{o}_r\|_2} \cdot \frac{\hat{\mathbf{o}}_r}{\|\hat{\mathbf{o}}_r\|_2} \right)

    Optimization Constraints:

    • Gradients from Lopen\mathcal{L}_{\text{open}} are detached from the density field σ\sigma to prevent 2D multi-view semantic inconsistencies from corrupting underlying 3D geometry.
    • Rays within a 10-pixel boundary from the image borders are masked out during training to suppress boundary artifacts produced by 2D open-vocabulary segmentation networks.
  3. Knowl 3 — Active Novel Viewpoint Selection via Open-Set Feature Uncertainty

    algorithm

    To improve open-vocabulary 3D representations in underexplored or ambiguous scene regions, OpenNeRF identifies regions with high semantic uncertainty across projected viewpoints and generates target-focused novel camera poses for synthetic rendering.

    Input: Posed RGB images {Ik}k=1N\{I_k\}_{k=1}^N, 2D OpenSeg feature maps {Fk}k=1N\{F_k\}_{k=1}^N, coarse point cloud P={pi}i=1MP = \{\mathbf{p}_i\}_{i=1}^M, NeRF model fθf_\theta.
    Output: Set of novel camera poses Pnovel\mathcal{P}_{\text{novel}}.
    1. for each point pi∈P\mathbf{p}_i \in P do
    2. Collect the set of projected feature vectors Si={Fk(πk(pi))}S_i = \{F_k(\pi_k(\mathbf{p}_i))\} for all views kk where pi\mathbf{p}_i is visible.
    3. Compute mean μi∈RD\boldsymbol{\mu}_i \in \mathbb{R}^D and covariance Σi∈RD×D\boldsymbol{\Sigma}_i \in \mathbb{R}^{D \times D} over SiS_i using Welford's single-pass online algorithm.
    4. Compute the uncertainty score ui=det⁡(Σi)u_i = \det(\boldsymbol{\Sigma}_i) as the generalized variance.
    5. end for
    6. Normalize uncertainty scores {ui}\{u_i\} to [0,1][0, 1].
    7. Initialize Pnovel←∅\mathcal{P}_{\text{novel}} \leftarrow \emptyset.
    8. for each point pi∈P\mathbf{p}_i \in P do
    9. Sample x∼Uniform(0,1)x \sim \text{Uniform}(0, 1).
    10. if x<uix < u_i then
    11. Set target point t←pi\mathbf{t} \leftarrow \mathbf{p}_i.
    12. Set camera center ccam←t+voffset+ϵnoise\mathbf{c}_{\text{cam}} \leftarrow \mathbf{t} + \mathbf{v}_{\text{offset}} + \boldsymbol{\epsilon}_{\text{noise}}, where voffset\mathbf{v}_{\text{offset}} is a fixed offset into the scene and ϵnoise\boldsymbol{\epsilon}_{\text{noise}} is random jitter.
    13. Query NeRF density σ(ccam)\sigma(\mathbf{c}_{\text{cam}}) and along the line-of-sight from ccam\mathbf{c}_{\text{cam}} to t\mathbf{t}.
    14. if σ(ccam)≈0\sigma(\mathbf{c}_{\text{cam}}) \approx 0 (free space) and t\mathbf{t} is unobstructed then
    15. Compute camera pose matrix T=LookAt(eye=ccam,target=t,up=[0,1,0]T)T = \text{LookAt}(\text{eye}=\mathbf{c}_{\text{cam}}, \text{target}=\mathbf{t}, \text{up}=[0, 1, 0]^T).
    16. Pnovel←Pnovel∪{T}\mathcal{P}_{\text{novel}} \leftarrow \mathcal{P}_{\text{novel}} \cup \{T\}.
    17. end if
    18. end if
    19. end for
    20. return Pnovel\mathcal{P}_{\text{novel}}
  4. Knowl 4 — Correlation Between Multi-View Semantic Feature Uncertainty and Ground-Truth Error

    empirical result

    The multi-view generalized variance ui=det⁡(Σi)∈Ru_i = \det(\boldsymbol{\Sigma}_i) \in \mathbb{R}, computed over projected 2D OpenSeg feature vectors at point pi\mathbf{p}_i, serves as a direct indicator of ground-truth semantic feature error ϵi\epsilon_i.

    The per-point error ϵi\epsilon_i is defined as the Euclidean distance between the mean projected feature vector μi∈RD\boldsymbol{\mu}_i \in \mathbb{R}^D and the ground-truth text feature vector oigt=CLIPtext(classi)∈RD\mathbf{o}_i^{\text{gt}} = \text{CLIP}_{\text{text}}(\text{class}_i) \in \mathbb{R}^D:

    ϵi=∥oigt−μi∥2\epsilon_i = \|\mathbf{o}_i^{\text{gt}} - \boldsymbol{\mu}_i\|_2

    Across all Replica scenes, the Pearson correlation coefficient between the generalized variance uiu_i and the error ϵi\epsilon_i yields an average correlation of rˉ=0.653\bar{r} = 0.653 (reaching r=0.69r = 0.69 in scene room0). This demonstrates that cross-view feature disagreement strongly indicates semantic prediction error and validates using generalized variance as an acquisition function for novel view rendering.

  5. Knowl 5 — Zero-Shot 3D Semantic Point Cloud Segmentation Inference

    model/method

    Zero-shot 3D point cloud semantic segmentation for arbitrary candidate text categories {c1,c2,…,cK}\{c_1, c_2, \dots, c_K\} is performed by projecting rendered 2D open-set feature maps onto 3D points and computing cosine similarities against pre-computed CLIP text encodings.

    1. Text Encoding: For each class name ckc_k, a normalized text embedding ek∈RD\mathbf{e}_k \in \mathbb{R}^D is extracted using the pre-trained CLIP text encoder (ViT-L/14@336):

    ek=CLIPtext(ck)∥CLIPtext(ck)∥2\mathbf{e}_k = \frac{\text{CLIP}_{\text{text}}(c_k)}{\|\text{CLIP}_{\text{text}}(c_k)\|_2}

    1. Feature Projection: Posed 2D open-set feature maps O^j\hat{O}_j are rendered from the trained NeRF and projected onto the 3D point cloud PP. For points pi∈P\mathbf{p}_i \in P visible across multiple viewpoints, the projected feature vectors are averaged to yield point feature oi∈RD\mathbf{o}_i \in \mathbb{R}^D.

    2. Classification: Each 3D point pi\mathbf{p}_i is assigned the semantic label k^(pi)\hat{k}(\mathbf{p}_i) maximizing cosine similarity:

    k^(pi)=arg⁡max⁡k∈{1,…,K}(oi∥oi∥2⋅ek)\hat{k}(\mathbf{p}_i) = \arg\max_{k \in \{1, \dots, K\}} \left( \frac{\mathbf{o}_i}{\|\mathbf{o}_i\|_2} \cdot \mathbf{e}_k \right)

  6. Knowl 6 — Open-Vocabulary 3D Semantic Segmentation Benchmark on Replica

    data/table

    Zero-shot open-vocabulary 3D semantic segmentation performance evaluated on the 8 scenes of the Replica dataset (office0-4, room0-2) across 51 semantic categories split into Head (17 most frequent), Common (17 intermediate), and Tail (17 rarest) subsets. Models are evaluated at 640×360640 \times 360 image resolution using OpenSeg 2D image features and CLIP text encoder ViT-L/14@336 averaged over three runs.

    All Head Common Tail
    Method mIoU mAcc mIoU mAcc mIoU mAcc mIoU mAcc
    LERF 10.5 25.8 19.2 28.1 10.1 31.2 2.3 17.3
    OpenScene (Distilled) 14.8 23.0 30.2 41.1 12.8 21.3 1.4 6.7
    OpenScene (Ensemble) 15.9 24.6 31.7 44.8 14.5 22.6 1.5 6.3
    OpenNeRF (Ours) 20.4 31.7 35.4 46.2 20.1 31.3 5.8 17.6

    OpenNeRF achieves a +4.5+4.5 mIoU and +7.1+7.1 mAcc overall improvement over the strongest baseline (OpenScene Ensemble) without requiring any 3D pre-training on Matterport. On tail classes, OpenNeRF achieves 5.85.8 mIoU, nearly quadruple the 1.51.5 mIoU of OpenScene Ensemble, enabling the identification of fine-grained small objects (such as wall-plugs, clocks, tissue-paper, tablets, and cloth) that are missed by explicit mesh-based convolutional baselines.

  7. Knowl 7 — Ablation of Feature Rendering, Depth Supervision, and Active Views

    data/table

    Ablation study assessing the individual contributions of 3D feature representation extraction, depth loss supervision, and active novel view rendering on Replica 3D semantic segmentation.

    Configuration mIoU mAcc
    LERF 10.4 25.5
    OpenScene 15.1 24.6
    (1) Direct 3D Feature Sampling 16.5 29.1
    (2) Render Project 18.5 29.8
    (3) + Depth supervision 19.4 30.8
    (4) + Rendered active novel views 20.4 31.7
    • Render & Project vs. Direct Sampling: Rendering 2D feature maps and back-projecting them to 3D point clouds achieves 18.518.5 mIoU compared to 16.516.5 mIoU from directly evaluating the MLP at discrete 3D sample coordinates (+2.0+2.0 mIoU gain), benefiting from continuous ray integration.
    • Depth Supervision: Supervising density with sensor depth via smooth L1L_1 loss improves mIoU from 18.518.5 to 19.419.4 by sharpening 3D volumetric geometry.
    • Active Novel Views: Uncertainty-guided view generation adds +1.0+1.0 mIoU (20.420.4 mIoU). In contrast, synthesizing views along the original camera trajectory yields no performance gain (19.419.4 mIoU), while generating novel views from randomly placed cameras drops performance drastically to 15.415.4 mIoU due to viewpoints colliding with scene geometry or observing uninformative scenes.

Coverage note — No substantial contributed material was omitted; qualitative visualizations on iPhone captures and open-set phrase query visualizations were omitted as standalone knowls because their quantitative performance and methodology are fully covered by the benchmark and ablation knowls.

References

  1. 1.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al. Do as I Can, Not as I Say: Grounding Language in Robotic Affordances. arXiv preprint arXiv:2204.01691, 2022.
  2. 2.Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2D-3D-Semantic Data for Indoor Scene Understanding. International Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  3. 3.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  4. 4.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. neurips, 2021.
  5. 5.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  6. 6.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2017a.
  7. 7.Angela Dai, Matthias Nießner, Michael Zollöfer, Shahram Izadi, and Christian Theobalt. BundleFusion: Real-time Globally Consistent 3D Reconstruction using On-the-fly Surface Re-integration. ACM Transactions on Graphics 2017 (TOG), 2017b.
  8. 8.Alexandros Delitzas, Ayca Takmaz, Federico Tombari, Robert Sumner, Marc Pollefeys, and Francis Engelmann. Scenefun3d: Fine-grained functionality and affordance understanding in 3d scenes. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  9. 9.E Knuth Donald et al. The Art of Computer Programming. Sorting and searching, 1999.
  10. 10.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scaling Open-Vocabulary Image Segmentation with Image-Level Labels. In European Conference on Computer Vision (ECCV), 2022.
  11. 11.Huy Ha and Shuran Song. Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models. In Conference on Robot Learning (CoRL), 2022.
  12. 12.Lei Han, Tian Zheng, Lan Xu, and Lu Fang. OccuSeg: Occupancy-Aware 3D Instance Segmentation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  13. 13.Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Runze Zhang, Jiayu Dong, Xin Wang, Guangyuan Sun, Hongbo Fu, and Chiew-Lan Tai. VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation. In International Conference on Computer Vision (ICCV), 2021.
  14. 14.Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu, Mohd Omama, Tao Chen, Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil Keetha, Ayush Tewari, et al. ConceptFusion: Open-Set Multimodal 3D Mapping. Robotics: Science and Systems (RSS), 2023.
  15. 15.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up Visual and Vision-Language Representation Learning with Noisy Text Supervision. In International Conference on Machine Learning (ICML), 2021.
  16. 16.Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. LERF: Language Embedded Radiance Fields. In International Conference on Computer Vision (ICCV), 2023.
  17. 17.Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann. Decomposing NeRF for Editing via Feature Field Distillation. International Conference on Neural Information Processing Systems (NeurIPS), 2022.
  18. 18.Lars Kreuzberg, Idil Esen Zulfikar, Sabarinath Mahadevan, Francis Engelmann, and Bastian Leibe. 4D-STOP: Panoptic Segmentation of 4D Lidar using Spatio-Temporal Object Proposal Generation and Aggregation. In European Conference on Computer Vision (ECCV) Workshops, 2022.
  19. 19.Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J Guibas, Andrea Tagliasacchi, Frank Dellaert, and Thomas Funkhouser. Panoptic neural fields: A semantic object-aware neural scene representation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  20. 20.Marc Levoy. Efficient ray tracing of volume data. ACM Transactions on Graphics (TOG), 1990.
  21. 21.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and René Ranftl. Language-Driven Semantic Segmentation. International Conference on Learning Representations (ICLR), 2022a.
  22. 22.Jinke Li, Xiao He, Yang Wen, Yuan Gao, Xiaoqiang Cheng, and Dan Zhang. Panoptic-phnet: Towards real-time and high-precision lidar panoptic segmentation via clustering pseudo heatmap. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022b.
  23. 23.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP. International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  24. 24.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision (ECCV), 2020.
  25. 25.Michael Oechsle, Songyou Peng, and Andreas Geiger. UNISURF: unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In International Conference on Computer Vision (ICCV), 2021.
  26. 26.Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. OpenScene: 3D Scene Understanding with Open Vocabularies. International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  27. 27.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. International Conference on Neural Information Processing Systems (NeurIPS), 2017.
  28. 28.Santhosh K Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. arXiv preprint arXiv:2109.08238, 2021.
  29. 29.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. DenseClip: Language-Guided Dense Prediction with Context-Aware Prompting. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  30. 30.David Rozenberszki, Or Litany, and Angela Dai. Language-Grounded Indoor 3D Semantic Segmentation in the Wild. In European Conference on Computer Vision (ECCV), 2022.
  31. 31.Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. Mask3D: Mask Transformer for 3D Instance Segmentation. International Conference on Robotics and Automation (ICRA), 2023.
  32. 32.Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulò, Norman Müller, Matthias Nießner, Angela Dai, and Peter Kontschieder. Panoptic Lifting for 3D Scene Understanding with Neural Fields. International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  33. 33.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. The Replica Dataset: A Digital Replica of Indoor Spaces. arXiv preprint arXiv:1906.05797, 2019.
  34. 34.Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. OpenMask3D: Open-Vocabulary 3D Instance Segmentation. In International Conference on Neural Information Processing Systems (NeurIPS), 2023a.
  35. 35.Ayça Takmaz, Jonas Schult, Irem Kaftan, Mertcan Akçay, Bastian Leibe, Robert Sumner, Francis Engelmann, and Siyu Tang. 3D Segmentation of Humans in Point Clouds with Synthetic Data. In International Conference on Computer Vision (ICCV), 2023b.
  36. 36.Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In International Conference on Neural Information Processing Systems (NeurIPS), 2020.
  37. 37.Vadim Tschernezki, Iro Laina, Diane Larlus, and Andrea Vedaldi. Neural Feature Fusion Fields: 3D distillation of self-supervised 2D image representations. In International Conference on 3D Vision (3DV), 2022.
  38. 38.Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  39. 39.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In International Conference on Neural Information Processing Systems (NeurIPS), 2021.
  40. 40.Samuel S Wilks. Certain Generalizations in the Analysis of Variance. Biometrika, 1932.
  41. 41.Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. In International Conference on Neural Information Processing Systems (NeurIPS), 2021.
  42. 42.Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler, and Andreas Geiger. Monosdf: Exploring monocular geometric cues for neural implicit surface reconstruction. In International Conference on Neural Information Processing Systems (NeurIPS), 2022.
  43. 43.Yuanwen Yue, Theodora Kontogianni, Konrad Schindler, and Francis Engelmann. Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  44. 44.Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and Andrew J Davison. In-place Scene Labeling and Understanding with Implicit Scene Representation. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  45. 45.Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R Oswald, and Marc Pollefeys. Nice-SLAM: Neural Implicit Scalable Encoding for SLAM. In International Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  46. 46.René Zurbrügg, Yifan Liu, Francis Engelmann, Suryansh Kumar, Marco Hutter, Vaishakh Patil, and Fisher Yu. ICGNet: A Unified Approach for Instance-Centric Grasping. In International Conference on Robotics and Automation (ICRA), 2024.

Citation

MLA
Engelmann, F., et al. “OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views”. ICLR 2024, 2024, http://arxiv.org/abs/2404.03650v1.
APA
Engelmann, F., Manhardt, F., Niemeyer, M., Tateno, K., Pollefeys, M., & Tombari, F. (2024). OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views. ICLR 2024. http://arxiv.org/abs/2404.03650v1
Chicago
Engelmann, F., F. Manhardt, M. Niemeyer, K. Tateno, M. Pollefeys, and F. Tombari. 2024. “OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views”. ICLR 2024. http://arxiv.org/abs/2404.03650v1.
Harvard
Engelmann, F. et al. (2024) “OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views”, ICLR 2024 [Preprint]. Available at: http://arxiv.org/abs/2404.03650v1.
Vancouver
1. Engelmann F, Manhardt F, Niemeyer M, Tateno K, Pollefeys M, Tombari F (2024) OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views. ICLR 2024

BibTeX

@article{engelmann2024opennerf,
  title = {OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views},
  author = {Engelmann, Francis and Manhardt, Fabian and Niemeyer, Michael and Tateno, Keisuke and Pollefeys, Marc and Tombari, Federico},
  year = {2024},
  journal = {ICLR 2024},
  url = {http://arxiv.org/abs/2404.03650v1},
  eprint = {2404.03650}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission