Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
Phuc D. A. NguyenTuan Duc NgoEvangelos KalogerakisChuang GanAnh Tuan TranCuong PhamKhoi Nguyen
Introduces Open3DIS, an open-vocabulary 3D instance segmentation framework that pairs class-agnostic 3D network proposals with aggregated multi-view 2D masks and pointwise CLIP feature extraction to accurately segment small, rare, and geometrically ambiguous objects.
Autonomous robotics, augmented reality, and virtual reality systems require the ability to perceive, locate, and segment physical objects in 3D spaces using natural language. Conventional 3D scene perception models rely on closed-set training, which severely limits their recognition to a small, predefined list of categories. While open-vocabulary methods have emerged to identify novel objects, current techniques struggle significantly when detecting small, rare, or geometrically ambiguous items in complex 3D environments.
The article demonstrates a new framework, named Open3DIS, designed to solve open-vocabulary 3D instance segmentation by generating precise 3D object masks for arbitrary object classes. It evaluates how effectively 2D image-based foundation models can be combined with 3D spatial representations to identify both standard and previously unseen objects from natural language descriptions.
The approach combines two complementary proposal mechanisms applied to multi-view color-depth video sequences and reconstructed 3D point clouds. First, a dedicated module aggregates 2D instance segmentations across camera frames and maps them onto cohesive 3D point clusters, known as superpoints, using hierarchical agglomerative clustering. Second, these proposals are merged with initial shape candidates produced by a 3D instance segmentation network. Finally, a pointwise feature extraction module pools multi-view vision-language features across the point cloud, weighting points by how frequently they appear across top camera views, to match 3D masks directly against open-ended text queries. The framework was evaluated across standard benchmarks including ScanNet200, Replica, and S3DIS.
The experimental findings show substantial improvements in 3D scene understanding. On the challenging 200-class ScanNet200 benchmark, Open3DIS achieved an average precision of 23.7, outperforming prior state-of-the-art open-vocabulary methods by approximately 1.5 times. It demonstrated particular strength on rare, tail-end object categories, where it scored 21.8 average precision—surpassing even fully supervised baselines on these uncommon classes. On the S3DIS dataset, the system achieved a novel-class detection precision of 26.3 to 29.0, more than doubling the scores of existing alternatives. Ablation tests confirmed that hierarchical agglomerative clustering and multi-view pointwise feature extraction were crucial to maximizing accuracy.
These results demonstrate that marrying the broad semantic recognition of 2D foundation models with 3D geometric clustering solves critical blind spots for small and rare objects without requiring costly, fully supervised 3D annotations. For technical leaders and product teams developing intelligent agents or spatial computing applications, this approach offers a path toward highly capable zero-shot 3D interaction. The findings indicate that systems can reliably understand diverse natural language instructions—such as locating items by their brand or intended function—while matching the performance of specialized supervised models.
Organizations developing spatial perception systems should consider adopting hybrid 2D-guided 3D proposal pipelines to improve recognition accuracy for tail categories. The main current limitation noted in the article is that the 2D-guided module and the 3D segmentation network operate independently before their candidate masks are merged. Future work should focus on end-to-end integration where 2D and 3D modules interactively reinforce each other, as well as optimizing computational efficiency across high-density video sequences.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). Learn how superpoints and query decoders are used to aggregate point clouds into 3D instance masks, providing the structural foundation that Open3DIS builds upon for proposal generation.
- Paper: SoftGroup for 3D Instance Segmentation on Point Clouds, Thang Vu et al. (2022). Understand the soft grouping and proposal refinement mechanisms in point cloud instance segmentation that motivate Open3DIS's hybrid 2D-guided and 3D network candidate proposals.
- Paper: CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data, Yihan Zeng et al. (2023). Explore how 2D vision-language representations are directly aligned with 3D point cloud clusters, establishing foundational multimodal cross-attention techniques essential for open-vocabulary 3D scene perception.
- Paper: ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding, Le Xue et al. (2023). Examine the principles of aligning 3D point representations with pre-trained 2D vision-language embeddings like CLIP to enable zero-shot 3D object identification.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). Review how 2D instance segmentations across multiple viewpoints can be lifted into consistent 3D volumetric representations without manual 3D annotations.
- Paper: CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection, Yang Cao et al. (2023). Discover how 2D vision-language priors and 3D geometric clustering are unified to detect unseen, novel 3D object categories in an open-vocabulary setting.
- Paper: Large-Scale Point Cloud Semantic Segmentation with Superpoint Graphs, Loic Landrieu et al. (2017). Familiarize yourself with the original superpoint graph formulation that partitions unstructured point clouds into geometrically homogeneous primitives for hierarchical 3D segmentation.
- Paper: GARField: Group Anything with Radiance Fields, Chung Min Kim et al. (2024). See how multi-view 2D segmentation masks are lifted into continuous neural fields by conditioning affinity on physical 3D scale for hierarchical 3D decomposition.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). Extend open-set 3D understanding from explicit point clouds to neural radiance fields by distilling 2D pixel-aligned features and rendering active novel views.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). Apply open-vocabulary 3D scene representations to downstream embodied robotic reasoning tasks such as spatial placement, compatibility, and relative configuration.
