Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

Phuc D. A. NguyenTuan Duc NgoEvangelos KalogerakisChuang GanAnh Tuan TranCuong PhamKhoi Nguyen

article2024CVPR124 citations

Introduces Open3DIS, an open-vocabulary 3D instance segmentation framework that pairs class-agnostic 3D network proposals with aggregated multi-view 2D masks and pointwise CLIP feature extraction to accurately segment small, rare, and geometrically ambiguous objects.

Listen

Autonomous robotics, augmented reality, and virtual reality systems require the ability to perceive, locate, and segment physical objects in 3D spaces using natural language. Conventional 3D scene perception models rely on closed-set training, which severely limits their recognition to a small, predefined list of categories. While open-vocabulary methods have emerged to identify novel objects, current techniques struggle significantly when detecting small, rare, or geometrically ambiguous items in complex 3D environments.

The article demonstrates a new framework, named Open3DIS, designed to solve open-vocabulary 3D instance segmentation by generating precise 3D object masks for arbitrary object classes. It evaluates how effectively 2D image-based foundation models can be combined with 3D spatial representations to identify both standard and previously unseen objects from natural language descriptions.

The approach combines two complementary proposal mechanisms applied to multi-view color-depth video sequences and reconstructed 3D point clouds. First, a dedicated module aggregates 2D instance segmentations across camera frames and maps them onto cohesive 3D point clusters, known as superpoints, using hierarchical agglomerative clustering. Second, these proposals are merged with initial shape candidates produced by a 3D instance segmentation network. Finally, a pointwise feature extraction module pools multi-view vision-language features across the point cloud, weighting points by how frequently they appear across top camera views, to match 3D masks directly against open-ended text queries. The framework was evaluated across standard benchmarks including ScanNet200, Replica, and S3DIS.

The experimental findings show substantial improvements in 3D scene understanding. On the challenging 200-class ScanNet200 benchmark, Open3DIS achieved an average precision of 23.7, outperforming prior state-of-the-art open-vocabulary methods by approximately 1.5 times. It demonstrated particular strength on rare, tail-end object categories, where it scored 21.8 average precision—surpassing even fully supervised baselines on these uncommon classes. On the S3DIS dataset, the system achieved a novel-class detection precision of 26.3 to 29.0, more than doubling the scores of existing alternatives. Ablation tests confirmed that hierarchical agglomerative clustering and multi-view pointwise feature extraction were crucial to maximizing accuracy.

These results demonstrate that marrying the broad semantic recognition of 2D foundation models with 3D geometric clustering solves critical blind spots for small and rare objects without requiring costly, fully supervised 3D annotations. For technical leaders and product teams developing intelligent agents or spatial computing applications, this approach offers a path toward highly capable zero-shot 3D interaction. The findings indicate that systems can reliably understand diverse natural language instructions—such as locating items by their brand or intended function—while matching the performance of specialized supervised models.

Organizations developing spatial perception systems should consider adopting hybrid 2D-guided 3D proposal pipelines to improve recognition accuracy for tail categories. The main current limitation noted in the article is that the 2D-guided module and the 3D segmentation network operate independently before their candidate masks are merged. Future work should focus on end-to-end integration where 2D and 3D modules interactively reinforce each other, as well as optimizing computational efficiency across high-density video sequences.

arXiv: 2312.10671
Cover for Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance

Abstract

We introduce Open3DIS, a novel solution designed to tackle the problem of Open-Vocabulary Instance Segmentation within 3D scenes. Objects within 3D environments exhibit diverse shapes, scales, and colors, making precise instance-level identification a challenging task. Recent advancements in Open-Vocabulary scene understanding have made significant strides in this area by employing class-agnostic 3D instance proposal networks for object localization and learning queryable features for each 3D mask. While these methods produce high-quality instance proposals, they struggle with identifying small-scale and geometrically ambiguous objects. The key idea of our method is a new module that aggregates 2D instance masks across frames and maps them to geometrically coherent point cloud regions as high-quality object proposals addressing the above limitations. These are then combined with 3D class-agnostic instance proposals to include a wide range of objects in the real world. To validate our approach, we conducted experiments on three prominent datasets, including ScanNet200, S3DIS, and Replica, demonstrating significant performance gains in segmenting objects with diverse categories over the state-of-the-art approaches.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. 2D-Guided-3D Instance Proposal Module
  • 3.2. 3D Instance Segmentation Network
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Comparison to prior work
  • 4.3. Ablation study
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Open3DIS Framework for Open-Vocabulary 3D Instance Segmentation

    model/method

    Open3DIS is a method for open-vocabulary 3D point cloud instance segmentation (OV-3DIS). Given an input 3D point cloud P={pn}n=1NP = \{p_n\}_{n=1}^N (with pn∈R6p_n \in \mathbb{R}^6 encoding 3D coordinates and RGB colors) and an RGB-D video sequence V={(It,Dt,Πt)}t=1TV = \{(I_t, D_t, \Pi_t)\}_{t=1}^T (where ItI_t is an RGB image, DtD_t is a depth map, and Πt\Pi_t is the known projection camera matrix), the goal is to predict a set of binary instance masks representing object instances of arbitrary classes described by open-vocabulary text queries.

    The framework consists of four primary components:

    1. 2D-Guided 3D Instance Proposal Module: Generates K1K_1 3D binary instance masks M1∈{0,1}K1×NM_1 \in \{0, 1\}^{K_1 \times N} by identifying geometrically coherent 3D superpoint clusters corresponding to 2D instance segmentations across frames and merging them hierarchically over multi-view images.
    2. Class-Agnostic 3D Instance Segmentation Network: Employs a pre-trained 3D instance segmenter (such as ISBNet or Mask3D) with its semantic classification heads stripped away, extracting K2K_2 geometric 3D object proposals M2∈{0,1}K2×NM_2 \in \{0, 1\}^{K_2 \times N} directly from point cloud geometry and pointwise feature maps.
    3. Proposal Combination: Merges proposal sets M1M_1 and M2M_2 into an aggregate set of KK candidate masks M∈{0,1}K×NM \in \{0, 1\}^{K \times N}, followed by 3D Non-Maximum Suppression (NMS) using an intersection-over-union (IoU) threshold τdup=0.5\tau_{\text{dup}} = 0.5 to eliminate near-duplicate proposals.
    4. Pointwise 3D-Aware Feature Extraction: Computes a dense CLIP feature representation across the 3D points by projecting 3D proposals onto their top-λ\lambda visible 2D views, accumulating multi-scale 2D CLIP features weighted by point visibility, and assigning text-query alignment scores via average point-feature cosine similarity.
  2. Knowl 2 — Per-Frame Superpoint Merging for 2D-Guided 3D Proposals

    algorithm

    To build initial 3D object proposals from single-view 2D predictions, 3D points are first grouped via an unsupervised graph-based oversegmentation into UU geometrically homogeneous regions termed superpoints, represented as binary masks {qu}u=1U∈{0,1}U×N\{q_u\}_{u=1}^U \in \{0, 1\}^{U \times N}. A 2D open-vocabulary segmenter (such as Grounded-SAM) extracts 2D masks {mi2D}\{m^{2D}_i\} for each RGB-D frame. For every 2D mask, superpoints are merged according to spatial 2D-3D projection overlap and 3D deep feature similarity.

    Input: 3D point cloud P∈RN×6P \in \mathbb{R}^{N \times 6}, superpoint masks {qu}u=1U\{q_u\}_{u=1}^U, per-frame 2D mask mi2Dm_i^{2D}, camera matrix Π\Pi, 3D backbone pointwise features F3D∈RN×D3DF^{3D} \in \mathbb{R}^{N \times D_{3D}}, thresholds τiou\tau_{iou} and τsim\tau_{sim}
    Output: Single-frame 3D proposal region ri⊆{1,…,U}r_i \subseteq \{1, \dots, U\}
    Compute average 3D feature for each superpoint: fu3D=1∣qu∣∑n∈quFn3Df_u^{3D} = \frac{1}{|q_u|} \sum_{n \in q_u} F_n^{3D}
    For each superpoint quq_u, project its points onto the image plane via Π\Pi and compute IoU ou,io_{u, i} with mi2Dm_i^{2D}
    Identify candidate superpoints Qcand={u∣ou,i>τiou}Q_{\text{cand}} = \{u \mid o_{u, i} > \tau_{iou}\}
    Initialize region ri←{arg⁡max⁡uou,i}r_i \leftarrow \{\arg\max_{u} o_{u, i}\}
    repeat
        found_addition ←\leftarrow False
        For each neighboring superpoint qu∈Qcand∖riq_u \in Q_{\text{cand}} \setminus r_i adjacent to the current region rir_i:
            Compute maximum feature cosine similarity: simax⁡=max⁡u′∈ricos⁡(fu′3D,fu3D)s_i^{\max} = \max_{u' \in r_i} \cos(f_{u'}^{3D}, f_u^{3D})
            if simax⁡>τsims_i^{\max} > \tau_{sim} and simax⁡s_i^{\max} is maximal among neighbors then
                ri←ri∪{u}r_i \leftarrow r_i \cup \{u\}
                found_addition ←\leftarrow True
    until not found_addition
    return rir_i
  3. Knowl 3 — Hierarchical Multi-Frame Agglomerative Region Merging

    model/method

    Because single-frame back-projected 3D regions capture only partial surfaces and suffer from occlusion, Open3DIS merges region sets across video frames in a bottom-up hierarchical binary tree structure using Agglomerative Clustering.

    Given two point cloud region sets {ri}i=1I\{r_i\}_{i=1}^I and {rj}j=1J\{r_j\}_{j=1}^J from different frames or merged hierarchical levels, an active set of size I+JI + J is constructed. The pairwise merging compatibility is evaluated using both 3D bounding geometry overlap and deep feature consistency. Specifically, for region pair (ri,rj)(r_i, r_j), the 3D intersection-over-union is oi,j′=IoU(ri,rj)o'_{i,j} = \text{IoU}(r_i, r_j) and the feature similarity is si,j′=cos⁡(fi3D,fj3D)s'_{i,j} = \cos(f_i^{3D}, f_j^{3D}), where fi3Df_i^{3D} is the mean 3D backbone feature across all points in rir_i.

    The binary cost matrix C∈{0,1}(I+J)×(I+J)C \in \{0, 1\}^{(I+J) \times (I+J)} is computed as: ci,j=I(oi,j′>τiou)⊙I(si,j′>τsim)c_{i,j} = \mathbb{I}(o'_{i,j} > \tau_{\text{iou}}) \odot \mathbb{I}(s'_{i,j} > \tau_{\text{sim}}) where I(⋅)\mathbb{I}(\cdot) is the indicator function and ⊙\odot is the logical AND operator. Regions with ci,j=1c_{i,j} = 1 are iteratively merged and the cost matrix is dynamically updated until no positive entries remain (C=0C = 0).

    Merging is scheduled hierarchically rather than sequentially: pairs of neighboring video frames are merged at level 0, their resulting region sets form pairs at level 1, and the process proceeds up the binary tree until all frames in the RGB-D sequence are aggregated into complete 3D object proposals.

  4. Knowl 4 — Pointwise 3D-Aware Visual Feature Extraction and Text Query Alignment

    equation

    For each candidate 3D proposal binary mask mk3D∈{0,1}Nm_k^{3D} \in \{0, 1\}^N, the 3D points are projected into all sequence viewpoints to select the top-λ\lambda views (where λ=5\lambda = 5) with the highest counts of visible projected points. For each view λ\lambda, a 2D crop around the projected mask is encoded into a normalized CLIP image embedding fλ,kCLIP∈RDCLIPf_{\lambda,k}^{\text{CLIP}} \in \mathbb{R}^{D_{\text{CLIP}}}, and νλ∈{0,1}N\nu_\lambda \in \{0, 1\}^N denotes the binary per-point visibility mask for view λ\lambda.

    The dense, instance-aware pointwise CLIP feature map FCLIP∈RN×DCLIPF^{\text{CLIP}} \in \mathbb{R}^{N \times D_{\text{CLIP}}} across all points in the point cloud is accumulated as: FCLIP=NV(∑k(∑λ(νλ∗fλ,kCLIP)∗mk3D))F^{\text{CLIP}} = \text{NV}\left( \sum_k \left( \sum_\lambda (\nu_\lambda * f_{\lambda,k}^{\text{CLIP}}) * m_k^{3D} \right) \right) where ∗* denotes element-wise multiplication with broadcasting, and NV(x)=x∥x∥2\text{NV}(x) = \frac{x}{\|x\|_2} is the L2L_2 normalization operator.

    Given a text prompt query ρ\rho with normalized CLIP text embedding eρ∈RDCLIPe_\rho \in \mathbb{R}^{D_{\text{CLIP}}}, the matching score sk,ρCLIPs_{k, \rho}^{\text{CLIP}} for the kk-th 3D proposal mask mk3Dm_k^{3D} is calculated as the average cosine similarity over all points belonging to the mask: sk,ρCLIP=1∣mk3D∣∑n=1Ncos⁡(FnCLIP∗mk,n3D,eρ)s_{k, \rho}^{\text{CLIP}} = \frac{1}{|m_k^{3D}|} \sum_{n=1}^N \cos(F^{\text{CLIP}}_n * m_{k,n}^{3D}, e_\rho) where ∣mk3D∣|m_k^{3D}| is the number of points in the kk-th mask.

  5. Knowl 5 — ScanNet200 Open-Vocabulary 3D Instance Segmentation Benchmark

    data/table

    Quantitative evaluation of Open3DIS against prior open-vocabulary and fully-supervised 3D instance segmentation methods on the ScanNet200 validation set (198 classes categorized into head, common, and tail splits). Evaluated metrics include average precision (AP) across IoU thresholds from 50% to 95% in 5% increments, AP50\text{AP}_{50}, AP25\text{AP}_{25}, and split-specific APs (APhead\text{AP}_{\text{head}}, APcom\text{AP}_{\text{com}}, APtail\text{AP}_{\text{tail}}).

    Method Setting 3D Proposal AP AP50\text{AP}_{50} AP25\text{AP}_{25} APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}}
    ISBNet Fully-sup – 24.5 32.7 37.6 38.6 20.5 12.5
    Mask3D Fully-sup – 26.9 36.2 41.4 39.8 21.7 17.9
    OpenScene + DBScan Open-vocab None 2.8 7.8 18.6 2.7 3.1 2.6
    OpenScene + Mask3D Open-vocab Mask3D 11.7 15.2 17.8 13.4 11.6 9.9
    SAM3D Open-vocab None 6.1 14.2 21.3 7.0 6.2 4.6
    OVIR-3D Open-vocab None 13.0 24.9 32.3 14.4 12.7 11.7
    OpenIns3D Open-vocab Mask3D 8.8 10.3 14.4 16.0 6.5 4.2
    OpenMask3D Open-vocab Mask3D 15.4 19.9 23.1 17.1 14.1 14.9
    Open3DIS (only 2D) Open-vocab None 18.2 26.1 31.4 18.9 16.5 19.2
    Open3DIS (only 3D) Open-vocab ISBNet 18.6 23.1 27.3 24.7 16.9 13.3
    Open3DIS (2D and 3D) Open-vocab ISBNet 23.7 29.4 32.8 27.8 21.2 21.8

    Open3DIS using both 2D and 3D proposals outperforms prior open-vocabulary methods by approximately 1.5×1.5\times in overall AP (23.7 vs 15.4 for OpenMask3D). Furthermore, in rare tail categories (APtail\text{AP}_{\text{tail}}), Open3DIS achieves 21.8 AP, outperforming both open-vocabulary baselines and fully-supervised models (ISBNet at 12.5 and Mask3D at 17.9).

  6. Knowl 6 — Cross-Dataset and Novel-Class Generalization on S3DIS, Replica, and ScanNet20-Pretrained ScanNet200

    data/table

    Generalization performance of Open3DIS evaluated across three settings: (1) ScanNet200 evaluated with 3D proposals pre-trained only on ScanNet20 (51 base classes, remaining novel classes); (2) Zero-shot evaluation on Replica (48 classes) using only 2D proposals without 3D training; (3) S3DIS evaluated on base and novel category splits (B8/N4 and B6/N6).

    ScanNet20-Pretrained on ScanNet200
    Method Pretrain APnovel\text{AP}_{\text{novel}} APbase\text{AP}_{\text{base}} AP
    PLA (Base 20) ScanNet20 0.3 15.8 4.5
    OpenScene + Mask3D ScanNet20 7.6 11.1 8.5
    OpenMask3D ScanNet20 11.9 14.3 12.6
    Open3DIS (Ours) ScanNet20 16.5 25.8 19.0
    Replica (Zero-Shot)
    Method 3D Proposal AP AP50\text{AP}_{50} AP25\text{AP}_{25}
    OpenMask3D Mask3D 13.1 18.4 24.2
    OVIR-3D None 11.1 20.5 27.5
    Open3DIS (only 2D) None 18.1 26.7 30.5
    Open3DIS (2D and 3D) ISBNet 18.5 24.5 28.2
    S3DIS Base/Novel Splits
    Method AP50B\text{AP}_{50}^B (B8/N4) AP50N\text{AP}_{50}^N (B8/N4) AP50B\text{AP}_{50}^B (B6/N6) AP50N\text{AP}_{50}^N (B6/N6)
    PLA 59.0 8.6 46.9 9.8
    Lowis3D 58.7 13.8 51.8 15.8
    Open3DIS (Ours) 60.8 26.3 50.0 29.0

    When pre-trained solely on ScanNet20, Open3DIS obtains 16.5 APnovel\text{AP}_{\text{novel}} on ScanNet200, outperforming OpenMask3D (11.9) and contrastive 3D methods like PLA (0.3). On Replica, Open3DIS achieves 18.1 AP with 2D-only proposals (outperforming OVIR-3D by +7.0 AP). On S3DIS, Open3DIS achieves 26.3 and 29.0 AP50N\text{AP}_{50}^N on novel categories, nearly doubling prior state-of-the-art results.

  7. Knowl 7 — Ablation on Feature Representation and Pointwise 3D-Aware Pooling

    data/table

    Comparison of different open-vocabulary feature extraction strategies applied to Open3DIS's candidate 3D proposals on the ScanNet200 validation set.

    Setting AP APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}}
    A1: OpenScene (distill) 3.3 5.5 2.4 1.7
    A2: OpenScene (fusion) 17.5 21.5 17.1 13.3
    A3: OpenScene (ensemble) 5.6 6.4 4.8 5.7
    B: Mask-wise Feature Pooling 22.2 25.9 19.3 21.4
    C: Point-wise Feature Accumulation (Open3DIS) 23.7 27.8 21.2 21.8

    Setting B projects 3D proposals onto top-λ\lambda views and pools features directly at the mask level (as in OpenMask3D), achieving 22.2 AP. Setting C (Open3DIS pointwise accumulation weighted by per-point visibility across top views) reaches 23.7 AP (+1.5 AP improvement), confirming that frequency-of-visibility weighting on individual 3D points yields more descriptive multi-view embeddings.

  8. Knowl 8 — Ablation on Superpoint Filtering, Merging Strategy, and Merging Order

    data/table

    Ablation results on ScanNet200 evaluating the internal design choices of the 2D-Guided 3D Instance Proposal Module (2D-G-3DIP) in isolation (without 3D network proposals).

    Superpoint Filtering Condition AP APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}}
    ✓ Deep Feature Cosine Similarity 18.2 18.9 16.5 19.2
    ✓ None 15.9 16.5 14.3 17.0
    ✓ Euclidean Distance 16.0 16.4 14.1 17.6
    None None (raw point grouping) 12.0 12.6 11.2 12.2
    Merging Strategy Merging Order AP APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}}
    Hungarian Sequential 13.2 13.9 11.3 14.7
    Hungarian Hierarchical 16.1 16.1 13.3 19.4
    Agglomerative Sequential 16.9 17.8 16.1 18.0
    Agglomerative Hierarchical 18.2 18.9 16.5 19.2

    Using superpoints combined with deep feature cosine filtering outperforms Euclidean distance filtering (18.2 AP vs 16.0 AP) and raw point grouping (12.0 AP). For multi-frame merging, Agglomerative Clustering paired with hierarchical tree ordering achieves the highest performance (18.2 AP), surpassing Hungarian bipartite matching (16.1 AP) and sequential frame merging (16.9 AP).

  9. Knowl 9 — Sensitivity to Thresholds, View Selection, and Segmenter Backbones

    data/table

    Hyperparameter sensitivity and modular compatibility of Open3DIS on the ScanNet200 validation set across varying IoU thresholds τiou\tau_{\text{iou}}, feature similarity thresholds τsim\tau_{\text{sim}}, view selection counts λ\lambda, and upstream 2D/3D segmenter choices.

    τiou\tau_{\text{iou}} Threshold 0.3 0.5 0.7 0.9 0.95
    AP 17.7 17.8 18.0 18.2 16.9
    AP50\text{AP}_{50} 25.4 25.8 25.9 26.1 24.1
    τsim\tau_{\text{sim}} Threshold 0.5 0.7 0.8 0.9 0.95
    AP 14.2 14.6 17.2 18.2 16.2
    AP50\text{AP}_{50} 21.0 21.8 25.1 26.1 23.8
    Top-λ\lambda View Selection Top 1 Top 5 Top 10 Top 20 All
    AP 21.2 23.7 22.6 22.5 22.5
    AP50\text{AP}_{50} 27.3 29.4 28.7 29.0 29.1
    2D Segmenter AP APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}} 3D Segmenter AP APhead\text{AP}_{\text{head}} APcom\text{AP}_{\text{com}} APtail\text{AP}_{\text{tail}}
    SEEM 21.5 26.5 19.6 18.0 Mask3D 23.7 26.4 22.5 21.9
    ODISE 21.6 26.0 19.5 19.1 ISBNet 23.7 27.8 21.2 21.8
    Detic 22.2 26.8 20.0 19.2 – – – – –
    Grounded-SAM 23.7 27.8 21.2 21.8 – – – – –

    Optimal proposal generation occurs at τiou=0.9\tau_{\text{iou}} = 0.9 and τsim=0.9\tau_{\text{sim}} = 0.9. For viewpoint pooling, λ=5\lambda = 5 achieves peak performance (23.7 AP); increasing beyond 5 introduces noise from occluded views. Both ISBNet and Mask3D serve as effective 3D backbones (23.7 AP each), while Grounded-SAM provides the strongest 2D segmentation guidance.

  10. Knowl 10 — Decoupled 2D-Guided and 3D Proposal Pipelines

    limitation

    In Open3DIS, the class-agnostic 3D proposal network and the 2D-Guided-3D Instance Proposal Module operate independently as separate pipelines. Their resulting candidate 3D masks are combined via simple concatenation followed by Non-Maximum Suppression (NMS), rather than interacting synergistically during the proposal generation process. A tighter, unified integration where 2D and 3D modalities iteratively guide and refine each other remains an open direction for future work.

Coverage note — None was omitted; all contributed algorithms, equations, empirical benchmarks (ScanNet200, S3DIS, Replica), ablation studies, and limitations are fully covered.

References

  1. 1.Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1534–1543, 2016.
  2. 2.Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese. Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105, 2017.
  3. 3.Yang Cao, Yihan Zeng, Hang Xu, and Dan Xu. Coda: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3d object detection. arXiv preprint arXiv:2310.02960, 2023.
  4. 4.Shaoyu Chen, Jiemin Fang, Qian Zhang, Wenyu Liu, and Xinggang Wang. Hierarchical aggregation for 3d instance segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15467–15476, 2021.
  5. 5.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017.
  6. 6.Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. Decoupling zero-shot semantic segmentation. 2022.
  7. 7.Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Lowis3d: Language-driven open-world instance-level 3d scene understanding. arXiv preprint arXiv:2308.00353, 2023.
  8. 8.Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Pla: Language-driven open-vocabulary 3d scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023.
  9. 9.Shichao Dong, Guosheng Lin, and Tzu-Yi Hung. Learning regional purity for instance segmentation on 3d point clouds. In European Conference on Computer Vision, pages 56–72. Springer, 2022.
  10. 10.Martin Ester, Hans-Peter Kriegel, Jörg Sander, Xiaowei Xu, et al. A density-based algorithm for discovering clusters in large spatial databases with noise. In kdd, pages 226–231, 1996.
  11. 11.Pedro F Felzenszwalb and Daniel P Huttenlocher. Efficient graph-based image segmentation. International journal of computer vision, 59:167–181, 2004.
  12. 12.Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, et al. Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning. arXiv preprint arXiv:2309.16650, 2023.
  13. 13.Shuting He, Henghui Ding, and Wei Jiang. Semantic-promoted debiasing and background disambiguation for zero-shot instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19498–19507, 2023.
  14. 14.Tong He, Chunhua Shen, and Anton van den Hengel. Dyco3d: Robust instance segmentation of 3d point clouds through dynamic convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 354–363, 2021.
  15. 15.Tong He, Wei Yin, Chunhua Shen, and Anton van den Hengel. Pointinst3d: Segmenting 3d instances by points. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part III, pages 286–302. Springer, 2022.
  16. 16.Deepti Hegde, Jeya Maria Jose Valanarasu, and Vishal M Patel. Clip goes 3d: Leveraging prompt tuning for language grounded 3d recognition. arXiv preprint arXiv:2303.11313, 2023.
  17. 17.Yining Hong, Chunru Lin, Yilun Du, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. 3d concept learning and reasoning from multi-view images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9202–9212, 2023.
  18. 18.Ji Hou, Angela Dai, and Matthias Nießner. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4421–4430, 2019.
  19. 19.Zhening Huang, Xiaoyang Wu, Xi Chen, Hengshuang Zhao, Lei Zhu, and Joan Lasenby. Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation. arXiv preprint, 2023.
  20. 20.Le Hui, Linghua Tang, Yaqi Shen, Jin Xie, and Jian Yang. Learning superpoint graph cut for 3d instance segmentation. In Advances in Neural Information Processing Systems, 2022.
  21. 21.Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, and Ehsan Elhamifar. Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7020–7031, 2022.
  22. 22.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021.
  23. 23.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4867–4876, 2020.
  24. 24.Prannay Kaul, Weidi Xie, and Andrew Zisserman. Multi-modal classifiers for open-vocabulary object detection. In International Conference on Machine Learning, 2023.
  25. 25.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023.
  26. 26.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  27. 27.Loic Landrieu and Mohamed Boussaha. Point cloud over-segmentation with graph-structured deep metric learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7440–7449, 2019.
  28. 28.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. In International Conference on Learning Representations, 2022.
  29. 29.Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Open-vocabulary object segmentation with diffusion models. 2023.
  30. 30.Zhihao Liang, Zhihao Li, Songcen Xu, Mingkui Tan, and Kui Jia. Instance segmentation in 3d scenes using semantic superpoint tree networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2783–2792, 2021.
  31. 31.Jiaheng Liu, Tong He, Honghui Yang, Rui Su, Jiayi Tian, Junran Wu, Hongcheng Guo, Ke Xu, and Wanli Ouyang. 3d-queryis: A query-based framework for 3d instance segmentation. arXiv preprint arXiv:2211.09375, 2022.
  32. 32.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  33. 33.Shiyang Lu, Haonan Chang, Eric Pu Jing, Abdeslam Boularias, and Kostas Bekris. Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data. In 7th Annual Conference on Robot Learning, 2023.
  34. 34.Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, and Shanghang Zhang. Open-vocabulary point-cloud object detection without 3d annotation. 2023.
  35. 35.Daniel Müllner. Modern hierarchical, agglomerative clustering algorithms. arXiv preprint arXiv:1109.2378, 2011.
  36. 36.Tuan Duc Ngo, Binh-Son Hua, and Khoi Nguyen. Isbnet: a 3d point cloud instance segmentation network with instance-aware sampling and box-aware dynamic convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13550–13559, 2023.
  37. 37.Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, Andrea Tagliasacchi, Marc Pollefeys, and Thomas Funkhouser. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  38. 38.Chau Pham, Truong Vu, and Khoi Nguyen. Lp-ovod: Open-vocabulary object detection by linear probing. arXiv preprint arXiv:2310.17109, 2023.
  39. 39.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  40. 40.Damien Robert, Hugo Raguet, and Loic Landrieu. Efficient 3d semantic segmentation with superpoint transformer. arXiv preprint arXiv:2306.08045, 2023.
  41. 41.David Rozenberszki, Or Litany, and Angela Dai. Language-grounded indoor 3d semantic segmentation in the wild. In Proceedings of the European Conference on Computer Vision (ECCV), 2022.
  42. 42.Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. Mask3d for 3d semantic instance segmentation. In International Conference on Robotics and Automation (ICRA), 2023.
  43. 43.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797, 2019.
  44. 44.Jiahao Sun, Chunmei Qing, Junpeng Tan, and Xiangmin Xu. Superpoint transformer for 3d scene instance segmentation. arXiv preprint arXiv:2211.15766, 2022.
  45. 45.Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. OpenMask3D: Open-Vocabulary 3D Instance Segmentation. In Advances in Neural Information Processing Systems (NeurIPS), 2023.
  46. 46.Vibashan VS, Ning Yu, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M Patel, and Ran Xu. Mask-free ovis: Open-vocabulary instance segmentation without manual mask annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23539–23549, 2023.
  47. 47.Thang Vu, Kookhoi Kim, Tung M. Luu, Xuan Thanh Nguyen, and Chang D. Yoo. Softgroup for 3d instance segmentation on 3d point clouds. In CVPR, 2022.
  48. 48.Luting Wang, Yi Liu, Penghui Du, Zihan Ding, Yue Liao, Qiaosong Qi, Biaolong Chen, and Si Liu. Object-aware distillation pyramid for open-vocabulary object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11186–11196, 2023.
  49. 49.Weiyue Wang, Ronald Yu, Qiangui Huang, and Ulrich Neumann. Sgpn: Similarity group proposal network for 3d point cloud instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2569–2578, 2018.
  50. 50.Jianzong Wu, Xiangtai Li, Henghui Ding, Xia Li, Guangliang Cheng, Yunhai Tong, and Chen Change Loy. Betrayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation. arXiv preprint arXiv:2301.00805, 2023.
  51. 51.Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, and Chunhua Shen. Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. arXiv preprint arXiv:2303.11681, 2023.
  52. 52.Yizheng Wu, Min Shi, Shuaiyuan Du, Hao Lu, Zhiguo Cao, and Weicai Zhong. 3d instances as 1d kernels. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIX, pages 235–252. Springer, 2022.
  53. 53.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18134–18144, 2022.
  54. 54.Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2955–2966, 2023.
  55. 55.Mutian Xu, Xingyilang Yin, Lingteng Qiu, Yang Liu, Xin Tong, and Xiaoguang Han. Sampro3d: Locating sam prompts in 3d for zero-shot scene segmentation. arXiv preprint arXiv:2311.17707, 2023.
  56. 56.Mi Yan, Jiazhao Zhang, Yan Zhu, and He Wang. Maskclustering: View consensus based mask graph clustering for open-vocabulary 3d instance segmentation, 2024.
  57. 57.Bo Yang, Jianan Wang, Ronald Clark, Qingyong Hu, Sen Wang, Andrew Markham, and Niki Trigoni. Learning object bounding boxes for 3d instance segmentation on point clouds. In Advances in Neural Information Processing Systems, pages 6737–6746, 2019.
  58. 58.Jihan Yang, Runyu Ding, Zhe Wang, and Xiaojuan Qi. Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding. arXiv preprint arXiv:2304.00962, 2023.
  59. 59.Yunhan Yang, Xiaoyang Wu, Tong He, Hengshuang Zhao, and Xihui Liu. Sam3d: Segment anything in 3d scenes. arXiv preprint arXiv:2306.03908, 2023.
  60. 60.Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, and Hang Xu. Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23497–23506, 2023.
  61. 61.Li Yi, Wang Zhao, He Wang, Minhyuk Sung, and Leonidas J Guibas. Gspn: Generative shape proposal network for 3d instance segmentation in point cloud. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3947–3956, 2019.
  62. 62.Yingda Yin, Yuzheng Liu, Yang Xiao, Daniel Cohen-Or, Jingwei Huang, and Baoquan Chen. Sai3d: Segment any instance in 3d scenes. arXiv preprint arXiv:2312.11557, 2023.
  63. 63.Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Open-vocabulary detr with conditional matching. 2022.
  64. 64.Cheng Zhang, Haocheng Wan, Shengqiang Liu, Xinyi Shen, and Zizhao Wu. Pvt: Point-voxel transformer for 3d deep learning. arXiv preprint arXiv:2108.06076, 2021.
  65. 65.Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianwei Yang, and Lei Zhang. A simple framework for open-vocabulary segmentation and detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1020–1031, 2023.
  66. 66.Zhuowen Tu Zheng Ding, Jieke Wang. Open-vocabulary universal image segmentation with maskclip. In International Conference on Machine Learning, 2023.
  67. 67.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16793–16803, 2022.
  68. 68.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbóhl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In ECCV, 2022.
  69. 69.Chenming Zhu, Wenwei Zhang, Tai Wang, Xihui Liu, and Kai Chen. Object2scene: Putting objects in context for open-vocabulary 3d detection. arXiv preprint arXiv:2309.09456, 2023.
  70. 70.Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. arXiv preprint arXiv:2304.06718, 2023.

Citation

MLA
Nguyen, P. D. A., et al. “Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance”. arXiv, 2023, http://arxiv.org/abs/2312.10671v3.
APA
Nguyen, P. D. A., Ngo, T. D., Kalogerakis, E., Gan, C., Tran, A., Pham, C., & Nguyen, K. (2023). Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance. arXiv. http://arxiv.org/abs/2312.10671v3
Chicago
Nguyen, P. D. A., T. D. Ngo, E. Kalogerakis, et al. 2023. “Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance”. arXiv. http://arxiv.org/abs/2312.10671v3.
Harvard
Nguyen, P.D.A. et al. (2023) “Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10671v3.
Vancouver
1. Nguyen PDA, Ngo TD, Kalogerakis E, Gan C, Tran A, Pham C, Nguyen K (2023) Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance. arXiv

BibTeX

@article{nguyen2023open3dis,
  title = {Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance},
  author = {Nguyen, Phuc D. A. and Ngo, Tuan Duc and Kalogerakis, Evangelos and Gan, Chuang and Tran, Anh and Pham, Cuong and Nguyen, Khoi},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10671v3},
  eprint = {2312.10671}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE