Superpoint Transformer for 3D Scene Instance Segmentation

Jiahao SunChunmei QingJunpeng TanXiangmin Xu

article2023AAAI206 citations

Proposes an end-to-end 3D instance segmentation framework that uses superpoint cross-attention within a transformer decoder to directly predict instance masks without relying on intermediate grouping steps or complex post-processing.

Listen

Accurate three-dimensional scene understanding is essential for emerging technologies such as autonomous driving, robotics, and augmented reality. A critical component of this understanding is three-dimensional instance segmentation, which involves identifying individual objects within spatial point cloud data and delineating their exact boundaries. Existing approaches typically rely on either generating bounding boxes or performing bottom-up point aggregation based on semantic predictions. However, these methods suffer from performance bottlenecks caused by inaccurate bounding boxes, error propagation from intermediate semantic predictions, and time-consuming aggregation steps.

The article evaluates a unified framework called Superpoint Transformer, or SPFormer, designed to perform end-to-end instance segmentation directly on three-dimensional scenes without relying on object detection bounding boxes or intermediate semantic segmentation tasks.

To overcome computational limits, the approach combines a bottom-up grouping stage with a top-down prediction pipeline. Raw three-dimensional points are first processed through a sparse neural network and pooled into geometric clusters called superpoints, which reduce hundreds of thousands of individual points into several hundred compact representations. A query decoder utilizing attention mechanisms then uses learnable vectors to focus on relevant superpoints and generate object masks, categories, and confidence scores. Optimal one-to-one matching between predictions and actual objects is performed using mask comparisons, eliminating the need for separate point aggregation and post-processing filters.

Evaluations across standard benchmark datasets demonstrate several key findings. First, the proposed framework achieves a mean average precision of 54.9% on the ScanNetv2 hidden test benchmark, exceeding the previous state of the art by 4.3 percentage points. Second, on the ScanNetv2 validation set, it achieves a mean average precision of 56.3%, outperforming the previous best result by 6.9 percentage points. Third, the system demonstrates the fastest processing time among evaluated models at 247 milliseconds per scene, even when accounting for initial superpoint generation on the processor. Finally, additional tests on the S3DIS indoor dataset confirm strong performance across diverse environments, achieving a leading 66.8% precision score at a 50% overlap threshold on Area 5.

These results demonstrate that direct mask matching and superpoint-based attention can deliver superior accuracy while reducing computational latency. For organizations deploying spatial vision systems, eliminating multi-stage aggregation and complex post-processing reduces algorithmic complexity, memory requirements, and processing overhead. This enables faster decision-making cycles in safety-critical applications like automated navigation and robotics.

Engineering and research teams should consider adopting direct superpoint-query architectures when designing real-time spatial segmentation pipelines. Further work should explore optimizing initial superpoint computation on graphics hardware to reduce overall inference latency even further, as well as evaluating model performance on large-scale outdoor datasets and diverse sensor configurations. While confidence in the indoor benchmark results is high, practitioners should validate the method under varying point densities and sensor noise conditions before full deployment.

arXiv: 2211.15766
Cover for Superpoint Transformer for 3D Scene Instance Segmentation

Abstract

Most existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of the overall 3D instance segmentation framework. 2) Existing methods require a time-consuming intermediate step of aggregation. To address these issues, this paper proposes a novel end-to-end 3D instance segmentation method based on Superpoint Transformer, named as SPFormer. It groups potential features from point clouds into superpoints, and directly predicts instances through query vectors without relying on the results of object detection or semantic segmentation. The key step in this framework is a novel query decoder with transformers that can capture the instance information through the superpoint cross-attention mechanism and generate the superpoint masks of the instances. Through bipartite matching based on superpoint masks, SPFormer can implement the network training without the intermediate aggregation step, which accelerates the network. Extensive experiments on ScanNetv2 and S3DIS benchmarks verify that our method is concise yet efficient. Notably, SPFormer exceeds compared state-of-the-art methods by 4.3% on ScanNetv2 hidden test set in terms of mAP and keeps fast inference speed (247ms per frame) simultaneously. Code is available at https://github.com/sunjiahao1999/SPFormer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Backbone and Superpoints
  • 3.2 Query Decoder
  • 3.3 Bipartite Matching and Loss Function
  • 3.4 Inference
  • 4 Experiments
  • 4.1 Benchmark Results
  • 4.2 Ablation Study
  • 4.3 Visualizations
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — SPFormer Architecture and Superpoint Feature Extraction

    model/method

    SPFormer is an end-to-end, two-stage 3D scene instance segmentation framework that eliminates intermediate geometric grouping and bounding-box detection steps.

    1. Bottom-Up Grouping Stage: Given an input point cloud P∈RN×6P \in \mathbb{R}^{N \times 6} with 3D coordinates (x,y,z)(x, y, z) and color channels (r,g,b)(r, g, b), the point cloud is voxelized and processed by a sparse 3D U-Net backbone built with submanifold sparse convolutions (SSC) and sparse convolutions (SC). The backbone extracts point-wise latent features P′∈RN×CP' \in \mathbb{R}^{N \times C}. A superpoint pooling layer computes superpoint features S∈RM×CS \in \mathbb{R}^{M \times C} by applying average pooling over point features within each of MM pre-computed geometric superpoints (M≪NM \ll N). Unlike traditional grouping pipelines, these superpoint features remain potential mid-level embeddings trained directly via instance-level supervision without auxiliary semantic or center-offset regression losses.

    2. Top-Down Proposal Stage: A query decoder processes superpoint features SS along two parallel branches:

    • Mask Branch: An MLP projects superpoint features to mask-aware representations Smask∈RM×DS_{\text{mask}} \in \mathbb{R}^{M \times D}, where DD is the feature dimension.
    • Instance Branch: A stacked transformer decoder refines KK learnable instance query vectors Zℓ∈RK×DZ_\ell \in \mathbb{R}^{K \times D} (where ℓ∈{1,2,… }\ell \in \{1, 2, \dots\} indexes the decoder layer) by cross-attending to projected superpoint features S′∈RM×DS' \in \mathbb{R}^{M \times D} using masked superpoint cross-attention.
  2. Knowl 2 — Superpoint Mask-Constrained Cross-Attention

    equation

    In the SPFormer transformer decoder, query vectors capture instance representations by attending exclusively to relevant foreground superpoints. At decoder layer ℓ\ell, cross-attention is defined as:

    Z^ℓ=softmax(QKTD+Aℓ−1)V\hat{Z}_\ell = \text{softmax}\left(\frac{Q K^T}{\sqrt{D}} + A_{\ell-1}\right) V

    where:

    • Q=ψQ(Zℓ−1)∈RK×DQ = \psi_Q(Z_{\ell-1}) \in \mathbb{R}^{K \times D} is the linear projection of query vectors Zℓ−1Z_{\ell-1} from the preceding layer.
    • K=ψK(S′)∈RM×DK = \psi_K(S') \in \mathbb{R}^{M \times D} and V=ψV(S′)∈RM×DV = \psi_V(S') \in \mathbb{R}^{M \times D} are linear projections of the superpoint features S′∈RM×DS' \in \mathbb{R}^{M \times D}.
    • DD denotes the feature embedding dimension.
    • Aℓ−1∈RK×MA_{\ell-1} \in \mathbb{R}^{K \times M} is a dynamic attention mask derived from the superpoint instance mask predictions Mℓ−1∈[0,1]K×MM_{\ell-1} \in [0, 1]^{K \times M} of the preceding layer:

    Aℓ−1(i,j)={0if Mℓ−1(i,j)≥τ−∞otherwiseA_{\ell-1}(i, j) = \begin{cases} 0 & \text{if } M_{\ell-1}(i, j) \ge \tau \\ -\infty & \text{otherwise} \end{cases}

    where τ\tau is a foreground threshold set empirically to τ=0.5\tau = 0.5. In each transformer decoder layer, the order of attention operations is modified to execute superpoint cross-attention prior to self-attention, and explicit query position encodings are omitted.

  3. Knowl 3 — Superpoint-Based Bipartite Matching Cost

    equation

    To train SPFormer end-to-end without hand-crafted anchors or region proposals, instance predictions are assigned to ground-truth instances using the Hungarian algorithm based on a pairwise matching cost CikC_{ik} between the ii-th predicted proposal and the kk-th ground-truth instance:

    Cik=−λcls⋅pi,ck+λmask⋅CikmaskC_{ik} = -\lambda_{\text{cls}} \cdot p_{i, c_k} + \lambda_{\text{mask}} \cdot C^{\text{mask}}_{ik}

    where pi,ckp_{i, c_k} is the predicted probability of the ground-truth category ckc_k for proposal ii, and coefficients are set to λcls=0.5\lambda_{\text{cls}} = 0.5 and λmask=1.0\lambda_{\text{mask}} = 1.0.

    The superpoint mask matching cost CikmaskC^{\text{mask}}_{ik} combines binary cross-entropy (BCE) and Dice loss with Laplace smoothing:

    Cikmask=BCE(mi,mkgt)+1−2(mi⋅mkgt)+1∣mi∣+∣mkgt∣+1C^{\text{mask}}_{ik} = \text{BCE}(m_i, m^{\text{gt}}_k) + 1 - \frac{2 (m_i \cdot m^{\text{gt}}_k) + 1}{|m_i| + |m^{\text{gt}}_k| + 1}

    where mi∈[0,1]Mm_i \in [0, 1]^M is the predicted superpoint mask probability vector and mkgt∈{0,1}Mm^{\text{gt}}_k \in \{0, 1\}^M is the binary ground-truth superpoint mask. A superpoint is assigned mk,jgt=1m^{\text{gt}}_{k, j} = 1 if more than 50% of its constituent raw points belong to ground-truth instance kk.

  4. Knowl 4 — SPFormer Multi-Task Training Objective and IoU-Aware Score Loss

    equation

    SPFormer trains all modules jointly using a multi-task loss L\mathcal{L} defined across all decoder layers via iterative prediction:

    L=βclsLcls+βsLs+βmask(Lbce+Ldice)\mathcal{L} = \beta_{\text{cls}} \mathcal{L}_{\text{cls}} + \beta_s \mathcal{L}_s + \beta_{\text{mask}} (\mathcal{L}_{\text{bce}} + \mathcal{L}_{\text{dice}})

    where weights are set to βcls=0.5\beta_{\text{cls}} = 0.5, βs=0.5\beta_s = 0.5, and βmask=1.0\beta_{\text{mask}} = 1.0.

    • Lcls\mathcal{L}_{\text{cls}} is the cross-entropy classification loss computed over all KK query proposals, treating unassigned proposals as a background ("no instance") class.
    • Lbce\mathcal{L}_{\text{bce}} and Ldice\mathcal{L}_{\text{dice}} are the binary cross-entropy and Dice losses computed on matched proposal-ground truth pairs.
    • Ls\mathcal{L}_s is an L2L_2 regression loss for an auxiliary score branch that predicts mask quality (IoU) to mitigate ranking misalignment caused by one-to-one Hungarian matching:

    Ls=1∑k=1NgtI{iouk>0.5}∑k=1NgtI{iouk>0.5}∥sk−iouk∥2\mathcal{L}_s = \frac{1}{\sum_{k=1}^{N_{\text{gt}}} \mathbb{I}_{\{\text{iou}_k > 0.5\}}} \sum_{k=1}^{N_{\text{gt}}} \mathbb{I}_{\{\text{iou}_k > 0.5\}} \|s_k - \text{iou}_k\|_2

    where sk∈[0,1]s_k \in [0, 1] is the predicted quality score, iouk\text{iou}_k is the superpoint mask intersection-over-union between the assigned proposal and its ground truth, NgtN_{\text{gt}} is the total number of ground-truth instances, and I{iouk>0.5}\mathbb{I}_{\{\text{iou}_k > 0.5\}} restricts loss calculation strictly to high-quality proposals whose IoU exceeds 50%.

  5. Knowl 5 — SPFormer Inference and Proposal Scoring Mechanism

    model/method

    During inference, SPFormer processes the point cloud through the backbone, superpoint pooling, and the final transformer decoder layer without any iterative refinement aggregation steps or Non-Maximum Suppression (NMS).

    For each of the KK query vectors ZZ, the shared prediction head generates:

    1. Class probabilities pi∈RNclass+1p_i \in \mathbb{R}^{N_{\text{class}} + 1}.
    2. IoU-aware quality scores si∈[0,1]s_i \in [0, 1].
    3. Superpoint mask predictions Mi=σ(Zi⋅SmaskT)∈[0,1]MM_i = \sigma(Z_i \cdot S_{\text{mask}}^T) \in [0, 1]^M, where σ\sigma denotes the sigmoid function.

    An additional mask confidence score msi∈[0,1]ms_i \in [0, 1] is calculated by averaging the probabilities of superpoints in mask MiM_i that exceed 0.5:

    msi=∑j=1MMi,j⋅I{Mi,j>0.5}∑j=1MI{Mi,j>0.5}ms_i = \frac{\sum_{j=1}^M M_{i, j} \cdot \mathbb{I}_{\{M_{i, j} > 0.5\}}}{\sum_{j=1}^M \mathbb{I}_{\{M_{i, j} > 0.5\}}}

    The final sorting score seise_i used to rank candidate instances is given by the geometric mean:

    sei=pi⋅si⋅msi3se_i = \sqrt[3]{p_i \cdot s_i \cdot ms_i}

    Point-level instance masks are retrieved by directly assigning the predicted superpoint labels to their respective constituent points.

  6. Knowl 6 — ScanNetv2 Benchmark Performance

    data/table

    SPFormer was evaluated on the ScanNetv2 3D instance segmentation benchmark (18 categories). On the hidden test set, SPFormer achieved 54.9% mAP, outperforming the previous state of the art by 4.3% absolute mAP, and setting top AP scores across 8 individual categories.

    Method mAP AP50\text{AP}_{50} AP25\text{AP}_{25}
    3D-BoNet 25.3 - -
    MTML 28.2 - -
    GICN 34.1 - -
    3D-MPA 35.5 - -
    Dyco3D 39.5 - -
    PE 39.6 - -
    PointGroup 40.7 - -
    HAIS 45.7 - -
    OccuSeg 48.6 - -
    SoftGroup 50.4 - -
    SSTNet 50.6 - -
    SPFormer (Ours) 54.9 - -

    On the ScanNetv2 validation set, SPFormer achieved 56.3% mAP, 73.9% AP50\text{AP}_{50}, and 82.9% AP25\text{AP}_{25}, exceeding the next best method (SSTNet at 49.4% mAP / SoftGroup at 46.0% mAP) by 6.9% mAP, 6.3% AP50\text{AP}_{50}, and 4.0% AP25\text{AP}_{25}.

  7. Knowl 7 — Inference Latency and Component-Wise Runtime Breakdown

    data/table

    Inference runtime per scene was benchmarked on the ScanNetv2 validation set using an NVIDIA RTX 3090 GPU, with sparse convolutional operations standardized using spconv v2.1.

    Method Component Time (ms) Total (ms)
    PointGroup Backbone (GPU): 48, Grouping (GPU+CPU): 218, ScoreNet (GPU): 106 372
    HAIS Backbone (GPU): 50, Hier. aggr. (GPU+CPU): 116, Intra-inst refinement (GPU): 90 256
    SoftGroup Backbone (GPU): 48, Soft grouping (GPU+CPU): 121, Top-down refinement (GPU): 97 266
    SSTNet S.p. extraction (CPU): 179, Backbone (GPU): 34, Tree Network (GPU+CPU): 148, ScoreNet (GPU): 58 419
    SPFormer S.p. extraction (CPU): 179, Backbone (GPU): 29, S.p. pooling (GPU): 18, Query decoder (GPU): 21 247

    By replacing complex point aggregation, tree splitting, and refinement sub-networks with a transformer query decoder (21 ms GPU time) and superpoint pooling (18 ms GPU time), SPFormer executes in 247 ms total per scan, achieving lower latency than all compared grouping- and proposal-based pipelines.

  8. Knowl 8 — Ablation of Core Architectural Components in SPFormer

    data/table

    An ablation study on the ScanNetv2 validation set assesses the contribution of superpoint pooling, mask-based bipartite matching versus bounding-box bipartite matching, and the IoU score prediction branch.

    Superpoint Pooling Matching Target Score Branch mAP AP50\text{AP}_{50} AP25\text{AP}_{25}
    Mask 34.3 54.7 72.9
    ✓ Box 49.9 68.1 78.9
    ✓ Mask 55.0 72.4 82.5
    ✓ Mask ✓ 56.3 73.9 82.9
    • Omitting superpoint pooling (feeding raw point backbone features directly into the transformer query decoder) reduces mAP from 55.0% to 34.3% because query cross-attention cannot effectively normalize across hundreds of thousands of point tokens.
    • Matching via 3D bounding boxes instead of superpoint masks drops performance by 5.1% mAP (49.9% vs 55.0%) due to unconstrained bounding box degrees of freedom and invisible geometric centers.
    • Adding the IoU-aware score branch improves mAP by +1.3% (55.0% to 56.3%) by realigning proposal confidence rankings.
  9. Knowl 9 — Ablation of Transformer Decoder Structural Design

    data/table

    The structural design choices of the SPFormer query decoder were ablated on the ScanNetv2 validation set against a standard transformer decoder baseline without position encodings.

    Iterative Prediction Attention Mask Position Encoding Cross-Attention First mAP AP50\text{AP}_{50} AP25\text{AP}_{25}
    51.0 69.6 79.8
    ✓ 52.5 71.4 81.6
    ✓ ✓ 56.0 73.3 82.6
    ✓ ✓ ✓ 55.6 72.7 82.0
    ✓ ✓ ✓ 56.3 73.9 82.9

    Key findings include:

    1. Superpoint attention masking Aℓ−1A_{\ell-1} provides the single largest gain (+3.5% mAP) by restricting query attention to predicted foreground regions.
    2. Query positional encodings degrade performance (-0.4% mAP) due to point cloud disorder and variable geometry.
    3. Swapping decoder sub-layers to perform superpoint cross-attention before self-attention provides a +0.3% mAP gain.
  10. Knowl 10 — Impact of Decoder Depth, Query Count, and Mask Loss Formulations

    empirical result

    Ablation experiments on ScanNetv2 validation set determine the sensitivity of SPFormer to transformer hyperparameters and mask loss choices:

    1. Decoder Depth: Testing decoder layer counts ℓ∈{1,3,6,12}\ell \in \{1, 3, 6, 12\} with K=400K = 400 query vectors yields mAPs of 49.1%, 54.7%, 56.3%, and 55.3%, respectively, indicating that 6 layers provide optimal capacity without over-parameterization.
    2. Query Vector Count: Evaluating K∈{100,200,400,800}K \in \{100, 200, 400, 800\} with 6 layers yields mAPs of 54.2%, 55.2%, 56.3%, and 55.9%, respectively. Model performance peaks around 400 queries, reflecting the higher instance density of 3D indoor scenes compared to typical 2D datasets.
    3. Mask Loss Components:
    • Focal loss alone: 23.1% mAP.
    • Binary cross-entropy (BCE) loss alone: 35.3% mAP.
    • Dice loss alone: 54.8% mAP.
    • Dice + Focal loss: 55.1% mAP.
    • Dice + BCE loss: 56.3% mAP (optimal configuration).

    Dice loss is essential for handling superpoint mask sparsity, and its combination with standard BCE produces the best overall segmentation quality.

Coverage note — No substantial contributed material was omitted. All key models, equations, algorithm mechanisms, benchmark data tables, runtime profiles, and architectural ablation studies were covered.

References

  1. 1.Armeni, I.; Sener, O.; Zamir, A. R.; Jiang, H.; Brilakis, I.; Fischer, M.; and Savarese, S. 2016. 3D Semantic Parsing of Large-Scale Indoor Spaces. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.
  2. 2.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-end object detection with transformers. In European conference on computer vision, 213–229. Springer.
  3. 3.Chen, S.; Fang, J.; Zhang, Q.; Liu, W.; and Wang, X. 2021. Hierarchical aggregation for 3d instance segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 15467–15476.
  4. 4.Cheng, B.; Misra, I.; Schwing, A. G.; Kirillov, A.; and Girdhar, R. 2022a. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1290–1299.
  5. 5.Cheng, B.; Schwing, A.; and Kirillov, A. 2021. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34: 17864–17875.
  6. 6.Cheng, T.; Wang, X.; Chen, S.; Zhang, W.; Zhang, Q.; Huang, C.; Zhang, Z.; and Liu, W. 2022b. Sparse Instance Activation for Real-Time Instance Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4433–4442.
  7. 7.Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T.; and Nießner, M. 2017. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, 5828–5839.
  8. 8.Dai, Z.; Cai, B.; Lin, Y.; and Chen, J. 2021. Up-detr: Unsupervised pre-training for object detection with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1601–1610.
  9. 9.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations.
  10. 10.Engelmann, F.; Bokeloh, M.; Fathi, A.; Leibe, B.; and Nießner, M. 2020. 3d-mpa: Multi-proposal aggregation for 3d semantic instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 9031–9040.
  11. 11.Fang, Y.; Yang, S.; Wang, X.; Li, Y.; Fang, C.; Shan, Y.; Feng, B.; and Liu, W. 2021. Instances as queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6910–6919.
  12. 12.Graham, B.; Engelcke, M.; and Van Der Maaten, L. 2018. 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 9224–9232.
  13. 13.Guo, R.; Niu, D.; Qu, L.; and Li, Z. 2021. Sotr: Segmenting objects with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 7157–7166.
  14. 14.He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, 2961–2969.
  15. 15.He, T.; Shen, C.; and van den Hengel, A. 2021. Dyco3d: Robust instance segmentation of 3d point clouds through dynamic convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 354–363.
  16. 16.Hou, J.; Dai, A.; and Nießner, M. 2019. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4421–4430.
  17. 17.Huang, Z.; Huang, L.; Gong, Y.; Huang, C.; and Wang, X. 2019. Mask scoring r-cnn. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6409–6418.
  18. 18.Jiang, L.; Zhao, H.; Shi, S.; Liu, S.; Fu, C.-W.; and Jia, J. 2020. Pointgroup: Dual-set point grouping for 3d instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4867–4876.
  19. 19.Kuhn, H. W. 1955. The Hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2): 83–97.
  20. 20.Lahoud, J.; Ghanem, B.; Pollefeys, M.; and Oswald, M. R. 2019. 3d instance segmentation via multi-task metric learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 9256–9266.
  21. 21.Landrieu, L.; and Simonovsky, M. 2018. Large-scale point cloud semantic segmentation with superpoint graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4558–4567.
  22. 22.Liang, Z.; Li, Z.; Xu, S.; Tan, M.; and Jia, K. 2021. Instance segmentation in 3d scenes using semantic superpoint tree networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2783–2792.
  23. 23.Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, 2980–2988.
  24. 24.Liu, S.-H.; Yu, S.-Y.; Wu, S.-C.; Chen, H.-T.; and Liu, T.-L. 2020. Learning gaussian instance segmentation in point clouds. arXiv preprint arXiv:2007.09860.
  25. 25.Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022.
  26. 26.Milletari, F.; Navab, N.; and Ahmadi, S.-A. 2016. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), 565–571. IEEE.
  27. 27.Narita, G.; Seno, T.; Ishikawa, T.; and Kaji, Y. 2019. Panopticfusion: Online volumetric semantic mapping at the level of stuff and things. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4205–4212. IEEE.
  28. 28.Park, K.-B.; Kim, M.; Choi, S. H.; and Lee, J. Y. 2020. Deep learning-based smart task assistance in wearable augmented reality. Robotics and Computer-Integrated Manufacturing, 63: 101887.
  29. 29.Qi, C. R.; Su, H.; Mo, K.; and Guibas, L. J. 2017a. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 652–660.
  30. 30.Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017b. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30.
  31. 31.Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, 10347–10357. PMLR.
  32. 32.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  33. 33.Vu, T.; Kim, K.; Luu, T. M.; Nguyen, T.; and Yoo, C. D. 2022. SoftGroup for 3D Instance Segmentation on Point Clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2708–2717.
  34. 34.Xie, C.; Xiang, Y.; Mousavian, A.; and Fox, D. 2021. Unseen object instance segmentation for robotic environments. IEEE Transactions on Robotics, 37(5): 1343–1359.
  35. 35.Yang, B.; Wang, J.; Clark, R.; Hu, Q.; Wang, S.; Markham, A.; and Trigoni, N. 2019. Learning object bounding boxes for 3D instance segmentation on point clouds. Advances in neural information processing systems, 32.
  36. 36.Yi, L.; Zhao, W.; Wang, H.; Sung, M.; and Guibas, L. J. 2019. Gspn: Generative shape proposal network for 3d instance segmentation in point cloud. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3947–3956.
  37. 37.Zhang, B.; and Wonka, P. 2021. Point cloud instance segmentation using probabilistic embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8883–8892.
  38. 38.Zhou, D.; Fang, J.; Song, X.; Liu, L.; Yin, J.; Dai, Y.; Li, H.; and Yang, R. 2020. Joint 3d instance segmentation and object detection for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1839–1849.

Citation

MLA
Sun, J., et al. “Superpoint Transformer for 3D Scene Instance Segmentation”. arXiv, 2022, http://arxiv.org/abs/2211.15766v1.
APA
Sun, J., Qing, C., Tan, J., & Xu, X. (2022). Superpoint Transformer for 3D Scene Instance Segmentation. arXiv. http://arxiv.org/abs/2211.15766v1
Chicago
Sun, J., C. Qing, J. Tan, and X. Xu. 2022. “Superpoint Transformer for 3D Scene Instance Segmentation”. arXiv. http://arxiv.org/abs/2211.15766v1.
Harvard
Sun, J. et al. (2022) “Superpoint Transformer for 3D Scene Instance Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.15766v1.
Vancouver
1. Sun J, Qing C, Tan J, Xu X (2022) Superpoint Transformer for 3D Scene Instance Segmentation. arXiv

BibTeX

@article{sun2022superpoint,
  title = {Superpoint Transformer for 3D Scene Instance Segmentation},
  author = {Sun, Jiahao and Qing, Chunmei and Tan, Junpeng and Xu, Xiangmin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.15766v1},
  eprint = {2211.15766}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF