Superpoint Transformer for 3D Scene Instance Segmentation
Jiahao SunChunmei QingJunpeng TanXiangmin Xu
Proposes an end-to-end 3D instance segmentation framework that uses superpoint cross-attention within a transformer decoder to directly predict instance masks without relying on intermediate grouping steps or complex post-processing.
Accurate three-dimensional scene understanding is essential for emerging technologies such as autonomous driving, robotics, and augmented reality. A critical component of this understanding is three-dimensional instance segmentation, which involves identifying individual objects within spatial point cloud data and delineating their exact boundaries. Existing approaches typically rely on either generating bounding boxes or performing bottom-up point aggregation based on semantic predictions. However, these methods suffer from performance bottlenecks caused by inaccurate bounding boxes, error propagation from intermediate semantic predictions, and time-consuming aggregation steps.
The article evaluates a unified framework called Superpoint Transformer, or SPFormer, designed to perform end-to-end instance segmentation directly on three-dimensional scenes without relying on object detection bounding boxes or intermediate semantic segmentation tasks.
To overcome computational limits, the approach combines a bottom-up grouping stage with a top-down prediction pipeline. Raw three-dimensional points are first processed through a sparse neural network and pooled into geometric clusters called superpoints, which reduce hundreds of thousands of individual points into several hundred compact representations. A query decoder utilizing attention mechanisms then uses learnable vectors to focus on relevant superpoints and generate object masks, categories, and confidence scores. Optimal one-to-one matching between predictions and actual objects is performed using mask comparisons, eliminating the need for separate point aggregation and post-processing filters.
Evaluations across standard benchmark datasets demonstrate several key findings. First, the proposed framework achieves a mean average precision of 54.9% on the ScanNetv2 hidden test benchmark, exceeding the previous state of the art by 4.3 percentage points. Second, on the ScanNetv2 validation set, it achieves a mean average precision of 56.3%, outperforming the previous best result by 6.9 percentage points. Third, the system demonstrates the fastest processing time among evaluated models at 247 milliseconds per scene, even when accounting for initial superpoint generation on the processor. Finally, additional tests on the S3DIS indoor dataset confirm strong performance across diverse environments, achieving a leading 66.8% precision score at a 50% overlap threshold on Area 5.
These results demonstrate that direct mask matching and superpoint-based attention can deliver superior accuracy while reducing computational latency. For organizations deploying spatial vision systems, eliminating multi-stage aggregation and complex post-processing reduces algorithmic complexity, memory requirements, and processing overhead. This enables faster decision-making cycles in safety-critical applications like automated navigation and robotics.
Engineering and research teams should consider adopting direct superpoint-query architectures when designing real-time spatial segmentation pipelines. Further work should explore optimizing initial superpoint computation on graphics hardware to reduce overall inference latency even further, as well as evaluating model performance on large-scale outdoor datasets and diverse sensor configurations. While confidence in the indoor benchmark results is high, practitioners should validate the method under varying point densities and sensor noise conditions before full deployment.
- Paper: Large-Scale Point Cloud Semantic Segmentation with Superpoint Graphs, Loic Landrieu et al. (2017). This paper establishes the concept of partitioning large-scale point clouds into geometrically homogeneous superpoints to structure 3D scene segmentation, which SPFormer adapts into its superpoint-level query transformer.
- Paper: SoftGroup for 3D Instance Segmentation on Point Clouds, Thang Vu et al. (2022). SoftGroup represents the prevailing bottom-up grouping and multi-stage refinement paradigm for 3D instance segmentation that SPFormer directly targets and outperforms by replacing multi-stage aggregation with direct mask matching.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). This work introduces query-based masked-attention mechanisms for direct mask classification, providing the conceptual foundation for SPFormer's query decoder on 3D superpoint representations.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer pioneers the universal mask classification paradigm using learnable queries and bipartite matching, which SPFormer translates into an end-to-end framework for 3D point cloud instance segmentation.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). It introduces submanifold sparse convolutional networks, the foundational sparse neural network backbone used to process raw 3D coordinates efficiently before superpoint pooling.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). This paper formulates the generalized sparse tensor convolutions widely adopted by modern 3D point cloud backbones to extract features from spatial sensor data.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). This foundational paper presents the first deep learning architecture designed to process raw, unordered 3D point sets directly with permutation invariance.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). This work extends 3D instance and panoptic scene understanding by lifting multi-view 2D predictions into consistent 3D volumetric neural fields without relying exclusively on supervised 3D point cloud grouping.
- Paper: CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data, Yihan Zeng et al. (2023). CLIP2 generalizes 3D point cloud perception to open-vocabulary recognition by aligning real-world 3D clusters directly with natural language representations.
