SoftGroup for 3D Instance Segmentation on Point Clouds
Thang VuKookhoi KimTung Minh LuuThanh Xuan NguyenChang D. Yoo
Proposes SoftGroup, a 3D instance segmentation method that associates points with multiple semantic classes during bottom-up grouping to prevent error propagation, achieving substantial accuracy gains and fast inference on ScanNet v2 and S3DIS.
Three-dimensional instance segmentation—identifying individual objects and their exact boundaries within 3D point cloud data—is a critical computer vision capability for autonomous driving, robotics, and augmented reality. Prevailing approaches rely on a bottom-up process that assigns each point strictly to a single object class before grouping nearby points into instances. However, because local parts of objects can be ambiguous, early classification mistakes propagate directly into the grouping phase. This creates incomplete object boundaries and generates false-positive detections, undermining perception accuracy.
The article demonstrates and evaluates SoftGroup, a two-stage method designed to resolve this error propagation. The primary objective is to improve 3D instance segmentation accuracy by enabling flexible, multi-class point associations followed by targeted object refinement.
The approach integrates a bottom-up grouping stage with a top-down refinement stage. Rather than committing each point to a single label, the bottom-up stage uses a soft probability threshold to allow ambiguous points to associate with multiple candidate classes, creating preliminary point-level proposals. The top-down stage then extracts features from each proposal and processes them through classification, segmentation, and mask-scoring branches to refine valid objects and classify erroneous predictions as background. The authors validated SoftGroup on two standard benchmarks, the ScanNet v2 dataset of indoor scans and the Stanford Large-Scale 3D Indoor Spaces (S3DIS) dataset, measuring precision across varying overlap thresholds.
The evaluation yielded several key findings. First, SoftGroup established a new state of the art, outperforming previous leading methods on the ScanNet v2 hidden test set by 6.2 percentage points in 50% overlap average precision (AP50), reaching 76.1%, and leading in 12 of 18 object categories. Second, on S3DIS Area 5, it outperformed the second-best method by 6.8 percentage points in AP50 (reaching 66.1%) and 8.9 percentage points in overall average precision. Third, ablation experiments demonstrated that the score threshold for soft grouping (optimized at 0.2) and the multi-branch top-down refinement work synergistically, boosting baseline performance by 6.5 percentage points. Finally, the system achieved this performance while maintaining high computational efficiency, processing a full scan in 345 milliseconds on standard hardware.
These findings demonstrate that deferring definitive classification until after proposal generation effectively resolves long-standing error propagation issues without sacrificing execution speed. For real-world systems in robotics and automated navigation, this translates to improved operational safety and reliability through reduced false alarms and better object boundary definition, with negligible latency trade-offs.
Organizations developing 3D perception pipelines should consider adopting two-stage soft-grouping frameworks over strict single-label grouping. As next steps, technical teams can utilize the authors' open-source codebase and pre-trained models to reproduce the results and evaluate performance on domain-specific 3D data. Further testing in outdoor environments and under dynamic conditions is recommended, as the study focuses strictly on benchmark indoor datasets.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). Submanifold Sparse Convolutional Networks provide the sparse 3D convolutional backbone and coordinate-based processing fundamental to voxel-based point cloud segmentation pipelines like SoftGroup.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). PointNet++ introduces hierarchical spatial aggregation and grouping principles on raw point sets that underpin modern bottom-up point cloud segmentation architectures.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet establishes the foundational deep learning architecture for directly processing unordered 3D point sets with permutation invariance.
- Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN establishes the paradigm of instance segmentation and proposal-based refinement that SoftGroup adapts into a top-down refinement stage for 3D point clouds.
- Paper: Deep Hough Voting for 3D Object Detection in Point Clouds, Charles R. Qi et al. (2019). VoteNet demonstrates deep geometric clustering and voting mechanisms in 3D point clouds, which directly inform instance-level grouping and centroid estimation.
- Paper: Hybrid Task Cascade for Instance Segmentation, Kai Chen et al. (2019). Hybrid Task Cascade introduces multi-stage progressive refinement linking bounding proposals and masks, motivating the top-down refinement design in SoftGroup.
- Paper: Large-Scale Point Cloud Semantic Segmentation with Superpoint Graphs, Loic Landrieu et al. (2017). This paper establishes superpoint graphs and geometric partitioning for point clouds, providing the conceptual foundation for grouping-based 3D scene segmentation.
- Paper: PointCNN: Convolution On X-Transformed Points, Yangyan Li et al. (2018). PointCNN presents an effective framework for applying convolutional operations directly on irregular 3D point clouds via X-transformations.
- Paper: Symphonize 3D Semantic Scene Completion with Contextual Instance Queries, Haoyi Jiang et al. (2024). Symphonies extends 3D instance-centric reasoning by incorporating contextual instance queries into 3D semantic scene completion.
