ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution
Tuan Duc NgoBinh-Son HuaKhoi Nguyen
Proposes a cluster-free 3D instance segmentation framework that couples instance-aware point sampling with box-guided dynamic convolutions to achieve state-of-the-art accuracy and fast inference on ScanNetV2, S3DIS, and STPLS3D benchmarks.
Three-dimensional instance segmentation—identifying individual objects and assigning semantic categories within 3D point cloud data—is a critical capability for technologies such as autonomous driving, augmented reality, and robotics. Existing bottom-up methods rely heavily on centroid clustering to group raw data points into distinct objects. However, these conventional approaches often fail when identical items are tightly packed together or when large, loosely connected objects are incorrectly fragmented into separate pieces.
The article introduces and evaluates ISBNet, a cluster-free framework designed to segment 3D point clouds accurately and efficiently. The primary objective is to demonstrate that combining instance-aware point sampling with box-aware dynamic convolution overcomes the clustering failures of prior systems and establishes superior performance across benchmark 3D datasets.
The authors evaluated ISBNet through empirical experiments on three standard indoor and outdoor datasets: ScanNetV2, S3DIS, and the aerial photogrammetry dataset STPLS3D. ISBNet replaces traditional grouping heuristics with an iterative sampling strategy that selects object candidate points while avoiding background and previously detected instances. These candidates aggregate local contextual features and jointly predict object bounding boxes. The network then integrates these predicted 3D bounding boxes as explicit geometric cues during dynamic convolution to isolate and generate final instance masks.
The findings show that ISBNet achieves state-of-the-art accuracy across all evaluated benchmarks while maintaining high processing speeds. First, ISBNet reached an average precision of 55.9 on the hidden ScanNetV2 test benchmark and exceeded the second-best method on the validation set by 3.7 points in average precision. Second, it surpassed existing state-of-the-art models on S3DIS cross-validation and STPLS3D, beating prior benchmarks by 3.4 and 3.0 points in average precision, respectively. Third, the proposed sampling technique achieved up to 100% instance candidate recall on ScanNetV2 validation data, compared to only 75.5% achieved by clustering methods. Finally, ISBNet demonstrated high computational efficiency, executing a complete scene analysis in 237 milliseconds on a single graphics processing unit, making it the fastest among evaluated state-of-the-art alternatives.
These results demonstrate that eliminating hand-tuned clustering in favor of direct instance-aware candidate sampling and bounding box geometric cues significantly enhances segmentation reliability. For real-world systems, this improvement reduces the risk of misidentifying overlapping obstacles or fragmenting complex assets, while the sub-second runtime supports the strict latency requirements necessary for safe, real-time autonomous navigation and mapping.
Organizations developing 3D perception pipelines should consider adopting cluster-free, dynamic convolution architectures over traditional clustering approaches when processing dense spatial data. Further development should focus on testing these models across additional real-world operational environments and exploring enhanced dynamic convolutions that incorporate richer geometric structures.
The confidence in these findings is strong across the evaluated standard datasets, though some limitations remain. The sampling step depends on the quality of intermediate instance predictions, meaning early errors could propagate through the pipeline. Additionally, axis-aligned bounding boxes do not always tightly fit irregularly shaped or intertwined objects, which can cause adjacent, touching objects—such as a counter and a refrigerator—to be erroneously merged.
- Paper: SoftGroup for 3D Instance Segmentation on Point Clouds, Thang Vu et al. (2022). It establishes the two-stage grouping and top-down mask refinement paradigm on point clouds that ISBNet directly aims to streamline and replace with cluster-free dynamic convolution.
- Paper: Deep Hough Voting for 3D Object Detection in Point Clouds, Charles R. Qi et al. (2019). It introduces deep Hough voting and candidate bounding box regression on raw 3D point sets, foundational concepts that ISBNet leverages for object candidate generation and box-aware cues.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). It provides the submanifold sparse convolutional backbone architecture widely used by 3D segmentation frameworks like ISBNet to efficiently extract voxel- and point-level features.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). It introduces hierarchical point set feature learning and farthest-point sampling strategies essential for understanding spatial candidate sampling in point clouds.
- Paper: RBGNet: Ray-based Grouping for 3D Object Detection, Haiyang Wang et al. (2022). It explores foreground-biased sampling and ray-based geometric grouping for 3D bounding box prediction, directly preceding ISBNet's instance-aware sampling strategy.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). It provides the foundational permutation-invariant deep learning architecture for directly processing raw, unordered 3D point sets.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). It explores an alternative end-to-end transformer-based approach for 3D scene instance segmentation that bypasses traditional clustering via superpoint query decoding.
- Paper: GARField: Group Anything with Radiance Fields, Chung Min Kim et al. (2024). It extends 3D instance grouping into multi-scale radiance fields and Gaussian splatting for open-ended, hierarchical scene decomposition.
- Paper: Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training, Xiaoyang Wu et al. (2024). It advances 3D segmentation architectures toward large-scale multi-dataset representation learning and prompt-based cross-domain generalization.
- Paper: Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields, Shijie Zhou et al. (2024). It builds upon explicit geometric representations to distill 2D foundation model features into 3D Gaussian splats for promptable 3D instance and semantic segmentation.
