Built independently by an author, for readers. Read the story and support ChapterPal

keyword

S3DIS dataset

The S3DIS dataset, short for the Stanford Large-Scale 3D Indoor Spaces dataset, is a benchmark point cloud dataset used in computer vision and machine learning for 3D indoor scene understanding tasks such as semantic and instance segmentation. Captured using 3D scanners across several university buildings, the dataset encompasses over 6,000 square meters of indoor space spanning hundreds of rooms across six distinct large-scale areas, including offices, conference rooms, hallways, and common areas. Each point in the scans includes three-dimensional spatial coordinates and RGB color information, paired with detailed point-level annotations corresponding to thirteen structural and object classes such as walls, floors, ceilings, chairs, tables, and clutter. Due to its scale, architectural diversity, and comprehensive annotations, it serves as a standard benchmark for evaluating neural network architectures and 3D perception algorithms.

4 items

SoftGroup for 3D Instance Segmentation on Point Clouds

SoftGroup for 3D Instance Segmentation on Point Clouds

Thang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen, Chang D. Yoo

OrganizationsKorea Advanced Institute of Science and Technology

Why you should read this

Proposes SoftGroup, a 3D instance segmentation method that associates points with multiple semantic classes during bottom-up grouping to prevent error propagation, achieving substantial accuracy gains and fast inference on ScanNet v2 and S3DIS.

Existing state-of-the-art 3D instance segmentation methods perform semantic segmentation followed by grouping. The hard predictions are made when performing semantic segmentation such that each point is associated with a single class. However, the errors stemming from hard decision propagate into grouping that results in (1) low overlaps between the predicted instance with the ground truth and (2) substantial false positives. To address the aforementioned problems, this paper proposes a 3D instance segmentation method referred to as SoftGroup by performing bottom-up soft grouping followed by top-down refinement. SoftGroup allows each point to be associated with multiple classes to mitigate the problems stemming from semantic prediction errors and suppresses false positive instances by learning to categorize them as background. Experimental results on different datasets and multiple evaluation metrics demonstrate the efficacy of SoftGroup. Its performance surpasses the strongest prior method by a significant margin of +6.2% on the ScanNet v2 hidden test set and +6.8% on S3DIS Area 5 in terms of AP_50. SoftGroup is also fast, running at 345ms per scan with a single Titan X on ScanNet v2 dataset. The source code and trained models for both datasets are available at \url{this https URL}.

Added

2026-09-26

Fast Point Transformer

Fast Point Transformer

Chunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik Park

OrganizationsPohang University of Science and Technology

Why you should read this

Proposes a lightweight local self-attention architecture with centroid-aware voxelization and hashing that encodes continuous 3D coordinates to process large-scale point clouds over a hundred times faster than standard point transformers.

The recent success of neural networks enables a better interpretation of 3D point clouds, but processing a large-scale 3D scene remains a challenging problem. Most current approaches divide a large-scale scene into small regions and combine the local predictions together. However, this scheme inevitably involves additional stages for pre- and post-processing and may also degrade the final output due to predictions in a local perspective. This paper introduces Fast Point Transformer that consists of a new lightweight self-attention layer. Our approach encodes continuous 3D coordinates, and the voxel hashing-based architecture boosts computational efficiency. The proposed method is demonstrated with 3D semantic segmentation and 3D detection. The accuracy of our approach is competitive to the best voxel-based method, and our network achieves 129 times faster inference time than the state-of-the-art, Point Transformer, with a reasonable accuracy trade-off in 3D semantic segmentation on S3DIS dataset.

Added

2026-09-26

ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution

ISBNet: a 3D Point Cloud Instance Segmentation Network with Instance-aware Sampling and Box-aware Dynamic Convolution

Tuan Duc Ngo, Binh-Son Hua, Khoi Nguyen

OrganizationsVinAI Research

Why you should read this

Proposes a cluster-free 3D instance segmentation framework that couples instance-aware point sampling with box-guided dynamic convolutions to achieve state-of-the-art accuracy and fast inference on ScanNetV2, S3DIS, and STPLS3D benchmarks.

Existing 3D instance segmentation methods are predomi- nated by the bottom-up design – manually fine-tuned algo- rithm to group points into clusters followed by a refinement network. However, by relying on the quality of the clus- ters, these methods generate susceptible results when (1) nearby objects with the same semantic class are packed together, or (2) large objects with loosely connected re- gions. To address these limitations, we introduce ISBNet, a novel cluster-free method that represents instances as ker- nels and decodes instance masks via dynamic convolution. To efficiently generate high-recall and discriminative ker- nels, we propose a simple strategy named Instance-aware Farthest Point Sampling to sample candidates and lever- age the local aggregation layer inspired by PointNet++ to encode candidate features. Moreover, we show that pre- dicting and leveraging the 3D axis-aligned bounding boxes in the dynamic convolution further boosts performance. Our method set new state-of-the-art results on ScanNetV2 (55.9), S3DIS (60.8), and STPLS3D (49.2) in terms of AP and retains fast inference time (237ms per scene on Scan- NetV2). The source code and trained models are available at https://github.com/VinAIResearch/ISBNet.

Added

2026-09-26