CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds
Haiyang WangLihe DingShaocong DongShaoshuai ShiAoxue LiJianan LiZhenguo LiLiwei Wang
Introduces a two-stage fully sparse convolutional framework for 3D indoor object detection that overcomes bottom-up grouping errors by enforcing semantic consistency during proposal generation and using an efficient sparse RoI pooling module to recover missed geometric features directly from the backbone.
Detecting 3D objects from raw, irregular 3D point cloud data is critical for emerging technologies such as autonomous driving, robotics, and augmented reality. In complex indoor environments, existing detection frameworks typically group nearby points together without considering their object category. This class-agnostic grouping often leads to errors in cluttered spaces, such as merging parts of adjacent but unrelated objects or using search areas that fail to capture the boundaries of large objects while adding noise to small ones. Furthermore, traditional secondary refinement modules rely on complex, memory-heavy pooling operations that degrade fine-grained structural information.
The article evaluates CAGroup3D, a fully convolutional two-stage 3D object detection framework designed to overcome these grouping and refinement weaknesses. The authors demonstrate how pairing category-aware point clustering with a lightweight, convolution-based refinement step substantially boosts 3D detection precision.
To demonstrate this capability, the authors conducted rigorous empirical experiments on two standard indoor 3D benchmarks: ScanNet V2, which contains rich 3D mesh scans across 18 object categories, and SUN RGB-D, which includes over 10,000 single-view RGB-D images across 10 categories. The framework first uses a dual-resolution 3D sparse convolutional network to extract detailed spatial features. It then applies a class-aware grouping strategy that shifts points toward estimated object centers and groups only points sharing the same predicted category within search regions scaled to average category dimensions. Finally, a fully sparse convolutional region pooling module directly samples and refines object bounding boxes from the extracted feature maps without relying on traditional max-pooling operations.
The evaluation produced four key findings in order of importance. First, CAGroup3D established a new state of the art in 3D object detection, achieving 75.1% mean average precision at an intersection-over-union threshold of 0.25 on ScanNet V2—a 3.6 percentage point gain over the prior leading method—and 66.8% on SUN RGB-D, outperforming both single-modality and multi-modal camera-plus-depth methods. Second, adding class-aware local grouping accounted for the largest performance leap, raising precision from 68.22% to 72.10% over the baseline. Third, the proposed region pooling module improved bounding box accuracy during the refinement stage, increasing precision at the stricter 0.50 threshold from 57.18% to 60.31%. Fourth, the new pooling mechanism reduced GPU memory consumption to 2,468 megabytes, requiring less than one-third of the memory consumed by common alternative pooling approaches.
These results demonstrate that incorporating category awareness directly into point grouping and utilizing memory-efficient sparse convolutions can substantially improve both detection accuracy and hardware efficiency. For operational systems in robotics and smart spaces, these improvements mean safer, more reliable spatial awareness and lower computing resource demands. Organizations building spatial perception pipelines should adopt category-adaptive grouping and sparse convolutional pooling architectures over older, class-agnostic voting pipelines.
The findings are supported with high confidence through multiple independent experimental trials. However, the current framework is primarily designed to address variations across different categories and does not account for size and shape variations among objects within the same category. Future development should focus on addressing this within-class variability to further improve real-world robustness.
- Paper: SoftGroup for 3D Instance Segmentation on Point Clouds, Thang Vu et al. (2022). Introduces soft semantic grouping and top-down refinement for point clouds to tackle bottom-up grouping errors, directly contextualizing CAGroup3D's class-aware grouping strategy.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). Pioneers submanifold sparse convolutions on 3D grids, establishing the foundational sparse backbone architecture utilized throughout CAGroup3D.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). Establishes a high-performing two-stage 3D detection paradigm with spatial RoI pooling that informs CAGroup3D's proposal generation and sparse RoI refinement.
- Paper: Deep Hough Voting for 3D Object Detection in Point Clouds, Charles R. Qi et al. (2019). Formulates deep Hough voting for generating 3D bounding box proposals directly from point clouds on indoor benchmarks like ScanNet and SUN RGB-D.
- Paper: PointRCNN: 3D Object Proposal Generation and Detection From Point Cloud, Shaoshuai Shi et al. (2019). Introduces the concept of bottom-up foreground point segmentation coupled with canonical 3D box proposal refinement.
- Paper: VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection, Yin Zhou et al. (2017). Presents the original voxel feature encoding and end-to-end 3D convolutional network pipeline for point-cloud-based detection.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). Advances bottom-up grouping for 3D point clouds by pooling points into superpoints and using attention-based query decoders for end-to-end prediction.
- Paper: OcTr: Octree-Based Transformer for 3D Object Detection, Chao Zhou et al. (2023). Extends sparse point cloud object detection by using dynamic octrees with transformer mechanisms to enhance receptive fields and long-range feature capture.
- Paper: Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection, Bo Zhang et al. (2023). Generalizes point-cloud 3D object detection pipelines across diverse datasets and sensor modalities using semantic coupling and data-level alignment.
- Paper: GD-MAE: Generative Decoder for MAE Pre-Training on LiDAR Point Clouds, Honghui Yang et al. (2023). Introduces generative masked autoencoding on sparse LiDAR representations to significantly boost downstream 3D object detection backbones.
