RBGNet: Ray-based Grouping for 3D Object Detection
Haiyang WangShaoshuai ShiZe YangRongyao FangQi QianHongsheng LiBernt SchieleLiwei Wang
Presents RBGNet, a point-based 3D object detection framework that improves bounding box prediction on point clouds by combining foreground-biased point sampling with ray-based feature grouping to model object surface geometry.
Three-dimensional object detection from spatial point clouds is a critical capability for autonomous systems, robotics, and augmented reality. Point clouds collected by depth sensors are naturally sparse, irregular, and unorganized, making it difficult for automated systems to accurately identify object boundaries and orientations. Existing point-based detection methods aggregate local points into object candidates, but they largely overlook fine-grained surface geometry and frequently waste computational capacity by sampling empty background areas instead of informative foreground surfaces.
The article demonstrates a single-stage 3D object detection framework, named RBGNet, designed to improve 3D bounding box estimation from raw point clouds. The core objective is to evaluate whether explicitly capturing foreground surface geometry and concentrating point sampling on object surfaces can significantly enhance detection accuracy without sacrificing processing efficiency.
To achieve this, the approach introduces two primary mechanisms into a voting-based detection architecture. First, a ray-based feature grouping module emits a uniform set of rays outward from predicted candidate centers to sample anchor points along object surfaces in a coarse-to-fine manner, effectively capturing geometric shape features. Second, a foreground biased sampling strategy classifies points early in the network and allocates 87.5% of downsampled points to probable foreground objects while retaining 12.5% from the background to preserve overall scene context. The framework was evaluated on two benchmark indoor datasets, ScanNet V2 and SUN RGB-D, across standard average precision metrics.
The experimental findings show that the proposed framework sets a new state of the art in 3D object detection. On the ScanNet V2 benchmark, the baseline model achieved a detection accuracy of 70.2% at a 0.25 intersection-over-union threshold and 54.2% at a 0.50 threshold, outperforming prior leading point-based methods by 2.5 and 3.3 percentage points, respectively. On the SUN RGB-D benchmark, it reached 64.1% accuracy at the 0.25 threshold, surpassing all previous geometry-only detectors. Ablation experiments revealed that adding ray-based grouping alone improved accuracy by several percentage points, and increasing ray density consistently enhanced object surface point recovery. Furthermore, the foreground sampling strategy drastically improved the concentration of object points during processing (reaching 87.8% foreground concentration in deep layers compared to 30.3% in standard sampling) while maintaining competitive inference speeds of 4.75 to 7.23 frames per second.
These results demonstrate that surface geometry and biased sampling provide vital spatial cues that resolve bounding box ambiguities without requiring multi-modal inputs such as standard RGB images. By generating tighter, more reliable 3D bounding boxes at competitive processing speeds, the approach enhances the safety, situational awareness, and operational precision of robotic navigation and spatial computing systems.
Organizations developing 3D perception pipelines should consider adopting foreground biased sampling and ray-based geometric grouping to upgrade point-cloud processing backbones. Engineering teams can select between ray configurations (such as 6 rays for higher frame rates versus 66 rays for maximum detection accuracy) to balance real-time latency requirements against detection performance. Future work should focus on validating the framework on outdoor autonomous driving environments, exploring edge-device optimizations, and evaluating performance under severe sensor noise or partial object occlusions.
- Paper: Deep Hough Voting for 3D Object Detection in Point Clouds, Charles R. Qi et al. (2019). Introduces VoteNet's deep Hough voting mechanism for generating object center proposals from sparse point clouds, which RBGNet directly builds upon with its ray-based grouping and biased sampling strategies.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). Presents hierarchical feature learning and sampling on point sets via PointNet++, providing the standard point-processing backbone foundational to vote-based 3D object detection architectures.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). Establishes the foundational deep learning architecture for directly processing irregular and unordered 3D point sets with permutation invariance.
- Paper: PointRCNN: 3D Object Proposal Generation and Detection From Point Cloud, Shaoshuai Shi et al. (2019). Introduces foreground point segmentation and direct 3D box proposal generation from point clouds, underlying the foreground-biased sampling principles utilized in RBGNet.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). Demonstrates keypoint-based set abstraction and multi-scale feature pooling for point cloud bounding box refinement, motivating RBGNet's fine-grained surface geometric grouping.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). Extends 3D scene understanding by grouping raw points into geometric superpoints and leveraging attention mechanisms for direct end-to-end instance segmentation without candidate bounding boxes.
- Paper: CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection, Yang Cao et al. (2023). Generalizes indoor 3D object detection to open-vocabulary settings by discovering novel 3D bounding boxes and aligning point cloud features with multimodal language priors.
- Paper: Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection, Bo Zhang et al. (2023). Addresses cross-dataset generalizability in 3D object detection by unifying feature representations and detection heads across diverse sensor setups and benchmark taxonomies.
- Paper: BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection, Zhenxin Li et al. (2024). Advances 3D object detection by reviving dense bird's-eye-view representations with perspective cross-attention and temporal feature fusion.
- Paper: RADIANT: Radar-Image Association Network for 3D Object Detection, Yunfei Long et al. (2023). Applies geometric offset learning between sparse radar returns and physical object centroids to improve 3D bounding box estimation accuracy.
