RBGNet: Ray-based Grouping for 3D Object Detection

Haiyang WangShaoshuai ShiZe YangRongyao FangQi QianHongsheng LiBernt SchieleLiwei Wang

article2022CVPR69 citations

Presents RBGNet, a point-based 3D object detection framework that improves bounding box prediction on point clouds by combining foreground-biased point sampling with ray-based feature grouping to model object surface geometry.

Listen

Three-dimensional object detection from spatial point clouds is a critical capability for autonomous systems, robotics, and augmented reality. Point clouds collected by depth sensors are naturally sparse, irregular, and unorganized, making it difficult for automated systems to accurately identify object boundaries and orientations. Existing point-based detection methods aggregate local points into object candidates, but they largely overlook fine-grained surface geometry and frequently waste computational capacity by sampling empty background areas instead of informative foreground surfaces.

The article demonstrates a single-stage 3D object detection framework, named RBGNet, designed to improve 3D bounding box estimation from raw point clouds. The core objective is to evaluate whether explicitly capturing foreground surface geometry and concentrating point sampling on object surfaces can significantly enhance detection accuracy without sacrificing processing efficiency.

To achieve this, the approach introduces two primary mechanisms into a voting-based detection architecture. First, a ray-based feature grouping module emits a uniform set of rays outward from predicted candidate centers to sample anchor points along object surfaces in a coarse-to-fine manner, effectively capturing geometric shape features. Second, a foreground biased sampling strategy classifies points early in the network and allocates 87.5% of downsampled points to probable foreground objects while retaining 12.5% from the background to preserve overall scene context. The framework was evaluated on two benchmark indoor datasets, ScanNet V2 and SUN RGB-D, across standard average precision metrics.

The experimental findings show that the proposed framework sets a new state of the art in 3D object detection. On the ScanNet V2 benchmark, the baseline model achieved a detection accuracy of 70.2% at a 0.25 intersection-over-union threshold and 54.2% at a 0.50 threshold, outperforming prior leading point-based methods by 2.5 and 3.3 percentage points, respectively. On the SUN RGB-D benchmark, it reached 64.1% accuracy at the 0.25 threshold, surpassing all previous geometry-only detectors. Ablation experiments revealed that adding ray-based grouping alone improved accuracy by several percentage points, and increasing ray density consistently enhanced object surface point recovery. Furthermore, the foreground sampling strategy drastically improved the concentration of object points during processing (reaching 87.8% foreground concentration in deep layers compared to 30.3% in standard sampling) while maintaining competitive inference speeds of 4.75 to 7.23 frames per second.

These results demonstrate that surface geometry and biased sampling provide vital spatial cues that resolve bounding box ambiguities without requiring multi-modal inputs such as standard RGB images. By generating tighter, more reliable 3D bounding boxes at competitive processing speeds, the approach enhances the safety, situational awareness, and operational precision of robotic navigation and spatial computing systems.

Organizations developing 3D perception pipelines should consider adopting foreground biased sampling and ray-based geometric grouping to upgrade point-cloud processing backbones. Engineering teams can select between ray configurations (such as 6 rays for higher frame rates versus 66 rays for maximum detection accuracy) to balance real-time latency requirements against detection performance. Future work should focus on validating the framework on outdoor autonomous driving environments, exploring edge-device optimizations, and evaluating performance under severe sensor noise or partial object occlusions.

arXiv: 2204.02251
Cover for RBGNet: Ray-based Grouping for 3D Object Detection

Abstract

As a fundamental problem in computer vision, 3D object detection is experiencing rapid growth. To extract the point-wise features from the irregularly and sparsely distributed points, previous methods usually take a feature grouping module to aggregate the point features to an object candidate. However, these methods have not yet leveraged the surface geometry of foreground objects to enhance grouping and 3D box generation. In this paper, we propose the RBGNet framework, a voting-based 3D detector for accurate 3D object detection from point clouds. In order to learn better representations of object shape to enhance cluster features for predicting 3D boxes, we propose a ray-based feature grouping module, which aggregates the point-wise features on object surfaces using a group of determined rays uniformly emitted from cluster centers. Considering the fact that foreground points are more meaningful for box estimation, we design a novel foreground biased sampling strategy in downsample process to sample more points on object surfaces and further boost the detection performance. Our model achieves state-of-the-art 3D detection performance on ScanNet V2 and SUN RGB-D with remarkable performance gains. Code will be available at https://github.com/Haiyang-W/RBGNet.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Overview
  • 3.2. Ray-based Feature Grouping
  • 3.2.1 Ray Point Representation
  • 3.2.2 Feature Enhancement by Determined Rays
  • 3.3. Foreground Biased Sampling
  • 3.4. Learning Objective
  • 4. Experiments
  • 4.1. Datasets and Evaluation Metric
  • 4.2. Implementation Details
  • 4.3. Comparison with state-of-the-art methods
  • 4.4. Ablation Studies and Discussions
  • 4.5. Inference Speed
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — RBGNet one-stage detection architecture

    model/method

    RBGNet is a one-stage point-cloud 3D detector built on a VoteNet-style voting pipeline. It receives 50,000 three-dimensional scene points, extracts hierarchical point features with a PointNet++ backbone, replaces farthest-point sampling in the second through fourth set-abstraction layers with Foreground Biased Sampling, and uses voting to form candidate object centers with associated cluster features. A ray-based feature grouping module operates around each candidate center, encodes the candidate object surface, and augments the cluster feature used for 3D box regression and semantic classification. The resulting proposals are filtered with 3D non-maximum suppression. The standard RBGNet configuration uses 66 rays per candidate and 256 object candidates; a larger configuration uses 512 candidates and twice the PointNet++ channel capacity.

  2. Knowl 2 — Determined spherical ray representation

    equation

    For each vote cluster, RBGNet represents a candidate as ci=[vi,fi]c_i=[v_i,f_i], where vi∈R3v_i\in\mathbb{R}^3 is the predicted cluster center, fi∈RCf_i\in\mathbb{R}^C is its CC-dimensional feature, and i∈{1,…,M}i\in\{1,\ldots,M\} indexes the MM clusters. The method emits rays from viv_i using PP uniformly spaced polar-angle bins. For bin p∈{0,…,P−1}p\in\{0,\ldots,P-1\}, the polar angle and number of rays are

    θp=πpP−1,Ap={1,p=0 or p=P−1,4p,0<p≤(P−1)/2,4(P−p−1),(P−1)/2<p<P−1.\theta_p=\frac{\pi p}{P-1},\qquad A_p=\begin{cases} 1,&p=0\text{ or }p=P-1,\\ 4p,&0<p\leq (P-1)/2,\\ 4(P-p-1),&(P-1)/2<p<P-1. \end{cases}

    The aa-th ray in bin pp, for a∈{0,…,Ap−1}a\in\{0,\ldots,A_p-1\}, has azimuth psi_{p,a}=2\pi a/A_p and polar angle theta_{p,a}=\theta_p. The total number of rays is N=∑p=0P−1ApN=\sum_{p=0}^{P-1}A_p; RBGNet uses P=9P=9, which gives N=66N=66 rays, with denser angular coverage near the equatorial plane. All rays from cluster ii have the same far-bound distance lil_i, predicted from fif_i as the candidate object scale. The scale is explicitly supervised by

    Lscale-reg=1I∑i∥li−li∗∥η 1[i is positive],\mathcal{L}_{\mathrm{scale\text{-}reg}}=\frac{1}{I}\sum_i \left\lVert l_i-l_i^*\right\rVert_{\eta}\,\mathbb{1}[i\text{ is positive}],

    where li∗l_i^* is the half-diagonal length of the assigned ground-truth box, a positive cluster is one whose center lies within 0.30.3 m of a ground-truth object center, II is the number of positive clusters, 1[⋅]\mathbb{1}[\cdot] is an indicator, and ∥⋅∥η\lVert\cdot\rVert_{\eta} denotes the smooth-ℓ1\ell_1 norm.

  3. Knowl 3 — Coarse-to-fine surface anchor generation

    model/method

    RBGNet samples anchor points along every determined ray in two stages. Before ray sampling, the seed-point representation is upsampled to 2,048 points by trilinear interpolation at the target positions used by the first PointNet++ set-abstraction layer. In the coarse stage, each ray is divided into KcK_c equal bins and one anchor point is sampled from each bin by stratified sampling. PointNet++ set abstraction aggregates nearby seed features around every coarse anchor. A binary mask head then predicts whether each coarse anchor belongs to the candidate object, using the anchor’s local feature together with the candidate cluster feature. During training, an anchor is positive when the ball-query region around it contains a surface point from its assigned ground-truth object.

    In the fine stage, RBGNet uses inverse-transform sampling to draw KfK_f anchors from the positive regions predicted for each ray by the coarse mask. This concentrates fine anchors in dense portions of the candidate object rather than allocating them uniformly to free space and background. Fine anchors receive local features and positive masks through the same set-abstraction and mask-prediction process. The procedure produces coarse and fine feature sets, masks, and three-dimensional anchor positions for all NN rays.

  4. Knowl 4 — Ordered ray feature enhancement

    model/method

    RBGNet converts the masked coarse and fine anchor features into an object-shape descriptor while preserving the predefined order of rays and the order of anchors along each ray. For either branch, the feature of every negative anchor is replaced by zero. In the coarse branch, the KcK_c masked features ρ^n,k(c)\hat{\rho}^{(c)}_{n,k} on ray nn are concatenated in depth order and projected to a 32-dimensional ray feature:

    rn(c)=Fpoint(c)({ρ^n,k(c)}k=1Kc,⊙),r_n^{(c)}=\mathcal{F}_{\mathrm{point}}^{(c)}\left(\{\hat{\rho}^{(c)}_{n,k}\}_{k=1}^{K_c},\odot\right),

    where ⊙\odot denotes concatenation. The NN ray features are then concatenated in the determined ray order and processed by a two-layer MLP to obtain a 128-dimensional coarse descriptor:

    μ(c)=Fray(c)({rn(c)}n=1N,⊙).\mu^{(c)}=\mathcal{F}_{\mathrm{ray}}^{(c)}\left(\{r_n^{(c)}\}_{n=1}^{N},\odot\right).

    The fine branch applies the same operations to produce a 128-dimensional descriptor μ(f)\mu^{(f)}. A fusion network combines the two descriptors, g=Ffuse(μ(c),μ(f))g=\mathcal{F}_{\mathrm{fuse}}(\mu^{(c)},\mu^{(f)}), and the fused surface-geometry feature gg is combined with the original cluster feature fif_i for box and semantic prediction. The authors report that changing the predefined ordering strategy does not materially affect performance, provided that the ordering is consistent within the feature aggregation process.

  5. Knowl 5 — Foreground Biased Sampling

    algorithm

    Foreground Biased Sampling replaces task-agnostic farthest-point sampling with separate sampling of high-confidence foreground points and the remaining background points.

    Input: Point set D with 2048 points, point coordinates, point features, target sample counts alpha and beta, and foreground quota kappa
    Output: Final sampled point set S
    For every point d_j in D, compute a foreground score e_j in [0, 1] from its feature and coordinates using a segmentation head
    Sort D by descending foreground score
    Put the top kappa points into foreground set D_f
    Put the remaining points into background set D_b
    Apply farthest-point sampling to D_f to obtain alpha points D_hat_f
    Apply farthest-point sampling to D_b to obtain beta points D_hat_b
    Return S as the concatenation of D_hat_f and D_hat_b

    The foreground segmentation head is supervised by the ground-truth 3D boxes with cross-entropy loss. In the standard setting, 87.5% of the target samples come from the foreground set and 12.5% come from the background set, preserving scene coverage while emphasizing object surfaces. For the 2,048-to-1,024 downsampling step, RBGNet uses κ=1024\kappa=1024, α=896\alpha=896, and β=128\beta=128. At inference time, the foreground score is obtained from the margin between the positive and negative class outputs. The same strategy is applied in the second through fourth PointNet++ set-abstraction layers.

  6. Knowl 6 — Joint training objective

    equation

    RBGNet is trained end-to-end with a weighted sum of six losses:

    L=λvote-regLvote-reg+λfbsLfbs+λrbfgLrbfg+λobj-clsLobj-cls+λboxLbox+λsem-clsLsem-cls.\mathcal{L}=\lambda_{\mathrm{vote\text{-}reg}}\mathcal{L}_{\mathrm{vote\text{-}reg}}+\lambda_{\mathrm{fbs}}\mathcal{L}_{\mathrm{fbs}}+\lambda_{\mathrm{rbfg}}\mathcal{L}_{\mathrm{rbfg}}+\lambda_{\mathrm{obj\text{-}cls}}\mathcal{L}_{\mathrm{obj\text{-}cls}}+\lambda_{\mathrm{box}}\mathcal{L}_{\mathrm{box}}+\lambda_{\mathrm{sem\text{-}cls}}\mathcal{L}_{\mathrm{sem\text{-}cls}}.

    Here the six terms respectively supervise vote-center regression, foreground-biased sampling, ray-based feature grouping, proposal objectness, 3D box estimation, and semantic classification; each λ\lambda is a nonnegative loss-balancing coefficient. The ray-based grouping loss is

    Lrbfg=λscale-regLscale-reg+λc-clsLc-cls+λf-clsLf-cls,\mathcal{L}_{\mathrm{rbfg}}=\lambda_{\mathrm{scale\text{-}reg}}\mathcal{L}_{\mathrm{scale\text{-}reg}}+\lambda_{\mathrm{c\text{-}cls}}\mathcal{L}_{\mathrm{c\text{-}cls}}+\lambda_{\mathrm{f\text{-}cls}}\mathcal{L}_{\mathrm{f\text{-}cls}},

    where Lscale-reg\mathcal{L}_{\mathrm{scale\text{-}reg}} is the smooth-ℓ1\ell_1 scale-regression loss, and Lc-cls\mathcal{L}_{\mathrm{c\text{-}cls}} and Lf-cls\mathcal{L}_{\mathrm{f\text{-}cls}} are cross-entropy losses for valid coarse and fine surface-anchor queries. The foreground-sampling loss Lfbs\mathcal{L}_{\mathrm{fbs}} is also cross entropy; the remaining voting, objectness, box, and semantic losses use the VoteNet label assignments and loss definitions.

  7. Knowl 7 — Training and evaluation protocol

    experimental setup

    RBGNet is evaluated on ScanNet V2 and SUN RGB-D using their standard data splits and the VoteNet evaluation protocol. ScanNet V2 contains 1,513 reconstructed indoor training scenes with axis-aligned boxes for 18 categories. SUN RGB-D contains approximately 5,000 RGB-D training images with oriented 3D boxes for 10 categories; depth images are converted to point clouds using the supplied camera parameters. Mean average precision is reported at 3D intersection-over-union thresholds of 0.25 and 0.50.

    Each training scene is subsampled to 50,000 input points. The network is optimized with AdamW, batch size 8 per GPU, for 360 epochs. The initial learning rate is 0.006 on ScanNet V2 and 0.004 on SUN RGB-D, and each learning rate is reduced by a factor of 10 at epochs 240 and 330. Gradient-norm clipping is used to stabilize training.

  8. Knowl 8 — State-of-the-art detection performance

    empirical result

    RBGNet improves point-only 3D detection on both evaluated indoor-scene benchmarks. On ScanNet V2 with a standard PointNet++ backbone and 66 rays, the 256-candidate configuration obtains 70.2 [email protected] and 54.2 [email protected]; its averages over 25 trials are 69.6 and 53.6. A stronger configuration with twice the channel capacity and 512 candidates obtains 70.6 and 55.2, with 25-trial averages of 69.9 and 54.7. The corresponding Group-free baseline reports 67.2 and 49.7 as its best scores, with averages of 66.6 and 49.0. Thus, the standard RBGNet configuration improves on the same-backbone prior comparison by 2.5 mAP points at IoU 0.25 and 3.3 points at IoU 0.50; the stronger configuration improves over the compared Group-free configuration by 1.5 and 2.4 points.

    On SUN RGB-D, the standard RBGNet configuration obtains 64.1 [email protected] and 47.2 [email protected], with 25-trial averages of 63.6 and 46.3. The corresponding Group-free scores are 63.0 and 45.2, with averages of 62.6 and 44.4. RBGNet uses geometric point-cloud input only and nevertheless outperforms the prior point-only methods reported in the comparison.

  9. Knowl 9 — Ablation evidence for ray grouping and foreground sampling

    empirical result

    On ScanNet V2, the two proposed modules provide complementary gains. A strong baseline without either module reaches 66.2 [email protected] and 48.2 [email protected]. Adding only Foreground Biased Sampling raises the scores to 67.1 and 49.0; adding only ray-based feature grouping raises them to 69.0 and 52.9; adding both reaches 69.6 and 53.6.

    The number of rays controls both surface coverage and accuracy. With 0, 6, 18, 38, 66, and 102 rays, the reported [email protected] values are respectively 67.1, 68.4, 68.7, 69.2, 69.6, and 69.9, while the [email protected] values are 49.0, 51.6, 52.0, 52.7, 53.6, and 53.9. The corresponding reported object-point recalls for 6, 18, 38, 66, and 102 rays are 38.1, 63.3, 75.8, 78.1, and 86.5 percent. RBGNet selects 66 rays as the accuracy-memory compromise.

    With all other settings fixed, replacing the baseline voting grouping with RoI pooling, back-tracing, Group-free grouping, or RBGNet ray grouping gives respectively (67.6,49.9)(67.6,49.9), (67.7,50.1)(67.7,50.1), (68.1,50.5)(68.1,50.5), and (69.6,53.6)(69.6,53.6) for [email protected] and [email protected]; the voting baseline is (67.1,49.0)(67.1,49.0). Foreground Biased Sampling also retains more foreground points through the backbone: at the second, third, and fourth downsampling layers, foreground percentages on ScanNet V2 are (51.2,73.2,87.8)(51.2,73.2,87.8) for RBGNet, compared with (31.1,30.8,30.3)(31.1,30.8,30.3) for standard FPS and (40.4,42.1,43.8)(40.4,42.1,43.8) for fused FPS. On SUN RGB-D, the corresponding percentages are (30.8,45.3,65.1)(30.8,45.3,65.1) for RBGNet, (17.9,17.8,17.7)(17.9,17.8,17.7) for FPS, and (21.3,22.9,23.5)(21.3,22.9,23.5) for fused FPS.

  10. Knowl 10 — Inference-speed trade-off across ray counts

    empirical result

    RBGNet maintains competitive inference speed while improving detection accuracy. On the same workstation—a single NVIDIA Tesla V100 GPU with 256 GB RAM and an Intel Xeon E5-2650 v3—the reported methods achieve the following best [email protected], best [email protected], and frame rate: MLCVNet (64.5,41.4,5.37)(64.5,41.4,5.37), BRNet (66.1,50.9,7.37)(66.1,50.9,7.37), H3DNet (67.2,48.1,3.75)(67.2,48.1,3.75), and Group-free (67.3,48.9,6.64)(67.3,48.9,6.64). RBGNet with 6, 18, 38, and 66 rays obtains respectively (69.0,52.3,7.23)(69.0,52.3,7.23), (69.0,52.6,5.70)(69.0,52.6,5.70), (69.7,53.3,5.27)(69.7,53.3,5.27), and (70.2,54.2,4.75)(70.2,54.2,4.75), where the third value in each tuple is frames per second. Increasing the number of rays improves accuracy but reduces throughput; the selected 66-ray configuration remains faster than H3DNet while delivering higher reported accuracy.

Coverage note — The preliminary ground-truth-feature diagnostic, qualitative box visualizations, and appendix-specific loss-weight details were omitted because they are supportive analyses rather than standalone load-bearing contributions.

References

  1. 1.David Acuna, Huan Ling, Amlan Kar, and Sanja Fidler. Efficient interactive annotation of segmentation datasets with polygon-rnn++. In CVPR, 2018. 3
  2. 2.Ronald T Azuma. A survey of augmented reality. Presence: teleoperators & virtual environments, 1997. 1
  3. 3.Mayank Bansal, Alex Krizhevsky, and Abhijit Ogale. Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst. 2018. 1
  4. 4.Mark Billinghurst, Adrian Clark, and Gun Lee. A survey of augmented reality. 2015. 1
  5. 5.Jintai Chen, Biwen Lei, Qingyu Song, Haochao Ying, Danny Z Chen, and Jian Wu. A hierarchical graph network for 3d object detection on point clouds. In CVPR, 2020. 4, 7
  6. 6.Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In CVPR, 2017. 1, 2
  7. 7.Bowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang, and Dong Xu. Back-tracing representative points for voting-based 3d object detection in point clouds. In CVPR, 2021. 1, 4, 7, 8
  8. 8.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 2, 6, 7
  9. 9.Jiajun Deng, Shaoshuai Shi, Peiwei Li, Wen gang Zhou, Yanyong Zhang, and Houqiang Li. Voxel r-cnn: Towards high performance voxel-based 3d object detection. In AAAI, 2021. 1
  10. 10.Oren Dovrat, Itai Lang, and Shai Avidan. Learning to sample. CoRR, 2018. 3
  11. 11.Francis Engelmann, Martin Bokeloh, Alireza Fathi, Bastian Leibe, and Matthias Nießner. 3d-mpa: Multi-proposal aggregation for 3d semantic instance segmentation. In CVPR, 2020. 7
  12. 12.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In CVPR, 2018. 2
  13. 13.Ji Hou, Angela Dai, and Matthias Nießner. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In CVPR, 2019. 7
  14. 14.Li Jiang, Hengshuang Zhao, Shu Liu, Xiaoyong Shen, Chi-Wing Fu, and Jiaya Jia. Hierarchical point-edge interaction network for point cloud semantic segmentation. In ICCV, 2019. 3
  15. 15.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR, 2020. 2
  16. 16.Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven L Waslander. Joint 3d proposal generation and object detection from view aggregation. In IROS, 2018. 1
  17. 17.Itai Lang, Asaf Manor, and Shai Avidan. Samplenet: Differentiable point cloud sampling. In CVPR, 2020. 3
  18. 18.Justin Liang, Namdar Homayounfar, Wei-Chiu Ma, Yuwen Xiong, Rui Hu, and Raquel Urtasun. Polytransform: Deep polygon transformer for instance segmentation. In CVPR, 2020. 3
  19. 19.Ming Liang, Bin Yang, Yun Chen, Rui Hu, and Raquel Urtasun. Multi-task multi-sensor fusion for 3d object detection. In CVPR, 2019. 1
  20. 20.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. 2021. 1, 3, 7, 8
  21. 21.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 4, 5
  22. 22.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In CVPR, 2021. 7
  23. 23.Ehsan Nezhadarya, Ehsan Taghavi, Ryan Razani, Bingbing Liu, and Jun Luo. Adaptive hierarchical down-sampling for point cloud classification. In CVPR, 2020. 3
  24. 24.Jinhyung Park, Xinshuo Weng, Yunze Man, and Kris Kitani. Multi-modality task cascade for 3d object detection. arXiv preprint arXiv:2107.04013, 2021. 7
  25. 25.Sida Peng, Wen Jiang, Huaijin Pi, Xiuli Li, Hujun Bao, and Xiaowei Zhou. Deep snake for real-time instance segmentation. In CVPR, 2020. 3
  26. 26.Hughes Perreault, Guillaume-Alexandre Bilodeau, Nicolas Saunier, and Maguelonne Heritier. Centerpoly: real-time instance segmentation using bounding polygons. In ICCV, 2021. 3
  27. 27.Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas. Imvotenet: Boosting 3d object detection in point clouds with image votes. In CVPR, 2020. 7
  28. 28.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, 2019. 1, 2, 3, 4, 6, 7, 8
  29. 29.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, 2018. 1, 7
  30. 30.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017. 1, 2
  31. 31.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017. 1, 3, 4, 5, 6
  32. 32.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In CVPR, 2020. 1, 2
  33. 33.Shaoshuai Shi, Li Jiang, Jiajun Deng, Zhe Wang, Chaoxu Guo, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn++: Point-voxel feature set abstraction with local vector representation for 3d object detection. arXiv preprint arXiv:2102.00463, 2021. 2
  34. 34.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointrcnn: 3d object proposal generation and detection from point cloud. In CVPR, 2019. 1, 3, 7
  35. 35.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. TPAMI, 2020. 1, 2
  36. 36.Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015. 2, 6, 7
  37. 37.Shuran Song and Jianxiong Xiao. Deep sliding shapes for amodal 3d object detection in rgb-d images. In CVPR, 2016. 1
  38. 38.Dequan Wang, Coline Devin, Qi-Zhi Cai, Philipp Krahenbühl, and Trevor Darrell. Monocular plan view networks for autonomous driving. In IROS, 2019. 1
  39. 39.Haiyang Wang, Wenguan Wang, Xizhou Zhu, Jifeng Dai, and Liwei Wang. Collaborative visual navigation. arXiv preprint arXiv:2107.01151, 2021. 1
  40. 40.Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu, Ding Liang, Chunhua Shen, and Ping Luo. Polarmask: Single shot instance segmentation with polar representation. In CVPR, 2020. 3
  41. 41.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Dening Lu, Mingqiang Wei, and Jun Wang. Venet: Voting enhancement network for 3d object detection. In ICCV, 2021. 7
  42. 42.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Yiming Zhang, Kai Xu, and Jun Wang. Mlcvnet: Multi-level context votenet for 3d object detection. In CVPR, 2020. 3, 4, 7, 8
  43. 43.Wenqiang Xu, Haiyang Wang, Fubo Qi, and Cewu Lu. Explicit shape encoding for real-time instance segmentation. In ICCV, 2019. 3
  44. 44.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 2018. 1, 2
  45. 45.Bin Yang, Ming Liang, and Raquel Urtasun. Hdnet: Exploiting hd maps for 3d object detection. In CoRL, 2018. 1
  46. 46.Bin Yang, Wenjie Luo, and Raquel Urtasun. Pixor: Real-time 3d object detection from point clouds. In CVPR, 2018. 1, 2
  47. 47.Zetong Yang, Yanan Sun, Shu Liu, and Jiaya Jia. 3dssd: Point-based 3d single stage object detector. In CVPR, 2020. 2, 3, 7
  48. 48.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3d object detection and tracking. In CVPR, 2021. 1, 2
  49. 49.Zaiwei Zhang, Bo Sun, Haitao Yang, and Qixing Huang. H3dnet: 3d object detection using hybrid geometric primitives. In ECCV, 2020. 1, 3, 7, 8
  50. 50.Hengshuang Zhao, Li Jiang, Chi-Wing Fu, and Jiaya Jia. Pointweb: Enhancing local neighborhood features for point cloud processing. In CVPR, 2019. 3
  51. 51.Wu Zheng, Weiliang Tang, Sijin Chen, Li Jiang, and Chi-Wing Fu. Cia-ssd: Confident iou-aware single-stage object detector from point cloud. AAAI, 2021. 2
  52. 52.Wu Zheng, Weiliang Tang, Li Jiang, and Chi-Wing Fu. Se-ssd: Self-ensembling single-stage object detector from point cloud. In CVPR, 2021. 2
  53. 53.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In CVPR, 2018. 1, 2
  54. 54.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In ICRA, 2017. 1

Citation

MLA
Wang, H., et al. “RBGNet: Ray-based Grouping for 3D Object Detection”. arXiv, 2022, http://arxiv.org/abs/2204.02251v1.
APA
Wang, H., Shi, S., Yang, Z., Fang, R., Qian, Q., Li, H., Schiele, B., & Wang, L. (2022). RBGNet: Ray-based Grouping for 3D Object Detection. arXiv. http://arxiv.org/abs/2204.02251v1
Chicago
Wang, H., S. Shi, Z. Yang, et al. 2022. “RBGNet: Ray-based Grouping for 3D Object Detection”. arXiv. http://arxiv.org/abs/2204.02251v1.
Harvard
Wang, H. et al. (2022) “RBGNet: Ray-based Grouping for 3D Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.02251v1.
Vancouver
1. Wang H, Shi S, Yang Z, Fang R, Qian Q, Li H, Schiele B, Wang L (2022) RBGNet: Ray-based Grouping for 3D Object Detection. arXiv

BibTeX

@article{wang2022rbgnet,
  title = {RBGNet: Ray-based Grouping for 3D Object Detection},
  author = {Wang, Haiyang and Shi, Shaoshuai and Yang, Ze and Fang, Rongyao and Qian, Qi and Li, Hongsheng and Schiele, Bernt and Wang, Liwei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.02251v1},
  eprint = {2204.02251}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE