CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds

Haiyang WangLihe DingShaocong DongShaoshuai ShiAoxue LiJianan LiZhenguo LiLiwei Wang

article2022NeurIPS110 citations

Introduces a two-stage fully sparse convolutional framework for 3D indoor object detection that overcomes bottom-up grouping errors by enforcing semantic consistency during proposal generation and using an efficient sparse RoI pooling module to recover missed geometric features directly from the backbone.

Listen

Detecting 3D objects from raw, irregular 3D point cloud data is critical for emerging technologies such as autonomous driving, robotics, and augmented reality. In complex indoor environments, existing detection frameworks typically group nearby points together without considering their object category. This class-agnostic grouping often leads to errors in cluttered spaces, such as merging parts of adjacent but unrelated objects or using search areas that fail to capture the boundaries of large objects while adding noise to small ones. Furthermore, traditional secondary refinement modules rely on complex, memory-heavy pooling operations that degrade fine-grained structural information.

The article evaluates CAGroup3D, a fully convolutional two-stage 3D object detection framework designed to overcome these grouping and refinement weaknesses. The authors demonstrate how pairing category-aware point clustering with a lightweight, convolution-based refinement step substantially boosts 3D detection precision.

To demonstrate this capability, the authors conducted rigorous empirical experiments on two standard indoor 3D benchmarks: ScanNet V2, which contains rich 3D mesh scans across 18 object categories, and SUN RGB-D, which includes over 10,000 single-view RGB-D images across 10 categories. The framework first uses a dual-resolution 3D sparse convolutional network to extract detailed spatial features. It then applies a class-aware grouping strategy that shifts points toward estimated object centers and groups only points sharing the same predicted category within search regions scaled to average category dimensions. Finally, a fully sparse convolutional region pooling module directly samples and refines object bounding boxes from the extracted feature maps without relying on traditional max-pooling operations.

The evaluation produced four key findings in order of importance. First, CAGroup3D established a new state of the art in 3D object detection, achieving 75.1% mean average precision at an intersection-over-union threshold of 0.25 on ScanNet V2—a 3.6 percentage point gain over the prior leading method—and 66.8% on SUN RGB-D, outperforming both single-modality and multi-modal camera-plus-depth methods. Second, adding class-aware local grouping accounted for the largest performance leap, raising precision from 68.22% to 72.10% over the baseline. Third, the proposed region pooling module improved bounding box accuracy during the refinement stage, increasing precision at the stricter 0.50 threshold from 57.18% to 60.31%. Fourth, the new pooling mechanism reduced GPU memory consumption to 2,468 megabytes, requiring less than one-third of the memory consumed by common alternative pooling approaches.

These results demonstrate that incorporating category awareness directly into point grouping and utilizing memory-efficient sparse convolutions can substantially improve both detection accuracy and hardware efficiency. For operational systems in robotics and smart spaces, these improvements mean safer, more reliable spatial awareness and lower computing resource demands. Organizations building spatial perception pipelines should adopt category-adaptive grouping and sparse convolutional pooling architectures over older, class-agnostic voting pipelines.

The findings are supported with high confidence through multiple independent experimental trials. However, the current framework is primarily designed to address variations across different categories and does not account for size and shape variations among objects within the same category. Future development should focus on addressing this within-class variability to further improve real-world robustness.

arXiv: 2210.04264
Cover for CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds

Abstract

We present a novel two-stage fully sparse convolutional 3D object detection framework, named CAGroup3D. Our proposed method first generates some high-quality 3D proposals by leveraging the class-aware local group strategy on the object surface voxels with the same semantic predictions, which considers semantic consistency and diverse locality abandoned in previous bottom-up approaches. Then, to recover the features of missed voxels due to incorrect voxel-wise segmentation, we build a fully sparse convolutional RoI pooling module to directly aggregate fine-grained spatial information from backbone for further proposal refinement. It is memory-and-computation efficient and can better encode the geometry-specific features of each 3D proposal. Our model achieves state-of-the-art 3D detection performance with remarkable gains of +3.6% on ScanNet V2 and +2.6% on SUN RGB-D in term of mAP@0.25. Code will be available at https://github.com/Haiyang-W/CAGroup3D.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 3D Voxel CNN for Point Cloud Feature Learning
  • 3.2 Class-Aware 3D Proposal Generation
  • 3.3 RoI-Conv point cloud feature pooling for 3D Proposal Refinement
  • 3.4 Learning Objective
  • 4 Experiments
  • 4.1 Datasets and Evaluation Metric
  • 4.2 Implementation Details
  • 4.3 Benchmarking Results
  • 4.4 Ablation Studies and Discussions
  • 5 Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — CAGroup3D two-stage sparse detection framework

    model/method

    CAGroup3D is a two-stage, fully sparse-convolutional detector for oriented 3D object boxes from unordered point clouds. A sparse 3D voxel backbone first produces voxel-wise geometric features. Stage I predicts semantic scores and center votes for non-empty surface voxels, forms class-specific vote spaces, and applies class-aware local grouping to generate proposal boxes. Stage II revisits the backbone features inside each proposal with a hierarchical RoI-Conv pooling module, recovering features from surface voxels missed by Stage-I semantic prediction and refining each proposal’s dimensions, center, and orientation. The architecture avoids hand-crafted point set abstraction and uses sparse convolutions throughout the feature extraction, proposal generation, and proposal refinement pipeline.

  2. Knowl 2 — Bilateral sparse voxel backbone

    model/method

    The feature extractor is BiResNet, a dual-resolution sparse 3D convolutional network based on ResNet-18. Its downsampling branch replaces all ordinary convolutions with sparse 3D convolutions to obtain large receptive fields and multiscale context. An auxiliary branch maintains a higher-resolution sparse feature map at one-half the input voxel resolution, performs no further downsampling, and exchanges information with the downsampling branch through bridge operations. Bilateral fusion therefore combines contextual features from the low-resolution branch with fine-grained geometric features from the high-resolution branch, which are used by both voxel-wise prediction and proposal generation.

  3. Knowl 3 — Voxel-wise semantic and voting predictions

    model/method

    Let the sparse backbone produce NN non-empty voxels. For voxel ii, its spatial coordinate is xi∈R3x_i\in\mathbb{R}^{3}, its feature is fi∈RCf_i\in\mathbb{R}^{C}, and oi=[xi;fi]o_i=[x_i;f_i] denotes their concatenation. A voting branch predicts a spatial center offset Δxi∈R3\Delta x_i\in\mathbb{R}^{3} and feature offset Δfi∈RC\Delta f_i\in\mathbb{R}^{C}, producing the voted voxel

    pi=[xi+Δxi ; fi+Δfi],i=1,…,N. p_i=[x_i+\Delta x_i\,;\,f_i+\Delta f_i],\qquad i=1,\ldots,N.

    The spatial offset is trained toward the displacement from the seed voxel to the center of its associated ground-truth 3D box. In parallel, a one-layer MLP semantic branch predicts a score vector over NclassN_{\mathrm{class}} object categories:

    si=MLP⁡sem(oi)∈[0,1]Nclass. s_i=\operatorname{MLP}_{\mathrm{sem}}(o_i)\in[0,1]^{N_{\mathrm{class}}}.

    The semantic scores are trained with focal loss, while the spatial votes use smooth-ℓ1\ell_1 regression. Both targets are obtained from bounding boxes rather than instance or semantic masks. If a voxel lies inside multiple ground-truth boxes, the box with the smallest volume supplies its target.

  4. Knowl 4 — Class-aware local grouping and adaptive re-voxelization

    model/method

    For each class jj, CAGroup3D retains every voted voxel whose predicted class-jj score exceeds a threshold τ\tau, rather than assigning each voxel to only its highest-scoring class. The class-specific vote set is

    Cj={pi:si(j)>τ, i=1,…,N},j=1,…,Nclass,\mathcal{C}_j=\{p_i:s_i^{(j)}>\tau,\ i=1,\ldots,N\},\qquad j=1,\ldots,N_{\mathrm{class}},

    where si(j)s_i^{(j)} is the class-jj component of the semantic score vector. Each Cj\mathcal{C}_j is independently voxelized with average feature pooling. If dj=(wj,hj,lj)d_j=(w_j,h_j,l_j) is the average spatial dimension of class jj in the training data, the voxel size for that class is αdj\alpha d_j, where α\alpha is a scalar scale factor. This produces a class-specific non-empty voxel set Vj\mathcal{V}_j.

    A sparse 3D convolution with kernel size k(a)k^{(a)} is centered at every voxel in Vj\mathcal{V}_j and uses only voxels from the same class-specific vote space:

    ai(j)=SparseConv⁡3D(j)(vi,Vj,k(a)),vi∈Vj. a_i^{(j)}=\operatorname{SparseConv}^{(j)}_{3D}\left(v_i,\mathcal{V}_j,k^{(a)}\right),\qquad v_i\in\mathcal{V}_j.

    A shared anchor-free prediction head then estimates class probabilities, box regression parameters, and confidence scores from the aggregated features. The grouping operation is class-dependent while retaining a common kernel size; adaptive voxel sizes give different categories different effective local regions, improving semantic consistency and accommodating category-level size variation.

  5. Knowl 5 — RoI-Conv hierarchical proposal refinement

    algorithm

    RoI-Conv pooling extracts proposal-specific features directly from sparse backbone voxels without ball queries, vector queries, or max pooling. Its input is a set of sparse voxels with features and a set of proposal boxes. For each proposal, uniformly sample Gx×Gy×GzG_x\times G_y\times G_z grid points in voxel coordinates. Merge coincident grid points from overlapping proposals into a unique set. For each unique grid point gkg_k, collect the neighboring input voxels within sparse-convolution kernel size k(p)k^{(p)}. If the neighborhood is non-empty, apply one shared sparse 3D convolution centered at gkg_k; otherwise discard the grid point. The remaining output voxels retain their spatial locations and encode local geometry.

    Input: sparse voxel set I, proposal boxes M, sampling resolution (Gx, Gy, Gz), kernel size kp
    Output: proposal-specific sparse feature set Q
    For each proposal in M:
        Uniformly sample Gx × Gy × Gz grid points inside the proposal
    Merge duplicate grid points from overlapping proposals
    For each unique grid point gk:
        Collect input voxels Nk within kernel size kp around gk
        If Nk is non-empty:
            Compute qk with the shared sparse convolution centered at gk
            Add qk to Q
        Otherwise:
            Discard gk
    Return Q

    CAGroup3D stacks two sparse abstraction blocks. The first samples a 7×7×77\times7\times7 grid and uses a 55-voxel sparse-convolution kernel; the second samples a 1×1×11\times1\times1 grid and uses a 77-voxel kernel. Before the final block, voxels associated with each oriented proposal are transformed into that proposal’s canonical coordinate system. The resulting feature for each proposal predicts residual corrections to its dimensions, center, and orientation.

  6. Knowl 6 — Joint training objective

    equation

    CAGroup3D is trained from scratch with a weighted sum of losses for voxel semantics, voting, proposal centerness, Stage-I box estimation, Stage-I classification, and Stage-II box refinement:

    L=βsemLsem+βvoteLvote+βcntrLcntr+βboxLbox+βclsLcls+βreboxLrebox,\mathcal{L}=\beta_{\mathrm{sem}}\mathcal{L}_{\mathrm{sem}}+\beta_{\mathrm{vote}}\mathcal{L}_{\mathrm{vote}}+\beta_{\mathrm{cntr}}\mathcal{L}_{\mathrm{cntr}}+\beta_{\mathrm{box}}\mathcal{L}_{\mathrm{box}}+\beta_{\mathrm{cls}}\mathcal{L}_{\mathrm{cls}}+\beta_{\mathrm{rebox}}\mathcal{L}_{\mathrm{rebox}},

    where each β\beta is a loss-balancing coefficient. Lsem\mathcal{L}_{\mathrm{sem}} is focal loss for voxel-wise semantic scores, Lvote\mathcal{L}_{\mathrm{vote}} is smooth-ℓ1\ell_1 loss for predicted center offsets, and the Stage-I centerness, box, and classification losses supervise proposal generation. Stage-II refinement uses residual regression and an IoU loss:

    Lrebox=∑r∈{x,y,z,l,h,w,θ}Lsmooth-ℓ1(Δr∗,Δr)+Liou.\mathcal{L}_{\mathrm{rebox}}=\sum_{r\in\{x,y,z,l,h,w,\theta\}}\mathcal{L}_{\mathrm{smooth}\text{-}\ell_1}(\Delta r^{*},\Delta r)+\mathcal{L}_{\mathrm{iou}}.

    Here x,y,zx,y,z are box-center coordinates, l,h,wl,h,w are box dimensions, θ\theta is orientation, Δr\Delta r is the predicted residual relative to a Stage-I proposal, and Δr∗\Delta r^{*} is the corresponding ground-truth residual.

  7. Knowl 7 — Datasets, training, and evaluation protocol

    experimental setup

    The detector is evaluated on ScanNet V2 and SUN RGB-D using their standard splits and mean average precision at 3D IoU thresholds 0.250.25 and 0.500.50. ScanNet V2 has 1,201 training scans, 312 validation scans, and 18 annotated object categories with axis-aligned boxes. SUN RGB-D has 10,355 RGB-D images, approximately 5,000 training images, and 10 categories with oriented boxes; depth images are converted to point clouds using the provided camera parameters.

    For both datasets, the input voxel size is 0.02 m0.02\,\mathrm{m} and the BiResNet high-resolution branch uses 0.04 m0.04\,\mathrm{m}. The class-aware grouping scale is α=0.15\alpha=0.15, the grouping kernel size is k(a)=9k^{(a)}=9, and the RoI-Conv blocks use sampling resolutions 7×7×77\times7\times7 and 1×1×11\times1\times1 with sparse kernels of sizes 55 and 77. The semantic threshold starts at τ=0.15\tau=0.15 and decreases by 0.020.02 every 10 epochs on ScanNet V2 or every 4 epochs on SUN RGB-D until reaching 0.050.05.

    Training uses AdamW with batch size 1616, initial learning rate 0.0010.001, and weight decay 0.00010.0001. ScanNet V2 is trained for 120 epochs with tenfold learning-rate reductions at epochs 80 and 110; SUN RGB-D is trained for 48 epochs with reductions at epochs 32 and 44. Training uses two NVIDIA Tesla V100 GPUs with 32 GB per card and gradient-norm clipping. The reported best and average results are obtained from five training runs with five evaluations per trained model, giving 25 trials.

  8. Knowl 8 — Benchmark performance on indoor 3D detection

    empirical result

    CAGroup3D achieves the strongest reported performance in the paper’s comparisons on both indoor benchmarks. On ScanNet V2, it obtains [email protected] of 75.175.1 and [email protected] of 61.361.3; the averages over 25 trials are 74.574.5 and 60.360.3, respectively. The corresponding FCAF3D results are 71.571.5 and 57.357.3 for the best runs, with averages of 70.770.7 and 56.056.0. Thus, CAGroup3D improves the best results by 3.63.6 and 4.04.0 mAP points at the two IoU thresholds.

    On SUN RGB-D, CAGroup3D obtains [email protected] of 66.866.8 and [email protected] of 50.250.2, with 25-trial averages of 66.466.4 and 49.549.5. FCAF3D obtains 64.264.2 and 48.948.9 for the best runs, with averages of 63.863.8 and 48.248.2. The gains are therefore 2.62.6 and 1.31.3 mAP points. The paper reports that CAGroup3D also exceeds the listed multi-sensor methods on these point-cloud benchmarks, despite SUN RGB-D having relatively poor point-cloud quality.

  9. Knowl 9 — Component, hyperparameter, and memory ablations

    empirical result

    On the ScanNet V2 validation set, averaged over 25 trials, a fully sparse VoteNet-style baseline scores 68.2268.22 [email protected] and 53.1753.17 [email protected]. Adding semantic prediction raises these values to 69.2469.24 and 54.0554.05. Adding class-specific diverse local grouping raises them further to 72.1072.10 and 57.0757.07. Replacing the FPN-style backbone with BiResNet gives 73.2173.21 and 57.1857.18, while adding RoI-Conv refinement produces the full-model result of 74.5074.50 and 60.3160.31.

    The class-aware grouping is sensitive but not narrowly tuned to its re-voxelization scale: with k(a)=9k^{(a)}=9, α=1.00\alpha=1.00 gives 36.38/24.6636.38/24.66, α=0.20\alpha=0.20 gives 74.21/58.7774.21/58.77, the selected α=0.15\alpha=0.15 gives 74.50/60.3174.50/60.31, and α=0.05\alpha=0.05 gives 72.98/58.1072.98/58.10 for [email protected]/[email protected]. For the semantic threshold, τ=0.06\tau=0.06 gives 74.51/60.2774.51/60.27, whereas an overly permissive τ=0.02\tau=0.02 drops performance to 73.53/58.8573.53/58.85.

    Replacing RoI-Conv with other pooling designs under otherwise unchanged settings yields the following ScanNet V2 results and peak memory usage: PointRCNN pooling, 73.65/57.8373.65/57.83 with 8,054 MB8{,}054\,\mathrm{MB}; Part-A2 pooling, 74.01/58.8974.01/58.89 with 6,540 MB6{,}540\,\mathrm{MB}; the authors’ set-abstraction implementation, 73.89/58.1473.89/58.14 with 11,508 MB11{,}508\,\mathrm{MB}; and sparse-convolution RoI-Conv, 74.50/60.3174.50/60.31 with 2,468 MB2{,}468\,\mathrm{MB}. These results support both the accuracy benefit of the two-stage refinement and the claimed memory efficiency of sparse-convolution pooling.

  10. Knowl 10 — Limitation of class-aware locality modeling

    limitation

    CAGroup3D explicitly models inter-category locality by assigning different grouping regions to different semantic classes, but it does not explicitly model intra-category locality. Objects belonging to the same category can still have different spatial dimensions because of incomplete point clouds and scale variation within a category. The learnable sparse convolution may handle some of this variation implicitly, but the paper identifies explicit intra-category adaptive grouping as an unresolved limitation.

Coverage note — No substantial contributed material was deliberately omitted; background and related-work discussion were excluded as non-contributory.

References

  1. 1.Ronald T Azuma. A survey of augmented reality. Presence: teleoperators & virtual environments, 1997.
  2. 2.Mayank Bansal, Alex Krizhevsky, and Abhijit Ogale. Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst. 2018.
  3. 3.Mark Billinghurst, Adrian Clark, and Gun Lee. A survey of augmented reality. 2015.
  4. 4.Jintai Chen, Biwen Lei, Qingyu Song, Haochao Ying, Danny Z Chen, and Jian Wu. A hierarchical graph network for 3d object detection on point clouds. In CVPR, 2020.
  5. 5.Bowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang, and Dong Xu. Back-tracing representative points for voting-based 3d object detection in point clouds. In CVPR, 2021.
  6. 6.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017.
  7. 7.Anton Konushin Danila Rukhovich, Anna Vorontsova. Fcaf3d: Fully convolutional anchor-free 3d object detection. arXiv preprint arXiv:2112.00322, 2021.
  8. 8.Jiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou, Yanyong Zhang, and Houqiang Li. Voxel r-cnn: Towards high performance voxel-based 3d object detection. In AAAI, 2021.
  9. 9.Shaocong Dong, Lihe Ding, Haiyang Wang, Tingfa Xu, Xinli Xu, Jie Wang, Ziyang Bian, Ying Wang, and Jianan Li. MsSVT: Mixed-scale sparse voxel transformer for 3d object detection on point clouds. In NeurIPS, 2022.
  10. 10.Francis Engelmann, Martin Bokeloh, Alireza Fathi, Bastian Leibe, and Matthias Nießner. 3d-mpa: Multi-proposal aggregation for 3d semantic instance segmentation. In CVPR, 2020.
  11. 11.Lue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang, Hang Zhao, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang. Embracing single stride 3d object detector with sparse transformer. In CVPR, 2022.
  12. 12.Benjamin Graham and Laurens van der Maaten. Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307, 2017.
  13. 13.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In CVPR, 2018.
  14. 14.JunYoung Gwak, Christopher Choy, and Silvio Savarese. Generative sparse detection networks for 3d single-shot object detection. In ECCV, 2020.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  16. 16.Yuanduo Hong, Huihui Pan, Weichao Sun, and Yisong Jia. Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes. arXiv preprint arXiv:2101.06085, 2021.
  17. 17.Ji Hou, Angela Dai, and Matthias Nießner. 3d-sis: 3d semantic instance segmentation of rgb-d scans. In CVPR, 2019.
  18. 18.Jikai Jin, Bohang Zhang, Haiyang Wang, and Liwei Wang. Non-convex distributionally robust optimization: Non-asymptotic analysis. In NeurIPS, 2021.
  19. 19.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. Pointpillars: Fast encoders for object detection from point clouds. In CVPR, 2019.
  20. 20.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  21. 21.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In ICCV, 2017.
  22. 22.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In ICCV, 2021.
  23. 23.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
  24. 24.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In CVPR, 2021.
  25. 25.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017.
  26. 26.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, 2018.
  27. 27.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, 2019.
  28. 28.Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas. Imvotenet: Boosting 3d object detection in point clouds with image votes. In CVPR, 2020.
  29. 29.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017.
  30. 30.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointrcnn: 3d object proposal generation and detection from point cloud. In CVPR, 2019.
  31. 31.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In CVPR, 2020.
  32. 32.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. TPAMI, 2020.
  33. 33.Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015.
  34. 34.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In CVPR, 2019.
  35. 35.Thang Vu, Kookhoi Kim, Tung M Luu, Thanh Nguyen, and Chang D Yoo. Softgroup for 3d instance segmentation on point clouds. In CVPR, 2022.
  36. 36.Dequan Wang, Coline Devin, Qi-Zhi Cai, Philipp Krähenbühl, and Trevor Darrell. Monocular plan view networks for autonomous driving. In IROS, 2019.
  37. 37.Haiyang Wang, Wenguan Wang, Xizhou Zhu, Jifeng Dai, and Liwei Wang. Collaborative visual navigation. arXiv preprint arXiv:2107.01151, 2021.
  38. 38.Haiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang, Qi Qian, Hongsheng Li, Bernt Schiele, and Liwei Wang. Rbgnet: Ray-based grouping for 3d object detection. In CVPR, 2022.
  39. 39.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. TPAMI, 2020.
  40. 40.Jun Wang, Shiyi Lan, Mingfei Gao, and Larry S Davis. Infofocus: 3d object detection for autonomous driving with dynamic information modeling. In ECCV, 2020.
  41. 41.Yikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang, Fuchun Sun, and Yunhe Wang. Multimodal token fusion for vision transformers. In CVPR, 2022.
  42. 42.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Yiming Zhang, Kai Xu, and Jun Wang. Mlcvnet: Multi-level context votenet for 3d object detection. In CVPR, 2020.
  43. 43.Qian Xie, Yu-Kun Lai, Jing Wu, Zhoutao Wang, Dening Lu, Mingqiang Wei, and Jun Wang. Venet: Voting enhancement network for 3d object detection. In ICCV, 2021.
  44. 44.Xinli Xu, Shaocong Dong, Lihe Ding, Jie Wang, Tingfa Xu, and Jianan Li. Fusionrcnn: Lidar-camera fusion for two-stage 3d object detection. arXiv preprint arXiv:2209.10733, 2022.
  45. 45.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 2018.
  46. 46.Bin Yang, Wenjie Luo, and Raquel Urtasun. Pixor: Real-time 3d object detection from point clouds. In CVPR, 2018.
  47. 47.Hao Yang, Chen Shi, Yihong Chen, and Liwei Wang. Boosting 3d object detection via object-focused image fusion. arXiv preprint arXiv:2207.10589, 2022.
  48. 48.Zetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen, and Jiaya Jia. Std: Sparse-to-dense 3d object detector for point cloud. In ICCV, 2019.
  49. 49.Li Yi, Wang Zhao, He Wang, Minhyuk Sung, and Leonidas Guibas. Gspn: Generative shape proposal network for 3d instance segmentation in point cloud. In CVPR, 2019.
  50. 50.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3d object detection and tracking. In CVPR, 2021.
  51. 51.Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang. Improved analysis of clipping algorithms for non-convex optimization. In NeurIPS, 2020.
  52. 52.Zaiwei Zhang, Bo Sun, Haitao Yang, and Qixing Huang. H3dnet: 3d object detection using hybrid geometric primitives. In ECCV, 2020.
  53. 53.Yu Zheng, Yueqi Duan, Jiwen Lu, Jie Zhou, and Qi Tian. Hyperdet3d: Learning a scene-conditioned 3d object detector. In CVPR, 2022.
  54. 54.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In CVPR, 2018.
  55. 55.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In ICRA, 2017.

Citation

MLA
Wang, H., et al. “CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 29975–88, https://proceedings.neurips.cc/paper_files/paper/2022/file/c1aaf7c3f306fe94f77236dc0756d771-Paper-Conference.pdf.
APA
Wang, H., Ding, L., Dong, S., Shi, S., Li, A., Li, J., Li, Z., & Wang, L. (2022). CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds. Advances in Neural Information Processing Systems, 35, 29975–29988. https://proceedings.neurips.cc/paper_files/paper/2022/file/c1aaf7c3f306fe94f77236dc0756d771-Paper-Conference.pdf
Chicago
Wang, H., L. Ding, S. Dong, et al. 2022. “CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds”. Advances in Neural Information Processing Systems 35: 29975–88. https://proceedings.neurips.cc/paper_files/paper/2022/file/c1aaf7c3f306fe94f77236dc0756d771-Paper-Conference.pdf.
Harvard
Wang, H. et al. (2022) “CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 29975–29988. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/c1aaf7c3f306fe94f77236dc0756d771-Paper-Conference.pdf.
Vancouver
1. Wang H, Ding L, Dong S, Shi S, Li A, Li J, Li Z, Wang L (2022) CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 29975–29988

BibTeX

@inproceedings{wang2022cagroup3d,
  title = {CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds},
  author = {Wang, Haiyang and Ding, Lihe and Dong, Shaocong and Shi, Shaoshuai and Li, Aoxue and Li, Jianan and Li, Zhenguo and Wang, Liwei},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {29975-29988},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/c1aaf7c3f306fe94f77236dc0756d771-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors