SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds

Qingdong HeZhengning WangHao ZengYi ZengYijun Liu

article2022AAAI131 citations

Proposes an end-to-end 3D object detection framework that couples spherical voxel complete graphs with global nearest-neighbor graph attention to effectively capture both local point interactions and cross-voxel context directly from raw LiDAR data.

Listen

Accurate 3D object detection from light detection and ranging (LiDAR) point cloud sensors is vital for autonomous driving and augmented reality systems. Traditional processing pipelines often convert irregular 3D spatial points into standard 2D views or structured voxel grids, which frequently discards critical fine-grained geometric structure and incurs massive computational overhead. Alternative direct-point methods also struggle to effectively relate neighboring points across local and global contexts.

The article demonstrates and evaluates Sparse Voxel-Graph Attention Network (SVGA-Net), an end-to-end deep learning framework designed to detect 3D objects and estimate precise 3D bounding boxes directly from raw, unprojected LiDAR point clouds using graph representations.

To achieve this, the system divides point clouds into fixed-radius spherical 3D voxels rather than standard rectangular grids. It builds a local complete graph within each spherical voxel to model fine-grained spatial relationships, alongside a global k-nearest neighbors graph connecting voxel centers to guide features with broad spatial context. A multi-scale sparse-to-dense regression module then aggregates high-level and low-level feature maps via upsampling, convolution, and element-wise addition to predict final object classes and 3D bounding boxes. The framework was evaluated on two standardized self-driving benchmarks: the KITTI dataset and the large-scale Waymo Open Dataset.

The evaluation produced several key findings. First, on the Waymo Open Dataset (Level 1 vehicle detection), SVGA-Net achieved 73.45% 3D mean Average Precision (mAP) and 83.52% Bird's Eye View mAP, outperforming established architectures such as PV-RCNN across both Level 1 and Level 2 evaluations. Second, on the KITTI benchmark, SVGA-Net outperformed leading multi-modal methods (which combine both camera images and LiDAR) on moderate and hard car detection by up to 7.50%, achieving 80.47% average precision on moderate difficulty cars. Third, ablation testing confirmed that omitting the global attention mechanism or the multi-scale sparse-to-dense regression connections caused sharp declines in detection accuracy. Finally, the framework demonstrated an average inference execution time of 62 milliseconds per sample, with feature aggregation accounting for roughly 66% of the processing duration.

These findings indicate that graph-based representations can effectively extract high-precision spatial and geometric insights directly from sparse LiDAR data without requiring expensive multi-sensor image fusion pipelines. This offers a path to lower hardware integration costs and streamlined onboard computation while preserving detection robustness under severe occlusion and poor lighting.

For engineering and development roadmaps, teams should consider adopting spherical graph-based feature aggregation in autonomous perception stacks, especially where direct LiDAR processing is prioritized. As next steps, the article recommends exploring the integration of camera RGB image features into the SVGA-Net architecture to further push bounding box precision.

Decision-makers should note that the system experiences performance limitations when detecting objects subject to extreme occlusion exceeding 80%, as insufficient point density inhibits the proper construction of local graphs. Within normal operational environments, however, the reported benchmarks demonstrate high confidence and competitive generalization across diverse driving datasets.

Cover for SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds

Abstract

Accurate 3D object detection from point clouds has become a crucial component in autonomous driving. However, the volumetric representations and the projection methods in previous works fail to establish the relationships between the local point sets. In this paper, we propose Sparse Voxel-Graph Attention Network (SVGA-Net), a novel end-to-end trainable network which mainly contains voxel-graph module and sparse-to-dense regression module to achieve comparable 3D detection tasks from raw LIDAR data. Specifically, SVGA-Net constructs the local complete graph within each divided 3D spherical voxel and global KNN graph through all voxels. The local and global graphs serve as the attention mechanism to enhance the extracted features. In addition, the novel sparse-to-dense regression module enhances the 3D box estimation accuracy through feature maps aggregation at different levels. Experiments on KITTI detection benchmark and Waymo Open dataset demonstrate the efficiency of extending the graph representation to 3D object detection and the proposed SVGA-Net can achieve decent detection accuracy.

Table of Contents

  • Introduction
  • Related Work
  • Proposed Method
  • Voxel-graph Network Architecture
  • Sparse-to-dense Regression
  • Loss Function
  • Experiments
  • Training
  • Comparing with State-of-the-art Methods
  • Qualitative Results
  • Ablation Studies
  • Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Spherical Voxel Grouping and Local Feature Extraction in SVGA-Net

    model/method

    Given an input raw LiDAR point cloud G={V,D}G = \{V, D\} with V={p1,p2,…,pn}V = \{p_1, p_2, \dots, p_n\} representing nn points in D=4D=4 dimensions (pi=[xi,yi,zi,si]Tp_i = [x_i, y_i, z_i, s_i]^T, where xi,yi,zix_i, y_i, z_i are Cartesian coordinates along the X, Y, Z axes and sis_i is laser reflection intensity):

    1. Iterative Farthest Point Sampling (FPS) selects NN keypoints P={pi=[vi,si]T∈R4}i=1NP = \{p_i = [v_i, s_i]^T \in \mathbb{R}^4\}_{i=1}^N, where vi=[xi,yi,zi]v_i = [x_i, y_i, z_i].

    2. For each sampled keypoint pi∈Pp_i \in P, a spherical voxel hih_i is defined by querying all points within a fixed Euclidean radius rr:

    hi={p1,p2,…,pj,⋯∣∥vi−vj∥2<r}h_i = \{p_1, p_2, \dots, p_j, \dots \mid \|v_i - v_j\|_2 < r\}

    partitioning the 3D space into NN spherical voxels H={h1,h2,…,hN}H = \{h_1, h_2, \dots, h_N\}.

    1. For each spherical voxel hi={pj=[xj,yj,zj,sj]T}j=1th_i = \{p_j = [x_j, y_j, z_j, s_j]^T\}_{j=1}^t containing tt points, local point-wise features are extracted via a shared Multi-Layer Perceptron (MLP):

    f(hi)=MLP(pj)j=1,2,…,tf(h_i) = \text{MLP}(p_j)_{j=1, 2, \dots, t}

    yielding the initial point-wise feature set F={fi}i=1NF = \{f_i\}_{i=1}^N across all voxels.

  2. Knowl 2 — Dual Local-Complete and Global-KNN Graph Attention Mechanism

    model/method

    SVGA-Net learns relational representations of point clouds by combining local intra-voxel complete graphs with a global inter-voxel KK-nearest neighbor (KNN) graph:

    1. Local Complete Graph: Within each spherical voxel hih_i, a fully connected directed graph is established over its constituent points. The local attention coefficient αj,k\alpha_{j,k} between point node pjp_j (with feature fjf_j) and point node pkp_k (with feature fj,kf_{j,k}) in the intra-voxel index set U(pj)\mathcal{U}(p_j) is computed via normalized dot-product attention:

    αj,k=exp⁡(fjTfj,k)∑k∈U(pj)exp⁡(fjTfj,k)\alpha_{j,k} = \frac{\exp(f_j^T f_{j,k})}{\sum_{k \in \mathcal{U}(p_j)} \exp(f_j^T f_{j,k})}

    1. Global Attention Layer: The physical centroid cic_i of each spherical voxel hih_i is mapped via a 3-layer MLP into an initial global descriptor fg,i∈Fg={fg,1,…,fg,N}f_{g,i} \in F_g = \{f_{g,1}, \dots, f_{g,N}\}. A KNN graph connects each voxel centroid fg,if_{g,i} to its KK nearest neighbor centroids ϖ(fg,i)\varpi(f_{g,i}). The attention weight between centroid fg,if_{g,i} and its ll-th neighbor fg,i,lf_{g,i,l} is:

    βm,i=fg,iTfg,i,l∑l∈ϖ(fg,i)fg,iTfg,i,l\beta_{m,i} = \frac{f_{g,i}^T f_{g,i,l}}{\sum_{l \in \varpi(f_{g,i})} f_{g,i}^T f_{g,i,l}}

    The final global modulation factor βm\beta_m in the mm-th point attention layer is the mean over all KK neighbors, ensuring node permutation invariance.

    1. Feature Aggregation: The updated node feature fj′f'_j dynamically incorporates local attention-weighted neighbor features and the global factor:

    fj′=βm⋅fj+∑k∈U(pj)αj,k⋅fj,kf'_j = \beta_m \cdot f_j + \sum_{k \in \mathcal{U}(p_j)} \alpha_{j,k} \cdot f_{j,k}

    Each updated feature is transformed by a 2-layer MLP with non-linear activations. Stacking nn such attention layers followed by channel-wise max pooling yields the final aggregated descriptor for each spherical voxel.

  3. Knowl 3 — Sparse-to-Dense Regression Module for 3D Bounding Box Proposal

    model/method

    The Sparse-to-Dense Regression (SDR) module converts aggregated spherical voxel descriptors into dense spatial feature representations across multiple scales for classification and 3D bounding box regression:

    1. Multi-Scale Downsampling: Three successive convolutional blocks process the input voxel features into hierarchical representations b1,b2,b3b_1, b_2, b_3. The initial layer of each block applies a stride of 2 to downsample the spatial resolution by half, followed by sequences of stride-1 convolutions, Batch Normalization, and ReLU activations.

    2. Cross-Scale Receptive Field Fusion: To combine high-resolution fine features with low-resolution large-receptive-field features, the feature map of the second block b2b_2 is concatenated with upsampled b1b_1, and b3b_3 is concatenated with upsampled b2b_2.

    3. Dense Scale Channels: Parallel convolutions and upsampling operations are applied to each of the three fused streams to produce three dense feature maps F1,F2,F3F_1, F_2, F_3 having identical spatial dimensions.

    4. Multi-Level Residual Addition: The original upsampled block outputs b1,b2,b3b_1, b_2, b_3 are added element-wise to the corresponding processed dense maps F1,F2,F3F_1, F_2, F_3.

    5. Final Head Aggregation: The three residual-fused feature maps are concatenated and processed by a 3×33 \times 3 convolutional layer to generate the final feature map FsF_s, which is routed to separate classification and 3D bounding box regression heads.

  4. Knowl 4 — Multi-Task Loss and Box Regression Parameterization in SVGA-Net

    equation

    In SVGA-Net, prior 3D anchors are parameterized as (xa,ya,za,la,wa,ha,θa)(x_a, y_a, z_a, l_a, w_a, h_a, \theta_a) and ground truth 3D bounding boxes as (xgt,ygt,zgt,lgt,wgt,hgt,θgt)(x_{gt}, y_{gt}, z_{gt}, l_{gt}, w_{gt}, h_{gt}, \theta_{gt}), where (x,y,z)(x, y, z) are centroid coordinates, (l,w,h)(l, w, h) are length, width, and height, and θ\theta is the heading orientation angle.

    The 7-dimensional regression residual targets Δt∈{Δx,Δy,Δz,Δw,Δl,Δh,Δθ}\Delta t \in \{\Delta x, \Delta y, \Delta z, \Delta w, \Delta l, \Delta h, \Delta \theta\} are parameterized as:

    Δx=xgt−xada,Δy=ygt−yada,Δz=zgt−zaha\Delta x = \frac{x_{gt} - x_a}{d_a}, \quad \Delta y = \frac{y_{gt} - y_a}{d_a}, \quad \Delta z = \frac{z_{gt} - z_a}{h_a}

    Δw=log⁡(wgtwa),Δl=log⁡(lgtla),Δh=log⁡(hgtha)\Delta w = \log\left(\frac{w_{gt}}{w_a}\right), \quad \Delta l = \log\left(\frac{l_{gt}}{l_a}\right), \quad \Delta h = \log\left(\frac{h_{gt}}{h_a}\right)

    Δθ=sin⁡(θgt−θa)\Delta \theta = \sin(\theta_{gt} - \theta_a)

    where da=wa2+la2d_a = \sqrt{w_a^2 + l_a^2} is the base diagonal length of the anchor.

    The overall multi-task training loss is:

    Ltotal=αLcls+β1Npos∑t∈{x,y,z,l,w,h,θ}Lreg(Δt∗,Δt)L_{total} = \alpha L_{cls} + \beta \frac{1}{N_{pos}} \sum_{t \in \{x, y, z, l, w, h, \theta\}} L_{reg}(\Delta t^*, \Delta t)

    where Δt∗\Delta t^* is the predicted residual, LregL_{reg} is the Smooth L1 loss, NposN_{pos} and NnegN_{neg} are the counts of positive and negative anchors, and balancing weights are fixed to α=1\alpha = 1 and β=2\beta = 2.

    The classification loss LclsL_{cls} is defined as a balanced binary cross-entropy loss:

    Lcls=γ11Npos∑iLcls(pipos,1)+γ21Nneg∑iLcls(pineg,0)L_{cls} = \gamma_1 \frac{1}{N_{pos}} \sum_i L_{cls}(p_i^{pos}, 1) + \gamma_2 \frac{1}{N_{neg}} \sum_i L_{cls}(p_i^{neg}, 0)

    where piposp_i^{pos} and pinegp_i^{neg} are the predicted probabilities for positive and negative anchors, with hyperparameters γ1=1.5\gamma_1 = 1.5 and γ2=1.0\gamma_2 = 1.0.

  5. Knowl 5 — 3D Object Detection Performance Comparison on the KITTI Test Benchmark

    data/table
    Method Modality APcarAP_{car} (%) APpedestrianAP_{pedestrian} (%) APcyclistAP_{cyclist} (%)
    Easy Mod. Hard Easy Mod. Hard Easy Mod. Hard
    MV3D R+L 71.09 62.35 55.12 - - - - - -
    F-Pointnet R+L 81.20 70.39 62.19 51.21 44.89 40.23 71.96 56.77 50.39
    AVOD-FPN R+L 81.94 71.88 66.38 50.80 42.81 40.88 64.00 52.18 46.61
    F-ConvNet R+L 85.88 76.51 68.08 52.37 45.61 41.49 79.58 64.68 57.03
    MMF R+L 86.81 76.75 68.41 - - - - - -
    Voxelnet L 77.47 65.11 57.73 39.48 33.69 31.51 61.22 48.36 44.37
    SECOND L 83.13 73.66 66.20 51.07 42.56 37.29 70.51 53.85 46.90
    PointPillars L 79.05 74.99 68.30 52.08 43.43 41.49 75.78 59.07 52.92
    PointRCNN L 85.94 75.76 68.32 49.43 41.78 38.63 73.93 59.60 53.59
    STD L 86.61 77.63 76.06 53.08 44.24 41.97 78.89 62.53 55.77
    3DSSD L 88.36 79.57 74.55 - - - - - -
    SA-SSD L 88.75 79.79 74.16 - - - - - -
    PV-RCNN L 90.25 81.43 76.82 - - - 78.60 63.71 57.65
    Point-GNN L 88.33 79.47 72.29 51.92 43.77 40.14 78.60 63.48 57.08
    SVGA-Net (ours) L 87.33 80.47 75.91 48.48 40.39 37.92 78.58 62.28 54.88

    The table compares 3D bounding box detection Average Precision (APAP) on the official KITTI test server across Car, Pedestrian, and Cyclist categories on Easy, Moderate, and Hard difficulty levels. IoU thresholds for evaluation are 0.7 for Car and 0.5 for Pedestrian and Cyclist. Modality 'L' indicates LiDAR-only input and 'R+L' indicates RGB plus LiDAR multi-modal input. SVGA-Net operates strictly on LiDAR data, outperforming the top multi-modal method MMF on car detection across all three difficulty levels (+0.52%, +3.72%, and +7.50% AP), and outperforming the graph neural network method Point-GNN on Moderate and Hard car detection.

  6. Knowl 6 — 3D and BEV Vehicle Detection Performance on Waymo Open Dataset

    data/table
    Difficulty Method 3D mAP (IoU=0.7) BEV mAP (IoU=0.7)
    Overall 0–30m 30–50m 50m–Inf Overall 0–30m 30–50m 50m–Inf
    LEVEL 1 PointPillars 56.62 81.01 51.75 27.94 75.57 92.10 74.06 55.47
    MVF 62.93 86.30 60.02 36.02 80.40 93.59 79.21 63.09
    PV-RCNN 70.30 91.92 69.21 42.17 82.96 97.35 82.99 64.97
    SVGA-Net 73.45 92.53 69.44 42.08 83.52 97.60 83.14 64.25
    LEVEL 2 PV-RCNN 65.36 91.58 65.13 36.46 77.45 94.64 80.39 55.39
    SVGA-Net 66.65 91.65 66.78 39.29 80.97 95.54 81.58 60.18

    The table evaluates vehicle detection on 202 validation sequences of the Waymo Open Dataset across distance ranges (0--30m, 30--50m, 50m--∞\infty) using 3D mean Average Precision (3D mAP) and Bird's Eye View mean Average Precision (BEV mAP) at an IoU threshold of 0.7 for LEVEL 1 and LEVEL 2 difficulty objects. SVGA-Net achieves 73.45% overall 3D mAP and 83.52% overall BEV mAP on LEVEL 1, surpassing PV-RCNN by +3.15% 3D mAP, and achieves 66.65% 3D mAP and 80.97% BEV mAP on LEVEL 2.

  7. Knowl 7 — 3D and Bird's Eye View Detection Performance for Car on KITTI Validation Set

    data/table
    Method Modality APcarAP_{car} 3D (%) APcarAP_{car} BEV (%)
    Easy Moderate Hard Easy Moderate Hard
    MV3D R+L 71.29 62.68 56.56 86.55 78.10 76.67
    F-Pointnet R+L 83.76 70.92 63.65 88.16 84.02 76.44
    AVOD-FPN R+L 84.41 74.44 68.65 - - -
    F-ConvNet R+L 89.02 78.80 77.09 90.23 88.79 86.84
    Voxelnet L 81.97 65.46 62.85 89.60 84.81 78.57
    SECOND L 87.43 76.48 69.10 89.96 87.07 79.66
    PointRCNN L 88.88 78.63 77.38 - - -
    Fast PointRCNN L 89.12 79.00 77.48 90.12 88.10 86.24
    STD L 89.70 79.80 79.30 90.50 88.50 88.10
    SA-SSD L 90.15 79.91 78.78 - - -
    3DSSD L 89.71 79.45 78.67 - - -
    Point-GNN L 87.89 78.34 77.38 89.82 88.31 87.16
    SVGA-Net (ours) L 90.59 80.23 79.15 90.27 89.16 88.11

    The table presents car detection performance on the KITTI validation split (3769 point cloud frames) for 3D bounding box detection APAP and Bird's Eye View (BEV) detection APAP at an IoU threshold of 0.7. SVGA-Net achieves 90.59% Easy, 80.23% Moderate, and 79.15% Hard 3D AP, and 90.27% Easy, 89.16% Moderate, and 88.11% Hard BEV AP, achieving the highest Moderate accuracy among the compared LiDAR and multi-modal models.

  8. Knowl 8 — Ablation Analysis of Attention Layers, Global Context, SDR Design, and KNN Size

    data/table
    Ablation Variable Setting 3DAPcar3D AP_{car} (%)
    Easy Moderate Hard
    Attention Layers (nn) n=1n = 1 86.77 75.37 74.19
    n=2n = 2 88.86 78.81 78.03
    n=3n = 3 90.59 80.23 79.15
    n=4n = 4 89.62 79.26 77.58
    Global Attention Layer w/o global attention 88.42 78.11 76.54
    with global attention 90.59 80.23 79.15
    Regression Design SR (no b1,b2b_1, b_2 concat) 87.53 77.81 76.22
    DR (no bi+Fib_i + F_i addition) 88.39 78.44 76.56
    Full SDR 90.59 80.23 79.15
    Centroid KNN Neighbors (KK) K=1K = 1 76.37 69.15 68.47
    K=2K = 2 84.53 75.61 71.92
    K=3K = 3 90.59 80.23 79.15
    K=4K = 4 88.91 79.22 77.86
    K=5K = 5 86.58 76.82 75.43

    The ablation study evaluates structural choices on the KITTI validation split for the Car class:

    1. Setting the number of point-attention layers to n=3n=3 achieves optimal performance (80.23% Moderate AP), whereas n=4n=4 degrades accuracy due to over-learning.
    2. Incorporating the global attention layer provides a +2.12% Moderate AP improvement over omitting it.
    3. In the Sparse-to-Dense Regression module, removing multi-scale concatenation (SR) drops Moderate AP by 2.42%, and removing residual element-wise addition (DR) drops Moderate AP by 1.79%.
    4. Selecting K=3K = 3 nearest neighbors for the global centroid graph achieves the highest accuracy; smaller KK values fail to incorporate sufficient global context, while larger KK values oversmooth the representation.
  9. Knowl 9 — Training Configuration, Hyperparameters, and Inference Latency Breakdown of SVGA-Net

    experimental setup

    SVGA-Net is trained and parameterized with category-specific network configurations:

    • Car Detection: Point set sample count N=1024N = 1024, spherical voxel radius r=1.8 mr = 1.8\text{ m}. Anchors are positive if IoU≥0.6\text{IoU} \ge 0.6 or if they have the highest IoU with a ground truth box, and negative if IoU<0.45\text{IoU} < 0.45. Non-Maximum Suppression (NMS) IoU threshold is set to 0.70.7.

    • Pedestrian and Cyclist Detection: Sample count N=512N = 512, spherical voxel radius r=0.8 mr = 0.8\text{ m}. Anchors are positive if IoU≥0.5\text{IoU} \ge 0.5 (or highest IoU) and negative if IoU<0.35\text{IoU} < 0.35. NMS IoU threshold is set to 0.60.6.

    • Layer Dimensions: Local point-wise and global attention MLPs use 3 layers of sizes (64,128,128)(64, 128, 128). The n=3n = 3 stacked point-attention layers are followed by 2-layer MLPs with channel sizes (128,128)(128, 128), (128,256)(128, 256), and (512,1024)(512, 1024) respectively.

    • Training Schedule: Trained using ADAM with initial learning rate 0.0010.001 for the first 140 epochs, decayed by a factor of 10 every 20 epochs thereafter, for a total of 200 epochs with batch size 16 across 4 NVIDIA GTX 1080 GPUs.

    • Inference Runtime: On an NVIDIA GTX 1080 GPU, the total inference time per point cloud sample is 62 ms62\text{ ms}, broken down into:

      • Data reading and pre-processing: 9 ms9\text{ ms} (14.5%14.5\%)
      • Local and global feature aggregation: 41 ms41\text{ ms} (66.1%66.1\%)
      • Final box proposal detection: 12 ms12\text{ ms} (19.4%19.4\%)
  10. Knowl 10 — Performance Degradation of SVGA-Net Under Severe Occlusion

    limitation

    SVGA-Net exhibits reduced detection accuracy on smaller, highly occluded objects such as pedestrians and cyclists. When an object's occlusion ratio exceeds 80%80\%, the sparsity of returned LiDAR points inside the fixed spherical voxel prevents the construction of a reliable local complete graph, leading to ineffective intra-voxel attention and degraded feature aggregation.

Coverage note — No substantial contributed material was omitted; the extraction captures the complete architecture, graph attention mechanism, sparse-to-dense regression module, loss formulation, test and validation benchmark evaluations on KITTI and Waymo, ablation studies, runtime profiling, and stated occlusion limitations.

References

  1. 1.Bi, Y.; Chadha, A.; Abbas, A.; Bourtsoulatze, E.; and Andreopoulos, Y. 2019. Graph-Based Object Classification for Neuromorphic Vision Sensing. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 491–501.
  2. 2.Chen, X.; Ma, H.; Wan, J.; Li, B.; and Xia, T. 2017. Multi-view 3d object detection network for autonomous driving. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1907–1915.
  3. 3.Chen, Y.; Liu, S.; Shen, X.; and Jia, J. 2019. Fast point r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, 9775–9784.
  4. 4.Geiger, A.; Lenz, P.; and Urtasun, R. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, 3354–3361. IEEE.
  5. 5.Girshick, R. 2015. Fast R-CNN. In 2015 IEEE International Conference on Computer Vision (ICCV), 1440–1448.
  6. 6.Kaul, C.; Pears, N.; and Manandhar, S. 2019. SAWNet: A Spatially Aware Deep Neural Network for 3D Point Cloud Processing. arXiv preprint arXiv:1905.07650.
  7. 7.Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  8. 8.Ku, J.; Mozifian, M.; Lee, J.; Harakeh, A.; and Waslander, S. L. 2018. Joint 3d proposal generation and object detection from view aggregation. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1–8. IEEE.
  9. 9.Landrieu, L.; and Simonovsky, M. 2018. Large-Scale Point Cloud Semantic Segmentation with Superpoint Graphs. In CVPR 2018.
  10. 10.Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O. 2019. PointPillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 12697–12705.
  11. 11.Li, B. 2017. 3d fully convolutional network for vehicle detection in point cloud. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1513–1518. IEEE.
  12. 12.Liang, M.; Yang, B.; Chen, Y.; Hu, R.; and Urtasun, R. 2019. Multi-task multi-sensor fusion for 3d object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7345–7353.
  13. 13.Liang, M.; Yang, B.; Wang, S.; and Urtasun, R. 2018. Deep continuous fusion for multi-sensor 3d object detection. In Proceedings of the European Conference on Computer Vision (ECCV), 641–656.
  14. 14.Liu, Z.; Tang, H.; Lin, Y.; and Han, S. 2019. Point-Voxel CNN for efficient 3D deep learning. In Advances in Neural Information Processing Systems, 963–973.
  15. 15.Liu, Z.; Zhao, X.; Huang, T.; Hu, R.; Zhou, Y.; and Bai, X. 2020. TANet: Robust 3D Object Detection from Point Clouds with Triple Attention. AAAI.
  16. 16.Park, Y.; Lepetit, V.; and Woo, W. 2008. Multiple 3D Object Tracking for Augmented Reality. In Proceedings of the 7th IEEE/ACM International Symposium on Mixed and Augmented Reality, ISMAR ’08, 117–120. Washington, DC, USA: IEEE Computer Society. ISBN 978-1-4244-2840-3.
  17. 17.Qi, C. R.; Liu, W.; Wu, C.; Su, H.; and Guibas, L. J. 2018. Frustum pointnets for 3d object detection from rgb-d data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 918–927.
  18. 18.Qi, C. R.; Su, H.; Mo, K.; and Guibas, L. J. 2017a. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 652–660.
  19. 19.Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017b. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in neural information processing systems, 5099–5108.
  20. 20.Qi, X.; Liao, R.; Jia, J.; Fidler, S.; and Urtasun, R. 2017. 3D Graph Neural Networks for RGBD Semantic Segmentation. In 2017 IEEE International Conference on Computer Vision (ICCV), 5209–5218.
  21. 21.Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 779–788.
  22. 22.Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, 91–99.
  23. 23.Shen, Y.; Feng, C.; Yang, Y.; and Tian, D. 2018. Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4548–4557.
  24. 24.Shi, S.; Guo, C.; Jiang, L.; Wang, Z.; Shi, J.; Wang, X.; and Li, H. 2020. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10529–10538.
  25. 25.Shi, S.; Wang, X.; and Li, H. 2019. Pointrcnn: 3d object proposal generation and detection from point cloud. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–779.
  26. 26.Shi, W.; and Rajkumar, R. R. 2020. Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  27. 27.Simon, M.; Amende, K.; Kraus, A.; Honer, J.; Samann, T.; Kaulbersch, H.; Milz, S.; and Michael Gross, H. 2019. Complexer-yolo: Real-time 3d object detection and tracking on semantic point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 0–0.
  28. 28.Sindagi, V. A.; Zhou, Y.; and Tuzel, O. 2019. MVX-Net: Multimodal VoxelNet for 3D Object Detection. In 2019 International Conference on Robotics and Automation (ICRA), 7276–7282.
  29. 29.Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B.; et al. 2020. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2446–2454.
  30. 30.Wang, Y.; Sun, Y.; Liu, Z.; Sarma, S. E.; Bronstein, M. M.; and Solomon, J. M. 2019. Dynamic Graph CNN for Learning on Point Clouds. ACM Trans. Graph., 38(5): 146:1–146:12.
  31. 31.Wang, Z.; and Jia, K. 2019. Frustum ConvNet: Sliding Frustums to Aggregate Local Point-Wise Features for Amodal. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1742–1749. IEEE.
  32. 32.Xie, S.; Liu, S.; Chen, Z.; and Tu, Z. 2018. Attentional ShapeContextNet for Point Cloud Recognition. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4606–4615.
  33. 33.Yan, Y.; Mao, Y.; and Li, B. 2018. Second: Sparsely embedded convolutional detection. Sensors, 18(10): 3337.
  34. 34.Yang, B.; Luo, W.; and Urtasun, R. 2018. Pixor: Real-time 3d object detection from point clouds. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 7652–7660.
  35. 35.Yang, Z.; Sun, Y.; Liu, S.; and Jia, J. 2020. 3dssd: Point-based 3d single stage object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11040–11048.
  36. 36.Yang, Z.; Sun, Y.; Liu, S.; Shen, X.; and Jia, J. 2019. Std: Sparse-to-dense 3d object detector for point cloud. In Proceedings of the IEEE International Conference on Computer Vision, 1951–1960.
  37. 37.Yu, Y.; Huang, Z.; Li, F.; Zhang, H.; and Le, X. 2020. Point Encoder GAN: A deep learning model for 3D point cloud inpainting. Neurocomputing, 384: 192–199.
  38. 38.Zhang, J.; Su, Q.; Wang, C.; and Gu, H. 2020. Monocular 3D Vehicle Detection with Multi-instance Depth and Geometry Reasoning for Autonomous Driving. Neurocomputing, 403: 182–192.
  39. 39.Zhou, Y.; Sun, P.; Zhang, Y.; Anguelov, D.; Gao, J.; Ouyang, T.; Guo, J.; Ngiam, J.; and Vasudevan, V. 2020. End-to-end multi-view fusion for 3d object detection in lidar point clouds. In Conference on Robot Learning, 923–932. PMLR.
  40. 40.Zhou, Y.; and Tuzel, O. 2018. Voxelnet: End-to-end learning for point cloud based 3d object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4490–4499.

Citation

MLA
He, Q., et al. “SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds”. arXiv, 2020, http://arxiv.org/abs/2006.04043v2.
APA
He, Q., Wang, Z., Zeng, H., Zeng, Y., & Liu, Y. (2020). SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds. arXiv. http://arxiv.org/abs/2006.04043v2
Chicago
He, Q., Z. Wang, H. Zeng, Y. Zeng, and Y. Liu. 2020. “SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds”. arXiv. http://arxiv.org/abs/2006.04043v2.
Harvard
He, Q. et al. (2020) “SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2006.04043v2.
Vancouver
1. He Q, Wang Z, Zeng H, Zeng Y, Liu Y (2020) SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds. arXiv

BibTeX

@article{he2020svga,
  title = {SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object Detection from Point Clouds},
  author = {He, Qingdong and Wang, Zhengning and Zeng, Hao and Zeng, Yi and Liu, Yijun},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2006.04043v2},
  eprint = {2006.04043}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF