PointConv: Deep Convolutional Networks on 3D Point Clouds

Wenxuan WuZhongang QiFuxin Li

article2018CVPR1,912 citations

Proposes PointConv, an efficient convolution operator that combines continuous weight functions with local density estimation to build scalable, deep neural networks directly on irregular 3D point clouds.

Listen

Three-dimensional point clouds generated by sensors like LIDAR are central to autonomous navigation, robotics, and augmented reality. Unlike standard two-dimensional images arranged on uniform grids, point clouds are irregular, unordered, and unevenly sampled. Traditional convolutional neural networks cannot process these sets directly without either discarding fine structural details or converting them into computationally prohibitive 3D volumetric grids.

The article aims to introduce and validate PointConv, a novel continuous convolution operation that runs directly on unordered 3D point sets while accounting for non-uniform sampling density. The researchers also set out to provide an efficient mathematical formulation capable of scaling these operations to deep neural network architectures.

The authors constructed PointConv by treating convolution filters as continuous functions of local relative coordinates, using multi-layer perceptrons to learn spatial weights alongside an inverse density estimation step to compensate for non-uniform sampling. To address the massive memory overhead of generating individual filter weights, the authors changed the summation order, reducing the operation to standard matrix multiplication and simple two-dimensional convolutions. The method was evaluated on synthetic benchmarks for 3D object classification and part segmentation, a real-world dataset of complex indoor scene scans, and a 2D image classification benchmark treated as a point cloud.

The evaluation produced several key findings. First, on real-world indoor scene segmentation, the approach achieved a mean intersection-over-union score of 55.6%, outperforming earlier methods that scored between 30.6% and 43.8%. Second, the efficient reformulation cut memory usage down to roughly 1/64th of the original version, making modern deep architectures feasible. Third, on 3D synthetic benchmarks, the model reached state-of-the-art results, scoring 92.5% accuracy on 40-class object recognition and an 85.7% instance average intersection-over-union on part segmentation. Finally, when tested on 2D image data treated as irregular points, the method matched the 93% accuracy of standard convolutional networks, showing it functions as a true general convolution.

These findings demonstrate that deep neural networks can process raw point cloud data directly with high accuracy and low memory usage, eliminating the need to project data into cumbersome voxel grids. By achieving both translation invariance and point order invariance, this architecture provides a reliable foundation for high-performance 3D spatial perception in autonomous vehicles, mobile robotics, and computer vision systems.

Organizations developing 3D perception pipelines should consider adopting PointConv-based architectures to improve accuracy and efficiency in high-resolution spatial tasks. Future development efforts should focus on integrating this convolution operation into deeper modern frameworks, such as residual and densely connected network backbones, to evaluate potential gains across larger operational pipelines.

While the method shows strong performance across synthetic and indoor datasets, its effectiveness relies on accurate local neighborhood estimation and offline kernel density calculations. Practitioners should test the approach in varying outdoor lighting, heavy sensor noise, and adverse weather conditions before deploying it in safety-critical autonomous platforms.

Cover for PointConv: Deep Convolutional Networks on 3D Point Clouds

Abstract

Unlike images which are represented in regular dense grids, 3D point clouds are irregular and unordered, hence applying convolution on them can be difficult. In this paper, we extend the dynamic filter to a new convolution operation, named PointConv. PointConv can be applied on point clouds to build deep convolutional networks. We treat convolution kernels as nonlinear functions of the local coordinates of 3D points comprised of weight and density functions. With respect to a given point, the weight functions are learned with multi-layer perceptron networks and density functions through kernel density estimation. The most important contribution of this work is a novel reformulation proposed for efficiently computing the weight functions, which allowed us to dramatically scale up the network and significantly improve its performance. The learned convolution kernel can be used to compute translation-invariant and permutation-invariant convolution on any point set in the 3D space. Besides, PointConv can also be used as deconvolution operators to propagate features from a subsampled point cloud back to its original resolution. Experiments on ModelNet40, ShapeNet, and ScanNet show that deep convolutional neural networks built on PointConv are able to achieve state-of-the-art on challenging semantic segmentation benchmarks on 3D point clouds. Besides, our experiments converting CIFAR-10 into a point cloud showed that networks built on PointConv can match the performance of convolutional networks in 2D images of a similar structure.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 PointConv
  • 3.1 Convolution on 3D Point Clouds
  • 3.2 Feature Propagation Using Deconvolution
  • 4 Efficient PointConv
  • 5 Experiments
  • 5.1 Classification on ModelNet40
  • 5.2 ShapeNet Part Segmentation
  • 5.3 Semantic Scene Labeling
  • 5.4 Classification on CIFAR-10
  • 6 Ablation Experiments and Visualizations
  • 6.1 The Structure of MLP
  • 6.2 Inverse Density Scale
  • 6.3 Ablation Studies on ScanNet
  • 6.4 Visualization
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — PointConv Operator for 3D Point Clouds

    model/method

    PointConv is a permutation-invariant and translation-invariant continuous convolution operation adapted for 3D point clouds with non-uniform point distributions. Given a reference point (x,y,z)∈R3(x, y, z) \in \mathbb{R}^3 and its local neighborhood GG containing points at relative coordinates (δx,δy,δz)(\delta_x, \delta_y, \delta_z), continuous 3D convolution is formulated via Monte Carlo approximation as:

    PointConv(S,W,F)xyz=∑(δx,δy,δz)∈GS(δx,δy,δz)W(δx,δy,δz)F(x+δx,y+δy,z+δz)PointConv(S, W, F)_{xyz} = \sum_{(\delta_x, \delta_y, \delta_z) \in G} S(\delta_x, \delta_y, \delta_z) W(\delta_x, \delta_y, \delta_z) F(x + \delta_x, y + \delta_y, z + \delta_z)

    where F(x+δx,y+δy,z+δz)∈RCinF(x + \delta_x, y + \delta_y, z + \delta_z) \in \mathbb{R}^{C_{in}} is the input feature vector at the neighbor point, W(δx,δy,δz)∈RCin×CoutW(\delta_x, \delta_y, \delta_z) \in \mathbb{R}^{C_{in} \times C_{out}} is a continuous weight function approximated by a multi-layer perceptron (MLP) shared across points taking 3D relative coordinates as input, and S(δx,δy,δz)∈RS(\delta_x, \delta_y, \delta_z) \in \mathbb{R} is an inverse density scale factor that compensates for non-uniform spatial sampling.

    For a local region around a centroid with KK neighboring points, input features Fin∈RK×CinF_{in} \in \mathbb{R}^{K \times C_{in}}, relative 3D coordinates Plocal∈RK×3P_{local} \in \mathbb{R}^{K \times 3}, and density scales S∈RKS \in \mathbb{R}^K, the output feature vector Fout∈RCoutF_{out} \in \mathbb{R}^{C_{out}} is computed discretely as:

    Fout=∑k=1K∑cin=1CinS(k)W(k,cin)Fin(k,cin)F_{out} = \sum_{k=1}^{K} \sum_{c_{in}=1}^{C_{in}} S(k) W(k, c_{in}) F_{in}(k, c_{in})

    where k∈{1,…,K}k \in \{1, \dots, K\} indexes neighbor points, cin∈{1,…,Cin}c_{in} \in \{1, \dots, C_{in}\} indexes input feature channels, and W(k,cin)∈RCoutW(k, c_{in}) \in \mathbb{R}^{C_{out}} is the weight vector predicted for the kk-th neighbor and cinc_{in}-th input channel.

  2. Knowl 2 — Memory-Efficient PointConv Reformulation

    theoretical result

    A naive implementation of PointConv computes and stores explicit filter weight tensors of size B×N×K×(Cin×Cout)B \times N \times K \times (C_{in} \times C_{out}) (where BB is batch size, NN is point count, KK is neighbors per local region, CinC_{in} is input channel count, and CoutC_{out} is output channel count), resulting in prohibitive memory overhead. PointConv can be equivalently reformulated by decomposing the last linear layer of the weight MLP and changing the summation order.

    Let the weight function MLP's last linear layer have weight parameters H∈RCmid×(Cin×Cout)H \in \mathbb{R}^{C_{mid} \times (C_{in} \times C_{out})}, and let M∈RK×CmidM \in \mathbb{R}^{K \times C_{mid}} be the intermediate activations input to this last layer for the KK neighbor points. Defining density-scaled input features F~in=S⋅Fin∈RK×Cin\tilde{F}_{in} = S \cdot F_{in} \in \mathbb{R}^{K \times C_{in}}, the PointConv output feature Fout∈RCoutF_{out} \in \mathbb{R}^{C_{out}} is equivalent to:

    Fout=Conv1×1(H,F~inTM)F_{out} = \text{Conv}_{1 \times 1}\left(H, \tilde{F}_{in}^T M\right)

    where F~inTM∈RCin×Cmid\tilde{F}_{in}^T M \in \mathbb{R}^{C_{in} \times C_{mid}} is computed via standard matrix multiplication, followed by a 1×11 \times 1 convolution with weight kernel HH.

    This factorization reduces the memory required for intermediate convolution filters from O(K⋅Cin⋅Cout)O(K \cdot C_{in} \cdot C_{out}) to O(K⋅Cmid)O(K \cdot C_{mid}), which is a reduction factor of:

    CmidK⋅Cout\frac{C_{mid}}{K \cdot C_{out}}

    For a representative setting with B=32,N=512,K=32,Cin=64,Cout=64B = 32, N = 512, K = 32, C_{in} = 64, C_{out} = 64, and Cmid=32C_{mid} = 32, the memory footprint for one layer drops from 8.0 GB to 0.1255 GB (a ∼64×\sim 64\times reduction).

  3. Knowl 3 — Inverse Density Scaling via Non-linear Transformed Kernel Density Estimation

    model/method

    Point clouds collected by physical sensors are frequently sampled non-uniformly, which biases discrete Monte Carlo approximations of continuous convolutions. PointConv corrects this sampling bias by incorporating an inverse density scale factor S(δx,δy,δz)S(\delta_x, \delta_y, \delta_z) computed via a two-stage process:

    1. Offline Kernel Density Estimation (KDE): For each point in the point cloud, spatial point density is calculated offline using kernelized density estimation.
    2. 1D Non-linear Transformation: The estimated density scalar for each point is fed into a 1D multi-layer perceptron (MLP) to produce the inverse density weight SS.

    Passing the KDE output through a learned non-linear MLP allows the network to adaptively modulate whether and how strongly to apply inverse density re-weighting across different feature layers and spatial scales.

  4. Knowl 4 — PointDeconv Operator for Hierarchical Feature Propagation

    model/method

    PointDeconv is a deconvolution operator designed to propagate coarse, subsampled point features back to higher-resolution point sets for dense point-wise prediction tasks like semantic scene segmentation.

    PointDeconv consists of two sequential operations:

    1. Distance-Weighted 3-Nearest-Neighbor Interpolation: Coarser point features from the preceding layer are mapped to higher-resolution point positions by linear interpolation using the 3 nearest neighbor points in 3D Euclidean space.
    2. Skip-Link Concatenation and PointConv Convolution: The interpolated features are concatenated along the channel dimension with encoding features of matching resolution via skip links. A PointConv operation is then applied to the concatenated feature representations to capture local geometric dependencies.

    This process is repeated across the decoder hierarchy until point features are fully restored to the original input resolution.

  5. Knowl 5 — Semantic Scene Segmentation on ScanNet Benchmark

    empirical result

    PointConv was evaluated on the ScanNet indoor 3D semantic segmentation benchmark (1513 training scans, 100 hidden test scans) using mean Intersection-over-Union (mIoU) across 20 object categories with 3D coordinates and RGB inputs.

    Method mIoU (%)
    ScanNet baseline 30.6
    PointNet++ 33.9
    SPLATNet 39.3
    Tangent Convolutions 43.8
    PointConv 55.6

    PointConv achieved an official test mIoU of 55.6%55.6\%, outperforming prior point-based and volumetric/tangent convolution baselines by a significant margin. On an NVIDIA GTX 1080Ti GPU, PointConv requires approximately 170 seconds per training epoch on ScanNet, with an evaluation time of roughly 0.5 seconds for 8×81928 \times 8192 points.

  6. Knowl 6 — 3D Object Classification Performance on ModelNet40

    empirical result

    PointConv was evaluated on the ModelNet40 shape classification benchmark (12,311 CAD models across 40 classes: 9,843 training and 2,468 testing models) using 1,024 points per object along with computed surface normal vectors.

    Method Input Representation Accuracy (%)
    Subvolume Voxels 89.2
    ECC Graphs 87.4
    Kd-Network 1024 points 91.8
    PointNet 1024 points 89.2
    PointNet++ 1024 points 90.2
    PointNet++ 5000 points + normal 91.9
    SpiderCNN 1024 points + normal 92.4
    PointConv 1024 points + normal 92.5

    PointConv achieves 92.5%92.5\% classification accuracy, achieving state-of-the-art performance among 3D point and graph input representations while utilizing only 1,024 points.

  7. Knowl 7 — 3D Part Segmentation Performance on ShapeNet

    empirical result

    PointConv was evaluated on the ShapeNet Part Segmentation dataset (16,881 shapes across 16 categories and 50 part annotations) with 3D point coordinates and surface normal features. Evaluation was conducted using class-average mean Intersection-over-Union (class mIoU) and instance-average mean Intersection-over-Union (instance mIoU).

    Method Class Avg. mIoU (%) Instance Avg. mIoU (%)
    Kd-net 77.4 82.3
    PointNet 80.4 83.7
    PointNet++ 81.9 85.1
    SSCNN 82.0 84.7
    SPLATNet3D_{3D} 82.0 84.6
    SpiderCNN 82.4 85.3
    SSCN – 86.0
    PointConv 82.8 85.7

    PointConv achieves 82.8%82.8\% class average mIoU and 85.7%85.7\% instance average mIoU, demonstrating competitive performance against specialized point-cloud and sparse-voxel segmentation networks.

  8. Knowl 8 — 2D Image Convolution Equivalence on CIFAR-10

    empirical result

    To test the property that PointConv acts as a continuous generalization of conventional 2D grid convolutions, CIFAR-10 images (32×3232 \times 32 pixels) were converted into 2D point clouds where each pixel is represented as a point with 2D coordinates (x,y)(x, y) scaled to the unit ball and RGB color channels as features.

    Network / Method Accuracy (%)
    Image Convolution (5-layer) 88.52
    AlexNet 89.00
    VGG19 93.60
    SpiderCNN 77.97
    PointCNN 80.22
    PointConv (5-layer) 89.13
    PointConv (VGG19) 93.19

    A 5-layer PointConv network achieves 89.13%89.13\% accuracy, matching a standard 5-layer 2D image convolution (88.52%88.52\%) and AlexNet (89.00%89.00\%), while outperforming alternative point architectures PointCNN (80.22%80.22\%) and SpiderCNN (77.97%77.97\%). When configured with the VGG19 topology, PointConv achieves 93.19%93.19\%, on par with the standard 2D image VGG19 (93.60%93.60\%).

  9. Knowl 9 — Ablation Analysis of Density Scaling and Sliding Window on ScanNet

    empirical result

    Ablation experiments on the ScanNet validation set evaluate the contributions of the inverse density scale SS, the non-linear 1D MLP density transformation, input modalities (xyzxyz vs. xyz+RGBxyz+\text{RGB}), and sliding window evaluation stride size (0.5 m,1.0 m,1.5 m0.5\,\text{m}, 1.0\,\text{m}, 1.5\,\text{m}).

    Input Stride Size (m) PointConv mIoU (%) No Density mIoU (%) Density (no MLP) mIoU (%)
    xyz 0.5 61.0 60.3 60.1
    xyz 1.0 59.0 58.2 57.7
    xyz 1.5 58.2 56.9 57.3
    xyz+RGB 0.5 60.8 58.9 –
    xyz+RGB 1.0 58.6 56.7 –
    xyz+RGB 1.5 57.5 56.1 –

    Incorporating learned inverse density scaling consistently improves segmentation mIoU across all sliding window stride sizes. Applying density directly without the 1D MLP non-linear transformation results in lower performance than omitting density scaling altogether, confirming that the learnable transformation is necessary to adaptively scale point weights.

Coverage note — Ablation sweeps on the number of MLP layers and intermediate channel size C_mid on a ScanNet classification subset (Section 6.1) and qualitative 2D/3D continuous filter visualizations (Section 6.4) were omitted as standalone knowls, as their essential takeaways are incorporated into the efficient PointConv reformulation and ablation knowls.

References

  1. 1.Michael M Bronstein and Iasonas Kokkinos. Scale-invariant heat kernel signatures for non-rigid shape recognition. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 1704–1711. IEEE, 2010.
  2. 2.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 2, 6, 7
  3. 3.Ding-Yun Chen, Xiao-Pei Tian, Yu-Te Shen, and Ming Ouhyoung. On visual similarity based 3d model retrieval. In Computer graphics forum, volume 22, pages 223–232. Wiley Online Library, 2003.
  4. 4.Hang Chu, Wei-Chiu Ma3 Kaustav Kundu, Raquel Urtasun, and Sanja Fidler. Surfconv: Bridging 3d and 2d convolution for rgbd images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3002–3011, 2018.
  5. 5.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), volume 1, 2017. 2, 6, 7, 8
  6. 6.Yi Fang, Jin Xie, Guoxian Dai, Meng Wang, Fan Zhu, Tiantian Xu, and Edward Wong. 3d deep shape descriptor. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2319–2328, 2015.
  7. 7.Benjamin Graham and Laurens van der Maaten. Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307, 2017. 2, 6, 7
  8. 8.Adrien Gressin, Clement Mallet, J er ome Demantk e, and Nicolas David. Towards 3d lidar point cloud registration improvement using optimal neighborhood knowledge. ISPRS journal of photogrammetry and remote sensing, 79:240–251, 2013.
  9. 9.Fabian Groh, Patrick Wieschollek, and Hendrik Lensch. Flex-convolution (deep learning beyond grid-worlds). arXiv preprint arXiv:1803.07289, 2018. 2
  10. 10.Kan Guo, Dongqing Zou, and Xiaowu Chen. 3d mesh labeling via deep convolutional neural networks. ACM Transactions on Graphics (TOG), 35(1):3, 2015.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  12. 12.Pedro Hermosilla, Tobias Ritschel, Pere-Pau Vazquez, Alvar Vinacua, and Timo Ropinski. Monte carlo convolution for learning on non-uniformly sampled point clouds. In SIGGRAPH Asia 2018 Technical Papers, page 235. ACM, 2018. 3, 4
  13. 13.Binh-Son Hua, Minh-Khoi Tran, and Sai-Kit Yeung. Pointwise convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 984–993, 2018. 2
  14. 14.Qiangui Huang, Weiyue Wang, and Ulrich Neumann. Recurrent slice networks for 3d segmentation of point clouds. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2626–2635, 2018.
  15. 15.Jorn-Henrik Jacobsen, Jan van Gemert, Zhongyou Lou, and Arnold WM Smeulders. Structured receptive fields in cnns. In Computer Vision and Pattern Recognition (CVPR), 2016 IEEE Conference on, pages 2610–2619. IEEE, 2016.
  16. 16.Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool. Dynamic filter networks. In Advances in Neural Information Processing Systems, pages 667–675, 2016. 1, 3, 4
  17. 17.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  18. 18.Roman Klokov and Victor Lempitsky. Escape from cells: Deep kd-networks for the recognition of 3d point cloud models. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 863–872. IEEE, 2017. 2, 6, 7
  19. 19.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009. 6
  20. 20.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 7
  21. 21.Yangyan Li, Rui Bu, Mingchao Sun, and Baoquan Chen. Pointcnn. arXiv preprint arXiv:1801.07791, 2018. 2, 7
  22. 22.Haibin Ling and David W Jacobs. Shape classification using the inner-distance. IEEE transactions on pattern analysis and machine intelligence, 29(2):286–299, 2007.
  23. 23.Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In Intelligent Robots and Systems (IROS), 2015 IEEE/RSJ International Conference on, pages 922–928. IEEE, 2015. 2
  24. 24.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1520–1528, 2015. 2, 4
  25. 25.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. arXiv preprint arXiv:1711.08488, 2017.
  26. 26.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 1(2):4, 2017. 2, 6
  27. 27.Charles R Qi, Hao Su, Matthias Nießner, Angela Dai, Mengyuan Yan, and Leonidas J Guibas. Volumetric and multi-view cnns for object classification on 3d data. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5648–5656, 2016. 2, 6
  28. 28.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, pages 5105–5114, 2017. 2, 4, 6, 7
  29. 29.Xiaojuan Qi, Renjie Liao, Jiaya Jia, Sanja Fidler, and Raquel Urtasun. 3d graph neural networks for rgbd semantic segmentation. In Proceedings of theqi IEEE Conference on Computer Vision and Pattern Recognition, pages 5199–5208, 2017.
  30. 30.Siamak Ravanbakhsh, Jeff Schneider, and Barnabas Poczos. Deep learning with sets and point clouds. arXiv preprint arXiv:1611.04500, 2016. 2
  31. 31.Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. Octnet: Learning deep 3d representations at high resolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 3, 2017. 2
  32. 32.Radu Bogdan Rusu, Nico Blodow, and Michael Beetz. Fast point feature histograms (fpfh) for 3d registration. In Robotics and Automation, 2009. ICRA’09. IEEE International Conference on, pages 3212–3217. IEEE, 2009.
  33. 33.Martin Simonovsky and Nikos Komodakis. Dynamic edgeconditioned filters in convolutional neural networks on graphs. In Proc. CVPR, 2017. 1, 3, 4, 5, 6
  34. 34.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 7
  35. 35.Hang Su, Varun Jampani, Deqing Sun, Subhransu Maji, Evangelos Kalogerakis, Ming-Hsuan Yang, and Jan Kautz. Splatnet: Sparse lattice networks for point cloud processing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2530–2539, 2018. 2, 6, 7
  36. 36.Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller. Multi-view convolutional neural networks for 3d shape recognition. In Proceedings of the IEEE international conference on computer vision, pages 945–953, 2015. 2
  37. 37.Maxim Tatarchenko, Jaesik Park, Vladlen Koltun, and QianYi Zhou. Tangent convolutions for dense prediction in 3d. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3887–3896, 2018. 2, 7
  38. 38.Berwin A Turlach. Bandwidth selection in kernel density estimation: A review. In CORE and Institut de Statistique. Citeseer, 1993. 4
  39. 39.Nitika Verma, Edmond Boyer, and Jakob Verbeek. Feastnet: Feature-steered graph convolutions for 3d shape analysis. In CVPR 2018-IEEE Conference on Computer Vision & Pattern Recognition, 2018. 2
  40. 40.Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, and Raquel Urtasun. Deep parametric continuous convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2589–2597, 2018. 3
  41. 41.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. arXiv preprint arXiv:1801.07829, 2018. 3
  42. 42.Zizhao Wu, Ruyang Shou, Yunhai Wang, and Xinguo Liu. Interactive shape co-segmentation via label propagation. Computers & Graphics, 38:248–254, 2014.
  43. 43.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1912–1920, 2015. 2, 6, 7
  44. 44.Yifan Xu, Tianqi Fan, Mingye Xu, Long Zeng, and Yu Qiao. Spidercnn: Deep learning on point sets with parameterized convolutional filters. arXiv preprint arXiv:1803.11527, 2018. 3, 6, 7
  45. 45.Li Yi, Hao Su, Xingwen Guo, and Leonidas Guibas. Syncspeccnn: Synchronized spectral cnn for 3d shape segmentation. In Computer Vision and Pattern Recognition (CVPR), 2017. 6, 7
  46. 46.Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe. Unsupervised learning of depth and ego-motion from video. In CVPR, volume 2, page 7, 2017.
  47. 47.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. arXiv preprint arXiv:1711.06396, 2017.

Citation

MLA
Wu, W., et al. “PointConv: Deep Convolutional Networks on 3D Point Clouds”. arXiv, 2018, http://arxiv.org/abs/1811.07246v3.
APA
Wu, W., Qi, Z., & Fuxin, L. (2018). PointConv: Deep Convolutional Networks on 3D Point Clouds. arXiv. http://arxiv.org/abs/1811.07246v3
Chicago
Wu, W., Z. Qi, and L. Fuxin. 2018. “PointConv: Deep Convolutional Networks on 3D Point Clouds”. arXiv. http://arxiv.org/abs/1811.07246v3.
Harvard
Wu, W., Qi, Z. and Fuxin, L. (2018) “PointConv: Deep Convolutional Networks on 3D Point Clouds”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1811.07246v3.
Vancouver
1. Wu W, Qi Z, Fuxin L (2018) PointConv: Deep Convolutional Networks on 3D Point Clouds. arXiv

BibTeX

@article{wu2018pointconv,
  title = {PointConv: Deep Convolutional Networks on 3D Point Clouds},
  author = {Wu, Wenxuan and Qi, Zhongang and Fuxin, Li},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1811.07246v3},
  eprint = {1811.07246}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE