PointConv: Deep Convolutional Networks on 3D Point Clouds
Wenxuan WuZhongang QiFuxin Li
Proposes PointConv, an efficient convolution operator that combines continuous weight functions with local density estimation to build scalable, deep neural networks directly on irregular 3D point clouds.
Three-dimensional point clouds generated by sensors like LIDAR are central to autonomous navigation, robotics, and augmented reality. Unlike standard two-dimensional images arranged on uniform grids, point clouds are irregular, unordered, and unevenly sampled. Traditional convolutional neural networks cannot process these sets directly without either discarding fine structural details or converting them into computationally prohibitive 3D volumetric grids.
The article aims to introduce and validate PointConv, a novel continuous convolution operation that runs directly on unordered 3D point sets while accounting for non-uniform sampling density. The researchers also set out to provide an efficient mathematical formulation capable of scaling these operations to deep neural network architectures.
The authors constructed PointConv by treating convolution filters as continuous functions of local relative coordinates, using multi-layer perceptrons to learn spatial weights alongside an inverse density estimation step to compensate for non-uniform sampling. To address the massive memory overhead of generating individual filter weights, the authors changed the summation order, reducing the operation to standard matrix multiplication and simple two-dimensional convolutions. The method was evaluated on synthetic benchmarks for 3D object classification and part segmentation, a real-world dataset of complex indoor scene scans, and a 2D image classification benchmark treated as a point cloud.
The evaluation produced several key findings. First, on real-world indoor scene segmentation, the approach achieved a mean intersection-over-union score of 55.6%, outperforming earlier methods that scored between 30.6% and 43.8%. Second, the efficient reformulation cut memory usage down to roughly 1/64th of the original version, making modern deep architectures feasible. Third, on 3D synthetic benchmarks, the model reached state-of-the-art results, scoring 92.5% accuracy on 40-class object recognition and an 85.7% instance average intersection-over-union on part segmentation. Finally, when tested on 2D image data treated as irregular points, the method matched the 93% accuracy of standard convolutional networks, showing it functions as a true general convolution.
These findings demonstrate that deep neural networks can process raw point cloud data directly with high accuracy and low memory usage, eliminating the need to project data into cumbersome voxel grids. By achieving both translation invariance and point order invariance, this architecture provides a reliable foundation for high-performance 3D spatial perception in autonomous vehicles, mobile robotics, and computer vision systems.
Organizations developing 3D perception pipelines should consider adopting PointConv-based architectures to improve accuracy and efficiency in high-resolution spatial tasks. Future development efforts should focus on integrating this convolution operation into deeper modern frameworks, such as residual and densely connected network backbones, to evaluate potential gains across larger operational pipelines.
While the method shows strong performance across synthetic and indoor datasets, its effectiveness relies on accurate local neighborhood estimation and offline kernel density calculations. Practitioners should test the approach in varying outdoor lighting, heavy sensor noise, and adverse weather conditions before deploying it in safety-critical autonomous platforms.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet establishes the foundational principles of processing unordered 3D point sets directly using permutation-invariant symmetric functions.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, C. Qi et al. (2017). PointNet++ introduces the hierarchical feature learning and local neighborhood querying paradigms that PointConv builds upon for continuous spatial convolutions.
- Paper: Deep Sets, Manzil Zaheer et al. (2017). Deep Sets provides the formal theoretical framework for defining permutation-invariant and permutation-equivariant operations over unordered point sets.
- Paper: Geometric Deep Learning: Going beyond Euclidean data, Michael M. Bronstein et al. (2016). This survey provides essential background on adapting classical convolutional operations and spatial filtering to non-Euclidean geometric data.
- Paper: KPConv: Flexible and Deformable Convolution for Point Clouds, Hugues Thomas et al. (2019). KPConv advances point cloud convolution by introducing rigid and deformable Euclidean kernel points that dynamically adapt to irregular local geometries.
- Paper: RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds, Qingyong Hu et al. (2019). RandLA-Net extends efficient point-based processing to massive outdoor scenes by pairing fast random downsampling with attentive local feature aggregation.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). PV-RCNN combines point-based continuous feature abstraction with 3D voxel convolutions to achieve highly accurate 3D object detection.
- Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). PCT shifts from continuous point convolutions to self-attention mechanisms and offset-attention to capture long-range geometric context in point clouds.
- Paper: Point Transformer, Nico Engel et al. (2020). Point Transformer replaces local convolutional filtering with cross-attention and local feature sorting directly on unstructured 3D points.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). Minkowski Engine generalizes sparse 3D point convolutions to high-dimensional spatio-temporal domains for 4D video perception.
