keyword
sparse convolution
Sparse convolution is a specialized neural network operation designed to process spatially sparse data, such as 3D point clouds and voxelized spatial grids, where the majority of locations are empty. Unlike standard dense convolutions that apply kernel computations uniformly across every coordinate in a regular grid, sparse convolution restricts feature storage and mathematical operations strictly to occupied or active spatial sites. By skipping computations on empty regions and maintaining hash tables or coordinate lists of non-zero elements, this operation prevents the cubic scaling of memory and computation typically associated with dense multi-dimensional grids. It enables deep neural architectures to efficiently extract geometric features from large-scale, irregular structures without diluting the inherent spatial sparsity of the input data.
3 items

Fast Point Transformer
Chunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik Park
Why you should read this
Proposes a lightweight local self-attention architecture with centroid-aware voxelization and hashing that encodes continuous 3D coordinates to process large-scale point clouds over a hundred times faster than standard point transformers.
The recent success of neural networks enables a better interpretation of 3D point clouds, but processing a large-scale 3D scene remains a challenging problem. Most current approaches divide a large-scale scene into small regions and combine the local predictions together. However, this scheme inevitably involves additional stages for pre- and post-processing and may also degrade the final output due to predictions in a local perspective. This paper introduces Fast Point Transformer that consists of a new lightweight self-attention layer. Our approach encodes continuous 3D coordinates, and the voxel hashing-based architecture boosts computational efficiency. The proposed method is demonstrated with 3D semantic segmentation and 3D detection. The accuracy of our approach is competitive to the best voxel-based method, and our network achieves 129 times faster inference time than the state-of-the-art, Point Transformer, with a reasonable accuracy trade-off in 3D semantic segmentation on S3DIS dataset.
Added
2026-09-26

SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
Pin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng, Xiangxuan Ren, Bailan Feng, Chao Ma
Why you should read this
Proposes an efficient vision-based 3D semantic occupancy prediction network that replaces dense volume processing with a purely sparse latent representation, cutting computation by up to 74.9% while improving accuracy by preventing feature hallucinations in empty space.
Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic time and space complexity, which limits its scalability in terms of perception range or spatial resolution. Existing approaches compress the dense representation using projections like Bird’s Eye View (BEV) or Tri-Perspective View (TPV). Although efficient, these projections result in information loss, especially for tasks like semantic occupancy prediction. To address this, we propose SparseOcc, an efficient occupancy network inspired by sparse point cloud processing. It utilizes a lossless sparse latent representation with three key innovations. Firstly, a 3D sparse diffuser performs latent completion using spatially decomposed 3D sparse convolutional kernels. Secondly, a feature pyramid and sparse interpolation enhance scales with information from others. Finally, the transformer head is redesigned as a sparse variant. SparseOcc achieves a remarkable 74.9% reduction on FLOPs over the dense baseline. Interestingly, it also improves accuracy, from 12.8% to 14.1% mIOU, which in part can be attributed to the sparse representation’s ability to avoid hallucinations on empty voxels.
Added
2026-09-26

GD-MAE: Generative Decoder for MAE Pre-Training on LiDAR Point Clouds
Honghui Yang, Tong He, Jiaheng Liu, Hua Chen, Boxi Wu, Binbin Lin, Xiaofei He, Wanli Ouyang
Why you should read this
Proposes a generative decoder framework for masked autoencoding on large-scale LiDAR point clouds that eliminates complex decoder designs, reduces pre-training latency by over 88%, and achieves state-of-the-art 3D detection performance using only a fraction of labeled data.
Despite the tremendous progress of Masked Autoencoders (MAE) in developing vision tasks such as image and video, exploring MAE in large-scale 3D point clouds remains challenging due to the inherent irregularity. In contrast to previous 3D MAE frameworks, which either design a complex decoder to infer masked information from maintained regions or adopt sophisticated masking strategies, we instead propose a much simpler paradigm. The core idea is to apply a Generative Decoder for MAE (GD-MAE) to automatically merges the surrounding context to restore the masked geometric knowledge in a hierarchical fusion manner. In doing so, our approach is free from introducing the heuristic design of decoders and enjoys the flexibility of exploring various masking strategies. The corresponding part costs less than 12% latency compared with conventional methods, while achieving better performance. We demonstrate the efficacy of the proposed method on several large-scale benchmarks: Waymo, KITTI, and ONCE. Consistent improvement on downstream detection tasks illustrates strong robustness and generalization capability. Not only our method reveals state-of-the-art results, but remarkably, we achieve comparable accuracy even with 20% of the labeled data on the Waymo dataset. Code will be released.
Added
2026-09-26
