Fast Point Transformer
Chunghyun ParkYoonwoo JeongMinsu ChoJaesik Park
Proposes a lightweight local self-attention architecture with centroid-aware voxelization and hashing that encodes continuous 3D coordinates to process large-scale point clouds over a hundred times faster than standard point transformers.
Interpreting large-scale 3D sensor data quickly and accurately is essential for emerging technologies such as autonomous robotics, augmented reality, and intelligent spatial agents. Traditional deep learning frameworks face severe bottlenecks: point-based models achieve high accuracy but require slow, computationally heavy neighbor searches and multi-stage stitching, whereas voxel-based models operate rapidly on regular spatial grids but lose fine geometric details due to quantization errors. Developing a 3D vision pipeline that balances real-time processing speed with fine-grained accuracy has remained a persistent industry challenge.
The article demonstrates the effectiveness of Fast Point Transformer, a lightweight 3D deep learning architecture designed to process large point clouds rapidly while preserving continuous geometric coordinate information. The researchers evaluate this approach across standard 3D semantic segmentation and 3D object detection benchmarks (S3DIS and ScanNet), assessing inference latency, segmentation accuracy, detection precision, and geometric consistency under spatial rotations and translations.
The proposed method bridges the gap between point and voxel representations through three core techniques: centroid-aware voxelization and devoxelization that encode continuous relative positions to prevent quantization loss, a decomposed lightweight self-attention layer using cosine similarity to minimize memory complexity, and an underlying voxel hashing structure enabling single-shot full-scene inference without expensive neighbor searches.
The experimental findings show substantial operational gains. First, Fast Point Transformer achieves an inference speed of 0.14 seconds per scene on the S3DIS dataset, running 129 times faster than the state-of-the-art Point Transformer baseline and at least 83 times faster than other standard point-based models. Second, it delivers competitive accuracy, achieving a 68.5% mean intersection-over-union on S3DIS (increasing to 70.1% with rotation averaging) and outperforming the leading voxel baseline, MinkowskiNet42. Third, it exhibits exceptional model compactness and resilience; reducing network parameters by up to 71.5% causes negligible accuracy loss (under 0.3 percentage points), unlike voxel convolutional networks that degrade sharply. Fourth, when integrated into 3D object detection frameworks on ScanNet, it outperforms previous point and voxel backbones, raising mean average precision scores significantly. Finally, evaluation via a novel consistency metric confirms the model maintains highly stable predictions regardless of rigid rotations and translations.
These results demonstrate that organizations deploying 3D computer vision do not need to choose between slow, high-accuracy point transformers and fast, error-prone voxel convolutions. Fast Point Transformer lowers computational and memory hardware overheads, unlocking viable single-shot, real-time 3D perception for latency-critical and edge-device applications.
Teams implementing 3D spatial systems should consider adopting lightweight self-attention and centroid-aware hashing mechanisms to reduce inference infrastructure costs. For future development, the authors note that exploring non-convolutional, native transformer network architectures could yield additional accuracy improvements at very fine voxel resolutions.
- Paper: Point Transformer, Nico Engel et al. (2020). Point Transformer introduced attention mechanisms tailored directly for unordered 3D point sets, establishing the foundational architecture and high computational bottlenecks that Fast Point Transformer specifically redesigns for speed.
- Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). This paper establishes key principles of attention and offset-attention architectures on irregular point sets, providing direct context for lightweight transformer design on 3D data.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). It introduces hash-table-based submanifold sparse convolutions on 3D grids, forming the foundational voxel processing structure that Fast Point Transformer combines with self-attention.
- Paper: 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, Christopher Choy et al. (2019). This work develops the Minkowski sparse convolutional backbone (MinkowskiNet), which serves as a primary voxel baseline outpaced by Fast Point Transformer.
- Paper: PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection, Shaoshuai Shi et al. (2019). PV-RCNN pioneered hybrid point-voxel feature integration to reconcile point precision with voxel speed, motivating the centroid-aware voxelization strategies explored in the source.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet is the seminal deep learning architecture for directly processing unordered 3D point sets with permutation invariance, establishing the core problem space.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). PointNet++ establishes hierarchical local feature aggregation across point neighborhoods, highlighting the expensive neighbor query bottlenecks that Fast Point Transformer resolves.
- Paper: RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds, Qingyong Hu et al. (2019). RandLA-Net analyzes computational bottlenecks in large-scale point cloud processing and proposes efficient local spatial encoding schemes.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). Superpoint Transformer builds upon efficient 3D transformer architectures to perform end-to-end 3D scene instance segmentation without relying on expensive dense grouping operations.
- Paper: SCTN: Sparse Convolution-Transformer Network for Scene Flow Estimation, Bing Li et al. (2022). SCTN extends hybrid sparse convolution and point transformer designs to dynamic temporal 3D tasks such as scene flow estimation.
- Paper: PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies, Guocheng Qian et al. (2022). PointNeXt investigates modern scaling and optimization recipes across standard 3D point cloud benchmarks like S3DIS and ScanNet.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). DUSt3R extends transformer-based 3D coordinate regression into dense multi-view geometric reconstruction without explicit camera calibration.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT scales transformer-driven 3D geometric prediction to end-to-end full scene understanding and feed-forward dense point tracking.
