Fast Point Transformer

Chunghyun ParkYoonwoo JeongMinsu ChoJaesik Park

article2022CVPR252 citations

Proposes a lightweight local self-attention architecture with centroid-aware voxelization and hashing that encodes continuous 3D coordinates to process large-scale point clouds over a hundred times faster than standard point transformers.

Listen

Interpreting large-scale 3D sensor data quickly and accurately is essential for emerging technologies such as autonomous robotics, augmented reality, and intelligent spatial agents. Traditional deep learning frameworks face severe bottlenecks: point-based models achieve high accuracy but require slow, computationally heavy neighbor searches and multi-stage stitching, whereas voxel-based models operate rapidly on regular spatial grids but lose fine geometric details due to quantization errors. Developing a 3D vision pipeline that balances real-time processing speed with fine-grained accuracy has remained a persistent industry challenge.

The article demonstrates the effectiveness of Fast Point Transformer, a lightweight 3D deep learning architecture designed to process large point clouds rapidly while preserving continuous geometric coordinate information. The researchers evaluate this approach across standard 3D semantic segmentation and 3D object detection benchmarks (S3DIS and ScanNet), assessing inference latency, segmentation accuracy, detection precision, and geometric consistency under spatial rotations and translations.

The proposed method bridges the gap between point and voxel representations through three core techniques: centroid-aware voxelization and devoxelization that encode continuous relative positions to prevent quantization loss, a decomposed lightweight self-attention layer using cosine similarity to minimize memory complexity, and an underlying voxel hashing structure enabling single-shot full-scene inference without expensive neighbor searches.

The experimental findings show substantial operational gains. First, Fast Point Transformer achieves an inference speed of 0.14 seconds per scene on the S3DIS dataset, running 129 times faster than the state-of-the-art Point Transformer baseline and at least 83 times faster than other standard point-based models. Second, it delivers competitive accuracy, achieving a 68.5% mean intersection-over-union on S3DIS (increasing to 70.1% with rotation averaging) and outperforming the leading voxel baseline, MinkowskiNet42. Third, it exhibits exceptional model compactness and resilience; reducing network parameters by up to 71.5% causes negligible accuracy loss (under 0.3 percentage points), unlike voxel convolutional networks that degrade sharply. Fourth, when integrated into 3D object detection frameworks on ScanNet, it outperforms previous point and voxel backbones, raising mean average precision scores significantly. Finally, evaluation via a novel consistency metric confirms the model maintains highly stable predictions regardless of rigid rotations and translations.

These results demonstrate that organizations deploying 3D computer vision do not need to choose between slow, high-accuracy point transformers and fast, error-prone voxel convolutions. Fast Point Transformer lowers computational and memory hardware overheads, unlocking viable single-shot, real-time 3D perception for latency-critical and edge-device applications.

Teams implementing 3D spatial systems should consider adopting lightweight self-attention and centroid-aware hashing mechanisms to reduce inference infrastructure costs. For future development, the authors note that exploring non-convolutional, native transformer network architectures could yield additional accuracy improvements at very fine voxel resolutions.

arXiv: 2112.04702
Cover for Fast Point Transformer

Abstract

The recent success of neural networks enables a better interpretation of 3D point clouds, but processing a large-scale 3D scene remains a challenging problem. Most current approaches divide a large-scale scene into small regions and combine the local predictions together. However, this scheme inevitably involves additional stages for pre- and post-processing and may also degrade the final output due to predictions in a local perspective. This paper introduces Fast Point Transformer that consists of a new lightweight self-attention layer. Our approach encodes continuous 3D coordinates, and the voxel hashing-based architecture boosts computational efficiency. The proposed method is demonstrated with 3D semantic segmentation and 3D detection. The accuracy of our approach is competitive to the best voxel-based method, and our network achieves 129 times faster inference time than the state-of-the-art, Point Transformer, with a reasonable accuracy trade-off in 3D semantic segmentation on S3DIS dataset.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Fast Point Transformer
  • 3.1. Overview
  • 3.2. Centroid-aware Voxel & Devoxelization
  • 3.3. Lightweight Self-Attention
  • 3.4. Network Architecture
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Baselines
  • 4.3. Consistency Test
  • 4.4. 3D Semantic Segmentation
  • 4.5. 3D Object Detection
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Fast Point Transformer Framework for Large-Scale 3D Scene Understanding

    model/method

    Fast Point Transformer is a 3D scene understanding architecture designed to process large-scale point clouds efficiently in a single-shot feed-forward pass without dividing scenes into small cropped blocks or performing expensive kk-nearest neighbor searches. The architecture operates in three successive stages:

    1. Centroid-Aware Voxelization: The continuous input point cloud Pin={(pn,in)}n=1N\mathcal{P}^{\text{in}} = \{(\mathbf{p}_n, \mathbf{i}_n)\}_{n=1}^N (with 3D coordinates pn\mathbf{p}_n and input features in\mathbf{i}_n) is converted into non-empty voxel representations V={(vi,fi,ci)}i=1I\mathcal{V} = \{(\mathbf{v}_i, \mathbf{f}_i, \mathbf{c}_i)\}_{i=1}^I, where vi∈Z3\mathbf{v}_i \in \mathbb{Z}^3 is the discrete voxel coordinate, ci∈R3\mathbf{c}_i \in \mathbb{R}^3 is the continuous centroid coordinate of the points inside voxel ii, and fi\mathbf{f}_i is the aggregated voxel feature incorporating continuous point-to-centroid positional encodings.

    2. Lightweight Self-Attention (LSA) Backbone: Local self-attention layers update voxel features fi\mathbf{f}_i to fi′\mathbf{f}'_i. Spatial neighbor queries are executed via voxel hashing with O(1)\mathcal{O}(1) time complexity per query, replacing O(Nlog⁡N)\mathcal{O}(N \log N) point-level kk-NN search. The network adopts a hierarchical U-Net architecture by integrating LSA layers with voxel downsampling and upsampling operations.

    3. Centroid-Aware Devoxelization: The output voxel features V′={(vi,fi′,ci)}i=1I\mathcal{V}' = \{(\mathbf{v}_i, \mathbf{f}'_i, \mathbf{c}_i)\}_{i=1}^I are projected back to the original continuous points to produce Pout={(pn,on)}n=1N\mathcal{P}^{\text{out}} = \{(\mathbf{p}_n, \mathbf{o}_n)\}_{n=1}^N, utilizing the continuous centroid-to-point positional encodings to preserve fine geometric detail.

  2. Knowl 2 — Centroid-Aware Voxelization and Devoxelization

    model/method

    To preserve fine continuous geometric structures and prevent quantization artifacts during voxel grid operations, Fast Point Transformer uses centroid-aware voxelization and devoxelization.

    For an input point cloud Pin={(pn,in)}n=1N\mathcal{P}^{\text{in}} = \{(\mathbf{p}_n, \mathbf{i}_n)\}_{n=1}^N, where pn∈R3\mathbf{p}_n \in \mathbb{R}^3 is the continuous point coordinate and in∈RDin\mathbf{i}_n \in \mathbb{R}^{D_{\text{in}}} is the raw input feature (e.g., RGB color), points are partitioned into non-empty voxels i∈{1,…,I}i \in \{1, \dots, I\}. Let M(i)\mathcal{M}(i) be the set of point indices mapped to voxel ii, and μ(n)\mu(n) denote the index of the voxel containing point nn. The continuous voxel centroid ci∈R3\mathbf{c}_i \in \mathbb{R}^3 is computed as: ci=1∣M(i)∣∑n∈M(i)pn\mathbf{c}_i = \frac{1}{|\mathcal{M}(i)|} \sum_{n \in \mathcal{M}(i)} \mathbf{p}_n

    A learnable continuous centroid-to-point positional encoding en∈RDenc\mathbf{e}_n \in \mathbb{R}^{D_{\text{enc}}} is computed using an encoding layer δenc:R3→RDenc\delta_{\text{enc}} : \mathbb{R}^3 \to \mathbb{R}^{D_{\text{enc}}}: en=δenc(pn−cμ(n))\mathbf{e}_n = \delta_{\text{enc}}(\mathbf{p}_n - \mathbf{c}_{\mu(n)})

    The voxel feature fi∈RDin+Denc\mathbf{f}_i \in \mathbb{R}^{D_{\text{in}} + D_{\text{enc}}} is formed by concatenating point features with positional encodings and aggregating over intra-voxel points using a permutation-invariant operator Ω\Omega (e.g., average pooling): fi=Ωn∈M(i)(in⊕en)\mathbf{f}_i = \Omega_{n \in \mathcal{M}(i)}(\mathbf{i}_n \oplus \mathbf{e}_n) where ⊕\oplus denotes vector concatenation.

    During devoxelization, given updated voxel features fi′∈RDout\mathbf{f}'_i \in \mathbb{R}^{D_{\text{out}}}, the continuous output point features on∈RDout\mathbf{o}_n \in \mathbb{R}^{D_{\text{out}}} are reconstructed via a multilayer perceptron MLP:RDout+Denc→RDout\text{MLP}: \mathbb{R}^{D_{\text{out}} + D_{\text{enc}}} \to \mathbb{R}^{D_{\text{out}}}: on=MLP(fμ(n)′⊕en)\mathbf{o}_n = \text{MLP}(\mathbf{f}'_{\mu(n)} \oplus \mathbf{e}_n)

  3. Knowl 3 — Coordinate Decomposition for Space Complexity Reduction in Continuous Positional Encoding

    theoretical result

    In local self-attention over continuous voxel centroids ci,cj∈R3\mathbf{c}_i, \mathbf{c}_j \in \mathbb{R}^3, directly computing continuous relative positional encodings δ(ci,cj)=MLP(ci−cj)\delta(\mathbf{c}_i, \mathbf{c}_j) = \text{MLP}(\mathbf{c}_i - \mathbf{c}_j) requires O(IKD)\mathcal{O}(IKD) space complexity, where II is the number of voxels, KK is the number of local neighbors per voxel, and DD is the encoding dimension. This overhead arises because continuous centroids produce up to O(IK)\mathcal{O}(IK) distinct pairwise relative displacement vectors (ci−cj)(\mathbf{c}_i - \mathbf{c}_j).

    Fast Point Transformer decomposes the relative centroid displacement using discrete voxel grid coordinates vi,vj∈Z3\mathbf{v}_i, \mathbf{v}_j \in \mathbb{Z}^3: ci−cj=(ci−vi)−(cj−vj)+(vi−vj)\mathbf{c}_i - \mathbf{c}_j = (\mathbf{c}_i - \mathbf{v}_i) - (\mathbf{c}_j - \mathbf{v}_j) + (\mathbf{v}_i - \mathbf{v}_j)

    This decomposition enables splitting the positional encoding into:

    1. A continuous intra-voxel offset δabs(ci−vi)∈RD\delta_{\text{abs}}(\mathbf{c}_i - \mathbf{v}_i) \in \mathbb{R}^D stored per voxel, requiring O(ID)\mathcal{O}(ID) space.
    2. A discretized inter-voxel relative displacement δrel(vi−vj)∈RD\delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j) \in \mathbb{R}^D, requiring O(KD)\mathcal{O}(KD) space because there are at most KK distinct discrete relative neighbor offsets in Z3\mathbb{Z}^3.

    Because δabs(cj−vj)\delta_{\text{abs}}(\mathbf{c}_j - \mathbf{v}_j) is reused across queries, the total space complexity for continuous relative positional encoding is reduced from O(IKD)\mathcal{O}(IKD) to O(ID+KD)\mathcal{O}(ID + KD).

  4. Knowl 4 — Lightweight Self-Attention (LSA) Layer Formulation

    equation

    The Lightweight Self-Attention (LSA) layer in Fast Point Transformer updates the feature representation of voxel centroid ii over its local neighbor set N(i)\mathcal{N}(i) using cosine similarity between centroid-aware query features and discretized relative positional encodings: fi′=∑j∈N(i)ϕ(gi)⋅δrel(vi−vj)∥ϕ(gi)∥∥δrel(vi−vj)∥ψ(gj)\mathbf{f}'_i = \sum_{j \in \mathcal{N}(i)} \frac{\phi(\mathbf{g}_i) \cdot \delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)}{\|\phi(\mathbf{g}_i)\|\|\delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)\|} \psi(\mathbf{g}_j) where:

    • gi=fi+δabs(ci−vi)\mathbf{g}_i = \mathbf{f}_i + \delta_{\text{abs}}(\mathbf{c}_i - \mathbf{v}_i) is the centroid-aware voxel feature, combining the voxel feature fi\mathbf{f}_i with the continuous intra-voxel positional encoding δabs(ci−vi)\delta_{\text{abs}}(\mathbf{c}_i - \mathbf{v}_i).
    • vi,vj∈Z3\mathbf{v}_i, \mathbf{v}_j \in \mathbb{Z}^3 are discrete voxel coordinates.
    • δrel(vi−vj)∈RD\delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j) \in \mathbb{R}^D is the discretized relative positional encoding.
    • ϕ(⋅)\phi(\cdot) and ψ(⋅)\psi(\cdot) denote linear projection functions for queries and values, respectively.
    • N(i)\mathcal{N}(i) denotes the set of local voxel neighbor indices queried via voxel hashing.

    Cosine similarity is utilized instead of softmax normalization softmax(ϕ(gi)⊤δrel(vi−vj))\text{softmax}(\phi(\mathbf{g}_i)^\top \delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)) because 3D point cloud voxel grids have variable neighborhood sizes ∣N(i)∣|\mathcal{N}(i)|. In sparse boundary regions where ∣N(i)∣=1|\mathcal{N}(i)| = 1, softmax normalizes the weight to 1.0 (reducing the attention mechanism to a static linear layer), whereas cosine similarity provides dynamic, magnitude-normalized weighting across varying voxel densities.

  5. Knowl 5 — Consistency Score (CScore) Metric for Rigid Transformation Invariance

    definition

    The Consistency Score (CScore) measures the semantic prediction coherence of a 3D point cloud model under rigid geometric transformations (rotations and translations). Given a set of point clouds S={Pin}\mathcal{S} = \{\mathcal{P}^{\text{in}}\}, a 3D semantic segmentation model f:Pin→Cf: \mathcal{P}^{\text{in}} \to \mathcal{C} mapping input points (pn,in)(\mathbf{p}_n, \mathbf{i}_n) to semantic classes C\mathcal{C}, and a set of rigid transformations T={Tm}\mathcal{T} = \{\mathbf{T}_m\}, the consistency score CScore(f;S,T)\text{CScore}(f; \mathcal{S}, \mathcal{T}) is defined as: CScore(f;S,T)=1∣S∣∑Pin∈S1∣Pin∣∣T∣∑n=1∣Pin∣∑m=1∣T∣I(f(pn,in),f(Tmpn,in))\text{CScore}(f; \mathcal{S}, \mathcal{T}) = \frac{1}{|\mathcal{S}|} \sum_{\mathcal{P}^{\text{in}} \in \mathcal{S}} \frac{1}{|\mathcal{P}^{\text{in}}| |\mathcal{T}|} \sum_{n=1}^{|\mathcal{P}^{\text{in}}|} \sum_{m=1}^{|\mathcal{T}|} \mathbb{I}\Big(f(\mathbf{p}_n, \mathbf{i}_n), f(\mathbf{T}_m \mathbf{p}_n, \mathbf{i}_n)\Big) where I(⋅,⋅)\mathbb{I}(\cdot, \cdot) is the indicator function returning 1 if the predicted class of the unperturbed point equals that of the transformed point, and 0 otherwise.

    In standard evaluation benchmarks on ScanNet validation with voxel size LL, T\mathcal{T} comprises 41 rigid transformations:

    • 26 3D translations: {0,L/3,2L/3}3∖{[0,0,0]}\{0, L/3, 2L/3\}^3 \setminus \{[0, 0, 0]\}.
    • 15 rotations around the gravity axis: {0.125π,0.25π,…,1.875π}\{0.125\pi, 0.25\pi, \dots, 1.875\pi\}.
  6. Knowl 6 — Inference Latency and Semantic Segmentation Performance on S3DIS Area 5

    data/table

    On the S3DIS dataset Area 5 test split, Fast Point Transformer achieves competitive accuracy to point-based transformer architectures while operating at single-shot voxel speeds, running 129 times faster than Point Transformer.

    Method Neighbor Prep. Neighbor Infer. Inference Style Latency (s) mAcc (%) mIoU (%)
    PointNet None None Crop-and-stitch 18.16 49.0 41.1
    SPGraph None None Crop-and-stitch 18.28 66.5 58.0
    PointWeb O(1)\mathcal{O}(1) O(MNK)\mathcal{O}(MNK) Crop-and-stitch 11.62 66.6 60.3
    MinkowskiNet32 (5cm) O(N)\mathcal{O}(N) O(M)\mathcal{O}(M) Single-shot 0.08 71.7 65.4
    KPConv deform O(Nlog⁡N)\mathcal{O}(N \log N) O(KMlog⁡N)\mathcal{O}(KM \log N) Multi-shot 105.15 72.8 67.1
    PAConv O(1)\mathcal{O}(1) O(MNlog⁡K)\mathcal{O}(MN \log K) Crop-and-stitch 28.13 73.0 66.6
    PointTransformer O(1)\mathcal{O}(1) O(MNlog⁡K)\mathcal{O}(MN \log K) Multi-shot 18.07 76.5 70.4
    Voxel size: 4cm
    MinkowskiNet42 O(N)\mathcal{O}(N) O(M)\mathcal{O}(M) Single-shot 0.08 74.4±0.874.4 \pm 0.8 67.1±0.167.1 \pm 0.1
    + rotation average O(N)\mathcal{O}(N) O(M)\mathcal{O}(M) Multi-shot 0.66 75.0±0.775.0 \pm 0.7 68.4±0.168.4 \pm 0.1
    FastPointTransformer (ours) O(N)\mathcal{O}(N) O(M)\mathcal{O}(M) Single-shot 0.14 76.5±0.676.5 \pm 0.6 68.5±0.268.5 \pm 0.2
    + rotation average O(N)\mathcal{O}(N) O(M)\mathcal{O}(M) Multi-shot 1.13 77.3±0.777.3 \pm 0.7 70.1±0.370.1 \pm 0.3

    Here NN is the number of dataset points, MM is the number of query points (or voxel centroids), and KK is the neighborhood size (M,N≫KM, N \gg K). Latency is measured per scene on an Intel Core i7-5930K CPU and an NVIDIA GeForce RTX 3090 GPU.

  7. Knowl 7 — Prediction Consistency and Accuracy Across Voxel Resolutions on ScanNet

    data/table

    Fast Point Transformer demonstrates significantly higher consistency scores (CScore) and semantic segmentation accuracy (mIoU) across different voxel resolutions on the ScanNet validation dataset compared to the baseline voxel network MinkowskiNet42, demonstrating robust suppression of quantization artifacts.

    Method CScore (%) - R CScore (%) - t CScore (%) - R t mIoU (%)
    Voxel size: 10cm
    MinkowskiNet42 92.2±0.192.2 \pm 0.1 92.0±0.192.0 \pm 0.1 92.0±0.192.0 \pm 0.1 60.5±0.260.5 \pm 0.2
    FastPointTransformer 94.7±0.394.7 \pm 0.3 94.6±0.194.6 \pm 0.1 94.6±0.194.6 \pm 0.1 65.9±0.665.9 \pm 0.6
    Voxel size: 5cm
    MinkowskiNet42 94.2±0.194.2 \pm 0.1 95.1±0.195.1 \pm 0.1 94.8±0.194.8 \pm 0.1 66.7±0.266.7 \pm 0.2
    FastPointTransformer 95.9±0.495.9 \pm 0.4 96.4±0.196.4 \pm 0.1 96.2±0.296.2 \pm 0.2 70.0±0.170.0 \pm 0.1
    Voxel size: 2cm
    MinkowskiNet42 95.9±0.695.9 \pm 0.6 96.9±0.396.9 \pm 0.3 96.6±0.196.6 \pm 0.1 71.9±0.271.9 \pm 0.2
    FastPointTransformer 96.9±0.396.9 \pm 0.3 97.4±0.497.4 \pm 0.4 97.2±0.197.2 \pm 0.1 72.1±0.372.1 \pm 0.3

    Transformations evaluated include 15 rotations around the gravity axis (R), 26 grid translations (t), and both combined (R & t). At larger voxel sizes (10cm and 5cm) where quantization artifacts are most severe, Fast Point Transformer outperforms MinkowskiNet42 by +5.4%+5.4\% and +3.3%+3.3\% mIoU, respectively.

  8. Knowl 8 — Model Parameter Scaling and Compression Resilience on ScanNet

    data/table

    When network capacity is reduced by decreasing the number of building blocks, Fast Point Transformer retains its segmentation performance on ScanNet validation, whereas MinkowskiNet degrades significantly.

    Method # Param. (M) Rel. Param. (%) mIoU (%) Δ\Delta mIoU (%)
    Voxel size: 10cm
    MinkowskiNet42 37.9 ±0.0\pm 0.0 60.5±0.260.5 \pm 0.2 ±0.0\pm 0.0
    MinkowskiNet (small) 21.7 −42.7-42.7 59.9±0.659.9 \pm 0.6 −0.6-0.6
    MinkowskiNet (smaller) 11.6 −69.4-69.4 58.2±0.958.2 \pm 0.9 −2.3-2.3
    FastPointTransformer 37.9 ±0.0\pm 0.0 65.9±0.665.9 \pm 0.6 ±0.0\pm 0.0
    FastPointTransformer (small) 20.2 −46.7-46.7 66.0±0.366.0 \pm 0.3 +0.1+0.1
    FastPointTransformer (smaller) 10.8 −71.5-71.5 65.7±0.165.7 \pm 0.1 −0.2-0.2
    Voxel size: 5cm
    MinkowskiNet42 37.9 ±0.0\pm 0.0 66.7±0.366.7 \pm 0.3 ±0.0\pm 0.0
    MinkowskiNet (small) 21.7 −42.7-42.7 66.0±0.166.0 \pm 0.1 −0.7-0.7
    MinkowskiNet (smaller) 11.6 −69.4-69.4 64.2±0.464.2 \pm 0.4 −2.5-2.5
    FastPointTransformer 37.9 ±0.0\pm 0.0 70.0±0.170.0 \pm 0.1 ±0.0\pm 0.0
    FastPointTransformer (small) 20.2 −46.7-46.7 70.3±0.270.3 \pm 0.2 +0.3+0.3
    FastPointTransformer (smaller) 10.8 −71.5-71.5 69.7±0.269.7 \pm 0.2 −0.3-0.3

    Even with a 71.5%71.5\% reduction in parameters (10.8M parameters), Fast Point Transformer maintains 65.7%65.7\% mIoU at 10cm resolution, outperforming the most lightweight MinkowskiNet at 5cm resolution (11.6M parameters, 64.2%64.2\% mIoU).

  9. Knowl 9 — 3D Object Detection Performance with VoteNet on ScanNet

    data/table

    Fast Point Transformer serves as an effective feature backbone for 3D object detection within the VoteNet architecture on the ScanNet validation dataset, outperforming both point-based and voxel-based backbones.

    Backbone [email protected] (%) [email protected] (%)
    PointNet++ 54.2 30.1
    RS-CNN 51.6 29.5
    KPConv 48.9 29.2
    MinkowskiNet 53.8 30.2
    MinkowskiNet (reproduced) 55.3±0.255.3 \pm 0.2 33.0±0.533.0 \pm 0.5
    FastPointTransformer 59.1±0.159.1 \pm 0.1 35.6±0.435.6 \pm 0.4

    Fast Point Transformer improves detection accuracy by +3.8%+3.8\% [email protected] and +2.6%+2.6\% [email protected] over the reproduced MinkowskiNet backbone without any changes to the downstream Hough voting or object bounding box proposal modules.

  10. Knowl 10 — Ablation Analysis of Positional Encoding, Attention Formulation, and Local Window Size

    empirical result

    Ablation studies on the ScanNet validation dataset at 10cm voxel resolution evaluate key design choices in Fast Point Transformer:

    1. Positional Encodings: Starting from a base network without continuous positional encodings (59.1% mIoU, 27.3M parameters), adding centroid-aware voxelization encoding δenc\delta_{\text{enc}} increases performance to 61.3% mIoU. Adding centroid-aware devoxelization increases performance to 62.7% mIoU, and incorporating the continuous intra-voxel feature offset δabs\delta_{\text{abs}} achieves the full performance of 65.3% mIoU (37.9M parameters). In contrast, adding δabs\delta_{\text{abs}} to MinkowskiNet42 fails to improve accuracy (65.1% to 65.0% mIoU), indicating that δabs\delta_{\text{abs}} specifically benefits dynamic self-attention.

    2. Attention Types: Cosine similarity on query and relative position cosine(ϕ(gi),δrel(vi−vj))\text{cosine}(\phi(\mathbf{g}_i), \delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)) achieves 65.3% mIoU, outperforming softmax attention softmax(ϕ(gi),δrel(vi−vj))\text{softmax}(\phi(\mathbf{g}_i), \delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)) at 61.0% mIoU and cosine attention including query-key feature similarity cosine(ϕ(gi),ξ(gj)+δrel(vi−vj))\text{cosine}(\phi(\mathbf{g}_i), \xi(\mathbf{g}_j) + \delta_{\text{rel}}(\mathbf{v}_i - \mathbf{v}_j)) at 62.1% mIoU.

    3. Local Window Size: A local window size of k=3k=3 gives 65.3% mIoU with 0.106s latency. Increasing the local window size degrades performance (k=5k=5: 62.4% mIoU, 0.127s latency; k=7k=7: 61.9% mIoU, 0.168s latency) due to the sparsity of 3D voxel data when utilizing learnable positional tokens.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Dan A Alcantara, Andrei Sharf, Fatemeh Abbasinejad, Shubhabrata Sengupta, Michael Mitzenmacher, John D Owens, and Nina Amenta. Real-time parallel hashing on the gpu. In ACM SIGGRAPH Asia 2009 papers, pages 1–9. 2009. 1
  2. 2.Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In CVPR, pages 1534–1543, 2016. 1, 2, 5, 6, 7, 8
  3. 3.Irwan Bello. Lambdanetworks: Modeling long-range interactions without attention. In ICLR, 2021. 3, 8
  4. 4.Thomas Chaton, Nicolas Chaulet, Sofiane Horache, and Loic Landrieu. Torch-points3d: A modular multi-task framework for reproducible deep learning on 3d point clouds. In 3DV, pages 1–10. IEEE, 2020. 5, 8
  5. 5.Hung-Yueh Chiang, Yen-Liang Lin, Yueh-Cheng Liu, and Winston H Hsu. A unified point-based framework for 3d segmentation. In 3DV, pages 155–163. IEEE, 2019. 5
  6. 6.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In CVPR, pages 3075–3084, 2019. 1, 2, 3, 5, 6, 7, 8
  7. 7.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richlyannotated 3d reconstructions of indoor scenes. In CVPR, pages 5828–5839, 2017. 2, 5, 6, 7, 8
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. 3
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 3
  10. 10.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In CVPR, pages 9224–9232, 2018. 1, 2, 3, 5, 7
  11. 11.Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, and Shi-Min Hu. Pct: Point cloud transformer. Computational Visual Media, 7(2):187–199, Apr 2021. 2, 3
  12. 12.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, NeurIPS, 2021. 3
  13. 13.Lei Han, Tian Zheng, Yinheng Zhu, Lan Xu, and Lu Fang. Live semantic 3d perception for immersive augmented reality. IEEE TVCG, 26(5):2012–2022, 2020. 1
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 5
  15. 15.Wenbo Hu, Hengshuang Zhao, Li Jiang, Jiaya Jia, and Tien-Tsin Wong. Bidirectional projection network for cross dimension scene understanding. In CVPR, pages 14373–14382, June 2021. 5
  16. 16.Abhijit Kundu, Xiaoqi Yin, Alireza Fathi, David Ross, Brian Brewington, Thomas Funkhouser, and Caroline Pantofaru. Virtual multi-view fusion for 3d semantic segmentation. In ECCV, pages 518–535. Springer, 2020. 5
  17. 17.Loic Landrieu and Martin Simonovsky. Large-scale point cloud semantic segmentation with superpoint graphs. In CVPR, pages 4558–4567, 2018. 5, 6, 7
  18. 18.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on χ-transformed points. In NeurIPS, pages 828–838, 2018. 1, 5
  19. 19.Yuchen Li, Qiwei Zhu, Zheng Lyu, Zhongdong Huang, and Jianling Sun. Dycuckoo: dynamic hash tables on gpus. In ICDE, pages 744–755. IEEE, 2021. 1
  20. 20.Yongcheng Liu, Bin Fan, Shiming Xiang, and Chunhong Pan. Relation-shape convolutional neural network for point cloud analysis. In CVPR, pages 8895–8904, 2019. 8
  21. 21.Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Pointvoxel cnn for efficient 3d deep learning. In NeurIPS, 2019. 2
  22. 22.Jiageng Mao, Xiaogang Wang, and Hongsheng Li. Interpolated convolutional networks for 3d point cloud understanding. In ICCV, October 2019. 1, 2
  23. 23.Jiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai, Jiashi Feng, Xiaodan Liang, Hang Xu, and Chunjing Xu. Voxel transformer for 3d object detection. In ICCV, pages 3164–3173, October 2021. 1, 2, 3, 5
  24. 24.Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Marc Stamminger. Real-time 3d reconstruction at scale using voxel hashing. ACM TOG, 32(6):1–11, 2013. 1
  25. 25.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In ICCV, pages 9277–9286, 2019. 2, 5, 8
  26. 26.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, pages 652–660, 2017. 1, 2, 5, 6, 7
  27. 27.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++ deep hierarchical feature learning on point sets in a metric space. In CVPR, pages 5105–5114, 2017. 1, 2, 5, 8
  28. 28.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 3
  29. 29.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. Stand-alone self-attention in vision models. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett, editors, NeurIPS, volume 32. Curran Associates, Inc., 2019. 2, 3, 4, 8
  30. 30.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, pages 234–241. Springer, 2015. 5, 8
  31. 31.Radu Alexandru Rosu, Peer Schütt, Jan Quenzel, and Sven Behnke. Latticenet: Fast point cloud segmentation using permutohedral lattices. In RSS, 2020. 3
  32. 32.Hang Su, Varun Jampani, Deqing Sun, Subhransu Maji, Evangelos Kalogerakis, Ming-Hsuan Yang, and Jan Kautz. Splatnet: Sparse lattice networks for point cloud processing. In CVPR, pages 2530–2539, 2018. 3
  33. 33.Haotian Tang, Zhijian Liu, Shengyu Zhao, Yujun Lin, Ji Lin, Hanrui Wang, and Song Han. Searching efficient 3d architectures with sparse point-voxel convolution. In ECCV, pages 685–702. Springer, 2020. 2
  34. 34.Maxim Tatarchenko, Jaesik Park, Vladlen Koltun, and Qian-Yi Zhou. Tangent convolutions for dense prediction in 3d. In CVPR, pages 3887–3896, 2018. 1, 2, 5
  35. 35.Lyne Tchapmi, Christopher Choy, Iro Armeni, JunYoung Gwak, and Silvio Savarese. Segcloud: Semantic segmentation of 3d point clouds. In 3DV, pages 537–547. IEEE, 2017. 1, 5
  36. 36.Matthias Teschner, Bruno Heidelberger, Matthias Müller, Danat Pomerantes, and Markus H Gross. Optimized spatial hashing for collision detection of deformable objects. In Vmv, volume 3, pages 47–54, 2003. 1
  37. 37.Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J Guibas. Kpconv: Flexible and deformable convolution for point clouds. In ICCV, pages 6411–6420, 2019. 1, 2, 5, 6, 7, 8
  38. 38.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In CVPR, pages 12894–12904, June 2021. 2, 3
  39. 39.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, NeurIPS, volume 30. Curran Associates, Inc., 2017. 3
  40. 40.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In CVPR, pages 9621–9630, 2019. 5, 7
  41. 41.Mutian Xu, Runyu Ding, Hengshuang Zhao, and Xiaojuan Qi. Paconv: Position adaptive convolution with dynamic kernel assembling on point clouds. In CVPR, pages 3173–3182, June 2021. 1, 4, 5, 6, 7
  42. 42.Xu Yan, Chaoda Zheng, Zhen Li, Sheng Wang, and Shuguang Cui. Pointasnl: Robust point clouds processing using nonlocal neural networks with adaptive sampling. In CVPR, pages 5589–5598, 2020. 5, 7
  43. 43.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis E.H. Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In ICCV, pages 558–567, October 2021. 3
  44. 44.Cheng Zhang, Haocheng Wan, Shengqiang Liu, Xinyi Shen, and Zizhao Wu. Pvt: Point-voxel transformer for 3d deep learning. arXiv preprint arXiv:2108.06076, 2021. 2
  45. 45.Feihu Zhang, Jin Fang, Benjamin Wah, and Philip Torr. Deep fusionnet for point cloud semantic segmentation. In ECCV, volume 2, page 6, 2020. 2, 3
  46. 46.Richard Zhang. Making convolutional networks shift-invariant again. In ICML, pages 7324–7334. PMLR, 2019. 2
  47. 47.Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring self-attention for image recognition. In CVPR, June 2020. 4
  48. 48.Hengshuang Zhao, Li Jiang, Chi-Wing Fu, and Jiaya Jia. Pointweb: Enhancing local neighborhood features for point cloud processing. In CVPR, pages 5565–5573, 2019. 5, 6, 7
  49. 49.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H.S. Torr, and Vladlen Koltun. Point transformer. In ICCV, pages 16259–16268, October 2021. 1, 2, 3, 4, 5, 6, 7, 8
  50. 50.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In CVPR, pages 4490–4499, 2018. 1
  51. 51.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable {detr}: Deformable transformers for end-to-end object detection. In ICLR, 2021. 4

Citation

MLA
Park, C., et al. “Fast Point Transformer”. arXiv, 2021, http://arxiv.org/abs/2112.04702v2.
APA
Park, C., Jeong, Y., Cho, M., & Park, J. (2021). Fast Point Transformer. arXiv. http://arxiv.org/abs/2112.04702v2
Chicago
Park, C., Y. Jeong, M. Cho, and J. Park. 2021. “Fast Point Transformer”. arXiv. http://arxiv.org/abs/2112.04702v2.
Harvard
Park, C. et al. (2021) “Fast Point Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.04702v2.
Vancouver
1. Park C, Jeong Y, Cho M, Park J (2021) Fast Point Transformer. arXiv

BibTeX

@article{park2021fast,
  title = {Fast Point Transformer},
  author = {Park, Chunghyun and Jeong, Yoonwoo and Cho, Minsu and Park, Jaesik},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.04702v2},
  eprint = {2112.04702}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE