Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Voxel-based Set Attention

Voxel-based set attention is a deep learning mechanism used in 3D computer vision to efficiently extract geometric features from point clouds partitioned into discrete spatial voxel grids. Rather than calculating full pairwise self-attention among points in each voxel, which scales quadratically with point density, the mechanism decomposes intra-voxel self-attention into two successive cross-attention operations mediated by a set of learnable latent codes. This latent-space formulation allows the network to process point clusters of arbitrary size in parallel with linear computational complexity relative to point count, while preserving permutation invariance and avoiding point dropout. By integrating voxel-based spatial partitioning with set-to-set attention, it effectively captures local 3D geometric interactions while maintaining computational scalability.

1 item

Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds

Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds

Chenhang He, Ruihuang Li, Shuai Li, Lei Zhang

Why you should read this

Proposes Voxel Set Transformer, a linear-complexity backbone that processes arbitrary-sized 3D point clusters in parallel through latent-code induced cross-attentions to achieve efficient and competitive 3D object detection on large-scale point clouds.

Transformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to compute the self-attention on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space. To solve this issue, existing methods usually compute self-attention locally by grouping the points into clusters of the same size, or perform convolutional self-attention on a discretized representation. However, the former results in stochastic point dropout, while the latter typically has narrow attention fields. In this paper, we propose a novel voxel-based architecture, namely Voxel Set Transformer (VoxSeT), to detect 3D objects from point clouds by means of set-to-set translation. VoxSeT is built upon a voxel-based set attention (VSA) module, which reduces the self-attention in each voxel by two cross-attentions and models features in a hidden space induced by a group of latent codes. With the VSA module, VoxSeT can manage voxelized point clusters with arbitrary size in a wide range, and process them in parallel with linear complexity. The proposed VoxSeT integrates the high performance of transformer with the efficiency of voxel-based model, which can be used as a good alternative to the convolutional and point-based backbones. VoxSeT reports competitive results on the KITTI and Waymo detection benchmarks. The source codes can be found at https://github.com/skyhehe123/VoxSeT.

Added

2026-10-05