Built independently by an author, for readers. Read the story and support ChapterPal

keyword

point density positional encoding

Point density positional encoding is a feature representation technique in 3D deep learning that augments standard geometric positional embeddings with measurements of local point cloud density. In attention mechanisms and neural networks that process spatial sensor data, standard positional encodings map the spatial coordinates of points or grid locations to provide relative or absolute spatial awareness. Point density positional encoding extends these spatial coordinates by integrating local density information, enabling self-attention layers to distinguish between dense and sparse regions within a non-uniformly sampled space. By encoding both geometric location and point concentration into a unified embedding, models can more effectively adapt to sensor-induced variations in sampling density across distances and complex spatial geometries.

1 item

Point Density-Aware Voxels for LiDAR 3D Object Detection

Point Density-Aware Voxels for LiDAR 3D Object Detection

Jordan S. K. Hu, Tianshu Kuai, Steven L. Waslander

Why you should read this

Presents an end-to-end two-stage 3D object detection framework that addresses non-uniform LiDAR point distributions by using voxel point centroids, density-aware region-of-interest grid pooling with self-attention, and density-guided confidence refinement to improve detection across varying distances.

LiDAR has become one of the primary 3D object detection sensors in autonomous driving. However, LiDAR's diverging point pattern with increasing distance results in a non-uniform sampled point cloud ill-suited to discretized volumetric feature extraction. Current methods either rely on voxelized point clouds or use inefficient farthest point sampling to mitigate detrimental effects caused by density variation but largely ignore point density as a feature and its predictable relationship with distance from the LiDAR sensor. Our proposed solution, Point Density-Aware Voxel network (PDV), is an end-to-end two stage LiDAR 3D object detection architecture that is designed to account for these point density variations. PDV efficiently localizes voxel features from the 3D sparse convolution backbone through voxel point centroids. The spatially localized voxel features are then aggregated through a density-aware RoI grid pooling module using kernel density estimation (KDE) and self-attention with point density positional encoding. Finally, we exploit LiDAR's point density to distance relationship to refine our final bounding box confidences. PDV outperforms all state-of-the-art methods on the Waymo Open Dataset and achieves competitive results on the KITTI dataset.

Added

2026-10-05