Built independently by an author, for readers. Read the story and support ChapterPal

keyword

OctPEG modules

OctPEG modules, or octree-based positional encoding generator modules, are neural network components designed to dynamically generate spatial position representations for octree-structured 3D point cloud data. Primarily utilized within Transformer architectures for point cloud compression and geometry modeling, these modules serve as an alternative to traditional absolute positional encodings. Instead of embedding fixed coordinate information that can introduce rigid biases, an OctPEG module derives positional features conditioned on the octree sequence, thereby improving translation invariance and enabling the model to learn localized spatial context more effectively across hierarchical tree representations.

1 item

OctFormer: Efficient Octree-Based Transformer for Point Cloud Compression with Local Enhancement

OctFormer: Efficient Octree-Based Transformer for Point Cloud Compression with Local Enhancement

Mingyue Cui, Junhua Long, Mingjian Feng, Boyang Li, Kai Huang

OrganizationsSun Yat-sen University

Why you should read this

Proposes an octree-based Transformer entropy model using non-overlapped context windows and local feature enhancement to significantly speed up point cloud compression while achieving superior compression ratios compared to voxel-based and attention-based baselines.

Point cloud compression with a higher compression ratio and tiny loss is essential for efficient data transportation. However, previous methods that depend on 3D convolution or frequent multi-head self-attention operations bring huge computations. To address this problem, we propose an octree-based Transformer compression method called OctFormer, which does not rely on the occupancy information of sibling nodes. Our method uses non-overlapped context windows to construct octree node sequences and share the result of a multi-head self-attention operation among a sequence of nodes. Besides, we introduce a locally-enhance module for exploiting the sibling features and a positional encoding generator for enhancing the translation invariance of the octree node sequence. Compared to the previous state-of-the-art works, our method obtains up to 17% Bpp savings compared to the voxel-context-based baseline and saves an overall 99% coding time compared to the attention-based baseline.

Added

2026-09-26