PointGPT: Auto-regressively Generative Pre-training from Point Clouds
Guangyan ChenMeiling WangYi YangKai YuLi YuanYufeng Yue
Proposes an autoregressive generative pre-training framework for 3D point clouds that arranges patches using Morton ordering and applies a dual masking strategy to prevent shape leakage, achieving state-of-the-art representation learning performance across standard benchmarks.
Three-dimensional point clouds are essential data structures for spatial intelligence applications such as robotics and autonomous driving. However, training effective point cloud models has historically required labor-intensive manual data annotation. While generative pre-training methods have achieved significant breakthroughs in natural language processing and 2D computer vision by learning directly from unlabeled data, adapting these techniques to 3D point clouds is hindered by fundamental data differences: point clouds lack a natural sequence order, contain heavy spatial redundancy, and suffer from task misalignment between fine-grained coordinate prediction and high-level semantic reasoning.
The article introduces and evaluates PointGPT, a self-supervised generative pre-training framework that adapts the auto-regressive pre-training concept to 3D point clouds. The primary objective is to demonstrate that an auto-regressive transformer can learn high-quality 3D geometric representations without relying on external teacher models, 2D images, or natural language inputs.
To achieve this, the approach partitions raw point clouds into localized patches and sequences them geometrically along a spatial space-filling curve, preserving local geometric structures without leaking overall global shape. The framework processes these sequences through an extractor-generator transformer decoder utilizing a dual masking strategy, which strategically masks preceding tokens to eliminate redundancy and force the network to understand holistic shapes. The model pre-trains by auto-regressively predicting subsequent patches using a combined coordinate distance loss. To scale model capacity, the evaluation incorporates an unlabeled hybrid pre-training dataset of approximately 300,000 point clouds, followed by an intermediate supervised alignment stage using a labeled hybrid dataset of roughly 200,000 point clouds across 87 categories.
The experimental findings show that PointGPT consistently outperforms existing 3D self-supervised and fully supervised models. On the real-world ScanObjectNN benchmark, PointGPT achieves state-of-the-art classification accuracy of 93.4% on the hardest setting, surpassing comparable transformer baselines and outperforming multi-modal teacher methods by at least 1.8%. On clean 3D CAD data from ModelNet40, the scaled model achieves a top accuracy of 94.9%. Furthermore, the framework sets new state-of-the-art benchmarks across all standard few-shot learning scenarios—especially in 10-shot tests—and achieves a leading 86.6% instance segmentation accuracy on ShapeNetPart, while ablation tests confirm that Morton curve ordering, dual masking, and relative direction prompts are all critical contributors to performance.
These results demonstrate that self-supervised auto-regressive generation is highly viable for 3D spatial data, offering substantial operational advantages. By eliminating the dependence on cross-modal teachers and labor-intensive annotations, the method simplifies the 3D development pipeline and reduces training complexity. Moreover, its strong few-shot learning performance implies that organizations can reliably deploy 3D vision systems in new target environments with minimal labeled training samples.
Organizations developing 3D perception pipelines should consider adopting pure 3D generative pre-training and Morton-based patching architectures over conventional masked autoencoders to prevent structural information leakage. Teams should also utilize intermediate multi-dataset alignment stages when scaling up model capacities to combat overfitting on small target datasets. Future engineering and research efforts should prioritize expanding 3D dataset curation to narrow the scale gap between 3D vision and language models.
Confidence in these findings is high across standard object-level classification and segmentation benchmarks, supported by systematic ablation analyses. However, readers should note that current 3D pre-training dataset scales remain several orders of magnitude smaller than those in natural language and 2D vision. Consequently, caution is advised before extrapolating these results directly to unbounded, complex outdoor autonomous driving scenes without further pilot validation.
- Paper: Generative Pretraining From Pixels, Mark Chen et al. (2020). It establishes the concept of auto-regressive generative pre-training for visual sequence data (Image GPT) that PointGPT directly translates to 3D point cloud representations.
- Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). It introduces transformer architectures tailored specifically for unstructured 3D point clouds, serving as foundational architectural background for PointGPT's transformer decoder.
- Paper: Point Transformer, Nico Engel et al. (2020). It demonstrates how self-attention mechanisms can capture local and global geometric relationships directly on point sets without voxelization or 2D projections.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). It provides the foundational hierarchical grouping and farthest point sampling techniques universally used to partition point clouds into localized patches.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). It introduces the fundamental deep learning framework for learning permutation-invariant features directly from raw 3D coordinate sets.
- Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). It establishes deep autoencoder representations and distance-based geometric loss formulations like Chamfer Distance for point cloud generation and reconstruction.
- Paper: Dynamic Graph CNN for Learning on Point Clouds, Yue Wang et al. (2018). It develops local neighborhood feature aggregation (EdgeConv) on point sets, which is central to building patch-level geometric representations in modern 3D architectures.
- Paper: Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training, Xiaoyang Wu et al. (2024). It addresses the multi-dataset alignment and domain transfer challenges highlighted by PointGPT's intermediate hybrid pre-training approach by introducing prompt-based multi-dataset training.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). It extends generative 3D modeling from isolated object-level point clouds to expansive, large-scale outdoor driving scenes using hierarchical representations.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It scales feed-forward 3D transformer architectures to predict dense 3D point geometry and scene structures directly from input views in a generalized foundation framework.
