PointGPT: Auto-regressively Generative Pre-training from Point Clouds

Guangyan ChenMeiling WangYi YangKai YuLi YuanYufeng Yue

article2023NeurIPS219 citations

Proposes an autoregressive generative pre-training framework for 3D point clouds that arranges patches using Morton ordering and applies a dual masking strategy to prevent shape leakage, achieving state-of-the-art representation learning performance across standard benchmarks.

Listen

Three-dimensional point clouds are essential data structures for spatial intelligence applications such as robotics and autonomous driving. However, training effective point cloud models has historically required labor-intensive manual data annotation. While generative pre-training methods have achieved significant breakthroughs in natural language processing and 2D computer vision by learning directly from unlabeled data, adapting these techniques to 3D point clouds is hindered by fundamental data differences: point clouds lack a natural sequence order, contain heavy spatial redundancy, and suffer from task misalignment between fine-grained coordinate prediction and high-level semantic reasoning.

The article introduces and evaluates PointGPT, a self-supervised generative pre-training framework that adapts the auto-regressive pre-training concept to 3D point clouds. The primary objective is to demonstrate that an auto-regressive transformer can learn high-quality 3D geometric representations without relying on external teacher models, 2D images, or natural language inputs.

To achieve this, the approach partitions raw point clouds into localized patches and sequences them geometrically along a spatial space-filling curve, preserving local geometric structures without leaking overall global shape. The framework processes these sequences through an extractor-generator transformer decoder utilizing a dual masking strategy, which strategically masks preceding tokens to eliminate redundancy and force the network to understand holistic shapes. The model pre-trains by auto-regressively predicting subsequent patches using a combined coordinate distance loss. To scale model capacity, the evaluation incorporates an unlabeled hybrid pre-training dataset of approximately 300,000 point clouds, followed by an intermediate supervised alignment stage using a labeled hybrid dataset of roughly 200,000 point clouds across 87 categories.

The experimental findings show that PointGPT consistently outperforms existing 3D self-supervised and fully supervised models. On the real-world ScanObjectNN benchmark, PointGPT achieves state-of-the-art classification accuracy of 93.4% on the hardest setting, surpassing comparable transformer baselines and outperforming multi-modal teacher methods by at least 1.8%. On clean 3D CAD data from ModelNet40, the scaled model achieves a top accuracy of 94.9%. Furthermore, the framework sets new state-of-the-art benchmarks across all standard few-shot learning scenarios—especially in 10-shot tests—and achieves a leading 86.6% instance segmentation accuracy on ShapeNetPart, while ablation tests confirm that Morton curve ordering, dual masking, and relative direction prompts are all critical contributors to performance.

These results demonstrate that self-supervised auto-regressive generation is highly viable for 3D spatial data, offering substantial operational advantages. By eliminating the dependence on cross-modal teachers and labor-intensive annotations, the method simplifies the 3D development pipeline and reduces training complexity. Moreover, its strong few-shot learning performance implies that organizations can reliably deploy 3D vision systems in new target environments with minimal labeled training samples.

Organizations developing 3D perception pipelines should consider adopting pure 3D generative pre-training and Morton-based patching architectures over conventional masked autoencoders to prevent structural information leakage. Teams should also utilize intermediate multi-dataset alignment stages when scaling up model capacities to combat overfitting on small target datasets. Future engineering and research efforts should prioritize expanding 3D dataset curation to narrow the scale gap between 3D vision and language models.

Confidence in these findings is high across standard object-level classification and segmentation benchmarks, supported by systematic ablation analyses. However, readers should note that current 3D pre-training dataset scales remain several orders of magnitude smaller than those in natural language and 2D vision. Consequently, caution is advised before extrapolating these results directly to unbounded, complex outdoor autonomous driving scenes without further pilot validation.

arXiv: 2305.11487
  • Paper: Generative Pretraining From Pixels, Mark Chen et al. (2020). It establishes the concept of auto-regressive generative pre-training for visual sequence data (Image GPT) that PointGPT directly translates to 3D point cloud representations.
  • Paper: PCT: Point cloud transformer, Meng-Hao Guo et al. (2020). It introduces transformer architectures tailored specifically for unstructured 3D point clouds, serving as foundational architectural background for PointGPT's transformer decoder.
  • Paper: Point Transformer, Nico Engel et al. (2020). It demonstrates how self-attention mechanisms can capture local and global geometric relationships directly on point sets without voxelization or 2D projections.
  • Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). It provides the foundational hierarchical grouping and farthest point sampling techniques universally used to partition point clouds into localized patches.
  • Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). It introduces the fundamental deep learning framework for learning permutation-invariant features directly from raw 3D coordinate sets.
  • Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). It establishes deep autoencoder representations and distance-based geometric loss formulations like Chamfer Distance for point cloud generation and reconstruction.
  • Paper: Dynamic Graph CNN for Learning on Point Clouds, Yue Wang et al. (2018). It develops local neighborhood feature aggregation (EdgeConv) on point sets, which is central to building patch-level geometric representations in modern 3D architectures.
Cover for PointGPT: Auto-regressively Generative Pre-training from Point Clouds

Abstract

Large language models (LLMs) based on the generative pre-training transformer (GPT) [46] have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decoder [27], with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Self-supervised Learning for NLP and Image Processing
  • 2.2 Self-supervised Learning for Point Cloud
  • 3 PointGPT
  • 3.1 Point Cloud Sequencer
  • 3.2 Transformer Decoder with a Dual Masking Strategy
  • 3.3 Generation Target
  • 3.4 Post-Pre-training
  • 4 Experiments
  • 4.1 Implementation and Pre-training Setups
  • 4.2 Downstream Tasks
  • 4.3 Ablation Studies
  • 5 Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — PointGPT Pre-training Framework

    model/method

    PointGPT adapts the generative pre-training transformer (GPT) paradigm to 3D point cloud representation learning without relying on 2D/cross-modal teacher models or leaking holistic object shape information.

    Given an unordered 3D point cloud X={x1,x2,…,xM}⊂R3X = \{x_1, x_2, \dots, x_M\} \subset \mathbb{R}^3, the framework executes representation learning via three interconnected components:

    1. Point Cloud Sequencer: Partitions the unordered point cloud into localized geometric patches, sorts the patches into an ordered 1D sequence using a geometric Morton space-filling curve, and embeds each patch into a continuous token vector.
    2. Extractor-Generator Transformer Decoder:
      • An Extractor consisting of transformer decoder blocks processes the sequence under a causal dual masking strategy to learn high-level semantic representations conditioned strictly on preceding tokens, augmented with sinusoidal absolute positional encodings.
      • A lightweight Generator decoder takes the extracted representations and relative direction prompts (encoding the vectors between consecutive patch centers) to predict subsequent point patches without disclosing global shape.
    3. Auto-regressive Point Generation: An MLP prediction head projects generator tokens into 3D coordinate space, supervised by an auto-regressive generation loss based on the Chamfer Distance to ground-truth next patches.

    After pre-training, the generator is discarded, and the extractor is directly transferred to downstream 3D recognition and segmentation tasks without the dual mask.

  2. Knowl 2 — Point Cloud Sequencer for Sequential Patch Generation

    model/method

    To adapt auto-regressive sequential modeling to unordered 3D point clouds, the Point Cloud Sequencer converts an input point cloud X={x1,x2,…,xM}⊂R3X = \{x_1, x_2, \dots, x_M\} \subset \mathbb{R}^3 into an ordered sequence of patch embeddings via three operations:

    1. Point Patch Partitioning: Farthest Point Sampling (FPS) samples nn patch center points C∈Rn×3C \in \mathbb{R}^{n \times 3} from XX. For each center point ci∈Cc_i \in C, the KK-Nearest Neighbors (KNN) algorithm identifies the kk closest points in XX to form nn irregular point patches P∈Rn×k×3P \in \mathbb{R}^{n \times k \times 3}: C=FPS⁡(X),C∈Rn×3C = \operatorname{FPS}(X), \quad C \in \mathbb{R}^{n \times 3} P=KNN⁡(C,X),P∈Rn×k×3P = \operatorname{KNN}(C, X), \quad P \in \mathbb{R}^{n \times k \times 3}

    2. Morton Order Sorting: To impose a sequential geometric order while preserving 3D spatial adjacency, center coordinates are converted into 1D Morton codes (Z-order curve) and sorted to define permutation indices O∈Rn×1O \in \mathbb{R}^{n \times 1}: O=argmax⁡(MortonCode⁡(C))O = \operatorname{argmax}(\operatorname{MortonCode}(C)) Cs=C[O]∈Rn×3,Ps=P[O]∈Rn×k×3C^s = C[O] \in \mathbb{R}^{n \times 3}, \quad P^s = P[O] \in \mathbb{R}^{n \times k \times 3} where CsC^s and PsP^s denote the sorted center points and sorted point patches, respectively.

    3. Patch Embedding: Each patch in PsP^s is normalized relative to its center coordinate and embedded into a DD-dimensional token TT using a lightweight PointNet network: T=PointNet⁡(Ps),T∈Rn×DT = \operatorname{PointNet}(P^s), \quad T \in \mathbb{R}^{n \times D}

  3. Knowl 3 — Transformer Decoder with Dual Masking Strategy and Extractor-Generator Architecture

    model/method

    To handle the low information density of point clouds and bridge the semantic gap between low-level point reconstruction and high-level downstream tasks, PointGPT utilizes a specialized transformer decoder architecture:

    1. Dual Masking Strategy: Standard causal masking allows tokens to attend to all preceding patches. Because raw point clouds have high geometric redundancy, a dual mask Md∈{0,1}n×nM^d \in \{0, 1\}^{n \times n} additionally masks a random proportion (e.g., 70%70\%) of the preceding tokens for each query token. The attention calculation is formulated as: SelfAttention⁡(T)=softmax⁡(QKTD−(1−Md)⋅∞)V\operatorname{SelfAttention}(T) = \operatorname{softmax}\left(\frac{QK^T}{\sqrt{D}} - (1 - M^d) \cdot \infty\right)V where Q,K,V∈Rn×DQ, K, V \in \mathbb{R}^{n \times D} are linear projections of input tokens TT, and unmasked entries in MdM^d are 11 while masked entries are 00.

    2. Extractor: Composed of transformer decoder blocks with dual masking. It learns representations T∈Rn×DT \in \mathbb{R}^{n \times D} using Absolute Positional Encodings (APE) computed from sorted center coordinates CsC^s via sinusoidal positional encoding PE⁡(⋅)\operatorname{PE}(\cdot): T=Extractor⁡(T+APE⁡)T = \operatorname{Extractor}(T + \operatorname{APE})

    3. Generator with Relative Direction Prompts (RDP): A shallower transformer decoder specialized for generation. To resolve directional ambiguity in auto-regressive prediction without leaking absolute coordinates or global shape, Relative Direction Prompts (RDP) representing unit vectors between consecutive patch centers are added to the first n′=n−1n' = n - 1 extracted tokens T1:n′T_{1:n'}: RDP⁡i=PE⁡(Ci+1s−Cis∥Ci+1s−Cis∥2),i∈{1,…,n′},RDP⁡∈Rn′×D\operatorname{RDP}_i = \operatorname{PE}\left(\frac{C^s_{i+1} - C^s_i}{\|C^s_{i+1} - C^s_i\|_2}\right), \quad i \in \{1, \dots, n'\}, \quad \operatorname{RDP} \in \mathbb{R}^{n' \times D} Tg=Generator⁡(T1:n′+RDP⁡),Tg∈Rn′×DT^g = \operatorname{Generator}(T_{1:n'} + \operatorname{RDP}), \quad T^g \in \mathbb{R}^{n' \times D}

    4. Prediction Head: A two-layer MLP with ReLU activation projects each generator token TgT^g into point coordinates for the subsequent patch: Ppd=Reshape⁡(MLP⁡(Tg)),Ppd∈Rn′×k×3P^{pd} = \operatorname{Reshape}(\operatorname{MLP}(T^g)), \quad P^{pd} \in \mathbb{R}^{n' \times k \times 3}

  4. Knowl 4 — Auto-regressive Generation Loss and Auxiliary Fine-tuning Objective

    equation

    The pre-training generation objective measures the difference between the predicted point patches Ppd∈Rn′×k×3P^{pd} \in \mathbb{R}^{n' \times k \times 3} and the ground-truth subsequent patches Pgt∈Rn′×k×3P^{gt} \in \mathbb{R}^{n' \times k \times 3} (the last n′=n−1n' = n - 1 patches in sorted sequence PsP^s). The loss Lg\mathcal{L}^g combines the L1L_1-form and L2L_2-form Chamfer Distance (CD): Lg=L1g+L2g\mathcal{L}^g = \mathcal{L}^g_1 + \mathcal{L}^g_2 where the LnL_n-form CD loss for n∈{1,2}n \in \{1, 2\} is defined as: Lng=1∣Ppd∣∑a∈Ppdmin⁡b∈Pgt∥a−b∥nn+1∣Pgt∣∑b∈Pgtmin⁡a∈Ppd∥a−b∥nn\mathcal{L}^g_n = \frac{1}{|P^{pd}|} \sum_{a \in P^{pd}} \min_{b \in P^{gt}} \|a - b\|_n^n + \frac{1}{|P^{gt}|} \sum_{b \in P^{gt}} \min_{a \in P^{pd}} \|a - b\|_n^n where ∣P∣|P| is point set cardinality and ∥⋅∥n\|\cdot\|_n is the LnL_n norm.

    During downstream fine-tuning, the auto-regressive generation loss is retained as an auxiliary regularization objective: Lf=Ld+λ×Lg\mathcal{L}^f = \mathcal{L}^d + \lambda \times \mathcal{L}^g where Ld\mathcal{L}^d denotes the task-specific downstream loss (e.g., cross-entropy), Lg\mathcal{L}^g is the generation loss, and λ\lambda balances the objectives (empirically optimal at λ=3\lambda = 3).

  5. Knowl 5 — Scaled Pre-training and Intermediate Post-Pre-training Protocol

    model/method

    To prevent overfitting when scaling point cloud transformer models to large capacities (ViT-B and ViT-L), PointGPT adopts a multi-dataset scaling and intermediate post-pre-training strategy:

    1. Unlabeled Hybrid Dataset (UHD): An aggregation of approximately 300,000 unlabeled point clouds combining synthetic objects, indoor scenes, and outdoor scans from ShapeNet, S3DIS, Semantic3D, and others for self-supervised pre-training.
    2. Labeled Hybrid Dataset (LHD): A unified collection of approximately 200,000 point clouds across 87 aligned semantic categories sourced from SUN RGB-D, PartNet, ShapeNet, ModelNet40, S3DIS, and ScanObjectNN.
    3. Sequential Training Stages:
      • Self-Supervised Pre-training: High-capacity models (PointGPT-B with ViT-B configuration and PointGPT-L with ViT-L configuration) are pre-trained auto-regressively on UHD.
      • Supervised Post-Pre-training: The models undergo intermediate supervised classification training on LHD to integrate multi-source high-level semantics.
      • Task-Specific Fine-tuning: The resulting weights are fine-tuned on target downstream tasks.

    For standard single-dataset comparisons, PointGPT-S (ViT-S configuration) is pre-trained exclusively on ShapeNet without post-pre-training.

  6. Knowl 6 — 3D Object Classification Results on ScanObjectNN and ModelNet40

    empirical result

    PointGPT was evaluated for 3D object classification on the real-world ScanObjectNN dataset across three standard splits (OBJ_BG, OBJ_ONLY, and PB_T50_RS) and on the synthetic ModelNet40 dataset with 1k and 8k input points (using voting evaluation).

    Method ScanObjectNN Accuracy (%) ModelNet40 Accuracy (%) Post-Pre-train
    OBJ_BG OBJ_ONLY PB_T50_RS 1k P 8k P
    PointNet 73.3 79.2 68.0 89.2 90.8 No
    DGCNN 82.8 86.2 78.1 92.9 - No
    PointMLP - - 85.4 94.5 - No
    PointNeXt - - 87.7 94.0 - No
    Point-BERT 87.4 88.1 83.1 93.2 93.8 No
    MaskPoint 89.3 88.1 84.3 93.8 - No
    Point-MAE 90.0 88.2 85.2 93.8 94.0 No
    Point-M2AE 91.2 88.8 86.4 94.0 - No
    PointGPT-S 91.6 90.0 86.9 94.0 94.2 No
    PointGPT-B 95.8 95.2 91.9 94.4 94.6 Yes (UHD+LHD)
    PointGPT-L 97.2 96.6 93.4 94.7 94.9 Yes (UHD+LHD)
    ACT (cross-modal) 93.3 91.9 88.2 93.7 94.0 -
    ReCon (cross-modal) 95.4 93.6 91.3 94.5 94.7 -

    When pre-trained solely on ShapeNet under single-modal conditions, PointGPT-S achieved 86.9%86.9\% on ScanObjectNN PB_T50_RS and 94.0%94.0\% on ModelNet40 1k P, outperforming previous single-modal masked autoencoders (Point-MAE at 85.2%85.2\% / 93.8%93.8\%, Point-M2AE at 86.4%86.4\% / 94.0%94.0\%). Scaled with UHD and LHD post-pre-training, PointGPT-L reached 93.4%93.4\% on ScanObjectNN PB_T50_RS and 94.9%94.9\% on ModelNet40 8k P, exceeding multi-modal teacher-guided models such as ReCon (91.3%91.3\% and 94.7%94.7\%).

  7. Knowl 7 — Few-Shot Classification Performance on ModelNet40

    empirical result

    PointGPT was evaluated under standard few-shot classification settings on ModelNet40 over 10 independent trials without post-pre-training, testing ww-way ss-shot configurations (w∈{5,10}w \in \{5, 10\} and s∈{10,20}s \in \{10, 20\}).

    Method 5-way Accuracy (%) 10-way Accuracy (%)
    10-shot 20-shot 10-shot 20-shot
    DGCNN 31.6 ±\pm 2.8 40.8 ±\pm 4.6 19.9 ±\pm 2.1 16.9 ±\pm 1.5
    OcCo 90.6 ±\pm 2.8 92.5 ±\pm 1.9 82.9 ±\pm 1.3 86.5 ±\pm 2.2
    Point-BERT 94.6 ±\pm 3.1 96.3 ±\pm 2.7 91.0 ±\pm 5.4 92.7 ±\pm 5.1
    MaskPoint 95.0 ±\pm 3.7 97.2 ±\pm 1.7 91.4 ±\pm 4.0 93.4 ±\pm 3.5
    Point-MAE 96.3 ±\pm 2.5 97.8 ±\pm 1.8 92.6 ±\pm 4.1 95.0 ±\pm 3.0
    Point-M2AE 96.8 ±\pm 1.8 98.3 ±\pm 1.4 92.3 ±\pm 4.5 95.0 ±\pm 3.0
    PointGPT-S 96.8 ±\pm 2.0 98.6 ±\pm 1.1 92.6 ±\pm 4.6 95.2 ±\pm 3.4
    PointGPT-B (UHD pre-train) 97.5 ±\pm 2.0 98.8 ±\pm 1.0 93.5 ±\pm 4.0 95.8 ±\pm 3.0
    PointGPT-L (UHD pre-train) 98.0 ±\pm 1.9 99.0 ±\pm 1.0 94.1 ±\pm 3.3 96.1 ±\pm 2.8
    ACT (cross-modal) 96.8 ±\pm 2.3 98.0 ±\pm 1.4 93.3 ±\pm 4.0 95.6 ±\pm 2.8
    ReCon (cross-modal) 97.3 ±\pm 1.9 98.9 ±\pm 1.2 93.3 ±\pm 3.9 95.8 ±\pm 3.0

    PointGPT-S pre-trained on ShapeNet matched or surpassed previous single-modal approaches across all benchmark splits. Pre-training on the larger UHD dataset enabled PointGPT-L to establish new state-of-the-art results: 98.0%98.0\% (5-way 10-shot), 99.0%99.0\% (5-way 20-shot), 94.1%94.1\% (10-way 10-shot), and 96.1%96.1\% (10-way 20-shot).

  8. Knowl 8 — 3D Part Segmentation Performance on ShapeNetPart

    empirical result

    PointGPT was evaluated on the ShapeNetPart dataset (16,881 objects across 16 categories, sampled to 2048 points) using class mean intersection over union (Cls. mIoU) and instance mean intersection over union (Inst. mIoU). Multi-scale feature representations were extracted by concatenating outputs from layers 13td\frac{1}{3}t_d, 23td\frac{2}{3}t_d, and tdt_d of the extractor (where tdt_d is the extractor depth), followed by average and max pooling, upsampling, and an MLP segmentation head.

    Method Cls. mIoU (%) Inst. mIoU (%)
    PointNet 80.4 83.7
    PointNet++ 81.9 85.1
    DGCNN 82.3 85.2
    PointMLP 84.6 86.1
    PointContrast - 85.1
    CrossPoint - 85.5
    Point-BERT 84.1 85.6
    Point-MAE - 86.1
    PointGPT-S 84.1 86.2
    PointGPT-B (UHD + LHD) 84.5 86.5
    PointGPT-L (UHD + LHD) 84.8 86.6
    ACT (cross-modal) 84.7 86.1
    ReCon (cross-modal) 84.8 86.4

    PointGPT-S achieved 86.2%86.2\% Inst. mIoU, outperforming Point-MAE (86.1%86.1\%) and Point-BERT (85.6%85.6\%). Scaled PointGPT-L achieved 84.8%84.8\% Cls. mIoU and 86.6%86.6\% Inst. mIoU, exceeding the cross-modal teacher-based method ReCon (84.8%84.8\% Cls. mIoU, 86.4%86.4\% Inst. mIoU).

  9. Knowl 9 — Ablation Analysis of PointGPT Architectural and Pre-training Components

    empirical result

    Ablation studies on PointGPT-S pre-trained on ShapeNet and fine-tuned on ModelNet40 (without post-pre-training) evaluated individual system components:

    1. Generator Depth: Increasing the generator depth from 0 blocks (93.85%93.85\%), 2 blocks (94.08%94.08\%), 4 blocks (94.21%94.21\%) to 6 blocks (94.24%94.24\%) confirmed that a dedicated generator decouples reconstruction from semantic representation learning. Depth 4 was chosen as default.
    2. Generation Targets: Predicting point coordinates (94.21%94.21\%) outperformed handcrafted FPFH geometric features (94.13%94.13\%). While two-stage features from pre-trained PointNet (94.31%94.31\%) or DGCNN (94.35%94.35\%) gave marginal improvements, direct coordinate prediction avoided teacher-model pre-training and inference cost.
    3. Auxiliary Generation Loss during Fine-tuning: Varying the auxiliary loss weight λ∈{0,1,3,5}\lambda \in \{0, 1, 3, 5\} resulted in classification accuracies of 94.01%94.01\%, 94.15%94.15\%, 94.21%94.21\%, and 94.05%94.05\%, showing that λ=3\lambda = 3 optimal regularization.
    4. Chamfer Loss Formulation: Combining L1L_1 and L2L_2 CD losses (94.21%94.21\%) was superior to L1L_1 only (93.66%93.66\%) or L2L_2 only (94.13%94.13\%).
    5. Generator Prompts: Relative Direction Prompts (RDP) achieved 94.21%94.21\%, outperforming absolute positional encoding (94.06%94.06\%) and no prompts (93.69%93.69\%) by preventing overfitting to the patch ordering.
    6. Dual Masking Ratio: Testing masking ratios {0,10%,30%,50%,70%,90%}\{0, 10\%, 30\%, 50\%, 70\%, 90\%\} produced accuracies {93.68%,93.70%,93.85%,94.01%,94.21%,93.66%}\{93.68\%, 93.70\%, 93.85\%, 94.01\%, 94.21\%, 93.66\%\}, indicating 70%70\% is optimal.
    7. Point Sorting Algorithm: Morton code sorting (91.6%91.6\% OBJ_BG, 86.9%86.9\% PB_T50_RS on ScanObjectNN; 94.0%94.0\% on ModelNet40) outperformed KD-tree sorting (81.3%81.3\% PB_T50_RS; 93.2%93.2\% ModelNet40) and Hilbert curve sorting (84.3%84.3\% PB_T50_RS; 93.4%93.4\% ModelNet40) due to superior preservation of 3D spatial adjacency in 1D.
  10. Knowl 10 — PointGPT Pre-training Hyperparameters and Experimental Setup

    experimental setup

    The experimental configuration for PointGPT pre-training and architecture variants is defined as follows:

    1. Point Cloud Sampling and Patching: Raw point clouds are uniformly sampled to M=1024M = 1024 points. Farthest Point Sampling extracts n=64n = 64 center points, and KK-Nearest Neighbors with k=32k = 32 points constructs n=64n = 64 patches.
    2. Model Variants:
      • PointGPT-S: Extractor has 12 transformer blocks, embedding dimension D=384D = 384, 6 attention heads (ViT-S configuration); Generator has 4 decoder blocks (D=384D = 384); total parameters ≈19.5M\approx 19.5\text{M}.
      • PointGPT-B: Extractor has 12 transformer blocks, embedding dimension D=768D = 768, 12 attention heads (ViT-B configuration); total parameters ≈82.1M\approx 82.1\text{M}.
      • PointGPT-L: Extractor has 24 transformer blocks, embedding dimension D=1024D = 1024, 16 attention heads (ViT-L configuration).
    3. Pre-training Optimization: Pre-trained for 300 epochs using the AdamW optimizer with initial learning rate 1×10−31 \times 10^{-3}, cosine learning rate decay, weight decay 0.050.05, and batch size 128128.
    4. Dual Masking and Loss Weights: Pre-training dual mask ratio is set to 0.700.70. Auxiliary generation loss weight during fine-tuning is set to λ=3\lambda = 3.

Coverage note — None was omitted; all primary contributions, including the point cloud sequencer, dual-masking extractor-generator decoder, auto-regressive pre-training and fine-tuning losses, dataset scaling, benchmark evaluations, and ablation studies, are represented.

Citation

MLA
Chen, G., et al. “PointGPT: Auto-regressively Generative Pre-training from Point Clouds”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 29667–79, https://proceedings.neurips.cc/paper_files/paper/2023/file/5ed5c3c846f684a54975ad7a2525199f-Paper-Conference.pdf.
APA
Chen, G., Wang, M., Yang, Y., Yu, K., Yuan, L., & Yue, Y. (2023). PointGPT: Auto-regressively Generative Pre-training from Point Clouds. Advances in Neural Information Processing Systems, 36, 29667–29679. https://proceedings.neurips.cc/paper_files/paper/2023/file/5ed5c3c846f684a54975ad7a2525199f-Paper-Conference.pdf
Chicago
Chen, G., M. Wang, Y. Yang, K. Yu, L. Yuan, and Y. Yue. 2023. “PointGPT: Auto-regressively Generative Pre-training from Point Clouds”. Advances in Neural Information Processing Systems 36: 29667–79. https://proceedings.neurips.cc/paper_files/paper/2023/file/5ed5c3c846f684a54975ad7a2525199f-Paper-Conference.pdf.
Harvard
Chen, G. et al. (2023) “PointGPT: Auto-regressively Generative Pre-training from Point Clouds”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 29667–29679. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/5ed5c3c846f684a54975ad7a2525199f-Paper-Conference.pdf.
Vancouver
1. Chen G, Wang M, Yang Y, Yu K, Yuan L, Yue Y (2023) PointGPT: Auto-regressively Generative Pre-training from Point Clouds. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 29667–29679

BibTeX

@inproceedings{chen2023pointgpt,
  title = {PointGPT: Auto-regressively Generative Pre-training from Point Clouds},
  author = {Chen, Guangyan and Wang, Meiling and Yang, Yi and Yu, Kai and Yuan, Li and Yue, Yufeng},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {29667-29679},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/5ed5c3c846f684a54975ad7a2525199f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors