Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ModelNet40 dataset

The ModelNet40 dataset is a widely used benchmark dataset in computer vision and deep learning designed for 3D object classification, shape retrieval, and spatial representation learning. Originating as a curated subset of the Princeton ModelNet project, it contains over 12,000 clean 3D computer-aided design models categorized into 40 distinct object classes, such as airplanes, chairs, cars, and tables. While the models are originally structured as 3D polygonal meshes, they are frequently converted into point clouds, volumetric voxel grids, or multi-view 2D image projections to train and evaluate diverse 3D deep learning architectures against standardized training and testing splits.

3 items

PointGPT: Auto-regressively Generative Pre-training from Point Clouds

PointGPT: Auto-regressively Generative Pre-training from Point Clouds

Guangyan Chen, Meiling Wang, Yi Yang, Kai Yu, Li Yuan, Yufeng Yue

OrganizationsBeijing Institute of TechnologyPeking University

Why you should read this

Proposes an autoregressive generative pre-training framework for 3D point clouds that arranges patches using Morton ordering and applies a dual masking strategy to prevent shape leakage, achieving state-of-the-art representation learning performance across standard benchmarks.

Large language models (LLMs) based on the generative pre-training transformer (GPT) [46] have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decoder [27], with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT.

Added

2026-09-26

Volumetric and Multi-view CNNs for Object Classification on 3D Data

Volumetric and Multi-view CNNs for Object Classification on 3D Data

Charles R. Qi, Hao Su, Matthias Niessner, Angela Dai, Mengyuan Yan, Leonidas J. Guibas

OrganizationsStanford University

Why you should read this

Proposes improved volumetric and multi-view convolutional neural network architectures to close the performance gap between 3D voxel and 2D view-based representations for 3D object classification.

3D shape models are becoming widely available and easier to capture, making available 3D information crucial for progress in object classification. Current state-of-the-art methods rely on CNNs to address this problem. Recently, we witness two types of CNNs being developed: CNNs based upon volumetric representations versus CNNs based upon multi-view representations. Empirical results from these two types of CNNs exhibit a large gap, indicating that existing volumetric CNN architectures and approaches are unable to fully exploit the power of 3D representations. In this paper, we aim to improve both volumetric CNNs and multi-view CNNs according to extensive analysis of existing approaches. To this end, we introduce two distinct network architectures of volumetric CNNs. In addition, we examine multi-view CNNs, where we introduce multi-resolution filtering in 3D. Overall, we are able to outperform current state-of-the-art methods for both volumetric CNNs and multi-view CNNs. We provide extensive experiments designed to evaluate underlying design choices, thus providing a better understanding of the space of methods available for object classification on 3D data.

Added

2026-09-24