keyword
ModelNet40 dataset
The ModelNet40 dataset is a widely used benchmark dataset in computer vision and deep learning designed for 3D object classification, shape retrieval, and spatial representation learning. Originating as a curated subset of the Princeton ModelNet project, it contains over 12,000 clean 3D computer-aided design models categorized into 40 distinct object classes, such as airplanes, chairs, cars, and tables. While the models are originally structured as 3D polygonal meshes, they are frequently converted into point clouds, volumetric voxel grids, or multi-view 2D image projections to train and evaluate diverse 3D deep learning architectures against standardized training and testing splits.
3 items

PointGPT: Auto-regressively Generative Pre-training from Point Clouds
Guangyan Chen, Meiling Wang, Yi Yang, Kai Yu, Li Yuan, Yufeng Yue
Why you should read this
Proposes an autoregressive generative pre-training framework for 3D point clouds that arranges patches using Morton ordering and applies a dual masking strategy to prevent shape leakage, achieving state-of-the-art representation learning performance across standard benchmarks.
Large language models (LLMs) based on the generative pre-training transformer (GPT) [46] have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, addressing the challenges associated with disorder properties, low information density, and task gaps. Specifically, a point cloud auto-regressive generation task is proposed to pre-train transformer models. Our method partitions the input point cloud into multiple point patches and arranges them in an ordered sequence based on their spatial proximity. Then, an extractor-generator based transformer decoder [27], with a dual masking strategy, learns latent representations conditioned on the preceding point patches, aiming to predict the next one in an auto-regressive manner. To explore scalability and enhance performance, a larger pre-training dataset is collected. Additionally, a subsequent post-pre-training stage is introduced, incorporating a labeled hybrid dataset. Our scalable approach allows for learning high-capacity models that generalize well, achieving state-of-the-art performance on various downstream tasks. In particular, our approach achieves classification accuracies of 94.9% on the ModelNet40 dataset and 93.4% on the ScanObjectNN dataset, outperforming all other transformer models. Furthermore, our method also attains new state-of-the-art accuracies on all four few-shot learning benchmarks. Codes are available at https://github.com/CGuangyan-BIT/PointGPT.
Added
2026-09-26

Volumetric and Multi-view CNNs for Object Classification on 3D Data
Charles R. Qi, Hao Su, Matthias Niessner, Angela Dai, Mengyuan Yan, Leonidas J. Guibas
Why you should read this
Proposes improved volumetric and multi-view convolutional neural network architectures to close the performance gap between 3D voxel and 2D view-based representations for 3D object classification.
3D shape models are becoming widely available and easier to capture, making available 3D information crucial for progress in object classification. Current state-of-the-art methods rely on CNNs to address this problem. Recently, we witness two types of CNNs being developed: CNNs based upon volumetric representations versus CNNs based upon multi-view representations. Empirical results from these two types of CNNs exhibit a large gap, indicating that existing volumetric CNN architectures and approaches are unable to fully exploit the power of 3D representations. In this paper, we aim to improve both volumetric CNNs and multi-view CNNs according to extensive analysis of existing approaches. To this end, we introduce two distinct network architectures of volumetric CNNs. In addition, we examine multi-view CNNs, where we introduce multi-resolution filtering in 3D. Overall, we are able to outperform current state-of-the-art methods for both volumetric CNNs and multi-view CNNs. We provide extensive experiments designed to evaluate underlying design choices, thus providing a better understanding of the space of methods available for object classification on 3D data.
Added
2026-09-24

PCT: Point cloud transformer
Meng-Hao Guo, Junxiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph Robert Martin, Shimin Hu
Why you should read this
Proposes a specialized Transformer architecture for 3D point clouds that utilizes permutation invariance and localized feature aggregation to achieve state-of-the-art accuracy across shape classification, part segmentation, and normal estimation benchmarks.
The irregular domain and lack of ordering make it challenging to design deep neural networks for point cloud processing. This paper presents a novel framework named Point Cloud Transformer(PCT) for point cloud learning. PCT is based on Transformer, which achieves huge success in natural language processing and displays great potential in image processing. It is inherently permutation invariant for processing a sequence of points, making it well-suited for point cloud learning. To better capture local context within the point cloud, we enhance input embedding with the support of farthest point sampling and nearest neighbor search. Extensive experiments demonstrate that the PCT achieves the state-of-the-art performance on shape classification, part segmentation and normal estimation tasks.
Added
2026-09-16
