Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training
Xiaoyang WuZhuotao TianXin WenBohao PengXihui LiuKaicheng YuHengshuang Zhao
Proposes Point Prompt Training to overcome negative transfer across diverse 3D point cloud datasets by using domain-specific prompts and language-guided label alignment to train a single high-performing representation model.
Three-dimensional deep learning has historically lagged behind 2D vision and language processing because 3D point cloud datasets are relatively small and expensive to collect. While combining multiple existing 3D datasets into a single training pool seems like a natural way to scale up models, direct joint training causes severe negative transfer. Differences in data density, sensor environments, and label spaces cause models trained on combined data to perform worse on individual datasets than models trained on a single source alone.
The article aims to overcome these multi-source domain barriers by developing Point Prompt Training, a unified training framework that enables a single 3D perception model to learn collaboratively from multiple datasets without degrading performance across specific tasks.
To achieve this, the authors designed two core mechanisms: Prompt-driven Normalization and Language-guided Categorical Alignment. The normalization module injects lightweight, learnable dataset-specific prompts into the model layers, adjusting feature scaling and shifting to absorb domain-specific variations while keeping shared backbone representations generalizable. The categorical alignment mechanism maps disparate dataset category names into a shared semantic language embedding using a pre-trained text encoder, unifying conflicting label definitions. The framework was evaluated across diverse indoor and outdoor benchmarks—including ScanNet, S3DIS, Structured3D, SemanticKITTI, nuScenes, and Waymo—using both convolutional and transformer backbones under supervised joint training, supervised pre-training, and unsupervised pre-training settings.
The experimental findings show that Point Prompt Training completely reverses negative transfer and consistently outperforms single-dataset and standard joint-training baselines. In indoor semantic segmentation, joint training increased accuracy on ScanNet from 68.9% to 75.7% and on S3DIS from 63.3% to 72.2% Mean Intersection over Union compared to naive joint baselines. In outdoor autonomous driving benchmarks, the unified model improved SemanticKITTI validation accuracy by over 7 percentage points over scratch baselines. Additionally, the representations transferred successfully to instance segmentation and showed high data efficiency, achieving strong performance even when training scenes had as few as 20 annotated points.
These results demonstrate that organizations can train a single, shared-weight 3D model across varied synthetic, real, indoor, and outdoor data sources rather than deploying fragmented, dataset-specific models. This unified approach reduces operational overhead and model maintenance while unlocking the benefits of synthetic-to-real transfer. It establishes that prompt tuning and language grounding can effectively bridge structural domain gaps during representation pre-training.
Organizations developing 3D perception systems should adopt domain prompting and language-based label unification when combining disparate point cloud data sources. Future engineering and research efforts should explore extending the framework to simultaneous multi-task training (such as combining segmentation with 3D object detection) and evaluate joint cross-domain training that mixes indoor and outdoor datasets directly.
Confidence in these findings is high given the extensive testing across standard benchmarks and diverse model architectures. However, decision-makers should note that the current study focuses primarily on dense semantic segmentation tasks, relies on pre-trained text encoders for category alignment, and has not yet evaluated a single model trained simultaneously on mixed indoor-outdoor domain extremes.
- Paper: Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection, Bo Zhang et al. (2023). This paper establishes the core challenges of data-level differences and taxonomy discrepancies when training unified 3D models across multiple point cloud datasets, providing the exact baseline context for Point Prompt Training.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). This work introduces Visual Prompt Tuning to adapt vision transformers using lightweight learnable prompt tokens, which underpins the prompt-driven layer adaptation mechanism utilized in the source paper.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). This foundational paper presents context optimization for vision-language models, establishing the prompt-based conditioning and text-feature alignment methods extended by the source for categorical alignment.
- Paper: CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data, Yihan Zeng et al. (2023). This paper demonstrates language-guided cross-modal pretraining for raw 3D point clouds, laying the groundwork for mapping disparate 3D category names into shared text embeddings.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). This seminal work establishes direct neural processing on unordered 3D point sets and defines standard indoor benchmark evaluations on ScanNet and S3DIS.
- Paper: Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer, René Ranftl et al. (2019). This study analyzes the difficulties of mixing heterogeneous visual datasets during training and developing invariant strategies to prevent negative transfer across diverse domains.
No sufficiently relevant recommendations were found.
