CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data
Yihan ZengChenhan JiangJiageng MaoJianhua HanChaoqiang YeQingqiu HuangDit-Yan YeungZhen YangXiaodan LiangHang Xu
Introduces a cross-modal pretraining framework that aligns real-world 3D point clouds with text and images using automatically generated triplet proxies, enabling strong zero-shot and few-shot 3D recognition without relying on intermediate 2D projection losses.
Three-dimensional (3D) point cloud perception provides accurate spatial geometry and resilience to illumination changes, making it vital for safety-critical applications such as autonomous driving and robotics. However, conventional 3D recognition methods rely heavily on closed, human-annotated taxonomies that fail to identify rare or unexpected objects in complex environments. While 2D vision-language models have achieved breakthrough open-vocabulary performance using vast Internet-scale text-image datasets, comparable 3D vision-language pretraining has stalled due to the severe scarcity and synthetic bias of paired 3D datasets, alongside previous attempts relying on 2D depth projections that discard crucial 3D geometric structures.
The article introduces and evaluates CLIP2 (Contrastive Language-Image-Point Cloud Pretraining), a framework designed to directly align raw 3D point cloud representations with open-vocabulary natural language without requiring manual 3D annotations.
The approach operates in two main stages using real-world indoor and outdoor datasets. First, the authors implement an automated Triplet Proxy Collection pipeline that leverages a 2D open-vocabulary detector and spatial calibration to extract paired text descriptions, 2D image proposals, and corresponding 3D point cloud clusters, yielding over 1.6 million unannotated training triplets. Second, the authors apply a cross-modal contrastive pretraining framework initialized from pretrained 2D vision-language embeddings. This joint objective aligns the 3D point cloud encoder at both the semantic text level and the instance image level.
Evaluation across multiple indoor, outdoor, and object-level benchmarks demonstrated significant performance gains. On zero-shot indoor recognition, CLIP2 attained a 61.3% mean Top-1 accuracy on SUN RGB-D and 43.8% on ScanNet, substantially outperforming prior intermediate-depth methods. On realistic object-level benchmarks, CLIP2 reached 39.1% zero-shot classification accuracy, a 16.1% relative improvement over the previous state-of-the-art, and outperformed competing models by 5.3% to 9.6% across few-shot settings. In outdoor autonomous driving scenarios, the framework achieved 37.8% average zero-shot accuracy across nuScenes and ONCE, outperforming baseline models by more than 20%, while successfully discovering unannotated long-tail hazards such as tires, road debris, and carried bags. Multi-modal ensembling further increased indoor recognition performance to 69.6%.
These findings indicate that aligning native 3D geometry directly with language spaces provides a viable, cost-effective path to open-world machine perception. Eliminating the need for manual 3D labeling reduces developmental overhead while significantly mitigating safety risks associated with uncataloged obstacles in autonomous navigation. The evidence confirms that preserving full 3D point cloud geometry is inherently superior to relying on intermediate 2D depth projections, especially in sparse outdoor LiDAR settings.
Organizations developing autonomous systems should consider integrating automated cross-modal proxy extraction and 3D contrastive pretraining to broaden perception capabilities beyond fixed taxonomies. Engineering teams can also implement multi-modal representation ensembling during inference whenever complementary camera feeds are operational. However, because CLIP2 does not yet generate tightly fitted 3D bounding boxes and exhibits low precision due to over-generating open-world object proposals, it should not yet operate as a standalone bounding-box detector. Future efforts should focus on integrating this open-vocabulary representation into downstream 3D bounding-box detection architectures and conducting pilot field validations in real-time navigation pipelines.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. It introduces the foundational contrastive language-image pretraining (CLIP) paradigm and visual-textual embedding space that CLIP2 directly builds upon and extends to 3D point clouds.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). It provides the foundational deep learning architecture for directly processing unordered 3D point clouds, establishing the structural representations required for native 3D geometry encoding in CLIP2.
- Paper: Grounded Language-Image Pre-training, Liunian Harold Li et al. (2022). It introduces grounded language-image pretraining to align object-level visual regions with textual concepts, underlying the open-vocabulary detection proxies used to generate CLIP2's training triplets.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). It demonstrates open-vocabulary object detection through vision-and-language knowledge distillation, establishing the 2D detector methodology that enables automated cross-modal proposal extraction in CLIP2.
- Paper: Frustum PointNets for 3D Object Detection from RGB-D Data, Charles R. Qi et al. (2018). It pioneers the technique of projecting 2D bounding box proposals into 3D viewing frustums to isolate point cloud clusters, a core geometric mechanism used in CLIP2's triplet extraction pipeline.
- Paper: Deep Learning for 3D Point Clouds: A Survey, Yulan Guo et al. (2019). It provides a comprehensive survey of deep learning paradigms and benchmarks for 3D point cloud understanding, framing the structural limitations that CLIP2 seeks to overcome.
- Paper: Context-aware Alignment and Mutual Masking for 3D-Language Pre-training, Zhao Jin et al. (2023). It extends multimodal 3D representation learning by incorporating context-aware alignment and mutual masking objectives for pretraining on 3D-language data.
- Paper: CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language, Aditya Sanghi et al. (2023). It leverages contrastive vision-language representations to synthesize diverse, high-fidelity 3D shapes directly from natural language prompts.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). It builds upon open-vocabulary multimodal representations to scale promptable 3D object detection across massive unconstrained environments.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). It explores lifting 2D image-level segmentations into coherent 3D scene-level volumetric neural fields for panoptic understanding.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). It synthesizes the broader landscape of adapting vision-language foundation models to visual and spatial perception tasks.
