CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

Yihan ZengChenhan JiangJiageng MaoJianhua HanChaoqiang YeQingqiu HuangDit-Yan YeungZhen YangXiaodan LiangHang Xu

article2023CVPR107 citations

Introduces a cross-modal pretraining framework that aligns real-world 3D point clouds with text and images using automatically generated triplet proxies, enabling strong zero-shot and few-shot 3D recognition without relying on intermediate 2D projection losses.

Listen

Three-dimensional (3D) point cloud perception provides accurate spatial geometry and resilience to illumination changes, making it vital for safety-critical applications such as autonomous driving and robotics. However, conventional 3D recognition methods rely heavily on closed, human-annotated taxonomies that fail to identify rare or unexpected objects in complex environments. While 2D vision-language models have achieved breakthrough open-vocabulary performance using vast Internet-scale text-image datasets, comparable 3D vision-language pretraining has stalled due to the severe scarcity and synthetic bias of paired 3D datasets, alongside previous attempts relying on 2D depth projections that discard crucial 3D geometric structures.

The article introduces and evaluates CLIP2 (Contrastive Language-Image-Point Cloud Pretraining), a framework designed to directly align raw 3D point cloud representations with open-vocabulary natural language without requiring manual 3D annotations.

The approach operates in two main stages using real-world indoor and outdoor datasets. First, the authors implement an automated Triplet Proxy Collection pipeline that leverages a 2D open-vocabulary detector and spatial calibration to extract paired text descriptions, 2D image proposals, and corresponding 3D point cloud clusters, yielding over 1.6 million unannotated training triplets. Second, the authors apply a cross-modal contrastive pretraining framework initialized from pretrained 2D vision-language embeddings. This joint objective aligns the 3D point cloud encoder at both the semantic text level and the instance image level.

Evaluation across multiple indoor, outdoor, and object-level benchmarks demonstrated significant performance gains. On zero-shot indoor recognition, CLIP2 attained a 61.3% mean Top-1 accuracy on SUN RGB-D and 43.8% on ScanNet, substantially outperforming prior intermediate-depth methods. On realistic object-level benchmarks, CLIP2 reached 39.1% zero-shot classification accuracy, a 16.1% relative improvement over the previous state-of-the-art, and outperformed competing models by 5.3% to 9.6% across few-shot settings. In outdoor autonomous driving scenarios, the framework achieved 37.8% average zero-shot accuracy across nuScenes and ONCE, outperforming baseline models by more than 20%, while successfully discovering unannotated long-tail hazards such as tires, road debris, and carried bags. Multi-modal ensembling further increased indoor recognition performance to 69.6%.

These findings indicate that aligning native 3D geometry directly with language spaces provides a viable, cost-effective path to open-world machine perception. Eliminating the need for manual 3D labeling reduces developmental overhead while significantly mitigating safety risks associated with uncataloged obstacles in autonomous navigation. The evidence confirms that preserving full 3D point cloud geometry is inherently superior to relying on intermediate 2D depth projections, especially in sparse outdoor LiDAR settings.

Organizations developing autonomous systems should consider integrating automated cross-modal proxy extraction and 3D contrastive pretraining to broaden perception capabilities beyond fixed taxonomies. Engineering teams can also implement multi-modal representation ensembling during inference whenever complementary camera feeds are operational. However, because CLIP2 does not yet generate tightly fitted 3D bounding boxes and exhibits low precision due to over-generating open-world object proposals, it should not yet operate as a standalone bounding-box detector. Future efforts should focus on integrating this open-vocabulary representation into downstream 3D bounding-box detection architectures and conducting pilot field validations in real-time navigation pipelines.

arXiv: 2303.12417
Cover for CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

Abstract

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remains an open problem. Existing works that leverage VLM for 3D understanding generally resort to constructing intermediate 2D representations for the 3D data, but at the cost of losing 3D geometry information. To take a step toward open-world 3D vision understanding, we propose Contrastive Language-Image-Point Cloud Pretraining (CLIP²) to directly learn the transferable 3D point cloud representation in realistic scenarios with a novel proxy alignment mechanism. Specifically, we exploit naturally-existed correspondences in 2D and 3D scenarios, and build well-aligned and instance-based text-image-point proxies from those complex scenarios. On top of that, we propose a cross-modal contrastive objective to learn semantic and instance-level aligned point cloud representation. Experimental results on both indoor and outdoor scenarios show that our learned 3D representation has great transfer ability in downstream tasks, including zero-shot and few-shot 3D recognition, which boosts the state-of-the-art methods by large margins. Furthermore, we provide analyses of the capability of different representations in real scenarios and present the optional ensemble scheme.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Triplet Proxy Collection
  • 3.2. Cross-Modal Contrastive Pretraining
  • 4. Experiment
  • 4.1. Zero-shot Transfer
  • 4.1.1 Indoor Scenarios
  • 4.1.2 Outdoor Scenarios
  • 4.2. Few-shot Classification
  • 4.3. Ablations and Analysis
  • 5. Limitation
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — CLIP2 Framework Architecture for Language-Image-Point Cloud Alignment

    model/method

    The Contrastive Language-Image-Point Cloud Pretraining (CLIP2^2) framework aligns 3D point cloud representations directly with the pre-trained 2D vision-language feature space of CLIP without converting 3D data into intermediate 2D representations such as projected depth maps.

    The framework consists of three modality encoders:

    1. Language Encoder (EθTE_\theta^T): Extracted from a pre-trained Vision-Language Model (CLIP) and kept frozen to encode text prompts into embeddings fT∈R1×CTf^T \in \mathbb{R}^{1 \times C_T}.
    2. Visual Image Encoder (EθIE_\theta^I): Extracted from pre-trained CLIP and kept frozen to encode cropped 2D object proposal images into embeddings fI∈R1×CIf^I \in \mathbb{R}^{1 \times C_I}.
    3. Point Cloud Encoder (EθPE_\theta^P): A trainable 3D backbone (such as PointNet++) that processes raw 3D point cloud instances into 3D geometric embeddings fP∈R1×CPf^P \in \mathbb{R}^{1 \times C_P}.

    Embeddings from all three encoders are projected to a shared CC-dimensional space. CLIP2^2 trains the point cloud encoder to align 3D geometry simultaneously with high-level category semantic descriptions from text and instance-level visual features from 2D images.

  2. Knowl 2 — Triplet Proxy Collection Pipeline

    algorithm

    To address the scarcity of aligned 3D-text pairs without human annotation, the Triplet Proxy Collection mechanism automatically constructs a multimodal pretraining dataset Dproxy={(XiT,XiI,XiP)}i=1M\mathcal{D}_{\text{proxy}} = \{ (X_i^T, X_i^I, X_i^P) \}_{i=1}^M of language proxies XTX^T, image proxies XIX^I, and 3D point cloud proxies XPX^P from unannotated multimodal scene datasets S={(Ps,Is)}s=1∣S∣\mathcal{S} = \{ (P_s, I_s) \}_{s=1}^{|\mathcal{S}|}, where Ps∈RNP×3P_s \in \mathbb{R}^{N_P \times 3} is the scene point cloud and Is∈RNI×H×W×3I_s \in \mathbb{R}^{N_I \times H \times W \times 3} is the scene image.

    Input: Multimodal scene dataset S={(Ps,Is)}s=1∣S∣\mathcal{S} = \{(P_s, I_s)\}_{s=1}^{|\mathcal{S}|}, open-world vocabulary list XT={t1,…,tV}X^T = \{t_1, \dots, t_V\}, open-vocabulary 2D detector MM, score threshold ϵ=0.3\epsilon = 0.3
    Output: Triplet proxy dataset $\mathcal{D}_{\text{proxy}} = \{(X_i^T, X_i^I, X_i^P)\}
    Initialize Dproxy←∅\mathcal{D}_{\text{proxy}} \leftarrow \emptyset
    for each scene (Ps,Is)∈S(P_s, I_s) \in \mathcal{S} do
        Extract 2D bounding box proposals and matched text labels using M(Is,XT)M(I_s, X^T) with threshold ϵ\epsilon:
        {(Xs,kI,Xs,kT)}k←M(Is,XT)\{(X_{s,k}^I, X_{s,k}^T)\}_k \leftarrow M(I_s, X^T)
        for each 2D proposal Xs,kIX_{s,k}^I with category text Xs,kTX_{s,k}^T do
            if Scene type is Indoor (RGB-D) then
                Segment foreground pixels in Xs,kIX_{s,k}^I using GrabCut
                Transform segmented foreground depth pixels from (u,v,d)(u,v,d) coordinates to Cartesian (x,y,z)(x,y,z) coordinates via camera intrinsics to form point cloud proxy Xs,kPX_{s,k}^P
            else if Scene type is Outdoor (LiDAR) then
                Extrude 2D bounding box Xs,kIX_{s,k}^I into a 3D view frustum using camera calibration
                Cluster LiDAR points within the 3D frustum using DBSCAN
                Select the principal point cloud cluster as Xs,kPX_{s,k}^P
            end if
            $\mathcal{D}_{\text{proxy}} \leftarrow \mathcal{D}_{\text{proxy}} \cup \{(X_{s,k}^T, X_{s,k}^I, X_{s,k}^P)\}
        end for
    end for
    return Dproxy\mathcal{D}_{\text{proxy}}

    This pipeline automatically generates approximately 220,000 proxy triplets for indoor scenes (SUN RGB-D) and 1.4 million proxy triplets for outdoor scenes (nuScenes) using an open-world vocabulary of V=1206V = 1206 categories from LVIS.

  3. Knowl 3 — Cross-Modal Contrastive Pretraining Loss Formulation

    equation

    CLIP2^2 trains the 3D point cloud encoder via a cross-modal contrastive learning objective LCM(T,I,P)L_{CM}(T, I, P) that jointly optimizes semantic-level language-to-3D alignment and instance-level image-to-3D alignment.

    For a training mini-batch of size NN with triplet proxies (XiT,XiI,XiP)(X_i^T, X_i^I, X_i^P) and normalized feature vectors fiT,fiI,fiP∈RCf_i^T, f_i^I, f_i^P \in \mathbb{R}^C under temperature coefficient τ\tau:

    1. Semantic-Level Language-3D Alignment Objective: l(i,T,P)=−log⁡exp⁡(fiT⋅fiP/τ)exp⁡(fiT⋅fiP/τ)+∑j∈N,XjT≠XiTexp⁡(fiT⋅fjP/τ)l(i, T, P) = -\log \frac{\exp\left(f_i^T \cdot f_i^P / \tau\right)}{\exp\left(f_i^T \cdot f_i^P / \tau\right) + \sum_{j \in N, X_j^T \neq X_i^T} \exp\left(f_i^T \cdot f_j^P / \tau\right)} L(T,P)=1N∑i=1Nl(i,T,P)L(T, P) = \frac{1}{N} \sum_{i=1}^N l(i, T, P)

    2. Instance-Level Image-3D Alignment Objective: l(i,I,P)=−log⁡exp⁡(fiI⋅fiP/τ)exp⁡(fiI⋅fiP/τ)+∑j∈N,j≠iexp⁡(fiI⋅fjP/τ)l(i, I, P) = -\log \frac{\exp\left(f_i^I \cdot f_i^P / \tau\right)}{\exp\left(f_i^I \cdot f_i^P / \tau\right) + \sum_{j \in N, j \neq i} \exp\left(f_i^I \cdot f_j^P / \tau\right)} L(I,P)=1N∑i=1Nl(i,I,P)L(I, P) = \frac{1}{N} \sum_{i=1}^N l(i, I, P)

    3. Combined Cross-Modal Objective: LCM(T,I,P)=λ1L(T,P)+λ2L(I,P)L_{CM}(T, I, P) = \lambda_1 L(T, P) + \lambda_2 L(I, P) where λ1=0.5\lambda_1 = 0.5 and λ2=0.5\lambda_2 = 0.5 balance the semantic and instance-level alignments.

  4. Knowl 4 — Zero-Shot 3D Object Recognition Formulation

    model/method

    Zero-shot classification of 3D point cloud objects is conducted by measuring the similarity between the learned 3D point cloud representation and pre-trained language embeddings representing arbitrary target categories.

    Given a set of KK candidate category names {c1,c2,…,cK}\{c_1, c_2, \dots, c_K\}, each category is inserted into a prompt template (e.g., "point cloud of a {classname}."). The frozen CLIP text encoder EθTE_\theta^T generates a normalized text feature matrix FK∈RK×CF_K \in \mathbb{R}^{K \times C}.

    For a query 3D point cloud instance, the trained point cloud encoder EθPE_\theta^P produces a normalized embedding fP∈R1×Cf^P \in \mathbb{R}^{1 \times C}. The zero-shot classification prediction is obtained via:

    logits=softmax(fP(FK)T)\text{logits} = \text{softmax}\left(f^P (F_K)^T\right)

    The predicted class is the category index corresponding to arg⁡max⁡(logits)\arg\max(\text{logits}).

  5. Knowl 5 — Zero-Shot 3D Recognition Performance on SUN RGB-D and ScanNet

    data/table

    Zero-shot recognition Top-1 accuracy (%) across categories on indoor scene datasets SUN RGB-D (10 classes) and ScanNet (17 classes). Baseline methods include PointCLIP and Clip2Point, evaluated with their original pretraining and when retrained using the Triplet Proxy (TP.) set. CLIP2^2 w/ En. denotes inference with multi-modal ensembling.

    Method Avg. Bed Bookshelf Chair Desk Sofa Table Toilet Bathtub Dresser Night Stand
    PointCLIP 11.5 0.0 94.0 0.0 0.0 0.0 14.7 0.0 0.0 6.1 0.0
    Clip2Point 18.6 10.9 20.6 64.3 34.4 13.8 14.1 26.2 0.0 1.4 0.0
    PointCLIP w/ TP. 38.0 45.3 100.0 62.5 48.5 44.4 4.8 55.2 16.3 3.3 0.0
    Clip2Point w/ TP. 56.9 78.0 87.6 36.2 36.6 64.7 37.4 82.1 77.5 67.6 1.2
    CLIP2^2 61.3 84.0 75.5 70.7 47.3 75.5 33.8 86.2 65.3 71.8 2.4
    CLIP2^2 w/ En. 69.6 87.3 94.3 70.7 54.1 79.3 47.0 91.7 85.7 82.2 3.6

    On ScanNet (17 classes), the average Top-1 accuracies are:

    • PointCLIP: 6.3% (26.1% w/ TP.)
    • Clip2Point: 24.9% (35.2% w/ TP.)
    • CLIP2^2: 38.5%

    Pretraining with the Triplet Proxy dataset improves both depth-based baselines, while CLIP2^2's direct point cloud representation achieves the highest standalone zero-shot recognition accuracy.

  6. Knowl 6 — Zero-Shot 3D Recognition Performance on Outdoor LiDAR Datasets (nuScenes and ONCE)

    data/table

    Zero-shot recognition Top-1 accuracy (%) on outdoor driving benchmarks: nuScenes (10 classes) and ONCE (5 classes). Avg. denotes the mean Top-1 accuracy across all categories across both benchmarks.

    Avg. nuScenes (10 classes) ONCE (5 classes)
    Method Car Truck Bus Ped. Bicy. Trail. C.V. Motor. Barr. T.C. Car Cyc. Ped. Truck Bus
    PointCLIP 11.7 0.0 0.0 0.0 29.1 41.8 3.4 0.0 0.1 42.5 0.0 0.0 13.8 79.2 0.2 7.61
    Clip2Point 12.4 0.4 0.3 0.1 13.5 31.0 1.8 9.3 1.5 66.2 0.3 17.3 11.8 95.4 35.7 4.0
    PointCLIP w/ TP. 28.3 18.8 0.0 5.5 74.0 17.9 57.0 1.9 4.5 2.1 29.7 51.8 9.2 99.8 5.0 46.8
    Clip2Point w/ TP. 33.0 26.7 16.8 51.2 45.2 15.8 13.9 20.0 5.7 10.5 34.2 39.4 27.8 95.5 40.6 51.7
    CLIP2^2 37.8 41.9 41.3 22.5 40.3 21.1 20.6 24.8 22.4 17.3 35.3 52.7 27.3 77.7 78.5 44.0

    *(Ped.: Pedestrian, Bicy./Cyc.: Bicycle/Cyclist, Trail.: Trailer, C.V.: Construction Vehicle, Motor.: Motorcycle, Barr.: Barrier, T.C.: Traffic Cone).

    On nuScenes alone, CLIP2^2 attains 28.8% Top-1 accuracy, and on ONCE it achieves 56.0%, outperforming depth-projection methods where LiDAR point sparsity causes substantial geometric information loss.

  7. Knowl 7 — Zero-Shot and Few-Shot Classification Results on ScanObjectNN

    data/table

    Classification performance on the realistic single-object benchmark ScanObjectNN under zero-shot and KK-way NN-shot few-shot settings.

    Method Zero-Shot 15-way (N-shot) 10-way (N-shot)
    Acc. (%) 4-S 8-S 16-S 10-S 20-S
    PointNet++ - 41.0 47.6 55.0 - -
    PointCLIP 15.4 46.0 50.0 55.6 - -
    Clip2Point 23.3 - - - - -
    CrossPoint - - - - 58.7 ±\pm 1.8 64.6 ±\pm 1.2
    CLIP2^2 39.1 51.3 59.6 62.5 60.6 ±\pm 2.5 66.3 ±\pm 3.2

    Under zero-shot transfer, CLIP2^2 reaches 39.1% accuracy, providing a 16.1% relative improvement over prior state-of-the-art representations. In few-shot classification, CLIP2^2 outperforms supervised PointNet++ and self-supervised models (PointCLIP, CrossPoint) across all tested sample regimes.

  8. Knowl 8 — Open-Vocabulary 3D Recognition and Localization Capabilities

    empirical result

    CLIP2^2 demonstrates vocabulary expansion and zero-shot 3D localization without using annotated 3D bounding boxes:

    1. Vocabulary Expansion (ScanNet): Evaluating instance Top-5 accuracy when the candidate vocabulary is enlarged:

      • 384 noisy classes: PointCLIP achieves 0.3%, Clip2Point achieves 6.4%, and CLIP2^2 achieves 22.0%.
      • 249 merged classes: PointCLIP achieves 0.4%, Clip2Point achieves 7.0%, and CLIP2^2 achieves 31.7%.
    2. Indoor Zero-Shot Localization (SUN RGB-D): Evaluating 3D proxy bounding boxes on unseen classes:

      • Supervised 3DETR: mAP25=1.3%\text{mAP}_{25} = 1.3\%.
      • OV3D (trained on seen classes): mAP25=13.0%\text{mAP}_{25} = 13.0\%, AR25=37.7%\text{AR}_{25} = 37.7\%.
      • CLIP2^2 (unsupervised proxies): mAP25=12.7%\text{mAP}_{25} = 12.7\%, AR25=43.0%\text{AR}_{25} = 43.0\%, and mIoU=52.8%\text{mIoU} = 52.8\%, yielding a +5.3%+5.3\% improvement in AR25\text{AR}_{25} over OV3D without 3D annotations.
    3. Outdoor Zero-Shot Localization (nuScenes): Matching predicted 3D proxy centers against ground truth within a distance threshold λ=2 m\lambda = 2\,\text{m} yields Recall =87.4%= 87.4\% and Precision =6.7%= 6.7\%. The low precision is driven by detection of unannotated open-world objects (such as road debris, plastic bags, and tires).

  9. Knowl 9 — Ablation of Pretraining Objectives and Representation Ensembling

    empirical result

    Ablations on training objectives and modality ensembling demonstrate the impact of different cross-modal alignments:

    1. Pretraining Objectives Comparison:

      • Depth + Lang-Depth: 38.0% (SUN RGB-D), 9.3% (nuScenes), 31.8% (ScanObjectNN).
      • Depth + Image-Depth: 56.4% (SUN RGB-D), 21.7% (nuScenes), 39.0% (ScanObjectNN).
      • Point Cloud + Lang-Point: 52.6% (SUN RGB-D), 26.8% (nuScenes), 37.7% (ScanObjectNN).
      • Point Cloud + Image-Point: 56.5% (SUN RGB-D), 24.4% (nuScenes), 37.8% (ScanObjectNN).
      • Point Cloud + Joint Lang-Image-Point (LCML_{CM}): 61.3% (SUN RGB-D), 28.8% (nuScenes), 39.4% (ScanObjectNN).
    2. Representation Ensembling (Inference Logit Summation):

      • Point Cloud (PC) only: 61.3% indoor, 28.8% outdoor.
      • Image only: 64.2% indoor, 41.1% outdoor.
      • Depth only: 56.9% indoor, 23.9% outdoor.
      • PC + Image: 68.7% indoor, 43.9% outdoor.
      • PC + Depth: 64.8% indoor, 30.4% outdoor.
      • PC + Image + Depth: 69.6% indoor, 42.3% outdoor.

    While depth features slightly improve indoor accuracy when combined (+0.9%), they degrade outdoor accuracy by 1.6% due to LiDAR projection information loss. The pure 3D point cloud representation remains robust when 2D images are absent.

  10. Knowl 10 — Limitations in 3D Bounding Box Tightness and Precision

    limitation

    While CLIP2^2 enables zero-shot 3D localization and recognition via proxy generation without 3D annotations, it presents two primary limitations:

    1. Coarse Bounding Boxes: The 3D proxies generated via 2D frustum extrusion and spatial clustering (e.g., GrabCut or DBSCAN) produce estimated maximum bounding boxes rather than oriented, tight bounding boxes fitted by standard 3D object detectors.
    2. Precision on Benchmark Evaluation: Because the open-vocabulary detector identifies realistic obstacles and long-tail objects not present in closed benchmark annotation taxonomies (e.g., plastic bags, vehicle tires, small debris), the evaluated precision is low despite high recall.

Coverage note — None was omitted; all contributed models, loss formulations, proxy creation algorithms, benchmark evaluations, ablation studies, and limitations are fully covered.

References

  1. 1.Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri, Kanchana Thilakarathna, and Ranga Rodrigo. Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding. In CVPR, 2022. 2, 3, 7, 8
  2. 2.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020. 2, 4, 5, 6, 7, 8
  3. 3.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 8
  4. 4.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019. 5, 6
  5. 5.Ali Cheraghian, Shafin Rahman, Dylan Campbell, and Lars Petersson. Mitigating the hubness problem for zero-shot learning of 3d objects. In BMVC, 2019. 2, 3
  6. 6.Ali Cheraghian, Shafin Rahman, Dylan Campbell, and Lars Petersson. Transductive zero-shot learning for 3d point cloud classification. In WACV, 2020. 2, 3
  7. 7.Ali Cheraghian, Shafin Rahman, Townim F Chowdhury, Dylan Campbell, and Lars Petersson. Zero-shot learning on 3d point cloud objects and beyond. IJCV, 2022. 2
  8. 8.Ali Cheraghian, Shafin Rahman, and Lars Petersson. Zero-shot learning of 3d point cloud objects. In MVA. IEEE, 2019. 2
  9. 9.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 2, 4, 5
  10. 10.Zhipeng Ding, Xu Han, and Marc Niethammer. Votenet: A deep learning label fusion method for multi-atlas segmentation. In MICCAI, 2019. 1, 6
  11. 11.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, pages 5356–5364, 2019. 4
  12. 12.Huy Ha and Shuran Song. Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models. In CoRL, 2022. 2
  13. 13.Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo. Clip2point: Transfer clip to point cloud classification with image-depth pre-training. arXiv preprint arXiv:2210.01055, 2022. 2, 3, 4, 5, 6, 7, 8
  14. 14.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML. PMLR, 2021. 2, 3
  15. 15.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. NeurIPS, 2021. 2
  16. 16.Hanxue Liang, Chenhan Jiang, Dapeng Feng, Xin Chen, Hang Xu, Xiaodan Liang, Wei Zhang, Zhenguo Li, and Luc Van Gool. Exploring geometry-aware contrast and clustering harmonization for self-supervised 3d object detection. In ICCV, 2021. 3
  17. 17.Haotian Liu, Mu Cai, and Yong Jae Lee. Masked discrimination for self-supervised learning on point clouds. In ECCV, 2022. 3
  18. 18.Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, and Shanghang Zhang. Open-vocabulary 3d detection via image-level class and debiased cross-modal contrastive learning. arXiv preprint arXiv:2207.01987, 2022. 2, 6
  19. 19.Jiageng Mao, Minzhe Niu, Chenhan Jiang, Hanxue Liang, Jingheng Chen, Xiaodan Liang, Yamin Li, Chaoqiang Ye, Wei Zhang, Zhenguo Li, et al. One million scenes for autonomous driving: Once dataset. arXiv preprint arXiv:2106.11037, 2021. 2, 4, 6
  20. 20.Jiageng Mao, Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. 3d object detection for autonomous driving: A review and new outlooks. arXiv preprint arXiv:2206.09474, 2022. 1
  21. 21.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In ICCV, 2021. 6
  22. 22.Anshul Paigwar, David Sierra-Gonzalez, Özgür Erkent, and Christian Laugier. Frustum-pointpillars: A multi-stage approach for 3d object detection using rgb camera and lidar. In ICCV, 2021. 4
  23. 23.Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. In ECCV, 2022. 3
  24. 24.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In CVPR, 2018. 4
  25. 25.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In CVPR, 2017. 1, 2, 7
  26. 26.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. NeurIPS, 2017. 5, 7
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 3, 4, 5
  28. 28.Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. "grabcut" interactive foreground extraction using iterated graph cuts. TOG, 2004. 4
  29. 29.Erich Schubert, Jörg Sander, Martin Ester, Hans Peter Kriegel, and Xiaowei Xu. Dbscan revisited, revisited: why and how you should (still) use dbscan. TODS, 42(3):1–21, 2017. 4
  30. 30.Charu Sharma and Manohar Kaul. Self-supervised few-shot learning on point clouds. NeurIPS, 2020. 3
  31. 31.Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. Pointrcnn: 3D object proposal generation and detection from point cloud. In CVPR, 2019. 1
  32. 32.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015. 2, 4, 5, 6, 7, 8
  33. 33.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020. 5
  34. 34.Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. In ICCV, 2019. 2, 5, 6, 7, 8
  35. 35.Kashi Venkatesh Vishwanath, Diwaker Gupta, Amin Vahdat, and Ken Yocum. Modelnet: Towards a datacenter emulation environment. In 2009 IEEE Ninth International Conference on Peer-to-Peer Computing. IEEE, 2009. 6
  36. 36.Hanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby, and Matt J Kusner. Unsupervised point cloud pre-training via occlusion completion. In ICCV, 2021. 2
  37. 37.Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany. Pointcontrast: Unsupervised pre-training for 3d point cloud understanding. In ECCV, 2020. 3
  38. 38.Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, and Hang Xu. Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection. In NeurIPS, 2022. 2, 4, 5
  39. 39.Yangyang Ye, Houjin Chen, Chi Zhang, Xiaoli Hao, and Zhaoxiang Zhang. Sarpnet: Shape attention regional proposal network for lidar-based 3d object detection. Neurocomputing, 2020. 3
  40. 40.Tianwei Yin, Xingyi Zhou, and Philipp Krähenbühl. Center-based 3d object detection and tracking. arXiv preprint arXiv:2006.11275, 2020. 1
  41. 41.Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In CVPR, 2022. 3
  42. 42.Renrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang, Bin Zhao, Dong Wang, Yu Qiao, and Hongsheng Li. Point-m2ae: Multi-scale masked autoencoders for hierarchical point cloud pre-training. arXiv preprint arXiv:2205.14401, 2022. 3
  43. 43.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In CVPR, 2022. 2, 3, 4, 5, 6, 7, 8
  44. 44.Zaiwei Zhang, Bo Sun, Haitao Yang, and Qixing Huang. H3dnet: 3d object detection using hybrid geometric primitives. In ECCV, 2020. 1, 6

Citation

MLA
Zeng, Y., et al. “CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data”. arXiv, 2023, http://arxiv.org/abs/2303.12417v2.
APA
Zeng, Y., Jiang, C., Mao, J., Han, J., Ye, C., Huang, Q., Yeung, D.-Y., Yang, Z., Liang, X., & Xu, H. (2023). CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data. arXiv. http://arxiv.org/abs/2303.12417v2
Chicago
Zeng, Y., C. Jiang, J. Mao, et al. 2023. “CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data”. arXiv. http://arxiv.org/abs/2303.12417v2.
Harvard
Zeng, Y. et al. (2023) “CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.12417v2.
Vancouver
1. Zeng Y, Jiang C, Mao J, Han J, Ye C, Huang Q, Yeung D-Y, Yang Z, Liang X, Xu H (2023) CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data. arXiv

BibTeX

@article{zeng2023clip,
  title = {CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data},
  author = {Zeng, Yihan and Jiang, Chenhan and Mao, Jiageng and Han, Jianhua and Ye, Chaoqiang and Huang, Qingqiu and Yeung, Dit-Yan and Yang, Zhen and Liang, Xiaodan and Xu, Hang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.12417v2},
  eprint = {2303.12417}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE