ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding

Le XueMingfei GaoChen XingRoberto Martín-MartínJiajun WuCaiming XiongRan XuJuan Carlos NieblesSilvio Savarese

article2023CVPR434 citations

Introduces a model-agnostic pre-training framework that aligns 3D point cloud encoders with frozen vision-language models using synthesized multimodal triplets, substantially boosting zero-shot and standard 3D recognition performance across various architectures.

Listen

Three-dimensional computer vision is critical for emerging technologies such as autonomous driving, robotics, and augmented reality. However, real-world 3D models face severe performance limitations because existing 3D datasets are small and cover only narrow, pre-defined categories. Collecting and labeling 3D spatial data is expensive and time-consuming, unlike 2D image domains that benefit from massive, diverse datasets.

The article introduces and evaluates ULIP (Unified Representation of Language, Images, and Point Clouds), a multimodal pre-training framework designed to enhance 3D visual understanding. The primary objective is to demonstrate that aligning 3D point cloud models with rich visual and textual concepts from existing vision-language models can substantially boost 3D recognition performance, even with minimal training data.

To achieve this without requiring extensive manual annotation, the framework leverages pre-trained vision-language models—specifically CLIP and its variant SLIP—which already understand millions of image-text relationships. The authors synthesized training triplets consisting of point clouds, computer-rendered multi-view images, and descriptive text prompts from roughly 52,500 models in the ShapeNet55 repository. While keeping the image and text encoders frozen to prevent catastrophic forgetting, the authors trained various standard 3D encoders to align spatial data into the shared image-text feature space using contrastive learning. The resulting models were evaluated across standard classification, zero-shot classification, and cross-modal retrieval tasks on the ModelNet40 and ScanObjectNN benchmarks.

The evaluation produced four key findings. First, ULIP established new state-of-the-art results across standard 3D classification tasks, boosting baseline models such as PointMLP and PointBERT by approximately 3 percentage points on real-world scanned objects. Second, in zero-shot classification—where models identify novel objects without specific training—ULIP outperformed the previous leading method, PointCLIP, by roughly 29 to 35 percentage points on top-1 accuracy across synthetic and scanned benchmarks. Third, ablation analyses confirmed that jointly aligning all three modalities (point clouds, images, and text) consistently produced superior results compared to aligning point clouds with only images or only text. Fourth, ULIP significantly enhanced data efficiency, allowing 3D models to achieve higher accuracy with only a fraction of the downstream labeled training data.

These findings have direct operational and economic implications. Organizations developing 3D applications can achieve higher recognition accuracy and generalize to rare, unseen categories without undertaking costly 3D data collection and labeling campaigns. Because ULIP modifies only the pre-training phase and does not alter the underlying 3D network architecture, it introduces zero additional computational latency or overhead during inference in production environments.

Based on these results, engineering teams should consider adopting ULIP pre-training as a plug-and-play enhancement for existing 3D architectures, such as PointNet++, PointBERT, and PointMLP. Organizations aiming to implement cross-modal functionality, such as image-to-3D object retrieval, can also utilize this unified representation space. As next steps, technical teams should pilot pre-trained ULIP models on internal, domain-specific 3D data to measure real-world performance gains.

The findings are supported by consistent results across standard academic benchmarks, providing high confidence in the methodology. However, decision-makers should note that the pre-training triplets relied on synthetic CAD models and automated multi-view image rendering rather than raw, noisy sensor scans. Validating performance on specialized, complex real-world point cloud environments remains an important consideration before full deployment.

Cover for ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding

Abstract

The recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from other modalities, such as language. Inspired by this, leveraging multimodal information for 3D modality could be promising to improve 3D understanding under the restricted data regime, but this line of research is not well studied. Therefore, we introduce ULIP to learn a unified representation of image, text, and 3D point cloud by pre-training with object triplets from the three modalities. To overcome the shortage of training triplets, ULIP leverages a pre-trained vision-language model that has already learned a common visual and textual space by training with massive image-text pairs. Then, ULIP learns a 3D representation space aligned with the common image-text space, using a small number of automatically synthesized triplets. ULIP is agnostic to 3D backbone networks and can easily be integrated into any 3D architecture. Experiments show that ULIP effectively improves the performance of multiple recent 3D backbones by simply pre-training them on ShapeNet55 using our framework, achieving state-of-the-art performance in both standard 3D classification and zero-shot 3D classification on ModelNet40 and ScanObjectNN. ULIP also improves the performance of PointMLP by around 3% in 3D classification on ScanObjectNN, and outperforms PointCLIP by 28.8% on top-1 accuracy for zero-shot 3D classification on ModelNet40. Our code and pre-trained models will be released.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Learning Unified Representation of Language, Image and Point Cloud
  • 3.1. Creating Training Triplets for ULIP
  • 3.2. Aligning Representations of Three Modalities
  • 4. Experiments
  • 4.1. 3D Backbone Networks
  • 4.2. Downstream Datasets
  • 4.3. Implementation Details
  • 4.4. Standard 3D Classification
  • 4.5. Zero-Shot 3D Classification
  • 4.6. Analyses
  • 4.7. Cross-Modal Retrieval
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Unified image–text–point-cloud pre-training

    model/method

    ULIP (Learning a Unified Representation of Language, Images, and Point Clouds) pre-trains a 3D point-cloud encoder so that its representations occupy the same feature space as image and language representations from a pre-trained vision–language model. Each training example contains an image, a text description, and a point cloud corresponding to the same 3D object. The image and text encoders are kept fixed, while the chosen 3D backbone is optimized to match the two pre-aligned CLIP-family representations through contrastive learning. The resulting 3D encoder can be fine-tuned for ordinary labeled 3D classification or used directly with text representations for zero-shot 3D classification. ULIP does not require a particular 3D architecture and can be applied to PointNet++, PointMLP, PointBERT, and other point-cloud backbones.

  2. Knowl 2 — Automatically synthesized training triplets

    model/method

    ULIP constructs training triplets from approximately 52.5K ShapeNet55 CAD models without manual triplet annotation. For CAD model ii, the triplet is Ti=(Ii,Si,Pi)T_i=(I_i,S_i,P_i), where IiI_i is a randomly selected rendered image or depth map, SiS_i is a set of prompted text descriptions, and PiP_i is a sampled point cloud. The point cloud is obtained by uniformly sampling NpN_p points from the original CAD point cloud and applying random point dropping, scaling, translation, and rotation during training. Each CAD model is rendered from viewpoints spaced every 1212 degrees, producing 30 RGB images and 30 depth maps; one of these 60 candidates is sampled at each pre-training iteration. The model metadata supplies a taxonomy word, which is inserted into 63 standard image–text prompts plus one 3D-specific prompt, yielding 64 text descriptions. The metadata word is randomly selected during each iteration, and the 64 resulting text embeddings are averaged.

  3. Knowl 3 — Cross-modal contrastive objective

    equation

    For a batch BB of matched objects, ULIP extracts image, text, and point-cloud embeddings hiI,hiS,hiP∈Rd\mathbf h_i^I,\mathbf h_i^S,\mathbf h_i^P\in\mathbb R^d for object ii, where II, SS, and PP denote image, sentence, and point-cloud modalities, respectively. For any ordered pair of modalities (M1,M2)(M_1,M_2), with a learnable temperature τ>0\tau>0, ULIP uses a symmetric in-batch contrastive loss:

    L(M1,M2)=∑i∈B[−12log⁡exp⁡((hiM1⋅hiM2)/τ)∑k∈Bexp⁡((hiM1⋅hkM2)/τ)−12log⁡exp⁡((hiM1⋅hiM2)/τ)∑k∈Bexp⁡((hkM1⋅hiM2)/τ)].L_{(M_1,M_2)}=\sum_{i\in B}\left[-\frac{1}{2}\log\frac{\exp\left((\mathbf h_i^{M_1}\cdot\mathbf h_i^{M_2})/\tau\right)}{\sum_{k\in B}\exp\left((\mathbf h_i^{M_1}\cdot\mathbf h_k^{M_2})/\tau\right)}-\frac{1}{2}\log\frac{\exp\left((\mathbf h_i^{M_1}\cdot\mathbf h_i^{M_2})/\tau\right)}{\sum_{k\in B}\exp\left((\mathbf h_k^{M_1}\cdot\mathbf h_i^{M_2})/\tau\right)}\right].

    The matched pair (i,i)(i,i) is positive, while other objects in the batch provide negatives. The total ULIP loss is

    Lfinal=αL(I,S)+βL(I,P)+θL(P,S).L_{\mathrm{final}}=\alpha L_{(I,S)}+\beta L_{(I,P)}+\theta L_{(P,S)}.

    The default coefficients are α=0\alpha=0 and β=θ=1\beta=\theta=1, so only image–point-cloud and text–point-cloud alignment losses update the 3D encoder. The image–text loss is unnecessary because the pre-trained vision–language model has already aligned those modalities.

  4. Knowl 4 — Frozen vision–language teacher and 3D-only optimization

    assumption

    During ULIP pre-training, the image encoder fIf_I and text encoder fSf_S of the pre-trained vision–language model remain frozen, and only the point-cloud encoder fPf_P is updated. The paper reports that updating the image and text encoders with the relatively small ShapeNet55 triplet set causes catastrophic forgetting and substantially degrades downstream performance. In the experiments, the frozen encoders come from SLIP, an advanced CLIP-family model, while the trainable component is the selected 3D backbone. Thus, ULIP transfers semantic structure learned from large-scale image–text data into 3D without requiring large-scale manually paired image–text–3D data.

  5. Knowl 5 — Representations used by ULIP

    equation

    For CAD object ii, the trainable point-cloud encoder fPf_P maps an augmented point cloud PiP_i containing NpN_p sampled 3D points to a feature vector, the frozen image encoder fIf_I maps a sampled RGB image or depth map IiI_i to a feature vector, and the frozen text encoder fSf_S maps the set SiS_i of 64 prompted descriptions to 64 feature vectors. ULIP defines

    hiP=fP(Pi),hiI=fI(Ii),hiS=Avg⁡(fS(Si)),\mathbf h_i^P=f_P(P_i),\qquad \mathbf h_i^I=f_I(I_i),\qquad \mathbf h_i^S=\operatorname{Avg}\bigl(f_S(S_i)\bigr),

    where hiP,hiI,hiS∈Rd\mathbf h_i^P,\mathbf h_i^I,\mathbf h_i^S\in\mathbb R^d and Avg⁡\operatorname{Avg} averages the 64 text-encoder outputs. The three vectors represent the same CAD object in the common embedding space used by the contrastive objective.

  6. Knowl 6 — Downstream standard and zero-shot classification

    model/method

    For standard 3D classification, ULIP first pre-trains a 3D backbone and then fine-tunes that unchanged backbone with labeled point clouds and a learned classification head. Because ULIP changes only the pre-training stage and does not add modules to the inference architecture, it introduces no additional inference latency relative to the original 3D backbone. For zero-shot 3D classification, the pre-trained 3D encoder is used without fine-tuning. For every candidate category cc, the frozen text encoder produces an embedding from the prompt a point cloud model of cc. The category whose text embedding is closest to the input point-cloud embedding in the shared feature space is selected as the prediction. This procedure allows the 3D encoder to classify categories without labeled examples from the downstream dataset.

  7. Knowl 7 — Experimental configuration

    experimental setup

    ULIP is evaluated with PointNet++, PointMLP, and PointBERT on ModelNet40 and ScanObjectNN. ModelNet40 contains 9,843 training CAD models and 2,468 test models from 40 categories. ScanObjectNN contains 2,902 real scanned objects from 15 categories and includes object-only, background-noise, and hardest variants; the experiments use the provided benchmark variants. During ULIP pre-training, the point-cloud input contains 1,024, 2,048, or 8,192 points depending on the backbone, training lasts 250 epochs, the batch size is 64, the learning rate is 10−310^{-3}, and AdamW is used. Pre-training uses eight A100 GPUs. Standard-classification fine-tuning uses one A100 GPU: on ModelNet40, PointNet++ uses learning rate 1.5×10−41.5\times10^{-4} for 200 epochs with batch size 24, while PointMLP uses learning rate 0.10.1 for 300 epochs with batch size 32; on ScanObjectNN, PointMLP uses learning rate 0.030.03 for 350 epochs and PointBERT uses learning rate 2×10−42\times10^{-4} for 300 epochs, both with batch size 32. Standard classification is measured by overall accuracy and class-mean accuracy; zero-shot classification is measured by top-1 and top-5 overall accuracy.

  8. Knowl 8 — Improvement on real-world ScanObjectNN classification

    data/table

    ULIP improves several 3D backbones on ScanObjectNN, a real-scanned-object benchmark. Overall accuracy is reported as OA and class-mean accuracy as mAcc. The strongest ULIP result uses PointMLP with 2,000 sampled points, whereas the other listed models use 1,000 points.

    Could not parse LaTeX table

    ULIP raises PointBERT by 3.3 OA points and PointMLP by 3.1 OA points in the 1,000-point setting. With 2,000 points, PointMLP + ULIP reaches 89.4 OA and 88.5 mAcc, exceeding the previous 86.0 OA best result by 3.4 points.

  9. Knowl 9 — Improvement on ModelNet40 classification

    data/table

    ULIP also improves standard classification on the easier synthetic ModelNet40 benchmark, where many recent methods already approach saturation near 94% overall accuracy. The key baseline and ULIP results are:

    Could not parse LaTeX table

    The asterisk denotes use of a voting technique. ULIP gives the largest OA gain to PointNet++ at 2.7 points, while PointMLP with ULIP obtains the best reported result in both OA and mAcc among the listed methods: 94.7 OA and 92.4 mAcc.

  10. Knowl 10 — Large zero-shot 3D classification gains

    data/table

    ULIP enables zero-shot classification by comparing point-cloud and category-text embeddings, without fine-tuning on the evaluation data. On ModelNet40, the paper evaluates three category sets: ALL contains all 40 categories; Medium removes categories whose exact names occur in ShapeNet55 pre-training metadata; Hard additionally removes categories with semantically similar counterparts in the pre-training categories. The results are:

    Could not parse LaTeX table

    On ScanObjectNN, the corresponding ALL-set results are:

    Could not parse LaTeX table

    The best PointBERT result improves over PointCLIP by 40.2 top-1 points on ModelNet40 ALL and by 28.8 points on the Hard set. The best ScanObjectNN result improves top-1 accuracy by 34.5 points. These Hard-set gains indicate that the zero-shot advantage is not explained solely by exact category overlap between ShapeNet55 and ModelNet40.

  11. Knowl 11 — Three-way alignment is better than aligning only two modalities

    empirical result

    An ablation compares aligning the point-cloud representation with text only (P+TP+T), image only (P+IP+I), or both image and text (P+I+TP+I+T). The three-modality objective consistently gives the best zero-shot results for both PointMLP and PointBERT.

    Could not parse LaTeX table

    The result supports using both complementary semantic sources: image alignment alone or text alignment alone is weaker than simultaneously aligning the 3D representation to the pre-aligned image–text space.

  12. Knowl 12 — Improved data efficiency during downstream fine-tuning

    empirical result

    ULIP pre-training improves performance when only a fraction of downstream labeled point clouds is available. In the paper's data-efficiency comparison, PointMLP and PointBERT are fine-tuned with varying percentages of training samples, and the ULIP-pre-trained versions achieve higher overall accuracy throughout the comparison, with especially clear gains in the low-data regime. PointBERT outperforms PointMLP when fewer than 20% of the downstream training samples are used. Although PointBERT already receives separate ShapeNet55 self-supervised pre-training, adding ULIP still produces a clear further improvement.

Coverage note — The qualitative real-image-to-point-cloud retrieval demonstration is omitted because it is a small-scale result without quantitative metrics; background, related work, and non-contributed material are also omitted.

References

  1. 1.Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1534–1543, 2016.
  2. 2.Cesar Cadena, Anthony R Dick, and Ian D Reid. Multimodal auto-encoders as joint estimators for robotics scene understanding. In Robotics: Science and systems, volume 5, 2016.
  3. 3.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015.
  4. 4.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer, 2020.
  5. 5.Zhimin Chen, Longlong Jing, Yang Liang, YingLi Tian, and Bing Li. Multimodal semi-supervised learning for 3d objects. arXiv preprint arXiv:2110.11601, 2021.
  6. 6.Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European conference on computer vision, pages 628–644. Springer, 2016.
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  9. 9.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pages 178–178. IEEE, 2004.
  10. 10.Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li, Ran Xu, Wenhao Liu, and Caiming Xiong. Open vocabulary object detection with pseudo bounding-box labels. arXiv preprint arXiv:2111.09452, 2021.
  11. 11.Ankit Goyal, Hei Law, Bowei Liu, Alejandro Newell, and Jia Deng. Revisiting point cloud shape classification with a simple and effective baseline. In International Conference on Machine Learning, pages 3809–3820. PMLR, 2021.
  12. 12.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9224–9232, 2018.
  13. 13.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921, 2021.
  14. 14.Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu. Pct: Point cloud transformer. Computational Visual Media, 7(2):187–199, 2021.
  15. 15.Abdullah Hamdi, Silvio Giancola, and Bernard Ghanem. Mvtn: Multi-view transformation network for 3d shape recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1–11, 2021.
  16. 16.Qingyong Hu, Bo Yang, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang, Niki Trigoni, and Andrew Markham. Randla-net: Efficient semantic segmentation of large-scale point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11108–11117, 2020.
  17. 17.Loic Landrieu and Martin Simonovsky. Large-scale point cloud semantic segmentation with superpoint graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4558–4567, 2018.
  18. 18.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34:9694–9705, 2021.
  19. 19.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019.
  20. 20.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10965–10975, 2022.
  21. 21.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer, 2020.
  22. 22.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. Pointcnn: Convolution on X-transformed points. arXiv preprint arXiv:1801.07791, 2018.
  23. 23.Yingwei Li, Adams Wei Yu, Tianjian Meng, Ben Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Yifeng Lu, Denny Zhou, Quoc V Le, et al. Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17182–17191, 2022.
  24. 24.Zhichao Li, Feng Wang, and Naiyan Wang. Lidar r-cnn: An efficient and universal 3d object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7546–7555, 2021.
  25. 25.Yongcheng Liu, Bin Fan, Gaofeng Meng, Jiwen Lu, Shiming Xiang, and Chunhong Pan. Densepoint: Learning densely contextual representation for efficient point cloud processing. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5239–5248, 2019.
  26. 26.Yongcheng Liu, Bin Fan, Shiming Xiang, and Chunhong Pan. Relation-shape convolutional neural network for point cloud analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8895–8904, 2019.
  27. 27.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2949–2958, 2021.
  28. 28.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019.
  29. 29.Xu Ma, Can Qin, Haoxuan You, Haoxi Ran, and Yun Fu. Rethinking network design and local geometry in point cloud: A simple residual mlp framework. arXiv preprint arXiv:2202.07123, 2022.
  30. 30.Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In 2015 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 922–928. IEEE, 2015.
  31. 31.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2906–2917, 2021.
  32. 32.Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie. Slip: Self-supervision meets language-image pre-training. In European Conference on Computer Vision, pages 529–544. Springer, 2022.
  33. 33.Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. arXiv preprint arXiv:2203.06604, 2022.
  34. 34.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2085–2094, 2021.
  35. 35.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017.
  36. 36.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017.
  37. 37.Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Abed Al Kader Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. arXiv preprint arXiv:2206.04670, 2022.
  38. 38.Shi Qiu, Saeed Anwar, and Nick Barnes. Geometric back-projection network for point cloud classification. IEEE Transactions on Multimedia, 24:1943–1955, 2021.
  39. 39.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  40. 40.Haoxi Ran, Jun Liu, and Chengjie Wang. Surface representation for point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18942–18952, 2022.
  41. 41.Haoxi Ran, Wei Zhuo, Jun Liu, and Li Lu. Learning inner-group relations on point clouds. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15477–15487, 2021.
  42. 42.Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10529–10538, 2020.
  43. 43.Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019.
  44. 44.Hugues Thomas, Charles R Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui, François Goulette, and Leonidas J Guibas. Kpconv: Flexible and deformable convolution for point clouds. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6411–6420, 2019.
  45. 45.Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.
  46. 46.Thang Vu, Kookhoi Kim, Tung M Luu, Thanh Nguyen, and Chang D Yoo. Softgroup for 3d instance segmentation on point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2708–2717, 2022.
  47. 47.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. Acm Transactions On Graphics (tog), 38(5):1–12, 2019.
  48. 48.Christian Wojek, Stefan Walk, Stefan Roth, and Bernt Schiele. Monocular 3d scene understanding with explicit occlusion reasoning. In CVPR 2011, pages 1993–2000. IEEE, 2011.
  49. 49.Bo Wu, Yang Liu, Bo Lang, and Lei Huang. Dgcnn: Disordered graph convolutional neural network based on the gaussian mixture model. Neurocomputing, 321:346–356, 2018.
  50. 50.Wenxuan Wu, Zhongang Qi, and Li Fuxin. Pointconv: Deep convolutional networks on 3d point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9621–9630, 2019.
  51. 51.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1912–1920, 2015.
  52. 52.Tiange Xiang, Chaoyi Zhang, Yang Song, Jianhui Yu, and Weidong Cai. Walk in the cloud: Learning curves for point clouds shape analysis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 915–924, 2021.
  53. 53.Chen Xing, Negar Rostamzadeh, Boris Oreshkin, and Pedro O O Pinheiro. Adaptive cross-modal few-shot learning. Advances in Neural Information Processing Systems, 32, 2019.
  54. 54.Mutian Xu, Junhao Zhang, Zhipeng Zhou, Mingye Xu, Xiaojuan Qi, and Yu Qiao. Learning geometry-disentangled representation for complementary understanding of 3d object point cloud. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3056–3064, 2021.
  55. 55.Yifan Xu, Tianqi Fan, Mingye Xu, Long Zeng, and Yu Qiao. Spidercnn: Deep learning on point sets with parameterized convolutional filters. In Proceedings of the European Conference on Computer Vision (ECCV), pages 87–102, 2018.
  56. 56.Xu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao, Ruimao Zhang, Shuguang Cui, and Zhen Li. Let images give you more: Point cloud cross-modal training for shape analysis. arXiv preprint arXiv:2210.04208, 2022.
  57. 57.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3d object detection and tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11784–11793, 2021.
  58. 58.Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19313–19322, 2022.
  59. 59.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8552–8562, 2022.
  60. 60.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16259–16268, 2021.

Citation

MLA
Xue, L., et al. “ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding”. arXiv, 2022, http://arxiv.org/abs/2212.05171v4.
APA
Xue, L., Gao, M., Xing, C., Martín-Martín, R., Wu, J., Xiong, C., Xu, R., Niebles, J. C., & Savarese, S. (2022). ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding. arXiv. http://arxiv.org/abs/2212.05171v4
Chicago
Xue, L., M. Gao, C. Xing, et al. 2022. “ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding”. arXiv. http://arxiv.org/abs/2212.05171v4.
Harvard
Xue, L. et al. (2022) “ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.05171v4.
Vancouver
1. Xue L, Gao M, Xing C, Martín-Martín R, Wu J, Xiong C, Xu R, Niebles JC, Savarese S (2022) ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding. arXiv

BibTeX

@article{xue2022ulip,
  title = {ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding},
  author = {Xue, Le and Gao, Mingfei and Xing, Chen and Martín-Martín, Roberto and Wu, Jiajun and Xiong, Caiming and Xu, Ran and Niebles, Juan Carlos and Savarese, Silvio},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.05171v4},
  eprint = {2212.05171}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE