GPT4Point: A Unified Framework for Point-Language Understanding and Generation

Zhangyang QiYe FangZeyi SunXiaoyang WuTong WuJiaqi WangDahua LinHengshuang Zhao

article2024CVPR94 citations

Presents GPT4Point, a unified point-language framework that connects point cloud features with large language models and diffusion models to perform both 3D object-level understanding and controllable text-to-3D generation.

Listen

While multimodal artificial intelligence systems have achieved significant success in interpreting and generating two-dimensional images and text, they struggle to accurately understand the physical geometry of three-dimensional objects. Most existing methods rely on 2D image projections or focus solely on broad scene coordinates, leading to a substantial loss of 3D spatial fidelity and an inability to detect physical anomalies. The article introduces GPT4Point, a unified framework designed to bridge language processing directly with 3D point cloud data for both multimodal comprehension and controlled 3D object generation.

To overcome the severe scarcity of high-quality paired 3D data and text, the authors developed Pyramid-XL, an automated data annotation engine that created over one million progressively detailed 3D text pairs from the Objaverse-XL dataset. Using this data, the GPT4Point framework aligns point clouds and text features via a Bert-based transformer without altering frozen large language models, significantly reducing training overhead. Aligned features are then channeled into a language model for text reasoning or into a diffusion pipeline for generating refined 3D objects conditioned on coarse point inputs.

The evaluation demonstrates that GPT4Point outperforms leading 2D vision-language and dedicated 3D models across multiple benchmarks. In zero-shot 3D object classification on ModelNet40, GPT4Point achieves a top-1 accuracy of 43.90%, outperforming InstructBLIP by 12.42 points and PointLLM by 2.57 points. In 3D question answering, it attains a zero-shot accuracy of 27.6%, exceeding InstructBLIP by more than 11 points and PointLLM by 4.2 points, while maintaining competitive scores across standard captioning metrics. Furthermore, in controllable text-to-3D generation, conditioning on GPT4Point features yields higher fidelity and geometric consistency than direct text-to-3D or single-image baselines, achieving lower error scores and superior human evaluation ratings.

These findings indicate that integrating native 3D geometry into language models provides superior spatial reasoning and quality control compared to traditional 2D-reliant approaches. This capability is critical for reducing failure rates and risks in downstream applications like autonomous robotics, spatial computing, and automated 3D asset production. Organizations evaluating multimodal AI should consider native 3D alignment strategies over purely image-based proxies for geometry-sensitive workflows. Future research should expand the framework beyond isolated objects to complex, multi-object indoor and outdoor environments, and explore end-to-end interactive editing pipelines.

arXiv: 2312.02980
Cover for GPT4Point: A Unified Framework for Point-Language Understanding and Generation

Abstract

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation. Still, their understanding of the 3D world needs to be improved, limiting progress in 3D language understanding and generation. To solve this problem, we introduce GPT4Point, an innovative, groundbreaking point-language multimodal model explicitly designed for unified 3D object understanding and generation within the MLLM framework. GPT4Point, as a powerful 3D MLLM, can seamlessly execute point-text reference tasks such as point-cloud captioning and Q&A. Additionally, GPT4Point is equipped with advanced capabilities for controllable 3D generation, and it can get high-quality results through a low-quality point-text feature that maintains geometric shapes and colors. We develop Pyramid-XL, a point-language dataset annotation engine, to support the expansive needs of 3D object-text pairs. It constructs a large-scale database of over 1M objects of varied text granularity levels from the Objaverse-XL dataset, essential for training GPT4Point. A comprehensive benchmark has been proposed to evaluate 3D point-language understanding capabilities. In extensive evaluations, GPT4Point has demonstrated superior performance in understanding and generation.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. Point-Language Dataset Annotation Engine
  • 3.2. Model Architecture
  • 4. Benchmarks and Evaluation
  • 4.1. Composition of Test Set
  • 4.2. 3D Object Recognition
  • 4.3. 3D Object Text Inference
  • 5. Experiments
  • 5.1. Training Details
  • 5.2. Evaluation and Diverse Tasks
  • 5.3. Assessing the Effectiveness of Pyramid-XL
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — GPT4Point Two-Stage Framework for 3D Multimodal Understanding and Generation

    model/method

    GPT4Point is a unified multimodal framework designed to perform 3D point cloud understanding (such as point-text retrieval, zero-shot classification, 3D captioning, and 3D question answering) and controllable text-to-3D generation. The architecture operates in two training stages:

    1. Stage 1 (Point-Text Feature Alignment): A point cloud P∈RN×6P \in \mathbb{R}^{N \times 6} (with NN points containing 3D XYZ coordinates and RGB color values) is encoded by a Point-BERT backbone E\mathcal{E} to produce point feature tokens T1p=E(P)T_1^p = \mathcal{E}(P). An accompanying text description is tokenized into text tokens T1tT_1^t. Both are fed into a BERT-based Point-QFormer FQF_Q which fuses point and textual representations. The model is jointly trained on three objectives: Point-Text Contrast (PTC), Point-Text Matching (PTM), and Point-Text Generation / Caption Generation (PTG).

    2. Stage 2 (Point Understanding and Generation):

    • Understanding Branch (3D MLLM): The point cloud is processed by E\mathcal{E} and FQF_Q to extract aligned point tokens T2p=FQ(E(P))T_2^p = F_Q(\mathcal{E}(P)), which are projected via a fully connected layer into the input space of a frozen Large Language Model (such as OPT or Flan-T5). The LLM is trained on point captioning and dialogue tasks.
    • Diffusion Branch (Controllable Text-to-3D): The aligned point features from FQF_Q are projected via a linear layer into the CLIP token embedding space (T3pT_3^p) and concatenated with text token embeddings (T3tT_3^t) generated by a CLIP tokenizer. The combined tokens are processed by a CLIP text encoder to condition a point cloud diffusion model (Point-E) for generating high-fidelity point clouds preserving geometry and color from low-quality point inputs.
  2. Knowl 2 — Pyramid-XL Automated Point-Language Dataset Annotation Engine

    model/method

    Pyramid-XL is a three-level hierarchical annotation pipeline designed to generate point-text pairs from 3D object models in datasets like Objaverse and Objaverse-XL, avoiding the errors Vision-Language Models (VLMs) make when directly parsing multi-view images of 3D objects:

    • Level 1 (Single-View Caption): BLIP-2 generates concise descriptions (~10 words) from individual 2D rendered views of the 3D object.
    • Level 2 (Multi-View Synthesis Caption): GPT-4 synthesizes multiple Level 1 descriptions across viewpoints into a coherent multi-view description (~30 words).
    • Level 3 (VLM Instruction Caption and QA Pair): The single view with the highest CLIP text-image similarity score relative to the Level 2 description is selected. An advanced VLM (Qwen-VL) generates detailed dense captions (~50 words) and object-specific instruction question-answer pairs.

    Pyramid-XL produced over 1,000,000 objects annotated with Level 1 captions, 660,000 objects with Level 2 captions, and 70,000 objects with Level 3 dense captions and QA pairs.

  3. Knowl 3 — Stage-Wise Training Objectives and Loss Functions of GPT4Point

    equation

    GPT4Point uses separate objective functions across its two training stages:

    Stage 1 Loss (L1L_1): Jointly optimizes Point-Text Contrast (PTC), Point-Text Matching (PTM), and Point Caption Generation (PTG) using the Point-QFormer FQF_Q with equal loss weight ratios (1:1:1):

    L1=FQ(T1p,T1t)=FQ(E(P),T1t)L_1 = F_Q(T_1^p, T_1^t) = F_Q(\mathcal{E}(P), T_1^t)

    where P∈RN×6P \in \mathbb{R}^{N \times 6} denotes the input point cloud with NN points (XYZ coordinates and RGB colors), E\mathcal{E} is the point encoder yielding point tokens T1pT_1^p, and T1tT_1^t denotes the text tokens.

    Stage 2 Loss (L2L_2): Trains the language model branch for text comprehension and generation using Point Caption loss alone:

    L2=FLLM(T2p,T2t)=FLLM(FQ(E(P)),T2t)L_2 = F_{\text{LLM}}(T_2^p, T_2^t) = F_{\text{LLM}}(F_Q(\mathcal{E}(P)), T_2^t)

    where T2p=FQ(E(P))T_2^p = F_Q(\mathcal{E}(P)) represents semantically aligned point features projected to match the LLM token dimension, T2tT_2^t represents target text tokens from the LLM tokenizer, and FLLMF_{\text{LLM}} represents the autoregressive language modeling loss.

  4. Knowl 4 — Controllable Text-to-3D Generation Mechanism via Aligned Point-Text Embeddings

    model/method

    GPT4Point performs controllable text-to-3D generation by conditioning a feed-forward point cloud diffusion model (Point-E) on aligned point-text embeddings:

    1. A coarse or low-quality point cloud PP is passed through the pretrained point encoder E\mathcal{E} and frozen Point-QFormer FQF_Q.
    2. A single fully connected layer projects the output tokens into the CLIP token embedding space, denoted as T3pT_3^p.
    3. T3pT_3^p is concatenated with original text embeddings T3tT_3^t obtained from the text prompt via the CLIP tokenizer.
    4. The concatenated sequence is fed through the CLIP text encoder to generate conditioned guidance embeddings.
    5. The frozen Point-E diffusion model generates high-quality point clouds matching the target text prompt while strictly preserving the geometric structure and color distribution specified by the low-quality point condition.
  5. Knowl 5 — ObjaverseXL-LVIS 3D Point Captioning and Question Answering Performance

    data/table

    GPT4Point was evaluated against 2D Vision-Language Models (BLIP-2, InstructBLIP, Qwen-VL) and 3D MLLMs (PointLLM) on the 1K ObjaverseXL-LVIS point cloud captioning and question answering test sets.

    Model #Trainable Params ObjaverseXL-LVIS Caption (1K) ObjaXL-LVIS QA (1K)
    BLEU1 BLEU4 METEOR ROUGE CIDEr Acc BLEU1 ROUGE
    BLIP-2 (OPT 2.7B) 188M 22.2 3.0 10.3 28.2 32.3 13.4 14.2 16.8
    BLIP-2 (OPT 6.7B) 188M 24.9 4.1 11.5 30.0 44.2 15.4 15.1 18.3
    InstructBLIP (Vicuna 7B) 202M 25.5 4.3 11.6 30.7 47.2 15.9 16.2 20.1
    Qwen-VL (Qwen 7B) 7.2B 27.1 4.9 13.1 31.3 63.8 18.2 19.5 24.4
    PointLLM (Vicuna 13B) 13.3B 26.2 4.9 11.9 31.3 50.9 23.4 22.3 26.2
    GPT4Point (OPT 2.7B) 110M 28.9 6.0 13.2 33.9 68.4 22.1 23.4 25.3
    GPT4Point (OPT 6.7B) 110M 31.5 7.2 13.8 35.4 78.7 27.1 26.2 30.4
    GPT4Point (Flan-T5-XL) 110M 32.2 7.2 14.2 35.5 78.0 27.6 26.3 31.3

    GPT4Point with frozen LLM backbones requires only 110M trainable parameters in the Point-QFormer and projection layers, outperforming 2D VLMs and PointLLM (13.3B trainable parameters) across all text inference metrics (e.g., CIDEr score of 78.7 vs 50.9 for PointLLM, and QA accuracy of 27.6% vs 23.4%).

  6. Knowl 6 — 3D Point-Text Retrieval and Zero-Shot Classification Results

    data/table

    GPT4Point was evaluated on bidirectional Point-Text Retrieval using the 1K ObjaverseXL-LVIS test dataset and on Zero-Shot Point Classification on the ModelNet40 benchmark (2,468 objects across 40 classes):

    Model Point →\rightarrow Text Retrieval Text →\rightarrow Point Retrieval ModelNet40
    R@1 R@5 R@10 R@1 R@5 R@10 Acc@1
    BLIP-2 17.56 41.16 52.82 16.72 40.20 52.56 35.62
    InstructBLIP 20.40 43.10 55.30 13.70 32.50 42.70 31.48
    PointLLM (Vicuna 7B) - - - - - - 41.33
    GPT4Point 32.20 64.00 81.30 89.70 98.10 98.90 43.90

    GPT4Point achieves 43.90% top-1 accuracy on ModelNet40 zero-shot classification, outperforming PointLLM by 2.57 points and InstructBLIP by 12.42 points. In 3D retrieval, GPT4Point scores 32.20% R@1 on Point →\rightarrow Text and 89.70% R@1 on Text →\rightarrow Point, substantially exceeding single-view 2D vision-language baselines.

  7. Knowl 7 — Benchmark Construction and Experimental Implementation Details

    experimental setup

    The 3D point-language benchmark and experimental implementation consist of:

    • Benchmark Dataset (ObjaverseXL-LVIS): Built by aligning Objaverse objects with LVIS categories and excluding complex indoor/outdoor scene scans. Validation and test sets each contain 1,000 single- or multi-object instances. Text annotations generated by Pyramid-XL were refined through multiple rounds of expert manual inspection.
    • Input Representation: 8,192 points sampled per 3D mesh object, formatted as N×6N \times 6 (XYZ spatial coordinates and normalized RGB color values).
    • Point Backbone: Point-BERT pretrained on point-cloud retrieval tasks via ULIP-2.
    • Language Backbones: Pretrained OPT (2.7B and 6.7B) and Flan-T5-XL.
    • Optimization: AdamW optimizer with an initial learning rate of 1×10−41 \times 10^{-4} and weight decay of 0.05. Batch size is 32. Both Stage 1 and Stage 2 are trained for 10 epochs on 8 NVIDIA A100 GPUs.
  8. Knowl 8 — Quantitative Evaluation of Controllable Text-to-3D Generation

    data/table

    Controllable text-to-3D generation using Point-QFormer feature conditioning was compared against direct text-to-3D and direct image-to-3D baselines on the Cap3D 2K test set using Frechet Inception Distance (FID), CLIP Score, and a 1-to-5 human user study rating:

    Text-to-3D Method Rendering FID ↓\downarrow CLIP Score ↑\uparrow User Study (1–5) ↑\uparrow
    Direct text-to-3D 34.7 74.9 3.98
    Direct image-to-3D 32.6 75.3 3.67
    Controllable text-to-3D (GPT4Point) 31.6 76.2 4.03

    Conditioning the Point-E generation on aligned Point-QFormer features achieves the lowest FID (31.6), the highest CLIP alignment score (76.2), and the highest human preference score (4.03).

  9. Knowl 9 — Ablation of Pyramid-XL Annotation Granularity on 3D Object QA

    data/table

    An ablation study evaluated the effect of pretraining data scale and annotation granularity from Pyramid-XL on downstream 3D Object Question Answering using an OPT-2.7B language model backbone evaluated on the ObjaverseXL-LVIS validation and test sets:

    Pyramid-XL Annotation Granularity 3D Object QA Accuracy (%)
    val test
    Level 2 22.3 22.1
    Level 1 + Level 2 25.6 25.4
    Level 1 + Level 2 + Level 3 (30%) 27.3 27.1
    Level 1 + Level 2 + Level 3 (70%) 28.3 28.2
    Level 1 + Level 2 + Level 3 (100%) 28.5 28.4

    Pretraining on high-volume coarse annotations (Level 1 + Level 2) raises QA accuracy by 3.3% over Level 2 alone. Incorporating fine-grained Level 3 instruction captions and QA pairs further scales accuracy from 25.4% to 28.4% on the test set.

Coverage note — None was omitted; all contributed architectural components, data engine stages, equations, benchmark setups, and empirical evaluation tables from the paper are represented.

References

  1. 1.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv:2308.12966, 2023.
  2. 2.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL Workshop, 2005.
  3. 3.Ching Chang, Wen-Chih Peng, and Tien-Fu Chen. Llm4ts: Two-stage fine-tuning for time-series forecasting with pretrained llms. arXiv:2308.08469, 2023.
  4. 4.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021.
  5. 5.Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al. Videollm: Modeling video sequence with large language models. arXiv:2305.13292, 2023.
  6. 6.Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3d: Disentangling geometry and appearance for highquality text-to-3d content creation. In ICCV, 2023.
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023.
  8. 8.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards generalpurpose vision-language models with instruction tuning. In NeurIPS, 2022.
  9. 9.Weinan Dai, Jinglei Tao, Xu Yan, Zhenyuan Feng, and Jinkun Chen. Addressing unintended bias in toxicity detection: An lstm and attention-based approach. In ICAICA, 2023.
  10. 10.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objaverse-xl: A universe of 10m+ 3d objects. In NeurIPS, 2023.
  11. 11.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In CVPR, 2023.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  13. 13.Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. In ICML, 2020.
  14. 14.Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, and Yu Qiao. Llamaadapter v2: Parameter-efficient visual instruction model. arXiv:2304.15010, 2023.
  15. 15.Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan. Planting a seed of vision in large language model. arXiv:2307.08041, 2023.
  16. 16.Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. Text-to-audio generation using instruction-tuned llm and latent diffusion model. In ACMMM, 2023.
  17. 17.Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. In CVPR, 2023.
  18. 18.Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al. Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following. arXiv:2309.00615, 2023.
  19. 19.Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al. Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following. arXiv:2309.00615, 2023.
  20. 20.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019.
  21. 21.Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models. In NeurIPS, 2023.
  22. 22.Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. Audiogpt: Understanding and generating speech, music, sound, and talking head. In AAAI, 2024.
  23. 23.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al. Language is not all you need: Aligning perception with language models. In NeurIPS, 2023.
  24. 24.Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo. Clip2point: Transfer clip to point cloud classification with image-depth pre-training. In ICCV, 2023.
  25. 25.Heewoo Jun and Alex Nichol. Shap-e: Generating conditional 3d implicit functions. arXiv:2305.02463, 2023.
  26. 26.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv:2305.03726, 2023.
  27. 27.Dongxu Li, Junnan Li, and Steven CH Hoi. Blip-diffusion: Pre-trained subject representation for controllable text-toimage generation and editing. In NeurIPS, 2023.
  28. 28.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023.
  29. 29.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv:2305.06355, 2023.
  30. 30.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023.
  31. 31.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, 2004.
  32. 32.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023.
  33. 33.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
  34. 34.Tiange Luo, Chris Rockwell, Honglak Lee, and Justin Johnson. Scalable 3d captioning with pretrained models. In NeurIPS, 2023.
  35. 35.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  36. 36.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv:2212.08751, 2022.
  37. 37.OpenAI. Chatgpt. https://openai.com/blog/chatgpt, 2022.
  38. 38.OpenAI. GPT-4 technical report, 2023.
  39. 39.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022.
  40. 40.Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. In ECCV, 2022.
  41. 41.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  42. 42.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824, 2023.
  43. 43.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR, 2023.
  44. 44.Zhangyang Qi, Jiaqi Wang, Xiaoyang Wu, and Hengshuang Zhao. Ocbev: Object-centric bev transformer for multi-view 3d object detection. In 3DV, 2024.
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  46. 46.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. In JMLR, 2020.
  47. 47.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  48. 48.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. In CVPR, 2023.
  49. 49.Zhan Shi, Xu Zhou, Xipeng Qiu, and Xiaodan Zhu. Improving image captioning with better use of captions. In ACL, 2020.
  50. 50.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multimodality. In ICLR, 2024.
  51. 51.InternLM Team. Internlm: A multilingual language model with progressively enhanced capabilities, 2023.
  52. 52.et al Tong Wu, Jiarui Zhang. Omniobject3d: Largevocabulary 3d object dataset for realistic perception, reconstruction and generation. In CVPR, 2023.
  53. 53.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv:2302.13971, 2023.
  54. 54.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. In NeurIPS, 2023.
  55. 55.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. In ICLR, 2021.
  56. 56.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. In CVPR, 2015.
  57. 57.Runsen Xu, Xiaolong Wang, Tai Wang, Yilun Chen, Jiangmiao Pang, and Dahua Lin. Pointllm: Empowering large language models to understand point clouds. arXiv:2308.16911, 2023.
  58. 58.Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese. Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding. In CVPR, 2023.
  59. 59.Le Xue, Ning Yu, Shu Zhang, Junnan Li, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese. Ulip-2: Towards scalable multimodal pre-training for 3d understanding. In CVPR, 2024.
  60. 60.Lili Yu, Bowen Shi, Ramakanth Pasunuru, Benjamin Muller, Olga Golovneva, Tianlu Wang, Arun Babu, Binh Tang, Brian Karrer, Shelly Sheynin, Candace Ross, Adam Polyak, Russell Howes, Vasu Sharma, Puxin Xu, Hovhannes Tamoyan, Oron Ashual, Uriel Singer, Shang-Wen Li, Susan Zhang, Richard James, Gargi Ghosh, Yaniv Taigman, Maryam Fazel-Zarandi, Asli Celikyilmaz, Luke Zettlemoyer, and Armen Aghajanyan. Scaling autoregressive multi-modal models: Pretraining and instruction tuning. arXiv:2309.02591, 2023.
  61. 61.Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In CVPR, 2022.
  62. 62.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, 2023.
  63. 63.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In CVPR, 2022.
  64. 64.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. In ICLR, 2024.
  65. 65.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv:2205.01068, 2022.
  66. 66.Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue. Meta-transformer: A unified framework for multimodal learning. arXiv:2307.10802, 2023.
  67. 67.Qihua Zhou, Song Guo, Jun Pan, Jiacheng Liang, Jingcai Guo, Zhenda Xu, and Jingren Zhou. Pass: Patch automatic skip scheme for efficient on-device video perception. TPAMI, 2024.
  68. 68.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592, 2023.
  69. 69.Zhu Ziyu, Ma Xiaojian, Chen Yixin, Deng Zhidong, Huang Siyuan, and Li Qing. 3d-vista: Pre-trained transformer for 3d vision and text alignment. In ICCV, 2023.

Citation

MLA
Qi, Z., et al. “GPT4Point: A Unified Framework for Point-Language Understanding and Generation”. arXiv, 2023, http://arxiv.org/abs/2312.02980v2.
APA
Qi, Z., Fang, Y., Sun, Z., Wu, X., Wu, T., Wang, J., Lin, D., & Zhao, H. (2023). GPT4Point: A Unified Framework for Point-Language Understanding and Generation. arXiv. http://arxiv.org/abs/2312.02980v2
Chicago
Qi, Z., Y. Fang, Z. Sun, et al. 2023. “GPT4Point: A Unified Framework for Point-Language Understanding and Generation”. arXiv. http://arxiv.org/abs/2312.02980v2.
Harvard
Qi, Z. et al. (2023) “GPT4Point: A Unified Framework for Point-Language Understanding and Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.02980v2.
Vancouver
1. Qi Z, Fang Y, Sun Z, Wu X, Wu T, Wang J, Lin D, Zhao H (2023) GPT4Point: A Unified Framework for Point-Language Understanding and Generation. arXiv

BibTeX

@article{qi2023gpt4point,
  title = {GPT4Point: A Unified Framework for Point-Language Understanding and Generation},
  author = {Qi, Zhangyang and Fang, Ye and Sun, Zeyi and Wu, Xiaoyang and Wu, Tong and Wang, Jiaqi and Lin, Dahua and Zhao, Hengshuang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.02980v2},
  eprint = {2312.02980}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE