Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Grounding

Zhihao YuanJinke RenChun-Mei FengHengshuang ZhaoShuguang CuiZhen Li

article2024CVPR76 citations

Proposes a visual programming framework that translates natural language descriptions into executable modular code using large language models, achieving zero-shot open-vocabulary 3D visual grounding without requiring dense annotations.

Listen

Identifying and localizing physical objects within 3D environments using natural language—known as 3D Visual Grounding—is essential for autonomous robotics, virtual reality, and spatial computing. However, conventional methods rely heavily on supervised training over human-annotated datasets, which is prohibitively costly and confines systems to fixed, predefined vocabularies. The article evaluates and demonstrates a training-free visual programming framework that leverages large language models and multi-modal perception tools to localize objects in complex 3D scenes in a zero-shot, open-vocabulary manner.

The framework converts free-form text descriptions into modular Python programs generated by a language model. It executes these programs across three specialized components: view-independent spatial modules, view-dependent modules mapped through a 2D egocentric camera perspective, and a language-object correlation module that merges 3D geometric point clouds with 2D visual appearance features. The system was benchmarked across the full validation sets of the ScanRefer and Nr3D datasets, comparing its performance against leading supervised and open-vocabulary baselines.

The findings show that the proposed zero-shot framework achieves strong localization accuracy, reaching an Acc@0.5 score of 32.7% on ScanRefer, which outperforms supervised baselines such as ScanRefer (24.3%) and TGNN (29.7%), and dramatically surpasses existing zero-shot open-vocabulary approaches like OpenScene (6.5%) and LERF (0.9%). On the Nr3D benchmark, the framework achieved a 39.0% overall top-1 accuracy, exceeding the supervised InstanceRefer model (38.8%) and outperforming the 3DVG-Transformer by 2.0% on view-dependent queries. Programmatic visual reasoning also proved vastly superior to conversational language model interaction, lowering computational token costs from 3.05to3.05 to 0.19 per batch on GPT-3.5 while boosting accuracy from 25.4% to 32.1%.

These results demonstrate that organizations can deploy high-performing 3D spatial intelligence without the massive time, data-collection expenses, and computational overhead required to train supervised models. Furthermore, by structuring visual reasoning into modular code and 2D camera projections, the framework resolves core spatial ambiguity issues while preserving the flexibility to swap in future foundational vision and language models as they improve.

Organizations developing spatial computing and robotics systems should transition from rigid, closed-vocabulary models toward modular programmatic architectures. Next steps include improving language model generation accuracy through expanded in-context prompt libraries and self-verification mechanisms, as program generation remains the primary source of error. While confidence in the framework's geometric and spatial reasoning is high across indoor benchmarks, practitioners should exercise caution in ambiguous edge cases involving subtle physical actions or domain gaps in 2D imagery until underlying perception and parsing models mature further.

arXiv: 2311.15383
Cover for Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Grounding

Abstract

3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary, which can be restrictive. To address this issue, we propose a novel visual programming approach for zero-shot open-vocabulary 3DVG, leveraging the capabilities of large language models (LLMs). Our approach begins with a unique dialog-based method, engaging with LLMs to establish a foundational understanding of zero-shot 3DVG. Building on this, we design a visual program that consists of three types of modules, i.e., view-independent, view-dependent, and functional modules. These modules, specifically tailored for 3D scenarios, work collaboratively to perform complex reasoning and inference. Furthermore, we develop an innovative language-object correlation module to extend the scope of existing 3D object detectors into open-vocabulary scenarios. Extensive experiments demonstrate that our zero-shot approach can outperform some supervised baselines, marking a significant stride towards effective 3DVG. Code is available at https://curryyuan.github.io/2SVG3D.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Dialog with LLM
  • 3.2. 3D Visual Programming
  • 3.3. Addressing View-Dependent Relations
  • 3.4. Language-Object Correlation Module
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Quantitative Results
  • 4.3. Qualitative Results
  • 4.4. Ablation Studies
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Zero-shot visual programming pipeline for 3D visual grounding

    model/method

    The paper introduces a zero-shot, open-vocabulary approach to 3D visual grounding (3DVG) that does not require object–text pair annotations or a fixed query vocabulary. Given an RGB-D point cloud and a free-form referring expression, a large language model (LLM) uses in-context examples to generate a structured 3D visual program. The program is translated into executable Python code, its operations are executed on scene proposals and associated images, and the resulting candidate is returned as the target bounding box. The pipeline separates language-level reasoning from geometric and visual execution: the LLM determines which object descriptions and relations are needed, while specialized modules perform localization, spatial comparison, and open-vocabulary recognition.

  2. Knowl 2 — Composable 3D visual-program representation and module taxonomy

    model/method

    A visual program is a sequence of operations, each consisting of a module name, input arguments, and an output variable that can be reused by later operations. A typical program is:

    BOX0 = LOC('round cocktail table')
    BOX1 = LOC('blue and yellow poster')
    TARGET = CLOSEST(targets=BOX0, anchors=BOX1)
    

    Here, LOC returns candidate 3D boxes matching a language query, while relation modules filter or rank candidate boxes relative to anchor boxes. The paper organizes modules into three groups:

    • View-independent modules: NEAR, CLOSE, NEXT_TO, FAR, ABOVE, BELOW, UNDER, TOP, ON, OPPOSITE, and MIDDLE; these operate on 3D spatial relationships that do not depend on the observer’s orientation.
    • View-dependent modules: FRONT, BEHIND, BACK, RIGHT, LEFT, FACING, LEFTMOST, RIGHTMOST, LOOKING, ACROSS, and BETWEEN; these require a viewpoint-consistent 2D interpretation of the scene.
    • Functional modules: MIN, MAX, SIZE, LENGTH, and WIDTH; these select objects using extremal or geometric criteria.

    The modules are composable: the output of one localization or relation operation can become the target or anchor input of another operation, allowing multi-step spatial reasoning without training a task-specific grounding network.

  3. Knowl 3 — Egocentric projection for resolving view-dependent relations

    equation

    To make relations such as left, right, front, behind, and between well-defined in a freely rotatable 3D scene, the method creates a virtual orthographic camera at the scene center and rotates it toward an anchor object. Let Pc∈R3P_c\in\mathbb{R}^{3} be the scene-center position, Pa∈R3P_a\in\mathbb{R}^{3} the anchor position, and up∈R3up\in\mathbb{R}^{3} an upward direction. The camera pose is computed as

    R,t=LookAt⁡(Pc,Pa,up),R,t=\operatorname{LookAt}(P_c,P_a,up),

    where R∈R3×3R\in\mathbb{R}^{3\times3} is a rotation matrix and t∈R3t\in\mathbb{R}^{3} is a translation vector. For a 3D point with homogeneous coordinates p=(x,y,z,1)T∈R4p=(x,y,z,1)^\mathsf{T}\in\mathbb{R}^{4} and an orthographic camera intrinsic matrix I∈R3×3I\in\mathbb{R}^{3\times3}, its egocentric projection is

    p~=(u,v,w)T=I [R∣t]p,\tilde p=(u,v,w)^\mathsf{T}=I\,[R\mid t]p,

    where uu and vv are horizontal and vertical image-plane coordinates and ww is depth. Candidate object centers with smaller uu are treated as left of the anchor and those with larger uu as right; depth ww distinguishes front from behind. The same projected coordinate system supports the BETWEEN relation. If a query does not provide a separate target set, the method uses the candidate objects themselves as anchors, enabling expressions such as “the leftmost window.”

  4. Knowl 4 — Language–object correlation for open-vocabulary localization

    model/method

    The language–object correlation (LOC) module extends a closed-vocabulary 3D instance segmentation model to open-vocabulary localization by combining 3D geometry with 2D appearance. For a query such as “round cocktail table,” the module first uses a 3D instance segmentation network to obtain object proposals and filter proposals whose closed-set prediction is table. Each remaining 3D proposal is mapped to corresponding RGB imagery, where fine-grained color, texture, shape, and appearance cues are evaluated against the full query.

    The paper implements three interchangeable kinds of 2D assistance: (1) an image-classification model such as CLIP, using a dynamic vocabulary containing both the query phrase and the coarse 3D class and ranking proposals by image–text cosine similarity; (2) a visual question-answering model such as ViLT, asked whether the image contains the queried object; and (3) a general vision-language model such as BLIP-2, asked the same question and evaluated by its generated response. The design is not tied to a particular 2D or 3D backbone, and the final LOC module retains the geometric discrimination of 3D proposals while adding open-vocabulary visual recognition.

  5. Knowl 5 — Direct LLM dialogue baseline and its failure modes

    limitation

    The paper first constructs a direct dialogue baseline that converts a colored 3D point cloud P∈RN×6P\in\mathbb{R}^{N\times6}, where NN is the number of points and each point contains 3D coordinates and color, into a textual scene description. Every detected object is represented by an identifier, category, 3D position, and dimensions, for example: “Object <id> is a <category> located at (x,y,z)(x,y,z) with sizes (width,length,height)(width,length,height).” The scene narrative and referring expression are supplied to an LLM, which identifies relevant target and anchor objects and explains its choice.

    This dialogue can solve simple view-independent cases such as selecting the keyboard closest to a door by comparing object distances. However, the paper finds that direct dialogue is unreliable for view-dependent relations because the interpretation of left and right changes with the observer’s viewpoint, while an LLM tends to compare raw coordinate axes. It is also unreliable for numerical operations such as distance calculation. These limitations motivate the explicit executable relation modules and external visual computations in the visual-programming system.

  6. Knowl 6 — Datasets and evaluation protocol

    experimental setup

    The experiments evaluate zero-shot 3DVG on the validation splits of ScanRefer and Nr3D. ScanRefer contains 51,500 referring sentences associated with 800 ScanNet scenes. Nr3D consists of human-written referring expressions collected through a two-player reference game; its examples are divided into easy cases with at most one same-class distractor and hard cases with multiple same-class distractors. Nr3D is also partitioned into view-dependent and view-independent expressions.

    Two evaluation settings are used. In the proposal-generation setting, the method must localize an object box, and [email protected] and [email protected] are the percentages of predictions whose intersection-over-union (IoU) with the ground-truth box exceeds 0.25 and 0.5, respectively; this is the default ScanRefer setting. In the classification-only setting, ground-truth object masks are supplied and the method must select the correct instance; this top-1 accuracy is the default Nr3D setting. Comparisons include supervised 3DVG systems, open-vocabulary scene-understanding systems, and ablations using only 2D or only 3D information.

  7. Knowl 7 — ScanRefer validation performance

    data/table

    On ScanRefer, the proposed zero-shot visual-programming system is compared with supervised and open-vocabulary methods across scenes containing a unique object of the queried class, scenes containing multiple same-class objects, and the complete validation set. The numbers are percentages for box accuracy at the indicated IoU threshold.

    Could not parse LaTeX table

    The full zero-shot system reaches 36.4 at IoU 0.25 and 32.7 at IoU 0.5 overall. It exceeds the supervised ScanRefer and TGNN baselines at IoU 0.5, and it substantially outperforms the other open-vocabulary baselines. Combining 3D and 2D information is important: the full system improves over the 3D-only variant by 3.3 points at IoU 0.5 overall and over the 2D-only variant by 15.1 points.

  8. Knowl 8 — Nr3D language-grounding performance

    data/table

    On Nr3D, ground-truth instance masks remove localization error, so the reported values are top-1 instance-selection accuracies in percent. The proposed zero-shot system is compared across easy, hard, view-dependent, view-independent, and overall subsets.

    Could not parse LaTeX table

    The full zero-shot system obtains 39.0 overall, exceeding InstanceRefer’s 38.8 and its own 3D-only ablation’s 36.7. On the view-dependent split it obtains 36.8, which is 2.0 points higher than 3DVG-Transformer’s 34.8, supporting the value of explicit relation modules for viewpoint-sensitive grounding. The system does not surpass the strongest supervised BUTD-DETR baseline, which reaches 54.6 overall.

  9. Knowl 9 — Visual programming improves accuracy, token use, and monetary cost

    empirical result

    The paper compares direct LLM dialogue with visual-program generation on 700 ScanRefer validation examples using GPT-3.5-turbo-0613 and GPT-4-0613. [email protected] is box accuracy at IoU threshold 0.5, Tokens is the total number of input and output tokens, and Cost is the reported API cost in US dollars.

    Could not parse LaTeX table

    For both language models, visual programming is more accurate and substantially cheaper than direct dialogue. GPT-4 improves over GPT-3.5 for both approaches, but the program-based system still uses far fewer tokens: 121k versus 1,959k with GPT-3.5 and 115k versus 1,916k with GPT-4. The remaining experiments use GPT-3.5 to reduce cost.

  10. Knowl 10 — Ablations validate explicit relations and multimodal perception

    empirical result

    Cumulative relation ablations show that explicit view-dependent and view-independent operations are important for grounding. The reported grounding accuracy increases as relation modules are added: for view-dependent relations, LEFT gives 26.5, LEFT+RIGHT gives 32.4, LEFT+RIGHT+FRONT gives 35.9, adding BEHIND gives 36.8, and adding BETWEEN gives 39.0. For view-independent relations, CLOSEST gives 18.8, CLOSEST+FARTHEST gives 30.7, adding LOWER gives 34.0, adding HIGHER gives 36.8, and using all four gives 39.0. Thus LEFT, RIGHT, and CLOSEST are especially influential.

    The perception ablations also support combining 3D and 2D evidence. On the ScanRefer unique/multiple split with [email protected], CLIP, ViLT, and BLIP-2 as 2D assistants obtain (62.5,27.1,35.7)(62.5,27.1,35.7), (60.3,27.1,35.1)(60.3,27.1,35.1), and (63.8,27.7,36.4)(63.8,27.7,36.4), respectively, where each tuple is unique, multiple, and overall accuracy. With different 3D backbones on Nr3D, PointNet++, PointBERT, and PointNeXt obtain view-dependent/view-independent/overall accuracies of (35.8,39.4,38.2)(35.8,39.4,38.2), (36.0,39.8,38.6)(36.0,39.8,38.6), and (36.8,40.0,39.0)(36.8,40.0,39.0), respectively. These results show that the framework is compatible with multiple 2D and 3D foundation models and benefits from both geometric and appearance information.

Coverage note — The prompt-size and voting curves, qualitative visual examples, and detailed error-breakdown percentages were omitted because they are secondary analyses rather than load-bearing components of the method or its principal quantitative claims.

References

  1. 1.Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In European Conference on Computer Vision, pages 422–440. Springer, 2020.
  2. 2.Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al. Do as i can, not as i say: Grounding language in robotic affordances. In Conference on Robot Learning, pages 287–318. PMLR, 2023.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  4. 4.Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In European Conference on Computer Vision, pages 202–221. Springer, 2020.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  6. 6.Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Language conditioned spatial relation reasoning for 3d object grounding. Advances in Neural Information Processing Systems, 35:20522–20535, 2022.
  7. 7.Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, and Angel X Chang. Unit3d: A unified transformer for 3d dense captioning and visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18109–18119, 2023.
  8. 8.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In CVPR, pages 3075–3084, 2019.
  9. 9.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, pages 5828–5839, 2017.
  10. 10.John David N Dionisio, William G Burns Iii, and Richard Gilbert. 3d virtual worlds and the metaverse: Current status and future possibilities. ACM Computing Surveys (CSUR), 45(3):1–38, 2013.
  11. 11.Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al. Faith and fate: Limits of transformers on compositionality. arXiv preprint arXiv:2305.18654, 2023.
  12. 12.Qi Feng, Vitaly Ablavsky, and Stan Sclaroff. Cityflow-nl: Tracking and retrieval of vehicles at city scale by natural language descriptions. arXiv preprint arXiv:2101.04741, 2021.
  13. 13.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scaling open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision, pages 540–557. Springer, 2022.
  14. 14.Zoey Guo, Yiwen Tang, Ray Zhang, Dong Wang, Zhigang Wang, Bin Zhao, and Xuelong Li. Viewrefer: Grasp the multi-view knowledge for 3d visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15372–15383, 2023.
  15. 15.Tanmay Gupta and Aniruddha Kembhavi. Visual programming: Compositional visual reasoning without training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14953–14962, 2023.
  16. 16.Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models. Advances in Neural Information Processing Systems, 36:20482–20494, 2023.
  17. 17.Joy Hsu, Jiayuan Mao, and Jiajun Wu. Ns3d: Neuro-symbolic grounding of 3d objects and relations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2614–2623, 2023.
  18. 18.Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu. Text-guided graph neural networks for referring 3d instance segmentation. 35th AAAI Conference on Artificial Intelligence, 2021.
  19. 19.Shijia Huang, Yilun Chen, Jiaya Jia, and Liwei Wang. Multi-view transformer for 3d visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15524–15533, 2022.
  20. 20.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR, 2022.
  21. 21.Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki. Bottom up top down detection transformers for language grounding in images and point clouds. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI, pages 417–433. Springer, 2022.
  22. 22.Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. Pointgroup: Dual-set point grouping for 3d instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4867–4876, 2020.
  23. 23.Zhao Jin, Munawar Hayat, Yuwei Yang, Yulan Guo, and Yinjie Lei. Context-aware alignment and mutual masking for 3d-language pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10984–10994, 2023.
  24. 24.Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embedded radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19729–19739, 2023.
  25. 25.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR, 2021.
  26. 26.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  27. 27.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. In International Conference on Learning Representations, 2022.
  28. 28.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  29. 29.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7061–7070, 2023.
  30. 30.Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500. IEEE, 2023.
  31. 31.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2949–2958, 2021.
  32. 32.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  33. 33.Stylianos Mystakidis. Metaverse. Encyclopedia, 2(1):486–497, 2022.
  34. 34.OpenAI OpenAI. Gpt-4 technical report. 2023.
  35. 35.Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 815–824, 2023.
  36. 36.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017.
  37. 37.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017.
  38. 38.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep hough voting for 3d object detection in point clouds. In proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9277–9286, 2019.
  39. 39.Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. Advances in Neural Information Processing Systems, 35:23192–23204, 2022.
  40. 40.Xing-yue Qiu, Chuang-Kai Chiu, Lu-Lu Zhao, Cai-Feng Sun, and Shu-jie Chen. Trends in vr/ar technology-supporting language learning from 2008 to 2019: A research perspective. Interactive Learning Environments, 31(4):2090–2113, 2023.
  41. 41.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  42. 42.Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox. Languagerefer: Spatial-language model for 3d visual grounding. In 5th Annual Conference on Robot Learning, 2021.
  43. 43.Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. Mask3d: Mask transformer for 3d semantic instance segmentation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 8216–8223. IEEE, 2023.
  44. 44.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 567–576, 2015.
  45. 45.Dídac Surís, Sachit Menon, and Carl Vondrick. Vipergpt: Visual inference via python execution for reasoning. arXiv preprint arXiv:2303.08128, 2023.
  46. 46.Ayça Takmaz, Elisabetta Fedele, Robert W Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. Openmask3d: Open-vocabulary 3d instance segmentation. arXiv preprint arXiv:2306.13631, 2023.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017.
  49. 49.Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor. Chatgpt for robotics: Design principles and model abilities. Microsoft Auton. Syst. Robot. Res, 2:20, 2023.
  50. 50.John Vince and John A Vince. Mathematics for computer graphics. Springer, 2006.
  51. 51.Thang Vu, Kookhoi Kim, Tung M Luu, Thanh Nguyen, and Chang D Yoo. Softgroup for 3d instance segmentation on point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2708–2717, 2022.
  52. 52.Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, and Matthias Nießner. Rio: 3d object instance relocalization in changing indoor environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7658–7667, 2019.
  53. 53.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6629–6638, 2019.
  54. 54.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  55. 55.Wei Wei. Research progress on virtual reality (vr) and augmented reality (ar) in tourism and hospitality: A critical review of publications from 2000 to 2018. Journal of Hospitality and Tourism Technology, 10(4):539–570, 2019.
  56. 56.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023.
  57. 57.Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, and Jian Zhang. Eda: Explicit text-decoupling and dense alignment for 3d visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19231–19242, 2023.
  58. 58.Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9068–9079, 2018.
  59. 59.Xu Yan, Zhihao Yuan, Yuhao Du, Yinghong Liao, Yao Guo, Shuguang Cui, and Zhen Li. Comprehensive visual question answering on point clouds through compositional scene manipulation. IEEE Transactions on Visualization and Computer Graphics, 2023.
  60. 60.Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo. Sat: 2d semantics assisted training for 3d visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1856–1866, 2021.
  61. 61.Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19313–19322, 2022.
  62. 62.Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Zhen Li, and Shuguang Cui. Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1791–1800, 2021.
  63. 63.Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li, Yao Guo, Shuguang Cui, and Zhen Li. Toward explainable and fine-grained 3d grounding through referring textual phrases. arXiv preprint arXiv:2207.01821, 2022.
  64. 64.Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Zhen Li, and Shuguang Cui. X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  65. 65.Yihan Zeng, Chenhan Jiang, Jiageng Mao, Jianhua Han, Chaoqiang Ye, Qingqiu Huang, Dit-Yan Yeung, Zhen Yang, Xiaodan Liang, and Hang Xu. Clip2: Contrastive language-image-point pretraining from real-world point cloud data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15244–15253, 2023.
  66. 66.Jiazhao Zhang, Chenyang Zhu, Lintao Zheng, and Kai Xu. Fusion-aware point convolution for online semantic 3d scene segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4534–4543, 2020.
  67. 67.Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8552–8562, 2022.
  68. 68.Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu. 3dvg-transformer: Relation modeling for visual grounding on point clouds. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2928–2937, 2021.
  69. 69.Denny Zhou, Nathanael Scharli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.
  70. 70.Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo, Ziyao Zeng, Zipeng Qin, Shanghang Zhang, and Peng Gao. Point-clip v2: Prompting clip and gpt for powerful 3d open-world learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2639–2650, 2023.

Citation

MLA
Yuan, Z., et al. “Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding”. arXiv, 2023, http://arxiv.org/abs/2311.15383v2.
APA
Yuan, Z., Ren, J., Feng, C.-M., Zhao, H., Cui, S., & Li, Z. (2023). Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding. arXiv. http://arxiv.org/abs/2311.15383v2
Chicago
Yuan, Z., J. Ren, C.-M. Feng, H. Zhao, S. Cui, and Z. Li. 2023. “Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding”. arXiv. http://arxiv.org/abs/2311.15383v2.
Harvard
Yuan, Z. et al. (2023) “Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.15383v2.
Vancouver
1. Yuan Z, Ren J, Feng C-M, Zhao H, Cui S, Li Z (2023) Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding. arXiv

BibTeX

@article{yuan2023visual,
  title = {Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding},
  author = {Yuan, Zhihao and Ren, Jinke and Feng, Chun-Mei and Zhao, Hengshuang and Cui, Shuguang and Li, Zhen},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.15383v2},
  eprint = {2311.15383}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE