Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Grounding
Zhihao YuanJinke RenChun-Mei FengHengshuang ZhaoShuguang CuiZhen Li
Proposes a visual programming framework that translates natural language descriptions into executable modular code using large language models, achieving zero-shot open-vocabulary 3D visual grounding without requiring dense annotations.
Identifying and localizing physical objects within 3D environments using natural language—known as 3D Visual Grounding—is essential for autonomous robotics, virtual reality, and spatial computing. However, conventional methods rely heavily on supervised training over human-annotated datasets, which is prohibitively costly and confines systems to fixed, predefined vocabularies. The article evaluates and demonstrates a training-free visual programming framework that leverages large language models and multi-modal perception tools to localize objects in complex 3D scenes in a zero-shot, open-vocabulary manner.
The framework converts free-form text descriptions into modular Python programs generated by a language model. It executes these programs across three specialized components: view-independent spatial modules, view-dependent modules mapped through a 2D egocentric camera perspective, and a language-object correlation module that merges 3D geometric point clouds with 2D visual appearance features. The system was benchmarked across the full validation sets of the ScanRefer and Nr3D datasets, comparing its performance against leading supervised and open-vocabulary baselines.
The findings show that the proposed zero-shot framework achieves strong localization accuracy, reaching an Acc@0.5 score of 32.7% on ScanRefer, which outperforms supervised baselines such as ScanRefer (24.3%) and TGNN (29.7%), and dramatically surpasses existing zero-shot open-vocabulary approaches like OpenScene (6.5%) and LERF (0.9%). On the Nr3D benchmark, the framework achieved a 39.0% overall top-1 accuracy, exceeding the supervised InstanceRefer model (38.8%) and outperforming the 3DVG-Transformer by 2.0% on view-dependent queries. Programmatic visual reasoning also proved vastly superior to conversational language model interaction, lowering computational token costs from 0.19 per batch on GPT-3.5 while boosting accuracy from 25.4% to 32.1%.
These results demonstrate that organizations can deploy high-performing 3D spatial intelligence without the massive time, data-collection expenses, and computational overhead required to train supervised models. Furthermore, by structuring visual reasoning into modular code and 2D camera projections, the framework resolves core spatial ambiguity issues while preserving the flexibility to swap in future foundational vision and language models as they improve.
Organizations developing spatial computing and robotics systems should transition from rigid, closed-vocabulary models toward modular programmatic architectures. Next steps include improving language model generation accuracy through expanded in-context prompt libraries and self-verification mechanisms, as program generation remains the primary source of error. While confidence in the framework's geometric and spatial reasoning is high across indoor benchmarks, practitioners should exercise caution in ambiguous edge cases involving subtle physical actions or domain gaps in 2D imagery until underlying perception and parsing models mature further.
- Paper: Multi-View Transformer for 3D Visual Grounding, Shijia Huang et al. (2022). This paper establishes fundamental multi-view modeling and view-dependent reasoning techniques for 3D visual grounding that the source directly builds upon in its visual programming modules.
- Paper: Language Conditioned Spatial Relation Reasoning for 3D Object Grounding, Shizhe Chen et al. (2022). This work introduces explicit spatial relation reasoning and benchmark evaluation protocols on datasets like Nr3D and Sr3D, which form the prerequisite basis for the source's zero-shot modular reasoning framework.
- Paper: CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data, Yihan Zeng et al. (2023). This paper presents methods for aligning point clouds with open-vocabulary text embeddings, providing essential foundational concepts for the source's language-object correlation module.
- Paper: ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding, Le Xue et al. (2023). This work develops unified multimodal representations across language, images, and 3D point clouds, which underpin the open-vocabulary transfer strategies adapted by the source.
- Paper: Context-aware Alignment and Mutual Masking for 3D-Language Pre-training, Zhao Jin et al. (2023). This study addresses viewpoint ambiguities and spatial relation representations in 3D-language pretraining, informing the view-independent and view-dependent decomposition in the source.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). This paper introduces seminal vision-language knowledge distillation for open-vocabulary detection, offering core methodology for adapting closed-set detectors to open vocabularies.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). This work provides an extensive benchmark analyzing how multimodal large language models perceive, remember, and reason over 3D spatial layouts, extending the zero-shot spatial capabilities explored in the source.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). This research extends open-vocabulary 3D scene understanding by mapping dense vision-language features directly into implicit 3D radiance fields for open-set segmentation.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). This paper applies multi-step visual chain-of-thought reasoning to embodied vision-language-action settings, extending the modular visual programming concept to physical robot manipulation.
