RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

Chan Hee SongValts BlukisJonathan TremblayStephen TyreeYu SuStan Birchfield

article2025CVPR200 citations

Introduces RoboSpatial, a large-scale multimodal dataset with three million annotations across multiple reference frames and spatial compatibility tasks, enabling vision-language models to achieve superior spatial reasoning for robotic manipulation.

Listen

Vision-language models that combine visual perception with language understanding are increasingly used to direct robots in real-world tasks such as object manipulation, navigation, and automated planning. However, existing models frequently fail at basic spatial reasoning because they are trained on general internet images lacking precise scale, depth, and viewpoint context. Most current systems cannot reliably interpret directional relationships relative to a specific object, identify available free space, or determine whether an item physically fits into a target area, which creates substantial operational bottlenecks for autonomous robotic systems.

The article demonstrates that this performance gap stems primarily from a lack of appropriate training data and introduces a large-scale multimodal dataset, ROBOSPATIAL, to teach foundational spatial reasoning to both two-dimensional and three-dimensional vision-language models. The dataset pairs five thousand real-world three-dimensional indoor and tabletop scans with one million egocentric images to generate roughly three million spatial question-answer pairs. The automated data generation pipeline systematically extracts three core relationship types: spatial context, which identifies empty coordinates for placement; spatial compatibility, which verifies whether an object physically fits with a minimum safety margin; and spatial configuration, which evaluates relative positions. Crucially, every question is formulated across three distinct reference frames—observer-centric, world-centric, and object-centric.

Rigorous evaluations show that training on this dataset substantially boosts spatial reasoning across all tested models. On a held-out validation benchmark, fine-tuned models improved their overall performance by roughly twenty to thirty percentage points over baseline versions. In physical robot manipulation trials, a standard vision-language model trained on the dataset more than doubled its task execution success rate from 23.7% to 52.6%, outperforming leading closed-source models such as GPT-4o (46.9%). Models also demonstrated strong out-of-domain transfer, successfully interpreting novel spatial prepositions like "under" and "next to" and inferring implicit object orientations from everyday language.

These findings indicate that targeted, geometrically grounded training data can resolve fundamental reasoning limitations without requiring massive, ungrounded web datasets. For organizations developing robotic automation, enhanced spatial understanding lowers physical error rates, reduces collision risks, and expands the range of complex manipulation tasks that robots can handle from natural language instructions. Organizations pursuing autonomous robotics should integrate multi-perspective spatial supervision into their visual reasoning pipelines rather than relying exclusively on off-the-shelf generalist vision models.

While three-dimensional models showed initial performance advantages over two-dimensional counterparts, direct architectural comparisons remain somewhat inconclusive due to overlapping pretraining datasets. Additionally, in physical deployments, minor two-dimensional pixel prediction shifts translated to physical errors of five to ten centimeters. Future research and technical pilots should focus on refining coordinate projection accuracy, assessing performance under varied camera viewpoints, and training models on partial, sparse sensor scans to facilitate seamless deployment on real-world mobile robots.

Cover for RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

Abstract

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant challenges in spatial reasoning tasks, as their training data are based on general-purpose image datasets that often lack sophisticated spatial understanding. For example, datasets frequently do not capture reference frame comprehension, yet effective spatial reasoning requires understanding whether to reason from ego-, world-, or object-centric perspectives. To address this issue, we introduce ROBOSPATIAL, a large-scale dataset for spatial understanding in robotics. It consists of real indoor and tabletop scenes, captured as 3D scans and egocentric images, and annotated with rich spatial information relevant to robotics. The dataset includes 1M images, 5k 3D scans, and 3M annotated spatial relationships, and the pairing of 2D egocentric images with 3D scans makes it both 2D- and 3D-ready. Our experiments show that models trained with ROBOSPATIAL outperform baselines on downstream tasks such as spatial affordance prediction, spatial relationship prediction, and robotics manipulation.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Spatial Relationships
  • 3.2. Dataset Generation
  • 3.2.1. Stage 1: 3D Spatial Relation Extraction
  • 3.2.2. Stage 2: 2D Spatial Point and Region Sampling
  • 3.2.3. Question-Answer Generation
  • 4. Experiments
  • 4.1. Setup
  • 4.1.1. Trained 2D/3D VLMs
  • 4.1.2. Spatial Understanding Evaluation
  • 4.1.3. Cross-Dataset Generalization Evaluation
  • 4.1.4. Out-of-Domain Evaluation
  • Spatial Context
  • Spatial Compatibility
  • 4.2. Results
  • Spatial Configuration
  • 4.3. Real Robot Experiments
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — ROBOSPATIAL: a 2D/3D robotics spatial-reasoning resource

    model/method

    ROBOSPATIAL is a large-scale spatial-understanding dataset for vision-language models used in robotics. It is constructed from real indoor and tabletop scenes and pairs egocentric RGB images with 3D scans, oriented 3D bounding boxes, and spatial question-answer annotations. The resource contains approximately 5,000 3D scans, 1 million images, and 3 million spatial question-answer pairs, making it usable for both RGB-only 2D VLMs and RGB-D or point-cloud 3D VLMs. The released resource also includes the held-out ROBOSPATIAL-Val benchmark, the manually collected ROBOSPATIAL-Home benchmark, and code for generating annotations from 3D scenes.

  2. Knowl 2 — Three actionable spatial-reasoning tasks and three reference frames

    definition

    ROBOSPATIAL decomposes robotic spatial reasoning into three task types. Spatial configuration asks whether a directional relation holds between two objects and returns a binary yes/no answer. Spatial context asks for image points in vacant space or on a support surface satisfying a directional relation to an anchor object and returns one or more 2D coordinates. Spatial compatibility asks whether a specified target object can fit in the queried region and returns a binary yes/no answer.

    Each task is generated under three reference frames: ego-centric, whose axes are determined by the camera pose; world-centric, whose axes are determined by the global scene coordinate system; and object-centric, whose axes are determined by the orientation of the anchor object. The directional vocabulary consists of left, right, above, below, front, and behind. Thus, a statement such as “in front of the car” can be evaluated relative to the camera, the global scene, or the car’s own front-facing direction.

  3. Knowl 3 — Automatic 3D-to-2D spatial annotation pipeline

    model/method

    The ROBOSPATIAL generation pipeline takes a source scene dataset containing RGB images, camera intrinsics and extrinsics, and semantic oriented 3D bounding boxes. It produces entries of the form di=(Ii,qi,ai,li)d_i=(I_i,q_i,a_i,l_i), where IiI_i is an RGB image, qiq_i is a spatial question, aia_i is its answer, and lil_i is one of the reference-frame labels ego, world, or object.

    The first stage extracts spatial relations from 3D geometry. It uses each bounding box’s 3D position and heading vector, with the heading vector treated as the object’s front direction, to determine directional relations independently in the three reference frames. Camera extrinsics transform positions between the camera, world, and anchor-object coordinate systems. For configuration questions, the pipeline evaluates visible object pairs that are uniquely identifiable in the image and records whether the relation is true or false.

    The second stage maps 3D spatial annotations into image-space targets. It builds a top-down occupancy map from the 3D bounding boxes, samples candidate points in unoccupied regions satisfying the required directional relation, projects those points into the image using the camera intrinsics, and removes candidates whose camera-to-point rays intersect an occupied box. The resulting context answer is a list of valid image coordinates.

  4. Knowl 4 — Geometry-based spatial compatibility criterion

    model/method

    For a spatial compatibility question, ROBOSPATIAL tests whether the target object can actually occupy the queried region rather than merely checking whether the region is empty. The pipeline places a virtual oriented bounding box with the target object’s dimensions at a sampled candidate location on the ground plane. The candidate is compatible only if the virtual box does not intersect any existing scene bounding box and leaves at least a 10 cm margin along every axis.

    The simulated object may translate and rotate within the plane. Compatibility is therefore recorded as true only when at least one placement satisfies collision avoidance and the 10 cm clearance requirement; otherwise it is recorded as false. This produces placement-oriented supervision for object rearrangement, assembly, and manipulation.

  5. Knowl 5 — Template-based questions with auxiliary object grounding

    algorithm

    ROBOSPATIAL converts extracted spatial relations into deterministic question-answer pairs using templates of the form “TARGET RELATION ANCHOR REFERENCE FRAME.” Configuration and compatibility templates produce binary answers, while context templates produce lists of 2D image coordinates.

    Input: RGB images, camera calibration, oriented 3D bounding boxes, extracted spatial relations
    Output: spatial question-answer pairs and object-grounding examples
    For each extracted relation:
        Select its target, anchor, directional relation, and reference frame
        Instantiate the deterministic template for configuration, context, or compatibility
        Attach the binary answer or projected coordinate list
        Project each referenced 3D bounding box into the image
        Add the object description and projected 2D bounding box to the grounding set
    Return all spatial question-answer pairs and grounding examples

    The auxiliary grounding set links object descriptions to their projected 2D bounding boxes. It is included during training to improve resolution of object references, but its annotations are not counted as spatial-reasoning evaluation targets. Deterministic wording is intended to make the supervision depend on visual grounding and geometry rather than linguistic ambiguity or commonsense priors.

  6. Knowl 6 — Training data splits and spatial evaluation protocol

    experimental setup

    The authors apply the generation pipeline to ScanNet, Matterport3D, 3RScan, HOPE, and GraspNet-1Billion, using indoor scenes for navigation-oriented layouts and tabletop scenes for object-centric manipulation. The resulting rounded dataset statistics are:

    • Indoor training: 4,916 scans, 883k images, and 2.7M question-answer pairs.
    • Indoor validation: 40 scans, 1k images, and 3k question-answer pairs.
    • Tabletop training: 190 scenes, 76k images, and 220k question-answer pairs.
    • Tabletop validation: 77 scenes, 355 images, and 3k question-answer pairs.

    ROBOSPATIAL-Val contains 6,000 heuristically generated questions from scans entirely unseen during training, with 2,000 questions each for configuration, context, and compatibility. Binary questions are evaluated by accuracy. For coordinate prediction, a predicted 3D location is counted as correct when it lies inside the convex hull of a reference point set derived from the scene geometry. The convex-hull rule gives a well-defined metric but is conservative because a point just outside the hull is marked incorrect.

    Out-of-domain evaluation uses ROBOSPATIAL-Home, which contains 350 manually written questions over RGB-D indoor scenes captured with an iPhone depth sensor; the spatial portion of BLINK, which tests binary relations such as next to, touching, and on top; and the position category of SpatialBench.

  7. Knowl 7 — 2D/3D VLM training and comparison protocol

    experimental setup

    The study evaluates RGB-only 2D VLMs and RGB-D or point-cloud 3D VLMs. The 2D models are VILA-1.5-8B, LLaVA-NeXT-8B, SpaceLLaVA-13B, RoboPoint-13B, Molmo-7B, and GPT-4o. SpaceLLaVA is a community implementation related to SpatialVLM, RoboPoint is specialized for spatial affordance points, and Molmo is designed for pointing and counting. GPT-4o is used as a closed-source baseline.

    The 3D models are 3D-LLM, which reconstructs colored 3D point clouds from multiple RGB views, and LEO, which operates on segmented colored point clouds of individual objects. Open-source models are evaluated both zero-shot and after fine-tuning on ROBOSPATIAL. Fine-tuning also uses the auxiliary object-grounding data to reduce failures caused by incorrect object reference resolution; the grounding data itself is not used to compute spatial-reasoning scores.

  8. Knowl 8 — ROBOSPATIAL substantially improves held-out in-domain spatial reasoning

    empirical result

    Fine-tuning on ROBOSPATIAL improves every reported task and environment score for every open-source model on ROBOSPATIAL-Val. The following values are averages over configuration, context, and compatibility, reported separately for indoor images, tabletop data, and the combined total:

    • VILA: indoor 43.1 to 64.8, tabletop 37.4 to 62.9, total 40.2 to 63.9.
    • LLaVA-NeXT: indoor 31.4 to 60.4, tabletop 29.2 to 60.5, total 30.3 to 60.5.
    • SpaceLLaVA: indoor 38.9 to 67.8, tabletop 46.2 to 63.6, total 43.6 to 65.7.
    • RoboPoint: indoor 39.6 to 71.0, tabletop 38.2 to 70.1, total 38.9 to 70.6.
    • 3D-LLM: indoor 37.6 to 63.1, tabletop 42.4 to 66.0, total 40.0 to 64.6.
    • LEO: indoor 41.9 to 73.1, tabletop 43.7 to 70.7, total 42.8 to 71.9.

    The un-fine-tuned reference scores for Molmo and GPT-4o are 50.1 and 50.8 overall, respectively. The strongest fine-tuned overall score is LEO at 71.9, while the strongest fine-tuned 2D score is RoboPoint at 70.6. These results show that the gains are not limited to a particular input modality or to a model already specialized for spatial pointing.

  9. Knowl 9 — Generalization to new environments, language, and implicit reference frames

    empirical result

    ROBOSPATIAL fine-tuning transfers across environment types: training on indoor scenes improves performance on tabletop data, and training on tabletop data improves performance on indoor data. The reported cross-environment scores increase from 38.7 to 48.9 and from 38.2 to 51.3 for RoboPoint, and from 41.9 to 47.2 and from 43.7 to 54.5 for LEO, for the two reported transfer directions.

    The gains also extend beyond the generated validation distribution. On ROBOSPATIAL-Home, VILA improves from 57.8 to 65.9 for configuration, 0.0 to 15.6 for context, and 69.0 to 78.0 for compatibility; LLaVA-NeXT improves from 60.2 to 76.4, 0.0 to 19.7, and 71.4 to 80.1; SpaceLLaVA improves from 61.0 to 71.6, 0.08 to 13.1, and 34.3 to 62.0; and RoboPoint improves from 59.4 to 70.0, 2.5 to 21.3, and 80.1 to 88.6. On BLINK, the corresponding accuracy changes are 72.7 to 79.7 for VILA, 71.3 to 79.0 for LLaVA-NeXT, 76.2 to 81.8 for SpaceLLaVA, and 63.6 to 70.6 for RoboPoint. On SpatialBench, they are 53.0 to 73.6, 55.9 to 70.6, 47.1 to 67.7, and 44.1 to 64.7, respectively.

    The authors report that ROBOSPATIAL-trained models can handle prepositions absent from training, such as under, next to, and far away, by mapping them to learned directional and proximity primitives. They also often infer an intended object-centric frame when ROBOSPATIAL-Home questions omit an explicit frame, suggesting that the training data associates object geometry and orientation with spatial language.

  10. Knowl 10 — Real-robot manipulation validation and deployment caveats

    empirical result

    The authors test spatial predictions in tabletop manipulation with a Kinova Jaco arm and a ZED2 RGB-D camera. The system uses a modular design: a VLM predicts a spatial answer or target point, and cuRobo separately plans collision-free robot motion for picking and placing. The experiment uses colored cubes, cylinders, food items, and toys to reduce object-recognition difficulty, and comprises more than 200 model queries.

    Reported task success rates are 23.7% for LLaVA-NeXT, 52.6% for LLaVA-NeXT fine-tuned on ROBOSPATIAL, 44.7% for RoboPoint, 46.2% for RoboPoint fine-tuned on ROBOSPATIAL, 43.8% for Molmo, and 46.9% for GPT-4o. ROBOSPATIAL-trained LLaVA-NeXT achieves the highest success rate. Qualitative trials indicate that fine-tuning helps align “in front of” with an object’s head direction and place objects at a scale-appropriate distance, whereas RoboPoint often places objects too far from the reference object.

    The paper identifies two important deployment limitations. A small 2-pixel error in a 2D image prediction can become a 5–10 cm physical placement error, and some 2D failures arise during projection from image coordinates into 3D. In addition, the apparent advantage of 3D VLMs is not conclusive because 3D-LLM and LEO have pretraining exposure to RGB-D indoor datasets that overlap with source environments. ROBOSPATIAL is therefore designed to support future controlled comparisons and 3D models that operate on partial rather than complete scans.

Coverage note — No substantial contributed material was omitted; detailed per-task entries in the in-domain tables were condensed into their task-averaged results, while the main cross-domain, out-of-domain, and robot findings were retained.

References

  1. 1.Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. Scanqa: 3d question answering for spatial scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  2. 2.Wenxiao Cai, Yaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo Zhao. Spatialbot: Precise spatial understanding with vision language models. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), 2025.
  3. 3.Alexandre Campeau-Lecours, Hugo Lamontagne, Simon Latour, Philippe Fauteux, Veronique Maheu, François Boucher, Charles Deguire, and Louis-Joseph Caron L’Ecuyer. Kinova modular robot arms for service robotics applications. Int. J. Robot. Appl. Technol., 5(2):49–71, 2017.
  4. 4.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. International Conference on 3D Vision (3DV), 2017.
  5. 5.Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. Spatialvlm: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14455–14465, 2024.
  6. 6.An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. Spatialrgpt: Grounded spatial reasoning in vision-language models. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  7. 7.Open X-Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Andrey Kolobov, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Scholkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Buchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Felipe Vieira Frujeri, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guangwen Yang, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homanga Bharadhwaj, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jasmine Hsu, Jay Vakil, Jeannette Bohg, Jeffrey Bingham, Jeffrey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, Joao Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi ”Jim” Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R Sanketi, Patrick ”Tree” Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Mart’in-Mart’in, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shubham Tulsiani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vikash Kumar, Vincent Vanhoucke, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiangyu Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yansong Pang, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yongqiang Dou, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, Zipeng Fu, and Zipeng Lin. Open X-Embodiment: Robotic learning datasets and RT-X models. https://arxiv.org/abs/2310.08864, 2023.
  8. 8.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017.
  9. 9.Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz, Aaron Sarnat, Byron Bischoff, Pete Walsh, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Jen Dumas, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  10. 10.Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang, and Zhongyu Wei. EmbSpatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 346–355, Bangkok, Thailand, 2024. Association for Computational Linguistics.
  11. 11.Jiafei Duan, Wilbert Pumacay, Nishanth Kumar, Yi Ru Wang, Shulin Tian, Wentao Yuan, Ranjay Krishna, Dieter Fox, Ajay Mandlekar, and Yijie Guo. AHA: A vision-language-model for detecting and reasoning over failures in robotic manipulation. In The Thirteenth International Conference on Learning Representations, 2025.
  12. 12.Hao-Shu Fang, Chenxi Wang, Minghao Gou, and Cewu Lu. Graspnet-1billion: A large-scale benchmark for general object grasping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11444–11453, 2020.
  13. 13.Kuan Fang, Fangchen Liu, Pieter Abbeel, and Sergey Levine. Moka: Open-world robotic manipulation through mark-based visual prompting. Robotics: Science and Systems (RSS), 2024.
  14. 14.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  15. 15.Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A. Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. In Proceedings of the European Conference on Computer Vision (ECCV), pages 148–166. Springer, 2024.
  16. 16.Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang, and Rakesh Ranjan. Mmg-ego4d: Multi-modal generalization in egocentric action recognition. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6481–6491, 2023.
  17. 17.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselassie, Cristina Gonzalez, James Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jachym Kolař, Satwik Kotur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Ziwei Zhao, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fuegen, Bernard Ghanem, Vamsi Krishna Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik. Ego4d: Around the world in 3,000 hours of egocentric video. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18973–18990, 2022.
  18. 18.Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Injecting the 3d world into large language models. In Advances in Neural Information Processing Systems, 2023. NeurIPS.
  19. 19.Haoxu Huang, Fanqi Lin, Yingdong Hu, Shengjie Wang, and Yang Gao. Copa: General robotic manipulation through spatial constraints of parts with foundation models. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 9488–9495, 2024.
  20. 20.Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. In Proceedings of the International Conference on Machine Learning (ICML), 2024.
  21. 21.Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Li Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation. In 8th Annual Conference on Robot Learning, 2024.
  22. 22.Yifei Huang, Jilan Xu, Baoqi Pei, Yuping He, Guo Chen, Mingfang Zhang, Lijin Yang, Zheng Nie, Jinyao Liu, Guoshun Fan, Dechen Lin, Fang Fang, Kunpeng Li, Chang Yuan, Xinyuan Chen, Yaohui Wang, Yali Wang, Yu Qiao, and Limin Wang. An egocentric vision-language model based portable real-time smart assistant, 2025.
  23. 23.Drew A. Hudson and Christopher D. Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6700–6709, 2019.
  24. 24.Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Sceneverse: Scaling 3d vision-language learning for grounded scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), 2024.
  25. 25.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  26. 26.Amita Kamath, Jack Hessel, and Kai-Wei Chang. What’s “up” with vision-language models? investigating their struggle with spatial reasoning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023.
  27. 27.Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla: An open-source vision-language-action model. In Proceedings of the 8th Conference on Robot Learning (CoRL), pages 2679–2713. PMLR, 2025.
  28. 28.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017.
  29. 29.Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500, 2023.
  30. 30.Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. VILA: On Pre-training for Visual Language Models . In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26679–26689, Los Alamitos, CA, USA, 2024. IEEE Computer Society.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollar. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV), 2014.
  32. 32.Xiongkun Linghu, Jiangyong Huang, Xuesong Niu, Xiaojian Ma, Baoxiong Jia, and Siyuan Huang. Multi-modal situated reasoning in 3d scenes. In Advances in Neural Information Processing Systems, 2024. NeurIPS.
  33. 33.Fangyu Liu, Guy Emerson, and Nigel Collier. Visual spatial reasoning. Transactions of the Association for Computational Linguistics, 11:635–651, 2023.
  34. 34.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems, 2023.
  35. 35.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296–26306, 2024.
  36. 36.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision (ECCV), 2024.
  37. 37.Xiaojian Ma, Silong Yong, Zilong Zheng, Qing Li, Yitao Liang, Song-Chun Zhu, and Siyuan Huang. Sqa3d: Situated question answering in 3d scenes. In International Conference on Learning Representations, 2023.
  38. 38.Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton, Sasha Sax, and Aravind Rajeswaran. Openeqa: Embodied question answering in the era of foundation models. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  39. 39.Yunze Man, Liang-Yan Gui, and Yu-Xiong Wang. Situational awareness matters in 3d vision language reasoning. In CVPR, 2024.
  40. 40.Soroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao, Jacky Liang, Ishita Dasgupta, Annie Xie, Danny Driess, Ayzaan Wahid, Zhuo Xu, Quan Vuong, Tingnan Zhang, Tsang-Wei Edward Lee, Kuang-Huei Lee, Peng Xu, Sean Kirmani, Yuke Zhu, Andy Zeng, Karol Hausman, Nicolas Heess, Chelsea Finn, Sergey Levine, and Brian Ichter. Pivot: iterative visual prompting elicits actionable knowledge for vlms. In Proceedings of the International Conference on Machine Learning (ICML). JMLR.org, 2024.
  41. 41.Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands, 2024.
  42. 42.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simon Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mely, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Ceron Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024.
  43. 43.Navid Rajabi and Jana Kosecka. Towards grounded visual spatial reasoning in multi-modal vision language models, 2024.
  44. 44.Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, and Tsung-Yu Lin. Learning to localize objects improves spatial reasoning in visual-llms. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12977–12987, 2024.
  45. 45.Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Radle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollar, and Christoph Feichtenhofer. SAM 2: Segment anything in images and videos. In The Thirteenth International Conference on Learning Representations, 2025.
  46. 46.Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch, and Zeynep Akata. Clevr-x: A visual reasoning dataset for natural language explanations. In xxAI - Beyond explainable Artificial Intelligence, pages 85–104. Springer, 2022.
  47. 47.Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu, Reza Haf, and Yuan-Fang Li. An empirical analysis on spatial reasoning capabilities of large multimodal models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21440–21455, Miami, Florida, USA, 2024. Association for Computational Linguistics.
  48. 48.Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 11523–11530, 2023.
  49. 49.Chan Hee Song, Jihyung Kil, Tai-Yu Pan, Brian M. Sadler, Wei-Lun Chao, and Yu Su. One step at a time: Long-horizon vision-and-language navigation with milestones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15482–15491, 2022.
  50. 50.Alessandro Suglia, Claudio Greco, Katie Baker, Jose L. Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon. AlanaVLM: A multimodal embodied AI foundation model for egocentric video understanding. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 11101–11122, Miami, Florida, USA, 2024. Association for Computational Linguistics.
  51. 51.Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. A corpus for reasoning about natural language grounded in photographs. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6418–6428, Florence, Italy, 2019. Association for Computational Linguistics.
  52. 52.Balakumar Sundaralingam, Siva Kumar Sastry Hari, Adam Fishman, Caelan Garrett, Karl Van Wyk, Valts Blukis, Alexander Millane, Helen Oleynikova, Ankur Handa, Fabio Ramos, et al. cuRobo: Parallelized collision-free minimum-jerk robot motion generation. arXiv preprint arXiv:2310.17274, 2023.
  53. 53.Emilia Szymanska, Mihai Dusmanu, Jan-Willem Buurlage, Mahdi Rad, and Marc Pollefeys. Space3D-Bench: Spatial 3D Question Answering Benchmark. In European Conference on Computer Vision (ECCV) Workshops, 2024.
  54. 54.Stephen Tyree, Jonathan Tremblay, Thang To, Jia Cheng, Terry Mosier, Jeffrey Smith, and Stan Birchfield. 6-dof pose estimation of household objects for robotic manipulation: An accessible dataset and benchmark. In International Conference on Intelligent Robots and Systems (IROS), 2022.
  55. 55.Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi. Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration. IEEE Robotics and Automation Letters, 9(11):10567–10574, 2024.
  56. 56.Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, and Matthias Niessner. Rio: 3d object instance re-localization in changing indoor environments. In Proceedings IEEE International Conference on Computer Vision (ICCV), 2019.
  57. 57.Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Yixuan Li, and Neel Joshi. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  58. 58.Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, and Jiangmiao Pang. Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  59. 59.Youngsun Wi, Mark Van der Merwe, Pete Florence, Andy Zeng, and Nima Fazeli. Calamari: Contact-aware and language conditioned spatial action mapping for contact-rich manipulation. In 7th Annual Conference on Robot Learning, 2023.
  60. 60.Liuchang Xu, Shuo Zhao, Qingming Lin, Luyao Chen, Qianqian Luo, Sensen Wu, Xinyue Ye, Hailin Feng, and Zhenhong Du. Evaluating large language models on spatial tasks: A multi-task benchmarking study, 2024.
  61. 61.Yutaro Yamada, Yihan Bao, Andrew K. Lampinen, Jungo Kasai, and Ilker Yildirim. Evaluating spatial understanding of large language models, 2024.
  62. 62.Wentao Yuan, Jiafei Duan, Valts Blukis, Wilbert Pumacay, Ranjay Krishna, Adithyavairavan Murali, Arsalan Mousavian, and Dieter Fox. Robopoint: A vision-language model for spatial affordance prediction in robotics. In 8th Annual Conference on Robot Learning, 2024.
  63. 63.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  64. 64.Yue Zhang, Zhiyang Xu, Ying Shen, Parisa Kordjamshidi, and Lifu Huang. SPARTUN3d: Situated spatial understanding of 3d world in large language model. In The Thirteenth International Conference on Learning Representations, 2025.
  65. 65.Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, Michael S. Ryoo, Krista Reymann, Kanishka Rao, Karl Pertsch, Igor Mordatch, Henryk Michalewski, Yao Lu, Sergey Levine, Lisa Lee, Tsang-Wei Edward Lee, Isabel Leal, Yuheng Kuang, Dmitry Kalashnikov, Ryan Julian, Nikhil J. Joshi, Alex Irpan, Brian Ichter, Jasmine Hsu, Alexander Herzog, Karol Hausman, Keerthana Gopalakrishnan, Chuyuan Fu, Pete Florence, Chelsea Finn, Kumar Avinava Dubey, Danny Driess, Tianli Ding, Krzysztof Marcin Choromanski, Xi Chen, Yevgen Chebotar, Justice Carbajal, Noah Brown, Anthony Brohan, Montserrat Gonzalez Arenas, and Kehang Han. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, pages 2165–2183. PMLR, 2023.
  66. 66.Alex Zook, Fan-Yun Sun, Josef Spjut, Valts Blukis, Stan Birchfield, and Jonathan Tremblay. Grs: Generating robotic simulation tasks from real-world images, 2024.

Citation

MLA
Song, C. H., et al. “RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 15768–80, https://doi.org/10.1109/CVPR52734.2025.01470.
APA
Song, C. H., Blukis, V., Tremblay, J., Tyree, S., Su, Y., & Birchfield, S. (2025). RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15768–15780. https://doi.org/10.1109/CVPR52734.2025.01470
Chicago
Song, C. H., V. Blukis, J. Tremblay, S. Tyree, Y. Su, and S. Birchfield. 2025. “RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15768–80. https://doi.org/10.1109/CVPR52734.2025.01470.
Harvard
Song, C.H. et al. (2025) “RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 15768–15780. Available at: https://doi.org/10.1109/CVPR52734.2025.01470.
Vancouver
1. Song CH, Blukis V, Tremblay J, Tyree S, Su Y, Birchfield S (2025) RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 15768–15780

BibTeX

@inproceedings{Song_2025, title={RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01470}, DOI={10.1109/cvpr52734.2025.01470}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Song, Chan Hee and Blukis, Valts and Tremblay, Jonathan and Tyree, Stephen and Su, Yu and Birchfield, Stan}, year={2025}, month=June, pages={15768–15780} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE