RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
Chan Hee SongValts BlukisJonathan TremblayStephen TyreeYu SuStan Birchfield
Introduces RoboSpatial, a large-scale multimodal dataset with three million annotations across multiple reference frames and spatial compatibility tasks, enabling vision-language models to achieve superior spatial reasoning for robotic manipulation.
Vision-language models that combine visual perception with language understanding are increasingly used to direct robots in real-world tasks such as object manipulation, navigation, and automated planning. However, existing models frequently fail at basic spatial reasoning because they are trained on general internet images lacking precise scale, depth, and viewpoint context. Most current systems cannot reliably interpret directional relationships relative to a specific object, identify available free space, or determine whether an item physically fits into a target area, which creates substantial operational bottlenecks for autonomous robotic systems.
The article demonstrates that this performance gap stems primarily from a lack of appropriate training data and introduces a large-scale multimodal dataset, ROBOSPATIAL, to teach foundational spatial reasoning to both two-dimensional and three-dimensional vision-language models. The dataset pairs five thousand real-world three-dimensional indoor and tabletop scans with one million egocentric images to generate roughly three million spatial question-answer pairs. The automated data generation pipeline systematically extracts three core relationship types: spatial context, which identifies empty coordinates for placement; spatial compatibility, which verifies whether an object physically fits with a minimum safety margin; and spatial configuration, which evaluates relative positions. Crucially, every question is formulated across three distinct reference frames—observer-centric, world-centric, and object-centric.
Rigorous evaluations show that training on this dataset substantially boosts spatial reasoning across all tested models. On a held-out validation benchmark, fine-tuned models improved their overall performance by roughly twenty to thirty percentage points over baseline versions. In physical robot manipulation trials, a standard vision-language model trained on the dataset more than doubled its task execution success rate from 23.7% to 52.6%, outperforming leading closed-source models such as GPT-4o (46.9%). Models also demonstrated strong out-of-domain transfer, successfully interpreting novel spatial prepositions like "under" and "next to" and inferring implicit object orientations from everyday language.
These findings indicate that targeted, geometrically grounded training data can resolve fundamental reasoning limitations without requiring massive, ungrounded web datasets. For organizations developing robotic automation, enhanced spatial understanding lowers physical error rates, reduces collision risks, and expands the range of complex manipulation tasks that robots can handle from natural language instructions. Organizations pursuing autonomous robotics should integrate multi-perspective spatial supervision into their visual reasoning pipelines rather than relying exclusively on off-the-shelf generalist vision models.
While three-dimensional models showed initial performance advantages over two-dimensional counterparts, direct architectural comparisons remain somewhat inconclusive due to overlapping pretraining datasets. Additionally, in physical deployments, minor two-dimensional pixel prediction shifts translated to physical errors of five to ten centimeters. Future research and technical pilots should focus on refining coordinate projection accuracy, assessing performance under varied camera viewpoints, and training models on partial, sparse sensor scans to facilitate seamless deployment on real-world mobile robots.
- Paper: Language Conditioned Spatial Relation Reasoning for 3D Object Grounding, Shizhe Chen et al. (2022). This paper establishes foundational techniques for grounding natural language spatial relations and geometric configurations in 3D scenes, which directly motivates the multi-frame spatial reasoning datasets explored in RoboSpatial.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Understanding how web-scale vision-language models can be adapted into vision-language-action policies highlights the spatial reasoning bottlenecks that RoboSpatial seeks to resolve.
- Paper: Matterport3D: Learning from RGB-D Data in Indoor Environments, Angel Chang et al. (2017). This work introduces large-scale real-world indoor 3D RGB-D scans and reconstructions that underpin foundational 3D scene understanding pipelines.
- Paper: ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D Understanding, Le Xue et al. (2023). It provides the essential cross-modal pretraining paradigm aligning 2D image models, natural language, and 3D point cloud representations.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). This paper introduces foundational discrete coordinate tokens and grounded visual reasoning for vision-language models that RoboSpatial scales to 3D robotics contexts.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). It details how multimodal language models incorporate continuous embodied sensor inputs, providing context for the spatial limitations inherent to generalist robot planners.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). This study analyzes how visual backbone designs and spatial representations affect downstream localization and spatial understanding in vision-language models.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). This work builds on foundational 3D spatial reasoning by benchmarking how multimodal large language models perceive, remember, and recall spatial layouts over extended video sequences.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). It applies spatial visual reasoning directly to robotic control by having vision-language-action models generate intermediate visual subgoals before executing manipulation steps.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). This research extends spatial understanding by combining high-level vision-language reasoning with low-level future visual prediction to guide precise robotic manipulation.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). It translates spatial relation reasoning into actionable robotic control by structuring visual and textual affordances such as placement space and collision-free paths.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This review offers a broader perspective on the architectural trends and real-world deployment challenges of vision-language-action systems that integrate spatial reasoning into robot control.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). This benchmark broadens spatial intelligence evaluation from discriminative perception to generative text-to-image models across complex layouts.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). It scales open-world 3D spatial detection from 2D and depth inputs across diverse, unconstrained environments.
