An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
Fatemeh ShiriXiao-Yu GuoMona FarXin YuReza HafYuan-Fang Li
Presents the Spatial-MM benchmark to expose critical weaknesses in large multimodal models, showing that while symbolic aids like bounding boxes improve performance, models still fail on human-perspective viewpoints and gain no benefit from chain-of-thought prompting on complex spatial questions.
Modern artificial intelligence systems increasingly rely on Large Multimodal Models (LMMs) to interpret both visual images and text. While these models perform well on broad vision-and-language tasks, their ability to understand spatial arrangements and reason through complex physical relationships remains unreliable. To address this critical gap, the article introduces Spatial-MM, a new benchmark designed to evaluate spatial reasoning across both single-step and multi-hop scenarios. The evaluation assesses leading commercial and open-source models—including GPT-4o, GPT-4 Vision, Gemini 1.5 Pro, and LLaVA-1.5—across different spatial viewpoints, visual grounding techniques, and structured reasoning paths.
The experimental findings show significant limitations in how current models handle spatial data. First, models perform markedly worse when required to answer questions from the perspective of an entity inside the image (such as an observed human) compared to the standard camera viewpoint; for example, GPT-4o's accuracy drops from over 63% from the camera view to 27.5% from an internal human perspective. Second, traditional text-based chain-of-thought prompting fails to improve spatial accuracy in multi-hop questions, with models often performing worse than under standard prompting. Third, detailed failure analysis demonstrates that when multi-hop reasoning fails, an incorrect spatial step is responsible roughly 91% of the time, whereas non-spatial steps rarely cause errors. Conversely, providing structured visual grounding—such as synthesized bounding boxes or scene graphs—substantially improves accuracy across all evaluated models.
These results demonstrate that while current AI systems excel at basic object recognition, they cannot be trusted to independently manage complex spatial relationships or perspective-shifting tasks. Relying on raw multimodal models for spatial decision-making in fields like robotics, autonomous systems, or spatial navigation creates substantial operational and safety risks. To mitigate these risks, systems requiring spatial reasoning should integrate explicit visual scaffolding, such as automated scene graphs or localized bounding boxes, rather than relying on chain-of-thought textual prompting alone. Future development should focus on expanding larger-scale multi-hop spatial training data and refining models' intrinsic perspective-taking capabilities.
- Paper: GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, Drew A. Hudson et al. (2019). Spatial-MM compares models against GQA-spatial, so GQA’s scene-graph-based benchmark provides the underlying framework for interpreting that comparison.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). RoboSpatial turns the source’s diagnosis of viewpoint-sensitive spatial errors into a robotics training approach that tests relations across observer-, world-, and object-centric frames.
- Paper: HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models, Huizhi Liang et al. (2026). HiSpatial extends the source’s benchmarked spatial deficits into a hierarchical 3D training framework spanning geometry, object relations, and multi-step reasoning.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). Thinking in Space carries the source’s evaluation of spatial reasoning into video, testing whether models can perceive and recall layouts across real indoor scenes.
- Paper: VLM4D: Towards Spatiotemporal Awareness in Vision Language Models, Shijie Zhou et al. (2025). VLM4D extends the source’s concern with perspective shifts to dynamic scenes by evaluating spatial and temporal reasoning across video viewpoints.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT develops bounding-box-guided visual reasoning, extending the source’s finding that explicit visual grounding can outperform text-only reasoning prompts.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). CoT-VLA applies visual intermediate reasoning to robot actions, continuing the source’s investigation of how reasoning methods affect spatial performance in physical tasks.
- Paper: CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates, Shresth Grover et al. (2025). CoSPlan carries the source’s scene-graph grounding recommendation into sequential planning, testing structured relations as models correct errors in evolving scenes.
- Paper: MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency, Dongzhi Jiang et al. (2025). MME-CoT extends the source’s warning about unreliable chain-of-thought by evaluating the quality, robustness, and efficiency of intermediate multimodal reasoning steps.
- Paper: EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents, Rui Yang 0010 et al. (2025). EmbodiedBench applies spatial-awareness evaluation to vision-driven agents, extending the source’s capability findings toward navigation and physical control.
