Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin QuHaozhe QiMihai DusmanuMahdi RadRui WangMarc Pollefeys
Proposes Loc3R-VLM, a framework that equips 2D vision-language models with 3D spatial reasoning and language-based localization capabilities from monocular video by coupling global layout reconstruction with viewpoint-aware situation modeling.
Artificial intelligence systems increasingly need to navigate and understand complex physical environments, such as in robotics and autonomous driving. However, modern multimodal vision-language models largely process visual information as isolated two-dimensional images. As a result, they struggle to maintain a persistent three-dimensional awareness of their surroundings and fail to reason effectively from dynamic, viewpoint-dependent perspectives. Existing approaches attempt to overcome this by relying on dense three-dimensional sensor data like point clouds or depth maps during operation, which are rarely available in real-world deployments and fail to scale.
The article introduces and evaluates Loc3R-VLM, a framework designed to give two-dimensional vision-language models human-like spatial reasoning and localization capabilities directly from standard monocular video, without requiring any three-dimensional data at inference time.
To achieve this, the authors developed a multimodal learning framework built on an open-source vision-language model. Drawing inspiration from human spatial cognition, the method integrates two auxiliary training tasks: reconstructing a global bird's-eye-view layout of the scene to serve as an internal mental map, and explicit situation modeling using dedicated tokens to predict an agent's position and orientation from natural language descriptions. The architecture also incorporates lightweight camera pose priors extracted from a pre-trained geometric foundation model to resolve scale ambiguity. The system was trained end-to-end using standard indoor benchmarks, such as ScanNet-based datasets, and evaluated across localization and spatial question-answering tasks.
The article demonstrates several key findings:
- In language-based localization on the SQA3D benchmark, the framework achieved state-of-the-art results, reaching 42.6% accuracy within 0.5 meters and 75.9% within 1.0 meter. It substantially outperformed prior systems relying on dense point clouds, exceeding the strongest point-cloud baseline by 25.2 percentage points at the 0.5-meter threshold and by 39.0 percentage points at the 1.0-meter threshold.
- In viewpoint-aware reasoning on VSI-Bench, the model achieved an overall accuracy score of 63.2%, outperforming leading generalist models such as GPT-4o and Gemini-1.5-Pro. The most pronounced gains appeared in viewpoint-dependent tasks, including relative direction, where accuracy improved by 36.1 percentage points over the next best generalist baseline.
- On general and situated question-answering benchmarks, including ScanQA, MSQA, and Beacon3D, the model consistently ranked highest among two-dimensional vision-language approaches, improving spatial subcategory accuracy on MSQA and Beacon3D by 11.1 and 9.4 percentage points, respectively.
- Ablation analyses confirmed that both the internal layout reconstruction and the situation modeling modules contribute synergistically to spatial grounding, while the camera pose priors are essential for metric-scale accuracy.
These findings indicate that explicit spatial supervision allows models to build reliable internal representations of three-dimensional space from ordinary video feeds. By eliminating the requirement for specialized three-dimensional sensors or ground-truth depth during deployment, this approach significantly reduces operational complexity and hardware costs for spatial artificial intelligence. It also provides a scalable foundation for embodied agents to safely interpret human verbal instructions in physical spaces.
Based on these results, engineering teams developing embodied agents, home robotics, or spatial reasoning assistants should consider integrating explicit spatial and situational supervision into video-language pipelines rather than relying solely on language-only instruction tuning. Before wide-scale operational deployment, further research and pilot testing are recommended to address key limitations: the current two-dimensional bird's-eye-view projection discards vertical granularity, which can limit reasoning across vertically stacked objects or multi-level environments; the fixed 32-frame sampling can create visual blind spots in large rooms; and the evaluation remains restricted to static indoor settings. Overall, the methodology shows high empirical reliability and confidence within the scope of static indoor environments.
- Paper: Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization, Siyan Dong et al. (2025). Reloc3r establishes foundational methods for fast and generalizable relative camera pose estimation that inform the camera pose priors and visual localization objectives utilized in Loc3R-VLM.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). This benchmark study exposes the critical spatial reasoning deficiencies of multimodal large language models on video inputs, establishing the specific problem setting and motivation addressed by Loc3R-VLM.
- Paper: RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics, Chan Hee Song et al. (2025). RoboSpatial investigates multi-coordinate spatial reasoning across observer-centric and world-centric frames, laying the groundwork for situation modeling and spatial supervision in vision-language architectures.
- Paper: An Embodied Generalist Agent in 3D World, Jiangyong Huang et al. (2024). LEO provides a comprehensive paradigm for training multimodal agents on 3D perception and situated question answering, serving as a primary point of comparison and foundation for Loc3R-VLM.
- Paper: Multi-View Transformer for 3D Visual Grounding, Shijia Huang et al. (2022). This paper presents multi-view representations to resolve human vantage-point discrepancies in 3D visual grounding, motivating Loc3R-VLM's explicit situation modeling from monocular video sequences.
- Paper: Language Conditioned Spatial Relation Reasoning for 3D Object Grounding, Shizhe Chen et al. (2022). ViL3DRel establishes explicit geometric relation modeling between objects in 3D space, informing how spatial grounding supervision can be structured alongside linguistic expressions.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision delivers a unified baseline architecture for video and single-image reasoning upon which monocular 3D supervision and reasoning frameworks build.
- Paper: HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models, Huizhi Liang et al. (2026). HiSpatial extends the paradigm of 3D-aware vision-language understanding by organizing spatial perception and reasoning into a formal four-level cognitive hierarchy supported by large-scale synthetic QA generation.
- Paper: SG2Loc: Sequential Visual Localization on 3D Scene Graphs, Nicole Damblon et al. (2026). SG2Loc advances sequential localization and spatial representation by utilizing lightweight 3D scene graphs instead of dense geometric reconstruction.
