Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

Kevin QuHaozhe QiMihai DusmanuMahdi RadRui WangMarc Pollefeys

article2026arXiv2 citations

Proposes Loc3R-VLM, a framework that equips 2D vision-language models with 3D spatial reasoning and language-based localization capabilities from monocular video by coupling global layout reconstruction with viewpoint-aware situation modeling.

Listen

Artificial intelligence systems increasingly need to navigate and understand complex physical environments, such as in robotics and autonomous driving. However, modern multimodal vision-language models largely process visual information as isolated two-dimensional images. As a result, they struggle to maintain a persistent three-dimensional awareness of their surroundings and fail to reason effectively from dynamic, viewpoint-dependent perspectives. Existing approaches attempt to overcome this by relying on dense three-dimensional sensor data like point clouds or depth maps during operation, which are rarely available in real-world deployments and fail to scale.

The article introduces and evaluates Loc3R-VLM, a framework designed to give two-dimensional vision-language models human-like spatial reasoning and localization capabilities directly from standard monocular video, without requiring any three-dimensional data at inference time.

To achieve this, the authors developed a multimodal learning framework built on an open-source vision-language model. Drawing inspiration from human spatial cognition, the method integrates two auxiliary training tasks: reconstructing a global bird's-eye-view layout of the scene to serve as an internal mental map, and explicit situation modeling using dedicated tokens to predict an agent's position and orientation from natural language descriptions. The architecture also incorporates lightweight camera pose priors extracted from a pre-trained geometric foundation model to resolve scale ambiguity. The system was trained end-to-end using standard indoor benchmarks, such as ScanNet-based datasets, and evaluated across localization and spatial question-answering tasks.

The article demonstrates several key findings:

  1. In language-based localization on the SQA3D benchmark, the framework achieved state-of-the-art results, reaching 42.6% accuracy within 0.5 meters and 75.9% within 1.0 meter. It substantially outperformed prior systems relying on dense point clouds, exceeding the strongest point-cloud baseline by 25.2 percentage points at the 0.5-meter threshold and by 39.0 percentage points at the 1.0-meter threshold.
  2. In viewpoint-aware reasoning on VSI-Bench, the model achieved an overall accuracy score of 63.2%, outperforming leading generalist models such as GPT-4o and Gemini-1.5-Pro. The most pronounced gains appeared in viewpoint-dependent tasks, including relative direction, where accuracy improved by 36.1 percentage points over the next best generalist baseline.
  3. On general and situated question-answering benchmarks, including ScanQA, MSQA, and Beacon3D, the model consistently ranked highest among two-dimensional vision-language approaches, improving spatial subcategory accuracy on MSQA and Beacon3D by 11.1 and 9.4 percentage points, respectively.
  4. Ablation analyses confirmed that both the internal layout reconstruction and the situation modeling modules contribute synergistically to spatial grounding, while the camera pose priors are essential for metric-scale accuracy.

These findings indicate that explicit spatial supervision allows models to build reliable internal representations of three-dimensional space from ordinary video feeds. By eliminating the requirement for specialized three-dimensional sensors or ground-truth depth during deployment, this approach significantly reduces operational complexity and hardware costs for spatial artificial intelligence. It also provides a scalable foundation for embodied agents to safely interpret human verbal instructions in physical spaces.

Based on these results, engineering teams developing embodied agents, home robotics, or spatial reasoning assistants should consider integrating explicit spatial and situational supervision into video-language pipelines rather than relying solely on language-only instruction tuning. Before wide-scale operational deployment, further research and pilot testing are recommended to address key limitations: the current two-dimensional bird's-eye-view projection discards vertical granularity, which can limit reasoning across vertically stacked objects or multi-level environments; the fixed 32-frame sampling can create visual blind spots in large rooms; and the evaluation remains restricted to static indoor settings. Overall, the methodology shows high empirical reliability and confidence within the scope of static indoor environments.

arXiv: 2603.18002kevinqu7/Loc3R-VLM
Cover for Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

Abstract

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input representations with geometric cues rather than explicitly teaching models to reason in 3D space. We introduce Loc3R-VLM, a framework that equips 2D Vision-Language Models with advanced 3D understanding capabilities from monocular video input. Inspired by human spatial cognition, Loc3R-VLM relies on two joint objectives: global layout reconstruction to build a holistic representation of the scene structure, and explicit situation modeling to anchor egocentric perspective. These objectives provide direct spatial supervision that grounds both perception and language in a 3D context. To ensure geometric consistency and metric-scale alignment, we leverage lightweight camera pose priors extracted from a pre-trained 3D foundation model. Loc3R-VLM achieves state-of-the-art performance in language-based localization and outperforms existing 2D- and video-based approaches on situated and general 3D question-answering benchmarks, demonstrating that our spatial supervision framework enables strong 3D understanding. Project page: this https URL

Citation

MLA
Qu, K., et al. “Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models”. arXiv, 2026, https://doi.org/10.48550/arxiv.2603.18002.
APA
Qu, K., Qi, H., Dusmanu, M., Rad, M., Wang, R., & Pollefeys, M. (2026). Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models. arXiv. https://doi.org/10.48550/arxiv.2603.18002
Chicago
Qu, K., H. Qi, M. Dusmanu, M. Rad, R. Wang, and M. Pollefeys. 2026. “Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2603.18002.
Harvard
Qu, K. et al. (2026) “Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2603.18002.
Vancouver
1. Qu K, Qi H, Dusmanu M, Rad M, Wang R, Pollefeys M (2026) Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models. https://doi.org/10.48550/arxiv.2603.18002

BibTeX

@misc{https://doi.org/10.48550/arxiv.2603.18002,
  doi = {10.48550/ARXIV.2603.18002},
  url = {https://arxiv.org/abs/2603.18002},
  author = {Qu, Kevin and Qi, Haozhe and Dusmanu, Mihai and Rad, Mahdi and Wang, Rui and Pollefeys, Marc},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/