OVI-MAP:Open-Vocabulary Instance-Semantic Mapping

Zilong DengFederico TombariMarc PollefeysJohanna WaldDániel Baráth

article2026arXiv4 citations

Presents a real-time open-vocabulary 3D instance mapping system that decouples geometry reconstruction from semantic inference using selective-view vision-language querying, eliminating dense feature fusion while maintaining high temporal consistency during online exploration.

Listen

Autonomous robots and spatial computing platforms require real-time 3D scene understanding that can recognize and segment novel objects without relying on predefined label sets. Existing open-vocabulary mapping methods frequently fail in operational settings because they store large, continuous language embeddings at every 3D point or tie object segmentation directly to fixed semantic classes. These practices result in severe memory consumption, high processing latency, and unstable object tracking when encountering unseen environments.

The article demonstrates OVI-MAP, an online 3D mapping pipeline that decouples geometric instance reconstruction from semantic inference. By evaluating this system on standard indoor benchmarks, the article shows that separating class-agnostic physical object mapping from vision-language reasoning delivers accurate, real-time, open-vocabulary scene understanding.

The approach constructs a globally consistent 3D instance map from streaming RGB-D video using class-agnostic 2D segmentation refined by depth boundaries and integrated via spatial voting. Instead of querying a large vision-language model on every incoming video frame, the system introduces an object-centric view selection mechanism that tracks spherical coverage around each 3D object and extracts language embeddings only when observing a substantially novel viewpoint. Semantic features from these selected views are merged using visibility weighting, and the entire framework was validated against leading online and offline mapping approaches across the Replica and ScanNet datasets on standard hardware.

Key findings show that OVI-MAP outperforms existing online open-vocabulary systems while maintaining strict real-time performance. First, in instance segmentation, the system achieved 50.8% average precision at 50% intersection-over-union on Replica, more than doubling the 23.6% scored by the online baseline OVO-SLAM. Second, in open-vocabulary semantic segmentation under a real-time 30 frames-per-second constraint, the system reached 27.0% mean intersection-over-union on Replica and 16.3% on ScanNet, exceeding online alternatives and rivaling complex offline pipelines. Third, the object-centric view coverage strategy reduced vision-language queries by 53% compared to conventional pixel-counting heuristics while matching or improving semantic accuracy.

These results demonstrate that robotics and mixed-reality systems can achieve flexible, open-world comprehension without costly high-end computing clusters or proprietary closed-set models. Eliminating dense per-voxel language storage substantially reduces compute risk, memory overhead, and frame drops, enabling reliable natural language queries and navigation in unconstrained real-world settings.

For engineering and operational deployment, teams should adopt decoupled architectures that isolate geometric tracking from semantic extraction and implement view-diversity filters to minimize inference overhead. Further development should prioritize running pilot evaluations on physical mobile robots to assess real-world latency, dynamic scene handling, and edge computing constraints.

Confidence in these findings is high for structured indoor environments, as supported by standard benchmark datasets. However, stakeholders should note that the pipeline's overall performance remains constrained by the accuracy of the underlying 2D image segmentor on small or visually complex objects, as well as the alignment quality of existing vision-language models.

arXiv: 2603.26541OVI-MAP/OVI-MAP
Cover for OVI-MAP:Open-Vocabulary Instance-Semantic Mapping

Abstract

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely on the closed-set assumption or dense per-pixel language fusion, which limits scalability and temporal consistency. We introduce OVI-MAP that decouples instance reconstruction from semantic inference. We propose to build a class-agnostic 3D instance map that is incrementally constructed from RGB-D input, while semantic features are extracted only from a small set of automatically selected views using vision-language models. This design enables stable instance tracking and zero-shot semantic labeling throughout online exploration. Our system operates in real time and outperforms state-of-the-art open-vocabulary mapping baselines on standard benchmarks.

Citation

MLA
Deng, Z., et al. “OVI-MAP:Open-Vocabulary Instance-Semantic Mapping”. arXiv, 2026, https://doi.org/10.48550/arxiv.2603.26541.
APA
Deng, Z., Tombari, F., Pollefeys, M., Wald, J., & Barath, D. (2026). OVI-MAP:Open-Vocabulary Instance-Semantic Mapping. arXiv. https://doi.org/10.48550/arxiv.2603.26541
Chicago
Deng, Z., F. Tombari, M. Pollefeys, J. Wald, and D. Barath. 2026. “OVI-MAP:Open-Vocabulary Instance-Semantic Mapping”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2603.26541.
Harvard
Deng, Z. et al. (2026) “OVI-MAP:Open-Vocabulary Instance-Semantic Mapping”. arXiv. Available at: https://doi.org/10.48550/arxiv.2603.26541.
Vancouver
1. Deng Z, Tombari F, Pollefeys M, Wald J, Barath D (2026) OVI-MAP:Open-Vocabulary Instance-Semantic Mapping. https://doi.org/10.48550/arxiv.2603.26541

BibTeX

@misc{https://doi.org/10.48550/arxiv.2603.26541,
  doi = {10.48550/ARXIV.2603.26541},
  url = {https://arxiv.org/abs/2603.26541},
  author = {Deng, Zilong and Tombari, Federico and Pollefeys, Marc and Wald, Johanna and Barath, Daniel},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {OVI-MAP:Open-Vocabulary Instance-Semantic Mapping},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/