OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
Zilong DengFederico TombariMarc PollefeysJohanna WaldDániel Baráth
Presents a real-time open-vocabulary 3D instance mapping system that decouples geometry reconstruction from semantic inference using selective-view vision-language querying, eliminating dense feature fusion while maintaining high temporal consistency during online exploration.
Autonomous robots and spatial computing platforms require real-time 3D scene understanding that can recognize and segment novel objects without relying on predefined label sets. Existing open-vocabulary mapping methods frequently fail in operational settings because they store large, continuous language embeddings at every 3D point or tie object segmentation directly to fixed semantic classes. These practices result in severe memory consumption, high processing latency, and unstable object tracking when encountering unseen environments.
The article demonstrates OVI-MAP, an online 3D mapping pipeline that decouples geometric instance reconstruction from semantic inference. By evaluating this system on standard indoor benchmarks, the article shows that separating class-agnostic physical object mapping from vision-language reasoning delivers accurate, real-time, open-vocabulary scene understanding.
The approach constructs a globally consistent 3D instance map from streaming RGB-D video using class-agnostic 2D segmentation refined by depth boundaries and integrated via spatial voting. Instead of querying a large vision-language model on every incoming video frame, the system introduces an object-centric view selection mechanism that tracks spherical coverage around each 3D object and extracts language embeddings only when observing a substantially novel viewpoint. Semantic features from these selected views are merged using visibility weighting, and the entire framework was validated against leading online and offline mapping approaches across the Replica and ScanNet datasets on standard hardware.
Key findings show that OVI-MAP outperforms existing online open-vocabulary systems while maintaining strict real-time performance. First, in instance segmentation, the system achieved 50.8% average precision at 50% intersection-over-union on Replica, more than doubling the 23.6% scored by the online baseline OVO-SLAM. Second, in open-vocabulary semantic segmentation under a real-time 30 frames-per-second constraint, the system reached 27.0% mean intersection-over-union on Replica and 16.3% on ScanNet, exceeding online alternatives and rivaling complex offline pipelines. Third, the object-centric view coverage strategy reduced vision-language queries by 53% compared to conventional pixel-counting heuristics while matching or improving semantic accuracy.
These results demonstrate that robotics and mixed-reality systems can achieve flexible, open-world comprehension without costly high-end computing clusters or proprietary closed-set models. Eliminating dense per-voxel language storage substantially reduces compute risk, memory overhead, and frame drops, enabling reliable natural language queries and navigation in unconstrained real-world settings.
For engineering and operational deployment, teams should adopt decoupled architectures that isolate geometric tracking from semantic extraction and implement view-diversity filters to minimize inference overhead. Further development should prioritize running pilot evaluations on physical mobile robots to assess real-world latency, dynamic scene handling, and edge computing constraints.
Confidence in these findings is high for structured indoor environments, as supported by standard benchmark datasets. However, stakeholders should note that the pipeline's overall performance remains constrained by the accuracy of the underlying 2D image segmentor on small or visually complex objects, as well as the alignment quality of existing vision-language models.
- Paper: Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance, Phuc D. A. Nguyen et al. (2024). It introduces 2D-guided open-vocabulary 3D instance segmentation that OVI-MAP builds upon and adapts into an efficient online mapping pipeline.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). It establishes foundational techniques for lifting frame-by-frame 2D segmentations into globally consistent 3D instance representations.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). It provides a clear baseline for dense 2D vision-language feature distillation into 3D space, which OVI-MAP overcomes by decoupling geometric reconstruction from semantic inference.
- Paper: SG2Loc: Sequential Visual Localization on 3D Scene Graphs, Nicole Damblon et al. (2026). It applies object-centric 3D semantic representations to enable efficient, lightweight visual localization without dense point clouds.
- Paper: Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models, Kevin Qu et al. (2026). It advances 3D spatial reasoning and localization in vision-language models by predicting mental spatial maps directly from camera feeds.
