SG2Loc: Sequential Visual Localization on 3D Scene Graphs

Nicole DamblonOlga VysotskaFederico TombariMarc PollefeysDániel Baráth

article2026arXiv0 citations

Introduces SG2Loc, a lightweight sequential visual localization approach that tracks camera poses using compact 3D scene graphs and particle filtering, drastically reducing map storage while eliminating the need for dense point clouds or image databases.

Listen

Visual localization is a foundational capability for autonomous robotics, navigation, and augmented reality systems. Traditional visual positioning pipelines typically rely on massive image databases or dense 3D point clouds, creating substantial storage, memory, and bandwidth bottlenecks on edge devices such as mobile robots and wearable headsets.

The article demonstrates and evaluates SG2Loc, a lightweight sequential visual localization framework designed to determine camera positions across image sequences using compact 3D scene graphs rather than heavy visual maps.

The framework represents indoor environments as scene graphs consisting of coarse object meshes and semantic embeddings. A multi-round particle filter tracks camera position and orientation over time by matching semantic patch features from query images against object projections derived from the coarse meshes. To refine accuracy, the framework incorporates complementary depth and photometric cues alongside a final standard geometric refinement step. The authors evaluated the approach on benchmark indoor datasets—3RScan across 30 test rooms and ScanNet across 48 scene pairs—under challenging conditions involving temporal layout changes and sequence lengths ranging from 5 to 25 frames.

The evaluation produced four key findings. First, SG2Loc matched or exceeded the accuracy of state-of-the-art baselines. On 3RScan across all sequence lengths, it achieved the highest joint position and orientation recall (e.g., 0.44 joint recall and 0.07-meter median error at 25 frames), while remaining competitive on ScanNet. Second, the framework drastically reduced map storage requirements, needing only 9.8 MB per scene on 3RScan compared to 57.6 MB for HLoc and 701.3 MB for MeshLoc, and 28.2 MB on ScanNet compared to over 2,290 MB and 10,590 MB for HLoc and MeshLoc, respectively. Third, extending scene graph matching sequentially for initial scene retrieval improved top-1 recall from 0.69 to 0.81 in 50-candidate room datasets. Fourth, all multi-modal supervision signals contributed measurably to performance, with semantic cues providing initial structural stability and depth and color constraints sharpening precision.

These findings indicate that autonomous agents do not need dense, multi-gigabyte visual databases for reliable global indoor relocalization. Transitioning to lightweight scene graphs lowers storage requirements by up to two orders of magnitude, making large-scale map deployment feasible across bandwidth- and memory-constrained hardware.

For practical deployment, organizations should adopt keyframing to integrate frames at realistic intervals, as the evaluation demonstrates that keyframed updates outperform standard single-frame baselines under online timing constraints. Prior to large-scale operational rollout, engineering teams should optimize the parallel raycasting code to reduce per-frame inference times (currently averaging 2.9 seconds) and run pilot evaluations in dynamic environments where viewpoint diversity may vary.

The primary limitation of the method is its computational latency during frame updates and a vulnerability to localization ambiguities when an input sequence contains visually uniform views with minimal perspective shift. Confidence in the reported storage savings and indoor localization precision remains high based on comprehensive cross-dataset testing.

No sufficiently relevant recommendations were found.

Cover for SG2Loc: Sequential Visual Localization on 3D Scene Graphs

Abstract

Visual localization in complex indoor environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose estimates are refined over time, is important for autonomous agents. However, traditional methods often require storing extensive image databases or point clouds, leading to significant overhead. This paper introduces a novel, lightweight approach to sequential visual localization using 3D scene graphs. Our method represents the environment with a compact scene graph, where nodes represent objects (with coarse meshes) and edges encode spatial relationships. For each image in the localization phase, we extract per-patch semantic features, predicting object identities. Localization is performed within a particle filter framework. Each particle, representing a camera pose, projects the coarse object meshes from the scene graph into the image, assigning object identities to patches based on visibility. The similarity of the per-patch features, in the input image, and object features from the scene graph determines the weight of a particle. Subsequent images are incorporated sequentially, refining the pose estimate. By leveraging a compact scene graph and efficient semantic matching, our method significantly reduces storage while maintaining performance on real-world datasets. The code will be available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Sequential Localization with Particle Filter
  • 3.1 Initialization
  • 3.2 Prediction (Motion Model)
  • 3.3 Update (Observation Model)
  • 3.4 Additional Supervision Signals
  • 3.5 Coarse-to-Fine Optimization
  • 3.6 Pose Refinement with PnP
  • 4 Sequential Scene Retrieval
  • 5 Experiments
  • 6 Conclusion
  • References
  • A Additional Ablations on Supervision Signals
  • B Qualitative Results
  • C Experiment using DROID-SLAM poses
  • D Large-scale environments

Citation

MLA
Damblon, N., et al. “SG2Loc: Sequential Visual Localization on 3D Scene Graphs”. arXiv, 2026, http://arxiv.org/abs/2606.11880v1.
APA
Damblon, N., Vysotska, O., Tombari, F., Pollefeys, M., & Barath, D. (2026). SG2Loc: Sequential Visual Localization on 3D Scene Graphs. arXiv. http://arxiv.org/abs/2606.11880v1
Chicago
Damblon, N., O. Vysotska, F. Tombari, M. Pollefeys, and D. Barath. 2026. “SG2Loc: Sequential Visual Localization on 3D Scene Graphs”. arXiv. http://arxiv.org/abs/2606.11880v1.
Harvard
Damblon, N. et al. (2026) “SG2Loc: Sequential Visual Localization on 3D Scene Graphs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.11880v1.
Vancouver
1. Damblon N, Vysotska O, Tombari F, Pollefeys M, Barath D (2026) SG2Loc: Sequential Visual Localization on 3D Scene Graphs. arXiv

BibTeX

@article{damblon2026sg2loc,
  title = {SG2Loc: Sequential Visual Localization on 3D Scene Graphs},
  author = {Damblon, Nicole and Vysotska, Olga and Tombari, Federico and Pollefeys, Marc and Barath, Daniel},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.11880v1},
  eprint = {2606.11880}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission