SG2Loc: Sequential Visual Localization on 3D Scene Graphs
Nicole DamblonOlga VysotskaFederico TombariMarc PollefeysDániel Baráth
Introduces SG2Loc, a lightweight sequential visual localization approach that tracks camera poses using compact 3D scene graphs and particle filtering, drastically reducing map storage while eliminating the need for dense point clouds or image databases.
Visual localization is a foundational capability for autonomous robotics, navigation, and augmented reality systems. Traditional visual positioning pipelines typically rely on massive image databases or dense 3D point clouds, creating substantial storage, memory, and bandwidth bottlenecks on edge devices such as mobile robots and wearable headsets.
The article demonstrates and evaluates SG2Loc, a lightweight sequential visual localization framework designed to determine camera positions across image sequences using compact 3D scene graphs rather than heavy visual maps.
The framework represents indoor environments as scene graphs consisting of coarse object meshes and semantic embeddings. A multi-round particle filter tracks camera position and orientation over time by matching semantic patch features from query images against object projections derived from the coarse meshes. To refine accuracy, the framework incorporates complementary depth and photometric cues alongside a final standard geometric refinement step. The authors evaluated the approach on benchmark indoor datasets—3RScan across 30 test rooms and ScanNet across 48 scene pairs—under challenging conditions involving temporal layout changes and sequence lengths ranging from 5 to 25 frames.
The evaluation produced four key findings. First, SG2Loc matched or exceeded the accuracy of state-of-the-art baselines. On 3RScan across all sequence lengths, it achieved the highest joint position and orientation recall (e.g., 0.44 joint recall and 0.07-meter median error at 25 frames), while remaining competitive on ScanNet. Second, the framework drastically reduced map storage requirements, needing only 9.8 MB per scene on 3RScan compared to 57.6 MB for HLoc and 701.3 MB for MeshLoc, and 28.2 MB on ScanNet compared to over 2,290 MB and 10,590 MB for HLoc and MeshLoc, respectively. Third, extending scene graph matching sequentially for initial scene retrieval improved top-1 recall from 0.69 to 0.81 in 50-candidate room datasets. Fourth, all multi-modal supervision signals contributed measurably to performance, with semantic cues providing initial structural stability and depth and color constraints sharpening precision.
These findings indicate that autonomous agents do not need dense, multi-gigabyte visual databases for reliable global indoor relocalization. Transitioning to lightweight scene graphs lowers storage requirements by up to two orders of magnitude, making large-scale map deployment feasible across bandwidth- and memory-constrained hardware.
For practical deployment, organizations should adopt keyframing to integrate frames at realistic intervals, as the evaluation demonstrates that keyframed updates outperform standard single-frame baselines under online timing constraints. Prior to large-scale operational rollout, engineering teams should optimize the parallel raycasting code to reduce per-frame inference times (currently averaging 2.9 seconds) and run pilot evaluations in dynamic environments where viewpoint diversity may vary.
The primary limitation of the method is its computational latency during frame updates and a vulnerability to localization ambiguities when an input sequence contains visually uniform views with minimal perspective shift. Confidence in the reported storage savings and indoor localization precision remains high based on comprehensive cross-dataset testing.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). Provides foundational concepts for visual scene graph generation that underpins the structured object and relationship representations utilized in SG2Loc.
- Paper: ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras, Raul Mur-Artal et al. (2016). Establishes standard visual SLAM and camera localization techniques that sequential visual localization methods build upon and improve.
- Paper: SUN RGB-D: A RGB-D scene understanding benchmark suite, Shuran Song et al. (2015). Supplies foundational indoor RGB-D 3D bounding box and semantic scene data essential for learning 3D spatial and object representations.
- Paper: PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization, Alex Kendall et al. (2015). Introduces end-to-end 6-DoF camera pose relocalization frameworks that motivate more compact, structured localization paradigms.
No sufficiently relevant recommendations were found.
