SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation
Ted de Vries LentschZimin XiaHolger CaesarJulian F. P. Kooij
Proposes a geometry-guided cross-view camera pose estimation method that splits the field of view into directional slices to aggregate aerial features via precomputed masks, cutting median localization error on the VIGOR benchmark by up to 50% while running at 150 frames per second.
Reliable vehicle and robot localization is critical for autonomous navigation, yet standard satellite navigation often fails in dense urban environments due to signal blockage. While building and maintaining detailed three-dimensional point clouds or high-definition semantic maps across vast areas is expensive, continuously updated overhead aerial imagery offers a scalable alternative. The challenge lies in accurately determining a ground camera's precise planar position and heading—its three-degrees-of-freedom pose—by matching a ground-level image to an overhead aerial view.
The article introduces and evaluates SliceMatch, a novel computer vision method designed to achieve both highly accurate and real-time cross-view camera pose estimation without requiring prior orientation knowledge or iterative optimization. The method divides the ground camera's horizontal field of view into vertical segments or slices and leverages known geometric projection rules to aggregate corresponding aerial features. By pairing cross-view attention with contrastive learning across thousands of candidate poses, the system evaluates all candidate locations and headings in parallel through precomputed masks.
The authors conducted comprehensive experiments on two standard benchmarks, the multi-city VIGOR panorama dataset and the vehicle-based KITTI dataset, using corrected ground-truth labels. The key findings demonstrate substantial improvements over previous approaches: SliceMatch achieves a 19% reduction in median localization error on VIGOR using a standard baseline network architecture, and a 50% error reduction when using a modern ResNet-50 network. In vehicle tests on KITTI without an initial heading prior, SliceMatch successfully localized targets where previous iterative methods failed entirely due to local optima. Furthermore, SliceMatch executes at over 150 frames per second, maintaining consistent inference speeds even when scaling up to one million candidate poses.
These results show that enforcing geometric structure and directional slicing within image descriptors bridges the domain gap between ground and aerial views effectively, overcoming the speed and convergence trade-offs of prior techniques. Operating at speeds well above typical real-time sensor requirements, this approach lowers the computational and maintenance costs associated with autonomous navigation mapping. Looking forward, organizations developing autonomous navigation stacks should consider integrating SliceMatch with downstream temporal filters or sensor fusion frameworks to resolve occasional multimodal orientation ambiguities caused by symmetrical urban layouts.
- Paper: NetVLAD: CNN Architecture for Weakly Supervised Place Recognition, Relja Arandjelović et al. (2015). It introduces deep descriptor aggregation and metric learning foundations for visual geo-localization that SliceMatch builds upon for cross-view pose matching.
- Paper: LoFTR: Detector-Free Local Feature Matching with Transformers, Jiaming Sun et al. (2021). It details attention-driven transformer mechanisms for dense cross-view feature correlation and matching without explicit detector keypoints.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). It provides the foundational geometry for projecting 2D camera view frustums onto top-down spatial coordinate frames.
- Paper: SuperGlue: Learning Feature Matching With Graph Neural Networks, Paul-Edouard Sarlin et al. (2020). It establishes graph-based attention and geometric contextual reasoning for matching visual features across disparate viewpoints.
- Paper: Deep Visual Geo-localization Benchmark, Gabriele Moreno Berton et al. (2022). It benchmarks visual geo-localization architectures, highlighting the trade-offs in backbone choices and feature aggregators essential for cross-view retrieval.
- Paper: Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization, Siyan Dong et al. (2025). It extends fast camera pose estimation by employing large-scale regression and motion averaging across multiple reference views without needing scene-specific fine-tuning.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It generalizes multi-view camera pose estimation into a unified feed-forward transformer framework predicting 3D scene geometry and camera extrinsics directly.
- Paper: Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence, Junyi Zhang et al. (2024). It explores geometry-aware semantic correspondences across significant viewpoint and orientation shifts to overcome orientation ambiguities in visual feature matching.
