SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

Ted de Vries LentschZimin XiaHolger CaesarJulian F. P. Kooij

article2023CVPR75 citations

Proposes a geometry-guided cross-view camera pose estimation method that splits the field of view into directional slices to aggregate aerial features via precomputed masks, cutting median localization error on the VIGOR benchmark by up to 50% while running at 150 frames per second.

Listen

Reliable vehicle and robot localization is critical for autonomous navigation, yet standard satellite navigation often fails in dense urban environments due to signal blockage. While building and maintaining detailed three-dimensional point clouds or high-definition semantic maps across vast areas is expensive, continuously updated overhead aerial imagery offers a scalable alternative. The challenge lies in accurately determining a ground camera's precise planar position and heading—its three-degrees-of-freedom pose—by matching a ground-level image to an overhead aerial view.

The article introduces and evaluates SliceMatch, a novel computer vision method designed to achieve both highly accurate and real-time cross-view camera pose estimation without requiring prior orientation knowledge or iterative optimization. The method divides the ground camera's horizontal field of view into vertical segments or slices and leverages known geometric projection rules to aggregate corresponding aerial features. By pairing cross-view attention with contrastive learning across thousands of candidate poses, the system evaluates all candidate locations and headings in parallel through precomputed masks.

The authors conducted comprehensive experiments on two standard benchmarks, the multi-city VIGOR panorama dataset and the vehicle-based KITTI dataset, using corrected ground-truth labels. The key findings demonstrate substantial improvements over previous approaches: SliceMatch achieves a 19% reduction in median localization error on VIGOR using a standard baseline network architecture, and a 50% error reduction when using a modern ResNet-50 network. In vehicle tests on KITTI without an initial heading prior, SliceMatch successfully localized targets where previous iterative methods failed entirely due to local optima. Furthermore, SliceMatch executes at over 150 frames per second, maintaining consistent inference speeds even when scaling up to one million candidate poses.

These results show that enforcing geometric structure and directional slicing within image descriptors bridges the domain gap between ground and aerial views effectively, overcoming the speed and convergence trade-offs of prior techniques. Operating at speeds well above typical real-time sensor requirements, this approach lowers the computational and maintenance costs associated with autonomous navigation mapping. Looking forward, organizations developing autonomous navigation stacks should consider integrating SliceMatch with downstream temporal filters or sensor fusion frameworks to resolve occasional multimodal orientation ambiguities caused by symmetrical urban layouts.

Cover for SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

Abstract

This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists of ground and aerial feature extractors, feature aggregators, and a pose predictor. The feature extractors extract dense features from the ground and aerial images. Given a set of candidate camera poses, the feature aggregators construct a single ground descriptor and a set of pose-dependent aerial descriptors. Notably, our novel aerial feature aggregator has a cross-view attention module for ground-view guided aerial feature selection and utilizes the geometric projection of the ground camera’s viewing frustum on the aerial image to pool features. The efficient construction of aerial descriptors is achieved using precomputed masks. SliceMatch is trained using contrastive learning and pose estimation is formulated as a similarity comparison between the ground descriptor and the aerial descriptors. Compared to the state-of-the-art, SliceMatch achieves a 19% lower median localization error on the VIGOR benchmark using the same VGG16 backbone at 150 frames per second, and a 50% lower error when using a ResNet50 backbone.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Cross-View Camera Pose Estimation
  • 3.2. SliceMatch Overview
  • 3.3. Geometry-Guided Cross-View Aggregation
  • 3.3.1 Ground Feature Aggregator
  • 3.3.2 Aerial Feature Aggregator
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Evaluation Metrics
  • 4.3. Implementation Details
  • 4.4. Baselines
  • 4.5. Ablation Study
  • 4.6. Same-Area Generalization
  • 4.7. Cross-Area Generalization
  • 4.8. Runtime Analysis
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — SliceMatch pose-matching formulation

    model/method

    SliceMatch estimates the 3-DoF pose of a ground-level image relative to a square overhead aerial image. The pose is ξ=(u,v,θ)\xi=(u,v,\theta), where (u,v)∈[0,1]2(u,v)\in[0,1]^2 is the normalized camera location in the aerial image and θ∈[0,360∘)\theta\in[0,360^\circ) is the clockwise angle from North to the camera’s forward direction. The method assumes small camera pitch and roll.

    Given a ground image IgI_g, an aerial image IaI_a, and a candidate set Ξ={ξ1,…,ξK}\Xi=\{\xi^1,\ldots,\xi^K\}, separate convolutional feature extractors produce

    zg=fg(Ig)∈RH×W×C,za=fa(Ia)∈RL×L×C,z_g=f_g(I_g)\in\mathbb{R}^{H\times W\times C},\qquad z_a=f_a(I_a)\in\mathbb{R}^{L\times L\times C},

    where HH and WW are the ground feature-map dimensions, L×LL\times L is the aerial feature-map size, and CC is the feature-channel dimension. The ground aggregator divides the ground view into NN azimuth slices and concatenates their normalized descriptors into dg∈RDd_g\in\mathbb{R}^{D}, where D=NCD=NC. The aerial aggregator constructs one pose-dependent descriptor dak∈RDd_a^k\in\mathbb{R}^{D} for each candidate pose ξk\xi^k.

    The predicted pose is the candidate with maximum cosine similarity to the ground descriptor:

    k∗=argmax⁡k∈{1,…,K}  cos⁡(dg,dak),ξ^=ξk∗.k^*=\underset{k\in\{1,\ldots,K\}}{\operatorname{argmax}}\;\operatorname{cos}(d_g,d_a^k),\qquad \hat\xi=\xi^{k^*}.

    Dense ground and aerial features are extracted only once; testing additional candidate poses primarily requires applying their precomputed geometric masks, aggregating features, and evaluating cosine similarities.

  2. Knowl 2 — Ground-view slice aggregation

    model/method

    The SliceMatch ground aggregator first suppresses ground-image features that are unlikely to correspond to the aerial image, such as sky or transient objects. It applies a learned spatial attention mask to the ground feature map:

    zg′=Mg⊙zg,Mg=Sigmoid⁡(Conv⁡1×1(zg)),z'_g=M_g\odot z_g,\qquad M_g=\operatorname{Sigmoid}(\operatorname{Conv}_{1\times1}(z_g)),

    where zg∈RH×W×Cz_g\in\mathbb{R}^{H\times W\times C} is the ground feature map, Mg∈[0,1]H×W×1M_g\in[0,1]^{H\times W\times1} is broadcast over the CC channels, ⊙\odot denotes element-wise multiplication, and Conv⁡1×1\operatorname{Conv}_{1\times1} is a learned pointwise convolutional mapping.

    The reweighted feature map zg′z'_g is partitioned horizontally into NN non-overlapping vertical regions, each representing an azimuth interval of the ground camera’s horizontal field of view. The features within each region are averaged and L2-normalized to produce a CC-dimensional slice descriptor d^gn\hat d_g^n for slice n∈{1,…,N}n\in\{1,\ldots,N\}. The ground global descriptor is the concatenation

    dg=Concat⁡(d^g1,…,d^gN)∈RNC.d_g=\operatorname{Concat}(\hat d_g^1,\ldots,\hat d_g^N)\in\mathbb{R}^{NC}.

    Thus, each component of the ground descriptor retains the visual information associated with a particular viewing direction rather than collapsing the entire ground image into one spatially agnostic vector.

  3. Knowl 3 — Ground-guided cross-view attention

    model/method

    For every ground slice descriptor, SliceMatch learns which aerial features are relevant to the content visible in that ground-view direction. Let zai,j∈RCz_a^{i,j}\in\mathbb{R}^{C} be the aerial feature at spatial location (i,j)(i,j), with 1≤i,j≤L1\leq i,j\leq L, and let d^gn∈RC\hat d_g^n\in\mathbb{R}^{C} be the descriptor of ground slice nn. A similarity map is computed as

    Si,jn=Sim⁡(d^gn,zai,j),S^n_{i,j}=\operatorname{Sim}(\hat d_g^n,z_a^{i,j}),

    where Sn∈RL×LS^n\in\mathbb{R}^{L\times L} and Sim⁡\operatorname{Sim} is the feature-similarity function used by the attention module.

    The similarity map is concatenated with the aerial feature map and passed through a learned pointwise network to obtain a slice-specific aerial attention mask:

    Man=Sigmoid⁡ ⁣(Conv⁡1×1(Concat⁡(za,Sn))),za′n=Man⊙za.M_a^n=\operatorname{Sigmoid}\!\left(\operatorname{Conv}_{1\times1}\big(\operatorname{Concat}(z_a,S^n)\big)\right), \qquad z_a^{\prime n}=M_a^n\odot z_a.

    Here Man∈[0,1]L×L×1M_a^n\in[0,1]^{L\times L\times1} is the aerial mask for ground slice nn, and za′n∈RL×L×Cz_a^{\prime n}\in\mathbb{R}^{L\times L\times C} is the reweighted aerial feature map. Separate masks are produced for all NN ground slices, so each viewing direction can select aerial content conditioned on the corresponding ground appearance.

  4. Knowl 4 — Geometry-guided aerial slice descriptors

    model/method

    SliceMatch represents the camera viewing frustum geometrically in the aerial feature map. For each candidate pose ξk\xi^k and azimuth slice nn, it precomputes a soft mask Pk,n∈[0,1]L×LP^{k,n}\in[0,1]^{L\times L}. The value Pi,jk,nP^{k,n}_{i,j} is proportional to the fraction of aerial feature-map cell (i,j)(i,j) intersected by the projection of that pose’s frustum slice: a value of 11 denotes full inclusion, 00 denotes no intersection, and intermediate values denote partial overlap.

    Using the cross-view-reweighted aerial feature map za′nz_a^{\prime n}, SliceMatch computes the descriptor for pose kk and slice nn by weighted spatial averaging followed by L2 normalization:

    d^ak,n=Norm⁡ ⁣(∑i,jPi,jk,n za,i,j′n∑i,jPi,jk,n),\hat d_a^{k,n}=\operatorname{Norm}\!\left(\frac{\sum_{i,j}P^{k,n}_{i,j}\,z_{a,i,j}^{\prime n}}{\sum_{i,j}P^{k,n}_{i,j}}\right),

    where d^ak,n∈RC\hat d_a^{k,n}\in\mathbb{R}^{C}, za,i,j′n∈RCz_{a,i,j}^{\prime n}\in\mathbb{R}^{C} is the feature vector at aerial location (i,j)(i,j), and the denominator is assumed to be nonzero. The complete descriptor for candidate pose ξk\xi^k is

    dak=Concat⁡(d^ak,1,…,d^ak,N)∈RNC.d_a^k=\operatorname{Concat}(\hat d_a^{k,1},\ldots,\hat d_a^{k,N})\in\mathbb{R}^{NC}.

    Because the masks encode all pose-dependent geometry before inference, changing the number of candidate poses does not require recomputing the image features or the cross-view attention maps. Additional candidates require only weighted averaging, normalization, and similarity computations that can be parallelized.

  5. Knowl 5 — Pose-discriminative contrastive loss

    equation

    SliceMatch trains the ground descriptor against aerial descriptors from different locations and orientations using a modified InfoNCE loss. Let dgd_g be the ground descriptor, let daGTd_a^{GT} be the aerial descriptor constructed at the ground-truth pose ξGT\xi^{GT}, and let dakd_a^k be the descriptor for the kk-th non-ground-truth training candidate pose. Define cGT=cos⁡(dg,daGT)c^{GT}=\operatorname{cos}(d_g,d_a^{GT}) and ck=cos⁡(dg,dak)c^k=\operatorname{cos}(d_g,d_a^k) as cosine similarities, let KK be the number of negative candidate poses used in one training example, let τ>0\tau>0 be the temperature, and let α>0\alpha>0 weight the aggregate contribution of the negative candidates. The loss is

    L=−log⁡(exp⁡(cGT/τ)αK∑k=1Kexp⁡(ck/τ)+exp⁡(cGT/τ)).\mathcal{L}=-\log\left(\frac{\exp(c^{GT}/\tau)}{\frac{\alpha}{K}\sum_{k=1}^{K}\exp(c^k/\tau)+\exp(c^{GT}/\tau)}\right).

    The original InfoNCE weighting is recovered when α=K\alpha=K. SliceMatch instead uses α=4\alpha=4, which emphasizes discrimination against the set of location-and-orientation alternatives without making the negative contribution grow directly with the number of candidates. The contrastive formulation trains the descriptors to distinguish both camera position and camera orientation.

  6. Knowl 6 — Datasets, implementation, and candidate-pose evaluation

    experimental setup

    SliceMatch is evaluated on VIGOR and KITTI. VIGOR contains geo-tagged ground panoramas and aerial images from four United States cities, with same-area and cross-area generalization splits. Each ground panorama has one positive aerial image and three semi-positive aerial images; the experiments use positive images and corrected ground-truth locations because the original aerial-image resolution caused location-label errors of up to approximately 33 m. KITTI contains limited-horizontal-field-of-view images captured from moving vehicles, augmented with aerial images; its training and Test1 sets cover the same region, while Test2 covers a different region.

    The main implementation uses separate ImageNet-pretrained VGG16 feature extractors truncated at stage 5, with the final pooling operation removed. The ground and aerial extractors do not share weights. The input sizes are 320×640320\times640 for VIGOR ground images, 256×1024256\times1024 for KITTI ground images, and 512×512512\times512 for aerial images. The resulting feature maps have sizes 20×4020\times40 for VIGOR ground images, 16×6416\times64 for KITTI ground images, and 32×3232\times32 for aerial images, with C=512C=512 channels. The 1×11\times1 convolutional modules contain two sequential pointwise convolutions with an intermediate ReLU. Training uses Adam with learning rate 10−510^{-5} and batch size 44.

    Training candidates form uniform pose grids of 7×77\times7 locations and 1616 orientations on VIGOR, giving Ktrain=784K_{\mathrm{train}}=784, and 5×55\times5 locations and 1616 orientations on KITTI, giving Ktrain=400K_{\mathrm{train}}=400. Inference uses 21×21×64=28,22421\times21\times64=28{,}224 candidates on VIGOR and 15×15×64=14,40015\times15\times64=14{,}400 candidates on KITTI. Localization is measured by mean and median position error in meters; orientation is measured by mean and median absolute angular error in degrees. KITTI additionally reports recall within 11 m and 55 m for lateral and longitudinal errors and within 1∘1^\circ and 5∘5^\circ for orientation.

  7. Knowl 7 — Slice and attention ablation

    data/table

    An ablation on the VIGOR tuning split measures the effect of the number of azimuth slices NN, cross-view attention, and the modified loss weighting. Location errors are in meters and orientation errors are in degrees; lower values are better. The row marked original loss uses the unmodified InfoNCE weighting α=K\alpha=K, whereas the other rows use the selected weighting α=4\alpha=4.

    N Cross-view attention Location Orientation
    Mean Median Mean Median
    1 No 12.73 11.51 - -
    4 Yes 9.47 7.47 51.49 32.96
    8 Yes 9.16 6.81 37.68 15.58
    16 Yes 7.60 5.23 29.27 9.22
    32 Yes 8.14 5.31 32.01 10.31
    16 original loss) Yes 8.08 5.44 31.05 11.02
    16 No 7.93 5.81 29.50 12.32

    Using one slice prevents orientation estimation, while increasing the number of slices to 1616 substantially improves both localization and orientation. Performance slightly degrades at 3232 slices, indicating saturation beyond 1616. At N=16N=16, adding cross-view attention improves median localization from 5.815.81 m to 5.235.23 m and median orientation error from 12.32∘12.32^\circ to 9.22∘9.22^\circ. The modified loss improves mean localization error from 8.088.08 m to 7.607.60 m relative to the original InfoNCE weighting. SliceMatch therefore uses N=16N=16, cross-view attention, and α=4\alpha=4.

  8. Knowl 8 — VIGOR localization and pose-estimation results

    data/table

    The following results compare SliceMatch with the global-descriptor baselines Cross-View Regression (CVR) and Multi-Class Classification (MCC) on VIGOR. An aligned ground image has known orientation and requires localization only; an unaligned image requires full 3-DoF pose estimation. The reported metrics are mean and median location error in meters and mean and median orientation error in degrees.

    Model Backbone Aligned Same-area location Same-area orientation Cross-area location Cross-area orientation
    images Mean Median Mean Median Mean Median Mean Median
    CVR VGG16 Yes 8.99 7.81 - - 8.89 7.73 - -
    MCC VGG16 Yes 6.94 3.64 - - 9.05 5.14 - -
    SliceMatch VGG16 Yes 5.18 2.58 - - 5.53 2.55 - -
    MCC VGG16 No 9.87 6.25 56.86 16.02 12.66 9.55 72.13 29.97
    SliceMatch VGG16 No 8.41 5.07 28.43 5.15 8.48 5.64 26.20 5.18
    SliceMatch ResNet50 No 6.49 3.13 25.46 4.71 7.22 3.31 25.97 4.51

    With the same VGG16 backbone and unknown orientation, SliceMatch reduces same-area median localization error from 6.256.25 m for MCC to 5.075.07 m, a 19%19\% reduction, and reduces median orientation error from 16.02∘16.02^\circ to 5.15∘5.15^\circ, a 68%68\% reduction. The gains persist across areas, where SliceMatch obtains a 5.645.64 m median location error and 5.18∘5.18^\circ median orientation error. Replacing VGG16 with ResNet50 further reduces same-area median location error to 3.133.13 m, approximately 50%50\% below MCC’s VGG16 result. The score map can remain multimodal in symmetric scenes, but the highest-scoring mode can occasionally be incorrect, producing large outlier errors and causing mean errors to exceed median errors substantially.

  9. Knowl 9 — KITTI localization and orientation results

    data/table

    SliceMatch is compared with DSM, a fine-grained retrieval method, and LM, an iterative pose-estimation method, on KITTI. Same-area and cross-area denote the evaluation regions. A 20∘20^\circ prior means that a noisy orientation prior is available during training and testing; X means no orientation prior. Location errors are in meters, orientation errors are in degrees, and recalls are percentages. The notation r@1mr@1\text{m}, for example, denotes the percentage of samples whose error is at most 11 m.

    Model Area Prior Location Lateral recall Longitudinal recall Orientation
    Mean Median r@1m r@5m r@1m r@5m Mean Median
    DSM Same 20^ - - 10.12 48.24 4.08 20.14 - -
    LM Same 20^ 12.08 11.42 35.54 80.36 5.22 26.13 3.72 2.83
    SliceMatch Same 20^ 7.96 4.39 49.09 98.52 15.19 57.35 4.12 3.65
    LM Same X 15.51 15.97 5.17 25.44 4.66 25.39 89.91 90.75
    SliceMatch Same X 9.39 5.41 39.73 87.92 13.63 49.22 8.71 4.42
    DSM Cross 20^ - - 10.77 48.24 3.87 19.50 - -
    LM Cross 20^ 12.58 12.11 27.82 72.89 5.75 26.48 3.95 3.03
    SliceMatch Cross 20^ 13.50 9.77 32.43 86.44 8.30 35.57 4.20 6.61
    LM Cross X 15.50 16.02 5.60 25.60 5.64 25.76 89.84 89.85
    SliceMatch Cross X 14.85 11.85 24.00 72.89 7.17 33.12 23.64 7.96

    On the same-area split with a 20∘20^\circ prior, SliceMatch improves over LM from 12.0812.08 m to 7.967.96 m mean location error and from 11.4211.42 m to 4.394.39 m median error, while also obtaining higher lateral and longitudinal localization recalls. Without a prior, SliceMatch remains usable whereas LM’s orientation error becomes approximately 90∘90^\circ. On the cross-area split with a prior, SliceMatch has lower median location error than LM but higher mean location and orientation errors. Without a prior, SliceMatch again has substantially better localization and orientation performance than LM, whose iterative optimization becomes trapped in poor local optima.

  10. Knowl 10 — Inference speed and scalability with candidate poses

    empirical result

    On a single NVIDIA Tesla V100 GPU, SliceMatch processes VIGOR ground–aerial image pairs at 167167 frames per second and KITTI pairs at 156156 frames per second. On VIGOR, the corresponding rates are 5050 frames per second for CVR, 2929 frames per second for MCC when performing localization only, and 33 frames per second for MCC when estimating the full pose. On KITTI, the iterative LM baseline runs at only 0.590.59 frames per second.

    SliceMatch’s inference time remains nearly constant as the number of candidate poses increases because feature extraction and cross-view attention are independent of the candidate count. The authors observed this behavior even when testing up to K=1×106K=1\times10^6 candidate poses; the additional work consists mainly of parallel weighted averaging with precomputed masks, descriptor normalization, and cosine-similarity evaluation.

Coverage note — No substantial contributed material was omitted; background, related work, acknowledgements, and future-work suggestions were excluded, while the reported ambiguity and failure behavior is included with the experimental results.

References

  1. 1.Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. NetVLAD: CNN architecture for weakly supervised place recognition. In CVPR, pages 5297–5307, 2016. 2
  2. 2.Boaz Ben-Moshe, Elazar Elkin, Harel Levi, and Ayal Weissman. Improving accuracy of GNSS devices in urban canyons. In CCCG, pages 511–515, 2011. 1
  3. 3.Hao Cai, Zhaozheng Hu, Gang Huang, Dunyao Zhu, and Xiaocong Su. Integration of GPS, monocular vision, and high definition (HD) map for accurate vehicle localization. Sensors, 18(10):3270, 2018. 1
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. 6
  5. 5.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 3
  6. 6.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The KITTI dataset. IJRR, 32(11):1231–1237, 2013. 6, 8
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 3
  8. 8.Henry Howard-Jenkins and Victor Adrian Prisacariu. LaLaLoc++: Global floor plan comprehension for layout localisation in unvisited environments. In ECCV, pages 693–709, 2022. 3
  9. 9.Henry Howard-Jenkins, Jose-Raul Ruiz-Sarmiento, and Victor Adrian Prisacariu. LaLaLoc: Latent layout localisation in dynamic, unvisited environments. In CVPR, pages 10107–10116, 2021. 3
  10. 10.Sixing Hu, Mengdan Feng, Rang MH Nguyen, and Gim Hee Lee. CVM-Net: Cross-view matching network for image-based ground-to-aerial geo-localization. In CVPR, pages 7258–7267, 2018. 2
  11. 11.Wenmiao Hu, Yichen Zhang, Yuxuan Liang, Yifang Yin, Andrei Georgescu, An Tran, Hannes Kruppa, See-Kiong Ng, and Roger Zimmermann. Beyond geo-localization: Fine-grained orientation of street-view images by cross-view matching with satellite imagery. In ACM Multimedia, pages 6155–6164, 2022. 3
  12. 12.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In NeurIPS, pages 18661–18673, 2020. 4
  13. 13.Thomas King, Holger Fußler, Matthias Transier, and Wolfgang Effelsberg. Dead-reckoning for position-based forwarding on highways. In WIT, pages 199–204, 2006. 1
  14. 14.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 6
  15. 15.Kenneth Levenberg. A method for the solution of certain non-linear problems in least squares. Quarterly of applied mathematics, 2(2):164–168, 1944. 3
  16. 16.Songlian Li, Zhigang Tu, Yujin Chen, and Tan Yu. Multi-scale attention encoder for street-to-aerial image geo-localization. CAAI TIT, 2022. 2
  17. 17.Tsung-Yi Lin, Yin Cui, Serge Belongie, and James Hays. Learning deep representations for ground-to-aerial geolocalization. In CVPR, pages 5007–5015, 2015. 2
  18. 18.Liu Liu and Hongdong Li. Lending orientation to neural networks for cross-view geo-localization. In CVPR, pages 5624–5633, 2019. 2
  19. 19.Stephanie Lowry, Niko Sunderhauf, Paul Newman, John J Leonard, David Cox, Peter Corke, and Michael J Milford. Visual place recognition: A survey. T-RO, 32(1):1–19, 2015. 1
  20. 20.Xiaohu Lu, Zuoyue Li, Zhaopeng Cui, Martin R Oswald, Marc Pollefeys, and Rongjun Qin. Geometry-aware satellite-to-ground image synthesis for urban areas. In CVPR, pages 859–867, 2020. 3
  21. 21.Xiufan Lu, Siqi Luo, and Yingying Zhu. It’s okay to be wrong: Cross-view geo-localization with step-adaptive iterative refinement. TGRS, 60:1–13, 2022. 2
  22. 22.Donald W Marquardt. An algorithm for least-squares estimation of nonlinear parameters. SIAM, 11(2):431–441, 1963. 3
  23. 23.Zhixiang Min, Naji Khosravan, Zachary Bessinger, Manjunath Narayana, Sing Bing Kang, Enrique Dunn, and Ivaylo Boyadzhiev. LASER: Latent space rendering for 2d visual localization. In CVPR, pages 11122–11131, 2022. 3
  24. 24.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. In arXiv preprint arXiv:1807.03748, 2018. 4, 7
  25. 25.Jonah Philion and Sanja Fidler. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In ECCV, pages 194–210, 2020. 3
  26. 26.Krishna Regmi and Mubarak Shah. Bridging the domain gap for ground-to-aerial image matching. In ICCV, pages 470–479, 2019. 2
  27. 27.Tyler GR Reid, Sarah E Houts, Robert Cammarata, Graham Mills, Siddharth Agarwal, Ankit Vora, and Gaurav Pandey. Localization requirements for autonomous vehicles. SAE IJ-CAV, 2(12-02-03-0012):173–190, 2019. 3
  28. 28.Thomas Roddick and Roberto Cipolla. Predicting semantic map representations from images using pyramid occupancy networks. In CVPR, pages 11138–11147, 2020. 3
  29. 29.Royston Rodrigues and Masahiro Tani. Global assists local: Effective aerial representations for field of view constrained image geo-localization. In WACV, pages 3871–3879, 2022. 2
  30. 30.Avishkar Saha, Oscar Mendez Maldonado, Chris Russell, and Richard Bowden. Translating images into maps. In ICRA, 2021. 3
  31. 31.Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Fast image-based localization using direct 2d-to-3d matching. In ICCV, pages 667–674, 2011. 1
  32. 32.Yujiao Shi, Dylan Campbell, Xin Yu, and Hongdong Li. Geometry-guided street-view panorama synthesis from satellite imagery. T-PAMI, 44(12):10009–10022, 2022. 3
  33. 33.Yujiao Shi and Hongdong Li. Beyond cross-view image retrieval: Highly accurate vehicle localization using satellite image. In CVPR, pages 17010–17020, 2022. 1, 2, 3, 6, 7, 8
  34. 34.Yujiao Shi, Liu Liu, Xin Yu, and Hongdong Li. Spatial-aware feature aggregation for cross-view image based geo-localization. In NeurIPS, pages 10090–10100, 2019. 2, 4
  35. 35.Yujiao Shi, Xin Yu, Dylan Campbell, and Hongdong Li. Where am I looking at? Joint location and orientation estimation by cross-view matching. In CVPR, pages 4064–4072, 2020. 2, 6, 7, 8
  36. 36.Yujiao Shi, Xin Yu, Liu Liu, Dylan Campbell, Piotr Koniusz, and Hongdong Li. Accurate 3-dof camera geo-localization via ground-to-satellite image matching. T-PAMI, 2022. 1, 2, 3
  37. 37.Yujiao Shi, Xin Yu, Liu Liu, Tong Zhang, and Hongdong Li. Optimal feature transport for cross-view image geo-localization. In AAAI, pages 11990–11997, 2020. 2
  38. 38.Yujiao Shi, Xin Yu, Shan Wang, and Hongdong Li. CVLNet: Cross-view semantic correspondence learning for video-based camera localization. In ACCV, 2022. 2
  39. 39.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 3, 6
  40. 40.Aysim Toker, Qunjie Zhou, Maxim Maximov, and Laura Leal-Taixe. Coming down to earth: Satellite-to-street view synthesis for geo-localization. In CVPR, pages 6488–6497, 2021. 2
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 2
  42. 42.Sebastiano Verde, Thiago Resek, Simone Milani, and Anderson Rocha. Ground-to-aerial viewpoint localization via landmark graphs matching. SPL, 27:1490–1494, 2020. 2
  43. 43.Huayou Wang, Changliang Xue, Yanxing Zhou, Feng Wen, and Hongbo Zhang. Visual semantic localization based on hd map for autonomous vehicles in urban scenarios. In ICRA, pages 11255–11261, 2021. 1
  44. 44.Shan Wang, Yanhao Zhang, and Hongdong Li. Satellite image based cross-view localization for autonomous vehicle. In arXiv preprint arXiv:2207.13506, 2022. 1, 2, 3
  45. 45.Tingyu Wang, Zhedong Zheng, Chenggang Yan, Jiyong Zhang, Yaoqi Sun, Bolun Zheng, and Yi Yang. Each part matters: Local patterns facilitate cross-view geo-localization. TCSVT, 32(2):867–879, 2021. 2
  46. 46.Scott Workman and Nathan Jacobs. On the location dependence of convolutional neural network features. In CVPR Workshops, pages 70–78, 2015. 2
  47. 47.Scott Workman, Richard Souvenir, and Nathan Jacobs. Wide-area image geolocalization with aerial reference imagery. In ICCV, pages 3961–3969, 2015. 2
  48. 48.Zimin Xia, Olaf Booij, Marco Manfredi, and Julian FP Kooij. Geographically local representation learning with a spatial prior for visual localization. In ECCV Workshops, pages 557–573, 2020. 1
  49. 49.Zimin Xia, Olaf Booij, Marco Manfredi, and Julian FP Kooij. Cross-view matching for vehicle localization by learning geographically local representations. RA-L, 6(3):5921–5928, 2021. 1
  50. 50.Zimin Xia, Olaf Booij, Marco Manfredi, and Julian FP Kooij. Visual cross-view metric localization with dense uncertainty estimates. In ECCV, pages 90–106, 2022. 1, 2, 3, 4, 5, 6, 7, 8
  51. 51.Hongji Yang, Xiufan Lu, and Yingying Zhu. Cross-view geo-localization with layer-to-layer transformer. In NeurIPS, 2021. 2
  52. 52.Menghua Zhai, Zachary Bessinger, Scott Workman, and Nathan Jacobs. Predicting ground-level scene layout from aerial imagery. In CVPR, pages 867–875, 2017. 2
  53. 53.Sijie Zhu, Mubarak Shah, and Chen Chen. TransGeo: Transformer is all you need for cross-view image geo-localization. In CVPR, pages 1162–1171, 2022. 2
  54. 54.Sijie Zhu, Taojiannan Yang, and Chen Chen. Revisiting street-to-aerial view image geo-localization and orientation estimation. In WACV, pages 756–765, 2021. 2
  55. 55.Sijie Zhu, Taojiannan Yang, and Chen Chen. VIGOR: Cross-view image geo-localization beyond one-to-one retrieval. In CVPR, pages 3640–3649, 2021. 1, 2, 3, 4, 5, 6, 7, 8

Citation

MLA
Lentsch, T., et al. “SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation”. arXiv, 2022, http://arxiv.org/abs/2211.14651v3.
APA
Lentsch, T., Xia, Z., Caesar, H., & Kooij, J. F. P. (2022). SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation. arXiv. http://arxiv.org/abs/2211.14651v3
Chicago
Lentsch, T., Z. Xia, H. Caesar, and J. F. P. Kooij. 2022. “SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation”. arXiv. http://arxiv.org/abs/2211.14651v3.
Harvard
Lentsch, T. et al. (2022) “SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.14651v3.
Vancouver
1. Lentsch T, Xia Z, Caesar H, Kooij JFP (2022) SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation. arXiv

BibTeX

@article{lentsch2022slicematch,
  title = {SliceMatch: Geometry-guided Aggregation for Cross-View Pose Estimation},
  author = {Lentsch, Ted and Xia, Zimin and Caesar, Holger and Kooij, Julian F. P.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.14651v3},
  eprint = {2211.14651}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE