Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

Siyan DongShuzhe WangShaohui LiuLulu CaiQingnan FanJuho KannalaYanchao Yang

article2025CVPR110 citations

Presents Reloc3r, a scalable framework that combines a symmetric relative pose regression network trained on eight million image pairs with minimalist motion averaging, achieving state-of-the-art camera localization accuracy in real time across unseen scenes.

Listen

Visual localization determines the exact camera position and orientation of a captured image relative to an existing map or image database, a capability essential for autonomous robotics, navigation, and augmented reality. Existing methods face persistent trade-offs: traditional geometric techniques achieve high accuracy but require intensive computation that restricts real-time deployment, while modern direct regression methods either fail to generalize to novel environments or require computationally expensive, scene-specific retraining.

The article evaluates whether a streamlined neural architecture, combined with large-scale pre-training across diverse environments, can achieve real-time speed, high localization accuracy, and zero-shot generalization to unseen scenes. It introduces Reloc3r, an end-to-end framework designed to resolve the historical performance trade-offs in visual camera localization.

To establish a scalable and robust system, the researchers paired an image-to-image relative pose regression model with a lightweight, parameter-free motion averaging module. The network employs a fully symmetric Vision Transformer backbone with shared weights to estimate relative rotation and translation directions without enforcing metric scale during network training. The motion averaging module then triangulates the absolute metric coordinates and calculates orientation across multiple retrieved reference images. The entire framework was trained once on a large dataset of approximately eight million image pairs spanning indoor, outdoor, and object-centric scenes, and subsequently evaluated across six public benchmark datasets without any scene-specific fine-tuning.

Experimental results show that Reloc3r sets a new performance standard across multiple benchmarks. In pairwise relative pose estimation, it outperforms prior regression models by large margins, delivering top accuracy across indoor and outdoor benchmarks while operating at 24 to 66 frames per second, which is up to 20 to 50 times faster than competing non-regression methods. In visual localization on unseen outdoor environments, Reloc3r roughly halved the median error of previous relative pose regression methods, achieving an average error of 0.38 meters and 0.52 degrees on Cambridge Landmarks. In indoor localization on the 7 Scenes benchmark, it attained an average median error of 0.04 meters and 1.02 degrees, matching or outperforming specialized models that required days of scene-specific training. Furthermore, on multi-view object datasets, the system achieved best-in-class accuracy (95.8% rotation accuracy within 15 degrees) using only pairwise evaluations.

These findings demonstrate that relative pose regression does not suffer from fundamental accuracy limits when supported by sufficient data diversity and scaled transformer architectures. For commercial and operational applications, this eliminates the need for costly per-scene 3D model construction and offline retraining cycles. The ability to achieve high precision with real-time inference latency of 15 to 42 milliseconds directly reduces compute infrastructure costs and enhances onboard autonomy in dynamic, real-world deployments.

Organizations developing vision-based navigation and augmented reality tools should consider adopting data-driven relative pose estimation as a viable alternative to complex structure-from-motion pipelines. For practical implementation, practitioners must ensure that image retrieval retrieval mechanisms select reference frames with adequate geometric baseline separation. Future engineering efforts should focus on integrating active viewpoint selection to mitigate rare geometric collinearity failures, where linear camera paths prevent absolute metric triangulation.

The results provide a high degree of confidence across standard indoor and outdoor operational domains. A known limitation occurs when the query image and all retrieved reference frames lie in a straight line, which mathematically degrades scale triangulation in the motion averaging step. Nevertheless, the framework demonstrates strong robustness across varied baseline separations and novel environments.

  • Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). Expands on feed-forward transformer-based camera pose and dense geometry estimation by jointly predicting intrinsics, extrinsics, depth, and point tracks across multi-view collections in a single unified architecture.
Cover for Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

Abstract

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either generalize well to new scenes or provide accurate camera pose estimates. To address these issues, we present Reloc3r, a simple yet effective visual localization framework. It consists of an elegantly designed relative pose regression network, and a minimalist motion averaging module for absolute pose estimation. Trained on approximately eight million posed image pairs, Reloc3r achieves surprisingly good performance and generalization ability. We conduct extensive experiments on six public datasets, consistently demonstrating the effectiveness and efficiency of the proposed method. It provides high-quality camera pose estimates in real time and generalizes to novel scenes. Code: https://github.com/ffrivera0/reloc3r.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Relative Camera Pose Regression
  • 3.2. Motion Averaging
  • 4. Experiments
  • 4.1. Relative Camera Pose Estimation
  • 4.2. Visual Localization
  • 4.3. Analyses
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Symmetric Vision Transformer Architecture for Relative Camera Pose Regression

    model/method

    The Reloc3r relative camera pose regression network takes an input pair of images (I1,I2)(I_1, I_2) of dimensions H×W×3H \times W \times 3 and predicts the relative 6-DoF camera poses P^I1,I2\hat{P}_{I_1, I_2} and P^I2,I1\hat{P}_{I_2, I_1} up to an unknown translation scale.

    The network comprises three main stages with shared weights between branches:

    1. ViT Encoder: Each image IiI_i (i∈{1,2}i \in \{1, 2\}) is divided into non-overlapping spatial patches and mapped to a sequence of TT tokens of feature dimension dd. Rotary Position Embeddings (RoPE) encode 2D relative spatial positions. The tokens pass through m=24m = 24 standard Vision Transformer encoder blocks consisting of multi-head self-attention and feed-forward networks, producing encoded token representations F1,F2∈RT×dF_1, F_2 \in \mathbb{R}^{T \times d}: Fi=Encoder(Patchify(Ii)),i∈{1,2}F_i = \text{Encoder}(\text{Patchify}(I_i)), \quad i \in \{1, 2\}

    2. Cross-Attention ViT Decoder: Feature tokens are decoded across n=12n = 12 Transformer decoder blocks. Each decoder block inserts a cross-attention layer between its self-attention and feed-forward layers to facilitate bidirectional feature exchange between the two views: G1=Decoder(F1,F2),G2=Decoder(F2,F1)G_1 = \text{Decoder}(F_1, F_2), \quad G_2 = \text{Decoder}(F_2, F_1)

    3. Pose Regression Head: A regression head containing h=2h = 2 convolutional layers followed by spatial average pooling outputs a 9D continuous rotation representation and a 3D translation direction vector: P^I1,I2=Head(G1),P^I2,I1=Head(G2)\hat{P}_{I_1, I_2} = \text{Head}(G_1), \quad \hat{P}_{I_2, I_1} = \text{Head}(G_2) The predicted 9D rotation representation is converted into an orthogonal rotation matrix R^∈SO(3)\hat{R} \in \mathrm{SO}(3) via Singular Value Decomposition (SVD) orthogonalization, and concatenated with the unit-normalized translation direction vector t^∈R3\hat{t} \in \mathbb{R}^3 to construct a 3×43 \times 4 relative pose matrix.

    Because the two branches share weights and decoders symmetrically, the model eliminates bias arising from image input ordering.

  2. Knowl 2 — Scale-Free Angular Loss Formulation for Relative Pose Regression

    equation

    Reloc3r trains its relative camera pose network by measuring angular discrepancies between predicted and ground-truth camera orientations and translation directions, omitting the metric translation magnitude.

    Let R^,R∈SO(3)\hat{R}, R \in \mathrm{SO}(3) be the predicted and ground-truth relative rotation matrices, and let t^,t∈R3∖{0}\hat{t}, t \in \mathbb{R}^3 \setminus \{\mathbf{0}\} be the predicted and ground-truth relative translation vectors. The total training objective L\mathcal{L} is defined as the unweighted sum of the geodesic rotation error ℓR\ell_R and the cosine translation angular error ℓt\ell_t:

    L=ℓR+ℓt\mathcal{L} = \ell_R + \ell_t

    where: ℓR=arccos⁡(tr(R^−1R)−12)\ell_R = \arccos\left( \frac{\mathrm{tr}(\hat{R}^{-1} R) - 1}{2} \right) ℓt=arccos⁡(t^⋅t∥t^∥∥t∥)\ell_t = \arccos\left( \frac{\hat{t} \cdot t}{\lVert \hat{t} \rVert \lVert t \rVert} \right)

    Here, tr(⋅)\mathrm{tr}(\cdot) denotes matrix trace, ⋅\cdot denotes the vector dot product, and ∥⋅∥\lVert \cdot \rVert is the Euclidean norm. Formulating translation error as an angle places both rotation and translation supervision in identical units (radians), removing the need for dataset-specific loss weighting hyperparameters and avoiding inconsistencies across datasets with differing metric scales.

  3. Knowl 3 — Motion Averaging Module for Absolute 6-DoF Camera Localization

    model/method

    To localize a query image IqI_q within a reference coordinate frame defined by a database of posed images D={(Idn,Pdn)}n=1N\mathcal{D} = \{(I_{d_n}, P_{d_n})\}_{n=1}^N, where each database pose Pdn=(Rdn,cdn)P_{d_n} = (R_{d_n}, c_{d_n}) consists of rotation Rdn∈SO(3)R_{d_n} \in \mathrm{SO}(3) and optical center cdn∈R3c_{d_n} \in \mathbb{R}^3, Reloc3r aggregates pair-wise relative predictions without RANSAC or robust non-linear optimizers:

    1. Pair Selection: The top-KK database images {Id1,…,IdK}\{I_{d_1}, \dots, I_{d_K}\} most similar to IqI_q are retrieved via an image retrieval method (such as NetVLAD, with K=10K = 10). For each pair (Iq,Idk)(I_q, I_{d_k}), the relative pose regressor infers relative rotation R^q,dk\hat{R}_{q, d_k} and unit translation direction t^q,dk\hat{t}_{q, d_k}.

    2. Rotation Averaging: Each pair yields an estimate of the absolute query rotation R^q(k)=RdkR^q,dk\hat{R}_q^{(k)} = R_{d_k} \hat{R}_{q, d_k}. To ensure robustness against outliers, the global query rotation R^q\hat{R}_q is obtained by computing the component-wise geometric median across all KK rotation estimates converted to quaternion representations.

    3. Camera Center Triangulation: The absolute query camera position cq∈R3c_q \in \mathbb{R}^3 is determined from the database camera centers cdkc_{d_k} and relative translation rays t^q,dk\hat{t}_{q, d_k}. A linear least-squares problem is formulated to find cqc_q minimizing the sum of squared Euclidean orthogonal distances to all KK rays: c^q=arg⁡min⁡c∑k=1K∥(I−t^q,dkt^q,dkT)(c−cdk)∥2\hat{c}_q = \arg\min_{c} \sum_{k=1}^K \left\lVert (I - \hat{t}_{q, d_k} \hat{t}_{q, d_k}^T) (c - c_{d_k}) \right\rVert^2 The closed-form solution for c^q\hat{c}_q is obtained directly using Singular Value Decomposition (SVD).

    The final absolute 6-DoF pose is P^q=(R^q,−R^qc^q)∈R3×4\hat{P}_q = (\hat{R}_q, -\hat{R}_q \hat{c}_q) \in \mathbb{R}^{3 \times 4}.

  4. Knowl 4 — Large-Scale Multi-Domain Pre-training Setup

    experimental setup

    Reloc3r is trained on approximately 8 million image pairs collated from seven diverse public datasets encompassing object-centric, indoor, and outdoor scenes:

    • CO3Dv2: ∼1.0M\sim 1.0\text{M} pairs (object-centric)
    • ScanNet++: ∼850K\sim 850\text{K} pairs (indoor)
    • ARKitScenes: ∼2.1M\sim 2.1\text{M} pairs (indoor)
    • BlendedMVS: ∼1.0M\sim 1.0\text{M} pairs (outdoor)
    • MegaDepth: ∼1.8M\sim 1.8\text{M} pairs (outdoor)
    • DL3DV: ∼1.1M\sim 1.1\text{M} pairs (indoor & outdoor)
    • RealEstate10K: ∼100K\sim 100\text{K} pairs (indoor & outdoor)

    All relative poses are converted to standard OpenCV coordinate frames. Images are centered at their principal points and resized to a width of 512 pixels (or 224 pixels for the fast variant).

    Training parameters: The architecture uses m=24m = 24 encoder blocks and n=12n = 12 decoder blocks, initialized with pre-trained DUSt3R 512-DPT weights (with the decoder initialized from DUSt3R's decoder2). It is trained on 8 AMD MI250x (40GB) GPUs using memory-efficient attention (reducing memory by 25% and increasing throughput by 14%), a per-GPU batch size of 8, and an Adam-style optimizer with learning rate decaying from 1×10−51 \times 10^{-5} down to 1×10−71 \times 10^{-7}.

  5. Knowl 5 — Benchmark Evaluation of Relative Pose Estimation on ScanNet1500, RealEstate10K, and ACID

    data/table

    Pair-wise relative camera pose regression accuracy is evaluated using AUC@5°, AUC@10°, and AUC@20° (the area under the cumulative error curve thresholded at maximum rotation and translation angular errors of 5°, 10°, and 20°), alongside per-pair inference runtime in milliseconds.

    Method ScanNet1500 RealEstate10K ACID Time (ms)
    AUC@5 AUC@10 AUC@20 AUC@5 AUC@10 AUC@20 AUC@5 AUC@10 AUC@20
    Non-PR
    Efficient LoFTR 19.20 37.00 53.60 - - - - - - 40
    ROMA 28.90 50.40 68.30 54.60 69.80 79.70 46.30 58.80 68.90 300
    DUSt3R 23.81 45.91 65.57 39.70 56.88 70.43 21.50 35.95 49.70 441
    MASt3R 28.01 50.24 68.83 63.54 76.39 84.50 52.12 64.54 73.61 294
    NoPoSplat 31.80 53.80 71.70 69.10 80.60 87.70 48.60 61.70 72.80 >2000
    PR
    Map-free (Regress-SN) 1.84 8.75 25.33 0.83 4.06 13.97 1.32 5.82 16.28 10
    Map-free (Regress-MF) 0.50 3.48 13.15 1.61 6.74 18.38 2.57 9.96 24.50 10
    ExReNet (SN) 2.30 10.71 26.13 2.17 7.94 20.43 1.90 7.53 18.69 17
    ExReNet (SUNCG) 1.61 7.00 18.03 3.27 12.06 27.85 4.14 13.43 27.70 17
    Reloc3r-224 (Ours) 28.34 52.60 71.56 59.70 75.05 84.71 28.25 47.34 62.54 15
    Reloc3r-512 (Ours) 34.79 58.37 75.56 66.70 80.20 88.39 38.18 56.39 70.34 25

    Reloc3r-512 outperforms all prior pose regression (PR) baselines by substantial margins across indoor (ScanNet1500), mixed indoor/outdoor (RealEstate10K), and completely unseen aerial outdoor (ACID) scenes. On ScanNet1500, Reloc3r-512 sets the top performance across all thresholds, outperforming state-of-the-art non-PR matchers while operating at 25 ms per pair (over 50×50\times faster than dense correspondence approaches like ROMA and NoPoSplat).

  6. Knowl 6 — Visual Localization Benchmark on 7 Scenes and Cambridge Landmarks

    data/table

    Absolute visual localization performance is evaluated by reporting median position error (meters) and median orientation error (degrees) across indoor scenes (7 Scenes) and outdoor scenes (Cambridge Landmarks). Reloc3r is evaluated zero-shot without per-scene training.

    Dataset / Metric APR (e.g. DFNet+NeFeS) Seen RPR (e.g. CamNet) Unseen RPR (ImageNet+NCM) Unseen RPR (ExReNet) Reloc3r-224 Reloc3r-512
    7 Scenes (Indoor)
    Chess 0.02 / 0.57^ 0.04 / 1.73^ - 0.05 / 1.63^ 0.03 / 0.99^ 0.03 / 0.88^
    Fire 0.02 / 0.74^ 0.03 / 1.74^ - 0.07 / 2.54^ 0.04 / 1.13^ 0.03 / 0.81^
    Heads 0.02 / 1.28^ 0.05 / 1.98^ - 0.03 / 2.71^ 0.02 / 1.23^ 0.01 / 0.95^
    Office 0.02 / 0.56^ 0.04 / 1.62^ - 0.06 / 1.75^ 0.05 / 0.88^ 0.04 / 0.88^
    Pumpkin 0.02 / 0.55^ 0.04 / 1.64^ - 0.07 / 2.04^ 0.07 / 1.14^ 0.06 / 1.10^
    RedKitchen 0.02 / 0.57^ 0.04 / 1.63^ - 0.07 / 2.10^ 0.05 / 1.23^ 0.04 / 1.26^
    Stairs 0.05 / 1.28^ 0.04 / 1.51^ - 0.19 / 4.87^ 0.12 / 2.25^ 0.07 / 1.26^
    Average (7S) 0.02 / 0.79^ 0.04 / 1.69^ 0.19 / 4.30^ 0.08 / 2.52^ 0.05 / 1.26^ 0.04 / 1.02^
    Cambridge (Outdoor)
    GreatCourt - - - 9.79 / 4.46^ 1.71 / 0.94^ 1.22 / 0.73^
    KingsCollege 0.37 / 0.54^ - - 2.33 / 2.48^ 0.47 / 0.41^ 0.42 / 0.36^
    OldHospital 0.52 / 0.88^ - - 3.54 / 3.49^ 0.87 / 0.66^ 0.62 / 0.55^
    ShopFacade 0.15 / 0.53^ - - 0.72 / 2.41^ 0.18 / 0.53^ 0.13 / 0.58^
    StMarysChurch 0.37 / 1.14^ - - 2.30 / 3.72^ 0.41 / 0.73^ 0.34 / 0.58^
    Average (4 Scenes) 0.35 / 0.77^ - 0.83 / 1.36^ 2.22 / 3.03^ 0.48 / 0.58^ 0.38 / 0.52^
    Average (5 Scenes) - 1.37 / 2.30^ - 3.74 / 3.31^ 0.73 / 0.65^ 0.55 / 0.56^

    On 7 Scenes, Reloc3r-512 reaches an average error of 0.04 m/1.02∘0.04\text{ m} / 1.02^\circ, outperforming previous unseen RPR methods (such as ExReNet at 0.08 m/2.52∘0.08\text{ m} / 2.52^\circ) and matching per-scene trained APR models without requiring scene-specific training. On Cambridge Landmarks, Reloc3r achieves 0.38 m/0.52∘0.38\text{ m} / 0.52^\circ across four standard scenes, halving the median error of prior unseen RPR methods and surpassing APR approaches in rotational accuracy.

  7. Knowl 7 — Multi-View Relative Pose Benchmark on CO3Dv2

    data/table

    Relative camera pose performance is evaluated across 41 categories of the object-centric CO3Dv2 dataset using 10 frames per sequence (evaluating all 45 pair combinations). Metrics include Relative Rotation Accuracy within 15° (RRA@15), Relative Translation Accuracy within 15° (RTA@15), and mean Average Accuracy at 30° (mAA@30 / AUC@30).

    Method RRA@15 (%) RTA@15 (%) mAA@30 (%)
    Non-PR
    PixSfM 33.7 32.9 30.1
    RelPose 57.1 - -
    PoseDiffusion 80.5 79.8 66.5
    RelPose++ 82.3 77.2 65.1
    RayDiffusion (8 frames) 93.3 - -
    VGGSfM 92.1 88.3 74.0
    DUSt3R (w/ PnP) 94.3 88.4 77.2
    MASt3R 94.6 91.9 81.1
    PR
    PoseReg 53.2 49.1 45.0
    RayReg (8 frames) 89.2 - -
    Reloc3r-224 (Ours) 93.6 91.9 79.1
    Reloc3r-512 (Ours) 95.8 93.7 82.9

    Reloc3r-512 achieves 95.8%95.8\% RRA@15, 93.7%93.7\% RTA@15, and 82.9%82.9\% mAA@30, outperforming all structure-based, diffusion-based, and feed-forward regression methods despite evaluating solely on isolated image pairs without joint multi-view bundle adjustment.

  8. Knowl 8 — Ablation on Network Symmetry and Scale-Free Supervision

    empirical result

    Ablation experiments on the ScanNet1500 dataset assess the architectural choice of branch symmetry and scale-free direction learning against asymmetric branches and metric translation regression:

    Model Variant AUC@5 AUC@10 AUC@20
    Reloc3r-512 (Asymmetric Architecture) 32.71 56.84 74.63
    Reloc3r-512 (Metric Pose Output) 25.70 50.20 70.07
    Reloc3r-512 (Default: Symmetric + Angular Loss) 34.79 58.37 75.56
    1. Asymmetric vs. Symmetric Architecture: Using independent decoders and regression heads for each image branch degrades AUC@5 from 34.7934.79 to 32.7132.71 while increasing model size and memory overhead. Weight sharing between symmetric branches enforces pose consistency P^I1,I2≈P^I2,I1−1\hat{P}_{I_1, I_2} \approx \hat{P}_{I_2, I_1}^{-1} and removes input ordering bias.

    2. Metric Pose vs. Scale-Free Pose Supervision: Direct regression of metric translations results in a sharp performance drop (AUC@5 drops to 25.7025.70). Regressing translation directions as scale-free unit vectors prevents training imbalance across datasets with disparate physical scales and leaves metric scaling to linear triangulation.

  9. Knowl 9 — Collinear Viewpoint Degeneracy in Triangulation

    limitation

    Reloc3r fails to determine absolute metric camera translation when the query viewpoint and all retrieved reference database camera centers are perfectly or near-perfectly collinear. Under strict collinearity, the line-of-sight translation vectors do not provide non-zero baseline parallax across intersecting planes, causing the linear least-squares triangulation system solved via SVD to become degenerate and rendering the global translation scale unconstrained.

Coverage note — No substantial contributed material was omitted. Qualitative cross-attention patch-matching observations mentioned in Section 4.3 were excluded as standalone knowls because they represent interpretability visualizations rather than quantitative results or primary algorithmic contributions.

References

  1. 1.Yehya Abouelnaga, Mai Bui, and Slobodan Ilic. Distillpose: Lightweight camera localization using auxiliary learning. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7919–7924. IEEE, 2021.
  2. 2.Relja Arandjelovic and Andrew Zisserman. Three things everyone should know to improve object retrieval. In 2012 IEEE conference on computer vision and pattern recognition, pages 2911–2918. IEEE, 2012.
  3. 3.Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. Netvlad: Cnn architecture for weakly supervised place recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5297–5307, 2016.
  4. 4.Eduardo Arnold, Jamie Wynn, Sara Vicente, Guillermo Garcia-Hernando, Aron Monszpart, Victor Prisacariu, Daniyar Turmukhambetov, and Eric Brachmann. Map-free visual relocalization: Metric pose relative to a single image. In European Conference on Computer Vision, pages 690–708. Springer, 2022.
  5. 5.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  6. 6.Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros. Sequential modeling enables scalable learning for large vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22861–22872, 2024.
  7. 7.Vassileios Balntas, Shuda Li, and Victor Prisacariu. Relocnet: Continuous metric learning relocalisation using neural nets. In Proceedings of the European conference on computer vision (ECCV), pages 751–767, 2018.
  8. 8.Daniel Barath and Jiˇr´ı Matas. Graph-cut ransac. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6733–6741, 2018.
  9. 9.Daniel Barath and Jiri Matas. Graph-cut ransac: Local optimization on spatially coherent structures. IEEE transactions on pattern analysis and machine intelligence, 44(9): 4961–4974, 2021.
  10. 10.Daniel Barath, Jiri Matas, and Jana Noskova. Magsac: marginalizing sample consensus. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10197–10205, 2019.
  11. 11.Daniel Barath, Jana Noskova, Maksym Ivashechkin, and Jiri Matas. Magsac++, a fast, reliable and accurate robust estimator. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1304–1312, 2020.
  12. 12.Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. arXiv preprint arXiv:2111.08897, 2021.
  13. 13.Eric Brachmann and Carsten Rother. Learning less is more-6d camera localization via 3d surface regression. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4654–4662, 2018.
  14. 14.Eric Brachmann and Carsten Rother. Visual camera relocalization from rgb and rgb-d images using dsac. IEEE transactions on pattern analysis and machine intelligence, 44(9):5847–5865, 2021.
  15. 15.Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. Dsac-differentiable ransac for camera localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6684–6692, 2017.
  16. 16.Eric Brachmann, Tommaso Cavallari, and Victor Adrian Prisacariu. Accelerated coordinate encoding: Learning to relocalize in minutes using rgb and poses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5044–5053, 2023.
  17. 17.G Bradski. The opencv library. Dr. Dobb’s Journal of Software Tools, 2000.
  18. 18.Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-aware learning of maps for camera localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2616–2625, 2018.
  19. 19.Avishek Chatterjee and Venu Madhav Govindu. Efficient and robust large-scale rotation averaging. In Proceedings of the IEEE international conference on computer vision, pages 521–528, 2013.
  20. 20.Shuai Chen, Zirui Wang, and Victor Prisacariu. Directposenet: Absolute pose regression with photometric consistency. In 2021 International Conference on 3D Vision (3DV), pages 1175–1185. IEEE, 2021.
  21. 21.Shuai Chen, Xinghui Li, Zirui Wang, and Victor A Prisacariu. Dfnet: Enhance absolute pose regression with direct feature matching. In European Conference on Computer Vision, pages 1–17. Springer, 2022.
  22. 22.Shuai Chen, Yash Bhalgat, Xinghui Li, Jia-Wang Bian, Kejie Li, Zirui Wang, and Victor Adrian Prisacariu. Neural refinement for absolute pose regression with feature synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20987–20996, 2024.
  23. 23.Shuai Chen, Tommaso Cavallari, Victor Adrian Prisacariu, and Eric Brachmann. Map-relative pose regression for visual re-localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20665–20674, 2024.
  24. 24.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1–113, 2023.
  25. 25.Ondˇrej Chum, Jiˇr´ı Matas, and Josef Kittler. Locally optimized ransac. In Pattern Recognition: 25th DAGM Symposium, Magdeburg, Germany, September 10-12, 2003. Proceedings 25, pages 236–243. Springer, 2003.
  26. 26.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
  27. 27.Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superpoint: Self-supervised interest point detection and description. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 224–236, 2018.
  28. 28.Jacob Devlin. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  29. 29.Mingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi, and Ping Luo. Camnet: Coarse-to-fine retrieval for camera relocalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2871–2880, 2019.
  30. 30.Siyan Dong, Shuzhe Wang, Yixin Zhuang, Juho Kannala, Marc Pollefeys, and Baoquan Chen. Visual localization via few-shot scene region classification. In 2022 International Conference on 3D Vision (3DV), pages 393–402. IEEE, 2022.
  31. 31.Siyan Dong, Shaohui Liu, Hengkai Guo, Baoquan Chen, and Marc Pollefeys. Lazy visual localization via motion averaging. arXiv preprint arXiv:2307.09981, 2023.
  32. 32.Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  33. 33.Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-net: A trainable cnn for joint description and detection of local features. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, pages 8092–8101, 2019.
  34. 34.Johan Edstedt, Qiyu Sun, Georg B¨okman, M˚arten Wadenb¨ack, and Michael Felsberg. Roma: Robust dense feature matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19790–19800, 2024.
  35. 35.Sovann En, Alexis Lechervy, and Fr´ed´eric Jurie. Rpnet: An end-to-end network for relative camera pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018.
  36. 36.Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
  37. 37.Khang Truong Giang, Soohwan Song, and Sungho Jo. Learning to produce semi-dense correspondences for visual localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19468–19478, 2024.
  38. 38.Richard Hartley and Andrew Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003.
  39. 39.Richard I Hartley and Peter Sturm. Triangulation. Computer vision and image understanding, 68(2):146–157, 1997.
  40. 40.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022.
  41. 41.Martin Humenberger, Yohann Cabon, Nicolas Guerin, Julien Morat, Vincent Leroy, J´erˆome Revaud, Philippe Rerole, No´e Pion, Cesar de Souza, and Gabriela Csurka. Robust image retrieval-based visual localization using kapture. arXiv preprint arXiv:2007.13867, 2020.
  42. 42.Alex Kendall and Roberto Cipolla. Modelling uncertainty in deep learning for camera relocalization. In 2016 IEEE international conference on Robotics and Automation (ICRA), pages 4762–4769. IEEE, 2016.
  43. 43.Alex Kendall and Roberto Cipolla. Geometric loss functions for camera pose regression with deep learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5974–5983, 2017.
  44. 44.Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. In Proceedings of the IEEE international conference on computer vision, pages 2938–2946, 2015.
  45. 45.Fadi Khatib, Yuval Margalit, Meirav Galun, and Ronen Basri. Leveraging image matching toward end-to-end relative camera pose regression. arXiv preprint arXiv:2211.14950, 2022.
  46. 46.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023.
  47. 47.Laurent Kneip, Davide Scaramuzza, and Roland Siegwart. A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation. In CVPR 2011, pages 2969–2976. IEEE, 2011.
  48. 48.Viktor Larsson and contributors. PoseLib - Minimal Solvers for Camera Pose Estimation, 2020.
  49. 49.Zakaria Laskar, Iaroslav Melekhov, Surya Kalia, and Juho Kannala. Camera relocalization by computing pairwise relative poses using convolutional neural network. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 929–938, 2017.
  50. 50.Karel Lebeda, Jirı Matas, and Ondrej Chum. Fixing the locally optimized ransac–full experimental evaluation. In British machine vision conference. Citeseer Princeton, NJ, USA, 2012.
  51. 51.Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Grounding image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024.
  52. 52.Jake Levinson, Carlos Esteves, Kefan Chen, Noah Snavely, Angjoo Kanazawa, Afshin Rostamizadeh, and Ameesh Makadia. An analysis of svd for deep rotation estimation. Advances in Neural Information Processing Systems, 33: 22554–22565, 2020.
  53. 53.Xiaotian Li, Juha Ylioinas, and Juho Kannala. Full-frame scene coordinate regression for image-based localization. arXiv preprint arXiv:1802.03237, 2018.
  54. 54.Xiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek, and Juho Kannala. Hierarchical scene coordinate classification and regression for visual localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11983–11992, 2020.
  55. 55.Yunpeng Li, Noah Snavely, and Daniel P Huttenlocher. Location recognition using prioritized feature matching. In Computer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part II 11, pages 791–804. Springer, 2010.
  56. 56.Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2041–2050, 2018.
  57. 57.Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani. Relpose++: Recovering 6d poses from sparse-view observations. In 2024 International Conference on 3D Vision (3DV), pages 106–115. IEEE, 2024.
  58. 58.Jingyu Lin, Jiaqi Gu, Bojian Wu, Lubin Fan, Renjie Chen, Ligang Liu, and Jieping Ye. Learning neural volumetric pose features for camera localization. arXiv preprint arXiv:2403.12800, 2024.
  59. 59.Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, and Marc Pollefeys. Pixel-perfect structure-from-motion with featuremetric refinement. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5987–5997, 2021.
  60. 60.Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys. Lightglue: Local feature matching at light speed. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 17627–17638, 2023.
  61. 61.Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22160–22169, 2024.
  62. 62.Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14458–14467, 2021.
  63. 63.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  64. 64.Shaohui Liu, Yidan Gao, Tianyi Zhang, R´emi Pautrat, Johannes L Sch¨onberger, Viktor Larsson, and Marc Pollefeys. Robust incremental structure-from-motion with hybrid features. In European Conference on Computer Vision, pages 249–269. Springer, 2025.
  65. 65.David G Lowe. Object recognition from local scale-invariant features. In Proceedings of the seventh IEEE international conference on computer vision, pages 1150–1157. Ieee, 1999.
  66. 66.Iaroslav Melekhov, Juha Ylioinas, Juho Kannala, and Esa Rahtu. Image-based localization using hourglass networks. In Proceedings of the IEEE international conference on computer vision workshops, pages 879–886, 2017.
  67. 67.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  68. 68.Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. Large language models: A survey. arXiv preprint arXiv:2402.06196, 2024.
  69. 69.Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, and Arnaud de La Fortelle. Lens: Localization enhanced by nerf synthesis. In Conference on Robot Learning, pages 1347–1356. PMLR, 2022.
  70. 70.Tayyab Naseer and Wolfram Burgard. Deep regression for monocular camera-based 6-dof global localization in outdoor environments. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1525–1530. IEEE, 2017.
  71. 71.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730–27744, 2022.
  72. 72.Vojtech Panek, Zuzana Kukelova, and Torsten Sattler. Meshloc: Mesh-based visual localization. In European Conference on Computer Vision, pages 589–609. Springer, 2022.
  73. 73.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023.
  74. 74.Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024.
  75. 75.Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10901–10911, 2021.
  76. 76.Jerome Revaud, Cesar De Souza, Martin Humenberger, and Philippe Weinzaepfel. R2d2: Reliable and repeatable detector and descriptor. Advances in neural information processing systems, 32, 2019.
  77. 77.Chris Rockwell, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park, Justin Johnson, and David F Fouhey. Far: Flexible accurate and robust 6dof relative camera pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19854–19864, 2024.
  78. 78.Soham Saha, Girish Varma, and CV Jawahar. Improved visual relocalization by discovering anchor points. arXiv preprint arXiv:1811.04370, 2018.
  79. 79.Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In CVPR, 2019.
  80. 80.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4938–4947, 2020.
  81. 81.Paul-Edouard Sarlin, Ajaykumar Unagar, Mans Larsson, Hugo Germain, Carl Toft, Viktor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, et al. Back to the feature: Learning robust camera localization from pixels to pose. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3247–3257, 2021.
  82. 82.Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Fast image-based localization using direct 2d-to-3d matching. In 2011 International Conference on Computer Vision, pages 667–674. IEEE, 2011.
  83. 83.Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Improving image-based localization by active correspondence search. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part I 12, pages 752–765. Springer, 2012.
  84. 84.Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE transactions on pattern analysis and machine intelligence, 39(9):1744–1756, 2016.
  85. 85.Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, and Tomas Pajdla. Are large-scale 3d models really necessary for accurate visual localization? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1637–1646, 2017.
  86. 86.Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, et al. Benchmarking 6dof outdoor visual localization in changing conditions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8601–8610, 2018.
  87. 87.Torsten Sattler, Qunjie Zhou, Marc Pollefeys, and Laura Leal-Taixe. Understanding the limitations of cnn-based absolute camera pose regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3302–3312, 2019.
  88. 88.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016.
  89. 89.Yoli Shavit, Ron Ferens, and Yosi Keller. Learning multi-scene absolute pose regression with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2733–2742, 2021.
  90. 90.Yoli Shavit, Ron Ferens, and Yosi Keller. Coarse-to-fine multi-scene pose regression with transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  91. 91.Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. Scene coordinate regression forests for camera relocalization in rgb-d images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2930–2937, 2013.
  92. 92.Brandon Smart, Chuanxia Zheng, Iro Laina, and Victor Adrian Prisacariu. Splatt3r: Zero-shot gaussian splatting from uncalibarated image pairs. arXiv preprint arXiv:2408.13912, 2024.
  93. 93.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1746–1754, 2017.
  94. 94.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568: 127063, 2024.
  95. 95.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8922–8931, 2021.
  96. 96.Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Akihiko Torii. Inloc: Indoor visual localization with dense matching and view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7199–7209, 2018.
  97. 97.Shitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu, and Ping Tan. Learning camera localization via dense scene matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1831–1841, 2021.
  98. 98.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. arXiv preprint arXiv:2406.16860, 2024.
  99. 99.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  100. 100.Diantao Tu, Hainan Cui, Xianwei Zheng, and Shuhan Shen. Panopose: Self-supervised relative pose estimation for panoramic images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20009–20018, 2024.
  101. 101.Mehmet Ozgur Turkoglu, Eric Brachmann, Konrad Schindler, Gabriel J Brostow, and Aron Monszpart. Visual camera re-localization using graph neural networks and relative pose supervision. In 2021 International Conference on 3D Vision (3DV), pages 145–155. IEEE, 2021.
  102. 102.A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017.
  103. 103.Florian Walch, Caner Hazirbas, Laura Leal-Taixe, Torsten Sattler, Sebastian Hilsenbeck, and Daniel Cremers. Image-based localization using lstms for structured feature correlation. In Proceedings of the IEEE international conference on computer vision, pages 627–637, 2017.
  104. 104.Bing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao, Niki Trigoni, and Andrew Markham. Atloc: Attention guided camera localization. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 10393–10401, 2020.
  105. 105.Hengyi Wang and Lourdes Agapito. 3d reconstruction with spatial memory. arXiv preprint arXiv:2408.16061, 2024.
  106. 106.Jianyuan Wang, Christian Rupprecht, and David Novotny. Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9773–9783, 2023.
  107. 107.Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21686–21697, 2024.
  108. 108.Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. arXiv preprint arXiv:2410.19115, 2024.
  109. 109.Shuzhe Wang, Juho Kannala, and Daniel Barath. Dgc-gnn: Leveraging geometry and color cues for visual descriptor-free 2d-3d matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20881–20891, 2024.
  110. 110.Shuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Xiaotian Li, Yi Zhao, Giorgos Tolias, and Juho Kannala. Hscnet++: Hierarchical scene coordinate classification and regression for visual localization with transformer. International Journal of Computer Vision, pages 1–21, 2024.
  111. 111.Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697–20709, 2024.
  112. 112.Yifan Wang, Xingyi He, Sida Peng, Dongli Tan, and Xiaowei Zhou. Efficient loftr: Semi-dense local feature matching with sparse-like speed. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21666–21675, 2024.
  113. 113.Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Br´egier, Yohann Cabon, Vaibhav Arora, Leonid Antsfeld, Boris Chidlovskii, Gabriela Csurka, and J´erˆome Revaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion. Advances in Neural Information Processing Systems, 35:3502–3516, 2022.
  114. 114.Philippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon, Vaibhav Arora, Romain Br´egier, Gabriela Csurka, Leonid Antsfeld, Boris Chidlovskii, and J´erˆome Revaud. Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 17969–17980, 2023.
  115. 115.Kyle Wilson and Noah Snavely. Robust global translations with 1dsfm. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part III 13, pages 61–75. Springer, 2014.
  116. 116.Dominik Winkelbauer, Maximilian Denninger, and Rudolph Triebel. Learning to localize in new environments from synthetic training data. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 5840–5846. IEEE, 2021.
  117. 117.Jian Wu, Liwei Ma, and Xiaolin Hu. Delving deeper into convolutional neural networks for camera relocalization. In 2017 IEEE International Conference on Robotics and Automation (ICRA), pages 5644–5651. IEEE, 2017.
  118. 118.Fei Xue, Xin Wu, Shaojun Cai, and Junqiu Wang. Learning multi-view camera relocalization with graph neural networks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11372–11381. IEEE, 2020.
  119. 119.Luwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li, Yasutaka Furukawa, and Ping Tan. Sanet: Scene agnostic network for camera localization. In Proceedings of the IEEE/CVF international conference on computer vision, pages 42–51, 2019.
  120. 120.Luwei Yang, Rakesh Shrestha, Wenbo Li, Shuaicheng Liu, Guofeng Zhang, Zhaopeng Cui, and Ping Tan. Scenesqueezer: Learning to compress scene for camera relocalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8259–8268, 2022.
  121. 121.Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1790–1799, 2020.
  122. 122.Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images. arXiv preprint arXiv:2410.24207, 2024.
  123. 123.Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d indoor scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023.
  124. 124.Yingda Yin, Yang Wang, He Wang, and Baoquan Chen. A laplace-inspired distribution on so(3) for probabilistic rotation estimation. In International Conference on Learning Representations (ICLR), 2023.
  125. 125.Bernhard Zeisl, Torsten Sattler, and Marc Pollefeys. Camera pose voting for large-scale image-based localization. In Proceedings of the IEEE International Conference on Computer Vision, pages 2704–2712, 2015.
  126. 126.Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. Monst3r: A simple approach for estimating geometry in the presence of motion. arXiv preprint arXiv:2410.03825, 2024.
  127. 127.Jason Y. Zhang, Deva Ramanan, and Shubham Tulsiani. RelPose: Predicting probabilistic relative rotation for single objects in the wild. In European Conference on Computer Vision, 2022.
  128. 128.Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. arXiv preprint arXiv:2402.14817, 2024.
  129. 129.Qunjie Zhou, Torsten Sattler, Marc Pollefeys, and Laura Leal-Taixe. To learn or not to learn: Visual localization from essential matrices. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 3319–3326. IEEE, 2020.
  130. 130.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. arXiv preprint arXiv:1805.09817, 2018.
  131. 131.Bingbing Zhuang, Loong-Fah Cheong, and Gim Hee Lee. Baseline desensitizing in translation averaging. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4539–4547, 2018.

Citation

MLA
Dong, S., et al. “Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 16739–52, https://doi.org/10.1109/CVPR52734.2025.01560.
APA
Dong, S., Wang, S., Liu, S., Cai, L., Fan, Q., Kannala, J., & Yang, Y. (2025). Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16739–16752. https://doi.org/10.1109/CVPR52734.2025.01560
Chicago
Dong, S., S. Wang, S. Liu, et al. 2025. “Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16739–52. https://doi.org/10.1109/CVPR52734.2025.01560.
Harvard
Dong, S. et al. (2025) “Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 16739–16752. Available at: https://doi.org/10.1109/CVPR52734.2025.01560.
Vancouver
1. Dong S, Wang S, Liu S, Cai L, Fan Q, Kannala J, Yang Y (2025) Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 16739–16752

BibTeX

@inproceedings{Dong_2025, title={Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01560}, DOI={10.1109/cvpr52734.2025.01560}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Dong, Siyan and Wang, Shuzhe and Liu, Shaohui and Cai, Lulu and Fan, Qingnan and Kannala, Juho and Yang, Yanchao}, year={2025}, month=June, pages={16739–16752} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE