Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization
Siyan DongShuzhe WangShaohui LiuLulu CaiQingnan FanJuho KannalaYanchao Yang
Presents Reloc3r, a scalable framework that combines a symmetric relative pose regression network trained on eight million image pairs with minimalist motion averaging, achieving state-of-the-art camera localization accuracy in real time across unseen scenes.
Visual localization determines the exact camera position and orientation of a captured image relative to an existing map or image database, a capability essential for autonomous robotics, navigation, and augmented reality. Existing methods face persistent trade-offs: traditional geometric techniques achieve high accuracy but require intensive computation that restricts real-time deployment, while modern direct regression methods either fail to generalize to novel environments or require computationally expensive, scene-specific retraining.
The article evaluates whether a streamlined neural architecture, combined with large-scale pre-training across diverse environments, can achieve real-time speed, high localization accuracy, and zero-shot generalization to unseen scenes. It introduces Reloc3r, an end-to-end framework designed to resolve the historical performance trade-offs in visual camera localization.
To establish a scalable and robust system, the researchers paired an image-to-image relative pose regression model with a lightweight, parameter-free motion averaging module. The network employs a fully symmetric Vision Transformer backbone with shared weights to estimate relative rotation and translation directions without enforcing metric scale during network training. The motion averaging module then triangulates the absolute metric coordinates and calculates orientation across multiple retrieved reference images. The entire framework was trained once on a large dataset of approximately eight million image pairs spanning indoor, outdoor, and object-centric scenes, and subsequently evaluated across six public benchmark datasets without any scene-specific fine-tuning.
Experimental results show that Reloc3r sets a new performance standard across multiple benchmarks. In pairwise relative pose estimation, it outperforms prior regression models by large margins, delivering top accuracy across indoor and outdoor benchmarks while operating at 24 to 66 frames per second, which is up to 20 to 50 times faster than competing non-regression methods. In visual localization on unseen outdoor environments, Reloc3r roughly halved the median error of previous relative pose regression methods, achieving an average error of 0.38 meters and 0.52 degrees on Cambridge Landmarks. In indoor localization on the 7 Scenes benchmark, it attained an average median error of 0.04 meters and 1.02 degrees, matching or outperforming specialized models that required days of scene-specific training. Furthermore, on multi-view object datasets, the system achieved best-in-class accuracy (95.8% rotation accuracy within 15 degrees) using only pairwise evaluations.
These findings demonstrate that relative pose regression does not suffer from fundamental accuracy limits when supported by sufficient data diversity and scaled transformer architectures. For commercial and operational applications, this eliminates the need for costly per-scene 3D model construction and offline retraining cycles. The ability to achieve high precision with real-time inference latency of 15 to 42 milliseconds directly reduces compute infrastructure costs and enhances onboard autonomy in dynamic, real-world deployments.
Organizations developing vision-based navigation and augmented reality tools should consider adopting data-driven relative pose estimation as a viable alternative to complex structure-from-motion pipelines. For practical implementation, practitioners must ensure that image retrieval retrieval mechanisms select reference frames with adequate geometric baseline separation. Future engineering efforts should focus on integrating active viewpoint selection to mitigate rare geometric collinearity failures, where linear camera paths prevent absolute metric triangulation.
The results provide a high degree of confidence across standard indoor and outdoor operational domains. A known limitation occurs when the query image and all retrieved reference frames lie in a straight line, which mathematically degrades scale triangulation in the motion averaging step. Nevertheless, the framework demonstrates strong robustness across varied baseline separations and novel environments.
- Paper: DUSt3R: Geometric 3D Vision Made Easy, Shuzhe Wang et al. (2023). Introduces the foundation for feed-forward pairwise geometric estimation and uncalibrated multi-view 3D vision using large-scale transformer architectures that Reloc3r builds upon and adapts for fast camera localization.
- Paper: PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization, Alex Kendall et al. (2015). Pioneers direct deep learning regression for 6-DOF camera relocalization and establishes standard evaluation benchmarks like Cambridge Landmarks and 7 Scenes.
- Paper: Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses, Eric Brachmann et al. (2023). Presents scene coordinate regression techniques and benchmark baselines on indoor and outdoor relocalization against which zero-shot regression models are evaluated.
- Paper: LoFTR: Detector-Free Local Feature Matching with Transformers, Jiaming Sun et al. (2021). Demonstrates transformer-based attention across image pairs for robust geometric correspondence, motivating Reloc3r's symmetric transformer backbone.
- Paper: An efficient solution to the five-point relative pose problem, David Nister (2004). Provides the foundational algebraic and geometric principles of relative pose estimation that neural relative pose regressors seek to approximate and accelerate.
- Paper: NetVLAD: CNN Architecture for Weakly Supervised Place Recognition, Relja Arandjelović et al. (2015). Establishes the image retrieval mechanisms essential for selecting reference database frames during visual localization pipelines.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). Expands on feed-forward transformer-based camera pose and dense geometry estimation by jointly predicting intrinsics, extrinsics, depth, and point tracks across multi-view collections in a single unified architecture.
