Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses
Eric BrachmannTommaso CavallariVictor Adrian Prisacariu
Presents Accelerated Coordinate Encoding, an approach that trains accurate visual relocalization models in under five minutes from posed RGB images alone, speeding up scene coordinate mapping by up to 300 times compared to existing methods.
Visual relocalization is essential for real-time applications such as augmented reality, mobile robotics, and navigation, where a system must quickly determine a camera's precise position and orientation within an environment. While learning-based methods achieve high accuracy and generate compact, privacy-preserving scene representations, they traditionally require hours or days of retraining for every single new scene. This extensive mapping delay creates a major operational bottleneck and incurs high computational and financial costs, severely limiting the real-world viability of learning-based relocalizers.
The article introduces Accelerated Coordinate Encoding (ACE), a visual relocalization pipeline designed to train scene coordinate regression networks in under five minutes while matching or exceeding the accuracy of state-of-the-art methods.
To achieve this, the article separates the relocalization architecture into a fixed, scene-agnostic convolutional feature extractor and a small, scene-specific multi-layer perceptron. In the first minute of mapping, the feature extractor processes standard RGB images and their known camera positions into an 8-million-feature training buffer. Over the next four minutes, the regression network samples randomized patches across thousands of distinct camera views in parallel batches. This process eliminates correlated gradient updates, allowing the system to use high learning rates safely. The framework also replaces slow end-to-end optimization with an adaptive curriculum over a reprojection loss, using a smoothly tightening error boundary to focus model capacity on reliable geometry without needing dedicated depth sensors or 3D meshes.
Evaluations across indoor and outdoor benchmarks demonstrate that ACE maps environments up to 300 times faster than the leading baseline, DSAC*. On indoor benchmarks, ACE mapped scenes in five minutes and achieved a 97.1% success rate on 7Scenes and 99.9% on 12Scenes, matching the accuracy of systems that took 11 to 15 hours to train. On outdoor phone scans from the Wayspots dataset, ACE attained a 52.2% average relocalization rate, outperforming the full baseline's 50.7%. Furthermore, ACE compresses complete scene environments into 4MB map files—reducing storage needs by up to seven times compared to prior neural relocalizers and hundreds of times compared to traditional point-cloud methods. Additionally, ACE proved resilient to budget compute hardware, incurring only a 10% training slowdown on low-cost GPUs where traditional baselines doubled their training times to 28 hours.
These results demonstrate that fast, accurate, and low-cost visual mapping can be achieved on standard consumer devices without requiring specialized depth cameras or costly 3D reconstruction pipelines. By lowering training compute time from 15 hours to under five minutes on budget hardware, the method reduces infrastructure costs, lowers energy consumption, and enables on-demand mapping for dynamic user environments.
Organizations developing spatial computing, augmented reality, or robotics platforms should consider adopting decoupled architectures and patch-decorrelation training workflows to eliminate scene pre-processing bottlenecks. When deploying to expansive outdoor environments, teams can explore multi-model ensembles, as demonstrated by the quad-model variant in the article, to balance map footprint against geometric coverage. Further engineering improvements, such as running buffer generation and network training on parallel execution threads, should be prioritized to reduce mapping latency even further.
Decision-makers should note that ACE relies on known initial mapping camera poses, typically supplied by standard visual odometry or structure-from-motion tools, and performance can still degrade in highly repetitive or untextured outdoor scenes. Nonetheless, the empirical evidence provides high confidence that ACE establishes a practical, highly scalable standard for learning-based visual relocalization.
- Paper: PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization, Alex Kendall et al. (2015). It pioneers learning-based camera relocalization directly from monocular RGB images, providing the foundational framing for learned visual localization systems.
- Paper: NetVLAD: CNN Architecture for Weakly Supervised Place Recognition, Relja Arandjelović et al. (2015). It establishes essential deep feature representations and pooling mechanisms for robust visual place recognition and localization across varying viewpoints.
- Paper: LoFTR: Detector-Free Local Feature Matching with Transformers, Jiaming Sun et al. (2021). It details modern transformer-based dense correspondence and feature matching techniques crucial for establishing accurate multi-view geometric constraints.
- Paper: An efficient solution to the five-point relative pose problem, David Nister (2004). It introduces the core minimal geometric relative pose solver that underpins hypothesis generation in camera pose estimation pipelines.
No sufficiently relevant recommendations were found.
