Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visual localization

Visual localization is the computer vision process of estimating the exact position and orientation, commonly referred to as the camera pose, from which an image was captured relative to a known environment or coordinate frame. It works by matching visual features from a query image against a pre-built spatial representation of the scene, such as a database of geo-referenced photographs, a three-dimensional point cloud, scene coordinate maps, or neural representations. Widely utilized in autonomous driving, robotics, and augmented reality, visual localization systems often employ hierarchical pipelines that combine coarse global image retrieval with fine local feature matching, or leverage learning-based regression models to achieve accurate spatial awareness under varying viewpoints, lighting, and environmental conditions.

7 items

Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Dror Aiger, Bingyi Cao, Andre Araujo, Kaifeng Chen

OrganizationsGoogle

Why you should read this

Inverts the standard image retrieval workflow by using scalable local feature search for initial candidate retrieval and multidimensional scaling to build query-time global embeddings for fast, highly accurate re-ranking on benchmark datasets.

The dominant paradigm in image retrieval systems today is to search large databases using global image features, and re-rank those initial results with local image feature matching techniques. This design, dubbed global-to-local, stems from the computational cost of local matching approaches, which can only be afforded for a small number of retrieved images. However, emerging efficient local feature search approaches have opened up new possibilities, in particular enabling detailed retrieval at large scale, to find partial matches which are often missed by global feature search. In parallel, global feature-based re-ranking has shown promising results with high computational efficiency. In this work, we leverage these building blocks to introduce a local-to-global retrieval paradigm, where efficient local feature search meets effective global feature re-ranking. Critically, we propose a re-ranking method where global features are computed on-the-fly, based on the local feature retrieval similarities. Such re-ranking-only global features leverage multidimensional scaling techniques to create embeddings which respect the local similarities obtained during search, enabling a significant re-ranking boost. Experimentally, we demonstrate solid retrieval performance, setting new state-of-the-art results on the Revisited Oxford and Paris datasets.

Added

2026-09-29

Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses

Accelerated Coordinate Encoding: Learning to Relocalize in Minutes Using RGB and Poses

Eric Brachmann, Tommaso Cavallari, Victor Adrian Prisacariu

OrganizationsNianticUniversity of Oxford

Why you should read this

Presents Accelerated Coordinate Encoding, an approach that trains accurate visual relocalization models in under five minutes from posed RGB images alone, speeding up scene coordinate mapping by up to 300 times compared to existing methods.

Learning-based visual relocalizers exhibit leading pose accuracy, but require hours or days of training. Since training needs to happen on each new scene again, long training times make learning-based relocalization impractical for most applications, despite its promise of high accuracy. In this paper we show how such a system can actually achieve the same accuracy in less than 5 minutes. We start from the obvious: a relocalization network can be split in a scene-agnostic feature backbone, and a scene-specific prediction head. Less obvious: using an MLP prediction head allows us to optimize across thousands of view points simultaneously in each single training iteration. This leads to stable and extremely fast convergence. Furthermore, we substitute effective but slow end-to-end training using a robust pose solver with a curriculum over a reprojection loss. Our approach does not require privileged knowledge, such as depth maps or a 3D model, for speedy training. Overall, our approach is up to 300x faster in mapping than state-of-the-art scene coordinate regression, while keeping accuracy on par. Code is available: https://nianticlabs.github.io/ace

Added

2026-09-26

Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, Yanchao Yang

OrganizationsAalto UniversityETH ZurichUniversity of Hong KongUniversity of OuluVivo

Why you should read this

Presents Reloc3r, a scalable framework that combines a symmetric relative pose regression network trained on eight million image pairs with minimalist motion averaging, achieving state-of-the-art camera localization accuracy in real time across unseen scenes.

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either generalize well to new scenes or provide accurate camera pose estimates. To address these issues, we present Reloc3r, a simple yet effective visual localization framework. It consists of an elegantly designed relative pose regression network, and a minimalist motion averaging module for absolute pose estimation. Trained on approximately eight million posed image pairs, Reloc3r achieves surprisingly good performance and generalization ability. We conduct extensive experiments on six public datasets, consistently demonstrating the effectiveness and efficiency of the proposed method. It provides high-quality camera pose estimates in real time and generalizes to novel scenes. Code: https://github.com/ffrivera0/reloc3r.

Added

2026-09-26

Text2Loc: 3D Point Cloud Localization from Natural Language

Text2Loc: 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers

OrganizationsLudwig Maximilian University of MunichMunich Center for Machine LearningTechnical University of MunichUniversity of Oxford

Why you should read this

Proposes Text2Loc, a coarse-to-fine framework that localizes natural language descriptions within city-scale 3D point clouds by combining a hierarchical transformer for cross-sentence context with a matching-free fine localization network.

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network, Text2Loc, that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place recognition, followed by fine localization. In global place recognition, relational dynamics among each textual hint are captured in a hierarchical transformer with max-pooling (HTM), whereas a balance between positive and negative pairs is maintained using text-submap contrastive learning. Moreover, we propose a novel matching-free fine localization method to further refine the location predictions, which completely removes the need for complicated text-instance matching and is lighter, faster, and more accurate than previous methods. Extensive experiments show that Text2Loc improves the localization accuracy by up to 2× over the state-of-the-art on the KITTI360Pose dataset. Our project page is publicly available at https://yan-xia.github.io/projects/text2loc/.

Added

2026-09-26

Renderable Neural Radiance Map for Visual Navigation

Renderable Neural Radiance Map for Visual Navigation

Obin Kwon, Jeongho Park, Songhwai Oh

Why you should read this

Proposes a 2D grid-based neural radiance map that embeds visual scene features into latent codes to enable real-time camera tracking, image-based localization, and image-goal movement across unseen environments without scene-specific retraining.

We propose a novel type of map for visual navigation, a renderable neural radiance map (RNR-Map), which is designed to contain the overall visual information of a 3D environment. The RNR-Map has a grid form and consists of latent codes at each pixel. These latent codes are embedded from image observations, and can be converted to the neural radiance field which enables image rendering given a camera pose. The recorded latent codes implicitly contain visual information about the environment, which makes the RNR-Map visually descriptive. This visual information in RNR-Map can be a useful guideline for visual localization and navigation. We develop localization and navigation frameworks that can effectively utilize the RNR-Map. We evaluate the proposed frameworks on camera tracking, visual localization, and image-goal navigation. Experimental results show that the RNR-Map-based localization framework can find the target location based on a single query image with fast speed and competitive accuracy compared to other baselines. Also, this localization framework is robust to environmental changes, and even finds the most visually similar places when a query image from a different environment is given. The proposed navigation framework outperforms the existing image-goal navigation methods in difficult scenarios, under odometry and actuation noises. The navigation framework shows 65.7% success rate in curved scenarios of the NRNS [21] dataset, which is an improvement of 18.6% over the current state-of-the-art. Project page: https://rllab-snu.github.io/projects/RNR-Map/

Added

2026-09-26

PATS: Patch Area Transportation with Subdivision for Local Feature Matching

PATS: Patch Area Transportation with Subdivision for Local Feature Matching

Junjie Ni, Yijin Li, Zhaoyang Huang, Hongsheng Li, Hujun Bao, Zhaopeng Cui, Guofeng Zhang

OrganizationsThe Chinese University of Hong KongZhejiang UniversityZJU-SenseTime Joint Lab of 3D Vision

Why you should read this

Proposes an optimal transport framework that resolves scale discrepancies in detector-free local feature matching by formulating many-to-many patch correspondences to enable self-supervised scale estimation and coarse-to-fine subdivision.

Local feature matching aims at establishing sparse correspondences between a pair of images. Recently, detector-free methods present generally better performance but are not satisfactory in image pairs with large scale differences. In this paper, we propose Patch Area Transportation with Subdivision (PATS) to tackle this issue. Instead of building an expensive image pyramid, we start by splitting the original image pair into equal-sized patches and gradually resizing and subdividing them into smaller patches with the same scale. However, estimating scale differences between these patches is non-trivial since the scale differences are determined by both relative camera poses and scene structures, and thus spatially varying over image pairs. Moreover, it is hard to obtain the ground truth for real scenes. To this end, we propose patch area transportation, which enables learning scale differences in a self-supervised manner. In contrast to bipartite graph matching, which only handles one-to-one matching, our patch area transportation can deal with many-to-many relationships. PATS improves both matching accuracy and coverage, and shows superior performance in downstream tasks, such as relative pose estimation, visual localization, and optical flow estimation. The source code is available at https://zju3dv.github.io/pats/.

Added

2026-09-26

SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

Ted de Vries Lentsch, Zimin Xia, Holger Caesar, Julian F. P. Kooij

OrganizationsDelft University of Technology

Why you should read this

Proposes a geometry-guided cross-view camera pose estimation method that splits the field of view into directional slices to aggregate aerial features via precomputed masks, cutting median localization error on the VIGOR benchmark by up to 50% while running at 150 frames per second.

This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists of ground and aerial feature extractors, feature aggregators, and a pose predictor. The feature extractors extract dense features from the ground and aerial images. Given a set of candidate camera poses, the feature aggregators construct a single ground descriptor and a set of pose-dependent aerial descriptors. Notably, our novel aerial feature aggregator has a cross-view attention module for ground-view guided aerial feature selection and utilizes the geometric projection of the ground camera’s viewing frustum on the aerial image to pool features. The efficient construction of aerial descriptors is achieved using precomputed masks. SliceMatch is trained using contrastive learning and pose estimation is formulated as a similarity comparison between the ground descriptor and the aerial descriptors. Compared to the state-of-the-art, SliceMatch achieves a 19% lower median localization error on the VIGOR benchmark using the same VGG16 backbone at 150 frames per second, and a 50% lower error when using a ResNet50 backbone.

Added

2026-09-26