Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visual transformers

Visual transformers are deep learning architectures based on the self-attention mechanism designed to process and interpret visual data, including images, videos, and three-dimensional point clouds. Unlike conventional convolutional neural networks that rely on localized sliding filters, visual transformers segment visual inputs into sequences of discrete tokens, such as image patches or spatial point groups, and compute relationships across all tokens simultaneously. Originally adapted from natural language processing models, this approach allows visual transformers to capture global context and long-range dependencies across an entire scene, serving as versatile frameworks for computer vision tasks such as image classification, object detection, semantic segmentation, and geometric shape analysis.

2 items

Deep Visual Geo-localization Benchmark

Deep Visual Geo-localization Benchmark

Gabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone, Gabriela Csurka, Torsten Sattler, Barbara Caputo

OrganizationsConsorzio Interuniversitario Nazionale per l'InformaticaCzech Technical University in PragueNaver Labs EuropePolitecnico di Torino

Why you should read this

Establishes a modular visual geo-localization benchmarking framework and conducts extensive evaluations across pipeline components, offering practical guidelines that balance retrieval accuracy with real-world computational and memory costs.

In this paper, we propose a new open-source benchmark-ing framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used ar-chitectures, with the flexibility to change individual compo-nents of a geo-localization pipeline. The purpose of this framework is twofold: i) gaining insights into how differ-ent components and design choices in a VG pipeline im-pact the final results, both in terms of performance (re-call@N metric) and system requirements (such as execu-tion time and memory consumption); ii) establish a system-atic evaluation protocol for comparing different methods. Using the proposed framework, we perform a large suite of experiments which provide criteria for choosing back-bone, aggregation and negative mining depending on the use-case and requirements. We also assess the impact of engineering techniques like pre/post-processing, data aug-mentation and image resizing, showing that better perfor-mance can be obtained through somewhat simple proce-dures: for example, downscaling the images' resolution to 80% can lead to similar results with a 36% savings in ex-traction time and dataset storage requirement. Code and trained models are available at dataset storage require-ment. https://deep-vg-bench.herokuapp.com/.

Added

2026-09-26