keyword
visual transformers
Visual transformers are deep learning architectures based on the self-attention mechanism designed to process and interpret visual data, including images, videos, and three-dimensional point clouds. Unlike conventional convolutional neural networks that rely on localized sliding filters, visual transformers segment visual inputs into sequences of discrete tokens, such as image patches or spatial point groups, and compute relationships across all tokens simultaneously. Originally adapted from natural language processing models, this approach allows visual transformers to capture global context and long-range dependencies across an entire scene, serving as versatile frameworks for computer vision tasks such as image classification, object detection, semantic segmentation, and geometric shape analysis.
2 items

Deep Visual Geo-localization Benchmark
Gabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone, Gabriela Csurka, Torsten Sattler, Barbara Caputo
Why you should read this
Establishes a modular visual geo-localization benchmarking framework and conducts extensive evaluations across pipeline components, offering practical guidelines that balance retrieval accuracy with real-world computational and memory costs.
In this paper, we propose a new open-source benchmark-ing framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used ar-chitectures, with the flexibility to change individual compo-nents of a geo-localization pipeline. The purpose of this framework is twofold: i) gaining insights into how differ-ent components and design choices in a VG pipeline im-pact the final results, both in terms of performance (re-call@N metric) and system requirements (such as execu-tion time and memory consumption); ii) establish a system-atic evaluation protocol for comparing different methods. Using the proposed framework, we perform a large suite of experiments which provide criteria for choosing back-bone, aggregation and negative mining depending on the use-case and requirements. We also assess the impact of engineering techniques like pre/post-processing, data aug-mentation and image resizing, showing that better perfor-mance can be obtained through somewhat simple proce-dures: for example, downscaling the images' resolution to 80% can lead to similar results with a 36% savings in ex-traction time and dataset storage requirement. Code and trained models are available at dataset storage require-ment. https://deep-vg-bench.herokuapp.com/.
Added
2026-09-26

PCT: Point cloud transformer
Meng-Hao Guo, Junxiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph Robert Martin, Shimin Hu
Why you should read this
Proposes a specialized Transformer architecture for 3D point clouds that utilizes permutation invariance and localized feature aggregation to achieve state-of-the-art accuracy across shape classification, part segmentation, and normal estimation benchmarks.
The irregular domain and lack of ordering make it challenging to design deep neural networks for point cloud processing. This paper presents a novel framework named Point Cloud Transformer(PCT) for point cloud learning. PCT is based on Transformer, which achieves huge success in natural language processing and displays great potential in image processing. It is inherently permutation invariant for processing a sequence of points, making it well-suited for point cloud learning. To better capture local context within the point cloud, we enhance input embedding with the support of farthest point sampling and nearest neighbor search. Extensive experiments demonstrate that the PCT achieves the state-of-the-art performance on shape classification, part segmentation and normal estimation tasks.
Added
2026-09-16
