keyword
local descriptors
Local descriptors are compact numerical or binary representations that capture the visual appearance, texture, or geometric structure of small, localized patches around specific interest points in an image. Unlike global descriptors that summarize an entire image in a single feature vector, local descriptors focus on distinct localized regions, allowing systems to establish fine-grained point correspondences between different views of a scene. They are typically designed or learned to remain robust against photometric and geometric variations, such as changes in illumination, scale, rotation, and camera perspective. Consequently, local descriptors serve as foundational components in computer vision for applications such as image matching, multi-view 3D reconstruction, visual localization, image retrieval, and object recognition.
10 items

Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
Dror Aiger, Bingyi Cao, Andre Araujo, Kaifeng Chen
Why you should read this
Inverts the standard image retrieval workflow by using scalable local feature search for initial candidate retrieval and multidimensional scaling to build query-time global embeddings for fast, highly accurate re-ranking on benchmark datasets.
The dominant paradigm in image retrieval systems today is to search large databases using global image features, and re-rank those initial results with local image feature matching techniques. This design, dubbed global-to-local, stems from the computational cost of local matching approaches, which can only be afforded for a small number of retrieved images. However, emerging efficient local feature search approaches have opened up new possibilities, in particular enabling detailed retrieval at large scale, to find partial matches which are often missed by global feature search. In parallel, global feature-based re-ranking has shown promising results with high computational efficiency. In this work, we leverage these building blocks to introduce a local-to-global retrieval paradigm, where efficient local feature search meets effective global feature re-ranking. Critically, we propose a re-ranking method where global features are computed on-the-fly, based on the local feature retrieval similarities. Such re-ranking-only global features leverage multidimensional scaling techniques to create embeddings which respect the local similarities obtained during search, enabling a significant re-ranking boost. Experimentally, we demonstrate solid retrieval performance, setting new state-of-the-art results on the Revisited Oxford and Paris datasets.
Added
2026-09-29

Learning to compare image patches via convolutional neural networks
Sergey Zagoruyko, Nikos Komodakis
Why you should read this
Demonstrates how specialized convolutional neural network architectures can learn image patch similarity directly from raw pixels to outperform traditional handcrafted feature matching across standard computer vision benchmarks.
In this paper we show how to learn directly from image data (i.e., without resorting to manually-designed features) a general similarity function for comparing image patches, which is a task of fundamental importance for many computer vision problems. To encode such a function, we opt for a CNN-based model that is trained to account for a wide variety of changes in image appearance. To that end, we explore and study multiple neural network architectures, which are specifically adapted to this task. We show that such an approach can significantly outperform the state-of-the-art on several problems and benchmark datasets.
Added
2026-09-25

DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline Stereo
Engin Tola, V. Lepetit, P. Fua
Why you should read this
Proposes a computationally efficient local image descriptor that uses Gaussian convolutions over gradient orientations to enable fast dense matching and accurate depth estimation for wide-baseline stereo pairs.
In this paper, we introduce a local image descriptor, DAISY, which is very efficient to compute densely. We also present an EM based algorithm to compute dense depth and occlusion maps from wide baseline image pairs using this descriptor. This yields much better results in wide baseline situations than the pixel and correlation based algorithms that are commonly used in narrow baseline stereo. Also, using a descriptor makes our algorithm robust against many photometric and geometric transformations. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance when used densely. It is important to note that our approach is the first algorithm that attempts to estimate dense depth maps from wide baseline image pairs and we show that it is a good one at that with many experiments for depth estimation accuracy, occlusion detection, and comparing it against other descriptors on laser scanned ground truth scenes. We also tested our approach on a variety of indoor and outdoor scenes with different photometric and geometric transformations and our experiments support our claim to being robust against these.
Added
2026-09-24

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid
Why you should read this
Presents Hamming embedding and weak geometric consistency to refine visual-word descriptor matching and filter geometrically inconsistent features directly within an inverted file, substantially increasing retrieval accuracy on million-scale image databases.
This paper improves recent methods for large scale image search. State-of-the-art methods build on the bag-of-features image representation. We, first, analyze bag-of-features in the framework of approximate nearest neighbor search. This shows the sub-optimality of such a representation for matching descriptors and leads us to derive a more precise representation based on 1) Hamming embedding (HE) and 2) weak geometric consistency constraints (WGC). HE provides binary signatures that refine the matching based on visual words. WGC filters matching descriptors that are not consistent in terms of angle and scale. HE and WGC are integrated within the inverted file and are efficiently exploited for all images, even in the case of very large datasets. Experiments performed on a dataset of one million of images show a significant improvement due to the binary signature and the weak geometric consistency constraints, as well as their efficiency. Estimation of the full geometric transformation, i.e., a re-ranking step on a short list of images, is complementary to our weak geometric consistency constraints and allows to further improve the accuracy.
Added
2026-09-16

SuperGlue: Learning Feature Matching With Graph Neural Networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
Why you should read this
Introduces SuperGlue, an attention-based graph neural network that solves differentiable optimal transport to match local image features and reject non-matchable points in real time for 3D pose estimation.
This paper introduces SuperGlue, a neural network that matches two sets of local features by jointly finding correspondences and rejecting non-matchable points. Assignments are estimated by solving a differentiable optimal transport problem, whose costs are predicted by a graph neural network. We introduce a flexible context aggregation mechanism based on attention, enabling SuperGlue to reason about the underlying 3D scene and feature assignments jointly. Compared to traditional, hand-designed heuristics, our technique learns priors over geometric transformations and regularities of the 3D world through end-to-end training from image pairs. SuperGlue outperforms other learned approaches and achieves state-of-the-art results on the task of pose estimation in challenging real-world indoor and outdoor environments. The proposed method performs matching in real-time on a modern GPU and can be readily integrated into modern SfM or SLAM systems. The code and trained weights are publicly available at this https URL.
Added
2026-09-14

SuperPoint: Self-Supervised Interest Point Detection and Description
Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
Why you should read this
Proposes a self-supervised fully convolutional network that simultaneously detects keypoints and extracts descriptors in a single forward pass, using homographic adaptation to train on real images without human annotations while outperforming traditional methods like SIFT and ORB.
This paper presents a self-supervised framework for training interest point detectors and descriptors suitable for a large number of multiple-view geometry problems in computer vision. As opposed to patch-based neural networks, our fully-convolutional model operates on full-sized images and jointly computes pixel-level interest point locations and associated descriptors in one forward pass. We introduce Homographic Adaptation, a multi-scale, multi-homography approach for boosting interest point detection repeatability and performing cross-domain adaptation (e.g., synthetic-to-real). Our model, when trained on the MS-COCO generic image dataset using Homographic Adaptation, is able to repeatedly detect a much richer set of interest points than the initial pre-adapted deep model and any other traditional corner detector. The final system gives rise to state-of-the-art homography estimation results on HPatches when compared to LIFT, SIFT and ORB.
Added
2026-09-11

A Comparison of Affine Region Detectors
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, F. Schaffalitzky, Timor Kadir, Luc Van Gool
Why you should read this
Establishes the standard benchmark dataset and quantitative evaluation protocol for assessing affine covariant region detectors across diverse imaging transformations, providing systematic performance comparisons of leading methods like MSER, Harris-Affine, and Hessian-Affine.
The paper gives a snapshot of the state of the art in affine covariant region detectors, and compares their performance on a set of test images under varying imaging conditions. Six types of detectors are included: detectors based on affine normalization around Harris (Mikolajczyk and Schmid, 2002; Schaffalitzky and Zisserman, 2002) and Hessian points (Mikolajczyk and Schmid, 2002), a detector of ‘maximally stable extremal regions’, proposed by Matas et al. (2002); an edge-based region detector (Tuytelaars and Van Gool, 1999) and a detector based on intensity extrema (Tuytelaars and Van Gool, 2000), and a detector of ‘salient regions’,
Added
2026-09-11

BRIEF: Binary Robust Independent Elementary Features
Michael Calonder, Vincent Lepetit, Christoph Strecha, Pascal Fua
Why you should read this
Introduces BRIEF, an ultra-fast local feature descriptor constructed directly from simple pairwise pixel intensity tests on smoothed patches, enabling rapid matching via the Hamming distance with minimal memory requirements while rivaling SURF in recognition performance.
We propose to use binary strings as an efficient feature point descriptor, which we call BRIEF. We show that it is highly discriminative even when using relatively few bits and can be computed using simple intensity difference tests. Furthermore, the descriptor similarity can be evaluated using the Hamming distance, which is very efficient to compute, instead of the L2 norm as is usually done. As a result, BRIEF is very fast both to build and to match. We compare it against SURF and U-SURF on standard benchmarks and show that it yields a similar or better recognition performance, while running in a fraction of the time required by either.
Added
2026-09-11

Visual categorization with bags of keypoints
Gabriella Csurka, Christopher R. Dance, Lixin Fan, Jutta Willamowski, Cédric Bray
Why you should read this
Proposes the bag-of-keypoints framework for visual object categorization, demonstrating that vector-quantized local image descriptors paired with support vector machines achieve high accuracy across cluttered scenes without relying on spatial geometry.
We present a novel method for generic visual categorization: the problem of identifying the object content of natural images while generalizing across variations inherent to the object class. This bag of keypoints method is based on vector quantization of affine invariant descriptors of image patches. We propose and compare two alternative implementations using different classifiers: Naïve Bayes and SVM. The main advantages of the method are that it is simple, computationally efficient and intrinsically invariant. We present results for simultaneously classifying seven semantic visual categories. These results clearly demonstrate that the method is robust to background clutter and produces good categorization accuracy even without exploiting geometric information.
Added
2026-09-09

A performance evaluation of local descriptors
Krystian Mikolajczyk, Cordelia Schmid
Why you should read this
Evaluates prominent local image descriptors across diverse transformations to demonstrate that SIFT-based representations consistently achieve the highest matching accuracy regardless of the interest region detector used.
In this paper, we compare the performance of descriptors computed for local interest regions, as, for example, extracted by the Harris-Affine detector [32]. Many different descriptors have been proposed in the literature. It is unclear which descriptors are more appropriate and how their performance depends on the interest region detector. The descriptors should be distinctive and at the same time robust to changes in viewing conditions as well as to errors of the detector. Our evaluation uses as criterion recall with respect to precision and is carried out for different image transformations. We compare shape context [3], steerable filters [12], PCA-SIFT [19], differential invariants [20], spin images [21], SIFT [26], complex filters [37], moment invariants [43], and cross-correlation for different types of interest regions. We also propose an extension of the SIFT descriptor and show that it outperforms the original method. Furthermore, we observe that the ranking of the descriptors is mostly independent of the interest region detector and that the SIFT-based descriptors perform best. Moments and steerable filters show the best performance among the low dimensional descriptors.
Added
2026-09-07
