Built independently by an author, for readers. Read the story and support ChapterPal

keyword

local descriptors

Local descriptors are compact numerical or binary representations that capture the visual appearance, texture, or geometric structure of small, localized patches around specific interest points in an image. Unlike global descriptors that summarize an entire image in a single feature vector, local descriptors focus on distinct localized regions, allowing systems to establish fine-grained point correspondences between different views of a scene. They are typically designed or learned to remain robust against photometric and geometric variations, such as changes in illumination, scale, rotation, and camera perspective. Consequently, local descriptors serve as foundational components in computer vision for applications such as image matching, multi-view 3D reconstruction, visual localization, image retrieval, and object recognition.

10 items

Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Dror Aiger, Bingyi Cao, Andre Araujo, Kaifeng Chen

OrganizationsGoogle

Why you should read this

Inverts the standard image retrieval workflow by using scalable local feature search for initial candidate retrieval and multidimensional scaling to build query-time global embeddings for fast, highly accurate re-ranking on benchmark datasets.

The dominant paradigm in image retrieval systems today is to search large databases using global image features, and re-rank those initial results with local image feature matching techniques. This design, dubbed global-to-local, stems from the computational cost of local matching approaches, which can only be afforded for a small number of retrieved images. However, emerging efficient local feature search approaches have opened up new possibilities, in particular enabling detailed retrieval at large scale, to find partial matches which are often missed by global feature search. In parallel, global feature-based re-ranking has shown promising results with high computational efficiency. In this work, we leverage these building blocks to introduce a local-to-global retrieval paradigm, where efficient local feature search meets effective global feature re-ranking. Critically, we propose a re-ranking method where global features are computed on-the-fly, based on the local feature retrieval similarities. Such re-ranking-only global features leverage multidimensional scaling techniques to create embeddings which respect the local similarities obtained during search, enabling a significant re-ranking boost. Experimentally, we demonstrate solid retrieval performance, setting new state-of-the-art results on the Revisited Oxford and Paris datasets.

Added

2026-09-29

DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline Stereo

DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline Stereo

Engin Tola, V. Lepetit, P. Fua

OrganizationsÉcole Polytechnique Fédérale de Lausanne

Why you should read this

Proposes a computationally efficient local image descriptor that uses Gaussian convolutions over gradient orientations to enable fast dense matching and accurate depth estimation for wide-baseline stereo pairs.

In this paper, we introduce a local image descriptor, DAISY, which is very efficient to compute densely. We also present an EM based algorithm to compute dense depth and occlusion maps from wide baseline image pairs using this descriptor. This yields much better results in wide baseline situations than the pixel and correlation based algorithms that are commonly used in narrow baseline stereo. Also, using a descriptor makes our algorithm robust against many photometric and geometric transformations. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance when used densely. It is important to note that our approach is the first algorithm that attempts to estimate dense depth maps from wide baseline image pairs and we show that it is a good one at that with many experiments for depth estimation accuracy, occlusion detection, and comparing it against other descriptors on laser scanned ground truth scenes. We also tested our approach on a variety of indoor and outdoor scenes with different photometric and geometric transformations and our experiments support our claim to being robust against these.

Added

2026-09-24

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search

Hervé Jégou, Matthijs Douze, Cordelia Schmid

OrganizationsINRIA

Why you should read this

Presents Hamming embedding and weak geometric consistency to refine visual-word descriptor matching and filter geometrically inconsistent features directly within an inverted file, substantially increasing retrieval accuracy on million-scale image databases.

This paper improves recent methods for large scale image search. State-of-the-art methods build on the bag-of-features image representation. We, first, analyze bag-of-features in the framework of approximate nearest neighbor search. This shows the sub-optimality of such a representation for matching descriptors and leads us to derive a more precise representation based on 1) Hamming embedding (HE) and 2) weak geometric consistency constraints (WGC). HE provides binary signatures that refine the matching based on visual words. WGC filters matching descriptors that are not consistent in terms of angle and scale. HE and WGC are integrated within the inverted file and are efficiently exploited for all images, even in the case of very large datasets. Experiments performed on a dataset of one million of images show a significant improvement due to the binary signature and the weak geometric consistency constraints, as well as their efficiency. Estimation of the full geometric transformation, i.e., a re-ranking step on a short list of images, is complementary to our weak geometric consistency constraints and allows to further improve the accuracy.

Added

2026-09-16

A performance evaluation of local descriptors

A performance evaluation of local descriptors

Krystian Mikolajczyk, Cordelia Schmid

OrganizationsINRIAUniversity of Oxford

Why you should read this

Evaluates prominent local image descriptors across diverse transformations to demonstrate that SIFT-based representations consistently achieve the highest matching accuracy regardless of the interest region detector used.

In this paper, we compare the performance of descriptors computed for local interest regions, as, for example, extracted by the Harris-Affine detector [32]. Many different descriptors have been proposed in the literature. It is unclear which descriptors are more appropriate and how their performance depends on the interest region detector. The descriptors should be distinctive and at the same time robust to changes in viewing conditions as well as to errors of the detector. Our evaluation uses as criterion recall with respect to precision and is carried out for different image transformations. We compare shape context [3], steerable filters [12], PCA-SIFT [19], differential invariants [20], spin images [21], SIFT [26], complex filters [37], moment invariants [43], and cross-correlation for different types of interest regions. We also propose an extension of the SIFT descriptor and show that it outperforms the original method. Furthermore, we observe that the ranking of the descriptors is mostly independent of the interest region detector and that the SIFT-based descriptors perform best. Moments and steerable filters show the best performance among the low dimensional descriptors.

Added

2026-09-07