topic
local image descriptors
Local image descriptors are compact numerical representations that capture the distinctive visual patterns of small, localized patches around interest points in a digital image. Typically computed after keypoint detection, these feature vectors encode local visual characteristics such as gradient orientations, textures, and intensity distributions while remaining robust or invariant to changes in scale, rotation, illumination, and viewpoint. By enabling reliable point-to-point correspondence between different images of the same scene or object, local image descriptors serve as fundamental components in computer vision applications, including panoramic image stitching, three-dimensional reconstruction, object recognition, and visual simultaneous localization and mapping.
2 items

DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline Stereo
Engin Tola, V. Lepetit, P. Fua
Why you should read this
Proposes a computationally efficient local image descriptor that uses Gaussian convolutions over gradient orientations to enable fast dense matching and accurate depth estimation for wide-baseline stereo pairs.
In this paper, we introduce a local image descriptor, DAISY, which is very efficient to compute densely. We also present an EM based algorithm to compute dense depth and occlusion maps from wide baseline image pairs using this descriptor. This yields much better results in wide baseline situations than the pixel and correlation based algorithms that are commonly used in narrow baseline stereo. Also, using a descriptor makes our algorithm robust against many photometric and geometric transformations. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance when used densely. It is important to note that our approach is the first algorithm that attempts to estimate dense depth maps from wide baseline image pairs and we show that it is a good one at that with many experiments for depth estimation accuracy, occlusion detection, and comparing it against other descriptors on laser scanned ground truth scenes. We also tested our approach on a variety of indoor and outdoor scenes with different photometric and geometric transformations and our experiments support our claim to being robust against these.
Added
2026-09-24

Aggregating Local Image Descriptors into Compact Codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, P. Pérez, Cordelia Schmid
Why you should read this
Proposes an image indexing framework that aggregates local descriptors into compact codes of just a few dozen bytes, enabling accurate visual search across 100 million images in roughly 250 milliseconds on a single processor core.
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image dataset takes about 250 ms on one processor core.
Added
2026-09-24
