keyword
local image descriptors
Local image descriptors are numerical vector representations that capture the visual characteristics, such as texture, gradients, and shape, of specific local patches or keypoint neighborhoods within an image. Designed to provide a distinct summary of a small region, these descriptors are engineered or learned to remain robust and invariant to common variations, including changes in scale, rotation, illumination, and viewpoint. They can be generated using classical handcrafted algorithms or learned end-to-end with deep neural networks. By enabling reliable point-to-point correspondence between different images, local image descriptors serve as foundational elements in computer vision applications such as feature matching, object recognition, 3D reconstruction, and large-scale image retrieval, where multiple local descriptors are often aggregated into compact global representations.
3 items

LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua
Why you should read this
Presents the first fully differentiable deep network architecture that unifies keypoint detection, orientation estimation, and descriptor generation into an end-to-end trainable feature extraction pipeline that outperforms classical methods across standard benchmarks.
We introduce a novel Deep Network architecture that implements the full feature point handling pipeline, that is, detection, orientation estimation, and feature description. While previous works have successfully tackled each one of these problems individually, we show how to learn to do all three in a unified manner while preserving end-to-end differentiability. We then demonstrate that our Deep pipeline outperforms state-of-the-art methods on a number of benchmark datasets, without the need of retraining.
Added
2026-09-25

Aggregating Local Image Descriptors into Compact Codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, P. Pérez, Cordelia Schmid
Why you should read this
Proposes an image indexing framework that aggregates local descriptors into compact codes of just a few dozen bytes, enabling accurate visual search across 100 million images in roughly 250 milliseconds on a single processor core.
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image dataset takes about 250 ms on one processor core.
Added
2026-09-24

AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
Gui-Song Xia, Jingwen Hu, Fan Hu, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang, Xiaoqiang Lu
Why you should read this
Introduces the Aerial Image Dataset (AID), a large-scale benchmark of over ten thousand annotated images that overcomes performance saturation in smaller datasets by establishing baseline evaluations for deep learning models in remote sensing scene classification.
Aerial scene classification, which aims to automatically label an aerial image with a specific semantic category, is a fundamental problem for understanding high-resolution remote sensing imagery. In recent years, it has become an active task in remote sensing area and numerous algorithms have been proposed for this task, including many machine learning and data-driven approaches. However, the existing datasets for aerial scene classification like UC-Merced dataset and WHU-RS19 are with relatively small sizes, and the results on them are already saturated. This largely limits the development of scene classification algorithms. This paper describes the Aerial Image Dataset (AID): a large-scale dataset for aerial scene classification. The goal of AID is to advance the state-of-the-arts in scene classification of remote sensing images. For creating AID, we collect and annotate more than ten thousands aerial scene images. In addition, a comprehensive review of the existing aerial scene classification techniques as well as recent widely-used deep learning methods is given. Finally, we provide a performance analysis of typical aerial scene classification and deep learning approaches on AID, which can be served as the baseline results on this benchmark.
Added
2026-09-16
