keyword
SIFT descriptors
A SIFT descriptor, short for Scale-Invariant Feature Transform descriptor, is a local visual feature representation used in computer vision to identify and match distinct points of interest across digital images. Introduced by David Lowe, the descriptor computes the distribution of local image intensity gradients within a localized region surrounding a detected keypoint, typically encoding these gradients into a 128-dimensional numerical vector based on orientation histograms. By normalizing gradient orientations relative to the keypoints dominant local orientation and measuring them at a scale determined by keypoint detection, the descriptor achieves robust invariance to image scaling, rotation, illumination variations, and moderate changes in viewpoint. These properties make SIFT descriptors fundamental tools for tasks such as object recognition, image stitching, 3D reconstruction, and visual category classification.
10 items

Matching with PROSAC - progressive sample consensus
Ondřej Chum, Jiri Matas
Why you should read this
Proposes a sample consensus algorithm that accelerates correspondence matching by orders of magnitude over RANSAC through progressive sampling of tentative matches ordered by similarity, while retaining identical worst-case convergence guarantees.
A new robust matching method is proposed. The Progressive Sample Consensus (PROSAC) algorithm exploits the linear ordering defined on the set of correspondences by a similarity function used in establishing tentative correspondences. Unlike RANSAC, which treats all correspondences equally and draws random samples uniformly from the full set, PROSAC samples are drawn from progressively larger sets of top-ranked correspondences. Under the mild assumption that the similarity measure predicts correctness of a match better than random guessing, we show that PROSAC achieves large computational savings. Experiments demonstrate it is often significantly faster (up to more than hundred times) than RANSAC. For the derived size of the sampled set of correspondences as a function of the number of samples already drawn, PROSAC converges towards RANSAC in the worst case. The power of the method is demonstrated on wide-baseline matching problems.
Added
2026-09-25

Learning To Count Objects in Images
V. Lempitsky, Andrew Zisserman
Why you should read this
Develops a supervised object counting framework that bypasses individual instance detection by learning an image density function using dot annotations, optimized with a maximum subarray-based loss for accurate and fast count estimation across image subregions.
We propose a new supervised learning framework for visual object counting tasks, such as estimating the number of cells in a microscopic image or the number of humans in surveillance video frames. We focus on the practically-attractive case when the training images are annotated with dots (one dot per object). Our goal is to accurately estimate the count. However, we evade the hard task of learning to detect and localize individual object instances. Instead, we cast the problem as that of estimating an image density whose integral over any image region gives the count of objects within that region. Learning to infer such density can be formulated as a minimization of a regularized risk quadratic cost function. We introduce a new loss function, which is well-suited for such learning, and at the same time can be computed efficiently via a maximum subarray algorithm. The learning can then be posed as a convex quadratic program solvable with cutting-plane optimization. The proposed framework is very flexible as it can accept any domain-specific visual features. Once trained, our system provides accurate object counts and requires a very small time overhead over the feature extraction step, making it a good candidate for applications involving real-time processing or dealing with huge amount of visual data.
Added
2026-09-25

A Theoretical Analysis of Feature Pooling in Visual Recognition
Y-Lan Boureau, Jean Ponce, Yann LeCun
Why you should read this
Establishes a rigorous theoretical framework explaining why max pooling outperforms average pooling for sparse visual features and demonstrates how pooling cardinality and dictionary size directly govern recognition accuracy.
Many modern visual recognition algorithms incorporate a step of spatial 'pooling', where the outputs of several nearby feature detectors are combined into a local or global 'bag of features', in a way that preserves task-related information while removing irrelevant details. Pooling is used to achieve invariance to image transformations, more compact representations, and better robustness to noise and clutter. Several papers have shown that the details of the pooling operation can greatly influence the performance, but studies have so far been purely empirical. In this paper, we show that the reasons underlying the performance of various pooling methods are obscured by several confounding factors, such as the link between the sample cardinality in a spatial pool and the resolution at which low-level features have been extracted. We provide a detailed theoretical analysis of max pooling and average pooling, and give extensive empirical comparisons for object recognition tasks.
Added
2026-09-25

Image Classification using Random Forests and Ferns
Anna Bosch, Andrew Zisserman, X. Muñoz
Why you should read this
Demonstrates how combining spatial pyramid descriptors of shape and appearance with automatic region-of-interest selection and random forest classifiers significantly improves multi-class image recognition on large-scale benchmarks like Caltech-256.
Added
2026-09-25

Aggregating Local Image Descriptors into Compact Codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, P. Pérez, Cordelia Schmid
Why you should read this
Proposes an image indexing framework that aggregates local descriptors into compact codes of just a few dozen bytes, enabling accurate visual search across 100 million images in roughly 250 milliseconds on a single processor core.
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image dataset takes about 250 ms on one processor core.
Added
2026-09-24

Improving the Fisher Kernel for Large-Scale Image Classification
Florent Perronnin, Jorge Sánchez, Thomas Mensink
Why you should read this
Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.
The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.
Added
2026-09-14

Visual categorization with bags of keypoints
Gabriella Csurka, Christopher R. Dance, Lixin Fan, Jutta Willamowski, Cédric Bray
Why you should read this
Proposes the bag-of-keypoints framework for visual object categorization, demonstrating that vector-quantized local image descriptors paired with support vector machines achieve high accuracy across cluttered scenes without relying on spatial geometry.
We present a novel method for generic visual categorization: the problem of identifying the object content of natural images while generalizing across variations inherent to the object class. This bag of keypoints method is based on vector quantization of affine invariant descriptors of image patches. We propose and compare two alternative implementations using different classifiers: Naïve Bayes and SVM. The main advantages of the method are that it is simple, computationally efficient and intrinsically invariant. We present results for simultaneously classifying seven semantic visual categories. These results clearly demonstrate that the method is robust to background clutter and produces good categorization accuracy even without exploiting geometric information.
Added
2026-09-09

Product Quantization for Nearest Neighbor Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid
Why you should read this
Formulates the mathematics of decomposing high-dimensional spaces into lower-dimensional Cartesian products to compress vectors and radically accelerate distance estimations.
This paper introduces a product quantization-based approach for approximate nearest neighbor search. The idea is to decompose the space into a Cartesian product of low-dimensional subspaces and to quantize each subspace separately. A vector is represented by a short code composed of its subspace quantization indices. The euclidean distance between two vectors can be efficiently estimated from their codes. An asymmetric version increases precision, as it computes the approximate distance between a vector and a code. Experimental results show that our approach searches for nearest neighbors efficiently, in particular in combination with an inverted file system. Results for SIFT and GIST image descriptors show excellent search accuracy, outperforming three state-of-the-art approaches. The scalability of our approach is validated on a data set of two billion vectors.
Added
2026-05-03

Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories
Svetlana Lazebnik, Cordelia Schmid, Jean Ponce
Why you should read this
Demonstrates that a simple method of hierarchically partitioning images and aggregating local feature statistics across multiple spatial scales can outperform complex geometric correspondence approaches for scene recognition while providing an efficient and interpretable alternative to standard bag-of-features representations.
This paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors.
Added
2026-02-21

Histograms of Oriented Gradients for Human Detection
Navneet Dalal, Bill Triggs
Why you should read this
Demonstrates that a simple dense grid of locally normalized histogram-based orientation features dramatically outperforms existing methods for detecting people in images, establishing a practical and interpretable foundation for robust visual object recognition.
We study the question of feature sets for robust visual object recognition; adopting linear SVM based human detection as a test case. After reviewing existing edge and gradient based descriptors, we show experimentally that grids of histograms of oriented gradient (HOG) descriptors significantly outperform existing feature sets for human detection. We study the influence of each stage of the computation on performance, concluding that fine-scale gradients, fine orientation binning, relatively coarse spatial binning, and high-quality local contrast normalization in overlapping descriptor blocks are all important for good results. The new approach gives near-perfect separation on the original MIT pedestrian database, so we introduce a more challenging dataset containing over 1800 annotated human images with a large range of pose variations and backgrounds.
Added
2026-02-21
