Built independently by an author, for readers. Read the story and support ChapterPal

keyword

SIFT descriptors

A SIFT descriptor, short for Scale-Invariant Feature Transform descriptor, is a local visual feature representation used in computer vision to identify and match distinct points of interest across digital images. Introduced by David Lowe, the descriptor computes the distribution of local image intensity gradients within a localized region surrounding a detected keypoint, typically encoding these gradients into a 128-dimensional numerical vector based on orientation histograms. By normalizing gradient orientations relative to the keypoints dominant local orientation and measuring them at a scale determined by keypoint detection, the descriptor achieves robust invariance to image scaling, rotation, illumination variations, and moderate changes in viewpoint. These properties make SIFT descriptors fundamental tools for tasks such as object recognition, image stitching, 3D reconstruction, and visual category classification.

10 items

Learning To Count Objects in Images

Learning To Count Objects in Images

V. Lempitsky, Andrew Zisserman

OrganizationsUniversity of Oxford

Why you should read this

Develops a supervised object counting framework that bypasses individual instance detection by learning an image density function using dot annotations, optimized with a maximum subarray-based loss for accurate and fast count estimation across image subregions.

We propose a new supervised learning framework for visual object counting tasks, such as estimating the number of cells in a microscopic image or the number of humans in surveillance video frames. We focus on the practically-attractive case when the training images are annotated with dots (one dot per object). Our goal is to accurately estimate the count. However, we evade the hard task of learning to detect and localize individual object instances. Instead, we cast the problem as that of estimating an image density whose integral over any image region gives the count of objects within that region. Learning to infer such density can be formulated as a minimization of a regularized risk quadratic cost function. We introduce a new loss function, which is well-suited for such learning, and at the same time can be computed efficiently via a maximum subarray algorithm. The learning can then be posed as a convex quadratic program solvable with cutting-plane optimization. The proposed framework is very flexible as it can accept any domain-specific visual features. Once trained, our system provides accurate object counts and requires a very small time overhead over the feature extraction step, making it a good candidate for applications involving real-time processing or dealing with huge amount of visual data.

Added

2026-09-25

Improving the Fisher Kernel for Large-Scale Image Classification

Improving the Fisher Kernel for Large-Scale Image Classification

Florent Perronnin, Jorge Sánchez, Thomas Mensink

OrganizationsXerox

Why you should read this

Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.

The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.

Added

2026-09-14

Product Quantization for Nearest Neighbor Search

Product Quantization for Nearest Neighbor Search

Hervé Jégou, Matthijs Douze, Cordelia Schmid

OrganizationsINRIA

Why you should read this

Formulates the mathematics of decomposing high-dimensional spaces into lower-dimensional Cartesian products to compress vectors and radically accelerate distance estimations.

This paper introduces a product quantization-based approach for approximate nearest neighbor search. The idea is to decompose the space into a Cartesian product of low-dimensional subspaces and to quantize each subspace separately. A vector is represented by a short code composed of its subspace quantization indices. The euclidean distance between two vectors can be efficiently estimated from their codes. An asymmetric version increases precision, as it computes the approximate distance between a vector and a code. Experimental results show that our approach searches for nearest neighbors efficiently, in particular in combination with an inverted file system. Results for SIFT and GIST image descriptors show excellent search accuracy, outperforming three state-of-the-art approaches. The scalability of our approach is validated on a data set of two billion vectors.

Added

2026-05-03

Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories

Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories

Svetlana Lazebnik, Cordelia Schmid, Jean Ponce

OrganizationsEcole Normale SupérieureINRIAUniversity of Illinois Urbana-Champaign

Why you should read this

Demonstrates that a simple method of hierarchically partitioning images and aggregating local feature statistics across multiple spatial scales can outperform complex geometric correspondence approaches for scene recognition while providing an efficient and interpretable alternative to standard bag-of-features representations.

This paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors.

Added

2026-02-21

Histograms of Oriented Gradients for Human Detection

Histograms of Oriented Gradients for Human Detection

Navneet Dalal, Bill Triggs

OrganizationsINRIA

Why you should read this

Demonstrates that a simple dense grid of locally normalized histogram-based orientation features dramatically outperforms existing methods for detecting people in images, establishing a practical and interpretable foundation for robust visual object recognition.

We study the question of feature sets for robust visual object recognition; adopting linear SVM based human detection as a test case. After reviewing existing edge and gradient based descriptors, we show experimentally that grids of histograms of oriented gradient (HOG) descriptors significantly outperform existing feature sets for human detection. We study the influence of each stage of the computation on performance, concluding that fine-scale gradients, fine orientation binning, relatively coarse spatial binning, and high-quality local contrast normalization in overlapping descriptor blocks are all important for good results. The new approach gives near-perfect separation on the original MIT pedestrian database, so we introduce a more challenging dataset containing over 1800 annotated human images with a large range of pose variations and backgrounds.

Added

2026-02-21