Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visual vocabulary

A visual vocabulary is a collection of representative local image patterns used in computer vision to discretize, index, and analyze digital images. It is generated by extracting local feature descriptors from a set of training images and grouping mathematically similar descriptors into distinct clusters, typically using algorithms such as k-means. Each cluster center serves as a discrete visual word. In downstream tasks, local features extracted from an input image are mapped to their nearest visual words in the vocabulary, converting the image into a histogram of visual word frequencies. This representation, commonly known as a bag of visual features, converts complex visual information into compact, structured data that enables efficient image retrieval, object recognition, and scene categorization.

2 items

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search

Hervé Jégou, Matthijs Douze, Cordelia Schmid

OrganizationsINRIA

Why you should read this

Presents Hamming embedding and weak geometric consistency to refine visual-word descriptor matching and filter geometrically inconsistent features directly within an inverted file, substantially increasing retrieval accuracy on million-scale image databases.

This paper improves recent methods for large scale image search. State-of-the-art methods build on the bag-of-features image representation. We, first, analyze bag-of-features in the framework of approximate nearest neighbor search. This shows the sub-optimality of such a representation for matching descriptors and leads us to derive a more precise representation based on 1) Hamming embedding (HE) and 2) weak geometric consistency constraints (WGC). HE provides binary signatures that refine the matching based on visual words. WGC filters matching descriptors that are not consistent in terms of angle and scale. HE and WGC are integrated within the inverted file and are efficiently exploited for all images, even in the case of very large datasets. Experiments performed on a dataset of one million of images show a significant improvement due to the binary signature and the weak geometric consistency constraints, as well as their efficiency. Estimation of the full geometric transformation, i.e., a re-ranking step on a short list of images, is complementary to our weak geometric consistency constraints and allows to further improve the accuracy.

Added

2026-09-16

Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories

Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories

Svetlana Lazebnik, Cordelia Schmid, Jean Ponce

OrganizationsEcole Normale SupérieureINRIAUniversity of Illinois Urbana-Champaign

Why you should read this

Demonstrates that a simple method of hierarchically partitioning images and aggregating local feature statistics across multiple spatial scales can outperform complex geometric correspondence approaches for scene recognition while providing an efficient and interpretable alternative to standard bag-of-features representations.

This paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors.

Added

2026-02-21