keyword
large scale image search
Large scale image search is the computational process of efficiently identifying and retrieving visually similar or relevant images from massive repositories containing millions to billions of digital pictures. Unlike small collections where exhaustive image-by-image comparisons are feasible, large-scale systems must balance high retrieval accuracy, sub-second query speeds, and minimal memory consumption. To operate at this magnitude, these systems convert visual content into compact numerical representations—such as aggregated local feature vectors or deep convolutional neural network embeddings—and organize them using scalable indexing techniques, including inverted files, vector quantization, and approximate nearest neighbor search.
3 items

Fine-Tuning CNN Image Retrieval with No Human Annotation
Filip Radenovic, Giorgos Tolias, Ondřej Chum
Why you should read this
Introduces an automated framework that leverages 3D structure-from-motion models to mine hard training examples without human annotation, while introducing trainable Generalized-Mean pooling to achieve state-of-the-art CNN image retrieval performance.
Image descriptors based on activations of Convolutional Neural Networks (CNNs) have become dominant in image retrieval due to their discriminative power, compactness of representation, and search efficiency. Training of CNNs, either from scratch or fine-tuning, requires a large amount of annotated data, where a high quality of annotation is often crucial. In this work, we propose to fine-tune CNNs for image retrieval on a large collection of unordered images in a fully automated manner. Reconstructed 3D models obtained by the state-of-the-art retrieval and structure-from-motion methods guide the selection of the training data. We show that both hard-positive and hard-negative examples, selected by exploiting the geometry and the camera positions available from the 3D models, enhance the performance of particular-object retrieval. CNN descriptor whitening discriminatively learned from the same training data outperforms commonly used PCA whitening. We propose a novel trainable Generalized-Mean (GeM) pooling layer that generalizes max and average pooling and show that it boosts retrieval performance. Applying the proposed method to the VGG network achieves state-of-the-art performance on the standard benchmarks: Oxford Buildings, Paris, and Holidays datasets.
Added
2026-09-24

Aggregating Local Image Descriptors into Compact Codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, P. Pérez, Cordelia Schmid
Why you should read this
Proposes an image indexing framework that aggregates local descriptors into compact codes of just a few dozen bytes, enabling accurate visual search across 100 million images in roughly 250 milliseconds on a single processor core.
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image dataset takes about 250 ms on one processor core.
Added
2026-09-24

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid
Why you should read this
Presents Hamming embedding and weak geometric consistency to refine visual-word descriptor matching and filter geometrically inconsistent features directly within an inverted file, substantially increasing retrieval accuracy on million-scale image databases.
This paper improves recent methods for large scale image search. State-of-the-art methods build on the bag-of-features image representation. We, first, analyze bag-of-features in the framework of approximate nearest neighbor search. This shows the sub-optimality of such a representation for matching descriptors and leads us to derive a more precise representation based on 1) Hamming embedding (HE) and 2) weak geometric consistency constraints (WGC). HE provides binary signatures that refine the matching based on visual words. WGC filters matching descriptors that are not consistent in terms of angle and scale. HE and WGC are integrated within the inverted file and are efficiently exploited for all images, even in the case of very large datasets. Experiments performed on a dataset of one million of images show a significant improvement due to the binary signature and the weak geometric consistency constraints, as well as their efficiency. Estimation of the full geometric transformation, i.e., a re-ranking step on a short list of images, is complementary to our weak geometric consistency constraints and allows to further improve the accuracy.
Added
2026-09-16
