keyword
compact codes
Compact codes are low-dimensional, highly compressed discrete or binary representations of high-dimensional feature vectors designed to minimize memory usage and accelerate data processing. Generated through techniques such as product quantization, vector quantization, and hashing, these representations compress complex numerical descriptors, such as image features or text embeddings, into compact byte sequences while preserving essential distance and semantic relationships. This encoding allows computational systems to store massive datasets directly in memory and execute rapid similarity search, nearest-neighbor retrieval, and machine learning inference with minimal loss of accuracy.
2 items

FastText.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hervé Jégou, Tomas Mikolov
Why you should read this
Proposes a product quantization approach to compress text classification models by two orders of magnitude with minimal accuracy loss, enabling fastText deployment on memory-constrained devices.
We consider the problem of producing compact architectures for text classification, such that the full model fits in a limited amount of memory. After considering different solutions inspired by the hashing literature, we propose a method built upon product quantization to store word embeddings. While the original technique leads to a loss in accuracy, we adapt this method to circumvent quantization artefacts. Our experiments carried out on several benchmarks show that our approach typically requires two orders of magnitude less memory than fastText while being only slightly inferior with respect to accuracy. As a result, it outperforms the state of the art by a good margin in terms of the compromise between memory usage and accuracy.
Added
2026-09-25

Aggregating Local Image Descriptors into Compact Codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, P. Pérez, Cordelia Schmid
Why you should read this
Proposes an image indexing framework that aggregates local descriptors into compact codes of just a few dozen bytes, enabling accurate visual search across 100 million images in roughly 250 milliseconds on a single processor core.
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image dataset takes about 250 ms on one processor core.
Added
2026-09-24
