keyword
Hamming distance
Hamming distance is a mathematical metric that measures the difference between two sequences of equal length by counting the number of positions at which their corresponding symbols differ. Originally introduced by Richard Hamming for error detection and correction in telecommunications, it quantifies the minimum number of substitutions required to transform one sequence into another. In computer science and digital information processing, it is commonly applied to binary vectors and bitstrings, where it can be evaluated efficiently using a bitwise exclusive-OR operation followed by a count of the set bits. Beyond coding theory and network communication, Hamming distance functions as a core similarity metric across diverse domains, including bioinformatics for comparing genetic sequences, cryptography, and computer vision and machine learning for nearest-neighbor search, hash-based retrieval, and binary feature matching.
3 items

Non-parametric Local Transforms for Computing Visual Correspondence
Ramin Zabih, John Woodfill
Why you should read this
Introduces the rank and census transforms to establish visual correspondence using relative pixel intensity ordering, providing illumination invariance and high accuracy near object boundaries where conventional correlation techniques fail.
We propose a new approach to the correspondence problem that makes use of non-parametric local transforms as the basis for correlation. Non-parametric local transforms rely on the relative ordering of local intensity values, and not on the intensity values themselves. Correlation using such transforms can tolerate a significant number of outliers. This can result in improved performance near object boundaries when compared with conventional methods such as normalized correlation. We introduce two non-parametric local transforms: the rank transform, which measures local intensity, and the census transform, which summarizes local image structure. We describe some properties of these transforms, and demonstrate their utility on both synthetic and real data.
Added
2026-09-17

Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search
Hervé Jégou, Matthijs Douze, Cordelia Schmid
Why you should read this
Presents Hamming embedding and weak geometric consistency to refine visual-word descriptor matching and filter geometrically inconsistent features directly within an inverted file, substantially increasing retrieval accuracy on million-scale image databases.
This paper improves recent methods for large scale image search. State-of-the-art methods build on the bag-of-features image representation. We, first, analyze bag-of-features in the framework of approximate nearest neighbor search. This shows the sub-optimality of such a representation for matching descriptors and leads us to derive a more precise representation based on 1) Hamming embedding (HE) and 2) weak geometric consistency constraints (WGC). HE provides binary signatures that refine the matching based on visual words. WGC filters matching descriptors that are not consistent in terms of angle and scale. HE and WGC are integrated within the inverted file and are efficiently exploited for all images, even in the case of very large datasets. Experiments performed on a dataset of one million of images show a significant improvement due to the binary signature and the weak geometric consistency constraints, as well as their efficiency. Estimation of the full geometric transformation, i.e., a re-ranking step on a short list of images, is complementary to our weak geometric consistency constraints and allows to further improve the accuracy.
Added
2026-09-16

Near-optimal hashing algorithms for approximate nearest neighbor in high dimensions
Alexandr Andoni, Piotr Indyk
Why you should read this
Explains the foundational theory behind locality-sensitive hashing and provides sub-linear query time guarantees for vector matching while confronting the curse of dimensionality.
In this article, we give an overview of efficient algorithms for the approximate and exact nearest neighbor problem. The goal is to preprocess a dataset of objects (e.g., images) so that later, given a new query object, one can quickly return the dataset object that is most similar to the query. The problem is of significant interest in a wide variety of areas.
Added
2026-05-03
