Built independently by an author, for readers. Read the story and support ChapterPal

keyword

VLAD representations

VLAD representations, short for Vector of Locally Aggregated Descriptors, are compact global feature vectors used in computer vision and image retrieval to summarize a set of local image descriptors into a single fixed-length vector. To generate a VLAD representation, a visual vocabulary of cluster centers is precomputed using clustering methods such as k-means, and each local descriptor from an image is mapped to its nearest cluster center. For every cluster, the residual differences between the assigned local descriptors and the cluster center are accumulated, preserving first-order statistical distribution information rather than merely counting word frequencies as in standard bag-of-words models. The resulting residual vectors from all clusters are concatenated, normalized, and frequently compressed to support highly efficient similarity matching and scalable nearest-neighbor search across large-scale datasets.

1 item