Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
Dror AigerBingyi CaoAndre AraujoKaifeng Chen
Inverts the standard image retrieval workflow by using scalable local feature search for initial candidate retrieval and multidimensional scaling to build query-time global embeddings for fast, highly accurate re-ranking on benchmark datasets.
Visual search systems underpin critical applications such as e-commerce product matching, visual question answering, and fine-grained entity identification. Standard architectures rely on a global-to-local design, which conducts a broad initial search using compact global image features and refines candidate matches using detailed local feature comparisons. This traditional workflow frequently fails when query images share only partial visual overlaps or occlusions with database targets, leading to early candidate loss that post-processing cannot recover.
The article demonstrates a novel local-to-global image retrieval architecture that inverts this convention. The objective is to evaluate whether using scalable local feature matching for initial candidate retrieval, paired with on-the-fly global feature re-ranking, improves retrieval precision across standard benchmarks without excessive computational overhead.
The evaluated approach uses Constrained Approximate Nearest Neighbors to perform fast initial retrieval directly on local image features. It then applies multidimensional scaling—specifically through the iterative SMACOF optimization algorithm—to convert non-metric, localized pairwise similarity scores into query-specific global vector embeddings on the fly. These generated embeddings are merged with global features to execute a refinement step across the candidate shortlist. The authors validated this pipeline using the standard Revisited Oxford and Revisited Paris benchmark datasets, supplemented with a 1-million-image distractor collection to assess large-scale performance.
The experiments show that the local-to-global paradigm establishes a new state of the art in image retrieval. In large-scale testing with 1 million distractor images under the most challenging evaluation settings, the method achieved mean average precision scores of 79.8% on Revisited Oxford and 83.4% on Revisited Paris, surpassing previous leading techniques by 2.1% and 1.4%, respectively. Ablation analyses confirmed that deriving global embeddings directly from local pairwise similarities via multidimensional scaling provides critical performance gains that standard global re-ranking alone cannot match.
These results demonstrate that localized search at scale successfully mitigates candidate drop-off caused by partial visual matches. Operating the entire search and re-ranking pipeline takes approximately 0.7 seconds per query on a 24-core CPU for an index of one million images, indicating that the system can support near real-time multimodal search infrastructure.
Organizations evaluating large-scale visual search architectures should consider adopting local-to-global retrieval pipelines for accuracy-sensitive applications. Further development should focus on applying this multidimensional scaling formulation to alternate similarity metrics and learned matching algorithms, as well as evaluating approximate dimension-reduction techniques to optimize query response times.
Decision-makers should account for trade-offs in storage footprint and computation. Storing uncompressed local descriptors requires approximately 21 kilobytes per image, compared to roughly 10 kilobytes for compressed baseline alternatives. While benchmark results demonstrate high reliability across diverse test cases, practitioners must validate memory capacity and latency requirements within their specific production environments before widespread deployment.
- Paper: Fine-Tuning CNN Image Retrieval with No Human Annotation, Filip Radenovic et al. (2017). Establishes foundational CNN representations with GeM pooling and unsupervised metric learning widely used for global descriptors and benchmarks like Revisited Oxford and Paris.
- Paper: Aggregating Local Image Descriptors into Compact Codes, Hervé Jégou et al. (2012). Introduces essential vector aggregation and indexing techniques that connect local descriptor distributions with scalable global search architectures.
- Paper: Hamming Embedding and Weak Geometric Consistency for Large Scale Image Search, Hervé Jégou et al. (2008). Pioneered inverted-file local feature search with geometric consistency, the classical foundation that the local-to-global paradigm seeks to invert and enhance.
- Paper: NetVLAD: CNN Architecture for Weakly Supervised Place Recognition, Relja Arandjelović et al. (2015). Develops end-to-end differentiable local aggregation into global image descriptors for visual retrieval, establishing the standard global-feature retrieval baseline.
- Paper: Re-ranking Person Re-identification with k-Reciprocal Encoding, Zhun Zhong et al. (2017). Demonstrates the efficacy of post-search reciprocal neighborhood graph re-ranking, providing the conceptual basis for re-ranking with search similarity structures.
- Paper: LoFTR: Detector-Free Local Feature Matching with Transformers, Jiaming Sun et al. (2021). Presents modern transformer-based dense local matching mechanisms that drive modern fine-grained partial matching in visual retrieval.
- Paper: Local Grayvalue Invariants for Image Retrieval, Cordelia Schmid et al. (1997). Introduces the classical invariant local feature matching and geometric verification framework for handling partial occlusions and viewpoint variations in image retrieval.
No sufficiently relevant recommendations were found.
