Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Dror AigerBingyi CaoAndre AraujoKaifeng Chen

article2025ICCVW3 citations

Inverts the standard image retrieval workflow by using scalable local feature search for initial candidate retrieval and multidimensional scaling to build query-time global embeddings for fast, highly accurate re-ranking on benchmark datasets.

Listen

Visual search systems underpin critical applications such as e-commerce product matching, visual question answering, and fine-grained entity identification. Standard architectures rely on a global-to-local design, which conducts a broad initial search using compact global image features and refines candidate matches using detailed local feature comparisons. This traditional workflow frequently fails when query images share only partial visual overlaps or occlusions with database targets, leading to early candidate loss that post-processing cannot recover.

The article demonstrates a novel local-to-global image retrieval architecture that inverts this convention. The objective is to evaluate whether using scalable local feature matching for initial candidate retrieval, paired with on-the-fly global feature re-ranking, improves retrieval precision across standard benchmarks without excessive computational overhead.

The evaluated approach uses Constrained Approximate Nearest Neighbors to perform fast initial retrieval directly on local image features. It then applies multidimensional scaling—specifically through the iterative SMACOF optimization algorithm—to convert non-metric, localized pairwise similarity scores into query-specific global vector embeddings on the fly. These generated embeddings are merged with global features to execute a refinement step across the candidate shortlist. The authors validated this pipeline using the standard Revisited Oxford and Revisited Paris benchmark datasets, supplemented with a 1-million-image distractor collection to assess large-scale performance.

The experiments show that the local-to-global paradigm establishes a new state of the art in image retrieval. In large-scale testing with 1 million distractor images under the most challenging evaluation settings, the method achieved mean average precision scores of 79.8% on Revisited Oxford and 83.4% on Revisited Paris, surpassing previous leading techniques by 2.1% and 1.4%, respectively. Ablation analyses confirmed that deriving global embeddings directly from local pairwise similarities via multidimensional scaling provides critical performance gains that standard global re-ranking alone cannot match.

These results demonstrate that localized search at scale successfully mitigates candidate drop-off caused by partial visual matches. Operating the entire search and re-ranking pipeline takes approximately 0.7 seconds per query on a 24-core CPU for an index of one million images, indicating that the system can support near real-time multimodal search infrastructure.

Organizations evaluating large-scale visual search architectures should consider adopting local-to-global retrieval pipelines for accuracy-sensitive applications. Further development should focus on applying this multidimensional scaling formulation to alternate similarity metrics and learned matching algorithms, as well as evaluating approximate dimension-reduction techniques to optimize query response times.

Decision-makers should account for trade-offs in storage footprint and computation. Storing uncompressed local descriptors requires approximately 21 kilobytes per image, compared to roughly 10 kilobytes for compressed baseline alternatives. While benchmark results demonstrate high reliability across diverse test cases, practitioners must validate memory capacity and latency requirements within their specific production environments before widespread deployment.

arXiv: 2509.04351

No sufficiently relevant recommendations were found.

Cover for Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

Abstract

The dominant paradigm in image retrieval systems today is to search large databases using global image features, and re-rank those initial results with local image feature matching techniques. This design, dubbed global-to-local, stems from the computational cost of local matching approaches, which can only be afforded for a small number of retrieved images. However, emerging efficient local feature search approaches have opened up new possibilities, in particular enabling detailed retrieval at large scale, to find partial matches which are often missed by global feature search. In parallel, global feature-based re-ranking has shown promising results with high computational efficiency. In this work, we leverage these building blocks to introduce a local-to-global retrieval paradigm, where efficient local feature search meets effective global feature re-ranking. Critically, we propose a re-ranking method where global features are computed on-the-fly, based on the local feature retrieval similarities. Such re-ranking-only global features leverage multidimensional scaling techniques to create embeddings which respect the local similarities obtained during search, enabling a significant re-ranking boost. Experimentally, we demonstrate solid retrieval performance, setting new state-of-the-art results on the Revisited Oxford and Paris datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Local-to-Global Image Retrieval
  • 3.1 Re-ranking Global Features
  • 3.2 Enhanced Re-ranking with MDS
  • 3.3 On-the-fly MDS per query
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Results
  • 4.3 Ablation Study
  • 4.4 Computational Cost and Memory
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Local-to-Global (L2G) Image Retrieval Paradigm

    model/method

    The Local-to-Global (L2G) retrieval architecture reverses the conventional Global-to-Local (G2L) image retrieval framework. In G2L systems, an initial search across large databases is conducted using compact global image features, after which the top retrieved candidates are refined using local feature matching (such as spatial geometric verification or Chamfer distance computation). This can lead to recall loss when query images share only partial visual overlap with database targets.

    In contrast, L2G structures the search pipeline as follows:

    1. Initial Local Feature Search: An efficient search mechanism (such as Constrained Approximate Nearest Neighbors, CANN) queries index local descriptors (e.g., FIRE local features) to identify the top-kk nearest candidate images based on fine-grained, localized visual similarities.

    2. On-the-Fly Metric Embedding via Multidimensional Scaling (MDS): Pairwise local feature dissimilarities between the query and its top-kk retrieved candidates are converted on-the-fly into Euclidean coordinates using multidimensional scaling (specifically SMACOF). This creates a query-specific metric space ("globalized local features") that preserves local similarity relations.

    3. Global Feature Re-ranking and Neighborhood Fusion: The query-specific MDS embeddings are linearly combined with pre-extracted global descriptors (e.g., SuperGlobal) and processed through a graph-based neighborhood aggregation re-ranker, producing the final refined ranking.

  2. Knowl 2 — On-the-Fly Query-Time MDS Re-ranking Algorithm

    algorithm

    The on-the-fly Multidimensional Scaling (MDS) re-ranking pipeline uses an offline pre-computation step to store sparse neighbor distances and an online query step to project localized non-metric dissimilarities into a unified Euclidean embedding space for global re-ranking.

    • Offline Stage: For each database image in the index, an efficient local feature search procedure (EFF-INDEX-QUERY, such as CANN) pre-computes the top-kk nearest index images and stores their pairwise distances in a sparse lookup table INDEX-SPARSE-DISTANCES.
    • Online Query Stage: Given a query image qq, EFF-INDEX-QUERY identifies the top-kk candidates C={c1,…,ck}C = \{c_1, \dots, c_k\}. A (k+1)×(k+1)(k+1) \times (k+1) pairwise dissimilarity matrix DD is constructed: distances from qq to each cic_i come from the query search, while candidate-to-candidate pairwise distances are fetched from INDEX-SPARSE-DISTANCES (missing pairs are set to the maximum distance 1.01.0). Distances are modulated by an exponential power pp, and SMACOF optimizes Euclidean coordinates X∈R(k+1)×dX \in \mathbb{R}^{(k+1) \times d}. The resulting embeddings are fused with global descriptors (SuperGlobal) via weight ww, followed by metric-space neighborhood re-ranking over the top MM candidates.
    Input: Query image qq, index database of local descriptors, precomputed sparse distance table INDEXSPARSEDISTANCESINDEX_SPARSE_DISTANCES, candidate count kk, neighborhood size MM, power parameter pp, convergence threshold ϵ\epsilon, fusion weight ww
    Output: Re-ranked candidate list of database images
    C←EFF-INDEX-QUERY(q,k)C \leftarrow \text{EFF-INDEX-QUERY}(q, k)
    Initialize (k+1)×(k+1)(k+1) \times (k+1) dissimilarity matrix DD for elements {q}∪C\{q\} \cup C
    for each ci∈Cc_i \in C do
        D[q,ci]←(LocalDissimilarity(q,ci))pD[q, c_i] \leftarrow (\text{LocalDissimilarity}(q, c_i))^p
        D[ci,q]←D[q,ci]D[c_i, q] \leftarrow D[q, c_i]
    end for
    for each pair (ci,cj)∈C×C(c_i, c_j) \in C \times C with i≠ji \neq j do
        if (ci,cj)∈INDEXSPARSEDISTANCES(c_i, c_j) \in INDEX_SPARSE_DISTANCES then
            D[ci,cj]←(INDEXSPARSEDISTANCES[ci,cj])pD[c_i, c_j] \leftarrow (INDEX_SPARSE_DISTANCES[c_i, c_j])^p
        else
            D[ci,cj]←1.0D[c_i, c_j] \leftarrow 1.0
        end if
    end for
    X←SMACOF(D,ϵ)X \leftarrow \text{SMACOF}(D, \epsilon)
    SMDS←ComputePairwiseCosineSimilarities(X)S_{MDS} \leftarrow \text{ComputePairwiseCosineSimilarities}(X)
    SGlobal←ComputePairwiseSimilarities(q,C,SuperGlobalFeatures)S_{Global} \leftarrow \text{ComputePairwiseSimilarities}(q, C, \text{SuperGlobalFeatures})
    SFused←(1−w)⋅SMDS+w⋅SGlobalS_{Fused} \leftarrow (1 - w) \cdot S_{MDS} + w \cdot S_{Global}
    RerankedList←NeighborhoodReRanking(SFused,M)RerankedList \leftarrow \text{NeighborhoodReRanking}(S_{Fused}, M)
    return RerankedListRerankedList
  3. Knowl 3 — SMACOF Multidimensional Scaling for Incomplete and Non-Metric Dissimilarities

    model/method

    Pairwise dissimilarities generated by local feature retrieval methods (such as asymmetric Chamfer distances produced by CANN) present two structural obstacles for downstream metric-space re-rankers:

    1. Non-metric Dissimilarities: Chamfer distances violate metric properties such as the triangle inequality and symmetry. Re-ranking algorithms that assume a metric space produce inconsistent results when directly supplied with non-metric dissimilarities.
    2. Incomplete / Sparse Distance Matrices: Calculating all O(N2)O(N^2) pairwise distances across a database of size NN is computationally prohibitive at scale. Therefore, only distances between each image and its nearest neighbors are retained.

    To map these dissimilarities into a metric space, the system uses the SMACOF (Scaling by MAjorizing a COmplicated Function) iterative optimization algorithm. SMACOF minimizes the stress function:

    σ(X)=∑i<jwij(dij−∥xi−xj∥2)2\sigma(X) = \sum_{i < j} w_{ij} (d_{ij} - \|x_i - x_j\|_2)^2

    where dijd_{ij} is the input pairwise dissimilarity between images ii and jj, xi,xj∈Rpx_i, x_j \in \mathbb{R}^p are the reconstructed coordinate vectors in pp-dimensional Euclidean space, and wij≥0w_{ij} \ge 0 is a weighting factor (set to assign default maximal distance dij=1.0d_{ij}=1.0 when distance pairs are unobserved in the sparse index matrix).

    SMACOF operates with a computational complexity of O(Nc2)O(N_c^2) per iteration, where Nc=k+1N_c = k+1 is the number of localized candidates (with convergence typically reached in 5–105\text{--}10 iterations for k=700k=700).

  4. Knowl 4 — State-of-the-Art Retrieval Benchmark Comparison

    data/table

    The Local-to-Global (L2G) retrieval pipeline (CANN-FIRE search with SMACOF MDS re-ranking) was benchmarked against existing global and hybrid retrieval methods on Revisited Oxford (ROxf) and Revisited Paris (RPar), evaluating base datasets and their large-scale extensions with 1 million distractors (ROxf+1M, RPar+1M) under Medium and Hard evaluation protocols using mean Average Precision (mAP).

    Medium Hard
    Method ROxf ROxf+1M RPar RPar+1M ROxf ROxf+1M RPar RPar+1M
    (1) Global feature retrieval
    RN50-DELG 73.6 60.6 85.7 68.6 51.0 32.7 71.5 44.4
    RN101-DELG 76.3 63.7 86.6 70.6 55.6 37.5 72.4 46.9
    RN50-DOLG 80.5 76.6 89.8 80.8 58.8 52.2 77.7 62.8
    RN101-DOLG 81.5 77.4 91.0 83.3 61.1 54.8 80.3 66.7
    RN50-CVNet 81.0 72.6 88.8 79.0 62.1 50.2 76.5 60.2
    RN101-CVNet 80.2 74.0 90.3 80.6 63.1 53.7 79.1 62.2
    RN50-SuperGlobal (No re-ranking) 83.9 74.7 90.5 81.3 67.7 53.6 80.3 65.2
    RN101-SuperGlobal (No re-ranking) 85.3 78.8 92.1 83.9 72.1 61.9 83.5 69.1
    (2) Global retrieval + Local re-ranking
    RN50-DELG (GV re-rank top 100) 78.3 67.2 85.7 69.6 57.9 43.6 71.0 45.7
    RN101-DELG (GV re-rank top 100) 81.2 69.1 87.2 71.5 64.0 47.5 72.8 48.7
    RN50-CVNet (Re-rank top 400) 87.9 80.7 90.5 82.4 75.6 65.1 80.2 67.3
    RN101-CVNet (Re-rank top 400) 87.2 81.9 91.2 83.8 75.9 67.4 81.1 69.3
    (3) SuperGlobal / AMES retrieval + Re-ranking
    RN50-SuperGlobal (Re-rank top 400) 88.8 80.0 92.0 83.4 77.1 64.2 84.4 68.7
    RN101-SuperGlobal (Re-rank top 400) 90.9 84.4 93.3 84.9 80.2 71.1 86.7 71.4
    AMES (600,600) (Re-rank top 1600) 93.6 88.2 95.3 90.1 84.8 77.7 90.7 82.0
    (4) Local retrieval + Global re-ranking
    L2G CANN-FIRE + MDS re-ranking (Ours) 92.9 90.5 97.1 92.1 83.0 79.8 91.7 83.4

    L2G achieves the highest mAP on the large-scale distractor benchmarks (+1M), obtaining 79.8% on ROxf+1M Hard (+2.1% over AMES) and 83.4% on RPar+1M Hard (+1.4% over AMES), as well as on RPar base benchmarks.

  5. Knowl 5 — Ablation Analysis of L2G Pipeline Components

    data/table

    The contribution of each pipeline component in the Local-to-Global (L2G) framework was evaluated on the Hard evaluation splits of Revisited Paris (RPar) and Revisited Oxford (ROxf) using mean Average Precision (mAP).

    Configuration RParis (Hard) ROxford (Hard)
    Full Model 91.7 83.0
    - Without SuperGlobal re-ranking process 84.6 73.3
    - Without merging final similarities with SuperGlobal (w=1.0w=1.0) 89.4 81.6
    - Replace MDS re-ranking by SuperGlobal re-ranking 86.4 72.4
    - Replace FIRE local similarity by AMES 93.6 85.2

    Key observations:

    • Removing the neighborhood graph refinement stage ("Without SuperGlobal re-ranking process") degrades mAP by 7.1% on RPar and 9.7% on ROxf, demonstrating that pairwise graph refinement over candidate sets is essential.
    • Omitting the fusion with global SuperGlobal features (w=1.0w=1.0) drops performance to 89.4% on RPar and 81.6% on ROxf, demonstrating complementarity between global semantics and MDS local embeddings.
    • Replacing MDS re-ranking with direct SuperGlobal re-ranking of raw local dissimilarities leads to a substantial performance drop (86.4% on RPar, 72.4% on ROxf), showing that non-metric ranking dissimilarities are incompatible with metric-space graph re-rankers without MDS projection.
    • Replacing FIRE local similarities with learned AMES local similarities further increases accuracy to 93.6% on RPar and 85.2% on ROxf, confirming framework modularity.
  6. Knowl 6 — Experimental Configuration and Hyperparameter Tuning

    experimental setup

    Experiments evaluated the L2G pipeline on Revisited Oxford (ROxf, 4,993 database images) and Revisited Paris (RPar, 6,322 database images), each with 70 query images, combined with the 1,000,000 distractor set (R1M). Performance was evaluated using mean Average Precision (mAP) under the Medium and Hard benchmarks.

    Component setup and hyperparameter tuning:

    • Local Features: 600 FIRE local descriptors per image.
    • Search Algorithm: CANN index using asymmetric Chamfer distance.
    • MDS Engine: SMACOF Multidimensional Scaling.
    • Global Features: SuperGlobal embeddings.

    All hyperparameters were tuned on a small validation sample of 1,000 Oxford images:

    • MDS candidate pool size: k=700k = 700 top-ranked initial candidates.
    • Chamfer power modulation: p=0.01p = 0.01 (controls scaling of large vs. small distances in the Chamfer metric).
    • MDS convergence threshold: ϵ=0.1\epsilon = 0.1.
    • Feature fusion weight: w=0.19w = 0.19 (weight balancing SuperGlobal global features against MDS embeddings).
    • Re-ranking candidate pool size: M=1600M = 1600 shortlisted candidates.
    • SuperGlobal neighborhood parameters: kSG=6k_{SG} = 6, β=0.31\beta = 0.31.
  7. Knowl 7 — Computational Latency and Memory Footprint

    empirical result

    The computational cost and storage footprint of the complete L2G pipeline evaluated on the 1-million image distractor benchmark (ROxf+1M) are:

    • Query Execution Time: Total query latency is approximately 0.7 seconds0.7\text{ seconds} running on a 24-core CPU, partitioned into:
      • Initial local feature search using CANN: ≈0.2 seconds\approx 0.2\text{ seconds}.
      • On-the-fly SMACOF MDS embedding computation and neighborhood re-ranking for the top 1,600 candidates: ≈0.5 seconds\approx 0.5\text{ seconds}.
    • Memory Footprint: Storing 600 uncompressed FIRE local features per image requires approximately 21 kB21\text{ kB} of memory per index image. While this exceeds the storage requirements of pure global feature models or AMES with PQ8 compression (approximately 10 kB10\text{ kB} per image), it enables scalable initial local search.
  8. Knowl 8 — Trade-offs and Limitations of the L2G Pipeline

    limitation

    The Local-to-Global (L2G) retrieval pipeline involves specific computational and storage trade-offs compared to purely global retrieval methods:

    1. Index Storage Footprint: Retaining uncompressed local descriptors (600 FIRE features per database image) requires approximately 21 kB21\text{ kB} per image, which is higher than storing single global feature vectors or compressed representations such as AMES PQ8 ( ≈10 kB\,\approx 10\text{ kB} per image).
    2. Re-ranking Latency: Query-time on-the-fly Multidimensional Scaling (MDS) adds an extra optimization step (approximately 0.5 seconds0.5\text{ seconds} on a 24-core CPU for 1,600 candidates) compared to standard global dot-product re-ranking, although it remains fast enough for real-time applications.

Coverage note — Qualitative visual query examples (Figure 3) and the candidate recall curve plot (Figure 4) were omitted as standalone knowls because their quantitative and architectural conclusions are fully conveyed by the benchmark comparison, ablation analysis, and method knowls.

References

  1. 1.D. Aiger, A. Araujo, and S. Lynen. Yes, we CANN: Constrained Approximate Nearest Neighbors for Local Feature-Based Visual Localization. In Proc. ICCV, 2023. 2, 3, 5
  2. 2.A. Babenko, A. Slesarev, A. Chigorin, and V. Lempitsky. Neural Codes for Image Retrieval. In Proc. ECCV, 2014. 2
  3. 3.H. Bay, A. Ess, T. Tuytelaars, and L. Van Gool. Speeded-Up Robust Features (SURF). CVIU, 2008. 2
  4. 4.B. Cao, A. Araujo, and J. Sim. Unifying Deep Local and Global Features for Image Search. In Proc. ECCV, 2020. 1, 2, 6
  5. 5.Y. Chen, H. Hu, Y. Luan, H. Sun, S. Changpinyo, A. Ritter, and M.-W. Chang. Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions? In Proc. EMNLP, 2023. 1
  6. 6.Jan De Leeuw. Applications of convex analysis to multidimensional scaling. In J. R. Barra, F. Brodeau, G. Romier, and B. Van Cutsem, editors, Recent Developments in Statistics, pages 133–145. North-Holland Publishing Company, 1977. 4, 5
  7. 7.Laxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee, and Vahab Mirrokni. Muvera: Multi-vector retrieval via fixed dimensional encodings. In Advances in Neural Information Processing Systems, 2023. 3
  8. 8.Christos Faloutsos and King-Ip Lin. Fastmap: A fast algorithm for indexing, data-mining and visualization of traditional and multimedia datasets. In Proceedings of the 1995 ACM SIGMOD international conference on Management of data, pages 163–174, 1995. 5
  9. 9.A. Gordo, J. Almazan, J. Revaud, and D. Larlus. End-to-end Learning of Deep Visual Representations for Image Retrieval. IJCV, 2017. 2
  10. 10.Z. Hu, A. Iscen, C. Sun, Z. Wang, K.-W. Chang, Y. Sun, C. Schmid, D. Ross, and A. Fathi. REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory. In Proc. CVPR, 2023. 1
  11. 11.H. Jegou, F. Perronnin, M. Douze, J. Sanchez, P. Perez, and C. Schmid. Aggregating Local Image Descriptors into Compact Codes. PAMI, 2012. 2
  12. 12.J. Krause, M. Stark, J. Deng, and L. Fei-Fei. 3D Object Representations for Fine-Grained Categorization. In Proc. ICCV Workshops, 2013. 1
  13. 13.S. Lee, H. Seong, S. Lee, and E. Kim. Correlation Verification for Image Retrieval. In Proc. CVPR, 2022. 1, 2, 6
  14. 14.Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang. Deepfashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations. In Proc. CVPR, 2016. 1
  15. 15.D. Lowe. Distinctive Image Features from Scale-Invariant Keypoints. IJCV, 2004. 2
  16. 16.T. Mensink, J. Uijlings, L. Castrejon, A. Goel, F. Cadar, H. Zhou, F. Sha, A. Araujo, and V. Ferrari. Encyclopedic VQA: Visual Questions About Detailed Properties of Fine-Grained Categories. In Proc. ICCV, 2023. 1
  17. 17.T. Ng, V. Balntas, Y. Tian, and K. Mikolajczyk. SOLAR: Second-Order Loss and Attention for Image Retrieval. In Proc. ECCV, 2020. 2
  18. 18.D. Nister and H. Stewenius. Scalable Recognition with a Vocabulary Tree. In Proc. CVPR, 2006. 2
  19. 19.H. Noh, A. Araujo, J. Sim, T. Weyand, and B. Han. Large-Scale Image Retrieval with Attentive Deep Local Features. In Proc. ICCV, 2017. 2
  20. 20.J. Peng, C. Xiao, and Y. Li. RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification. arXiv:2006.12634, 2021. 1
  21. 21.J. Philbin, O. Chum, M. Isard, J. Sivic, and A. Zisserman. Object Retrieval with Large Vocabularies and Fast Spatial Matching. In Proc. CVPR, 2007. 2, 5
  22. 22.J. Philbin, O. Chum, M. Isard, J. Sivic, and A. Zisserman. Lost in Quantization: Improving Particular Object Retrieval in Large Scale Image Databases. In Proc. CVPR, 2008. 5
  23. 23.F. Radenovic, A. Iscen, G. Tolias, Y. Avrithis, and O. Chum. Revisiting Oxford and Paris: Large-Scale Image Retrieval Benchmarking. In Proc. CVPR, 2018. 2, 5
  24. 24.J. Revaud, J. Almazan, R. S. Rezende, and C. R. Souza. Learning With Average Precision: Training Image Retrieval With a Listwise Loss. In Proc. ICCV, October 2019. 2
  25. 25.N. Saeed, H. Nam, M. Haq, and D. Saqib. A Survey on Multidimensional Scaling. ACM Comput. Surv., 2018. 2, 4
  26. 26.S. Shao, K. Chen, A. Karpur, Q. Cui, A. Araujo, and B. Cao. Global Features are All You Need for Image Retrieval and Reranking. 2023. 2, 3, 4, 5, 6, 7
  27. 27.V. d. Silva and J. B. Tenenbaum. Sparse multidimensional scaling using landmark points. Technical Report (Stanford University), 2004. 4
  28. 28.J. Sivic and A. Zisserman. Video Google: A Text Retrieval Approach to Object Matching in Videos. In ICCV, 2003. 2
  29. 29.H. Song, Y. Xiang, S. Jegelka, and S. Savarese. Deep Metric Learning via Lifted Structured Feature Embedding. In Proc. CVPR, 2016. 1
  30. 30.P. Suma, G. Kordopatis-Zilos, A. Iscen, and G. Tolias. AMES: Asymmetric and Memory-Efficient Similarity Estimation for Instance-level Retrieval. In Proc. ECCV, 2024. 1, 2, 6, 8
  31. 31.F. Tan, J. Yuan, and V. Ordonez. Instance-level Image Retrieval using Reranking Transformers. In Proc. ICCV, 2021. 1, 2
  32. 32.M. Teichmann, A. Araujo, M. Zhu, and J. Sim. Detect-to-Retrieve: Efficient Regional Aggregation for Image Search. In CVPR, 2019. 2
  33. 33.G. Tolias, Y. Avrithis, and H. Jegou. Image Search with Selective Match Kernels: Aggregation Across Single and Multiple Images. IJCV, 2015. 2
  34. 34.G. Tolias, T. Jenicek, and O. Chum. Learning and Aggregating Deep Local Descriptors for Instance-Level Recognition. In ECCV, 2020. 2
  35. 35.Jarkko Venna, Jaakko Peltonen, Kristian Nybo, Helena Aidos, and Samuel Kaski. Global versus local methods in nonlinear dimensionality reduction. Neural Networks, 23(1):125–136, 2010. 4
  36. 36.P. Weinzaepfel, T. Lucas, D. Larlus, and Y. Kalantidis. Learning Super-Features for Image Retrieval. In ICLR, 2022. 2, 5, 7
  37. 37.M. Yang, D. He, M. Fan, B. Shi, X. Xue, F. Li, E. Ding, and J. Huang. DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features. In Proc. ICCV, 2021. 2, 6
  38. 38.N.-A. Ypsilantis, K. Chen, B. Cao, M. Lipovsky, P. Dogan-Schonberger, G. Makosa, B. Bluntschli, M. Seyedhosseini, O. Chum, and A. Araujo. Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations. In Proc. ICCV, 2023. 1
  39. 39.N.-A. Ypsilantis, N. Garcia, G. Han, S. Ibrahimi, N. Van Noord, and G. Tolias. The Met Dataset: Instance-level Recognition for Artworks. In Proc. NeurIPS Datasets and Benchmarks Track, 2021. 1

Citation

MLA
Aiger, D., et al. “Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking”. arXiv, 2025, http://arxiv.org/abs/2509.04351v2.
APA
Aiger, D., Cao, B., Chen, K., & Araujo, A. (2025). Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking. arXiv. http://arxiv.org/abs/2509.04351v2
Chicago
Aiger, D., B. Cao, K. Chen, and A. Araujo. 2025. “Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking”. arXiv. http://arxiv.org/abs/2509.04351v2.
Harvard
Aiger, D. et al. (2025) “Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2509.04351v2.
Vancouver
1. Aiger D, Cao B, Chen K, Araujo A (2025) Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking. arXiv

BibTeX

@article{aiger2025global,
  title = {Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking},
  author = {Aiger, Dror and Cao, Bingyi and Chen, Kaifeng and Araujo, Andre},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2509.04351v2},
  eprint = {2509.04351}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/