NetVLAD: CNN Architecture for Weakly Supervised Place Recognition

Relja ArandjelovićPetr GronatAkihiko ToriiTomas PajdlaJosef Sivic

article2015CVPR3,274 citations

Introduces NetVLAD, a differentiable aggregation layer and weakly supervised ranking loss that allow standard convolutional neural networks to be trained end-to-end for visual place recognition and image retrieval.

Listen

The research addresses the challenge of large-scale visual place recognition, in which a query photograph must be matched to its correct geographic location within a database of millions of images despite large changes in viewpoint, illumination, season, and scene content such as moving vehicles or people. Accurate, efficient recognition supports applications including autonomous driving, augmented reality, and archival image geolocation, yet existing hand-engineered descriptors and off-the-shelf convolutional networks still fall short on challenging benchmarks.

The work set out to create a convolutional architecture that can be trained end-to-end directly for place recognition, rather than relying on features optimized for unrelated tasks such as object classification.

The authors replaced the final pooling stage of standard networks with a new, differentiable NetVLAD layer that aggregates mid-level convolutional features in a manner inspired by the Vector of Locally Aggregated Descriptors (VLAD) representation. They trained the resulting networks on large collections of Google Street View Time Machine imagery using a weakly supervised ranking loss that requires only approximate GPS labels. Training and evaluation used geographically disjoint splits of the Pittsburgh and Tokyo datasets containing up to 250,000 database images and thousands of queries.

The trained NetVLAD representations raised recall@1 by roughly 47 percent relative to the best off-the-shelf convolutional descriptors on the Pittsburgh benchmark and established new state-of-the-art results for compact descriptors on the Tokyo 24/7 benchmark, which includes night and sunset queries. The same networks also improved mean average precision by about 20 percent relative on the Oxford 5k image-retrieval task when reduced to 256 dimensions. VLAD-style aggregation consistently outperformed simple max pooling, and training the pooling layer itself produced the largest single gain.

These gains demonstrate that task-specific training on weakly labeled, temporally diverse street-level imagery yields descriptors that are both more discriminative and more robust to real-world appearance change than generic features. The resulting compact vectors enable faster and more reliable retrieval while remaining compatible with existing indexing methods.

The NetVLAD layer and the weakly supervised ranking loss are modular components that can be inserted into other convolutional architectures or applied to additional ranking problems that possess only coarse labels. Further gains are likely from training on more diverse geographic and scene content and from joint optimization with dimensionality-reduction stages.

The primary limitations are that all training data came from urban street-level panoramas, so generalization to non-urban or indoor environments remains untested, and modest overfitting was observed when lower network layers were updated. Results are nevertheless consistent across multiple datasets and architectures, supporting high confidence in the reported improvements for the evaluated conditions.

arXiv: 1511.07247relja/netvlad

No sufficiently relevant recommendations were found.

Cover for NetVLAD: CNN Architecture for Weakly Supervised Place Recognition

Abstract

We tackle the problem of large scale visual place recognition, where the task is to quickly and accurately recognize the location of a given query photograph. We present the following three principal contributions. First, we develop a convolutional neural network (CNN) architecture that is trainable in an end-to-end manner directly for the place recognition task. The main component of this architecture, NetVLAD, is a new generalized VLAD layer, inspired by the "Vector of Locally Aggregated Descriptors" image representation commonly used in image retrieval. The layer is readily pluggable into any CNN architecture and amenable to training via backpropagation. Second, we develop a training procedure, based on a new weakly supervised ranking loss, to learn parameters of the architecture in an end-to-end manner from images depicting the same places over time downloaded from Google Street View Time Machine. Finally, we show that the proposed architecture significantly outperforms non-learnt image representations and off-the-shelf CNN descriptors on two challenging place recognition benchmarks, and improves over current state-of-the-art compact image representations on standard image retrieval benchmarks.

Table of Contents

  • 1. Introduction
  • 1.1. Related work
  • 2. Method overview
  • 3. Deep architecture for place recognition
  • 3.1. NetVLAD: A Generalized VLAD layer (fVLADf_{VLAD})
  • 4. Learning from Time Machine data
  • 5. Experiments
  • 5.1. Datasets and evaluation methodology
  • 5.2. Results and discussion
  • 5.3. Qualitative evaluation
  • 5.4. Image retrieval
  • 6. Conclusions
  • Appendices
  • A. Implementation details
  • B. Google Street View Time Machine datasets
  • C. Additional results and discussions

Knowls

  1. Knowl 1 — NetVLAD Differentiable Pooling Layer

    model/method

    The NetVLAD layer is a differentiable, trainable pooling layer that maps an input array of NN local DD-dimensional feature vectors {xi}i=1N\{\mathbf{x}_i\}_{i=1}^N (extracted as the H×W×DH \times W \times D activation map of a CNN convolutional layer) into a fixed-length descriptor matrix VRK×DV \in \mathbb{R}^{K \times D}, where KK is the number of learned visual clusters.

    The (j,k)(j, k)-th entry of VV, representing the jj-th feature dimension (j{1,,D}j \in \{1, \dots, D\}) for cluster kk (k{1,,K}k \in \{1, \dots, K\}), is computed as:

    V(j,k)=i=1Naˉk(xi)(xi(j)ck(j))V(j, k) = \sum_{i=1}^N \bar{a}_k(\mathbf{x}_i) (x_i(j) - c_k(j))

    where ckRD\mathbf{c}_k \in \mathbb{R}^D is the cluster anchor center, xi(j)x_i(j) and ck(j)c_k(j) are the jj-th dimensions of descriptor xi\mathbf{x}_i and anchor ck\mathbf{c}_k, and aˉk(xi)\bar{a}_k(\mathbf{x}_i) is a soft-assignment weight given by:

    aˉk(xi)=ewkTxi+bkk=1KewkTxi+bk\bar{a}_k(\mathbf{x}_i) = \frac{e^{\mathbf{w}_k^T \mathbf{x}_i + b_k}}{\sum_{k'=1}^K e^{\mathbf{w}_{k'}^T \mathbf{x}_i + b_{k'}}}

    The parameters {wkRD}\{\mathbf{w}_k \in \mathbb{R}^D\} and {bkR}\{b_k \in \mathbb{R}\} are trainable filter weights and biases decoupled from the cluster anchor centers {ck}\{\mathbf{c}_k\}.

    The layer is implemented in standard CNN components as a 1×11 \times 1 convolution with KK filters {wk}\{\mathbf{w}_k\} and biases {bk}\{b_k\}, followed by a channel-wise spatial softmax to compute assignment weights aˉk(xi)\bar{a}_k(\mathbf{x}_i), and a residual accumulation core that evaluates the sum of weighted differences (xick)(\mathbf{x}_i - \mathbf{c}_k). Following accumulation, the matrix VV undergoes intra-normalization (L2 normalization of each DD-dimensional column independently), is serialized into a (K×D)(K \times D)-dimensional vector, and undergoes global L2 normalization.

  2. Knowl 2 — Weakly Supervised Triplet Ranking Loss

    model/method

    When training place recognition models using GPS-tagged database imagery where exact visual correspondences are unknown, supervision is weak: a query image qq has a set of potential positive images {piq}i=1P\{p_i^q\}_{i=1}^{P} (images within a geographical proximity threshold of qq) and a set of definite negative images {njq}j=1M\{n_j^q\}_{j=1}^{M} (images beyond a separation threshold from qq).

    To handle the fact that not all geographically proximate images depict overlapping visual content, the training objective identifies the single best matching potential positive image piqp_{i^*}^q under the current model parameters θ\theta:

    piq=argminpiqdθ(q,piq)p_{i^*}^q = \arg\min_{p_i^q} d_\theta(q, p_i^q)

    where dθ(a,b)=fθ(a)fθ(b)2d_\theta(a, b) = \|f_\theta(a) - f_\theta(b)\|_2 is the Euclidean distance between the L2-normalized image embeddings extracted by the network fθf_\theta.

    The weakly supervised ranking loss LθL_\theta for a training tuple (q,{piq},{njq})(q, \{p_i^q\}, \{n_j^q\}) is defined as:

    Lθ=j=1Mmax(0,minidθ2(q,piq)+mdθ2(q,njq))L_\theta = \sum_{j=1}^M \max\left(0, \, \min_{i} d_\theta^2(q, p_i^q) + m - d_\theta^2(q, n_j^q)\right)

    where m>0m > 0 is a constant margin parameter. The loss incurs a penalty whenever the squared Euclidean distance between the query and a negative image fails to exceed the squared Euclidean distance to the closest positive by at least mm.

  3. Knowl 3 — Cached Hard Negative Mining Algorithm for Weakly Supervised Triplet Ranking

    algorithm

    Training with weakly supervised tuples (q,{piq},{njq})(q, \{p_i^q\}, \{n_j^q\}) involves mining hard negatives from thousands of candidates. Computing full forward passes through the CNN for all negative candidates at every SGD step is computationally prohibitive. To scale training, image representations are precomputed across the dataset and cached, and backpropagation is executed only on the mined hard negatives.

    Input: Training query qq, potential positives {piq}\{p_i^q\}, database pool Dneg\mathcal{D}_{neg} (images >25 m> 25\text{ m} from qq)
    Input: Cache update interval NcacheN_{cache}, margin mm, network parameters θ\theta
    Output: Updated network parameters θ\theta
    Initialize cache of image descriptors C={fθ(I)ID}\mathcal{C} = \{f_\theta(I) \mid I \in \mathcal{D}\}
    for each training iteration step tt do
        if tmodNcache==0t \bmod N_{cache} == 0 then
            Recompute descriptor cache C\mathcal{C} for all database and query images using fθf_\theta
        end if
        
        Compute exact representations fθ(q)f_\theta(q) and {fθ(piq)}\{f_\theta(p_i^q)\} for current query and all potential positives
        Determine best matching positive: piq=argminpiqfθ(q)fθ(piq)2p_{i^*}^q = \arg\min_{p_i^q} \|f_\theta(q) - f_\theta(p_i^q)\|_2
        
        Sample a random candidate set SDneg\mathcal{S} \subset \mathcal{D}_{neg} of size 1000
        Using cached representations from C\mathcal{C}, identify the 10 hardest negatives from S\mathcal{S} and include the 10 hardest negatives from the previous epoch
        
        Execute forward and backward passes on qq, {piq}\{p_i^q\}, and the selected 10 hardest negatives {njq}j=110\{n_j^q\}_{j=1}^{10}
        Compute loss: Lθ=j=110max(0,fθ(q)fθ(piq)22+mfθ(q)fθ(njq)22)L_\theta = \sum_{j=1}^{10} \max(0, \|f_\theta(q) - f_\theta(p_{i^*}^q)\|_2^2 + m - \|f_\theta(q) - f_\theta(n_j^q)\|_2^2)
        Update parameters θ\theta via SGD step
    end for
  4. Knowl 4 — NetVLAD Parameter Initialization and Decoupling

    model/method

    In standard VLAD, the soft-assignment weights aˉk(xi)eαxick2\bar{a}_k(\mathbf{x}_i) \propto e^{-\alpha \|\mathbf{x}_i - \mathbf{c}_k\|^2} couple the cluster center ck\mathbf{c}_k used for soft assignment to the anchor used for residual difference (xick)(\mathbf{x}_i - \mathbf{c}_k). Expanding the squared Euclidean distance yields:

    aˉk(xi)=ewkTxi+bkkewkTxi+bk\bar{a}_k(\mathbf{x}_i) = \frac{e^{\mathbf{w}_k^T \mathbf{x}_i + b_k}}{\sum_{k'} e^{\mathbf{w}_{k'}^T \mathbf{x}_i + b_{k'}}}

    where wk=2αck\mathbf{w}_k = 2\alpha \mathbf{c}_k and bk=αck2b_k = -\alpha \|\mathbf{c}_k\|^2.

    NetVLAD decouples the parameters by treating {wk}\{\mathbf{w}_k\}, {bk}\{b_k\}, and {ck}\{\mathbf{c}_k\} as three independent sets of learnable parameters for each cluster k{1,,K}k \in \{1, \dots, K\}. This decoupling provides additional flexibility: the anchor ck\mathbf{c}_k acts as a local coordinate origin that can be shifted during end-to-end training to minimize residual inner products between non-matching images, while wk\mathbf{w}_k and bkb_k control the partitioning of the descriptor space.

    To initialize NetVLAD to reproduce conventional VLAD prior to gradient descent:

    1. Run kk-means clustering on conv5 local descriptors sampled from the training set to obtain initial cluster centers {ck}k=1K\{\mathbf{c}_k\}_{k=1}^K.
    2. Set wk=2αck\mathbf{w}_k = 2\alpha \mathbf{c}_k and bk=αck2b_k = -\alpha \|\mathbf{c}_k\|^2.
    3. Choose the scalar scaling factor α>0\alpha > 0 to be sufficiently large such that soft assignments are sparse; specifically, α\alpha is set so that the ratio of the largest to the second largest assignment weight aˉk(xi)\bar{a}_k(\mathbf{x}_i) is on average 100 across descriptors.
  5. Knowl 5 — Ablation of Training CNN Layer Depths for Place Recognition

    data/table

    The effect of varying the depth to which backpropagation updates network weights was evaluated using AlexNet on the Pitts30k-val dataset. Training the NetVLAD pooling layer while keeping convolutional layers fixed provides the largest individual performance gain over off-the-shelf baselines, while fine-tuning down to mid-level convolutional layers (conv3) yields the highest recognition accuracy. Backpropagating all the way to early layers (conv1/conv2) results in mild overfitting.

    Lowest trained fmaxf_{\text{max}} fVLADf_{\text{VLAD}}
    layer r@1 r@5 r@10 r@1 r@5 r@10
    none (off-the-shelf) 33.5 57.3 68.4 54.5 69.8 76.1
    NetVLAD 80.5 91.8 95.2
    conv5 63.8 83.8 89.0 84.1 94.6 95.5
    conv4 62.1 83.6 89.2 85.1 94.4 96.1
    conv3 69.8 86.7 90.3 85.5 94.6 96.5
    conv2 69.1 87.6 91.5 84.5 94.6 96.6
    conv1 (full) 68.5 86.2 90.8 84.2 94.7 96.1

    In the table, Lowest trained layer indicates that weights in that layer and all layers above it are learned, while layers below remain fixed to their pretrained states. r@N denotes Recall@N (percentage of queries with at least one correct match within 25 meters among the top NN ranked candidates).

  6. Knowl 6 — Necessity of Multi-Temporal Street View Time Machine Supervision

    data/table

    Training place recognition networks requires multi-temporal imagery depicting the same geographical locations across different times, seasons, and lighting conditions. When trained on images captured at the identical time, the network overfits to transient visual elements (such as parked cars or identical lighting) and fails to generalize to test queries.

    Training data Recall@1 (%) Recall@10 (%)
    Pretrained on ImageNet 33.5 68.5
    Pretrained on Places205 24.8 54.4
    Trained without Time Machine 38.7 68.1
    Trained with Time Machine 68.5 90.8

    Results are reported for Max-pooling (fmaxf_{\text{max}}) with AlexNet evaluated on Pitts30k-val. In the Trained without Time Machine condition, training queries and database images are sampled from imagery captured at the same time. Training with multi-temporal Google Street View Time Machine data improves Recall@1 from 38.7% to 68.5%.

  7. Knowl 7 — Visual Place Recognition Performance on Pitts250k and Tokyo 24/7

    empirical result

    The end-to-end trained NetVLAD architecture (fVLADf_{\text{VLAD}} with K=64K=64) combined with PCA dimensionality reduction, whitening, and L2 normalization achieves state-of-the-art visual place recognition across large-scale benchmarks:

    1. Pittsburgh (Pitts250k-test): End-to-end trained AlexNet with NetVLAD achieves 81.0% Recall@1, outperforming off-the-shelf AlexNet with standard VLAD (55.0% Recall@1) by a 47% relative margin. Trained VGG-16 with NetVLAD and PCA whitening achieves the top performance on Pitts250k.

    2. Tokyo 24/7 (Day/Sunset/Night mobile queries vs. Day Street View database): Under severe day-to-night lighting and viewpoint shifts, VGG-16 NetVLAD with whitening sets the state of the art, outperforming previous hand-engineered methods (RootSIFT + VLAD + whitening and the view-synthesis approach of Torii et al., CVPR 2015).

    3. Compactness vs. Max Pooling: 128-dimensional NetVLAD achieves 42.9% Recall@1 on Tokyo 24/7, exceeding the 38.4% Recall@1 of a 4-times larger 512-dimensional Max-pooling descriptor (fmaxf_{\text{max}}).

  8. Knowl 8 — Performance on Compact Image Retrieval Benchmarks

    data/table

    When trained on Pittsburgh street imagery, VGG-16 with NetVLAD and PCA whitening reduced to 256 dimensions achieves state-of-the-art retrieval accuracy (mean Average Precision, mAP) on standard object and instance retrieval benchmarks (Oxford 5k, Paris 6k, and Holidays).

    Method Oxford 5k Paris 6k Holidays
    full crop full crop orig rot
    Jégou and Zisserman 47.2 65.7 65.7
    Gordo et al. 78.3
    Razavian et al. 53.3 67.0 74.2
    Babenko and Lempitsky 58.9 53.1 80.2
    NetVLAD off-the-shelf 53.4 55.5 64.3 67.7 82.1 86.0
    NetVLAD trained 62.5 63.5 72.0 73.5 79.9 84.3

    full uses the uncropped query image, while crop crops the query to the region of interest (ROI). On Oxford 5k (crop), trained NetVLAD reaches 63.5% mAP (a +20% relative improvement over prior compact 256-D representations). On Holidays, where images contain non-urban scenes (landscapes, underwater, animals), off-the-shelf features outperform the Pittsburgh-trained model (86.0% vs 84.3% on rotated images) due to urban domain specificity.

Coverage note — None was omitted; all key contributions (NetVLAD layer formulation, weakly supervised ranking loss, caching algorithm, parameter initialization/decoupling, depth ablations, Time Machine supervision ablation, place recognition benchmarks, and instance retrieval results) are covered.

References

  1. 1.Project webpage (code/networks). http://www.di.ens.fr/willow/research/netvlad/. 6, 11
  2. 2.R. Arandjelović and A. Zisserman. Three things everyone should know to improve object retrieval. In Proc. CVPR, 2012. 2, 7
  3. 3.R. Arandjelović and A. Zisserman. All about VLAD. In Proc. CVPR, 2013. 1, 2, 3, 4, 7
  4. 4.R. Arandjelović and A. Zisserman. DisLocation: Scalable descriptor distinctiveness for location recognition. In Proc. ACCV, 2014. 1, 2, 6
  5. 5.M. Aubry, B. C. Russell, and J. Sivic. Painting-to-3D model alignment via discriminative visual elements. ACM Transactions on Graphics (TOG), 33(2):14, 2014. 1
  6. 6.H. Azizpour, A. Razavian, J. Sullivan, A. Maki, and S. Carlsson. Factors of transferability from a generic ConvNet representation. CoRR, abs/1406.5774, 2014. 2, 3, 5, 6, 7, 12, 13
  7. 7.A. Babenko and V. Lempitsky. Aggregating local deep features for image retrieval. In Proc. ICCV, 2015. 2, 3, 6, 7, 12, 13, 14
  8. 8.A. Babenko, A. Slesarev, A. Chigorin, and V. Lempitsky. Neural codes for image retrieval. In Proc. ECCV, 2014. 2, 14
  9. 9.S. Cao and N. Snavely. Graph-based discriminative learning for location recognition. In Proc. CVPR, 2013. 1, 2
  10. 10.D. M. Chen, G. Baatz, K. Koeser, S. S. Tsai, R. Vedantham, T. Pylvanainen, K. Roimela, X. Chen, J. Bach, M. Pollefeys, B. Girod, and R. Grzeszczuk. City-scale landmark identification on mobile devices. In Proc. CVPR, 2011. 1, 2, 5, 11
  11. 11.O. Chum, A. Mikulik, M. Perdoch, and J. Matas. Total recall II: Query expansion revisited. In Proc. CVPR, 2011. 2
  12. 12.O. Chum, J. Philbin, J. Sivic, M. Isard, and A. Zisserman. Total recall: Automatic query expansion with a generative feature model for object retrieval. In Proc. ICCV, 2007. 2
  13. 13.M. Cimpoi, S. Maji, and A. Vedaldi. Deep filter banks for texture recognition and segmentation. In Proc. CVPR, 2015. 3, 4
  14. 14.G. Csurka, C. Bray, C. Dance, and L. Fan. Visual categorization with bags of keypoints. In Workshop on Statistical Learning in Computer Vision, ECCV, pages 1–22, 2004. 3
  15. 15.M. Cummins and P. Newman. FAB-MAP: Probabilistic localization and mapping in the space of appearance. The International Journal of Robotics Research, 2008. 1, 2
  16. 16.M. Cummins and P. Newman. Highly scalable appearance-only SLAM - FAB-MAP 2.0. In RSS, 2009. 1, 2
  17. 17.J. Delhumeau, P.-H. Gosselin, H. Jégou, and P. Pérez. Revisiting the VLAD image representation. In Proc. ACMM, 2013. 2
  18. 18.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proc. CVPR, 2009. 6
  19. 19.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. DeCAF: A deep convolutional activation feature for generic visual recognition. CoRR, abs/1310.1531, 2013. 1
  20. 20.J. Foulds and E. Frank. A review of multi-instance learning assumptions. The Knowledge Engineering Review, 25(01):1–25, 2010. 6
  21. 21.R. B. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proc. CVPR, 2014. 1, 13
  22. 22.Y. Gong, L. Wang, R. Guo, and S. Lazebnik. Multi-scale orderless pooling of deep convolutional activation features. In Proc. ECCV, 2014. 2, 3, 4, 6, 7, 12, 13
  23. 23.A. Gordo, J. A. Rodríguez-Serrano, F. Perronnin, and E. Valveny. Leveraging category-level labels for instance-level image retrieval. In Proc. CVPR, pages 3045–3052, 2012. 14
  24. 24.P. Gronat, G. Obozinski, J. Sivic, and T. Pajdla. Learning and calibrating per-location classifiers for visual place recognition. In Proc. CVPR, 2013. 1, 2, 5, 6, 11
  25. 25.H. Jégou and O. Chum. Negative evidences and co-occurrences in image retrieval: the benefit of PCA and whitening. In Proc. ECCV, 2012. 2, 7
  26. 26.H. Jégou, M. Douze, and C. Schmid. Hamming embedding and weak geometric consistency for large scale image search. In Proc. ECCV, pages 304–317, 2008. 2, 8, 14
  27. 27.H. Jégou, M. Douze, and C. Schmid. On the burstiness of visual elements. In Proc. CVPR, Jun 2009. 2
  28. 28.H. Jégou, M. Douze, and C. Schmid. Product quantization for nearest neighbor search. IEEE PAMI, 2011. 1
  29. 29.H. Jégou, M. Douze, C. Schmid, and P. Pérez. Aggregating local descriptors into a compact image representation. In Proc. CVPR, 2010. 1, 2, 3, 7
  30. 30.H. Jégou, H. Harzallah, and C. Schmid. A contextual dissimilarity measure for accurate and efficient image search. In Proc. CVPR, 2007. 2
  31. 31.H. Jégou, F. Perronnin, M. Douze, J. Sánchez, P. Pérez, and C. Schmid. Aggregating local images descriptors into compact codes. IEEE PAMI, 2012. 1
  32. 32.H. Jégou and A. Zisserman. Triangulation embedding and democratic aggregation for image search. In Proc. CVPR, 2014. 2, 14
  33. 33.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proc. CVPR, 2015. 8
  34. 34.A. Kendall, M. Grimes, and R. Cipolla. PoseNet: A convolutional network for real-time 6-DOF camera relocalization. In Proc. ICCV, 2015. 2
  35. 35.J. Knopp, J. Sivic, and T. Pajdla. Avoiding confusing features in place recognition. In Proc. ECCV, 2010. 1, 2, 5, 11
  36. 36.D. Kotzias, M. Denil, P. Blunsom, and N. de Freitas. Deep multi-instance transfer learning. CoRR, abs/1411.3128, 2014. 6
  37. 37.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In NIPS, pages 1106–1114, 2012. 1, 2, 6, 11, 12
  38. 38.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, 1989. 1
  39. 39.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 1
  40. 40.Y. Li, N. Snavely, D. Huttenlocher, and P. Fua. Worldwide pose estimation using 3D point clouds. In Proc. ECCV, 2012. 1
  41. 41.T.-Y. Lin, Y. Cui, S. Belongie, and J. Hays. Learning deep representations for ground-to-aerial geolocalization. In Proc. CVPR, 2015. 2
  42. 42.T.-Y. Lin, A. RoyChowdhury, and S. Maji. Bilinear CNN models for fine-grained visual recognition. In Proc. ICCV, 2015. 4
  43. 43.D. Lowe. Distinctive image features from scale-invariant keypoints. IJCV, 60(2):91–110, 2004. 1, 3, 7
  44. 44.W. Maddern and S. Vidas. Towards robust night and day place recognition using visible and thermal imaging. In Proc. Intl. Conf. on Robotics and Automation, 2014. 1, 2
  45. 45.A. Makadia. Feature tracking for wide-baseline image retrieval. In Proc. ECCV, 2010. 2
  46. 46.C. McManus, W. Churchill, W. Maddern, A. Stewart, and P. Newman. Shady dealings: Robust, long-term visual localisation using illumination invariance. In Proc. Intl. Conf. on Robotics and Automation, 2014. 1, 2
  47. 47.S. Middelberg, T. Sattler, O. Untzelmann, and L. Kobbelt. Scalable 6-DOF localization on mobile devices. In Proc. ECCV, 2014. 1
  48. 48.A. Mikulik, M. Perdoch, O. Chum, and J. Matas. Learning a fine vocabulary. In Proc. ECCV, 2010. 2
  49. 49.M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Learning and transferring mid-level image representations using convolutional neural networks. In Proc. CVPR, 2014. 1, 13
  50. 50.M. Paulin, M. Douze, Z. Harchaoui, J. Mairal, F. Perronnin, and C. Schmid. Local convolutional features with unsupervised training for image retrieval. In Proc. ICCV, 2015. 2
  51. 51.F. Perronnin and C. Dance. Fisher kernels on visual vocabularies for image categorization. In Proc. CVPR, 2007. 2
  52. 52.F. Perronnin, Y. Liu, J. Sanchez, and H. Poirier. Large-scale image retrieval with compressed fisher vectors. In Proc. CVPR, 2010. 1, 2
  53. 53.J. Philbin, O. Chum, M. Isard, J. Sivic, and A. Zisserman. Object retrieval with large vocabularies and fast spatial matching. In Proc. CVPR, 2007. 1, 2, 8, 14
  54. 54.J. Philbin, O. Chum, M. Isard, J. Sivic, and A. Zisserman. Lost in quantization: Improving particular object retrieval in large scale image databases. In Proc. CVPR, 2008. 2, 8, 14
  55. 55.J. Philbin, M. Isard, J. Sivic, and A. Zisserman. Descriptor learning for efficient retrieval. In Proc. ECCV, 2010. 2
  56. 56.D. Qin, X. Chen, M. Guillaumin, and L. V. Gool. Quantized kernel learning for feature matching. In NIPS, 2014. 2
  57. 57.D. Qin, Y. Chen, M. Guillaumin, and L. V. Gool. Learning to rank bag-of-word histograms for large-scale object retrieval. In Proc. BMVC., 2014. 2
  58. 58.D. Qin, S. Gammeter, L. Bossard, T. Quack, and L. Van Gool. Hello neighbor: accurate object retrieval with k-reciprocal nearest neighbors. In Proc. CVPR, 2011. 2
  59. 59.D. Qin, C. Wengert, and L. V. Gool. Query adaptive similarity for large scale object retrieval. In Proc. CVPR, 2013. 2
  60. 60.A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. CNN features off-the-shelf: An astounding baseline for recognition. CoRR, abs/1403.6382, 2014. 2, 5, 6, 7, 12, 13
  61. 61.A. S. Razavian, J. Sullivan, A. Maki, and S. Carlsson. A baseline for visual instance retrieval with deep convolutional networks. CoRR, abs/1412.6574v2, 2014. 14
  62. 62.A. S. Razavian, J. Sullivan, A. Maki, and S. Carlsson. A baseline for visual instance retrieval with deep convolutional networks. In Proc. ICLR, 2015. 2, 3, 5, 6, 7, 12, 13, 14
  63. 63.T. Sattler, M. Havlena, F. Radenovic, K. Schindler, and M. Pollefeys. Hyperpoints and fine vocabularies for large-scale location recognition. In Proc. ICCV, 2015. 1, 2
  64. 64.T. Sattler, B. Leibe, and L. Kobbelt. Fast image-based localization using direct 2D–to–3D matching. In Proc. ICCV, 2011. 1, 2
  65. 65.T. Sattler, T. Weyand, B. Leibe, and L. Kobbelt. Image retrieval for image-based localization revisited. In Proc. BMVC., 2012. 1, 2, 6
  66. 66.G. Schindler, M. Brown, and R. Szeliski. City-scale location recognition. In Proc. CVPR, 2007. 1, 2
  67. 67.F. Schroff, D. Kalenichenko, and J. Philbin. FaceNet: A unified embedding for face recognition and clustering. In Proc. CVPR, 2015. 6
  68. 68.M. Schultz and T. Joachims. Learning a distance metric from relative comparisons. In NIPS, 2004. 6
  69. 69.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. OverFeat: Integrated recognition, localization and detection using convolutional networks. CoRR, abs/1312.6229, 2013. 1
  70. 70.E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, and F. Moreno-Noguer. Fracking deep convolutional image descriptors. CoRR, abs/1412.6537, 2014. 2
  71. 71.K. Simonyan, A. Vedaldi, and A. Zisserman. Descriptor learning using convex optimisation. In Proc. ECCV, 2012. 2
  72. 72.K. Simonyan, A. Vedaldi, and A. Zisserman. Deep Fisher networks for large-scale image classification. In NIPS, 2013. 4
  73. 73.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In Proc. ICLR, 2015. 1, 6, 11
  74. 74.J. Sivic and A. Zisserman. Video Google: A text retrieval approach to object matching in videos. In Proc. ICCV, volume 2, pages 1470–1477, 2003. 1, 3
  75. 75.N. Sunderhauf, S. Shirazi, A. Jacobson, E. Pepperell, F. Dayoub, B. Upcroft, and M. Milford. Place recognition with ConvNet landmarks: Viewpoint-robust, condition-robust, training-free. In Robotics: Science and Systems, 2015. 1, 2
  76. 76.V. Sydorov, M. Sakurada, and C. Lampert. Deep fisher kernels – end to end learning of the fisher kernel GMM parameters. In Proc. CVPR, 2014. 4
  77. 77.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proc. CVPR, 2014. 1
  78. 78.G. Tolias, Y. Avrithis, and H. Jégou. To aggregate or not to aggregate: Selective match kernels for image search. In Proc. ICCV, 2013. 2
  79. 79.G. Tolias and H. Jégou. Visual query expansion with or without geometry: refining local descriptors by feature aggregation. Pattern Recognition, 2014. 2
  80. 80.A. Torii, R. Arandjelović, J. Sivic, M. Okutomi, and T. Pajdla. 24/7 place recognition by view synthesis. In Proc. CVPR, 2015. 1, 2, 6, 7, 11, 12, 15
  81. 81.A. Torii, J. Sivic, T. Pajdla, and M. Okutomi. Visual place recognition with repetitive structures. In Proc. CVPR, 2013. 1, 2, 5, 6, 11, 12
  82. 82.T. Turcot and D. G. Lowe. Better matching with fewer features: The selection of useful features in large database recognition problems. In ICCV Workshop on Emergent Issues in Large Amounts of Visual Data (WS-LAVD), 2009. 2
  83. 83.T. Tuytelaars and K. Mikolajczyk. Local invariant feature detectors: A survey. Foundations and Trends® in Computer Graphics and Vision, 3(3):177–280, 2008. 1
  84. 84.A. Vedaldi and K. Lenc. Matconvnet – convolutional neural networks for matlab. 2015. 11
  85. 85.P. Viola, J. C. Platt, and C. Zhang. Multiple instance boosting for object detection. In NIPS, 2005. 6
  86. 86.J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu. Learning fine-grained image similarity with deep ranking. In Proc. CVPR, 2014. 6
  87. 87.K. Q. Weinberger, J. Blitzer, and L. Saul. Distance metric learning for large margin nearest neighbor classification. In NIPS, 2006. 6
  88. 88.S. Winder, G. Hua, and M. Brown. Picking the best DAISY. In Proc. CVPR, pages 178–185, 2009. 2
  89. 89.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. In Proc. ECCV, 2014. 1, 8, 13
  90. 90.J. Zepeda and P. Pérez. Exemplar SVMs as visual feature encoders. In Proc. CVPR, 2015. 2
  91. 91.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva. Learning deep features for scene recognition using places database. In NIPS, 2014. 1, 6, 12

Citation

MLA
Arandjelović, R., et al. “NetVLAD: CNN Architecture for Weakly Supervised Place Recognition”. arXiv, 2015, http://arxiv.org/abs/1511.07247v3.
APA
Arandjelović, R., Gronat, P., Torii, A., Pajdla, T., & Sivic, J. (2015). NetVLAD: CNN architecture for weakly supervised place recognition. arXiv. http://arxiv.org/abs/1511.07247v3
Chicago
Arandjelović, R., P. Gronat, A. Torii, T. Pajdla, and J. Sivic. 2015. “NetVLAD: CNN Architecture for Weakly Supervised Place Recognition”. arXiv. http://arxiv.org/abs/1511.07247v3.
Harvard
Arandjelović, R. et al. (2015) “NetVLAD: CNN architecture for weakly supervised place recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1511.07247v3.
Vancouver
1. Arandjelović R, Gronat P, Torii A, Pajdla T, Sivic J (2015) NetVLAD: CNN architecture for weakly supervised place recognition. arXiv

BibTeX

@article{arandjelovic2015netvlad,
  title = {NetVLAD: CNN architecture for weakly supervised place recognition},
  author = {Arandjelović, Relja and Gronat, Petr and Torii, Akihiko and Pajdla, Tomas and Sivic, Josef},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1511.07247v3},
  eprint = {1511.07247}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF