SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks

Bo LiWei WuQiang WangFangyi ZhangJunliang XingJunjie Yan

article2018CVPR2,264 citations

Proposes a spatial-aware sampling strategy and multi-layer feature aggregation to overcome translation invariance limitations, successfully enabling deep ResNet backbones in Siamese visual tracking to achieve state-of-the-art accuracy across major benchmarks.

Listen

Visual object tracking is a foundational computer vision capability with applications spanning autonomous navigation, surveillance, robotics, and augmented reality. Tracking algorithms must operate reliably despite real-world complications like lighting shifts, severe occlusions, and background clutter. Siamese neural networks have emerged as popular solutions due to their high processing speed; however, they have historically suffered from an accuracy gap compared to leading tracking systems. Although other vision disciplines have advanced rapidly by adopting deep network architectures such as ResNet, Siamese trackers remained restricted to older, shallow architectures because deep networks consistently degraded tracking accuracy.

The article set out to identify the root cause of this performance degradation and demonstrate how to successfully train very deep Siamese networks for high-precision, real-time visual tracking.

The authors conducted theoretical analyses and experimental investigations across large-scale video datasets, including COCO, ImageNet, and YouTube-BoundingBoxes. They discovered that the zero-padding used in modern deep networks disrupts strict translation invariance, causing trackers to develop an artificial positional bias toward the image center. To overcome this limitation, the authors developed a spatial-aware sampling strategy that trains deep networks with random translational shifts. They also designed an advanced tracker, termed SiamRPN++, which incorporates multi-layer feature aggregation to combine fine spatial details with high-level semantics, along with a lightweight depth-wise cross-correlation mechanism that reduces parameters by a factor of ten.

Rigorous evaluations demonstrated that SiamRPN++ establishes new state-of-the-art results across major tracking benchmarks (OTB2015, VOT2018, UAV123, LaSOT, and TrackingNet). On the challenging VOT2018 benchmark, the model surpassed the challenge winner by 6.4% in overall overlap performance while operating at real-time speeds of 35 frames per second. On the large-scale TrackingNet dataset, it outperformed previous leading Siamese models by 9.5% in success rate and 10.3% in precision. Ablation experiments confirmed that multi-layer aggregation provided a 4.0% performance boost over single-layer baselines, and lightweight variants utilizing MobileNet achieved competitive accuracy at speeds exceeding 70 frames per second.

These findings resolve a longstanding barrier in visual tracking by proving that deep, off-the-shelf neural network architectures can be effectively utilized without performance collapse. By balancing high accuracy with real-time operational efficiency, this framework reduces compute trade-offs for production systems, allowing high-performance computer vision pipelines to operate in latency-sensitive, resource-constrained environments.

Organizations developing computer vision applications should adopt spatial-aware shift augmentations when fine-tuning deep backbones and replace standard correlation layers with depth-wise alternatives to improve model stability and efficiency. Depending on system constraints, deployment teams can choose between the primary ResNet-50 backbone for maximum accuracy (35 frames per second) or the MobileNet variant for high-throughput mobile applications (70 frames per second).

The reported findings carry high confidence across diverse short-term and long-term benchmarks. However, the authors note a minor limitation: because Siamese trackers evaluate instances without continuous online model updating, their robustness remains slightly lower than computationally heavier correlation-filter methods in select edge cases. Further research into efficient online model adaptation is recommended to close this remaining robustness gap.

  • Paper: Exploring Simple Siamese Representation Learning, Xinlei Chen et al. (2021). This paper builds on Siamese representation learning by exploring simpler optimization strategies without negative pairs, continuing the trajectory of feature design for Siamese architectures.
Cover for SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks

Abstract

Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Siamese Tracking with Very Deep Networks
  • 3.1 Analysis on Siamese Networks for Tracking
  • 3.2 ResNet-driven Siamese Tracking
  • 3.3 Layer-wise Aggregation
  • 3.4 Depthwise Cross Correlation
  • 4 Experimental Results
  • 4.1 Training Dataset and Evaluation
  • 4.2 Implementation Details
  • 4.3 Ablation Experiments
  • 4.4 Comparison with the state-of-the-art
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Spatial-Aware Sampling Strategy for Deep Siamese Tracking

    model/method

    Siamese visual tracking traditionally relies on strict translation invariance f(z,x[Δτ])=f(z,x)[Δτ]f(z, x[\Delta \tau]) = f(z, x)[\Delta \tau] between the exemplar template zz and the search patch xx. When modern deep backbones (such as ResNet-50 or MobileNet) with zero-padding are used, this strict translation invariance is broken, causing the network to learn a strong spatial center bias where locations near the image borders receive near-zero tracking probability regardless of visual content.

    To eliminate this spatial bias and allow end-to-end training of deep architectures, a spatial-aware sampling strategy applies random translation shifts during data augmentation. Specifically, the target center is shifted by an offset sampled uniformly within a maximum range of ±64\pm 64 pixels. This prevents the model from collapsing into center-biased trivial solutions, uniformly distributing positive sample priors across the search region and restoring tracking performance (e.g., boosting the Expected Average Overlap (EAO) on VOT2018 from 0.14 with zero shift to over 0.40 with suitable shift).

  2. Knowl 2 — Depth-wise Cross-Correlation Layer

    model/method

    In Siamese tracking, cross-correlation embeds information from the target template branch zz and search area branch xx. The depth-wise cross-correlation (DW-XCorr) module computes channel-by-channel correlation between the template and search features, significantly reducing parameters and computational cost compared to Up-Channel cross-correlation (UP-XCorr).

    Given feature maps from the backbone, the module operates as follows:

    1. Feature adjustment: The template feature map Fl(z)F_l(z) and search feature map Fl(x)F_l(x) from layer ll pass through separate, non-shared convolutional-batch normalization (conv-bn) blocks to account for asymmetry in classification and bounding box regression.
    2. Channel-wise cross-correlation: The two adjusted feature maps, both possessing identical channel dimension (e.g., 256 channels), undergo depth-wise cross-correlation where each channel of the template acts as a convolutional filter applied exclusively to the corresponding channel of the search feature.
    3. Feature fusion and head output: The resulting correlation map is passed through a conv-bn-relu block to mix channel information, followed by a final 1×11 \times 1 convolutional layer to generate the classification scores or bounding box regression offsets.

    DW-XCorr contains approximately 10 times fewer parameters than UP-XCorr, balances parameter counts between the template and search branches, stabilizes training convergence, and produces channel features that exhibit semantic specialization.

  3. Knowl 3 — Multi-Layer Feature Aggregation in SiamRPN++

    equation

    To leverage visual representations across multiple abstraction levels, SiamRPN++ aggregates predictions from multiple residual blocks (specifically, conv3, conv4, and conv5 of ResNet-50, denoted as stages l∈{3,4,5}l \in \{3, 4, 5\}). Each layer branch features an independent Siamese Region Proposal Network (RPN) block providing classification score map SlS_l and bounding box regression map BlB_l. Because the outputs from all three RPN blocks share the same spatial resolution, they are combined via weighted fusion:

    Sall=∑l=35αl∗SlS_{\text{all}} = \sum_{l=3}^{5} \alpha_l * S_l

    Ball=∑l=35βl∗BlB_{\text{all}} = \sum_{l=3}^{5} \beta_l * B_l

    where:

    • SlS_l is the classification score map output from the ll-th RPN stage.
    • BlB_l is the bounding box regression map output from the ll-th RPN stage.
    • αl∈R\alpha_l \in \mathbb{R} is the learned fusion weight for classification at layer ll.
    • βl∈R\beta_l \in \mathbb{R} is the learned fusion weight for bounding box regression at layer ll.
    • SallS_{\text{all}} and BallB_{\text{all}} are the aggregated classification and bounding box prediction maps.

    The weights {αl}\{\alpha_l\} and {βl}\{\beta_l\} are separated for classification and regression and optimized end-to-end via gradient descent along with the entire network.

  4. Knowl 4 — ResNet Backbone Adaptation for Siamese Tracking

    model/method

    Standard classification backbones like ResNet-50 have a total stride of 32 pixels, which produces spatial feature maps that are too coarse for dense Siamese tracking. SiamRPN++ modifies ResNet-50 as follows:

    • Stride reduction and dilation: The spatial stride of the conv4 and conv5 residual blocks is reduced from 2 to 1 (retaining an effective total stride of 8 pixels from input to output), and dilated convolutions with increased receptive field are applied to compensate for the reduced stride.
    • Channel projection: A 1×11 \times 1 convolutional layer is appended to the output of each selected block (conv3, conv4, conv5) to project channel dimensions to 256.
    • Template cropping: With zero padding preserved, the template branch feature map size expands to 15×1515 \times 15; the center 7×77 \times 7 region is cropped to serve as the template feature map, ensuring each cell covers the target while reducing correlation cost.
    • Differential learning rates: During end-to-end training, the ResNet backbone feature extractor is trained with a learning rate 10 times smaller than that of the RPN tracking heads, enabling effective representation fine-tuning without destabilizing the tracking objectives.
  5. Knowl 5 — SiamRPN++ Training and Optimization Setup

    experimental setup

    SiamRPN++ is trained end-to-end on four large-scale tracking and detection datasets:

    • Training datasets: COCO, ImageNet DET, ImageNet VID, and YouTube-BoundingBoxes.
    • Input dimensions: Exemplar template patches are cropped to 127×127×3127 \times 127 \times 3 pixels; search region patches are cropped to 255×255×3255 \times 255 \times 3 pixels.
    • Anchor configuration: The Siamese RPN heads use 5 anchor boxes with distinct aspect ratios at each spatial position.
    • Optimization: Synchronized Stochastic Gradient Descent (SGD) across 8 GPUs with a minibatch size of 128 pairs (16 pairs per GPU) for a total of 20 epochs (taking ~12 hours to converge).
    • Learning rate schedule: A 5-epoch warmup stage at learning rate 0.001 trains only the RPN branches while keeping the backbone frozen. For the subsequent 15 epochs, the entire network is trained end-to-end with the learning rate decaying exponentially from 0.005 to 0.0005, with the backbone learning rate scaled to 0.1 of the RPN learning rate. Momentum is set to 0.9 and weight decay to 0.0005.
    • Loss function: The overall training loss is L=Lcls+Lreg\mathcal{L} = \mathcal{L}_{\text{cls}} + \mathcal{L}_{\text{reg}}, combining cross-entropy classification loss Lcls\mathcal{L}_{\text{cls}} and smooth L1L_1 bounding box regression loss Lreg\mathcal{L}_{\text{reg}}.
  6. Knowl 6 — Component Ablations on VOT2018 and OTB2015

    data/table

    Ablation experiments evaluated the effect of backbone choice (AlexNet vs. ResNet-50), backbone fine-tuning, correlation type (Up-Channel UP-XCorr vs. Depth-wise DW-XCorr), and multi-layer aggregation (combinations of conv3 (L3), conv4 (L4), and conv5 (L5)) on VOT2018 (Expected Average Overlap, EAO) and OTB2015 (Area Under Curve, AUC).

    Backbone L3 L4 L5 Finetune Corr VOT2018 (EAO) OTB2015 (AUC)
    AlexNet UP 0.332 0.658
    AlexNet DW 0.355 0.666
    ResNet-50 ✓ ✓ ✓ UP 0.371 0.664
    ResNet-50 ✓ ✓ ✓ ✓ UP 0.390 0.684
    ResNet-50 ✓ ✓ DW 0.331 0.669
    ResNet-50 ✓ ✓ DW 0.374 0.678
    ResNet-50 ✓ ✓ DW 0.320 0.646
    ResNet-50 ✓ ✓ ✓ DW 0.346 0.677
    ResNet-50 ✓ ✓ ✓ DW 0.336 0.674
    ResNet-50 ✓ ✓ ✓ DW 0.383 0.683
    ResNet-50 ✓ ✓ ✓ DW 0.395 0.673
    ResNet-50 ✓ ✓ ✓ ✓ DW 0.414 0.696

    The ablation shows that:

    1. Backbone fine-tuning provides substantial gains (e.g., from 0.395 to 0.414 EAO for 3-layer DW ResNet-50).
    2. Depth-wise correlation outperforms up-channel correlation on ResNet-50 by 2.4% EAO (0.414 vs. 0.390) while requiring 10x fewer parameters.
    3. Layer-wise aggregation of all three layers (L3+L4+L5) achieves an EAO of 0.414, which is 4.0% higher than the best single-layer baseline (L4 alone at 0.374).
  7. Knowl 7 — Tracking Benchmark Comparison on VOT2018

    data/table

    SiamRPN++ was benchmarked against the top state-of-the-art visual trackers on the VOT2018 dataset (60 challenging sequences) using Expected Average Overlap (EAO, ↑\uparrow), Accuracy (A, ↑\uparrow), Robustness / failure rate (R, ↓\downarrow), and Average Overlap under no-reset evaluation (AO, ↑\uparrow).

    Tracker EAO ↑\uparrow Accuracy ↑\uparrow Robustness ↓\downarrow AO ↑\uparrow
    DLSTpp 0.325 0.543 0.224 0.495
    DaSiamRPN 0.326 0.569 0.337 0.398
    SA_Siam_R 0.337 0.566 0.258 0.429
    CPT 0.339 0.506 0.239 0.379
    DeepSTRCF 0.345 0.523 0.215 0.436
    DRT 0.356 0.519 0.201 0.426
    RCO 0.376 0.507 0.155 0.384
    UPDT 0.378 0.536 0.184 0.454
    SiamRPN 0.383 0.586 0.276 0.472
    MFT 0.385 0.505 0.140 0.393
    LADCF 0.389 0.503 0.159 0.421
    SiamRPN++ (Ours) 0.414 0.600 0.234 0.498

    SiamRPN++ achieves the top overall EAO (0.414), representing a 6.4% relative improvement over the VOT2018 competition winner LADCF (0.389) and a 9.5% absolute gain in accuracy over MFT (0.600 vs. 0.505). Against the baseline DaSiamRPN, SiamRPN++ reduces the failure rate (robustness) by 10.3% (0.234 vs. 0.337) while running in real-time at 35 FPS on an NVIDIA Titan Xp GPU. Lightweight variants (using ResNet-18 or MobileNetV2) operate at over 70 FPS with competitive accuracy.

  8. Knowl 8 — Tracking Performance on TrackingNet, LaSOT, UAV123, and OTB2015

    empirical result

    SiamRPN++ was evaluated across several large-scale and specialized tracking benchmarks:

    • TrackingNet (511 test videos): SiamRPN++ achieves state-of-the-art performance across all three metrics: Success AUC of 73.3%73.3\%, Precision P=69.4%P = 69.4\%, and Normalized Precision Pnorm=80.0%P_{\text{norm}} = 80.0\%, outperforming DaSiamRPN (AUC 63.8%63.8\%, P=59.1%P = 59.1\%, Pnorm=73.4%P_{\text{norm}} = 73.4\%), MDNet (AUC 60.6%60.6\%), CFNet (57.8%57.8\%), SiamFC (57.1%57.1\%), ECO (55.4%55.4\%), and CSRDCF (53.4%53.4\%).
    • LaSOT (280 test videos): SiamRPN++ achieves an AUC score of 49.6%49.6\%, improving normalized distance precision and AUC relatively by 23.7%23.7\% and 24.9%24.9\% over MDNet.
    • UAV123 (123 aerial sequences): SiamRPN++ achieves an overall success score of 0.6130.613, substantially outperforming DaSiamRPN (0.5860.586), ECO (0.5250.525), and SiamRPN.
    • OTB2015 (100 video sequences): SiamRPN++ achieves leading overlap success rate and precision, surpassing DaSiamRPN by 3.8%3.8\% in overlap AUC and 3.4%3.4\% in precision score.
  9. Knowl 9 — Long-Term Visual Tracking Performance on VOT2018-LT

    empirical result

    On the VOT2018 Long-Term (VOT2018-LT) benchmark—consisting of 35 long sequences where targets frequently leave the field of view or undergo prolonged full occlusion—SiamRPN++ equipped with a long-term tracking strategy achieves an overall FF-score of 0.6290.629.

    This performance surpasses the VOT2018-LT challenge winner MBMD (F=0.610F = 0.610, 1.9% higher) and DaSiam_LT (F=0.607F = 0.607, 2.2% higher). Furthermore, the long-term version of SiamRPN++ runs at 21 FPS, which is roughly 8 times faster than MBMD.

  10. Knowl 10 — Semantic Channel Orthogonality in Depth-wise Cross-Correlation

    empirical result

    In depth-wise cross-correlation (DW-XCorr), individual channels of the correlation output exhibit orthogonal semantic selectivity:

    • Feature activations for objects of the same category cluster in the same subset of correlation channels while the remaining channels are strongly suppressed.
    • For example, in the conv4 DW-XCorr feature representation (256 total channels), the 148th channel consistently produces high activation responses on cars with low responses on persons/faces; the 222nd channel responds strongly to persons; and the 226th channel responds strongly to human faces.
    • In contrast, up-channel cross-correlation (UP-XCorr) generates response maps lacking clear channel-level semantic interpretability.

Coverage note — None omitted; all primary theoretical insights, model architectural designs, sampling strategies, optimization protocols, ablation studies, and benchmark results are represented.

References

  1. 1.L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr. Fully-convolutional siamese networks for object tracking. In ECCV Workshops, 2016. 1, 2, 3, 5, 8
  2. 2.G. Bhat, J. Johnander, M. Danelljan, F. Shahbaz Khan, and M. Felsberg. Unveiling the power of deep tracking. In ECCV, September 2018. 7
  3. 3.D. Bolme, J. Beveridge, B. Draper, and Y. Lui. Visual object tracking using adaptive correlation filters. In CVPR, 2010. 2
  4. 4.L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 2
  5. 5.M. Danelljan, G. Bhat, F. Shahbaz Khan, and M. Felsberg. Eco: Efficient convolution operators for tracking. In CVPR, 2017. 1, 2, 7, 8
  6. 6.M. Danelljan, G. Hager, F. S. Khan, and M. Felsberg. Learning spatially regularized correlation filters for visual tracking. In ICCV, 2015. 2
  7. 7.M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg. Convolutional features for correlation filter based visual tracking. In ICCV Workshops, 2015. 2
  8. 8.M. Danelljan, F. S. Khan, M. Felsberg, and J. V. De Weijer. Adaptive color attributes for real-time visual tracking. In CVPR, 2014. 2
  9. 9.M. Danelljan, A. Robinson, F. S. Khan, and M. Felsberg. Beyond correlation filters: Learning continuous convolution operators for visual tracking. In ECCV, 2016. 2, 7
  10. 10.H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y. Xu, C. Liao, and H. Ling. Lasot: A high-quality benchmark for large-scale single object tracking, 2018. 2, 6, 8
  11. 11.Q. Guo, W. Feng, C. Zhou, R. Huang, L. Wan, and S. Wang. Learning dynamic siamese network for visual object tracking. In ICCV, 2017. 1
  12. 12.Q. Guo, W. Feng, C. Zhou, R. Huang, L. Wan, and S. Wang. Learning dynamic siamese network for visual object tracking. In ICCV, 2017. 2
  13. 13.K. He, G. Gkioxari, P. Dollar, and R. Girshick. Mask r-cnn. In The IEEE International Conference on Computer Vision (ICCV), Oct 2017. 6
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 2, 4, 6
  15. 15.D. Held, S. Thrun, and S. Savarese. Learning to track at 100 fps with deep regression networks. In ECCV, 2016. 1, 2
  16. 16.J. Henriques, R. Caseiro, P. Martins, and J. Batista. High-speed tracking with kernelized correlation filters. TPAMI, 2015. 2
  17. 17.Z. Hong, Z. Chen, C. Wang, X. Mei, D. Prokhorov, and D. Tao. Multi-store tracker (muster): A cognitive psychology inspired approach to object tracking. In CVPR, 2015. 2
  18. 18.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
  19. 19.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, and L. Cehovin Zajc. The visual object tracking vot2016 challenge results. In ECCV Workshops, 2015. 2
  20. 20.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, and L. Cehovin Zajc. The visual object tracking vot2017 challenge results. In ICCV, 2017. 2
  21. 21.M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pfugfelder, L. C. Zajc, T. Vojir, G. Bhat, A. Lukezic, A. Eldesokey, G. Fernandez, and et al. The sixth visual object tracking vot2018 challenge results. In ECCV Workshops, 2018. 2, 6, 7, 8
  22. 22.M. Kristan, J. Matas, A. Leonardis, M. Felsberg, L. ˇ Cehovin, and G. Fern´ . The visual object tracking vot2015 challenge results. In ICCV Workshops, 2015. 2
  23. 23.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 1, 2
  24. 24.B. Li, J. Yan, W. Wu, Z. Zhu, and X. Hu. High performance visual tracking with siamese region proposal network. In CVPR, 2018. 1, 2, 3, 4, 5, 8
  25. 25.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 6
  26. 26.L. Liu, J. Xing, H. Ai, and X. Ruan. Hand posture recognition using finger geometric feature. In ICIP, 2012. 1
  27. 27.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 4, 6
  28. 28.A. Lukezic, T. Vojir, L. Cehovin Zajc, J. Matas, and M. Kristan. Discriminative correlation filter with channel and spatial reliability. In CVPR, 2017. 8
  29. 29.M. Mueller, N. Smith, and B. Ghanem. A benchmark and simulator for uav tracking. In ECCV, pages 445–461. Springer, 2016. 8
  30. 30.M. Muller, A. Bibi, S. Giancola, S. Al-Subaihi, and ¨ B. Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. ECCV, 2018. 2, 6, 8
  31. 31.M. Muller, N. Smith, and B. Ghanem. A benchmark and ¨ simulator for uav tracking. In ECCV, 2016. 2, 6
  32. 32.H. Nam and B. Han. Learning multi-domain convolutional neural networks for visual tracking. In CVPR, 2016. 2, 8
  33. 33.C. Peng, T. Xiao, Z. Li, Y. Jiang, X. Zhang, K. Jia, G. Yu, and J. Sun. Megdet: A large mini-batch object detector. In CVPR, 2018. 2
  34. 34.R. Pflugfelder. An in-depth analysis of visual tracking with siamese neural networks. arXiv:1707.00569, 2017. 3
  35. 35.E. Real, J. Shlens, S. Mazzocchi, X. Pan, and V. Vanhoucke. Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. In Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, pages 7464–7473. IEEE, 2017. 6
  36. 36.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. IJCV, 2015. 6
  37. 37.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 2
  38. 38.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015. 2
  39. 39.W. Tang, P. Yu, and Y. Wu. Deeply learned compositional models for human pose estimation. In ECCV, 2018. 2
  40. 40.R. Tao, E. Gavves, and A. W. M. Smeulders. Siamese instance search for tracking. In CVPR, 2016. 1, 2, 3
  41. 41.J. Valmadre, L. Bertinetto, J. F. Henriques, A. Vedaldi, and P. H. Torr. End-to-end representation learning for correlation filter based tracking. In CVPR, 2017. 1, 2, 3, 4, 8
  42. 42.Q. Wang, J. Gao, J. Xing, M. Zhang, and W. Hu. Dcfnet: Discriminant correlation filters network for visual tracking. In arXiv:1704.04057, 2017. 1, 2, 3
  43. 43.Q. Wang, Z. Teng, J. Xing, J. Gao, W. Hu, and S. Maybank. Learning attentions: Residual attentional siamese network for high performance online visual tracking. In CVPR, 2018. 1, 2
  44. 44.Q. Wang, M. Zhang, J. Xing, J. Gao, W. Hu, and S. Maybank. Do not lose the details: Reinforced representation learning for high performance visual tracking. In IJCAI, 2018. 1, 2
  45. 45.Y. Wu, J. Lim, and M.-H. Yang. Online object tracking: A benchmark. In CVPR, 2013. 2
  46. 46.Y. Wu, J. Lim, and M.-H. Yang. Object tracking benchmark. TPAMI, 2015. 1, 2, 5, 6, 7
  47. 47.J. Xing, H. Ai, and S. Lao. Multiple human tracking based on multi-view upper-body detection and discriminative learning. In ICPR, 2010. 1
  48. 48.G. Zhang and P. Vela. Good features to track for visual slam. In CVPR, 2015. 1
  49. 49.M. Zhang, Q. Wang, J. Xing, J. Gao, P. Peng, W. Hu, and S. Maybank. Visual tracking via spatially aligned correlation filters network. In ECCV, 2016. 2
  50. 50.M. Zhang, J. Xing, J. Gao, and W. Hu. Robust visual tracking using joint scale-spatial correlation filters. In ICIP, 2015. 2
  51. 51.M. Zhang, J. Xing, J. Gao, X. Shi, Q. Wang, and W. Hu. Joint scale-spatial correlation tracking with adaptive rotation estimation. In ICCV Workshops, 2015. 2
  52. 52.Z. Zhu, Q. Wang, B. Li, W. Wu, J. Yan, and W. Hu. Distractor-aware siamese networks for visual object tracking. In ECCV, 2018. 1, 2, 3, 6, 7, 8

Citation

MLA
Li, B., et al. “SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks”. arXiv, 2018, http://arxiv.org/abs/1812.11703v1.
APA
Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., & Yan, J. (2018). SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks. arXiv. http://arxiv.org/abs/1812.11703v1
Chicago
Li, B., W. Wu, Q. Wang, F. Zhang, J. Xing, and J. Yan. 2018. “SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks”. arXiv. http://arxiv.org/abs/1812.11703v1.
Harvard
Li, B. et al. (2018) “SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.11703v1.
Vancouver
1. Li B, Wu W, Wang Q, Zhang F, Xing J, Yan J (2018) SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks. arXiv

BibTeX

@article{li2018siamrpn,
  title = {SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks},
  author = {Li, Bo and Wu, Wei and Wang, Qiang and Zhang, Fangyi and Xing, Junliang and Yan, Junjie},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.11703v1},
  eprint = {1812.11703}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE