Transformer Tracking with Cyclic Shifting Window Attention

Zikai SongJunqing YuYi-Ping Phoebe ChenWei Yang

article2022CVPR239 citations

Proposes a multi-scale cyclic shifting window attention transformer that shifts visual tracking from pixel-level to window-level matching, preserving object integrity while reducing computation and achieving state-of-the-art accuracy across five major tracking benchmarks.

Listen

Visual object tracking is a foundational computer vision capability essential for autonomous navigation, surveillance, and automated monitoring systems. Standard tracking architectures based on neural attention mechanisms typically compute pixel-by-pixel relationships across flattened image features. This conventional setup discards critical structural boundaries and spatial positioning, leading to degraded tracking accuracy when targets face partial occlusion or move near similar visual distractors.

The article develops and evaluates CSWinTT, a transformer tracking architecture that elevates cross-attention from the individual pixel level to multi-scale image windows. Its primary goal is to preserve object integrity and precise spatial geometry by matching structured visual patches rather than disordered individual pixels.

To achieve this without degrading resolution or generating excessive computational overhead, the approach splits extracted image features into multi-scale windows and applies a cyclic shifting mechanism that systematically translates patches across multiple directions. A spatially regularized masking filter penalizes severe boundary distortions, while three targeted computational optimizations—removing query translations, halving redundant shifting periods, and applying coordinate-based matrix indexing—streamline processing. The tracker was trained end-to-end on major tracking datasets using a standard ResNet-50 backbone and evaluated against state-of-the-art methods across five challenging public benchmarks comprising thousands of video sequences.

CSWinTT achieved top-ranked performance across all five benchmarks, establishing new state-of-the-art tracking accuracy. On the UAV123 benchmark, the architecture achieved a 70.5% Area Under the Curve (AUC) and 90.3% precision, outperforming leading models such as STARK by 1.3% in AUC and 2.1% in precision. On the large-scale LaSOT and TrackingNet datasets, it attained leading AUC scores of 66.2% and 81.9%, respectively. In ablation testing, adding cyclic shifts to basic window attention improved tracking AUC by 15.3 percentage points, demonstrating that positional sample expansion is critical. Furthermore, the efficiency optimizations improved execution throughput from an unusable 1.0 frame per second up to 12.4 frames per second on a single GPU.

These findings demonstrate that preserving structural boundaries and local context during cross-image feature matching substantially improves tracking reliability in complex operational environments, such as scenes with dense background clutter and intermittent occlusions. For operational deployments, the enhanced accuracy directly lowers the risk of target loss, though the speed of approximately 12 frames per second represents a trade-off that may require dedicated hardware acceleration for strictly real-time embedded systems.

Organizations developing automated tracking workflows should consider adopting multi-scale window-level attention matching when tracking robustness under clutter is paramount. Next steps should focus on adapting the core multi-scale cyclic shifting mechanism to broader computer vision domains, including object recognition and stereo matching, while exploring further model compression to achieve higher frame rates on edge devices. Confidence in these conclusions is high given the consistent top-tier results across diverse standard benchmarks; however, practitioners should note that the current architecture outputs bounding boxes rather than pixel-level segmentation masks and requires adequate GPU capacity to achieve practical operational speeds.

arXiv: 2205.03806
Cover for Transformer Tracking with Cyclic Shifting Window Attention

Abstract

Transformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel-to-pixel attention strategy on flattened image features and unavoidably ignore the integrity of objects. In this paper, we propose a new transformer architecture with multi-scale cyclic shifting window attention for visual object tracking, elevating the attention from pixel to window level. The cross-window multi-scale attention has the advantage of aggregating attention at different scales and generates the best fine-scale match for the target object. Furthermore, the cyclic shifting strategy brings greater accuracy by expanding the window samples with positional information, and at the same time saves huge amounts of computational power by removing redundant calculations. Extensive experiments demonstrate the superior performance of our method, which also sets the new state-of-the-art records on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-10k benchmarks. Our project is available at https://github.com/SkyeSong38/CSWinTT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Multi-Scale Cyclic Shifting Window Attention
  • 3.2. Efficient Computation
  • 3.3. Tracking with Window Transformer
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. State-of-the-art Comparison
  • 4.3. Ablation Study
  • 4.4. Qualitative Analysis
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — CSWinTT window-level transformer tracker

    model/method

    CSWinTT tracks an object by matching a template patch from the initial frame with a search region from the current frame. A ResNet-50 backbone followed by a bottleneck layer produces feature maps fz∈Rd×Hz×Wzf_z \in \mathbb{R}^{d \times H_z \times W_z} and fs∈Rd×Hs×Wsf_s \in \mathbb{R}^{d \times H_s \times W_s} for the template and search region, respectively.

    For transformer head ii, a feature map is partitioned into non-overlapping windows of side length rir_i, with did_i channels per head. The numbers of template and search windows are

    Nz(i)=HzriWzri,Ns(i)=HsriWsri.N_z^{(i)} = \frac{H_z}{r_i}\frac{W_z}{r_i}, \qquad N_s^{(i)} = \frac{H_s}{r_i}\frac{W_s}{r_i}.

    Each window is treated as an indivisible local unit rather than flattening all pixels into a single pixel sequence. The resulting template and search window sequences are projected into queries, keys, and values and processed by a six-layer transformer matching module. Outputs from the multi-scale heads are concatenated and passed to a corner-based bounding-box head. This design performs cross-window matching while retaining the spatial organization and local integrity within each window.

  2. Knowl 2 — Window-level multi-head attention

    equation

    For attention head ii, let Qi∈Rnq×dkQ_i \in \mathbb{R}^{n_q \times d_k} be the query-window representations, Ki∈Rnk×dkK_i \in \mathbb{R}^{n_k \times d_k} the key-window representations, and Vi∈Rnk×dvV_i \in \mathbb{R}^{n_k \times d_v} the value-window representations. Here nqn_q and nkn_k are the numbers of query and key windows, dkd_k is the key dimension, and dvd_v is the value dimension. CSWinTT uses the standard scaled dot-product attention separately in each head:

    Hi=softmax⁡(QiKiTdk)Vi,H_i = \operatorname{softmax}\left(\frac{Q_iK_i^{\mathsf T}}{\sqrt{d_k}}\right)V_i,

    where the softmax is applied row-wise over key windows. The outputs of nhn_h heads are concatenated:

    MultiHead⁡(Q,K,V)=Concat⁡(H1,…,Hnh).\operatorname{MultiHead}(Q,K,V)=\operatorname{Concat}(H_1,\ldots,H_{n_h}).

    In CSWinTT, the template supplies the query windows and the search region supplies the key-value windows for cross-window matching. Compared with pixel-level attention, this reduces the attention resolution from pairwise pixels to pairwise windows while preserving the relative arrangement of pixels inside each window.

  3. Knowl 3 — Multi-scale cyclic shifting of windows

    model/method

    CSWinTT expands each window-level match by cyclically shifting every window through multiple integer offsets. For a base window of size r×rr \times r, let shift⁡(x,y)\operatorname{shift}(x,y) translate the window by xx pixels horizontally and yy pixels vertically, with cyclic wrapping at feature-map boundaries. The offsets satisfy x,y∈{−r+1,…,r−1}x,y \in \{-r+1,\ldots,r-1\}, so one base window produces (2r−1)2(2r-1)^2 shifted samples.

    The shifts are performed independently for each window, so pixels from different original windows are not mixed. They provide additional alignments when the same object part occupies different positions inside two windows, while the shift offset supplies positional information. The conceptual attention operation shifts both query and key-value windows; the implementation later removes redundant query shifts without changing the relative alignment between the query and key samples.

  4. Knowl 4 — Spatially regularized attention mask

    equation

    Cyclic shifts near the edge of the shift range can break object continuity and create boundary artifacts. CSWinTT adds a spatial penalty to the attention score for a shifted sample. For a window of side length rr, an offset (x,y)(x,y) in pixels receives the mask value

    M(x,y)=−(xr)2−(yr)2,x,y∈{−r+1,…,r−1}.M(x,y)=-\left(\frac{x}{r}\right)^2-\left(\frac{y}{r}\right)^2, \qquad x,y\in\{-r+1,\ldots,r-1\}.

    The mask is added to the scaled dot-product attention score associated with that shift:

    S(Q,K)=QKTdk+M,S(Q,K)=\frac{QK^{\mathsf T}}{\sqrt{d_k}}+M,

    where QQ and KK are query and key window representations and dkd_k is the key dimension. The unshifted or nearby samples therefore receive larger relative weights, whereas samples farther from the base window center are penalized more strongly. This regularization is intended to preserve the reliability of spatially close object configurations while retaining the additional alignments created by cyclic shifting.

  5. Knowl 5 — Redundant-computation removal for cyclic attention

    model/method

    Cyclic shifting substantially enlarges the attention computation. For a feature map of height HH, width WW, and channel dimension dd, with window side length rr, the paper gives the standard flattened attention-score cost as O(HW×dHW)O(HW \times dHW) and the feature-fusion cost as O(HW×HW×d)O(HW \times HW \times d). With cyclic window samples, the attention-score cost becomes

    O([HrWr(2r−1)2]2r2d).O\left(\left[\frac{H}{r}\frac{W}{r}(2r-1)^2\right]^2 r^2d\right).

    CSWinTT uses three optimizations. First, it cyclically shifts only the key-value windows and leaves the query windows unchanged, because applying the same shift to both sides does not alter their relative alignment. Second, it halves the shifting periods because shifts generated in opposite directions contain duplicated configurations. Third, it implements cyclic translations by permuting matrix coordinates rather than physically copying and translating feature matrices. These operations remove redundant calculations while preserving the intended relative window matching.

  6. Knowl 6 — Multi-scale head configuration and temporal templates

    model/method

    The implemented transformer has eight attention heads with window sizes

    [r1,…,r8]=[1,2,4,8,1,2,4,8].[r_1,\ldots,r_8]=[1,2,4,8,1,2,4,8].

    The repeated window sizes in the second half of the heads use a search-feature translation by (ri/2,ri/2)(r_i/2,r_i/2) pixels before non-overlapping partitioning. The two partitions provide complementary window contents, reducing the chance that an object is repeatedly cut at the same boundary.

    The tracker uses two templates of identical input size. One remains fixed to the initial-frame template, while the other is updated online from a high-confidence recent tracking result. A score head controls whether the online template is updated. During inference, the initial template and its backbone feature are created in the first frame; each later search region is passed through the tracker, and the predicted bounding box is returned.

  7. Knowl 7 — Training and implementation protocol

    experimental setup

    CSWinTT is trained end-to-end on LaSOT, GOT-10k, and TrackingNet using image pairs sampled from the same video sequence. Brightness jitter and horizontal flipping are used as data augmentation. The template input is 128×128128\times128 pixels; the search region covers 525^2 times the target-box area and is resized to 384×384384\times384 pixels.

    The ResNet-50 backbone is initialized from ImageNet-pretrained weights, while the remaining parameters use Xavier uniform initialization. Training minimizes a weighted combination of L1 bounding-box loss and generalized IoU loss, with weights λL1=5\lambda_{\mathrm{L1}}=5 and λGIoU=2\lambda_{\mathrm{GIoU}}=2. AdamW uses learning rates 10−510^{-5} for backbone parameters and 10−410^{-4} for other parameters. The reported training uses two Nvidia Tesla T4 GPUs, 600 epochs, 4×1044\times10^4 images per epoch, and mini-batches of 64 images. The implementation uses Python 3.7 and PyTorch 1.6; online tracking runs at approximately 12 frames per second on one GPU.

  8. Knowl 8 — Benchmark performance across five tracking datasets

    data/table

    CSWinTT was evaluated on UAV123, LaSOT, TrackingNet, GOT-10k, and VOT2020. AUC denotes the area-under-the-success curve, PP denotes center-distance precision, PNormP_{\mathrm{Norm}} denotes normalized precision, AO denotes average overlap, SR0.5SR_{0.5} and SR0.75SR_{0.75} denote success rates at IoU thresholds 0.5 and 0.75, and EAO, Accuracy, and Robustness are the VOT2020 metrics. The comparison below records CSWinTT and the strongest preceding comparator emphasized by the paper for each dataset.

    Could not parse LaTeX table

    CSWinTT reports the highest AUC on UAV123, LaSOT, TrackingNet, and GOT-10k among the compared methods, with gains over the preceding AUC leaders of 1.3 percentage points on LaSOT, 0.7 on TrackingNet, and 0.6 on GOT-10k. On VOT2020, where the tracker predicts bounding boxes rather than segmentation masks, it obtains the highest reported EAO and robustness among the compared trackers. GOT-10k training follows the benchmark restriction of using only its training split.

  9. Knowl 9 — Component ablation on UAV123

    data/table

    The contribution of window attention, cyclic shifting, spatial regularization, and relative-position encoding was evaluated on UAV123 using AUC and precision. Win means multi-scale window attention, CS means cyclic shifting, SR means the spatially regularized mask, and Pos means relative-position encoding.

    Could not parse LaTeX table

    Window attention without shifting reduces AUC from 66.2 to 54.4 because it makes the similarity map too coarse. Adding cyclic shifts raises AUC to 69.7, a 15.3-point increase over window attention alone and 3.5 points over the original transformer. The spatial mask raises AUC from 69.7 to 70.1, indicating that it alleviates boundary artifacts. Relative-position encoding produces only a small additional effect: the configuration without the mask reaches 69.8 AUC, while the full configuration reaches 70.5 AUC.

  10. Knowl 10 — Window-scale selection and runtime ablations

    empirical result

    On UAV123, cyclic shifting was evaluated with the same window size in all eight heads and with the proposed multi-scale configuration. The results were:

    Could not parse LaTeX table

    The 4×44\times4 single-scale setting is the strongest single-window choice, but combining 1×11\times1, 2×22\times2, 4×44\times4, and 8×88\times8 heads gives the best result, supporting the use of multiple matching scales.

    The runtime effect of the three computational optimizations was measured in frames per second:

    Could not parse LaTeX table

    Without optimization, cyclic shifting is nearly unusable at 1.0 FPS. Removing query shifts provides the largest improvement, and the complete optimization set reaches 12.4 FPS. It remains slower than the original transformer because cyclic shifting introduces additional computation.

  11. Knowl 11 — Qualitative discrimination under occlusion and distractors

    empirical result

    Attention heat maps from the final matching layer were compared between the original pixel-level transformer and CSWinTT on sequences containing occlusion or visually similar distractors. CSWinTT concentrated attention more strongly on the target than the original transformer in these cases.

    The paper attributes this behavior to two properties of the proposed representation. Window partitioning keeps local object parts intact, so an unobscured part can still match even when another part is occluded. Multi-scale heads expose the matcher to different occlusion extents. Cyclic shifts additionally align an object part that lies near the center of a template window with the same part near the boundary of a search window; the shift magnitude supplies a positional cue that helps distinguish the target from nearby distractors.

Coverage note — The full rows of every baseline comparison were not reproduced; only the CSWinTT results and the principal preceding comparator were retained because the additional baseline rows do not add distinct contributed findings.

References

  1. 1.Luca Bertinetto, Jack Valmadre, Joao F. Henriques, Andrea Vedaldi, and Philip H. S. Torr. Fully-convolutional siamese networks for object tracking. In Proceedings of the ECCV, pages 850–865. Springer, 2016. 1, 2, 6
  2. 2.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning discriminative model prediction for tracking. In Proceedings of the ICCV, pages 6182–6191. IEEE, October 2019. 2, 6
  3. 3.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Know your surroundings: Exploiting scene information for object tracking. In Proceedings of the ECCV. Springer, 2020. 6
  4. 4.Goutam Bhat, Joakim Johnander, Martin Danelljan, Fahad Shahbaz Khan, and Michael Felsberg. Unveiling the power of deep tracking. In Proceedings of the ECCV, pages 483–498. Springer, September 2018. 2, 6
  5. 5.David S Bolme, J Ross Beveridge, Bruce A Draper, and Yui Man Lui. Visual object tracking using adaptive correlation filters. In Proceedings of the CVPR, pages 2544–2550. IEEE, June 2010. 2
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229. Springer, 2020. 1, 2
  7. 7.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In Proceedings of the CVPR, pages 8126–8135, 2021. 1, 2, 3, 6
  8. 8.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. Atom: Accurate tracking by overlap maximization. In Proceedings of the CVPR, pages 4660–4669. IEEE, June 2019. 2, 6
  9. 9.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. Eco: Efficient convolution operators for tracking. In Proceedings of the CVPR, pages 6638–6646. IEEE, July 2017. 2, 6
  10. 10.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic regression for visual tracking. In Proceedings of the CVPR, pages 7183–7192, 2020. 2, 6
  11. 11.Martin Danelljan, Andreas Robinson, Fahad Shahbaz Khan, and Michael Felsberg. Beyond correlation filters: Learning continuous convolution operators for visual tracking. In Proceedings of the ECCV, pages 472–488. Springer, 2016. 2
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021. 1, 2
  13. 13.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. Lasot: A high-quality benchmark for large-scale single object tracking. In Proceedings of the CVPR. IEEE, June 2019. 5, 6
  14. 14.Dongyan Guo, Yanyan Shao, Ying Cui, Zhenhua Wang, Liyan Zhang, and Chunhua Shen. Graph attention tracking. In Proceedings of the CVPR, pages 9543–9552, 2021. 2, 6
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the CVPR, pages 770–778. IEEE, June 2016. 3, 5
  16. 16.Joao F Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. High-speed tracking with kernelized correlation filters. IEEE TPAMI, 37(3):583–596, March 2015. 2, 6
  17. 17.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE TPAMI, 2019. 5, 6
  18. 18.Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. Cotr: Correspondence transformer for matching across images. In Proceedings of the ICCV, 2021. 1
  19. 19.Ilchae Jung, Jeany Son, Mooyeol Baek, and Bohyung Han. Real-time mdnet. In Proceedings of the ECCV, pages 83–98. Springer, September 2018. 2
  20. 20.Hamed Kiani Galoogahi, Ashton Fagg, and Simon Lucey. Learning background-aware correlation filters for visual tracking. In Proceedings of the ICCV, pages 1135–1143. IEEE, Oct 2017. 2
  21. 21.Matej Kristan, Ales Leonardis, Jiří Matas, Michael Felsberg, Roman Pflugfelder, Joni-Kristian Kämäräinen, Martin Danelljan, Luka Cehovin Zajc, Alan Lukežič, Ondrej Drbohlav, et al. The eighth visual object tracking vot2020 challenge results. In ECCV, pages 547–601. Springer, 2020. 5, 6, 7
  22. 22.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. Siamrpn++: Evolution of siamese visual tracking with very deep networks. In Proceedings of the CVPR, pages 4282–4291. IEEE, June 2019. 1, 2, 6
  23. 23.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High performance visual tracking with siamese region proposal network. In Proceedings of the CVPR, pages 8971–8980. IEEE, June 2018. 1, 2
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 2
  25. 25.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the ICCV, 2021. 1, 2, 7
  26. 26.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proceedings of the ICLR, 2018. 5
  27. 27.Alan Lukezic, Jiri Matas, and Matej Kristan. D3s-a discriminative single shot segmentation tracker. In Proceedings of the CVPR, pages 7133–7142, 2020. 6
  28. 28.Alan Lukezic, Tomas Vojir, Luka Čehovin Zajc, Jiri Matas, and Matej Kristan. Discriminative correlation filter with channel and spatial reliability. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6309–6318, 2017. 6
  29. 29.Ziang Ma, Linyuan Wang, Haitao Zhang, Wei Lu, and Jun Yin. Rpt: Learning point set representation for siamese visual tracking. In Proceedings of the ECCVW. Springer, 08 2020. 2
  30. 30.Matthias Mueller, Neil Smith, and Bernard Ghanem. A benchmark and simulator for uav tracking. In Proceedings of the ECCV, pages 445–461. Springer, 2016. 5, 6, 7
  31. 31.Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In Proceedings of the ECCV, September 2018. 5, 6
  32. 32.Hyeonseob Nam and Bohyung Han. Learning multi-domain convolutional neural networks for visual tracking. In Proceedings of the CVPR, pages 4293–4302. IEEE, June 2016. 2
  33. 33.Shi Pu, Yibing Song, Chao Ma, Honggang Zhang, and Ming-Hsuan Yang. Deep attentive tracking via reciprocative learning. In NeurIPS, pages 1931–1941. 2018. 2
  34. 34.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the CVPR, pages 658–666, 2019. 5
  35. 35.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 115(3):211–252, 2015. 5
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pages 5998–6008, 2017. 1, 2, 4
  37. 37.Paul Voigtlaender, Jonathon Luiten, Philip HS Torr, and Bastian Leibe. Siam r-cnn: Visual tracking by re-detection. In Proceedings of the CVPR, pages 6578–6588, 2020. 1, 6
  38. 38.Guangting Wang, Chong Luo, Xiaoyan Sun, Zhiwei Xiong, and Wenjun Zeng. Tracking by instance detection: A meta-learning approach. In Proceedings of the CVPR, pages 6288–6297, 2020. 6
  39. 39.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In Proceedings of the CVPR, pages 5463–5474, 2021. 1, 2
  40. 40.Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li. Transformer meets tracker: Exploiting temporal context for robust visual tracking. In Proceedings of the CVPR, pages 1571–1580, 2021. 1, 2, 3, 6
  41. 41.Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H.S. Torr. Fast online object tracking and segmentation: A unifying approach. In Proceedings of the CVPR, pages 1328–1338. IEEE, June 2019. 1, 2
  42. 42.Tianyang Xu, Zhen-Hua Feng, Xiao-Jun Wu, and Josef Kittler. Learning adaptive discriminative correlation filters via temporal consistency preserving spatial feature selection for robust visual object tracking. IEEE TIP, 28(11):5596–5609, 2019. 2
  43. 43.Yihong Xu, Yutong Ban, Guillaume Delorme, Chuang Gan, Daniela Rus, and Xavier Alameda-Pineda. Transcenter: Transformers with dense queries for multiple-object tracking. arXiv preprint arXiv:2103.15145, 2021. 2
  44. 44.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines. In Proceedings of the AAAI, volume 34, pages 12549–12556, 2020. 2, 6
  45. 45.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In Proceedings of the ICCV, 2021. 1, 2, 3, 5, 6
  46. 46.Yuechen Yu, Yilei Xiong, Weilin Huang, and Matthew R Scott. Deformable siamese attention networks for visual object tracking. In Proceedings of the CVPR, pages 6728–6737, 2020. 2, 6
  47. 47.Zhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li, and Weiming Hu. Learn to match: Automatic matching network design for visual tracking. In Proceedings of the ICCV, pages 13339–13348, 2021. 6
  48. 48.Zhipeng Zhang and Houwen Peng. Deeper and wider siamese networks for real-time visual tracking. In Proceedings of the CVPR, pages 4591–4600, 2019. 2
  49. 49.Zhipeng Zhang, Houwen Peng, Jianlong Fu, Bing Li, and Weiming Hu. Ocean: Object-aware anchor-free tracking. In Proceedings of the ECCV, pages 771–787, 2020. 6
  50. 50.Moju Zhao, Kei Okada, and Masayuki Inaba. Trtr: Visual tracking with transformer. arXiv preprint arXiv:2105.03817, 2021. 1
  51. 51.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. IJCV, 127(3):302–321, 2019. 2
  52. 52.Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu. Distractor-aware siamese networks for visual object tracking. In Proceedings of the ECCV, pages 101–117. Springer, September 2018. 2

Citation

MLA
Song, Z., et al. “Transformer Tracking with Cyclic Shifting Window Attention”. arXiv, 2022, http://arxiv.org/abs/2205.03806v1.
APA
Song, Z., Yu, J., Chen, Y.-P. P., & Yang, W. (2022). Transformer Tracking with Cyclic Shifting Window Attention. arXiv. http://arxiv.org/abs/2205.03806v1
Chicago
Song, Z., J. Yu, Y.-P. P. Chen, and W. Yang. 2022. “Transformer Tracking with Cyclic Shifting Window Attention”. arXiv. http://arxiv.org/abs/2205.03806v1.
Harvard
Song, Z. et al. (2022) “Transformer Tracking with Cyclic Shifting Window Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.03806v1.
Vancouver
1. Song Z, Yu J, Chen Y-PP, Yang W (2022) Transformer Tracking with Cyclic Shifting Window Attention. arXiv

BibTeX

@article{song2022transformer,
  title = {Transformer Tracking with Cyclic Shifting Window Attention},
  author = {Song, Zikai and Yu, Junqing and Chen, Yi-Ping Phoebe and Yang, Wei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.03806v1},
  eprint = {2205.03806}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE