TCTrack: Temporal Contexts for Aerial Tracking

Ziang CaoZiyuan HuangLiang PanShiwei ZhangZiwei LiuChanghong Fu

article2022CVPR263 citations

Proposes a real-time aerial tracking framework that exploits multi-level temporal contexts through dynamically calibrated convolutions and memory-efficient transformer refinement to achieve superior tracking accuracy on UAV edge devices.

Listen

Visual tracking from unmanned aerial vehicles is critical for applications like surveying, localization, and motion analysis, but it faces major operational bottlenecks. Drones suffer from abrupt camera motion, fast-moving targets, severe occlusions, and motion blur. Simultaneously, their onboard hardware is heavily constrained by battery life and compute power, making heavy state-of-the-art tracking models impractical. Traditional lightweight trackers frequently lose targets because they process individual video frames in isolation, discarding historical temporal context that reveals object trajectory.

The article demonstrates and evaluates TCTrack, a lightweight visual tracking framework designed to exploit temporal context across consecutive frames without exceeding edge hardware constraints. The objective is to achieve tracking robustness and real-time processing speeds on resource-limited aerial platforms.

The authors designed a two-level temporal architecture. First, an online temporally adaptive convolutional network dynamically calibrates its feature-extraction weights using feature history from past frames. Second, an adaptive temporal transformer filters out noise and refines target similarity maps by updating a compact memory of prior knowledge. The system was evaluated on four standardized aerial tracking benchmarks against 51 existing trackers and tested in real-world flight trials using an onboard NVIDIA Jetson AGX Xavier computing module.

The key findings demonstrate clear performance gains over existing solutions. On standard aerial benchmarks, the framework achieved top accuracy and precision, delivering a 3% to 5% improvement in area-under-the-curve performance compared to the next-best lightweight trackers. Ablation experiments showed that integrating filtered temporal knowledge improved tracking success in fast-motion conditions by roughly 12% to 15% and in partial-occlusion scenarios by about 11%. When compared to heavy, deep-network trackers, the proposed model achieved competitive tracking quality while operating 2.49 times faster than the leading deep alternative, reaching 125.6 frames per second on desktop hardware. In real-world aerial field tests, the system sustained over 27 frames per second on onboard hardware with low resource consumption (averaging 46% GPU and 12.43% CPU utilization) while maintaining high tracking accuracy under diverse lighting and camera motion.

These findings prove that temporal context can compensate for lightweight neural network architectures, closing the performance gap with computationally expensive deep models. For operational teams, this reduces compute hardware costs, lowers payload power consumption, and enhances mission reliability in complex aerial environments. Organizations deploying aerial vision systems should consider incorporating continuous temporal context architectures rather than relying solely on deeper backbones or single-frame detection.

Decision-makers should proceed with field pilots for target operations, although further development is warranted. The authors note that the framework's training regimen focused on short temporal sequences (four frames), meaning its behavior during extended, long-duration occlusions requires further validation. Recommended next steps include converting the pipeline to specialized runtimes like TensorRT and ONNX for further execution speed improvements, alongside governance controls to prevent unauthorized surveillance deployment.

arXiv: 2203.01885
Cover for TCTrack: Temporal Contexts for Aerial Tracking

Abstract

Temporal contexts among consecutive frames are far from being fully utilized in existing visual trackers. In this work, we present TCTrack1, a comprehensive framework to fully exploit temporal contexts for aerial tracking. The temporal contexts are incorporated at two levels: the extraction of features and the refinement of similarity maps. Specifically, for feature extraction, an online temporally adaptive convolution is proposed to enhance the spatial features using temporal information, which is achieved by dynamically calibrating the convolution weights according to the previous frames. For similarity map refinement, we propose an adaptive temporal transformer, which first effectively encodes temporal knowledge in a memory-efficient way, before the temporal knowledge is decoded for accurate adjustment of the similarity map. TCTrack is effective and efficient: evaluation on four aerial tracking benchmarks shows its impressive performance; real-world UAV tests show its high speed of over 27 FPS on NVIDIA Jetson AGX Xavier.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Temporal Contexts for Aerial Tracking
  • 3.1. Feature extraction with online TAdaConv
  • 3.2. Similarity Refinement with AT-Trans
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparison with Light-Weight Trackers
  • 4.3. Ablation Study
  • 4.4. Comparison with Deep Trackers
  • 5. Real-world Tests
  • 6. Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — TCTrack uses temporal context at two stages of Siamese aerial tracking

    model/method

    TCTrack is a Siamese tracking framework that injects temporal information both before and after target-search correlation. Given a template image feature ZZ and the search-region feature XtX_t at frame tt, an online temporally adaptive backbone ϕtada\phi_{\mathrm{tada}} produces features whose convolutional parameters depend on preceding frames. Depth-wise correlation produces a similarity map RtR_t, which is transformed by a convolution into FtF_t. An adaptive temporal transformer then combines FtF_t with a persistent temporal state Mt−1M_{t-1} and outputs a refined map Ft∗F_t^*. Classification and bounding-box regression heads operate on Ft∗F_t^* to produce the tracking result.

    The streaming pipeline is therefore: adaptively extract current features, correlate the template and search features, encode the current map with the previous temporal state, decode the temporal state to refine the current map, predict classification and regression outputs, and retain only the updated temporal state MtM_t for the next frame. This places temporal modeling at both the feature-extraction level through TAdaCNN and the similarity-map level through AT-Trans.

  2. Knowl 2 — Online TAdaConv calibrates convolution weights from a short history of frame descriptors

    model/method

    For an input feature map XtX_t at frame tt, online temporally adaptive convolution produces X~t\widetilde{X}_t using frame-dependent convolution parameters:

    X~t=Wt∗Xt+bt,Wt=Wb⊙αtw,bt=bb⊙αtb.\widetilde{X}_t = W_t * X_t + b_t, \qquad W_t = W_b \odot \alpha_t^w, \qquad b_t = b_b \odot \alpha_t^b.

    Here ∗* is spatial convolution, WbW_b and bbb_b are learned base convolution weights and bias, ⊙\odot denotes channel-wise parameter modulation with broadcasting over the convolution dimensions, and αtw\alpha_t^w and αtb\alpha_t^b are calibration factors for frame tt. The factors are generated from a queue of global descriptors:

    X^t=GAP⁡(Xt)∈RC,X^=Cat⁡(X^t,X^t−1,…,X^t−L+1)∈RL×C,\widehat{X}_t=\operatorname{GAP}(X_t)\in\mathbb{R}^{C}, \qquad \widehat{X}=\operatorname{Cat}(\widehat{X}_t,\widehat{X}_{t-1},\ldots,\widehat{X}_{t-L+1})\in\mathbb{R}^{L\times C}, αtw=Fw(X^)+1,αtb=Fb(X^)+1.\alpha_t^w=\mathcal{F}_w(\widehat{X})+1, \qquad \alpha_t^b=\mathcal{F}_b(\widehat{X})+1.

    CC is the number of feature channels, LL is the temporal queue length, GAP⁡\operatorname{GAP} is global average pooling, and Fw\mathcal{F}_w and Fb\mathcal{F}_b are one-dimensional convolutions whose kernel spans the LL descriptors. The calibration-convolution weights are initialized to zero, so the layer initially behaves as an ordinary convolution. Because the tracker processes one frame at a time, the calibration uses only past and current descriptors rather than future frames. When fewer than LL frames are available, the descriptor from the first frame fills the missing queue entries. The final tracker uses L=3L=3 and replaces the last two AlexNet convolutional layers with online TAdaConv.

  3. Knowl 3 — AT-Trans encodes a filtered temporal prior with an adaptive transformer encoder

    model/method

    The adaptive temporal encoder receives the current similarity feature map FtF_t and the previous temporal prior Mt−1M_{t-1}. Its first attention layer uses the previous prior as the query and the current map as both key and value, thereby emphasizing current-frame evidence; a second self-attention layer further integrates the resulting representation:

    Et1=Norm⁡(Ft+MHA⁡(Mt−1,Ft,Ft)),E_t^{1}=\operatorname{Norm}\left(F_t+\operatorname{MHA}(M_{t-1},F_t,F_t)\right), Et2=Norm⁡(Et1+MHA⁡(Et1,Et1,Et1)).E_t^{2}=\operatorname{Norm}\left(E_t^{1}+\operatorname{MHA}(E_t^{1},E_t^{1},E_t^{1})\right).

    Norm⁡\operatorname{Norm} is layer normalization and MHA⁡(Q,K,V)\operatorname{MHA}(Q,K,V) is multi-head attention over spatial tokens. AT-Trans uses six attention heads; for an input channel dimension CiC_i, each head has dimension Ci/6C_i/6. For head hh, the attention uses the standard scaled operation softmax⁡(QWqh(KWkh)⊤/dh)VWvh\operatorname{softmax}(QW_q^h(KW_k^h)^\top/\sqrt{d_h})VW_v^h, followed by concatenation and an output projection.

    To suppress unreliable context caused by blur or occlusion, a channel-wise temporal filter is generated from the global descriptor of Et1E_t^1:

    αt=FFN⁡(GAP⁡(C(Et1))),\alpha_t=\operatorname{FFN}\left(\operatorname{GAP}(\mathcal{C}(E_t^{1}))\right), Etf=Et2+C(Cat⁡(Et2,Et1))⊙αt,E_t^{f}=E_t^{2}+\mathcal{C}\left(\operatorname{Cat}(E_t^{2},E_t^{1})\right)\odot\alpha_t,

    where C\mathcal{C} is a convolution, FFN⁡\operatorname{FFN} is a feed-forward network, and ⊙\odot is channel-wise multiplication. The updated temporal prior is obtained by self-attending to the filtered representation:

    Mt=Norm⁡(Etf+MHA⁡(Etf,Etf,Etf)).M_t=\operatorname{Norm}\left(E_t^{f}+\operatorname{MHA}(E_t^{f},E_t^{f},E_t^{f})\right).

    Only MtM_t is retained for the next frame; the intermediate temporal features from earlier frames are not stored. For the first frame, the prior is initialized from the target-specific initial similarity map using M0=Cinit(R1)M_0=\mathcal{C}_{\mathrm{init}}(R_1) rather than a target-independent random tensor.

  4. Knowl 4 — The AT-Trans decoder uses temporal attention to refine each similarity map

    model/method

    The adaptive temporal decoder refines the current similarity feature map FtF_t using the encoded temporal prior MtM_t. It first applies self-attention to the current map, then uses the resulting representation as a query and the temporal prior as both key and value:

    Dt3=Norm⁡(Ft+MHA⁡(Ft,Ft,Ft)),D_t^{3}=\operatorname{Norm}\left(F_t+\operatorname{MHA}(F_t,F_t,F_t)\right), Dt4=Norm⁡(Dt3+MHA⁡(Dt3,Mt,Mt)),D_t^{4}=\operatorname{Norm}\left(D_t^{3}+\operatorname{MHA}(D_t^{3},M_t,M_t)\right), Ft∗=Norm⁡(Dt4+FFN⁡(Dt4)).F_t^*=\operatorname{Norm}\left(D_t^{4}+\operatorname{FFN}(D_t^{4})\right).

    The attention map selectively extracts useful information from the temporal prior while retaining the current spatial evidence. The refined feature Ft∗F_t^*, rather than the unrefined similarity map, is passed to the classification and regression heads. Qualitative similarity maps reported by the paper show that this refinement is especially useful under camera motion, fast motion, and occlusion.

  5. Knowl 5 — TCTrack is trained as a lightweight multi-frame Siamese tracker

    experimental setup

    The tracker uses AlexNet initialized from ImageNet because aerial deployment prioritizes low latency. Training examples are four-frame videos sampled from VID, LaSOT, and GOT-10K. TCTrack is trained for 100 epochs on two NVIDIA TITAN RTX GPUs with SGD, momentum 0.90.9, and mini-batches of 124 template-search pairs. The learning rate decreases logarithmically from 0.0050.005 to 0.00050.0005; the AlexNet backbone is frozen for the first 10 epochs and trained thereafter. Template and search inputs are 127×127127\times127 and 287×287287\times287 pixels, respectively. Online TAdaConv replaces the last two convolutional layers, while AT-Trans is randomly initialized.

    The framework is evaluated on UAV123, UAV123@10fps, DTB70, and the long-term UAVTrack112_L benchmark. Comparisons include 51 existing trackers, with Siamese trackers evaluated using the same AlexNet backbone where applicable.

  6. Knowl 6 — TCTrack improves precision and success across four aerial tracking benchmarks

    empirical result

    The full TCTrack obtains the following reported overall scores, where precision is the precision-plot score and success is the success-plot area-under-the-curve score:

    • UAV123: precision 0.8000.800, success 0.6040.604.
    • UAV123@10fps: precision 0.7740.774, success 0.5880.588.
    • DTB70: precision 0.8130.813, success 0.6220.622.
    • UAVTrack112_L: precision 0.7860.786, success 0.5820.582.

    On UAV123, the paper reports an AUC improvement of 3% over HiFT and 4.3% over SiamRPN++. On DTB70, which emphasizes severe motion, TCTrack ranks first with a reported 5% AUC improvement over the strongest competing tracker. On UAVTrack112_L, TCTrack exceeds SiamRPN++ (0.5590.559 success, 0.7730.773 precision) and HiFT (0.5510.551 success, 0.7340.734 precision). Attribute evaluations also place TCTrack first in the reported fast-camera-motion, background-clutter, partial-occlusion, and deformation subsets, supporting its robustness in the conditions targeted by the temporal design.

  7. Knowl 7 — Ablations show that filtering, target-specific initialization, and prior-based queries are essential

    empirical result

    On UAV123, the transformer without the temporal information filter and trained with single-frame examples obtains overall precision 0.7500.750 and success 0.5500.550. Adding the temporal information filter raises these values to 0.7650.765 and 0.5730.573. Multi-frame training with convolutional initialization but without the filter instead gives 0.7320.732 precision and 0.5080.508 success, showing that simply propagating unfiltered temporal information can harm tracking.

    With the filter and multi-frame training, random temporal-prior initialization gives 0.7720.772 precision and 0.5860.586 success, whereas convolutional initialization from the first-frame similarity map gives 0.8000.800 precision and 0.6040.604 success. Using the current similarity map as the encoder query gives 0.7710.771 precision and 0.5800.580 success, while using the previous temporal prior as the query gives the best overall result, 0.8000.800 precision and 0.6040.604 success. For the best configuration, the reported precision/success pairs are 0.810/0.6150.810/0.615 for camera motion, 0.793/0.5860.793/0.586 for fast motion, and 0.710/0.5100.710/0.510 for partial occlusion.

    When TAdaConv is added to the transformer baseline, increasing its descriptor-history length improves overall performance: L=1L=1 gives 0.7490.749 precision and 0.5610.561 success, L=2L=2 gives 0.7740.774 and 0.5730.573, and L=3L=3 gives 0.7760.776 and 0.5800.580, compared with 0.7500.750 and 0.5500.550 without TAdaConv. The adopted history length is therefore L=3L=3.

  8. Knowl 8 — The online temporal design preserves real-time efficiency on desktop and embedded hardware

    empirical result

    The lightweight AlexNet-based implementation reaches a reported rate of up to 125.6 FPS in the desktop comparison while remaining competitive with trackers using substantially deeper backbones; the paper reports that it runs 2.49 times faster than TransT in that comparison. On an NVIDIA Jetson AGX Xavier, AlexNet inference alone takes 3.4 ms for a 287×287×3287\times287\times3 input and has 2.47 million parameters, lower latency than the other evaluated backbone choices.

    For real-world tests, the tracker runs onboard an NVIDIA Jetson AGX Xavier connected to a Pixhawk2 flight controller. Without TensorRT acceleration, it maintains more than 27 FPS. Average resource usage is 15.29% RAM, 3% GPU VRAM, 46% GPU utilization, and 12.43% CPU utilization. Tracking is evaluated with center location error, using 20 pixels as the success threshold. The recorded UAV sequences include scale variation, partial occlusion, motion blur, low illumination, and camera motion; the tracker maintains stable trajectories in these conditions.

  9. Knowl 9 — The paper does not establish long-term temporal modeling under extended occlusion

    limitation

    The training procedure uses short four-frame videos, so the paper explicitly states that the potential of TCTrack for very long-term temporal modeling and prolonged occlusion is not fully explored. The authors also identify the absence of TensorRT and ONNX versions as an engineering limitation for deployment. Finally, they note a negative societal risk: the efficiency and robustness that make TCTrack suitable for UAV deployment could facilitate unauthorized surveillance.

Coverage note — No substantial contributed material was omitted; detailed per-sequence plots, full competitor rankings, and qualitative image panels were condensed because their main evidence is represented by the reported benchmark, ablation, and deployment results.

References

  1. 1.Luca Bertinetto, Jack Valmadre, Stuart Golodetz, Ondrej Miksik, and Philip HS Torr. Staple: Complementary Learners for Real-Time Tracking. In CVPR, pages 1401–1409, 2016. 5
  2. 2.Luca Bertinetto, Jack Valmadre, Joao F. Henriques, Andrea Vedaldi, and Philip H. S. Torr. Fully-Convolutional Siamese Networks for Object Tracking. In ECCV, pages 850–865, 2016. 1, 2, 5, 6
  3. 3.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Know Your Surroundings: Exploiting Scene Information for Object Tracking. In ECCV, pages 205–221, 2020. 2
  4. 4.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning Discriminative Model Prediction for Tracking. In ICCV, pages 6181–6190, 2019. 1, 5
  5. 5.David S Bolme, J Ross Beveridge, Bruce A Draper, and Yui Man Lui. Visual Object Tracking Using Adaptive Correlation Filters. In CVPR, pages 2544–2550, 2010. 2
  6. 6.Ziang Cao, Changhong Fu, Junjie Ye, Bowen Li, and Yiming Li. HiFT: Hierarchical Feature Transformer for Aerial Tracking. In ICCV, pages 15437–15446, 2021. 1, 5, 6
  7. 7.Ziang Cao, Changhong Fu, Junjie Ye, Bowen Li, and Yiming Li. SiamAPN++: Siamese Attentional Aggregation Network for Real-Time UAV Tracking. In IROS, pages 1–7, 2021. 1, 2, 5, 6
  8. 8.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer Tracking. In CVPR, pages 8126–8135, 2021. 5
  9. 9.Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. Siamese Box Adaptive Network for Visual Tracking. In CVPR, pages 6668–6677, 2020. 2, 5
  10. 10.Kenan Dai, Dong Wang, Huchuan Lu, Chong Sun, and Jianhua Li. Visual Tracking via Adaptive Spatially-Regularized Correlation Filters. In CVPR, pages 4670–4679, 2019. 2
  11. 11.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ATOM: Accurate Tracking by Overlap Maximization. In CVPR, pages 4655–4664, 2019. 1, 5
  12. 12.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ECO: Efficient Convolution Operators for Tracking. In CVPR, pages 6931–6939, 2017. 5, 6
  13. 13.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic Regression for Visual Tracking. In CVPR, pages 7181–7190, 2020. 5
  14. 14.Martin Danelljan, Gustav Hager, Fahad Khan, and Michael Felsberg. Accurate Scale Estimation for Robust Visual Tracking. In BMVC, 2014. 5
  15. 15.Martin Danelljan, Gustav Hager, Fahad Shahbaz Khan, and Michael Felsberg. Discriminative Scale Space Tracking. PAMI, 39(8):1561–1575, 2016. 5
  16. 16.Martin Danelljan, Gustav Hager, Fahad Shahbaz Khan, and Michael Felsberg. Learning Spatially Regularized Correlation Filters for Visual Tracking. In ICCV, pages 4310–4318, 2015. 1, 2, 5, 6
  17. 17.Martin Danelljan, Andreas Robinson, Fahad Shahbaz Khan, and Michael Felsberg. Beyond Correlation Filters: Learning Continuous Convolution Operators for Visual Tracking. In ECCV, pages 472–488, 2016. 5, 6
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR, 2021. 4
  19. 19.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. Lasot: A High-Quality Benchmark for Large-Scale Single Object Tracking. In CVPR, pages 5374–5383, 2019. 5
  20. 20.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast Networks for Video Recognition. In ICCV, pages 6202–6211, 2019. 2
  21. 21.Changhong Fu, Ziang Cao, Yiming Li, Junjie Ye, and Chen Feng. Onboard Real-Time Aerial Tracking With Efficient Siamese Anchor Proposal Network. TGRS, pages 1–13, 2021. 1, 2, 5, 6
  22. 22.Changhong Fu, Ziang Cao, Yiming Li, Junjie Ye, and Chen Feng. Siamese Anchor Proposal Network for High-Speed Aerial Tracking. In ICRA, pages 1–7, 2021. 1, 2, 5
  23. 23.Zhihong Fu, Qingjie Liu, Zehua Fu, and Yunhong Wang. STMTrack: Template-free Visual Tracking with Space-time Memory Networks. In CVPR, pages 13774–13783, 2021. 2, 5
  24. 24.Junyu Gao, Tianzhu Zhang, and Changsheng Xu. Graph Convolutional Tracking. In CVPR, pages 4649–4659, 2019. 2
  25. 25.Dongyan Guo, Yanyan Shao, Ying Cui, Zhenhua Wang, Liyan Zhang, and Chunhua Shen. Graph Attention Tracking. In CVPR, pages 1–10, 2021. 5
  26. 26.Dongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang, and Shengyong Chen. SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking. In CVPR, pages 6268–6276, 2020. 2, 5
  27. 27.Qing Guo, Wei Feng, Ce Zhou, Rui Huang, Liang Wan, and Song Wang. Learning Dynamic Siamese Network for Visual Object Tracking. In ICCV, pages 1781–1789, 2017. 2, 5, 6
  28. 28.Tengda Han, Weidi Xie, and Andrew Zisserman. Video Representation Learning by Dense Predictive Coding. In CVPRW, pages 1–10, 2019. 2
  29. 29.Tengda Han, Weidi Xie, and Andrew Zisserman. Memory-Augmented Dense Predictive Coding for Video Representation Learning. In ECCV, pages 312–329, 2020. 2
  30. 30.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR, pages 770–778, 2016. 5
  31. 31.Joao F. Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. High-Speed Tracking with Kernelized Correlation Filters. PAMI, pages 583–596, 2015. 1, 2
  32. 32.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A Large High-Diversity Benchmark for Generic Object Tracking in The Wild. TPAMI, 43(5):1562–1577, 2019. 5
  33. 33.Ziyuan Huang, Changhong Fu, Yiming Li, Fuling Lin, and Peng Lu. Learning Aberrance Repressed Correlation Filters for Real-Time UAV Tracking. In ICCV, pages 2891–2900, Nov. 2019. 1, 2, 5, 6
  34. 34.Ziyuan Huang, Shiwei Zhang, Jianwen Jiang, Mingqian Tang, Rong Jin, and Marcelo H Ang. Self-Supervised Motion Learning from Static Images. In CVPR, pages 1276–1285, 2021. 2
  35. 35.Ziyuan Huang, Shiwei Zhang, Liang Pan, Zhiwu Qing, Mingqian Tang, Ziwei Liu, and Marcelo H Ang Jr. TAda! Temporally-Adaptive Convolutions for Video Understanding. In ICLR, 2022. 2, 3
  36. 36.Yuqi Huo, Mingyu Ding, Haoyu Lu, Ziyuan Huang, Mingqian Tang, Zhiwu Lu, and Tao Xiang. Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw. In IJCAI, pages 751–757, 2021. 2
  37. 37.Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer. SqueezeNet: AlexNet-Level Accuracy with 50x Fewer Parameters and¡ 0.5 MB Model Size. arXiv preprint arXiv:1602.07360, 2016. 5
  38. 38.Hamed Kiani Galoogahi, Ashton Fagg, and Simon Lucey. Learning Background-Aware Correlation Filters for Visual Tracking. In ICCV, pages 1135–1143, 2017. 1, 2, 5, 6
  39. 39.Dahun Kim, Donghyeon Cho, and In So Kweon. Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles. In AAAI, volume 33, pages 8545–8552, 2019. 2
  40. 40.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet Classification with Deep Convolutional Neural Networks. In NIPS, pages 1097–1105, 2012. 5
  41. 41.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks. In CVPR, pages 4277–4286, 2019. 1, 2, 3, 5, 6
  42. 42.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High Performance Visual Tracking with Siamese Region Proposal Network. In CVPR, pages 8971–8980, 2018. 1, 2
  43. 43.Feng Li, Cheng Tian, Wangmeng Zuo, Lei Zhang, and Ming-Hsuan Yang. Learning Spatial-Temporal Regularized Correlation Filters for Visual Tracking. In CVPR, pages 4904–4913, 2018. 2, 5, 6
  44. 44.Feng Li, Yingjie Yao, Peihua Li, David Zhang, Wangmeng Zuo, and Ming-Hsuan Yang. Integrating Boundary and Center Correlation Filters for Visual Tracking with Aspect Ratio Variation. In ICCVW, pages 2001–2009, 2017. 5
  45. 45.Siyi Li and Dit-Yan Yeung. Visual Object Tracking for Unmanned Aerial Vehicles: A Benchmark and New Motion Models. In AAAI, pages 1–7, 2017. 5, 6
  46. 46.Xin Li, Chao Ma, Baoyuan Wu, Zhenyu He, and Ming-Hsuan Yang. Target-Aware Deep Tracking. In CVPR, pages 1369–1378, 2019. 5, 6
  47. 47.Yiming Li, Changhong Fu, Fangqiang Ding, Ziyuan Huang, and Geng Lu. AutoTrack: Towards High-Performance Visual Tracking for UAV With Automatic Spatio-Temporal Regularization. In CVPR, pages 11920–11929, Jun. 2020. 1, 2, 5, 6
  48. 48.Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal Shift Module for Efficient Video Understanding. In ICCV, pages 7083–7093, 2019. 2
  49. 49.Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu. TAM: Temporal Adaptive Module for Video Recognition. In ICCV, pages 13708–13718, 2021. 2
  50. 50.Alan Lukezic, Jiri Matas, and Matej Kristan. D3S-A Discriminative Single Shot Segmentation Tracker. In CVPR, pages 7133–7142, 2020. 5
  51. 51.Alan Lukezic, Tomas Vojir, Luka Cehovin Zajc, Jiri Matas, and Matej Kristan. Discriminative Correlation Filter with Channel and Spatial Reliability. In CVPR, pages 6309–6318, 2017. 5
  52. 52.Chao Ma, Jia-Bin Huang, Xiaokang Yang, and Ming-Hsuan Yang. Hierarchical Convolutional Features for Visual Tracking. In ICCV, pages 3074–3082, 2015. 5
  53. 53.Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool. Learning Target Candidate Association to Keep Track of What Not to Track. In ICCV, pages 13444–13454, 2021. 5
  54. 54.Matthias Mueller, Neil Smith, and Bernard Ghanem. A Benchmark and Simulator for UAV Tracking. In ECCV, pages 445–461, 2016. 5, 6, 7
  55. 55.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet Large Scale Visual Recognition Challenge. International journal of computer vision, 115(3):211–252, 2015. 5
  56. 56.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In CVPR, pages 4510–4520, 2018. 5
  57. 57.Jia Shao, Bo Du, Chen Wu, and Lefei Zhang. Tracking Objects from Satellite Videos: A Velocity Feature Based Correlation Filter. TGRS, 57(10):7860–7871, 2019. 1
  58. 58.Karen Simonyan and Andrew Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556, 2014. 5
  59. 59.Ivan Sosnovik, Artem Moskalev, and Arnold WM Smeulders. Scale Equivariance Improves Siamese Tracking. In WACV, pages 2765–2774, 2021. 5
  60. 60.Mingxing Tan and Quoc Le. Efficientnet: Rethinking Model Scaling for Convolutional Neural Networks. In ICML, pages 6105–6114, 2019. 5
  61. 61.Mani Thomas, Chandra Kambhamettu, and Cathleen A Geiger. Motion Tracking of Discontinuous Sea Ice. TGRS, 49(12):5064–5079, 2011. 1
  62. 62.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning Spatiotemporal Features with 3D Convolutional Networks. In ICCV, pages 4489–4497, 2015. 2
  63. 63.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A Closer Look at Spatiotemporal Convolutions for Action Recognition. In CVPR, pages 6450–6459, 2018. 2
  64. 64.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pages 6000–6010, 2017. 4
  65. 65.Chen Wang, Le Zhang, Lihua Xie, and Junsong Yuan. Kernel Cross-Correlator. In AAAI, volume 32, 2018. 5
  66. 66.Ning Wang, Yibing Song, Chao Ma, Wengang Zhou, Wei Liu, and Houqiang Li. Unsupervised Deep Tracking. In CVPR, pages 1308–1317, 2019. 5, 6
  67. 67.Ning Wang, Wengang Zhou, Qi Tian, Richang Hong, Meng Wang, and Houqiang Li. Multi-Cue Correlation Filters for Robust Visual Tracking. In CVPR, pages 4844–4853, 2018. 5
  68. 68.Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li. Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking. In CVPR, pages 1571–1580, 2021. 2, 5
  69. 69.Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip HS Torr. Fast Online Object Tracking and Segmentation: A Unifying Approach. In CVPR, pages 1328–1338, 2019. 5
  70. 70.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, pages 7794–7803, 2018. 2
  71. 71.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines. In AAAI, volume 34, pages 12549–12556, 2020. 5
  72. 72.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning Spatio-Temporal Transformer for Visual Tracking. In CVPR, pages 1–10, 2021. 2
  73. 73.Tianyu Yang and Antoni B Chan. Learning Dynamic Memory Networks for Object Tracking. In ECCV, pages 152–167, 2018. 2
  74. 74.Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan, and Fahad Shahbaz Khan. Learning the Model Update for Siamese Trackers. In ICCV, pages 4010–4019, 2019. 2, 5
  75. 75.Le Zhang and Ponnuthurai Nagaratnam Suganthan. Robust Visual Tracking via Co-Trained Kernelized Correlation Filters. PR, 69:82–93, 2017. 5, 6
  76. 76.Tianzhu Zhang, Changsheng Xu, and Ming-Hsuan Yang. Multi-Task Correlation Particle Filter for Robust Object Tracking. In CVPR, pages 4335–4343, 2017. 5
  77. 77.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. In CVPR, pages 6848–6856, 2018. 5
  78. 78.Zhipeng Zhang and Houwen Peng. Deeper and Wider Siamese Networks for Real-time Visual Tracking. In CVPR, pages 4591–4600, 2019. 2, 5
  79. 79.Zhipeng Zhang, Houwen Peng, Jianlong Fu, Bing Li, and Weiming Hu. Ocean: Object-Aware Anchor-Free Tracking. In ECCV, pages 771–787, 2020. 5
  80. 80.Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu. Distractor-Aware Siamese Networks for Visual Object Tracking. In ECCV, pages 101–117, 2018. 5, 6

Citation

MLA
Cao, Z., et al. “TCTrack: Temporal Contexts for Aerial Tracking”. arXiv, 2022, http://arxiv.org/abs/2203.01885v3.
APA
Cao, Z., Huang, Z., Pan, L., Zhang, S., Liu, Z., & Fu, C. (2022). TCTrack: Temporal Contexts for Aerial Tracking. arXiv. http://arxiv.org/abs/2203.01885v3
Chicago
Cao, Z., Z. Huang, L. Pan, S. Zhang, Z. Liu, and C. Fu. 2022. “TCTrack: Temporal Contexts for Aerial Tracking”. arXiv. http://arxiv.org/abs/2203.01885v3.
Harvard
Cao, Z. et al. (2022) “TCTrack: Temporal Contexts for Aerial Tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.01885v3.
Vancouver
1. Cao Z, Huang Z, Pan L, Zhang S, Liu Z, Fu C (2022) TCTrack: Temporal Contexts for Aerial Tracking. arXiv

BibTeX

@article{cao2022tctrack,
  title = {TCTrack: Temporal Contexts for Aerial Tracking},
  author = {Cao, Ziang and Huang, Ziyuan and Pan, Liang and Zhang, Shiwei and Liu, Ziwei and Fu, Changhong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.01885v3},
  eprint = {2203.01885}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE