Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking

Yongxin LiMengyuan LiuYou WuXucheng WangXiangyang YangShuiwang Li

article2024ICML61 citations

Presents AVTrack, an efficient UAV tracking framework that combines dynamic transformer block activation with mutual information maximization to achieve real-time tracking speeds of roughly 220 FPS while resisting severe viewing angle variations.

Listen

Deploying computer vision models on unmanned aerial vehicles (UAVs) requires balancing high tracking accuracy with extreme computational efficiency, as drones operate under strict onboard processing and battery constraints. While modern vision transformer models deliver superior tracking precision in complex aerial environments, their computational overhead has largely prevented reliable real-time deployment on lightweight drone hardware.

The article develops and evaluates AVTrack, an adaptive and view-invariant vision transformer tracking framework designed specifically for real-time UAV applications.

The approach introduces two key mechanisms into a single-stream transformer architecture: an Activation Module that dynamically skips unnecessary computational blocks based on scene complexity, and a view-invariant representation learning loss that maximizes mutual information between different perspectives of the target during training without adding inference costs. The authors conducted extensive evaluations across five standard aerial tracking benchmarks—DTB70, UAVDT, VisDrone2018, UAV123, and UAV123@10fps—benchmarking against 13 lightweight trackers and 14 deep tracking systems, followed by embedded deployment on an onboard NVIDIA Jetson AGX Xavier platform.

The findings demonstrate substantial improvements across efficiency, accuracy, and generalizability. AVTrack achieved real-time speeds on a standard computer, reaching between 250 and 283 frames per second (FPS) on a graphics processing unit (GPU) and approximately 60 FPS on a central processing unit (CPU), running over 1.4 times faster than previous adaptive transformer trackers. In accuracy, AVTrack delivered state-of-the-art results, achieving an average precision of 84.1% and a success rate of 64.3% across benchmarks, outperforming conventional correlation filters and lightweight deep networks. Crucially, learning view-invariant representations boosted baseline precision by roughly 2.5% to 4.6% across models, directly resolving drone-specific challenges such as severe viewing angle shifts. When integrated into other leading tracking frameworks, the core modules consistently boosted inference speeds by 14% to 22% with negligible impact on accuracy, and real-world embedded flight testing confirmed smooth operation at 42.4 FPS with low resource utilization (39.7% GPU and 13.5% CPU).

These results establish that structured block-level conditional computation and mutual-information-based training allow high-capacity transformer models to run efficiently on resource-constrained aerial hardware without sacrificing tracking robustness. This reduces the risk of target loss during sharp drone maneuvers, improves battery life through lower computational loads, and eliminates the traditional trade-off between tracking accuracy and operational speed.

Organizations developing autonomous drone systems should adopt structured conditional activation and view-invariant loss objectives when deploying onboard vision transformers. Engineering teams can integrate these plug-and-play components into existing tracker architectures or deploy AVTrack directly on embedded edge hardware, tuning the initial layer activation parameter depending on specific platform speed-versus-accuracy requirements.

While empirical confidence is high across standard public datasets and the tested Jetson Xavier platform, real-time performance guarantees remain subject to the capabilities of specific onboard hardware. Further testing across broader environmental extremes, such as nighttime conditions and highly degraded weather, is recommended prior to mission-critical operational deployment.

Cover for Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking

Abstract

Harnessing transformer-based models, visual tracking has made substantial strides. However, the sluggish performance of current trackers limits their practicality on devices with constrained computational capabilities, especially for real-time unmanned aerial vehicle (UAV) tracking. Addressing this challenge, we introduce AVTrack, an adaptive computation framework tailored to selectively activate transformer blocks for real-time UAV tracking in this work. Our novel Activation Module (AM) dynamically optimizes ViT architecture, selectively engaging relevant components and enhancing inference efficiency without compromising much tracking performance. Moreover, we bolster the effectiveness of ViTs, particularly in addressing challenges arising from extreme changes in viewing angles commonly encountered in UAV tracking, by learning view-invariant representations through mutual information maximization. Extensive experiments on five tracking benchmarks affirm the effectiveness and versatility of our approach, positioning it as a state-of-the-art solution in visual tracking. Code is released at: https://github.com/wuyou3474/AVTrack.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Visual Tracking
  • 2.2. Efficient Vision Transformer
  • 2.3. View-Invariant Feature Representation
  • 3. Method
  • 3.1. Overview
  • 3.2. Activation Module (AM)
  • 3.3. View-Invariant Representations (VIR) via Mutual Information Maximization
  • 3.4. Prediction Head and Training Loss
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparison with Lightweight Trackers
  • 4.3. Comparison with Deep Trackers
  • 4.4. Attribute-Based Evaluation
  • 4.5. Ablation Study
  • 4.6. Qualitative Results
  • 5. Conclusions
  • Acknowledgements
  • Impact Statement
  • References
  • Appendices
  • A. Comparison with Deep Trackers
  • B. Impact of Weighting the Loss for Learning View-Invariant Feature Representations
  • C. Impact of Activation Module (AM) and View-Invariant Representations (VIR).
  • D. Application to SOTA Trackers
  • E. Study on When to Initiate the AM
  • F. Real-world Tests
  • G. Attribute-Based Evaluation

Knowls

  1. Knowl 1 — AVTrack Single-Stream Transformer Tracking Framework

    model/method

    AVTrack is an end-to-end single-stream visual tracking framework tailored for real-time unmanned aerial vehicle (UAV) applications. The architecture unifies feature extraction and cross-attention fusion within a lightweight pre-trained Vision Transformer (ViT) backbone (such as ViT-tiny, DeiT-tiny, or EVA-tiny) coupled with a fully convolutional prediction head.

    The framework processes a pair of inputs: a target template image Z∈R3×Hz×WzZ \in \mathbb{R}^{3 \times H_z \times W_z} and a search image X∈R3×Hx×WxX \in \mathbb{R}^{3 \times H_x \times W_x} (typically sized at 128×128128 \times 128 and 256×256256 \times 256 pixels, respectively). Images are divided into non-overlapping patches of size P×PP \times P, yielding Pz=(Hz×Wz)/P2P_z = (H_z \times W_z)/P^2 template tokens and Px=(Hx×Wx)/P2P_x = (H_x \times W_x)/P^2 search tokens, with a total token sequence length of K=Pz+PxK = P_z + P_x and embedding dimension dd.

    The tokens are jointly processed through NN transformer blocks. To ensure computational efficiency, layers beyond the first nfn_f layers incorporate an Activation Module that dynamically selects whether to execute or bypass subsequent transformer blocks. The final output tokens corresponding to the search area are reshaped into a 2D spatial feature map and passed to a prediction head H\mathcal{H} comprising four stacked Conv-BN-ReLU layers. The head predicts:

    1. A local coordinate offset map o∈[0,1]2×(Hx/P)×(Wx/P)o \in [0, 1]^{2 \times (H_x/P) \times (W_x/P)},
    2. A normalized bounding box dimension map s∈[0,1]2×(Hx/P)×(Wx/P)s \in [0, 1]^{2 \times (H_x/P) \times (W_x/P)},
    3. A target classification score map p∈[0,1](Hx/P)×(Wx/P)p \in [0, 1]^{(H_x/P) \times (W_x/P)}.

    During inference, the classification score map is element-wise multiplied by a Hanning window matching the spatial dimensions to inject a positional prior. The predicted center position is chosen via (xc,yc)=argmax⁡(x,y)p(x,y)(x_c, y_c) = \operatorname{argmax}_{(x, y)} p(x, y), and the final target bounding box is computed as: {(xt,yt);(w,h)}={(xc,yc)+o(xc,yc);s(xc,yc)}\{(x_t, y_t); (w, h)\} = \{(x_c, y_c) + o(x_c, y_c); s(x_c, y_c)\}

  2. Knowl 2 — Activation Module for Dynamic Transformer Block Skipping

    model/method

    The Activation Module (AM) introduces structured conditional computation into Vision Transformers by adaptively activating or skipping entire transformer blocks based on the input context, thereby avoiding the irregular memory access latencies caused by dynamic token-pruning operations.

    For any transformer layer ii where i>nfi > n_f (with nfn_f denoting the count of mandatory initial active layers), let t1:Ki−1(Z,X)∈RK×dt_{1:K}^{i-1}(Z, X) \in \mathbb{R}^{K \times d} denote the KK output tokens of dimension dd produced by layer i−1i-1. To minimize computation within the AM itself, the module extracts a 1D token slice by projecting with the first standard basis vector e1=[1,0,…,0]T∈Rde_1 = [1, 0, \dots, 0]^T \in \mathbb{R}^d: ri−1=e1Tt1:Ki−1(Z,X)∈RKr^{i-1} = e_1^T t_{1:K}^{i-1}(Z, X) \in \mathbb{R}^K

    The activation probability pi∈[0,1]p^i \in [0, 1] for the ii-th transformer block is computed via a learnable linear layer Li:RK→R\mathcal{L}^i: \mathbb{R}^K \to \mathbb{R} and a sigmoid activation function σ(x)=1/(1+e−x)\sigma(x) = 1 / (1 + e^{-x}): pi=σ(Li(ri−1))p^i = \sigma(\mathcal{L}^i(r^{i-1}))

    During forward inference, the block execution decision follows a thresholding rule with hyperparameter β∈(0.5,1)\beta \in (0.5, 1):

    • If pi>βp^i > \beta, the ii-th transformer block is executed.
    • If pi≤βp^i \le \beta, the ii-th transformer block is bypassed, and the token sequence t1:Ki−1(Z,X)t_{1:K}^{i-1}(Z, X) is routed directly to the input of block i+1i+1.

    To preserve foundational low-level feature extraction and cross-attention between the template and search image, the first nfn_f transformer blocks are permanently active (set to nf=1n_f = 1 by default).

  3. Knowl 3 — Block Sparsity Regularization Loss

    equation

    To prevent the network from trivially activating all adaptive transformer blocks during training to minimize tracking error, AVTrack employs a block sparsity regularization loss Lspar\mathcal{L}_{spar}. The loss penalizes deviations of the mean layer activation probabilities from a target sparsity setpoint:

    Lspar=∣1N−nf∑i=nf+1Npi−ζ∣\mathcal{L}_{spar} = \left| \frac{1}{N - n_f} \sum_{i=n_f+1}^N p^i - \zeta \right|

    where:

    • NN is the total number of transformer blocks in the Vision Transformer backbone,
    • nfn_f is the number of initial mandatory active transformer blocks,
    • pi∈[0,1]p^i \in [0, 1] is the activation probability computed by the Activation Module for the ii-th transformer block,
    • ζ∈[0,1]\zeta \in [0, 1] is a target block sparsity constant. Lower values of ζ\zeta encourage greater model sparsity by driving the mean activation probability toward zero.
  4. Knowl 4 — View-Invariant Feature Learning via Mutual Information Maximization

    model/method

    To improve tracking robustness against severe viewpoint and camera angle variations inherent to aerial tracking, AVTrack enforces view-invariant representation (VIR) learning by maximizing the mutual information (MI) between feature representations of the same target under different viewpoints.

    Let t1:K∞(Z,X)=tKZ∞(Z,X)∪tKX∞(Z,X)t_{1:K}^\infty(Z, X) = t_{K_Z}^\infty(Z, X) \cup t_{K_X}^\infty(Z, X) be the final token representations output by the backbone, where tKZ∞t_{K_Z}^\infty corresponds to the template image ZZ and tKX∞t_{K_X}^\infty corresponds to the search image XX. Using the ground-truth target bounding box Z′Z' within the search frame during training, linear interpolation extracts the target patch tokens tKZ′∞(Z,X)⊂tKX∞(Z,X)t_{K_{Z'}}^\infty(Z, X) \subset t_{K_X}^\infty(Z, X).

    The mutual information between the template view tKZ∞t_{K_Z}^\infty and search view tKZ′∞t_{K_{Z'}}^\infty is maximized using the Deep InfoMax estimator based on the Jensen-Shannon Divergence (JSD). The VIR loss Lvir\mathcal{L}_{vir} is formulated as: Lvir=−I^Θ(JSD)(tKZ′∞(Z,X),tKZ∞(Z,X))\mathcal{L}_{vir} = - \hat{I}_\Theta^{(JSD)}(t_{K_{Z'}}^\infty(Z, X), t_{K_Z}^\infty(Z, X)) where the estimator I^Θ(JSD)(a,b)\hat{I}_\Theta^{(JSD)}(a, b) is defined as: I^Θ(JSD)(a,b)=Ep(a,b)[−α(−TΘ(a,b))]−Ep(a)p(b)[α(TΘ(a,b))]\hat{I}_\Theta^{(JSD)}(a, b) = \mathbb{E}_{p(a, b)}[-\alpha(-T_\Theta(a, b))] - \mathbb{E}_{p(a)p(b)}[\alpha(T_\Theta(a, b))] with α(z)=log⁡(1+ez)\alpha(z) = \log(1 + e^z) denoting the softplus function, TΘT_\Theta denoting a neural network discriminator parameterized by Θ\Theta, p(a,b)p(a, b) denoting the joint distribution, and p(a)p(b)p(a)p(b) denoting the product of marginal distributions.

    Because Lvir\mathcal{L}_{vir} is computed only during model training, it introduces zero computational overhead during real-time tracking inference.

  5. Knowl 5 — Multi-Task Training Objective for AVTrack

    equation

    The overall end-to-end multi-task training objective Ltotal\mathcal{L}_{total} for AVTrack combines classification, bounding box regression, block sparsity, and view-invariance losses:

    Ltotal=Lcls+λiouLiou+λL1LL1+γLspar+κLvir\mathcal{L}_{total} = \mathcal{L}_{cls} + \lambda_{iou}\mathcal{L}_{iou} + \lambda_{L1}\mathcal{L}_{L1} + \gamma \mathcal{L}_{spar} + \kappa \mathcal{L}_{vir}

    where:

    • Lcls\mathcal{L}_{cls} is the weighted focal loss applied to target classification,
    • Liou\mathcal{L}_{iou} is the Generalized Intersection over Union (GIoU) loss for bounding box localization,
    • LL1\mathcal{L}_{L1} is the L1L_1 bounding box coordinate regression loss,
    • Lspar\mathcal{L}_{spar} is the block sparsity loss,
    • Lvir\mathcal{L}_{vir} is the Jensen-Shannon mutual information maximization loss for view-invariant representation learning,
    • Hyperparameter coefficients are set to λiou=2\lambda_{iou} = 2, λL1=5\lambda_{L1} = 5, γ=50\gamma = 50, and κ=1×10−4\kappa = 1 \times 10^{-4} (0.00010.0001).
  6. Knowl 6 — Benchmark Evaluation on Lightweight Aerial Trackers

    data/table

    The tracking performance of AVTrack variants (AVTrack-ViT, AVTrack-EVA, and AVTrack-DeiT) was evaluated against 13 lightweight UAV tracking methods across five standard benchmarks: DTB70, UAVDT, VisDrone2018, UAV123, and UAV123@10fps. Metrics include Precision (Prec. %), Success Rate (Succ. %), GPU FPS, and CPU FPS measured on an Intel i9-10850K CPU and NVIDIA TitanX GPU.

    Tracker DTB70 UAVDT VisDrone2018 UAV123 Avg. Avg. FPS
    Prec. Succ. Prec. Succ. Prec. Succ. Prec. Succ. Prec. Succ. GPU CPU
    AVTrack-ViT 81.3 63.3 79.9 57.7 86.4 65.9 84.0 66.2 82.9 63.8 250.2 59.7
    AVTrack-EVA 82.6 64.0 78.8 57.2 83.4 62.5 83.0 64.7 81.8 62.3 283.7 62.8
    AVTrack-DeiT 84.3 65.0 82.1 58.7 86.0 65.3 84.8 66.8 84.1 64.3 256.8 59.5
    Aba-ViTrack 85.9 66.4 83.4 59.9 86.1 65.3 86.4 66.4 85.3 64.7 181.5 50.3
    LiteTrack 82.5 63.9 81.6 59.3 79.7 61.4 84.2 65.9 82.2 63.1 141.6 -
    SMAT 81.9 63.8 80.8 58.7 82.5 63.4 81.8 64.6 81.5 62.8 124.2 -
    HiFT 80.2 59.4 65.2 47.5 71.9 52.6 78.7 59.0 74.2 55.1 160.3 -
    TCTrack 81.2 62.2 72.5 53.0 79.9 59.4 80.0 60.5 78.3 59.0 139.6 -
    UDAT 80.6 61.8 80.1 59.2 81.6 61.9 76.1 59.0 79.2 60.1 33.7 -
    SGDViT 78.5 60.4 65.7 48.0 72.1 52.1 75.4 57.5 75.6 56.8 110.5 -
    ABDNet 76.8 59.6 75.5 55.3 75.0 57.2 79.3 60.7 76.7 59.1 130.2 -
    DRCI 81.4 61.8 84.0 59.0 83.4 60.0 76.7 59.7 79.8 59.1 281.3 62.4
    AutoTrack 71.6 47.8 71.8 45.0 78.8 57.3 68.9 47.2 71.6 49.0 - 57.8
    RACF 72.6 50.5 77.3 49.4 83.4 60.0 70.2 47.7 74.6 81.2 - 35.6

    AVTrack-DeiT achieves an average precision of 84.1% and success rate of 64.3% while running at 256.8 GPU FPS and 59.5 CPU FPS. AVTrack-EVA reaches the highest GPU tracking speed at 283.7 FPS. All AVTrack variants maintain real-time tracking performance on a single CPU (>59 FPS).

  7. Knowl 7 — Comparative Performance Against SOTA Deep Trackers

    data/table

    AVTrack-DeiT was compared with 14 state-of-the-art generic deep visual trackers on VisDrone2018 as well as across four UAV tracking datasets (DTB70, UAVDT, VisDrone2018, and UAV123@10fps).

    Tracker VisDrone2018 Prec. VisDrone2018 Succ. VisDrone2018 FPS 4-Dataset Avg. (Prec., Succ.) 4-Dataset Avg. FPS
    AVTrack-DeiT (Ours) 86.0 65.3 220.0 (83.9, 63.7) 253.4
    ROMTrack 86.3 66.7 51.1 (85.1, 65.9) 53.1
    SeqTrack 85.3 65.8 11.0 (84.6, 64.8) 17.6
    SLT-TransT 85.6 65.3 29.5 (84.5, 65.2) 32.6
    TransT 85.9 65.2 51.7 (84.2, 65.4) 55.0
    OSTrack 84.2 64.8 66.0 (84.4, 65.8) 65.8
    ToMP 84.1 64.4 21.6 (85.7, 65.8) 23.8
    KeepTrack 84.0 63.5 18.7 (85.2, 64.1) 20.3
    MAT 81.6 62.2 71.2 (81.8, 62.5) 72.3
    SparseTT 81.4 62.1 28.3 (82.2, 64.6) 31.5
    AutoMatch 78.1 59.6 62.1 (81.9, 62.3) 35.2
    PrDiMP50 79.4 59.7 41.3 (81.6, 59.7) 42.3
    SiamRPN++ 79.1 60.0 55.1 (79.9, 60.5) 57.6

    While achieving precision within 0.3% and success rate within 1.4% of top-performing heavy deep trackers on VisDrone2018 (e.g., ROMTrack at 86.3% Prec. and 66.7% Succ.), AVTrack-DeiT runs at 220.0 FPS, operating 4.3×\times faster than ROMTrack (51.1 FPS), 7.5×\times faster than SLT-TransT (29.5 FPS), and over 10×\times faster than ToMP (21.6 FPS) and KeepTrack (18.7 FPS).

  8. Knowl 8 — Ablation and Generalizability Analysis of Activation Module and View-Invariant Representations

    empirical result

    Ablation experiments on VisDrone2018 demonstrate the individual and combined effects of integrating the View-Invariant Representations (VIR) loss and the Activation Module (AM) across different Vision Transformer backbones and existing state-of-the-art visual trackers:

    1. Backbone Ablations on VisDrone2018:

      • AVTrack-ViT: Baseline achieves 83.0% Precision (Prec.), 62.7% Success Rate (Succ.) at 188.3 FPS. Adding VIR increases Prec. to 87.1% (+4.1%) and Succ. to 66.4% (+3.7%). Adding AM achieves 86.4% Prec., 65.9% Succ., and accelerates GPU tracking speed to 238.6 FPS (+26.0% speedup).
      • AVTrack-EVA: Baseline achieves 79.7% Prec., 60.7% Succ. at 235.6 FPS. Adding VIR boosts Prec. to 84.3% (+4.6%) and Succ. to 63.2% (+2.5%). Adding AM achieves 83.4% Prec., 62.5% Succ. at 285.6 FPS (+21.0% speedup).
      • AVTrack-DeiT: Baseline achieves 82.3% Prec., 63.2% Succ. at 192.7 FPS. Adding VIR improves Prec. to 86.7% (+4.4%) and Succ. to 65.9% (+2.7%). Adding AM yields 86.0% Prec., 65.3% Succ. at 220.0 FPS (+14.0% speedup).
    2. Application to External SOTA Trackers (with ViT-Tiny backbones):

      • ARTrack: Baseline (77.7% Prec., 59.5% Succ., 77.5 FPS) →\to +VIR (79.9% Prec., 60.8% Succ.) →\to +VIR+AM (79.3% Prec., 60.4% Succ., 94.5 FPS, +22% speedup).
      • DropTrack: Baseline (81.5% Prec., 62.7% Succ., 177.4 FPS) →\to +VIR (83.1% Prec., 64.0% Succ.) →\to +VIR+AM (82.7% Prec., 63.6% Succ., 214.5 FPS, +21% speedup).
      • GRM: Baseline (82.7% Prec., 63.4% Succ., 198.5 FPS) →\to +VIR (84.1% Prec., 64.5% Succ.) →\to +VIR+AM (83.7% Prec., 64.1% Succ., 234.7 FPS, +18% speedup).

    Across all architectures, VIR consistently improves precision by 1.4%–4.6%, while AM introduces a 14%–26% speedup with minimal degradation in tracking accuracy (<0.7% drop).

  9. Knowl 9 — Sensitivity Analysis of Mandatory Active Layers and View-Invariance Loss Weight

    empirical result

    Hyperparameter evaluations on AVTrack across five UAV tracking benchmarks (DTB70, UAVDT, VisDrone2018, UAV123, UAV123@10fps) reveal key sensitivity dynamics:

    1. Mandatory Initial Layers (nfn_f): When varying the number of initial mandatory active transformer blocks nfn_f from 1 to 8 in AVTrack-DeiT:

      • nf=1n_f = 1 yields 84.1% Avg. Precision, 64.3% Avg. Success, and 256.8 GPU FPS.
      • Increasing nfn_f to 8 modestly improves Avg. Precision to 84.8% and Avg. Success to 64.7%, but severely degrades tracking speed down to 148.5 GPU FPS.
      • Each single increment in nfn_f causes a >5%>5\% reduction in FPS. Setting nf=1n_f = 1 provides the optimal trade-off between inference frame rate and accuracy.
    2. View-Invariant Loss Weight (κ\kappa): Varying κ\kappa across the range [0.5,1.5]×10−4[0.5, 1.5] \times 10^{-4}:

      • Setting κ=1.0×10−4\kappa = 1.0 \times 10^{-4} achieves the best overall performance (84.1% Avg. Precision, 64.3% Avg. Success).
      • Deviations to suboptimal weights (e.g., κ=0.5×10−4\kappa = 0.5 \times 10^{-4}) lead to performance drops of up to 1.6% in Precision (82.5%) and 1.2% in Success Rate (63.1%).
  10. Knowl 10 — Real-Time Onboard UAV Deployment on Embedded Hardware

    empirical result

    AVTrack-DeiT was validated in physical flight experiments by deploying the tracking pipeline onto an embedded NVIDIA Jetson AGX Xavier (32GB) onboard processor mounted on a standard UAV platform.

    • AVTrack-DeiT maintained an average tracking speed of 42.4 FPS with 39.7% GPU utilization and 13.5% CPU utilization during active flight.
    • Baseline Tracker (AVTrack-DeiT without AM and VIR) operated at an average speed of 36.7 FPS with 43.3% GPU utilization and 16.7% CPU utilization.

    The real-world embedded flight tests confirm that AVTrack achieves lower hardware utilization, runs 15.5% faster on embedded edge hardware, and tracks accurately under severe viewpoint changes during continuous UAV operations.

Coverage note — None was omitted; all contributed models, loss formulations, benchmark results, ablations, hyperparameter studies, and hardware deployment evaluations are covered.

References

  1. 1.Bertinetto, L., Valmadre, J., and et al. Staple: Complementary learners for real-time tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  2. 2.Bracci, S., Caramazza, A., and Peelen, M. V. View-invariant representation of hand postures in the human lateral occipitotemporal cortex. NeuroImage, 2018.
  3. 3.Cai, Y., Liu, J., Tang, J., and Wu, G. Robust object modeling for visual tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
  4. 4.Cao, Z., Fu, C., Ye, J., Li, B., and Li, Y. Hift: Hierarchical feature transformer for aerial tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  5. 5.Cao, Z., Huang, Z., Pan, L., Zhang, S., Liu, Z., and Fu, C. Tctrack: Temporal contexts for aerial tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  6. 6.Chen, B., Li, P., Bai, L., Qiao, L., Shen, Q., Li, B., Gan, W., Wu, W., and Ouyang, W. Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking. In European Conference on Computer Vision (ECCV), 2022.
  7. 7.Chen, X., Yan, B., Zhu, J., Wang, D., Yang, X., and Lu, H. Transformer tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021a.
  8. 8.Chen, X., Peng, H., Wang, D., Lu, H., and Hu, H. Seqtrack: Sequence to sequence learning for visual object tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  9. 9.Chen, Y., Dai, X., Chen, D., Liu, M., Dong, X., Yuan, L., and Liu, Z. Mobile-former: Bridging mobilenet and transformer. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021b.
  10. 10.Cui, Y., Jiang, C., Wang, L., and Wu, G. Mixformer: End-to-end tracking with iterative mixed attention. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  11. 11.Danelljan, M., Khan, F. S., Felsberg, M., and van de Weijer, J. Adaptive color attributes for real-time visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014.
  12. 12.Danelljan, M., Bhat, G., Shahbaz Khan, F., and Felsberg, M. Eco: Efficient convolution operators for tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  13. 13.Danelljan, M., Hager, G., Khan, F. S., and Felsberg, M. Discriminative scale space tracking. IEEE transactions on pattern analysis and machine intelligence (TPAMI), 2017.
  14. 14.Danelljan, M., Gool, L. V., and Timofte, R. Probabilistic regression for visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  15. 15.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  16. 16.Du, D., Qi, Y., Yu, H., Yang, Y.-F., Duan, K., Li, G., Zhang, W., Huang, Q., and Tian, Q. The unmanned aerial vehicle benchmark: Object detection and tracking. In European Conference on Computer Vision (ECCV), 2018.
  17. 17.Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., and Ling, H. Lasot: A high-quality benchmark for large-scale single object tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  18. 18.Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y. Eva: Exploring the limits of masked visual representation learning at scale. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  19. 19.Feng, C., Jie, Z., Zhong, Y., Chu, X., and Ma, L. Aedet: Azimuth-invariant multi-view 3d object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  20. 20.Fu, Z., Fu, Z., Liu, Q., Cai, W., and Wang, Y. Sparsett: Visual tracking with sparse transformers. In International Joint Conference on Artificial Intelligence (IJCAI), 2022.
  21. 21.Gao, L., Ji, Y., Gedamu, K., Zhu, X., Xu, X., and Shen, H. T. View-invariant human action recognition via view transformation network (vtn). IEEE Transactions on Multimedia, 2022.
  22. 22.Gao, S., Zhou, C., and Zhang, J. Generalized relation modeling for transformer tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  23. 23.Gopal, G. Y. and Amer, M. A. Separable self and mixed attention transformers for efficient object tracking. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2024.
  24. 24.Guo, D., Shao, Y., Cui, Y., Wang, Z., Zhang, L., and Shen, C. Graph attention tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  25. 25.Henriques, J. F., Caseiro, R., and et al. High-speed tracking with kernelized correlation filters. IEEE transactions on pattern analysis and machine intelligence (TPAMI), 2015.
  26. 26.Huang, L., Zhao, X., and Huang, K. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE transactions on pattern analysis and machine intelligence (PAMI), 2021.
  27. 27.Huang, Z., Fu, C., and et al. Learning aberrance repressed correlation filters for real-time uav tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
  28. 28.Ji, X. and Liu, H. Advances in view-invariant human motion analysis: A review. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 2010.
  29. 29.Kim, M., Lee, S., Ok, J., Han, B., and Cho, M. Towards sequence-level training for visual tracking. In European Conference on Computer Vision (ECCV), 2022.
  30. 30.Kumie, G. A., Habtie, M. A., Ayall, T. A., Zhou, C., Liu, H., Seid, A. M., and Erbad, A. Dual-attention network for view-invariant action recognition. Complex & Intelligent Systems, 2024.
  31. 31.Law, H. and Deng, J. Cornernet: Detecting objects as paired keypoints. In European conference on computer vision (ECCV), 2018.
  32. 32.Li, B., Wu, W., and et al. Siamrpn++: Evolution of siamese visual tracking with very deep networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  33. 33.Li, C., Min, X., Sun, S., Lin, W., and Tang, Z. Deepgait: A learning deep convolutional representation for view-invariant gait recognition using joint bayesian. Applied Sciences, 2017.
  34. 34.Li, S. and Yeung, D. Y. Visual object tracking for unmanned aerial vehicles: A benchmark and new motion models. In AAAI Conference on Artificial Intelligence (AAAI), 2017.
  35. 35.Li, S., Jiang, Q., Zhao, Q., Lu, L., and Feng, Z. Asymmetric discriminative correlation filters for visual tracking. Frontiers of Information Technology & Electronic Engineering, 21(10):1467–1484, 2020a.
  36. 36.Li, S., Liu, Y., Zhao, Q., and Feng, Z. Learning residue-aware correlation filters and refining scale estimates with the grabcut for real-time uav tracking. In 2021 International Conference on 3D Vision (3DV), pp. 1238–1248. IEEE, 2021a.
  37. 37.Li, S., Zhao, Q., Feng, Z., and Lu, L. Equivalence of correlation filter and convolution filter in visual tracking. In Image and Graphics, pp. 623–634, Cham, 2021b. Springer International Publishing.
  38. 38.Li, S., Liu, Y., Zhao, Q., and Feng, Z. Learning residue-aware correlation filters and refining scale for real-time uav tracking. Pattern Recognition (PR), 2022a.
  39. 39.Li, S., Yang, Y., Zeng, D., and Wang, X. Adaptive and background-aware vision transformer for real-time uav tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
  40. 40.Li, Y. and Zhu, J. A scale adaptive kernel correlation filter tracker with feature integration. In European Conference on Computer Vision (ECCV), 2015.
  41. 41.Li, Y., Fu, C., and et al. Autotrack: Towards high-performance visual tracking for uav with automatic spatio-temporal regularization. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020b.
  42. 42.Li, Y., Yuan, G., Wen, Y., Hu, E., Evangelidis, G., Tulyakov, S., Wang, Y., and Ren, J. Efficientformer: Vision transformers at mobilenet speed. In Advances in Neural Information Processing Systems (NeurIPS), 2022b.
  43. 43.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European Conference on Computer Vision (ECCV), 2014.
  44. 44.Liu, M., Wang, Y., Sun, Q., and Li, S. Global filter pruning with self-attention for real-time uav tracking. In British Machine Vision Conference (BMVC), 2022a.
  45. 45.Liu, Z., Feng, R., Chen, H., Wu, S., Gao, Y., Gao, Y., and Wang, X. Temporal feature alignment and mutual information maximization for video-based human pose estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022b.
  46. 46.M., D. and et al. Adaptive decontamination of the training set: A unified formulation for discriminative visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  47. 47.Ma, S., Liu, Y., Zeng, D., Liao, Y., Xu, X., and Li, S. Learning disentangled representation in pruning for real-time uav tracking. In Asian Conference on Machine Learning (ACML), 2023.
  48. 48.MacKay, D. J. Information theory, inference and learning algorithms. Cambridge university press, 2004.
  49. 49.Mao, J., Yang, H., Li, A., Li, H., and Chen, Y. Tprune: Efficient transformer pruning for mobile devices. ACM Transactions on Cyber-Physical Systems (TCPS), 2021.
  50. 50.Mayer, C., Danelljan, M., Paudel, D. P., and Van Gool, L. Learning target candidate association to keep track of what not to track. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  51. 51.Mayer, C., Danelljan, M., Bhat, G., Paul, M., Paudel, D. P., Yu, F., and Gool, L. V. Transforming model prediction for tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  52. 52.Mueller, M., Smith, N., and Ghanem, B. A benchmark and simulator for uav tracking. In European Conference on Computer Vision (ECCV), 2016.
  53. 53.Mueller, M., Smith, N., and Ghanem, B. Context-aware correlation filter tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  54. 54.Muller, M., Bibi, A., Giancola, S., Alsubaihi, S., and Ghanem, B. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In European Conference on Computer Vision (ECCV), 2018.
  55. 55.Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G. On variational bounds of mutual information. In International Conference on Machine Learning (ICML), 2019.
  56. 56.Rao, C., Yilmaz, A., and Shah, M. View-invariant representation and recognition of actions. International journal of computer vision (IJCV), 2002.
  57. 57.Rao, Y., Zhao, W., Liu, B., Lu, J., Zhou, J., and Hsieh, C.-J. Dynamicvit: Efficient vision transformers with dynamic token sparsification. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  58. 58.R.D., H. and et al. Learning deep representations by mutual information estimation and maximization. In International Conference on Learning Representations (ICLR), 2019.
  59. 59.Rezatofighi, S. H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I. D., and Savarese, S. Generalized intersection over union: A metric and a loss for bounding box regression. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  60. 60.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L. Imagenet large scale visual recognition challenge. International journal of computer vision (IJCV), 2014.
  61. 61.Shiraga, K., Makihara, Y., Muramatsu, D., Echigo, T., and Yagi, Y. Geinet: View-invariant gait recognition using a convolutional neural network. In International conference on biometrics (ICB), 2016.
  62. 62.Steuer, R., Kurths, J., Daub, C. O., Weise, J., and Selbig, J. The mutual information: detecting and evaluating dependencies between variables. Bioinformatics, 2002.
  63. 63.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning (ICML), 2021.
  64. 64.Wang, N., gang Zhou, W., Tian, Q., Hong, R., Wang, M., and Li, H. Multi-cue correlation filters for robust visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  65. 65.Wang, N., Zhou, W., Wang, J., and Li, H. Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  66. 66.Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768, 2020.
  67. 67.Wang, X., Zeng, D., Zhao, Q., and Li, S. Rank-based filter pruning for real-time uav tracking. In IEEE International Conference on Multimedia and Expo (ICME), 2022.
  68. 68.Wang, X., Yang, X., Ye, H., and Li, S. Learning disentangled representation with mutual information maximization for real-time uav tracking. In IEEE International Conference on Multimedia and Expo (ICME), 2023.
  69. 69.Wei, Q., Zeng, B., Liu, J., He, L., and Zeng, G. Litetrack: Layer pruning with asynchronous feature extraction for lightweight and efficient visual tracking. arXiv preprint arXiv:2309.09249, 2023a.
  70. 70.Wei, X., Bai, Y., Zheng, Y., Shi, D., and Gong, Y. Autoregressive visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023b.
  71. 71.Wu, Q., Yang, T., Liu, Z., Wu, B., Shan, Y., and Chan, A. B. Dropmae: Masked autoencoders with spatial-attention dropout for tracking tasks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  72. 72.Wu, W., Zhong, P., and Li, S. Fisher pruning for real-time uav tracking. In International Joint Conference on Neural Networks (IJCNN), 2022.
  73. 73.Xia, L., Chen, C.-C., and Aggarwal, J. K. View invariant human action recognition using histograms of 3d joints. In IEEE computer society conference on computer vision and pattern recognition workshops (CVPRW), 2012.
  74. 74.Xie, F., Wang, C., Wang, G., Yang, W., and Zeng, W. Learning tracking representations via dual-branch fully transformer networks. In IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2021.
  75. 75.Xie, F., Wang, C., Wang, G., Cao, Y., Yang, W., and Zeng, W. Correlation-aware deep tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  76. 76.Yang, X., Yan, J., Cheng, Y., and Zhang, Y. Learning deep generative clustering via mutual information maximization. IEEE Transactions on Neural Networks and Learning Systems (TNNLS), 2022.
  77. 77.Yao, L., Fu, C., and et al. Sgdvit: Saliency-guided dynamic vision transformer for uav tracking. In IEEE International Conference on Robotics and Automation (ICRA), 2023.
  78. 78.Ye, B., Chang, H., Ma, B., Shan, S., and Chen, X. Joint feature learning and relation modeling for tracking: A one-stream framework. In European Conference on Computer Vision (ECCV), 2022a.
  79. 79.Ye, J., Fu, C., Zheng, G., Paudel, D. P., and Chen, G. Unsupervised domain adaptation for nighttime aerial tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022b.
  80. 80.Yin, H., Vahdat, A., Alvarez, J. M., Mallya, A., Kautz, J., and Molchanov, P. A-vit: Adaptive tokens for efficient vision transformer. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  81. 81.Zeng, D., Zou, M., Wang, X., and Li, S. Towards discriminative representations with contrastive instances for real-time uav tracking. In IEEE International Conference on Multimedia and Expo (ICME), 2023.
  82. 82.Zhang, J., Peng, H., Wu, K., Liu, M., Xiao, B., Fu, J., and Yuan, L. Minivit: Compressing vision transformers with weight multiplexing. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022a.
  83. 83.Zhang, Z., Peng, H., Fu, J., Li, B., and Hu, W. Ocean: Object-aware anchor-free tracking. In European Conference on Computer Vision (ECCV), 2020.
  84. 84.Zhang, Z., Liu, Y., Wang, X., Li, B., and Hu, W. Learn to match: Automatic matching network design for visual tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  85. 85.Zhang, Z., Wu, F., Qiu, Y., Liang, J., and Li, S. Tracking small and fast moving objects: A benchmark. In Asian Conference on Computer Vision (ACCV), 2022b.
  86. 86.Zhao, H., Wang, D., and Lu, H. Representation learning for visual object tracking by masked appearance transfer. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  87. 87.Zhong, P., Zeng, D., Wang, X., and Li, S. Efficiency and precision trade-offs in uav tracking with filter pruning and dynamic channel weighting. In Fuzzy Systems and Data Mining (FSDM), 2022.
  88. 88.Zhong, P., Wu, W., Dai, X., Zhao, Q., and Li, S. Fisher pruning for developing real-time uav trackers. Journal of Real-Time Image Processing, 2023.
  89. 89.Zhou, Z., Pei, W., Li, X., Wang, H., Zheng, F., and He, Z. Saliency-associated object tracking. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  90. 90.Zhu, P., Wen, L., and et al. Visdrone-sot2018: The vision meets drone single-object tracking challenge results. In European Conference on Computer Vision (ECCV), 2018.
  91. 91.Zuo, H., Fu, C., Li, S., Lu, K., Li, Y., and Feng, C. Adversarial blur-deblur network for robust uav tracking. IEEE Robotics and Automation Letters (RAL), 2023.

Citation

MLA
Wu, Y., et al. “Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking”. arXiv, 2024, http://arxiv.org/abs/2412.20002v3.
APA
Wu, Y., Li, Y., Liu, M., Wang, X., Yang, X., Ye, H., Zeng, D., Zhao, Q., & Li, S. (2024). Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking. arXiv. http://arxiv.org/abs/2412.20002v3
Chicago
Wu, Y., Y. Li, M. Liu, et al. 2024. “Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking”. arXiv. http://arxiv.org/abs/2412.20002v3.
Harvard
Wu, Y. et al. (2024) “Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2412.20002v3.
Vancouver
1. Wu Y, Li Y, Liu M, Wang X, Yang X, Ye H, Zeng D, Zhao Q, Li S (2024) Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking. arXiv

BibTeX

@article{wu2024learning,
  title = {Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking},
  author = {Wu, You and Li, Yongxin and Liu, Mengyuan and Wang, Xucheng and Yang, Xiangyang and Ye, Hengzhou and Zeng, Dan and Zhao, Qijun and Li, Shuiwang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2412.20002v3},
  eprint = {2412.20002}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/