Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking
Yongxin LiMengyuan LiuYou WuXucheng WangXiangyang YangShuiwang Li
Presents AVTrack, an efficient UAV tracking framework that combines dynamic transformer block activation with mutual information maximization to achieve real-time tracking speeds of roughly 220 FPS while resisting severe viewing angle variations.
Deploying computer vision models on unmanned aerial vehicles (UAVs) requires balancing high tracking accuracy with extreme computational efficiency, as drones operate under strict onboard processing and battery constraints. While modern vision transformer models deliver superior tracking precision in complex aerial environments, their computational overhead has largely prevented reliable real-time deployment on lightweight drone hardware.
The article develops and evaluates AVTrack, an adaptive and view-invariant vision transformer tracking framework designed specifically for real-time UAV applications.
The approach introduces two key mechanisms into a single-stream transformer architecture: an Activation Module that dynamically skips unnecessary computational blocks based on scene complexity, and a view-invariant representation learning loss that maximizes mutual information between different perspectives of the target during training without adding inference costs. The authors conducted extensive evaluations across five standard aerial tracking benchmarks—DTB70, UAVDT, VisDrone2018, UAV123, and UAV123@10fps—benchmarking against 13 lightweight trackers and 14 deep tracking systems, followed by embedded deployment on an onboard NVIDIA Jetson AGX Xavier platform.
The findings demonstrate substantial improvements across efficiency, accuracy, and generalizability. AVTrack achieved real-time speeds on a standard computer, reaching between 250 and 283 frames per second (FPS) on a graphics processing unit (GPU) and approximately 60 FPS on a central processing unit (CPU), running over 1.4 times faster than previous adaptive transformer trackers. In accuracy, AVTrack delivered state-of-the-art results, achieving an average precision of 84.1% and a success rate of 64.3% across benchmarks, outperforming conventional correlation filters and lightweight deep networks. Crucially, learning view-invariant representations boosted baseline precision by roughly 2.5% to 4.6% across models, directly resolving drone-specific challenges such as severe viewing angle shifts. When integrated into other leading tracking frameworks, the core modules consistently boosted inference speeds by 14% to 22% with negligible impact on accuracy, and real-world embedded flight testing confirmed smooth operation at 42.4 FPS with low resource utilization (39.7% GPU and 13.5% CPU).
These results establish that structured block-level conditional computation and mutual-information-based training allow high-capacity transformer models to run efficiently on resource-constrained aerial hardware without sacrificing tracking robustness. This reduces the risk of target loss during sharp drone maneuvers, improves battery life through lower computational loads, and eliminates the traditional trade-off between tracking accuracy and operational speed.
Organizations developing autonomous drone systems should adopt structured conditional activation and view-invariant loss objectives when deploying onboard vision transformers. Engineering teams can integrate these plug-and-play components into existing tracker architectures or deploy AVTrack directly on embedded edge hardware, tuning the initial layer activation parameter depending on specific platform speed-versus-accuracy requirements.
While empirical confidence is high across standard public datasets and the tested Jetson Xavier platform, real-time performance guarantees remain subject to the capabilities of specific onboard hardware. Further testing across broader environmental extremes, such as nighttime conditions and highly degraded weather, is recommended prior to mission-critical operational deployment.
- Paper: MixFormerV2: Efficient Fully Transformer Tracking, Yutao Cui et al. (2023). MixFormerV2 establishes the single-stream, fully transformer visual tracking paradigm and model reduction strategies that AVTrack adapts and accelerates for UAV-specific deployment.
- Paper: Transformer Tracking, Xin Chen et al. (2021). TransT introduces the core mechanism of replacing correlation operations with attention-driven transformer architectures for visual object tracking.
- Paper: A Benchmark and Simulator for UAV Tracking, Matthias Mueller et al. (2016). This paper establishes the foundational UAV123 benchmark dataset and simulation methodologies used to evaluate real-time aerial tracking performance.
- Paper: SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks, Bo Li et al. (2018). SiamRPN++ outlines the critical challenges and architectural techniques for training deep backbone networks for high-precision real-time visual tracking.
- Paper: Learning Discriminative Model Prediction for Tracking, Goutam Bhat et al. (2019). DiMP introduces an end-to-end discriminative formulation for runtime target appearance modeling that informs modern deep tracking baselines.
- Paper: ATOM: Accurate Tracking by Overlap Maximization, Martin Danelljan et al. (2018). ATOM demonstrates the decoupling of bounding box estimation and online target classification, which forms the architectural basis of modern visual tracking frameworks.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). ViT provides the fundamental patch-based vision transformer architecture that serves as the foundation for transformer-based tracking models.
- Paper: Unifying Visual and Vision-Language Tracking via Contrastive Learning, Yinchao Ma et al. (2024). UVLTrack extends the principles of contrastive learning and multi-modal feature adaptation to unify bounding box visual tracking with natural language descriptions.
