Unifying Visual and Vision-Language Tracking via Contrastive Learning
Yinchao MaYuyang TangWenfei YangTianzhu ZhangJinpeng ZhangMengxue Kang
Proposes a unified tracking framework, UVLTrack, that utilizes multi-modal contrastive learning and a dynamic box head to track targets across bounding box, natural language, and combined reference settings using a single set of parameters.
Target tracking in video is essential for autonomous systems, robotics, and intelligent surveillance. In practical deployments, user inputs specifying what to track vary widely: systems may receive an initial visual bounding box, a natural language text description, or a combination of both. Historically, computer vision tracking models have specialized in only one or two of these input types. Because of the semantic gap between image features and language representations, trackers designed for natural language often struggle when provided only with bounding boxes, while visual-only trackers cannot utilize text to resolve visual ambiguity.
The article demonstrates that a single, unified deep learning architecture can achieve state-of-the-art tracking performance across all three input modalities (bounding box, natural language, and language plus bounding box) simultaneously without changing network parameters. To accomplish this, the authors introduce "UVLTrack," a framework combining a modality-unified feature extractor with a modality-adaptive box head.
The authors evaluated the framework across thirteen public benchmarks, covering seven visual tracking datasets, three vision-language tracking datasets, and three visual grounding benchmarks. The method isolates low-level visual and textual processing in early Transformer layers before merging them in deeper layers, aligning these representations using a multi-modal contrastive loss. To avoid rigid target estimation, the dynamic box head samples historical video frames to distinguish true targets from potential visual distractors and background elements.
The evaluation produced several key findings: First, UVLTrack established state-of-the-art results on all seven visual tracking benchmarks, demonstrating that cross-modal capabilities do not degrade standard visual performance. Second, under pure natural language tracking, the larger model variant (UVLTrack-L) outperformed the previous best model, JointNLT, by significant margins across all benchmarks, including an improvement in tracking success from 54.6% to 58.2% on the TNL2K benchmark. Third, the base model (UVLTrack-B) operated at 57 to 58 frames per second, running approximately 1.46 times faster than JointNLT while improving accuracy. Finally, ablation studies showed that aligning features with contrastive loss yielded consistent gains of 1.2% to 2.3% across all input modalities, while dynamic distractor modeling added up to 2.1% in success rates over static detection heads.
These results show that engineering teams do not need to maintain multiple specialized tracking pipelines for different operational inputs. A single, shared architecture reduces computational maintenance costs, streamlines model lifecycle management, and increases system robustness in real-time edge or server environments.
Organizations developing vision-based tracking systems should consider adopting unified Transformer architectures that integrate contrastive alignment and dynamic distractor modeling. To maximize performance, implementations should initialize textual encoders with dedicated pre-trained language models rather than relying solely on visual pre-training. While the model shows high confidence across benchmark scenarios, real-world deployment should be preceded by domain-specific pilot testing, particularly under severe environmental visibility constraints or extreme computational budget limits.
- Paper: Transformer Tracking, Xin Chen et al. (2021). This foundational work demonstrates how attention-based Transformer architectures replace traditional correlation modules for feature fusion in visual tracking, establishing the tracking paradigm that UVLTrack builds upon.
- Paper: MixFormer: End-to-End Tracking with Iterative Mixed Attention, Yutao Cui et al. (2022). It introduces mixed attention for simultaneous feature extraction and target integration, serving as an architectural precursor to unified Transformer tracking frameworks.
- Paper: MixFormerV2: Efficient Fully Transformer Tracking, Yutao Cui et al. (2023). It establishes fully Transformer-based tracking architectures with unified prediction tokens, providing key context for modern Transformer tracker design.
- Paper: End-to-End Referring Video Object Segmentation with Multimodal Transformers, Adam Botach et al. (2022). It explores multimodal Transformer architectures for tracking and segmenting targets from natural language queries, motivating cross-modal alignment in tracking.
- Paper: LaSOT: A High-Quality Benchmark for Large-Scale Single Object Tracking, Heng Fan et al. (2018). This benchmark paper establishes large-scale visual tracking datasets that include natural language descriptions, providing the empirical foundation for vision-language tracking evaluations.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). It introduces universal cross-modal Transformer pre-training with contrastive and matching objectives, establishing principles for aligning image and text representations utilized in UVLTrack.
- Paper: ATOM: Accurate Tracking by Overlap Maximization, Martin Danelljan et al. (2018). It introduces target bounding box estimation decoupled from online classification to prevent tracking drift, influencing the design of dynamic distractor-aware box heads.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This comprehensive survey extends unified vision-language understanding into embodied robotic control and decision-making within Vision-Language-Action frameworks.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). It broadens multimodal video-language pre-training by decoupling visual keyframes and motion tokens for generative and comprehending multimodal large models.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). It scales unified multimodal reasoning and spatial-temporal grounding to large language model architectures across extensive visual and video domains.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It investigates native multimodal pre-training recipes that jointly optimize visual grounding and linguistic reasoning at scale in large language models.
