Single-Model and Any-Modality for Video Object Tracking
Zongwei WuJilai ZhengXiangxuan RenFlorin-Alexandru VasluianuChao MaDanda Pani PaudelLuc Van GoolRadu Timofte
Presents Un-Track, a unified transformer framework that tracks objects across diverse auxiliary modalities like depth, thermal, and event data using a single parameter set by learning a shared low-rank embedding from paired RGB-X inputs with minimal computational overhead.
Standard video tracking systems rely heavily on standard visual cameras, which often fail in real-world conditions like total darkness, heavy occlusions, or rapid movement. While adding auxiliary sensors—such as thermal imaging, depth sensors, or high-speed event cameras—solves these failure modes, current solutions require separate, customized software models for each sensor type. This requirement significantly increases operational costs, maintenance overhead, and system memory footprints.
The article evaluates a unified framework, named Un-Track, designed to handle any auxiliary sensor using a single set of model parameters. The primary objective is to demonstrate that a single tracker can adapt to depth, thermal, or event inputs during live deployment without needing specialized per-sensor retraining or fine-tuning.
The authors develop an approach combining explicit edge cues with low-rank mathematical factorization to project disparate sensor streams into a shared common space. This design relies on a lightweight prompting mechanism that enhances uncertain visual features using auxiliary data while keeping the core pre-trained vision transformer model frozen and adapting it with parameter-efficient fine-tuning. The framework was trained exclusively on paired two-stream data and evaluated across five standard benchmark datasets encompassing diverse sensor domains.
The evaluation yields several key findings. First, Un-Track sets new performance records on major benchmark datasets, achieving an F-score of 0.610 on the DepthTrack test set and outperforming specialized, sensor-specific models. Second, the single-model variant outperforms previous state-of-the-art unified architectures across depth, thermal, and event benchmarks, showing notable absolute precision gains such as a 3.8% increase on thermal benchmarks. Third, this performance is achieved with minimal computational overhead, adding only 6.6 million parameters to the 92-million-parameter baseline and increasing compute by less than 10%. Finally, the system demonstrates strong cross-modal generalization, leveraging geometric and motion priors during thermal tracking and maintaining superior tracking accuracy even when auxiliary sensor data is entirely missing.
These results show that engineering teams can consolidate multiple tracking pipelines into a single deployable model, lowering memory requirements and reducing training infrastructure demands. Deploying a single unified model also eliminates operational failure risks caused by missing or malfunctioning sensor feeds in complex multi-sensor hardware systems.
Organizations developing autonomous systems, security platforms, or robotics should evaluate unified multimodal tracking architectures to streamline sensor integration. Before enterprise-scale deployment, teams should conduct pilot testing under dynamic operational conditions where auxiliary sensors may disconnect or experience degraded signals. The findings are backed by consistent experimental evidence across multiple benchmark datasets, offering high confidence in the framework's effectiveness across standard visual, thermal, depth, and event modalities.
- Paper: FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space, Shengzhong Liu et al. (2023). Introduces the foundational framework for learning factorized, shared latent representations across multimodal sensing signals that directly underpins the shared latent space formulation used in Un-Track.
- Paper: MixFormerV2: Efficient Fully Transformer Tracking, Yutao Cui et al. (2023). Establishes modern, fully transformer-based single-stream tracking architectures that Un-Track builds upon for its unified multi-modal tracking framework.
- Paper: MixFormer: End-to-End Tracking with Iterative Mixed Attention, Yutao Cui et al. (2022). Presents end-to-end attention-driven target and search feature fusion in vision transformers, providing architectural context for single-model transformer tracking.
- Paper: Transformer Tracking, Xin Chen et al. (2021). Pioneers attention-based feature interaction in visual tracking, establishing the core transformer foundation for subsequent single-model trackers.
- Paper: U2Fusion: A Unified Unsupervised Image Fusion Network, Han Xu et al. (2020). Provides a unified deep model for diverse multimodal sensor fusion without task-specific tuning, establishing key concepts for handling heterogeneous sensory inputs.
- Paper: Unifying Visual and Vision-Language Tracking via Contrastive Learning, Yinchao Ma et al. (2024). Extends the concept of unified, single-model tracking across disparate input modalities by unifying visual bounding boxes and natural language descriptions via contrastive learning.
- Paper: Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking, Yongxin Li et al. (2024). Applies adaptive vision transformer tracking principles to resource-constrained real-time UAV deployment using dynamic token activation.
