Single-Model and Any-Modality for Video Object Tracking

Zongwei WuJilai ZhengXiangxuan RenFlorin-Alexandru VasluianuChao MaDanda Pani PaudelLuc Van GoolRadu Timofte

article2024CVPR168 citations

Presents Un-Track, a unified transformer framework that tracks objects across diverse auxiliary modalities like depth, thermal, and event data using a single parameter set by learning a shared low-rank embedding from paired RGB-X inputs with minimal computational overhead.

Listen

Standard video tracking systems rely heavily on standard visual cameras, which often fail in real-world conditions like total darkness, heavy occlusions, or rapid movement. While adding auxiliary sensors—such as thermal imaging, depth sensors, or high-speed event cameras—solves these failure modes, current solutions require separate, customized software models for each sensor type. This requirement significantly increases operational costs, maintenance overhead, and system memory footprints.

The article evaluates a unified framework, named Un-Track, designed to handle any auxiliary sensor using a single set of model parameters. The primary objective is to demonstrate that a single tracker can adapt to depth, thermal, or event inputs during live deployment without needing specialized per-sensor retraining or fine-tuning.

The authors develop an approach combining explicit edge cues with low-rank mathematical factorization to project disparate sensor streams into a shared common space. This design relies on a lightweight prompting mechanism that enhances uncertain visual features using auxiliary data while keeping the core pre-trained vision transformer model frozen and adapting it with parameter-efficient fine-tuning. The framework was trained exclusively on paired two-stream data and evaluated across five standard benchmark datasets encompassing diverse sensor domains.

The evaluation yields several key findings. First, Un-Track sets new performance records on major benchmark datasets, achieving an F-score of 0.610 on the DepthTrack test set and outperforming specialized, sensor-specific models. Second, the single-model variant outperforms previous state-of-the-art unified architectures across depth, thermal, and event benchmarks, showing notable absolute precision gains such as a 3.8% increase on thermal benchmarks. Third, this performance is achieved with minimal computational overhead, adding only 6.6 million parameters to the 92-million-parameter baseline and increasing compute by less than 10%. Finally, the system demonstrates strong cross-modal generalization, leveraging geometric and motion priors during thermal tracking and maintaining superior tracking accuracy even when auxiliary sensor data is entirely missing.

These results show that engineering teams can consolidate multiple tracking pipelines into a single deployable model, lowering memory requirements and reducing training infrastructure demands. Deploying a single unified model also eliminates operational failure risks caused by missing or malfunctioning sensor feeds in complex multi-sensor hardware systems.

Organizations developing autonomous systems, security platforms, or robotics should evaluate unified multimodal tracking architectures to streamline sensor integration. Before enterprise-scale deployment, teams should conduct pilot testing under dynamic operational conditions where auxiliary sensors may disconnect or experience degraded signals. The findings are backed by consistent experimental evidence across multiple benchmark datasets, offering high confidence in the framework's effectiveness across standard visual, thermal, depth, and event modalities.

arXiv: 2311.15851
Cover for Single-Model and Any-Modality for Video Object Tracking

Abstract

In the realm of video object tracking, auxiliary modalities such as depth, thermal, or event data have emerged as valuable assets to complement the RGB trackers. In practice, most existing RGB trackers learn a single set of parameters to use them across datasets and applications. However, a similar single-model unification for multi-modality tracking presents several challenges. These challenges stem from the inherent heterogeneity of inputs – each with modality-specific representations, the scarcity of multi-modal datasets, and the absence of all the modalities at all times. In this work, we introduce Un-Track, a Unified Tracker of a single set of parameters for any modality. To handle any modality, our method learns their common latent space through low-rank factorization and reconstruction techniques. More importantly, we use only the RGB-X pairs to learn the common latent space. This unique shared representation seamlessly binds all modalities together, enabling effective unification and accommodating any missing modality, all within a single transformer-based architecture. Our Un-Track achieves +8.1 absolute F-score gain, on the DepthTrack dataset, by introducing only +2.14 (over 21.50) GFLOPs with +6.6M (over 93M) parameters, through a simple yet efficient prompting strategy. Extensive comparisons on five benchmark datasets with different modalities show that Un-Track surpasses both SOTA unified trackers and modality-specific counterparts, validating our effectiveness and practicality. The source code is publicly available at https://github.com/Zongwei97/UnTrack.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methods
  • 3.1. Overall Framework
  • 3.2. Shared Embedding
  • 3.3. Outer Modal Prompting
  • 3.4. Inner Finetuning
  • 4. Experiments
  • 4.1. Training Data
  • 4.2. Within distribution Evaluation
  • 4.3. Generalization Across Datasets
  • 5. Ablation Studies
  • 6. Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Un-Track Single-Model Multi-Modal Video Tracking Framework

    model/method

    Un-Track is a unified video object tracking framework that operates across diverse auxiliary modalities (such as depth, thermal, or event data paired with RGB) using a single, uniform set of model parameters without requiring modality-specific fine-tuning.

    The framework is built around a pre-trained transformer RGB tracker backbone (OSTrack) with frozen parameters and comprises three main components:

    1. Shared Embedding Module: Projects heterogeneous auxiliary modalities into a shared, modality-agnostic feature space by combining explicit edge priors with implicit low-rank matrix decomposition and reconstruction.
    2. Outer Modal Prompting Module: Enhances RGB tokens with cross-modal awareness from the shared auxiliary representation at each transformer stage via dynamic token shrinkage and low-rank token recovery.
    3. Inner Fine-Tuning with LoRA: Injects trainable low-rank adaptation matrices into the frozen multi-head attention layers of the RGB transformer to adapt the backbone efficiently.

    During training, Un-Track learns exclusively from mixed pairs of RGB and one auxiliary modality (M={MD,MT,ME}M = \{M^D, M^T, M^E\}) without requiring all modalities to co-occur. During inference, the same set of parameters seamlessly tracks targets with any single available RGB-X modality or in an RGB-only setting.

  2. Knowl 2 — Shared Auxiliary Embedding via Low-Rank Factorization and Edge Guidance

    model/method

    The Shared Embedding module transforms heterogeneous auxiliary inputs (depth DD, thermal TT, event EE) into a unified modality-agnostic representation FF through explicit edge guidance combined with implicit low-rank matrix reconstruction.

    1. Explicit Edge Guidance: Horizontal and vertical spatial gradient maps are computed for both RGB and the auxiliary input by calculating differences between adjacent pixels along the xx- and yy-axes. These gradient maps are merged with visual features to yield an explicit gradient feature GG, from which a low-rank matrix GkG_k is derived.

    2. In-Domain Low-Rank Approximation: For input feature M∈RcM \in \mathbb{R}^c split into domain-specific subset features D,T,ED, T, E from datasets MD,MT,MEM^D, M^T, M^E, modality-specific multilayer perceptrons (MLPs) σd,σt,σe\sigma_d, \sigma_t, \sigma_e project channels to a lower rank dimension kk (k<ck < c): Dk=σd(D),Tk=σt(T),Ek=σe(E)D_k = \sigma_d(D), \quad T_k = \sigma_t(T), \quad E_k = \sigma_e(E)

    3. Global Shared Low-Rank Matrix: The subset low-rank matrices are concatenated along the channel dimension and integrated with the low-rank gradient guidance GkG_k: Mk=φR1([Dk,Tk,Ek])+φR2(Gk)M_k = \varphi_{R_1}([D_k, T_k, E_k]) + \varphi_{R_2}(G_k) where [⋅][ \cdot ] denotes channel concatenation, and φR1,φR2\varphi_{R_1}, \varphi_{R_2} are projection MLPs mapping into the low-rank latent space.

    4. Embedding Reconstruction: The global low-rank representation MkM_k is mapped back to the original feature dimension and added to the full-resolution explicit edge feature GG: F=ΦR(Mk)+GF = \Phi_R(M_k) + G where ΦR\Phi_R is a reconstruction MLP. This ensures both cross-modal geometric alignment and domain-specific feature preservation.

  3. Knowl 3 — Implicit Shared Embedding Reconstruction Algorithm

    algorithm

    The implicit shared embedding algorithm produces a reconstructed, modality-unified feature representation from mixed multimodal inputs and explicit gradient features.

    Input: Mixed multimodal feature tensor MM, explicit gradient binding tensor GG
    Output: Reconstructed unified feature tensor FF
    1. Partition input feature tensor MM into modality-specific subset features DD (depth), TT (thermal), and EE (event).
    2. For each modality x∈{d,t,e}x \in \{d, t, e\}, apply modality-specific mapping σx\sigma_x to project from channel dimension cc to low-rank dimension kk (k<ck < c), yielding in-domain low-rank representations Dk=σd(D)D_k = \sigma_d(D), Tk=σt(T)T_k = \sigma_t(T), and Ek=σe(E)E_k = \sigma_e(E).
    3. Project explicit gradient binding tensor GG into low-rank gradient matrix GkG_k.
    4. Concatenate low-rank representations along the channel dimension as [Dk,Tk,Ek][D_k, T_k, E_k] and project via MLP φR1\varphi_{R_1}.
    5. Project low-rank gradient GkG_k via MLP φR2\varphi_{R_2}.
    6. Compute global shared low-rank matrix Mk=φR1([Dk,Tk,Ek])+φR2(Gk)M_k = \varphi_{R_1}([D_k, T_k, E_k]) + \varphi_{R_2}(G_k).
    7. Project MkM_k back to original feature space via reconstruction MLP ΦR\Phi_R and add original explicit gradient GG: F=ΦR(Mk)+GF = \Phi_R(M_k) + G.
    8. Return reconstructed feature tensor FF.
  4. Knowl 4 — Outer Modal Prompting via Token Shrinkage and Low-Rank Factorization

    model/method

    The outer modal prompting module injects cross-modal cues from the shared auxiliary embedding FF into the primary RGB feature tokens II by categorizing tokens based on reliability and recovering them via low-rank factorization.

    1. Token Partitioning: A dynamic scoring function s(I)s(I) computes confidence scores to partition tokens in II into negative (mnm_n), uncertain (mum_u), and positive (mpm_p) binary masks according to score percentiles (e.g., top 25% positive, bottom 25% negative, and remaining 50% uncertain).

    2. Low-Rank Token Approximation:

    • Negative tokens in II are replaced with corresponding tokens from FF, while positive tokens are retained: Il1=σc(mn⋅F+mp⋅I)I_{l_1} = \sigma_c(m_n \cdot F + m_p \cdot I) where σc\sigma_c maps the modified tokens into a low-rank space of dimension ll.
    • Uncertain tokens from both modalities are combined to filter noise and recover missing cues: Il2=σn(mu⋅F+mu⋅I)I_{l_2} = \sigma_n(m_u \cdot F + m_u \cdot I) where σn\sigma_n is a low-rank approximation function.
    1. Intra-Modality Low-Rank Fusion: The low-rank representations are concatenated and fused: Il=φP([Il1,Il2])I_l = \varphi_P([I_{l_1}, I_{l_2}]) where φP\varphi_P is a learnable MLP.

    2. Cross-Modal Fusion and Reconstruction: The auxiliary embedding FF is processed through the symmetric token shrinkage procedure to produce low-rank matrix FlF_l. The fused output OO in original feature space is: O=ΦP(Il+Fl)O = \Phi_P(I_l + F_l) where ΦP\Phi_P is a reconstruction MLP.

  5. Knowl 5 — Parameter-Efficient Backbone Adaptation via LoRA

    model/method

    To fine-tune a pre-trained transformer tracking backbone without retraining its full weights or overfitting on small multimodal tracking datasets, Low-Rank Adaptation (LoRA) is incorporated into each multi-head self-attention module.

    For a pre-trained attention projection weight matrix W0∈Rd×kW_0 \in \mathbb{R}^{d \times k}, W0W_0 remains frozen while two trainable rank-rr parameter matrices B∈Rd×rB \in \mathbb{R}^{d \times r} and A∈Rr×kA \in \mathbb{R}^{r \times k} (with rank r≪min⁡(d,k)r \ll \min(d, k)) are introduced. Given an input feature vector xx, the updated attention output hh is computed as: h=W0x+BAxh = W_0 x + B A x

    Combined with outer modal prompting, LoRA adaptation allows the model to achieve unified multi-modal tracking with only +6.65M additional trainable parameters over the 92.08M parameter baseline.

  6. Knowl 6 — Within-Distribution Multimodal Tracking Performance on DepthTrack, LasHeR, and VisEvent

    empirical result

    Un-Track was evaluated on three standard multimodal tracking benchmarks: DepthTrack (RGB-D), LasHeR (RGB-T), and VisEvent (RGB-E), under both modality-specific fine-tuning ("X-Specific") and a single unified parameter checkpoint ("Uni-model").

    Method DepthTrack (RGB-D) LasHeR (RGB-T) VisEvent (RGB-E)
    F-score ↑\uparrow Re ↑\uparrow Pr ↑\uparrow PR ↑\uparrow SR ↑\uparrow Precision ↑\uparrow Success ↑\uparrow
    Modality-Specific Parameters
    Stark 0.397 0.406 0.388 0.418 0.333 0.612 0.446
    ProTrack 0.578 0.573 0.583 0.538 0.420 0.632 0.471
    ViPT 0.594 0.596 0.592 0.651 0.525 0.758 0.592
    Un-Track (Specific) 0.612 0.610 0.613 0.667 0.536 0.763 0.597
    Uni-model with a Single Set of Parameters
    AiATrack 0.515 0.526 0.505 0.463 0.365 0.626 0.444
    OSTrack 0.569 0.582 0.557 0.530 0.422 0.691 0.525
    SeqTrack 0.590 0.600 0.580 0.582 0.441 0.665 0.504
    ViPT (Uni-model) 0.561 0.562 0.560 0.608 0.490 0.740 0.579
    Un-Track (Uni-model) 0.610 0.610 0.610 0.646 0.513 0.755 0.589

    When trained as a single unified checkpoint across mixed modalities, Un-Track achieves 0.610 F-score on DepthTrack, 0.646 PR / 0.513 SR on LasHeR, and 0.755 Precision / 0.589 Success on VisEvent. It outperforms previous uni-models (surpassing ViPT Uni-model by +4.9% F-score on DepthTrack and +3.8% PR on LasHeR) and also surpasses previous modality-specific state-of-the-art models (such as depth-specific ViPT with 0.594 F-score).

  7. Knowl 7 — Cross-Dataset Generalization and Missing-Modality Tracking Performance

    empirical result

    The generalization capability of Un-Track's single-checkpoint model was tested on out-of-distribution datasets (VOT-RGBD2022 and RGBT234) and on DepthTrack when auxiliary data is completely absent (dummy input).

    Method VOT-RGBD2022 RGBT234 DepthTrack (Dummy Depth)
    EAO ↑\uparrow Acc ↑\uparrow Rob ↑\uparrow MPR ↑\uparrow MSR ↑\uparrow F-score ↑\uparrow GFLOPs ↓\downarrow Params (M) ↓\downarrow
    RGB Baseline (OSTrack) 0.666 0.808 0.814 0.755 0.569 0.529 21.50 92.08
    SeqTrack (Uni-model) 0.679 0.802 0.846 0.806 0.599 - - -
    ViPT (Specific) 0.721 0.815 0.871 0.835 0.617 - - -
    ViPT (Uni-model) - - - - - 0.542 21.80 92.96
    Un-Track (Uni-model) 0.718 0.820 0.864 0.842 0.625 0.558 23.64 98.73

    On VOT-RGBD2022, Un-Track achieves 0.820 accuracy and 0.718 EAO, matching or exceeding depth-specific models. On RGBT234, Un-Track achieves 0.842 MPR and 0.625 MSR without specific thermal training, outperforming thermal-specific ViPT (0.835 MPR, 0.617 MSR). When the auxiliary modality is missing and replaced with dummy inputs on DepthTrack, Un-Track attains an F-score of 0.558, outperforming the RGB baseline (0.529) and ViPT (0.542) with an addition of only 2.14 GFLOPs and 6.65M parameters.

  8. Knowl 8 — Ablation of Framework Modules and Shared Embedding Components

    empirical result

    An ablation study on DepthTrack under the single parameter set configuration evaluated the necessity of each architectural component and design choice in the shared embedding module.

    Configuration F-score ↑\uparrow Recall ↑\uparrow Precision ↑\uparrow
    Full Un-Track (Uni-model) 0.610 0.608 0.611
    w/o Shared Embedding 0.599 0.602 0.597
    Replacing Prompting with Fovea Attention (ViPT) 0.579 0.575 0.584
    w/o LoRA Finetuning 0.594 0.598 0.596
    Shared Embedding Internal Variants:
    w/o Explicit Edge Guidance 0.600 0.602 0.598
    w/o Implicit Learning (Edge-Only Embedding) 0.604 0.609 0.599
    w/o In-Domain Low-Rank Approximation 0.581 0.583 0.579

    Key findings:

    1. Removing the shared embedding drops F-score from 0.610 to 0.599 due to direct exposure to heterogeneous modal distributions.
    2. Replacing token shrinkage prompting with fovea-attention prompting causes a major performance drop to 0.579.
    3. Removing explicit edge guidance from the shared embedding drops F-score to 0.600, while relying solely on edge features without implicit learning drops F-score to 0.604.
    4. Bypassing in-domain low-rank approximation and directly decomposing mixed-modality features causes the largest degradation (F-score 0.581), showing that domain-specific low-rank projection is essential prior to cross-modal concatenation.
  9. Knowl 9 — Hyperparameter Sensitivity for Low-Rank Dimensions and Token Percentiles

    empirical result

    The sensitivity of Un-Track's tracking performance on DepthTrack was analyzed across varying low-rank dimensions for shared embedding (kk), modal prompting (ll), LoRA fine-tuning (rr), and token selection percentiles.

    Metric Shared Rank kk Prompt Rank ll LoRA Rank rr Token Percentile
    22 44 88 44 88 1616 22 44 88 1/81/8 1/41/4 1/31/3
    F-score ↑\uparrow 0.607 0.610 0.602 0.596 0.610 0.606 0.601 0.610 0.600 0.604 0.610 0.595
    Recall ↑\uparrow 0.606 0.608 0.601 0.593 0.608 0.609 0.599 0.608 0.598 0.606 0.608 0.593
    Precision ↑\uparrow 0.608 0.611 0.604 0.599 0.611 0.604 0.602 0.611 0.602 0.602 0.611 0.596

    Optimal configuration:

    • k=4k=4 for shared embedding rank (lower ranks are overly sparse; higher ranks preserve modality-specific noise).
    • l=8l=8 for modal prompting rank.
    • r=4r=4 for LoRA adaptation rank.
    • Percentile =1/4= 1/4 (top 25% positive tokens retained, bottom 25% negative tokens replaced, middle 50% uncertain tokens fused), providing the best trade-off between token exchange and token fusion.

Coverage note — No substantial contributed material was omitted. All primary architecture designs, mathematical formulations, algorithms, cross-benchmark evaluation results, generalization experiments, and ablation analyses are included.

References

  1. 1.Simon Arridge, Pascal Fernsel, and Andreas Hauptmann. Joint reconstruction and low-rank decomposition for dynamic inverse problems. Inverse Problems and Imaging, 16(3):483–523, 2022. 2
  2. 2.Boyu Chen, Peixia Li, Lei Bai, Lei Qiao, Qiuhong Shen, Bo Li, Weihao Gan, Wei Wu, and Wanli Ouyang. Backbone is all your need: A simplified architecture for visual object tracking. In ECCV, 2022. 2
  3. 3.Qinyu Chen, Zuowen Wang, Shih-Chii Liu, and Chang Gao. 3et: Efficient event-based eye tracking using a change-based convlstm network. arXiv preprint arXiv:2308.11771, 2023. 6
  4. 4.Wanli Chen, Xinge Zhu, Ruoqi Sun, Junjun He, Ruiyu Li, Xiaoyong Shen, and Bei Yu. Tensor low-rank reconstruction for semantic segmentation. In ECCV, 2020. 2
  5. 5.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In CVPR, 2021. 2, 6
  6. 6.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In CVPR, 2021. 2
  7. 7.Xin Chen, Bin Yan, Jiawen Zhu, Huchuan Lu, Xiang Ruan, and Dong Wang. High-performance transformer tracking. TPAMI, 2022. 2
  8. 8.Xin Chen, Houwen Peng, Dong Wang, Huchuan Lu, and Han Hu. Seqtrack: Sequence to sequence learning for visual object tracking. In CVPR, 2023. 5, 6, 7
  9. 9.Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. Siamese box adaptive network for visual tracking. In CVPR, 2020. 1
  10. 10.Anthony Cioppa, Silvio Giancola, Adrien Deliege, Le Kang, Xin Zhou, Zhiyu Cheng, Bernard Ghanem, and Marc Van Droogenbroeck. Soccernet-tracking: Multiple object tracking dataset and benchmark in soccer videos. In CVPR, 2022. 1
  11. 11.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Mixformer: End-to-end tracking with iterative mixed attention. In CVPR, 2022. 2
  12. 12.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ATOM: Accurate tracking by overlap maximization. In CVPR, 2019. 1
  13. 13.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic regression for visual tracking. In CVPR, 2020. 6
  14. 14.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. LaSOT: A high-quality benchmark for large-scale single object tracking. In CVPR, 2019. 2, 5
  15. 15.Yingkai Fu, Meng Li, Wenxi Liu, Yuanchen Wang, Jiqing Zhang, Baocai Yin, Xiaopeng Wei, and Xin Yang. Distractor-aware event-based tracking. TIP, 2023. 1
  16. 16.Shang Gao, Jinyu Yang, Zhe Li, Feng Zheng, Alesˇ Leonardis, and Jingkuan Song. Learning dual-fused modality-aware representations for rgbd tracking. In ECCV, 2022. 2
  17. 17.Shenyuan Gao, Chunluan Zhou, Chao Ma, Xinggang Wang, and Junsong Yuan. Aiatrack: Attention in attention for transformer visual tracking. In ECCV, 2022. 5, 6, 7
  18. 18.Yuan Gao, Chenglong Li, Yabin Zhu, Jin Tang, Tao He, and Futian Wang. Deep adaptive fusion network for high performance RGBT tracking. In ICCVW, 2019. 6, 7
  19. 19.Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. In CVPR, 2023. 2
  20. 20.Botao He, Haojia Li, Siyuan Wu, Dong Wang, Zhiwei Zhang, Qianli Dong, Chao Xu, and Fei Gao. Fast-dynamic-vision: Detection and tracking dynamic objects with event and depth sensing. In IROS, 2021. 1
  21. 21.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In ICLR, 2022. 3, 4, 5
  22. 22.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. TPAMI, 2019. 2, 5
  23. 23.Sajid Javed, Martin Danelljan, Fahad Shahbaz Khan, Muhammad Haris Khan, Michael Felsberg, and Jiri Matas. Visual object tracking with discriminative filters and siamese networks: a survey and outlook. TPAMI, 45(5):6552–6574, 2022. 1
  24. 24.I-Hong Jhuo, Dong Liu, DT Lee, and Shih-Fu Chang. Robust visual domain adaptation with low-rank reconstruction. In CVPR, 2012. 2
  25. 25.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. In ECCV, 2022. 2, 4
  26. 26.Ben Kang, Xin Chen, Dong Wang, Houwen Peng, and Huchuan Lu. Exploring lightweight hierarchical vision transformers for efficient visual tracking. In ICCV, 2023. 2
  27. 27.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. In CVPR, 2023. 2
  28. 28.Matej Kristan, Ales Leonardis, Ji ˇ ˇr´ı Matas, Michael Felsberg, Roman Pflugfelder, Joni-Kristian Kam¨ ar¨ ainen, Martin ¨ Danelljan, Luka Cehovin Zajc, Alan Luke ˇ ziˇ c, Ondrej Dr- ˇ bohlav, et al. The eighth visual object tracking vot2020 challenge results. In ECCVW, 2020. 5
  29. 29.Matej Kristan, Jiˇr´ı Matas, Ales Leonardis, Michael Felsberg, ˇ Roman Pflugfelder, Joni-Kristian Kam¨ ar¨ ainen, Hyung Jin ¨ Chang, Martin Danelljan, Luka Cehovin, Alan Lukeziˇ c, et al. ˇ The ninth visual object tracking vot2021 challenge results. In ICCVW, 2021. 7
  30. 30.Matej Kristan, Ales Leonardis, Ji ˇ ˇr´ı Matas, Michael Felsberg, Roman Pflugfelder, Joni-Kristian Kam¨ ar¨ ainen, Hyung Jin ¨ Chang, Martin Danelljan, Luka Cehovin Zajc, Alan Luke ˇ ziˇ c,ˇ et al. The tenth visual object tracking vot2022 challenge results. In ECCVW, 2023. 6, 7
  31. 31.Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, and Chen-Yu Lee. Multimodal prompting with missing modalities for visual recognition. In CVPR, 2023. 2, 3
  32. 32.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High performance visual tracking with siamese region proposal network. In CVPR, 2018. 2
  33. 33.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. SiamRPN++: Evolution of siamese visual tracking with very deep networks. In CVPR, 2019. 2
  34. 34.Chenglong Li, Nan Zhao, Yijuan Lu, Chengli Zhu, and Jin Tang. Weighted sparse representation regularized graph learning for RGB-T object tracking. In ACM MM, 2017. 6, 7
  35. 35.Chenglong Li, Xinyan Liang, Yijuan Lu, Nan Zhao, and Jin Tang. RGB-T object tracking: Benchmark and baseline. PR, 96:106977, 2019. 2, 6, 7
  36. 36.Chenglong Li, Wanlin Xue, Yaqing Jia, Zhichen Qu, Bin Luo, Jin Tang, and Dengdi Sun. Lasher: A large-scale high-diversity benchmark for RGBT tracking. TIP, 31:392–404, 2021. 2, 5, 6
  37. 37.Qiao Liu, Xin Li, Zhenyu He, Nana Fan, Di Yuan, and Hongpeng Wang. Learning deep multi-level similarity for thermal infrared object tracking. TMM, 23:2114–2126, 2020. 2
  38. 38.Cheng Long Li, Andong Lu, Ai Hua Zheng, Zhengzheng Tu, and Jin Tang. Multi-adapter rgbt tracking. In ICCVW, 2019. 2
  39. 39.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. In ICLR, 2023. 2
  40. 40.Alan Lukezic, Ugur Kart, Jani Kapyla, Ahmed Durmush, Joni-Kristian Kamarainen, Jiri Matas, and Matej Kristan. Cdtb: A color and depth visual object tracking dataset and benchmark. In ICCV, 2019. 1
  41. 41.Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng. Smil: Multimodal learning with severely missing modality. In AAAI, 2021. 2
  42. 42.Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem. TrackingNet: A large-scale dataset and benchmark for object tracking in the wild. In ECCV, 2018. 2, 5
  43. 43.Yansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun, and Feng Wu. Get: Group event transformer for event-based vision. In ICCV, 2023. 6
  44. 44.Yansheng Qiu, Ziyuan Zhao, Hongdou Yao, Delin Chen, and Zheng Wang. Modal-aware visual prompting for incomplete multi-modal brain tumor segmentation. In ACM MM, 2023. 2
  45. 45.Esteban Real, Jonathon Shlens, Stefano Mazzocchi, Xin Pan, and Vincent Vanhoucke. Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. In CVPR, 2017. 2
  46. 46.Yibing Song, Chao Ma, Xiaohe Wu, Lijun Gong, Linchao Bao, Wangmeng Zuo, Chunhua Shen, Rynson WH Lau, and Ming-Hsuan Yang. Vital: Visual tracking via adversarial learning. In CVPR, 2018. 6
  47. 47.Daniel Stadler and Jurgen Beyerer. Improving multiple pedestrian tracking by track management and occlusion handling. In CVPR, 2021. 1
  48. 48.Chuanming Tang, Xiao Wang, Ju Huang, Bo Jiang, Lin Zhu, Jianlin Zhang, Yaowei Wang, and Yonghong Tian. Revisiting color-event based tracking: A unified network, dataset, and metric. arXiv preprint arXiv:2211.11010, 2022. 1
  49. 49.Chuanming Tang, Xiao Wang, Ju Huang, Bo Jiang, Lin Zhu, Jianlin Zhang, Yaowei Wang, and Yonghong Tian. Revisiting color-event based tracking: A unified network, dataset, and metric. arXiv preprint arXiv:2211.11010, 2022. 2
  50. 50.Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr, and Bastian Leibe. Siam R-CNN: Visual tracking by re-detection. In CVPR, 2020. 2
  51. 51.Zhexiong Wan, Yuxin Mao, Jing Zhang, and Yuchao Dai. Rpeflow: Multimodal fusion of rgb-pointcloud-event for joint optical flow and scene flow estimation. In ICCV, 2023. 2
  52. 52.Chaoqun Wang, Chunyan Xu, Zhen Cui, Ling Zhou, Tong Zhang, Xiaoya Zhang, and Jian Yang. Cross-modal pattern-propagation for RGB-T tracking. In CVPR, 2020. 7
  53. 53.Hu Wang, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, and Gustavo Carneiro. Multi-modal learning with missing modality via shared-specific feature modelling. In CVPR, 2023. 3
  54. 54.Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li. Transformer meets tracker: Exploiting temporal context for robust visual tracking. In CVPR, 2021. 2
  55. 55.Xiao Wang, Jianing Li, Lin Zhu, Zhipeng Zhang, Zhe Chen, Xin Li, Yaowei Wang, Yonghong Tian, and Feng Wu. Visevent: Reliable object tracking via collaboration of frame and event flows. arXiv preprint arXiv:2108.05015, 2021. 2, 5, 6
  56. 56.Xiao Wang, Xiujun Shu, Shilliang Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, and Feng Wu. Mfgnet: Dynamic modality-aware filter generation for rgb-t tracking. TMM, 2022. 2
  57. 57.Zuowen Wang, Yuhuang Hu, and Shih-Chii Liu. Exploiting spatial sparsity for event cameras with visual transformers. In ICIP, 2022. 2
  58. 58.David Wisth, Marco Camurri, Sandipan Das, and Maurice Fallon. Unified multi-modal landmark tracking for tightly coupled lidar-visual-inertial odometry. RA-L, 6(2):1004–1011, 2021. 2
  59. 59.Yun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu, and Jin Tang. Attribute-based progressive fusion network for RGBT tracking. In AAAI, 2022. 7
  60. 60.Yun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu, and Jin Tang. Attribute-based progressive fusion network for rgbt tracking. In AAAI, 2022. 1, 2
  61. 61.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. SiamFC++: Towards robust and accurate visual tracking with target estimation guidelines. In AAAI, 2020. 2
  62. 62.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In ICCV, 2021. 5, 6, 7
  63. 63.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In ICCV, 2021. 2
  64. 64.Bin Yan, Houwen Peng, Kan Wu, Dong Wang, Jianlong Fu, and Huchuan Lu. Lighttrack: Finding lightweight neural networks for object tracking via one-shot architecture search. In CVPR, 2021. 2
  65. 65.Bin Yan, Yi Jiang, Jiannan Wu, Dong Wang, Ping Luo, Zehuan Yuan, and Huchuan Lu. Universal instance perception as object discovery and retrieval. In CVPR, 2023. 5
  66. 66.Song Yan, Jinyu Yang, Jani Kapyl ¨ a, Feng Zheng, Ale ¨ sˇ Leonardis, and Joni-Kristian Kam¨ ar¨ ainen. Depthtrack: Un- ¨ veiling the power of RGBD tracking. In ICCV, 2021. 2, 5, 7
  67. 67.Song Yan, Jinyu Yang, Ales Leonardis, and Joni-Kristian Kamarainen. Depth-only object tracking. In BMVC, 2021. 2
  68. 68.Jinyu Yang, Zhe Li, Song Yan, Feng Zheng, Ales Leonardis, ˇ Joni-Kristian Kam¨ ar¨ ainen, and Ling Shao. Rgbd ob- ¨ ject tracking: An in-depth review. arXiv preprint arXiv:2203.14134, 2022. 1
  69. 69.Jinyu Yang, Zhe Li, Feng Zheng, Ales Leonardis, and Jingkuan Song. Prompting for multi-modal tracking. In ACMMM, 2022. 1, 2, 5, 6, 7
  70. 70.Jinyu Yang, Shang Gao, Zhe Li, Feng Zheng, and Alesˇ Leonardis. Resource-efficient rgbd aerial tracking. In CVPR, 2023. 2
  71. 71.Rui Yao, Guosheng Lin, Shixiong Xia, Jiaqi Zhao, and Yong Zhou. Video object segmentation and tracking: A survey. ACM TIST, 11(4):1–47, 2020. 2
  72. 72.Botao Ye, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Joint feature learning and relation modeling for tracking: A one-stream framework. In ECCV, 2022. 3, 5, 6, 7
  73. 73.Jie Yin, Ang Li, Tao Li, Wenxian Yu, and Danping Zou. M2dgr: A multi-sensor and multi-scenario slam dataset for ground robots. RA-L, 7(2):2266–2273, 2021. 2
  74. 74.Yuechen Yu, Yilei Xiong, Weilin Huang, and Matthew R Scott. Deformable siamese attention networks for visual object tracking. In CVPR, 2020. 2
  75. 75.Jiandian Zeng, Tianyi Liu, and Jiantao Zhou. Tag-assisted multimodal sentiment analysis under uncertain missing modalities. In ACM SIGIR, 2022. 2
  76. 76.Chunhui Zhang, Xin Sun, Yiqian Yang, Li Liu, Qiong Liu, Xi Zhou, and Yanfeng Wang. All in one: Exploring unified vision-language tracking with multi-modal alignment. In ACM MM, 2023. 2
  77. 77.Hui Zhang, Lei Zhang, Li Zhuo, and Jing Zhang. Object tracking in RGB-T videos using modal-aware attention network and competitive learning. Sensors, 20(2):393, 2020. 6, 7
  78. 78.Jiqing Zhang, Xin Yang, Yingkai Fu, Xiaopeng Wei, Baocai Yin, and Bo Dong. Object tracking by jointly exploiting frame and event domain. In ICCV, 2021. 2
  79. 79.Jiqing Zhang, Bo Dong, Haiwei Zhang, Jianchuan Ding, Felix Heide, Baocai Yin, and Xin Yang. Spiking transformers for event-based single object tracking. In CVPR, 2022. 1
  80. 80.Jiaming Zhang, Ruiping Liu, Hao Shi, Kailun Yang, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, and Rainer Stiefelhagen. Delivering arbitrary-modal semantic segmentation. In CVPR, 2023. 2
  81. 81.Lichao Zhang, Martin Danelljan, Abel Gonzalez-Garcia, Joost van de Weijer, and Fahad Shahbaz Khan. Multi-modal fusion for end-to-end RGB-T tracking. In ICCVW, 2019. 6, 7
  82. 82.Pengyu Zhang, Dong Wang, and Huchuan Lu. Multi-modal visual tracking: Review and experimental comparison. arXiv preprint arXiv:2012.04176, 2020. 1
  83. 83.Pengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu, and Xiang Ruan. Visible-thermal uav tracking: A large-scale benchmark and new baseline. In CVPR, 2022. 1
  84. 84.Wenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang, Jianping Shi, and Chen Change Loy. Robust multi-modality multi-object tracking. In ICCV, 2019. 1
  85. 85.Zhipeng Zhang and Houwen Peng. Deeper and wider siamese networks for real-time visual tracking. In CVPR, 2019. 2
  86. 86.Haojie Zhao, Dong Wang, and Huchuan Lu. Representation learning for visual object tracking by masked appearance transfer. In CVPR, 2023. 1
  87. 87.Jinjian Zhao, Xiaohan Zhang, and Pengyu Zhang. A unified approach for tracking uavs in infrared. In ICCV, 2021. 1
  88. 88.Shaochuan Zhao, Tianyang Xu, Xiao-Jun Wu, and Xue-Feng Zhu. Adaptive feature fusion for visual object tracking. PR, 111:107679, 2021. 2
  89. 89.Aihua Zheng, Zi Wang, Zihan Chen, Chenglong Li, and Jin Tang. Robust multi-modality person re-identification. In AAAI, 2021. 1
  90. 90.Jiawen Zhu, Zhenyu Chen, Zeqi Hao, Shijie Chang, Lu Zhang, Dong Wang, Huchuan Lu, Bin Luo, Jun-Yan He, Jin-Peng Lan, et al. Tracking anything in high quality. arXiv preprint arXiv:2307.13974, 2023. 1
  91. 91.Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, and Huchuan Lu. Visual prompt multi-modal tracking. In CVPR, 2023. 1, 2, 5, 6, 7
  92. 92.Jiawen Zhu, Huayi Tang, Zhi-Qi Cheng, Jun-Yan He, Bin Luo, Shihao Qiu, Shengming Li, and Huchuan Lu. Dcpt: Darkness clue-prompted tracking in nighttime uavs. arXiv preprint arXiv:2309.10491, 2023. 1
  93. 93.Xue-Feng Zhu, Xiao-Jun Wu, Tianyang Xu, Zhen-Hua Feng, and Josef Kittler. Robust visual object tracking via adaptive attribute-aware discriminative correlation filters. TMM, 24: 301–312, 2021. 2
  94. 94.Xue-Feng Zhu, Tianyang Xu, Zhangyong Tang, Zucheng Wu, Haodong Liu, Xiao Yang, Xiao-Jun Wu, and Josef Kittler. RGBD1K: A large-scale dataset and benchmark for RGB-D object tracking. AAAI, 2023. 2, 5, 7
  95. 95.Xue-Feng Zhu, Tianyang Xu, Zhangyong Tang, Zucheng Wu, Haodong Liu, Xiao Yang, Xiao-Jun Wu, and Josef Kittler. Rgbd1k: A large-scale dataset and benchmark for rgb-d object tracking. In AAAI, 2023. 2
  96. 96.Yabin Zhu, Chenglong Li, Jin Tang, and Bin Luo. Quality-aware feature aggregation network for robust RGBT tracking. IEEE TIV, 6(1):121–130, 2020. 6, 7
  97. 97.Zhiyu Zhu, Junhui Hou, and Xianqiang Lyu. Learning graph-embedded key-event back-tracing for object tracking in event clouds. NeurIPS, 2022. 2
  98. 98.Zhiyu Zhu, Junhui Hou, and Dapeng Oliver Wu. Cross-modal orthogonal high-rank augmentation for rgb-event transformer-trackers. In ICCV, 2023. 1, 6

Citation

MLA
Wu, Z., et al. “Single-Model and Any-Modality for Video Object Tracking”. arXiv, 2023, http://arxiv.org/abs/2311.15851v3.
APA
Wu, Z., Zheng, J., Ren, X., Vasluianu, F.-A., Ma, C., Paudel, D. P., Gool, L. V., & Timofte, R. (2023). Single-Model and Any-Modality for Video Object Tracking. arXiv. http://arxiv.org/abs/2311.15851v3
Chicago
Wu, Z., J. Zheng, X. Ren, et al. 2023. “Single-Model and Any-Modality for Video Object Tracking”. arXiv. http://arxiv.org/abs/2311.15851v3.
Harvard
Wu, Z. et al. (2023) “Single-Model and Any-Modality for Video Object Tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.15851v3.
Vancouver
1. Wu Z, Zheng J, Ren X, Vasluianu F-A, Ma C, Paudel DP, Gool LV, Timofte R (2023) Single-Model and Any-Modality for Video Object Tracking. arXiv

BibTeX

@article{wu2023single,
  title = {Single-Model and Any-Modality for Video Object Tracking},
  author = {Wu, Zongwei and Zheng, Jilai and Ren, Xiangxuan and Vasluianu, Florin-Alexandru and Ma, Chao and Paudel, Danda Pani and Gool, Luc Van and Timofte, Radu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.15851v3},
  eprint = {2311.15851}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE