Unifying Visual and Vision-Language Tracking via Contrastive Learning

Yinchao MaYuyang TangWenfei YangTianzhu ZhangJinpeng ZhangMengxue Kang

article2024AAAI59 citations

Proposes a unified tracking framework, UVLTrack, that utilizes multi-modal contrastive learning and a dynamic box head to track targets across bounding box, natural language, and combined reference settings using a single set of parameters.

Listen

Target tracking in video is essential for autonomous systems, robotics, and intelligent surveillance. In practical deployments, user inputs specifying what to track vary widely: systems may receive an initial visual bounding box, a natural language text description, or a combination of both. Historically, computer vision tracking models have specialized in only one or two of these input types. Because of the semantic gap between image features and language representations, trackers designed for natural language often struggle when provided only with bounding boxes, while visual-only trackers cannot utilize text to resolve visual ambiguity.

The article demonstrates that a single, unified deep learning architecture can achieve state-of-the-art tracking performance across all three input modalities (bounding box, natural language, and language plus bounding box) simultaneously without changing network parameters. To accomplish this, the authors introduce "UVLTrack," a framework combining a modality-unified feature extractor with a modality-adaptive box head.

The authors evaluated the framework across thirteen public benchmarks, covering seven visual tracking datasets, three vision-language tracking datasets, and three visual grounding benchmarks. The method isolates low-level visual and textual processing in early Transformer layers before merging them in deeper layers, aligning these representations using a multi-modal contrastive loss. To avoid rigid target estimation, the dynamic box head samples historical video frames to distinguish true targets from potential visual distractors and background elements.

The evaluation produced several key findings: First, UVLTrack established state-of-the-art results on all seven visual tracking benchmarks, demonstrating that cross-modal capabilities do not degrade standard visual performance. Second, under pure natural language tracking, the larger model variant (UVLTrack-L) outperformed the previous best model, JointNLT, by significant margins across all benchmarks, including an improvement in tracking success from 54.6% to 58.2% on the TNL2K benchmark. Third, the base model (UVLTrack-B) operated at 57 to 58 frames per second, running approximately 1.46 times faster than JointNLT while improving accuracy. Finally, ablation studies showed that aligning features with contrastive loss yielded consistent gains of 1.2% to 2.3% across all input modalities, while dynamic distractor modeling added up to 2.1% in success rates over static detection heads.

These results show that engineering teams do not need to maintain multiple specialized tracking pipelines for different operational inputs. A single, shared architecture reduces computational maintenance costs, streamlines model lifecycle management, and increases system robustness in real-time edge or server environments.

Organizations developing vision-based tracking systems should consider adopting unified Transformer architectures that integrate contrastive alignment and dynamic distractor modeling. To maximize performance, implementations should initialize textual encoders with dedicated pre-trained language models rather than relying solely on visual pre-training. While the model shows high confidence across benchmark scenarios, real-world deployment should be preceded by domain-specific pilot testing, particularly under severe environmental visibility constraints or extreme computational budget limits.

  • Paper: Transformer Tracking, Xin Chen et al. (2021). This foundational work demonstrates how attention-based Transformer architectures replace traditional correlation modules for feature fusion in visual tracking, establishing the tracking paradigm that UVLTrack builds upon.
  • Paper: MixFormer: End-to-End Tracking with Iterative Mixed Attention, Yutao Cui et al. (2022). It introduces mixed attention for simultaneous feature extraction and target integration, serving as an architectural precursor to unified Transformer tracking frameworks.
  • Paper: MixFormerV2: Efficient Fully Transformer Tracking, Yutao Cui et al. (2023). It establishes fully Transformer-based tracking architectures with unified prediction tokens, providing key context for modern Transformer tracker design.
  • Paper: End-to-End Referring Video Object Segmentation with Multimodal Transformers, Adam Botach et al. (2022). It explores multimodal Transformer architectures for tracking and segmenting targets from natural language queries, motivating cross-modal alignment in tracking.
  • Paper: LaSOT: A High-Quality Benchmark for Large-Scale Single Object Tracking, Heng Fan et al. (2018). This benchmark paper establishes large-scale visual tracking datasets that include natural language descriptions, providing the empirical foundation for vision-language tracking evaluations.
  • Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). It introduces universal cross-modal Transformer pre-training with contrastive and matching objectives, establishing principles for aligning image and text representations utilized in UVLTrack.
  • Paper: ATOM: Accurate Tracking by Overlap Maximization, Martin Danelljan et al. (2018). It introduces target bounding box estimation decoupled from online classification to prevent tracking drift, influencing the design of dynamic distractor-aware box heads.
Cover for Unifying Visual and Vision-Language Tracking via Contrastive Learning

Abstract

Single object tracking aims to locate the target object in a video sequence according to the state specified by different modal references, including the initial bounding box (BBOX), natural language (NL), or both (NL+BBOX). Due to the gap between different modalities, most existing trackers are designed for single or partial of these reference settings and overspecialize on the specific modality. Differently, we present a unified tracker called UVLTrack, which can simultaneously handle all three reference settings (BBOX, NL, NL+BBOX) with the same parameters. The proposed UVLTrack enjoys several merits. First, we design a modality-unified feature extractor for joint visual and language feature learning and propose a multi-modal contrastive loss to align the visual and language features into a unified semantic space. Second, a modality-adaptive box head is proposed, which makes full use of the target reference to mine ever-changing scenario features dynamically from video contexts and distinguish the target in a contrastive way, enabling robust performance in different reference settings. Extensive experimental results demonstrate that UVLTrack achieves promising performance on seven visual tracking datasets, three vision-language tracking datasets, and three visual grounding datasets. Codes and models will be open-sourced at https://github.com/OpenSpaceAI/UVLTrack.

Table of Contents

  • Introduction
  • Related Work
  • Visual Tracking
  • Vision-Language Tracking
  • Method
  • Tracking Architecture
  • Modality-Unified Feature Extractor
  • Modality-Adaptive Box Head
  • Training Objective
  • Experiment
  • Implementation Details
  • State-of-the-art Comparisons
  • Ablation Study
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Modality-Unified Tracking Architecture

    model/method

    The UVLTrack architecture is a unified single-object tracker capable of processing three reference modal settings using a single set of network parameters: initial bounding box only (BBOX), natural language description only (NL), or both (NL+BBOX).

    Given a text description ll, a target template image z∈R3×Hz×Wzz \in \mathbb{R}^{3 \times H_z \times W_z} cropped using the initial bounding box, and a search region image x∈R3×Hx×Wxx \in \mathbb{R}^{3 \times H_x \times W_x} (or full image for grounding), the inputs are tokenized as follows:

    1. Text ll is embedded into token embeddings El0∈RNl×CE_l^0 \in \mathbb{R}^{N_l \times C}, where NlN_l is the maximum text sequence length and CC is the feature dimension. A learnable language semantic token Tl0∈R1×CT_l^0 \in \mathbb{R}^{1 \times C} is prepended to capture global text semantics.
    2. Template zz and search region xx are partitioned into non-overlapping p×pp \times p patches, linearly projected, and combined with learnable position embeddings to form template embeddings Ez0∈RNz×CE_z^0 \in \mathbb{R}^{N_z \times C} and search embeddings Ex0∈RNx×CE_x^0 \in \mathbb{R}^{N_x \times C}, where Nz=HzWzp2N_z = \frac{H_z W_z}{p^2} and Nx=HxWxp2N_x = \frac{H_x W_x}{p^2}. A learnable visual semantic token Tv0∈R1×CT_v^0 \in \mathbb{R}^{1 \times C} is prepended to capture global visual semantics.

    The feature extractor consists of NN shallow Transformer encoder layers followed by MM deep Transformer encoder layers. In the shallow layers, visual tokens [Tv;Ez;Ex][T_v; E_z; E_x] and language tokens [Tl;El][T_l; E_l] are processed independently to avoid corruption of unimodal low-level feature representations. In the deep layers, visual and language tokens interact via attention mechanisms for high-level semantic fusion. When a reference modality is not provided in a given task setting, its corresponding token slots are filled with zeros and masked out.

  2. Knowl 2 — Task-Oriented Multi-Head Attention Mechanism

    model/method

    To enable parallel training and execution across bounding box, language, and joint reference modalities within a shared backbone, UVLTrack employs Task-Oriented Multi-Head Attention (TMHA) with input-dependent attention masking.

    For encoder layer i∈{1,…,N+M}i \in \{1, \dots, N+M\}, key KiK^i, query QiQ^i, and value ViV^i are computed from the layer input Ei−1E^{i-1} via layer normalization LN(⋅)\text{LN}(\cdot) and linear projections. The layer update is formulated as: E^i=Softmax(Qi(Ki)⊤C+Ma)Vi+Ei−1\hat{E}^i = \text{Softmax}\left( \frac{Q^i (K^i)^\top}{\sqrt{C}} + M_a \right) V^i + E^{i-1} Ei=MLP(LN(E^i))+E^iE^i = \text{MLP}(\text{LN}(\hat{E}^i)) + \hat{E}^i where CC is the channel dimension, MLP(⋅)\text{MLP}(\cdot) is a multi-layer perceptron, and Ma∈{0,−∞}M_a \in \{0, -\infty\} is a modality-specific attention mask matrix.

    In shallow layers (11 to NN), MaM_a blocks all attention interactions between text tokens [Tl,El][T_l, E_l] and visual tokens [Tv,Ez,Ex][T_v, E_z, E_x]. In deep layers (N+1N+1 to N+MN+M), MaM_a allows all-to-all cross-modal attention between present modalities while setting attention weights to −∞-\infty for any unavailable (zero-padded) reference modality tokens.

  3. Knowl 3 — Multi-Modal Contrastive Loss for Modality Alignment

    equation

    To project visual and natural language features into a shared semantic space and prevent the tracker from degrading when language is absent, a Multi-Modal Contrastive (MMC) loss is applied across each encoder layer i∈{1,…,N+M}i \in \{1, \dots, N+M\}.

    Let Ti∈R1×CT^i \in \mathbb{R}^{1 \times C} denote the semantic token at layer ii (TliT_l^i for language reference, TviT_v^i for visual reference) and Exi=[fi,1,fi,2,…,fi,Nx]∈RNx×CE_x^i = [f^{i,1}, f^{i,2}, \dots, f^{i,N_x}] \in \mathbb{R}^{N_x \times C} be the search region patch embeddings. The cosine similarity scores Si=[si,1,si,2,…,si,Nx]S^i = [s^{i,1}, s^{i,2}, \dots, s^{i,N_x}] are computed by: si,j=sim(Ti,fi,j)τ,sim(Ti,fi,j)=Ti(fi,j)⊤∥Ti∥2∥fi,j∥2s^{i,j} = \frac{\text{sim}(T^i, f^{i,j})}{\tau}, \quad \text{sim}(T^i, f^{i,j}) = \frac{T^i (f^{i,j})^\top}{\|T^i\|_2 \|f^{i,j}\|_2} where τ\tau is a temperature hyperparameter and ∥⋅∥2\|\cdot\|_2 denotes the Euclidean norm.

    Let spis_p^i be the similarity score at the center patch of the ground-truth target bounding box (positive sample), and let [sni,k]k=1Nneg[s_{n}^{i,k}]_{k=1}^{N_{neg}} denote the top NnegN_{neg} highest similarity scores among patches strictly outside the ground-truth target bounding box (hard negative samples). The layer-wise MMC loss is defined as: Lmmci=−log⁡(exp⁡(spi)exp⁡(spi)+∑k=1Nnegexp⁡(sni,k))\mathcal{L}_{mmc}^i = -\log\left( \frac{\exp(s_p^i)}{\exp(s_p^i) + \sum_{k=1}^{N_{neg}} \exp(s_{n}^{i,k})} \right)

    The total MMC loss summed over all N+MN+M encoder layers is added to the training objective.

  4. Knowl 4 — Modality-Adaptive Box Head

    model/method

    The Modality-Adaptive Box Head (MABH) combines a static convolutional prediction network with a dynamic scenario-based prototype comparator to localize targets stably across disparate reference modalities.

    The search region embeddings from the final encoder layer ExN+ME_x^{N+M} are reshaped into a 2D spatial feature map of shape Hxp×Wxp\frac{H_x}{p} \times \frac{W_x}{p} and processed by a three-branch convolutional network to generate:

    1. A center score map C^∈(0,1)Hxp×Wxp\hat{C} \in (0, 1)^{\frac{H_x}{p} \times \frac{W_x}{p}}
    2. An offset map O^∈[0,1)2×Hxp×Wxp\hat{O} \in [0, 1)^{2 \times \frac{H_x}{p} \times \frac{W_x}{p}}
    3. A normalized box size map S^∈(0,1)2×Hxp×Wxp\hat{S} \in (0, 1)^{2 \times \frac{H_x}{p} \times \frac{W_x}{p}}

    In parallel, dynamic scenario prototype matching computes a target similarity map L^∈(0,1)Hxp×Wxp\hat{L} \in (0, 1)^{\frac{H_x}{p} \times \frac{W_x}{p}}. The target center coordinate (xc,yc)(x_c, y_c) is determined by modulating the static center score map with the dynamic target similarity map: (xc,yc)=arg⁡max⁡(x,y)(C^(x,y)⋅L^(x,y))(x_c, y_c) = \arg\max_{(x, y)} \left( \hat{C}(x, y) \cdot \hat{L}(x, y) \right)

    The final estimated target bounding box b^=(x^,y^,w^,h^)\hat{b} = (\hat{x}, \hat{y}, \hat{w}, \hat{h}) is then computed via: (x^,y^)=((xc+O^(0,xc,yc))⋅p,  (yc+O^(1,xc,yc))⋅p)(\hat{x}, \hat{y}) = \left( (x_c + \hat{O}(0, x_c, y_c)) \cdot p, \; (y_c + \hat{O}(1, x_c, y_c)) \cdot p \right) (w^,h^)=(S^(0,xc,yc)⋅Hx,  S^(1,xc,yc)⋅Wx)(\hat{w}, \hat{h}) = \left( \hat{S}(0, x_c, y_c) \cdot H_x, \; \hat{S}(1, x_c, y_c) \cdot W_x \right)

  5. Knowl 5 — Distribution-Based Dynamic Prototype Mining

    equation

    To discriminate the target from distractors and background in video contexts, dynamic prototypes are mined from template embeddings EzN+ME_z^{N+M} and historical high-confidence context embeddings EcN+ME_c^{N+M}, concatenated as Et=[EzN+M;EcN+M]∈R(Nz+Nc)×CE_t = [E_z^{N+M}; E_c^{N+M}] \in \mathbb{R}^{(N_z + N_c) \times C}, with target mask Mt=[Mz;Mc]∈R1×(Nz+Nc)M_t = [M_z; M_c] \in \mathbb{R}^{1 \times (N_z + N_c)} (00 for inside-target positions, −∞-\infty outside) and its complement M~t\tilde{M}_t.

    In-box and out-box cross-attention weights between final semantic token TN+MT^{N+M} and EtE_t are: Ain=Softmax(TN+MEt⊤C+Mt),Aout=Softmax(TN+MEt⊤C+M~t)A_{in} = \text{Softmax}\left( \frac{T^{N+M} E_t^\top}{\sqrt{C}} + M_t \right), \quad A_{out} = \text{Softmax}\left( \frac{T^{N+M} E_t^\top}{\sqrt{C}} + \tilde{M}_t \right)

    The target update token is Tt=AinEtT_t = A_{in} E_t. For out-box patches, probabilities in AoutA_{out} are sorted in descending order and cumulatively summed. Patches where the cumulative sum is below threshold β\beta form distractor mask MdM_d (00 for distractors, −∞-\infty otherwise), and the remaining out-box patches form background mask M~d\tilde{M}_d. Distractor and background tokens are computed via: Td=Softmax(TN+MEt⊤C+M~t+Md)Et,Tb=Softmax(TN+MEt⊤C+M~t+M~d)EtT_d = \text{Softmax}\left( \frac{T^{N+M} E_t^\top}{\sqrt{C}} + \tilde{M}_t + M_d \right) E_t, \quad T_b = \text{Softmax}\left( \frac{T^{N+M} E_t^\top}{\sqrt{C}} + \tilde{M}_t + \tilde{M}_d \right) E_t

    Prototypes are updated by adding tokens to learnable prototypes P^t=TN+M\hat{P}_t = T^{N+M}, P^d\hat{P}_d, and P^b\hat{P}_b: Pt=P^t+Tt,Pd=P^d+Td,Pb=P^b+TbP_t = \hat{P}_t + T_t, \quad P_d = \hat{P}_d + T_d, \quad P_b = \hat{P}_b + T_b

    For search region patch features ExN+M=[f1,…,fNx]E_x^{N+M} = [f^1, \dots, f^{N_x}], similarity scores L^=[αt1,…,αtNx]\hat{L} = [\alpha_t^1, \dots, \alpha_t^{N_x}] are computed with temperature τ\tau: α^ti=sim(fi,Pt)τ,α^bi=max⁡(sim(fi,Pd)τ,sim(fi,Pb)τ,0),αti=exp⁡(α^ti)exp⁡(α^ti)+exp⁡(α^bi)\hat{\alpha}_t^i = \frac{\text{sim}(f^i, P_t)}{\tau}, \quad \hat{\alpha}_b^i = \max\left( \frac{\text{sim}(f^i, P_d)}{\tau}, \frac{\text{sim}(f^i, P_b)}{\tau}, 0 \right), \quad \alpha_t^i = \frac{\exp(\hat{\alpha}_t^i)}{\exp(\hat{\alpha}_t^i) + \exp(\hat{\alpha}_b^i)} where the constant 00 in the background term prevents unseen objects from receiving artificially high target scores.

  6. Knowl 6 — Unified Tracking Loss Objective

    equation

    The overall training objective function L\mathcal{L} for UVLTrack combines target score map supervision, center classification loss, bounding box regression loss, and multi-modal contrastive losses across all N+MN+M encoder layers: L=Ltgt+Lcls+Lbox+λmmc∑i=1N+MLmmci\mathcal{L} = \mathcal{L}_{tgt} + \mathcal{L}_{cls} + \mathcal{L}_{box} + \lambda_{mmc} \sum_{i=1}^{N+M} \mathcal{L}_{mmc}^i where:

    1. Ltgt=Lbce(L^,L)\mathcal{L}_{tgt} = \mathcal{L}_{bce}(\hat{L}, L) is the binary cross-entropy loss between the predicted dynamic target similarity map L^\hat{L} and ground-truth binary target mask LL (11 for patches inside the target bounding box, 00 outside).
    2. Lcls\mathcal{L}_{cls} is the focal classification loss on the predicted center score map C^\hat{C}.
    3. Lbox=λ1L1+λgiouLgiou\mathcal{L}_{box} = \lambda_1 \mathcal{L}_1 + \lambda_{giou} \mathcal{L}_{giou} is the bounding box regression loss combining L1L_1 loss and Generalized Intersection-over-Union (GIoU) loss.
    4. Hyperparameter weights are set to λgiou=2.0\lambda_{giou} = 2.0, λ1=5.0\lambda_1 = 5.0, and λmmc=0.1\lambda_{mmc} = 0.1.
  7. Knowl 7 — UVLTrack Model Variants and Experimental Setup

    experimental setup

    UVLTrack is evaluated using two architectural configurations:

    • UVLTrack-B: N=6N = 6 shallow encoder layers, M=6M = 6 deep encoder layers (1212 layers total). Image backbone parameters are initialized from ViT-B pretrained with Masked Autoencoders (MAE). Language branch shallow layers are initialized with BERT-base (uncased).
    • UVLTrack-L: N=12N = 12 shallow encoder layers, M=12M = 12 deep encoder layers (2424 layers total). Image backbone parameters are initialized from ViT-L pretrained with MAE. Language branch shallow layers are initialized with BERT-base (uncased).

    Image template patches are cropped to 128×128128 \times 128 (22×2^2 \times bounding box area) and search patches to 256×256256 \times 256 (42×4^2 \times bounding box area) with patch size p=16p = 16. Maximum text sentence length is Nl=40N_l = 40. In the dynamic head, distractor threshold β=0.75\beta = 0.75 and negative sample count Nneg=9N_{neg} = 9.

    Models are trained jointly across all reference modalities on the training sets of LaSOT, GOT-10k, COCO2017, TrackingNet, TNL2K, OTB99, and RefCOCOg-google using data augmentations (translation, horizontal flip, color jittering).

  8. Knowl 8 — Visual Tracking Performance Across Standard Benchmarks

    data/table

    When initialized exclusively with the initial bounding box (BBOX), UVLTrack outperforms existing visual-only trackers across seven visual tracking benchmarks evaluated via Success Area Under the Curve (AUC) and Precision (P):

    Method TNL2K LaSOT LaSOText TrackingNet NFS UAV123
    AUC P AUC P AUC P AUC P AUC AUC
    UVLTrack-L 64.8 68.8 71.3 78.3 51.2 59.0 84.1 82.9 67.6 71.0
    OSTrack-384 55.9 - 71.1 77.6 50.5 57.6 83.9 83.2 66.5 70.7
    MixFormer-L - - 70.1 76.3 - - 83.9 83.1 - -
    SimTrack-L/14 55.6 55.7 70.5 - - - 83.4 - - -
    UVLTrack-B 62.7 65.4 69.4 74.9 49.2 55.8 83.4 82.1 65.9 69.3
    OSTrack-256 54.3 - 69.1 75.2 47.4 53.3 83.1 82.0 64.7 68.3
    MixFormer-22k - - 69.2 74.7 - - 83.1 81.6 - -
    SimTrack-B/16 54.8 53.8 69.3 - - - 82.3 - - -
    AiATrack - - 69.0 73.8 47.7 55.4 82.7 80.4 - -
    STARK - - 66.4 71.2 - - 81.3 78.1 - -
    TransT 50.7 51.7 64.9 73.8 - - 81.4 80.3 65.3 68.1

    When evaluated without language on LaSOT, earlier vision-language trackers suffered significant drops (JointNLT: 54.5% AUC, VLTTT: 53.4% AUC), whereas UVLTrack maintains robust visual tracking (69.4% / 71.3% AUC).

  9. Knowl 9 — Vision-Language Tracking and Visual Grounding Performance

    data/table

    UVLTrack was evaluated on vision-language tracking benchmarks under natural language only (NL) and joint language plus bounding box (NL+BBOX) initialization, as well as on visual grounding datasets:

    TNL2K LaSOT OTB99
    Method AUC P AUC P AUC P
    NL Setting
    UVLTrack-L 58.2 60.9 59.6 63.9 63.5 83.2
    UVLTrack-B 55.7 57.2 57.2 61.0 60.1 79.1
    JointNLT 54.6 55.0 56.9 59.3 59.2 77.6
    CTRNLT 14.0 9.0 52.0 51.0 53.0 72.0
    NL+BBOX Setting
    UVLTrack-L 64.9 69.3 71.4 78.7 71.1 92.0
    UVLTrack-B 63.1 66.7 69.4 75.9 69.3 89.9
    JointNLT 56.9 58.1 60.4 63.6 65.3 85.6
    VLTTT 53.1 53.3 67.3 72.1 76.4 93.1
    SNLT 27.6 41.9 54.0 57.6 66.6 80.4

    On Visual Grounding Top-1 accuracy, UVLTrack achieves state-of-the-art results across datasets:

    • RefCOCO: val 85.47%, testA 87.56%, testB 81.73% (vs. VLTVG: 84.77%, 87.24%, 80.49%)
    • RefCOCO+: val 74.60%, testA 79.70%, testB 65.64% (vs. VLTVG: 74.19%, 78.93%, 65.17%)
    • RefCOCOg: val-g 73.86%, test-u 74.86% (vs. VLTVG: 72.98%, 74.18%)

    In inference speed, UVLTrack-B operates at 58 FPS (visual) / 57 FPS (vision-language), running 1.46×1.46\times faster than JointNLT (39 FPS).

  10. Knowl 10 — Ablation Analysis on Architecture, Losses, and Distractor Mining

    empirical result

    Ablation experiments conducted on TNL2K with UVLTrack-B demonstrate the impact of key architectural and objective components:

    1. Component Contributions:

      • Baseline (UVLTrack without MMCLoss, using static anchor-free head): BBOX AUC 59.4%, NL AUC 51.6%, NL+BBOX AUC 59.7%.
      • Adding MMCLoss: BBOX AUC increases by +1.2% (to 60.6%), NL by +2.2% (to 53.8%), and NL+BBOX by +2.3% (to 62.0%), confirming that multimodal alignment benefits all tracking modes.
      • Adding MABH: BBOX AUC increases by +2.1% (to 62.7%), NL by +1.9% (to 55.7%), and NL+BBOX by +1.1% (to 63.1%).
    2. Shallow (NN) vs. Deep (MM) Encoder Layers Split (total 12 layers):

      • All deep fusion (N=0,M=12N=0, M=12): BBOX 61.3%, NL 55.1%, NL+BBOX 62.3% (early fusion impairs low-level unimodal feature modeling).
      • Almost all shallow (N=11,M=1N=11, M=1): BBOX 62.5%, NL 46.2%, NL+BBOX 61.9% (insufficient fusion harms language tracking).
      • Equal split (N=6,M=6N=6, M=6): Achieves optimal balance across all reference settings (62.7% / 55.7% / 63.1%).
    3. Distractor Separation Threshold β\beta:

      • Merging all non-target background into a single token without distractor thresholding (β=0.0\beta = 0.0) degrades tracking (BBOX 61.3%, NL 54.3%, NL+BBOX 61.8%). Distractor mining is optimal at β=0.75\beta = 0.75 (62.7% / 55.7% / 63.1%).
    4. Sampling Strategy in MMCLoss:

      • Selecting the center target patch as positive sample and the top-99 hard negatives outside the box achieves superior performance (AUC 62.7% / 55.7% / 63.1%) compared to average target feature pooling (62.1% / 55.2% / 62.3%) or random negative sampling (61.8% / 54.9% / 62.1%).

Coverage note — None was omitted; all key architectural components, mathematical formulations, experimental settings, empirical comparisons, and ablation studies have been captured.

References

  1. 1.Alper, Y.; Omar, J.; Mubarak, S.; et al. 2006. Object tracking: A survey. ACM Computing Surveys, 38(4): 13–es.
  2. 2.Bertinetto, L.; Valmadre, J.; Henriques, J. F.; Vedaldi, A.; and Torr, P. H. S. 2016. Fully-Convolutional Siamese Networks for Object Tracking. In Proceedings of the European Conference on Computer Vision Workshops.
  3. 3.Bhat, G.; Danelljan, M.; Gool, L. V.; and Timofte, R. 2019. Learning Discriminative Model Prediction for Tracking. In Proceedings of the IEEE International Conference on Computer Vision.
  4. 4.Chen, B.; Li, P.; Bai, L.; Qiao, L.; Shen, Q.; Li, B.; Gan, W.; Wu, W.; and Ouyang, W. 2022a. Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking. In Proceedings of the European Conference on Computer Vision.
  5. 5.Chen, F.; Wang, X.; Zhao, Y.; Lv, S.; and Niu, X. 2022b. Visual object tracking: A survey. Computer Vision and Image Understanding, 222: 103508.
  6. 6.Chen, L.; Ma, W.; Xiao, J.; Zhang, H.; and Chang, S.-F. 2021a. Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding. In Proceedings of the AAAI conference on artificial intelligence, volume 35, 1036–1044.
  7. 7.Chen, X.; Yan, B.; Zhu, J.; Wang, D.; Yang, X.; and Lu, H. 2021b. Transformer Tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  8. 8.Cui, Y.; Cheng, J.; Wang, L.; and Wu, G. 2022. MixFormer: End-to-End Tracking with Iterative Mixed Attention. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  9. 9.Danelljan, M.; Bhat, G.; Khan, F. S.; and Felsberg, M. 2019. ATOM: Accurate Tracking by Overlap Maximization. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  10. 10.Danelljan, M.; Gool, L. V.; Timofte, R.; et al. 2020. Probabilistic regression for visual tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  11. 11.Deng, J.; Yang, Z.; Chen, T.; Zhou, W.; and Li, H. 2021. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1769–1779.
  12. 12.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT.
  13. 13.Fan, H.; Bai, H.; Lin, L.; Yang, F.; Chu, P.; Deng, G.; Yu, S.; Huang, M.; Liu, J.; Xu, Y.; et al. 2021. Lasot: A high-quality large-scale single object tracking benchmark. International Journal of Computer Vision, 129(2): 439–461.
  14. 14.Fan, H.; Lin, L.; Yang, F.; Chu, P.; Deng, G.; Yu, S.; Bai, H.; Xu, Y.; Liao, C.; and Ling, H. 2019. LaSOT: A High-Quality Benchmark for Large-Scale Single Object Tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  15. 15.Feng, Q.; Ablavsky, V.; Bai, Q.; and Sclaroff, S. 2021. Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5851–5860.
  16. 16.Fu, Z.; Liu, Q.; Fu, Z.; and Wang, Y. 2021. Stmtrack: Template-free visual tracking with space-time memory networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 13774–13783.
  17. 17.Gao, S.; Zhou, C.; Ma, C.; Wang, X.; and Yuan, J. 2022. Aiatrack: Attention in attention for transformer visual tracking. In Proceedings of the European Conference on Computer Vision, 146–164. Springer.
  18. 18.Glorot, X.; et al. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the International Conference on Artificial Intelligence and Statistics.
  19. 19.Guo, D.; Wang, J.; Cui, Y.; Wang, Z.; and Chen, S. 2020. SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  20. 20.Guo, M.; Zhang, Z.; Fan, H.; and Jing, L. 2022. Divert more attention to vision-language tracking. Advances in Neural Information Processing Systems, 35: 4446–4460.
  21. 21.He, K.; Chen, X.; Xie, S.; Li, Y.; Dollar, P.; and Girshick, R. 2022. Masked autoencoders are scalable vision learners. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 16000–16009.
  22. 22.Huang, B.; Lian, D.; Luo, W.; and Gao, S. 2021. Look before you leap: Learning landmark features for one-stage visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16888–16897.
  23. 23.Huang, L.; Zhao, X.; Huang, K.; et al. 2019. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  24. 24.Kiani, H., Galoogahi; Fagg, A.; Huang, C.; Ramanan, D.; Lucey, S.; et al. 2017. Need for speed: A benchmark for higher frame rate object tracking. In Proceedings of the IEEE International Conference on Computer Vision.
  25. 25.Li, B.; Wu, W.; Wang, Q.; Zhang, F.; Xing, J.; and Yan, J. 2019. SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition.
  26. 26.Li, Y.; Yu, J.; Cai, Z.; and Pan, Y. 2022. Cross-modal target retrieval for tracking by natural language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4931–4940.
  27. 27.Li, Z.; Tao, R.; Gavves, E.; Snoek, C. G.; and Smeulders, A. W. 2017. Tracking by natural language specification. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6495–6503.
  28. 28.Lin, T.-Y.; Maire, M.; Belongie, S. J.; Bourdev, L. D.; Girshick, R. B.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; and Zitnick, C. L. 2014. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision.
  29. 29.Liu, D.; Zhang, H.; Wu, F.; and Zha, Z.-J. 2019. Learning to assemble neural module tree networks for visual grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4673–4682.
  30. 30.Ma, Y.; He, J.; Yang, D.; Zhang, T.; and Wu, F. 2023. Adaptive Part Mining for Robust Visual Tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  31. 31.Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 11–20.
  32. 32.Mayer, C.; Danelljan, M.; Paudel, D. P.; and Van Gool, L. 2021. Learning target candidate association to keep track of what not to track. In Proceedings of the IEEE International Conference on Computer Vision, 13444–13454.
  33. 33.Mueller, M.; Smith, N.; Ghanem, B.; et al. 2016. A benchmark and simulator for UAV tracking. In Proceedings of the European Conference on Computer Vision.
  34. 34.Muller, M.; Bibi, A.; Giancola, S.; Alsubaihi, S.; and Ghanem, B. 2018. TrackingNet: A large-scale dataset and benchmark for object tracking in the wild. In Proceedings of the European Conference on Computer Vision.
  35. 35.Noman, M.; Ghallabi, W. A.; Najiha, D.; Mayer, C.; Dudhane, A.; Danelljan, M.; Cholakkal, H.; Khan, S.; Van Gool, L.; and Khan, F. S. 2022. Avist: A benchmark for visual object tracking in adverse visibility. arXiv preprint arXiv:2208.06888.
  36. 36.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is all you need. In Advances of Neural Information Processing Systems.
  37. 37.Wang, N.; Zhou, W.; Wang, J.; and Li, H. 2021a. Transformer meets tracker: Exploiting temporal context for robust visual tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 1571–1580.
  38. 38.Wang, X.; Shu, X.; Zhang, Z.; Jiang, B.; Wang, Y.; Tian, Y.; and Wu, F. 2021b. Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 13763–13773.
  39. 39.Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; and Zhang, L. 2021. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE International Conference on Computer Vision, 22–31.
  40. 40.Xu, Y.; Wang, Z.; Li, Z.; Yuan, Y.; and Yu, G. 2020. SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines. In Proceedings of the AAAI Conference on Artificial Intelligence.
  41. 41.Yan, B.; Peng, H.; Fu, J.; Wang, D.; and Lu, H. 2021. Learning spatio-temporal transformer for visual tracking. In Proceedings of the IEEE International Conference on Computer Vision, 10448–10457.
  42. 42.Yang, L.; Xu, Y.; Yuan, C.; Liu, W.; Li, B.; and Hu, W. 2022. Improving visual grounding with visual-linguistic verification and iterative reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9499–9508.
  43. 43.Yang, Z.; Chen, T.; Wang, L.; and Luo, J. 2020a. Improving one-stage visual grounding by recursive sub-query construction. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16, 387–404. Springer.
  44. 44.Yang, Z.; Kumar, T.; Chen, T.; Su, J.; and Luo, J. 2020b. Grounding-tracking-integration. IEEE Transactions on Circuits and Systems for Video Technology, 31(9): 3433–3443.
  45. 45.Ye, B.; Chang, H.; Ma, B.; and Shan, S. 2022a. Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework. Proceedings of the European Conference on Computer Vision.
  46. 46.Ye, J.; Tian, J.; Yan, M.; Yang, X.; Wang, X.; Zhang, J.; He, L.; and Lin, X. 2022b. Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15502–15512.
  47. 47.Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016. Modeling context in referring expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, 69–85. Springer.
  48. 48.Zhang, Z.; Peng, H.; Fu, J.; Li, B.; and Hu, W. 2020. Ocean: Object-aware Anchor-free Tracking. In Proceedings of the European Conference on Computer Vision.
  49. 49.Zhou, L.; Zhou, Z.; Mao, K.; and He, Z. 2023. Joint Visual Grounding and Tracking with Natural Language Specification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 23151–23160.
  50. 50.Zhu, C.; Zhou, Y.; Shen, Y.; Luo, G.; Pan, X.; Lin, M.; Chen, C.; Cao, L.; Sun, X.; and Ji, R. 2022. Seqtr: A simple yet universal network for visual grounding. In European Conference on Computer Vision, 598–615. Springer.

Citation

MLA
Ma, Y., et al. “Unifying Visual and Vision-Language Tracking via Contrastive Learning”. arXiv, 2024, http://arxiv.org/abs/2401.11228v1.
APA
Ma, Y., Tang, Y., Yang, W., Zhang, T., Zhang, J., & Kang, M. (2024). Unifying Visual and Vision-Language Tracking via Contrastive Learning. arXiv. http://arxiv.org/abs/2401.11228v1
Chicago
Ma, Y., Y. Tang, W. Yang, T. Zhang, J. Zhang, and M. Kang. 2024. “Unifying Visual and Vision-Language Tracking via Contrastive Learning”. arXiv. http://arxiv.org/abs/2401.11228v1.
Harvard
Ma, Y. et al. (2024) “Unifying Visual and Vision-Language Tracking via Contrastive Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.11228v1.
Vancouver
1. Ma Y, Tang Y, Yang W, Zhang T, Zhang J, Kang M (2024) Unifying Visual and Vision-Language Tracking via Contrastive Learning. arXiv

BibTeX

@article{ma2024unifying,
  title = {Unifying Visual and Vision-Language Tracking via Contrastive Learning},
  author = {Ma, Yinchao and Tang, Yuyang and Yang, Wenfei and Zhang, Tianzhu and Zhang, Jinpeng and Kang, Mengxue},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.11228v1},
  eprint = {2401.11228}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF