Transformer Tracking

Xin ChenBin YanJiawen ZhuDong WangXiaoyun YangHuchuan Lu

article2021CVPR1,390 citations

Introduces TransT, a high-speed visual tracking framework that replaces traditional cross-correlation with a Transformer-based attention fusion network to capture global semantic dependencies and achieve top performance across major tracking benchmarks at 50 fps.

Listen

Visual object tracking—the task of estimating the location and bounding box of a target throughout a video sequence—is essential for autonomous driving, robotics, and surveillance systems. Most prevailing tracking architectures rely heavily on correlation operations to match the initial target template with subsequent search regions. However, correlation acts as a local linear comparison that causes substantial loss of semantic information and lacks global context, making systems vulnerable to occlusions, deformations, and similar background distractors.

The article evaluates whether replacing traditional correlation with an attention-based network can resolve these accuracy bottlenecks. The authors demonstrate a novel tracking architecture, termed Transformer Tracking (TransT), which completely eliminates correlation operations in favor of attention-driven feature fusion.

The proposed framework extracts visual features using a modified standard convolutional network and fuses them through stacked attention modules before passing them to a prediction head. The fusion engine employs two mechanisms: ego-context augment modules using self-attention to enrich local context within each branch, and cross-feature augment modules using cross-attention to dynamically associate the template and search areas. The tracker was evaluated against roughly twenty competing methods across six diverse, large-scale visual tracking benchmarks (including LaSOT, TrackingNet, and GOT-10k) and trained offline on extensive standard datasets.

The experimental findings show that TransT establishes a new state-of-the-art performance level across major benchmarks. On the large-scale LaSOT dataset, TransT achieved a 64.9% success rate (AUC), outperforming established models like Ocean (56.0%) and DiMP (56.9%). On the TrackingNet and GOT-10k benchmarks, it achieved top-tier scores of 81.4% AUC and 72.3% average overlap, respectively. Ablation experiments revealed that removing attention mechanisms and substituting back correlation led to steep drops in accuracy, confirming the necessity of the attention modules. Furthermore, TransT operates at approximately 50 frames per second on a single GPU, easily exceeding real-time deployment requirements and proving more than ten times faster than competing high-accuracy frameworks such as SiamR-CNN, which runs at under 5 frames per second.

These results indicate that computer vision systems can achieve superior tracking robustness without the computational overhead of complex online update routines or anchor adjustments. By shifting to attention-based fusion, engineering teams can build tracking pipelines that maintain precision during severe environmental disruptions while preserving high operational throughput and lower computational latency.

Organizations developing real-time computer vision applications should consider adopting attention-based fusion architectures in place of correlation-centric pipelines. Teams can implement the open-source codebase to evaluate performance on domain-specific edge devices. Because the evaluation was conducted primarily on curated academic benchmarks and requires high-end GPU acceleration for 50 frames per second processing, stakeholders should conduct targeted pilot tests in edge environments with constrained hardware or specialized optical sensors before broad operational deployment.

  • Paper: Fully-Convolutional Siamese Networks for Object Tracking, Luca Bertinetto et al. (2016). This paper establishes the foundational fully-convolutional Siamese tracking framework based on cross-correlation feature matching that TransT seeks to replace with an attention-based fusion network.
  • Paper: SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks, Bo Li et al. (2018). This work introduces deep Siamese tracking architectures and depth-wise cross-correlation fusion, providing the direct baseline and performance bottleneck analyzed by TransT.
  • Paper: SuperGlue: Learning Feature Matching With Graph Neural Networks, Paul-Edouard Sarlin et al. (2020). This paper introduces the concept of combining self-attention and cross-attention modules for matching visual features across distinct image regions, directly inspiring TransT's feature fusion architecture.
  • Paper: On the Relationship between Self-Attention and Convolutional Layers, Jean-Baptiste Cordonnier et al. (2020). This theoretical study demonstrates how self-attention layers can generalize and surpass standard convolutional operations, establishing the conceptual basis for replacing correlation with attention in vision.
  • Paper: Object Tracking Benchmark, Yi Wu et al. (2015). This paper provides the standard benchmarking methodology and evaluation protocols widely used to assess visual tracking performance across diverse challenges.
Cover for Transformer Tracking

Abstract

Correlation acts as a critical role in the tracking field, especially in recent popular Siamese-based trackers. The correlation operation is a simple fusion manner to consider the similarity between the template and the search region. However, the correlation operation itself is a local linear matching process, leading to lose semantic information and fall into local optimum easily, which may be the bottleneck of designing high-accuracy tracking algorithms. Is there any better feature fusion method than correlation? To address this issue, inspired by Transformer, this work presents a novel attention-based feature fusion network, which effectively combines the template and search region features solely using attention. Specifically, the proposed method includes an ego-context augment module based on self-attention and a cross-feature augment module based on cross-attention. Finally, we present a Transformer tracking (named TransT) method based on the Siamese-like feature extraction backbone, the designed attention-based fusion mechanism, and the classification and regression head. Experiments show that our TransT achieves very promising results on six challenging datasets, especially on large-scale LaSOT, TrackingNet, and GOT-10k benchmarks. Our tracker runs at approximatively 50 fps on GPU. Code and models are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Transformer Tracking
  • 3.1 Overall Architecture
  • 3.2 Ego-Context Augment and Cross-Feature Augment Modules
  • 3.3 Training Loss
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Evaluation on TrackingNet, LaSOT and GOT-10k Datasets
  • 4.3 Ablation Study and Analysis
  • 4.4 Evaluation on Other Datasets
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — TransT Transformer Tracking Architecture

    model/method

    Transformer Tracking (TransT) is a visual object tracking framework that replaces cross-correlation operations with an attention-based feature fusion network. The pipeline consists of three core components:

    1. Feature Extraction Backbone: Given a template image patch z∈R3×Hz0×Wz0z \in \mathbb{R}^{3 \times H_{z0} \times W_{z0}} (centered on the target in the initial frame and expanded by twice the target side length, cropped and resized to 128×128128 \times 128) and a search region patch x∈R3×Hx0×Wx0x \in \mathbb{R}^{3 \times H_{x0} \times W_{x0}} (centered on the target in the previous frame and expanded by four times the target side length, cropped and resized to 256×256256 \times 256), a modified ResNet-50 backbone extracts feature maps. The fifth stage of ResNet-50 is removed, the downsampling stride of the fourth stage is reduced from 2 to 1 to preserve spatial resolution, and the 3×33 \times 3 convolutions in the fourth stage employ dilated convolutions with a dilation rate/stride of 2 to expand the receptive field. The backbone outputs feature maps fz∈RC×Hz×Wzf_z \in \mathbb{R}^{C \times H_z \times W_z} and fx∈RC×Hx×Wxf_x \in \mathbb{R}^{C \times H_x \times W_x}, where C=1024C = 1024, Hz=Wz=16H_z = W_z = 16, and Hx=Wx=32H_x = W_x = 32.

    2. Feature Fusion Network: A 1×11 \times 1 convolution reduces the channel dimension of fzf_z and fxf_x from C=1024C = 1024 to d=256d = 256, producing fz0∈Rd×Hz×Wzf_{z0} \in \mathbb{R}^{d \times H_z \times W_z} and fx0∈Rd×Hx×Wxf_{x0} \in \mathbb{R}^{d \times H_x \times W_x}. These feature maps are flattened spatially into sets of dd-dimensional feature vectors fz1∈Rd×(HzWz)f_{z1} \in \mathbb{R}^{d \times (H_z W_z)} and fx1∈Rd×(HxWx)f_{x1} \in \mathbb{R}^{d \times (H_x W_x)}. The fusion network enhances intra-branch context via Ego-Context Augment (ECA) modules and fuses cross-branch information via Cross-Feature Augment (CFA) modules across N=4N = 4 stacked fusion layers, followed by a final decoding CFA module that outputs an enhanced search feature representation f∈Rd×(HxWx)f \in \mathbb{R}^{d \times (H_x W_x)}.

    3. Prediction Head: An anchor-free head receives the fused feature map ff and directly predicts foreground/background classification probabilities and bounding box coordinates for each of the HxWx=1024H_x W_x = 1024 spatial positions.

  2. Knowl 2 — Ego-Context Augment Module

    model/method

    The Ego-Context Augment (ECA) module enhances contextual representations within an individual feature branch (either template or search region) using multi-head self-attention with a residual connection.

    Given an input feature sequence X∈Rd×NxX \in \mathbb{R}^{d \times N_x} (where dd is the feature dimension and NxN_x is the number of spatial positions) and spatial positional encodings Px∈Rd×NxP_x \in \mathbb{R}^{d \times N_x} computed via sinusoidal functions, the ECA module computes:

    XEC=X+MultiHead(X+Px,X+Px,X)X_{EC} = X + \text{MultiHead}(X + P_x, X + P_x, X)

    where XEC∈Rd×NxX_{EC} \in \mathbb{R}^{d \times N_x} is the output. The multi-head attention operation is defined as:

    MultiHead(Q,K,V)=Concat(H1,…,Hnh)WO\text{MultiHead}(Q, K, V) = \text{Concat}(H_1, \dots, H_{n_h}) W^O

    Hi=Attention(QWiQ,KWiK,VWiV)H_i = \text{Attention}(Q W_i^Q, K W_i^K, V W_i^V)

    Attention(Q′,K′,V′)=softmax(Q′(K′)⊤dk)V′\text{Attention}(Q', K', V') = \text{softmax}\left(\frac{Q' (K')^\top}{\sqrt{d_k}}\right) V'

    where nh=8n_h = 8 is the number of attention heads, d=dm=256d = d_m = 256, dk=dv=dm/nh=32d_k = d_v = d_m / n_h = 32, and WiQ∈Rdm×dkW_i^Q \in \mathbb{R}^{d_m \times d_k}, WiK∈Rdm×dkW_i^K \in \mathbb{R}^{d_m \times d_k}, WiV∈Rdm×dvW_i^V \in \mathbb{R}^{d_m \times d_v}, and WO∈Rnhdv×dmW^O \in \mathbb{R}^{n_h d_v \times d_m} are learnable projection parameter matrices.

  3. Knowl 3 — Cross-Feature Augment Module

    model/method

    The Cross-Feature Augment (CFA) module adaptively fuses features from two distinct feature branches using multi-head cross-attention followed by a feed-forward network (FFN) with residual connections.

    Let Xq∈Rd×NqX_q \in \mathbb{R}^{d \times N_q} be the feature sequence of the query branch with corresponding sinusoidal spatial positional encoding Pq∈Rd×NqP_q \in \mathbb{R}^{d \times N_q}, and let Xkv∈Rd×NkvX_{kv} \in \mathbb{R}^{d \times N_{kv}} be the feature sequence of the key/value branch with spatial positional encoding Pkv∈Rd×NkvP_{kv} \in \mathbb{R}^{d \times N_{kv}}, where d=256d = 256.

    The CFA module performs cross-attention and feed-forward transformation as follows:

    X~CF=Xq+MultiHead(Xq+Pq,Xkv+Pkv,Xkv)\tilde{X}_{CF} = X_q + \text{MultiHead}(X_q + P_q, X_{kv} + P_{kv}, X_{kv})

    XCF=X~CF+FFN(X~CF)X_{CF} = \tilde{X}_{CF} + \text{FFN}(\tilde{X}_{CF})

    where XCF∈Rd×NqX_{CF} \in \mathbb{R}^{d \times N_q} is the module output. The feed-forward network FFN(⋅)\text{FFN}(\cdot) consists of two linear transformations with a ReLU activation between them:

    FFN(u)=max⁡(0,uW1+b1)W2+b2\text{FFN}(u) = \max(0, u W_1 + b_1) W_2 + b_2

    where W1,W2W_1, W_2 are weight matrices and b1,b2b_1, b_2 are bias vectors.

  4. Knowl 4 — Structure of the Attention-Based Feature Fusion Network

    model/method

    The TransT feature fusion network takes flattened feature maps fz1∈Rd×(HzWz)f_{z1} \in \mathbb{R}^{d \times (H_z W_z)} (template branch, length Nz=16×16=256N_z = 16 \times 16 = 256, d=256d = 256) and fx1∈Rd×(HxWx)f_{x1} \in \mathbb{R}^{d \times (H_x W_x)} (search region branch, length Nx=32×32=1024N_x = 32 \times 32 = 1024, d=256d = 256).

    The core of the network consists of N=4N = 4 identical fusion layers arranged in sequence. Each fusion layer contains:

    1. Two parallel Ego-Context Augment (ECA) modules: one applied to the template branch to capture intra-template context, and one applied to the search region branch to capture intra-search context.
    2. Two parallel Cross-Feature Augment (CFA) modules: one receiving the enhanced template features as queries and the enhanced search features as keys/values, and the second receiving the enhanced search features as queries and the enhanced template features as keys/values.

    Following the N=4N = 4 stacked fusion layers, an additional decoding CFA module is applied. This final CFA takes the search region branch features as the query input (XqX_q) and the template branch features as the key/value input (XkvX_{kv}), generating the final fused search representation f∈Rd×(HxWx)f \in \mathbb{R}^{d \times (H_x W_x)} containing 10241024 feature vectors of dimension d=256d = 256 for subsequent bounding box prediction.

  5. Knowl 5 — TransT Prediction Head and Direct Coordinate Regression

    model/method

    The prediction head in TransT evaluates the fused feature map f∈Rd×(HxWx)f \in \mathbb{R}^{d \times (H_x W_x)} (d=256d = 256, HxWx=1024H_x W_x = 1024) without using predefined anchor boxes or anchor points.

    The head consists of two parallel branches:

    1. Classification Branch: A 3-layer multilayer perceptron (MLP) with hidden dimension d=256d = 256 and ReLU activations that outputs HxWx=1024H_x W_x = 1024 binary classification probabilities corresponding to foreground/background assignment for each spatial position.
    2. Regression Branch: A 3-layer MLP with hidden dimension d=256d = 256 and ReLU activations that directly outputs HxWx=1024H_x W_x = 1024 4-dimensional vectors representing normalized bounding box coordinates (xmin⁡,ymin⁡,xmax⁡,ymax⁡)(x_{\min}, y_{\min}, x_{\max}, y_{\max}) relative to the search region size.

    By regressing normalized coordinates directly from each feature vector, the framework eliminates heuristic anchor matching and scale/ratio hyperparameter tuning.

  6. Knowl 6 — TransT Multi-Task Training Loss

    equation

    TransT is trained end-to-end using a multi-task loss combining binary cross-entropy classification loss Lcls\mathcal{L}_{cls} and bounding box regression loss Lreg\mathcal{L}_{reg}.

    Predictions corresponding to spatial positions located within the ground-truth target bounding box are labeled as positive samples (yj=1y_j = 1), while all other positions are negative samples (yj=0y_j = 0). To address class imbalance between foreground and background locations, the loss contributed by negative samples is down-weighted by a factor of 16.

    The classification loss is:

    Lcls=−∑j[yjlog⁡(pj)+116(1−yj)log⁡(1−pj)]\mathcal{L}_{cls} = -\sum_{j} \left[ y_j \log(p_j) + \frac{1}{16}(1 - y_j) \log(1 - p_j) \right]

    where pj∈[0,1]p_j \in [0, 1] is the predicted foreground probability for sample jj.

    The regression loss is calculated exclusively over positive samples (yj=1y_j = 1) as a linear combination of ℓ1\ell_1-norm loss and Generalized Intersection over Union (GIoU) loss:

    Lreg=∑j1{yj=1}[λGLGIoU(bj,b^)+λ1L1(bj,b^)]\mathcal{L}_{reg} = \sum_{j} \mathbf{1}_{\{y_j = 1\}} \left[ \lambda_G \mathcal{L}_{GIoU}(b_j, \hat{b}) + \lambda_1 \mathcal{L}_1(b_j, \hat{b}) \right]

    where bjb_j is the predicted normalized bounding box for positive sample jj, b^\hat{b} is the normalized ground-truth bounding box, λG=2\lambda_G = 2, and λ1=5\lambda_1 = 5.

  7. Knowl 7 — TransT Online Tracking Inference with Window Penalty

    algorithm

    During online tracking, TransT processes each video frame sequentially and applies a 2D Hanning window penalty to suppress distractor locations far from the target's previous state.

    Input: Initial video frame with ground-truth target bounding box b0b_0, subsequent frame sequence {It}t=1T\{I_t\}_{t=1}^T
    Input: Hanning window weighting parameter w=0.49w = 0.49
    Output: Predicted bounding boxes {bt}t=1T\{b_t\}_{t=1}^T
    Crop template patch zz of size 128×128128 \times 128 centered on target from I0I_0 (expanded by 2×2\times target size)
    Extract backbone features fzf_z from zz
    for each frame t=1t = 1 to TT do
        Crop search region patch xtx_t of size 256×256256 \times 256 centered on target position from bt−1b_{t-1} (expanded by 4×4\times target size)
        Extract backbone features fx,tf_{x,t} from xtx_t
        Pass fzf_z and fx,tf_{x,t} into Feature Fusion Network to obtain fused feature map ft∈R256×1024f_t \in \mathbb{R}^{256 \times 1024}
        Compute classification scores score∈R32×32\text{score} \in \mathbb{R}^{32 \times 32} and bounding boxes {bj}j=11024\{b_j\}_{j=1}^{1024} via Prediction Head
        Construct 2D Hanning window scoreh∈R32×32\text{score}_h \in \mathbb{R}^{32 \times 32}
        Compute penalized confidence scores:
            scorew=(1−w)⋅score+w⋅scoreh\text{score}_w = (1 - w) \cdot \text{score} + w \cdot \text{score}_h
        Select bounding box bt=bj∗b_t = b_{j^*} corresponding to index j∗=arg⁡max⁡j(scorew)jj^* = \arg\max_j (\text{score}_w)_j
        Map normalized coordinates btb_t back to image coordinate space of ItI_t
    end for

    The tracker operates at approximately 50 frames per second (fps) on a single GPU.

  8. Knowl 8 — Tracking Performance on Large-Scale Benchmarks

    data/table

    TransT was evaluated against state-of-the-art visual object trackers on three large-scale tracking benchmarks: LaSOT (1400 videos; evaluated via Success AUC, Normalized Precision PNormP_{Norm}, and Precision PP), TrackingNet (511 test sequences; evaluated via AUC, PNormP_{Norm}, and PP), and GOT-10k (180 test sequences; evaluated via Average Overlap AO, Success Rate at 0.5 threshold SR0.5\text{SR}_{0.5}, and Success Rate at 0.75 threshold SR0.75\text{SR}_{0.75} under the zero-overlap evaluation protocol). TransT-GOT indicates training restricted solely to the GOT-10k training split.

    Method LaSOT TrackingNet GOT-10k
    AUC PNormP_{Norm} PP AUC PNormP_{Norm} PP AO SR0.5\text{SR}_{0.5} SR0.75\text{SR}_{0.75}
    TransT (Ours) 64.9 73.8 69.0 81.4 86.7 80.3 72.3 82.4 68.2
    TransT-GOT (Ours) - - - - - - 67.1 76.8 60.9
    SiamR-CNN 64.8 72.2 - 81.2 85.4 80.0 64.9 72.8 59.7
    Ocean 56.0 65.1 56.6 - - - 61.1 72.1 47.3
    KYS 55.4 63.3 - 74.0 80.0 68.8 63.6 75.1 51.5
    DCFST - - - 75.2 80.9 70.0 63.8 75.3 49.8
    SiamFC++ 54.4 62.3 54.7 75.4 80.0 70.5 59.5 69.5 47.9
    PrDiMP 59.8 68.8 60.8 75.8 81.6 70.4 63.4 73.8 54.3
    CGACD 51.8 62.6 - 71.1 80.0 69.3 - - -
    SiamAttn 56.0 64.8 - 75.2 81.7 - - - -
    MAML 52.3 - - 75.7 82.2 72.5 - - -
    D3S - - - 72.8 76.8 66.4 59.7 67.6 46.2
    SiamCAR 50.7 60.0 51.0 - - - 56.9 67.0 41.5
    SiamBAN 51.4 59.8 52.1 - - - - - -
    DiMP 56.9 65.0 56.7 74.0 80.1 68.7 61.1 71.7 49.2
    SiamRPN++ 49.6 56.9 49.1 73.3 80.0 69.4 51.7 61.6 32.5
    ATOM 51.5 57.6 50.5 70.3 77.1 64.8 55.6 63.4 40.2
    ECO 32.4 33.8 30.1 55.4 61.8 49.2 31.6 30.9 11.1
    MDNet 39.7 46.0 37.3 60.6 70.5 56.5 29.9 30.3 9.9
    SiamFC 33.6 42.0 33.9 57.1 66.3 53.3 34.8 35.3 9.8

    TransT outperforms existing correlation-based trackers across all three benchmarks. On GOT-10k, TransT-GOT (trained strictly following the GOT-10k protocol) outperforms SiamR-CNN by 2.2% in AO. While SiamR-CNN achieves comparable accuracy on LaSOT and TrackingNet, it runs at under 5 fps, whereas TransT operates in real time at approximately 50 fps.

  9. Knowl 9 — Ablation Analysis on Transformer Fusion, Correlation, and Post-Processing

    data/table

    Ablation studies were conducted across LaSOT, TrackingNet, and GOT-10k to evaluate the components of TransT: Ego-Context Augment (ECA), Cross-Feature Augment (CFA), the original standard encoder-decoder Transformer structure (TransT(ori)), and Hanning window post-processing (models without post-processing are denoted by -np). In models replacing CFA with correlation, cross-attention is removed and the final fusion layer applies depth-wise cross-correlation.

    Method ECA CFA Corr. LaSOT TrackingNet GOT-10k
    AUC PNormP_{Norm} PP AUC PNormP_{Norm} PP AO SR0.5\text{SR}_{0.5} SR0.75\text{SR}_{0.75}
    TransT ✓ ✓ 64.9 73.8 69.0 81.4 86.7 80.3 72.3 82.4 68.2
    TransT ✓ 62.9 71.9 66.2 81.1 86.2 79.1 70.6 81.2 65.7
    TransT ✓ ✓ 57.7 65.4 59.5 77.5 82.2 74.0 62.8 72.2 54.8
    TransT ✓ 47.7 48.6 41.7 68.8 71.4 60.9 50.9 58.0 33.3
    TransT-np ✓ ✓ 62.9 71.5 66.9 81.1 86.4 80.0 71.5 81.5 67.5
    TransT-np ✓ 61.0 69.6 64.5 80.0 85.0 77.9 68.1 78.3 64.0
    TransT-np ✓ ✓ 57.3 65.2 58.8 76.2 80.8 72.8 61.4 70.7 53.7
    TransT-np ✓ 35.3 17.9 20.1 46.5 40.3 27.4 38.2 36.8 7.0
    TransT(ori) - - - 62.3 71.1 66.2 81.3 86.1 78.9 70.3 80.2 65.8
    TransT(ori)-np - - - 60.9 69.4 64.8 80.9 85.6 78.4 68.6 78.2 65.1

    Key takeaways:

    1. Attention vs. Correlation: Replacing CFA cross-attention with depth-wise correlation drops LaSOT AUC from 64.9% to 57.7% (with ECA) and to 47.7% (without ECA). Without post-processing (-np), the pure correlation model drops drastically to 35.3% AUC on LaSOT, showing that correlation-based trackers rely heavily on spatial window priors due to degraded semantic representations.
    2. TransT vs. Original Transformer: TransT outperforms TransT(ori) (which feeds template to encoder and search region to decoder) across all metrics (LaSOT AUC 64.9% vs. 62.3%), confirming the advantage of the alternating ECA/CFA architecture.
    3. Ego-Context Augmentation: Removing ECA decreases LaSOT AUC by 2.0% (64.9% to 62.9%), demonstrating the value of intra-feature context modeling.
  10. Knowl 10 — Tracking Performance on NFS, OTB2015, and UAV123 Datasets

    data/table

    TransT was evaluated on three standard tracking benchmarks: Need for Speed (NFS, 30 fps version, evaluating tracking of fast-moving objects), Object Tracking Benchmark (OTB2015 / OTB100, 100 sequences), and UAV123 (123 low-altitude aerial videos). Tracking performance is reported using overall Success Area Under Curve (AUC %).

    Dataset Ours (TransT) PrDiMP DiMP SiamRPN++ ATOM ECO MDNet
    NFS 65.7 63.5 62.0 50.2 58.4 46.6 42.2
    OTB2015 69.4 69.6 68.4 69.6 66.9 69.1 67.8
    UAV123 69.1 68.0 65.3 61.3 64.2 53.2 52.8

    TransT achieves the top AUC score on NFS (65.7%, exceeding PrDiMP by 2.2%) and on UAV123 (69.1%, exceeding PrDiMP by 1.1%), while achieving performance competitive with state-of-the-art algorithms on OTB2015 (69.4% vs. 69.6% for PrDiMP and SiamRPN++).

Coverage note — Qualitative visualizations of intermediate attention map patterns (Figure 4) and dataset attribute breakdown radar charts (Figure 5) were omitted as standalone knowls because their core scientific findings are fully captured within the architectural, algorithmic, and tabular ablation knowls.

References

  1. 1.Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip H S Torr. Fully-convolutional siamese networks for object tracking. In ECCVW, 2016. 1, 2, 3, 6, 7
  2. 2.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning discriminative model prediction for tracking. In ICCV, 2019. 1, 2, 6, 7, 8
  3. 3.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Know Your Surroundings: Exploiting scene information for object tracking. In ECCV, 2020. 6
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2, 4, 5
  5. 5.Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. Siamese box adaptive network for visual tracking. In CVPR, 2020. 6, 7
  6. 6.Jongwon Choi, Hyung Jin Chang, Sangdoo Yun, Tobias Fischer, Yiannis Demiris, and Jin Young Choi. Attentional correlation filter network for adaptive visual tracking. In CVPR, 2017. 2
  7. 7.Janghoon Choi, Junseok Kwon, and Kyoung Mu Lee. Deep meta learning for real-time target-aware visual tracking. In ICCV, 2019. 1, 2
  8. 8.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ECO: Efficient convolution operators for tracking. In CVPR, 2017. 2, 6, 7, 8
  9. 9.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ATOM: Accurate tracking by overlap maximization. In CVPR, 2019. 1, 2, 6, 7, 8
  10. 10.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic regression for visual tracking. In CVPR, 2020. 1, 6, 7, 8
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, 2019. 2
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. CoRR, abs/2010.11929, 2020. 2
  13. 13.Fei Du, Peng Liu, Wei Zhao, and Xianglong Tang. Correlation-guided attention for corner detection based visual tracking. In CVPR, 2020. 1, 3, 6, 7
  14. 14.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. LaSOT: A high-quality benchmark for large-scale single object tracking. In CVPR, 2019. 6, 7, 8
  15. 15.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In ICAIS, 2010. 6
  16. 16.Dongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang, and Shengyong Chen. SiamCAR: Siamese fully convolutional classification and regression for visual tracking. In CVPR, 2020. 6, 7
  17. 17.Anfeng He, Chong Luo, Xinmei Tian, and Wenjun Zeng. A twofold siamese network for real-time object tracking. In CVPR, 2018. 2
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 3, 6
  19. 19.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. TPAMI, 2019. 6, 7, 8
  20. 20.Hamed Kiani Galoogahi, Ashton Fagg, Chen Huang, Deva Ramanan, and Simon Lucey. Need for speed: A benchmark for higher frame rate object tracking. In ICCV, 2017. 8
  21. 21.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. SiamRPN++: Evolution of siamese visual tracking with very deep networks. In CVPR, 2019. 1, 2, 6, 7, 8
  22. 22.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High performance visual tracking with siamese region proposal network. In CVPR, 2018. 1, 2, 3, 7
  23. 23.Peixia Li, Dong Wang, Lijun Wang, and Huchuan Lu. Deep visual tracking: Review and experimental comparison. PR, 2018. 1
  24. 24.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 2, 6
  25. 25.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2018. 6
  26. 26.Alan Lukezic, Jiri Matas, and Matej Kristan. D3S - A discriminative single shot segmentation tracker. In CVPR, 2020. 6, 7
  27. 27.Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney. RWTH ASR Systems for LibriSpeech: hybrid vs attention. In INTERSPEECH, 2019. 2
  28. 28.Seyed Mojtaba Marvasti-Zadeh, Li Cheng, Hossein Ghanei-Yakhdan, and Shohreh Kasaei. Deep learning for visual tracking: A comprehensive survey. CoRR, abs/1912.00535, 2019. 1
  29. 29.Matthias Mueller, Neil Smith, and Bernard Ghanem. A benchmark and simulator for UAV tracking. In ECCV, 2016. 8
  30. 30.Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem. TrackingNet: A large-scale dataset and benchmark for object tracking in the wild. In ECCV, 2018. 6, 7, 8
  31. 31.Hyeonseob Nam and Bohyung Han. Learning multi–domain convolutional neural networks for visual tracking. In CVPR, 2016. 6, 7, 8
  32. 32.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In ICML, 2018. 2
  33. 33.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R–CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 2
  34. 34.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian D. Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, 2019. 5
  35. 35.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, and Michael Bernstein. ImageNet Large scale visual recognition challenge. IJCV, 2015. 6
  36. 36.Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Edouard Grave, Tatiana Likhomanenko, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert. End-to-end ASR: from supervised to semi-supervised learning with modern architectures. CoRR, abs/1911.08460, 2019. 2
  37. 37.Ran Tao, Efstratios Gavves, and Arnold W. M. Smeulders. Siamese instance search for tracking. In CVPR, 2016. 2
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017. 2, 4
  39. 39.Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr, and Bastian Leibe. Siam R-CNN: Visual tracking by re-detection. In CVPR, 2020. 6, 7
  40. 40.Guangting Wang, Chong Luo, Xiaoyan Sun, Zhiwei Xiong, and Wenjun Zeng. Tracking by Instance Detection: A meta-learning approach. In CVPR, 2020. 6, 7
  41. 41.Qiang Wang, Zhu Teng, Junliang Xing, Jin Gao, Weiming Hu, and Stephen J. Maybank. Learning Attentions: Residual attentional siamese network for high performance online visual tracking. In CVPR, 2018. 2
  42. 42.Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H. S. Torr. Fast online object tracking and segmentation: A unifying approach. In CVPR, 2019. 2
  43. 43.Yi Wu, Jongwoo Lim, and Ming Hsuan Yang. Object tracking benchmark. TPAMI, 2015. 8
  44. 44.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. SiamFC++: Towards robust and accurate visual tracking with target estimation guidelines. In AAAI, 2020. 1, 2, 6, 7
  45. 45.Bin Yan, Xinyu Zhang, Dong Wang, Huchuan Lu, and Xiaoyun Yang. Alpha-refine: Boosting tracking performance by precise bounding box estimation. In CVPR, 2021. 2
  46. 46.Yuechen Yu, Yilei Xiong, Weilin Huang, and Matthew R. Scott. Deformable siamese attention networks for visual object tracking. In CVPR, 2020. 1, 2, 6, 7
  47. 47.Lichao Zhang, Abel Gonzalez-Garcia, Joost van de Weijer, Martin Danelljan, and Fahad Shahbaz Khan. Learning the model update for siamese trackers. In ICCV, 2019. 1
  48. 48.Zhipeng Zhang, Houwen Peng, Jianlong Fu, Bing Li, and Weiming Hu. Ocean: Object-aware anchor-free tracking. In ECCV, 2020. 1, 2, 6, 7
  49. 49.Linyu Zheng, Ming Tang, Yingying Chen, Jinqiao Wang, and Hanqing Lu. Learning feature embeddings for discriminant model based tracking. In ECCV, 2020. 6
  50. 50.Zheng Zhu, Wei Wu, Wei Zou, and Junjie Yan. End-to-end flow correlation tracking with spatial-temporal attention. In CVPR, 2018. 2

Citation

MLA
Chen, X., et al. “Transformer Tracking”. arXiv, 2021, http://arxiv.org/abs/2103.15436v1.
APA
Chen, X., Yan, B., Zhu, J., Wang, D., Yang, X., & Lu, H. (2021). Transformer Tracking. arXiv. http://arxiv.org/abs/2103.15436v1
Chicago
Chen, X., B. Yan, J. Zhu, D. Wang, X. Yang, and H. Lu. 2021. “Transformer Tracking”. arXiv. http://arxiv.org/abs/2103.15436v1.
Harvard
Chen, X. et al. (2021) “Transformer Tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2103.15436v1.
Vancouver
1. Chen X, Yan B, Zhu J, Wang D, Yang X, Lu H (2021) Transformer Tracking. arXiv

BibTeX

@article{chen2021transformer,
  title = {Transformer Tracking},
  author = {Chen, Xin and Yan, Bin and Zhu, Jiawen and Wang, Dong and Yang, Xiaoyun and Lu, Huchuan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2103.15436v1},
  eprint = {2103.15436}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE