SwinTrack: A Simple and Strong Baseline for Transformer Tracking

Liting LinHeng FanZhipeng ZhangYong XuHaibin Ling

article2022NeurIPS555 citations

Presents a fully attentional Siamese tracking baseline that unifies Transformer-based feature extraction, fusion, and historical trajectory encoding to achieve state-of-the-art visual tracking accuracy and speed.

Listen

Visual object tracking is a critical capability in computer vision applications such as autonomous navigation, surveillance, and robotics. Recent advances have adapted attention-based Transformer models to improve tracking accuracy, but prior systems largely rely on hybrid designs that use traditional convolutional neural networks to extract image features and only apply Transformers for secondary feature fusion. This hybrid approach fails to realize the full benefits of attention mechanisms across the entire image representation process. At the same time, existing trackers often struggle to integrate temporal context efficiently without incurring substantial computational overhead.

The article demonstrates and evaluates SwinTrack, a streamlined, fully attentional tracking framework that eliminates convolutional feature extractors. The authors design an end-to-end architecture using a Swin Transformer backbone for both representation extraction and feature fusion, complemented by an ultra-lightweight "motion token" that embeds historical target coordinates to supply critical temporal context.

The evaluation was conducted across five major public tracking benchmarks, including LaSOT, LaSOText, TrackingNet, GOT-10k, and TNL2k. The authors assessed tracking accuracy, precision, and processing speeds across two primary variants: a high-capacity base model (SwinTrack-B-384) and an efficiency-focused lightweight model (SwinTrack-T-224), supported by systematic ablation studies on feature extractors, fusion schemes, and decoder designs.

The analysis reveals several key findings. First, the full-capacity SwinTrack-B-384 set a new state-of-the-art record on the demanding LaSOT benchmark with a 71.3% success score, exceeding previous best-in-class methods by 3.1 to 4.2 absolute percentage points while maintaining real-time processing at roughly 45 frames per second. Second, the lightweight SwinTrack-T-224 matched or outperformed existing top models across benchmarks while operating at approximately 98 frames per second, which is two to five times faster than competing state-of-the-art trackers. Third, ablation testing showed that replacing a standard ResNet backbone with a Transformer backbone yielded substantial accuracy gains, boosting LaSOT success scores by 2.5 percentage points and LaSOText scores by 5.1 percentage points. Finally, incorporating the historical motion token consistently improved accuracy across all datasets—particularly when resolving confusing distractors—with virtually no computational overhead.

These findings indicate that unifying image representation and fusion within a pure Transformer architecture significantly improves target discrimination while simplifying overall pipeline design. By discarding complicated components such as multi-scale feature hierarchies, query-based decoders, and continuous template updates, engineering teams can achieve superior tracking accuracy with reduced structural complexity and lower operational latency.

For technical leaders and system architects, the article supports adopting fully attentional architectures for production tracking pipelines. Deployments with strict latency and compute constraints should utilize the lightweight configuration (SwinTrack-T-224) to achieve high-throughput processing at near 100 frames per second, whereas accuracy-critical applications should leverage the base variant. Practitioners should also integrate trajectory-based motion embeddings as an inexpensive mechanism to enhance robustness against visual distractors. Future engineering efforts can explore incorporating richer multi-frame contextual signals into the sequence architecture.

Confidence in these findings is high given the consistent gains demonstrated across multiple independent benchmarks. However, leaders should note that optimal performance depends on standard motion assumptions, such as local temporal continuity, and the availability of frame-rate metadata to properly scale trajectory sampling intervals during inference.

Cover for SwinTrack: A Simple and Strong Baseline for Transformer Tracking

Abstract

Recently Transformer has been largely explored in tracking and shown state-of-the-art (SOTA) performance. However, existing efforts mainly focus on fusing and enhancing features generated by convolutional neural networks (CNNs). The potential of Transformer in representation learning remains under-explored. In this paper, we aim to further unleash the power of Transformer by proposing a simple yet efficient fully-attentional tracker, dubbed SwinTrack, within classic Siamese framework. In particular, both representation learning and feature fusion in SwinTrack leverage the Transformer architecture, enabling better feature interactions for tracking than pure CNN or hybrid CNN-Transformer frameworks. Besides, to further enhance robustness, we present a novel motion token that embeds historical target trajectory to improve tracking by providing temporal context. Our motion token is lightweight with negligible computation but brings clear gains. In our thorough experiments, SwinTrack exceeds existing approaches on multiple benchmarks. Particularly, on the challenging LaSOT, SwinTrack sets a new record with 0.713 SUC score. It also achieves SOTA results on other benchmarks. We expect SwinTrack to serve as a solid baseline for Transformer tracking and facilitate future research. Our codes and results are released at https://github.com/LitingLin/SwinTrack.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Tracking via Vision-Motion Transformer
  • 3.1 Swin-Transformer for Feature Extraction
  • 3.2 Vision-Motion Representation Learning
  • 3.3 Discussion
  • 3.4 Head and Loss
  • 4 Experiments
  • 4.1 Implementation
  • 4.2 Comparisons to State-of-the-arts
  • 4.3 Ablation Experiment
  • 5 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — SwinTrack Fully Attentional Siamese Tracking Architecture

    model/method

    SwinTrack is a pure Transformer-based Siamese tracking framework comprising three primary components:

    1. Feature Extraction Backbone: A Swin Transformer extracting hierarchical visual tokens from both the target template z∈RHz×Wz×3z \in \mathbb{R}^{H_z \times W_z \times 3} and the search region x∈RHx×Wx×3x \in \mathbb{R}^{H_x \times W_x \times 3}. The outputs from stage 3 of the backbone with stride s=16s=16 yield template tokens φ(z)∈RHzsWzs×C\varphi(z) \in \mathbb{R}^{\frac{H_z}{s}\frac{W_z}{s} \times C} and search tokens φ(x)∈RHxsWxs×C\varphi(x) \in \mathbb{R}^{\frac{H_x}{s}\frac{W_x}{s} \times C}, where CC is the hidden channel dimension of the network.

    2. Vision-Motion Encoder-Decoder Feature Fusion:

    • An Encoder consisting of NN Transformer blocks using multi-head self-attention (MSA) and feed-forward networks (FFN) applied to concatenated template and search tokens, enabling symmetric cross-feature interaction and self-enhancement with shared weights.
    • A single-layer Decoder using multi-head cross-attention (MCA) that integrates visual representations with a historical Motion Token representing the target's past trajectory, producing the vision-motion search representation fvm∈RHxs×Wxs×Cf_{vm} \in \mathbb{R}^{\frac{H_x}{s} \times \frac{W_x}{s} \times C}.
    1. Prediction Head: Two 3-layer multi-layer perceptron (MLP) branches receiving fvmf_{vm} to predict an IoU-aware classification score map rcls∈R(Hx×Wx)×1r_{cls} \in \mathbb{R}^{(H_x \times W_x) \times 1} and a bounding box coordinate regression map rreg∈R(Hx×Wx)×4r_{reg} \in \mathbb{R}^{(H_x \times W_x) \times 4}.
  2. Knowl 2 — Historical Motion Token Formulation in SwinTrack

    model/method

    To capture temporal dynamics with minimal computational overhead, SwinTrack represents target trajectory context as a single compact motion token Emotion∈R1×dE_{\text{motion}} \in \mathbb{R}^{1 \times d}.

    1. Trajectory Sampling: Given past bounding boxes T={o1,…,ot}T = \{o_1, \dots, o_t\} where oi=(oix1,oiy1,oix2,oiy2)o_i = (o_i^{x1}, o_i^{y1}, o_i^{x2}, o_i^{y2}) represents the top-left and bottom-right corner coordinates at frame ii, nn historical states are sampled at a fixed frame interval Δ\Delta: T={os(1),os(2),…,os(n)},where s(i)=max⁡(t−i×Δ,1)T = \{o_{s(1)}, o_{s(2)}, \dots, o_{s(n)}\}, \quad \text{where } s(i) = \max(t - i \times \Delta, 1) If sequence frame rate ff is available, the interval is adjusted to Δ30f\frac{\Delta}{30} f assuming a standard 30 fps.

    2. Crop-Invariant Coordinate Transformation: To match the search frame pre-processing (where a patch is cropped with center (ix,iy)(i_x, i_y) and scale factors (sx,sy)(s_x, s_y), and mapped to center (ox,oy)(o_x, o_y) in the cropped region via xo=(xi−ix)sx+oxx_o = (x_i - i_x)s_x + o_x), the same affine mapping is applied to sampled coordinates, producing transformed coordinates Tˉ={oˉs(1),…,oˉs(n)}\bar{T} = \{\bar{o}_{s(1)}, \dots, \bar{o}_{s(n)}\}.

    3. Coordinate Quantization and Embedding Lookup: Let (w,h)=(Wxs,Hxs)(w, h) = (\frac{W_x}{s}, \frac{H_x}{s}) denote the spatial dimensions of the search feature map, and gg denote quantization granularity. Four embedding matrices W∈R(g+1)×deW \in \mathbb{R}^{(g+1) \times d_e} embed each coordinate element independently. The coordinate index quantization function is: n(o,l)={⌊ol×g⌋if valid (target present and within search frame)g+1else (padding/invalid token index)n(o, l) = \begin{cases} \lfloor \frac{o}{l} \times g \rfloor & \text{if valid (target present and within search frame)} \\ g + 1 & \text{else (padding/invalid token index)} \end{cases} For each sampled box oˉs(i)\bar{o}_{s(i)}, the index tuple is: o^s(i)=[n(oˉs(i)x1,w),  n(oˉs(i)y1,h),  n(oˉs(i)x2,w),  n(oˉs(i)y2,h)]\hat{o}_{s(i)} = \left[ n(\bar{o}_{s(i)}^{x1}, w), \; n(\bar{o}_{s(i)}^{y1}, h), \; n(\bar{o}_{s(i)}^{x2}, w), \; n(\bar{o}_{s(i)}^{y2}, h) \right]

    4. Concatenation: Looked-up coordinate embeddings for all nn sampled trajectory steps are concatenated into the single motion token Emotion∈R1×dE_{\text{motion}} \in \mathbb{R}^{1 \times d} (where d=4n⋅ded = 4n \cdot d_e). Construction requires only table lookups and vector concatenation, adding negligible computational cost.

  3. Knowl 3 — Concatenation-Based Vision-Motion Feature Fusion

    model/method

    SwinTrack fuses template, search, and motion features using a concatenation-based Transformer encoder and a cross-attention decoder:

    1. Transformer Encoder: The template tokens φ(z)\varphi(z) and search region tokens φ(x)\varphi(x) are concatenated along the spatial token dimension: fm1=Concat(φ(z),φ(x))f_m^1 = \text{Concat}(\varphi(z), \varphi(x)) For layers l=1,…,Ll = 1, \dots, L: fml′=fml+MSA(LN(fml))f_m^{l\prime} = f_m^l + \text{MSA}(\text{LN}(f_m^l)) fml+1=fml′+FFN(LN(fml′))f_m^{l+1} = f_m^{l\prime} + \text{FFN}(\text{LN}(f_m^{l\prime})) where MSA\text{MSA} denotes Multi-Head Self-Attention, LN\text{LN} is Layer Normalization, and FFN\text{FFN} is a two-layer MLP with GELU activation. After layer LL, the mixed sequence is decoupled: fzL,fxL=DeConcat(fmL)f_z^L, f_x^L = \text{DeConcat}(f_m^L) This concatenation-based self-attention computes intra-template, intra-search, and cross-attention symmetrically while sharing parameters and reducing computational latency.

    2. Vision-Motion Decoder: The decoder applies a single Multi-Head Cross-Attention (MCA) block to fuse search tokens fxLf_x^L with the concatenation of the motion token EmotionE_{\text{motion}}, template tokens fzLf_z^L, and search tokens fxLf_x^L: fmD=Concat(Emotion,fzL,fxL)f_m^D = \text{Concat}(E_{\text{motion}}, f_z^L, f_x^L) fvm′=fxL+MCA(LN(fxL),LN(fmD))f_{vm}^\prime = f_x^L + \text{MCA}(\text{LN}(f_x^L), \text{LN}(f_m^D)) fvm=fvm′+FFN(LN(fvm′))f_{vm} = f_{vm}^\prime + \text{FFN}(\text{LN}(f_{vm}^\prime)) where fxLf_x^L acts as the query and fmDf_m^D acts as key and value. The resulting representation fvm∈RHxs×Wxs×Cf_{vm} \in \mathbb{R}^{\frac{H_x}{s} \times \frac{W_x}{s} \times C} is forwarded to the prediction heads.

  4. Knowl 4 — IoU-Aware Classification and Bounding Box Regression Objectives

    model/method

    SwinTrack utilizes two 3-layer MLP heads on top of the vision-motion feature representation fvmf_{vm}:

    • Classification Head predicts classification response map rcls∈R(Hx×Wx)×1r_{cls} \in \mathbb{R}^{(H_x \times W_x) \times 1}.
    • Regression Head predicts bounding box offsets rreg∈R(Hx×Wx)×4r_{reg} \in \mathbb{R}^{(H_x \times W_x) \times 4}.
    1. Classification Loss (LclsL_{cls}): Rather than binary cross-entropy on discrete 0/1 targets, the classification branch is trained to predict the IoU-Aware Classification Score (IACS), where the continuous target is the Intersection over Union IoU(b,b^)\text{IoU}(b, \hat{b}) between the predicted bounding box bb and ground truth box b^\hat{b}. Training uses the Varifocal Loss (LVFLL_{\text{VFL}}): Lcls=LVFL(p,IoU(b,b^))L_{cls} = L_{\text{VFL}}(p, \text{IoU}(b, \hat{b})) where pp is the predicted IACS, bb is the predicted bounding box, and b^\hat{b} is the ground-truth box.

    2. Regression Loss (LregL_{reg}): Bounding box regression uses Generalized IoU (GIoU) loss, weighted by the predicted classification score pp to prioritize high-confidence locations: Lreg=∑j1{IoU(bj,b^)>0}[p⋅LGIoU(bj,b^)]L_{reg} = \sum_j \mathbf{1}_{\{\text{IoU}(b_j, \hat{b}) > 0\}} \left[ p \cdot L_{\text{GIoU}}(b_j, \hat{b}) \right] where jj indexes spatial positions on the feature map, bjb_j is the predicted bounding box at position jj, and negative samples (where IoU(bj,b^)≤0\text{IoU}(b_j, \hat{b}) \le 0) are ignored.

  5. Knowl 5 — SwinTrack Online Tracking Inference Procedure

    algorithm

    During online inference, SwinTrack maintains a trajectory buffer of recent target bounding boxes to construct motion tokens and applies a Hanning penalty window for spatial smoothing.

    Input: Initial video frame I1I_1 with target bounding box b1b_1, video frames ItI_t for t=2,…,Tt = 2, \dots, T, confidence threshold θconf\theta_{\text{conf}}, penalty weight γ\gamma, sampling parameters n,Δn, \Delta.
    Output: Predicted target bounding boxes btb_t for t=2,…,Tt = 2, \dots, T.
    Initialize trajectory buffer with o1=b1o_1 = b_1.
    Crop template patch zz from I1I_1 centered at target with background area factor 2.
    Extract template tokens φ(z)\varphi(z) using Swin-Transformer backbone.
    for t=2t = 2 to TT do:
        Crop search region xtx_t from ItI_t centered at center of bt−1b_{t-1} with background factor 4.
        Extract search tokens φ(xt)\varphi(x_t) using Swin-Transformer backbone.
        
        Sample nn past boxes {os(1),…,os(n)}\{o_{s(1)}, \dots, o_{s(n)}\} from trajectory history where s(i)=max⁡(t−1−i⋅Δ,1)s(i) = \max(t - 1 - i \cdot \Delta, 1).
        Apply search-crop affine transformation to map sampled boxes to search coordinate frame.
        Quantize transformed coordinates into indices and lookup embeddings to form EmotionE_{\text{motion}}.
        
        Encode fused tokens fmL=Encoder(Concat(φ(z),φ(xt)))f_m^L = \text{Encoder}(\text{Concat}(\varphi(z), \varphi(x_t))).
        Deconcatenate fmLf_m^L into fzLf_z^L and fxLf_x^L.
        Decode vision-motion features fvm=Decoder(fxL,Concat(Emotion,fzL,fxL))f_{vm} = \text{Decoder}(f_x^L, \text{Concat}(E_{\text{motion}}, f_z^L, f_x^L)).
        
        Predict classification score map rcls∈R(Hx×Wx)×1r_{\text{cls}} \in \mathbb{R}^{(H_x \times W_x) \times 1} and regression map rreg∈R(Hx×Wx)×4r_{\text{reg}} \in \mathbb{R}^{(H_x \times W_x) \times 4}.
        Apply Hanning window penalty: rcls′=(1−γ)⋅rcls+γ⋅hr_{\text{cls}}' = (1 - \gamma) \cdot r_{\text{cls}} + \gamma \cdot h, where hh is a 2D Hanning window matching rclsr_{\text{cls}}.
        
        Find optimal location j∗=arg⁡max⁡jrcls′[j]j^* = \arg\max_j r_{\text{cls}}'[j] with raw confidence score ct=rcls[j∗]c_t = r_{\text{cls}}[j^*].
        Derive target bounding box btb_t from regression prediction rreg[j∗]r_{\text{reg}}[j^*].
        
        if ct≥θconfc_t \ge \theta_{\text{conf}} then:
            Append ot=bto_t = b_t to trajectory history.
        else:
            Append ot=−∞o_t = -\infty (marked as invalid/lost) to trajectory history.
        end if
    end for
  6. Knowl 6 — SwinTrack Model Configurations and Training Protocol

    experimental setup

    SwinTrack is instantiated in two primary configurations:

    • SwinTrack-T-224: Swin Transformer-Tiny backbone (pretrained on ImageNet-1k), template image size 112×112112 \times 112, search region size 224×224224 \times 224, hidden feature dimension C=384C = 384, number of encoder blocks N=4N = 4, quantization granularity g=14g = 14.
    • SwinTrack-B-384: Swin Transformer-Base backbone (pretrained on ImageNet-22k), template image size 192×192192 \times 192, search region size 384×384384 \times 384, hidden feature dimension C=512C = 512, number of encoder blocks N=8N = 8, quantization granularity g=24g = 24.

    Both configurations use the stage-3 features of Swin Transformer with stride s=16s = 16. The motion token samples n=16n = 16 historical boxes with sampling interval Δ=15\Delta = 15 (adjusted by sequence frame rate ff as 1530f\frac{15}{30}f). Confidence threshold θconf=0.4\theta_{\text{conf}} = 0.4 for LaSOT and 0.30.3 for other datasets. For GOT-10k specific evaluation, n=8,Δ=8n = 8, \Delta = 8 without frame rate adjustment.

    Training Details:

    • Datasets: Training splits of LaSOT, TrackingNet, GOT-10k (with 1,000 overlapping videos removed), and COCO 2017. For the GOT-10k protocol, only the GOT-10k train split is used.
    • Optimization: AdamW optimizer with initial learning rate 5×10−45 \times 10^{-4} (5×10−55 \times 10^{-5} for backbone), weight decay 10−410^{-4}, on 8 NVIDIA V100 GPUs.
    • Schedule: 300 epochs total (131,072 samples/epoch), 3-epoch linear warmup, learning rate decayed by a factor of 10 after epoch 210. (For GOT-10k-only protocol: 150 epochs total, decayed after epoch 120).
    • Regularization: DropPath rate of 0.1 applied to backbone and encoder.
  7. Knowl 7 — Benchmark Tracking Performance on Five Datasets

    data/table

    SwinTrack was evaluated across five tracking benchmarks: LaSOT (280 test videos), LaSOText\text{LaSOT}_{\text{ext}} (150 challenging test videos with distractors), TrackingNet (test split), GOT-10k (180 test videos evaluated under the train-split only protocol), and TNL2k (700 test videos). Metrics reported are Success score (SUC, %), Precision (P, %), Average Overlap (AO, %), and Success Rates at thresholds 0.50 (SR0.5\text{SR}_{0.5}) and 0.75 (SR0.75\text{SR}_{0.75}).

    Tracker LaSOT LaSOText\text{LaSOT}_{\text{ext}} TrackingNet GOT-10k TNL2k
    SUC P SUC P SUC P AO SR0.5\text{SR}_{0.5} SR0.75\text{SR}_{0.75} SUC P
    C-RPN 45.5 44.3 27.5 32.0 66.9 61.9 - - - - -
    SiamRPN++ 49.6 49.1 34.0 39.6 73.3 69.4 51.7 61.6 32.5 41.3 41.2
    Ocean 56.0 56.6 - - - - 61.1 72.1 47.3 38.4 37.7
    DiMP 56.9 56.7 39.2 45.1 74.0 68.7 61.1 71.7 49.2 44.7 43.4
    LTMU 57.2 57.2 41.4 47.3 - - - - - 48.5 47.3
    SiamR-CNN 64.8 - - - 81.2 80.0 64.9 72.8 59.7 52.3 52.8
    STMTrack 60.6 63.3 - - 80.3 76.7 64.2 73.7 57.5 - -
    AutoMatch 58.3 59.9 37.6 43.0 76.0 72.6 65.2 76.6 54.3 - -
    TrDiMP 63.9 61.4 - - 78.4 73.1 67.1 77.7 58.3 - -
    TransT 64.9 69.0 - - 81.4 80.3 67.1 76.8 60.9 51.0 -
    STARK 67.1 - - - 82.0 - 68.8 78.1 64.1 - -
    KeepTrack 67.1 70.2 48.2 - - - - - - - -
    SwinTrack-T-224 67.2 70.8 47.6 53.9 81.1 78.4 71.3 81.9 64.5 53.0 53.2
    SwinTrack-B-384 71.3 76.5 49.1 55.6 84.0 82.8 72.4 80.5 67.8 55.9 57.1

    SwinTrack-B-384 achieves state-of-the-art results across all five datasets, notably achieving 71.3% SUC on LaSOT (surpassing STARK by 4.2 percentage points and crossing the 70% threshold) and 72.4% AO on GOT-10k. The compact SwinTrack-T-224 matches or exceeds previous complex models (e.g., KeepTrack, STARK) while operating at real-time speeds.

  8. Knowl 8 — Ablation Study on SwinTrack Structural and Algorithmic Components

    data/table

    Ablation experiments evaluated the impact of individual architectural choices on SwinTrack-T-224 (without motion token) across LaSOT, LaSOText\text{LaSOT}_{\text{ext}}, TrackingNet, and GOT-10k (full training dataset protocol).

    Configuration LaSOT SUC (%) LaSOText\text{LaSOT}_{\text{ext}} SUC (%) TrackingNet SUC (%) GOT-10k mAO (%) Speed (fps) Params (M)
    Baseline (SwinTrack-T-224, no motion token) 66.7 46.9 80.8 70.9 98 22.7
    Replace Swin backbone with ResNet-50 64.2 41.8 79.5 68.2 121 20.0
    Replace concat-fusion with cross-attention fusion 66.6 45.4 80.2 69.3 72 34.6
    Replace decoder with target query-based decoder 66.6 43.2 79.6 69.0 91 25.3
    Replace untied pos. enc. with absolute sine pos. enc. 65.7 45.0 80.0 70.0 103 21.6
    Replace Varifocal loss with binary cross-entropy (BCE) 66.2 46.7 79.4 68.2 98 22.7
    Remove Hanning penalty window at inference 65.7 46.0 80.0 69.6 98 22.7

    Key observations:

    1. Backbone: Swin Transformer provides substantial gains over ResNet-50 (+2.5% SUC on LaSOT, +5.1% SUC on LaSOText\text{LaSOT}_{\text{ext}}).
    2. Fusion Strategy: Concatenation-based fusion outperforms cross-attention fusion while using fewer parameters (22.7M vs 34.6M) and operating faster (98 fps vs 72 fps).
    3. Query-Based Decoder: A target-query decoder degrades tracking accuracy across all benchmarks compared to the proposed encoder-based matching structure.
    4. Positional Encoding: Multi-source untied positional encoding outperforms standard absolute sine positional encoding by 0.8% to 1.9% SUC across datasets.
    5. Loss: Varifocal loss targeting IoU-aware classification scores improves mAO on GOT-10k by 2.7% over standard BCE loss.
  9. Knowl 9 — Ablation Analysis of the Motion Token and Trajectory Embedding

    data/table

    The contribution of the historical motion token was evaluated on SwinTrack-T-224 and SwinTrack-B-384 across four tracking benchmarks: LaSOT, LaSOText\text{LaSOT}_{\text{ext}}, TrackingNet, and GOT-10k.

    Variant / Configuration LaSOT SUC (%) LaSOText\text{LaSOT}_{\text{ext}} SUC (%) TrackingNet SUC (%) GOT-10k mAO (%) Speed (fps)
    SwinTrack-T-224 (with motion token) 67.2 47.6 81.1 71.3 96
    SwinTrack-T-224 (without motion token) 66.7 47.0 80.8 70.0 98
    SwinTrack-T-224 (learnable embedding token) 66.3 45.2 81.2 70.0 96
    SwinTrack-B-384 (with motion token) 71.3 49.1 84.0 72.4 45
    SwinTrack-B-384 (without motion token) 70.2 48.5 84.0 70.7 45

    The historical motion token consistently improves performance across models, providing gains of +0.5% to +1.1% SUC on LaSOT and +1.3% to +1.7% mAO on GOT-10k with virtually no impact on inference speed (e.g., 96 fps vs 98 fps on SwinTrack-T-224, 45 fps on SwinTrack-B-384). Replacing the motion token with a generic learnable embedding vector degrades performance (66.3% SUC on LaSOT), confirming that the improvements stem directly from the embedded temporal trajectory dynamics rather than added network capacity.

  10. Knowl 10 — Tracking Speed, MACs, and Parameter Efficiency Comparison

    data/table

    Inference throughput (frames per second, fps), Multiply-Accumulate operations (MACs in Giga-operations), and parameter counts (Millions) were measured and compared against state-of-the-art Transformer-based trackers.

    Tracker Speed (fps) MACs (G) Params (M)
    TrDiMP 26 - -
    TransT 50 - 23
    STARK-ST50 42 10.9 24
    STARK-ST101 32 18.5 42
    SwinTrack-T-224 98 6.4 23
    SwinTrack-B-384 45 69.7 91

    SwinTrack-T-224 operates at 98 fps with 6.4 G MACs and 23M parameters, running more than 2×2\times faster than STARK-ST50 (42 fps, 10.9 G MACs) and 3×3\times faster than STARK-ST101 (32 fps, 18.5 G MACs) with comparable model size. SwinTrack-B-384, despite its larger capacity (91M parameters, 69.7 G MACs), maintains a real-time frame rate of 45 fps, which is faster than STARK-ST101.

Coverage note — The detailed generalization formulas for extending 1D untied positional encoding to multi-dimensional multi-source token inputs described in the supplementary appendix were omitted as they adapt existing prior formulation (Ke et al., 2021).

References

  1. 1.Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., Torr, P.H., 2016. Fully-convolutional siamese networks for object tracking, in: ECCVW.
  2. 2.Bhat, G., Danelljan, M., Gool, L.V., Timofte, R., 2019. Learning discriminative model prediction for tracking, in: ICCV.
  3. 3.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S., 2020. End-to-end object detection with transformers, in: ECCV.
  4. 4.Chen, C.F.R., Fan, Q., Panda, R., 2021a. Crossvit: Cross-attention multi-scale vision transformer for image classification, in: ICCV.
  5. 5.Chen, X., Yan, B., Zhu, J., Wang, D., Yang, X., Lu, H., 2021b. Transformer tracking, in: CVPR.
  6. 6.Dai, K., Zhang, Y., Wang, D., Li, J., Lu, H., Yang, X., 2020. High-performance long-term tracking with meta-updater, in: CVPR.
  7. 7.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2021. An image is worth 16x16 words: Transformers for image recognition at scale, in: ICLR.
  8. 8.Fan, H., Bai, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Huang, M., Liu, J., Xu, Y., et al., 2021. Lasot: A high-quality large-scale single object tracking benchmark. International Journal of Computer Vision 129, 439–461.
  9. 9.Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., Ling, H., 2019. Lasot: A high-quality benchmark for large-scale single object tracking, in: CVPR.
  10. 10.Fan, H., Ling, H., 2019. Siamese cascaded region proposal networks for real-time visual tracking, in: CVPR.
  11. 11.Fan, H., Ling, H., 2021. Cract: Cascaded regression-align-classification for robust visual tracking, in: IROS.
  12. 12.Fu, Z., Liu, Q., Fu, Z., Wang, Y., 2021. Stmtrack: Template-free visual tracking with space-time memory networks, in: CVPR.
  13. 13.Han, W., Dong, X., Khan, F.S., Shao, L., Shen, J., 2021. Learning to fuse asymmetric feature maps in siamese trackers, in: CVPR.
  14. 14.He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: CVPR.
  15. 15.Huang, L., Zhao, X., Huang, K., 2019. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 1562–1577.
  16. 16.Ke, G., He, D., Liu, T.Y., 2021. Rethinking positional encoding in language pre-training, in: ICLR.
  17. 17.Krizhevsky, A., Sutskever, I., Hinton, G.E., 2012. Imagenet classification with deep convolutional neural networks. NIPS .
  18. 18.Larsson, G., Maire, M., Shakhnarovich, G., 2016. Fractalnet: Ultra-deep neural networks without residuals, in: ICLR.
  19. 19.Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., Yan, J.S., 2019. Evolution of siamese visual tracking with very deep networks, in: CVPR.
  20. 20.Li, B., Yan, J., Wu, W., Zhu, Z., Hu, X., 2018. High performance visual tracking with siamese region proposal network, in: CVPR.
  21. 21.Li, X., Wang, W., Wu, L., Chen, S., Hu, X., Li, J., Tang, J., Yang, J., 2020. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection, in: NeurIPS.
  22. 22.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014. Microsoft coco: Common objects in context, in: ECCV.
  23. 23.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B., 2021. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV .
  24. 24.Loshchilov, I., Hutter, F., 2019. Decoupled weight decay regularization, in: ICLR.
  25. 25.Mayer, C., Danelljan, M., Paudel, D.P., Van Gool, L., 2021. Learning target candidate association to keep track of what not to track, in: ICCV.
  26. 26.Muller, M., Bibi, A., Giancola, S., Alsubaihi, S., Ghanem, B., 2018. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild, in: ECCV.
  27. 27.Ren, S., He, K., Girshick, R., Sun, J., 2015. Faster r-cnn: Towards real-time object detection with region proposal networks, in: NIPS.
  28. 28.Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., Savarese, S., 2019. Generalized intersection over union .
  29. 29.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H., 2021. Training data-efficient image transformers & distillation through attention, in: ICML.
  30. 30.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017. Attention is all you need, in: NeurIPS.
  31. 31.Voigtlaender, P., Luiten, J., Torr, P.H., Leibe, B., 2020. Siam r-cnn: Visual tracking by re-detection, in: CVPR.
  32. 32.Wang, N., Zhou, W., Wang, J., Li, H., 2021a. Transformer meets tracker: Exploiting temporal context for robust visual tracking, in: CVPR.
  33. 33.Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L., 2021b. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions, in: ICCV.
  34. 34.Wang, X., Shu, X., Zhang, Z., Jiang, B., Wang, Y., Tian, Y., Wu, F., 2021c. Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark, in: CVPR.
  35. 35.Xu, Y., Wang, Z., Li, Z., Yuan, Y., Yu, G., 2020. Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines, in: AAAI.
  36. 36.Yan, B., Peng, H., Fu, J., Wang, D., Lu, H., 2021. Learning spatio-temporal transformer for visual tracking, in: ICCV.
  37. 37.Yu, Y., Xiong, Y., Huang, W., Scott, M.R., 2020. Deformable siamese attention networks for visual object tracking, in: CVPR.
  38. 38.Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.H., Tay, F.E., Feng, J., Yan, S., 2021. Tokens-to-token vit: Training vision transformers from scratch on imagenet, in: ICCV.
  39. 39.Zhang, H., Wang, Y., Dayoub, F., Sünderhauf, N., 2021a. Varifocalnet: An iou-aware dense object detector, in: CVPR.
  40. 40.Zhang, Z., Liu, Y., Wang, X., Li, B., Hu, W., 2021b. Learn to match: Automatic matching network design for visual tracking, in: ICCV.
  41. 41.Zhang, Z., Peng, H., Fu, J., Li, B., Hu, W., 2020. Ocean: Object-aware anchor-free tracking, in: ECCV.

Citation

MLA
Lin, L., et al. “SwinTrack: A Simple and Strong Baseline for Transformer Tracking”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 16743–54, https://proceedings.neurips.cc/paper_files/paper/2022/file/6a5c23219f401f3efd322579002dbb80-Paper-Conference.pdf.
APA
Lin, L., Fan, H., Zhang, Z., Xu, Y., & Ling, H. (2022). SwinTrack: A Simple and Strong Baseline for Transformer Tracking. Advances in Neural Information Processing Systems, 35, 16743–16754. https://proceedings.neurips.cc/paper_files/paper/2022/file/6a5c23219f401f3efd322579002dbb80-Paper-Conference.pdf
Chicago
Lin, L., H. Fan, Z. Zhang, Y. Xu, and H. Ling. 2022. “SwinTrack: A Simple and Strong Baseline for Transformer Tracking”. Advances in Neural Information Processing Systems 35: 16743–54. https://proceedings.neurips.cc/paper_files/paper/2022/file/6a5c23219f401f3efd322579002dbb80-Paper-Conference.pdf.
Harvard
Lin, L. et al. (2022) “SwinTrack: A Simple and Strong Baseline for Transformer Tracking”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 16743–16754. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/6a5c23219f401f3efd322579002dbb80-Paper-Conference.pdf.
Vancouver
1. Lin L, Fan H, Zhang Z, Xu Y, Ling H (2022) SwinTrack: A Simple and Strong Baseline for Transformer Tracking. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 16743–16754

BibTeX

@inproceedings{lin2022swintrack,
  title = {SwinTrack: A Simple and Strong Baseline for Transformer Tracking},
  author = {Lin, Liting and Fan, Heng and Zhang, Zhipeng and Xu, Yong and Ling, Haibin},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {16743-16754},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/6a5c23219f401f3efd322579002dbb80-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors