MixFormerV2: Efficient Fully Transformer Tracking

Yutao CuiTianhui SongGangshan WuLimin Wang

article2023NeurIPS145 citations

Presents MixFormerV2, a fully transformer tracking framework that eliminates dense convolutional operations by using learnable prediction tokens and novel knowledge distillation strategies to achieve real-time tracking speeds on both GPU and CPU platforms without sacrificing accuracy.

Listen

Visual object tracking—the task of estimating the location of an object across video frames given an initial bounding box—is critical for real-world technologies such as autonomous vehicles, robotics, and automated surveillance. Recent advances using transformer architectures have established state-of-the-art tracking accuracy. However, these models remain too computationally heavy and slow for practical deployment on standard central processing units (CPUs) or resource-constrained graphics processing units (GPUs) due to complex convolutional prediction heads and separate sample-quality estimation modules.

The article introduces and evaluates MixFormerV2, an efficient, fully transformer-based tracking architecture designed to eliminate dense convolutional operations entirely. The objective is to demonstrate that a streamlined transformer framework, coupled with a novel knowledge-distillation and model-reduction strategy, can achieve top-tier tracking accuracy while running at high speeds across both GPU and CPU hardware.

The researchers evaluated MixFormerV2 through extensive experiments across multiple standard tracking benchmarks, including LaSOT, TrackingNet, UAV123, TNL2K, and VOT2022. The method introduces four special learnable prediction tokens into a unified transformer backbone to jointly compress target and search information, followed by simple multi-layer feed-forward networks (MLPs) to predict coordinate probability distributions and target quality scores. To compress the architecture, the study applies a multi-stage distillation paradigm: dense-to-sparse distillation to transfer localization knowledge from heavy convolutional heads to sparse token heads, progressive depth pruning to drop transformer layers smoothly without starting training from scratch, and intermediate-teacher supervision combined with internal dimension reduction for lightweight CPU models.

The evaluation produced several key findings: First, the GPU-focused model (MixFormerV2-B) achieved an Area Under the Curve (AUC) of 70.6% on the LaSOT benchmark and 56.7% on TNL2k while operating at 165 frames per second (FPS), surpassing prior one-stream transformer trackers like OSTrack by 1.5% in AUC and roughly 57% in processing speed. Second, the compact version (MixFormerV2-S) set a new benchmark for lightweight tracking by operating at real-time speeds on standard CPUs (30 FPS) and 325 FPS on GPUs, outperforming leading efficient architectures like FEAR-L by 2.7% AUC on LaSOT. Third, the progressive model depth pruning strategy outperformed standard initialization techniques by 1.9% AUC, confirming that smoothly decaying redundant layers preserves vital learned representations during compression.

These results carry significant practical implications for deployment. By eliminating custom convolutional operators and region-of-interest pooling layers, MixFormerV2 provides a unified, hardware-friendly architecture that is simpler to maintain and port across edge devices. Organizations deploying computer vision systems can reduce hardware expenditure and energy costs while maintaining state-of-the-art tracking precision. Furthermore, achieving real-time performance on standard CPUs broadens the feasibility of advanced vision models in edge environments where dedicated GPU acceleration is cost-prohibitive.

Decision-makers and engineering teams should consider adopting this streamlined transformer design for applications requiring high-throughput or low-power video tracking. When deploying on high-end edge GPUs, MixFormerV2-B offers an optimal balance of top-tier accuracy and high throughput, while MixFormerV2-S serves as the primary candidate for CPU-only systems. For teams planning model compression pipelines, adopting progressive depth pruning rather than retraining pruned models from scratch is strongly recommended. Future development should focus on testing these models within operational vehicle and robotic platforms.

While confidence in the empirical benchmark results is high, some operational limitations remain. The multi-stage distillation process requires substantial upfront training time and compute—exceeding 100 hours on high-end hardware clusters for full model reduction. Additionally, qualitative analysis indicates that extreme visual occlusions or closely situated distracting objects can still cause prediction errors, warranting cautious testing in mission-critical or safety-sensitive operational settings.

No sufficiently relevant recommendations were found.

Cover for MixFormerV2: Efficient Fully Transformer Tracking

Abstract

Transformer-based trackers have achieved high accuracy on standard benchmarks. However, their efficiency remains an obstacle to practical deployment on both GPU and CPU platforms. In this paper, to mitigate this issue, we propose a fully transformer tracking framework based on the successful MixFormer tracker [14], coined as MixFormerV2, without any dense convolutional operation or complex score prediction module. We introduce four special prediction tokens and concatenate them with those from target template and search area. Then, we apply a unified transformer backbone on these mixed token sequence. These prediction tokens are able to capture the complex correlation between target template and search area via mixed attentions. Based on them, we can easily predict the tracking box and estimate its confidence score through simple MLP heads. To further improve the efficiency of MixFormerV2, we present a new distillation-based model reduction paradigm, including dense-to-sparse distillation and deep-to-shallow distillation. The former one aims to transfer knowledge from the dense-head based MixViT to our fully transformer tracker, while the latter one is for pruning the backbone layers. We instantiate two MixForemrV2 trackers, where the MixFormerV2-B achieves an AUC of 70.6% on LaSOT and AUC of 56.7% on TNL2k with a high GPU speed of 165 FPS, and the MixFormerV2-S surpasses FEAR-L by 2.7% AUC on LaSOT with a real-time CPU speed.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Fully Transformer Tracking: MixFormerV2
  • 3.2 Distillation-Based Model Reduction
  • 3.2.1 Dense-to-Sparse Distillation
  • 3.2.2 Deep-to-Shallow Distillation
  • 3.3 Training of MixFormerV2
  • 4 Experiments
  • 4.1 Implemented Details
  • 4.2 Exploration Studies
  • 4.2.1 Analysis on MixFormerV2 Framework
  • 4.2.2 Analysis on Dense-to-Sparse Distillation
  • 4.2.3 Analysis on Deep-to-Shallow Distillation
  • 4.2.4 Model Pruning Route
  • 4.3 Comparison with the Previous Methods
  • 5 Conclusion
  • Acknowledgement
  • References
  • Appendix
  • S.1 Details of Training Time
  • S.2 More Results on VOT2020 and GOT10k
  • S.3 More Ablation Studies
  • S.4 Visualization Results

Knowls

  1. Knowl 1 — MixFormerV2 Fully Transformer Tracking Architecture

    model/method

    MixFormerV2 is an end-to-end fully transformer visual tracking framework that eliminates dense convolutional regression heads (such as corner heads) and complex online score prediction modules (SPMs). The architecture comprises:

    1. Input Sequence: A mixed sequence consisting of target template image tokens, current search area image tokens, and four special learnable prediction tokens (tokenT,tokenL,tokenB,tokenRtoken_T, token_L, token_B, token_R) corresponding to the top, left, bottom, and right bounding box boundaries.
    2. Backbone: NN stacked Prediction-Token-Involved Mixed Attention Modules (P-MAM) that jointly extract features and model cross-attention across template, search area, and prediction tokens without convolutions.
    3. Distribution-based Localization Head: A shared two-layer Multi-Layer Perceptron (MLP) applied directly to each of the four output prediction tokens to predict 1D probability distributions for the target's bounding box coordinates.
    4. Score Head: A two-layer MLP head that operates on the mean vector of the four output prediction tokens to predict a scalar target quality/confidence score s∈Rs \in \mathbb{R} used for dynamic online template updating.
  2. Knowl 2 — Prediction-Token-Involved Mixed Attention Module

    equation

    Given template tokens tt, search area tokens ss, and four learnable prediction tokens ee, let qt,kt,vtq_t, k_t, v_t denote the query, key, and value representations for the template; qs,ks,vsq_s, k_s, v_s for the search area; and qe,ke,veq_e, k_e, v_e for the prediction tokens. The concatenated keys and values across all tokens are defined as:

    ktse=Concat(kt,ks,ke),vtse=Concat(vt,vs,ve)k_{tse} = \text{Concat}(k_t, k_s, k_e), \quad v_{tse} = \text{Concat}(v_t, v_s, v_e)

    The attention outputs for each token group within each transformer block are computed via an asymmetric mixed attention scheme:

    Attent=Softmax(qtktTd)vt\text{Atten}_t = \text{Softmax}\left(\frac{q_t k_t^T}{\sqrt{d}}\right) v_t

    Attens=Softmax(qsktseTd)vtse\text{Atten}_s = \text{Softmax}\left(\frac{q_s k_{tse}^T}{\sqrt{d}}\right) v_{tse}

    Attene=Softmax(qektseTd)vtse\text{Atten}_e = \text{Softmax}\left(\frac{q_e k_{tse}^T}{\sqrt{d}}\right) v_{tse}

    where dd denotes the feature channel dimension. The template tokens only attend to themselves to maintain computational efficiency during online tracking, while the search tokens and prediction tokens attend to the entire mixed token sequence.

  3. Knowl 3 — Token-Based Probability Distribution Regression and Confidence Scoring

    equation

    Instead of directly regressing absolute coordinate offsets, MixFormerV2 estimates bounding box boundaries as 1D probability density functions. For each coordinate X∈{T,L,B,R}X \in \{T, L, B, R\} (top, left, bottom, and right boundaries), a shared MLP predicts a probability distribution P^X(x)\hat{P}_X(x) over possible coordinate positions x∈Rx \in \mathbb{R}:

    P^X(x)=MLP(tokenX),X∈{T,L,B,R}\hat{P}_X(x) = \text{MLP}(token_X), \quad X \in \{T, L, B, R\}

    The final continuous bounding box coordinate BXB_X is calculated as the expected value over the predicted distribution:

    BX=EP^X[X]=∫RxP^X(x)dxB_X = \mathbb{E}_{\hat{P}_X}[X] = \int_{\mathbb{R}} x \hat{P}_X(x) dx

    The tracking quality score s∈Rs \in \mathbb{R} is predicted by passing the mean of the four prediction tokens through an MLP score head:

    s=MLP(14∑X∈{T,L,B,R}tokenX)s = \text{MLP}\left(\frac{1}{4} \sum_{X \in \{T, L, B, R\}} token_X\right)

  4. Knowl 4 — Dense-to-Sparse Knowledge Distillation via Marginal Distributions

    model/method

    To transfer dense localization knowledge from a teacher tracker with a 2D convolutional corner head (MixViT) to the sparse token-based MixFormerV2, the teacher's 2D joint corner probability maps PTL(x,y)P_{TL}(x, y) (top-left corner) and PBR(x,y)P_{BR}(x, y) (bottom-right corner) are projected into four 1D marginal coordinate distributions:

    PT(x)=∫RPTL(x,y)dy,PL(y)=∫RPTL(x,y)dxP_T(x) = \int_{\mathbb{R}} P_{TL}(x, y)dy, \quad P_L(y) = \int_{\mathbb{R}} P_{TL}(x, y)dx

    PB(x)=∫RPBR(x,y)dy,PR(y)=∫RPBR(x,y)dxP_B(x) = \int_{\mathbb{R}} P_{BR}(x, y)dy, \quad P_R(y) = \int_{\mathbb{R}} P_{BR}(x, y)dx

    These 1D marginal distributions serve as soft labels to supervise the student MixFormerV2's predicted coordinate distributions P^X(x)\hat{P}_X(x) for X∈{T,L,B,R}X \in \{T, L, B, R\} using Kullback-Leibler (KL) divergence loss:

    Lloc=∑X∈{T,L,B,R}LKL(P^X,PX)L_{loc} = \sum_{X \in \{T, L, B, R\}} L_{KL}(\hat{P}_X, P_X)

  5. Knowl 5 — Progressive Model Depth Pruning for Transformer Backbones

    model/method

    Progressive Model Depth Pruning (PMDP) compresses the transformer backbone by initializing the student network as an exact copy of the teacher network and progressively decaying the weights of a subset of layers E\mathcal{E} to zero during training.

    For any layer i∈Ei \in \mathcal{E} to be eliminated, its output calculation at training epoch tt is scaled by a cosine decay factor γ(t)\gamma(t):

    xi=γ(t)⋅(FFN(ATTN(xi−1)+xi−1)+ATTN(xi−1))+xi−1x_i = \gamma(t) \cdot \left(\text{FFN}(\text{ATTN}(x_{i-1}) + x_{i-1}) + \text{ATTN}(x_{i-1})\right) + x_{i-1}

    where ATTN\text{ATTN} is multi-head self-attention, FFN\text{FFN} is the feed-forward network, and Layer Normalization is omitted for clarity. The decay factor γ(t)\gamma(t) over the first mm epochs follows:

    γ(t)={0.5×(1+cos⁡(tmπ)),t≤m0,t>m\gamma(t) = \begin{cases} 0.5 \times \left(1 + \cos\left(\frac{t}{m}\pi\right)\right), & t \le m \\ 0, & t > m \end{cases}

    When t>mt > m, the layer acts as an identity mapping xi=xi−1x_i = x_{i-1} and is removed from the network. During training, the remaining student layers are supervised via feature mimicking (L2L_2 loss between intermediate student and teacher representations FiSF_i^S and FjTF_j^T) and output logits distillation.

  6. Knowl 6 — Total Training Loss for Distillation-Based Model Reduction

    equation

    The overall loss function LL for distillation training of a student tracker SS supervised by teacher tracker TT and ground-truth bounding box BgtB^{gt} is formulated as:

    L=λ1L1(BS,Bgt)+λ2Lciou(BS,Bgt)+λ3Ldist(S,T)L = \lambda_1 L_1(B^S, B^{gt}) + \lambda_2 L_{ciou}(B^S, B^{gt}) + \lambda_3 L_{dist}(S, T)

    where:

    • BSB^S is the bounding box derived from the expectation of the student's predicted coordinate distributions P^X\hat{P}_X.
    • L1(BS,Bgt)L_1(B^S, B^{gt}) is the ℓ1\ell_1 bounding box regression loss.
    • Lciou(BS,Bgt)L_{ciou}(B^S, B^{gt}) is the Complete Intersection-over-Union (CIoU) loss.
    • Ldist(S,T)=Lloc+LfeatL_{dist}(S, T) = L_{loc} + L_{feat}, where Lloc=∑X∈{T,L,B,R}LKL(P^XS,PXT)L_{loc} = \sum_{X \in \{T, L, B, R\}} L_{KL}(\hat{P}_X^S, P_X^T) is the coordinate logits KL-divergence loss and Lfeat=∑(i,j)∈ML2(FiS,FjT)L_{feat} = \sum_{(i,j) \in \mathcal{M}} L_2(F_i^S, F_j^T) is the intermediate feature mimicking loss across matched layer pairs M\mathcal{M}.
    • λ1,λ2,λ3\lambda_1, \lambda_2, \lambda_3 are balancing hyperparameters.
  7. Knowl 7 — Multi-Stage Distillation Pipeline for Base and Real-Time CPU MixFormerV2

    model/method

    MixFormerV2 is instantiated into two variants using distinct multi-stage distillation pathways:

    1. MixFormerV2-B (8 P-MAM layers, MLP ratio 4.0, search image 288×288288 \times 288, template image 128×128128 \times 128, 58.8M parameters):

      • Stage 1 (Dense-to-Sparse Distillation): A 12-layer MixFormerV2 is trained using soft labels from a 12-layer MixViT-L (or MixViT-B) teacher.
      • Stage 2 (Deep-to-Shallow Distillation): The 12-layer MixFormerV2 is pruned down to 8 layers using progressive model depth pruning (PMDP) with m=40m=40 decay epochs.
      • Stage 3: The Score Prediction MLP is trained for 50 epochs.
    2. MixFormerV2-S (4 P-MAM layers, MLP ratio 1.0, search image 224×224224 \times 224, template image 112×112112 \times 112, 16.2M parameters for real-time CPU tracking):

      • Intermediate Teacher Distillation: To bridge the representation gap, the 12-layer MixFormerV2 is distilled to an 8-layer intermediate model, which is then distilled to a 4-layer model (MLP ratio 4.0).
      • MLP Dimension Reduction: The hidden feature dimension of the feed-forward network is reduced from ratio 4.0 to ratio 1.0 by initializing student weights as a truncated submatrix of the teacher weights (w′=w[:d1′,:d2′]w' = w[:d_1', :d_2']) and training with feature mimicking and logits distillation.
  8. Knowl 8 — State-of-the-Art Tracking Performance Comparison of MixFormerV2-B

    data/table

    MixFormerV2-B achieves state-of-the-art tracking accuracy while operating at 165 FPS on an Nvidia RTX 8000 GPU, outperforming previous one-stream transformer trackers.

    Method LaSOT LaSOText TNL2K TrackingNet Speed
    AUC PNormP_{Norm} P AUC P AUC P AUC PNormP_{Norm} P GPU (FPS)
    MixFormerV2-B 70.6 80.8 76.2 50.6 56.9 57.4 58.4 83.4 88.1 81.6 165
    MixFormerV2-B* 69.5 79.1 75.0 - - 56.6 57.1 82.9 87.6 81.0 165
    MixFormer 69.2 78.7 74.7 - - - - 83.1 88.1 81.6 25
    OSTrack-256 69.1 78.7 75.2 47.4 53.3 54.3 - 83.1 87.8 82.0 105
    SimTrack-B 69.3 78.5 - - - 54.8 53.8 82.3 86.5 - 40
    CTTrack-B 67.8 77.8 74.0 - - - - 82.5 87.1 80.3 40
    SwinTrack-T 67.2 - 70.8 47.6 53.9 53.0 53.2 81.1 - 78.4 98
    TransT 64.9 73.8 69.0 - - 50.7 51.7 81.4 86.7 80.3 50

    Note: MixFormerV2-B uses MixViT-L as the dense-to-sparse distillation teacher, whereas MixFormerV2-B uses MixViT-B as the teacher.

  9. Knowl 9 — CPU Real-Time Performance Comparison of MixFormerV2-S

    data/table

    MixFormerV2-S achieves real-time CPU tracking (30 FPS on an Intel Xeon Gold 6230R CPU @ 2.10GHz) and 325 FPS on GPU, outperforming prior lightweight and real-time trackers across multiple benchmarks.

    Method LaSOT LaSOText TNL2K TrackingNet Speed
    AUC PNormP_{Norm} P AUC P AUC P AUC PNormP_{Norm} P CPU (FPS)
    MixFormerV2-S 60.6 69.9 60.4 43.6 46.2 48.3 43.0 75.8 81.1 70.4 30
    FEAR-L 57.9 68.6 60.9 - - - - - - - -
    FEAR-XS 53.5 64.1 54.5 - - - - - - - 26
    HCAT 59.0 68.3 60.5 - - - - 76.6 82.6 72.9 45
    E.T.Track 59.1 - - - - - - 74.5 80.3 70.6 42
    LightTrack-LargeA 55.5 - 56.1 - - - - 73.6 78.8 70.0 -
    LightTrack-Mobile 53.8 - 53.7 - - - - 72.5 77.9 69.5 36
    STARK-Lightning 58.6 69.0 57.9 - - - - - - - 42
    DiMP 56.9 65.0 56.7 - - - - 74.0 80.1 68.7 15
    SiamFC++ 54.4 62.3 54.7 - - - - 75.4 80.0 70.5 20
  10. Knowl 10 — FLOPs Comparison Between Token Localization Head and Pyramidal Corner Head

    empirical result

    The computational complexity of the token-based localization head (T4T4) is over 1500×1500\times lower than the 2D convolutional Pyramidal Corner Head (Py-CornerPy\text{-}Corner) used in MixViT:

    • Token Localization Head (T4T4): Contains 4 coordinate MLPs (one for each boundary token), each comprising two linear layers with channel dimensions from 768 to 72: LoadT4=4×(768×768+768×72)=2,580,480≈2.58 MFLOPs\text{Load}_{T4} = 4 \times (768 \times 768 + 768 \times 72) = 2{,}580{,}480 \approx 2.58\text{ MFLOPs}

    • Pyramidal Corner Head (Py-CornerPy\text{-}Corner): Employs 24 convolutional layers with varying feature resolutions (ranging from 18×1818\times 18 to 72×7272\times 72) and 3×33\times 3 kernel size: LoadPy-Corner=3,902,587,776≈3.90 GFLOPs\text{Load}_{Py\text{-}Corner} = 3{,}902{,}587{,}776 \approx 3.90\text{ GFLOPs}

    Replacing the pyramidal corner head with the token-based head increases tracking frame rate on an 8-layer MixViT-B backbone from 90 FPS to 166 FPS on an Nvidia RTX 8000 GPU while reducing overall model compute from 27.2 GFLOPs to 22.5 GFLOPs.

  11. Knowl 11 — Ablation Analysis of Distillation, Student Initialization, and Regression Designs

    empirical result

    Ablation experiments on the LaSOT dataset demonstrate the efficacy of key design choices in MixFormerV2:

    1. Regression Head Design (12-layer ViT-B without distillation or score prediction):

      • Direct box prediction with 1 token (T1): 63.1% AUC, 112 FPS.
      • Distribution-based prediction with 4 tokens (T4): 67.5% AUC, 110 FPS.
      • Pyramidal convolutional corner head: 69.0% AUC, 92 FPS.
    2. Score Prediction Head (8-layer MixFormerV2-B):

      • Without score prediction: 68.9% AUC, 166 FPS.
      • With token-based MLP score head: 70.6% AUC (+1.7% AUC), 165 FPS (only 0.6% latency increase).
      • MixViT-B SPM baseline drops frame rate by 13.0% (92 FPS to 80 FPS) for +0.6% AUC.
    3. Dense-to-Sparse Distillation (12-layer MixFormerV2 student):

      • Without distillation: 67.5% AUC.
      • MixViT-B teacher (69.0% AUC): 68.9% AUC (+1.4% AUC).
      • MixViT-L teacher (71.5% AUC): 69.6% AUC (+2.1% AUC).
    4. Student Depth Initialization (4-layer student from 12-layer MixViT-B teacher):

      • MAE pre-trained first 4 layers (MAE-fir4): 62.9% AUC (LaSOT), 45.2% (LaSOText), 65.7% (UAV123).
      • Skipped teacher layers (Tea-skip4 - layers 3, 6, 9, 12): 64.4% AUC (LaSOT), 46.1% (LaSOText), 66.6% (UAV123).
      • Progressive Model Depth Pruning (PMDP, m=40m=40): 64.8% AUC (LaSOT), 47.1% (LaSOText), 67.5% (UAV123).
    5. Intermediate Teacher Distillation (12L teacher →\rightarrow 4L student):

      • Direct distillation (12L →\rightarrow 4L): 64.8% AUC.
      • With 8-layer intermediate teacher (12L →\rightarrow 8L →\rightarrow 4L): 65.5% AUC (+0.7% AUC).
  12. Knowl 12 — Training Computational Overhead of Multi-Stage Distillation

    limitation

    The primary limitation of MixFormerV2 is the substantial training overhead required by the sequential distillation and pruning pipeline, particularly for the lightweight MixFormerV2-S model. Producing MixFormerV2-S requires four sequential training stages of 500 epochs each on 8 Nvidia RTX 8000 GPUs:

    1. Dense-to-sparse distillation (12-layer MixViT to 12-layer MixFormerV2): ~43 hours.
    2. Deep-to-shallow distillation stage 1 (12 layers to 8 layers): ~42 hours.
    3. Deep-to-shallow distillation stage 2 (8 layers to 4 layers): ~35 hours.
    4. MLP hidden dimension pruning (MLP ratio 4.0 to 1.0), followed by 50 epochs of Score Head training.

    This multi-stage training pipeline requires over 120 total GPU hours on 8 high-end GPUs, making retraining and architectural exploration resource-intensive.

Coverage note — None was omitted; all contributed tracking models (MixFormerV2-B and MixFormerV2-S), attention formulations (P-MAM), distillation strategies (dense-to-sparse marginalization, PMDP, intermediate teacher, and MLP reduction), experimental benchmarks (LaSOT, LaSOText, TNL2K, TrackingNet, UAV123, VOT2022, VOT2020, GOT-10k), FLOPs analysis, and training limitations are fully captured.

References

  1. 1.Romero Adriana, Ballas Nicolas, K Samira Ebrahimi, Chassang Antoine, Gatta Carlo, and B Yoshua. Fitnets: Hints for thin deep nets. Proc. ICLR, 2, 2015.
  2. 2.Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr. Fully-convolutional siamese networks for object tracking. In Proceedings of the European Conference on Computer Vision, ECCV Workshops, 2016.
  3. 3.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning discriminative model prediction for tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, pages 6182–6191, 2019.
  4. 4.Philippe Blatter, Menelaos Kanakis, Martin Danelljan, and Luc Van Gool. Efficient visual tracking with exemplar transformers. arXiv preprint arXiv:2112.09686, 2021.
  5. 5.Vasyl Borsuk, Roman Vei, Orest Kupyn, Tetiana Martyniuk, Igor Krashenyi, and Jiˇri Matas. Fear: Fast, efficient, accurate and robust visual tracker. arXiv preprint arXiv:2112.07957, 2021.
  6. 6.Arnav Chavan, Zhiqiang Shen, Zhuang Liu, Zechun Liu, Kwang-Ting Cheng, and Eric P Xing. Vision transformer slimming: Multi-dimension searching in continuous optimization space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4931–4941, 2022.
  7. 7.Boyu Chen, Peixia Li, Lei Bai, Lei Qiao, Qiuhong Shen, Bo Li, Weihao Gan, Wei Wu, and Wanli Ouyang. Backbone is all your need: A simplified architecture for visual object tracking. In Proceedings of the European Conference on Computer Vision, ECCV, 2022.
  8. 8.Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling. Autoformer: Searching transformers for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12270–12280, 2021.
  9. 9.Xin Chen, Dong Wang, Dongdong Li, and Huchuan Lu. Efficient visual tracking via hierarchical cross-attention transformer. arXiv preprint arXiv:2203.13537, 2022.
  10. 10.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2021.
  11. 11.Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. Siamese box adaptive network for visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2020.
  12. 12.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Target transformed regression for accurate tracking. CoRR, 2021.
  13. 13.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Fully convolutional online tracking. Computer Vision and Image Understanding, 224:103547, 2022.
  14. 14.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Mixformer: End-to-end tracking with iterative mixed attention. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2022.
  15. 15.Yutao Cui, Cheng Jiang, Gangshan Wu, and Limin Wang. Mixformer: End-to-end tracking with iterative mixed attention, 2023.
  16. 16.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ATOM: accurate tracking by overlap maximization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019.
  17. 17.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic regression for visual tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2020.
  18. 18.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, ICLR, 2021.
  19. 19.Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. The Journal of Machine Learning Research, 20(1):1997–2017, 2019.
  20. 20.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. Lasot: A high-quality benchmark for large-scale single object tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019.
  21. 21.Chengyue Gong, Dilin Wang, Meng Li, Xinlei Chen, Zhicheng Yan, Yuandong Tian, Vikas Chandra, et al. Nasvit: Neural architecture search for efficient vision transformers with gradient conflict aware supernet training. In International Conference on Learning Representations, 2021.
  22. 22.Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. Compressing deep convolutional networks using vector quantization. arXiv preprint arXiv:1412.6115, 2014.
  23. 23.Jianyuan Guo, Kai Han, Yunhe Wang, Han Wu, Xinghao Chen, Chunjing Xu, and Chang Xu. Distilling object detectors via decoupled features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2154–2164, 2021.
  24. 24.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2022.
  25. 25.Yihui He, Xiangyu Zhang, and Jian Sun. Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE international conference on computer vision, pages 1389–1397, 2017.
  26. 26.João F. Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. High-speed tracking with kernelized correlation filters. IEEE Trans. Pattern Anal. Mach. Intell., 37(3):583–596, 2015.
  27. 27.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network (2015). arXiv preprint arXiv:1503.02531, 2, 2015.
  28. 28.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE Trans. Pattern Anal. Mach. Intell., 43(5):1562–1577, 2021.
  29. 29.Matej Kristan, Ales Leonardis, and et. al. The eighth visual object tracking VOT2020 challenge results. In Adrien Bartoli and Andrea Fusiello, editors, Proceedings of the European Conference on Computer Vision, ECCV Workshops, 2020.
  30. 30.Matej Kristan, Aleš Leonardis, Jiˇrí Matas, Michael Felsberg, Roman Pflugfelder, Joni-Kristian Kämäräinen, Hyung Jin Chang, Martin Danelljan, Luka Cehovin Zajc, Alan Lukeži ˇ c, et al. The tenth visual object ˇ tracking vot2022 challenge results. In ECCV 2022 Workshops, pages 431–460, 2023.
  31. 31.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. Siamrpn++: Evolution of siamese visual tracking with very deep networks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2019.
  32. 32.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High performance visual tracking with siamese region proposal network. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2018.
  33. 33.Quanquan Li, Shengying Jin, and Junjie Yan. Mimicking very efficient network for object detection. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 6356–6364, 2017.
  34. 34.Liting Lin, Heng Fan, Yong Xu, and Haibin Ling. Swintrack: A simple and strong baseline for transformer tracking. Neural Information Processing Systems, NIPS, 2022.
  35. 35.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In Proceedings of the European Conference on Computer Vision, ECCV, 2014.
  36. 36.Benlin Liu, Yongming Rao, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Metadistiller: Network self-boosting via meta-learned top-down distillation. In European Conference on Computer Vision, pages 694–709. Springer, 2020.
  37. 37.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2021.
  38. 38.Alan Lukezic, Jiri Matas, and Matej Kristan. D3S - A discriminative single shot segmentation tracker. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2020.
  39. 39.Christoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul, Danda Pani Paudel, Fisher Yu, and Luc Van Gool. Transforming model prediction for tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2022.
  40. 40.Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool. Learning target candidate association to keep track of what not to track. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2021.
  41. 41.Matthias Mueller, Neil Smith, and Bernard Ghanem. A benchmark and simulator for UAV tracking. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, Proceedings of the European Conference on Computer Vision, ECCV, 2016.
  42. 42.Matthias Müller, Adel Bibi, Silvio Giancola, Salman Al-Subaihi, and Bernard Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In Proceedings of the European Conference on Computer Vision, ECCV, 2018.
  43. 43.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  44. 44.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949, 2021.
  45. 45.Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. Compact transformer tracker with correlative masked modeling. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), February 2023.
  46. 46.Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, and Wei Yang. Transformer tracking with cyclic shifting window attention. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2022.
  47. 47.Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. Haq: Hardware-aware automated quantization with mixed precision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8612–8620, 2019.
  48. 48.Xiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, and Feng Wu. Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13763–13773, 2021.
  49. 49.Fei Xie, Chunyu Wang, Guangting Wang, Yue Cao, Wankou Yang, and Wenjun Zeng. Correlation-aware deep tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2022.
  50. 50.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines. In Proceedings of the AAAI Conference on Artificial Intelligence, AAAI, 2020.
  51. 51.Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, and Xing Sun. Evo-vit: Slow-fast token evolution for dynamic vision transformer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2964–2972, 2022.
  52. 52.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2021.
  53. 53.Bin Yan, Houwen Peng, Kan Wu, Dong Wang, Jianlong Fu, and Huchuan Lu. Lighttrack: Finding lightweight neural networks for object tracking via one-shot architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15180–15189, 2021.
  54. 54.Zhendong Yang, Zhe Li, Ailing Zeng, Zexian Li, Chun Yuan, and Yu Li. Vitkd: Practical guidelines for vit feature knowledge distillation. arXiv preprint arXiv:2209.02432, 2022.
  55. 55.Botao Ye, Hong Chang, Bingpeng Ma, and Shiguang Shan. Joint feature learning and relation modeling for tracking: A one-stream framework. Proceedings of the European Conference on Computer Vision, ECCV, 2022.
  56. 56.Jinnian Zhang, Houwen Peng, Kan Wu, Mengchen Liu, Bin Xiao, Jianlong Fu, and Lu Yuan. Minivit: Compressing vision transformers with weight multiplexing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12145–12154, 2022.
  57. 57.Zhaohui Zheng, Rongguang Ye, Qibin Hou, Dongwei Ren, Ping Wang, Wangmeng Zuo, and Ming-Ming Cheng. Localization distillation for object detection. arXiv preprint arXiv:2204.05957, 2022.

Citation

MLA
Cui, Y., et al. “MixFormerV2: Efficient Fully Transformer Tracking”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 58736–51, https://proceedings.neurips.cc/paper_files/paper/2023/file/b7870bd43b2d133a1ed95582ae5d82a4-Paper-Conference.pdf.
APA
Cui, Y., Song, T., Wu, G., & Wang, L. (2023). MixFormerV2: Efficient Fully Transformer Tracking. Advances in Neural Information Processing Systems, 36, 58736–58751. https://proceedings.neurips.cc/paper_files/paper/2023/file/b7870bd43b2d133a1ed95582ae5d82a4-Paper-Conference.pdf
Chicago
Cui, Y., T. Song, G. Wu, and L. Wang. 2023. “MixFormerV2: Efficient Fully Transformer Tracking”. Advances in Neural Information Processing Systems 36: 58736–51. https://proceedings.neurips.cc/paper_files/paper/2023/file/b7870bd43b2d133a1ed95582ae5d82a4-Paper-Conference.pdf.
Harvard
Cui, Y. et al. (2023) “MixFormerV2: Efficient Fully Transformer Tracking”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 58736–58751. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/b7870bd43b2d133a1ed95582ae5d82a4-Paper-Conference.pdf.
Vancouver
1. Cui Y, Song T, Wu G, Wang L (2023) MixFormerV2: Efficient Fully Transformer Tracking. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 58736–58751

BibTeX

@inproceedings{cui2023mixformerv2,
  title = {MixFormerV2: Efficient Fully Transformer Tracking},
  author = {Cui, Yutao and Song, Tianhui and Wu, Gangshan and Wang, Limin},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {58736-58751},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/b7870bd43b2d133a1ed95582ae5d82a4-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors