MixFormer: End-to-End Tracking with Iterative Mixed Attention

Yutao CuiCheng JiangLimin WangGangshan Wu

article2022CVPR759 citations

Proposes MixFormer, an end-to-end transformer tracker that unifies feature extraction and target information integration through iterative mixed attention to set new state-of-the-art performance across five major visual tracking benchmarks.

Listen

Visual object tracking is a foundational computer vision capability essential for autonomous systems, surveillance, and human-computer interaction. Conventional trackers rely on a complex, multi-stage pipeline: first using a generic convolutional network to extract features, then passing these features through a dedicated integration module to compare the target with the search area, and finally predicting bounding boxes with task-specific heads. This separated approach creates computational bottlenecks and limits accuracy, as generic feature extractors often miss fine-grained, target-specific details required to handle severe occlusions, object deformations, and visual distractors.

The article introduces and evaluates MixFormer, a streamlined tracking framework that unifies generic feature extraction and target information integration into a single end-to-end architecture. The primary objective is to demonstrate that coupling these processes via iterative mixed attention improves tracking accuracy and robustness while simplifying system design.

The authors designed a Mixed Attention Module that simultaneously executes self-attention within target and search regions and cross-attention between them. To maximize practical efficiency, they developed an asymmetric attention mechanism that eliminates unnecessary computations and paired it with a score prediction module to reliably update target templates over time. MixFormer was rigorously benchmarked against leading trackers across five standard evaluation datasets—including LaSOT, TrackingNet, VOT2020, GOT-10k, and UAV123—alongside thorough ablation studies evaluating module configurations, attention types, and regression heads.

The evaluation produced several decisive findings. First, MixFormer established a new state-of-the-art across all five benchmarks; for example, the large variant (MixFormer-L) achieved a top Expected Average Overlap score of 0.555 on VOT2020 (outperforming STARK by 5.0%) and normalized precision scores of 79.9% on LaSOT and 88.9% on TrackingNet. Second, ablations confirmed that simultaneous feature extraction and integration substantially outperforms decoupled architectures, yielding an 8.6% boost over standard separate attention pipelines while utilizing fewer parameters and floating-point operations. Third, the asymmetric attention design increased processing throughput by approximately 24% without sacrificing accuracy. Finally, the standard MixFormer model operated in real time at 25 frames per second on standard commercial hardware (a GTX 1080Ti GPU).

These results demonstrate that dedicated feature integration networks are unnecessary in modern vision systems. By unifying feature learning and target interaction within a single transformer backbone, practitioners can achieve higher tracking precision and improved robustness against target deformation, while reducing architectural complexity, memory overhead, and post-processing dependencies.

Organizations developing computer vision pipelines should consider replacing multi-stage tracking networks with unified transformer architectures like MixFormer, particularly for applications requiring real-time execution and resilience against visual distractors. Teams should evaluate model size trade-offs: the standard MixFormer offers optimal throughput (25 FPS) for real-time edge deployments, whereas MixFormer-L is suited for offline or compute-heavy environments where maximal precision is paramount.

The findings are supported with high confidence by extensive benchmark evaluations and clear ablation comparisons. However, limitations remain: the framework is currently designed for single-object short-term tracking and requires moderate computing resources for training. Future research and development should focus on extending this unified attention paradigm to multi-object tracking and testing performance under extreme edge-device constraints.

Cover for MixFormer: End-to-End Tracking with Iterative Mixed Attention

Abstract

Tracking often uses a multi-stage pipeline of feature extraction, target information integration, and bounding box estimation. To simplify this pipeline and unify the process of feature extraction and target information integration, we present a compact tracking framework, termed as MixFormer, built upon transformers. Our core design is to utilize the flexibility of attention operations, and propose a Mixed Attention Module (MAM) for simultaneous feature extraction and target information integration. This synchronous modeling scheme allows to extract target-specific discriminative features and perform extensive communication between target and search area. Based on MAM, we build our MixFormer tracking framework simply by stacking multiple MAMs with progressive patch embedding and placing a localization head on top. In addition, to handle multiple target templates during online tracking, we devise an asymmetric attention scheme in MAM to reduce computational cost, and propose an effective score prediction module to select high-quality templates. Our MixFormer sets a new state-of-the-art performance on five tracking benchmarks, including LaSOT, TrackingNet, VOT2020, GOT-10k, and UAV123. In particular, our MixFormer-L achieves NP score of 79.9% on LaSOT, 88.9% on TrackingNet and EAO of 0.555 on VOT2020. We also perform in-depth ablation studies to demonstrate the effectiveness of simultaneous feature extraction and information integration. Code and trained models are publicly available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Mixed Attention Module (MAM)
  • 3.2 MixFormer for Tracking
  • 3.3 Training and Inference
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Comparison with the state-of-the-art trackers
  • 4.3 Exploration Studies
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Mixed Attention Module (MAM)

    model/method

    The Mixed Attention Module (MAM) unifies generic feature extraction and cross-region information integration into a single attention operation. MAM takes token sequences from two image sources: target templates (TT) and the search area (SS).

    To preserve local spatial context, the input tokens of target and search are first reshaped to 2D feature maps, passed through a separable depth-wise convolutional projection layer (with downsampling applied to keys and values for computational efficiency), flattened, and projected linearly into queries (qt,qsq_t, q_s), keys (kt,ksk_t, k_s), and values (vt,vsv_t, v_s). Target and search keys/values are concatenated along the token sequence dimension:

    km=Concat(kt,ks),vm=Concat(vt,vs)k_m = \text{Concat}(k_t, k_s), \quad v_m = \text{Concat}(v_t, v_s)

    The mixed attention operations are defined as:

    Attentiont=Softmax(qtkmTd)vm\text{Attention}_t = \text{Softmax}\left(\frac{q_t k_m^T}{\sqrt{d}}\right) v_m

    Attentions=Softmax(qskmTd)vm\text{Attention}_s = \text{Softmax}\left(\frac{q_s k_m^T}{\sqrt{d}}\right) v_m

    where dd is the key feature dimension, qt∈RNt×dq_t \in \mathbb{R}^{N_t \times d}, qs∈RNs×dq_s \in \mathbb{R}^{N_s \times d}, and km,vm∈R(Nt+Ns)×dk_m, v_m \in \mathbb{R}^{(N_t + N_s) \times d}. Each attention operation computes both self-attention (within the same region) and cross-attention (between target and search regions) simultaneously in a single matrix multiplication.

  2. Knowl 2 — Asymmetric Mixed Attention Scheme

    equation

    In visual tracking, cross-attention from target queries to search keys can incorporate unnecessary background distractors into the template representation and incurs substantial computational overhead when handling multiple online templates. The asymmetric mixed attention scheme eliminates the target-to-search cross-attention while preserving search-to-target cross-attention:

    Attentiont=Softmax(qtktTd)vt\text{Attention}_t = \text{Softmax}\left(\frac{q_t k_t^T}{\sqrt{d}}\right) v_t

    Attentions=Softmax(qskmTd)vm\text{Attention}_s = \text{Softmax}\left(\frac{q_s k_m^T}{\sqrt{d}}\right) v_m

    where km=Concat(kt,ks)k_m = \text{Concat}(k_t, k_s), vm=Concat(vt,vs)v_m = \text{Concat}(v_t, v_s), dd is the feature dimension of the key tokens, qt,kt,vtq_t, k_t, v_t are the target query, key, and value matrices, and qs,ks,vsq_s, k_s, v_s are the search query, key, and value matrices.

    Under this formulation, the template tokens within each MAM depend solely on target tokens and remain invariant across changing search regions during tracking, allowing template feature caching and enabling high-efficiency online inference.

  3. Knowl 3 — MixFormer Tracking Architecture and Backbone

    model/method

    MixFormer is an end-to-end visual tracker that couples target-specific feature extraction and target-search feature integration directly inside a transformer backbone, eliminating the need for an explicit downstream integration module or multi-scale feature aggregation.

    The tracker receives TT template images (one static initial template and T−1T-1 online dynamic templates) of resolution Ht×Wt=128×128H_t \times W_t = 128 \times 128 and a search region of resolution Hs×Ws=320×320H_s \times W_s = 320 \times 320. The backbone consists of three progressive stages:

    1. Stage 1: An overlapping convolutional patch embedding layer (kernel size 7×77 \times 7, stride 4) maps target and search images to initial token representations of channel dimension CC. The flattened tokens are concatenated into a sequence of length (T⋅Ht4Wt4+Hs4Ws4)=(T⋅32⋅32+80⋅80)(T \cdot \frac{H_t}{4}\frac{W_t}{4} + \frac{H_s}{4}\frac{W_s}{4}) = (T \cdot 32 \cdot 32 + 80 \cdot 80) and processed through N1N_1 Target-Search MAM blocks.
    2. Stage 2: Tokens are reshaped and downsampled via a convolutional patch embedding layer (kernel size 3×33 \times 3, stride 2) to channel dimension 3C3C, followed by N2N_2 MAM blocks on sequence length (T⋅16⋅16+40⋅40)(T \cdot 16 \cdot 16 + 40 \cdot 40).
    3. Stage 3: Tokens are downsampled via a convolutional patch embedding layer (kernel size 3×33 \times 3, stride 2) to channel dimension 6C6C, followed by N3N_3 MAM blocks on sequence length (T⋅8⋅8+20⋅20)(T \cdot 8 \cdot 8 + 20 \cdot 20).

    Finally, the search tokens of dimension 20×20×6C20 \times 20 \times 6C are extracted and passed directly to the localization head.

  4. Knowl 4 — Structural Specifications of MixFormer Variants

    data/table

    MixFormer is instantiated in two model variants: MixFormer (initialized with CvT-21) and MixFormer-L (initialized with the first 16 layers of CvT24-W). Both operate on target templates of shape 128×128×3128 \times 128 \times 3 and search regions of shape 320×320×3320 \times 320 \times 3.

    Stage Output Size (S,TS, T) Layer Parameters MixFormer MixFormer-L
    Stage 1 S:80×80S: 80 \times 80, T:32×32T: 32 \times 32 Conv. Embed 7×7,647 \times 7, 64, stride 4 7×7,1927 \times 7, 192, stride 4
    MAM + MLP [H1=1,D1=64,R1=4]×1\left[H_1=1, D_1=64, R_1=4\right] \times 1 [H1=3,D1=192,R1=4]×2\left[H_1=3, D_1=192, R_1=4\right] \times 2
    Stage 2 S:40×40S: 40 \times 40, T:16×16T: 16 \times 16 Conv. Embed 3×3,1923 \times 3, 192, stride 2 3×3,7683 \times 3, 768, stride 2
    MAM + MLP [H2=3,D2=192,R2=4]×4\left[H_2=3, D_2=192, R_2=4\right] \times 4 [H2=12,D2=768,R2=4]×2\left[H_2=12, D_2=768, R_2=4\right] \times 2
    Stage 3 S:20×20S: 20 \times 20, T:8×8T: 8 \times 8 Conv. Embed 3×3,3843 \times 3, 384, stride 2 3×3,10243 \times 3, 1024, stride 2
    MAM + MLP [H3=6,D3=384,R3=4]×16\left[H_3=6, D_3=384, R_3=4\right] \times 16 [H3=16,D3=1024,R3=4]×12\left[H_3=16, D_3=1024, R_3=4\right] \times 12
    Parameters (MACs) 35.61 M 183.89 M
    FLOPs 23.04 G 127.81 G
    Speed (GTX 1080Ti) 25 FPS 18 FPS

    HiH_i, DiD_i, and RiR_i denote the number of attention heads, embedding channel dimension, and MLP feature dimension expansion ratio in the ii-th stage, respectively.

  5. Knowl 5 — Score Prediction Module and Online Template Update

    model/method

    To handle object deformation and avoid corrupting the tracker with poor-quality online templates, MixFormer uses a Score Prediction Module (SPM) to dynamically evaluate candidate online templates.

    SPM Architecture:

    1. A learnable score token s∈R1×Cs \in \mathbb{R}^{1 \times C} serves as query and performs cross-attention over search region-of-interest (RoI) tokens to encode candidate target feature information.
    2. The updated score token then serves as query and performs cross-attention over all token positions of the initial ground-truth target template, implicitly comparing candidate appearance against the initial reference.
    3. The resulting token passes through a three-layer MLP followed by a sigmoid activation to output a confidence score pi∈[0,1]p_i \in [0, 1].

    Online Update Mechanism: During online tracking, templates are updated periodically (default update interval: 200 frames). Among all frames within the interval whose predicted score pi≥0.5p_i \ge 0.5, the candidate template with the highest confidence score is selected to replace the dynamic online template.

  6. Knowl 6 — Corner-Based and Query-Based Localization Heads

    model/method

    MixFormer supports two target bounding box prediction heads applied directly on top of the backbone search tokens (20×20×6C20 \times 20 \times 6C):

    1. Corner-Based Localization Head (Default): A fully convolutional network composed of sequential Conv-BN-ReLU layers that predict probability distributions for the top-left and bottom-right bounding box corners directly. Bounding box coordinates are computed by taking the mathematical expectation over the predicted corner probability distributions, avoiding complex decoders or heuristic post-processing.
    2. Query-Based Localization Head: A pure transformer localization mechanism where a learnable regression token is appended to the token sequence in the final backbone stage. This regression token acts as an anchor aggregating global context across target and search area tokens. The resulting output token is passed to a three-layer Feed-Forward Network (FFN) to directly regress the normalized 4D bounding box coordinates.
  7. Knowl 7 — MixFormer Training Protocol and Loss Functions

    experimental setup

    MixFormer is trained in a two-stage process:

    Stage 1: Backbone and Localization Head Training The backbone weights are initialized from pretrained CvT models (CvT-21 for MixFormer, CvT24-W for MixFormer-L). The backbone and localization head are trained jointly for 500 epochs using the Adam optimizer with weight decay 10−410^{-4}, an initial learning rate of 10−410^{-4} reduced to 10−510^{-5} at epoch 400, and a gradient clipping norm of 0.1. Batch normalization layers in the backbone are frozen. The loss function is a combination of L1L_1 loss and Generalized IoU (GIoUGIoU) loss:

    Lloc=λL1L1(Bi,B^i)+λgiouLgiou(Bi,B^i)L_{loc} = \lambda_{L1} L_1(B_i, \hat{B}_i) + \lambda_{giou} L_{giou}(B_i, \hat{B}_i)

    where λL1=5\lambda_{L1} = 5, λgiou=2\lambda_{giou} = 2, BiB_i is the ground-truth bounding box, and B^i\hat{B}_i is the predicted bounding box.

    Stage 2: Score Prediction Module Training With the backbone and localization head frozen, the SPM is trained for 40 epochs (batch size 32, learning rate decayed at epoch 30) using binary cross-entropy loss:

    Lscore=−[yilog⁡(pi)+(1−yi)log⁡(1−pi)]L_{score} = -\left[y_i \log(p_i) + (1 - y_i) \log(1 - p_i)\right]

    where yi∈{0,1}y_i \in \{0, 1\} is the ground-truth target label and pip_i is the predicted confidence score.

  8. Knowl 8 — State-of-the-Art Benchmark Tracking Results

    data/table

    MixFormer was evaluated against prior state-of-the-art visual object trackers across five tracking benchmarks: LaSOT, TrackingNet, GOT-10k, UAV123, and VOT2020.

    Method LaSOT TrackingNet GOT-10k UAV123 VOT2020
    AUC PNormP_{Norm} P AUC PNormP_{Norm} P AO SR0.5SR_{0.5} SR0.75SR_{0.75} AUC P EAO
    MixFormer-L 70.1 79.9 76.3 83.9 88.9 83.1 75.6 85.7 72.8 69.5 91.0 0.555
    MixFormer-22k 69.2 78.7 74.7 83.1 88.1 81.6 72.6 82.2 68.8 70.4 91.8 0.535
    MixFormer-1k 67.9 77.3 73.9 82.6 87.7 81.2 73.2 83.2 70.2 68.7 89.5 0.527
    STARK 67.1 77.0 - 82.0 86.9 - 68.8 78.1 64.1 - - 0.505
    KeepTrack 67.1 77.2 70.2 - - - - - - 69.7 - -
    TransT 64.9 73.8 69.0 81.4 86.7 80.3 67.1 76.8 60.9 69.1 - -
    TrDiMP 63.9 - 61.4 78.4 83.3 73.1 67.1 77.7 58.3 67.5 - -
    STMTrack 60.6 69.3 63.3 80.3 85.1 76.7 64.2 73.7 57.5 64.7 - -
    DiMP 56.9 65.0 56.7 74.0 80.1 68.7 61.1 71.7 49.2 65.4 - -

    All metrics are in percentages (%) except VOT2020 EAO. MixFormer-L surpasses STARK by 5.0% in EAO on VOT2020, 2.9% in PNormP_{Norm} on LaSOT, and 2.0% in PNormP_{Norm} on TrackingNet.

  9. Knowl 9 — Ablation Study: Simultaneous Mixed Attention vs. Decoupled Processing

    empirical result

    To evaluate the advantage of unifying feature extraction and target integration within the backbone over separate multi-stage pipelines, MixFormer variants were tested on LaSOT:

    1. Decoupled Architecture with Self-Attention Backbone + Cross-Attention Integration: A 21-layer self-attention (SAM) backbone followed by 1 cross-attention module (CAM) achieved 59.8% AUC (37.35M parameters, 20.69 GFLOPs). Using 3 CAMs achieved 60.5% AUC (40.92M parameters, 22.20 GFLOPs). Using TransT's ECA+CFA(4) integration module achieved 66.9% AUC (49.75M parameters, 27.81 GFLOPs).
    2. Progressive Mixed Attention Integration: Replacing SAM layers with MAM layers progressively improved performance:
      • 20 SAM + 1 MAM: 65.8% AUC
      • 15 SAM + 6 MAM: 66.2% AUC
      • 10 SAM + 11 MAM: 67.4% AUC
      • 5 SAM + 16 MAM: 68.1% AUC
      • 21 MAM (MixFormer-Base): 68.4% AUC (35.97M parameters, 20.85 GFLOPs).

    Simultaneously performing feature extraction and mutual cross-attention throughout all 21 backbone layers outperforms decoupled processing by up to 8.6% AUC while using fewer parameters and FLOPs.

  10. Knowl 10 — Ablation Study: Asymmetric Attention, Score Prediction, and Heads

    empirical result

    Ablation experiments on LaSOT evaluated the contributions of specific MixFormer components:

    • Asymmetric Mixed Attention: Applying asymmetric attention (pruning target-to-search cross-attention) increased inference speed from 19 FPS to 25 FPS on a GTX 1080Ti (a 31.6% frame-rate improvement) with minimal AUC change (68.4% vs. 68.1%).
    • Score Prediction Module (SPM):
      • Static template only (no online updates): 68.1% AUC
      • Online template updates at fixed intervals without SPM: 66.6% AUC (performance drops due to low-quality/distractor templates)
      • Online template updates guided by SPM confidence scores: 69.2% AUC (+1.1% over static baseline, +2.6% over naive online updating).
    • Localization Head Generalization: The fully convolutional corner head achieved 68.4% AUC, while the query-based sparse regression head achieved 66.0% AUC (outperforming query-based STARK-ST's 63.7%), demonstrating that the MAM backbone generalizes across different head paradigms.

Coverage note — None was omitted. All key contributed components (MAM formulation, asymmetric MAM, backbone architecture and variants, score prediction module, online template update strategy, localization heads, training losses, empirical benchmark results, and ablation studies) have been thoroughly covered.

References

  1. 1.Luca Bertinetto, Jack Valmadre, Stuart Golodetz, Ondrej Miksik, and Philip H. S. Torr. Staple: Complementary learners for real-time tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 1
  2. 2.Luca Bertinetto, Jack Valmadre, Joao F. Henriques, Andrea Ȱ Vedaldi, and Philip H. S. Torr. Fully-convolutional siamese networks for object tracking. In ECCV Workshops, 2016. 1, 2, 7
  3. 3.Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. Learning discriminative model prediction for tracking. In ICCV, 2019. 1, 2, 6, 7
  4. 4.Goutam Bhat, Joakim Johnander, Martin Danelljan, Fahad Shahbaz Khan, and Michael Felsberg. Unveiling the power of deep tracking. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss, editors, ECCV, 2018. 1
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. In European Conference on Computer Vision, 2020. 3, 5
  6. 6.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 4, 7, 8
  7. 7.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2021. 2, 3, 5, 7
  8. 8.Zedu Chen, Bineng Zhong, Guorong Li, Shengping Zhang, and Rongrong Ji. Siamese box adaptive network for visual tracking. In CVPR, June 2020. 1
  9. 9.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Fully convolutional online tracking. CoRR, 2020. 1, 2, 7
  10. 10.Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu. Target transformed regression for accurate tracking. CoRR, 2021. 1, 2, 3, 7
  11. 11.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ECO: efficient convolution operators for tracking. In CVPR, 2017. 2
  12. 12.Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg. ATOM: accurate tracking by overlap maximization. In CVPR, 2019. 1, 2, 7
  13. 13.Martin Danelljan, Luc Van Gool, and Radu Timofte. Probabilistic regression for visual tracking. In CVPR, 2020. 7
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 2009. 6
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021. 2
  16. 16.Fei Du, Peng Liu, Wei Zhao, and Xianglong Tang. Correlation-guided attention for corner detection based visual tracking. In CVPR, 2020. 3, 7
  17. 17.Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling. Lasot: A high-quality benchmark for large-scale single object tracking. In CVPR, 2019. 2, 6, 7, 9
  18. 18.Heng Fan and Haibin Ling. Siamese cascaded region proposal networks for real-time visual tracking. In CVPR, 2019. 1
  19. 19.Zhihong Fu, Qingjie Liu, Zehua Fu, and Yunhong Wang. Stmtrack: Template-free visual tracking with space-time memory networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 2, 7
  20. 20.Dongyan Guo, Yanyan Shao, Ying Cui, Zhenhua Wang, Liyan Zhang, and Chunhua Shen. Graph attention tracking. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2021. 7
  21. 21.Dongyan Guo, Jun Wang, Ying Cui, Zhenhua Wang, and Shengyong Chen. Siamcar: Siamese fully convolutional classification and regression for visual tracking. In CVPR, June 2020. 2
  22. 22.Joao F. Henriques, Rui Caseiro, Pedro Martins, and Jorge Ȱ Batista. High-speed tracking with kernelized correlation filters. IEEE Trans. Pattern Anal. Mach. Intell., 37(3):583–596, 2015. 1, 2, 6
  23. 23.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE Trans. Pattern Anal. Mach. Intell., 43(5):1562–1577, 2021. 2, 6, 7
  24. 24.Hamed Kiani Galoogahi, Ashton Fagg, and Simon Lucey. Learning background-aware correlation filters for visual tracking. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017. 1, 2
  25. 25.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, ICLR, 2015. 6
  26. 26.Matej Kristan, Ales Leonardis, and et. al. The eighth visual object tracking VOT2020 challenge results. In Adrien Bartoli and Andrea Fusiello, editors, Computer Vision - ECCV 2020 Workshops, 2020. 2, 6
  27. 27.Matej Kristan, Jiř'ı Matas, Ales Leonardis, Michael Felsberg, Ć Roman Pflugfelder, Joni-Kristian Kam¨ ar¨ ainen, Hyung Jin ¨ Chang, Martin Danelljan, Luka Cehovin, Alan LukeziȰ c, On- Ć drej Drbohlav, Jani Kapyl ¨ a, Gustav H ¨ ager, Song Yan, Jinyu ¨ Yang, Zhongqun Zhang, and Gustavo Fernandez. The ninth ´ visual object tracking vot2021 challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, October 2021. 3
  28. 28.Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan. Siamrpn++: Evolution of siamese visual tracking with very deep networks. In CVPR, 2019. 2, 5, 7
  29. 29.Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu. High performance visual tracking with siamese region proposal network. In CVPR, 2018. 1, 2
  30. 30.Feng Li, Cheng Tian, Wangmeng Zuo, Lei Zhang, and MingHsuan Yang. Learning spatial-temporal regularized correlation filters for visual tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 2
  31. 31.Xiang Li, Wenhai Wang, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss V2: learning reliable localization quality estimation for dense object detection. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR. Computer Vision Foundation / IEEE, 2021. 5
  32. 32.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. CoRR, abs/2104.05707, 2021. 2
  33. 33.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and ´ C. Lawrence Zitnick. Microsoft COCO: common objects in context. In ECCV, 2014. 6
  34. 34.Liwei Liu, Junliang Xing, Haizhou Ai, and Xiang Ruan. Hand posture recognition using finger geometric feature. In ICPR, 2012. 1
  35. 35.Alan Lukezic, Jiri Matas, and Matej Kristan. D3S - A discriminative single shot segmentation tracker. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, pages 7131–7140. Computer Vision Foundation / IEEE, 2020. 6, 7
  36. 36.Alan Lukezic, Tomas Vojir, Luka Cehovin Zajc, Jiri Matas, and Matej Kristan. Discriminative correlation filter with channel and spatial reliability. In CVPR, 2017. 1, 2
  37. 37.Alan Lukezic, Tomas Vojir, Luka Cehovin Zajc, Jiri Matas, and Matej Kristan. Discriminative correlation filter with channel and spatial reliability. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017. 1
  38. 38.Ziang Ma, Linyuan Wang, Haitao Zhang, Wei Lu, and Jun Yin. RPT: learning point set representation for siamese visual tracking. In Adrien Bartoli and Andrea Fusiello, editors, Computer Vision - ECCV 2020 Workshops, Lecture Notes in Computer Science, 2020. 6
  39. 39.Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool. Learning target candidate association to keep track of what not to track. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 7, 8
  40. 40.Matthias Mueller, Neil Smith, and Bernard Ghanem. A benchmark and simulator for UAV tracking. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, ECCV, 2016. 2, 6, 7
  41. 41.Matthias Muller, Adel Bibi, Silvio Giancola, Salman Al- ¨ Subaihi, and Bernard Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In ECCV, 2018. 2, 6, 7
  42. 42.Hyeonseob Nam and Bohyung Han. Learning multi-domain convolutional neural networks for visual tracking. In CVPR, 2016. 1, 7
  43. 43.Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV. IEEE, 2019. 6
  44. 44.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian D. Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR. Computer Vision Foundation / IEEE, 2019. 5
  45. 45.Yibing Song, Chao Ma, Xiaohe Wu, Lijun Gong, Linchao Bao, Wangmeng Zuo, Chunhua Shen, Rynson W. H. Lau, and Ming-Hsuan Yang. VITAL: visual tracking via adversarial learning. In CVPR, 2018. 1
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, NIPS, 2017. 1, 3
  47. 47.Paul Voigtlaender, Jonathon Luiten, Philip H.S. Torr, and Bastian Leibe. Siam r-cnn: Visual tracking by re-detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 7
  48. 48.Guangting Wang, Chong Luo, Xiaoyan Sun, Zhiwei Xiong, and Wenjun Zeng. Tracking by instance detection: A metalearning approach. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2020. 7
  49. 49.Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li. Transformer meets tracker: Exploiting temporal context for robust visual tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 2, 7
  50. 50.Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H. S. Torr. Fast online object tracking and segmentation: A unifying approach. In CVPR, 2019. 6
  51. 51.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 2
  52. 52.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. CoRR, abs/2103.15808, 2021. 2, 5, 6
  53. 53.Yi Wu, Jongwoo Lim, and Ming-Hsuan Yang. Object tracking benchmark. IEEE Trans. Pattern Anal. Mach. Intell., 37(9):1834–1848, 2015. 9
  54. 54.Fei Xie, Chunyu Wang, Guangting Wang, Wankou Yang, and Wenjun Zeng. Learning tracking representations via dual-branch fully transformer networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, October 2021. 6, 7
  55. 55.Junliang Xing, Haizhou Ai, and Shihong Lao. Multiple human tracking based on multi-view upper-body detection and discriminative learning. In ICPR, 2010. 1
  56. 56.Yinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan, and Gang Yu. Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020. 1, 7
  57. 57.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 1, 2, 3, 4, 5, 6, 7, 8
  58. 58.Bin Yan, Xinyu Zhang, Dong Wang, Huchuan Lu, and Xiaoyun Yang. Alpha-refine: Boosting tracking performance by precise bounding box estimation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR. Computer Vision Foundation / IEEE, 2021. 6
  59. 59.Bin Yu, Ming Tang, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Hao Feng, Xuetao Feng, and Hanqing Lu. High-performance discriminative tracking with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 1, 7
  60. 60.Yuechen Yu, Yilei Xiong, Weilin Huang, and Matthew R. Scott. Deformable siamese attention networks for visual object tracking. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 6727–6736. Computer Vision Foundation / IEEE, 2020. 3
  61. 61.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 2
  62. 62.Zhipeng Zhang, Yufan Liu, Bing Li, Weiming Hu, and Houwen Peng. Toward accurate pixelwise object tracking via attention retrieval. IEEE Transactions on Image Processing, 30:8553–8566, 2021. 6
  63. 63.Zhipeng Zhang, Yihao Liu, Xiao Wang, Bing Li, and Weiming Hu. Learn to match: Automatic matching network design for visual tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 7
  64. 64.Zhipeng Zhang, Houwen Peng, Jianlong Fu, Bing Li, and Weiming Hu. Ocean: Object-aware anchor-free tracking. In ECCV, 2020. 1, 7
  65. 65.Zikun Zhou, Wenjie Pei, Xin Li, Hongpeng Wang, Feng Zheng, and Zhenyu He. Saliency-associated object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 7
  66. 66.Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu. Distractor-aware siamese networks for visual object tracking. In ECCV, 2018. 2

Citation

MLA
Cui, Y., et al. “MixFormer: End-to-End Tracking with Iterative Mixed Attention”. arXiv, 2022, http://arxiv.org/abs/2203.11082v2.
APA
Cui, Y., Jiang, C., Wang, L., & Wu, G. (2022). MixFormer: End-to-End Tracking with Iterative Mixed Attention. arXiv. http://arxiv.org/abs/2203.11082v2
Chicago
Cui, Y., C. Jiang, L. Wang, and G. Wu. 2022. “MixFormer: End-to-End Tracking with Iterative Mixed Attention”. arXiv. http://arxiv.org/abs/2203.11082v2.
Harvard
Cui, Y. et al. (2022) “MixFormer: End-to-End Tracking with Iterative Mixed Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.11082v2.
Vancouver
1. Cui Y, Jiang C, Wang L, Wu G (2022) MixFormer: End-to-End Tracking with Iterative Mixed Attention. arXiv

BibTeX

@article{cui2022mixformer,
  title = {MixFormer: End-to-End Tracking with Iterative Mixed Attention},
  author = {Cui, Yutao and Jiang, Cheng and Wang, Limin and Wu, Gangshan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.11082v2},
  eprint = {2203.11082}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE