Global Tracking Transformers

Xingyi ZhouTianwei YinVladlen KoltunPhilipp Krähenbühl

article2022CVPR220 citations

Introduces a transformer-based multi-object tracking framework that directly groups detections across video sequences into global trajectories via trajectory queries, eliminating heuristic pairwise association while enabling end-to-end joint training with modern object detectors.

Listen

Multi-object tracking is a critical capability for autonomous systems, robotics, and intelligent video analytics, where machines must accurately identify and follow multiple moving entities over time. Conventional tracking-by-detection systems primarily rely on greedy, frame-by-frame pairwise associations or separate, computationally heavy graph-based optimization algorithms. These existing paradigms often fail during complex scenarios involving long occlusions, drastic appearance changes, or large object vocabularies, while also requiring cumbersome multi-stage engineering pipelines.

The article demonstrates that global multi-object tracking can be unified and formulated directly within an end-to-end differentiable deep neural network using a specialized transformer architecture. The authors evaluate this framework, termed the Global Tracking Transformer, to establish whether global object associations across multi-frame temporal windows can outperform traditional local matching methods while maintaining high computational efficiency.

The authors designed a lightweight architecture that takes detected object features across a short sequence of video frames and uses detection features from a single frame as trajectory queries. These queries interact with all detected objects across the temporal window via attention layers to produce complete, globally consistent trajectories in a single forward pass. The framework was evaluated across standard computer vision benchmarks, including MOT17 for dense pedestrian tracking and the large-scale TAO benchmark spanning 488 object classes under challenging real-world conditions.

The analysis reveals several key findings. First, the proposed global tracking method significantly outperformed existing published approaches on the large-vocabulary TAO benchmark, achieving a 20.1 tracking mean average precision compared to the prior state of the art at 12.4—a relative improvement of approximately 62%. Second, on the MOT17 pedestrian benchmark, the system achieved competitive top-tier performance with 75.3 MOTA and 59.1 HOTA, surpassing most contemporary transformer-based tracking models. Third, the transformer association head proved remarkably efficient, requiring only a single encoder layer and single decoder layer without positional embeddings, adding merely 3 to 4 milliseconds of computation per frame on top of the base detector.

These results demonstrate that long-range temporal reasoning does not require complex offline combinatorial solvers or heavy multi-layer architectures. By operating directly on extracted object features rather than raw pixels, the framework integrates seamlessly with standard object detectors, reducing system complexity and enabling real-time deployment. For organizations developing video understanding or robotic perception systems, this approach reduces engineering overhead and improves tracking reliability in dynamic environments.

Organizations evaluating or deploying multi-object tracking should consider adopting query-based transformer heads to replace heuristic pairwise association modules, especially when tracking across diverse object categories. For immediate implementation, using a 16-to-32 frame sliding window offers an optimal balance between trajectory consistency and memory usage. When deploying in high-frame-rate settings, combining learned association likelihoods with spatial bounding-box overlap provides an additional accuracy boost.

The study notes clear operational boundaries: the model operates within a maximum temporal window of 32 frames due to GPU memory constraints, meaning it cannot natively recover from continuous occlusions lasting longer than 32 frames without sliding-window linking heuristics. Furthermore, due to the scarcity of richly annotated multi-category tracking datasets, the model relied partly on synthetic video generated from static images. Confidence in the reported performance is high across standard benchmarks, though teams deploying in production should validate long-term identity retention under extended occlusions.

  • Paper: TrackFormer: Multi-Object Tracking with Transformers, Tim Meinhardt et al. (2022). TrackFormer pioneered end-to-end multi-object tracking via transformer object queries and autoregressive tracking attention, establishing the foundational transformer tracking-by-attention paradigm that Global Tracking Transformers streamlines into global temporal windows.
  • Paper: ByteTrack: Multi-Object Tracking by Associating Every Detection Box, Yifu Zhang et al. (2021). ByteTrack provides the essential modern benchmark baseline and heuristic association paradigm in multi-object tracking that Global Tracking Transformers aims to replace with learned, end-to-end query-based attention.
  • Paper: HOTA: A Higher Order Metric for Evaluating Multi-object Tracking, Jonathon Luiten et al. (2020). HOTA defines the primary unified evaluation metric balancing detection, association, and localization accuracy used to measure and validate the performance of Global Tracking Transformers.
  • Paper: Tracking Objects as Points, Xingyi Zhou et al. (2020). CenterTrack formulated simultaneous object detection and tracking as point displacement estimation, introducing direct object-feature tracking concepts that inform lightweight multi-object tracking architectures.
  • Paper: FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking, Yifu Zhang et al. (2020). FairMOT established the importance of joint detection and re-identification feature balancing in single-network multi-object tracking architectures.
  • Paper: Simple online and realtime tracking with a deep association metric, Nicolai Wojke et al. (2017). Deep SORT represents the classical multi-stage tracking-by-detection paradigm combining Kalman filtering and learned deep appearance descriptors that end-to-end transformer trackers seek to modernize.
  • Paper: Simple online and realtime tracking, Alex Bewley et al. (2016). SORT established the foundational real-time tracking-by-detection framework using bipartite matching and bounding box overlap heuristics against which modern transformer trackers are benchmarked.
  • Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). MOT16 defined the standardized multi-object tracking benchmark and rigorous annotation protocols foundational to the MOT17 benchmark evaluated in the paper.
Cover for Global Tracking Transformers

Abstract

We present a novel transformer-based architecture for global multi-object tracking. Our network takes a short sequence of frames as input and produces global trajectories for all objects. The core component is a global tracking transformer that operates on objects from all frames in the sequence. The transformer encodes object features from all frames, and uses trajectory queries to group them into trajectories. The trajectory queries are object features from a single frame and naturally produce unique trajectories. Our global tracking transformer does not require intermediate pairwise grouping or combinatorial association, and can be jointly trained with an object detector. It achieves competitive performance on the popular MOT17 benchmark, with 75.3 MOTA and 59.1 HOTA. More importantly, our framework seamlessly integrates into state-of-the-art large-vocabulary detectors to track any objects. Experiments on the challenging TAO dataset show that our framework consistently improves upon baselines that are based on pairwise association, outperforming published work by a significant 7 tracking mAP. Code is available at https://github.com/xingyizhou/GTR.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Preliminaries
  • 4. Global tracking transformers
  • 4.1. Tracking transformers
  • 4.2. Training
  • 4.3. Online Inference
  • 4.4. Network architecture
  • 4.5. Connection to embedding learning and ReID
  • 5. Experiments
  • 5.1. Evaluation metrics
  • 5.2. Training and inference details
  • 5.3. Global versus local association
  • 5.4. Comparison to the state-of-the-art
  • 5.5. Design choice experiments
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Global Tracking Transformer for direct trajectory prediction

    model/method

    The Global Tracking Transformer (GTR) converts multi-object tracking into global grouping within a temporal window. For TT consecutive frames, an object detector independently produces detections, bounding boxes, and feature vectors. GTR concatenates the features of all detections from all frames, encodes them jointly, and uses trajectory queries to assign detections to trajectories. Each query produces one complete trajectory by selecting at most one detection, or no detection, in every frame.

    The trajectory queries are detection features rather than fixed learned parameters. Consequently, the model can adapt its queries to the objects present in the input video and can be trained jointly with the detector. GTR directly predicts query-to-detection assignments in one forward pass, avoiding intermediate pairwise matching and offline graph-based combinatorial optimization.

  2. Knowl 2 — Differentiable probabilistic association over complete trajectories

    equation

    Let frame tt contain NtN_t detected objects, with object features collected in F∈RN×DF\in\mathbb{R}^{N\times D}, where N=∑t=1TNtN=\sum_{t=1}^{T}N_t is the total number of detections and DD is the feature dimension. Let qk∈RDq_k\in\mathbb{R}^{D} be trajectory query kk. GTR produces a real association logit git(qk,F)g_i^t(q_k,F) for detection ii in frame tt and uses a special no-association state ∅\emptyset with logit g∅t(qk,F)=0g_\emptyset^t(q_k,F)=0. The per-frame association distribution is

    PA(αkt=i∣qk,F)=exp⁡ ⁣(git(qk,F))∑j∈{∅,1,…,Nt}exp⁡ ⁣(gjt(qk,F)),P_A(\alpha_k^t=i\mid q_k,F)= \frac{\exp\!\left(g_i^t(q_k,F)\right)} {\displaystyle\sum_{j\in\{\emptyset,1,\ldots,N_t\}}\exp\!\left(g_j^t(q_k,F)\right)},

    where αkt∈{∅,1,…,Nt}\alpha_k^t\in\{\emptyset,1,\ldots,N_t\} identifies the detection assigned to query kk at time tt. The corresponding distribution over a trajectory state is obtained by mapping each detection index to its bounding box, and the trajectory distribution factorizes across frames as

    PT(τ∣qk,F)=∏t=1TPt(τt∣qk,F).P_T(\tau\mid q_k,F)=\prod_{t=1}^{T}P_t(\tau^t\mid q_k,F).

    Thus, association is differentiable and can be optimized by likelihood maximization while reasoning over all detections in the temporal window.

  3. Knowl 3 — Detection features as adaptive trajectory queries

    model/method

    For a ground-truth trajectory τ^k\hat\tau_k, GTR assigns detections using an intersection-over-union threshold of 0.50.5. Let α^kt\hat\alpha_k^t denote the assigned detection index in frame tt, and let bitb_i^t and τ^kt\hat\tau_k^t be the detected and ground-truth bounding boxes. The assignment is

    α^kt={∅,τ^kt=∅ or max⁡iIoU⁡(bit,τ^kt)<0.5,arg⁡max⁡iIoU⁡(bit,τ^kt),otherwise.\hat\alpha_k^t= \begin{cases} \emptyset, & \hat\tau_k^t=\emptyset\ \text{or}\ \max_i\operatorname{IoU}(b_i^t,\hat\tau_k^t)<0.5,\\ \displaystyle\arg\max_i\operatorname{IoU}(b_i^t,\hat\tau_k^t), & \text{otherwise}. \end{cases}

    Any feature corresponding to a matched detection of trajectory kk can serve as its query during training. In practice, GTR uses every post-NMS detection feature from every frame as a query. Queries matched to a ground-truth trajectory are trained to recover that trajectory, while unmatched features are background queries trained to select ∅\emptyset in every frame. Multiple training queries may represent the same trajectory, so training does not require a one-to-one query-to-trajectory matching. At inference, only detection features from one frame are used as queries; because post-NMS detections within a frame are distinct, this produces one trajectory candidate per current-frame detection without duplicate query templates.

  4. Knowl 4 — Joint association and detector training objective

    equation

    Let Fjs∈RDF_j^s\in\mathbb{R}^{D} be the feature of detection jj in frame ss, let FF be all features in the temporal window, and let α^kt\hat\alpha_k^t be the IoU-based assignment for ground-truth trajectory kk. Every nonempty matched assignment α^ks\hat\alpha_k^s supplies a query q=Fα^kssq=F_{\hat\alpha_k^s}^{s}. The association loss for trajectory kk is

    ℓk=−∑s:α^ks≠∅∑t=1Tlog⁡PA ⁣(αkt=α^kt∣q=Fα^kss,F).\ell_k=-\sum_{s:\hat\alpha_k^s\neq\emptyset}\sum_{t=1}^{T} \log P_A\!\left(\alpha_k^t=\hat\alpha_k^t\mid q=F_{\hat\alpha_k^s}^{s},F\right).

    Let Bs\mathcal{B}_s be the set of detections in frame ss that are not assigned to any ground-truth trajectory. Background queries are trained to produce empty trajectories using

    ℓbg=−∑s=1T∑j∈Bs∑t=1Tlog⁡PA ⁣(αt=∅∣q=Fjs,F).\ell_{\mathrm{bg}}=-\sum_{s=1}^{T}\sum_{j\in\mathcal{B}_s}\sum_{t=1}^{T} \log P_A\!\left(\alpha^t=\emptyset\mid q=F_j^s,F\right).

    For the set of KK ground-truth trajectories, the total association loss is

    Lasso=ℓbg+∑k=1Kℓk.L_{\mathrm{asso}}=\ell_{\mathrm{bg}}+\sum_{k=1}^{K}\ell_k.

    GTR minimizes this loss jointly with the detector's classification and bounding-box regression losses, with optional second-stage classification and regression losses for multiclass tracking. The same IoU assignment is therefore used to supervise both detection outputs and global trajectory assignments.

  5. Knowl 5 — Sliding-window online inference and track linking

    algorithm

    GTR performs online inference with a temporal window of T=32T=32 frames and stride 11. It retains the detector boxes and features from the most recent 32 frames, uses the detections in the current frame as trajectory queries, and generates one global trajectory candidate per current detection. Candidates are linked to persistent tracks across windows using a one-to-one Hungarian matching step.

    Input: Video frames, detector, trained GTR, window size T=32, new-track threshold θ=0.2
    Output: Persistent object tracks
    For each incoming frame t:
        Run the detector and retain the current detections and their features.
        Add the current detections and features to a buffer containing at most T frames.
        Use every current-frame detection feature as a trajectory query.
        Run GTR on all buffered detections and the current-frame queries.
        For each current query, obtain its candidate trajectory and per-frame association likelihoods.
        If t is the first frame:
            Initialize one persistent track from each current detection.
        Otherwise:
            Compute the average association likelihood between every current candidate and every existing track.
            Use the Hungarian algorithm to obtain a unique candidate-to-track assignment.
            For each current candidate:
                If its best matched score is below θ:
                    Start a new persistent track from its current detection.
                Otherwise:
                    Append the current detection underlying the candidate query to the matched persistent track.
    After processing the video:
        Remove trajectories shorter than five detections.
        Return the remaining persistent tracks.

    For MOT17, the implementation combines the trajectory association score with the box-to-track IoU by taking their maximum; this location cue is not used for TAO. The MOT output score threshold is 0.550.55, the TAO proposal threshold is 0.40.4, and the new-track threshold is θ=0.2\theta=0.2.

  6. Knowl 6 — Lightweight one-layer transformer architecture

    model/method

    GTR takes an encoder input F∈RN×DF\in\mathbb{R}^{N\times D} containing all detection features and a decoder query matrix Q∈RM×DQ\in\mathbb{R}^{M\times D} containing MM trajectory queries. A single self-attention encoder layer contextualizes the detections. A single cross-attention decoder layer uses the trajectory queries as queries and the encoded detection features as keys and values. A final linear projection produces an association matrix G∈RM×NG\in\mathbb{R}^{M\times N}.

    The default network has one encoder layer and one decoder layer, uses no self-attention among decoder queries, no Layer Normalization, and no spatial or temporal positional embeddings. It contains ten linear layers in total and remains lightweight even with hundreds of queries, running in only a fraction of the backbone detector's runtime. The architecture operates on detected object features rather than pixels, allowing it to be attached to existing object detectors.

  7. Knowl 7 — Training and evaluation protocol

    experimental setup

    The experiments evaluate GTR on TAO, which tracks 488 object classes in long-tail scenes, and MOT17, which tracks pedestrians in crowded videos. TAO contains approximately 0.5k, 1k, and 1.5k training, validation, and test videos, with roughly 40 annotated frames per video at one frame per second. Because the TAO training annotations are incomplete and training directly on them degraded detection, the TAO model is trained on LVIS and COCO for detection and on synthetic videos generated from static images: two independently augmented versions of an image are treated as endpoints, and intermediate images and annotations are linearly interpolated. The tracking head is fine-tuned end-to-end using clips of 88 frames.

    The TAO detector uses a Res2Net backbone with deformable convolution and CenterNet2 proposal and cascaded RoI heads. The MOT17 detector uses CenterNet with a DLA-34 backbone, BiFPN upsampling, and RoIAlign features. It is pretrained on CrowdHuman and then fine-tuned using CrowdHuman and MOT data in a 1:11:1 ratio, also with 88-frame training clips. Testing uses the official TAO tracking [email protected] metric and MOT17 MOTA; HOTA, DetA, and AssA are additionally reported. On the authors' hardware, the tracking transformer takes 44 ms per MOT17 frame and 33 ms per TAO frame, compared with detector times of 4747 ms and 8686 ms, respectively.

  8. Knowl 8 — Global association improves over local tracking

    data/table

    This validation experiment compares location-only, ReID-based, and joint local trackers with GTR using the same detector outputs. TAO reports tracking mAP, HOTA, DetA, and AssA; MOT17 reports MOTA, IDF1, HOTA, DetA, and AssA. Higher values are better. Increasing GTR's temporal window generally improves association accuracy, demonstrating that global access to a longer sequence helps recover associations that local pairwise matching misses.

    Could not parse LaTeX table

    With T=32T=32, GTR exceeds the FairMOT-style IoU+ReID baseline on MOT17 by 1.81.8 AssA and 1.71.7 IDF1. On TAO, performance saturates around T=16T=16, which the authors associate with the large appearance and layout changes between its sparsely sampled frames.

  9. Knowl 9 — Performance on TAO and MOT17 benchmarks

    empirical result

    On TAO, GTR obtains 22.522.5 validation tracking [email protected], with HOTA 45.845.8, DetA 36.836.8, and AssA 57.557.5. On the test set it obtains 20.120.1 tracking [email protected] at 11.211.2 FPS. The corresponding SORT and QDTrack test results are 10.210.2 and 12.412.4 mAP, so GTR improves over the published QDTrack result by 7.77.7 mAP. When GTR uses QDTrack's detector rather than its own detector, it still reaches 20.420.4 validation mAP, HOTA 40.740.7, DetA 30.130.1, and AssA 55.655.6, improving over QDTrack with the same detector by 4.34.3 mAP and 1.91.9 AssA.

    On the MOT17 private-detection test set, GTR achieves MOTA 75.375.3, IDF1 71.571.5, HOTA 59.159.1, DetA 61.661.6, and AssA 57.057.0, with 26,79326{,}793 false positives, 109,854109{,}854 false negatives, 2,8592{,}859 identity switches, and 19.619.6 FPS. These results are competitive with or better than most transformer-based and association-based trackers evaluated in the paper. GTR is below the reported AOA TAO challenge winner, which reaches 27.527.5 test mAP, but AOA's complete pipeline takes 989989 ms per image on the authors' machine, whereas GTR's TAO detector-plus-transformer processing takes approximately 8989 ms per image.

  10. Knowl 10 — Ablations validate attention and global association design choices

    empirical result

    The design ablations use T=32T=32 on the MOT17 validation set unless otherwise specified. Replacing the transformer association head with a direct dot product reduces HOTA from 63.063.0 to 61.361.3 and AssA from 66.266.2 to 63.6;italsoreducesDetAfrom63.6; it also reduces DetA from 60.4toto59.5andMOTAfromand MOTA from71.3toto70.5.Addingdecoderself−attentionontopofencoderattentiondoesnothelp:encoder−onlyattentiongivesHOTA. Adding decoder self-attention on top of encoder attention does not help: encoder-only attention gives HOTA 63.0,DetA, DetA 60.4,AssA, AssA 66.2,andMOTA, and MOTA 71.3,whileencoder−plus−decoderattentiongivesHOTA, while encoder-plus-decoder attention gives HOTA 62.3,DetA, DetA 60.5,AssA, AssA 64.5,andMOTA, and MOTA 71.2$.

    Learned spatial positional embeddings and additional temporal embeddings are also unnecessary. The no-embedding default obtains HOTA 63.063.0, DetA 60.460.4, AssA 66.266.2, and MOTA 71.371.3; adding spatial embeddings gives HOTA 62.562.5, DetA 60.760.7, AssA 65.065.0, and MOTA 71.771.7; adding both spatial and temporal embeddings gives HOTA 62.462.4, DetA 60.760.7, AssA 64.664.6, and MOTA 71.771.7. One encoder and one decoder layer are sufficient: the default 11-encoder/11-decoder configuration gives HOTA 63.063.0 and AssA 66.266.2, while 11-encoder/22-decoder gives HOTA 62.762.7 and AssA 65.065.0, and 22-encoder/11-decoder gives HOTA 63.063.0 and AssA 66.066.0.

    Using location during inference improves MOT17 AssA from 63.363.3 to 66.266.2 without changing MOTA from 71.371.3; on TAO, location does not change tracking mAP, which remains 22.522.5. This supports using temporal association alone for the low-frame-rate TAO videos while adding location cues for high-frame-rate MOT17.

Coverage note — The paper's explicit limitations—GPU-limited 32-frame memory, inability to recover from misses or occlusions lasting more than 32 frames, and TAO training only on static imagery—were omitted to prioritize the load-bearing method, objective, inference, and validation knowls; detailed competitor rows and backbone-specific optimization schedules were likewise condensed.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv:1607.06450, 2016. 5
  2. 2.Sara Beery, Guanhang Wu, Vivek Rathod, Ronny Votel, and Jonathan Huang. Context r-cnn: Long term temporal context for per-camera object detection. In CVPR, 2020. 3
  3. 3.Jerome Berclaz, Francois Fleuret, Engin Turetken, and Pascal Fua. Multiple object tracking using k-shortest paths optimization. TPAMI, 2011. 1
  4. 4.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In ICCV, 2019. 1, 2, 3
  5. 5.Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple online and realtime tracking. In ICIP, 2016. 1, 2, 3, 6, 7
  6. 6.Guillem Brasó and Laura Leal-Taixe. Learning a neural solver for multiple object tracking. In CVPR, 2020. 1, 2, 3, 6
  7. 7.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018. 6
  8. 8.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. In ECCV, 2020. 2, 3, 4, 5, 6, 7
  9. 9.Yihong Chen, Yue Cao, Han Hu, and Liwei Wang. Memory enhanced global-local aggregation for video object detection. In CVPR, 2020. 3
  10. 10.Peng Chu, Jiang Wang, Quanzeng You, Haibin Ling, and Zicheng Liu. Transmot: Spatial-temporal graph transformer for multiple object tracking. arXiv:2104.00194, 2021. 7, 8
  11. 11.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In ICCV, 2017. 6
  12. 12.Peng Dai, Renliang Weng, Wongun Choi, Changshui Zhang, Zhangping He, and Wei Ding. Learning a proposal classifier for multiple object tracking. CVPR, 2021. 2
  13. 13.Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan. Tao: A large-scale benchmark for tracking any object. In ECCV, 2020. 3, 4, 5, 6, 7
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 2
  15. 15.Fei Du, Bo Xu, Jiasheng Tang, Yuqi Zhang, Fan Wang, and Hao Li. 1st place solution to ECCV-TAO-2020: Detect and represent any object for tracking. arXiv:2101.08040, 2021. 7
  16. 16.Davi Frossard and Raquel Urtasun. End-to-end learning of multi-sensor 3D tracking by detection. In ICRA, 2018. 3
  17. 17.Shanghua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip HS Torr. Res2net: A new multi-scale backbone architecture. TPAMI, 2019. 6
  18. 18.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 7
  19. 19.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In CVPR, 2019. 5, 6
  20. 20.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- ́ shick. Mask r-cnn. In ICCV, 2017. 1, 2, 3, 6
  21. 21.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. TPAMI, 2019. 7
  22. 22.Chao Liang, Zhipeng Zhang, Yi Lu, Xue Zhou, Bing Li, Xiyong Ye, and Jianxiao Zou. Rethinking the competition between detection and reid in multi-object tracking. arXiv:2010.12138, 2020. 8
  23. 23.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ́ ICCV, 2017. 4
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ́ Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 5, 6
  25. 25.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv:2103.14030, 2021. 2
  26. 26.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3D object detection via transformers. arXiv:2104.00678, 2021. 4
  27. 27.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2017. 6
  28. 28.Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixe, and Bastian Leibe. ́ Hota: A higher order metric for evaluating multi-object tracking. IJCV, 2021. 5
  29. 29.Hao Luo, Youzhi Gu, Xingyu Liao, Shenqi Lai, and Wei Jiang. Bag of tricks and a strong baseline for deep person re-identification. In CVPR Workshops, 2019. 5
  30. 30.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. Trackformer: Multi-object tracking with transformers. arXiv:2101.02702. 2, 3, 7, 8
  31. 31.A. Milan, L. Leal-Taixe, I. Reid, S. Roth, and K. ́ Schindler. MOT16: A benchmark for multi-object tracking. arXiv:1603.00831, 2016. 2, 3, 5, 6
  32. 32.Jiangmiao Pang, Linlu Qiu, Xia Li, Haofeng Chen, Qi Li, Trevor Darrell, and Fisher Yu. Quasi-dense similarity learning for multiple object tracking. In CVPR, 2021. 2, 7, 8
  33. 33.Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu. Chained-tracker: Chaining paired attentive regression results for end-to-end joint multiple-object detection and tracking. In ECCV, 2020. 8
  34. 34.Jinlong Peng, Tao Wang, Weiyao Lin, Jian Wang, John See, Shilei Wen, and Erui Ding. Tpm: Multiple object tracking with tracklet-plane matching. Pattern Recognition, 2020. 2
  35. 35.Esteban Real, Jonathon Shlens, Stefano Mazzocchi, Xin Pan, and Vincent Vanhoucke. Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. In CVPR, 2017. 7
  36. 36.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015. 1, 3, 4
  37. 37.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015. 3, 7
  38. 38.Chaobing Shan, Chunbo Wei, Bing Deng, Jianqiang Huang, Xian-Sheng Hua, Xiaoliang Cheng, and Kewei Liang. Tracklets predicting based adaptive graph tracking. arXiv:2010.09015, 2020. 8
  39. 39.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv:1805.00123, 2018. 6
  40. 40.Peize Sun, Jinkun Cao, Yi Jiang, Rufeng Zhang, Enze Xie, Zehuan Yuan, Changhu Wang, and Ping Luo. Transtrack: Multiple-object tracking with transformer. arXiv:2012.15460, 2020. 2, 3, 6, 7, 8
  41. 41.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. Sparse r-cnn: End-to-end object detection with learnable proposals. CVPR, 2021. 4
  42. 42.Zhiqing Sun, Shengcao Cao, Yiming Yang, and Kris Kitani. Rethinking transformer-based set prediction for object detection. arXiv:2011.10881, 2020. 4
  43. 43.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In CVPR, 2020. 6
  44. 44.Siyu Tang, Mykhaylo Andriluka, Bjoern Andres, and Bernt Schiele. Multiple people tracking by lifted multicut and person re-identification. In CVPR, 2017. 1, 2
  45. 45.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. FCOS: Fully convolutional one-stage object detection. In ICCV, 2019. 3
  46. 46.Pavel Tokmakov, Jie Li, Wolfram Burgard, and Adrien Gaidon. Learning to track with object permanence. In ICCV, 2021. 2
  47. 47.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé J ́ egou. Training data-efficient image transformers & distillation through attention. arXiv:2012.12877, 2020. 2
  48. 48.ultralytics. Yolov5. https : / / github . com / ultralytics/yolov5, 2020. 7
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017. 2, 5
  50. 50.Qiang Wang, Yun Zheng, Pan Pan, and Yinghui Xu. Multiple object tracking with correlation learning. CVPR, 2021. 8
  51. 51.Weiyao Wang, Matt Feiszli, Heng Wang, and Du Tran. Unidentified video objects: A benchmark for dense, openworld segmentation. arXiv:2104.04691, 2021. 8
  52. 52.Yongxin Wang, Kris Kitani, and Xinshuo Weng. Joint object detection and multi-object tracking with graph neural networks. In ICRA, 2021. 8
  53. 53.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 2021. 2
  54. 54.Zhongdao Wang, Liang Zheng, Yixuan Liu, and Shengjin Wang. Towards real-time multi-object tracking. In ECCV, 2020. 1, 2, 5
  55. 55.Nicolai Wojke and Alex Bewley. Deep cosine metric learning for person re-identification. In WACV, 2018. 1, 2
  56. 56.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. Long-term feature banks for detailed video understanding. In CVPR, 2019. 3
  57. 57.Haiping Wu, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Sequence level semantics aggregation for video object detection. In ICCV, 2019. 3
  58. 58.Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang, Ming Yang, and Junsong Yuan. Track to detect and segment: An online multi-object tracker. In CVPR, 2021. 8
  59. 59.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github. com/facebookresearch/detectron2, 2019. 6
  60. 60.Jiarui Xu, Yue Cao, Zheng Zhang, and Han Hu. Spatialtemporal relation networks for multi-object tracking. In ICCV, 2019. 1, 2
  61. 61.Yihong Xu, Yutong Ban, Guillaume Delorme, Chuang Gan, Daniela Rus, and Xavier Alameda-Pineda. Transcenter: Transformers with dense queries for multiple-object tracking. arXiv:2103.15145, 2021. 2, 7, 8
  62. 62.Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell. Deep layer aggregation. In CVPR, 2018. 6
  63. 63.Qian Yu, Gerard Medioni, and Isaac Cohen. Multiple target tracking using spatio-temporal markov chain monte carlo data association. In CVPR, 2007. 1
  64. 64.Fangao Zeng, Bin Dong, Tiancai Wang, Cheng Chen, Xiangyu Zhang, and Yichen Wei. End-to-end multiple-object tracking with transformer. arXiv:2105.03247, 2021. 2, 3, 7, 8
  65. 65.Li Zhang, Yuan Li, and Ramakant Nevatia. Global data association for multi-object tracking using network flows. In CVPR, 2008. 1, 2, 3
  66. 66.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and reidentification in multiple object tracking. arXiv:2004.01888, 2020. 1, 2, 3, 5, 6, 7, 8
  67. 67.Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring self-attention for image recognition. In CVPR, 2020. 2
  68. 68.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Tracking objects as points. ECCV, 2020. 1, 2, 3, 5, 6, 8
  69. 69.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Probabilistic two-stage detection. In arXiv:2103.07461, 2021. 2, 6
  70. 70.Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. arXiv:1904.07850, 2019. 1, 3, 4, 6
  71. 71.Tianyu Zhu, Markus Hiller, Mahsa Ehsanpour, Rongkai Ma, Tom Drummond, and Hamid Rezatofighi. Looking beyond two frames: End-to-end multi-object tracking using spatial and temporal transformers. arXiv:2103.14829, 2021. 3
  72. 72.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR, 2021. 2, 4

Citation

MLA
Zhou, X., et al. “Global Tracking Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.13250v2.
APA
Zhou, X., Yin, T., Koltun, V., & Krähenbühl, P. (2022). Global Tracking Transformers. arXiv. http://arxiv.org/abs/2203.13250v2
Chicago
Zhou, X., T. Yin, V. Koltun, and P. Krähenbühl. 2022. “Global Tracking Transformers”. arXiv. http://arxiv.org/abs/2203.13250v2.
Harvard
Zhou, X. et al. (2022) “Global Tracking Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.13250v2.
Vancouver
1. Zhou X, Yin T, Koltun V, Krähenbühl P (2022) Global Tracking Transformers. arXiv

BibTeX

@article{zhou2022global,
  title = {Global Tracking Transformers},
  author = {Zhou, Xingyi and Yin, Tianwei and Koltun, Vladlen and Krähenbühl, Philipp},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.13250v2},
  eprint = {2203.13250}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE