TrackFormer: Multi-Object Tracking with Transformers

Tim MeinhardtAlexander KirillovLaura Leal-TaixéChristoph Feichtenhofer

article2022CVPR1,134 citations

Introduces TrackFormer, an end-to-end multi-object tracking framework that unifies detection and frame-to-frame data association into a single Transformer-based set prediction architecture using autoregressive track queries.

Listen

Multi-object tracking is critical for computer vision applications such as autonomous navigation, robotics, and automated surveillance. Traditional approaches separate the task into distinct steps—first detecting objects on individual frames and then linking those detections across time using complex graph algorithms, motion models, or appearance matching. These multi-step pipelines often struggle in crowded scenes, suffer from identity switches during occlusions, and involve computationally heavy optimization routines that limit real-time deployment.

The article introduces and evaluates TrackFormer, an end-to-end framework based on a Transformer encoder-decoder architecture that performs multi-object tracking and segmentation. The main objective is to demonstrate that object detection and frame-to-frame data association can be unified into a single "tracking-by-attention" process without relying on external motion models, appearance heuristics, or complex post-processing graphs.

The approach formulates multi-object tracking as a continuous set prediction problem across video frames. The model uses a standard convolutional neural network backbone to extract image features and an attention-based Transformer to reason about objects. It initializes new tracks using static object queries and maintains existing tracks over time via autoregressive track queries that embed identity and spatial location. The authors evaluated the system on standard benchmark datasets, including the MOT17 tracking benchmark and the MOTS20 segmentation challenge, using simulated frame pairs from person detection datasets to train the attention mechanisms effectively.

Across the evaluations, the article reports three key findings. First, TrackFormer achieved state-of-the-art results among online tracking methods trained on comparable public data, reaching a 74.1 tracking accuracy score and a 68.0 identity preservation score on MOT17 private detections. Second, when extended to video object segmentation on MOTS20, the model set a new benchmark by raising the tracking and segmentation accuracy from 40.6 to 54.9 and identity score from 42.4 to 63.6 compared to previous baseline models. Third, ablation analyses confirmed that autoregressive track queries and targeted data augmentations—such as temporal frame perturbation—are crucial, preventing tracking accuracy drops of more than 10 to 14 points.

These findings indicate that unifying detection and association into a single Transformer model significantly streamlines the tracking pipeline. Eliminating secondary graph solvers and handcrafted matching rules reduces engineering complexity and mitigates failure points in dense environments. For decision-makers and engineering leads, this architecture provides a simpler, highly competitive baseline that reduces the overhead of maintaining multi-component tracking software.

Organizations developing video analytics pipelines should consider adopting Transformer-based set prediction architectures to simplify system design and improve tracking consistency. Future work should focus on extending the training and inference pipeline beyond adjacent two-frame pairs to multi-frame temporal reasoning, which can further strengthen long-term tracking stability.

The primary limitation of the current approach is its dependence on large amounts of training data and its focus on short-term rather than long-term re-identification across prolonged occlusions with extensive movement. TrackFormer also runs at roughly 7.4 frames per second in the reported setup, which requires optimization for strict real-time edge applications. Nevertheless, confidence in the reported tracking improvements remains high due to rigorous benchmarking on standard industry datasets.

  • Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Learn the foundational DEtection TRansformer (DETR) framework, which introduces the set prediction formulation and object query decoding mechanism that TrackFormer extends to video multi-object tracking.
  • Paper: Transformer Tracking, Xin Chen et al. (2021). Understand how Transformer attention replaces traditional correlation operations for visual tracking, establishing the conceptual bridge from convolutional feature matching to attention-based tracking.
  • Paper: Tracking Objects as Points, Xingyi Zhou et al. (2020). Examine how single-stage simultaneous detection and tracking through continuous point association works before seeing how TrackFormer reformulates it autoregressively with Transformer queries.
  • Paper: FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking, Yifu Zhang et al. (2020). Discover how joint detection and appearance-based re-identification are balanced in single-network multi-object tracking pipelines prior to TrackFormer's query-based tracking-by-attention paradigm.
  • Paper: Simple online and realtime tracking, Alex Bewley et al. (2016). Review the classic tracking-by-detection baseline that TrackFormer aims to supersede by replacing heuristic Kalman filtering and bipartite data association with end-to-end attention.
  • Paper: HOTA: A Higher Order Metric for Evaluating Multi-object Tracking, Jonathon Luiten et al. (2020). Study the Higher Order Tracking Accuracy (HOTA) evaluation metric used to benchmark and balance detection and association performance in multi-object tracking methods like TrackFormer.
  • Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). Explore the standard MOT benchmark and evaluation protocols that define the multi-object tracking challenges TrackFormer is evaluated on.
Cover for TrackFormer: Multi-Object Tracking with Transformers

Abstract

The challenging task of multi-object tracking (MOT) requires simultaneous reasoning about track initialization, identity, and spatio-temporal trajectories. We formulate this task as a frame-to-frame set prediction problem and introduce TrackFormer, an end-to-end trainable MOT approach based on an encoder-decoder Transformer architecture. Our model achieves data association between frames via attention by evolving a set of track predictions through a video sequence. The Transformer decoder initializes new tracks from static object queries and autoregressively follows existing tracks in space and time with the conceptually new and identity preserving track queries. Both query types benefit from self- and encoder-decoder attention on global frame-level features, thereby omitting any additional graph optimization or modeling of motion and/or appearance. TrackFormer introduces a new tracking-by-attention paradigm and while simple in its design is able to achieve state-of-the-art performance on the task of multi-object tracking (MOT17) and segmentation (MOTS20). The code is available at https://github.com/timmeinhardt/trackformer

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. TrackFormer
  • 3.1. MOT as a set prediction problem
  • 3.2. Tracking-by-attention with queries
  • 3.3. TrackFormer training
  • 4. Experiments
  • 4.1. MOT benchmarks and metrics
  • 4.2. Implementation details
  • 4.3. Benchmark results
  • 4.4. Ablation study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — TrackFormer Architecture and Tracking-by-Attention Paradigm

    model/method

    TrackFormer is an end-to-end trainable multi-object tracking (MOT) model based on an encoder-decoder Transformer architecture that performs joint object detection and frame-to-frame data association via attention. Given an input video, the architecture processes video frames sequentially in four steps:

    1. Feature Extraction: A convolutional neural network (CNN) backbone (such as ResNet-50) extracts 2D feature maps from the current frame.
    2. Feature Encoding: A Transformer encoder processes frame features via multi-head self-attention.
    3. Query Decoding: A Transformer decoder updates a set of N=Nobject+NtrackN = N_{\text{object}} + N_{\text{track}} queries using self-attention over all queries and cross-attention over stacked feature maps of the previous and current frame.
    4. Prediction Mapping: Multi-layer perceptrons (MLPs) map the updated decoder output embeddings to bounding box coordinates and categorical class probabilities.

    The queries decoded at frame tt comprise two distinct types:

    • Static Object Queries (NobjectN_{\text{object}}): A fixed set of learned positional embeddings responsible for detecting newly appearing objects in the frame.
    • Autoregressive Track Queries (NtrackN_{\text{track}}): Embeddings initialized directly from the output embeddings of active tracks in frame t−1t-1 (those whose classification score exceeded an initialization threshold σobject\sigma_{\text{object}}). Each track query follows an object across time, carrying over its identity information.

    Decoder self-attention across the combined set of Nobject+NtrackN_{\text{object}} + N_{\text{track}} queries enables simultaneous track propagation and new object initialization, while inherently preventing duplicate detections of already tracked objects without external appearance/motion models or graph optimization. Tracks are removed if their classification score drops below σtrack\sigma_{\text{track}} or via a non-maximum suppression (NMS) step with an IoU threshold σNMS\sigma_{\text{NMS}} applied to remove duplicate detections.

  2. Knowl 2 — TrackFormer Frame-to-Frame Set Prediction Loss and Bipartite Matching

    model/method

    TrackFormer trains frame-to-frame tracking over adjacent video frame pairs (t−1,t)(t-1, t) by extending the Hungarian matching set prediction objective to multi-object tracking.

    Let KtK_t and Kt−1K_{t-1} denote the sets of ground-truth object identities present at frames tt and t−1t-1, respectively. The N=Nobject+NtrackN = N_{\text{object}} + N_{\text{track}} model predictions y^={y^j}j=1N\hat{y} = \{\hat{y}_j\}_{j=1}^N at frame tt are assigned to ground truth objects y={yi}y = \{y_i\} (where each yi=(bi,ci,ki)y_i = (b_i, c_i, k_i) contains bounding box bib_i, class cic_i, and identity kik_i) or the background class (c=0c=0) according to three disjoint sets:

    1. Persisting Tracks (Kt∩Kt−1K_t \cap K_{t-1}): Ground truth objects tracked from frame t−1t-1 are deterministically matched to their corresponding track query outputs by track identity kk.
    2. Terminated/Occluded Tracks (Kt−1∖KtK_{t-1} \setminus K_t): Track queries corresponding to objects that left the scene or became occluded are assigned to the background class (ci=0c_i = 0).
    3. New Objects (Kobject=Kt∖Kt−1K_{\text{object}} = K_t \setminus K_{t-1}): Newly appeared ground truth objects are matched injectively to the NobjectN_{\text{object}} object query predictions via minimum-cost bipartite matching: σ^=arg⁡min⁡σ∑ki∈KobjectCmatch(yi,y^σ(i))\hat{\sigma} = \arg\min_{\sigma} \sum_{k_i \in K_{\text{object}}} \mathcal{C}_{\text{match}}(y_i, \hat{y}_{\sigma(i)})

    The pair-wise matching cost Cmatch\mathcal{C}_{\text{match}} between ground truth yiy_i and prediction y^σ(i)\hat{y}_{\sigma(i)} is: Cmatch(yi,y^σ(i))=−λclsp^σ(i)(ci)+Cbox(bi,b^σ(i))\mathcal{C}_{\text{match}}(y_i, \hat{y}_{\sigma(i)}) = -\lambda_{\text{cls}} \hat{p}_{\sigma(i)}(c_i) + \mathcal{C}_{\text{box}}(b_i, \hat{b}_{\sigma(i)}) where p^σ(i)(ci)\hat{p}_{\sigma(i)}(c_i) is the predicted probability for class cic_i, and the box cost Cbox\mathcal{C}_{\text{box}} combines ℓ1\ell_1 loss and Generalized Intersection over Union (GIoU) cost Ciou\mathcal{C}_{\text{iou}}: Cbox(bi,b^σ(i))=λℓ1∥bi−b^σ(i)∥1+λiouCiou(bi,b^σ(i))\mathcal{C}_{\text{box}}(b_i, \hat{b}_{\sigma(i)}) = \lambda_{\ell_1} \|b_i - \hat{b}_{\sigma(i)}\|_1 + \lambda_{\text{iou}} \mathcal{C}_{\text{iou}}(b_i, \hat{b}_{\sigma(i)}) with hyperparameters λcls,λℓ1,λiou∈R\lambda_{\text{cls}}, \lambda_{\ell_1}, \lambda_{\text{iou}} \in \mathbb{R}.

    Given the full assignment mapping π\pi, the total multi-object tracking loss LMOT\mathcal{L}_{\text{MOT}} is computed over all NN queries: LMOT(y,y^,π)=∑i=1NLquery(y,y^i,π)\mathcal{L}_{\text{MOT}}(y, \hat{y}, \pi) = \sum_{i=1}^N \mathcal{L}_{\text{query}}(y, \hat{y}_i, \pi) Lquery(y,y^i,π)={−λclslog⁡p^i(cπ=i)+Lbox(bπ=i,b^i),if i∈π−λclslog⁡p^i(0),if i∉π\mathcal{L}_{\text{query}}(y, \hat{y}_i, \pi) = \begin{cases} -\lambda_{\text{cls}} \log \hat{p}_i(c_{\pi=i}) + \mathcal{L}_{\text{box}}(b_{\pi=i}, \hat{b}_i), & \text{if } i \in \pi \\ -\lambda_{\text{cls}} \log \hat{p}_i(0), & \text{if } i \notin \pi \end{cases} where Lbox\mathcal{L}_{\text{box}} is defined identically to Cbox\mathcal{C}_{\text{box}}.

  3. Knowl 3 — Track Query Re-Identification Mechanism

    model/method

    TrackFormer performs short-term track re-identification entirely through its Transformer decoder without requiring separate appearance re-ID networks or metric learning heads.

    When a track is lost (due to occlusion or temporary drop in classification score below σtrack\sigma_{\text{track}}), its track query is not discarded immediately. Instead, the model continues decoding the inactive track query for up to Ttrack-reidT_{\text{track-reid}} consecutive frames (the patience window). During this period:

    • The inactive track query does not output active bounding box predictions to the resulting trajectory.
    • If in any subsequent frame within the window the track query's classification score exceeds a threshold σtrack-reid\sigma_{\text{track-reid}}, the track is re-identified and reactivated with its original track identity preserved.

    Because track queries retain localized spatial and identity representations, this mechanism recovers tracks through short-term occlusions while relying on the same unified attention mechanism used for regular tracking.

  4. Knowl 4 — Track Augmentation Strategies for Frame-to-Frame Transformer Training

    model/method

    Training TrackFormer on strictly adjacent video frames limits the diversity of spatial displacements and occlusion events observed by track queries. Three track augmentation techniques are applied during training on frame pairs (t−1,t)(t-1, t):

    1. Temporal Range Sampling: Frame t−1t-1 is sampled from a temporal window around frame tt (rather than strictly t−1t-1). This simulates camera motion, fast object motion, and lower frame rates.
    2. False Negative Sampling: Track queries from frame t−1t-1 are randomly discarded with probability pFNp_{\text{FN}} before decoding frame tt. The corresponding ground truth objects in frame tt must then be detected as new tracks by static object queries, preventing degenerate reliance on track queries alone.
    3. False Positive Sampling: Background output embeddings from frame t−1t-1 (especially those with high spatial overlap with true objects) are added to the set of track queries for frame tt with probability pFPp_{\text{FP}}. This trains the decoder to classify spurious track queries as background (c=0c=0) and terminate them in occlusion scenarios.
  5. Knowl 5 — TrackFormer Extension to Multi-Object Tracking and Segmentation (MOTS)

    model/method

    TrackFormer extends directly from multi-object bounding box tracking to multi-object tracking and segmentation (MOTS) by incorporating an instance mask head.

    The mask head generates spatial attention maps between encoded frame feature maps and decoder output embeddings (for both object and track queries). Upscaling and convolutional layers convert these attention maps into pixel-level instance segmentation masks for each active track.

    For segmentation on MOTS20, TrackFormer uses single-scale DETR attention rather than multi-scale deformable attention to reduce GPU memory consumption on full-resolution feature maps and avoid mask degradation associated with sparse deformable attention. Training follows a three-stage schedule: (1) private detection pretraining on CrowdHuman and MOT17, (2) freezing the backbone and Transformer to train the segmentation head on person instances from COCO, and (3) end-to-end fine-tuning on MOTS20.

  6. Knowl 6 — MOT17 Multi-Object Tracking Benchmark Results

    data/table

    TrackFormer was evaluated on the MOT17 benchmark test set across both public detection and private detection tracks. Public detection evaluation enforces track initialization filtering with public detections based on a minimum IoU threshold.

    Method Data FPS ↑\uparrow MOTA ↑\uparrow IDF1 ↑\uparrow MT ↑\uparrow ML ↓\downarrow FP ↓\downarrow FN ↓\downarrow ID Sw. ↓\downarrow
    Public Detections (Online)
    FAMNet – – 52.0 48.7 450 787 14138 253616 3072
    Tracktor++ M+C 1.3 56.3 55.1 498 831 8866 235449 1987
    GSM M+C – 56.4 57.8 523 813 14379 230174 1485
    CenterTrack – 17.7 60.5 55.7 580 777 11599 208577 2540
    TMOH – – 62.1 62.8 633 739 10951 201195 1897
    TrackFormer – 7.4 62.3 57.6 688 638 16591 192123 4018
    Private Detections (Online)
    TubeTK JTA – 63.0 58.6 735 468 27060 177483 4137
    CTracker – – 66.6 57.4 759 570 22284 160491 5529
    CenterTrack CH 17.7 67.8 64.7 816 579 18498 160332 3039
    QuasiDense – – 68.7 66.3 957 516 26589 146643 3378
    TraDeS CH – 69.1 63.9 858 507 20892 150060 3555
    GSDT 6M – 73.2 66.5 981 411 26397 120666 3891
    FairMOT CH+PD – 73.7 72.3 1017 408 27507 117477 3303
    PermaTrack CH+PD – 73.8 68.9 1032 405 28998 115104 3699
    GRTU CH+6M – 75.5 76.9 1158 495 27813 108690 1572
    TLR CH+6M – 76.5 73.6 1122 300 29808 99510 3369
    TrackFormer CH 7.4 74.1 68.0 1113 246 34602 108777 2829

    Training datasets: CH = CrowdHuman, PD = Parallel Domain, 6M = 6 tracking datasets, M = Market1501, C = CUHK03, JTA = Joint Track Auto. Metrics reported: Multiple Object Tracking Accuracy (MOTA ↑\uparrow), Identity F1 Score (IDF1 ↑\uparrow), Mostly Tracked targets (MT ↑\uparrow), Mostly Lost targets (ML ↓\downarrow), False Positives (FP ↓\downarrow), False Negatives (FN ↓\downarrow), Identity Switches (ID Sw. ↓\downarrow), Frames Per Second (FPS ↑\uparrow).

    Under private detections, TrackFormer achieves 74.1 MOTA and 68.0 IDF1 among methods trained exclusively on CrowdHuman (surpassing CenterTrack by +6.3 MOTA and +3.3 IDF1). Under public detections, TrackFormer achieves 62.3 MOTA without CrowdHuman pretraining.

  7. Knowl 7 — MOTS20 Benchmark Tracking and Segmentation Performance

    data/table

    TrackFormer was evaluated on the Multi-Object Tracking and Segmentation (MOTS20) benchmark across 4-fold cross-validation on the train set and on the official test set.

    Method TbD sMOTSA ↑\uparrow IDF1 ↑\uparrow FP ↓\downarrow FN ↓\downarrow ID Sw. ↓\downarrow
    Train set (4-fold cross-validation)
    MHT_DAM ×\times 48.0 – – – –
    FWT ×\times 49.3 – – – –
    MOTDT ×\times 47.8 – – – –
    jCC ×\times 48.3 – – – –
    TrackRCNN 52.7 – – – –
    MOTSNet 56.8 – – – –
    PointTrack 58.1 – – – –
    TrackFormer 58.7 – – – –
    Test set
    Track R-CNN 40.6 42.4 1261 12641 567
    TrackFormer 54.9 63.6 2233 7195 278

    Methods marked with TbD perform tracking-by-detection on public detections without segmentation, followed by Mask R-CNN mask generation. Metrics include soft Multi-Object Tracking and Segmentation Accuracy (sMOTSA), Identity F1 (IDF1), False Positives (FP), False Negatives (FN), and Identity Switches (ID Sw.).

    TrackFormer achieves 58.7 sMOTSA on the 4-fold cross-validation train set and 54.9 sMOTSA with 63.6 IDF1 on the MOTS20 test set, outperforming Track R-CNN by +14.3 sMOTSA and +21.2 IDF1 while reducing ID switches from 567 to 278.

  8. Knowl 8 — Ablation of TrackFormer Components and Augmentations

    data/table

    The contribution of individual TrackFormer components, pretraining, and training augmentations was evaluated on a 50-50 train/val frame split of the MOT17 training set under private detection settings.

    Method MOTA ↑\uparrow Δ\Delta IDF1 ↑\uparrow Δ\Delta
    TrackFormer (Full) 71.3 73.4
    w\ o Pretraining on CrowdHuman 69.3 -2.0 71.8 -1.6
    w\ o Track query re-identification 69.2 -0.1 70.4 -1.4
    w\ o Track augmentations (FP) 68.4 -0.8 70.0 -0.4
    w\ o Track augmentations (Range) 64.0 -4.4 59.2 -10.8
    w\ o Track queries (greedy matching) 61.0 -3.0 45.1 -14.1

    Key takeaways from the ablation data:

    • Track Queries vs. Greedy Matching: Removing track queries and associating object-query detections across frames via greedy center-distance matching decreases MOTA by 3.0 points and IDF1 by 14.1 points, demonstrating the necessity of the tracking-by-attention query mechanism.
    • Temporal Range Augmentations: Sampling frames over a wider temporal range during training provides substantial robustness against motion, improving MOTA by 4.4 points and IDF1 by 10.8 points.
    • Track Query Re-identification: Retaining inactive track queries over short temporal windows yields a 1.4 point increase in IDF1.
    • False Positive (FP) Query Augmentations: Augmenting track queries with background embeddings prevents overfitting to false positives and improves MOTA by 0.8 points.
  9. Knowl 9 — Synergy of Joint Instance Mask Training on Bounding Box Tracking

    data/table

    The effect of joint instance mask segmentation training on standard bounding box tracking metrics was evaluated using 4-fold cross-validation on the MOTS20 training set.

    Method Mask training MOTA ↑\uparrow IDF1 ↑\uparrow
    TrackFormer ×\times 61.9 54.8
    TrackFormer ✓ 61.9 56.0

    Evaluation is conducted using standard bounding box ground-truth matching rather than mask IoU. Joint training with the segmentation mask prediction head improves IDF1 by +1.2 points (from 54.8 to 56.0) while keeping MOTA unchanged at 61.9. Pixel-level mask supervision provides auxiliary spatial cues that help the decoder resolve ambiguous occlusion events and maintain identity consistency during tracking.

Coverage note — No substantial contributed material was omitted; the knowls cover the TrackFormer architecture, set prediction formulation, track query mechanics, re-identification, training augmentations, segmentation extension, benchmark evaluations on MOT17 and MOTS20, and all ablation studies.

References

  1. 1.Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social lstm: Human trajectory prediction in crowded spaces. IEEE Conf. Comput. Vis. Pattern Recog., 2016. 2
  2. 2.Anton Andriyenko and Konrad Schindler. Multi-target tracking by continuous energy minimization. IEEE Conf. Comput. Vis. Pattern Recog., 2011. 2
  3. 3.Jerome Berclaz, Francois Fleuret, Engin Turetken, and Pascal Fua. Multiple object tracking using k-shortest paths optimization. IEEE Trans. Pattern Anal. Mach. Intell., 2011. 2
  4. 4.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In Int. Conf. Comput. Vis., 2019. 1, 2, 7
  5. 5.Keni Bernardin and Rainer Stiefelhagen. Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008, 2008. 6
  6. 6.Guillem Braso and Laura Leal-Taix e. Learning a neural solver for multiple object tracking. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 1, 2, 7
  7. 7.Nicolas Carion, F. Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. Eur. Conf. Comput. Vis., 2020. 1, 2, 3, 5, 6
  8. 8.Long Chen, Haizhou Ai, Zijie Zhuang, and Chong Shang. Real-time multiple people tracking with deeply learned candidate selection and person re-identification. In Int. Conf. Multimedia and Expo, 2018. 1, 2, 7
  9. 9.Wongun Choi and Silvio Savarese. Multiple target tracking in world coordinate with single, minimally calibrated camera. Eur. Conf. Comput. Vis., 2010. 2
  10. 10.Peng Chu and Haibin Ling. Famnet: Joint learning of feature, affinity and multi-dimensional assignment for online multiple object tracking. In Int. Conf. Comput. Vis., 2019. 2, 7
  11. 11.Qi Chu, Wanli Ouyang, Hongsheng Li, Xiaogang Wang, Bin Liu, and Nenghai Yu. Online multi-object tracking using cnn-based single object tracker with spatial-temporal attention mechanism. In Proceedings of the IEEE International Conference on Computer Vision, pages 4836–4845, 2017. 1, 2
  12. 12.Patrick Dendorfer, Aljosa Osep, Anton Milan, Daniel Cremers, Ian Reid, Stefan Roth, and Laura Leal-Taixe. Motchallenge: A benchmark for single-camera multiple target tracking. Int. J. Comput. Vis., 2020. 1
  13. 13.Matteo Fabbri, Fabio Lanzi, Simone Calderara, Andrea Palazzi, Roberto Vezzani, and Rita Cucchiara. Learning to detect and track visible and occluded body joints in a virtual world. In European Conference on Computer Vision (ECCV), 2018. 7
  14. 14.Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. Detect to track and track to detect. In ICCV, 2017. 2
  15. 15.Pedro F. Felzenszwalb, Ross B. Girshick, David A. McAllester, and Deva Ramanan. Object detection with discriminatively trained part based models. IEEE Trans. Pattern Anal. Mach. Intell., 2009. 6
  16. 16.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 2, 7
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 2, 3, 6
  18. 18.Roberto Henschel, Laura Leal-Taixe, Daniel Cremers, and Bodo Rosenhahn. Improvements to frank-wolfe optimization for multi-detector multi-object tracking. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 1, 2, 7
  19. 19.Andrea Hornakova, Roberto Henschel, Bodo Rosenhahn, and Paul Swoboda. Lifted disjoint paths with application in multiple object tracking. In Int. Conf. Mach. Learn., 2020. 2, 7
  20. 20.Hao Jiang, Sidney S. Fels, and James J. Little. A linear programming approach for multiple object tracking. IEEE Conf. Comput. Vis. Pattern Recog., 2007. 2
  21. 21.Margret Keuper, Siyu Tang, Bjoern Andres, Thomas Brox, and Bernt Schiele. Motion segmentation & multiple object tracking by correlation co-clustering. In IEEE Trans. Pattern Anal. Mach. Intell., 2018. 1, 2, 7
  22. 22.Chanho Kim, Fuxin Li, Arridhana Ciptadi, and James M. Rehg. Multiple hypothesis tracking revisited. In Int. Conf. Comput. Vis., 2015. 1, 2, 7
  23. 23.Laura Leal-Taixe, Cristian Canton-Ferrer, and Konrad Schindler. Learning by tracking: siamese cnn for robust target association. IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2016. 1, 2
  24. 24.Laura Leal-Taixe, Michele Fenzi, Alina Kuznetsova, Bodo Rosenhahn, and Silvio Savarese. Learning an image-based motion context for multiple people tracking. IEEE Conf. Comput. Vis. Pattern Recog., 2014. 2
  25. 25.Laura Leal-Taixe, Gerard Pons-Moll, and Bodo Rosenhahn. Everybody needs somebody: Modeling social and grouping behavior on a linear programming multiple people tracker. Int. Conf. Comput. Vis. Workshops, 2011. 1, 2
  26. 26.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deepreid: Deep filter pairing neural network for person reidentification. In IEEE Conf. Comput. Vis. Pattern Recog., 2014. 7
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollar. Microsoft coco: Common objects in context. arXiv:1405.0312, 2014. 6
  28. 28.Qiankun Liu, Qi Chu, Bin Liu, and Nenghai Yu. Gsm: Graph similarity model for multi-object tracking. In Int. Joint Conf. Art. Int., 2020. 1, 2, 7
  29. 29.Anton Milan, Laura Leal-Taixe, Ian D. Reid, Stefan Roth, and Konrad Schindler. Mot16: A benchmark for multi-object tracking. arXiv:1603.00831, 2016. 2, 6, 7, 8
  30. 30.Aljosa Oˇ sep, Wolfgang Mehner, Paul Voigtlaender, and Bastian Leibe. Track, then decide: Category-agnostic visionbased multi-object tracking. IEEE Int. Conf. Rob. Aut., 2018. 2
  31. 31.Bo Pang, Yizhuo Li, Yifan Zhang, Muchen Li, and Cewu Lu. Tubetk: Adopting tubes to track multi-object in a one-step training model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 7
  32. 32.Jiangmiao Pang, Linlu Qiu, Xia Li, Haofeng Chen, Qi Li, Trevor Darrell, and Fisher Yu. Quasi-dense similarity learning for multiple object tracking. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2021. 2, 7
  33. 33.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. arXiv preprint arXiv:1802.05751, 2018. 2
  34. 34.Stefano Pellegrini, Andreas Ess, Konrad Schindler, and Luc Van Gool. You’ll never walk alone: modeling social behavior for multi-target tracking. Int. Conf. Comput. Vis., 2009. 2
  35. 35.Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu. Chained-tracker: Chaining paired attentive regression results for end-to-end joint multiple-object detection and tracking. In Proceedings of the European Conference on Computer Vision, 2020. 7
  36. 36.Hamed Pirsiavash, Deva Ramanan, and Charless C. Fowlkes. Globally-optimal greedy algorithms for tracking a variable number of objects. IEEE Conf. Comput. Vis. Pattern Recog., 2011. 2
  37. 37.Lorenzo Porzi, Markus Hofinger, Idoia Ruiz, Joan Serrat, Samuel Rota Bulo, and Peter Kontschieder. Learning multiobject tracking and segmentation from automatic annotations. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 2, 7
  38. 38.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Adv. Neural Inform. Process. Syst., 2015. 1, 6
  39. 39.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 5
  40. 40.Ergys Ristani, Francesco Solera, Roger S. Zou, Rita Cucchiara, and Carlo Tomasi. Performance measures and a data set for multi-target, multi-camera tracking. In Eur. Conf. Comput. Vis. Workshops, 2016. 6
  41. 41.Ergys Ristani and Carlo Tomasi. Features for multi-target multi-camera tracking and re-identification. IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
  42. 42.Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese. Learning social etiquette: Human trajectory prediction. Eur. Conf. Comput. Vis., 2016. 2
  43. 43.Paul Scovanner and Marshall F. Tappen. Learning pedestrian dynamics from the real world. Int. Conf. Comput. Vis., 2009. 2
  44. 44.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv:1805.00123, 2018. 6, 7
  45. 45.H. Sheng, Y. Zhang, J. Chen, Z. Xiong, and J. Zhang. Heterogeneous association graph fusion for target association in multiple object tracking. IEEE Transactions on Circuits and Systems for Video Technology, 2019. 2, 7
  46. 46.Daniel Stadler and Jurgen Beyerer. Improving multiple pedestrian tracking by track management and occlusion handling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10958–10967, June 2021. 7
  47. 47.Russell Stewart, Mykhaylo Andriluka, and Andrew Y Ng. End-to-end people detection in crowded scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2325–2333, 2016. 2, 5
  48. 48.Siyu Tang, Mykhaylo Andriluka, Bjoern Andres, and Bernt Schiele. Multiple people tracking by lifted multicut and person re-identification. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 2
  49. 49.Pavel Tokmakov, Jie Li, Wolfram Burgard, and Adrien Gaidon. Learning to track with object permanence. In Int. Conf. Comput. Vis., 2021. 2, 7
  50. 50.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Adv. Neural Inform. Process. Syst., 2017. 1, 2, 3, 4
  51. 51.Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe. Mots: Multi-object tracking and segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 2, 6, 7, 8
  52. 52.Qiang Wang, Yun Zheng, Pan Pan, and Yinghui Xu. Multiple object tracking with correlation learning. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 7
  53. 53.Shuai Wang, Hao Sheng, Yang Zhang, Yubin Wu, and Zhang Xiong. A general recurrent tracking framework without real data. In Int. Conf. Comput. Vis., 2021. 7
  54. 54.Yongxin Wang, Kris Kitani, and Xinshuo Weng. Joint object detection and multi-object tracking with graph neural networks. In IEEE Int. Conf. Rob. Aut., May 2021. 2, 7
  55. 55.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. Endto-end video instance segmentation with transformers. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2021. 6
  56. 56.Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang, Ming Yang, and Junsong Yuan. Track to detect and segment: An online multi-object tracker. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021. 2, 7
  57. 57.Zhenbo Xu, Wei Zhang, Xiao Tan, Wei Yang, Huan Huang, Shilei Wen, Errui Ding, and Liusheng Huang. Segment as points for efficient online multi-object tracking and segmentation. In Eur. Conf. Comput. Vis., 2020. 2, 7
  58. 58.Kota Yamaguchi, Alexander C. Berg, Luis E. Ortiz, and Tamara L. Berg. Who are you with and where are you going? IEEE Conf. Comput. Vis. Pattern Recog., 2011. 2
  59. 59.Fan Yang, Wongun Choi, and Yuanqing Lin. Exploit all the layers: Fast and accurate cnn object detector with scale dependent pooling and cascaded rejection classifiers. IEEE Conf. Comput. Vis. Pattern Recog., 2016. 6
  60. 60.Fan Yang, Wongun Choi, and Yuanqing Lin. Exploit all the layers: Fast and accurate cnn object detector with scale dependent pooling and cascaded rejection classifiers. IEEE Conf. Comput. Vis. Pattern Recog., pages 2129–2137, 2016. 7
  61. 61.Qian Yu, Gerard Medioni, and Isaac Cohen. Multiple target tracking using spatio-temporal markov chain monte carlo data association. IEEE Conf. Comput. Vis. Pattern Recog., 2007. 2
  62. 62.Li Zhang, Yuan Li, and Ramakant Nevatia. Global data association for multi-object tracking using network flows. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008. 2
  63. 63.Y. Zhang, H. Sheng, Y. Wu, S. Wang, W. Lyu, W. Ke, and Z. Xiong. Long-term tracking with deep tracklet association. IEEE Trans. Image Process., 2020. 2, 7
  64. 64.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and re-identification in multiple object tracking. International Journal of Computer Vision, pages 1–19, 2021. 7
  65. 65.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. In Int. Conf. Comput. Vis., 2015. 7
  66. 66.Xingyi Zhou, Vladlen Koltun, and Philipp Krahenb uhl. Tracking objects as points. ECCV, 2020. 1, 2, 5, 6, 7, 8
  67. 67.Ji Zhu, Hua Yang, Nian Liu, Minyoung Kim, Wenjun Zhang, and Ming-Hsuan Yang. Online multi-object tracking with dual matching attention networks. In Eur. Conf. Comput. Vis., 2018. 1, 2
  68. 68.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. Int. Conf. Learn. Represent., 2021. 2, 6

Citation

MLA
Meinhardt, T., et al. “TrackFormer: Multi-Object Tracking with Transformers”. arXiv, 2021, http://arxiv.org/abs/2101.02702v3.
APA
Meinhardt, T., Kirillov, A., Leal-Taixe, L., & Feichtenhofer, C. (2021). TrackFormer: Multi-Object Tracking with Transformers. arXiv. http://arxiv.org/abs/2101.02702v3
Chicago
Meinhardt, T., A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer. 2021. “TrackFormer: Multi-Object Tracking with Transformers”. arXiv. http://arxiv.org/abs/2101.02702v3.
Harvard
Meinhardt, T. et al. (2021) “TrackFormer: Multi-Object Tracking with Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2101.02702v3.
Vancouver
1. Meinhardt T, Kirillov A, Leal-Taixe L, Feichtenhofer C (2021) TrackFormer: Multi-Object Tracking with Transformers. arXiv

BibTeX

@article{meinhardt2021trackformer,
  title = {TrackFormer: Multi-Object Tracking with Transformers},
  author = {Meinhardt, Tim and Kirillov, Alexander and Leal-Taixe, Laura and Feichtenhofer, Christoph},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2101.02702v3},
  eprint = {2101.02702}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE