DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

Peize SunJinkun CaoYi JiangZehuan YuanSong BaiKris KitaniPing Luo

article2022CVPR446 citations

Introduces DanceTrack, a large-scale multi-human tracking benchmark of group dancing scenes featuring uniform appearance and complex motion patterns, exposing the vulnerabilities of conventional appearance-based tracking models and directing research toward motion-centric association.

Listen

Multi-object tracking systems are widely used in video surveillance, autonomous driving, and robotics to detect targets and maintain their identities over time. Most existing algorithms rely heavily on visual re-identification to distinguish and associate targets across video frames. However, current benchmark datasets are skewed toward scenarios where people have distinct clothing and follow simple, predictable paths. In real-world environments where people wear uniform attire and move unpredictably, visual re-identification fails, exposing a critical vulnerability in current vision systems.

The article introduces DanceTrack, a large-scale video benchmark designed to evaluate multi-object tracking in scenarios characterized by uniform visual appearance, complex non-linear motion, and frequent occlusions. The article aims to demonstrate the shortcomings of current appearance-reliant tracking algorithms and provide empirical evidence on alternative strategies that improve tracking robustness.

To evaluate tracking systems, the authors assembled 100 group dancing videos comprising over 100,000 frames—nearly ten times the size of standard multi-human tracking benchmarks. The dataset emphasizes similar clothing, large-scale body deformation, and frequent crossovers. The authors conducted oracle analyses using ground-truth object bounding boxes to isolate association failures from detection errors. They also evaluated seven state-of-the-art tracking algorithms on both the established MOT17 dataset and DanceTrack, alongside ablation studies testing different motion models and fine-grained visual features.

The evaluation revealed several key findings. First, state-of-the-art tracking algorithms suffer a severe drop in tracking and association accuracy when moving from standard benchmarks to DanceTrack; for example, ByteTrack's Higher Order Tracking Accuracy dropped from 63.1% on MOT17 to 47.7% on DanceTrack, and its association accuracy fell by nearly half from 62.0% to 32.1%. Second, detection performance remained consistently high across all models, proving that target localization is not the limiting factor. Third, relying on appearance re-identification in uniform scenarios actually degraded performance; an oracle association using simple bounding-box overlap achieved an association accuracy of 53.6%, whereas incorporating visual re-identification reduced it to 43.2%. Finally, incorporating temporal motion dynamics and fine-grained visual representations such as human pose estimation and segmentation masks provided clear performance boosts.

These findings indicate that existing benchmarks have created an over-reliance on appearance-matching shortcuts, masking fundamental deficiencies in tracking logic. For practical applications where individuals share visual traits—such as sports, uniform workspaces, or crowded public venues—current computer vision systems risk frequent identity swaps and tracking errors. Relying solely on standard benchmark results introduces operational risks for deployment in complex environments.

The article recommends shifting tracking algorithm development away from pure appearance re-identification toward robust non-linear motion modeling and temporal dynamics. Practitioners should incorporate fine-grained features like human pose estimation into association pipelines. Furthermore, organizations deploying multi-object tracking in complex, uniform-appearance settings should benchmark their systems on diverse datasets rather than standard pedestrian datasets before full-scale deployment.

A current limitation of the DanceTrack dataset is that its official annotations are confined to bounding boxes rather than native segmentation masks or pose keypoints, requiring auxiliary datasets to train multi-task representations. In addition, the article focuses on evaluating existing models and demonstrating dataset utility rather than proposing a wholly new state-of-the-art tracking architecture. Nevertheless, the experimental results consistently demonstrate high confidence that appearance-only tracking is insufficient in complex visual settings.

No sufficiently relevant recommendations were found.

Cover for DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

Abstract

A typical pipeline for multi-object tracking (MOT) is to use a detector for object localization, and following re-identification (re-ID) for object association. This pipeline is partially motivated by recent progress in both object detection and re-ID, and partially motivated by biases in existing tracking datasets, where most objects tend to have distinguishing appearance and re-ID models are sufficient for establishing associations. In response to such bias, we would like to re-emphasize that methods for multi-object tracking should also work when object appearance is not sufficiently discriminative. To this end, we propose a large-scale dataset for multi-human tracking, where humans have similar appearance, diverse motion and extreme articulation. As the dataset contains mostly group dancing videos, we name it “DanceTrack”. We expect DanceTrack to provide a better platform to develop more MOT algorithms that rely less on visual discrimination and depend more on motion analysis. We benchmark several state-of-the-art trackers on our dataset and observe a significant performance drop on DanceTrack when compared against existing benchmarks. The dataset, project code and competition is released at: https://github.com/DanceTrack.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. DanceTrack
  • 3.1. Dataset Construction
  • 3.2. Dataset Statistic
  • 3.3. Evaluation Metrics
  • 3.4. Limitation
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Oracle Analysis
  • 4.3. Benchmark Results
  • 4.4. Association Strategy
  • 4.5. Analysis of More Modalities
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — DanceTrack Dataset Specification and Meta-Information

    definition

    DanceTrack is a large-scale multi-human object tracking dataset designed to evaluate multi-object tracking (MOT) algorithms in scenarios characterized by uniform visual appearance and diverse, non-linear motion patterns. In DanceTrack, targets (dancers across genres such as street dance, pop dance, classical dance, gymnastics, Kung Fu, and cheerleading) wear identical or highly similar clothing and undergo frequent crossovers, dynamic occlusions, and severe body articulations.

    The dataset consists of 100 video sequences recorded at 20 frames per second (FPS), totaling 105,855 image frames (approximately 10×\times the frame count of MOT17/MOT20) and 5,292 seconds of video. It contains 990 distinct trajectories with an average duration of 52.9 seconds and an average of 9 tracks per video (with up to 40 individuals simultaneously). The dataset is partitioned into:

    • Training set: 40 videos
    • Validation set: 25 videos
    • Test set: 35 videos
    Dataset MOT17 MOT20 DanceTrack
    Videos 14 8 100
    Avg. tracks 96 432 9
    Total tracks 1342 3456 990
    Avg. len. (s) 35.4 66.8 52.9
    Total len. (s) 463 535 5292
    FPS 30 25 20
    Total images 11,235 13,410 105,855

    Annotations consist of 2D bounding boxes and identity labels across time. For partially occluded humans, full-body bounding boxes are annotated. Fully occluded targets are not annotated during occlusion; when they reappear, their original identity is restored.

  2. Knowl 2 — Benchmark Evaluation of State-of-the-Art MOT Algorithms on MOT17 vs. DanceTrack

    data/table

    Leading multi-object tracking algorithms benchmarked under the private detection setting on the test sets of MOT17 and DanceTrack exhibit a severe drop in tracking and association accuracy on DanceTrack, despite maintaining high detection accuracy.

    MOT17 DanceTrack (Proposed)
    Methods HOTA DetA AssA MOTA IDF1 HOTA DetA AssA MOTA IDF1
    CenterTrack 52.2 53.8 51.0 67.8 64.7 41.8 78.1 22.6 86.8 35.7
    FairMOT 59.3 60.9 58.0 73.7 72.3 39.7 66.7 23.8 82.2 40.8
    QDTrack 53.9 55.6 52.7 68.7 66.3 45.7 72.1 29.2 83.0 44.8
    TransTrack 54.1 61.6 47.9 75.2 63.5 45.5 75.9 27.5 88.4 45.2
    TraDes 52.7 55.2 50.8 69.1 63.9 43.3 74.5 25.4 86.2 41.2
    MOTR 57.2 58.9 55.8 71.9 68.4 54.2 73.5 40.2 79.7 51.5
    ByteTrack 63.1 64.5 62.0 80.3 77.3 47.7 71.0 32.1 89.6 53.9

    While detection metrics on DanceTrack are consistently higher than on MOT17 (MOTA ranges between 79.7%79.7\% and 89.6%89.6\%, DetA between 66.7%66.7\% and 78.1%78.1\%), association performance collapses (AssA drops to 22.6%–40.2%22.6\%\text{--}40.2\%, IDF1 to 35.7%–53.9%35.7\%\text{--}53.9\%), resulting in depressed HOTA scores (39.7%–54.2%39.7\%\text{--}54.2\%). This confirms that tracking failure on DanceTrack stems from association difficulty in uniform-appearance and complex-motion scenarios rather than localization/detection failures.

  3. Knowl 3 — Oracle Association Analysis on Ground-Truth Detections

    data/table

    To isolate the association bottleneck from detection errors, tracking methods were evaluated using oracle (ground-truth) bounding box detections on the validation sets of MOT17 and DanceTrack across different association modalities: bounding-box IoU matching, linear motion modeling (Kalman Filter), and appearance matching (pre-trained re-ID feature cosine similarity).

    MOT17 DanceTrack (Proposed)
    Appearance IoU Motion HOTA DetA AssA MOTA IDF1 HOTA DetA AssA MOTA IDF1
    ✓ 98.1 98.9 97.3 98.0 97.8 72.8 98.9 53.6 98.7 63.5
    ✓ ✓ 96.4 97.1 95.8 99.7 98.1 69.4 87.9 54.8 99.4 71.3
    ✓ ✓ ✓ 95.0 94.7 95.4 99.3 98.8 59.7 82.5 43.2 97.2 60.5
    ✓ ✓ 93.3 99.0 87.9 98.9 90.9 68.0 97.7 47.4 97.9 58.7

    On MOT17, naive IoU matching with perfect detections reaches 98.198.1 HOTA and 97.397.3 AssA, showing that standard benchmarks are dominated by simple trajectory association. On DanceTrack, oracle IoU matching achieves only 72.872.8 HOTA and 53.653.6 AssA. Adding appearance features drastically degrades performance (HOTA drops from 72.872.8 to 59.7–68.059.7\text{--}68.0), proving that visual re-ID embeddings for uniform-dressed targets act as noise and induce false associations.

  4. Knowl 4 — Quantitative Metrics for Video Appearance Similarity and Relative Motion Diversity

    equation

    To quantify the distinctiveness of object appearance and the complexity of target motion within a video sequence, three metrics are formulated:

    1. Average Intra-Video Appearance Distance (VV):
    V=1T∑t=1T1Nt2∑i=1Nt∑j≠iNt(1−cos⁡⟨F(Bit),F(Bjt)⟩)V = \frac{1}{T} \sum_{t=1}^{T} \frac{1}{N_t^2} \sum_{i=1}^{N_t} \sum_{j \neq i}^{N_t} \left(1 - \cos\langle F(B_i^t), F(B_j^t)\rangle\right)

    where T∈NT \in \mathbb{N} is the number of frames in the video sequence, Nt∈NN_t \in \mathbb{N} is the number of visible objects in frame tt, BitB_i^t denotes the bounding box of object ii in frame tt, F(Bit)∈RdF(B_i^t) \in \mathbb{R}^d is the feature embedding from a pre-trained re-ID model, and cos⁡⟨u,v⟩=u⋅v∥u∥2∥v∥2\cos\langle \mathbf{u}, \mathbf{v}\rangle = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\|_2 \|\mathbf{v}\|_2} is the cosine similarity. A smaller VV indicates higher appearance uniformity among co-existing targets.

    1. Adjacent Frame IoU (UU):
    U=1N(T−1)∑i=1N∑t=1T−1IoU(Bit,Bit+1)U = \frac{1}{N(T-1)} \sum_{i=1}^{N} \sum_{t=1}^{T-1} \text{IoU}(B_i^t, B_i^{t+1})

    where N∈NN \in \mathbb{N} is the total number of objects, TT is the number of frames, and IoU(A,B)=∣A∩B∣∣A∪B∣\text{IoU}(A, B) = \frac{|A \cap B|}{|A \cup B|} denotes the Intersection-over-Union between bounding boxes in consecutive frames. UU measures displacement scale relative to frame rate.

    1. Frequency of Relative Position Switch (SS):
    S=∑i=1N∑j≠iN∑t=1T−1sw(Bit,Bjt,Bit+1,Bjt+1)2N(T−1)(N−1)S = \frac{\sum_{i=1}^{N} \sum_{j \neq i}^{N} \sum_{t=1}^{T-1} \text{sw}(B_i^t, B_j^t, B_i^{t+1}, B_j^{t+1})}{2N(T-1)(N-1)}

    where sw(Bit,Bjt,Bit+1,Bjt+1)=1\text{sw}(B_i^t, B_j^t, B_i^{t+1}, B_j^{t+1}) = 1 if objects ii and jj spatially overlap (bounding-box IoU>0\text{IoU} > 0) on frame tt or t+1t+1 and swap their relative horizontal (xx-coordinate) or vertical (yy-coordinate) center ordering between frame tt and frame t+1t+1, and sw(⋅)=0\text{sw}(\cdot) = 0 otherwise. A higher SS quantifies more frequent crossovers and non-linear trajectory interactions.

  5. Knowl 5 — Evaluation of Association Algorithms with a Fixed YOLOX Detector

    data/table

    Decoupling the detection stage from tracking by fixing the detector to a YOLOX model trained on DanceTrack training data reveals the isolated performance of different association schemes on the DanceTrack validation set:

    Association HOTA DetA AssA MOTA IDF1
    IoU 44.7 79.6 25.3 87.3 36.8
    SORT 47.8 74.0 31.0 88.2 48.3
    DeepSORT 45.8 70.9 29.7 87.1 46.8
    MOTDT 39.2 68.8 22.5 84.3 39.6
    BYTE 47.1 70.5 31.5 88.2 51.9

    Key takeaways:

    • Incorporating re-ID appearance matching into SORT to form DeepSORT reduces HOTA from 47.847.8 to 45.845.8 and IDF1 from 48.348.3 to 46.846.8, demonstrating that appearance cues degrade association when appearances are uniform.
    • MOTDT, which uses tracklets to guide bounding-box detection, achieves the lowest performance (39.239.2 HOTA), as detection on DanceTrack is already strong and tracking guidance introduces error propagation.
    • BYTE, which uses two-stage matching to recover low-confidence detections, achieves the highest AssA (31.531.5) and IDF1 (51.951.9).
  6. Knowl 6 — Comparison of Temporal Motion Models for Association on DanceTrack

    data/table

    Evaluating different motion models to capture temporal dynamics with fixed YOLOX detections on the DanceTrack validation set demonstrates the necessity of non-linear motion modeling:

    Motion Model HOTA DetA AssA MOTA IDF1
    None (IoU) 44.7 79.6 25.3 87.3 36.8
    Kalman filter 47.8 74.0 31.0 88.2 48.3
    LSTM 51.6 78.2 34.2 89.2 50.8

    Compared to frame-by-frame IoU without temporal modeling (44.744.7 HOTA, 25.325.3 AssA):

    • A linear Kalman filter boosts HOTA to 47.847.8 and AssA to 31.031.0 (+5.7+5.7 AssA).
    • A non-linear LSTM motion model achieves the highest overall accuracy, raising HOTA to 51.651.6 (+6.9+6.9 over IoU) and AssA to 34.234.2 (+8.9+8.9 over IoU). This demonstrates that non-linear temporal dynamics provide strong association cues when visual appearance is uninformative.
  7. Knowl 7 — Impact of Multi-Task Auxiliary Modalities on Association in DanceTrack

    data/table

    Because DanceTrack provides only bounding-box annotations, multi-task joint training with external datasets (COCO for instance masks and keypoints; KITTI for depth) was evaluated using a CenterNet detector and BYTE association on the DanceTrack validation set:

    Data Ass. Modality HOTA DetA AssA MOTA IDF1
    DanceTrack box 36.9 63.6 21.6 78.8 39.2
    + COCOmask box 38.1 (+1.2) 64.5 (+0.9) 22.6 (+1.0) 80.6 (+1.8) 40.3 (+1.1)
    + COCOmask + mask 39.2 (+1.1) 64.9 (+0.4) 23.9 (+1.3) 80.7 (+0.1) 41.6 (+0.3)
    DanceTrack box 36.9 63.6 21.6 78.8 39.2
    + COCOpose box 40.6 (+3.7) 65.5 (+1.9) 25.3 (+3.7) 82.9 (+4.1) 42.9 (+3.7)
    + COCOpose + pose 41.0 (+0.4) 65.9 (+0.4) 25.6 (+0.3) 83.1 (+0.3) 43.9 (+1.0)
    DanceTrack box 36.9 63.6 21.6 78.8 39.2
    + KITTI box 34.4 (-2.5) 57.8 (-5.8) 20.7 (-0.9) 72.9 (-5.9) 38.5 (-0.7)
    + KITTI + depth 35.1 (+0.7) 57.3 (-0.5) 21.6 (+0.9) 72.8 (-0.1) 40.2 (+1.7)

    Key observations:

    1. Fine-grained representations: Joint training with segmentation masks improves HOTA from 36.936.9 to 39.239.2 (+2.3+2.3). Human pose keypoints yield an even larger improvement, lifting HOTA to 41.041.0 (+4.1+4.1) and AssA to 25.625.6, as keypoint representations remain robust during severe body articulation and occlusion where bounding boxes fail.
    2. Depth cues: Joint training on KITTI causes domain degradation (due to KITTI's vehicle dominance), but incorporating depth directly into association improves IDF1 (+1.7+1.7) and HOTA (+0.7+0.7) relative to the joint-training box baseline.
  8. Knowl 8 — Limitations of the DanceTrack Benchmark

    limitation

    The authors highlight two main limitations of this work:

    1. Lack of a dedicated tracker: The work introduces the benchmark and empirical diagnoses showing why appearance- and linear-motion-based trackers fail, but does not present a novel multi-object tracking architecture specifically designed to solve the non-linear uniform-appearance tracking problem.
    2. Coarse annotation modalities: Annotations in DanceTrack are currently limited to 2D bounding boxes and identity labels. Dense human pose keypoints and instance segmentation masks, which the ablation studies prove to be beneficial for complex body deformations and occlusions, are not natively annotated in the dataset.

Coverage note — None was omitted; all contributed dataset definitions, diagnostic metrics, oracle experiments, benchmark comparisons, association/motion ablations, multi-modality analyses, and stated limitations are fully represented.

References

  1. 1.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 941–951, 2019.
  2. 2.Keni Bernardin and Rainer Stiefelhagen. Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008:1–10, 2008.
  3. 3.Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple online and realtime tracking. In IEEE international conference on image processing, pages 3464–3468, 2016.
  4. 4.Jinkun Cao, Xin Wang, Trevor Darrell, and Fisher Yu. Instance-aware predictive navigation in multi-agent environments. In IEEE International Conference on Robotics and Automation, pages 5096–5102, 2021.
  5. 5.Mohamed Chaabane, Peter Zhang, Ross Beveridge, and Stephen O’Hara. Deft: Detection embeddings for tracking. arXiv preprint arXiv:2102.02267, 2021.
  6. 6.Tatjana Chavdarova, Pierre Baque, Stéphane Bouquet, Andrii Maksai, Cijo Jose, Timur Bagautdinov, Louis Lettry, Pascal Fua, Luc Van Gool, and François Fleuret. Wildtrack: A multi-camera hd dataset for dense unscripted pedestrian detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5030–5039, 2018.
  7. 7.Long Chen, Haizhou Ai, Zijie Zhuang, and Chong Shang. Real-time multiple people tracking with deeply learned candidate selection and person re-identification. In IEEE international conference on multimedia and expo, pages 1–6, 2018.
  8. 8.Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan. Tao: A large-scale benchmark for tracking any object. In European Conference on Computer Vision, pages 436–454. Springer, 2020.
  9. 9.Patrick Dendorfer, Hamid Rezatofighi, Anton Milan, Javen Shi, Daniel Cremers, Ian Reid, Stefan Roth, Konrad Schindler, and Laura Leal-Taixe. Mot20: A benchmark for multi object tracking in crowded scenes. arXiv preprint arXiv:2003.09003, 2020.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  11. 11.James Ferryman and Ali Shahrokni. Pets2009: Dataset and challenge. In IEEE international workshop on performance evaluation of tracking and surveillance, pages 1–6, 2009.
  12. 12.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021.
  13. 13.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3354–3361, 2012.
  14. 14.Fredrik Gustafsson. Particle filter theory and practice with positioning applications. IEEE Aerospace and Electronic Systems Magazine, 25(7):53–82, 2010.
  15. 15.Rooji Jinan and Tara Raveendran. Particle filters for multiple target tracking. Procedia Technology, 24:980–987, 2016.
  16. 16.Rudolph Emil Kalman. A new approach to linear filtering and prediction problems. 1960.
  17. 17.Laura Leal-Taixe, Anton Milan, Ian Reid, Stefan Roth, and Konrad Schindler. Motchallenge 2015: Towards a benchmark for multi-target tracking. arXiv preprint arXiv:1504.01942, 2015.
  18. 18.Yiyi Liao, Jun Xie, and Andreas Geiger. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. arXiv preprint arXiv:2109.13410, 2021.
  19. 19.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755. Springer, 2014.
  20. 20.Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixe, and Bastian Leibe. Hota: A higher order metric for evaluating multi-object tracking. International Journal of Computer Vision, 129(2):548–578, 2021.
  21. 21.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. Trackformer: Multi-object tracking with transformers. arXiv preprint arXiv:2101.02702, 2021.
  22. 22.Anton Milan, Laura Leal-Taixe, Ian Reid, Stefan Roth, and Konrad Schindler. Mot16: A benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831, 2016.
  23. 23.Jiangmiao Pang, Linlu Qiu, Xia Li, Haofeng Chen, Qi Li, Trevor Darrell, and Fisher Yu. Quasi-dense similarity learning for multiple object tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 164–173, 2021.
  24. 24.Ziqiang Pei. Deepsort pytorch. https://github.com/ZQPei/deep_sort_pytorch, 2019.
  25. 25.Akshay Rangesh and Mohan Manubhai Trivedi. No blind spots: Full-surround multi-object tracking for autonomous vehicles using cameras and lidars. IEEE Transactions on Intelligent Vehicles, 4(4):588–599, 2019.
  26. 26.Ergys Ristani, Francesco Solera, Roger Zou, Rita Cucchiara, and Carlo Tomasi. Performance measures and a data set for multi-target, multi-camera tracking. In European Conference on Computer Vision, pages 17–35. Springer, 2016.
  27. 27.Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and Ping Luo. Transtrack: Multiple-object tracking with transformer. arXiv preprint arXiv:2012.15460, 2020.
  28. 28.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2446–2454, 2020.
  29. 29.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  30. 30.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  31. 31.Zhongdao Wang, Liang Zheng, Yixuan Liu, Yali Li, and Shengjin Wang. Towards real-time multi-object tracking. In European Conference on Computer Vision, pages 107–122. Springer, 2020.
  32. 32.Nicolai Wojke, Alex Bewley, and Dietrich Paulus. Simple online and realtime tracking with a deep association metric. In IEEE international conference on image processing, pages 3645–3649, 2017.
  33. 33.Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang, Ming Yang, and Junsong Yuan. Track to detect and segment: An online multi-object tracker. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12352–12361, 2021.
  34. 34.Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas Huang. Youtube-vos: A large-scale video object segmentation benchmark. arXiv preprint arXiv:1809.03327, 2018.
  35. 35.Alper Yilmaz, Omar Javed, and Mubarak Shah. Object tracking: A survey. Acm computing surveys (CSUR), 38(4):13–es, 2006.
  36. 36.Fisher Yu, Wenqi Xian, Yingying Chen, Fangchen Liu, Mike Liao, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving video database with scalable annotation tooling. arXiv preprint arXiv:1805.04687, 2(5):6, 2018.
  37. 37.Fangao Zeng, Bin Dong, Tiancai Wang, Cheng Chen, Xiangyu Zhang, and Yichen Wei. Motr: End-to-end multiple-object tracking with transformer. arXiv preprint arXiv:2105.03247, 2021.
  38. 38.Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. Bytetrack: Multi-object tracking by associating every detection box. arXiv preprint arXiv:2110.06864, 2021.
  39. 39.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and re-identification in multiple object tracking. International Journal of Computer Vision, 129(11):3069–3087, 2021.
  40. 40.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Tracking objects as points. In European Conference on Computer Vision, pages 474–490. Springer, 2020.
  41. 41.Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. arXiv preprint arXiv:1904.07850, 2019.

Citation

MLA
Sun, P., et al. “DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion”. arXiv, 2021, http://arxiv.org/abs/2111.14690v3.
APA
Sun, P., Cao, J., Jiang, Y., Yuan, Z., Bai, S., Kitani, K., & Luo, P. (2021). DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion. arXiv. http://arxiv.org/abs/2111.14690v3
Chicago
Sun, P., J. Cao, Y. Jiang, et al. 2021. “DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion”. arXiv. http://arxiv.org/abs/2111.14690v3.
Harvard
Sun, P. et al. (2021) “DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.14690v3.
Vancouver
1. Sun P, Cao J, Jiang Y, Yuan Z, Bai S, Kitani K, Luo P (2021) DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion. arXiv

BibTeX

@article{sun2021dancetrack,
  title = {DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion},
  author = {Sun, Peize and Cao, Jinkun and Jiang, Yi and Yuan, Zehuan and Bai, Song and Kitani, Kris and Luo, Ping},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.14690v3},
  eprint = {2111.14690}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE