GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild

Lianghua HuangXin ZhaoKaiqi Huang

article2018TPAMI1,882 citations

Introduces GOT-10k, a large-scale generic object tracking benchmark spanning over 560 object classes structured via WordNet, featuring a zero-overlap evaluation protocol to measure how well deep trackers generalize to unseen objects in the wild.

Listen

Generic visual object tracking—locating an arbitrary moving object across a video sequence without prior category knowledge—is a foundational technology for surveillance, autonomous robotics, biology, and augmented reality. Despite rapid advances driven by deep learning, progress has been constrained by existing benchmark datasets that contain narrow object distributions and evaluate models primarily on the same object classes used during training. This overlap introduces significant evaluation bias and masks whether tracking algorithms can reliably generalize to novel, real-world targets.

The article aims to resolve these limitations by introducing GOT-10k, a large-scale, high-diversity benchmark designed to systematically evaluate the generalization ability of generic object trackers in unconstrained environments. The researchers set out to demonstrate how dataset scale, semantic diversity, and unseen-class evaluation protocols impact model performance.

To build GOT-10k, the authors used the lexical database WordNet to guide an unbiased semantic selection of 563 moving object classes across five major categories and 87 distinct motion types. The resulting dataset comprises over 10,000 video segments containing more than 1.5 million manually verified bounding box annotations, supplemented with fine-grained labels for object visibility ratios and frame absences. The authors established a strict zero-overlap protocol between the 9,335 training videos and the 420 test videos, ensuring that test classes are completely unseen during training. Using this platform, the study systematically retrained and benchmarked 39 baseline tracking algorithms under unified conditions and introduced class-balanced evaluation metrics to prevent dominant classes from distorting rankings.

The benchmarking revealed several key findings regarding real-world tracking performance and dataset design. First, tracking arbitrary objects in the wild remains largely unsolved; the highest-performing baseline achieved a mean average overlap score of only 46.0%, and baseline accuracy degraded sharply during severe occlusion, fast motion, deformation, and low object resolution. Second, evaluating models on unseen object classes resulted in a measurable performance drop across all deep tracking architectures, confirming that traditional overlapping benchmarks overestimate tracker capability. Third, increasing training data scale and semantic diversity produced dramatic improvements—up to roughly a 15% gain—for fully trainable deep architectures, whereas smaller architectures initialized from pre-trained weights quickly plateaued. Finally, experimental stability analysis demonstrated that a curated test set of 420 videos across 84 unseen classes provides highly stable performance rankings without requiring costly, repetitive evaluations.

These findings indicate that real-world deployment risks for vision systems are higher than previously suggested by legacy benchmarks. Systems optimized on narrow or overlapping datasets are likely to underperform when encountering unfamiliar targets in operational environments. The divergence in rankings between GOT-10k and older benchmarks confirms that high performance on small, familiar datasets does not equate to robust real-world generalization. For practitioners, investing in architectures capable of learning from large, diverse motion and object pools is critical for long-term tracking reliability.

Based on the evidence, organizations developing visual tracking solutions should adopt strict unseen-class protocols for validation and prioritize training data diversity over simple sequence repetition. Tracking models designed for operational deployment should incorporate explicit mechanisms to handle occlusions and scale variations, such as learned memory networks and bounding box regression modules. Research and development teams should utilize the publicly available evaluation server, standardized toolkits, and private test set annotations provided by the GOT-10k platform to benchmark future tracker variants without parameter over-tuning.

The conclusions are supported with high confidence due to rigorous, multi-stage quality control procedures and consistent empirical rankings across dozens of baseline algorithms. However, users should note that the dataset naturally exhibits an imbalanced, long-tailed distribution across classes, reflective of real-world video availability. Tracking speeds on GOT-10k are also lower than on older benchmarks due to higher native video resolutions, which stakeholders must account for when sizing target deployment hardware.

  • Paper: Object Tracking Benchmark, Yi Wu et al. (2015). This seminal work establishes the standardized evaluation protocols, precision/success metrics, and tracking benchmarks upon which GOT-10k builds and expands.
  • Paper: Fully-Convolutional Siamese Networks for Object Tracking, Luca Bertinetto et al. (2016). It introduced Siamese deep trackers trained offline on large-scale video datasets, defining the modern deep visual tracking paradigm evaluated and advanced by GOT-10k.
  • Paper: ECO: Efficient Convolution Operators for Tracking, Martin Danelljan et al. (2017). It presents ECO, one of the primary state-of-the-art baseline trackers rigorously evaluated and benchmarked on the GOT-10k dataset.
  • Paper: ImageNet: A large-scale hierarchical image database, Jia Deng et al. (2009). It pioneered the methodology of structuring large-scale computer vision datasets around the semantic taxonomy of WordNet, a direct design basis for GOT-10k's class population.
  • Paper: High-Speed Tracking with Kernelized Correlation Filters, João F. Henriques et al. (2014). It establishes the Kernelized Correlation Filter framework that underlies classical and hybrid discriminative trackers compared in the GOT-10k benchmark.
  • Paper: Learning Multi-domain Convolutional Neural Networks for Visual Tracking, Hyeonseob Nam et al. (2016). It introduced multi-domain convolutional networks (MDNet) for generic visual tracking, representing a core deep tracking baseline analyzed in GOT-10k.
  • Paper: Tracking-Learning-Detection, Zdenek Kalal et al. (2012). It formalized the tracking-learning-detection paradigm for generic long-term tracking of arbitrary objects, providing foundational concepts for generic tracking in the wild.
Cover for GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild

Abstract

We introduce here a large tracking database that offers an unprecedentedly wide coverage of common moving objects in the wild, called GOT-10k. Specifically, GOT-10k is built upon the backbone of WordNet structure and it populates the majority of over 560 classes of moving objects and 87 motion patterns, magnitudes wider than the most recent similar-scale counterparts. The contributions of this paper are summarized in the following: (1) GOT-10k offers over 10,000 video segments with more than 1.5 million manually labeled bounding boxes, enabling unified training and stable evaluation of deep trackers. (2) GOT-10k is by far the first video trajectory dataset that uses the semantic hierarchy of WordNet to guide class population. (3) For the first time, GOT-10k introduces the one-shot protocol for tracker evaluation, where the training and test classes are zero-overlapped. The protocol avoids biased evaluation results towards familiar objects and it promotes generalization in tracker development. (4) We conduct extensive tracking experiments with 39 typical tracking algorithms on GOT-10k and analyze their results in this paper. (5) Finally, we develop a comprehensive platform for the tracking community that offers full-featured evaluation toolkits, an online evaluation server, and a responsive leaderboard. The annotations of GOT-10k's test data are kept private to avoid tuning parameters on it. The database, toolkits, evaluation server and baseline results are available at this http URL.

Table of Contents

  • I Introduction
  • II Related Work
  • II-A Evaluation Datasets for Tracking
  • II-B Training Datasets for Tracking
  • III Construction of GOT-10k
  • III-A Collection of Videos
  • III-B Annotation of Trajectories
  • III-C Dataset Splitting
  • IV Experiments
  • IV-A Baseline Models
  • IV-B Evaluation Methodology
  • IV-C Overall Performance
  • IV-D Evaluation by Challenges
  • IV-E Evaluation by Object and Motion Classes
  • IV-F Impact of Training Data
  • V Conclusion
  • References

Knowls

  1. Knowl 1 — GOT-10k Dataset Specification and WordNet Semantic Hierarchy

    definition

    GOT-10k is a large-scale visual tracking dataset for generic, short-term single-object tracking. It contains over 10,000 video segments (9,335 training, 180 validation, 420 test) and more than 1.5 million manually annotated bounding boxes sampled at 10 frames per second. Video durations range from 0.4 seconds to 148 seconds, with an average length of 15 seconds.

    The dataset's object and motion taxonomies are structured using the semantic hierarchy of WordNet:

    • Moving Object Classes (563 total): Populated by expanding five WordNet noun subtrees: animal (3.8k targets, 382 sub-classes, 360k bounding boxes), vehicle (2.4k targets, 154 sub-classes, 380k bounding boxes), person (2.5k targets, 1 sub-class, 487k bounding boxes), passive motion object (0.5k targets, 11 sub-classes, 70k bounding boxes), and object part (1.0k targets, 15 sub-classes, 214k bounding boxes).
    • Motion Classes (87 total): Populated by expanding WordNet verb subtrees for locomotion, action, and sport (with six domain-specific exceptions: camera motion, dragon-lion dance, parkour, powered cartwheeling, taking off, and uneven bars).

    In addition to axis-aligned bounding boxes, each frame is annotated with target absence flags and continuous visible ratios v∈[0,1]v \in [0, 1] quantized into seven 15% intervals (0%–15%0\%\text{--}15\%, 15%–30%15\%\text{--}30\%, ..., 90%–100%90\%\text{--}100\%). In the dataset, 15.43% of frames exhibit target occlusion or truncation (v<0.90v < 0.90), 1.86% exhibit heavy occlusion/truncation (v<0.45v < 0.45), and 0.43% exhibit full absence or out-of-view status.

  2. Knowl 2 — One-Shot Evaluation Protocol and Split Design in GOT-10k

    experimental setup

    To evaluate the generic, class-agnostic generalization capability of tracking algorithms on previously unseen object categories, GOT-10k implements a one-shot evaluation protocol where the object classes between the training and test partitions are strictly non-overlapping (zero overlap).

    The dataset partition parameters are:

    • Training Set: 9,335 video segments covering 480 object classes and 69 motion classes.
    • Validation Set: 180 video segments sampled uniformly across 150 object classes, covering 15 motion classes.
    • Test Set: 420 video segments covering 84 object classes and 31 motion classes.

    The person class is an exception present in both training and test sets due to its visual diversity and practical importance; to maintain a domain gap, the motion classes of person sequences in the training and test sets are strictly disjoint. To avoid high-frequency classes dominating test performance, the maximum number of sequences per object class in the test set is restricted to 8 (at most 1.9% of the test set). For stochastic trackers, evaluations are averaged over 3 independent runs.

  3. Knowl 3 — Class-Balanced Tracking Evaluation Metrics: mAO and mSR

    equation

    To prevent tracking benchmark scores from being biased toward frequent object classes that have more evaluation sequences, the mean average overlap (mAO) and mean success rate (mSR) are defined by first calculating the metric within each class and then averaging over all classes.

    For a benchmark with CC distinct object classes, where ScS_c denotes the set of test sequences belonging to class cc, and ∣Sc∣|S_c| is the number of sequences in that class, the class-balanced mean average overlap (mAO) is defined as:

    mAO=1C∑c=1C(1∣Sc∣∑i∈ScAOi)\text{mAO} = \frac{1}{C} \sum_{c=1}^C \left( \frac{1}{|S_c|} \sum_{i \in S_c} \text{AO}_i \right)

    where AOi\text{AO}_i is the average bounding-box Intersection-over-Union (IoU) overlap across all frames of sequence ii.

    Similarly, the class-balanced mean success rate at overlap threshold τ\tau (typically τ∈{0.50,0.75}\tau \in \{0.50, 0.75\}), denoted mSRτ\text{mSR}_\tau, is calculated as:

    mSRτ=1C∑c=1C(1∣Sc∣∑i∈ScSRi,τ)\text{mSR}_\tau = \frac{1}{C} \sum_{c=1}^C \left( \frac{1}{|S_c|} \sum_{i \in S_c} \text{SR}_{i, \tau} \right)

    where SRi,τ\text{SR}_{i, \tau} represents the proportion of frames in sequence ii whose overlap with the groundtruth bounding box exceeds the threshold τ\tau.

  4. Knowl 4 — Continuous Difficulty Indicators for Visual Tracking Challenge Attribution

    equation

    To quantify tracking difficulty objectively and reproducibly without subjective per-frame categorical labels, six continuous indicators are computed directly from trajectory bounding-box annotations:

    1. Occlusion/Truncation Degree: degree=1−vi\text{degree} = 1 - v_i where vi∈[0,1]v_i \in [0, 1] is the annotated visible ratio of the target at frame ii.

    2. Scale Variation: SVi=max⁡(sisi−T,si−Tsi)\text{SV}_i = \max\left(\frac{s_i}{s_{i-T}}, \frac{s_{i-T}}{s_i}\right) where si=wihis_i = \sqrt{w_i h_i} is the geometric mean size of the bounding box (width wiw_i, height hih_i) at frame ii, assessed over a time window T=5T = 5 frames.

    3. Aspect Ratio Variation: ARVi=max⁡(riri−T,ri−Tri)\text{ARV}_i = \max\left(\frac{r_i}{r_{i-T}}, \frac{r_{i-T}}{r_i}\right) where ri=hi/wir_i = h_i / w_i is the target aspect ratio at frame ii and T=5T = 5.

    4. Fast Motion (Relative Speed): di=∥pi−pi−1∥2sisi−1d_i = \frac{\|p_i - p_{i-1}\|_2}{\sqrt{s_i s_{i-1}}} where pi=(xi,yi)p_i = (x_i, y_i) is the target bounding-box center location in pixel coordinates and si=wihis_i = \sqrt{w_i h_i}.

    5. Illumination Variation: ui=∥ci−ci−1∥1u_i = \|c_i - c_{i-1}\|_1 where ci∈[0,1]3c_i \in [0, 1]^3 is the average RGB color vector of the groundtruth target region at frame ii.

    6. Low Resolution (Small Object Indicator): LRi=sismedianfor si≤smedian\text{LR}_i = \frac{s_i}{s^{\text{median}}} \quad \text{for } s_i \le s^{\text{median}} where smedians^{\text{median}} is the median object size wh\sqrt{wh} across all test frames.

  5. Knowl 5 — Overall Tracking Benchmark Results on GOT-10k

    data/table

    Evaluation of 39 baseline tracking algorithms on the 420 test sequences of GOT-10k. All deep trackers were retrained exclusively on the GOT-10k training set (with ImageNet initialization permitted for backbones). Trackers are evaluated on class-balanced mean average overlap (mAO), mean success rate at 0.50 overlap (mSR50\text{mSR}_{50}), mean success rate at 0.75 overlap (mSR75\text{mSR}_{75}), and tracking speed (fps).

    Tracker mAO mSR50\text{mSR}_{50} mSR75\text{mSR}_{75} Speed (fps) CF DL
    MemTracker 0.460 0.524 0.193 0.353@GPU ✓
    DeepSTRCF 0.449 0.481 0.169 1.07@GPU ✓ ✓
    SASiamP 0.445 0.491 0.165 25.4@GPU ✓
    SASiamR 0.443 0.492 0.160 5.13@GPU ✓
    SiamFCv2 0.434 0.481 0.190 19.6@GPU ✓
    GOTURN 0.418 0.475 0.163 70.1@GPU ✓
    DSiam 0.417 0.461 0.149 3.78@GPU ✓
    SiamFCIncep22 0.411 0.456 0.154 12@GPU ✓
    DAT 0.411 0.432 0.145 0.0774@GPU ✓
    CCOT 0.406 0.415 0.161 0.57@CPU ✓
    MetaSDNet 0.404 0.423 0.156 0.526@GPU ✓
    RT-MDNet 0.404 0.424 0.147 7.85@GPU ✓
    SiamFCNext22 0.398 0.430 0.151 12.2@GPU ✓
    ECO 0.395 0.407 0.170 2.21@CPU ✓
    SiamFC 0.392 0.426 0.135 32.6@GPU ✓
    STRCF 0.377 0.387 0.151 3.06@CPU ✓
    ECOhc 0.363 0.359 0.154 34.7@CPU ✓
    MDNet 0.352 0.367 0.137 0.951@GPU ✓
    BACF 0.346 0.361 0.149 3.22@CPU ✓
    KCF 0.279 0.263 0.099 11.5@CPU ✓
    IVT 0.171 0.114 0.031 47.3@CPU

    The highest mAO achieved among all evaluated baseline trackers is 0.460 (MemTracker), demonstrating that unconstrained generic object tracking under zero-overlap conditions remains challenging. Furthermore, relative rankings on GOT-10k differ substantially from older benchmarks such as OTB2015 (e.g., ECO achieves the highest performance on OTB2015 with 0.677 AUC, but drops to 14th on GOT-10k with 0.395 mAO).

  6. Knowl 6 — Generalization Gap between Seen and Unseen Object Classes

    data/table

    To quantify the performance drop caused by evaluating on unseen classes compared to familiar (seen) classes under the one-shot protocol, four deep trackers were trained on 4,000 randomly sampled videos from GOT-10k and evaluated on 240 videos of seen classes versus 240 videos of unseen classes across three random sampling trials.

    Tracker Seen Test Classes (mAO) Unseen Test Classes (mAO)
    MemTracker 0.595 / 0.589 / 0.612 0.588 / 0.567 / 0.598
    SiamFCv2 0.584 / 0.593 / 0.571 0.525 / 0.569 / 0.550
    GOTURN 0.489 / 0.491 / 0.485 0.440 / 0.462 / 0.457
    MDNet 0.443 / 0.437 / 0.450 0.439 / 0.441 / 0.449

    All deep trackers experience an observable performance degradation (ranging from 0.1% to 5.9% in mAO) when tested on novel object categories compared to seen categories, highlighting that deep models partially overfit to familiar semantic classes.

  7. Knowl 7 — Impact of Training Data Scale, Object Diversity, and Motion Diversity on Deep Trackers

    empirical result

    Ablation experiments on four representative deep trackers (MemTracker, SiamFCv2, GOTURN, and MDNet) reveal differing dependencies on training data attributes:

    • Training Data Scale (15 to 9,335 videos): Trackers trained entirely from scratch (MemTracker, SiamFCv2) display continuous performance increases as data scale grows, with mAO improving steadily up to 9,335 videos without saturation. Trackers initialized with ImageNet pretrained backbones and frozen early layers (MDNet, GOTURN) saturate very early (MDNet saturates at approximately 15 training videos; GOTURN's performance plateaus and slightly declines at high scale).
    • Object Diversity (5 to 405 classes, fixed at 2,000 videos): MemTracker and SiamFCv2 achieve steep mAO gains (nearly 15% absolute improvement) when object diversity increases from 5 to 405 classes without converging at 405, demonstrating that diverse object categories are crucial for learning generalizable representation. MDNet is largely insensitive to object class diversity.
    • Motion Diversity (4 to 64 motion classes, fixed at 2,000 videos): SiamFCv2 reaches peak performance at 16 motion classes, while MemTracker continues to improve up to 64 motion classes due to its dynamic memory update mechanism learning richer temporal transitions.
  8. Knowl 8 — Impact of Training and Test Class Imbalance on Tracker Performance and Ranking Stability

    data/table

    To evaluate the effect of the long-tailed class distribution in GOT-10k, four deep trackers were trained on 2,000 sequences sampled across 100 object classes under two regimes: a balanced regime (exactly 20 sequences per class) and an imbalanced regime (retaining the natural long-tailed distribution). Evaluation stability was also tested by comparing the standard deviation of 25 baseline tracker ranks on balanced versus imbalanced test sets.

    Tracker Balanced Training Data (mAO) Imbalanced Training Data (mAO)
    MemTracker 0.431 0.412
    SiamFCv2 0.399 0.407
    GOTURN 0.408 0.421
    MDNet 0.357 0.340
    Metric Balanced Test Data (std) Imbalanced Test Data (std) Abs. Difference between Ranks
    mAO 1.102 1.213 1.259
    mSR 1.114 1.241 1.302

    Class imbalance during training produces only marginal performance differences (±1%–2%\pm 1\%\text{--}2\% mAO), indicating that overall visual diversity (motion, attributes, scenes) dominates over strict uniform class balance. On the evaluation side, class-balanced metrics (mAO and mSR) ensure ranking stability across balanced and slightly imbalanced test sets with an average rank difference of only 1.25 to 1.30.

  9. Knowl 9 — Cross-Dataset Training Generalization: GOT-10k versus ImageNet-VID

    data/table

    Deep trackers trained on GOT-10k versus the widely used ImageNet-VID dataset were evaluated across both OTB2015 and GOT-10k test sets to test cross-dataset transferability.

    Tracker VID →\rightarrow OTB GOT-10k →\rightarrow OTB VID →\rightarrow GOT-10k GOT-10k →\rightarrow GOT-10k
    MemTracker 0.625 0.636 0.447 0.460
    SiamFCv2 0.613 0.621 0.423 0.434
    GOTURN 0.427 0.413 0.396 0.418
    MDNet 0.673 0.637 0.341 0.352

    When evaluated on GOT-10k, all four trackers show substantial absolute gains (1.1% to 2.2% mAO) when trained on GOT-10k compared to ImageNet-VID. Because ImageNet-VID contains only 30 object categories and frequent shot transitions/incomplete objects, models trained on it lack the visual and motion diversity required to generalize to broad unconstrained tracking scenarios.

Coverage note — None was omitted; all key contributions—dataset taxonomy/scale, one-shot protocol, class-balanced evaluation metrics, continuous challenge formulations, full benchmark comparisons, seen vs. unseen generalization, data diversity ablations, class balance studies, and cross-dataset transfer—are represented.

References

  1. 1.G. A. Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995.
  2. 2.Kristan, Matej, et al., The visual object tracking challenge VOT2019. http://www.votchallenge.net/vot2019/.
  3. 3.Kristan, Matej, et al., The sixth visual object tracking vot2018 challenge results. European Conference on Computer Vision Workshops (ECCVW). Springer, Cham, 2018, pp. 3-53.
  4. 4.M. Kristan et al., The Visual Object Tracking VOT2017 Challenge Results, In IEEE International Conference on Computer Vision Workshops (ICCVW), Venice, 2017, pp. 1949-1972.
  5. 5.M. Kristan et al., The Visual Object Tracking VOT2016 Challenge Results. In European Conference on Computer Vision Workshops (ECCVW), Amsterdam, 2016. (Lecture Notes in Computer Science vol 9914. Springer, Cham)
  6. 6.M. Kristan et al., The Visual Object Tracking VOT2015 Challenge Results, IEEE International Conference on Computer Vision Workshop (ICCVW), Santiago, 2015, pp. 564-586.
  7. 7.M. Kristan et al., The Visual Object Tracking VOT2014 Challenge Results. In European Conference on Computer Vision Workshops (ECCVW), 2014.. (Lecture Notes in Computer Science, vol 8926. Springer, Cham)
  8. 8.M. Kristan et al., ”The Visual Object Tracking VOT2013 Challenge Results,” 2013 IEEE International Conference on Computer Vision Workshops, Sydney, NSW, 2013, pp. 98-111.
  9. 9.Y. Wu, J. Lim, and M.-H. Yang. Online object tracking: A benchmark. In IEEE Conference on Computer vision and pattern recognition, pages 2411–2418. IEEE, 2013.
  10. 10.ehovin L, Kristan M, Leonardis A. Is my new tracker really better than yours? IEEE Winter Conference on Applications of Computer Vision. IEEE, 2014: 540-547.
  11. 11.ehovin L, Leonardis A, Kristan M. Visual object tracking performance measures revisited. IEEE Transactions on Image Processing, 2016, 25(3): 1261-1274.
  12. 12.Y. Wu, J. Lim, and M.-H. Yang. Object tracking benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(9):1834–1848, 2015.
  13. 13.J. Valmadre, L. Bertinetto, J. F. Henriques, R. Tao, A. Vedaldi, A. Smeulders, P. Torr, and E. Gavves. Long-term tracking in the wild: A benchmark. arXiv preprint arXiv:1803.09502, 2018.
  14. 14.H. K. Galoogahi, A. Fagg, C. Huang, D. Ramanan, and S. Lucey. Need for speed: A benchmark for higher frame rate object tracking. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 1134–1143. IEEE, 2017.
  15. 15.M. Mueller, N. Smith, and B. Ghanem. A benchmark and simulator for uav tracking. In European Conference on Computer Vision, pages 445–461. Springer, 2016.
  16. 16.P. Liang, E. Blasch, and H. Ling. Encoding color information for visual tracking: Algorithms and benchmark. IEEE Transactions on Image Processing, 24(12):5630–5644, 2015.
  17. 17.A. Li, M. Lin, Y. Wu, M.-H. Yang, and S. Yan. Nus-pro: A new visual tracking challenge. IEEE transactions on pattern analysis and machine intelligence, 38(2):335–349, 2016.
  18. 18.S. Song and J. Xiao. Tracking revisited using rgbd camera: Unified benchmark and baselines. In IEEE International Conference on Computer Vision, pages 233–240. IEEE, 2013.
  19. 19.Mller, M., Bibi, A., Giancola, S., Al-Subaihi, S. and Ghanem, B. TrackingNet: A Large-Scale Dataset and Benchmark for Object Tracking in the Wild. European Conference on Computer Vision, 2018.
  20. 20.Heng Fan, Liting Lin, Fan Yang, et al. LaSOT: A High-quality Benchmark for Large-scale Single Object Tracking. IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  21. 21.A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, and M. Shah. Visual tracking: An experimental survey. IEEE transactions on pattern analysis and machine intelligence, 36(7):1442–1468, 2014.
  22. 22.Liu, Qiao, and Zhenyu He. ”PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark.” arXiv preprint arXiv:1801.05944 (2018).
  23. 23.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  24. 24.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255. IEEE, 2009.
  25. 25.H. Su, J. Deng, and L. Fei-Fei. Crowdsourcing annotations for visual object detection. In Workshops at the Twenty-Sixth AAAI Conference on Artificial Intelligence, volume 1, 2012.
  26. 26.E. Real, J. Shlens, S. Mazzocchi, X. Pan, and V. Vanhoucke. Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. IEEE conference on Computer Vision and Pattern Recognition, 2017.
  27. 27.Kuznetsova A, Rom H, Alldrin N, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. arXiv preprint arXiv:1811.00982, 2018.
  28. 28.Milan, A., Leal-Taix, L., Reid, I., Roth, S., and Schindler, K.. MOT16: A benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831, 2016.
  29. 29.Leal-Taix, L., Milan, A., Reid, I., Roth, S., and Schindler, K.. Motchallenge 2015: Towards a benchmark for multi-target tracking. arXiv preprint arXiv:1504.01942, 2015.
  30. 30.Geiger, A., Lenz, P., and Urtasun, R.. Are we ready for autonomous driving? the kitti vision benchmark suite. In IEEE Conference on Computer Vision and Pattern Recognition, 2012 (pp. 3354-3361).
  31. 31.X. Wang, K. He, and A. Gupta. Transitive invariance for selfsupervised visual representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1329–1338, 2017.
  32. 32.X. Wang and A. Gupta. Unsupervised learning of visual representations using videos. In Proceedings of the 2015 IEEE International Conference on Computer Vision, pages 2794–2802. IEEE Computer Society, 2015.
  33. 33.Kang, K., Ouyang, W., Li, H., and Wang, X. Object detection from video tubelets with convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition pp. 817-825.
  34. 34.Kang, K., Li, H., Yan, J., Zeng, X., Yang, B., Xiao, T., et al. T-CNN: Tubelets with convolutional neural networks for object detection from videos. In IEEE Transactions on Circuits and Systems for Video Technology. 2017.
  35. 35.Cheng, J., Tsai, Y. H., Hung, W. C., Wang, S., and Yang, M. H. Fast and Accurate Online Video Object Segmentation via Tracking Parts. In arXiv preprint arXiv:1806.02323. 2018.
  36. 36.Chu, Q., Ouyang, W., Li, H., Wang, X., Liu, B., and Yu, N. Online multi-object tracking using cnn-based single object tracker with spatial-temporal attention mechanism. In IEEE International Conference on Computer Vision. pp. 4846-4855. 2017.
  37. 37.Wang, Yu-Xiong, Deva Ramanan, and Martial Hebert. Learning to model the tail. Advances in Neural Information Processing Systems. 2017.
  38. 38.Cui Y, Jia M, Lin T Y, et al. Class-Balanced Loss Based on Effective Number of Samples[J]. arXiv preprint arXiv:1901.05555, 2019.
  39. 39.Bengio, S. Sharing representations for long tail computer vision problems. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction (pp. 1-1). ACM, 2015.
  40. 40.Huang, C., Li, Y., Change Loy, C., & Tang, X. Learning deep representation for imbalanced classification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5375-5384), 2016.
  41. 41.D. Held, D. Guillory, B. Rebsamen, S. Thrun, and S. Savarese. A probabilistic framework for real-time 3d segmentation using spatial, temporal, and semantic cues. in Proceedings of Robotics: Science and Systems, 2016.
  42. 42.G. Zhang and P. A. Vela. Good features to track for visual slam. in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015.
  43. 43.K. Bozek, L. Hebert, A. S. Mikheyev, and G. J. Stephens. Towards dense object tracking in a 2d honeybee hive. in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
  44. 44.Roman Pflugfelder. An In-Depth Analysis of Visual Tracking with Siamese Neural Networks ArXiv 2018.
  45. 45.L. Bertinetto, J. Valmadre, S. Golodetz, O. Miksik, and P. H. Torr. Staple: Complementary learners for real-time tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1401–1409, 2016.
  46. 46.L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. Torr. Fully-convolutional siamese networks for object tracking. In European conference on computer vision, pages 850–865. Springer, 2016.
  47. 47.M. Danelljan, G. Bhat, F. Khan, and M. Felsberg. Eco: Efficient convolution operators for tracking. In 30th IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6931–6939. IEEE, 2017.
  48. 48.M. Danelljan, G. Hager, F. Khan, and M. Felsberg. Accurate scale ¨ estimation for robust visual tracking. In British Machine Vision Conference, Nottingham, September 1-5. BMVA Press, 2014.
  49. 49.Zhang, Jianming, Shugao Ma, and Stan Sclaroff. MEEM: robust tracking via multiple experts using entropy minimization. In European Conference on Computer Vision. Cham, 2014.
  50. 50.M. Danelljan, G. Hager, F. S. Khan, and M. Felsberg. Discrimina- ¨ tive scale space tracking. IEEE transactions on pattern analysis and machine intelligence, 39(8):1561–1575, 2017.
  51. 51.M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg. Learning spatially regularized correlation filters for visual tracking. In Proceedings of the IEEE International Conference on Computer Vision, pages 4310–4318, 2015.
  52. 52.M. Danelljan, G. Hager, F. Shahbaz Khan, and M. Felsberg. Adaptive decontamination of the training set: A unified formulation for discriminative visual tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1430–1438, 2016.
  53. 53.M. Danelljan, A. Robinson, F. S. Khan, and M. Felsberg. Beyond correlation filters: Learning continuous convolution operators for visual tracking. In European Conference on Computer Vision, pages 472–488. Springer, 2016.
  54. 54.D. Held, S. Thrun, and S. Savarese. Learning to track at 100 fps with deep regression networks. In European Conference on Computer Vision, pages 749–765. Springer, 2016.
  55. 55.J. F. Henriques, R. Caseiro, P. Martins, and J. Batista. Exploiting the circulant structure of tracking-by-detection with kernels. In European conference on computer vision, pages 702–715. Springer, 2012.
  56. 56.J. F. Henriques, R. Caseiro, P. Martins, and J. Batista. High-speed tracking with kernelized correlation filters. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(3):583–596, 2015.
  57. 57.Bazzani, L., Larochelle, H., Murino, V., Ting, J. A., & Freitas, N. D. Learning attentional policies for tracking and recognition in video with deep networks. In Proceedings of the 28th International Conference on Machine Learning, pp. 937-944.
  58. 58.Y. Li and J. Zhu. A scale adaptive kernel correlation filter tracker with feature integration. In European Conference on Computer Vision, pages 254–265. Springer, 2014.
  59. 59.C. Ma, J.-B. Huang, X. Yang, and M.-H. Yang. Hierarchical convolutional features for visual tracking. In Proceedings of the IEEE International Conference on Computer Vision, pages 3074–3082, 2015.
  60. 60.H. Nam and B. Han. Learning multi-domain convolutional neural networks for visual tracking. In Computer Vision and Pattern Recognition (CVPR), 2016 IEEE Conference on, pages 4293–4302. IEEE, 2016.
  61. 61.E. Park and A. C. Berg. Meta-tracker: Fast and robust online adaptation for visual object trackers. arXiv preprint arXiv:1801.03049, 2018.
  62. 62.H. Possegger, T. Mauthner, and H. Bischof. In defense of colorbased model-free tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2113–2120, 2015.
  63. 63.J. Valmadre, L. Bertinetto, J. Henriques, A. Vedaldi, and P. H. Torr. End-to-end representation learning for correlation filter based tracking. In Computer Vision and Pattern Recognition (CVPR), 2017 IEEE Conference on, pages 5000–5008. IEEE, 2017.
  64. 64.N. Wang, J. Shi, D.-Y. Yeung, and J. Jia. Understanding and diagnosing visual tracking systems. In Computer Vision (ICCV), 2015 IEEE International Conference on, pages 3101–3109. IEEE, 2015.
  65. 65.Li B, Yan J, Wu W, et al. High Performance Visual Tracking With Siamese Region Proposal Network, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018: 8971-8980.
  66. 66.Ross, D. A., Lim, J., Lin, R. S., and Yang, M. H. (2008). Incremental learning for robust visual tracking. International journal of computer vision, 77(1-3), 125-141.
  67. 67.Bao, C., Wu, Y., Ling, H., and Ji, H. (2012, June). Real time robust l1 tracker using accelerated proximal gradient approach. In IEEE Conference on Computer Vision and Pattern Recognition, 2012, pp. 1830-1837.
  68. 68.Bolme, D. S., Beveridge, J. R., Draper, B. A., and Lui, Y. M. (2010, June). Visual object tracking using adaptive correlation filters. In IEEE Conference on Computer Vision and Pattern Recognition, 2010, pp. 2544-2550.
  69. 69.Galoogahi, H. K., Fagg, A., and Lucey, S.. Learning backgroundaware correlation filters for visual tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA (pp. 21-26), 2017.
  70. 70.Ren, S., He, K., Girshick, R. and Sun, J.. Faster r-cnn: Towards realtime object detection with region proposal networks. In Advances in neural information processing systems (pp. 91-99), 2015.
  71. 71.Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks[C]//Advances in neural information processing systems. 2012: 1097-1105.
  72. 72.Guo, Q., Feng, W., Zhou, C., Huang, R., Wan, L., and Wang, S.. Learning dynamic siamese network for visual object tracking. In The IEEE International Conference on Computer Vision. (Oct 2017).
  73. 73.Bhat, G., Johnander, J., Danelljan, M., Khan, F. S., and Felsberg, M. Unveiling the Power of Deep Tracking. arXiv preprint arXiv:1804.06833, 2018.
  74. 74.Shi, J, and C. Tomasi. Good features to track. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pp. 593-600, 2002.
  75. 75.Wang, Q., Gao, J., Xing, J., Zhang, M., & Hu, W. Dcfnet: Discriminant correlation filters network for visual tracking. arXiv preprint arXiv:1704.04057, 2017.
  76. 76.He A, Luo C, Tian X, et al. A twofold siamese network for real-time object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4834-4843, 2018.
  77. 77.Jung, I., Son, J., Baek, M., & Han, B. Real-time mdnet. In Proceedings of the European Conference on Computer Vision (ECCV), pp. 83-98, 2018.
  78. 78.Zhipeng, Z., Houwen, P., & Qiang, W. (2019). Deeper and Wider Siamese Networks for Real-Time Visual Tracking. In Proceedings of the Computer Vision and Pattern Recognition, 2019.
  79. 79.Pu, S., Song, Y., Ma, C., Zhang, H., & Yang, M. H. Deep attentive tracking via reciprocative learning. In Advances in Neural Information Processing Systems, pp. 1931-1941, 2018.
  80. 80.Yang, Tianyu, and Antoni B. Chan. Learning dynamic memory networks for object tracking. In Proceedings of the European Conference on Computer Vision, pp. 152-167, 2018.
  81. 81.Li, F., Tian, C., Zuo, W., Zhang, L., & Yang, M. H. Learning spatial-temporal regularized correlation filters for visual tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4904-4913, 2018.
  82. 82.Li, Y., Zhu, J., Hoi, S. C., Song, W., Wang, Z., & Liu, H. Robust Estimation of Similarity Transformation for Visual Object Tracking. The Conference on Association for the Advancement of Artificial Intelligence (AAAI), 2019.

Citation

MLA
Huang, L., et al. “GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 5, 2021, pp. 1562–77, https://doi.org/10.1109/TPAMI.2019.2957464.
APA
Huang, L., Zhao, X., & Huang, K. (2021). GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5), 1562–1577. https://doi.org/10.1109/TPAMI.2019.2957464
Chicago
Huang, L., X. Zhao, and K. Huang. 2021. “GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild”. IEEE Transactions on Pattern Analysis and Machine Intelligence 43 (5): 1562–77. https://doi.org/10.1109/TPAMI.2019.2957464.
Harvard
Huang, L., Zhao, X. and Huang, K. (2021) “GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(5), pp. 1562–1577. Available at: https://doi.org/10.1109/TPAMI.2019.2957464.
Vancouver
1. Huang L, Zhao X, Huang K (2021) GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild. IEEE Transactions on Pattern Analysis and Machine Intelligence 43:1562–1577

BibTeX

@article{Huang_2021, title={GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild}, volume={43}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2019.2957464}, DOI={10.1109/tpami.2019.2957464}, number={5}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Huang, Lianghua and Zhao, Xin and Huang, Kaiqi}, year={2021}, month=May, pages={1562–1577} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF