GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
Lianghua HuangXin ZhaoKaiqi Huang
Introduces GOT-10k, a large-scale generic object tracking benchmark spanning over 560 object classes structured via WordNet, featuring a zero-overlap evaluation protocol to measure how well deep trackers generalize to unseen objects in the wild.
Generic visual object tracking—locating an arbitrary moving object across a video sequence without prior category knowledge—is a foundational technology for surveillance, autonomous robotics, biology, and augmented reality. Despite rapid advances driven by deep learning, progress has been constrained by existing benchmark datasets that contain narrow object distributions and evaluate models primarily on the same object classes used during training. This overlap introduces significant evaluation bias and masks whether tracking algorithms can reliably generalize to novel, real-world targets.
The article aims to resolve these limitations by introducing GOT-10k, a large-scale, high-diversity benchmark designed to systematically evaluate the generalization ability of generic object trackers in unconstrained environments. The researchers set out to demonstrate how dataset scale, semantic diversity, and unseen-class evaluation protocols impact model performance.
To build GOT-10k, the authors used the lexical database WordNet to guide an unbiased semantic selection of 563 moving object classes across five major categories and 87 distinct motion types. The resulting dataset comprises over 10,000 video segments containing more than 1.5 million manually verified bounding box annotations, supplemented with fine-grained labels for object visibility ratios and frame absences. The authors established a strict zero-overlap protocol between the 9,335 training videos and the 420 test videos, ensuring that test classes are completely unseen during training. Using this platform, the study systematically retrained and benchmarked 39 baseline tracking algorithms under unified conditions and introduced class-balanced evaluation metrics to prevent dominant classes from distorting rankings.
The benchmarking revealed several key findings regarding real-world tracking performance and dataset design. First, tracking arbitrary objects in the wild remains largely unsolved; the highest-performing baseline achieved a mean average overlap score of only 46.0%, and baseline accuracy degraded sharply during severe occlusion, fast motion, deformation, and low object resolution. Second, evaluating models on unseen object classes resulted in a measurable performance drop across all deep tracking architectures, confirming that traditional overlapping benchmarks overestimate tracker capability. Third, increasing training data scale and semantic diversity produced dramatic improvements—up to roughly a 15% gain—for fully trainable deep architectures, whereas smaller architectures initialized from pre-trained weights quickly plateaued. Finally, experimental stability analysis demonstrated that a curated test set of 420 videos across 84 unseen classes provides highly stable performance rankings without requiring costly, repetitive evaluations.
These findings indicate that real-world deployment risks for vision systems are higher than previously suggested by legacy benchmarks. Systems optimized on narrow or overlapping datasets are likely to underperform when encountering unfamiliar targets in operational environments. The divergence in rankings between GOT-10k and older benchmarks confirms that high performance on small, familiar datasets does not equate to robust real-world generalization. For practitioners, investing in architectures capable of learning from large, diverse motion and object pools is critical for long-term tracking reliability.
Based on the evidence, organizations developing visual tracking solutions should adopt strict unseen-class protocols for validation and prioritize training data diversity over simple sequence repetition. Tracking models designed for operational deployment should incorporate explicit mechanisms to handle occlusions and scale variations, such as learned memory networks and bounding box regression modules. Research and development teams should utilize the publicly available evaluation server, standardized toolkits, and private test set annotations provided by the GOT-10k platform to benchmark future tracker variants without parameter over-tuning.
The conclusions are supported with high confidence due to rigorous, multi-stage quality control procedures and consistent empirical rankings across dozens of baseline algorithms. However, users should note that the dataset naturally exhibits an imbalanced, long-tailed distribution across classes, reflective of real-world video availability. Tracking speeds on GOT-10k are also lower than on older benchmarks due to higher native video resolutions, which stakeholders must account for when sizing target deployment hardware.
- Paper: Object Tracking Benchmark, Yi Wu et al. (2015). This seminal work establishes the standardized evaluation protocols, precision/success metrics, and tracking benchmarks upon which GOT-10k builds and expands.
- Paper: Fully-Convolutional Siamese Networks for Object Tracking, Luca Bertinetto et al. (2016). It introduced Siamese deep trackers trained offline on large-scale video datasets, defining the modern deep visual tracking paradigm evaluated and advanced by GOT-10k.
- Paper: ECO: Efficient Convolution Operators for Tracking, Martin Danelljan et al. (2017). It presents ECO, one of the primary state-of-the-art baseline trackers rigorously evaluated and benchmarked on the GOT-10k dataset.
- Paper: ImageNet: A large-scale hierarchical image database, Jia Deng et al. (2009). It pioneered the methodology of structuring large-scale computer vision datasets around the semantic taxonomy of WordNet, a direct design basis for GOT-10k's class population.
- Paper: High-Speed Tracking with Kernelized Correlation Filters, João F. Henriques et al. (2014). It establishes the Kernelized Correlation Filter framework that underlies classical and hybrid discriminative trackers compared in the GOT-10k benchmark.
- Paper: Learning Multi-domain Convolutional Neural Networks for Visual Tracking, Hyeonseob Nam et al. (2016). It introduced multi-domain convolutional networks (MDNet) for generic visual tracking, representing a core deep tracking baseline analyzed in GOT-10k.
- Paper: Tracking-Learning-Detection, Zdenek Kalal et al. (2012). It formalized the tracking-learning-detection paradigm for generic long-term tracking of arbitrary objects, providing foundational concepts for generic tracking in the wild.
- Paper: SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks, Bo Li et al. (2018). It leverages modern large-scale tracking benchmarks like GOT-10k to effectively train and evaluate deep Siamese trackers with deep backbone architectures.
