Object Tracking Benchmark
Yi WuJongwoo LimMing-Hsuan Yang
Establishes a standardized visual tracking evaluation platform featuring 50 fully annotated video sequences categorized by 11 challenge attributes, an integrated code library of 29 algorithms, and systematic spatial and temporal perturbation metrics to assess tracking performance.
Object tracking remains a core challenge in computer vision for applications such as surveillance and medical imaging, yet inconsistent datasets, varying initial conditions, and non-uniform code have made it difficult to compare algorithms reliably or identify what drives robust performance.
This paper set out to create a standardized benchmark that would allow large-scale, fair evaluation of recent online single-target trackers and reveal which design choices matter most under realistic conditions.
The authors assembled a library of 29 publicly available trackers with uniform input and output formats, collected and fully annotated 50 video sequences with ground-truth bounding boxes plus 11 attributes that commonly affect tracking, and ran extensive tests using precision plots at a 20-pixel threshold and success plots measured by area under the curve. They evaluated each tracker more than 660,000 times by applying one-pass evaluation plus two new robustness protocols that perturb the starting frame and the initial bounding-box location or scale.
The evaluation shows that SCM, Struck, and ASLA rank at the top overall, with SCM leading in one-pass tests and Struck proving more stable when initialization varies. Local sparse representations outperform holistic sparse templates on occlusion and deformation; trackers that explicitly exploit background context or use structured output learning handle partial occlusions better; and dense-sampling methods with discriminative models cope more effectively with fast motion than particle-filter approaches. Performance drops noticeably when the initial box is enlarged by 20 percent or shifted, confirming that most trackers remain sensitive to scale and position errors introduced by detectors. Finally, the results indicate that motion or dynamic models receive far less attention than representation and search mechanisms, even though they strongly influence both accuracy and efficiency on fast or abrupt motion.
These findings imply that future trackers can gain the most by combining local appearance models, implicit or explicit background modeling, and stronger motion prediction, rather than refining any single component in isolation. The benchmark itself supplies a reusable platform that removes much of the ambiguity that has hindered progress.
The authors recommend extending both the dataset and the code library with additional sequences and trackers, and they note that improving dynamic models and testing under detector-driven initialization would yield the next practical advances. The work is limited to online single-target tracking with publicly released code and to the 50 sequences chosen; results could shift with different attribute distributions or parameter settings outside the defaults supplied by each method. The scale of the experiments and the consistency of the evaluation protocols give high confidence in the relative rankings and component analysis within the stated scope.
- Paper: The Pascal Visual Object Classes Challenge: A Retrospective, M. Everingham et al. (2014). Reading the PASCAL VOC retrospective first provides the foundational benchmark context and evaluation standards necessary for understanding large-scale object detection challenges.
- Paper: ImageNet Large Scale Visual Recognition Challenge, Olga Russakovsky et al. (2014). Understanding the ImageNet Large Scale Visual Recognition Challenge offers crucial background on the evolution of standardized visual benchmarking that preceded contemporary tracking and detection libraries.
- Paper: Simple online and realtime tracking, Alex Bewley et al. (2016). This work extends the benchmark's evaluation framework by integrating fast detection into online multi-object tracking pipelines.
- Paper: Simple online and realtime tracking with a deep association metric, Nicolai Wojke et al. (2017). This paper builds directly upon the benchmark's tracking evaluations by incorporating deep appearance descriptors to resolve identity switches.
