keyword
long-term tracking
Long-term tracking is a computer vision task in which an algorithm continuously follows and localizes a designated target object across an extended video sequence or stream. Unlike short-term tracking, which assumes the target remains visible within a localized neighborhood from frame to frame, long-term tracking is designed to operate over sustained durations where the target may undergo severe appearance changes, experience full occlusions, or completely leave and re-enter the camera field of view. To operate reliably in these environments, long-term tracking systems must assess target presence, explicitly declare target absence when the object is lost or out of frame, and incorporate wide-range redetection and re-identification mechanisms to automatically resume tracking once the target reappears.
6 items

Generalized Relation Modeling for Transformer Tracking
Shenyuan Gao, Chunluan Zhou, Jun Zhang
Why you should read this
Proposes a generalized relation modeling framework with adaptive token division that dynamically selects relevant search tokens to interact with template tokens, preventing target-background confusion and setting a new state of the art in real-time Transformer visual tracking.
Compared with previous two-stream trackers, the recent one-stream tracking pipeline, which allows earlier interaction between the template and search region, has achieved a remarkable performance gain. However, existing one-stream trackers always let the template interact with all parts inside the search region throughout all the encoder layers. This could potentially lead to target-background confusion when the extracted feature representations are not sufficiently discriminative. To alleviate this issue, we propose a generalized relation modeling method based on adaptive token division. The proposed method is a generalized formulation of attention-based relation modeling for Transformer tracking, which inherits the merits of both previous two-stream and one-stream pipelines whilst enabling more flexible relation modeling by selecting appropriate search tokens to interact with template tokens. An attention masking strategy and the Gumbel-Softmax technique are introduced to facilitate the parallel computation and end-to-end learning of the token division module. Extensive experiments show that our method is superior to the two-stream and one-stream pipelines and achieves state-of-the-art performance on six challenging benchmarks with a real-time running speed. Code and models are publicly available at https://github.com/Little-Podi/GRM.
Added
2026-10-05

LaSOT: A High-Quality Benchmark for Large-Scale Single Object Tracking
Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, Haibin Ling
Why you should read this
Introduces a large-scale visual tracking benchmark of 1,400 long-term sequences with over 3.5 million densely annotated frames and natural language descriptions, providing a standard platform to train deep trackers and evaluate algorithms against real-world challenges.
In this paper, we present LaSOT, a high-quality benchmark for Large-scale Single Object Tracking. LaSOT consists of 1,400 sequences with more than 3.5M frames in total. Each frame in these sequences is carefully and manually annotated with a bounding box, making LaSOT the largest, to the best of our knowledge, densely annotated tracking benchmark. The average video length of LaSOT is more than 2,500 frames, and each sequence comprises various challenges deriving from the wild where target objects may disappear and re-appear again in the view. By releasing LaSOT, we expect to provide the community with a large-scale dedicated benchmark with high quality for both the training of deep trackers and the veritable evaluation of tracking algorithms. Moreover, considering the close connections of visual appearance and natural language, we enrich LaSOT by providing additional language specification, aiming at encouraging the exploration of natural linguistic feature for tracking. A thorough experimental evaluation of 35 tracking algorithms on LaSOT is presented with detailed analysis, and the results demonstrate that there is still a big room for improvements.
Added
2026-09-24

GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
Lianghua Huang, Xin Zhao, Kaiqi Huang
Why you should read this
Introduces GOT-10k, a large-scale generic object tracking benchmark spanning over 560 object classes structured via WordNet, featuring a zero-overlap evaluation protocol to measure how well deep trackers generalize to unseen objects in the wild.
We introduce here a large tracking database that offers an unprecedentedly wide coverage of common moving objects in the wild, called GOT-10k. Specifically, GOT-10k is built upon the backbone of WordNet structure and it populates the majority of over 560 classes of moving objects and 87 motion patterns, magnitudes wider than the most recent similar-scale counterparts. The contributions of this paper are summarized in the following: (1) GOT-10k offers over 10,000 video segments with more than 1.5 million manually labeled bounding boxes, enabling unified training and stable evaluation of deep trackers. (2) GOT-10k is by far the first video trajectory dataset that uses the semantic hierarchy of WordNet to guide class population. (3) For the first time, GOT-10k introduces the one-shot protocol for tracker evaluation, where the training and test classes are zero-overlapped. The protocol avoids biased evaluation results towards familiar objects and it promotes generalization in tracker development. (4) We conduct extensive tracking experiments with 39 typical tracking algorithms on GOT-10k and analyze their results in this paper. (5) Finally, we develop a comprehensive platform for the tracking community that offers full-featured evaluation toolkits, an online evaluation server, and a responsive leaderboard. The annotations of GOT-10k's test data are kept private to avoid tuning parameters on it. The database, toolkits, evaluation server and baseline results are available at this http URL.
Added
2026-09-18

A Benchmark and Simulator for UAV Tracking
Matthias Mueller, Neil G. Smith, Bernard Ghanem
Why you should read this
Presents a dedicated low-altitude aerial tracking benchmark of 123 fully annotated HD video sequences alongside an Unreal Engine-based simulator for real-time evaluation and synthetic data generation.
In this paper, we propose a new aerial video dataset and benchmark for low altitude UAV target tracking, as well as, a photo-realistic UAV simulator that can be coupled with tracking methods. Our benchmark provides the first evaluation of many state-of-the-art and popular trackers on 123 new and fully annotated HD video sequences captured from a low-altitude aerial perspective. Among the compared trackers, we determine which ones are the most suitable for UAV tracking both in terms of tracking accuracy and run-time. The simulator can be used to evaluate tracking algorithms in real-time scenarios before they are deployed on a UAV “in the field”, as well as, generate synthetic but photo-realistic tracking datasets with automatic ground truth annotations to easily extend existing real-world datasets. Both the benchmark and simulator are made publicly available to the vision community on our website to further research in the area of object tracking from UAVs. (https://ivul.kaust.edu.sa/Pages/pub-benchmark-simulator-uav.aspx.).
Added
2026-09-18

SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks
Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, Junjie Yan
Why you should read this
Proposes a spatial-aware sampling strategy and multi-layer feature aggregation to overcome translation invariance limitations, successfully enabling deep ResNet backbones in Siamese visual tracking to achieve state-of-the-art accuracy across major benchmarks.
Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem.
Added
2026-09-16

Tracking-Learning-Detection
Zdenek Kalal, K. Mikolajczyk, Jiri Matas
Why you should read this
Proposes a real-time framework that decomposes long-term visual tracking into tracking, detection, and self-correcting P-N learning to sustain tracking of unknown objects through full occlusions, camera disappearances, and severe appearance changes.
This paper investigates long-term tracking of unknown objects in a video stream. The object is defined by its location and extent in a single frame. In every frame that follows, the task is to determine the object's location and extent or indicate that the object is not present. We propose a novel tracking framework (TLD) that explicitly decomposes the long-term tracking task into tracking, learning and detection. The tracker follows the object from frame to frame. The detector localizes all appearances that have been observed so far and corrects the tracker if necessary. The learning estimates detector's errors and updates it to avoid these errors in the future. We study how to identify detector's errors and learn from them. We develop a novel learning method (P-N learning) which estimates the errors by a pair of “experts”: (i) P-expert estimates missed detections, and (ii) N-expert estimates false alarms. The learning process is modeled as a discrete dynamical system and the conditions under which the learning guarantees improvement are found. We describe our real-time implementation of the TLD framework and the P-N learning. We carry out an extensive quantitative evaluation which shows a significant improvement over state-of-the-art approaches.
Source
http://vision.stanford.edu/teaching/cs231b_spring1415/papers/PAMI2011_KalalMikolajczykMatas.pdfAdded
2026-09-12
