Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiple object tracking

Multiple object tracking is a computer vision task that involves detecting multiple targets of interest in a video sequence and maintaining their unique identities across consecutive frames to generate continuous trajectories. Typically addressed through paradigms such as tracking by detection or joint detection and embedding, the process relies on core components including object localization, feature extraction or re-identification, and data association. By predicting target locations and linking detections over time despite challenges such as occlusions, interactions among targets, motion blur, and variations in appearance, multiple object tracking provides automated spatial and temporal analysis essential for applications like autonomous driving, video surveillance, and robotics.

2 items

UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu

OrganizationsInstitute of Automation, Chinese Academy of SciencesNanjing University of Posts and TelecommunicationsUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a unified multiple object tracking framework that couples detection, feature embedding, and identity association through an identity-aware feature enhancement module, creating a mutual feedback loop that improves both object localization and association accuracy across frames.

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM.

Added

2026-09-26

FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking

FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking

Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, Wenyu Liu

OrganizationsHuazhong University of Science and TechnologyMicrosoft

Why you should read this

Proposes FairMOT, an anchor-free tracking framework that balances object detection and re-identification within a single network to eliminate task competition and achieve state-of-the-art multi-object tracking performance.

Multi-object tracking (MOT) is an important problem in computer vision which has a wide range of applications. Formulating MOT as multi-task learning of object detection and re-ID in a single network is appealing since it allows joint optimization of the two tasks and enjoys high computation efficiency. However, we find that the two tasks tend to compete with each other which need to be carefully addressed. In particular, previous works usually treat re-ID as a secondary task whose accuracy is heavily affected by the primary detection task. As a result, the network is biased to the primary detection task which is not fair to the re-ID task. To solve the problem, we present a simple yet effective approach termed as FairMOT based on the anchor-free object detection architecture CenterNet. Note that it is not a naive combination of CenterNet and re-ID. Instead, we present a bunch of detailed designs which are critical to achieve good tracking results by thorough empirical studies. The resulting approach achieves high accuracy for both detection and tracking. The approach outperforms the state-of-the-art methods by a large margin on several public datasets. The source code and pre-trained models are released at this https URL.

Added

2026-09-18