Tracking Objects as Points
Xingyi ZhouVladlen KoltunPhilipp Krähenbühl
Introduces CenterTrack, a simple real-time framework that tracks objects as points across consecutive frames, eliminating complex temporal association pipelines while achieving state-of-the-art accuracy on 2D and 3D tracking benchmarks.
Real-time visual tracking of multiple objects is a vital capability for autonomous driving, robotics, and intelligent surveillance. Traditional systems often suffer from high computational overhead and latency because they rely on complex, multi-stage detection pipelines and computationally expensive pairwise feature comparisons across frames.
The article evaluates a streamlined approach, CenterTrack, which tracks objects as single points using a single-pass deep network. It demonstrates the framework's tracking accuracy, processing speed, and architectural modifications across standard two-dimensional and three-dimensional multi-object tracking benchmarks.
The evaluation relies on quantitative experiments across established vision benchmarks, including MOT16, MOT17, KITTI, and nuScenes, alongside pretraining on the CrowdHuman dataset. The methodology applies a single neural network pass per frame combined with a fast greedy matching algorithm based on center-point distances, supporting both private and public detection modes while avoiding complex combinatorial matching schemes.
The findings show that the proposed method achieves competitive accuracy while operating significantly faster than prior systems. On the MOT16 private benchmark, the model ranked second among published methods with a 69.6% tracking accuracy score (MOTA), running online at 17 frames per second (57 ms runtime), whereas competing approaches required significantly more time per frame plus separate detection overhead. Pretraining on the CrowdHuman dataset proved essential, boosting MOT17 validation accuracy from 60.7% to 66.1% by reducing missed detections. On the KITTI benchmark, the full tracking model achieved an 88.7% accuracy score, matching complex state-of-the-art systems while showing that simple greedy association performed identically to the more computationally demanding Hungarian matching algorithm. Additionally, on three-dimensional detection benchmarks, camera-based regression achieved performance comparable to existing monocular methods, though it trailed specialized LiDAR-based systems.
These results indicate that multi-object tracking can be executed efficiently without heavy pairwise re-identification networks or slow optimization steps. By unifying detection and point-based tracking into a single online pass, organizations can lower compute hardware costs, decrease processing latency, and simplify deployment in time-critical systems such as automated vehicles.
Engineering teams deploying this framework should leverage large, static human-detection datasets for pretraining to reduce false negatives. Teams should also adopt simpler greedy association logic over complex optimization algorithms and calibrate output and rendering confidence thresholds (such as 0.4 and 0.5) to balance false alarms and missed tracks. For three-dimensional tracking environments where high positional precision is safety-critical, practitioners should note the performance gap between camera-only inputs and LiDAR sensors before relying purely on monocular data.
- Paper: Objects as Points, Xingyi Zhou et al. (2019). CenterTrack directly builds upon CenterNet's anchor-free formulation of detecting objects as center points and extends it across time.
- Paper: Simple online and realtime tracking, Alex Bewley et al. (2016). SORT establishes the foundational baseline for real-time online multi-object tracking using simple spatial and motion cues that CenterTrack seeks to simplify and surpass.
- Paper: Simple online and realtime tracking with a deep association metric, Nicolai Wojke et al. (2017). DeepSORT introduces deep visual appearance features into online multi-object tracking, representing the standard tracking-by-detection paradigm that CenterTrack re-evaluates and streamlines.
- Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). This paper presents the MOT16 benchmark suite and standardized evaluation protocols that are central to measuring CenterTrack's multi-object tracking performance.
- Paper: CornerNet: Detecting Objects as Paired Keypoints, Hei Law et al. (2018). CornerNet pioneered keypoint-based object detection, providing the conceptual foundation for representing objects as points without anchor boxes.
- Paper: Center-based 3D Object Detection and Tracking, Tianwei Yin et al. (2020). CenterPoint adapts the center-based representation and tracking-as-points philosophy of CenterTrack to point cloud 3D object detection and tracking in LiDAR and driving scenarios.
- Paper: FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking, Yifu Zhang et al. (2020). FairMOT extends point-based tracking architectures by investigating how to extract balanced re-identification and detection features directly at object centers.
- Paper: ByteTrack: Multi-Object Tracking by Associating Every Detection Box, Yifu Zhang et al. (2021). ByteTrack advances simple and effective multi-object association beyond point-based center tracking by utilizing low-confidence detection boxes to prevent track fragmentation.
- Paper: HOTA: A Higher Order Metric for Evaluating Multi-object Tracking, Jonathon Luiten et al. (2020). HOTA provides a higher-order evaluation metric designed to balance detection and association accuracy, directly addressing evaluation challenges seen in online trackers like CenterTrack.
