UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation

Kefu YiKai LuoXiaolei LuoJiangui HuangHao WuRongdong HuWei Hao

article2024AAAI155 citations

Presents a purely motion-based multi-object tracker that avoids expensive frame-by-frame camera motion compensation by mapping Kalman filter state estimation and probability distributions onto the ground plane, achieving state-of-the-art accuracy at over 1000 FPS on a single CPU.

Listen

Real-time multi-object tracking in video is essential for autonomous systems and intelligent surveillance, yet it struggles when cameras move or jitter. Traditional solutions either extract visual appearance features or calculate camera motion compensation on every single frame. While effective, these techniques impose severe computational overhead, which prevents tracking systems from operating in real-time on edge devices.

The article demonstrates an efficient multi-object tracker, named UCMCTrack, that achieves state-of-the-art accuracy relying strictly on motion cues while using uniform camera motion compensation parameters across an entire video sequence.

To accomplish this, the authors shifted object motion modeling from the 2D image plane to the 3D ground plane using a standard linear camera projection model. Instead of relying on conventional bounding box overlap metrics like Intersection over Union, the method introduces the Mapped Mahalanobis Distance combined with a Correlated Measurement Distribution, which accurately accounts for spatial projection uncertainties. Motion noise from camera movements is absorbed into a standard Kalman filter's process noise covariance matrix. The system was evaluated across established tracking benchmarks, including MOT17, MOT20, DanceTrack, and KITTI, using standard detections from baseline detectors.

The evaluation produced several significant findings. First, UCMCTrack matches or outperforms complex trackers on MOT17, achieving 65.8 Higher Order Tracking Accuracy (HOTA) and 81.1 identification score (IDF1) when paired with traditional compensation, surpassing leading appearance-and-motion models. Second, on the complex DanceTrack dataset featuring irregular movements, it outperformed existing methods by 2.3 in HOTA, 3.4 in IDF1, and 5.5 in association accuracy. Third, in low frame-rate driving conditions on KITTI, standalone UCMCTrack proved more robust than traditional frame-by-frame compensation, which suffered from parameter estimation inaccuracies. Finally, the tracking association step runs at over 1,000 frames per second on a single CPU given detections, providing orders-of-magnitude computational efficiency improvements.

These findings indicate that complex, resource-intensive appearance networks and frame-by-frame image alignment are not required to handle camera motion. Ground-plane motion modeling successfully absorbs camera movement as system noise. This drastically lowers computing requirements and system latency, enabling high-performance tracking on low-power embedded hardware in robotics and vehicles.

Organizations developing computer vision systems should consider adopting ground-plane motion modeling to replace expensive alignment pipelines. For implementation, teams must tailor process compensation factors based on whether a scene is predominantly static or dynamic. Future development should explore combining this metric with appearance cues or bounding box heights for complex multi-level environments.

Decision-makers should note that the approach relies on the assumption of a single flat ground plane and requires camera calibration parameters. While the tracker is robust to detector noise, large errors in estimated camera tilt and roll can degrade tracking accuracy, warranting adequate camera parameter calibration during deployment.

No sufficiently relevant recommendations were found.

Cover for UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation

Abstract

Multi-object tracking (MOT) in video sequences remains a challenging task, especially in scenarios with significant camera movements. This is because targets can drift considerably on the image plane, leading to erroneous tracking outcomes. Addressing such challenges typically requires supplementary appearance cues or Camera Motion Compensation (CMC). While these strategies are effective, they also introduce a considerable computational burden, posing challenges for real-time MOT. In response to this, we introduce UCMCTrack, a novel motion model-based tracker robust to camera movements. Unlike conventional CMC that computes compensation parameters frame-by-frame, UCMCTrack consistently applies the same compensation parameters throughout a video sequence. It employs a Kalman filter on the ground plane and introduces the Mapped Mahalanobis Distance (MMD) as an alternative to the traditional Intersection over Union (IoU) distance measure. By leveraging projected probability distributions on the ground plane, our approach efficiently captures motion patterns and adeptly manages uncertainties introduced by homography projections. Remarkably, UCMCTrack, relying solely on motion cues, achieves state-of-the-art performance across a variety of challenging datasets, including MOT17, MOT20, DanceTrack and KITTI. More details and code are available at https://github.com/corfyi/UCMCTrack.

Table of Contents

  • Introduction
  • Related Work
  • Distance Measures
  • Motion Models
  • Camera Motion Compensation
  • Method
  • Motion Modeling on Ground Plane
  • Correlated Measurement Distribution
  • Mapped Mahalanobis Distance
  • Process Noise Compensation
  • Experiments
  • Setting
  • Benchmark Evaluation
  • Ablation Studies on UCMC
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Ground-Plane Motion Modeling in UCMCTrack

    model/method

    UCMCTrack models the movement of tracked targets directly on the 2D ground plane rather than on the 2D image plane, enabling the tracker to capture motion dynamics without suffering from bounding box overlap failures caused by camera jitter or high target speed.

    The target state vector at frame kk is defined by:

    x=[x,x˙,y,y˙]T\mathbf{x} = [x, \dot{x}, y, \dot{y}]^T

    where (x,y)(x, y) denotes the 2D ground coordinates and (x˙,y˙)(\dot{x}, \dot{y}) denotes the velocities along the ground plane axes.

    The observation vector z=[x,y]T\mathbf{z} = [x, y]^T is obtained by projecting the midpoint of the bottom edge of the 2D image bounding box (u,v)T(u, v)^T onto the ground plane via a linear camera projection model:

    [uv1]=A1γ[xy1]\begin{bmatrix} u \\ v \\ 1 \end{bmatrix} = \mathbf{A} \frac{1}{\gamma} \begin{bmatrix} x \\ y \\ 1 \end{bmatrix}

    where γ\gamma is a scale factor and A∈R3×3\mathbf{A} \in \mathbb{R}^{3 \times 3} is the transformation matrix formed by multiplying the camera intrinsic and extrinsic parameter matrices.

  2. Knowl 2 — Correlated Measurement Distribution on the Ground Plane

    equation

    While 2D detector errors on the image plane are assumed independent, their projection onto the ground plane introduces directional correlations that rotate the uncertainty ellipses according to the camera viewpoint.

    Let the image-plane detection noise covariance matrix be:

    Rkuv=[(σmwk)200(σmhk)2]\mathbf{R}_k^{uv} = \begin{bmatrix} (\sigma_m w_k)^2 & 0 \\ 0 & (\sigma_m h_k)^2 \end{bmatrix}

    where σm\sigma_m is the detection noise hyperparameter factor, and wkw_k and hkh_k denote the detected bounding box width and height on the image plane at frame kk.

    Given the inverse projection matrix A−1=[aij]3×3\mathbf{A}^{-1} = [a_{ij}]_{3 \times 3} mapping image coordinates (u,v)(u, v) back to ground coordinates (x,y)(x, y) with scale factor γ\gamma, the linear transformation Jacobian matrix C∈R2×2\mathbf{C} \in \mathbb{R}^{2 \times 2} is given by:

    C=[γa11−a31γxγa12−a32γxγa21−a31γyγa22−a32γy]\mathbf{C} = \begin{bmatrix} \gamma a_{11} - a_{31} \gamma x & \gamma a_{12} - a_{32} \gamma x \\ \gamma a_{21} - a_{31} \gamma y & \gamma a_{22} - a_{32} \gamma y \end{bmatrix}

    The projected, non-diagonal measurement error covariance matrix Rk\mathbf{R}_k on the ground plane is computed as:

    Rk=CRkuvCT\mathbf{R}_k = \mathbf{C} \mathbf{R}_k^{uv} \mathbf{C}^T

  3. Knowl 3 — Mapped Mahalanobis Distance for Ground-Plane Data Association

    equation

    Mapped Mahalanobis Distance (MMD) replaces the traditional Intersection over Union (IoU) metric by calculating matching costs in the ground-plane coordinate space while penalizing projection uncertainties.

    For a ground-plane observation z∈R2\mathbf{z} \in \mathbb{R}^2 and a Kalman filter predicted state x^∈R4\mathbf{\hat{x}} \in \mathbb{R}^4 with observation matrix H\mathbf{H}:

    1. Measurement residual: ϵ=z−Hx^\boldsymbol{\epsilon} = \mathbf{z} - \mathbf{H}\mathbf{\hat{x}}

    2. Residual covariance matrix: S=HPHT+Rk\mathbf{S} = \mathbf{H}\mathbf{P}\mathbf{H}^T + \mathbf{R}_k where P∈R4×4\mathbf{P} \in \mathbb{R}^{4 \times 4} is the Kalman filter predicted state error covariance matrix, and Rk∈R2×2\mathbf{R}_k \in \mathbb{R}^{2 \times 2} is the mapped ground-plane measurement noise covariance matrix.

    3. Normalized Mahalanobis Distance cost: D=ϵTS−1ϵ+ln⁡∣S∣D = \boldsymbol{\epsilon}^T \mathbf{S}^{-1} \boldsymbol{\epsilon} + \ln |\mathbf{S}| where ∣S∣|\mathbf{S}| denotes the determinant of S\mathbf{S} and ln⁡\ln is the natural logarithm. The term ln⁡∣S∣\ln |\mathbf{S}| accounts for the uncertainty and geometric distortion of the projected distribution.

  4. Knowl 4 — Process Noise Compensation for Uniform Camera Motion

    model/method

    Rather than calculating frame-by-frame image registration transformations (such as ECC or optical flow matching), Uniform Camera Motion Compensation (UCMC) models camera jitter and motion as acceleration noise within a constant velocity (CV) Kalman filter on the ground plane. The same process noise parameters are applied across the entire video sequence.

    The target position and velocity perturbations due to camera motion noise over frame interval Δt\Delta t are modeled as Δx=12σ(Δt)2\Delta x = \frac{1}{2} \sigma (\Delta t)^2 and Δv=σΔt\Delta v = \sigma \Delta t. In matrix form, the noise transition matrix is:

    G=[(Δt)220Δt00(Δt)220Δt]\mathbf{G} = \begin{bmatrix} \frac{(\Delta t)^2}{2} & 0 \\ \Delta t & 0 \\ 0 & \frac{(\Delta t)^2}{2} \\ 0 & \Delta t \end{bmatrix}

    The process noise covariance matrix Qk\mathbf{Q}_k of the ground-plane Kalman filter is computed as:

    Qk=Gdiag⁡(σx,σy)GT\mathbf{Q}_k = \mathbf{G} \operatorname{diag}(\sigma_x, \sigma_y) \mathbf{G}^T

    where σx\sigma_x and σy\sigma_y are constant process compensation factors representing motion noise along the xx and yy ground axes (corresponding to camera tilt and rotation/pan movements, respectively).

  5. Knowl 5 — Benchmark Results on MOT17 and MOT20 Test Sets

    data/table

    The tracking performance of UCMCTrack (motion-only) and UCMCTrack+ (UCMCTrack combined with frame-by-frame ECC-based camera motion compensation) was evaluated on the test sets of MOT17 and MOT20 against state-of-the-art motion-based and appearance-and-motion-based trackers using ByteTrack YOLOX detections.

    Tracker MOT17 MOT20
    HOTA↑\uparrow IDF1↑\uparrow MOTA↑\uparrow AssA↑\uparrow HOTA↑\uparrow IDF1↑\uparrow MOTA↑\uparrow AssA↑\uparrow
    appearance motion:
    FCG 62.6 77.7 76.7 63.4 57.3 69.7 68.0 58.1
    Quo Vadis 63.1 77.7 80.3 62.1 61.5 75.7 77.8 59.9
    GHOST 62.8 77.1 78.7 - 61.2 75.2 73.7 -
    Bot-SORT 65.0 80.2 80.5 65.5 63.3 77.5 77.8 62.9
    StrongSORT 64.4 79.5 79.6 64.4 62.6 77.0 73.8 64.0
    Deep OCSORT 64.9 80.6 79.4 65.9 63.9 79.2 75.6 65.7
    motion only:
    ByteTrack 63.1 77.3 80.3 62.0 61.3 75.2 77.8 59.6
    C-BIoU 64.1 79.7 81.1 63.7 - - - -
    MotionTrack 65.1 80.1 81.1 65.1 62.8 76.5 78.0 61.8
    SparseTrack 65.1 80.1 81.0 65.1 63.4 77.3 78.2 62.8
    OCSORT 63.2 77.5 78.0 63.4 62.4 76.3 75.7 62.5
    UCMCTrack (Ours) 64.3 79.0 79.0 64.6 62.8 77.4 75.5 63.5
    UCMCTrack+ (Ours) 65.8 81.1 80.5 66.6 62.8 77.4 75.7 63.4

    UCMCTrack+ achieves 65.8 HOTA and 81.1 IDF1 on MOT17, outperforming all previous motion-only and appearance-augmented trackers.

  6. Knowl 6 — Benchmark Results on DanceTrack and KITTI Test Sets

    data/table

    UCMCTrack was evaluated on the DanceTrack test set (characterized by uniform appearance, irregular motion, and minor camera jitter) and the KITTI test set (characterized by high-speed vehicle motion, 10 FPS rate, and intense ego-motion).

    DanceTrack-test
    Tracker HOTA↑\uparrow IDF1↑\uparrow MOTA↑\uparrow AssA↑\uparrow
    appearance motion:
    FCG 48.7 46.5 89.9 29.9
    GHOST 56.7 57.7 91.3 39.8
    StrongSORT 55.6 55.2 91.1 38.6
    Deep OCSORT 61.3 61.5 92.3 45.8
    motion only:
    ByteTrack 47.3 52.5 89.5 31.4
    C-BIoU 60.6 61.6 91.6 45.4
    MotionTrack 58.2 58.6 91.3 41.7
    SparseTrack 55.5 58.3 91.3 39.1
    OCSORT 55.1 54.9 92.2 40.4
    UCMCTrack (Ours) 63.4 65.0 88.8 51.1
    UCMCTrack+ (Ours) 63.6 65.0 88.9 51.3
    KITTI-test
    Tracker Car Pedestrian
    HOTA↑\uparrow MOTA↑\uparrow AssA↑\uparrow HOTA↑\uparrow MOTA↑\uparrow AssA↑\uparrow
    appearance motion:
    QD-3DT 72.8 85.9 72.2 41.1 51.8 38.8
    TuSimple 71.6 86.3 71.1 45.9 57.6 47.6
    StrongSORT 77.8 90.4 78.2 54.5 67.4 57.3
    motion only:
    CenterTrack 73.0 88.8 71.2 40.4 53.8 36.9
    TrackMPNN 72.3 87.3 70.6 39.4 52.1 35.5
    OCSORT 76.5 90.3 76.4 54.7 65.1 59.1
    UCMCTrack (Ours) 77.1 90.4 77.2 55.2 67.4 58.0
    UCMCTrack+ (Ours) 74.2 90.2 71.7 54.3 67.2 56.3

    On DanceTrack, UCMCTrack outperforms previous state-of-the-art methods by 2.1+ in HOTA and 3.4+ in IDF1. On KITTI, standalone UCMCTrack outperforms UCMCTrack+ due to inaccuracy in frame-by-frame CMC parameters during intense camera motion.

  7. Knowl 7 — Ablation Study of UCMCTrack Core Components

    data/table

    An ablation study evaluated the progressive addition of Mapped Mahalanobis Distance (MMD), Correlated Measurement Distribution (CMD), and Process Noise Compensation (PNC) relative to the baseline ByteTrack tracker on MOT17 and DanceTrack validation sets.

    Method IoU MMD CMD PNC IDF1↑\uparrow HOTA↑\uparrow
    MOT17 Validation Set
    baseline ✓ - - - 77.10 68.43
    UCMCTrack-v1 - ✓ - - 75.88 68.09
    UCMCTrack-v2 - ✓ ✓ - 79.68 70.44
    UCMCTrack-v3 - ✓ ✓ ✓ 82.20 71.96
    DanceTrack Validation Set
    baseline ✓ - - - 47.27 47.93
    UCMCTrack-v1 - ✓ - - 43.76 46.32
    UCMCTrack-v2 - ✓ ✓ - 53.93 55.06
    UCMCTrack-v3 - ✓ ✓ ✓ 62.64 60.42

    Replacing IoU with unweighted MMD (v1) initially lowers performance because MMD discards bounding box height cues. Adding CMD (v2) to capture directional uncertainty ellipses increases MOT17 HOTA to 70.44 and DanceTrack HOTA to 55.06. Adding PNC (v3) to compensate for camera motion noise achieves the highest scores (71.96 HOTA on MOT17 and 60.42 HOTA on DanceTrack).

  8. Knowl 8 — Ablation on Combining UCMC and Traditional Frame-by-Frame CMC

    data/table

    The interaction between Uniform Camera Motion Compensation (UCMC) and traditional frame-by-frame Camera Motion Compensation (CMC, using ECC image alignment) was ablated on MOT17 and DanceTrack validation sets.

    Method CMC UCMC IDF1↑\uparrow HOTA↑\uparrow
    MOT17 Validation Set
    baseline - - 77.10 68.43
    baseline+CMC ✓ - 81.07 70.97
    UCMCTrack - ✓ 82.20 71.96
    UCMCTrack+ ✓ ✓ 84.05 72.97
    DanceTrack Validation Set
    baseline - - 47.27 47.93
    baseline+CMC ✓ - 46.74 47.55
    UCMCTrack - ✓ 62.64 60.42
    UCMCTrack+ ✓ ✓ 62.52 59.18

    On MOT17, UCMCTrack alone achieves higher tracking accuracy than baseline + CMC (71.96 vs 70.97 HOTA), and combining both (UCMCTrack+) yields 72.97 HOTA. On DanceTrack, where viewpoint changes are minor, frame-by-frame CMC slightly degrades performance due to registration offset errors, whereas standalone UCMC provides a large performance boost (60.42 HOTA vs 47.93 baseline).

  9. Knowl 9 — Hyperparameter Sensitivity and Calibration Error Robustness of UCMCTrack

    empirical result

    Analysis of key parameters and errors in UCMCTrack reveals the following properties:

    1. Process Compensation Factor (σ\sigma): Static and dynamic scenes exhibit distinct responses. For static camera scenes, a smaller compensation factor σ\sigma (ln(σ)≈−1\\ln(\sigma) \approx -1) is optimal, whereas dynamic camera scenes require larger compensation factors (ln⁡(σ)≥3\ln(\sigma) \ge 3) to absorb camera motion noise into the Kalman process model.

    2. Detector Noise Factor (σm\sigma_m): Tracking performance peaks at σm=0.05\sigma_m = 0.05. Performance across σm∈[0.04,0.10]\sigma_m \in [0.04, 0.10] remains stable (HOTA variation <1.0< 1.0), demonstrating low sensitivity to σm\sigma_m.

    3. Extrinsic Camera Calibration Errors: Perturbing camera rotation angles around the XX, YY, and ZZ axes indicates that errors along the YY axis (yaw) have minimal effect on HOTA/IDF1 because yaw does not significantly deform the estimated ground plane. In contrast, rotational errors along the XX (pitch) and ZZ (roll) axes cause significant ground-plane perspective distortion, degrading tracking performance when errors exceed 2∘2^\circ--4∘4^\circ.

    4. Inference Speed: Given bounding box detections, UCMCTrack runs at over 1000 FPS on a single CPU.

  10. Knowl 10 — Ground Plane Assumption Limitation of UCMCTrack

    limitation

    An inherent limitation of UCMCTrack is its assumption that all tracked targets move on a single, planar ground surface. When tracking in non-planar environments (e.g., stairs, ramps, multi-level structures, or uncalibrated uneven terrain), the homography projection from image bounding box coordinates to ground coordinates deforms, leading to localization and state estimation errors.

Coverage note — No substantial contributed material was omitted. All core mathematical formulations, system components, benchmark datasets, ablation experiments, parameter analyses, and stated limitations are covered.

References

  1. 1.Aharon, N.; Orfaig, R.; and Bobrovsky, B.-Z. 2022. BoT-SORT: Robust associations multi-pedestrian tracking. arXiv preprint arXiv:2206.14651.
  2. 2.Babaee, M.; Li, Z.; and Rigoll, G. 2019. A dual CNN–RNN for multiple people tracking. Neurocomputing, 368: 69–83.
  3. 3.Bergmann, P.; Meinhardt, T.; and Leal-Taixe, L. 2019. Tracking Without Bells and Whistles. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  4. 4.Bernardin, K.; and Stiefelhagen, R. 2008. Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008: 1–10.
  5. 5.Bewley, A.; Ge, Z.; Ott, L.; Ramos, F.; and Upcroft, B. 2016. Simple Online and Realtime Tracking. In 2016 IEEE International Conference on Image Processing (ICIP), 3464–3468.
  6. 6.Cao, J.; Pang, J.; Weng, X.; Khirodkar, R.; and Kitani, K. 2023. Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9686–9696.
  7. 7.Cetintas, O.; Braso, G.; and Leal-Taix e, L. 2023. Unifying Short and Long-Term Tracking With Graph Hierarchies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 22877–22887.
  8. 8.Choi, W. 2015. Near-online multi-target tracking with aggregated local flow descriptor. In Proceedings of the IEEE international conference on computer vision, 3029–3037.
  9. 9.Dendorfer, P.; Rezatofighi, H.; Milan, A.; Shi, J.; Cremers, D.; Reid, I.; Roth, S.; Schindler, K.; and Leal-Taixe, L. 2020. Mot20: A benchmark for multi object tracking in crowded scenes. arXiv preprint arXiv:2003.09003.
  10. 10.Dendorfer, P.; Yugay, V.; Osep, A.; and Leal-Taixe, L. 2022. Quo Vadis: Is Trajectory Forecasting the Key Towards Long-Term Multi-Object Tracking? Advances in Neural Information Processing Systems, 35: 15657–15671.
  11. 11.Du, Y.; Wan, J.; Zhao, Y.; Zhang, B.; Tong, Z.; and Dong, J. 2021. GIAOTracker: A Comprehensive Framework for MC-MOT With Global Information and Optimizing Strategies in VisDrone 2021. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, 2809–2819.
  12. 12.Du, Y.; Zhao, Z.; Song, Y.; Zhao, Y.; Su, F.; Gong, T.; and Meng, H. 2023. Strongsort: Make deepsort great again. IEEE Transactions on Multimedia.
  13. 13.Evangelidis, G. D.; and Psarakis, E. Z. 2008a. Parametric image alignment using enhanced correlation coefficient maximization. IEEE transactions on pattern analysis and machine intelligence, 30(10): 1858–1865.
  14. 14.Evangelidis, G. D.; and Psarakis, E. Z. 2008b. Parametric Image Alignment Using Enhanced Correlation Coefficient Maximization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 30(10): 1858–1865.
  15. 15.Feichtenhofer, C.; Pinz, A.; and Zisserman, A. 2017. Detect to track and track to detect. In Proceedings of the IEEE international conference on computer vision, 3038–3046.
  16. 16.Fischler, M. A.; and Bolles, R. C. 1981. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6): 381–395.
  17. 17.Ge, Z.; Liu, S.; Wang, F.; Li, Z.; and Sun, J. 2021. YOLOX: Exceeding YOLO Series in 2021. CoRR, abs/2107.08430.
  18. 18.Geiger, A.; Lenz, P.; Stiller, C.; and Urtasun, R. 2013. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 32(11): 1231–1237.
  19. 19.Girbau, A.; Marques, F.; and Satoh, S. 2022. Multiple Object Tracking from appearance by hierarchically clustering tracklets. arXiv preprint arXiv:2210.03355.
  20. 20.Han, S.; Huang, P.; Wang, H.; Yu, E.; Liu, D.; and Pan, X. 2022. MAT: Motion-aware multi-object tracking. Neurocomputing, 476: 75–86.
  21. 21.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  22. 22.Held, D.; Thrun, S.; and Savarese, S. 2016. Learning to track at 100 fps with deep regression networks. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, 749–765. Springer.
  23. 23.Hu, H.-N.; Yang, Y.-H.; Fischer, T.; Darrell, T.; Yu, F.; and Sun, M. 2022. Monocular quasi-dense 3d object tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2): 1992–2008.
  24. 24.Khurana, T.; Dave, A.; and Ramanan, D. 2021. Detecting Invisible People. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 3174–3184.
  25. 25.Liu, J.; Wang, Z.; and Xu, M. 2020. DeepMTT: A deep learning maneuvering target-tracking algorithm based on bidirectional LSTM network. Information Fusion, 53: 289–304.
  26. 26.Liu, Q.; Chu, Q.; Liu, B.; and Yu, N. 2020. GSM: Graph Similarity Model for Multi-Object Tracking. In IJCAI, 530–536.
  27. 27.Liu, Z.; Wang, X.; Wang, C.; Liu, W.; and Bai, X. 2023. SparseTrack: Multi-Object Tracking by Performing Scene Decomposition based on Pseudo-Depth. arXiv:2306.05238.
  28. 28.Luiten, J.; Osep, A.; Dendorfer, P.; Torr, P.; Geiger, A.; Leal-Taixe, L.; and Leibe, B. 2021. Hota: A higher order metric for evaluating multi-object tracking. International journal of computer vision, 129: 548–578.
  29. 29.Maggiolino, G.; Ahmad, A.; Cao, J.; and Kitani, K. 2023. Deep oc-sort: Multi-pedestrian tracking by adaptive reidentification. arXiv preprint arXiv:2302.11813.
  30. 30.Marinello, N.; Proesmans, M.; and Van Gool, L. 2022. TripletTrack: 3D Object Tracking Using Triplet Embeddings and LSTM. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 4500–4510.
  31. 31.Milan, A.; Leal-Taixe, L.; Reid, I.; Roth, S.; and Schindler, K. 2016. MOT16: A benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831.
  32. 32.Rangesh, A.; Maheshwari, P.; Gebre, M.; Mhatre, S.; Ramezani, V.; and Trivedi, M. M. 2021. TrackMPNN: A message passing graph neural architecture for multi-object tracking. arXiv preprint arXiv:2101.04206.
  33. 33.Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; and Savarese, S. 2019. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 658–666.
  34. 34.Ristani, E.; Solera, F.; Zou, R.; Cucchiara, R.; and Tomasi, C. 2016. Performance measures and a data set for multi-target, multi-camera tracking. In Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part II, 17–35. Springer.
  35. 35.Rublee, E.; Rabaud, V.; Konolige, K.; and Bradski, G. 2011. ORB: An efficient alternative to SIFT or SURF. In 2011 International conference on computer vision, 2564–2571. Ieee.
  36. 36.Seidenschwarz, J.; Braso, G.; Serrano, V. C.; Elezi, I.; and Leal-Taixe, L. 2023. Simple Cues Lead to a Strong Multi-Object Tracker. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13813–13823.
  37. 37.Sun, P.; Cao, J.; Jiang, Y.; Yuan, Z.; Bai, S.; Kitani, K.; and Luo, P. 2022. DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20993–21002.
  38. 38.Tokmakov, P.; Li, J.; Burgard, W.; and Gaidon, A. 2021. Learning To Track With Object Permanence. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10860–10869.
  39. 39.Wojke, N.; Bewley, A.; and Paulus, D. 2017. Simple Online and Realtime Tracking with a Deep Association Metric. In 2017 IEEE International Conference on Image Processing (ICIP), 3645–3649.
  40. 40.Xiao, B.; Wu, H.; and Wei, Y. 2018. Simple baselines for human pose estimation and tracking. In Proceedings of the European conference on computer vision (ECCV), 466–481.
  41. 41.Xiao, C.; Cao, Q.; Zhong, Y.; Lan, L.; Zhang, X.; Cai, H.; Luo, Z.; and Tao, D. 2023. MotionTrack: Learning Motion Predictor for Multiple Object Tracking. arXiv:2306.02585.
  42. 42.Yang, F.; Odashima, S.; Masui, S.; and Jiang, S. 2023. Hard to track objects with irregular motions and similar appearances? make it easier by buffering the matching space. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 4799–4808.
  43. 43.Yang, J.; Ge, H.; Su, S.; and Liu, G. 2022. Transformer-based two-source motion model for multi-object tracking. Applied Intelligence, 1–13.
  44. 44.Yeh, C.-H.; Lin, C.-Y.; Muchtar, K.; Lai, H.-E.; and Sun, M.-T. 2017. Three-pronged compensation and hysteresis thresholding for moving object detection in real-time video surveillance. IEEE Transactions on Industrial Electronics, 64(6): 4945–4955.
  45. 45.Yu, J.; and McMillan, L. 2004. General linear cameras. In Computer Vision-ECCV 2004: 8th European Conference on Computer Vision, Prague, Czech Republic, May 11-14, 2004. Proceedings, Part II 8, 14–27. Springer.
  46. 46.Yu, Y.; Kurnianggoro, L.; and Jo, K.-H. 2019. Moving object detection for a moving camera based on global motion compensation and adaptive background model. International Journal of Control, Automation and Systems, 17: 1866–1874.
  47. 47.Zhang, Y.; Sun, P.; Jiang, Y.; Yu, D.; Weng, F.; Yuan, Z.; Luo, P.; Liu, W.; and Wang, X. 2022. Bytetrack: Multi-object tracking by associating every detection box. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXII, 1–21. Springer.
  48. 48.Zhang, Y.; Wang, T.; and Zhang, X. 2023. MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 22056–22065.
  49. 49.Zheng, Z.; Wang, P.; Liu, W.; Li, J.; Ye, R.; and Ren, D. 2020. Distance-IoU loss: Faster and better learning for bounding box regression. In Proceedings of the AAAI conference on artificial intelligence, volume 34, 12993–13000.
  50. 50.Zhou, X.; Koltun, V.; and Krahenb uhl, P. 2020. Tracking objects as points. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV, 474–490. Springer.

Citation

MLA
Yi, K., et al. “UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation”. arXiv, 2023, http://arxiv.org/abs/2312.08952v2.
APA
Yi, K., Luo, K., Luo, X., Huang, J., Wu, H., Hu, R., & Hao, W. (2023). UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation. arXiv. http://arxiv.org/abs/2312.08952v2
Chicago
Yi, K., K. Luo, X. Luo, et al. 2023. “UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation”. arXiv. http://arxiv.org/abs/2312.08952v2.
Harvard
Yi, K. et al. (2023) “UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.08952v2.
Vancouver
1. Yi K, Luo K, Luo X, Huang J, Wu H, Hu R, Hao W (2023) UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation. arXiv

BibTeX

@article{yi2023ucmctrack,
  title = {UCMCTrack: Multi-Object Tracking with Uniform Camera Motion Compensation},
  author = {Yi, Kefu and Luo, Kai and Luo, Xiaolei and Huang, Jiangui and Wu, Hao and Hu, Rongdong and Hao, Wei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.08952v2},
  eprint = {2312.08952}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF