Tracking Objects as Points

Xingyi ZhouVladlen KoltunPhilipp Krähenbühl

article2020ECCV1,402 citations

Introduces CenterTrack, a simple real-time framework that tracks objects as points across consecutive frames, eliminating complex temporal association pipelines while achieving state-of-the-art accuracy on 2D and 3D tracking benchmarks.

Listen

Real-time visual tracking of multiple objects is a vital capability for autonomous driving, robotics, and intelligent surveillance. Traditional systems often suffer from high computational overhead and latency because they rely on complex, multi-stage detection pipelines and computationally expensive pairwise feature comparisons across frames.

The article evaluates a streamlined approach, CenterTrack, which tracks objects as single points using a single-pass deep network. It demonstrates the framework's tracking accuracy, processing speed, and architectural modifications across standard two-dimensional and three-dimensional multi-object tracking benchmarks.

The evaluation relies on quantitative experiments across established vision benchmarks, including MOT16, MOT17, KITTI, and nuScenes, alongside pretraining on the CrowdHuman dataset. The methodology applies a single neural network pass per frame combined with a fast greedy matching algorithm based on center-point distances, supporting both private and public detection modes while avoiding complex combinatorial matching schemes.

The findings show that the proposed method achieves competitive accuracy while operating significantly faster than prior systems. On the MOT16 private benchmark, the model ranked second among published methods with a 69.6% tracking accuracy score (MOTA), running online at 17 frames per second (57 ms runtime), whereas competing approaches required significantly more time per frame plus separate detection overhead. Pretraining on the CrowdHuman dataset proved essential, boosting MOT17 validation accuracy from 60.7% to 66.1% by reducing missed detections. On the KITTI benchmark, the full tracking model achieved an 88.7% accuracy score, matching complex state-of-the-art systems while showing that simple greedy association performed identically to the more computationally demanding Hungarian matching algorithm. Additionally, on three-dimensional detection benchmarks, camera-based regression achieved performance comparable to existing monocular methods, though it trailed specialized LiDAR-based systems.

These results indicate that multi-object tracking can be executed efficiently without heavy pairwise re-identification networks or slow optimization steps. By unifying detection and point-based tracking into a single online pass, organizations can lower compute hardware costs, decrease processing latency, and simplify deployment in time-critical systems such as automated vehicles.

Engineering teams deploying this framework should leverage large, static human-detection datasets for pretraining to reduce false negatives. Teams should also adopt simpler greedy association logic over complex optimization algorithms and calibrate output and rendering confidence thresholds (such as 0.4 and 0.5) to balance false alarms and missed tracks. For three-dimensional tracking environments where high positional precision is safety-critical, practitioners should note the performance gap between camera-only inputs and LiDAR sensors before relying purely on monocular data.

arXiv: 2004.01177
  • Paper: Objects as Points, Xingyi Zhou et al. (2019). CenterTrack directly builds upon CenterNet's anchor-free formulation of detecting objects as center points and extends it across time.
  • Paper: Simple online and realtime tracking, Alex Bewley et al. (2016). SORT establishes the foundational baseline for real-time online multi-object tracking using simple spatial and motion cues that CenterTrack seeks to simplify and surpass.
  • Paper: Simple online and realtime tracking with a deep association metric, Nicolai Wojke et al. (2017). DeepSORT introduces deep visual appearance features into online multi-object tracking, representing the standard tracking-by-detection paradigm that CenterTrack re-evaluates and streamlines.
  • Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). This paper presents the MOT16 benchmark suite and standardized evaluation protocols that are central to measuring CenterTrack's multi-object tracking performance.
  • Paper: CornerNet: Detecting Objects as Paired Keypoints, Hei Law et al. (2018). CornerNet pioneered keypoint-based object detection, providing the conceptual foundation for representing objects as points without anchor boxes.
Cover for Tracking Objects as Points

Abstract

Tracking has traditionally been the art of following interest points through space and time. This changed with the rise of powerful deep networks. Nowadays, tracking is dominated by pipelines that perform object detection followed by temporal association, also known as tracking-by-detection. In this paper, we present a simultaneous detection and tracking algorithm that is simpler, faster, and more accurate than the state of the art. Our tracker, CenterTrack, applies a detection model to a pair of images and detections from the prior frame. Given this minimal input, CenterTrack localizes objects and predicts their associations with the previous frame. That's it. CenterTrack is simple, online (no peeking into the future), and real-time. It achieves 67.3% MOTA on the MOT17 challenge at 22 FPS and 89.4% MOTA on the KITTI tracking benchmark at 15 FPS, setting a new state of the art on both datasets. CenterTrack is easily extended to monocular 3D tracking by regressing additional 3D attributes. Using monocular video input, it achieves 28.3% AMOTA@0.2 on the newly released nuScenes 3D tracking benchmark, substantially outperforming the monocular baseline on this benchmark while running at 28 FPS.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Preliminaries
  • 4 Tracking objects as points
  • 4.1 Tracking-conditioned detection
  • 4.2 Association through offsets
  • 4.3 Training on video data
  • 4.4 Training on static image data
  • 4.5 End-to-end 3D object tracking
  • 5 Experiments
  • 5.1 Datasets and evaluation metrics
  • 5.2 Implementation details
  • 5.3 Public detection
  • 5.4 Main results
  • 5.5 Ablation studies
  • 5.6 Comparison to alternative motion models
  • 6 Conclusion
  • References
  • 0.A Tracking algorithms
  • 0.A.1 Private tracking
  • 0.A.2 Public tracking
  • 0.B Results on MOT16
  • 0.C 3D detection
  • 0.D Amodal bounding box regression
  • 0.E CrowdHuman dataset
  • 0.F Pretraining experiments
  • 0.G Additional experiments on KITTI
  • 0.H Output and rendering threshold

Knowls

  1. Knowl 1 — Greedy Center-Distance Association Algorithm for Private Multi-Object Tracking

    algorithm

    In CenterTrack's private tracking mode, object identities are propagated from frame t−1t-1 to frame tt using a greedy matching algorithm based on the Euclidean distance between predicted displaced object centers and current detections.

    Input: Tracked objects in previous frame T(t−1)={(pj(t−1),sj(t−1),idj(t−1))}j=1MT^{(t-1)} = \{(p_j^{(t-1)}, s_j^{(t-1)}, \text{id}_j^{(t-1)})\}_{j=1}^M, where pjp_j is center position, sj=(wj,hj)s_j = (w_j, h_j) is bounding box size; Heatmap peak detections in current frame B^(t)={(p^i(t),d^i(t))}i=1N\hat{B}^{(t)} = \{(\hat{p}_i^{(t)}, \hat{d}_i^{(t)})\}_{i=1}^N with 2D displacement offset d^i(t)\hat{d}_i^{(t)}, sorted in descending confidence order.
    Output: Tracked objects in current frame T(t)={(p^i(t),s^i(t),idi(t))}i=1NT^{(t)} = \{(\hat{p}_i^{(t)}, \hat{s}_i^{(t)}, \text{id}_i^{(t)})\}_{i=1}^N.
    Initialize T(t)←∅T^{(t)} \leftarrow \emptyset
    Initialize matched track index set S←∅S \leftarrow \emptyset
    Compute cost matrix W∈RN×MW \in \mathbb{R}^{N \times M} where Wij=∥p^i(t)−d^i(t)−pj(t−1)∥2W_{ij} = \|\hat{p}_i^{(t)} - \hat{d}_i^{(t)} - p_j^{(t-1)}\|_2
    for i←1i \leftarrow 1 to NN do
        j←arg⁡min⁡j∉SWijj \leftarrow \arg\min_{j \notin S} W_{ij}
        κ←min⁡(w^ih^i,wjhj)\kappa \leftarrow \min(\sqrt{\hat{w}_i \hat{h}_i}, \sqrt{w_j h_j})
        if Wij<κW_{ij} < \kappa then
            T(t)←T(t)∪{(p^i(t),s^i(t),idj(t−1))}T^{(t)} \leftarrow T^{(t)} \cup \{(\hat{p}_i^{(t)}, \hat{s}_i^{(t)}, \text{id}_j^{(t-1)})\}
            S←S∪{j}S \leftarrow S \cup \{j\}
        else
            T(t)←T(t)∪{(p^i(t),s^i(t),NewId)}T^{(t)} \leftarrow T^{(t)} \cup \{(\hat{p}_i^{(t)}, \hat{s}_i^{(t)}, \text{NewId})\}
        end if
    end for
    return T(t)T^{(t)}

    The algorithm computes the distance matrix WijW_{ij} between backward-displaced current center positions p^i(t)−d^i(t)\hat{p}_i^{(t)} - \hat{d}_i^{(t)} and prior frame centers pj(t−1)p_j^{(t-1)}. In descending confidence order of current peaks, each object ii is greedily matched to the nearest unmatched prior track j∉Sj \notin S. If WijW_{ij} is less than the scale threshold κ=min⁡(w^ih^i,wjhj)\kappa = \min(\sqrt{\hat{w}_i \hat{h}_i}, \sqrt{w_j h_j}), the identity is propagated; otherwise, a new track with a new identity is created. The same greedy association is applied to both 2D and 3D tracking.

  2. Knowl 2 — Public Detection Tracking via Public Detection Gated Track Initialization

    algorithm

    Under the public detection protocol for multi-object tracking, CenterTrack extends its greedy center-distance association by restricting track birth: new tracks are only initialized if they are physically close to an externally provided public detection.

    Input: Tracked objects from previous frame T(t−1)={(pj(t−1),sj(t−1),idj(t−1))}j=1MT^{(t-1)} = \{(p_j^{(t-1)}, s_j^{(t-1)}, \text{id}_j^{(t-1)})\}_{j=1}^M; Current frame heatmap peaks B^(t)={(p^i(t),d^i(t))}i=1N\hat{B}^{(t)} = \{(\hat{p}_i^{(t)}, \hat{d}_i^{(t)})\}_{i=1}^N sorted in descending confidence; External public detections D^(t)={(pk(t),sk(t))}k=1K\hat{D}^{(t)} = \{(p_k^{(t)}, s_k^{(t)})\}_{k=1}^K with center pk(t)p_k^{(t)} and size sk(t)=(wk,hk)s_k^{(t)} = (w_k, h_k).
    Output: Tracked objects in current frame T(t)T^{(t)}.
    Initialize T(t)←∅T^{(t)} \leftarrow \emptyset
    Initialize matched track index set S←∅S \leftarrow \emptyset
    Compute cost matrix W∈RN×MW \in \mathbb{R}^{N \times M} where Wij=∥p^i(t)−d^i(t)−pj(t−1)∥2W_{ij} = \|\hat{p}_i^{(t)} - \hat{d}_i^{(t)} - p_j^{(t-1)}\|_2
    Compute public detection distance matrix W′∈RN×KW' \in \mathbb{R}^{N \times K} where Wik′=∥p^i(t)−pk(t)∥2W'_{ik} = \|\hat{p}_i^{(t)} - p_k^{(t)}\|_2
    for i←1i \leftarrow 1 to NN do
        j←arg⁡min⁡j∉SWijj \leftarrow \arg\min_{j \notin S} W_{ij}
        κ←min⁡(w^ih^i,wjhj)\kappa \leftarrow \min(\sqrt{\hat{w}_i \hat{h}_i}, \sqrt{w_j h_j})
        if Wij<κW_{ij} < \kappa then
            T(t)←T(t)∪{(p^i(t),s^i(t),idj(t−1))}T^{(t)} \leftarrow T^{(t)} \cup \{(\hat{p}_i^{(t)}, \hat{s}_i^{(t)}, \text{id}_j^{(t-1)})\}
            S←S∪{j}S \leftarrow S \cup \{j\}
        else
            k←arg⁡min⁡k=1KWik′k \leftarrow \arg\min_{k=1}^K W'_{ik}
            κ′←min⁡(w^ih^i,wkhk)\kappa' \leftarrow \min(\sqrt{\hat{w}_i \hat{h}_i}, \sqrt{w_k h_k})
            if Wik′<κ′W'_{ik} < \kappa' then
                T(t)←T(t)∪{(p^i(t),s^i(t),NewId)}T^{(t)} \leftarrow T^{(t)} \cup \{(\hat{p}_i^{(t)}, \hat{s}_i^{(t)}, \text{NewId})\}
            end if
        end if
    end for
    return T(t)T^{(t)}

    Association with active existing tracks uses the same greedy center distance as private tracking. When a peak cannot be matched to an existing track, it is compared against all public detections D^(t)\hat{D}^{(t)}. A new track is created only if the minimum distance Wik′W'_{ik} to a public detection is within κ′=min⁡(w^ih^i,wkhk)\kappa' = \min(\sqrt{\hat{w}_i \hat{h}_i}, \sqrt{w_k h_k}); otherwise, the candidate peak is rejected.

  3. Knowl 3 — Monocular 3D Bounding Box Center Offset Regression

    model/method

    In monocular 3D object detection, perspective projection causes the 2D bounding box center to deviate from the projection of the 3D bounding box center. To resolve this offset, the network predicts a 2D offset vector field F^∈RWR×HR×2\hat{F} \in \mathbb{R}^{\frac{W}{R} \times \frac{H}{R} \times 2} from the 2D bounding box center to the projected 3D bounding box center, where WW and HH are the input image dimensions and RR is the downsampling stride.

    The predicted offset f^k=F^pk\hat{f}_k = \hat{F}_{p_k} at the ground-truth 2D center location pkp_k is supervised with an L1L_1 loss: Loff3d=1N∑k=1N∣f^k−fk∣L_{off3d} = \frac{1}{N} \sum_{k=1}^N \left| \hat{f}_k - f_k \right| where fk∈R2f_k \in \mathbb{R}^2 is the ground-truth offset between the 2D bounding box center and the projected 3D bounding box center for object kk, and NN is the number of labeled objects.

    In addition to F^\hat{F}, the 3D detection heads regress object depth D^∈RWR×HR\hat{D} \in \mathbb{R}^{\frac{W}{R} \times \frac{H}{R}}, 3D bounding box extent Γ^∈RWR×HR×3\hat{\Gamma} \in \mathbb{R}^{\frac{W}{R} \times \frac{H}{R} \times 3}, and 3D orientation vector A^∈RWR×HR×8\hat{A} \in \mathbb{R}^{\frac{W}{R} \times \frac{H}{R} \times 8}.

  4. Knowl 4 — 4-Channel Border Distance Regression for Amodal Bounding Boxes

    model/method

    In multi-object tracking datasets, target bounding boxes are annotated amodally, meaning the geometric center of the complete bounding box can fall outside the visible image frame for truncated objects. Because point-based detectors require the center keypoint to lie within the image boundaries, the 2-channel (w,h)(w, h) size regression head is extended to a 4-channel head: A^∈RWR×HR×4\hat{A} \in \mathbb{R}^{\frac{W}{R} \times \frac{H}{R} \times 4} where WW and HH are image dimensions and RR is the downsampling stride. The four channels represent distances from the detected in-frame keypoint center to the top, left, bottom, and right borders of the amodal bounding box.

    This formulation allows predicted amodal bounding boxes to be asymmetric with respect to the detected in-frame center. The head is trained with an L1L_1 loss: Lamodal_size=1N∑i=1N∣A^pi−ai∣L_{amodal\_size} = \frac{1}{N} \sum_{i=1}^N \left| \hat{A}_{p_i} - a_i \right| where pip_i is the detected in-frame center location of object ii, ai∈R4a_i \in \mathbb{R}^4 is the ground truth border distance vector, and NN is the number of objects.

  5. Knowl 5 — MOT16 Benchmark Evaluation for Private Detection

    empirical result

    On the MOT16 test set under the private detection benchmark, CenterTrack operates online at 17 FPS (57 ms total per frame) with a single forward pass through a single network, avoiding separate detection pipelines (+D) or O(n2)\mathcal{O}(n^2) pairwise Siamese Re-ID networks.

    Method Time (ms) MOTA ↑\uparrow IDF1 ↑\uparrow FP ↓\downarrow FN ↓\downarrow IDSW ↓\downarrow
    SORT 36+D 60.4 56.1 11183 59867 1135
    DeepSORT 59+D 61.4 62.2 12852 56668 781
    POI 100+D 66.1 65.1 5061 55914 805
    KNDT 1428+D 68.2 60.0 11479 45605 933
    LMP_p 2000+D 71.0 70.1 7880 44564 434
    Ours (CenterTrack Private) 57 69.6 60.7 10458 42805 2124

    CenterTrack achieves 69.6% MOTA, ranking second among published methods on the MOT16 leaderboard while running significantly faster than offline or matching-heavy methods (e.g., LMP_p at 2000+ ms and 71.0% MOTA).

  6. Knowl 6 — nuScenes Monocular 3D Detection Benchmark Evaluation

    empirical result

    Monocular 3D detection with CenterNet augmented by 2D-to-3D projected center offset regression was evaluated on the nuScenes test benchmark across 3D mean Average Precision (mAP), mean Translation Error (mATE), mean Size Error (mASE), mean Orientation Error (mAOE), mean Velocity Error (mAVE), mean Attribute Error (mAAE), and nuScenes Detection Score (NDS, weighted average giving weight 5 to mAP and 1 to each error metric).

    Method Modality mAP ↑\uparrow mATE ↓\downarrow mASE ↓\downarrow mAOE ↓\downarrow mAVE ↓\downarrow mAAE ↓\downarrow NDS ↑\uparrow
    Megvii LiDAR 52.8 0.300 0.247 0.379 0.245 0.140 63.3
    PointPillars LiDAR 30.5 0.517 0.290 0.500 0.316 0.368 45.3
    Mapillary Camera 30.4 0.738 0.263 0.546 1.0 0.134 38.4
    CenterNet (with offset) Camera 33.8 0.658 0.255 0.629 1.0 0.141 40.1

    CenterNet with 3D offset prediction reaches 33.8% mAP and 40.1 NDS using monocular camera imagery alone, outperforming camera-based Mapillary (30.4% mAP, 38.4 NDS) and achieving higher mAP than the LiDAR-based PointPillars baseline (30.5% mAP), though remaining below state-of-the-art LiDAR methods (Megvii at 52.8% mAP, 63.3 NDS).

  7. Knowl 7 — Ablation of CrowdHuman Pretraining and Detection Modes on MOT17

    empirical result

    Pretraining on the CrowdHuman static image dataset (using input resolution 512×512512 \times 512, false positive augmentation ratio λfp=0.1\lambda_{fp} = 0.1, false negative ratio λfn=0.4\lambda_{fn} = 0.4, random scaling ratio 0.05, and random translation ratio 0.05) evaluated on the MOT17 validation set demonstrates substantial reduction in false negatives.

    Configuration MOTA ↑\uparrow IDF1 ↑\uparrow MT ↑\uparrow ML ↓\downarrow FP ↓\downarrow FN ↓\downarrow IDSW ↓\downarrow
    Ours (Full model) 66.1 64.2 41.3 21.2 4.5% 28.4% 1.0%
    Only CrowdHuman 52.2 53.8 33.6 25.1 6.7% 39.7% 1.4%
    Scratch (Private) 60.7 62.8 33.0 22.4 4.0% 34.2% 1.0%
    Scratch (Public) 57.4 59.6 31.1 27.1 2.1% 39.6% 1.0%

    Key results:

    • Pretraining on CrowdHuman boosts MOTA from 60.7% to 66.1% and IDF1 from 62.8% to 64.2%, with false negatives dropping from 34.2% to 28.4%.
    • Evaluated zero-shot on MOT without seeing any MOT training frames, the model trained only on CrowdHuman achieves 52.2% MOTA.
    • The public detection configuration trained from scratch achieves 57.4% MOTA with a reduced false positive rate (2.1%).
  8. Knowl 8 — Component Ablations on the KITTI Multi-Object Tracking Benchmark

    empirical result

    An ablation study on the KITTI tracking validation set investigates the role of video training, prior heatmap noise simulation, association algorithms, and pretraining datasets.

    Configuration MOTA ↑\uparrow MOTP ↑\uparrow MT ↑\uparrow ML ↓\downarrow FP ↓\downarrow FN ↓\downarrow IDSW ↓\downarrow
    Ours (Full model) 88.7 86.7 90.3 2.1 5.4% 5.8% 0.1%
    Static image only 86.8 86.5 88.5 2.2 4.8% 7.9% 0.4%
    w.o. noisy heatmap 80.1 85.3 76.2 7.6 3.8% 16.1% 0.1%
    Hungarian matching 88.7 86.7 90.3 2.1 5.4% 5.8% 0.1%
    Scratch (w.o. nuScenes) 84.5 83.2 83.4 2.8 5.7% 9.6% 0.3%

    Findings on KITTI:

    • Removing noise simulation on the prior heatmap degrades MOTA from 88.7% to 80.1%, with false negatives increasing from 5.8% to 16.1%.
    • Training on static images (86.8% MOTA) underperforms training on video (88.7% MOTA) due to large inter-frame camera motion in KITTI driving scenes.
    • Using bipartite matching via the Hungarian algorithm yields the exact same performance as greedy nearest-neighbor matching (88.7% MOTA, 0.1% IDSW).
    • Training from scratch without nuScenes pretraining achieves 84.5% MOTA.
  9. Knowl 9 — Detection Output and Prior Heatmap Rendering Threshold Sensitivity

    empirical result

    CenterTrack tracking performance on the MOT validation set is evaluated across different detection output confidence thresholds θ\theta and prior heatmap rendering thresholds τ\tau.

    θ\theta τ\tau MOTA ↑\uparrow IDF1 ↑\uparrow MT ↑\uparrow ML ↓\downarrow FP ↓\downarrow FN ↓\downarrow IDSW ↓\downarrow
    0.4 0.4 62.6 64.9 44.0 18.9 10.3% 26.4% 0.7%
    0.4 0.6 65.5 63.2 38.6 22.4 2.5% 30.5% 1.5%
    0.4 0.5 66.1 64.2 41.3 21.2 4.5% 28.4% 1.0%
    0.3 0.5 66.2 64.3 43.1 19.2 5.7% 26.9% 1.2%
    0.5 0.5 65.2 62.1 39.8 23.0 3.7% 30.2% 0.9%

    Increasing both thresholds decreases output density, reducing false positives while increasing false negatives. The parameter pair θ=0.4\theta = 0.4 and τ=0.5\tau = 0.5 achieves the strongest overall balance across MOTA (66.1%), IDF1 (64.2%), and ID switches (1.0%).

Coverage note — None was omitted; all contributed algorithms, method formulations, and experimental evaluations in the supplementary document have been fully captured.

References

  1. 1.Bergmann, P., Meinhardt, T., Leal-Taixe, L.: Tracking without bells and whistles. In: ICCV (2019)
  2. 2.Bewley, A., Ge, Z., Ott, L., Ramos, F., Upcroft, B.: Simple online and realtime tracking. In: ICIP (2016)
  3. 3.Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuScenes: A multimodal dataset for autonomous driving. In: CVPR (2020)
  4. 4.Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? The KITTI vision benchmark suite. In: CVPR (2012)
  5. 5.Hu, H.N., Cai, Q.Z., Wang, D., Lin, J., Sun, M., Krahenbuhl, P., Darrell, T., Yu, F.: Joint monocular 3D detection and tracking. In: ICCV (2019)
  6. 6.Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., Beijbom, O.: Pointpillars: Fast encoders for object detection from point clouds. In: CVPR (2019)
  7. 7.Milan, A., Leal-Taixe, L., Reid, I., Roth, S., Schindler, K.: MOT16: A benchmark for multiobject tracking. arXiv:1603.00831 (2016)
  8. 8.Ren, J., Chen, X., Liu, J., Sun, W., Pang, J., Yan, Q., Tai, Y.W., Xu, L.: Accurate single stage detector using recurrent rolling convolution. In: CVPR (2017)
  9. 9.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. In: NIPS (2015)
  10. 10.Shao, S., Zhao, Z., Li, B., Xiao, T., Yu, G., Zhang, X., Sun, J.: Crowdhuman: A benchmark for detecting human in a crowd. arXiv:1805.00123 (2018)
  11. 11.Sharma, S., Ansari, J.A., Murthy, J.K., Krishna, K.M.: Beyond pixels: Leveraging geometry and shape cues for online multi-object tracking. In: ICRA (2018)
  12. 12.Simonelli, A., Bulo, S.R.R., Porzi, L., Lopez-Antequera, M., Kontschieder, P.: Disentangling monocular 3d object detection. In: ICCV (2019)
  13. 13.Tang, S., Andriluka, M., Andres, B., Schiele, B.: Multiple people tracking by lifted multicut and person re-identification. In: CVPR (2017)
  14. 14.Weng, X., Kitani, K.: A baseline for 3d multi-object tracking. arXiv:1907.03961 (2019)
  15. 15.Wojke, N., Bewley, A., Paulus, D.: Simple online and realtime tracking with a deep association metric. In: ICIP (2017)
  16. 16.Yu, F., Li, W., Li, Q., Liu, Y., Shi, X., Yan, J.: Poi: Multiple object tracking with high performance detection and appearance feature. In: ECCV Workshops (2016)
  17. 17.Zhou, X., Wang, D., Krahenbuhl, P.: Objects as points. arXiv:1904.07850 (2019)
  18. 18.Zhu, B., Jiang, Z., Zhou, X., Li, Z., Yu, G.: Class-balanced grouping and sampling for point cloud 3D object detection. arXiv:1908.09492 (2019)

Citation

MLA
Zhou, X., et al. “Tracking Objects as Points”. arXiv, 2020, http://arxiv.org/abs/2004.01177v2.
APA
Zhou, X., Koltun, V., & Krähenbühl, P. (2020). Tracking Objects as Points. arXiv. http://arxiv.org/abs/2004.01177v2
Chicago
Zhou, X., V. Koltun, and P. Krähenbühl. 2020. “Tracking Objects as Points”. arXiv. http://arxiv.org/abs/2004.01177v2.
Harvard
Zhou, X., Koltun, V. and Krähenbühl, P. (2020) “Tracking Objects as Points”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2004.01177v2.
Vancouver
1. Zhou X, Koltun V, Krähenbühl P (2020) Tracking Objects as Points. arXiv

BibTeX

@article{zhou2020tracking,
  title = {Tracking Objects as Points},
  author = {Zhou, Xingyi and Koltun, Vladlen and Krähenbühl, Philipp},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2004.01177v2},
  eprint = {2004.01177}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF