Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking

Ziqi PangJie LiPavel TokmakovDian ChenSergey ZagoruykoYu-Xiong Wang

article2023CVPR91 citations

Presents an end-to-end multi-camera 3D multi-object tracking framework that models both past temporal contexts and predicted future trajectories to maintain object continuity through long occlusions, reducing identity switches on the nuScenes benchmark by ninety percent.

Listen

Reliable 3D multi-object tracking is essential for the safe navigation of autonomous vehicles. While laser-based LiDAR systems are widely used, their high cost and hardware constraints make camera-only vision systems an attractive alternative. However, vision-only tracking faces significant hurdles, including depth ambiguity, noisy single-frame detections, and track loss during object occlusions or camera switches.

The article introduces and evaluates "Past-and-Future reasoning for Tracking" (PF-Track), an end-to-end framework designed to unify 3D object detection, tracking, and trajectory prediction. The primary objective is to demonstrate that integrating historical observations with future motion forecasting substantially improves tracking accuracy and maintains spatio-temporal continuity across multi-camera setups.

The approach operates within an attention-based tracking architecture using 3D object queries that persist over time. A "Past Reasoning" module refines object features and 3D bounding boxes by cross-referencing data across previous timeframes and neighboring objects. Simultaneously, a "Future Reasoning" module forecasts long-term trajectories, which are used both to advance query positions across frames and to sustain tracks through a strategy called track extension when objects are temporarily occluded. The methodology was evaluated on the large-scale nuScenes autonomous driving benchmark across 1,000 multi-camera video sequences spanning seven moving object categories.

The evaluation produced four key findings. First, PF-Track achieved state-of-the-art tracking accuracy on the nuScenes benchmark, reaching an Average Multi-Object Tracking Accuracy (AMOTA) score of 0.434 on the test set, outperforming existing camera-based methods. Second, the system reduced identity switches by approximately 90% compared to prior approaches, cutting ID switches by an order of magnitude (from several thousands down to a few hundred). Third, trajectory forecasting directly aided tracking, with track extensions of up to two seconds effectively preventing track loss during severe occlusions without requiring a dedicated re-identification module. Fourth, predicting future paths directly from high-dimensional query features reduced displacement errors compared to conventional methods that forecast strictly from low-level positional coordinates.

These findings indicate that end-to-end multi-task modeling resolves fundamental weaknesses in camera-based perception pipelines. By maintaining consistent object identities through visual occlusions and across different camera views, the framework enhances road safety and lowers system vulnerability. This demonstrates that camera-only autonomous navigation can achieve high spatio-temporal tracking fidelity without depending on expensive LiDAR hardware.

Organizations developing vision-based autonomous systems should transition from modular detection-and-tracking pipelines toward unified architectures that model both past context and future trajectories. Adopting learned motion forecasting to maintain occluded tracks offers an effective alternative to heuristic filters or complex re-identification models. Future work should focus on scaling the model for real-time edge hardware, testing across diverse sensor configurations, and evaluating performance in adverse environmental conditions.

Cover for Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking

Abstract

This work proposes an end-to-end multi-camera 3D multi-object tracking (MOT) framework. It emphasizes spatio-temporal continuity and integrates both past and future reasoning for tracked objects. Thus, we name it “Past-and-Future reasoning for Tracking” (PF-Track). Specifically, our method adopts the “tracking by attention” framework and represents tracked instances coherently over time with object queries. To explicitly use historical cues, our “Past Reasoning” module learns to refine the tracks and enhance the object features by cross-attending to queries from previous frames and other objects. The “Future Reasoning” module digests historical information and predicts robust future trajectories. In the case of long-term occlusions, our method maintains the object positions and enables re-association by integrating motion predictions. On the nuScenes dataset, our method improves AMOTA by a large margin and remarkably reduces ID-Switches by 90% compared to prior approaches, which is an order of magnitude less. The code and models are made available at https://github.com/TRI-ML/PF-Track.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method: PF-Track
  • 3.1. PF-Track Pipeline
  • 3.2. Past Reasoning
  • 3.3. Future Reasoning
  • 3.4. Loss Functions
  • 4. Experiments
  • 4.1. Datasets and Metrics
  • 4.2. Implementation Details
  • 4.3. State-of-the-art Comparison on nuScenes
  • 4.4. Ablation Studies
  • 4.5. Qualitative Results
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — PF-Track Framework for Multi-Camera 3D Multi-Object Tracking

    model/method

    PF-Track (Past-and-Future reasoning for Tracking) is an end-to-end, vision-only multi-camera 3D multi-object tracking (MOT) framework based on the tracking-by-attention paradigm. The system iteratively maintains and updates a set of 3D object queries across video timestamps.

    At each timestamp tt, given multi-view camera images {Itk}k=1K\{\mathbf{I}_t^k\}_{k=1}^K, the tracking objective is to output 3D bounding boxes Bt={bti}\mathbf{B}_t = \{\mathbf{b}_t^i\} associated with unique instance identity ii.

    Each query qti∈Qt\mathbf{q}_t^i \in \mathbf{Q}_t represents an individual 3D instance and is characterized by a feature embedding fti∈Rd\mathbf{f}_t^i \in \mathbb{R}^d and a 3D center location cti=(xti,yti,zti)∈R3\mathbf{c}_t^i = (x_t^i, y_t^i, z_t^i) \in \mathbb{R}^3. To track existing instances and detect newly appearing objects, the query set Qt\mathbf{Q}_t consists of:

    1. Track queries propagated from the preceding frame Qt−1\mathbf{Q}_{t-1} via Qt←Prop(Qt−1)\mathbf{Q}_t \leftarrow \mathbf{Prop}(\mathbf{Q}_{t-1}).
    2. A fixed set of 500 learnable detection queries initialized as learnable embeddings.

    The overall tracking iteration proceeds in three sequential stages:

    1. Query Decoding: An attention-based 3D detector (such as PETR) lifts query 3D reference positions cti\mathbf{c}_t^i into positional embeddings to attend to multi-view image features Ft\mathbf{F}_t, decoding initial 3D bounding boxes BtD\mathbf{B}_t^D and updated query features QtD\mathbf{Q}_t^D: BtD,QtD←Decoder(Ft,Qt)\mathbf{B}_t^D, \mathbf{Q}_t^D \leftarrow \mathbf{Decoder}(\mathbf{F}_t, \mathbf{Q}_t)
    2. Past Reasoning (PR\mathbf{PR}): A module aggregates historical query representations stored in a temporal query queue over a history window of τh\tau_h frames, Qt−τh:t−1\mathbf{Q}_{t-\tau_h:t-1}, generating refined query features QtR\mathbf{Q}_t^R and refined bounding boxes BtR\mathbf{B}_t^R: QtR,BtR←PR(QtD,BtD,Qt−τh:t−1)\mathbf{Q}_t^R, \mathbf{B}_t^R \leftarrow \mathbf{PR}(\mathbf{Q}_t^D, \mathbf{B}_t^D, \mathbf{Q}_{t-\tau_h:t-1})
    3. Future Reasoning (FR\mathbf{FR}): A motion forecasting module digests historical features to estimate future trajectory displacements Mt:t+τf\mathbf{M}_{t:t+\tau_f} over τf\tau_f future frames, updates query spatial coordinates to frame t+1t+1, and applies track extension to maintain states of occluded objects: Qt+1,Mt:t+τf←FR(QtR,Qt−τh:t−1)\mathbf{Q}_{t+1}, \mathbf{M}_{t:t+\tau_f} \leftarrow \mathbf{FR}(\mathbf{Q}_t^R, \mathbf{Q}_{t-\tau_h:t-1})

    The refined bounding boxes BtR\mathbf{B}_t^R serve as the final 3D tracking output at timestamp tt.

  2. Knowl 2 — Query Refinement via Decoupled Cross-Frame and Cross-Object Attention

    model/method

    To enhance object query representations using historical context without incurring heavy computational overhead, the Past Reasoning module in PF-Track decouples temporal and inter-object attention into two successive operations on the query queue:

    1. Cross-Frame Attention (Temporal Aggregation): For each tracked object instance ii, the query feature fti\mathbf{f}_t^i cross-attends across its own historical features within a temporal window of τh\tau_h frames: fti←CrossFrameAttn(Q=fti,K=ft−τh:ti,V=ft−τh:ti,PE=Pos(t−τh:t))\mathbf{f}_t^i \leftarrow \mathbf{CrossFrameAttn}(\text{Q}=\mathbf{f}_t^i, \text{K}=\mathbf{f}_{t-\tau_h:t}^i, \text{V}=\mathbf{f}_{t-\tau_h:t}^i, \text{PE}=\mathbf{Pos}(t-\tau_h:t)) where Pos(t−τh:t)\mathbf{Pos}(t-\tau_h:t) is a temporal positional encoding. History slots with empty features (e.g., newly initialized objects) are masked out.

    2. Cross-Object Attention (Spatial Context Interplay): The temporally aggregated features for all NtN_t active instances at timestamp tt, denoted ft1:Nt\mathbf{f}_t^{1:N_t}, attend to each other guided by their 3D spatial locations: ft1:Nt←CrossObjectAttn(Q, K, V=ft1:Nt,PE=Pos(ct1:Nt))\mathbf{f}_t^{1:N_t} \leftarrow \mathbf{CrossObjectAttn}(\text{Q, K, V}=\mathbf{f}_t^{1:N_t}, \text{PE}=\mathbf{Pos}(\mathbf{c}_t^{1:N_t})) where Pos(ct1:Nt)\mathbf{Pos}(\mathbf{c}_t^{1:N_t}) is the 3D spatial positional encoding of center coordinates ct\mathbf{c}_t.

    Computational Complexity: Decoupling reduces computational complexity from O(Nt2τh2)\mathcal{O}(N_t^2 \tau_h^2) for global spatio-temporal attention to O(Nt2+Ntτh2)\mathcal{O}(N_t^2 + N_t \tau_h^2), while allowing dedicated temporal and spatial positional encodings for each stage.

  3. Knowl 3 — Track Refinement Parameterization and Bounding Box Adjustment

    equation

    In the Past Reasoning module of PF-Track, the refined query feature fti∈QtR\mathbf{f}_t^i \in \mathbf{Q}_t^R is passed through a Multi-Layer Perceptron (MLP) to predict residual adjustments and refined attributes for the 3D bounding box of object ii:

    (Δxti,Δyti,Δzti,lti,wti,hti,θti,vti,sti)=MLP(fti)(\Delta x_t^i, \Delta y_t^i, \Delta z_t^i, l_t^i, w_t^i, h_t^i, \theta_t^i, \mathbf{v}_t^i, s_t^i) = \mathbf{MLP}(\mathbf{f}_t^i)

    where:

    • (Δxti,Δyti,Δzti)∈R3(\Delta x_t^i, \Delta y_t^i, \Delta z_t^i) \in \mathbb{R}^3 are 3D center coordinate residuals.
    • lti,wti,hti∈R+l_t^i, w_t^i, h_t^i \in \mathbb{R}^+ denote predicted 3D box length, width, and height.
    • θti∈[−π,π]\theta_t^i \in [-\pi, \pi] is the bounding box yaw orientation.
    • vti∈R2\mathbf{v}_t^i \in \mathbb{R}^2 represents the estimated velocity vector.
    • sti∈[0,1]s_t^i \in [0, 1] represents the classification confidence score.

    The final refined 3D bounding box bti∈BtR\mathbf{b}_t^i \in \mathbf{B}_t^R is constructed by adding the residuals to the coarse detected center coordinates (xti,yti,zti)(x_t^i, y_t^i, z_t^i) from the initial detection box BtD\mathbf{B}_t^D:

    bti=(xti+Δxti,yti+Δyti,zti+Δzti,lti,wti,hti,θti,vti,sti)\mathbf{b}_t^i = (x_t^i + \Delta x_t^i, y_t^i + \Delta y_t^i, z_t^i + \Delta z_t^i, l_t^i, w_t^i, h_t^i, \theta_t^i, \mathbf{v}_t^i, s_t^i)

  4. Knowl 4 — Future Motion Prediction and Single-Step Query Propagation

    model/method

    PF-Track models future motion using an attention-based forecasting head to capture dynamics directly from historical query features. This predicted motion drives query position propagation to the next timestamp.

    For each tracked object ii, motion query embeddings mft:t+τfi\mathbf{mf}_{t:t+\tau_f}^i covering a forecast horizon of τf\tau_f timestamps are initialized to zero and updated via cross-frame attention over historical query features ft−τh:ti\mathbf{f}_{t-\tau_h:t}^i:

    mft:t+τfi←CrossFrameAttn(Q=mft:t+τfi,K=ft−τh:ti,V=ft−τh:ti,PE=Pos(t−τh:t+τf))\mathbf{mf}_{t:t+\tau_f}^i \leftarrow \mathbf{CrossFrameAttn}(\text{Q}=\mathbf{mf}_{t:t+\tau_f}^i, \text{K}=\mathbf{f}_{t-\tau_h:t}^i, \text{V}=\mathbf{f}_{t-\tau_h:t}^i, \text{PE}=\mathbf{Pos}(t-\tau_h:t+\tau_f))

    where Pos(t−τh:t+τf)\mathbf{Pos}(t-\tau_h:t+\tau_f) denotes the temporal positional encoding covering both past and future horizons.

    The frame-by-frame 3D movement offsets mt:t+τfi={mt:t+ki}k=1τf\mathbf{m}_{t:t+\tau_f}^i = \{\mathbf{m}_{t:t+k}^i\}_{k=1}^{\tau_f} are decoded using an MLP: mt:t+τfi=MLP(mft:t+τfi)\mathbf{m}_{t:t+\tau_f}^i = \mathbf{MLP}(\mathbf{mf}_{t:t+\tau_f}^i)

    Single-Step Propagation: The 3D center location cti\mathbf{c}_t^i of query ii is propagated to timestamp t+1t+1 using the predicted single-step movement mt:t+1i∈R3\mathbf{m}_{t:t+1}^i \in \mathbb{R}^3: ct+1i=cti+mt:t+1i\mathbf{c}_{t+1}^i = \mathbf{c}_t^i + \mathbf{m}_{t:t+1}^i This propagation replaces reliance on coarse velocity vectors output by single-frame decoders.

  5. Knowl 5 — Track Extension Mechanism for Occlusion Handling without Re-Identification

    model/method

    To handle missing detections caused by occlusion, camera switching, or noisy sensor observations, PF-Track employs a track extension strategy that maintains track continuity without an explicit appearance-based Re-ID module.

    When a tracked instance fails to produce a high-confidence detection box at timestamp tt (e.g., due to occlusion or low detector score), the tracker ignores the low-confidence detection output and instead infers the object position using the multi-step future trajectory mtc:tc+τfi\mathbf{m}_{t_c:t_c+\tau_f}^i predicted at the most recent confident timestamp tc<tt_c < t:

    cti=ctci+mtc:ti\mathbf{c}_t^i = \mathbf{c}_{t_c}^i + \mathbf{m}_{t_c:t}^i

    This extrapolation maintains object queries across up to τf−1\tau_f - 1 consecutive occluded frames. When the object re-emerges (even in a different camera view), the query's spatially extrapolated position aligns with the re-appearing object, preserving identity association and preventing identity switches (IDS) caused by premature track termination or incorrect matching.

  6. Knowl 6 — PF-Track Multi-Task Training Objective

    equation

    PF-Track is trained end-to-end using a joint multi-task loss function combining detection, track refinement, and future trajectory forecasting:

    L=λclsDLclsD+λboxDLboxD+λclsRLclsR+λboxRLboxR+λfLf\mathcal{L} = \lambda_{\text{cls}}^D \mathcal{L}_{\text{cls}}^D + \lambda_{\text{box}}^D \mathcal{L}_{\text{box}}^D + \lambda_{\text{cls}}^R \mathcal{L}_{\text{cls}}^R + \lambda_{\text{box}}^R \mathcal{L}_{\text{box}}^R + \lambda_f \mathcal{L}_f

    where:

    • LclsD\mathcal{L}_{\text{cls}}^D and LclsR\mathcal{L}_{\text{cls}}^R are focal classification losses supervising the category scores of initial detections BtD\mathbf{B}_t^D and refined detections BtR\mathbf{B}_t^R, weighted by λclsD\lambda_{\text{cls}}^D and λclsR\lambda_{\text{cls}}^R.
    • LboxD\mathcal{L}_{\text{box}}^D and LboxR\mathcal{L}_{\text{box}}^R are L1L_1 bounding box regression losses supervising 3D box attributes (center, size, orientation, velocity) of BtD\mathbf{B}_t^D and BtR\mathbf{B}_t^R, weighted by λboxD\lambda_{\text{box}}^D and λboxR\lambda_{\text{box}}^R.
    • Lf\mathcal{L}_f is an L1L_1 trajectory forecasting loss measuring the deviation between predicted future displacements mt:t+τfi\mathbf{m}_{t:t+\tau_f}^i and ground-truth future movements over τf\tau_f timestamps, weighted by λf\lambda_f.

    Ground-truth assignment couples queries to consistent ground-truth instances across consecutive temporal frames during training to enforce track identity consistency.

  7. Knowl 7 — nuScenes Multi-Camera 3D MOT Benchmark Performance

    data/table

    PF-Track was evaluated against state-of-the-art vision-based 3D MOT algorithms on the nuScenes dataset across 7 mobile object categories. Evaluation metrics follow official nuScenes protocols: Average Multi-Object Tracking Accuracy (AMOTA), Average Multi-Object Tracking Precision (AMOTP in meters), Recall, Multi-Object Tracking Accuracy (MOTA), and Identity Switches (IDS).

    Method AMOTA ↑\uparrow AMOTP ↓\downarrow Recall ↑\uparrow MOTA ↑\uparrow IDS ↓\downarrow
    Validation Split
    DEFT 0.201 N/A N/A 0.171 N/A
    QD3DT 0.242 1.518 39.9% 0.218 5646
    MUTR3D 0.294 1.498 42.7% 0.267 3822
    TripletTrack 0.285 1.485 N/A N/A N/A
    CC-3DT 0.429 1.257 53.4% 0.385 2219
    PF-Track-S (Ours) 0.408 1.343 50.7% 0.376 166
    PF-Track-F (Ours) 0.479 1.227 59.0% 0.435 181
    Test Split
    CenterTrack 0.046 1.543 23.3% 0.043 3807
    PermaTrack 0.066 1.491 18.9% 0.060 3598
    DEFT 0.177 1.564 33.8% 0.156 6901
    QD3DT 0.217 1.550 37.5% 0.198 6856
    MUTR3D 0.270 1.494 41.1% 0.245 6018
    TripletTrack 0.268 1.504 40.0% 0.245 1144
    CC-3DT 0.410 1.274 53.8% 0.357 3334
    PF-Track-F (Ours) 0.434 1.252 53.8% 0.378 249

    PF-Track-S denotes the small-resolution setting (800×320800 \times 320) and PF-Track-F denotes the full-resolution setting (1600×6401600 \times 640). On the test split, PF-Track-F achieves 0.434 AMOTA (a 2.4-point improvement over CC-3DT and a 16.4-point improvement over MUTR3D) while reducing ID switches to 249, eliminating over 92% of ID switches relative to CC-3DT (3334) and MUTR3D (6018).

  8. Knowl 8 — Ablation Analysis of Past and Future Reasoning Modules

    data/table

    An ablation study evaluated the individual contributions of Query Refinement (QR) and Track Refinement (TR) from Past Reasoning, and Motion Prediction (Pred) and Track Extension (Ext) from Future Reasoning, on the nuScenes validation split using the small-resolution setting (800×320800 \times 320):

    Index Past Future AMOTA ↑\uparrow AMOTP ↓\downarrow IDS ↓\downarrow
    QR TR Pred Ext
    1 0.368 1.421 507
    2 ✓ 0.378 1.414 453
    3 ✓ ✓ 0.380 1.408 400
    4 ✓ 0.374 1.402 469
    5 ✓ ✓ 0.391 1.360 155
    6 ✓ ✓ ✓ ✓ 0.408 1.343 166

    Key takeaways:

    • Incorporating Past Reasoning alone (QR + TR, row 3) improves AMOTA from 0.368 to 0.380, AMOTP from 1.421 to 1.408, and reduces IDS from 507 to 400.
    • Future Reasoning with Track Extension (Pred + Ext, row 5) drastically reduces ID switches from 507 to 155 (a 69.4% decrease) and improves AMOTA to 0.391.
    • Combining both past and future reasoning modules (row 6) yields the highest overall AMOTA (0.408) and lowest AMOTP (1.343).
  9. Knowl 9 — PF-Track vs. Tracking-by-Detection Baselines on Vision Detections

    data/table

    To evaluate end-to-end tracking against modular tracking-by-detection pipelines in the camera domain, classical 3D MOT methods were paired with PETR 3D bounding box detections on the nuScenes validation split:

    Method AMOTA ↑\uparrow AMOTP ↓\downarrow IDS ↓\downarrow
    AB3DMOT 0.292 1.333 2419
    AB3DMOTΨ^\Psi 0.329 1.388 2677
    CenterPoint 0.233 1.270 2715
    CenterPointΨ^\Psi 0.383 1.329 3082
    SimpleTrack 0.320 1.295 1606
    SimpleTrackΨ^\Psi 0.402 1.324 2053
    PF-Track (Ours) 0.408 1.343 166

    $\Psi$ indicates baselines whose hyperparameters (such as association thresholds and Kalman filter parameters) were tuned for AMOTA specifically on PETR detections. While SimpleTrackΨ^\Psi achieves 0.402 AMOTA, its ID switches remain high at 2053 due to localization jitter and association fragility. In contrast, PF-Track achieves 0.408 AMOTA with only 166 ID switches (a 91.9% reduction compared to tuned SimpleTrack).

  10. Knowl 10 — Trajectory Forecasting Horizon and Feature-Based Prediction Quality

    data/table

    Future trajectory forecasting in PF-Track was evaluated for tracking utility across prediction horizons and for forecasting accuracy compared to state-based prediction models on true-positive tracks from the nuScenes validation split.

    Prediction Horizon Ablation on MOT (Small-Resolution):

    Horizon Track Extension AMOTA ↑\uparrow AMOTP ↓\downarrow IDS ↓\downarrow
    2.0s (4 frames) No 0.392 1.376 604
    2.0s (4 frames) Yes 0.402 1.342 217
    3.0s (6 frames) No 0.392 1.372 540
    3.0s (6 frames) Yes 0.402 1.340 208
    4.0s (8 frames) No 0.391 1.387 471
    4.0s (8 frames) Yes 0.408 1.343 166

    Extending prediction horizon up to 4.0 seconds (8 keyframes at 2Hz) consistently decreases ID switches when combined with track extension.

    Motion Forecasting Error at 4.0s Horizon (No HD-Maps):

    Forecasting Method ADE ↓\downarrow (@4.0s in m) FDE ↓\downarrow (@4.0s in m)
    LSTM (from 3D coordinates) 2.32 2.87
    VectorNet (from 3D coordinates) 2.01 2.48
    Constant Velocity Model 2.10 2.64
    PF-Track (Ours, from Query Features) 1.88 2.38

    PF-Track's end-to-end forecasting directly from rich learned query features achieves lower Average Displacement Error (ADE: 1.88 m) and Final Displacement Error (FDE: 2.38 m) than post-hoc trajectory models trained on discrete 3D bounding box coordinates.

Coverage note — None was omitted; all key architectural components, decoupled spatio-temporal mechanisms, loss formulations, tracking extension strategies, and empirical benchmark/ablation tables have been faithfully covered.

References

  1. 1.Adil Kaan Akan and Fatma Güney. Stretchbev: Stretching future instance prediction spatially and temporally. In ECCV, 2022. 3
  2. 2.Keni Bernardin and Rainer Stiefelhagen. Evaluating multiple object tracking performance: the clear MOT metrics. EURASIP Journal on Image and Video Processing, 2008. 6
  3. 3.Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple online and realtime tracking. In ICIP, 2016. 2
  4. 4.Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuScenes: A multi-modal dataset for autonomous driving. In CVPR, 2020. 1, 2, 5, 6
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2, 3
  6. 6.Sergio Casas, Abbas Sadat, and Raquel Urtasun. MP3: A unified model to map, perceive, predict and plan. In CVPR, 2021. 3
  7. 7.Mohamed Chaabane, Peter Zhang, J. Ross Beveridge, and Stephen O’Hara. DEFT: Detection embeddings for tracking. arXiv preprint arXiv:2102.02267, 2021. 1, 6
  8. 8.Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, and James Hays. Argoverse: 3D tracking and forecasting with rich maps. In CVPR, 2019. 1, 3, 7, 8
  9. 9.Hsu-kuang Chiu, Jie Li, Rareș Ambrusș, and Jeannette Bohg. Probabilistic 3D multi-modal, multi-object tracking for autonomous driving. In ICRA, 2021. 2
  10. 10.Patrick Dendorfer, Vladimir Yugay, Aljoša Ošep, and Laura Leal-Taixé. Quo Vadis: Is trajectory forecasting the key to-wards long-term multi-object tracking? In NeurIPS, 2022. 3
  11. 11.Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R Qi, Yin Zhou, et al. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In ICCV, 2021. 1, 3
  12. 12.Tobias Fischer, Yung-Hsu Yang, Suryansh Kumar, Min Sun, and Fisher Yu. CC-3DT: Panoramic 3D object tracking via cross-camera fusion. In CoRL, 2022. 1, 2, 6
  13. 13.Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. Deep ordinal regression network for monocular depth estimation. In CVPR, 2018. 2
  14. 14.Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid. VectorNet: Encoding HD maps and agent dynamics from vectorized representation. In CVPR, 2020. 3, 4, 7, 8
  15. 15.Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. Digging into self-supervised monocular depth estimation. In ICCV, 2019. 2
  16. 16.Junru Gu, Chenxu Hu, Tianyuan Zhang, Xuanyao Chen, Yilun Wang, Yue Wang, and Hang Zhao. ViP3D: End-to-end visual trajectory prediction via 3D agent queries. arXiv preprint arXiv:2208.01582, 2022. 1, 8
  17. 17.Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon. 3D packing for self-supervised monocular depth estimation. In CVPR, 2020. 2
  18. 18.Anthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas, Jeffrey Hawke, Vijay Badrinarayanan, Roberto Cipolla, and Alex Kendall. FIERY: Future instance prediction in bird’s-eye view from surround monocular cameras. In ICCV, 2021. 3
  19. 19.Hou-Ning Hu, Yung-Hsu Yang, Tobias Fischer, Trevor Darrell, Fisher Yu, and Min Sun. Monocular quasi-dense 3D object tracking. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. 1, 2, 6
  20. 20.Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. BEVDet: High-performance multi-camera 3D object detection in bird-eye-view. arXiv preprint arXiv:2112.11790, 2021. 1, 2
  21. 21.Boris Ivanovic, Kuan-Hui Lee, Pavel Tokmakov, Blake Wulfe, Rowan McAllister, Adrien Gaidon, and Marco Pavone. Heterogeneous-agent trajectory forecasting incorporating class uncertainty. arXiv preprint arXiv:2104.12446, 2021. 1
  22. 22.Boris Ivanovic and Marco Pavone. The trajectron: Probabilistic multi-agent trajectory modeling with dynamic spatiotemporal graphs. In ICCV, 2019. 3
  23. 23.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. PointPillars: Fast encoders for object detection from point clouds. In CVPR, 2019. 2
  24. 24.Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran Wang, Yukang Shi, Jianjian Sun, and Zeming Li. BEVDepth: Acquisition of reliable depth for multi-view 3d object detection. arXiv preprint arXiv:2206.10092, 2022. 1, 2
  25. 25.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In ECCV, 2022. 1, 2
  26. 26.Ming Liang, Bin Yang, Wenyuan Zeng, Yun Chen, Rui Hu, Sergio Casas, and Raquel Urtasun. PnPNet: End-to-end perception and prediction with tracking in the loop. In CVPR, 2020. 3
  27. 27.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In ICCV, 2017. 5
  28. 28.Yingfei Liu, Tiancai Wang, Xiangyu Zhang, and Jian Sun. PETR: Position embedding transformation for multi-view 3D object detection. In ECCV, 2022. 1, 2, 4, 6, 7
  29. 29.Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun. PETRv2: A unified framework for 3D perception from multi-camera images. arXiv preprint arXiv:2206.01256, 2022. 2
  30. 30.Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, and Bolei Zhou. Multimodal motion prediction with stacked transformers. In CVPR, 2021. 3
  31. 31.Jonathon Luiten, Tobias Fischer, and Bastian Leibe. Track to reconstruct and reconstruct to track. Robotics and Automation Letters, 2020. 1, 2
  32. 32.Wenjie Luo, Bin Yang, and Raquel Urtasun. Fast and furious: Real time end-to-end 3D detection, tracking and motion forecasting with a single convolutional net. In CVPR, 2018. 3, 8
  33. 33.Nicola Marinello, Marc Proesmans, and Luc Van Gool. TripletTrack: 3D object tracking using triplet embeddings and lstm. In CVPR, 2022. 1, 6
  34. 34.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. TrackFormer: Multi-object tracking with transformers. In CVPR, 2022. 2, 3
  35. 35.Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan, Zhengdong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, et al. Scene transformer: A unified multi-task model for behavior prediction and planning. In ICLR, 2021. 3, 4, 5
  36. 36.Ziqi Pang, Zhichao Li, and Naiyan Wang. Simpletrack: Understanding and rethinking 3D multi-object tracking. arXiv preprint arXiv:2111.09621, 2021. 1, 2, 5, 7
  37. 37.Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li, and Adrien Gaidon. Is pseudo-LiDAR needed for monocular 3D object detection? In ICCV, 2021. 2
  38. 38.Chu Peng, Wang Jiang, You Quanzeng, Ling Haibin, and Liu Zicheng. TransMOT: Spatial-temporal graph transformer for multiple object tracking. In CVPR, 2021. 2
  39. 39.Neehar Peri, Jonathon Luiten, Mengtian Li, Aljoša Ošep, and Laura Leal-Taixé. Forecasting from LiDAR via future object detection. In CVPR, 2022. 3
  40. 40.John Phillips, Julieta Martinez, Ioan Andrei Bârsan, Sergio Casas, Abbas Sadat, and Raquel Urtasun. Deep multi-task learning for joint localization, perception, and prediction. In CVPR, 2021. 3
  41. 41.Cody Reading, Ali Harakeh, Julia Chae, and Steven L. Waslander. Categorical depth distributionnetwork for monocular 3D object detection. CVPR, 2021. 2
  42. 42.Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data. In ECCV, 2020. 3
  43. 43.Samuel Scheidegger, Joachim Benjaminsson, Emil Rosenberg, Amrit Krishnan, and Karl Granström. Mono-camera 3D multi-object tracking using deep learning detections and pmbm filtering. In IEEE Intelligent Vehicles Symposium, 2018. 1, 2
  44. 44.Meet Shah, Zhiling Huang, Ankit Laddha, Matthew Langford, Blake Barber, Sidney Zhang, Carlos Vallespi-Gonzalez, and Raquel Urtasun. LiRaNet: End-to-end trajectory prediction using spatio-temporal radar fusion. arXiv preprint arXiv:2010.00731, 2020. 3
  45. 45.Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Motion transformer with global intention localization and local movement refinement. In NeurIPS, 2022. 3
  46. 46.Colton Stearns, Davis Rempe, Jie Li, Rares Ambrus, Sergey Zakharov, Vitor Guizilini, Yanchao Yang, and Leonidas J Guibas. SpOT: Spatiotemporal modeling for 3D object tracking. In ECCV, 2022. 2
  47. 47.Peize Sun, Jinkun Cao, Yi Jiang, Rufeng Zhang, Enze Xie, Zehuan Yuan, Changhu Wang, and Ping Luo. TransTrack: Multiple object tracking with transformer. arXiv preprint arXiv:2012.15460, 2020. 2
  48. 48.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020. 3
  49. 49.Siyu Tang, Mykhaylo Andriluka, Bjoern Andres, and Bernt Schiele. Multiple people tracking by lifted multicut and person re-identification. In CVPR, 2017. 2
  50. 50.Pavel Tokmakov, Jie Li, Wolfram Burgard, and Adrien Gaidon. Learning to track with object permanence. In ICCV, 2021. 2, 6
  51. 51.Qitai Wang, Yuntao Chen, Ziqi Pang, Naiyan Wang, and Zhaoxiang Zhang. Immortal tracker: Tracklet never dies. arXiv preprint arXiv:2111.13672, 2021. 2
  52. 52.Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. FCOS3D: Fully convolutional one-stage monocular 3D object detection. In ICCV, 2021. 2
  53. 53.Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. DETR3D: 3D object detection from multi-view images via 3D-to-2D queries. In CoRL, 2022. 1, 2
  54. 54.Xinshuo Weng, Boris Ivanovic, Kris Kitani, and Marco Pavone. Whose track is it anyway? Improving robustness to tracking errors with affinity-based trajectory prediction. In CVPR, 2022. 3
  55. 55.Xinshuo Weng, Boris Ivanovic, and Marco Pavone. MTP: Multi-hypothesis tracking and prediction for reduced error propagation. In IEEE Intelligent Vehicles Symposium, 2022. 3, 8
  56. 56.Xinshuo Weng, Jianren Wang, David Held, and Kris Kitani. 3D multi-object tracking: A baseline and new evaluation metrics. In IROS, 2020. 2, 6, 7
  57. 57.Xinshuo Weng, Jianren Wang, Sergey Levine, Kris Kitani, and Nicholas Rhinehart. Inverting the pose forecasting pipeline with SPF2: Sequential pointcloud forecasting for sequential pose forecasting. In CoRL, 2021. 3
  58. 58.Xinshuo Weng, Yongxin Wang, Yunze Man, and Kris M Kitani. GNN3DMOT: Graph neural network for 3D multi-object tracking with 2D-3D multi-feature learning. In CVPR, 2020. 1, 2
  59. 59.Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, et al. Argoverse 2: Next generation datasets for self-driving perception and forecasting. In NeurIPS, 2022. 1, 3
  60. 60.Nicolai Wojke, Alex Bewley, and Dietrich Paulus. Simple online and realtime tracking with a deep association metric. In ICIP, 2017. 2
  61. 61.Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. Center-based 3D object detection and tracking. In CVPR, 2021. 2, 7
  62. 62.Ye Yuan, Xinshuo Weng, Yanglan Ou, and Kris M Kitani. AgentFormer: Agent-aware transformers for socio-temporal multi-agent forecasting. In ICCV, 2021. 3
  63. 63.Jan-Nico Zaech, Alexander Liniger, Dengxin Dai, Martin Danelljan, and Luc Van Gool. Learnable online graph representations for 3D multi-object tracking. Robotics and Automation Letters, 2022. 1, 2
  64. 64.Fangao Zeng, Bin Dong, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. MOTR: End-to-end multiple-object tracking with transformer. In ECCV, 2022. 2, 3
  65. 65.Tianyuan Zhang, Xuanyao Chen, Yue Wang, Yilun Wang, and Hang Zhao. MUTR3D: A multi-camera tracking framework via 3D-to-2D queries. In CVPRW, 2022. 1, 2, 3, 6
  66. 66.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. FairMOT: On the fairness of detection and re-identification in multiple object tracking. International Journal of Computer Vision, 2021. 2
  67. 67.Yunpeng Zhang, Zheng Zhu, Wenzhao Zheng, Junjie Huang, Guan Huang, Jie Zhou, and Jiwen Lu. BEVerse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving. arXiv preprint arXiv:2205.09743, 2022. 3
  68. 68.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Tracking objects as points. In ECCV, 2020. 2, 6
  69. 69.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. In ICLR, 2021. 3

Citation

MLA
Pang, Z., et al. “Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking”. arXiv, 2023, http://arxiv.org/abs/2302.03802v2.
APA
Pang, Z., Li, J., Tokmakov, P., Chen, D., Zagoruyko, S., & Wang, Y.-X. (2023). Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking. arXiv. http://arxiv.org/abs/2302.03802v2
Chicago
Pang, Z., J. Li, P. Tokmakov, D. Chen, S. Zagoruyko, and Y.-X. Wang. 2023. “Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking”. arXiv. http://arxiv.org/abs/2302.03802v2.
Harvard
Pang, Z. et al. (2023) “Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.03802v2.
Vancouver
1. Pang Z, Li J, Tokmakov P, Chen D, Zagoruyko S, Wang Y-X (2023) Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking. arXiv

BibTeX

@article{pang2023standing,
  title = {Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking},
  author = {Pang, Ziqi and Li, Jie and Tokmakov, Pavel and Chen, Dian and Zagoruyko, Sergey and Wang, Yu-Xiong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.03802v2},
  eprint = {2302.03802}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE