UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

Sisi YouHantao YaoBing-Kun BaoChangsheng Xu

article2023CVPR76 citations

Proposes a unified multiple object tracking framework that couples detection, feature embedding, and identity association through an identity-aware feature enhancement module, creating a mutual feedback loop that improves both object localization and association accuracy across frames.

Listen

Multiple Object Tracking is essential for real-world computer vision applications, including visual surveillance, autonomous vehicles, virtual reality, and human-computer interaction. Conventional tracking systems typically follow multi-step approaches where object detection, appearance embedding extraction, and identity association operate independently. This separation creates a significant structural bottleneck: the rich historical identity and trajectory clues established during identity association cannot flow backward to help locate occluded targets or refine visual representations in future video frames.

The article demonstrates the Unified Tracking Model, a framework designed to bridge detection, embedding, and association into a mutually beneficial positive feedback loop. It evaluates this unified architecture against standard benchmarks to determine whether propagating identity-aware trajectory knowledge back into the core feature representations measurably enhances overall tracking accuracy, identity consistency, and robustness against visual occlusions.

The approach introduces an Identity-Aware Feature Enhancement mechanism comprising two complementary components: boosting attention, which reinforces features matching historical tracklets, and erasing attention, which actively suppresses distracting background noise. These enhanced representations feed the detection and embedding branches before entering a graph-matching identity association stage, while an adaptive memory bank aggregates historical representations to mitigate identity switches. The authors conducted rigorous experimental evaluations across three standard benchmark datasets—MOT16, MOT17, and MOT20—under both public and private object detection conditions.

The experimental findings show that the proposed unified model consistently outperforms existing multi-step tracking systems. Under public detection benchmarks, the model achieved Higher Order Tracking Accuracy improvements of 7.7% to 11.2% over standard tracking-by-detection baselines and improved identity association scores by up to 14.3%. Under private detection settings, it established top-tier performance, achieving an overall accuracy of 81.8% on MOT17 and outperforming comparable joint methods by 16.4% on the dense MOT20 benchmark. Furthermore, ablation analyses demonstrated that the model maintains significantly higher tracking success rates when objects experience heavy visual occlusion exceeding 50%.

These results confirm that closing the loop between trajectory matching and lower-level feature extraction substantially reduces tracking errors, false negatives, and identity confusion without requiring completely separate tracking pipelines. For operational decision-makers in robotics and automated surveillance, adopting unified feedback architectures can noticeably improve system reliability and situational awareness in crowded or visually complex environments. Organizations seeking higher tracking fidelity should prioritize integrated multi-task architectures over disconnected, multi-stage pipelines.

To build upon these findings, future development should optimize the visual embedding module to mitigate training constraints associated with small batch sizes during end-to-end optimization. Acknowledging that the current evaluations primarily focus on pedestrian tracking datasets under controlled benchmark conditions, stakeholders should pilot and validate the model across target operational domains—such as varying weather conditions or specialized vehicle tracking—before broad deployment.

Cover for UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

Abstract

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Separate Detection and Embedding
  • 2.2. Joint Detection and Embedding
  • 3. Methodology
  • 3.1. Identity-Aware Feature Enhancement
  • 3.2. Unified Tracking Model
  • 4. Experiments
  • 4.1. Benchmark Evaluation
  • 4.2. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Unified Tracking Model Architecture

    model/method

    The Unified Tracking Model (UTM) integrates object detection, feature embedding, and identity association into a closed-loop framework with mutual reinforcement.

    Given the tt-th video frame ItI_t, a convolutional backbone extracts raw backbone features FtF_t. The architecture proceeds through five coordinated modules:

    1. Candidate Proposal Generation: Candidate proposals Pt={Pt,1,…,Pt,n}P_t = \{P_{t,1}, \dots, P_{t,n}\} are formed by combining public detector proposals of frame ItI_t with the tracked object bounding boxes from the previous frame t−1t-1.
    2. Identity-Aware Feature Enhancement (IAFE): Uses the history tracklet feature set Et−1={Et−1,1,…,Et−1,m}E_{t-1} = \{E_{t-1,1}, \dots, E_{t-1,m}\} to enhance the backbone feature Ft,∗F_{t,*} corresponding to candidate proposal Pt,∗P_{t,*}, generating enhanced feature maps F~t\tilde{F}_t.
    3. Detection Branch: Built on Faster R-CNN, a bounding box regression head refines candidate proposals PtP_t into candidate boxes Bt={Bt,1,…,Bt,n}B_t = \{B_{t,1}, \dots, B_{t,n}\}, and a classification head predicts object confidence scores using F~t\tilde{F}_t.
    4. Embedding Branch: Applies RoI-Align and a convolutional sub-network on F~t\tilde{F}_t and BtB_t to produce identity appearance embeddings F^t={F^t,1,…,F^t,n}\hat{F}_t = \{\hat{F}_{t,1}, \dots, \hat{F}_{t,n}\}.
    5. Identity Association Branch: Formulates cross-frame identity association as graph matching between a candidate detection graph and a history tracklet graph, incorporating high-order context and cross-graph message passing.
    6. Memory Aggregation Module: Maintains a memory bank of recent embeddings and uses a learnable module to update tracklet features Et={Et,1,…,Et,m}E_t = \{E_{t,1}, \dots, E_{t,m}\}, which are passed to the IAFE module in frame t+1t+1.
  2. Knowl 2 — Identity-Aware Feature Enhancement Formulation

    equation

    The Identity-Aware Feature Enhancement (IAFE) module updates the candidate proposal backbone feature Ft,∗∈RC×H×WF_{t,*} \in \mathbb{R}^{C \times H \times W} of candidate proposal Pt,∗P_{t,*} using the tracklet feature set Et−1={Et−1,1,…,Et−1,m}E_{t-1} = \{E_{t-1,1}, \dots, E_{t-1,m}\} via combined boosting and erasing attention:

    F~t,∗=Ft,∗⊕[f(Ft,∗,Et−1)⊖g(Ft,∗,Et−1)]\tilde{F}_{t,*} = F_{t,*} \oplus \left[ f(F_{t,*}, E_{t-1}) \ominus g(F_{t,*}, E_{t-1}) \right]

    where ⊕\oplus and ⊖\ominus denote element-wise addition and subtraction, f(⋅)f(\cdot) is the Identity-Aware Boosting Attention (IABA), and g(⋅)g(\cdot) is the Identity-Aware Erasing Attention (IAEA).

    Identity-Aware Boosting Attention (IABA): For spatial index i∈[1,HW]i \in [1, HW], IABA enhances regions consistent with historical tracklets:

    f(Ft,∗i,Et−1)=∑k=1mλ∗,k∑j=1HWh(Ft,∗i,Et−1,kj)ρ(Et−1,kj)∑j=1HWh(Ft,∗i,Et−1,kj)f(F_{t,*}^i, E_{t-1}) = \sum_{k=1}^m \lambda_{*,k} \frac{\sum_{j=1}^{HW} h(F_{t,*}^i, E_{t-1,k}^j) \rho(E_{t-1,k}^j)}{\sum_{j=1}^{HW} h(F_{t,*}^i, E_{t-1,k}^j)}

    where Ft,∗i,Et−1,kj∈RCF_{t,*}^i, E_{t-1,k}^j \in \mathbb{R}^C, h(xi,xj)=exp⁡(ψ(xi)Tφ(xj))h(x_i, x_j) = \exp(\psi(x_i)^T \varphi(x_j)), ψ,φ,ρ\psi, \varphi, \rho are 1×11 \times 1 convolution layers, and λ∗,k\lambda_{*,k} is an indicator function:

    λ∗,k={1if IoU(Pt,∗,Bt−1,k)>λiou0otherwise\lambda_{*,k} = \begin{cases} 1 & \text{if } \text{IoU}(P_{t,*}, B_{t-1,k}) > \lambda_{\text{iou}} \\ 0 & \text{otherwise} \end{cases}

    where Bt−1,kB_{t-1,k} is the last bounding box of the kk-th tracklet and λiou\lambda_{\text{iou}} is a geometric threshold (set to 0.70.7).

    Identity-Aware Erasing Attention (IAEA): IAEA suppresses background and distracted regions:

    g(Ft,∗,Et−1)=∑k=1mλ∗,kFt,∗⊙R(Ft,∗,ϕ(Et−1,k))⊙M(Ft,∗,ϕ(Et−1,k))g(F_{t,*}, E_{t-1}) = \sum_{k=1}^m \lambda_{*,k} F_{t,*} \odot \mathcal{R}(F_{t,*}, \phi(E_{t-1,k})) \odot \mathcal{M}(F_{t,*}, \phi(E_{t-1,k}))

    where ϕ\phi denotes spatial average pooling, ⊙\odot is the Hadamard product, and R(Ft,∗[i,j],ϕ(Et−1,k))=(Ft,∗[i,j])Tϕ(Et−1,k)\mathcal{R}(F_{t,*}[i, j], \phi(E_{t-1,k})) = (F_{t,*}[i, j])^T \phi(E_{t-1,k}) computes dot-product spatial correlation. M\mathcal{M} is a binary mask generated by a 3×33 \times 3 sliding Block Binarization Layer (stride 1) that locates the continuous block with the maximal sum of correlation values, setting its values to 00 and other spatial positions to 11.

  3. Knowl 3 — Cross-Graph Message Passing and High-Order Identity Association

    model/method

    Identity association in the Unified Tracking Model (UTM) is cast as a quadratic graph matching problem between candidate detection graph GD=(VD,ED)G_D = (V_D, E_D) and tracklet graph GT=(VT,ET)G_T = (V_T, E_T).

    • Detection Graph: Nodes VD={(Bt,i;F^t,i)}i=1nV_D = \{(B_{t,i}; \hat{F}_{t,i})\}_{i=1}^n with bounding boxes Bt,iB_{t,i} and embeddings F^t,i=E(F~t,Bt,i)\hat{F}_{t,i} = \mathcal{E}(\tilde{F}_t, B_{t,i}). Directed edges are represented by feature concatenation ED={[F^t,i1,F^t,i2]}E_D = \{[\hat{F}_{t,i_1}, \hat{F}_{t,i_2}]\}.
    • Tracklet Graph: Nodes VT={(Bt−1,j;Et−1,j)}j=1mV_T = \{(B_{t-1,j}; E_{t-1,j})\}_{j=1}^m with last boxes Bt−1,jB_{t-1,j} and tracklet features Et−1,jE_{t-1,j}. Edges ET={[Et−1,j1,Et−1,j2]}E_T = \{[E_{t-1,j_1}, E_{t-1,j_2}]\}.

    Cross-Graph Message Passing: With initial node features F^t,i(0)=F^t,i\hat{F}_{t,i}^{(0)} = \hat{F}_{t,i} and Et−1,j(0)=Et−1,jE_{t-1,j}^{(0)} = E_{t-1,j}, node representations are iteratively updated across L=3L=3 steps via Type 3 aggregation:

    F^t,i(l+1)=Nv(F^t,i(l)+∑j=1mwi,j(l)Et−1,j(l))\hat{F}_{t,i}^{(l+1)} = \mathcal{N}_v\left( \hat{F}_{t,i}^{(l)} + \sum_{j=1}^m w_{i,j}^{(l)} E_{t-1,j}^{(l)} \right)

    where Nv\mathcal{N}_v is a multilayer perceptron, and weight wi,j(l)=cos⁡(F^t,i(l),Et−1,j(l))+IoU(Bt,i,Bt−1,j)w_{i,j}^{(l)} = \cos(\hat{F}_{t,i}^{(l)}, E_{t-1,j}^{(l)}) + \text{IoU}(B_{t,i}, B_{t-1,j}) combines appearance cosine similarity and bounding box geometric overlap. Symmetric updates are executed on GTG_T.

    Matching Optimization: An affinity matrix M∈Rnm×nmM \in \mathbb{R}^{nm \times nm} is constructed using first-order node-to-node cosine similarities and second-order edge-to-edge cosine similarities. The binary assignment permutation matrix Y∗∈{0,1}n×mY^* \in \{0, 1\}^{n \times m} is solved by:

    Y∗=arg⁡max⁡YYTMYY^* = \arg\max_Y Y^T M Y

    Matches with affinity lower than threshold γ=0.6\gamma = 0.6 are discarded. The association branch is trained using weighted binary cross-entropy loss.

  4. Knowl 4 — Tracklet Memory Aggregation Module

    model/method

    To prevent noisy embeddings caused by occlusion and identity switches from corrupting tracklet representations, UTM incorporates a memory aggregation module θ\theta.

    A memory bank stores the previous η=30\eta = 30 historical identity embeddings for each target jj:

    Ft−1,jm={F^t−η,j,…,F^t−1,j}F_{t-1,j}^m = \{\hat{F}_{t-\eta,j}, \dots, \hat{F}_{t-1,j}\}

    When a new detection embedding F^t,j\hat{F}_{t,j} is matched, the updated tracklet feature Et,jE_{t,j} is computed by adaptively selecting and weighting historical embeddings:

    Et,j=θ(Ft−1,jm,F^t,j)=ReLU(W[Ft−1,jm,F^t,j]+b)E_{t,j} = \theta(F_{t-1,j}^m, \hat{F}_{t,j}) = \text{ReLU}(\mathbf{W} [F_{t-1,j}^m, \hat{F}_{t,j}] + \mathbf{b})

    where W\mathbf{W} and b\mathbf{b} parameterize a learnable linear layer. The output Et,jE_{t,j} serves as the robust identity-aware representation utilized by the IAFE module in frame t+1t+1.

  5. Knowl 5 — Three-Stage Optimization Strategy for UTM

    algorithm

    To ensure fast and stable convergence across interconnected branches, the Unified Tracking Model (UTM) is trained in three sequential stages.

    Input: Video frames with ground-truth bounding boxes and identity labels
    Output: Optimized UTM tracking model
    Stage 1: Joint Training of Detection, Embedding, and Memory Modules
    Initialize Faster R-CNN backbone (pre-trained on COCO or CrowdHuman)
    Initialize embedding branch (pre-trained on Market-1501 and CUHK03)
    for each training iteration in MOT datasets do
        Compute enhanced features F~t\tilde{F}_t via IAFE using tracklet memory Et−1E_{t-1}
        Predict bounding box displacements and classification logits in detection branch
        Extract appearance embeddings F^t\hat{F}_t and update tracklet features EtE_t
        Compute detection loss: Ldet=LL1+LCE\mathcal{L}_{\text{det}} = \mathcal{L}_{\text{L1}} + \mathcal{L}_{\text{CE}}
        Compute embedding and memory loss: Lemb=LCE+Ltriplet\mathcal{L}_{\text{emb}} = \mathcal{L}_{\text{CE}} + \mathcal{L}_{\text{triplet}}
        Update detection, embedding, and memory weights using Adam optimizer
    end for
    Stage 2: Training Identity Association Branch
    Freeze weights of detection branch, embedding branch, and memory aggregation
    for each training iteration in MOT datasets do
        Construct graphs GDG_D and GTG_T from extracted boxes and embeddings
        Execute L=3L=3 message passing iterations
        Compute cross-graph affinity matrix MM and solve assignment Y∗Y^*
        Compute weighted binary cross-entropy loss Lassoc\mathcal{L}_{\text{assoc}}
        Update graph matching module weights
    end for
    Stage 3: End-to-End Fine-Tuning
    Unfreeze all branches
    for each training iteration up to 30 epochs do
        Execute full forward pass through IAFE, detection, embedding, association, and memory
        Compute combined loss Ltotal=Ldet+Lemb+Lassoc\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{det}} + \mathcal{L}_{\text{emb}} + \mathcal{L}_{\text{assoc}}
        Update all network parameters with learning rate 0.0020.002 (decay factor 0.50.5 every 3 epochs)
    end for
    return Optimized UTM model
  6. Knowl 6 — Performance Comparison on MOT Benchmarks with Public Detections

    data/table

    Tracking performance on the public detection benchmarks MOT16, MOT17, and MOT20 using Tracktor-refined public detections. Bold numbers indicate the best performance on the benchmark. Offline methods are designated with (O) and post-processing methods with *.

    Methods Refined MOTA ↑\uparrow HOTA ↑\uparrow IDF1 ↑\uparrow FP ↓\downarrow FN ↓\downarrow IDS ↓\downarrow
    MOT16
    MPNT(O) Tracktor 58.6 48.9 61.7 4,949 70,252 354
    LPC(O) Tracktor 58.8 51.7 67.6 6,167 68,432 435
    GMTsI(O) Tracktor 61.1 51.2 66.6 3,891 66,550 503
    DeepMOT Tracktor 54.8 42.2 53.4 2,955 78,765 645
    GMT Tracktor 55.9 48.9 63.9 2,371 77,545 531
    Tracktor Tracktor 56.2 44.6 54.9 2,394 76,844 617
    ArTIST Tracktor 56.6 - 57.8 3,532 75,031 519
    LifTsI Tracktor 57.5 49.6 64.1 4,249 72,868 335
    TADAM Tracktor 59.1 - 59.5 2,540 71,542 529
    TMOH* Tracktor 63.2 50.7 63.5 3,122 63,376 635
    UTM Tracktor 63.8 53.1 67.1 8,328 57,269 428
    MOT17
    LifTsI(O) Tracktor 58.2 50.7 65.2 16,850 217,944 1,022
    MPNT(O) Tracktor 58.8 49.0 61.7 17,413 213,594 1,185
    LPC(O) Tracktor 59.0 51.7 66.8 23,102 206,947 1,122
    GMTsI(O) Tracktor 59.0 51.1 65.9 20,395 209,553 1,105
    GMT Tracktor 56.2 49.1 63.8 8,719 236,541 1,778
    Tracktor Tracktor 56.3 44.8 55.1 8,866 235,449 1,987
    ArTIST Tracktor 56.7 - 57.5 12,353 230,437 1,756
    TADAM Tracktor 59.7 - 58.7 9,676 216,029 1,930
    TMOH* Tracktor 62.1 50.4 62.8 10,951 201,195 1,897
    UTM Tracktor 63.5 52.5 65.1 33,683 170,352 1,686
    MOT20
    LPC(O) Tracktor 56.3 49.0 62.5 11,726 213,056 1,562
    MPNT(O) Tracktor 57.6 46.8 59.1 16,953 210,384 1,210
    Tracktor Tracktor 52.6 42.1 52.7 6,930 236,680 1,648
    ArTIST Tracktor 53.6 - 51.0 7,765 230,576 1,531
    TADAM Tracktor 56.6 - 51.6 38,407 182,520 2,690
    TMOH* Tracktor 60.1 48.9 61.2 38,043 165,899 2,342
    UTM Tracktor 64.4 53.3 65.9 82,726 98,974 2,592

    UTM outperforms both SDE (Tracktor: +8.5%+8.5\% HOTA on MOT16, +7.7%+7.7\% on MOT17, +11.2%+11.2\% on MOT20) and JDE (TADAM: +7.6%+7.6\% IDF1 on MOT16, +6.4%+6.4\% on MOT17, +14.3%+14.3\% on MOT20) paradigms by bridging detection, embedding, and tracking via IAFE.

  7. Knowl 7 — Performance Comparison on MOT Challenge Private Detection Benchmarks

    data/table

    Comparison of UTM with state-of-the-art MOT methods under the private detection setting on MOT16, MOT17, and MOT20 benchmarks. Bold entries indicate best results.

    Methods Detector MOTA ↑\uparrow HOTA ↑\uparrow IDF1 ↑\uparrow FP ↓\downarrow FN ↓\downarrow IDS ↓\downarrow
    MOT16
    FairMOT CenterNet 75.7 61.6 75.3 16,163 27,442 621
    GRTU CenterNet 76.5 62.6 75.9 11,438 30,866 584
    TLR CenterNet 76.6 61.0 74.3 10,860 30,756 979
    UTM FRCNN 81.1 64.1 79.0 11,722 22,367 440
    MOT17
    FairMOT CenterNet 73.7 59.3 72.3 27,507 117,477 3,303
    PermaTrack CenterNet 73.8 55.5 68.9 28,998 115,104 3,699
    GRTU CenterNet 74.9 62.0 75.0 32,007 107,616 1,812
    TLR CenterNet 76.5 60.7 73.6 29,808 99,510 3,369
    MAA FRCNN 79.4 62.0 75.9 37,320 77,661 1,452
    ByteTrack YOLOX 80.3 63.1 77.3 25,491 83,721 2,196
    UTM FRCNN 81.8 64.0 78.7 25,077 76,298 1,431
    MOT20
    FairMOT CenterNet 61.8 54.6 67.3 103,440 88,901 5,243
    MAA FRCNN 73.9 57.3 71.2 24,942 108,744 1,331
    ReMOT CenterNet 77.4 61.2 73.1 28,351 86,659 1,789
    ByteTrack YOLOX 77.8 61.3 75.2 26,249 87,594 1,223
    UTM FRCNN 78.2 62.5 76.9 29,964 81,516 1,228

    UTM achieves 81.1%81.1\% MOTA on MOT16 (+4.5%+4.5\% over prior state-of-the-art), 81.8%81.8\% MOTA on MOT17, and 78.2%78.2\% MOTA on MOT20 (a +16.4%+16.4\% gain over CenterNet-based FairMOT).

  8. Knowl 8 — Ablation of Architectural Paradigms and IAFE Components

    empirical result

    Ablation experiments conducted on the MOT16 validation dataset evaluate the contributions of tracking paradigms and individual components within UTM:

    Comparison of Tracking Paradigms:

    • (a) SDE Baseline (Tracktor + ResNet-101 + Hungarian): MOTA 62.5%62.5\%, IDF1 67.4%67.4\%, HOTA 59.4%59.4\%.
    • (b) JDE Baseline (Joint Detection & Embedding + Hungarian): MOTA 62.6%62.6\%, IDF1 64.0%64.0\%, HOTA 58.5%58.5\%.
    • (c) Proposed UTM: MOTA 64.5%64.5\%, IDF1 73.1%73.1\%, HOTA 63.8%63.8\%.

    UTM yields improvements of +2.0%+2.0\% MOTA, +5.7%+5.7\% IDF1, and +4.4%+4.4\% HOTA over SDE, and +1.9%+1.9\% MOTA, +9.1%+9.1\% IDF1, and +5.3%+5.3\% HOTA over JDE.

    Component Ablation in UTM:

    • Without IAFE (w/o IAFE): MOTA drops to 62.6%62.6\%, IDF1 to 65.9%65.9\%, HOTA to 59.8%59.8\%, and false negatives (FN) increase from 38,00238,002 to 40,26240,262.
    • Without Boosting Attention (w/o IABA): MOTA is 63.9%63.9\%, IDF1 is 70.6%70.6\%, HOTA is 62.2%62.2\%, IDS rises to 518518.
    • Without Erasing Attention (w/o IAEA): MOTA is 63.6%63.6\%, IDF1 is 67.7%67.7\%, HOTA is 61.5%61.5\%.
    • Without Memory Aggregation (w/o Memory, using simple average pooling): MOTA is 64.2%64.2\%, IDF1 is 72.7%72.7\%, HOTA is 63.5%63.5\%.
    • Full UTM: MOTA 64.5%64.5\%, IDF1 73.1%73.1\%, HOTA 63.8%63.8\%, MT 186186, ML 117117, FP 878878, FN 38,00238,002, IDS 285285.
  9. Knowl 9 — Impact of Graph Matching Hyperparameters and Node Aggregation Rules

    empirical result

    Ablation experiments on MOT16 validation establish optimal structural and hyperparameter configurations for identity association in UTM:

    1. Matching Algorithm: High-order Graph Matching (GM) achieves MOTA 64.5%64.5\%, IDF1 73.1%73.1\%, and HOTA 63.8%63.8\%, outperforming the traditional Hungarian algorithm (MOTA 63.7%63.7\%, IDF1 68.0%68.0\%, HOTA 61.3%61.3\%) by +5.1%+5.1\% in IDF1 due to second-order edge similarity modeling.
    2. Node Aggregation Rules in Cross-Graph Message Passing:
      • Type 1 (Identity): MOTA 63.4%63.4\%, IDF1 63.8%63.8\%, HOTA 58.4%58.4\%, IDS 496496.
      • Type 2 (Uniform Average): MOTA 63.6%63.6\%, IDF1 66.7%66.7\%, HOTA 60.3%60.3\%, IDS 203203.
      • Type 3 (Similarity-Weighted Sum): MOTA 64.5%64.5\%, IDF1 73.1%73.1\%, HOTA 63.8%63.8\%, IDS 285285.
    3. Message Passing Steps (LL): Tracking accuracy improves monotonically as LL increases from 00 to 33 and saturates thereafter; L=3L = 3 is chosen.
    4. Geometric Threshold (λiou\lambda_{\text{iou}}): IDF1 and HOTA peak at λiou=0.7\lambda_{\text{iou}} = 0.7, decreasing for lower or higher thresholds.
    5. Affinity Matching Threshold (γ\gamma): Performance increases over γ∈[0.1,0.6]\gamma \in [0.1, 0.6] and degrades for γ>0.6\gamma > 0.6; optimal setting is γ=0.6\gamma = 0.6.
    6. Memory Length (η\eta): Metric performance steadily increases up to η=30\eta = 30 historical frames.
  10. Knowl 10 — End-to-End Metric Learning Sensitivity to Small Mini-Batch Size

    limitation

    Due to the memory footprint of concurrently running the backbone, IAFE attention modules, detection head, embedding head, memory aggregation bank, and quadratic graph matching in an end-to-end setting, the training batch size is constrained to a small mini-batch size (mini-batch size of 2 on an RTX 3090 GPU). This limited mini-batch size restricts the diversity of positive and negative identity pairings available per batch, dampening the optimization efficiency of triplet loss metric learning in the embedding branch.

Coverage note — None was omitted; all core architectural modules (IAFE, boosting/erasing attention, detection/embedding heads, graph matching, memory aggregation), equations, algorithms, empirical benchmarks, ablations, and stated limitations are fully represented.

References

  1. 1.Seung-Hwan Bae and Kuk-Jin Yoon. Confidence-based data association and discriminative deep appearance learning for robust online multi-object tracking. PAMI, 40(3):595–610, 2017.
  2. 2.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In ICCV, pages 941– 951, 2019.
  3. 3.Keni Bernardin and Rainer Stiefelhagen. Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008:1– 10, 2008.
  4. 4.Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft. Simple online and realtime tracking. In ICIP, pages 3464–3468, 2016.
  5. 5.Guillem Braso and Laura Leal-Taix ´ e. Learning a neural ´ solver for multiple object tracking. In CVPR, pages 6247– 6257, 2020.
  6. 6.Peng Chu and Haibin Ling. Famnet: Joint learning of feature, affinity and multi-dimensional assignment for online multiple object tracking. In ICCV, pages 6172–6181, 2019.
  7. 7.Qi Chu, Wanli Ouyang, Hongsheng Li, Xiaogang Wang, Bin Liu, and Nenghai Yu. Online multi-object tracking using cnn-based single object tracker with spatial-temporal attention mechanism. In ICCV, pages 4836–4845, 2017.
  8. 8.Peng Dai, Renliang Weng, Wongun Choi, Changshui Zhang, Zhangping He, and Wei Ding. Learning a proposal classifier for multiple object tracking. In CVPR, pages 2443–2452, 2021.
  9. 9.Patrick Dendorfer, Hamid Rezatofighi, Anton Milan, Javen Shi, Daniel Cremers, Ian Reid, Stefan Roth, Konrad Schindler, and Laura Leal-Taixe. Mot20: A benchmark ´ for multi object tracking in crowded scenes. arXiv preprint arXiv:2003.09003, 2020.
  10. 10.Caglayan Dicle, Octavia I Camps, and Mario Sznaier. The way they move: Tracking multiple targets with similar appearance. In ICCV, pages 2304–2311, 2013.
  11. 11.Piotr Dollar, Ron Appel, Serge Belongie, and Pietro Per- ´ ona. Fast feature pyramids for object detection. PAMI, 36(8):1532–1545, 2014.
  12. 12.Andreas Ess, Bastian Leibe, and Luc Van Gool. Depth and appearance for mobile scene analysis. In ICCV, pages 1–8, 2007.
  13. 13.Pedro F. Felzenszwalb, Ross B. Girshick, David A. McAllester, and Deva Ramanan. Object detection with discriminatively trained part-based models. PAMI, 32(9):1627– 1645, 2010.
  14. 14.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021.
  15. 15.Song Guo, Jingya Wang, Xinchao Wang, and Dacheng Tao. Online multiple object tracking with cross-task synergy. In CVPR, pages 8136–8145, 2021.
  16. 16.Jiawei He, Zehao Huang, Naiyan Wang, and Zhaoxiang Zhang. Learnable graph matching: Incorporating graph partitioning with deep feature learning for multiple object tracking. In CVPR, pages 5299–5309, 2021.
  17. 17.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- ´ shick. Mask r-cnn. In ICCV, pages 2961–2969, 2017.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  19. 19.Andrea Hornakova, Roberto Henschel, Bodo Rosenhahn, and Paul Swoboda. Lifted disjoint paths with application in multiple object tracking. In ICML, pages 4364–4375, 2020.
  20. 20.Andrea Hornakova, Timo Kaiser, Paul Swoboda, Michal Ro- ˇ linek, Bodo Rosenhahn, and Roberto Henschel. Making higher order mot scalable: An efficient approximate solver for lifted disjoint paths. In ICCV, pages 6330–6340, 2021.
  21. 21.Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Temporal complementary learning for video person re-identification. In ECCV, pages 388–405, 2020.
  22. 22.Rudolph Emil Kalman. A new approach to linear filtering and prediction problems. ASME Journal of Basic Engineering, 82(1):35–45, 1960.
  23. 23.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  24. 24.Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955.
  25. 25.Laura Leal-Taixe, Cristian Canton-Ferrer, and Konrad ´ Schindler. Learning by tracking: Siamese cnn for robust target association. In CVPRW, pages 33–40, 2016.
  26. 26.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deepreid: Deep filter pairing neural network for person reidentification. In CVPR, pages 152–159, 2014.
  27. 27.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, ´ Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, pages 2117–2125, 2017.
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In Proceedings of the ECCV, pages 740–755, 2014.
  29. 29.Zhichao Lu, Vivek Rathod, Ronny Votel, and Jonathan Huang. Retinatrack: Online single stage joint detection and tracking. In CVPR, pages 14668–14678, 2020.
  30. 30.Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixe, and Bastian Leibe. ´ Hota: A higher order metric for evaluating multi-object tracking. IJCV, pages 1–31, 2020.
  31. 31.Anton Milan, Laura Leal-Taixe, Ian Reid, Stefan Roth, and ´ Konrad Schindler. Mot16: A benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831, 2016.
  32. 32.Bo Pang, Yizhuo Li, Yifan Zhang, Muchen Li, and Cewu Lu. Tubetk: Adopting tubes to track multi-object in a one-step training model. In CVPR, pages 6308–6318, 2020.
  33. 33.Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu. Chained-tracker: Chaining paired attentive regression results for end-to-end joint multiple-object detection and tracking. In ECCV, pages 145–161, 2020.
  34. 34.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
  35. 35.Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. PAMI, 39(6):1137–1149, 2017.
  36. 36.Weihong Ren, Xinchao Wang, Jiandong Tian, Yandong Tang, and Antoni B Chan. Tracking-by-counting: Using network flows on crowd density maps for tracking multiple targets. TIP, 30:1439–1452, 2020.
  37. 37.Ergys Ristani, Francesco Solera, Roger Zou, Rita Cucchiara, and Carlo Tomasi. Performance measures and a data set for multi-target, multi-camera tracking. In ECCVW, pages 17– 35, 2016.
  38. 38.Ergys Ristani and Carlo Tomasi. Features for multi-target multi-camera tracking and re-identification. In CVPR, pages 6036–6046, 2018.
  39. 39.Fatemeh Saleh, Sadegh Aliakbarian, Hamid Rezatofighi, Mathieu Salzmann, and Stephen Gould. Probabilistic tracklet scoring and inpainting for multiple object tracking. In CVPR, pages 14329–14339, 2021.
  40. 40.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123, 2018.
  41. 41.Daniel Stadler and Jurgen Beyerer. Improving multiple pedestrian tracking by track management and occlusion handling. In CVPR, pages 10958–10967, 2021.
  42. 42.Daniel Stadler and J¨urgen Beyerer. Modelling ambiguous assignments for multi-person tracking in crowds. In WACV, pages 133–142, 2022.
  43. 43.Siyu Tang, Mykhaylo Andriluka, Bjoern Andres, and Bernt Schiele. Multiple people tracking by lifted multicut and person re-identification. In CVPR, pages 3539–3548, 2017.
  44. 44.Pavel Tokmakov, Jie Li, Wolfram Burgard, and Adrien Gaidon. Learning to track with object permanence. In ICCV, pages 10860–10869, 2021.
  45. 45.Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe. Mots: Multi-object tracking and segmentation. In CVPR, pages 7942–7951, 2019.
  46. 46.Xingyu Wan, Jiakai Cao, Sanping Zhou, Jinjun Wang, and Nanning Zheng. Tracking beyond detection: Learning a global response map for end-to-end multi-object tracking. TIP, 30:8222–8235, 2021.
  47. 47.Qiang Wang, Yun Zheng, Pan Pan, and Yinghui Xu. Multiple object tracking with correlation learning. In CVPR, pages 3876–3886, 2021.
  48. 48.Shuai Wang, Hao Sheng, Yang Zhang, Yubin Wu, and Zhang Xiong. A general recurrent tracking framework without real data. In ICCV, pages 13219–13228, 2021.
  49. 49.Yongxin Wang, Kris Kitani, and Xinshuo Weng. Joint object detection and multi-object tracking with graph neural networks. In ICRA, pages 13708–13715. IEEE, 2021.
  50. 50.Zhongdao Wang, Liang Zheng, Yixuan Liu, and Shengjin Wang. Towards real-time multi-object tracking. In ECCV, pages 107–122, 2020.
  51. 51.Xinshuo Weng, Yongxin Wang, Yunze Man, and Kris M Kitani. Gnn3dmot: Graph neural network for 3d multi-object tracking with 2d-3d multi-feature learning. In CVPR, pages 6499–6508, 2020.
  52. 52.Nicolai Wojke, Alex Bewley, and Dietrich Paulus. Simple online and realtime tracking with a deep association metric. In ICIP, pages 3645–3649, 2017.
  53. 53.Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang, Ming Yang, and Junsong Yuan. Track to detect and segment: An online multi-object tracker. In CVPR, pages 12352–12361, 2021.
  54. 54.Yihong Xu, Aljosa Osep, Yutong Ban, Radu Horaud, Laura Leal-Taixe, and Xavier Alameda-Pineda. How to train your ´ deep multi-object tracker. In CVPR, pages 6787–6796, 2020.
  55. 55.Bo Yang, Chang Huang, and Ram Nevatia. Learning affinities and dependencies for multi-target tracking using a crf model. In CVPR, pages 1233–1240, 2011.
  56. 56.Fan Yang, Xin Chang, Sakriani Sakti, Yang Wu, and Satoshi Nakamura. Remot: A model-agnostic refinement for multiple object tracking. Image and Vision Computing, 106:104091, 2021.
  57. 57.Junbo Yin, Wenguan Wang, Qinghao Meng, Ruigang Yang, and Jianbing Shen. A unified object motion and affinity model for online multi-object tracking. In CVPR, pages 6768–6777, 2020.
  58. 58.Sisi You, Hantao Yao, and Changsheng Xu. Multi-target multi-camera tracking with optical-based pose association. CSVT, 31(8):3105–3117, 2021.
  59. 59.Sisi You, Hantao Yao, and Changsheng Xu. Multiobject tracking with spatial-temporal topology-based detector. CSVT, 32(5):3023–3035, 2022.
  60. 60.Andrei Zanfir and Cristian Sminchisescu. Deep learning of graph matching. In CVPR, pages 2684–2693, 2018.
  61. 61.Li Zhang, Yuan Li, and Ramakant Nevatia. Global data association for multi-object tracking using network flows. In CVPR, pages 1–8, 2008.
  62. 62.Yang Zhang, Hao Sheng, Yubin Wu, Shuai Wang, Weifeng Lyu, Wei Ke, and Zhang Xiong. Long-term tracking with deep tracklet association. TIP, 29:6694–6706, 2020.
  63. 63.Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. Bytetrack: Multi-object tracking by associating every detection box. In ECCV, pages 1–21. Springer, 2022.
  64. 64.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and reidentification in multiple object tracking. IJCV, pages 1–19, 2021.
  65. 65.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. In ICCV, pages 1116–1124, 2015.
  66. 66.Xingyi Zhou, Dequan Wang, and Philipp Kr¨ahenb¨uhl. Objects as points. arXiv preprint arXiv:1904.07850, 2019.
  67. 67.Xingyi Zhou, Dequan Wang, and Philipp Kr¨ahenb¨uhl. Objects as points. arXiv preprint arXiv:1904.07850, 2019.
  68. 68.Ji Zhu, Hua Yang, Nian Liu, Minyoung Kim, Wenjun Zhang, and Ming-Hsuan Yang. Online multi-object tracking with dual matching attention networks. In ECCV, pages 366–382, 2018.

Citation

MLA
You, S., et al. “UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 21876–86, https://doi.org/10.1109/CVPR52729.2023.02095.
APA
You, S., Yao, H., Bao, B.-. kun ., & Xu, C. (2023). UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 21876–21886. https://doi.org/10.1109/CVPR52729.2023.02095
Chicago
You, S., H. Yao, B.-. kun . Bao, and C. Xu. 2023. “UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 21876–86. https://doi.org/10.1109/CVPR52729.2023.02095.
Harvard
You, S. et al. (2023) “UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 21876–21886. Available at: https://doi.org/10.1109/CVPR52729.2023.02095.
Vancouver
1. You S, Yao H, Bao B-kun, Xu C (2023) UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 21876–21886

BibTeX

@inproceedings{You_2023, title={UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement}, url={http://dx.doi.org/10.1109/CVPR52729.2023.02095}, DOI={10.1109/cvpr52729.2023.02095}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={You, Sisi and Yao, Hantao and Bao, Bing-kun and Xu, Changsheng}, year={2023}, month=June, pages={21876–21886} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE