Large-Scale Pre-training for Person Re-identification with Noisy Labels

Dengpan FuDongdong ChenHao YangJianmin BaoLu YuanLei ZhangHouqiang LiFang WenDong Chen

article2022CVPR74 citations

Introduces a scalable pre-training framework that learns transferable person re-identification representations directly from uncurated video tracklets by combining prototype-based label rectification with label-guided contrastive learning on a ten-million-image noisy dataset.

Listen

Person re-identification systems, which match individuals across different camera views, face major deployment bottlenecks because manually annotating thousands of identities across complex video environments is costly, time-consuming, and difficult to scale. While generic pre-trained visual models offer a starting point, they suffer from a substantial domain gap when applied to person-focused surveillance tasks. The article addresses this challenge by evaluating whether scalable model pre-training can learn directly from raw, uncurated street-view videos without any human annotation effort.

The article demonstrates an automated data generation and training strategy. First, the authors built a massive benchmark called LUPerson-NL by applying an automated multi-object tracking algorithm and pose estimation filtering to raw video footage, yielding over 10.6 million person images covering approximately 434,000 automatically generated identities across 21,697 scenes. Because automated tracking inevitably introduces labeling errors—such as splitting a single individual into multiple tracks or merging different people into one—the authors devised a pre-training framework utilizing noisy labels. This framework integrates standard supervised classification, prototype-based contrastive learning to progressively correct erroneous labels, and label-guided contrastive learning to align matching representations.

The experimental findings show clear performance gains across multiple standard benchmarks. First, models pre-trained with this noisy-label framework established new state-of-the-art results across major evaluation datasets, including CUHK03, Market1501, DukeMTMC, and MSMT17, consistently outperforming models initialized from both supervised general-purpose datasets and large-scale unsupervised datasets. Second, when integrated with strong baseline architectures, the method achieved significant accuracy improvements, raising mean Average Precision on CUHK03 by 5.7 percentage points and on MSMT17 by 2.3 percentage points compared to purely unsupervised pre-training. Third, the pre-trained models demonstrated remarkable sample efficiency: under restricted data settings with only 10% of training identities or frames available, accuracy improvements exceeded 15 percentage points over unsupervised baselines.

These results demonstrate that automated, weakly supervised video tracking provides sufficient structure to train highly transferable visual representations without human labeling costs. The internal label-rectification mechanism successfully repairs tracking errors during training, showing that massive datasets with imperfect labels can outperform smaller, meticulously curated datasets. This approach substantially lowers data collection costs and deployment timelines for visual search and tracking applications, particularly when fine-tuning on limited domain-specific operational data.

Organizations developing or deploying visual re-identification technologies should adopt automated video tracklet generation and label-correcting contrastive pre-training as a cost-effective default pipeline. However, decision-makers should note that the dataset relies on public street-view video collections, and tracking accuracy remains bounded by the quality of the underlying detection algorithms. The pre-trained models are publicly available for scientific research, and operational pilots are advised to assess transfer performance under varying real-world camera resolutions and environmental conditions.

Cover for Large-Scale Pre-training for Person Re-identification with Noisy Labels

Abstract

This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object tracking system on raw videos of an existing un-labeled Re-ID dataset “LUPerson” and build the Noisy Labeled variant called “LUPerson-NL”. Since theses ID labels automatically derived from tracklets inevitably contain noises, we develop a large-scale Pre-training frame-work utilizing Noisy Labels (PNL), which consists of three learning modules: supervised Re-ID learning, prototype-based contrastive learning, and label-guided contrastive learning. In principle, joint learning of these three mod-ules not only clusters similar examples to one prototype, but also rectifies noisy labels based on the prototype as-signment. We demonstrate that learning directly from raw videos is a promising alternative for pre-training, which utilizes spatial and temporal correlations as weak super-vision. This simple pre-training task provides a scalable way to learn SOTA Re-ID representations from scratch on “LUPerson-NL” without bells and whistles. For example, by applying on the same supervised Re-ID method MGN, our pre-trained model improves the mAP over the unsu-pervised pre-training counterpart by 5.7%, 2.2%, 2.3% on CUHK03, DukeMTMC, and MSMT17 respectively. Under the small-scale or few-shot setting, the performance gain is even more significant, suggesting a better transferability of the learned representation. Code is available at https://github.com/DengpanFu/LUPerson-NL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. LUPerson-NL: LUPerson With Noisy Labels
  • 3.1. Constructing LUPerson-NL
  • 3.2. Properties of LUPerson-NL
  • 4. PNL: Pre-training with Noisy Labels for Person Re-ID
  • 4.1. Supervised Classification
  • 4.2. Label Rectification with Prototypes
  • 4.3. Prototype Based Contrastive Learning
  • 4.4. Label-Guided Contrastive Learning
  • 5. Experiments
  • 5.1. Implementation
  • 5.2. Improving Supervised Re-ID
  • 5.3. Improving Unsupervised Re-ID Methods
  • 5.4. Comparison on Small-scale and Few-shot
  • 5.5. Comparison with other pre-training methods
  • 5.6. Ablation Study
  • 5.7. Label Correction
  • 5.8. Comparison with State-of-the-Art Methods
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Pre-training with Noisy Labels Framework for Person Re-Identification

    model/method

    The Pre-training framework utilizing Noisy Labels (PNL) trains visual representations for person re-identification (Re-ID) using weakly supervised pseudo-labels derived from video tracking. Given an input image xix_i from a dataset with noisy tracklet identity labels yi∈{1,…,K}y_i \in \{1, \dots, K\}, two augmented views (x~i,x~i′)(\tilde{x}_i, \tilde{x}'_i) are passed into a Siamese network consisting of a query encoder EqE_q and a momentum key encoder EkE_k.

    The query encoder extracts a feature qi=Eq(x~i)\bm{q}_i = E_q(\tilde{x}_i), and a classifier projects qi\bm{q}_i into class probabilities pi∈RK\bm{p}_i \in \mathbb{R}^K. Using a rectified hard identity label y^i∈{1,…,K}\hat{y}_i \in \{1, \dots, K\} computed during training, PNL optimizes a unified objective consisting of three losses:

    Li=Lcei+λproLproi+λlgcLlgci\mathcal{L}^i = \mathcal{L}^i_{ce} + \lambda_{pro}\mathcal{L}^i_{pro} + \lambda_{lgc}\mathcal{L}^i_{lgc}

    where λpro=1\lambda_{pro} = 1, λlgc=1\lambda_{lgc} = 1, Lproi\mathcal{L}^i_{pro} is the prototype-based contrastive loss, Llgci\mathcal{L}^i_{lgc} is the label-guided contrastive loss, and Lcei\mathcal{L}^i_{ce} is the supervised cross-entropy classification loss defined by:

    Lcei=−log⁡(pi[y^i])\mathcal{L}^i_{ce} = -\log(\bm{p}_i[\hat{y}_i])

    The parameters of the momentum key encoder EkE_k are updated at each iteration via an exponential moving average of EqE_q's weights with momentum coefficient m=0.999m = 0.999.

  2. Knowl 2 — Prototype-Based Pseudo-Label Rectification Mechanism

    model/method

    To mitigate tracking label noise—such as fragmented tracklets of the same individual (Noise-I) or identity switches across different individuals (Noise-II)—PNL dynamically computes a rectified identity label y^i\hat{y}_i for each sample xix_i during training.

    A set of KK class prototypes {c1,c2,…,cK}⊂Rd\{\bm{c}_1, \bm{c}_2, \dots, \bm{c}_K\} \subset \mathbb{R}^d is maintained, representing class-wise feature centroids. The similarity score siks_i^k between the query feature qi\bm{q}_i and prototype ck\bm{c}_k is computed using a softmax temperature τ\tau:

    sik=exp⁡(qi⋅ck/τ)∑j=1Kexp⁡(qi⋅cj/τ)s_i^k = \frac{\exp(\bm{q}_i \cdot \bm{c}_k / \tau)}{\sum_{j=1}^K \exp(\bm{q}_i \cdot \bm{c}_j / \tau)}

    Let si={sik}k=1K\bm{s}_i = \{s_i^k\}_{k=1}^K and pi∈RK\bm{p}_i \in \mathbb{R}^K denote the prototype similarity distribution and the supervised classification probability vector, respectively. A soft pseudo-label li\bm{l}_i is constructed by averaging both distributions:

    li=12(pi+si)\bm{l}_i = \frac{1}{2}(\bm{p}_i + \bm{s}_i)

    The hard rectified label y^i\hat{y}_i is assigned according to confidence threshold TT:

    y^i={arg⁡max⁡jlijif max⁡jlij>T,yiotherwise\hat{y}_i = \begin{cases} \arg\max_j \bm{l}_i^j & \text{if } \max_j \bm{l}_i^j > T, \\ y_i & \text{otherwise} \end{cases}

    where yiy_i is the original raw noisy label. In the default configuration, the temperature is set to τ=0.1\tau = 0.1 and the confidence threshold is set to T=0.8T = 0.8.

  3. Knowl 3 — Label-Guided Contrastive Learning with Rectified Feature Queue

    model/method

    Label-guided contrastive learning extends instance-level contrastive learning by utilizing rectified identity labels to define positive and negative pairs across different samples. A memory queue Q=[(kjt,y^jt)]t=1M\mathcal{Q} = [(\bm{k}_{j_t}, \hat{y}_{j_t})]_{t=1}^M of size MM stores past key features generated by the momentum encoder EkE_k along with their rectified labels y^jt\hat{y}_{j_t}.

    For a query feature qi\bm{q}_i with rectified label y^i\hat{y}_i and positive key feature ki=Ek(x~i′)\bm{k}_i = E_k(\tilde{x}'_i), the positive feature set P(i)\mathcal{P}(i) and negative feature set N(i)\mathcal{N}(i) are defined as:

    P(i)={kjt∣y^jt=y^i,∀(kjt,y^jt)∈Q}∪{ki}\mathcal{P}(i) = \{\bm{k}_{j_t} \mid \hat{y}_{j_t} = \hat{y}_i, \forall (\bm{k}_{j_t}, \hat{y}_{j_t}) \in \mathcal{Q}\} \cup \{\bm{k}_i\}

    N(i)={kjt∣y^jt≠y^i,∀(kjt,y^jt)∈Q}\mathcal{N}(i) = \{\bm{k}_{j_t} \mid \hat{y}_{j_t} \neq \hat{y}_i, \forall (\bm{k}_{j_t}, \hat{y}_{j_t}) \in \mathcal{Q}\}

    The label-guided contrastive loss Llgci\mathcal{L}^i_{lgc} is formulated as:

    Llgci=−1∣P(i)∣log⁡∑k+∈P(i)exp⁡(qi⋅k+/τ)∑k+∈P(i)exp⁡(qi⋅k+/τ)+∑k−∈N(i)exp⁡(qi⋅k−/τ)\mathcal{L}^i_{lgc} = \frac{-1}{|\mathcal{P}(i)|} \log \frac{\sum_{\bm{k}^+ \in \mathcal{P}(i)} \exp(\bm{q}_i \cdot \bm{k}^+ / \tau)}{\sum_{\bm{k}^+ \in \mathcal{P}(i)} \exp(\bm{q}_i \cdot \bm{k}^+ / \tau) + \sum_{\bm{k}^- \in \mathcal{N}(i)} \exp(\bm{q}_i \cdot \bm{k}^- / \tau)}

    where τ=0.1\tau = 0.1 is the temperature hyperparameter. At each step, the current key feature ki\bm{k}_i and its rectified label y^i\hat{y}_i are enqueued while the oldest entry is dequeued.

  4. Knowl 4 — Prototype-Based Contrastive Learning and Momentum Updates

    model/method

    Prototype-based contrastive learning pulls instance features towards their assigned identity centroid prototype while pushing them away from other class prototypes. For a query feature qi\bm{q}_i and its rectified class label y^i\hat{y}_i, the prototype contrastive loss is given by:

    Lproi=−log⁡exp⁡(qi⋅cy^i/τ)∑j=1Kexp⁡(qi⋅cj/τ)\mathcal{L}^i_{pro} = -\log \frac{\exp(\bm{q}_i \cdot \bm{c}_{\hat{y}_i} / \tau)}{\sum_{j=1}^{K} \exp(\bm{q}_i \cdot \bm{c}_j / \tau)}

    where cj∈Rd\bm{c}_j \in \mathbb{R}^d is the prototype vector for identity jj, KK is the total number of identities in the dataset, and τ=0.1\tau = 0.1 is the temperature parameter.

    The prototype corresponding to the rectified label y^i\hat{y}_i is updated step-wise using an exponential moving average with momentum coefficient m=0.999m = 0.999:

    cy^i←mcy^i+(1−m)qi\bm{c}_{\hat{y}_i} \leftarrow m \bm{c}_{\hat{y}_i} + (1 - m)\bm{q}_i

  5. Knowl 5 — LUPerson-NL Dataset Construction and Properties

    definition

    LUPerson-NL is a large-scale, weakly supervised person Re-ID dataset containing 10,683,716 images across 433,997 noisy identity labels extracted from 21,697 street-view video scenes.

    The dataset curation pipeline consists of four stages:

    1. Multi-object tracking: FairMOT is executed frame-by-frame on raw street-view videos to detect pedestrians and generate tracklets, with each tracklet assigned a unique initial ID.
    2. Pose-based bounding box filtering: HRNet human pose estimation detects body landmarks to filter out cropped or partial boxes lacking heads or upper bodies.
    3. Tracklet length thresholding: Person identities that appear in fewer than 200 total frames are removed.
    4. Temporal subsampling: Within each remaining tracklet, images are subsampled at a rate of 1 frame per 20 frames to eliminate near-identical frames while guaranteeing at least 10 images per identity.

    Identity distribution analysis shows that approximately 75% of identities have between 10 and 25 images, while only 6.4% (27,767 identities) contain more than 50 images.

  6. Knowl 6 — Downstream Fine-Tuning Performance Across Supervised Re-ID Baselines

    data/table

    The PNL pre-trained model on LUPerson-NL was evaluated by fine-tuning three standard supervised Re-ID frameworks: Triplet loss baseline (Trip), IDE (classification baseline), and Multiple Granularity Network (MGN). Evaluations were conducted on CUHK03 (labeled protocol), Market-1501, DukeMTMC-reID, and MSMT17 using mean Average Precision (mAP) and Cumulative Matching Characteristics top-1 accuracy (cmc1).

    Pre-train Model CUHK03 Market1501 DukeMTMC MSMT17
    Triplet Baseline (Trip)
    ImageNet Supervised 45.2 / 63.8 76.2 / 89.7 65.2 / 80.7 34.3 / 54.8
    ImageNet Unsupervised 55.5 / 61.2 75.1 / 88.5 65.4 / 81.1 34.4 / 55.4
    LUPerson Unsupervised 62.6 / 67.6 79.8 / 71.5 69.8 / 83.1 36.6 / 57.1
    LUPerson-NL (PNL) 69.1 / 73.1 81.2 / 91.4 71.0 / 84.7 41.4 / 61.6
    IDE Baseline
    ImageNet Supervised 50.6 / 55.9 74.1 / 90.2 62.8 / 80.8 36.2 / 66.2
    ImageNet Unsupervised 52.5 / 57.7 74.5 / 89.3 63.4 / 81.6 37.6 / 67.3
    LUPerson Unsupervised 57.6 / 62.3 77.9 / 91.0 65.9 / 82.2 39.8 / 68.9
    LUPerson-NL (PNL) 68.3 / 73.5 82.4 / 92.8 70.3 / 85.0 44.0 / 72.0
    MGN Baseline
    ImageNet Supervised 70.5 / 71.2 87.5 / 95.1 79.4 / 89.0 63.7 / 85.1
    ImageNet Unsupervised 67.1 / 67.0 88.2 / 95.3 79.5 / 89.1 62.7 / 84.3
    LUPerson Unsupervised 74.7 / 75.4 91.0 / 96.4 82.1 / 91.0 65.7 / 85.5
    LUPerson-NL (PNL) 80.4 / 80.9 91.9 / 96.6 84.3 / 92.0 68.0 / 86.0

    PNL pre-training consistently outperforms both ImageNet supervised/unsupervised models and LUPerson unsupervised pre-training across all baselines and datasets. For the strongest baseline (MGN), PNL improves mAP over LUPerson unsupervised pre-training by 5.7% on CUHK03, 0.9% on Market-1501, 2.2% on DukeMTMC, and 2.3% on MSMT17.

  7. Knowl 7 — Small-Scale and Few-Shot Downstream Re-ID Performance

    data/table

    The transferability of PNL representations was evaluated under restricted downstream data regimes using MGN on Market1501, DukeMTMC, and MSMT17. In the small-scale setting, the percentage of available training identities was varied from 10% to 90%. In the few-shot setting, the percentage of available training images per identity was varied from 10% to 90%. Results are reported in mAP / cmc1 (%).

    Setting Dataset Pre-train 10% 50% 90%
    Small-scale Market1501 IN sup. 53.1 / 76.9 81.5 / 93.5 86.9 / 95.2
    LUP unsup. 64.6 / 85.5 85.8 / 94.9 90.5 / 96.4
    LUPnl pnl. 72.4 / 88.8 88.3 / 95.5 91.3 / 96.4
    DukeMTMC IN sup. 45.1 / 65.3 71.8 / 84.6 78.0 / 88.3
    LUP unsup. 53.5 / 72.0 75.6 / 86.7 81.1 / 90.0
    LUPnl pnl. 60.6 / 75.8 78.8 / 88.3 83.3 / 91.2
    MSMT17 IN sup. 23.2 / 50.2 50.3 / 76.9 61.9 / 84.2
    LUP unsup. 25.5 / 51.1 53.0 / 77.7 63.7 / 85.0
    LUPnl pnl. 28.2 / 51.1 55.5 / 77.2 66.1 / 84.8
    Few-shot Market1501 IN sup. 21.1 / 41.8 80.2 / 92.8 86.7 / 94.6
    LUP unsup. 26.4 / 47.5 84.2 / 93.9 90.4 / 96.3
    LUPnl pnl. 42.0 / 61.6 88.1 / 95.2 91.6 / 96.4
    DukeMTMC IN sup. 31.5 / 47.1 73.9 / 85.7 79.1 / 88.8
    LUP unsup. 35.8 / 50.2 77.7 / 87.4 82.0 / 90.6
    LUPnl pnl. 52.2 / 64.1 81.1 / 89.6 84.1 / 91.3
    MSMT17 IN sup. 14.7 / 34.1 56.2 / 79.5 63.4 / 84.5
    LUP unsup. 17.0 / 36.0 57.4 / 80.5 65.0 / 85.1
    LUPnl pnl. 24.5 / 42.7 62.2 / 81.0 67.4 / 85.3

    The performance gains of PNL over LUPerson unsupervised pre-training are largest in the most data-constrained regimes: at 10% identity scale, PNL gains +7.8% mAP on Market1501 and +7.1% on DukeMTMC; at 10% few-shot image scale, PNL gains +15.6% mAP on Market1501 and +16.4% on DukeMTMC.

  8. Knowl 8 — Ablation Study of PNL Framework Modules and Label Correction

    data/table

    The contribution of each component in PNL was ablated on MSMT17 under small-scale settings (20%, 40%, and 100% available data) using ResNet-50. Evaluated components include supervised cross-entropy classification (cece), instance-wise contrastive learning (icic), prototype-based contrastive learning and label rectification (propro), and label-guided contrastive learning (lgclgc). Results are shown as mAP / cmc1 (%).

    # cece icic propro lgclgc 20% Data 40% Data 100% Data
    1 ✓ 32.0 / 56.1 45.0 / 69.5 62.7 / 83.0
    2 ✓ 34.5 / 59.5 47.9 / 72.6 65.3 / 84.0
    3 ✓ ✓ 37.6 / 62.6 49.6 / 73.5 66.5 / 84.7
    4 ✓ ✓ 35.7 / 59.1 48.5 / 72.4 65.8 / 84.1
    5 ✓ ✓ 38.5 / 63.0 50.9 / 74.5 67.1 / 85.2
    6 ✓ ✓ ✓ 39.0 / 63.4 51.7 / 74.4 67.4 / 85.4
    7 ✓ ✓ ✓ 39.6 / 63.7 51.9 / 75.0 68.0 / 86.0

    Directly training with raw noisy labels using only classification (row 1) performs worse than label-free instance contrastive learning (row 2, 62.7% vs 65.3% mAP at 100%), confirming the destructive impact of label noise. Replacing standard instance contrastive learning (icic) with label-guided contrastive learning (lgclgc) combined with prototype label rectification (propro) achieves the highest accuracy across all data scales.

    An isolated ablation on label correction (lclc) on MSMT17 further demonstrates its impact:

    • ce+proce + pro: without lclc yields 64.8 / 83.4 vs with lclc yielding 65.8 / 84.1.
    • ce+pro+lgcce + pro + lgc: without lclc yields 66.7 / 85.0 vs with lclc yielding 68.0 / 86.0.
  9. Knowl 9 — Downstream Transfer to Unsupervised and Domain-Adaptive Re-ID

    data/table

    The PNL pre-trained representation was tested on unsupervised person Re-ID using the Self-paced Contrastive Learning (SpCL) method under Pure Unsupervised Learning (USL) on Market1501 (M) and DukeMTMC (D), and Unsupervised Domain Adaptation (UDA) across datasets (DukeMTMC →\to Market1501 and Market1501 →\to DukeMTMC). Results are reported in mAP / cmc1 (%).

    Pre-train Model USL (M) USL (D) UDA (D →\to M) UDA (M →\to D)
    ImageNet Supervised 72.4 / 87.8 64.9 / 80.3 76.4 / 90.1 67.9 / 82.3
    ImageNet Unsupervised 72.9 / 88.6 62.6 / 78.8 77.1 / 90.6 66.3 / 81.6
    LUPerson Unsupervised 76.2 / 90.2 67.1 / 81.6 79.2 / 91.7 69.1 / 83.2
    LUPerson-NL (PNL) 75.6 / 89.3 68.1 / 82.0 80.7 / 92.2 72.2 / 84.9

    PNL pre-training achieves the highest performance on both UDA adaptation benchmarks (80.7% mAP on D →\to M and 72.2% mAP on M →\to D) and USL on DukeMTMC (68.1% mAP), demonstrating that noisy-label video pre-training benefits unannotated target domains.

  10. Knowl 10 — State-of-the-Art Person Re-Identification Benchmark Comparison

    data/table

    When fine-tuning standard Re-ID backbones (ResNet-50) initialized with PNL weights on public person Re-ID benchmarks without re-ranking post-processing, the method establishes new state-of-the-art results across CUHK03, Market-1501, DukeMTMC-reID, and MSMT17.

    Method CUHK03 Market1501 DukeMTMC MSMT17
    MGN (2018) 70.5 / 71.2 87.5 / 95.1 79.4 / 89.0 63.7 / 85.1
    BOT (2019) – 85.9 / 94.5 76.4 / 86.4 –
    DSA (2019) 75.2 / 78.9 87.6 / 95.7 74.3 / 86.2 –
    ABDNet (2019) – 88.3 / 95.6 78.6 / 89.0 60.8 / 82.3
    SCAL (2019) 72.3 / 74.8 89.3 / 95.8 79.6 / 89.0 –
    SONA (2019) 79.2 / 81.8 88.8 / 95.6 78.3 / 89.4 –
    SAN (2020) 76.4 / 80.1 88.0 / 96.1 75.5 / 87.9 55.7 / 79.2
    ISP (2020) 74.1 / 76.5 88.6 / 95.3 80.0 / 89.6 –
    LUPerson (2020) 79.6 / 81.9 91.0 / 96.4 82.1 / 91.0 65.7 / 85.5
    Ours + MGN 80.4 / 80.9 91.9 / 96.6 84.3 / 92.0 68.0 / 86.0
    Ours + BDB 82.3 / 84.7 88.4 / 95.4 79.0 / 89.2 53.4 / 79.0

    Results are reported in mAP / cmc1 (%). Ours+MGN achieves the highest scores across Market1501 (91.9% mAP), DukeMTMC (84.3% mAP), and MSMT17 (68.0% mAP), while Ours+BDB achieves 82.3% mAP on CUHK03.

Coverage note — None was omitted; all contributed methodology, dataset properties, mathematical definitions, loss equations, downstream experimental benchmarks, and ablation studies from the main text are represented.

References

  1. 1.Binghui Chen, Weihong Deng, and Jiani Hu. Mixed high-order attention network for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, pages 371–381, 2019.
  2. 2.Guangyi Chen, Chunze Lin, Liangliang Ren, Jiwen Lu, and Jie Zhou. Self-critical attention learning for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, pages 9637–9646, 2019.
  3. 3.Tianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan, Wuyang Chen, Yang Yang, Zhou Ren, and Zhangyang Wang. Abd-net: Attentive but diverse person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, pages 8351–8361, 2019.
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020.
  5. 5.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020.
  6. 6.Weihua Chen, Xiaotang Chen, Jianguo Zhang, and Kaiqi Huang. Beyond triplet loss: a deep quadruplet network for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 403–412, 2017.
  7. 7.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020.
  8. 8.Zhirui Chen, Jianheng Li, and Wei-Shi Zheng. Weakly supervised tracklet person re-identification by deep feature-wise mutual learning. arXiv preprint arXiv:1910.14333, 2019.
  9. 9.Zuozhuo Dai, Mingqiang Chen, Xiaodong Gu, Siyu Zhu, and Ping Tan. Batch dropblock network for person re-identification and beyond. In Proceedings of the IEEE International Conference on Computer Vision, pages 3691–3701, 2019.
  10. 10.Piotr Dollár, Ron Appel, Serge Belongie, and Pietro Perona. Fast feature pyramids for object detection. IEEE transactions on pattern analysis and machine intelligence, 36(8):1532–1545, 2014.
  11. 11.Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan. Object detection with discriminatively trained part-based models. IEEE transactions on pattern analysis and machine intelligence, 32(9):1627–1645, 2009.
  12. 12.Dengpan Fu, Dongdong Chen, Jianmin Bao, Hao Yang, Lu Yuan, Lei Zhang, Houqiang Li, and Dong Chen. Unsupervised pre-training for person re-identification. Proceedings of the IEEE conference on computer vision and pattern recognition, 2021.
  13. 13.Dengpan Fu, Bo Xin, Jingdong Wang, Dongdong Chen, Jianmin Bao, Gang Hua, and Houqiang Li. Improving person re-identification with iterative impression aggregation. IEEE Transactions on Image Processing, 29:9559–9571, 2020.
  14. 14.Yixiao Ge, Dapeng Chen, and Hongsheng Li. Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification. In International Conference on Learning Representations, 2019.
  15. 15.Yixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao, and Hongsheng Li. Self-paced contrastive learning with hybrid memory for domain adaptive object re-id. In Advances in Neural Information Processing Systems, 2020.
  16. 16.Douglas Gray and Hai Tao. Viewpoint invariant pedestrian recognition with an ensemble of localized features. In European conference on computer vision, pages 262–275. Springer, 2008.
  17. 17.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  18. 18.Xinqian Gu, Bingpeng Ma, Hong Chang, Shiguang Shan, and Xilin Chen. Temporal knowledge propagation for image-to-video person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9647–9656, 2019.
  19. 19.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020.
  20. 20.Lingxiao He and Wu Liu. Guided saliency feature learning for person re-identification in crowded scenes. In European Conference on Computer Vision, pages 357–373. Springer, 2020.
  21. 21.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737, 2017.
  22. 22.Yan Huang, Qiang Wu, JingSong Xu, and Yi Zhong. Sbsgan: Suppression of inter-domain background shift for person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9527–9536, 2019.
  23. 23.Xin Jin, Cuiling Lan, Wenjun Zeng, Zhibo Chen, and Li Zhang. Style normalization and restitution for generalizable person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3143–3152, 2020.
  24. 24.Xin Jin, Cuiling Lan, Wenjun Zeng, Guoqiang Wei, and Zhibo Chen. Semantics-aligned representation learning for person re-identification. In AAAI, pages 11173–11180, 2020.
  25. 25.Srikrishna Karanam, Mengran Gou, Ziyan Wu, Angels Rates-Borras, Octavia Camps, and Richard J Radke. A comprehensive evaluation and benchmark for person re-identification: Features, metrics, and datasets. arXiv preprint arXiv:1605.09653, 2(3):5, 2016.
  26. 26.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. Advances in Neural Information Processing Systems, 33, 2020.
  27. 27.Jianing Li, Jingdong Wang, Qi Tian, Wen Gao, and Shiliang Zhang. Global-local temporal representations for video person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3958–3967, 2019.
  28. 28.Junnan Li, Caiming Xiong, and Steven Hoi. Mopro: Webly supervised learning with momentum prototypes. In International Conference on Learning Representations, 2020.
  29. 29.Minxian Li, Xiatian Zhu, and Shaogang Gong. Unsupervised tracklet person re-identification. IEEE transactions on pattern analysis and machine intelligence, 42(7):1770–1782, 2019.
  30. 30.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deepreid: Deep filter pairing neural network for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 152–159, 2014.
  31. 31.Yu-Jhe Li, Yun-Chun Chen, Yen-Yu Lin, Xiaofei Du, and Yu-Chiang Frank Wang. Recover and identify: A generative dual model for cross-resolution person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8090–8099, 2019.
  32. 32.Yu-Jhe Li, Ci-Siang Lin, Yan-Bo Lin, and Yu-Chiang Frank Wang. Cross-dataset person re-identification via unsupervised pose disentanglement and adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7919–7929, 2019.
  33. 33.Yutian Lin, Xuanyi Dong, Liang Zheng, Yan Yan, and Yi Yang. A bottom-up clustering approach to unsupervised person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 8738–8745, 2019.
  34. 34.Fangyi Liu and Lei Zhang. View confusion feature learning for person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6639–6648, 2019.
  35. 35.Chen Change Loy, Chunxiao Liu, and Shaogang Gong. Person re-identification by manifold ranking. In 2013 IEEE International Conference on Image Processing, pages 3567–3571. IEEE, 2013.
  36. 36.Chuanchen Luo, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Spectral feature transformation for person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4976–4985, 2019.
  37. 37.Hao Luo, Wei Jiang, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, and Jianyang Gu. A strong baseline and batch normalization neck for deep person re-identification. IEEE Transactions on Multimedia, 2019.
  38. 38.Jingke Meng, Sheng Wu, and Wei-Shi Zheng. Weakly supervised person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 760–769, 2019.
  39. 39.Hyunjong Park and Bumsub Ham. Relation network for person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11839–11847, 2020.
  40. 40.Xuelin Qian, Yanwei Fu, Tao Xiang, Wenxuan Wang, Jie Qiu, Yang Wu, Yu-Gang Jiang, and Xiangyang Xue. Posenormalized image generation for person re-identification. In Proceedings of the European conference on computer vision (ECCV), pages 650–667, 2018.
  41. 41.Ruijie Quan, Xuanyi Dong, Yu Wu, Linchao Zhu, and Yi Yang. Auto-reid: Searching for a part-aware convnet for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, pages 3750–3759, 2019.
  42. 42.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, pages 91–99, 2015.
  43. 43.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  44. 44.Dong Shen, Shuai Zhao, Jinming Hu, Hao Feng, Deng Cai, and Xiaofei He. Es-net: Erasing salient parts to learn more in re-identification. IEEE Transactions on Image Processing, 2020.
  45. 45.Yantao Shen, Hongsheng Li, Shuai Yi, Dapeng Chen, and Xiaogang Wang. Person re-identification with deep similarity-guided graph neural network. In Proceedings of the European conference on computer vision (ECCV), pages 486–504, 2018.
  46. 46.Yumin Suh, Jingdong Wang, Siyu Tang, Tao Mei, and Kyoung Mu Lee. Part-aligned bilinear representations for person re-identification. In Proceedings of the European Conference on Computer Vision (ECCV), pages 402–419, 2018.
  47. 47.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In CVPR, 2019.
  48. 48.Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). In ECCV, 2018.
  49. 49.Dongkai Wang and Shiliang Zhang. Unsupervised person re-identification via multi-label classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10981–10990, 2020.
  50. 50.Guangrun Wang, Guangcong Wang, Xujie Zhang, Jianhuang Lai, Zhengtao Yu, and Liang Lin. Weakly supervised person re-id: Differentiable graphical learning and a new benchmark. IEEE Transactions on Neural Networks and Learning Systems, 32(5):2142–2156, 2020.
  51. 51.Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, and Xi Zhou. Learning discriminative features with multiple granularities for person re-identification. In 2018 ACM Multimedia Conference on Multimedia Conference, pages 274–282. ACM, 2018.
  52. 52.Longhui Wei, Shiliang Zhang, Wen Gao, and Qi Tian. Person transfer gan to bridge domain gap for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 79–88, 2018.
  53. 53.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3733–3742, 2018.
  54. 54.Bryan Ning Xia, Yuan Gong, Yizhe Zhang, and Christian Poellabauer. Second-order non-local attention networks for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, pages 3760–3769, 2019.
  55. 55.Ye Yuan, Wuyang Chen, Yang Yang, and Zhangyang Wang. In defense of the triplet loss again: Learning robust person re-identification with fast approximated triplet loss and label distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 354–355, 2020.
  56. 56.Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu. Fairmot: On the fairness of detection and re-identification in multiple object tracking. arXiv e-prints, pages arXiv–2004, 2020.
  57. 57.Zhizheng Zhang, Cuiling Lan, Wenjun Zeng, and Zhibo Chen. Densely semantically aligned person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 667–676, 2019.
  58. 58.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. 2015 IEEE International Conference on Computer Vision (ICCV), pages 1116–1124, 2015.
  59. 59.Liang Zheng, Hengheng Zhang, Shaoyan Sun, Manmohan Chandraker, Yi Yang, and Qi Tian. Person re-identification in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1367–1376, 2017.
  60. 60.Zhedong Zheng, Liang Zheng, and Yi Yang. A discriminatively learned cnn embedding for person reidentification. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14(1):1–20, 2017.
  61. 61.Zhedong Zheng, Liang Zheng, and Yi Yang. Unlabeled samples generated by gan improve the person re-identification baseline in vitro. In Proceedings of the IEEE International Conference on Computer Vision, 2017.
  62. 62.Zhun Zhong, Liang Zheng, Donglin Cao, and Shaozi Li. Re-ranking person re-identification with k-reciprocal encoding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1318–1327, 2017.
  63. 63.Kuan Zhu, Haiyun Guo, Zhiwei Liu, Ming Tang, and Jinqiao Wang. Identity-guided human semantic parsing for person re-identification. ECCV, 2020.

Citation

MLA
Fu, D., et al. “Large-Scale Pre-training for Person Re-identification with Noisy Labels”. arXiv, 2022, http://arxiv.org/abs/2203.16533v2.
APA
Fu, D., Chen, D., Yang, H., Bao, J., Yuan, L., Zhang, L., Li, H., Wen, F., & Chen, D. (2022). Large-Scale Pre-training for Person Re-identification with Noisy Labels. arXiv. http://arxiv.org/abs/2203.16533v2
Chicago
Fu, D., D. Chen, H. Yang, et al. 2022. “Large-Scale Pre-training for Person Re-identification with Noisy Labels”. arXiv. http://arxiv.org/abs/2203.16533v2.
Harvard
Fu, D. et al. (2022) “Large-Scale Pre-training for Person Re-identification with Noisy Labels”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16533v2.
Vancouver
1. Fu D, Chen D, Yang H, Bao J, Yuan L, Zhang L, Li H, Wen F, Chen D (2022) Large-Scale Pre-training for Person Re-identification with Noisy Labels. arXiv

BibTeX

@article{fu2022large,
  title = {Large-Scale Pre-training for Person Re-identification with Noisy Labels},
  author = {Fu, Dengpan and Chen, Dongdong and Yang, Hao and Bao, Jianmin and Yuan, Lu and Zhang, Lei and Li, Houqiang and Wen, Fang and Chen, Dong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16533v2},
  eprint = {2203.16533}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE