Large-Scale Pre-training for Person Re-identification with Noisy Labels
Dengpan FuDongdong ChenHao YangJianmin BaoLu YuanLei ZhangHouqiang LiFang WenDong Chen
Introduces a scalable pre-training framework that learns transferable person re-identification representations directly from uncurated video tracklets by combining prototype-based label rectification with label-guided contrastive learning on a ten-million-image noisy dataset.
Person re-identification systems, which match individuals across different camera views, face major deployment bottlenecks because manually annotating thousands of identities across complex video environments is costly, time-consuming, and difficult to scale. While generic pre-trained visual models offer a starting point, they suffer from a substantial domain gap when applied to person-focused surveillance tasks. The article addresses this challenge by evaluating whether scalable model pre-training can learn directly from raw, uncurated street-view videos without any human annotation effort.
The article demonstrates an automated data generation and training strategy. First, the authors built a massive benchmark called LUPerson-NL by applying an automated multi-object tracking algorithm and pose estimation filtering to raw video footage, yielding over 10.6 million person images covering approximately 434,000 automatically generated identities across 21,697 scenes. Because automated tracking inevitably introduces labeling errors—such as splitting a single individual into multiple tracks or merging different people into one—the authors devised a pre-training framework utilizing noisy labels. This framework integrates standard supervised classification, prototype-based contrastive learning to progressively correct erroneous labels, and label-guided contrastive learning to align matching representations.
The experimental findings show clear performance gains across multiple standard benchmarks. First, models pre-trained with this noisy-label framework established new state-of-the-art results across major evaluation datasets, including CUHK03, Market1501, DukeMTMC, and MSMT17, consistently outperforming models initialized from both supervised general-purpose datasets and large-scale unsupervised datasets. Second, when integrated with strong baseline architectures, the method achieved significant accuracy improvements, raising mean Average Precision on CUHK03 by 5.7 percentage points and on MSMT17 by 2.3 percentage points compared to purely unsupervised pre-training. Third, the pre-trained models demonstrated remarkable sample efficiency: under restricted data settings with only 10% of training identities or frames available, accuracy improvements exceeded 15 percentage points over unsupervised baselines.
These results demonstrate that automated, weakly supervised video tracking provides sufficient structure to train highly transferable visual representations without human labeling costs. The internal label-rectification mechanism successfully repairs tracking errors during training, showing that massive datasets with imperfect labels can outperform smaller, meticulously curated datasets. This approach substantially lowers data collection costs and deployment timelines for visual search and tracking applications, particularly when fine-tuning on limited domain-specific operational data.
Organizations developing or deploying visual re-identification technologies should adopt automated video tracklet generation and label-correcting contrastive pre-training as a cost-effective default pipeline. However, decision-makers should note that the dataset relies on public street-view video collections, and tracking accuracy remains bounded by the quality of the underlying detection algorithms. The pre-trained models are publicly available for scientific research, and operational pilots are advised to assess transfer performance under varying real-world camera resolutions and environmental conditions.
- Paper: Deep Learning for Person Re-Identification: A Survey and Outlook, Mang Ye et al. (2020). Provides a foundational and comprehensive taxonomy of deep learning methodologies, evaluation protocols, and standard benchmarks in person re-identification.
- Paper: Bag of Tricks and a Strong Baseline for Deep Person Re-Identification, Hao Luo et al. (2019). Establishes essential baseline training practices and optimization tricks that serve as the standard foundation for modern person re-identification models.
- Paper: Unsupervised Learning of Visual Features by Contrasting Cluster Assignments, Mathilde Caron et al. (2020). Introduces online prototype-based self-supervised contrastive clustering, providing the core conceptual mechanism adapted for label rectification in noisy pre-training.
- Paper: Improved Baselines with Momentum Contrastive Learning, Xinlei Chen et al. (2020). Details foundational momentum contrastive learning techniques used for scalable self-supervised and semi-supervised visual representation learning.
- Paper: Exploring the Limits of Weakly Supervised Pretraining, Dhruv Mahajan et al. (2018). Demonstrates the power and scalability of pre-training computer vision architectures on massive, uncurated datasets with noisy web labels.
- Paper: Person Transfer GAN to Bridge Domain Gap for Person Re-identification, Longhui Wei et al. (2017). Introduces the large-scale MSMT17 benchmark and investigates visual domain gaps that motivate pre-training directly on realistic multi-camera surveillance data.
- Paper: Learning Discriminative Features with Multiple Granularities for Person Re-Identification, Guanshuo Wang et al. (2018). Presents multi-granularity feature extraction and combined classification-ranking objectives widely adopted across modern person re-identification frameworks.
- Paper: Tracking Objects as Points, Xingyi Zhou et al. (2020). Exemplifies automated multi-object tracking mechanisms that generate the initial uncurated pedestrian tracklets used to construct massive pre-training datasets.
- Paper: HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining, Shixiang Tang et al. (2023). Generalizes scalable human-centric pre-training across multiple diverse perception tasks, including person re-identification, pose estimation, and pedestrian detection.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). Extends scalable pre-training with noisy synthetic annotations to multi-modal text-to-image person retrieval using large language models.
- Paper: MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID, Jianyang Gu et al. (2023). Designs compact, specialized neural architectures using contrastive search tailored specifically for open-set object and person re-identification.
- Paper: Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification, Shengcai Liao et al. (2022). Develops graph-sampling techniques to optimize mini-batch creation for scalable metric learning and generalizable person re-identification.
