Deep Learning for Person Re-Identification: A Survey and Outlook

Mang YeJianbing ShenGaojie LinTao XiangLing ShaoSteven C. H. Hoi

article2020TPAMI2,253 citations

Presents a comprehensive taxonomy of closed- and open-world person re-identification along with a competitive AGW baseline evaluated across twelve datasets and the mINP metric to measure practical search costs.

Listen

Surveillance systems increasingly rely on automated person re-identification to match individuals across non-overlapping camera views for public safety and operational monitoring. While artificial intelligence and deep learning have achieved rapid technical milestones on controlled benchmarks, practical deployment faces major hurdles. Real-world applications encounter shifting camera networks, unconstrained outdoor lighting, occlusions, varying resolutions, heterogeneous image modalities (such as infrared night imaging or text descriptions), and imperfect or unavailable manual annotations.

The article systematically reviews the progression of deep learning in pedestrian retrieval, categorizing the field into closed-world (ideal, controlled assumptions) and open-world (realistic, unconstrained conditions) settings. It evaluates existing methodologies across representation learning, deep metric learning, and ranking optimization, while introducing a high-performing baseline model and a practical metric designed to assess the total effort required to locate true matches.

The review synthesizes findings across dozens of major benchmarks spanning image, video, cross-modality, and partial-visibility datasets. The authors also experimentally validate their unified baseline architectureAttention Generalized-mean pooling with Weighted triplet loss (AGW)—across twelve public benchmarks, comparing it against established baselines and specialized state-of-the-art methods.

The key findings are as follows:

  1. Controlled Benchmarks Near Saturation: Closed-world image and video benchmarks have reached human-level or near-saturated retrieval rates. On standard single-query visible benchmarks like Market-1501, top-1 accuracy routinely exceeds 96%, shifting the primary research frontier toward unconstrained open-world challenges.
  2. Architectural and Optimization Synergy Drives Performance: Across modalities, top-performing models consistently combine part-level or localized feature aggregations, attention mechanisms (to suppress background noise and model cross-view relationships), and joint multi-loss training (unifying classification and relative distance metrics).
  3. Unified AGW Baseline Establishes Strong Cross-Task Performance: The introduced AGW framework achieves competitive or state-of-the-art results across single-modality image, video tracklet, visible-infrared cross-modality, and partial-occlusion datasets without requiring task-specific structural redesigns.
  4. Real-World Evaluation Requires Total-Effort Metrics: Standard metrics like top-1 accuracy and mean Average Precision mask investigator workload when retrieving multiple appearances of a person. The proposed mean Inverse Negative Penalty (mINP) effectively captures the penalty associated with finding the hardest true match.

These findings indicate that while core identification algorithms are mature under clean, single-camera modality setups, deploying these systems into uncontrolled enterprise environments introduces substantial performance drops and operational risks. System designers can leverage unified baselines like AGW to reduce deployment overhead, but must account for the significant manual review costs that arise when difficult matches are buried deep in retrieved lists.

Decision-makers and engineering teams should transition evaluation protocols from top-1 precision to holistic workload metrics such as mINP when assessing operational tools. Future initiatives should prioritize unsupervised domain adaptation, dynamic model updates for newly added camera feeds, and automated multi-modal fusion rather than solely optimizing closed-world classification accuracy. Further pilot testing in complex, dynamic camera environments remains necessary before deploying fully automated person-search pipelines at scale.

The analysis notes that current benchmark results may overstate operational efficacy because standard datasets rely heavily on pre-cropped bounding boxes and static gallery sizes. Confidence in the reported algorithmic comparisons is high across the evaluated academic datasets, but practitioners should exercise caution when deploying systems in scenarios involving heavy disguise, clothing changes over extended intervals, or extreme crowd densities.

Cover for Deep Learning for Person Re-Identification: A Survey and Outlook

Abstract

Person re-identification (Re-ID) aims at retrieving a person of interest across multiple non-overlapping cameras. With the advancement of deep neural networks and increasing demand of intelligent video surveillance, it has gained significantly increased interest in the computer vision community. By dissecting the involved components in developing a person Re-ID system, we categorize it into the closed-world and open-world settings. The widely studied closed-world setting is usually applied under various research-oriented assumptions, and has achieved inspiring success using deep learning techniques on a number of datasets. We first conduct a comprehensive overview with in-depth analysis for closed-world person Re-ID from three different perspectives, including deep feature representation learning, deep metric learning and ranking optimization. With the performance saturation under closed-world setting, the research focus for person Re-ID has recently shifted to the open-world setting, facing more challenging issues. This setting is closer to practical applications under specific scenarios. We summarize the open-world Re-ID in terms of five different aspects. By analyzing the advantages of existing methods, we design a powerful AGW baseline, achieving state-of-the-art or at least comparable performance on twelve datasets for FOUR different Re-ID tasks. Meanwhile, we introduce a new evaluation metric (mINP) for person Re-ID, indicating the cost for finding all the correct matches, which provides an additional criteria to evaluate the Re-ID system for real applications. Finally, some important yet under-investigated open issues are discussed.

Table of Contents

  • I Introduction
  • II Closed-world Person Re-Identification
  • II-A Feature Representation Learning
  • II-A1 Global Feature Representation Learning
  • II-A2 Local Feature Representation Learning
  • II-A3 Auxiliary Feature Representation Learning
  • II-A4 Video Feature Representation Learning
  • II-A5 Architecture Design
  • II-B Deep Metric Learning
  • II-B1 Loss Function Design
  • II-B2 Training strategy
  • II-C Ranking Optimization
  • II-C1 Re-ranking
  • II-C2 Rank Fusion
  • II-D Datasets and Evaluation
  • II-D1 Datasets and Evaluation Metrics
  • II-D2 In-depth Analysis on State-of-The-Arts
  • III Open-world Person Re-Identification
  • III-A Heterogeneous Re-ID
  • III-A1 Depth-based Re-ID
  • III-A2 Text-to-Image Re-ID
  • III-A3 Visible-Infrared Re-ID
  • III-A4 Cross-Resolution Re-ID
  • III-B End-to-End Re-ID
  • III-C Semi-supervised and Unsupervised Re-ID
  • III-C1 Unsupervised Re-ID
  • III-C2 Unsupervised Domain Adaptation
  • III-C3 State-of-The-Arts for Unsupervised Re-ID
  • III-D Noise-Robust Re-ID
  • III-E Open-set Re-ID and Beyond
  • IV An Outlook: Re-ID in Next Era
  • IV-A mINP: A New Evaluation Metric for Re-ID
  • IV-B A New Baseline for Single-/Cross-Modality Re-ID
  • IV-C Under-Investigated Open Issues
  • IV-C1 Uncontrollable Data Collection
  • IV-C2 Human Annotation Minimization
  • IV-C3 Domain-Specific/Generalizable Architecture Design
  • IV-C4 Dynamic Model Updating
  • IV-C5 Efficient Model Deployment
  • V Concluding Remarks
  • References

Knowls

  1. Knowl 1 — mINP (mean Inverse Negative Penalty) Evaluation Metric

    equation

    The mean Inverse Negative Penalty (mINP) measures the retrieval effort and search penalty required to find all correct matches of a query person by evaluating the rank position of the hardest ground-truth match.

    For a query image ii, let RihardN1R_i^{\text{hard}} \in \mathbb{N}_{\ge 1} denote the rank position of the hardest (lowest-ranked) correct match in the gallery ranking list, and let GiN1|G_i| \in \mathbb{N}_{\ge 1} denote the total number of ground-truth matches for query ii in the gallery set. The Negative Penalty (NPiNP_i) is defined as:

    NPi=RihardGiRihardNP_i = \frac{R_i^{\text{hard}} - |G_i|}{R_i^{\text{hard}}}

    The Inverse Negative Penalty (INPiINP_i) is defined as:

    INPi=1NPi=GiRihardINP_i = 1 - NP_i = \frac{|G_i|}{R_i^{\text{hard}}}

    Across all nn queries in the evaluation set, the mean Inverse Negative Penalty (mINPmINP) is computed as:

    mINP=1ni=1n(1NPi)=1ni=1nGiRihardmINP = \frac{1}{n} \sum_{i=1}^n (1 - NP_i) = \frac{1}{n} \sum_{i=1}^n \frac{|G_i|}{R_i^{\text{hard}}}

    A higher mINP(0,1]mINP \in (0, 1] corresponds to better retrieval performance, reaching 1.01.0 when all Gi|G_i| ground-truth matches occupy the top Gi|G_i| ranks. Unlike Cumulative Matching Characteristics (CMC), which evaluates only the first retrieved match, and mean Average Precision (mAP), which can be dominated by top-ranked easy matches, mINP directly reflects inspector workload in multi-camera tracking by penalizing late-retrieved hard matches.

  2. Knowl 2 — Weighted Regularization Triplet (WRT) Loss

    equation

    The Weighted Regularization Triplet (WRT) loss optimizes relative pairwise distances between positive and negative pairs within a training mini-batch without introducing an explicit margin hyperparameter:

    Lwrt(i)=log(1+exp(jPiwijpdijpkNiwikndikn))L_{\text{wrt}}(i) = \log\left(1 + \exp\left(\sum_{j \in P_i} w_{ij}^p d_{ij}^p - \sum_{k \in N_i} w_{ik}^n d_{ik}^n\right)\right)

    where ii is an anchor sample, PiP_i is the set of positive samples (matching the identity of ii), NiN_i is the set of negative samples (different identities), and dijp=fifj2d_{ij}^p = \|f_i - f_j\|_2 and dikn=fifk2d_{ik}^n = \|f_i - f_k\|_2 are the pairwise Euclidean distances between feature embeddings fi,fj,fkf_i, f_j, f_k. The adaptive pair weights wijpw_{ij}^p and wiknw_{ik}^n are computed via normalized softmax operations:

    wijp=exp(dijp)jPiexp(dijp),wikn=exp(dikn)kNiexp(dikn)w_{ij}^p = \frac{\exp(d_{ij}^p)}{\sum_{j' \in P_i} \exp(d_{ij'}^p)}, \quad w_{ik}^n = \frac{\exp(-d_{ik}^n)}{\sum_{k' \in N_i} \exp(-d_{ik'}^n)}

    Under this weighting scheme, hard positive instances with larger distances receive larger positive weights wijpw_{ij}^p, and hard negative instances with smaller distances receive larger negative weights wiknw_{ik}^n, dynamically scaling the gradient contributions according to sample hardness.

  3. Knowl 3 — AGW Baseline Architecture for Person Re-Identification

    model/method

    AGW (Attention Generalized mean pooling with Weighted triplet loss) is a strong, unified baseline network designed for person re-identification across single-modality, video-based, partial, and cross-modality tasks. Built on an ImageNet-pretrained ResNet-50 backbone, AGW introduces three key architectural improvements:

    1. Modified Backbone and BNNeck: The spatial downsampling stride of the final residual stage (conv5_x) is changed from 2 to 1, producing a spatial feature map of size 16×816 \times 8 for an input resolution of 256×128256 \times 128. A Batch Normalization neck (BNNeck) is placed after the pooling layer. The feature prior to BNNeck is used to compute metric losses during training, whereas the normalized feature after BNNeck is used as the representation for distance computation during inference.
    2. Non-Local Attention Blocks: Five non-local attention blocks (dot-product variant with a 512-channel bottleneck) are inserted into the residual stages after layers conv3_3, conv3_4, conv4_4, conv4_5, and conv4_6. Each block computes a position-wise weighted sum representation zi=Wzϕ(xi)+xiz_i = W_z \phi(x_i) + x_i with residual addition, followed by a BatchNorm layer whose affine parameters are initialized to zero.
    3. Generalized-Mean (GeM) Pooling: The standard Global Average Pooling layer is replaced by a learnable GeM pooling layer with initial power parameter pk=3.0p_k = 3.0.
    4. Optimization Objective: AGW combines cross-entropy identity classification loss LidL_{\text{id}} with label smoothing (epsilon=0.1\\epsilon = 0.1), center loss LctL_{\text{ct}}, and Weighted Regularization Triplet loss LwrtL_{\text{wrt}}:

    Ltotal=Lid+β1Lct+β2LwrtL_{\text{total}} = L_{\text{id}} + \beta_1 L_{\text{ct}} + \beta_2 L_{\text{wrt}}

    with balance coefficients set to β1=0.0005\beta_1 = 0.0005 and β2=1.0\beta_2 = 1.0.

  4. Knowl 4 — Generalized-Mean (GeM) Pooling in Feature Representation Learning

    equation

    In the AGW baseline for person re-identification, domain-specific discriminative pooling over spatial feature maps is performed via Generalized-Mean (GeM) pooling:

    f=[f1,,fk,,fK]T,fk=(1XkxiXkxipk)1pk\mathbf{f} = [f_1, \dots, f_k, \dots, f_K]^T, \quad f_k = \left(\frac{1}{|X_k|} \sum_{x_i \in X_k} x_i^{p_k}\right)^{\frac{1}{p_k}}

    where KK denotes the number of feature channels in the final convolutional layer, XkX_k is the set of W×HW \times H spatial activation values for channel kk (with Xk=W×H|X_k| = W \times H), and pkp_k is a pooling parameter learned via back-propagation during model training (initialized to pk=3.0p_k = 3.0).

    GeM pooling continuously interpolates between standard pooling operations: it is equivalent to standard Average Pooling when pk=1p_k = 1, and approaches Max Pooling as pkp_k \to \infty, allowing the network to adaptively focus on salient, identity-discriminative feature activations.

  5. Knowl 5 — AGW Adaptation for Video-Based Person Re-Identification

    model/method

    To handle video sequences containing multiple frames and temporal variations, AGW is extended into two video-specific models:

    1. AGW (Video Baseline): Given a pedestrian tracklet, 4 frames are selected using a constraint random sampling strategy during training. Frame-level feature vectors are extracted by the ResNet-50 backbone with non-local attention and GeM pooling, and then averaged via temporal pooling to form a single video-level feature representation prior to the BNNeck layer. Training is conducted for 400 epochs with the learning rate decreased by a factor of 10 every 100 epochs.
    2. AGW+ (Dense Video Variant): AGW+ enhances temporal modeling by employing a dense sampling strategy during inference, passing all frames of a tracklet through the network to generate the aggregated sequence representation. In training, AGW+ removes the warm-up learning rate schedule and adds a dropout layer directly before the linear classification layer to improve temporal representation robustness.
  6. Knowl 6 — Two-Stream AGW Architecture for Cross-Modality Visible-Infrared Re-ID

    model/method

    For cross-modality matching between daytime visible (RGB) and nighttime infrared (thermal) images, AGW is adapted into a two-stream network structure:

    1. Network Topology: The first residual block (conv1 / stage 1) has modality-specific parameters for visible and infrared inputs separately to capture modality-specific low-level cues. Subsequent blocks (conv2_x through conv5_x) share weights across both modalities to project features into a unified embedding space. A single shared identity classifier is used for both modalities.
    2. Batch Construction: Each mini-batch is constructed by randomly selecting 8 identities, with 4 visible and 4 infrared images sampled per identity, yielding 64 images per batch (32 visible, 32 infrared). This ensures balanced cross-modality and intra-modality positive and negative pairs for triplet mining.
    3. Training Protocol: Input images are resized to 288×144288 \times 144 and randomly cropped to 256×128256 \times 128. Infrared images are expanded to 3 channels. The model is trained for 60 epochs using Stochastic Gradient Descent (SGD) with momentum 0.9, an initial learning rate of 0.1 (decayed by 0.1 at epoch 20 and 0.01 at epoch 50), and loss Ltotal=Lid+LwrtL_{\text{total}} = L_{\text{id}} + L_{\text{wrt}} with GeM pooling parameter pk=3.0p_k = 3.0.
  7. Knowl 7 — Empirical Performance of AGW on Single-Modality Image Datasets

    data/table
    Method Market-1501 DukeMTMC
    Rank-1 (%) mAP (%) mINP (%) Rank-1 (%) mAP (%) mINP (%)
    BagTricks 94.5 85.9 59.4 86.4 76.4 40.7
    ABD-Net 95.6 88.3 66.2 89.0 78.6 42.1
    Base (B) 94.2 85.4 58.3 86.1 76.1 40.3
    B + Att 94.9 86.9 62.2 87.5 77.6 41.9
    B + WRT 94.6 86.8 61.9 87.1 77.0 41.4
    B + GeM 94.4 86.3 60.1 87.3 77.3 41.9
    B + WRT + GeM 94.9 87.1 62.5 88.2 78.1 43.4
    AGW (Full) 95.1 87.8 65.0 89.0 79.6 45.7
    Method CUHK03 MSMT17
    Rank-1 (%) mAP (%) mINP (%) Rank-1 (%) mAP (%) mINP (%)
    BagTricks 58.0 56.6 43.8 63.4 45.1 12.4
    AGW (Full) 63.6 62.0 50.3 68.3 49.3 14.7

    Ablation analysis and comparison on image-based benchmarks without re-ranking or additional annotations. Each component of AGW—non-local attention (Att), Weighted Regularization Triplet loss (WRT), and Generalized-Mean pooling (GeM)—consistently contributes to gains in Rank-1, mAP, and mINP over the baseline (B). On DukeMTMC, AGW achieves 45.7% mINP, outperforming ABD-Net (42.1%) and BagTricks (40.7%). On CUHK03 (detected bounding box setting) and MSMT17, AGW improves Rank-1 accuracy by +5.6% and +4.9% respectively over BagTricks.

  8. Knowl 8 — Empirical Performance of AGW and AGW+ on Video-Based Person Re-ID

    data/table
    Method MARS DukeVideo
    R1 (%) R5 (%) mAP (%) mINP (%) R1 (%) R5 (%) mAP (%) mINP (%)
    BagTricks 85.8 95.2 81.6 62.0 92.6 98.9 92.4 88.3
    CoSeg 84.9 95.5 79.9 57.8 95.4 99.3 94.1 89.8
    AGW 87.0 95.7 82.2 62.8 94.6 99.1 93.4 89.2
    AGW+ 87.6 85.8 83.0 63.9 95.4 99.3 94.9 91.9
    Method PRID2011 iLIDS-VID
    R1 (%) R5 (%) R20 (%) mINP (%) R1 (%) R5 (%) R20 (%) mINP (%)
    BagTricks 84.3 93.3 98.0 88.5 74.0 93.3 99.1 82.2
    AGW 87.8 96.6 98.9 91.7 78.0 97.0 99.5 85.5
    AGW+ 94.4 98.4 100.0 95.4 83.2 98.3 99.7 89.0

    Performance comparison of AGW and the dense-sampling video variant AGW+ on four video Re-ID datasets. By incorporating dense frame aggregation at inference and removing learning rate warmup with added dropout before classification, AGW+ achieves state-of-the-art results: 87.6% Rank-1 and 83.0% mAP on MARS, 95.4% Rank-1 and 94.9% mAP on DukeVideo, 94.4% Rank-1 and 95.4% mINP on PRID2011, and 83.2% Rank-1 and 89.0% mINP on iLIDS-VID.

  9. Knowl 9 — Empirical Performance of AGW on Visible-Infrared and Partial Re-ID

    data/table
    Method RegDB (Vis-to-Th) RegDB (Th-to-Vis) SYSU-MM01 (All Search)
    Rank-1 (%) mAP (%) Rank-1 (%) mAP (%) Rank-1 (%) mAP (%)
    Zero-Pad 17.75 18.90 16.63 17.82 14.80 15.95
    eBDTR 34.62 33.46 34.21 32.49 27.82 28.42
    HSME 50.85 47.00 50.15 46.16 20.68 23.12
    AlignGAN 57.90 53.60 56.30 53.40 42.40 40.70
    AGW 70.05 66.37 70.49 65.90 47.50 47.65
    Method Partial-REID Partial-iLIDS
    Rank-1 (%) Rank-3 (%) mINP (%) Rank-1 (%) Rank-3 (%) mINP (%)
    DSR 50.7 70.0 - 58.8 67.2 -
    VPM 67.7 81.9 - 67.2 76.5 -
    BagTricks 62.0 74.0 45.4 58.8 73.9 68.7
    AGW 69.7 80.0 56.7 64.7 79.8 73.3

    Cross-modality results demonstrate that two-stream AGW achieves 70.05% Rank-1 and 66.37% mAP on RegDB (Visible to Thermal), 70.49% Rank-1 and 65.90% mAP on RegDB (Thermal to Visible), and 47.50% Rank-1 and 47.65% mAP on SYSU-MM01 (All Search), outperforming GAN-based generation methods (such as AlignGAN) without requiring image synthesis. For partial person Re-ID (evaluated on models trained on Market-1501), AGW attains 69.7% Rank-1 on Partial-REID (surpassing VPM's 67.7%) and 64.7% Rank-1 on Partial-iLIDS.

  10. Knowl 10 — Closed-World versus Open-World Taxonomy for Person Re-Identification

    definition

    Person re-identification (Re-ID) systems are categorized into closed-world and open-world paradigms based on their assumptions across five system stages:

    1. Raw Data Collection: Closed-world assumes single-modality visible spectrum RGB cameras. Open-world involves heterogeneous sensor data, including infrared (thermal) images, depth maps, 3D body models, sketches, and natural language text queries.
    2. Bounding Box Generation: Closed-world operates on manually cropped or pre-detected bounding boxes focusing strictly on pedestrian appearance. Open-world performs end-to-end person search directly from raw surveillance video frames and multi-camera pedestrian tracking.
    3. Training Data Annotation: Closed-world assumes sufficient identity-labeled training data across camera pairs. Open-world addresses unavailable or limited cross-camera labels through unsupervised domain adaptation, semi-supervised learning, and one-shot learning.
    4. Annotation Quality and Visibility: Closed-world assumes clean labels and full-body pedestrian visibility. Open-world handles label noise, sample noise (poor detection/tracking crops), and severe occlusion (partial Re-ID).
    5. Pedestrian Retrieval: Closed-world operates under a closed-set assumption where the query identity is guaranteed to exist in the gallery. Open-world encompasses open-set Re-ID (identity verification with rejection of non-occurring queries), group Re-ID, and continuous adaptation across dynamically expanding camera networks.
  11. Knowl 11 — Open Issues and Limitations in Person Re-Identification

    limitation

    Key unresolved challenges and technical limitations in person re-identification include:

    • Clothing Change: Existing representations rely heavily on cloth color and pattern, causing severe failure when target persons change clothes across camera views. Invariant biometric cues (e.g., body shape, gait, face) remain difficult to extract reliably under surveillance conditions.
    • mINP Sensitivity with Massive Gallery Scale: While mINP effectively identifies the rank position of the hardest true positive match, numerical mINP differences compress as gallery sizes grow to very large scales.
    • Cross-Domain Generalization Gap: Deep models trained on standard benchmarks experience sharp performance degradation when evaluated on unseen domains or under adversarial attacks, indicating an urgent need for domain-generalizable representations rather than domain-specific fitting.
    • Dynamic Camera Network Adaptation: Real surveillance systems dynamically add, remove, or reposition camera nodes. Updating Re-ID models continuously to accommodate new camera distributions without catastrophic forgetting or full retraining from scratch remains unresolved.
    • Scalable and Fast Deployment: Fast binary hashing, model distillation, and resource-aware adaptive computation (e.g., dynamic early-exit routing) are needed to enable real-time retrieval over millions of gallery images on resource-constrained hardware.

Coverage note — No substantial contributed material from the paper was omitted; the knowls cover the mINP metric, the AGW baseline architecture and its video/cross-modality extensions, all empirical benchmark tables, and the taxonomies and open challenges.

References

  1. 1.Y.-C. Chen, X. Zhu, W.-S. Zheng, and J.-H. Lai, “Person re-identification by camera correlation aware feature augmentation,” IEEE TPAMI, vol. 40, no. 2, pp. 392–408, 2018.
  2. 2.L. Zheng, Y. Yang, and A. G. Hauptmann, “Person re-identification: Past, present and future,” arXiv preprint arXiv:1610.02984, 2016.
  3. 3.N. Gheissari, T. B. Sebastian, and R. Hartley, “Person reidentification using spatiotemporal appearance,” in CVPR, 2006, pp. 1528–1535.
  4. 4.J. Almazan, B. Gajic, N. Murray, and D. Larlus, “Re-id done right: towards good practices for person re-identification,” arXiv preprint arXiv:1801.05339, 2018.
  5. 5.L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in ICCV, 2015, pp. 1116–1124.
  6. 6.N. Martinel, G. Luca Foresti, and C. Micheloni, “Aggregating deep pyramidal representations for person re-identification,” in CVPR Workshops, 2019, pp. 0–0.
  7. 7.T. Wang, S. Gong, X. Zhu, and S. Wang, “Person re-identification by video ranking,” in ECCV, 2014, pp. 688–703.
  8. 8.L. Zheng, Z. Bie, Y. Sun, J. Wang, C. Su, S. Wang, and Q. Tian, “Mars: A video benchmark for large-scale person re-identification,” in ECCV, 2016.
  9. 9.M. Ye, C. Liang, Z. Wang, Q. Leng, J. Chen, and J. Liu, “Specific person retrieval via incomplete text description,” in ACM ICMR, 2015, pp. 547–550.
  10. 10.S. Li, T. Xiao, H. Li, W. Yang, and X. Wang, “Identity-aware textual-visual matching with latent co-attention,” in ICCV, 2017, pp. 1890–1899.
  11. 11.S. Karanam, Y. Li, and R. J. Radke, “Person re-identification with discriminatively trained viewpoint invariant dictionaries,” in ICCV, 2015, pp. 4516–4524.
  12. 12.S. Bak, S. Zaidenberg, B. Boulay, and F. Bremond, “Improving person re-identification by viewpoint cues,” in AVSS, 2014, pp. 175–180.
  13. 13.X. Li, W.-S. Zheng, X. Wang, T. Xiang, and S. Gong, “Multi-scale learning for low-resolution person re-identification,” in ICCV, 2015, pp. 3765–3773.
  14. 14.Y. Wang, L. Wang, Y. You, X. Zou, V. Chen, S. Li, G. Huang, B. Hariharan, and K. Q. Weinberger, “Resource aware person re-identification across multiple resolutions,” in CVPR, 2018, pp. 8042–8051.
  15. 15.Y. Huang, Z.-J. Zha, X. Fu, and W. Zhang, “Illumination-invariant person re-identification,” in ACM MM, 2019, pp. 365–373.
  16. 16.Y.-J. Cho and K.-J. Yoon, “Improving person re-identification via pose-aware multi-shot matching,” in CVPR, 2016, pp. 1354–1362.
  17. 17.H. Zhao, M. Tian, S. Sun, and et al, “Spindle net: Person re-identification with human body region guided feature decomposition and fusion,” in CVPR, 2017, pp. 1077–1085.
  18. 18.M. S. Sarfraz, A. Schumann, A. Eberle, and R. Stiefelhagen, “A pose-sensitive embedding for person re-identification with expanded cross neighborhood re-ranking,” in CVPR, 2018, pp. 420–429.
  19. 19.H. Huang, D. Li, Z. Zhang, X. Chen, and K. Huang, “Adversarially occluded samples for person re-identification,” in CVPR, 2018, pp. 5098–5107.
  20. 20.R. Hou, B. Ma, H. Chang, X. Gu, S. Shan, and X. Chen, “Vrstc: Occlusion-free video person re-identification,” in CVPR, 2019, pp. 7183–7192.
  21. 21.A. Wu, W.-s. Zheng, H.-X. Yu, S. Gong, and J. Lai, “Rgb-infrared cross-modality person re-identification,” in ICCV, 2017, pp. 5380–5389.
  22. 22.C. Song, Y. Huang, W. Ouyang, and L. Wang, “Mask-guided contrastive attention model for person re-identification,” in CVPR, 2018, pp. 1179–1188.
  23. 23.A. Das, R. Panda, and A. K. Roy-Chowdhury, “Continuous adaptation of multi-camera person identification models through sparse non-redundant representative selection,” CVIU, vol. 156, pp. 66–78, 2017.
  24. 24.N. Martinel, A. Das, C. Micheloni, and A. K. Roy-Chowdhury, “Temporal model adaptation for person re-identification,” in ECCV, 2016, pp. 858–877.
  25. 25.J. Garcia, N. Martinel, A. Gardel, I. Bravo, G. L. Foresti, and C. Micheloni, “Discriminant context information analysis for post-ranking person re-identification,” IEEE Transactions on Image Processing, vol. 26, no. 4, pp. 1650–1665, 2017.
  26. 26.W.-S. Zheng, S. Gong, and T. Xiang, “Towards open-world person re-identification by one-shot group-based verification,” IEEE TPAMI, vol. 38, no. 3, pp. 591–606, 2015.
  27. 27.A. Das, R. Panda, and A. Roy-Chowdhury, “Active image pair selection for continuous person re-identification,” in ICIP, 2015, pp. 4263–4267.
  28. 28.J. Song, Y. Yang, Y.-Z. Song, T. Xiang, and T. M. Hospedales, “Generalizable person re-identification by domain-invariant mapping network,” in CVPR, 2019, pp. 719–728.
  29. 29.A. Das, A. Chakraborty, and A. K. Roy-Chowdhury, “Consistent re-identification in a camera network,” in ECCV, 2014, pp. 330–345.
  30. 30.Q. Yang, A. Wu, and W. Zheng, “Person re-identification by contour sketch under moderate clothing change.” IEEE TPAMI, 2019.
  31. 31.D. Gray and H. Tao, “Viewpoint invariant pedestrian recognition with an ensemble of localized features,” in ECCV, 2008, pp. 262–275.
  32. 32.M. Farenzena, L. Bazzani, A. Perina, V. Murino, and M. Cristani, “Person re-identification by symmetry-driven accumulation of local features,” in CVPR, 2010, pp. 2360–2367.
  33. 33.Y. Yang, J. Yang, J. Yan, S. Liao, D. Yi, and S. Z. Li, “Salient color names for person re-identification,” in ECCV, 2014, pp. 536–551.
  34. 34.S. Liao, Y. Hu, X. Zhu, and S. Z. Li, “Person re-identification by local maximal occurrence representation and metric learning,” in CVPR, 2015, pp. 2197–2206.
  35. 35.T. Matsukawa, T. Okabe, E. Suzuki, and Y. Sato, “Hierarchical gaussian descriptor for person re-identification,” in CVPR, 2016, pp. 1363–1372.
  36. 36.M. Kostinger, M. Hirzer, P. Wohlhart, and et al, “Large scale metric learning from equivalence constraints,” in CVPR, 2012, pp. 2288–2295.
  37. 37.W.-S. Zheng, S. Gong, and T. Xiang, “Person re-identification by probabilistic relative distance comparison,” in CVPR, 2011, pp. 649–656.
  38. 38.F. Xiong, M. Gou, O. Camps, and M. Sznaier, “Person re-identification using kernel-based metric learning methods,” in ECCV, 2014, pp. 1–16.
  39. 39.M. Hirzer, P. M. Roth, M. Kostinger, and H. Bischof, “Relaxed ¨ pairwise learned metric for person re-identification,” in ECCV, 2012, pp. 780–793.
  40. 40.S. Liao and S. Z. Li, “Efficient psd constrained asymmetric metric learning for person re-identification,” in ICCV, 2015, pp. 3685–3693.
  41. 41.H.-X. Yu, A. Wu, and W.-S. Zheng, “Unsupervised person re-identification by deep asymmetric metric embedding,” IEEE TPAMI, 2018.
  42. 42.Z. Zheng, L. Zheng, and Y. Yang, “Unlabeled samples generated by gan improve the person re-identification baseline in vitro,” in ICCV, 2017, pp. 3754–3762.
  43. 43.W. Li, R. Zhao, T. Xiao, and X. Wang, “Deepreid: Deep filter pairing neural network for person re-identification,” in CVPR, 2014, pp. 152–159.
  44. 44.L. Wei, S. Zhang, W. Gao, and Q. Tian, “Person transfer gan to bridge domain gap for person re-identification,” in CVPR, 2018, pp. 79–88.
  45. 45.Q. Leng, M. Ye, and Q. Tian, “A survey of open-world person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), 2019.
  46. 46.D. Wu, S.-J. Zheng, X.-P. Zhang, C.-A. Yuan, F. Cheng, Y. Zhao, Y.-J. Lin, Z.-Q. Zhao, Y.-L. Jiang, and D.-S. Huang, “Deep learning-based methods for person re-identification: A comprehensive review,” Neurocomputing, 2019.
  47. 47.B. Lavi, M. F. Serj, and I. Ullah, “Survey on deep learning techniques for person re-identification task,” arXiv preprint arXiv:1807.05284, 2018.
  48. 48.X. Wang, “Intelligent multi-camera video surveillance: A review,” Pattern recognition letters, vol. 34, no. 1, pp. 3–19, 2013.
  49. 49.D. Geronimo, A. M. Lopez, A. D. Sappa, and T. Graf, “Survey of pedestrian detection for advanced driver assistance systems,” IEEE TPAMI, no. 7, pp. 1239–1258, 2009.
  50. 50.P. Dollar, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: A benchmark,” in CVPR, 2009, pp. 304–311.
  51. 51.E. Insafutdinov, M. Andriluka, L. Pishchulin, S. Tang, E. Levinkov, B. Andres, and B. Schiele, “Arttrack: Articulated multi-person tracking in the wild,” in CVPR, 2017, pp. 6457–6465.
  52. 52.E. Ristani and C. Tomasi, “Features for multi-target multi-camera tracking and re-identification,” in CVPR, 2018, pp. 6036–6046.
  53. 53.A. J. Ma, P. C. Yuen, and J. Li, “Domain transfer support vector ranking for person re-identification without target camera label information,” in ICCV, 2013, pp. 3567–3574.
  54. 54.T. Xiao, H. Li, W. Ouyang, and X. Wang, “Learning deep feature representations with domain guided dropout for person re-identification,” in CVPR, 2016, pp. 1249–1258.
  55. 55.L. Zheng, H. Zhang, S. Sun, M. Chandraker, Y. Yang, and Q. Tian, “Person re-identification in the wild,” in CVPR, 2017, pp. 1367–1376.
  56. 56.D. Yi, Z. Lei, S. Liao, and S. Z. Li, “Deep metric learning for person re-identification,” in ICPR, 2014, pp. 34–39.
  57. 57.A. Hermans, L. Beyer, and B. Leibe, “In defense of the triplet loss for person re-identification,” arXiv preprint arXiv:1703.07737, 2017.
  58. 58.Z. Zhong, L. Zheng, D. Cao, and S. Li, “Re-ranking person re-identification with k-reciprocal encoding,” in CVPR, 2017, pp. 1318–1327.
  59. 59.M. Ye, C. Liang, Y. Yu, Z. Wang, Q. Leng, C. Xiao, J. Chen, and R. Hu, “Person reidentification via ranking aggregation of similarity pulling and dissimilarity pushing,” IEEE Transactions on Multimedia (TMM), vol. 18, no. 12, pp. 2553–2566, 2016.
  60. 60.D. T. Nguyen, H. G. Hong, K. W. Kim, and K. R. Park, “Person recognition system based on a combination of body images from visible light and thermal cameras,” Sensors, vol. 17, no. 3, p. 605, 2017.
  61. 61.W.-H. Li, Z. Zhong, and W.-S. Zheng, “One-pass person re-identification by sketch online discriminant analysis,” Pattern Recognition, vol. 93, pp. 237–250, 2019.
  62. 62.A. Wu, W.-S. Zheng, and J.-H. Lai, “Robust depth-based person re-identification,” IEEE Transactions on Image Processing (TIP), vol. 26, no. 6, pp. 2588–2603, 2017.
  63. 63.S. Li, T. Xiao, H. Li, B. Zhou, D. Yue, and X. Wang, “Person search with natural language description,” in CVPR, 2017, pp. 1345–1353.
  64. 64.T. Xiao, S. Li, B. Wang, L. Lin, and X. Wang, “Joint detection and identification feature learning for person search,” in CVPR, 2017, pp. 3415–3424.
  65. 65.X. Liu, M. Song, D. Tao, X. Zhou, C. Chen, and J. Bu, “Semi-supervised coupled dictionary learning for person re-identification,” in CVPR, 2014, pp. 3550–3557.
  66. 66.R. Zhao, W. Ouyang, and X. Wang, “Unsupervised salience learning for person re-identification,” in CVPR, 2013, pp. 3586–3593.
  67. 67.Y. Sun, Q. Xu, Y. Li, C. Zhang, Y. Li, S. Wang, and J. Sun, “Perceive where to focus: Learning visibility-aware part-level features for partial person re-identification,” in CVPR, 2019, pp. 393–402.
  68. 68.X. Wang, G. Doretto, T. Sebastian, J. Rittscher, and P. Tu, “Shape and appearance context modeling,” in ICCV, 2007, pp. 1–8.
  69. 69.H. Wang, X. Zhu, T. Xiang, and S. Gong, “Towards unsupervised open-set person re-identification,” in ICIP, 2016, pp. 769–773.
  70. 70.X. Zhu, B. Wu, D. Huang, and W.-S. Zheng, “Fast open-world person re-identification,” IEEE Transactions on Image Processing (TIP), vol. 27, no. 5, pp. 2286 – 2300, 2018.
  71. 71.C. Su, S. Zhang, J. Xing, W. Gao, and Q. Tian, “Deep attributes driven multi-camera person re-identification,” in ECCV, 2016, pp. 475–491.
  72. 72.Y. Lin, L. Zheng, Z. Zheng, Y. Wu, and Y. Yang, “Improving person re-identification by attribute and identity learning,” arXiv preprint arXiv:1703.07220, 2017.
  73. 73.K. Liu, B. Ma, W. Zhang, and R. Huang, “A spatio-temporal appearance representation for viceo-based pedestrian re-identification,” in ICCV, 2015, pp. 3810–3818.
  74. 74.J. Dai, P. Zhang, D. Wang, H. Lu, and H. Wang, “Video person re-identification by temporal residual learning,” IEEE Transactions on Image Processing (TIP), vol. 28, no. 3, pp. 1366–1377, 2018.
  75. 75.L. Zhao, X. Li, Y. Zhuang, and J. Wang, “Deeply-learned part-aligned representations for person re-identification,” in CVPR, 2017, pp. 3219–3228.
  76. 76.H. Yao, S. Zhang, R. Hong, Y. Zhang, C. Xu, and Q. Tian, “Deep representation learning with part loss for person re-identification,” IEEE Transactions on Image Processing (TIP), 2019.
  77. 77.Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling,” in ECCV, 2018, pp. 480–496.
  78. 78.T. Matsukawa and E. Suzuki, “Person re-identification using cnn features learned from combination of attributes,” in ICPR, 2016, pp. 2428–2433.
  79. 79.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014.
  80. 80.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778.
  81. 81.F. Wang, W. Zuo, L. Lin, D. Zhang, and L. Zhang, “Joint learning of single-image and cross-image representations for person re-identification,” in CVPR, 2016, pp. 1288–1296.
  82. 82.Y. Sun, L. Zheng, W. Deng, and S. Wang, “Svdnet for pedestrian retrieval,” in ICCV, 2017, pp. 3800–3808.
  83. 83.M. Ye, X. Lan, and P. C. Yuen, “Robust anchor embedding for unsupervised video person re-identification in the wild,” in ECCV, 2018, pp. 170–186.
  84. 84.X. Qian, Y. Fu, Y.-G. Jiang, T. Xiang, and X. Xue, “Multi-scale deep learning architectures for person re-identification,” in ICCV, 2017, pp. 5399–5408.
  85. 85.F. Yang, K. Yan, S. Lu, H. Jia, X. Xie, and W. Gao, “Attention driven person re-identification,” Pattern Recognition, vol. 86, pp. 143–155, 2019.
  86. 86.W. Li, X. Zhu, and S. Gong, “Harmonious attention network for person re-identification,” in CVPR, 2018, pp. 2285–2294.
  87. 87.C. Wang, Q. Zhang, C. Huang, W. Liu, and X. Wang, “Mancs: A multi-task attentional network with curriculum sampling for person re-identification,” in ECCV, 2018, pp. 365–381.
  88. 88.Y. Shen, T. Xiao, H. Li, S. Yi, and X. Wang, “End-to-end deep kronecker-product matching for person re-identification,” in CVPR, 2018, pp. 6886–6895.
  89. 89.Y. Wang, Z. Chen, F. Wu, and G. Wang, “Person re-identification with cascaded pairwise convolutions,” in CVPR, 2018, pp. 1470–1478.
  90. 90.G. Chen, C. Lin, L. Ren, J. Lu, and J. Zhou, “Self-critical attention learning for person re-identification,” in ICCV, 2019, pp. 9637–9646.
  91. 91.J. Si, H. Zhang, C.-G. Li, J. Kuen, X. Kong, A. C. Kot, and G. Wang, “Dual attention matching network for context-aware feature sequence based person re-identification,” in CVPR, 2018, pp. 5363–5372.
  92. 92.M. Zheng, S. Karanam, Z. Wu, and R. J. Radke, “Re-identification with consistent attentive siamese networks,” in CVPR, 2019, pp. 5735–5744.
  93. 93.S. Zhou, F. Wang, Z. Huang, and J. Wang, “Discriminative feature learning with consistent attention regularization for person re-identification,” in ICCV, 2019, pp. 8040–8049.
  94. 94.D. Chen, D. Xu, H. Li, N. Sebe, and X. Wang, “Group consistent similarity learning via deep crf for person re-identification,” in CVPR, 2018, pp. 8649–8658.
  95. 95.C. Luo, Y. Chen, N. Wang, and Z. Zhang, “Spectral feature transformation for person re-identification,” in ICCV, 2019, pp. 4976–4985.
  96. 96.R. R. Varior, B. Shuai, J. Lu, D. Xu, and G. Wang, “A siamese long short-term memory architecture for human re-identification,” in ECCV, 2016, pp. 135–153.
  97. 97.Y. Suh, J. Wang, S. Tang, T. Mei, and K. Mu Lee, “Part-aligned bilinear representations for person re-identification,” in ECCV, 2018, pp. 402–419.
  98. 98.L. Zhao, X. Li, Y. Zhuang, and J. Wang, “Deeply-learned part-aligned representations for person re-identification,” in ICCV, 2017, pp. 3219–3228.
  99. 99.D. Cheng, Y. Gong, S. Zhou, J. Wang, and N. Zheng, “Person re-identification by multi-channel parts-based cnn with improved triplet loss function,” in CVPR, 2016, pp. 1335–1344.
  100. 100.D. Li, X. Chen, Z. Zhang, and K. Huang, “Learning deep context-aware features over body and latent parts for person re-identification,” in CVPR, 2017, pp. 384–393.
  101. 101.C. Su, J. Li, S. Zhang, J. Xing, W. Gao, and Q. Tian, “Pose-driven deep convolutional model for person re-identification,” in ICCV, 2017, pp. 3960–3969.
  102. 102.J. Xu, R. Zhao, F. Zhu, H. Wang, and W. Ouyang, “Attention-aware compositional network for person re-identification,” in CVPR, 2018, pp. 2119–2128.
  103. 103.Z. Zhang, C. Lan, W. Zeng, and Z. Chen, “Densely semantically aligned person re-identification,” in CVPR, 2019, pp. 667–676.
  104. 104.J. Guo, Y. Yuan, L. Huang, C. Zhang, J.-G. Yao, and K. Han, “Beyond human parts: Dual part-aligned representations for person re-identification,” in ICCV, 2019, pp. 3642–3651.
  105. 105.X. Sun and L. Zheng, “Dissecting person re-identification from the viewpoint of viewpoint,” in CVPR, 2019, pp. 608–617.
  106. 106.Z. Zhong, L. Zheng, Z. Luo, S. Li, and Y. Yang, “Invariance matters: Exemplar memory for domain adaptive person re-identification,” in CVPR, 2019, pp. 598–607.
  107. 107.B. N. Xia, Y. Gong, Y. Zhang, and C. Poellabauer, “Second-order non-local attention networks for person re-identification,” in ICCV, 2019, pp. 3760–3769.
  108. 108.R. Hou, B. Ma, H. Chang, X. Gu, S. Shan, and X. Chen, “Interaction-and-aggregation network for person re-identification,” in CVPR, 2019, pp. 9317–9326.
  109. 109.C.-P. Tay, S. Roy, and K.-H. Yap, “Aanet: Attribute attention network for person re-identifications,” in CVPR, 2019, pp. 7134–7143.
  110. 110.Y. Zhao, X. Shen, Z. Jin, H. Lu, and X.-s. Hua, “Attribute-driven feature disentangling and temporal aggregation for video person re-identification,” in CVPR, 2019, pp. 4913–4922.
  111. 111.W. Jingya, Z. Xiatian, G. Shaogang, and L. Wei, “Transferable joint attribute-identity deep learning for unsupervised person re-identification,” in CVPR, 2018, pp. 2275–2284.
  112. 112.X. Chang, T. M. Hospedales, and T. Xiang, “Multi-level factorisation net for person re-identification,” in CVPR, 2018, pp. 2109–2118.
  113. 113.F. Liu and L. Zhang, “View confusion feature learning for person re-identification,” in ICCV, 2019, pp. 6639–6648.
  114. 114.Z. Zhu, X. Jiang, F. Zheng, X. Guo, F. Huang, W. Zheng, and X. Sun, “Aware loss with angular regularization for person re-identification,” in AAAI, 2020.
  115. 115.J. Lin, L. Ren, J. Lu, J. Feng, and J. Zhou, “Consistent-aware deep learning for person re-identification in a camera network,” in CVPR, 2017, pp. 5771–5780.
  116. 116.J. Liu, B. Ni, Y. Yan, and et al., “Pose transferrable person re-identification,” in CVPR, 2018, pp. 4099–4108.
  117. 117.X. Qian, Y. Fu, T. Xiang, W. Wang, J. Qiu, Y. Wu, Y.-G. Jiang, and X. Xue, “Pose-normalized image generation for person re-identification,” in ECCV, 2018, pp. 650–667.
  118. 118.Z. Zhong, L. Zheng, Z. Zheng, and et al., “Camera style adaptation for person re-identification,” in CVPR, 2018, pp. 5157–5166.
  119. 119.Z. Zheng, X. Yang, Z. Yu, L. Zheng, Y. Yang, and J. Kautz, “Joint discriminative and generative learning for person re-identification,” in CVPR, 2019, pp. 2138–2147.
  120. 120.W. Deng, L. Zheng, Q. Ye, G. Kang, Y. Yang, and J. Jiao, “Image-image domain adaptation with preserved self-similarity and domain-dissimilarity for person re-identification,” in CVPR, 2018, pp. 994–1003.
  121. 121.Y.-J. Li, C.-S. Lin, Y.-B. Lin, and Y.-C. F. Wang, “Cross-dataset person re-identification via unsupervised pose disentanglement and adaptation,” in ICCV, 2019, pp. 7919–7929.
  122. 122.H. Luo, W. Jiang, Y. Gu, F. Liu, X. Liao, S. Lai, and J. Gu, “A strong baseline and batch normneuralization neck for deep person re-identification,” arXiv preprint arXiv:1906.08332, 2019.
  123. 123.Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” arXiv preprint arXiv:1708.04896, 2017.
  124. 124.Z. Dai, M. Chen, X. Gu, S. Zhu, and P. Tan, “Batch dropblock network for person re-identification and beyond,” in ICCV, 2019, pp. 3691–3701.
  125. 125.S. Bak, P. Carr, and J.-F. Lalonde, “Domain adaptation through synthesis for unsupervised person re-identification,” in ECCV, 2018, pp. 189–205.
  126. 126.M. Hirzer, C. Beleznai, P. M. Roth, and H. Bischof, “Person re-identification by descriptive and discriminative classification,” in Image Analysis, 2011, pp. 91–102.
  127. 127.N. McLaughlin, J. Martinez del Rincon, and P. Miller, “Recurrent convolutional network for video-based person re-identification,” in CVPR, 2016, pp. 1325–1334.
  128. 128.D. Chung, K. Tahboub, and E. J. Delp, “A two stream siamese convolutional neural network for person re-identification,” in ICCV, 2017, pp. 1983–1991.
  129. 129.Y. Yan, B. Ni, Z. Song, C. Ma, Y. Yan, and X. Yang, “Person re-identification via recurrent feature aggregation,” in ECCV, 2016, pp. 701–716.
  130. 130.Z. Zhou, Y. Huang, W. Wang, L. Wang, and T. Tan, “See the forest for the trees: Joint spatial and temporal recurrent neural networks for video-based person re-identification,” in CVPR, 2017, pp. 4747–4756.
  131. 131.S. Xu, Y. Cheng, K. Gu, Y. Yang, S. Chang, and P. Zhou, “Jointly attentive spatial-temporal pooling networks for video-based person re-identification,” in ICCV, 2017, pp. 4733–4742.
  132. 132.A. Subramaniam, A. Nambiar, and A. Mittal, “Co-segmentation inspired attention networks for video-based person re-identification,” in ICCV, 2019, pp. 562–572.
  133. 133.S. Li, S. Bak, P. Carr, and X. Wang, “Diversity regularized spatiotemporal attention for video-based person re-identification,” in CVPR, 2018, pp. 369–378.
  134. 134.D. Chen, H. Li, T. Xiao, S. Yi, and X. Wang, “Video person re-identification with competitive snippet-similarity aggregation and co-attentive snippet embedding,” in CVPR, 2018, pp. 1169–1178.
  135. 135.Y. Fu, X. Wang, Y. Wei, and T. Huang, “Sta: Spatial-temporal attention for large-scale video-based person re-identification,” in AAAI, 2019.
  136. 136.J. Li, J. Wang, Q. Tian, W. Gao, and S. Zhang, “Global-local temporal representations for video person re-identification,” in ICCV, 2019, pp. 3958–3967.
  137. 137.Y. Guo and N.-M. Cheung, “Efficient and deep person re-identification using multi-level similarity,” in CVPR, 2018, pp. 2335–2344.
  138. 138.K. Zhou, Y. Yang, A. Cavallaro, and T. Xiang, “Omni-scale feature learning for person re-identification,” in ICCV, 2019, pp. 3702–3712.
  139. 139.R. Quan, X. Dong, Y. Wu, L. Zhu, and Y. Yang, “Auto-reid: Searching for a part-aware convnet for person re-identification,” in ICCV, 2019, pp. 3750–3759.
  140. 140.M. Tian, S. Yi, H. Li, S. Li, X. Zhang, J. Shi, J. Yan, and X. Wang, “Eliminating background-bias for robust person re-identification,” in CVPR, 2018, pp. 5794–5803.
  141. 141.Z. Zheng, L. Zheng, and Y. Yang, “A discriminatively learned cnn embedding for person re-identification,” arXiv preprint arXiv:1611.05666, 2016.
  142. 142.M. Ye, X. Lan, Z. Wang, and P. C. Yuen, “Bi-directional center-constrained top-ranking for visible thermal person re-identification,” IEEE Transactions on Information Forensics and Security (TIFS), vol. 15, pp. 407–419, 2020.
  143. 143.P. Moutafis, M. Leng, and I. A. Kakadiaris, “An overview and empirical comparison of distance metric learning methods,” IEEE Transactions on Cybernetics, vol. 47, no. 3, pp. 612–625, 2016.
  144. 144.Y. Wu, Y. Lin, X. Dong, Y. Yan, W. Ouyang, and Y. Yang, “Exploit the unknown gradually: One-shot video-based person re-identification by stepwise learning,” in CVPR, 2018, pp. 5177–5186.
  145. 145.M. Ye, X. Zhang, P. C. Yuen, and S.-F. Chang, “Unsupervised embedding learning via invariant and spreading instance feature,” in CVPR, 2019, pp. 6210–6219.
  146. 146.N. Wojke and A. Bewley, “Deep cosine metric learning for person re-identification,” in WACV, 2018, pp. 748–756.
  147. 147.X. Fan, W. Jiang, H. Luo, and M. Fei, “Spherereid: Deep hypersphere manifold embedding for person re-identification,” JVCIR, vol. 60, pp. 51–58, 2019.
  148. 148.R. Muller, S. Kornblith, and G. Hinton, “When does label smoothing help?” arXiv preprint arXiv:1906.02629, 2019.
  149. 149.H. Shi, Y. Yang, X. Zhu, S. Liao, Z. Lei, W. Zheng, and S. Z. Li, “Embedding deep metric for person re-identification: A study against large variations,” in ECCV, 2016, pp. 732–748.
  150. 150.S. Zhou, J. Wang, J. Wang, Y. Gong, and N. Zheng, “Point to set similarity based deep feature learning for person re-identification,” in CVPR, 2017, pp. 3741–3750.
  151. 151.R. Yu, Z. Dou, S. Bai, Z. Zhang, Y. Xu, and X. Bai, “Hard-aware point-to-set deep metric for person re-identification,” in ECCV, 2018, pp. 188–204.
  152. 152.W. Chen, X. Chen, J. Zhang, and K. Huang, “Beyond triplet loss: a deep quadruplet network for person re-identification,” in CVPR, 2017, pp. 403–412.
  153. 153.Q. Yang, H.-X. Yu, A. Wu, and W.-S. Zheng, “Patch-based discriminative feature learning for unsupervised person re-identification,” in CVPR, 2019, pp. 3633–3642.
  154. 154.Z. Liu, J. Wang, S. Gong, H. Lu, and D. Tao, “Deep reinforcement active learning for human-in-the-loop person re-identification,” in ICCV, 2019, pp. 6122–6131.
  155. 155.J. Zhou, B. Su, and Y. Wu, “Easy identification from better constraints: Multi-shot person re-identification from reference constraints,” in CVPR, 2018, pp. 5373–5381.
  156. 156.F. Zheng, C. Deng, X. Sun, X. Jiang, X. Guo, Z. Yu, F. Huang, and R. Ji, “Pyramidal person re-identification via multi-loss dynamic training,” in CVPR, 2019, pp. 8514–8522.
  157. 157.M. Ye, C. Liang, Z. Wang, Q. Leng, and J. Chen, “Ranking optimization for person re-identification via similarity and dissimilarity,” in ACM Multimedia (ACM MM), 2015, pp. 1239–1242.
  158. 158.C. Liu, C. Change Loy, S. Gong, and G. Wang, “Pop: Person re-identification post-rank optimisation,” in ICCV, 2013, pp. 441–448.
  159. 159.H. Wang, S. Gong, X. Zhu, and T. Xiang, “Human-in-the-loop person re-identification,” in ECCV, 2016, pp. 405–422.
  160. 160.S. Paisitkriangkrai, C. Shen, and A. Van Den Hengel, “Learning to rank in person re-identification with metric ensembles,” in CVPR, 2015, pp. 1846–1855.
  161. 161.S. Bai, P. Tang, P. H. Torr, and L. J. Latecki, “Re-ranking via metric fusion for object retrieval and person re-identification,” in CVPR, 2019, pp. 740–749.
  162. 162.S. Bai, X. Bai, and Q. Tian, “Scalable person re-identification on supervised smoothed manifold,” in CVPR, 2017, pp. 2530–2539.
  163. 163.A. J. Ma and P. Li, “Query based adaptive re-ranking for person re-identification,” in ACCV, 2014, pp. 397–412.
  164. 164.J. Zhou, P. Yu, W. Tang, and Y. Wu, “Efficient online local metric adaptation via negative samples for person re-identification,” in ICCV, 2017, pp. 2420–2428.
  165. 165.L. Zheng, S. Wang, L. Tian, F. He, Z. Liu, and Q. Tian, “Query-adaptive late fusion for image search and person re-identification,” in CVPR, 2015, pp. 1741–1750.
  166. 166.A. Barman and S. K. Shah, “Shape: A novel graph theoretic algorithm for making consensus-based decisions in person re-identification systems,” in ICCV, 2017, pp. 1115–1124.
  167. 167.W.-S. Zheng, S. Gong, and T. Xiang, “Associating groups of people,” in BMVC, 2009, pp. 1–23.
  168. 168.C. C. Loy, C. Liu, and S. Gong, “Person re-identification by manifold ranking,” in ICIP, 2013, pp. 3567–3571.
  169. 169.M. Gou, Z. Wu, A. Rates-Borras, O. Camps, R. J. Radke et al., “A systematic evaluation and benchmark for person re-identification: Features, metrics, and datasets,” IEEE TPAMI, vol. 41, no. 3, pp. 523–536, 2018.
  170. 170.M. Li, X. Zhu, and S. Gong, “Unsupervised person re-identification by deep learning tracklet association,” in ECCV, 2018, pp. 737–753.
  171. 171.G. Song, B. Leng, Y. Liu, C. Hetang, and S. Cai, “Region-based quality estimation network for large-scale person re-identification,” in AAAI, 2018, pp. 7347–7354.
  172. 172.G. Wang, Y. Yuan, X. Chen, J. Li, and X. Zhou, “Learning discriminative features with multiple granularities for person re-identification,” in ACM MM, 2018, pp. 274–282.
  173. 173.T. Chen, S. Ding, J. Xie, Y. Yuan, W. Chen, Y. Yang, Z. Ren, and Z. Wang, “Abd-net: Attentive but diverse person re-identification,” in ICCV, 2019, pp. 8351–8361.
  174. 174.B. Chen, W. Deng, and J. Hu, “Mixed high-order attention network for person re-identification,” in ICCV, 2019, pp. 371–381.
  175. 175.X. Zhang, H. Luo, X. Fan, W. Xiang, Y. Sun, Q. Xiao, W. Jiang, C. Zhang, and J. Sun, “Alignedreid: Surpassing human-level performance in person re-identification,” arXiv preprint arXiv:1711.08184, 2017.
  176. 176.M. Ye, A. J. Ma, L. Zheng, J. Li, and P. C. Yuen, “Dynamic label graph matching for unsupervised video re-identification,” in ICCV, 2017, pp. 5142–5150.
  177. 177.Z. Wang, S. Zheng, M. Song, Q. Wang, A. Rahimpour, and H. Qi, “advpattern: Physical-world attacks on deep person re-identification via adversarially transformable patterns,” in ICCV, 2019, pp. 8341–8350.
  178. 178.J. Zhang, N. Wang, and L. Zhang, “Multi-shot pedestrian re-identification via sequential decision making,” in CVPR, 2018, pp. 6781–6789.
  179. 179.A. Haque, A. Alahi, and L. Fei-Fei, “Recurrent attention models for depth-based person identification,” in CVPR, 2016, pp. 1229–1238.
  180. 180.N. Karianakis, Z. Liu, Y. Chen, and S. Soatto, “Reinforced temporal attention and split-rate transfer for depth-based person re-identification,” in ECCV, 2018, pp. 715–733.
  181. 181.I. B. Barbosa, M. Cristani, A. Del Bue, L. Bazzani, and V. Murino, “Re-identification with rgb-d sensors,” in ECCV Workshop, 2012, pp. 433–442.
  182. 182.D. Chen, H. Li, X. Liu, Y. Shen, J. Shao, Z. Yuan, and X. Wang, “Improving deep visual representation for person re-identification by global and local image-language association,” in ECCV, 2018, pp. 54–70.
  183. 183.Y. Zhang and H. Lu, “Deep cross-modal projection learning for image-text matching,” in ECCV, 2018, pp. 686–701.
  184. 184.J. Liu, Z.-J. Zha, R. Hong, M. Wang, and Y. Zhang, “Deep adversarial graph attention convolution network for text-based person search,” in ACM MM, 2019, pp. 665–673.
  185. 185.M. Ye, Z. Wang, X. Lan, and P. C. Yuen, “Visible thermal person re-identification via dual-constrained top-ranking,” in IJCAI, 2018, pp. 1092–1099.
  186. 186.M. Ye, X. Lan, J. Li, and P. C. Yuen, “Hierarchical discriminative learning for visible thermal person re-identification,” in AAAI, 2018, pp. 7501–7508.
  187. 187.Y. Hao, N. Wang, J. Li, and X. Gao, “Hsme: Hypersphere manifold embedding for visible thermal person re-identification,” in AAAI, 2019, pp. 8385–8392.
  188. 188.M. Ye, J. Shen, and L. Shao, “Visible-infrared person re-identification via homogeneous augmented tri-modal learning,” IEEE TIFS, 2020.
  189. 189.Z. Wang, Z. Wang, Y. Zheng, Y.-Y. Chuang, and S. Satoh, “Learning to reduce dual-level discrepancy for infrared-visible person re-identification,” in CVPR, 2019, pp. 618–626.
  190. 190.G. Wang, T. Zhang, J. Cheng, S. Liu, Y. Yang, and Z. Hou, “Rgb-infrared cross-modality person re-identification via joint pixel and feature alignment,” in ICCV, 2019, pp. 3623–3632.
  191. 191.S. Choi, S. Lee, Y. Kim, T. Kim, and C. Kim, “Hi-cmd: Hierarchical cross-modality disentanglement for visible-infrared person re-identification,” in CVPR, 2020, pp. 10 257–10 266.
  192. 192.M. Ye, J. Shen, D. J. Crandall, L. Shao, and J. Luo, “Dynamic dual-attentive aggregation learning for visible-infrared person re-identification,” in ECCV, 2020.
  193. 193.Z. Wang, M. Ye, F. Yang, X. Bai, and S. Satoh, “Cascaded sr-gan for scale-adaptive low resolution person re-identification.” in IJCAI, 2018, pp. 3891–3897.
  194. 194.Y.-J. Li, Y.-C. Chen, Y.-Y. Lin, X. Du, and Y.-C. F. Wang, “Recover and identify: A generative dual model for cross-resolution person re-identification,” in ICCV, 2019, pp. 8090–8099.
  195. 195.H. Liu, J. Feng, Z. Jie, K. Jayashree, B. Zhao, M. Qi, J. Jiang, and S. Yan, “Neural person search machines,” in ICCV, 2017, pp. 493–501.
  196. 196.Y. Yan, Q. Zhang, B. Ni, W. Zhang, M. Xu, and X. Yang, “Learning context graph for person search,” in CVPR, 2019, pp. 2158–2167.
  197. 197.B. Munjal, S. Amin, F. Tombari, and F. Galasso, “Query-guided end-to-end person search,” in CVPR, 2019, pp. 811–820.
  198. 198.C. Han, J. Ye, Y. Zhong, X. Tan, C. Zhang, C. Gao, and N. Sang, “Re-id driven localization refinement for person search,” in ICCV, 2019, pp. 9814–9823.
  199. 199.X. Lan, H. Wang, S. Gong, and X. Zhu, “Deep reinforcement learning attention selection for person re-identification,” in BMVC, 2017.
  200. 200.M. Yamaguchi, K. Saito, Y. Ushiku, and T. Harada, “Spatio-temporal person retrieval via natural language queries,” in ICCV, 2017, pp. 1453–1462.
  201. 201.S. Tang, M. Andriluka, B. Andres, and B. Schiele, “Multiple people tracking by lifted multicut and person re-identification,” in CVPR, 2017, pp. 3539–3548.
  202. 202.Y. Hou, L. Zheng, Z. Wang, and S. Wang, “Locality aware appearance metric for multi-target multi-camera tracking,” arXiv preprint arXiv:1911.12037, 2019.
  203. 203.E. Kodirov, T. Xiang, Z. Fu, and S. Gong, “Person re-identification by unsupervised l1 graph learning,” in ECCV, 2016, pp. 178–195.
  204. 204.Z. Liu, D. Wang, and H. Lu, “Stepwise metric promotion for unsupervised video person re-identification,” in ICCV, 2017, pp. 2429–2438.
  205. 205.H. Fan, L. Zheng, and Y. Yang, “Unsupervised person re-identification: Clustering and fine-tuning,” arXiv preprint arXiv:1705.10444, 2017.
  206. 206.M. Ye, J. Li, A. J. Ma, L. Zheng, and P. C. Yuen, “Dynamic graph co-matching for unsupervised video-based person re-identification,” IEEE TIP, vol. 28, no. 6, pp. 2976–2990, 2019.
  207. 207.X. Wang, R. Panda, M. Liu, Y. Wang, and et al., “Exploiting global camera network constraints for unsupervised video person re-identification,” arXiv preprint arXiv:1908.10486, 2019.
  208. 208.K. Zeng, M. Ning, Y. Wang, and Y. Guo, “Hierarchical clustering with hard-batch triplet loss for person re-identification,” in CVPR, 2020, pp. 13 657–13 665.
  209. 209.H.-X. Yu, W.-S. Zheng, A. Wu, X. Guo, S. Gong, and J.-H. Lai, “Unsupervised person re-identification by soft multilabel learning,” in CVPR, 2019, pp. 2148–2157.
  210. 210.A. Wu, W.-S. Zheng, and J.-H. Lai, “Unsupervised person re-identification by camera-aware similarity consistency learning,” in ICCV, 2019, pp. 6922–6931.
  211. 211.J. Wu, Y. Yang, H. Liu, S. Liao, Z. Lei, and S. Z. Li, “Unsupervised graph association for person re-identification,” in ICCV, 2019, pp. 8321–8330.
  212. 212.Y. Fu, Y. Wei, G. Wang, Y. Zhou, H. Shi, and T. S. Huang, “Self-similarity grouping: A simple unsupervised cross domain adaptation approach for person re-identification,” in ICCV, 2019, pp. 6112–6121.
  213. 213.S. Bak and P. Carr, “One-shot metric learning for person re-identification,” in CVPR, 2017, pp. 2990–2999.
  214. 214.X. Wang, S. Paul, D. S. Raychaudhuri, and at al., “Learning person re-identification models from videos with weak supervision,” arXiv preprint arXiv:2007.10631, 2020.
  215. 215.Z. Zhong, L. Zheng, S. Li, and Y. Yang, “Generalizing a person retrieval model hetero-and homogeneously,” in ECCV, 2018, pp. 172–188.
  216. 216.J. Liu, Z.-J. Zha, D. Chen, R. Hong, and M. Wang, “Adaptive transfer network for cross-domain person re-identification,” in CVPR, 2019, pp. 7202–7211.
  217. 217.Y. Huang, Q. Wu, J. Xu, and Y. Zhong, “Sbsgan: Suppression of inter-domain background shift for person re-identification,” in ICCV, 2019, pp. 9527–9536.
  218. 218.Y. Chen, X. Zhu, and S. Gong, “Instance-guided context rendering for cross-domain person re-identification,” in ICCV, 2019, pp. 232–242.
  219. 219.Y. Ge, D. Chen, and H. Li, “Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification,” in ICLR, 2020.
  220. 220.Y. Wang, S. Liao, and L. Shao, “Surpassing real-world source training data: Random 3d characters for generalizable person re-identification,” in ACM MM, 2020, pp. 3422–3430.
  221. 221.L. Qi, L. Wang, J. Huo, L. Zhou, Y. Shi, and Y. Gao, “A novel unsupervised camera-aware domain adaptation framework for person re-identification,” in ICCV, 2019, pp. 8080–8089.
  222. 222.X. Zhang, J. Cao, C. Shen, and M. You, “Self-training with progressive augmentation for unsupervised cross-domain person re-identification,” in ICCV, 2019, pp. 8222–8231.
  223. 223.Y. Ge, F. Zhu, D. Chen, R. Zhao, and H. Li, “Self-paced contrastive learning with hybrid memory for domain adaptive object re-id,” in NeurIPS, 2020.
  224. 224.J. Lv, W. Chen, Q. Li, and C. Yang, “Unsupervised cross-dataset person re-identification by transfer learning of spatial-temporal patterns,” in CVPR, 2018, pp. 7948–7956.
  225. 225.S. Liao and L. Shao, “Interpretable and generalizable person re-identification with query-adaptive convolution and temporal lifting,” in ECCV, 2020.
  226. 226.H.-X. Yu, A. Wu, and W.-S. Zheng, “Cross-view asymmetric metric learning for unsupervised person re-identification,” in ICCV, 2017, pp. 994–1002.
  227. 227.X. Jin, C. Lan, W. Zeng, Z. Chen, and L. Zhang, “Style normalization and restitution for generalizable person re-identification,” in CVPR, 2020, pp. 3143–3152.
  228. 228.Y. Zhai, Q. Ye, S. Lu, M. Jia, R. Ji, and Y. Tian, “Multiple expert brainstorming for domain adaptive person re-identification,” in ECCV, 2020.
  229. 229.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR, 2020.
  230. 230.M. Ye, J. Shen, X. Zhang, P. C. Yuen, and S.-F. Chang, “Augmentation invariant and instance spreading feature for softmax embedding,” IEEE TPAMI, 2020.
  231. 231.W.-S. Zheng, X. Li, T. Xiang, S. Liao, J. Lai, and S. Gong, “Partial person re-identification,” in ICCV, 2015, pp. 4678–4686.
  232. 232.L. He, J. Liang, H. Li, and Z. Sun, “Deep spatial feature reconstruction for partial person re-identification: Alignment-free approach,” in CVPR, 2018, pp. 7073–7082.
  233. 233.L. He, Y. Wang, W. Liu, X. Liao, H. Zhao, Z. Sun, and J. Feng, “Foreground-aware pyramid reconstruction for alignment-free occluded person re-identification,” in ICCV, 2019, pp. 8450–8459.
  234. 234.J. Miao, Y. Wu, P. Liu, Y. Ding, and Y. Yang, “Pose-guided feature alignment for occluded person re-identification,” in ICCV, 2019, pp. 542–551.
  235. 235.T. Yu, D. Li, Y. Yang, T. Hospedales, and T. Xiang, “Robust person re-identification by modelling feature uncertainty,” in ICCV, 2019, pp. 552–561.
  236. 236.M. Ye and P. C. Yuen, “Purifynet: A robust person re-identification model with noisy labels,” IEEE TIFS, 2020.
  237. 237.X. Li, A. Wu, and W.-S. Zheng, “Adversarial open-world person re-identification,” in ECCV, 2018, pp. 280–296.
  238. 238.M. Golfarelli, D. Maio, and D. Malton, “On the error-reject trade-off in biometric verification systems,” IEEE TPAMI, vol. 19, no. 7, pp. 786–796, 1997.
  239. 239.G. Lisanti, N. Martinel, A. Del Bimbo, and G. Luca Foresti, “Group re-identification via unsupervised transfer of sparse features encoding,” in ICCV, 2017, pp. 2449–2458.
  240. 240.Y. Cai, V. Takala, and M. Pietikainen, “Matching groups of people by covariance descriptor,” in ICPR, 2010, pp. 2744–2747.
  241. 241.H. Xiao, W. Lin, B. Sheng, K. Lu, J. Yan, and et al., “Group re-identification: Leveraging and integrating multi-grain information,” in ACM MM, 2018, pp. 192–200.
  242. 242.Z. Huang, Z. Wang, W. Hu, C.-W. Lin, and S. Satoh, “Dot-gnn: Domain-transferred graph neural network for group re-identification,” in ACM MM, 2019, pp. 1888–1896.
  243. 243.Y. Shen, H. Li, S. Yi, D. Chen, and X. Wang, “Person re-identification with deep similarity-guided graph neural network,” in ECCV, 2018, pp. 486–504.
  244. 244.R. Panda, A. Bhuiyan, V. Murino, and A. K. Roy-Chowdhury, “Unsupervised adaptive re-identification in open world dynamic camera networks,” in CVPR, 2017, pp. 7054–7063.
  245. 245.S. M. Assari, H. Idrees, and M. Shah, “Human re-identification in crowd videos using personal, social and environmental constraints,” in ECCV, 2016, pp. 119–136.
  246. 246.X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR, 2018, pp. 7794–7803.
  247. 247.F. Radenovic, G. Tolias, and O. Chum, “Fine-tuning cnn image retrieval with no human annotation,” IEEE TPAMI, vol. 41, no. 7, pp. 1655–1668, 2018.
  248. 248.X. Wang, X. Han, W. Huang, D. Dong, and M. R. Scott, “Multi-similarity loss with general pair weighting for deep metric learning,” in CVPR, 2019, pp. 5022–5030.
  249. 249.L. He, Z. Sun, Y. Zhu, and Y. Wang, “Recognizing partial biometric patterns.” arXiv preprint arXiv:1810.07399, 2018.
  250. 250.J. Xue, Z. Meng, K. Katipally, H. Wang, and K. van Zon, “Clothing change aware person identification,” in CVPR Workshops, 2018, pp. 2112–2120.
  251. 251.F. Wan, Y. Wu, X. Qian, and Y. Fu, “When person re-identification meets changing clothes,” arXiv preprint arXiv:2003.04070, 2020.
  252. 252.S. Roy, S. Paul, N. E. Young, and A. K. Roy-Chowdhury, “Exploiting transitivity for learning person re-identification models on a budget,” in CVPR, 2018, pp. 7064–7072.
  253. 253.Z. Zheng and Y. Yang, “Person re-identification in the 3d space,” arXiv preprint arXiv:2006.04569, 2020.
  254. 254.Y. Hu, D. Yi, S. Liao, Z. Lei, and S. Z. Li, “Cross dataset person re-identification,” in ACCV, 2014, pp. 650–664.
  255. 255.G. Wu and S. Gong, “Decentralised learning from independent multi-domain labels for person re-identification,” arXiv preprint arXiv:2006.04150, 2020.
  256. 256.P. Bhargava, “Incremental learning in person re-identification,” arXiv preprint arXiv:1808.06281, 2018.
  257. 257.F. Zhu, X. Kong, L. Zheng, H. Fu, and Q. Tian, “Part-based deep hashing for large-scale person re-identification,” IEEE Transactions on Image Processing (TIP), vol. 26, no. 10, pp. 4806–4817, 2017.
  258. 258.J. Chen, Y. Wang, J. Qin, L. Liu, and L. Shao, “Fast person re-identification via cross-camera semantic binary transformation,” in CVPR, 2017, pp. 3873–3882.
  259. 259.G. Wang, S. Gong, J. Cheng, and Z. Hou, “Faster person re-identification,” in ECCV, 2020, pp. 275–292.
  260. 260.A. Wu, W.-S. Zheng, X. Guo, and J.-H. Lai, “Distilled person re-identification: Towards a more scalable system,” in CVPR, 2019, pp. 1187–1196.
  261. 261.M. Ye, X. Lan, and Q. Leng, “Modality-aware collaborative learning for visible thermal person re-identification,” in ACM MM, 2019, pp. 347–355.
  262. 262.Z. Feng, J. Lai, and X. Xie, “Learning modality-specific representations for visible-infrared person re-identification,” IEEE Transactions on Image Processing (TIP), vol. 29, pp. 579–590, 2020.
  263. 263.D. Li, X. Wei, X. Hong, and Y. Gong, “Infrared-visible cross-modal person re-identification with an x modality.” in AAAI, 2020, pp. 4610–4617.
  264. 264.Y. Lu, Y. Wu, B. Liu, T. Zhang, B. Li, Q. Chu, and N. Yu, “Cross-modality person re-identification with shared-specific feature transfer,” in CVPR, 2020, pp. 13 379–13 389.

Citation

MLA
Ye, M., et al. “Deep Learning for Person Re-identification: A Survey and Outlook”. arXiv, 2020, http://arxiv.org/abs/2001.04193v2.
APA
Ye, M., Shen, J., Lin, G., Xiang, T., Shao, L., & Hoi, S. C. H. (2020). Deep Learning for Person Re-identification: A Survey and Outlook. arXiv. http://arxiv.org/abs/2001.04193v2
Chicago
Ye, M., J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi. 2020. “Deep Learning for Person Re-identification: A Survey and Outlook”. arXiv. http://arxiv.org/abs/2001.04193v2.
Harvard
Ye, M. et al. (2020) “Deep Learning for Person Re-identification: A Survey and Outlook”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2001.04193v2.
Vancouver
1. Ye M, Shen J, Lin G, Xiang T, Shao L, Hoi SCH (2020) Deep Learning for Person Re-identification: A Survey and Outlook. arXiv

BibTeX

@article{ye2020deep,
  title = {Deep Learning for Person Re-identification: A Survey and Outlook},
  author = {Ye, Mang and Shen, Jianbing and Lin, Gaojie and Xiang, Tao and Shao, Ling and Hoi, Steven C. H.},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2001.04193v2},
  eprint = {2001.04193}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF