Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification

Shengcai LiaoLing Shao

article2022CVPR113 citationsFirst Prize of Best Paper Award

Proposes an efficient graph-based mini-batch sampling method that builds nearest neighbor class graphs to mine informative hard examples before batch construction, drastically cutting large-scale training time while boosting generalizable person re-identification accuracy.

Listen

Deploying automated person re-identification across different surveillance environments requires vision models that generalize reliably to unseen locations and camera setups. While training on large and diverse datasets significantly improves model generalization, current deep metric learning pipelines face severe computational bottlenecks. Standard approaches either maintain expensive class memory structures or rely on random mini-batch sampling strategies that fail to provide challenging, informative training examples.

The main objective of the article is to demonstrate an efficient mini-batch sampling technique, termed Graph Sampling, that accelerates large-scale pairwise metric learning while substantially improving cross-dataset generalization accuracy.

The authors conducted an empirical evaluation using four benchmark datasets, encompassing real-world surveillance data from CUHK03, Market-1501, and MSMT17, as well as synthetic imagery from RandPerson containing 8,000 distinct identities. Instead of using random selection, the Graph Sampling approach constructs a nearest-neighbor similarity graph across all identity classes at the start of each training epoch using only one sample per identity. Training batches are then formed by grouping each identity with its most visually similar neighboring classes. This sampling was paired with a modified query-adaptive convolution architecture, a hard triplet loss, and gradient clipping to stabilize training.

The experimental findings show substantial improvements in both operational efficiency and identification accuracy. First, the method reduced the training time on the 8,000-identity RandPerson dataset from 25.4 hours down to 2.0 hours, while graph construction itself introduced minimal overhead of only tens to hundreds of seconds per epoch. Second, when trained on RandPerson and tested on MSMT17, the proposed framework improved top-rank matching accuracy by 25.1 percentage points over the existing baseline. Third, under direct cross-dataset evaluations between Market-1501 and MSMT17, the framework outperformed prior state-of-the-art benchmarks by 20.6 percentage points in top-rank accuracy. Finally, the Graph Sampling strategy consistently outperformed both standard random sampling and subspace clustering methods across all evaluated benchmarks.

These results demonstrate that shifting hard-example mining directly into the data sampling stage yields highly discriminative, compact representations without the prohibitive computational costs of full-memory matching. For practitioners, this dramatically lowers the compute expenses and turnaround times required to train high-performing computer vision models on massive or synthetic datasets, enabling broader real-world deployment across previously unseen camera networks.

Organizations developing large-scale visual search and surveillance systems should adopt relationship-aware graph sampling over traditional random mini-batch samplers. When applying this technique, engineering teams should incorporate gradient norm clipping and limit within-class sample counts per batch to prevent optimization instabilities and mitigate overfitting on smaller training sets. Further exploration is recommended to evaluate whether the graph sampling framework provides comparable efficiency and accuracy gains in related computer vision domains, such as large-scale face recognition and general product image retrieval.

Readers should note that the performance advantages were evaluated specifically on person re-identification architectures and supervised pairwise metric learning. While confidence in the benchmarked scenarios is high due to consistent gains across both real and synthetic datasets, performance remains sensitive to hyperparameter choices such as triplet loss margins and gradient clipping thresholds, particularly when training on smaller or less diverse source datasets.

Cover for Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification

Abstract

Recent studies show that, both explicit deep feature matching as well as large-scale and diverse training data can significantly improve the generalization of person re-identification. However, the efficiency of learning deep matchers on large-scale data has not yet been adequately studied. Though learning with classification parameters or class memory is a popular way, it incurs large memory and computational costs. In contrast, pairwise deep metric learning within mini batches would be a better choice. However, the most popular random sampling method, the well-known PK sampler, is not informative and efficient for deep metric learning. Though online hard example mining has improved the learning efficiency to some extent, the mining in mini batches after random sampling is still limited. This inspires us to explore the use of hard example mining earlier, in the data sampling stage. To do so, in this paper, we propose an efficient mini-batch sampling method, called graph sampling (GS), for large-scale deep metric learning. The basic idea is to build a nearest neighbor relationship graph for all classes at the beginning of each epoch. Then, each mini batch is composed of a randomly selected class and its nearest neighboring classes so as to provide informative and challenging examples for learning. Together with an adapted competitive baseline, we improve the state of the art in generalizable person re-identification significantly, by 25.1% in Rank-1 on MSMT17 when trained on RandPerson. Besides, the proposed method also outperforms the competitive baseline, by 6.8% in Rank-1 on CUHK03-NP when trained on MSMT17. Meanwhile, the training time is significantly reduced, from 25.4 hours to 2 hours when trained on RandPerson with 8,000 identities. Code is available at https://github.com/ShengcaiLiao/QAConv.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Deep Metric Learning
  • 4. Graph Sampling
  • 4.1. Motivation
  • 4.2. GS Sampler
  • 4.3. Loss Function
  • 4.4. Gradient Clipping
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. Datasets
  • 5.3. Comparison to the State of the Art
  • 5.4. Ablation Study
  • 5.4.1 Comparison of QAConv variants
  • 5.4.2 Comparison of different sampling methods
  • 5.4.3 Parameter analysis
  • 5.4.4 Effect of gradient clipping
  • 5.4.5 Visualization of GS
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Graph Sampling Mini-Batch Construction

    model/method

    Graph Sampling (GS) is a class-level nearest-neighbor mini-batch sampling method for deep metric learning designed to provide informative hard negative samples during the data sampling stage rather than relying solely on post-sampling online mining.

    In deep metric learning, conventional random class sampling (such as the PKPK sampler, which randomly selects PP classes and KK images per class) produces mini-batches where instances are uniformly scattered across the dataset. Consequently, online hard example mining (OHEM) within such randomly assembled batches encounters few truly challenging negative pairs, reducing gradient informativeness.

    Graph Sampling operates as follows:

    1. Epoch Initialization: At the start of each epoch, exactly one image per class is randomly selected from all CC training classes to form an anchor set. Deep feature embeddings X∈RC×dX \in \mathbb{R}^{C \times d} are extracted using the latest network parameters fθf_\theta.

    2. Graph Construction: Pairwise class distances or similarities are evaluated among all CC extracted representations using the underlying similarity metric (e.g., Query-Adaptive Convolution / QAConv), yielding a distance matrix dist∈RC×C\text{dist} \in \mathbb{R}^{C \times C}. For each class c∈{1,2,…,C}c \in \{1, 2, \dots, C\}, its top P−1P - 1 nearest neighboring classes are retrieved, denoted as N(c)={x1,x2,…,xP−1}\mathcal{N}(c) = \{x_1, x_2, \dots, x_{P-1}\}. A directed nearest-neighbor graph G=(V,E)G = (V, E) is defined over vertices V={1,2,…,C}V = \{1, 2, \dots, C\} and directed edges E={(c1,c2)∣c2∈N(c1)}E = \{(c_1, c_2) \mid c_2 \in \mathcal{N}(c_1)\}.

    3. Batch Generation: Each training epoch comprises exactly CC mini-batch iterations. For each class cc chosen as an anchor, its neighborhood set A={c}∪{x∣(c,x)∈E}A = \{c\} \cup \{x \mid (c, x) \in E\} with ∣A∣=P|A| = P is retrieved. For each class in AA, KK instances are randomly drawn from the dataset, producing a mini-batch of size B=P×KB = P \times K.

    Because the batch size and composition focus on mutually similar classes, instances within each batch reside near class decision boundaries, forcing the metric learner to distinguish fine-grained visual differences.

  2. Knowl 2 — Graph Sampling Algorithm

    algorithm

    The Graph Sampling (GS) procedure constructs class nearest-neighbor graphs and produces mini-batches for deep metric learning across one training epoch.

    Input: Training dataset D\mathcal{D} partitioned into CC identity classes {D1,D2,…,DC}\{\mathcal{D}_1, \mathcal{D}_2, \dots, \mathcal{D}_C\}
    Input: Feature extractor fθf_\theta with parameters θ\theta
    Input: Pairwise similarity metric s(⋅,⋅)s(\cdot, \cdot)
    Input: Number of classes per batch PP, number of images per class KK
    Output: Ordered sequence of CC mini-batches {B1,B2,…,BC}\{\mathcal{B}_1, \mathcal{B}_2, \dots, \mathcal{B}_C\}
    # Step 1: Representative Feature Extraction
    for each class c∈{1,2,…,C}c \in \{1, 2, \dots, C\} do
        sample one image xc∼Dcx_c \sim \mathcal{D}_c
        compute embedding zc=fθ(xc)z_c = f_\theta(x_c)
    end for
    # Step 2: Distance Matrix Computation
    initialize distance matrix dist∈RC×C\text{dist} \in \mathbb{R}^{C \times C}
    for i=1i = 1 to CC do
        for j=1j = 1 to CC do
            dist[i,j]=−s(zi,zj)\text{dist}[i, j] = -s(z_i, z_j)
        end for
    end for
    # Step 3: Graph Construction
    for each class c∈{1,2,…,C}c \in \{1, 2, \dots, C\} do
        N(c)←\mathcal{N}(c) \leftarrow indices of the P−1P - 1 smallest non-self entries in dist[c,:]\text{dist}[c, :]
        Ac←{c}∪N(c)\mathcal{A}_c \leftarrow \{c\} \cup \mathcal{N}(c)
    end for
    # Step 4: Batch Assembly
    for each anchor class c∈{1,2,…,C}c \in \{1, 2, \dots, C\} do
        initialize batch Bc←∅\mathcal{B}_c \leftarrow \emptyset
        for each class id u∈Acu \in \mathcal{A}_c do
            sample KK images {xu,1,xu,2,…,xu,K}∼Du\{x_{u, 1}, x_{u, 2}, \dots, x_{u, K}\} \sim \mathcal{D}_u
            Bc←Bc∪{xu,1,…,xu,K}\mathcal{B}_c \leftarrow \mathcal{B}_c \cup \{x_{u, 1}, \dots, x_{u, K}\}
        end for
    end for
    return {B1,B2,…,BC}\{\mathcal{B}_1, \mathcal{B}_2, \dots, \mathcal{B}_C\}

    The total number of mini-batches per epoch is fixed at CC. Extracting 1 image per class and constructing the graph at epoch initialization takes 4 seconds on Market-1501 (C=751C=751), 9 seconds on MSMT17 training set (C=1,041C=1,041), 40 seconds on full MSMT17 (C=4,101C=4,101), and 138 seconds on RandPerson (C=8,000C=8,000) on an NVIDIA V100 GPU.

  3. Knowl 3 — Batch Online Hard Example Mining Triplet Loss for Metric Learning

    equation

    In Graph Sampling deep metric learning, feature matching and loss evaluation are performed exclusively within mini-batches using the batch Online Hard Example Mining (OHEM) triplet loss without requiring parametric classification layers or global class memory banks:

    ℓ(θ;X)=∑i=1P∑a=1K[m−min⁡p=1…Ks(fθ(xia),fθ(xip))+max⁡j=1…Pj≠imax⁡n=1…Ks(fθ(xia),fθ(xjn))]+\ell(\theta; X) = \sum_{i=1}^P \sum_{a=1}^K \left[ m - \min_{p=1\dots K} s(f_\theta(x_i^a), f_\theta(x_i^p)) + \max_{\substack{j=1\dots P \\ j \ne i}} \max_{n=1\dots K} s(f_\theta(x_i^a), f_\theta(x_j^n)) \right]_+

    where:

    • X={xia∣i∈{1,…,P},a∈{1,…,K}}X = \{x_i^a \mid i \in \{1, \dots, P\}, a \in \{1, \dots, K\}\} is the mini-batch containing PP classes with KK image instances each (total batch size B=P×KB = P \times K).
    • θ\theta represents the trainable network parameters.
    • fθ(⋅)f_\theta(\cdot) is the deep feature extraction network.
    • s(fθ(u),fθ(v))∈(−∞,+∞)s(f_\theta(u), f_\theta(v)) \in (-\infty, +\infty) denotes the pairwise similarity score between images uu and vv (e.g., computed via Query-Adaptive Convolution feature map cross-matching).
    • m∈R+m \in \mathbb{R}^+ is a fixed margin scalar (default m=16m = 16).
    • [z]+=max⁡(0,z)[z]_+ = \max(0, z) is the standard hinge rectifier.

    For each anchor xiax_i^a, the loss identifies the hardest positive instance xipx_i^p (minimum similarity within class ii) and the hardest negative instance xjnx_j^n (maximum similarity across all other P−1P-1 classes in the batch). Because Graph Sampling constructs mini-batches from top nearest-neighbor classes, the negative terms are informative, allowing the triplet loss to train the network without an auxiliary identity classification cross-entropy loss.

  4. Knowl 4 — Optimization Stabilization via Gradient Clipping and Instance Limitation

    model/method

    Combining Graph Sampling (which samples the most similar classes) with the batch OHEM triplet loss (which selects the hardest negative samples within the batch) produces extreme learning constraints that can cause gradient instability and poor convergence. Two techniques are employed to stabilize optimization:

    1. Instance Constraint (K=2K = 2): Restricting the number of sampled instances per class to K=2K = 2 (with P=32P = 32 for a batch size B=64B = 64) reduces the intra-class combinatorial difficulty while ensuring each anchor has a valid positive pair and multiple hard negative candidates from P−1=31P - 1 = 31 neighboring classes.

    2. Global Gradient Norm Clipping: During backpropagation, the global parameter gradient vector g=∇θℓg = \nabla_\theta \ell is scaled according to:

    g←min⁡(1,T∥g∥2)⋅gg \leftarrow \min\left(1, \frac{T}{\|g\|_2}\right) \cdot g

    where ∥g∥2\|g\|_2 is the Euclidean norm of the concatenated parameter gradients and TT is a threshold hyperparameter (default T=8T = 8).

    In addition to preventing training divergence caused by stiff boundary constraints, gradient clipping filters out noisy gradients from source domains, serving as a regularizer that prevents overfitting to source-specific artifacts and improves cross-dataset generalization on target domains.

  5. Knowl 5 — Generalizable Person Re-Identification Benchmark Comparison

    data/table

    Cross-dataset evaluation benchmarks generalizability by training on one dataset and evaluating on unseen target datasets using single-query Rank-1 accuracy (%) and mean Average Precision (mAP (%)). When MSMT17 is used with all images regardless of splits, it is denoted as MSMT17 (all).

    Method Venue Training CUHK03-NP Market-1501 MSMT17
    Rank-1 mAP Rank-1 mAP Rank-1 mAP
    M3L CVPR'21 Multi 33.1 32.1 75.9 50.2 36.9 14.7
    MGN ACMMM'18 Market-1501 8.5 7.4 95.7 86.9 - -
    MuDeep TPAMI'20 Market-1501 10.3 9.1 95.3 84.7 - -
    QAConv ECCV'20 Market-1501 9.9 8.6 - - 22.6 7.0
    OSNet-AIN TPAMI'21 Market-1501 - - 94.2 84.4 23.5 8.2
    CBN ECCV'20 Market-1501 - - 91.3 77.3 25.3 9.5
    QAConv-GS Ours Market-1501 19.1 18.1 91.6 75.5 45.9 17.2
    PCB ECCV'18 MSMT17 - - 52.7 26.7 - -
    MGN ACMMM'18 MSMT17 - - 48.7 25.1 - -
    ADIN WACV'20 MSMT17 - - 59.1 30.3 - -
    SNR CVPR'20 MSMT17 - - 70.1 41.4 - -
    CBN ECCV'20 MSMT17 - - 73.7 45.0 72.8 42.9
    QAConv-GS Ours MSMT17 20.9 20.6 79.1 49.5 79.2 50.9
    OSNet-IBN CVPR'19 MSMT17 (all) - - 66.5 37.2 - -
    OSNet-AIN TPAMI'21 MSMT17 (all) - - 70.1 43.3 - -
    QAConv ECCV'20 MSMT17 (all) 25.3 22.6 72.6 43.1 - -
    QAConv-GS Ours MSMT17 (all) 27.6 28.0 82.4 56.9 - -
    RP Baseline ACMMM'20 RandPerson 13.4 10.8 55.6 28.8 20.1 6.3
    CBN ECCV'20 RandPerson - - 64.7 39.3 20.0 6.8
    QAConv-GS Ours RandPerson 18.4 16.1 76.7 46.7 45.1 15.5

    QAConv-GS outperforms previous methods across cross-domain transfer tasks. On Market-1501 →\rightarrow MSMT17, Rank-1 increases from 25.3% (CBN) to 45.9% (+20.6%). On RandPerson →\rightarrow MSMT17, Rank-1 increases from 20.0% (CBN) to 45.1% (+25.1%). On MSMT17 (all) →\rightarrow Market-1501, Rank-1 improves from 72.6% (QAConv) to 82.4% (+9.8%). Although M3L trains on three datasets simultaneously, QAConv-GS trained on Market-1501 alone outperforms M3L on MSMT17 by 9.0% in Rank-1 and 2.5% in mAP.

  6. Knowl 6 — Comparison of QAConv Variants and Training Time Efficiency

    data/table

    The table compares three variants of QAConv across training wall-clock time (in hours on a single NVIDIA V100 GPU) and generalization performance (Rank-1 (%) and mAP (%)): (1) Ori: the original QAConv relying on a class memory bank holding full feature maps for all identities; (2) Base: an adapted baseline using pairwise mini-batch matching with PK sampling; and (3) GS: the proposed pairwise matching with Graph Sampling.

    Training Data Hours CUHK03 Market MSMT17
    R1 mAP R1 mAP R1 mAP
    Ori Market 1.33 9.9 8.6 - - 22.6 7.0
    Base Market 0.47 14.6 14.6 88.7 71.4 42.6 15.8
    GS Market 0.25 19.1 18.1 91.6 75.5 45.9 17.2
    Base MSMT 1.33 14.1 15.7 73.7 44.7 72.5 43.4
    GS MSMT 0.73 20.9 20.6 79.1 49.5 79.2 50.9
    Ori MS-all 26.9 25.3 22.6 72.6 43.1 - -
    Base MS-all 15.0 23.4 23.1 80.1 53.2 - -
    GS MS-all 3.42 27.6 28.0 82.4 56.9 - -
    Base RP 25.4 15.2 14.6 75.9 46.0 44.4 15.5
    GS RP 2.00 18.4 16.1 76.7 46.7 45.1 15.5

    When training on large datasets such as MSMT17 (all) (4,101 identities) and RandPerson (8,000 identities), the original class memory approach scales poorly because every iteration computes cross-convolutional maps against all stored class signatures. Removing class memory and using Graph Sampling reduces training time on RandPerson from 25.4 hours to 2.0 hours (a 12.7×12.7\times speedup) while improving cross-dataset performance across CUHK03, Market-1501, and MSMT17.

  7. Knowl 7 — Comparison of Mini-Batch Sampling Strategies

    data/table

    The table compares three mini-batch sampling methods under identical backbone architectures, QAConv matcher, and hard triplet loss: (1) PK Sampler: random selection of PP classes and KK instances; (2) Cluster: subspace spectral clustering on averaged class representations (M=10M = 10 subspaces) followed by intra-cluster random sampling; and (3) GS: Graph Sampling. Training time is reported in seconds per epoch.

    Method Data Time (s/ep) CUHK03 Market MSMT17
    R1 mAP R1 mAP R1 mAP
    PK Market 99 17.9 17.0 - - 43.3 15.6
    Cluster Market 117 18.4 17.3 - - 44.0 15.8
    GS Market 100 19.1 18.1 - - 45.9 17.2
    PK MSMT 141 18.6 18.8 75.7 46.1 - -
    Cluster MSMT 196 18.4 19.2 77.2 47.6 - -
    GS MSMT 145 20.9 20.6 79.1 49.5 - -
    PK MS-all 669 24.5 24.6 78.7 52.1 - -
    Cluster MS-all 881 26.3 26.3 80.4 54.2 - -
    GS MS-all 685 27.6 28.0 82.4 56.9 - -
    PK RP 1,150 16.9 14.7 73.2 43.5 40.3 13.1
    Cluster RP 1,922 17.3 15.0 73.3 43.3 40.4 13.4
    GS RP 1,397 18.4 16.1 76.7 46.7 45.1 15.5

    PK sampling yields the lowest generalization because random batches lack challenging negatives. Cluster sampling improves informativeness but requires forward passes over the entire dataset to compute class averages and perform spectral clustering, which increases per-epoch runtime (e.g., 1,922 s vs. 1,397 s on RandPerson). GS achieves the highest accuracy across all configurations (up to +4.7% Rank-1 and +3.4% mAP over Cluster on MSMT17 →\rightarrow CUHK03) because it guarantees that every anchor class is paired with its exact top-(P−1)(P-1) nearest neighbors.

  8. Knowl 8 — QAConv-GS Implementation and Training Hyperparameters

    experimental setup

    The QAConv-GS model is implemented in PyTorch with the following architectural and optimization configurations:

    • Network Backbone: ResNet-50 augmented with Instance-Batch Normalization (IBN-b) layers. Feature maps are extracted from layer3, followed by a 1×11 \times 1 neck convolution outputting 128 feature channels.
    • Input Preprocessing: Images are resized to 384×128384 \times 128 pixels. Augmentations include random cropping, horizontal flipping, random occlusion, and color jittering.
    • Optimization: Stochastic Gradient Descent (SGD) with Automatic Mixed Precision (AMP). The initial learning rate is set to 0.00050.0005 for the backbone and 0.0050.005 for newly added layers (neck conv and matcher).
    • Learning Rate Decay & Stopping: Maximum epochs set to 60. When the training loss drops to a factor of 0.70.7 relative to its initial value, learning rates are decayed by 0.1×0.1\times. Training triggers early stopping after an additional number of epochs equal to half of the elapsed epochs.
    • Sampling and Loss Parameters: Mini-batch size B=64B = 64 with K=2K = 2 instances per class (P=32P = 32 classes per batch). Margin parameter for the batch OHEM triplet loss is m=16m = 16. Gradient norm clipping threshold is T=8T = 8.
  9. Knowl 9 — Sensitivity to Batch Size, Margin Parameter, and Gradient Clipping Threshold

    empirical result

    Empirical sensitivity analysis on MSMT17, measuring mean accuracy (mAcc (%), the average of Rank-1 and mAP across all target test datasets over four runs), reveals distinct behavior across key hyperparameters:

    1. Batch Size (BB): Evaluated across B∈{4,8,16,32,64,128,256}B \in \{4, 8, 16, 32, 64, 128, 256\} with K=2K = 2. Accuracy increases steadily from ~26% at B=4B = 4 to ~42% at B=64B = 64, beyond which performance saturates (B=128B = 128 and B=256B = 256 yield similar mAcc around 42%-43%).

    2. Triplet Margin (mm): Evaluated across m∈{1,2,4,8,16,32,64}m \in \{1, 2, 4, 8, 16, 32, 64\}. The similarity score s(u,v)s(u, v) in QAConv spans (−∞,+∞)(-\infty, +\infty). Accuracy improves from ~37% at m=1m = 1 to a peak of ~42% at m=16m = 16. When m≥32m \ge 32, performance degrades rapidly (falling to ~22% at m=64m = 64) because the margin constraint becomes excessively hard, impeding optimization convergence.

    3. Gradient Clipping Threshold (TT): Evaluated across T∈{1,2,4,8,16,32,64,128,256,512,1024,∞}T \in \{1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, \infty\} on Market-1501, MSMT17, and RandPerson. On the comprehensive MSMT17 dataset, the model is resilient across T≥8T \ge 8 up to T=∞T = \infty (unclipped). However, on smaller (Market-1501) or synthetic (RandPerson) datasets, unclipped training suffers from overfitting; setting T=8T = 8 provides regularizing gradient bounds that maximize target-domain cross-dataset generalization.

Coverage note — Deliberately omitted qualitative visualization details (Fig. 4 / App. F) as they illustrate the same properties captured in the quantitative results, and omitted external UDA / SpCL adaptations mentioned only in the appendix.

References

  1. 1.Ejaz Ahmed, Michael Jones, and Tim K Marks. An improved deep learning architecture for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3908–3916, 2015. 1, 2
  2. 2.Yan Bai, Jile Jiao, Wang Ce, Jun Liu, Yihang Lou, Xuetao Feng, and Ling-Yu Duan. Person30k: A dual-meta generalization network for person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2123–2132, 2021. 1, 3
  3. 3.Seokeon Choi, Taekyung Kim, Minki Jeong, Hyoungseob Park, and Changick Kim. Meta batch-instance normalization for generalizable person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3425–3435, 2021. 3
  4. 4.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4690–4699, 2019. 2
  5. 5.Weijian Deng, Liang Zheng, Qixiang Ye, Guoliang Kang, Yi Yang, and Jianbin Jiao. Image-Image Domain Adaptation with Preserved Self-Similarity and Domain-Dissimilarity for Person Re-identification. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2018. 1, 2
  6. 6.Yixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao, et al. Self-paced contrastive learning with hybrid memory for domain adaptive object re-id. Advances in Neural Information Processing Systems, 33:11309–11321, 2020. 7
  7. 7.Ben Harwood, Vijay Kumar BG, Gustavo Carneiro, Ian Reid, and Tom Drummond. Smart mining for deep metric learning. In Proceedings of the IEEE International Conference on Computer Vision, pages 2821–2829, 2017. 3
  8. 8.K. He, X. zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 5
  9. 9.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification, 2017. 1, 2, 3, 4
  10. 10.Yang Hu, Dong Yi, Shengcai Liao, Zhen Lei, and Stan Z. Li. Cross dataset person Re-identification. In ACCV Workshop on Human Identification for Surveillance (HIS), pages 650–664, 2014. 1, 3, 5
  11. 11.Jieru Jia, Qiuqi Ruan, and Timothy M Hospedales. Frustratingly easy person re-identification: Generalizing person re-id in practice. In British Machine Vision Conference, 2019. 1, 3, 5
  12. 12.Xin Jin, Cuiling Lan, Wenjun Zeng, Zhibo Chen, and Li Zhang. Style Normalization and Restitution for Generalizable Person Re-identification. In CVPR, feb 2020. 1, 3, 5, 6
  13. 13.M Kostinger, Martin Hirzer, Paul Wohlhart, Peter M Roth, and Horst Bischof. Large scale metric learning from equivalence constraints. In IEEE Conference on Computer Vision and Pattern Recognition, 2012. 2
  14. 14.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. DeepReID: Deep filter pairing neural network for person re-identification. In IEEE Conference on Computer Vision and Pattern Recognition, 2014. 1, 2, 5
  15. 15.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deepreid: Deep filter pairing neural network for person re-identification. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2014. 3
  16. 16.Shengcai Liao, Yang Hu, Xiangyu Zhu, and Stan Z. Li. Person re-identification by local maximal occurrence representation and metric learning. In IEEE Conference on Computer Vision and Pattern Recognition, 2015. 2
  17. 17.Shengcai Liao and Ling Shao. Interpretable and Generalizable Person Re-Identification with Query-Adaptive Convolution and Temporal Lifting. In European Conference on Computer Vision (ECCV), 2020. 1, 2, 3, 5, 6
  18. 18.Shengcai Liao and Ling Shao. TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification. In Advances in Neural Information Processing Systems, 2021. 3, 7
  19. 19.W. Liu, Y. Wen, Z. Yu, M. Li, R. Bhiksha, and L. Song. Sphereface: Deep hypersphere embedding for face recognition. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, volume 1, page 1, 2017. 2
  20. 20.Xingang Pan, Ping Luo, Jianping Shi, and Xiaoou Tang. Two at once: Enhancing learning and generalization capacities via ibn-net. In Proceedings of the European Conference on Computer Vision (ECCV), pages 464–479, 2018. 5
  21. 21.Xuelin Qian, Yanwei Fu, Tao Xiang, Yu-Gang Jiang, and Xiangyang Xue. Leader-based Multi-Scale Attention Deep Architecture for Person Re-identification. TPAMI, 2020. 1, 5, 6
  22. 22.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 815–823, 2015. 2
  23. 23.Florian Schroff, Dmitry Kalenichenko, and James Philbin. FaceNet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2015. 3
  24. 24.Yantao Shen, Tong Xiao, Hongsheng Li, Shuai Yi, and Xiaogang Wang. End-to-end deep kronecker-product matching for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6886–6895, 2018. 1, 2
  25. 25.Evgeny Smirnov, Aleksandr Melnikov, Sergey Novoselov, Eugene Luckyanets, and Galina Lavrentyeva. Doppelganger mining for face representation learning. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 1916–1923, 2017. 3
  26. 26.Jifei Song, Yongxin Yang, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Generalizable person re-identification by domain-invariant mapping network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 719–728, 2019. 1, 3
  27. 27.Yumin Suh, Bohyung Han, Wonsik Kim, and Kyoung Mu Lee. Stochastic class-based hard example mining for deep metric learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7251–7259, 2019. 3
  28. 28.Yumin Suh, Jingdong Wang, Siyu Tang, Tao Mei, and Kyoung Mu Lee. Part-aligned bilinear representations for person re-identification. Proceedings of the European Conference on Computer Vision (ECCV), pages 402–419, 2018. 1, 2
  29. 29.Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). In Proceedings of the European Conference on Computer Vision (ECCV), 2018. 6
  30. 30.Chong Wang, Xue Zhang, and Xipeng Lan. How to train triplet networks with 100k identities? In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 1907–1915, 2017. 3, 7
  31. 31.Guangrun Wang, Guangcong Wang, Xujie Zhang, Jianhuang Lai, Zhengtao Yu, and Liang Lin. Weakly supervised person re-id: Differentiable graphical learning and a new benchmark. IEEE Transactions on Neural Networks and Learning Systems, 32(5):2142–2156, 2020. 1
  32. 32.Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, and Xi Zhou. Learning discriminative features with multiple granularities for person re-identification. In 2018 ACM Multimedia Conference on Multimedia Conference, pages 274–282. ACM, 2018. 6
  33. 33.Yanan Wang, Shengcai Liao, and Ling Shao. Surpassing Real-World Source Training Data: Random 3D Characters for Generalizable Person Re-Identification. In 28th ACM International Conference on Multimedia (ACMMM), 2020. 1, 3, 5, 6
  34. 34.Longhui Wei, Shiliang Zhang, Wen Gao, and Qi Tian. Person transfer gan to bridge domain gap for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 79–88, 2018. 5
  35. 35.Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. A discriminative feature learning approach for deep face recognition. In European conference on computer vision, pages 499–515. Springer, 2016. 2
  36. 36.Chao-Yuan Wu, R Manmatha, Alexander J Smola, and Philipp Krahenbuhl. Sampling matters in deep embedding learning. In Proceedings of the IEEE International Conference on Computer Vision, pages 2840–2848, 2017. 3
  37. 37.Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang, Ling Shao, and Steven C. H. Hoi. Deep Learning for Person Re-identification: A Survey and Outlook. arXiv preprint arXiv:2001.04193, 2020. 2, 3
  38. 38.Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z. Li. Deep metric learning for person re-identification. In International Conference on Pattern Recognition, pages 34–39, Dec. 2014. 1, 2, 3, 5
  39. 39.Ye Yuan, Wuyang Chen, Tianlong Chen, Yang Yang, Zhou Ren, Zhangyang Wang, and Gang Hua. Calibrated Domain-Invariant Learning for Highly Generalizable Large Scale Re-Identification. WACV, pages 3578–3587, nov 2020. 3, 5, 6
  40. 40.Tianyu Zhang, Lingxi Xie, Longhui Wei, Zijie Zhuang, Yongfei Zhang, Bo Li, and Qi Tian. Unrealperson: An adaptive pipeline towards costless person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11506–11515, 2021. 6
  41. 41.Yuyang Zhao, Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, and Nicu Sebe. Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6277–6286, 2021. 3, 5, 6
  42. 42.L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian. Scalable person re-identification: A benchmark. In Proceedings of IEEE International Conference on Computer Vision, 2015. 5
  43. 43.Liang Zheng, Hengheng Zhang, Shaoyan Sun, Manmohan Chandraker, Yi Yang, and Qi Tian. Person re-identification in the Wild. In Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, 2017. 1, 2, 3
  44. 44.Wei-Shi Zheng, Shaogang Gong, and Tao Xiang. Person re-identification by probabilistic relative distance comparison. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 649–656, 2011. 2
  45. 45.Zhun Zhong, Liang Zheng, Donglin Cao, and Shaozi Li. Re-ranking person re-identification with k-reciprocal encoding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1318–1327, 2017. 5
  46. 46.Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, and Tao Xiang. Omni-scale feature learning for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, 2019. 1, 3, 5, 6, 7
  47. 47.Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, and Tao Xiang. Learning generalisable omni-scale representations for person re-identification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021. 5, 6
  48. 48.Zijie Zhuang, Longhui Wei, Lingxi Xie, Tianyu Zhang, Hengheng Zhang, Haozhe Wu, Haizhou Ai, and Qi Tian. Rethinking the Distribution Gap of Person Re-identification with Camera-based Batch Normalization. In ECCV, pages 140–157, jan 2020. 1, 3, 5, 6

Citation

MLA
Liao, S., and L. Shao. “Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification”. arXiv, 2021, http://arxiv.org/abs/2104.01546v4.
APA
Liao, S., & Shao, L. (2021). Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification. arXiv. http://arxiv.org/abs/2104.01546v4
Chicago
Liao, S., and L. Shao. 2021. “Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification”. arXiv. http://arxiv.org/abs/2104.01546v4.
Harvard
Liao, S. and Shao, L. (2021) “Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.01546v4.
Vancouver
1. Liao S, Shao L (2021) Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification. arXiv

BibTeX

@article{liao2021graph,
  title = {Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification},
  author = {Liao, Shengcai and Shao, Ling},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.01546v4},
  eprint = {2104.01546}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE