Learning Discriminative Features with Multiple Granularities for Person Re-Identification

Guanshuo WangYufeng YuanXiong ChenJiwei LiXi Zhou

article2018ACM Multimedia1,519 citations

Introduces the Multiple Granularity Network (MGN), a multi-branch deep architecture that combines global features with multi-level stripe partitions to surpass semantic part-based methods across standard person re-identification benchmarks.

Listen

Identifying and tracking individuals across non-overlapping surveillance camera networks—a task known as person re-identification—is essential for modern public safety and security operations. However, automated systems often struggle due to variations in pedestrian poses, occlusions, background clutter, and image resolution. Previous methods attempted to solve this by locating specific body parts or using complex attention mechanisms, which added significant computational complexity, lacked robustness to dramatic appearance changes, and required multi-stage training pipelines.

The article demonstrates an end-to-end deep learning framework, termed the Multiple Granularity Network, designed to overcome these challenges. The system extracts and combines full-body global visual features with multi-level local part details using simple, uniform horizontal image stripes rather than complex body-part detectors.

To evaluate the system, the authors conducted extensive experiments on three benchmark surveillance datasets: Market-1501, DukeMTMC-reID, and CUHK03. The architecture modifies a standard visual backbone into three parallel processing branches: one capturing full-body global features, a second dividing feature maps into two horizontal stripes, and a third dividing them into three stripes. The model combines classification objectives for individual identity prediction with ranking objectives (batch-hard metric loss) using a specialized coarse-to-fine supervisory design during training.

The findings show that the Multiple Granularity Network significantly outperforms previous approaches across all benchmarks. On the Market-1501 dataset, the model achieved a 95.7% Rank-1 retrieval accuracy and an 86.9% mean average precision (improving to 96.6% and 94.2% when applying an automated re-ranking step), exceeding prior top systems by 1.9% and 5.3% respectively. On the challenging DukeMTMC-reID dataset, it reached 88.7% Rank-1 accuracy and 78.4% mean average precision, beating existing benchmarks by 3.5% and 5.6%. Ablation analyses confirmed that combining global, two-part, and three-part representations in a single coordinated network yielded substantially higher precision than training separate standalone networks or simply increasing model size.

These results indicate that automated surveillance systems can achieve high identification reliability without relying on complex, failure-prone human pose estimators or external semantic detectors. The multi-granularity representation effectively captures subtle local visual cues—such as logos, straps, or specific clothing patterns—even under partial occlusion, viewpoint shifts, or low image resolution. This architectural simplicity reduces engineering overhead while improving operational accuracy.

Organizations deploying automated visual tracking systems should consider adopting multi-granularity feature extraction pipelines to maximize matching precision. As an actionable next step, decision-makers should also invest in high-quality upstream person-detection algorithms; the experimental results showed a measurable accuracy decline on automatically detected bounding boxes compared to perfectly cropped images, highlighting bounding-box quality as a primary real-world operational bottleneck.

Cover for Learning Discriminative Features with Multiple Granularities for Person Re-Identification

Abstract

The combination of global and partial features has been an essential solution to improve discriminative performances in person re-identification (Re-ID) tasks. Previous part-based methods mainly focus on locating regions with specific pre-defined semantics to learn local representations, which increases learning difficulty but not efficient or robust to scenarios with large variances. In this paper, we propose an end-to-end feature learning strategy integrating discriminative information with various granularities. We carefully design the Multiple Granularity Network (MGN), a multi-branch deep network architecture consisting of one branch for global feature representations and two branches for local feature representations. Instead of learning on semantic regions, we uniformly partition the images into several stripes, and vary the number of parts in different local branches to obtain local feature representations with multiple granularities. Comprehensive experiments implemented on the mainstream evaluation datasets including Market-1501, DukeMTMC-reid and CUHK03 indicate that our method has robustly achieved state-of-the-art performances and outperformed any existing approaches by a large margin. For example, on Market-1501 dataset in single query mode, we achieve a state-of-the-art result of Rank-1/mAP=96.6%/94.2% after re-ranking.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Multiple Granularity Network
  • 3.1 Network Architecture
  • 3.2 Loss Functions
  • 3.3 Discussions
  • 4 Experiment
  • 4.1 Implementation
  • 4.2 Datasets and Protocols
  • 4.3 Comparison with State-of-the-Art Methods
  • 4.4 Effectiveness of Components
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Multiple Granularity Network Architecture for Person Re-Identification

    model/method

    The Multiple Granularity Network (MGN) is an end-to-end multi-branch deep neural network architecture designed for person re-identification. Using ResNet-50 as its backbone, the network processes an input image of size 384×128384 \times 128 up to the res_conv4_1 residual block, after which it splits into three independent branches that duplicate the remaining residual stages (res_conv4_2 through res_conv5_3) with separate, non-shared parameters:

    1. Global Branch: Employs standard down-sampling with a stride-2 convolution in res_conv5_1, yielding an output feature map of size 12×4×204812 \times 4 \times 2048. A Global Max-Pooling (GMP) operation produces a 2048-dimensional global feature vector zgGz_g^G. This vector is reduced to a 256-dimensional feature vector fgGf_g^G via a 1×11 \times 1 convolution layer followed by Batch Normalization and ReLU.

    2. Part-2 Branch: Omits down-sampling in res_conv5_1 (stride 1), producing an output feature map of size 24×8×204824 \times 8 \times 2048. This feature map is utilized in two ways:

      • Global representation: GMP across the entire feature map yields a 2048-dimensional vector zgP2z_g^{P2}, reduced by a 1×11 \times 1 convolution (with Batch Normalization and ReLU) to a 256-dimensional vector fgP2f_g^{P2}.
      • Local representations: The feature map is uniformly split horizontally into N=2N=2 equal stripes (12×812 \times 8 each). GMP is applied to each stripe independently to obtain two 2048-dimensional vectors zp0P2,zp1P2z_{p0}^{P2}, z_{p1}^{P2}, which are each reduced by separate 1×11 \times 1 convolution layers (with Batch Normalization and ReLU) to 256-dimensional local feature vectors fp0P2,fp1P2f_{p0}^{P2}, f_{p1}^{P2}.
    3. Part-3 Branch: Also omits down-sampling in res_conv5_1 (stride 1), producing a 24×8×204824 \times 8 \times 2048 feature map. Similar to the Part-2 branch, it produces:

      • Global representation: GMP over the full map yields zgP3z_g^{P3}, reduced via 1×11 \times 1 convolution to a 256-dimensional vector fgP3f_g^{P3}.
      • Local representations: The feature map is uniformly split horizontally into N=3N=3 equal stripes (8×88 \times 8 each). GMP on each stripe yields three 2048-dimensional vectors zp0P3,zp1P3,zp2P3z_{p0}^{P3}, z_{p1}^{P3}, z_{p2}^{P3}, which are reduced by separate 1×11 \times 1 convolution layers to 256-dimensional vectors fp0P3,fp1P3,fp2P3f_{p0}^{P3}, f_{p1}^{P3}, f_{p2}^{P3}.

    During inference, all 8 reduced 256-dimensional feature vectors are concatenated into a single 2048-dimensional feature embedding vector: f=[fgG,fgP2,fp0P2,fp1P2,fgP3,fp0P3,fp1P3,fp2P3]f = [f_g^G, f_g^{P2}, f_{p0}^{P2}, f_{p1}^{P2}, f_g^{P3}, f_{p0}^{P3}, f_{p1}^{P3}, f_{p2}^{P3}] which represents the pedestrian image by integrating global appearance and multi-granularity local stripe features.

  2. Knowl 2 — Classification-Before-Metric Supervisory Scheme in MGN

    model/method

    Multiple Granularity Network (MGN) employs an asymmetric coarse-to-fine multi-loss supervision strategy combining multi-class classification (softmax loss) and deep metric learning (batch-hard triplet loss):

    1. Softmax Loss Placement:

      • Applied to the three unreduced, 2048-dimensional globally-pooled features: zgGz_g^G (Global Branch), zgP2z_g^{P2} (Part-2 Branch), and zgP3z_g^{P3} (Part-3 Branch).
      • Applied to the five reduced, 256-dimensional local stripe features: fp0P2,fp1P2f_{p0}^{P2}, f_{p1}^{P2} from Part-2 Branch, and fp0P3,fp1P3,fp2P3f_{p0}^{P3}, f_{p1}^{P3}, f_{p2}^{P3} from Part-3 Branch.
      • Each target feature is fed to an independent linear classifier (fully connected layer without bias) corresponding to identity prediction over the training classes.
    2. Batch-Hard Triplet Loss Placement:

      • Applied exclusively to the three reduced, 256-dimensional global feature vectors: fgG,fgP2,fgP3f_g^G, f_g^{P2}, f_g^{P3}.
    3. Design Rationale:

      • Classification-before-metric: Unreduced 2048-dimensional features provide coarse, high-capacity representations for classification, while the reduced 256-dimensional features provide compact representations regularized by metric learning. This avoids convergence difficulties and manual loss-weight tuning associated with enforcing both losses on the exact same feature level.
      • Omission of Triplet Loss on Local Features: Triplet loss is deliberately not applied to local stripe features. Because bounding-box misalignment and pose variations cause stripe contents to shift across different samples of the same identity, local triplet constraints tend to destabilize and corrupt training.
  3. Knowl 3 — Formulation of Bias-Free Softmax and Batch-Hard Triplet Losses

    equation

    MGN is trained with a combination of bias-free softmax classification loss and batch-hard triplet loss:

    1. Bias-Free Softmax Loss: For a mini-batch of NN feature vectors where fif_i denotes the feature of the ii-th sample with ground-truth class label yi∈{1,…,C}y_i \in \{1, \dots, C\}, the classification loss is defined by removing the linear bias terms: Lsoftmax=−∑i=1Nlog⁡eWyiTfi∑k=1CeWkTfiL_{\text{softmax}} = -\sum_{i=1}^N \log \frac{e^{W_{y_i}^T f_i}}{\sum_{k=1}^C e^{W_k^T f_i}} where WkW_k denotes the classifier weight vector for class kk, CC is the total number of person identity classes in the training set, and NN is the total number of samples in the mini-batch.

    2. Batch-Hard Triplet Loss: For a mini-batch constructed with PP randomly selected pedestrian identities and KK randomly sampled images per identity (total batch size N=P×KN = P \times K), the batch-hard triplet loss selects the hardest positive and hardest negative pairs for each anchor: Ltriplet=∑i=1P∑a=1K[α+max⁡p=1,…,K∥fa(i)−fp(i)∥2−min⁡j=1,…,Pj≠imin⁡n=1,…,K∥fa(i)−fn(j)∥2]+L_{\text{triplet}} = \sum_{i=1}^P \sum_{a=1}^K \left[ \alpha + \max_{p=1,\dots,K} \|f_a^{(i)} - f_p^{(i)}\|_2 - \min_{\substack{j=1,\dots,P \\ j \neq i}} \min_{n=1,\dots,K} \|f_a^{(i)} - f_n^{(j)}\|_2 \right]_+ where fa(i)f_a^{(i)} is the anchor feature of the aa-th image of identity ii, fp(i)f_p^{(i)} is a positive feature belonging to the same identity ii, fn(j)f_n^{(j)} is a negative feature belonging to a different identity j≠ij \neq i, [x]+=max⁡(x,0)[x]_+ = \max(x, 0), and α>0\alpha > 0 is a fixed scalar margin parameter controlling the distance boundary between intra-class and inter-class representations.

  4. Knowl 4 — MGN Training Configuration and Evaluation Protocol

    experimental setup

    The standard implementation and training configuration for MGN is as follows:

    • Input Preprocessing: Pedestrian bounding box images are resized to 384×128384 \times 128 pixels. Random horizontal flipping is applied during training as the sole data augmentation technique.
    • Network Initialization: ResNet-50 backbone is initialized with ImageNet pre-trained weights. The three independent branches after res_conv4_1 are initialized using the identical pre-trained weights corresponding to layers res_conv4_2 through res_conv5_3.
    • Mini-Batch Construction: Each mini-batch is sampled by selecting P=16P = 16 distinct person identities and K=4K = 4 random images per identity, yielding a batch size of 64.
    • Optimization: Stochastic Gradient Descent (SGD) with a momentum of 0.90.9 and an L2L_2 weight decay factor of 0.00050.0005. Training runs for 8080 epochs in total.
    • Learning Rate Schedule: Initial learning rate is set to 0.010.01, and decayed by a factor of 10 to 0.0010.001 at epoch 40, and to 0.00010.0001 at epoch 60.
    • Loss Hyperparameters: Triplet loss margin is set to α=1.2\alpha = 1.2.
    • Evaluation Procedure: At test time, feature representations are extracted for both the original test image and its horizontally flipped version. The element-wise average of these two vectors forms the final feature vector. Person retrieval is evaluated using Cumulative Matching Characteristics (Rank-1, Rank-5, Rank-10) and mean Average Precision (mAP) computed over Euclidean distances in the 2048-dimensional embedding space.
  5. Knowl 5 — Person Re-Identification Performance on Standard Benchmarks

    empirical result

    MGN achieves state-of-the-art performance across three standard person re-identification benchmarks (Market-1501, DukeMTMC-reID, and CUHK03):

    1. Market-1501:

      • Single-query without re-ranking: Rank-1 accuracy of 95.7%95.7\% and mAP of 86.9%86.9\%, outperforming the previous best part-based method (PCB+RPP at 93.8%93.8\% Rank-1 / 81.6%81.6\% mAP).
      • Single-query with kk-reciprocal re-ranking (RK): Rank-1 accuracy of 96.6%96.6\% and mAP of 94.2%94.2\%.
      • Multiple-query: Rank-1 of 96.9%96.9\% / mAP of 90.7%90.7\% without re-ranking, and Rank-1 of 97.1%97.1\% / mAP of 95.9%95.9\% with re-ranking.
    2. DukeMTMC-reID:

      • Achieves Rank-1 accuracy of 88.7%88.7\% and mAP of 78.4%78.4\%, exceeding the previous state-of-the-art GP-reid (85.2%85.2\% Rank-1 / 72.8%72.8\% mAP) by +3.5%+3.5\% Rank-1 and +5.6%+5.6\% mAP.
    3. CUHK03 (evaluated under the 767/700 train/test identity split protocol):

      • Labeled bounding boxes: Rank-1 of 68.0%68.0\% and mAP of 67.4%67.4\%.
      • Detected (DPM) bounding boxes: Rank-1 of 66.8%66.8\% and mAP of 66.0%66.0\%, outperforming PCB+RPP (63.7%63.7\% Rank-1 / 57.5%57.5\% mAP on detected).
  6. Knowl 6 — Ablation Analysis of MGN Components, Granularities, and Multi-Branch Architecture

    data/table

    Ablation experiments conducted on the Market-1501 dataset in single-query mode evaluate individual sub-branches, multi-branch cooperation, partition granularities, and the role of triplet loss:

    Model Rank-1 (%) Rank-5 (%) Rank-10 (%) mAP (%)
    ResNet-50 Baseline 87.5 94.9 96.7 71.4
    ResNet-101 Baseline 90.4 95.7 97.2 78.0
    ResNet-50 + TP (Global Single) 88.7 96.0 97.2 75.0
    Global (Branch) 89.8 95.8 97.5 78.5
    Part-2 (Single) 92.6 97.1 98.0 80.2
    Part-2 (Branch) 94.4 97.9 98.8 83.9
    Part-3 (Single) 93.1 97.6 98.7 82.1
    Part-3 (Branch) 94.4 98.2 98.8 84.1
    G+P2+P3 (Single Ensemble) 94.4 97.6 98.5 85.2
    MGN w/o Part-3 94.4 97.9 98.7 85.7
    MGN w/ Part-4 95.1 98.3 98.9 86.1
    MGN (Part2+4) 94.8 98.2 98.9 85.6
    MGN (Part3+4) 95.0 98.1 98.8 86.1
    MGN w/o TP 95.3 97.9 98.7 86.2
    MGN (Full) 95.7 98.3 99.0 86.9

    Key observations include:

    • Architecture vs. Parameter Count: MGN without triplet loss outperforms ResNet-101 (95.3%95.3\% vs 90.4%90.4\% Rank-1, 86.2%86.2\% vs 78.0%78.0\% mAP), demonstrating that performance gains stem from the multi-granularity structure rather than merely adding parameters.
    • Multi-Branch vs. Ensemble: MGN full model outperforms the ensemble of three separately trained single networks (G+P2+P3 (Single Ensemble)) by +1.3%+1.3\% in Rank-1 and +1.7%+1.7\% in mAP, because the shared lower layers allow branches to collaboratively reinforce discriminative feature learning.
    • Sub-Branch Synergy: Features extracted from individual sub-branches within MGN consistently outperform the corresponding standalone single networks (e.g., Part-2 Branch achieves 94.4%94.4\% Rank-1 vs. 92.6%92.6\% for Part-2 Single), indicating mutual enhancement through joint backbone training.
    • Granularity Configurations: The default Part-2 + Part-3 partition scheme achieves the best performance. Skipping Part-3 (Part-2 + Part-4) degrades Rank-1 by 0.9%0.9\% and mAP by 1.3%1.3\%. Comparing Part-3 + Part-4 to Part-2 + Part-4 shows that overlapping partition boundaries (introduced by 3 and 4 splits) improve cross-stripe correlation learning compared to non-overlapping boundaries (2 and 4 splits).
    • Triplet Loss Contribution: Adding batch-hard triplet loss to reduced global features improves MGN by +0.4%+0.4\% Rank-1 and +0.7%+0.7\% mAP (and improves the ResNet-50 baseline by +1.2%+1.2\% Rank-1 and +3.6%+3.6\% mAP), with a larger relative impact on mAP due to its metric ranking regularization.
  7. Knowl 7 — Specialization of Spatial Feature Responses Across Granularities

    empirical result

    Analysis of spatial activation response maps (computed as the L2L_2-norm of feature channels across spatial locations at the output layers of each branch) shows distinct response preferences:

    1. Global Branch: Feature responses concentrate broadly on the central torso and main pedestrian body while filtering out background clutter, but often suppress peripheral body regions such as limbs, waist, and feet.
    2. Part-2 Branch: When down-sampling is omitted and the feature map is split into two horizontal stripes, activations shift from global body centers to intermediate structural landmarks, showing localized concentration on shoulders and major joints.
    3. Part-3 Branch: With three horizontal stripes, activation maps become more dispersed and finely focused on specific high-saliency semantic patterns (such as logos, t-shirt emblems, or backpack straps).

    This confirms that dividing unreduced feature representations into increasing numbers of uniform horizontal stripes forces deep networks to extract fine-grained discriminative details without requiring explicit semantic part estimators or pose detectors.

Coverage note — None. All core methodological contributions (architecture, multi-loss design, training scheme), empirical results, and ablation findings have been covered.

References

  1. 1.Ejaz Ahmed, Michael Jones, and Tim K Marks. 2015. An improved deep learning architecture for person re-identification. In CVPR. 3908–3916.
  2. 2.Jon Almazan, Bojana Gajic, Naila Murray, and Diane Larlus. 2018. Re-ID done right: towards good practices for person re-identification. arXiv preprint arXiv:1801.05339 (2018).
  3. 3.Xiang Bai, Mingkun Yang, Tengteng Huang, Zhiyong Dou, Rui Yu, and Yongchao Xu. 2017. Deep-Person: Learning Discriminative Deep Features for Person Re-Identification. arXiv preprint arXiv:1711.10658 (2017).
  4. 4.Xiaobin Chang, Timothy M. Hospedales, and Tao Xiang. 2018. Multi-Level Factorisation Net for Person Re-Identification. In CVPR. 2109–2118.
  5. 5.Dapeng Chen, Dan Xu, Hongsheng Li, Nicu Sebe, and Xiaogang Wang. 2018. Group Consistent Similarity Learning via Deep CRF for Person Re-Identification. In CVPR. 8649–8658.
  6. 6.Weihua Chen, Xiaotang Chen, Jianguo Zhang, and Kaiqi Huang. 2017. Beyond triplet loss: a deep quadruplet network for person re-identification. In CVPR. 403–412.
  7. 7.Yanbei Chen, Xiatian Zhu, and Shaogang Gong. 2017. Person Re-Identification by Deep Learning Multi-Scale Representations. In ICCV. 2590–2600.
  8. 8.De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, and Nanning Zheng. 2016. Person re-identification by multi-channel parts-based cnn with improved triplet loss function. In CVPR. 1335–1344.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In CVPR. 248–255.
  10. 10.Pedro Felzenszwalb, David McAllester, and Deva Ramanan. 2008. A discriminatively trained, multiscale, deformable part model. In CVPR. 1–8.
  11. 11.Ross Girshick. 2015. Fast r-cnn. In ICCV. 1440–1448.
  12. 12.Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006. Dimensionality reduction by learning an invariant mapping. In CVPR, Vol. 2. 1735–1742.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR. 770–778.
  14. 14.Alexander Hermans, Lucas Beyer, and Bastian Leibe. 2017. In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737 (2017).
  15. 15.Elad Hoffer and Nir Ailon. 2015. Deep metric learning using triplet network. In International Workshop on Similarity-Based Pattern Recognition. Springer, 84–92.
  16. 16.Houjing Huang, Dangwei Li, Zhang Zhang, Xiaotang Chen, and Kaiqi Huang. 2018. Adversarially Occluded Samples for Person Re-Identification. In CVPR. 5098–5107.
  17. 17.Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML. 448–456.
  18. 18.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. 2015. Spatial transformer networks. In NIPS. 2017–2025.
  19. 19.Dangwei Li, Xiaotang Chen, Zhang Zhang, and Kaiqi Huang. 2017. Learning deep context-aware features over body and latent parts for person re-identification. In CVPR. 384–393.
  20. 20.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. 2014. Deepreid: Deep filter pairing neural network for person re-identification. In CVPR. 152–159.
  21. 21.Wei Li, Xiatian Zhu, and Shaogang Gong. 2017. Person re-identification by deep joint learning of multi-loss classification. In IJCAI. 2194–2200.
  22. 22.Wei Li, Xiatian Zhu, and Shaogang Gong. 2018. Harmonious Attention Network for Person Re-Identification. In CVPR. 2285–2294.
  23. 23.Shengcai Liao, Yang Hu, Xiangyu Zhu, and Stan Z Li. 2015. Person re-identification by local maximal occurrence representation and metric learning. In CVPR. 2197–2206.
  24. 24.Hao Liu, Jiashi Feng, Meibin Qi, Jianguo Jiang, and Shuicheng Yan. 2017. End-to-end comparative attention networks for person re-identification. IEEE Transactions on Image Processing 26, 7 (2017), 3492–3506.
  25. 25.Xihui Liu, Haiyu Zhao, Maoqing Tian, Lu Sheng, Jing Shao, Shuai Yi, Junjie Yan, and Xiaogang Wang. 2017. HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis. In CVPR. 350–359.
  26. 26.Ergys Ristani, Francesco Solera, Roger Zou, Rita Cucchiara, and Carlo Tomasi. 2016. Performance Measures and a Data Set for Multi-Target, Multi-Camera Tracking. In ECCV workshop on Benchmarking Multi-Target Tracking. 17–35.
  27. 27.M. Saquib Sarfraz, Arne Schumann, Andreas Eberle, and Rainer Stiefelhagen. 2018. A Pose-Sensitive Embedding for Person Re-Identification With Expanded Cross Neighborhood Re-Ranking. In CVPR. 420–429.
  28. 28.Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. Facenet: A unified embedding for face recognition and clustering. In CVPR. 815–823.
  29. 29.Yantao Shen, Hongsheng Li, Tong Xiao, Shuai Yi, Dapeng Chen, and Xiaogang Wang. 2018. Deep Group-Shuffling Random Walk for Person Re-Identification. In CVPR. 2265–2274.
  30. 30.Yantao Shen, Tong Xiao, Hongsheng Li, Shuai Yi, and Xiaogang Wang. 2018. End-to-End Deep Kronecker-Product Matching for Person Re-Identification. In CVPR. 6886–6895.
  31. 31.Jianlou Si, Honggang Zhang, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex C. Kot, and Gang Wang. 2018. Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification. In CVPR. 5363–5372.
  32. 32.Hyun Oh Song, Yu Xiang, Stefanie Jegelka, and Silvio Savarese. 2016. Deep metric learning via lifted structured feature embedding. In CVPR. 4004–4012.
  33. 33.Chi Su, Jianing Li, Shiliang Zhang, Junliang Xing, Wen Gao, and Qi Tian. 2017. Pose-driven Deep Convolutional Model for Person Re-identification. In ICCV. 3980–3989.
  34. 34.Yi Sun, Xiaogang Wang, and Xiaoou Tang. 2015. Deeply learned face representations are sparse, selective, and robust. In CVPR. 2892–2900.
  35. 35.Yifan Sun, Liang Zheng, Weijian Deng, and Shengjin Wang. 2017. SVDNet for Pedestrian Retrieval. In ICCV. 2590–2600.
  36. 36.Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. 2018. Beyond Part Models: Person Retrieval with Refined Part Pooling. In ECCV. In press.
  37. 37.Rahul Rama Varior, Mrinal Haloi, and Gang Wang. 2016. Gated siamese convolutional neural network architecture for human re-identification. In ECCV. Springer, 791–808.
  38. 38.Feng Wang, Xiang Xiang, Jian Cheng, and Alan Loddon Yuille. 2017. Normface: l2 hypersphere embedding for face verification. In 2017 ACM on Multimedia Conference. 1041–1049.
  39. 39.Tong Xiao, Hongsheng Li, Wanli Ouyang, and Xiaogang Wang. 2016. Learning deep feature representations with domain guided dropout for person re-identification. In CVPR. 1249–1258.
  40. 40.Jing Xu, Rui Zhao, Feng Zhu, Huaming Wang, and Wanli Ouyang. 2018. Attention-Aware Compositional Network for Person Re-Identification. In CVPR. 2119–2128.
  41. 41.Hantao Yao, Shiliang Zhang, Yongdong Zhang, Jintao Li, and Qi Tian. 2017. Deep representation learning with part loss for person re-identification. arXiv preprint arXiv:1707.00798 (2017).
  42. 42.Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z Li. 2014. Deep metric learning for person re-identification. In ICPR. 34–39.
  43. 43.Xuan Zhang, Hao Luo, Xing Fan, Weilai Xiang, Yixiao Sun, Qiqi Xiao, Wei Jiang, Chi Zhang, and Jian Sun. 2017. Alignedreid: Surpassing human-level performance in person re-identification. arXiv preprint arXiv:1711.08184 (2017).
  44. 44.Haiyu Zhao, Maoqing Tian, Shuyang Sun, Jing Shao, Junjie Yan, Shuai Yi, Xiaogang Wang, and Xiaoou Tang. 2017. Spindle net: Person re-identification with human body region guided feature decomposition and fusion. In CVPR. 1077–1085.
  45. 45.Liming Zhao, Xi Li, Jingdong Wang, and Yueting Zhuang. 2017. Deeply-learned part-aligned representations for person re-identification. In ICCV. 3219–3228.
  46. 46.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. 2015. Scalable Person Re-identification: A Benchmark. In ICCV. 1116–1124.
  47. 47.Liang Zheng, Yi Yang, and Alexander G Hauptmann. 2016. Person re-identification: Past, present and future. arXiv preprint arXiv:1610.02984 (2016).
  48. 48.Zhedong Zheng, Liang Zheng, and Yi Yang. 2017. Pedestrian alignment network for large-scale person re-identification. arXiv preprint arXiv:1707.00408 (2017).
  49. 49.Zhedong Zheng, Liang Zheng, and Yi Yang. 2017. Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in vitro. In ICCV. 3774–3782.
  50. 50.Zhun Zhong, Liang Zheng, Donglin Cao, and Shaozi Li. 2017. Re-ranking person re-identification with k-reciprocal encoding. In CVPR. 3652–3661.

Citation

MLA
Wang, G., et al. “Learning Discriminative Features with Multiple Granularities for Person Re-Identification”. Proceedings of the 26th ACM International Conference on Multimedia, 2018, pp. 274–82, https://doi.org/10.1145/3240508.3240552.
APA
Wang, G., Yuan, Y., Chen, X., Li, J., & Zhou, X. (2018). Learning Discriminative Features with Multiple Granularities for Person Re-Identification. Proceedings of the 26th ACM International Conference on Multimedia, 274–282. https://doi.org/10.1145/3240508.3240552
Chicago
Wang, G., Y. Yuan, X. Chen, J. Li, and X. Zhou. 2018. “Learning Discriminative Features with Multiple Granularities for Person Re-Identification”. Proceedings of the 26th ACM International Conference on Multimedia, 274–82. https://doi.org/10.1145/3240508.3240552.
Harvard
Wang, G. et al. (2018) “Learning Discriminative Features with Multiple Granularities for Person Re-Identification”, Proceedings of the 26th ACM international conference on Multimedia. ACM, pp. 274–282. Available at: https://doi.org/10.1145/3240508.3240552.
Vancouver
1. Wang G, Yuan Y, Chen X, Li J, Zhou X (2018) Learning Discriminative Features with Multiple Granularities for Person Re-Identification. In: Proceedings of the 26th ACM international conference on Multimedia. ACM, pp 274–282

BibTeX

@inproceedings{Wang_2018, series={MM ’18}, title={Learning Discriminative Features with Multiple Granularities for Person Re-Identification}, url={http://dx.doi.org/10.1145/3240508.3240552}, DOI={10.1145/3240508.3240552}, booktitle={Proceedings of the 26th ACM international conference on Multimedia}, publisher={ACM}, author={Wang, Guanshuo and Yuan, Yufeng and Chen, Xiong and Li, Jiwei and Zhou, Xi}, year={2018}, month=Oct, pages={274–282}, collection={MM ’18} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF