Learning Discriminative Features with Multiple Granularities for Person Re-Identification
Guanshuo WangYufeng YuanXiong ChenJiwei LiXi Zhou
Introduces the Multiple Granularity Network (MGN), a multi-branch deep architecture that combines global features with multi-level stripe partitions to surpass semantic part-based methods across standard person re-identification benchmarks.
Identifying and tracking individuals across non-overlapping surveillance camera networks—a task known as person re-identification—is essential for modern public safety and security operations. However, automated systems often struggle due to variations in pedestrian poses, occlusions, background clutter, and image resolution. Previous methods attempted to solve this by locating specific body parts or using complex attention mechanisms, which added significant computational complexity, lacked robustness to dramatic appearance changes, and required multi-stage training pipelines.
The article demonstrates an end-to-end deep learning framework, termed the Multiple Granularity Network, designed to overcome these challenges. The system extracts and combines full-body global visual features with multi-level local part details using simple, uniform horizontal image stripes rather than complex body-part detectors.
To evaluate the system, the authors conducted extensive experiments on three benchmark surveillance datasets: Market-1501, DukeMTMC-reID, and CUHK03. The architecture modifies a standard visual backbone into three parallel processing branches: one capturing full-body global features, a second dividing feature maps into two horizontal stripes, and a third dividing them into three stripes. The model combines classification objectives for individual identity prediction with ranking objectives (batch-hard metric loss) using a specialized coarse-to-fine supervisory design during training.
The findings show that the Multiple Granularity Network significantly outperforms previous approaches across all benchmarks. On the Market-1501 dataset, the model achieved a 95.7% Rank-1 retrieval accuracy and an 86.9% mean average precision (improving to 96.6% and 94.2% when applying an automated re-ranking step), exceeding prior top systems by 1.9% and 5.3% respectively. On the challenging DukeMTMC-reID dataset, it reached 88.7% Rank-1 accuracy and 78.4% mean average precision, beating existing benchmarks by 3.5% and 5.6%. Ablation analyses confirmed that combining global, two-part, and three-part representations in a single coordinated network yielded substantially higher precision than training separate standalone networks or simply increasing model size.
These results indicate that automated surveillance systems can achieve high identification reliability without relying on complex, failure-prone human pose estimators or external semantic detectors. The multi-granularity representation effectively captures subtle local visual cues—such as logos, straps, or specific clothing patterns—even under partial occlusion, viewpoint shifts, or low image resolution. This architectural simplicity reduces engineering overhead while improving operational accuracy.
Organizations deploying automated visual tracking systems should consider adopting multi-granularity feature extraction pipelines to maximize matching precision. As an actionable next step, decision-makers should also invest in high-quality upstream person-detection algorithms; the experimental results showed a measurable accuracy decline on automatically detected bounding boxes compared to perfectly cropped images, highlighting bounding-box quality as a primary real-world operational bottleneck.
- Paper: Beyond Part Models: Person Retrieval with Refined Part Pooling, Yifan Sun et al. (2017). Introduces uniform horizontal stripe partitioning and refined pooling for part-based person re-identification, providing the direct structural foundation for the Multiple Granularity Network.
- Paper: In Defense of the Triplet Loss for Person Re-Identification, Alexander Hermans et al. (2017). Establishes the modern batch-hard triplet loss framework for training deep person re-identification backbones that the Multiple Granularity Network directly employs across its branches.
- Paper: Scalable Person Re-identification: A Benchmark, Liang Zheng et al. (2015). Introduces the Market-1501 dataset and key evaluation benchmarks that define the standard experimental framework and baselines used in this work.
- Paper: Re-ranking Person Re-identification with k-Reciprocal Encoding, Zhun Zhong et al. (2017). Develops the k-reciprocal encoding re-ranking post-processing algorithm that is critical to achieving the state-of-the-art results reported in this paper.
- Paper: Deep Metric Learning via Lifted Structured Feature Embedding, Hyun Oh Song et al. (2015). Provides fundamental techniques in structured deep metric learning for image retrieval across mini-batches.
- Paper: Improved Deep Metric Learning with Multi-class N-pair Loss Objective, Kihyuk Sohn (2016). Develops multi-class tuple-based metric learning losses that paved the way for effective multi-branch deep feature supervision.
- Paper: Person re-identification by Local Maximal Occurrence representation and metric learning, Shengcai Liao et al. (2014). Pioneers early horizontal stripe occurrence representations and cross-view metric learning on standard person re-identification benchmarks.
- Paper: Deep Learning for Person Re-Identification: A Survey and Outlook, Mang Ye et al. (2020). Synthesizes multi-granularity part models and ranking optimization into a comprehensive survey and introduces a unified benchmark baseline.
- Paper: FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking, Yifu Zhang et al. (2020). Applies joint feature representations from deep person re-identification to unified multi-object tracking architectures.
