Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function

De ChengYihong GongSanping ZhouJinjun WangN. Zheng

article2016CVPR1,327 citations

Presents a multi-channel convolutional architecture paired with an improved triplet loss that enforces an upper bound on positive pair distances, jointly learning global and body-part representations to improve person re-identification across non-overlapping camera views.

Listen

Identifying the same individual across multiple non-overlapping surveillance cameras remains a critical yet difficult task in automated security, robotics, and video analytics. The core challenges arise from severe variations in camera angles, lighting conditions, human body poses, and occlusions, combined with the fact that facial features are rarely clear enough for standard biometrics. Historically, systems tackled this by treating visual feature extraction and distance comparison as separate steps, which limited overall matching performance. The main objective of the article is to demonstrate an integrated deep learning approach that jointly learns both full-body and body-part visual features alongside an optimized distance metric using an enhanced loss function.

The authors evaluated their approach using standard experimental benchmarks on four public surveillance datasets: i-LIDS, VIPeR, PRID2011, and CUHK01. Their technical strategy uses a single multi-channel deep neural network that simultaneously processes a global full-body channel and four regional body-part channels with distinct filter sizes. This model was trained on image triplets—matching pairs versus non-matching individuals—using an improved triplet loss formulation. Unlike standard triplet loss, which only ensures that non-matching images are farther apart than matching ones, the improved loss explicitly enforces a compact margin on matching pairs, pulling features of the same individual much closer together.

The experimental findings show significant performance gains across all benchmark datasets. First, the proposed framework achieved top-ranked identification rates, outperforming leading traditional, deep learning, and metric ensemble methods across all four datasets. Second, incorporating body-part channels significantly boosted matching accuracy, providing up to a 13% improvement over full-body-only models. Third, replacing standard triplet loss with the improved loss function alone increased identification accuracy by up to 4%. Finally, detailed regional analysis revealed that the upper body (face and shoulders) provides the most discriminative and reliable features, while the lower body (legs and feet) contributes the least due to high movement variability.

These results imply that surveillance and retrieval systems can achieve substantially higher automated tracking accuracy without manual intervention or multi-stage pipelines. By capturing both holistic body appearance and fine-grained local parts within a single unified network, the model improves recognition reliability under extreme viewpoint and posture shifts. For operational decision-makers, this translates to reduced false-match rates and improved real-time tracking performance across disconnected camera networks.

Moving forward, stakeholders should consider adopting joint global-and-local feature learning frameworks when upgrading automated video surveillance infrastructure. The article suggests extending this multi-channel triplet architecture to broader multimedia tasks, such as large-scale image and video retrieval. Future work should validate performance on larger, unconstrained real-world deployments, as the current evaluation relies on standardized benchmark splits and requires tuning network depth and loss weighting parameters to match specific dataset scales.

Cover for Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function

Abstract

Person re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The Proposed Person Re-Id Method
  • 3.1. The Overall Framework
  • 3.2. Multi-Channel Parts-based CNN Model
  • 3.3. Improved Triplet Loss Function
  • 3.4. The Training Algorithm
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Experimental Evaluations
  • 4.3. Analysis of different body parts
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Multi-Channel Parts-Based CNN Architecture for Person Re-Identification

    model/method

    The multi-channel parts-based convolutional neural network (CNN) jointly extracts global full-body representations and local body-part representations for person re-identification. Given an input pedestrian image (resized to 100×250100 \times 250 pixels and cropped to 80×23080 \times 230 pixels), the network architecture consists of the following components:

    • Global Convolution Layer (G-conv1): A shared initial convolution layer with 32 feature maps, a kernel size of 7×7×37 \times 7 \times 3, and a stride of 3 pixels, followed by a rectified linear unit (ReLU) activation.
    • Full-Body Channel: Takes the full output of G-conv1, applies max pooling with a 3×33 \times 3 kernel (B-pool1), feeds into a full-body convolution layer (B-conv2) with 32 feature maps of kernel size 5×55 \times 5 (stride 1, ReLU), passes through a second 3×33 \times 3 max pooling layer (B-pool2), and connects to a fully connected layer (B-fc) producing a 400-dimensional feature vector.
    • Four Body-Part Channels: The feature map output of G-conv1 is vertically divided into four equal non-overlapping horizontal parts PiP_i for i∈{1,2,3,4}i \in \{1, 2, 3, 4\}. Each part forms the input to an independent body-part channel comprising a convolution layer (PiP_i-conv2) with 32 feature maps of kernel size 3×33 \times 3 (stride 1, ReLU, no pooling) and a channel-wise fully connected layer (PiP_i-fc) producing a 100-dimensional feature vector.
    • Network-Wise Fusion (N-fc): The 400-dimensional vector from the full-body channel and the four 100-dimensional vectors from the body-part channels are concatenated into an 800-dimensional vector and passed to a final network-wise fully connected layer that outputs the final 800-dimensional representation ϕw(I)\phi_w(I).

    For larger datasets such as CUHK01, a deeper variant is used where each of the five individual channels contains two convolution layers instead of one before its fully connected layer.

  2. Knowl 2 — Improved Triplet Loss Function for Person Re-Identification

    equation

    Given a triplet of training images Ii=⟨Iio,Ii+,Ii−⟩I_i = \langle I_i^o, I_i^+, I_i^- \rangle, where IioI_i^o is an anchor image, Ii+I_i^+ is a positive match of the same person, and Ii−I_i^- is a negative match of a different person, let ϕw(I)\phi_w(I) denote the feature embedding produced by the network with parameter set ww. The squared L2L_2-norm distance between two feature representations is:

    d(ϕw(Ia),ϕw(Ib))=∥ϕw(Ia)−ϕw(Ib)∥22d(\phi_w(I_a), \phi_w(I_b)) = \|\phi_w(I_a) - \phi_w(I_b)\|_2^2

    The improved triplet loss enforces both an inter-class margin condition and an intra-class compactness condition across NN triplet training examples:

    L(I,w)=1N∑i=1N(max⁡{dn(Iio,Ii+,Ii−,w),τ1}+βmax⁡{dp(Iio,Ii+,w),τ2})L(I, w) = \frac{1}{N} \sum_{i=1}^N \left( \max\{d_n(I_i^o, I_i^+, I_i^-, w), \tau_1\} + \beta \max\{d_p(I_i^o, I_i^+, w), \tau_2\} \right)

    where:

    • dn(Iio,Ii+,Ii−,w)=d(ϕw(Iio),ϕw(Ii+))−d(ϕw(Iio),ϕw(Ii−))d_n(I_i^o, I_i^+, I_i^-, w) = d(\phi_w(I_i^o), \phi_w(I_i^+)) - d(\phi_w(I_i^o), \phi_w(I_i^-)) is the relative inter-class difference, and τ1<0\tau_1 < 0 (set to τ1=−1\tau_1 = -1) enforces that the intra-class distance is smaller than the inter-class distance by at least ∣τ1∣|\tau_1|.
    • dp(Iio,Ii+,w)=d(ϕw(Iio),ϕw(Ii+))d_p(I_i^o, I_i^+, w) = d(\phi_w(I_i^o), \phi_w(I_i^+)) is the intra-class distance, and τ2>0\tau_2 > 0 (set to τ2=0.01\tau_2 = 0.01, with τ2≪∣τ1∣\tau_2 \ll |\tau_1|) restricts positive pairs to lie within a small distance threshold τ2\tau_2 to prevent wide intra-class dispersion.
    • β≥0\beta \ge 0 (set to β=0.002\beta = 0.002) is a weighting parameter balancing the intra-class compactness penalty relative to the inter-class ranking penalty.
  3. Knowl 3 — Triplet-Based Stochastic Gradient Descent Algorithm for Improved Triplet Loss

    algorithm

    The parameters ww of the multi-channel CNN are trained by stochastic gradient descent minimizing the improved triplet loss function. The gradient with respect to network parameters ww is formulated as:

    ∂L(I,w)∂w=1N∑i=1Nh1(Ii,w)+1N∑i=1Nh2(Ii,w)\frac{\partial L(I, w)}{\partial w} = \frac{1}{N} \sum_{i=1}^N h_1(I_i, w) + \frac{1}{N} \sum_{i=1}^N h_2(I_i, w)

    where the loss-component sub-gradients are:

    h1(Ii,w)={∂dn(Iio,Ii+,Ii−,w)∂w,if dn(Iio,Ii+,Ii−,w)>τ10,if dn(Iio,Ii+,Ii−,w)≤τ1h_1(I_i, w) = \begin{cases} \frac{\partial d_n(I_i^o, I_i^+, I_i^-, w)}{\partial w}, & \text{if } d_n(I_i^o, I_i^+, I_i^-, w) > \tau_1 \\ 0, & \text{if } d_n(I_i^o, I_i^+, I_i^-, w) \le \tau_1 \end{cases} h2(Ii,w)={β∂dp(Iio,Ii+,w)∂w,if dp(Iio,Ii+,w)>τ20,if dp(Iio,Ii+,w)≤τ2h_2(I_i, w) = \begin{cases} \beta \frac{\partial d_p(I_i^o, I_i^+, w)}{\partial w}, & \text{if } d_p(I_i^o, I_i^+, w) > \tau_2 \\ 0, & \text{if } d_p(I_i^o, I_i^+, w) \le \tau_2 \end{cases}

    with the distance gradients calculated from forward activations ϕw(⋅)\phi_w(\cdot) and backward propagation derivatives ∂ϕw(⋅)∂w\frac{\partial \phi_w(\cdot)}{\partial w}:

    ∂dp∂w=2(ϕw(Iio)−ϕw(Ii+))∂ϕw(Iio)−∂ϕw(Ii+)∂w\frac{\partial d_p}{\partial w} = 2(\phi_w(I_i^o) - \phi_w(I_i^+)) \frac{\partial \phi_w(I_i^o) - \partial \phi_w(I_i^+)}{\partial w} ∂dn∂w=2(ϕw(Iio)−ϕw(Ii+))∂ϕw(Iio)−∂ϕw(Ii+)∂w−2(ϕw(Iio)−ϕw(Ii−))∂ϕw(Iio)−∂ϕw(Ii−)∂w\frac{\partial d_n}{\partial w} = 2(\phi_w(I_i^o) - \phi_w(I_i^+)) \frac{\partial \phi_w(I_i^o) - \partial \phi_w(I_i^+)}{\partial w} - 2(\phi_w(I_i^o) - \phi_w(I_i^-)) \frac{\partial \phi_w(I_i^o) - \partial \phi_w(I_i^-)}{\partial w}

    The full training procedure is structured as follows:

    Input: Training triplet set {Ii}\{I_i\}, learning rate schedule λt\lambda_t, maximum iterations TT, thresholds τ1,τ2\tau_1, \tau_2, weight β\beta
    Output: Learned network parameters ww
    t←0t \leftarrow 0
    while t<Tt < T do
        t←t+1t \leftarrow t + 1
        ∂L(I,w)∂w←0\frac{\partial L(I, w)}{\partial w} \leftarrow 0
        for all training triplet samples Ii=⟨Iio,Ii+,Ii−⟩I_i = \langle I_i^o, I_i^+, I_i^- \rangle in the mini-batch do
            Calculate ϕw(Iio)\phi_w(I_i^o), ϕw(Ii+)\phi_w(I_i^+), ϕw(Ii−)\phi_w(I_i^-) by forward propagation
            Calculate ∂ϕw(Iio)∂w\frac{\partial \phi_w(I_i^o)}{\partial w}, ∂ϕw(Ii+)∂w\frac{\partial \phi_w(I_i^+)}{\partial w}, ∂ϕw(Ii−)∂w\frac{\partial \phi_w(I_i^-)}{\partial w} by backpropagation
            Calculate ∂dp∂w\frac{\partial d_p}{\partial w} and ∂dn∂w\frac{\partial d_n}{\partial w}
            Calculate h1(Ii,w)h_1(I_i, w) and h2(Ii,w)h_2(I_i, w) and accumulate gradient into ∂L(I,w)∂w\frac{\partial L(I, w)}{\partial w}
        end for
        Update network parameters: wt←wt−1−λt∂L(I,w)∂ww^t \leftarrow w^{t-1} - \lambda_t \frac{\partial L(I, w)}{\partial w}
    end while
  4. Knowl 4 — Training Protocols, Triplet Sampling, and Evaluation Setup for Person Re-Identification

    experimental setup
    • Data Augmentation: All input pedestrian images are normalized to 100×250100 \times 250 pixels. During training, a bounding box of 80×23080 \times 230 pixels is cropped around the center with a small random spatial perturbation.
    • Weight Initialization: Network weights are initialized from two zero-mean Gaussian distributions with standard deviations of 0.010.01 and 0.0010.001, respectively; bias terms are initialized to 0.
    • Batch Construction and Triplet Sampling: Each mini-batch contains 100 instances constructed by selecting 5 distinct identities and generating 20 triplets per identity per iteration. In each triplet ⟨Io,I+,I−⟩\langle I^o, I^+, I^- \rangle, the positive match I+I^+ is randomly chosen from the same identity class as IoI^o, and the negative match I−I^- is randomly chosen from the remaining identity classes. Hyperparameters in the improved triplet loss are set to τ1=−1\tau_1 = -1, τ2=0.01\tau_2 = 0.01, and β=0.002\beta = 0.002.
    • Evaluation Protocol: Evaluations follow standard Cumulative Match Characteristic (CMC) curves. For each dataset, approximately half of the identities are randomly sampled for training and the remaining half for testing. In two-camera setups, one image from camera A acts as the query and one from camera B acts as the gallery (one gallery image per person). Matching distances between query and gallery embeddings ϕw(I)\phi_w(I) are computed using L2L_2 distance. Top-kk accuracy is reported as the average over 10 random train/test splits across four datasets: i-LIDS (119 persons, 479 images), PRID2011 (200 persons appearing in both static cameras A and B), VIPeR (632 persons, 2 camera views), and CUHK01 (971 persons, 4 images each from 2 camera views).
  5. Knowl 5 — Ablation Analysis of Multi-Channel Architecture and Improved Triplet Loss

    empirical result

    To evaluate the individual contributions of the parts-based multi-channel structure and the intra-class constraint in the improved triplet loss, four model configurations were compared on the i-LIDS, PRID2011, VIPeR, and CUHK01 datasets:

    • OursT / Ours3T: Global full-body channel only, trained using the standard triplet loss function without the intra-class constraint (β=0\beta = 0).
    • OursTC / Ours3TC: Global full-body channel only, trained using the improved triplet loss function (β=0.002,τ2=0.01\beta = 0.002, \tau_2 = 0.01).
    • OursTP / Ours3TP: Complete multi-channel architecture (1 full-body channel + 4 body-part channels), trained using the standard triplet loss function.
    • OursTCP / Ours3TCP: Complete multi-channel architecture, trained using the improved triplet loss function.

    The comparative results establish two key findings:

    1. Effect of Improved Triplet Loss: Adding the intra-class distance constraint improves Rank-1 accuracy by up to 4%4\% across both single-channel and multi-channel architectures (e.g., on i-LIDS, Rank-1 increases from 43.2%43.2\% for OursT to 47.3%47.3\% for OursTC, and from 57.2%57.2\% for OursTP to 60.4%60.4\% for OursTCP; on VIPeR, Rank-1 increases from 34.3%34.3\% to 37.2%37.2\% for single-channel and from 43.8%43.8\% to 47.8%47.8\% for multi-channel).
    2. Effect of Body-Part Channels: Jointly learning body-part features alongside the full body yields an accuracy boost of up to 13.1%13.1\% over the full-body channel alone (e.g., on i-LIDS, Rank-1 increases from 47.3%47.3\% for OursTC to 60.4%60.4\% for OursTCP; on VIPeR, Rank-1 increases from 37.2%37.2\% for OursTC to 47.8%47.8\% for OursTCP).
  6. Knowl 6 — Cumulative Match Characteristic Performance on Standard Re-Identification Datasets

    data/table

    The proposed full model (OursTCP for i-LIDS, PRID2011, and VIPeR; Ours3TCP for CUHK01) was evaluated against representative metric learning, salience-based, and deep CNN methods using the Cumulative Match Characteristic (CMC) ranking metric across top-1 to top-30/100 ranks under standard 50/50 train/test splits averaged over 10 trials.

    i-LIDS Dataset Top1 Top5 Top10 Top15 Top20 Top30
    Adaboost 29.6 55.2 68.1 77.0 82.4 92.1
    LMNN 28.0 53.8 66.1 75.5 82.3 91.0
    ITML 29.0 54.0 70.5 81.0 86.7 95.0
    MCC 31.3 59.3 75.6 84.0 88.3 95.0
    PRDC 37.8 63.7 75.1 82.8 88.4 95.0
    Sakrapee (Metric Ensemble) 50.3 – – – – –
    Ding (Deep Relative Comparison) 52.1 68.2 78.0 83.6 88.8 95.0
    OursTCP 60.4 82.7 90.7 96.4 97.8 99.3
    PRID2011 Dataset Top1 Top10 Top20 Top50 Top100
    KISSME 15.0 39.0 52.0 68.0 80.0
    EIML 16.0 39.0 51.0 68.0 81.0
    LMNN 10.0 30.0 42.0 59.0 73.0
    ITML 12.0 36.0 47.0 64.0 79.0
    Maha 16.0 41.0 51.0 64.0 76.0
    DeepM 17.9 45.9 55.4 71.4 –
    Sakrapee 17.9 – – – –
    OursTCP 22.0 47.0 57.0 76.0 83.0
    VIPeR Dataset Top1 Top5 Top10 Top15 Top20 Top30
    MtMCML 28.8 59.3 75.8 83.4 88.5 93.5
    SDALF 19.9 38.4 49.4 58.5 66.0 74.4
    eSDC 26.3 46.4 58.6 66.6 72.8 80.5
    SalMatch 30.2 52.3 66.0 73.4 79.2 86.0
    Ding 40.5 60.8 70.4 78.3 84.4 90.9
    mFilter+LADF 43.4 – – – – –
    Sakrapee 45.9 – – – – –
    OursTCP 47.8 74.7 84.8 89.2 91.1 94.3
    CUHK01 Dataset Top1 Top5 Top10 Top15 Top20 Top30
    mFilter 34.3 55.0 65.3 70.5 – –
    SalMatch 28.5 46.3 57.2 64.1 – –
    FPNN 27.9 – – – – –
    Ejaz 47.5 – – – – –
    Sakrapee 53.4 76.4 84.4 – 90.5 –
    Ours3TCP 53.7 84.3 91.0 93.3 96.3 98.3

    The proposed method outperforms all baseline traditional and deep network models across all ranking measures on the four benchmark datasets, outperforming the previous state-of-the-art ensemble method by 10.1%10.1\% Rank-1 accuracy on i-LIDS (60.4%60.4\% vs. 50.3%50.3\%), 4.1%4.1\% on PRID2011 (22.0%22.0\% vs. 17.9%17.9\%), and 1.9%1.9\% on VIPeR (47.8%47.8\% vs. 45.9%45.9\%).

  7. Knowl 7 — Sensitivity of Re-Identification Accuracy to the Intra-Class Loss Weight Hyperparameter

    data/table

    The hyperparameter β\beta in the improved triplet loss balances the relative importance of the intra-class compactness constraint and the inter-class separation constraint. A cross-validation evaluation on the VIPeR dataset demonstrates how identification accuracy (%) varies across different values of β\beta:

    β\beta Top1 Top5 Top10 Top15 Top20 Top30
    0 43.8 69.5 79.7 81.0 85.4 90.2
    0.001 45.9 73.4 81.9 87.0 93.0 95.6
    0.002 47.8 74.7 84.8 89.2 91.1 94.3
    0.003 45.6 75.3 85.4 87.6 90.5 94.6
    0.004 43.7 73.1 81.5 87.9 91.1 93.4

    Setting β=0\beta = 0 corresponds to the conventional triplet loss without an intra-class constraint. Re-identification performance peaks within the interval β∈[0.001,0.003]\beta \in [0.001, 0.003], with β=0.002\beta = 0.002 providing the highest Rank-1 identification rate (47.8%47.8\%). Increasing β\beta beyond 0.0030.003 causes performance to decline, as an excessive intra-class penalty impairs the model's ability to maintain sufficient inter-class margin separation.

  8. Knowl 8 — Spatial Body-Part Contribution to Re-Identification Accuracy

    empirical result

    To evaluate how different spatial body regions contribute to re-identification accuracy, four distinct two-channel network models were trained on the VIPeR dataset. Each model combines the full-body channel with exactly one of the four body-part channels (P1,P2,P3,P4P_1, P_2, P_3, P_4, partitioned from top to bottom of the pedestrian):

    • Ours-Part1 (Face and Shoulders): Produces the largest performance improvement over the full-body baseline (OursT).
    • Ours-Part2 (Upper Torso / Chest): Provides the second largest performance gain.
    • Ours-Part3 (Lower Torso / Hips): Yields moderate performance gain.
    • Ours-Part4 (Legs and Feet): Provides the lowest performance improvement among all four parts.

    The marginal benefit of local part features decreases monotonically from the top to the bottom of the body. This is because the head and upper torso maintain relatively stable appearance and shape across camera viewpoints, whereas the legs and feet undergo large deformations, motion blur, and pose variations during walking, yielding less reliable features for person re-identification. Combining all four parts with the full-body channel (OursTP) surpasses every single-part variant.

Coverage note — All substantial contributed materials—including the multi-channel network architecture, the improved triplet loss formulation, the SGD training algorithm, the experimental protocol and parameters, ablation analyses, benchmark comparison tables, hyperparameter sensitivity analysis, and the body-part contribution study—have been included. The qualitative visualization of convolution activations was omitted as it merely visualizes the learned features without adding standalone technical content.

References

  1. 1.E. Ahmed, M. Jones, and T. K. Marks. An improved deep learning architecture for person re-identification. CVPR, 5:25, 2015.
  2. 2.S. Bak, E. Corvee, F. Bremond, and M. Thonnat. Person re-identification using spatial covariance regions of human body parts. In Advanced Video and Signal Based Surveillance (AVSS), 2010 Seventh IEEE International Conference on, pages 435–440, 2010.
  3. 3.X. Chang, Y. Yang, E. P. Xing, and Y.-L. Yu. Complex event detection using semantic saliency and nearly-isotonic svm. In International Conference on Machine Learning (ICML), 2015.
  4. 4.D. S. Cheng, M. Cristani, M. Stoppa, L. Bazzani, and V. Murino. Custom pictorial structures for re-identification. In BMVC, volume 1, page 6, 2011.
  5. 5.J. V. Davis, B. Kulis, P. Jain, S. Sra, and I. S. Dhillon. Information-theoretic metric learning. In ICML, pages 209–216, 2007.
  6. 6.S. Ding, L. Lin, G. Wang, and H. Chao. Deep feature learning with relative distance comparison for person re-identification. Pattern Recognition, 2015.
  7. 7.P. Dollár, Z. Tu, H. Tao, and S. Belongie. Feature mining for image classification. In CVPR, pages 1–8, 2007.
  8. 8.M. Farenzena, L. Bazzani, A. Perina, V. Murino, and M. Cristani. Person re-identification by symmetry-driven accumulation of local features. In CVPR, pages 2360–2367, 2010.
  9. 9.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 32(9):1627–1645, 2010.
  10. 10.N. Gheissari, T. B. Sebastian, and R. Hartley. Person reidentification using spatiotemporal appearance. In CVPR, volume 2, pages 1528–1535, 2006.
  11. 11.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, pages 580–587, 2014.
  12. 12.A. Globerson and S. T. Roweis. Metric learning by collapsing classes. In NIPS, pages 451–458, 2005.
  13. 13.D. Gray, S. Brennan, and H. Tao. Evaluating appearance models for recognition, reacquisition, and tracking. In Proc. IEEE International Workshop on Performance Evaluation for Tracking and Surveillance (PETS), volume 3, 2007.
  14. 14.D. Gray and H. Tao. Viewpoint invariant pedestrian recognition with an ensemble of localized features. In ECCV, pages 262–275. 2008.
  15. 15.M. Guillaumin, J. Verbeek, and C. Schmid. Is that you? metric learning approaches for face identification. In CVPR, pages 498–505, 2009.
  16. 16.J. Han, D. Zhang, S. Wen, L. Guo, T. Liu, and X. Li. Two-stage learning to predict human eye fixations via sdaes. 2015.
  17. 17.M. Hirzer, C. Beleznai, P. M. Roth, and H. Bischof. Person re-identification by descriptive and discriminative classification. In Image Analysis, pages 91–102. 2011.
  18. 18.M. Hirzer, P. M. Roth, and H. Bischof. Person re-identification by efficient impostor-based metric learning. In Advanced Video and Signal-Based Surveillance (AVSS), 2012 IEEE Ninth International Conference on, pages 203–208, 2012.
  19. 19.W. Hu, M. Hu, X. Zhou, T. Tan, J. Lou, and S. Maybank. Principal axis-based correspondence between multiple cameras for people tracking. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 28(4):663–671, 2006.
  20. 20.S. Khamis, C.-H. Kuo, V. K. Singh, V. D. Shet, and L. S. Davis. Joint learning for attribute-consistent person re-identification. In ECCV, pages 134–146, 2014.
  21. 21.M. Koestinger, M. Hirzer, P. Wohlhart, P. M. Roth, and H. Bischof. Large scale metric learning from equivalence constraints. In CVPR, pages 2288–2295, 2012.
  22. 22.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, pages 1097–1105, 2012.
  23. 23.W. Li and X. Wang. Locally aligned feature transforms across views. In CVPR, pages 3594–3601, 2013.
  24. 24.W. Li, R. Zhao, and X. Wang. Human reidentification with transferred metric learning. In ACCV, pages 31–44, 2012.
  25. 25.W. Li, R. Zhao, T. Xiao, and X. Wang. Deepreid: Deep filter pairing neural network for person re-identification. In CVPR, pages 152–159, 2014.
  26. 26.Z. Li, S. Chang, F. Liang, T. S. Huang, L. Cao, and J. R. Smith. Learning locally-adaptive decision functions for person verification. In CVPR, pages 3610–3617, 2013.
  27. 27.C. Liu, S. Gong, C. C. Loy, and X. Lin. Person re-identification: What features are important? In ECCV, pages 391–401, 2012.
  28. 28.B. Ma, Y. Su, and F. Jurie. Bicov: a novel image representation for person re-identification and face verification. In BMVC, pages 11–pages, 2012.
  29. 29.L. Ma, X. Yang, and D. Tao. Person re-identification over camera networks using multi-task distance metric learning. Image Processing, IEEE Transactions on, 23(8):3656–3670, 2014.
  30. 30.A. Mignon and F. Jurie. Pcca: A new approach for distance learning from sparse pairwise constraints. In CVPR, pages 2666–2672, 2012.
  31. 31.S. Paisitkriangkrai, C. Shen, and A. v. d. Hengel. Learning to rank in person re-identification with metric ensembles. arXiv preprint arXiv:1503.01543, 2015.
  32. 32.U. Park, A. K. Jain, I. Kitahara, K. Kogure, and N. Hagita. Vise: Visual search engine using multiple networked cameras. In ICPR, volume 3, pages 1204–1207, 2006.
  33. 33.P. M. Roth, M. Hirzer, M. Köstinger, C. Beleznai, and H. Bischof. Mahalanobis distance learning for person re-identification. In Person Re-Identification, pages 247–267. 2014.
  34. 34.F. Schroff, D. Kalenichenko, and J. Philbin. Facenet: A unified embedding for face recognition and clustering. arXiv preprint arXiv:1503.03832, 2015.
  35. 35.W. R. Schwartz and L. S. Davis. Learning discriminative appearance-based models using partial least squares. In Computer Graphics and Image Processing (SIBGRAPI), 2009 XXII Brazilian Symposium on, pages 322–329, 2009.
  36. 36.Y. Sun, Y. Chen, X. Wang, and X. Tang. Deep learning face representation by joint identification-verification. In NIPS, pages 1988–1996, 2014.
  37. 37.A. Toshev and C. Szegedy. Deeppose: Human pose estimation via deep neural networks. In CVPR, pages 1653–1660, 2014.
  38. 38.UK. Home office i-lids multiple camera tracking scenario definition. In 2008.
  39. 39.J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu. Learning fine-grained image similarity with deep ranking. In CVPR, pages 1386–1393, 2014.
  40. 40.X. Wang, G. Doretto, T. Sebastian, J. Rittscher, and P. Tu. Shape and appearance context modeling. In ICCV, pages 1–8, 2007.
  41. 41.K. Q. Weinberger, J. Blitzer, and L. K. Saul. Distance metric learning for large margin nearest neighbor classification. In NIPS, pages 1473–1480, 2005.
  42. 42.E. P. Xing, M. I. Jordan, S. Russell, and A. Y. Ng. Distance metric learning with application to clustering with side-information. In NIPS, pages 505–512, 2002.
  43. 43.F. Xiong, M. Gou, O. Camps, and M. Sznaier. Person re-identification using kernel-based metric learning methods. In ECCV, pages 1–16. 2014.
  44. 44.Y. Yang, J. Yang, J. Yan, S. Liao, D. Yi, and S. Z. Li. Salient color names for person re-identification. In ECCV, pages 536–551. 2014.
  45. 45.D. Yi, Z. Lei, S. Liao, and S. Z. Li. Deep metric learning for person re-identification. In ICPR, pages 34–39, 2014.
  46. 46.R. Zhao, W. Ouyang, and X. Wang. Person re-identification by salience matching. In ICCV, pages 2528–2535, 2013.
  47. 47.R. Zhao, W. Ouyang, and X. Wang. Unsupervised salience learning for person re-identification. In CVPR, pages 3586–3593, 2013.
  48. 48.R. Zhao, W. Ouyang, and X. Wang. Learning mid-level filters for person re-identification. In CVPR, pages 144–151, 2014.
  49. 49.W.-S. Zheng, S. Gong, and T. Xiang. Person re-identification by probabilistic relative distance comparison. In CVPR, pages 649–656, 2011.

Citation

MLA
Cheng, D., et al. “Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1335–44, https://doi.org/10.1109/CVPR.2016.149.
APA
Cheng, D., Gong, Y., Zhou, S., Wang, J., & Zheng, N. (2016). Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1335–1344. https://doi.org/10.1109/CVPR.2016.149
Chicago
Cheng, D., Y. Gong, S. Zhou, J. Wang, and N. Zheng. 2016. “Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1335–44. https://doi.org/10.1109/CVPR.2016.149.
Harvard
Cheng, D. et al. (2016) “Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 1335–1344. Available at: https://doi.org/10.1109/CVPR.2016.149.
Vancouver
1. Cheng D, Gong Y, Zhou S, Wang J, Zheng N (2016) Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 1335–1344

BibTeX

@inproceedings{Cheng_2016, title={Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function}, url={http://dx.doi.org/10.1109/CVPR.2016.149}, DOI={10.1109/cvpr.2016.149}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Cheng, De and Gong, Yihong and Zhou, Sanping and Wang, Jinjun and Zheng, Nanning}, year={2016}, month=June, pages={1335–1344} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE