In Defense of the Triplet Loss for Person Re-Identification

Alexander HermansLucas BeyerBastian Leibe

article2017arXiv3,654 citations

Demonstrates that end-to-end deep metric learning with a triplet loss variant outperforms conventional surrogate classification pipelines by a wide margin, challenging the prevailing belief that triplet loss is ineffective for person re-identification.

Listen

In the rapidly advancing field of computer vision, person re-identification aims to match images of the same individual across different cameras despite variations in pose, lighting, and clothing. A widespread assumption has held that the triplet loss performs poorly for this task compared with surrogate objectives such as classification or verification losses followed by a separate metric-learning stage. The paper challenges that view by showing that a carefully designed triplet loss supports effective end-to-end metric learning and delivers substantially higher accuracy than prior published methods on three large benchmarks.

The authors set out to evaluate whether variants of the triplet loss could eliminate the need for costly offline hard-triplet mining while matching or exceeding the performance of contemporary approaches. They conducted controlled ablation studies on a held-out validation split of the MARS dataset, then measured final performance on the CUHK03, Market-1501, and MARS test sets using both a pretrained ResNet-50 and a compact network trained from scratch. The core technical change replaces traditional triplet sampling with batch construction that draws multiple images per identity; within each batch the hardest positive and negative examples are selected automatically, yielding either abatch-hardorbatch-allloss that remains fully differentiable.

The strongest configuration, batch-hard loss with a soft-margin formulation, produced the highest validation scores and was therefore used for all subsequent runs. On Market-1501 the resulting TriNet model reached 69.14 percent mean average precision and 84.92 percent rank-1 accuracy under single-query evaluation, a gain of roughly 2528 points over the previous best triplet-based result and several points above the strongest classification-plus-metric-learning baselines. Comparable improvements appeared on MARS (67.70 percent mAP) and CUHK03 (87.58 percent rank-1 on the detected setting). The smaller network trained entirely from scratch achieved 60.71 percent mAP on Market-1501, remaining competitive with most published methods while using only one-fifth the parameters of the pretrained model. Additional experiments demonstrated that performance degrades far more gracefully for the from-scratch network when input resolution is reduced, underscoring its suitability for embedded hardware.

These outcomes indicate that the triplet loss, when paired with appropriate batch construction and margin handling, directly optimizes the embedding space needed for retrieval and therefore removes the requirement for a post-hoc metric-learning step. The gains translate into higher recall at modest computational cost and open the possibility of task-specific architectures that would be difficult to obtain from pretrained classification backbones. Practitioners should therefore consider the batch-hard soft-margin loss as a default starting point for new re-identification pipelines; when memory or latency constraints are tight, training a lightweight model from scratch offers a practical alternative that retains most of the accuracy.

The principal limitations are that all results were obtained on three established datasets with standard evaluation protocols, and that the study did not explore very large-scale or long-term tracking scenarios. Within these bounds, however, the systematic comparison of loss variants and the consistent outperformance of prior art provide strong evidence that the triplet loss deserves renewed attention for person re-identification.

Cover for In Defense of the Triplet Loss for Person Re-Identification

Abstract

In the past few years, the field of computer vision has gone through a revolution fueled mainly by the advent of large datasets and the adoption of deep convolutional neural networks for end-to-end learning. The person re-identification subfield is no exception to this. Unfortunately, a prevailing belief in the community seems to be that the triplet loss is inferior to using surrogate losses (classification, verification) followed by a separate metric learning step. We show that, for models trained from scratch as well as pretrained ones, using a variant of the triplet loss to perform end-to-end deep metric learning outperforms most other published methods by a large margin.

Table of Contents

  • 1. Introduction
  • 2. Learning Metric Embeddings, the Triplet Loss, and the Importance of Mining
  • 3. Experiments
  • 3.1. Datasets
  • 3.2. Training
  • 3.3. Network Architectures
  • 3.4. Triplet Loss
  • 3.5. Performance Evaluation
  • 3.6. To Pretrain or not to Pretrain?
  • 4. Discussion
  • 5. Conclusion
  • References
  • A. Test-time Augmentation
  • B. Hard Positives, Hard Negatives and Outliers
  • C. Experiments with Distractors
  • D. Notes on Network Training
  • E. Extended Comparison Tables
  • F. LuNet's Architecture
  • G. Full t-SNE Visualization

Knowls

  1. Knowl 1 — Batch Hard Triplet Loss

    equation

    The Batch Hard triplet loss computes the triplet loss over a mini-batch XX by selecting, for each anchor sample, the hardest positive and the hardest negative sample present within that batch:

    LBH(θ;X)=i=1Pa=1K[m+maxp=1KD(fθ(xai),fθ(xpi))minj=1Pn=1KjiD(fθ(xai),fθ(xnj))]+\mathcal{L}_{BH}(\theta; X) = \sum_{i=1}^{P} \sum_{a=1}^{K} \left[ m + \max_{p=1 \dots K} D\left(f_\theta(x_a^i), f_\theta(x_p^i)\right) - \min_{\substack{j=1 \dots P \\ n=1 \dots K \\ j \neq i}} D\left(f_\theta(x_a^i), f_\theta(x_n^j)\right) \right]_+

    where:

    • PP is the number of distinct classes (person identities) sampled in the mini-batch.
    • KK is the number of distinct images sampled per identity, resulting in a total batch size of PKPK.
    • xaiRFx_a^i \in \mathbb{R}^F is the aa-th image of the ii-th person in the batch, acting as the anchor.
    • xpiRFx_p^i \in \mathbb{R}^F is the pp-th image of the ii-th person in the batch, providing candidate positive matches.
    • xnjRFx_n^j \in \mathbb{R}^F is the nn-th image of the jj-th person (jij \neq i), providing candidate negative matches.
    • fθ:RFRDf_\theta: \mathbb{R}^F \to \mathbb{R}^D is the embedding function parameterized by θ\theta.
    • D(u,v)=uv2D(u, v) = \|u - v\|_2 is the metric distance function in the embedding space RD\mathbb{R}^D.
    • m>0m > 0 is the distance margin hyperparameter.
    • [z]+=max(0,z)[z]_+ = \max(0, z) denotes the standard hinge function.

    In total, PKPK loss terms contribute to LBH(θ;X)\mathcal{L}_{BH}(\theta; X), with the batch-level hard mining functioning as moderate hard mining over the global dataset.

  2. Knowl 2 — Soft-Margin Formulation for Triplet Loss

    model/method

    The soft-margin formulation of the triplet loss replaces the hard-threshold hinge function [m+z]+=max(0,m+z)[m + z]_+ = \max(0, m + z) with a smooth softplus approximation:

    s(z)=ln(1+ez)s(z) = \ln\left(1 + e^z\right)

    Applying this smooth approximation to the Batch Hard loss yields:

    LBH,soft(θ;X)=i=1Pa=1Kln(1+exp(maxp=1KD(fθ(xai),fθ(xpi))minj=1Pn=1KjiD(fθ(xai),fθ(xnj))))\mathcal{L}_{BH,\text{soft}}(\theta; X) = \sum_{i=1}^{P} \sum_{a=1}^{K} \ln \left( 1 + \exp\left( \max_{p=1 \dots K} D\left(f_\theta(x_a^i), f_\theta(x_p^i)\right) - \min_{\substack{j=1 \dots P \\ n=1 \dots K \\ j \neq i}} D\left(f_\theta(x_a^i), f_\theta(x_n^j)\right) \right) \right)

    where fθf_\theta is the embedding function mapping input image xx to RD\mathbb{R}^D, and D(u,v)=uv2D(u, v) = \|u - v\|_2 is the Euclidean distance.

    This formulation removes the margin hyperparameter mm and avoids a hard cutoff, continuously penalizing and pulling together positive pairs while pushing negative pairs apart with exponentially decaying gradient magnitude as separation grows.

  3. Knowl 3 — PK Mini-Batch Sampling for In-Batch Triplet Mining

    algorithm

    The PKPK mini-batch sampling algorithm generates batches structured to enable in-batch hard negative and hard positive mining without requiring offline triplet pre-mining.

    Input: Dataset D={(xk,yk)}D = \{(x_k, y_k)\} with images xkx_k and identity labels yk{1,,C}y_k \in \{1, \dots, C\}, number of classes per batch PP, number of images per class KK
    Output: Mini-batch XX of size P×KP \times K
    Sample PP distinct class identities uniformly at random without replacement from {1,,C}\{1, \dots, C\}
    for each sampled class identity i{1,,P}i \in \{1, \dots, P\} do
        if class ii contains at least KK images then
            Sample KK images of class ii uniformly at random without replacement
        else
            Sample all available images of class ii and replicate images with replacement until KK images are selected
        end if
    end for
    Assemble all selected samples into mini-batch X={xaii{1,,P},a{1,,K}}X = \{x_a^i \mid i \in \{1, \dots, P\}, a \in \{1, \dots, K\}\}
    return XX

    This construction allows forming PKPK batch-hard triplets or PK(PKK)(K1)PK(PK - K)(K - 1) batch-all triplets from PKPK forward passes.

  4. Knowl 4 — TriNet Architecture and Training Protocol

    model/method

    TriNet adapts an ImageNet-pretrained ResNet-50 for deep metric embedding in person re-identification:

    1. Architecture: The final 1000-way classification layer of ResNet-50 is discarded and replaced by two fully connected layers: the first has 1024 units followed by Batch Normalization and a ReLU activation; the second projects down to 128 units, which forms the final embedding dimension D=128D = 128. Output normalization (e.g., L2L_2 unit-sphere projection) is omitted.
    2. Distance Function: Non-squared Euclidean distance D(u,v)=uv2D(u, v) = \|u - v\|_2 is used directly.
    3. Optimization: Training uses the Adam optimizer with initial parameters ϵ0=3×104\epsilon_0 = 3 \times 10^{-4}, β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, mini-batch size N=72N = 72 (P=18P = 18 identities, K=4K = 4 images per identity).
    4. Learning Rate Schedule: The learning rate ϵ(t)\epsilon(t) is held constant at ϵ0\epsilon_0 for tt0=15000t \le t_0 = 15000 iterations, then exponentially decayed according to ϵ(t)=ϵ00.001tt0t1t0\epsilon(t) = \epsilon_0 \cdot 0.001^{\frac{t - t_0}{t_1 - t_0}} until stopping at t1=25000t_1 = 25000 iterations, with β1\beta_1 reduced to 0.50.5 for tt0t \ge t_0.
    5. Input Dimensions: Training images are resized to 118(H×W)1\frac{1}{8}(H \times W) and cropped to H×WH \times W (256×128256 \times 128 for Market-1501 and MARS; 256×96256 \times 96 for CUHK03) with random horizontal flips.
  5. Knowl 5 — LuNet Architecture for Training ReID Models from Scratch

    model/method

    LuNet is a 5.00M-parameter convolutional network designed to be trained from scratch for person re-identification on 128×64128 \times 64 images:

    Type Output / Kernel Dimensions
    Conv 128×7×7×3128 \times 7 \times 7 \times 3
    Res-block 12832128128 \to 32 \to 128
    MaxPool 3×33 \times 3, stride (2×2)(2 \times 2), padding (1×1)(1 \times 1)
    Res-block 12832128128 \to 32 \to 128
    Res-block 12832128128 \to 32 \to 128
    Res-block 12864256128 \to 64 \to 256
    MaxPool 3×33 \times 3, stride (2×2)(2 \times 2), padding (1×1)(1 \times 1)
    Res-block 25664256256 \to 64 \to 256
    Res-block 25664256256 \to 64 \to 256
    MaxPool 3×33 \times 3, stride (2×2)(2 \times 2), padding (1×1)(1 \times 1)
    Res-block 25664256256 \to 64 \to 256
    Res-block 25664256256 \to 64 \to 256
    Res-block 256128512256 \to 128 \to 512
    MaxPool 3×33 \times 3, stride (2×2)(2 \times 2), padding (1×1)(1 \times 1)
    Res-block 512128512512 \to 128 \to 512
    Res-block 512128512512 \to 128 \to 512
    MaxPool 3×33 \times 3, stride (2×2)(2 \times 2), padding (1×1)(1 \times 1)
    Res-block 512×(3×3×512),128×(3×3×512)512 \times (3 \times 3 \times 512), 128 \times (3 \times 3 \times 512)
    Linear 1024×5121024 \times 512
    Batch-Norm 512512
    LeakyReLU slope 0.30.3
    Linear 512×128512 \times 128

    Key architectural design principles:

    • Follows ResNet-v2 pre-activation structure.
    • All activation functions are LeakyReLU with a negative slope of 0.30.3.
    • Downsampling is performed strictly via 3×33 \times 3 max-pooling with stride 2 and padding (1,1)(1, 1), rather than strided convolutions.
    • Standard global average pooling is omitted; instead, spatial dimensions are reduced using a channel-reducing final residual block before flattening into the linear layers.
    • Training uses mini-batches of size 128128 (P=32,K=4P = 32, K = 4) with the Adam optimizer (ϵ0=103\epsilon_0 = 10^{-3}).
  6. Knowl 6 — Convex Mean Aggregation of Metric Embeddings

    model/method

    When combining multiple feature embeddings of the same identity (e.g., across multi-crop test-time augmentations, multi-query image sets, or video tracklets), embeddings are aggregated using the arithmetic mean (a convex combination):

    fˉ=1Nn=1Nfθ(xn)\bar{f} = \frac{1}{N} \sum_{n=1}^N f_\theta(x_n)

    Because the triplet loss trains the convex hull of positive instances to exclude negative samples, a convex combination of positive embeddings cannot become closer to any negative embedding than the constituent positive points were.

    In contrast, non-convex pooling operations such as element-wise max-pooling break this geometric guarantee; for example, max-pooling tracklet embeddings on the MARS dataset degrades the mean Average Precision (mAP) by 11.411.4 percentage points compared to mean aggregation.

  7. Knowl 7 — Empirical Comparison of Triplet Loss Formulations and Mining Strategies

    data/table

    Evaluation of multiple triplet loss formulations on the MARS validation split (150 validation identities, 475 training identities) using LuNet trained from scratch without data augmentation on half-resolution images:

    margin 0.1 margin 0.2 margin 0.5 margin 1.0 soft margin
    Loss Formulation mAP rank-1 mAP rank-1 mAP rank-1 mAP rank-1 mAP rank-1
    Triplet (Ltri\mathcal{L}_{\text{tri}}) 40.80 59.23 41.71 60.78 43.51 60.87 43.61 61.63 48.40 66.37
    Triplet (Ltri\mathcal{L}_{\text{tri}}) + OHM 16.6* 36.6* 61.40 82.95 32.0* 57.1* 41.45 59.42 46.63 65.43
    Batch hard (LBH\mathcal{L}_{BH}) 65.09 83.51 65.27 84.55 65.12 83.39 63.78 82.48 65.77 84.69
    Batch hard (LBH0\mathcal{L}_{BH \neq 0}) 63.10 83.04 64.19 83.42 63.71 82.29 64.06 84.50 - -
    Batch all (LBA\mathcal{L}_{BA}) 59.43 79.24 60.48 79.99 60.30 79.52 62.08 80.55 61.04 80.65
    Batch all (LBA0\mathcal{L}_{BA \neq 0}) 63.29 83.65 64.31 83.37 64.41 83.98 64.06 82.90 - -
    Lifted 3-pos. (LLG\mathcal{L}_{LG}) 64.00 82.71 63.87 82.86 63.61 84.55 64.02 84.17 - -
    Lifted 1-pos. (LL\mathcal{L}_{L}) 61.95 81.35 63.68 81.73 63.01 82.48 62.28 82.34 - -
    • Asterisks (*) denote runs that became trapped in bad local optima during offline hard mining (OHM).
    • Averaging only non-zero loss terms (L0\mathcal{L}_{\neq 0}) prevents trivial triplets from diluting gradients in Batch All.
    • The Batch Hard formulation with a soft margin achieves the highest performance (65.77% mAP, 84.69% rank-1) without requiring offline mining.
  8. Knowl 8 — Person Re-Identification Performance on Market-1501 and MARS

    data/table

    Performance comparison on the Market-1501 (single-query SQ and multi-query MQ) and MARS datasets using mean Average Precision (mAP, %) and Cumulative Matching Characteristic (CMC, %) rank-1 and rank-5:

    Market-1501 SQ Market-1501 MQ MARS
    Method Type mAP rank-1 rank-5 mAP rank-1 rank-5 mAP rank-1 rank-5
    TriNet E 69.14 84.92 94.21 76.42 90.53 96.29 67.70 79.80 91.36
    LuNet E 60.71 81.38 92.34 69.07 87.11 95.16 60.48 75.56 89.70
    IDE (R) + ML ours I 58.06 78.50 91.18 67.48 85.45 94.12 57.42 72.42 86.21
    LOMO + Null Space E 29.87 55.43 - 46.03 71.56 - - - -
    Gated siamese CNN V 39.55 65.88 - 48.45 76.04 - - - -
    CAN E 35.9 60.3 - 47.9 72.1 - - - -
    JLML I 65.5 85.1 - 74.5 89.7 - - - -
    ResNet 50 (I+V) I+V 59.87 79.51 90.91 70.33 85.84 94.54 - - -
    DTL E 41.5 63.3 - 49.7 72.4 - - - -
    DTL I+V 65.5 83.7 - 73.8 89.6 - - - -
    APR (R, 751) I 64.67 84.29 93.20 - - - - - -
    Latent Parts (Fusion) I 57.53 80.31 - 66.70 86.79 - 56.05 71.77 86.57
    IDE (R) + ML I 49.05 73.60 - - - - 55.12 70.51 -
    spatial temporal RNN E - - - - - - 50.7 70.6 90.0
    CNN + Video I - - - - - - - 55.5 70.2
    TriNet (Re-ranked) E 81.07 86.67 93.38 87.18 91.75 95.78 77.43 81.21 90.76
    LuNet (Re-ranked) E 75.62 84.59 91.89 82.61 89.31 94.48 73.68 78.48 88.74
    IDE (R) + ML ours (Re-ra.) I 71.38 81.62 89.88 79.78 86.79 92.96 69.50 74.39 85.86
    IDE (R) + ML (Re-ra.) I 63.63 77.11 - - - - 68.45 73.94 -

    Optimization criteria types: E = Embedding directly learned via metric loss, I = Identification classification loss, V = Verification pairwise loss. All TriNet and LuNet models include test-time augmentation.

  9. Knowl 9 — Person Re-Identification Performance on CUHK03

    data/table

    Performance comparison on the CUHK03 dataset across both Labeled and Detected bounding box splits using single-shot evaluation averaged over the 20 standard train/test splits:

    Labeled Detected
    Method Type rank-1 rank-5 rank-1 rank-5
    TriNet E 89.63 99.01 87.58 98.17
    Gated siamese CNN V - - 61.8 86.7
    LOMO + Null Space E 62.55 90.05 54.70 84.75
    CAN E 77.6 95.2 69.2 88.5
    Latent Parts (Fusion) I 74.21 94.33 67.99 91.04
    Spindle Net I 88.5 97.8 - -
    JLML I 83.2 98.0 80.6 96.9
    DTL I+V 85.4 - 84.1 -
    ResNet 50 (I+V) I+V - - 83.4 97.1

    Types denote: E = Embedding model, I = Identification model, V = Verification model. TriNet outperforms all prior identification, verification, and metric learning models on both splits.

  10. Knowl 10 — Impact of Input Image Resolution on Pretrained vs From-Scratch Networks

    data/table

    Evaluation of the effect of decreasing input image resolution on Market-1501 performance for the pretrained TriNet model versus the from-scratch LuNet model:

    TriNet (Pretrained) LuNet (From Scratch)
    Input Size (H×WH \times W) mAP rank-1 rank-5 mAP rank-1 rank-5
    256×128256 \times 128 69.14 84.92 94.21 - - -
    128×64128 \times 64 62.52 79.45 91.06 60.71 81.38 92.34
    64×3264 \times 32 47.42 68.08 85.84 57.18 78.21 90.94

    When input resolution is reduced from 128×64128 \times 64 to 64×3264 \times 32, TriNet suffers a sharp performance decline (mAP drops by 15.1015.10 percentage points from 62.52%62.52\% to 47.42%47.42\%, and rank-1 drops by 11.3711.37 percentage points). In contrast, LuNet exhibits minor degradation (mAP drops by 3.533.53 percentage points from 60.71%60.71\% to 57.18%57.18\%), outperforming the pretrained network by 9.76%9.76\% mAP and 10.13%10.13\% rank-1 at 64×3264 \times 32 resolution.

  11. Knowl 11 — Effect of Test-Time Augmentation on Metric and Classification Models

    empirical result

    Test-time augmentation (TTA) using combinations of deterministic multi-cropping (44 corner crops and 11 center crop of size H×WH \times W) and horizontal flips on Market-1501:

    TriNet (Embedding) LuNet (Embedding) IDE (R) Ours (Classification)
    TTA Setting mAP rank-1 mAP rank-1 mAP rank-1
    original (scaled) 65.48 82.51 55.61 76.72 56.29 77.55
    center 66.29 84.06 57.08 77.94 56.26 77.08
    center + flip 68.32 84.47 59.31 79.78 57.80 78.71
    5 crops 67.86 84.83 59.00 79.42 56.73 77.29
    5 crops + flip 69.14 84.92 60.71 81.38 58.06 78.50

    Full test-time augmentation (5 crops + flip, averaging 10 embeddings) provides a 3%\sim 3\% boost in mAP for embedding models trained with triplet loss (TriNet increases from 66.29%66.29\% to 69.14%69.14\%; LuNet increases from 57.08%57.08\% to 60.71%60.71\%), compared to a 1.8%\sim 1.8\% mAP gain for the classification baseline IDE (R).

  12. Knowl 12 — Scalability and Robustness Against 500k Distractor Images on Market-1501

    data/table

    Evaluation of model retrieval performance when scaling the Market-1501 gallery set from the base 19,732 images up to 519,732 images using 500k distractor images:

    TriNet LuNet Res50 (I+V)
    Gallery Size mAP rank-1 mAP rank-1 mAP rank-1
    19 732 69.14 84.92 60.71 81.38 59.87 79.51
    119 732 61.93 79.69 52.73 75.65 52.28 73.78
    219 732 58.74 77.88 49.44 73.40 49.11 71.50
    319 732 56.58 76.34 47.17 71.85 - -
    419 732 54.97 75.50 45.57 70.52 - -
    519 732 53.63 74.70 44.26 69.74 45.24 68.26

    With 500k additional distractors added to the gallery, TriNet maintains a 74.70%74.70\% rank-1 retrieval rate and 53.63%53.63\% mAP, outperforming the Res50 (I+V) baseline (45.24%45.24\% mAP, 68.26%68.26\% rank-1) by 8.398.39 percentage points in mAP.

Coverage note — Qualitative t-SNE visualizations (Figures 1 and 9), individual dataset outlier query illustrations (Figures 2 and 3), and fine-grained training trajectory plots (Figures 5-8) were omitted in favor of self-contained methodological descriptions, architectural specifications, and quantitative benchmark tables.

References

  1. 1.E. Ahmed, M. Jones, and T. K. Marks. An Improved Deep Learning Architecture for Person Re-Identification. In CVPR, pages 3908–3916, 2015.
  2. 2.J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer Normalization. arXiv preprint arXiv:1607.06450, 2016.
  3. 3.S. Bai, X. Bai, and Q. Tian. Scalable person re-identification on supervised smoothed manifold. CVPR, 2017.
  4. 4.I. B. Barbosa, M. Cristani, B. Caputo, A. Rognhaugen, and T. Theoharis. Looking Beyond Appearances: Synthetic Training Data for Deep CNNs in Re-identification. arXiv preprint arXiv:1701.03153, 2017.
  5. 5.F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. J. Goodfellow, A. Bergeron, N. Bouchard, and Y. Bengio. Theano: new features and speed improvements. In NIPS W., 2012.
  6. 6.W. Chen, X. Chen, J. Zhang, and K. Huang. A Multi-task Deep Network for Person Re-identification. AAAI, 2017.
  7. 7.Y. Chen, X. Zhu, and G. Shaogang. Person Re-Identification by Deep Learning Multi-Scale Representations. In ICCV W. on Cross-domain Human Identification, 2016.
  8. 8.D. Cheng, Y. Gong, S. Zhou, J. Wang, and N. Zheng. Person Re-Identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function. In CVPR, 2016.
  9. 9.S. Ding, L. Lin, G. Wang, and H. Chao. Deep feature learning with relative distance comparison for person re-identification. Pattern Recognition, 48(10):2993–3003, 2015.
  10. 10.M. Geng, Y. Wang, T. Xiang, and Y. Tian. Deep Transfer Learning for Person Re-identification. arXiv preprint arXiv:1611.05244, 2016.
  11. 11.X. Glorot, A. Bordes, and Y. Bengio. Deep Sparse Rectifier Neural Networks. In AISTATS, 2011.
  12. 12.N. Hawes, C. Burbridge, F. Jovan, L. Kunze, B. Lacerda, L. Mudrová, J. Young, J. Wyatt, D. Hebesberger, T. Kórner, et al. The STRANDS Project: Long-Term Autonomy in Everyday Environments. RAM, 24(3):146–156, 2017.
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Delving Deep Into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. In ICCV, pages 1026–1034, 2015.
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016.
  15. 15.K. He, X. Zhang, S. Ren, and J. Sun. Identity Mappings in Deep Residual Networks. In ECCV, 2016.
  16. 16.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  17. 17.S. Khamis, C.-H. Kuo, V. K. Singh, V. D. Shet, and L. S. Davis. Joint Learning for Attribute-Consistent Person Re-Identification. In ECCV, 2014.
  18. 18.D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. In ICLR, 2015.
  19. 19.A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In NIPS, pages 1097–1105, 2012.
  20. 20.D. Li, X. Chen, Z. Zhang, and K. Huang. Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification. In CVPR, 2017.
  21. 21.W. Li, R. Zhao, T. Xiao, and X. Wang. DeepReID: Deep Filter Pairing Neural Network for Person Re-Identification. In CVPR, 2014.
  22. 22.W. Li, X. Zhu, and S. Gong. Person Re-Identification by Deep Joint Learning of Multi-Loss Classification. In IJCAI, 2017.
  23. 23.S. Liao, Y. Hu, X. Zhu, and S. Z. Li. Person Re-identification by Local Maximal Occurrence Representation and Metric Learning. In CVPR, 2015.
  24. 24.Y. Lin, L. Zheng, Z. Zheng, Y. Wu, and Y. Yang. Improving person re-identification by attribute and identity learning. arXiv preprint arXiv:1703.07220, 2017.
  25. 25.H. Liu, J. Feng, M. Qi, J. Jiang, and S. Yan. End-to-End Comparative Attention Networks for Person Re-identification. Trans. Image Proc., 26(7):3492–3506, 2017.
  26. 26.J. Liu, Z.-J. Zha, Q. Tian, D. Liu, T. Yao, Q. Ling, and T. Mei. Multi-Scale Triplet CNN for Person Re-Identification. In ACMMM, 2016.
  27. 27.A. L. Maas, A. Y. Hannun, and A. Y. Ng. Rectifier Nonlinearities Improve Neural Network Acoustic Models. In ICML, 2013.
  28. 28.S. Paisitkriangkrai, C. Shen, and A. van den Hengel. Learning to rank in person re-identification with metric ensembles. In CVPR, 2015.
  29. 29.F. Schroff, D. Kalenichenko, and J. Philbin. FaceNet: A Unified Embedding for Face Recognition and Clustering. In CVPR, 2015.
  30. 30.A. Schumann and R. Stiefelhagen. Person Re-Identification by Deep Learning Attribute-Complementary Information. In CVPR Workshops, 2017.
  31. 31.H. Shi, Y. Yang, X. Zhu, S. Liao, Z. Lei, W. Zheng, and S. Z. Li. Embedding Deep Metric for Person Re-identification: A Study Against Large Variations. In ECCV, 2016.
  32. 32.H. O. Song, Y. Xiang, S. Jegelka, and S. Savarese. Deep Metric Learning via Lifted Structured Feature Embedding. In CVPR, 2016.
  33. 33.C. Su, S. Zhang, J. Xing, W. Gao, and Q. Tian. Deep Attributes Driven Multi-camera Person Re-identification. In ECCV, 2016.
  34. 34.A. Subramaniam, M. Chatterjee, and A. Mittal. Deep Neural Networks with Inexact Matching for Person Re-Identification. In NIPS, pages 2667–2675, 2016.
  35. 35.Y. Sun, L. Zheng, W. Deng, and S. Wang. SVDNet for Pedestrian Retrieval. ICCV, 2017.
  36. 36.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  37. 37.R. Triebel, K. Arras, R. Alami, L. Beyer, S. Breuers, R. Chatila, M. Chetouani, D. Cremers, V. Evers, M. Fiore, et al. SPENCER: A Socially Aware Service Robot for Passenger Guidance and Help in Busy Airports. In Field and Service Robotics, pages 607–622. Springer, 2016.
  38. 38.L. Van Der Maaten. Accelerating t-SNE using Tree-Based Algorithms. JMLR, 15(1):3221–3245, 2014.
  39. 39.R. R. Varior, M. Haloi, and G. Wang. Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification. In ECCV, 2016.
  40. 40.F. Wang, W. Zuo, L. Lin, D. Zhang, and L. Zhang. Joint Learning of Single-image and Cross-image Representations for Person Re-identification. In CVPR, 2016.
  41. 41.K. Q. Weinberger and L. K. Saul. Distance Metric Learning for Large Margin Nearest Neighbor Classification. JMLR, 10:207–244, 2009.
  42. 42.T. Xiao, H. Li, W. Ouyang, and X. Wang. Learning Deep Feature Representations with Domain Guided Dropout for Person Re-identification. In CVPR, 2016.
  43. 43.R. Yu, Z. Zhou, S. Bai, and X. Bai. Divide and Fuse: A Re-ranking Approach for Person Re-identification. BMVC, 2017.
  44. 44.M. D. Zeiler. ADADELTA: An Adaptive Learning Rate Method. arXiv preprint arXiv:1212.5701, 2012.
  45. 45.L. Zhang, T. Xiang, and S. Gong. Learning a Discriminative Null Space for Person Re-identification. In CVPR, 2016.
  46. 46.W. Zhang, S. Hu, and K. Liu. Learning Compact Appearance Representation for Video-based Person Re-Identification. arXiv preprint arXiv:1702.06294, 2017.
  47. 47.Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu. Deep Mutual Learning. arXiv preprint arXiv:1706.00384, 2017.
  48. 48.H. Zhao, M. Tian, S. Sun, J. Shao, J. Yan, S. Yi, X. Wang, and X. Tang. Spindle Net: Person Re-identification with Human Body Region Guided Feature Decomposition and Fusion. In CVPR, 2017.
  49. 49.L. Zheng, Z. Bie, Y. Sun, J. Wang, C. Su, S. Wang, and Q. Tian. MARS: A Video Benchmark for Large-Scale Person Re-Identification. In ECCV, 2016.
  50. 50.L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian. Scalable Person Re-Identification: A Benchmark. In ICCV, 2015.
  51. 51.L. Zheng, Y. Yang, and A. G. Hauptmann. Person Re-identification: Past, Present and Future. arXiv preprint arXiv:1610.02984, 2016.
  52. 52.Z. Zheng, L. Zheng, and Y. Yang. A Discriminatively Learned CNN Embedding for Person Re-identification. arXiv preprint arXiv:1611.05666, 2016.
  53. 53.Z. Zheng, L. Zheng, and Y. Yang. Pedestrian Alignment Network for Large-scale Person Re-identification. arXiv preprint arXiv:1707.00408, 2017.
  54. 54.Z. Zheng, L. Zheng, and Y. Yang. Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in vitro. ICCV, 2017.
  55. 55.Z. Zhong, L. Zheng, D. Cao, and S. Li. Re-ranking Person Re-identification with k-reciprocal Encoding. CVPR, 2017.
  56. 56.Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang. Random Erasing Data Augmentation. arXiv preprint arXiv:1708.04896, 2017.
  57. 57.Z. Zhou, Y. Huang, W. Wang, L. Wang, and T. Tan. See the Forest for the Trees: Joint Spatial and Temporal Recurrent Neural Networks for Video-based Person Re-identification. In CVPR, 2017.

Citation

MLA
Hermans, A., et al. “In Defense of the Triplet Loss for Person Re-Identification”. arXiv, 2017, http://arxiv.org/abs/1703.07737v4.
APA
Hermans, A., Beyer, L., & Leibe, B. (2017). In Defense of the Triplet Loss for Person Re-Identification. arXiv. http://arxiv.org/abs/1703.07737v4
Chicago
Hermans, A., L. Beyer, and B. Leibe. 2017. “In Defense of the Triplet Loss for Person Re-Identification”. arXiv. http://arxiv.org/abs/1703.07737v4.
Harvard
Hermans, A., Beyer, L. and Leibe, B. (2017) “In Defense of the Triplet Loss for Person Re-Identification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1703.07737v4.
Vancouver
1. Hermans A, Beyer L, Leibe B (2017) In Defense of the Triplet Loss for Person Re-Identification. arXiv

BibTeX

@article{hermans2017defense,
  title = {In Defense of the Triplet Loss for Person Re-Identification},
  author = {Hermans, Alexander and Beyer, Lucas and Leibe, Bastian},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1703.07737v4},
  eprint = {1703.07737}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors