Person re-identification by Local Maximal Occurrence representation and metric learning

Shengcai LiaoYang HuXiangyu ZhuS. Li

article2014CVPR2,144 citations
Listen

Automated surveillance systems often struggle to track individuals across multiple camera views due to significant differences in lighting, pedestrian orientation, and low image resolution. Conventional techniques typically address this by separating visual feature extraction from statistical distance matching into disconnected steps, which frequently leads to degraded matching accuracy and high computational overhead.

The article evaluates a unified framework that combines a resilient visual representation, termed Local Maximal Occurrence (LOMO), with an integrated metric learning technique, Cross-view Quadratic Discriminant Analysis (XQDA). The objective is to demonstrate that simultaneously optimizing dimensionality reduction and cross-camera matching substantially improves identification accuracy and processing efficiency.

To evaluate this framework, the authors conducted empirical testing across four standard public benchmark datasets: VIPeR (632 image pairs), QMUL GRID (250 pedestrian pairs plus 775 distractor images), CUHK Campus (971 identities), and CUHK03 (13,164 images across 1,360 pedestrians). The LOMO approach extracts multi-scale color and texture features, using illumination correction and horizontal occurrence maximization to handle angle shifts. XQDA then projects these high-dimensional features into a compact subspace while directly learning a cross-view distance metric using closed-form statistical decomposition.

The findings show substantial improvements over existing state-of-the-art methods across all test environments. On the VIPeR benchmark, the proposed framework achieved a top-rank identification rate of 40.00%, surpassing the prior best benchmark of 37.80%. On the challenging underground GRID dataset, the approach reached 16.56% top-rank accuracy without camera-specific tuning, compared to 12.24% for competing models. Gains were especially pronounced on larger benchmarks: top-rank accuracy reached 63.21% on CUHK Campus (an absolute improvement of 28.91 percentage points over the prior 34.30% record) and 52.20% on CUHK03 (a 31.55 percentage point increase over deep-learning baselines).

Beyond accuracy, the framework demonstrates significant computational efficiency. LOMO processes an image in approximately 0.012 seconds, while XQDA model training completes in 1.86 seconds on a standard desktop computervastly faster than iterative metric learning methods that require tens to hundreds of seconds. These results indicate that high-performing person re-identification can be deployed in near real-time operational environments without requiring expensive specialized computing infrastructure or cumbersome camera-by-camera retraining.

Organizations implementing automated multi-camera tracking should adopt the joint subspace and metric learning architecture while utilizing generalized models rather than camera-pair-specific configurations, which are difficult to maintain in dynamic networks. Next steps include exploring additional local texture and color descriptors within the occurrence-maximization pipeline and extending the matching algorithm to related visual search tasks such as cross-domain face recognition. While confidence in the benchmark results is high, practitioners should note that operational accuracy will naturally be lower in real-world deployments involving severe occlusions, imperfect automated bounding boxes, or extremely crowded scenes.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Local Maximal Occurrence Feature
  • 3.1. Dealing with Illumination Variations
  • 3.2. Dealing with Viewpoint Changes
  • 4. Cross-view Quadratic Discriminant Analysis
  • 4.1. Bayesian Face and KISSME Revisit
  • 4.2. XQDA
  • 4.3. Practical Computation
  • 5. Experiments
  • 5.1. Experiments on VIPeR
  • 5.1.1 Comparison of Metric Learning Algorithms
  • 5.1.2 Comparison of Features
  • 5.1.3 Comparison to the State of the Art
  • 5.2. Experiments on QMUL GRID
  • 5.3. Experiments on CUHK Campus
  • 5.4. Experiments on CUHK03
  • 5.5. Analysis of the Proposed Method
  • 5.5.1 Role of Retinex
  • 5.5.2 Role of Local Maximal Occurrence
  • 5.5.3 Subspace Dimensions
  • 5.5.4 Running Time
  • 6. Summary and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Local Maximal Occurrence (LOMO) Feature Representation

    model/method

    The Local Maximal Occurrence (LOMO) descriptor extracts a viewpoint-invariant and illumination-robust appearance representation from pedestrian images (normalized to 128×48128 \times 48 pixels).

    1. Illumination Handling: The image is preprocessed using a multiscale center/surround Retinex transform with two Gaussian filter scales (sigma1=5\\sigma_1 = 5 and sigma2=20\\sigma_2 = 20). Intensities are dynamically stretched to [0,255][0, 255]. In addition to Retinex-corrected HSV color, Scale Invariant Local Ternary Patterns (SILTP) are employed to achieve robustness to illumination scale variations and image noise.

    2. Local Feature Extraction: A sliding subwindow of size 10×1010 \times 10 pixels with a stride of 5 pixels is moved across the image. In each subwindow, three histograms are extracted:

    • An 8×8×8=5128 \times 8 \times 8 = 512-bin joint HSV color histogram.
    • Two scales of SILTP histograms, SILTP4,30.3\text{SILTP}_{4,3}^{0.3} and SILTP4,50.3\text{SILTP}_{4,5}^{0.3}, each containing 34=813^4 = 81 bins (162 bins total).
    1. Horizontal Maximal Pooling: To achieve invariance to horizontal viewpoint changes while preserving vertical spatial layout, the maximum occurrence value of each histogram bin bb is selected across all subwindows located at the same horizontal band/row: hrow(b)=maxkrowhk(b)h_{\text{row}}(b) = \max_{k \in \text{row}} h_k(b) where hk(b)h_k(b) represents the value of bin bb in subwindow kk of that horizontal row.

    2. Multiscale Pyramid and Normalization: A three-scale pyramid is constructed by downsampling the original 128×48128 \times 48 image via two successive 2×22 \times 2 local average pooling operations, yielding images of sizes 128×48128 \times 48, 64×2464 \times 24, and 32×1232 \times 12. The numbers of horizontal subwindow groups at each level are 24, 11, and 5, totaling 40 horizontal bands. The concatenated feature vector has dimensionality: (512+162)×(24+11+5)=674×40=26,960(512 + 162) \times (24 + 11 + 5) = 674 \times 40 = 26,960 A logarithmic transform f(v)=log(1+v)f(v) = \log(1 + v) is applied to suppress large bin counts, and the HSV and SILTP feature subvectors are independently normalized to unit L2L_2 norm.

  2. Knowl 2 — Cross-view Quadratic Discriminant Analysis (XQDA) Formulation and Metric Learning

    model/method

    Cross-view Quadratic Discriminant Analysis (XQDA) jointly learns a low-dimensional projection subspace WRd×rW \in \mathbb{R}^{d \times r} (r<dr < d) and a quadratic discriminant distance metric for cross-view feature vectors xRdx \in \mathbb{R}^d and zRdz \in \mathbb{R}^d across cc identity classes.

    Pairwise difference vectors Δ=xizj\Delta = x_i - z_j are divided into intrapersonal differences ΩI\Omega_I (where identity labels match: yi=ljy_i = l_j) and extrapersonal differences ΩE\Omega_E (where yiljy_i \neq l_j). Both classes are assumed to follow zero-mean multivariate Gaussian distributions: P(ΔΩI)=1(2π)d/2ΣI1/2exp(12ΔTΣI1Δ)P(\Delta | \Omega_I) = \frac{1}{(2\pi)^{d/2}|\Sigma_I|^{1/2}} \exp\left(-\frac{1}{2} \Delta^T \Sigma_I^{-1} \Delta\right) P(ΔΩE)=1(2π)d/2ΣE1/2exp(12ΔTΣE1Δ)P(\Delta | \Omega_E) = \frac{1}{(2\pi)^{d/2}|\Sigma_E|^{1/2}} \exp\left(-\frac{1}{2} \Delta^T \Sigma_E^{-1} \Delta\right) where ΣI,ΣERd×d\Sigma_I, \Sigma_E \in \mathbb{R}^{d \times d} are the intrapersonal and extrapersonal covariance matrices.

    In the projected subspace defined by transformation matrix W=(w1,w2,,wr)Rd×rW = (w_1, w_2, \dots, w_r) \in \mathbb{R}^{d \times r}, the projected covariance matrices are ΣI=WTΣIW\Sigma_I' = W^T \Sigma_I W and ΣE=WTΣEW\Sigma_E' = W^T \Sigma_E W. The distance function between xx and zz in the subspace is: dW(x,z)=(xz)TW((WTΣIW)1(WTΣEW)1)WT(xz)d_W(x, z) = (x - z)^T W \left( (W^T \Sigma_I W)^{-1} - (W^T \Sigma_E W)^{-1} \right) W^T (x - z)

    Because both ΩI\Omega_I and ΩE\Omega_E are zero-mean, traditional Fisher linear discriminant analysis is not applicable. Instead, the projection vector ww is optimized to maximize the ratio of projected extrapersonal variance σE(w)=wTΣEw\sigma_E(w) = w^T \Sigma_E w to intrapersonal variance σI(w)=wTΣIw\sigma_I(w) = w^T \Sigma_I w, formulated as the Generalized Rayleigh Quotient: J(w)=wTΣEwwTΣIwJ(w) = \frac{w^T \Sigma_E w}{w^T \Sigma_I w} subject to wTΣIw=1w^T \Sigma_I w = 1. The optimal projection vectors w1,,wrw_1, \dots, w_r are the eigenvectors corresponding to the rr largest eigenvalues of ΣI1ΣE\Sigma_I^{-1} \Sigma_E, obtained via generalized eigenvalue decomposition.

  3. Knowl 3 — Closed-Form Covariance Computation for Cross-View Metric Learning

    theoretical result

    Given cross-view training sets X=(x1,x2,,xn)Rd×nX = (x_1, x_2, \dots, x_n) \in \mathbb{R}^{d \times n} from one camera view and Z=(z1,z2,,zm)Rd×mZ = (z_1, z_2, \dots, z_m) \in \mathbb{R}^{d \times m} from another camera view across cc classes, where class kk contains nkn_k samples in XX and mkm_k samples in ZZ, the intrapersonal covariance ΣI\Sigma_I and extrapersonal covariance ΣE\Sigma_E can be computed without explicitly generating the n×mn \times m difference pairs.

    The covariance matrices are given in closed form by: nIΣI=X~X~T+Z~Z~TSRTRSTn_I \Sigma_I = \tilde{X}\tilde{X}^T + \tilde{Z}\tilde{Z}^T - S R^T - R S^T nEΣE=mXXT+nZZTsrTrsTnIΣIn_E \Sigma_E = m X X^T + n Z Z^T - s r^T - r s^T - n_I \Sigma_I where:

    • nI=k=1cnkmkn_I = \sum_{k=1}^c n_k m_k is the total number of intrapersonal sample pairs.
    • nE=nmnIn_E = n m - n_I is the total number of extrapersonal sample pairs.
    • X~=(m1x1,,m1xn1,,mcxn)Rd×n\tilde{X} = (\sqrt{m_1} x_1, \dots, \sqrt{m_1} x_{n_1}, \dots, \sqrt{m_c} x_n) \in \mathbb{R}^{d \times n}.
    • Z~=(n1z1,,n1zm1,,nczm)Rd×m\tilde{Z} = (\sqrt{n_1} z_1, \dots, \sqrt{n_1} z_{m_1}, \dots, \sqrt{n_c} z_m) \in \mathbb{R}^{d \times m}.
    • S=(i:yi=1xi,i:yi=2xi,,i:yi=cxi)Rd×cS = \left(\sum_{i: y_i=1} x_i, \sum_{i: y_i=2} x_i, \dots, \sum_{i: y_i=c} x_i\right) \in \mathbb{R}^{d \times c} is the matrix of class-wise sums of XX.
    • R=(j:lj=1zj,j:lj=2zj,,j:lj=czj)Rd×cR = \left(\sum_{j: l_j=1} z_j, \sum_{j: l_j=2} z_j, \dots, \sum_{j: l_j=c} z_j\right) \in \mathbb{R}^{d \times c} is the matrix of class-wise sums of ZZ.
    • s=i=1nxiRds = \sum_{i=1}^n x_i \in \mathbb{R}^d and r=j=1mzjRdr = \sum_{j=1}^m z_j \in \mathbb{R}^d are the total dataset sample sum vectors.

    This formulation reduces the computational complexity of evaluating ΣI\Sigma_I and ΣE\Sigma_E from O(nmd2)O(n m d^2) to O(Nd2)O(N d^2), where N=max(n,m)N = \max(n, m).

  4. Knowl 4 — Subspace Dimension Selection and Regularization in XQDA

    model/method

    Cross-view Quadratic Discriminant Analysis (XQDA) includes two specific procedures for numerical stability and automated hyperparameter selection:

    1. Covariance Regularization: When ΣI\Sigma_I is singular or near-singular, computing ΣI1\Sigma_I^{-1} is unstable. A positive regularization parameter γ\gamma is added to the diagonal elements of ΣI\Sigma_I: ΣIΣI+γI\Sigma_I \leftarrow \Sigma_I + \gamma I When input feature vectors are L2L_2-normalized to unit length, γ=0.001\gamma = 0.001 provides robust and smooth estimation.

    2. Automated Dimension Selection Criterion: Each eigenvalue λi\lambda_i of ΣI1ΣE\Sigma_I^{-1} \Sigma_E equals the variance ratio σE(wi)/σI(wi)\sigma_E(w_i) / \sigma_I(w_i) along the corresponding eigenvector wiw_i. Eigenvalues λi1\lambda_i \le 1 correspond to projection directions where extrapersonal variation is smaller than or equal to intrapersonal variation, providing no discriminative capability. Therefore, the subspace dimensionality rr is automatically determined by retaining all eigenvectors whose eigenvalues are strictly greater than 1: r=iI(λi>1)r = \sum_i \mathbb{I}(\lambda_i > 1)

  5. Knowl 5 — Person Re-identification Performance on the CUHK03 Benchmark

    data/table

    The CUHK03 dataset contains 13,164 images across 1,360 pedestrians observed from 6 camera views. Evaluation follows the standard single-shot protocol with 20 random splits (1,160 training identities, 100 testing identities, P=100P=100) on both manually labeled bounding boxes and automated pedestrian detector bounding boxes.

    Method Labeled Rank-1 (%) Detected Rank-1 (%)
    LOMO + XQDA (Proposed) 52.20 46.25
    DeepReID 20.65 19.89
    KISSME 14.17 11.70
    LDML 13.51 10.92
    eSDC 8.76 7.68
    LMNN 7.29 6.25
    ITML 5.53 5.14
    SDALF 5.60 4.87

    LOMO + XQDA achieves 52.20% rank-1 accuracy on labeled boxes and 46.25% on detected boxes, outperforming the prior state-of-the-art DeepReID by 31.55 percentage points (labeled) and 26.36 percentage points (detected).

  6. Knowl 6 — Person Re-identification Performance on the CUHK Campus Benchmark

    empirical result

    On the CUHK Campus dataset (971 pedestrians captured in two camera views, with Camera A capturing frontal/back views and Camera B capturing side views; images normalized to 160×60160 \times 60 pixels), the standard multi-shot experimental setting partitions the dataset into 485 training identities and 486 testing identities (P=486,M=2P=486, M=2).

    LOMO + XQDA achieves a rank-1 cumulative matching accuracy of 63.21%, improving over prior state-of-the-art methods by 28.91 percentage points:

    • LOMO + XQDA: 63.21%
    • Mid-level Filter: 34.30%
    • SalMatch: 28.45%
    • GenericMetric: 20.00%
    • eSDC: 19.67%
    • ITML: 15.98%
    • LMNN: 13.45%
    • L1L_1-norm: 10.33%
    • SDALF: 9.90%
    • L2L_2-norm: 9.84%
  7. Knowl 7 — Person Re-identification Performance on the VIPeR Benchmark

    data/table

    The VIPeR benchmark consists of 632 person image pairs captured by two cameras outdoors, scaled to 128×48128 \times 48 pixels. Experiments follow the standard protocol of 10 random 50/50 splits (316 training identities, 316 testing identities, P=316P=316).

    Method Rank 1 (%) Rank 10 (%) Rank 20 (%)
    LOMO + XQDA (Proposed) 40.00 80.51 91.08
    SCNCD 37.80 81.20 90.40
    kBiCov 31.11 70.71 82.45
    LADF 30.22 78.92 90.44
    SalMatch 30.16 65.54 79.15
    Mid-level Filter 29.11 65.95 79.87
    MtMCML 28.83 75.82 88.51
    RPLM 27.00 69.00 83.00
    LDFV 26.53 70.88 84.63
    SSCDL 25.60 68.10 83.60
    ColorInv 24.21 57.09 69.65
    LF 24.18 67.12 82.00
    SDALF 19.87 49.37 65.73
    KISSME 19.60 62.20 77.00
    PCCA 19.27 64.91 80.28
    WELF6 + PRDC 16.14 50.98 65.95
    PRDC 15.66 53.86 70.09
    ELF 12.00 44.00 61.00

    LOMO + XQDA achieves 40.00% rank-1 accuracy, outperforming the prior best method (SCNCD at 37.80%) by 2.20 percentage points. Furthermore, fusing LOMO + XQDA with LADF yields a 50.32% rank-1 identification rate on VIPeR.

  8. Knowl 8 — Person Re-identification Performance on the QMUL GRID Benchmark

    data/table

    The QMUL GRID dataset consists of 250 pedestrian image pairs from 8 camera views in an underground station and 775 distractor gallery images. Evaluation is performed over 10 random trials using 125 pairs for training and 125 pairs plus 775 distractors (P=900P=900) for testing.

    Without Camera Network Information
    Method Rank 1 (%) Rank 10 (%) Rank 20 (%)
    ELF6 + L1-norm 4.40 16.24 24.80
    ELF6 + RankSVM 10.24 33.28 43.68
    ELF6 + PRDC 9.68 32.96 44.32
    ELF6 + MRank-RankSVM 12.24 36.32 46.56
    ELF6 + MRank-PRDC 11.12 35.76 46.56
    ELF6 + XQDA 10.48 38.64 52.56
    LOMO + XQDA 16.56 41.84 52.40
    With Camera Network Information
    Method Rank 1 (%) Rank 10 (%) Rank 20 (%)
    ELF6 + MtMCML 14.08 45.84 59.84
    ELF6 + XQDA 16.32 40.72 51.76
    LOMO + XQDA 18.96 52.56 62.24

    When evaluated without camera network information, LOMO + XQDA reaches 16.56% rank-1 accuracy, outperforming the previous best result (12.24% by ELF6 + MRank-RankSVM) by 4.32 percentage points. When camera network pair-specific models are trained, LOMO + XQDA achieves 18.96% rank-1 accuracy.

  9. Knowl 9 — Ablation Analysis of Retinex and Horizontal Max-Pooling in LOMO

    empirical result

    An ablation study on VIPeR (P=316P=316) isolates the effects of Retinex illumination correction and horizontal local maximal pooling under direct Cosine similarity matching and XQDA metric learning:

    • Direct Cosine Matching:

      • Full LOMO feature: 20.25% rank-1 identification rate.
      • LOMO without Retinex: 12.97% rank-1 identification rate (a drop of 7.28 percentage points).
      • LOMO without horizontal maximal occurrence pooling: 11.39% rank-1 identification rate (a drop of 8.86 percentage points).
    • XQDA Metric Learning:

      • Full LOMO + XQDA: 37.34% rank-1 identification rate.
      • LOMO without Retinex + XQDA: 34.18% rank-1 identification rate.
      • LOMO without horizontal maximal occurrence pooling + XQDA: 28.16% rank-1 identification rate (a drop of 9.18 percentage points).

    These results confirm that Retinex improves color consistency across camera viewpoints, and horizontal local maximal pooling provides critical invariance against pose and viewpoint shifts.

  10. Knowl 10 — Computational Efficiency and Metric Learning Training Times

    data/table

    Computational benchmarks on VIPeR (128×48128 \times 48 images) executed in MATLAB on an Intel Core i5-2400 @ 3.10GHz CPU:

    • LOMO Feature Extraction Speed: Extracting the 26,960-dimensional LOMO representation takes an average of 0.012 seconds per image.
    • Metric Learning Training Times (averaged over 10 random trials on VIPeR):
    Algorithm Training Time (seconds)
    XQDA 1.86
    KISSME 1.34
    RLDA 1.53
    ITML 36.78
    LMNN 265.28

    Methods with closed-form generalized eigenvalue solutions (XQDA, KISSME, RLDA) require less than 2 seconds of total training time, whereas iterative optimization techniques (ITML, LMNN) require orders of magnitude more computation.

Coverage note — No substantial contributed material was omitted from the paper.

References

  1. 1.B. Alipanahi, M. Biggs, A. Ghodsi, et al. Distance metric learning vs. fisher discriminant analysis. In International conference on Artificial intelligence, 2008. 2
  2. 2.L. Bazzani, M. Cristani, and V. Murino. Symmetry-driven accumulation of local features for human characterization and re-identification. Computer Vision and Image Understanding, 117(2):130–144, 2013. 1, 6, 7
  3. 3.D. S. Cheng, M. Cristani, M. Stoppa, L. Bazzani, and V. Murino. Custom pictorial structures for re-identification. In BMVC, volume 2, page 6, 2011. 2
  4. 4.J. V. Davis, B. Kulis, P. Jain, S. Sra, and I. S. Dhillon. Information-theoretic metric learning. In Proceedings of the 24th international conference on Machine learning, pages 209–216. ACM, 2007. 1, 2, 5, 7
  5. 5.M. Dikmen, E. Akbas, T. S. Huang, and N. Ahuja. Pedestrian recognition with a learned metric. In Computer Vision–ACCV 2010, pages 501–512. Springer, 2011. 1, 2
  6. 6.M. Farenzena, L. Bazzani, A. Perina, V. Murino, and M. Cristani. Person re-identification by symmetry-driven accumulation of local features. In CVPR, pages 2360–2367, 2010. 2
  7. 7.N. Gheissari, T. B. Sebastian, and R. Hartley. Person reidentification using spatiotemporal appearance. In CVPR (2), pages 1528–1535, 2006. 2
  8. 8.S. Gong, M. Cristani, S. Yan, and C. C. Loy. Person Re-Identification. Springer, 2014. 1
  9. 9.D. Gray, S. Brennan, and H. Tao. Evaluating appearance models for recognition, reacquisition, and tracking. In IEEE International workshop on performance evaluation of tracking and surveillance, 2007. 2, 5, 6
  10. 10.D. Gray and H. Tao. Viewpoint invariant pedestrian recognition with an ensemble of localized features. In European Conference on Computer Vision, 2008. 1, 2, 4, 5, 6
  11. 11.M. Guillaumin, J. Verbeek, and C. Schmid. Is that you? metric learning approaches for face identification. In International Conference on Computer Vision, 2009. 2, 7
  12. 12.O. Hamdoun, F. Moutarde, B. Stanciulescu, and B. Steux. Person re-identification in multi-camera system by signature based on interest point descriptors collected on short video sequences. In ICDSC, pages 1–6, 2008. 2
  13. 13.T. Hastie, R. Tibshirani, and J. Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, Second Edition. Springer, 2009. 2
  14. 14.M. Hirzer, P. M. Roth, M. Köstinger, and H. Bischof. Relaxed pairwise learned metric for person re-identification. In European Conference on Computer Vision. 2012. 1, 2, 6
  15. 15.Y. Hu, S. Liao, Z. Lei, D. Yi, and S. Z. Li. Exploring structural information and fusing multiple features for person reidentification. 2
  16. 16.D. J. Jobson, Z.-U. Rahman, and G. A. Woodell. A multiscale retinex for bridging the gap between color images and the human observation of scenes. Image Processing, IEEE Transactions on, 6(7):965–976, 1997. 2
  17. 17.D. J. Jobson, Z.-U. Rahman, and G. A. Woodell. Properties and performance of a center/surround retinex. Image Processing, IEEE Transactions on, 6(3):451–462, 1997. 2
  18. 18.M. Köstinger, M. Hirzer, P. Wohlhart, P. M. Roth, and H. Bischof. Large scale metric learning from equivalence constraints. In IEEE Conference on Computer Vision and Pattern Recognition, 2012. 1, 2, 3, 4, 5, 6, 7
  19. 19.I. Kviatkovsky, A. Adam, and E. Rivlin. Color invariants for person reidentification. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 35(7):1622–1634, 2013. 6
  20. 20.E. H. Land and J. McCann. Lightness and retinex theory. JOSA, 61(1):1–11, 1971. 2
  21. 21.W. Li and X. Wang. Locally aligned feature transforms across views. In IEEE Conference on Computer Vision and Pattern Recognition, 2013. 2
  22. 22.W. Li, R. Zhao, and X. Wang. Human reidentification with transferred metric learning. In Assian Conference on Computer Vision, 2012. 7
  23. 23.W. Li, R. Zhao, T. Xiao, and X. Wang. DeepReID: Deep filter pairing neural network for person re-identification. In IEEE Conference on Computer Vision and Pattern Recognition, 2014. 7
  24. 24.Z. Li, S. Chang, F. Liang, T. S. Huang, L. Cao, and J. R. Smith. Learning locally-adaptive decision functions for person verification. In IEEE Conference on Computer Vision and Pattern Recognition, 2013. 1, 2, 5, 6
  25. 25.S. Liao, D. Yi, Z. Lei, R. Qin, and S. Z. Li. Heterogeneous face recognition from local structures of normalized appearance. In International Conference on Biometrics, 2009. 4
  26. 26.S. Liao, G. Zhao, V. Kellokumpu, M. Pietikäinen, and S. Z. Li. Modeling pixel process with scale invariant local patterns for background subtraction in complex scenes. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA, June 2010. 3
  27. 27.C. Liu, S. Gong, C. C. Loy, and X. Lin. Person re-identification: what features are important? In Computer Vision–ECCV 2012. Workshops and Demonstrations, pages 391–401. Springer, 2012. 3, 6
  28. 28.X. Liu, M. Song, D. Tao, X. Zhou, C. Chen, and J. Bu. Semi-supervised coupled dictionary learning for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014. 6
  29. 29.Y. Liu, Y. Shao, and F. Sun. Person re-identification based on visual saliency. In Intelligent Systems Design and Applications (ISDA), 2012 12th International Conference on, pages 884–889. IEEE, 2012. 2
  30. 30.C. C. Loy, C. Liu, and S. Gong. Person re-identification by manifold ranking. In IEEE International Conference on Image Processing, volume 20, 2013. 6
  31. 31.C. C. Loy, T. Xiang, and S. Gong. Multi-camera activity correlation analysis. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pages 1988–1995. IEEE, 2009. 6
  32. 32.B. Ma, Y. Su, and F. Jurie. Local descriptors encoded by fisher vectors for person re-identification. In European Conference on Computer Vision Workshops, 2012. 1, 2, 6
  33. 33.B. Ma, Y. Su, and F. Jurie. Covariance descriptor based on bio-inspired features for person re-identification and face verification. Image and Vision Computing, 32(6):379–390, 2014. 1, 5, 6
  34. 34.L. Ma, X. Yang, and D. Tao. Person re-identification over camera networks using multi-task distance metric learning. IEEE Transactions on Image Processing, 2014. 6, 7
  35. 35.A. Mignon and F. Jurie. Pcca: A new approach for distance learning from sparse pairwise constraints. In IEEE Conference on Computer Vision and Pattern Recognition, 2012. 6
  36. 36.B. Moghaddam, T. Jebara, and A. Pentland. Bayesian face recognition. Pattern Recognition, 33(11):1771–1782, 2000. 2, 3, 4
  37. 37.T. Ojala, M. Pietikäinen, and D. Harwood. “A comparative study of texture measures with classification based on feature distributions”. Pattern Recognition, 29(1):51–59, January 1996. 3
  38. 38.S. Pedagadi, J. Orwell, S. Velastin, and B. Boghossian. Local fisher discriminant analysis for pedestrian re-identification. In IEEE Conference on Computer Vision and Pattern Recognition, 2013. 1, 2, 6
  39. 39.B. Prosser, W.-S. Zheng, S. Gong, T. Xiang, and Q. Mary. Person re-identification by support vector ranking. In BMVC, 2010. 2, 3, 5, 6
  40. 40.R. Vezzani, D. Baltieri, and R. Cucchiara. People reidentification in surveillance and forensics: A survey. ACM Computing Surveys (CSUR), 46(2):29, 2013. 1
  41. 41.X. Wang, G. Doretto, T. Sebastian, J. Rittscher, and P. H. Tu. Shape and appearance context modeling. In ICCV, pages 1–8, 2007. 2
  42. 42.X. Wang, M. Yang, S. Zhu, and Y. Lin. Regionlets for generic object detection. In IEEE International Conference on Computer Vision, 2013. 2
  43. 43.K. Weinberger, J. Blitzer, and L. Saul. Distance metric learning for large margin nearest neighbor classification. Advances in neural information processing systems, 2006. 1, 2, 5, 7
  44. 44.Y. Yang, J. Yang, J. Yan, S. Liao, D. Yi, and S. Z. Li. Salient color names for person re-identification. In Proceedings of the European Conference on Computer Vision, 2014. 5, 6
  45. 45.J. Ye, T. Xiong, Q. Li, R. Janardan, J. Bi, V. Cherkassky, and C. Kambhamettu. ‘‘Efficient model selection for regularized linear discriminant analysis’’. In Proceedings of the ACM Conference on Information and Knowledge Management, pages 532–539, 2006. 5
  46. 46.R. Zhao, W. Ouyang, and X. Wang. Person re-identification by salience matching. In International Conference on Computer Vision, 2013. 1, 2, 5, 6
  47. 47.R. Zhao, W. Ouyang, and X. Wang. Unsupervised salience learning for person re-identification. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pages 3586–3593. IEEE, 2013. 2, 7
  48. 48.R. Zhao, W. Ouyang, and X. Wang. Learning mid-level filters for person re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014. 1, 6, 7
  49. 49.W.-S. Zheng, S. Gong, and T. Xiang. Person re-identification by probabilistic relative distance comparison. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pages 649–656. IEEE, 2011. 1, 2, 3, 5, 6
  50. 50.W.-S. Zheng, S. Gong, and T. Xiang. Reidentification by relative distance comparison. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 35(3):653–668, 2013. 5, 6

Citation

MLA
Liao, S., et al. “Person Re-identification by Local Maximal Occurrence Representation and Metric Learning”. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 2197–206, https://doi.org/10.1109/CVPR.2015.7298832.
APA
Liao, S., Hu, Y., Xiangyu Zhu, & Li, S. Z. (2015). Person re-identification by Local Maximal Occurrence representation and metric learning. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2197–2206. https://doi.org/10.1109/CVPR.2015.7298832
Chicago
Liao, S., Y. Hu, Xiangyu Zhu, and S. Z. Li. 2015. “Person Re-identification by Local Maximal Occurrence Representation and Metric Learning”. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2197–2206. https://doi.org/10.1109/CVPR.2015.7298832.
Harvard
Liao, S. et al. (2015) “Person re-identification by Local Maximal Occurrence representation and metric learning”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2197–2206. Available at: https://doi.org/10.1109/CVPR.2015.7298832.
Vancouver
1. Liao S, Hu Y, Xiangyu Zhu, Li SZ (2015) Person re-identification by Local Maximal Occurrence representation and metric learning. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2197–2206

BibTeX

@inproceedings{Liao_2015, title={Person re-identification by Local Maximal Occurrence representation and metric learning}, url={http://dx.doi.org/10.1109/CVPR.2015.7298832}, DOI={10.1109/cvpr.2015.7298832}, booktitle={2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Liao, Shengcai and Hu, Yang and Xiangyu Zhu and Li, Stan Z.}, year={2015}, month=June, pages={2197–2206} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE