MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID

Jianyang GuKai WangHao LuoChen ChenWei JiangYuqiang FangShanghang ZhangYang YouJian Zhao

article2023CVPR75 citations

Proposes an open-set Neural Architecture Search framework incorporating a Twins Contrastive Mechanism and a multi-scale interaction search space to discover lightweight, highly accurate backbone architectures for object re-identification.

Listen

Object re-identification—the task of matching specific people or vehicles across disparate camera views—is crucial for intelligent surveillance, smart city management, and automated security systems. Traditionally, developers have relied on computer vision backbones designed for standard image classification. However, standard classification assumes identical categories in training and testing, whereas re-identification operates in open-set environments requiring the detection of subtle, fine-grained visual distinctions under variable angles and lighting. While automated Neural Architecture Search (NAS) offers a route to custom models, existing search schemes still replicate classification workflows, resulting in suboptimal architectures for re-identification deployments.

The article demonstrates an automated architecture search framework specifically aligned with real-world re-identification mechanics. Its core objective is to design and evaluate a compact, high-performance neural network, termed Multi-Scale Interaction Net (MSINet), capable of surpassing conventional and heavier models across standard and cross-camera domain scenarios.

To achieve this, the authors introduced a Twins Contrastive Mechanism that unbinds training and validation classes during the search phase, using dual memory banks to simulate open-set retrieval. They also developed a Multi-Scale Interaction search space that dynamically tests how low-level contours and high-level semantics communicate across network layers, alongside a Spatial Alignment Module to enforce robust visual attention across varying camera views. The approach was systematically evaluated across four benchmark datasets for person and vehicle retrieval under fully supervised, unsupervised, and cross-domain settings.

The findings confirm substantial performance and efficiency gains. First, MSINet requires only 2.3 million parameters—roughly one-tenth the size of standard ResNet50 baseline models—while running at approximately 71% of standard inference time. Second, in supervised person retrieval on the MSMT17 dataset, MSINet outperformed standard ResNet50 by approximately 9% in mean Average Precision (59.6% vs. 50.4%) and outperformed specialized re-identification networks like OSNet and CDNet. Third, in cross-domain transfer (MSMT17 to Market-1501), MSINet demonstrated a 16% improvement in mean Average Precision over ResNet50 (48.4% vs. 31.8% when enhanced with the alignment module). Finally, the compact network delivered accuracy matching or exceeding heavy Vision Transformer baselines that contain up to 86 million parameters, while avoiding their steep computational requirements.

These results demonstrate that aligning architecture search with operational conditions yields highly discriminative, lightweight models. In practical terms, this lowers memory overhead and edge-computing hardware costs, decreases retrieval latency in large video feeds, and mitigates cross-camera performance degradation caused by perspective shifts. The work also proves that specialized convolutional interactions can rival complex transformer architectures in fine-grained retrieval tasks without the accompanying processing burden.

Organizations implementing automated surveillance and retrieval pipelines should consider adopting lightweight, interaction-based backbones over generic classification backbones to optimize real-time throughput. For multi-camera networks with high visual domain shifts, integrating spatial alignment modules during training is recommended to boost consistency. Further development should test these architectures on larger video streaming pilots and edge devices to evaluate latency trade-offs prior to full-scale enterprise rollout.

Confidence in these findings is supported by consistent cross-dataset benchmarking across both person and vehicle retrieval tasks. However, users should note that the architectural search was conducted primarily on a single benchmark dataset, and extreme occlusions or environmental conditions beyond standard academic benchmarks may require further domain adaptation and tuning.

Cover for MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID

Abstract

Neural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both in-domain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methods
  • 3.1. Twins Contrastive Mechanism
  • 3.2. Multi-Scale Interaction Space
  • 3.3. Spatial Alignment Module
  • 4. Experiments
  • 4.1. Datasets and Evaluation Metrics
  • 4.2. Architecture Search
  • 4.3. Comparison with Other Backbones
  • 4.4. Comparison with State-of-the-art Methods
  • 4.5. Ablation Studies
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Twins Contrastive Mechanism (TCM) for ReID Architecture Search

    model/method

    Traditional Neural Architecture Search (NAS) frameworks designed for image classification assume identical category sets between the training and validation splits. In contrast, object re-identification (ReID) is an open-set retrieval task where training and validation identities are disjoint. To address this discrepancy, the Twins Contrastive Mechanism (TCM) eliminates the need for a shared linear classification layer during differentiable architecture search by employing two decoupled auxiliary memory banks: Ctr\mathcal{C}_{\text{tr}} for training data and Cval\mathcal{C}_{\text{val}} for validation data.

    Each memory bank is initialized with the centroid feature of each category (computed by averaging the feature vectors belonging to that identity). At each search iteration, model parameters ω\omega are updated using the training contrastive loss, formulated for an embedded feature vector f∈RDf \in \mathbb{R}^D with ground-truth category identity jj as:

    Lclstr=−log⁡exp⁡(f⋅cjtr/τ)∑n=0Nctrexp⁡(f⋅cntr/τ)L_{\text{cls}}^{\text{tr}} = - \log \frac{\exp\left(f \cdot c_j^{\text{tr}} / \tau\right)}{\sum_{n=0}^{N_c^{\text{tr}}} \exp\left(f \cdot c_n^{\text{tr}} / \tau\right)}

    where cntrc_n^{\text{tr}} is the memorized centroid feature for category nn, NctrN_c^{\text{tr}} is the total number of training categories, and τ\tau is the temperature hyperparameter set to τ=0.05\tau = 0.05.

    Following the model parameter update, the centroid feature cjtrc_j^{\text{tr}} in memory is updated via momentum averaging with a momentum coefficient β=0.2\beta = 0.2:

    cjtr←βcjtr+(1−β)fc_j^{\text{tr}} \leftarrow \beta c_j^{\text{tr}} + (1 - \beta) f

    Architecture parameters α\alpha are subsequently updated on the validation split using an identical contrastive loss evaluated against the validation memory bank Cval\mathcal{C}_{\text{val}} containing NcvalN_c^{\text{val}} categories. By decoupling Ctr\mathcal{C}_{\text{tr}} and Cval\mathcal{C}_{\text{val}}, TCM allows adjusting the category overlap ratio between splits to simulate realistic ReID deployment while stabilizing gradient updates for α\alpha.

  2. Knowl 2 — Multi-Scale Interaction (MSI) Search Space and Candidate Operations

    model/method

    The Multi-Scale Interaction (MSI) search space configures feature extraction within each network cell by splitting the input into two parallel branches with distinct receptive field scales at a fixed scale ratio ρ=3:1\rho = 3:1. Within each branch, scale-specific representations are built by stacking a 1×11\times 1 convolution and depthwise 3×33\times 3 convolutions. The branches do not share parameters except within their Interaction Modules (IM), which facilitate direct cross-branch information exchange.

    Let (x1,x2)(x_1, x_2) denote the feature maps of the two branches, where x1,x2∈RC×H×Wx_1, x_2 \in \mathbb{R}^{C \times H \times W}. The candidate operations in the IM search space O\mathcal{O} comprise four choices:

    1. None: Employs no parameters and passes the representations unmodified: (x1,x2)(x_1, x_2).
    2. Exchange: Swaps the feature representations between the two branches without introducing trainable parameters: (x2,x1)(x_2, x_1).
    3. Channel Gate: Uses a shared 2-layer Multi-Layer Perceptron (MLP) with a sigmoid activation σ\sigma to calculate channel-wise attention:

    G(x)=σ(MLP(x))G(x) = \sigma(\text{MLP}(x))

    and outputs (G(x1)⋅x1,G(x2)⋅x2)(G(x_1) \cdot x_1, G(x_2) \cdot x_2). 4. Cross Attention: Reshapes each branch feature x∈RC×H×Wx \in \mathbb{R}^{C \times H \times W} into x~∈RC×N\tilde{x} \in \mathbb{R}^{C \times N}, where N=H×WN = H \times W. Instead of computing self-attention within a single branch, the key representations of the two branches are swapped to compute cross-branch correlation matrices, which are converted into spatial-channel attention masks and added back to the original features with a learnable weighting parameter.

    In differentiable search, candidate operation outputs are combined via softmax weighting:

    f(x)=∑o∈Oexp⁡(αio)∑o′∈Oexp⁡(αio′)⋅o(x)f(x) = \sum_{o \in \mathcal{O}} \frac{\exp(\alpha_i^o)}{\sum_{o' \in \mathcal{O}} \exp(\alpha_i^{o'})} \cdot o(x)

    where αio\alpha_i^o is the architecture parameter for operation oo at layer ii. After interaction, the two branches are fused via element-wise summation.

  3. Knowl 3 — Spatial Alignment Module (SAM) and Loss Formulation

    model/method

    The Spatial Alignment Module (SAM) enforces spatial attention consistency across images exhibiting variations in camera viewpoint, pose, illumination, and occlusion. For an anchor image ii and any batch image jj with reshaped feature representations x~i,x~j∈RC×N\tilde{x}_i, \tilde{x}_j \in \mathbb{R}^{C \times N} (where N=H×WN = H \times W), a mutual position-wise correlation matrix is calculated via Mutual Conv:

    A(i,j)=x~j⊤x~iA(i, j) = \tilde{x}_j^\top \tilde{x}_i

    The maximum correlation activation vector a(i,j)∈RNa(i, j) \in \mathbb{R}^N for each spatial position of sample ii is obtained by:

    a(i,j)=max⁡dim=1A(i,j)a(i, j) = \max_{\text{dim}=1} A(i, j)

    To avoid suppressing identity-discriminative cues while reducing background noise, SAM decouples the alignment of positive sample pairs from negative sample pairs. For negative pairs, mutual activation vectors are aligned with one another. For positive pairs, activation vectors are aligned with an explicit target activation a^(i)\hat{a}(i) produced by an auxiliary Position Activation Module (PAM).

    The spatial alignment loss for sample ii is defined as:

    Lsa(i)=1N+∑p∈I+(1−S(a^(i),a(i,p)))+1N−∑n1,n2∈I−(1−S(a(i,n1),a(i,n2)))L_{\text{sa}}(i) = \frac{1}{N_+} \sum_{p \in I_+} \left(1 - S(\hat{a}(i), a(i, p))\right) + \frac{1}{N_-} \sum_{n_1, n_2 \in I_-} \left(1 - S(a(i, n_1), a(i, n_2))\right)

    where I+I_+ and I−I_- represent the indices of positive and negative samples relative to sample ii with set sizes N+N_+ and N−N_-, and S(u,v)=u⋅v∥u∥2∥v∥2S(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2} denotes cosine similarity. The total ReID objective is L=Lid+Ltri+λsaLsaL = L_{\text{id}} + L_{\text{tri}} + \lambda_{\text{sa}} L_{\text{sa}}, where λsa=2.0\lambda_{\text{sa}} = 2.0.

  4. Knowl 4 — Searched MSINet Architecture Specification

    model/method

    The architecture discovered by the Twins Contrastive Mechanism on the MSMT17 dataset is designated as MSINet. The network consists of a stem module (7×77 \times 7 convolution and 3×33 \times 3 max pooling with stride 2), followed by a sequence of 6 MSI cells interspersed with down-sample blocks (a 1×11 \times 1 convolution followed by stride-2 average pooling). Each MSI cell contains two internal interaction positions, yielding 12 searchable Interaction Module (IM) slots across the network.

    The discrete interaction operations selected at each of the 12 positions are:

    • Cell #1: Position 1 = Channel Gate (G), Position 2 = Channel Gate (G)
    • Cell #2: Position 3 = Exchange (E), Position 4 = Channel Gate (G)
    • Cell #3: Position 5 = Cross Attention (A), Position 6 = Channel Gate (G)
    • Cell #4: Position 7 = Channel Gate (G), Position 8 = None (N)
    • Cell #5: Position 9 = Channel Gate (G), Position 10 = Cross Attention (A)
    • Cell #6: Position 11 = Exchange (E), Position 12 = Cross Attention (A)

    Shallow layers favor Channel Gate operations to suppress background clutter, whereas deeper layers increasingly select Cross Attention to exchange high-level semantic information across scales. MSINet contains 2.3M parameters (2.4M when integrated with SAM).

  5. Knowl 5 — Supervised In-Domain and Cross-Domain Object ReID Benchmark Performance

    data/table

    The performance of MSINet was evaluated across four ReID benchmarks: Market-1501 (M), MSMT17 (MS), VeRi-776 (VR), and VehicleID (VID), under both training from scratch and ImageNet pre-trained initialization, as well as cross-domain direct transfer settings (MS →\to M and VR →\to VID).

    Method Params Inf. Time M MS VR MS→\toM
    R-1 mAP R-1 mAP R-1 mAP R-1 mAP
    Trained from Scratch
    ResNet50 ∼\sim24M 1.00×\times 85.7 68.3 48.0 25.7 92.8 69.9 - -
    OSNet 2.2M 0.79×\times 93.6 81.0 71.0 43.3 95.4 72.8 - -
    CDNet 1.8M 0.67×\times 93.7 83.7 73.7 48.5 94.3 73.0 - -
    MSINet 2.3M 0.71×\times 94.6 87.0 76.0 52.5 95.9 75.0 - -
    ImageNet Pre-trained
    ResNet50 ∼\sim24M 1.00×\times 94.5 85.9 75.5 50.4 94.5 73.6 58.8 31.8
    OSNet 2.2M 0.79×\times 94.8 84.9 78.7 52.9 95.5 76.4 66.6 37.5
    CDNet 1.8M 0.67×\times 95.1 86.0 78.9 54.7 - - - -
    MSINet 2.3M 0.71×\times 95.3 89.6 81.0 59.6 96.8 78.8 74.9 46.2
    MSINet-SAM 2.4M 0.71×\times 95.5 89.9 80.7 59.5 96.7 79.0 76.3 48.4

    MSINet trained from scratch (2.3M parameters) outperforms ImageNet pre-trained ResNet50 (∼\sim24M parameters) on all metrics across datasets. In direct domain transfer (MS →\to M), MSINet-SAM achieves 76.3% Rank-1 and 48.4% mAP, improving over ResNet50 by +17.5% Rank-1 and +16.6% mAP.

  6. Knowl 6 — Comparison of MSINet with State-of-the-Art ReID Backbones

    data/table

    MSINet was benchmarked against established specialized ReID architectures on Market-1501 (M) and MSMT17 (MS) under supervised settings.

    Method Market-1501 MSMT17
    Rank-1 (%) mAP (%) Rank-1 (%) mAP (%)
    PCB 93.8 81.6 68.2 40.4
    MGN 95.7 86.9 76.9 52.1
    OSNet 93.6 81.0 71.0 43.3
    IANet 94.4 83.1 75.5 46.8
    DGNet 94.8 86.0 77.2 52.3
    Auto-ReID 94.5 85.1 - -
    SAN 96.1 88.0 79.2 55.7
    CDNet 95.1 86.0 78.9 54.7
    BAT-Net 95.1 87.4 79.5 56.8
    SFT 94.1 87.5 79.0 58.3
    CTF 94.8 87.7 - -
    RGA-SC 96.1 88.4 80.3 57.5
    MSINet 95.3 89.6 81.0 59.6

    MSINet achieves 89.6% mAP on Market-1501 and 59.6% mAP on MSMT17, outperforming both manually designed multi-branch/attention architectures (e.g., MGN, RGA-SC) and existing NAS ReID approaches (Auto-ReID with 13M parameters, CDNet with 1.8M parameters).

  7. Knowl 7 — MSINet as a Backbone in Unsupervised and Domain Adaptive ReID

    empirical result

    Replacing standard ResNet50 backbones with MSINet in state-of-the-art unsupervised learning (USL) and unsupervised domain adaptation (UDA) methods yields consistent performance gains:

    • USL on Market-1501 with Group Sampling (GS): Replacing ResNet50 with MSINet shifts Rank-1 from 92.3% to 91.7% while boosting mAP from 79.2% to 81.5% (+2.3% mAP).
    • USL on Market-1501 with HDCRL: MSINet improves Rank-1 from 92.4% to 92.9% (+0.5%) and mAP from 81.7% to 82.7% (+1.0%).
    • UDA on MSMT17 →\to Market-1501 with IDM: Substituting ResNet50 with MSINet increases Rank-1 from 61.3% to 66.0% (+4.7%) and mAP from 33.5% to 37.8% (+4.3%).
  8. Knowl 8 — Performance and Efficiency Comparison Between MSINet and Vision Transformers

    data/table

    MSINet was compared against Vision Transformer baselines on MSMT17 (MS) and VeRi-776 (VR) benchmarks in terms of model parameter count, relative inference latency, Rank-1 accuracy (%), and mean Average Precision (mAP %).

    Method Params Inf. Time MS VR
    R-1 (%) mAP (%) R-1 (%) mAP (%)
    DeiT-S ∼\sim22M 0.97×\times 76.3 55.2 95.5 76.3
    DeiT-B ∼\sim86M 1.79×\times 81.9 61.4 95.9 78.4
    ViT-B ∼\sim86M 1.79×\times 81.8 61.0 96.5 78.2
    MSINet 2.3M 0.71×\times 81.0 59.6 96.8 78.8

    With 2.3M parameters and a 0.71×0.71\times relative inference latency, MSINet outperforms DeiT-S across all metrics and surpasses both DeiT-B (86M) and ViT-B (86M) on the VeRi-776 vehicle ReID benchmark (96.8% Rank-1 and 78.8% mAP vs. 96.5% Rank-1 and 78.2% mAP for ViT-B), while achieving competitive accuracy on MSMT17 with less than 3%3\% of the parameter count.

  9. Knowl 9 — Ablation Study on Search Schemes, Architecture Choices, and Hyperparameters

    empirical result

    Systematic ablation experiments demonstrate the impact of individual design choices:

    1. Search Scheme: On MSMT17, searching under standard Cross-Entropy with shared classes between train/val ("CE Overlap") yields suboptimal retrieval performance. Replacing cross-entropy with contrastive loss ("TCM Overlap") provides minor gains. Decoupling category identities between train and validation via the full TCM mechanism yields substantial improvements in final validation accuracy.
    2. Interaction Operators: Homogeneous models restricted to a single interaction type across all cells show that None and Exchange (parameter-free) perform worst, Channel Gate performs best among single operations, and Cross Attention degrades if applied uniformly across all layers. MSINet's searched heterogeneous configuration surpasses all single-operation baselines as well as random architecture configurations.
    3. Branch Scale Ratio ρ\rho: Varying the receptive field scale ratio ρ∈{1:1,2:1,3:1,4:1}\rho \in \{1:1, 2:1, 3:1, 4:1\} within MSI cells reveals that introducing multi-scale disparity significantly improves performance, with gains plateauing beyond ρ=3:1\rho = 3:1. A ratio of 3:13:1 achieves the optimal trade-off between parameter scale and retrieval performance.
    4. Fusion Operation: Fusing branch outputs via element-wise summation ("Sum") or subtraction ("Minus") achieves comparable high accuracy, whereas element-wise multiplication ("Mul") degrades performance.
    5. SAM Alignment Loss Weight λsa\lambda_{\text{sa}}: Evaluating λsa∈[0,5]\lambda_{\text{sa}} \in [0, 5] on the VR →\to VID cross-domain transfer setting identifies λsa=2.0\lambda_{\text{sa}} = 2.0 as optimal for maximizing transfer generalizability.

Coverage note — None was omitted; all key theoretical concepts, search formulations, structural definitions, and empirical benchmark comparisons from the paper have been converted into standalone knowls.

References

  1. 1.Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le. Neural optimizer search with reinforcement learning. In ICML, pages 459–468, 2017. 2
  2. 2.Han Cai, Ligeng Zhu, and Song Han. Proxylessnas: Direct neural architecture search on target task and hardware. In ICLR, 2018. 2
  3. 3.Hao Chen, Yaohui Wang, Benoit Lagadec, Antitza Dantcheva, and Francois Bremond. Joint generative and contrastive learning for unsupervised person re-identification. In CVPR, pages 2004–2013, 2021. 6
  4. 4.Tsai-Shien Chen, Chih-Ting Liu, Chih-Wei Wu, and Shao-Yi Chien. Orientation-aware vehicle re-identification with semantics-guided part attention network. In ECCV, pages 330–346, 2020. 1, 3
  5. 5.Xin Chen, Lingxi Xie, Jun Wu, and Qi Tian. Progressive differentiable architecture search: Bridging the depth gap between search and evaluation. In ICCV, pages 1294–1303, 2019. 2
  6. 6.De Cheng, Jingyu Zhou, Nannan Wang, and Xinbo Gao. Hybrid dynamic contrast and probability distillation for unsupervised person re-id. IEEE Transactions on Image Processing, 31:3334–3346, 2022. 6, 7
  7. 7.Xiangxiang Chu, Tianbao Zhou, Bo Zhang, and Jixiang Li. Fair darts: Eliminating unfair advantages in differentiable architecture search. In ECCV, pages 465–480, 2020. 2
  8. 8.Xiyang Dai, Dongdong Chen, Mengchen Liu, Yinpeng Chen, and Lu Yuan. Da-nas: Data adapted pruning for efficient neural architecture search. In ECCV, pages 584–600, 2020. 2
  9. 9.Yongxing Dai, Jun Liu, Yifan Sun, Zekun Tong, Chi Zhang, and Ling-Yu Duan. Idm: An intermediate domain module for domain adaptive person re-id. In ICCV, pages 11864–11874, 2021. 6, 7
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009. 1, 6
  11. 11.Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter. Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves. In IJCAI, 2015. 2
  12. 12.Xuanyi Dong and Yi Yang. Searching for a robust neural architecture in four gpu hours. In CVPR, pages 1761–1770, 2019. 1
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020. 8
  14. 14.Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Lars Petersson, and Mehrtash Harandi. Bilinear attention networks for person retrieval. In ICCV, pages 8030–8039, 2019. 2, 6
  15. 15.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In CVPR, pages 3146–3154, 2019. 4
  16. 16.Yixiao Ge, Dapeng Chen, and Hongsheng Li. Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification. In ICLR, 2019. 6
  17. 17.Yixiao Ge, Dapeng Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Self-paced contrastive learning with hybrid memory for domain adaptive object re-id. In NeurIPS, 2020. 1, 3, 6
  18. 18.Yiluan Guo and Ngai-Man Cheung. Efficient and deep person re-identification using multi-level similarity. In CVPR, pages 2335–2344, 2018. 2
  19. 19.Xumeng Han, Xuehui Yu, Nan Jiang, Guorong Li, Jian Zhao, Qixiang Ye, and Zhenjun Han. Group sampling for unsupervised person re-identification. arXiv preprint arXiv:2107.03024, 2021. 6, 7
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 1, 2
  21. 21.Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. Transreid: Transformer-based object re-identification. In ICCV, pages 15013–15022, 2021. 8
  22. 22.Ruibing Hou, Bingpeng Ma, Hong Chang, Xinqian Gu, Shiguang Shan, and Xilin Chen. Interaction-and-aggregation network for person re-identification. In CVPR, pages 9317–9326, 2019. 6
  23. 23.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In CVPR, pages 4700–4708, 2017. 2
  24. 24.Haoxuanye Ji, Le Wang, Sanping Zhou, Wei Tang, Nanning Zheng, and Gang Hua. Meta pairwise relationship distillation for unsupervised person re-identification. In ICCV, pages 3661–3670, 2021. 6
  25. 25.Xin Jin, Cuiling Lan, Wenjun Zeng, Zhibo Chen, and Li Zhang. Style normalization and restitution for generalizable person re-identification. In CVPR, pages 3143–3152, 2020. 6
  26. 26.Xin Jin, Cuiling Lan, Wenjun Zeng, Guoqiang Wei, and Zhibo Chen. Semantics-aligned representation learning for person re-identification. In AAAI, pages 11173–11180, 2020. 6
  27. 27.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2014. 6
  28. 28.Hanjun Li, Gaojie Wu, and Wei-Shi Zheng. Combined depth space based architecture search for person re-identification. In CVPR, pages 6729–6738, 2021. 1, 2, 3, 5, 6
  29. 29.Jianing Li and Shiliang Zhang. Joint visual and temporal consistency for unsupervised domain adaptive person re-identification. In ECCV, pages 483–499, 2020. 6
  30. 30.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deep-reid: Deep filter pairing neural network for person re-identification. In CVPR, pages 152–159, 2014. 2
  31. 31.Wei Li, Xiatian Zhu, and Shaogang Gong. Harmonious attention network for person re-identification. In CVPR, pages 2285–2294, 2018. 2
  32. 32.Yulin Li, Jianfeng He, Tianzhu Zhang, Xiang Liu, Yongdong Zhang, and Feng Wu. Diverse part discovery: Occluded person re-identification with part-aware transformer. In CVPR, pages 2898–2907, 2021. 8
  33. 33.Shengcai Liao and Ling Shao. Interpretable and generalizable person re-identification with query-adaptive convolution and temporal lifting. In ECCV, pages 456–474, 2020. 6
  34. 34.Chenxi Liu, Barret Zoph, Maxim Neumann, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, and Kevin Murphy. Progressive neural architecture search. In ECCV, pages 19–34, 2018. 2
  35. 35.Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu. Hierarchical representations for efficient architecture search. In ICLR, 2018. 2
  36. 36.Hanxiao Liu, Karen Simonyan, and Yiming Yang. Darts: Differentiable architecture search. In ICLR, 2018. 1, 2
  37. 37.Hongye Liu, Yonghong Tian, Yaowei Yang, Lu Pang, and Tiejun Huang. Deep relative distance learning: Tell the difference between similar vehicles. In CVPR, pages 2167–2175, 2016. 5
  38. 38.Xinchen Liu, Wu Liu, Huadong Ma, and Huiyuan Fu. Large-scale vehicle re-identification in urban surveillance videos. In ICME, pages 1–6, 2016. 2, 5
  39. 39.Xinchen Liu, Wu Liu, Tao Mei, and Huadong Ma. A deep learning-based approach to progressive vehicle re-identification for urban surveillance. In ECCV, pages 869–884, 2016. 1, 2, 5
  40. 40.Xinchen Liu, Wu Liu, Tao Mei, and Huadong Ma. Provid: Progressive and multimodal vehicle reidentification for large-scale urban surveillance. T-MM, 20(3):645–658, 2017. 1
  41. 41.Chuanchen Luo, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Spectral feature transformation for person re-identification. In ICCV, pages 4976–4985, 2019. 6
  42. 42.Hao Luo, Wei Jiang, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, and Jianyang Gu. A strong baseline and batch normalization neck for deep person re-identification. T-MM, 22(10):2597–2609, 2019. 1, 5
  43. 43.Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, and Tie-Yan Liu. Neural architecture optimization. In NeurIPS, 2018. 2
  44. 44.Xuelin Qian, Yanwei Fu, Yu-Gang Jiang, Tao Xiang, and Xiangyang Xue. Multi-scale deep learning architectures for person re-identification. In ICCV, pages 5399–5408, 2017. 1, 3
  45. 45.Ruijie Quan, Xuanyi Dong, Yu Wu, Linchao Zhu, and Yi Yang. Auto-reid: Searching for a part-aware convnet for person re-identification. In ICCV, pages 3750–3759, 2019. 1, 2, 6, 7
  46. 46.Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. Regularized evolution for image classifier architecture search. In AAAI, pages 4780–4789, 2019. 2
  47. 47.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, pages 4510–4520, 2018. 2
  48. 48.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In CVPR, pages 815–823, 2015. 1
  49. 49.Yantao Shen, Tong Xiao, Hongsheng Li, Shuai Yi, and Xiaogang Wang. Learning deep neural networks for vehicle re-id with visual-spatio-temporal path proposals. In ICCV, pages 1900–1909, 2017. 1
  50. 50.Liangchen Song, Cheng Wang, Lefei Zhang, Bo Du, Qian Zhang, Chang Huang, and Xinggang Wang. Unsupervised domain adaptive re-identification: Theory and practice. PR, 102:107173, 2020. 1
  51. 51.Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). In ECCV, pages 480–496, 2018. 3, 6
  52. 52.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In CVPR, pages 2818–2826, 2016. 1, 2
  53. 53.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve´ Jegou. Training data-efficient image transformers & distillation through attention. In ICML, pages 10347–10357, 2021. 8
  54. 54.Dongkai Wang and Shiliang Zhang. Unsupervised person re-identification via multi-label classification. In CVPR, pages 10981–10990, 2020. 6
  55. 55.Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, and Xi Zhou. Learning discriminative features with multiple granularities for person re-identification. In ACMMM, pages 274–282, 2018. 6
  56. 56.Yicheng Wang, Zhenzhong Chen, Feng Wu, and Gang Wang. Person re-identification with cascaded pairwise convolutions. In CVPR, pages 1470–1478, 2018. 2
  57. 57.Zhikang Wang, Lihuo He, Xiaoguang Tu, Jian Zhao, Xinbo Gao, Shengmei Shen, and Jiashi Feng. Robust video-based person re-identification by hierarchical mining. T-CSVT, 2021. 1
  58. 58.Zhongdao Wang, Luming Tang, Xihui Liu, Zhuliang Yao, Shuai Yi, Jing Shao, Junjie Yan, Shengjin Wang, Hongsheng Li, and Xiaogang Wang. Orientation invariant feature embedding and spatial temporal regularization for vehicle re-identification. In ICCV, pages 379–387, 2017. 1
  59. 59.Zhongdao Wang, Jingwei Zhang, Liang Zheng, Yixuan Liu, Yifan Sun, Yali Li, and Shengjin Wang. Cycas: Self-supervised cycle association for learning re-identifiable descriptions. In ECCV, pages 72–88, 2020. 6
  60. 60.Longhui Wei, Shiliang Zhang, Wen Gao, and Qi Tian. Person transfer gan to bridge domain gap for person re-identification. In CVPR, pages 79–88, 2018. 2, 5
  61. 61.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In ECCV, pages 3–19, 2018. 4
  62. 62.Yuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen, Guo-Jun Qi, Qi Tian, and Hongkai Xiong. Pc-darts: Partial channel connections for memory-efficient architecture search. In ICLR, 2019. 2
  63. 63.Cheng Yan, Guansong Pang, Xiao Bai, Changhong Liu, Ning Xin, Lin Gu, and Jun Zhou. Beyond triplet loss: person re-identification with fine-grained difference-aware pairwise loss. T-MM, 2021. 1
  64. 64.Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang, Ling Shao, and Steven C.H. Hoi. Deep learning for person re-identification: A survey and outlook. T-PAMI, 2021. 1
  65. 65.Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z. Li. Deep metric learning for person re-identification. In ICPR, pages 34–39, 2014. 1
  66. 66.Anguo Zhang, Yueming Gao, Yuzhen Niu, Wenxi Liu, and Yongcheng Zhou. Coarse-to-fine person re-identification with auxiliary-domain classification and second-order information bottleneck. In CVPR, pages 598–607, 2021. 6
  67. 67.Zhizheng Zhang, Cuiling Lan, Wenjun Zeng, Xin Jin, and Zhibo Chen. Relation-aware global attention for person re-identification. In CVPR, pages 3186–3195, 2020. 2, 6, 7
  68. 68.Aihua Zheng, Xianmin Lin, Jiacheng Dong, Wenzhong Wang, Jin Tang, and Bin Luo. Multi-scale attention vehicle re-identification. Neural Computing and Applications, 32(23):17489–17503, 2020. 1, 3
  69. 69.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. In ICCV, pages 1116–1124, 2015. 2, 5
  70. 70.Liang Zheng, Yi Yang, and Alexander G Hauptmann. Person re-identification: Past, present and future. arXiv preprint arXiv:1610.02984, 2016. 1
  71. 71.Liang Zheng, Yi Yang, and Qi Tian. Sift meets cnn: A decade survey of instance retrieval. T-PAMI, 40(5):1224–1244, 2017. 1
  72. 72.Meng Zheng, Srikrishna Karanam, Ziyan Wu, and Richard J Radke. Re-identification with consistent attentive siamese networks. In CVPR, pages 5735–5744, 2019. 3
  73. 73.Zhedong Zheng, Xiaodong Yang, Zhiding Yu, Liang Zheng, Yi Yang, and Jan Kautz. Joint discriminative and generative learning for person re-identification. In CVPR, pages 2138–2147, 2019. 6
  74. 74.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In AAAI, pages 13001–13008, 2020. 6
  75. 75.Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, and Tao Xiang. Omni-scale feature learning for person re-identification. In ICCV, pages 3702–3712, 2019. 1, 2, 3, 4, 5, 6
  76. 76.Sanping Zhou, Fei Wang, Zeyi Huang, and Jinjun Wang. Discriminative feature learning with consistent attention regularization for person re-identification. In ICCV, pages 8040–8049, 2019. 3
  77. 77.Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le. Learning transferable architectures for scalable image recognition. In CVPR, pages 8697–8710, 2018. 2
  78. 78.Yang Zou, Xiaodong Yang, Zhiding Yu, BVK Vijaya Kumar, and Jan Kautz. Joint disentangling and adaptation for cross-domain person re-identification. In ECCV, pages 87–104, 2020. 1

Citation

MLA
Gu, J., et al. “MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID”. arXiv, 2023, http://arxiv.org/abs/2303.07065v1.
APA
Gu, J., Wang, K., Luo, H., Chen, C., Jiang, W., Fang, Y., Zhang, S., You, Y., & Zhao, J. (2023). MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID. arXiv. http://arxiv.org/abs/2303.07065v1
Chicago
Gu, J., K. Wang, H. Luo, et al. 2023. “MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID”. arXiv. http://arxiv.org/abs/2303.07065v1.
Harvard
Gu, J. et al. (2023) “MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.07065v1.
Vancouver
1. Gu J, Wang K, Luo H, Chen C, Jiang W, Fang Y, Zhang S, You Y, Zhao J (2023) MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID. arXiv

BibTeX

@article{gu2023msinet,
  title = {MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID},
  author = {Gu, Jianyang and Wang, Kai and Luo, Hao and Chen, Chen and Jiang, Wei and Fang, Yuqiang and Zhang, Shanghang and You, Yang and Zhao, Jian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.07065v1},
  eprint = {2303.07065}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE