Built independently by an author, for readers. Read the story and support ChapterPal

keyword

person re-identification

Person re-identification is a computer vision and video surveillance task that involves recognizing and matching an individual across multiple camera views, locations, or time intervals, typically where camera fields of view do not overlap. Operating primarily as an image or video retrieval process, the system compares a query image of a target individual against a gallery of candidates to find matches belonging to the same identity. To perform this association accurately, person re-identification systems utilize feature extraction and metric learning techniques to capture invariant visual representations of full-body appearance and localized body parts, allowing the models to overcome real-world challenges such as changes in illumination, viewing perspectives, human poses, occlusions, and camera resolutions.

15 items

Large-Scale Pre-training for Person Re-identification with Noisy Labels

Large-Scale Pre-training for Person Re-identification with Noisy Labels

Dengpan Fu, Dongdong Chen, Hao Yang, Jianmin Bao, Lu Yuan, Lei Zhang, Houqiang Li, Fang Wen, Dong Chen

OrganizationsInternational Digital Economy AcademyMicrosoftUniversity of Science and Technology of China

Why you should read this

Introduces a scalable pre-training framework that learns transferable person re-identification representations directly from uncurated video tracklets by combining prototype-based label rectification with label-guided contrastive learning on a ten-million-image noisy dataset.

This paper aims to address the problem of pre-training for person re-identification (Re-ID) with noisy labels. To setup the pre-training task, we apply a simple online multi-object tracking system on raw videos of an existing un-labeled Re-ID dataset “LUPerson” and build the Noisy Labeled variant called “LUPerson-NL”. Since theses ID labels automatically derived from tracklets inevitably contain noises, we develop a large-scale Pre-training frame-work utilizing Noisy Labels (PNL), which consists of three learning modules: supervised Re-ID learning, prototype-based contrastive learning, and label-guided contrastive learning. In principle, joint learning of these three mod-ules not only clusters similar examples to one prototype, but also rectifies noisy labels based on the prototype as-signment. We demonstrate that learning directly from raw videos is a promising alternative for pre-training, which utilizes spatial and temporal correlations as weak super-vision. This simple pre-training task provides a scalable way to learn SOTA Re-ID representations from scratch on “LUPerson-NL” without bells and whistles. For example, by applying on the same supervised Re-ID method MGN, our pre-trained model improves the mAP over the unsu-pervised pre-training counterpart by 5.7%, 2.2%, 2.3% on CUHK03, DukeMTMC, and MSMT17 respectively. Under the small-scale or few-shot setting, the performance gain is even more significant, suggesting a better transferability of the learned representation. Code is available at https://github.com/DengpanFu/LUPerson-NL.

Added

2026-09-26

MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID

MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID

Jianyang Gu, Kai Wang, Hao Luo, Chen Chen, Wei Jiang, Yuqiang Fang, Shanghang Zhang, Yang You, Jian Zhao

OrganizationsAlibaba GroupInstitute of North Electronic EquipmentNational University of SingaporeOPPOPeking UniversitySpace Engineering UniversityZhejiang University

Why you should read this

Proposes an open-set Neural Architecture Search framework incorporating a Twins Contrastive Mechanism and a multi-scale interaction search space to discover lightweight, highly accurate backbone architectures for object re-identification.

Neural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both in-domain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet.

Added

2026-09-26

UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

UTM: A Unified Multiple Object Tracking Model with Identity-Aware Feature Enhancement

Sisi You, Hantao Yao, Bing-Kun Bao, Changsheng Xu

OrganizationsInstitute of Automation, Chinese Academy of SciencesNanjing University of Posts and TelecommunicationsUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a unified multiple object tracking framework that couples detection, feature embedding, and identity association through an identity-aware feature enhancement module, creating a mutual feedback loop that improves both object localization and association accuracy across frames.

Recently, Multiple Object Tracking has achieved great success, which consists of object detection, feature embedding, and identity association. Existing methods apply the three-step or two-step paradigm to generate robust trajectories, where identity association is independent of other components. However, the independent identity association results in the identity-aware knowledge contained in the tracklet not be used to boost the detection and embedding modules. To overcome the limitations of existing methods, we introduce a novel Unified Tracking Model (UTM) to bridge those three components for generating a positive feedback loop with mutual benefits. The key insight of UTM is the Identity-Aware Feature Enhancement (IAFE), which is applied to bridge and benefit these three components by utilizing the identity-aware knowledge to boost detection and embedding. Formally, IAFE contains the Identity-Aware Boosting Attention (IABA) and the Identity-Aware Erasing Attention (IAEA), where IABA enhances the consistent regions between the current frame feature and identity-aware knowledge, and IAEA suppresses the distracted regions in the current frame feature. With better detections and embeddings, higher-quality tracklets can also be generated. Extensive experiments of public and private detections on three benchmarks demonstrate the robustness of UTM.

Added

2026-09-26

Noisy Correspondence Learning with Meta Similarity Correction

Noisy Correspondence Learning with Meta Similarity Correction

Haochen Han, Kaiyao Miao, Qinghua Zheng, Minnan Luo

OrganizationsXi'an Jiaotong University

Why you should read this

Proposes a meta-learning framework that trains a correction network on clean and mismatched meta-data to rectify similarity scores and filter out mismatched cross-modal pairs during retrieval training.

Despite the success of multimodal learning in cross-modal retrieval task, the remarkable progress relies on the correct correspondence among multimedia data. However, collecting such ideal data is expensive and time-consuming. In practice, most widely used datasets are harvested from the Internet and inevitably contain mismatched pairs. Training on such noisy correspondence datasets causes performance degradation because the cross-modal retrieval methods can wrongly enforce the mismatched data to be similar. To tackle this problem, we propose a Meta Similarity Correction Network (MSCN) to provide reliable similarity scores. We view a binary classification task as the meta-process that encourages the MSCN to learn discrimination from positive and negative meta-data. To further alleviate the influence of noise, we design an effective data purification strategy using meta-data as prior knowledge to remove the noisy samples. Extensive experiments are conducted to demonstrate the strengths of our method in both synthetic and real-world noises, including Flickr30K, MS-COCO, and Conceptual Captions. Our code is publicly available.

Added

2026-09-26

One More Check: Making "Fake Background" Be Tracked Again

One More Check: Making "Fake Background" Be Tracked Again

Chao Liang, Zhipeng Zhang, Xue Zhou, Bing Li, Weiming Hu

OrganizationsInstitute of Automation, Chinese Academy of SciencesIntelligent Terminal Key Laboratory of Sichuan ProvinceUniversity of Electronic Science and Technology of China

Why you should read this

Proposes a plug-and-play re-check network for one-shot multi-object trackers that reuses identity embeddings for temporal motion forecasting to recover targets mistakenly missed as background by single-frame detectors.

The one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be unreliable when facing disastrous visual degradation, e.g., motion blur, occlusions. Once a target bounding box is mistakenly classified as background by the detector, the temporal consistency of its corresponding tracklet will be no longer maintained. In this paper, we set out to restore the bounding boxes misclassified as "fake background" by proposing a re-check network. The re-check network innovatively expands the role of ID embedding from data association to motion forecasting by effectively propagating previous tracklets to the current frame with a small overhead. Note that the propagation results are yielded by an independent and efficient embedding search, preventing the model from over-relying on detection results. Eventually, it helps to reload the "fake background" and repair the broken tracklets. Building on a strong baseline CSTrack, we construct a new one-shot tracker and achieve favorable gains by 70.7 → 76.4, 70.6 → 76.3 MOTA on MOT16 and MOT17, respectively. It also reaches a new state-of-the-art MOTA and IDF1 performance. Code is released at https://github.com/JudasDie/SOTS.

Added

2026-09-26

Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function

Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function

De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, N. Zheng

OrganizationsInstitute of Artificial Intelligence and RoboticsXi'an Jiaotong University

Why you should read this

Presents a multi-channel convolutional architecture paired with an improved triplet loss that enforces an upper bound on positive pair distances, jointly learning global and body-part representations to improve person re-identification across non-overlapping camera views.

Person re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets.

Added

2026-09-25

Harmonious Attention Network for Person Re-identification

Harmonious Attention Network for Person Re-identification

Wei Li, Xiatian Zhu, Shaogang Gong

OrganizationsQueen Mary University of LondonVision Semantics Ltd.

Why you should read this

Proposes a Harmonious Attention convolutional neural network that jointly optimizes soft pixel-level and hard regional attention alongside feature representations to accurately re-identify poorly aligned person images with large pose variations.

Existing person re-identification (re-id) methods either assume the availability of well-aligned person bounding box images as model input or rely on constrained attention selection mechanisms to calibrate misaligned images. They are therefore sub-optimal for re-id matching in arbitrarily aligned person images potentially with large human pose variations and unconstrained auto-detection errors. In this work, we show the advantages of jointly learning attention selection and feature representation in a Convolutional Neural Network (CNN) by maximising the complementary information of different levels of visual attention subject to re-id discriminative learning constraints. Specifically, we formulate a novel Harmonious Attention CNN (HA-CNN) model for joint learning of soft pixel attention and hard regional attention along with simultaneous optimisation of feature representations, dedicated to optimise person re-id in uncontrolled (misaligned) images. Extensive comparative evaluations validate the superiority of this new HA-CNN model for person re-id over a wide variety of state-of-the-art methods on three large-scale benchmarks including CUHK03, Market-1501, and DukeMTMC-ReID.

Added

2026-09-25

Learning Discriminative Features with Multiple Granularities for Person Re-Identification

Learning Discriminative Features with Multiple Granularities for Person Re-Identification

Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, Xi Zhou

OrganizationsCloudWalk TechnologyShanghai Jiao Tong University

Why you should read this

Introduces the Multiple Granularity Network (MGN), a multi-branch deep architecture that combines global features with multi-level stripe partitions to surpass semantic part-based methods across standard person re-identification benchmarks.

The combination of global and partial features has been an essential solution to improve discriminative performances in person re-identification (Re-ID) tasks. Previous part-based methods mainly focus on locating regions with specific pre-defined semantics to learn local representations, which increases learning difficulty but not efficient or robust to scenarios with large variances. In this paper, we propose an end-to-end feature learning strategy integrating discriminative information with various granularities. We carefully design the Multiple Granularity Network (MGN), a multi-branch deep network architecture consisting of one branch for global feature representations and two branches for local feature representations. Instead of learning on semantic regions, we uniformly partition the images into several stripes, and vary the number of parts in different local branches to obtain local feature representations with multiple granularities. Comprehensive experiments implemented on the mainstream evaluation datasets including Market-1501, DukeMTMC-reid and CUHK03 indicate that our method has robustly achieved state-of-the-art performances and outperformed any existing approaches by a large margin. For example, on Market-1501 dataset in single query mode, we achieve a state-of-the-art result of Rank-1/mAP=96.6%/94.2% after re-ranking.

Added

2026-09-25

Re-ranking Person Re-identification with k-Reciprocal Encoding

Re-ranking Person Re-identification with k-Reciprocal Encoding

Zhun Zhong, Liang Zheng, Donglin Cao, Shaozi Li

OrganizationsFujian Key Laboratory of Brain-inspired Computing Technique and ApplicationsUniversity of Technology SydneyXiamen University

Why you should read this

Proposes an unsupervised re-ranking method that improves person re-identification accuracy across large-scale datasets by encoding k-reciprocal nearest neighbors through Jaccard distance without requiring manual intervention or labeled data.

When considering person re-identification (re-ID) as a retrieval process, re-ranking is a critical step to improve its accuracy. Yet in the re-ID community, limited effort has been devoted to re-ranking, especially those fully automatic, unsupervised solutions. In this paper, we propose a k-reciprocal encoding method to re-rank the re-ID results. Our hypothesis is that if a gallery image is similar to the probe in the k-reciprocal nearest neighbors, it is more likely to be a true match. Specifically, given an image, a k-reciprocal feature is calculated by encoding its k-reciprocal nearest neighbors into a single vector, which is used for re-ranking under the Jaccard distance. The final distance is computed as the combination of the original distance and the Jaccard distance. Our re-ranking method does not require any human interaction or any labeled data, so it is applicable to large-scale datasets. Experiments on the large-scale Market-1501, CUHK03, MARS, and PRW datasets confirm the effectiveness of our method.

Added

2026-09-24

Person Transfer GAN to Bridge Domain Gap for Person Re-identification

Person Transfer GAN to Bridge Domain Gap for Person Re-identification

Longhui Wei, Shiliang Zhang, Wen Gao, Qi Tian

OrganizationsPeking UniversityUniversity of Texas at San Antonio

Why you should read this

Introduces the large-scale MSMT17 benchmark and a Person Transfer GAN that bridges cross-dataset domain gaps by translating labeled person images into new camera environments without requiring target-domain annotations.

Although the performance of person Re-Identification (ReID) has been significantly boosted, many challenging issues in real scenarios have not been fully investigated, e.g., the complex scenes and lighting variations, viewpoint and pose changes, and the large number of identities in a camera network. To facilitate the research towards conquering those issues, this paper contributes a new dataset called MSMT17 with many important features, e.g., 1) the raw videos are taken by an 15-camera network deployed in both indoor and outdoor scenes, 2) the videos cover a long period of time and present complex lighting variations, and 3) it contains currently the largest number of annotated identities, i.e., 4,101 identities and 126,441 bounding boxes. We also observe that, domain gap commonly exists between datasets, which essentially causes severe performance drop when training and testing on different datasets. This results in that available training data cannot be effectively leveraged for new testing domains. To relieve the expensive costs of annotating new training samples, we propose a Person Transfer Generative Adversarial Network (PTGAN) to bridge the domain gap. Comprehensive experiments show that the domain gap could be substantially narrowed-down by the PTGAN.

Added

2026-09-17

Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in Vitro

Unlabeled Samples Generated by GAN Improve the Person Re-identification Baseline in Vitro

Zhedong Zheng, Liang Zheng, Yi Yang

OrganizationsUniversity of Technology Sydney

Why you should read this

Introduces label smoothing regularization for outliers to incorporate GAN-generated unlabeled images into supervised training, consistently improving person re-identification baseline accuracy without requiring additional data collection.

The main contribution of this paper is a simple semi-supervised pipeline that only uses the original training set without collecting extra data. It is challenging in 1) how to obtain more training data only from the training set and 2) how to use the newly generated data. In this work, the generative adversarial network (GAN) is used to generate unlabeled samples. We propose the label smoothing regularization for outliers (LSRO). This method assigns a uniform label distribution to the unlabeled images, which regularizes the supervised model and improves the baseline. We verify the proposed method on a practical problem: person re-identification (re-ID). This task aims to retrieve a query person from other cameras. We adopt the deep convolutional generative adversarial network (DCGAN) for sample generation, and a baseline convolutional neural network (CNN) for representation learning. Experiments show that adding the GAN-generated data effectively improves the discriminative ability of learned CNN embeddings. On three large-scale datasets, Market-1501, CUHK03 and DukeMTMC-reID, we obtain +4.37%, +1.6% and +2.46% improvement in rank-1 precision over the baseline CNN, respectively. We additionally apply the proposed method to fine-grained bird recognition and achieve a +0.6% improvement over a strong baseline. The code is available at this https URL.

Added

2026-09-17

Person re-identification by Local Maximal Occurrence representation and metric learning

Person re-identification by Local Maximal Occurrence representation and metric learning

Shengcai Liao, Yang Hu, Xiangyu Zhu, S. Li

OrganizationsInstitute of Automation, Chinese Academy of Sciences

Why you should read this

Proposes the Local Maximal Occurrence feature representation and Cross-view Quadratic Discriminant Analysis metric learning framework to address severe viewpoint and illumination variations in surveillance video, achieving massive rank-1 accuracy gains across four major person re-identification benchmarks.

Person re-identification is an important technique towards automatic search of a person's presence in a surveillance video. Two fundamental problems are critical for person re-identification, feature representation and metric learning. An effective feature representation should be robust to illumination and viewpoint changes, and a discriminant metric should be learned to match various person images. In this paper, we propose an effective feature representation called Local Maximal Occurrence (LOMO), and a subspace and metric learning method called Cross-view Quadratic Discriminant Analysis (XQDA). The LOMO feature analyzes the horizontal occurrence of local features, and maximizes the occurrence to make a stable representation against viewpoint changes. Besides, to handle illumination variations, we apply the Retinex transform and a scale invariant texture operator. To learn a discriminant metric, we propose to learn a discriminant low dimensional subspace by cross-view quadratic discriminant analysis, and simultaneously, a QDA metric is learned on the derived subspace. We also present a practical computation method for XQDA, as well as its regularization. Experiments on four challenging person re-identification databases, VIPeR, QMUL GRID, CUHK Campus, and CUHK03, show that the proposed method improves the state-of-the-art rank-1 identification rates by 2.2%, 4.88%, 28.91%, and 31.55% on the four databases, respectively.

Added

2026-09-16

Deep Learning for Person Re-Identification: A Survey and Outlook

Deep Learning for Person Re-Identification: A Survey and Outlook

Mang Ye, Jianbing Shen, Gaojie Lin, Tao Xiang, Ling Shao, Steven C. H. Hoi

OrganizationsBeijing Institute of TechnologyInception Institute of AISalesforceSingapore Management UniversityUniversity of SurreyWuhan University

Why you should read this

Presents a comprehensive taxonomy of closed- and open-world person re-identification along with a competitive AGW baseline evaluated across twelve datasets and the mINP metric to measure practical search costs.

Person re-identification (Re-ID) aims at retrieving a person of interest across multiple non-overlapping cameras. With the advancement of deep neural networks and increasing demand of intelligent video surveillance, it has gained significantly increased interest in the computer vision community. By dissecting the involved components in developing a person Re-ID system, we categorize it into the closed-world and open-world settings. The widely studied closed-world setting is usually applied under various research-oriented assumptions, and has achieved inspiring success using deep learning techniques on a number of datasets. We first conduct a comprehensive overview with in-depth analysis for closed-world person Re-ID from three different perspectives, including deep feature representation learning, deep metric learning and ranking optimization. With the performance saturation under closed-world setting, the research focus for person Re-ID has recently shifted to the open-world setting, facing more challenging issues. This setting is closer to practical applications under specific scenarios. We summarize the open-world Re-ID in terms of five different aspects. By analyzing the advantages of existing methods, we design a powerful AGW baseline, achieving state-of-the-art or at least comparable performance on twelve datasets for FOUR different Re-ID tasks. Meanwhile, we introduce a new evaluation metric (mINP) for person Re-ID, indicating the cost for finding all the correct matches, which provides an additional criteria to evaluate the Re-ID system for real applications. Finally, some important yet under-investigated open issues are discussed.

Added

2026-09-16

Beyond Part Models: Person Retrieval with Refined Part Pooling

Beyond Part Models: Person Retrieval with Refined Part Pooling

Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, Shengjin Wang

OrganizationsAustralian National UniversityHuaweiTsinghua UniversityUniversity of Technology SydneyUniversity of Texas at San Antonio

Why you should read this

Proposes a part-based convolutional baseline and a refined part pooling method that dynamically reassigns misaligned feature outliers across uniform horizontal partitions, setting a high standard for person re-identification without requiring external pose estimation.

Employing part-level features for pedestrian image description offers fine-grained information and has been verified as beneficial for person retrieval in very recent literature. A prerequisite of part discovery is that each part should be well located. Instead of using external cues, e.g., pose estimation, to directly locate parts, this paper lays emphasis on the content consistency within each part. Specifically, we target at learning discriminative part-informed features for person retrieval and make two contributions. (i) A network named Part-based Convolutional Baseline (PCB). Given an image input, it outputs a convolutional descriptor consisting of several part-level features. With a uniform partition strategy, PCB achieves competitive results with the state-of-the-art methods, proving itself as a strong convolutional baseline for person retrieval. (ii) A refined part pooling (RPP) method. Uniform partition inevitably incurs outliers in each part, which are in fact more similar to other parts. RPP re-assigns these outliers to the parts they are closest to, resulting in refined parts with enhanced within-part consistency. Experiment confirms that RPP allows PCB to gain another round of performance boost. For instance, on the Market-1501 dataset, we achieve (77.4+4.2)% mAP and (92.3+1.5)% rank-1 accuracy, surpassing the state of the art by a large margin.

Added

2026-09-14