Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID

Wentao TanChangxing DingJiayu JiangFei WangYibing ZhanDapeng Tao

article2024CVPR91 citations

Proposes a scalable framework for transferable text-to-image person re-identification that uses multimodal large language models to generate diverse text annotations via dynamic templates while filtering out hallucinated descriptions with noise-aware masking.

Listen

Text-to-image person re-identification—retrieving surveillance images of pedestrians based on descriptive text queries—is critical for security, crowd management, and social media analysis when photo probes are unavailable. However, existing models suffer from severe cross-domain performance degradation and struggle to transfer to new environments because current datasets rely on slow, expensive manual text annotation and are too small. While automated captioning via multi-modal large language models offers a scalable alternative, generated descriptions often suffer from repetitive sentence structures and hallucinated or inaccurate descriptive errors, causing models to overfit narrow language patterns or learn incorrect visual-text alignments.

The article evaluates whether multi-modal large language models can automatically generate high-volume training data to create a robust, transferable text-to-image retrieval system that directly deploys across diverse unseen benchmarks without fine-tuning on target domain data.

The authors curated a massive dataset of one million pedestrian images paired with four million synthetic captions generated by two public vision-language models. To solve repetitive sentence patterns, the team used dialogue-prompted language models to generate 46 structured phrasing templates that dynamically varied image descriptions. To address model hallucinations and inaccurate descriptive tokens, the team introduced a noise-aware masking method. This technique measures similarity between text tokens and image patch embeddings during training, identifying mismatched words and masking them out during training epochs rather than attempting to predict errors.

Evaluation on standard industry benchmarks demonstrated substantial gains. Models pre-trained on this curated dataset and evaluated under a direct transfer setting outperformed existing pre-training baselines by wide margins, lifting top-1 accuracy on standard benchmarks from previous baselines of roughly 7% to 22% up to 38% to 58%. In traditional fine-tuning configurations, the approach established new state-of-the-art results, raising top-1 retrieval performance on the RSTPReid benchmark by over 8% and mean average precision by nearly 6%. The ablation studies further confirmed that dynamically varying sentence templates improved top-1 retrieval by 3% to 4%, while noise-aware masking added an additional 2% to 3.5% gain across benchmarks, effectively proving that managing caption noise is vital for automated data pipelines.

These findings indicate that automated synthetic dataset generation can replace costly human annotation pipelines for fine-grained computer vision tasks, significantly lowering implementation costs and deployment timelines. Furthermore, mitigating noise by dynamically suppressing mismatched tokens proves vastly superior to standard predictive language modeling when working with imperfect, machine-generated annotations.

Stakeholders deploying automated surveillance and retrieval systems should transition toward synthetic multi-modal data pipelines while integrating structural prompt variation and error-filtering mechanisms to improve model generalization across camera networks. Before operational deployment, teams should conduct pilot evaluations to calibrate optimal masking ratios and expand captioning template sets to reflect specific operational phrasing.

The study's primary limitations stem from reliance on a fixed set of 46 sentence templates, which may not capture all natural linguistic variations, and occasional failures of the masking mechanism to detect subtle caption errors. Nevertheless, the substantial and consistent performance gains across multiple established benchmarks provide strong confidence in the viability of the approach.

No sufficiently relevant recommendations were found.

Cover for Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID

Abstract

Text-to-image person re-identification (ReID) retrieves pedestrian images according to textual descriptions. Manually annotating textual descriptions is time-consuming, restricting the scale of existing datasets and therefore the generalization ability of ReID models. As a result, we study the transferable text-to-image ReID problem, where we train a model on our proposed large-scale database and directly deploy it to various datasets for evaluation. We obtain substantial training data via Multi-modal Large Language Models (MLLMs). Moreover, we identify and address two key challenges in utilizing the obtained textual descriptions. First, an MLLM tends to generate descriptions with similar structures, causing the model to overfit specific sentence patterns. Thus, we propose a novel method that uses MLLMs to caption images according to various templates. These templates are obtained using a multi-turn dialogue with a Large Language Model (LLM). Therefore, we can build a large-scale dataset with diverse textual descriptions. Second, an MLLM may produce incorrect descriptions. Hence, we introduce a novel method that automatically identifies words in a description that do not correspond with the image. This method is based on the similarity between one text and all patch token embeddings in the image. Then, we mask these words with a larger probability in the subsequent training epoch, alleviating the impact of noisy textual descriptions. The experimental results demonstrate that our methods significantly boost the direct transfer text-to-image ReID performance. Benefiting from the pre-trained model weights, we also achieve state-of-the-art performance in the traditional evaluation settings.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methods
  • 3.1. Generating Diverse Descriptions
  • 3.2. Noise-Aware Masking
  • 3.3. Optimization
  • 4. Experiments
  • 4.1. Datasets and Settings
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparisons with State-of-the-Art Methods
  • 5. Conclusion and Limitations
  • References

Knowls

  1. Knowl 1 — Noise-Aware Masking for Filtering MLLM Textual Hallucinations

    model/method

    Multi-modal Large Language Models (MLLMs) can generate rich descriptive texts for pedestrian images, but frequently produce hallucinations or inaccurate attributes (such as incorrect colors or non-existent accessories). Rather than discarding entire image-text pairs containing noisy words—which wastes valid semantic details in matching tokens—Noise-Aware Masking (NAM) dynamically estimates the noisiness of each individual text token and selectively masks suspected tokens during training.

    For a full textual description TfullT^{full} tokenized into NN tokens and an image tokenized into MM non-overlapping patches, NAM calculates the cosine similarity between the intermediate embedding of each text token and all visual patch token embeddings. A text token whose embedding shows consistently low similarity across all image patches is identified as unrelated or noisy. Each token is assigned a masking probability proportional to its noise level, with the average sequence masking rate calibrated to a constant hyperparameter pp. The resulting masked text TnamT^{nam} is fed into the text encoder to produce the global textual representation used for matching loss computation, preventing the model from fitting incorrect text-image correspondences.

  2. Knowl 2 — Cross-Modal Token Similarity and Noise-Level Calibration in NAM

    equation

    Let Ft=[t1l,au2l,auNl]∈RN×dF_t = [t^l_1, au^l_2, au^l_N] \in \mathbb{R}^{N \times d} and Fv=[v1l,v2l,…,vMl]∈RM×dF_v = [v^l_1, v^l_2, \dots, v^l_M] \in \mathbb{R}^{M \times d} denote the intermediate token embeddings extracted from the ll-th layer of the text encoder and image encoder, respectively, where dd is the feature dimension, NN is the number of text tokens, and MM is the number of image patch tokens.

    The token-wise cross-modal cosine similarity matrix S∈RN×MS \in \mathbb{R}^{N \times M} is calculated as:

    S=Ft⊤FvS = F_t^\top F_v

    where sijs_{ij} is the cosine similarity between the ii-th text token and the jj-th visual patch token. The raw noise level rir_i of the ii-th text token is computed as:

    ri=1−(max⁡1≤j≤Msij)r_i = 1 - \left( \max_{1 \le j \le M} s_{ij} \right)

    To prevent over-masking in early training phases when similarity scores are generally lower, the noise vector r=[r1,…,rN]r = [r_1, \dots, r_N] is centered around the target average masking ratio p∈(0,1)p \in (0, 1) via:

    Er=1N∑i=1NriE_r = \frac{1}{N} \sum_{i=1}^N r_i

    ri′=ri−Er+pr'_i = r_i - E_r + p

    The calibrated value ri′r'_i is used as the probability of masking the ii-th token in the subsequent training epoch. In the initial training epoch, r′r' is initialized uniformly with constant value pp.

  3. Knowl 3 — Template-based Diversity Enhancement (TDE) for MLLM Pedestrian Captioning

    model/method

    When prompted with fixed static instructions, Multi-modal Large Language Models (MLLMs) tend to generate descriptions with uniform, repetitive sentence structures. This lack of structural diversity leads downstream text-to-image person ReID models to overfit specific sentence patterns, impairing zero-shot transfer to real-world natural language queries.

    Template-based Diversity Enhancement (TDE) resolves this by dynamically incorporating diverse sentence structures into the MLLM generation prompt. Using multi-turn dialogue with a Large Language Model (e.g., ChatGPT), an initial set of descriptive captions is analyzed to extract underlying syntactic patterns, which are expanded into 46 distinct sentence templates covering variations in ordering, phrasing, and grouping of attributes (such as clothing, footwear, hairstyle, gender, and belongings).

    During synthetic dataset generation, each image is captioned using both a static prompt and a dynamic prompt that randomly samples one of the 46 templates: "Generate a description about the overall appearance of the person, including clothing, shoes, hairstyle, gender, and belongings, in a style similar to the template: '{template}'. If some requirements in the template are not visible, you can ignore them. Do not imagine any contents that are not in the image." Captions generated under both prompts across multiple MLLMs (e.g., Qwen-VL and Shikra) provide complementary descriptive viewpoints and diverse sentence patterns.

  4. Knowl 4 — Similarity Distribution Matching Objective on Masked Descriptions

    equation

    The text-to-image ReID model is optimized using bidirectional Similarity Distribution Matching (SDM) loss calculated between the global visual feature vcls∈Rdv_{cls} \in \mathbb{R}^d (the [CLS] token of the last visual encoder layer) and the masked text feature teos′∈Rdt'_{eos} \in \mathbb{R}^d (the [EOS] token of the masked description TnamT^{nam} from the last text encoder layer).

    For a mini-batch of BB image-text pairs with ground-truth binary associations yi,j∈{0,1}y_{i,j} \in \{0, 1\}, the ground-truth matching distribution for the ii-th image is defined as qi,j=yi,j/∑b=1Byi,bq_{i,j} = y_{i,j} / \sum_{b=1}^B y_{i,b}. The predicted image-to-text probability distribution pip_i is:

    pi,j=exp⁡(sim(vclsi,teos′j)/τ)∑b=1Bexp⁡(sim(vclsi,teos′b)/τ)p_{i,j} = \frac{\exp(\text{sim}(v_{cls}^i, {t'_{eos}}^j)/\tau)}{\sum_{b=1}^B \exp(\text{sim}(v_{cls}^i, {t'_{eos}}^b)/\tau)}

    where sim(u,v)=u⊤v∥u∥∥v∥\text{sim}(u, v) = \frac{u^\top v}{\|u\| \|v\|} is cosine similarity and τ>0\tau > 0 is a learnable temperature coefficient (set to 0.020.02). The image-to-text SDM loss is:

    Li2t=1B∑i=1BKL(pi∥qi)=1B∑i=1B∑j=1Bpi,jlog⁡(pi,jqi,j+ϵ)\mathcal{L}_{i2t} = \frac{1}{B} \sum_{i=1}^B \text{KL}(\mathbf{p}_i \parallel \mathbf{q}_i) = \frac{1}{B} \sum_{i=1}^B \sum_{j=1}^B p_{i,j} \log\left( \frac{p_{i,j}}{q_{i,j} + \epsilon} \right)

    where ϵ\epsilon prevents division by zero. The text-to-image loss Lt2i\mathcal{L}_{t2i} is formulated symmetrically by swapping vclsv_{cls} and teos′t'_{eos}. The overall loss is:

    Lsdm=Li2t+Lt2i\mathcal{L}_{sdm} = \mathcal{L}_{i2t} + \mathcal{L}_{t2i}

  5. Knowl 5 — Direct Transfer Performance Comparison Across Pre-Training Datasets

    data/table

    When models are trained exclusively on pre-training datasets and directly evaluated on downstream benchmarks without domain-specific fine-tuning (zero-shot direct transfer setting using a CLIP-ViT/B-16 backbone), LUPerson-MLLM substantially outperforms prior datasets such as MALS and LUPerson-T, even when trained with only 0.1M images.

    Pretrain Dataset CUHK-PEDES ICFG-PEDES RSTPReid
    R1 (%) mAP (%) R1 (%) mAP (%) R1 (%) mAP (%)
    None (Raw CLIP) 12.65 11.15 6.67 2.51 13.45 10.31
    MALS (1.5 M) 19.36 18.62 7.93 3.52 22.85 17.11
    LUPerson-T (0.95 M) 21.88 19.96 11.46 4.56 22.40 17.08
    LUPerson-MLLM (0.1 M) 52.64 46.48 32.61 16.53 47.75 34.73
    LUPerson-MLLM (1.0 M) 57.61 51.44 38.36 20.43 51.50 37.34

    The strong direct transfer capability stems from TDE providing syntactic diversity matching varied human search queries, and NAM filtering out noisy/hallucinated attribute tokens during training.

  6. Knowl 6 — Ablation on Component Contributions to Direct Transfer Text-to-Image ReID

    data/table

    Ablation study evaluating the cumulative effects of static texts (TsT^s), dynamic texts generated via Template-based Diversity Enhancement (TdT^d), and Noise-Aware Masking (NAM) using 0.1M training images under the direct transfer evaluation protocol.

    Configuration CUHK-PEDES ICFG-PEDES RSTPReid
    R1 R5 mAP R1 R5 mAP R1 R5 mAP
    CLIP Baseline 12.65 27.16 11.15 6.67 17.91 2.51 13.45 33.85 10.31
    TqwensT^s_{qwen} 37.65 57.86 33.40 23.78 42.77 11.18 36.30 60.60 26.25
    TshikrasT^s_{shikra} 39.70 62.60 36.09 19.02 35.63 9.67 36.90 62.65 28.33
    Tqwens+TshikrasT^s_{qwen} + T^s_{shikra} 46.00 66.82 41.27 26.74 44.22 13.23 41.10 66.95 30.21
    TqwendT^d_{qwen} 40.72 62.36 37.21 24.16 41.24 11.32 38.65 64.70 28.81
    TshikradT^d_{shikra} 43.63 65.46 39.08 22.07 39.57 11.35 38.80 63.45 28.60
    Tqwend+TshikradT^d_{qwen} + T^d_{shikra} 48.86 69.41 44.09 28.43 46.37 14.23 44.25 66.15 32.99
    TDE (Ts+TdT^s + T^d) 50.32 71.36 45.74 29.12 47.96 15.13 45.70 70.75 33.23
    TDE + NAM 52.64 71.62 46.48 32.61 50.79 16.48 47.75 70.75 34.73

    Dynamic text generation (TdT^d) consistently outperforms static text generation (TsT^s) by 3.07%–3.93% Rank-1 accuracy per MLLM. Replacing standard equal-probability token masking with NAM provides an additional +2.32%, +3.49%, and +2.05% Rank-1 improvement across CUHK-PEDES, ICFG-PEDES, and RSTPReid, respectively.

  7. Knowl 7 — Detrimental Effect of Masked Language Modeling (MLM) Loss on Noisy MLLM Texts

    empirical result

    While Masked Language Modeling (MLM) loss—predicting masked tokens from image and context features—is beneficial when training on clean, human-annotated text-image pairs, adding MLM loss to Noise-Aware Masking (NAM) on MLLM-generated descriptions significantly degrades retrieval performance across all datasets:

    • CUHK-PEDES: Rank-1 drops from 52.64% (NAM) to 48.79% (NAM w/ MLM loss); mAP drops from 46.48% to 43.86%.
    • ICFG-PEDES: Rank-1 drops from 32.61% to 27.36%; mAP drops from 16.48% to 14.16%.
    • RSTPReid: Rank-1 drops from 47.75% to 44.45%; mAP drops from 34.73% to 33.07%.

    Because MLLM descriptions contain hallucinated or noisy words, forcing the model to reconstruct masked tokens from image features forces incorrect cross-modal associations, verifying that noisy tokens should only be masked out rather than predicted.

  8. Knowl 8 — Layer Selection and Masking Ratio Optimization in Noise-Aware Masking

    empirical result

    Empirical evaluation of two design hyperparameters in NAM:

    1. Extraction Layer ll for Similarity Computation: Evaluating token representations from layer 7 to 12 in a 12-layer Transformer encoder reveals that the 10th layer yields optimal performance across CUHK-PEDES (52.64% Rank-1), ICFG-PEDES (32.61% Rank-1), and RSTPReid (48.30% Rank-1). The 10th layer provides richer fine-grained local alignment details than the final (12th) layer, which is biased toward global task abstraction.
    2. Overall Average Masking Probability pp: Comparing NAM against Equal Masking (EM) across p∈[0.10,0.30]p \in [0.10, 0.30] demonstrates that NAM consistently outperforms EM at every ratio. The optimal performance is achieved at p≈0.15p \approx 0.15.
  9. Knowl 9 — Fine-Tuning State-of-the-Art Performance with LUPerson-MLLM Pre-Training

    data/table

    Pre-training backbones on the 1.0M LUPerson-MLLM dataset before fine-tuning on specific target benchmarks sets new state-of-the-art results across text-to-image person ReID benchmarks under traditional in-domain and cross-domain fine-tuning evaluation.

    Method Backbone CUHK-PEDES ICFG-PEDES RSTPReid
    R1 R5 mAP R1 R5 mAP R1 R5 mAP
    IRRA (Baseline) CLIP-ViT/B 73.38 89.93 66.10 63.46 80.25 38.06 60.20 81.30 47.17
    MALS (1.5 M) + IRRA CLIP-ViT/B 74.05 89.48 66.57 64.37 80.75 38.85 61.90 80.60 48.08
    LUPerson-T (0.95 M) + IRRA CLIP-ViT/B 74.37 89.51 66.60 64.50 80.24 38.22 62.20 83.30 48.33
    Ours (1.0 M) + IRRA CLIP-ViT/B 76.82 91.16 69.55 67.05 82.16 41.51 68.50 87.15 53.02
    APTM (Baseline) Swin-B / BERT 76.53 90.04 66.91 68.51 82.99 41.22 67.50 85.70 52.56
    Ours (1.0 M) + APTM Swin-B / BERT 78.13 91.19 68.75 69.37 83.55 42.42 69.95 87.35 54.17

    When transferred across datasets (e.g., pre-training on ICFG-PEDES and testing on CUHK-PEDES), initializing from LUPerson-MLLM outperforms initialization from MALS and LUPerson-T by 20.82% and 26.13% in Rank-1 accuracy, respectively.

  10. Knowl 10 — Limitations of Template Expansion and Cross-Modal Noise Detection

    limitation

    The methodology has two primary limitations:

    1. Template Diversity Ceiling: The syntactic diversity of generated captions is bounded by the finite set of 46 LLM-generated templates. While this prevents overfitting to a single sentence structure, it does not achieve true arbitrary natural language variation.
    2. Noise Detection Failure Modes: Noise-Aware Masking relies entirely on intermediate cross-modal token-patch cosine similarity. It can occasionally misidentify correct but abstract/contextual words as noise or fail to mask subtle semantic hallucinations whose token embeddings have spurious high similarity to unrelated image patches.

Coverage note — None was omitted. All major conceptual contributions (TDE, NAM, dataset construction, SDM loss formulation, ablation studies, and comparative evaluations in direct transfer and fine-tuning settings) are fully covered.

References

  1. 1.Surbhi Aggarwal, Venkatesh Babu Radhakrishnan, and Anirban Chakraborty. Text-based person search via attributeaided matching. In WACV, 2020. 2
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. NeurIPS, 2022. 3
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023. 1, 2, 3
  4. 4.Yang Bai, Min Cao, Daming Gao, Ziqiang Cao, Chen Chen, Zhenfeng Fan, Liqiang Nie, and Min Zhang. Rasa: Relation and sensitivity aware representation learning for text-based person search. IJCAI, 2023. 1, 2, 5, 8
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. NeurIPS, 2020. 3
  6. 6.Maryam Bukhari, Sadaf Yasmin, Sheneela Naz, Muazzam Maqsood, Jehyeok Rew, and Seungmin Rho. Language and vision based person re-identification for surveillance systems using deep learning with lip layers. Image and Vision Computing, 2023. 1
  7. 7.Cuiqun Chen, Mang Ye, and Ding Jiang. Towards modality-agnostic person re-identification with descriptive query. In CVPR, 2023. 6
  8. 8.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023. 2, 3
  9. 9.Tianlang Chen, Chenliang Xu, and Jiebo Luo. Improving text-based person search by spatial matching and adaptive threshold. In WACV, 2018. 2
  10. 10.Yuhao Chen, Guoqing Zhang, Yujiang Lu, Zhenxing Wang, and Yuhui Zheng. Tipcb: A simple but effective part-based convolutional baseline for text-based person search. Neurocomputing, 2022. 2, 8
  11. 11.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 3
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2
  13. 13.Changxing Ding and Dacheng Tao. Trunk-branch ensemble convolutional neural networks for video-based face recognition. IEEE TPAMI, 2018. 1
  14. 14.Changxing Ding, Kan Wang, Pengfei Wang, and Dacheng Tao. Multi-task learning with coarse priors for robust partaware person re-identification. IEEE TPAMI, 2022.
  15. 15.Zefeng Ding, Changxing Ding, Zhiyin Shao, and Dacheng Tao. Semantically self-aligned network for text-toimage part-aware person re-identification. arXiv preprint arXiv:2107.12666, 2021. 1, 2, 3, 5, 8
  16. 16.Ammarah Farooq, Muhammad Awais, Josef Kittler, and Syed Safwan Khalid. Axm-net: Implicit cross-modal feature alignment for person re-identification. In AAAI, 2022. 1, 2, 8
  17. 17.Dengpan Fu, Dongdong Chen, Jianmin Bao, Hao Yang, Lu Yuan, Lei Zhang, Houqiang Li, and Dong Chen. Unsupervised pre-training for person re-identification. In CVPR, 2021. 2, 3, 5
  18. 18.Hiren Galiyawala and Mehul S Raval. Person retrieval in surveillance using textual query: a review. Multimedia Tools and Applications, 2021. 1
  19. 19.Chenyang Gao, Guanyu Cai, Xinyang Jiang, Feng Zheng, Jun Zhang, Yifei Gong, Pai Peng, Xiaowei Guo, and Xing Sun. Contextual non-local alignment over full-scale representation for text-based person search. arXiv preprint arXiv:2101.03036, 2021. 2
  20. 20.Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, et al. Imagebind-llm: Multi-modality instruction tuning. arXiv preprint arXiv:2309.03905, 2023. 3
  21. 21.Xiao Han, Sen He, Li Zhang, and Tao Xiang. Textbased person search with limited data. arXiv preprint arXiv:2110.10807, 2021. 2, 8
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 2
  23. 23.Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and Xiaojuan Qi. Is synthetic data from generative models ready for image recognition? In ICLR, 2023. 4
  24. 24.Ding Jiang and Mang Ye. Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval. In CVPR, 2023. 1, 2, 5, 6, 7, 8
  25. 25.Ya Jing, Chenyang Si, Junbo Wang, Wei Wang, Liang Wang, and Tieniu Tan. Pose-guided multi-granularity attention network for text-based person search. In AAAI, 2020. 2
  26. 26.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In ECCV, 2018. 2
  27. 27.Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. Multimodal foundation models: From specialists to general-purpose assistants. arXiv preprint arXiv:2309.10020, 2023. 3
  28. 28.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. NeurIPS, 2021. 2, 8
  29. 29.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022. 2, 3, 4, 7
  30. 30.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 3
  31. 31.Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang. Person search with natural language description. In CVPR, 2017. 1, 2, 3, 5
  32. 32.Shiping Li, Min Cao, and Min Zhang. Learning semanticaligned feature representation for text-based person search. In ICASSP, 2022. 2, 8
  33. 33.Zechao Li, Hao Tang, Zhimao Peng, Guo-Jun Qi, and Jinhui Tang. Knowledge-guided semantic transfer network for fewshot image recognition. TNNLS. 2
  34. 34.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 3
  35. 35.Long Short-Term Memory. Long short-term memory. Neural computation, 2010. 2
  36. 36.Kai Niu, Yan Huang, Wanli Ouyang, and Liang Wang. Improving description-based person re-identification by multigranularity image-text alignments. TIP, 2020. 2
  37. 37.OpenAI. Chatgpt. https://openai.com/blog/chatgpt/, 2022. 2, 3
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 1, 2, 6, 7, 8
  39. 39.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj̈orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 3
  40. 40.Nikolaos Sarafianos, Xiang Xu, and Ioannis A Kakadiaris. Adversarial representation learning for text-to-image matching. In ICCV, 2019. 2
  41. 41.Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin, Jian Wang, and Changxing Ding. Learning granularity-unified representations for text-to-image person re-identification. In ACM MM, 2022. 1, 2, 3, 8
  42. 42.Zhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang, and Jingdong Wang. Unified pre-training with pseudo texts for text-to-image person re-identification. In ICCV, 2023. 1, 2, 5, 7, 8
  43. 43.Xiujun Shu, Wei Wen, Haoqian Wu, Keyu Chen, Yiran Song, Ruizhi Qiao, Bo Ren, and Xiao Wang. See finer, see more: Implicit modality alignment for text-based person retrieval. In ECCV, 2022. 2, 8
  44. 44.Wentao Tan, Changxing Ding, Pengfei Wang, Mingming Gong, and Kui Jia. Style interleaved learning for generalizable person re-identification. IEEE TMM, 2024. 1
  45. 45.Hao Tang, Chengcheng Yuan, Zechao Li, and Jinhui Tang. Learning attention-guided pyramidal features for few-shot fine-grained recognition. Pattern Recognition. 2
  46. 46.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste ´ Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. ` Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 3
  47. 47.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 3
  48. 48.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. TMLR, 2022. 3
  49. 49.Liwei Wang, Yin Li, and Svetlana Lazebnik. Learning deep structure-preserving image-text embeddings. In CVPR, 2016. 2
  50. 50.Pengfei Wang, Changxing Ding, Wentao Tan, Mingming Gong, Kui Jia, and Dacheng Tao. Uncertainty-aware clustering for unsupervised domain adaptive object reidentification. IEEE TMM, 2023. 1
  51. 51.Yuyu Wang, Chunjuan Bo, Dong Wang, Shuang Wang, Yunwei Qi, and Huchuan Lu. Language person search with mutually connected classification loss. In ICASSP. IEEE, 2019. 2
  52. 52.Zhe Wang, Zhiyuan Fang, Jun Wang, and Yezhou Yang. Vitaa: Visual-textual attributes alignment in person search by natural language. In ECCV, 2020. 1, 2, 8
  53. 53.Zijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan, Chao Liu, Tian Wang, and Yifeng Li. Caibc: Capturing all-round information beyond color for text-based person retrieval. In ACM MM, 2022. 2, 8
  54. 54.Zijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan, Chao Liu, Tian Wang, and Yifeng Li. Look before you leap: Improving text-based person retrieval by learning a consistent crossmodal common manifold. In ACM MM, 2022. 2, 8
  55. 55.Yushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li, Changqing Zou, and Shuguang Cui. Lapscore: languageguided person search via color reasoning. In ICCV, 2021. 8
  56. 56.Ziqiang Wu, Bingpeng Ma, Hong Chang, and Shiguang Shan. Refined knowledge transfer for language-based person search. TMM, 2023. 2
  57. 57.Yi Xie, Jianqing Zhu, Huanqiang Zeng, Canhui Cai, and Lixin Zheng. Learning matching behavior differences for compressing vehicle re-identification models. In VCIP, 2020. 1
  58. 58.Yi Xie, Fei Shen, Jianqing Zhu, and Huanqiang Zeng. Viewpoint robust knowledge distillation for accelerating vehicle re-identification. EURASIP J ADV SIG PR, 2021. 1
  59. 59.Yi Xie, Hanxiao Wu, Fei Shen, Jianqing Zhu, and Huanqiang Zeng. Object re-identification using teacher-like and light students. In BMVC, 2021. 2
  60. 60.Yi Xie, Huaidong Zhang, Xuemiao Xu, Jianqing Zhu, and Shengfeng He. Towards a smaller student: Capacity dynamic distillation for efficient image retrieval. In CVPR, 2023. 1
  61. 61.Shuanglin Yan, Neng Dong, Jun Liu, Liyan Zhang, and Jinhui Tang. Learning comprehensive representations with richer self for text-to-image person re-identification. In ACM MM, 2023. 1, 8
  62. 62.Shuanglin Yan, Hao Tang, Liyan Zhang, and Jinhui Tang. Image-specific information suppression and implicit local alignment for text-based person search. TNNLS, 2023. 2
  63. 63.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, et al. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305, 2023. 3
  64. 64.Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu, and Yujiao Wu. Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. In ACM MM, 2023. 1, 2, 3, 5, 7, 8
  65. 65.Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549, 2023. 2, 3
  66. 66.Ying Zhang and Huchuan Lu. Deep cross-modal projection learning for image-text matching. In ECCV, 2018. 1, 2, 8
  67. 67.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023. 3
  68. 68.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. NeurIPS, 2023. 3
  69. 69.Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, Mingliang Xu, and Yi-Dong Shen. Dual-path convolutional image-text embeddings with instance loss. ACM TOMM, 2020. 1, 8
  70. 70.Aichun Zhu, Zijie Wang, Yifeng Li, Xili Wan, Jing Jin, Tian Wang, Fangqiang Hu, and Gang Hua. Dssl: Deep surroundings-person separation learning for text-based person retrieval. In ACM MM, 2021. 1, 2, 5
  71. 71.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 3

Citation

MLA
Tan, W., et al. “Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID”. arXiv, 2024, http://arxiv.org/abs/2405.04940v3.
APA
Tan, W., Ding, C., Jiang, J., Wang, F., Zhan, Y., & Tao, D. (2024). Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID. arXiv. http://arxiv.org/abs/2405.04940v3
Chicago
Tan, W., C. Ding, J. Jiang, F. Wang, Y. Zhan, and D. Tao. 2024. “Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID”. arXiv. http://arxiv.org/abs/2405.04940v3.
Harvard
Tan, W. et al. (2024) “Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.04940v3.
Vancouver
1. Tan W, Ding C, Jiang J, Wang F, Zhan Y, Tao D (2024) Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID. arXiv

BibTeX

@article{tan2024harnessing,
  title = {Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID},
  author = {Tan, Wentao and Ding, Changxing and Jiang, Jiayu and Wang, Fei and Zhan, Yibing and Tao, Dapeng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.04940v3},
  eprint = {2405.04940}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE