Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID
Wentao TanChangxing DingJiayu JiangFei WangYibing ZhanDapeng Tao
Proposes a scalable framework for transferable text-to-image person re-identification that uses multimodal large language models to generate diverse text annotations via dynamic templates while filtering out hallucinated descriptions with noise-aware masking.
Text-to-image person re-identification—retrieving surveillance images of pedestrians based on descriptive text queries—is critical for security, crowd management, and social media analysis when photo probes are unavailable. However, existing models suffer from severe cross-domain performance degradation and struggle to transfer to new environments because current datasets rely on slow, expensive manual text annotation and are too small. While automated captioning via multi-modal large language models offers a scalable alternative, generated descriptions often suffer from repetitive sentence structures and hallucinated or inaccurate descriptive errors, causing models to overfit narrow language patterns or learn incorrect visual-text alignments.
The article evaluates whether multi-modal large language models can automatically generate high-volume training data to create a robust, transferable text-to-image retrieval system that directly deploys across diverse unseen benchmarks without fine-tuning on target domain data.
The authors curated a massive dataset of one million pedestrian images paired with four million synthetic captions generated by two public vision-language models. To solve repetitive sentence patterns, the team used dialogue-prompted language models to generate 46 structured phrasing templates that dynamically varied image descriptions. To address model hallucinations and inaccurate descriptive tokens, the team introduced a noise-aware masking method. This technique measures similarity between text tokens and image patch embeddings during training, identifying mismatched words and masking them out during training epochs rather than attempting to predict errors.
Evaluation on standard industry benchmarks demonstrated substantial gains. Models pre-trained on this curated dataset and evaluated under a direct transfer setting outperformed existing pre-training baselines by wide margins, lifting top-1 accuracy on standard benchmarks from previous baselines of roughly 7% to 22% up to 38% to 58%. In traditional fine-tuning configurations, the approach established new state-of-the-art results, raising top-1 retrieval performance on the RSTPReid benchmark by over 8% and mean average precision by nearly 6%. The ablation studies further confirmed that dynamically varying sentence templates improved top-1 retrieval by 3% to 4%, while noise-aware masking added an additional 2% to 3.5% gain across benchmarks, effectively proving that managing caption noise is vital for automated data pipelines.
These findings indicate that automated synthetic dataset generation can replace costly human annotation pipelines for fine-grained computer vision tasks, significantly lowering implementation costs and deployment timelines. Furthermore, mitigating noise by dynamically suppressing mismatched tokens proves vastly superior to standard predictive language modeling when working with imperfect, machine-generated annotations.
Stakeholders deploying automated surveillance and retrieval systems should transition toward synthetic multi-modal data pipelines while integrating structural prompt variation and error-filtering mechanisms to improve model generalization across camera networks. Before operational deployment, teams should conduct pilot evaluations to calibrate optimal masking ratios and expand captioning template sets to reflect specific operational phrasing.
The study's primary limitations stem from reliance on a fixed set of 46 sentence templates, which may not capture all natural linguistic variations, and occasional failures of the masking mechanism to detect subtle caption errors. Nevertheless, the substantial and consistent performance gains across multiple established benchmarks provide strong confidence in the viability of the approach.
- Paper: Improving CLIP Training with Language Rewrites, Lijie Fan et al. (2023). This paper demonstrates how leveraging language models to generate diverse caption rewrites prevents vision-language models from overfitting to rigid sentence structures, establishing the conceptual basis for the source's multi-template text generation strategy.
- Paper: BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models, Junnan Li et al. (2023). This work introduces the foundational architecture for utilizing frozen large language models to generate rich visual descriptions and cross-modal representations, which the source relies on to automatically synthesize person ReID training data.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). This study establishes methods for aligning visual features and textual tokens while mitigating the effects of noisy supervision in vision-language pre-training, directly motivating the source's text-patch similarity filtering and masking pipeline.
- Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). This paper provides the foundational cross-attention framework for computing fine-grained similarities between image regions and individual words, underlying the source's method for identifying non-corresponding tokens.
No sufficiently relevant recommendations were found.
