keyword
image-text matching
Image-text matching is a multimodal machine learning task that measures the semantic correspondence and similarity between visual data and natural language descriptions. The primary objective is to bridge the gap between vision and language by determining whether an image and a text accurately depict the same scene, objects, or concepts. To accomplish this, computational models typically extract features from both modalities, such as visual regions and textual tokens, and map them into a shared representation space or align them using cross-attention mechanisms. By quantifying the degree of relevance between visual content and written sentences, image-text matching serves as a foundational component for bidirectional cross-modal retrieval, including searching for images using text queries and retrieving descriptive text for given images.
3 items

Noisy Correspondence Learning with Meta Similarity Correction
Haochen Han, Kaiyao Miao, Qinghua Zheng, Minnan Luo
Why you should read this
Proposes a meta-learning framework that trains a correction network on clean and mismatched meta-data to rectify similarity scores and filter out mismatched cross-modal pairs during retrieval training.
Despite the success of multimodal learning in cross-modal retrieval task, the remarkable progress relies on the correct correspondence among multimedia data. However, collecting such ideal data is expensive and time-consuming. In practice, most widely used datasets are harvested from the Internet and inevitably contain mismatched pairs. Training on such noisy correspondence datasets causes performance degradation because the cross-modal retrieval methods can wrongly enforce the mismatched data to be similar. To tackle this problem, we propose a Meta Similarity Correction Network (MSCN) to provide reliable similarity scores. We view a binary classification task as the meta-process that encourages the MSCN to learn discrimination from positive and negative meta-data. To further alleviate the influence of noise, we design an effective data purification strategy using meta-data as prior knowledge to remove the noisy samples. Extensive experiments are conducted to demonstrate the strengths of our method in both synthetic and real-world noises, including Flickr30K, MS-COCO, and Conceptual Captions. Our code is publicly available.
Added
2026-09-26

Negative-Aware Attention Framework for Image-Text Matching
Kun Zhang, Zhendong Mao, Quan Wang, Yongdong Zhang
Why you should read this
Proposes a negative-aware attention framework that explicitly mines mismatched word-region fragments alongside matched clues to prevent false-positive alignments and achieve state-of-the-art image-text matching accuracy.
Image-text matching, as a fundamental task, bridges the gap between vision and language. The key of this task is to accurately measure similarity between these two modalities. Prior work measuring this similarity mainly based on matched fragments (i.e., word/region with high relevance), while underestimating or even ignoring the effect of mismatched fragments (i.e., word/region with low relevance), e.g., via a typical LeakyReLU or ReLU operation that forces negative scores close or exact to zero in attention. This work argues that mismatched textual fragments, which contain rich mismatching clues, are also crucial for image-text matching. We thereby propose a novel Negative-Aware Attention Framework (NAAF), which explicitly exploits both the positive effect of matched fragments and the negative effect of mismatched fragments to jointly infer image-text similarity. NAAF (1) delicately designs an iterative optimization method to maximally mine the mismatched fragments, facilitating more discriminative and robust negative effects, and (2) devises the two-branch matching mechanism to precisely calculate similarity/dissimilarity degrees for matched/mismatched fragments with different masks. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superior effectiveness of our NAAF, achieving state-of-the-art performance. Code will be released at: https://github.com/CrossmodalGroup/NAAF.
Added
2026-09-26

Stacked Cross Attention for Image-Text Matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, Xiaodong He
Why you should read this
Proposes a stacked cross-attention network that aligns visual regions with corresponding sentence words, achieving significant gains in bidirectional image-text retrieval accuracy on MS-COCO and Flickr30K.
In this paper, we study the problem of image-text matching. Inferring the latent semantic alignment between objects or other salient stuff (e.g. snow, sky, lawn) and the corresponding words in sentences allows to capture fine-grained interplay between vision and language, and makes image-text matching more interpretable. Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to capture limited number of semantic alignments which is less interpretable. In this paper, we present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. Our approach achieves the state-of-the-art results on the MS-COCO and Flickr30K datasets. On Flickr30K, our approach outperforms the current best methods by 22.1% relatively in text retrieval from image query, and 18.2% relatively in image retrieval with text query (based on Recall@1). On MS-COCO, our approach improves sentence retrieval by 17.8% relatively and image retrieval by 16.6% relatively (based on Recall@1 using the 5K test set). Code has been made available at: this https URL.
Added
2026-09-25
