Built independently by an author, for readers. Read the story and support ChapterPal

keyword

word-region fragment pairs

Word-region fragment pairs are cross-modal combinations of individual textual elements, such as words or tokens, and localized visual components, such as detected object bounding boxes or image patches. In vision-language processing and image-text matching tasks, these pairs represent the smallest functional units used to evaluate fine-grained semantic alignment between visual and textual modalities. Machine learning models analyze the semantic correspondence within each pair to quantify degrees of similarity or detect mismatching clues, subsequently aggregating these local interactions across an entire sentence and image to determine overall cross-modal relevance and coherence.

1 item

Negative-Aware Attention Framework for Image-Text Matching

Negative-Aware Attention Framework for Image-Text Matching

Kun Zhang, Zhendong Mao, Quan Wang, Yongdong Zhang

OrganizationsBeijing University of Posts and TelecommunicationsUniversity of Science and Technology of China

Why you should read this

Proposes a negative-aware attention framework that explicitly mines mismatched word-region fragments alongside matched clues to prevent false-positive alignments and achieve state-of-the-art image-text matching accuracy.

Image-text matching, as a fundamental task, bridges the gap between vision and language. The key of this task is to accurately measure similarity between these two modalities. Prior work measuring this similarity mainly based on matched fragments (i.e., word/region with high relevance), while underestimating or even ignoring the effect of mismatched fragments (i.e., word/region with low relevance), e.g., via a typical LeakyReLU or ReLU operation that forces negative scores close or exact to zero in attention. This work argues that mismatched textual fragments, which contain rich mismatching clues, are also crucial for image-text matching. We thereby propose a novel Negative-Aware Attention Framework (NAAF), which explicitly exploits both the positive effect of matched fragments and the negative effect of mismatched fragments to jointly infer image-text similarity. NAAF (1) delicately designs an iterative optimization method to maximally mine the mismatched fragments, facilitating more discriminative and robust negative effects, and (2) devises the two-branch matching mechanism to precisely calculate similarity/dissimilarity degrees for matched/mismatched fragments with different masks. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superior effectiveness of our NAAF, achieving state-of-the-art performance. Code will be released at: https://github.com/CrossmodalGroup/NAAF.

Added

2026-09-26