Built independently by an author, for readers. Read the story and support ChapterPal

keyword

negative-aware attention framework

A negative-aware attention framework is a cross-modal deep learning architecture that measures the correspondence between different modalities, such as images and text, by jointly evaluating both matching and mismatching elements. While conventional cross-modal attention mechanisms typically prioritize highly relevant, aligned fragments and suppress or discard low-relevance associations, a negative-aware framework explicitly mines and incorporates the negative signals from unaligned components. By employing specialized matching pathways to compute similarity scores for aligned fragments alongside dissimilarity penalties for mismatched fragments, this framework prevents false-positive associations and provides a more discriminative, fine-grained assessment of overall cross-modal alignment.

1 item

Negative-Aware Attention Framework for Image-Text Matching

Negative-Aware Attention Framework for Image-Text Matching

Kun Zhang, Zhendong Mao, Quan Wang, Yongdong Zhang

OrganizationsBeijing University of Posts and TelecommunicationsUniversity of Science and Technology of China

Why you should read this

Proposes a negative-aware attention framework that explicitly mines mismatched word-region fragments alongside matched clues to prevent false-positive alignments and achieve state-of-the-art image-text matching accuracy.

Image-text matching, as a fundamental task, bridges the gap between vision and language. The key of this task is to accurately measure similarity between these two modalities. Prior work measuring this similarity mainly based on matched fragments (i.e., word/region with high relevance), while underestimating or even ignoring the effect of mismatched fragments (i.e., word/region with low relevance), e.g., via a typical LeakyReLU or ReLU operation that forces negative scores close or exact to zero in attention. This work argues that mismatched textual fragments, which contain rich mismatching clues, are also crucial for image-text matching. We thereby propose a novel Negative-Aware Attention Framework (NAAF), which explicitly exploits both the positive effect of matched fragments and the negative effect of mismatched fragments to jointly infer image-text similarity. NAAF (1) delicately designs an iterative optimization method to maximally mine the mismatched fragments, facilitating more discriminative and robust negative effects, and (2) devises the two-branch matching mechanism to precisely calculate similarity/dissimilarity degrees for matched/mismatched fragments with different masks. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superior effectiveness of our NAAF, achieving state-of-the-art performance. Code will be released at: https://github.com/CrossmodalGroup/NAAF.

Added

2026-09-26