Negative-Aware Attention Framework for Image-Text Matching

Kun ZhangZhendong MaoQuan WangYongdong Zhang

article2022CVPR159 citations

Proposes a negative-aware attention framework that explicitly mines mismatched word-region fragments alongside matched clues to prevent false-positive alignments and achieve state-of-the-art image-text matching accuracy.

Listen

Cross-modal search between images and text is a foundational capability for multimodal artificial intelligence, enabling systems to retrieve relevant visual content given a textual query and vice versa. Existing matching approaches predominantly rely on positive alignment between words and image regions while suppressing or ignoring mismatched elements. This unilateral focus frequently causes false-positive errors, as an image and text pair containing multiple matching items can achieve a top retrieval rank despite containing critical descriptive words that do not exist in the visual scene.

The article demonstrates and evaluates a novel Negative-Aware Attention Framework (NAAF). The primary objective is to improve image-text matching accuracy by jointly factoring in both the positive contribution of matched elements and the negative penalty of mismatched textual elements.

To achieve this, the approach introduces a two-branch matching architecture alongside an iterative optimization mechanism. The model calculates positive alignment through standard attention while isolating mismatched words to compute explicit dissimilarity penalties that downgrade false positives. Because no manual word-level annotations exist to distinguish matches from mismatches, the method adaptively learns a dynamic decision boundary from sampled similarity distributions. The framework was evaluated on two widely used benchmark datasets, Flickr30K and MS-COCO, across standard top-rank recall metrics.

Key findings show that NAAF significantly improves cross-modal retrieval performance over existing state-of-the-art models. On the Flickr30K benchmark, the framework achieved an overall recall sum of 513.2, representing an absolute improvement of 13.6 points over previous top methods and a 48.2% relative gain compared to baseline cross-attention models. On the larger MS-COCO dataset, NAAF outperformed comparable techniques across both the 1,000-image and 5,000-image test splits, yielding near-term top-1 recall improvements of roughly 1% to 4% across directions. Ablation analyses confirmed that removing the negative attention branch caused a substantial drop in performance, proving that dissimilarity penalties are critical for cross-modal precision.

These results indicate that penalizing mismatched semantic concepts is essential for building reliable cross-modal retrieval and search systems. By explicitly downgrading high-similarity false positives, organizations can reduce search error rates and deliver more accurate multimodal applications. For future implementations, engineering teams should adopt two-branch attention architectures and dynamically tuned decision boundaries when developing text-to-image retrieval pipelines.

Confidence in these findings is high given the consistent gains across multiple benchmark splits and clear ablation studies. However, the framework focuses primarily on mismatched textual fragments rather than image-to-text mismatches, because visual scenes naturally contain unrelated background objects. Operational deployments will require evaluating performance across specialized, domain-specific vocabularies and broader enterprise datasets.

Cover for Negative-Aware Attention Framework for Image-Text Matching

Abstract

Image-text matching, as a fundamental task, bridges the gap between vision and language. The key of this task is to accurately measure similarity between these two modalities. Prior work measuring this similarity mainly based on matched fragments (i.e., word/region with high relevance), while underestimating or even ignoring the effect of mismatched fragments (i.e., word/region with low relevance), e.g., via a typical LeakyReLU or ReLU operation that forces negative scores close or exact to zero in attention. This work argues that mismatched textual fragments, which contain rich mismatching clues, are also crucial for image-text matching. We thereby propose a novel Negative-Aware Attention Framework (NAAF), which explicitly exploits both the positive effect of matched fragments and the negative effect of mismatched fragments to jointly infer image-text similarity. NAAF (1) delicately designs an iterative optimization method to maximally mine the mismatched fragments, facilitating more discriminative and robust negative effects, and (2) devises the two-branch matching mechanism to precisely calculate similarity/dissimilarity degrees for matched/mismatched fragments with different masks. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superior effectiveness of our NAAF, achieving state-of-the-art performance. Code will be released at: https://github.com/CrossmodalGroup/NAAF.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Negative-aware Attention
  • 3.1.1 Discriminative Mismatch Mining
  • 3.1.2 Neg-Pos Branch Matching
  • 3.1.3 Sampling and Updating Strategy
  • 3.2. Objective Function
  • 3.3. Feature Extraction
  • 4. Experiments
  • 4.1. Dataset and Implementation Details
  • 4.2. Comparison Results
  • 4.3. Ablation Study
  • 4.4. Visualization and Case Study
  • 5. Conclusion
  • 6. Acknowledgements
  • References

Knowls

  1. Knowl 1 — Negative-aware attention framework

    model/method

    NAAF is a local-level image-text matching framework that jointly models evidence for matching and mismatching instead of retaining only highly relevant word-region pairs. For an image and a sentence, it first extracts visual-region and word features, then uses two complementary modules: discriminative mismatch mining learns an adaptive boundary separating matched from mismatched word-region similarities, and Neg-Pos branch matching computes positive similarity from matched fragments and negative dissimilarity from mismatched textual fragments. The final image-text score combines both effects, allowing a few strongly mismatched words to downgrade otherwise false-positive pairs.

  2. Knowl 2 — Adaptive mismatch-distribution boundary

    model/method

    During training, NAAF maintains dynamically updated samples Sk−={s1−,s2−,…}S_k^- = \{s_1^-,s_2^-,\ldots\} and Sk+={s1+,s2+,…}S_k^+ = \{s_1^+,s_2^+,\ldots\} of mismatched and matched word-region similarities, respectively, where kk indexes the training update. It models these samples as Gaussian similarity distributions with means and standard deviations (μk−,σk−)(\mu_k^-,\sigma_k^-) and (μk+,σk+)(\mu_k^+,\sigma_k^+), where σk−,σk+>0\sigma_k^-,\sigma_k^+>0. A threshold t≥0t\geq 0 classifies a word-region pair as mismatched below the threshold and matched above it. The threshold minimizes a weighted overlap error:

    min⁡t≥0  α∫t+∞fk−(s) ds+∫−∞tfk+(s) ds,\min_{t\geq 0}\; \alpha\int_t^{+\infty} f_k^-(s)\,ds + \int_{-\infty}^{t} f_k^+(s)\,ds,

    where s∈Rs\in\mathbb{R} is similarity, α>0\alpha>0 penalizes mismatched fragments incorrectly classified as matched, and

    fk−(s)=1σk−2πexp⁡ ⁣(−(s−μk−)22(σk−)2),fk+(s)=1σk+2πexp⁡ ⁣(−(s−μk+)22(σk+)2).f_k^-(s)=\frac{1}{\sigma_k^-\sqrt{2\pi}}\exp\!\left(-\frac{(s-\mu_k^-)^2}{2(\sigma_k^-)^2}\right),\qquad f_k^+(s)=\frac{1}{\sigma_k^+\sqrt{2\pi}}\exp\!\left(-\frac{(s-\mu_k^+)^2}{2(\sigma_k^+)^2}\right).

    The resulting boundary is

    tk=[(β2k)2−4β1kβ3k−β2k2β1k]+,t_k=\left[\frac{\sqrt{(\beta_2^k)^2-4\beta_1^k\beta_3^k}-\beta_2^k}{2\beta_1^k}\right]_+,

    where [x]+=max⁡(x,0)[x]_+=\max(x,0), β1k=(σk+)2−(σk−)2\beta_1^k=(\sigma_k^+)^2-(\sigma_k^-)^2, β2k=2(μk+(σk−)2−μk−(σk+)2)\beta_2^k=2(\mu_k^+(\sigma_k^-)^2-\mu_k^-(\sigma_k^+)^2), and

    β3k=(σk+μk−)2−(σk−μk+)2+2(σk+σk−)2ln⁡ ⁣(σk−ασk+).\beta_3^k=(\sigma_k^+\mu_k^-)^2-(\sigma_k^-\mu_k^+)^2+2(\sigma_k^+\sigma_k^-)^2\ln\!\left(\frac{\sigma_k^-\alpha}{\sigma_k^+}\right).

    The learned boundary is fed back into attention matching while the distributions are resampled and updated, producing an iterative optimization that progressively separates matched and mismatched similarities.

  3. Knowl 3 — Penalty condition for accurate mismatch mining

    theoretical result

    For the Gaussian matched and mismatched similarity distributions used by NAAF, the paper gives a condition for adjusting the mismatch-error penalty so that the learned boundary approaches a state with high mining accuracy. With α>0\alpha>0 the penalty parameter, μk−\mu_k^- and μk+\mu_k^+ the distribution means, and σk−>0\sigma_k^->0 and σk+>0\sigma_k^+>0 their standard deviations, the adjustment target is

    α∗=σk−[σk+exp⁡ ⁣(β4k2((σk+)2−(σk−)2))]−1,\alpha^*=\sigma_k^-\left[\sigma_k^+\exp\!\left(\frac{\beta_4^k}{2\big((\sigma_k^+)^2-(\sigma_k^-)^2\big)}\right)\right]^{-1},

    where

    β4k=[σk+(μk+−μk−)σk−−3((σk+)2−(σk−)2)]2−(μk+−μk−)2.\beta_4^k=\left[\frac{\sigma_k^+(\mu_k^+-\mu_k^-)}{\sigma_k^-}-3\big((\sigma_k^+)^2-(\sigma_k^-)^2\big)\right]^2-(\mu_k^+-\mu_k^-)^2.

    NAAF uses the initial penalty during early training and adjusts it toward this value later. The intended effect is to balance stronger separation of mismatched fragments against the risk of incorrectly treating matched fragments as mismatches.

  4. Knowl 4 — Negative attention from mismatched words

    model/method

    Let a sentence contain word features U={ui}i=1mU=\{u_i\}_{i=1}^m, an image contain region features V={vj}j=1nV=\{v_j\}_{j=1}^n, and each feature lie in Rd\mathbb{R}^d. NAAF first computes cosine relevance between word uiu_i and region vjv_j:

    sij=uivjT∥ui∥ ∥vj∥.s_{ij}=\frac{u_i v_j^{\mathsf T}}{\|u_i\|\,\|v_j\|}.

    Using the learned mismatch boundary tkt_k, the strongest adjusted cross-modal relevance for word uiu_i is

    si=max⁡1≤j≤n(sij−tk).s_i=\max_{1\leq j\leq n}(s_{ij}-t_k).

    The negative effect of the word is sineg=si Mask⁡neg(si)s_i^{\mathrm{neg}}=s_i\,\operatorname{Mask}_{\mathrm{neg}}(s_i), where the mask equals 11 when its scalar input is negative and 00 otherwise. Thus, words whose best image-region match remains below the learned boundary contribute a negative score, while matched words contribute no negative term.

    To account for semantic relationships among words, NAAF propagates matching degrees within the sentence. For a temperature or scaling factor λ>0\lambda>0,

    wilintra=softmax⁡λ({uiulT∥ui∥ ∥ul∥}l=1m),s^i=∑l=1mwilintrasl.w_{il}^{\mathrm{intra}}=\operatorname{softmax}_{\lambda}\left(\left\{\frac{u_i u_l^{\mathsf T}}{\|u_i\|\,\|u_l\|}\right\}_{l=1}^{m}\right),\qquad \hat{s}_i=\sum_{l=1}^{m}w_{il}^{\mathrm{intra}}s_l.

    The propagated value s^i\hat{s}_i replaces sis_i when the negative effect is calculated at inference time, so semantically related words receive consistent mismatch information.

  5. Knowl 5 — Positive attention from matched fragments

    model/method

    For each word uiu_i, NAAF's positive branch masks out image regions whose adjusted relevance sij−tks_{ij}-t_k is nonpositive. Define Mask⁡pos(x)=x\operatorname{Mask}_{\mathrm{pos}}(x)=x if x>0x>0 and Mask⁡pos(x)=−∞\operatorname{Mask}_{\mathrm{pos}}(x)=-\infty otherwise. The cross-modal attention weights are

    wijinter=softmax⁡λ({Mask⁡pos(sij−tk)}j=1n),w_{ij}^{\mathrm{inter}}=\operatorname{softmax}_{\lambda}\left(\left\{\operatorname{Mask}_{\mathrm{pos}}(s_{ij}-t_k)\right\}_{j=1}^{n}\right),

    where λ>0\lambda>0 is the scaling factor. The attended visual feature and feature-level similarity for word uiu_i are

    v^i=∑j=1nwijintervj,sif=uiv^iT∥ui∥ ∥v^i∥.\hat{v}_i=\sum_{j=1}^{n}w_{ij}^{\mathrm{inter}}v_j,\qquad s_i^f=\frac{u_i\hat{v}_i^{\mathsf T}}{\|u_i\|\,\|\hat{v}_i\|}.

    NAAF also aggregates the raw word-region similarities. Let [x]+=max⁡(x,0)[x]_+=\max(x,0) and define

    sˉij=[sij]+∑r=1m[srj]+2,wijrelev=softmax⁡λ({sˉij}j=1n),\bar{s}_{ij}=\frac{[s_{ij}]_+}{\sqrt{\sum_{r=1}^{m}[s_{rj}]_+^2}},\qquad w_{ij}^{\mathrm{relev}}=\operatorname{softmax}_{\lambda}(\{\bar{s}_{ij}\}_{j=1}^{n}),

    where rr indexes words. The relevance-level positive score is sir=∑j=1nwijrelevsijs_i^r=\sum_{j=1}^{n}w_{ij}^{\mathrm{relev}}s_{ij}, and the total positive contribution of word uiu_i is sipos=sif+sirs_i^{\mathrm{pos}}=s_i^f+s_i^r. For the complete pair, NAAF jointly combines both branches:

    S(U,V)=1m∑i=1m(sineg+sipos).S(U,V)=\frac{1}{m}\sum_{i=1}^{m}\left(s_i^{\mathrm{neg}}+s_i^{\mathrm{pos}}\right).

    Here S(U,V)S(U,V) is the scalar image-text similarity, with negative mismatches lowering the score and matched fragments raising it.

  6. Knowl 6 — Pseudo-label sampling for fragment distributions

    algorithm

    NAAF has no word-region alignment annotations, so it constructs pseudo-similarity samples from image-text instance-level labels during training. For a ground-truth sentence-image pair, let vj+v_j^+ denote regions of the correct image; for a mismatched pair, let vj−v_j^- denote regions of an incorrect image. For every word feature uiu_i, it samples the largest cosine similarity in each case:

    si+=max⁡1≤j≤nvj+uiT∥vj+∥ ∥ui∥,si−=max⁡1≤j≤nvj−uiT∥vj−∥ ∥ui∥.s_i^+=\max_{1\leq j\leq n}\frac{v_j^+u_i^{\mathsf T}}{\|v_j^+\|\,\|u_i\|},\qquad s_i^-=\max_{1\leq j\leq n}\frac{v_j^-u_i^{\mathsf T}}{\|v_j^-\|\,\|u_i\|}.

    The positive maximum represents the assumption that every word in a correctly paired description has at least one matching region. Although all regions in an incorrect image are mismatched for a word, the largest incorrect-image similarity is used because it is the upper bound and therefore the most discriminative mismatched value. These positive and negative samples update the Gaussian distributions for each training mini-batch; the paper updates them only when the current instance-level similarity ranking is judged correct. Sampling and distribution updating are training-only operations.

  7. Knowl 7 — Hard-negative bidirectional ranking objective

    equation

    NAAF is trained end-to-end with a bidirectional triplet ranking loss. For a ground-truth pair (U,V)(U,V), let S(U,V)S(U,V) be the NAAF similarity, let γ>0\gamma>0 be the margin, and select the hardest incorrect image and sentence as

    V′=arg max⁡p≠V  S(U,p),U′=arg max⁡q≠U  S(q,V).V'=\underset{p\ne V}{\operatorname{arg\,max}}\;S(U,p),\qquad U'=\underset{q\ne U}{\operatorname{arg\,max}}\;S(q,V).

    The loss over ground-truth pairs is

    L=∑(U,V)[γ−S(U,V)+S(U,V′)]++[γ−S(U,V)+S(U′,V)]+,L=\sum_{(U,V)}\left[\gamma-S(U,V)+S(U,V')\right]_+ + \left[\gamma-S(U,V)+S(U',V)\right]_+,

    where [x]+=max⁡(x,0)[x]_+=\max(x,0). The first hinge term requires the correct image to outrank the hardest incorrect image for sentence UU; the second requires the correct sentence to outrank the hardest incorrect sentence for image VV.

  8. Knowl 8 — Visual and textual feature construction

    experimental setup

    For each image, NAAF uses bottom-up visual attention: a Faster R-CNN detector pretrained on Visual Genome proposes the top K=36K=36 salient regions. Mean-pooled convolutional features from pretrained ResNet-101 are mapped by a fully connected layer to d=1024d=1024 dimensions, yielding region features vj∈R1024v_j\in\mathbb{R}^{1024}. For each sentence, every word is represented by a one-hot vector, embedded with pretrained GloVe, and processed by a bidirectional gated recurrent unit. The final word feature ui∈R1024u_i\in\mathbb{R}^{1024} is the average of the forward and backward hidden states.

    Experiments use Flickr30K with 31,000 images and 155,000 sentences, split into 29,000 training, 1,000 validation, and 1,000 test images, and MS-COCO with 123,287 images and 616,435 sentences, split into 113,287 training, 5,000 validation, and 5,000 test images. MS-COCO is evaluated both by averaging five 1,000-image test folds and on the full 5,000-image test set. Retrieval is measured by image-to-text and text-to-image R@1R@1, R@5R@5, and R@10R@10, where R@KR@K is the percentage of queries whose ground-truth item occurs in the top KK results; rSumrSum is the sum of all six recall values. Training uses Adam, initial learning rate 0.00050.0005 with a 10% decay every 10 epochs, 20 epochs, batch sizes 128 for Flickr30K and 256 for MS-COCO, scaling factor λ=20\lambda=20, initial penalty α=2.0\alpha=2.0, penalty adjustment at epoch 15, and ranking margin γ=0.2\gamma=0.2.

  9. Knowl 9 — State-of-the-art retrieval performance

    empirical result

    NAAF substantially improves image-text retrieval on both benchmarks under the stated recall metrics. On Flickr30K, it obtains image-to-text recalls (R@1,R@5,R@10)=(81.9,96.1,98.3)(R@1,R@5,R@10)=(81.9,96.1,98.3) and text-to-image recalls (61.0,85.3,90.6)(61.0,85.3,90.6), giving rSum=513.2rSum=513.2. The strongest listed prior model, SGRAF, obtains (77.8,94.1,97.4)(77.8,94.1,97.4) for image-to-text, (58.5,83.0,88.8)(58.5,83.0,88.8) for text-to-image, and rSum=499.6rSum=499.6.

    On the MS-COCO 1K test evaluation, NAAF obtains image-to-text (80.5,96.5,98.8)(80.5,96.5,98.8), text-to-image (64.1,90.7,96.5)(64.1,90.7,96.5), and rSum=527.2rSum=527.2, compared with the strongest listed prior rSum=524.3rSum=524.3. On the full MS-COCO 5K test set, NAAF obtains image-to-text (58.9,85.2,92.0)(58.9,85.2,92.0), text-to-image (42.5,70.9,81.4)(42.5,70.9,81.4), and rSum=430.9rSum=430.9. The gains, especially in R@1R@1, support the claim that explicitly penalizing mismatched textual fragments reduces false-positive retrievals while retaining positive matching evidence.

  10. Knowl 10 — Ablation evidence for mismatch modeling

    empirical result

    Flickr30K ablations show that both negative matching and the mechanisms used to make it discriminative are important. The full NAAF model reaches image-to-text (81.9,96.1,98.3)(81.9,96.1,98.3), text-to-image (61.0,85.3,90.6)(61.0,85.3,90.6). Removing the negative branch reduces these values to (66.2,91.0,96.2)(66.2,91.0,96.2) and (53.4,80.0,87.8)(53.4,80.0,87.8), while removing the positive/negative masks gives (74.2,93.0,96.4)(74.2,93.0,96.4) and (57.2,82.9,88.4)(57.2,82.9,88.4). Removing intra-text propagation gives (79.3,96.1,98.0)(79.3,96.1,98.0) and (59.2,83.9,90.0)(59.2,83.9,90.0); removing GloVe gives (75.9,93.6,97.7)(75.9,93.6,97.7) and (55.5,81.0,87.9)(55.5,81.0,87.9).

    With initial mismatch penalty α=2.0\alpha=2.0 and later adjustment toward α∗\alpha^*, NAAF reaches (79.6,96.3,98.3)(79.6,96.3,98.3) for image-to-text and (59.3,83.9,90.2)(59.3,83.9,90.2) for text-to-image in the single-model ablation. Omitting the adjustment lowers performance to (78.7,95.4,97.6)(78.7,95.4,97.6) and (59.3,83.4,89.7)(59.3,83.4,89.7), indicating that overly aggressive mismatch mining can misclassify matched fragments. Using batch sizes 32, 64, and 128 produces image-to-text R@1R@1 values of 76.076.0, 78.378.3, and 79.679.6, respectively; omitting the maximum-based negative sampling produces R@1=76.3R@1=76.3, supporting the proposed distribution-update and upper-bound sampling strategies.

Coverage note — The paper's qualitative false-positive visualizations and individual word-region case studies were omitted because they provide supporting examples rather than additional load-bearing method, theory, or quantitative results; no explicit limitation was stated by the paper.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR, pages 6077–6086, 2018. 6
  2. 2.Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval. In CVPR, pages 12655–12663, 2020. 1, 2, 3, 6, 7
  3. 3.Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, and Changhu Wang. Learning the best pooling strategy for visual semantic embedding. In CVPR, pages 15789–15798, 2021. 1, 3, 6, 7
  4. 4.Tianlang Chen, Jiajun Deng, and Jiebo Luo. Adaptive offline quintuplet loss for image-text matching. In ECCV, pages 549–565, 2020. 2
  5. 5.Tianlang Chen and Jiebo Luo. Expressing objects just like words: Recurrent visual embedding for image-text matching. In AAAI, volume 34, pages 10583–10590, 2020. 1, 2, 3, 6, 7
  6. 6.Hui Cui, Lei Zhu, Jingjing Li, Yang Yang, and Liqiang Nie. Scalable deep hashing for large-scale social image retrieval. In IEEE Transactions on image processing, volume 29, pages 1271–1284, 2019. 2
  7. 7.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching. In AAAI, 2021. 1, 2, 3, 6, 7, 8
  8. 8.Aviv Eisenschtat and Lior Wolf. Linking image and text with 2-way nets. In CVPR, pages 4601–4611, 2017. 2
  9. 9.Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. Vse++: Improving visual-semantic embeddings with hard negatives. In arXiv preprint arXiv:1707.05612, 2017. 1, 2
  10. 10.Xuri Ge, Fuhai Chen, Joemon M Jose, Zhilong Ji, Zhongqin Wu, and Xiao Liu. Structured multi-modal feature embedding and alignment for image-sentence retrieval. In ACM MM, pages 5185–5193, 2021. 6, 7
  11. 11.Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang. Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models. In CVPR, pages 7181–7189, 2018. 3
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 6
  13. 13.Zhibin Hu, Yongsheng Luo, Jiong Lin, Yan Yan, and Jian Chen. Multi-level visual-semantic alignments with relation-wise dual attention network for image and text matching. In IJCAI, pages 789–795, 2019. 1, 2, 3
  14. 14.Yan Huang and Liang Wang. Acmm: Aligned cross-modal memory for few-shot image and sentence matching. In ICCV, pages 5774–5783, 2019. 3
  15. 15.Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang. Learning semantic concepts and order for image and sentence matching. In CVPR, pages 6163–6171, 2018. 3, 5
  16. 16.Zhong Ji, Kexin Chen, and Haoran Wang. Step-wise hierarchical alignment network for image-text matching. In IJCAI, 2021. 1, 2, 3, 6, 7
  17. 17.Zhong Ji, Haoran Wang, Jungong Han, and Yanwei Pang. Saliency-guided attention network for image-sentence matching. In ICCV, pages 5754–5763, 2019. 3
  18. 18.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, pages 3128–3137, 2015. 2, 5, 6
  19. 19.Andrej Karpathy, Armand Joulin, and Li Fei-Fei. Deep fragment embeddings for bidirectional image sentence mapping. In NeurIPS, 2014. 1, 3, 5, 6
  20. 20.Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. Unifying visual-semantic embeddings with multimodal neural language models. ICML, 2014. 3
  21. 21.Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf. Associating neural word embeddings with deep image representations using fisher vectors. In CVPR, pages 4437–4446, 2015. 2
  22. 22.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalanditis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. In IJCV, volume 123, pages 32–73, 2017. 6
  23. 23.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In ECCV, pages 201–216, 2018. 1, 2, 3, 6, 7, 8
  24. 24.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In ICCV, pages 4654–4662, 2019. 1, 3
  25. 25.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 6
  26. 26.Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang. Focus your attention: A bidirectional focal attention network for image-text matching. In ACM MM, pages 3–11, 2019. 1, 2, 3, 6, 7, 8
  27. 27.Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. Graph structured network for image-text matching. In CVPR, pages 10921–10930, 2020. 1, 2, 3, 6, 7, 8
  28. 28.Xu Lu, Lei Zhu, Zhiyong Cheng, Liqiang Nie, and Huaxiang Zhang. Online multi-modal hashing with dynamic query-adaption. In ACM SIGIR, pages 715–724, 2019. 2
  29. 29.Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li. Multimodal convolutional neural networks for matching image and sentence. In ICCV, pages 2623–2631, 2015. 3
  30. 30.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, pages 1532–1543, 2014. 6
  31. 31.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV, pages 2641–2649, 2015. 6
  32. 32.Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. Dynamic modality interaction modeling for image-text retrieval. In ACM SIGIR, pages 1104–1113, 2021. 3
  33. 33.Botian Shi, Lei Ji, Pan Lu, Zhendong Niu, and Nan Duan. Knowledge aware semantic concept expansion for image-text matching. In IJCAI, volume 1, page 2, 2019. 3
  34. 34.Yale Song and Mohammad Soleymani. Polysemous visual-semantic embedding for cross-modal retrieval. In CVPR, pages 1979–1988, 2019. 7
  35. 35.Haoran Wang, Ying Zhang, Zhong Ji, Yanwei Pang, and Lin Ma. Consensus-aware visual-semantic embedding for image-text matching. In ECCV, pages 18–34. Springer, 2020. 6, 7
  36. 36.Liwei Wang, Yin Li, and Svetlana Lazebnik. Learning deep structure-preserving image-text embeddings. In CVPR, pages 5005–5013, 2016. 2
  37. 37.Sijin Wang, Ruiping Wang, Ziwei Yao, Shiguang Shan, and Xilin Chen. Cross-modal scene graph matching for relationship-aware image-text retrieval. In WACV, pages 1508–1517, 2020. 3, 6, 7
  38. 38.Tan Wang, Xing Xu, Yang Yang, Alan Hanjalic, Heng Tao Shen, and Jingkuan Song. Matching images and text with multi-modal tensor fusion and re-ranking. In ACM MM, pages 12–20, 2019. 3
  39. 39.Yaxiong Wang, Hao Yang, Xueming Qian, Lin Ma, Jing Lu, Biao Li, and Xin Fan. Position focused attention network for image-text matching. In IJCAI, pages 3792–3798, 2019. 1, 2, 3
  40. 40.Jiwei Wei, Xing Xu, Yang Yang, Yanli Ji, Zheng Wang, and Heng Tao Shen. Universal weighting metric learning for cross-modal matching. In CVPR, pages 13005–13014, 2020. 2
  41. 41.Yiling Wu, Shuhui Wang, and Qingming Huang. Online asymmetric similarity learning for cross-modal retrieval. In CVPR, pages 4269–4278, 2017. 2
  42. 42.Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang. Learning fragment self-attention embeddings for image-text matching. In ACM MM, pages 2088–2096, 2019. 3
  43. 43.Shiyang Yan, Li Yu, and Yuan Xie. Discrete-continuous action space policy gradient-based attention for image-text matching. In CVPR, pages 8096–8105, 2021. 1
  44. 44.Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z Li. Context-aware attention network for image-text retrieval. In CVPR, pages 3536–3545, 2020. 3
  45. 45.Mo Zhou, Zhenxing Niu, Le Wang, Zhanning Gao, Qilin Zhang, and Gang Hua. Ladder loss for coherent visual-semantic embedding. In AAAI, volume 34, pages 13050–13057, 2020. 2
  46. 46.Lei Zhu, Xu Lu, Zhiyong Cheng, Jingjing Li, and Huaxiang Zhang. Deep collaborative multi-view hashing for large-scale image search. In IEEE Transactions on Image Processing, volume 29, pages 4643–4655, 2020. 2

Citation

MLA
Zhang, K., et al. “Negative-Aware Attention Framework for Image-Text Matching”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15640–49, https://doi.org/10.1109/CVPR52688.2022.01521.
APA
Zhang, K., Mao, Z., Wang, Q., & Zhang, Y. (2022). Negative-Aware Attention Framework for Image-Text Matching. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15640–15649. https://doi.org/10.1109/CVPR52688.2022.01521
Chicago
Zhang, K., Z. Mao, Q. Wang, and Y. Zhang. 2022. “Negative-Aware Attention Framework for Image-Text Matching”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15640–49. https://doi.org/10.1109/CVPR52688.2022.01521.
Harvard
Zhang, K. et al. (2022) “Negative-Aware Attention Framework for Image-Text Matching”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 15640–15649. Available at: https://doi.org/10.1109/CVPR52688.2022.01521.
Vancouver
1. Zhang K, Mao Z, Wang Q, Zhang Y (2022) Negative-Aware Attention Framework for Image-Text Matching. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 15640–15649

BibTeX

@inproceedings{Zhang_2022, title={Negative-Aware Attention Framework for Image-Text Matching}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01521}, DOI={10.1109/cvpr52688.2022.01521}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhang, Kun and Mao, Zhendong and Wang, Quan and Zhang, Yongdong}, year={2022}, month=June, pages={15640–15649} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE