Fine-grained Image-text Matching by Cross-modal Hard Aligning Network

Zhengxin PanFangyu WuBailing Zhang

article2023CVPR126 citations

Proposes a cross-modal hard aligning network that reframes fine-grained image-text matching as a hard assignment coding problem, eliminating noisy region-word associations to improve retrieval accuracy while substantially cutting memory and computation costs compared to standard cross-attention methods.

Listen

Modern digital systems increasingly require automated tools to match visual content with descriptive text across massive multimedia collections. Existing fine-grained image-text retrieval methods typically rely on cross-attention mechanisms to align textual words with image regions. However, standard cross-attention introduces noisy, irrelevant alignments and requires calculating dense attention matrices, which creates severe computational bottlenecks, degrades retrieval accuracy, and leads to excessive memory and latency overhead in practical applications.

The article demonstrates that cross-modal fragment alignment can be reframed through an information coding perspective, evaluating whether hard assignment coding can replace soft cross-attention to deliver higher retrieval accuracy and substantially greater computational efficiency. To validate this concept, the authors developed the Cross-modal Hard Aligning Network (CHAN), which treats sentence words as queries and image regions as visual codewords, retaining only the single most relevant region-word alignment while discarding all redundant pairs. The approach was evaluated through extensive experimental benchmarks on the standard MS-COCO and Flickr30K datasets across both bidirectional image-to-text and text-to-image retrieval tasks.

The findings confirm three critical results. First, CHAN significantly outperforms previous state-of-the-art methods in retrieval accuracy on both benchmarks, achieving an RSUM of 518.5 on Flickr30K and 532.6 on MS-COCO 5-fold 1K using a standard language model backbone without requiring ensemble modeling. Second, CHAN provides over a tenfold speedup in total inference time compared to recent competing alignment architectures and is more than three times faster than standard baseline implementations. Third, ablation testing reveals that querying visual codebooks with text queries combined with LogSumExp pooling yields optimal alignment quality, and increasing the number of visual regions steadily improves accuracy without the degradation seen in older models.

These results demonstrate that dense soft alignments are largely redundant for cross-modal similarity matching. Retaining only the primary corresponding fragment reduces algorithmic memory complexity and eliminates the need for iterative batch processing during inference. Consequently, deployment of this hard aligning framework offers immediate performance gains, lower cloud infrastructure costs, and reduced latency for cross-modal search platforms. Organizations deploying visual-text search systems should consider transitioning from soft cross-attention architectures to hard assignment alignment mechanisms.

Future research should expand this coding framework toward information-theoretic objectives such as maximizing mutual information between modalities. While confidence in the reported experimental results is high across established benchmark datasets, real-world implementations should conduct domain-specific testing to confirm that visual object extractors capture sufficient region granularity when processing complex, cluttered, or out-of-domain imagery.

  • Paper: Stacked Cross Attention for Image-Text Matching, Kuang-Huei Lee et al. (2018). Lee et al. established the standard stacked cross-attention paradigm for image-text matching, providing the baseline architecture and dense alignment problem that CHAN replaces with hard assignment coding.
  • Paper: Negative-Aware Attention Framework for Image-Text Matching, Kun Zhang et al. (2022). Zhang et al. address the pitfalls of fine-grained cross-modal attention by introducing negative penalties for mismatched fragments, motivating CHAN's approach to eliminating redundant and noisy region-word pairs.
  • Paper: Cross-Modal Discrete Representation Learning, Alexander H. Liu et al. (2022). Liu et al. formulate cross-modal matching through shared discrete codebooks and vector quantization, establishing the conceptual foundation for treating cross-modal fragment alignment as an information coding problem.
  • Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). Karpathy and Fei-Fei introduced the foundational framework for learning fine-grained latent alignments between visual image regions and descriptive sentence fragments for cross-modal retrieval.
  • Paper: Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics, Micah Hodosh et al. (2013). Hodosh et al. formalized cross-modal image description and retrieval as a ranking task evaluated on benchmark datasets like Flickr, creating the standard evaluation framework used in CHAN.
Cover for Fine-grained Image-text Matching by Cross-modal Hard Aligning Network

Abstract

Current state-of-the-art image-text matching methods implicitly align the visual-semantic fragments, like regions in images and words in sentences, and adopt cross-attention mechanism to discover fine-grained cross-modal semantic correspondence. However, the cross-attention mechanism may bring redundant or irrelevant region-word alignments, degenerating retrieval accuracy and limiting efficiency. Although many researchers have made progress in mining meaningful alignments and thus improving accuracy, the problem of poor efficiency remains unresolved. In this work, we propose to learn fine-grained image-text matching from the perspective of information coding. Specifically, we suggest a coding framework to explain the fragments aligning process, which provides a novel view to reexamine the cross-attention mechanism and analyze the problem of redundant alignments. Based on this framework, a Cross-modal Hard Aligning Network (CHAN) is designed, which comprehensively exploits the most relevant region-word pairs and eliminates all other alignments. Extensive experiments conducted on two public datasets, MS-COCO and Flickr30K, verify that the relevance of the most associated word-region pairs is discriminative enough as an indicator of the image-text similarity, with superior accuracy and efficiency over the state-of-the-art approaches on the bidirectional image and text retrieval tasks. Our code will be available at https://github.com/ppanzx/CHAN.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Cross-modal Hard Aligning Network
  • 3.1. Coding Framework for Fragment Alignment
  • 3.2. Hard Assignment Coding
  • 3.3. Cross-modal Hard Alignment Network
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Comparison Results
  • 4.3. Ablation Study
  • 4.4. Visualization and Case Study
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Coding framework for fine-grained image-text matching

    model/method

    CHAN formulates fine-grained image-text matching as information coding. A sentence is represented by word features T={ti∣i=1,…,L}T=\{t_i\mid i=1,\ldots,L\} with ti∈Rdt_i\in\mathbb{R}^d, and an image is represented by salient-region features V={vj∣j=1,…,K}V=\{v_j\mid j=1,\ldots,K\} with vj∈Rdv_j\in\mathbb{R}^d. Each word feature tit_i is treated as a query and each image-region feature vjv_j as a visual codeword.

    The word-region similarity is cosine similarity,

    sij=ti⊤vj∥ti∥ ∥vj∥.s_{ij}=\frac{t_i^{\top}v_j}{\lVert t_i\rVert\,\lVert v_j\rVert}.

    A general coding mechanism reconstructs an attended query as t^i=∑j=1Kωijvj\hat t_i=\sum_{j=1}^{K}\omega_{ij}v_j, where ωij\omega_{ij} is the assignment weight of region vjv_j to word tit_i. The word-image similarity is s(ti,V)=S(ti,t^i)s(t_i,V)=S(t_i,\hat t_i), where SS is cosine similarity. Word-image similarities are aggregated with LogSumExp pooling:

    s(T,V)=1λlog⁡(∑i=1Lexp⁡(λs(ti,V))),s(T,V)=\frac{1}{\lambda}\log\left(\sum_{i=1}^{L}\exp\bigl(\lambda s(t_i,V)\bigr)\right),

    where λ>0\lambda>0 controls the emphasis on highly relevant word-image pairs. This framework interprets conventional cross-attention as soft assignment coding because it combines multiple visual codewords to reconstruct each word query.

  2. Knowl 2 — Hard assignment coding retains only the best region for each word

    theoretical result

    Under the CHAN assumption that a semantically matched image contains a region that best describes each word, the most relevant codeword for word tit_i is vkiv_{k_i}, where ki∈{1,…,K}k_i\in\{1,\ldots,K\} satisfies ki=arg⁡max⁡jsijk_i=\arg\max_j s_{ij}. Hard assignment coding sets the assignment weights to

    ωij={1,j=ki,0,j≠ki.\omega_{ij}=\begin{cases} 1,&j=k_i,\\ 0,&j\ne k_i. \end{cases}

    This is the limiting form of the softmax assignment ωij=exp⁡(sij/τ)/∑j′=1Kexp⁡(sij′/τ)\omega_{ij}=\exp(s_{ij}/\tau)/\sum_{j'=1}^{K}\exp(s_{ij'}/\tau) as the temperature τ>0\tau>0 approaches zero. The resulting word-image similarity is exactly the maximum word-region similarity:

    s(ti,V)=ti⊤vki∥ti∥ ∥vki∥=siki=max⁡j=1,…,Ksij.s(t_i,V)=\frac{t_i^{\top}v_{k_i}}{\lVert t_i\rVert\,\lVert v_{k_i}\rVert}=s_{ik_i}=\max_{j=1,\ldots,K}s_{ij}.

    Thus, hard assignment discards redundant region-word alignments, avoids constructing an attended query from irrelevant regions, and uses the single most informative region to determine each word's contribution.

  3. Knowl 3 — Probabilistic lower-bound interpretation of hard alignment

    theoretical result

    CHAN gives hard assignment coding a probabilistic interpretation. Let P(ti,V)P(t_i,V) denote the probability that word tit_i is semantically present in image codebook VV, and let P(ti,vj)P(t_i,v_j) denote the semantic-consistency probability between word tit_i and region vjv_j. The model assumes these probabilities are proportional to the corresponding similarities: P(ti,V)∝s(ti,V)P(t_i,V)\propto s(t_i,V) and P(ti,vj)∝sijP(t_i,v_j)\propto s_{ij}.

    For a subset of RR mutually independent codewords that includes the best codeword vkiv_{k_i}, the semantic-presence probability satisfies

    P(ti,V)=1−∏j=1R(1−P(ti,vj))≥P(ti,vki).P(t_i,V)=1-\prod_{j=1}^{R}\bigl(1-P(t_i,v_j)\bigr)\ge P(t_i,v_{k_i}).

    Consequently, the most relevant word-region pair provides a lower bound on the probability that the word's semantics occur in the image. CHAN uses this most informative pair as a discriminative approximation instead of combining potentially dependent and irrelevant codewords through soft assignment.

  4. Knowl 4 — Cross-modal Hard Aligning Network architecture

    model/method

    CHAN contains visual representation, text representation, hard assignment coding, and ranking-loss modules. The visual branch extracts the top-KK region features from each image using Faster R-CNN pretrained on Visual Genome with bottom-up/top-down attention. A fully connected layer maps each region to dd dimensions, and a self-attention layer injects contextual information before the KK regions form the visual codebook V={vj}j=1KV=\{v_j\}_{j=1}^{K}.

    CHAN supports two text branches. In the BiGRU version, each sentence is tokenized, its words are initialized with pretrained GloVe embeddings, and a bidirectional GRU produces one query vector per word by averaging the forward and backward hidden states. In the BERT version, word-level vectors from the last layer of pretrained BERT are projected through a fully connected layer into the common dd-dimensional space.

    After ℓ2\ell_2-normalizing all word queries and visual codewords, CHAN computes the cosine-similarity matrix S∈RL×KS\in\mathbb{R}^{L\times K} by S=TV⊤S=TV^{\top}. It applies row-wise maximum pooling to obtain max⁡j=1,…,KSij\max_{j=1,\ldots,K}S_{ij} for every word and then applies LogSumExp pooling across words:

    s(T,V)=1λlog⁡(∑i=1Lexp⁡(λmax⁡j=1,…,KSij)).s(T,V)=\frac{1}{\lambda}\log\left(\sum_{i=1}^{L}\exp\left(\lambda\max_{j=1,\ldots,K}S_{ij}\right)\right).

    The architecture overview on page 5 depicts this process as a similarity matrix followed by row-wise max-pooling and word-level LogSumExp pooling.

  5. Knowl 5 — Bidirectional hard-negative triplet objective

    equation

    CHAN is trained with a bidirectional hinge-based triplet ranking loss and online hard-negative mining. Let DD be the training set of matched image-sentence pairs (T,V)(T,V), let α>0\alpha>0 be the margin, and let [x]+=max⁡(x,0)[x]_+=\max(x,0). Using the CHAN similarity score s(T,V)s(T,V), define the hardest mismatched image and sentence within the current minibatch as

    V^=arg⁡max⁡V′≠Vs(T,V′),T^=arg⁡max⁡T′≠Ts(T′,V).\hat V=\arg\max_{V'\ne V}s(T,V'),\qquad \hat T=\arg\max_{T'\ne T}s(T',V).

    The optimization objective is

    L=∑(T,V)∼D([α+s(T,V^)−s(T,V)]++[α+s(T^,V)−s(T,V)]+).\mathcal{L}=\sum_{(T,V)\sim D}\left([\alpha+s(T,\hat V)-s(T,V)]_+ + [\alpha+s(\hat T,V)-s(T,V)]_+\right).

    The first term separates a matched sentence-image pair from its hardest competing image, while the second separates it from its hardest competing sentence.

  6. Knowl 6 — Memory and computation advantage over soft assignment

    theoretical result

    Suppose a batch contains B1B_1 images represented by KK visual regions in dd dimensions and B2B_2 captions represented by LL word queries in dd dimensions. Both hard and soft assignment compute an image-caption assignment matrix A∈RB1×B2×K×LA\in\mathbb{R}^{B_1\times B_2\times K\times L}, requiring time O(B1B2KLd)O(B_1B_2KLd).

    Soft assignment additionally constructs attended text features T^∈RB1×B2×L×d\hat T\in\mathbb{R}^{B_1\times B_2\times L\times d}, giving spatial complexity O(B1B2Ld)O(B_1B_2Ld). Hard assignment needs only the assignment matrix, with spatial complexity O(B1B2KL)O(B_1B_2KL). Because the number of visual regions is much smaller than the embedding dimension in the stated setting (K≪dK\ll d), hard assignment substantially reduces memory consumption. It also eliminates attended-query construction and the repeated iterations that soft-attention retrieval may require when memory cannot hold all cross-modal attention weights.

  7. Knowl 7 — Retrieval datasets and evaluation protocol

    experimental setup

    CHAN is evaluated on MS-COCO and Flickr30K for bidirectional image-text retrieval. MS-COCO contains 123,287 images with five captions per image; the split uses 113,287 training images, 5,000 validation images, and 5,000 test images. Results are reported both by averaging over five 1K-image test folds and on the full 5K-image test set. Flickr30K contains 31,783 images with five captions per image; 1,014 images are used for validation, 1,000 for testing, and the remainder for training.

    Retrieval performance is measured by R@KR@K, the percentage of queries whose correct match appears among the top KK retrieved instances. Higher values are better. The paper also reports RSUM\mathrm{RSUM}, the sum of the six recalls from image-to-text and text-to-image retrieval at K∈{1,5,10}K\in\{1,5,10\}. The main comparisons use single CHAN models rather than ensembles, whereas several competing methods in the benchmark tables use two-model ensembles.

  8. Knowl 8 — CHAN improves retrieval accuracy on MS-COCO and Flickr30K

    data/table

    The reported benchmark results show that CHAN obtains the best or near-best retrieval scores while using a single model. The table gives the key reported metrics: image-to-text R@1R@1, text-to-image R@1R@1, and RSUM\mathrm{RSUM}, where RSUM\mathrm{RSUM} sums all six directional recalls. An asterisk marks a competing ensemble result; CHAN entries are single-model results.

    Could not parse LaTeX table

    On the full COCO 5K test set, CHAN-BiGRU exceeds NAAF by 2.5 RSUM\mathrm{RSUM} points, and CHAN-BERT exceeds the TERAN ensemble by 0.5 points. On Flickr30K, CHAN-BiGRU slightly exceeds NAAF and CHAN-BERT exceeds TERAN by 5.1 RSUM\mathrm{RSUM} points. The results support the claim that retaining only the most relevant region-word correspondence can improve matching accuracy despite discarding dense cross-attention alignments.

  9. Knowl 9 — Inference efficiency of hard alignment

    empirical result

    The inference-efficiency comparison on the COCO 5-fold 1K, COCO 5K, and Flickr30K test sets shows that both CHAN-BiGRU and CHAN-BERT achieve the strongest accuracy-efficiency trade-off among the compared fragment-aligning systems. The paper reports that the two CHAN variants are more than 10 times faster than other recent methods and approximately 3 times faster than a reimplemented SCAN model using the same BiGRU representation and soft assignment coding, while also obtaining the best accuracy on all three test sets.

    The comparison plot on page 6 places retrieval quality (RSUM\mathrm{RSUM}) against total inference time in milliseconds per image-caption pair. The result is attributed to hard alignment's avoidance of attended-query construction and its lower memory requirement, rather than to a reduction in the asymptotic cost of forming the region-word similarity matrix.

  10. Knowl 10 — Ablation evidence for codebook direction and pooling

    data/table

    Ablations on the full COCO 5K test set use BiGRU-based CHAN as the baseline and compare coding directions and pooling operators. Image regions used as the visual codebook with hard assignment outperform soft cross-attention and using words as the codebook. LogSumExp pooling is the strongest aggregation method, while max-pooling is particularly harmful because it retains only one word-level score and discards the contributions of other informative words.

    Could not parse LaTeX table

    Visual-codebook hard assignment improves RSUM\mathrm{RSUM} from 418.6 for cross-attention to 433.4, whereas textual-codebook hard assignment decreases it to 398.0. Among the pooling choices, LogSumExp obtains the highest RSUM\mathrm{RSUM} of 433.4; average, sum, and softmax pooling obtain 431.9, 428.1, and 420.5, respectively.

Coverage note — The page-8 qualitative attention-map case study, including the four word-region examples, is omitted because it provides illustrative visual confirmation of hard alignment rather than an additional method, theorem, or quantitative result.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR, pages 6077–6086, 2018. 5, 7
  2. 2.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021. 3
  3. 3.Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval. In CVPR, pages 12655–12663, 2020. 3, 7
  4. 4.Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, and Changhu Wang. Learning the best pooling strategy for visual semantic embedding. In CVPR, pages 15789–15798, 2021. 2, 3, 7
  5. 5.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015. 3, 6, 7
  6. 6.Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen, and Peilin Liu. Cross-modal graph matching network for image-text retrieval. TOMM, 18(4):1–23, 2022. 3, 7
  7. 7.Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. Attention-based models for speech recognition. NeurIPS, 28, 2015. 4
  8. 8.Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio De Rezende, Yannis Kalantidis, and Diane Larlus. Probabilistic embeddings for cross-modal retrieval. In CVPR, pages 8415–8424, 2021. 3, 5
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 5, 7
  10. 10.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching. Technical report, AAAI, 2021. 2, 5, 6, 7
  11. 11.Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. Vse++:improving visual-semantic embeddings with hard negatives. In BMVC, 2017. 3, 5, 6, 7
  12. 12.Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. Devise: A deep visual-semantic embedding model. NeurIPS, 26, 2013. 1, 3
  13. 13.Xuri Ge, Fuhai Chen, Joemon M Jose, Zhilong Ji, Zhongqin Wu, and Xiao Liu. Structured multi-modal feature embedding and alignment for image-sentence retrieval. In ACMMM, pages 5185–5193, 2021. 5
  14. 14.Jan C van Gemert, Jan-Mark Geusebroek, Cor J Veenman, and Arnold WM Smeulders. Kernel codebooks for scene categorization. In ECCV, pages 696–709. Springer, 2008. 2, 3, 4
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 7
  16. 16.Yan Huang, Wei Wang, and Liang Wang. Instance-aware image and sentence matching with selective multimodal lstm. In CVPR, pages 2310–2318, 2017. 2, 3
  17. 17.Yongzhen Huang, Zifeng Wu, Liang Wang, and Tieniu Tan. Feature coding in image classification: A comprehensive study. PAMI, 36(3):493–506, 2013. 2, 3
  18. 18.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, pages 4904–4916. PMLR, 2021. 3
  19. 19.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, pages 3128–3137, 2015. 2, 3
  20. 20.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 123(1):32–73, 2017. 5
  21. 21.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In ECCV, pages 201–216, 2018. 1, 2, 3, 4, 5, 6, 7
  22. 22.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In CVPR, pages 4654–4662, 2019. 2, 3, 7
  23. 23.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Image-text embedding learning via visual and textual semantic reasoning. PAMI, 2022. 1, 3, 7
  24. 24.Alex Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, and James Glass. Cross-modal discrete representation learning. In ACL, pages 3013–3035, 2022. 3
  25. 25.Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang. Focus your attention: A bidirectional focal attention network for image-text matching. In ACMMM, pages 3–11, 2019. 2
  26. 26.Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. Graph structured network for image-text matching. In CVPR, pages 10921–10930, 2020. 3
  27. 27.Lingqiao Liu, Lei Wang, and Xinwang Liu. In defense of soft-assignment coding. In ICCV, pages 2486–2493. IEEE, 2011. 2, 3, 4
  28. 28.Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, and Stéphane Marchand-Maillet. Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders. TOMM, 17(4):1–23, 2021. 6, 7
  29. 29.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, pages 1532–1543, 2014. 5
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021. 3
  31. 31.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28, 2015. 5
  32. 32.Josef Sivic and Andrew Zisserman. Video google: A text retrieval approach to object matching in videos. In ICCV, volume 3, pages 1470–1470, 2003. 2, 3
  33. 33.Thomas Theodoridis, Theocharis Chatzis, Vassilios Solachidis, Kosmas Dimitropoulos, and Petros Daras. Cross-modal variational alignment of latent spaces. In CVPRW, pages 960–961, 2020. 3
  34. 34.Jan C Van Gemert, Cor J Veenman, Arnold WM Smeulders, and Jan-Mark Geusebroek. Visual word ambiguity. PAMI, 32(7):1271–1283, 2009. 3, 4, 6
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 30, 2017. 5
  36. 36.Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. Order-embeddings of images and language. arXiv preprint arXiv:1511.06361, 2015. 3
  37. 37.Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen. Adversarial cross-modal retrieval. In ACMMM, pages 154–162, 2017. 3
  38. 38.Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik. Learning two-branch neural networks for image-text matching tasks. PAMI, 41(2):394–407, 2018. 3
  39. 39.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022. 3
  40. 40.Jiwei Wei, Yang Yang, Xing Xu, Xiaofeng Zhu, and Heng Tao Shen. Universal weighting metric learning for cross-modal retrieval. PAMI, 2021. 3
  41. 41.Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. Multi-modality cross attention network for image and sentence matching. In CVPR, pages 10941–10950, 2020. 3, 7
  42. 42.Max Welling and Thomas N Kipf. Semi-supervised classification with graph convolutional networks. In ICLR, 2016. 3
  43. 43.Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang. Learning fragment self-attention embeddings for image-text matching. In ACMMM, pages 2088–2096, 2019. 2, 3
  44. 44.Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. Vision-language pre-training with triple contrastive learning. In CVPR, pages 15671–15680, 2022. 3
  45. 45.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2:67–78, 2014. 3, 6
  46. 46.Kun Zhang, Zhendong Mao, Quan Wang, and Yongdong Zhang. Negative-aware attention framework for image-text matching. In CVPR, pages 15661–15670, 2022. 2, 3, 5, 6, 7

Citation

MLA
Pan, Z., et al. “Fine-grained Image-text Matching by Cross-modal Hard Aligning Network”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 19275–84, https://doi.org/10.1109/CVPR52729.2023.01847.
APA
Pan, Z., Wu, F., & Zhang, B. (2023). Fine-grained Image-text Matching by Cross-modal Hard Aligning Network. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19275–19284. https://doi.org/10.1109/CVPR52729.2023.01847
Chicago
Pan, Z., F. Wu, and B. Zhang. 2023. “Fine-grained Image-text Matching by Cross-modal Hard Aligning Network”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19275–84. https://doi.org/10.1109/CVPR52729.2023.01847.
Harvard
Pan, Z., Wu, F. and Zhang, B. (2023) “Fine-grained Image-text Matching by Cross-modal Hard Aligning Network”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 19275–19284. Available at: https://doi.org/10.1109/CVPR52729.2023.01847.
Vancouver
1. Pan Z, Wu F, Zhang B (2023) Fine-grained Image-text Matching by Cross-modal Hard Aligning Network. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 19275–19284

BibTeX

@inproceedings{Pan_2023, title={Fine-grained Image-text Matching by Cross-modal Hard Aligning Network}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01847}, DOI={10.1109/cvpr52729.2023.01847}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Pan, Zhengxin and Wu, Fangyu and Zhang, Bailing}, year={2023}, month=June, pages={19275–19284} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE