Stacked Cross Attention for Image-Text Matching

Kuang-Huei LeeXi ChenGang HuaHoudong HuXiaodong He

article2018ECCV1,409 citations

Proposes a stacked cross-attention network that aligns visual regions with corresponding sentence words, achieving significant gains in bidirectional image-text retrieval accuracy on MS-COCO and Flickr30K.

Listen

Cross-modal retrieval—the ability to search images using descriptive text queries and retrieve accurate text descriptions from image inputs—is a critical capability for modern search engines, multimedia databases, and vision-language systems. A fundamental challenge in this domain is accurately aligning visual elements, such as objects and background scenes, with corresponding words in descriptive sentences. Earlier systems either aggregated region-word similarities without considering context or used step-limited attention processes that restricted interpretability. The article addresses this challenge by introducing Stacked Cross Attention, a framework designed to fully capture latent semantic alignments between visual regions and individual words to infer overall image-text similarity.

The main objective of the article is to develop, evaluate, and demonstrate the Stacked Cross Attention Network (SCAN). The approach uses context from both images and sentences in two complementary formulations: Image-Text (which aligns sentence words to each visual region to determine region importance) and Text-Image (which aligns image regions to each word to determine word importance). To achieve this, salient visual regions are extracted using object detection models, while sentences are processed via bidirectional recurrent neural networks to capture word order and linguistic context. The model is trained using a triplet ranking loss that focuses on hard negative samples, and it is evaluated on standard benchmark datasets, specifically Flickr30K (31,000 images) and MS-COCO (123,287 images).

Across extensive empirical evaluations, the proposed method significantly outperforms existing state-of-the-art approaches. On the Flickr30K dataset, the model improves top-1 retrieval accuracy by 22.1% relatively for sentence retrieval and 18.2% relatively for image retrieval over previous leading methods. On the 5,000-image MS-COCO test set, an ensemble combining both attention formulations improves top-1 sentence retrieval by 17.8% relatively and image retrieval by 16.6% relatively. Ablation studies demonstrate that incorporating hard negative sampling during training delivers a massive performance boost—improving sentence top-1 recall by 48.2%—and confirm that bidirectional language processing consistently outperforms unidirectional alternatives.

These results indicate that fine-grained, contextual cross-attention significantly boosts accuracy and system transparency. By surfacing explicit attention alignments between words and image regions, the architecture allows system operators to inspect and interpret the underlying reasoning behind match decisions, reducing the risks associated with opaque multi-modal systems. Decision-makers and engineering teams seeking to optimize cross-modal search workflows should consider adopting stacked cross-attention mechanisms combined with hard negative training strategies. Further development should focus on improving the representation of dynamic interactions and actions, which remain challenging to extract from static visual features.

Cover for Stacked Cross Attention for Image-Text Matching

Abstract

In this paper, we study the problem of image-text matching. Inferring the latent semantic alignment between objects or other salient stuff (e.g. snow, sky, lawn) and the corresponding words in sentences allows to capture fine-grained interplay between vision and language, and makes image-text matching more interpretable. Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to capture limited number of semantic alignments which is less interpretable. In this paper, we present Stacked Cross Attention to discover the full latent alignments using both image regions and words in a sentence as context and infer image-text similarity. Our approach achieves the state-of-the-art results on the MS-COCO and Flickr30K datasets. On Flickr30K, our approach outperforms the current best methods by 22.1% relatively in text retrieval from image query, and 18.2% relatively in image retrieval with text query (based on Recall@1). On MS-COCO, our approach improves sentence retrieval by 17.8% relatively and image retrieval by 16.6% relatively (based on Recall@1 using the 5K test set). Code has been made available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Learning Alignments with Stacked Cross Attention
  • 3.1 Stacked Cross Attention
  • 3.2 Alignment Objective
  • 3.3 Representing images with Bottom-Up Attention
  • 3.4 Representing Sentences
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Results on Flickr30K
  • 4.3 Results on MS-COCO
  • 4.4 Ablation Studies
  • 5 Visualization and Analysis
  • 5.1 Visualizing Attention
  • 5.2 Image and Sentence Retrieval
  • 6 Conclusions
  • 0.A Details of Training
  • 0.B Details of Bottom-up Attention
  • 0.C Additional Examples
  • References

Knowls

  1. Knowl 1 — Image-Text Stacked Cross Attention Formulation

    model/method

    The Image-Text (i-t) Stacked Cross Attention formulation computes global image-sentence similarity in two consecutive attention stages, attending to sentence words with respect to each detected image region.

    Given an image II represented by kk region vectors V={v1,…,vk}⊂RhV = \{v_1, \dots, v_k\} \subset \mathbb{R}^h and a sentence TT represented by nn word vectors E={e1,…,en}⊂RhE = \{e_1, \dots, e_n\} \subset \mathbb{R}^h, the cosine similarity between every region-word pair is defined as:

    sij=viTej∥vi∥∥ej∥,i∈{1,…,k},  j∈{1,…,n}s_{ij} = \frac{v_i^T e_j}{\|v_i\| \|e_j\|}, \quad i \in \{1, \dots, k\}, \; j \in \{1, \dots, n\}

    The similarities are thresholded at zero with [x]+=max⁡(x,0)[x]_+ = \max(x, 0) and normalized across image regions:

    sˉij=[sij]+∑i=1k[sij]+2\bar{s}_{ij} = \frac{[s_{ij}]_+}{\sqrt{\sum_{i=1}^k [s_{ij}]_+^2}}

    In Stage 1, the attention weight of the jj-th word with respect to the ii-th region is obtained via softmax with inverse temperature parameter λ1\lambda_1:

    αij=exp⁡(λ1sˉij)∑j=1nexp⁡(λ1sˉij)\alpha_{ij} = \frac{\exp(\lambda_1 \bar{s}_{ij})}{\sum_{j=1}^n \exp(\lambda_1 \bar{s}_{ij})}

    The attended sentence representation ait∈Rha_i^t \in \mathbb{R}^h for region viv_i is the weighted sum of word features:

    ait=∑j=1nαijeja_i^t = \sum_{j=1}^n \alpha_{ij} e_j

    In Stage 2, the relevance score R(vi,ait)R(v_i, a_i^t) between region viv_i and the attended sentence context aita_i^t is computed via cosine similarity:

    R(vi,ait)=viTait∥vi∥∥ait∥R(v_i, a_i^t) = \frac{v_i^T a_i^t}{\|v_i\| \|a_i^t\|}

    The global image-text similarity S(I,T)S(I, T) is aggregated across all kk regions using either LogSumExp (LSE) pooling with scaling parameter λ2\lambda_2:

    SLSE(I,T)=log⁡(∑i=1kexp⁡(λ2R(vi,ait)))1/λ2S_{\text{LSE}}(I, T) = \log \left( \sum_{i=1}^k \exp\left(\lambda_2 R(v_i, a_i^t)\right) \right)^{1/\lambda_2}

    or average (AVG) pooling:

    SAVG(I,T)=1k∑i=1kR(vi,ait)S_{\text{AVG}}(I, T) = \frac{1}{k} \sum_{i=1}^k R(v_i, a_i^t)

  2. Knowl 2 — Text-Image Stacked Cross Attention Formulation

    model/method

    The Text-Image (t-i) Stacked Cross Attention formulation operates symmetrically to the Image-Text model by attending to salient image regions with respect to each word in the sentence.

    Given kk image region features V={v1,…,vk}⊂RhV = \{v_1, \dots, v_k\} \subset \mathbb{R}^h and nn word features E={e1,…,en}⊂RhE = \{e_1, \dots, e_n\} \subset \mathbb{R}^h, cosine similarities sij=viTej∥vi∥∥ej∥s_{ij} = \frac{v_i^T e_j}{\|v_i\| \|e_j\|} are zero-thresholded and normalized across words in the sentence:

    sˉij′=[sij]+∑j=1n[sij]+2\bar{s}'_{ij} = \frac{[s_{ij}]_+}{\sqrt{\sum_{j=1}^n [s_{ij}]_+^2}}

    where [x]+=max⁡(x,0)[x]_+ = \max(x, 0). In Stage 1, the attention weight of the ii-th image region with respect to the jj-th word is computed using inverse temperature λ1\lambda_1:

    αij′=exp⁡(λ1sˉij′)∑i=1kexp⁡(λ1sˉij′)\alpha'_{ij} = \frac{\exp(\lambda_1 \bar{s}'_{ij})}{\sum_{i=1}^k \exp(\lambda_1 \bar{s}'_{ij})}

    The attended visual representation ajv∈Rha_j^v \in \mathbb{R}^h for word eje_j is given by:

    ajv=∑i=1kαij′via_j^v = \sum_{i=1}^k \alpha'_{ij} v_i

    In Stage 2, relevance between word eje_j and attended visual context ajva_j^v is evaluated via cosine similarity:

    R′(ej,ajv)=ejTajv∥ej∥∥ajv∥R'(e_j, a_j^v) = \frac{e_j^T a_j^v}{\|e_j\| \|a_j^v\|}

    Global image-text similarity S′(I,T)S'(I, T) is aggregated over all nn words using LogSumExp (LSE) pooling:

    SLSE′(I,T)=log⁡(∑j=1nexp⁡(λ2R′(ej,ajv)))1/λ2S'_{\text{LSE}}(I, T) = \log \left( \sum_{j=1}^n \exp\left(\lambda_2 R'(e_j, a_j^v)\right) \right)^{1/\lambda_2}

    or average (AVG) pooling:

    SAVG′(I,T)=1n∑j=1nR′(ej,ajv)S'_{\text{AVG}}(I, T) = \frac{1}{n} \sum_{j=1}^n R'(e_j, a_j^v)

  3. Knowl 3 — Hard Negative Triplet Ranking Loss for Visual-Semantic Alignment

    model/method

    To train cross-modal visual-semantic representations, the alignment objective employs a hinge-based triplet ranking loss focused exclusively on the hardest negatives within each training mini-batch.

    For a ground-truth positive image-sentence pair (I,T)(I, T), the hardest negative sentence in the mini-batch is T^h=arg⁡max⁡d≠TS(I,d)\hat{T}_h = \arg\max_{d \neq T} S(I, d) and the hardest negative image is I^h=arg⁡max⁡m≠IS(m,T)\hat{I}_h = \arg\max_{m \neq I} S(m, T), where SS is a cross-modal similarity function. The triplet loss is defined as:

    lhard(I,T)=[α−S(I,T)+S(I,T^h)]++[α−S(I,T)+S(I^h,T)]+l_{\text{hard}}(I, T) = [\alpha - S(I, T) + S(I, \hat{T}_h)]_+ + [\alpha - S(I, T) + S(\hat{I}_h, T)]_+

    where [x]+≡max⁡(x,0)[x]_+ \equiv \max(x, 0) and α>0\alpha > 0 is a fixed margin hyperparameter.

  4. Knowl 4 — Bottom-Up Visual Region and Bidirectional Recurrent Text Feature Extraction

    model/method

    The Stacked Cross Attention Network (SCAN) maps visual regions and linguistic tokens into a shared hh-dimensional semantic space:

    1. Visual Encoding via Bottom-Up Attention: Salient image regions (objects, salient entities, and visual attributes) are detected using a Faster R-CNN with a ResNet-101 backbone pre-trained on Visual Genome. For each detected region i∈{1,…,k}i \in \{1, \dots, k\}, mean-pooled 2048-dimensional convolutional features fif_i are projected through a fully connected layer to produce the region feature vector vi∈Rhv_i \in \mathbb{R}^h:

    vi=Wvfi+bvv_i = W_v f_i + b_v

    where Wv∈Rh×2048W_v \in \mathbb{R}^{h \times 2048} and bv∈Rhb_v \in \mathbb{R}^h.

    1. Sentence Encoding via Bidirectional GRU: Each word token wiw_i in a sentence T=(w1,…,wn)T = (w_1, \dots, w_n) is represented as a one-hot vector and projected into a 300-dimensional word embedding xi=Wewix_i = W_e w_i. A bidirectional Gated Recurrent Unit (Bi-GRU) processes the sequence forward and backward:

    h⃗i=GRU→(xi),h←i=GRU←(xi)\vec{h}_i = \overrightarrow{\text{GRU}}(x_i), \quad \overleftarrow{h}_i = \overleftarrow{\text{GRU}}(x_i)

    The final contextualized word embedding ei∈Rhe_i \in \mathbb{R}^h centered around wiw_i is computed as the element-wise average of the forward and backward hidden states:

    ei=h⃗i+h←i2e_i = \frac{\vec{h}_i + \overleftarrow{h}_i}{2}

  5. Knowl 5 — Cross-Modal Retrieval Performance on Flickr30K

    data/table

    Cross-modal retrieval evaluation on the Flickr30K test set (1,000 images, 5 captions per image). Performance is measured using Recall@KK (R@KK, in percent) for sentence retrieval given an image query and image retrieval given a text query.

    Method Sentence retrieval Image retrieval
    R@1 R@5 R@10 R@1 R@5 R@10
    DVSA (R-CNN, AlexNet) 22.2 48.2 61.4 15.2 37.7 50.5
    HM-LSTM (R-CNN, AlexNet) 38.1 - 76.5 27.7 - 68.8
    SM-LSTM (VGG) 42.5 71.9 81.5 30.2 60.4 72.3
    2WayNet (VGG) 49.8 67.5 - 36.0 55.6 -
    DAN (ResNet) 55.0 81.8 89.0 39.4 69.2 79.1
    VSE++ (ResNet) 52.9 - 87.2 39.6 - 79.5
    DPC (ResNet) 55.6 81.9 89.5 39.1 69.2 80.9
    SCO (ResNet) 55.5 82.0 89.3 41.1 70.5 80.1
    Ours (Faster R-CNN, ResNet):
    SCAN t-i LSE (λ1=9,λ2=6\lambda_1 = 9, \lambda_2 = 6) 61.1 85.4 91.5 43.3 71.9 80.9
    SCAN t-i AVG (λ1=9\lambda_1 = 9) 61.8 87.5 93.7 45.8 74.4 83.0
    SCAN i-t LSE (λ1=4,λ2=5\lambda_1 = 4, \lambda_2 = 5) 67.7 88.9 94.0 44.0 74.2 82.6
    SCAN i-t AVG (λ1=4\lambda_1 = 4) 67.9 89.0 94.4 43.9 74.2 82.8
    SCAN t-i AVG + i-t LSE 67.4 90.3 95.8 48.6 77.7 85.2

    All single SCAN variants outperform previous state-of-the-art models across all metrics. The best individual sentence retrieval model is SCAN i-t AVG (R@1 of 67.9%, a 22.1% relative improvement over DPC's 55.6%), while the ensemble of SCAN t-i AVG and SCAN i-t LSE achieves the best overall image retrieval (R@1 of 48.6%, an 18.2% relative improvement over SCO's 41.1%).

  6. Knowl 6 — Cross-Modal Retrieval Performance on MS-COCO

    data/table

    Cross-modal retrieval evaluation on MS-COCO under two testing settings: 5-fold average over 1K test images and evaluation on the full 5K test image set.

    Method Sentence retrieval Image retrieval
    R@1 R@5 R@10 R@1 R@5 R@10
    1K test images
    DVSA (R-CNN, AlexNet) 38.4 69.9 80.5 27.4 60.2 74.8
    HM-LSTM (R-CNN, AlexNet) 43.9 - 87.8 36.1 - 86.7
    Order-embeddings (VGG) 46.7 - 88.9 37.9 - 85.9
    SM-LSTM (VGG) 53.2 83.1 91.5 40.7 75.8 87.4
    2WayNet (VGG) 55.8 75.2 - 39.7 63.3 -
    VSE++ (ResNet) 64.6 - 95.7 52.0 - 92.0
    DPC (ResNet) 65.6 89.8 95.5 47.1 79.9 90.0
    GXN (ResNet) 68.5 - 97.9 56.6 - 94.5
    SCO (ResNet) 69.9 92.9 97.5 56.7 87.5 94.8
    SCAN t-i LSE (λ1=9,λ2=6\lambda_1 = 9, \lambda_2 = 6) 67.5 92.9 97.6 53.0 85.4 92.9
    SCAN t-i AVG (λ1=9\lambda_1 = 9) 70.9 94.5 97.8 56.4 87.0 93.9
    SCAN i-t LSE (λ1=4,λ2=20\lambda_1 = 4, \lambda_2 = 20) 68.4 93.9 98.0 54.8 86.1 93.3
    SCAN i-t AVG (λ1=4\lambda_1 = 4) 69.2 93.2 97.5 54.4 86.0 93.6
    SCAN t-i LSE + i-t AVG 72.7 94.8 98.4 58.8 88.4 94.8
    5K test images
    Order-embeddings (VGG) 23.3 - 84.7 31.7 - 74.6
    VSE++ (ResNet) 41.3 - 81.2 30.3 - 72.4
    DPC (ResNet) 41.2 70.5 81.1 25.3 53.4 66.4
    GXN (ResNet) 42.0 - 84.7 31.7 - 74.6
    SCO (ResNet) 42.8 72.3 83.0 33.1 62.9 75.5
    SCAN i-t LSE 46.4 77.4 87.2 34.4 63.7 75.7
    SCAN t-i AVG + i-t LSE 50.4 82.2 90.0 38.6 69.3 80.4

    On the 1K test set, the ensemble (SCAN t-i LSE + i-t AVG) achieves 72.7% sentence R@1 and 58.8% image R@1, outperforming SCO by 4.0% and 3.7% relative improvements respectively. On the 5K test set, SCAN t-i AVG + i-t LSE achieves 50.4% sentence R@1 and 38.6% image R@1, yielding relative improvements of 17.8% (sentence retrieval) and 16.6% (image retrieval) over SCO.

  7. Knowl 7 — Ablation of Latent Region-Word Alignment Mechanisms

    data/table

    To evaluate the benefit of discovering fine-grained latent visual-semantic correspondences versus global embedding matching and simple aggregation, unweighted baseline models (Sum-Max Text-Image and Sum-Max Image-Text) were evaluated on Flickr30K alongside SCAN and single-vector baselines.

    Sum-Max Text-Image computes similarity without attention by taking the maximum region similarity for each word: SSM′(I,T)=∑j=1nmax⁡i(sij)S'_{SM}(I, T) = \sum_{j=1}^n \max_i(s_{ij}). Sum-Max Image-Text computes SSM(I,T)=∑i=1kmax⁡j(sij)S_{SM}(I, T) = \sum_{i=1}^k \max_j(s_{ij}), where sij=viTejs_{ij} = v_i^T e_j.

    Method Sentence retrieval Image retrieval
    R@1 R@5 R@10 R@1 R@5 R@10
    VSE++ (fixed ResNet, 1 crop) 31.9 - 68.0 23.1 - 60.7
    Sum-Max t-i 59.6 85.2 92.9 44.1 70.0 79.0
    Sum-Max i-t 56.7 83.5 89.7 36.8 65.6 74.9
    SCO (current state-of-the-art) 55.5 82.0 89.3 41.1 70.5 80.1
    SCAN t-i AVG (λ1=9\lambda_1 = 9) 61.8 87.5 93.7 45.8 74.4 83.0
    SCAN i-t AVG (λ1=10\lambda_1 = 10) 67.9 89.0 94.4 43.9 74.2 82.8

    Inferring region-word alignments (Sum-Max models) significantly outperforms matching global image and sentence vectors (VSE++). Incorporating Stacked Cross Attention (SCAN) further improves sentence retrieval R@1 from 56.7% (Sum-Max i-t) to 67.9% (SCAN i-t AVG) and image retrieval R@1 from 44.1% (Sum-Max t-i) to 45.8% (SCAN t-i AVG).

  8. Knowl 8 — Ablation of Architectural and Training Components in SCAN

    data/table

    Ablation study on the Flickr30K dataset evaluating the impact of individual architectural components and loss configurations relative to the baseline SCAN Image-Text model with average pooling (SCAN i-t AVG, λ1=4\lambda_1 = 4).

    Method Sentence retrieval Image retrieval
    R@1 R@5 R@10 R@1 R@5 R@10
    Baseline: SCAN i-t AVG 67.9 89.0 94.4 43.9 74.2 82.8
    No hard negatives 45.8 77.8 86.2 33.9 63.7 73.4
    Not normalize image embedding 67.8 89.3 94.6 43.3 73.7 82.7
    SCAN i-t SUM 63.9 89.0 93.9 45.0 73.1 82.0
    SCAN i-t MAX 59.7 83.9 90.8 43.3 72.0 80.9
    One-directional GRU 63.6 87.7 93.7 43.2 73.1 82.3

    Key observations:

    • Hard Negatives: Using hard negatives in the triplet ranking loss is critical, improving sentence retrieval R@1 by 48.2% relatively (from 45.8% to 67.9%).
    • Pooling Function: Average pooling (R@1 = 67.9%) outperforms both summation (SUM, R@1 = 63.9%) and max pooling (MAX, R@1 = 59.7%).
    • Recurrent Context: Bidirectional GRU provides a 4.3% R@1 gain in sentence retrieval and a 0.7% gain in image retrieval over a unidirectional GRU.
    • Feature Normalization: Omitting region embedding normalization in cosine similarity computation has minimal impact on retrieval performance (67.8% vs. 67.9% R@1).
  9. Knowl 9 — Ensembling Complementary Cross Attention Formulations

    empirical result

    Ensembling the Text-Image (t-i) and Image-Text (i-t) formulations by averaging their predicted similarity scores yields higher retrieval accuracy than either single formulation alone on both Flickr30K and MS-COCO.

    On Flickr30K, the ensemble of SCAN t-i AVG and SCAN i-t LSE achieves 48.6% R@1 on image retrieval compared to 45.8% for t-i AVG and 44.0% for i-t LSE. On MS-COCO (5K test set), the ensemble achieves 50.4% sentence R@1 and 38.6% image R@1, compared to 46.4% and 34.4% for the best single model (SCAN i-t LSE). This gain demonstrates that attending from image regions to sentence words and attending from words to image regions capture complementary visual-semantic alignment cues.

Coverage note — Qualitative attention visualization figures and specific retrieval ranking examples (Figures 4, 5, and 6) were omitted as standalone knowls because their core scientific implications (interpretability and alignment behavior) are fully reflected in the methodology and empirical results.

References

  1. 1.Anderson, P., et al.: Bottom-up and top-down attention for image captioning and VQA. In: CVPR (2018)
  2. 2.Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv preprint arXiv:1607.06450 (2016)
  3. 3.Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: ICLR (2015)
  4. 4.Buschman, T.J., Miller, E.K.: Top-down versus bottom-up control of attention in the prefrontal and posterior parietal cortices. Science 315(5820), 1860–1862 (2007)
  5. 5.Chorowski, J.K., Bahdanau, D., Serdyuk, D., Cho, K., Bengio, Y.: Attention-based models for speech recognition. In: NIPS (2015)
  6. 6.Corbetta, M., Shulman, G.L.: Control of goal-directed and stimulus-driven attention in the brain. Nat. Rev. Neurosci. 3(3), 201 (2002)
  7. 7.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: a large-scale hierarchical image database. In: CVPR (2009)
  8. 8.Devlin, J., et al.: Language models for image captioning: the quirks and what works. In: ACL (2015)
  9. 9.Eisenschtat, A., Wolf, L.: Linking image and text with 2-way nets. In: CVPR (2017)
  10. 10.Faghri, F., Fleet, D.J., Kiros, J.R., Fidler, S.: VSE++: improved visual-semantic embeddings. arXiv preprint arXiv:1707.05612 (2017)
  11. 11.Fang, H., et al.: From captions to visual concepts and back. In: CVPR (2015)
  12. 12.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR (2014)
  13. 13.Gu, J., Cai, J., Joty, S., Niu, L., Wang, G.: Look, imagine and match: improving textual-visual cross-modal retrieval with generative models. In: CVPR (2018)
  14. 14.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  15. 15.He, X., Deng, L., Chou, W.: Discriminative learning in sequential pattern recognition. IEEE Sig. Process. Mag. 25(5), 1436 (2008)
  16. 16.Huang, Y., Wang, W., Wang, L.: Instance-aware image and sentence matching with selective multimodal LSTM. In: CVPR (2017)
  17. 17.Huang, Y., Wu, Q., Wang, L.: Learning semantic concepts and order for image and sentence matching. In: CVPR (2018)
  18. 18.Juang, B.H., Hou, W., Lee, C.H.: Minimum classification error rate methods for speech recognition. IEEE Trans. Speech Audio process. 5(3), 257–265 (1997)
  19. 19.Karpathy, A., Fei-Fei, L.: Deep visual-semantic alignments for generating image descriptions. In: CVPR (2015)
  20. 20.Karpathy, A., Joulin, A., Fei-Fei, L.: Deep fragment embeddings for bidirectional image sentence mapping. In: NIPS (2014)
  21. 21.Katsuki, F., Constantinidis, C.: Bottom-up and top-down attention: different processes and overlapping neural systems. Neuroscientist 20(5), 509–521 (2014)
  22. 22.Kiros, R., Salakhutdinov, R., Zemel, R.S.: Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539 (2014)
  23. 23.Klein, B., Lev, G., Sadeh, G., Wolf, L.: Associating neural word embeddings with deep image representations using fisher vectors. In: CVPR (2015)
  24. 24.Krishna, R., et al.: Visual Genome: connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 123(1), 32–73 (2017)
  25. 25.Kumar, A., et al.: Ask me anything: dynamic memory networks for natural language processing. In: ICML (2016)
  26. 26.Lee, K.H., He, X., Zhang, L., Yang, L.: CleanNet: transfer learning for scalable image classifier training with label noise. In: CVPR (2018)
  27. 27.Lev, G., Sadeh, G., Klein, B., Wolf, L.: RNN Fisher vectors for action recognition and image annotation. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9910, pp. 833–850. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46466-4_50
  28. 28.Li, J., Luong, M.T., Jurafsky, D.: A hierarchical neural autoencoder for paragraphs and documents. In: ACL (2015)
  29. 29.Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48
  30. 30.Luong, M.T., Pham, H., Manning, C.D.: Effective approaches to attention-based neural machine translation. In: EMNLP (2015)
  31. 31.Nam, H., Ha, J.W., Kim, J.: Dual attention networks for multimodal reasoning and matching. In: CVPR (2017)
  32. 32.Niu, Z., Zhou, M., Wang, L., Gao, X., Hua, G.: Hierarchical multimodal LSTM for dense visual-semantic embedding. In: ICCV (2017)
  33. 33.Peng, Y., Qi, J., Yuan, Y.: CM-GANs: cross-modal generative adversarial networks for common representation learning. arXiv preprint arXiv:1710.05106 (2017)
  34. 34.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: towards real-time object detection with region proposal networks. In: NIPS (2015)
  35. 35.Rush, A.M., Chopra, S., Weston, J.: A neural attention model for abstractive sentence summarization. In: EMNLP (2015)
  36. 36.Schuster, M., Paliwal, K.K.: Bidirectional recurrent neural networks. IEEE Trans. Sig. Process. 45(11), 2673–2681 (1997)
  37. 37.Socher, R., Karpathy, A., Le, Q.V., Manning, C.D., Ng, A.Y.: Grounded compositional semantics for finding and describing images with sentences. In: ACL (2014)
  38. 38.Vendrov, I., Kiros, R., Fidler, S., Urtasun, R.: Order-embeddings of images and language. In: ICLR (2016)
  39. 39.Wang, L., Li, Y., Lazebnik, S.: Learning deep structure-preserving image-text embeddings. In: CVPR (2016)
  40. 40.Xu, K., et al.: Show, attend and tell: neural image caption generation with visual attention. In: ICML (2015)
  41. 41.Xu, T., et al.: AttnGAN: fine-grained text to image generation with attentional generative adversarial networks. In: CVPR (2018)
  42. 42.Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., Hovy, E.: Hierarchical attention networks for document classification. In: NAACL-HLT (2016)
  43. 43.Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions. In: ACL (2014)
  44. 44.Zheng, Z., Zheng, L., Garrett, M., Yang, Y., Shen, Y.D.: Dual-path convolutional image-text embedding. arXiv preprint arXiv:1711.05535 (2017)

Citation

MLA
Lee, K.-H., et al. “Stacked Cross Attention for Image-Text Matching”. arXiv, 2018, http://arxiv.org/abs/1803.08024v2.
APA
Lee, K.-H., Chen, X., Hua, G., Hu, H., & He, X. (2018). Stacked Cross Attention for Image-Text Matching. arXiv. http://arxiv.org/abs/1803.08024v2
Chicago
Lee, K.-H., X. Chen, G. Hua, H. Hu, and X. He. 2018. “Stacked Cross Attention for Image-Text Matching”. arXiv. http://arxiv.org/abs/1803.08024v2.
Harvard
Lee, K.-H. et al. (2018) “Stacked Cross Attention for Image-Text Matching”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.08024v2.
Vancouver
1. Lee K-H, Chen X, Hua G, Hu H, He X (2018) Stacked Cross Attention for Image-Text Matching. arXiv

BibTeX

@article{lee2018stacked,
  title = {Stacked Cross Attention for Image-Text Matching},
  author = {Lee, Kuang-Huei and Chen, Xi and Hua, Gang and Hu, Houdong and He, Xiaodong},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.08024v2},
  eprint = {1803.08024}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF