Stacked Attention Networks for Image Question Answering

Zichao YangXiaodong HeJianfeng GaoLi DengAlex Smola

article2015CVPR2,026 citations

Introduces stacked attention networks that perform multi-step visual reasoning by iteratively querying an image to progressively pinpoint the visual evidence required for question answering.

Listen

Automated image question answering requires artificial intelligence systems to interpret natural language questions and locate specific visual evidence within an image to predict accurate answers. Conventional models typically combine a question with a single global image summary, which introduces background noise and often fails when questions require multi-step reasoning over fine-grained, localized details.

The article evaluates whether introducing stacked attention networks—which query an image over multiple progressive reasoning steps—can improve question-answering accuracy across diverse visual benchmarks. The proposed architecture uses deep neural networks to extract spatial visual features and question representations, then applies multiple visual attention layers where each layer refines the query and hones in on increasingly specific image regions before predicting an answer.

To evaluate this framework, the authors conducted extensive experiments across four standard benchmark datasets: DAQUAR-ALL, DAQUAR-REDUCED, COCO-QA, and the large-scale VQA dataset comprising hundreds of thousands of question-answer pairs. The evaluations compared one-layer and two-layer stacked attention networks against established baseline architectures using standard classification accuracy and taxonomy-based similarity metrics.

The analysis yielded four major findings. First, two-layer stacked attention networks consistently outperformed prior state-of-the-art models across all four benchmarks, achieving absolute accuracy improvements of 5.9% on DAQUAR-ALL (reaching 29.3%), 6.5% on DAQUAR-REDUCED (reaching 46.2%), 5.1% to 6.6% on COCO-QA (reaching 61.6%), and 4.8% on the official VQA test benchmark (reaching 58.9%). Second, models utilizing two attention layers systematically outperformed single-layer variants across all datasets, confirming the value of progressive reasoning. Third, the most substantial performance gains occurred on fine-grained visual queries involving object types, colors, and locations (such as a 7.2% gain on color queries in COCO-QA and a 9.7% gain on open-ended descriptive questions in VQA), whereas binary yes/no questions showed minimal improvement. Fourth, empirical testing demonstrated diminishing returns beyond two reasoning steps, as networks with three or more attention layers did not yield additional accuracy gains.

These results imply that iterative spatial attention is a highly effective mechanism for reducing irrelevant visual noise and resolving complex multi-object relationships. For system design and deployment, adopting a two-layer attention architecture delivers optimal performance without incurring the computational overhead or training risks associated with deeper attention stacks. The findings also indicate that language-dominated question types, such as binary queries, require different modeling strategies than spatially grounded questions.

Based on these findings, teams developing multimodal visual question-answering systems should implement multi-stage visual attention mechanisms to achieve significant accuracy gains on descriptive and localized tasks. However, decision-makers should recognize existing limitations: an error analysis of misclassified samples revealed that 42% involved identifying the correct visual region but predicting the wrong answer word, while 31% involved ambiguous labels and 5% contained human labeling errors in benchmark datasets. While confidence in the multi-step attention mechanism's superior performance is high, addressing downstream answer prediction precision and refining dataset label quality remain essential priorities for future development.

  • Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This foundational paper establishes the Visual Question Answering task and benchmark dataset that the source paper directly targets and seeks to improve upon.
  • Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). It introduces spatial visual attention over convolutional feature maps for multimodal vision-language tasks, serving as the direct inspiration for attention-based visual reasoning in the source.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). It introduces the concept of multi-hop memory querying for iterative reasoning, which directly underlies the stacked, multi-layer reasoning architecture developed in the source paper.
  • Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). It presents fundamental methods for aligning natural language sentence representations with local visual features extracted from convolutional networks.
  • Paper: Recurrent Models of Visual Attention, Volodymyr Mnih et al. (2014). It introduces core principles of sequential visual attention mechanisms in deep learning that motivate multi-step visual query processing.
Cover for Stacked Attention Networks for Image Question Answering

Abstract

This paper presents stacked attention networks (SANs) that learn to answer natural language questions from images. SANs use semantic representation of a question as query to search for the regions in an image that are related to the answer. We argue that image question answering (QA) often requires multiple steps of reasoning. Thus, we develop a multiple-layer SAN in which we query an image multiple times to infer the answer progressively. Experiments conducted on four image QA data sets demonstrate that the proposed SANs significantly outperform previous state-of-the-art approaches. The visualization of the attention layers illustrates the progress that the SAN locates the relevant visual clues that lead to the answer of the question layer-by-layer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Stacked Attention Networks (SANs)
  • 3.1 Image Model
  • 3.2 Question Model
  • 3.2.1 LSTM based question model
  • 3.2.2 CNN based question model
  • 3.3 Stacked Attention Networks
  • 4 Experiments
  • 4.1 Data sets
  • 4.2 Baselines and evaluation methods
  • 4.3 Model configuration and training
  • 4.4 Results and analysis
  • 4.5 Visualization of attention layers
  • 4.6 Errors analysis
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Stacked Attention Network Architecture for Multi-Step Visual Question Answering

    model/method

    The Stacked Attention Network (SAN) performs multi-step visual reasoning by iteratively querying image regions using a question-guided query vector that is updated across KK successive attention layers.

    Let vI=[v1,v2,…,vm]∈Rd×mv_I = [v_1, v_2, \dots, v_m] \in \mathbb{R}^{d \times m} denote the matrix of transformed visual feature vectors for mm spatial image regions, and let vQ∈Rdv_Q \in \mathbb{R}^d be the question representation vector. The query vector is initialized as u0=vQu^0 = v_Q.

    For each attention layer k∈{1,…,K}k \in \{1, \dots, K\}, the intermediate attention representation hAk∈Rk′×mh_A^k \in \mathbb{R}^{k' \times m} and the spatial attention probability vector pIk=[pI,1k,…,pI,mk]T∈Rmp_I^k = [p_{I,1}^k, \dots, p_{I,m}^k]^T \in \mathbb{R}^m are computed via:

    hAk=tanh⁡(WI,AkvI⊕(WQ,Akuk−1+bAk))h_A^k = \tanh\left(W_{I,A}^k v_I \oplus \left(W_{Q,A}^k u^{k-1} + b_A^k\right)\right)

    pIk=softmax(WPkhAk+bPk)p_I^k = \text{softmax}\left(W_P^k h_A^k + b_P^k\right)

    where WI,Ak,WQ,Ak∈Rk′×dW_{I,A}^k, W_{Q,A}^k \in \mathbb{R}^{k' \times d}, bAk∈Rk′b_A^k \in \mathbb{R}^{k'}, WPk∈R1×k′W_P^k \in \mathbb{R}^{1 \times k'}, and bPk∈Rb_P^k \in \mathbb{R}. The symbol ⊕\oplus denotes the addition of a vector to each column of a matrix. The softmax function normalizes across the mm spatial regions.

    The attended image feature vector v~Ik∈Rd\tilde{v}_I^k \in \mathbb{R}^d is computed as the weighted sum of region features:

    v~Ik=∑i=1mpI,ikvi\tilde{v}_I^k = \sum_{i=1}^m p_{I,i}^k v_i

    The query vector is updated by combining the visual context with the previous query:

    uk=v~Ik+uk−1u^k = \tilde{v}_I^k + u^{k-1}

    After KK layers of attention, the final combined query vector uKu^K is passed to a softmax classifier over candidate answer classes:

    pans=softmax(WuuK+bu)p_{\text{ans}} = \text{softmax}\left(W_u u^K + b_u\right)

    where WuW_u and bub_u are the weight matrix and bias vector of the output classification layer.

  2. Knowl 2 — Question Encoders in SAN: LSTM and Multi-Scale CNN

    model/method

    Natural language questions q=[q1,q2,…,qT]q = [q_1, q_2, \dots, q_T] consisting of one-hot word vectors qtq_t are first mapped to continuous embeddings xt=Weqt∈Rdex_t = W_e q_t \in \mathbb{R}^{d_e} via embedding matrix WeW_e. SAN supports two alternative question representation models:

    1. LSTM Question Model: Embeddings xtx_t are fed sequentially to a Long Short-Term Memory network:

    it=σ(Wxixt+Whiht−1+bi)i_t = \sigma(W_{xi} x_t + W_{hi} h_{t-1} + b_i)

    ft=σ(Wxfxt+Whfht−1+bf)f_t = \sigma(W_{xf} x_t + W_{hf} h_{t-1} + b_f)

    ot=σ(Wxoxt+Whoht−1+bo)o_t = \sigma(W_{xo} x_t + W_{ho} h_{t-1} + b_o)

    ct=ftct−1+ittanh⁡(Wxcxt+Whcht−1+bc)c_t = f_t c_{t-1} + i_t \tanh(W_{xc} x_t + W_{hc} h_{t-1} + b_c)

    ht=ottanh⁡(ct)h_t = o_t \tanh(c_t)

    where σ\sigma is the logistic sigmoid function, and it,ft,ot,ct,hti_t, f_t, o_t, c_t, h_t denote input gate, forget gate, output gate, memory cell state, and hidden state, respectively. The final hidden state serves as the question representation: vQ=hTv_Q = h_T.

    1. CNN Question Model: Convolutions are applied across the concatenated word embedding matrix x1:T=[x1,…,xT]x_{1:T} = [x_1, \dots, x_T] using three filter window sizes c∈{1,2,3}c \in \{1, 2, 3\} corresponding to unigram, bigram, and trigram phrases:

    hc,t=tanh⁡(Wcxt:t+c−1+bc)h_{c,t} = \tanh(W_c x_{t:t+c-1} + b_c)

    where WcW_c and bcb_c are the convolution weights and biases for window size cc. A coordinate-wise max-pooling over time operation is applied for each filter size:

    h~c=max⁡t[hc,1,hc,2,…,hc,T−c+1]\tilde{h}_c = \max_{t} [h_{c,1}, h_{c,2}, \dots, h_{c,T-c+1}]

    The pooled vectors are concatenated to yield the complete question representation: vQ=[h~1,h~2,h~3]v_Q = [\tilde{h}_1, \tilde{h}_2, \tilde{h}_3].

  3. Knowl 3 — Spatial Feature Extraction and Region Projection from VGGNet

    model/method

    Rather than using a global feature vector from a fully connected layer, SAN extracts spatial feature maps from raw images II using the final pooling layer of a pre-trained VGGNet:

    fI=CNNvgg(I)f_I = \text{CNN}_{\text{vgg}}(I)

    Input images are rescaled to 448×448448 \times 448 pixels, producing a feature map fIf_I of shape 512×14×14512 \times 14 \times 14. This feature map decomposes the image into m=14×14=196m = 14 \times 14 = 196 spatial grid regions, where each 512-dimensional vector fI,if_{I,i} (i∈{0,…,195}i \in \{0, \dots, 195\}) corresponds to a 32×3232 \times 32 pixel receptive field.

    To align the visual features with the question representation dimension dd, each region vector is projected using a single-layer perceptron with non-linear activation:

    vI=tanh⁡(WIfI+bI)v_I = \tanh(W_I f_I + b_I)

    where WI∈Rd×512W_I \in \mathbb{R}^{d \times 512} and bI∈Rdb_I \in \mathbb{R}^d. The resulting matrix vI∈Rd×196v_I \in \mathbb{R}^{d \times 196} contains the visual feature vector viv_i for each image region ii in its columns.

  4. Knowl 4 — Evaluation Results on DAQUAR-ALL and DAQUAR-REDUCED Benchmarks

    data/table

    Performance of Stacked Attention Networks (SAN) evaluated against prior image question answering models on DAQUAR-ALL (6,795 train / 5,673 test questions across indoor scenes) and DAQUAR-REDUCED (3,876 train / 297 test questions constrained to 37 object classes). Metrics include top-1 classification accuracy and Wu-Palmer similarity scores at thresholds 0.9 (WUPS0.9) and 0.0 (WUPS0.0).

    DAQUAR-ALL (%) DAQUAR-REDUCED (%)
    Methods Accuracy WUPS0.9 WUPS0.0 Accuracy WUPS0.9 WUPS0.0
    Multi-World 7.9 11.9 38.8 12.7 18.2 51.5
    Ask-Your-Neurons (Language) 19.1 25.2 65.1 31.7 38.4 80.1
    Ask-Your-Neurons (Language + IMG) 21.7 28.0 65.0 34.7 40.8 79.5
    GUESS (VSE) - - - 18.2 29.7 77.6
    BOW (VSE) - - - 32.7 43.2 81.3
    LSTM (VSE) - - - 32.7 43.5 81.6
    IMG+BOW (VSE) - - - 34.2 45.0 81.5
    VIS+LSTM (VSE) - - - 34.4 46.1 82.2
    2-VIS+BLSTM (VSE) - - - 35.8 46.8 82.2
    IMG-CNN 23.4 29.6 63.0 39.7 44.9 83.1
    SAN(1, LSTM) 28.9 34.7 68.5 45.2 49.6 84.0
    SAN(1, CNN) 29.2 35.1 67.8 45.2 49.6 83.7
    SAN(2, LSTM) 29.3 34.9 68.1 46.2 51.2 85.1
    SAN(2, CNN) 29.3 35.1 68.6 45.5 50.2 83.6
    Human 50.2 50.8 67.3 60.3 61.0 79.0

    On DAQUAR-ALL, two-layer SANs outperform the best prior baseline (IMG-CNN) by 5.9 percentage points in accuracy and 5.5 points in WUPS0.9. On DAQUAR-REDUCED, SAN(2, LSTM) outperforms IMG-CNN by 6.5 percentage points in accuracy and 6.3 points in WUPS0.9.

  5. Knowl 5 — Evaluation and Question-Type Breakdown on COCO-QA Benchmark

    data/table

    Performance of Stacked Attention Networks (SAN) on COCO-QA (78,736 train / 38,948 test samples generated from MS-COCO captions) overall and across question types: Objects (70%), Number (7%), Color (17%), and Location (6%).

    Methods Accuracy (%) WUPS0.9 (%) WUPS0.0 (%) Objects (%) Number (%) Color (%) Location (%)
    GUESS 6.7 17.4 73.4 2.1 35.8 13.9 8.9
    BOW 37.5 48.5 82.8 37.3 43.6 34.8 40.8
    LSTM 36.8 47.6 82.3 35.9 45.3 36.3 38.4
    IMG 43.0 58.6 85.9 40.4 29.3 42.7 44.2
    IMG+BOW 55.9 66.8 89.0 58.7 44.1 52.0 49.4
    VIS+LSTM 53.3 63.9 88.3 56.5 46.1 45.9 45.5
    2-VIS+BLSTM 55.1 65.3 88.6 58.2 44.8 49.5 47.3
    CNN (Question) 32.7 44.3 80.9 - - - -
    IMG-CNN 55.0 65.4 88.6 - - - -
    SAN(1, LSTM) 59.6 69.6 90.1 62.5 49.0 54.8 51.6
    SAN(1, CNN) 60.7 70.6 90.5 63.6 48.7 56.7 52.7
    SAN(2, LSTM) 61.0 71.0 90.7 63.6 49.8 57.9 52.8
    SAN(2, CNN) 61.6 71.6 90.9 64.5 48.6 57.9 54.0

    SAN(2, CNN) achieves a 5.7 percentage point overall accuracy gain over the strongest baseline (IMG+BOW). The largest category-specific improvements over IMG+BOW are observed in Color (+5.9%), Objects (+5.8%), Location (+4.6%), and Number (+4.5%). Two-layer SAN models outperform single-layer SAN models across all categories (e.g., Color improves by 2.2% from layer 1 to layer 2).

  6. Knowl 6 — Evaluation Results on the Open-Ended VQA Benchmark

    data/table

    Performance of SAN evaluated on the VQA dataset (248,349 training and 121,512 validation questions over MS-COCO images, evaluated with human consensus metric min⁡(# matching human labels/3,1)\min(\text{\# matching human labels} / 3, 1)). Results are reported on the official test-dev and test-std server evaluation as well as a local validation partition (val2).

    test-dev (%) test-std (%)
    Methods All Yes/No Number Other All
    Question only 48.1 75.7 36.7 27.1 -
    Image only 28.1 64.0 0.4 3.8 -
    Q+I 52.6 75.6 33.7 37.4 -
    LSTM Q 48.8 78.2 35.7 26.6 -
    LSTM Q+I 53.7 78.9 35.2 36.4 54.1
    SAN(2, CNN) 58.7 79.3 36.6 46.1 58.9

    On the local val2 partition, model variations perform as follows:

    Methods All (%) Yes/No (36%) Number (10%) Other (54%)
    SAN(1, LSTM) 56.6 78.1 41.6 44.8
    SAN(1, CNN) 56.9 78.8 42.0 45.0
    SAN(2, LSTM) 57.3 78.3 42.2 45.9
    SAN(2, CNN) 57.6 78.6 41.8 46.4

    On test-dev, SAN(2, CNN) outperforms the baseline LSTM Q+I by 5.0% absolute overall, with the largest gain in the 'Other' category (+9.7%), compared to +1.4% on 'Number' and +0.4% on 'Yes/No'.

  7. Knowl 7 — Layer-by-Layer Visual Attention Progressive Focus

    empirical result

    Visualizing the attention weight distributions across successive layers reveals a coarse-to-fine reasoning pattern:

    1. In the first attention layer (k=1k=1), attention is distributed broadly across multiple objects and contextual concepts mentioned in the question query (for example, identifying both a bicycle and a basket).
    2. In the second attention layer (k=2k=2), utilizing the updated query vector u1=v~I1+u0u^1 = \tilde{v}_I^1 + u^0, the attention sharply concentrates on the precise image sub-regions containing visual evidence required for the answer (such as focusing specifically on the dogs inside the basket, or the horns on a person's head).

    This multi-step progression allows the network to filter out irrelevant visual noise that single-layer attention or global image representations cannot exclude.

  8. Knowl 8 — Error Taxonomy for Stacked Attention Networks on COCO-QA

    empirical result

    Qualitative error analysis of 100 randomly sampled mispredicted instances from the COCO-QA test set identified four distinct failure modes:

    1. Incorrect Attention Region (22%): The model focuses attention on image regions irrelevant to answering the question.
    2. Correct Region but Incorrect Prediction (42%): The model attends to the correct visual region but fails in recognition or attribute classification, outputting the wrong word label.
    3. Ambiguous / Plausible Answers (31%): The model outputs a prediction that disagrees with the ground-truth annotation but is visually plausible (e.g., predicting 'vase' when the dataset ground-truth label is 'pot').
    4. Erroneous Dataset Labels (5%): The ground-truth dataset label is itself factually incorrect (e.g., the ground-truth label is 'cars' for an image showing a train).
  9. Knowl 9 — Training Configuration and Model Hyperparameters

    experimental setup

    The experimental setup and hyperparameters for training SAN models are as follows:

    • Visual backbone: VGGNet pre-trained features are extracted from the last pooling layer (512×14×14512 \times 14 \times 14) on 448×448448 \times 448 images; VGGNet parameters remain frozen during training.
    • DAQUAR and COCO-QA question encoders: Word embedding dimension is 500. LSTM hidden size is 500. For the CNN question encoder, filter counts for unigram, bigram, and trigram convolutions are 128, 256, and 256 respectively, yielding a 640-dimensional question representation.
    • VQA question encoder: For the larger VQA dataset, LSTM hidden size and CNN filter counts are doubled to accommodate vocabulary and output space scale.
    • Optimization: Stochastic gradient descent (SGD) with momentum 0.9, mini-batch size of 100, learning rate selected via grid search, gradient clipping, and dropout regularization.
  10. Knowl 10 — Performance Saturation with Increasing Attention Layers

    limitation

    Empirical evaluations comparing stacked attention network architectures with different stack depths show that two-layer SAN models (K=2K=2) consistently outperform single-layer models (K=1K=1) across all datasets. However, stacking three or more attention layers (K≥3K \ge 3) yields no further improvement in question answering accuracy.

Coverage note — No substantial contributed material was omitted. All architectural equations, question models, visual feature extraction procedures, full benchmark results across DAQUAR, COCO-QA, and VQA, layer visualizations, error analyses, training setups, and depth limitations are fully covered.

References

  1. 1.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. Vqa: Visual question answering. arXiv preprint arXiv:1505.00468, 2015. 1, 2, 5, 6, 7
  2. 2.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014. 1
  3. 3.J. Berant and P. Liang. Semantic parsing via paraphrasing. In Proceedings of ACL, volume 7, page 92, 2014. 1
  4. 4.A. Bordes, S. Chopra, and J. Weston. Question answering with subgraph embeddings. arXiv preprint arXiv:1406.3676, 2014. 1
  5. 5.X. Chen and C. L. Zitnick. Learning a recurrent visual representation for image caption generation. arXiv preprint arXiv:1411.5654, 2014. 2
  6. 6.H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al. From captions to visual concepts and back. arXiv preprint arXiv:1411.4952, 2014. 2, 3, 5
  7. 7.H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu. Are you talking to a machine? dataset and methods for multilingual image question answering. arXiv preprint arXiv:1505.05612, 2015. 1, 2
  8. 8.A. Graves. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850, 2013. 5
  9. 9.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. 1
  10. 10.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. arXiv preprint arXiv:1412.2306, 2014. 2
  11. 11.Y. Kim. Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882, 2014. 3
  12. 12.R. Kiros, R. Salakhutdinov, and R. S. Zemel. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539, 2014. 2
  13. 13.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2
  14. 14.A. Kumar, O. Irsoy, J. Su, J. Bradbury, R. English, B. Pierce, P. Ondruska, I. Gulrajani, and R. Socher. Ask me anything: Dynamic memory networks for natural language processing. arXiv preprint arXiv:1506.07285, 2015. 1
  15. 15.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 1
  16. 16.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014, pages 740–755. Springer, 2014. 5
  17. 17.L. Ma, Z. Lu, and H. Li. Learning to answer questions from image using convolutional neural network. arXiv preprint arXiv:1506.00333, 2015. 2, 5, 6
  18. 18.M. Malinowski and M. Fritz. A multi-world approach to question answering about real-world scenes based on uncertain input. In Advances in Neural Information Processing Systems, pages 1682–1690, 2014. 1, 2, 4, 5, 6
  19. 19.M. Malinowski, M. Rohrbach, and M. Fritz. Ask your neurons: A neural-based approach to answering questions about images. arXiv preprint arXiv:1505.01121, 2015. 1, 2, 5, 6
  20. 20.J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). arXiv preprint arXiv:1412.6632, 2014. 2
  21. 21.M. Ren, R. Kiros, and R. Zemel. Exploring models and data for image question answering. arXiv preprint arXiv:1505.02074, 2015. 1, 2, 5, 6
  22. 22.Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil. A latent semantic model with convolutional-pooling structure for information retrieval. In Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management, pages 101–110. ACM, 2014. 3
  23. 23.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 2
  24. 24.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958, 2014. 5
  25. 25.I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pages 3104–3112, 2014. 3
  26. 26.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. arXiv preprint arXiv:1409.4842, 2014. 2
  27. 27.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. arXiv preprint arXiv:1411.4555, 2014. 2
  28. 28.J. Weston, S. Chopra, and A. Bordes. Memory networks. arXiv preprint arXiv:1410.3916, 2014. 1
  29. 29.Z. Wu and M. Palmer. Verbs semantics and lexical selection. In Proceedings of the 32nd annual meeting on Association for Computational Linguistics, pages 133–138. Association for Computational Linguistics, 1994. 5
  30. 30.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044, 2015. 1, 2
  31. 31.W.-t. Yih, M.-W. Chang, X. He, and J. Gao. Semantic parsing via staged query graph generation: Question answering with knowledge base. In Proceedings of the Joint Conference of the 53rd Annual Meeting of the ACL and the 7th International Joint Conference on Natural Language Processing of the AFNLP, 2015. 1
  32. 32.W.-t. Yih, X. He, and C. Meek. Semantic parsing for singlerelation question answering. In Proceedings of ACL, 2014. 1

Citation

MLA
Yang, Z., et al. “Stacked Attention Networks for Image Question Answering”. arXiv, 2015, http://arxiv.org/abs/1511.02274v2.
APA
Yang, Z., He, X., Gao, J., Deng, L., & Smola, A. (2015). Stacked Attention Networks for Image Question Answering. arXiv. http://arxiv.org/abs/1511.02274v2
Chicago
Yang, Z., X. He, J. Gao, L. Deng, and A. Smola. 2015. “Stacked Attention Networks for Image Question Answering”. arXiv. http://arxiv.org/abs/1511.02274v2.
Harvard
Yang, Z. et al. (2015) “Stacked Attention Networks for Image Question Answering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1511.02274v2.
Vancouver
1. Yang Z, He X, Gao J, Deng L, Smola A (2015) Stacked Attention Networks for Image Question Answering. arXiv

BibTeX

@article{yang2015stacked,
  title = {Stacked Attention Networks for Image Question Answering},
  author = {Yang, Zichao and He, Xiaodong and Gao, Jianfeng and Deng, Li and Smola, Alex},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1511.02274v2},
  eprint = {1511.02274}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE