Hierarchical Question-Image Co-Attention for Visual Question Answering

Jiasen LuJianwei YangDhruv BatraDevi Parikh

article2016NeurIPS1,752 citations

Introduces a hierarchical co-attention model for visual question answering that jointly computes attention over relevant image regions and multi-level language structures across word, phrase, and question representations.

Listen

Visual Question Answering enables computing systems to interpret both text-based questions and visual content to produce accurate answers. While existing technologies effectively locate relevant spatial regions in an image, they largely ignore linguistic nuance by treating questions uniformly. In real-world environments, however, language contains significant variations and filler words that can degrade automated reasoning if not properly prioritized.

The main objective of the article is to demonstrate a joint attention framework that simultaneously determines which visual regions to observe and which question elements to prioritize. Specifically, it evaluates whether coupling visual and linguistic attention across multiple hierarchical levels improves answer accuracy over standard models.

To accomplish this, the authors developed a hierarchical co-attention architecture and evaluated it across two benchmark datasets: the primary VQA dataset comprising over 6.1 million question-answer pairs and the COCO-QA dataset containing roughly 117,000 question samples. The approach represents text hierarchically at the word, phrase, and full-question levels. It links these linguistic representations with image features using two co-attention strategies: parallel co-attention, which generates visual and linguistic attention simultaneously, and alternating co-attention, which alternates attention between modalities sequentially.

The evaluation yielded several key findings. First, the hierarchical co-attention model established new state-of-the-art results, raising overall benchmark accuracy on the VQA dataset to 62.1% from a previous baseline of 60.4%, and on COCO-QA to 65.4% from 61.6%. Second, the architecture showed notable improvements on difficult query types, including a 3.4 percentage point gain on open-ended descriptive questions and a 1.4 percentage point gain on numerical counting questions. Third, ablation analyses confirmed that linguistic attention is critical: removing question-level attention caused the largest performance decline of 1.7 percentage points, while omitting phrase-level and word-level attention reduced performance by 0.3 and 0.2 percentage points, respectively. Finally, incorporating advanced visual feature extractors further elevated results, outperforming competing models using the same underlying vision networks.

These findings indicate that treating text and vision with equal, symmetric importance significantly improves multimodal reasoning and provides greater robustness against linguistic phrasing differences. By demonstrating that question attention is as vital as visual focus, the work highlights that multimodal applications cannot rely solely on better vision models to drive accuracy improvements.

For practitioners and decision-makers implementing vision-language systems, adopting hierarchical co-attention frameworks provides a clear pathway to higher accuracy and better interpretability. Organizations should weigh the operational trade-offs between parallel co-attention, which can be harder to optimize due to joint matrix operations, and alternating co-attention, which carries modest risks of sequential error accumulation. Future development should focus on extending this co-attention framework to other complex multimodal applications beyond static question answering.

Confidence in these findings is supported by rigorous evaluations across established large-scale datasets and systematic component testing. However, decision-makers should note that the evaluation is bounded by closed-vocabulary classification setups that restrict answers to the most frequent training categories, meaning additional validation is needed before deploying into entirely unconstrained operational settings.

Cover for Hierarchical Question-Image Co-Attention for Visual Question Answering

Abstract

A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant to answering the question. In this paper, we argue that in addition to modeling "where to look" or visual attention, it is equally important to model "what words to listen to" or question attention. We present a novel co-attention model for VQA that jointly reasons about image and question attention. In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN). Our model improves the state-of-the-art on the VQA dataset from 60.3% to 60.5%, and from 61.6% to 63.3% on the COCO-QA dataset. By using ResNet, the performance is further improved to 62.1% for VQA and 65.4% for COCO-QA.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Notation
  • 3.2 Question Hierarchy
  • 3.3 Co-Attention
  • 3.4 Encoding for Predicting Answers
  • 4 Experiment
  • 4.1 Datasets
  • 4.2 Setup
  • 4.3 Results and Analysis
  • 4.4 Ablation Study
  • 4.5 Qualitative Results
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Hierarchical Question Representation via Multi-Scale 1D Convolutions and Recurrent Encoding

    model/method

    The hierarchical question representation encodes a question of TT words at three distinct linguistic levels: word level, phrase level, and question level.

    1. Word Level: Given one-hot encodings Q={q1,…,qT}Q = \{q_1, \dots, q_T\}, words are embedded into a continuous vector space through an end-to-end learned embedding matrix, yielding word representations:

    Qw={q1w,q2w,…,qTw},qtw∈RdQ^w = \{q_1^w, q_2^w, \dots, q_T^w\}, \quad q_t^w \in \mathbb{R}^d

    1. Phrase Level: To capture contextual phrase structures of varying lengths, 1D convolutions with temporal filter window sizes s∈{1,2,3}s \in \{1, 2, 3\} (corresponding to unigram, bigram, and trigram windows) are applied over the word vectors. For word position tt, the convolution response for window size ss is:

    q^s,tp=tanh⁡(Wcsqt:t+s−1w),s∈{1,2,3}\hat{q}_{s,t}^p = \tanh(W_c^s q_{t:t+s-1}^w), \quad s \in \{1, 2, 3\}

    where WcsW_c^s denotes the convolution weight parameters and QwQ^w is zero-padded to preserve sequence length TT. Phrase features at each position tt are obtained via max-pooling across the different nn-gram representations:

    qtp=max⁡(q^1,tp,q^2,tp,q^3,tp),t∈{1,2,…,T}q_t^p = \max(\hat{q}_{1,t}^p, \hat{q}_{2,t}^p, \hat{q}_{3,t}^p), \quad t \in \{1, 2, \dots, T\}

    This pooling adaptively selects the most salient nn-gram feature at each time step while preserving the original sequence length and token order.

    1. Question Level: A Long Short-Term Memory (LSTM) network takes the max-pooled phrase sequence {qtp}t=1T\{q_t^p\}_{t=1}^T as input. The question-level feature at each step tt is the LSTM hidden state:

    qts=LSTM(qtp,qt−1s),t∈{1,2,…,T}q_t^s = \text{LSTM}(q_t^p, q_{t-1}^s), \quad t \in \{1, 2, \dots, T\}

    This results in three sequence feature matrices Qw,Qp,Qs∈Rd×TQ^w, Q^p, Q^s \in \mathbb{R}^{d \times T} representing the question at word, phrase, and sentence levels.

  2. Knowl 2 — Parallel Question-Image Co-Attention Mechanism

    model/method

    Parallel co-attention simultaneously calculates spatial visual attention on the image and linguistic attention across question tokens by constructing an affinity matrix connecting all image regions and question words.

    Given an image feature map V={v1,…,vN}∈Rd×NV = \{v_1, \dots, v_N\} \in \mathbb{R}^{d \times N} across NN spatial locations and a question representation matrix Q={q1,…,qT}∈Rd×TQ = \{q_1, \dots, q_T\} \in \mathbb{R}^{d \times T} at a given hierarchical level, the affinity matrix C∈RT×NC \in \mathbb{R}^{T \times N} is computed as:

    C=tanh⁡(QTWbV)C = \tanh(Q^T W_b V)

    where Wb∈Rd×dW_b \in \mathbb{R}^{d \times d} is a learnable parameter matrix. The affinity matrix CC maps question attention space to image attention space (and CTC^T maps image attention space to question attention space).

    Rather than applying a max operation over affinities, CC is treated as a feature transformation to compute attention spaces:

    Hv=tanh⁡(WvV+(WqQ)C),Hq=tanh⁡(WqQ+(WvV)CT)H^v = \tanh(W_v V + (W_q Q)C), \quad H^q = \tanh(W_q Q + (W_v V) C^T)

    av=softmax(whvTHv),aq=softmax(whqTHq)a^v = \text{softmax}(w_{hv}^T H^v), \quad a^q = \text{softmax}(w_{hq}^T H^q)

    where Wv,Wq∈Rk×dW_v, W_q \in \mathbb{R}^{k \times d} and whv,whq∈Rkw_{hv}, w_{hq} \in \mathbb{R}^k are learnable weights. The resulting vectors av∈RNa^v \in \mathbb{R}^N and aq∈RTa^q \in \mathbb{R}^T represent attention probabilities over image locations and question tokens, respectively.

    The co-attended feature representations v^∈Rd\hat{v} \in \mathbb{R}^d and q^∈Rd\hat{q} \in \mathbb{R}^d are computed as probability-weighted sums:

    v^=∑n=1Nanvvn,q^=∑t=1Tatqqt\hat{v} = \sum_{n=1}^N a_n^v v_n, \quad \hat{q} = \sum_{t=1}^T a_t^q q_t

    Parallel co-attention is applied independently across word, phrase, and question hierarchy levels to produce pairs (v^w,q^w)(\hat{v}^w, \hat{q}^w), (v^p,q^p)(\hat{v}^p, \hat{q}^p), and (v^s,q^s)(\hat{v}^s, \hat{q}^s).

  3. Knowl 3 — Alternating Question-Image Co-Attention Mechanism

    model/method

    Alternating co-attention sequentially alternates between generating image and question attention through a generic attention operator x^=A(X;g)\hat{x} = \mathcal{A}(X; g), which takes a feature matrix X={x1,…,xM}∈Rd×MX = \{x_1, \dots, x_M\} \in \mathbb{R}^{d \times M} and an attention guidance vector g∈Rdg \in \mathbb{R}^d as inputs:

    H=tanh⁡(WxX+(Wgg)1T)H = \tanh(W_x X + (W_g g)\mathbf{1}^T)

    ax=softmax(whxTH)a^x = \text{softmax}(w_{hx}^T H)

    x^=∑i=1Maixxi\hat{x} = \sum_{i=1}^M a_i^x x_i

    where 1∈RM\mathbf{1} \in \mathbb{R}^M is a vector of ones, Wx,Wg∈Rk×dW_x, W_g \in \mathbb{R}^{k \times d}, whx∈Rkw_{hx} \in \mathbb{R}^k, and ax∈RMa^x \in \mathbb{R}^M is the attention distribution over the elements of XX.

    For an image feature map V∈Rd×NV \in \mathbb{R}^{d \times N} and a question sequence matrix Q∈Rd×TQ \in \mathbb{R}^{d \times T} at any hierarchy level, alternating co-attention executes three successive steps:

    1. Question Summarization: Summarize the question into a single vector s^\hat{s} without external guidance:

    s^=A(Q;0)\hat{s} = \mathcal{A}(Q; \mathbf{0})

    1. Image Attention: Attend to spatial image regions using the question summary s^\hat{s} as guidance:

    v^=A(V;s^)\hat{v} = \mathcal{A}(V; \hat{s})

    1. Image-Guided Question Attention: Re-attend to the question tokens using the attended image representation v^\hat{v} as guidance:

    q^=A(Q;v^)\hat{q} = \mathcal{A}(Q; \hat{v})

    This process is executed independently at the word, phrase, and question levels to obtain (v^w,q^w)(\hat{v}^w, \hat{q}^w), (v^p,q^p)(\hat{v}^p, \hat{q}^p), and (v^s,q^s)(\hat{v}^s, \hat{q}^s).

  4. Knowl 4 — Hierarchical Multi-Modal Feature Fusion for Answer Prediction

    model/method

    Visual Question Answering is formulated as a multi-class classification problem over a candidate answer vocabulary. Co-attended image and question feature pairs from the three hierarchical levels—word level (v^w,q^w)(\hat{v}^w, \hat{q}^w), phrase level (v^p,q^p)(\hat{v}^p, \hat{q}^p), and question level (v^s,q^s)(\hat{v}^s, \hat{q}^s)—are fused recursively from bottom to top using a multi-layer perceptron:

    hw=tanh⁡(Ww(q^w+v^w))h^w = \tanh(W_w(\hat{q}^w + \hat{v}^w))

    hp=tanh⁡(Wp[(q^p+v^p),hw])h^p = \tanh(W_p[(\hat{q}^p + \hat{v}^p), h^w])

    hs=tanh⁡(Ws[(q^s+v^s),hp])h^s = \tanh(W_s[(\hat{q}^s + \hat{v}^s), h^p])

    p=softmax(Whhs)p = \text{softmax}(W_h h^s)

    where [⋅,⋅][\cdot, \cdot] denotes vector concatenation, Ww∈Rd×dW_w \in \mathbb{R}^{d \times d}, Wp∈Rd×2dW_p \in \mathbb{R}^{d \times 2d}, Ws∈Rdhidden×2dW_s \in \mathbb{R}^{d_{\text{hidden}} \times 2d}, and Wh∈RC×dhiddenW_h \in \mathbb{R}^{C \times d_{\text{hidden}}} are learnable weight matrices (with biases omitted), CC is the number of candidate answer classes, and p∈RCp \in \mathbb{R}^C is the predicted probability distribution over the answer set.

  5. Knowl 5 — Experimental and Training Configuration for Hierarchical Co-Attention

    experimental setup

    The hierarchical co-attention architecture is trained with the following hyperparameters and settings:

    • Optimization: RMSprop optimizer with a base learning rate of 4×10−44 \times 10^{-4}, momentum of 0.990.99, and weight decay of 1×10−81 \times 10^{-8}.
    • Batch Size and Epochs: Batch size of 300, trained for up to 256 epochs with early stopping if validation accuracy does not improve for 5 consecutive epochs.
    • Image Preprocessing and Features: Images are rescaled to 448×448448 \times 448. Visual features V∈Rd×196V \in \mathbb{R}^{d \times 196} (14×1414 \times 14 grid, N=196N=196) are extracted from the activations of the last pooling layer of VGGNet or ResNet.
    • Dimensions: Word embedding dimension and intermediate hidden layer dimensions dd and kk are set to 512. The hidden dimension of WsW_s (dhiddend_{\text{hidden}}) is set to 512 for COCO-QA and 1024 for the larger VQA dataset.
    • Regularization: Dropout with probability 0.50.5 is applied to each layer.
    • Answer Vocabularies: For the VQA dataset, the top 1000 most frequent training+validation answers are used as candidate classes (covering 86.54% of answer occurrences).
  6. Knowl 6 — Benchmark Evaluation Results on the VQA Dataset

    data/table

    Hierarchical co-attention models were evaluated on the VQA benchmark under open-ended and multiple-choice settings across answer subsets: Yes/No (Y/N), Number (Num), and Other.

    Open-Ended Multiple-Choice
    test-dev test-std test-dev test-std
    Method Y/N Num Other All All Y/N Num Other All All
    LSTM Q+I 80.5 36.8 43.0 57.8 58.2 80.5 38.2 53.0 62.7 63.1
    Region Sel. - - - - - 77.6 34.3 55.8 62.4 -
    SMem 80.9 37.3 43.1 58.0 58.2 - - - - -
    SAN 79.3 36.6 46.1 58.7 58.9 - - - - -
    FDA 81.1 36.2 45.8 59.2 59.5 81.5 39.0 54.7 64.0 64.2
    DMN+ 80.5 36.8 48.3 60.3 60.4 - - - - -
    Oursp^p + VGG 79.5 38.7 48.3 60.1 - 79.5 39.8 57.4 64.6 -
    Oursa^a + VGG 79.6 38.4 49.1 60.5 - 79.7 40.1 57.9 64.9 -
    Oursa^a + ResNet 79.7 38.7 51.7 61.8 62.1 79.7 40.0 59.8 65.8 66.1

    Alternating co-attention with ResNet (Oursa^a + ResNet) achieves 62.1% on open-ended test-std and 66.1% on multiple-choice test-std, outperforming previous methods (DMN+ at 60.4% open-ended and FDA at 64.2% multiple-choice). Using identical ResNet features, Oursa^a + ResNet outperforms FDA by 1.8% on open-ended test-dev.

  7. Knowl 7 — Benchmark Evaluation Results on the COCO-QA Dataset

    data/table

    Hierarchical co-attention models were evaluated on the COCO-QA dataset across four question categories (Object, Number, Color, Location) reporting classification accuracy as well as Wu-Palmer similarity thresholds (WUPS 0.9 and WUPS 0.0).

    Method Object Number Color Location Accuracy WUPS0.9 WUPS0.0
    2-VIS+BLSTM 58.2 44.8 49.5 47.3 55.1 65.3 88.6
    IMG-CNN - - - - 58.4 68.5 89.7
    SAN (2, CNN) 64.5 48.6 57.9 54.0 61.6 71.6 90.9
    Oursp^p + VGG 65.6 49.6 61.5 56.8 63.3 73.0 91.3
    Oursa^a + VGG 65.6 48.9 59.8 56.7 62.9 72.8 91.3
    Oursa^a + ResNet 68.0 51.0 62.9 58.8 65.4 75.1 92.0

    Parallel co-attention with VGG (Oursp^p + VGG) outperforms alternating co-attention (Oursa^a + VGG) on COCO-QA (63.3% vs 62.9%). Utilizing ResNet visual features elevates accuracy to 65.4%, improving upon the prior state of the art (SAN at 61.6%) by 3.8%.

  8. Knowl 8 — Ablation Analysis of Question Hierarchy and Attention Modalities

    data/table

    An ablation study conducted on the VQA validation dataset using alternating co-attention with VGG features (Oursa^a + VGG) evaluates the contribution of each attention mechanism and linguistic hierarchy level.

    Method Y/N Num Other All
    LSTM Q+I (Baseline) 79.8 32.9 40.7 54.3
    Image Atten (No question attention) 79.8 33.9 43.6 55.9
    Question Atten (No image attention) 79.4 33.3 41.7 54.8
    W/O Q-Atten (Uniform question-level attention) 79.6 32.1 42.9 55.3
    W/O P-Atten (Uniform phrase-level attention) 79.5 34.1 45.4 56.7
    W/O W-Atten (Uniform word-level attention) 79.6 34.4 45.6 56.8
    Full Model 79.6 35.0 45.7 57.0

    Key observations from this ablation include:

    • Attention Modalities: Image attention alone (55.9%) outperforms the holistic baseline (54.3%), and question attention alone (54.8%) also improves over baseline. Combining both into full co-attention yields 57.0% (a 1.1% gain over image attention alone).
    • Hierarchy Sensitivity: Removing question-level attention causes the largest accuracy drop (1.7% down to 55.3%), removing phrase-level attention causes a 0.3% drop (to 56.7%), and removing word-level attention causes a 0.2% drop (to 56.8%). This indicates that levels closer to the answer prediction layer have the highest performance impact.
    • Multi-Round Iteration: Performing one additional round of alternating co-attention improves validation accuracy by an additional 0.3%.
  9. Knowl 9 — Comparative Dynamics and Trade-Offs of Parallel versus Alternating Co-Attention

    empirical result

    Parallel and alternating co-attention present contrasting optimization characteristics and performance trade-offs:

    • Parallel Co-Attention: Computes a dense token-region affinity matrix C=tanh⁡(QTWbV)C = \tanh(Q^T W_b V) via dot products between every word embedding and image feature vector. This pairwise compression of two vectors into scalar similarity values makes parallel co-attention harder to optimize during backpropagation. However, it prevents order-dependent bias between modalities and achieves higher performance on COCO-QA (63.3% vs 62.9% accuracy with VGG).
    • Alternating Co-Attention: Decouples cross-modal reasoning into sequential attention steps guided by intermediate single-vector summaries (question summary s^→\hat{s} \to attended image v^→\hat{v} \to attended question q^\hat{q}). While easier to train, alternating co-attention can accumulate attention errors sequentially across iterations. It achieves slightly higher performance on the larger VQA dataset (60.5% vs 60.1% test-dev open-ended accuracy with VGG).

Coverage note — Qualitative attention visualizations and sample success/failure visualizations were omitted as they provide visual illustrations of the empirical results without contributing independent quantitative findings or new model components.

References

  1. 1.Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Deep compositional question answering with neural module networks. In CVPR, 2016.
  2. 2.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In ICCV, 2015.
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In ICLR, 2015.
  4. 4.R. Collobert, K. Kavukcuoglu, and C. Farabet. Torch7: A matlab-like environment for machine learning. In BigLearn, NIPS Workshop, 2011.
  5. 5.Abhishek Das, Harsh Agrawal, C Lawrence Zitnick, Devi Parikh, and Dhruv Batra. Human attention in visual question answering: Do humans and deep networks look at the same regions? arXiv preprint arXiv:1606.03556, 2016.
  6. 6.Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv preprint arXiv:1606.01847, 2016.
  7. 7.Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. Are you talking to a machine? dataset and methods for multilingual image question answering. In NIPS, 2015.
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  9. 9.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. In NIPS, 2015.
  10. 10.Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen. Convolutional neural network architectures for matching natural language sentences. In NIPS, 2014.
  11. 11.Ilija Ilievski, Shuicheng Yan, and Jiashi Feng. A focused dynamic attention model for visual question answering. arXiv:1604.01485, 2016.
  12. 12.Jin-Hwa Kim, Sang-Woo Lee, Dong-Hyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Multimodal residual learning for visual qa. arXiv preprint arXiv:1606.01455, 2016.
  13. 13.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. arXiv preprint arXiv:1602.07332, 2016.
  14. 14.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  15. 15.Lin Ma, Zhengdong Lu, and Hang Li. Learning to answer questions from image using convolutional neural network. In AAAI, 2016.
  16. 16.Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. Ask your neurons: A neural-based approach to answering questions about images. In ICCV, 2015.
  17. 17.Mengye Ren, Ryan Kiros, and Richard Zemel. Exploring models and data for image question answering. In NIPS, 2015.
  18. 18.Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, and Phil Blunsom. Reasoning about entailment with neural attention. In ICLR, 2016.
  19. 19.Cicero dos Santos, Ming Tan, Bing Xiang, and Bowen Zhou. Attentive pooling networks. arXiv preprint arXiv:1602.03609, 2016.
  20. 20.Kevin J Shih, Saurabh Singh, and Derek Hoiem. Where to look: Focus regions for visual question answering. In CVPR, 2016.
  21. 21.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  22. 22.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  23. 23.Caiming Xiong, Stephen Merity, and Richard Socher. Dynamic memory networks for visual and textual question answering. In ICML, 2016.
  24. 24.Huijuan Xu and Kate Saenko. Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. arXiv preprint arXiv:1511.05234, 2015.
  25. 25.Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In CVPR, 2016.
  26. 26.Wenpeng Yin, Hinrich Schütze, Bing Xiang, and Bowen Zhou. Abcnn: Attention-based convolutional neural network for modeling sentence pairs. In ACL, 2016.
  27. 27.Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Yin and yang: Balancing and answering binary visual questions. arXiv preprint arXiv:1511.05099, 2015.
  28. 28.Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7w: Grounded question answering in images. In CVPR, 2016.
  29. 29.C Lawrence Zitnick, Aishwarya Agrawal, Stanislaw Antol, Margaret Mitchell, Dhruv Batra, and Devi Parikh. Measuring machine intelligence through visual question answering. AI Magazine, 37(1), 2016.

Citation

MLA
Lu, J., et al. “Hierarchical Question-Image Co-Attention for Visual Question Answering”. arXiv, 2016, http://arxiv.org/abs/1606.00061v5.
APA
Lu, J., Yang, J., Batra, D., & Parikh, D. (2016). Hierarchical Question-Image Co-Attention for Visual Question Answering. arXiv. http://arxiv.org/abs/1606.00061v5
Chicago
Lu, J., J. Yang, D. Batra, and D. Parikh. 2016. “Hierarchical Question-Image Co-Attention for Visual Question Answering”. arXiv. http://arxiv.org/abs/1606.00061v5.
Harvard
Lu, J. et al. (2016) “Hierarchical Question-Image Co-Attention for Visual Question Answering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1606.00061v5.
Vancouver
1. Lu J, Yang J, Batra D, Parikh D (2016) Hierarchical Question-Image Co-Attention for Visual Question Answering. arXiv

BibTeX

@article{lu2016hierarchical,
  title = {Hierarchical Question-Image Co-Attention for Visual Question Answering},
  author = {Lu, Jiasen and Yang, Jianwei and Batra, Dhruv and Parikh, Devi},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1606.00061v5},
  eprint = {1606.00061}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors