Image Captioning with Semantic Attention

Quanzeng YouHailin JinZhaowen WangChen FangJiebo Luo

article2016CVPR1,797 citations

Proposes a semantic attention framework for image captioning that selectively fuses bottom-up concept proposals with top-down visual features in recurrent neural networks to generate more accurate descriptions.

Listen

Automatically generating natural language descriptions for visual media is essential for real-world applications such as assistive technologies for visually impaired users and broad multimedia analysis. Prior automated captioning methods rely either on top-down processing, which translates a whole-image overview into sentences but misses subtle elements, or bottom-up processing, which identifies distinct objects and terms without a unified framework for generating natural text. The article sets out to evaluate and demonstrate an automated image captioning framework that bridges this gap by combining global visual overviews with localized semantic concept detection through a selective attention feedback system.

The researchers developed an architecture that blends top-down and bottom-up information inside a sequence-generating language model. A vision neural network first captures a global visual feature to initialize the sentence generator, while a separate visual attribute detector extracts candidate keywords and concepts. At each word-generation step, the system applies two complementary attention layers: an input attention model that weights candidate visual attributes according to preceding words, and an output attention model that weights attributes against the internal state before predicting the next word. The evaluation tested parametric and non-parametric attribute detectors across standard benchmark datasets containing tens of thousands of human-annotated images, assessing accuracy across established language-generation metrics.

The findings show that combining semantic attention with attribute detection outperforms existing state-of-the-art captioning systems. First, using fully convolutional networks to detect local attributes yielded superior caption accuracy compared to neighbor-retrieval or ranking-loss approaches. Second, applying dynamic attention to detected concepts generated better descriptions than simple static fusion methods, such as concatenating or taking maximum values of attribute vectors. Third, integrating both input and output attention layers consistently enhanced caption quality over using either layer alone. In public benchmark tests, the proposed model achieved top-ranking performance against established competitive baselines.

These results demonstrate that semantic attention resolves the long-standing trade-off between global visual context and fine-grained object recognition in automated description tasks. By operating over detected words and concepts rather than fixed spatial image locations, systems can leverage external image data and text semantics more flexibly. This offers substantial improvements in output relevance and descriptive accuracy for visual computing systems without requiring entirely re-engineered visual models.

Decision-makers and practitioners deploying visual accessibility or search systems should consider hybrid semantic attention pipelines that incorporate local attribute detection. The primary risk and limitation identified is that false positive attribute predictions can mislead the language model into hallucinating incorrect background objects. Organizations should focus subsequent efforts on enhancing visual concept accuracy, exploring phrase-level representations, and validating attention pipelines across diverse operational domains before production deployment.

Cover for Image Captioning with Semantic Attention

Abstract

Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fields: computer vision and natural language processing. Existing approaches are either top-down, which start from a gist of an image and convert it into words, or bottom-up, which come up with words describing various aspects of an image and then combine them. In this paper, we propose a new algorithm that combines both approaches through a model of semantic attention. Our algorithm learns to selectively attend to semantic concept proposals and fuse them into hidden states and outputs of recurrent neural networks. The selection and fusion form a feedback connecting the top-down and bottom-up computation. We evaluate our algorithm on two public benchmarks: Microsoft COCO and Flickr30K. Experimental results show that our algorithm significantly outperforms the state-of-the-art approaches consistently across different evaluation metrics.

Table of Contents

  • 1 Introduction
  • 1.1 Main contributions
  • 2 Related work
  • 3 Semantic attention for image captioning
  • 3.1 Overall framework
  • 3.2 Input attention model
  • 3.3 Output attention model
  • 3.4 Model learning
  • 4 Visual attribute prediction
  • 4.1 Non-parametric attribute prediction
  • 4.2 Parametric attribute prediction
  • 5 Experiments
  • 5.1 Datasets and settings
  • 5.2 Performance on MS-COCO
  • 5.3 Performance on Flickr30k
  • 5.4 Visualization of attended attributes
  • 5.5 Analysis of attention model
  • 5.6 The role of visual attributes
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Semantic Attention Architecture for Image Captioning

    model/method

    The semantic attention architecture combines a top-down global visual representation vv with a bottom-up set of detected semantic visual attributes or concepts {Ai}i=1N\{A_i\}_{i=1}^N to generate image captions using a Recurrent Neural Network (RNN).

    Let Y\mathcal{Y} be the vocabulary dictionary. The global visual feature vv extracted from a Convolutional Neural Network (CNN) initializes the RNN input at step t=0t=0, providing visual context. At each subsequent time step t>0t > 0, the state transition and word probability distribution are computed as follows:

    x0=φ0(v)=Wx,vvx_0 = \varphi_0(v) = W_{x,v} v

    ht=RNN(ht−1,xt)h_t = \text{RNN}(h_{t-1}, x_t)

    Yt∼pt=ϕ(ht,{Ai})Y_t \sim p_t = \phi(h_t, \{A_i\})

    xt=φ(Yt−1,{Ai}),t>0x_t = \varphi(Y_{t-1}, \{A_i\}), \quad t > 0

    where:

    • vv is the global visual feature vector extracted from the CNN.
    • Wx,vW_{x,v} is a linear projection weight matrix for initializing input x0x_0.
    • ht∈Rnh_t \in \mathbb{R}^n is the RNN hidden state vector at step tt.
    • xt∈Rmx_t \in \mathbb{R}^m is the input vector fed into the RNN at step tt.
    • Yt∈YY_t \in \mathcal{Y} is the word sampled at step tt.
    • pt∈R∣Y∣p_t \in \mathbb{R}^{|\mathcal{Y}|} is the probability distribution over the vocabulary dictionary Y\mathcal{Y} at step tt.
    • φ\varphi is the input semantic attention model that modulates visual attributes based on the previously generated word Yt−1Y_{t-1}.
    • ϕ\phi is the output semantic attention model that modulates visual attributes based on the current hidden state hth_t.
  2. Knowl 2 — Input Semantic Attention Model Formulation

    model/method

    The input semantic attention model φ\varphi computes relevance scores between the previously generated word Yt−1Y_{t-1} and each detected visual attribute Ai∈{Ai}i=1NA_i \in \{A_i\}_{i=1}^N, and fuses the attended attributes into the RNN input vector xtx_t for t>0t > 0.

    Let yt−1∈R∣Y∣y_{t-1} \in \mathbb{R}^{|\mathcal{Y}|} and yi∈R∣Y∣y^i \in \mathbb{R}^{|\mathcal{Y}|} denote the one-hot representations of the word Yt−1Y_{t-1} and the attribute AiA_i within vocabulary Y\mathcal{Y}, and let E∈Rd×∣Y∣E \in \mathbb{R}^{d \times |\mathcal{Y}|} denote a word embedding matrix (d≪∣Y∣d \ll |\mathcal{Y}|). The attention score αti\alpha_t^i for attribute AiA_i at time step tt is calculated via a bilinear form and normalized over all candidate attributes using a softmax function:

    αti∝exp⁡(yt−1TETUEyi)\alpha_t^i \propto \exp\left(y_{t-1}^T E^T U E y^i\right)

    where U∈Rd×dU \in \mathbb{R}^{d \times d} is a learnable bilinear parameter matrix.

    The input vector xt∈Rmx_t \in \mathbb{R}^m is formed by projecting the sum of the embedded previous word and the weighted attribute embeddings:

    xt=Wx,Y(Eyt−1+diag(wx,A)∑i=1NαtiEyi)x_t = W_{x,Y} \left( E y_{t-1} + \text{diag}(w^{x,A}) \sum_{i=1}^N \alpha_t^i E y^i \right)

    where:

    • Wx,Y∈Rm×dW_{x,Y} \in \mathbb{R}^{m \times d} is the projection matrix mapping to the RNN input space.
    • diag(wx,A)\text{diag}(w^{x,A}) is a diagonal matrix parameterized by wx,A∈Rdw^{x,A} \in \mathbb{R}^d, scaling the relative importance of visual attributes across word embedding dimensions.
  3. Knowl 3 — Output Semantic Attention Model Formulation

    model/method

    The output semantic attention model ϕ\phi modulates the contribution of detected visual attributes {Ai}i=1N\{A_i\}_{i=1}^N using the current RNN hidden state ht∈Rnh_t \in \mathbb{R}^n to produce the word prediction distribution ptp_t at time step tt.

    Let yi∈R∣Y∣y^i \in \mathbb{R}^{|\mathcal{Y}|} be the one-hot vector for attribute AiA_i, E∈Rd×∣Y∣E \in \mathbb{R}^{d \times |\mathcal{Y}|} be the word embedding matrix, and σ\sigma be the activation function connecting the input node to the hidden state of the RNN. The attention weight βti\beta_t^i for attribute AiA_i is computed via a bilinear function with hth_t:

    βti∝exp⁡(htTVσ(Eyi))\beta_t^i \propto \exp\left( h_t^T V \sigma(E y^i) \right)

    where V∈Rn×dV \in \mathbb{R}^{n \times d} is a bilinear parameter matrix, and βti\beta_t^i is normalized over all i∈{1,…,N}i \in \{1, \dots, N\} with softmax.

    The predictive distribution pt∈R∣Y∣p_t \in \mathbb{R}^{|\mathcal{Y}|} over vocabulary Y\mathcal{Y} is generated using transposed weight sharing ETE^T:

    pt∝exp⁡(ETWY,h(ht+diag(wY,A)∑i=1Nβtiσ(Eyi)))p_t \propto \exp\left( E^T W^{Y,h} \left( h_t + \text{diag}(w^{Y,A}) \sum_{i=1}^N \beta_t^i \sigma(E y^i) \right) \right)

    where:

    • WY,h∈Rd×nW^{Y,h} \in \mathbb{R}^{d \times n} is a projection matrix.
    • diag(wY,A)\text{diag}(w^{Y,A}) is a diagonal matrix defined by parameter vector wY,A∈Rnw^{Y,A} \in \mathbb{R}^n, which balances the relative importance of visual attributes in each dimension of the RNN state space.
  4. Knowl 4 — Attention Regularization for Completeness and Sparsity

    equation

    To train the semantic attention image captioning system, the negative log-likelihood of ground-truth target words is combined with regularizers on the input attention weight matrix α\alpha and output attention weight matrix β\beta:

    min⁡ΘA,ΘR−∑tlog⁡p(Yt)+g(α)+g(β)\min_{\Theta_A, \Theta_R} -\sum_{t} \log p(Y_t) + g(\alpha) + g(\beta)

    where ΘA={U,V,W∗,∗,w∗,∗}\Theta_A = \{U, V, W^{*,*}, w^{*,*}\} denotes the attention model parameters, ΘR\Theta_R denotes the RNN parameters, and α,β∈RT×N\alpha, \beta \in \mathbb{R}^{T \times N} are matrices whose (t,i)(t,i)-th entries are αti\alpha_t^i and βti\beta_t^i respectively (TT is sentence length, NN is the number of candidate attributes).

    The regularization function g(α)g(\alpha) (and analogously g(β)g(\beta)) enforces completeness of attribute coverage across the whole sentence and sparsity of attention at each individual time step using mixed matrix norms:

    g(α)=∥α∥1,p+∥αT∥q,1=[∑i=1N(∑t=1Tαti)p]1/p+∑t=1T[∑i=1N(αti)q]1/qg(\alpha) = \|\alpha\|_{1,p} + \|\alpha^T\|_{q,1} = \left[ \sum_{i=1}^N \left( \sum_{t=1}^T \alpha_t^i \right)^p \right]^{1/p} + \sum_{t=1}^T \left[ \sum_{i=1}^N (\alpha_t^i)^q \right]^{1/q}

    with hyperparameters p>1p > 1 and 0<q<10 < q < 1. The first term with p>1p > 1 penalizes excessive accumulated attention on any single attribute across the entire sequence. The second term with 0<q<10 < q < 1 penalizes diverting attention over multiple attributes at any single time step. The default parameters are set to p=2p = 2 and q=0.5q = 0.5.

  5. Knowl 5 — Visual Attribute Prediction Methods

    model/method

    Semantic visual attributes {Ai}\{A_i\} are detected as candidate concepts from input images using either non-parametric retrieval or parametric deep models:

    1. Non-parametric Attribute Prediction (kk-NN): Given a query image, its nearest visual neighbors in the training set are retrieved based on Euclidean distance in GoogleNet feature space. Term-Frequency (TF) is computed over the ground-truth captions of the retrieved neighbor images, and the most frequent words are selected as candidate attributes.

    2. Parametric Multi-Label Classification (Ranking Loss, RK): A set of predefined visual attributes is constructed from the most frequent vocabulary words in training captions. A deep convolutional network is trained using a multi-label ranking loss to assign relevance scores between an input image and predefined attribute categories.

    3. Parametric Patch-Level Attribute Detection (Fully Convolutional Network, FCN): A Fully Convolutional Network is trained on local image patches to predict visual attributes, producing patch-level relevance scores that are pooled across the image.

    The top N=10N = 10 detected attributes ranked by detection confidence are retained to construct {Ai}\{A_i\}.

  6. Knowl 6 — Comparative Image Captioning Performance on MS-COCO and Flickr30k

    data/table

    The semantic attention model combined with FCN visual attribute detection (Ours-ATT-FCN) outperforms state-of-the-art baselines (Google NIC, m-RNN, LRCN, MSR/CMU, and Toronto spatial attention) across BLEU (B-1 to B-4) and METEOR on both Flickr30k and MS-COCO validation splits.

    Flickr30k MS-COCO
    Model B-1 B-2 B-3 B-4 METEOR B-1 B-2 B-3 B-4 METEOR
    Google NIC 0.663 0.423 0.277 0.183 – 0.666 0.451 0.304 0.203 –
    m-RNN 0.60 0.41 0.28 0.19 – 0.67 0.49 0.35 0.25 –
    LRCN 0.587 0.39 0.25 0.165 – 0.628 0.442 0.304 0.21 –
    MSR/CMU – – – 0.126 0.164 – – – 0.19 0.204
    Toronto 0.669 0.439 0.296 0.199 0.185 0.718 0.504 0.357 0.250 0.230
    Ours-CON-kk-NN 0.619 0.426 0.291 0.197 0.179 0.675 0.503 0.373 0.279 0.227
    Ours-CON-RK 0.623 0.432 0.295 0.200 0.179 0.647 0.472 0.338 0.237 0.204
    Ours-CON-FCN 0.639 0.447 0.309 0.213 0.188 0.700 0.532 0.398 0.300 0.238
    Ours-MAX-kk-NN 0.622 0.426 0.287 0.193 0.178 0.673 0.501 0.371 0.279 0.227
    Ours-MAX-RK 0.623 0.429 0.294 0.202 0.178 0.655 0.478 0.344 0.245 0.208
    Ours-MAX-FCN 0.633 0.444 0.306 0.21 0.181 0.699 0.530 0.398 0.301 0.240
    Ours-ATT-kk-NN 0.618 0.428 0.290 0.195 0.172 0.676 0.505 0.375 0.281 0.227
    Ours-ATT-RK 0.617 0.424 0.286 0.193 0.177 0.679 0.506 0.375 0.282 0.231
    Ours-ATT-FCN 0.647 0.460 0.324 0.230 0.189 0.709 0.537 0.402 0.304 0.243

    Among attribute detectors, FCN consistently delivers superior features compared to kk-NN and Ranking Loss (RK). Among concept integration strategies (CON: concatenation, MAX: element-wise max, ATT: semantic attention), dynamic attention achieves the highest accuracy across higher n-gram BLEU scores and METEOR.

  7. Knowl 7 — MS-COCO Online Server Evaluation Benchmark

    data/table

    Performance of the semantic attention model (ATT, corresponding to Ours-ATT-FCN) on the official online MS-COCO test server under 5 reference captions (c5) and 40 reference captions (c40), compared against top competing methods (OV: Oriol Vinyals, MSR Cap: MSR Captivator, mRNN: mRNN share.JMao).

    B-1 B-2 B-3 B-4 METEOR ROUGE-L CIDEr
    Alg c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40
    ATT 0.731 0.920 0.565 0.815 0.424 0.709 0.316 0.599 0.250 0.335 0.535 0.682 0.943 0.958
    OV 0.713 0.895 0.542 0.802 0.407 0.694 0.309 0.587 0.254 0.346 0.530 0.682 0.943 0.946
    MSR Cap 0.715 0.907 0.543 0.819 0.407 0.710 0.308 0.601 0.248 0.339 0.526 0.680 0.931 0.937
    mRNN 0.716 0.890 0.545 0.798 0.404 0.687 0.299 0.575 0.242 0.325 0.521 0.666 0.917 0.935

    The semantic attention model achieved top-1 rankings across several evaluation metrics at submission time, showing particularly strong performance on BLEU-1 through BLEU-4 (c5), ROUGE-L, and CIDEr.

  8. Knowl 8 — Ablation of Input and Output Attention Mechanisms

    empirical result

    Evaluating the input attention model (modulating previous word Yt−1Y_{t-1} and attributes) and output attention model (modulating RNN hidden state hth_t and attributes) independently on the MS-COCO validation dataset using ground-truth attributes reveals complementary benefits:

    Model Variant B-1 B-2 B-3 B-4 METEOR ROUGE-L CIDEr
    Input Only 0.88 0.75 0.62 0.50 0.33 0.65 1.56
    Output Only 0.89 0.76 0.62 0.50 0.33 0.65 1.58
    Full (Input + Output) 0.91 0.79 0.65 0.53 0.34 0.67 1.68

    While output attention alone performs slightly better than input attention alone (e.g., CIDEr 1.58 vs 1.56), combining both attention modules leads to a substantial improvement across all metrics (e.g., CIDEr increases to 1.68 and B-4 to 0.53). This demonstrates that input and output layers attend to different aspects of semantic concepts during caption analysis and generation.

  9. Knowl 9 — Captioning Performance Limits with Ground-Truth Attributes

    data/table

    Supplying ground-truth visual attributes (the most frequent words in reference captions) to the captioning pipeline provides an empirical upper bound on performance across Flickr30k and MS-COCO:

    Dataset Model B-1 B-2 B-3 B-4 METEOR ROUGE-L CIDEr
    Flickr30k Ours-GT-ATT 0.824 0.679 0.534 0.412 0.269 0.588 0.949
    Ours-GT-MAX 0.719 0.542 0.396 0.283 0.230 0.529 0.747
    Ours-GT-CON 0.708 0.534 0.388 0.276 0.222 0.516 0.685
    Google NIC 0.663 0.423 0.277 0.183 – – –
    Toronto 0.669 0.439 0.296 0.199 0.185 – –
    MS-COCO Ours-GT-ATT 0.910 0.786 0.654 0.534 0.341 0.667 1.685
    Ours-GT-MAX 0.790 0.635 0.494 0.379 0.279 0.580 1.161
    Ours-GT-CON 0.766 0.617 0.484 0.377 0.279 0.582 1.237
    Google NIC 0.666 0.451 0.304 0.203 – – –
    Toronto 0.718 0.504 0.357 0.250 0.230 – –

    When supplied with high-quality visual attributes, semantic attention (Ours-GT-ATT) achieves a CIDEr score of 1.685 and B-4 of 0.534 on MS-COCO, demonstrating that the semantic attention mechanism can leverage visual attributes to achieve significant performance gains over simple pooling schemes (MAX, CON) and previous models.

  10. Knowl 10 — Vulnerability to Erroneous Visual Attributes

    limitation

    The semantic attention model is sensitive to false or misleading visual attributes detected in the bottom-up stage. Irrelevant detected attributes can disrupt attention and cause the LSTM to generate inaccurate descriptions. For example:

    1. Detecting background concepts such as 'clock' can distract the model from salient foreground objects, leading to hallucinated phrases like 'clock tower'.
    2. Erroneously detecting 'tower' can cause the model to generate 'building' when describing an image of a train on tracks.

    Because the language generation process directly attends to and conditions on the candidate attribute pool {Ai}\{A_i\}, detector errors propagate directly into lexical selection errors in the generated sentences.

Coverage note — None was omitted; all key contributions including the semantic attention architecture (input/output models), matrix norm regularization, visual attribute prediction pipelines, full quantitative benchmark results, ablation study, and qualitative error modes are covered.

References

  1. 1.J. Ba, V. Mnih, and K. Kavukcuoglu. Multiple object recognition with visual attention. Proceedings of the International Conference on Learning Representations (ICLR), 2015. 3
  2. 2.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. Proceedings of the International Conference on Learning Representations (ICLR), 2014. 2
  3. 3.X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015. 5
  4. 4.X. Chen and C. L. Zitnick. Mind’s eye: A recurrent visual representation for image caption generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2422–2431, 2015. 1, 6
  5. 5.K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. Conference on Empirical Methods in Natural Language Processing (EMNLP 2014), 2014. 2
  6. 6.M. Denil, L. Bazzani, H. Larochelle, and N. de Freitas. Learning where to attend with deep architectures for image tracking. Neural computation, 24(8):2151–2184, 2012. 3
  7. 7.J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick. Exploring nearest neighbor approaches for image captioning. arXiv preprint arXiv:1505.04467, 2015. 5
  8. 8.J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2626–2634, 2015. 1, 2, 6
  9. 9.D. Elliott and F. Keller. Image description using visual dependency representations. In EMNLP, pages 1292–1302, 2013. 1, 2
  10. 10.V. Escorcia, J. C. Niebles, and B. Ghanem. On the relationship between visual attributes and convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1256–1264, 2015. 4
  11. 11.H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al. From captions to visual concepts and back. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1473–1482, 2015. 1, 2, 5
  12. 12.A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth. Every picture tells a story: Generating sentences from images. In Computer Vision–ECCV 2010, pages 15–29. Springer, 2010. 1, 2
  13. 13.Y. Gong, Y. Jia, T. Leung, A. Toshev, and S. Ioffe. Deep convolutional ranking for multilabel image annotation. Proceedings of the International Conference on Learning Representations (ICLR), 2014. 5
  14. 14.Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik. Improving image-sentence embeddings using large weakly annotated photo collections. In Computer Vision–ECCV 2014, pages 529–545. Springer, 2014. 5
  15. 15.K. Gregor, I. Danihelka, A. Graves, and D. Wierstra. Draw: A recurrent neural network for image generation. arXiv preprint arXiv:1502.04623, 2015. 3
  16. 16.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 1, 2, 5
  17. 17.C. Koch and S. Ullman. Shifts in selective visual attention: towards the underlying neural circuitry. In Matters of intelligence, pages 115–141. Springer, 1987. 2
  18. 18.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2, 4
  19. 19.G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg. Baby talk: Understanding and generating image descriptions. In Proceedings of the 24th CVPR. Citeseer, 2011. 1, 2
  20. 20.P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi. Collective generation of natural image descriptions. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume 1, pages 359–368. Association for Computational Linguistics, 2012. 1, 2
  21. 21.H. Larochelle and G. E. Hinton. Learning to combine foveal glimpses with a third-order boltzmann machine. In Advances in neural information processing systems, pages 1243–1251, 2010. 2
  22. 22.R. Lebret, P. O. Pinheiro, and R. Collobert. Simple image description generator via a linear phrase-based approach. Proceedings of the International Conference on Learning Representations (ICLR), 2015. 1, 2
  23. 23.S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi. Composing simple image descriptions using web-scale n-grams. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, pages 220–228. Association for Computational Linguistics, 2011. 1, 2
  24. 24.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 5
  25. 25.J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille. Learning like a child: Fast novel visual concept learning from sentence descriptions of images. In ICCV, 2015. 1, 2, 4
  26. 26.J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). arXiv preprint arXiv:1412.6632, 2014. 1, 2, 6
  27. 27.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119, 2013. 3
  28. 28.V. Mnih, N. Heess, A. Graves, et al. Recurrent models of visual attention. In Advances in Neural Information Processing Systems, pages 2204–2212, 2014. 3
  29. 29.J. Pennington, R. Socher, and C. D. Manning. Glove: Global vectors for word representation. Proceedings of the Empiricial Methods in Natural Language Processing (EMNLP 2014), 12:1532–1543, 2014. 3, 5
  30. 30.M. W. Spratling and M. H. Johnson. A feedback model of visual attention. Journal of cognitive neuroscience, 16(2):219–237, 2004. 2
  31. 31.I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pages 3104–3112, 2014. 2
  32. 32.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 5
  33. 33.Y. Tang, N. Srivastava, and R. R. Salakhutdinov. Learning generative models with visual attention. In Advances in Neural Information Processing Systems, pages 1808–1816, 2014. 2
  34. 34.T. Tieleman and G. Hinton. Lecture 6.5 - rmsprop, coursera: Neural networks for machine learning. 2012. 5
  35. 35.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3156–3164, 2015. 1, 2, 5, 6
  36. 36.K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044, 2015. 1, 2, 3, 6
  37. 37.B. Zhou, V. Jagadeesh, and R. Piramuthu. Conceptlearner: Discovering visual concepts from weakly labeled image collections. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. 5

Citation

MLA
You, Q., et al. “Image Captioning with Semantic Attention”. arXiv, 2016, http://arxiv.org/abs/1603.03925v1.
APA
You, Q., Jin, H., Wang, Z., Fang, C., & Luo, J. (2016). Image Captioning with Semantic Attention. arXiv. http://arxiv.org/abs/1603.03925v1
Chicago
You, Q., H. Jin, Z. Wang, C. Fang, and J. Luo. 2016. “Image Captioning with Semantic Attention”. arXiv. http://arxiv.org/abs/1603.03925v1.
Harvard
You, Q. et al. (2016) “Image Captioning with Semantic Attention”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1603.03925v1.
Vancouver
1. You Q, Jin H, Wang Z, Fang C, Luo J (2016) Image Captioning with Semantic Attention. arXiv

BibTeX

@article{you2016image,
  title = {Image Captioning with Semantic Attention},
  author = {You, Quanzeng and Jin, Hailin and Wang, Zhaowen and Fang, Chen and Luo, Jiebo},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1603.03925v1},
  eprint = {1603.03925}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE