SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning

Long ChenHanwang ZhangJun XiaoLiqiang NieJian ShaoWei LiuTat-Seng Chua

article2016CVPR1,848 citations

Proposes a convolutional neural network that integrates multi-layer spatial and channel-wise attention to capture both what and where to attend, substantially outperforming standard visual attention models on image captioning benchmarks.

Listen

Automated image captioning plays a crucial role in bridging computer vision and natural language processing for real-world applications such as visual search, accessibility tools, and automated media tagging. While visual attention mechanisms have become standard to help systems focus on relevant image areas as sentences are generated, existing approaches rely primarily on basic 2D spatial attention at the final layer of a deep neural network. This conventional method overlooks critical dimensions of feature representation, specifically the semantic attributes captured across individual network channels and the varying levels of abstraction formed across multiple layers.

The article evaluates a novel architecture called Spatial and Channel-wise Attention in Convolutional Networks (SCA-CNN). The objective is to demonstrate that incorporating channel-wise attention (which selects semantic concepts or "what" to look at) alongside spatial attention ("where" to look) across multiple network layers significantly improves image caption generation compared to standard spatial-only approaches.

To evaluate this framework, the authors conducted extensive experiments across three standard benchmark datasets: Flickr8K (8,000 images), Flickr30K (31,000 images), and MSCOCO (over 123,000 images). They implemented the approach using two standard deep learning backbones, VGG-19 and ResNet-152, paired with a Long Short-Term Memory (LSTM) network for sentence decoding. Performance was measured using standard automated language metrics, including BLEU, METEOR, ROUGE-L, and CIDEr, alongside official evaluations on the MSCOCO test server.

The evaluation yielded several key findings. First, integrating spatial and channel-wise attention outperforms standard spatial models, boosting the primary BLEU-4 translation score by approximately 4.8% over foundational spatial baseline models. Second, channel-wise attention delivers substantial gains when applied to networks with large channel capacities; for example, on ResNet-152 (which contains 2,048 channels), channel attention alone significantly outperformed pure spatial attention. Third, applying attention across multiple layers further improves caption quality by capturing both low-level shapes and high-level concepts, with two-layer configurations consistently delivering peak performance across benchmarks. Finally, as a single model, SCA-CNN achieved competitive results against complex ensemble systems on official benchmark leaderboards.

These findings indicate that treating neural network feature representations as three-dimensional—incorporating spatial position, semantic channel selection, and multi-layer depth—produces richer, more accurate descriptions without requiring external attribute classifiers. Implementing this decoupled attention mechanism improves system performance while keeping computational and memory costs manageable for practical deployments.

For practical implementation and future work, technical teams adopting image captioning systems should replace last-layer spatial attention with multi-layer spatial and channel-wise attention pipelines, prioritizing two-layer setups to maximize accuracy without overfitting. The article suggests extending this architecture to temporal dimensions for video captioning and investigating training strategies that allow deeper multi-layer attention on smaller datasets.

Readers should note certain limitations: adding excessive attention layers on smaller datasets increases the risk of overfitting, and model performance depends partly on the underlying network backbone. Nonetheless, the consistent gains across multiple standard benchmarks provide high confidence in the effectiveness of the SCA-CNN framework for image captioning tasks.

Cover for SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning

Abstract

Visual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN encoding an input image. However, we argue that such spatial attention does not necessarily conform to the attention mechanism --- a dynamic feature extractor that combines contextual fixations over time, as CNN features are naturally spatial, channel-wise and multi-layer. In this paper, we introduce a novel convolutional neural network dubbed SCA-CNN that incorporates Spatial and Channel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN dynamically modulates the sentence generation context in multi-layer feature maps, encoding where (i.e., attentive spatial locations at multiple layers) and what (i.e., attentive channels) the visual attention is. We evaluate the proposed SCA-CNN architecture on three benchmark image captioning datasets: Flickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN significantly outperforms state-of-the-art visual attention-based image captioning methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Spatial and Channel-wise Attention CNN
  • 3.1 Overview
  • 3.2 Spatial Attention
  • 3.3 Channel-wise Attention
  • 4 Experiments
  • 4.1 Dataset and Metric
  • 4.2 Setup
  • 4.3 Evaluations of Channel-wise Attention (Q1)
  • 4.4 Evaluations of Multi-layer Attention (Q2)
  • 4.5 Comparison with State-of-The-Arts (Q3)
  • 4.6 Visualization of Spatial and Channel-wise Attention
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — SCA-CNN Multi-Layer Attention Framework for Image Captioning

    model/method

    SCA-CNN modulates convolutional neural network (CNN) multi-layer feature representations dynamically according to the decoding language context in an encoder-decoder image captioning framework.

    At caption generation step tt, given the previous Long Short-Term Memory (LSTM) decoder hidden state ht−1∈Rd\mathbf{h}_{t-1} \in \mathbb{R}^d and the ll-th convolutional layer feature map Vl∈RWl×Hl×Cl\mathbf{V}^l \in \mathbb{R}^{W^l \times H^l \times C^l} (for l=1,…,Ll = 1, \dots, L), SCA-CNN computes spatial and channel-wise attention weights γl∈RWl×Hl×Cl\boldsymbol{\gamma}^l \in \mathbb{R}^{W^l \times H^l \times C^l} and produces an attentive feature map Xl\mathbf{X}^l via recurrent multi-layer modulation:

    Vl=CNN(Xl−1)\mathbf{V}^l = \text{CNN}(\mathbf{X}^{l-1}) γl=Φ(ht−1,Vl)\boldsymbol{\gamma}^l = \Phi(\mathbf{h}_{t-1}, \mathbf{V}^l) Xl=f(Vl,γl)=Vl⊙γl\mathbf{X}^l = f(\mathbf{V}^l, \boldsymbol{\gamma}^l) = \mathbf{V}^l \odot \boldsymbol{\gamma}^l

    where X0\mathbf{X}^0 is the input image, Φ(⋅)\Phi(\cdot) is the attention function, and ⊙\odot represents element-wise multiplication. The final layer modulated feature XL\mathbf{X}^L is fed into the LSTM to decode the tt-th caption word yt\mathbf{y}_t:

    ht=LSTM(ht−1,XL,yt−1)\mathbf{h}_t = \text{LSTM}(\mathbf{h}_{t-1}, \mathbf{X}^L, \mathbf{y}_{t-1}) yt∼pt=softmax(ht,yt−1)\mathbf{y}_t \sim \mathbf{p}_t = \text{softmax}(\mathbf{h}_t, \mathbf{y}_{t-1})

    where pt∈R∣D∣\mathbf{p}_t \in \mathbb{R}^{|\mathcal{D}|} is the output probability distribution over vocabulary D\mathcal{D}. To avoid the O(WlHlClk)O(W^l H^l C^l k) memory requirement of direct 3D joint attention (where kk is the projection dimension), SCA-CNN factorizes γl\boldsymbol{\gamma}^l into separate spatial attention weights αl\boldsymbol{\alpha}^l and channel-wise attention weights βl\boldsymbol{\beta}^l.

  2. Knowl 2 — Spatial Attention Mechanism

    model/method

    The spatial attention function Φs\Phi_s assigns attention weights across the spatial locations of a feature map based on decoding state context.

    Given a feature tensor V∈RW×H×C\mathbf{V} \in \mathbb{R}^{W \times H \times C}, it is reshaped across spatial dimensions to form V=[v1,v2,…,vm]∈RC×m\mathbf{V} = [\mathbf{v}_1, \mathbf{v}_2, \dots, \mathbf{v}_m] \in \mathbb{R}^{C \times m}, where m=W⋅Hm = W \cdot H and each vi∈RC\mathbf{v}_i \in \mathbb{R}^C is the feature vector at spatial location ii. Conditioned on the decoder hidden state ht−1∈Rd\mathbf{h}_{t-1} \in \mathbb{R}^d, spatial attention weights α∈Rm\boldsymbol{\alpha} \in \mathbb{R}^m are computed by:

    a=tanh⁡((WsV+bs)⊕Whsht−1)\mathbf{a} = \tanh((\mathbf{W}_s \mathbf{V} + \mathbf{b}_s) \oplus \mathbf{W}_{hs} \mathbf{h}_{t-1}) α=softmax(Wia+bi)\boldsymbol{\alpha} = \text{softmax}(\mathbf{W}_i \mathbf{a} + b_i)

    where Ws∈Rk×C\mathbf{W}_s \in \mathbb{R}^{k \times C}, Whs∈Rk×d\mathbf{W}_{hs} \in \mathbb{R}^{k \times d}, Wi∈R1×k\mathbf{W}_i \in \mathbb{R}^{1 \times k}, bs∈Rk\mathbf{b}_s \in \mathbb{R}^k, and bi∈Rb_i \in \mathbb{R} are trainable parameters, kk is the common projection dimension, and ⊕\oplus represents column-wise matrix-vector addition. The computational complexity of computing spatial attention is O(WHk)O(W H k).

  3. Knowl 3 — Channel-wise Attention Mechanism

    model/method

    The channel-wise attention function Φc\Phi_c weights individual feature channels, functioning as dynamic semantic attribute selection conditioned on sentence context.

    Given a feature tensor V∈RW×H×C\mathbf{V} \in \mathbb{R}^{W \times H \times C} viewed as CC channel feature slices U=[u1,u2,…,uC]\mathbf{U} = [\mathbf{u}_1, \mathbf{u}_2, \dots, \mathbf{u}_C] with ui∈RW×H\mathbf{u}_i \in \mathbb{R}^{W \times H}, spatial mean pooling is applied to each channel to obtain the channel descriptor vector v=[v1,v2,…,vC]T∈RC\mathbf{v} = [v_1, v_2, \dots, v_C]^T \in \mathbb{R}^C:

    vi=1W⋅H∑j=1W⋅Hui,jv_i = \frac{1}{W \cdot H} \sum_{j=1}^{W \cdot H} u_{i, j}

    Given the decoder hidden state ht−1∈Rd\mathbf{h}_{t-1} \in \mathbb{R}^d, the channel attention distribution β∈RC\boldsymbol{\beta} \in \mathbb{R}^C is computed as:

    b=tanh⁡((Wc⊗v+bc)⊕Whcht−1)\mathbf{b} = \tanh((\mathbf{W}_c \otimes \mathbf{v} + \mathbf{b}_c) \oplus \mathbf{W}_{hc} \mathbf{h}_{t-1}) β=softmax(Wi′b+bi′)\boldsymbol{\beta} = \text{softmax}(\mathbf{W}'_i \mathbf{b} + b'_i)

    where Wc∈Rk\mathbf{W}_c \in \mathbb{R}^k, Whc∈Rk×d\mathbf{W}_{hc} \in \mathbb{R}^{k \times d}, Wi′∈R1×k\mathbf{W}'_i \in \mathbb{R}^{1 \times k}, bc∈Rk\mathbf{b}_c \in \mathbb{R}^k, and bi′∈Rb'_i \in \mathbb{R} are trainable parameters, ⊗\otimes denotes outer product producing a matrix in Rk×C\mathbb{R}^{k \times C}, and ⊕\oplus denotes broadcasting addition. The computational complexity of channel-wise attention is O(Ck)O(C k).

  4. Knowl 4 — Channel-Spatial and Spatial-Channel Attention Orders

    model/method

    Spatial attention Φs\Phi_s and channel-wise attention Φc\Phi_c can be composed sequentially in two configurations within each convolutional layer:

    1. Channel-Spatial (C-S): Channel-wise attention is computed first and applied to modulate the initial feature map V\mathbf{V}, after which spatial attention is computed over the channel-modulated feature map: β=Φc(ht−1,V)\boldsymbol{\beta} = \Phi_c(\mathbf{h}_{t-1}, \mathbf{V}) α=Φs(ht−1,fc(V,β))\boldsymbol{\alpha} = \Phi_s(\mathbf{h}_{t-1}, f_c(\mathbf{V}, \boldsymbol{\beta})) X=f(V,α,β)\mathbf{X} = f(\mathbf{V}, \boldsymbol{\alpha}, \boldsymbol{\beta}) where fc(V,β)f_c(\mathbf{V}, \boldsymbol{\beta}) scales each channel cc of V\mathbf{V} by weight βc\beta_c, and f(V,α,β)f(\mathbf{V}, \boldsymbol{\alpha}, \boldsymbol{\beta}) applies element-wise multiplication along both channel and spatial dimensions.

    2. Spatial-Channel (S-C): Spatial attention is computed first, and the spatially modulated feature map is passed to the channel attention module: α=Φs(ht−1,V)\boldsymbol{\alpha} = \Phi_s(\mathbf{h}_{t-1}, \mathbf{V}) β=Φc(ht−1,fs(V,α))\boldsymbol{\beta} = \Phi_c(\mathbf{h}_{t-1}, f_s(\mathbf{V}, \boldsymbol{\alpha})) X=f(V,α,β)\mathbf{X} = f(\mathbf{V}, \boldsymbol{\alpha}, \boldsymbol{\beta}) where fs(V,α)f_s(\mathbf{V}, \boldsymbol{\alpha}) scales each spatial region ii across all channels by weight αi\alpha_i.

    In empirical evaluations, Channel-Spatial (C-S) consistently achieves slightly superior captioning performance compared to Spatial-Channel (S-C).

  5. Knowl 5 — Layer-by-Layer Warm-Start Training for Multi-Layer Attention

    algorithm

    To train multi-layer SCA-CNN architectures without suffering from optimization instability or long training times, models with KK attentive layers are trained incrementally using the weights of models trained with K−1K-1 attentive layers as initialization.

    Input: Training dataset D={(In,Yn)}n=1ND = \{(I_n, Y_n)\}_{n=1}^N, target number of attentive layers K∈{1,2,3}K \in \{1, 2, 3\}, CNN backbone architecture.
    Output: Trained parameter set ΘK\Theta_K for the KK-layer attentive SCA-CNN.
    Initialize 1-layer SCA-CNN with pretrained CNN backbone weights.
    Train 1-layer SCA-CNN on DD with Adadelta until early stopping to obtain Θ1\Theta_1.
    for k=2k = 2 to KK do
        Initialize the attention parameters of the first k−1k-1 attentive layers in Θk\Theta_k using the trained parameters from Θk−1\Theta_{k-1}.
        Initialize the kk-th attentive layer's parameters randomly.
        Train the kk-layer model Θk\Theta_k end-to-end on DD using Adadelta until convergence or early stopping.
    end for
    return ΘK\Theta_K
  6. Knowl 6 — Experimental Configuration for Image Captioning Evaluation

    experimental setup

    SCA-CNN is evaluated on three benchmark datasets:

    • Flickr8k: 8,000 images (6,000 train, 1,000 validation, 1,000 test).
    • Flickr30k: 31,000 images (29,000 train, 1,000 validation, 1,000 test).
    • MSCOCO: 82,783 training images, 5,000 validation images, 5,000 test images (Karpathy split), and 40,775 test images on the official evaluation server.

    Architectural Details and Hyperparameters:

    • CNN Image Encoders: VGG-19 (attentive layers at conv5_4, conv5_3, conv5_2) and ResNet-152 (attentive layers at res5c, res5c_branch2b, res5c_branch2a).
    • Language Decoder: 1-layer LSTM with hidden dimension d=1000d = 1000 and word embedding dimension 100.
    • Attention dimension: k=512k = 512 for both spatial and channel attention.
    • Optimization: Adadelta with dropout and early stopping.
    • Batch size: 16 on Flickr8k; 64 on Flickr30k and MSCOCO.
    • Inference: Beam search with beam size 5 (without length normalization).
    • Metrics: BLEU (B@1 to B@4), METEOR (MT), ROUGE-L (RG), and CIDEr (CD).
  7. Knowl 7 — Single-Layer Attention Mechanism Comparison across Architectures

    data/table

    Comparison of single-layer attention configurations on Flickr8k, Flickr30k, and MSCOCO using VGG-19 and ResNet-152 backbones demonstrates the role of channel dimension in channel attention efficacy and the superiority of joint channel-spatial modeling.

    Dataset Network Method B@4 MT RG CD
    Flickr8k VGG-19 S 23.0 21.0 49.1 60.6
    SAT 21.3 20.3 – –
    C 22.6 20.3 48.7 58.7
    S-C 22.6 20.9 48.7 60.6
    C-S 23.5 21.1 49.2 60.3
    ResNet-152 S 20.5 19.6 47.4 49.9
    SAT 21.7 20.1 48.4 55.5
    C 24.4 21.5 50.0 65.5
    S-C 24.8 22.2 50.5 65.1
    C-S 25.7 22.1 50.9 66.5
    Flickr30k VGG-19 S 21.1 18.4 43.1 39.5
    SAT 19.9 18.5 – –
    C 20.1 18.0 42.7 38.0
    S-C 20.8 17.8 42.9 38.2
    C-S 21.0 18.0 43.3 38.5
    ResNet-152 S 20.5 17.4 42.8 35.3
    SAT 20.1 17.8 42.9 36.3
    C 21.5 18.4 43.8 42.2
    S-C 21.9 18.5 44.0 43.1
    C-S 22.1 19.0 44.6 42.5
    MSCOCO VGG-19 S 28.2 23.3 51.0 85.7
    SAT 25.0 23.0 – –
    C 27.3 22.7 50.1 83.4
    S-C 28.0 23.0 50.6 84.9
    C-S 28.1 23.5 50.9 84.7
    ResNet-152 S 28.3 23.1 51.2 84.0
    SAT 28.4 23.2 51.2 84.9
    C 29.5 23.7 51.8 91.0
    S-C 29.8 23.9 52.0 91.2
    C-S 30.4 24.5 52.5 91.7

    Pure channel attention (C) yields much larger performance gains on ResNet-152 (2048 channels) than on VGG-19 (512 channels), demonstrating that channel attention benefits from richer filter diversity. Combining channel and spatial attention in C-S format achieves top scores across datasets.

  8. Knowl 8 — Ablation of Attention Layer Depth across S and C-S Models

    data/table

    Performance of pure spatial attention (S) and Channel-Spatial attention (C-S) across 1, 2, and 3 attentive layers on Flickr8k, Flickr30k, and MSCOCO.

    Dataset Network Attentive Layers S B@4 S CD C-S B@4 C-S CD
    Flickr8k VGG-19 1-layer 23.0 60.6 23.5 60.3
    2-layers 22.8 60.4 22.8 62.1
    3-layers 21.6 54.5 22.7 62.3
    ResNet-152 1-layer 20.5 49.9 25.7 66.5
    2-layers 22.9 58.8 25.8 67.1
    3-layers 23.9 61.7 25.3 67.5
    Flickr30k VGG-19 1-layer 21.1 39.5 21.0 38.5
    2-layers 21.9 39.5 21.8 41.4
    3-layers 20.8 38.5 20.7 39.2
    ResNet-152 1-layer 20.5 35.3 22.1 42.5
    2-layers 20.6 39.7 22.3 44.7
    3-layers 21.0 43.5 22.0 42.8
    MSCOCO VGG-19 1-layer 28.2 85.7 28.1 84.7
    2-layers 29.0 87.4 29.8 89.7
    3-layers 27.4 80.8 29.4 88.4
    ResNet-152 1-layer 28.3 84.0 30.4 91.7
    2-layers 29.7 91.1 31.1 95.2
    3-layers 29.6 90.3 30.9 94.7

    Adding a second attentive layer consistently improves metric scores (reaching 31.1 B@4 and 95.2 CIDEr on MSCOCO for 2-layer C-S ResNet-152). Increasing depth to 3 layers leads to overfitting and performance drops on smaller datasets like Flickr8k and Flickr30k.

  9. Knowl 9 — Benchmark Captioning Comparison against State-of-the-Art Methods

    data/table

    Comparison of 2-layer C-S SCA-CNN against state-of-the-art multimodal, spatial attention, and semantic attention models on Flickr8k, Flickr30k, and MSCOCO datasets.

    Flickr8k Flickr30k MSCOCO
    Model B@1 B@2 B@3 B@4 MT B@1 B@2 B@3 B@4 MT B@1 B@2 B@3 B@4 MT
    Deep VS 57.9 38.3 24.5 16.0 – 57.3 36.9 24.0 15.7 – 62.5 45.0 32.1 23.0 19.5
    Google NIC†^\dagger 63.0 41.0 27.0 – – 66.3 42.3 27.7 18.3 – 66.6 46.1 32.9 24.6 –
    m-RNN – – – – – 60.0 41.0 28.0 19.0 – 67.0 49.0 35.0 25.0 –
    Soft-Attention 67.0 44.8 29.9 19.5 18.9 66.7 43.4 28.8 19.1 18.5 70.7 49.2 34.4 24.3 23.9
    Hard-Attention 67.0 45.7 31.4 21.3 20.3 66.9 43.9 29.6 19.9 18.5 71.8 50.4 35.7 25.0 23.0
    emb-gLSTM 64.7 45.9 31.8 21.2 20.6 64.6 44.6 30.5 20.6 17.9 67.0 49.1 35.8 26.4 22.7
    ATT†^\dagger – – – – – 64.7 46.0 32.4 23.0 18.9 70.9 53.7 40.2 30.4 24.3
    SCA-CNN-VGG 65.5 46.6 32.6 22.8 21.6 64.6 45.3 31.7 21.8 18.8 70.5 53.3 39.7 29.8 24.2
    SCA-CNN-ResNet 68.2 49.6 35.9 25.8 22.4 66.2 46.8 32.5 22.3 19.5 71.9 54.8 41.1 31.1 25.0

    (†\dagger denotes ensemble model results; -- indicates metric not reported).

    SCA-CNN-ResNet outperforms the single-model spatial baseline Hard-Attention by 4.5% B@4 on Flickr8k (25.8 vs 21.3), 2.4% B@4 on Flickr30k (22.3 vs 19.9), and 6.1% B@4 on MSCOCO (31.1 vs 25.0). As a single model, it also exceeds the ensemble model ATT on MSCOCO across BLEU4 (31.1 vs 30.4) and METEOR (25.0 vs 24.3).

  10. Knowl 10 — MSCOCO Official Online Test Server Evaluation

    data/table

    Performance of single-model 2-layer C-S SCA-CNN (ResNet-152 backbone) evaluated on the MSCOCO Image Captioning Challenge test server with 5 reference captions (c5) and 40 reference captions (c40).

    B@1 B@2 B@3 B@4 METEOR ROUGE-L CIDEr
    Model c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40
    SCA-CNN 71.2 89.4 54.2 80.2 40.4 69.1 30.2 57.9 24.4 33.1 52.4 67.4 91.2 92.1
    Hard-Attention 70.5 88.1 52.8 77.9 38.3 65.8 27.7 53.7 24.1 32.2 51.6 65.4 86.5 89.3
    ATT†^\dagger 73.1 90.0 56.5 81.5 42.4 70.9 31.6 59.9 25.0 33.5 53.5 68.2 95.3 95.8
    Google NIC†^\dagger 71.3 89.5 54.2 80.2 40.7 69.4 30.9 58.7 25.4 34.6 53.0 68.2 94.3 94.6

    (†\dagger denotes ensemble model results).

    SCA-CNN improves upon single-model Hard-Attention by +2.5 B@4 (c5) and +4.2 B@4 (c40), as well as +4.7 CIDEr (c5) and +2.8 CIDEr (c40), performing competitively with ensemble models.

  11. Knowl 11 — Limitations in Layer Scaling and Temporal Modeling

    limitation

    SCA-CNN has two documented limitations:

    1. Overfitting with Deeper Attention Stacks: Stacking attention over more than 2 convolutional layers increases parameter count and leads to performance degradation on smaller training corpora (such as Flickr8k).
    2. Lack of Temporal Attention: The formulation operates strictly across spatial, channel, and layer dimensions of static 2D images, lacking temporal attention mechanisms needed to handle sequential video frames for video captioning.

Coverage note — None was omitted; all substantive architectural designs, mathematical formulations, training algorithms, ablation studies, benchmark results, and stated limitations are included.

References

  1. 1.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh. Vqa: Visual question answering. In ICCV, 2015. 2
  2. 2.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. In ICLR, 2014. 2
  3. 3.S. Banerjee and A. Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In ACL, 2005. 5
  4. 4.K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia. Abc-cnn: An attention based convolutional neural network for visual question answering. In CVPR, 2016. 1
  5. 5.M. Corbetta and G. L. Shulman. Control of goal-directed and stimulus-driven attention in the brain. Nature reviews neuroscience, 2002. 1
  6. 6.J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell. Long-term recurrent convolutional networks for visual recognition and description. In CVPR, 2015. 2
  7. 7.H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu. Are you talking to a machine? dataset and methods for multilingual image question. In NIPS, 2015. 2
  8. 8.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. 2016. 1, 2, 3, 5
  9. 9.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 1997. 2, 5
  10. 10.M. Hodosh, P. Young, and J. Hockenmaier. Framing image description as a ranking task: Data, models and evaluation metrics. JAIR, 2013. 5
  11. 11.X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars. Guiding the long-short term memory model for image caption generation. In ICCV, 2015. 2, 5, 7
  12. 12.X. Jiang, F. Wu, X. Li, Z. Zhao, W. Lu, S. Tang, and Y. Zhuang. Deep compositional cross-modal learning to rank via local-global alignment. In ACM MM, pages 69–78, 2015. 2
  13. 13.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015. 2, 5, 7
  14. 14.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2016. 2
  15. 15.C.-Y. Lin. Rouge: A package for automatic evaluation of summaries. In ACL, 2004. 5
  16. 16.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 5
  17. 17.M. Malinowski, M. Rohrbach, and M. Fritz. Ask your neurons: A neural-based approach to answering questions about images. In ICCV, 2015. 2
  18. 18.J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille. Deep captioning with multimodal recurrent neural networks (m-rnn). In ICLR, 2015. 7
  19. 19.V. Mnih, N. Heess, A. Graves, et al. Recurrent models of visual attention. In NIPS, 2014. 1
  20. 20.K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002. 5
  21. 21.M. Ren, R. Kiros, and R. Zemel. Exploring models and data for image question answering. In NIPS, 2015. 2
  22. 22.P. H. Seo, Z. Lin, S. Cohen, X. Shen, and B. Han. Hierarchical attention networks. arXiv preprint arXiv:1606.02393, 2016. 2
  23. 23.F. Shen, C. Shen, W. Liu, and H. Tao Shen. Supervised discrete hashing. In CVPR, pages 37–45, 2015. 2
  24. 24.F. Shen, C. Shen, Q. Shi, A. Van Den Hengel, and Z. Tang. Inductive hashing on manifolds. In CVPR, pages 1562–1569, 2013. 2
  25. 25.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 1, 2, 3, 5
  26. 26.M. F. Stollenga, J. Masci, F. Gomez, and J. Schmidhuber. Deep networks with internal selective attention through feedback connections. In NIPS, 2014. 1
  27. 27.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In CVPR, pages 2818–2826, 2016. 7
  28. 28.R. Vedantam, C. Lawrence Zitnick, and D. Parikh. Cider: Consensus-based image description evaluation. In CVPR, 2015. 5
  29. 29.S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko. Sequence to sequence-video to text. In ICCV, 2015. 2
  30. 30.S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko. Translating videos to natural language using deep recurrent neural networks. In NAACL-HLT, 2015. 2
  31. 31.O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: A neural image caption generator. In CVPR, 2015. 2, 5, 7
  32. 32.Y. Wei, W. Xia, M. Lin, J. Huang, B. Ni, J. Dong, Y. Zhao, and S. Yan. Hcp: A flexible cnn framework for multi-label image classification. TPAMI, 2016. 1
  33. 33.H. Xu and K. Saenko. Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In ECCV, 2016. 1, 2
  34. 34.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015. 1, 2, 3, 5, 7
  35. 35.Z. Yang, X. He, J. Gao, L. Deng, and A. Smola. Stacked attention networks for image question answering. In CVPR, 2016. 1, 2
  36. 36.L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville. Describing videos by exploiting temporal structure. In ICCV, 2015. 1
  37. 37.Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo. Image captioning with semantic attention. In CVPR, 2016. 2, 7
  38. 38.P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2014. 5
  39. 39.M. D. Zeiler. Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012. 5
  40. 40.M. D. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. In ECCV, 2014. 1, 2, 8
  41. 41.H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua. Visual translation embedding network for visual relation detection. In CVPR, 2017. 2
  42. 42.Z. Zhao, H. Lu, C. Deng, X. He, and Y. Zhuang. Partial multimodal sparse coding via adaptive similarity structure regularization. In ACM MM, pages 152–156, 2016. 2
  43. 43.Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei. Visual7w: Grounded question answering in images. In CVPR, 2016. 2

Citation

MLA
Chen, L., et al. “SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning”. arXiv, 2016, http://arxiv.org/abs/1611.05594v2.
APA
Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., & Chua, T.-S. (2016). SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning. arXiv. http://arxiv.org/abs/1611.05594v2
Chicago
Chen, L., H. Zhang, J. Xiao, et al. 2016. “SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning”. arXiv. http://arxiv.org/abs/1611.05594v2.
Harvard
Chen, L. et al. (2016) “SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1611.05594v2.
Vancouver
1. Chen L, Zhang H, Xiao J, Nie L, Shao J, Liu W, Chua T-S (2016) SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning. arXiv

BibTeX

@article{chen2016sca,
  title = {SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning},
  author = {Chen, Long and Zhang, Hanwang and Xiao, Jun and Nie, Liqiang and Shao, Jian and Liu, Wei and Chua, Tat-Seng},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1611.05594v2},
  eprint = {1611.05594}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE