Residual Attention Network for Image Classification

Fei WangMengqing JiangChen QianShuo YangCheng LiHonggang ZhangXiaogang WangXiaoou Tang

article2017CVPR3,670 citations

Proposes attention residual learning to scale attention-aware convolutional neural networks to hundreds of layers, delivering state-of-the-art image classification accuracy while significantly reducing computational cost compared to standard deep residual networks.

Listen

The Residual Attention Network integrates an attention mechanism into deep convolutional networks for image classification by stacking multiple Attention Modules, each containing a trunk branch for feature processing and a mask branch that applies soft, adaptive weighting via a bottom-up top-down structure. This design was developed because prior deep feedforward networks such as ResNet improved accuracy through greater depth yet lacked explicit mechanisms to emphasize relevant features while suppressing noise, and earlier attention approaches had not scaled effectively within end-to-end trainable feedforward architectures on large-scale benchmarks.

The work evaluated this architecture through systematic experiments on CIFAR-10, CIFAR-100, and ImageNet, comparing variants with and without attention residual learning, different mask structures, and alternative trunk units such as ResNeXt and Inception. Networks were trained with standard SGD schedules on the full training sets and assessed via single-crop top-1 and top-5 error on held-out validation images, with additional tests introducing controlled label noise.

The principal results show that attention residual learning prevents feature degradation when modules are stacked, enabling networks up to 452 layers deep. Attention-452 reaches 3.90 percent error on CIFAR-10 and 20.45 percent on CIFAR-100, outperforming ResNet-1001 while using fewer parameters. On ImageNet, Attention-92 achieves 19.5 percent top-1 and 4.8 percent top-5 error, improving on ResNet-200 by 0.6 percent top-1 while requiring only 46 percent trunk depth and 69 percent of the forward FLOPs. Mixed attention without spatial or channel constraints performs best, and the networks maintain substantially lower error than ResNet under label noise levels from 10 to 70 percent.

These outcomes indicate that embedding adaptive attention inside residual blocks yields both higher accuracy and greater computational efficiency than simply deepening or widening existing architectures. The noise robustness further suggests practical value in settings where training labels are imperfect. The approach can be applied to other base units without architectural overhaul, confirming its generality.

Further development should test the same modules on detection and segmentation tasks to determine whether the observed gains transfer. Additional gains may come from exploring larger or more diverse datasets and from combining the attention modules with recent normalization or regularization techniques. The main limitations are that all reported ImageNet numbers use single-crop evaluation and that the largest models were trained only on the standard splits; broader validation across multiple runs and additional domains would strengthen confidence in the efficiency claims.

arXiv: 1704.06904
Cover for Residual Attention Network for Image Classification

Abstract

In this work, we propose "Residual Attention Network", a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which generate attention-aware features. The attention-aware features from different modules change adaptively as layers going deeper. Inside each Attention Module, bottom-up top-down feedforward structure is used to unfold the feedforward and feedback attention process into a single feedforward process. Importantly, we propose attention residual learning to train very deep Residual Attention Networks which can be easily scaled up to hundreds of layers. Extensive analyses are conducted on CIFAR-10 and CIFAR-100 datasets to verify the effectiveness of every module mentioned above. Our Residual Attention Network achieves state-of-the-art object recognition performance on three benchmark datasets including CIFAR-10 (3.90% error), CIFAR-100 (20.45% error) and ImageNet (4.8% single model and single crop, top-5 error). Note that, our method achieves 0.6% top-1 accuracy improvement with 46% trunk depth and 69% forward FLOPs comparing to ResNet-200. The experiment also demonstrates that our network is robust against noisy labels.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Residual Attention Network
  • 3.1. Attention Residual Learning
  • 3.2. Soft Mask Branch
  • 3.3. Spatial Attention and Channel Attention
  • 4. Experiments
  • 4.1. CIFAR and Analysis
  • 4.2. ImageNet Classification
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Attention Module and Residual Attention Network Architecture

    model/method

    The Residual Attention Network is a convolutional neural network constructed by stacking multiple Attention Modules. Each Attention Module splits an input feature map xx into two parallel processing pathways:

    • Trunk Branch: Computes standard feature representations, denoted F(x)F(x), using deep convolutional blocks such as pre-activation Residual Units, ResNeXt units, or Inception units.
    • Soft Mask Branch: Employs a bottom-up top-down fully convolutional structure to compute an attention mask M(x)M(x) of identical spatial and channel dimensions to F(x)F(x), with output values normalized to [0,1][0, 1].

    Stacking Attention Modules enables multi-stage feature refinement: shallow modules capture low-level attention (such as background color filtering), while deeper modules capture high-level semantic attention (such as object part or instance localization).

  2. Knowl 2 — Attention Residual Learning Formulation

    equation

    In a Residual Attention Network, the output Hi,c(x)H_{i,c}(x) of an Attention Module at spatial position ii and channel index c∈{1,…,C}c \in \{1, \dots, C\} is formulated as:

    Hi,c(x)=(1+Mi,c(x))⋅Fi,c(x)H_{i,c}(x) = (1 + M_{i,c}(x)) \cdot F_{i,c}(x)

    where Fi,c(x)F_{i,c}(x) is the feature tensor computed by the trunk branch, and Mi,c(x)∈[0,1]M_{i,c}(x) \in [0, 1] is the soft mask computed by the mask branch.

    In contrast to naive attention multiplication Hi,c(x)=Mi,c(x)⋅Fi,c(x)H_{i,c}(x) = M_{i,c}(x) \cdot F_{i,c}(x), which causes severe feature signal attenuation and disrupts the identity mapping of residual units when repeated across many layers, attention residual learning enforces an identity mapping baseline: as Mi,c(x)→0M_{i,c}(x) \to 0, Hi,c(x)→Fi,c(x)H_{i,c}(x) \to F_{i,c}(x). This property ensures that adding Attention Modules does not degrade performance relative to the underlying trunk architecture and allows training networks with hundreds of layers.

  3. Knowl 3 — Bottom-Up Top-Down Soft Mask Branch Structure

    model/method

    The soft mask branch in each Attention Module is structured as an encoder-decoder network with skip connections to combine multi-scale spatial context with fine-grained local features:

    • Bottom-Up Encoder: Applies a sequence of Residual Units and stride-2 max pooling layers to rapidly increase receptive field size and extract global image representations until reaching the minimum spatial resolution (7×77 \times 7).
    • Top-Down Decoder: Restores the original spatial dimensions through symmetrical bilinear interpolation upsampling layers interleaved with Residual Units.
    • Skip Connections: Connect features at matching spatial resolutions from the bottom-up path to the top-down path via element-wise addition to preserve localized detail.
    • Mask Normalization: The decoder output passes through two consecutive 1×11 \times 1 convolutional layers followed by a sigmoid activation, yielding mask values in [0,1][0, 1].

    The structure is parameterized by three hyperparameters: pp, the number of pre-processing Residual Units prior to branch splitting; tt, the number of Residual Units in the trunk branch; and rr, the number of Residual Units between adjacent pooling layers in the mask branch. The standard setting on ImageNet is {p=1,t=2,r=1}\{p=1, t=2, r=1\}.

  4. Knowl 4 — Soft Mask Activation Functions and Attention Types

    equation

    The soft mask branch transforms pre-activation feature values xi,cx_{i,c} at spatial position ii and channel index c∈{1,…,C}c \in \{1, \dots, C\} into an attention mask using one of three normalization functions:

    1. Mixed Attention (f1f_1): Applies an unconstrained channel- and position-wise sigmoid:

    f1(xi,c)=11+exp⁡(−xi,c)f_1(x_{i,c}) = \frac{1}{1 + \exp(-x_{i,c})}

    1. Channel Attention (f2f_2): Removes spatial information by applying L2L_2 normalization across all channels at each spatial position ii:

    f2(xi,c)=xi,c∥xi∥2f_2(x_{i,c}) = \frac{x_{i,c}}{\|x_i\|_2}

    where xi=[xi,1,xi,2,…,xi,C]Tx_i = [x_{i,1}, x_{i,2}, \dots, x_{i,C}]^T.

    1. Spatial Attention (f3f_3): Normalizes spatially within each channel cc before applying sigmoid to isolate spatial saliency:

    f3(xi,c)=11+exp⁡(−xi,c−meancstdc)f_3(x_{i,c}) = \frac{1}{1 + \exp\left(-\frac{x_{i,c} - \text{mean}_c}{\text{std}_c}\right)}

    where meanc\text{mean}_c and stdc\text{std}_c are the spatial mean and standard deviation of channel cc.

    When evaluated on CIFAR-10 with Attention-56, mixed attention (f1f_1) achieves a top-1 error of 5.52%5.52\%, outperforming channel attention (f2f_2, 6.24%6.24\%) and spatial attention (f3f_3, 6.33%6.33\%).

  5. Knowl 5 — Gradient Filtering Property and Noisy Label Robustness

    theoretical result

    In an Attention Module with soft mask branch parameters θ\theta and trunk branch parameters ϕ\phi, the forward output feature is M(x,θ)⋅T(x,ϕ)M(x, \theta) \cdot T(x, \phi) (or (1+M(x,θ))⋅T(x,ϕ)(1 + M(x, \theta)) \cdot T(x, \phi) under attention residual learning). During backpropagation, the gradient of the loss with respect to the trunk parameters ϕ\phi is scaled by the mask:

    ∂(M(x,θ)T(x,ϕ))∂ϕ=M(x,θ)∂T(x,ϕ)∂ϕ\frac{\partial (M(x, \theta) T(x, \phi))}{\partial \phi} = M(x, \theta) \frac{\partial T(x, \phi)}{\partial \phi}

    Because M(x,θ)∈[0,1]M(x, \theta) \in [0, 1] directly modulates the incoming error gradient ∂T(x,ϕ)∂ϕ\frac{\partial T(x, \phi)}{\partial \phi}, the mask branch acts as a gradient filter. When training on images with noisy or incorrect labels, regions identified by the mask as background or task-irrelevant receive near-zero mask values, preventing erroneous supervisory gradients from updating the trunk parameters.

  6. Knowl 6 — ImageNet Architecture Specifications for Attention-56 and Attention-92

    model/method

    Residual Attention Networks for ImageNet classification (224×224224 \times 224 input) are structured in four resolution stages using pre-activation Residual Units:

    • Conv1 & Pooling: 7×77 \times 7 convolution (64 filters, stride 2) and 3×33 \times 3 max pooling (stride 2), outputting 56×5656 \times 56.
    • Stage 1 (56×5656 \times 56): 1 Residual Unit [1×1,64;3×3,64;1×1,256][1\times1, 64; 3\times3, 64; 1\times1, 256], followed by Attention Module(s) (×1\times 1 for Attention-56, imes1 imes 1 for Attention-92; mask branch uses 3 max-pooling layers down to 7×77 \times 7).
    • Stage 2 (28×2828 \times 28): 1 Residual Unit [1×1,128;3×3,128;1×1,512][1\times1, 128; 3\times3, 128; 1\times1, 512], followed by Attention Module(s) (×1\times 1 for Attention-56, imes2 imes 2 for Attention-92; mask branch uses 2 max-pooling layers down to 7×77 \times 7).
    • Stage 3 (14×1414 \times 14): 1 Residual Unit [1×1,256;3×3,256;1×1,1024][1\times1, 256; 3\times3, 256; 1\times1, 1024], followed by Attention Module(s) (×1\times 1 for Attention-56, imes3 imes 3 for Attention-92; mask branch uses 1 max-pooling layer down to 7×77 \times 7).
    • Stage 4 (7×77 \times 7): 3 Residual Units [1×1,512;3×3,512;1×1,2048][1\times1, 512; 3\times3, 512; 1\times1, 2048].
    • Head: 7×77 \times 7 global average pooling (stride 1) and 1000-class fully connected softmax.

    Attention-56 contains 31.9 million parameters, 6.2 GFLOPs, and a trunk depth of 56. Attention-92 contains 51.3 million parameters, 10.4 GFLOPs, and a trunk depth of 92.

  7. Knowl 7 — ImageNet Validation Benchmarking and Backbone Adaptability

    data/table

    On the ImageNet LSVRC 2012 validation set (single-crop evaluation), Residual Attention Networks achieve lower classification error rates than baseline architectures while requiring fewer parameters and FLOPs. The Attention Module also integrates across different building blocks, including ResNeXt and Inception.

    Network Parameters (×106\times 10^6) FLOPs (×109\times 10^9) Top-1 Error (%) Top-5 Error (%)
    ResNet-152 60.2 11.3 22.16 6.16
    Attention-56 31.9 6.3 21.76 5.90
    ResNeXt-101 44.5 7.8 21.20 5.60
    AttentionNeXt-56 31.9 6.3 21.20 5.60
    Inception-ResNet-v1 — — 21.30 5.50
    AttentionInception-56 31.9 6.3 20.36 5.29
    ResNet-200 64.7 15.0 20.10 4.80
    Inception-ResNet-v2 — — 19.90 4.90
    Attention-92 51.3 10.4 19.50 4.80

    Attention-56 achieves lower error than ResNet-152 with a 0.40% top-1 improvement while requiring 53.0% of the parameters and 55.8% of the computational cost. Attention-92 reduces top-1 error by 0.60% relative to ResNet-200 while using 20.7% fewer parameters and 30.7% fewer forward FLOPs.

  8. Knowl 8 — CIFAR-10 and CIFAR-100 Classification Performance

    data/table

    Residual Attention Networks evaluated on the CIFAR-10 and CIFAR-100 benchmarks outperform state-of-the-art residual and wide residual networks, showing consistent accuracy gains as attention depth is scaled.

    Network Parameters (×106\times 10^6) CIFAR-10 Error (%) CIFAR-100 Error (%)
    ResNet-164 1.7 5.46 24.33
    ResNet-1001 10.3 4.64 22.71
    WRN-16-8 11.0 4.81 22.07
    WRN-28-10 36.5 4.17 20.50
    Attention-92 1.9 4.99 21.71
    Attention-236 5.1 4.14 21.16
    Attention-452 8.6 3.90 20.45

    Attention-452 (configured with hyper-parameters {p=2,t=4,r=3}\{p=2, t=4, r=3\} and 6 Attention Modules per stage) reaches 3.90% error on CIFAR-10 and 20.45% error on CIFAR-100. Attention-236 outperforms the 1001-layer ResNet-1001 on both datasets while using approximately half the parameter budget (5.1M vs 10.3M).

  9. Knowl 9 — Attention Residual Learning vs. Naive Attention Learning Ablation

    empirical result

    Ablation experiments on CIFAR-10 compare Attention Residual Learning (ARL, H=(1+M)⋅FH = (1+M) \cdot F) against Naive Attention Learning (NAL, H=M⋅FH = M \cdot F) across network depths:

    • Attention-56: ARL achieves 5.52% top-1 error; NAL achieves 5.89% top-1 error.
    • Attention-92: ARL achieves 4.99% top-1 error; NAL achieves 5.35% top-1 error.
    • Attention-128: ARL achieves 4.44% top-1 error; NAL achieves 5.57% top-1 error.
    • Attention-164: ARL achieves 4.31% top-1 error; NAL degrades to 7.18% top-1 error.

    While ARL test error decreases monotonically with increasing depth, NAL performance degrades as layers deepen. This degradation occurs because repeated multiplicative masking in NAL attenuates feature magnitude across successive modules, causing the mean absolute feature response to vanish by stage 2 in Attention-164.

  10. Knowl 10 — Robustness of Residual Attention Networks to Label Noise

    empirical result

    The noise tolerance of Attention-92 was compared against ResNet-164 on CIFAR-10 under varying label noise proportions 1−r∈{10%,30%,50%,70%}1-r \in \{10\%, 30\%, 50\%, 70\%\}, using a confusion transition matrix QQ where true class labels are retained with probability rr and uniformly assigned to any of the other 9 classes with probability 1−r9\frac{1-r}{9}:

    • 10% Noise Level: ResNet-164 error is 5.93%; Attention-92 error is 5.15%.
    • 30% Noise Level: ResNet-164 error is 6.61%; Attention-92 error is 5.79%.
    • 50% Noise Level: ResNet-164 error is 8.35%; Attention-92 error is 7.27%.
    • 70% Noise Level: ResNet-164 error is 17.21%; Attention-92 error is 15.75%.

    Attention-92 achieves lower test error across all noise ratios and degrades more slowly than ResNet-164 as noise increases, validating the gradient filtering effect of the soft mask branch.

  11. Knowl 11 — Soft Mask Architecture: Encoder-Decoder vs. Local Convolutions

    empirical result

    On CIFAR-10, an Attention-56 model using a multi-scale bottom-up top-down encoder-decoder soft mask branch was compared to an equivalent model using local convolutions (three stacked Residual Units without spatial pooling or upsampling) constrained to the same FLOP count:

    • Attention-Local-Conv-56 (Local Convolutions): Top-1 error of 6.48%.
    • Attention-Encoder-Decoder-56 (Encoder-Decoder): Top-1 error of 5.52%.

    The encoder-decoder architecture provides an absolute error reduction of 0.94%, demonstrating that multi-scale receptive field expansion and global context aggregation are essential for generating effective soft attention masks.

Coverage note — None was omitted; all key contributions, mathematical formulations, structural designs, ablation studies, and empirical benchmark results from the paper are represented.

References

  1. 1.V. Badrinarayanan, A. Handa, and R. Cipolla. Segnet: A deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling. arXiv preprint arXiv:1505.07293, 2015. 2, 3
  2. 2.C. Cao, X. Liu, Y. Yang, Y. Yu, J. Wang, Z. Wang, Y. Huang, L. Wang, C. Huang, W. Xu, D. Ramanan, and T. S. Huang. Look and think twice: Capturing top-down visual attention with feedback convolutional neural networks. In ICCV, 2015. 3
  3. 3.L.-C. Chen, Y. Yang, J. Wang, W. Xu, and A. L. Yuille. Attention to scale: Scale-aware semantic image segmentation. arXiv preprint arXiv:1511.03339, 2015. 3, 5
  4. 4.J. Dai, K. He, and J. Sun. Convolutional feature masking for joint object and stuff segmentation. In CVPR, 2015. 2
  5. 5.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR. IEEE. 1, 5, 7
  6. 6.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov. Scalable object detection using deep neural networks. In CVPR, 2014. 2
  7. 7.K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra. Draw: A recurrent neural network for image generation. In ICML, 2015. 2
  8. 8.B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik. Simultaneous detection and segmentation. In ECCV, 2014. 2
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015. 6
  10. 10.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 2, 3, 5, 6, 7, 8
  11. 11.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. arXiv preprint arXiv:1603.05027, 2016. 3, 5, 7, 8
  12. 12.L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell. Deep compositional captioning: Describing novel object categories without paired training data. arXiv preprint arXiv:1511.05284, 2015. 2
  13. 13.S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural computation, 1997. 2
  14. 14.G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger. Deep networks with stochastic depth. arXiv preprint arXiv:1603.09382, 2016. 3
  15. 15.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 3
  16. 16.L. Itti and C. Koch. Computational modelling of visual attention. Nature reviews neuroscience, 2001. 1
  17. 17.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In NIPS, 2015. 3, 5
  18. 18.J.-H. Kim, S.-W. Lee, D. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang. Multimodal residual learning for visual qa. In Advances in Neural Information Processing Systems, pages 361–369, 2016. 2
  19. 19.A. Krizhevsky. Learning multiple layers of features from tiny images. 2009. 5
  20. 20.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 7
  21. 21.H. Larochelle and G. E. Hinton. Learning to combine foveal glimpses with a third-order boltzmann machine. In NIPS, 2010. 2, 4
  22. 22.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 2, 3
  23. 23.V. Mnih, N. Heess, A. Graves, et al. Recurrent models of visual attention. In NIPS, 2014. 1, 2
  24. 24.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. arXiv preprint arXiv:1603.06937, 2016. 2, 3
  25. 25.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. In ICCV, 2015. 2, 3
  26. 26.A. Shrivastava and A. Gupta. Contextual priming and feedback for faster r-cnn. In ECCV, 2016. 2
  27. 27.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. ICLR, 2015. 1, 3
  28. 28.N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. JMLR, 2014. 3
  29. 29.R. K. Srivastava, K. Greff, and J. Schmidhuber. Training very deep networks. In NIPS, 2015. 2, 3
  30. 30.M. F. Stollenga, J. Masci, F. Gomez, and J. Schmidhuber. Deep networks with internal selective attention through feedback connections. In NIPS, 2014. 3
  31. 31.S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus. Training convolutional networks with noisy labels. arXiv preprint arXiv:1406.2080, 2014. 7
  32. 32.C. Szegedy, S. Ioffe, and V. Vanhoucke. Inception-v4, inception-resnet and the impact of residual connections on learning. CoRR, abs/1602.07261, 2016. 3, 8
  33. 33.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015. 1, 3, 7
  34. 34.D. Walther, L. Itti, M. Riesenhuber, T. Poggio, and C. Koch. Attentional selection for object recognitiona gentle way. In International Workshop on Biologically Motivated Computer Vision, pages 472–479. Springer, 2002. 1
  35. 35.T. Xiao, Y. Xu, K. Yang, J. Zhang, Y. Peng, and Z. Zhang. The application of two-level attention models in deep convolutional neural network for fine-grained image classification. In CVPR, 2015. 2
  36. 36.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. arXiv preprint arXiv:1611.05431, 2016. 3, 8
  37. 37.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015. 2
  38. 38.S. Yang, P. Luo, C. C. Loy, and X. Tang. From facial parts responses to face detection: A deep learning approach. In ICCV, 2015. 2
  39. 39.S. Zagoruyko and N. Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016. 7
  40. 40.B. Zhao, X. Wu, J. Feng, Q. Peng, and S. Yan. Diversified visual attention networks for fine-grained object classification. arXiv preprint arXiv:1606.08572, 2016. 1

Citation

MLA
Wang, F., et al. “Residual Attention Network for Image Classification”. arXiv, 2017, http://arxiv.org/abs/1704.06904v1.
APA
Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., & Tang, X. (2017). Residual Attention Network for Image Classification. arXiv. http://arxiv.org/abs/1704.06904v1
Chicago
Wang, F., M. Jiang, C. Qian, et al. 2017. “Residual Attention Network for Image Classification”. arXiv. http://arxiv.org/abs/1704.06904v1.
Harvard
Wang, F. et al. (2017) “Residual Attention Network for Image Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1704.06904v1.
Vancouver
1. Wang F, Jiang M, Qian C, Yang S, Li C, Zhang H, Wang X, Tang X (2017) Residual Attention Network for Image Classification. arXiv

BibTeX

@article{wang2017residual,
  title = {Residual Attention Network for Image Classification},
  author = {Wang, Fei and Jiang, Mengqing and Qian, Chen and Yang, Shuo and Li, Cheng and Zhang, Honggang and Wang, Xiaogang and Tang, Xiaoou},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1704.06904v1},
  eprint = {1704.06904}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE