Residual Attention Network for Image Classification
Fei WangMengqing JiangChen QianShuo YangCheng LiHonggang ZhangXiaogang WangXiaoou Tang
Proposes attention residual learning to scale attention-aware convolutional neural networks to hundreds of layers, delivering state-of-the-art image classification accuracy while significantly reducing computational cost compared to standard deep residual networks.
The Residual Attention Network integrates an attention mechanism into deep convolutional networks for image classification by stacking multiple Attention Modules, each containing a trunk branch for feature processing and a mask branch that applies soft, adaptive weighting via a bottom-up top-down structure. This design was developed because prior deep feedforward networks such as ResNet improved accuracy through greater depth yet lacked explicit mechanisms to emphasize relevant features while suppressing noise, and earlier attention approaches had not scaled effectively within end-to-end trainable feedforward architectures on large-scale benchmarks.
The work evaluated this architecture through systematic experiments on CIFAR-10, CIFAR-100, and ImageNet, comparing variants with and without attention residual learning, different mask structures, and alternative trunk units such as ResNeXt and Inception. Networks were trained with standard SGD schedules on the full training sets and assessed via single-crop top-1 and top-5 error on held-out validation images, with additional tests introducing controlled label noise.
The principal results show that attention residual learning prevents feature degradation when modules are stacked, enabling networks up to 452 layers deep. Attention-452 reaches 3.90 percent error on CIFAR-10 and 20.45 percent on CIFAR-100, outperforming ResNet-1001 while using fewer parameters. On ImageNet, Attention-92 achieves 19.5 percent top-1 and 4.8 percent top-5 error, improving on ResNet-200 by 0.6 percent top-1 while requiring only 46 percent trunk depth and 69 percent of the forward FLOPs. Mixed attention without spatial or channel constraints performs best, and the networks maintain substantially lower error than ResNet under label noise levels from 10 to 70 percent.
These outcomes indicate that embedding adaptive attention inside residual blocks yields both higher accuracy and greater computational efficiency than simply deepening or widening existing architectures. The noise robustness further suggests practical value in settings where training labels are imperfect. The approach can be applied to other base units without architectural overhaul, confirming its generality.
Further development should test the same modules on detection and segmentation tasks to determine whether the observed gains transfer. Additional gains may come from exploring larger or more diverse datasets and from combining the attention modules with recent normalization or regularization techniques. The main limitations are that all reported ImageNet numbers use single-crop evaluation and that the largest models were trained only on the standard splits; broader validation across multiple runs and additional domains would strengthen confidence in the efficiency claims.
- Paper: Deep Residual Learning for Image Recognition, Kaiming He et al. (2016). Residual Attention Network builds directly upon the foundational deep residual learning framework introduced by He et al.
- Paper: ImageNet Classification with Deep Convolutional Neural Networks, Alex Krizhevsky et al. (2012). Understanding AlexNet's pioneering convolutional architecture provides the baseline image classification context assumed by subsequent network enhancements.
- Paper: Recurrent Models of Visual Attention, Volodymyr Mnih et al. (2014). Recurrent Models of Visual Attention established the conceptual foundations for applying attention mechanisms to visual processing tasks.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). Show, Attend and Tell popularized neural attention mechanisms within deep vision pipelines, directly preceding spatial attention designs in convolutional networks.
- Paper: CBAM: Convolutional Block Attention Module, Sanghyun Woo et al. (2018). CBAM extends the principles of feature recalibration and spatial attention explored in Residual Attention Networks into a modular, lightweight design.
- Paper: ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks, Qilong Wang et al. (2019). ECA-Net continues the line of channel-attention research by proposing an even more efficient mechanism without dimensionality reduction.
- Paper: Dual Attention Network for Scene Segmentation, Jun Fu et al. (2019). Dual Attention Network builds upon spatial and channel attention ideas to tackle dense prediction tasks like scene segmentation.
- Paper: Image Super-Resolution Using Very Deep Residual Channel Attention Networks, Yulun Zhang et al. (2018). RCAN extends attention-guided feature extraction into very deep residual architectures tailored for image super-resolution.
