Attention mechanisms in computer vision: A survey
Meng-Hao GuoTianhan XuJiangjiang LiuZheng-Ning LiuPeng-Tao JiangTai-Jiang MuSong-Hai ZhangRalph Robert MartinMing-Ming ChengShimin Hu
Systematizes visual attention models into channel, spatial, temporal, and branch categories, clarifying how dynamic feature weighting improves performance across classification, segmentation, and multimodal vision tasks.
Modern computer vision systems must process massive volumes of complex visual data, yet standard deep learning architectures often struggle to balance computational efficiency with the ability to capture broader contextual relationships. The human visual system solves this problem by dynamically focusing on the most informative regions while ignoring irrelevant background noise. The article evaluates how visual attention mechanisms imitate this dynamic selection process to enhance deep neural network performance across tasks such as image classification, object detection, and video understanding.
The authors conducted a comprehensive review of roughly a decade of deep learning research, establishing a unified mathematical formulation and categorizing attention methods based on their operational data domains rather than specific downstream applications. The analysis identifies four historical development phases—progressing from early recurrent neural networks to explicit region transformers, implicit feature recalibration, and modern self-attention models. The article organizes existing techniques into six core domains: channel attention (identifying what to focus on), spatial attention (where to focus), temporal attention (when to focus), branch attention (which network pathway to select), and two hybrid categories combining spatial with channel or temporal domains.
The findings demonstrate that attention mechanisms substantially improve representational power, noise suppression, and transformation invariance. While initial self-attention models introduced quadratic computational complexity, subsequent architectural innovations effectively reduced computation to manageable linear or near-linear levels. Furthermore, pure attention-based vision transformers have demonstrated the ability to match or outperform conventional convolutional networks, especially when trained on large-scale datasets.
These results indicate that adopting attention mechanisms can deliver significant gains in computer vision accuracy without necessarily inflating parameter counts or infrastructure costs. Engineering teams should deploy domain-tailored modules, such as lightweight channel recalibration for classification or hybrid spatial-temporal attention for video streams. Future initiatives should focus on developing general-purpose attention blocks, specialized training optimizers, and streamlined deployment methods for edge devices. Stakeholders should note that current attention maps offer intuitive rather than mathematically verifiable explanations, warranting cautious validation in safety-critical settings such as autonomous driving and medical diagnosis.
- Paper: Squeeze-and-Excitation Networks, Jie Hu et al. (2018). Introduces the seminal Squeeze-and-Excitation channel attention mechanism that foundational visual attention surveys build upon and categorize.
- Paper: CBAM: Convolutional Block Attention Module, Sanghyun Woo et al. (2018). Presents CBAM, a foundational architecture combining sequential channel and spatial attention modules that forms a core paradigm reviewed in the survey.
- Paper: Dual Attention Network for Scene Segmentation, J. Fu et al. (2019). Establishes dual spatial and channel self-attention networks for visual feature modeling, serving as a primary representative of combined spatial-channel attention in the survey.
- Paper: ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks, Qilong Wang et al. (2019). Provides the lightweight Efficient Channel Attention (ECA) module directly discussed in taxonomy overviews of efficient channel-based attention.
- Paper: Selective Kernel Networks, Xiang Li et al. (2019). Introduces Selective Kernel Networks (SKNet), establishing the concept of branch-level dynamic attention highlighted in the survey's taxonomy.
- Paper: Residual Attention Network for Image Classification, Fei Wang et al. (2017). Introduces residual attention learning in deep networks, demonstrating how trunk-and-mask branch attention mechanisms stabilize deep visual representations.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). Pioneers soft and hard visual attention mechanisms in deep vision-language modeling, establishing the groundwork for modern spatial attention.
- Paper: CCNet: Criss-Cross Attention for Semantic Segmentation, Zilong Huang et al. (2019). Proposes CCNet to solve computational complexity in full spatial self-attention, representing a key efficient spatial attention innovation.
- Paper: A Model of Saliency-Based Visual Attention for Rapid Scene Analysis, L. Itti et al. (1998). Defines the biologically inspired computational saliency model that motivated the historical adoption of attention mechanisms in computer vision.
- Paper: A Survey on Vision Transformer, Kai Han et al. (2020). Provides a comprehensive architectural survey of self-attention and Vision Transformers that complements and contextualizes the survey's scope.
- Paper: YOLOv12: Attention-Centric Real-Time Object Detectors, Yunjie Tian et al. (2025). Applies modern area-attention mechanisms to real-time object detection backbones to overcome traditional attention latency bottlenecks.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). Extends beyond standard attention formulations by introducing selective state space models to achieve dynamic global receptive fields with linear computational complexity.
- Paper: Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models, Xuyang Liu et al. (2026). Utilizes hierarchical attention scores to dynamically prune and compress high-resolution tokens in modern large vision-language models.
