Attention mechanisms in computer vision: A survey

Meng-Hao GuoTianhan XuJiangjiang LiuZheng-Ning LiuPeng-Tao JiangTai-Jiang MuSong-Hai ZhangRalph Robert MartinMing-Ming ChengShimin Hu

article2021Computational Visual Media2,506 citations

Systematizes visual attention models into channel, spatial, temporal, and branch categories, clarifying how dynamic feature weighting improves performance across classification, segmentation, and multimodal vision tasks.

Listen

Modern computer vision systems must process massive volumes of complex visual data, yet standard deep learning architectures often struggle to balance computational efficiency with the ability to capture broader contextual relationships. The human visual system solves this problem by dynamically focusing on the most informative regions while ignoring irrelevant background noise. The article evaluates how visual attention mechanisms imitate this dynamic selection process to enhance deep neural network performance across tasks such as image classification, object detection, and video understanding.

The authors conducted a comprehensive review of roughly a decade of deep learning research, establishing a unified mathematical formulation and categorizing attention methods based on their operational data domains rather than specific downstream applications. The analysis identifies four historical development phases—progressing from early recurrent neural networks to explicit region transformers, implicit feature recalibration, and modern self-attention models. The article organizes existing techniques into six core domains: channel attention (identifying what to focus on), spatial attention (where to focus), temporal attention (when to focus), branch attention (which network pathway to select), and two hybrid categories combining spatial with channel or temporal domains.

The findings demonstrate that attention mechanisms substantially improve representational power, noise suppression, and transformation invariance. While initial self-attention models introduced quadratic computational complexity, subsequent architectural innovations effectively reduced computation to manageable linear or near-linear levels. Furthermore, pure attention-based vision transformers have demonstrated the ability to match or outperform conventional convolutional networks, especially when trained on large-scale datasets.

These results indicate that adopting attention mechanisms can deliver significant gains in computer vision accuracy without necessarily inflating parameter counts or infrastructure costs. Engineering teams should deploy domain-tailored modules, such as lightweight channel recalibration for classification or hybrid spatial-temporal attention for video streams. Future initiatives should focus on developing general-purpose attention blocks, specialized training optimizers, and streamlined deployment methods for edge devices. Stakeholders should note that current attention maps offer intuitive rather than mathematically verifiable explanations, warranting cautious validation in safety-critical settings such as autonomous driving and medical diagnosis.

Cover for Attention mechanisms in computer vision: A survey

Abstract

Humans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multi-modal tasks and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention and branch attention; a related repository this https URL is dedicated to collecting related work. We also suggest future directions for attention mechanism research.

Table of Contents

  • I Introduction
  • II Other surveys
  • III Attention methods in computer vision
  • III-A General form
  • III-B Channel Attention
  • III-B1 SENet
  • III-B2 GSoP-Net
  • III-B3 SRM
  • III-B4 GCT
  • III-B5 ECANet
  • III-B6 FcaNet
  • III-B7 EncNet
  • III-B8 Bilinear Attention
  • III-C Spatial Attention
  • III-C1 RAM
  • III-C2 Glimpse Network
  • III-C3 Hard and soft attention
  • III-C4 Attention Gate
  • III-C5 STN
  • III-C6 Deformable Convolutional Networks
  • III-C7 Self-attention and variants
  • III-C8 Vision Transformers
  • III-C9 GENet
  • III-C10 PSANet
  • III-D Temporal Attention
  • III-D1 Self-attention and variants
  • III-D2 TAM
  • III-E Branch Attention
  • III-E1 Highway networks
  • III-E2 SKNet
  • III-E3 CondConv
  • III-E4 Dynamic Convolution
  • III-F Channel & Spatial Attention
  • III-F1 Residual Attention Network
  • III-F2 CBAM
  • III-F3 BAM
  • III-F4 scSE
  • III-F5 Triplet Attention
  • III-F6 SimAM
  • III-F7 Coordinate attention
  • III-F8 DANet
  • III-F9 RGA
  • III-F10 Self-Calibrated Convolutions
  • III-F11 SPNet
  • III-F12 SCA-CNN
  • III-F13 GALA
  • III-G Spatial & Temporal Attention
  • III-G1 STA-LSTM
  • III-G2 RSTAN
  • III-G3 STA
  • III-G4 STGCN
  • IV Future directions
  • IV-A Necessary and sufficient condition for attention
  • IV-B General attention block
  • IV-C Characterisation and interpretability
  • IV-D Sparse activation
  • IV-E Attention-based pre-trained models
  • IV-F Optimization
  • IV-G Deployment
  • V Conclusions
  • References

Knowls

  1. Knowl 1 — Unified General Formulation of Visual Attention Mechanisms

    definition

    An attention mechanism in computer vision can be viewed as an adaptive selection process that dynamically weights features based on the saliency of the input. Most visual attention models can be represented under a unified general formulation:

    Attention=f(g(x),x)\text{Attention} = f(g(x), x)

    where:

    • xx denotes the input feature representation (such as an intermediate feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, where CC is the channel dimension and H,WH, W are spatial height and width).
    • g(x)g(x) represents the attention generation function, which computes attention weights or gating coefficients that identify discriminative regions, channels, or temporal slices.
    • f(g(x),x)f(g(x), x) represents the feature processing or modulation function that scales, filters, or aggregates the input features xx conditioned on the generated attention g(x)g(x).

    For instance:

    • In scaled dot-product self-attention, projections Q,K,V=Linear(x)Q, K, V = \text{Linear}(x) are computed, where g(x)=Softmax(QKT)g(x) = \text{Softmax}(QK^T) and f(g(x),x)=g(x)Vf(g(x), x) = g(x)V.
    • In Squeeze-and-Excitation (SE) channel attention, g(x)=σ(MLP(GAP(x)))g(x) = \sigma(\text{MLP}(\text{GAP}(x))) and f(g(x),x)=g(x)⊙xf(g(x), x) = g(x) \odot x, where GAP\text{GAP} denotes global average pooling, MLP\text{MLP} is a multi-layer perceptron, σ\sigma is the sigmoid activation function, and ⊙\odot denotes channel-wise multiplication.
  2. Knowl 2 — Data-Domain Taxonomy of Computer Vision Attention Mechanisms

    definition

    Visual attention mechanisms can be systematically categorized based on the data domain across which attention masks and selection operations operate:

    1. Channel Attention ("What to pay attention to"): Generates attention masks across feature channels (CC) to adaptively recalibrate the importance of distinct channel feature maps, treating different channels as distinct semantic pattern detectors.
    2. Spatial Attention ("Where to pay attention"): Generates attention masks across spatial coordinates (H×WH \times W) or predicts spatial transformations and sampling offsets directly to emphasize task-relevant spatial regions.
    3. Temporal Attention ("When to pay attention"): Generates dynamic weights across the temporal dimension (TT) in video sequences to highlight salient keyframes and suppress temporal redundancy.
    4. Branch Attention ("Which branch to pay attention to"): Operates across multiple network pathways, layer connections, or convolution kernels to dynamically select or fuse branches in multi-branch architectures.
    5. Channel & Spatial Attention: Combines channel and spatial selection either sequentially, in parallel, or via joint 3D tensors to determine simultaneously which features and where in the image plane to focus on.
    6. Spatial & Temporal Attention: Jointly or sequentially computes spatial and temporal attention weights to select informative regions across frames in video understanding tasks.
  3. Knowl 3 — Four Developmental Phases of Visual Attention in Deep Learning

    definition

    The evolution of deep-learning-based visual attention mechanisms spans four historical phases:

    1. Phase 1: RNN-based Recurrent Attention (2014–2015): Pioneered by models such as the Recurrent Attention Model (RAM), these approaches combine recurrent neural networks (RNNs) with reinforcement learning policy gradients to sequentially choose glimpse locations over an image in non-differentiable or semi-differentiable pipelines.
    2. Phase 2: Explicit Region Prediction and Geometric Transformation (2015–2017): Introduced by Spatial Transformer Networks (STN) and Deformable Convolutional Networks (DCN), these methods use differentiable sub-networks to explicitly predict spatial transformations (such as affine transformation matrices or continuous sampling offsets) to achieve geometric transformation invariance.
    3. Phase 3: Implicit Channel and Feature Recalibration (2017–2019): Initiated by Squeeze-and-Excitation Networks (SENet) and extended by CBAM and ECANet, this phase utilizes lightweight sub-networks to implicitly compute soft attention weights over channels or spatial maps to recalibrate features.
    4. Phase 4: Self-Attention and Visual Transformers (2018–Present): Began with Non-Local Neural Networks adapting natural language processing self-attention to capture long-range visual dependencies, culminating in pure transformer architectures such as the Vision Transformer (ViT) that replace standard convolutions with multi-head self-attention.
  4. Knowl 4 — Channel Attention Mechanisms and Excitation Variants

    model/method

    Channel attention recalibrates the weight of each channel in an intermediate feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}. In standard Squeeze-and-Excitation (SENet), spatial dimensions are squeezed via global average pooling (GAP\text{GAP}) and excited via a two-layer multi-layer perceptron (MLP) with reduction ratio rr:

    s=σ(W2δ(W1GAP(X))),Y=s⊙Xs = \sigma(W_2 \delta(W_1 \text{GAP}(X))), \quad Y = s \odot X

    where W1∈RCr×CW_1 \in \mathbb{R}^{\frac{C}{r} \times C}, W2∈RC×CrW_2 \in \mathbb{R}^{C \times \frac{C}{r}}, δ\delta is ReLU, σ\sigma is the sigmoid function, and ⊙\odot denotes channel-wise multiplication.

    Subsequent channel attention methods improve either the squeeze module, the excitation module, or both:

    • Enhanced Squeeze Modules:
      • GSoP-Net: Replaces first-order pooling with global second-order pooling (Cov(Conv(X))\text{Cov}(\text{Conv}(X))) to capture pairwise feature covariance.
      • FcaNet: Analyzes GAP in the frequency domain as the lowest frequency component of the 2D Discrete Cosine Transform (DCT) and uses multi-spectral DCT components across grouped channels: s=σ(W2δ(W1[DCT(Group(X))]))s = \sigma(W_2 \delta(W_1 [\text{DCT}(\text{Group}(X))])).
      • EncNet (CEM): Computes soft-assignment residual descriptors against KK learnable cluster centers D={d1,…,dK}D = \{d_1, \dots, d_K\}.
    • Enhanced Excitation Modules:
      • ECANet: Avoids dimensionality reduction by replacing fully-connected layers with a local 1D convolution of adaptive kernel size k=ψ(C)=∣log⁡2(C)γ+bγ∣oddk = \psi(C) = \left| \frac{\log_2(C)}{\gamma} + \frac{b}{\gamma} \right|_{\text{odd}}, yielding s=σ(Conv1Dk(GAP(X)))s = \sigma(\text{Conv1D}_k(\text{GAP}(X))).
    • Joint Squeeze and Excitation Enhancements:
      • SRM (Style-based Recalibration Module): Uses style pooling combining mean and standard deviation, followed by channel-wise fully connected (CFC) layers and batch normalization: s=σ(BN(CFC(SP(X))))s = \sigma(\text{BN}(\text{CFC}(\text{SP}(X)))).
      • GCT (Gated Channel Transformation): Computes spatial ℓ2\ell_2-norms per channel, applies channel normalization CN\text{CN}, a tanh⁡\tanh gating function, and adds an identity connection: Y=tanh⁡(γCN(αNorm(X))+β)⊙X+XY = \tanh(\gamma \text{CN}(\alpha \text{Norm}(X)) + \beta) \odot X + X.
  5. Knowl 5 — Spatial Attention and Self-Attention Formulations in Vision

    model/method

    Spatial attention selects relevant spatial locations within a feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}. Major formulations include:

    1. Implicit Soft Mask Attention:

      • GENet: Uses spatial pooling or depthwise convolutions gather(X)\text{gather}(X) followed by spatial interpolation and sigmoid activation: s=σ(Interp(fgather(X)))s = \sigma(\text{Interp}(f_{\text{gather}}(X))) and Y=s⊙XY = s \odot X.
      • PSANet: Decomposes point-to-point aggregation into bi-directional position-sensitive terms (collect and distribute): zi=∑j∈Ω(i)FΔij(xi)xj+∑j∈Ω(i)FΔij(xj)xjz_i = \sum_{j \in \Omega(i)} F_{\Delta_{ij}}(x_i) x_j + \sum_{j \in \Omega(i)} F_{\Delta_{ij}}(x_j) x_j.
    2. Explicit Spatial Transformation:

      • Spatial Transformer Networks (STN): Computes an affine transformation matrix θ=floc(U)\theta = f_{\text{loc}}(U) to map output grid coordinates (xis,yis)(x_i^s, y_i^s) to input coordinates (xit,yit)(x_i^t, y_i^t) using bilinear sampling.
      • Deformable Convolutional Networks (DCN): Predicts 2D sampling offsets Δpi\Delta p_i per location: Y(p0)=∑pi∈Rw(pi)X(p0+pi+Δpi)Y(p_0) = \sum_{p_i \in \mathcal{R}} w(p_i) X(p_0 + p_i + \Delta p_i).
    3. Self-Attention and Vision Transformers:

      • Given feature map F∈RC×H×WF \in \mathbb{R}^{C \times H \times W}, queries QQ, keys KK, and values VV in RC×N\mathbb{R}^{C \times N} (N=HWN = HW) are linearly projected. The full self-attention operation is:

      A=Softmax(QKT)∈RN×N,Y=AVA = \text{Softmax}(Q K^T) \in \mathbb{R}^{N \times N}, \quad Y = A V

      • Local Self-Attention (e.g., SASA) restricts aggregation to a local neighborhood Nk(i,j)\mathcal{N}_k(i, j) with relative positional embeddings ra−i,b−jr_{a-i, b-j}:

      Yi,j=∑a,b∈Nk(i,j)Softmaxa,b(qi,jTka,b+qi,jTra−i,b−j)va,bY_{i, j} = \sum_{a, b \in \mathcal{N}_k(i, j)} \text{Softmax}_{a, b}(q_{i, j}^T k_{a, b} + q_{i, j}^T r_{a-i, b-j}) v_{a, b}

      • Vision Transformer (ViT) flattens non-overlapping image patches, prepends a learnable class token, adds positional embeddings, and applies stacked Multi-Head Attention (MHA\text{MHA}) blocks with Layer Normalization and MLP layers.
  6. Knowl 6 — Branch Attention Mechanisms for Multi-Path and Dynamic Convolutions

    model/method

    Branch attention dynamically selects or weights computation across different architectural pathways, receptive field scales, or convolution kernels:

    1. Highway Networks: Uses an input-dependent gating function Tl(X)=σ(WlTX+bl)T_l(X) = \sigma(W_l^T X + b_l) to adaptively route information between non-linear transformations Hl(Xl)H_l(X_l) and identity skip-connections:

      Yl=Hl(Xl)Tl(Xl)+Xl(1−Tl(Xl))Y_l = H_l(X_l) T_l(X_l) + X_l (1 - T_l(X_l))

    2. Selective Kernel Networks (SKNet): Dynamically adjusts neuron receptive field sizes across KK branches with different kernel sizes (Uk=Fk(X)U_k = F_k(X)). Features are fused via summation U=∑k=1KUkU = \sum_{k=1}^K U_k, squeezed into vector z=δ(BN(WGAP(U)))z = \delta(\text{BN}(W \text{GAP}(U))), and branch selection softmax weights sk(c)s_k^{(c)} are computed:

      sk(c)=eWk(c)z∑j=1KeWj(c)z,Y=∑k=1Ksk⊙Uks_k^{(c)} = \frac{e^{W_k^{(c)} z}}{\sum_{j=1}^K e^{W_j^{(c)} z}}, \quad Y = \sum_{k=1}^K s_k \odot U_k

    3. Dynamic / Conditionally Parameterized Convolutions (CondConv, Dynamic Conv): Rather than routing feature maps, branch attention dynamically aggregates KK parallel convolution kernels {W1,…,WK}\{W_1, \dots, W_K\} before performing convolution:

      α=Softmax(W2δ(W1GAP(X))),Wdyn=∑k=1KαkWk,Y=Wdyn∗X\alpha = \text{Softmax}(W_2 \delta(W_1 \text{GAP}(X))), \quad W_{\text{dyn}} = \sum_{k=1}^K \alpha_k W_k, \quad Y = W_{\text{dyn}} * X

      This enables an ensemble-like capacity increase with negligible computational overhead during inference compared to multi-branch convolution execution.

  7. Knowl 7 — Hybrid Channel and Spatial Attention Architectures

    model/method

    Hybrid channel and spatial attention mechanisms combine channel selection ("what") and spatial localization ("where"):

    1. Sequential Channel and Spatial Attention (CBAM):

      • Computes channel attention using dual pooling (average and max pooling): sc=σ(W2δ(W1GAPs(X))+W2δ(W1GMPs(X)))s_c = \sigma(W_2 \delta(W_1 \text{GAP}_s(X)) + W_2 \delta(W_1 \text{GMP}_s(X)))
      • Follows channel modulation X′=sc⊙XX' = s_c \odot X with a spatial attention sub-module using channel-wise pooling and large-kernel convolution: ss=σ(Conv([GAPc(X′);GMPc(X′)])),Y=ss⊙X′s_s = \sigma(\text{Conv}([\text{GAP}_c(X'); \text{GMP}_c(X')])), \quad Y = s_s \odot X'
    2. Parallel Dual Attention (BAM and scSE):

      • BAM: Computes channel branch scs_c and spatial branch sss_s (via dilated convolutions) independently and combines them: Y=(1+σ(Expand(sc)+Expand(ss)))⊙XY = (1 + \sigma(\text{Expand}(s_c) + \text{Expand}(s_s))) \odot X.
      • scSE: Evaluates spatial squeeze ss=σ(Conv1×1(X))s_s = \sigma(\text{Conv}^{1\times 1}(X)) and channel squeeze sc=σ(MLP(GAP(X)))s_c = \sigma(\text{MLP}(\text{GAP}(X))) in parallel, fusing outputs via addition or multiplication.
    3. Cross-Domain Interaction (Triplet Attention):

      • Avoids independent processing of dimensions by modeling three two-domain interactions ((C,H)(C, H), (C,W)(C, W), (H,W)(H, W)) using rotation transformations Pm1,Pm2P_{m1}, P_{m2}, along with ZZ-pooling ([GMP(⋅);GAP(⋅)][\text{GMP}(\cdot); \text{GAP}(\cdot)]): Y=13(s0⊙X+Pm1−1(s1⊙Pm1(X))+Pm2−1(s2⊙Pm2(X)))Y = \frac{1}{3} \left( s_0 \odot X + P_{m1}^{-1}(s_1 \odot P_{m1}(X)) + P_{m2}^{-1}(s_2 \odot P_{m2}(X)) \right)
    4. Coordinate Attention:

      • Factorizes 2D spatial pooling into 1D horizontal (GAPh\text{GAP}^h) and vertical (GAPw\text{GAP}^w) pooling to preserve spatial positional coordinate structures along orthogonal axes: zh=GAPh(X),zw=GAPw(X),f=δ(BN(Conv1×1([zh;zw])))z^h = \text{GAP}^h(X), \quad z^w = \text{GAP}^w(X), \quad f = \delta(\text{BN}(\text{Conv}^{1\times 1}([z^h; z^w]))) sh=σ(Convh1×1(fh)),sw=σ(Convw1×1(fw)),Y=X⊙sh⊙sws^h = \sigma(\text{Conv}_h^{1\times 1}(f^h)), \quad s^w = \sigma(\text{Conv}_w^{1\times 1}(f^w)), \quad Y = X \odot s^h \odot s^w
    5. Self-Attention Based Dual Attention (DANet):

      • Computes a spatial position attention matrix Apos=Softmax(QKT)A_{\text{pos}} = \text{Softmax}(Q K^T) and a channel attention matrix Achn=Softmax(XXT)A_{\text{chn}} = \text{Softmax}(X X^T) in parallel, summing their aggregated outputs.
    6. Parameter-Free 3D Energy Attention (SimAM):

      • Derives analytical neuron importance based on visual neuroscience surround suppression: et∗=4(σ^2+λ)(t−μ^)2+2σ^2+2λ,Y=Sigmoid(1E)⊙Xe_t^* = \frac{4(\hat{\sigma}^2 + \lambda)}{(t - \hat{\mu})^2 + 2\hat{\sigma}^2 + 2\lambda}, \quad Y = \text{Sigmoid}\left(\frac{1}{E}\right) \odot X where μ^\hat{\mu} and σ^2\hat{\sigma}^2 are channel-wise mean and variance.
  8. Knowl 8 — Spatiotemporal Attention Formulations for Video Processing

    model/method

    Spatiotemporal attention mechanisms jointly model spatial regions and temporal frame dependencies in video representation learning:

    1. Separated Spatial and Temporal Attention:

      • STA-LSTM: Uses LSTM hidden state ht−1sh_{t-1}^s to guide spatial attention over joints or regions (st=Ustanh⁡(WxsXt+Whsht−1s+bsi)+bsos_t = U_s \tanh(W_{xs} X_t + W_{hs} h_{t-1}^s + b_{si}) + b_{so}), and computes temporal gating βt=δ(WxpXt+Whpht−1p+bp)\beta_t = \delta(W_{xp} X_t + W_{hp} h_{t-1}^p + b_p) to select keyframes.
      • RSTAN: Uses recurrent hidden states to first compute a spatial location-score map αt∗(n,k)\alpha_t^*(n, k) for each frame nn at step tt, aggregates frame features ln=∑kαt∗(n,k)X(n,k)l_n = \sum_{k} \alpha_t^*(n, k) X(n, k), and applies temporal attention βt∗(n)\beta_t^*(n) across {l1,…,lT}\{l_1, \dots, l_T\} to produce output ϕt=∑n=1Tβt∗(n)ln\phi_t = \sum_{n=1}^T \beta_t^*(n) l_n.
    2. Non-Parametric Joint Spatiotemporal Attention (STA):

      • Computes frame-level energy maps using channel ℓ2\ell_2-norms: gn(h,w)=∥∑c=1CXn(c,h,w)2∥2∑h=1H∑w=1W∥∑c=1CXn(c,h,w)2∥2g_n(h, w) = \frac{\left\| \sum_{c=1}^C X_n(c, h, w)^2 \right\|_2}{\sum_{h=1}^H \sum_{w=1}^W \left\| \sum_{c=1}^C X_n(c, h, w)^2 \right\|_2}
      • Partitions frames into KK spatial parts, computes part attention scores sn,ks_{n, k}, and normalizes them across all video frames using ℓ1\ell_1-normalization: S(n,k)=sn,k∑n=1N∥sn,k∥1S(n, k) = \frac{s_{n, k}}{\sum_{n=1}^N \|s_{n, k}\|_1}. Outputs combine max-attended and weighted-average part descriptors.
    3. Spatiotemporal Graph Convolution (STGCN):

      • Partitions each frame into PP patches (N=TPN = TP total nodes) and defines an affinity matrix E(i,j)=(Wϕxi)TWϕxjE(i, j) = (W_\phi x_i)^T W_\phi x_j. A normalized adjacency matrix A^=D−12(A+I)D−12\hat{A} = D^{-\frac{1}{2}} (A + I) D^{-\frac{1}{2}} models long-range cross-frame and intra-frame patch interactions through graph convolution: Xm=A^Xm−1Wm+Xm−1X^m = \hat{A} X^{m-1} W^m + X^{m-1}
  9. Knowl 9 — Analysis of Necessary versus Sufficient Conditions for Visual Attention

    theoretical result

    The algebraic relation:

    Attention=f(g(x),x)\text{Attention} = f(g(x), x)

    serves as a necessary condition for visual attention mechanisms, but it is not a sufficient condition.

    Counterexample and Evidence: Standard multi-branch or gating architectures (such as GoogLeNet / Inception modules) satisfy the functional form f(g(x),x)f(g(x), x) through linear projections, concatenation, and branch combinations. However, they do not constitute attention mechanisms because they do not dynamically select or weight features according to input saliency. Thus, defining a mathematically rigorous necessary and sufficient condition that completely demarcates attention mechanisms from general feedforward transformations remains an open theoretical problem.

  10. Knowl 10 — Research Challenges and Future Directions in Visual Attention

    limitation

    Several fundamental challenges and future research directions characterize visual attention models in computer vision:

    1. Lack of a General Attention Block: Existing attention designs are specialized for particular tasks (e.g., channel attention for classification, spatial self-attention for dense prediction). Designing a task-agnostic, general attention module capable of dynamically selecting between channel, spatial, temporal, and branch attention remains an open challenge.
    2. Interpretability Beyond Heatmap Visualization: Saliency heatmaps provide only intuitive visualizations rather than formal interpretability or failure-mode guarantees, which limits safety-critical deployment in medical imaging and autonomous driving.
    3. Sparse Activation and Human Vision Simulation: Empirical visualizations demonstrate that attention mechanisms naturally induce sparse activations, mirroring human visual cognition. Developing architectures that explicitly exploit structured sparsity could improve both computational efficiency and cognitive fidelity.
    4. Dedicated Optimizers for Visual Transformers: Standard CNN optimizers (e.g., SGD, standard Adam) often underperform on transformer-based visual models compared to AdamW or Sharpness-Aware Minimization (SAM), indicating a need for optimization strategies tailored specifically to attention landscapes.
    5. Edge Deployment and Hardware Acceleration: While pure attention models often achieve higher accuracy than CNNs, their irregular memory access patterns and variable matrix multiplications hinder efficient optimization and deployment on edge devices relative to uniform convolution operations.

Coverage note — None was omitted; the knowls cover the unified mathematical definition, data-domain taxonomy, historical phases, mathematical formulations for all six domain categories, theoretical necessary-vs-sufficient analysis, and future research directions.

References

  1. 1.Itti, L.; Koch, C.; Niebur, E. A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence Vol. 20, No. 11, 1254–1259, 1998.
  2. 2.Hayhoe, M.; Ballard, D. Eye movements in natural behavior. Trends in Cognitive Sciences Vol. 9, No. 4, 188–194, 2005
  3. 3.Rensink, R. A. The dynamic representation of scenes. Visual Cognition Vol. 7, Nos. 1–3, 17–42, 2000.
  4. 4.Corbetta, M.; Shulman, G. L. Control of goal-directed and stimulus-driven attention in the brain. Nature Reviews Neuroscience Vol. 3, No. 3, 201–215, 2002.
  5. 5.Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Wu, E. H. Squeeze-and-excitation networks. IEEE Transactions on Pattern Analysis and Machine Intelligence Vol. 42, No. 8, 2011–2023, 2020.
  6. 6.Woo, S.; Park, J.; Lee, J.; Kweon, I. S. CBAM: Convolutional block attention module. In: Computer Vision – ECCV 2018. Lecture Notes in Computer Science, Vol. 11211. Ferrari, V.; Hebert, M.; Sminchisescu, C.; Weiss, Y. Eds. Springer Cham, 3–19, 2018.
  7. 7.Dai, J. F.; Qi, H. Z.; Xiong, Y. W.; Li, Y.; Zhang, G. D.; Hu, H.; Wei, Y. Deformable convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision, 764–773, 2017.
  8. 8.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In: Computer Vision – ECCV 2020. Lecture Notes in Computer Science, Vol. 12346. Vedaldi, A.; Bischof, H.; Brox, T.; Frahm, J. M. Eds. Springer Cham, 213–229, 2020.
  9. 9.Yuan, Y.; Wang, J. OCNet: Object context network for scene parsing. arXiv preprint arXiv:1809.00916, 2018.
  10. 10.Fu, J.; Liu, J.; Tian, H. J.; Li, Y.; Bao, Y. J.; Fang, Z. W.; Lu, H. Dual attention network for scene segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3141–3149, 2019.
  11. 11.Yang, J. L.; Ren, P. R.; Zhang, D. Q.; Chen, D.; Wen, F.; Li, H. D.; Hua, G. Neural aggregation network for video face recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5216–5225, 2017.
  12. 12.Wang, Q. C.; Wu, T. Y.; Zheng, H.; Guo, G. D. Hierarchical pyramid diverse attention networks for face recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8323–8332, 2020.
  13. 13.Li, W.; Zhu, X. T.; Gong, S. G. Harmonious attention network for person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2285–2294, 2018.
  14. 14.Chen, B. H.; Deng, W. H.; Hu, J. N. Mixed high-order attention network for person re-identification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 371–381, 2019.
  15. 15.Wang, X. L.; Girshick, R.; Gupta, A.; He, K. M. Non-local neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7794–7803, 2018.
  16. 16.Du, W. B.; Wang, Y. L.; Qiao, Y. Recurrent spatial-temporal attention network for action recognition in videos. IEEE Transactions on Image Processing Vol. 27, No. 3, 1347–1360, 2018.
  17. 17.Peng, Y. X.; He, X. T.; Zhao, J. J. Object-part attention model for fine-grained image classification. IEEE Transactions on Image Processing Vol. 27, No. 3, 1487–1500, 2018.
  18. 18.He, P.; Huang, W. L.; He, T.; Zhu, Q. L.; Qiao, Y.; Li, X. L. Single shot text detector with regional attention. In: Proceedings of the IEEE International Conference on Computer Vision, 3066–3074, 2017.
  19. 19.Oktay, O.; Schlemper, J.; Folgoc, L. L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N. Y.; Kainz, B.; et al. Attention U-Net: Learning where to look for the pancreas. arXiv preprint arXiv:1804.03999, 2018.
  20. 20.Guan, Q.; Huang, Y.; Zhong, Z.; Zheng, Z.; Zheng, L.; Yang, Y. Diagnose like a radiologist: Attention guided convolutional neural network for thorax disease classification. arXiv preprint arXiv:1801.09927, 2018.
  21. 21.Gregor, K.; Danihelka, I.; Graves, A.; Wierstra, D. DRAW: A recurrent neural network for image generation. In: Proceedings of the 32nd International Conference on Machine Learning, 1462–1471, 2015.
  22. 22.Zhang, H.; Goodfellow, I. J.; Metaxas, D. N.; Odena, A. Self-attention generative adversarial networks. In: Proceedings of the 36th International Conference on Machine Learning, 7354–7363, 2019.
  23. 23.Chu, X.; Yang, W.; Ouyang, W. L.; Ma, C.; Yuille, A. L.; Wang, X. G. Multi-context attention for human pose estimation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5669–5678, 2017.
  24. 24.Dai, T.; Cai, J. R.; Zhang, Y. B.; Xia, S. T.; Zhang, L. Second-order attention network for single image super-resolution. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11057–11066, 2019.
  25. 25.Zhang, Y. L.; Li, K. P.; Li, K.; Wang, L. C.; Zhong, B. N.; Fu, Y. Image super-resolution using very deep residual channel attention networks. In: Computer Vision – ECCV 2018. Lecture Notes in Computer Science, Vol. 11211. Ferrari, V.; Hebert, M.; Sminchisescu, C.; Weiss, Y. Eds. Springer Cham, 294–310, 2018.
  26. 26.Xie, S. N.; Liu, S. N.; Chen, Z. Y.; Tu, Z. W. Attentional ShapeContextNet for point cloud recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4606–4615, 2018.
  27. 27.Guo, M. H.; Cai, J. X.; Liu, Z. N.; Mu, T. J.; Martin, R. R.; Hu, S. M. PCT: Point cloud transformer. Computational Visual Media Vol. 7, No. 2, 187–199, 2021.
  28. 28.Su, W. J.; Zhu, X. Z.; Cao, Y.; Li, B.; Lu, L. W.; Wei, F. R.; Dai, J. L-BERT: Pre-training of generic visual-linguistic representations. In: Proceedings of the International Conference on Learning Representations, 2020.
  29. 29.Xu, T.; Zhang, P. C.; Huang, Q. Y.; Zhang, H.; Gan, Z.; Huang, X. L.; He, X. AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1316–1324, 2018.
  30. 30.Wu, Y. X.; He, K. M. Group normalization. International Journal of Computer Vision Vol. 128, No. 3, 742–755, 2020.
  31. 31.Mnih, V.; Heess, N.; Graves, A.; Kavukcuoglu, K. Recurrent models of visual attention. In: Proceedings of the 27th International Conference on Neural Information Processing Systems, Vol. 2, 2204–2212, 2014.
  32. 32.Jaderberg, M.; Simonyan, K.; Zisserman, A.; Kavukcuoglu, K. Spatial transformer networks. In: Proceedings of the 28th International Conference on Neural Information Processing Systems, Vol. 2, 2017–2025, 2015.
  33. 33.Vaswani, A.; Shazeer, N. M.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing System, 6000–6010, 2017.
  34. 34.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16×16 words: Transformers for image recognition at scale. In: Proceedings of the 9th International Conference on Learning Representations, 2021.
  35. 35.Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhutdinov, R.;. Zemel, R.; Bengio, Y. Show, attend and tell: Neural image caption generation with visual attention. In: Proceedings of the 32nd International Conference on Machine Learning, 2048–2057, 2015.
  36. 36.Zhu, X. Z.; Hu, H.; Lin, S.; Dai, J. F. Deformable ConvNets V2: More deformable, better results. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9300–9308, 2019.
  37. 37.Wang, Q. L.; Wu, B. G.; Zhu, P. F.; Li, P. H.; Zuo, W. M.; Hu, Q. H. ECA-net: Efficient channel attention for deep convolutional neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11531–11539, 2020.
  38. 38.Devlin, J.; Chang, M. W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  39. 39.Yang, Z. L.; Dai, Z. H.; Yang, Y. M.; Carbonell, J. G.; Salakhutdinov, R.; Le, Q. V. XLNet: Generalized autoregressive pretraining for language understanding. In: Proceedings of the 33rd Conference on Neural Information Processing Systems, 2019.
  40. 40.Li, X.; Zhong, Z. S.; Wu, J. L.; Yang, Y. B.; Lin, Z. C.; Liu, H. Expectation-maximization attention networks for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9166–9175, 2019.
  41. 41.Huang, Z. L.; Wang, X. G.; Huang, L. C.; Huang, C.; Wei, Y. C.; Liu, W. Y. CCNet: Criss-cross attention for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence doi: 10.1109/TPAMI.2020.3007032, 2020.
  42. 42.Geng, Z.; Guo, M.-H.; Chen, H.; Li, X.; Wei, K.; Lin, Z. Is attention better than matrix decomposition? In: Proceedings of the International Conference on Learning Representations, 2021.
  43. 43.Ramachandran, P.; Parmar, N.; Vaswani, A.; Bello, I.; Levskaya, A.; Shlens, J. Stand-alone self-attention in vision models. In: Proceedings of the 33rd Conference on Neural Information Processing Systems, 2019.
  44. 44.Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.-H.; Tay, F. E.; Feng, J.; Yan, S. Tokens-to-Token ViT: Training vision transformers from scratch on ImageNet. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 558–567, 2021.
  45. 45.Wang, W. H.; Xie, E. Z.; Li, X.; Fan, D. P.; Song, K. T.; Liang, D.; Lu, T.; Luo, P.; Shao, L. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE/CVF International Conference on Computer Visio, 568–578, 2021.
  46. 46.Liu, Z.; Lin, Y. T.; Cao, Y.; Hu, H.; Guo, B. N. Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 10012–10022, 2021.
  47. 47.Wu, H.; Xiao, B.; Codella, N.; Liu, M.; Dai, X.; Yuan, L.; Zhang, L. CvT: Introducing convolutions to vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 22–31, 2021.
  48. 48.Yuan, L.; Hou, Q. B.; Jiang, Z. H.; Feng, J. S.; Yan, S. C. VOLO: Vision outlooker for visual recognition. arXiv preprint arXiv:2106.13112, 2021.
  49. 49.Dai, Z. H.; Liu, H. X.; Le, Q. V.; Tan, M. X. CoAtNet: Marrying convolution and attention for all data sizes. arXiv preprint arXiv:2106.04803, 2021.
  50. 50.Chen, L.; Zhang, H. W.; Xiao, J.; Nie, L. Q.; Shao, J.; Liu, W.; Chua, T. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6298–6306, 2017.
  51. 51.Nair, V.; Hinton, G. E. Rectified linear units improve restricted Boltzmann machines. In: Proceedings of the 27th International Conference on Machine Learning, 807–814, 2010.
  52. 52.Ioffe, S.; Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Proceedings of the 32nd International Conference on International Conference on Machine Learning, Vol. 37, 448–456, 2015.
  53. 53.Zhang, H.; Dana, K.; Shi, J. P.; Zhang, Z. Y.; Wang, X. G.; Tyagi, A.; Agrawal, A. Context encoding for semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7151–7160, 2018.
  54. 54.Gao, Z. L.; Xie, J. T.; Wang, Q. L.; Li, P. H. Global second-order pooling convolutional networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3019–3028, 2019.
  55. 55.Lee, H.; Kim, H. E.; Nam, H. SRM: A style-based recalibration module for convolutional neural networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 1854–1862, 2019.
  56. 56.Yang, Z. X.; Zhu, L. C.; Wu, Y.; Yang, Y. Gated channel transformation for visual recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11791–11800, 2020.
  57. 57.Qin, Z. Q.; Zhang, P. Y.; Wu, F.; Li, X. FcaNet: Frequency channel attention networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 783–792, 2021.
  58. 58.Diba, A. L.; Fayyaz, M.; Sharma, V.; Arzani, M. M.; Yousefzadeh, R.; Gall, J.; van Gool, L. Spatio-temporal channel correlation networks for action classification. In: Computer Vision – ECCV 2018. Lecture Notes in Computer Science, Vol. 11208. Ferrari, V.; Hebert, M.; Sminchisescu, C.; Weiss, Y. Eds. Springe Cham, 299–315, 2018.
  59. 59.Chen, Z. R.; Li, Y.; Bengio, S.; Si, S. You look twice: GaterNet for dynamic filter selection in CNNs. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9164–9172, 2019.
  60. 60.Shi, H. Y.; Lin, G. S.; Wang, H.; Hung, T. Y.; Wang, Z. H. SpSequenceNet: Semantic segmentation network on 4D point clouds. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4573–4582, 2020.
  61. 61.Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Vedaldi, A. Gather-excite: Exploiting feature context in convolutional neural networks. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 9423–9433, 2018.
  62. 62.Yan, X.; Zheng, C. D.; Li, Z.; Wang, S.; Cui, S. G. PointASNL: Robust point clouds processing using nonlocal neural networks with adaptive sampling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5588–5597, 2020.
  63. 63.Hu, H.; Gu, J. Y.; Zhang, Z.; Dai, J. F.; Wei, Y. C. Relation networks for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3588–3597, 2018.
  64. 64.Zhang, H.; Zhang, H.; Wang, C. G.; Xie, J. Y. Co-occurrent features in semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 548–557, 2019.
  65. 65.Bello, I.; Zoph, B.; Le, Q.; Vaswani, A.; Shlens, J. Attention augmented convolutional networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 3285–3294, 2019.
  66. 66.Zhu, X. Z.; Cheng, D. Z.; Zhang, Z.; Lin, S.; Dai, J. F. An empirical study of spatial attention mechanisms in deep networks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 6687–6696, 2019.
  67. 67.Li, X.; Yang, Y. B.; Zhao, Q. J.; Shen, T. C.; Lin, Z. C.; Liu, H. Spatial pyramid based graph reasoning for semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8947–8956, 2020.
  68. 68.Zhu, Z.; Xu, M. D.; Bai, S.; Huang, T. T.; Bai, X. Asymmetric non-local neural networks for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 593–602, 2019.
  69. 69.Cao, Y.; Xu, J. R.; Lin, S.; Wei, F. Y.; Hu, H. GCNet: Non-local networks meet squeeze-excitation networks and beyond. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop, 1971–1980, 2019.
  70. 70.Chen, Y.; Kalantidis, Y.; Li, J.; Yan, S.; Feng, J. A2-nets: Double attention networks. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 350–359, 2018.
  71. 71.Chen, Y. P.; Rohrbach, M.; Yan, Z. C.; Yan, S. C.; Feng, J. S.; Kalantidis, Y. Graph-based global reasoning networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 433–442, 2019.
  72. 72.Zhang, S. Y.; Yan, S. P.; He, X. M. LatentGNN: Learning efficient non-local relations for visual recognition. In: Proceedings of the 36th International Conference on Machine Learning, 7374–7383, 2019.
  73. 73.Yuan, Y.; Chen, X.; Chen, X.; Wang, J. Segmentation transformer: Object-contextual representations for semantic segmentation. arXiv preprint arXiv: 1909.11065, 2019.
  74. 74.Yin, M. H.; Yao, Z. L.; Cao, Y.; Li, X.; Zhang, Z.; Lin, S.; Hu, H. Disentangled non-local neural networks. In: Computer Vision – ECCV 2020. Lecture Notes in Computer Science, Vol. 12360. Vedaldi, A.; Bischof, H.; Brox, T.; Frahm, J. M. Eds. Springer Cham, 191–207, 2020.
  75. 75.Guo, M. H.; Liu, Z. N.; Mu, T. J.; Hu, S. M. Beyond self-attention: External attention using two linear layers for visual tasks. arXiv preprint arXiv:2105.02358, 2021.
  76. 76.Hu, H.; Zhang, Z.; Xie, Z. D.; Lin, S. Local relation networks for image recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 3463–3472, 2019.
  77. 77.Zhao, H. S.; Jia, J. Y.; Koltun, V. Exploring self-attention for image recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10073–10082, 2020.
  78. 78.Chen, M.; Radford, A.; Child, R.; Wu, J.; Jun, H.; Luan, D.; Sutskever, I. Generative pretraining from pixels. In: Proceedings of the 37th International Conference on Machine Learning, 1691–1703, 2020.
  79. 79.Chen, H. T.; Wang, Y. H.; Guo, T. Y.; Xu, C.; Deng, Y. P.; Liu, Z. H.; Ma, S.; Xu, C.; Xu, C.; Gao, W. Pre-trained image processing transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12294–12305, 2021.
  80. 80.Zhao, H.; Jiang, L.; Jia, J.; Torr, P.; Koltun, V. Point transformer. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 16259–16268, 2021.
  81. 81.Zheng, S. X.; Lu, J. C.; Zhao, H. S.; Zhu, X. T.; Luo, Z. K.; Wang, Y. B.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6877–6886, 2021.
  82. 82.Han, K.; Xiao, A.; Wu, E.; Guo, J.; Xu, C.; Wang, Y. Transformer in transformer. arXiv preprint arXiv:2103.00112, 2021.
  83. 83.Liu, S. L.; Zhang, L.; Yang, X.; Su, H.; Zhu, J. Query2Label: A simple transformer way to multi-label classification. arXiv preprint arXiv:2107.10834, 2021.
  84. 84.Chen, X. L.; Xie, S. N.; He, K. M. An empirical study of training self-supervised visual transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9640–9649, 2021.
  85. 85.Bao, H. B.; Dong, L.; Wei, F. R. BEiT: BERT pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
  86. 86.Xie, E. Z.; Wang, W. H.; Yu, Z. D.; Anandkumar, A.; Alvarez, J.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. arXiv preprint arXiv:2105.15203, 2021.
  87. 87.Zhao, H.; Zhang, Y.; Liu, S.; Shi, J.; Loy, C. C.; Lin, D.; Jia, J. PSANet: Point-wise spatial attention network for scene parsing. In: Computer Vision – ECCV 2018. Lecture Notes in Computer Science, Vol. 11213. Ferrari, V.; Hebert, M.; Sminchisescu, C.; Weiss, Y. Eds. Springer Cham, 270–286, 2018.
  88. 88.Ba, J.; Mnih, V.; Kavukcuoglu, K. Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755, 2014.
  89. 89.Sharma, S.; Kiros, R.; Salakhutdinov, R. Action recognition using visual attention. arXiv preprint arXiv:1511.04119, 2015.
  90. 90.Girdhar, R.; Ramanan, D. Attentional pooling for action recognition. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, 33–44, 2017.
  91. 91.Li, Z. Y.; Gavrilyuk, K.; Gavves, E.; Jain, M.; Snoek, C. G. M. VideoLSTM convolves, attends and flows for action recognition. Computer Vision and Image Understanding Vol. 166, 41–50, 2018.
  92. 92.Yue, K. Y.; Sun, M.; Yuan, Y. C.; Zhou, F.; Ding, E. R.; Xu, F. X. Compact generalized non-local network. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 6511–6520, 2018.
  93. 93.Liu, X. H.; Han, Z. Z.; Wen, X.; Liu, Y. S.; Zwicker, M. L2G auto-encoder: Understanding point clouds by local-to-global reconstruction with hierarchical self-attention. In: Proceedings of the 27th ACM International Conference on Multimedia, 989–997, 2019.
  94. 94.Paigwar, A.; Erkent, O.; Wolf, C.; Laugier, C. Attentional PointNet for 3D-object detection in point clouds. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 1297–1306, 2019.
  95. 95.Wen, X.; Han, Z. Z.; Youk, G.; Liu, Y. S. CF-SIS: Semantic-instance segmentation of 3D point clouds by context fusion with self-attention. In: Proceedings of the 28th ACM International Conference on Multimedia, 1661–1669, 2020.
  96. 96.Yang, J. C.; Zhang, Q.; Ni, B. B.; Li, L. G.; Liu, J. X.; Zhou, M. D.; Tian, Q. Modeling point clouds with self-attention and gumbel subset sampling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3318–3327, 2019.
  97. 97.Xu, J.; Zhao, R.; Zhu, F.; Wang, H. M.; Ouyang, W. L. Attention-aware compositional network for person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2119–2128, 2018.
  98. 98.Liu, H.; Feng, J. S.; Qi, M. B.; Jiang, J. G.; Yan, S. C. End-to-end comparative attention networks for person re-identification. IEEE Transactions on Image Processing Vol. 26, No. 7, 3492–3506, 2017.
  99. 99.Zheng, Z. D.; Zheng, L.; Yang, Y. Pedestrian alignment network for large-scale person re-identification. IEEE Transactions on Circuits and Systems for Video Technology Vol. 29, No. 10, 3037–3045, 2019.
  100. 100.Li, K. P.; Wu, Z. Y.; Peng, K. C.; Ernst, J.; Fu, Y. Tell me where to look: Guided attention inference network. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9215–9223, 2018.
  101. 101.Zhang, Z. Z.; Lan, C. L.; Zeng, W. J.; Jin, X.; Chen, Z. B. Relation-aware global attention for person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3183–3192, 2020.
  102. 102.Zhao, B.; Wu, X.; Feng, J. S.; Peng, Q.; Yan, S. C. Diversified visual attention networks for fine-grained object classification. IEEE Transactions on Multimedia Vol. 19, No. 6, 1245–1256, 2017.
  103. 103.Bryan, B.; Gong, Y.; Zhang, Y. Z.; Poellabauer, C. Second-order non-local attention networks for person re-identification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 3759–3768, 2019.
  104. 104.Zheng, H. L.; Fu, J. L.; Mei, T.; Luo, J. B. Learning multi-attention convolutional neural network for fine-grained image recognition. In: Proceedings of the IEEE International Conference on Computer Vision, 5219–5227, 2017.
  105. 105.Fu, J. L.; Zheng, H. L.; Mei, T. Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4476–4484, 2017.
  106. 106.Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic anchor boxes are better queries for DETR. arXiv preprint arXiv:2201.12329, 2022.
  107. 107.Yang, G. Y.; Li, X. L.; Martin, R.; Hu, S. M. Sampling equivariant self-attention networks for object detection in aerial images. arXiv preprint arXiv:2111.03420, 2021.
  108. 108.Zheng, H. L.; Fu, J. L.; Zha, Z. J.; Luo, J. B. Looking for the devil in the details: Learning trilinear attention sampling network for fine-grained image recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5007–5016, 2019.
  109. 109.Lee, J.; Lee, Y.; Kim, J.; Kosiorek, A. R.; Choi, S.; Teh Y. W. Set transformer: A framework for attention-based permutation-invariant neural networks. In: Proceedings of the 36th International Conference on Machine Learning, 3744–3753, 2019.
  110. 110.Xu, S. J.; Cheng, Y.; Gu, K.; Yang, Y.; Chang, S. Y.; Zhou, P. Jointly attentive spatial-temporal pooling networks for video-based person re-identification. In: Proceedings of the IEEE International Conference on Computer Vision, 4743–4752, 2017.
  111. 111.Zhang, R. M.; Li, J. Y.; Sun, H. B.; Ge, Y. Y.; Luo, P.; Wang, X. G.; Lin, L. SCAN: Self-and-collaborative attention network for video person re-identification. IEEE Transactions on Image Processing Vol. 28, No. 10, 4870–4882, 2019.
  112. 112.Chen, D. P.; Li, H. S.; Xiao, T.; Yi, S.; Wang, X. G. Video person re-identification with competitive snippet-similarity aggregation and co-attentive snippet embedding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1169–1178, 2018.
  113. 113.Srivastava, R. K.; Greff, K.; Schmidhuber, J. Training very deep networks. In: Proceedings of the 28th International Conference on Neural Information Processing Systems, Vol. 2, 2377–2385, 2015.
  114. 114.Li, X.; Wang, W. H.; Hu, X. L.; Yang, J. Selective kernel networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 510–519, 2019.
  115. 115.Zhang, H.; Wu, C.; Zhang, Z.; Zhu, Y.; Lin, H.; Zhang, Z.; Sun, Y.; He, T.; Mueller, J.; Manmatha, R.; et al. ResNeSt: Split-attention networks. arXiv preprint arXiv:2004.08955, 2020.
  116. 116.Chen, Y. P.; Dai, X. Y.; Liu, M. C.; Chen, D. D.; Yuan, L.; Liu, Z. C. Dynamic convolution: attention over convolution kernels. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11027–11036, 2020.
  117. 117.Park, J.; Woo, S.; Lee, J.-Y.; Kweon, I. S. BAM: Bottleneck attention module. arXiv preprint arXiv:1807.06514, 2018.
  118. 118.Yang, L.; Zhang, R.-Y.; Li, L.; Xie, X. SimAM: A simple, parameter-free attention module for convolutional neural networks. In: Proceedings of the 38th International Conference on Machine Learning, 11863–11874, 2021.
  119. 119.Wang, F.; Jiang, M. Q.; Qian, C.; Yang, S.; Li, C.; Zhang, H. G.; Wang, X.; Tang, X. Residual attention network for image classification. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6450–6458, 2017.
  120. 120.Guo, M.-H.; Lu, C.-Z.; Liu, Z.-N.; Cheng, M.-M.; Hu, S.-M. Visual attention network. arXiv preprint arXiv:2202.09741, 2022.
  121. 121.Liu, J. J.; Hou, Q. B.; Cheng, M. M.; Wang, C. H.; Feng, J. S. Improving convolutional networks with self-calibrated convolutions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10093–10102, 2020.
  122. 122.Misra, D.; Nalamada, T.; Arasanipalai, A. U.; Hou, Q. B. Rotate to attend: Convolutional triplet attention module. In: Proceedings of the IEEE Winter Conference on Applications of Computer Vision, 3138–3147, 2021.
  123. 123.Linsley, .; Shiebler, D.; Eberhardt, S.; Serre, T. Learning what and where to attend. In: Proceedings of the 7th International Conference on Learning Representations, 2019.
  124. 124.Roy, A. G.; Navab, N.; Wachinger, C. Recalibrating fully convolutional networks with spatial and channel “squeeze and excitation” blocks. IEEE Transactions on Medical Imaging Vol. 38, No. 2, 540–549, 2019.
  125. 125.Hou, Q. B.; Zhang, L.; Cheng, M. M.; Feng, J. S. Strip pooling: Rethinking spatial pooling for scene parsing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4002–4011, 2020.
  126. 126.You, H. X.; Feng, Y. F.; Ji, R. R.; Gao, Y. PVNet: A joint convolutional network of point cloud and multi-view for 3D shape recognition. In: Proceedings of the 26th ACM International Conference on Multimedia, 1310–1318, 2018.
  127. 127.Xie, Q.; Lai, Y. K.; Wu, J.; Wang, Z. T.; Zhang, Y. M.; Xu, K.; Wang, J. MLCVNet: Multi-level context VoteNet for 3D object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10444–10453, 2020.
  128. 128.Wang, C.; Zhang, Q.; Huang, C.; Liu, W.; Wang, X. Mancs: A multi-task attentional network with curriculum sampling for person re-identification. In: Computer Vision – ECCV 2018. Lecture Notes in Computer Science, Vol. 11208. Ferrari, V.; Hebert, M.; Sminchisescu, C.; Weiss, Y. Eds. Springer Cham, 384–400, 2018.
  129. 129.Chen, T. L.; Ding, S. J.; Xie, J. Y.; Yuan, Y.; Chen, W. Y.; Yang, Y.; Ren, Z.; Wang, Z. ABD-net: Attentive but diverse person re-identification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 8350–8360, 2019.
  130. 130.Hou, Q. B.; Zhou, D. Q.; Feng, J. S. Coordinate attention for efficient mobile network design. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13708–13717, 2021.
  131. 131.Song, S.; Lan, C.; Xing, J.; Zeng, W.; Liu, J. An end-to-end spatio-temporal attention model for human action recognition from skeleton data. In: Proceedings of the 31st AAAI Conference on Artificial Intelligence, 4263–4270, 2017.
  132. 132.Fu, Y.; Wang, X. Y.; Wei, Y. C.; Huang, T. STA: Spatial-temporal attention for large-scale video-based person re-identification. Proceedings of the AAAI Conference on Artificial Intelligence Vol. 33, 8287–8294, 2019.
  133. 133.Gao, L. L.; Li, X. P.; Song, J. K.; Shen, H. T. Hierarchical LSTMs with adaptive attention for visual captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence Vol. 42, No. 5, 1112–1131, 2020.
  134. 134.Yan, C. G.; Tu, Y. B.; Wang, X. Z.; Zhang, Y. B.; Hao, X. H.; Zhang, Y. D.; Dai, Q. STAT: Spatial-temporal attention mechanism for video captioning. IEEE Transactions on Multimedia Vol. 22, No. 1, 229–241, 2020.
  135. 135.Meng, L. L.; Zhao, B.; Chang, B.; Huang, G.; Sun, W.; Tung, F.; Sigal, L. Interpretable spatio-temporal attention for video action recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop, 1513–1522, 2019.
  136. 136.He, B.; Yang, X. T.; Wu, Z. X.; Chen, H.; Shrivastava, A. GTA: Global temporal attention for video action understanding. arXiv preprint arXiv:2012.08510, 2020.
  137. 137.Li, S.; Bak, S.; Carr, P.; Wang, X. G. Diversity regularized spatiotemporal attention for video-based person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 369–378, 2018.
  138. 138.Zhang, Z. Z.; Lan, C. L.; Zeng, W. J.; Chen, Z. B. Multi-granularity reference-aided attentive feature aggregation for video-based person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10404–10413, 2020.
  139. 139.Shim, M.; Ho, H. I.; Kim, J.; Wee, D. READ: Reciprocal attention discriminator for image-to-video re-identification. In: Computer Vision – ECCV 2020. Lecture Notes in Computer Science, Vol. 12359. Vedaldi, A.; Bischof, H.; Brox, T.; Frahm, J. M. Eds. Springer Cham, 335–350, 2020.
  140. 140.Liu, R.; Deng, H. M.; Huang, Y. Y.; Shi, X. Y.; Li, H. S. Decoupled spatial-temporal transformer for video inpainting. arXiv preprint arXiv:2104.06637, 2021.
  141. 141.Chaudhari, S.; Mithal, V.; Polatkan, G.; Ramanath, R. An attentive survey of attention models. ACM Transactions on Intelligent Systems and Technology Vol. 12, No. 5, Article No. 53, 2021.
  142. 142.Xu, Y. F.; Wei, H. P.; Lin, M. X.; Deng, Y. Y.; Sheng, K. K.; Zhang, M. D.; Tang, F.; Dong, W.; Huang, F.; Xu, C. Transformers in computational visual media: A survey. Computational Visual Media Vol. 8, No. 1, 33–62, 2022.
  143. 143.Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y.; et al. A survey on visual transformer. arXiv preprint arXiv:2012.12556, 2020.
  144. 144.Khan, S.; Naseer, M.; Hayat, M.; Zamir, S. W.; Khan, F. S.; Shah, M. Transformers in vision: A survey. ACM Computing Surveys https://doi.org/10.1145/3505244, 2022.
  145. 145.Wang, F.; Tax, D. M. J. Survey on the attention based RNN model and its applications in computer vision. arXiv preprint arXiv:1601.06823, 2016.
  146. 146.He, K. M.; Zhang, X. Y.; Ren, S. Q.; Sun, J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778, 2016.
  147. 147.Fang, P. F.; Zhou, J. M.; Roy, S.; Petersson, L.; Harandi, M. Bilinear attention networks for person retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 8029–8038, 2019.
  148. 148.Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Computation Vol. 9, No. 8, 1735–1780, 1997.
  149. 149.Sutton, R. S.; McAllester, D. A.; Singh, S. P.; Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In: Proceedings of the 12th International Conference on Neural Information Processing Systems, 1057–1063, 1999.
  150. 150.Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014.
  151. 151.Lin, Z. H.; Feng, M. W.; Santos, C. N. D.; Yu, M.; Bengio, Y. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130, 2017.
  152. 152.Dai, Z. H.; Yang, Z. L.; Yang, Y. M.; Carbonell, J.; Le, Q.; Salakhutdinov, R. Transformer-XL: Attentive language models beyond a fixed-length context. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2978–2988, 2019.
  153. 153.Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X. Y.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  154. 154.Zhu, X. Z.; Su, W. J.; Lu, L. W.; Li, B.; Wang, X. G.; Dai, J. F. Deformable DETR: Deformable transformers for end-to-end object detection. In: Proceedings of the International Conference on Learning Representations, 2021.
  155. 155.Liu, W.; Rabinovich, A.; Berg, A. C. ParseNet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015.
  156. 156.Peng, C.; Zhang, X. Y.; Yu, G.; Luo, G. M.; Sun, J. Large kernel matters—Improve semantic segmentation by global convolutional network. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1743–1751, 2017.
  157. 157.Zhao, H. S.; Shi, J. P.; Qi, X. J.; Wang, X. G.; Jia, J. Y. Pyramid scene parsing network. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6230–6239, 2017.
  158. 158.He, K. M.; Zhang, X. Y.; Ren, S. Q.; Sun, J. Spatial pyramid pooling in deep convolutional networks for visual recognition. In: Computer Vision – ECCV 2014. Lecture Notes in Computer Science, Vol. 8691. Fleet, D.; Pajdla, T.; Schiele, B.; Tuytelaars, T. Eds. Springer Cham, 346–361, 2014.
  159. 159.Tolstikhin, I.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X. H.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; et al. MLP-mixer: An all-MLP architecture for vision. In: Proceedings of the 35th Conference on Neural Information Processing Systems, 2021.
  160. 160.Touvron, H.; Bojanowski, P.; Caron, M.; Cord, M.; El-Nouby, A.; Grave, E.; Izacard, G.; Joulin, A.; Synnaeve, G.; Verbeek, J.; et al. ResMLP: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv: 2105.03404, 2021.
  161. 161.Shaw, P.; Uszkoreit, J.; Vaswani, A. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018.
  162. 162.Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. In: Proceedings of the 34th Conference on Neural Information Processing Systems, 2020.
  163. 163.Ba, J. L.; Kiros, J. R.; Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  164. 164.Hendrycks, D.; Gimpel, K. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016.
  165. 165.Sun, C.; Shrivastava, A.; Singh, S.; Gupta, A. Revisiting unreasonable effectiveness of data in deep learning era. In: Proceedings of the IEEE International Conference on Computer Vision, 843–852, 2017.
  166. 166.Deng, J.; Dong, W.; Socher, R.; Li, L. J.; Kai, L.; Li, F. F. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 248–255, 2009.
  167. 167.Zhou, D. Q.; Kang, B. Y.; Jin, X. J.; Yang, L. J.; Lian, X. C.; Jiang, Z. H.; Hou, Q. B.; Feng, J. S. DeepViT: Towards deeper vision transformer. arXiv preprint arXiv:2103.11886, 2021.
  168. 168.Touvron, H.; Cord, M.; Sablayrolles, A.; Synnaeve, G.; J´egou, H. Going deeper with image transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 32–42, 2021.
  169. 169.Liu, R.; Deng, H. M.; Huang, Y. Y.; Shi, X. Y.; Lu, L. W.; Sun, W. X.; Wang, X.; Dai, J.; Li, H. FuseFormer: Fusing fine-grained information in transformers for video inpainting. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 14040–14049, 2021.
  170. 170.He, K. M.; Chen, X. L.; Xie, S. N.; Li, Y. H.; Doll´ar, P.; Girshick, R. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021.
  171. 171.Guo, M. H.; Liu, Z. N.; Mu, T. J.; Liang, D.; Martin, R. R.; Hu, S. M. Can attention enable MLPs to catch up with CNNs? Computational Visual Media Vol. 7, No. 3, 283–288, 2021.
  172. 172.Li, J. N.; Zhang, S. L.; Wang, J. D.; Gao, W.; Tian, Q. Global–local temporal representations for video person re-identification. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 3957–3966, 2019.
  173. 173.Liu, Z. Y.; Wang, L. M.; Wu, W.; Qian, C.; Lu, T. TAM: Temporal adaptive module for video recognition. arXiv preprint arXiv:2005.06803, 2020.
  174. 174.Yang, B.; Bender, G.; Le, Q. V.; Ngiam, J. CondConv: Conditionally parameterized convolutions for efficient inference. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems, Article No. 117, 1307–1318, 2019.
  175. 175.Spillmann, L.; Dresp-Langley, B.; Tseng, C. H. Beyond the classical receptive field: The effect of contextual stimuli. Journal of Vision Vol. 15, No. 9, 7, 2015.
  176. 176.Xie, S. N.; Girshick, R.; Doll´ar, P.; Tu, Z. W.; He, K. M. Aggregated residual transformations for deep neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5987–5995, 2017.
  177. 177.Webb, B. S.; Dhruv, N. T.; Solomon, S. G.; Tailby, C.; Lennie, P. Early and late mechanisms of surround suppression in striate cortex of macaque. Journal of Neuroscience Vol. 25, No. 50, 11666–11675, 2005.
  178. 178.Yang, J. R.; Zheng, W. S.; Yang, Q. Z.; Chen, Y. C.; Tian, Q. Spatial-temporal graph convolutional network for video-based person re-identification. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3286–3296, 2020.
  179. 179.Szegedy, C.; Liu, W.; Jia, Y. Q.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1–9, 2015.
  180. 180.Caron, M.; Touvron, H.; Misra, I.; J´egou, H.; Mairal, J.; Bojanowski, P.; Joulin, A. Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 9650–9660, 2021.
  181. 181.Qian, N. On the momentum term in gradient descent learning algorithms. Neural Networks Vol. 12, No. 1, 145–151, 1999.
  182. 182.Kingma, D. P.; Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  183. 183.Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  184. 184.Chen, X. N.; Hsieh, C. J.; Gong, B. Q. When vision transformers outperform ResNets without pretraining or strong data augmentations. arXiv preprint arXiv:2106.01548, 2021.
  185. 185.Foret, P.; Kleiner, A.; Mobahi, H.; Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. arXiv preprint arXiv:2010.01412, 2020.

Citation

MLA
Guo, M.-H., et al. “Attention Mechanisms in Computer Vision: A Survey”. Computational Visual Media, vol. 8, no. 3, 2022, pp. 331–68, https://doi.org/10.1007/s41095-022-0271-y.
APA
Guo, M.-H., Xu, T.-X., Liu, J.-J., Liu, Z.-N., Jiang, P.-T., Mu, T.-J., Zhang, S.-H., Martin, R. R., Cheng, M.-M., & Hu, S.-M. (2022). Attention mechanisms in computer vision: A survey. Computational Visual Media, 8(3), 331–368. https://doi.org/10.1007/s41095-022-0271-y
Chicago
Guo, M.-H., T.-X. Xu, J.-J. Liu, et al. 2022. “Attention Mechanisms in Computer Vision: A Survey”. Computational Visual Media 8 (3): 331–68. https://doi.org/10.1007/s41095-022-0271-y.
Harvard
Guo, M.-H. et al. (2022) “Attention mechanisms in computer vision: A survey”, Computational Visual Media, 8(3), pp. 331–368. Available at: https://doi.org/10.1007/s41095-022-0271-y.
Vancouver
1. Guo M-H, Xu T-X, Liu J-J, Liu Z-N, Jiang P-T, Mu T-J, Zhang S-H, Martin RR, Cheng M-M, Hu S-M (2022) Attention mechanisms in computer vision: A survey. Computational Visual Media 8:331–368

BibTeX

@article{Guo_2022, title={Attention mechanisms in computer vision: A survey}, volume={8}, ISSN={2096-0433}, url={http://dx.doi.org/10.1007/s41095-022-0271-y}, DOI={10.1007/s41095-022-0271-y}, number={3}, journal={Computational Visual Media}, publisher={Tsinghua University Press}, author={Guo, Meng-Hao and Xu, Tian-Xing and Liu, Jiang-Jiang and Liu, Zheng-Ning and Jiang, Peng-Tao and Mu, Tai-Jiang and Zhang, Song-Hai and Martin, Ralph R. and Cheng, Ming-Ming and Hu, Shi-Min}, year={2022}, month=Sept, pages={331–368} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF