SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks

Lingxiao YangRu-Yuan ZhangLida LiXiaohua Xie

article2021ICML1,857 citations

Proposes a neuroscience-inspired, parameter-free attention module that infers true 3-D weights for convolutional neural networks via an analytical closed-form solution implemented in under ten lines of code.

Listen

Deep convolutional neural networks are widely used across computer vision applications, but maximizing their accuracy often requires increasing model size or incorporating specialized components called attention modules. Existing attention mechanisms typically refine features by focusing separately on channel or spatial dimensions and rely heavily on heuristic, hand-tuned architectures that add extra parameters and computational overhead. Developing a simple, unified method to learn full three-dimensional attention weights without inflating network size remains a key challenge for deploying efficient vision models.

The article aims to design and evaluate a simple, parameter-free attention module that infers three-dimensional importance weights for individual neurons based on neuroscience principles. The objective is to demonstrate that this module improves representation quality and task performance across diverse vision architectures without introducing extra learnable parameters.

The authors developed an energy function grounded in the visual neuroscience concept of spatial suppression, where distinctive, informative neurons suppress neighboring neuronal activity. By deriving an analytical closed-form solution to this energy function, the module directly calculates neuron-level importance using channel-wise mean and variance without requiring iterative optimization or complex pooling operations. The approach was evaluated by integrating the module into various standard network backbones, including ResNet variants and MobileNetV2, across benchmark image classification, object detection, and instance segmentation tasks.

The findings show that the proposed module consistently improves classification accuracy across all tested architectures without adding any parameters or floating-point operations. On standard image classification benchmarks, the module improved top-1 accuracy across small and large network backbones, outperforming or matching heavier attention mechanisms such as Squeeze-and-Excitation while retaining high inference throughput (such as 147 frames per second on ResNet-18). In downstream object detection and instance segmentation tasks, integrating the module into standard detectors improved detection accuracy by approximately 1.4 to 1.6 points over baseline models, matching or slightly exceeding the performance of alternative modules that added between 2.5 million and 4.7 million parameters.

These results indicate that computer vision models can achieve superior feature selectivity and accuracy without paying a penalty in parameter count or requiring architectural search for attention sub-blocks. Because the module adds zero parameters and can be expressed in less than ten lines of standard code, it significantly simplifies model development pipelines and eases memory constraints for resource-sensitive deployments.

Engineering and research teams working on visual recognition tasks should consider adopting this module as a lightweight, plug-and-play enhancement for existing convolutional backbones. Before broad operational rollout, practitioners should perform standard cross-validation to select the regularization hyperparameter for their specific datasets and evaluate execution latency on their target hardware platforms.

Confidence in these findings is high across standard academic benchmarks, though the evaluation relies on the assumption that neurons within a single channel share underlying statistical distributions. While inference speed was measured on specific graphical processing hardware, future work could further investigate performance on specialized edge accelerators and across video or non-vision domains.

  • Paper: Squeeze-and-Excitation Networks, Jie Hu et al. (2018). Read SE-Net first to understand the channel-recalibration baseline SimAM contrasts with when arguing for neuron-level attention without learned parameters.
  • Paper: CBAM: Convolutional Block Attention Module, Sanghyun Woo et al. (2018). CBAM establishes the combined channel-and-spatial attention design that makes SimAM’s unified, three-dimensional neuron weighting easier to place.
  • Paper: BAM: Bottleneck Attention Module, Jongchan Park et al. (2018). BAM provides an earlier channel-plus-spatial module and efficiency trade-off against which SimAM’s parameter-free construction can be understood.
  • Paper: ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks, Qilong Wang et al. (2019). ECA-Net shows how prior lightweight channel attention reduced overhead while retaining learnable weights, clarifying the design distinction SimAM makes.

No sufficiently relevant recommendations were found.

Cover for SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks

Abstract

In this paper, we propose a conceptually simple but very effective attention module for Convolutional Neural Networks (ConvNets). In contrast to existing channel-wise and spatial-wise attention modules, our module instead infers 3-D attention weights for the feature map in a layer without adding parameters to the original networks. Specifically, we base on some well-known neuroscience theories and propose to optimize an energy function to find the importance of each neuron. We further derive a fast closed-form solution for the energy function, and show that the solution can be implemented in less than ten lines of code. Another advantage of the module is that most of the operators are selected based on the solution to the defined energy function, avoiding too many efforts for structure tuning. Quantitative evaluations on various visual tasks demonstrate that the proposed module is flexible and effective to improve the representation ability of many ConvNets. Our code is available at Pytorch-SimAM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overview of existing attention modules
  • 3.2. Our attention module
  • 4. Experiments
  • 4.1. CIFAR Classification
  • 4.2. ImageNet Classification
  • 4.3. Object Detection and Instance Segmentation
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — SimAM 3D Attention Mechanism

    model/method

    SimAM (Simple Attention Module) is a parameter-free, 3-D attention module designed to refine feature representations in Convolutional Neural Networks (ConvNets). Unlike 1-D channel attention (which assigns a single scalar per channel) or 2-D spatial attention (which assigns a single scalar per spatial location across channels), SimAM computes full 3-D attention weights W∈RC×H×WW \in \mathbb{R}^{C \times H \times W} for an input tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, assigning a distinct importance weight to every single neuron across both spatial and channel dimensions.

    The design is grounded in visual neuroscience principles of spatial suppression: the most informative neurons exhibit firing patterns distinct from surrounding neurons and suppress background activity. SimAM formulates the importance of each individual neuron by measuring its linear separability from surrounding neurons in the same channel via an energy optimization problem with an exact, closed-form solution. The resulting energy values directly determine the 3-D attention scaling weights without adding any learnable parameters to the backbone network.

  2. Knowl 2 — Closed-Form Minimal Energy Formulation for Neuron Importance

    equation

    For an input feature tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, let t∈Rt \in \mathbb{R} denote a target neuron and xi∈Rx_i \in \mathbb{R} denote other neurons in the same channel, where i∈{1,…,M−1}i \in \{1, \dots, M-1\} indexes spatial locations and M=H×WM = H \times W. Under a linear transformation with scalar weight wtw_t and bias btb_t, the energy function measuring the linear separability of target neuron tt from background neurons (with binary targets 11 and −1-1) with L2L_2 regularization weight λ>0\lambda > 0 is:

    et(wt,bt,y,xi)=1M−1∑i=1M−1(−1−(wtxi+bt))2+(1−(wtt+bt))2+λwt2e_t(w_t, b_t, y, x_i) = \frac{1}{M-1}\sum_{i=1}^{M-1}\left(-1 - (w_t x_i + b_t)\right)^2 + \left(1 - (w_t t + b_t)\right)^2 + \lambda w_t^2

    Under the assumption that all neurons in a single channel share the channel mean μ^=1M∑i=1Mxi\hat{\mu} = \frac{1}{M}\sum_{i=1}^{M} x_i and channel variance σ^2=1M∑i=1M(xi−μ^)2\hat{\sigma}^2 = \frac{1}{M}\sum_{i=1}^{M} (x_i - \hat{\mu})^2, the minimal energy et∗e_t^* of neuron tt attains the closed-form solution:

    et∗=4(σ^2+λ)(t−μ^)2+2σ^2+2λe_t^* = \frac{4(\hat{\sigma}^2 + \lambda)}{(t - \hat{\mu})^2 + 2\hat{\sigma}^2 + 2\lambda}

    A lower value of et∗e_t^* indicates higher linear separability and greater distinctiveness of neuron tt relative to its spatial surround, meaning 1/et∗1 / e_t^* quantifies the visual importance of neuron tt.

  3. Knowl 3 — Feature Refinement and Attention Modulation Function

    equation

    Following the neurobiological principle of sensory gain control, where attentional modulation acts as a multiplicative gain on neural responses, the refined feature map X~∈RC×H×W\tilde{X} \in \mathbb{R}^{C \times H \times W} is computed by scaling the input feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W} with the inverse of the grouped minimal energy tensor E∈RC×H×WE \in \mathbb{R}^{C \times H \times W} (which contains all minimal energies et∗e_t^* computed across channel and spatial dimensions):

    X~=sigmoid(1E)⊙X\tilde{X} = \text{sigmoid}\left(\frac{1}{E}\right) \odot X

    where ⊙\odot represents the Hadamard (element-wise) product, sigmoid(z)=11+e−z\text{sigmoid}(z) = \frac{1}{1 + e^{-z}} restricts excessively large energy inverse values while monotonically preserving relative neuron importance, and the inverse energy for target neuron tt evaluates element-wise to:

    1et∗=(t−μ^)24(σ^2+λ)+12\frac{1}{e_t^*} = \frac{(t - \hat{\mu})^2}{4(\hat{\sigma}^2 + \lambda)} + \frac{1}{2}

    with channel mean μ^\hat{\mu}, channel variance σ^2\hat{\sigma}^2, and regularizer λ>0\lambda > 0.

  4. Knowl 4 — Parameter-Free SimAM Forward Computation

    algorithm

    The forward computation of SimAM takes a batch of feature maps and an energy regularizer λ\lambda, calculates spatial statistics per channel, and returns attended feature maps using element-wise operations without maintaining any trainable parameters.

    Input: Input feature tensor XX of shape (N,C,H,W)(N, C, H, W), regularization scalar λ\lambda
    Output: Attended feature tensor X~\tilde{X} of shape (N,C,H,W)(N, C, H, W)
    n←H×W−1n \leftarrow H \times W - 1
    μ←mean of X across spatial dimensions (H,W)\mu \leftarrow \text{mean of } X \text{ across spatial dimensions } (H, W)
    d←(X−μ)2d \leftarrow (X - \mu)^2
    v←1n∑spatialdv \leftarrow \frac{1}{n} \sum_{\text{spatial}} d
    Einv←d4(v+λ)+0.5E_{\text{inv}} \leftarrow \frac{d}{4(v + \lambda)} + 0.5
    X~←X⊙sigmoid(Einv)\tilde{X} \leftarrow X \odot \text{sigmoid}(E_{\text{inv}})
    return X~\tilde{X}

    The computational complexity of this procedure is O(N⋅C⋅H⋅W)O(N \cdot C \cdot H \cdot W) in time and requires no parameter storage.

  5. Knowl 5 — Spatial Distribution Sharing Assumption in Channel Feature Maps

    assumption

    In the analytical derivation of the per-neuron minimal energy et∗e_t^*, it is assumed that all M=H×WM = H \times W neurons within a single channel of an input feature map X∈RC×H×WX \in \mathbb{R}^{C \times H \times W} follow a single shared spatial distribution characterized by the channel-wide mean μ^=1M∑i=1Mxi\hat{\mu} = \frac{1}{M}\sum_{i=1}^{M} x_i and channel-wide variance σ^2=1M∑i=1M(xi−μ^)2\hat{\sigma}^2 = \frac{1}{M}\sum_{i=1}^{M} (x_i - \hat{\mu})^2.

    This assumption replaces the strictly local statistics μt=1M−1∑i=1,xi≠tM−1xi\mu_t = \frac{1}{M-1}\sum_{i=1, x_i \ne t}^{M-1} x_i and σt2=1M−1∑i=1,xi≠tM−1(xi−μt)2\sigma_t^2 = \frac{1}{M-1}\sum_{i=1, x_i \ne t}^{M-1}(x_i - \mu_t)^2 with the global channel statistics μ^\hat{\mu} and σ^2\hat{\sigma}^2, allowing mean and variance to be calculated once per channel and reused across all spatial positions. This reduces the computational cost of solving the MM individual energy minimization problems from O(M2)O(M^2) to O(M)O(M) operations per channel.

  6. Knowl 6 — ImageNet-1K Classification Performance and Parameter Efficiency

    data/table

    Performance of various attention modules plugged into standard backbone architectures on ImageNet-1K classification (1.2M training images, 50K validation images, input resolution 224×224224 \times 224). Top-1 and Top-5 validation accuracies (%), parameter counts, parameter increase relative to baseline (+Params), FLOPs, and inference speed (FPS on a single NVIDIA GTX 1080 Ti with batch size 1 over 500 images) are compared across SE, CBAM, ECA, SRM, and SimAM (lambda=0.1\\lambda = 0.1).

    Model Top-1 Acc. Top-5 Acc. # Parameters + Params # FLOPs Inference Speed
    ResNet-18 70.33% 89.58% 11.69 M 0 1.82 G 215 FPS
    + SE 71.19% 90.21% 11.78 M 0.087 M 1.82 G 144 FPS
    + CBAM 71.24% 90.04% 11.78 M 0.090 M 1.82 G 78 FPS
    + ECA 70.71% 89.85% 11.69 M 36 1.82 G 148 FPS
    + SRM 71.09% 89.98% 11.69 M 0.004 M 1.82 G 115 FPS
    + SimAM 71.31% 89.88% 11.69 M 0 1.82 G 147 FPS
    ResNet-34 73.75% 91.60% 21.80 M 0 3.67 G 119 FPS
    + SE 74.32% 91.99% 21.95 M 0.157 M 3.67 G 81 FPS
    + CBAM 74.41% 91.85% 21.96 M 0.163 M 3.67 G 38 FPS
    + ECA 74.03% 91.73% 21.80 M 74 3.67 G 82 FPS
    + SRM 74.49% 92.01% 21.81 M 0.008 M 3.67 G 59 FPS
    + SimAM 74.46% 92.02% 21.80 M 0 3.67 G 78 FPS
    ResNet-50 76.34% 93.12% 25.56 M 0 4.11 G 89 FPS
    + SE 77.51% 93.74% 28.07 M 2.515 M 4.12 G 64 FPS
    + CBAM 77.63% 93.88% 28.09 M 2.533 M 4.12 G 33 FPS
    + ECA 77.17% 93.52% 25.56 M 88 4.12 G 64 FPS
    + SRM 77.51% 93.06% 25.59 M 0.030 M 4.11 G 56 FPS
    + SimAM 77.45% 93.66% 25.56 M 0 4.11 G 64 FPS
    ResNet-101 77.82% 93.85% 44.55 M 0 7.83 G 47 FPS
    + SE 78.39% 94.13% 49.29 M 4.743 M 7.85 G 33 FPS
    + CBAM 78.57% 94.18% 49.33 M 4.781 M 7.85 G 14 FPS
    + ECA 78.46% 94.12% 44.55 M 171 7.84 G 33 FPS
    + SRM 78.58% 94.15% 44.68 M 0.065 M 7.83 G 25 FPS
    + SimAM 78.65% 94.11% 44.55 M 0 7.83 G 32 FPS
    ResNeXt-50 77.47% 93.52% 25.03 M 0 4.26 G 70 FPS
    + SE 77.96% 93.93% 27.54 M 2.51 M 4.27 G 53 FPS
    + CBAM 78.06% 94.07% 27.56 M 2.53 M 4.27 G 32 FPS
    + ECA 77.74% 93.87% 25.03 M 86 4.27 G 54 FPS
    + SRM 78.04% 93.91% 25.06 M 0.030 M 4.26 G 46 FPS
    + SimAM 78.00% 93.93% 25.03 M 0 4.26 G 53 FPS
    MobileNetV2 71.90% 90.51% 3.50 M 0 0.31 G 99 FPS
    + SE 72.46% 90.85% 3.53 M 0.028 M 0.31 G 65 FPS
    + CBAM 72.49% 90.78% 3.54 M 0.032 M 0.32 G 35 FPS
    + ECA 72.01% 90.46% 3.50 M 59 0.31 G 66 FPS
    + SRM 72.32% 90.70% 3.51 M 0.003 M 0.31 G 53 FPS
    + SimAM 72.36% 90.74% 3.50 M 0 0.31 G 66 FPS

    The data shows that SimAM improves baseline Top-1 accuracy across all architectures (e.g., +0.98%+0.98\% on ResNet-18, +1.11%+1.11\% on ResNet-50, +0.83%+0.83\% on ResNet-101) while adding 0 parameters, keeping FLOPs constant, and matching the throughput of SE and ECA while operating substantially faster than 2D/3D combination methods like CBAM.

  7. Knowl 7 — CIFAR-10 and CIFAR-100 Classification Performance

    data/table

    Top-1 accuracy (mean ±\pm standard deviation over 5 trials in %) for standard ConvNet architectures with different attention modules evaluated on CIFAR-10 (C10) and CIFAR-100 (C100), with λ=10−4\lambda = 10^{-4} for SimAM:

    Attention ResNet-20 ResNet-56 ResNet-110 MobileNetV2
    Module C10 C100 C10 C100 C10 C100 C10 C100
    Baseline 92.330.19 68.880.15 93.580.20 72.240.37 94.510.31 75.540.24 91.860.12 71.320.09
    + SE 92.420.14 69.450.11 93.690.17 72.840.51 94.680.22 76.560.30 91.790.20 71.540.31
    + CBAM 92.600.31 69.470.35 93.820.10 72.470.52 94.830.18 76.450.54 91.880.16 71.790.22
    + ECA 92.350.35 68.890.57 93.680.08 72.450.38 94.720.15 76.330.65 92.340.23 71.240.45
    + GC 92.470.19 69.160.48 93.580.08 72.500.50 94.780.25 76.210.17 91.730.14 71.780.28
    + SimAM 92.730.18 69.570.40 93.760.13 72.820.25 94.720.18 76.420.27 92.360.20 72.080.28
    Attention PreResNet-20 PreResNet-56 PreResNet-110 WideResNet-20x10
    Module C10 C100 C10 C100 C10 C100 C10 C100
    Baseline 92.140.25 68.700.30 93.710.24 71.830.23 94.220.18 75.950.22 95.780.10 81.310.39
    + SE 92.240.06 68.700.21 93.570.15 72.570.32 94.400.18 76.680.30 96.240.04 81.300.08
    + CBAM 92.190.11 68.760.56 93.670.08 72.160.12 94.370.33 76.010.57 95.980.17 80.540.23
    + ECA 92.160.25 68.310.46 93.780.17 72.430.45 94.700.31 76.110.54 96.120.15 80.350.22
    + GC 92.190.20 68.960.48 93.770.12 72.440.19 94.850.19 75.880.28 96.120.18 79.980.17
    + SimAM 92.470.12 69.130.50 93.800.30 72.360.19 94.900.19 76.240.31 96.090.21 81.510.25

    SimAM consistently outperforms the baseline models across all 8 networks on both datasets without adding any parameters, achieving top performance on small models (ResNet-20, PreResNet-20) and large models (MobileNetV2, WideResNet-20x10).

  8. Knowl 8 — COCO Object Detection and Instance Segmentation Performance

    data/table

    Evaluation of SimAM on MS COCO dataset (trained on coco_2017_train, evaluated on coco_2017_val) using Faster R-CNN and Mask R-CNN with Feature Pyramid Networks (FPN). All backbone models are pre-trained on ImageNet-1K. Metrics include bounding box Average Precision (AP) for detection, mask AP for instance segmentation, and additional backbone parameters (+Params):

    Backbone AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L + Params
    Faster RCNN for Object Detection
    ResNet-50 37.8 58.5 40.8 21.9 41.7 48.2 0
    + SE 39.4 60.6 43.0 23.6 43.5 50.4 2.5 M
    + SimAM 39.2 60.7 42.8 22.8 43.0 50.6 0
    ResNet-101 39.6 60.3 43.0 22.5 43.7 51.4 0
    + SE 41.1 62.0 45.2 24.1 45.4 53.0 4.7 M
    + SimAM 41.2 62.4 45.0 24.0 45.6 52.8 0
    Mask RCNN for Object Detection
    ResNet-50 38.1 58.9 41.3 22.2 41.5 49.4 0
    + SE 39.9 61.1 43.4 24.6 43.6 51.3 2.5 M
    + SimAM 39.8 61.0 43.4 23.1 43.7 51.4 0
    ResNet-101 40.3 60.8 44.0 22.8 44.2 52.9 0
    + SE 41.8 62.6 45.5 24.3 46.3 54.1 4.7 M
    + SimAM 41.8 62.8 46.0 24.8 46.2 53.9 0
    Mask RCNN for Instance Segmentation
    ResNet-50 34.6 55.6 36.7 18.8 37.7 46.8 0
    + SE 36.0 57.8 38.1 20.7 39.4 48.5 2.5 M
    + SimAM 36.0 57.9 38.2 19.1 39.7 48.6 0
    ResNet-101 36.3 57.6 38.8 18.8 39.9 49.6 0
    + SE 37.2 59.4 39.7 20.2 41.2 50.4 4.7 M
    + SimAM 37.6 59.5 40.1 20.5 41.5 50.8 0

    SimAM achieves comparable detection accuracy (+1.4+1.4 to +1.6+1.6 AP over baseline) and superior instance segmentation performance (+1.3+1.3 AP on ResNet-101 over baseline and +0.4+0.4 AP over SE) while saving 2.5M to 4.7M parameters compared to SE.

  9. Knowl 9 — Hyperparameter Regularization and Architectural Integration of SimAM

    experimental setup

    SimAM is integrated as a plug-and-play module placed immediately after the second convolutional layer inside each residual or convolutional block of a ConvNet.

    The only hyperparameter in SimAM is the regularization coefficient λ\lambda in the closed-form energy equation. Cross-validation across values λ∈{10−1,10−2,10−3,10−4,10−5,10−6}\lambda \in \{10^{-1}, 10^{-2}, 10^{-3}, 10^{-4}, 10^{-5}, 10^{-6}\} demonstrates that:

    1. SimAM provides consistent accuracy gains across this entire broad range.
    2. λ=10−4\lambda = 10^{-4} provides the optimal trade-off between mean top-1 accuracy and variance on CIFAR datasets (searched on a 45k/5k split of CIFAR-100 using ResNet-20).
    3. λ=0.1\lambda = 0.1 provides the best trade-off between accuracy and robustness on ImageNet-1K (searched using ResNet-18 on half-resolution inputs across 50 epochs).

Coverage note — None was omitted; all contributed models, closed-form derivations, algorithmic implementations, and empirical results across ImageNet, CIFAR, and COCO are fully covered.

References

  1. 1.Aubry, M., Russell, B. C., and Sivic, J. Painting-to-3D Model Alignment via Discriminative Visual Elements. ACM Transactions on Graphics (ToG), 33(2):1–14, 2014.
  2. 2.Bengio, Y., Simard, P., and Frasconi, P. Learning Long-term Dependencies with Gradient Descent is Difficult. IEEE Transactions on Neural Networks, 5(2):157–166, 1994.
  3. 3.Cao, Y., Xu, J., Lin, S., Wei, F., and Hu, H. Global Context Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  4. 4.Carrasco, M. Visual Attention: The Past 25 Years. Vision Research, 51(13):1484–1525, 2011.
  5. 5.Chatfield, K., Simonyan, K., Vedaldi, A., and Zisserman, A. Return of the Devil in the Details: Delving Deep into Convolutional Nets. arXiv:1405.3531, 2014.
  6. 6.Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., Zhang, Z., Cheng, D., Zhu, C., Cheng, T., Zhao, Q., Li, B., Lu, X., Zhu, R., Wu, Y., Dai, J., Wang, J., Shi, J., Ouyang, W., Loy, C. C., and Lin, D. MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv:1906.07155, 2019.
  7. 7.Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1251–1258, 2017.
  8. 8.Chun, M. M., Golomb, J. D., and Turk-Browne, N. B. A Taxonomy of External and Internal Attention. Annual Review of Psychology, 62:73–101, 2011.
  9. 9.Dong, X. and Yang, Y. Searching for a Robust Neural Architecture in Four GPU Hours. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1761–1770, 2019.
  10. 10.Feichtenhofer, C. X3D: Expanding Architectures for Efficient Video Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 203–213, 2020.
  11. 11.Feichtenhofer, C., Pinz, A., and Zisserman, A. Convolutional Two-stream Network Fusion for Video Action Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1933–1941, 2016.
  12. 12.Glorot, X. and Bengio, Y. Understanding the Difficulty of Training Deep Feedforward Neural Networks. In Thirteenth International Conference on Artificial Intelligence and Statistics, pp. 249–256. JMLR Workshop and Conference Proceedings, 2010.
  13. 13.Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. Accurate, Large Minibatch SGD: Training Imagenet in 1 Hour. arXiv:1706.02677, 2017.
  14. 14.Guo, Z., Zhang, X., Mu, H., Heng, W., Liu, Z., Wei, Y., and Sun, J. Single Path One-Shot Neural Architecture Search with Uniform Sampling. In European Conference on Computer Vision, pp. 544–560. Springer, 2020.
  15. 15.Hariharan, B., Malik, J., and Ramanan, D. Discriminative Decorrelation for Clustering and Classification. In European Conference on Computer Vision, pp. 459–472. Springer, 2012.
  16. 16.He, K., Zhang, X., Ren, S., and Sun, J. Identity Mappings in Deep Residual Networks. In European Conference on Computer Vision, pp. 630–645. Springer, 2016a.
  17. 17.He, K., Zhang, X., Ren, S., and Sun, J. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016b.
  18. 18.He, K., Gkioxari, G., Dollár, P., and Girshick, R. Mask R-CNN. In IEEE International Conference on Computer Vision, pp. 2961–2969, 2017.
  19. 19.Hillyard, S. A., Vogel, E. K., and Luck, S. J. Sensory Gain Control (Amplification) as a Mechanism of Selective Attention: Electrophysiological and Neuroimaging evidence. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences, 353(1373): 1257–1270, 1998.
  20. 20.Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et al. Searching for MobileNetV3. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1314–1324, 2019.
  21. 21.Hu, J., Shen, L., Albanie, S., Sun, G., and Vedaldi, A. Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks. arXiv:1810.12348, 2018a.
  22. 22.Hu, J., Shen, L., and Sun, G. Squeeze-and-Excitation Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 7132–7141, 2018b.
  23. 23.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely Connected Convolutional Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700–4708, 2017.
  24. 24.Huang, G., Liu, S., Van der Maaten, L., and Weinberger, K. Q. CondenseNet: An Efficient Densenet using Learned Group Convolutions. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 2752–2761, 2018.
  25. 25.Krizhevsky, A., Hinton, G., et al. Learning Multiple Layers of Features from Tiny Images. 2009.
  26. 26.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems, 25: 1097–1105, 2012.
  27. 27.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based Learning Applied to Document Recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  28. 28.Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z. Deeply-Supervised Nets. In Artificial Intelligence and Statistics, pp. 562–570. PMLR, 2015.
  29. 29.Lee, H., Kim, H.-E., and Nam, H. SRM: A Style-based Recalibration Module for Convolutional Neural Networks. In IEEE International Conference on Computer Vision, pp. 1854–1862, 2019.
  30. 30.Li, X., Sun, W., and Wu, T. Attentive Normalization. arXiv:1908.01259, 2019a.
  31. 31.Li, X., Wang, W., Hu, X., and Yang, J. Selective Kernel Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 510–519, 2019b.
  32. 32.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision, pp. 740–755. Springer, 2014.
  33. 33.Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. Feature Pyramid Networks for Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117–2125, 2017.
  34. 34.Lin, X., Ma, L., Liu, W., and Chang, S.-F. Context-Gated Convolution. In European Conference on Computer Vision, pp. 701–718. Springer, 2020.
  35. 35.Liu, C., Zoph, B., Neumann, M., Shlens, J., Hua, W., Li, L.-J., Fei-Fei, L., Yuille, A., Huang, J., and Murphy, K. Progressive Neural Architecture Search. In European Conference on Computer Vision, pp. 19–34, 2018a.
  36. 36.Liu, C., Chen, L.-C., Schroff, F., Adam, H., Hua, W., Yuille, A. L., and Fei-Fei, L. Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 82–92, 2019.
  37. 37.Liu, H., Simonyan, K., and Yang, Y. DARTs: Differentiable Architecture Search. arXiv:1806.09055, 2018b.
  38. 38.Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A. C. SSD: Single Shot Multibox Detector. In European Conference on Computer Vision, pp. 21–37. Springer, 2016.
  39. 39.Ren, S., He, K., Girshick, R., and Sun, J. Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks. arXiv:1506.01497, 2015.
  40. 40.Reynolds, J. H. and Chelazzi, L. Attentional Modulation of Visual Processing. Annu. Rev. Neurosci., 27:611–647, 2004.
  41. 41.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3): 211–252, 2015.
  42. 42.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 4510–4520, 2018.
  43. 43.Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 618–626, 2017.
  44. 44.Simonyan, K. and Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556, 2014.
  45. 45.Srivastava, R. K., Greff, K., and Schmidhuber, J. Highway Networks. arXiv:1505.00387, 2015.
  46. 46.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. Going Deeper with Convolutions. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–9, 2015.
  47. 47.Tan, M. and Le, Q. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In International Conference on Machine Learning, pp. 6105–6114. PMLR, 2019.
  48. 48.Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., and Le, Q. V. MnasNet: Platform-aware Neural Architecture Search for Mobile. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 2820–2828, 2019.
  49. 49.Tan, M., Pang, R., and Le, Q. V. EfficientDet: Scalable and Efficient Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 10781–10790, 2020.
  50. 50.Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. Residual Attention Network for Image Classification. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 3156–3164, 2017.
  51. 51.Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. Temporal Segment Networks for Action Recognition in Videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740–2755, 2018a.
  52. 52.Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 11534–11542, 2020.
  53. 53.Wang, X., Girshick, R., Gupta, A., and He, K. Non-Local Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 7794–7803, 2018b.
  54. 54.Webb, B. S., Dhruv, N. T., Solomon, S. G., Tailby, C., and Lennie, P. Early and Late Mechanisms of Surround Suppression in Striate Cortex of Macaque. Journal of Neuroscience, 25(50):11666–11675, 2005.
  55. 55.Woo, S., Park, J., Lee, J.-Y., and So Kweon, I. CBAM: Convolutional Block Attention Module. In European Conference on Computer Vision, pp. 3–19, 2018.
  56. 56.Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., and Keutzer, K. FBNet: Hardware-aware Efficient ConvNet Design via Differentiable Neural Architecture Search. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 10734–10742, 2019.
  57. 57.Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K. Aggregated Residual Transformations for Deep Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 1492–1500, 2017.
  58. 58.Yang, Z., Zhu, L., Wu, Y., and Yang, Y. Gated Channel Transformation for Visual Recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 11794–11803, 2020.
  59. 59.Zagoruyko, S. and Komodakis, N. Wide Residual Networks. arXiv:1605.07146, 2016.
  60. 60.Zeiler, M. D. and Fergus, R. Visualizing and Understanding Convolutional Networks. In European Conference on Computer Vision, pp. 818–833. Springer, 2014.
  61. 61.Zhang, R. Making Convolutional Networks Shift-Invariant Again. In International Conference on Machine Learning, pp. 7324–7334. PMLR, 2019.
  62. 62.Zoph, B. and Le, Q. V. Neural Architecture Search with Reinforcement Learning. arXiv:1611.01578, 2016.

Citation

MLA
Yang, L., et al. “SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks”. International Conference on Machine Learning, vol. 139, 2021, pp. 11863–74, https://proceedings.mlr.press/v139/yang21o.html.
APA
Yang, L., Zhang, R.-Y., Li, L., & Xie, X. (2021). SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks. International Conference on Machine Learning, 139, 11863–11874. https://proceedings.mlr.press/v139/yang21o.html
Chicago
Yang, L., R.-Y. Zhang, L. Li, and X. Xie. 2021. “SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks”. International Conference on Machine Learning 139: 11863–74. https://proceedings.mlr.press/v139/yang21o.html.
Harvard
Yang, L. et al. (2021) “SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks”, International Conference on Machine Learning. PMLR, pp. 11863–11874. Available at: https://proceedings.mlr.press/v139/yang21o.html.
Vancouver
1. Yang L, Zhang R-Y, Li L, Xie X (2021) SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks. In: International Conference on Machine Learning. PMLR, pp 11863–11874

BibTeX

@InProceedings{pmlr-v139-yang21o,
  title = 	 {SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural Networks},
  author =       {Yang, Lingxiao and Zhang, Ru-Yuan and Li, Lida and Xie, Xiaohua},
  booktitle = 	 {Proceedings of the 38th International Conference on Machine Learning},
  pages = 	 {11863--11874},
  year = 	 {2021},
  editor = 	 {Meila, Marina and Zhang, Tong},
  volume = 	 {139},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {18--24 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v139/yang21o/yang21o.pdf},
  url = 	 {https://proceedings.mlr.press/v139/yang21o.html},
  abstract = 	 {In this paper, we propose a conceptually simple but very effective attention module for Convolutional Neural Networks (ConvNets). In contrast to existing channel-wise and spatial-wise attention modules, our module instead infers 3-D attention weights for the feature map in a layer without adding parameters to the original networks. Specifically, we base on some well-known neuroscience theories and propose to optimize an energy function to find the importance of each neuron. We further derive a fast closed-form solution for the energy function, and show that the solution can be implemented in less than ten lines of code. Another advantage of the module is that most of the operators are selected based on the solution to the defined energy function, avoiding too many efforts for structure tuning. Quantitative evaluations on various visual tasks demonstrate that the proposed module is flexible and effective to improve the representation ability of many ConvNets. Our code is available at Pytorch-SimAM.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/