Stand-Alone Self-Attention in Vision Models

Prajit RamachandranNiki ParmarAshish VaswaniIrwan BelloAnselm LevskayaJonathon Shlens

article2019NeurIPS1,386 citations

Demonstrates that replacing spatial convolutions entirely with stand-alone self-attention in ResNet models achieves superior ImageNet accuracy and competitive COCO detection performance with significantly fewer parameters and FLOPs.

Listen

Modern computer vision heavily relies on convolutional neural networks, which process visual information through fixed, local filters. While convolutions scale well computationally, they struggle to capture long-range contextual relationships across an image. Although recent research has added attention mechanisms—which weigh interactions between different elements based on their content—on top of convolutional models, attention has rarely been considered as a complete replacement for convolutions across an entire network.

The article evaluates whether local self-attention can serve as an effective, stand-alone building block for computer vision models. Specifically, it examines whether replacing standard spatial convolutions with content-based self-attention layers maintains or improves model accuracy while reducing computational overhead.

To test this approach, the researchers substituted spatial convolutions with local self-attention layers within standard vision architectures, primarily ResNet for image classification and RetinaNet for object detection. They evaluated these models on standard benchmarks, including the ImageNet classification dataset of over 1.2 million images and the COCO object detection benchmark, assessing performance across varying model depths, widths, and structural configurations.

The findings demonstrate that stand-alone attention is a viable and efficient primitive. On ImageNet, a fully attentional ResNet-50 model outperformed the baseline convolutional model by 0.5% top-1 accuracy while requiring 12% fewer floating point operations and 29% fewer parameters. On the COCO object detection task, an attention-based model matched baseline detection performance while utilizing 39% fewer floating point operations and 34% fewer parameters. In layer-by-layer analyses, the authors found that self-attention provides the greatest benefit in later network stages where high-level semantic integration occurs, whereas traditional convolutions remain advantageous in the initial stem layers for low-level feature extraction. Additionally, relative positional encodings proved essential, boosting accuracy by roughly 2% over absolute positional encodings.

These results show that computer vision systems can achieve superior representational efficiency with significantly fewer parameters and operations by using content-based interactions. For technical leaders and engineers, this presents an opportunity to deploy lighter-weight models with lower theoretical compute costs. However, current software and hardware accelerators lack optimized operations for these attention layers, meaning that practical execution time (wall-clock latency) is currently slower than established convolutional networks.

Organizations evaluating this approach should consider hybrid architectures as the immediate next step, combining convolutional initial layers with attention-driven later stages. Before migrating production workloads to pure attention-based models, teams must conduct hardware profiling to confirm that efficiency gains on paper translate into practical runtime savings. Further research should focus on hardware-optimized kernels and automated architecture searches designed natively around attention primitives.

arXiv: 1906.05909
  • Paper: Non-local Neural Networks, Xiaolong Wang et al. (2018). This foundational work introduced non-local self-attention blocks as plug-and-play augmentations for convolutional backbones, providing the exact architectural baseline that the source seeks to replace with pure, stand-alone self-attention.
  • Paper: Image Transformer, Niki Parmar et al. (2018). This paper establishes local 2D self-attention mechanisms over pixel neighborhoods for visual data, supplying key mathematical formulations adapted by the source for discriminative vision backbones.
Cover for Stand-Alone Self-Attention in Vision Models

Abstract

Convolutions are a fundamental building block of modern computer vision systems. Recent approaches have argued for going beyond convolutions in order to capture long-range dependencies. These efforts focus on augmenting convolutional models with content-based interactions, such as self-attention and non-local means, to achieve gains on a number of vision tasks. The natural question that arises is whether attention can be a stand-alone primitive for vision models instead of serving as just an augmentation on top of convolutions. In developing and testing a pure self-attention vision model, we verify that self-attention can indeed be an effective stand-alone layer. A simple procedure of replacing all instances of spatial convolutions with a form of self-attention applied to ResNet model produces a fully self-attentional model that outperforms the baseline on ImageNet classification with 12% fewer FLOPS and 29% fewer parameters. On COCO object detection, a pure self-attention model matches the mAP of a baseline RetinaNet while having 39% fewer FLOPS and 34% fewer parameters. Detailed ablation studies demonstrate that self-attention is especially impactful when used in later layers. These results establish that stand-alone self-attention is an important addition to the vision practitioner's toolbox.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Convolutions
  • 2.2 Self-Attention
  • 3 Fully Attentional Vision Models
  • 3.1 Replacing Spatial Convolutions
  • 3.2 Replacing the Convolutional Stem
  • 4 Experiments
  • 4.1 ImageNet Classification
  • 4.2 COCO Object Detection
  • 4.3 Where is stand-alone attention most useful?
  • 4.4 Which components are important in attention?
  • 4.4.1 Effect of spatial extent of self-attention
  • 4.4.2 Importance of positional information
  • 4.4.3 Importance of spatially-aware attention stem
  • 5 Discussion
  • References
  • A Appendix
  • A.1 Attention Stem
  • A.2 ImageNet Training Details
  • A.3 Object Detection Training Details

Knowls

  1. Knowl 1 — Local Spatial-Relative Self-Attention Mechanism for Images

    model/method

    Local spatial-relative self-attention replaces spatial convolutions by computing attention over a local neighborhood of pixels rather than globally across the entire image. Given an input feature map x∈Rh×w×dinx \in \mathbb{R}^{h \times w \times d_{\text{in}}} with height hh, width ww, and dind_{\text{in}} channels, for each pixel location ijij, a local memory block Nk(i,j)\mathcal{N}_k(i, j) of spatial extent kk centered around (i,j)(i, j) is extracted:

    Nk(i,j)={(a,b)∣∣a−i∣≤⌊k/2⌋,∣b−j∣≤⌊k/2⌋}\mathcal{N}_k(i, j) = \left\{ (a, b) \mid |a - i| \le \lfloor k/2 \rfloor, |b - j| \le \lfloor k/2 \rfloor \right\}

    Single-headed spatial-relative self-attention transforms the pixel at (i,j)(i, j) into a query qij=WQxij∈Rdoutq_{ij} = W_Q x_{ij} \in \mathbb{R}^{d_{\text{out}}} and all neighborhood pixels at (a,b)∈Nk(i,j)(a, b) \in \mathcal{N}_k(i, j) into keys kab=WKxab∈Rdoutk_{ab} = W_K x_{ab} \in \mathbb{R}^{d_{\text{out}}} and values vab=WVxab∈Rdoutv_{ab} = W_V x_{ab} \in \mathbb{R}^{d_{\text{out}}}, where WQ,WK,WV∈Rdout×dinW_Q, W_K, W_V \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}} are learned linear projections.

    Relative 2D spatial position is incorporated by factorizing the displacement from (i,j)(i, j) to (a,b)(a, b) into a row offset a−ia - i and a column offset b−jb - j. Learned row embeddings ra−i∈R12doutr_{a-i} \in \mathbb{R}^{\frac{1}{2}d_{\text{out}}} and column embeddings rb−j∈R12doutr_{b-j} \in \mathbb{R}^{\frac{1}{2}d_{\text{out}}} are concatenated to form the 2D relative position embedding ra−i,b−j=[ra−i;rb−j]∈Rdoutr_{a-i, b-j} = [r_{a-i}; r_{b-j}] \in \mathbb{R}^{d_{\text{out}}}. The output yij∈Rdouty_{ij} \in \mathbb{R}^{d_{\text{out}}} is computed as:

    yij=∑a,b∈Nk(i,j)softmaxab(qij⊤kab+qij⊤ra−i,b−j)vaby_{ij} = \sum_{a,b \in \mathcal{N}_k(i, j)} \text{softmax}_{ab}\left( q_{ij}^\top k_{ab} + q_{ij}^\top r_{a-i, b-j} \right) v_{ab}

    where softmaxab(⋅)\text{softmax}_{ab}(\cdot) denotes a softmax normalized over all (a,b)∈Nk(i,j)(a, b) \in \mathcal{N}_k(i, j). Incorporating relative position offsets preserves translation equivariance.

    For multi-head attention with NN heads, the input feature vector xijx_{ij} is partitioned depthwise into NN segments xijn∈Rdin/Nx_{ij}^n \in \mathbb{R}^{d_{\text{in}} / N}, each head applies the single-head attention operation using head-specific projection matrices WQn,WKn,WVn∈R(dout/N)×(din/N)W_Q^n, W_K^n, W_V^n \in \mathbb{R}^{(d_{\text{out}} / N) \times (d_{\text{in}} / N)}, and the resulting head outputs are concatenated along the channel dimension to produce yij∈Rdouty_{ij} \in \mathbb{R}^{d_{\text{out}}}.

  2. Knowl 2 — Spatially-Aware Attention Stem for Vision Networks

    model/method

    Standard self-attention struggles when applied directly to raw RGB input pixels at the stem of a vision network because raw pixel content is heavily correlated and lacks the abstract semantic features required for content-content matching. To allow the initial network layer to learn localized features such as edge detectors without full spatial convolutions, distance-based spatial awareness is injected directly into the value transformation of the stem attention layer.

    Over an input patch within a 4×44 \times 4 pooling window, the pointwise linear value transformation WVxabW_V x_{ab} is replaced by a spatially-varying mixture of MM learned linear transformations:

    v~ab=(∑m=1Mp(a,b,m)WVm)xab\tilde{v}_{ab} = \left( \sum_{m=1}^M p(a, b, m) W_V^m \right) x_{ab}

    where each WVm∈Rdout×dinW_V^m \in \mathbb{R}^{d_{\text{out}} \times d_{\text{in}}} is a value projection matrix, and the convex mixing coefficients p(a,b,m)p(a, b, m) are determined by the pixel's coordinates (a,b)(a, b) within the local window:

    p(a,b,m)=softmaxm((embrow(a)+embcol(b))⊤νm)p(a, b, m) = \text{softmax}_m\left( \left(\text{emb}_{\text{row}}(a) + \text{emb}_{\text{col}}(b)\right)^\top \nu^m \right)

    Here, embrow(a)\text{emb}_{\text{row}}(a) and embcol(b)\text{emb}_{\text{col}}(b) are learned embeddings aligned with the row and column coordinates inside the pooling window, νm\nu^m is a learned per-mixture embedding vector, and the softmax is evaluated across all mixture components m∈{1,…,M}m \in \{1, \dots, M\}. The mixture coefficients p(a,b,m)p(a, b, m) are shared across the attention heads.

    The complete attention stem executes this spatially-aware self-attention within non-overlapping 4×44 \times 4 spatial blocks of the input image, followed by batch normalization and a 4×44 \times 4 max pooling operation.

  3. Knowl 3 — Stand-Alone Self-Attention Architecture for ResNet and RetinaNet

    model/method

    Fully attentional convolutional architectures are constructed by substituting spatial convolutions (convolutions with kernel size k>1k > 1) with local spatial-relative self-attention layers while retaining 1×11 \times 1 pointwise convolutions and residual connectivity.

    In ResNet bottleneck blocks—which consist of a 1×11 \times 1 down-projection convolution, a 3×33 \times 3 spatial convolution, a 1×11 \times 1 up-projection convolution, and an identity skip connection—the 3×33 \times 3 spatial convolution is replaced by a local spatial-relative self-attention layer with spatial extent k=7k = 7 and N=8N = 8 attention heads. When spatial downsampling is required between layer groups, a 2×22 \times 2 average pooling operation with stride 2 is placed immediately after the attention layer.

    In the fully attentional ResNet variant, the standard 7×77 \times 7 convolution stem (stride 2 followed by 3×33 \times 3 max pooling) is replaced by the spatially-aware attention stem with 4×44 \times 4 self-attention blocks and 4×44 \times 4 max pooling.

    For RetinaNet object detection, the classification backbone is replaced by the attention-based ResNet, and all 3×33 \times 3 convolutions in the Feature Pyramid Network (FPN) and detection heads (which have output channels dout=256d_{\text{out}} = 256) are replaced with local self-attention layers (k=7,N=8k = 7, N = 8). Strided convolutions in the FPN are replaced with self-attention followed by 2×22 \times 2 average pooling with stride 2. In the shared classification and box regression subnetworks, an additional pointwise 1×11 \times 1 convolution is added at the end of each head to mix the attentional head representations.

  4. Knowl 4 — Parameter and Computational Scaling: Local Self-Attention vs Convolution

    theoretical result

    For an input with dind_{\text{in}} channels, an output with doutd_{\text{out}} channels, and a local spatial window of extent k×kk \times k:

    1. Parameter Count: A spatial convolution requires k2dindoutk^2 d_{\text{in}} d_{\text{out}} parameters, growing quadratically O(k2)\mathcal{O}(k^2) with the spatial window size kk. In contrast, local self-attention requires 3dindout3 d_{\text{in}} d_{\text{out}} parameters for the projection matrices WQ,WK,WVW_Q, W_K, W_V, plus kdoutk d_{\text{out}} parameters for the factorized 2D relative position embeddings ra−ir_{a-i} and rb−jr_{b-j}. As a result, the primary parameter count of self-attention is independent of the spatial extent kk.

    2. Computational Complexity: The FLOP cost of local self-attention grows substantially slower as a function of spatial extent kk than spatial convolutions under standard channel dimensions din,doutd_{\text{in}}, d_{\text{out}}. When din=dout=128d_{\text{in}} = d_{\text{out}} = 128, a standard spatial convolution with k=3k = 3 consumes the same total floating-point operations as a local self-attention layer with k=19k = 19.

  5. Knowl 5 — ImageNet Classification Performance Across Depths and Scaling

    data/table

    ImageNet classification performance of standard ResNet baselines compared against two attention configurations: Conv-stem + Attention (which retains the standard convolutional stem and uses local self-attention in all bottleneck blocks) and Full Attention (which uses the spatially-aware attention stem and local self-attention everywhere). Attention layers use spatial extent k=7k = 7 and N=8N = 8 heads. Models are evaluated at depths of 26, 38, and 50 layers (with 1 FLOP defined as 2 operations: multiply and add).

    ResNet-26 ResNet-38 ResNet-50
    Model FLOPS (B) Params (M) Acc. (%) FLOPS (B) Params (M) Acc. (%) FLOPS (B) Params (M) Acc. (%)
    Baseline 4.7 13.7 74.5 6.5 19.6 76.2 8.2 25.6 76.9
    Conv-stem + Attention 4.5 10.3 75.8 5.7 14.1 77.1 7.0 18.0 77.4
    Full Attention 4.7 10.3 74.8 6.0 14.1 76.9 7.2 18.0 77.6

    Full Attention ResNet-50 surpasses the convolutional ResNet-50 baseline by 0.7% top-1 accuracy (77.6% vs 76.9%) while requiring 12% fewer FLOPS (7.2B vs 8.2B) and 29% fewer parameters (18.0M vs 25.6M). Conv-stem + Attention outperforms the baseline across all tested network depths (26, 38, and 50) and across all scaled widths.

  6. Knowl 6 — COCO Object Detection Performance with Attentional RetinaNet

    data/table

    Object detection performance evaluated on the COCO dataset using RetinaNet with various combinations of convolutional and attention-based backbones, Feature Pyramid Networks (FPN), and detection heads. All attention layers use spatial extent k=7k = 7 and N=8N = 8 heads.

    Detection Heads + FPN Backbone FLOPS (B) Params (M) mAPcoco/50/75\text{mAP}_{\text{coco}/50/75} mAPs/m/l\text{mAP}_{s/m/l}
    Convolution Baseline 182 33.4 36.5 / 54.3 / 39.0 18.3 / 40.6 / 51.7
    Convolution Conv-stem + Attention 173 25.9 36.8 / 54.6 / 39.3 18.4 / 41.1 / 51.7
    Convolution Full Attention 173 25.9 36.2 / 54.0 / 38.7 17.5 / 40.3 / 51.7
    Attention Conv-stem + Attention 111 22.0 36.6 / 54.3 / 39.1 19.0 / 40.7 / 51.1
    Attention Full Attention 110 22.0 36.6 / 54.5 / 39.2 18.5 / 40.6 / 51.6

    Replacing only the backbone with Conv-stem + Attention improves COCO mAP from 36.5 to 36.8 while saving 9B FLOPS and 7.5M parameters. Replacing all spatial convolutions across the backbone, FPN, and detection heads with Full Attention achieves 36.6 mAP—matching the convolutional baseline while reducing computational cost by 39% (110B vs 182B FLOPS) and parameters by 34% (22.0M vs 33.4M).

  7. Knowl 7 — Layer Group Placement Analysis of Convolutions vs Attention

    data/table

    Ablation on ResNet-50 (with a convolutional stem) investigating the effect of assigning standard convolutions versus local spatial-relative self-attention (k=7k = 7) across the four sequential layer groups (Groups 1 to 4, delineated by spatial downsampling). Accuracies are reported on the ImageNet validation set.

    Convolution Groups Attention Groups FLOPS (B) Params (M) Top-1 Acc. (%)
    - 1, 2, 3, 4 7.0 18.0 80.2
    1 2, 3, 4 7.3 18.1 80.7
    1, 2 3, 4 7.5 18.5 80.7
    1, 2, 3 4 8.0 20.8 80.2
    1, 2, 3, 4 - 8.2 25.6 79.5
    2, 3, 4 1 7.9 25.5 79.7
    3, 4 1, 2 7.8 25.0 79.6
    4 1, 2, 3 7.2 22.7 79.9

    The highest classification accuracy (80.7%) is achieved when using convolutions in the early layers (Group 1 or Groups 1–2) and attention in the later layers (Groups 2–4 or Groups 3–4), outperforming all-convolution (79.5%) and all-attention (80.2%). Inverting this order—placing attention in early layers and convolutions in later layers—causes accuracy to drop to 79.6%–79.9% while significantly increasing parameters (22.7M–25.5M). Convolutions excel at extracting low-level local features, whereas attention layers are more effective at integrating high-level semantic context.

  8. Knowl 8 — Ablation of Positional Encodings and Attention Logit Interactions

    data/table

    Ablation experiments on ImageNet classification evaluating different positional encoding formulations and the decomposition of attention logits in a ResNet-50 model with a convolutional stem.

    Positional Encoding Type FLOPS (B) Params (M) Top-1 Acc. (%)
    None 6.9 18.0 77.6
    Absolute (Sinusoidal) 6.9 18.0 78.2
    Relative (2D Factorized) 7.0 18.0 80.2
    Attention Logit Type FLOPS (B) Params (M) Top-1 Acc. (%)
    q⊤rq^\top r (Relative position only) 6.1 16.7 76.9
    q⊤k+q⊤rq^\top k + q^\top r (Content + Relative position) 7.0 18.0 77.4

    2D factorized relative positional embeddings outperform absolute sinusoidal positional embeddings by 2.0% top-1 accuracy (80.2% vs 78.2%) and outperform no positional encoding by 2.6%. Furthermore, discarding content-content interactions (q⊤kq^\top k) entirely and relying solely on content-relative position interactions (q⊤rq^\top r) results in only a 0.5% drop in accuracy (76.9% vs 77.4%) while reducing parameters from 18.0M to 16.7M and FLOPS from 7.0B to 6.1B.

  9. Knowl 9 — Effect of Spatial Extent k on Self-Attention Performance

    data/table

    Ablation measuring the effect of varying the attention receptive field spatial extent k×kk \times k in a ResNet-50 architecture (using a convolutional stem) on ImageNet classification.

    Spatial Extent (k×kk \times k) FLOPS (B) Top-1 Acc. (%)
    3×33 \times 3 6.6 76.4
    5×55 \times 5 6.7 77.2
    7×77 \times 7 7.0 77.4
    9×99 \times 9 7.3 77.7
    11×1111 \times 11 7.7 77.6

    Because the parameter count in local self-attention is decoupled from the neighborhood size kk, total network parameter count remains identical across all spatial extent variations. Small attention windows (3×33 \times 3) yield poor accuracy (76.4%), while increasing kk steadily improves performance up to 9×99 \times 9 (77.7%), beyond which accuracy gains plateau.

  10. Knowl 10 — Ablation of Attention Stem Architectural Variants

    data/table

    Comparison of different architectural formulations for the initial stem layer of an otherwise fully attentional ResNet-50 evaluated on ImageNet classification.

    Attention Stem Type FLOPS (B) Top-1 Acc. (%)
    Stand-alone attention (Equation 3) 7.1 76.2
    Spatial convolution for values 7.4 77.2
    Spatially-aware values (Equations 8–9) 7.2 77.6

    Using standard local self-attention in the stem results in lower accuracy (76.2%). Introducing spatially-aware value transformations via position-dependent mixtures of point-wise projections improves accuracy by 1.4% to 77.6% at 7.2B FLOPS. Generating values using a spatial convolution achieves 77.2% top-1 accuracy but incurs higher computation (7.4B FLOPS).

  11. Knowl 11 — Wall-Clock Execution Time Bottleneck on Accelerator Hardware

    limitation

    Despite requiring fewer theoretical floating-point operations (FLOPS) and fewer parameters than convolutional baselines, stand-alone local self-attention models run slower in actual wall-clock execution time during training and inference on current hardware accelerators (such as GPUs and TPUs). This discrepancy arises because standard convolutional layers benefit from decades of low-level hardware optimizations and highly efficient GEMM implementations, whereas local sliding-window self-attention lacks purpose-built, highly optimized hardware kernel routines.

Coverage note — None omitted; all core architectural formulations, complexity analyses, ImageNet classification and COCO detection results, layer placement studies, component ablations, and hardware limitations contributed in the paper have been converted into knowls.

References

  1. 1.R. C. Gonzalez, R. E. Woods, et al., “Digital image processing [m],” Publishing house of electronics industry, vol. 141, no. 7, 2002.
  2. 2.K. Fukushima, “Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position,” Biological cybernetics, vol. 36, no. 4, pp. 193–202, 1980.
  3. 3.K. Fukushima, “Neocognitron: A hierarchical neural network capable of visual pattern recognition,” Neural networks, vol. 1, no. 2, pp. 119–130, 1988.
  4. 4.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural computation, vol. 1, no. 4, pp. 541–551, 1989.
  5. 5.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, 1998.
  6. 6.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2009.
  7. 7.J. Nickolls and W. J. Dally, “The gpu computing era,” IEEE micro, vol. 30, no. 2, pp. 56–69, 2010.
  8. 8.A. Krizhevsky, “Learning multiple layers of features from tiny images,” tech. rep., University of Toronto, 2009.
  9. 9.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing System, 2012.
  10. 10.Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature, vol. 521, no. 7553, p. 436, 2015.
  11. 11.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  12. 12.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception architecture for computer vision,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision, 2016.
  14. 14.S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  15. 15.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  16. 16.B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le, “Learning transferable architectures for scalable image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8697–8710, 2018.
  17. 17.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  18. 18.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision, pp. 2980–2988, 2017.
  19. 19.S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems, pp. 91–99, 2015.
  20. 20.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 4, pp. 834–848, 2018.
  21. 21.L.-C. Chen, M. Collins, Y. Zhu, G. Papandreou, B. Zoph, F. Schroff, H. Adam, and J. Shlens, “Searching for efficient multi-scale architectures for dense image prediction,” in Advances in Neural Information Processing Systems, pp. 8713–8724, 2018.
  22. 22.K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision, pp. 2961–2969, 2017.
  23. 23.E. P. Simoncelli and B. A. Olshausen, “Natural image statistics and neural representation,” Annual review of neuroscience, vol. 24, no. 1, pp. 1193–1216, 2001.
  24. 24.D. L. Ruderman and W. Bialek, “Statistics of natural images: Scaling in the woods,” in Advances in neural information processing systems, pp. 551–558, 1994.
  25. 25.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, pp. 5998–6008, 2017.
  26. 26.Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., “Google’s neural machine translation system: Bridging the gap between human and machine translation,” arXiv preprint arXiv:1609.08144, 2016.
  27. 27.J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in neural information processing systems, pp. 577–585, 2015.
  28. 28.W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4960–4964, IEEE, 2016.
  29. 29.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning, pp. 2048–2057, 2015.
  30. 30.J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  31. 31.M. Tan, B. Chen, R. Pang, V. Vasudevan, and Q. V. Le, “Mnasnet: Platform-aware neural architecture search for mobile,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  32. 32.X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7794–7803, 2018.
  33. 33.I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le, “Attention augmented convolutional networks,” CoRR, vol. abs/1904.09925, 2019.
  34. 34.J. Hu, L. Shen, S. Albanie, G. Sun, and A. Vedaldi, “Gather-excite: Exploiting feature context in convolutional neural networks,” in Advances in Neural Information Processing Systems, pp. 9423–9433, 2018.
  35. 35.H. Hu, Z. Zhang, Z. Xie, and S. Lin, “Local relation networks for image recognition,” arXiv preprint arXiv:1904.11491, 2019.
  36. 36.A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” arXiv preprint arXiv:1609.03499, 2016.
  37. 37.T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, “PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications,” arXiv preprint arXiv:1701.05517, 2017.
  38. 38.J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin, “Convolutional sequence to sequence learning,” CoRR, vol. abs/1705.03122, 2017.
  39. 39.L. Sifre and S. Mallat, “Rigid-motion scattering for image classification,” PhD thesis, Ph. D. thesis, vol. 1, p. 3, 2014.
  40. 40.S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Learning Representations, 2015.
  41. 41.F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  42. 42.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
  43. 43.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4510–4520, 2018.
  44. 44.S. Bartunov, A. Santoro, B. Richards, L. Marris, G. E. Hinton, and T. Lillicrap, “Assessing the scalability of biologically-motivated deep learning algorithms and architectures,” in Advances in Neural Information Processing Systems, pp. 9368–9378, 2018.
  45. 45.D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in International Conference on Learning Representations, 2015.
  46. 46.C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, C. Hawthorne, A. M. Dai, M. D. Hoffman, and D. Eck, “Music transformer,” in Advances in Neural Processing Systems, 2018.
  47. 47.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog, vol. 1, p. 8, 2019.
  48. 48.J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” CoRR, vol. abs/1810.04805, 2018.
  49. 49.N. Parmar, A. Vaswani, J. Uszkoreit, Ł. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International Conference on Machine Learning, 2018.
  50. 50.N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, R. Sepassi, and B. A. Hechtman, “Mesh-tensorflow: Deep learning for supercomputers,” CoRR, vol. abs/1811.02084, 2018.
  51. 51.P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” arXiv preprint arXiv:1803.02155, 2018.
  52. 52.A. Buades, B. Coll, and J.-M. Morel, “A non-local algorithm for image denoising,” in Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) - Volume 2 - Volume 02, CVPR ’05, (Washington, DC, USA), pp. 60–65, IEEE Computer Society, 2005.
  53. 53.Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “Aˆ 2-nets: Double attention networks,” in Advances in Neural Information Processing Systems, pp. 352–361, 2018.
  54. 54.B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations, 2017.
  55. 55.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li, “Imagenet large scale visual recognition challenge,” CoRR, vol. abs/1409.0575, 2014.
  56. 56.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision, pp. 740–755, Springer, 2014.
  57. 57.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117–2125, 2017.
  58. 58.T. S. Cohen, M. Geiger, J. Köhler, and M. Welling, “Spherical cnns,” arXiv preprint arXiv:1801.10130, 2018.
  59. 59.T. S. Cohen, M. Weiler, B. Kicanaoglu, and M. Welling, “Gauge equivariant convolutional networks and the icosahedral cnn,” arXiv preprint arXiv:1902.04615, 2019.
  60. 60.G. Ghiasi, T.-Y. Lin, R. Pang, and Q. V. Le, “Nas-fpn: Learning scalable feature pyramid architecture for object detection,” arXiv preprint arXiv:1904.07392, 2019.
  61. 61.F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli, “Pay less attention with lightweight and dynamic convolutions,” arXiv preprint arXiv:1901.10430, 2019.
  62. 62.X. Zhu, D. Cheng, Z. Zhang, S. Lin, and J. Dai, “An empirical study of spatial attention mechanisms in deep networks,” arXiv preprint arXiv:1904.05873, 2019.
  63. 63.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 4, pp. 834–848, 2017.
  64. 64.L.-C. Chen, A. Hermans, G. Papandreou, F. Schroff, P. Wang, and H. Adam, “Masklab: Instance segmentation by refining object detection with semantic and direction features,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4013–4022, 2018.
  65. 65.D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self-supervised interest point detection and description,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 224–236, 2018.
  66. 66.A. Toshev and C. Szegedy, “Deeppose: Human pose estimation via deep neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1653–1660, 2014.
  67. 67.A. Newell, K. Yang, and J. Deng, “Stacked hourglass networks for human pose estimation,” in European Conference on Computer Vision, pp. 483–499, Springer, 2016.
  68. 68.Y. E. NESTEROV, “A method for solving the convex programming problem with convergence rate o(1/k2 ),” Dokl. Akad. Nauk SSSR, vol. 269, pp. 543–547, 1983.
  69. 69.I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International Conference on Machine Learning, 2013.
  70. 70.I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” arXiv preprint arXiv:1608.03983, 2016.
  71. 71.N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” SIGARCH Comput. Archit. News, vol. 45, pp. 1–12, June 2017.
  72. 72.B. Polyak and A. Juditsky, “Acceleration of stochastic approximation by averaging,” SIAM Journal on Control and Optimization, vol. 30, no. 4, pp. 838–855, 1992.
  73. 73.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Computer Vision and Pattern Recognition (CVPR), 2015.

Citation

MLA
Ramachandran, P., et al. “Stand-Alone Self-Attention in Vision Models”. arXiv, 2019, http://arxiv.org/abs/1906.05909v1.
APA
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., & Shlens, J. (2019). Stand-Alone Self-Attention in Vision Models. arXiv. http://arxiv.org/abs/1906.05909v1
Chicago
Ramachandran, P., N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens. 2019. “Stand-Alone Self-Attention in Vision Models”. arXiv. http://arxiv.org/abs/1906.05909v1.
Harvard
Ramachandran, P. et al. (2019) “Stand-Alone Self-Attention in Vision Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1906.05909v1.
Vancouver
1. Ramachandran P, Parmar N, Vaswani A, Bello I, Levskaya A, Shlens J (2019) Stand-Alone Self-Attention in Vision Models. arXiv

BibTeX

@article{ramachandran2019stand,
  title = {Stand-Alone Self-Attention in Vision Models},
  author = {Ramachandran, Prajit and Parmar, Niki and Vaswani, Ashish and Bello, Irwan and Levskaya, Anselm and Shlens, Jonathon},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1906.05909v1},
  eprint = {1906.05909}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission