CoAtNet: Marrying Convolution and Attention for All Data Sizes

Zihang DaiHanxiao LiuQuoc V. LeMingxing Tan

article2021NeurIPS1,714 citationsOutstanding Paper Award

Presents CoAtNet, a hybrid architecture that integrates depthwise convolution with self-attention to achieve state-of-the-art image classification performance while matching massive Vision Transformers using 23 times less training data.

Listen

Deep learning architectures for computer vision face a fundamental trade-off: traditional Convolutional Neural Networks generalize well on limited data due to built-in spatial assumptions, whereas Transformer models offer massive capacity on huge datasets but struggle with data efficiency and slower training convergence. Prior attempts to combine both approaches have largely relied on ad-hoc designs without a clear systematic foundation.

The article designs and evaluates a hybrid network family, named CoAtNet, to combine depthwise convolution and self-attention into a unified, scalable vision architecture. The goal is to maximize both generalization on smaller datasets and learning capacity on web-scale datasets under constrained computational budgets.

The authors conducted extensive empirical evaluations across three standardized dataset tiers: ImageNet-1K (1.28 million images), ImageNet-21K (12.7 million images), and the massive proprietary JFT dataset (up to 3 billion images). They benchmarked multiple architectural layouts by training models with matching parameter scales and analyzing the gap between training loss and test accuracy, as well as downstream transfer performance across various image resolutions.

The findings show that depthwise convolution merges naturally into self-attention via relative position bias, significantly improving generalization with minimal extra computation. Second, an architecture layout placing convolutional blocks in early stages and Transformer blocks in later stages (a C-C-T-T structure) achieves the best trade-off between capacity, transferability, and hardware efficiency. Third, when pre-trained on ImageNet-21K, CoAtNet achieved an 88.56% top-1 accuracy on ImageNet-1K, matching the performance of a Vision Transformer model pre-trained on a 23 times larger dataset (JFT-300M). Finally, when scaled up with the JFT-3B dataset, CoAtNet reached a state-of-the-art 90.88% top-1 accuracy while requiring roughly 1.5 to 4 times less computation than competing large models.

These results demonstrate that organizations can achieve leading visual recognition accuracy with significantly reduced training time, dataset size requirements, and computational costs. Rather than replacing convolutions with pure attention mechanisms, strategically hybridizing both allows models to generalize faster on standard datasets while scaling efficiently to massive data volumes.

Organizations developing computer vision systems should adopt hybrid architectures—placing convolution in early layers and self-attention in deeper layers—especially when operating under budget or data constraints. Machine learning teams should also ensure data augmentations are introduced during pre-training rather than only during fine-tuning to avoid performance degradation from distribution shifts. Next steps should focus on extending this hybrid approach beyond image classification to dense prediction tasks such as object detection and semantic segmentation.

The primary limitation of this study is its exclusive evaluation on image classification benchmarks and dependence on specialized hardware accelerators (TPUs) for timing measurements. However, given the systematic ablation studies and consistent performance gains across diverse data scales, confidence in the architecture's core efficiency and generalization advantages is high.

arXiv: 2106.04803
  • Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). It investigates whether modernizing pure convolutional architectures with transformer-inspired design choices can match hybrid models like CoAtNet without explicit self-attention.
  • Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). It extends modern convolutional architectures to large-scale masked autoencoding regimes to compete with transformer and hybrid scaling behaviors.
  • Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It explores architectural techniques for scaling attention-based vision backbones to billions of parameters and extremely high resolutions, addressing challenges encountered in large-scale models like CoAtNet.
  • Paper: CoCa: Contrastive Captioners are Image-Text Foundation Models, Jiahui Yu et al. (2022). It scales vision transformer backbones on multi-billion image datasets (JFT-3B) for unified multimodal foundation models, building on the scaling limits established by CoAtNet.
  • Paper: MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Sachin Mehta et al. (2021). It adapts the principle of combining local convolutions with global transformer attention into lightweight architectures tailored specifically for mobile and edge deployment.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). It introduces parameter-efficient tuning methods for large-scale pre-trained visual backbones, avoiding the expensive full fine-tuning typical of billion-scale vision models.
Cover for CoAtNet: Marrying Convolution and Attention for All Data Sizes

Abstract

Transformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transformers tend to have larger model capacity, their generalization can be worse than convolutional networks due to the lack of the right inductive bias. To effectively combine the strengths from both architectures, we present CoAtNets(pronounced "coat" nets), a family of hybrid models built from two key insights: (1) depthwise Convolution and self-Attention can be naturally unified via simple relative attention; (2) vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency. Experiments show that our CoAtNets achieve state-of-the-art performance under different resource constraints across various datasets: Without extra data, CoAtNet achieves 86.0% ImageNet top-1 accuracy; When pre-trained with 13M images from ImageNet-21K, our CoAtNet achieves 88.56% top-1 accuracy, matching ViT-huge pre-trained with 300M images from JFT-300M while using 23x less data; Notably, when we further scale up CoAtNet with JFT-3B, it achieves 90.88% top-1 accuracy on ImageNet, establishing a new state-of-the-art result.

Table of Contents

  • 1 Introduction
  • 2 Model
  • 2.1 Merging Convolution and Self-Attention
  • 2.2 Vertical Layout Design
  • 3 Related Work
  • 4 Experiments
  • 4.1 Experiment Setting
  • 4.2 Main Results
  • 4.3 Ablation Studies
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Model Details
  • A.2 Hyper-Parameters
  • A.3 Complete Comparison

Knowls

  1. Knowl 1 — Unified Relative Self-Attention Formulation

    equation

    Depthwise convolution and self-attention both compute spatial transformations via per-dimension weighted sums over defined receptive fields. Depthwise convolution applies an input-independent static kernel over a local neighborhood L(i)\mathcal{L}(i):

    yi=∑j∈L(i)wi−j⊙xjy_i = \sum_{j \in \mathcal{L}(i)} w_{i-j} \odot x_j

    where xi,yi∈RDx_i, y_i \in \mathbb{R}^D are input and output representation vectors at spatial position ii, and wi−j∈RDw_{i-j} \in \mathbb{R}^D is a translation-equivariant weight kernel depending solely on the relative coordinate shift i−ji - j.

    Standard self-attention dynamically computes input-adaptive weights over the global receptive field G\mathcal{G} via dot-product similarity:

    yi=∑j∈Gexp⁡(xi⊤xj)∑k∈Gexp⁡(xi⊤xk)xjy_i = \sum_{j \in \mathcal{G}} \frac{\exp(x_i^\top x_j)}{\sum_{k \in \mathcal{G}} \exp(x_i^\top x_k)} x_j

    CoAtNet unifies depthwise convolution and self-attention into a single pre-normalization relative attention module:

    yi=∑j∈Gexp⁡(xi⊤xj+wi−j)∑k∈Gexp⁡(xi⊤xk+wi−k)xjy_i = \sum_{j \in \mathcal{G}} \frac{\exp\left(x_i^\top x_j + w_{i-j}\right)}{\sum_{k \in \mathcal{G}} \exp\left(x_i^\top x_k + w_{i-k}\right)} x_j

    where wi−j∈Rw_{i-j} \in \mathbb{R} is a static scalar parameter indexing the relative spatial displacement between position ii and position jj. This formulation combines the translation equivariance inductive bias of convolution with the input-adaptive weighting and global receptive field of self-attention. In multi-head self-attention, queries, keys, and values are linearly projected per head, and independent relative bias matrices are maintained for each head.

  2. Knowl 2 — 2D Relative Positional Bias Parameterization and Resolution Scaling

    model/method

    For a 2D spatial feature map of height HH and width WW, the displacement between any query coordinate (i,j)(i, j) and key coordinate (i′,j′)(i', j') spans −H<i−i′<H-H < i - i' < H vertically and −W<j−j′<W-W < j - j' < W horizontally. For each attention head, CoAtNet defines a trainable parameter matrix P∈R(2H−1)×(2W−1)P \in \mathbb{R}^{(2H - 1) \times (2W - 1)}. Under 1-based array indexing, the scalar relative bias for coordinate pair ((i,j),(i′,j′))((i, j), (i', j')) is:

    w(i,j)−(i′,j′)=Pi−i′+H, j−j′+Ww_{(i,j)-(i',j')} = P_{i - i' + H, \, j - j' + W}

    To compute this on Tensor Processing Units (TPUs), two successive einsum operations along the height and width axes index the H2W2H^2 W^2 relative bias elements with computational complexity O(HW(H+W))O(HW(H + W)), which is strictly subsumed by the O(H2W2D)O(H^2 W^2 D) cost of computing pairwise dot-product attention (where DD is the per-head feature dimension). On GPUs, indexing is implemented via memory gather operations. At inference time, the full (H2×W2)(H^2 \times W^2) pairwise bias matrix is precomputed and cached.

    When transferring or fine-tuning a model on a higher spatial resolution H′×W′H' \times W' (where H′>HH' > H or W′>WW' > W), the parameter tensor PP is resized from (2H−1)×(2W−1)(2H - 1) \times (2W - 1) to (2H′−1)×(2W′−1)(2H' - 1) \times (2W' - 1) using 2D bilinear interpolation.

  3. Knowl 3 — Multi-Stage Vertical Architecture Layout (C-C-T-T)

    model/method

    CoAtNet is organized into a 5-stage progressive downsampling hierarchy (S0,S1,S2,S3,S4S0, S1, S2, S3, S4), where spatial resolution is halved and feature channels are expanded at the beginning of each successive stage (S1S1 at 1/41/4, S2S2 at 1/81/8, S3S3 at 1/161/16, and S4S4 at 1/321/32 of input image resolution):

    • Stage S0S0: 2-layer standard convolutional stem operating on full resolution.
    • Stage S1S1: Inverted residual bottleneck convolutional blocks (MBConv) with Squeeze-and-Excitation (SE) modules.
    • Stage S2S2: MBConv blocks with SE (Convolutional stage).
    • Stage S3S3: Relative Transformer blocks (Self-Attention stage).
    • Stage S4S4: Relative Transformer blocks (Self-Attention stage).

    This vertical layout, denoted C-C-T-T (Convolution-Convolution-Transformer-Transformer), is selected based on generalization, capacity, and efficiency:

    1. Early stages (S1,S2S1, S2) process fine-grained spatial maps where local patterns dominate; convolutional MBConv blocks provide high inductive bias, avoid quadratic memory overhead, and speed up training convergence.
    2. Later stages (S3,S4S3, S4) process downsampled representations where global semantic relations dominate; global relative Transformer blocks provide high model capacity to capture long-range contextual dependencies.
  4. Knowl 4 — Pre-Activation and Independent Downsampling Block Design

    model/method

    CoAtNet applies a unified pre-activation structure across both MBConv convolutional blocks and relative Transformer blocks:

    x←x+Module(Norm(x))x \leftarrow x + \text{Module}(\text{Norm}(x))

    where Norm\text{Norm} denotes BatchNorm for MBConv modules and LayerNorm for Self-Attention and Feed-Forward Network (FFN) modules. Gaussian Error Linear Units (GELU) are used as activation functions throughout.

    At the boundary of each downsampling stage from S1S1 to S4S4, spatial resolution is reduced by a factor of 2 independently along both the residual transformation branch and the identity skip branch:

    • In a downsampling relative Transformer block:

    x←Proj(Pool(x))+Attention(Pool(Norm(x)))x \leftarrow \text{Proj}(\text{Pool}(x)) + \text{Attention}(\text{Pool}(\text{Norm}(x)))

    where Pool\text{Pool} denotes 2×22 \times 2 spatial max pooling with stride 2, and Proj\text{Proj} is a linear projection matching the enlarged channel dimension.

    • In a downsampling MBConv block:

    x←Proj(Pool(x))+Conv1×1(DepthConv3×3(Conv1×1(Norm(x),stride=2)))x \leftarrow \text{Proj}(\text{Pool}(x)) + \text{Conv}_{1 \times 1}(\text{DepthConv}_{3 \times 3}(\text{Conv}_{1 \times 1}(\text{Norm}(x), \text{stride}=2)))

    where the first 1×11 \times 1 convolution applies stride 2 downsampling directly on the normalized inputs prior to depthwise convolution and projection.

  5. Knowl 5 — Architectural Configurations of the CoAtNet Model Family

    data/table

    The CoAtNet architecture family scales across depths LL (number of stacked blocks per stage) and channel widths DD (hidden channel dimension). All MBConv blocks use kernel size 3, inverted bottleneck expansion rate E=4E=4, and squeeze-and-excitation shrink rate 0.250.25. All Transformer blocks use FFN expansion rate E=4E=4. Attention head size is 32 for CoAtNet-0 to CoAtNet-4, 64 for CoAtNet-5, and 128 for CoAtNet-6 and CoAtNet-7.

    Stage Spatial Size CoAtNet-0 CoAtNet-1 CoAtNet-2 CoAtNet-3
    S0-Conv 1/21/2 L=2,D=64L=2, D=64 L=2,D=64L=2, D=64 L=2,D=128L=2, D=128 L=2,D=192L=2, D=192
    S1-MBConv 1/41/4 L=2,D=96L=2, D=96 L=2,D=96L=2, D=96 L=2,D=128L=2, D=128 L=2,D=192L=2, D=192
    S2-MBConv 1/81/8 L=3,D=192L=3, D=192 L=6,D=192L=6, D=192 L=6,D=256L=6, D=256 L=6,D=384L=6, D=384
    S3-TFMRel 1/161/16 L=5,D=384L=5, D=384 L=14,D=384L=14, D=384 L=14,D=512L=14, D=512 L=14,D=768L=14, D=768
    S4-TFMRel 1/321/32 L=2,D=768L=2, D=768 L=2,D=768L=2, D=768 L=2,D=1024L=2, D=1024 L=2,D=1536L=2, D=1536

    Configurations for larger variants:

    • CoAtNet-4: S0 (L=2,D=192L=2, D=192), S1 (L=2,D=192L=2, D=192), S2 (L=12,D=384L=12, D=384), S3 (L=28,D=768L=28, D=768), S4 (L=2,D=1536L=2, D=1536).
    • CoAtNet-5: S0 (L=2,D=192L=2, D=192), S1 (L=2,D=256L=2, D=256), S2 (L=12,D=512L=12, D=512), S3 (L=28,D=1280L=28, D=1280), S4 (L=2,D=2048L=2, D=2048).
    • CoAtNet-6: S0 (L=2,D=192L=2, D=192), S1 (L=2,D=192L=2, D=192), S2 (L=4,D=384L=4, D=384), S3 (L=8L=8 MBConv blocks with D=768D=768 followed by L=42L=42 TFMRel blocks with D=1536D=1536), S4 (L=2,D=2048L=2, D=2048).
    • CoAtNet-7: S0 (L=2,D=192L=2, D=192), S1 (L=2,D=256L=2, D=256), S2 (L=4,D=512L=4, D=512), S3 (L=8L=8 MBConv blocks with D=1024D=1024 followed by L=42L=42 TFMRel blocks with D=2048D=2048), S4 (L=2,D=3072L=2, D=3072).
  6. Knowl 6 — ImageNet-1K Classification Performance Without External Pre-Training

    data/table

    When trained directly on ImageNet-1K (1.28M images) from scratch for 300 epochs, CoAtNet models achieve top-1 accuracies matching or exceeding state-of-the-art ConvNets (EfficientNetV2, NFNet) and outperforming pure Vision Transformer models (DeiT, Swin, CaiT) across matching parameter and FLOP budgets.

    Model Input Resolution Parameters FLOPs Top-1 Accuracy (%)
    DeiT-B 3842384^2 86M 55.4B 83.1
    Swin-B 3842384^2 88M 47.0B 84.2
    CaiT-S-36 3842384^2 68M 48.0B 85.0
    EfficientNetV2-L 4802480^2 121M 53.0B 85.7
    NFNet-F5 5442544^2 377M 289.8B 86.0
    CoAtNet-0 2242224^2 25M 4.2B 81.6
    CoAtNet-0 3842384^2 25M 13.4B 83.9
    CoAtNet-1 2242224^2 42M 8.4B 83.3
    CoAtNet-1 3842384^2 42M 27.4B 85.1
    CoAtNet-2 2242224^2 75M 15.7B 84.1
    CoAtNet-2 3842384^2 75M 49.8B 85.7
    CoAtNet-2 5122512^2 75M 96.7B 85.9
    CoAtNet-3 2242224^2 168M 34.7B 84.5
    CoAtNet-3 3842384^2 168M 107.4B 85.8
    CoAtNet-3 5122512^2 168M 203.1B 86.0

    CoAtNet-3 achieves 86.0% ImageNet-1K top-1 accuracy at 512×512512 \times 512 resolution with 168M parameters and 203.1B FLOPs, matching NFNet-F5 (377M parameters, 289.8B FLOPs).

  7. Knowl 7 — Transfer Learning Performance from ImageNet-21K Pre-Training

    data/table

    Models pre-trained on ImageNet-21K (12.7M images) for 90 or 150 epochs at 2242224^2 and fine-tuned for 30 epochs on ImageNet-1K show significant accuracy and data-efficiency gains over ConvNets and Transformers.

    Model Fine-tuning Resolution Parameters FLOPs Top-1 Accuracy (%)
    ViT-L/16 3842384^2 304M 190.7B 85.3
    Swin-L 3842384^2 197M 103.9B 86.4
    EfficientNetV2-L 4802480^2 121M 53.0B 86.8
    CvT-W24 3842384^2 277M 193.2B 87.7
    CoAtNet-2 3842384^2 75M 49.8B 87.1
    CoAtNet-2 5122512^2 75M 96.7B 87.3
    CoAtNet-3 3842384^2 168M 107.4B 87.6
    CoAtNet-3 5122512^2 168M 203.1B 87.9
    CoAtNet-4 3842384^2 275M 189.5B 87.9
    CoAtNet-4 (+ PT-RA) 3842384^2 275M 189.5B 88.3
    CoAtNet-4 (+ PT-RA-E150) 3842384^2 275M 189.5B 88.4
    CoAtNet-4 5122512^2 275M 360.9B 88.1
    CoAtNet-4 (+ PT-RA) 5122512^2 275M 360.9B 88.4
    CoAtNet-4 (+ PT-RA-E150) 5122512^2 275M 360.9B 88.56

    CoAtNet-4 (+ PT-RA-E150) achieves 88.56% top-1 accuracy on ImageNet-1K with ImageNet-21K pre-training, matching ViT-Huge (88.55%) which required pre-training on the 23×23\times larger JFT-300M dataset (300M images) with 2.3×2.3\times more parameters (632M) and 2.2×2.2\times more training steps. (PT-RA denotes RandAugment during pre-training; E150 denotes 150 pre-training epochs).

  8. Knowl 8 — Large-Scale Pre-Training on JFT-300M and JFT-3B

    data/table

    When pre-trained on large weakly-labeled image datasets (JFT-300M and JFT-3B) for 14 epochs and fine-tuned for 30 epochs on ImageNet-1K, CoAtNet models establish state-of-the-art accuracies with lower computational cost.

    Model Pre-training Data Parameters FLOPs TPUv3 Core-Days ImageNet Top-1 (%)
    ViT-L/16 JFT-300M 307M 364B 0.68K 87.76
    ViT-H/14 JFT-300M 632M 1021B 2.5K 88.55
    NFNet-F4+ JFT-300M 527M 367B 1.86K 89.20
    CoAtNet-3 JFT-300M 168M 214B 0.58K 88.81
    CoAtNet-4 JFT-300M 275M 361B 0.95K 89.11
    CoAtNet-5 JFT-300M 688M 812B 1.82K 89.77
    ViT-G/14 JFT-3B 1.84B 5160B >30K 90.45
    CoAtNet-6 JFT-3B 1.47B 1521B 6.6K 90.45
    CoAtNet-7 JFT-3B 2.44B 2586B 20.1K 90.88

    All models are evaluated at 512×512512 \times 512 resolution (except ViT-H/14 and ViT-G/14 at 518×518518 \times 518). On JFT-3B, CoAtNet-6 matches ViT-G/14 top-1 accuracy (90.45%) with 4.5×4.5\times fewer TPUv3 core-days (6.6K6.6\text{K} vs >30K>30\text{K}). CoAtNet-7 achieves 90.88% top-1 accuracy on ImageNet-1K using 1.5×1.5\times less compute than ViT-G/14.

  9. Knowl 9 — Generalization and Transfer Impact of Relative Self-Attention

    empirical result

    In controlled ablations comparing relative attention (adding scalar offset wi−jw_{i-j} prior to Softmax) with standard attention (relying solely on absolute positional encodings):

    • In ImageNet-1K training from scratch (evaluated with CoAtNet-2):
      • At 224×224224 \times 224: Relative attention achieves 84.1% top-1 accuracy vs 83.8% for standard attention (+0.3%+0.3\%).
      • At 384×384384 \times 384: Relative attention achieves 85.7% top-1 accuracy vs 85.3% for standard attention (+0.4%+0.4\%).
    • In ImageNet-21K pre-training and fine-tuning on ImageNet-1K (evaluated with CoAtNet-3):
      • Pre-training precision@1 on ImageNet-21K (2242224^2): 53.0% with relative attention vs 52.8% without relative attention (+0.2%+0.2\%).
      • Downstream fine-tuned accuracy on ImageNet-1K (3842384^2): 87.9% with relative attention vs 87.4% without relative attention (+0.5%+0.5\%).

    While pre-training precision on massive datasets is virtually identical between the two variants, relative attention delivers substantial accuracy gains during transfer fine-tuning and training from scratch, demonstrating that translation-equivariant relative attention acts primarily as a generalization enhancer rather than a capacity expander.

  10. Knowl 10 — Regularization Interaction Between Pre-Training and Fine-Tuning

    model/method

    Completely disabling data augmentation or regularization during large-scale pre-training (e.g., ImageNet-21K or JFT) and subsequently activating it during downstream ImageNet-1K fine-tuning degrades fine-tuning performance due to an induced data distribution shift.

    To ensure positive transfer:

    1. A mild degree of data augmentation and stochastic regularization—specifically RandAugment (magnitude parameters 2, 5) and stochastic depth (0.10.1 to 0.30.3)—is applied during upstream pre-training on ImageNet-21K and JFT.
    2. Although this mild regularization slightly lowers upstream pre-training precision, it prevents severe distribution shifts and allows the full deployment of regularization (RandAugment, MixUp, stochastic depth) during downstream fine-tuning.

    For CoAtNet-4 pre-trained on ImageNet-21K, adding mild pre-training RandAugment (+PT-RA) boosts fine-tuned ImageNet-1K top-1 accuracy from 87.9% to 88.3% at 384×384384 \times 384 and from 88.1% to 88.4% at 512×512512 \times 512.

Coverage note — Ablation details comparing BatchNorm versus LayerNorm execution throughput on TPU accelerators and minor intermediate stage block distributions (such as V1 and V2 in the stage allocation search) were omitted in favor of the primary model specifications and main empirical findings.

References

  1. 1.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, pages 1097–1105, 2012.
  2. 2.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  3. 3.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  4. 4.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  5. 5.Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. ICML, 2019.
  6. 6.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  8. 8.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  9. 9.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018.
  10. 10.Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. Attention augmented convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3286–3295, 2019.
  11. 11.Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. arXiv preprint arXiv:2101.11605, 2021.
  12. 12.Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. Efficient attention: Attention with linear complexities. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3531–3539, 2021.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  15. 15.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision, pages 843–852, 2017.
  16. 16.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  17. 17.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. arXiv preprint arXiv:2103.17239, 2021.
  18. 18.Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Qibin Hou, and Jiashi Feng. Deepvit: Towards deeper vision transformer. arXiv preprint arXiv:2103.11886, 2021.
  19. 19.Mingxing Tan and Quoc V Le. Efficientnetv2: Smaller models and faster training. ICML, 2021.
  20. 20.Andrew Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. arXiv preprint arXiv:2102.06171, 2021.
  21. 21.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. arXiv preprint arXiv:2103.12731, 2021.
  22. 22.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  23. 23.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. arXiv preprint arXiv:2103.15808, 2021.
  24. 24.Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. arXiv preprint arXiv:2104.01136, 2021.
  25. 25.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021.
  26. 26.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. arXiv preprint arXiv:2106.04560, 2021.
  27. 27.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018.
  28. 28.Laurent Sifre. Rigid-motion scattering for image classification. Ph.D. thesis section 6.2, 2014.
  29. 29.Mirgahney Mohamed, Gabriele Cesa, Taco S Cohen, and Max Welling. A data and compute efficient design for limited-resources deep learning. arXiv preprint arXiv:2004.09691, 2020.
  30. 30.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  32. 32.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020.
  33. 33.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  34. 34.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909, 2019.
  35. 35.Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2820–2828, 2019.
  36. 36.Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. A survey on visual transformer. arXiv preprint arXiv:2012.12556, 2020.
  37. 37.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. Transformers in vision: A survey. arXiv preprint arXiv:2101.01169, 2021.
  38. 38.Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. Music transformer. arXiv preprint arXiv:1809.04281, 2018.
  39. 39.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  40. 40.Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. Transformer dissection: A unified understanding of transformer’s attention via the lens of kernel. arXiv preprint arXiv:1908.11775, 2019.
  41. 41.Irwan Bello. Lambdanetworks: Modeling long-range interactions without attention. arXiv preprint arXiv:2102.08602, 2021.
  42. 42.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018.
  43. 43.Kun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou, Fengwei Yu, and Wei Wu. Incorporating convolution designs into visual transformers. arXiv preprint arXiv:2103.11816, 2021.
  44. 44.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021.
  45. 45.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  46. 46.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  47. 47.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pages 646–661. Springer, 2016.
  48. 48.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  49. 49.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  50. 50.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In European conference on computer vision, pages 630–645. Springer, 2016.
  51. 51.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  52. 52.Zihang Dai, Guokun Lai, Yiming Yang, and Quoc V Le. Funnel-transformer: Filtering out sequential redundancy for efficient language processing. arXiv preprint arXiv:2006.03236, 2020.

Citation

MLA
Dai, Z., et al. “CoAtNet: Marrying Convolution and Attention for All Data Sizes”. arXiv, 2021, http://arxiv.org/abs/2106.04803v2.
APA
Dai, Z., Liu, H., Le, Q. V., & Tan, M. (2021). CoAtNet: Marrying Convolution and Attention for All Data Sizes. arXiv. http://arxiv.org/abs/2106.04803v2
Chicago
Dai, Z., H. Liu, Q. V. Le, and M. Tan. 2021. “CoAtNet: Marrying Convolution and Attention for All Data Sizes”. arXiv. http://arxiv.org/abs/2106.04803v2.
Harvard
Dai, Z. et al. (2021) “CoAtNet: Marrying Convolution and Attention for All Data Sizes”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.04803v2.
Vancouver
1. Dai Z, Liu H, Le QV, Tan M (2021) CoAtNet: Marrying Convolution and Attention for All Data Sizes. arXiv

BibTeX

@article{dai2021coatnet,
  title = {CoAtNet: Marrying Convolution and Attention for All Data Sizes},
  author = {Dai, Zihang and Liu, Hanxiao and Le, Quoc V. and Tan, Mingxing},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.04803v2},
  eprint = {2106.04803}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission