CoAtNet: Marrying Convolution and Attention for All Data Sizes
Zihang DaiHanxiao LiuQuoc V. LeMingxing Tan
Presents CoAtNet, a hybrid architecture that integrates depthwise convolution with self-attention to achieve state-of-the-art image classification performance while matching massive Vision Transformers using 23 times less training data.
Deep learning architectures for computer vision face a fundamental trade-off: traditional Convolutional Neural Networks generalize well on limited data due to built-in spatial assumptions, whereas Transformer models offer massive capacity on huge datasets but struggle with data efficiency and slower training convergence. Prior attempts to combine both approaches have largely relied on ad-hoc designs without a clear systematic foundation.
The article designs and evaluates a hybrid network family, named CoAtNet, to combine depthwise convolution and self-attention into a unified, scalable vision architecture. The goal is to maximize both generalization on smaller datasets and learning capacity on web-scale datasets under constrained computational budgets.
The authors conducted extensive empirical evaluations across three standardized dataset tiers: ImageNet-1K (1.28 million images), ImageNet-21K (12.7 million images), and the massive proprietary JFT dataset (up to 3 billion images). They benchmarked multiple architectural layouts by training models with matching parameter scales and analyzing the gap between training loss and test accuracy, as well as downstream transfer performance across various image resolutions.
The findings show that depthwise convolution merges naturally into self-attention via relative position bias, significantly improving generalization with minimal extra computation. Second, an architecture layout placing convolutional blocks in early stages and Transformer blocks in later stages (a C-C-T-T structure) achieves the best trade-off between capacity, transferability, and hardware efficiency. Third, when pre-trained on ImageNet-21K, CoAtNet achieved an 88.56% top-1 accuracy on ImageNet-1K, matching the performance of a Vision Transformer model pre-trained on a 23 times larger dataset (JFT-300M). Finally, when scaled up with the JFT-3B dataset, CoAtNet reached a state-of-the-art 90.88% top-1 accuracy while requiring roughly 1.5 to 4 times less computation than competing large models.
These results demonstrate that organizations can achieve leading visual recognition accuracy with significantly reduced training time, dataset size requirements, and computational costs. Rather than replacing convolutions with pure attention mechanisms, strategically hybridizing both allows models to generalize faster on standard datasets while scaling efficiently to massive data volumes.
Organizations developing computer vision systems should adopt hybrid architectures—placing convolution in early layers and self-attention in deeper layers—especially when operating under budget or data constraints. Machine learning teams should also ensure data augmentations are introduced during pre-training rather than only during fine-tuning to avoid performance degradation from distribution shifts. Next steps should focus on extending this hybrid approach beyond image classification to dense prediction tasks such as object detection and semantic segmentation.
The primary limitation of this study is its exclusive evaluation on image classification benchmarks and dependence on specialized hardware accelerators (TPUs) for timing measurements. However, given the systematic ablation studies and consistent performance gains across diverse data scales, confidence in the architecture's core efficiency and generalization advantages is high.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It introduces the Vision Transformer (ViT) architecture and establishes how attention models scale on massive image datasets, which CoAtNet directly builds upon to address the lack of inductive biases.
- Paper: On the Relationship between Self-Attention and Convolutional Layers, Jean-Baptiste Cordonnier et al. (2020). It provides the foundational theoretical and empirical analysis demonstrating how multi-head self-attention with relative positional encodings can express and unify convolutional operations.
- Paper: EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks, Mingxing Tan et al. (2019). It establishes principled convolutional scaling and inverted bottleneck design principles that inspire CoAtNet's convolutional stages and stage-wise capacity allocations.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It details how training optimizations and distillation allow vision transformers to generalize effectively under limited data regimes, motivating CoAtNet's hybrid approach across data scales.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). It presents an early hybrid paradigm incorporating convolutional token embeddings and projections into vision transformers to enhance spatial inductive bias.
- Paper: Designing Network Design Spaces, Ilija Radosavovic et al. (2020). It outlines systematic network design space exploration and stage configurations that inform the vertical stacking decisions evaluated in CoAtNet.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). It investigates whether modernizing pure convolutional architectures with transformer-inspired design choices can match hybrid models like CoAtNet without explicit self-attention.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). It extends modern convolutional architectures to large-scale masked autoencoding regimes to compete with transformer and hybrid scaling behaviors.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It explores architectural techniques for scaling attention-based vision backbones to billions of parameters and extremely high resolutions, addressing challenges encountered in large-scale models like CoAtNet.
- Paper: CoCa: Contrastive Captioners are Image-Text Foundation Models, Jiahui Yu et al. (2022). It scales vision transformer backbones on multi-billion image datasets (JFT-3B) for unified multimodal foundation models, building on the scaling limits established by CoAtNet.
- Paper: MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Sachin Mehta et al. (2021). It adapts the principle of combining local convolutions with global transformer attention into lightweight architectures tailored specifically for mobile and edge deployment.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). It introduces parameter-efficient tuning methods for large-scale pre-trained visual backbones, avoiding the expensive full fine-tuning typical of billion-scale vision models.
