Do Vision Transformers See Like Convolutional Neural Networks?
Maithra RaghuThomas UnterthinerSimon KornblithChiyuan ZhangAlexey Dosovitskiy
Demonstrates how Vision Transformers build visual representations distinct from convolutional networks by integrating global context early and preserving spatial information through strong residual connections across layers.
Convolutional neural networks have served as the standard foundation for computer vision systems, relying on built-in spatial assumptions. Recently, Vision Transformers adapted from natural language processing have matched or exceeded convolutional networks on image classification. This development raises a critical question: are Vision Transformers solving visual tasks using the same internal mechanisms as convolutional networks, or are they forming fundamentally different visual representations?
The article aims to evaluate the differences in internal feature representations, information propagation, and spatial properties between Vision Transformers and standard convolutional neural networks, as well as to determine how training dataset scale influences these representations.
To perform this evaluation, the authors conducted an empirical study comparing several representative Vision Transformer models and ResNet convolutional models. The models were pretrained on massive datasets such as JFT-300M and standard benchmarks like ImageNet. The researchers used Centered Kernel Alignment, a statistical similarity metric that compares layer activations within and across networks, complemented by effective receptive field measurements, architectural interventions, and linear classification probes across multiple datasets.
The analysis reveals five primary findings. First, Vision Transformers exhibit a highly uniform representation structure across all layers, showing strong similarity between lower and higher layers, whereas convolutional networks process representations across distinctly separated stages. Second, self-attention allows early Vision Transformer layers to aggregate both local and global information simultaneously, while convolutional networks strictly process local information in lower layers. Third, skip connections are substantially more influential in Vision Transformers, where removing a skip connection causes an approximate four percent drop in performance and breaks representation continuity. Fourth, Vision Transformers that use a classification token preserve precise spatial location data into their final layers, whereas convolutional networks and pooled Transformer models disperse spatial information. Fifth, dataset scale is crucial for large Vision Transformers; models pretrained on massive datasets achieve roughly a 30 percent higher linear probe accuracy in intermediate layers compared to those trained on smaller datasets, which also fail to learn necessary local attention early on.
These findings demonstrate that Vision Transformers do not merely mimic convolutional networks; they employ distinct operational dynamics characterized by early global processing and intense shortcut feature reuse. For technical decision-makers, this indicates that Vision Transformers possess strong inherent potential for tasks requiring precise spatial localization, such as object detection. However, this architecture also incurs significant computational and data costs, as large-scale data is required to learn basic local features that convolutional networks encode by design.
Organizations evaluating these architectures should prioritize Vision Transformers when large pretraining datasets and computing resources are available, especially for multimodal or localization-heavy pipelines. Conversely, convolutional networks or hybrid architectures remain preferable in low-data regimes where hardcoded spatial assumptions prevent performance degradation. Future efforts should evaluate Vision Transformers on dense prediction benchmarks like object detection and semantic segmentation, while exploring hybrid models that combine convolutional efficiency with Transformer flexibility.
The findings are supported by consistent results across multiple models and probe datasets. However, readers should note that the primary metric, Centered Kernel Alignment, aggregates complex multidimensional feature representations into single scalar values. In addition, the largest model advantages depend heavily on access to proprietary, large-scale pretraining datasets.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This foundational paper introduces the Vision Transformer (ViT) architecture whose internal representations and mechanisms are directly investigated and compared against CNNs in the source.
- Paper: On the Relationship between Self-Attention and Convolutional Layers, Jean-Baptiste Cordonnier et al. (2020). This study establishes the mathematical and empirical equivalence between self-attention and convolutional operations, providing theoretical context for understanding how ViTs and CNNs process spatial features.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). This work demonstrates how data-efficient Vision Transformers (DeiT) learn effectively on ImageNet via training strategies and distillation from CNN teachers, highlighting practical differences in how both architectures learn visual features.
- Paper: MLP-Mixer: An all-MLP Architecture for Vision, I. Tolstikhin et al. (2021). This paper presents the MLP-Mixer architecture, providing essential context for the source's comparative analysis of representation uniformity across attention, convolutional, and pure MLP vision backbones.
- Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). Building on comparative insights between ViT and CNN representations, this work modernizes classical convolutional architectures by incorporating structural designs learned from Vision Transformers.
- Paper: CoAtNet: Marrying Convolution and Attention for All Data Sizes, Zihang Dai et al. (2021). Leveraging the distinct representational strengths of CNNs and self-attention, this paper develops hybrid architectures that integrate depthwise convolutions in early stages with attention blocks in later stages.
- Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). This work applies findings on ViT's lack of local inductive biases by directly embedding convolutional operations into vision transformer tokenization and projection layers.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). This paper extends comparative architectural insights by co-designing pure convolutional networks to effectively utilize masked autoencoder pre-training paradigms originally developed for Vision Transformers.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). This work extends Vision Transformer representational principles beyond discriminative classification into generative modeling by replacing standard convolutional U-Nets with transformer backbones in diffusion architectures.
