Do Vision Transformers See Like Convolutional Neural Networks?

Maithra RaghuThomas UnterthinerSimon KornblithChiyuan ZhangAlexey Dosovitskiy

article2021NeurIPS1,515 citationsOutstanding Paper Award

Demonstrates how Vision Transformers build visual representations distinct from convolutional networks by integrating global context early and preserving spatial information through strong residual connections across layers.

Listen

Convolutional neural networks have served as the standard foundation for computer vision systems, relying on built-in spatial assumptions. Recently, Vision Transformers adapted from natural language processing have matched or exceeded convolutional networks on image classification. This development raises a critical question: are Vision Transformers solving visual tasks using the same internal mechanisms as convolutional networks, or are they forming fundamentally different visual representations?

The article aims to evaluate the differences in internal feature representations, information propagation, and spatial properties between Vision Transformers and standard convolutional neural networks, as well as to determine how training dataset scale influences these representations.

To perform this evaluation, the authors conducted an empirical study comparing several representative Vision Transformer models and ResNet convolutional models. The models were pretrained on massive datasets such as JFT-300M and standard benchmarks like ImageNet. The researchers used Centered Kernel Alignment, a statistical similarity metric that compares layer activations within and across networks, complemented by effective receptive field measurements, architectural interventions, and linear classification probes across multiple datasets.

The analysis reveals five primary findings. First, Vision Transformers exhibit a highly uniform representation structure across all layers, showing strong similarity between lower and higher layers, whereas convolutional networks process representations across distinctly separated stages. Second, self-attention allows early Vision Transformer layers to aggregate both local and global information simultaneously, while convolutional networks strictly process local information in lower layers. Third, skip connections are substantially more influential in Vision Transformers, where removing a skip connection causes an approximate four percent drop in performance and breaks representation continuity. Fourth, Vision Transformers that use a classification token preserve precise spatial location data into their final layers, whereas convolutional networks and pooled Transformer models disperse spatial information. Fifth, dataset scale is crucial for large Vision Transformers; models pretrained on massive datasets achieve roughly a 30 percent higher linear probe accuracy in intermediate layers compared to those trained on smaller datasets, which also fail to learn necessary local attention early on.

These findings demonstrate that Vision Transformers do not merely mimic convolutional networks; they employ distinct operational dynamics characterized by early global processing and intense shortcut feature reuse. For technical decision-makers, this indicates that Vision Transformers possess strong inherent potential for tasks requiring precise spatial localization, such as object detection. However, this architecture also incurs significant computational and data costs, as large-scale data is required to learn basic local features that convolutional networks encode by design.

Organizations evaluating these architectures should prioritize Vision Transformers when large pretraining datasets and computing resources are available, especially for multimodal or localization-heavy pipelines. Conversely, convolutional networks or hybrid architectures remain preferable in low-data regimes where hardcoded spatial assumptions prevent performance degradation. Future efforts should evaluate Vision Transformers on dense prediction benchmarks like object detection and semantic segmentation, while exploring hybrid models that combine convolutional efficiency with Transformer flexibility.

The findings are supported by consistent results across multiple models and probe datasets. However, readers should note that the primary metric, Centered Kernel Alignment, aggregates complex multidimensional feature representations into single scalar values. In addition, the largest model advantages depend heavily on access to proprietary, large-scale pretraining datasets.

arXiv: 2108.08810
  • Paper: A ConvNet for the 2020s, Zhuang Liu et al. (2022). Building on comparative insights between ViT and CNN representations, this work modernizes classical convolutional architectures by incorporating structural designs learned from Vision Transformers.
  • Paper: CoAtNet: Marrying Convolution and Attention for All Data Sizes, Zihang Dai et al. (2021). Leveraging the distinct representational strengths of CNNs and self-attention, this paper develops hybrid architectures that integrate depthwise convolutions in early stages with attention blocks in later stages.
  • Paper: CvT: Introducing Convolutions to Vision Transformers, Haiping Wu et al. (2021). This work applies findings on ViT's lack of local inductive biases by directly embedding convolutional operations into vision transformer tokenization and projection layers.
  • Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). This paper extends comparative architectural insights by co-designing pure convolutional networks to effectively utilize masked autoencoder pre-training paradigms originally developed for Vision Transformers.
  • Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). This work extends Vision Transformer representational principles beyond discriminative classification into generative modeling by replacing standard convolutional U-Nets with transformer backbones in diffusion architectures.
Cover for Do Vision Transformers See Like Convolutional Neural Networks?

Abstract

Convolutional neural networks (CNNs) have so far been the de-facto model for visual data. Recent work has shown that (Vision) Transformer models (ViT) can achieve comparable or even superior performance on image classification tasks. This raises a central question: how are Vision Transformers solving these tasks? Are they acting like convolutional networks, or learning entirely different visual representations? Analyzing the internal representation structure of ViTs and CNNs on image classification benchmarks, we find striking differences between the two architectures, such as ViT having more uniform representations across all layers. We explore how these differences arise, finding crucial roles played by self-attention, which enables early aggregation of global information, and ViT residual connections, which strongly propagate features from lower to higher layers. We study the ramifications for spatial localization, demonstrating ViTs successfully preserve input spatial information, with noticeable effects from different classification methods. Finally, we study the effect of (pretraining) dataset scale on intermediate features and transfer learning, and conclude with a discussion on connections to new architectures such as the MLP-Mixer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background and Experimental Setup
  • 4 Representation Structure of ViTs and Convolutional Networks
  • 5 Local and Global Information in Layer Representations
  • 6 Representation Propagation through Skip Connections
  • 7 Spatial Information and Localization
  • 8 Effects of Scale on Transfer Learning
  • 9 Discussion
  • References
  • A Additional details on Methods and the Experimental Setup
  • B Additional Representation Structure Results
  • C Additional Local/Global Information Results
  • D Localization
  • E Additional Representation Propagation Results
  • F Additional results on linear probes
  • G Effects of Scale on Transfer Learning
  • H Preliminary Results on MLP-Mixer

Knowls

  1. Knowl 1 — Uniform Representation Structure in Vision Transformers Versus Stage-Based Structure in ResNets

    empirical result

    Analysis of layer-to-layer representation similarity using Centered Kernel Alignment (CKA) reveals fundamental structural differences between Vision Transformers (ViTs, including ViT-B/16, ViT-L/16, and ViT-H/14) and Convolutional Neural Networks (ResNet-50, ResNet-152):

    • ViT Representation Structure: ViTs exhibit a uniform internal similarity structure across depth, resulting in a grid-like CKA heatmap with substantial similarity between early and late layers.
    • ResNet Representation Structure: ResNets exhibit distinct stage-wise transitions, where representations within the same residual stage are similar, but similarity between early and late stages is very low.
    • Cross-Model Mapping: In cross-model CKA comparisons, the lower half of ResNet layers aligns only with approximately the lowest 25% of ViT layers. The upper half of ResNet corresponds to approximately the middle third of ViT layers, while the highest ViT layers develop representations that are dissimilar to all ResNet layers.
  2. Knowl 2 — Early Layer Self-Attention Distance and Scale Dependency in Vision Transformers

    empirical result

    The spatial span of self-attention in Vision Transformers is quantified by computing the mean attention distance for each attention head, defined as the average pixel distance between the query patch position and attended positions weighted by the attention matrix across samples.

    • Mix of Local and Global Heads: When pretrained at scale (e.g., on JFT-300M), early ViT layers (encoder blocks 0 and 1) learn a mixture of local attention heads (small mean pixel distance) and global attention heads (large mean pixel distance). In deep layers, all attention heads attend purely globally. This contrasts with CNNs, whose lower layers are structurally hardcoded to have strictly local receptive fields.
    • Dependency on Dataset Scale: When large Vision Transformers (such as ViT-L/16 and ViT-H/14) are trained on smaller datasets (ImageNet-1k alone without large-scale pretraining), early attention layers fail to learn local attention heads, which directly correlates with degraded classification performance. Smaller models like ViT-B/32 are able to learn early local heads even when trained only on ImageNet.
  3. Knowl 3 — Phase Transition and Dominance of Skip Connections in Vision Transformers

    empirical result

    The propagation of representations through Vision Transformer layers is characterized by the norm ratio ∥zi∥∥f(zi)∥\frac{\|z_i\|}{\|f(z_i)\|}, where ziz_i denotes the representation entering layer ii from the identity/skip connection and f(zi)f(z_i) is the transformation produced by the long branch (multi-head self-attention or MLP block).

    • Skip Connection Strength: ViTs exhibit substantially larger norm ratios across all layers compared to ResNets, showing that identity paths carry a much larger proportion of the representation norm than in CNNs.
    • Phase Transition: ViT representations undergo an architectural phase transition across depth:
      1. In the first half of the model, the classification token (CLS\text{CLS}, token 0) is propagated primarily through the skip connection (high norm ratio), while spatial patch tokens receive significant updates from the long branch (lower norm ratio).
      2. In the second half of the model, this dynamic reverses: spatial patch token representations are propagated largely unchanged along the skip connections, while the CLS\text{CLS} token is heavily transformed by the long branch (predominantly via MLP sublayers).
    • Ablation Effect: Removing skip connections from an intermediate block partitions the network's representation similarity before and after the ablated block and incurs an accuracy drop of approximately 4%.
  4. Knowl 4 — Spatial Localization Preservation in Vision Transformers Versus ResNets

    empirical result

    Evaluating the CKA similarity between individual token representations at deeper layers and non-overlapping input image patches demonstrates that Vision Transformers maintain spatial localization through to the final layer:

    • ViT Token Localization: In standard ViTs trained with a classification (CLS\text{CLS}) token, patch tokens located in the image interior exhibit the highest similarity specifically to their corresponding input image patch in the penultimate and final blocks. Tokens corresponding to image borders show high similarity to their matching patch as well as other edge locations.
    • ResNet Localization: Spatial feature vectors in deep ResNet layers become spatially diffuse and exhibit comparable similarity across broad regions of the input, demonstrating significantly weaker spatial discriminability compared to ViT tokens.
  5. Knowl 5 — Impact of Classification Strategy on Spatial Token Specialization: CLS Token vs Global Average Pooling

    empirical result

    The choice of classification aggregation mechanism directly controls the spatial characteristics of learned representations in Vision Transformers:

    • CLS Token Models: In ViTs trained with a dedicated CLS\text{CLS} token, spatial patch tokens remain localized in deeper layers and do not aggregate global image features. As a result, linear probes trained on individual spatial tokens in later layers achieve low classification accuracy, while probes trained on the CLS\text{CLS} token achieve high accuracy.
    • Global Average Pooling (GAP) Models: In ViTs trained with GAP (averaging all final layer tokens to compute classification logits without a CLS\text{CLS} token), individual spatial tokens lose spatial localization in deep layers and aggregate global class information. Linear probes trained on single spatial tokens from late GAP layers achieve classification performance comparable to pooling all tokens together, mirroring the behavior observed in ResNet spatial feature maps.
  6. Knowl 6 — Centered Kernel Alignment for Neural Network Hidden Representations

    equation

    Let X∈Rm×p1X \in \mathbb{R}^{m \times p_1} and Y∈Rm×p2Y \in \mathbb{R}^{m \times p_2} denote the activation matrices of two neural network layers with p1p_1 and p2p_2 neurons, respectively, evaluated on the same mm input examples. Let K=XXT∈Rm×mK = XX^T \in \mathbb{R}^{m \times m} and L=YYT∈Rm×mL = YY^T \in \mathbb{R}^{m \times m} denote the corresponding Gram matrices.

    Centered Kernel Alignment (CKA) between the two layers is computed as:

    CKA(K,L)=HSIC(K,L)HSIC(K,K) HSIC(L,L)\text{CKA}(K, L) = \frac{\text{HSIC}(K, L)}{\sqrt{\text{HSIC}(K, K) \, \text{HSIC}(L, L)}}

    where HSIC\text{HSIC} is the Hilbert-Schmidt Independence Criterion. Given the centering matrix H=Im−1m11TH = I_m - \frac{1}{m}\mathbf{1}\mathbf{1}^T and the centered Gram matrices K′=HKHK' = HKH and L′=HLHL' = HLH, the empirical HSIC estimator is:

    HSIC(K,L)=vec(K′)⋅vec(L′)(m−1)2\text{HSIC}(K, L) = \frac{\text{vec}(K') \cdot \text{vec}(L')}{(m - 1)^2}

    CKA is invariant to orthogonal transformations of representations (including neuron permutation) and invariant to isotropic scaling. In practice, minibatch sampling (e.g., batch size 1024 across 10,240 examples averaged over 20 trials) provides an unbiased approximation of HSIC at scale.

  7. Knowl 7 — Effect of Pretraining Dataset Scale on Intermediate Representation Quality for Transfer Learning

    empirical result

    Pretraining Vision Transformers on large-scale datasets (JFT-300M) substantially enhances the quality and transferability of intermediate-layer representations compared to training on ImageNet-1k alone:

    • Linear Probe Transfer Gap: Linear probes evaluated on intermediate layer representations of JFT-300M pretrained ViT-L/16 and ViT-H/14 achieve up to a 30% absolute accuracy improvement on downstream ImageNet, CIFAR-10, and CIFAR-100 classification over the same architectures pretrained only on ImageNet.
    • Intermediate Feature Superiority Over CNNs: JFT-300M pretrained ViTs develop stronger intermediate representations than ResNets across middle and upper layers when evaluated on downstream classification tasks via linear probes.
    • Data Requirements by Layer Depth: Subsetting JFT-300M (from 3% to 100%) reveals that lower ViT layers achieve near-maximal CKA similarity to the fully trained model with as little as 3% to 10% of the dataset, whereas learning high-quality higher-layer representations strictly requires large dataset scale.
  8. Knowl 8 — Monotonic Relationship Between ViT Attention Head Locality and ResNet Lower-Layer Similarity

    empirical result

    When subsets of self-attention heads from the first encoder block of a Vision Transformer (ViT-L/16 or ViT-H/14) are grouped by their mean attention distance (from most local to most global), their CKA similarity to lower ResNet layers decreases monotonically as the attention distance increases.

    This indicates that:

    1. The representations learned by local attention heads in early ViT layers correspond closely to the local representations produced by convolutional layers in ResNets.
    2. Global attention heads in early ViT layers learn quantitatively distinct representations that are not present in early CNN layers.
  9. Knowl 9 — Effective Receptive Field Dynamics in Vision Transformers Versus ResNets

    empirical result

    Effective Receptive Fields (ERFs)—computed as the absolute gradient of the feature map center location with respect to input pixels across channels—exhibit distinct progression patterns in ViTs versus ResNets:

    • ResNets: Receptive fields begin highly localized in early blocks and expand gradually across stages throughout the depth of the network.
    • Vision Transformers: Lower layer ERFs are larger than corresponding ResNet layers, and midway through the network (e.g., around block 6 in a 12-block model), the ERF shifts rapidly from local to global, covering the entire input image.
    • Influence of Skip Connections: Post-residual ViT ERFs maintain a pronounced concentration at the central patch across all depths due to strong identity propagation. In contrast, pre-residual attention sublayer ERFs show broader global integration without central patch dominance.
  10. Knowl 10 — Linear Probing Protocol for Layer-Wise Neural Network Representations

    experimental setup

    Layer-wise linear probing evaluates intermediate representations using closed-form regularized least-squares regression:

    1. Target Mapping: Class labels for kk-shot training examples are mapped to target vectors Y∈{−1,1}NY \in \{-1, 1\}^N, where NN is the number of classes, and the linear probe weights are recovered via ridge regression in closed form.
    2. Vision Transformer Evaluation: Probes are trained either on individual token vectors or on the global average of token vectors extracted from the output of each transformer block (after self-attention, MLP, and residual additions).
    3. ResNet Evaluation and Dimension Matching: Feature maps are extracted at the output of each residual block. Because spatial resolution decreases and channel count increases with depth in ResNets, early feature maps are partitioned into non-overlapping patches and flattened to approximate the channel dimensionality of the final block before being pooled to match the final spatial resolution.

Coverage note — Preliminary representation heatmaps for MLP-Mixer architectures and fine-tuning transfer similarity heatmaps across domain-specific datasets (medical/satellite) were omitted as they reinforce the primary findings without introducing distinct core analytical methodologies.

References

  1. 1.G. Alain and Y. Bengio. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644, 2016.
  2. 2.I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le. Attention augmented convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3286–3295, 2019.
  3. 3.S. Bhojanapalli, A. Chakrabarti, D. Glasner, D. Li, T. Unterthiner, and A. Veit. Understanding robustness of transformers for image classification. arXiv preprint arXiv:2103.14586, 2021.
  4. 4.N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  5. 5.M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  6. 6.M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever. Generative pretraining from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR, 2020.
  7. 7.X. Chen, S. Xie, and K. He. An empirical study of training self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021.
  8. 8.A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. In ACL, 2018.
  9. 9.J.-B. Cordonnier, A. Loukas, and M. Jaggi. On the relationship between self-attention and convolutional layers. arXiv preprint arXiv:1911.03584, 2019.
  10. 10.C. Cortes, M. Mohri, and A. Rostamizadeh. Algorithms for learning kernels based on centered alignment. The Journal of Machine Learning Research, 13(1):795–828, 2012.
  11. 11.S. d'Ascoli, H. Touvron, M. Leavitt, A. Morcos, G. Biroli, and L. Sagun. Convit: Improving vision transformers with soft convolutional inductive biases. arXiv preprint arXiv:2103.10697, 2021.
  12. 12.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  13. 13.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  15. 15.A. Gretton, K. Fukumizu, C. H. Teo, L. Song, B. Schölkopf, A. J. Smola, et al. A kernel statistical test of independence. In Nips, volume 20, pages 585–592. Citeseer, 2007.
  16. 16.A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby. Big transfer (bit): General visual representation learning. arXiv preprint arXiv:1912.11370, 6(2):8, 2019.
  17. 17.S. Kornblith, M. Norouzi, H. Lee, and G. Hinton. Similarity of neural network representations revisited. In ICML, 2019.
  18. 18.S. Kornblith, H. Lee, T. Chen, and M. Norouzi. What’s in a loss function for image classification? arXiv preprint arXiv:2010.16402, 2020.
  19. 19.N. Kriegeskorte, M. Mur, and P. A. Bandettini. Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in systems neuroscience, 2:4, 2008.
  20. 20.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097–1105, 2012.
  21. 21.S. R. Kudugunta, A. Bapna, I. Caswell, N. Arivazhagan, and O. Firat. Investigating multilingual nmt representations at scale. arXiv preprint arXiv:1909.02197, 2019.
  22. 22.G. W. Lindsay. Convolutional neural networks as a model of the visual system: past, present, and future. Journal of cognitive neuroscience, pages 1–15, 2020.
  23. 23.W. Luo, Y. Li, R. Urtasun, and R. Zemel. Understanding the effective receptive field in deep convolutional neural networks. arXiv preprint arXiv:1701.04128, 2017.
  24. 24.N. Maheswaranathan, A. H. Williams, M. D. Golub, S. Ganguli, and D. Sussillo. Universality and individuality in neural dynamics across large populations of recurrent networks. Advances in neural information processing systems, 2019:15629, 2019.
  25. 25.A. Merchant, E. Rahimtoroghi, E. Pavlick, and I. Tenney. What happens to bert embeddings during fine-tuning? arXiv preprint arXiv:2004.14448, 2020.
  26. 26.A. S. Morcos, M. Raghu, and S. Bengio. Insights on representational similarity in neural networks with canonical correlation. arXiv preprint arXiv:1806.05759, 2018.
  27. 27.B. Mustafa, A. Loh, J. Freyberg, P. MacWilliams, M. Wilson, S. M. McKinney, M. Sieniek, J. Winkens, Y. Liu, P. Bui, et al. Supervised transfer learning at scale for medical imaging. arXiv preprint arXiv:2101.05913, 2021.
  28. 28.M. Naseer, K. Ranasinghe, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang. Intriguing properties of vision transformers, 2021.
  29. 29.T. Nguyen, M. Raghu, and S. Kornblith. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. arXiv preprint arXiv:2010.15327, 2020.
  30. 30.N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran. Image transformer. In International Conference on Machine Learning, pages 4055–4064. PMLR, 2018.
  31. 31.S. Paul and P.-Y. Chen. Vision transformers are robust learners. arXiv preprint arXiv:2105.07581, 2021.
  32. 32.M. E. Peters, M. Neumann, L. Zettlemoyer, and W.-t. Yih. Dissecting contextual word embeddings: Architecture and representation. In EMNLP, 2018.
  33. 33.A. Raghu, M. Raghu, S. Bengio, and O. Vinyals. Rapid learning or feature reuse? towards understanding the effectiveness of maml. arXiv preprint arXiv:1909.09157, 2019.
  34. 34.M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein. Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. arXiv preprint arXiv:1706.05806, 2017.
  35. 35.M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio. Transfusion: Understanding transfer learning for medical imaging. arXiv preprint arXiv:1902.07208, 2019.
  36. 36.P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens. Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909, 2019.
  37. 37.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  38. 38.J. Shi, E. Shea-Brown, and M. Buice. Comparison against task driven artificial neural networks reveals functional properties in mouse visual cortex. Advances in Neural Information Processing Systems, 32:5764–5774, 2019.
  39. 39.L. Song, A. Smola, A. Gretton, J. Bedo, and K. Borgwardt. Feature selection via dependence maximization. Journal of Machine Learning Research, 13(5), 2012.
  40. 40.C. Sun, A. Shrivastava, S. Singh, and A. Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision, pages 843–852, 2017.
  41. 41.Y. Tay, M. Dehghani, J. Gupta, D. Bahri, V. Aribandi, Z. Qin, and D. Metzler. Are pre-trained convolutions better than pre-trained transformers? arXiv preprint arXiv:2105.03322, 2021.
  42. 42.I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, D. Keysers, J. Uszkoreit, M. Lucic, et al. Mlp-mixer: An all-mlp architecture for vision. arXiv preprint arXiv:2105.01601, 2021.
  43. 43.H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, A. Joulin, G. Synnaeve, J. Verbeek, and H. Jégou. Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404, 2021.
  44. 44.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  45. 45.E. Voita, R. Sennrich, and I. Titov. The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives. In EMNLP, 2019.
  46. 46.B. Wu, C. Xu, X. Dai, A. Wan, P. Zhang, Z. Yan, M. Tomizuka, J. Gonzalez, K. Keutzer, and P. Vajda. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677, 2020.
  47. 47.J. M. Wu, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass. Similarity analysis of contextual word representation models. arXiv preprint arXiv:2005.01172, 2020.
  48. 48.S. Wu, A. Conneau, H. Li, L. Zettlemoyer, and V. Stoyanov. Emerging cross-lingual structure in pretrained language models. arXiv preprint arXiv:1911.01464, 2019.
  49. 49.L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z. Jiang, F. E. Tay, J. Feng, and S. Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021.
  50. 50.X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, et al. The visual task adaptation benchmark. 2019.

Citation

MLA
Raghu, M., et al. “Do Vision Transformers See Like Convolutional Neural Networks?”. Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 12116–28, https://proceedings.neurips.cc/paper_files/paper/2021/file/652cf38361a209088302ba2b8b7f51e0-Paper.pdf.
APA
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., & Dosovitskiy, A. (2021). Do Vision Transformers See Like Convolutional Neural Networks?. Advances in Neural Information Processing Systems, 34, 12116–12128. https://proceedings.neurips.cc/paper_files/paper/2021/file/652cf38361a209088302ba2b8b7f51e0-Paper.pdf
Chicago
Raghu, M., T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy. 2021. “Do Vision Transformers See Like Convolutional Neural Networks?”. Advances in Neural Information Processing Systems 34: 12116–28. https://proceedings.neurips.cc/paper_files/paper/2021/file/652cf38361a209088302ba2b8b7f51e0-Paper.pdf.
Harvard
Raghu, M. et al. (2021) “Do Vision Transformers See Like Convolutional Neural Networks?”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 12116–12128. Available at: https://proceedings.neurips.cc/paper_files/paper/2021/file/652cf38361a209088302ba2b8b7f51e0-Paper.pdf.
Vancouver
1. Raghu M, Unterthiner T, Kornblith S, Zhang C, Dosovitskiy A (2021) Do Vision Transformers See Like Convolutional Neural Networks?. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 12116–12128

BibTeX

@inproceedings{raghu2021vision,
  title = {Do Vision Transformers See Like Convolutional Neural Networks?},
  author = {Raghu, Maithra and Unterthiner, Thomas and Kornblith, Simon and Zhang, Chiyuan and Dosovitskiy, Alexey},
  year = {2021},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {34},
  pages = {12116-12128},
  url = {https://proceedings.neurips.cc/paper_files/paper/2021/file/652cf38361a209088302ba2b8b7f51e0-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission