Going deeper with Image Transformers

Hugo TouvronMatthieu CordAlexandre SablayrollesGabriel SynnaeveHervé Jégou

article2021ICCV1,393 citations

Develops architectural modifications that prevent performance saturation in deep vision transformers, achieving state-of-the-art ImageNet accuracy with fewer parameters and no external training data.

Listen

Vision transformers have emerged as a powerful alternative to traditional convolutional networks for image classification tasks. However, training deeper transformer networks has historically suffered from optimization instability and early performance saturation when trained without massive external datasets. The article addresses these training bottlenecks to enable vision transformers to successfully scale with depth and achieve higher accuracy solely using standard image datasets.

The main objective of the article is to develop and evaluate architectural and optimization modifications that stabilize the training of deep vision transformers and improve image classification performance without relying on external data.

The authors conducted an extensive empirical study using standard image classification benchmarks, primarily ImageNet and several transfer learning datasets. They introduced two core modifications: a per-channel scaling technique called LayerScale to stabilize deep residual blocks, and a specialized architecture named Class-Attention in Image Transformers (CaiT) that separates image patch processing from final class token extraction. The experimental evaluations tested models of varying depths (ranging up to 48 layers) across standard resolutions and paired them with optimization techniques like distillation and stochastic depth.

The evaluation yielded several critical findings. First, LayerScale successfully stabilized deep vision transformers up to 48 layers, preventing optimization failure and enabling performance to scale with depth. Second, explicitly separating patch self-attention from class-attention layers resolved the conflicting roles of early class tokens, improving classification accuracy while reducing computational complexity. Third, the resulting CaiT models established state-of-the-art results on ImageNet without external training data, achieving 86.5% top-1 accuracy on standard validation, while also setting new records on the ImageNet-Real and ImageNet-V2 benchmarks. Finally, the proposed architecture demonstrated strong transfer learning capability, outperforming leading convolutional networks across several downstream datasets.

These findings imply that vision transformers can match or exceed top-tier convolutional networks without requiring specialized external pre-training datasets. For practitioners and decision-makers, this translates to improved classification performance and reduced computational and parameter overhead at higher accuracy regimes, offering a viable path for deploying efficient, deep transformer backbones.

Organizations developing computer vision systems should consider adopting LayerScale when training deep transformers to avoid optimization collapse. When designing classification pipelines, adopting dedicated class-attention stages can improve computational efficiency. For deployment scenarios with strict compute limitations at low model sizes, decision-makers should weigh the trade-offs, as traditional convolutional networks remain more efficient at smaller scales.

The findings are supported by high experimental confidence across multiple standard benchmarks. However, the study focuses predominantly on image classification and transfer learning. Readers should exercise caution before generalizing these results to other vision domains, such as dense object detection or segmentation, without further empirical validation.

arXiv: 2103.17239
  • Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It builds directly on the stability challenges of deep vision transformers by proposing post-normalization and scaled cosine attention to scale models up to billions of parameters.
  • Paper: Scaling Vision Transformers, Xiaohua Zhai et al. (2021). It advances the principles of transformer scaling and stabilization to train an unprecedented 22-billion-parameter Vision Transformer.
  • Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). It applies large-scale, deeply optimized vision transformer architectures to self-supervised feature learning across massive datasets without requiring fine-tuning.
  • Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). It addresses residual connection depth degradation and scaling limitations in deep vision transformers by introducing biomimetic aggregated attention mechanisms.
  • Paper: DINOv3, Oriane Siméoni et al. (2025). It extends multi-billion-parameter vision transformer optimization and stability techniques to create generalized foundation models across diverse downstream tasks.
Cover for Going deeper with Image Transformers

Abstract

Transformers have been recently adapted for large scale image classification, achieving high scores shaking up the long supremacy of convolutional neural networks. However the optimization of image transformers has been little studied so far. In this work, we build and optimize deeper transformer networks for image classification. In particular, we investigate the interplay of architecture and optimization of such dedicated transformers. We make two transformers architecture changes that significantly improve the accuracy of deep transformers. This leads us to produce models whose performance does not saturate early with more depth, for instance we obtain 86.5% top-1 accuracy on Imagenet when training with no external data, we thus attain the current SOTA with less FLOPs and parameters. Moreover, our best model establishes the new state of the art on Imagenet with Reassessed labels and Imagenet-V2 / match frequency, in the setting with no additional training data. We share our code and models.

Table of Contents

  • 1 Introduction
  • 2 Deeper image transformers with LayerScale
  • 3 Specializing layers for class attention
  • 4 Experiments
  • 4.1 Preliminary analysis with deeper architectures
  • 4.1.1 Adjusting the drop-rate of stochastic depth.
  • 4.1.2 Comparison of normalization strategies
  • 4.1.3 Analysis of Layerscale
  • 4.2 Class-attention layers
  • 4.3 Our CaiT models
  • 4.4 Results
  • 4.4.1 Performance/complexity of CaiT models
  • 4.4.2 Comparison with the state of the art on Imagenet
  • 4.4.3 Transfer learning
  • 4.5 Ablation
  • 4.5.1 Step by step from DeiT-Small to CaiT-S36
  • 4.5.2 Optimization of the number of heads
  • 4.5.3 Adaptation of the crop-ratio
  • 4.5.4 Longer training schedules
  • 5 Visualizations
  • 5.1 Attention map
  • 5.2 Illustration of saliency in class-attention
  • 6 Related work
  • 7 Conclusion
  • 8 Acknowledgments
  • References
  • A Variations on LayerScale init
  • B Design of the class-attention stage

Knowls

  1. Knowl 1 — LayerScale Residual Block Formulation

    model/method

    LayerScale is an architectural modification for deep vision transformers that stabilizes optimization when scaling depth. For a pre-layer-normalization transformer block containing a Self-Attention (SA) layer and a Feed-Forward Network (FFN), LayerScale introduces a learnable per-channel diagonal weighting matrix at the output of each residual sub-block:

    xl′=xl+diag(λl,1,…,λl,d)×SA(η(xl))x'_l = x_l + \text{diag}(\lambda_{l,1}, \dots, \lambda_{l,d}) \times \text{SA}(\eta(x_l))

    xl+1=xl′+diag(λl,1′,…,λl,d′)×FFN(η(xl′))x_{l+1} = x'_l + \text{diag}(\lambda'_{l,1}, \dots, \lambda'_{l,d}) \times \text{FFN}(\eta(x'_l))

    where η\eta is the LayerNorm operator, dd is the embedding dimension, and λl,i,λl,i′∈R\lambda_{l,i}, \lambda'_{l,i} \in \mathbb{R} are learnable scaling parameters initialized to a small positive constant ε\varepsilon.

    The initialization hyperparameter ε\varepsilon is set based on transformer depth:

    • ε=0.1\varepsilon = 0.1 for depth ≤18\le 18
    • ε=10−5\varepsilon = 10^{-5} for depth 2424
    • ε=10−6\varepsilon = 10^{-6} for depth ≥36\ge 36

    By initializing λ\lambda close to 0 while keeping pre-LayerNorm and learning rate warmup intact, LayerScale starts the network close to an identity mapping and allows the model to integrate layer updates progressively. Because the diagonal matrix is linear, it introduces no extra operations at inference time as it can be absorbed into the preceding projection weight matrices.

  2. Knowl 2 — Class-Attention in Image Transformers (CaiT) Architecture

    model/method

    Class-Attention in Image Transformers (CaiT) is a two-stage vision transformer architecture designed to decouple patch representation learning from class-level feature aggregation.

    In standard Vision Transformers (ViT), a learnable class token (CLS) is prepended to the image patch tokens at the input layer and processed through all self-attention layers, forcing attention weights to balance patch-to-patch spatial interactions and patch-to-class summary extraction simultaneously. CaiT replaces this with two distinct stages:

    1. Self-Attention (SA) Stage: A stack of NN standard transformer blocks operating exclusively on image patch vectors xpatchesx_{\text{patches}}, with no class token present.
    2. Class-Attention (CA) Stage: A stack of MM specialized blocks (typically M=2M=2) alternating Multi-Head Class-Attention (CA) and FFN layers. A learnable class token xclassx_{\text{class}} is introduced at this stage and acts as the query attending to the frozen patch representations output by the SA stage. Information flows strictly from patches to the class embedding; patch tokens are not updated in this stage.

    The final updated xclassx_{\text{class}} vector is fed directly to a linear classification head.

  3. Knowl 3 — Multi-Head Class-Attention Formulation and Complexity

    equation

    Let pp be the number of patch embeddings, dd the embedding dimension, and hh the number of attention heads. Given the patch embeddings xpatches∈Rp×dx_{\text{patches}} \in \mathbb{R}^{p \times d} and the class token xclass∈R1×dx_{\text{class}} \in \mathbb{R}^{1 \times d}, the sequence is augmented as z=[xclass,xpatches]∈R(p+1)×dz = [x_{\text{class}}, x_{\text{patches}}] \in \mathbb{R}^{(p+1) \times d}.

    With projection matrices Wq,Wk,Wv,Wo∈Rd×dW_q, W_k, W_v, W_o \in \mathbb{R}^{d \times d} and bias vectors bq,bk,bv,bo∈Rdb_q, b_k, b_v, b_o \in \mathbb{R}^d, the Multi-Head Class-Attention (CA) is defined by:

    Q=Wqxclass+bq∈R1×dQ = W_q x_{\text{class}} + b_q \in \mathbb{R}^{1 \times d}

    K=Wkz+bk∈R(p+1)×dK = W_k z + b_k \in \mathbb{R}^{(p+1) \times d}

    V=Wvz+bv∈R(p+1)×dV = W_v z + b_v \in \mathbb{R}^{(p+1) \times d}

    A=Softmax(QKTd/h)∈Rh×1×(p+1)A = \text{Softmax}\left(\frac{Q K^T}{\sqrt{d/h}}\right) \in \mathbb{R}^{h \times 1 \times (p+1)}

    outCA=Wo(AV)+bo∈R1×d\text{out}_{\text{CA}} = W_o (A V) + b_o \in \mathbb{R}^{1 \times d}

    The residual update modifies only the class embedding: xclass←xclass+diag(λ)×outCAx_{\text{class}} \leftarrow x_{\text{class}} + \text{diag}(\lambda) \times \text{out}_{\text{CA}}.

    Complexity: In standard Self-Attention, Q∈Rp×dQ \in \mathbb{R}^{p \times d} leads to an attention map QKT∈Rh×p×pQ K^T \in \mathbb{R}^{h \times p \times p} with O(p2d)O(p^2 d) computational and memory cost. In Class-Attention, Q∈R1×dQ \in \mathbb{R}^{1 \times d} generates QKT∈Rh×1×(p+1)Q K^T \in \mathbb{R}^{h \times 1 \times (p+1)}, reducing the complexity in the CA stage to O(pd)O(p d), which scales linearly with the number of patches pp.

  4. Knowl 4 — CaiT Model Family Architectural Configurations

    data/table

    CaiT models are parameterized by depth (number of self-attention blocks + number of class-attention blocks) and embedding dimension dd. The dimension dd is set such that the dimension per head is fixed at d/h=48d/h = 48, and each model employs LayerScale initialization ε\varepsilon and stochastic depth drop-rate drd_r.

    Model Depth (SA+CA) dd Heads hh Params FLOPs (224) Drop-rate drd_r
    XXS-24 24 + 2 192 4 12.0M 2.5B 0.05
    XXS-36 36 + 2 192 4 17.3M 3.8B 0.10
    XS-24 24 + 2 288 6 26.6M 5.4B 0.05
    XS-36 36 + 2 288 6 38.6M 8.1B 0.10
    S-24 24 + 2 384 8 46.9M 9.4B 0.10
    S-36 36 + 2 384 8 68.2M 13.9B 0.20
    S-48 48 + 2 384 8 89.5M 18.6B 0.30
    M-24 24 + 2 768 16 185.9M 36.0B 0.20
    M-36 36 + 2 768 16 270.9M 53.7B 0.30
    M-48 48 + 2 768 16 356.0M 70.3B 0.40

    For LayerScale initialization, models with 24 SA layers use ε=10−5\varepsilon = 10^{-5}, while models with 36 or 48 SA layers use ε=10−6\varepsilon = 10^{-6}. Models incorporate talking-heads attention and use an evaluation crop-ratio of 1.0.

  5. Knowl 5 — ImageNet Classification Performance and Comparison with State-of-the-Art

    data/table

    When trained on ImageNet-1k without external training data, deep CaiT models achieve state-of-the-art accuracy across ImageNet-1k validation, ImageNet-Real, and ImageNet-V2 (matched frequency).

    Network Params FLOPs Image size ImNet top-1 Real top-1 V2 top-1
    NFNet-F6+SAM 438M 377.3B 448 86.5% 89.9% 75.8%
    DeiT-B↑\uparrow384Υ\Upsilon 87M 55.5B 384 85.2% 89.3% 75.2%
    ViT-B/16 86M 55.4B 384 77.9% 83.6% -
    CaiT-S36 68M 13.9B 224 83.3% 88.0% 72.5%
    CaiT-S36↑\uparrow384 68M 48.0B 384 85.0% 89.2% 75.0%
    CaiT-S36Υ\Upsilon 68M 13.9B 224 84.0% 88.9% 74.1%
    CaiT-S36↑\uparrow384Υ\Upsilon 68M 48.0B 384 85.4% 89.8% 76.2%
    CaiT-M36↑\uparrow384Υ\Upsilon 271M 173.3B 384 86.1% 90.0% 76.3%
    CaiT-M36↑\uparrow448Υ\Upsilon 271M 247.8B 448 86.3% 90.2% 76.7%
    CaiT-M48↑\uparrow448Υ\Upsilon 356M 329.6B 448 86.5% 90.2% 76.9%

    Υ\Upsilon denotes hard distillation using a RegNetY-16GF teacher model, and ↑\uparrow indicates fine-tuning at the specified higher resolution. CaiT-M48↑\uparrow448Υ\Upsilon matches the 86.5% top-1 accuracy of NFNet-F6+SAM while using 19% fewer parameters (356M vs 438M) and 13% fewer FLOPs (329.6B vs 377.3B), while surpassing it on ImageNet-Real (90.2% vs 89.9%) and ImageNet-V2 (76.9% vs 75.8%).

  6. Knowl 6 — Comparison of Normalization and Residual Scaling Methods Across Depths

    data/table

    Training vision transformers at increasing depths using the standard DeiT optimization procedure fails above 18 layers without modifications. Evaluating different residual normalization and scaling methods on DeiT-Small (embedding dimension d=384d=384) across depths demonstrates that per-channel LayerScale provides superior stability and top-1 accuracy on ImageNet-1k.

    Depth Baseline (dr=0.05d_r=0.05) Baseline (adjusted drd_r) ReZero (adapted) T-Fixup (adapted) Fixup (adapted) Scalar α=ε\alpha = \varepsilon LayerScale
    12 79.9% 79.9% [dr=0.05d_r=0.05] 78.3% 79.4% 80.7% 80.4% 80.5%
    18 80.1% 80.7% [dr=0.10d_r=0.10] 80.1% 81.7% 82.0% 81.6% 81.7%
    24 78.9%†^\dagger 81.0% [dr=0.20d_r=0.20] 80.8% 81.5% 82.3% 81.1% 82.4%
    36 78.9%†^\dagger 81.9% [dr=0.25d_r=0.25] 81.6% 82.1% 82.4% 81.6% 82.9%

    †\dagger indicates training failure before completion. Unadapted ReZero, Fixup, and T-Fixup fail to converge on DeiT; reintroducing LayerNorm and warmup enables convergence. LayerScale outperforms single-scalar scaling (α=ε\alpha = \varepsilon) and Fixup variants at depths 24 and 36 while requiring only a single scalar initialization parameter ε\varepsilon rather than specialized per-layer initialization rules.

  7. Knowl 7 — Ablation on Class Token Insertion Depth and Class-Attention Design

    data/table

    Evaluating the position and mechanism for class token aggregation on DeiT-Small (d=384d=384, 12 total layers, without LayerScale) on ImageNet-1k reveals the impact of separating self-attention from class aggregation.

    Configuration Depth (SA+CA) Insertion Layer Top-1 Acc. Params FLOPs
    DeiT-S Baseline 12 + 0 0 79.9% 22M 4.6B
    Average Pooling 12 + 0 n/a 80.3% 22M 4.6B
    Late Insertion 12 + 0 2 80.0% 22M 4.6B
    Late Insertion 12 + 0 4 80.0% 22M 4.6B
    Late Insertion 12 + 0 8 80.0% 22M 4.6B
    Late Insertion 12 + 0 10 80.5% 22M 4.6B
    Late Insertion 12 + 0 11 80.3% 22M 4.6B
    CaiT Stage 9 + 3 9 79.6% 22M 3.6B
    CaiT Stage 10 + 2 10 80.3% 22M 4.0B
    CaiT Stage 11 + 1 11 80.6% 22M 4.3B
    CaiT Stage 12 + 1 12 80.8% 24M 4.7B
    CaiT Stage 12 + 2 12 80.8% 26M 4.7B
    CaiT Stage 12 + 3 12 80.6% 27M 4.8B

    Delaying class token insertion to layer 10 improves performance over insertion at layer 0 (80.5% vs 79.9%). Using specialized Class-Attention (CA) layers where patch tokens are frozen further reduces computation while improving accuracy: 10 SA + 2 CA matches average pooling (80.3%) at lower compute (4.0B vs 4.6B FLOPs), and 12 SA + 2 CA reaches 80.8% top-1 accuracy.

  8. Knowl 8 — LayerScale Effect on Residual Branch Activation Norms and Training Dynamics

    empirical result

    Analysis of residual activation norms reveals how LayerScale stabilizes deep transformer training:

    1. Uniform Branch Contribution: Measuring the ratio of the residual activation norm to the main branch norm, ∥gl(x)∥2/∥x∥2\|g_l(x)\|_2 / \|x\|_2, across all 36 layers of a transformer shows that LayerScale maintains this ratio at an average of ~20% uniformly across all layers. Without LayerScale, the contribution fluctuates heavily across layers and drops substantially in deeper layers.
    2. Dynamic Adaptation vs. Static Scaling: In a control experiment where DeiT-S architectures of various depths are trained using fixed scaling factors set to the final values learned by LayerScale, the models converge but underperform dynamic LayerScale training:
    Depth 12 18 24 36
    LayerScale (learnable) 80.5% 81.7% 82.4% 82.9%
    Re-trained with fixed weights 80.6% 81.5% 81.2% 81.6%

    This indicates that the dynamic evolution of the per-channel scaling parameters throughout gradient descent is necessary to unlock the benefits of greater depth.

  9. Knowl 9 — Transfer Learning Performance of CaiT

    data/table

    When fine-tuned at resolution 224 with a crop-ratio of 0.875 on downstream datasets, CaiT models demonstrate strong transfer learning capabilities, outperforming both convolutional and previous transformer baselines.

    Model CIFAR-10 CIFAR-100 Flowers-102 Stanford Cars iNat-18 iNat-19 FLOPs
    EfficientNet-B7 98.9% 91.7% 98.8% 94.7% - - 37.0B
    ViT-B/16 98.1% 87.1% 89.5% - - - 55.5B
    ViT-L/16 97.9% 86.4% 89.7% - - - 190.7B
    DeiT-B 224 99.1% 90.8% 98.4% 92.1% 73.2% 77.7% 17.5B
    CaiT-S-36 224 99.2% 92.2% 98.8% 93.5% 77.1% 80.6% 13.9B
    CaiT-M-36 224 99.3% 93.3% 99.0% 93.5% 76.9% 81.7% 53.7B
    CaiT-S-36 Υ\Upsilon 224 99.2% 92.2% 99.0% 94.1% 77.0% 81.4% 13.9B
    CaiT-M-36 Υ\Upsilon 224 99.4% 93.1% 99.1% 94.2% 78.0% 81.8% 53.7B

    Fine-tuning employs the AdamW optimizer with learning rates reduced by 10×10\times (for Cars, Flowers, iNaturalist) or 100×100\times (for CIFAR-10, CIFAR-100) and trained for 1000 epochs (CIFAR, Flowers, Cars) or 360 epochs (iNaturalist).

  10. Knowl 10 — Cumulative Ablation Path from DeiT-Small to CaiT-S36

    data/table

    Step-by-step ablation tracking the transformation of baseline DeiT-Small (d=384d=384, 300 epochs) into CaiT-S36 on ImageNet-1k:

    Improvement Step Top-1 Acc. Params FLOPs
    DeiT-S [d=384d=384, 300 epochs] 79.9% 22M 4.6B
    + More heads [8 heads, dim/head = 48] 80.0% 22M 4.6B
    + Talking-heads attention 80.5% 22M 4.6B
    + Depth [36 blocks] (training failed) 69.9%†^\dagger 64M 13.8B
    + LayerScale [init ε=10−6\varepsilon = 10^{-6}] 80.5% 64M 13.8B
    + Stochastic depth adaptation [dr=0.2d_r = 0.2] 83.0% 64M 13.8B
    + CaiT architecture [specialized class-attention] 83.2% 68M 13.9B
    + Longer training [400 epochs] 83.4% 68M 13.9B
    + Inference at higher resolution [256] 83.8% 68M 18.6B
    + Fine-tuning at higher resolution [384] 84.8% 68M 48.0B
    + Hard distillation [teacher: RegNetY-16GF] 85.2% 68M 48.0B
    + Adjust crop ratio [0.875 →\to 1.0] 85.4% 68M 48.0B

    Each ingredient provides complementary accuracy gains. Notably, increasing depth from 12 to 36 fails without LayerScale and stochastic depth tuning, but with both reaches 83.0%, while specialized class attention, distillation, and resolution scaling elevate final top-1 accuracy to 85.4%.

Coverage note — Qualitative attention map visualizations (Figures 6 and 7) and secondary appendix ablations testing distillation tokens within class attention and zero vs. uniform LayerScale initializations were omitted in favor of the primary quantitative and architectural contributions.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Thomas C. Bachlechner, Bodhisattwa Prasad Majumder, H. H. Mao, G. Cottrell, and Julian McAuley. Rezero is all you need: Fast convergence at large depth. arXiv preprint arXiv:2003.04887, 2020.
  3. 3.Irwan Bello. Lambdanetworks: Modeling long-range interactions without attention. In International Conference on Learning Representations, 2021.
  4. 4.Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le. Attention augmented convolutional networks. In Conference on Computer Vision and Pattern Recognition, 2019.
  5. 5.Maxim Berman, Hervé Jégou, Andrea Vedaldi, Iasonas Kokkinos, and Matthijs Douze. Multigrain: a unified image embedding for classes and instances. arXiv preprint arXiv:1902.05509, 2019.
  6. 6.Lucas Beyer, Olivier J. Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aaron van den Oord. Are we done with imagenet? arXiv preprint arXiv:2006.07159, 2020.
  7. 7.Andrew Brock, Soham De, and Samuel L Smith. Characterizing signal propagation to close the performance gap in unnormalized resnets. arXiv preprint arXiv:2101.08692, 2021.
  8. 8.A. Brock, Soham De, S. L. Smith, and K. Simonyan. High-performance large-scale image recognition without normalization. arXiv preprint arXiv:2102.06171, 2021.
  9. 9.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  10. 10.Antoni Buades, Bartomeu Coll, and J-M Morel. A non-local algorithm for image denoising. In Conference on Computer Vision and Pattern Recognition, 2005.
  11. 11.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, 2020.
  12. 12.Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic convolution: Attention over convolution kernels. In Conference on Computer Vision and Pattern Recognition, 2020.
  13. 13.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  14. 14.Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. Randaugment: Practical automated data augmentation with a reduced search space. arXiv preprint arXiv:1909.13719, 2019.
  15. 15.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Conference on Computer Vision and Pattern Recognition, 2017.
  16. 16.Soham De and Samuel L Smith. Batch normalization biases residual blocks towards the identity function in deep networks. arXiv e-prints, pages arXiv–2002, 2020.
  17. 17.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations, 2021.
  20. 20.Alaaeldin El-Nouby, Natalia Neverova, Ivan Laptev, and Hervé Jégou. Training vision transformers for image retrieval. arXiv preprint arXiv:2102.05644, 2021.
  21. 21.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019. ICLR 2020.
  22. 22.Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin. Training with quantization noise for extreme model compression. arXiv preprint arXiv:2004.07320, 2020.
  23. 23.Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635, 2018.
  24. 24.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In AISTATS, 2010.
  25. 25.Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. Non-autoregressive neural machine translation. arXiv preprint arXiv:1711.02281, 2017.
  26. 26.Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. arXiv preprint arXiv:2103.00112, 2021.
  27. 27.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition, June 2016.
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. arXiv preprint arXiv:1603.05027, 2016.
  29. 29.Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry. Augment your batch: Improving generalization through instance repetition. In Conference on Computer Vision and Pattern Recognition, 2020.
  30. 30.Grant Van Horn, Oisin Mac Aodha, Yang Song, Alexander Shepard, Hartwig Adam, Pietro Perona, and Serge J. Belongie. The inaturalist challenge 2018 dataset. arXiv preprint arXiv:1707.06642, 2018.
  31. 31.Grant Van Horn, Oisin Mac Aodha, Yang Song, Alexander Shepard, Hartwig Adam, Pietro Perona, and Serge J. Belongie. The inaturalist challenge 2019 dataset. arXiv preprint arXiv:1707.06642, 2019.
  32. 32.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. arXiv preprint arXiv:1709.01507, 2017.
  33. 33.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. Deep networks with stochastic depth. In European Conference on Computer Vision, 2016.
  34. 34.Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs. Improving transformer optimization through better initialization. In International Conference on Machine Learning, pages 4475–4483. PMLR, 2020.
  35. 35.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning, 2015.
  36. 36.Shigeki Karita, Nanxin Chen, Tomoki Hayashi, et al. A comparative study on transformer vs rnn in speech applications. arXiv preprint arXiv:1909.06317, 2019.
  37. 37.Diederik P Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. arXiv preprint arXiv:1807.03039, 2018.
  38. 38.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), 2013.
  39. 39.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, CIFAR, 2009.
  40. 40.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  41. 41.Jason Lee, Elman Mansimov, and Kyunghyun Cho. Deterministic non-autoregressive neural sequence modeling by iterative refinement. arXiv preprint arXiv:1802.06901, 2018.
  42. 42.Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. Selective kernel networks. Conference on Computer Vision and Pattern Recognition, 2019.
  43. 43.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  44. 44.I. Loshchilov and F. Hutter. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 2017.
  45. 45.Christoph Lüscher, Eugen Beck, Kazuki Irie, et al. Rwth asr systems for librispeech: Hybrid vs attention. Interspeech 2019, Sep 2019.
  46. 46.M-E. Nilsback and A. Zisserman. Automated flower classification over a large number of classes. In Proceedings of the Indian Conference on Computer Vision, Graphics and Image Processing, 2008.
  47. 47.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In International Conference on Machine Learning, pages 4055–4064. PMLR, 2018.
  48. 48.H. Pham, Qizhe Xie, Zihang Dai, and Quoc V. Le. Meta pseudo labels. arXiv preprint arXiv:2003.10580, 2020.
  49. 49.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners.
  50. 50.Ilija Radosavovic, Raj Prateek Kosaraju, Ross B. Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. Conference on Computer Vision and Pattern Recognition, 2020.
  51. 51.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, I. Bello, Anselm Levskaya, and Jonathon Shlens. Stand-alone self-attention in vision models. In Neurips, 2019.
  52. 52.B. Recht, Rebecca Roelofs, L. Schmidt, and V. Shankar. Do imagenet classifiers generalize to imagenet? arXiv preprint arXiv:1902.10811, 2019.
  53. 53.A. Romero, Nicolas Ballas, S. Kahou, Antoine Chassang, C. Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2015.
  54. 54.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. International journal of Computer Vision, 2015.
  55. 55.Noam Shazeer, Zhenzhong Lan, Youlong Cheng, N. Ding, and L. Hou. Talking-heads attention. arXiv preprint arXiv:2003.02436, 2020.
  56. 56.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015.
  57. 57.A. Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, P. Abbeel, and Ashish Vaswani. Bottleneck transformers for visual recognition. arXiv preprint arXiv:2101.11605, 2021.
  58. 58.R. Srivastava, Klaus Greff, and J. Schmidhuber. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  59. 59.R. Srivastava, Klaus Greff, and J. Schmidhuber. Training very deep networks. In NIPS, 2015.
  60. 60.Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert. End-to-end asr: from supervised to semi-supervised learning with modern architectures. arXiv preprint arXiv:1911.08460, 2019.
  61. 61.C. Szegedy, Wei Liu, Yangqing Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Conference on Computer Vision and Pattern Recognition, 2015.
  62. 62.Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946, 2019.
  63. 63.Hugo Touvron, M. Cord, M. Douze, F. Massa, Alexandre Sablayrolles, and H. Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020.
  64. 64.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herve Jegou. Fixing the train-test resolution discrepancy. Neurips, 2019.
  65. 65.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou. Fixing the train-test resolution discrepancy: Fixefficientnet. arXiv preprint arXiv:2003.08237, 2020.
  66. 66.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  67. 67.X. Wang, Ross B. Girshick, A. Gupta, and Kaiming He. Non-local neural networks. Conference on Computer Vision and Pattern Recognition, 2018.
  68. 68.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  69. 69.Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, and Peter Vajda. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677, 2020.
  70. 70.L. Xiao, Y. Bahri, Jascha Sohl-Dickstein, S. Schoenholz, and Jeffrey Pennington. Dynamical isometry and a mean field theory of cnns: How to train 10, 000-layer vanilla convolutional neural networks. arXiv preprint arXiv:1806.05393, 2018.
  71. 71.Cihang Xie, Mingxing Tan, Boqing Gong, Jiang Wang, A. Yuille, and Quoc V. Le. Adversarial examples improve image recognition. Conference on Computer Vision and Pattern Recognition, 2020.
  72. 72.Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. Conference on Computer Vision and Pattern Recognition, 2017.
  73. 73.Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. Self-training and pre-training are complementary for speech recognition. arXiv preprint arXiv:2010.11430, 2020.
  74. 74.L. Yuan, Y. Chen, Tao Wang, Weihao Yu, Yujun Shi, F. Tay, Jiashi Feng, and S. Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021.
  75. 75.Hongyi Zhang, Yann Dauphin, and Tengyu Ma. Fixup initialization: Residual learning without normalization. arXiv preprint arXiv:1901.09321, 2019.
  76. 76.Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi-Li Zhang, Haibin Lin, Yu e Sun, Tong He, Jonas Mueller, R. Manmatha, M. Li, and Alex Smola. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955, 2020.
  77. 77.Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V Le, and Yonghui Wu. Pushing the limits of semi-supervised learning for automatic speech recognition. arXiv preprint arXiv:2010.10504, 2020.
  78. 78.Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. Exploring self-attention for image recognition. In Conference on Computer Vision and Pattern Recognition, 2020.

Citation

MLA
Touvron, H., et al. “Going Deeper with Image Transformers”. arXiv, 2021, http://arxiv.org/abs/2103.17239v2.
APA
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., & Jégou, H. (2021). Going deeper with Image Transformers. arXiv. http://arxiv.org/abs/2103.17239v2
Chicago
Touvron, H., M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou. 2021. “Going Deeper with Image Transformers”. arXiv. http://arxiv.org/abs/2103.17239v2.
Harvard
Touvron, H. et al. (2021) “Going deeper with Image Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2103.17239v2.
Vancouver
1. Touvron H, Cord M, Sablayrolles A, Synnaeve G, Jégou H (2021) Going deeper with Image Transformers. arXiv

BibTeX

@article{touvron2021going,
  title = {Going deeper with Image Transformers},
  author = {Touvron, Hugo and Cord, Matthieu and Sablayrolles, Alexandre and Synnaeve, Gabriel and Jégou, Hervé},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2103.17239v2},
  eprint = {2103.17239}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE