Going deeper with Image Transformers
Hugo TouvronMatthieu CordAlexandre SablayrollesGabriel SynnaeveHervé Jégou
Develops architectural modifications that prevent performance saturation in deep vision transformers, achieving state-of-the-art ImageNet accuracy with fewer parameters and no external training data.
Vision transformers have emerged as a powerful alternative to traditional convolutional networks for image classification tasks. However, training deeper transformer networks has historically suffered from optimization instability and early performance saturation when trained without massive external datasets. The article addresses these training bottlenecks to enable vision transformers to successfully scale with depth and achieve higher accuracy solely using standard image datasets.
The main objective of the article is to develop and evaluate architectural and optimization modifications that stabilize the training of deep vision transformers and improve image classification performance without relying on external data.
The authors conducted an extensive empirical study using standard image classification benchmarks, primarily ImageNet and several transfer learning datasets. They introduced two core modifications: a per-channel scaling technique called LayerScale to stabilize deep residual blocks, and a specialized architecture named Class-Attention in Image Transformers (CaiT) that separates image patch processing from final class token extraction. The experimental evaluations tested models of varying depths (ranging up to 48 layers) across standard resolutions and paired them with optimization techniques like distillation and stochastic depth.
The evaluation yielded several critical findings. First, LayerScale successfully stabilized deep vision transformers up to 48 layers, preventing optimization failure and enabling performance to scale with depth. Second, explicitly separating patch self-attention from class-attention layers resolved the conflicting roles of early class tokens, improving classification accuracy while reducing computational complexity. Third, the resulting CaiT models established state-of-the-art results on ImageNet without external training data, achieving 86.5% top-1 accuracy on standard validation, while also setting new records on the ImageNet-Real and ImageNet-V2 benchmarks. Finally, the proposed architecture demonstrated strong transfer learning capability, outperforming leading convolutional networks across several downstream datasets.
These findings imply that vision transformers can match or exceed top-tier convolutional networks without requiring specialized external pre-training datasets. For practitioners and decision-makers, this translates to improved classification performance and reduced computational and parameter overhead at higher accuracy regimes, offering a viable path for deploying efficient, deep transformer backbones.
Organizations developing computer vision systems should consider adopting LayerScale when training deep transformers to avoid optimization collapse. When designing classification pipelines, adopting dedicated class-attention stages can improve computational efficiency. For deployment scenarios with strict compute limitations at low model sizes, decision-makers should weigh the trade-offs, as traditional convolutional networks remain more efficient at smaller scales.
The findings are supported by high experimental confidence across multiple standard benchmarks. However, the study focuses predominantly on image classification and transfer learning. Readers should exercise caution before generalizing these results to other vision domains, such as dense object detection or segmentation, without further empirical validation.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This seminal paper introduces the Vision Transformer (ViT) architecture, establishing the patch-tokenization and self-attention paradigm that the source work directly seeks to deepen and stabilize.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It establishes data-efficient training strategies and attention-based distillation for ViTs on ImageNet-1k, providing the foundational training setup and baseline architecture improved upon by the source.
- Paper: Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet, Li Yuan et al. (2021). It analyzes the structural bottlenecks of shallow-wide ViT backbones trained from scratch and explores deeper, narrower architectures that the source builds on to overcome depth saturation.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It builds directly on the stability challenges of deep vision transformers by proposing post-normalization and scaled cosine attention to scale models up to billions of parameters.
- Paper: Scaling Vision Transformers, Xiaohua Zhai et al. (2021). It advances the principles of transformer scaling and stabilization to train an unprecedented 22-billion-parameter Vision Transformer.
- Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). It applies large-scale, deeply optimized vision transformer architectures to self-supervised feature learning across massive datasets without requiring fine-tuning.
- Paper: TransNeXt: Robust Foveal Visual Perception for Vision Transformers, Dai Shi (2024). It addresses residual connection depth degradation and scaling limitations in deep vision transformers by introducing biomimetic aggregated attention mechanisms.
- Paper: DINOv3, Oriane Siméoni et al. (2025). It extends multi-billion-parameter vision transformer optimization and stability techniques to create generalized foundation models across diverse downstream tasks.
