Scaling Vision Transformers to 22 Billion Parameters
Mostafa DehghaniJosip DjolongaBasil MustafaPiotr PadlewskiJonathan HeekJustin GilmerAndreas Peter SteinerMathilde CaronRobert GeirhosIbrahim Alabdulmohsin
Introduces an efficient training recipe for a 22-billion-parameter Vision Transformer, demonstrating that language-model-scale vision architectures achieve strong downstream performance, improved fairness tradeoffs, and closer alignment with human visual perception.
While scaling Transformer models to hundreds of billions of parameters has driven major breakthroughs in natural language processing, computer vision architectures have trailed behind. Prior to this work, the largest dense Vision Transformer contained only 4 billion parameters. The article addresses this gap by introducing ViT-22B, a 22-billion-parameter Vision Transformer, to demonstrate that the benefits of massive scale can be realized in visual recognition.
To overcome the severe training instabilities and hardware bottlenecks common at this scale, the researchers implemented key architectural modifications. They introduced parallel attention and multi-layer perceptron blocks, removed unnecessary bias terms, and applied normalization to query and key projections to prevent gradient explosion. The model was trained on 4 billion semi-automatically annotated images using a 2D mesh of 1,024 tensor processing chips, which achieved a high hardware utilization rate of 54.9% by overlapping communication with computation.
The resulting model established new performance benchmarks across multiple visual tasks, even when used simply as a frozen feature extractor. ViT-22B achieved 89.5% top-1 accuracy on standard image recognition benchmarks and reached state-of-the-art results on out-of-distribution tests, such as 74.3% on the challenging ObjectNet evaluation. It also set strong baselines in video classification and dense spatial tasks like semantic segmentation and depth estimation. Beyond raw accuracy, the model demonstrated an 87% shape bias—greatly surpassing typical models that rely on texture and coming closer to human-like visual perception—alongside improved calibration, robustness, and reduced performance disparities across demographic subgroups.
These findings prove that visual models follow scaling trends similar to language models, meaning larger pre-trained backbones can serve as versatile foundations for many downstream tasks without requiring costly full-model fine-tuning. For practical deployment, the authors showed that knowledge distillation allows these benefits to be compressed: a smaller, highly efficient student model distilled from ViT-22B achieved a state-of-the-art 88.6% accuracy on standard benchmarks.
Decision-makers and engineering teams looking to leverage these capabilities should adopt a frozen-backbone or distillation strategy, as fine-tuning all 22 billion parameters is computationally expensive. Because the model was evaluated within an academic and research scope on curated datasets, organizations should conduct domain-specific testing and safety assessments before deploying it in high-stakes operational environments such as surveillance, healthcare, or autonomous driving.
- Paper: Scaling Vision Transformers, Xiaohua Zhai et al. (2021). It introduces the primary ViT scaling principles, architectural stabilization techniques (such as QK-normalization), and large-scale pre-training recipes that form the foundation for scaling dense Vision Transformers.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It introduces the fundamental Vision Transformer architecture and demonstrates the initial data and compute scaling properties that enable transformer-based visual representations.
- Paper: Scaling Vision with Sparse Mixture of Experts, Carlos Riquelme et al. (2021). It establishes early billion-parameter vision scaling dynamics and distributed training challenges on massive image datasets prior to dense multi-billion scaling.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It details essential techniques to combat activation explosion and training instabilities when scaling vision transformers beyond billions of parameters.
- Paper: Big Transfer (BiT): General Visual Representation Learning, Alexander Kolesnikov et al. (2019). It establishes the underlying transfer learning paradigms and hyperparameter rules for evaluating large-scale visual representations across diverse downstream tasks.
- Paper: Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism, Mohammad Shoeybi et al. (2019). It provides the model-parallel and distributed training framework required to train multi-billion-parameter transformer architectures across accelerator clusters.
- Paper: Revisiting Unreasonable Effectiveness of Data in Deep Learning Era, Chen Sun et al. (2017). It demonstrates the empirical scaling laws and performance improvements unlocked by pre-training on hundreds of millions of images like JFT.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). It provides the foundational empirical analysis of internal representations and visual perceptual biases in scaled Vision Transformers compared to convolutional networks.
- Paper: Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, Zhe Chen et al. (2024). It scales vision transformer encoders to multi-billion parameter sizes and integrates them into large multimodal architectures for visual-linguistic tasks.
- Paper: From Sparse to Soft Mixtures of Experts, Joan Puigcerver et al. (2024). It extends massive vision transformer pre-training on JFT-4B by exploring soft mixture-of-experts routing up to 54 billion parameters to further improve compute efficiency.
- Paper: DINOv2: Learning Robust Visual Features without Supervision, Maxime Oquab et al. (2023). It explores whether scaling vision transformer backbones to billion-parameter sizes can produce general-purpose visual representations purely through self-supervised pre-training.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It leverages scaled visual encoders in conjunction with large language models to study progressive multimodal model, data, and test-time scaling.
- Paper: DINOv3, Oriane Siméoni et al. (2025). It scales self-supervised Vision Transformers beyond prior limits to multi-billion parameter regimes while resolving patch feature degradation across massive datasets.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It incorporates dynamic resolution scaling and multimodal positional embeddings to scale vision-language models up to 72 billion parameters.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It investigates advanced native pre-training and scaling recipes up to 78 billion parameters to seamlessly unify visual and linguistic capabilities.
