Scaling Vision Transformers to 22 Billion Parameters

Mostafa DehghaniJosip DjolongaBasil MustafaPiotr PadlewskiJonathan HeekJustin GilmerAndreas Peter SteinerMathilde CaronRobert GeirhosIbrahim Alabdulmohsin

article2023ICML836 citations

Introduces an efficient training recipe for a 22-billion-parameter Vision Transformer, demonstrating that language-model-scale vision architectures achieve strong downstream performance, improved fairness tradeoffs, and closer alignment with human visual perception.

Listen

While scaling Transformer models to hundreds of billions of parameters has driven major breakthroughs in natural language processing, computer vision architectures have trailed behind. Prior to this work, the largest dense Vision Transformer contained only 4 billion parameters. The article addresses this gap by introducing ViT-22B, a 22-billion-parameter Vision Transformer, to demonstrate that the benefits of massive scale can be realized in visual recognition.

To overcome the severe training instabilities and hardware bottlenecks common at this scale, the researchers implemented key architectural modifications. They introduced parallel attention and multi-layer perceptron blocks, removed unnecessary bias terms, and applied normalization to query and key projections to prevent gradient explosion. The model was trained on 4 billion semi-automatically annotated images using a 2D mesh of 1,024 tensor processing chips, which achieved a high hardware utilization rate of 54.9% by overlapping communication with computation.

The resulting model established new performance benchmarks across multiple visual tasks, even when used simply as a frozen feature extractor. ViT-22B achieved 89.5% top-1 accuracy on standard image recognition benchmarks and reached state-of-the-art results on out-of-distribution tests, such as 74.3% on the challenging ObjectNet evaluation. It also set strong baselines in video classification and dense spatial tasks like semantic segmentation and depth estimation. Beyond raw accuracy, the model demonstrated an 87% shape bias—greatly surpassing typical models that rely on texture and coming closer to human-like visual perception—alongside improved calibration, robustness, and reduced performance disparities across demographic subgroups.

These findings prove that visual models follow scaling trends similar to language models, meaning larger pre-trained backbones can serve as versatile foundations for many downstream tasks without requiring costly full-model fine-tuning. For practical deployment, the authors showed that knowledge distillation allows these benefits to be compressed: a smaller, highly efficient student model distilled from ViT-22B achieved a state-of-the-art 88.6% accuracy on standard benchmarks.

Decision-makers and engineering teams looking to leverage these capabilities should adopt a frozen-backbone or distillation strategy, as fine-tuning all 22 billion parameters is computationally expensive. Because the model was evaluated within an academic and research scope on curated datasets, organizations should conduct domain-specific testing and safety assessments before deploying it in high-stakes operational environments such as surveillance, healthcare, or autonomous driving.

arXiv: 2302.05442
Cover for Scaling Vision Transformers to 22 Billion Parameters

Abstract

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.

Table of Contents

  • 1 Introduction
  • 2 Model Architecture
  • 3 Training Infrastructure and Efficiency
  • 4 Experiments
  • 4.1 Training details
  • 4.2 Transfer to image classification
  • 4.2.1 Linear probing
  • 4.2.2 Zero-shot via locked-image tuning
  • 4.2.3 Out-of-distribution
  • 4.3 Transfer to dense prediction
  • 4.3.1 Semantic segmentation
  • 4.3.2 Monocular depth estimation
  • 4.4 Transfer to video classification
  • 4.5 Beyond accuracy on downstream tasks
  • 4.5.1 Fairness
  • 4.5.2 Human Alignment
  • 4.5.3 Plex - pretrained large model extensions
  • 4.5.4 Calibration
  • 4.5.5 Distillation
  • 5 Conclusion
  • References
  • A Zero-shot Classification Examples
  • B Scalability
  • C Model Card
  • D Transfer to image classification: More results and addition details
  • D.1 Linear probing with L-BFGS
  • D.2 Out of distribution classification
  • D.3 Head2Toe
  • D.4 Few-shot
  • E Transfer to dense prediction: More results and addition details.
  • E.1 Semantic segmentation: frozen versus fine-tuning.
  • E.2 Monocular Depth Estimation
  • E.2.1 Dataset
  • E.2.2 Decoder Architectures
  • E.2.3 Training Details
  • E.2.4 Metrics
  • E.2.5 Qualitative Results
  • F Video Classification
  • G Fairness
  • H Calibration
  • I Plex
  • I.1 Details about the evaluation
  • I.2 Details about the Plex architecture
  • I.3 Details about the hyperparameters
  • I.4 Results of Plex-22B and challenges
  • J Error Consistency & Human Alignment
  • K Perceptual similarity
  • L Feature attribution analysis

Knowls

  1. Knowl 1 — ViT-22B Architecture Specifications

    model/method

    ViT-22B is a dense, encoder-only Vision Transformer scaling the architecture to 21.74 billion parameters (21,743M parameters).

    The architectural hyperparameters are defined as follows:

    • Hidden dimension (width dmodeld_{\text{model}}): 6144
    • Depth (number of Transformer encoder blocks): 48
    • MLP intermediate dimension: 24,576 (expansion factor of 4 relative to width)
    • Number of attention heads: 48 (head dimension dk=128d_k = 128)
    • Patch resolution: 14×1414 \times 14 pixels
    • Default input resolution: 224×224224 \times 224 pixels, yielding 256 visual tokens per image
    • Positional embeddings: Learned 1D positional embeddings, interpolated via 2D spatial interpolation when transferring to higher input resolutions.
    • Representation pooling: Multi-head attention pooling (MAP) replaces the class token to aggregate per-token sequence representations into a global feature vector.
  2. Knowl 2 — Query-Key Normalization for Attention Stability

    model/method

    When scaling Vision Transformers to large parameter scales (approximately 8 billion parameters and beyond) with the Adam optimizer at standard learning rates (10−310^{-3}), training encounters severe loss divergence within the first few thousand optimization steps. The divergence is caused by unchecked growth of attention logits (reaching magnitudes exceeding 50,000), which forces the post-softmax attention weights into an almost one-hot distribution with near-zero entropy, triggering exploding gradients and parameter instability.

    To prevent logit explosion and stabilize optimization, Layer Normalization without learnable scale/bias and without mean-centering is applied directly to the query (QQ) and key (KK) tensors prior to computing dot products: Attention(X)=softmax(1dLN(XWQ)(LN(XWK))T)XWV\text{Attention}(X) = \text{softmax}\left( \frac{1}{\sqrt{d}} \text{LN}(X W^Q) (\text{LN}(X W^K))^T \right) X W^V where X∈RN×dmodelX \in \mathbb{R}^{N \times d_{\text{model}}} is the input token sequence (NN tokens, width dmodeld_{\text{model}}), dd is the head dimension, WQ,WK,WV∈Rdmodel×dW^Q, W^K, W^V \in \mathbb{R}^{d_{\text{model}} \times d} are the linear projection matrices for queries, keys, and values, and LN(⋅)\text{LN}(\cdot) denotes layer normalization. This stabilizes training dynamics across scale, allowing ViT-22B to train stably with an unreduced learning rate of 10−310^{-3}.

  3. Knowl 3 — Parallel Layer Formulation and Projection Fusion

    model/method

    ViT-22B alters the standard sequential Transformer encoder block by executing the Multi-Head Self-Attention and Multi-Layer Perceptron (MLP) sub-layers in parallel from a single shared pre-layer normalization: y′=LayerNorm(x)y' = \text{LayerNorm}(x) y=x+MLP(y′)+Attention(y′)y = x + \text{MLP}(y') + \text{Attention}(y') where xx represents the input activation tensor to the encoder layer and yy represents the output activation tensor.

    This formulation enables fusing linear operations across the two sub-layers:

    1. The linear projections for query, key, value (WQ,WK,WVW^Q, W^K, W^V) and the first linear layer of the MLP (up-projection to intermediate dimension 24,576) are combined into a single matrix multiplication operation.
    2. The attention output projection and the second linear layer of the MLP (down-projection to dimension 6144) are combined into a single matrix multiplication operation.

    Biases are omitted on the QKV projections and on all LayerNorms (no bias and no mean centering), while bias terms are retained on the MLP dense layers. This parallel structure speeds up training by approximately 15% without quality degradation.

  4. Knowl 4 — Linear Probing Performance on ImageNet Benchmarks

    empirical result

    When used as a frozen visual feature extractor evaluated via linear probing (using SGD with momentum for 10 epochs at 224px resolution), ViT-22B outperforms prior dense Vision Transformers on in-distribution and out-of-distribution classification tasks:

    Model ImageNet ReaL ImageNet-v2 ObjectNet ImageNet-R ImageNet-A
    ViT-B/32 80.18 86.00 69.56 46.03 75.03 31.20
    ViT-B/16 84.20 88.79 75.07 56.01 82.50 52.67
    ViT-L/16 86.66 90.05 78.57 63.84 89.92 67.96
    ViT-g/14 88.51 90.50 81.10 68.84 92.33 77.51
    ViT-G/14 88.98 90.60 81.32 69.55 91.74 78.79
    ViT-e/14 89.26 90.74 82.51 71.54 94.33 81.56
    ViT-22B 89.51 90.94 83.15 74.30 94.27 83.80

    All backbones are frozen and pre-trained on JFT. Linear probing on frozen ViT-22B features reaches 89.51% top-1 accuracy on ImageNet-1k, approaching or exceeding full fine-tuning performance of smaller high-resolution architectures (e.g., fine-tuned ViT-L/16 at 88.5% and MaxViT-XL at 89.53%). Furthermore, linear probing accuracy on the challenging ObjectNet benchmark scales strongly with parameter count, reaching 74.30%.

  5. Knowl 5 — Zero-Shot Classification via Locked-Image Tuning (LiT)

    empirical result

    By applying Locked-image Tuning (LiT), a text Transformer tower (matching the capacity of ViT-g) is trained contrastively on the English subset of WebLI for 1M steps with a 32k batch size while keeping the ViT-22B vision backbone frozen. Zero-shot transfer performance on ImageNet and out-of-distribution variants demonstrates strong scaling benefits:

    Model ImageNet ImageNet-v2 ImageNet-R ImageNet-A ObjectNet ReaL
    CLIP 76.2 70.1 88.9 77.2 72.3 -
    ALIGN 76.4 70.1 92.2 75.8 72.2 -
    BASIC 85.7 80.6 95.7 85.6 78.9 -
    CoCa 86.3 80.7 96.5 90.2 82.7 -
    LiT-ViT-g/14 85.2 79.8 94.9 81.8 82.5 88.6
    LiT-ViT-e/14 85.4 80.6 96.1 88.0 84.9 88.4
    LiT-ViT-22B 85.9 80.9 96.0 90.1 87.6 88.6

    LiT-ViT-22B achieves 85.9% zero-shot accuracy on ImageNet-1k, 90.1% on ImageNet-A, and 87.6% on ObjectNet, establishing a new zero-shot state-of-the-art on ObjectNet (+2.7% over LiT-ViT-e/14 and +4.9% over CoCa).

  6. Knowl 6 — Knowledge Distillation to Compact Vision Transformers

    empirical result

    Using an ImageNet-finetuned ViT-22B model (fine-tuned at 384px resolution) as a distillation teacher, smaller student architectures (ViT-B/16 and ViT-L/16) achieve state-of-the-art accuracy at their respective scales. Students are initialized from JFT pre-trained checkpoints and trained for 1000 epochs by minimizing the Kullback-Leibler (KL) divergence to the teacher's soft prediction logits across 500 random augmentations and MixUp transforms per training image.

    Student Architecture Training / Distillation Setup ImageNet-1k Top-1 (%)
    ViT-B/16 Standard pretraining on JFT 84.2
    ViT-B/16 Scaled pretraining on JFT 86.6
    ViT-B/16 DeiT III training on ImageNet-21k 86.7
    ViT-B/16 Distilled from ViT-22B (JFT init) 88.6
    ViT-L/16 Standard pretraining on JFT 87.1
    ViT-L/16 Scaled pretraining on JFT 88.5
    ViT-L/16 DeiT III training on ImageNet-21k 87.7
    ViT-L/16 Distilled from ViT-22B (JFT init) 89.6

    Distillation from ViT-22B sets new state-of-the-art results for ViT-B/16 (88.6%) and ViT-L/16 (89.6%) on ImageNet-1k evaluated at 384px resolution.

  7. Knowl 7 — Perceptual Alignment and Shape Bias at 22B Scale

    empirical result

    Evaluating ViT-22B against the model-vs-human benchmark reveals human-like perceptual properties:

    1. Shape vs. Texture Bias: While standard ImageNet-trained convolutional networks and standard vision transformers show a pronounced texture bias (20–30% shape bias vs. 70–80% texture bias), humans exhibit 96% shape bias (4% texture bias). ViT-22B fine-tuned on ImageNet at 384px resolution achieves an 87% shape bias (13% texture bias), the highest recorded for any machine learning vision model.
    2. Human Error Consistency: Aggregated over 17 out-of-distribution image distortion datasets, ViT-22B fine-tuned models outperform previous models across all model-vs-human metrics: ViT-22B at 224px resolution achieves the highest out-of-distribution accuracy, ViT-22B at 384px resolution minimizes the classification accuracy gap to humans, and ViT-22B at 560px resolution achieves the highest error consistency (systematic agreement with human errors above chance levels).
  8. Knowl 8 — Dense Prediction Transfer: Semantic Segmentation and Monocular Depth Estimation

    empirical result

    ViT-22B representations transfer effectively to dense spatial prediction tasks:

    1. Semantic Segmentation Few-Shot Transfer: When transferring to ADE20k using an end-to-end fine-tuned linear decoder with limited training masks, ViT-22B exhibits substantial data-efficiency gains. On 1/16 of the ADE20k training data (1,200 images), ViT-22B attains 44.7 mIoU, outperforming DeiT-III Large (36.1 mIoU) by +8.6 mIoU and ViT-G (42.4 mIoU) by +2.3 mIoU. With full training data, ViT-22B achieves 54.9 mIoU with a linear decoder and 55.3 mIoU with UperNet.
    2. Monocular Depth Estimation on Waymo Open: Using a Dense Prediction Transformer (DPT) decoder on frozen backbone features, ViT-22B reaches an MSE of 0.021, absolute relative error (AbsRel) of 0.095, and inlier threshold δ<1.1\delta < 1.1 of 0.702 (compared to AbsRel 0.112 and δ<1.1\delta < 1.1 0.631 for ViT-e, and AbsRel 0.121 and δ<1.1\delta < 1.1 0.594 for ViT-L). With a purely linear readout decoder, ViT-22B achieves AbsRel 0.166 and δ<1.1\delta < 1.1 0.412, outperforming ViT-e (0.204 / 0.332).
  9. Knowl 9 — Video Action Recognition Transfer with Frozen Spatial Features

    empirical result

    ViT-22B transfers to video action recognition via a factorised spatial-temporal architecture where the spatial transformer is initialized with frozen ViT-22B weights and a lightweight temporal transformer (63.7M parameters) processes a single pooled token representation per video frame.

    On Kinetics 400 (128 frames, stride 2), frozen ViT-22B achieves 88.0% top-1 accuracy, exceeding the 4B-parameter ViT-e baseline (86.5%) by 1.5% and matching CoCA (88.0%), which uses multi-token unpooled spatial representations and higher resolution (576px576\text{px} vs 224px224\text{px}). On Moments in Time (32 frames, stride 2), frozen ViT-22B achieves 44.9% top-1 accuracy, outperforming ViT-e (43.6%) by 1.3%.

  10. Knowl 10 — Fairness Trade-offs and Subgroup Disparity Reductions

    empirical result

    Evaluating demographic parity (DP) fairness on CelebA with binary gender as the sensitive attribute and attributes "attractive" and "smiling" as target labels shows that model scaling improves fairness-accuracy characteristics:

    1. Pareto Trade-off: When debiasing frozen features using the Randomized Threshold Optimizer (RTO) across target DP levels ({0.0,0.025,0.05,0.075,0.1}\{0.0, 0.025, 0.05, 0.075, 0.1\}), ViT-22B achieves higher classification accuracy at every prescribed bias constraint than smaller ViT architectures (ViT-L, ViT-g, ViT-G, ViT-e).
    2. Disparity Reduction Across Subgroups: Prior to debiasing, ViT-22B reduces performance disparities between male and female subgroups. For the "attractive" target, the absolute accuracy gap between gender subgroups is reduced to ≈0.003\approx 0.003 (compared to 0.005–0.0080.005\text{--}0.008 for smaller ViTs). For "smiling", the gap decreases to ≈0.009\approx 0.009 (down from 0.012–0.0170.012\text{--}0.017). Similar gap reductions are observed for Expected Calibration Error (ECE) and Oracle Collaborative AUC (OC-AUC).
  11. Knowl 11 — Training Infrastructure, Model Parallelism, and Hardware Utilization

    model/method

    ViT-22B is trained across 1024 TPUv4 chips using JAX/Flax and Scenic via jax.xmap over a 2D logical device mesh of size t×kt \times k (data-parallel axis tt, model-parallel axis kk):

    1. Asynchronous Parallel Linear Operations (y=Axy = Ax): Matrix multiplications are overlapped asynchronously with neighbor chip communication. For a matrix A∈Rm×nA \in \mathbb{R}^{m \times n} partitioned over kk devices, row-sharding requires communicating (k−1)(n/k)(k-1)(n/k) input floats, whereas column-sharding requires scatter-reducing (k−1)(m/k)(k-1)(m/k) output floats. ViT-22B leverages this asymmetry by applying column-sharding to the Transformer MLP output projection (n=4mn = 4m) to minimize communication volume, while applying row-sharding elsewhere.
    2. Asynchronous Parameter Sharding: Large parameter tensors are sharded over the data axis and gathered asynchronously prior to forward execution, then scattered during the backward pass.
    3. Hardware Utilization: ViT-22B processes 1.15k tokens per second per core (forward + backward pass), achieving a Model FLOPs Utilization (MFU) of 54.9% on TPUv4, compared to 46.2% reported for PaLM-540B and 44.0% for ViT-e.
  12. Knowl 12 — Head2Toe Multi-Layer Feature Probing vs. Linear Probing

    empirical result

    Head2Toe probing evaluates frozen visual representations by concatenating intermediate features across all 48 encoder blocks, post-positional embedding features, post-pooling head features, pre-logits, and logits (token-averaged to form a 349,081-dimensional vector) and training a linear classifier with no feature selection. On the 19-task VTAB-1k benchmark and four standard vision datasets:

    Method VTAB-Avg Natural Specialized Structured CIFAR-10 CIFAR-100 Flowers Pets
    Fine-tuning 76.71 89.09 87.08 61.83 99.63 95.96 97.59 99.75
    Linear (6144-d) 63.15 80.86 87.05 35.70 99.37 93.39 99.75 98.15
    Head2Toe 70.12 84.60 88.61 48.19 99.45 94.11 99.69 97.46

    Head2Toe achieves 70.12% VTAB-1k average accuracy, gaining +6.97% points over standard pre-logit linear probing (63.15%), with the largest gain observed on Structured tasks (+12.49% points), recovering over half of the performance gap to full fine-tuning (76.71%) while requiring substantially less computation and memory.

Coverage note — Omitted minor secondary analyses including qualitative Integrated Gradients attribution visualizations (Appendix L), exact hyperparameter sweeps for Plex variants (Appendix I), and repetitive lists of dataset metadata (Appendix C / Table 13) which do not alter the main scientific conclusions.

References

  1. 1.Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi. Exploring the limits of large scale pre-training. arXiv preprint arXiv:2110.02095, 2021.
  2. 2.Thomas Adler, Johannes Brandstetter, Michael Widrich, Andreas Mayr, David P. Kreil, Michael Kopp, Günter Klambauer, and Sepp Hochreiter. Cross-domain few-shot learning by representation fusion. arXiv preprint arXiv:2010.06498, 2020.
  3. 3.Osman Aka, Ken Burke, Alex Bauerle, Christina Greer, and Margaret Mitchell. Measuring model biases in the absence of ground truth. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 327–335, 2021.
  4. 4.Ibrahim Alabdulmohsin and Mario Lucic. A near optimal algorithm for debiasing trained machine learning models. In NeurIPS, 2021.
  5. 5.Anders Andreassen, Yasaman Bahri, Behnam Neyshabur, and Rebecca Roelofs. The evolution of out-of-distribution robustness throughout fine-tuning. arXiv preprint arXiv:2106.15831, 2021.
  6. 6.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. ViViT: A video vision transformer. In CVPR, 2021.
  7. 7.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  8. 8.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In NeurIPS, pages 9448–9458, 2019.
  9. 9.Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al. Deepmind lab. arXiv preprint arXiv:1612.03801, 2016.
  10. 10.Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord. Are we done with imagenet? arXiv preprint arXiv:2006.07159, 2020.
  11. 11.Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. arXiv preprint arXiv:2212.08013, 2022a.
  12. 12.Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. Knowledge distillation: A good teacher is patient and consistent. In CVPR, pages 10925–10934, 2022b.
  13. 13.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  14. 14.Joy Buolamwini and Timnit Gebru. Gender shades: Intersectional accuracy disparities in commercial gender classification. In FAccT, 2018.
  15. 15.Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu. A limited memory algorithm for bound constrained optimization. SIAM Journal on scientific computing, 16(5):1190–1208, 1995.
  16. 16.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334), 2017.
  17. 17.Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. PaLI: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022.
  18. 18.Gong Cheng, Junwei Han, and Xiaoqiang Lu. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, 2017.
  19. 19.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  20. 20.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  21. 21.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In CVPR, pages 3606–3613, 2014.
  22. 22.Mark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton, and Jesse Berent. Correlated input-dependent label noise in large-scale image classification. In CVPR, pages 1551–1560, 2021.
  23. 23.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relationship between self-attention and convolutional layers. arXiv preprint arXiv:1911.03584, 2019.
  24. 24.Yin Cui, Yang Song, Chen Sun, Andrew Howard, and Serge Belongie. Large scale fine-grained categorization and domain-specific transfer learning. In CVPR, pages 4109–4118, 2018.
  25. 25.Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay. The efficiency misnomer. arXiv preprint arXiv:2110.12894, 2021a.
  26. 26.Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The benchmark lottery. arXiv preprint arXiv:2107.07002, 2021b.
  27. 27.Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, and Yi Tay. Scenic: A jax library for computer vision research and beyond. In CVPR, 2022.
  28. 28.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009.
  29. 29.Jessica Deuschel, Bettina Finzel, and Ines Rieger. Uncovering the bias in facial expressions. arXiv preprint arXiv:2011.11311, 2020.
  30. 30.Josip Djolonga, Frances Hubis, Matthias Minderer, Zachary Nado, Jeremy Nixon, Rob Romijnders, Dustin Tran, and Mario Lucic. Robustness Metrics, 2020. URL https://github.com/google-research/robustness_metrics.
  31. 31.Josip Djolonga, Jessica Yung, Michael Tschannen, Rob Romijnders, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Matthias Minderer, Alexander D’Amour, Dan Moldovan, Sylvain Gelly, Neil Houlsby, Xiaohua Zhai, and Mario Lucic. On robustness and transferability of convolutional neural networks. In CVPR, pages 16458–16468, 2021.
  32. 32.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  33. 33.Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. Fairness through awareness. In Innovations in Theoretical Computer Science, 2012.
  34. 34.David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. In NeurIPS, 2014.
  35. 35.Ran El-Yaniv and Yair Wiener. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(5), 2010.
  36. 36.Utku Evci, Vincent Dumoulin, Hugo Larochelle, and Michael C Mozer. Head2Toe: Utilizing intermediate representations for better transfer learning. In ICML, 2022.
  37. 37.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (VOC) challenge. IJCV, 2010.
  38. 38.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021.
  39. 39.Stanislav Fort, Jie Ren, and Balaji Lakshminarayanan. Exploring the limits of out-of-distribution detection. In NeurIPS, pages 7068–7081, 2021.
  40. 40.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 32(11):1231–1237, 2013.
  41. 41.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In ICLR, 2019.
  42. 42.Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Partial success in closing the gap between human and machine vision. In NeurIPS, pages 23885–23899, 2021.
  43. 43.Justin Gilmer, Andrea Schioppa, and Jeremy Cohen. Intriguing Properties of Transformer Training Instabilities, 2023. To appear.
  44. 44.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML, 2017.
  45. 45.Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee. Flax: A neural network library and ecosystem for JAX, 2020. URL http://github.com/google/flax.
  46. 46.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019.
  47. 47.Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. Women also snowboard: Overcoming bias in captioning models. In ECCV, 2018.
  48. 48.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261, 2019.
  49. 49.Dan Hendrycks, Steven Basart, Mantas Mazeika, Mohammadreza Mostajabi, Jacob Steinhardt, and Dawn Song. Scaling out-of-distribution detection for real-world settings. arXiv preprint arXiv:1911.11132, 2019.
  50. 50.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. arXiv preprint arXiv:2006.16241, 2020.
  51. 51.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In CVPR, pages 15262–15271, 2021.
  52. 52.Max Hermann, Boitumelo Ruf, Martin Weinmann, and Stefan Hinz. Self-supervised learning for monocular depth estimation from aerial imagery. arXiv preprint arXiv:2008.07246, 2020.
  53. 53.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  54. 54.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, pages 4904–4916, 2021.
  55. 55.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, pages 2901–2910, 2017.
  56. 56.Norman P Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson. A domain-specific supercomputer for training deep neural networks. Communications of the ACM, 63(7):67–78, 2020.
  57. 57.Kaggle and EyePacs. Kaggle diabetic retinopathy detection, 2015. URL https://www.kaggle.com/c/diabetic-retinopathy-detection/data.
  58. 58.Jakob Nikolas Kather, Cleo-Aron Weis, Francesco Bianconi, Susanne M Melchers, Lothar R Schad, Timo Gaiser, Alexander Marx, and Frank Gerrit Zöllner. Multi-class texture analysis in colorectal cancer histology. Scientific reports, 6:27988, 2016.
  59. 59.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  60. 60.Amr Khalifa, Michael C. Mozer, Hanie Sedghi, Behnam Neyshabur, and Ibrahim Alabdulmohsin. Layer-stack temperature scaling. arXiv preprint arXiv:2211.10193, 2022.
  61. 61.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  62. 62.Ian D Kivlichan, Zi Lin, Jeremiah Liu, and Lucy Vasserman. Measuring and improving model-moderator collaboration using uncertainty estimation. arXiv preprint arXiv:2107.04212, 2021.
  63. 63.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big Transfer (BiT): General visual representation learning. In ECCV, pages 491–507, 2020.
  64. 64.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition, 2013.
  65. 65.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images, 2009. Technical Report, University of Toronto.
  66. 66.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In EMNLP, pages 66–71, November 2018.
  67. 67.Manoj Kumar, Neil Houlsby, Nal Kalchbrenner, and Ekin Dogus Cubuk. Do better imagenet classifiers assess perceptual similarity better? Transactions on Machine Learning Research, 2022.
  68. 68.Yann LeCun, Fu Jie Huang, and Leon Bottou. Learning methods for generic object recognition with invariance to pose and lighting. In CVPR, volume 2, 2004.
  69. 69.Fei-Fei Li, Marco Andreeto, Marc’Aurelio Ranzato, and Pietro Perona. Caltech 101, 2022. CaltechDATA, doi: 10.22002/D1.20086.
  70. 70.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In ICCV, 2015.
  71. 71.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In ICLR, 2017.
  72. 72.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, pages 181–196, 2018.
  73. 73.Loic Matthey, Irina Higgins, Demis Hassabis, and Alexander Lerchner. dSprites: Disentanglement testing sprites dataset, 2017. URL https://github.com/deepmind/dsprites-dataset/.
  74. 74.Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic. Revisiting the calibration of modern neural networks. NeurIPS, 34:15682–15694, 2021.
  75. 75.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In FAccT, pages 220–229, 2019.
  76. 76.Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al. Moments in time dataset: one million videos for event understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2), 2019.
  77. 77.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In CVPR, 2014.
  78. 78.Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In AAAI, 2015.
  79. 79.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng. Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, 2011.
  80. 80.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729, 2008.
  81. 81.Sinno Jialin Pan and Qiang Yang. A survey on transfer learning. IEEE Transactions on Knowledge and Data Engineering, 22(10):1345–1359, 2010.
  82. 82.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012.
  83. 83.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V Le. Combined scaling for zero-shot transfer learning. arXiv preprint arXiv:2111.10050, 2021.
  84. 84.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently scaling transformer inference. arXiv preprint arXiv:2211.05102, 2022.
  85. 85.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763, 2021.
  86. 86.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019.
  87. 87.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In CVPR, pages 12179–12188, 2021.
  88. 88.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In ICML, 2019.
  89. 89.Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby. Scaling vision with sparse mixture of experts. In NeurIPS, volume 34, pages 8583–8595, 2021.
  90. 90.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 2022.
  91. 91.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In ICCV, 2021.
  92. 92.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In CVPR, pages 843–852, 2017.
  93. 93.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020.
  94. 94.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In ICML, pages 3319–3328, 2017.
  95. 95.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  96. 96.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. In NeurIPS, 2020.
  97. 97.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022.
  98. 98.Eu Wern Teh and Graham W Taylor. Metric learning for patch classification in digital pathology. In International Conference on Medical Imaging with Deep Learning–Extended Abstract Track, 2019.
  99. 99.Hugo Touvron, Matthieu Cord, and Hervé Jégou. DeiT III: Revenge of the ViT. In ECCV, 2022.
  100. 100.Dustin Tran, Jeremiah Liu, Michael W Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet, Huiyi Hu, et al. Plex: Towards reliability using pretrained large model extensions. arXiv preprint arXiv:2207.07411, 2022.
  101. 101.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011.
  102. 102.Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  103. 103.Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, Sameer Kumar, Tongfei Guo, Yuanzhong Xu, and Zongwei Zhou. Overlap communication with dependent computation via decomposition in large deep learning models. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems, pages 93––106, 2022a.
  104. 104.Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, Limin Wang, and Yu Qiao. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022b.
  105. 105.Zeyu Wang, Klint Qinami, Ioannis Christos Karakozis, Kyle Genova, Prem Nair, Kenji Hata, and Olga Russakovsky. Towards fairness in visual recognition: Effective strategies for bias mitigation. In CVPR, 2020.
  106. 106.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  107. 107.Yeming Wen, Dustin Tran, and Jimmy Ba. Batchensemble: an alternative approach to efficient ensemble and lifelong learning. In ICLR, 2019.
  108. 108.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In CVPR, pages 3485–3492, 2010.
  109. 109.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018.
  110. 110.Yi Yang and Shawn Newsam. Bag-of-visual-words and spatial extensions for land-use classification. In SIGSPATIAL international conference on advances in geographic information systems, pages 270–279, 2010.
  111. 111.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research, 2022a.
  112. 112.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. Transactions on Machine Learning Research, 2022b.
  113. 113.Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, and Krishna P Gummadi. Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment. In International Conference on World Wide Web, 2017.
  114. 114.Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, et al. A large-scale study of representation learning with the visual task adaptation benchmark. arXiv preprint arXiv:1910.04867, 2019.
  115. 115.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In CVPR, pages 12104–12113, 2022a.
  116. 116.Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. LiT: Zero-shot transfer with locked-image text tuning. In CVPR, pages 18123–18133, 2022b.
  117. 117.Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  118. 118.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
  119. 119.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. arXiv preprint arXiv:1707.09457, 2017.
  120. 120.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6): 1452–1464, 2017a.
  121. 121.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. In CVPR, 2017b.

Citation

MLA
Dehghani, M., et al. “Scaling Vision Transformers to 22 Billion Parameters”. arXiv, 2023, http://arxiv.org/abs/2302.05442v1.
APA
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Riquelme, C., Minderer, M., Puigcerver, J., Evci, U., … Houlsby, N. (2023). Scaling Vision Transformers to 22 Billion Parameters. arXiv. http://arxiv.org/abs/2302.05442v1
Chicago
Dehghani, M., J. Djolonga, B. Mustafa, et al. 2023. “Scaling Vision Transformers to 22 Billion Parameters”. arXiv. http://arxiv.org/abs/2302.05442v1.
Harvard
Dehghani, M. et al. (2023) “Scaling Vision Transformers to 22 Billion Parameters”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.05442v1.
Vancouver
1. Dehghani M, Djolonga J, Mustafa B, et al (2023) Scaling Vision Transformers to 22 Billion Parameters. arXiv

BibTeX

@article{dehghani2023scaling,
  title = {Scaling Vision Transformers to 22 Billion Parameters},
  author = {Dehghani, Mostafa and Djolonga, Josip and Mustafa, Basil and Padlewski, Piotr and Heek, Jonathan and Gilmer, Justin and Steiner, Andreas and Caron, Mathilde and Geirhos, Robert and Alabdulmohsin, Ibrahim and Jenatton, Rodolphe and Beyer, Lucas and Tschannen, Michael and Arnab, Anurag and Wang, Xiao and Riquelme, Carlos and Minderer, Matthias and Puigcerver, Joan and Evci, Utku and Kumar, Manoj and Steenkiste, Sjoerd van and Elsayed, Gamaleldin F. and Mahendran, Aravindh and Yu, Fisher and Oliver, Avital and Huot, Fantine and Bastings, Jasmijn and Collier, Mark Patrick and Gritsenko, Alexey and Birodkar, Vighnesh and Vasconcelos, Cristina and Tay, Yi and Mensink, Thomas and Kolesnikov, Alexander and Pavetić, Filip and Tran, Dustin and Kipf, Thomas and Lučić, Mario and Zhai, Xiaohua and Keysers, Daniel and Harmsen, Jeremiah and Houlsby, Neil},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.05442v1},
  eprint = {2302.05442}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/