Mimetic Initialization of Self-Attention Layers

Asher TrockmanJ. Zico Kolter

article2023ICML64 citations

Proposes a simple, learning-free weight initialization strategy for self-attention layers that mimics weight patterns observed in pretrained models, boosting classification accuracy by up to 5% when training vanilla Vision Transformers from scratch on small and medium datasets.

Listen

Modern transformer models achieve state-of-the-art performance across artificial intelligence tasks, but they typically require massive datasets and costly pre-training to perform well. When trained from scratch on smaller datasets, standard vision transformers lag significantly behind traditional convolutional networks unless practitioners introduce complex architectural changes, hybrid layers, or specialized auxiliary training routines. This creates high computational costs and deployment barriers for organizations working with limited data or constrained computing budgets.

The article demonstrates that standard, unmodified transformers can be trained effectively from scratch on smaller datasets simply by initializing their self-attention layers with structured mathematical patterns that mimic pre-trained models. The primary objective is to evaluate whether a compute-free, learning-free initialization scheme—termed "mimetic initialization"—can deliver the benefits of pre-training without requiring architectural changes, additional data, or complex training pipelines.

To evaluate this method, the authors conducted empirical experiments using standard vision transformers across multiple computer vision datasets, including CIFAR-10, CIFAR-100, Tiny ImageNet, SVHN, and ImageNet-1k, as well as language modeling tasks on Penn TreeBank and WikiText-103. The technique constructs weight matrices using closed-form singular value decompositions so that the product of query and key weights approximates an identity matrix, while the product of value and projection weights approximates a negative identity matrix. These weights are coupled with standard sinusoidal position embeddings, avoiding any pre-training computation.

The findings show that mimetic initialization consistently and substantially improves model accuracy. On image classification tasks, the method yields accuracy gains of up to 7.8% on CIFAR-10, 6.4% on CIFAR-100, 5.6% on Tiny ImageNet, and up to 4.1% on ImageNet-1k when using standard training pipelines. The benefits are especially pronounced in larger model configurations and when combined with scaled sinusoidal position embeddings. On language benchmarks, the method demonstrates more modest but consistent improvements, reducing perplexity from 28.87 to 28.21 on WikiText-103 and showing slight error reductions on Penn TreeBank.

These results demonstrate that a significant portion of the performance advantage typically attributed to large-scale pre-training stems from establishing favorable initial attention patterns rather than learned features alone. Organizations can use this insight to train vanilla vision transformers on modest datasets using standard training procedures, lowering computational costs, shortening development timelines, and removing the need for custom hybrid architectures. In contrast to prevailing assumptions that transformers require convolutional modifications to handle small vision datasets, proper weight initialization provides a viable, low-complexity alternative.

Practitioners training vision transformers from scratch should adopt mimetic initialization alongside standard sinusoidal position embeddings as a simple, zero-cost default. For language processing, teams should conduct small-scale pilot validations before full deployment, as language gains are less dramatic. Future efforts should explore tailoring mimetic initialization formulas specifically for language structures and investigating whether similar initialization principles apply to non-transformer architectures.

The evidence supporting vision tasks is strong and consistent across multiple benchmarks and architectural sizes, providing high confidence for visual recognition applications. However, confidence should be tempered for language modeling, where improvements are small, and for extremely reduced datasets, where data efficiency gains did not scale inversely with dataset size as expected.

arXiv: 2305.09828
  • Paper: Revealing the Dark Secrets of Masked Image Modeling, Zhenda Xie et al. (2023). This paper analyzes the internal representation and attention head diversity of pre-trained vision models, deepening the understanding of attention patterns that mimetic initialization approximates.
  • Paper: In-context Convergence of Transformers, Yu Huang et al. (2024). This study theoretically models the convergence and optimization dynamics of softmax self-attention layers during gradient descent from specific weight alignments.
  • Paper: Kolmogorov-Arnold Transformer, Xingyi Yang et al. (2025). This work builds on transformer initialization and architectural scaling by designing specialized variance-preserving weight initialization schemes for alternative transformer layers.
Cover for Mimetic Initialization of Self-Attention Layers

Abstract

It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they “look” more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique mimetic initialization.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Observations
  • 4. Method
  • 5. Experiments
  • 5.1. CIFAR-10
  • 5.2. ImageNet
  • 5.3. Other Datasets
  • 6. Why does this initialization work?
  • 7. Language Modeling
  • 8. Conclusion
  • References
  • A. Additional results

Knowls

  1. Knowl 1 — Mimetic Initialization Algorithm for Transformer Self-Attention Layers

    algorithm

    Mimetic initialization sets the self-attention weights of a Transformer so that the product of the query and key projection matrices approximates a positive diagonal matrix, while the product of the value and output projection matrices approximates a negative diagonal matrix, mimicking the empirical weight patterns observed in pretrained Vision Transformers without requiring training.

    For a Transformer with model embedding dimension dd and attention head dimension k=d/hk = d / h (where hh is the number of attention heads), the initialization procedure per self-attention head is defined as follows:

    Input: Model dimension dd, head dimension kk, hyperparameters α1,β1,α2,β2∈[0,1]\alpha_1, \beta_1, \alpha_2, \beta_2 \in [0, 1]
    Output: Query matrix WQ∈Rd×kW_Q \in \mathbb{R}^{d \times k}, Key matrix WK∈Rd×kW_K \in \mathbb{R}^{d \times k}, Value matrix WV∈Rd×dW_V \in \mathbb{R}^{d \times d}, Projection matrix Wproj∈Rd×dW_{proj} \in \mathbb{R}^{d \times d}
    Sample random Gaussian noise matrices:
      Z1∼N(0,1dId)Z_1 \sim \mathcal{N}(0, \frac{1}{d} I_d)
      Z2∼N(0,1dId)Z_2 \sim \mathcal{N}(0, \frac{1}{d} I_d)
    Construct target matrices:
      MQK=α1Z1+β1IdM_{QK} = \alpha_1 Z_1 + \beta_1 I_d
      MVP=α2Z2−β2IdM_{VP} = \alpha_2 Z_2 - \beta_2 I_d
    Compute Singular Value Decompositions:
      U1,Σ1,V1T=SVD(MQK)U_1, \Sigma_1, V_1^T = \text{SVD}(M_{QK})
      U2,Σ2,V2T=SVD(MVP)U_2, \Sigma_2, V_2^T = \text{SVD}(M_{VP})
    Compute factorized weights:
      WQ=U1[:,:k] (Σ1[:k,:k])1/2W_Q = U_1[:, :k] \, (\Sigma_1[:k, :k])^{1/2}
      WK=V1[:,:k] (Σ1[:k,:k])1/2W_K = V_1[:, :k] \, (\Sigma_1[:k, :k])^{1/2}
      WV=U2 (Σ2)1/2W_V = U_2 \, (\Sigma_2)^{1/2}
      Wproj=V2 (Σ2)1/2W_{proj} = V_2 \, (\Sigma_2)^{1/2}
    return WQ,WK,WV,WprojW_Q, W_K, W_V, W_{proj}

    For Vision Transformers, default hyperparameters are α1=β1=0.7\alpha_1 = \beta_1 = 0.7 and α2=β2=0.4\alpha_2 = \beta_2 = 0.4. The matrix Z1Z_1 is sampled independently for each attention head, and the model requires standard sinusoidal positional embeddings.

  2. Knowl 2 — Expected Attention Map Structure Under Mimetic Initialization

    theoretical result

    Let X∈Rn×dX \in \mathbb{R}^{n \times d} represent a sequence of nn LayerNorm-normalized token embeddings of dimension dd, assumed to have zero mean E[X]=0\mathbb{E}[X] = 0 and covariance E[XXT]=In\mathbb{E}[X X^T] = I_n. Let P∈Rn×dP \in \mathbb{R}^{n \times d} denote fixed additive sinusoidal positional embeddings.

    When the query and key projection matrices WQ,WK∈Rd×dW_Q, W_K \in \mathbb{R}^{d \times d} are initialized such that WQWKT=α1Z1+β1IdW_Q W_K^T = \alpha_1 Z_1 + \beta_1 I_d with Z1∼N(0,1dId)Z_1 \sim \mathcal{N}(0, \frac{1}{d} I_d), the expectation of the unnormalized attention logit matrix prior to Softmax is:

    E[(X+P)WQWKT(X+P)T]=E[(X+P)(α1Z1+β1Id)(X+P)T]=β1dIn+β1PPT\mathbb{E}\left[(X + P) W_Q W_K^T (X + P)^T\right] = \mathbb{E}\left[(X + P)(\alpha_1 Z_1 + \beta_1 I_d)(X + P)^T\right] = \beta_1 d I_n + \beta_1 P P^T

    For an attention head dimension kk, this yields expected attention maps of the form:

    Softmax(1k(β1dIn+β1PPT))\text{Softmax}\left(\frac{1}{\sqrt{k}}(\beta_1 d I_n + \beta_1 P P^T)\right)

    The term β1dIn\beta_1 d I_n creates a positive diagonal bias that encourages self-focus (avoiding initial rank collapse while still propagating gradients), while the inner-product term PPTP P^T biases attention toward spatially adjacent tokens according to their positional encodings.

  3. Knowl 3 — CIFAR-10 Classification Accuracy Across Vision Transformer Configurations

    data/table

    When training vanilla Vision Transformers from scratch on CIFAR-10 for 100 epochs using RandAugment, Cutout, AdamW (learning rate 3×10−33 \times 10^{-3}, weight decay 0.01), batch size 512, patch size 2×22 \times 2, and sinusoidal position embeddings, mimetic initialization (α1=β1=0.7,α2=β2=0.4 \alpha_1=\beta_1=0.7, \alpha_2=\beta_2=0.4) provides consistent top-1 accuracy gains across varying widths, depths, and attention head counts compared to standard normal initialization.

    Width Depth Heads Baseline Acc. (%) Mimetic Init Acc. (%) Δ\Delta Acc. (%)
    96 6 3 84.75 87.90 +3.15
    96 12 3 84.75 88.84 +4.09
    192 6 3 85.85 89.68 +4.63
    192 12 1 85.25 89.88 +4.63
    192 12 3 86.07 90.78 +4.71
    192 12 6 86.74 91.38 +4.64
    192 24 3 86.36 91.85 +5.49
    384 12 3 86.26 91.56 +5.30
    384 12 6 84.40 92.17 +7.77
    384 12 12 86.39 92.30 +5.91

    The accuracy improvement is larger for models with higher capacity: a ViT with width 384, depth 12, and 6 heads achieves a +7.77% gain, compared to a +4.09% gain for a width 96 model of equal depth.

  4. Knowl 4 — Ablation Analysis of Mimetic Initialization Components

    data/table

    Ablation experiments on a ViT-Tiny model (width 192, depth 12, 3 heads) trained from scratch on CIFAR-10 for 100 epochs show the relative necessity of each component in the mimetic initialization framework.

    Configuration Top-1 Accuracy (%)
    Full Mimetic Initialization 91.38
    Random Positional Embeddings (with Mimetic Init) 88.70
    No Init (Sinusoidal Positional Embeddings only) 87.39
    Initialize only WQ,WKW_Q, W_K 89.17
    Initialize only WV,WprojW_V, W_{proj} 87.23
    Positive Diagonal Value-Projection (WVWproj∝+cIW_V W_{proj} \propto +cI) 89.65
    GPSA (Gated Positional Self-Attention, 8 heads) 90.03
    GPSA (Gated Positional Self-Attention, 4 heads) 90.83
    GPSA (4 heads) +WVWproj∝−cI+ W_V W_{proj} \propto -cI 91.21
    Pretrained ImageNet Weights (WQ,WK,WV,WprojW_Q, W_K, W_V, W_{proj} + pos. embed) 91.15

    The results highlight that:

    1. Sinusoidal position embeddings are indispensable; switching to learnable random embeddings reduces accuracy by 2.68%.
    2. Initializing WVWprojW_V W_{proj} with a negative diagonal is essential: flipping it to a positive diagonal (+cI+cI) causes a 1.73% drop in accuracy.
    3. Mimetic initialization matches or marginally exceeds direct weight transfer of self-attention and position embedding layers from an ImageNet-pretrained ViT-Tiny (91.38% vs. 91.15%).
  5. Knowl 5 — ImageNet-1k Performance in ResNet-Style vs. DeiT Training Pipelines

    data/table

    On ImageNet-1k (224×224224 \times 224 input resolution, patch size 16×1616 \times 16), mimetic initialization substantially improves training from scratch in standard ResNet-style pipelines and provides steady improvements in long DeiT training schedules.

    Architecture Pipeline Epochs Batch Size Base Acc. (%) Init Acc. (%) Δ\Delta Acc. (%)
    ViT-Tiny ResNet-style 150 640 70.28 73.08 +2.80
    ViT-Tiny ResNet-style 150 1024 67.80 71.92 +4.12
    ViT-Tiny DeiT-style 300 1024 72.08 72.65 +0.57
    ViT-Small DeiT-style 300 1024 79.83 80.36 +0.53

    In the 150-epoch ResNet-style pipeline with cross-entropy loss, where Vision Transformers typically underperform due to weak convolutional inductive biases, mimetic initialization improves top-1 accuracy by 2.80% to 4.12%. In the 300-epoch DeiT pipeline, it yields an accuracy increase of ~0.55%.

  6. Knowl 6 — Generalization Across Additional Vision Benchmarks

    data/table

    Training a ViT-Tiny model from scratch for 100 epochs on diverse image classification datasets confirms that mimetic initialization is effective beyond CIFAR-10 and ImageNet.

    Dataset Patch Size Baseline Acc. (%) Mimetic Init Acc. (%) Δ\Delta Acc. (%)
    Tiny ImageNet 4×44 \times 4 45.24 50.87 +5.63
    CIFAR-100 2×22 \times 2 60.94 67.33 +6.39
    SVHN 2×22 \times 2 96.40 96.79 +0.39

    Mimetic initialization achieves gains of +5.63% on Tiny ImageNet and +6.39% on CIFAR-100, while delivering a +0.39% gain on SVHN where base performance is already near saturation (96.40%).

  7. Knowl 7 — Mimetic Initialization for Autoregressive Language Modeling

    data/table

    Applying mimetic initialization to autoregressive language modeling with vanilla Transformers using sinusoidal positional embeddings and weight tying yields improvements across character-level and word-level tasks. For language models, hyperparameters are tuned to α1=0,β1=0.5\alpha_1 = 0, \beta_1 = 0.5 for WQWKTW_Q W_K^T and α2=0.2,β2=0.2\alpha_2 = 0.2, \beta_2 = 0.2 for WVWprojW_V W_{proj}.

    Dataset Model Architecture Metric Baseline Mimetic Init
    Char-level Penn Treebank (PTB) 12 layers, width 384, 8 heads BPC (↓\downarrow) 1.233 1.210
    Word-level Penn Treebank (PTB) 12 layers, width 384, 8 heads Perplexity (↓\downarrow) 84.84 82.34
    WikiText-103 16 layers, width 410, 10 heads Perplexity (↓\downarrow) 28.87 28.21

    On character-level PTB (100 epochs), bits-per-character (BPC) improves from 1.233 to 1.210. On word-level PTB (100 epochs, with word-level embedding dropout), perplexity decreases by 2.50 points. On the larger WikiText-103 dataset (50 epochs), perplexity drops from 28.87 to 28.21.

  8. Knowl 8 — Role of Negative Diagonal in Value-Projection Spatial Filtering

    model/method

    Initializing the value-projection product to approximate a negative identity matrix (WVWproj≈α2Z2−β2IW_V W_{proj} \approx \alpha_2 Z_2 - \beta_2 I) provides an edge-detector-like spatial filtering mechanism in the residual stream.

    When the attention map approximates the identity InI_n and LayerNorm satisfies η(Xl)≈Xl\eta(X_l) \approx X_l, the residual update for layer ll simplifies to: Xl+1≈Xl+Inη(Xl)WVWproj≈Xl+Xl(α2Z2−β2I)=(1−β2)Xl+α2XlZ2X_{l+1} \approx X_l + I_n \eta(X_l) W_V W_{proj} \approx X_l + X_l(\alpha_2 Z_2 - \beta_2 I) = (1 - \beta_2) X_l + \alpha_2 X_l Z_2

    An empirical comparison using a doubly-block circulant convolution matrix CC with a learnable scalar γ\gamma on CIFAR-10 reveals the necessity of negative spatial coefficients:

    1. Adding convolution inside Softmax, Softmax(XWQWKTXT+γC)\text{Softmax}(X W_Q W_K^T X^T + \gamma C), achieves 87.5% accuracy (below the 88.1% vanilla self-attention baseline).
    2. Adding convolution outside Softmax, Softmax(XWQWKTXT)+γC\text{Softmax}(X W_Q W_K^T X^T) + \gamma C, improves accuracy to 89.9%.
    3. Restricting CC to strictly positive values via C′=Softmax(γC)C' = \text{Softmax}(\gamma C) degrades accuracy to 75.0%.

    This shows that negative spatial mixing coefficients are necessary for effective visual representation learning from scratch, providing the functional basis for the negative diagonal in WVWprojW_V W_{proj}.

  9. Knowl 9 — Impact of Positional Embedding Scaling Factor

    empirical result

    Scaling the sinusoidal positional embeddings by a positive constant multiplier γ\gamma, such that the initial token representation is X+γPX + \gamma P (where X,P∈Rn×dX, P \in \mathbb{R}^{n \times d}), controls the magnitude of the local attention prior γ2PPT\gamma^2 P P^T relative to the token embeddings.

    On CIFAR-10 using a ViT-Tiny model trained for 100 epochs:

    • At default unit scale (γ=1.0\gamma = 1.0), test accuracy is approximately 90.8%.
    • Increasing the scaling multiplier to γ≈2.0\gamma \approx 2.0 improves test accuracy to approximately 91.4% (a ~0.5% boost).
    • Increasing γ>3.0\gamma > 3.0 causes test accuracy to decline.

    This gain arises because larger values of γ\gamma strengthen the relative prominence of the spatial proximity bias β1γ2PPT\beta_1 \gamma^2 P P^T in the expected attention logits.

  10. Knowl 10 — Robustness Across Patch Sizes and Input Resolutions

    data/table

    On CIFAR-10 (100 training epochs on ViT-Tiny), mimetic initialization maintains consistent performance improvements across a wide sweep of input image resolutions (32×3232 \times 32 to 256×256256 \times 256) and patch sizes (2×22 \times 2 to 32×3232 \times 32).

    Patch Size Input Resolution Baseline Acc. (%) Mimetic Init Acc. (%) Δ\Delta Acc. (%)
    2 32×3232 \times 32 85.79 90.46 +4.67
    4 64×6464 \times 64 90.24 92.38 +2.14
    8 128×128128 \times 128 90.03 92.49 +2.46
    16 256×256256 \times 256 88.85 92.74 +3.89
    4 32×3232 \times 32 88.43 90.47 +2.03
    8 64×6464 \times 64 88.00 90.96 +2.96
    16 128×128128 \times 128 87.90 91.90 +4.00
    32 256×256256 \times 256 86.27 90.15 +3.88

    Mimetic initialization outperforms standard initialization across all combinations of patch size and resolution, with accuracy gains ranging from +2.03% to +4.67%.

Coverage note — None was omitted; all primary algorithms, mathematical derivations of expected attention structure, vision and language empirical benchmarks, ablations, and mechanistic analyses are fully captured.

References

  1. 1.Bai, S., Kolter, J. Z., and Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018.
  2. 2.Cao, Y.-H., Yu, H., and Wu, J. Training vision transformers with only 2040 images. arXiv preprint arXiv:2201.10728, 2022.
  3. 3.Cordonnier, J.-B., Loukas, A., and Jaggi, M. On the relationship between self-attention and convolutional layers. arXiv preprint arXiv:1911.03584, 2019.
  4. 4.Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  5. 5.Dai, Z., Liu, H., Le, Q. V., and Tan, M. Coatnet: Marrying convolution and attention for all data sizes. Advances in Neural Information Processing Systems, 34:3965–3977, 2021.
  6. 6.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  7. 7.d’Ascoli, S., Touvron, H., Leavitt, M. L., Morcos, A. S., Biroli, G., and Sagun, L. Convit: Improving vision transformers with soft convolutional inductive biases. In International Conference on Machine Learning, pp. 2286–2296. PMLR, 2021.
  8. 8.Gani, H., Naseer, M., and Yaqub, M. How to train vision transformer on small-scale datasets? arXiv preprint arXiv:2210.07240, 2022.
  9. 9.Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H. Escaping the big data paradigm with compact transformers. arXiv preprint arXiv:2104.05704, 2021.
  10. 10.He, B., Martens, J., Zhang, G., Botev, A., Brock, A., Smith, S. L., and Teh, Y. W. Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation. arXiv preprint arXiv:2302.10322, 2023.
  11. 11.Huang, X. S., Perez, F., Ba, J., and Volkovs, M. Improving transformer optimization through better initialization. In International Conference on Machine Learning, pp. 4475–4483. PMLR, 2020.
  12. 12.Lee, S. H., Lee, S., and Song, B. C. Vision transformer for small-size datasets. arXiv preprint arXiv:2112.13492, 2021.
  13. 13.Liu, Y., Sangineto, E., Bi, W., Sebe, N., Lepri, B., and Nadai, M. Efficient training of visual transformers with small datasets. Advances in Neural Information Processing Systems, 34:23818–23830, 2021.
  14. 14.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pp. 10347–10357. PMLR, 2021a.
  15. 15.Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jegou, H. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 32–42, 2021b.
  16. 16.Trockman, A. and Kolter, J. Z. Patches are all you need? arXiv preprint arXiv:2201.09792, 2022.
  17. 17.Trockman, A., Willmott, D., and Kolter, J. Z. Understanding the covariance structure of convolutional filters. arXiv preprint arXiv:2210.03651, 2022.
  18. 18.Wightman, R., Touvron, H., and Jegou, H. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021.
  19. 19.Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L. Cvt: Introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 22–31, 2021.
  20. 20.Yuan, K., Guo, S., Liu, Z., Zhou, A., Yu, F., and Wu, W. Incorporating convolution designs into visual transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 579–588, 2021.
  21. 21.Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T. Unveiling transformers with lego: a synthetic reasoning task. arXiv preprint arXiv:2206.04301, 2022.
  22. 22.Zhao, J., Schafer, F., and Anandkumar, A. Zero initialization: Initializing residual networks with only zeros and ones. arXiv preprint arXiv:2110.12661, 2021.

Citation

MLA
Trockman, A., and J. Z. Kolter. “Mimetic Initialization of Self-Attention Layers”. International Conference on Machine Learning, vol. 202, 2023, pp. 34456–68, https://proceedings.mlr.press/v202/trockman23a.html.
APA
Trockman, A., & Kolter, J. Z. (2023). Mimetic Initialization of Self-Attention Layers. International Conference on Machine Learning, 202, 34456–34468. https://proceedings.mlr.press/v202/trockman23a.html
Chicago
Trockman, A., and J. Z. Kolter. 2023. “Mimetic Initialization of Self-Attention Layers”. International Conference on Machine Learning 202: 34456–68. https://proceedings.mlr.press/v202/trockman23a.html.
Harvard
Trockman, A. and Kolter, J.Z. (2023) “Mimetic Initialization of Self-Attention Layers”, International Conference on Machine Learning. PMLR, pp. 34456–34468. Available at: https://proceedings.mlr.press/v202/trockman23a.html.
Vancouver
1. Trockman A, Kolter JZ (2023) Mimetic Initialization of Self-Attention Layers. In: International Conference on Machine Learning. PMLR, pp 34456–34468

BibTeX

@InProceedings{pmlr-v202-trockman23a,
  title = 	 {Mimetic Initialization of Self-Attention Layers},
  author =       {Trockman, Asher and Kolter, J Zico},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {34456--34468},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/trockman23a/trockman23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/trockman23a.html},
  abstract = 	 {It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-trained Transformers (particularly for vision) to attempt to find reasons for this discrepancy. Surprisingly, we find that simply initializing the weights of self-attention layers so that they "look" more like their pre-trained counterparts allows us to train vanilla Transformers faster and to higher final accuracies, particularly on vision tasks such as CIFAR-10 and ImageNet classification, where we see gains in accuracy of over 5% and 4%, respectively. Our initialization scheme is closed form, learning-free, and very simple: we set the product of the query and key weights to be approximately the identity, and the product of the value and projection weights to approximately the negative identity. As this mimics the patterns we saw in pre-trained Transformers, we call the technique "mimetic initialization".}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/