LayerNorm Induces Recency Bias in Transformer Decoders

Junu KimXiao LiuZheng-Wen LinLei JiYeyun GongEdward Choi

article2025ACL1 citations

Demonstrates that LayerNorm turns the early-token attention bias of causal transformers into recency bias, resolving a fundamental contradiction in transformer behavior and informing positional encoding design.

Listen

Modern sequence modeling architectures, particularly Transformer decoders used in large language models, rely heavily on positional mechanisms to track token order and generalize to long texts. Prior theoretical work suggested that stacking causal self-attention layers inherently causes models to focus on earlier tokens in a sequence. However, empirical observations of real-world Transformer decoders show the opposite effect: models exhibit a recency bias, systematically assigning higher attention weights to more recent tokens. Understanding why this happens without explicit positional encodings is critical for diagnosing model behavior and designing architectures that extrapolate well over long context windows.

This article aims to resolve this contradiction by demonstrating both theoretically and empirically how specific architectural components drive positional behavior. Specifically, it evaluates how layer normalization, residual connections, and non-uniform (anisotropic) input token distributions interact with causal self-attention to induce recency bias.

To conduct this evaluation, the authors combine mathematical proofs with extensive numerical simulations of Transformer decoders lacking learnable parameters and positional encodings. They track attention behavior across varying hidden dimensions, layers, and input anisotropy levels, measuring recency using a recency probability metric—the probability that attention scores assign higher weight to a closer preceding token than to a more distant one across 10 million simulation trials.

The analysis reveals four key findings. First, layer normalization is the core driver of recency bias; stacking causal self-attention with layer normalization provably produces recency bias in the second layer and beyond, whereas stacking self-attention alone yields no recency bias (recency probability remains around 0.50). Second, the directional bias of input token embeddings strongly amplifies this effect. Under realistic anisotropic embeddings (anisotropy factor of 0.5), recency probability jumps substantially (from 0.50 to roughly 0.64 in lower-dimensional second-layer settings and up to nearly 0.91 in deeper layers). Third, residual connections moderate the effect, reducing the recency probability (e.g., from roughly 0.64 down to 0.59 in second-layer tests) by blending in unshared vector components. Fourth, this emergent recency bias extends structurally across modern multi-head and grouped-query attention variants.

These findings imply that Transformer decoders inherently encode positional preferences through standard normalization layers rather than purely through explicit positional encodings. However, the resulting recency bias is non-uniform across sequence positions and fails to satisfy strict relative distance invariance. This uneven positional interaction may actively degrade long-context generalization and extrapolation performance in production language models, creating unintended performance risks for applications requiring balanced attention across lengthy documents.

To address this, engineering and research teams developing foundation models should account for normalization-driven biases when designing positional schemes. A promising next step is to explore mitigation strategies, such as modifying the causal attention mask to counteract uneven positional weighting, before deploying new long-context architectures.

The study’s conclusions are robust within its mathematical and simulation framework, though confidence should be contextualized by noted boundary conditions. The analysis simplifies real-world architectures by omitting feed-forward networks, learnable parameter updates, and modern explicit positional encodings such as rotary position embeddings. Further empirical validation on fully trained, production-scale models is recommended to measure the downstream impact on task accuracy.

arXiv: 2509.21042starmpcc/layernorm_recency_bias

No sufficiently relevant recommendations were found.

Cover for LayerNorm Induces Recency Bias in Transformer Decoders

Abstract

Causal self-attention provides positional information to Transformer decoders. Prior work has shown that stacks of causal self-attention layers alone induce a positional bias in attention scores toward earlier tokens. However, this differs from the bias toward later tokens typically observed in Transformer decoders, known as recency bias. We address this discrepancy by analyzing the interaction between causal self-attention and other architectural components. We show that stacked causal self-attention layers combined with LayerNorm induce recency bias. Furthermore, we examine the effects of residual connections and the distribution of input token embeddings on this bias. Our results provide new theoretical insights into how positional information interacts with architectural components and suggest directions for improving positional encoding strategies.

Table of Contents

  • 1 Introduction
  • 2 Theoretical Analysis
  • 2.1 Preliminaries
  • 2.2 LayerNorm
  • 2.3 Residual Connection
  • 2.4 Distribution of Input Token Embeddings
  • 3 Empirical Analysis
  • 4 Conclusion
  • References
  • A Proofs
  • A.1 Proof of Theorem
  • A.2 Proof of Proposition
  • A.3 Proof of Proposition and
  • B Extended Results

Knowls

  1. Knowl 1 — LayerNorm makes the second causal-attention layer recency-biased

    theoretical result

    Consider a sequence of nn independent input vectors xi(0)∈Rdx_i^{(0)}\in\mathbb{R}^d, each distributed as N(0,Id/d)\mathcal{N}(0,I_d/d). In each layer, normalize each token vector with LayerNorm, apply causal self-attention with identity query, key, and value projections, and use the attention output as the next-layer representation; omit residual connections, feed-forward networks, positional encodings, and learned parameters. For large dd, LayerNorm is approximated by yi≈d xi/∥xi∥2y_i\approx \sqrt{d}\,x_i/\|x_i\|_2. In the first layer, the diagonal attention logit is approximately d\sqrt d and the off-diagonal logits approximately zero. The resulting second-layer logits, for query position ii and earlier key position j<ij<i, are approximately S_{ij}^{(2)}=\frac{\sqrt d,(e^{\sqrt d}+j-1)}{\sqrt{e^{2\sqrt d}+i-1},\sqrt{e^{2\sqrt d}+j-1}}.$$ Here Sij(2)S_{ij}^{(2)} is the unnormalized attention score, and ee is the base of the natural logarithm. For fixed ii, this expression increases strictly with jj; the diagonal score is Sii(2)=dS_{ii}^{(2)}=\sqrt d and exceeds the score immediately before it. Thus, in this high-dimensional approximation, stacking causal attention with LayerNorm induces recency bias by the second layer, even in this simplified architecture.

  2. Knowl 2 — Recency bias is defined by score ordering, with RP measuring its frequency

    definition

    Let SijS_{ij} be the raw attention score for query position ii and key position jj in a causal sequence, with positions indexed from earlier to later. The score matrix has recency bias when Sij>SikS_{ij}>S_{ik} for every i≥j>ki\ge j>k: for a fixed query, a closer key scores higher than a more distant key. For simulation results, the paper uses recency probability (RP), the probability over eligible comparisons and simulation draws that Sij>SikS_{ij}>S_{ik} for i>j>ki>j>k, excluding diagonal comparisons. An RP of 0.50.5 is the no-preference baseline; values above 0.50.5 indicate a tendency toward more recent keys.

  3. Knowl 3 — The second-layer recency effect requires LayerNorm in the isotropic control

    theoretical result

    For independent Gaussian input embeddings xi(0)∼N(0,Id/d)x_i^{(0)}\sim\mathcal{N}(0,I_d/d) and large hidden size dd, the paper's simplified stack of causal self-attention layers does not exhibit recency bias in the second-layer scores when LayerNorm is removed. In the high-dimensional approximation, the off-diagonal score for a fixed query does not vary with the earlier key position, so it does not increase toward more recent keys. This contrasts with the same setup using LayerNorm, which yields strictly increasing second-layer scores for earlier keys as their positions approach the query.

  4. Knowl 4 — Second-layer recency bias persists with shared anisotropy and residual connections

    theoretical result

    The paper models anisotropic input vectors as xi(0)=ϵi+α/(1−α) vx_i^{(0)}=\epsilon_i+\sqrt{\alpha/(1-\alpha)}\,v, where ϵi\epsilon_i and the shared vector vv are independent draws from N(0,Id/d)\mathcal{N}(0,I_d/d) and 0≤α<10\le\alpha<1. In the simplified high-dimensional, LayerNorm-based causal-attention stack, let γ=0\gamma=0 denote no residual connection and γ=1\gamma=1 denote adding the input through a residual connection. For either residual setting and any allowed anisotropy level, the second-layer score for a fixed query increases strictly with the earlier key position. Thus, neither adding this shared embedding component nor including the residual connection removes the theoretical recency ordering.

  5. Knowl 5 — Simulation isolates attention, normalization, residuals, and embedding anisotropy

    experimental setup

    The simulations use an attention-only Transformer with no learned parameters, positional encodings, or feed-forward modules. Each causal-attention layer uses LayerNorm and identity projections; residual connections are either absent or included. The input has 10 token vectors drawn from the isotropic or shared-component Gaussian distribution xi(0)=ϵi+α/(1−α)vx_i^{(0)}=\epsilon_i+\sqrt{\alpha/(1-\alpha)}v. The main comparison uses hidden sizes d∈{16,64}d\in\{16,64\} and averages attention scores over 10,000,000 simulations. The study also examines layers 3 and 4 and varies α\alpha to assess how embedding anisotropy and depth affect recency probability.

  6. Knowl 6 — LayerNorm and anisotropy determine the size of the measured second-layer bias

    empirical result

    In the 10-token simulations, RP is the fraction of eligible off-diagonal comparisons in which a nearer key receives a higher score than a farther key; 0.50.5 is the neutral baseline. Without LayerNorm, second-layer RP is 0.50000.5000 for both d=16d=16 and d=64d=64. With LayerNorm and isotropic inputs (α=0\alpha=0), it is 0.50150.5015 for d=16d=16 and 0.50000.5000 for d=64d=64, indicating little measurable bias under those conditions. With LayerNorm and anisotropy α=0.5\alpha=0.5, second-layer RP rises to 0.63820.6382 for d=16d=16 and 0.55440.5544 for d=64d=64. Adding residual connections at α=0.5\alpha=0.5 lowers those values to 0.59310.5931 and 0.54570.5457, respectively. The measurements support the predicted recency ordering while showing that its measured strength depends on dimension, anisotropy, and residual connections.

  7. Knowl 7 — Bias strengthens with depth and anisotropy, while residuals moderate it

    empirical result

    In the simulations with α=0.5\alpha=0.5, second-layer recency probability increases further in later layers. For d=16d=16, RP is 0.80020.8002 in layer 3 and 0.90790.9079 in layer 4 without residual connections, versus 0.66100.6610 and 0.70780.7078 with residual connections. For d=64d=64, the corresponding values are 0.62740.6274 and 0.70730.7073 without residuals, and 0.57910.5791 and 0.60530.6053 with residuals. Simulations varying α\alpha from 0.20.2 to 0.80.8 show that, from layer 2 onward, stronger anisotropy is associated with higher RP; adding residual connections reduces the measured bias across the tested anisotropy levels.

  8. Knowl 8 — The argument extends qualitatively to multi-head attention variants

    theoretical result

    The paper states that its theoretical framework extends to multi-head attention, multi-query attention, and grouped-query attention. In the dimensional terms of the analysis, the full hidden size dd is replaced by the per-head dimension dhd_h; in particular, the scaling term d\sqrt d becomes dh\sqrt{d_h}. The paper claims that this replacement preserves the derivation's structure and its qualitative conclusion about recency bias. This is a theoretical extension, not an empirical evaluation of those attention variants in full decoder layers.

  9. Knowl 9 — The induced bias is positional but not purely relative

    theoretical result

    The attention-score pattern produced by causal masking and LayerNorm is not a relative positional encoding in the sense of depending only on the distance between a query and key. In the paper's analysis, a score depends on the query position and key position separately, so pairs at the same separation need not have identical scores. The authors suggest that this uneven positional interaction could adversely affect length generalization, but they do not establish that performance effect experimentally.

  10. Knowl 10 — The study does not test full-model effects or performance consequences

    limitation

    The analysis and simulations do not cover feed-forward networks, other learned Transformer parameters, or multi-head attention as components of full decoder layers. They also do not evaluate how the induced recency bias affects overall model performance. Finally, although the paper notes that most modern Transformer decoders use RoPE, it leaves the interaction between RoPE and the positional information induced by causal self-attention unexamined.

Coverage note — Only algebraic proof steps and redundant visualizations are omitted; they add no standalone contributed results beyond the theoretical claims and simulation findings captured here.

References

  1. 1.Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sang￾hai. 2023. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901.
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  3. 3.Federico Barbero, Alex Vitvitskyi, Christos Perivolaropoulos, Razvan Pascanu, and Petar Velickovič. 2025. Round and round we go! what makes rotary positional encodings useful? In The Thirteenth International Conference on Learning Representations.
  4. 4.Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen, Alexander Rudnicky, and Peter Ramadge. 2023. Latent positional information is in the self-attention variance of transformer language models without positional embeddings. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1183–1193.
  5. 5.Philipp Dufter, Martin Schmitt, and Hinrich Schütze. 2022. Position information in transformers: An overview. Computational Linguistics, 48(3):733–763.
  6. 6.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 55–65, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. 2022. Transformer language models without positional encodings still learn positional information. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1382–1390.
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  9. 9.Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023. The impact of positional encoding on length generalization in transformers. Advances in Neural Information Processing Systems, 36:24892–24928.
  10. 10.Ofir Press, Noah Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations.
  11. 11.Gautam Reddy. 2024. The mechanistic basis of data dependence and abrupt learning in an in-context classification task. In The Twelfth International Conference on Learning Representations.
  12. 12.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155.
  13. 13.Noam Shazeer. 2019. Fast transformer decoding: One write-head is all you need. arXiv preprint arXiv:1911.02150.
  14. 14.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063.
  15. 15.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  16. 16.Xinyi Wu, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie. 2025. On the emergence of position bias in transformers. In Forty-second International Conference on Machine Learning.
  17. 17.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020. On layer normalization in the transformer architecture. In International conference on machine learning, pages 10524–10533. PMLR.
  18. 18.Liang Zhao, Xiachong Feng, Xiaocheng Feng, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin, and Ting Liu. 2024. Length extrapolation of transformers: A survey from the perspective of positional encoding. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9959–9977.
  19. 19.Chunsheng Zuo, Pavel Guerzhoy, and Michael Guerzhoy. 2025. Position information emerges in causal transformers without positional encodings via similarity of nearby embeddings. In Proceedings of the 31st International Conference on Computational Linguistics, pages 9418–9430.

Citation

MLA
Kim, J., et al. “LayerNorm Induces Recency Bias in Transformer Decoders”. arXiv, 2025, https://doi.org/10.48550/arxiv.2509.21042.
APA
Kim, J., Liu, X., Lin, Z., Ji, L., Gong, Y., & Choi, E. (2025). LayerNorm Induces Recency Bias in Transformer Decoders. arXiv. https://doi.org/10.48550/arxiv.2509.21042
Chicago
Kim, J., X. Liu, Z. Lin, L. Ji, Y. Gong, and E. Choi. 2025. “LayerNorm Induces Recency Bias in Transformer Decoders”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2509.21042.
Harvard
Kim, J. et al. (2025) “LayerNorm Induces Recency Bias in Transformer Decoders”. arXiv. Available at: https://doi.org/10.48550/arxiv.2509.21042.
Vancouver
1. Kim J, Liu X, Lin Z, Ji L, Gong Y, Choi E (2025) LayerNorm Induces Recency Bias in Transformer Decoders. https://doi.org/10.48550/arxiv.2509.21042

BibTeX

@misc{https://doi.org/10.48550/arxiv.2509.21042,
  doi = {10.48550/ARXIV.2509.21042},
  url = {https://arxiv.org/abs/2509.21042},
  author = {Kim, Junu and Liu, Xiao and Lin, Zhenghao and Ji, Lei and Gong, Yeyun and Choi, Edward},
  keywords = {Computation and Language (cs.CL), Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {LayerNorm Induces Recency Bias in Transformer Decoders},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution Non Commercial No Derivatives 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/