LayerNorm Induces Recency Bias in Transformer Decoders
Junu KimXiao LiuZheng-Wen LinLei JiYeyun GongEdward Choi
Demonstrates that LayerNorm turns the early-token attention bias of causal transformers into recency bias, resolving a fundamental contradiction in transformer behavior and informing positional encoding design.
Modern sequence modeling architectures, particularly Transformer decoders used in large language models, rely heavily on positional mechanisms to track token order and generalize to long texts. Prior theoretical work suggested that stacking causal self-attention layers inherently causes models to focus on earlier tokens in a sequence. However, empirical observations of real-world Transformer decoders show the opposite effect: models exhibit a recency bias, systematically assigning higher attention weights to more recent tokens. Understanding why this happens without explicit positional encodings is critical for diagnosing model behavior and designing architectures that extrapolate well over long context windows.
This article aims to resolve this contradiction by demonstrating both theoretically and empirically how specific architectural components drive positional behavior. Specifically, it evaluates how layer normalization, residual connections, and non-uniform (anisotropic) input token distributions interact with causal self-attention to induce recency bias.
To conduct this evaluation, the authors combine mathematical proofs with extensive numerical simulations of Transformer decoders lacking learnable parameters and positional encodings. They track attention behavior across varying hidden dimensions, layers, and input anisotropy levels, measuring recency using a recency probability metric—the probability that attention scores assign higher weight to a closer preceding token than to a more distant one across 10 million simulation trials.
The analysis reveals four key findings. First, layer normalization is the core driver of recency bias; stacking causal self-attention with layer normalization provably produces recency bias in the second layer and beyond, whereas stacking self-attention alone yields no recency bias (recency probability remains around 0.50). Second, the directional bias of input token embeddings strongly amplifies this effect. Under realistic anisotropic embeddings (anisotropy factor of 0.5), recency probability jumps substantially (from 0.50 to roughly 0.64 in lower-dimensional second-layer settings and up to nearly 0.91 in deeper layers). Third, residual connections moderate the effect, reducing the recency probability (e.g., from roughly 0.64 down to 0.59 in second-layer tests) by blending in unshared vector components. Fourth, this emergent recency bias extends structurally across modern multi-head and grouped-query attention variants.
These findings imply that Transformer decoders inherently encode positional preferences through standard normalization layers rather than purely through explicit positional encodings. However, the resulting recency bias is non-uniform across sequence positions and fails to satisfy strict relative distance invariance. This uneven positional interaction may actively degrade long-context generalization and extrapolation performance in production language models, creating unintended performance risks for applications requiring balanced attention across lengthy documents.
To address this, engineering and research teams developing foundation models should account for normalization-driven biases when designing positional schemes. A promising next step is to explore mitigation strategies, such as modifying the causal attention mask to counteract uneven positional weighting, before deploying new long-context architectures.
The study’s conclusions are robust within its mathematical and simulation framework, though confidence should be contextualized by noted boundary conditions. The analysis simplifies real-world architectures by omitting feed-forward networks, learnable parameter updates, and modern explicit positional encodings such as rotary position embeddings. Further empirical validation on fully trained, production-scale models is recommended to measure the downstream impact on task accuracy.
- Paper: On Layer Normalization in the Transformer Architecture, Ruibin Xiong et al. (2020). Its analysis of how Pre-LN and Post-LN alter Transformer behavior supplies essential context for understanding the source’s account of normalization-driven attention bias.
No sufficiently relevant recommendations were found.
