On the Emergence of Position Bias in Transformers
Xinyi WuYifei WangStefanie JegelkaAli Jadbabaie
Develops a graph-theoretic framework that mathematically explains how multi-layer causal masking and relative positional encodings interact across depth to produce systematic position biases such as attention sinks and the lost-in-the-middle effect in transformers.
Transformer-based language models frequently exhibit position bias, where the model prioritizes information based on its location in a sequence rather than its semantic relevance. This leads to well-known failure modes such as the "lost-in-the-middle" problem—where retrieval accuracy drops sharply for context placed in the center of long inputs—as well as extreme sensitivity to example ordering in prompts and the formation of uninformative "attention sinks" at initial positions. The article aims to establish a rigorous, graph-theoretic framework to analyze how attention masks and relative positional encodings shape and amplify these positional biases across multi-layer networks.
To evaluate these dynamics, the researchers modeled multi-layer attention mathematically as paths in directed graphs, tracking the cumulative flow of contextual information across successive layers. They paired these theoretical bounds with controlled numerical experiments using synthetic in-context classification tasks. The empirical setup systematically varied model depth, attention masks (causal, sliding-window, and prefix), relative positional encodings (decay masks and rotary positional encoding, or RoPE), and the underlying training data distributions across 10,000 evaluation sequences per condition.
The analysis produced three primary findings. First, causal masking inherently drives positional bias toward the beginning of a sequence: across deep layers, contextual representations exponentially converge toward the first token because initial positions are repeatedly incorporated into intermediate computations. Second, while relative positional encodings introduce distance-based decay that favors recent tokens within individual layers, multi-layer accumulation counteracts this effect; this creates a non-monotonic trade-off where deeper networks still increasingly favor early tokens. Third, experimental tests showed that the causal mask alone cannot simulate positional encodings across arbitrary positions, but instead strictly predisposes the model to early sequence locations, reproducing performance gaps up to 30 to 45 percentage points in favor of early tokens over middle and ending positions.
These findings indicate that position biases are mathematical artifacts inherent to standard transformer designs rather than accidental training anomalies. For practitioners and decision-makers, this highlights an architectural tension: increasing model depth to improve expressive power simultaneously magnifies bias toward initial tokens and reduces model sensitivity to content in the middle or end of contexts. Furthermore, empirical results reveal that the classic "lost-in-the-middle" failure pattern is actively induced when training data disproportionately emphasizes the start and end of sequences, showing that data composition and architectural choices interact directly to influence reliability.
Based on these insights, system designers should avoid assuming that causal masking alone or basic relative encodings eliminate positional vulnerabilities. Teams developing long-context retrieval or reasoning systems should evaluate alternative attention mechanisms—such as prefix masks to distribute attention over initial context blocks or carefully tuned decay parameters to balance local and global focus. Before major architectural overhauls, practitioners should conduct controlled pilot evaluations to assess how token embedding geometry and dataset position distributions impact their specific workloads.
The conclusions are mathematically sound within bounded linear and multi-layer attention formulations, but the experimental validation relies primarily on synthetic classification tasks and simplified attention architectures. While the findings strongly align with observed behaviors in production-scale language models, caution is warranted when extrapolating the exact numerical magnitudes to models with complex component interactions, such as full multilayer perceptron blocks and varied residual connection topologies.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). Read this account of relative position representations first to understand the positional mechanisms whose interactions with attention masks the source analyzes.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Its graph-based attention rollout and flow methods provide useful groundwork for the source’s analysis of information pathways across Transformer layers.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). Its introduction of ALiBi’s distance-dependent attention bias gives concrete context for the source’s analysis of positional decay and long-range attention.
- Paper: LayerNorm Induces Recency Bias in Transformer Decoders, Junu Kim et al. (2025). It extends the source’s structural account by showing how layer normalization and residual connections can produce recency bias in practical decoder settings.
- Paper: Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation, Zichong Li et al. (2026). Building on the source’s diagnosis of positional brittleness, it tests a RoPE-perturbation training method to improve retrieval across input positions.
- Paper: RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings, Jarod Lévy et al. (2026). It carries the source’s analysis of positional-encoding trade-offs into a data-aware RoPE design aimed at improving long-context retrieval and extrapolation.
