On the Emergence of Position Bias in Transformers

Xinyi WuYifei WangStefanie JegelkaAli Jadbabaie

article2025ICML42 citations

Develops a graph-theoretic framework that mathematically explains how multi-layer causal masking and relative positional encodings interact across depth to produce systematic position biases such as attention sinks and the lost-in-the-middle effect in transformers.

Listen

Transformer-based language models frequently exhibit position bias, where the model prioritizes information based on its location in a sequence rather than its semantic relevance. This leads to well-known failure modes such as the "lost-in-the-middle" problem—where retrieval accuracy drops sharply for context placed in the center of long inputs—as well as extreme sensitivity to example ordering in prompts and the formation of uninformative "attention sinks" at initial positions. The article aims to establish a rigorous, graph-theoretic framework to analyze how attention masks and relative positional encodings shape and amplify these positional biases across multi-layer networks.

To evaluate these dynamics, the researchers modeled multi-layer attention mathematically as paths in directed graphs, tracking the cumulative flow of contextual information across successive layers. They paired these theoretical bounds with controlled numerical experiments using synthetic in-context classification tasks. The empirical setup systematically varied model depth, attention masks (causal, sliding-window, and prefix), relative positional encodings (decay masks and rotary positional encoding, or RoPE), and the underlying training data distributions across 10,000 evaluation sequences per condition.

The analysis produced three primary findings. First, causal masking inherently drives positional bias toward the beginning of a sequence: across deep layers, contextual representations exponentially converge toward the first token because initial positions are repeatedly incorporated into intermediate computations. Second, while relative positional encodings introduce distance-based decay that favors recent tokens within individual layers, multi-layer accumulation counteracts this effect; this creates a non-monotonic trade-off where deeper networks still increasingly favor early tokens. Third, experimental tests showed that the causal mask alone cannot simulate positional encodings across arbitrary positions, but instead strictly predisposes the model to early sequence locations, reproducing performance gaps up to 30 to 45 percentage points in favor of early tokens over middle and ending positions.

These findings indicate that position biases are mathematical artifacts inherent to standard transformer designs rather than accidental training anomalies. For practitioners and decision-makers, this highlights an architectural tension: increasing model depth to improve expressive power simultaneously magnifies bias toward initial tokens and reduces model sensitivity to content in the middle or end of contexts. Furthermore, empirical results reveal that the classic "lost-in-the-middle" failure pattern is actively induced when training data disproportionately emphasizes the start and end of sequences, showing that data composition and architectural choices interact directly to influence reliability.

Based on these insights, system designers should avoid assuming that causal masking alone or basic relative encodings eliminate positional vulnerabilities. Teams developing long-context retrieval or reasoning systems should evaluate alternative attention mechanisms—such as prefix masks to distribute attention over initial context blocks or carefully tuned decay parameters to balance local and global focus. Before major architectural overhauls, practitioners should conduct controlled pilot evaluations to assess how token embedding geometry and dataset position distributions impact their specific workloads.

The conclusions are mathematically sound within bounded linear and multi-layer attention formulations, but the experimental validation relies primarily on synthetic classification tasks and simplified attention architectures. While the findings strongly align with observed behaviors in production-scale language models, caution is warranted when extrapolating the exact numerical magnitudes to models with complex component interactions, such as full multilayer perceptron blocks and varied residual connection topologies.

Cover for On the Emergence of Position Bias in Transformers

Abstract

Recent studies have revealed various manifestations of position bias in transformer architectures, from the “lost-in-the-middle” phenomenon to attention sinks, yet a comprehensive theoretical understanding of how attention masks and positional encodings shape these biases remains elusive. This paper presents a graph-theoretic framework for analyzing position bias in multi-layer attention. Modeling attention masks as directed graphs, we quantify how tokens interact with contextual information based on their sequential positions. We uncover two key insights: First, causal masking inherently biases attention toward earlier positions, as tokens in deeper layers attend to increasingly more contextualized representations of earlier tokens. Second, we characterize the competing effects of the causal mask and relative positional encodings, such as the decay mask and rotary positional encoding (RoPE): while both mechanisms introduce distance-based decay within individual attention maps, their aggregate effect across multiple attention layers—coupled with the causal mask—leads to a trade-off between the long-term decay effects and the cumulative importance of early sequence positions. Through controlled numerical experiments, we not only validate our theoretical findings but also reproduce position biases observed in real-world LLMs. Our framework offers a principled foundation for understanding positional biases in transformers, shedding light on the complex interplay of attention mechanism components and guiding more informed architectural design.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Setup
  • 3.1 (Masked) Attention Mechanism
  • 3.2 Attention Update
  • 3.3 Relative Positional Encoding
  • 4 Main Results
  • 4.1 Attention Masks: A Graph-Theoretic View
  • 4.2 Relative PEs: A Competing Decay Effect
  • 4.3 A Closer Look at RoPE
  • 5 Experiments
  • 5.1 The Effects of Depth and Relative PEs
  • 5.2 Can Causal Mask Induce Usable Positional Information?
  • 6 Conclusion
  • References
  • A Proof of
  • A.1 Auxiliary results
  • A.2 Proof of
  • B Proof of
  • B.1 Auxiliary results
  • B.2 Proof of
  • C Proof of
  • D Proof of
  • E Proof of
  • F Proof of
  • G Proof of
  • H Implicit differentiation of xx with respect to θ1\theta_{1} and tt
  • I The effect of RoPE: case for d≥2d\geq 2
  • I.1 Results
  • I.2 Proofs of
  • I.3 Proof of
  • J Experiments
  • K Additional Experimental Results
  • K.1 The effect of depth and relative PEs
  • K.2 The effect of residual connections on position bias
  • K.3 The role of training data on positional bias
  • K.4 Attention sinks
  • L Additional Results under a Fixed-Vocabulary Setting

Knowls

  1. Knowl 1 — Graph-based attention rollout framework

    model/method

    The paper analyzes a single-head, attention-only network with NN token representations X(t)∈RN×dX^{(t)}\in\mathbb{R}^{N\times d} at layer tt. A fixed directed graph GG encodes the attention mask: an edge (j,i)(j,i) means token ii may attend to token jj. For query and key matrices WQ(t),WK(t)∈Rd×dW_Q^{(t)},W_K^{(t)}\in\mathbb{R}^{d\times d} and value matrix WV(t)∈Rd×dW_V^{(t)}\in\mathbb{R}^{d\times d}, the masked attention matrix A(t)A^{(t)} is the row-wise softmax of X(t)WQ(t)(X(t)WK(t))⊤X^{(t)}W_Q^{(t)}(X^{(t)}W_K^{(t)})^\top over allowed edges, with zero entries elsewhere, and X(t+1)=A(t)X(t)WV(t)X^{(t+1)}=A^{(t)}X^{(t)}W_V^{(t)}. The analysis assumes uniformly bounded query and key matrix operator norms, sup⁡tmax⁡(∥WQ(t)∥2,∥WK(t)∥2)≤C\sup_t\max(\|W_Q^{(t)}\|_2,\|W_K^{(t)}\|_2)\le C, and bounded operator norms of the cumulative value-matrix products WV(0)⋯WV(k)W_V^{(0)}\cdots W_V^{(k)} for all k≥0k\ge0.

    The cumulative attention rollout is P(t)=A(t)A(t−1)⋯A(0)P^{(t)}=A^{(t)}A^{(t-1)}\cdots A^{(0)}. Its entry Pij(t)P^{(t)}_{ij} is the probability that token ii's context at depth tt traces back to input token jj; equivalently, it tracks the total direct and indirect attention flow through all paths permitted by the mask. These assumptions and results concern attention-only dynamics; residual connections and MLPs are not part of the theoretical model.

  2. Knowl 2 — Causal attention concentrates context on the first token

    theoretical result

    Let NN tokens use a causal attention mask, so token ii attends only to tokens j≤ij\le i. Under the bounded-weight assumptions stated for the graph-based attention model, for any initial representations and every target token i∈{1,…,N}i\in\{1,\ldots,N\}, cumulative attention rollout converges to the first input token as depth grows:

    lim⁡t→∞Pi1(t)=1.\lim_{t\to\infty}P^{(t)}_{i1}=1.

    Moreover, there exist constants C,ϵ∈(0,1)C,\epsilon\in(0,1) with Nϵ<1N\epsilon<1 such that, for every 1<j≤i1<j\le i and integer depth t≥0t\ge0,

    Pij(t)≤C(1−(j−1)ϵ)t.P^{(t)}_{ij}\le C\bigl(1-(j-1)\epsilon\bigr)^t.

    Thus the causal mask alone creates an increasingly strong preference for early context, independently of the semantic content of the tokens. The paper attributes the effect to repeated attention over already contextualized representations: early tokens influence later tokens both directly and through intermediate positions.

  3. Knowl 3 — Sliding-window and prefix masks determine the limiting context positions

    theoretical result

    Under the same bounded-weight assumptions as the graph-based attention model, alternative masks change which early positions accumulate context. For a sliding-window mask of width w≥2w\ge2, let r=⌈(N−1)/(w−1)⌉r=\lceil(N-1)/(w-1)\rceil. For every target token ii, rollout converges to the first token, lim⁡t→∞Pi1(t)=1\lim_{t\to\infty}P^{(t)}_{i1}=1. There are constants C,ϵ∈(0,1)C,\epsilon\in(0,1) with Nϵr<1N\epsilon^r<1 such that for 1<j≤i1<j\le i and t≥0t\ge0,

    Pij(t)≤C(1−(j−1)ϵr)t/(2r).P^{(t)}_{ij}\le C\bigl(1-(j-1)\epsilon^r\bigr)^{t/(2r)}.

    A narrower window moderates the convergence rate but does not prevent eventual concentration on the first token.

    For a prefix mask whose first KK tokens form the prefix, rollout instead concentrates on the prefix: lim⁡t→∞∑k=1KPik(t)=1\lim_{t\to\infty}\sum_{k=1}^{K}P^{(t)}_{ik}=1 for every ii. Each prefix token retains nonzero limiting influence: there is a constant κ>0\kappa>0 such that lim inf⁡t→∞Pik(t)≥κ\liminf_{t\to\infty}P^{(t)}_{ik}\ge\kappa for every k≤Kk\le K. For non-prefix positions, constants C,ϵ∈(0,1)C,\epsilon\in(0,1) exist such that Pij(t)≤C(1−(j−K)ϵ)tP^{(t)}_{ij}\le C(1-(j-K)\epsilon)^t for K<j≤iK<j\le i. These results identify the first token or the prefix tokens as the positions that ultimately dominate context under the respective masks.

  4. Knowl 4 — A decay mask imposes exponential distance decay within each layer

    theoretical result

    Consider causal attention with a decay-mask bias Dij=−(i−j)mD_{ij}=-(i-j)m for j≤ij\le i, where m>0m>0 is the decay strength, and let Adecay(t)A^{(t)}_{\mathrm{decay}} be the resulting row-softmax attention matrix. Under the bounded-weight assumptions of the graph-based attention model, there are constants Cmin⁡,Cmax⁡>0C_{\min},C_{\max}>0, independent of token positions and layer, such that for all 1≤j≤i≤N1\le j\le i\le N and t≥0t\ge0,

    Cmin⁡e−(i−j)m≤(Adecay(t))ij≤Cmax⁡e−(i−j)m.C_{\min}e^{-(i-j)m}\le (A^{(t)}_{\mathrm{decay}})_{ij}\le C_{\max}e^{-(i-j)m}.

    Thus each individual attention map gives exponentially less weight to more distant earlier tokens, with the rate set by mm.

  5. Knowl 5 — Layerwise decay and cumulative causal paths yield non-monotonic distance preference

    theoretical result

    With causal attention and the decay mask of strength m>0m>0, consider the rollout probability Pdecay,ij(t)P^{(t)}_{\mathrm{decay},ij} from source token jj to target token ii after t+1t+1 attention layers. Under the bounded-weight assumptions and for any fixed maximum depth TT, the paper establishes, for j≤ij\le i and 0≤t≤T0\le t\le T,

    Pdecay,ij(t)=Θ ⁣((t+i−ji−j)e−(i−j)m).P^{(t)}_{\mathrm{decay},ij}=\Theta\!\left(\binom{t+i-j}{i-j}e^{-(i-j)m}\right).

    Here Θ\Theta denotes equality up to positive multiplicative bounds in the stated finite-depth regime. The binomial factor counts the cumulative causal paths, while the exponential factor is the distance penalty accumulated along them. Under Stirling’s approximation, the log of this expression has its maximum at distance x∗=t/(em−1)x^*=t/(e^m-1), where x=i−jx=i-j. Increasing mm moves the preferred distance nearer to the target; increasing depth tt moves it farther away, toward earlier positions. Consequently, the per-layer monotone decay can become a non-monotonic aggregate preference across layers.

  6. Knowl 6 — RoPE induces a weak Gaussian-like distance decay in individual layers

    theoretical result

    For causal attention using RoPE, restrict the analysis to the two-dimensional feature pair with the slowest base rotation angle θ1\theta_1. Let qi(t)=Xi,:(t)WQ(t)q_i^{(t)}=X^{(t)}_{i,:}W_Q^{(t)} and kj(t)=Xj,:(t)WK(t)k_j^{(t)}=X^{(t)}_{j,:}W_K^{(t)} be the unrotated query and key, and let ϕij(t)\phi_{ij}^{(t)} be the angle between them. Assume the bounded-weight conditions, nonzero query and key norms, ∣ϕij(t)∣≤δθ1|\phi_{ij}^{(t)}|\le\delta\theta_1 for some δ>0\delta>0, and (δ+N−1)θ1≤π(\delta+N-1)\theta_1\le\pi. Then positive constants Cmin⁡,Cmax⁡,c,c′C_{\min},C_{\max},c,c' exist such that for 1≤j≤i≤N1\le j\le i\le N,

    Cmin⁡e−c(i−j)2θ12≤(ARoPE(t))ij≤Cmax⁡e−c′(i−j)2θ12.C_{\min}e^{-c(i-j)^2\theta_1^2}\le (A^{(t)}_{\mathrm{RoPE}})_{ij}\le C_{\max}e^{-c'(i-j)^2\theta_1^2}.

    In this slow-rotation regime, RoPE therefore creates a distance-dependent decay within each attention layer. The paper characterizes it as more gradual than decay-mask attenuation when the base angle is small.

  7. Knowl 7 — Across layers, RoPE decay competes with causal path accumulation

    theoretical result

    Under the causal mask and the slowest-rotating two-dimensional RoPE conditions stated for the single-layer result, fix a maximum depth T>0T>0. For rollout from token jj to token ii at any integer depth 0≤t≤T0\le t\le T, the paper obtains a constant c>0c>0 such that

    PRoPE,ij(t)=Θ ⁣((t+i−ji−j)e−c(i−j)2θ12),j≤i.P^{(t)}_{\mathrm{RoPE},ij}=\Theta\!\left(\binom{t+i-j}{i-j}e^{-c(i-j)^2\theta_1^2}\right),\qquad j\le i.

    The binomial term represents the number of causal paths and favors greater distances as depth grows; the exponential term penalizes distance, with strength set by the RoPE base angle θ1\theta_1. Under the paper’s Stirling-based characterization of the maximizing distance, the preferred distance increases with depth and decreases as θ1\theta_1 increases. Thus stronger RoPE distance decay favors nearby context, while deeper attention shifts aggregate influence toward earlier positions.

  8. Knowl 8 — Controlled retrieval experiments show depth and relative encodings alter position bias

    empirical result

    The experiments use an in-context retrieval task: a sequence contains n=8n=8 item-label pairs followed by a query item, and a classifier predicts the query’s label. Items are generated from a Gaussian mixture with K=2048K=2048 classes and L=32L=32 labels; the reported setup uses embedding dimension d=64d=64, γ=0.75\gamma=0.75, and burstiness B=4B=4. A single-head attention-only network feeds a three-layer ReLU MLP classifier. Models use AdamW with learning rate 10−310^{-3}, weight decay 10−610^{-6}, batch size 128, and 100,000 training iterations. The decay-mask strength is m=−log⁡(0.8)≈0.223m=-\log(0.8)\approx0.223. Position bias is measured on 10,000 paired test sequences per comparison: the two sequences contain identical vectors at positions aa and bb, but swap which position carries the correct answer. The accuracy gap [a,b]−[b,a][a,b]-[b,a] is positive when the model favors the earlier of the compared positions.

    With position-unbiased training data and no residual connections, increasing attention depth from 2 to 6 layers increased the first-versus-middle accuracy gap from 0.0590.059 to 0.0910.091 without positional encoding, and the first-versus-last gap from 0.0770.077 to 0.1060.106. For RoPE, the first-versus-middle gap rose from 0.0020.002 at depth 2 to 0.0260.026 at depth 6; for the decay mask it shifted from −0.055-0.055 to 0.0200.020. Across the tested positional encodings, deeper pure-attention models increased preference for earlier positions, while relative encodings shifted attention toward more recent tokens; the recent-token effect was more pronounced for the decay mask than for RoPE. With residual connections, the relationship between depth and bias was non-monotonic and depended on the encoding and depth regime.

    The authors also tested fixed-vocabulary inputs with anisotropic class vectors (λ=0.75\lambda=0.75). Most trends persisted, and depth consistently amplified bias in the tested residual-connected models. That setting made retrieval more difficult than the Gaussian-mixture setting, so the paper notes that embedding geometry may affect both task difficulty and observed positional bias.

  9. Knowl 9 — Training-position structure affects learned bias and the lost-in-the-middle pattern

    empirical result

    In a two-layer residual-connected retrieval model trained on sequences biased equally toward the first or last position, the authors compared causal versus unmasked attention and no positional encoding, sinusoidal absolute encoding, or RoPE. They averaged five runs and measured accuracy gaps on 10,000 test sequences per comparison; a positive gap favors the earlier position. With no positional encoding, causal attention produced first-versus-middle and first-versus-last gaps of 0.2760.276 and 0.3080.308, while the corresponding unmasked gaps were −0.002-0.002 and 0.0000.000. The causal mask without positional encoding therefore introduced an early-position preference, but did not provide the flexible positional signal needed to learn arbitrary location-specific biases. With sinusoidal encoding, the causal gaps were 0.3160.316 (first versus middle) and 0.0780.078 (first versus last); with RoPE they were 0.4720.472 and 0.4330.433. The positional encodings allowed the model to capture biases at both trained endpoints.

    In the tested synthetic conditions, the lost-in-the-middle pattern—lower retrieval performance for middle positions than for sequence ends—appeared when training data favored both the first and last positions, but not when training data lacked positional bias or favored only one location. The paper therefore finds that learned position preferences depend on both the architecture and the positional structure of the training data; it does not establish that this synthetic pattern explains lost-in-the-middle behavior in real language models.

  10. Knowl 10 — Mask-graph center positions align with observed attention sinks

    empirical result

    A center node in an attention-mask graph is a token from which every token is reachable by a directed path. The theoretical mask results identify the first token as the center for causal and sliding-window masks, and the prefix tokens as the center set for a prefix mask; cumulative context converges toward those positions. In controlled two-layer experiments with position-unbiased sequences, the authors observed attention sinks at the absolute first token under causal attention, a tendency for the first token to be a sink under sliding-window attention (especially for larger windows), and sinks on the prefix tokens under prefix attention rather than only on the first token. Sink measurements used a threshold of τ=0.2\tau=0.2 over 10,000 sequences. These observations support the paper’s interpretation that mask structure helps determine sink locations, while the theoretical convergence result concerns cumulative attention rollout rather than directly proving the empirical sink statistic.

Coverage note — No substantial contributed material has been omitted. The principal theory, mask variants, relative-position results, controlled retrieval findings, and attention-sink observations are represented; scope restrictions on the theory and the synthetic nature of the experiments are stated within the relevant knowls.

References

  1. 1.Abnar, S. and Zuidema, W. Quantifying attention flow in transformers. In ACL, 2020.
  2. 2.Ait-Saada, M. and Nadif, M. Is anisotropy truly harmful? a case study on text clustering. In ACL, 2023.
  3. 3.Alman, J. and Song, Z. Fast attention requires bounded entries. In NeurIPS, 2023.
  4. 4.Bahdanau, D., Cho, K., and Bengio, Y. Neural machine translation by jointly learning to align and translate. In ICLR, 2015.
  5. 5.Barbero, F., Banino, A., Kapturowski, S., Kumaran, D., Ara’ujo, J. G., Vitvitskyi, A., Pascanu, R., and Velivckovi’c, P. Transformers need glasses! information over-squashing in language tasks. In NeurIPS, 2024a.
  6. 6.Barbero, F., Vitvitskyi, A., Perivolaropoulos, C., Pascanu, R., and Velivckovi’c, P. Round and round we go! what makes rotary positional encodings useful? ArXiv, abs/2410.06205, 2024b.
  7. 7.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. ArXiv, 2020.
  8. 8.Cai, X., Huang, J., Bian, Y.-L., and Church, K. W. Isotropy in the contextual embedding space: Clusters and manifolds. In ICLR, 2021.
  9. 9.Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. I. Kerple: Kernelized relative positional embedding for length extrapolation. In NeurIPS, 2022.
  10. 10.Dubey, A. and et al. The llama 3 herd of models. ArXiv, abs/2407.21783, 2024.
  11. 11.Ethayarajh, K. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and gpt-2 embeddings. In EMNLP, 2019.
  12. 12.Fang, L., Yifei Wang, K. G., Fang, L., and Wang, Y. Rethinking invariance in in-context learning. In ICLR, 2025.
  13. 13.Gao, J., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y. Representation degeneration problem in training natural language generation models. In ICLR, 2019.
  14. 14.Glanzer, M. and Cunitz, A. R. Two storage mechanisms in free recall. Journal of Verbal Learning and Verbal Behavior, 1966.
  15. 15.Godey, N., de la Clergerie, E. V., and Sagot, B. Anisotropy is inherent to self-attention in transformers. ArXiv, abs/2401.12143, 2024.
  16. 16.Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., and Lin, M. When attention sink emerges in language models: An empirical view. In ICLR, 2025.
  17. 17.Guo, T., Pai, D., Bai, Y., Jiao, J., Jordan, M. I., and Mei, S. Active-dormant attention heads: Mechanistically demystifying extreme-token phenomena in llms. 2024.
  18. 18.Guo, X. and Vosoughi, S. Serial position effects of large language models. ArXiv, 2024.
  19. 19.Halliday, M. A. An Introduction to Functional Grammar. 2004.
  20. 20.Hollenstein, N., Pirovano, F., Zhang, C., Jager, L. A., and Beinborn, L. Multilingual language models predict human reading behavior. In NAACL, 2021.
  21. 21.Hou, Y., Zhang, J., Lin, Z., Lu, H., Xie, R., McAuley, J., and Zhao, W. X. Large language models are zero-shot rankers for recommender systems. In ECIR, 2024.
  22. 22.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b. ArXiv, 2023.
  23. 23.Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S. The impact of positional encoding on length generalization in transformers. In NeurIPS, 2023.
  24. 24.Kim, Y., Denton, C., Hoang, L., and Rush, A. M. Structured attention networks. In ICLR, 2017.
  25. 25.Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K. Attention is not only a weight: Analyzing transformers with vector norms. In EMNLP, 2020.
  26. 26.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., rahman Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL, 2020.
  27. 27.Li, Y., Pazdera, J. K., and Kahana, M. J. Eeg decoders track memory dynamics. Nature Communications, 2024a.
  28. 28.Li, Z., Liu, H., Zhou, D., and Ma, T. Chain of thought empowers transformers to solve inherently serial problems. ArXiv, abs/2402.12875, 2024b.
  29. 29.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 2024.
  30. 30.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In ICLR, 2019.
  31. 31.Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In ACL, 2022.
  32. 32.Merrill, W. and Sabharwal, A. The parallelism tradeoff: Limitations of log-precision transformers. Transactions of the Association for Computational Linguistics, 11:531–545, 2022.
  33. 33.Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. Rethinking the role of demonstrations: What makes in-context learning work? In EMNLP, 2022.
  34. 34.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
  35. 35.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. In ICLR, 2022.
  36. 36.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020.
  37. 37.Reddy, G. The mechanistic basis of data dependence and abrupt learning in an in-context classification task. In ICLR, 2024.
  38. 38.Sanford, C., Hsu, D., and Telgarsky, M. One-layer transformers fail to solve the induction heads task. ArXiv, abs/2408.14332, 2024.
  39. 39.Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. 2023.
  40. 40.Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In NeurIPS, 2017.
  41. 41.Wang, Z., Zhang, H., Li, X., Huang, K.-H., Han, C., Ji, S., Kakade, S. M., Peng, H., and Ji, H. Eliminating position bias of language models: A mechanistic approach. ArXiv, abs/2407.01100, 2024.
  42. 42.Wu, X., Ajorlou, A., Wu, Z., and Jadbabaie, A. Demystifying oversmoothing in attention-based graph neural networks. In NeurIPS, 2023.
  43. 43.Wu, X., Ajorlou, A., Wang, Y., Jegelka, S., and Jadbabaie, A. On the role of attention masks and layernorm in transformers. In NeurIPS, 2024.
  44. 44.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. In ICLR, 2024.
  45. 45.Yu, Y., Jiang, H., Luo, X., Wu, Q., Lin, C.-Y., Li, D., Yang, Y., Huang, Y., and Qiu, L. Mitigate position bias in large language models via scaling a single dimension. ArXiv, abs/2406.02536, 2024.
  46. 46.Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S. Are transformers universal approximators of sequence-to-sequence functions? In ICLR, 2020a.
  47. 47.Yun, C., Chang, Y.-W., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S. O(n) connections are expressive enough: Universal approximability of sparse transformers. In NeurIPS, 2020b.
  48. 48.Zhang, Z. A., Chen, R., Liu, S., Yao, Z., Ruwase, O., Chen, B., Wu, X., and Wang, Z. Found in the middle: How language models use long contexts better via plug-and-play positional encoding. ArXiv, abs/2403.04797, 2024.
  49. 49.Zhao, T. Z., Wallace, E., Feng, S., Klein, D., and Singh, S. Calibrate before use: Improving few-shot performance of language models. In ICML, 2021.
  50. 50.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS, 2023.

Citation

MLA
Wu, X., et al. “On the Emergence of Position Bias in Transformers”. arXiv, 2025, http://arxiv.org/abs/2502.01951v4.
APA
Wu, X., Wang, Y., Jegelka, S., & Jadbabaie, A. (2025). On the Emergence of Position Bias in Transformers. arXiv. http://arxiv.org/abs/2502.01951v4
Chicago
Wu, X., Y. Wang, S. Jegelka, and A. Jadbabaie. 2025. “On the Emergence of Position Bias in Transformers”. arXiv. http://arxiv.org/abs/2502.01951v4.
Harvard
Wu, X. et al. (2025) “On the Emergence of Position Bias in Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.01951v4.
Vancouver
1. Wu X, Wang Y, Jegelka S, Jadbabaie A (2025) On the Emergence of Position Bias in Transformers. arXiv

BibTeX

@article{wu2025the,
  title = {On the Emergence of Position Bias in Transformers},
  author = {Wu, Xinyi and Wang, Yifei and Jegelka, Stefanie and Jadbabaie, Ali},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.01951v4},
  eprint = {2502.01951}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/