Built independently by an author, for readers. Read the story and support ChapterPal

keyword

positional encodings

Positional encodings are mathematical representations added to token vectors in neural network architectures, such as transformers, to convey the sequential order or spatial arrangement of elements within an input sequence. Because standard self-attention mechanisms process all tokens simultaneously and are inherently permutation-invariant, models cannot naturally distinguish the relative or absolute placement of words or symbols without auxiliary information. Positional encodings address this limitation by injecting order-aware signals into the network, implemented either as absolute coordinate vectors assigned to specific token indices, relative distance metrics computed between pairs of tokens, or geometric operations like rotary embeddings applied directly to query and key representations. These encodings can be generated using deterministic mathematical functions, learned as trainable parameters, or incorporated as structural attention biases, enabling models to accurately capture syntactic, semantic, and contextual relationships across sequential data.

16 items

On the Emergence of Position Bias in Transformers

On the Emergence of Position Bias in Transformers

Xinyi Wu, Yifei Wang, Stefanie Jegelka, Ali Jadbabaie

OrganizationsMassachusetts Institute of TechnologyTechnical University of Munich

Why you should read this

Develops a graph-theoretic framework that mathematically explains how multi-layer causal masking and relative positional encodings interact across depth to produce systematic position biases such as attention sinks and the lost-in-the-middle effect in transformers.

Recent studies have revealed various manifestations of position bias in transformer architectures, from the “lost-in-the-middle” phenomenon to attention sinks, yet a comprehensive theoretical understanding of how attention masks and positional encodings shape these biases remains elusive. This paper presents a graph-theoretic framework for analyzing position bias in multi-layer attention. Modeling attention masks as directed graphs, we quantify how tokens interact with contextual information based on their sequential positions. We uncover two key insights: First, causal masking inherently biases attention toward earlier positions, as tokens in deeper layers attend to increasingly more contextualized representations of earlier tokens. Second, we characterize the competing effects of the causal mask and relative positional encodings, such as the decay mask and rotary positional encoding (RoPE): while both mechanisms introduce distance-based decay within individual attention maps, their aggregate effect across multiple attention layers—coupled with the causal mask—leads to a trade-off between the long-term decay effects and the cumulative importance of early sequence positions. Through controlled numerical experiments, we not only validate our theoretical findings but also reproduce position biases observed in real-world LLMs. Our framework offers a principled foundation for understanding positional biases in transformers, shedding light on the complex interplay of attention mechanism components and guiding more informed architectural design.

Added

2026-10-05

Word Order Does Matter and Shuffled Language Models Know It

Word Order Does Matter and Shuffled Language Models Know It

Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, Anders Søgaard

OrganizationsUniversity of CopenhagenUniversity of OsloUppsala University

Why you should read this

Reveals how language models trained on scrambled text still recover natural word order through subword segmentation artifacts and statistical dependencies between sentence length and token frequencies, explaining why position embeddings remain vital even under shuffled pre-training.

Recent studies have shown that language models pretrained and/or fine-tuned on randomly permuted sentences exhibit competitive performance on GLUE, putting into question the importance of word order information. Somewhat counter-intuitively, some of these studies also report that position embeddings appear to be crucial for models’ good performance with shuffled text. We probe these language models for word order information and investigate what position embeddings learned from shuffled text encode, showing that these models retain information pertaining to the original, naturalistic word order. We show this is in part due to a subtlety in how shuffling is implemented in previous work – before rather than after subword segmentation. Surprisingly, we find even Language models trained on text shuffled after subword segmentation retain some semblance of information about word order because of the statistical dependencies between sentence length and unigram probabilities. Finally, we show that beyond GLUE, a variety of language understanding tasks do require word order information, often to an extent that cannot be learned through fine-tuning.

Added

2026-10-03

Why are Sensitive Functions Hard for Transformers?

Why are Sensitive Functions Hard for Transformers?

Michael Hahn, Mark Rofin

OrganizationsSaarland Informatics CampusSaarland University

Why you should read this

Proves that transformers computing highly sensitive functions occupy extremely sharp, isolated parameter regions, mathematically explaining why these models inherently struggle to learn and generalize functions like PARITY despite having the expressive capacity to represent them.

Empirical studies have identified a range of learnability biases and limitations of transformers, such as a persistent difficulty in learning to compute simple formal languages such as PARITY, and a bias towards low-degree functions. However, theoretical understanding remains limited, with existing expressiveness theory either overpredicting or underpredicting realistic learning abilities. We prove that, under the transformer architecture, the loss landscape is constrained by the input-space sensitivity: Transformers whose output is sensitive to many parts of the input string inhabit isolated points in parameter space, leading to a low-sensitivity bias in generalization. We show theoretically and empirically that this theory unifies a broad array of empirical observations about the learning abilities and biases of transformers, such as their generalization bias towards low sensitivity and low degree, and difficulty in length generalization for PARITY. This shows that understanding transformers' inductive biases requires studying not just their in-principle expressivity, but also their loss landscape.

Added

2026-10-03

Overcoming a Theoretical Limitation of Self-Attention

Overcoming a Theoretical Limitation of Self-Attention

David Chiang, Peter Cholak

OrganizationsUniversity of Notre Dame

Why you should read this

Proves that standard transformers can recognize challenging regular languages like PARITY with perfect accuracy and shows that scaling attention logits by the logarithm of sequence length resolves severe length generalization failures in practice.

Although transformers are remarkably effective for many tasks, there are some surprisingly easy-looking regular languages that they struggle with. Hahn shows that for languages where acceptance depends on a single input symbol, a transformer's classification decisions become less and less confident (that is, with cross-entropy approaching 1 bit per string) as input strings get longer and longer. We examine this limitation using two languages: PARITY, the language of bit strings with an odd number of 1s, and FIRST, the language of bit strings starting with a 1. We demonstrate three ways of overcoming the limitation suggested by Hahn's lemma. First, we settle an open question by constructing a transformer that recognizes PARITY with perfect accuracy, and similarly for FIRST. Second, we use layer normalization to bring the cross-entropy of both models arbitrarily close to zero. Third, when transformers need to focus on a single position, as for FIRST, we find that they can fail to generalize to longer strings; we offer a simple remedy to this problem that also improves length generalization in machine translation.

Added

2026-10-01

Looped Transformers as Programmable Computers

Looped Transformers as Programmable Computers

Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D. Lee, Dimitris Papailiopoulos

OrganizationsPrinceton UniversityUniversity of Wisconsin MadisonYonsei University

Why you should read this

Demonstrates how constant-depth looped transformers can function as universal computers by executing instruction sets directly from input prompts to run iterative algorithms, linear algebra operations, and in-context backpropagation.

We present a framework for using transformer networks as universal computers by programming them with specific weights and placing them in a loop. Our input sequence acts as a punch-card, consisting of instructions and memory for data read/writes. We demonstrate that a constant number of encoder layers can emulate basic computing blocks, including lexicographic operations, non-linear functions, function calls, program counters, and conditional branches. Using this framework, we emulate a computer using a simple instruction-set architecture, which allows us to map iterative algorithms to programs that can be executed by a constant depth looped transformer network. We show how a single frozen transformer, instructed by its input, can emulate a basic calculator, a basic linear algebra library, and even a full backpropagation, in-context learning algorithm. Our findings reveal the potential of transformer networks as programmable compute units and offer insight into the mechanics of attention. 3

Added

2026-09-30

Randomized Positional Encodings Boost Length Generalization of Transformers

Randomized Positional Encodings Boost Length Generalization of Transformers

Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, Joel Veness

OrganizationsDalle Molle Institute for Artificial Intelligence ResearchGoogle

Why you should read this

Introduces a randomized positional encoding scheme that subsamples ordered positions from an extended range during training, enabling Transformers trained on short sequences to generalize effectively to unseen sequence lengths across algorithmic reasoning tasks.

Transformers have impressive generalization capabilities on tasks with a fixed context length. However, they fail to generalize to sequences of arbitrary length, even for seemingly simple tasks such as duplicating a string. Moreover, simply training on longer sequences is inefficient due to the quadratic computation complexity of the global attention mechanism. In this work, we demonstrate that this failure mode is linked to positional encodings being out-of-distribution for longer sequences (even for relative encodings) and introduce a novel family of positional encodings that can overcome this problem. Concretely, our randomized positional encoding scheme simulates the positions of longer sequences and randomly selects an ordered subset to fit the sequence's length. Our large-scale empirical evaluation of 6000 models across 15 algorithmic reasoning tasks shows that our method allows Transformers to generalize to sequences of unseen length (increasing test accuracy by 12.0% on average).

Added

2026-09-26

Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers

Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurélien Lucchi, Thomas Hofmann

OrganizationsCSEM SAETH ZurichUniversity of Basel

Why you should read this

Proposes a learnable context-pruning mechanism for autoregressive language models that discards up to 80% of uninformative tokens during generation to cut inference memory and double throughput without hurting downstream accuracy.

Autoregressive Transformers adopted in Large Language Models (LLMs) are hard to scale to long sequences. Despite several works trying to reduce their computational cost, most of LLMs still adopt attention layers between all pairs of tokens in the sequence, thus incurring a quadratic cost. In this study, we present a novel approach that dynamically prunes contextual information while preserving the model’s expressiveness, resulting in reduced memory and computational requirements during inference. Our method employs a learnable mechanism that determines which uninformative tokens can be dropped from the context at any point across the generation process. By doing so, our approach not only addresses performance concerns but also enhances interpretability, providing valuable insight into the model’s decision-making process. Our technique can be applied to existing pre-trained models through a straightforward fine-tuning process, and the pruning strength can be specified by a sparsity parameter. Notably, our empirical findings demonstrate that we can effectively prune up to 80% of the context without significant performance degradation on downstream tasks, offering a valuable tool for mitigating inference costs. Our reference implementation achieves up to 2× increase in inference throughput and even greater memory savings.

Added

2026-09-26

DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, Lei Zhang

OrganizationsInternational Digital Economy AcademyPeng Cheng LaboratoryThe Hong Kong University of Science and TechnologyTsinghua University

Why you should read this

Proposes using dynamic anchor boxes as queries in detection transformers to accelerate training convergence and improve object detection accuracy on COCO through explicit spatial priors and scale-aware positional attention updated layer by layer.

We present in this paper a novel query formulation using dynamic anchor boxes for DETR (DEtection TRansformer) and offer a deeper understanding of the role of queries in DETR. This new formulation directly uses box coordinates as queries in Transformer decoders and dynamically updates them layer-by-layer. Using box coordinates not only helps using explicit positional priors to improve the query-to-feature similarity and eliminate the slow training convergence issue in DETR, but also allows us to modulate the positional attention map using the box width and height information. Such a design makes it clear that queries in DETR can be implemented as performing soft ROI pooling layer-by-layer in a cascade manner. As a result, it leads to the best performance on MS-COCO benchmark among the DETR-like detection models under the same setting, e.g., AP 45.7\% using ResNet50-DC5 as backbone trained in 50 epochs. We also conducted extensive experiments to confirm our analysis and verify the effectiveness of our methods. Code is available at \url{this https URL}.

Added

2026-09-25

Image Transformer

Image Transformer

Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran

OrganizationsGoogleUniversity of California Berkeley

Why you should read this

Adapts the Transformer architecture to autoregressive image generation by restricting self-attention to local neighborhoods, outperforming convolutional networks in both density estimation on ImageNet and large-scale super-resolution.

Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art.

Added

2026-09-18

LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

LeRoPE: Learnable RoPE Frequencies Improve Language Modeling

Petros Karypis, Sean O'Brien, Shreyas Kadekodi, Rui Zhu, Julian McAuley

OrganizationsUniversity of California, San Diego

Why you should read this

Develops LeRoPE, a novel modification to Rotary Positional Encodings that learns a scalar per frequency, consistently outperforming standard and partial RoPE across various model scales while requiring less computational effort.

Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models. RoPE rotates two-dimensional chunks of query and key vectors, operating as a function of their relative positional offset. The position-wise rates of rotation in RoPE typically follow a geometric sequence specified by a fixed base-frequency hyperparameter. Prior work has improved performance by either increasing this parameter to slow rotation or by applying RoPE to only a subset of QK dimensions. In this work we modify RoPE by learning a scalar per frequency, treating frequencies as learnable parameters rather than hyperparameters. We validate Learned RoPE by training a ladder of language models from scratch, ranging from 52M to 2.5B parameters. We observe and analyze the emergence of a high-norm, positional LeRoPE band. LeRoPE consistently outperforms RoPE and partial RoPE across all scales, with RoPE requiring 3.4% more compute (FLOPs) to match LeRoPE at the largest scale.

Added

2026-07-30

Creative Commons License
Big Bird: Transformers for Longer Sequences

Big Bird: Transformers for Longer Sequences

Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Yang Li, Amr Ahmed

OrganizationsGoogle

Why you should read this

Provides the theoretical proof that sparse attention involving random, window, and global connections is a universal approximator of sequence functions.

Transformers-based models, such as BERT, have been one of the most successful deep learning models for NLP. Unfortunately, one of their core limitations is the quadratic dependency (mainly in terms of memory) on the sequence length due to their full attention mechanism. To remedy this, we propose, BigBird, a sparse attention mechanism that reduces this quadratic dependency to linear. We show that BigBird is a universal approximator of sequence functions and is Turing complete, thereby preserving these properties of the quadratic, full attention model. Along the way, our theoretical analysis reveals some of the benefits of having O(1)O(1) global tokens (such as CLS), that attend to the entire sequence as part of the sparse attention mechanism. The proposed sparse attention can handle sequences of length up to 8x of what was previously possible using similar hardware. As a consequence of the capability to handle longer context, BigBird drastically improves performance on various NLP tasks such as question answering and summarization. We also propose novel applications to genomics data.

Added

2026-02-11

Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov

OrganizationsCarnegie Mellon UniversityGoogle

Why you should read this

Proposes segment-level recurrence and a relative positional encoding scheme to allow the modeling of dependencies beyond a fixed context window.

Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-of-the-art results of bpc/perplexity to 0.99 on enwiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch.

Added

2026-02-09