Built independently by an author, for readers. Read the story and support ChapterPal

keyword

attention with linear biases

Attention with linear biases is a positional representation method for transformer neural networks that encodes relative sequence order by adding a static, distance-proportional penalty directly to attention score computations. Instead of injecting learned or fixed positional embeddings into token representations at the input layer, this approach modifies the dot product between queries and keys by subtracting a penalty scaled linearly by the distance between the corresponding tokens and a head-specific constant slope. This mechanism imposes an inductive bias favoring nearby tokens without requiring additional learned positional parameters. Consequently, it enables models to efficiently extrapolate to input sequence lengths substantially longer than those seen during training while reducing memory usage and computational overhead.

1 item

Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Ofir Press, Noah A. Smith, Mike Lewis

OrganizationsAllen Institute for AIMetaUniversity of Washington

Why you should read this

Introduces a static positional bias that allows Transformers to generalize to sequence lengths far beyond those encountered during training.

Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.

Added

2026-02-09