keyword
attention with linear biases
Attention with linear biases is a positional representation method for transformer neural networks that encodes relative sequence order by adding a static, distance-proportional penalty directly to attention score computations. Instead of injecting learned or fixed positional embeddings into token representations at the input layer, this approach modifies the dot product between queries and keys by subtracting a penalty scaled linearly by the distance between the corresponding tokens and a head-specific constant slope. This mechanism imposes an inductive bias favoring nearby tokens without requiring additional learned positional parameters. Consequently, it enables models to efficiently extrapolate to input sequence lengths substantially longer than those seen during training while reducing memory usage and computational overhead.
1 item

