keyword
rotary positional embedding
Rotary positional embedding is a method used in transformer neural networks to encode the position of tokens within a sequence by rotating their feature representations in multi-dimensional space. Unlike traditional positional encodings that add fixed or learned vectors directly to token embeddings, rotary positional embedding applies a rotation matrix to the query and key representations within the attention mechanism, scaling the rotation angle proportionally to each token position. This mathematical formulation naturally incorporates relative distance directly into the attention computation, allowing the model to focus on the spatial or sequential gaps between elements. Consequently, rotary positional embedding offers strong relative position awareness, enables smooth decay of attention weights over distance, and supports effective generalization to sequence lengths or spatial dimensions beyond those observed during training.
1 item

