A local self-attention layer is a neural network component that calculates attention weights between an input element and a restricted neighborhood of surrounding elements rather than across the entire input feature map or sequence. In vision and sequential modeling, this layer applies the self-attention mechanism within localized spatial or temporal windows, serving as an alternative to spatial convolutions or global attention. By constraining each query to interact only with nearby keys and values, local self-attention retains the ability to dynamically weight features based on content while avoiding the quadratic computational and memory complexity associated with full-context global attention.