keyword
hybrid SWA
Hybrid sliding window attention, often abbreviated as hybrid SWA, is an architectural mechanism in transformer-based neural networks that interleaves local sliding window attention layers with full or global attention layers. In this configuration, sliding window layers restrict each token to attend only to a fixed-size local neighborhood of neighboring tokens, which substantially lowers computational complexity and key-value cache memory usage. To prevent the loss of broader context, the model periodically places global attention layers throughout its depth, allowing information to propagate across the entire sequence. This hybrid strategy allows language models to efficiently handle long context windows while retaining the ability to capture both fine-grained local patterns and comprehensive, long-range dependencies.
1 item

