keyword
Ring Attention
Ring Attention is a distributed computing technique for transformer models that enables neural networks to process exceptionally long input sequences across multiple hardware devices. In traditional transformer architectures, calculating self-attention requires memory that scales quadratically with sequence length, creating a memory bottleneck on individual processors. Ring Attention addresses this limitation by partitioning the input sequence into smaller blocks and assigning each block to a separate device arranged in a logical ring topology. During computation, each device retains its assigned query data while circulating key and value blocks around the ring, concurrently overlapping inter-device communication with blockwise attention calculations. This distributed approach allows the context length to scale proportionally with the number of participating devices without relying on approximations or incurring severe communication latency.
2 items

Gated Linear Attention Transformers with Hardware-Efficient Training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim
Why you should read this
Presents an I/O-aware chunkwise parallel algorithm and data-dependent gating mechanism for linear attention that outperforms FlashAttention-2 in training speed while matching standard Transformers and Mamba on language modeling benchmarks.
Transformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness and are thus slower than highly optimized implementations of softmax attention. This work describes a hardware-efficient algorithm for linear attention that trades off memory movement against parallelizability. The resulting implementation, dubbed FLASHLINEARATTENTION, is faster than FLASHATTENTION-2 (Dao, 2023) as a standalone layer even on short sequence lengths (e.g., 1K). We then generalize this algorithm to a more expressive variant of linear attention with data-dependent gates. When used as a replacement for the standard attention layer in Transformers, the resulting gated linear attention (GLA) Transformer is found to perform competitively against the LLaMA-architecture Transformer (Touvron et al., 2023) as well recent linear-time-inference baselines such as RetNet (Sun et al., 2023a) and Mamba (Gu & Dao, 2023) on moderate-scale language modeling experiments. GLA Transformer is especially effective at length generalization, enabling a model trained on 2K to generalize to sequences longer than 20K without significant perplexity degradations. For training speed, the GLA Transformer has higher throughput than a similarly-sized Mamba model.
Added
2026-09-28

Ring Attention with Blockwise Transformers for Near-Infinite Context
Hao Liu, Matei Zaharia, Pieter Abbeel
Why you should read this
Overlaps the communication of key-value blocks with local computation over a mesh ring, theoretically allowing context lengths bounded only by the total cluster memory.
Transformers have emerged as the architecture of choice for many state-of-the-art AI models, showcasing exceptional performance across a wide range of AI applications. However, the memory demands imposed by Transformers limit their ability to handle long sequences, thereby posing challenges in utilizing videos, actions, and other long-form sequences and modalities in complex environments. We present a novel approach, Ring Attention with Blockwise Transformers (Ring Attention), which leverages blockwise computation of self-attention and feedforward to distribute long sequences across multiple devices while fully overlapping the communication of key-value blocks with the computation of blockwise attention. Our approach enables training and inference of sequences that are up to device count times longer than those achievable by prior memory-efficient Transformers, without resorting to approximations or incurring additional communication and computation overheads. Extensive experiments on language modeling and reinforcement learning tasks demonstrate the effectiveness of our approach in allowing millions of tokens context size and improving performance.
Added
2026-03-13
