Gated Linear Attention Transformers with Hardware-Efficient Training
Songlin YangBailin WangYikang ShenRameswar PandaYoon Kim
Presents an I/O-aware chunkwise parallel algorithm and data-dependent gating mechanism for linear attention that outperforms FlashAttention-2 in training speed while matching standard Transformers and Mamba on language modeling benchmarks.
Modern large language models face computational bottlenecks because standard attention mechanisms require processing time and memory that scale quadratically with sequence length. While linear attention architectures offer efficient linear-time generation during deployment, they have historically lagged behind standard models in accuracy. Furthermore, existing implementations fail to optimize memory movement between fast and slow GPU memory, causing them to run slower in practice than standard attention on typical sequence lengths.
The article develops a hardware-efficient training algorithm, named FlashLinearAttention, and introduces a gated linear attention (GLA) Transformer that incorporates input-dependent gating. The primary objective is to evaluate whether GLA can match the task accuracy of leading Transformer architectures while maintaining superior computational speed and memory efficiency during both training and long-sequence processing.
The authors designed an input-output-aware training method that organizes computations into chunks and sub-chunks, balancing parallel execution with hardware-accelerated matrix multiplication on specialized GPU tensor cores. To evaluate the approach, they trained language models at scales of 340 million and 1.3 billion parameters on subsets of up to 100 billion tokens from the SlimPajama dataset. They benchmarked the GLA Transformer against strong standard Transformers and modern linear-time sequence models across standard language modeling, commonsense reasoning, and information retrieval tasks.
The evaluation revealed several key findings. First, FlashLinearAttention operated substantially faster than FlashAttention-2 as an attention layer, even on sequences as short as 1,000 tokens. Second, the 1.3 billion-parameter GLA Transformer matched the overall performance of the standard Transformer baseline, achieving a 51.0% average zero-shot accuracy compared to 50.9% for the baseline, while outperforming linear-time competitors like RetNet at 48.9% and Mamba at 50.0%. Third, GLA demonstrated superior length generalization, maintaining low perplexity when scaled from 2,000 training tokens to sequences exceeding 20,000 tokens, where standard Transformers fail. Fourth, GLA achieved higher training throughput than Mamba on single GPUs, particularly at sequence lengths beyond 4,000 tokens, while maintaining superior accuracy on memory- and recall-intensive extraction benchmarks.
These findings indicate that organizations can train and deploy large language models with linear inference costs and faster long-context training without sacrificing output quality. By effectively leveraging specialized GPU matrix hardware and data-dependent gating, GLA removes the traditional trade-off between the efficiency of recurrent architectures and the modeling accuracy of standard Transformers, potentially lowering long-term compute and infrastructure costs.
Organizations training foundation models should consider adopting hardware-efficient gated linear attention for long-context workflows and scalable inference pipelines. However, because the empirical evaluations were capped at 1.3 billion parameters and 100 billion tokens due to compute constraints, engineering teams should conduct internal pilot studies at larger model scales (such as 7 billion parameters and above) and explore multi-modal data before committing to full-scale architectural transitions.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). It introduces the fundamental formulation of linear transformers and shows their equivalence to recurrent neural networks with 2D hidden states, which the source builds upon and generalizes.
- Paper: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, Tri Dao (2023). It establishes the I/O-aware tiling and work-partitioning techniques that inspire and serve as the baseline speed benchmark for the hardware-efficient FLASHLINEARATTENTION algorithm.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). It introduces hardware-aware selective state-space modeling with data-dependent gating, providing direct architectural motivation and a key performance baseline for Gated Linear Attention.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). It introduces IO-aware exact attention computations on GPU SRAM/HBM hierarchies, laying the groundwork for memory-efficient attention algorithms.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). It establishes key kernel approximation methods for linear-complexity attention that motivate modern linear attention formulations.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). It generalizes the connection between linear attention models like GLA and selective state-space models into a unified State Space Duality framework with hardware-efficient block algorithms.
- Paper: Dynamic Linear Attention, Xin Wang et al. (2026). It extends linear attention architectures, including gated variants, by introducing dynamic, content-aware memory state creation and fixed-size cache merging to reduce compression errors across long contexts.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). It develops a unifying test-time regression framework that formally derives gated linear attention and state-space models as regression updates for associative recall.
- Paper: Fast Weight Attention for Continual Learning, Yifan Zhang et al. (2026). It extends gated recurrent state updates like those in linear attention to explicit normalized next-latent regression rules for stable continual learning.
- Paper: Simple linear attention language models balance the recall-throughput tradeoff, Simran Arora et al. (2024). It investigates and expands the theoretical and practical recall-throughput tradeoffs inherent to recurrent linear attention architectures like GLA.
- Paper: Learning to (Learn at Test Time): RNNs with Expressive Hidden States, Yu Sun 0020 et al. (2025). It advances sub-quadratic recurrent sequence modeling by replacing fixed or gated linear recurrent states with expressive hidden models trained at test time.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). It provides a comprehensive survey comparing gated linear models, state-space architectures, and IO-aware attention kernels across the modern landscape of efficient large language models.
