Gated Linear Attention Transformers with Hardware-Efficient Training

Songlin YangBailin WangYikang ShenRameswar PandaYoon Kim

article2024ICML565 citations

Presents an I/O-aware chunkwise parallel algorithm and data-dependent gating mechanism for linear attention that outperforms FlashAttention-2 in training speed while matching standard Transformers and Mamba on language modeling benchmarks.

Listen

Modern large language models face computational bottlenecks because standard attention mechanisms require processing time and memory that scale quadratically with sequence length. While linear attention architectures offer efficient linear-time generation during deployment, they have historically lagged behind standard models in accuracy. Furthermore, existing implementations fail to optimize memory movement between fast and slow GPU memory, causing them to run slower in practice than standard attention on typical sequence lengths.

The article develops a hardware-efficient training algorithm, named FlashLinearAttention, and introduces a gated linear attention (GLA) Transformer that incorporates input-dependent gating. The primary objective is to evaluate whether GLA can match the task accuracy of leading Transformer architectures while maintaining superior computational speed and memory efficiency during both training and long-sequence processing.

The authors designed an input-output-aware training method that organizes computations into chunks and sub-chunks, balancing parallel execution with hardware-accelerated matrix multiplication on specialized GPU tensor cores. To evaluate the approach, they trained language models at scales of 340 million and 1.3 billion parameters on subsets of up to 100 billion tokens from the SlimPajama dataset. They benchmarked the GLA Transformer against strong standard Transformers and modern linear-time sequence models across standard language modeling, commonsense reasoning, and information retrieval tasks.

The evaluation revealed several key findings. First, FlashLinearAttention operated substantially faster than FlashAttention-2 as an attention layer, even on sequences as short as 1,000 tokens. Second, the 1.3 billion-parameter GLA Transformer matched the overall performance of the standard Transformer baseline, achieving a 51.0% average zero-shot accuracy compared to 50.9% for the baseline, while outperforming linear-time competitors like RetNet at 48.9% and Mamba at 50.0%. Third, GLA demonstrated superior length generalization, maintaining low perplexity when scaled from 2,000 training tokens to sequences exceeding 20,000 tokens, where standard Transformers fail. Fourth, GLA achieved higher training throughput than Mamba on single GPUs, particularly at sequence lengths beyond 4,000 tokens, while maintaining superior accuracy on memory- and recall-intensive extraction benchmarks.

These findings indicate that organizations can train and deploy large language models with linear inference costs and faster long-context training without sacrificing output quality. By effectively leveraging specialized GPU matrix hardware and data-dependent gating, GLA removes the traditional trade-off between the efficiency of recurrent architectures and the modeling accuracy of standard Transformers, potentially lowering long-term compute and infrastructure costs.

Organizations training foundation models should consider adopting hardware-efficient gated linear attention for long-context workflows and scalable inference pipelines. However, because the empirical evaluations were capped at 1.3 billion parameters and 100 billion tokens due to compute constraints, engineering teams should conduct internal pilot studies at larger model scales (such as 7 billion parameters and above) and explore multi-modal data before committing to full-scale architectural transitions.

Cover for Gated Linear Attention Transformers with Hardware-Efficient Training

Abstract

Transformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness and are thus slower than highly optimized implementations of softmax attention. This work describes a hardware-efficient algorithm for linear attention that trades off memory movement against parallelizability. The resulting implementation, dubbed FLASHLINEARATTENTION, is faster than FLASHATTENTION-2 (Dao, 2023) as a standalone layer even on short sequence lengths (e.g., 1K). We then generalize this algorithm to a more expressive variant of linear attention with data-dependent gates. When used as a replacement for the standard attention layer in Transformers, the resulting gated linear attention (GLA) Transformer is found to perform competitively against the LLaMA-architecture Transformer (Touvron et al., 2023) as well recent linear-time-inference baselines such as RetNet (Sun et al., 2023a) and Mamba (Gu & Dao, 2023) on moderate-scale language modeling experiments. GLA Transformer is especially effective at length generalization, enabling a model trained on 2K to generalize to sequences longer than 20K without significant perplexity degradations. For training speed, the GLA Transformer has higher throughput than a similarly-sized Mamba model.

Table of Contents

  • 1 Introduction
  • 2 Background: Linear Attention
  • 2.1 Parallel and Recurrent Forms
  • 2.2 Chunkwise Parallel Form
  • 3 Hardware-Efficient Linear Attention
  • 3.1 Principles of Hardware-Efficient Algorithms
  • 3.2 Hardware Considerations for Linear Attention
  • 3.3 FLASHLINEARATTENTION: Hardware-Efficient Linear Attention with the Chunkwise Form
  • 4 Gated Linear Attention
  • 4.1 Recurrent and Parallel Form of GLA
  • 4.2 Chunkwise Parallel Form of GLA
  • 4.3 Hardware-Efficient GLA
  • 4.4 GLA Transformer
  • 5 Empirical Study
  • 5.1 Experimental Setup
  • 5.2 Main Results
  • 5.3 Training Efficiency
  • 5.4 Limitations & Future Work
  • 6 Related Work
  • 7 Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A Extended Related Work
  • A.1 Linear Attention
  • A.2 Sequence parallelism
  • A.3 Hardware-efficient algorithm
  • B Details for Chunkwise (Gated) Linear Attention
  • C General Gated Linear Attention
  • C.1 Parallel form
  • C.2 Chunkwise parallel form
  • D Additional Experimental Results

Knowls

  1. Knowl 1 — Gated Linear Attention Recurrent Formulation and Gate Parameterization

    model/method

    Gated Linear Attention (GLA) enhances causal linear attention by equipping the recurrent hidden state update with a time-varying, data-dependent vector decay (forget gate) αt∈(0,1)1×dk\alpha_t \in (0, 1)^{1 \times d_k} for sequence positions t∈[1,L]t \in [1, L]:

    St=Diag(αt)St−1+kt⊤vt=(αt⊤1)⊙St−1+kt⊤vtS_t = \text{Diag}(\alpha_t) S_{t-1} + k_t^\top v_t = (\alpha_t^\top \mathbf{1}) \odot S_{t-1} + k_t^\top v_t

    ot=qtSto_t = q_t S_t

    where xt∈R1×dx_t \in \mathbb{R}^{1 \times d} is the input token representation, qt=xtWQ∈R1×dkq_t = x_t W_Q \in \mathbb{R}^{1 \times d_k} is the query vector, kt=xtWK∈R1×dkk_t = x_t W_K \in \mathbb{R}^{1 \times d_k} is the key vector, vt=xtWV∈R1×dvv_t = x_t W_V \in \mathbb{R}^{1 \times d_v} is the value vector, St∈Rdk×dvS_t \in \mathbb{R}^{d_k \times d_v} is the matrix-valued hidden state (S0=0S_0 = 0), ot∈R1×dvo_t \in \mathbb{R}^{1 \times d_v} is the layer output vector, and 1∈R1×dv\mathbf{1} \in \mathbb{R}^{1 \times d_v} is a vector of all ones.

    The vector gate αt\alpha_t is computed from the input representation xtx_t via a low-rank projection followed by an elementwise sigmoid activation and a temperature factor:

    αt=σ(xtWα,1Wα,2+bα)1/τ∈(0,1)1×dk\alpha_t = \sigma(x_t W_{\alpha, 1} W_{\alpha, 2} + b_\alpha)^{1/\tau} \in (0, 1)^{1 \times d_k}

    where Wα,1∈Rd×16W_{\alpha, 1} \in \mathbb{R}^{d \times 16}, Wα,2∈R16×dkW_{\alpha, 2} \in \mathbb{R}^{16 \times d_k}, bα∈R1×dkb_\alpha \in \mathbb{R}^{1 \times d_k}, and τ=16\tau = 16 is a fixed temperature parameter that biases gate activations closer to 11 to maintain slower historical forgetting. Parameter allocation sets dk=d/2d_k = d/2 and dv=dd_v = d, matching the total parameter footprint of standard softmax self-attention layers (approx4d2 approx 4d^2 parameters per layer).

  2. Knowl 2 — Memory-Efficient Closed-Form Gradient for Data-Dependent Gating

    theoretical result

    In gated linear attention where the recurrent state evolves as St=Diag(αt)St−1+kt⊤vtS_t = \text{Diag}(\alpha_t) S_{t-1} + k_t^\top v_t, let bt=∏j=1tαj∈(0,1)1×dkb_t = \prod_{j=1}^t \alpha_j \in (0, 1)^{1 \times d_k} denote the cumulative gate product up to step tt. Given the backpropagated gradients with respect to queries dqt∈R1×dkdq_t \in \mathbb{R}^{1 \times d_k} and keys dkt∈R1×dkdk_t \in \mathbb{R}^{1 \times d_k}, the gradient with respect to the logarithm of the cumulative gate log⁡bt\log b_t satisfies the exact closed form:

    dlog⁡bt=qt⊙dqt−kt⊙dktd\log b_t = q_t \odot dq_t - k_t \odot dk_t

    Because log⁡bt=∑i=1tlog⁡αi\log b_t = \sum_{i=1}^t \log \alpha_i, the gradient with respect to the per-step log-gate log⁡αt\log \alpha_t is computed via a reverse cumulative sum:

    dlog⁡αt=∑i=tLdlog⁡bid\log \alpha_t = \sum_{i=t}^L d\log b_i

    This closed-form formulation computes the full gradient vector dαtd\alpha_t without materializing the sequence of 3D matrix-valued hidden states S∈RL×dk×dvS \in \mathbb{R}^{L \times d_k \times d_v} in GPU global memory (HBM), reducing backward-pass activation memory requirements from O(Ldkdv)O(L d_k d_v) to O(Ldk)O(L d_k).

  3. Knowl 3 — Two-Level Secondary Chunking Algorithm for Gated Linear Attention

    algorithm

    Because gated linear attention requires elementwise log-space accumulation (Pij=∑k=1dkQikKjkexp⁡(log⁡Bik−log⁡Bjk)P_{ij} = \sum_{k=1}^{d_k} Q_{ik} K_{jk} \exp(\log B_{ik} - \log B_{jk})) to prevent numerical underflow/overflow, the intra-chunk attention computation cannot be directly mapped to a single half-precision GEMM. The two-level secondary chunking algorithm partitions each primary chunk of size CC into smaller sub-chunks of size cc (e.g., c=16c = 16). Inter-sub-chunk interactions across different sub-chunks within a chunk are computed using fast half-precision matrix multiplications on GPU Tensor Cores, while intra-sub-chunk interactions within the same sub-chunk are computed in full precision (FP32) log space.

    Input: Q, K in R^{L x d_k}, V in R^{L x d_v}, a in R^{L x d_k} (log forget gate), primary chunk size C, sub-chunk size c
    Output: O in R^{L x d_v}
    Initialize S = 0 in R^{d_k x d_v}
    Compute intra-chunk cumulative log gates B in R^{L x d_k}
    for i = 0 to L/C - 1 do
        bq = Q[i*C : (i+1)*C], bk = K[i*C : (i+1)*C]
        bv = V[i*C : (i+1)*C], bb = B[i*C : (i+1)*C]
        b_end = bb[C - 1] (last row of chunk log decay)
        q_chunk = bq * exp(bb)
        k_chunk = bk * exp(b_end - bb)
        g_chunk = exp(b_end)
        o_inter = q_chunk S
        S = (g_chunk^T 1) * S + k_chunk^T bv
        for j = 0 to C/c - 1 do
            q_sub = bq[j*c : (j+1)*c], k_sub = bk[j*c : (j+1)*c]
            v_sub = bv[j*c : (j+1)*c], b_sub = bb[j*c : (j+1)*c]
            Compute intra-sub-chunk attention p_intra in R^{c x c} in FP32 log space without matmul
            o_sub = p_intra v_sub
            z_ref = b_sub[0]
            q_scaled = q_sub * exp(b_sub - z_ref)
            for u = 0 to j - 1 do
                k_prev = bk[u*c : (u+1)*c], v_prev = bv[u*c : (u+1)*c], b_prev = bb[u*c : (u+1)*c]
                p_inter = q_scaled @ (k_prev * exp(z_ref - b_prev))^T (computed via Tensor Core GEMM)
                o_sub = o_sub + p_inter @ v_prev
            end for
            O[i*C + j*c : i*C + (j+1)*c] = o_inter[j*c : (j+1)*c] + o_sub
        end for
    end for
    return O
  4. Knowl 4 — FlashLinearAttention Chunkwise Training Algorithm

    algorithm

    FlashLinearAttention is an I/O-aware chunkwise algorithm for standard causal linear attention (St=St−1+kt⊤vtS_t = S_{t-1} + k_t^\top v_t, ot=qtSto_t = q_t S_t) that divides the sequence length LL into N=L/CN = L/C blocks of length CC and computes intermediate states on SRAM.

    In the materialization version, sequence-level parallelism is achieved by running an initial recurrent sweep to compute and store chunk-level boundary states S[n]=SnC∈Rd×dS_{[n]} = S_{nC} \in \mathbb{R}^{d \times d} into HBM, followed by a parallel loop (parfor) across all NN chunks where SRAM tiling computes intra-chunk and inter-chunk output contributions simultaneously. In the non-materialization version, hidden states are retained in SRAM and computed sequentially across chunks n=1,…,Nn = 1, \dots, N.

    Input: Q, K, V in R^{L x d}, chunk size C in [L], materialize in {True, False}
    Output: O in R^{L x d}, and optionally S = {S_[1], ..., S_[N]}
    Divide Q, K, V into N = L/C blocks {Q_[1]...Q_[N]}, {K_[1]...K_[N]}, {V_[1]...V_[N]} of size C x d each
    Initialize S = 0 in R^{d x d} on SRAM
    Construct causal mask M in R^{C x C} on chip
    if materialize then
        for n = 1 to N do
            Store S to HBM as S_[n]
            Load K_[n], V_[n] from HBM to SRAM
            Compute on chip: S = S + K_[n]^T V_[n]
        end for
        parfor n = 1 to N do
            Load Q_[n], K_[n], V_[n], S_[n] from HBM to SRAM
            Compute on chip: O'_n = Q_[n] S_[n] + ( (Q_[n] K_[n]^T) * M ) V_[n]
            Store O'_n to HBM as O_[n]
        end parfor
        return O = {O_[1]...O_[N]}, S = {S_[1]...S_[N]}
    else
        for n = 1 to N do
            Load Q_[n], K_[n], V_[n] from HBM to SRAM
            Compute on chip: O'_n = Q_[n] S + ( (Q_[n] K_[n]^T) * M ) V_[n]
            Compute on chip: S = S + K_[n]^T V_[n]
            Store O'_n to HBM as O_[n]
        end for
        return O = {O_[1]...O_[N]}
    end if
  5. Knowl 5 — Multi-Head Gated Linear Attention Transformer Architecture

    model/method

    The Gated Linear Attention (GLA) Transformer extends the GLA layer to HH attention heads and integrates it into a Transformer backbone with feed-forward networks (FFNs).

    For each head h∈{1,…,H}h \in \{1, \dots, H\}, head dimensions are defined as dk′=dk/Hd'_k = d_k / H and dv′=dv/Hd'_v = d_v / H:

    Sth=(αth)⊤1⊙St−1h+(kth)⊤vth∈Rdk′×dv′S_t^h = (\alpha_t^h)^\top \mathbf{1} \odot S_{t-1}^h + (k_t^h)^\top v_t^h \in \mathbb{R}^{d'_k \times d'_v}

    oth=qthSth∈R1×dv′o_t^h = q_t^h S_t^h \in \mathbb{R}^{1 \times d'_v}

    Head outputs are normalized individually with LayerNorm (LN) before concatenation, modulated by an input-dependent gating vector rtr_t, and projected through an output matrix WO∈Rdv×dW_O \in \mathbb{R}^{d_v \times d}:

    ot′=concat(LN(ot1),…,LN(otH))∈R1×dvo'_t = \text{concat}(\text{LN}(o_t^1), \dots, \text{LN}(o_t^H)) \in \mathbb{R}^{1 \times d_v}

    rt=Swish(xtWr+br)∈R1×dvr_t = \text{Swish}(x_t W_r + b_r) \in \mathbb{R}^{1 \times d_v}

    yt=(rt⊙ot′)WO∈R1×dy_t = (r_t \odot o'_t) W_O \in \mathbb{R}^{1 \times d}

    where Wr∈Rd×dvW_r \in \mathbb{R}^{d \times d_v} and br∈R1×dvb_r \in \mathbb{R}^{1 \times d_v}. The contextual representation X(l)X^{(l)} at layer ll is transformed through interleaved multi-head GLA and SwiGLU FFN layers with pre-layer normalization and residual connections:

    Y(l)=GLA(LN(X(l)))+X(l)Y^{(l)} = \text{GLA}(\text{LN}(X^{(l)})) + X^{(l)}

    X(l+1)=SwiGLU(LN(Y(l)))+Y(l)X^{(l+1)} = \text{SwiGLU}(\text{LN}(Y^{(l)})) + Y^{(l)}

    SwiGLU(Z)=(Swish(ZW1)⊙ZW2)W3\text{SwiGLU}(Z) = (\text{Swish}(Z W_1) \odot Z W_2) W_3

  6. Knowl 6 — Chunkwise Parallel Forward Pass for Gated Linear Attention

    equation

    In chunkwise gated linear attention, the input sequence is divided into non-overlapping chunks of length CC, where S[i]∈Rdk×dvS_{[i]} \in \mathbb{R}^{d_k \times d_v} denotes the chunk-level hidden state after chunk ii (S[i]=SiCS_{[i]} = S_{iC}). Let bt=∏j=1tαj∈(0,1)1×dkb_t = \prod_{j=1}^t \alpha_j \in (0, 1)^{1 \times d_k} be the cumulative decay. Intra-chunk scaling vectors are defined for position j∈[1,C]j \in [1, C] within chunk i+1i+1 as:

    ΛiC+j=biC+jbiC,ΓiC+j=b(i+1)CbiC+j,γi+1=b(i+1)CbiC\Lambda_{iC+j} = \frac{b_{iC+j}}{b_{iC}}, \quad \Gamma_{iC+j} = \frac{b_{(i+1)C}}{b_{iC+j}}, \quad \gamma_{i+1} = \frac{b_{(i+1)C}}{b_{iC}}

    The inter-chunk hidden state recurrence and inter-chunk output component O[i+1]inter∈RC×dvO_{[i+1]}^{\text{inter}} \in \mathbb{R}^{C \times d_v} are given by:

    S[i+1]=(γi+1⊤1)⊙S[i]+(K[i+1]⊙Γ[i+1])⊤V[i+1]S_{[i+1]} = (\gamma_{i+1}^\top \mathbf{1}) \odot S_{[i]} + (K_{[i+1]} \odot \Gamma_{[i+1]})^\top V_{[i+1]}

    O[i+1]inter=(Q[i+1]⊙Λ[i+1])S[i]O_{[i+1]}^{\text{inter}} = (Q_{[i+1]} \odot \Lambda_{[i+1]}) S_{[i]}

    Total chunk output combines inter-chunk and intra-chunk contributions:

    O[i+1]=O[i+1]inter+O[i+1]intraO_{[i+1]} = O_{[i+1]}^{\text{inter}} + O_{[i+1]}^{\text{intra}}

    where O[i+1]intra=(((Q[i+1]⊙Λ[i+1])(K[i+1]/Λ[i+1])⊤)⊙M)V[i+1]O_{[i+1]}^{\text{intra}} = ( ( (Q_{[i+1]} \odot \Lambda_{[i+1]}) (K_{[i+1]} / \Lambda_{[i+1]})^\top ) \odot M ) V_{[i+1]} with causal mask M∈{0,1}C×CM \in \{0, 1\}^{C \times C}.

  7. Knowl 7 — Language Modeling and Downstream Task Zero-Shot Evaluation

    data/table

    Comparison of GLA Transformer against Transformer++ (LLaMA architecture with RoPE, SwiGLU, RMSNorm), RetNet, and Mamba across language modeling and zero-shot downstream benchmarks. All models were trained from scratch on identical subsets of SlimPajama (340M models on 15B tokens; 1.3B models on 100B tokens) using the Mistral tokenizer.

    Scale Model Wiki. LMB. LMB. PIQA Hella. Wino. ARC-e ARC-c Avg.
    ppl ↓\downarrow ppl ↓\downarrow acc ↑\uparrow acc ↑\uparrow acc norm ↑\uparrow acc ↑\uparrow acc ↑\uparrow acc norm ↑\uparrow ↑\uparrow
    340M Transformer++ 28.39 42.69 31.0 63.3 34.0 50.4 44.5 24.2 41.2
    15B Tok RetNet 32.33 49.19 28.6 63.5 33.5 52.5 44.5 23.4 41.0
    Mamba 28.39 39.66 30.6 65.0 35.4 50.1 46.3 23.6 41.8
    GLA 28.65 43.35 30.3 64.8 34.5 51.4 45.1 22.7 41.5
    1.3B Transformer++ 16.85 13.44 48.9 70.8 49.6 53.6 56.0 26.5 50.9
    100B Tok RetNet 18.64 17.27 43.3 70.0 47.3 52.5 54.8 25.6 48.9
    Mamba 17.06 13.89 46.2 72.2 40.1 54.1 59.0 28.2 50.0
    GLA 17.22 14.47 46.9 71.8 49.8 53.9 57.2 26.6 51.0

    The results demonstrate that data-dependent gating enables GLA to consistently outperform data-independent linear attention (RetNet) across all downstream tasks, achieving an overall downstream accuracy (51.0% at 1.3B) competitive with Transformer++ (50.9%) and Mamba (50.0%).

  8. Knowl 8 — Performance on Recall-Intensive Evaluation Benchmarks

    empirical result

    Subquadratic models with matrix-valued hidden states (GLA and RetNet) and state-space models (Mamba) were evaluated on the synthetic Multi-Query Associative Recall (MQAR) task and three real-world recall-intensive information extraction benchmarks (FDA, SWDE, and SQuAD). Standard quadratic attention achieves perfect scores across these synthetic associative recall settings.

    On the synthetic MQAR benchmark across sequence lengths of 256 (16 key-value pairs) and 512 (64 key-value pairs) with varying model dimensions (64 to 512), GLA consistently achieves higher recall accuracy than RetNet, Mamba, Hyena, and RWKV-4.

    On real-world recall tasks, model accuracy scores are reported below:

    Scale Model FDA ↑\uparrow SWDE ↑\uparrow SQuAD ↑\uparrow
    340M Params Transformer++ 21.4 42.2 22.1
    15B Tokens RetNet 2.9 13.3 27.6
    Mamba 2.1 12.4 23.0
    GLA 8.1 18.6 27.2
    1.3B Params Transformer++ 27.4 66.6 31.5
    100B Tokens RetNet 14.3 42.8 34.7
    Mamba 6.2 41.4 35.2
    GLA 19.9 50.6 42.6

    GLA significantly outperforms competing subquadratic models (Mamba and RetNet) on recall-intensive information extraction (FDA and SWDE) and reading comprehension (SQuAD), driven by its expanded dk×dvd_k \times d_v recurrent memory state and data-dependent gating selection mechanism.

  9. Knowl 9 — Context Length Extrapolation and Truncated BPTT Training

    empirical result

    1.3B-parameter models trained on SlimPajama for 100B tokens under 2K context length were evaluated for length extrapolation on the test sets of SlimPajama and PG19 without additional fine-tuning:

    1. When pretrained on 2K context lengths, GLA successfully extrapolates up to 18K positions on SlimPajama and maintains low perplexity across position buckets up to 25K on PG19, outperforming Mamba (which experiences steep perplexity degradation beyond 4K context) and RetNet.
    2. When training on long sequences (24K context), training GLA via Truncated Backpropagation Through Time (TBPTT) across 12 segments of length 2K—where hidden states are passed forward across segments but gradients are not backpropagated across segment boundaries—achieves perplexity improvements virtually identical to direct 8K context training while avoiding the memory and compute overhead of long-context gradient propagation.
  10. Knowl 10 — Training Throughput and Hardware Efficiency

    empirical result

    Training speed and memory footprint were benchmarked on a single NVIDIA H100 GPU:

    1. Standalone Layer Speed: At batch size 32, 16 heads, head dimension 64, and chunk size 64, FlashLinearAttention (both materialization and non-materialization Triton implementations) achieves faster execution times than FlashAttention-2 (CUDA) across all sequence lengths from 292^9 (512) to 2152^{15} (32768), and is substantially faster than an I/O-unaware PyTorch chunkwise linear attention implementation.
    2. 1.3B Model Training Throughput: When training 1.3B parameter models on an H100 GPU across sequence lengths from 2048 to 16284 (with corresponding batch sizes 8 to 1), GLA Transformer achieves higher training throughput (tokens per second) than Mamba, with the relative throughput advantage expanding at sequence lengths beyond 4096. GPU memory footprints among Transformer++, Mamba, and GLA remain comparable across all evaluated sequence lengths.
  11. Knowl 11 — Ablations on Gating Granularity and Head Dimension in GLA

    data/table

    Ablation study evaluating the effect of gating parameterization and attention head count on 340M parameter GLA variants pretrained for 7B tokens on SlimPajama. Evaluation metric is the average training perplexity over the final 200 optimization steps.

    Model Variant Training Ppl.
    GLA Transformer (4 heads) 14.77
    No gate (i.e., Linear Attention) 23.21
    Data-independent scalar decay (i.e., RetNet) 16.55
    Data-dependent scalar gate 15.56
    Small head dimension (8 heads) 15.29
    Large head dimension (1 head) 14.61

    The ablation indicates that:

    1. Removing the gate causes a severe degradation in perplexity (23.2123.21 vs 14.7714.77).
    2. Moving from data-independent decay (16.5516.55) to a data-dependent scalar gate improves perplexity to 15.5615.56, but fine-grained vector gating αt∈(0,1)1×dk\alpha_t \in (0, 1)^{1 \times d_k} provides further gains (14.7714.77).
    3. Larger head dimensions (fewer heads) yield superior perplexity (14.6114.61 for 1 head vs 15.2915.29 for 8 heads), with 4 heads providing an optimal balance between perplexity and GPU SRAM memory footprint.
  12. Knowl 12 — General Gated Linear Attention Formulation with Bilinear Forgetting

    model/method

    General Gated Linear Attention parameterizes the recurrent state transition using a full rank-1 outer-product forget gate matrix Gt=αt⊤βtG_t = \alpha_t^\top \beta_t with αt∈(0,1)1×dk\alpha_t \in (0, 1)^{1 \times d_k} and βt∈(0,1)1×dv\beta_t \in (0, 1)^{1 \times d_v}:

    St=(αt⊤βt)⊙St−1+kt⊤vt,ot=qtStS_t = (\alpha_t^\top \beta_t) \odot S_{t-1} + k_t^\top v_t, \quad o_t = q_t S_t

    Letting bt=∏j=1tαj∈(0,1)1×dkb_t = \prod_{j=1}^t \alpha_j \in (0, 1)^{1 \times d_k} and dt=∏j=1tβj∈(0,1)1×dvd_t = \prod_{j=1}^t \beta_j \in (0, 1)^{1 \times d_v}, the equivalent parallel attention formulation over sequence length LL with causal mask M∈{0,1}L×LM \in \{0, 1\}^{L \times L} is given by:

    Q~=Q⊙B,K~=K/B,V~=V/D\tilde{Q} = Q \odot B, \quad \tilde{K} = K / B, \quad \tilde{V} = V / D

    O~=((Q~K~⊤)⊙M)V~,O=O~⊙D\tilde{O} = ( ( \tilde{Q} \tilde{K}^\top ) \odot M ) \tilde{V}, \quad O = \tilde{O} \odot D

    where B∈(0,1)L×dkB \in (0, 1)^{L \times d_k} and D∈(0,1)L×dvD \in (0, 1)^{L \times d_v} are obtained by stacking btb_t and dtd_t.

    The closed-form gradients with respect to log⁡αt\log \alpha_t and log⁡βt\log \beta_t are:

    dlog⁡bt=kt⊙dkt−qt⊙dqt,dlog⁡αt=∑i=tLdlog⁡bid\log b_t = k_t \odot dk_t - q_t \odot dq_t, \quad d\log \alpha_t = \sum_{i=t}^L d\log b_i

    dlog⁡dt=ot⊙dot−vt⊙dvt,dlog⁡βt=∑i=tLdlog⁡did\log d_t = o_t \odot do_t - v_t \odot dv_t, \quad d\log \beta_t = \sum_{i=t}^L d\log d_i

Coverage note — None was omitted; all key algorithmic contributions (FlashLinearAttention, GLA two-level chunking, closed-form gate gradients), architectural designs, theoretical formulations (general bilinear gating), and core empirical evaluations (language modeling, recall tasks, extrapolation, speed/memory benchmarks, and ablations) are included.

References

  1. 1.Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Re, C. Zoology: Measuring and improving recall in efficient language models. CoRR, abs/2312.04927, 2023a.
  2. 2.Arora, S., Yang, B., Eyuboglu, S., Narayan, A., Hojel, A., Trummer, I., and Re, C. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes, April 2023b. URL http://arxiv.org/abs/2304.09433. arXiv:2304.09433 [cs].
  3. 3.Arora, S., Eyuboglu, S., Zhang, M., Timalsina, A., Alberti, S., Zinsley, D., Zou, J., Rudra, A., and R’e, C. Simple linear attention language models balance the recall-throughput tradeoff. ArXiv, abs/2402.18668, 2024.
  4. 4.Auer, S., Barone, D. A. C., Bartz, C., Cortes, E. G., Jaradeh, M. Y., Karras, O., Koubarakis, M., Mouromtsev, D., Pliukhin, D., Radyush, D., Shilin, I., Stocker, M., and Tsalapati, E. The sciqa scientific question answering benchmark for scholarly knowledge. Scientific Reports, 13(1):7240, May 2023. ISSN 2045-2322. doi: 10.1038/s41598-023-33607-z.
  5. 5.Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C. Using fast weights to attend to the recent past. Advances in neural information processing systems, 29, 2016.
  6. 6.Beck, M., Poppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024.
  7. 7.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer. arXiv preprint arXiv: Arxiv-2004.05150, 2020. URL https://arxiv.org/abs/2004.05150v2.
  8. 8.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  9. 9.Blelloch, G. E. Prefix sums and their applications. 1990.
  10. 10.Brandon, W., Nrusimha, A., Qian, K., Ankner, Z., Jin, T., Song, Z., and Ragan-Kelley, J. Striped attention: Faster ring attention for causal transformers. ArXiv, abs/2311.09431, 2023.
  11. 11.Buckman, J. and Gelada, C. Linear Transformers Are Faster After All, 2024.
  12. 12.Chaurasia, G., Ragan-Kelley, J., Paris, S., Drettakis, G., and Durand, F. Compiling high performance recursive filters. In High Performance Graphics, 2015.
  13. 13.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. PREPRINT, 2019. URL https://arxiv.org/abs/1904.10509v1.
  14. 14.Cho, K., Van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. Learning phrase representations using rnn encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078, 2014.
  15. 15.Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A. Rethinking attention with performers. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  16. 16.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  17. 17.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  18. 18.Dao, T. Flashattention-2: Faster attention with better parallelism and work partitioning. CoRR, abs/2307.08691, 2023. doi: 10.48550/ARXIV.2307.08691.
  19. 19.Dao, T. and Gu, A. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024.
  20. 20.Dao, T., Chen, B., Sohoni, N. S., Desai, A. D., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Re, C. Monarch: Expressive structured matrices for efficient and accurate training. In International Conference on Machine Learning, 2022a.
  21. 21.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. In NeurIPS, 2022b.
  22. 22.Fu, D. Y., Arora, S., Grogan, J., Johnson, I., Eyuboglu, S., Thomas, A. W., Spector, B., Poli, M., Rudra, A., and R’e, C. Monarch mixer: A simple sub-quadratic gemm-based architecture. ArXiv, abs/2310.12109, 2023a.
  23. 23.Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C. Hungry hungry hippos: Towards language modeling with state space models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023b.
  24. 24.Fu, D. Y., Epstein, E. L., Nguyen, E., Thomas, A., Zhang, M., Dao, T., Rudra, A., and Re, C. Simple hardware-efficient long convolutions for sequence modeling. International Conference on Machine Learning, 2023c. doi: 10.48550/arXiv.2302.06646. URL https://arxiv.org/abs/2302.06646v1.
  25. 25.Fu, D. Y., Kumbong, H., Nguyen, E., and Re, C. Flashfftconv: Efficient convolutions for long sequences with tensor cores. CoRR, abs/2311.05908, 2023d.
  26. 26.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, September 2021.
  27. 27.Gers, F. A., Schmidhuber, J., and Cummins, F. A. Learning to forget: Continual prediction with LSTM. Neural Comput., 12(10):2451–2471, 2000.
  28. 28.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. 2023.
  29. 29.Gu, A., Goel, K., and R’e, C. Efficiently modeling long sequences with structured state spaces. International Conference On Learning Representations, 2021a.
  30. 30.Gu, A., Johnson, I., Goel, K., Saab, K. K., Dao, T., Rudra, A., and R’e, C. Combining recurrent, convolutional, and continuous-time models with linear state-space layers. Neural Information Processing Systems, 2021b. URL https://arxiv.org/abs/2110.13985v1.
  31. 31.Gu, A., Goel, K., and Re, C. Efficiently modeling long sequences with structured state spaces. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022.
  32. 32.Gupta, A. and Berant, J. Diagonal state spaces are as effective as structured state spaces. ARXIV.ORG, 2022. doi: 10.48550/arXiv.2203.14343.
  33. 33.Hasani, R., Lechner, M., Wang, T.-H., Chahine, M., Amini, A., and Rus, D. Liquid structural state-space models. arXiv preprint arXiv:2209.12951, 2022.
  34. 34.Hinton, G. E. and Plaut, D. C. Using fast weights to deblur old memories. In Proceedings of the ninth annual conference of the Cognitive Science Society, pp. 177–186, 1987.
  35. 35.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
  36. 36.Hooker, S. The hardware lottery. Communications of the ACM, 64:58 – 65, 2020.
  37. 37.Hua, W., Dai, Z., Liu, H., and Le, Q. V. Transformer quality in linear time. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 9099–9117. PMLR, 2022.
  38. 38.Irie, K., Schlag, I., Csordas, R., and Schmidhuber, J. Going beyond linear transformers with recurrent fast weight programmers. Advances in Neural Information Processing Systems, 34:7703–7717, 2021.
  39. 39.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. ArXiv preprint, abs/2310.06825, 2023.
  40. 40.Kacham, P., Mirrokni, V., and Zhong, P. Polysketchformer: Fast transformers via sketching polynomial kernels, 2023.
  41. 41.Kasai, J., Peng, H., Zhang, Y., Yogatama, D., Ilharco, G., Pappas, N., Mao, Y., Chen, W., and Smith, N. A. Finetuning pretrained transformers into RNNs. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 10630–10643, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.830.
  42. 42.Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning, pp. 5156–5165. PMLR, 2020.
  43. 43.Katsch, T. Gateloop: Fully data-controlled linear recurrence for sequence modeling. ArXiv, abs/2311.01927, 2023.
  44. 44.Kitaev, N., Kaiser, L., and Levskaya, A. Reformer: The efficient transformer. International Conference On Learning Representations, 2020. URL https://arxiv.org/abs/2001.04451v2.
  45. 45.Li, D., Shao, R., Xie, A., Xing, E. P., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H. Lightseq: Sequence level parallelism for distributed training of long context transformers. ArXiv, abs/2310.03294, 2023a.
  46. 46.Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y. Sequence parallelism: Long sequence training from system perspective. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, July 2023b. Association for Computational Linguistics.
  47. 47.Li, S., You, C., Guruganesh, G., Ainslie, J., Ontanon, S., Zaheer, M., Sanghai, S., Yang, Y., Kumar, S., and Bhojanapalli, S. Functional interpolation for relative positions improves long context transformers. arXiv preprint arXiv:2310.04418, 2023c.
  48. 48.Li, Y., Cai, T., Zhang, Y., Chen, D., and Dey, D. What makes convolutional models great on long sequence modeling? In The Eleventh International Conference on Learning Representations, 2023d. URL https://openreview.net/forum?id=TGJSPbRpJX-.
  49. 49.Lingle, L. D. Transformer-vq: Linear-time transformers via vector quantization. CoRR, abs/2309.16354, 2023. doi: 10.48550/ARXIV.2309.16354.
  50. 50.Liu, H., Zaharia, M., and Abbeel, P. Ring attention with blockwise transformers for near-infinite context. ArXiv, abs/2310.01889, 2023.
  51. 51.Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., and Liu, Y. Vmamba: Visual state space model. arXiv preprint arXiv:2401.10166, 2024.
  52. 52.Lockard, C., Shiralkar, P., and Dong, X. L. OpenCeres: When Open Information Extraction Meets the Semi-Structured Web. In Burstein, J., Doran, C., and Solorio, T. (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 3047–3056, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19- 1309. URL https://aclanthology.org/N19-1309.
  53. 53.Loshchilov, I. and Hutter, F. Fixing weight decay regularization in adam. 2018.
  54. 54.Ma, J., Li, F., and Wang, B. U-mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv preprint arXiv:2401.04722, 2024.
  55. 55.Ma, X., Zhou, C., Kong, X., He, J., Gui, L., Neubig, G., May, J., and Zettlemoyer, L. Mega: Moving average equipped gated attention. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=qNLe3iq2El.
  56. 56.Mao, H. H. Fine-tuning pre-trained transformers into decaying fast weights. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 10236–10242, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.697.
  57. 57.Martin, E. and Cundy, C. Parallelizing linear recurrent neural nets over sequence length. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018.
  58. 58.Massaroli, S., Poli, M., Fu, D. Y., Kumbong, H., Parnichkun, R. N., Timalsina, A., Romero, D. W., McIntyre, Q., Chen, B., Rudra, A., Zhang, C., Re, C., Ermon, S., and Bengio, Y. Laughing hyena distillery: Extracting compact recurrences from convolutions. NEURIPS, 2023. URL https://arxiv.org/abs/2310.18780v1.
  59. 59.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  60. 60.Nahshan, Y., Kampeas, J., and Haleva, E. Linear log-normal attention with unbiased concentration, 2023.
  61. 61.Oren, M., Hassid, M., Adi, Y., and Schwartz, R. Transformers are multi-state rnns. ArXiv, abs/2401.06104, 2024.
  62. 62.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  63. 63.Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., V., K. K. G., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhu, J., and Zhu, R. RWKV: reinventing rnns for the transformer era. CoRR, abs/2305.13048, 2023. doi: 10.48550/ARXIV.2305.13048.
  64. 64.Peng, B., Goldstein, D., Anthony, Q., Albalak, A., Alcaide, E., Biderman, S., Cheah, E., Ferdinan, T., Hou, H., Kazienko, P., et al. Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence. arXiv preprint arXiv:2404.05892, 2024.
  65. 65.Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L. Random feature attention. arXiv preprint arXiv:2103.02143, 2021.
  66. 66.Peng, H., Kasai, J., Pappas, N., Yogatama, D., Wu, Z., Kong, L., Schwartz, R., and Smith, N. A. ABC: Attention with bounded-memory control. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 2022. Association for Computational Linguistics.
  67. 67.Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Re, C. Hyena hierarchy: Towards larger convolutional language models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pp. 28043–28078. PMLR, 2023.
  68. 68.Pramanik, S., Elelimy, E., Machado, M. C., and White, A. Recurrent linear transformers. CoRR, abs/2310.15719, 2023.
  69. 69.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021.
  70. 70.Qin, Z., Han, X., Sun, W., Li, D., Kong, L., Barnes, N., and Zhong, Y. The devil in linear transformer. arXiv preprint arXiv:2210.10340, 2022.
  71. 71.Qin, Z., Han, X., Sun, W., He, B., Li, D., Li, D., Dai, Y., Kong, L., and Zhong, Y. Toeplitz neural network for sequence modeling. In The Eleventh International Conference on Learning Representations, 2023a. URL https://openreview.net/forum?id=IxmWsm4xrua.
  72. 72.Qin, Z., Li, D., Sun, W., Sun, W., Shen, X., Han, X., Wei, Y., Lv, B., Yuan, F., Luo, X., et al. Scaling transnormer to 175 billion parameters. arXiv preprint arXiv:2307.14995, 2023b.
  73. 73.Qin, Z., Yang, S., and Zhong, Y. Hierarchically gated recurrent neural network for sequence modeling. CoRR, abs/2311.04823, 2023c. doi: 10.48550/ARXIV.2311.04823.
  74. 74.Qin, Z., Sun, W., Li, D., Shen, X., Sun, W., and Zhong, Y. Lightning attention-2: A free lunch for handling unlimited sequence lengths in large language models. 2024a.
  75. 75.Qin, Z., Yang, S., Sun, W., Shen, X., Li, D., Sun, W., and Zhong, Y. Hgrn2: Gated linear rnns with state expansion. arXiv preprint arXiv:2404.07904, 2024b.
  76. 76.Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P. Compressive transformers for long-range sequence modelling. arXiv preprint, 2019.
  77. 77.Rajpurkar, P., Jia, R., and Liang, P. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Gurevych, I. and Miyao, Y. (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 784–789, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-2124. URL https://aclanthology.org/P18-2124.
  78. 78.Ren, L., Liu, Y., Wang, S., Xu, Y., Zhu, C., and Zhai, C. Sparse modular activation for efficient sequence modeling. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=TfbzX6I14i.
  79. 79.Roemmele, M., Bejan, C. A., and Gordon, A. S. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series, 2011. URL https://people.ict.usc.edu/∼gordon/publications/AAAI-SPRING11A.PDF.
  80. 80.Romero, D. W., Kuzina, A., Bekkers, E. J., Tomczak, J. M., and Hoogendoorn, M. Ckconv: Continuous kernel convolution for sequential data. arXiv preprint arXiv: 2102.02611, 2021. URL https://arxiv.org/abs/2102.02611v3.
  81. 81.Roy, A., Saffar, M., Vaswani, A., and Grangier, D. Efficient content-based sparse attention with routing transformers. International Conference On Topology, Algebra And Categories In Logic, 2020. doi: 10.1162/tacl a 00353. URL https://arxiv.org/abs/2003.05997v5.
  82. 82.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  83. 83.Saphra, N., Fleisig, E., Cho, K., and Lopez, A. First tragedy, then parse: History repeats itself in the new era of large language models. ArXiv, abs/2311.05020, 2023.
  84. 84.Schlag, I., Irie, K., and Schmidhuber, J. Linear transformers are secretly fast weight programmers. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 9355–9366. PMLR, 2021.
  85. 85.Schmidhuber, J. Learning to control fast-weight memories: An alternative to dynamic recurrent networks. Neural Computation, 4(1):131–139, 1992.
  86. 86.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  87. 87.Smith, J. T. H., Warrington, A., and Linderman, S. W. Simplified state space layers for sequence modeling. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023.
  88. 88.Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, 2023.
  89. 89.Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. CoRR, abs/2104.09864, 2021.
  90. 90.Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621, 2023a.
  91. 91.Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F. A lengthextrapolatable transformer. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 14590–14604. Association for Computational Linguistics, 2023b. doi: 10.18653/V1/2023.ACL-LONG.816.
  92. 92.Sun, Y., Dong, L., Zhu, Y., Huang, S., Wang, W., Ma, S., Zhang, Q., Wang, J., and Wei, F. You only cache once: Decoder-decoder architectures for language models. 2024.
  93. 93.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  94. 94.van der Westhuizen, J. and Lasenby, J. The unreasonable effectiveness of the forget gate. CoRR, abs/1804.04849, 2018.
  95. 95.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  96. 96.Wang, C., Tsepa, O., Ma, J., and Wang, B. Graph-mamba: Towards long-range graph sequence modeling with selective state spaces. arXiv preprint arXiv:2402.00789, 2024a.
  97. 97.Wang, J., Yan, J. N., Gu, A., and Rush, A. M. Pretraining without attention. CoRR, abs/2212.10544, 2022.
  98. 98.Wang, J., Gangavarapu, T., Yan, J. N., and Rush, A. M. Mambabyte: Token-free selective state space model. arXiv preprint arXiv:2401.13660, 2024b.
  99. 99.Wu, F., Fan, A., Baevski, A., Dauphin, Y., and Auli, M. Pay less attention with lightweight and dynamic convolutions. International Conference on Learning Representations, 2019. URL https://arxiv.org/abs/1901.10430v2.
  100. 100.Xing, Z., Ye, T., Yang, Y., Liu, G., and Zhu, L. Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation. arXiv preprint arXiv:2401.13560, 2024.
  101. 101.Yan, J. N., Gu, J., and Rush, A. M. Diffusion models without attention. 2023.
  102. 102.Yang, Y., Xing, Z., and Zhu, L. Vivim: a video vision mamba for medical video object segmentation. arXiv preprint arXiv:2401.14168, 2024.
  103. 103.Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020. URL https://arxiv.org/abs/2007.14062v2.
  104. 104.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  105. 105.Zhang, B. and Sennrich, R. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  106. 106.Zhang, J., Jiang, S., Feng, J., Zheng, L., and Kong, L. Linear attention via orthogonal memory, 2023.
  107. 107.Zhang, M., Bhatia, K., Kumbong, H., and Re, C. The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry, 2024.
  108. 108.Zhang, Y. and Cai, D. Linearizing transformer with key-value memory. In Goldberg, Y., Kozareva, Z., and Zhang, Y. (eds.), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics.
  109. 109.Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., and Wang, X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417, 2024.

Citation

MLA
Yang, S., et al. “Gated Linear Attention Transformers with Hardware-Efficient Training”. arXiv, 2023, http://arxiv.org/abs/2312.06635v6.
APA
Yang, S., Wang, B., Shen, Y., Panda, R., & Kim, Y. (2023). Gated Linear Attention Transformers with Hardware-Efficient Training. arXiv. http://arxiv.org/abs/2312.06635v6
Chicago
Yang, S., B. Wang, Y. Shen, R. Panda, and Y. Kim. 2023. “Gated Linear Attention Transformers with Hardware-Efficient Training”. arXiv. http://arxiv.org/abs/2312.06635v6.
Harvard
Yang, S. et al. (2023) “Gated Linear Attention Transformers with Hardware-Efficient Training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.06635v6.
Vancouver
1. Yang S, Wang B, Shen Y, Panda R, Kim Y (2023) Gated Linear Attention Transformers with Hardware-Efficient Training. arXiv

BibTeX

@article{yang2023gated,
  title = {Gated Linear Attention Transformers with Hardware-Efficient Training},
  author = {Yang, Songlin and Wang, Bailin and Shen, Yikang and Panda, Rameswar and Kim, Yoon},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.06635v6},
  eprint = {2312.06635}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/