Express Language Modeling
Albert GongAnnabelle Michael CarrellRaaz DwivediLester Mackey
Develops Express, a method for converting non-causal attention approximations into causal ones with provable theoretical guarantees, delivering faster execution speeds than FlashAttention 2 across key long-context prefill and decoding bottlenecks.
Modern artificial intelligence relies heavily on large language models, but processing long text sequences is computationally expensive. Standard attention mechanisms suffer from quadratic scaling, meaning compute and memory requirements explode as context lengths increase. While various approximation methods speed up unmasked computations, language modeling requires causal masking, where each word can only attend to preceding text. This constraint causes severe memory and compute bottlenecks across long-context processing, key-value cache retention, and multi-step reasoning.
The article introduces Express, a general framework designed to convert high-quality unmasked attention approximations into causal approximations. By pairing this framework with the kernel halving algorithm, the article demonstrates Thinformer Express, an attention approximation technique providing strong mathematical error bounds, fixed memory scaling, and significant speedups.
The authors developed an input/output-aware GPU kernel implementation using Triton, incorporating memory tiling and parallelization to eliminate redundant data copies. They evaluated the framework across established long-context understanding benchmarks, such as LongBench-E, and complex mathematical reasoning tests using MATH-500. The testing suite benchmarked the approach against state-of-the-art exact and approximate baselines, including FlashAttention 2 and HyperAttention, across multiple open models like Llama 3.1 and DeepSeek-R1-Distill-Llama-8B.
The evaluation revealed several critical findings. First, in long-context processing, Thinformer Express achieved up to an 82-fold speedup over FlashAttention 2 at 512,000 tokens without running out of memory. Second, when integrated with leading key-value cache compression methods, it substantially reduced overall runtime while preserving baseline task accuracy. Third, in long-form step-by-step reasoning, Thinformer Express matched exact attention accuracy using only 61% of the typical cache size. Fourth, it delivered matching exact-attention accuracy in reasoning tasks while requiring only 56% of the computation time, maintaining very low compression overhead.
These results demonstrate that organizations can significantly reduce hardware costs and runtime latency for long-context language tasks without sacrificing model performance. Operating with a compact, dynamic cache lowers GPU memory requirements, allowing longer context windows on constrained hardware. This enables broader deployment of high-performing reasoning models and decreases energy consumption per inference query.
Engineering and deployment teams should consider piloting Thinformer Express to accelerate long-context pipelines and optimize key-value cache storage in generative workflows. Organizations deploying heavy reasoning models can adopt this approach to cut serving infrastructure costs. Future development should focus on extending the Triton implementation to support newer hardware features, such as 8-bit floating-point precision and advanced memory accelerators, as well as evaluating performance across non-English languages and specialized industrial domains.
The primary limitations include empirical validation limited to English, Chinese, and mathematical reasoning, alongside a software kernel that does not yet leverage newer architectural GPU features like specialized tensor memory accelerators. Nevertheless, because the theoretical guarantees are rigorous and backed by consistent benchmark results across diverse models, there is high confidence in the reported speed and memory improvements for standard transformer architectures.
- Paper: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, Tri Dao (2023). FlashAttention-2 establishes the causal and non-causal attention kernels and hardware-efficiency baseline that Express’s Triton implementation improves upon.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). FlashAttention introduces the I/O-aware attention approach underlying the efficiency context for Express’s implementation and speed comparisons.
No sufficiently relevant recommendations were found.
