SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
Jintao ZhangChendong XiangHaofeng HuangJia WeiHaocheng XiJun ZhuJianfei Chen
Proposes SpargeAttn, a training-free sparse attention mechanism that uses a two-stage online filter to accelerate inference across language, image, and video models without degrading output quality.
As modern artificial intelligence models process increasingly large context windows—such as high-resolution video frames and 128,000-token text sequences—the standard attention mechanism creates severe computational bottlenecks. Because attention scales quadratically with sequence length, generating long-form content requires massive GPU memory and latency. While sparse attention techniques attempt to reduce this overhead by skipping near-zero values, existing methods generally suffer from two flaws: they rely on task-specific patterns that fail to generalize across text, image, and video domains, or they impose prediction overheads that erode practical speed gains.
The article introduces and evaluates SpargeAttn, a universal, training-free, sparse, and quantized attention framework. The primary objective is to demonstrate that a single, pattern-free sparse attention mechanism can significantly accelerate inference across diverse generative modalities without retraining or degrading model output quality.
To achieve this, the authors developed a two-stage online filtering system integrated with low-precision arithmetic. In the first stage, the method predicts sparse blocks on the fly by selectively compressing token blocks that exhibit high internal similarity, while leaving non-similar blocks fully computed to protect essential information. In the second stage, an online filter identifies and omits negligible matrix updates during computation at the GPU warp level. For visual tasks, a space-filling Hilbert curve reordering improves token locality and sparsity. The authors validated SpargeAttn across multiple benchmarks and architectures, including Llama 3.1 for language, Flux and Stable Diffusion 3.5 for image generation, and CogVideoX, Mochi, and Open-Sora-Plan for video synthesis.
The evaluation yielded several critical findings. First, SpargeAttn achieved kernel-level computation speeds 2.5 to 5 times faster than standard full attention and existing sparse baselines, reaching up to 708 tera operations per second on long-context language tasks. Second, it delivered substantial end-to-end inference acceleration across real-world workloads, including a 1.83x end-to-end speedup on the Mochi video generation model (reducing generation time from 1,897 to 1,037 seconds on an NVIDIA L40 GPU) and a 1.73x speedup on Llama 3.1 with 128K context. Third, SpargeAttn maintained full baseline quality across all domains—preserving text perplexity, retrieval fidelity, image realism, and video temporal consistency—whereas competing sparse approaches caused severe visual artifacts or failed retrieval. Fourth, online sparsity scaled positively with context length, increasing from roughly 7% at 8K tokens to 54% at 128K tokens while introducing less than 1% prediction overhead.
These results demonstrate that organizations can sharply cut serving latency and hardware costs for generative AI workloads without risking degradation in model accuracy. Because SpargeAttn is entirely training-free and compatible with low-precision arithmetic, technical leaders can deploy it directly into existing inference pipelines without expensive model fine-tuning or architectural redesigns.
Engineering teams supporting long-context large language models or visual diffusion models should consider piloting SpargeAttn as a drop-in replacement for standard attention kernels. When deploying to vision models, teams should adopt space-filling Hilbert curve token ordering to maximize sparsity benefits. Further tuning should calibrate threshold parameters per model layer and attention head, as sparsity varies across the network architecture. Confidence in these findings is high across the tested hardware and architectures, though production teams should validate parameter settings when porting the kernel to non-NVIDIA hardware or custom attention variants.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). FlashAttention’s IO-aware tiling explains how attention kernels can reduce memory traffic while computing exact attention, a key systems backdrop for SpargeAttention’s kernel design.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). Sparse Transformer introduces the central strategy of reducing attention cost by selectively computing token interactions, providing foundational context for SpargeAttention’s sparse blocks.
- Paper: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision, Jay Shah et al. (2024). FlashAttention-3 shows how asynchronous GPU execution and low-precision arithmetic accelerate attention, helping readers understand the hardware optimizations paired with SpargeAttention’s sparsity.
No sufficiently relevant recommendations were found.
