SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

Jintao ZhangChendong XiangHaofeng HuangJia WeiHaocheng XiJun ZhuJianfei Chen

article2025ICML63 citations

Proposes SpargeAttn, a training-free sparse attention mechanism that uses a two-stage online filter to accelerate inference across language, image, and video models without degrading output quality.

Listen

As modern artificial intelligence models process increasingly large context windows—such as high-resolution video frames and 128,000-token text sequences—the standard attention mechanism creates severe computational bottlenecks. Because attention scales quadratically with sequence length, generating long-form content requires massive GPU memory and latency. While sparse attention techniques attempt to reduce this overhead by skipping near-zero values, existing methods generally suffer from two flaws: they rely on task-specific patterns that fail to generalize across text, image, and video domains, or they impose prediction overheads that erode practical speed gains.

The article introduces and evaluates SpargeAttn, a universal, training-free, sparse, and quantized attention framework. The primary objective is to demonstrate that a single, pattern-free sparse attention mechanism can significantly accelerate inference across diverse generative modalities without retraining or degrading model output quality.

To achieve this, the authors developed a two-stage online filtering system integrated with low-precision arithmetic. In the first stage, the method predicts sparse blocks on the fly by selectively compressing token blocks that exhibit high internal similarity, while leaving non-similar blocks fully computed to protect essential information. In the second stage, an online filter identifies and omits negligible matrix updates during computation at the GPU warp level. For visual tasks, a space-filling Hilbert curve reordering improves token locality and sparsity. The authors validated SpargeAttn across multiple benchmarks and architectures, including Llama 3.1 for language, Flux and Stable Diffusion 3.5 for image generation, and CogVideoX, Mochi, and Open-Sora-Plan for video synthesis.

The evaluation yielded several critical findings. First, SpargeAttn achieved kernel-level computation speeds 2.5 to 5 times faster than standard full attention and existing sparse baselines, reaching up to 708 tera operations per second on long-context language tasks. Second, it delivered substantial end-to-end inference acceleration across real-world workloads, including a 1.83x end-to-end speedup on the Mochi video generation model (reducing generation time from 1,897 to 1,037 seconds on an NVIDIA L40 GPU) and a 1.73x speedup on Llama 3.1 with 128K context. Third, SpargeAttn maintained full baseline quality across all domains—preserving text perplexity, retrieval fidelity, image realism, and video temporal consistency—whereas competing sparse approaches caused severe visual artifacts or failed retrieval. Fourth, online sparsity scaled positively with context length, increasing from roughly 7% at 8K tokens to 54% at 128K tokens while introducing less than 1% prediction overhead.

These results demonstrate that organizations can sharply cut serving latency and hardware costs for generative AI workloads without risking degradation in model accuracy. Because SpargeAttn is entirely training-free and compatible with low-precision arithmetic, technical leaders can deploy it directly into existing inference pipelines without expensive model fine-tuning or architectural redesigns.

Engineering teams supporting long-context large language models or visual diffusion models should consider piloting SpargeAttn as a drop-in replacement for standard attention kernels. When deploying to vision models, teams should adopt space-filling Hilbert curve token ordering to maximize sparsity benefits. Further tuning should calibrate threshold parameters per model layer and attention head, as sparsity varies across the network architecture. Confidence in these findings is high across the tested hardware and architectures, though production teams should validate parameter settings when porting the kernel to non-NVIDIA hardware or custom attention variants.

Zhang et al (2025).pdf

No sufficiently relevant recommendations were found.

Cover for SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

Abstract

An efficient attention implementation is essential for large models due to its quadratic time complexity. Fortunately, attention commonly exhibits sparsity, i.e., many values in the attention map are near zero, allowing for the omission of corresponding computations. Many studies have utilized the sparse pattern to accelerate attention. However, most existing works focus on optimizing attention within specific models by exploiting certain sparse patterns of the attention map. A universal sparse attention that guarantees both the speedup and end-to-end performance of diverse models remains elusive. In this paper, we propose SpargeAttn, a universal sparse and quantized attention for any model. Our method uses a two-stage online filter: in the first stage, we rapidly and accurately predict the attention map, enabling the skip of some matrix multiplications in attention. In the second stage, we design an online softmax-aware filter that incurs no extra overhead and further skips some matrix multiplications. Experiments show that our method significantly accelerates diverse models, including language, image, and video generation, without sacrificing end-to-end metrics. The code is available at https://github.com/thu-ml/SpargeAttn.

Citation

MLA
Zhang, J., et al. “SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.18137.
APA
Zhang, J., Xiang, C., Huang, H., Wei, J., Xi, H., Zhu, J., & Chen, J. (2025). SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference. arXiv. https://doi.org/10.48550/arxiv.2502.18137
Chicago
Zhang, J., C. Xiang, H. Huang, et al. 2025. “SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.18137.
Harvard
Zhang, J. et al. (2025) “SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.18137.
Vancouver
1. Zhang J, Xiang C, Huang H, Wei J, Xi H, Zhu J, Chen J (2025) SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference. https://doi.org/10.48550/arxiv.2502.18137

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.18137,
  doi = {10.48550/ARXIV.2502.18137},
  url = {https://arxiv.org/abs/2502.18137},
  author = {Zhang, Jintao and Xiang, Chendong and Huang, Haofeng and Wei, Jia and Xi, Haocheng and Zhu, Jun and Chen, Jianfei},
  keywords = {Machine Learning (cs.LG), Artificial Intelligence (cs.AI), Computer Vision and Pattern Recognition (cs.CV), Performance (cs.PF), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/