XAttention: Block Sparse Attention with Antidiagonal Scoring
Ruyi XuGuangxuan XiaoHaofeng HuangJunxian GuoSong Han
Proposes a plug-and-play block-sparse attention framework that leverages strided antidiagonal summation to accurately identify critical attention blocks, speeding up long-context attention computation by up to 13.5 times without sacrificing model accuracy across language and video benchmarks.
Deploying modern artificial intelligence models that process extremely long sequences—such as lengthy documents, hour-long videos, or complex video generation tasks—faces a major bottleneck. The standard mathematical mechanism used to compare information across tokens, known as attention, experiences computational costs that scale quadratically with sequence length. While existing selective or sparse computation methods reduce this burden by evaluating only critical blocks of information, their complex selection processes create heavy overhead that often erodes efficiency or degrades task accuracy.
The article evaluates a new framework called XAttention, designed to accelerate long-context processing without sacrificing model accuracy or requiring retraining. The authors evaluate this approach across diverse language benchmarks (RULER and LongBench), long-duration video understanding (Video-MME), and video generation benchmarks (VBench using HunyuanVideo and Wan2.1).
To solve the efficiency bottleneck, XAttention estimates block importance by summing sampled values along antidiagonal paths (from lower-left to upper-right) across attention blocks. Because these antidiagonal lines naturally intersect the dominant vertical and diagonal patterns in attention maps, they capture essential information without missing critical signals. The framework then prunes non-essential blocks based on an attention threshold, and it optimizes thresholds across individual attention heads using dynamic programming. Crucially, this requires no model fine-tuning or retraining.
The empirical findings demonstrate significant performance and speed advantages. Pattern selection overhead was reduced by up to 24.9 times compared to previous sparse methods. In attention computation, XAttention achieved up to a 13.5-times speedup at sequence lengths of 256,000 tokens while computing only about 7% of total attention values. End-to-end processing speeds improved by approximately 2.6 to 5.1 times across varying context lengths. Across benchmark tasks, accuracy remained on par with, and in several long-context cases exceeded, standard full attention—achieving an average RULER score of 88.47% compared to 87.52% for full attention and 84.15% for prior sparse baselines.
These findings indicate that organizations can significantly cut compute costs, lower inference latency, and deploy advanced long-context multimodal systems onto existing hardware without costly retraining cycles. In video generation workflows, incorporating a brief initial full-attention warmup phase preserves high visual fidelity while delivering over 50% computational sparsity.
Decision-makers considering long-context deployment should evaluate XAttention as a drop-in acceleration layer for high-throughput language and video inference pipelines. Teams should tune the stride sampling parameter (such as stride 8 or 16) to achieve the optimal trade-off between computational sparsity and precision for specific workloads.
Confidence in these findings is supported by consistent validation across multiple model architectures, including LLaMA-3.1, Qwen2-VL, Mistral Nemo, and diffusion-based video models. However, users should note that setting overly large stride intervals (e.g., stride 64) degrades pattern detection accuracy, and generative video tasks require a brief multi-step warmup to prevent initial layout drift.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). This paper establishes the foundational IO-aware tiling and block-sparse GPU computation methods that underpin hardware-efficient attention frameworks like XAttention.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). This work introduces foundational sparse attention factorizations to break the quadratic complexity of long-sequence Transformer processing.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). This paper demonstrates how dynamic contextual sparsity can be identified and exploited to accelerate Transformer inference without compromising model accuracy.
- Paper: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, Tri Dao (2023). This work optimizes parallel work partitioning and memory tiling strategies essential for implementing high-throughput custom attention kernels.
- Paper: S2-Attention: Hardware-Aware Context Sharding Among Attention Heads, Xihui Lin et al. (2024). This work addresses practical GPU memory access bottlenecks in sparse attention by developing hardware-aware tiled execution kernels.
- Paper: LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing, Wen Zan et al. (2026). This work tackles the quadratic scoring overhead and scattered memory bottlenecks in long-context sparse attention via streaming-aware hierarchical indexing.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). This paper extends hardware-aligned sparse attention principles into a natively trainable, hierarchical architecture optimized for ultra-long context LLMs.
- Paper: Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference, Dhruv Deshmukh et al. (2025). This work introduces layer-wise reuse of sparse attention patterns to further reduce dynamic scoring overhead during long-context inference.
- Paper: Language Models Can Control Their Own Attention, Namgyu Ho et al. (2026). This paper explores an alternative paradigm where models dynamically declare their own sparse attention spans at decode time without auxiliary scoring proxies.
