XAttention: Block Sparse Attention with Antidiagonal Scoring

Ruyi XuGuangxuan XiaoHaofeng HuangJunxian GuoSong Han

article2025ICML150 citations

Proposes a plug-and-play block-sparse attention framework that leverages strided antidiagonal summation to accurately identify critical attention blocks, speeding up long-context attention computation by up to 13.5 times without sacrificing model accuracy across language and video benchmarks.

Listen

Deploying modern artificial intelligence models that process extremely long sequences—such as lengthy documents, hour-long videos, or complex video generation tasks—faces a major bottleneck. The standard mathematical mechanism used to compare information across tokens, known as attention, experiences computational costs that scale quadratically with sequence length. While existing selective or sparse computation methods reduce this burden by evaluating only critical blocks of information, their complex selection processes create heavy overhead that often erodes efficiency or degrades task accuracy.

The article evaluates a new framework called XAttention, designed to accelerate long-context processing without sacrificing model accuracy or requiring retraining. The authors evaluate this approach across diverse language benchmarks (RULER and LongBench), long-duration video understanding (Video-MME), and video generation benchmarks (VBench using HunyuanVideo and Wan2.1).

To solve the efficiency bottleneck, XAttention estimates block importance by summing sampled values along antidiagonal paths (from lower-left to upper-right) across attention blocks. Because these antidiagonal lines naturally intersect the dominant vertical and diagonal patterns in attention maps, they capture essential information without missing critical signals. The framework then prunes non-essential blocks based on an attention threshold, and it optimizes thresholds across individual attention heads using dynamic programming. Crucially, this requires no model fine-tuning or retraining.

The empirical findings demonstrate significant performance and speed advantages. Pattern selection overhead was reduced by up to 24.9 times compared to previous sparse methods. In attention computation, XAttention achieved up to a 13.5-times speedup at sequence lengths of 256,000 tokens while computing only about 7% of total attention values. End-to-end processing speeds improved by approximately 2.6 to 5.1 times across varying context lengths. Across benchmark tasks, accuracy remained on par with, and in several long-context cases exceeded, standard full attention—achieving an average RULER score of 88.47% compared to 87.52% for full attention and 84.15% for prior sparse baselines.

These findings indicate that organizations can significantly cut compute costs, lower inference latency, and deploy advanced long-context multimodal systems onto existing hardware without costly retraining cycles. In video generation workflows, incorporating a brief initial full-attention warmup phase preserves high visual fidelity while delivering over 50% computational sparsity.

Decision-makers considering long-context deployment should evaluate XAttention as a drop-in acceleration layer for high-throughput language and video inference pipelines. Teams should tune the stride sampling parameter (such as stride 8 or 16) to achieve the optimal trade-off between computational sparsity and precision for specific workloads.

Confidence in these findings is supported by consistent validation across multiple model architectures, including LLaMA-3.1, Qwen2-VL, Mistral Nemo, and diffusion-based video models. However, users should note that setting overly large stride intervals (e.g., stride 64) degrades pattern detection accuracy, and generative video tasks require a brief multi-step warmup to prevent initial layout drift.

arXiv: 2503.16428mit-han-lab/x-attention
Cover for XAttention: Block Sparse Attention with Antidiagonal Scoring

Abstract

Long-Context Transformer Models (LCTMs) are vital for real-world applications but suffer high computational costs due to attention's quadratic complexity. Block-sparse attention mitigates this by focusing computation on critical regions, yet existing methods struggle with balancing accuracy and efficiency due to costly block importance measurements. In this paper, we introduce XAttention, a plug-and-play framework that dramatically accelerates long-context inference in Transformers models using sparse attention. XAttention's key innovation is the insight that the sum of antidiagonal values (i.e., from the lower-left to upper-right) in the attention matrix provides a powerful proxy for block importance. This allows for precise identification and pruning of non-essential blocks, resulting in high sparsity and dramatically accelerated inference. Across comprehensive evaluations on demanding long-context benchmarks—including RULER and LongBench for language, VideoMME for video understanding, and VBench for video generation—XAttention achieves accuracy comparable to full attention while delivering substantial computational gains. We demonstrate up to 13.5× acceleration in attention computation. These results underscore XAttention's ability to unlock the practical potential of block sparse attention, paving the way for scalable and efficient deployment of LCTMs in real-world applications.

Table of Contents

  • 1. Introduction
  • 2. Method
  • 2.1. Importance Prediction
  • 2.2. Threshold Block Selection
  • 2.3. Minimum Threshold Prediction
  • 3. Experiments
  • 3.1. Experimental Setup
  • 3.2. Accuracy Results
  • 3.3. Efficiency Results
  • 3.4. Ablation Study
  • 4. Related Work
  • 4.1. Long-Context Large Language Models
  • 4.2. Sparse Attention
  • 4.3. LLM Inference Acceleration
  • 4.4. Recent Works
  • 5. Conclusion
  • Impact Statement
  • ACKNOWLEDGMENTS
  • References
  • A. Video Generation Results on Wan 2.1
  • B. Various Language Model Results on Ruler

Knowls

  1. Knowl 1 — Antidiagonal Scoring Principle for Block-Sparse Attention

    model/method

    XAttention estimates the importance of attention blocks without computing all token-pair dot products or running expensive pattern searches by scoring elements along block antidiagonals.

    In standard Transformer self-attention maps, critical interactions manifest primarily as vertical patterns (representing attention sink tokens or globally salient keys) and slash/diagonal patterns (representing local or relative position dependencies). Sampling along the antidiagonal (from lower-left to upper-right) with stride SS within each attention block of size B×BB \times B ensures two structural guarantees:

    1. Information Preservation: Every token index contributes to at least one antidiagonal sum within the sampled structure.
    2. Pattern Detection: Any vertical line (column) or slash line (subdiagonal) within a block geometrically intersects the antidiagonal slice, ensuring that localized high-attention activations are detected.

    The sum of elements sampled along these strided antidiagonals serves as an efficient proxy for the cumulative attention mass within the block.

  2. Knowl 2 — XAttention Block Selection Algorithm

    algorithm

    The XAttention block selection algorithm partitions the sequence length LL into NB=⌊L/B⌋N_B = \lfloor L/B \rfloor blocks of size BB, gathers strided antidiagonal query and key slices with stride SS, approximates attention probabilities, and identifies the minimal set of blocks exceeding an attention threshold τ\tau.

    Given the full query matrix Q∈RL×dQ \in \mathbb{R}^{L \times d}, key matrix K∈RL×dK \in \mathbb{R}^{L \times d}, block size BB, stride SS, head dimension dhd_h, and cumulative threshold τ∈(0,1]\tau \in (0, 1], the block selection function identifies the smallest set of blocks B\mathcal{B} whose cumulative approximate attention mass satisfies:

    find_blocks(A,τ)=arg⁡min⁡B{∣B∣:∑b∈B∑(i,j)∈bAi,j≥τ}\text{find\_blocks}(A, \tau) = \arg\min_{\mathcal{B}} \left\{ |\mathcal{B}| : \sum_{b \in \mathcal{B}} \sum_{(i,j) \in b} A_{i,j} \ge \tau \right\}

    Input: Query matrix Q∈RL×dQ \in \mathbb{R}^{L \times d}, Key matrix K∈RL×dK \in \mathbb{R}^{L \times d}, block size BB, stride SS, head dimension dhd_h, threshold τ\tau
    Output: Sparse mask MM
    NB←⌊L/B⌋N_B \leftarrow \lfloor L/B \rfloor
    for b=0b = 0 to NB−1N_B - 1 do
        Qslice←Q[b⋅B:(b+1)⋅B,:]Q_{slice} \leftarrow Q[b \cdot B : (b + 1) \cdot B, :]
        Qreshaped←[]Q_{reshaped} \leftarrow []
        for i=S−1i = S - 1 down to 00 do
            QreshapedQ_{reshaped}.append(Qslice[i::S,:]Q_{slice}[i :: S, :]
        end for
        Kreshaped←[]K_{reshaped} \leftarrow []
        for i=0i = 0 to S−1S - 1 do
            KreshapedK_{reshaped}.append(K[i::S,:]K[i :: S, :]
        end for
        Aapprox←Softmax(QreshapedKreshapedTdh⋅S)A_{approx} \leftarrow \text{Softmax}\left( \frac{Q_{reshaped} K_{reshaped}^T}{\sqrt{d_h \cdot S}} \right)
        Mb←find_blocks(Aapprox,τ)M_b \leftarrow \text{find\_blocks}(A_{approx}, \tau)
    end for
    M←concatenate(M0,M1,…,MNB−1)M \leftarrow \text{concatenate}(M_0, M_1, \dots, M_{N_B - 1})
    return MM
  3. Knowl 3 — Minimum Threshold Prediction via Dynamic Programming

    model/method

    XAttention optimizes the trade-off between model accuracy and attention sparsity by dynamically assigning distinct attention retention thresholds τh\tau_h to each attention head h∈{1,2,…,H}h \in \{1, 2, \dots, H\} through dynamic programming.

    Let D[h][m]D[h][m] be a dynamic programming table storing the maximal performance achievable across the first hh heads after performing exactly m∈{1,2,…,M}m \in \{1, 2, \dots, M\} threshold reduction steps. The recurrence relation is:

    D[h][m]=max⁡(D[h−1][m],P(h,m))D[h][m] = \max(D[h - 1][m], P(h, m))

    where P(h,m)P(h, m) denotes the downstream task performance when head hh's threshold is reduced by one step relative to the state D[h−1][m−1]D[h - 1][m - 1].

    At each adjustment step mm, a head's threshold is decreased multiplicatively by 10%:

    th(m)=th(m−1)×0.9th(m) = th(m - 1) \times 0.9

    Running this procedure with M=1000M = 1000 steps yields head-specific thresholds that lower the average threshold from an initial τ=0.9\tau = 0.9 down to an average of ≈0.8\approx 0.8, reducing compute density while preserving task accuracy.

  4. Knowl 4 — Full-Attention Denoising Warmup for Video Diffusion Transformers

    model/method

    When applying XAttention to Diffusion Transformer (DiT) architectures (such as HunyuanVideo or Wan2.1) operating with non-causal attention across long spatiotemporal video token sequences, applying block sparsity from the very first denoising step causes layout shifts and lower fidelity compared to full attention.

    Because initial diffusion denoising steps establish coarse spatial layouts and global scene structure, XAttention incorporates a full-attention warmup phase: the model executes dense full attention for the first 5 denoising steps, and then switches to XAttention with antidiagonal scoring (e.g., stride S=8S = 8, τ∈{0.90,0.95}\tau \in \{0.90, 0.95\}) for the remaining denoising steps (e.g., steps 6 through 50). This eliminates layout drift while retaining over 50% attention sparsity across the generation run.

  5. Knowl 5 — Long-Context Language Benchmark Evaluation on RULER and LongBench

    data/table

    Evaluated on the synthetic RULER long-context benchmark and the real-world LongBench benchmark using Llama-3.1-8B-Instruct, XAttention configured with stride S∈{8,16}S \in \{8, 16\} and dynamically predicted thresholds outperforms dense FlashAttention and training-free sparse attention baselines (FlexPrefill, MInference, SeerAttention).

    Method 4k 8k 16k 32k 64k 128k Avg.
    Full Attention 96.74 94.03 92.02 84.17 81.32 76.89 87.52
    FlexPrefill 95.99 93.67 92.73 88.14 81.14 74.67 87.72
    MInference 96.54 94.06 91.37 85.79 83.03 54.12 84.15
    SeerAttn 95.32 92.14 92.20 88.05 83.30 72.37 87.23
    XAttn S=8 96.83 94.07 93.17 90.75 84.08 72.31 88.47
    XAttn S=16 96.11 93.95 93.56 90.64 83.12 71.11 88.08

    On the LongBench suite across 16 tasks (Single-Doc QA, Multi-Doc QA, Summarization, Few-shot Learning, Code), XAttention achieves an average score of 40.60, outperforming Full Attention (40.34), MInference (40.30), and FlexPrefill (36.83).

  6. Knowl 6 — Prefill Acceleration and End-to-End Speedup

    data/table

    XAttention achieves significant computational speedups during the prefill stage across context lengths from 8k to 256k tokens. Its strided antidiagonal block selection is training-free and computes block masks without search overhead, executing pattern selection up to 24.9×\times faster than MInference and 5.9×\times faster than FlexPrefill.

    At 256k tokens, prefill attention computation achieves up to a 13.5×\times speedup (S=16S=16, 7.32% density) and 9.8×\times speedup (S=8S=8, 6.89% density) relative to FlashInfer FlashAttention.

    End-to-end prefill speedups on Llama-3.1-8B-Instruct on RULER across context lengths are:

    Setting 8k 16k 32k 64k 128k 256k
    XAttn S=8 2.59× 3.04× 3.96× 4.38× 4.67× 4.93×
    XAttn S=16 2.89× 3.60× 4.40× 4.82× 4.94× 5.12×

    Attention block density decreases with sequence length, dropping from ∼\sim52% at 4k to ∼\sim6.9% (S=8S=8) and ∼\sim7.3% (S=16S=16) at 128k context length.

  7. Knowl 7 — Video Understanding Performance on Video-MME

    data/table

    Evaluated on the Video-MME benchmark (900 videos up to 1 hour, sampled at 1 frame per second) using Qwen2-VL-7B-Instruct with stride S=16S = 16 and threshold τ=0.9\tau = 0.9, XAttention outperforms other sparse attention approaches and matches or exceeds full dense attention, particularly on long-duration videos.

    Method Short (%) Medium (%) Long (%) Overall (%)
    w/o sub w/ sub w/o sub w/ sub w/o sub w/ sub w/o sub w/ sub
    Full Attention 72.1 78.1 63.9 69.4 55.1 60.2 63.7 69.2
    MInference 71.7 77.6 62.3 67.9 55.2 59.8 63.1 68.4
    FlexPrefill 71.4 77.4 62.6 68.3 53.8 57.3 62.6 67.7
    XAttention 71.9 78.8 62.6 68.5 55.7 60.3 63.3 69.1
  8. Knowl 8 — Video Generation Fidelity and Sparsity on HunyuanVideo and Wan2.1

    data/table

    Evaluated across 946 VBench text prompts generating 720×1280720 \times 1280 resolution, 129-frame videos using 50 denoising steps with a 5-step full-attention warmup and stride S=8S = 8, XAttention maintains high fidelity compared to dense full attention across Diffusion Transformer architectures while achieving over 50% attention sparsity.

    Model Threshold τ\tau PSNR (↑\uparrow) SSIM (↑\uparrow) LPIPS (↓\downarrow) Density (%, ↓\downarrow)
    HunyuanVideo 0.90 21.5 0.767 0.215 34.4%
    HunyuanVideo 0.95 23.5 0.822 0.155 45.5%
    Wan2.1 0.90 21.2 0.745 0.231 39.2%
    Wan2.1 0.95 22.7 0.819 0.129 33.6%
  9. Knowl 9 — Ablation of Sparsity Pattern, Stride, and Selection Mechanisms

    data/table

    Ablations on Llama-3.1-8B-Instruct on the RULER benchmark validate the core design choices of XAttention:

    1. Pattern Geometry (at 32k / average accuracy and density): The antidiagonal pattern substantially outperforms diagonal and random sampling while selecting fewer blocks.

      • Antidiagonal (S=8S=8): 90.75% (32k), 88.47% (Avg), 20.97% density.
      • Diagonal (S=8S=8): 76.47% (32k), 81.06% (Avg), 24.47% density.
      • Random (S=8S=8): 82.53% (32k), 82.48% (Avg), 27.57% density.
    2. Stride Size SS:

      • S=4S=4: 88.89% Avg, 21.09% density.
      • S=8S=8: 88.47% Avg, 20.97% density.
      • S=16S=16: 88.08% Avg, 27.93% density.
      • S=64S=64: 81.21% Avg, 39.88% density (too coarse to capture slash attention patterns).
    3. Selection Strategy (S=8S=8):

      • Dynamic Threshold: 88.47% Avg, 20.97% density.
      • Top-K (K=8192K=8192): 84.13% Avg, 19.92% density.
      • Top-Ratio (Ratio=27%): 85.42% Avg, 21.00% density.
    4. Threshold Prediction (S=8S=8):

      • Dynamically predicted minimum τ\tau: 88.47% Avg, 20.97% density.
      • Fixed τ=0.9\tau = 0.9: 84.96% Avg, 26.13% density.
  10. Knowl 10 — Generalization Across Multiple Large Language Model Architectures

    data/table

    XAttention generalizes across multiple long-context LLM architectures on the RULER benchmark (4k to 128k context lengths), consistently matching dense attention performance and outperforming existing sparse baselines.

    Model Method Avg (4k–128k) Delta
    Mistral Nemo 12B Full Attention 67.97 –
    MInference 64.49 -3.48
    FlexPrefill 64.61 -3.36
    XAttn S=4 67.92 -0.05
    XAttn S=16 67.47 -0.50
    Phi-3.5 Mini 3.8B Full Attention 84.68 –
    MInference 81.89 -2.79
    FlexPrefill 82.83 -1.85
    XAttn S=4 84.86 +0.18
    XAttn S=16 83.82 -0.86
    Qwen2.5 7B Full Attention 77.84 –
    MInference 74.02 -3.82
    FlexPrefill 75.10 -2.74
    XAttn S=4 77.75 -0.09
    XAttn S=16 77.21 -0.63

Coverage note — None was omitted; all substantive algorithmic designs, formulas, experimental benchmarks across text, video understanding, and video generation, efficiency profiles, ablations, and generalization results have been fully represented.

References

  1. 1.Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508, 2023.
  2. 2.Beltagy, I., Peters, M. E., and Cohan, A. Longformer: The long-document transformer, 2020. arXiv:2004.05150.
  3. 3.Burtsev, M. S., Kuratov, Y., Peganov, A., and Sapunov, G. V. Memory transformer, 2021. URL https://arxiv.org/abs/2006.11527.
  4. 4.Chen, S., Wong, S., Chen, L., and Tian, Y. Extending context window of large language models via positional interpolation, 2023. arXiv: 2306.15595.
  5. 5.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. 2019a.
  6. 6.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers, 2019b. URL https://arxiv.org/abs/1904.10509.
  7. 7.Dao, T. FlashAttention-2: Faster attention with better parallelism and work partitioning, 2023.
  8. 8.Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C. FlashAttention: Fast and memory-efficient exact attention with IO-awareness, 2022. arXiv:2205.14135.
  9. 9.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Roziere, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J., Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., El-Arini, K., Iyer, K., Malik, K., Chiu, K., Bhalla, K., Rantala-Yeary, L., van der Maaten, L., Chen, L., Tan, L., Jenkins, L., Martin, L., Madaan, L., Malo, L., Blecher, L., Landzaat, L., de Oliveira, L., Muzzi, M., Pasupuleti, M., Singh, M., Paluri, M., Kardas, M., Oldham, M., Rita, M., Pavlova, M., Kambadur, M., Lewis, M., Si, M., Singh, M. K., Hassan, M., Goyal, N., Torabi, N., Bashlykov, N., Bogoychev, N., Chatterji, N., Duchenne, O., Celebi, O., Alrassy, P., Zhang, P., Li, P., Vasic, P., Weng, P., Bhargava, P., Dubal, P., Krishnan, P., Koura, P. S., Xu, P., He, Q., Dong, Q., Srinivasan, R., Ganapathy, R., Calderer, R., Cabral, R. S., Stojnic, R., Raileanu, R., Girdhar, R., Patel, R., Sauvestre, R., Polidoro, R., Sumbaly, R., Taylor, R., Silva, R., Hou, R., Wang, R., Hosseini, S., Chennabasappa, S., Singh, S., Bell, S., Kim, S. S., Edunov, S., Nie, S., Narang, S., Raparthy, S., Shen, S., Wan, S., Bhosale, S., Zhang, S., Vandenhende, S., Batra, S., Whitman, S., Sootla, S., Collot, S., Gururangan, S., Borodinsky, S., Herman, T., Fowler, T., Sheasha, T., Georgiou, T., Scialom, T., Speckbacher, T., Mihaylov, T., Xiao, T., Karn, U., Goswami, V., Gupta, V., Ramanathan, V., Kerkez, V., Gonguet, V., Do, V., Vogeti, V., Petrovic, V., Chu, W., Xiong, W., Fu, W., Meers, W., Martinet, X., Wang, X., Tan, X. E., Xie, X., Jia, X., Wang, X., Goldschlag, Y., Gaur, Y., Babaei, Y., Wen, Y., Song, Y., Zhang, Y., Li, Y., Mao, Y., Coudert, Z. D., Yan, Z., Chen, Z., Papakipos, Z., Singh, A., Grattafiori, A., Jain, A., Kelsey, A., Shajnfeld, A., Gangidi, A., Victoria, A., Goldstand, A., Menon, A., Sharma, A., Boesenberg, A., Vaughan, A., Baevski, A., Feinstein, A., Kallet, A., Sangani, A., Yunus, A., Lupu, A., Alvarado, A., Caples, A., Gu, A., Ho, A., Poulton, A., Ryan, A., Ramchandani, A., Franco, A., Saraf, A., Chowdhury, A., Gabriel, A., Bharambe, A., Eisenman, A., Yazdan, A., James, B., Maurer, B., Leonhardi, B., Huang, B., Loyd, B., Paola, B. D., Paranjape, B., Liu, B., Wu, B., Ni, B., Hancock, B., Wasti, B., Spence, B., Stojkovic, B., Gamido, B., Montalvo, B., Parker, C., Burton, C., Mejia, C., Wang, C., Kim, C., Zhou, C., Hu, C., Chu, C.-H., Cai, C., Tindal, C., Feichtenhofer, C., Civin, D., Beaty, D., Kreymer, D., Li, D., Wyatt, D., Adkins, D., Xu, D., Testuggine, D., David, D., Parikh, D., Liskovich, D., Foss, D., Wang, D., Le, D., Holland, D., Dowling, E., Jamil, E., Montgomery, E., Presani, E., Hahn, E., Wood, E., Brinkman, E., Arcaute, E., Dunbar, E., Smothers, E., Sun, F., Kreuk, F., Tian, F., Ozgenel, F., Caggioni, F., Guzman, F., Kanayet, F., Seide, F., Florez, G. M., Schwarz, G., Badeer, G., Swee, G., Halpern, G., Thattai, G., Herman, G., Sizov, G., Guangyi, Zhang, Lakshminarayanan, G., Shojanazeri, H., Zou, H., Wang, H., Zha, H., Habeeb, H., Rudolph, H., Suk, H., Aspegren, H., Goldman, H., Damlaj, I., Molybog, I., Tufanov, I., Veliche, I.-E., Gat, I., Weissman, J., Geboski, J., Kohli, J., Asher, J., Gaya, J.-B., Marcus, J., Tang, J., Chan, J., Zhen, J., Reizenstein, J., Teboul, J., Zhong, J., Jin, J., Yang, J., Cummings, J., Carvill, J., Shepard, J., McPhie, J., Torres, J., Ginsburg, J., Wang, J., Wu, K., U, K. H., Saxena, K., Prasad, K., Khandelwal, K., Zand, K., Matosich, K., Veeraraghavan, K., Michelena, K., Li, K., Huang, K., Chawla, K., Lakhotia, K., Huang, K., Chen, L., Garg, L., A, L., Silva, L., Bell, L., Zhang, L., Guo, L., Yu, L., Moshkovich, L., Wehrstedt, L., Khabsa, M., Avalani, M., Bhatt, M., Tsimpoukelli, M., Mankus, M., Hasson, M., Lennie, M., Reso, M., Groshev, M., Naumov, M., Lathi, M., Keneally, M., Seltzer, M. L., Valko, M., Restrepo, M., Patel, M., Vyatskov, M., Samvelyan, M., Clark, M., Macey, M., Wang, M., Hermoso, M. J., Metanat, M., Rastegari, M., Bansal, M., Santhanam, N., Parks, N., White, N., Bawa, N., Singhal, N., Egebo, N., Usunier, N., Laptev, N. P., Dong, N., Zhang, N., Cheng, N., Chernoguz, O., Hart, O., Salpekar, O., Kalinli, O., Kent, P., Parekh, P., Saab, P., Balaji, P., Rittner, P., Bontrager, P., Roux, P., Dollar, P., Zvyagina, P., Ratanchandani, P., Yuvraj, P., Liang, Q., Alao, R., Rodriguez, R., Ayub, R., Murthy, R., Nayani, R., Mitra, R., Li, R., Hogan, R., Battey, R., Wang, R., Maheswari, R., Howes, R., Rinott, R., Bondu, S. J., Datta, S., Chugh, S., Hunt, S., Dhillon, S., Sidorov, S., Pan, S., Verma, S., Yamamoto, S., Ramaswamy, S., Lindsay, S., Lindsay, S., Feng, S., Lin, S., Zha, S. C., Shankar, S., Zhang, S., Zhang, S., Wang, S., Agarwal, S., Sajuyigbe, S., Chintala, S., Max, S., Chen, S., Kehoe, S., Satterfield, S., Govindaprasad, S., Gupta, S., Cho, S., Virk, S., Subramanian, S., Choudhury, S., Goldman, S., Remez, T., Glaser, T., Best, T., Kohler, T., Robinson, T., Li, T., Zhang, T., Matthews, T., Chou, T., Shaked, T., Vontimitta, V., Ajayi, V., Montanez, V., Mohan, V., Kumar, V. S., Mangla, V., Albiero, V., Ionescu, V., Poenaru, V., Mihailescu, V. T., Ivanov, V., Li, W., Wang, W., Jiang, W., Bouaziz, W., Constable, W., Tang, X., Wang, X., Wu, X., Wang, X., Xia, X., Wu, X., Gao, X., Chen, Y., Hu, Y., Jia, Y., Qi, Y., Li, Y., Zhang, Y., Zhang, Y., Adi, Y., Nam, Y., Yu, Wang, Hao, Y., Qian, Y., He, Y., Rait, Z., DeVito, Z., Rosnbrick, Z., Wen, Z., Yang, Z., and Zhao, Z. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  10. 10.Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., Zhang, M., et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024.
  11. 11.Gao, Y., Zeng, Z., Du, D., Cao, S., So, H. K.-H., Cao, T., Yang, F., and Yang, M. Seerattention: Learning intrinsic sparse attention in your llms. arXiv preprint arXiv:2410.13276, 2024.
  12. 12.Guo, J., Tang, H., Yang, S., Zhang, Z., Liu, Z., and Han, S. Block Sparse Attention. https://github.com/mit-han-lab/Block-Sparse-Attention, 2024.
  13. 13.Hong, K., Dai, G., Xu, J., Mao, Q., Li, X., Liu, J., Chen, K., Dong, Y., and Wang, Y. Flashdecoding++: Faster large language model inference on gpus, 2024.
  14. 14.Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., and Ginsburg, B. Ruler: What’s the real context size of your long-context language models? arXiv preprint arXiv:2404.06654, 2024.
  15. 15.Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., and Liu, Z. VBench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.
  16. 16.Jiang, H., Li, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L. Minference 1.0: Accelerating pre-filling for longcontext llms via dynamic sparse attention. arXiv preprint arXiv:2407.02490, 2024.
  17. 17.Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Yuan, J., Long, Y., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., He, Z., Xu, Z., Zhou, Z., Xu, Z., Tao, Y., Lu, Q., Liu, S., Zhou, D., Wang, H., Yang, Y., Wang, D., Liu, Y., Jiang, J., and Zhong, C. Hunyuanvideo: A systematic framework for large video generative models, 2025. URL https://arxiv.org/abs/2412.03603.
  18. 18.Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. Efficient memory management for large language model serving with pagedattention, 2023.
  19. 19.Lai, X., Lu, J., Luo, Y., Ma, Y., and Zhou, X. Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=OfjIlbelrT.
  20. 20.Leviathan, Y., Kalman, M., and Matias, Y. Selective attention improves transformer, 2024. URL https://arxiv.org/abs/2410.02703.
  21. 21.Li, M., Cai, T., Cao, J., Zhang, Q., Cai, H., Bai, J., Jia, Y., Liu, M.-Y., Li, K., and Han, S. Distrifusion: Distributed parallel inference for high-resolution diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  22. 22.Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L. Video-llava: Learning united visual representation by alignment before projection, 2023.
  23. 23.Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration, 2024.
  24. 24.Lin*, Y., Tang*, H., Yang*, S., Zhang, Z., Xiao, G., Gan, C., and Han, S. Qserve: W4a8kv4 quantization and system co-design for efficient llm serving. arXiv preprint arXiv:2405.04532, 2024.
  25. 25.Liu, H., Zaharia, M., and Abbeel, P. Ring attention with blockwise transformers for near-infinite context, 2023.
  26. 26.Lu, E., Jiang, Z., Liu, J., Du, Y., Jiang, T., Hong, C., Liu, S., He, W., Yuan, E., Wang, Y., Huang, Z., Yuan, H., Xu, S., Xu, X., Lai, G., Chen, Y., Zheng, H., Yan, J., Su, J., Wu, Y., Zhang, N. Y., Yang, Z., Zhou, X., Zhang, M., and Qiu, J. Moba: Mixture of block attention for long-context llms, 2025. URL https://arxiv.org/abs/2502.13189.
  27. 27.OpenAI. Gpt-4 technical report, 2023.
  28. 28.Oren, M., Hassid, M., Yarden, N., Adi, Y., and Schwartz, R. Transformers are multi-state rnns, 2024. URL https://arxiv.org/abs/2401.06104.
  29. 29.Peebles, W. and Xie, S. Scalable diffusion models with transformers, 2023. URL https://arxiv.org/abs/2212.09748.
  30. 30.Peng, B., Quesnelle, J., Fan, H., and Shippole, E. Yarn: Efficient context window extension of large language models, 2023.
  31. 31.Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S. Quest: Query-aware sparsity for efficient long-context llm inference, 2024.
  32. 32.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  33. 33.Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  34. 34.Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y. Retrieval head mechanistically explains long-context factuality, 2024.
  35. 35.Xi, H., Yang, S., Zhao, Y., Xu, C., Li, M., Li, X., Lin, Y., Cai, H., Zhang, J., Li, D., Chen, J., Stoica, I., Keutzer, K., and Han, S. Sparse videogen: Accelerating video diffusion transformers with spatial-temporal sparsity, 2025. URL https://arxiv.org/abs/2502.01776.
  36. 36.Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M. Infllm: Training-free longcontext extrapolation for llms with an efficient context memory, 2024a. URL https://arxiv.org/abs/2402.04617.
  37. 37.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, 2023a.
  38. 38.Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. Efficient streaming language models with attention sinks. arXiv, 2023b.
  39. 39.Xiao, G., Yin, T., Freeman, W. T., Durand, F., and Han, S. Fastcomposer: Tuning-free multi-subject image generation with localized attention, 2023c. URL https://arxiv.org/abs/2305.10431.
  40. 40.Xiao, G., Tang, J., Zuo, J., Guo, J., Yang, S., Tang, H., Fu, Y., and Han, S. Duoattention: Efficient long-context llm inference with retrieval and streaming heads. arXiv, 2024b.
  41. 41.Yang, X., Chen, T., and Chen, B. Ape: Faster and longer context-augmented generation via adaptive parallel encoding. In ICLR 2025, 2025.
  42. 42.Ye, Z., Lai, R., Lu, R., Lin, C.-Y., Zheng, S., Chen, L., Chen, T., and Ceze, L. Cascade inference: Memory bandwidth efficient shared prefix batch decoding. https://flashinfer.ai/2024/01/08/cascade-inference.html, Jan 2024. URL https://flashinfer.ai/2024/01/08/cascade-inference.html. Accessed on 2024-02-01.
  43. 43.Yuan, J., Gao, H., Dai, D., Luo, J., Zhao, L., Zhang, Z., Xie, Z., Wei, Y. X., Wang, L., Xiao, Z., Wang, Y., Ruan, C., Zhang, M., Liang, W., and Zeng, W. Native sparse attention: Hardware-aligned and natively trainable sparse attention, 2025. URL https://arxiv.org/abs/2502.11089.
  44. 44.Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A. Big bird: Transformers for longer sequences. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.-F., and Lin, H.-T. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020. Curran Associates, Inc., 2020.
  45. 45.Zhang, P., Chen, Y., Su, R., Ding, H., Stoica, I., Liu, Z., and Zhang, H. Fast video generation with sliding tile attention, 2025. URL https://arxiv.org/abs/2502.04507.
  46. 46.Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Re, C., Barrett, C., Wang, Z., and Chen, B. H2o: Heavy-hitter oracle for efficient generative inference of large language models, 2023.

Citation

MLA
Xu, R., et al. “XAttention: Block Sparse Attention with Antidiagonal Scoring”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.16428.
APA
Xu, R., Xiao, G., Huang, H., Guo, J., & Han, S. (2025). XAttention: Block Sparse Attention with Antidiagonal Scoring. arXiv. https://doi.org/10.48550/arxiv.2503.16428
Chicago
Xu, R., G. Xiao, H. Huang, J. Guo, and S. Han. 2025. “XAttention: Block Sparse Attention with Antidiagonal Scoring”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.16428.
Harvard
Xu, R. et al. (2025) “XAttention: Block Sparse Attention with Antidiagonal Scoring”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.16428.
Vancouver
1. Xu R, Xiao G, Huang H, Guo J, Han S (2025) XAttention: Block Sparse Attention with Antidiagonal Scoring. https://doi.org/10.48550/arxiv.2503.16428

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.16428,
  doi = {10.48550/ARXIV.2503.16428},
  url = {https://arxiv.org/abs/2503.16428},
  author = {Xu, Ruyi and Xiao, Guangxuan and Huang, Haofeng and Guo, Junxian and Han, Song},
  keywords = {Computation and Language (cs.CL), Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {XAttention: Block Sparse Attention with Antidiagonal Scoring},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/