SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Yizhao GaoShuming GuoShijie CaoYuqing XiaYu ChengLei WangLingxiao MaYutao SunTianzhu YeLi Dong

article2026arXiv28 citations

Introduces SeerAttention-R, a lightweight plug-in sparse attention framework for auto-regressive decoding that preserves reasoning accuracy on benchmarks like AIME while achieving up to a 9x inference speedup over FlashAttention-3 on H100 GPUs.

Listen

Modern artificial intelligence models solve complex problems by generating extended step-by-step reasoning sequences. While these longer responses substantially enhance accuracy, they introduce serious computational bottlenecks. As the generated text grows, the model must scan an expanding memory history at each step, causing memory demand and computing costs to grow dramatically. Sparse attention can mitigate these expenses by focusing only on essential tokens, but existing sparse techniques struggle to maintain precision during prolonged generation steps.

To address this challenge, the article introduces SeerAttention-R, a sparse attention framework designed to accelerate the step-by-step decoding phase in reasoning models. The objective of the article is to demonstrate that a lightweight, learned gating module can accurately select critical memory blocks during inference without retraining the underlying base model, thereby preserving reasoning accuracy while significantly speeding up computation.

To evaluate this approach, the authors trained a lightweight attention gating module on a dataset of 400 million tokens while freezing the original model parameters. They evaluated four open-source reasoning models—including Qwen3 variants ranging from 4 billion to 14 billion parameters and DeepSeek-R1-Distill-Qwen-14B—across challenging mathematical and general reasoning benchmarks such as AIME24, AIME25, MATH-500, and GPQA-Diamond. The team also implemented a specialized hardware execution kernel using TileLang and Triton to benchmark processing speedups on NVIDIA H100 GPUs against standard full attention baselines.

The findings show that attention in reasoning models is naturally sparse; focusing on only 2,000 to 4,000 important tokens is sufficient to match the accuracy of standard full-attention models that evaluate the entire context. In head-to-head comparisons, SeerAttention-R consistently outperformed leading training-free sparse methods, maintaining near-lossless accuracy even under large memory block sizes of 64 and 128 tokens. Furthermore, larger models demonstrated greater resilience to sparsity, closing performance gaps more easily than smaller models. At the hardware execution level, the customized decoding kernel achieved near-theoretical speedups of up to 9 times over standard FlashAttention-3 at 90 percent sparsity on long sequences.

These results demonstrate that organizations can dramatically reduce the compute time and operational costs of serving long-reasoning artificial intelligence models without sacrificing task performance. The lightweight nature of the training process—requiring only around 12 GPU hours for an 8-billion-parameter model—allows existing pretrained models to adopt sparse decoding with minimal post-training investment. Moreover, the framework avoids accuracy penalties that typically force competing sparse methods to generate excessively long, error-prone reasoning chains.

Decision-makers should consider piloting SeerAttention-R as an add-on optimization for large-scale reasoning deployments, particularly where long-sequence generation drives high inference costs. Next development steps should focus on integrating this sparse kernel into production serving systems, evaluating automatic threshold tuning across varying task complexities, and unifying prefill and decoding sparsity into a single end-to-end framework.

While the findings demonstrate high confidence in kernel-level speedups and reasoning accuracy across standard benchmarks, full end-to-end system throughput in live production environments remains to be confirmed. Continued validation across broader reasoning tasks and heterogeneous hardware setups is recommended before sweeping deployment.

Cover for SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Abstract

We introduce SeerAttention-R, a sparse attention framework specifically tailored for the long decoding of reasoning models. Extended from SeerAttention, SeerAttention-R retains the design of learning attention sparsity through a self-distilled gating mechanism, while removing query pooling to accommodate auto-regressive decoding. With a lightweight plug-in gating, SeerAttention-R is flexible and can be easily integrated into existing pretrained model without modifying the original parameters. We demonstrate that SeerAttention-R, trained on just 0.4B tokens, maintains near-lossless reasoning accuracy with 4K token budget in AIME benchmark under large sparse attention block sizes (64/128). Using TileLang, we develop a highly optimized sparse decoding kernel that achieves near-theoretical speedups of up to 9x over FlashAttention-3 on H100 GPU at 90% sparsity. Code is available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 SeerAttention-R
  • 2.1 A Recap of SeerAttention
  • 2.2 SeerAttention-R: AttnGate for Sparse Decoding
  • 2.3 Distillation/Training
  • 3 Inference of SeerAttention-R
  • 3.1 Sparsify Methods: Token Budget vs Threshold
  • 3.2 K Compression Cache
  • 3.3 Block Sparse Flash Decoding Kernel
  • 4 Experiments
  • 4.1 Experiments Setup
  • 4.2 Oracle Sparse Accuracy: How Sparse is Attention in Reasoning Models?
  • 4.3 Results of SeerAttention-R and Quest
  • 4.4 Kernel Speedup
  • 5 Ablation Studies
  • 5.1 Block Size for Sparse Attention
  • 5.2 Hybrid Dense Attention in the First Two Layers
  • 5.3 Threshold VS Token Budgets
  • 5.4 Impact of Sparse Attention on Generate Length
  • 5.5 Training Budget
  • 6 Limitation and Future Work
  • 6.1 End-to-end Speedup
  • 6.2 Adaptive Sparsity Ratio
  • 6.3 Unify Sparse Prefill and Decoding
  • 7 Related Works
  • 7.1 Training-free vs. Training-based Sparse Attention
  • 7.2 KV Cache Compression: A Sparse Attention Perspective
  • 7.3 Other Efficient Attention Algorithms
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — AttnGate Architecture for Sparse Autoregressive Decoding

    model/method

    SeerAttention-R adapts the learned attention gating mechanism (AttnGate) for autoregressive decoding in reasoning large language models without altering the original model parameters. Unlike prefill gating, the decoding AttnGate eliminates sequence-level pooling on the Query (QQ) tensor to support single-token autoregressive generation.

    To align with Grouped Query Attention (GQA), the AttnGate query branch uses a linear projection that maps each group of g=Hq/Hkvg = H_q / H_{kv} query heads into a single shared query head of dimension dgated_{gate} (where HqH_q is the total number of query heads, HkvH_{kv} is the number of key-value heads, and dd is the model's head dimension). The Key (KK) tensor is compressed along the sequence dimension using non-overlapping chunk pooling of kernel and stride size equal to block size bb. To capture outliers and preserve the distribution, sequence pooling concatenates Max, Min, and Average pooling outputs before applying a linear projection of input dimension 3d3d to dimension dgated_{gate}. Rotary Position Embeddings (RoPE) are reapplied to pre-RoPE QQ and KK tensors within the AttnGate, assigning each compressed key block the position index of its initial token.

  2. Knowl 2 — Mathematical Formulation of the Decoding Attention Gate

    equation

    Let Qnope∈RB×Hq×1×dQ_{nope} \in \mathbb{R}^{B \times H_q \times 1 \times d} and Knope∈RB×Hkv×L×dK_{nope} \in \mathbb{R}^{B \times H_{kv} \times L \times d} denote the un-rotated (pre-RoPE) query and key tensors at a decoding step, where BB is batch size, HqH_q is the query head count, HkvH_{kv} is the key-value head count, LL is sequence length, dd is head dimension, and g=Hq/Hkvg = H_q / H_{kv} is the GQA group size. For a block size bb, the block-level attention prediction score matrix S∈RB×Hkv×1×⌈L/b⌉S \in \mathbb{R}^{B \times H_{kv} \times 1 \times \lceil L/b \rceil} is calculated as:

    Qgate=RoPE(Wgateq reshape(Qnope,[…,g⋅d]))Q_{gate} = \text{RoPE}\left(W_{gate}^q \, \text{reshape}(Q_{nope}, [\dots, g \cdot d])\right)

    Kgate=RoPE(Wgatek concat[Pmax(Knope),Pmin(Knope),Pavg(Knope)])K_{gate} = \text{RoPE}\left(W_{gate}^k \, \text{concat}\left[P_{max}(K_{nope}), P_{min}(K_{nope}), P_{avg}(K_{nope})\right]\right)

    S=softmax(QgateKgate⊤dgate)S = \text{softmax}\left(\frac{Q_{gate} K_{gate}^\top}{\sqrt{d_{gate}}}\right)

    where Wgateq∈Rdgate×(g⋅d)W_{gate}^q \in \mathbb{R}^{d_{gate} \times (g \cdot d)} and Wgatek∈Rdgate×3dW_{gate}^k \in \mathbb{R}^{d_{gate} \times 3d} are learnable gating weights, dgated_{gate} is the AttnGate hidden dimension per head, and Pmax,Pmin,PavgP_{max}, P_{min}, P_{avg} represent sequence-dimension max-, min-, and average-pooling operations with kernel and stride size bb.

  3. Knowl 3 — Self-Distillation Objective and Ground Truth Generation for Decoding Sparsity

    model/method

    SeerAttention-R trains the lightweight AttnGate via self-distillation while keeping all original transformer weights frozen. During training on token sequences packed up to 32k tokens, target ground truth attention maps are generated by the original model.

    To construct the ground truth for autoregressive decoding without query sequence compression, the full attention score map is column-wise 1D max-pooled across keys in chunks of block size bb. To accommodate shared sparsity across GQA query groups, the column-pooled scores are further max-pooled across all query heads within each GQA group, producing a block-level ground truth corresponding to key-value heads. The resulting target vector is normalized to sum to 1. The AttnGate parameters are trained by minimizing the Kullback-Leibler (KL) divergence loss between the predicted AttnGate probability distribution SS and this normalized ground truth distribution.

  4. Knowl 4 — Fused Ground Truth Extraction in Attention Forward Kernel

    algorithm

    To prevent high GPU memory overhead from explicitly constructing full O(L2)O(L^2) attention matrices during distillation, SeerAttention-R integrates ground truth block-score extraction directly into the FlashAttention-2 forward kernel by reusing intermediate tile-level row maximums:

    Input: Query tensor QQ, Key tensor KK, Value tensor VV
    Output: Attention Output OO, Ground Truth Block Scores GTGT
    for ii from 1 to TrT_r
      Load qiq_i
      for jj from 1 to TcT_c
        Load kjk_j, vjv_j
        Compute sij=dot(qi,kj)s_{ij} = \text{dot}(q_i, k_j)
        rij=rowmax(sij)r_{ij} = \text{rowmax}(s_{ij})
        Store rijr_{ij}
        Update mij=max⁡(mi(j−1),rij)m_{ij} = \max(m_{i(j-1)}, r_{ij}), lijl_{ij}, and oijo_{ij}
      Compute final lil_i, mim_i, and OiO_i
      for jj from 1 to TcT_c
        Load rijr_{ij}
        Rescale gtij=exp⁡(rij−mi)/ligt_{ij} = \exp(r_{ij} - m_i) / l_i
        Store gtijgt_{ij}
    return OO, GTGT

    Here TrT_r and TcT_c denote the number of row (query) and column (key/value) tiles, rijr_{ij} stores the maximum pre-softmax attention score in block (i,j)(i,j), and mim_i and lil_i denote the row maximum and softmax normalizer.

  5. Knowl 5 — K Compression Cache and Sparse Decoding Inference Pipeline

    model/method

    During autoregressive inference, SeerAttention-R maintains a K Compression Cache containing the compressed key vectors (KgateK_{gate}) alongside the main KV cache to eliminate redundant computation for past tokens.

    For a block size of bb (e.g., b=64b=64), the K Compression Cache updates once every bb newly generated tokens by passing the latest bb keys through the pooling and linear layers. When the current sequence length is not an exact multiple of bb, the incomplete final block is unconditionally selected to prevent accuracy loss. For b=64b=64, the K Compression Cache consumes only 1/1281/128 (<1%) of the memory required by the original KV cache.

    To sparsify attention, AttnGate predictions are converted to discrete block selections via one of two strategies:

    1. Token budget: A fixed token budget BB is converted to a block budget k=⌈B/b⌉k = \lceil B/b \rceil, and the top-kk block indices are selected using a Top-k kernel on un-softmaxed logits.
    2. Threshold: Blocks whose predicted scores exceed a threshold τ\tau are activated, avoiding sorting and providing dynamic head-level sparsity.
  6. Knowl 6 — Load-Balanced Block-Sparse Flash Decoding Kernel

    model/method

    To execute block-sparse attention on GPU hardware efficiently, SeerAttention-R introduces a customized decoding kernel implemented in TileLang with a corresponding Triton implementation.

    1. Grid Scheduling: The kernel schedules threads over a 3D launch space of (batch, heads_kv, num_split).
    2. Workload Balancing: In the presence of variable active block counts across sequences and heads, the key/value dimension is partitioned along num_split using max_selected_blocks rather than the total sequence block count. The kernel loops exclusively through the active block indices provided by AttnGate (tensor shape [batch, heads_kv, max_selected_blocks]), skipping unselected memory blocks.
    3. Hardware Optimization: On NVIDIA H100 GPUs, query head groups are padded to 64 to leverage Tensor Core wgmma instructions. TileLang applies automated memory swizzling, warp specialization, rasterization, and software pipelining.
  7. Knowl 7 — Intrinsic Attention Sparsity in Reasoning Models Under Oracle Selection

    empirical result

    Evaluating Qwen3-14B on reasoning benchmarks (AIME24, AIME25, MATH-500, and GPQA-Diamond) using oracle block selection (ground truth sparse KV blocks extracted from full attention) confirms that reasoning models exhibit high intrinsic sparsity.

    Across block sizes of 32, 64, and 128 tokens, an oracle token budget of 2k tokens is sufficient to achieve completely lossless accuracy compared to the dense baseline on all four benchmarks. At an aggressive budget of 1k tokens, minor accuracy degradation occurs only on the most difficult benchmarks (AIME24 and AIME25) when paired with the largest block size of 128, while block sizes of 32 and 64 remain near-lossless.

  8. Knowl 8 — Reasoning Benchmark Performance of SeerAttention-R vs. Quest

    empirical result

    SeerAttention-R was evaluated against dense full attention and the training-free sparse baseline Quest across four open-source reasoning models (Qwen3-4B, Qwen3-8B, Qwen3-14B, and DeepSeek-R1-Distill-Qwen-14B) on AIME24, AIME25, MATH-500, and GPQA-Diamond, using a 32,768 token maximum generation limit, a block size of 64, and sparse attention across all layers.

    SeerAttention-R achieves near-lossless accuracy with a 4k token budget on AIME24/AIME25 and a 2k token budget on MATH-500/GPQA-Diamond across all models. In contrast, Quest suffers substantial accuracy drops at block size 64 and fails to reach dense performance even with an 8k token budget. Larger models (14B) show higher tolerance to sparsity than smaller models (4B, 8B), closing the gap to full attention more easily.

  9. Knowl 9 — TileLang Block-Sparse Flash Decoding Kernel Speedup on NVIDIA H100

    empirical result

    Microbenchmarks of the TileLang-based block-sparse flash decoding kernel on an NVIDIA H100 GPU (evaluating sequence lengths from 8k to 128k, batch sizes 1 to 16, 64 query heads, 8 KV heads, head dimension 128, and sparsity from 0.5 to 0.9) demonstrate significant latency reductions relative to FlashAttention-3 (FA3) and a Triton sparse baseline:

    1. At batch size 16 and sequence length ≥32k\ge 32\text{k} with 90% sparsity (0.9), the TileLang kernel achieves near-theoretical speedups of up to 8.6×8.6\times (and up to 9×9\times) compared to FA3, outperforming the Triton baseline by 1.7×1.7\times.
    2. For moderate memory footprints (e.g., batch size 4 and sequence length 32k32\text{k} at 90% sparsity), the kernel delivers up to a 6×6\times speedup over FA3.
  10. Knowl 10 — Sparsity Robustness Across Attention Block Sizes

    empirical result

    In an ablation on Qwen3-4B and Qwen3-8B on the AIME24 benchmark with a 4k token budget across block sizes 16, 32, 64, and 128:

    1. SeerAttention-R maintains steady accuracy across block sizes from 32 to 128 (retaining ~70% accuracy on Qwen3-4B and ~72% on Qwen3-8B), despite enforcing shared sparsity across GQA groups.
    2. Quest suffers steep accuracy degradation as block size increases (dropping from ~35% at block size 16 to ~0% at block size 128 on Qwen3-4B, and from ~60% to ~20% on Qwen3-8B).
    3. A block size of 16 was excluded from SeerAttention-R training and inference due to excessive intermediate attention map memory overhead causing out-of-memory errors.
  11. Knowl 11 — Impact of Sparse Attention Accuracy on Generated Reasoning Length

    data/table

    Inaccurate sparse attention selection causes reasoning models to generate significantly longer reasoning chains due to error accumulation. The table compares AIME24 pass@1 accuracy and average generated token sequence length for Qwen3-8B (where the full-attention baseline achieves 74.5% accuracy and 15.1k average generated tokens):

    Method / Metric 2k Budget 4k Budget 6k Budget 8k Budget
    Quest Accuracy (%) 13.3 44.2 52.5 59.6
    Quest Gen. Length (k tokens) 30.0 22.9 19.6 17.2
    SeerAttention-R Accuracy (%) 56.6 72.3 74.2 75.1
    SeerAttention-R Gen. Length (k tokens) 19.8 16.3 15.3 15.1

    When sparse attention is imprecise (e.g., Quest at 2k budget), generated length doubles to 30.0k tokens. In contrast, SeerAttention-R stabilizes to the dense baseline length (15.1k tokens) and matches dense accuracy (75.1%) at an 8k budget.

  12. Knowl 12 — AttnGate Distillation Computational Budget

    data/table

    Because only the AttnGate parameters are trained while original model weights are frozen, self-distillation in SeerAttention-R requires minimal compute. Models were trained for 800 steps with a global batch size of 16 on 32k packed sequences (0.4B total tokens) from OpenR1-MATH-220k on AMD MI300x GPUs using DeepSpeed Stage 2 and AdamW (learning rate 1e-3, cosine decay):

    Model GPU Hours (0.4B Tokens)
    Qwen3-4B 10.9
    Qwen3-8B 12.2
    Qwen3-14B 18.6

    Distilling gating modules for an 8B model requires only 12.2 GPU hours.

  13. Knowl 13 — Effect of Hybrid Dense Attention in Early Transformer Layers

    empirical result

    An ablation study on Qwen3-4B on the AIME24 benchmark (block size 64) comparing all-sparse attention layers against keeping the first two layers fully dense demonstrates divergent behavior between methods:

    1. Quest relies heavily on hybrid dense layers; keeping the first two layers dense improves Quest's accuracy by up to 70 percentage points at a 2k budget.
    2. SeerAttention-R achieves virtually identical accuracy whether using pure sparsity across all layers or hybrid dense attention in the first two layers (~60% at 2k budget, ~75% at 8k budget), showing that the learned AttnGate makes accurate sparse predictions throughout all network layers.
  14. Knowl 14 — System Limitations and Future Directions of SeerAttention-R

    limitation

    The authors identify three primary limitations of the current framework:

    1. End-to-end integration: The evaluation focuses on kernel-level execution; realizing end-to-end inference speedups requires integration with serving engines (e.g., vLLM, SGLang, Lserve), support for PagedAttention, and combining AttnGate with KV cache CPU offloading.
    2. Fixed sparsity allocation: SeerAttention-R uses static token budgets or fixed score thresholds rather than dynamically adjusting sparsity according to query difficulty or reasoning depth (e.g., via Nucleus/Top-pp sampling).
    3. Gating divergence between phases: SeerAttention (prefill) and SeerAttention-R (decoding) are trained separately with distinct AttnGate architectures due to differing query parallelism constraints.

Coverage note — No substantial contributed material was omitted; all core architectural equations, distillation algorithms, inference mechanisms, empirical accuracy results, GPU kernel speedup measurements, ablations, and stated limitations are included.

References

  1. 1.TileLang. URL https://github.com/tile-ai/tilelang.
  2. 2.Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant J Nair, Ilya Soloveychik, and Purushotham Kamath. Keyformer: Kv cache reduction through key tokens selection for efficient generative inference. Proceedings of Machine Learning and Systems, 6:114–127, 2024.
  3. 3.Joshua Ainslie, James Lee-Thorp, Michiel De Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023.
  4. 4.Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024.
  5. 5.Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai, Zhiding Yu, and Alexey Tumanov. Rocketkv: Accelerating long-context llm inference via two-stage kv cache compression. arXiv preprint arXiv:2502.14051, 2025.
  6. 6.Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
  7. 7.William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, and Jonathan Ragan-Kelley. Reducing transformer key-value cache size with cross-layer attention. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  8. 8.Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Anima Anandkumar, Abedelkadir Asi, and Junjie Hu. R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration. arXiv preprint arXiv:2505.24133, 2025.
  9. 9.Guoxuan Chen, Han Shi, Jiawei Li, Yihang Gao, Xiaozhe Ren, Yimeng Chen, Xin Jiang, Zhenguo Li, Weiyang Liu, and Chao Huang. Sepllm: Accelerate large language models by compressing one segment into one separator. arXiv preprint arXiv:2412.12094, 2024.
  10. 10.Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jingjia Luo, Di Liu, Huiqiang Jiang, Qi Chen, Jing Liu, Bailu Ding, Xiao Yan, Jiawei Jiang, Chen Chen, Mingxing Zhang, Yuqing Yang, Fan Yang, and Mao Yang. Retroinfer: A vector-storage approach for scalable long-context llm inference, 2025. URL https://arxiv.org/abs/2505.02922.
  11. 11.Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye, Yang Zhou, Jianyu Zhang, Niklas Nolte, Yuandong Tian, Matthijs Douze, Leon Bottou, Zhihao Jia, et al. Magicpig: Lsh sampling for efficient llm generation. arXiv preprint arXiv:2410.16179, 2024.
  12. 12.Yu Cheng, Lei Wang, Yining Shi, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang, Zhiwen Mo, Feiyang Chen, Fan Yang, Mao Yang, and Zhi Yang. PipeThreader: Software-defined pipelining for efficient dnn execution. In 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25), 2025. URL https://www.usenix.org/conference/osdi25/presentation/cheng.
  13. 13.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  14. 14.Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. 2023. URL https://arxiv.org/abs/2307.08691.
  15. 15.Tri Dao and Albert Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality. arXiv preprint arXiv:2405.21060, 2024.
  16. 16.Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, et al. Hymba: A hybrid-head architecture for small language models. arXiv preprint arXiv:2411.13676, 2024.
  17. 17.Hugging Face. Open r1: A fully open reproduction of deepseek-r1, January 2025. URL https://github.com/huggingface/open-r1.
  18. 18.Tianyu Fu, Haofeng Huang, Xuefei Ning, Genghan Zhang, Boju Chen, Tianqi Wu, Hongyi Wang, Zixiao Huang, Shiyao Li, Shengen Yan, et al. Moa: Mixture of sparse attention for automatic large language model compression. arXiv preprint arXiv:2406.14909, 2024.
  19. 19.Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, et al. Seerattention: Learning intrinsic sparse attention in your llms. arXiv preprint arXiv:2410.13276, 2024.
  20. 20.Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. Model tells you what to discard: Adaptive kv cache compression for llms. arXiv preprint arXiv:2310.01801, 2023.
  21. 21.Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve. Better & faster large language models via multi-token prediction. arXiv preprint arXiv:2404.19737, 2024.
  22. 22.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
  23. 23.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  24. 24.Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, and Sheng Guo. Omnikv: Dynamic context selection for efficient long-context llms. In The Thirteenth International Conference on Learning Representations, 2025.
  25. 25.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  26. 26.Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, June Paik, Michael W Mahoney, Kurt Keutzer, and Amir Gholami. Squeezed attention: Accelerating long context length llm inference. arXiv preprint arXiv:2411.09688, 2024.
  27. 27.Junhao Hu, Wenrui Huang, Weidong Wang, Zhenwen Li, Tiancheng Hu, Zhixia Liu, Xusheng Chen, Tao Xie, and Yizhou Shan. Efficient long-decoding inference with reasoning-aware attention sparsity. arXiv preprint arXiv:2502.11147, 2025.
  28. 28.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  29. 29.Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H Abdi, Dongsheng Li, Chin-Yew Lin, et al. Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention. arXiv preprint arXiv:2407.02490, 2024.
  30. 30.James M Joyce. Kullback-leibler divergence. In International encyclopedia of statistical science, pages 720–722. Springer, 2011.
  31. 31.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning, pages 5156–5165. PMLR, 2020.
  32. 32.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  33. 33.Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma, and Xun Zhou. Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference. arXiv preprint arXiv:2502.20766, 2025.
  34. 34.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR, 2023.
  35. 35.Aonian Li, Bangwei Gong, Bo Yang, Boji Shan, Chang Liu, Cheng Zhu, Chunhao Zhang, Congchao Guo, Da Chen, Dong Li, et al. Minimax-01: Scaling foundation models with lightning attention. arXiv preprint arXiv:2501.08313, 2025.
  36. 36.Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, et al. Scbench: A kv cache-centric analysis of long-context methods. arXiv preprint arXiv:2412.10319, 2024.
  37. 37.Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. Snapkv: Llm knows what you are looking for before generation. Advances in Neural Information Processing Systems, 37:22947–22970, 2024.
  38. 38.Chaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang, Tian Tang, Boyu Tian, Ion Stoica, Song Han, and Mingyu Gao. Twilight: Adaptive attention sparsity with hierarchical top-p pruning. arXiv preprint arXiv:2502.02770, 2025.
  39. 39.Zhixuan Lin, Johan Obando-Ceron, Xu Owen He, and Aaron Courville. Adaptive computation pruning for the forgetting transformer. arXiv preprint arXiv:2504.06949, 2025.
  40. 40.Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al. Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434, 2024.
  41. 41.Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024.
  42. 42.Di Liu, Meng Chen, Baotong Lu, Huiqiang Jiang, Zhenhua Han, Qianxi Zhang, Qi Chen, Chengruidong Zhang, Bailu Ding, Kai Zhang, et al. Retrievalattention: Accelerating long-context llm inference via vector retrieval. arXiv preprint arXiv:2409.10516, 2024.
  43. 43.Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng Yu, Chun Yuan, and Lu Hou. Quantization hurts reasoning? an empirical study on quantized reasoning models. arXiv preprint arXiv:2504.04823, 2025.
  44. 44.Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava. Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time. Advances in Neural Information Processing Systems, 36:52342–52364, 2023.
  45. 45.Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Yutao Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu. Moba: Mixture of block attention for long-context llms. arXiv preprint arXiv:2502.13189, 2025.
  46. 46.Pierre-Emmanuel Mazaré, Gergely Szilvasy, Maria Lomeli, Francisco Massa, Naila Murray, Hervé Jégou, and Matthijs Douze. Inference-time sparse attention with asymmetric indexing. arXiv preprint arXiv:2502.08246, 2025.
  47. 47.Pierre-Emmanuel Mazaré, Gergely Szilvasy, Maria Lomeli, Francisco Massa, Naila Murray, Hervé Jégou, and Matthijs Douze. Inference-time sparse attention with asymmetric indexing. arXiv preprint arXiv:2502.08246, 2025.
  48. 48.Art of Problem Solving. Aime problems and solutions. https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions.
  49. 49.Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048, 2023.
  50. 50.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024.
  51. 51.Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. Advances in Neural Information Processing Systems, 37:68658–68685, 2024.
  52. 52.Noam Shazeer. Fast transformer decoding: One write-head is all you need. arXiv preprint arXiv:1911.02150, 2019.
  53. 53.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  54. 54.Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models. arXiv preprint arXiv:2307.08621, 2023.
  55. 55.Yutao Sun, Li Dong, Yi Zhu, Shaohan Huang, Wenhui Wang, Shuming Ma, Quanlu Zhang, Jianyong Wang, and Furu Wei. You only cache once: Decoder-decoder architectures for language models. Advances in Neural Information Processing Systems, 37:7339–7361, 2024.
  56. 56.Yutao Sun, Tianzhu Ye, Dong Li, Yuqing Xia, Jian Chen, Yizhao Gao, Shijie Cao, Jianyong Wang, and Furu Wei. Rectified sparse attention. arXiv preprint arXiv:2506.04108, 2025.
  57. 57.Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. Quest: Query-aware sparsity for efficient long-context llm inference. arXiv preprint arXiv:2406.10774, 2024.
  58. 58.MiniCPM Team. Minicpm4: Ultra-efficient llms on end devices. 2025.
  59. 59.Philippe Tillet, Hsiang-Tsung Kung, and David Cox. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pages 10–19, 2019.
  60. 60.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  61. 61.Lei Wang, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi, Ningxin Zheng, Ziming Miao, Fan Yang, Ting Cao, Yuqing Yang, and Mao Yang. Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), pages 307–323, Santa Clara, CA, July 2024. USENIX Association. ISBN 978-1-939133-40-3. URL https://www.usenix.org/conference/osdi24/presentation/wang-lei.
  62. 62.Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, and Maosong Sun. Infllm: Training-free long-context extrapolation for llms with an efficient context memory. arXiv preprint arXiv:2402.04617, 2024.
  63. 63.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023.
  64. 64.Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. Duoattention: Efficient long-context llm inference with retrieval and streaming heads. arXiv preprint arXiv:2410.10819, 2024.
  65. 65.Ruyi Xu, Guangxuan Xiao, Haofeng Huang, Junxian Guo, and Song Han. Xattention: Block sparse attention with antidiagonal scoring. arXiv preprint arXiv:2503.16428, 2025.
  66. 66.An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  67. 67.Shang Yang, Junxian Guo, Haotian Tang, Qinghao Hu, Guangxuan Xiao, Jiaming Tang, Yujun Lin, Zhijian Liu, Yao Lu, and Song Han. Lserve: Efficient long-sequence llm serving with unified sparse attention. arXiv preprint arXiv:2502.14866, 2025.
  68. 68.Shuo Yang, Ying Sheng, Joseph E Gonzalez, Ion Stoica, and Lianmin Zheng. Post-training sparse attention with double sparsity. arXiv preprint arXiv:2408.07092, 2024.
  69. 69.Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. arXiv preprint arXiv:2312.06635, 2023.
  70. 70.Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. arXiv preprint arXiv:2406.06484, 2024.
  71. 71.Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, YX Wei, Lean Wang, Zhiping Xiao, et al. Native sparse attention: Hardware-aligned and natively trainable sparse attention. arXiv preprint arXiv:2502.11089, 2025.
  72. 72.Ted Zadouri, Hubert Strauss, and Tri Dao. Hardware-efficient attention for fast decoding. arXiv preprint arXiv:2505.21487, 2025.
  73. 73.Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. Big bird: Transformers for longer sequences. Advances in neural information processing systems, 33:17283–17297, 2020.
  74. 74.Zihao Zeng, Bokai Lin, Tianqi Hou, Hao Zhang, and Zhijie Deng. In-context kv-cache eviction for llms via attention-gate. arXiv preprint arXiv:2410.12876, 2024.
  75. 75.Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui. Pqcache: Product quantization-based kvcache for long context llm inference. arXiv preprint arXiv:2407.12820, 2024.
  76. 76.Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei, Haocheng Xi, Jun Zhu, and Jianfei Chen. Spargeattn: Accurate sparse attention accelerating any model inference. In International Conference on Machine Learning (ICML), 2025.
  77. 77.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Re, Clark Barrett, Zhangyang Wang, and Beidi Chen. H2o: Heavy-hitter oracle for efficient generative inference of large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=RkRrPp7GKO.
  78. 78.Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al. Sglang: Efficient execution of structured language model programs. Advances in Neural Information Processing Systems, 37:62557–62583, 2024.
  79. 79.Hongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke, Haoyu Li, Chen Zhang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Wei Cui, Fan Yang, Mao Yang, Lidong Zhou, Asaf Cidon, and Gennady Pekhimenko. ROLLER: Fast and efficient tensor compilation for deep learning. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pages 233–248, Carlsbad, CA, July 2022. USENIX Association. ISBN 978-1-939133-28-1. URL https://www.usenix.org/conference/osdi22/presentation/zhu.

Citation

MLA
Gao, Y., et al. “SeerAttention-R: Sparse Attention Adaptation for Long Reasoning”. arXiv, 2025, http://arxiv.org/abs/2506.08889v1.
APA
Gao, Y., Guo, S., Cao, S., Xia, Y., Cheng, Y., Wang, L., Ma, L., Sun, Y., Ye, T., Dong, L., So, H. K.-H., Hua, Y., Cao, T., Yang, F., & Yang, M. (2025). SeerAttention-R: Sparse Attention Adaptation for Long Reasoning. arXiv. http://arxiv.org/abs/2506.08889v1
Chicago
Gao, Y., S. Guo, S. Cao, et al. 2025. “SeerAttention-R: Sparse Attention Adaptation for Long Reasoning”. arXiv. http://arxiv.org/abs/2506.08889v1.
Harvard
Gao, Y. et al. (2025) “SeerAttention-R: Sparse Attention Adaptation for Long Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.08889v1.
Vancouver
1. Gao Y, Guo S, Cao S, et al (2025) SeerAttention-R: Sparse Attention Adaptation for Long Reasoning. arXiv

BibTeX

@article{gao2025seerattention,
  title = {SeerAttention-R: Sparse Attention Adaptation for Long Reasoning},
  author = {Gao, Yizhao and Guo, Shuming and Cao, Shijie and Xia, Yuqing and Cheng, Yu and Wang, Lei and Ma, Lingxiao and Sun, Yutao and Ye, Tianzhu and Dong, Li and So, Hayden Kwok-Hay and Hua, Yu and Cao, Ting and Yang, Fan and Yang, Mao},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.08889v1},
  eprint = {2506.08889}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission