SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Yizhao GaoShuming GuoShijie CaoYuqing XiaYu ChengLei WangLingxiao MaYutao SunTianzhu YeLi Dong
Introduces SeerAttention-R, a lightweight plug-in sparse attention framework for auto-regressive decoding that preserves reasoning accuracy on benchmarks like AIME while achieving up to a 9x inference speedup over FlashAttention-3 on H100 GPUs.
Modern artificial intelligence models solve complex problems by generating extended step-by-step reasoning sequences. While these longer responses substantially enhance accuracy, they introduce serious computational bottlenecks. As the generated text grows, the model must scan an expanding memory history at each step, causing memory demand and computing costs to grow dramatically. Sparse attention can mitigate these expenses by focusing only on essential tokens, but existing sparse techniques struggle to maintain precision during prolonged generation steps.
To address this challenge, the article introduces SeerAttention-R, a sparse attention framework designed to accelerate the step-by-step decoding phase in reasoning models. The objective of the article is to demonstrate that a lightweight, learned gating module can accurately select critical memory blocks during inference without retraining the underlying base model, thereby preserving reasoning accuracy while significantly speeding up computation.
To evaluate this approach, the authors trained a lightweight attention gating module on a dataset of 400 million tokens while freezing the original model parameters. They evaluated four open-source reasoning models—including Qwen3 variants ranging from 4 billion to 14 billion parameters and DeepSeek-R1-Distill-Qwen-14B—across challenging mathematical and general reasoning benchmarks such as AIME24, AIME25, MATH-500, and GPQA-Diamond. The team also implemented a specialized hardware execution kernel using TileLang and Triton to benchmark processing speedups on NVIDIA H100 GPUs against standard full attention baselines.
The findings show that attention in reasoning models is naturally sparse; focusing on only 2,000 to 4,000 important tokens is sufficient to match the accuracy of standard full-attention models that evaluate the entire context. In head-to-head comparisons, SeerAttention-R consistently outperformed leading training-free sparse methods, maintaining near-lossless accuracy even under large memory block sizes of 64 and 128 tokens. Furthermore, larger models demonstrated greater resilience to sparsity, closing performance gaps more easily than smaller models. At the hardware execution level, the customized decoding kernel achieved near-theoretical speedups of up to 9 times over standard FlashAttention-3 at 90 percent sparsity on long sequences.
These results demonstrate that organizations can dramatically reduce the compute time and operational costs of serving long-reasoning artificial intelligence models without sacrificing task performance. The lightweight nature of the training process—requiring only around 12 GPU hours for an 8-billion-parameter model—allows existing pretrained models to adopt sparse decoding with minimal post-training investment. Moreover, the framework avoids accuracy penalties that typically force competing sparse methods to generate excessively long, error-prone reasoning chains.
Decision-makers should consider piloting SeerAttention-R as an add-on optimization for large-scale reasoning deployments, particularly where long-sequence generation drives high inference costs. Next development steps should focus on integrating this sparse kernel into production serving systems, evaluating automatic threshold tuning across varying task complexities, and unifying prefill and decoding sparsity into a single end-to-end framework.
While the findings demonstrate high confidence in kernel-level speedups and reasoning accuracy across standard benchmarks, full end-to-end system throughput in live production environments remains to be confirmed. Continued validation across broader reasoning tasks and heterogeneous hardware setups is recommended before sweeping deployment.
- Paper: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision, Jay Shah et al. (2024). FlashAttention-3 establishes the state-of-the-art GPU kernel baseline against which SeerAttention-R measures its decoding speedups and GPU throughput.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). FlashAttention provides the foundational IO-aware tiling and block-sparse GPU execution concepts that SeerAttention-R's optimized TileLang sparse kernels build upon.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). Deja Vu introduces dynamic contextual sparsity and lightweight prediction mechanisms for transformer inference that underpin learned sparse attention gating.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). Native Sparse Attention details hardware-aligned block selection and gating architectures for long contexts, providing valuable context for SeerAttention-R's design.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). XAttention investigates block-sparse attention patterns and importance scoring mechanisms that SeerAttention-R optimizes for autoregressive reasoning decoding.
- Paper: Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference, Dhruv Deshmukh et al. (2025). Kascade develops practical GPU-tiled sparse attention for decoding and prefill phases, directly preceding SeerAttention-R's focus on long-context decoding efficiency.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). PagedAttention introduces the foundational memory management and KV-cache paging models required to deploy long-context attention kernels efficiently.
- Paper: LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing, Wen Zan et al. (2026). LongCat Sparse Attention extends sparse attention acceleration to ultra-long contexts up to one million tokens using streaming-aware and hierarchical indexing.
- Paper: Prefix Sliding for efficient test-time scaling, Niklas Muennighoff et al. (2026). Prefix Sliding explores an alternative test-time context scaling strategy by retaining prompt prefixes and sliding local windows during long reasoning rollouts.
- Paper: Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning, Heng Wang et al. (2026). Random Attention investigates eviction and sparsity baselines across extended reasoning traces, providing a complementary perspective on reasoning KV-cache management.
- Paper: Language Models Can Control Their Own Attention, Namgyu Ho et al. (2026). Declarative Attention builds on the need for efficient KV cache access during decoding by allowing models to dynamically declare their attention spans via text tokens.
