ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
Lu YeZe TaoYong HuangYang Li
Proposes ChunkAttention, a self-attention module that uses a prefix-tree structure and a two-phase partition kernel to share key/value caches across requests with matching system prompts, accelerating attention computation by up to 4.8× in multi-tenant LLM serving.
Serving large language models in multi-tenant environments requires substantial memory and computational resources. The self-attention mechanism, a core computational component of these models, suffers from high latency and memory bottlenecks during real-time token generation because it repeatedly processes large key-value (KV) data caches. At the same time, many user requests share identical system instructions, examples, or tool definitions at the beginning of their prompts. Storing duplicate KV caches for these shared prefixes wastes significant hardware memory and limits overall serving capacity.
The article evaluates ChunkAttention, a self-attention module designed to automatically identify shared prompt prefixes across requests and share their KV cache representations in memory during inference. The authors develop and assess this system to reduce memory consumption and accelerate decoding speeds without requiring manual prompt pre-configuration.
The approach introduces a prefix-tree structure that divides monolithic KV caches into smaller chunks and links shared prefix paths dynamically across active requests. On top of this tree structure, the authors implement a two-phase partition self-attention kernel that batches shared computations across sequences before processing sequence-specific tokens. The authors benchmarked this system against industry-standard baselines, including PagedAttention and FlashAttention, using microkernel tests and end-to-end evaluations on a 7-billion-parameter language model hosted on NVIDIA A100 GPUs under varying batch sizes and arrival rates.
The evaluation yields several key findings. First, ChunkAttention accelerates the self-attention microkernel by 3.2× to 4.8× compared to standard PagedAttention when shared system prompts range from 1,024 to 4,096 tokens. Second, end-to-end serving evaluations demonstrate a 70% to 90% reduction in peak KV cache memory usage when handling long shared prefixes. Third, the system achieves 1.6× to 2.3× higher request throughput while maintaining an average generation latency under 40 milliseconds per token. Finally, when no prompt tokens are shared, ChunkAttention exhibits no performance regression compared to existing baseline systems.
These findings indicate that dynamic prefix-aware caching significantly lowers hardware serving costs and enhances system capacity for applications with repetitive context, such as customer service chatbots, tool-augmented models, and batch benchmark evaluations. The automated runtime detection eliminates operational overhead for infrastructure teams, as developers do not need to manually pre-register static system prompts.
Organizations operating large language model infrastructure should consider adopting dynamic prefix-aware KV caching to improve throughput and reduce memory footprints in shared workloads. Engineering teams should keep shared system instructions at the very beginning of prompts to maximize cache sharing benefits. Further community development is recommended to generalize these low-level GPU kernel optimizations across different model architectures and diverse hardware accelerators.
The reported performance gains depend heavily on shared text appearing at the exact beginning of prompt sequences; any modifications or variations in the initial tokens eliminate prefix matching benefits. Additionally, because the current implementation is heavily tuned for specific GPU architectures and standard attention dimensions, additional testing and optimization are necessary before deploying the solution across non-standard hardware environments.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). PagedAttention introduces the foundational non-contiguous KV-cache management framework and serving architecture that ChunkAttention directly builds upon and benchmarks against.
- Paper: FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, Tri Dao (2023). FlashAttention-2 provides the core hardware-aware IO-tiling and attention work-partitioning techniques foundational to developing ChunkAttention's two-phase GPU kernel.
- Paper: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills, Amey Agrawal et al. (2023). SARATHI introduces chunked prefill strategies for LLM inference that ChunkAttention builds upon for prefix-aware chunking and computation partitioning.
- Paper: Orca: A Distributed Serving System for Transformer-Based Generative Models, Gyeong-In Yu et al. (2022). Orca establishes iteration-level scheduling for generative transformer serving, providing the architectural foundation for modern multi-tenant batch serving systems.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). ThinK explores complementary memory reduction by pruning redundancy in the KV cache channel dimension rather than sharing across prompt sequence prefixes.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). PolarQuant investigates coordinate-transformation-based KV cache quantization to achieve orthogonal memory compression alongside prefix-sharing serving systems.
- Paper: Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference, Dhruv Deshmukh et al. (2025). Kascade develops cross-layer sparse attention to cut attention computation during long-context prefill and decode stages.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). QJL presents extreme 1-bit JL-transform quantization for KV caches to tackle inference memory bottlenecks from a numeric precision perspective.
- Paper: Language Models Can Control Their Own Attention, Namgyu Ho et al. (2026). Declarative Attention enables models to dynamically scope and mask their own KV cache reads at decode time within serving engines like vLLM.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). XAttention extends long-context inference acceleration through block-sparse attention scoring along antidiagonal patterns.
- Paper: Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention, Jingyang Yuan et al. (2025). Native Sparse Attention introduces hardware-aligned, trainable sparse attention kernels to natively alleviate the quadratic scaling of standard KV caches.
