ThinK: Thinner Key Cache by Query-Driven Pruning
Yuhui XuZhanming JieHanze DongLei WangXudong LuAojun ZhouAmrita SahaCaiming XiongDoyen Sahoo
Proposes ThinK, a query-driven KV cache channel pruning technique that reduces memory consumption for long-context large language models and supports up to five times larger inference batch sizes on a single GPU without loss of accuracy.
Deploying large language models for long-sequence tasks presents severe memory and computational bottlenecks during inference. The key-value cache—which stores intermediate attention representations—expands linearly with both batch size and sequence length, driving up infrastructure costs and limiting request throughput. While prior optimization efforts focus on dropping tokens along the sequence dimension or quantizing values, they overlook redundancy within the key cache channel dimensions.
The article demonstrates that key cache channels contain significant informational redundancy and introduces THINK, a query-driven channel pruning framework designed to cut key-value memory costs without sacrificing model accuracy.
To evaluate this approach, the researchers conducted extensive empirical experiments across standard benchmarks including LongBench and Needle-in-a-Haystack. They evaluated modern open-source models, specifically LLaMA-2-7B, LLaMA-3 (8B and 70B), and Mistral-7B, across tasks such as document question answering, summarization, and code generation. The method was tested both as a standalone technique and paired with existing compression methods, including token eviction (H2O, SnapKV) and quantization (KIVI).
The analysis yielded several key findings. First, key cache channel magnitudes are highly uneven, with singular value analysis revealing that top components capture over 90% of attention energy. Second, pruning up to 40% of key cache channels using the proposed query-driven criterion maintained or improved baseline accuracy while reducing overall key-value cache memory footprints by 20%. Third, when integrated with low-precision quantization, the method achieved a 2.8x peak memory reduction and increased achievable batch sizes from 4x to 5x on a single GPU compared to baseline configurations. Finally, generation speed improved, decreasing time-per-output-token from 0.27 to 0.25 milliseconds and raising throughput from 5,168 to 5,518 tokens per second under equivalent workloads.
These findings indicate that channel-level sparsity provides an effective path to lower operational expenses and alleviate hardware memory constraints. Because the approach functions as a plug-and-play optimization, enterprise teams can integrate it directly with existing eviction or quantization pipelines without retraining model weights.
Organizations serving long-context workloads should consider adopting query-driven channel pruning to improve hardware density and throughput. A moderate pruning ratio of approximately 40% is recommended, as it delivers consistent memory savings while preserving accuracy. Practitioners should implement the method first on key caches, while holding a small buffer of recent tokens unpruned to preserve response quality.
Regarding limitations, calculating channel importance adds a minor overhead to the initial time-to-first-token during the prefill phase. Additionally, extending channel pruning to value caches yields smaller efficiency gains because value channels exhibit more uniform energy distributions. Confidence in the reported key cache improvements is high across the evaluated architectures and standard benchmarks, though organizations should validate performance on specialized enterprise workloads prior to full-scale deployment.
- Paper: Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao et al. (2023). This paper establishes foundational insights into attention sink phenomena and token-level KV cache retention, motivating ThinK's focus on complementary channel-dimension redundancy.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). This work introduces PagedAttention and addresses KV cache memory bottlenecks during LLM serving, setting the systems context that ThinK aims to optimize along the channel dimension.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This foundational paper analyzes the KV cache memory bandwidth bottleneck during incremental decoding and introduces multi-query sharing of key-value representations.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). This paper demonstrates contextual and dynamic sparsity in transformer activations at inference time, providing conceptual precedent for query-dependent dynamic pruning.
- Paper: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, Zhihong Shao et al. (2024). This work introduces Multi-Head Latent Attention to compress key-value representations into low-rank latent vectors, directly relating to the low-rank structure ThinK exploits in key channels.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). This work provides essential background on activation outlier distribution and channel-dependent magnitude variations in large language models.
- Paper: A Probabilistic Interpretation of KV Cache Eviction, Renato Geh et al. (2026). This book provides a theoretical and probabilistic framework for KV cache reduction and eviction strategies, expanding on deterministic and heuristic cache compression techniques like ThinK.
- Paper: Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning, Heng Wang et al. (2026). This work critically examines complex KV cache eviction and scoring heuristics, offering a complementary perspective on memory compression during long reasoning traces.
- Paper: LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing, Wen Zan et al. (2026). This paper extends efficient long-context attention serving by co-designing hardware-aware cross-layer indexing and hierarchical sparse attention.
- Paper: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, DeepSeek-AI (2026). This monograph investigates extreme scaling of KV cache compression through composite sparse attention and ultra-low-bit formats for million-token context workloads.
- Paper: Prefix Sliding for efficient test-time scaling, Niklas Muennighoff et al. (2026). This text studies runtime context compression and sliding-window cache boundaries tailored specifically for long test-time reasoning traces.
