ThinK: Thinner Key Cache by Query-Driven Pruning

Yuhui XuZhanming JieHanze DongLei WangXudong LuAojun ZhouAmrita SahaCaiming XiongDoyen Sahoo

article2025ICLR77 citations

Proposes ThinK, a query-driven KV cache channel pruning technique that reduces memory consumption for long-context large language models and supports up to five times larger inference batch sizes on a single GPU without loss of accuracy.

Listen

Deploying large language models for long-sequence tasks presents severe memory and computational bottlenecks during inference. The key-value cache—which stores intermediate attention representations—expands linearly with both batch size and sequence length, driving up infrastructure costs and limiting request throughput. While prior optimization efforts focus on dropping tokens along the sequence dimension or quantizing values, they overlook redundancy within the key cache channel dimensions.

The article demonstrates that key cache channels contain significant informational redundancy and introduces THINK, a query-driven channel pruning framework designed to cut key-value memory costs without sacrificing model accuracy.

To evaluate this approach, the researchers conducted extensive empirical experiments across standard benchmarks including LongBench and Needle-in-a-Haystack. They evaluated modern open-source models, specifically LLaMA-2-7B, LLaMA-3 (8B and 70B), and Mistral-7B, across tasks such as document question answering, summarization, and code generation. The method was tested both as a standalone technique and paired with existing compression methods, including token eviction (H2O, SnapKV) and quantization (KIVI).

The analysis yielded several key findings. First, key cache channel magnitudes are highly uneven, with singular value analysis revealing that top components capture over 90% of attention energy. Second, pruning up to 40% of key cache channels using the proposed query-driven criterion maintained or improved baseline accuracy while reducing overall key-value cache memory footprints by 20%. Third, when integrated with low-precision quantization, the method achieved a 2.8x peak memory reduction and increased achievable batch sizes from 4x to 5x on a single GPU compared to baseline configurations. Finally, generation speed improved, decreasing time-per-output-token from 0.27 to 0.25 milliseconds and raising throughput from 5,168 to 5,518 tokens per second under equivalent workloads.

These findings indicate that channel-level sparsity provides an effective path to lower operational expenses and alleviate hardware memory constraints. Because the approach functions as a plug-and-play optimization, enterprise teams can integrate it directly with existing eviction or quantization pipelines without retraining model weights.

Organizations serving long-context workloads should consider adopting query-driven channel pruning to improve hardware density and throughput. A moderate pruning ratio of approximately 40% is recommended, as it delivers consistent memory savings while preserving accuracy. Practitioners should implement the method first on key caches, while holding a small buffer of recent tokens unpruned to preserve response quality.

Regarding limitations, calculating channel importance adds a minor overhead to the initial time-to-first-token during the prefill phase. Additionally, extending channel pruning to value caches yields smaller efficiency gains because value channels exhibit more uniform energy distributions. Confidence in the reported key cache improvements is high across the evaluated architectures and standard benchmarks, though organizations should validate performance on specialized enterprise workloads prior to full-scale deployment.

Cover for ThinK: Thinner Key Cache by Query-Driven Pruning

Abstract

Large Language Models (LLMs) have revolutionized the field of natural language processing, achieving unprecedented performance across a variety of applications. However, their increased computational and memory demands present significant challenges, especially when handling long sequences. This paper focuses on the long-context scenario, addressing the inefficiencies in KV cache memory consumption during inference. Unlike existing approaches that optimize the memory based on the sequence length, we identify substantial redundancy in the channel dimension of the KV cache, as indicated by an uneven magnitude distribution and a low-rank structure in the attention weights. In response, we propose ThinK, a novel query-dependent KV cache pruning method designed to minimize attention weight loss while selectively pruning the least significant channels. Our approach not only maintains or enhances model accuracy but also achieves a reduction in KV cache memory costs by over 20% compared with vanilla KV cache eviction and quantization methods. For instance, ThinK integrated with KIVI can achieve a 2.8x reduction in peak memory usage while maintaining nearly the same quality, enabling up to a 5x increase in batch size when using a single GPU. Extensive evaluations on the LLaMA and Mistral models across various long-sequence datasets verified the efficiency of ThinK, establishing a new baseline algorithm for efficient LLM deployment without compromising performance. Our code has been made available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Observations
  • 3 ThinK
  • 3.1 Preliminary Study of KV Cache Optimization
  • 3.2 Query-Driven Pruning
  • 3.3 Implementation of ThinK
  • 4 Experiment Results
  • 4.1 Settings
  • 4.2 Results on LongBench
  • 4.3 Results on Needle-in-a-Haystack
  • 4.4 Ablation Studies
  • 5 Related Work
  • 6 Conclusion and Limitations
  • References
  • A Observations
  • B Needle-in-a-Haystack test performance comparison
  • C Implementations
  • C.1 Implementation with quantization
  • D Value Cache Pruning
  • E Pruning Key Cache on Vanilla Models
  • F Comparisons of Generation Speed and Throughput
  • G Comparisons with SVD based Methods

Citation

MLA
Xu, Y., et al. “ThinK: Thinner Key Cache by Query-Driven Pruning”. arXiv, 2024, http://arxiv.org/abs/2407.21018v3.
APA
Xu, Y., Jie, Z., Dong, H., Wang, L., Lu, X., Zhou, A., Saha, A., Xiong, C., & Sahoo, D. (2024). ThinK: Thinner Key Cache by Query-Driven Pruning. arXiv. http://arxiv.org/abs/2407.21018v3
Chicago
Xu, Y., Z. Jie, H. Dong, et al. 2024. “ThinK: Thinner Key Cache by Query-Driven Pruning”. arXiv. http://arxiv.org/abs/2407.21018v3.
Harvard
Xu, Y. et al. (2024) “ThinK: Thinner Key Cache by Query-Driven Pruning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.21018v3.
Vancouver
1. Xu Y, Jie Z, Dong H, Wang L, Lu X, Zhou A, Saha A, Xiong C, Sahoo D (2024) ThinK: Thinner Key Cache by Query-Driven Pruning. arXiv

BibTeX

@article{xu2024think,
  title = {ThinK: Thinner Key Cache by Query-Driven Pruning},
  author = {Xu, Yuhui and Jie, Zhanming and Dong, Hanze and Wang, Lei and Lu, Xudong and Zhou, Aojun and Saha, Amrita and Xiong, Caiming and Sahoo, Doyen},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.21018v3},
  eprint = {2407.21018}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors