KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Zirui LiuJiayi YuanHongye JinShaochen (Henry) ZhongZhaozhuo XuVladimir BravermanBeidi ChenXia Hu
Introduces KIVI, a tuning-free asymmetric 2-bit key-value cache quantization method that quantizes keys per-channel and values per-token to reduce peak memory by 2.6x and increase large language model inference throughput up to 3.47x without sacrificing generation quality.
Deploying large language models for real-world applications is highly demanding and costly, requiring substantial hardware resources. To reduce operating costs, service providers batch multiple user requests together. However, during generation, the memory required to store previously computed conversational states—known as the key-value cache—grows rapidly with larger batch sizes and longer context lengths. This cache quickly becomes the primary bottleneck for both memory capacity and processing speed, causing hardware computational cores to sit idle while large volumes of cache data are loaded.
The article introduces and evaluates KIVI, a plug-and-play algorithm that compresses the key-value cache down to 2-bit representation without requiring any additional model fine-tuning or retraining. The primary objective is to demonstrate that asymmetric compression across different cache dimensions drastically reduces memory footprints and increases request processing throughput while preserving model output quality.
To establish these results, the authors analyzed how numerical values are distributed inside the memory caches of prominent open-source models, including Llama, Mistral, and Falcon. Based on these mathematical insights, they designed a system that compresses the key cache across feature channels and the value cache across individual tokens, while maintaining a small sliding window of recent tokens in full precision. The approach was tested across standard benchmarks covering short and long contexts, complex reasoning, information retrieval, and simulated production workloads on high-end hardware.
The investigation produced several critical findings. First, the asymmetric compression strategy reduces overall peak memory usage by approximately 2.6 times on tested models, such as Llama-2-7B. Second, this substantial memory reduction allows servers to handle up to 4 times larger batch sizes simultaneously, increasing total inference throughput by 2.35 to 3.47 times on real-world workloads. Third, the compression causes negligible accuracy loss—typically within 2% of uncompressed models—even on challenging reasoning and long-context benchmarks. Fourth, maintaining a small full-precision buffer for the most recent tokens proved critical for preserving accuracy on multi-step reasoning tasks.
These findings have direct operational and financial implications for organizations running language model infrastructure. By dramatically cutting cache memory demands, infrastructure operators can significantly lower hardware provisioning costs, reduce energy consumption, and serve substantially more user requests on existing servers without degrading user experience. The method departs from conventional uniform compression techniques by recognizing that different components of the attention mechanism exhibit distinct error sensitivities.
Organizations serving large language models should consider evaluating asymmetric cache quantization in their serving pipelines to improve throughput and cost efficiency. For standard architectures, adopting 2-bit quantization offers significant capacity gains, whereas architectures that already employ heavily compressed single-head attention mechanisms (such as Falcon) achieve more reliable stability using 4-bit settings. The authors also recommend fusing cache quantization steps directly into earlier computational operations in future software updates to further eliminate latency overheads.
Confidence in these findings is high for standard transformer architectures across standard context lengths up to 32,000 tokens. However, decision-makers should exercise caution when applying extreme 2-bit compression to models with highly condensed native attention designs without prior validation, as structural differences can heighten sensitivity to precision loss.
- Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention, Woosuk Kwon et al. (2023). This work establishes PagedAttention and memory-efficient KV cache management in serving systems, forming the direct baseline problem space KIVI targets for memory reduction.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This seminal paper introduces multi-query attention to mitigate the memory bandwidth bottlenecks of reloading key-value caches during autoregressive decoding.
- Paper: FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU, Ying Sheng et al. (2023). This work demonstrates high-throughput LLM inference with fine-grained 4-bit quantization for key-value caches, providing foundational context for KV cache compression.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). This work demonstrates how activation outlier distributions affect low-bit LLM quantization, motivating KIVI's per-channel and per-token element analysis.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). This foundational paper analyzes emergent outlier features across transformer channels and tokens, motivating asymmetric quantization strategies.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). This paper establishes post-training low-bit quantization for generative language models, which KIVI adapts specifically for dynamic KV cache states.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). This paper directly integrates query-driven channel pruning with KIVI's quantization scheme to further compress key cache memory.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). This work advances beyond standard per-channel/per-token KV cache quantization by applying polar coordinate transformations and compares directly against KIVI.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). This work extends low-bit KV cache compression to extreme 1-bit regimes using random projections and unbiased attention estimators.
- Paper: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate, Amir Zandieh et al. (2026). This research develops online vector quantization with near-optimal distortion bounds specifically tailored for real-time KV cache management.
- Paper: Evaluating Quantized Large Language Models, Shiyao Li et al. (2024). This work conducts an extensive empirical evaluation of post-training quantization tolerance across KV caches and long-context task domains.
