KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Zirui LiuJiayi YuanHongye JinShaochen (Henry) ZhongZhaozhuo XuVladimir BravermanBeidi ChenXia Hu

article2024ICML410 citations

Introduces KIVI, a tuning-free asymmetric 2-bit key-value cache quantization method that quantizes keys per-channel and values per-token to reduce peak memory by 2.6x and increase large language model inference throughput up to 3.47x without sacrificing generation quality.

Listen

Deploying large language models for real-world applications is highly demanding and costly, requiring substantial hardware resources. To reduce operating costs, service providers batch multiple user requests together. However, during generation, the memory required to store previously computed conversational states—known as the key-value cache—grows rapidly with larger batch sizes and longer context lengths. This cache quickly becomes the primary bottleneck for both memory capacity and processing speed, causing hardware computational cores to sit idle while large volumes of cache data are loaded.

The article introduces and evaluates KIVI, a plug-and-play algorithm that compresses the key-value cache down to 2-bit representation without requiring any additional model fine-tuning or retraining. The primary objective is to demonstrate that asymmetric compression across different cache dimensions drastically reduces memory footprints and increases request processing throughput while preserving model output quality.

To establish these results, the authors analyzed how numerical values are distributed inside the memory caches of prominent open-source models, including Llama, Mistral, and Falcon. Based on these mathematical insights, they designed a system that compresses the key cache across feature channels and the value cache across individual tokens, while maintaining a small sliding window of recent tokens in full precision. The approach was tested across standard benchmarks covering short and long contexts, complex reasoning, information retrieval, and simulated production workloads on high-end hardware.

The investigation produced several critical findings. First, the asymmetric compression strategy reduces overall peak memory usage by approximately 2.6 times on tested models, such as Llama-2-7B. Second, this substantial memory reduction allows servers to handle up to 4 times larger batch sizes simultaneously, increasing total inference throughput by 2.35 to 3.47 times on real-world workloads. Third, the compression causes negligible accuracy loss—typically within 2% of uncompressed models—even on challenging reasoning and long-context benchmarks. Fourth, maintaining a small full-precision buffer for the most recent tokens proved critical for preserving accuracy on multi-step reasoning tasks.

These findings have direct operational and financial implications for organizations running language model infrastructure. By dramatically cutting cache memory demands, infrastructure operators can significantly lower hardware provisioning costs, reduce energy consumption, and serve substantially more user requests on existing servers without degrading user experience. The method departs from conventional uniform compression techniques by recognizing that different components of the attention mechanism exhibit distinct error sensitivities.

Organizations serving large language models should consider evaluating asymmetric cache quantization in their serving pipelines to improve throughput and cost efficiency. For standard architectures, adopting 2-bit quantization offers significant capacity gains, whereas architectures that already employ heavily compressed single-head attention mechanisms (such as Falcon) achieve more reliable stability using 4-bit settings. The authors also recommend fusing cache quantization steps directly into earlier computational operations in future software updates to further eliminate latency overheads.

Confidence in these findings is high for standard transformer architectures across standard context lengths up to 32,000 tokens. However, decision-makers should exercise caution when applying extreme 2-bit compression to models with highly condensed native attention designs without prior validation, as structural differences can heighten sensitivity to precision loss.

arXiv: 2402.02750jy-yuan/KIVI
Cover for KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Abstract

Efficiently serving large language models (LLMs) requires batching of many requests to reduce the cost per request. Yet, with larger batch sizes and longer context lengths, the key-value (KV) cache, which stores attention keys and values to avoid re-computations, significantly increases memory demands and becomes the new bottleneck in speed and memory usage. Additionally, the loading of the KV cache causes the computational core to be idle, which limits the inference speed. A straightforward and effective solution to reduce KV cache size is quantization, which decreases the total bytes taken by KV cache. However, there is a lack of in-depth studies that explore the element distribution of KV cache to understand the hardness and limitation of KV cache quantization. To fill the gap, we conducted a comprehensive study on the element distribution in KV cache of popular LLMs. Our findings indicate that the key cache should be quantized per-channel, i.e., group elements along the channel dimension and quantize them together. In contrast, the value cache should be quantized per-token. From this analysis, we developed a tuning-free 2bit KV cache quantization algorithm named KIVI. With hardware-friendly implementation, KIVI can enable Llama, Falcon, and Mistral models to maintain almost the same quality while using 2.6×\mathbf{2.6\times} less peak memory (including model weight). This reduction in memory usage enables up to 4×\mathbf{4\times} larger batch size, bringing 2.35×∼3.47×\mathbf{2.35\times \sim 3.47\times} throughput on real LLM inference workload. The source code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background: Attention Inference-Time Workflow
  • 3 Methodology
  • 3.1 Preliminary Study of KV Cache Quantization
  • 3.2 Why Key and Value Cache Should Quantize Along Different Dimensions?
  • 3.3 KIVI: Algorithm and System Support
  • 4 Experiments
  • 4.1 Settings
  • 4.2 Accuracy and Efficiency Analysis
  • 4.2.1 Comparison Between Different Quantization Configurations
  • 4.2.2 Accuracy Comparison on Generation Tasks
  • 4.2.3 Ablation
  • 4.2.4 Efficiency Comparison
  • 5 Related Work
  • 6 Conclusion and Future Work
  • References
  • A Detailed Implementations
  • B NIAH Setting
  • C More Ablation Results
  • D More Experimental Results

Knowls

  1. Knowl 1 — Asymmetric Key-Value Cache Quantization Principle

    model/method

    In large language model (LLM) inference, the key cache and value cache exhibit fundamentally different numerical distributions and roles in the self-attention operation, demanding asymmetric quantization dimensions:

    1. Key Cache (XKX_K) via Per-Channel Quantization: Key activations have persistent, cross-token outliers concentrated in a small number of fixed channel dimensions. Applying group-wise quantization along the channel dimension (grouping elements along channels across tokens) isolates large-magnitude outliers to their specific channels. This prevents dynamic range stretching from distorting standard channels and avoids large reconstruction errors in attention logits (A=Softmax(tQXK⊤)A = \text{Softmax}(t_Q X_K^\top)).

    2. Value Cache (XVX_V) via Per-Token Quantization: Value activations exhibit no channel-specific outlier patterns. Because the attention output vector is computed as a weighted sum over value tokens (tO=AXV=∑j=1lAij[XV]j∗t_O = A X_V = \sum_{j=1}^{l} A_{ij} [X_V]_{j*}), and attention weights AA are highly sparse, the output is dominated by a tiny subset of key tokens. Quantizing the value cache along the token dimension (grouping elements within each token row) confines quantization error strictly within individual tokens. Quantization noise in non-essential tokens is thus prevented from bleeding into the value representations of important tokens.

  2. Knowl 2 — KIVI Quantization and Streaming Inference Algorithm

    algorithm

    The KIVI algorithm performs group-wise asymmetric quantization (BB-bit, e.g., 2-bit or 4-bit) for the key-value (KV) cache during both prefill and autoregressive decoding phases. It maintains an unquantized residual sliding window of length RR for the most recent tokens to preserve local precision and accommodate the streaming arrival of new tokens.

    procedure Prefill:
        Input: Prompt activation tensor X∈Rlprompt×dX \in \mathbb{R}^{l_{\text{prompt}} \times d}, Key weight WK∈Rd×dW_K \in \mathbb{R}^{d \times d}, Value weight WV∈Rd×dW_V \in \mathbb{R}^{d \times d}, Group size GG, Residual length RR
        Output: Unquantized prompt key XKX_K and value XVX_V, Quantized KV cache (Q(XKg),XKr,Q(XVg),XVr)(Q(X_{K_g}), X_{K_r}, Q(X_{V_g}), X_{V_r})
        XK←XWKX_K \leftarrow X W_K
        XV←XWVX_V \leftarrow X W_V
        XVg←XV[:lprompt−R]X_{V_g} \leftarrow X_V[: l_{\text{prompt}} - R]
        XVr←XV[lprompt−R:]X_{V_r} \leftarrow X_V[l_{\text{prompt}} - R :]
        Q(XVg)←GroupQuant(XVg,dim=token,numGroup=d//G)Q(X_{V_g}) \leftarrow \text{GroupQuant}(X_{V_g}, \text{dim}=\text{token}, \text{numGroup}=d // G)
        r←lprompt(modR)r \leftarrow l_{\text{prompt}} \pmod R
        XKg←XK[:lprompt−r]X_{K_g} \leftarrow X_K[: l_{\text{prompt}} - r]
        XKr←XK[lprompt−r:]X_{K_r} \leftarrow X_K[l_{\text{prompt}} - r :]
        Q(XKg)←GroupQuant(XKg,dim=channel,numGroup=lprompt//G)Q(X_{K_g}) \leftarrow \text{GroupQuant}(X_{K_g}, \text{dim}=\text{channel}, \text{numGroup}=l_{\text{prompt}} // G)
        return XK,XVX_K, X_V, cache (Q(XKg),XKr,Q(XVg),XVr)(Q(X_{K_g}), X_{K_r}, Q(X_{V_g}), X_{V_r})
    procedure Decoding:
        Input: New token embedding t∈R1×dt \in \mathbb{R}^{1 \times d}, Query weight WQW_Q, Key weight WKW_K, Value weight WVW_V, Cache (Q(XKg),XKr,Q(XVg),XVr)(Q(X_{K_g}), X_{K_r}, Q(X_{V_g}), X_{V_r}), Group size GG, Residual length RR
        Output: Attention output tO∈R1×dt_O \in \mathbb{R}^{1 \times d}, Updated cache
        tQ←tWQ,  tK←tWK,  tV←tWVt_Q \leftarrow t W_Q, \; t_K \leftarrow t W_K, \; t_V \leftarrow t W_V
        XKr←Concat([XKr,tK],dim=token)X_{K_r} \leftarrow \text{Concat}([X_{K_r}, t_K], \text{dim}=\text{token})
        XVr←Concat([XVr,tV],dim=token)X_{V_r} \leftarrow \text{Concat}([X_{V_r}, t_V], \text{dim}=\text{token})
        if len(XKr)=R\text{len}(X_{K_r}) = R then
            Q(XKr)←GroupQuant(XKr,dim=channel,numGroup=R//G)Q(X_{K_r}) \leftarrow \text{GroupQuant}(X_{K_r}, \text{dim}=\text{channel}, \text{numGroup}=R // G)
            Q(XKg)←Concat([Q(XKg),Q(XKr)],dim=token)Q(X_{K_g}) \leftarrow \text{Concat}([Q(X_{K_g}), Q(X_{K_r})], \text{dim}=\text{token})
            XKr←empty tensorX_{K_r} \leftarrow \text{empty tensor}
        end if
        if len(XVr)>R\text{len}(X_{V_r}) > R then
            Q(XVr′)←GroupQuant(XVr[:−R],dim=token,numGroup=d//G)Q(X'_{V_r}) \leftarrow \text{GroupQuant}(X_{V_r}[: -R], \text{dim}=\text{token}, \text{numGroup}=d // G)
            Q(XVg)←Concat([Q(XVg),Q(XVr′)],dim=token)Q(X_{V_g}) \leftarrow \text{Concat}([Q(X_{V_g}), Q(X'_{V_r})], \text{dim}=\text{token})
            XVr←XVr[−R:]X_{V_r} \leftarrow X_{V_r}[-R :]
        end if
        Ag←tQQ(XKg)⊤A_g \leftarrow t_Q Q(X_{K_g})^\top
        Ar←tQXKr⊤A_r \leftarrow t_Q X_{K_r}^\top
        A←Concat([Ag,Ar],dim=token)A \leftarrow \text{Concat}([A_g, A_r], \text{dim}=\text{token})
        S←Softmax(A)S \leftarrow \text{Softmax}(A)
        Sg←S[:−len(XKr)],  Sr←S[−len(XKr):]S_g \leftarrow S[: -\text{len}(X_{K_r})], \; S_r \leftarrow S[-\text{len}(X_{K_r}) :]
        tO←SgQ(XVg)+SrXVrt_O \leftarrow S_g Q(X_{V_g}) + S_r X_{V_r}
        return tOt_O, cache (Q(XKg),XKr,Q(XVg),XVr)(Q(X_{K_g}), X_{K_r}, Q(X_{V_g}), X_{V_r})

    The operation tQQ(XKg)⊤t_Q Q(X_{K_g})^\top is computed via a fused mixed-precision matrix multiplication CUDA kernel that dequantizes cached keys on chip at the tiling level.

  3. Knowl 3 — Group-Wise Affine Quantization Formulation

    equation

    For a tensor slice XX belonging to a quantization group (of size G=32G=32), uniform asymmetric round-to-nearest integer quantization to BB bits (where B∈{2,4}B \in \{2, 4\}) and its corresponding dequantization X′X' are defined by:

    Q(X)=⌊X−zXsX⌉,X′=Q(X)⋅sX+zXQ(X) = \left\lfloor \frac{X - z_X}{s_X} \right\rceil, \quad X' = Q(X) \cdot s_X + z_X

    where ⌊⋅⌉\lfloor \cdot \rceil denotes the round-to-nearest integer operator, zX∈Rz_X \in \mathbb{R} is the zero-point, and sX∈Rs_X \in \mathbb{R} is the scaling factor:

    zX=min⁡(X),sX=max⁡(X)−min⁡(X)2B−1z_X = \min(X), \quad s_X = \frac{\max(X) - \min(X)}{2^B - 1}

    For the key cache XK∈Rl×dX_K \in \mathbb{R}^{l \times d}, grouping is performed per-channel along the token sequence dimension with l//Gl // G groups of size GG per channel. For the value cache XV∈Rl×dX_V \in \mathbb{R}^{l \times d}, grouping is performed per-token along the hidden channel dimension with d//Gd // G groups of size GG per token.

  4. Knowl 4 — Reconstruction and Attention Error Under Different Quantization Orientations

    empirical result

    Empirical evaluation on Llama-2-13B demonstrates the mathematical divergence between per-token and per-channel quantization for Key and Value caches:

    1. Key Cache: Quantizing key cache per-token produces an average relative reconstruction error ∥XK−XK′∥F∥XK∥F=13.67\frac{\|X_K - X'_K\|_F}{\|X_K\|_F} = 13.67 and an attention matrix relative error ∥A−A′∥F∥A∥F=47.00\frac{\|A - A'\|_F}{\|A\|_F} = 47.00. By contrast, per-channel key quantization reduces the reconstruction error to 4.554.55 and attention error to 9.609.60 (a ∼5×\sim 5\times reduction in attention error) under an average attention sparsity of 84.3%84.3\%.

    2. Value Cache: While the relative reconstruction error ∥XV−XV′∥F∥XV∥F\frac{\|X_V - X'_V\|_F}{\|X_V\|_F} is comparable between per-token (4.574.57) and per-channel (3.733.73), the relative attention output error Δ=∥AXV−AXV′∥F∥AXV∥F\Delta = \frac{\|A X_V - A X'_V\|_F}{\|A X_V\|_F} is 3.553.55 for per-token quantization versus 49.8949.89 for per-channel quantization. Quantizing value cache per-channel causes an approximately 14×14\times to 15×15\times higher error on the attention output vector due to the destruction of sparsely weighted token values.

  5. Knowl 5 — Performance Comparison Across Quantization Strategies on LM-Eval Benchmarks

    data/table

    The performance of 16-bit baseline, simulated (fake) group-wise quantized KV cache configurations (group size 32, zero-padded across all tokens), and KIVI (group size 32, residual length 128) across CoQA (Exact Match), TruthfulQA (BLEU), and GSM8K (Exact Match) is reported below. In the table, 'K - C' / 'V - C' denote per-channel quantization and 'K - T' / 'V - T' denote per-token quantization.

    Model Method CoQA TruthfulQA GSM8K
    Llama-2-7B 16bit 63.88 30.76 13.50
    4bit (K - T, V - T) 64.82 29.85 12.28
    2bit (K - C, V - T) [Simulated] 59.08 33.10 5.76
    2bit (K - T, V - T) [Simulated] 39.88 18.29 0.83
    2bit (K - C, V - C) [Simulated] 3.60 0.27 0.00
    2bit (K - T, V - C) [Simulated] 1.30 0.49 0.08
    KIVI-4 63.78 30.80 13.80
    KIVI-2 63.05 33.95 12.74
    Llama-2-13B 16bit 66.37 29.53 22.67
    4bit (K - T, V - T) 66.73 29.14 20.92
    2bit (K - C, V - T) [Simulated] 63.53 28.60 12.21
    2bit (K - T, V - T) [Simulated] 52.93 24.98 4.55
    2bit (K - C, V - C) [Simulated] 2.88 0.74 0.00
    2bit (K - T, V - C) [Simulated] 2.80 0.26 0.08
    KIVI-4 66.38 29.49 23.65
    KIVI-2 66.23 29.84 20.77
    Falcon-7B 16bit 59.83 23.20 4.55
    4bit (K - T, V - T) 58.53 22.94 3.26
    2bit (K - C, V - T) [Simulated] 43.93 20.82 1.29
    2bit (K - T, V - T) [Simulated] 25.72 0.91 0.53
    2bit (K - C, V - C) [Simulated] 41.95 17.11 1.52
    2bit (K - T, V - C) [Simulated] 19.53 0.94 0.15
    KIVI-4 59.67 22.58 4.47
    KIVI-2 57.48 24.98 3.41
    Mistral-7B 16bit 67.40 30.45 38.36
    4bit (K - T, V - T) 67.80 29.83 36.85
    2bit (K - C, V - T) [Simulated] 61.65 29.64 26.46
    2bit (K - T, V - T) [Simulated] 54.55 25.86 5.00
    2bit (K - C, V - C) [Simulated] 24.40 24.86 2.27
    2bit (K - T, V - C) [Simulated] 10.73 19.12 0.99
    KIVI-4 66.95 30.49 37.30
    KIVI-2 66.35 32.17 36.01

    The table demonstrates three key findings: (1) Symmetric per-channel value quantization (V-C) leads to catastrophic accuracy failure across all architectures; (2) Asymmetric quantization (K-C, V-T) outperforms all other 2-bit simulated baselines; (3) KIVI's inclusion of a full-precision residual sliding window (R=128R=128) recovers the performance degradation observed in simulated (K-C, V-T) on reasoning tasks like GSM8K, bringing KIVI-2 to within ∼1–2%\sim 1\text{--}2\% of the 16-bit baseline across Llama and Mistral models.

  6. Knowl 6 — Long-Context Benchmark Results on LongBench

    data/table

    KIVI performance across long-context tasks in LongBench evaluated on Llama-2 (7B, 13B, 7B-Chat, 13B-Chat), Falcon-7B, and Mistral-7B using group size G=32G=32 and residual window length R=128R=128. Maximum context length was set to 8192 for Mistral-7B and 4096 for other architectures.

    Model Precision Qasper QMSum MultiNews TREC TriviaQA SAMSum LCC RepoBench-P Average
    Llama2-7B 16bit 9.52 21.28 3.51 66.00 87.72 41.69 66.66 59.82 44.52
    KIVI-4 9.28 21.42 3.88 66.00 87.72 41.82 66.80 59.83 44.59
    KIVI-2 9.31 20.50 1.14 66.00 87.42 42.71 66.88 60.23 44.27
    Llama2-13B 16bit 9.32 21.38 3.71 70.00 87.87 43.55 66.61 56.42 44.85
    KIVI-4 9.16 20.86 3.21 69.00 86.97 44.26 65.30 57.08 44.48
    KIVI-2 8.58 20.69 6.19 69.50 87.78 44.30 65.08 55.46 44.69
    Llama2-7B-Chat 16bit 19.65 20.54 26.36 63.00 84.28 41.12 59.75 52.93 45.95
    KIVI-4 19.62 20.70 25.49 63.00 84.13 40.87 59.27 53.56 45.83
    KIVI-2 19.32 20.46 25.48 63.00 84.84 40.60 58.71 52.97 45.67
    Llama2-13B-Chat 16bit 24.18 20.37 25.69 67.50 86.90 42.18 50.23 50.64 45.96
    KIVI-4 23.00 20.36 26.06 67.50 87.20 42.04 52.55 52.77 46.44
    KIVI-2 23.59 20.76 25.25 67.50 87.17 41.56 49.93 48.45 45.52
    Falcon-7B 16bit 1.48 2.35 11.09 13.00 5.84 2.44 23.86 9.69 8.71
    KIVI-4 1.04 2.41 11.98 13.00 5.84 2.36 23.72 9.92 8.78
    KIVI-2 1.98 3.61 6.78 10.00 6.24 2.73 22.18 10.12 7.95
    Mistral-7B 16bit 8.12 19.98 19.99 67.50 89.80 41.69 66.59 58.99 46.58
    KIVI-4 7.89 20.06 20.58 67.50 89.80 41.56 66.45 58.62 46.56
    KIVI-2 6.92 19.71 17.92 66.50 89.63 41.66 65.52 58.99 45.85

    The data shows that KIVI-2 maintains nearly identical average performance to unquantized 16-bit models across long-context QA, summarization, few-shot learning, and code completion benchmarks (e.g., 44.2744.27 vs 44.5244.52 on Llama2-7B, 44.6944.69 vs 44.8544.85 on Llama2-13B, and 45.8545.85 vs 46.5846.58 on Mistral-7B).

  7. Knowl 7 — Serving Peak Memory Reduction and Throughput Speedup

    empirical result

    In real LLM serving workloads synthesized from ShareGPT (average prompt length lprompt=161l_{\text{prompt}} = 161, average generation length lgen=338l_{\text{gen}} = 338) on a single NVIDIA A100 GPU (80GB) hosting Llama-2-7B:

    1. Memory Compression: KIVI-2 achieves a 2.6×2.6\times reduction in overall peak GPU memory consumption (including model weights and KV cache).
    2. Batch Size Scaling: Under identical GPU memory ceilings, KIVI enables serving batch sizes up to 4×4\times larger than FP16 baselines before encountering out-of-memory errors.
    3. Throughput Gain: By unlocking larger batch sizes and reducing GPU memory bandwidth bottlenecks during cache loading, KIVI yields an end-to-end token generation throughput improvement of 2.35×∼3.47×2.35\times \sim 3.47\times.
  8. Knowl 8 — Ablation of Group Size and Residual Length

    data/table

    The effect of quantization group size GG and residual sliding window length RR on multi-step reasoning accuracy was ablated on GSM8K using Llama-2-13B:

    Group Size GG (R=128R=128) GSM8K Residual Length RR (G=32G=32) GSM8K
    32 20.77 32 20.62
    64 21.00 64 19.86
    128 17.29 96 20.55
    – – 128 20.77

    Performance remains stable when scaling group size from G=32G=32 to G=64G=64, but drops by 3.48%3.48\% at G=128G=128 due to excessive dynamic range within large channel/token groups. Varying residual length RR between 3232 and 128128 retains comparable reasoning accuracy ({20.62,19.86,20.55,20.77}\{20.62, 19.86, 20.55, 20.77\}), showing that a moderate full-precision window (R≥32R \ge 32) is sufficient to safeguard reasoning capability while optimizing memory.

  9. Knowl 9 — Needle-in-a-Haystack Retrieval Robustness

    empirical result

    In the Needle-in-a-Haystack (NIAH) retrieval test evaluated on Llama-3-8B-Instruct (up to 27K tokens / 20K words) and Mistral-7B-Instruct-v0.2 (up to 30K tokens / 20K words) using a 7-digit passkey embedded in Paul Graham essays:

    • Both KIVI-2 and KIVI-4 maintain near 100%100\% passkey retrieval accuracy (score of 1.0) across all document depth levels (0.00.0 to 1.01.0) and document lengths (0.5K0.5\text{K} to 20K20\text{K} words).
    • Performance under 2-bit KV cache quantization matches the unquantized FP16 baseline across the entire context window, indicating that extreme 2-bit asymmetric quantization preserves deep retrieval capabilities without attention degradation.

Coverage note — Omitted supplementary benchmark tables for specific instruction-tuned variants (Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.2, LongChat-7B-v1.5 in Tables 8-10, and Table 7 residual ablations) as they reproduce the identical qualitative trends and conclusions captured in the primary LongBench and LM-Eval knowls.

References

  1. 1.ShareGPT Team. https://sharegpt.com/, 2023.
  2. 2.Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023.
  3. 3.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508, 2023.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  5. 5.Yu-Neng Chuang, Songchen Li, Jiayi Yuan, Guanchu Wang, Kwei-Herng Lai, Leisheng Yu, Sirui Ding, Chia-Yuan Chang, Qiaoyu Tan, Daochen Zha, et al. Understanding different design choices in training large time series models. arXiv preprint arXiv:2406.14045, 2024.
  6. 6.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  7. 7.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  8. 8.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al. A framework for few-shot language model evaluation. Version v0. 0.1. Sept, page 8, 2021.
  9. 9.Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. Kvquant: Towards 10 million context length llm inference with kv cache quantization. arXiv preprint arXiv:2401.18079, 2024.
  10. 10.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  11. 11.Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei. S3: Increasing gpu utilization during generative inference for higher throughput. arXiv preprint arXiv:2306.06000, 2023.
  12. 12.Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023.
  13. 13.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611–626, 2023.
  14. 14.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  15. 15.Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, and Anshumali Shrivastava. Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time. Advances in Neural Information Processing Systems, 36, 2024.
  16. 16.Amirkeivan Mohtashami and Martin Jaggi. Landmark attention: Random-access infinite context length for transformers. arXiv preprint arXiv:2305.16300, 2023.
  17. 17.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116, 2023.
  18. 18.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. Efficiently scaling transformer inference. Proceedings of Machine Learning and Systems, 5, 2023.
  19. 19.Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  20. 20.Noam Shazeer. Fast transformer decoding: One write-head is all you need. arXiv preprint arXiv:1911.02150, 2019.
  21. 21.Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. Flexgen: High-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning, pages 31094–31116. PMLR, 2023.
  22. 22.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022.
  23. 23.Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du. Scan and snap: Understanding training dynamics and token composition in 1-layer transformer. arXiv preprint arXiv:2305.16380, 2023.
  24. 24.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  25. 25.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  26. 26.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pages 38087–38099. PMLR, 2023a.
  27. 27.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023b.
  28. 28.Zhaozhuo Xu, Zirui Liu, Beidi Chen, Yuxin Tang, Jue Wang, Kaixiong Zhou, Xia Hu, and Anshumali Shrivastava. Compress, then prompt: Improving accuracy-efficiency trade-off of llm inference with transferable prompt. arXiv preprint arXiv:2305.11186, 2023.
  29. 29.Jiayi Yuan, Ruixiang Tang, Xiaoqian Jiang, and Xia Hu. Large language models for healthcare data augmentation: An example on patient-trial matching. In AMIA Annual Symposium Proceedings, volume 2023, page 1324. American Medical Informatics Association, 2023.
  30. 30.Jiayi Yuan, Hongyi Liu, Shaochen Zhong, Yu-Neng Chuang, Songchen Li, Guanchu Wang, Duy Le, Hongye Jin, Vipin Chaudhary, Zhaozhuo Xu, Zirui Liu, and Xia Hu. Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches. arXiv preprint arXiv:2407.01527, 2024.
  31. 31.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al. H _2 o: Heavy-hitter oracle for efficient generative inference of large language models. arXiv preprint arXiv:2306.14048, 2023.
  32. 32.Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. Atom: Low-bit quantization for efficient and accurate llm serving. Proceedings of Machine Learning and Systems, 6:196–209, 2024.

Citation

MLA
Zirui Liu, et al. “KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization”. Unpublished, 2023, https://doi.org/10.13140/RG.2.2.28167.37282.
APA
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Braverman, V., Beidi Chen, & Hu, X. (2023). KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization. Unpublished. https://doi.org/10.13140/RG.2.2.28167.37282
Chicago
Zirui Liu, Jiayi Yuan, Hongye Jin, et al. 2023. “KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization”. Unpublished, ahead of print. https://doi.org/10.13140/RG.2.2.28167.37282.
Harvard
Zirui Liu et al. (2023) “KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization”, Unpublished [Preprint]. Available at: https://doi.org/10.13140/RG.2.2.28167.37282.
Vancouver
1. Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Braverman V, Beidi Chen, Hu X (2023) KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization. Unpublished. https://doi.org/10.13140/RG.2.2.28167.37282

BibTeX

@article{https://doi.org/10.13140/rg.2.2.28167.37282,
  doi = {10.13140/RG.2.2.28167.37282},
  url = {https://www.researchgate.net/doi/10.13140/RG.2.2.28167.37282},
  author = {{Zirui Liu} and {Jiayi Yuan} and {Hongye Jin} and {Shaochen Zhong} and {Zhaozhuo Xu} and Braverman, Vladimir and {Beidi Chen} and Hu, Xia},
  language = {en},
  title = {KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization},
  publisher = {Unpublished},
  year = {2023}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/