Evaluating Quantized Large Language Models

Shiyao LiXuefei NingLuning WangTengxuan LiuXiangsheng ShiShengen YanGuohao DaiHuazhong YangYu Wang

article2024ICML112 citations

Presents a comprehensive empirical evaluation of post-training quantization across 11 large language model families, revealing critical sensitivity patterns across weights, activations, and KV caches to guide optimal bit-width selection for diverse tasks.

Listen

Deploying large language models presents significant operational hurdles due to immense memory demands and computational overhead. Post-training quantization—a compression method that converts high-precision numerical values to lower bit-widths—offers a practical way to curb hardware requirements and operational costs. However, quantization is inherently lossy, and its practical trade-offs across different model architectures, tensor types, and diverse operational tasks have not been systematically understood.

The article provides a comprehensive evaluation of post-training quantization across model weights, activations, and key-value attention caches. It evaluates 11 model families ranging from 125 million to 180 billion parameters across five task domains: basic language processing, emergent capabilities, model trustworthiness, multi-turn dialogue, and long-context comprehension.

The authors conducted extensive empirical testing comparing standard uniform quantization and state-of-the-art recovery techniques against full-precision baselines across dozens of public benchmark datasets. Statistical properties, including tensor outliers and standard deviations, were analyzed to explain model behavior under reduced precision.

The evaluation yielded several critical findings regarding quantization tolerance. First, model size creates divergent sensitivities: larger models tolerate aggressive weight and key-value cache quantization better due to fewer outlier values, but they show substantially lower tolerance to activation quantization due to heavy-tailed outliers. Second, task difficulty dictates vulnerability; multi-step mathematical reasoning and self-calibration degrade sharply at lower bit-widths, primarily driven by logical reasoning errors rather than arithmetic mistakes, whereas basic language understanding remains robust. Third, multi-turn dialogue and long-context processing are exceptionally sensitive to precision loss; long-context tasks of 4,000 tokens or more experience severe performance drops under key-value cache compression below 8 bits, while dialogues collapse into repetitive sentences and random tokens below 4 bits. Fourth, architectural scaling methods like Mixture-of-Experts boost raw capability but fail to improve quantization tolerance compared to dense models of similar parameter scale. Finally, popular recovery techniques like AWQ and SmoothQuant fail to restore model viability under extreme low-bit regimes such as 2-bit weights or 4-bit activations.

These findings provide clear guidance for managing the trade-offs between inference speed, memory footprint, and model accuracy. Organizations can safely implement 4-bit weight, 8-bit activation, and 4-bit key-value cache quantization for routine language understanding and short-context workflows while keeping accuracy loss under 2 percent. Conversely, deploying lower precisions for complex reasoning, extended context windows, or small models under 13 billion parameters introduces significant operational risk and capability failure.

Decision-makers should establish task-specific quantization policies: maintain at least 8-bit precision across all tensors for models under 13 billion parameters handling reasoning tasks, preserve 8-bit key-value caches for contexts exceeding 4,000 tokens, and limit aggressive 4-bit compression to large, general-purpose models. Where hardware budgets are constrained, adopting a larger model compressed to 3-bit weights often outperforms a smaller model running at full precision.

Confidence in these findings is high for the evaluated architectures, though the specific bit-width thresholds may not generalize universally to unexamined model families. Because the article exclusively examines post-training quantization without fine-tuning, organizations pursuing extreme compression below 4 bits must conduct targeted pilot evaluations using quantization-aware training before deploying to production.

arXiv: 2402.18158
Cover for Evaluating Quantized Large Language Models

Abstract

Post-training quantization (PTQ) has emerged as a promising technique to reduce the cost of large language models (LLMs). Specifically, PTQ can effectively mitigate memory consumption and reduce computational overhead in LLMs. To meet the requirements of both high efficiency and performance across diverse scenarios, a comprehensive evaluation of quantized LLMs is essential to guide the selection of quantization methods. This paper presents a thorough evaluation of these factors by evaluating the effect of PTQ on Weight, Activation, and KV Cache on 11 model families, including OPT, LLaMA2, Falcon, Bloomz, Mistral, ChatGLM, Vicuna, LongChat, StableLM, Gemma, and Mamba, with parameters ranging from 125M to 180B. The evaluation encompasses five types of tasks: basic NLP, emergent ability, trustworthiness, dialogue, and long-context tasks. Moreover, we also evaluate the state-of-the-art (SOTA) quantization methods to demonstrate their applicability. Based on the extensive experiments, we systematically summarize the effect of quantization, provide recommendations to apply quantization techniques, and point out future directions. The code can be found in https://github.com/thu-nics/qllm-eval.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Quantization
  • 2.2. Benchmarks and Models
  • 2.3. Statistical Analysis
  • 3. Evaluation on Basic NLP Tasks
  • 3.1. Experimental Setups
  • 3.2. Effects of Quantization on Three Tensor Types
  • 3.3. Effects of Quantization on Different LLMs
  • 3.4. Effects of Quantization on Different Tasks
  • 4. Evaluation on Emergent Abilities
  • 4.1. Experimental Setups
  • 4.2. Experimental Results
  • 5. Evaluation on Trustworthiness
  • 5.1. Experimental Setups
  • 5.2. Experimental Results
  • 6. Evaluation on Dialogue Tasks
  • 6.1. Experimental Setups
  • 6.2. Experimental Results
  • 7. Evaluation on Long-Context Tasks
  • 7.1. Experimental Setups
  • 7.2. Experimental Results
  • 8. Limitations
  • Acknowledgement
  • Impact Statement
  • References
  • Appendix
  • A. Additional Preliminaries
  • A.1. Large Language Model Inference
  • A.2. Quantization
  • A.3. Experimental Setup Details
  • B. Additional Details of Evaluation on Basic NLP Abilities
  • B.1. Introduction of Datasets
  • B.2. Introduction of Metrics
  • B.3. Additional Results on Different Tensor Types
  • B.4. Additional Results on Different LLMs
  • B.5. Additional Results on Different Tasks
  • B.6. Additional Results on Different Quantization methods
  • B.7. Addition Results on Statistical Analysis
  • C. Additional Details of Evaluation on Emergent Abilities
  • C.1. Introduction of Datasets
  • C.2. Introduction of Metrics
  • C.3. Additional Results
  • D. Additional Details of Evaluation on Trustworthiness
  • D.1. Introduction of Datasets
  • D.2. Introduction of Metrics
  • D.3. Effects of Quantization on Ethics
  • D.4. Effects of Quantization on Hallucination
  • D.5. Effects of Quantization on Adversarial Robustness
  • E. Additional Details of Evaluation on Dialogue Abilities
  • E.1. Introduction of Datasets
  • E.2. Introduction of Metrics
  • E.3. Additional Case Study
  • F. Additional Details of Evaluation on Long-Context
  • F.1. Introduction of Datasets
  • F.2. Introduction of Metrics
  • F.3. Additional Results on LongEval dataset
  • F.4. Additional Results on Multi-document QA dataset

Knowls

  1. Knowl 1 — Task-Specific Bit-Width Recommendations for Post-Training Quantization of LLMs

    empirical result

    Across an evaluation of 11 large language model (LLM) families (including OPT, LLaMA2, Falcon, Bloomz, Mistral, Mixtral, ChatGLM, Vicuna, LongChat, StableLM, Gemma, and Mamba) across diverse tasks, post-training quantization (PTQ) bit-widths that preserve model performance within a 2% accuracy loss threshold vary substantially based on task type, model parameter scale, and context length:

    Task Category Model Size / Condition Recommended Bit-Widths (≤2%\le 2\% Loss)
    Basic NLP Tasks All model sizes W4, W4A8, KV4, W8KV4
    Emergent Abilities Model size <13B< 13\text{B} W8, W8A8, KV8
    Model size ≥13B\ge 13\text{B} W4, W4A8, KV4
    Trustworthiness Model size <7B< 7\text{B} W8, W8A8, KV8
    Model size ≥7B\ge 7\text{B} W4, W4A8, KV4
    Dialogue All model sizes W8, W8A8, KV4
    Long-Context Context length <4K< 4\text{K} tokens W4, W4A8, KV4
    Context length ≥4K\ge 4\text{K} tokens W4, W4A8, KV8

    Here, W, A, and KV represent Weight, Activation, and Key-Value Cache bit-widths, respectively (e.g., W4A8 denotes 4-bit weights and 8-bit activations; W8KV4 denotes 8-bit weights and 4-bit KV Cache).

  2. Knowl 2 — Opposing Model Scale Effects on Weight, KV Cache, and Activation Quantization Tolerance

    empirical result

    The tolerance of large language models to quantization scales in opposing directions depending on the tensor type being quantized:

    1. Weight and KV Cache Quantization: As model parameter scale increases within a model family, tolerance to Weight-only and KV Cache quantization increases. For example, 3-bit weight quantization (W3) degrades LLaMA2-7B accuracy severely on language understanding tasks, whereas W3 LLaMA2-70B exhibits negligible decline. This resilience is explained by tensor kurtosis: K=1n∑i=1n(xi−μσ)4K = \frac{1}{n} \sum_{i=1}^n \left( \frac{x_i - \mu}{\sigma} \right)^4 where xix_i are tensor elements, nn is the number of elements, μ\mu is the mean, and σ\sigma is the standard deviation. Weight kurtosis decreases as model size grows (e.g., OPT weight kurtosis decreases from 13.16 in OPT-1.3B to 8.74 in OPT-6.7B and 5.19 in OPT-66B; LLaMA2 weight kurtosis decreases from 4.93 in 7B to 4.83 in 70B), indicating fewer outlier weights in larger networks.

    2. Activation Quantization: Conversely, as model scale increases, tolerance to Activation quantization decreases. Activation kurtosis is orders of magnitude larger than that of weights (>1000>1000 versus ∼10\sim 10) and escalates dramatically with model size (OPT activation kurtosis rises from 544.97 in OPT-1.3B to 1562.67 in OPT-6.7B and 4945.32 in OPT-66B; LLaMA2 activation kurtosis rises from 1167.38 in 7B to 1279.15 in 70B). Consequently, larger models experience worse performance collapse under activation quantization (e.g., OPT-66B suffers severe degradation even at W8A8).

  3. Knowl 3 — Layer-Wise Outlier Asymmetry and Hybrid Static-Dynamic Activation Quantization

    empirical result

    Statistical profiling reveals extreme kurtosis variation across different linear projections in Transformer and State-Space architectures:

    • Feed-Forward Network (FFN) Down-Projection: Activation kurtosis in down-projection layers is disproportionately large compared to all other layers (1.54×1051.54 \times 10^5 for LLaMA2-7B, 3.84×1053.84 \times 10^5 for LLaMA2-13B, and 3.59×1053.59 \times 10^5 for LLaMA2-70B, compared to 1515–303303 in attention projections).
    • Attention Out-Projection (OO): Exhibits the highest weight kurtosis (e.g., 55.69 in OPT-1.3B and 33.80 in OPT-6.7B) but the lowest activation kurtosis among attention layers (e.g., 99.04 in OPT-1.3B and 62.23 in OPT-6.7B).
    • State Space Out-Projection: In Mamba-2.8B, activation kurtosis of the output projection reaches 26836.5426836.54, compared to 871.29871.29 in input projections (X_ProjX\text{\_Proj}).

    Because activation outliers are concentrated in FFN down-projection layers, fully static activation quantization (using offline calibration) causes large perplexity spikes, whereas dynamic activation quantization (computing scaling factors per-token at runtime) is computationally heavier. A hybrid quantization scheme—applying dynamic quantization exclusively to FFN down-projection layers and static quantization to all other layers—recovers nearly the entire performance of full dynamic quantization:

    Model FP16 PPL W4A8 Dynamic W4A8 Static W4A8 Static w.o. Down (Hybrid)
    LLaMA2-7B 11.71 12.51 40.78 12.63
    LLaMA2-13B 10.22 10.59 12.10 10.57
    LLaMA2-70B 6.87 7.23 16.89 7.26
  4. Knowl 4 — Quantization Sensitivity of Sparse Mixture-of-Experts Models

    empirical result

    Increasing total parameter count via Mixture-of-Experts (MoE) architectures does not impart the higher quantization tolerance observed when scaling dense models.

    While FP16 Mixtral-8x7B achieves task performance comparable to dense LLaMA2-70B, its sensitivity to Weight-only (W3, W2) and KV Cache (KV3, KV2) quantization is significantly higher than that of LLaMA2-70B or Falcon-40B/180B. Instead, the quantization degradation rate of Mixtral-8x7B closely tracks that of smaller dense models, namely Mistral-7B and LLaMA2-7B.

    Furthermore, retaining the MoE routing/gating linear layers in FP16 precision while quantizing all expert weights yields no measurable accuracy gain on language understanding benchmarks (such as LAMBADA), indicating that the quantization vulnerability stems from the compressed expert feed-forward weights rather than routing errors.

  5. Knowl 5 — Vulnerability Hierarchy of Emergent Abilities and Mathematical Reasoning Error Taxonomy

    empirical result

    Across evaluated emergent abilities, large language models display a clear sensitivity hierarchy under post-training quantization:

    1. High Sensitivity: Multi-Step Mathematical Reasoning (GSM8K) and Self-Calibration (evaluating model confidence on MMLU questions) exhibit the lowest quantization tolerance. Small models (<13B<13\text{B}) experience near-total collapse in self-calibration and mathematical reasoning at W3 or KV3 bit-widths.
    2. Moderate to High Robustness: In-Context Learning (MMLU, CEval), Instruction-Following (ARC, Hellaswag), and Commonsense Multi-Step Reasoning (StrategyQA) are substantially more robust, retaining high normalized performance down to W4, W4A8, and KV4.

    An error classification of failure cases on mathematical multi-step reasoning (GSM8K) for LLaMA2-70B demonstrates that loss of reasoning capability is dominated by broken problem-solving logic rather than arithmetic failure:

    • Incorrect Logic: Accounts for 49.4%49.4\% (under KV3), 55.8%55.8\% (under W3), and 56.9%56.9\% (under W4A8) of total errors.
    • Calculation Error: Accounts for 18.3%18.3\% (W3), 21.5%21.5\% (W4A8), and 23.7%23.7\% (KV3) of errors.
    • Condition Missing: Accounts for 13.8%13.8\% to 16.9%16.9\% of errors (omitting problem constraints).
    • Copy Mistake: Accounts for 7.7%7.7\% to 10.0%10.0\% of errors (transcription mistakes from previously generated intermediate equations).
  6. Knowl 6 — Long-Context Degradation and KV Cache Quantization Sensitivity

    empirical result

    Evaluation on long-context tasks (such as LongEval line retrieval up to 16K tokens and Multi-Document Question Answering up to 6K tokens) shows that sequence length directly amplifies quantization vulnerability:

    1. Length-Induced Sensitivity: While models quantized to W4, W4A8, or KV4 retain near-lossless accuracy on short contexts (<4K<4\text{K} tokens), performance degrades rapidly as text length extends to ≥4K\ge 4\text{K} tokens.
    2. KV Cache vs. Weight Sensitivity: For long texts, models are markedly more sensitive to KV Cache quantization than to Weight-only or Weight-Activation quantization at identical bit-widths. For instance, in LongChat (LLaMA-based), even KV8 quantization induces severe degradation beyond 6K tokens, whereas W8 remains intact.
    3. Family Variations: Vicuna-7B/13B and ChatGLM2/3-6B can maintain accuracy at KV8 across long contexts but suffer sharp drops at KV4 (e.g., on LongEval 16K, Vicuna-7B accuracy drops from 57.80% in FP16 to 37.00% under W8KV4). The Mistral family demonstrates higher tolerance to KV Cache quantization, maintaining stability at KV4 across 32K context lengths.
  7. Knowl 7 — Degradation Progression and Failure Modes in Quantized Multi-Turn Dialogue

    empirical result

    Evaluation of chat-tuned LLMs on multi-turn dialogue (MT-Bench, scored from 1 to 10 by GPT-4) reveals distinct structural failure modes across quantization precisions:

    1. Weight vs. KV Tolerance: Dialogue capability is more tolerant to KV Cache quantization than Weight quantization. Most models preserve GPT-4 dialogue scores within 2% loss at W8, W8A8, and KV4. Quantizing weights to W4 causes a notable degradation (>0.3>0.3 score drop on LLaMA2-13B and LLaMA2-70B).
    2. Progressive Failure Patterns:
      • 3-bit (W3, KV3): Models retain semantic intent but produce sentence-level repetition loops and occasional unescaped unicode tokens (e.g., \u2013, \u0101).
      • 2-bit (W2, KV2, W4A4): Models undergo catastrophic breakdown, producing continuous token-level repetitions or random token noise. Exceptions include ChatGLM3-6B, Falcon-40B, and Falcon-180B, which generate syntactically fluent sentences under KV2, though the outputs lack factual or substantive content.
    3. Multi-Turn Dynamics: The second turn of dialogue experiences steeper degradation under severe KV cache compression (e.g., on LLaMA2-13B at KV3, turn 1 drops by 0.31 points while turn 2 drops by 1.19 points relative to W8A8).
  8. Knowl 8 — Refusal Inversion in Ethical Judgment Under Weight vs. KV Cache Quantization

    empirical result

    On ethical judgment benchmarks (ETHICS: commonsense moral and virtue tasks), small models (<7B<7\text{B}, such as LLaMA2-7B and Mistral-7B) display opposing behavioral shifts depending on whether weights or KV caches are quantized:

    • Weight Quantization (W3): Produces a non-intuitive increase in benchmark accuracy. In FP16, safety alignment causes small models to frequently output canned refusal responses (e.g., "I apologize, but I cannot provide a straightforward answer..."), which are graded as incorrect. W3 quantization disrupts these refusal guardrails, prompting the model to answer the ethical queries directly. Attention map analysis shows that W3 quantization increases the attention weight assigned to the input question tokens.
    • KV Cache Quantization (KV3): Decreases benchmark accuracy by intensifying refusal and evasiveness. Attention map analysis reveals that KV3 quantization decreases attention directed at the prompt question tokens, leading to uninformative responses.
    • Scale Dependency: This refusal inversion occurs exclusively in small models (<7B<7\text{B}). For models ≥7B\ge 7\text{B}, accuracy degrades monotonically as bit-width decreases across all tensor types.
  9. Knowl 9 — Rank Preservation and Multiple-Choice Evaluation Metric Artifacts Under Quantization

    empirical result

    Evaluation across diverse model families demonstrates two key methodological properties of quantized LLM benchmarking:

    1. Rank Preservation: When bit-widths remain at or above W4, W4A8, and KV4, the relative performance ranking of LLMs strongly mirrors that of the full-precision FP16 models. The Spearman rank correlation coefficient: ρ=1−6∑i=1ndi2n(n2−1)\rho = 1 - \frac{6 \sum_{i=1}^n d_i^2}{n(n^2 - 1)} (where did_i is the rank difference for the ii-th model and nn is the number of evaluated models) exceeds 0.90.9 across benchmark datasets for W4, W4A8, and KV4.
    2. Multiple-Choice Perplexity Artifacts: Evaluating quantized LLMs via perplexity (PPL) on multiple-choice options can produce false accuracy improvements (e.g., Mistral-7B accuracy on RACE rising by 3.5% after W8A8 quantization). PPL analysis reveals that mean perplexity on correct answers consistently worsens (increases) after quantization. However, on questions where the FP16 model exhibited high prediction uncertainty (characterized by low standard deviation of PPL across candidate options), quantization perturbations act as random noise that can flip incorrect choices to correct ones by chance.
  10. Knowl 10 — Uniform Post-Training Quantization Scheme and Tensor Granularity

    model/method

    The uniform post-training quantization mapping a 16-bit floating-point value XFP16X_{\text{FP16}} to an NN-bit integer XINTX_{\text{INT}} is defined by: XINT=⌊XFP16−ZS⌉X_{\text{INT}} = \left\lfloor \frac{X_{\text{FP16}} - Z}{S} \right\rceil S=max⁡(XFP16)−min⁡(XFP16)2N−1S = \frac{\max(X_{\text{FP16}}) - \min(X_{\text{FP16}})}{2^N - 1} where SS is the positive scale factor, ZZ is the zero-point offset, and ⌊⋅⌉\lfloor \cdot \rceil denotes rounding to the nearest integer. In symmetric quantization, Z=0Z = 0; in asymmetric quantization, Z=min⁡(XFP16)Z = \min(X_{\text{FP16}}).

    The evaluation applies distinct granularity configurations tailored to inference hardware constraints:

    1. Weight-Only Quantization: Asymmetric uniform group-wise quantization. The weight tensor WW is partitioned along the input channel dimension into groups whose size equals the attention head dimension (e.g., 128 for LLaMA2, Mistral, Vicuna, LongChat, ChatGLM; 64 for Falcon; 64/80/96/128 for OPT and Bloomz; 80 for StableLM-3B).
    2. Weight-Activation Quantization: Asymmetric group-wise quantization for weights WW, paired with symmetric per-token quantization for activation tensors XX (Z=0Z = 0, single scale factor per token vector).
    3. KV Cache Quantization: Asymmetric group-wise quantization applied independently to the key (KK) and value (VV) cache tensors in each self-attention block.
  11. Knowl 11 — Inefficacy of Post-Training Quantization at 2-Bit Weight and W4A4 Precision

    limitation

    State-of-the-art post-training quantization (PTQ) algorithms, including Activation-aware Weight Quantization (AWQ) and SmoothQuant, fail to restore functional performance at ultra-low bit-widths:

    • 2-Bit Weights (W2): While AWQ provides noticeable accuracy recovery at 3-bit weights (e.g., improving LLaMA2-70B GSM8K accuracy from 50.87% under Round-to-Nearest to 56.41%), both Round-to-Nearest (RTN) and AWQ result in total model failure (0.00% accuracy on GSM8K and LAMBADA) under W2 quantization across LLaMA2 models.
    • 4-Bit Weight-Activation (W4A4): RTN quantization causes near-complete performance loss (<1%<1\% accuracy on GSM8K). While SmoothQuant partially mitigates activation clipping, accuracy remains severely depressed (e.g., LLaMA2-70B achieves only 5.53% on GSM8K and 38.11% on LAMBADA compared to FP16 baselines of 59.14% and 78.96%, respectively).

    PTQ is therefore insufficient for sub-3-bit weights or 4-bit activations, necessitating Quantization-Aware Training (QAT) or specialized non-uniform/window-based quantization schemes.

Coverage note — None. All primary empirical findings, statistical analyses, task-specific bit-width recommendations, architectural evaluations (dense vs. MoE, layer-wise kurtosis), error distributions, and methodological limitations contributed by the paper are fully covered.

References

  1. 1.Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G. Falcon-40B: an open large language model with state-of-the-art performance. 2023.
  2. 2.Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pp. 7432–7439, 2020.
  3. 3.Bondarenko, Y., Nagel, M., and Blankevoort, T. Quantizable transformers: Removing outliers by helping attention heads do nothing. arXiv preprint arXiv:2306.12929, 2023.
  4. 4.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  5. 5.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  6. 6.Contributors, O. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
  7. 7.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  8. 8.Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 320–335, 2022.
  9. 9.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pretrained transformers, 2023.
  10. 10.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  11. 11.Gemma Team, T. M., Hardin, C., Dadashi, R., Bhupatiraju, S., Sifre, L., Riviere, M., Kale, M. S., Love, J., Tafti, P., Hussenot, L., and et al. Gemma. 2024. doi: 10.34740/KAGGLE/M/3301. URL https://www.kaggle.com/m/3301.
  12. 12.Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361, 2021.
  13. 13.GitHub. https://github.com/features/copilot. 2023.
  14. 14.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces, 2023.
  15. 15.Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., and Steinhardt, J. Aligning ai with shared human values. Proceedings of the International Conference on Learning Representations (ICLR), 2021a.
  16. 16.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021b.
  17. 17.Huang, Y., Bai, Y., Zhu, Z., Zhang, J., Zhang, J., Su, T., Liu, J., Lv, C., Zhang, Y., Lei, J., Fu, Y., Sun, M., and He, J. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322, 2023.
  18. 18.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  19. 19.Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221, 2022.
  20. 20.Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K. Squeezellm: Dense-and-sparse quantization. arXiv preprint arXiv:2306.07629, 2023.
  21. 21.Krishnamoorthi, R. Quantizing deep convolutional networks for efficient inference: A whitepaper. arXiv preprint arXiv:1806.08342, 2018.
  22. 22.Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683, 2017.
  23. 23.Lee, C., Jin, J., Kim, T., Kim, H., and Park, E. Owq: Lessons learned from activation outliers for weight quantization in large language models. arXiv preprint arXiv:2306.02272, 2023.
  24. 24.Levesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  25. 25.Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J., Stoica, I., Ma, X., and Zhang, H. How long can context length of open-source llms truly promise? In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023.
  26. 26.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration, 2023.
  27. 27.Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods, 2021.
  28. 28.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172, 2023a.
  29. 29.Liu, P., Liu, Z., Gao, Z.-F., Gao, D., Zhao, W. X., Li, Y., Ding, B., and Wen, J.-R. Do emergent abilities exist in quantized large language models: An empirical study. arXiv preprint arXiv:2307.08072, 2023b.
  30. 30.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888, 2023c.
  31. 31.Liu, Z., Yuan, J., Jin, H., Zhong, S., Xu, Z., Braverman, V., Chen, B., and Hu, X. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750, 2024.
  32. 32.Mattern, J. and Hohr, K. Mamba-chat. GitHub, 2023. URL https://github.com/havenhq/mamba-chat.
  33. 33.Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., Van Baalen, M., and Blankevoort, T. A white paper on neural network quantization. arXiv preprint arXiv:2106.08295, 2021.
  34. 34.OpenAI. Gpt-4 technical report, 2023.
  35. 35.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  36. 36.Park, G., Park, B., Kim, M., Lee, S., Kim, J., Kwon, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models, 2023.
  37. 37.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  38. 38.Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019.
  39. 39.Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Fu, D. Y., Xie, Z., Chen, B., Barrett, C., Gonzalez, J. E., Liang, P., Re, C., Stoica, I., and Zhang, C. Flexgen: High-throughput generative inference of large language models with a single gpu, 2023.
  40. 40.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  41. 41.Tow, J., Bellagente, M., Mahan, D., and Riquelme, C. Stablelm 3b 4e1t. URL https://huggingface.co/stabilityai/stablelm-3b-4e1t.
  42. 42.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  43. 43.Wan, Z., Wang, X., Liu, C., Alam, S., Zheng, Y., et al. Efficient large language models: A survey. arXiv preprint arXiv:2312.03863, 1, 2023.
  44. 44.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. 2019. In the Proceedings of ICLR.
  45. 45.Wang, B., Xu, C., Wang, S., Gan, Z., Cheng, Y., Gao, J., Awadallah, A. H., and Li, B. Adversarial glue: A multitask benchmark for robustness evaluation of language models. ArXiv, abs/2111.02840, 2021.
  46. 46.Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2022a.
  47. 47.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022b.
  48. 48.Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X. Outlier suppression: Pushing the limit of low-bit transformer language models. Advances in Neural Information Processing Systems, 35:17402–17414, 2022c.
  49. 49.Workshop, B., Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagn e, R., Luccioni, A. S., Yvon,  F., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  50. 50.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023.
  51. 51.Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. Advances in Neural Information Processing Systems, 35: 27168–27183, 2022.
  52. 52.Yao, Z., Wu, X., Li, C., Youn, S., and He, Y. Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation. arXiv preprint arXiv:2303.08302, 2023.
  53. 53.Yuan, T., Ning, X., Zhou, D., Yang, Z., Li, S., Zhuang, M., Tan, Z., Yao, Z., Lin, D., Li, B., et al. Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k. arXiv preprint arXiv:2402.05136, 2024.
  54. 54.Yuan, Z., Niu, L., Liu, J., Liu, W., Wang, X., Shang, Y., Sun, G., Wu, Q., Wu, J., and Wu, B. Rptq: Reorder-based post-training quantization for large language models. arXiv preprint arXiv:2304.01089, 2023.
  55. 55.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  56. 56.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  57. 57.Zheng, C., Huang, M., and Sun, A. Chid: A large-scale chinese idiom dataset for cloze test. arXiv preprint arXiv:1906.01265, 2019.
  58. 58.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023a.
  59. 59.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023b.
  60. 60.Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., et al. A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294, 2024.

Citation

MLA
Li, S., et al. “Evaluating Quantized Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.18158v2.
APA
Li, S., Ning, X., Wang, L., Liu, T., Shi, X., Yan, S., Dai, G., Yang, H., & Wang, Y. (2024). Evaluating Quantized Large Language Models. arXiv. http://arxiv.org/abs/2402.18158v2
Chicago
Li, S., X. Ning, L. Wang, et al. 2024. “Evaluating Quantized Large Language Models”. arXiv. http://arxiv.org/abs/2402.18158v2.
Harvard
Li, S. et al. (2024) “Evaluating Quantized Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.18158v2.
Vancouver
1. Li S, Ning X, Wang L, Liu T, Shi X, Yan S, Dai G, Yang H, Wang Y (2024) Evaluating Quantized Large Language Models. arXiv

BibTeX

@article{li2024evaluating,
  title = {Evaluating Quantized Large Language Models},
  author = {Li, Shiyao and Ning, Xuefei and Wang, Luning and Liu, Tengxuan and Shi, Xiangsheng and Yan, Shengen and Dai, Guohao and Yang, Huazhong and Wang, Yu},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.18158v2},
  eprint = {2402.18158}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/