SqueezeLLM: Dense-and-Sparse Quantization

Sehoon KimColeman HooperAmir GholamiZhen DongXiuyu LiSheng ShenMichael W. MahoneyKurt Keutzer

article2024ICML382 citations

Presents SqueezeLLM, a post-training quantization framework combining sensitivity-based non-uniform quantization and dense-and-sparse matrix decomposition to achieve near-lossless 3-bit weight compression and up to 2.3× faster inference on large language models.

Listen

Deploying generative large language models for real-world tasks is heavily constrained by high computing resource requirements and memory limitations. During single-batch generative inference, performance is primarily constrained by memory bandwidth—the rate at which data moves from memory to processing cores—rather than raw arithmetic throughput. While quantizing weights to lower bit precisions reduces memory footprint, conventional low-bit compression methods frequently lead to steep drops in output quality and accuracy.

The article introduces and evaluates SqueezeLLM, a post-training quantization framework designed to enable lossless weight compression down to 3-bit precision while maintaining high inference speed and model accuracy without requiring full retraining.

The authors conducted empirical benchmarks, hardware profiling, and ablation studies across multiple model families, including LLaMA, LLaMA-2, OPT, and Vicuna, spanning sizes from 1.3 billion to 70 billion parameters. The method relies on two core mechanisms: sensitivity-based non-uniform quantization, which uses second-order loss sensitivity (approximated via Fisher information) to assign quantization levels near the most critical weights, and a dense-and-sparse decomposition, which isolates a tiny fraction of sensitive and outlier values into a full-precision sparse format while compressing the remaining dense matrix into low bitwidths.

The evaluation yielded several key findings. First, SqueezeLLM achieves significant compression with minimal quality loss; for example, on 3-bit LLaMA models, it reduces the perplexity gap to the 16-bit baseline by up to 2.1 times compared to previous state-of-the-art techniques. Second, retaining merely 0.45% of parameters in full-precision sparse format substantially improves quantization resolution, outperforming traditional grouping methods. Third, hardware deployment on A6000 and A100 GPUs demonstrated up to a 2.4-fold inference speedup and an approximate 4-fold reduction in peak memory usage relative to the uncompressed 16-bit baseline. Finally, the framework preserves performance on downstream tasks, maintaining high zero-shot and few-shot accuracy on standard knowledge benchmarks and instruction-following evaluations.

These findings indicate that generative model deployment costs and hardware barriers can be substantially reduced without compromising model capability. By addressing memory bandwidth bottlenecks rather than arithmetic operations, organizations can run larger models on fewer or smaller hardware instances, mitigating operational expense and avoiding complex multi-chip serving setups.

Decision-makers seeking to optimize large language model serving pipelines should consider adopting dense-and-sparse non-uniform quantization formats for single-batch and low-latency deployments. Where maximum accuracy is needed at ultra-low bitwidths, configuring a small sparsity allocation (such as 0.45%) provides the best balance between memory overhead and model quality compared to standard grouping approaches.

Confidence in these findings is supported by extensive empirical testing across varied architectures and actual hardware profiling. However, the study's primary scope focuses on generative decoder-only architectures and single-batch inference regimes. Readers should exercise caution when extrapolating these throughput gains to large-batch processing, where compute throughput becomes the dominant constraint, or to encoder-based network architectures.

Cover for SqueezeLLM: Dense-and-Sparse Quantization

Abstract

Generative Large Language Models (LLMs) have demonstrated remarkable results for a wide range of tasks. However, deploying these models for inference has been a significant challenge due to their unprecedented resource requirements. This has forced existing deployment frameworks to use multi-GPU inference pipelines, which are often complex and costly, or to use smaller and less performant models. In this work, we demonstrate that the main bottleneck for generative inference with LLMs is memory bandwidth, rather than compute, specifically for single batch inference. While quantization has emerged as a promising solution by representing weights with reduced precision, previous efforts have often resulted in notable performance degradation. To address this, we introduce SqueezeLLM, a post-training quantization framework that not only enables lossless compression to ultra-low precisions of up to 3-bit, but also achieves higher quantization performance under the same memory constraint. Our framework incorporates two novel ideas: (i) sensitivity-based non-uniform quantization, which searches for the optimal bit precision assignment based on second-order information; and (ii) the Dense-and-Sparse decomposition that stores outliers and sensitive weight values in an efficient sparse format. When applied to the LLaMA models, our 3-bit quantization significantly reduces the perplexity gap from the FP16 baseline by up to 2.1× as compared to the state-of-the-art methods with the same memory requirement. Furthermore, when deployed on an A6000 GPU, our quantized models achieve up to 2.3× speedup compared to the baseline. Our code is available at https://github.com/SqueezeAILab/SqueezeLLM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Memory Wall
  • 4. Methodology
  • 4.1. Sensitivity-Based Non-uniform Quantization
  • 4.2. Dense-and-Sparse Quantization
  • 4.3. Dense-and-Sparse Kernel Implementation
  • 5. Evaluations
  • 5.1. Experiment Setup
  • 5.2. Main Results
  • 5.3. Quantization of Instruction Following Models
  • 5.4. Hardware Deployment and Profiling
  • 6. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Related Works on Transformer Quantization
  • B. Experiment Setup (Details)
  • C. Data Skew in Per-channel Sparsity Pattern
  • D. Ablation Studies
  • D.1. Sensitivity-Based Quantization.
  • D.2. Impact of Sparsity Levels on SqueezeLLM
  • D.3. Impact of Grouping on SqueezeLLM
  • D.4. Comparison of Optimization Objectives for Non-uniform Quantization: Minimizing Layer-wise Perturbation versus Final Output Perturbation
  • D.5. Impact of Non-uniform Quantization versus Dense-and-Sparse Decomposition
  • D.6. Impact of Dense-and-Sparse Decomposition versus Precision
  • E. Quantization Cost Analysis
  • E.1. Memory Requirement
  • E.2. Quantization Time
  • E.3. Data Efficiency
  • F. Comparison with Other Weight-only Quantization Methods
  • F.1. Comparison with QuIP
  • F.2. Comparison with OmniQuant
  • G. Additional Hardware Profiling Results
  • H. Additional Experiment Results
  • H.1. Perplexity Evaluation
  • H.2. 5-shot MMLU Evaluation
  • I. Limitations

Knowls

  1. Knowl 1 — Sensitivity-Based Non-Uniform Quantization Objective

    model/method

    In SqueezeLLM, low-bit post-training quantization is formulated to minimize the second-order perturbation of the global model loss rather than individual layer activations. Applying a second-order Taylor expansion of the loss function L(W)\mathcal{L}(W) around converged weights WW gives:

    L(WQ)≈L(W)−g⊤(W−WQ)+12(W−WQ)⊤H(W−WQ)\mathcal{L}(W_Q) \approx \mathcal{L}(W) - g^\top (W - W_Q) + \frac{1}{2} (W - W_Q)^\top H (W - W_Q)

    where g=∇WL(W)g = \nabla_W \mathcal{L}(W) is the gradient, H=E[∇W2L(W)]H = \mathbb{E}[\nabla_W^2 \mathcal{L}(W)] is the Hessian matrix, and WQW_Q is the quantized weight tensor. Assuming convergence (g≈0g \approx 0), the optimal quantization configuration Q(w)∗Q(w)^* minimizes (W−WQ)⊤H(W−WQ)(W - W_Q)^\top H (W - W_Q).

    To make computation scalable, the Hessian is approximated by the empirical Fisher Information Matrix F=1∣D∣∑d∈Dgdgd⊤F = \frac{1}{|\mathcal{D}|} \sum_{d \in \mathcal{D}} g_d g_d^\top over a calibration dataset D\mathcal{D}, where gd=∇WL(W;d)g_d = \nabla_W \mathcal{L}(W; d). Assuming cross-weight interactions are negligible simplifies FF to its diagonal elements diag(F)\text{diag}(F). This reduces the quantization of each output channel to a 1-dimensional sensitivity-weighted kk-means clustering problem:

    Q(w)∗≈arg⁡min⁡Q∑i=1NFii(wi−Q(wi))2Q(w)^* \approx \arg\min_Q \sum_{i=1}^N F_{ii} (w_i - Q(w_i))^2

    where wi∈Ww_i \in W is an individual weight, Q(wi)∈{q1,…,qk}Q(w_i) \in \{q_1, \dots, q_k\} is its assigned cluster centroid from a codebook of size k=2bk = 2^b for bb-bit precision, and FiiF_{ii} acts as an importance weight pulling centroids closer to parameters that cause higher loss degradation when perturbed.

  2. Knowl 2 — Dense-and-Sparse Weight Matrix Decomposition

    model/method

    Large language model weights exhibit heavy-tailed distributions where approximately 99.9%99.9\% of values reside in a narrow range spanning roughly 10%10\% of the entire dynamic range, with a small fraction of large outliers and highly sensitive parameters. SqueezeLLM resolves this by decomposing each weight matrix W∈Rm×nW \in \mathbb{R}^{m \times n} into a dense component DD and a sparse component SS:

    W=D+SW = D + S

    where:

    D=W⋅I[Tmin⁡≤w≤Tmax⁡]⋅I[w∉Ωsens]D = W \cdot \mathbb{I}[T_{\min} \le w \le T_{\max}] \cdot \mathbb{I}[w \notin \Omega_{\text{sens}}]

    S=W⋅(I[w<Tmin⁡∨w>Tmax⁡]+I[w∈Ωsens])S = W \cdot (\mathbb{I}[w < T_{\min} \lor w > T_{\max}] + \mathbb{I}[w \in \Omega_{\text{sens}}])

    Here, I[⋅]\mathbb{I}[\cdot] is the indicator function, and thresholds Tmin⁡T_{\min} and Tmax⁡T_{\max} filter out distribution outliers based on extreme percentiles (e.g., 0.40%0.40\% of weights). The set Ωsens\Omega_{\text{sens}} contains the most sensitive weights identified by the largest diagonal Fisher information values FiiF_{ii} (e.g., 0.05%0.05\% of weights), resulting in a combined sparsity of 0.45%0.45\%.

    The dense matrix DD has a significantly reduced dynamic range and is quantized to bb bits (e.g., 3-bit or 4-bit) using sensitivity-weighted kk-means clustering with channel-wise lookup tables (LUTs). The sparse matrix SS retains its values in unquantized FP16 format and is stored in Compressed Sparse Row (CSR) format. During inference, matrix-vector multiplication is performed as WX=DX+SXWX = DX + SX, where XX is the input activation vector, avoiding decompression of SS into dense storage.

  3. Knowl 3 — Balanced Sparse-Dense GPU Matrix-Vector Multiplication Kernel

    model/method

    To accelerate single-batch LLM inference under the memory bandwidth bottleneck, SqueezeLLM implements custom CUDA kernels for the WX=DX+SXWX = DX + SX operation:

    1. Dense LUT Dequantization Kernel: Evaluates DXDX for 3-bit and 4-bit weights. The kernel loads packed 3/4-bit integer indices from global memory and performs on-the-fly dequantization to FP16 using register- and shared-memory-cached channel-wise lookup tables (LUTs), executing inner products in FP16 arithmetic.
    2. Balanced CSR Sparse Kernel: Evaluates SXSX where SS is stored in CSR format. Because outlier entries are heavily skewed across output channels (rows), standard 1-thread-per-row assignments create severe thread divergence. SqueezeLLM implements a balanced hybrid CSR kernel that assigns a fixed budget of 10 non-zero elements per GPU thread, distributing rows across multiple threads with synchronized reduction.
    3. Kernel Fusion: The dense non-uniform kernel and the balanced sparse kernel are launched within a single CUDA kernel call, eliminating kernel launch latency and avoiding redundant global memory read/write cycles for summing DX+SXDX + SX.
  4. Knowl 4 — Perplexity Comparison of SqueezeLLM on LLaMA Models

    data/table

    Perplexity comparison on C4 and WikiText-2 (sequence chunk length 2048) for LLaMA-7B and LLaMA-13B models across 3-bit and 4-bit post-training quantization methods:

    Model / Method Target Bits Avg Bits C4 PPL (↓\downarrow) Wiki PPL (↓\downarrow) Speedup (↑\uparrow) Mem (GB, ↓\downarrow)
    LLaMA-7B
    Baseline FP16 16 16.00 7.08 5.68 1.0×\times 12.7
    RTN 3 3.00 28.26 25.61 2.3×\times 2.9
    GPTQ 3 3.00 9.55 7.55 2.3×\times 2.9
    SqueezeLLM (0%) 3 3.02 7.75 6.32 2.1×\times 2.9
    GPTQ (g128) 3 3.24 7.89 6.27 0.2×\times 3.0
    AWQ (g128) 3 3.24 7.90 6.44 2.0×\times 3.0
    SqueezeLLM (0.45%) 3 3.24 7.56 6.13 1.9×\times 3.1
    RTN 4 4.00 7.73 6.29 2.0×\times 3.7
    GPTQ 4 4.00 7.43 5.94 2.0×\times 3.7
    SpQR 4 3.94 7.28 5.87 1.2×\times N/A
    SqueezeLLM (0%) 4 4.05 7.21 5.79 1.8×\times 3.8
    GPTQ (g128) 4 4.24 7.21 5.78 0.4×\times 3.8
    AWQ (g128) 4 4.24 7.22 5.82 1.6×\times 3.8
    SqueezeLLM (0.45%) 4 4.27 7.18 5.77 1.7×\times 4.0
    LLaMA-13B
    Baseline FP16 16 16.00 6.61 5.09 1.0×\times 24.6
    RTN 3 3.00 13.24 11.78 2.7×\times 5.3
    GPTQ 3 3.00 8.22 6.22 2.7×\times 5.3
    SqueezeLLM (0%) 3 3.02 7.08 5.60 2.4×\times 5.4
    GPTQ (g128) 3 3.25 7.12 5.47 0.2×\times 5.6
    AWQ (g128) 3 3.25 7.08 5.52 2.2×\times 5.7
    SqueezeLLM (0.45%) 3 3.24 6.92 5.45 2.2×\times 5.8
    RTN 4 4.00 6.99 5.53 2.3×\times 6.8
    GPTQ 4 4.00 6.84 5.29 2.3×\times 6.8
    SpQR 4 3.96 6.72 5.22 1.2×\times N/A
    SqueezeLLM (0%) 4 4.04 6.71 5.18 2.0×\times 6.9
    GPTQ (g128) 4 4.25 6.70 5.17 0.4×\times 7.0
    AWQ (g128) 4 4.25 6.70 5.21 1.9×\times 7.2
    SqueezeLLM (0.45%) 4 4.26 6.68 5.17 1.9×\times 7.3

    Dense SqueezeLLM at 3.02 bits reduces C4 perplexity from 9.55 (standard GPTQ) to 7.75 on LLaMA-7B. With 0.45% sparse values added, SqueezeLLM reaches 7.56 PPL at 3.24 effective bits, outperforming both grouped GPTQ (7.89 PPL) and grouped AWQ (7.90 PPL).

  5. Knowl 5 — Final Output Loss Perturbation vs. Layer-Wise Reconstruction Objectives

    empirical result

    Existing PTQ frameworks (including GPTQ, AWQ, and SpQR) optimize layer-wise activation reconstruction arg⁡min⁡Q∥WX−WQX∥22\arg\min_Q \|WX - W_Q X\|_2^2, which weights weight error by input activation magnitudes. SqueezeLLM instead minimizes perturbation to the end-to-end model loss via Fisher information weights arg⁡min⁡Q∑iFii(wi−Q(wi))2\arg\min_Q \sum_i F_{ii} (w_i - Q(w_i))^2.

    Comparing both objectives on 3-bit quantization of LLaMA-7B across varying sparsity levels on the C4 benchmark shows that minimizing final output loss perturbation consistently outperforms layer-wise perturbation minimization by up to 0.3 perplexity points across all sparsity thresholds.

  6. Knowl 6 — Suboptimality of Grouping in Non-Uniform Quantization

    empirical result

    In uniform quantization, dividing channels into small groups (e.g., group size 128) incurs minimal storage overhead because only one FP16 scale and zero point are stored per group. In non-uniform quantization, however, assigning separate quantization bins to each group requires storing a full kk-entry FP16 lookup table (LUT) per group (16×2b16 \times 2^b bits per group), creating substantial memory overhead.

    Evaluating 3-bit LLaMA-7B on C4 reveals that:

    • Non-uniform quantization with group sizes of 512 or 1024, as well as a hybrid configuration (group size 1024 with 0.05% sparsity), achieves strictly inferior perplexity-to-model-size trade-offs compared to channel-wise Dense-and-Sparse decomposition.
    • Channel-wise Dense-and-Sparse decomposition (0.05% sensitive values with variable outlier sparsity) achieves ≈7.56\approx 7.56 perplexity at a relative model size of 0.194, whereas grouping configurations yield perplexity values above 7.61 at relative model sizes exceeding 0.198.
  7. Knowl 7 — Single-Batch Inference Latency and GPTQ Activation Reordering Overhead

    empirical result

    Profiling 128-token generation for 3-bit LLaMA models on an NVIDIA RTX A6000 GPU demonstrates that SqueezeLLM achieves 1.8×\times to 2.4×\times speedups over the FP16 baseline:

    • For LLaMA-7B, latency drops from 3.2 s (FP16, 12.7 GB memory) to 1.5 s for dense SqueezeLLM (2.9 GB) and 1.7 s for SqueezeLLM with 0.45% sparsity (3.1 GB).
    • For LLaMA-13B, latency drops from 5.6 s (FP16, 24.6 GB) to 2.4 s for dense SqueezeLLM (5.4 GB) and 2.5 s for SqueezeLLM with 0.45% sparsity (5.8 GB).
    • For LLaMA-30B and LLaMA-65B, FP16 encounters Out-Of-Memory (OOM) on a single 48 GB GPU, whereas SqueezeLLM runs in 4.0 s (12.5 GB) and 7.6 s (24.5 GB) for dense configurations, and 4.4 s (14.7 GB) and 8.8 s (28.0 GB) for 0.45% sparse configurations.

    In contrast, grouped GPTQ with activation reordering suffers severe slowdowns (13.7 s on 7B, 24.2 s on 13B, 61.9 s on 30B, 117.8 s on 65B). Activation reordering assigns elements within the same channel to different scaling groups accessed via index lookups, destroying coalesced memory access patterns on the GPU.

  8. Knowl 8 — Zero-Shot MMLU Accuracy on Quantized Vicuna Models

    data/table

    Zero-shot weighted accuracy (%) and peak GPU memory (GB) on the Massive Multitask Language Understanding (MMLU) benchmark for Vicuna v1.1 and v1.3 models quantized with AWQ and SqueezeLLM:

    Method Avg Bits Vicuna-7B (v1.1) Vicuna-13B (v1.1) Vicuna-7B (v1.3) Vicuna-13B (v1.3) Vicuna-33B (v1.3)
    Baseline FP16 16.00 39.1% (12.7G) 41.2% (24.6G) 40.2% (12.7G) 43.3% (24.6G) 49.5% (OOM)
    AWQ (g128) 4.25 38.0% (3.8G) 40.4% (7.2G) 39.6% (3.8G) 42.2% (7.2G) 49.5% (17.2G)
    SqueezeLLM (0%) 4.05 38.8% (3.8G) 39.2% (6.9G) 39.3% (3.8G) 44.1% (6.9G) 48.0% (17.5G)
    SqueezeLLM (0.45%) 4.26 39.4% (4.0G) 41.0% (7.3G) 39.5% (4.0G) 43.8% (7.3G) 49.9% (18.7G)
    AWQ (g128) 3.25 36.5% (3.0G) 37.6% (5.7G) 37.4% (3.0G) 40.7% (5.7G) 46.4% (13.2G)
    SqueezeLLM (0%) 3.02 36.0% (2.9G) 37.2% (5.4G) 35.1% (2.9G) 40.5% (5.4G) 46.2% (12.5G)
    SqueezeLLM (0.45%) 3.24 37.7% (3.1G) 39.4% (5.8G) 37.6% (3.1G) 40.8% (5.8G) 47.7% (14.7G)

    At 3-bit precision with matched effective bitwidths (≈3.24\approx 3.24 bits), SqueezeLLM (0.45%) outperforms AWQ (g128) across all configurations (e.g., 39.4% vs 37.6% on Vicuna-13B v1.1, and 47.7% vs 46.4% on Vicuna-33B v1.3). At 4 bits, SqueezeLLM (0.45%) matches or exceeds the FP16 baseline accuracy across all evaluated models.

  9. Knowl 9 — Ablation of Sensitivity-Weighted Clustering and Sparsity Allocation

    empirical result

    Ablation on 3-bit LLaMA-7B evaluated on C4 perplexity (FP16 baseline: 7.08) demonstrates the impact of sensitivity weighting and sparse parameter isolation:

    1. Sensitivity-Agnostic vs. Sensitivity-Based Clustering:

      • At 0% sparsity: unweighted kk-means yields 18.08 PPL, whereas sensitivity-weighted kk-means yields 7.75 PPL (compared to 28.26 for RTN uniform quantization).
      • At 0.05% sparsity: unweighted kk-means yields 8.10 PPL, while sensitivity-weighted yields 7.67 PPL.
      • At 0.45% sparsity: unweighted kk-means yields 7.61 PPL, while sensitivity-weighted yields 7.56 PPL.
    2. Sensitive Value Sparsity Budget: Isolating sensitive parameters into the sparse matrix without outliers shows diminishing perplexity improvements beyond 0.05% sparsity (C4 perplexity drops from 7.75 at 0% to 7.67 at 0.05%, but remains flat at 7.66 at 0.20%).

    3. Combined Outlier and Sensitivity Extraction: Allocating 0.05% of parameters to sensitive values and 0.40% to outlier values consistently outperforms allocating the entire 0.45% budget to outliers alone.

  10. Knowl 10 — Calibration Sample Efficiency and Computational Quantization Overhead

    empirical result

    SqueezeLLM's offline post-training quantization requires two phases:

    1. Fisher Information Calibration: Calibration set size sweeps on LLaMA-2 7B show that as few as 10 samples are sufficient for convergence, achieving 7.72 C4 perplexity and 6.17 WikiText-2 perplexity (matching 100 samples at 7.72 C4 and 6.18 WikiText-2). Computing the diagonal Fisher matrix on an NVIDIA A100 GPU takes 0.3 min (7B), 0.6 min (13B), 1.3 min (30B), and 2.5 min (65B).
    2. Sensitivity-Weighted K-Means: 1D weighted clustering on an Intel Xeon Gold 6126 (48 cores) requires 11 min (7B), 17 min (13B), 45 min (30B), and 80 min (65B), making total quantization time comparable to GPTQ (10 min for 7B, 96 min for 65B).

    Peak host RAM during full-model gradient backpropagation ranges from 33 GB for LLaMA-7B to 292 GB for LLaMA-65B.

  11. Knowl 11 — Low-Bit Quantization Performance Comparison with QuIP and OmniQuant

    data/table

    Perplexity on WikiText-2 for LLaMA-2 models quantized to 4, 3, and 2 bits compared with QuIP (evaluated at sequence length 4096) and OmniQuant (evaluated at sequence length 2048):

    Method Bit Width Avg Bits LLaMA-2 13B (seq 4096) LLaMA-2 70B (seq 4096)
    Baseline FP16 16-bit 16.00 4.81 3.32
    QuIP (reported) 4-bit 4.00 — 3.53
    SqueezeLLM (dense) 4-bit 4.05 4.67 3.21
    QuIP (reported) 3-bit 3.00 — 3.85
    SqueezeLLM (dense) 3-bit 3.02 5.01 3.55
    QuIP (reproduced) 2-bit 2.00 20.54 6.20
    SqueezeLLM (dense) 2-bit 2.01 61.25 10.86
    SqueezeLLM (0.1% sparse) 2-bit 2.05 7.91 5.04
    SqueezeLLM (0.45% sparse) 2-bit 2.22 7.43 4.71

    At 4-bit and 3-bit precision, dense SqueezeLLM outperforms QuIP on LLaMA-2 70B (3.21 vs 3.53 at 4-bit; 3.55 vs 3.85 at 3-bit). At 2-bit precision, dense-only quantization degrades significantly (10.86 PPL on 70B), but adding 0.1% sparsity reduces 2-bit LLaMA-2 70B perplexity to 5.04, outperforming QuIP's 6.20.

  12. Knowl 12 — Limitations of SqueezeLLM

    limitation

    The methodology and empirical evaluation of SqueezeLLM have the following limitations:

    1. Evaluation Architecture: Experiments are focused on autoregressive decoder-only Transformer architectures (LLaMA, LLaMA-2, OPT, Vicuna) for text generation, leaving encoder-only and encoder-decoder architectures unvalidated.
    2. Roofline Modeling Assumptions: The arithmetic intensity and runtime speedup models rely on simulation-based roofline analysis under single-batch conditions, making simplified assumptions about GPU pipeline scheduling and memory controllers.
    3. Host Memory Footprint During Quantization: Unlike layer-wise PTQ methods that only process forward activations, computing the Fisher information matrix requires back-propagating gradients through the entire model, demanding up to 292 GB of host memory for a 65B parameter model.

Coverage note — Full perplexity tables for OPT models across all sizes and individual GPT-4 instruction-following win-rate bar charts were omitted as they reinforce the primary findings demonstrated in the LLaMA/Vicuna benchmark tables.

References

  1. 1.Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I. BinaryBERT: Pushing the limit of BERT quantization. arXiv preprint arXiv:2012.15701, 2020.
  2. 2.Bondarenko, Y., Nagel, M., and Blankevoort, T. Understanding and overcoming the challenges of efficient Transformer quantization. arXiv preprint arXiv:2109.12948, 2021.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  4. 4.Cai, Y., Yao, Z., Dong, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. ZeroQ: A novel zero shot quantization framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13169–13178, 2020.
  5. 5.Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. M. Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems, 36, 2024.
  6. 6.Chen, B., Dao, T., Winsor, E., Song, Z., Rudra, A., and Re,́ C. Scatterbrain: Unifying sparse and low-rank attention. Advances in Neural Information Processing Systems, 34: 17413–17426, 2021.
  7. 7.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  8. 8.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  9. 9.Chung, I., Kim, B., Choi, Y., Kwon, S. J., Jeon, Y., Park, B., Kim, S., and Lee, D. Extremely low bit transformer quantization for on-device neural machine translation. arXiv preprint arXiv:2009.07453, 2020.
  10. 10.Dass, J., Wu, S., Shi, H., Li, C., Ye, Z., Wang, Z., and Lin, Y. Vitality: Unifying low-rank and sparse approximation for vision transformer acceleration with a linear taylor attention. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp. 415–428. IEEE, 2023.
  11. 11.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems.
  12. 12.Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D. SpQR: A sparse-quantized representation for near-lossless LLM weight compression. arXiv preprint arXiv:2306.03078, 2023.
  13. 13.Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems, 36, 2024.
  14. 14.Dong, Z., Yao, Z., Arfeen, D., Gholami, A., Mahoney, M. W., and Keutzer, K. HAWQ-V2: Hessian Aware trace-Weighted Quantization of neural networks. NeurIPS’19 workshop on Beyond First-Order Optimization Methods in Machine Learning., 2019.
  15. 15.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. GLAM: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pp. 5547–5569. PMLR, 2022.
  16. 16.Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D. Extreme compression of large language models via additive quantization. arXiv preprint arXiv:2401.06118, 2024.
  17. 17.Evtushenko, G. Sparse Matrix-Vector Multiplication with CUDA. https://medium.com/analytics-vidhya/sparse-matrix-vector-multiplication-with-cuda-42d191878e8f, 2019.
  18. 18.Flegar, G. and Quintana-Ort´ı, E. S. Balanced csr sparse matrix-vector product on graphics processors. In Euro-Par 2017: Parallel Processing: 23rd International Conference on Parallel and Distributed Computing, Santiago de Compostela, Spain, August 28–September 1, 2017, Proceedings 23, pp. 697–709. Springer, 2017.
  19. 19.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  20. 20.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., et al. A framework for few-shot language model evaluation, 2021.
  21. 21.Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  22. 22.Gholami, A., Yao, Z., Kim, S., Hooper, C., Mahoney, M. W., and Keutzer, K. Ai and memory wall. IEEE Micro, pp. 1–5, 2024.
  23. 23.GPTQ-For-LLaMA. https://github.com/qwopqwop200/gptq-for-llama.
  24. 24.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. International Conference on Learning Representations, 2016.
  25. 25.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021.
  26. 26.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  27. 27.Huang, Y., Yang, H., Dong, Z., Gudovskiy, D., Okuno, T., Nakata, Y., Du, Y., Zhang, S., and Keutzer, K. Output sensitivity-aware detr quantization. 2023.
  28. 28.Jeon, Y., Lee, C., Cho, E., and Ro, Y. Mr. BiQ: Post-training non-uniform quantization based on minimizing the reconstruction error. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12329–12338, 2022.
  29. 29.Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K. I-BERT: Integer-only bert quantization. arXiv preprint arXiv:2101.01321, 2021.
  30. 30.Kim, S., Hooper, C., Wattanawong, T., Kang, M., Yan, R., Genc, H., Dinh, G., Huang, Q., Keutzer, K., Mahoney, M. W., Shao, S., and Gholami, A. Full stack optimization of transformer inference: a survey. arXiv preprint arXiv:2302.14017, 2023.
  31. 31.Kovaleva, O., Kulshreshtha, S., Rogers, A., and Rumshisky, A. Bert busters: Outlier dimensions that disrupt transformers. arXiv preprint arXiv:2105.06990, 2021.
  32. 32.LeCun, Y., Denker, J. S., and Solla, S. A. Optimal brain damage. In Advances in neural information processing systems, pp. 598–605, 1990.
  33. 33.Li, X., Liu, Y., Lian, L., Yang, H., Dong, Z., Kang, D., Zhang, S., and Keutzer, K. Q-diffusion: Quantizing diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 17535–17545, October 2023.
  34. 34.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. 2023.
  35. 35.Liu, Y., Yang, H., Dong, Z., Keutzer, K., Du, L., and Zhang, S. NoisyQuant: Noisy bias-enhanced post-training activation quantization for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20321–20330, 2023.
  36. 36.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models, 2016.
  37. 37.Oh, S., Sim, H., Kim, J., and Lee, J. Non-uniform step size quantization for accurate post-training quantization. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XI, pp. 658–673. Springer, 2022.
  38. 38.Patterson, D. A. Latency lags bandwith. Communications of the ACM, 47(10):71–75, 2004.
  39. 39.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  40. 40.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagne, R., Luccioni, A. S., Yvon, F., Gall é, M., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  41. 41.Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P. Omniquant: Omnidirectionally calibrated quantization for large language models. arXiv preprint arXiv:2308.13137, 2023.
  42. 42.Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K. Q-BERT: Hessian based ultra low precision quantization of bert. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 8815–8821, 2020.
  43. 43.Shomron, G., Gabbay, F., Kurzum, S., and Weiser, U. Post-training sparsity-aware quantization. Advances in Neural Information Processing Systems, 34:17737–17748, 2021.
  44. 44.Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990, 2022.
  45. 45.Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  46. 46.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  47. 47.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  48. 48.Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X. Outlier suppression: Pushing the limit of low-bit transformer language models. arXiv preprint arXiv:2209.13325, 2022.
  49. 49.Wei, X., Zhang, Y., Li, Y., Zhang, X., Gong, R., Guo, J., and Liu, X. Outlier suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling. arXiv preprint arXiv:2304.09145, 2023.
  50. 50.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 38087–38099. PMLR, 23–29 Jul 2023.
  51. 51.Xu, Y., Wang, Y., Zhou, A., Lin, W., and Xiong, H. Deep neural network compression with single and multiple level quantization, 2018.
  52. 52.Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. arXiv preprint arXiv:2206.01861, 2022.
  53. 53.Yuan, Z., Niu, L., Liu, J., Liu, W., Wang, X., Shang, Y., Sun, G., Wu, Q., Wu, J., and Wu, B. RPTQ: Reorder-based post-training quantization for large language models. arXiv preprint arXiv:2304.01089, 2023.
  54. 54.Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A. GOBO: Quantizing attention-based nlp models for low latency and energy efficient inference. In 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pp. 811–824. IEEE, 2020.
  55. 55.Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M. Q8BERT: Quantized 8bit bert. arXiv preprint arXiv:1910.06188, 2019.
  56. 56.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  57. 57.Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q. TernaryBERT: Distillation-aware ultra-low bit bert. arXiv preprint arXiv:2009.12812, 2020.
  58. 58.Zhang, Y., Dong, Z., Yang, H., Lu, M., Tseng, C.-C., Guo, Y., Keutzer, K., Du, L., and Zhang, S. Qd-bev: Quantization-aware view-guided distillation for multi-view 3d object detection. 2023.
  59. 59.Zhao, R., Hu, Y., Dotzel, J., De Sa, C., and Zhang, Z. Improving neural network quantization without retraining using outlier channel splitting. In International conference on machine learning, pp. 7543–7552. PMLR, 2019.

Citation

MLA
Kim, S., et al. “SqueezeLLM: Dense-and-Sparse Quantization”. arXiv, 2023, http://arxiv.org/abs/2306.07629v4.
APA
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., & Keutzer, K. (2023). SqueezeLLM: Dense-and-Sparse Quantization. arXiv. http://arxiv.org/abs/2306.07629v4
Chicago
Kim, S., C. Hooper, A. Gholami, et al. 2023. “SqueezeLLM: Dense-and-Sparse Quantization”. arXiv. http://arxiv.org/abs/2306.07629v4.
Harvard
Kim, S. et al. (2023) “SqueezeLLM: Dense-and-Sparse Quantization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.07629v4.
Vancouver
1. Kim S, Hooper C, Gholami A, Dong Z, Li X, Shen S, Mahoney MW, Keutzer K (2023) SqueezeLLM: Dense-and-Sparse Quantization. arXiv

BibTeX

@article{kim2023squeezellm,
  title = {SqueezeLLM: Dense-and-Sparse Quantization},
  author = {Kim, Sehoon and Hooper, Coleman and Gholami, Amir and Dong, Zhen and Li, Xiuyu and Shen, Sheng and Mahoney, Michael W. and Keutzer, Kurt},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.07629v4},
  eprint = {2306.07629}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/