SqueezeLLM: Dense-and-Sparse Quantization
Sehoon KimColeman HooperAmir GholamiZhen DongXiuyu LiSheng ShenMichael W. MahoneyKurt Keutzer
Presents SqueezeLLM, a post-training quantization framework combining sensitivity-based non-uniform quantization and dense-and-sparse matrix decomposition to achieve near-lossless 3-bit weight compression and up to 2.3× faster inference on large language models.
Deploying generative large language models for real-world tasks is heavily constrained by high computing resource requirements and memory limitations. During single-batch generative inference, performance is primarily constrained by memory bandwidth—the rate at which data moves from memory to processing cores—rather than raw arithmetic throughput. While quantizing weights to lower bit precisions reduces memory footprint, conventional low-bit compression methods frequently lead to steep drops in output quality and accuracy.
The article introduces and evaluates SqueezeLLM, a post-training quantization framework designed to enable lossless weight compression down to 3-bit precision while maintaining high inference speed and model accuracy without requiring full retraining.
The authors conducted empirical benchmarks, hardware profiling, and ablation studies across multiple model families, including LLaMA, LLaMA-2, OPT, and Vicuna, spanning sizes from 1.3 billion to 70 billion parameters. The method relies on two core mechanisms: sensitivity-based non-uniform quantization, which uses second-order loss sensitivity (approximated via Fisher information) to assign quantization levels near the most critical weights, and a dense-and-sparse decomposition, which isolates a tiny fraction of sensitive and outlier values into a full-precision sparse format while compressing the remaining dense matrix into low bitwidths.
The evaluation yielded several key findings. First, SqueezeLLM achieves significant compression with minimal quality loss; for example, on 3-bit LLaMA models, it reduces the perplexity gap to the 16-bit baseline by up to 2.1 times compared to previous state-of-the-art techniques. Second, retaining merely 0.45% of parameters in full-precision sparse format substantially improves quantization resolution, outperforming traditional grouping methods. Third, hardware deployment on A6000 and A100 GPUs demonstrated up to a 2.4-fold inference speedup and an approximate 4-fold reduction in peak memory usage relative to the uncompressed 16-bit baseline. Finally, the framework preserves performance on downstream tasks, maintaining high zero-shot and few-shot accuracy on standard knowledge benchmarks and instruction-following evaluations.
These findings indicate that generative model deployment costs and hardware barriers can be substantially reduced without compromising model capability. By addressing memory bandwidth bottlenecks rather than arithmetic operations, organizations can run larger models on fewer or smaller hardware instances, mitigating operational expense and avoiding complex multi-chip serving setups.
Decision-makers seeking to optimize large language model serving pipelines should consider adopting dense-and-sparse non-uniform quantization formats for single-batch and low-latency deployments. Where maximum accuracy is needed at ultra-low bitwidths, configuring a small sparsity allocation (such as 0.45%) provides the best balance between memory overhead and model quality compared to standard grouping approaches.
Confidence in these findings is supported by extensive empirical testing across varied architectures and actual hardware profiling. However, the study's primary scope focuses on generative decoder-only architectures and single-batch inference regimes. Readers should exercise caution when extrapolating these throughput gains to large-batch processing, where compute throughput becomes the dominant constraint, or to encoder-based network architectures.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). This paper establishes the second-order post-training quantization framework for large language models that SqueezeLLM directly builds upon and enhances with non-uniform sensitivity analysis.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). This work introduces the foundational concept of isolating outlier channels in transformer hidden states, motivating SqueezeLLM's Dense-and-Sparse decomposition.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). This work details activation outlier mitigation in post-training quantization, providing necessary context for why sensitive values must be handled separately in low-bit LLM regimes.
- Paper: LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation, Yixiao Li et al. (2023). This paper demonstrates decomposing weight matrices into dense and sparse components to balance compression and expressiveness, foundational to SqueezeLLM's structural design.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). This comprehensive survey outlines the mathematical foundations and hardware trade-offs of low-precision and non-uniform quantization schemes.
- Paper: Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through Estimation, Zechun Liu et al. (2022). This work examines the trade-offs and bridging techniques between non-uniform bit representations and hardware-efficient execution.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). This work advances ultra-low-bit LLM quantization below 3 bits by replacing explicit sparse outlier storage with randomized Hadamard transformations and lattice codebooks.
- Paper: BiLLM: Pushing the Limit of Post-Training Quantization for LLMs, Wei Huang et al. (2024). This paper extends post-training compression into the extreme 1-bit binarization regime using structured residual approximations for highly sensitive parameters.
- Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). This work builds on second-order sensitivity quantization by incorporating Kronecker-factored gradient statistics alongside activation covariances for sub-3-bit compression.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). This paper explores an alternative paradigm to retain sub-4-bit accuracy by augmenting outlier-aware initialization with self-distillation.
- Paper: Evaluating Quantized Large Language Models, Shiyao Li et al. (2024). This study conducts a systematic downstream benchmark evaluating the real-world trade-offs and error compounding of low-bit quantization across complex reasoning and context lengths.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). This paper investigates the safety, robustness, and ethical alignment impacts of deploying LLMs compressed with post-training quantization methods.
- Paper: Exploiting LLM Quantization, Kazuki Egashira et al. (2024). This research analyzes security vulnerabilities uniquely introduced by post-training quantization, demonstrating how quantization boundaries can be exploited.
