Evaluating Quantized Large Language Models
Shiyao LiXuefei NingLuning WangTengxuan LiuXiangsheng ShiShengen YanGuohao DaiHuazhong YangYu Wang
Presents a comprehensive empirical evaluation of post-training quantization across 11 large language model families, revealing critical sensitivity patterns across weights, activations, and KV caches to guide optimal bit-width selection for diverse tasks.
Deploying large language models presents significant operational hurdles due to immense memory demands and computational overhead. Post-training quantization—a compression method that converts high-precision numerical values to lower bit-widths—offers a practical way to curb hardware requirements and operational costs. However, quantization is inherently lossy, and its practical trade-offs across different model architectures, tensor types, and diverse operational tasks have not been systematically understood.
The article provides a comprehensive evaluation of post-training quantization across model weights, activations, and key-value attention caches. It evaluates 11 model families ranging from 125 million to 180 billion parameters across five task domains: basic language processing, emergent capabilities, model trustworthiness, multi-turn dialogue, and long-context comprehension.
The authors conducted extensive empirical testing comparing standard uniform quantization and state-of-the-art recovery techniques against full-precision baselines across dozens of public benchmark datasets. Statistical properties, including tensor outliers and standard deviations, were analyzed to explain model behavior under reduced precision.
The evaluation yielded several critical findings regarding quantization tolerance. First, model size creates divergent sensitivities: larger models tolerate aggressive weight and key-value cache quantization better due to fewer outlier values, but they show substantially lower tolerance to activation quantization due to heavy-tailed outliers. Second, task difficulty dictates vulnerability; multi-step mathematical reasoning and self-calibration degrade sharply at lower bit-widths, primarily driven by logical reasoning errors rather than arithmetic mistakes, whereas basic language understanding remains robust. Third, multi-turn dialogue and long-context processing are exceptionally sensitive to precision loss; long-context tasks of 4,000 tokens or more experience severe performance drops under key-value cache compression below 8 bits, while dialogues collapse into repetitive sentences and random tokens below 4 bits. Fourth, architectural scaling methods like Mixture-of-Experts boost raw capability but fail to improve quantization tolerance compared to dense models of similar parameter scale. Finally, popular recovery techniques like AWQ and SmoothQuant fail to restore model viability under extreme low-bit regimes such as 2-bit weights or 4-bit activations.
These findings provide clear guidance for managing the trade-offs between inference speed, memory footprint, and model accuracy. Organizations can safely implement 4-bit weight, 8-bit activation, and 4-bit key-value cache quantization for routine language understanding and short-context workflows while keeping accuracy loss under 2 percent. Conversely, deploying lower precisions for complex reasoning, extended context windows, or small models under 13 billion parameters introduces significant operational risk and capability failure.
Decision-makers should establish task-specific quantization policies: maintain at least 8-bit precision across all tensors for models under 13 billion parameters handling reasoning tasks, preserve 8-bit key-value caches for contexts exceeding 4,000 tokens, and limit aggressive 4-bit compression to large, general-purpose models. Where hardware budgets are constrained, adopting a larger model compressed to 3-bit weights often outperforms a smaller model running at full precision.
Confidence in these findings is high for the evaluated architectures, though the specific bit-width thresholds may not generalize universally to unexamined model families. Because the article exclusively examines post-training quantization without fine-tuning, organizations pursuing extreme compression below 4 bits must conduct targeted pilot evaluations using quantization-aware training before deploying to production.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). It introduces SmoothQuant, a fundamental post-training activation-weight migration method evaluated directly as a recovery baseline in the source paper.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). It introduces Activation-aware Weight Quantization (AWQ), which provides the foundation for the weight-activation protection techniques evaluated in the source.
- Paper: LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, Tim Dettmers et al. (2022). It reveals the emergence and structural role of activation outliers in large language models, explaining the theoretical cause of activation quantization sensitivity analyzed throughout the source.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). It establishes GPTQ, a standard second-order post-training weight quantization framework that serves as a cornerstone baseline for the empirical comparisons in the source.
- Paper: FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU, Ying Sheng et al. (2023). It introduces 4-bit key-value attention cache and weight quantization for high-throughput LLM inference, establishing key mechanics of the KV cache compression evaluated in the source.
- Paper: A Survey of Quantization Methods for Efficient Neural Network Inference, Amir Gholami et al. (2021). It offers a comprehensive foundational survey on uniform versus non-uniform quantization trade-offs across neural network weights and activations.
- Paper: DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs, Haokun Lin et al. (2024). It introduces dual transformations to overcome the extreme activation outlier failures at low bit-widths that the source identifies in standard post-training quantization.
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, Albert Tseng et al. (2024). It designs randomized Hadamard incoherence transforms and lattice codebooks to preserve accuracy in the ultra-low 2-bit and 3-bit regimes where the source finds standard methods collapse.
- Paper: PolarQuant: Quantizing KV Caches with Polar Transformation, Insu Han et al. (2025). It develops a coordinate-transform compression method to solve the severe long-context key-value cache degradation below 8 bits documented in the source.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). It applies 1-bit Johnson-Lindenstrauss random projections to address the memory bottlenecks and precision loss in key-value caches evaluated by the source.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). It expands the source's findings on trustworthiness and safety degradation by rigorously auditing bias, toxicity, and adversarial robustness under post-training compression.
- Paper: BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation, Dayou Du et al. (2024). It explores self-distillation to restore reasoning performance in sub-4-bit regimes after post-training methods fail at low precision.
- Paper: BiLLM: Pushing the Limit of Post-Training Quantization for LLMs, Wei Huang et al. (2024). It pushes post-training quantization to the extreme 1-bit boundary using residual binarization, directly confronting the low-bit limits identified in the source.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). It complements the source's key-value cache quantization findings by introducing query-driven channel pruning to reduce KV cache memory without precision-related degradation.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). It uses low-rank plus quantized matrix decomposition to recover model capabilities in sub-4-bit settings where pure post-training quantization fails.
- Paper: Scaling FP8 training to trillion-token LLMs, Maxim Fishman et al. (2025). It addresses low-precision activation instabilities during training, offering a complementary training-time perspective on the outlier behaviors examined during inference in the source.
