Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
Junyuan HongJinhao DuanChenhui ZhangZhangheng LiChulin XieKelsey LiebermanJames DiffenderferBrian R. BartoldsonAjay Kumar JaiswalKaidi Xu
Reveals how model compression affects large language model trustworthiness across eight safety dimensions, demonstrating that 4-bit quantization preserves or even improves safety profiles while pruning and extreme 3-bit quantization severely degrade them despite maintaining benign task performance.
Organizations increasingly deploy compressed large language models to reduce hardware and operational costs while speeding up response times. However, conventional benchmarks only evaluate standard language capability, ignoring how compression alters essential safety, compliance, and ethical behaviors in production settings.
The article evaluates how modern compression techniques affect language model trustworthiness across eight critical dimensions: stereotype bias, toxicity, privacy preservation, fairness, machine ethics, adversarial robustness, out-of-distribution robustness, and resilience to adversarial demonstrations. It aims to identify the optimal balance between efficiency, task performance, and safety risks.
The authors conducted empirical evaluations across three 13-billion-parameter foundation models (LLAMA2, LLAMA2 Chat, and Vicuna Chat) using five compression techniques. These techniques included two post-training weight quantization methods (GPTQ and AWQ) and three 50% pruning methods (Magnitude, SparseGPT, and Wanda). Evaluations were executed across the 57-task Massive Multitask Language Understanding (MMLU) benchmark and 33 specific test cases from the DecodingTrust benchmark.
First, quantization preserves trustworthiness significantly better than pruning; 8-bit quantization maintains original baseline performance within 3 points, whereas 50% pruning degrades key safety metrics by more than 5 to 40 points. Second, 4-bit quantization serves as the optimal operational target, maintaining benign task performance while preserving overall trust scores within a 5-point threshold. Third, moderate 4-bit quantization unexpectedly improves specific dimensions, raising machine ethics scores by up to 22 points and reducing demographic unfairness without increasing silent refusals. Fourth, aggressive 3-bit quantization triggers severe safety failures, including a drop of roughly 50 points in toxicity defense under GPTQ, caused by the model losing its ability to follow system safety prompts. Finally, these critical safety deteriorations cannot be detected through standard capability benchmarks like MMLU alone.
These findings indicate that teams can achieve hardware efficiency without compromising alignment by adopting moderate quantization, whereas aggressive pruning introduces hidden operational and compliance risks. Standard performance checks are insufficient to verify model safety before deployment. Organizations should select well-aligned source models and favor activation-aware quantization (such as AWQ) at 4 bits. Before deploying models compressed with random calibration data or reduced to extreme bit rates, teams must conduct comprehensive multi-dimensional trust evaluations.
Confidence in these findings is high for post-training quantization on open-source foundation models. However, readers should note that results vary depending on calibration data, and base model alignment heavily shapes the compressed model's final reliability.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Read AWQ first to understand the activation-aware quantization method that this study evaluates and recommends for compressed-model deployment.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). GPTQ establishes the post-training quantization method used in this study, making its compression approach and bit-width tradeoffs easier to interpret.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT explains one of the pruning methods evaluated here, providing the necessary context for interpreting its pruning-related trustworthiness results.
No sufficiently relevant recommendations were found.
