Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression

Junyuan HongJinhao DuanChenhui ZhangZhangheng LiChulin XieKelsey LiebermanJames DiffenderferBrian R. BartoldsonAjay Kumar JaiswalKaidi Xu

article2024ICML79 citations

Reveals how model compression affects large language model trustworthiness across eight safety dimensions, demonstrating that 4-bit quantization preserves or even improves safety profiles while pruning and extreme 3-bit quantization severely degrade them despite maintaining benign task performance.

Listen

Organizations increasingly deploy compressed large language models to reduce hardware and operational costs while speeding up response times. However, conventional benchmarks only evaluate standard language capability, ignoring how compression alters essential safety, compliance, and ethical behaviors in production settings.

The article evaluates how modern compression techniques affect language model trustworthiness across eight critical dimensions: stereotype bias, toxicity, privacy preservation, fairness, machine ethics, adversarial robustness, out-of-distribution robustness, and resilience to adversarial demonstrations. It aims to identify the optimal balance between efficiency, task performance, and safety risks.

The authors conducted empirical evaluations across three 13-billion-parameter foundation models (LLAMA2, LLAMA2 Chat, and Vicuna Chat) using five compression techniques. These techniques included two post-training weight quantization methods (GPTQ and AWQ) and three 50% pruning methods (Magnitude, SparseGPT, and Wanda). Evaluations were executed across the 57-task Massive Multitask Language Understanding (MMLU) benchmark and 33 specific test cases from the DecodingTrust benchmark.

First, quantization preserves trustworthiness significantly better than pruning; 8-bit quantization maintains original baseline performance within 3 points, whereas 50% pruning degrades key safety metrics by more than 5 to 40 points. Second, 4-bit quantization serves as the optimal operational target, maintaining benign task performance while preserving overall trust scores within a 5-point threshold. Third, moderate 4-bit quantization unexpectedly improves specific dimensions, raising machine ethics scores by up to 22 points and reducing demographic unfairness without increasing silent refusals. Fourth, aggressive 3-bit quantization triggers severe safety failures, including a drop of roughly 50 points in toxicity defense under GPTQ, caused by the model losing its ability to follow system safety prompts. Finally, these critical safety deteriorations cannot be detected through standard capability benchmarks like MMLU alone.

These findings indicate that teams can achieve hardware efficiency without compromising alignment by adopting moderate quantization, whereas aggressive pruning introduces hidden operational and compliance risks. Standard performance checks are insufficient to verify model safety before deployment. Organizations should select well-aligned source models and favor activation-aware quantization (such as AWQ) at 4 bits. Before deploying models compressed with random calibration data or reduced to extreme bit rates, teams must conduct comprehensive multi-dimensional trust evaluations.

Confidence in these findings is high for post-training quantization on open-source foundation models. However, readers should note that results vary depending on calibration data, and base model alignment heavily shapes the compressed model's final reliability.

arXiv: 2403.15447

No sufficiently relevant recommendations were found.

Cover for Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Assessing the Trustworthiness of Compressed LLMs
  • 4. Revisiting Paths to 7B-sized LLMs: Training Smaller, or Compressing Larger?
  • 5. From Moderate to High Compression Rates: The (Unexpected) Gains and Losses
  • 5.1. Finding the Essential Compression Rates and Induced Gains for Trustworthiness
  • 5.2. The Losses on the Extreme Compression Rate
  • 6. Bag of Tricks for Trustworthy Compression
  • 7. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Additional Related Works
  • B. Additional Experimental Results
  • C. Detailed Breakdown Results of DecodingTrust Benchmark
  • C.1. AdvGLUE++
  • C.2. Adversarial Demonstration
  • C.3. Out-of-Distribution (OOD)
  • C.4. Fairness
  • C.5. Machine Ethics
  • C.6. Privacy
  • C.7. Stereotype
  • C.8. Toxicity

Knowls

  1. Knowl 1 — Quantization is the most reliable compression route for joint efficiency and trustworthiness

    empirical result

    Across LLAMA2 13b, LLAMA2 13b Chat, and Vicuna 13b Chat, weight quantization generally preserves trustworthiness more reliably than 50% hardware-friendly pruning at comparable compressed size. Eight-bit quantization causes less than a 3-point decrease from the corresponding dense model across MMLU and all eight trustworthiness dimensions, whereas 2:4 pruning produces losses exceeding 5 points in at least three dimensions and varies substantially across source models and pruning methods. Four-bit quantization is usually the practical trustworthiness sweet spot, while three-bit quantization can create severe trust failures that are not visible in MMLU alone. The trustworthiness of a compressed model therefore depends strongly on both the compression method and the dense model from which it was produced.

  2. Knowl 2 — Comprehensive evaluation protocol for compressed LLMs

    experimental setup

    The study evaluates three 13-billion-parameter dense models: LLAMA2 13b, LLAMA2 13b Chat, and Vicuna 13b Chat. The models cover a base model and two chat-oriented models, providing different alignment and trustworthiness profiles. Each dense or compressed model is evaluated using MMLU average accuracy over 57 language-understanding tasks and eight DecodingTrust dimensions: AdvGLUE++ adversarial robustness, out-of-distribution robustness, robustness to adversarial demonstrations, machine ethics, fairness, toxicity, privacy, and stereotype. The experiments compare dense models with 7-billion-parameter models, 7-billion-sized compressed models, and models compressed to 3-, 4-, and 8-bit precision or 50% sparsity. Additional experiments use GPTQ-compressed LLAMA2 Chat models with 7, 13, and 70 billion parameters to test whether source-model size changes compression reliability.

  3. Knowl 3 — Post-training compression configurations

    model/method

    The evaluated pruning methods are one-shot magnitude pruning, SparseGPT, and Wanda. Each uses hardware-friendly 2:4 semi-structured sparsity, meaning that two weights remain nonzero within every consecutive group of four weights. Magnitude pruning uses no calibration data; SparseGPT calibrates and updates weights using 128 examples; Wanda selects weights using weight magnitude combined with activation magnitude and also uses 128 examples. The evaluated quantizers are GPTQ and AWQ at 3-, 4-, and 8-bit weight precision. GPTQ uses a weight-based calibration criterion and AWQ uses activation-aware calibration; both use 128 calibration examples. No method receives additional post-compression training. For the 50% SparseGPT comparisons, three independently sampled C4 calibration sets are used and the reported result is their average; the 3- and 4-bit quantization experiments are also repeated with three random calibration seeds.

  4. Knowl 4 — Trustworthiness scoring and refusal handling

    definition

    All benchmark results are reported on the DecodingTrust 0–100 normalized scale, with MMLU reported as average accuracy. Refusal rates are measured alongside the main trustworthiness scores because an answer outside the expected response set can reflect either inability to answer or a safety refusal. For classification accuracy tasks such as AdvGLUE++, a refusal is treated as an incorrect answer; for privacy, a refusal is treated as safe because it does not reveal private information. To reduce biases from differing refusal rates among open-source models, the study includes refused outputs in fairness evaluation and treats them as fair failures, while ethics evaluation treats refusal to identify an immoral action as a negative response that lowers the false-positive rate. Consequently, a lower refusal rate is not automatically safer: its effect depends on the trustworthiness dimension and task.

  5. Knowl 5 — Pre-training a 7-billion model versus compressing a 13-billion model

    empirical result

    The comparison between dense 7-billion and 13-billion models shows that smaller pre-trained models can outperform larger ones on selected trust dimensions. The 13-billion models are consistently better on MMLU, robustness to adversarial demonstrations, and ethics, but 7-billion models are often better on out-of-distribution robustness, AdvGLUE++ robustness, and fairness; fairness favors the 7-billion counterpart across all three model families. In the LLAMA2 Chat comparison, the 7-billion model exceeds the 13-billion model by more than 5 points on out-of-distribution robustness, AdvGLUE++, and fairness. For the non-chat LLAMA2 model, differences between 7-billion and 13-billion versions range from roughly 10 to 52 points across dimensions, suggesting that alignment affects how trustworthiness changes with model size. Despite these isolated advantages of pre-trained small models, eight-bit quantization of a 13-billion model provides a more consistent 7-billion-sized substitute.

  6. Knowl 6 — Behavior of 50%-compressed 7-billion-sized models

    empirical result

    When 13-billion models are compressed to approximately the memory and computation size of a 7-billion, 16-bit model, eight-bit quantization is consistently close to the original 13-billion dense model in both benign performance and trustworthiness. The compressed model largely inherits the source model's strengths and weaknesses, and this preservation also occurs for the non-chat LLAMA2 model, so it does not depend on conversational alignment. In contrast, 2:4 pruning produces serious losses in at least three dimensions and inconsistent changes across source models. Pruning can improve an individual dimension, such as AdvGLUE++ or fairness for a particular model, but these gains are not reliable across models. Four-bit quantization of a 13-billion model is more memory-efficient and more accurate on MMLU than a dense 7-billion model, while also outperforming the 7-billion model on several trust dimensions for which the dense 13-billion source model is weak.

  7. Knowl 7 — Four-bit quantization can preserve and sometimes improve trustworthiness

    empirical result

    For LLAMA2 13b Chat, four-bit quantization is the highest compression rate tested that keeps every evaluated trustworthiness dimension within 5 points of the 16-bit dense model while retaining useful MMLU performance. The compressed model can also improve dimensions on which the dense source model performs poorly. In the reported ethics aggregate, the dense model scores 54.1, whereas four-bit GPTQ and AWQ score 76.3 and 62.8, respectively. GPTQ four-bit quantization also substantially reduces unfairness in few-shot fairness tests: the equalized-odds difference decreases by more than 0.2 in the 16-shot setting without increasing refusal relative to the dense model. The gains are not universal across prompts, but similar four-bit improvements are observed in the other two evaluated model families, indicating that moderate quantization can produce a low-cost efficiency–trustworthiness improvement.

  8. Knowl 8 — Extreme three-bit quantization creates hidden trust risks

    empirical result

    Three-bit quantization can retain apparently acceptable benign performance while substantially degrading trustworthiness. AWQ three-bit remains within approximately 5 points of the dense model on many dimensions, but it exhibits significant degradation and high variance in robustness to adversarial demonstrations and fairness. GPTQ is more unstable: its worst observed losses include roughly 30 points in out-of-distribution robustness and 50 points in toxicity, despite an MMLU decrease of only about 8 points. In toxicity tests, three-bit GPTQ answers roughly 80% of toxic prompts instead of refusing at the dense model's approximately 10–60% rate, indicating that it often ignores the LLAMA2 Chat system instruction not to produce toxic content. In out-of-distribution tests, GPTQ frequently fails to follow the required answer format and emits random or empty outputs. The corresponding MT-Bench instruction-following scores for LLAMA2 13b Chat are 2.89, 6.55, 6.85, and 7.00 for GPTQ at 3, 4, 8, and 16 bits, compared with 6.42, 6.73, 6.99, and 7.00 for AWQ. These results support the paper's conjecture that GPTQ's catastrophic three-bit failures are largely related to lost instruction-following ability, while showing that standard benign evaluation cannot reliably expose the risk.

  9. Knowl 9 — Compression reliability depends strongly on source-model size

    data/table

    The following results compare dense 16-bit LLAMA2 Chat models with their four-bit GPTQ versions. The columns are MMLU, AdvGLUE++, out-of-distribution robustness, adversarial-demonstration robustness, ethics, fairness, toxicity, privacy, and stereotype scores, all reported in points. Four-bit GPTQ is relatively close to the 13-billion and 70-billion dense models, but it causes larger trustworthiness losses for the 7-billion source model even when MMLU changes little.

    Could not parse LaTeX table

    The 7-billion source model loses 12.9 ethics points, 9.1 fairness points, and 8.1 out-of-distribution points after four-bit GPTQ, while its MMLU loss is only 1.6 points. This demonstrates that smaller source models contain less redundancy for quantization and that MMLU alone understates their trustworthiness degradation.

  10. Knowl 10 — Trust effects are dimension-specific and algorithm-dependent

    empirical result

    The detailed trustworthiness evaluations show that compression does not produce a single monotonic change across all dimensions. AWQ is generally more stable than GPTQ on adversarial demonstrations, out-of-distribution robustness, and adversarial robustness. In adversarial-demonstration tests, AWQ remains close to the dense model, whereas GPTQ degrades especially at low bit widths; compression can nevertheless improve counterfactual-demonstration robustness. Quantization usually harms out-of-distribution performance on transformed input styles and unknown knowledge, although in-context examples can restore style robustness toward dense-model levels. Privacy effects are task-dependent: low-bit AWQ produces about 10% more personally identifiable-information leakage while better recognizing privacy-sensitive words and events, whereas GPTQ shows the opposite pattern and becomes worse at private-event recognition at higher precision. Quantization induces more stereotype bias in ordinary benign prompting, while malicious system prompts in targeted and untargeted tests often trigger refusals that make the measured stereotype score appear robust. Thus, refusal rates and aggregate trust scores must be inspected separately for each dimension.

  11. Knowl 11 — Operational safeguards for trustworthy compression

    limitation

    The study recommends starting from a dense model that is already trustworthy, because four- and eight-bit quantization generally preserve the source model's trust profile. Quantization should usually be preferred to pruning when the goal is to retain trustworthiness at a given efficiency level. Heavily compressed models must be evaluated with the full trustworthiness suite after calibration: random calibration data can produce more than 5 points of variance in fairness, ethics, and adversarial-demonstration scores even at four-bit GPTQ, and variance reaches about 15 points for three-bit GPTQ; AWQ is more stable but still variable at three bits. This uncertainty is not reliably predicted by MMLU. Finally, deployment should use a multi-dimensional trade-off rather than optimizing one score, because improving one trust dimension can coincide with deterioration in another, such as higher out-of-distribution robustness accompanied by worse stereotype performance.

Coverage note — Detailed per-task plots and individual sub-scenario definitions from all 33 DecodingTrust test cases were not separately expanded because their load-bearing findings are consolidated in the dimension-specific and refusal-aware knowls.

References

  1. 1.Ahmadian, A., Dash, S., Chen, H., Venkitesh, B., Gou, S., Blunsom, P., Üstün, A., and Hooker, S. Intriguing properties of quantization at scale. arXiv preprint arXiv:2305.19268, 2023.
  2. 2.Bartoldson, B. R., Kailkhura, B., and Blalock, D. Compute-efficient deep learning: Algorithmic trends and opportunities. Journal of Machine Learning Research, 24:1–77, 2023.
  3. 3.Boratko, M., Padigela, H., Mikkilineni, D., Yuvraj, P., Das, R., McCallum, A., Chang, M., Fokoue-Nkoutche, A., Kapanipathi, P., Mattei, N., et al. A systematic classification of knowledge, reasoning, and context within the arc dataset. arXiv preprint arXiv:1806.00358, 2018.
  4. 4.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  5. 5.Chee, J., Cai, Y., Kuleshov, V., and De Sa, C. Quip: 2-bit quantization of large language models with guarantees. arXiv preprint arXiv:2307.13304, 2023.
  6. 6.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  7. 7.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In NAACL, 2019.
  8. 8.Demszky, D., Yang, D., Yeager, D. S., Bryan, C. J., Clapper, M., Chandhok, S., Eichstaedt, J. C., Hecht, C., Jamieson, J., Johnson, M., et al. Using large language models in psychology. Nature Reviews Psychology, 2(11):688–701, 2023.
  9. 9.Dettmers, T. and Zettlemoyer, L. Sparse networks from scratch: Faster training without losing performance. arXiv preprint arXiv:1907.04840, 2019.
  10. 10.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  11. 11.Diffenderfer, J. and Kailkhura, B. Multi-prize lottery ticket hypothesis: Finding accurate binary neural networks by pruning a randomly weighted network. In International Conference on Learning Representations, 2020.
  12. 12.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  13. 13.Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning, pp. 2943–2952. PMLR, 2020.
  14. 14.Frantar, E. and Alistarh, D. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  15. 15.Frantar, E. and Alistarh, D. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pp. 10323–10337. PMLR, 2023.
  16. 16.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  17. 17.Gale, T., Elsen, E., and Hooker, S. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574, 2019.
  18. 18.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015.
  19. 19.Hasan, A., Rugina, I., and Wang, A. Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning. arXiv preprint arXiv:2401.10862, 2024.
  20. 20.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  21. 21.Huang, H., Zhao, Z., Backes, M., Shen, Y., and Zhang, Y. Composite backdoor attacks against large language models. arXiv preprint arXiv:2310.07676, 2023a.
  22. 22.Huang, Y., Zhang, Q., Sun, L., et al. Trustgpt: A benchmark for trustworthy and responsible large language models. arXiv preprint arXiv:2306.11507, 2023b.
  23. 23.Jaiswal, A., Gan, Z., Du, X., Zhang, B., Wang, Z., and Yang, Y. Compressing llms: The truth is rarely pure and never simple. arXiv preprint arXiv:2310.01382, 2023a.
  24. 24.Jaiswal, A., Liu, S., Chen, T., and Wang, Z. The emergence of essential sparsity in large pre-trained models: The weights that matter. arXiv preprint arXiv:2306.03805, 2023b.
  25. 25.Jaiswal, A. K., Ma, H., Chen, T., Ding, Y., and Wang, Z. Training your sparse neural network better with any mask. In International Conference on Machine Learning, pp. 9833–9844. PMLR, 2022.
  26. 26.Jaiswal, A. K., Liu, S., Chen, T., Ding, Y., and Wang, Z. Instant soup: Cheap pruning ensembles in a single pass can draw lottery tickets from large models. In International Conference on Machine Learning, pp. 14691–14701. PMLR, 2023c.
  27. 27.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  28. 28.Li, J., Cheng, X., Zhao, W. X., Nie, J., and rong Wen, J. Halueval: A large-scale hallucination evaluation benchmark for large language models. Conference on Empirical Methods in Natural Language Processing, 2023. doi: 10.48550/arXiv.2305.11747.
  29. 29.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  30. 30.Lin, T., Stich, S. U., Barba, L., Dmitriev, D., and Jaggi, M. Dynamic model pruning with feedback. In International Conference on Learning Representations, 2020.
  31. 31.Liu, S., Chen, T., Zhang, Z., Chen, X., Huang, T., Jaiswal, A., and Wang, Z. Sparsity may cry: Let us fail (current) sparse neural networks together! arXiv preprint arXiv:2303.02141, 2023a.
  32. 32.Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Cheng, R. G. H., Klochkov, Y., Taufiq, M. F., and Li, H. Trustworthy llms: a survey and guideline for evaluating large language models’ alignment. arXiv preprint arXiv:2308.05374, 2023b.
  33. 33.Ma, X., Fang, G., and Wang, X. Llm-pruner: On the structural pruning of large language models. arXiv preprint arXiv:2305.11627, 2023.
  34. 34.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP, 2018.
  35. 35.Mo, L., Wang, B., Chen, M., and Sun, H. How trustworthy are open-source llms? an assessment under malicious demonstrations shows their vulnerabilities. arXiv preprint arXiv:2311.09447, 2023.
  36. 36.Mostafa, H. and Wang, X. Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization. In International Conference on Machine Learning, 2019.
  37. 37.Namburi, S. S. S., Sreedhar, M., Srinivasan, S., and Sala, F. The cost of compression: Investigating the impact of compression on parametric knowledge in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 5255–5273, 2023.
  38. 38.Nvidia. Nvidia a100 tensor core gpu architecture. https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf, 2020.
  39. 39.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  40. 40.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R. The lambada dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  41. 41.Perez, E., Ringer, S., Lukošiut¯e, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. Discovering language model behaviors with model-written evaluations. arXiv preprint arXiv:2212.09251, 2022.
  42. 42.Qiu, H., Zhang, S., Li, A., He, H., and Lan, Z. Latent jailbreak: A benchmark for evaluating text safety and output robustness of large language models. arXiv preprint arXiv: 2307.08487, 2023.
  43. 43.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints, 2019.
  44. 44.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  45. 45.Singh, S. P. and Alistarh, D. Woodfisher: Efficient second-order approximation for neural network compression. Advances in Neural Information Processing Systems, 33:18098–18109, 2020.
  46. 46.Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024.
  47. 47.Sun, M., Liu, Z., Bair, A., and Kolter, J. Z. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023.
  48. 48.Tata, S. and Patel, J. M. Piqa: An algebra for querying protein data sets. In 15th International Conference on Scientific and Statistical Database Management, 2003., pp. 141–150. IEEE, 2003.
  49. 49.Timiryasov, I. and Tastet, J. Baby llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty. Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, 2023. doi: 10.48550/arXiv.2308.02019.
  50. 50.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  51. 51.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  52. 52.Tseng, A., Chee, J., Sun, Q., Kuleshov, V., and De Sa, C. Quip#: with lattice codebooks, 2023. URL https://cornell-relaxml.github.io/quip-sharp/. Accessed: 2024-01-24.
  53. 53.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In ICLR, 2019.
  54. 54.Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. arXiv preprint arXiv:2306.11698, 2023a.
  55. 55.Wang, S., Zhao, Z., Ouyang, X., Wang, Q., and Shen, D. Chatcad: Interactive computer-aided diagnosis on medical image using large language models. arXiv preprint arXiv:2302.07257, 2023b.
  56. 56.Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P. Assessing the brittleness of safety alignment via pruning and low-rank modifications. arXiv preprint arXiv:2402.05162, 2024.
  57. 57.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  58. 58.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023.
  59. 59.Xu, M., Xu, Y. L., and Mandic, D. P. Tensorgpt: Efficient compression of the embedding layer in llms based on the tensor-train decomposition. arXiv preprint arXiv:2307.00526, 2023.
  60. 60.Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S. Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity. arXiv preprint arXiv:2310.05175, 2023.
  61. 61.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  62. 62.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  63. 63.Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H. Learning n: m fine-grained structured sparse neural networks from scratch. arXiv preprint arXiv:2102.04010, 2021.
  64. 64.Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., Zhang, Y., et al. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv preprint arXiv:2306.04528, 2023.
  65. 65.Zhu, M. and Gupta, S. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878, 2017.

Citation

MLA
Hong, J., et al. “Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression”. arXiv, 2024, http://arxiv.org/abs/2403.15447v3.
APA
Hong, J., Duan, J., Zhang, C., Li, Z., Xie, C., Lieberman, K., Diffenderfer, J., Bartoldson, B., Jaiswal, A., Xu, K., Kailkhura, B., Hendrycks, D., Song, D., Wang, Z., & Li, B. (2024). Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression. arXiv. http://arxiv.org/abs/2403.15447v3
Chicago
Hong, J., J. Duan, C. Zhang, et al. 2024. “Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression”. arXiv. http://arxiv.org/abs/2403.15447v3.
Harvard
Hong, J. et al. (2024) “Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.15447v3.
Vancouver
1. Hong J, Duan J, Zhang C, et al (2024) Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression. arXiv

BibTeX

@article{hong2024decoding,
  title = {Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression},
  author = {Hong, Junyuan and Duan, Jinhao and Zhang, Chenhui and Li, Zhangheng and Xie, Chulin and Lieberman, Kelsey and Diffenderfer, James and Bartoldson, Brian and Jaiswal, Ajay and Xu, Kaidi and Kailkhura, Bhavya and Hendrycks, Dan and Song, Dawn and Wang, Zhangyang and Li, Bo},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.15447v3},
  eprint = {2403.15447}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/