SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Xin WangYu ZhengZhongwei WanMi Zhang

article2025ICLR302 citations

Develops SVD-LLM, a post-training compression framework that pairs truncation-aware data whitening with sequential low-rank parameter updates to preserve large language model accuracy at high compression ratios.

Listen

Large Language Models deliver impressive capabilities across natural language processing tasks, but their immense memory and computational footprints make deployment costly and challenging. While Singular Value Decomposition—a mathematical technique that approximates large weight matrices with smaller, low-rank matrices—provides an attractive pathway for post-training compression without specialized hardware constraints, existing approaches suffer severe accuracy drops at moderate to high compression ratios. These degradations stem from two core issues: truncating smaller singular values often fails to minimize overall compression loss, and previous methods do not update model weights after compression to recover lost performance.

The article introduces and evaluates SVD-LLM, a novel post-training compression framework designed to maintain model accuracy across wide compression ranges. The method pairs a truncation-aware data whitening process—which guarantees a direct mathematical alignment between truncated singular values and minimal compression loss—with a sequential parameter update strategy that fine-tunes the resulting low-rank matrices to restore generative fidelity.

To establish credibility and broad applicability, the authors evaluated SVD-LLM across seven distinct models spanning three prominent model families (LLaMA, OPT, and Mistral) at scales ranging from 7 billion to 30 billion parameters. Testing covered 10 standard evaluation benchmarks measuring language modeling perplexity, common sense reasoning, and complex generation tasks, alongside physical hardware benchmarks on both NVIDIA A100 graphics processing units and AMD EPYC central processing units.

The findings show that SVD-LLM significantly outperforms existing decomposition baselines. At high compression ratios of 40% and above, SVD-LLM achieves more than 400% higher reasoning accuracy and reduces language perplexity by over 99% compared to prior decomposition methods, which completely lost text generation capabilities. Furthermore, SVD-LLM surpasses state-of-the-art structured pruning methods—reducing perplexity by up to 56% under identical memory constraints—and outperforms 1-bit post-training quantization methods. When combined with 2-bit quantization, it matches or beats compute-heavy 1-bit methods that require complete model retraining. On hardware, it demonstrates near-linear memory reductions and substantial throughput gains on both central and graphics processors, while also compressing runtime key-value cache memory.

These results demonstrate that organizations can significantly lower inference costs, reduce memory requirements, and accelerate deployment timelines on standard hardware without undergoing resource-intensive model retraining. For decision-makers, SVD-LLM offers a practical deployment strategy, either as a standalone compression tool or combined with post-training quantization for extreme memory efficiency. Practitioners should consider running small-scale pilot deployments for targeted workloads, focusing on latency-memory trade-offs when enabling key-value cache compression.

Confidence in the reported outcomes is supported by extensive multi-model benchmarking and mathematical proofs. However, users should note specific operational limitations: accuracy still degrades at aggressive compression levels above 60% to 80%, runtime reconstruction of key-value cache states can introduce latency overheads, and compressed models occasionally produce repetitive text outputs. Future refinements will focus on reducing latency during memory recovery and further enhancing generative quality under extreme compression budgets.

No sufficiently relevant recommendations were found.

Abstract

The advancements in Large Language Models (LLMs) have been hindered by their substantial sizes, which necessitates LLM compression methods for practical deployment. Singular Value Decomposition (SVD) offers a promising solution for LLM compression. However, state-of-the-art SVD-based LLM compression methods have two key limitations: truncating smaller singular values may lead to higher compression loss, and the lack of update on the compressed weights after SVD truncation. In this work, we propose SVD-LLM, a SVD-based post-training LLM compression method that addresses the limitations of existing methods. SVD-LLM incorporates a truncation-aware data whitening technique to ensure a direct mapping between singular values and compression loss. Moreover, SVD-LLM adopts a parameter update with sequential low-rank approximation to compensate for the accuracy degradation after SVD compression. We evaluate SVD-LLM on 10 datasets and seven models from three different LLM families at three different scales. Our results demonstrate the superiority of SVD-LLM over state-of-the-arts, especially at high model compression ratios. Our code is available at this https URL

Citation

MLA
Wang, X., et al. “SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression”. arXiv, 2024, http://arxiv.org/abs/2403.07378v5.
APA
Wang, X., Zheng, Y., Wan, Z., & Zhang, M. (2024). SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression. arXiv. http://arxiv.org/abs/2403.07378v5
Chicago
Wang, X., Y. Zheng, Z. Wan, and M. Zhang. 2024. “SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression”. arXiv. http://arxiv.org/abs/2403.07378v5.
Harvard
Wang, X. et al. (2024) “SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.07378v5.
Vancouver
1. Wang X, Zheng Y, Wan Z, Zhang M (2024) SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression. arXiv

BibTeX

@article{wang2024svd,
  title = {SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression},
  author = {Wang, Xin and Zheng, Yu and Wan, Zhongwei and Zhang, Mi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.07378v5},
  eprint = {2403.07378}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors