SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
Xin WangYu ZhengZhongwei WanMi Zhang
Develops SVD-LLM, a post-training compression framework that pairs truncation-aware data whitening with sequential low-rank parameter updates to preserve large language model accuracy at high compression ratios.
Large Language Models deliver impressive capabilities across natural language processing tasks, but their immense memory and computational footprints make deployment costly and challenging. While Singular Value Decomposition—a mathematical technique that approximates large weight matrices with smaller, low-rank matrices—provides an attractive pathway for post-training compression without specialized hardware constraints, existing approaches suffer severe accuracy drops at moderate to high compression ratios. These degradations stem from two core issues: truncating smaller singular values often fails to minimize overall compression loss, and previous methods do not update model weights after compression to recover lost performance.
The article introduces and evaluates SVD-LLM, a novel post-training compression framework designed to maintain model accuracy across wide compression ranges. The method pairs a truncation-aware data whitening process—which guarantees a direct mathematical alignment between truncated singular values and minimal compression loss—with a sequential parameter update strategy that fine-tunes the resulting low-rank matrices to restore generative fidelity.
To establish credibility and broad applicability, the authors evaluated SVD-LLM across seven distinct models spanning three prominent model families (LLaMA, OPT, and Mistral) at scales ranging from 7 billion to 30 billion parameters. Testing covered 10 standard evaluation benchmarks measuring language modeling perplexity, common sense reasoning, and complex generation tasks, alongside physical hardware benchmarks on both NVIDIA A100 graphics processing units and AMD EPYC central processing units.
The findings show that SVD-LLM significantly outperforms existing decomposition baselines. At high compression ratios of 40% and above, SVD-LLM achieves more than 400% higher reasoning accuracy and reduces language perplexity by over 99% compared to prior decomposition methods, which completely lost text generation capabilities. Furthermore, SVD-LLM surpasses state-of-the-art structured pruning methods—reducing perplexity by up to 56% under identical memory constraints—and outperforms 1-bit post-training quantization methods. When combined with 2-bit quantization, it matches or beats compute-heavy 1-bit methods that require complete model retraining. On hardware, it demonstrates near-linear memory reductions and substantial throughput gains on both central and graphics processors, while also compressing runtime key-value cache memory.
These results demonstrate that organizations can significantly lower inference costs, reduce memory requirements, and accelerate deployment timelines on standard hardware without undergoing resource-intensive model retraining. For decision-makers, SVD-LLM offers a practical deployment strategy, either as a standalone compression tool or combined with post-training quantization for extreme memory efficiency. Practitioners should consider running small-scale pilot deployments for targeted workloads, focusing on latency-memory trade-offs when enabling key-value cache compression.
Confidence in the reported outcomes is supported by extensive multi-model benchmarking and mathematical proofs. However, users should note specific operational limitations: accuracy still degrades at aggressive compression levels above 60% to 80%, runtime reconstruction of key-value cache states can introduce latency overheads, and compressed models occasionally produce repetitive text outputs. Future refinements will focus on reducing latency during memory recovery and further enhancing generative quality under extreme compression budgets.
- Paper: LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation, Yixiao Li et al. (2023). LoSparse shows how SVD-based low-rank factors can compress language-model weights, giving useful context for SVD-LLM’s low-rank compression approach.
No sufficiently relevant recommendations were found.