LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation
Yixiao LiYifan YuQingru ZhangChen LiangPengcheng HeWeizhu ChenTuo Zhao
Proposes a structured compression framework that decomposes Transformer weight matrices into low-rank and sparse components, preventing the expressive capacity loss of conventional pruning while outperforming standard compression baselines across understanding and generation tasks.
Large language models deliver outstanding results across language tasks, but their immense size creates severe memory and computational bottlenecks during practical deployment. Standard compression techniques face notable trade-offs: pruning removes entire neurons and risks discarding critical expressive information at high compression rates, whereas low-rank approximation captures only shared coherent features and loses the diverse, distinct behaviors of individual neurons. To address these limitations, the article introduces LoSparse, a structured compression method that represents model weight matrices as the sum of a low-rank matrix and a sparse matrix to compress coherent shared components while retaining expressive neuron diversity.
The article evaluates LoSparse against leading structured pruning and compression approaches across language understanding, question answering, and text generation benchmarks. Experiments were conducted using standard transformer architectures—DeBERTa, BERT, and BART—across multiple datasets (GLUE, SQuAD, XSum, and CNN/DailyMail) under various parameter retention ratios ranging from 50% down to 5%. In high-level terms, the process initializes the low-rank component using singular value decomposition to capture shared structure, assigns the remainder to a sparse component, and then iteratively prunes non-expressive parts of the sparse matrix during fine-tuning.
The findings show that LoSparse consistently outperforms existing pruning methods, especially at high compression levels. When retaining only 10% of model parameters on natural language understanding benchmarks, LoSparse achieves up to 2.0% higher accuracy than baseline iterative pruning. In question-answering tasks at an extreme 5% parameter retention level, LoSparse surpasses standard pruning by 3.0 points in F1 score. On abstractive text generation, LoSparse outperforms baseline pruning by nearly 3.0 points in ROUGE-1 score at a 30% retention ratio. In addition, LoSparse avoids the training divergence observed in baseline methods and integrates effectively with knowledge distillation and other multi-level pruning frameworks to achieve further performance gains.
These results demonstrate that combining low-rank and sparse approximations allows organizations to compress large language models aggressively without incurring catastrophic accuracy loss. This capability significantly reduces hardware and operational costs for serving models in resource-constrained production environments. Technical leaders should consider adopting composite low-rank and sparse structured compression for model deployment pipelines and evaluating its integration with existing knowledge distillation workflows.
Decision-makers should note that the evaluation in the article is focused primarily on base and large model sizes (up to hundreds of millions of parameters) and specific natural language processing benchmarks. Further validation on modern multi-billion-parameter generative models and domain-specific production workloads is recommended prior to full-scale enterprise rollout.
- Paper: LoRA: Low-Rank Adaptation of Large Language Models, Edward J. Hu et al. (2022). This work establishes the low-rank parameter-efficient adaptation paradigm that LoSparse directly integrates and builds upon for structured matrix approximation.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). It introduces scalable, post-training sparse pruning techniques for large language models, providing the foundation for the pruning component in LoSparse.
- Paper: ZipLM: Inference-Aware Structured Pruning of Language Models, Eldar Kurtic et al. (2023). This paper establishes structured pruning frameworks for transformers, offering essential context on structured parameter removal.
- Paper: AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning, Qingru Zhang et al. (2023). It introduces adaptive singular value pruning and low-rank allocation, demonstrating how low-rank decomposition interacts with parameter pruning.
- Paper: Compressing Transformers: Features Are Low-Rank, but Weights Are Not!, Hao Yu et al. (2023). It analyzes the low-rank characteristics of internal transformer representations versus weights, providing foundational motivation for low-rank model compression.
- Paper: Learning Structured Sparsity in Deep Neural Networks, Wei Wen et al. (2016). This paper establishes the foundational principles of structured sparsity and group regularization in neural networks.
- Paper: Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding, Song Han et al. (2015). It provides seminal groundwork on neural network compression pipelines utilizing weight pruning and sparse representations.
- Paper: RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation, Mahdi Nikdan et al. (2024). RoSA extends the principle of joint low-rank and sparse decomposition specifically to robust parameter-efficient fine-tuning for generative LLMs.
- Paper: LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning, Han Guo et al. (2024). LQ-LoRA explores low-rank matrix decomposition combined with quantization, extending structured decomposition techniques beyond sparse approximations.
- Paper: SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression, Xin Wang 0120 et al. (2025). SVD-LLM advances low-rank matrix decomposition for LLMs by designing truncation-aware SVD and weight updates for post-training compression.
- Paper: SliceGPT: Compress Large Language Models by Deleting Rows and Columns, Saleh Ashkboos et al. (2024). SliceGPT develops structured post-training compression by systematically deleting rows and columns, offering an alternative structured pruning formulation.
- Paper: DoRA: Weight-Decomposed Low-Rank Adaptation, Shih-Yang Liu et al. (2024). DoRA builds upon low-rank parameter adaptation by decoupling weight updates into directional low-rank and magnitude components.
