LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation

Yixiao LiYifan YuQingru ZhangChen LiangPengcheng HeWeizhu ChenTuo Zhao

article2023ICML120 citations

Proposes a structured compression framework that decomposes Transformer weight matrices into low-rank and sparse components, preventing the expressive capacity loss of conventional pruning while outperforming standard compression baselines across understanding and generation tasks.

Listen

Large language models deliver outstanding results across language tasks, but their immense size creates severe memory and computational bottlenecks during practical deployment. Standard compression techniques face notable trade-offs: pruning removes entire neurons and risks discarding critical expressive information at high compression rates, whereas low-rank approximation captures only shared coherent features and loses the diverse, distinct behaviors of individual neurons. To address these limitations, the article introduces LoSparse, a structured compression method that represents model weight matrices as the sum of a low-rank matrix and a sparse matrix to compress coherent shared components while retaining expressive neuron diversity.

The article evaluates LoSparse against leading structured pruning and compression approaches across language understanding, question answering, and text generation benchmarks. Experiments were conducted using standard transformer architectures—DeBERTa, BERT, and BART—across multiple datasets (GLUE, SQuAD, XSum, and CNN/DailyMail) under various parameter retention ratios ranging from 50% down to 5%. In high-level terms, the process initializes the low-rank component using singular value decomposition to capture shared structure, assigns the remainder to a sparse component, and then iteratively prunes non-expressive parts of the sparse matrix during fine-tuning.

The findings show that LoSparse consistently outperforms existing pruning methods, especially at high compression levels. When retaining only 10% of model parameters on natural language understanding benchmarks, LoSparse achieves up to 2.0% higher accuracy than baseline iterative pruning. In question-answering tasks at an extreme 5% parameter retention level, LoSparse surpasses standard pruning by 3.0 points in F1 score. On abstractive text generation, LoSparse outperforms baseline pruning by nearly 3.0 points in ROUGE-1 score at a 30% retention ratio. In addition, LoSparse avoids the training divergence observed in baseline methods and integrates effectively with knowledge distillation and other multi-level pruning frameworks to achieve further performance gains.

These results demonstrate that combining low-rank and sparse approximations allows organizations to compress large language models aggressively without incurring catastrophic accuracy loss. This capability significantly reduces hardware and operational costs for serving models in resource-constrained production environments. Technical leaders should consider adopting composite low-rank and sparse structured compression for model deployment pipelines and evaluating its integration with existing knowledge distillation workflows.

Decision-makers should note that the evaluation in the article is focused primarily on base and large model sizes (up to hundreds of millions of parameters) and specific natural language processing benchmarks. Further validation on modern multi-billion-parameter generative models and domain-specific production workloads is recommended prior to full-scale enterprise rollout.

arXiv: 2306.11222
Cover for LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation

Abstract

Transformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To reduce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse approximation), a novel model compression technique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low-rank approximations and pruning, while avoiding their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approximations, and low-rank approximation prevents pruning from losing too many expressive neurons. We evaluate our method on natural language understanding, question answering, and natural language generation tasks. We show that it significantly outperforms existing compression methods. Our code is publicly available at https://github.com/yxli2123/LoSparse

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Transformer Models
  • 2.2. Importance Scores for Pruning
  • 2.3. Structured Pruning
  • 3. Method
  • 3.1. Approximation by Low-rank and Sparse Matrices
  • 3.2. Algorithm
  • 4. Experiments
  • 4.1. Natural Language Understanding
  • 4.2. Question Answering
  • 4.3. Natural Language Generation
  • 4.4. Analysis
  • 4.5. Combination with Knowledge Distillation
  • 4.6. Embed with Other Compression Method
  • 6. Conclusion
  • 5. Discussion
  • References
  • A. GLUE Dataset Statistics
  • B. Natural Language Understanding
  • B.1. Training Details
  • Hyper-parameter Details
  • C. Question Answering
  • C.1. Dataset
  • C.2. Training Details
  • D. Natural Language Generation
  • D.1. Training Details
  • E. Combination with Knowledge Distillation
  • E.1. Teacher Models
  • E.2. Training Details
  • F. Combination with CoFi
  • F.1. Teacher Models
  • F.2. Training Details

Knowls

  1. Knowl 1 — Low-rank plus structured-sparse representation

    model/method

    LoSparse represents each transformer weight matrix with a low-rank component and a structured sparse residual. For a weight matrix W∈Rd1×d2W\in\mathbb{R}^{d_1\times d_2}, the compressed representation is

    W^=UV+S,\widehat W=UV+S,

    where U∈Rd1×rU\in\mathbb{R}^{d_1\times r}, V∈Rr×d2V\in\mathbb{R}^{r\times d_2}, and rr is the chosen low rank. The residual S∈Rd1×d2S\in\mathbb{R}^{d_1\times d_2} is pruned column-wise, so removing a column corresponds to removing a neuron rather than individual unstructured entries. For an input matrix XX, the forward computation is implemented as Y=(XU)V+XSY=(XU)V+XS, allowing the low-rank and sparse terms to be evaluated in parallel.

    The low-rank term is intended to preserve expressive, shared structure across neurons, while the sparse term preserves diverse, neuron-specific structure and allows non-expressive residual neurons to be removed. The initial factorization reconstructs the pretrained matrix exactly, but subsequent training learns the factors and pruned residual as a compressed approximation.

  2. Knowl 2 — LoSparse initialization, importance scoring, and pruning algorithm

    algorithm

    LoSparse is applied to each selected pretrained weight matrix W(0)W^{(0)} using a target rank rr, total training steps TT, learning rate α\alpha, initial warm-up duration tit_i, final fine-tuning duration tft_f, final retained-neuron fraction pTp_T, and exponential-moving-average coefficient β\beta. In the reported experiments, all backbone weight matrices were compressed except LayerNorm parameters and the final prediction head.

    For each matrix, compute its singular value decomposition and let σ1≥⋯≥σr\sigma_1\geq\cdots\geq\sigma_r be the top rr singular values, with corresponding left singular vectors uk∈Rd1u_k\in\mathbb{R}^{d_1} and right singular vectors vk∈Rd2v_k\in\mathbb{R}^{d_2}. Initialize

    U(0)=[σ1u1,…,σrur],V(0)=[σ1v1,…,σrvr]⊤,U^{(0)}=[\sqrt{\sigma_1}u_1,\ldots,\sqrt{\sigma_r}u_r],\qquad V^{(0)}=[\sqrt{\sigma_1}v_1,\ldots,\sqrt{\sigma_r}v_r]^{\top},

    and set S(0)=W(0)−U(0)V(0)S^{(0)}=W^{(0)}-U^{(0)}V^{(0)}.

    At training step tt, update U(t)U^{(t)}, V(t)V^{(t)}, and a temporary sparse matrix S~(t)=S(t)−α∇S(t)L\widetilde S^{(t)}=S^{(t)}-\alpha\nabla_{S^{(t)}}L, where LL is the task loss. For a sparse entry sij(t)s_{ij}^{(t)}, compute the sensitivity score Iij(t)=∣sij(t) ∂L(t)/∂sij(t)∣I_{ij}^{(t)}=\left|s_{ij}^{(t)}\,\partial L^{(t)}/\partial s_{ij}^{(t)}\right| and smooth it as I‾ij(t)=βI‾ij(t−1)+(1−β)Iij(t)\overline I_{ij}^{(t)}=\beta\overline I_{ij}^{(t-1)}+(1-\beta)I_{ij}^{(t)}. The importance of column ii of S(t)S^{(t)} is the mean smoothed entry score, Γ(S∗i(t))=d1−1∑j=1d1I‾ji(t)\Gamma(S_{*i}^{(t)})=d_1^{-1}\sum_{j=1}^{d_1}\overline I_{ji}^{(t)}.

    Only the columns whose scores are among the top ptp_t fraction are retained: S∗i(t+1)=S~∗i(t)S_{*i}^{(t+1)}=\widetilde S_{*i}^{(t)} for retained columns and S∗i(t+1)=0S_{*i}^{(t+1)}=0 otherwise. The retained fraction follows a cubic schedule,

    pt={1,0≤t<ti,pT+(1−pT)(1−t−tiT−ti−tf)3,ti≤t<T−tf,pT,T−tf≤t≤T.p_t=\begin{cases}1,&0\leq t<t_i,\\ p_T+(1-p_T)\left(1-\dfrac{t-t_i}{T-t_i-t_f}\right)^3,&t_i\leq t<T-t_f,\\ p_T,&T-t_f\leq t\leq T.\end{cases}

    Thus, the sparse component is left intact during warm-up, gradually pruned during the middle of training, and fine-tuned at the final target sparsity while the low-rank factors continue to receive gradient updates throughout.

  3. Knowl 3 — Mechanism for separating shared and diverse neuron structure

    model/method

    The method is motivated by two empirical properties of transformer weight matrices. The singular-value plots on page 3 show a rapid initial spectral decay in BART-large decoder weights and DeBERTaV3-large encoder weights, indicating that many neurons share a common subspace that can be represented efficiently by a low-rank factorization. The same spectra retain longer, heavy-tailed components, which the paper interprets as neuron-specific or incoherent structure that a low-rank approximation would discard.

    LoSparse therefore uses UVUV to capture the shared coherent subspace and applies structured pruning only to the residual SS. The importance-score histograms on page 4, obtained for selected DeBERTaV3-base query and key projections on SST-2, show that this decoupling produces many low-importance residual neurons under LoSparse compared with direct iterative pruning. The paper argues that this makes pruning safer: the low-rank component protects expressive shared information, while pruning removes mainly non-expressive information in the residual.

  4. Knowl 4 — Natural-language-understanding compression results

    empirical result

    LoSparse was evaluated on GLUE by compressing DeBERTaV3-base at 20%, 15%, and 10% remaining-weight ratios, excluding WNLI, and by compressing BERT-base on MNLI, RTE, and QNLI. Movement pruning and iterative pruning (ITP) were the baselines; N.A. denotes non-convergence. All backbone matrices except LayerNorm and the final prediction head were compressed.

    For DeBERTaV3-base, LoSparse matched or exceeded the baselines on every reported GLUE dataset and every remaining ratio. At 10% remaining weights, its scores were MNLI-m/mm 81.7/81.881.7/81.8 versus ITP 79.7/79.679.7/79.6, RTE 66.066.0 versus ITP non-convergence, QNLI 86.186.1 versus 82.382.3, MRPC accuracy/F1 82.3/87.482.3/87.4 versus 78.5/84.378.5/84.3, QQP accuracy/F1 89.5/86.089.5/86.0 versus 88.3/84.488.3/84.4, SST-2 89.289.2 versus 88.388.3, CoLA Matthews correlation 40.040.0 versus 38.038.0, and STS-B Pearson/Spearman correlation 87.2/87.087.2/87.0 versus 86.3/86.086.3/86.0.

    On BERT-base at 10% remaining weights, LoSparse achieved MNLI-m/mm 78.3/77.878.3/77.8, RTE 63.063.0, and QNLI 84.884.8, compared with ITP scores of 77.7/78.377.7/78.3, 61.861.8, and 83.983.9. At the 20% ratio, LoSparse likewise exceeded ITP on MNLI-m/mm (80.4/80.380.4/80.3 versus 80.1/79.880.1/79.8), RTE (65.265.2 versus 64.464.4), and QNLI (86.986.9 versus 86.586.5). The paper attributes the improved stability under extreme sparsity to retaining and continuously updating a low-rank component in every matrix rather than allowing whole matrices to fluctuate between nonzero and zero states.

  5. Knowl 5 — Question-answering results under extreme sparsity

    empirical result

    On SQuAD v1.1, LoSparse compressed DeBERTaV3-base and BERT-base and was compared with ITP; Movement pruning was also reported for BERT-base. The evaluation metrics were exact match (EM) and token-level F1, and the remaining-weight ratios were 5%, 10%, 20%, 30%, 40%, and 50%.

    For DeBERTaV3-base, the pretrained scores were 87.7/93.587.7/93.5 for EM/F1. LoSparse obtained 69.3/79.169.3/79.1, 72.9/82.872.9/82.8, 76.8/85.876.8/85.8, 80.2/88.080.2/88.0, 82.1/89.482.1/89.4, and 82.3/90.382.3/90.3 at the six ratios respectively, while ITP obtained 65.2/76.165.2/76.1, 70.9/80.370.9/80.3, 75.0/83.975.0/83.9, 78.2/86.278.2/86.2, 80.5/87.580.5/87.5, and 81.5/89.681.5/89.6. Thus, LoSparse was better at every reported ratio, including a 3.0-point F1 advantage at 5% remaining weights and a 1.9-point advantage at 40%.

    For BERT-base, LoSparse produced 57.6/70.657.6/70.6, 65.2/76.865.2/76.8, 69.7/80.469.7/80.4, 73.0/82.973.0/82.9, 74.6/84.274.6/84.2, and 75.8/85.175.8/85.1 across the same ratios. ITP produced 54.0/67.354.0/67.3, 62.5/74.262.5/74.2, 66.8/78.066.8/78.0, 72.3/82.472.3/82.4, 74.5/84.274.5/84.2, and 76.0/85.176.0/85.1; Movement pruning was non-convergent at 5% and produced 51.4/64.651.4/64.6, 63.3/74.563.3/74.5, 68.8/79.068.8/79.0, 73.0/82.473.0/82.4, and 76.2/84.176.2/84.1 at the remaining ratios from 10% through 50%. LoSparse was especially advantageous at low remaining ratios, although ITP slightly exceeded it in BERT-base EM at 50%.

  6. Knowl 6 — Summarization results with BART-large

    empirical result

    LoSparse was evaluated by compressing all encoder and decoder weight matrices of BART-large on XSum and CNN/DailyMail. Training used a batch size of 32, 12 epochs, beam-search length 8, and ROUGE-1/2/L evaluation. Results were reported at 50%, 40%, and 30% remaining weights because the baseline methods did not exceed the Lead-3 baseline at lower ratios.

    On XSum, ITP versus LoSparse produced ROUGE-1/2/L scores of 38.42/16.32/31.4338.42/16.32/31.43 versus 39.18/16.91/31.6239.18/16.91/31.62 at 50%, 36.71/14.96/29.8636.71/14.96/29.86 versus 38.30/16.02/30.7238.30/16.02/30.72 at 40%, and 34.42/13.15/27.9934.42/13.15/27.99 versus 37.41/15.42/30.0237.41/15.42/30.02 at 30%. On CNN/DailyMail, the corresponding scores were 40.76/18.30/37.6540.76/18.30/37.65 versus 41.54/19.04/38.5841.54/19.04/38.58, 40.52/18.10/37.3140.52/18.10/37.31 versus 41.42/19.00/38.4741.42/19.00/38.47, and 40.35/17.98/37.1540.35/17.98/37.15 versus 41.21/18.84/38.2141.21/18.84/38.21.

    LoSparse therefore outperformed ITP at every tested ratio and metric. Its ROUGE-1 advantage on XSum grew from 0.760.76 points at 50% remaining weights to 2.992.99 points at 30%, whereas the corresponding CNN/DailyMail advantage at 30% was 0.860.86 points. The larger gain on XSum was reported as evidence that the method is particularly useful for the more abstractive summarization task.

  7. Knowl 7 — Ablation of sparse approximation and allocation

    empirical result

    The paper tested whether the sparse residual is essential by comparing LoSparse with two low-rank-only variants on SQuAD v1.1, MNLI-m, and XSum. Low-rank I discarded SS and fine-tuned only the SVD-initialized product UVUV. Low-rank II used the full initialization W(0)=U(0)V(0)+S(0)W^{(0)}=U^{(0)}V^{(0)}+S^{(0)} but gradually pruned the initialized residual SS completely to zero. LoSparse instead retained and trained a structured-pruned residual throughout.

    Across the remaining ratios plotted on page 8, LoSparse outperformed both variants on all three tasks. Low-rank II was consistently better than Low-rank I, showing that retaining the sparse residual during the transition from the pretrained matrix to the low-rank representation is more effective than immediately discarding it. The paper explains this as preservation of pretrained knowledge that is lost when the model is initialized only from the truncated SVD.

    A separate allocation study varied the proportion of all pretrained parameters assigned to low-rank factors among 1%, 2%, 3%, and 5%, while adjusting the sparse allocation to maintain a fixed total remaining ratio. On SST-2, SQuAD v1.1, and QNLI, performance remained comparatively stable across these allocations at the tested total ratios of 10%–50%. This indicates that the low-rank and sparse components make roughly comparable contributions and that LoSparse is not highly sensitive to the precise allocation.

  8. Knowl 8 — Compatibility with knowledge distillation

    empirical result

    LoSparse was combined with layer-wise knowledge distillation by using a fine-tuned model as teacher and a compressed model as student. In the DeBERTaV3-base experiment, the student was compressed to 20% of the original size before distillation. The teacher achieved MNLI 90.590.5, SQuAD EM/F1 87.6/93.587.6/93.5, SST-2 96.196.1, and RTE 82.082.0. ITP plus distillation achieved 85.985.9, 77.9/87.277.9/87.2, 92.092.0, and 58.158.1, whereas LoSparse plus distillation achieved MNLI-m/mm 84.6/84.784.6/84.7, SQuAD 81.0/89.881.0/89.8, SST-2 93.293.2, and RTE 71.171.1.

    On BERT-base, LoSparse plus distillation was compared with several published distillation methods at 50% and 25% remaining weights. Its MNLI-m/mm, RTE, QNLI, and SST-2 scores were 85.1/85.385.1/85.3, 75.875.8, 92.292.2, and 93.293.2 at 50%, and 84.6/84.784.6/84.7, 72.272.2, 91.491.4, and 92.392.3 at 25%. These results were on par with or better than most listed distillation methods, including low-rank factorization combined with distillation. The experiments support the paper's claim that LoSparse compression and knowledge distillation are complementary rather than competing procedures.

  9. Knowl 9 — Embedding LoSparse into CoFi

    model/method

    The paper embedded LoSparse into CoFi, a coarse-to-fine compression method that masks layers and attention/feed-forward sublayers, then attention heads, and finally neurons. LoSparse replaced CoFi's final neuron-level masks while preserving CoFi's layer- and head-level masks. The pretrained matrices were first decomposed into low-rank and sparse components, after which the combined model was trained using the CoFi procedure and its distillation setup.

    Because CoFi controls a fraction rCoFir_{\mathrm{CoFi}} of the parameters and LoSparse controls a fraction rLoSparser_{\mathrm{LoSparse}} within the retained structure, the total remaining ratio is rCoFirLoSparser_{\mathrm{CoFi}}r_{\mathrm{LoSparse}}. For the reported 10% model, the authors used rCoFi=0.2r_{\mathrm{CoFi}}=0.2 and rLoSparse=0.5r_{\mathrm{LoSparse}}=0.5.

    On BERT-base, the uncompressed teacher scored MNLI 84.4184.41, MRPC accuracy/F1 87.74/91.3587.74/91.35, RTE 72.5672.56, QNLI 91.5491.54, and SST-2 92.4392.43. CoFi alone scored 80.0080.00, 84.07/88.5084.07/88.50, 67.5167.51, 86.6786.67, and 90.6090.60, while CoFi+LoSparse scored 82.5682.56, 85.54/89.4585.54/89.45, 68.2368.23, 89.6689.66, and 91.5191.51, respectively. Thus, replacing CoFi's final masks with LoSparse improved every reported task, including approximately one point of accuracy on MRPC, RTE, and SST-2.

  10. Knowl 10 — Stated limitation for convolutional-network compression

    limitation

    The paper limits its strongest claims to transformer-style models and explains why directly transferring prior low-rank-plus-sparse CNN constructions may be ineffective. In those CNN methods, sparse convolutional kernels and added low-rank convolutional layers do not directly form an approximation of the same weight matrix, so the two components are not explicitly coupled as UV+SUV+S. In addition, CNN kernels have fewer dimensions and therefore intrinsically lower rank than the large transformer matrices targeted by LoSparse. Their low-rank component consequently has less capacity to provide useful compression at very high compression rates.

Coverage note — No substantial contributed material was omitted; appendix-level dataset statistics, task-specific hyperparameter tables, and routine implementation details were excluded because the load-bearing method, results, analyses, integrations, and stated limitation are represented above.

References

  1. 1.Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I. The second pascal recognising textual entailment challenge. 2006.
  2. 2.Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D. The fifth pascal recognizing textual entailment challenge. In TAC, 2009.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  4. 4.Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pp. 1–14, Vancouver, Canada, August 2017. Association for Computational Linguistics. doi: 10.18653/v1/S17-2001.
  5. 5.Dagan, I., Glickman, O., and Magnini, B. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop, 2007.
  6. 6.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  7. 7.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005.
  8. 8.Fan, A., Grave, E., and Joulin, A. Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556, 2019.
  9. 9.Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. The third PASCAL recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pp. 1–9, Prague, June 2007. Association for Computational Linguistics.
  10. 10.Hajimolahoseini, H., Rezagholizadeh, M., Partovinia, V., Tahaei, M., Awad, O. M., and Liu, Y. Compressing pre-trained language models using progressive low rank decomposition.
  11. 11.Han, S., Pool, J., Tran, J., and Dally, W. Learning both weights and connections for efficient neural network. Advances in neural information processing systems, 28, 2015.
  12. 12.Hawkins, C., Yang, H., Li, M., Lai, L., and Chandra, V. Low-rank+ sparse tensor compression for neural networks. arXiv preprint arXiv:2111.01697, 2021.
  13. 13.He, P., Liu, X., Gao, J., and Chen, W. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020.
  14. 14.He, P., Gao, J., and Chen, W. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021.
  15. 15.Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P. Teaching machines to read and comprehend. Advances in neural information processing systems, 28, 2015.
  16. 16.Hinton, G., Vinyals, O., Dean, J., et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  17. 17.Hsu, Y.-C., Hua, T., Chang, S., Lou, Q., Shen, Y., and Jin, H. Language model compression with weighted low-rank factorization. arXiv preprint arXiv:2207.00112, 2022a.
  18. 18.Hsu, Y.-C., Hua, T., Chang, S.-E., Lou, Q., Shen, Y., and Jin, H. Language model compression with weighted low-rank factorization. ArXiv, abs/2207.00112, 2022b.
  19. 19.Jalali, A., Sanghavi, S., Ruan, C., and Ravikumar, P. A dirty model for multi-task learning. Advances in neural information processing systems, 23, 2010.
  20. 20.Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M. Block pruning for faster transformers. arXiv preprint arXiv:2109.04838, 2021.
  21. 21.Levesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning, 2012.
  22. 22.Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871–7880, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.703.
  23. 23.Liang, C., Zuo, S., Chen, M., Jiang, H., Liu, X., He, P., Zhao, T., and Chen, W. Super tickets in pre-trained language models: From model compression to improving generalization. arXiv preprint arXiv:2105.12002, 2021.
  24. 24.Liang, K. J., Hao, W., Shen, D., Zhou, Y., Chen, W., Chen, C., and Carin, L. Mixkd: Towards efficient distillation of large-scale language models. ArXiv, abs/2011.00593, 2020.
  25. 25.Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics.
  26. 26.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  27. 27.Louizos, C., Welling, M., and Kingma, D. P. Learning sparse neural networks through l 0 regularization. arXiv preprint arXiv:1712.01312, 2017.
  28. 28.McCarley, J., Chakravarti, R., and Sil, A. Structured pruning of a bert-based question answering model. arXiv preprint arXiv:1910.06360, 2019.
  29. 29.Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J. Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440, 2016.
  30. 30.Molchanov, P., Mallya, A., Tyree, S., Frosio, I., and Kautz, J. Importance estimation for neural network pruning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11264–11272, 2019.
  31. 31.Morris, J., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., and Qi, Y. Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 119–126, 2020.
  32. 32.Narayan, S., Cohen, S. B., and Lapata, M. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. ArXiv, abs/1808.08745, 2018.
  33. 33.Noach, M. B. and Goldberg, Y. Compressing pre-trained language models by matrix decomposition. In AACL, 2020.
  34. 34.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  35. 35.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  36. 36.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392, Austin, Texas, November 2016a. Association for Computational Linguistics. doi: 10.18653/v1/D16-1264.
  37. 37.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016b.
  38. 38.Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014.
  39. 39.Sanh, V., Wolf, T., and Rush, A. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389, 2020.
  40. 40.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1631–1642, Seattle, Washington, USA, October 2013. Association for Computational Linguistics.
  41. 41.Sun, S., Cheng, Y., Gan, Z., and Liu, J. Patient knowledge distillation for bert model compression. In Conference on Empirical Methods in Natural Language Processing, 2019.
  42. 42.Sun, S., Gan, Z., Cheng, Y., Fang, Y., Wang, S., and Liu, J. Contrastive distillation on intermediate representations for language model compression. In Conference on Empirical Methods in Natural Language Processing, 2020.
  43. 43.Tahaei, M. S., Charlaix, E., Nia, V. P., Ghodsi, A., and Rezagholizadeh, M. Kroneckerbert: Learning kronecker decomposition for pre-trained language models via knowledge distillation. arXiv preprint arXiv:2109.06243, 2021.
  44. 44.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019.
  45. 45.Warstadt, A., Singh, A., and Bowman, S. R. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019. doi: 10.1162/tacl_a_00290.
  46. 46.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1112–1122, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1101.
  47. 47.Xia, M., Zhong, Z., and Chen, D. Structured pruning learns compact and accurate models. arXiv preprint arXiv:2204.00408, 2022.
  48. 48.Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M. Bert-of-theseus: Compressing bert by progressive module replacing. In Conference on Empirical Methods in Natural Language Processing, 2020.
  49. 49.Yu, X., Liu, T., Wang, X., and Tao, D. On compressing deep models by low rank and sparse decomposition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7370–7379, 2017.
  50. 50.Zhang, Q., Zuo, S., Liang, C., Bukharin, A., He, P., Chen, W., and Zhao, T. Platon: Pruning large transformer models with upper confidence bound of weight importance. In International Conference on Machine Learning, pp. 26809–26823. PMLR, 2022.
  51. 51.Zhu, M. and Gupta, S. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878, 2017.

Citation

MLA
Li, Y., et al. “LoSparse: Structured Compression of Large Language Models Based on Low-Rank and Sparse Approximation”. International Conference on Machine Learning, vol. 202, 2023, pp. 20336–50, https://proceedings.mlr.press/v202/li23ap.html.
APA
Li, Y., Yu, Y., Zhang, Q., Liang, C., He, P., Chen, W., & Zhao, T. (2023). LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation. International Conference on Machine Learning, 202, 20336–20350. https://proceedings.mlr.press/v202/li23ap.html
Chicago
Li, Y., Y. Yu, Q. Zhang, et al. 2023. “LoSparse: Structured Compression of Large Language Models Based on Low-Rank and Sparse Approximation”. International Conference on Machine Learning 202: 20336–50. https://proceedings.mlr.press/v202/li23ap.html.
Harvard
Li, Y. et al. (2023) “LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation”, International Conference on Machine Learning. PMLR, pp. 20336–20350. Available at: https://proceedings.mlr.press/v202/li23ap.html.
Vancouver
1. Li Y, Yu Y, Zhang Q, Liang C, He P, Chen W, Zhao T (2023) LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation. In: International Conference on Machine Learning. PMLR, pp 20336–20350

BibTeX

@InProceedings{pmlr-v202-li23ap,
  title = 	 {{L}o{S}parse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation},
  author =       {Li, Yixiao and Yu, Yifan and Zhang, Qingru and Liang, Chen and He, Pengcheng and Chen, Weizhu and Zhao, Tuo},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {20336--20350},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/li23ap/li23ap.pdf},
  url = 	 {https://proceedings.mlr.press/v202/li23ap.html},
  abstract = 	 {Transformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To re- duce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse ap- proximation), a novel model compression tech- nique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low- rank approximations and pruning, while avoid- ing their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approxima- tions, and low-rank approximation prevents prun- ing from losing too many expressive neurons. We evaluate our method on natural language under- standing, question answering, and natural lan- guage generation tasks. We show that it signif- icantly outperforms existing compression meth- ods. Our code is publicly available at https: //github.com/yxli2123/LoSparse}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/