ZipLM: Inference-Aware Structured Pruning of Language Models

Eldar KurticElias FrantarDan Alistarh

article2023NeurIPS52 citations

Proposes an inference-aware structured pruning method that optimizes the loss-runtime trade-off to generate an entire family of accurate, speedup-guaranteed encoder and decoder language models in a single training run.

Listen

Modern large language models deliver strong performance across various language tasks, but their high computational demands lead to steep infrastructure costs and slow execution during deployment. Structural compression—which removes entire internal components such as attention heads or matrix columns—allows models to run faster on standard hardware without requiring specialized software. However, existing structural pruning approaches often cause severe accuracy drops, demand expensive retraining, require tedious manual tuning, and fail to guarantee actual runtime speedups across different target hardware environments.

The article introduces and evaluates ZipLM, a structured compression framework designed to generate smaller, faster, and highly accurate language models tailored to target inference hardware. The primary objective is to demonstrate that an inference-aware pruning algorithm can consistently deliver state-of-the-art accuracy while precisely meeting user-defined speedup and latency requirements.

To accomplish this, the authors developed a mathematically grounded method that evaluates the trade-off between loss and actual runtime for each component, removing parts one at a time and compensating for their removal across the remaining model weights. The system uses real hardware latency benchmarks to guide pruning and incorporates a token-level knowledge distillation mechanism that transfers capabilities from the original model without requiring manual layer alignment. The researchers tested the approach on both encoder (BERT) and decoder (GPT-2) architectures across standard natural language processing benchmarks, covering both one-shot post-training pruning and gradual pruning with fine-tuning on GPUs and CPUs.

The empirical findings highlight four major results. First, ZipLM outperforms prior structured pruning and distillation techniques across standard benchmarks: on BERT models, it achieves up to 6x to 15x inference speedups while retaining at least 99% of original model accuracy. Second, the method matches the performance of heavily customized architectures, such as MobileBERT, by pruning the standard base model without requiring complex architectural redesigns. Third, on generative models, ZipLM produced a GPT-2 variant that outperformed DistilGPT2 while being 60% smaller and 30% faster. Finally, the framework is highly efficient and predictable: it creates an entire suite of compressed models across multiple speedup targets in a single training run—using roughly one-fifth of the computational training epochs required by competing state-of-the-art methods—with actual measured on-device speedup deviating by less than 5.3% from target specifications.

These results demonstrate that organizations can drastically cut the operational costs, energy consumption, and turnaround times of deploying language models without sacrificing predictive quality. Because ZipLM optimizes directly for specific hardware and deployment modes (such as high-throughput batching versus low-latency interactive generation), engineering teams can avoid costly trial-and-error tuning cycles and predictably meet strict service-level agreements.

Decision-makers and engineering teams should consider adopting hardware-aware structured pruning pipelines over conventional sparsity techniques when preparing models for production. For edge environments with CPU constraints, the article recommends combining ZipLM structured pruning with unstructured pruning and quantization, which yielded up to 50x speedups in testing. Before wide deployment, teams should run pilot benchmarks on their specific target hardware to calibrate the latency tables.

The findings are supported by comprehensive benchmarks on established English-language datasets. However, the study has limitations: the evaluations focus exclusively on English corpora, meaning performance on lower-resource or non-English languages requires further experimental validation. Additionally, the broader availability of highly compressed models increases the need to maintain strong safety, alignment, and watermarking safeguards to prevent misuse.

Cover for ZipLM: Inference-Aware Structured Pruning of Language Models

Abstract

The breakthrough performance of large language models (LLMs) comes with major computational footprints and high deployment costs. In this paper, we progress towards resolving this problem by proposing a novel structured compression approach for LLMs, called ZipLM. ZipLM achieves state-of-the-art accuracy-vs-speedup, while matching a set of desired target runtime speedups in any given inference environment. Specifically, given a model, a dataset, an inference environment, as well as a set of speedup targets, ZipLM iteratively identifies and removes components with the worst loss-runtime trade-off. Unlike prior methods that specialize in either the post-training/one-shot or the gradual compression setting, and only for specific families of models such as BERT (encoder) or GPT (decoder), ZipLM produces state-of-the-art compressed models across all these settings. Furthermore, ZipLM achieves superior results for a fraction of the computational cost relative to prior distillation and pruning techniques, making it a cost-effective approach for generating an entire family of smaller, faster, and highly accurate models, guaranteed to meet the desired inference specifications. In particular, ZipLM outperforms all prior BERTbase distillation and pruning techniques, such as CoFi, MiniLM, and TinyBERT. Moreover, it matches the performance of the heavily optimized MobileBERT model, obtained via extensive architecture search, by simply pruning the baseline BERTlarge model. When compressing GPT2, ZipLM outperforms DistilGPT2 while being 60% smaller and 30% faster. Our code is available at: https://github.com/IST-DASLab/ZipLM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 The ZipLM Structured Pruning Algorithm (Local Correlations)
  • 3.2 Inference-Aware Structured Pruning (Global Correlations)
  • 3.3 Layer-wise Token Distillation
  • 4 Experiments
  • 4.1 Gradual Structured Pruning
  • 4.2 On the Importance of Inference-Awareness
  • 4.3 Post-training/One-shot Structured Pruning
  • 5 Discussion and Extensions
  • References
  • A Compound Compression for Edge Deployment
  • B Ablation Studies
  • C Additional GLUE Results
  • D Additional Validation
  • E Latency Table Used for ZipLM Pruning
  • F Speedup Evaluations
  • G Structure of Pruned Models
  • H Experiments - Additional Results
  • I Hyper-parameters for Reproducibility
  • J Broader Impact and Limitations

Knowls

  1. Knowl 1 — Layer-wise structured reconstruction objective

    model/method

    ZipLM compresses a Transformer layer by reconstructing its original outputs on a small calibration set. Let X∈Rdin×nX\in\mathbb{R}^{d_{\mathrm{in}}\times n} be the matrix of calibration inputs, W∈Rdout×dinW\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}} the original layer weights, and WcW_c the compressed weights constrained to a structured-compression set C\mathcal{C}. ZipLM solves

    argmin⁡Wc∈C  ∥WcX−WX∥F2.\underset{W_c\in\mathcal{C}}{\operatorname{argmin}}\;\left\|W_cX-WX\right\|_F^2.

    The structured constraint requires the same input-weight indices to be removed from every row of WW, so pruning a structure such as a column or an attention head couples the row-wise reconstruction problems. For a selected structure SS, ZipLM simultaneously chooses the structure to remove and an update to the remaining weights that compensates for the removed weights. In Transformer layers, the method removes attention heads, intermediate-dimension columns from the second feed-forward linear layer, or complete attention/feed-forward residual modules; removing columns from the second linear layer permits the corresponding rows in the preceding layer to be deleted without changing the layer output.

  2. Knowl 2 — Second-order one-at-a-time structured pruning

    algorithm

    For a layer with calibration-input matrix XX, ZipLM uses the input Hessian H=XX⊤H=XX^\top for the squared reconstruction objective and applies a regularized inverse in practice, H−1=(2XX⊤+λI)−1H^{-1}=(2XX^\top+\lambda I)^{-1}, where λ≥0\lambda\geq 0 is a damping parameter and II is the identity matrix. For a candidate structure SS with pruning mask MSM_S, let Wi,MSW_{i,M_S} denote the weights selected by the mask in row ii, let W:,MSW_{:,M_S} denote the selected columns across all rows, and let (H−1)MS,MS(H^{-1})_{M_S,M_S} and (H−1)MS,:(H^{-1})_{M_S,:} denote the corresponding submatrices. ZipLM assigns the structure the joint saliency

    E(S)=∑i=1doutWi,MS((H−1)MS,MS)−1Wi,MS⊤,E(S)=\sum_{i=1}^{d_{\mathrm{out}}}W_{i,M_S}\left((H^{-1})_{M_S,M_S}\right)^{-1}W_{i,M_S}^{\top},

    and uses the compensating update

    δS=−W:,MS((H−1)MS,MS)−1(H−1)MS,:.\delta_S=-W_{:,M_S}\left((H^{-1})_{M_S,M_S}\right)^{-1}(H^{-1})_{M_S,:}.

    The significant structures are removed one at a time rather than by independently ranking the original model structures. After each removal, ZipLM applies W←W+δSW\leftarrow W+\delta_S and updates the inverse Hessian by block Gaussian elimination:

    H−1←H−1−H:,MS−1((H−1)MS,MS)−1HMS,:−1.H^{-1}\leftarrow H^{-1}-H^{-1}_{:,M_S}\left((H^{-1})_{M_S,M_S}\right)^{-1}H^{-1}_{M_S,:}.
    Input: Weight matrix WW, calibration inputs XX, candidate structures R\mathcal{R}, number of structures kk, damping λ\lambda
    Output: Structurally pruned weight matrix WW
    Compute H−1=(2XX⊤+λI)−1H^{-1}=(2XX^\top+\lambda I)^{-1}
    for t=1t=1 to kk do
      For every remaining structure SS in R\mathcal{R}, compute E(S)E(S)
      Select the structure SS with minimum E(S)E(S)
      Compute δS=−W:,MS((H−1)MS,MS)−1(H−1)MS,:\delta_S=-W_{:,M_S}((H^{-1})_{M_S,M_S})^{-1}(H^{-1})_{M_S,:}
      Update W←W+δSW\leftarrow W+\delta_S
      Update H−1H^{-1} by block Gaussian elimination
      Remove SS from R\mathcal{R}
    end for
    Set all weights selected by the final pruning mask exactly to zero
    return WW

    The update makes correlated or redundant structures less attractive after a neighboring structure has been removed. For non-overlapping structures of size ∣MS∣|M_S|, the inverse-Hessian update costs O(∣MS∣dcol2)O(|M_S|d_{\mathrm{col}}^2) per step, while the additional small matrix inversions cost O(dcol∣MS∣2)O(d_{\mathrm{col}}|M_S|^2) when reused across rows. Attention heads are treated as blocks of consecutive columns, typically 64 columns, whereas feed-forward intermediate structures are individual columns.

  3. Knowl 3 — Inference-aware speedup-constrained search

    model/method

    ZipLM selects structures using measured runtime rather than sparsity alone. For a target inference environment, it builds a latency lookup table containing end-to-end Transformer attention-block runtimes for each possible number of remaining attention heads and feed-forward-block runtimes for intermediate dimensions sampled as 0.9i0.9^i relative reductions, i=0,…,42i=0,\ldots,42, covering relative steps of approximately 10% up to about 99% sparsity. The table includes implementation overheads and is measured on the target hardware, batch size, and sequence length.

    For each layer, ZipLM uses the one-at-a-time pruning procedure to precompute candidate versions at multiple sparsities. It then searches over layer-wise candidate configurations whose estimated total runtime satisfies a requested speedup while minimizing the accumulated reconstruction loss. The search adapts SPDY to structured pruning by assigning a layer-wise prior

    ps=∥Wc,sX−WX∥F2∥WX∥F2,p_s=\frac{\left\|W_{c,s}X-WX\right\|_F^2}{\left\|WX\right\|_F^2},

    where ss is the structured sparsity level and Wc,sW_{c,s} is the corresponding compressed layer. This prior reaches 1 when the entire layer is dropped, unlike a quadratic sparsity prior that can incorrectly treat near-total structural removal as only marginally harder than moderate pruning. ZipLM performs 1000 fixed search steps, randomly mutating an expected 10% of the layer-sensitivity coefficients at each step. The same run produces a family of models for multiple target speedups, such as 2x through 15x.

  4. Knowl 4 — Layer-wise token-level distillation

    equation

    Because ZipLM preserves the Transformer hidden dimension, it distills token representations from the teacher and student at every compatible unpruned layer without manually matching teacher and student layers. For input xx, student parameters θs\theta^s, teacher parameters θt\theta^t, task loss LtaskL_{\mathrm{task}}, output-logit KL-divergence LlogitL_{\mathrm{logit}}, and token loss LtokenL_{\mathrm{token}}, the training objective is

    L(θs,θt∣x)=λ1Ltask(θs∣x)+λ2Llogit(θs,θt∣x)+λ3Ltoken(θs,θt∣x).L(\theta^s,\theta^t\mid x)=\lambda_1L_{\mathrm{task}}(\theta^s\mid x)+\lambda_2L_{\mathrm{logit}}(\theta^s,\theta^t\mid x)+\lambda_3L_{\mathrm{token}}(\theta^s,\theta^t\mid x).

    At layer kk, let hk,js,hk,jt∈Rdhh^s_{k,j},h^t_{k,j}\in\mathbb{R}^{d_h} be the student and teacher hidden vectors for token position jj, let PP be the set of padding-token positions, and let Nnp=∑j1[j∉P]N_{\mathrm{np}}=\sum_j\mathbf{1}[j\notin P] be the number of non-padding tokens. The token-level loss is the average Euclidean distance over non-padding tokens and over all unpruned layers:

    Ltokenk=1Nnp∑j1[j∉P] ∥hk,js−hk,jt∥2.L^k_{\mathrm{token}}=\frac{1}{N_{\mathrm{np}}}\sum_j\mathbf{1}[j\notin P]\,\left\|h^s_{k,j}-h^t_{k,j}\right\|_2.

    Thus, the student is encouraged to reproduce the teacher’s representation for every input token while avoiding the manual layer selection and learned shape-matching transformations used by many earlier structured-distillation methods.

  5. Knowl 5 — Evaluation across encoder and decoder models

    experimental setup

    ZipLM was evaluated in both gradual and post-training structured-pruning settings. Encoder experiments used pretrained BERT-base and BERT-large on SQuADv1.1 and the GLUE tasks SST-2, QNLI, MNLI, and QQP. For direct comparisons with prior structural-pruning work, BERT inference used one NVIDIA V100 16 GB GPU, batch size 128, and sequence lengths 384 for SQuAD and 128 for GLUE. Gradual BERT runs were fine-tuned before pruning and between pruning steps, and produced target speedups from 2x through 15x.

    Decoder experiments pruned the 124M-parameter GPT-2 model using OpenWebTextCorpus, followed by zero-shot evaluation on WikiText-103 without downstream fine-tuning. Throughput pruning used batch size 16 and sequence length 1024, with target speedups of 1.5x, 2x, 2.5x, and 3x. Latency pruning used batch size 1 and variable-length text-generation prompts with Top-KK sampling. The GPT-2 runs did not use knowledge distillation because of its memory overhead.

  6. Knowl 6 — BERT accuracy-speedup results

    empirical result

    ZipLM produced stronger accuracy-speedup trade-offs than prior structured-pruning and distillation methods on BERT-base and BERT-large. On SQuADv1.1 with BERT-base, ZipLM exceeded CoFi and TinyBERT by about 3 F1 points at comparable speedups, or achieved at least 60% higher speedup at comparable F1. On the GLUE tasks, ZipLM consistently improved both accuracy and speedup over the strongest competing methods. It preserved the dense BERT-base accuracy while reaching up to 6x speedup on QQP and 10x on SST-2.

    At the MLPerf threshold of more than 99% recovery of dense-model accuracy, ZipLM obtained BERT-base speedups of 5x on SQuADv1.1, 6x on QNLI, 6x on MNLI, 13x on SST-2, and 15x on QQP. On BERT-large SQuADv1.1, it maintained the dense model’s F1 at 4x speedup and reached 6x speedup at 99% recovery. ZipLM also matched the performance of MobileBERT by pruning the standard BERT-large architecture, without MobileBERT’s custom bottlenecks, factorized embeddings, operator substitutions, or architecture-search procedure.

  7. Knowl 7 — GPT-2 throughput and latency compression

    data/table

    ZipLM produced different GPT-2 architectures for throughput-constrained and latency-constrained inference, even when the requested speedups were similar. Throughput pruning favored reducing matrix dimensions while preserving depth, whereas latency pruning favored dropping complete Transformer modules while retaining more width. The zero-shot WikiText-103 perplexities, measured after pruning GPT-2 on OpenWebTextCorpus, were:

    Could not parse LaTeX table

    At comparable throughput speedup, ZipGPT2 at 1.5x had perplexity 35.4 versus DistilGPT2’s 43.0. ZipGPT2 at 2.1x reduced the decoder to 26.5M parameters, a 60% reduction relative to the 42.5M-parameter DistilGPT2, while improving speedup from 1.6x to 2.1x and obtaining perplexity 41.5. At comparable latency speedup, ZipGPT2 at 2.0x used 39.2M parameters and achieved perplexity 41.2, compared with 42.5M parameters and perplexity 43.0 for DistilGPT2 at 1.9x. The dense GPT-2 reference was trained by OpenAI on a larger private dataset, so the direct compressed-model comparison is primarily with DistilGPT2.

  8. Knowl 8 — Post-training one-shot pruning and calibration robustness

    data/table

    ZipLM also works without retraining after pruning. On BERT-base, its one-shot results exceeded the prior post-training method of Kwon et al. at both tested speedups on SQuADv1.1, QQP, and MNLI:

    Could not parse LaTeX table

    The method was robust to calibration-set size on one-shot BERT-base SQuADv1.1 pruning. At 1.5x and 2.0x speedups, ZipLM obtained F1 scores of 82.3 and 48.4 with 4 calibration samples, 86.8 and 82.6 with 32 samples, 86.8 and 83.6 with 128 samples, 86.8 and 84.1 with 512 samples, 87.1 and 84.1 with 2048 samples, and 87.6 and 84.7 with 4096 samples. The competing Kwon et al. method used 2048 samples and obtained 86.2 and 76.5.

  9. Knowl 9 — Measured runtime fidelity and the value of inference-awareness

    data/table

    Runtime measurements showed that equal sparsity does not imply equal speedup across inference devices. For shrinking a Transformer feed-forward intermediate dimension, the measured speedups were:

    Could not parse LaTeX table

    For example, reducing the intermediate dimension from 3072 to 302, approximately 90% sparsity, produced 6.9x speedup on a V100 but only 3.1x on an A100. In the target-speedup experiments, the difference between requested and measured speedup was at most 5.28%: for BERT-base SQuADv1.1, targets of 2x, 4x, 6x, 8x, 10x, 12x, and 14x achieved 1.98x, 4.05x, 6.16x, 8.25x, 10.36x, 12.31x, and 14.33x; for BERT-large, they achieved 2.01x, 4.05x, 6.09x, 8.27x, 10.33x, 12.46x, and 14.74x. On BERT-base SQuADv1.1, pruning toward measured speedup improved F1 by as much as 10 points over pruning toward the same nominal sparsity, especially at high compression.

  10. Knowl 10 — CPU deployment and computational efficiency

    empirical result

    ZipLM can replace layer dropping in a compound edge-compression pipeline. The evaluated pipeline first applies ZipLM structured pruning, then oBERT unstructured pruning to 80% sparsity, and finally INT8 quantization-aware training. On a single Intel Cascade Lake CPU core using the DeepSparse inference engine, ZipLM plus oBERT plus quantization-aware training improved the speedup at full SQuADv1.1 accuracy recovery from 3x for the layer-dropping pipeline to 13x. At the largest compression ratio, it improved speedup from 30x to 50x.

    ZipLM was also substantially cheaper than distillation-based compression. Producing the full BERT-base family of models with speedup targets from 2x through 15x required 115 total training epochs for ZipLM versus 560 epochs for CoFi, making ZipLM 4.87 times more efficient in that comparison. On one RTX A6000, generating the family required approximately 35 hours on larger datasets such as MNLI and approximately 10 hours on smaller datasets such as SST-2. ZipLM used one hyperparameter set to produce the entire family, whereas competing methods generally require separate tuning for each compressed model.

  11. Knowl 11 — Extreme-compression scaling laws

    theoretical result

    ZipLM produced structuredly pruned BERT models without model collapse at extreme target speedups of up to 75x for BERT-base and 55x for BERT-large. Across the evaluated speedup range, the accuracy-speedup curves were approximately linear, and the fitted SQuADv1.1 relationships were

    F1large≈92.1−0.3 speeduplarge,F1_{\mathrm{large}}\approx 92.1-0.3\,\mathrm{speedup}_{\mathrm{large}}, F1base≈90.3−0.6 speedupbase,F1_{\mathrm{base}}\approx 90.3-0.6\,\mathrm{speedup}_{\mathrm{base}},

    where each speedup variable is the inference-speedup factor relative to the corresponding dense BERT model and each F1F1 value is in percentage points. The fitted degradation rate for BERT-base is approximately twice that of BERT-large, which the paper attributes to greater redundant representational capacity in the larger model. At 15x speedup, the average BERT-base model retained about 2% of its intermediate dimension and 6% of its attention heads, corresponding to approximately 2.9M encoder parameters while retaining more than 95% of the dense model’s accuracy.

Coverage note — The paper’s English-only benchmark limitation and broader-impact discussion were omitted because they are stated limitations rather than load-bearing parts of the ZipLM method or empirical contribution.

References

  1. 1.Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. Once-for-all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791, 2019.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.
  3. 3.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. In International Conference on Learning Representations, 2019.
  4. 4.Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833, 2018.
  5. 5.Elias Frantar and Dan Alistarh. SPDY: Accurate pruning with speedup guarantees. arXiv preprint arXiv:2201.13096, 2022.
  6. 6.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  7. 7.Elias Frantar, Eldar Kurtic, and Dan Alistarh. M-fac: Efficient matrix-free approximations of second-order information. Advances in Neural Information Processing Systems, 34, 2021.
  8. 8.Elias Frantar, Sidak Pal Singh, and Dan Alistarh. Optimal Brain Compression: A framework for accurate post-training quantization and pruning. arXiv preprint arXiv:2208.11580, 2022. Accepted to NeurIPS 2022, to appear.
  9. 9.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. arXiv preprint arXiv:2103.13630, 2021.
  10. 10.Aaron Gokaslan and Vanya Cohen. Openwebtext corpus, 2019.
  11. 11.Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6):1789–1819, 2021.
  12. 12.Zichao Guo, Xiangyu Zhang, Haoyuan Mu, Wen Heng, Zechun Liu, Yichen Wei, and Jian Sun. Single path one-shot neural architecture search with uniform sampling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVI 16, pages 544–560. Springer, 2020.
  13. 13.Babak Hassibi and David Stork. Second order derivatives for network pruning: Optimal brain surgeon. Advances in neural information processing systems, 5, 1992.
  14. 14.Yang He, Ping Liu, Ziwei Wang, Zhilan Hu, and Yi Yang. Filter pruning via geometric median for deep convolutional neural networks acceleration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4340–4349, 2019.
  15. 15.Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. Amc: Automl for model compression and acceleration on mobile devices. In Proceedings of the European conference on computer vision (ECCV), pages 784–800, 2018.
  16. 16.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7), 2015.
  17. 17.Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research, 22(241):1–124, 2021.
  18. 18.Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. Dynabert: Dynamic bert with adaptive width and depth. Advances in Neural Information Processing Systems, 33:9782–9793, 2020.
  19. 19.Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry. Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks. Advances in Neural Information Processing Systems, 34:21099–21111, 2021.
  20. 20.Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704–2713, 2018.
  21. 21.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174, 2020.
  22. 22.Aran Komatsuzaki. One epoch is all you need. arXiv preprint arXiv:1906.06669, 2019.
  23. 23.Eldar Kurtic and Dan Alistarh. Gmp*: Well-tuned global magnitude pruning can outperform most bert-pruning methods. arXiv preprint arXiv:2210.06384, 2022.
  24. 24.Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4163––4181, 2022.
  25. 25.Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Bill Nell, Nir Shavit, and Dan Alistarh. Inducing and exploiting activation sparsity for fast inference on deep neural networks. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 5533–5543, Virtual, 13–18 Jul 2020. PMLR.
  26. 26.Woosuk Kwon, Sehoon Kim, Michael W Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers. arXiv preprint arXiv:2204.09656, 2022.
  27. 27.François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush. Block pruning for faster transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10619–10629, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
  28. 28.Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. Datasets: A community library for natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184. Association for Computational Linguistics, November 2021.
  29. 29.Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710, 2016.
  30. 30.Yawei Li, Kamil Adamczewski, Wen Li, Shuhang Gu, Radu Timofte, and Luc Van Gool. Revisiting random channel pruning for neural network compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 191–201, 2022.
  31. 31.Lucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman, and Daniela Rus. Provable filter pruning for efficient neural networks. arXiv preprint arXiv:1911.07412, 2019.
  32. 32.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  33. 33.Liyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang. Group fisher pruning for practical network compression. In International Conference on Machine Learning, pages 7021–7032. PMLR, 2021.
  34. 34.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE international conference on computer vision, pages 2736–2744, 2017.
  35. 35.Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al. Deja vu: Contextual sparsity for efficient llms at inference time. In International Conference on Machine Learning, pages 22137–22176. PMLR, 2023.
  36. 36.Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, et al. Mlperf: An industry standard benchmark suite for machine learning performance. IEEE Micro, 40(2):8–16, 2020.
  37. 37.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016.
  38. 38.Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or down? adaptive rounding for post-training quantization. In International Conference on Machine Learning, pages 7197–7206. PMLR, 2020.
  39. 39.NeuralMagic. Deep sparse: A fast cpu inference engine, 2021.
  40. 40.Matan Ben Noach and Yoav Goldberg. Compressing pre-trained language models by matrix decomposition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 884–889, 2020.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  42. 42.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In EMNLP, 2016.
  43. 43.Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. Poor man’s bert: Smaller and faster transformer models. arXiv preprint arXiv:2004.03844, 2020.
  44. 44.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019.
  45. 45.Victor Sanh, Thomas Wolf, and Alexander Rush. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378–20389, 2020.
  46. 46.S. Shankar. Identifying quora question pairs having the same intent. 2017.
  47. 47.Sidak Pal Singh and Dan Alistarh. Woodfisher: Efficient second-order approximation for neural network compression. Advances in Neural Information Processing Systems, 33, 2020.
  48. 48.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642, 2013.
  49. 49.Yang Sui, Miao Yin, Yi Xie, Huy Phan, Saman Aliari Zonouz, and Bo Yuan. Chip: Channel independence-based pruning for compact neural networks. Advances in Neural Information Processing Systems, 34:24604–24616, 2021.
  50. 50.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. Patient knowledge distillation for bert model compression. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4323–4332, 2019.
  51. 51.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. Mobilebert: a compact task-agnostic bert for resource-limited devices. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2158–2170, 2020.
  52. 52.Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Well-read students learn better: The impact of student initialization on knowledge distillation. ArXiv, abs/1908.08962, 2019.
  53. 53.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  54. 54.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. ArXiv, abs/1804.07461, 2018.
  55. 55.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems, 33:5776–5788, 2020.
  56. 56.Ziheng Wang, Jeremy Wohlwend, and Tao Lei. Structured pruning of large language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6151–6162, 2020.
  57. 57.Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics, 2018.
  58. 58.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, October 2020. Association for Computational Linguistics.
  59. 59.Mengzhou Xia, Zexuan Zhong, and Danqi Chen. Structured pruning learns compact and accurate models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1513–1528, 2022.
  60. 60.Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. Bert-of-theseus: Compressing bert by progressive module replacing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7859–7869, 2020.

Citation

MLA
Kurtić, E., et al. “ZipLM: Inference-Aware Structured Pruning of Language Models”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 65597–617, https://proceedings.neurips.cc/paper_files/paper/2023/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf.
APA
Kurtić, E., Frantar, E., & Alistarh, D. (2023). ZipLM: Inference-Aware Structured Pruning of Language Models. Advances in Neural Information Processing Systems, 36, 65597–65617. https://proceedings.neurips.cc/paper_files/paper/2023/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf
Chicago
Kurtić, E., E. Frantar, and D. Alistarh. 2023. “ZipLM: Inference-Aware Structured Pruning of Language Models”. Advances in Neural Information Processing Systems 36: 65597–617. https://proceedings.neurips.cc/paper_files/paper/2023/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf.
Harvard
Kurtić, E., Frantar, E. and Alistarh, D. (2023) “ZipLM: Inference-Aware Structured Pruning of Language Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 65597–65617. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf.
Vancouver
1. Kurtić E, Frantar E, Alistarh D (2023) ZipLM: Inference-Aware Structured Pruning of Language Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 65597–65617

BibTeX

@inproceedings{kurtic2023ziplm,
  title = {ZipLM: Inference-Aware Structured Pruning of Language Models},
  author = {Kurtić, Eldar and Frantar, Elias and Alistarh, Dan},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {65597-65617},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/ced46a50befedcb884ccf0cbe8c3ad23-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors