A Simple and Effective Pruning Approach for Large Language Models

Mingjie SunZhuang LiuAnna BairJ. Kolter

article2023ICLR977 citations

Introduces Wanda, a post-training pruning method that sparsifies large language models by evaluating the product of weight magnitudes and input activations, matching computationally expensive alternatives without requiring retraining or weight updates.

Listen

Large Language Models deliver state-of-the-art natural language capabilities but demand substantial computational and memory resources due to their massive parameter scale. Model pruning—the practice of setting unneeded weights to zero—offers an effective path toward compression, yet existing techniques remain impractical for modern architectures. Traditional pruning requires extensive retraining or fine-tuning, while recent one-shot methods rely on complex, computationally expensive weight updates. Meanwhile, conventional magnitude-based pruning fails dramatically when applied directly to modern language models.

The article introduces and evaluates Wanda (Pruning by Weights and activations), a simple, one-shot pruning method designed to induce high sparsity in pretrained language models without requiring model retraining or weight updates. Motivated by the emergence of outlier features with exceptionally large values in billion-scale architectures, the method determines weight importance by multiplying each individual weight's magnitude by the norm of its corresponding input activation, evaluating these scores locally on a per-output neuron basis.

The researchers evaluated Wanda across the LLaMA and LLaMA-2 model families spanning 7B to 70B parameters, with additional testing on OPT, BLOOM, and Pythia architectures. Across zero-shot reasoning benchmarks, few-shot evaluations, and language modeling perplexity tests, the approach was compared directly against standard magnitude pruning and the leading second-order baseline, SparseGPT, under unstructured (50%, 60%, 80%) and hardware-friendly structured (2:4 and 4:8) sparsity patterns using modest calibration data.

The findings show that Wanda substantially outperforms standard magnitude pruning across all evaluated architectures and sparsity levels without modifying retained weights. At 50% unstructured sparsity, Wanda achieves performance competitive with SparseGPT, and in the largest models (LLaMA-65B and LLaMA-2-70B), the 50% sparse models match the zero-shot accuracy of their original dense counterparts. Computationally, calculating pruning scores with Wanda is up to 300 times faster than SparseGPT, executing in a single forward pass with high robustness even when calibrated on very small sample sizes. Furthermore, structured 2:4 sparsity delivers an approximate 1.6-fold speedup in core matrix multiplications on standard graphics hardware.

These results demonstrate that large language models contain effective, exact sparse subnetworks that can be uncovered without costly weight reconstruction or iterative updates. For organizations deploying generative artificial intelligence, this technique significantly reduces the computational overhead, runtime latency, and engineering costs of model compression. It also shows that the common drop in accuracy observed after pruning can be largely recovered through lightweight fine-tuning methods such as low-rank adaptation.

Engineering teams looking to reduce inference costs and memory footprints should adopt Wanda as an efficient baseline for model sparsification, particularly for larger model tiers where performance degradation is minimal. When higher accuracy is mandatory on smaller models, teams should pair Wanda pruning with parameter-efficient fine-tuning. Further work should explore applying Wanda to dynamic sparse training from scratch and evaluating its effectiveness across non-transformer architectures, domain-specific tasks, and extreme sparsity thresholds exceeding 70%.

Cover for A Simple and Effective Pruning Approach for Large Language Models

Abstract

As their size increases, Large Languages Models (LLMs) are natural candidates for network pruning methods: approaches that drop a subset of network weights while striving to preserve performance. Existing methods, however, require either retraining, which is rarely affordable for billion-scale LLMs, or solving a weight reconstruction problem reliant on second-order information, which may also be computationally expensive. In this paper, we introduce a novel, straightforward yet effective pruning method, termed Wanda (Pruning by Weights and activations), designed to induce sparsity in pretrained LLMs. Motivated by the recent observation of emergent large magnitude features in LLMs, our approach prunes weights with the smallest magnitudes multiplied by the corresponding input activations, on a per-output basis. Notably, Wanda requires no retraining or weight update, and the pruned LLM can be used as is. We conduct a thorough evaluation of our method Wanda on LLaMA and LLaMA-2 across various language benchmarks. Wanda significantly outperforms the established baseline of magnitude pruning and performs competitively against recent method involving intensive weight update. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Wanda: Pruning by Weights and Activations
  • 4 Experiments
  • 4.1 Zero-Shot Tasks
  • 4.2 Language Modeling
  • 4.3 Speedup
  • 5 Analysis
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Image classifiers
  • B Wanda on previous LLMs
  • C Additional Baselines
  • D complementary Experimental Results
  • D.1 Number of Calibration Samples
  • D.2 Robustness Analysis
  • D.3 Higher Sparsity
  • D.4 Few-shot results on MMLU
  • D.5 Fine-tuning
  • D.6 Zero-Shot Tasks

Knowls

  1. Knowl 1 — Wanda scores weights using both weight magnitude and input activation

    model/method

    For a linear layer with weight matrix W∈RCout×CinW\in\mathbb{R}^{C_{\text{out}}\times C_{\text{in}}} and calibration activations X∈R(NL)×CinX\in\mathbb{R}^{(N L)\times C_{\text{in}}}, Wanda assigns weight WijW_{ij} the score

    Sij=∣Wij∣ ∥X:,j∥2.S_{ij}=|W_{ij}|\,\|X_{:,j}\|_2.

    Here ii indexes an output and jj an input feature; X:,jX_{:,j} contains that feature's activation across the NN batch elements and LL token positions. Lower scores indicate weights to prune. Unlike magnitude-only scoring, this gives greater importance to weights connected to high-magnitude input features, even when those weights themselves have small magnitudes. The paper reports that the activation norm can be estimated from a modest calibration set, and that the ℓ2\ell_2 norm worked better in its experiments than ℓ1\ell_1 or ℓ∞\ell_\infty norms.

  2. Knowl 2 — Wanda compares weights separately within each output

    model/method

    For every output row ii of a linear layer, Wanda compares the scores of the weights in that row only. The comparison group for WijW_{ij} is Gij={Wiv:1≤v≤Cin}G_{ij}=\{W_{i v}:1\leq v\leq C_{\text{in}}\}. At target sparsity ss, the lowest-scoring fraction ss of the weights in each output row is set to zero. Thus each output receives the same pruning fraction, rather than weights being ranked across an entire layer. The authors find that this per-output grouping is consistently better than layer-wise grouping for LLMs, including when the scores use ordinary weight magnitude instead of Wanda's activation-aware metric.

  3. Knowl 3 — One-pass, layerwise Wanda pruning requires no weight updates

    algorithm

    Wanda prunes a pretrained model in a forward traversal of its linear layers. Calibration activations are used to compute each input feature's ℓ2\ell_2 norm; after pruning one layer, the next layer's activations are obtained by forwarding the calibration inputs through the already-pruned prefix. For each output, the algorithm zeros the requested fraction of lowest-scoring input weights. It does not update surviving weights, use gradients, or iteratively refine the mask. The paper's main LLM experiments use 128 calibration sequences sampled from C4. Pruning focuses on linear layers, excluding the initial embedding layer and final classification head.

    Input: Pretrained model, calibration sequences, target sparsity ss
    Output: Model with pruned linear-layer weights
    For each linear layer in forward order:
        Run calibration sequences through the current model prefix
        Record layer inputs XX with shape (NL,Cin)(N L, C_{in})
        For each output row ii:
            For each input feature jj, compute Sij=∣Wij∣ ∥X:,j∥2S_{ij}=|W_{ij}|\,\|X_{:,j}\|_2
            Select the ⌊sCin⌋\lfloor s C_{in}\rfloor input indices with lowest scores
            Set the selected weights in row ii to zero
    Return the pruned model
  4. Knowl 4 — Wanda extends to structured N:M sparsity

    model/method

    For structured N:MN:M sparsity, each group of MM consecutive weights is constrained to contain at most NN nonzero weights. Wanda applies the same weight-and-activation score used for unstructured pruning, but compares scores within each group of MM consecutive weights connected to an output, pruning the lower-priority weights in that group. This adapts Wanda to a hardware-supported sparsity pattern without changing its importance metric.

  5. Knowl 5 — Wanda preserves zero-shot accuracy competitively at 50% and structured sparsities

    empirical result

    The table reports mean accuracy across seven zero-shot tasks for LLaMA and LLaMA-2 models. All pruned models use uniform sparsity in linear layers and no fine-tuning; SparseGPT uses weight updates, whereas Wanda and magnitude pruning do not. Wanda substantially exceeds magnitude pruning and is generally competitive with SparseGPT. At unstructured 50% sparsity, Wanda reaches 66.67% on LLaMA-65B and 67.03% on LLaMA-2-70B, close to dense results of 66.97% and 67.08%, respectively.

    Method Sparsity LLaMA 7B 13B 30B 65B LLaMA-2 7B 13B 70B
    Dense 0% 59.99 62.59 65.38 66.97 59.71 63.03 67.08
    Magnitude 50% 46.94 47.61 53.83 62.74 51.14 52.85 60.93
    SparseGPT 50% 54.94 58.61 63.09 66.30 56.24 60.72 67.28
    Wanda 50% 54.21 59.33 63.60 66.67 56.24 60.83 67.03
    Magnitude 4:8 46.03 50.53 53.53 62.17 50.64 52.81 60.28
    SparseGPT 4:8 52.80 55.99 60.79 64.87 53.80 59.15 65.84
    Wanda 4:8 52.76 56.09 61.00 64.97 52.49 58.75 66.06
    Magnitude 2:4 44.73 48.00 53.16 61.28 45.58 49.89 59.95
    SparseGPT 2:4 50.60 53.22 58.91 62.57 50.94 54.86 63.89
    Wanda 2:4 48.53 52.30 59.21 62.84 48.75 55.03 64.14
  6. Knowl 6 — Wanda matches SparseGPT on many WikiText perplexity results without updating weights

    empirical result

    These are WikiText validation perplexities for pruned LLaMA and LLaMA-2 models. Lower is better. Pruning uses the same calibration set for the methods that require calibration, and applies uniform sparsity to linear layers; Wanda does not update surviving weights. At unstructured 50% sparsity, Wanda is close to SparseGPT across model sizes and greatly improves over magnitude pruning—for example, LLaMA-7B has perplexity 7.26 with Wanda versus 17.29 with magnitude pruning. Wanda's comparison with SparseGPT under structured sparsity varies by model size. At much higher unstructured sparsity, performance degrades sharply: the reported best 80%-sparse model has perplexity 25.86, far worse than the 5.68 perplexity of dense LLaMA-7B.

    Method Sparsity LLaMA 7B 13B 30B 65B LLaMA-2 7B 13B 70B
    Dense 0% 5.68 5.09 4.77 3.56 5.12 4.57 3.12
    Magnitude 50% 17.29 20.21 7.54 5.90 14.89 6.37 4.98
    SparseGPT 50% 7.22 6.21 5.31 4.57 6.51 5.63 3.98
    Wanda 50% 7.26 6.15 5.24 4.57 6.42 5.56 3.98
    Magnitude 4:8 16.84 13.84 7.62 6.36 16.48 6.76 5.54
    SparseGPT 4:8 8.61 7.40 6.17 5.38 8.12 6.60 4.59
    Wanda 4:8 8.57 7.40 5.97 5.30 7.97 6.55 4.47
    Magnitude 2:4 42.13 18.37 9.10 7.11 54.59 8.33 6.33
    SparseGPT 2:4 11.00 9.11 7.16 6.28 10.17 8.32 5.40
    Wanda 2:4 11.53 9.58 6.90 6.25 11.02 8.27 5.16
  7. Knowl 7 — Per-output comparison is the best tested Wanda grouping for LLaMA-7B

    data/table

    This ablation measures WikiText validation perplexity after pruning LLaMA-7B to unstructured 50% sparsity, with no weight-update procedure. It varies the comparison group while holding each score metric fixed. For Wanda's activation-aware metric, comparing weights within each output row gives the lowest perplexity, 7.26; layer-wise comparison gives 7.95. Per-output grouping also improves the tested magnitude and SparseGPT-style metrics relative to layer-wise comparison, showing that grouping choice matters in addition to the score itself.

    Pruning metric Layer Input, 1 Input, 128 Output, 1 Output, 128
    Magnitude: ∣Wij∣|W_{ij}| 17.29 8.86 16.82 13.41 17.47
    SparseGPT-style metric 7.91 8.86 8.02 7.41 7.74
    Wanda: ∣Wij∣ ∥X:,j∥2|W_{ij}|\,\|X_{:,j}\|_2 7.95 8.86 8.12 7.26 7.71
  8. Knowl 8 — Wanda reduces pruning time and structured sparsity accelerates matrix multiplication

    empirical result

    On NVIDIA A6000 GPUs, the reported time for computing pruning metrics—excluding the forward pass shared by both methods—is substantially lower for Wanda than for SparseGPT. The paper gives theoretical costs of O(dhidden2)O(d_{\text{hidden}}^2) for Wanda and O(dhidden3)O(d_{\text{hidden}}^3) for SparseGPT, where dhiddend_{\text{hidden}} is the model hidden-state dimension. For inference, a CUTLASS GEMM simulation of LLaMA-65B linear layers at batch size 1 finds 1.54–1.63×\times speedups with structured 2:4 sparsity. The authors also report 1.24×\times end-to-end speedup for LLaMA-7B, with latency reduced from 312 ms to 251 ms.

    Pruning metric time (seconds) LLaMA 7B 13B 30B 65B
    SparseGPT 203.1 339.0 810.3 1353.4
    Wanda 0.54 0.91 2.9 5.6
  9. Knowl 9 — Wanda's calibration estimates are stable with few samples

    empirical result

    The paper varies the number of calibration sequences for unstructured 50% pruning and evaluates WikiText validation perplexity. Wanda is comparatively insensitive to calibration-set size: for LLaMA-7B its perplexity ranges from 7.46 to 7.25 over the tested sizes, while SparseGPT ranges from 10.22 to 7.19. For LLaMA-2-7B, Wanda remains between 6.53 and 6.45, whereas SparseGPT ranges from 8.63 to 6.49. These results support the claim that Wanda's activation-norm statistics can be estimated reliably from limited calibration data; the main experiments use 128 sequences.

    Model Method 1 16 32 64 128 256 512 1024 2048
    LLaMA-7B SparseGPT 10.22 7.61 7.36 7.29 7.26 7.20 7.19 7.23 7.20
    LLaMA-7B Wanda 7.46 7.27 7.28 7.28 7.26 7.30 7.26 7.25 7.26
    LLaMA-2-7B SparseGPT 8.63 6.67 6.62 6.61 6.53 6.52 6.50 6.49 6.49
    LLaMA-2-7B Wanda 6.53 6.45 6.46 6.45 6.45 6.45 6.45 6.45 6.45
  10. Knowl 10 — Wanda also prunes OPT, Pythia, and BLOOM models

    empirical result

    The authors test Wanda beyond LLaMA on OPT, Pythia, and BLOOM. At unstructured 50% sparsity, magnitude pruning can fail catastrophically on these families, while Wanda produces finite perplexities close to SparseGPT. In the examples below, values are WikiText perplexities; lower is better. Wanda is slightly worse than SparseGPT for Pythia-12B, but slightly better for BLOOM-7.1B. Across OPT and BLOOM sizes, the paper reports that Wanda's gap from SparseGPT generally narrows as model size increases.

    Model Dense Magnitude SparseGPT Wanda
    Pythia-12B 8.59 3e53\mathrm{e}5 11.02 11.27
    BLOOM-7.1B 11.37 2e62\mathrm{e}6 13.96 13.55

Coverage note — Omitted the auxiliary fine-tuning recovery study, image-classifier comparison, and the interpretive connection to diagonal Hessian pruning because they are secondary to Wanda's method and its primary LLM pruning results.

References

  1. 1.Arash Ahmadian, Saurabh Dash, Hongyu Chen, Bharat Venkitesh, Stephen Gou, Phil Blunsom, Ahmet Üstün, and Sara Hooker. Intriguing properties of quantization at scale. In NeurIPS, 2023.
  2. 2.Mohammad Babaeizadeh, Paris Smaragdis, and Roy H. Campbell. Noiseout: A simple way to prune neural networks. arXiv preprint arXiv:1611.06211, 2016.
  3. 3.Hritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, and Dan Roth. Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale. In Association for Computational Linguistics (ACL), 2023.
  4. 4.Kayhan Behdin, Ayan Archarya, Aman Gupta, Sathiya Keerthi, and Rahul Mazumder. Quantease: Optimization-based quantization for language models – an efficient and intuitive algorithm. arXiv preprint arXiv:2309.01885, 2023.
  5. 5.Riade Benbaki, Wenyu Chen, Xiang Meng, Hussein Hazimeh, Natalia Ponomareva, Zhe Zhao, and Rahul Mazumder. Fast as chita: Neural network pruning with combinatorial optimization. In ICML, 2023.
  6. 6.Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A suite for analyzing large language models across training and scaling. arXiv preprint arXiv:2304.01373, 2023.
  7. 7.Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. What is the state of neural network pruning? In Proceedings of Machine Learning and Systems, 2020.
  8. 8.Michael Bommarito and Daniel Martin Katz. Gpt takes the bar exam. arXiv preprint arXiv:2212.14402, 2022.
  9. 9.Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. Understanding and overcoming the challenges of efficient transformer quantization. arXiv:2109.12948, 2021.
  10. 10.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  11. 11.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  12. 12.Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. The lottery ticket hypothesis for pre-trained bert networks. NeurIPS, 2020.
  13. 13.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  14. 14.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009.
  16. 16.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM.int8(): 8-bit matrix multiplication for transformers at scale. In NeurIPS, 2022.
  17. 17.Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless llm weight compression. arXiv preprint arXiv:2306.03078, 2023.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanove. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  19. 19.Guneet S. Dhillon, Kamyar Azizzadenesheli, Zachary C. Lipton, Jeremy Bernstein, Jean Kossaifi, Aran Khanna, and Anima Anandkumar. Stochastic activation pruning for robust adversarial defense. In International Conference on Learning Representations, 2018.
  20. 20.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  21. 21.Abhimanyu Dubey, Moitreya Chatterjee, and Narendra Ahuja. Coreset-based neural network compression. In ECCV, 2018.
  22. 22.Christoforos Nalmpantis Elena Voita, Javier Ferrando. Neurons in large language models: Dead, n-gram, positional. arXiv preprint arXiv:2309.04827, 2023.
  23. 23.Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. Rigging the lottery: Making all tickets winners. In ICML, 2020.
  24. 24.Angela Fan, Edouard Grave, and Armand Joulin. Reducing transformer depth on demand with structured dropout. In International Conference on Learning Representations, 2020.
  25. 25.Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang. Depgraph: Towards any structural pruning. In Conference on Computer Vision and Pattern Recognition, 2023.
  26. 26.Jonathan Frankle and Carbin Michael. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In ICLR, 2019.
  27. 27.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. Stabilizing the lottery ticket hypothesis. In ICML, 2020.
  28. 28.Elias Frantar and Dan Alistarh. Spdy: Accurate pruning with speedup guarantees. In ICML, 2022.
  29. 29.Elias Frantar and Dan Alistarh. SparseGPT: Massive language models can be accurately pruned in one-shot. In ICML, 2023.
  30. 30.Elias Frantar, Sidak Pal Singh, and Dan Alistarh. Optimal Brain Compression: A framework for accurate post-training quantization and pruning. In NeurIPS, 2022.
  31. 31.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training compression for generative pretrained transformers. In ICLR, 2023a.
  32. 32.Elias Frantar, Carlos Riquelme, Neil Houlsby, Dan Alistarh, and Utku Evci. Scaling laws for sparsely-connected foundation models. arXiv preprint arXiv:2309.08520, 2023b.
  33. 33.Advait Gadhikar, Sohom Mukherjee, and Rebekka Burkholz. Why random pruning is all we need to start sparse. In ICML, 2023.
  34. 34.Trevor Gale, Erich Elsen, and Sara Hooker. The state of sparsity in deep neural networks. In ICML, 2019.
  35. 35.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPoFi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation. Version v0. 0.1. Sept, 2021.
  36. 36.Song Han, Jeff Pool, John Tran, and William J Dally. Learning both weights and connections for efficient neural networks. In NeurIPS, 2015.
  37. 37.Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In ICLR, 2016.
  38. 38.Babak Hassibi, David G Stork, and Gregory J Wolff. Optimal brain surgeon and general network pruning. In IEEE International Conference on Neural Networks, 1993.
  39. 39.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In ICLR, 2021.
  40. 40.Duc Hoang, Shiwei Liu, Radu Marculescu, and Zhangyang Wang. Revisiting pruning at initialization through the lens of ramanujan graph? In ICLR, 2023.
  41. 41.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, and et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  42. 42.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2021.
  43. 43.Hengyuan Hu, Rui Peng, Yu-Wing Tai, and Chi-Keung Tang. Network trimming: A data-driven neuron pruning approach towards efficient deep architectures. arXiv preprint arXiv:1607.03250, 2016.
  44. 44.Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Seffi Naor, and Daniel Soudry. Accelerated sparse neural training: A provable and efficient method to find N:M transposable masks. In NeurIPS, 2021.
  45. 45.Tian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle, and Gintare Karolina Dziugaite. Pruning’s effect on generalization through the lens of training and regularization. In NeurIPS, 2022.
  46. 46.Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky. Bert busters: Outlier dimensions that disrupt transformers. In ACL Findings, 2021.
  47. 47.Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi. Soft threshold weight reparameterization for learnable sparsity. In ICML, 2020.
  48. 48.Denis Kuznedelev, Eldar Kurtic, Eugenia Iofinova, Elias Frantar, Alexandra Peste, and Dan Alistarh. Accurate neural network pruning requires rethinking sparse optimization. arXiv preprint arXiv:2308.02060, 2023.
  49. 49.Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. A fast post-training pruning framework for transformers. In NeurIPS, 2022.
  50. 50.Yann LeCun, John S Denker, and Sara A Solla. Optimal brain damage. In NeurIPS, 1989.
  51. 51.Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr. Snip: Single-shot network pruning based on connection sensitivity. In ICLR, 2018.
  52. 52.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  53. 53.Shiwei Liu, Tianlong Chen, Zhenyu Zhang, Xuxi Chen, Tianjin Huang, Ajay Jaiswal, and Zhangyang Wang. Sparsity may cry: Let us fail (current) sparse neural networks together! In ICLR, 2023a.
  54. 54.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In ICCV, 2017.
  55. 55.Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. Rethinking the value of network pruning. In ICLR, 2019.
  56. 56.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In CVPR, 2022.
  57. 57.Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, and Beidi Chen. Deja vu: Contextual sparsity for efficient llms at inference time. In ICML, 2023b.
  58. 58.Christos Louizos, Max Welling, and Diederik P. Kingma. Learning sparse neural networks through l0 regularization. In ICLR, 2018.
  59. 59.Ziyang Luo, Artur Kulmizev, and Xiaoxi Mao. Positional artefacts propagate through masked language model embeddings. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, August 2021.
  60. 60.Xinyin Ma, Gongfan Fang, and Xinchao Wang. Llm-pruner: On the structural pruning of large language models. arXiv preprint arXiv:2305.11627, 2023.
  61. 61.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  62. 62.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  63. 63.Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.
  64. 64.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. Pruning convolutional neural networks for resource efficient inference. In ICLR, 2017.
  65. 65.Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Froscio, and Jan Kautz Kautz. Importance estimation for neural network pruning. In CVPR, 2019.
  66. 66.Azade Nova, Hanjun Dai, and Dale Schuurmans. Gradient-free structured pruning with unlabeled data. In International Conference on Machine Learning, 2023.
  67. 67.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  68. 68.Mansheej Paul, Feng Chen, Brett W. Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite. Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask? In ICLR, 2023.
  69. 69.Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh. AC/DC: Alternating compressed/decompressed training of deep neural networks. In NeurIPS, 2021.
  70. 70.Giovanni Puccetti, Anna Rogers, Aleksandr Drozd, and Felice Dell’Orletta. Outliers dimensions that disrupt transformers are driven by frequency. arXiv preprint arXiv:2205.11380, 2022.
  71. 71.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020.
  72. 72.Siyu Ren and Kenny Q. Zhu. Pruning pre-trained language models with principled importance and self-regularization. In ACL, 2023.
  73. 73.Alex Renda, Jonathan Frankle, and Michael Carbin. Comparing rewinding and fine-tuning in neural network pruning. In ICLR, 2020.
  74. 74.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. arXiv preprint arXiv:1907.10641, 2019.
  75. 75.Victor Sanh, Thomas Wolf, and Alexander M. Rush. Movement pruning: Adaptive sparsity by fine-tuning. In NeurIPS, 2020.
  76. 76.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  77. 77.Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? arXiv preprint arXiv:2304.15004, 2023.
  78. 78.Maying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao, Jianna Liu, and Jose M Alvarez. Structural pruning via latency-saliency knapsack. NeurIPS, 2022.
  79. 79.Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, and et al. High-throughput generative inference of large language models with a single gpu. In ICML, 2023.
  80. 80.Sidak Pal Singh and Dan Alistarh. Woodfisher: Efficient second-order approximation for neural network compression. In NeurIPS, 2020.
  81. 81.William Timkey and Marten van Schijndel. All bark and no bite: Rogue dimensions in transformer language models obscure representational quality. arXiv:2109.04404, 2021.
  82. 82.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  83. 83.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  84. 84.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  85. 85.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, and et al. Emergent abilities of large language models. In Transactions on Machine Learning Research, 2022a.
  86. 86.Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. Outlier suppression: Pushing the limit of low-bit transformer language models. In NeurIPS, 2022b.
  87. 87.Mengzhou Xia, Zexuan Zhong, and Danqi Chen. Structured pruning learns compact and accurate models. In Association for Computational Linguistics (ACL), 2022.
  88. 88.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In ICML, 2023.
  89. 89.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  90. 90.Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao. Platon: Pruning large transformer models with upper confidence bound of weight importance. In ICML, 2021.
  91. 91.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  92. 92.Yefan Zhou, Yaoqing Yang, Arin Chang, and Michael W. Mahoney. A three-regime model of network pruning. In ICML, 2023.
  93. 93.Michael Zhu and Suyog Gupta. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878, 2017.

Citation

MLA
Sun, M., et al. “A Simple and Effective Pruning Approach for Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2306.11695v3.
APA
Sun, M., Liu, Z., Bair, A., & Kolter, J. Z. (2023). A Simple and Effective Pruning Approach for Large Language Models. arXiv. http://arxiv.org/abs/2306.11695v3
Chicago
Sun, M., Z. Liu, A. Bair, and J. Z. Kolter. 2023. “A Simple and Effective Pruning Approach for Large Language Models”. arXiv. http://arxiv.org/abs/2306.11695v3.
Harvard
Sun, M. et al. (2023) “A Simple and Effective Pruning Approach for Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.11695v3.
Vancouver
1. Sun M, Liu Z, Bair A, Kolter JZ (2023) A Simple and Effective Pruning Approach for Large Language Models. arXiv

BibTeX

@article{sun2023simple,
  title = {A Simple and Effective Pruning Approach for Large Language Models},
  author = {Sun, Mingjie and Liu, Zhuang and Bair, Anna and Kolter, J. Zico},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.11695v3},
  eprint = {2306.11695}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors