SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Elias FrantarDan Alistarh

article2023ICML1,206 citations

Proposes SparseGPT, a one-shot pruning method that reduces 175-billion-parameter language models to over 50% sparsity in under five hours without requiring retraining or sacrificing accuracy.

Listen

Large Language Models deliver state-of-the-art performance across diverse natural language tasks but are exceptionally expensive to deploy. Top-performing models with roughly 175 billion parameters require hundreds of gigabytes of memory and multiple high-end accelerators just for inference. While pruning—the removal of redundant model weights—is an established compression strategy, existing post-training methods require massive retraining to recover accuracy or scale poorly to massive models. As a result, pruning has remained impractical for models containing tens to hundreds of billions of parameters.

The article demonstrates that massive generative pretrained models can be pruned accurately in a single step without any retraining or fine-tuning. It evaluates a new post-training compression method, called SparseGPT, designed to efficiently prune modern transformer models at the scale of 10 to 100+ billion parameters while maintaining near-original model accuracy.

The approach formulates layer-by-layer weight pruning as a large-scale regression problem. By synchronizing inverse Hessian matrices across weight rows and applying adaptive mask selection across blocks of columns, the method optimizes weight updates locally without computing global network gradients. Credibility is supported by end-to-end evaluations on the largest publicly available model families (OPT up to 175B and BLOOM at 176B) using only 128 generic calibration samples (2,048 tokens each) and executing on a single NVIDIA A100 GPU. Evaluations benchmark perplexity on standard text corpora and accuracy across multiple zero-shot evaluation tasks.

The findings establish that SparseGPT can prune 50% to 60% of weights from 175-billion-parameter models in under 4.5 hours with negligible loss in accuracy and perplexity. Crucially, larger models prove significantly more compressible than smaller ones, dropping virtually zero accuracy at 50% sparsity. Standard baseline approaches collapse completely beyond 10% to 30% sparsity, whereas SparseGPT removes over 100 billion parameters successfully. Furthermore, the method extends seamlessly to hardware-friendly semi-structured patterns (such as 2:4 sparsity) and combines with 4-bit weight quantization in a single pass to outperform standalone 3-bit quantization.

These results demonstrate that the extreme computational and financial overhead of running massive generative models can be reduced substantially post-training. Organizations can lower memory footprints and operational costs while maintaining model accuracy without expensive retraining cycles. Practical inference benchmarks show 1.54x to 1.82x speedups on CPUs and GPUs, confirming that substantial operational savings are immediately achievable.

Decision-makers should consider adopting single-pass pruning workflows alongside quantization pipelines to compress production models. When applying structured 2:4 sparsity, sensitivity analysis indicates that pruning earlier layers while retaining later layers provides the best accuracy-performance trade-off. Future efforts should focus on engineering customized sparse inference kernels and investigating the theoretical mechanisms behind why larger models compress more easily.

The evaluation relies on a limited calibration set of 128 text segments and local layer-wise approximations rather than global optimization. While zero-shot task metrics can exhibit variance across individual benchmarks, perplexity evaluations across multiple datasets demonstrate high consistency and robust statistical confidence across the reported results.

Cover for SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Abstract

We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy. This is achieved via a new pruning method called SparseGPT, specifically designed to work efficiently and accurately on massive GPT-family models. We can execute SparseGPT on the largest available open-source models, OPT-175B and BLOOM-176B, in under 4.5 hours, and can reach 60% unstructured sparsity with negligible increase in perplexity: remarkably, more than 100 billion weights from these models can be ignored at inference time. SparseGPT generalizes to semi-structured (2:4 and 4:8) patterns, and is compatible with weight quantization approaches. The code is available at: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 The SparseGPT Algorithm
  • 3.1 Fast Approximate Reconstruction
  • 3.2 Adaptive Mask Selection
  • 3.3 Extension to Semi-Structured Sparsity
  • 3.4 Full Algorithm Pseudocode
  • 3.5 Joint Sparsification & Quantization
  • 4 Experiments
  • 4.1 Results
  • 5 Related Work
  • 6 Discussion
  • 7 Acknowledgements
  • References
  • A Ablation Studies
  • A.1 Approximation Quality
  • B Evaluation Details
  • C Additional Results
  • D Partial 2:4 Results
  • E Sparsity Acceleration

Knowls

  1. Knowl 1 — The SparseGPT Algorithm for One-Shot Post-Training Pruning

    algorithm

    SparseGPT is a post-training, one-shot pruning algorithm designed to prune large language models with tens to hundreds of billions of parameters without any retraining or fine-tuning. It compresses the model layer-by-layer by solving a localized squared error minimization problem for each layer ℓ\ell with weight matrix Wℓ∈Rdrow×dcolW_\ell \in \mathbb{R}^{d_{\text{row}} \times d_{\text{col}}} and layer inputs Xℓ∈Rdcol×nX_\ell \in \mathbb{R}^{d_{\text{col}} \times n}:

    arg⁡min⁡Mℓ,W^ℓ∥WℓXℓ−(Mℓ⊙W^ℓ)Xℓ∥22\arg\min_{M_\ell, \widehat{W}_\ell} \|W_\ell X_\ell - (M_\ell \odot \widehat{W}_\ell) X_\ell\|_2^2

    where Mℓ∈{0,1}drow×dcolM_\ell \in \{0, 1\}^{d_{\text{row}} \times d_{\text{col}}} is the binary pruning mask and W^ℓ\widehat{W}_\ell represents the updated unpruned weights. SparseGPT iterates column-by-column through WW, freezes pruned weights, and applies error compensation updates to all remaining unpruned columns using precomputed inverse Hessian information. It utilizes batched lazy updates with block size BB and adaptive mask selection with block size BsB_s to maximize compute efficiency and hardware utilization, reducing overall layer complexity from O(drow⋅dcol3)O(d_{\text{row}} \cdot d_{\text{col}}^3) to O(dcol3+drowdcol2)=O(dhidden3)O(d_{\text{col}}^3 + d_{\text{row}} d_{\text{col}}^2) = O(d_{\text{hidden}}^3).

    Input: Layer weight matrix W∈Rdrow×dcolW \in \mathbb{R}^{d_{\text{row}} \times d_{\text{col}}}, sample inputs XX, dampening factor λ=0.01\lambda = 0.01, batch update block size B=128B = 128, mask selection block size Bs=128B_s = 128, target unstructured sparsity fraction pp
    Output: Sparsified weight matrix WW with pp sparsity
    H←2XX⊤+λ⋅mean(diag(2XX⊤))⋅IH \leftarrow 2 X X^\top + \lambda \cdot \text{mean}(\text{diag}(2 X X^\top)) \cdot I
    H−1←Cholesky(H−1)⊤H^{-1} \leftarrow \text{Cholesky}(H^{-1})^\top
    M←1drow×dcolM \leftarrow \mathbf{1}_{d_{\text{row}} \times d_{\text{col}}}
    E←0drow×BE \leftarrow \mathbf{0}_{d_{\text{row}} \times B}
    for i=0,B,2B,…,dcol−1i = 0, B, 2B, \dots, d_{\text{col}} - 1 do
        for j=i,…,i+B−1j = i, \dots, i + B - 1 do
            if j mod Bs==0j \bmod B_s == 0 then
                M:,j:(j+Bs)←M_{:, j:(j+B_s)} \leftarrow mask preserving (1−p)(1 - p) fraction of weights wc∈W:,j:(j+Bs)w_c \in W_{:, j:(j+B_s)} with largest wc2/[H−1]cc2w_c^2 / [H^{-1}]_{cc}^2
            end if
            E:,j−i←W:,j/[H−1]jjE_{:, j-i} \leftarrow W_{:, j} / [H^{-1}]_{jj}
            E:,j−i←(1−M:,j)⊙E:,j−iE_{:, j-i} \leftarrow (1 - M_{:, j}) \odot E_{:, j-i}
            W:,j:(i+B)←W:,j:(i+B)−E:,j−i⋅Hj,j:(i+B)−1W_{:, j:(i+B)} \leftarrow W_{:, j:(i+B)} - E_{:, j-i} \cdot H^{-1}_{j, j:(i+B)}
        end for
        W:,(i+B):←W:,(i+B):−E⋅Hi:(i+B),(i+B):−1W_{:, (i+B):} \leftarrow W_{:, (i+B):} - E \cdot H^{-1}_{i:(i+B), (i+B):}
    end for
    W←W⊙MW \leftarrow W \odot M
    return WW
  2. Knowl 2 — Hessian Synchronization via Recursive Gaussian Elimination

    model/method

    In layer-wise post-training pruning, exact unpruned weight reconstruction for a row wiw_i given a mask MiM_i requires inverting the masked Hessian HMi=XMiXMi⊤H_{M_i} = X_{M_i} X_{M_i}^\top. Because each row typically has a distinct sparsity mask MiM_i, and (HMi)−1≠(H−1)Mi(H_{M_i})^{-1} \neq (H^{-1})_{M_i}, exact reconstruction necessitates drowd_{\text{row}} separate matrix inversions of size O(dcol×dcol)O(d_{\text{col}} \times d_{\text{col}}), scaling at O(drow⋅dcol3)O(d_{\text{row}} \cdot d_{\text{col}}^3).

    SparseGPT resolves this by defining a sequence of nested index subsets Uj⊆{1,…,dcol}U_j \subseteq \{1, \dots, d_{\text{col}}\} defined recursively as:

    Uj+1=Uj∖{j},U1={1,…,dcol}U_{j+1} = U_j \setminus \{j\}, \quad U_1 = \{1, \dots, d_{\text{col}}\}

    Rather than computing row-specific inverse Hessians, all rows share the same sequence of subset inverse Hessians (HUj)−1=((XX⊤)Uj)−1(H_{U_j})^{-1} = ((X X^\top)_{U_j})^{-1}. The updated inverse (HUj+1)−1(H_{U_{j+1}})^{-1} is computed recursively in O(dcol2)O(d_{\text{col}}^2) time by removing the first row and column corresponding to feature jj from B=(HUj)−1B = (H_{U_j})^{-1} via one step of Gaussian elimination:

    (HUj+1)−1=(B−1[B]11B:,1B1,:)2:,2:(H_{U_{j+1}})^{-1} = \left( B - \frac{1}{[B]_{11}} B_{:, 1} B_{1, :} \right)_{2:, 2:}

    with initial condition (HU1)−1=H−1=(XX⊤+λI)−1(H_{U_1})^{-1} = H^{-1} = (X X^\top + \lambda I)^{-1}. This recursion produces the complete sequence of dcold_{\text{col}} inverse Hessians in O(dcol3)O(d_{\text{col}}^3) total time, identical in complexity to a single matrix inversion.

  3. Knowl 3 — Adaptive Pruning Mask Selection via Iterative Blocking

    model/method

    Pruning weights changes the values of remaining unpruned weights through error compensation updates, making a static mask selected upfront suboptimal. However, selecting a pruning mask strictly column-by-column enforces an equal sparsity ratio across all columns, which degrades model accuracy because large language models possess a small set of highly sensitive outlier activation channels.

    SparseGPT applies adaptive mask selection via iterative blocking with block size Bs=128B_s = 128. For each block of BsB_s columns W:,j:(j+Bs)W_{:, j:(j+B_s)}, the pruning mask M:,j:(j+Bs)M_{:, j:(j+B_s)} is selected simultaneously across the entire block based on the Optimal Brain Surgeon (OBS) error metric:

    εc=wc2[H−1]cc2\varepsilon_c = \frac{w_c^2}{[H^{-1}]_{cc}^2}

    where wcw_c is the current weight value (reflecting updates from all prior processed columns) and [H−1]cc[H^{-1}]_{cc} is the corresponding diagonal entry from the inverse Hessian sequence. Within each row of the block, the p%p\% of weights with the smallest error metric εc\varepsilon_c are selected for pruning. This allows non-uniform distribution of pruned weights across individual columns within each block while incorporating prior compensation updates.

  4. Knowl 4 — Unified Joint Post-Training Sparsification and Quantization

    model/method

    SparseGPT frames pruning as a column-wise greedy weight freezing process, which shares the underlying Cholesky decomposition of H−1H^{-1} and lazy batched updates with the GPTQ quantization algorithm. SparseGPT merges weight sparsification and weight quantization into a single compression pass with zero extra computational overhead.

    In this joint procedure, all weights that are not zeroed by the pruning mask MM are quantized to the nearest grid point via a quantization function quant(⋅)\text{quant}(\cdot). The generalized error column E:,j−iE_{:, j-i} propagated to subsequent columns is computed as:

    E:,j−i=W:,j−M:,j⊙quant(W:,j)[H−1]jjE_{:, j-i} = \frac{W_{:, j} - M_{:, j} \odot \text{quant}(W_{:, j})}{[H^{-1}]_{jj}}

    where ⊙\odot is the element-wise product and [H−1]jj[H^{-1}]_{jj} is the diagonal entry of the Cholesky-factored inverse Hessian. By unifying both operations in one loop, later pruning decisions adapt dynamically to prior quantization errors, and later quantization steps compensate for prior pruning errors.

  5. Knowl 5 — Extension of SparseGPT to Semi-Structured N:M Sparsity

    model/method

    SparseGPT supports fine-grained semi-structured N:MN:M sparsity patterns, where every contiguous block of MM weights must contain exactly NN zeros (such as 2:42:4 or 4:84:8 sparsity patterns accelerated by NVIDIA Ampere Tensor Cores).

    To enforce an N:MN:M pattern, the adaptive mask selection block size is set to Bs=MB_s = M. For each row in the block of MM consecutive columns, the algorithm evaluates the Optimal Brain Surgeon pruning error:

    εm=wm2[H−1]mm\varepsilon_m = \frac{w_m^2}{[H^{-1}]_{mm}}

    for each of the MM candidate weights and sets the NN weights with the lowest error εm\varepsilon_m to zero in the mask M:,j:(j+M)M_{:, j:(j+M)}. The remaining M−NM - N weights in each group are retained and subsequently updated to compensate for the removed weights.

  6. Knowl 6 — Perplexity Scaling and Sparsification Across Model Sizes on the OPT Family

    data/table

    One-shot post-training pruning performance was evaluated across the entire OPT model family (from 125M to 175B parameters) on the raw-WikiText2 language modeling benchmark (measuring test perplexity, where lower is better). Pruning was applied uniformly across all linear layers (excluding embeddings and head) to 50% unstructured sparsity, 4:8 semi-structured sparsity, and 2:4 semi-structured sparsity, and compared against Magnitude Pruning (50%) and AdaPrune (50%).

    Model Variant 125M 350M 1.3B 2.7B 6.7B 13B 30B 66B 175B
    Dense (0%) 27.66 22.00 14.62 12.47 10.86 10.13 9.56 9.34 8.35
    Magnitude (50%) 193.0 97.80 1.7×1041.7 \times 10^4 265.0 969.0 1.2×1041.2 \times 10^4 168.0 4.2×1034.2 \times 10^3 4.3×1044.3 \times 10^4
    AdaPrune (50%) 58.66 48.46 32.52 – – – – – –
    SparseGPT (50%) 36.85 31.58 17.46 13.48 11.55 11.17 9.79 9.32 8.21
    SparseGPT (4:8) – – – 14.98 12.56 11.77 10.30 9.65 8.45
    SparseGPT (2:4) – – – 17.18 14.20 12.96 10.90 10.09 8.74

    The results establish that larger models exhibit greater compressibility: the perplexity gap between dense and sparsified models decreases monotonically as parameter count grows. At 175B scale, SparseGPT achieves 50% unstructured sparsity with a perplexity of 8.21 (improving slightly upon the 8.35 dense baseline on this dataset) and drops only 0.10 PPL for 4:8 sparsity (8.45) and 0.39 PPL for 2:4 sparsity (8.74). In contrast, magnitude pruning collapses completely across all model sizes (e.g., perplexity of 4.3×1044.3 \times 10^4 on OPT-175B).

  7. Knowl 7 — Zero-Shot Task Accuracy of Sparsified OPT-175B

    data/table

    Zero-shot accuracy was measured for OPT-175B across five standard NLP evaluation benchmarks: LAMBADA (word prediction requiring broad context), PIQA (physical reasoning), ARC-Easy (elementary science questions), ARC-Challenge (difficult science questions), and StoryCloze. Evaluations compared the dense baseline against 50% magnitude pruning, 50% unstructured SparseGPT, 4:8 semi-structured SparseGPT, and 2:4 semi-structured SparseGPT.

    Method Sparsity LAMBADA PIQA ARC-e ARC-c StoryCloze Average
    Dense Baseline 0% 75.59% 81.07% 71.04% 43.94% 79.82% 70.29%
    Magnitude 50% 0.02% 54.73% 28.03% 25.60% 47.10% 31.10%
    SparseGPT 50% 78.47% 80.63% 70.45% 43.94% 79.12% 70.52%
    SparseGPT 4:8 80.30% 79.54% 68.85% 41.30% 78.10% 69.62%
    SparseGPT 2:4 80.92% 79.54% 68.77% 39.25% 77.08% 69.11%

    SparseGPT preserves zero-shot performance within 1.2% of the dense model across all evaluated patterns (averaging 70.52% for 50% unstructured, 69.62% for 4:8, and 69.11% for 2:4, compared to 70.29% for dense), whereas magnitude pruning collapses to near-random performance (31.10% average).

  8. Knowl 8 — Joint 50% Sparsity and 4-Bit Quantization Scaling vs Size-Equivalent Precision

    empirical result

    Storing a model with 50% weight sparsity and 4-bit integer quantization (using non-zero values accompanied by a 1-bit presence bitmask) requires 3.0 bits per weight in total memory footprint. Comparing joint 50% unstructured sparsity + 4-bit quantization against size-equivalent 3-bit post-training quantization (using GPTQ) across OPT models with ≥2.7B\ge 2.7\text{B} parameters shows that joint sparsification and quantization yields lower raw-WikiText2 perplexity at every scale.

    On OPT-175B, 50% sparsity combined with 4-bit quantization achieves a perplexity of 8.29, outperforming pure 3-bit GPTQ (8.68 PPL). Furthermore, combining 2:4 and 4:8 semi-structured sparsity with 4-bit quantization on OPT-175B results in perplexities of 8.55 and 8.85, respectively, indicating that 4-bit quantization introduces only an ≈0.1\approx 0.1 perplexity increase over non-quantized semi-structured sparse models. At more extreme compression, 50% sparsity + 3-bit quantization (equivalent to 2.5 bits/weight) achieves 8.60 PPL on OPT-175B, outperforming 2.5-bit GPTQ (8.94 PPL).

  9. Knowl 9 — Layer Sensitivity and Single-Pass Partial N:M Sparsification

    empirical result

    In 175B-scale language models (OPT-175B and BLOOM-176B), layer sensitivity analysis for partial 2:4 semi-structured pruning reveals that later layers in the network are significantly more sensitive to sparsification than earlier layers. When sparsifying 2/3 of the total layers to 2:4 sparsity, skipping the final 1/3 of layers preserves the lowest perplexity (8.38 PPL on raw-WikiText2 for OPT-175B and 8.52 PPL for BLOOM-176B) compared to skipping the front 1/3 (8.78 and 9.17 PPL) or skipping specific layer types like attention or feed-forward projections.

    Because SparseGPT compresses layers strictly sequentially from input to output, this front-to-back resilience allows generating a sequence of progressively sparsified models (e.g., 1/2, 2/3, 3/4, 4/5 sparse fractions) in a single pruning execution by saving intermediate layer checkpoints and concatenating the first xx sparsified layers with the remaining nlayers−xn_{\text{layers}} - x dense layers.

  10. Knowl 10 — Inference Acceleration of Sparse LLMs on CPU and GPU

    data/table

    Real-world inference acceleration of SparseGPT-compressed models was benchmarked using existing deployment tools without model-specific custom kernel engineering.

    For CPU inference, end-to-end latency of unstructured sparse OPT-2.7B models was measured on an 18-core Intel Core i9-7980XE CPU (@ 2.60 GHz) using the DeepSparse engine for a batch size of 400 tokens:

    Unstructured Sparsity 40% 50% 60%
    End-to-End Speedup over Dense 1.57×\times 1.82×\times 2.16×\times

    For GPU inference, individual linear layer execution times of 2:4 semi-structured sparse matrices were evaluated on an NVIDIA Ampere GPU using NVIDIA CUTLASS kernels against dense cuBLAS kernels at batch size 2048 for the layer dimensions of OPT-175B:

    Layer Matrix Shape Q/K/V/Out FC1 FC2
    Dense Runtime (cuBLAS) 2.84 ms 10.26 ms 10.23 ms
    2:4 Sparse Runtime (CUTLASS) 1.59 ms 6.15 ms 6.64 ms
    Speedup 1.79×\times 1.67×\times 1.54×\times

    Across individual layers, 2:4 GPU acceleration delivers between 54%54\% and 79%79\% speedup over optimized dense kernels.

  11. Knowl 11 — SparseGPT Hyperparameter Robustness and Reconstruction Error Bound

    empirical result

    Ablation experiments on OPT-2.7B at 50% uniform sparsity identify key operational bounds for SparseGPT:

    1. Calibration Data Size: Perplexity improves rapidly as calibration data increases from 2 to 64 samples (each of 2048 tokens from the C4 dataset), flattening at 128 samples, which provides stable performance across model scales.
    2. Hessian Dampening: Adding a diagonal dampening term λ⋅mean(diag(H))⋅I\lambda \cdot \text{mean}(\text{diag}(H)) \cdot I to the Hessian matrix with λ=0.01\lambda = 0.01 (1%) prevents numerical instability during matrix inversions without degrading reconstruction accuracy.
    3. Calibration Randomness: Across 5 independent calibration samplings with different random seeds, OPT-2.7B achieves a raw-WikiText2 perplexity of 13.52±0.07513.52 \pm 0.075 (mean ±\pm standard deviation), showing negligible sensitivity to calibration sampling.
    4. Reconstruction Quality: The layer-wise squared reconstruction error of SparseGPT's partial update scheme is on average only ≈20%\approx 20\% higher than that of exact, computationally intractable Optimal Brain Compression (OBC) reconstruction on OPT-2.7B, dropping to ≈10%\approx 10\% on large fully-connected layers.

Coverage note — None. All primary theoretical, algorithmic, experimental, and ablation contributions presented in the paper have been represented as standalone knowls.

References

  1. 1.Blumensath, T. and Davies, M. E. Iterative thresholding for sparse approximations. Journal of Fourier Analysis and Applications, 14(5-6):629–654, 2008.
  2. 2.Boratko, M., Padigela, H., Mikkilineni, D., Yuvraj, P., Das, R., McCallum, A., Chang, M., Fokoue-Nkoutche, A., Kapanipathi, P., Mattei, N., et al. A systematic classification of knowledge, reasoning, and context within the ARC dataset. arXiv preprint arXiv:1806.00358, 2018.
  3. 3.Dettmers, T. and Zettlemoyer, L. The case for 4-bit precision: k-bit inference scaling laws. arXiv preprint arXiv:2212.09720, 2022.
  4. 4.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. LLM.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  5. 5.EleutherAI. EleutherAI LM Evaluation Harness, 2022. URL https://github.com/EleutherAI/lm-evaluation-harness.
  6. 6.Elsen, E., Dukhan, M., Gale, T., and Simonyan, K. Fast sparse convnets. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  7. 7.Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning (ICML), 2020.
  8. 8.Frantar, E. and Alistarh, D. SPDY: Accurate pruning with speedup guarantees. arXiv preprint arXiv:2201.13096, 2022.
  9. 9.Frantar, E., Kurtic, E., and Alistarh, D. M-FAC: Efficient matrix-free approximations of second-order information. In Conference on Neural Information Processing Systems (NeurIPS), 2021.
  10. 10.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. GPTQ: Accurate post-training compression for generative pretrained transformers. arXiv preprint arXiv:2210.17323, 2022a.
  11. 11.Frantar, E., Singh, S. P., and Alistarh, D. Optimal Brain Compression: A framework for accurate post-training quantization and pruning. arXiv preprint arXiv:2208.11580, 2022b. Accepted to NeurIPS 2022, to appear.
  12. 12.Gale, T., Elsen, E., and Hooker, S. The state of sparsity in deep neural networks. In International Conference on Machine Learning (ICML), 2019.
  13. 13.Hagiwara, M. A simple and effective method for removal of hidden units and weights. Neurocomputing, 6(2):207 – 218, 1994. ISSN 0925-2312. Backpropagation, Part IV.
  14. 14.Han, S., Pool, J., Tran, J., and Dally, W. J. Learning both weights and connections for efficient neural networks. In Conference on Neural Information Processing Systems (NeurIPS), 2015.
  15. 15.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations (ICLR), 2016.
  16. 16.Hassibi, B., Stork, D. G., and Wolff, G. J. Optimal brain surgeon and general network pruning. In IEEE International Conference on Neural Networks, 1993.
  17. 17.He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S. AMC: AutoML for model compression and acceleration on mobile devices. In European Conference on Computer Vision (ECCV), 2018.
  18. 18.Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. arXiv preprint arXiv:2102.00554, 2021.
  19. 19.Hubara, I., Chmiel, B., Island, M., Banner, R., Naor, S., and Soudry, D. Accelerated sparse neural training: A provable and efficient method to find N:M transposable masks. In Conference on Neural Information Processing Systems (NeurIPS), 2021a.
  20. 20.Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D. Accurate post training quantization with small calibration sets. In International Conference on Machine Learning (ICML), 2021b.
  21. 21.HuggingFace. HuggingFace Perplexity Calculation, 2022. URL https://huggingface.co/docs/transformers/perplexity.
  22. 22.Kingdon, J. Hypothesising Neural Nets, pp. 81–106. Springer London, London, 1997. ISBN 978-1-4471-0949-5. doi: 10.1007/978-1-4471-0949-5 5.
  23. 23.Kurtic, E. and Alistarh, D. Gmp*: Well-tuned global magnitude pruning can outperform most bert-pruning methods. arXiv preprint arXiv:2210.06384, 2022.
  24. 24.Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D. The Optimal BERT Surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259, 2022.
  25. 25.Kurtz, M., Kopinsky, J., Gelashvili, R., Matveev, A., Carr, J., Goin, M., Leiserson, W., Moore, S., Nell, B., Shavit, N., and Alistarh, D. Inducing and exploiting activation sparsity for fast inference on deep neural networks. In International Conference on Machine Learning (ICML), 2020.
  26. 26.Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A. A fast post-training pruning framework for transformers. arXiv preprint arXiv:2204.09656, 2022.
  27. 27.LeCun, Y., Denker, J. S., and Solla, S. A. Optimal brain damage. In Conference on Neural Information Processing Systems (NeurIPS), 1989.
  28. 28.Li, Y., Gong, R., Tan, X., Yang, Y., Hu, P., Zhang, Q., Yu, F., Wang, W., and Gu, S. BRECQ: Pushing the limit of post-training quantization by block reconstruction. In International Conference on Learning Representations (ICLR), 2021.
  29. 29.Liu, L., Zhang, S., Kuang, Z., Zhou, A., Xue, J.-H., Wang, X., Chen, Y., Yang, W., Liao, Q., and Zhang, W. Group fisher pruning for practical network compression. In International Conference on Machine Learning (ICML), 2021.
  30. 30.Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B. The penn treebank: Annotating predicate argument structure. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994.
  31. 31.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  32. 32.Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.
  33. 33.Mostafazadeh, N., Roth, M., Louis, A., Chambers, N., and Allen, J. Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pp. 46–51, 2017.
  34. 34.Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T. Up or down? Adaptive rounding for post-training quantization. In International Conference on Machine Learning (ICML), 2020.
  35. 35.NeuralMagic. DeepSparse, 2022. URL https://github.com/neuralmagic/deepsparse.
  36. 36.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The LAMBADA dataset: Word prediction requiring a broad discourse context. arXiv preprint arXiv:1606.06031, 2016.
  37. 37.Park, G., Park, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D. nuQmm: Quantized matmul for efficient inference of large-scale generative language models. arXiv preprint arXiv:2206.09557, 2022a.
  38. 38.Park, M., You, J., Nagel, M., and Chang, S. Quadapter: Adapter for gpt-2 quantization. arXiv preprint arXiv:2211.16912, 2022b.
  39. 39.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. Pytorch: An imperative style, high-performance deep learning library. In Conference on Neural Information Processing Systems (NeurIPS), 2019.
  40. 40.Peste, A., Iofinova, E., Vladu, A., and Alistarh, D. AC/DC: Alternating compressed/decompressed training of deep neural networks. In Conference on Neural Information Processing Systems (NeurIPS), 2021.
  41. 41.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 (140):1–67, 2020.
  42. 42.Sanh, V., Wolf, T., and Rush, A. M. Movement pruning: Adaptive sparsity by fine-tuning. arXiv preprint arXiv:2005.07683, 2020.
  43. 43.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagne, R., Luccioni, A. S., Yvon, F., Galle, M., et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  44. 44.Singh, S. P. and Alistarh, D. WoodFisher: Efficient second-order approximation for neural network compression. In Conference on Neural Information Processing Systems (NeurIPS), 2020.
  45. 45.Tata, S. and Patel, J. M. PiQA: An algebra for querying protein data sets. In International Conference on Scientific and Statistical Database Management, 2003.
  46. 46.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  47. 47.Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438, 2022.
  48. 48.Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y. ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. arXiv preprint arXiv:2206.01861, 2022.
  49. 49.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  50. 50.Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H. Learning N:M fine-grained structured sparse neural networks from scratch. In International Conference on Learning Representations (ICLR), 2021.
  51. 51.Zhu, M. and Gupta, S. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878, 2017.

Citation

MLA
Frantar, E., and D. Alistarh. “SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot”. arXiv, 2023, http://arxiv.org/abs/2301.00774v3.
APA
Frantar, E., & Alistarh, D. (2023). SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot. arXiv. http://arxiv.org/abs/2301.00774v3
Chicago
Frantar, E., and D. Alistarh. 2023. “SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot”. arXiv. http://arxiv.org/abs/2301.00774v3.
Harvard
Frantar, E. and Alistarh, D. (2023) “SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.00774v3.
Vancouver
1. Frantar E, Alistarh D (2023) SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot. arXiv

BibTeX

@article{frantar2023sparsegpt,
  title = {SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot},
  author = {Frantar, Elias and Alistarh, Dan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.00774v3},
  eprint = {2301.00774}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/