SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Saleh AshkboosMaximilian L. CrociMarcelo Gennari Do NascimentoTorsten HoeflerJames Hensman

article2024ICLR334 citations

Introduces SliceGPT, a post-training compression method that exploits computational invariance to slice transformer weight matrices into smaller dense arrays, cutting parameters by up to 25% and accelerating inference on standard GPUs with minimal loss in zero-shot accuracy.

Listen

Deploying pre-trained large language models incurs significant financial and hardware costs because their massive parameter sizes demand extensive memory and multi-GPU infrastructure for inference. Existing compression techniques, such as traditional pruning or semi-structured sparsity, often require specialized sparse hardware libraries, add data structure overhead, or necessitate costly fine-tuning to recover lost accuracy.

The article demonstrates and evaluates SliceGPT, a post-training structured pruning method designed to reduce model size and accelerate inference without requiring custom sparse code or complex retraining. The scheme achieves this by exploiting a mathematical property called computational invariance to systematically delete entire rows and columns from weight matrices, effectively shrinking the network's internal embedding dimension.

The approach was evaluated across multiple open-source model families—including LLAMA-2 (up to 70B parameters), OPT (up to 66B parameters), and Phi-2 (2.7B parameters)—using standard calibration datasets on modern hardware configurations (Quadro RTX6000, A100, and H100 GPUs). The compression process relies on principal component analysis computed on a single GPU in a few hours, followed by optional lightweight recovery fine-tuning.

The analysis produced several key findings: First, SliceGPT successfully removed up to 25% of parameters across large models while maintaining 90% to 99% of original zero-shot task performance. Second, compressed models delivered substantial compute and memory savings; for LLAMA-2 70B, inference compute dropped to 64% on consumer-level GPUs (reducing the GPU requirement from 7 to 5) and 66% on A100 GPUs (reducing GPU count from 4 to 3). Third, sliced models achieved up to 1.55x throughput improvements at 25% slicing on H100 GPUs, and larger slicing levels enabled massive batch-size scaling. Finally, SliceGPT consistently outperformed competitive semi-structured sparsity schemes (such as 2:4 sparsity) in perplexity on large models while utilizing standard dense matrix operations.

These findings have direct operational and financial implications for enterprise AI deployment. Organizations can lower capital and operating expenditures by reducing the number of high-end GPUs needed for serving models, decreasing energy consumption, and improving response latency. Because SliceGPT produces smaller dense matrices rather than sparse formats, it integrates seamlessly into existing deployment pipelines without requiring custom sparse kernels or software rewrites.

Engineering and deployment teams should consider evaluating SliceGPT on large production models to lower hosting footprints. When deploying smaller networks (such as Phi-2 or models under 13B parameters), applying lightweight recovery fine-tuning with representative task data is recommended to restore accuracy. Furthermore, organizations can explore combining SliceGPT with complementary techniques, such as quantization, to capture cumulative efficiency gains.

Confidence in these findings is strong across large-scale architectures, supported by consistent empirical evaluations on standard benchmarks. However, decision-makers should note that smaller models (13B parameters or fewer) experience higher relative accuracy degradation from pruning than larger models. Additionally, optimal calibration requires selecting representative datasets and sufficient sequence lengths to avoid unintended performance drops.

Cover for SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Abstract

Large language models have become the cornerstone of natural language processing, but their use comes with substantial costs in terms of compute and memory resources. Sparsification provides a solution to alleviate these resource constraints, and recent works have shown that trained models can be sparsified post-hoc. Existing sparsification techniques face challenges as they need additional data structures and offer constrained speedup with current hardware. In this paper we present SliceGPT, a new post-training sparsification scheme which replaces each weight matrix with a smaller (dense) matrix, reducing the embedding dimension of the network. Through extensive experimentation, we show that SliceGPT can remove up to 25% of the model parameters (including embeddings) for LLAMA2-70B, OPT 66B and Phi-2 models while maintaining 99%, 99% and 90% zero-shot task performance of the dense model respectively. Our sliced models run on fewer GPUs and run faster without any additional code optimization: on 24GB consumer GPUs we reduce the total compute for inference on LLAMA2-70B to 64% of that of the dense model; on 40GB A100 GPUs we reduce it to 66%. We offer a new insight, computational invariance in transformer networks, which enables SliceGPT and we hope it will inspire and enable future avenues to reduce memory and computation demands for pre-trained models. Code is available at: this https URL

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Transformer Networks
  • 2.2 Related work
  • 3 SliceGPT
  • 3.1 Computational invariance in transformer networks
  • 3.2 LayerNorm Transformers can be converted to RMSNorm
  • 3.3 A transformation per block
  • 3.4 Slicing
  • 4 Experimental Validation
  • 4.1 Results
  • 5 Conclusion and Future Work
  • References
  • A Appendix
  • A.1 Proof of Equation
  • A.2 Single Precision Eigenvalue Calculation
  • A.3 Sensitivity to the calibration set size and sequence length
  • A.4 Spectrum Analysis of Llama- 2 and OPT Models
  • A.5 Detailed Zero-shot Results
  • A.6 Benchmarking Throughput Experiment
  • A.7 Benchmarking Inference Time of SliceGPT against SparseGPT

Knowls

  1. Knowl 1 — Computational Invariance of RMSNorm-Connected Transformer Networks

    theoretical result

    In a transformer network where layers are interconnected via Root Mean Square Normalization (RMSNorm) and residual connections, the end-to-end model computation is invariant under orthogonal basis transformations of the residual stream.

    Let Q∈RD×DQ \in \mathbb{R}^{D \times D} be an orthogonal matrix such that Q⊤Q=QQ⊤=IQ^\top Q = Q Q^\top = I, where DD is the embedding dimension. Because multiplying a row vector x∈RDx \in \mathbb{R}^D by QQ preserves its Euclidean norm (∥xQ∥=∥x∥\lVert x Q \rVert = \lVert x \rVert), the RMSNorm operation commutes with orthogonal transformations:

    RMSNorm(XQ)Q⊤=RMSNorm(X)\text{RMSNorm}(X Q) Q^\top = \text{RMSNorm}(X)

    for any signal matrix X∈RN×DX \in \mathbb{R}^{N \times D} with sequence length NN.

    Consequently, given an RMSNorm-connected transformer with embedding matrix WembdW_{\text{embd}}, linear layer blocks with input matrices WinℓW_{\text{in}}^\ell and biases binℓb_{\text{in}}^\ell, output matrices WoutℓW_{\text{out}}^\ell and biases boutℓb_{\text{out}}^\ell, and language modeling head WheadW_{\text{head}} and bias bheadb_{\text{head}}, the following transformed parameters produce output logits identical to the original network:

    W~embd=WembdQ\tilde{W}_{\text{embd}} = W_{\text{embd}} Q W~inℓ=Q⊤Winℓ\tilde{W}_{\text{in}}^\ell = Q^\top W_{\text{in}}^\ell W~outℓ=WoutℓQ\tilde{W}_{\text{out}}^\ell = W_{\text{out}}^\ell Q b~outℓ=Q⊤boutℓ\tilde{b}_{\text{out}}^\ell = Q^\top b_{\text{out}}^\ell W~head=Q⊤Whead\tilde{W}_{\text{head}} = Q^\top W_{\text{head}} b~inℓ=binℓ,b~head=bhead\tilde{b}_{\text{in}}^\ell = b_{\text{in}}^\ell, \quad \tilde{b}_{\text{head}} = b_{\text{head}}

  2. Knowl 2 — SliceGPT Post-Training Compression via Layerwise PCA and Slicing

    algorithm

    SliceGPT is a post-training structured pruning algorithm that reduces the embedding dimension DD of a transformer to Dsmall<DD_{\text{small}} < D by transforming the activations at each layer using Principal Component Analysis (PCA) and deleting the rows and columns corresponding to the least significant principal components.

    Because activation signals between different transformer blocks are not aligned, SliceGPT calculates a distinct orthogonal transformation matrix Qℓ∈RD×DQ_\ell \in \mathbb{R}^{D \times D} for each block ℓ\ell. To maintain mathematical consistency across blocks with different rotations, an explicit linear transformation Qℓ−1⊤QℓQ_{\ell-1}^\top Q_\ell is introduced into the residual skip-connection between block ℓ−1\ell-1 and block ℓ\ell.

    Input: Pre-trained transformer parameters WembdW_{\text{embd}}, {Winℓ,binℓ,Woutℓ,boutℓ}ℓ=1L\{W_{\text{in}}^\ell, b_{\text{in}}^\ell, W_{\text{out}}^\ell, b_{\text{out}}^\ell\}_{\ell=1}^L, Whead,bheadW_{\text{head}}, b_{\text{head}}; calibration dataset {si}i=1B\{s_i\}_{i=1}^B; target reduced dimension DsmallD_{\text{small}}
    Output: Compressed model with embedding dimension DsmallD_{\text{small}}
    1. Convert all LayerNorm operations to RMSNorm by absorbing mean subtraction M=I−1D11⊤M = I - \frac{1}{D}\mathbf{1}\mathbf{1}^\top into preceding linear output weights and diagonal gains diag(α)\text{diag}(\alpha) into succeeding linear input weights.
    2. Set Q0=IDQ_0 = I_D.
    3. for ℓ=1\ell = 1 to LL do:
        a. Feed calibration sequences through the network up to the ℓ\ell-th RMSNorm block to obtain activation signals Xℓ,i∈RN×DX_{\ell, i} \in \mathbb{R}^{N \times D} for each calibration sample ii.
        b. Compute the uncentered sample covariance matrix in double precision (FP64): Cℓ=∑i=1BXℓ,i⊤Xℓ,iC_\ell = \sum_{i=1}^B X_{\ell, i}^\top X_{\ell, i}.
        c. Compute the eigendecomposition Cℓ=QℓΛℓQℓ⊤C_\ell = Q_\ell \Lambda_\ell Q_\ell^\top, where Qℓ∈RD×DQ_\ell \in \mathbb{R}^{D \times D} contains orthonormal eigenvectors sorted by descending eigenvalues.
        d. Let Dproj∈RD×DsmallD_{\text{proj}} \in \mathbb{R}^{D \times D_{\text{small}}} be the truncation matrix containing the first DsmallD_{\text{small}} columns of the identity matrix IDI_D.
        e. Slice input weights for block ℓ\ell: W~inℓ=Dproj⊤Qℓ−1⊤Winℓ\tilde{W}_{\text{in}}^\ell = D_{\text{proj}}^\top Q_{\ell-1}^\top W_{\text{in}}^\ell.
        f. Slice output weights and bias for block ℓ\ell: W~outℓ=WoutℓQℓDproj\tilde{W}_{\text{out}}^\ell = W_{\text{out}}^\ell Q_\ell D_{\text{proj}} and b~outℓ=Dproj⊤Qℓ⊤boutℓ\tilde{b}_{\text{out}}^\ell = D_{\text{proj}}^\top Q_\ell^\top b_{\text{out}}^\ell.
        g. Form the sliced residual transformation: Rℓ=Dproj⊤Qℓ−1⊤QℓDproj∈RDsmall×DsmallR_\ell = D_{\text{proj}}^\top Q_{\ell-1}^\top Q_\ell D_{\text{proj}} \in \mathbb{R}^{D_{\text{small}} \times D_{\text{small}}}.
    4. Slice embedding matrix: W~embd=WembdQ0Dproj\tilde{W}_{\text{embd}} = W_{\text{embd}} Q_0 D_{\text{proj}}.
    5. Slice language model head: W~head=Dproj⊤QL⊤Whead\tilde{W}_{\text{head}} = D_{\text{proj}}^\top Q_L^\top W_{\text{head}}.
    6. return Sliced model parameters.
  3. Knowl 3 — Converting LayerNorm-Based Transformers to RMSNorm Form

    model/method

    Standard LayerNorm operations subtract the mean and divide by standard deviation before applying learned scale α∈RD\alpha \in \mathbb{R}^D and offset β∈RD\beta \in \mathbb{R}^D:

    LayerNorm(X)=RMSNorm(XM)diag(α)D+1Nβ⊤\text{LayerNorm}(X) = \text{RMSNorm}(X M) \text{diag}(\alpha) \sqrt{D} + \mathbf{1}_N \beta^\top

    where X∈RN×DX \in \mathbb{R}^{N \times D} is the activation matrix, and M=I−1D11⊤M = I - \frac{1}{D} \mathbf{1} \mathbf{1}^\top is the D×DD \times D mean-subtraction matrix.

    To apply orthogonal transformations and slicing to architectures originally trained with standard LayerNorm (such as the OPT family), the network is first converted to an exact RMSNorm equivalent by absorbing the linear operations into adjacent weight matrices:

    1. The mean-subtraction matrix MM is post-multiplied into the preceding block output matrix: Wout←WoutMW_{\text{out}} \leftarrow W_{\text{out}} M, and the embedding matrix is replaced by Wembd←WembdMW_{\text{embd}} \leftarrow W_{\text{embd}} M.
    2. The scale matrix diag(α)\text{diag}(\alpha) (scaled by D\sqrt{D}) is pre-multiplied into the succeeding block input matrix: Win←diag(α)WinW_{\text{in}} \leftarrow \text{diag}(\alpha) W_{\text{in}}, and the LM head is pre-multiplied by the final LayerNorm scale.

    This reordering leaves the mathematical output of the transformer unchanged while transforming the layer interface into pure RMSNorm, making it invariant to subsequent orthogonal transformations.

  4. Knowl 4 — WikiText-2 Language Modeling Perplexity Across Slicing Levels

    data/table

    SliceGPT was evaluated on language generation perplexity using WikiText-2 across OPT (125M to 66B) and LLAMA-2 (7B to 70B) models. Calibration was performed using 1024 sequences of length 2048 from the WikiText-2 training set.

    Method OPT LLAMA-2
    125M 1.3B 2.7B 6.7B 13B 30B 66B 7B 13B 70B
    Dense 27.64 14.61 12.46 10.85 10.12 9.56 9.33 5.47 4.88 3.32
    SparseGPT 2:4 45.07 29.61 14.90 13.00 11.80 10.53 10.22 8.69 7.07 4.98
    SliceGPT (10%) 29.34 15.10 12.75 10.92 10.27 9.65 9.43 5.89 5.21 3.69
    SliceGPT (20%) 34.26 16.43 13.73 11.48 10.66 9.87 9.57 6.64 5.81 4.25
    SliceGPT (25%) 37.74 17.46 14.56 11.90 10.94 10.04 9.68 7.24 6.30 4.60
    SliceGPT (30%) 43.98 19.09 15.83 12.51 11.33 10.27 9.85 8.12 6.99 5.05

    The data shows that SliceGPT at 25% slicing achieves better perplexity than the SparseGPT 2:4 structured sparsity baseline across all LLAMA-2 models (e.g., 4.60 vs 4.98 for LLAMA-2 70B). In the OPT family, 30% slicing outperforms SparseGPT 2:4 for all models with 6.7B parameters or more. Larger model variants retain a much higher fraction of their dense perplexity after slicing compared to smaller variants.

  5. Knowl 5 — Downstream Zero-Shot Accuracy and LoRA Recovery Fine-Tuning

    empirical result

    SliceGPT was evaluated on five zero-shot downstream benchmarks: PIQA, WinoGrande, HellaSwag, ARC-easy, and ARC-challenge using the LM Evaluation Harness.

    Without fine-tuning, 25% slicing retains 99% of dense performance on OPT 66B (65.17% vs 66.16% dense average accuracy), 96% on LLAMA-2 70B (73.59% vs 76.57% dense average accuracy), and 87% on Phi-2 (62.52% vs 72.24% dense average accuracy).

    Applying post-slicing parameter-efficient recovery fine-tuning (RFT) using LoRA (r=32r=32, α=10\alpha=10, sequence length 1024, trained on 5k sequences from Alpaca) restores accuracy:

    • LLAMA-2 70B sliced by 30% (reducing parameters from ~70B to 51.6B) reaches 74.30% mean zero-shot accuracy with RFT (compared to 76.57% dense and 71.67% un-tuned sliced).
    • Phi-2 (2.8B parameters) sliced by 25% (down to ~2.2B parameters) achieves 65.24% mean accuracy with RFT (compared to 72.24% dense), retaining 90.3% of the original model's zero-shot accuracy.

    RFT with Alpaca yields significantly higher accuracy recovery than RFT with WikiText-2 due to greater alignment with downstream task distributions. Attempted RFT on OPT models did not yield accuracy improvements.

  6. Knowl 6 — Inference Latency and Compute Reductions on Target GPUs

    data/table

    Per-token generation time (batch size 1, sequence length 128) and total GPU compute required for dense and 25%-sliced models on NVIDIA A100 (40GB) and Quadro RTX6000 (24GB) GPUs.

    GPU Type Slicing OPT 66B LLAMA-2 70B
    Per-Token Time Total Compute Per-Token Time Total Compute
    A100 (40GB) Dense 114ms on 4 GPUs 456 GPUms 125ms on 4 GPUs 500 GPUms
    25% 102ms on 3 GPUs 306 GPUms 110ms on 3 GPUs 330 GPUms
    Quadro RTX6000 Dense 237ms on 6 GPUs 1422 GPUms 252ms on 7 GPUs 1764 GPUms
    (24GB) 25% 204ms on 5 GPUs 1020 GPUms 215ms on 5 GPUs 1075 GPUms

    Slicing by 25% reduces the required GPU count by 1 GPU on A100s and 1 to 2 GPUs on RTX6000s. For LLAMA-2 70B, total inference compute is reduced to 66% (330 vs 500 GPUms) on A100 GPUs and 64% (1075 vs 1764 GPUms) on RTX6000 GPUs. These gains are realized using standard, unoptimized dense GEMM operations because the underlying matrix dimensions are reduced directly.

  7. Knowl 7 — Double Precision Requirement for PCA Computation

    empirical result

    Computing the orthogonal transformation matrices QℓQ_\ell in SliceGPT requires double precision (FP64) arithmetic during the covariance matrix accumulation and eigendecomposition steps.

    When PCA is computed using single precision (FP32) via standard numerical libraries (torch.linalg), numerical errors during eigenvector calculation cause substantial degradation in downstream language model perplexity on larger models. For example, on LLAMA-2 70B with a calibration set of 128 samples and sequence length 2048 on WikiText-2:

    • Slicing at 25% with FP64 yields a perplexity of 4.89, whereas FP32 PCA yields 7.01.
    • Slicing at 30% with FP64 yields a perplexity of 5.42, whereas FP32 PCA degrades to 8.75.

    Computing QℓQ_\ell in double precision for LLAMA-2 70B takes approximately 3.5 hours on a single 80GB NVIDIA H100 GPU.

  8. Knowl 8 — Batch Throughput Scaling on NVIDIA H100 GPUs

    data/table

    Maximum token generation throughput on NVIDIA H100 80GB GPUs evaluated at sequence length 128 by scaling batch size up to the GPU memory limit.

    Model Slicing GPUs Batch Size Throughput (Tokens/s)
    OPT 13B Dense 1 512 2518
    25% 1 512 2846 (1.13×1.13\times)
    50% 1 512 3071 (1.22×1.22\times)
    OPT 66B Dense 2 16 141
    25% 2 16 152 (1.08×1.08\times)
    50% 1 32 441 (6.26×6.26\times)
    LLAMA-2 13B Dense 1 512 2707
    25% 1 512 2878 (1.06×1.06\times)
    50% 1 512 3122 (1.15×1.15\times)
    LLAMA-2 70B Dense 2 128 541
    25% 2 256 839 (1.55×1.55\times)
    50% 1 128 1014 (3.75×3.75\times)

    Because SliceGPT reduces the embedding dimension of both the weights and the activation tensors, memory footprints shrink significantly. Slicing at 50% allows OPT 66B and LLAMA-2 70B to fit entirely on a single 80GB GPU instead of two GPUs, yielding end-to-end throughput increases of 6.26×6.26\times (normalized for a 2-GPU allocation) for OPT 66B and 3.75×3.75\times for LLAMA-2 70B.

  9. Knowl 9 — Layerwise Activation Spectrum Characteristics and Non-Uniform Slicing

    empirical result

    Spectral analysis of the activation covariance matrices across transformer layers reveals structural differences between model families:

    1. For both OPT 6.7B and LLAMA-2 7B, the eigenvalue spectrum exhibits faster decay in early layers than in deeper layers, indicating that earlier layers have fewer dominant components.
    2. The embedding spectrum of LLAMA-2 is more uniformly distributed than that of OPT, lacking heavily dominant principal components and making LLAMA-2 more sensitive to uniform truncation.

    Varying the slicing ratio per layer by discarding a fixed target variance during each layer's PCA step (yielding ~24-25% total parameter reduction) has opposite effects across architectures:

    • For OPT models, non-uniform layerwise slicing improves WikiText-2 perplexity over constant 25% slicing (OPT 6.7B improves by 0.16; OPT 13B by 0.28; OPT 30B by 0.18; OPT 66B by 0.12).
    • For LLAMA-2 models, non-uniform layerwise slicing degrades WikiText-2 perplexity compared to constant 25% slicing (LLAMA-2 7B worsens by 0.79; LLAMA-2 13B by 0.17; LLAMA-2 70B by 0.19).

Coverage note — None was omitted; all key theoretical claims, conversion mechanics, core algorithms, empirical generation and zero-shot results, fine-tuning dynamics, and hardware runtime benchmarks are fully covered.

References

  1. 1.Saleh Ashkboos, Ilia Markov, Elias Frantar, Tingxuan Zhong, Xincheng Wang, Jie Ren, Torsten Hoefler, and Dan Alistarh. Towards end-to-end 4-bit inference on generative large language models. arXiv preprint arXiv:2310.09259, 2023.
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  3. 3.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020.
  4. 4.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457, 2018. URL https://api.semanticscholar.org/CorpusID:3922816.
  5. 5.Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM. int8 (): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339, 2022.
  6. 6.Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. Spqr: A sparse-quantized representation for near-lossless LLM weight compression. arXiv preprint arXiv:2306.03078, 2023.
  7. 7.Elias Frantar and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022.
  8. 8.Elias Frantar and Dan Alistarh. SparseGPT: Massive language models can be accurately pruned in one-shot. 2023.
  9. 9.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022.
  10. 10.Trevor Gale, Erich Elsen, and Sara Hooker. The state of sparsity in deep neural networks, 2019.
  11. 11.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al. A framework for few-shot language model evaluation. Version v0. 0.1. Sept, 2021.
  12. 12.Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. A survey of quantization methods for efficient neural network inference. CoRR, abs/2103.13630, 2021. URL https://arxiv.org/abs/2103.13630.
  13. 13.Manish Gupta and Puneet Agrawal. Compression of deep learning models for text: A survey, 2021.
  14. 14.Song Han, Huizi Mao, and William J. Dally. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, 2016.
  15. 15.Babak Hassibi, David G Stork, and Gregory J Wolff. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pp. 293–299. IEEE, 1993.
  16. 16.Yihui He, Xiangyu Zhang, and Jian Sun. Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE international conference on computer vision, pp. 1389–1397, 2017.
  17. 17.Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. CoRR, abs/2102.00554, 2021. URL https://arxiv.org/abs/2102.00554.
  18. 18.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021.
  19. 19.Zehao Huang and Naiyan Wang. Data-driven sparse structure selection for deep neural networks. In Proceedings of the European conference on computer vision (ECCV), pp. 304–320, 2018.
  20. 20.Yann LeCun, John Denker, and Sara Solla. Optimal brain damage. Advances in neural information processing systems, 2, 1989.
  21. 21.Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE international conference on computer vision, pp. 2736–2744, 2017.
  22. 22.Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. Thinet: A filter level pruning method for deep neural network compression. In Proceedings of the IEEE international conference on computer vision, pp. 5058–5066, 2017.
  23. 23.Xinyin Ma, Gongfan Fang, and Xinchao Wang. Llm-pruner: On the structural pruning of large language models. arXiv preprint arXiv:2305.11627, 2023a. URL https://arxiv.org/pdf/2305.11627.pdf.
  24. 24.Xinyin Ma, Gongfan Fang, and Xinchao Wang. LLM-pruner: On the structural pruning of large language models, 2023b.
  25. 25.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. Compacter: Efficient low-rank hypercomplex adapter layers, 2021.
  26. 26.Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  27. 27.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  28. 28.Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.
  29. 29.Matan Ben Noach and Yoav Goldberg. Compressing pre-trained language models by matrix decomposition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pp. 884–889, Suzhou, China, December 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.aacl-main.88.
  30. 30.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
  31. 31.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018.
  32. 32.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  33. 33.Sidak Pal Singh and Dan Alistarh. Woodfisher: Efficient second-order approximation for neural network compression. Advances in Neural Information Processing Systems, 33:18098–18109, 2020.
  34. 34.Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023.
  35. 35.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023.
  36. 36.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023.
  37. 37.Murad Tukan, Alaa Maalouf, Matan Weksler, and Dan Feldman. Compressed deep networks: Goodbye SVD, hello robust low-rank approximation. arXiv preprint arXiv:2009.05647, 2020.
  38. 38.Tycho FA van der Ouderaa, Markus Nagel, Mart van Baalen, Yuki M Asano, and Tijmen Blankevoort. The llm surgeon. arXiv preprint arXiv:2312.17244, 2023.
  39. 39.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  40. 40.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  41. 41.Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning, pp. 38087–38099. PMLR, 2023.
  42. 42.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  43. 43.Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  44. 44.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  45. 45.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
  46. 46.Michael Zhu and Suyog Gupta. To prune, or not to prune: exploring the efficacy of pruning for model compression, 2017.
  47. 47.Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. A survey on model compression for large language models. arXiv preprint arXiv:2308.07633, 2023.

Citation

MLA
Ashkboos, S., et al. “SliceGPT: Compress Large Language Models by Deleting Rows and Columns”. arXiv, 2024, http://arxiv.org/abs/2401.15024v2.
APA
Ashkboos, S., Croci, M. L., Nascimento, M. G. do ., Hoefler, T., & Hensman, J. (2024). SliceGPT: Compress Large Language Models by Deleting Rows and Columns. arXiv. http://arxiv.org/abs/2401.15024v2
Chicago
Ashkboos, S., M. L. Croci, M. G. do . Nascimento, T. Hoefler, and J. Hensman. 2024. “SliceGPT: Compress Large Language Models by Deleting Rows and Columns”. arXiv. http://arxiv.org/abs/2401.15024v2.
Harvard
Ashkboos, S. et al. (2024) “SliceGPT: Compress Large Language Models by Deleting Rows and Columns”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.15024v2.
Vancouver
1. Ashkboos S, Croci ML, Nascimento MG do, Hoefler T, Hensman J (2024) SliceGPT: Compress Large Language Models by Deleting Rows and Columns. arXiv

BibTeX

@article{ashkboos2024slicegpt,
  title = {SliceGPT: Compress Large Language Models by Deleting Rows and Columns},
  author = {Ashkboos, Saleh and Croci, Maximilian L. and Nascimento, Marcelo Gennari do and Hoefler, Torsten and Hensman, James},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.15024v2},
  eprint = {2401.15024}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission