SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Saleh AshkboosMaximilian L. CrociMarcelo Gennari Do NascimentoTorsten HoeflerJames Hensman
Introduces SliceGPT, a post-training compression method that exploits computational invariance to slice transformer weight matrices into smaller dense arrays, cutting parameters by up to 25% and accelerating inference on standard GPUs with minimal loss in zero-shot accuracy.
Deploying pre-trained large language models incurs significant financial and hardware costs because their massive parameter sizes demand extensive memory and multi-GPU infrastructure for inference. Existing compression techniques, such as traditional pruning or semi-structured sparsity, often require specialized sparse hardware libraries, add data structure overhead, or necessitate costly fine-tuning to recover lost accuracy.
The article demonstrates and evaluates SliceGPT, a post-training structured pruning method designed to reduce model size and accelerate inference without requiring custom sparse code or complex retraining. The scheme achieves this by exploiting a mathematical property called computational invariance to systematically delete entire rows and columns from weight matrices, effectively shrinking the network's internal embedding dimension.
The approach was evaluated across multiple open-source model families—including LLAMA-2 (up to 70B parameters), OPT (up to 66B parameters), and Phi-2 (2.7B parameters)—using standard calibration datasets on modern hardware configurations (Quadro RTX6000, A100, and H100 GPUs). The compression process relies on principal component analysis computed on a single GPU in a few hours, followed by optional lightweight recovery fine-tuning.
The analysis produced several key findings: First, SliceGPT successfully removed up to 25% of parameters across large models while maintaining 90% to 99% of original zero-shot task performance. Second, compressed models delivered substantial compute and memory savings; for LLAMA-2 70B, inference compute dropped to 64% on consumer-level GPUs (reducing the GPU requirement from 7 to 5) and 66% on A100 GPUs (reducing GPU count from 4 to 3). Third, sliced models achieved up to 1.55x throughput improvements at 25% slicing on H100 GPUs, and larger slicing levels enabled massive batch-size scaling. Finally, SliceGPT consistently outperformed competitive semi-structured sparsity schemes (such as 2:4 sparsity) in perplexity on large models while utilizing standard dense matrix operations.
These findings have direct operational and financial implications for enterprise AI deployment. Organizations can lower capital and operating expenditures by reducing the number of high-end GPUs needed for serving models, decreasing energy consumption, and improving response latency. Because SliceGPT produces smaller dense matrices rather than sparse formats, it integrates seamlessly into existing deployment pipelines without requiring custom sparse kernels or software rewrites.
Engineering and deployment teams should consider evaluating SliceGPT on large production models to lower hosting footprints. When deploying smaller networks (such as Phi-2 or models under 13B parameters), applying lightweight recovery fine-tuning with representative task data is recommended to restore accuracy. Furthermore, organizations can explore combining SliceGPT with complementary techniques, such as quantization, to capture cumulative efficiency gains.
Confidence in these findings is strong across large-scale architectures, supported by consistent empirical evaluations on standard benchmarks. However, decision-makers should note that smaller models (13B parameters or fewer) experience higher relative accuracy degradation from pruning than larger models. Additionally, optimal calibration requires selecting representative datasets and sufficient sequence lengths to avoid unintended performance drops.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT establishes foundational post-training pruning techniques for large language models, providing the essential baseline context and motivation for SliceGPT's dense matrix dimension reduction.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). Deja Vu analyzes contextual sparsity and structural parameter bottlenecks in modern transformer architectures, framing the hardware-efficient model compression challenges addressed by SliceGPT.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). DepGraph formalizes the architectural dependency constraints across transformer layers that motivate SliceGPT's computational invariance framework for structured dimension slicing.
- Paper: Pruning Filters for Efficient ConvNets, Hao Li et al. (2016). This seminal work on structured filter pruning provides the conceptual prerequisite for deleting entire matrix rows and columns to achieve hardware-native speedups.
- Paper: SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, Guangxuan Xiao et al. (2023). SmoothQuant demonstrates how mathematical equivalence transformations can redistribute internal activation variances across transformer layers, conceptually setting the stage for computational invariance in SliceGPT.
- Paper: SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression, Xin Wang 0120 et al. (2025). SVD-LLM extends post-training matrix reduction techniques by combining singular value decomposition with truncation-aware weight updating to compress transformer representations.
- Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). This empirical study evaluates whether structural and width-pruning compression techniques like SliceGPT transfer useful initialization knowledge compared to training smaller models from scratch.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This survey provides a comprehensive synthesis of post-training compression and efficient LLM architectures, situating structured matrix slicing within the broader design space.
