SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias FrantarDan Alistarh
Proposes SparseGPT, a one-shot pruning method that reduces 175-billion-parameter language models to over 50% sparsity in under five hours without requiring retraining or sacrificing accuracy.
Large Language Models deliver state-of-the-art performance across diverse natural language tasks but are exceptionally expensive to deploy. Top-performing models with roughly 175 billion parameters require hundreds of gigabytes of memory and multiple high-end accelerators just for inference. While pruning—the removal of redundant model weights—is an established compression strategy, existing post-training methods require massive retraining to recover accuracy or scale poorly to massive models. As a result, pruning has remained impractical for models containing tens to hundreds of billions of parameters.
The article demonstrates that massive generative pretrained models can be pruned accurately in a single step without any retraining or fine-tuning. It evaluates a new post-training compression method, called SparseGPT, designed to efficiently prune modern transformer models at the scale of 10 to 100+ billion parameters while maintaining near-original model accuracy.
The approach formulates layer-by-layer weight pruning as a large-scale regression problem. By synchronizing inverse Hessian matrices across weight rows and applying adaptive mask selection across blocks of columns, the method optimizes weight updates locally without computing global network gradients. Credibility is supported by end-to-end evaluations on the largest publicly available model families (OPT up to 175B and BLOOM at 176B) using only 128 generic calibration samples (2,048 tokens each) and executing on a single NVIDIA A100 GPU. Evaluations benchmark perplexity on standard text corpora and accuracy across multiple zero-shot evaluation tasks.
The findings establish that SparseGPT can prune 50% to 60% of weights from 175-billion-parameter models in under 4.5 hours with negligible loss in accuracy and perplexity. Crucially, larger models prove significantly more compressible than smaller ones, dropping virtually zero accuracy at 50% sparsity. Standard baseline approaches collapse completely beyond 10% to 30% sparsity, whereas SparseGPT removes over 100 billion parameters successfully. Furthermore, the method extends seamlessly to hardware-friendly semi-structured patterns (such as 2:4 sparsity) and combines with 4-bit weight quantization in a single pass to outperform standalone 3-bit quantization.
These results demonstrate that the extreme computational and financial overhead of running massive generative models can be reduced substantially post-training. Organizations can lower memory footprints and operational costs while maintaining model accuracy without expensive retraining cycles. Practical inference benchmarks show 1.54x to 1.82x speedups on CPUs and GPUs, confirming that substantial operational savings are immediately achievable.
Decision-makers should consider adopting single-pass pruning workflows alongside quantization pipelines to compress production models. When applying structured 2:4 sparsity, sensitivity analysis indicates that pruning earlier layers while retaining later layers provides the best accuracy-performance trade-off. Future efforts should focus on engineering customized sparse inference kernels and investigating the theoretical mechanisms behind why larger models compress more easily.
The evaluation relies on a limited calibration set of 128 text segments and local layer-wise approximations rather than global optimization. While zero-shot task metrics can exhibit variance across individual benchmarks, perplexity evaluations across multiple datasets demonstrate high consistency and robust statistical confidence across the reported results.
- Paper: Second Order Derivatives for Network Pruning: Optimal Brain Surgeon, Babak Hassibi et al. (1992). Introduces the Optimal Brain Surgeon framework that SparseGPT scales and reformulates into an efficient, layer-wise one-shot solver for massive language models.
- Paper: Optimal Brain Damage, Yann LeCun et al. (1989). Establishes the foundational second-order Hessian Taylor approximation for evaluating weight saliency during neural network pruning.
- Paper: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Elias Frantar et al. (2023). Presents the fast second-order matrix update scheme (GPTQ/OBQ) on transformer layers that SparseGPT directly builds upon and adapts for weight pruning.
- Paper: OPT: Open Pre-trained Transformer Language Models, Susan Zhang et al. (2022). Introduces the open-source OPT transformer family up to 175B parameters, which serves as the primary benchmark target for demonstrating SparseGPT's extreme-scale pruning.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). Provides the foundational paradigm of network pruning and sparse neural network efficiency that modern post-training pruning methods seek to accelerate.
- Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). Evaluates whether models pruned via techniques like SparseGPT serve as better initialization points for downstream retraining compared to training smaller models from scratch.
- Paper: AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration, Ji Lin et al. (2024). Extends post-training large language model compression by introducing activation-aware quantization techniques complementary to one-shot pruning.
- Paper: KronQ: LLM Quantization via Kronecker-Factored Hessian, Donghyun Lee et al. (2026). Advances the second-order compression framework by incorporating gradient statistics into a Kronecker-factored Hessian for ultra-low-bit LLM compression.
