keyword
magnitude pruning
Magnitude pruning is a neural network compression technique that removes parameters by setting the weights with the smallest absolute values to zero, based on the premise that weights closest to zero contribute the least to model performance. By inducing sparsity, this approach reduces the model memory footprint and computational requirements while seeking to preserve predictive accuracy. It can be applied globally across all layers simultaneously or locally within individual layers, and it can be used to generate unstructured sparse representations of individual weights or structured patterns such as pruned channels and blocks. Practitioners often employ magnitude pruning either as a post-training one-shot procedure or iteratively alongside fine-tuning cycles to recover accuracy lost during weight removal.
4 items

Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Gen Li, Ajay Kumar Jaiswal, Mykola Pechenizkiy, Yi Liang, Michael Bendersky, Zhangyang Wang, Shiwei Liu
Why you should read this
Proposes a non-uniform layerwise pruning method that leverages emergent activation outlier distributions in large language models to enable high-sparsity compression up to seventy percent with minimal perplexity degradation and significant inference acceleration.
Added
2026-10-05

A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, J. Kolter
Why you should read this
Introduces Wanda, a post-training pruning method that sparsifies large language models by evaluating the product of weight magnitudes and input activations, matching computationally expensive alternatives without requiring retraining or weight updates.
As their size increases, Large Languages Models (LLMs) are natural candidates for network pruning methods: approaches that drop a subset of network weights while striving to preserve performance. Existing methods, however, require either retraining, which is rarely affordable for billion-scale LLMs, or solving a weight reconstruction problem reliant on second-order information, which may also be computationally expensive. In this paper, we introduce a novel, straightforward yet effective pruning method, termed Wanda (Pruning by Weights and activations), designed to induce sparsity in pretrained LLMs. Motivated by the recent observation of emergent large magnitude features in LLMs, our approach prunes weights with the smallest magnitudes multiplied by the corresponding input activations, on a per-output basis. Notably, Wanda requires no retraining or weight update, and the pruned LLM can be used as is. We conduct a thorough evaluation of our method Wanda on LLaMA and LLaMA-2 across various language benchmarks. Wanda significantly outperforms the established baseline of magnitude pruning and performs competitively against recent method involving intensive weight update. Code is available at this https URL.
Added
2026-10-05

SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar, Dan Alistarh
Why you should read this
Proposes SparseGPT, a one-shot pruning method that reduces 175-billion-parameter language models to over 50% sparsity in under five hours without requiring retraining or sacrificing accuracy.
We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy. This is achieved via a new pruning method called SparseGPT, specifically designed to work efficiently and accurately on massive GPT-family models. We can execute SparseGPT on the largest available open-source models, OPT-175B and BLOOM-176B, in under 4.5 hours, and can reach 60% unstructured sparsity with negligible increase in perplexity: remarkably, more than 100 billion weights from these models can be ignored at inference time. SparseGPT generalizes to semi-structured (2:4 and 4:8) patterns, and is compatible with weight quantization approaches. The code is available at: this https URL.
Added
2026-09-25

Comparing Rewinding and Fine-tuning in Neural Network Pruning
Alex Renda, Jonathan Frankle, Michael Carbin
Why you should read this
Demonstrates that "learning rate rewinding"—resetting the learning rate schedule rather than just the weights—drastically improves the accuracy of pruned models compared to standard fine-tuning.
Many neural network pruning algorithms proceed in three steps: train the network to completion, remove unwanted structure to compress the network, and retrain the remaining structure to recover lost accuracy. The standard retraining technique, fine-tuning, trains the unpruned weights from their final trained values using a small fixed learning rate. In this paper, we compare fine-tuning to alternative retraining techniques. Weight rewinding (as proposed by Frankle et al., (2019)), rewinds unpruned weights to their values from earlier in training and retrains them from there using the original training schedule. Learning rate rewinding (which we propose) trains the unpruned weights from their final values using the same learning rate schedule as weight rewinding. Both rewinding techniques outperform fine-tuning, forming the basis of a network-agnostic pruning algorithm that matches the accuracy and compression ratios of several more network-specific state-of-the-art techniques.
Added
2026-02-26
