Built independently by an author, for readers. Read the story and support ChapterPal

keyword

magnitude pruning

Magnitude pruning is a neural network compression technique that removes parameters by setting the weights with the smallest absolute values to zero, based on the premise that weights closest to zero contribute the least to model performance. By inducing sparsity, this approach reduces the model memory footprint and computational requirements while seeking to preserve predictive accuracy. It can be applied globally across all layers simultaneously or locally within individual layers, and it can be used to generate unstructured sparse representations of individual weights or structured patterns such as pruned channels and blocks. Practitioners often employ magnitude pruning either as a post-training one-shot procedure or iteratively alongside fine-tuning cycles to recover accuracy lost during weight removal.

4 items

A Simple and Effective Pruning Approach for Large Language Models

A Simple and Effective Pruning Approach for Large Language Models

Mingjie Sun, Zhuang Liu, Anna Bair, J. Kolter

OrganizationsBosch Center for AICarnegie Mellon UniversityMeta

Why you should read this

Introduces Wanda, a post-training pruning method that sparsifies large language models by evaluating the product of weight magnitudes and input activations, matching computationally expensive alternatives without requiring retraining or weight updates.

As their size increases, Large Languages Models (LLMs) are natural candidates for network pruning methods: approaches that drop a subset of network weights while striving to preserve performance. Existing methods, however, require either retraining, which is rarely affordable for billion-scale LLMs, or solving a weight reconstruction problem reliant on second-order information, which may also be computationally expensive. In this paper, we introduce a novel, straightforward yet effective pruning method, termed Wanda (Pruning by Weights and activations), designed to induce sparsity in pretrained LLMs. Motivated by the recent observation of emergent large magnitude features in LLMs, our approach prunes weights with the smallest magnitudes multiplied by the corresponding input activations, on a per-output basis. Notably, Wanda requires no retraining or weight update, and the pruned LLM can be used as is. We conduct a thorough evaluation of our method Wanda on LLaMA and LLaMA-2 across various language benchmarks. Wanda significantly outperforms the established baseline of magnitude pruning and performs competitively against recent method involving intensive weight update. Code is available at this https URL.

Added

2026-10-05

Comparing Rewinding and Fine-tuning in Neural Network Pruning

Comparing Rewinding and Fine-tuning in Neural Network Pruning

Alex Renda, Jonathan Frankle, Michael Carbin

OrganizationsMassachusetts Institute of Technology

Why you should read this

Demonstrates that "learning rate rewinding"—resetting the learning rate schedule rather than just the weights—drastically improves the accuracy of pruned models compared to standard fine-tuning.

Many neural network pruning algorithms proceed in three steps: train the network to completion, remove unwanted structure to compress the network, and retrain the remaining structure to recover lost accuracy. The standard retraining technique, fine-tuning, trains the unpruned weights from their final trained values using a small fixed learning rate. In this paper, we compare fine-tuning to alternative retraining techniques. Weight rewinding (as proposed by Frankle et al., (2019)), rewinds unpruned weights to their values from earlier in training and retrains them from there using the original training schedule. Learning rate rewinding (which we propose) trains the unpruned weights from their final values using the same learning rate schedule as weight rewinding. Both rewinding techniques outperform fine-tuning, forming the basis of a network-agnostic pruning algorithm that matches the accuracy and compression ratios of several more network-specific state-of-the-art techniques.

Added

2026-02-26