A Simple and Effective Pruning Approach for Large Language Models
Mingjie SunZhuang LiuAnna BairJ. Kolter
Introduces Wanda, a post-training pruning method that sparsifies large language models by evaluating the product of weight magnitudes and input activations, matching computationally expensive alternatives without requiring retraining or weight updates.
Large Language Models deliver state-of-the-art natural language capabilities but demand substantial computational and memory resources due to their massive parameter scale. Model pruning—the practice of setting unneeded weights to zero—offers an effective path toward compression, yet existing techniques remain impractical for modern architectures. Traditional pruning requires extensive retraining or fine-tuning, while recent one-shot methods rely on complex, computationally expensive weight updates. Meanwhile, conventional magnitude-based pruning fails dramatically when applied directly to modern language models.
The article introduces and evaluates Wanda (Pruning by Weights and activations), a simple, one-shot pruning method designed to induce high sparsity in pretrained language models without requiring model retraining or weight updates. Motivated by the emergence of outlier features with exceptionally large values in billion-scale architectures, the method determines weight importance by multiplying each individual weight's magnitude by the norm of its corresponding input activation, evaluating these scores locally on a per-output neuron basis.
The researchers evaluated Wanda across the LLaMA and LLaMA-2 model families spanning 7B to 70B parameters, with additional testing on OPT, BLOOM, and Pythia architectures. Across zero-shot reasoning benchmarks, few-shot evaluations, and language modeling perplexity tests, the approach was compared directly against standard magnitude pruning and the leading second-order baseline, SparseGPT, under unstructured (50%, 60%, 80%) and hardware-friendly structured (2:4 and 4:8) sparsity patterns using modest calibration data.
The findings show that Wanda substantially outperforms standard magnitude pruning across all evaluated architectures and sparsity levels without modifying retained weights. At 50% unstructured sparsity, Wanda achieves performance competitive with SparseGPT, and in the largest models (LLaMA-65B and LLaMA-2-70B), the 50% sparse models match the zero-shot accuracy of their original dense counterparts. Computationally, calculating pruning scores with Wanda is up to 300 times faster than SparseGPT, executing in a single forward pass with high robustness even when calibrated on very small sample sizes. Furthermore, structured 2:4 sparsity delivers an approximate 1.6-fold speedup in core matrix multiplications on standard graphics hardware.
These results demonstrate that large language models contain effective, exact sparse subnetworks that can be uncovered without costly weight reconstruction or iterative updates. For organizations deploying generative artificial intelligence, this technique significantly reduces the computational overhead, runtime latency, and engineering costs of model compression. It also shows that the common drop in accuracy observed after pruning can be largely recovered through lightweight fine-tuning methods such as low-rank adaptation.
Engineering teams looking to reduce inference costs and memory footprints should adopt Wanda as an efficient baseline for model sparsification, particularly for larger model tiers where performance degradation is minimal. When higher accuracy is mandatory on smaller models, teams should pair Wanda pruning with parameter-efficient fine-tuning. Further work should explore applying Wanda to dynamic sparse training from scratch and evaluating its effectiveness across non-transformer architectures, domain-specific tasks, and extreme sparsity thresholds exceeding 70%.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT provides the second-order one-shot pruning baseline that Wanda directly compares against, clarifying the accuracy and speed trade-offs Wanda is designed to improve.
- Paper: Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models, Peijie Dong et al. (2024). Pruner-Zero continues post-training LLM pruning by automating the discovery of importance metrics, providing a next step beyond Wanda’s hand-designed weight-and-activation score.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). This study applies Wanda alongside other compression methods to test how pruning affects LLM trustworthiness, extending the efficiency results to safety and reliability concerns.
