Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Lu YinYou WuZhenyu ZhangCheng-Yu HsiehYaqing WangYiling JiaGen LiAjay Kumar JaiswalMykola PechenizkiyYi Liang
Proposes a non-uniform layerwise pruning method that leverages emergent activation outlier distributions in large language models to enable high-sparsity compression up to seventy percent with minimal perplexity degradation and significant inference acceleration.
Deploying large language models presents significant financial, computational, and environmental challenges due to their massive parameter sizes. While network pruning—removing redundant parameters to compress models—is an effective solution, standard techniques often require expensive retraining that is impractical at billion-parameter scales. Recent post-training pruning methods allow single-step compression without fine-tuning, but they conventionally apply a uniform pruning ratio across every layer. This uniform approach risks damaging critical model components and overlooks how individual layers function within large architectures.
The article introduces and evaluates Outlier Weighed Layerwise Sparsity (OWL), a novel compression framework designed to allocate non-uniform pruning ratios across model layers. The study's primary objective is to demonstrate that aligning layerwise sparsity with the internal distribution of outlier features—exceptionally large activation values that are critical to language model performance—enables substantially higher compression rates without sacrificing accuracy.
To develop and validate this method, the authors conducted comprehensive empirical evaluations using leading model families, including LLaMA (7B to 65B), LLaMA-2, Vicuna, OPT, and Mistral. They analyzed layerwise outlier distributions across standard datasets, such as WikiText for language modeling quality and seven benchmark datasets for zero-shot reasoning tasks. The team integrated OWL into leading pruning techniques, notably Wanda and SparseGPT, and measured end-to-end execution speed on central processing units using the DeepSparse inference engine.
The article reports several critical findings. First, outlier features in dense language models follow a non-uniform, U-shaped distribution across layers, meaning that initial and final layers house significantly higher concentrations of critical weights than middle layers. Second, OWL consistently outperforms uniform pruning baselines, especially at extreme sparsity: at 70% parameter removal on LLaMA-7B, OWL reduces language modeling perplexity by 61.22 points when combined with Wanda and 6.80 points when combined with SparseGPT. Third, OWL delivers average accuracy gains of 2.19% to 4.72% across common-sense zero-shot evaluation benchmarks at 70% sparsity. Fourth, this compression translates into real-world efficiency, achieving a 2.6-fold CPU inference speedup at 70% sparsity and up to a 3.9-fold speedup at 90% sparsity, while adding negligible computational overhead (under two seconds) during the pruning process. Finally, minimal post-pruning fine-tuning with only 30,000 tokens rapidly recovers performance losses.
These findings challenge the prevailing assumption that uniform layerwise pruning is optimal for large language models. By retaining more weights in layers with high outlier density and pruning more aggressively in layers with low outlier density, organizations can cut inference latency and hardware memory requirements dramatically. Furthermore, the principles of outlier-guided layer weighting transfer effectively to other compression techniques, including low-rank matrix approximation, structured group pruning, and mixed-precision quantization.
Decision-makers and technical teams should transition from uniform layer pruning to outlier-aware allocation strategies when deploying compressed models for production. When extreme compression (70% or higher) is necessary, pairing OWL with lightweight parameter fine-tuning is strongly recommended to restore baseline accuracy. The authors note, however, that OWL’s benefits are tied directly to the emergence of activation outliers in textual models; tests on vision models showed minimal improvement due to the absence of pronounced outlier phenomena. Organizations should confidently adopt OWL for Transformer-based language models while conducting application-specific latency and accuracy benchmarking on target hardware prior to full deployment.
- Paper: A Simple and Effective Pruning Approach for Large Language Models, Mingjie Sun et al. (2023). Read Wanda first to understand the activation-aware pruning baseline that OWL extends with outlier-guided, non-uniform layer sparsity.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT supplies the second post-training pruning baseline that OWL adapts to its layerwise sparsity allocation.
No sufficiently relevant recommendations were found.
