Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Lu YinYou WuZhenyu ZhangCheng-Yu HsiehYaqing WangYiling JiaGen LiAjay Kumar JaiswalMykola PechenizkiyYi Liang

article2024ICML230 citations

Proposes a non-uniform layerwise pruning method that leverages emergent activation outlier distributions in large language models to enable high-sparsity compression up to seventy percent with minimal perplexity degradation and significant inference acceleration.

Listen

Deploying large language models presents significant financial, computational, and environmental challenges due to their massive parameter sizes. While network pruning—removing redundant parameters to compress models—is an effective solution, standard techniques often require expensive retraining that is impractical at billion-parameter scales. Recent post-training pruning methods allow single-step compression without fine-tuning, but they conventionally apply a uniform pruning ratio across every layer. This uniform approach risks damaging critical model components and overlooks how individual layers function within large architectures.

The article introduces and evaluates Outlier Weighed Layerwise Sparsity (OWL), a novel compression framework designed to allocate non-uniform pruning ratios across model layers. The study's primary objective is to demonstrate that aligning layerwise sparsity with the internal distribution of outlier features—exceptionally large activation values that are critical to language model performance—enables substantially higher compression rates without sacrificing accuracy.

To develop and validate this method, the authors conducted comprehensive empirical evaluations using leading model families, including LLaMA (7B to 65B), LLaMA-2, Vicuna, OPT, and Mistral. They analyzed layerwise outlier distributions across standard datasets, such as WikiText for language modeling quality and seven benchmark datasets for zero-shot reasoning tasks. The team integrated OWL into leading pruning techniques, notably Wanda and SparseGPT, and measured end-to-end execution speed on central processing units using the DeepSparse inference engine.

The article reports several critical findings. First, outlier features in dense language models follow a non-uniform, U-shaped distribution across layers, meaning that initial and final layers house significantly higher concentrations of critical weights than middle layers. Second, OWL consistently outperforms uniform pruning baselines, especially at extreme sparsity: at 70% parameter removal on LLaMA-7B, OWL reduces language modeling perplexity by 61.22 points when combined with Wanda and 6.80 points when combined with SparseGPT. Third, OWL delivers average accuracy gains of 2.19% to 4.72% across common-sense zero-shot evaluation benchmarks at 70% sparsity. Fourth, this compression translates into real-world efficiency, achieving a 2.6-fold CPU inference speedup at 70% sparsity and up to a 3.9-fold speedup at 90% sparsity, while adding negligible computational overhead (under two seconds) during the pruning process. Finally, minimal post-pruning fine-tuning with only 30,000 tokens rapidly recovers performance losses.

These findings challenge the prevailing assumption that uniform layerwise pruning is optimal for large language models. By retaining more weights in layers with high outlier density and pruning more aggressively in layers with low outlier density, organizations can cut inference latency and hardware memory requirements dramatically. Furthermore, the principles of outlier-guided layer weighting transfer effectively to other compression techniques, including low-rank matrix approximation, structured group pruning, and mixed-precision quantization.

Decision-makers and technical teams should transition from uniform layer pruning to outlier-aware allocation strategies when deploying compressed models for production. When extreme compression (70% or higher) is necessary, pairing OWL with lightweight parameter fine-tuning is strongly recommended to restore baseline accuracy. The authors note, however, that OWL’s benefits are tied directly to the emergence of activation outliers in textual models; tests on vision models showed minimal improvement due to the absence of pronounced outlier phenomena. Organizations should confidently adopt OWL for Transformer-based language models while conducting application-specific latency and accuracy benchmarking on target hardware prior to full deployment.

arXiv: 2310.05175

No sufficiently relevant recommendations were found.

Cover for Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Outlier Weighed Layerwise Sparsity
  • 3.1 Rationale
  • 3.2 Empirical Study
  • 3.3 Outlier Weighed Layerwise Sparsity (OWL)
  • 4 Experiments
  • 4.1 Main Experiments
  • 4.2 More Advanced LLMs
  • 4.3 More Practical Applications of OWL
  • 5 Analysis
  • 6 Conclusion
  • 7 Impact Statements
  • 8 Acknowledgement
  • References
  • A Vision Model Pruning
  • B More Practical Applications of OWL
  • B.1 N:M Sparsity
  • B.2 Structured Pruning
  • B.3 Mixed-Precision Quantization
  • C Per-Block Vs. Per-Layer
  • D Hyperparameters

Knowls

  1. Knowl 1 — Outlier Weighed Layerwise Sparsity (OWL) Framework

    model/method

    Outlier Weighed Layerwise Sparsity (OWL) is a layerwise sparsity allocation method for one-shot pruning of Large Language Models (LLMs) without retraining. Rather than applying a uniform sparsity ratio across all layers or globally thresholding weights across the entire model, OWL allocates non-uniform target sparsity ratios to transformer blocks in proportion to the presence of activation and weight outliers.

    Given an LL-block LLM with a target model-level weight sparsity S∈(0,1)S \in (0, 1), OWL evaluates the Layerwise Outlier Distribution (LOD) vector D=[D1,D2,…,DL]D = [D_1, D_2, \dots, D_L], where DlD_l denotes the outlier proportion in block ll. Guided by the principle that outlier-rich blocks must undergo less pruning to prevent performance degradation, the target sparsity SlS_l for block ll is set inversely to its outlier density: Sl∝1−DlS_l \propto 1 - D_l

    To prevent severe inter-layer sparsity discrepancies (which cause structural collapse in extreme layers), OWL enforces a bounded search interval around the overall target SS using a hyperparameter λ∈[0.02,0.20]\lambda \in [0.02, 0.20]: Sl∈[S−λ,S+λ]S_l \in [S - \lambda, S + \lambda] subject to the global constraint that the average parameter sparsity satisfies 1L∑l=1LSl=S\frac{1}{L} \sum_{l=1}^L S_l = S. Once layerwise targets are assigned, weights within each block are pruned using local unstructured pruning metrics such as Wanda or SparseGPT.

  2. Knowl 2 — Layerwise Outlier Distribution (LOD) Formulation

    equation

    The Layerwise Outlier Distribution (LOD) measures the layer-by-layer density of outlier weights in a large language model. For a linear layer with weight matrix W∈RCout×CinW \in \mathbb{R}^{C_{\text{out}} \times C_{\text{in}}} and calibration input feature activations X∈R(N×L)×CinX \in \mathbb{R}^{(N \times L) \times C_{\text{in}}} across batch dimension NN and sequence length LL, the outlier score AijA_{ij} for weight WijW_{ij} is computed as: Aij=∥Xj∥2⋅∣Wij∣A_{ij} = \|X_j\|_2 \cdot |W_{ij}| where ∥Xj∥2\|X_j\|_2 represents the ℓ2\ell_2-norm of the input feature column corresponding to input channel jj.

    The layerwise outlier ratio DlD^l for layer ll is the proportion of weights whose outlier scores exceed the layer average by a factor of at least MM: Dl=∑i=1Cout∑j=1CinI(Aijl>M⋅Aˉl)CinCoutD^l = \frac{\sum_{i=1}^{C_{\text{out}}} \sum_{j=1}^{C_{\text{in}}} \mathbb{I}\left(A^l_{ij} > M \cdot \bar{A}^l\right)}{C_{\text{in}} C_{\text{out}}} where Aˉl=1CinCout∑i=1Cout∑j=1CinAijl\bar{A}^l = \frac{1}{C_{\text{in}} C_{\text{out}}} \sum_{i=1}^{C_{\text{out}}} \sum_{j=1}^{C_{\text{in}}} A^l_{ij} is the arithmetic mean of outlier scores across layer ll, I(⋅)\mathbb{I}(\cdot) is the indicator function returning 1 if the condition holds and 0 otherwise, and MM is an outlier multiplier threshold (typically set to M=5M=5 or M=7M=7). The entire network's outlier profile is given by LOD=[D1,D2,…,Dn]\text{LOD} = [D^1, D^2, \dots, D^n].

  3. Knowl 3 — WikiText Language Modeling Perplexity at 70% Sparsity

    data/table

    Language modeling evaluations on the WikiText validation dataset demonstrate that integrating OWL into Wanda and SparseGPT substantially reduces perplexity at 70% unstructured weight sparsity across LLaMA-V1 (7B, 13B, 30B, 65B) and OPT-6.7B architectures compared to uniform layerwise pruning:

    Method Layerwise Sparsity LLaMA-7B LLaMA-13B LLaMA-30B LLaMA-65B OPT-6.7B
    Dense Baseline - 5.68 5.09 4.10 4.77 10.13
    Magnitude Uniform 48419.12 84539.45 977.73 46.89 290985.03
    Wanda Uniform 85.77 55.90 17.37 15.23 162.92
    OWL w. Wanda Non-Uniform 24.55 17.17 10.75 8.61 40.22
    SparseGPT Uniform 26.30 19.24 12.56 10.45 20.29
    OWL w. SparseGPT Non-Uniform 19.49 14.55 10.28 8.28 22.48

    OWL with Wanda achieves a 61.22 perplexity reduction on LLaMA-7B (from 85.77 to 24.55) and a 38.73 point reduction on LLaMA-13B. When applied to SparseGPT, OWL achieves a 6.81 perplexity reduction on LLaMA-7B (from 26.30 to 19.49). The performance advantage of OWL grows larger as the model size decreases.

  4. Knowl 4 — Zero-Shot Common Sense Reasoning Accuracy at 70% Sparsity

    data/table

    Zero-shot task accuracy across seven benchmarks (BoolQ, RTE, HellaSwag, WinoGrande, ARC-Easy, ARC-Challenge, OpenBookQA) demonstrates that OWL consistently outperforms uniform sparsity baselines at 70% unstructured sparsity across the LLaMA-V1 family:

    Model Method BoolQ RTE HellaSwag WinoGrande ARC-e ARC-c OBQA Mean
    LLaMA-7B Dense 75.14 66.43 74.80 70.01 67.67 41.38 41.40 62.40
    Magnitude 38.29 52.71 24.68 51.46 26.98 22.35 25.80 34.61
    Wanda 55.11 57.40 31.83 51.38 34.22 19.80 26.00 39.39
    OWL w. Wanda 62.48 58.48 44.79 58.72 45.03 26.19 29.60 46.47
    SparseGPT 64.53 53.79 42.11 58.64 43.06 24.57 27.80 44.93
    OWL w. SparseGPT 67.13 53.43 48.56 62.03 45.41 27.65 32.00 48.03
    LLaMA-13B Dense 77.86 70.40 78.08 72.77 69.19 47.18 43.80 65.61
    Wanda 61.71 52.71 34.31 52.33 37.16 20.90 29.60 41.25
    OWL w. Wanda 62.69 52.71 51.03 63.14 49.54 28.67 34.40 48.88
    SparseGPT 66.94 52.71 47.91 62.90 45.03 27.99 35.20 48.38
    OWL w. SparseGPT 64.95 53.07 54.39 66.54 48.86 30.12 38.00 50.85
    LLaMA-30B Dense 82.69 66.79 81.19 75.85 73.48 50.77 44.60 67.91
    Wanda 66.12 57.76 58.84 67.32 59.26 33.11 40.20 54.66
    OWL w. Wanda 66.42 52.35 62.94 69.30 61.83 35.84 40.00 55.53
    SparseGPT 66.51 63.90 60.38 69.85 58.54 33.70 40.60 55.78
    OWL w. SparseGPT 67.58 58.48 64.88 70.72 60.82 35.07 42.20 57.11
    LLaMA-65B Dense 84.86 69.68 82.94 77.35 75.08 52.56 44.20 69.52
    Wanda 76.30 56.68 61.26 70.48 63.47 35.67 39.40 57.61
    OWL w. Wanda 80.12 58.84 66.16 73.56 65.45 39.93 42.20 60.89
    SparseGPT 80.64 59.57 66.42 72.61 60.52 38.57 40.80 59.88
    OWL w. SparseGPT 82.63 67.15 68.52 75.06 60.10 39.59 39.00 61.72

    Compared to uniform layerwise sparsity with Wanda and SparseGPT alone, OWL achieves an average accuracy gain of 4.72% and 2.19% across all 7 downstream tasks and 4 model scales.

  5. Knowl 5 — Comparison of Layerwise Sparsity Strategies for LLM Pruning

    empirical result

    Comparing alternative layerwise sparsity strategies applied to LLaMA-7B using Wanda pruning on WikiText validation perplexity reveals that heuristic allocations commonly used in vision models fail on LLMs at high sparsity:

    Sparsity Strategy 10% 20% 30% 40% 50% 60% 70% 80%
    Global 14.11 3134 10293 10762 14848 17765 5147 39918.56
    Erdős-Rényi (ER) 5.69 5.80 6.02 6.55 7.74 12.16 112.03 11151.18
    ER-plus 5.70 5.82 6.05 6.62 8.00 14.04 229.17 6013.91
    Uniform 5.69 5.81 5.99 6.38 7.26 10.70 85.77 3499.88
    OWL-inverse (1−LOD1 - \text{LOD}) 5.72 5.83 6.04 6.51 8.03 26.05 822.23 9616.08
    OWL (Ours) 5.70 5.80 6.01 6.39 7.22 9.35 24.54 1002.87

    Key observations:

    1. At low sparsity levels (≤40%\le 40\%), all layerwise heuristics except Global pruning perform comparably (perplexity ~5.7 to 6.6).
    2. Global thresholding causes catastrophic collapse even at 20% sparsity (perplexity 3134).
    3. Graph-theoretic allocations like Erd\H{o}s-R'enyi (ER) and ER-plus degrade severely beyond 60% sparsity compared to Uniform.
    4. Inverting OWL's outlier weights (OWL-inverse, which increases sparsity on outlier-dense layers) causes a sharp perplexity spike (822.23 at 70% sparsity), verifying that preserving outlier-rich layers is the operative cause of OWL's performance.
  6. Knowl 6 — Layerwise Outlier Structure and Pruning Metric Impact

    empirical result

    In dense LLMs, the distribution of activation/weight outliers follows a non-uniform, loosely U-shaped curve across model depth: the earliest transformer layers (near input) and latest transformer layers (near output) exhibit a substantially higher concentration of outliers, while middle layers show lower, monotonically decreasing outlier densities.

    Pruning efficacy is directly correlated with a method's ability to preserve this outlier structure:

    • In LLaMA-13B at 70% uniform sparsity (dense baseline LOD: 5.432%, perplexity: 5.090), SparseGPT retains the highest fraction of outliers, increasing total LOD to 6.645% (ΔLOD=+1.213%\Delta\text{LOD} = +1.213\%, perplexity: 19.235).
    • Wanda preserves outliers by increasing LOD to 5.716% (ΔLOD=+0.284%\Delta\text{LOD} = +0.284\%, perplexity: 55.900).
    • Naive magnitude pruning fails to protect outliers, decreasing total LOD to 5.322% (ΔLOD=−0.110%\Delta\text{LOD} = -0.110\%, perplexity: 84539.445), which triggers model collapse.
  7. Knowl 7 — Per-Block Versus Per-Layer Sparsity Allocation in OWL

    empirical result

    Allocating non-uniform sparsity at the granularity of entire Transformer blocks is substantially superior to allocating sparsity on individual linear projection layers.

    In a Transformer block containing 7 linear layers (q_proj, k_proj, v_proj, o_proj, gate_proj, down_proj, and up_proj):

    • Applying OWL on an independent per-layer basis produces nearly uniform sparsity across depth for specific projection types (such as v_proj, gate_proj, and up_proj), resulting in a degraded WikiText validation perplexity of 86.285 on LLaMA-7B at 70% sparsity.
    • Applying OWL on a per-block basis (calculating an aggregate outlier density for all 7 layers within each block and assigning a unified block-level target sparsity) preserves inter-block heterogeneity and achieves a WikiText validation perplexity of 24.55 on LLaMA-7B at 70% sparsity.
  8. Knowl 8 — Inference Acceleration on CPUs with DeepSparse

    empirical result

    Evaluating LLaMA-V2-7B-chat-hf pruned with OWL on the DeepSparse CPU inference engine using an Intel Xeon Platinum 8360Y 36-core processor yields substantial end-to-end token decode speedups:

    Metric Dense 10% 20% 30% 40% 50% 60% 70% 80% 90%
    Latency (ms) 213.83 216.86 221.62 218.01 167.54 121.25 101.41 81.89 64.57 54.24
    Throughput (tok/s) 4.68 4.61 4.51 4.59 5.97 8.25 9.86 12.21 15.48 18.43
    Speedup 1.0× 1.0× 1.0× 1.0× 1.3× 1.8× 2.1× 2.6× 3.3× 3.9×

    While sparsity below 40% does not yield speedup over dense execution, OWL delivers 2.6×2.6\times speedup (81.89 ms latency, 12.21 tokens/s) at 70% sparsity, 3.3×3.3\times at 80% sparsity, and 3.9×3.9\times at 90% sparsity. The additional runtime required to compute the OWL metric during pruning is negligible (≤2.0\le 2.0 seconds on an NVIDIA A100 GPU for a 65B parameter model).

  9. Knowl 9 — Application of LOD to SVD, N:M Sparsity, Structured Pruning, and Mixed Precision

    model/method

    The Layerwise Outlier Distribution (LOD) can serve as a universal layer importance metric across diverse LLM compression regimes beyond unstructured pruning:

    1. SVD Low-Rank Approximation: For weight matrix W∈Rd1×d2W \in \mathbb{R}^{d_1 \times d_2} with preserved rank rr, compression ratio is rmin⁡(d1,d2)\frac{r}{\min(d_1, d_2)}. Using LOD to assign lower compression to outlier-heavy layers reduces WikiText perplexity on un-finetuned LLaMA-V1-7B: at 40% rank reduction, OWL-SVD achieves 43.02 perplexity versus 1909.34 for uniform SVD; at 30% reduction, it achieves 12.92 versus 17.23.
    2. Mixed N:M Sparsity: Allowing variable non-zero weights NN across layers under an N:8N:8 scheme guided by OWL yields lower perplexity on LLaMA-7B: at mixed 3:8, OWL achieves 21.49 perplexity versus 42.56 for uniform 3:8; at mixed 2:8, OWL achieves 331.37 versus 2962.00.
    3. Structured Pruning: When applied to LLM-Pruner (removing entire neurons/heads), OWL layer allocation reduces LLaMA-7B perplexity at 60% structured sparsity from 90.02 to 76.99 on WikiText, and from 192.06 to 150.16 on PTB.
    4. Mixed-Precision Quantization: Allocating bit precision (e.g., 3/4-bit or 2/3/4-bit) according to OWL scores achieves 9.09 perplexity for mixed 3/4-bit (versus 14.61 using L1L_1-norm) and 190.28 for mixed 2/3/4-bit (versus 13959.42 using L1L_1-norm).
  10. Knowl 10 — Performance Recovery via LoRA Fine-Tuning on Sparse LLMs

    empirical result

    Applying parameter-efficient fine-tuning via LoRA with an unmerged adapter on OWL-pruned models restores language modeling fidelity using minimal training data (30,000 tokens sampled from the C4 training dataset).

    On WikiText validation perplexity at 70% sparsity:

    • For LLaMA-7B pruned with OWL w. SparseGPT, LoRA fine-tuning reduces perplexity from 19.49 to 11.15.
    • For LLaMA-13B pruned with OWL w. SparseGPT, LoRA fine-tuning reduces perplexity from 14.55 to 9.00.
  11. Knowl 11 — Limitation of OWL on Vision Models

    limitation

    When evaluated on computer vision models on ImageNet-1K, OWL provides modest accuracy gains on vision transformers (DeiT-Base achieves 54.24% top-1 accuracy at 70% sparsity with OWL w. Wanda vs. 49.20% with uniform Wanda) but yields no benefit on convolutional models (ConvNeXt-Base achieves 68.28% at 70% sparsity with OWL w. Wanda vs. 68.18% with uniform Wanda).

    This behavior is attributed to the origin of outlier features: outliers in transformer language models are causally linked to high-frequency tokens in natural language pre-training corpora. Because continuous-valued vision data lacks comparable discrete high-frequency token distributions, outlier features are less prominent, diminishing the advantages of outlier-weighted layerwise sparsity.

Coverage note — None was omitted; all key theoretical definitions, formulas, empirical pruning benchmarks, ablation studies, extensions (SVD, N:M, structured pruning, quantization, LoRA recovery), hardware acceleration metrics, and limitations are fully represented.

References

  1. 1.Bhojanapalli, S., Chakrabarti, A., Veit, A., Lukasik, M., Jain, H., Liu, F., Chang, Y.-W., and Kumar, S. Leveraging redundancy in attention with reuse transformers. arXiv preprint arXiv:2110.06821, 2021.
  2. 2.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems (NeurIPs), 33:1877–1901, 2020.
  3. 3.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019.
  4. 4.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018.
  5. 5.DeepSparse. NeuralMagic DeepSparse Inference Engine, 2021. URL https://github.com/neuralmagic/deepsparse.
  6. 6.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 248–255, 2009.
  7. 7.Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. Llm. int8 (): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems (NeurIPs), 2022.
  8. 8.Erdős, P. and Rényi, A. On random graphs i. Publicationes Mathematicae (Debrecen), 6:290–297, 1959.
  9. 9.Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning (ICML), pp. 2943–2952, 2020.
  10. 10.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations (ICLR), 2019.
  11. 11.Frantar, E. and Alistarh, D. Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning (ICML), 2023.
  12. 12.Gale, T., Elsen, E., and Hooker, S. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574, 2019.
  13. 13.Gale, T., Zaharia, M., Young, C., and Elsen, E. Sparse gpu kernels for deep learning. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1–14. IEEE, 2020.
  14. 14.Han, S., Pool, J., Tran, J., and Dally, W. Learning both weights and connections for efficient neural network. In Advances in Neural Information Processing Systems (NeurIPS), pp. 1135–1143, 2015.
  15. 15.Hassibi, B., Stork, D. G., and Wolff, G. J. Optimal brain surgeon and general network pruning. In IEEE international conference on neural networks, pp. 293–299. IEEE, 1993.
  16. 16.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  17. 17.Jaiswal, A., Gan, Z., Du, X., Zhang, B., Wang, Z., and Yang, Y. Compressing llms: The truth is rarely pure and never simple. arXiv preprint arXiv:2310.01382, 2023a.
  18. 18.Jaiswal, A., Liu, S., Chen, T., and Wang, Z. The emergence of essential sparsity in large pre-trained models: The weights that matter. arXiv preprint arXiv:2306.03805, 2023b.
  19. 19.Janowsky, S. A. Pruning versus clipping in neural networks. Physical Review A, 39(12):6600, 1989.
  20. 20.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  21. 21.Kovaleva, O., Kulshreshtha, S., Rogers, A., and Rumshisky, A. Bert busters: Outlier dimensions that disrupt transformers. arXiv preprint arXiv:2105.06990, 2021.
  22. 22.Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D. The optimal bert surgeon: Scalable and accurate second-order pruning for large language models. arXiv preprint arXiv:2203.07259, 2022.
  23. 23.LeCun, Y., Denker, J., and Solla, S. Optimal brain damage. In Advances in Neural Information Processing Systems (NeurIPS), pp. 598–605, 1989.
  24. 24.Lee, J., Park, S., Mo, S., Ahn, S., and Shin, J. Layer-adaptive sparsity for the magnitude-based pruning. arXiv preprint arXiv:2010.07611, 2020.
  25. 25.Lee, N., Ajanthan, T., and Torr, P. Snip: Single-shot network pruning based on connection sensitivity. In International Conference on Learning Representations (ICLR), 2019.
  26. 26.Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978, 2023.
  27. 27.Lin, S., Ji, R., Yan, C., Zhang, B., Cao, L., Ye, Q., Huang, F., and Doermann, D. Towards optimal structured cnn pruning via generative adversarial learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2790–2799, 2019.
  28. 28.Liu, S. and Wang, Z. Ten lessons we have learned in the new" sparseland": A short handbook for sparse neural network researchers. arXiv preprint arXiv:2302.02596, 2023.
  29. 29.Liu, S., Chen, T., Chen, X., Atashgahi, Z., Yin, L., Kou, H., Shen, L., Pechenizkiy, M., Wang, Z., and Mocanu, D. C. Sparse training via boosting pruning plasticity with neuroregeneration. In Advances in Neural Information Processing Systems (NeurIPS), 2021a.
  30. 30.Liu, S., Yin, L., Mocanu, D. C., and Pechenizkiy, M. Do we actually need dense over-parameterization? in-time over-parameterization in sparse training. In International Conference on Machine Learning, pp. 6989–7000. PMLR, 2021b.
  31. 31.Liu, S., Chen, T., Chen, X., Shen, L., Mocanu, D. C., Wang, Z., and Pechenizkiy, M. The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training. arXiv preprint arXiv:2202.02643, 2022a.
  32. 32.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11976–11986, 2022b.
  33. 33.Luccioni, A. S., Viguier, S., and Ligozat, A.-L. Estimating the carbon footprint of bloom, a 176b parameter language model. arXiv preprint arXiv:2211.02001, 2022.
  34. 34.Ma, X., Fang, G., and Wang, X. Llm-pruner: On the structural pruning of large language models. arXiv preprint arXiv:2305.11627, 2023.
  35. 35.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016a.
  36. 36.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016b.
  37. 37.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018.
  38. 38.Mittal, D., Bhardwaj, S., Khapra, M. M., and Ravindran, B. Studying the plasticity in deep convolutional neural networks using random pruning. Machine Vision and Applications, 30(2):203–216, 2019.
  39. 39.Mocanu, D. C., Mocanu, E., Nguyen, P. H., Gibescu, M., and Liotta, A. A topological insight into restricted boltzmann machines. Machine Learning, 104(2):243–270, Sep 2016. ISSN 1573-0565. doi: 10.1007/s10994-016-5570-z. URL https://doi.org/10.1007/s10994-016-5570-z.
  40. 40.Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature Communications, 9:1–12, 2018.
  41. 41.Mozer, M. C. and Smolensky, P. Skeletonization: A technique for trimming the fat from a network via relevance assessment. In Advances in Neural Information Processing Systems (NeurIPS), pp. 107–115, 1989.
  42. 42.Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  43. 43.Puccetti, G., Rogers, A., Drozd, A., and Dell’Orletta, F. Outliers dimensions that disrupt transformers are driven by frequency. arXiv preprint arXiv:2205.11380, 2022.
  44. 44.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  45. 45.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande: An adversarial winograd schema challenge at scale. arXiv preprint arXiv:1907.10641, 2019.
  46. 46.Sanh, V., Wolf, T., and Rush, A. M. Movement pruning: Adaptive sparsity by fine-tuning. arXiv preprint arXiv:2005.07683, 2020.
  47. 47.Schaeffer, R., Miranda, B., and Koyejo, S. Are emergent abilities of large language models a mirage? arXiv preprint arXiv:2304.15004, 2023.
  48. 48.Sun, M., Liu, Z., Bair, A., and Kolter, J. Z. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023.
  49. 49.Sun, W., Zhou, A., Stuijk, S., Wijnhoven, R., Nelson, A. O., Corporaal, H., et al. Dominosearch: Find layer-wise fine-grained n: M sparse schemes from dense neural networks. Advances in neural information processing systems, 34:20721–20732, 2021.
  50. 50.Tang, C., Ouyang, K., Wang, Z., Zhu, Y., Ji, W., Wang, Y., and Zhu, W. Mixed-precision neural network quantization via learned layer-wise importance. In European Conference on Computer Vision, pp. 259–275. Springer, 2022.
  51. 51.Timkey, W. and van Schijndel, M. All bark and no bite: Rogue dimensions in transformer language models obscure representational quality. arXiv preprint arXiv:2109.04404, 2021.
  52. 52.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. Training data-efficient image transformers & distillation through attention. In International conference on machine learning, pp. 10347–10357. PMLR, 2021.
  53. 53.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  54. 54.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  55. 55.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  56. 56.Wang, C., Zhang, G., and Grosse, R. Picking winning tickets before training by preserving gradient flow. In International Conference on Learning Representations (ICLR), 2020.
  57. 57.Wang, W. and Tu, Z. Rethinking the value of transformer components. arXiv preprint arXiv:2011.03803, 2020.
  58. 58.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022.
  59. 59.Wen, W., He, Y., Rajbhandari, S., Zhang, M., Wang, W., Liu, F., Hu, B., Chen, Y., and Li, H. Learning intrinsic sparse structures within long short-term memory. arXiv preprint arXiv:1709.05027, 2017.
  60. 60.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning (ICML), pp. 38087–38099. PMLR, 2023.
  61. 61.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019.
  62. 62.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  63. 63.Zhang, Y., Zhao, L., Lin, M., Sun, Y., Yao, Y., Han, X., Tanner, J., Liu, S., and Ji, R. Dynamic sparse no training: Training-free fine-tuning for sparse llms. arXiv preprint arXiv:2310.08915, 2023.
  64. 64.Zhu, M. and Gupta, S. To prune, or not to prune: exploring the efficacy of pruning for model compression. In International Conference on Learning Representations Workshop (ICLRW), 2017.
  65. 65.Zimmer, M., Andoni, M., Spiegel, C., and Pokutta, S. Perp: Rethinking the prune-retrain paradigm in the era of llms. arXiv preprint arXiv:2312.15230, 2023.

Citation

MLA
Yin, L., et al. “Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity”. International Conference on Machine Learning, vol. 235, 2024, pp. 57101–15, https://proceedings.mlr.press/v235/yin24e.html.
APA
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Li, G., Jaiswal, A. K., Pechenizkiy, M., Liang, Y., Bendersky, M., Wang, Z., & Liu, S. (2024). Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity. International Conference on Machine Learning, 235, 57101–57115. https://proceedings.mlr.press/v235/yin24e.html
Chicago
Yin, L., Y. Wu, Z. Zhang, et al. 2024. “Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity”. International Conference on Machine Learning 235: 57101–15. https://proceedings.mlr.press/v235/yin24e.html.
Harvard
Yin, L. et al. (2024) “Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity”, International Conference on Machine Learning. PMLR, pp. 57101–57115. Available at: https://proceedings.mlr.press/v235/yin24e.html.
Vancouver
1. Yin L, Wu Y, Zhang Z, et al (2024) Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity. In: International Conference on Machine Learning. PMLR, pp 57101–57115

BibTeX

@InProceedings{pmlr-v235-yin24e,
  title = 	 {Outlier Weighed Layerwise Sparsity ({OWL}): A Missing Secret Sauce for Pruning {LLM}s to High Sparsity},
  author =       {Yin, Lu and Wu, You and Zhang, Zhenyu and Hsieh, Cheng-Yu and Wang, Yaqing and Jia, Yiling and Li, Gen and Jaiswal, Ajay Kumar and Pechenizkiy, Mykola and Liang, Yi and Bendersky, Michael and Wang, Zhangyang and Liu, Shiwei},
  booktitle = 	 {Proceedings of the 41st International Conference on Machine Learning},
  pages = 	 {57101--57115},
  year = 	 {2024},
  editor = 	 {Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix},
  volume = 	 {235},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {21--27 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v235/main/assets/yin24e/yin24e.pdf},
  url = 	 {https://proceedings.mlr.press/v235/yin24e.html},
  abstract = 	 {Large Language Models (LLMs), renowned for their remarkable performance across diverse domains, present a challenge due to their colossal model size when it comes to practical deployment. In response to this challenge, efforts have been directed toward the application of traditional network pruning techniques to LLMs, uncovering a massive number of parameters can be pruned in one-shot without hurting performance. Building upon insights gained from pre-LLM models, particularly BERT-level language models, prevailing LLM pruning strategies have consistently adhered to the practice of uniformly pruning all layers at equivalent sparsity levels, resulting in robust performance. However, this observation stands in contrast to the prevailing trends observed in the field of vision models, where non-uniform layerwise sparsity typically yields substantially improved results. To elucidate the underlying reasons for this disparity, we conduct a comprehensive analysis of the distribution of token features within LLMs. In doing so, we discover a strong correlation with the emergence of outliers, defined as features exhibiting significantly greater magnitudes compared to their counterparts in feature dimensions. Inspired by this finding, we introduce a novel LLM pruning methodology that incorporates a tailored set of non-uniform layerwise sparsity ratios specifically designed for LLM pruning, termed as Outlier Weighed Layerwise sparsity (OWL). The sparsity ratio of OWL is directly proportional to the outlier ratio observed within each layer, facilitating a more effective alignment between layerwise weight sparsity and outlier ratios. Our empirical evaluation, conducted across the LLaMA-V1/V2, Vicuna, OPT, and Mistral, spanning various benchmarks, demonstrates the distinct advantages offered by OWL over previous methods. For instance, OWL exhibits a remarkable performance gain, surpassing the state-of-the-art Wanda and SparseGPT by 61.22 and 6.80 perplexity at a high sparsity level of 70%, respectively, while delivering 2.6$\times$ end-to-end inference speed-up in the DeepSparse inference engine. Code is available at https://github.com/luuyin/OWL.git.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/