Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models

Peijie DongLujun LiZhenheng TangXiang LiuXinglin PanQiang WangXiaowen Chu

article2024ICML75 citations

Presents an automated framework that uses genetic programming to discover effective symbolic post-training pruning metrics from scratch, achieving state-of-the-art compression performance on large language models without requiring retraining or weight updates.

Listen

Large Language Models deliver state-of-the-art natural language capabilities but demand immense computational and memory resources, creating substantial bottlenecks for deployment. While network pruning reduces model size by removing redundant parameters, traditional techniques require costly retraining cycles. Recent post-training pruning methods eliminate the need for retraining, yet they rely on handcrafted importance metrics derived through extensive human trial and error. Furthermore, minor format changes in these formulas can cause severe performance instability, making manual formula discovery inefficient and unreliable.

To overcome these hurdles, the article introduces Pruner-Zero, an automated search framework that evolves symbolic pruning metrics from scratch using genetic programming. The authors formulate pruning metric discovery as a symbolic regression problem, operating over mathematical primitives alongside network statistics including weights, activations, and pre-computed gradients while excluding computationally prohibitive second-order curvature calculations. To prevent the search space from becoming clogged with mathematically redundant expressions, the framework incorporates an opposing operation simplification strategy that detects and removes counteracting mathematical operations.

The search discovered an optimal metric combining squared absolute weights and min-max scaled gradient magnitudes. Evaluated on the LLaMA, LLaMA-2, and OPT model families across language modeling and standard zero-shot reasoning benchmarks, Pruner-Zero demonstrates substantial performance gains. Under a 50% unstructured parameter reduction on LLaMA models, it consistently achieves lower language perplexity than leading post-training baselines such as Wanda and SparseGPT without requiring any weight updates or retraining. In structured 2:4 and 4:8 hardware-friendly pruning patterns, Pruner-Zero also preserves predictive quality more effectively than existing methods, with performance degradation diminishing on larger models such as 70-billion parameter variants. Additionally, Pruner-Zero achieves these results while pruning models in roughly half the execution time demanded by second-order optimization methods.

These findings indicate that automated, data-driven discovery can construct superior pruning metrics compared to human intuition alone. By delivering lower operational degradation at 50% sparsity without requiring iterative weight recalculations, Pruner-Zero offers organizations a practical, cost-effective avenue to compress massive models for production hardware. For resource-constrained deployments, the authors further demonstrate that lightweight parameter-efficient fine-tuning can rapidly recover residual accuracy losses after pruning.

Organizations seeking to deploy compressed large language models should consider adopting Pruner-Zero for post-training compression workflows, particularly when deployment speed and budget preclude full model retraining. However, decision-makers should note that the automated search was performed primarily using 50% sparsity on a single model family, and evaluations centered mainly on perplexity and standard zero-shot benchmarks. Further testing and empirical validation are recommended on domain-specific workloads, structured pruning targets, and advanced reasoning tasks before wide-scale deployment.

arXiv: 2406.02924
Cover for Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models

Abstract

Despite the remarkable capabilities, Large Language Models (LLMs) face deployment challenges due to their extensive size. Pruning methods drop a subset of weights to accelerate, but many of them require retraining, which is prohibitively expensive and computationally demanding. Recently, post-training pruning approaches introduced novel metrics, enabling the pruning of LLMs without retraining. However, these metrics require the involvement of human experts and tedious trial and error. To efficiently identify superior pruning metrics, we develop an automatic framework for searching symbolic pruning metrics using genetic programming. In particular, we devise an elaborate search space encompassing the existing pruning metrics to discover the potential symbolic pruning metric. We propose an opposing operation simplification strategy to increase the diversity of the population. In this way, Pruner-Zero allows auto-generation of symbolic pruning metrics. Based on the searched results, we explore the correlation between pruning metrics and performance after pruning and summarize some principles. Extensive experiments on LLaMA and LLaMA-2 on language modeling and zero-shot tasks demonstrate that our Pruner-Zero obtains superior performance than SOTA post-training pruning methods. Code at: https://github.com/pprp/Pruner-Zero.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Language Model Pruning
  • 2.2. Symbolic Regression
  • 3. Pruner-Zero Framework
  • 3.1. Pruner-Zero Search Space Design
  • 3.2. Genetic Programming Framework
  • 4. Experiments
  • 4.1. Models and Implementation Details
  • 4.2. Language Modeling
  • 4.3. Zero-shot Tasks
  • 4.4. Evaluation of In-Context Learning
  • 4.5. Ablation Study
  • 4.6. Analysis
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Related Work
  • B. Motivation of Searching Pruning Metric
  • C. Detailed Search Space
  • C.1. Search Space Composition
  • C.2. Operation Vocabulary
  • C.3. Searched Metrics with Equations
  • C.4. More Discussion over the Discovered SPM
  • D. Genetic Programming in SPM Optimization
  • E. Expanding the Results
  • E.1. Detailed Correlation Analysis
  • E.2. Expanding the Zero-Shot Tasks
  • E.3. Comparison with Previous Pruning Methods
  • E.4. Comparative Analysis of Pruning Techniques
  • E.5. Pruning Time Comparison
  • E.6. Performance on Higher Sparsity Ratio
  • E.7. Robustness Assessment
  • F. Searched Metrics with Expressions
  • G. Limitations and Future Works

Knowls

  1. Knowl 1 — Layer-wise post-training pruning objective

    model/method

    Pruner-Zero addresses post-training pruning by solving an independent reconstruction problem for each linear layer. For layer ll, let WlW_l be the pretrained weight matrix, XlX_l the calibration activations, MlM_l a binary mask with the prescribed sparsity ratio, and ⊙\odot element-wise multiplication. The layer mask is chosen to minimize the change in layer output:

    arg⁡min⁡Ml∥WlXl−(Ml⊙Wl)Xl∥22.\arg\min_{M_l}\left\|W_lX_l-(M_l\odot W_l)X_l\right\|_2^2.

    The mask is induced by a symbolic pruning metric SS: a ranking function f(S,Wl,ϕ)f(S,W_l,\phi) assigns an importance score to every weight using WlW_l, XlX_l, and gradient statistics GlG_l, then retains the highest-scoring weights until the target sparsity ratio ϕ\phi is reached. Unlike methods that update surviving weights during pruning, the main Pruner-Zero procedure only ranks and removes weights.

  2. Knowl 2 — Unified symbolic pruning metric search space

    model/method

    A symbolic pruning metric is represented as an expression tree whose leaves are LLM statistics and whose internal nodes are primitive mathematical operations. The leaf inputs are weights WW, calibration activations XX, and gradients GG; the Hessian is excluded because its computation scales quadratically with hidden dimension, approximately O(dhidden2)O(d_{\mathrm{hidden}}^2). The vocabulary contains 17 primitives: 13 unary operations, including square, negation, absolute value, logarithm, exponential, square root, tanh, power, skip, min-max scaling, z-score scaling, ℓ2\ell_2 norm, and ℓ1\ell_1 norm; and four binary operations: addition, subtraction, multiplication, and division.

    Every expression tree produces a matrix with the same shape as the layer weights so that its entries can directly rank individual weights. Unary operations have one child and binary operations have two; a placeholder is used for the unused second child of a unary operation. Because activations do not generally have the same shape as weights, the initialization procedure uses activation norms such as ∥X∥2\|X\|_2 and prevents incompatible unary-operation placements. Gradients are collected from 128 calibration samples and stored before the search, so the search phase need not recompute them.

  3. Knowl 3 — Genetic-programming search with opposing-operation simplification

    algorithm

    Pruner-Zero evolves symbolic expression trees using genetic programming. The search uses an initial population of 50 trees with depths from 3 to 5, 300 iterations, tournament selection from the 10 best candidates, and mutation probability p=0.5p=0.5. Fitness is the WikiText2 perplexity of LLaMA-2-7B after 50% unstructured pruning; lower perplexity is better.

    Input: Search-space primitives, population size 50, sample ratio r, top-k 10, maximum iterations 300, tree depths 3 to 5, mutation probability 0.5
    Output: Symbolic pruning tree with the lowest post-pruning perplexity
    Initialize 50 symbolic trees with valid shapes and depths from 3 to 5
    Evaluate the initial population by post-training pruning and perplexity
    for each of 300 iterations do
        Sample r times the population size candidate trees
        Keep the 10 candidates with lowest perplexity
        Randomly select two parents from these 10 candidates
        Select random subtrees in the two parents and exchange them
        Mutate operation nodes with probability 0.5, preserving unary or binary type
        Apply opposing-operation simplification to the offspring
        if the simplified offspring is equivalent to either parent then
            Replace it with a newly sampled valid random tree
        end if
        Evaluate the offspring after pruning and append it to the population
        Remove the tree with the highest perplexity
    end for
    return the best tree

    Opposing-operation simplification removes mathematically redundant antagonistic pairs such as exponential with logarithm or subtraction with negation. This prevents a small population from filling with different trees that compute equivalent metrics. In the reported 50-tree, 300-iteration search, omitting this simplification increased the best perplexity from 6.7079 with simplification to 6.8395 without it.

  4. Knowl 4 — Discovered Pruner-Zero pruning metric

    equation

    The best symbolic metric found by the evolutionary search is reported as

    SPZ(W,G)=∣W∣⊙∣W∣⊙σ(∣G∣),S_{\mathrm{PZ}}(W,G)=\sqrt{|W|\odot|W|}\odot\sigma(|G|),

    where WW and GG are element-wise weight and gradient matrices for a layer, ∣⋅∣|\cdot| is element-wise absolute value, ⊙\odot is element-wise multiplication, the square root is element-wise, and σ\sigma is min-max scaling of gradient magnitudes to a common range. The resulting matrix has one saliency value per weight; the smallest saliency values are pruned. The metric requires calibration-derived gradients but no weight update or retraining during pruning.

  5. Knowl 5 — Importance analysis and design principles for pruning metrics

    theoretical result

    The paper relates weight importance to the error caused by setting one parameter wmw_m to zero. If gmg_m is the gradient associated with wmw_m, a tractable first-order approximation gives the importance score Im(1)=(gmwm)2I_m^{(1)}=(g_mw_m)^2, or in vector form I(W)=(W⊙G)2I(W)=(W\odot G)^2. The second-order Hessian formulation is considered but rejected for the automated search because of its computational cost.

    For LLaMA-2-7B, the measured kurtosis of the weights is 0.7734 and that of the gradients is 5.8203, indicating substantially different distributions and motivating explicit scaling of their contributions. Analysis of low-perplexity evolved metrics found three recurring design principles: multiplication is associated with better pruning metrics; square and square-root operations frequently occur together and can have counteractive effects; and min-max scaling has low correlation with other operations and behaves as a comparatively stable normalization component.

  6. Knowl 6 — Experimental protocol for evaluating Pruner-Zero

    experimental setup

    The main evaluation applies the searched metric to LLaMA-7B, 13B, 30B, and 65B and LLaMA-2-7B, 13B, and 70B. Pruning is applied uniformly to all linear layers using unstructured 50% sparsity and N:M structured 4:8 and 2:4 sparsity, each retaining half of the entries in the relevant block. WikiText2 calibration data provide activation and gradient statistics, with 128 calibration samples used for the main comparisons. Language-model quality is measured by WikiText2 validation perplexity; generalization is assessed with mean zero-shot accuracy over BoolQ, RTE, HellaSwag, WinoGrande, ARC-Easy, ARC-Challenge, and OpenBookQA. Baselines are magnitude pruning, SparseGPT, and Wanda; SparseGPT updates weights, whereas magnitude pruning, Wanda, and Pruner-Zero do not.

  7. Knowl 7 — WikiText2 language-modeling results

    data/table

    The reported WikiText2 comparison on page 6 tests dense models and three pruning regimes. Lower perplexity is better. Pruner-Zero consistently improves on magnitude pruning and Wanda without weight updates, and generally improves on SparseGPT despite SparseGPT using weight updates.

    Method Update Sparsity LLaMA-7B LLaMA-13B LLaMA-30B LLaMA-65B LLaMA-2-7B LLaMA-2-13B LLaMA-2-70B
    Dense – 0% 5.68 5.09 4.77 3.56 5.12 4.57 3.12
    Magnitude No 50% 17.29 20.21 7.54 5.90 14.89 6.37 4.98
    SparseGPT Yes 50% 7.22 6.21 5.31 4.57 6.51 5.63 3.98
    Wanda No 50% 7.26 6.15 5.24 4.57 6.42 5.56 3.98
    Pruner-Zero No 50% 6.95 5.94 5.01 4.33 6.26 5.36 3.82
    Magnitude No 4:8 16.84 13.84 7.62 6.36 16.48 6.76 5.54
    SparseGPT Yes 4:8 8.61 7.40 6.17 5.38 8.12 6.60 4.59
    Wanda No 4:8 8.57 7.40 5.97 5.30 7.97 6.55 4.47
    Pruner-Zero No 4:8 8.12 6.81 5.65 4.92 7.67 6.10 4.31
    Magnitude No 2:4 42.13 18.37 9.10 7.11 54.59 8.33 6.33
    SparseGPT Yes 2:4 11.00 9.11 7.16 6.28 10.17 8.32 5.40
    Wanda No 2:4 11.53 9.58 6.90 6.25 11.02 8.27 5.16
    Pruner-Zero No 2:4 10.61 8.11 6.51 5.67 10.52 7.41 4.81

    At 50% unstructured sparsity, Pruner-Zero reaches perplexities of 6.95, 5.94, 5.01, and 4.33 on LLaMA-7B through 65B, and 6.26, 5.36, and 3.82 on LLaMA-2-7B through 70B. Its advantage over Wanda is also retained under both structured formats.

  8. Knowl 8 — Zero-shot accuracy after pruning

    data/table

    The paper's zero-shot comparison on page 7 reports mean accuracy over seven common-sense tasks. Pruner-Zero is evaluated without weight updates and generally gives the highest mean accuracy among the pruned methods, sometimes exceeding the dense baseline for larger models.

    Method Update Sparsity LLaMA-7B LLaMA-13B LLaMA-30B LLaMA-65B LLaMA-2-7B LLaMA-2-13B LLaMA-2-70B
    Dense – 0% 59.99 62.59 65.38 66.97 59.71 63.03 67.08
    Magnitude No 50% 46.94 47.61 53.83 62.74 51.14 52.85 60.93
    SparseGPT Yes 50% 54.94 58.61 63.09 66.30 56.24 60.72 67.28
    Wanda No 50% 54.21 59.33 63.60 66.67 56.24 60.83 67.03
    Pruner-Zero No 50% 59.56 62.67 67.49 69.81 58.87 64.83 71.10
    Magnitude No 4:8 46.03 50.53 53.53 62.17 50.64 52.81 60.28
    SparseGPT Yes 4:8 52.80 55.99 60.79 64.87 53.80 59.15 65.84
    Wanda No 4:8 52.76 56.09 61.00 64.97 52.49 58.75 66.06
    Pruner-Zero No 4:8 56.24 59.03 64.04 68.04 55.82 61.97 69.94
    Magnitude No 2:4 44.73 48.00 53.16 61.28 45.58 49.89 59.95
    SparseGPT Yes 2:4 50.60 53.22 58.91 62.57 50.94 54.86 63.89
    Wanda No 2:4 48.53 52.30 59.21 62.84 48.75 55.03 64.14
    Pruner-Zero No 2:4 52.06 56.78 62.00 65.42 52.02 58.38 67.69

    At 50% unstructured sparsity, Pruner-Zero obtains mean accuracies of 59.56, 62.67, 67.49, and 69.81 on LLaMA-7B through 65B, compared with dense scores of 59.99, 62.59, 65.38, and 66.97. On LLaMA-2, its means are 58.87, 64.83, and 71.10 for 7B, 13B, and 70B, with the 70B pruned model exceeding the dense score of 67.08.

  9. Knowl 9 — Generalization to other model families and sparsity levels

    empirical result

    The searched metric transfers beyond the LLaMA-2-7B search model. On OPT models with 50% unstructured sparsity and no weight update, Pruner-Zero achieves WikiText2 perplexities of 37.69, 35.91, 18.19, 13.85, 11.86, and 11.32 for 125M, 350M, 1.3B, 2.7B, 6.7B, and 13B models, respectively. The corresponding Wanda values are 38.96, 35.92, 19.12, 14.28, 11.94, and 11.42. On Tiny-LLaMA, the perplexities for Pruner-Zero are 10.65 at unstructured 50%, 22.17 at 2:4, and 14.47 at 4:8 sparsity, compared with 11.21, 27.17, and 16.18 for Wanda.

    At 60% unstructured sparsity, Pruner-Zero remains competitive on the larger LLaMA and LLaMA-2 models. Its perplexities are 9.83, 7.66, 6.22, and 5.31 on LLaMA-7B, 13B, 30B, and 65B, and 9.58, 6.90, and 4.64 on LLaMA-2-7B, 13B, and 70B. These results support transfer across architectures, model sizes, and more aggressive sparsity, although the paper notes that the main search was conducted at 50% unstructured sparsity.

  10. Knowl 10 — Search and pruning efficiency

    empirical result

    Evolutionary search converges in fewer than 100 iterations in the reported comparison, whereas random search requires approximately 300 iterations. Varying WikiText2 calibration size from 8 to 256 samples improves all methods; Pruner-Zero has perplexity below 7 once more than 16 samples are used, and 128 samples are selected for the main experiments.

    For LLaMA-2-7B at 50% unstructured sparsity, the measured pruning-only times and WikiText2 perplexities are:

    Method Pruning time (s) Perplexity
    Magnitude 0.92 17.29
    Wanda 402.25 7.26
    SparseGPT 1178.62 7.22
    Pruner-Zero 444.12 6.95

    Thus Pruner-Zero is reported as roughly twice as fast as SparseGPT and about 10% slower than Wanda, while obtaining lower perplexity than both. Its offline gradient preprocessing avoids real-time gradient computation during pruning.

Coverage note — Detailed per-task zero-shot matrices, GSM8K in-context learning, LoRA fine-tuning, BERT-transfer experiments, seed-variation tables, the complete list of evolved expressions, and the paper's extended limitation/future-work discussion were omitted because they are supporting analyses rather than the ten most load-bearing contributions.

References

  1. 1.Abdin, M., Jacobs, S. A., and etc., A. A. A. Phi-3 technical report: A highly capable language model locally on your phone. ArXiv, abs/2404.14219, 2024.
  2. 2.Akhauri, Y., Munoz, J. P., Jain, N., and Iyer, R. EZNAS: Evolving zero-cost proxies for neural architecture scoring. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), NeurIPS, 2022.
  3. 3.Ashkboos, S., Croci, M. L., do Nascimento, M. G., Hoefler, T., and Hensman, J. SliceGPT: Compress large language models by deleting rows and columns. In ICLR, 2024.
  4. 4.Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M. The lottery ticket hypothesis for pretrained bert networks. In NeurIPS, 2020.
  5. 5.Chen, T., Liang, L., DING, T., Zhu, Z., and Zharkov, I. OTOv2: Automatic, generic, user-friendly. In ICLR, 2023.
  6. 6.Chijiwa, D., Yamaguchi, S. y., Ida, Y., Umakoshi, K., and INOUE, T. Pruning randomly initialized neural networks with iterative randomization. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), NeurIPS, volume 34, pp. 4503–4513. Curran Associates, Inc., 2021.
  7. 7.Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K. Boolq: Exploring the surprising difficulty of natural yes/no questions. ArXiv, abs/1905.10044, 2019.
  8. 8.Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv, abs/1803.05457, 2018.
  9. 9.Das, R. J., Ma, L., and Shen, Z. Beyond size: How gradients shape pruning decisions in large language models. ArXiv, abs/2311.04902, 2023.
  10. 10.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  11. 11.Diao, E., Wang, G., Zhang, J., Yang, Y., Ding, J., and Tarokh, V. Pruning deep neural networks from a sparsity perspective. In ICLR, 2023.
  12. 12.Dong, P., Niu, X., Li, L., Xie, L., Zou, W., Ye, T., Wei, Z., and Pan, H. Prior-guided one-shot neural architecture search. arXiv preprint arXiv:2206.13329, 2022.
  13. 13.Dong, P., Li, L., and Wei, Z. Diswot: Student architecture search for distillation without training. In CVPR, 2023a.
  14. 14.Dong, P., Li, L., Wei, Z., Niu, X., Tian, Z., and Pan, H. Emq: Evolving training-free proxies for automated mixed-precision quantization. In ICCV, pp. 17076–17086, 2023b.
  15. 15.Dong, P., Li, L., Pan, X., Wei, Z., Liu, X., Wang, Q., and Chu, X. Parzc: Parametric zero-cost proxies for efficient nas. arXiv preprint arXiv:2402.02105, 2024.
  16. 16.Fang, G., Ma, X., Song, M., Mi, M. B., and Wang, X. Depgraph: Towards any structural pruning. CVPR, pp. 16091–16101, 2023.
  17. 17.Frantar, E. and Alistarh, D. Sparsegpt: Massive language models can be accurately pruned in one-shot. In ICML, 2023.
  18. 18.Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. Gptq: Accurate post-training quantization for generative pretrained transformers. In ICLR, 2023.
  19. 19.Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, September 2021.
  20. 20.Han, S., Pool, J., Tran, J., and Dally, W. J. Learning both weights and connections for efficient neural network. In NeurIPS, 2015.
  21. 21.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. ICLR, 2016a.
  22. 22.Han, S., Mao, H., and Dally, W. J. Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding. In ICLR, 2016b.
  23. 23.He, X., Yao, J., Wang, Y., Tang, Z., Cheung, K. C., See, S., Han, B., and Chu, X. Nas-lid: Efficient neural architecture search with local intrinsic dimension. In AAAI, 2022.
  24. 24.Hoang, D. N., Liu, S., Marculescu, R., and Wang, Z. revisiting pruning at initialization through the lens of ramanujan graph. In ICLR, 2023.
  25. 25.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In ICLR, 2022.
  26. 26.Hu, Y., Wang, X., Li, L., and Gu, Q. Improving one-shot nas with shrinking-and-expanding supernet. Pattern Recognition, 2021.
  27. 27.Koza, J. R. Genetic programming as a means for programming computers by natural selection. Statistics and Computing, 4:87–112, 1994.
  28. 28.Kwon, W., Kim, S., Mahoney, M. W., Hassoun, J., Keutzer, K., and Gholami, A. A fast post-training pruning framework for transformers. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), NeurIPS, 2022.
  29. 29.Lee, N., Ajanthan, T., and Torr, P. H. Snip: Single-shot network pruning based on connection sensitivity. In ICLR, 2019.
  30. 30.Li, L. Self-regulated feature learning via teacher-free feature distillation. In ECCV, 2022.
  31. 31.Li, L. and Jin, Z. Shadow knowledge distillation: Bridging offline and online knowledge transfer. In NeuIPS, 2022.
  32. 32.Li, L., Dong, P., Wei, Z., and Yang, Y. Automated knowledge distillation via monte carlo tree search. In ICCV, 2023.
  33. 33.Li, L., Dong, P., Li, A., Wei, Z., and Yang, Y. Kd-zero: Evolving knowledge distiller for any teacher-student pairs. NeuIPS, 2024a.
  34. 34.Li, S., Han, X., and Bai, J. Nuteprune: Efficient progressive pruning with numerous teachers for large language models. ArXiv, abs/2402.09773, 2024b.
  35. 35.Li, S., Ning, X., Wang, L., Liu, T., Shi, X., Yan, S., Dai, G., Yang, H., and Wang, Y. Evaluating quantized large language models. ICML, abs/2402.18158, 2024c.
  36. 36.Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S. Awq: Activation-aware weight quantization for llm compression and acceleration. In MLSys, 2024.
  37. 37.Liu, S.-y., Liu, Z., Huang, X., Dong, P., and Cheng, K.-T. LLM-FP4: 4-bit floating-point quantized transformers. In EMNLP, pp. 592–605, Singapore, December 2023a. Association for Computational Linguistics. URL https://aclanthology.org/2023.emnlp-main.39.
  38. 38.Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T. Rethinking the value of network pruning. In ICLR, 2019.
  39. 39.Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V. Llm-qat: Data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888, 2023b.
  40. 40.Louizos, C., Welling, M., and Kingma, D. P. Learning sparse neural networks through l 0 regularization. In ICLR, 2018.
  41. 41.Lu, L., Chen, Z., Lu, X., Rao, Y., Li, L., and Pang, S. Uniads: Universal architecture-distiller search for distillation gap. In AAAI, 2024.
  42. 42.Lu, M., Luo, X., Chen, T., Chen, W., Liu, D., and Wang, Z. Learning pruning-friendly networks via frank-wolfe: One-shot, any-sparsity, and no retraining. In ICLR, 2022.
  43. 43.Luccioni, A. S., Viguier, S., and Ligozat, A.-L. Estimating the carbon footprint of bloom, a 176b parameter language model. Journal of Machine Learning Research, 24(253): 1–15, 2023.
  44. 44.Ma, X., Fang, G., and Wang, X. Llm-pruner: On the structural pruning of large language models. In NeurIPS, 2023.
  45. 45.Martius, G. and Lampert, C. H. Extrapolation and learning equations. ArXiv, abs/1610.02995, 2016.
  46. 46.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=Byj72udxe.
  47. 47.Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. Can a suit of armor conduct electricity? a new dataset for open book question answering. In EMNLP, 2018.
  48. 48.ModelTC. Lightllm. https://github.com/ModelTC/lightllm, 2023. A Python-based LLM inference and serving framework.
  49. 49.OpenAI. Gpt-4 technical report, 2023.
  50. 50.Pham, H., Guan, M. Y., Zoph, B., Le, Q. V., and Dean, J. Efficient neural architecture search via parameter sharing. In ICML, 2018.
  51. 51.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints, 2019.
  52. 52.Real, E., Liang, C., So, D., and Le, Q. Automl-zero: Evolving machine learning algorithms from scratch. In ICML, 2020.
  53. 53.Ren, S. and Zhu, K. Pruning pre-trained language models with principled importance and self-regularization. In ACL, pp. 8995–9008, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.573.
  54. 54.Sahoo, S. S., Lampert, C. H., and Martius, G. Learning equations for extrapolation and control. In ICML, 2018.
  55. 55.Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. Winogrande. Communications of the ACM, 64:99–106, 2019.
  56. 56.Sanh, V., Wolf, T., and Rush, A. Movement pruning: Adaptive sparsity by fine-tuning. NeurIPS, 33:20378–20389, 2020.
  57. 57.Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O. Green ai. Communications of the ACM, 63(12):54–63, 2020.
  58. 58.Sharma, P., Ash, J. T., and Misra, D. The truth is in there: Improving reasoning in language models with layer-selective rank reduction. In ICLR, 2024.
  59. 59.Sreenivasan, K., yong Sohn, J., Yang, L., Grinde, M., Nagle, A., Wang, H., Xing, E., Lee, K., and Papailiopoulos, D. Rare gems: Finding lottery tickets at initialization. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), NeurIPS, 2022.
  60. 60.Srinivas, S. and Babu, R. V. Data-free parameter pruning for deep neural networks. In BMVC, 2015.
  61. 61.Sun, M., Liu, Z., Bair, A., and Kolter, J. Z. A simple and effective pruning approach for large language models. In ICLR, 2024.
  62. 62.Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S. Pruning neural networks without any data by iteratively conserving synaptic flow. NeurIPS, 33:6377–6389, 2020.
  63. 63.Tang, Z., Wang, Y., Wang, Q., and Chu, X. The impact of gpu dvfs on the energy and performance of deep learning: An empirical study. In Proceedings of the Tenth ACM International Conference on Future Energy Systems, e-Energy ’19, pp. 315–325, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450366717. doi: 10.1145/3307772.3328315.
  64. 64.Tang, Z., Shi, S., Chu, X., Wang, W., and Li, B. Communication-efficient distributed deep learning: A comprehensive survey. arXiv preprint arXiv:2003.06307, 2020.
  65. 65.Tang, Z., Wang, Y., He, X., Zhang, L., Pan, X., Wang, Q., Zeng, R., Zhao, K., Shi, S., He, B., et al. Fusionai: Decentralized training and deploying llms with massive consumer-level gpus. arXiv preprint arXiv:2309.01172, 2023.
  66. 66.Tang, Z., Zhang, Y., Shi, S., Tian, X., Liu, T., Han, B., and Chu, X. Fedimpro: Measuring and improving client update in federated learning. In ICLR, 2024.
  67. 67.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., `Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  68. 68.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  69. 69.van der Ouderaa, T. F. A., Nagel, M., van Baalen, M., Asano, Y. M., and Blankevoort, T. The llm surgeon. In ICLR, 2024.
  70. 70.Virgolin, M. and Pissis, S. P. Symbolic regression is np-hard. arXiv preprint arXiv:2207.01018, 2022.
  71. 71.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. Glue: A multi-task benchmark and analysis platform for natural language understanding. In BlackboxNLP@EMNLP, 2018.
  72. 72.Wang, H., Qin, C., Zhang, Y., and Fu, Y. R. Neural pruning via growing regularization. ICLR, abs/2012.09243, 2020.
  73. 73.Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=_VjQlMeSB_J.
  74. 74.Wei, Z., Pan, H., Li, L., Dong, P., Tian, Z., Niu, X., and Li, D. Tvt: Training-free vision transformer search on tiny datasets. arXiv preprint arXiv:2311.14337, 2023.
  75. 75.Wei, Z., Dong, P., Hui, Z., Li, A., Li, L., Lu, M., Pan, H., and Li, D. Auto-prox: Training-free vision transformer architecture search via automatic proxy discovery. In AAAI, 2024.
  76. 76.Werner, M., Junginger, A., Hennig, P., and Martius, G. Informed equation learning. ArXiv, abs/2105.06331, 2021.
  77. 77.Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S. SmoothQuant: Accurate and efficient post-training quantization for large language models. In ICML, 2023.
  78. 78.Xiaolong, L., Lujun, L., Chao, L., and Yao, A. Norm: Knowledge distillation via n-to-one representation matching. In ICLR, 2023.
  79. 79.Xu, P., Shao, W., Chen, M., Tang, S., Zhang, K., Gao, P., An, F., Qiao, Y., and Luo, P. BESA: Pruning large language models with blockwise parameter-efficient sparsity allocation. In ICLR, 2024.
  80. 80.Yang, Z., Cui, Y., Yao, X., and Wang, S. Gradient-based intra-attention pruning on pre-trained language models. In ACL, 2022.
  81. 81.Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), NeurIPS, volume 35, pp. 27168–27183. Curran Associates, Inc., 2022.
  82. 82.Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S. Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning LLMs to high sparsity. In ICML, 2024.
  83. 83.Yvinec, E., Dapogny, A., Cord, M., and Bailly, K. Red++ : Data-free pruning of deep neural networks via input splitting and output merging. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45:3664–3676, 2021.
  84. 84.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In ACL, 2019.
  85. 85.Zhang, P., Zeng, G., Wang, T., and Lu, W. Tinylama: An open-source small language model. ArXiv, abs/2401.02385, 2024a.
  86. 86.Zhang, Q., Zuo, S., Liang, C., Bukharin, A., He, P., Chen, W., and Zhao, T. Platon: Pruning large transformer models with upper confidence bound of weight importance. In ICML, pp. 26809–26823. PMLR, 2022a.
  87. 87.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M. T., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. Opt: Open pre-trained transformer language models. ArXiv, abs/2205.01068, 2022b.
  88. 88.Zhang, Y., Bai, H., Lin, H., Zhao, J., Hou, L., and Cannistraci, C. V. Plug-and-play: An efficient post-training pruning method for large language models. In ICLR, 2024b.
  89. 89.Zhu, C., Li, L., Wu, Y., and Sun, Z. Saswot: Real-time semantic segmentation architecture search without training. In AAAI, 2024.
  90. 90.Zimmer, M., Andoni, M., Spiegel, C., and Pokutta, S. Perp: Rethinking the prune-retrain paradigm in the era of llms. ArXiv, abs/2312.15230, 2023.
  91. 91.Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V. Learning transferable architectures for scalable image recognition. In CVPR, 2018.

Citation

MLA
Dong, P., et al. “Pruner-Zero: Evolving Symbolic Pruning Metric from Scratch for Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2406.02924v1.
APA
Dong, P., Li, L., Tang, Z., Liu, X., Pan, X., Wang, Q., & Chu, X. (2024). Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models. arXiv. http://arxiv.org/abs/2406.02924v1
Chicago
Dong, P., L. Li, Z. Tang, et al. 2024. “Pruner-Zero: Evolving Symbolic Pruning Metric from Scratch for Large Language Models”. arXiv. http://arxiv.org/abs/2406.02924v1.
Harvard
Dong, P. et al. (2024) “Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.02924v1.
Vancouver
1. Dong P, Li L, Tang Z, Liu X, Pan X, Wang Q, Chu X (2024) Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models. arXiv

BibTeX

@article{dong2024pruner,
  title = {Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models},
  author = {Dong, Peijie and Li, Lujun and Tang, Zhenheng and Liu, Xiang and Pan, Xinglin and Wang, Qiang and Chu, Xiaowen},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.02924v1},
  eprint = {2406.02924}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/