Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models
Peijie DongLujun LiZhenheng TangXiang LiuXinglin PanQiang WangXiaowen Chu
Presents an automated framework that uses genetic programming to discover effective symbolic post-training pruning metrics from scratch, achieving state-of-the-art compression performance on large language models without requiring retraining or weight updates.
Large Language Models deliver state-of-the-art natural language capabilities but demand immense computational and memory resources, creating substantial bottlenecks for deployment. While network pruning reduces model size by removing redundant parameters, traditional techniques require costly retraining cycles. Recent post-training pruning methods eliminate the need for retraining, yet they rely on handcrafted importance metrics derived through extensive human trial and error. Furthermore, minor format changes in these formulas can cause severe performance instability, making manual formula discovery inefficient and unreliable.
To overcome these hurdles, the article introduces Pruner-Zero, an automated search framework that evolves symbolic pruning metrics from scratch using genetic programming. The authors formulate pruning metric discovery as a symbolic regression problem, operating over mathematical primitives alongside network statistics including weights, activations, and pre-computed gradients while excluding computationally prohibitive second-order curvature calculations. To prevent the search space from becoming clogged with mathematically redundant expressions, the framework incorporates an opposing operation simplification strategy that detects and removes counteracting mathematical operations.
The search discovered an optimal metric combining squared absolute weights and min-max scaled gradient magnitudes. Evaluated on the LLaMA, LLaMA-2, and OPT model families across language modeling and standard zero-shot reasoning benchmarks, Pruner-Zero demonstrates substantial performance gains. Under a 50% unstructured parameter reduction on LLaMA models, it consistently achieves lower language perplexity than leading post-training baselines such as Wanda and SparseGPT without requiring any weight updates or retraining. In structured 2:4 and 4:8 hardware-friendly pruning patterns, Pruner-Zero also preserves predictive quality more effectively than existing methods, with performance degradation diminishing on larger models such as 70-billion parameter variants. Additionally, Pruner-Zero achieves these results while pruning models in roughly half the execution time demanded by second-order optimization methods.
These findings indicate that automated, data-driven discovery can construct superior pruning metrics compared to human intuition alone. By delivering lower operational degradation at 50% sparsity without requiring iterative weight recalculations, Pruner-Zero offers organizations a practical, cost-effective avenue to compress massive models for production hardware. For resource-constrained deployments, the authors further demonstrate that lightweight parameter-efficient fine-tuning can rapidly recover residual accuracy losses after pruning.
Organizations seeking to deploy compressed large language models should consider adopting Pruner-Zero for post-training compression workflows, particularly when deployment speed and budget preclude full model retraining. However, decision-makers should note that the automated search was performed primarily using 50% sparsity on a single model family, and evaluations centered mainly on perplexity and standard zero-shot benchmarks. Further testing and empirical validation are recommended on domain-specific workloads, structured pruning targets, and advanced reasoning tasks before wide-scale deployment.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT establishes the post-training regression and Hessian-based pruning formulation for billion-scale LLMs that Pruner-Zero directly seeks to automate and outperform.
- Paper: AMC: AutoML for Model Compression and Acceleration on Mobile Devices, Yihui He et al. (2018). AMC pioneers using automated search and reinforcement learning for model compression, laying the foundational paradigm for automated pruning space optimization.
- Paper: Second Order Derivatives for Network Pruning: Optimal Brain Surgeon, Babak Hassibi et al. (1992). Optimal Brain Surgeon defines the classical second-order saliency metrics that form the theoretical basis of the symbolic operators explored by Pruner-Zero.
- Paper: Optimal Brain Damage, Yann LeCun et al. (1989). Optimal Brain Damage introduces Hessian-based saliency metrics for network weight removal, serving as an essential precursor to modern LLM pruning metrics.
- Paper: SNIP: Single-shot Network Pruning based on Connection Sensitivity, Namhoon Lee et al. (2018). SNIP introduces connection sensitivity metrics based on gradients for one-shot pruning, supplying fundamental metric concepts evaluated in symbolic search spaces.
- Paper: Pruning neural networks without any data by iteratively conserving synaptic flow, Hidenori Tanaka et al. (2020). SynFlow establishes theoretical conservation laws and gradient-based flow metrics for data-free pruning that inform the primitive search space of pruning operations.
- Paper: Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression, Junyuan Hong et al. (2024). This study evaluates the trustworthiness, safety, and robustness consequences of applying post-training LLM pruning methods like those discovered in Pruner-Zero.
- Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). This paper investigates whether pruning strategies on modern LLMs provide superior initialization compared to training small dense models from scratch.
