SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Jiwon SongKyungseok OhTaesu KimHyungjun KimYulhwa KimJae-Joon Kim
Proposes a block-level pruning method called SLEB that removes redundant transformer blocks based on layer output similarity, achieving real hardware acceleration and low perplexity degradation without requiring specialized sparse matrix support.
Large language models provide powerful natural language processing capabilities, but their immense size creates substantial memory and computational challenges during real-world deployment. Traditional compression techniques like weight pruning often fail to deliver practical acceleration because modern hardware struggles with sparse operations, while dynamic methods like early exiting cannot reduce memory footprints and perform poorly when processing multiple user requests simultaneously.
The article demonstrates and evaluates SLEB, a streamlined pruning method designed to accelerate large language models by identifying and removing entire redundant transformer blocks without requiring expensive model retraining.
The authors developed an iterative, training-free approach that evaluates the importance of each architectural block based on total network output impact, pruning the most redundant blocks one at a time using a small set of 128 calibration text samples. They tested this method across several model architectures, including the OPT and LLaMA-2 model families ranging from 6.7 billion to 70 billion parameters, and benchmarked processing latency, multi-user throughput, language quality, and performance on standard reasoning tasks against leading pruning alternatives.
The evaluation yielded several key findings:
- Removing entire blocks translates directly into runtime speedups. On a 70-billion-parameter model, pruning 20% of the blocks yielded a 1.27 times improvement in throughput during generation and a 1.26 times reduction in prompt latency, whereas existing weight pruning methods produced negligible speedup or even slowed down generation.
- Language quality and reasoning accuracy remained well-preserved up to a 20% pruning ratio, outperforming existing fine-grained pruning methods across zero-shot benchmarks.
- The method exhibited high stability across datasets and executed quickly, compressing a 70-billion-parameter model in approximately 1.5 hours on two enterprise graphics processors.
- Pruning entire blocks demonstrated complete compatibility with 4-bit post-training quantization, enabling compound compression without additional loss in language fluency.
These findings indicate that architectural redundancy at the block level can be exploited to achieve tangible operational cost reductions and lower response latency. Because the approach permanently deletes entire blocks, organizations can simultaneously shrink hardware memory footprints and improve batched request efficiency without needing specialized sparsity hardware or costly retraining cycles.
Organizations seeking to optimize large language model serving should consider adopting block-level pruning as an alternative to fine-grained weight pruning, particularly when targeting compression rates up to 20%. For maximum memory savings, teams should combine this approach with standard 4-bit weight quantization.
Confidence in these findings is high for models within the 6.7-billion to 70-billion parameter range up to moderate pruning targets. However, stakeholders should note that performance degrades significantly if more than 20% to 30% of blocks are removed without fine-tuning, and results may vary when applied to fundamentally different model architectures.
- Paper: A Fast Post-Training Pruning Framework for Transformers, Woosuk Kwon et al. (2022). Its training-free, calibration-based framework for pruning Transformer components provides a direct methodological precursor to SLEB’s post-training block removal.
- Paper: Structured Pruning Learns Compact and Accurate Models, Mengzhou Xia et al. (2022). CoFi’s joint pruning of entire Transformer layers and finer components introduces the structured-pruning choices that SLEB narrows to redundant blocks.
No sufficiently relevant recommendations were found.
