Structured Pruning Learns Compact and Accurate Models
Mengzhou XiaZexuan ZhongDanqi Chen
Proposes CoFi, a structured pruning method that jointly eliminates coarse- and fine-grained Transformer components to achieve over tenfold inference speedups while matching the accuracy of expensive distillation baselines without requiring unlabeled pre-training data.
Modern natural language processing relies heavily on large pre-trained language models that demand substantial memory, storage, and computing power. To deploy these models effectively in production, organizations typically turn to model compression techniques: knowledge distillation, which trains smaller models to mimic larger ones, or model pruning, which removes redundant parameters from existing models. While distillation achieves high inference speedups, it requires computationally expensive pre-training over billions of unlabeled tokens. Conversely, existing structured pruning techniques create flexible sub-networks but struggle to achieve speedups beyond twofold to threefold acceleration.
The article introduces and evaluates CoFi (Coarse- and Fine-grained Pruning), a task-specific structured pruning framework designed to deliver highly parallelizable, compact models. The objective is to demonstrate that structured pruning can match or exceed the accuracy and inference latency of state-of-the-art distillation methods without requiring massive unlabeled datasets or prolonged pre-training cycles.
The approach operates by jointly learning pruning decisions across multiple granularities simultaneously, including entire attention and feed-forward layers, individual attention heads, and hidden dimensions. These components are regularized using sparsity constraints that allow the optimization process to dynamically determine the most efficient network shape. To preserve model quality during pruning, the framework introduces a dynamic layerwise distillation technique that automatically aligns and transfers intermediate representations from the full teacher model to the evolving pruned student model. The method was evaluated across eight GLUE benchmark tasks and the SQuAD question-answering dataset using standard base Transformer architectures.
The evaluation yielded several key findings. First, CoFi achieves over 10-fold inference speedups on GPUs with a 95% parameter sparsity rate while preserving over 90% of the baseline model accuracy. Second, it matches or outperforms leading distillation baselines like TinyBERT while slashing training time from roughly 350 GPU hours down to under 20 GPU hours on a single GPU. Third, joint coarse-and-fine pruning proved crucial: omitting whole-layer pruning severely degraded speedups from 12.1-fold to 7.0–8.3-fold at high sparsity, while omitting hidden-dimension pruning degraded model accuracy. Finally, analysis revealed that feed-forward layers contain substantially more redundancy than attention layers, showing a 71% reduction in intermediate dimensions compared to a 39% reduction in attention heads at 60% overall sparsity.
These results demonstrate that task-specific structured pruning provides an efficient, low-cost path to production-ready language models. Organizations can bypass the heavy computational overhead, engineering complexity, and data management risks associated with large-scale unlabeled data distillation. By adapting network depth and width directly to specific downstream tasks, teams can deploy models that run over ten times faster on standard hardware without significant predictive degradation.
Decision-makers and engineering teams should consider adopting multi-granularity structured pruning as a preferred compression pipeline for task-specific deployments, particularly where compute budgets or training turnarounds are constrained. For existing models, pruning fine-tuned weights directly with dynamic intermediate distillation offers the best balance of speed and retention. Future work should pilot this approach on broader generative architectures and investigate upstream task-agnostic pruning to establish general-purpose compact base models.
The primary limitation of the study is its focus on task-specific compression of encoder architectures, meaning the pruned models cannot be universally reused across distinct tasks without retraining. High confidence in these findings is supported by consistent empirical improvements across multiple standard benchmarks and clear ablation studies.
- Paper: Are Sixteen Heads Really Better than One?, Paul Michel et al. (2019). This work establishes the redundancy of multi-headed attention in Transformers and introduces gradient-based head importance ranking, providing the foundational basis for fine-grained structured pruning in CoFi.
- Paper: Learning Efficient Convolutional Networks through Network Slimming, Zhuang Liu et al. (2017). This paper establishes the structured pruning of network channels using sparse scaling masks, which directly informs CoFi's mask-based optimization across structural components.
- Paper: MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers, Wenhui Wang et al. (2020). MiniLM demonstrates how layerwise intermediate distillation from a teacher Transformer guides student representation learning, an approach adapted in CoFi to supervise pruned sub-networks.
- Paper: DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, Victor Sanh et al. (2019). DistilBERT introduces layer-dropping and task distillation principles for Transformer compression, providing the direct architectural benchmark that CoFi aims to match in latency without unlabeled pre-training data.
- Paper: Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al. (2015). This seminal text introduces knowledge distillation from larger models to smaller architectures, establishing the core distillation objective utilized during CoFi's optimization.
- Paper: Pruning Convolutional Neural Networks for Resource Efficient Inference, Pavlo Molchanov et al. (2016). This work introduces first-order Taylor expansion criteria for estimating module importance in structured pruning, underlying the gradient-based saliency methods used in fine-grained Transformer pruning.
- Paper: ZipLM: Inference-Aware Structured Pruning of Language Models, Eldar Kurtic et al. (2023). ZipLM advances structured pruning for language models by integrating real hardware latency benchmarks directly into the component removal and distillation objective.
- Paper: A Fast Post-Training Pruning Framework for Transformers, Woosuk Kwon et al. (2022). This paper develops a fast post-training structured pruning framework for Transformers that eliminates full retraining cycles through Fisher information search and layer reconstruction.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). DepGraph generalizes multi-granularity structured pruning across arbitrary neural architectures through automatic dependency modeling.
- Paper: LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation, Yixiao Li et al. (2023). LoSparse builds upon Transformer structural compression by combining structured pruning with low-rank decomposition to retain expressive neuron diversity.
- Paper: SliceGPT: Compress Large Language Models by Deleting Rows and Columns, Saleh Ashkboos et al. (2024). SliceGPT extends post-training structured reduction to modern large language models by slicing entire weight matrix rows and columns using computational invariance.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). SparseGPT scales one-shot post-training pruning techniques to modern generative LLMs with tens to hundreds of billions of parameters.
