Built independently by an author, for readers. Read the story and support ChapterPal

keyword

inference-aware structured pruning

Inference-aware structured pruning is a neural network compression technique that removes entire architectural components, such as attention heads, hidden dimensions, channels, or full layers, while explicitly optimizing for measured execution speed and latency in a target deployment environment. Unlike traditional structured pruning methods that rely on parameter counts or theoretical floating-point operations as proxies for efficiency, this approach directly evaluates the trade-off between task accuracy and real-world runtime on specific hardware. By iteratively identifying and eliminating the structural units that contribute least effectively to performance relative to their execution cost, it generates smaller, dense models guaranteed to achieve practical speedups and meet concrete inference latency requirements.

1 item

ZipLM: Inference-Aware Structured Pruning of Language Models

ZipLM: Inference-Aware Structured Pruning of Language Models

Eldar Kurtic, Elias Frantar, Dan Alistarh

OrganizationsInstitute of Science and Technology AustriaIST Austria & Neural Magic

Why you should read this

Proposes an inference-aware structured pruning method that optimizes the loss-runtime trade-off to generate an entire family of accurate, speedup-guaranteed encoder and decoder language models in a single training run.

The breakthrough performance of large language models (LLMs) comes with major computational footprints and high deployment costs. In this paper, we progress towards resolving this problem by proposing a novel structured compression approach for LLMs, called ZipLM. ZipLM achieves state-of-the-art accuracy-vs-speedup, while matching a set of desired target runtime speedups in any given inference environment. Specifically, given a model, a dataset, an inference environment, as well as a set of speedup targets, ZipLM iteratively identifies and removes components with the worst loss-runtime trade-off. Unlike prior methods that specialize in either the post-training/one-shot or the gradual compression setting, and only for specific families of models such as BERT (encoder) or GPT (decoder), ZipLM produces state-of-the-art compressed models across all these settings. Furthermore, ZipLM achieves superior results for a fraction of the computational cost relative to prior distillation and pruning techniques, making it a cost-effective approach for generating an entire family of smaller, faster, and highly accurate models, guaranteed to meet the desired inference specifications. In particular, ZipLM outperforms all prior BERTbase distillation and pruning techniques, such as CoFi, MiniLM, and TinyBERT. Moreover, it matches the performance of the heavily optimized MobileBERT model, obtained via extensive architecture search, by simply pruning the baseline BERTlarge model. When compressing GPT2, ZipLM outperforms DistilGPT2 while being 60% smaller and 30% faster. Our code is available at: https://github.com/IST-DASLab/ZipLM.

Added

2026-09-26