ZipLM is an inference-aware structured pruning framework designed to compress large language models while achieving specific target runtime speedups in a designated hardware environment. The approach works by iteratively identifying and removing model components that have the least favorable trade-off between task loss and inference latency. By directly incorporating deployment hardware specifications into the pruning process, ZipLM can generate optimized variants of both encoder and decoder architectures that meet strict speed and accuracy requirements at a lower computational cost than traditional architecture search or extensive distillation methods.