Inference-aware structured pruning is a neural network compression technique that removes entire architectural components, such as attention heads, hidden dimensions, channels, or full layers, while explicitly optimizing for measured execution speed and latency in a target deployment environment. Unlike traditional structured pruning methods that rely on parameter counts or theoretical floating-point operations as proxies for efficiency, this approach directly evaluates the trade-off between task accuracy and real-world runtime on specific hardware. By iteratively identifying and eliminating the structural units that contribute least effectively to performance relative to their execution cost, it generates smaller, dense models guaranteed to achieve practical speedups and meet concrete inference latency requirements.