keyword
global pruning
Global pruning is a neural network compression technique in which model components, such as individual weights, channels, or entire sub-modules, are evaluated and removed based on a single unified importance metric applied across the entire architecture rather than within each layer independently. In contrast to local or layer-wise pruning, which enforces a fixed reduction ratio separately for every individual layer, global pruning compares the relative significance of components across all layers simultaneously. This perspective allows the algorithm to dynamically allocate varying degrees of sparsity throughout the network, automatically removing a greater proportion of redundant elements from less critical layers while preserving essential structures in more sensitive layers to maximize overall model efficiency and accuracy.
3 items

Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model
Yeskendir Koishekenov, Alexandre Berard, Vassilina Nikoulina
Why you should read this
Proposes an inference-time pruning method that removes up to 80% of experts from the 54.5-billion-parameter NLLB-200 translation model without fine-tuning, reducing hardware requirements from multiple devices down to a single 32GB GPU while maintaining translation quality.
The recently released NLLB-200 is a set of multilingual Neural Machine Translation models that cover 202 languages. The largest model is based on a Mixture of Experts architecture and achieves SoTA results across many language pairs. It contains 54.5B parameters and requires at least four 32GB GPUs just for inference. In this work, we propose a pruning method that enables the removal of up to 80% of experts without further finetuning and with a negligible loss in translation quality, which makes it feasible to run the model on a single 32GB GPU. Further analysis suggests that our pruning metrics can identify language-specific experts.
Added
2026-10-03

Does a Global Perspective Help Prune Sparse MoEs Elegantly?
Zeliang Zhang, Nikhil Ghosh, Jiani Liu, Bin Yu, Xiaodong Liu
Why you should read this
Proposes GRAPE, a global redundancy-aware pruning method for sparse Mixture-of-Experts models that dynamically allocates pruning budgets across layers based on cross-layer redundancy to consistently outperform uniform pruning baselines.
Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial memory consumption. Existing pruning methods typically allocate budgets uniformly across layers, overlooking the heterogeneous redundancy that arises in sparse MoEs. We propose GRAPE (Global Redundancy-Aware Pruning of Experts, a global pruning strategy that dynamically allocates pruning budgets based on cross-layer redundancy. Experiments on Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS show that, under the same pruning budget, GRAPE consistently achieves the best average performance. On the three main models reported in the paper, it improves average accuracy over the strongest local baseline by 1.40% on average across pruning settings, with gains of up to 2.45%.
Added
2026-09-29

Learning Efficient Convolutional Networks through Network Slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, Changshui Zhang
Why you should read this
Proposes a hardware-efficient structured pruning method that enforces channel-level sparsity using batch normalization scaling factors, allowing for immediate inference speedups.
The deployment of deep convolutional neural networks (CNNs) in many real world applications is largely hindered by their high computational cost. In this paper, we propose a novel learning scheme for CNNs to simultaneously 1) reduce the model size; 2) decrease the run-time memory footprint; and 3) lower the number of computing operations, without compromising accuracy. This is achieved by enforcing channel-level sparsity in the network in a simple but effective way. Different from many existing approaches, the proposed method directly applies to modern CNN architectures, introduces minimum overhead to the training process, and requires no special software/hardware accelerators for the resulting models. We call our approach network slimming, which takes wide and large networks as input models, but during training insignificant channels are automatically identified and pruned afterwards, yielding thin and compact models with comparable accuracy. We empirically demonstrate the effectiveness of our approach with several state-of-the-art CNN models, including VGGNet, ResNet and DenseNet, on various image classification datasets. For VGGNet, a multi-pass version of network slimming gives a 20× reduction in model size and a 5× reduction in computing operations.
Added
2026-02-26
