Built independently by an author, for readers. Read the story and support ChapterPal

keyword

global pruning

Global pruning is a neural network compression technique in which model components, such as individual weights, channels, or entire sub-modules, are evaluated and removed based on a single unified importance metric applied across the entire architecture rather than within each layer independently. In contrast to local or layer-wise pruning, which enforces a fixed reduction ratio separately for every individual layer, global pruning compares the relative significance of components across all layers simultaneously. This perspective allows the algorithm to dynamically allocate varying degrees of sparsity throughout the network, automatically removing a greater proportion of redundant elements from less critical layers while preserving essential structures in more sensitive layers to maximize overall model efficiency and accuracy.

3 items

Does a Global Perspective Help Prune Sparse MoEs Elegantly?

Does a Global Perspective Help Prune Sparse MoEs Elegantly?

Zeliang Zhang, Nikhil Ghosh, Jiani Liu, Bin Yu, Xiaodong Liu

OrganizationsFlatiron InstituteMicrosoftUniversity of California BerkeleyUniversity of Rochester

Why you should read this

Proposes GRAPE, a global redundancy-aware pruning method for sparse Mixture-of-Experts models that dynamically allocates pruning budgets across layers based on cross-layer redundancy to consistently outperform uniform pruning baselines.

Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial memory consumption. Existing pruning methods typically allocate budgets uniformly across layers, overlooking the heterogeneous redundancy that arises in sparse MoEs. We propose GRAPE (Global Redundancy-Aware Pruning of Experts, a global pruning strategy that dynamically allocates pruning budgets based on cross-layer redundancy. Experiments on Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS show that, under the same pruning budget, GRAPE consistently achieves the best average performance. On the three main models reported in the paper, it improves average accuracy over the strongest local baseline by 1.40% on average across pruning settings, with gains of up to 2.45%.

Added

2026-09-29

Learning Efficient Convolutional Networks through Network Slimming

Learning Efficient Convolutional Networks through Network Slimming

Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, Changshui Zhang

OrganizationsCornell UniversityFudan UniversityIntelTsinghua University

Why you should read this

Proposes a hardware-efficient structured pruning method that enforces channel-level sparsity using batch normalization scaling factors, allowing for immediate inference speedups.

The deployment of deep convolutional neural networks (CNNs) in many real world applications is largely hindered by their high computational cost. In this paper, we propose a novel learning scheme for CNNs to simultaneously 1) reduce the model size; 2) decrease the run-time memory footprint; and 3) lower the number of computing operations, without compromising accuracy. This is achieved by enforcing channel-level sparsity in the network in a simple but effective way. Different from many existing approaches, the proposed method directly applies to modern CNN architectures, introduces minimum overhead to the training process, and requires no special software/hardware accelerators for the resulting models. We call our approach network slimming, which takes wide and large networks as input models, but during training insignificant channels are automatically identified and pruned afterwards, yielding thin and compact models with comparable accuracy. We empirically demonstrate the effectiveness of our approach with several state-of-the-art CNN models, including VGGNet, ResNet and DenseNet, on various image classification datasets. For VGGNet, a multi-pass version of network slimming gives a 20× reduction in model size and a 5× reduction in computing operations.

Added

2026-02-26