Does a Global Perspective Help Prune Sparse MoEs Elegantly?
Zeliang ZhangNikhil GhoshJiani LiuBin YuXiaodong Liu
Proposes GRAPE, a global redundancy-aware pruning method for sparse Mixture-of-Experts models that dynamically allocates pruning budgets across layers based on cross-layer redundancy to consistently outperform uniform pruning baselines.
Large language models that use sparse Mixture-of-Experts (MoE) architectures achieve high computational efficiency during inference by activating only a small subset of specialized sub-networks, known as experts, for each input. However, housing numerous experts creates massive overall parameter counts, resulting in high memory requirements and infrastructure costs. To reduce this footprint, existing compression methods remove or merge redundant experts, but they almost universally apply uniform pruning targets across all layers, assuming redundancy is evenly distributed throughout the network.
The article aims to evaluate whether a global pruning approach—one that allocates reduction targets across the entire model based on varying layer-by-layer redundancy—can compress sparse MoEs more effectively than standard uniform layer-by-layer methods without degrading performance.
To test this, the authors developed GRAPE (Global Redundancy-Aware Pruning of Experts). The approach measures cross-layer expert similarity and uses an entropy-regularized greedy selection process with a restart mechanism, ensuring that pruning naturally concentrates in highly redundant layers while preventing over-pruning that could destabilize the model. The method was evaluated without task-specific fine-tuning across five major open-source MoE models—Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS—across standardized reasoning, language understanding, and question-answering benchmarks under various compression budgets.
The findings show that expert redundancy is highly uneven across model layers, with later layers generally exhibiting greater redundancy than earlier ones. Across the primary evaluated models, GRAPE consistently achieved the highest average accuracy under identical global pruning budgets, outperforming the strongest uniform local baseline by an average of 1.40% across settings and achieving accuracy gains of up to 2.45% on Mixtral-8x22B. Additionally, the advantage of global redundancy-aware pruning widened under more aggressive pruning targets, maintaining model retention where uniform methods suffered sharp performance drops.
These results demonstrate that treating network compression globally rather than layer-by-layer significantly improves the accuracy-to-memory trade-off in sparse MoE models. Organizations can achieve greater memory savings without paying the typical performance penalty associated with model slimming, directly reducing the hardware and hosting costs required to deploy large-scale models in production environments.
For engineering and product teams managing MoE deployments, adopting global, cross-layer pruning strategies is recommended over uniform layer-wise pruning pipelines. However, practitioners should proceed cautiously: as observed in models like DeepSeek-MoE, extreme redundancy imbalances in specific layers can trigger model collapse if unconstrained. Teams should utilize safeguards like entropy thresholds during compression and conduct further pilot analyses to establish optimal redundancy metrics and layer-protection tolerances before applying aggressive pruning to production workloads.
- Paper: Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models, Xudong Lu et al. (2024). Establishes the foundational layer-by-layer expert pruning and skipping paradigms for Mixture-of-Experts Large Language Models that the source extends via global cross-layer budget allocation.
- Paper: DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models, Damai Dai et al. (2024). Introduces fine-grained expert division and shared expert architectures, serving as a primary target MoE architecture evaluated in the source paper.
- Paper: Scaling Laws for Fine-Grained Mixture of Experts, Jan Ludziejewski et al. (2024). Provides critical theoretical and empirical scaling laws regarding expert granularity and redundancy in sparse Mixture-of-Experts models.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Introduces the foundational sparsely-gated Mixture-of-Experts architecture and gating mechanics essential to understanding modern sparse MoE LLMs.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). Demonstrates accurate one-shot post-training parameter pruning for massive transformer models, establishing the standard post-training compression framework.
- Paper: DepGraph: Towards Any Structural Pruning, Gongfan Fang et al. (2023). Presents dependency graph modeling for structural pruning across interconnected layers, providing vital conceptual grounding for non-uniform cross-layer pruning.
- Paper: DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, Samyam Rajbhandari et al. (2022). Details the inference and training scaling constraints of large-scale MoE models that motivate post-training expert parameter reduction.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Analyzes routing dynamics and expert capacity factors in sparse models, explaining the behavioral nuances of router-allocated experts.
- Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). Investigates whether pruned models retain distinct knowledge advantages over training from scratch across equivalent compute budgets, providing a macro-level evaluation of LLM pruning efficacy.
- Paper: Scaling Embeddings Outperforms Scaling Experts in Language Models, Hong Liu et al. (2026). Examines embedding scaling as an alternative paradigm to expert scaling in sparse language models, extending the inquiry into parameter allocation and architectural efficiency beyond MoE pruning.
