Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
Xudong LuQi LiuYuhui XuAojun ZhouSiyuan HuangBo ZhangJunchi YanHongsheng Li
Proposes post-training expert pruning and dynamic skipping methods for mixture-of-experts large language models, halving memory requirements and boosting inference speed with minimal performance loss.
Mixture-of-Experts large language models achieve state-of-the-art performance with lower computation per token by routing inputs to specialized sub-networks, known as experts. However, their massive static parameter counts create major hardware bottlenecks. For instance, running popular models like Mixtral 8x7B requires at least two high-end enterprise graphics processing units solely to store the full set of expert weights in memory. Existing weight-pruning methods require specialized custom hardware to realize practical speedups, leaving a significant barrier to cost-effective, standard deployment.
The article evaluates hardware-friendly, post-training methods that remove or skip entire experts rather than modifying internal weight matrices. Specifically, it demonstrates how layer-by-layer expert pruning and dynamic on-the-fly expert skipping can substantially reduce memory requirements and accelerate inference speed while preserving core model capabilities across general and specialized tasks.
The researchers developed two complementary techniques: permanent expert pruning and dynamic expert skipping. To prune experts without retraining, the method passes a small calibration dataset through the model and systematically selects the subset of experts in each layer that minimizes token reconstruction loss. For dynamic skipping, the model evaluates routing confidence during inference and skips secondary experts when their relative weights fall below a calibrated threshold. The authors evaluated these techniques on both the base and instruction-tuned versions of Mixtral 8x7B across standard general reasoning benchmarks and mathematical problem-solving tasks.
The evaluation produced four key findings. First, permanently pruning two out of eight experts per layer reduced memory usage from approximately 90 gigabytes to 68 gigabytes (a 24% parameter reduction), cutting the required hardware from two high-end graphics processing units down to a single device while improving inference speed by 1.20× with only a 2.9-point average performance drop. Second, pruning four experts cut memory by nearly 48% (down to 47 gigabytes) and achieved a 1.27× speedup, outperforming conventional weight-pruning baselines in both task accuracy and latency. Third, combining two-expert pruning with dynamic skipping matched the 1.27×–1.33× speed of four-expert pruning while preserving significantly higher benchmark performance. Fourth, for specialized domains like mathematics, calibrating on domain-specific data and applying lightweight task-specific fine-tuning largely eliminated performance losses, with a pruned seven-expert model slightly exceeding the full eight-expert baseline on math benchmarks.
These findings indicate that significant redundant capacity exists within multi-expert language models. Deploying these methods translates directly to operational cost reductions by halving the infrastructure footprint needed for model serving, lowering inter-chip communication overhead, and increasing throughput without requiring specialized hardware. The results also show that aligning calibration data to the target deployment domain is critical, as pre-training data alone leads to suboptimal expert selection for domain-specific applications.
For engineering teams seeking to optimize existing Mixture-of-Experts deployments, the article supports adopting two-expert pruning combined with dynamic skipping to enable single-device serving with minimal accuracy loss. When deploying for specialized domains, teams should calibrate expert pruning against domain-relevant data and run targeted fine-tuning. However, decision-makers should note that the current combinatorial search approach is best suited for models with a moderate number of experts (such as four to eight per layer) and may require algorithmic adaptation for architectures with substantially larger expert counts.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). It introduces the foundational sparsely-gated Mixture-of-Experts layer and gating mechanisms that the source model architecture relies upon and seeks to prune.
- Paper: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, William Fedus et al. (2022). It establishes key principles for routing tokens efficiently across sparse Transformer expert layers, which form the direct architectural baseline for the source paper's expert pruning and skipping.
- Paper: DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, Samyam Rajbhandari et al. (2022). It highlights the deployment and inference bottlenecks of large MoE architectures and provides critical background on MoE model compression.
- Paper: SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Elias Frantar et al. (2023). It presents one-shot post-training pruning using calibration data for large language models, providing the core methodological framing that the source adapts to coarse expert-level pruning.
- Paper: A Fast Post-Training Pruning Framework for Transformers, Woosuk Kwon et al. (2022). It outlines post-training structured pruning techniques for Transformers without retraining, offering key foundational concepts for structured layer-wise pruning.
- Paper: MoEfication: Transformer Feed-forward Layers are Mixtures of Experts, Zhengyan Zhang et al. (2022). It demonstrates how Transformer feed-forward networks function as modular sparse experts, underpinning the intuition for whole-expert pruning and skipping.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). It introduces dynamic, input-dependent skipping and sparsity during LLM inference, directly motivating dynamic routing thresholding and expert skipping.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). It details routing stability, fine-tuning behavior, and transferability dynamics in sparse expert models that inform the source's downstream task calibration.
- Paper: Scaling Laws for Fine-Grained Mixture of Experts, Jan Ludziejewski et al. (2024). It extends sparse expert architectural design by establishing scaling laws for fine-grained expert division, addressing the granularity limitations of expert pruning discussed in the source.
- Paper: DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models, Damai Dai et al. (2024). It designs a fine-grained expert architecture with shared and routed experts to minimize capacity redundancy directly during pretraining rather than via post-training pruning.
- Paper: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, Zhihong Shao et al. (2024). It applies fine-grained expert segmentation and device-limited routing at scale to optimize MoE inference efficiency and serving memory footprints.
- Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). It builds on LLM pruning methodologies by rigorously analyzing whether pruned models outperform dense models trained from scratch under token-matched retraining budgets.
- Paper: Scaling Embeddings Outperforms Scaling Experts in Language Models, Hong Liu et al. (2026). It investigates alternatives to expert over-parameterization by evaluating embedding scaling in high-sparsity language models.
