Does a Global Perspective Help Prune Sparse MoEs Elegantly?

Zeliang ZhangNikhil GhoshJiani LiuBin YuXiaodong Liu

article2026arXiv2 citations

Proposes GRAPE, a global redundancy-aware pruning method for sparse Mixture-of-Experts models that dynamically allocates pruning budgets across layers based on cross-layer redundancy to consistently outperform uniform pruning baselines.

Listen

Large language models that use sparse Mixture-of-Experts (MoE) architectures achieve high computational efficiency during inference by activating only a small subset of specialized sub-networks, known as experts, for each input. However, housing numerous experts creates massive overall parameter counts, resulting in high memory requirements and infrastructure costs. To reduce this footprint, existing compression methods remove or merge redundant experts, but they almost universally apply uniform pruning targets across all layers, assuming redundancy is evenly distributed throughout the network.

The article aims to evaluate whether a global pruning approach—one that allocates reduction targets across the entire model based on varying layer-by-layer redundancy—can compress sparse MoEs more effectively than standard uniform layer-by-layer methods without degrading performance.

To test this, the authors developed GRAPE (Global Redundancy-Aware Pruning of Experts). The approach measures cross-layer expert similarity and uses an entropy-regularized greedy selection process with a restart mechanism, ensuring that pruning naturally concentrates in highly redundant layers while preventing over-pruning that could destabilize the model. The method was evaluated without task-specific fine-tuning across five major open-source MoE models—Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS—across standardized reasoning, language understanding, and question-answering benchmarks under various compression budgets.

The findings show that expert redundancy is highly uneven across model layers, with later layers generally exhibiting greater redundancy than earlier ones. Across the primary evaluated models, GRAPE consistently achieved the highest average accuracy under identical global pruning budgets, outperforming the strongest uniform local baseline by an average of 1.40% across settings and achieving accuracy gains of up to 2.45% on Mixtral-8x22B. Additionally, the advantage of global redundancy-aware pruning widened under more aggressive pruning targets, maintaining model retention where uniform methods suffered sharp performance drops.

These results demonstrate that treating network compression globally rather than layer-by-layer significantly improves the accuracy-to-memory trade-off in sparse MoE models. Organizations can achieve greater memory savings without paying the typical performance penalty associated with model slimming, directly reducing the hardware and hosting costs required to deploy large-scale models in production environments.

For engineering and product teams managing MoE deployments, adopting global, cross-layer pruning strategies is recommended over uniform layer-wise pruning pipelines. However, practitioners should proceed cautiously: as observed in models like DeepSeek-MoE, extreme redundancy imbalances in specific layers can trigger model collapse if unconstrained. Teams should utilize safeguards like entropy thresholds during compression and conduct further pilot analyses to establish optimal redundancy metrics and layer-protection tolerances before applying aggressive pruning to production workloads.

arXiv: 2604.06542
  • Paper: Small LLMs: Pruning vs. Training from Scratch, Yufeng Xu et al. (2026). Investigates whether pruned models retain distinct knowledge advantages over training from scratch across equivalent compute budgets, providing a macro-level evaluation of LLM pruning efficacy.
  • Paper: Scaling Embeddings Outperforms Scaling Experts in Language Models, Hong Liu et al. (2026). Examines embedding scaling as an alternative paradigm to expert scaling in sparse language models, extending the inquiry into parameter allocation and architectural efficiency beyond MoE pruning.
Cover for Does a Global Perspective Help Prune Sparse MoEs Elegantly?

Abstract

Empirical scaling laws for language models have encouraged the development of ever-larger LLMs, despite their growing computational and memory costs. Sparse Mixture-of-Experts (MoEs) offer a promising alternative by activating only a subset of experts per forward pass, improving efficiency without sacrificing performance. However, the large number of expert parameters still leads to substantial memory consumption.

Existing pruning methods typically allocate budgets uniformly across layers, overlooking the heterogeneous redundancy that arises in sparse MoEs. We propose GRAPE (Global Redundancy-Aware Pruning of Experts, a global pruning strategy that dynamically allocates pruning budgets based on cross-layer redundancy. Experiments on Mixtral-8x7B, Mixtral-8x22B, DeepSeek-MoE, Qwen-MoE, and GPT-OSS show that, under the same pruning budget, GRAPE consistently achieves the best average performance. On the three main models reported in the paper, it improves average accuracy over the strongest local baseline by 1.40% on average across pruning settings, with gains of up to 2.45%.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Methodology
  • 3.1 Preliminary
  • 3.2 Not all MoE layers are equally redundant
  • 3.3 Globally Pruning the MoEs
  • 4 Evaluations
  • 4.1 Experiment setup
  • 4.2 Experiment Results
  • 5 Conclusion
  • References
  • A More Results on MoE Pruning

Knowls

  1. Knowl 1 — GRAPE: One-Shot Entropy-Aware Global MoE Pruning Algorithm with Restart

    algorithm

    GRAPE (Global Redundancy-Aware Pruning of Experts) performs greedy cross-layer expert merging under a global parameter budget while using layerwise entropy thresholds and a restart mechanism to prevent over-concentrating pruning in individual layers.

    Input: Pairwise expert similarity matrices {Dl}l=1L\{D^l\}_{l=1}^L with Dl∈RN×ND^l \in \mathbb{R}^{N \times N}, total target retained experts K<LNK < LN, entropy tolerance parameter γ∈[0,1]\gamma \in [0, 1]
    Output: Expert cluster partition {Cl}l=1L\{C^l\}_{l=1}^L across layers such that ∑l=1L∣Cl∣=K\sum_{l=1}^L |C^l| = K
    Initialize Cl←{{0},{1},…,{N−1}}C^l \leftarrow \{\{0\}, \{1\}, \dots, \{N-1\}\} for all l∈{1,…,L}l \in \{1, \dots, L\}
    Initialize layer residual redundancy Rl←∑i≠jDijlR^l \leftarrow \sum_{i \neq j} D^l_{ij} for all l∈{1,…,L}l \in \{1, \dots, L\}
    Compute initial layer fractions pl←∣Cl∣/∑l′∣Cl′∣p_l \leftarrow |C^l| / \sum_{l'} |C^{l'}|
    Compute initial global entropy E←−∑l=1Lpllog⁡plE \leftarrow -\sum_{l=1}^L p_l \log p_l
    Set entropy threshold E^←E(1−γ)\hat{E} \leftarrow E(1 - \gamma)
    Initialize frozen layer set F←∅F \leftarrow \emptyset
    while ∑l=1L∣Cl∣>K\sum_{l=1}^L |C^l| > K do
        if F={1,…,L}F = \{1, \dots, L\} then
            F←∅F \leftarrow \emptyset # Restart: reset frozen set if all layers are frozen
        end if
        Select layer with maximum residual redundancy: l⋆←arg⁡max⁡l∉FRll^\star \leftarrow \arg\max_{l \notin F} R^l
        Select most similar pair of expert clusters: (i⋆,j⋆)←arg⁡max⁡i≠jDijl⋆(i^\star, j^\star) \leftarrow \arg\max_{i \neq j} D^{l^\star}_{ij}
        Cl⋆←Union(Cl⋆,i⋆,j⋆)C^{l^\star} \leftarrow \text{Union}(C^{l^\star}, i^\star, j^\star)
        Rl⋆←Rl⋆−2Di⋆,j⋆l⋆R^{l^\star} \leftarrow R^{l^\star} - 2 D^{l^\star}_{i^\star, j^\star}
        Di⋆,j⋆l⋆←0D^{l^\star}_{i^\star, j^\star} \leftarrow 0, Dj⋆,i⋆l⋆←0D^{l^\star}_{j^\star, i^\star} \leftarrow 0
        Update layer fractions pl←∣Cl∣/∑l′=1L∣Cl′∣p_l \leftarrow |C^l| / \sum_{l'=1}^L |C^{l'}|
        E←−∑l=1Lpllog⁡plE \leftarrow -\sum_{l=1}^L p_l \log p_l
        if E<E^E < \hat{E} then
            F←F∪{l⋆}F \leftarrow F \cup \{l^\star\} # Freeze layer to avoid excessive layer imbalance
        end if
    end while
    return {Cl}l=1L\{C^l\}_{l=1}^L
  2. Knowl 2 — Global Entropy Safeguarding for Cross-Layer MoE Pruning

    model/method

    In a sparse Mixture-of-Experts (MoE) architecture with LL layers and NN experts per layer, global pruning can be formulated by constructing a block-diagonal pairwise similarity matrix A=blockdiag(D1,D2,…,DL)∈RLN×LNA = \text{blockdiag}(D^1, D^2, \dots, D^L) \in \mathbb{R}^{LN \times LN}, where Dl∈RN×ND^l \in \mathbb{R}^{N \times N} encodes the intra-layer similarity between experts in layer ll. The unconstrained global pruning objective finds a binary mask M=blockdiag(M1,…,ML)∈{0,1}LN×LNM = \text{blockdiag}(M^1, \dots, M^L) \in \{0, 1\}^{LN \times LN} satisfying:

    arg⁡min⁡M∑i≠jAij⋅Mij\arg\min_M \sum_{i \neq j} A_{ij} \cdot M_{ij}

    Directly minimizing this objective can cause degenerate pruning: when certain layers have disproportionately high redundancy, a purely similarity-driven budget allocation removes too many experts from those few layers, leading to layer collapse.

    To prevent this, GRAPE defines the global entropy EE over the layer-wise fraction of retained experts plp_l:

    pl=∣Cl∣∑l′=1L∣Cl′∣,E=−∑l=1Lpllog⁡plp_l = \frac{|C^l|}{\sum_{l'=1}^L |C^{l'}|}, \quad E = -\sum_{l=1}^L p_l \log p_l

    where ∣Cl∣|C^l| is the number of remaining expert clusters in layer ll, and ∑l=1Lpl=1\sum_{l=1}^L p_l = 1. A higher entropy corresponds to an even distribution of retained experts across layers, while a lower entropy signifies severe pruning concentration in specific layers. GRAPE imposes an entropy floor E^=Einitial(1−γ)\hat{E} = E_{\text{initial}}(1 - \gamma), parameterized by tolerance γ∈[0,1]\gamma \in [0, 1]. When pruning in a layer causes E<E^E < \hat{E}, that layer is frozen from further expert removals until the target budget is satisfied.

  3. Knowl 3 — Cross-Layer MoE Redundancy and Normalized Relative Redundancy Score

    definition

    For an MoE model with LL sparse layers and NN experts per layer, let Dl∈RN×ND^l \in \mathbb{R}^{N \times N} denote the pairwise similarity matrix of experts in the ll-th MoE layer, computed via metrics such as Centered Kernel Alignment (CKA), mean squared error, or expert output representations.

    The average intra-layer redundancy score RlR^l for layer ll is defined as:

    Rl=1N(N−1)∑i≠jDijlR^l = \frac{1}{N(N - 1)} \sum_{i \neq j} D^l_{ij}

    The normalized relative redundancy score across layers R~l∈[0,1]\tilde{R}^l \in [0, 1] is defined as:

    R~l=Rl−min⁡l′Rl′max⁡l′Rl′−min⁡l′Rl′\tilde{R}^l = \frac{R^l - \min_{l'} R^{l'}}{\max_{l'} R^{l'} - \min_{l'} R^{l'}}

    Empirical evaluation of R~l\tilde{R}^l reveals that redundancy is heterogeneous across layers: earlier layers typically exhibit lower redundancy than later layers, but the trend across intermediate layers is non-monotonic. Consequently, uniform layer-wise pruning allocations fail to reflect the underlying distribution of model redundancy.

  4. Knowl 4 — Benchmark Results on Pruning Mixtral-8x22B and DeepSeek-MoE

    data/table

    The table below compares GRAPE against uniform layer-wise pruning baselines on Mixtral-8x22B and DeepSeek-MoE under average pruning budgets of 2 experts per layer (2e) and 4 experts per layer (4e) across MMLU (Humanities, Social Science, STEM, Other), BoolQ, OpenBookQA, and RTE in zero-shot / task-agnostic settings (no fine-tuning).

    Model Method MMLU BoolQ OpenBookQA RTE Average
    Hum. Soc. STEM Other
    Mixtral-8x22B Original 68.6 84.1 67.1 78.7 87.9 35.8 71.2 70.4
    Router-guided (2e/4e) 27.3/22.7 25.4/25.8 24.4/24.0 27.9/23.4 62.8/62.7 12.8/13.0 54.2/49.5 33.5/31.6
    Count-guided (2e/4e) 58.0/45.7 74.9/57.7 54.1/42.0 70.2/45.7 81.5/74.4 35.2/27.0 69.3/57.4 63.3/50.0
    Enumerate (2e/4e) 60.4/53.9 78.0/67.2 59.5/52.3 73.0/64.2 87.4/80.5 35.0/31.1 70.1/67.9 66.2/59.6
    DEK (2e/4e) 62.3/57.8 78.5/69.7 60.2/51.3 73.4/64.2 87.6/83.1 35.8/33.2 71.1/68.1 67.0/61.1
    GRAPE (Ours, 2e/4e) 64.1/58.4 80.4/72.9 62.7/54.6 75.3/67.9 88.0/84.1 35.2/32.0 71.4/68.5 68.2/62.6
    DeepSeek-MoE Original 40.4 47.9 36.1 49.5 77.2 32.8 66.0 50.0
    Router-guided (2e/4e) 38.1/34.3 46.4/41.6 33.9/33.6 47.7/43.1 71.3/72.1 33.2/31.4 60.4/60.2 47.3/45.2
    Count-guided (2e/4e) 38.3/35.9 47.4/42.5 34.1/32.9 47.4/45.6 76.2/75.9 33.8/32.4 64.9/67.5 48.9/47.5
    DEK (2e/4e) 38.7/39.2 47.3/47.0 34.9/33.7 46.9/46.6 77.4/76.6 32.2/32.2 64.6/66.4 48.8/48.8
    GRAPE (Ours, 2e/4e) 39.7/39.7 48.3/47.6 35.8/35.3 50.0/50.0 77.8/77.5 32.0/31.2 65.0/65.5 49.8/49.5

    On Mixtral-8x22B, GRAPE achieves average accuracies of 68.2% (2e) and 62.6% (4e), outperforming DEK (the strongest local baseline) by 1.20% and 1.50% absolute (1.79% and 2.45% relative). On DeepSeek-MoE, GRAPE achieves 49.8% (2e) and 49.5% (4e), outperforming the strongest local baselines by 1.84% and 1.43% relative.

  5. Knowl 5 — Benchmark Results on Pruning GPT-OSS

    data/table

    The table below compares GRAPE against uniform per-layer pruning baselines on GPT-OSS (gpt-oss-20b) with medium reasoning effort under global budgets equivalent to pruning 2 and 4 experts per layer (2e/4e).

    Model Method MMLU BoolQ OpenBookQA RTE Average
    GPT-OSS Original 85.6 88.7 93.0 92.8 90.0
    Router-guided (2e/4e) 83.4 / 81.3 88.9 / 87.3 89.6 / 85.6 90.6 / 89.5 88.1 / 85.9
    Count-guided (2e/4e) 84.8 / 82.5 88.2 / 89.3 93.6 / 93.2 92.4 / 91.3 89.8 / 89.1
    Enumerate (2e/4e) 83.7 / 82.5 88.8 / 88.7 93.4 / 92.8 93.5 / 92.1 89.9 / 89.0
    DEK (2e/4e) 83.6 / 80.8 89.0 / 88.1 94.6 / 92.4 92.4 / 90.6 89.9 / 88.0
    GRAPE (Ours, 2e/4e) 85.3 / 83.4 89.0 / 89.0 94.2 / 93.6 92.6 / 91.9 90.3 / 89.5

    Under a 2-expert per layer global budget, GRAPE achieves an average accuracy of 90.3% (surpassing the unpruned model's 90.0% and beating the strongest local baseline at 89.9%). Under a 4-expert per layer budget, GRAPE maintains 89.5% average accuracy, exceeding the best uniform baseline (Count-guided at 89.1%) by 0.45% relative.

  6. Knowl 6 — Experimental Setup for Global vs Local MoE Pruning

    experimental setup

    Experiments evaluate MoE pruning across multiple architectures under a strictly task-agnostic protocol where no post-pruning fine-tuning is performed:

    1. Evaluated Models:

      • Mixtral-8x22B: 56 MoE layers, 8 experts per layer, Top-2 routing (2 experts activated per token).
      • DeepSeek-MoE-16B: 27 MoE layers, each comprising 64 private experts (top-6 activated) and 2 shared experts (both always activated).
      • GPT-OSS (gpt-oss-20b): 24 MoE layers, 32 experts per layer, Top-4 routing.
      • Mixtral-8x7B and Qwen-MoE (evaluated in appendix).
    2. Evaluation Benchmarks:

      • MMLU (covering Humanities, Social Science, STEM, and Other subfields; full dataset for Mixtral and DeepSeek-MoE, randomly sampled 1,000 examples for GPT-OSS due to compute costs).
      • BoolQ, OpenBookQA, and RTE (evaluated on full test sets).
    3. Evaluation Protocol:

      • Mixtral and DeepSeek-MoE are prompted to directly emit the final answer.
      • GPT-OSS uses its default chain-of-thought generation style with reasoning effort set to medium.
    4. Baseline Methods: Four layer-wise uniform pruning baselines are compared under identical total expert reduction budgets: Router-guided (using routing probabilities/weights), Count-guided (expert visitation frequency), Enumerate (calibration loss landscape search), and DEK (expert output representation similarity).

  7. Knowl 7 — Performance Retention on Mixtral-8x7B and Qwen-MoE

    empirical result

    When measuring retained performance (defined as the ratio of pruned model accuracy to unpruned model accuracy Accuracypruned/Accuracyoriginal\text{Accuracy}_{\text{pruned}} / \text{Accuracy}_{\text{original}}) on Mixtral-8x7B and Qwen-MoE:

    • On Mixtral-8x7B, global greedy pruning exhibits a pronounced advantage over uniform layer-wise baselines, especially under aggressive 4-expert pruning where uniform per-layer methods experience large accuracy drops (particularly on MMLU).
    • On Qwen-MoE, where expert removal is generally less destructive overall, global greedy pruning maintains the highest or near-highest retention ratio across individual tasks and achieves the best overall average retention.

    This demonstrates that dynamic cross-layer budget allocation preserves performance more effectively than uniform pruning even on smaller or less redundant MoE architectures.

  8. Knowl 8 — Vulnerability to Layer Imbalance and Collapse in Global MoE Pruning

    limitation

    In certain MoE architectures, redundancy is disproportionately concentrated in specific layers. For instance, the final layer of DeepSeek-MoE has an average intra-layer redundancy score substantially higher than all preceding layers.

    Without explicit regularization (such as GRAPE's entropy threshold and restart mechanism), a pure global greedy similarity objective allocates nearly the entire global pruning quota to these few high-redundancy layers. Removing too many experts from a single layer induces catastrophic layer imbalance and can cause the entire model to collapse. Designing universally robust metrics to measure layer-level redundancy across heterogeneous MoE architectures remains an open problem.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. 2025. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925.
  2. 2.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45.
  3. 3.Tianlong Chen, Zhenyu Zhang, Ajay Jaiswal, Shiwei Liu, and Zhangyang Wang. 2023. Sparse moe as the new dropout: Scaling dense and self-slimmable transformers. arXiv preprint arXiv:2303.01610.
  4. 4.Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. Task-specific expert pruning for sparse mixture-of-experts. arXiv preprint arXiv:2206.00277.
  5. 5.Mohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang, Pin-Yu Chen, and Christopher Carothers. 2024. A provably effective method for pruning experts in fine-tuned sparse mixture-of-experts. arXiv preprint arXiv:2405.16646.
  6. 6.Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al. 2024. Deepseek-moe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066.
  7. 7.MohammadReza Davari, Stefan Horoi, Amine Natik, Guillaume Lajoie, Guy Wolf, and Eugene Belilovsky. Reliability of cka as a similarity measure in deep learning. In The Eleventh International Conference on Learning Representations.
  8. 8.Shwai He, Daize Dong, Liang Ding, and Ang Li. 2024. Demystifying the compression of mixture-of-experts through a unified framework. arXiv preprint arXiv:2406.02500.
  9. 9.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  10. 10.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  11. 11.Jaeseong Lee, Aurick Qiao, Daniel F Campos, Zhewei Yao, Yuxiong He, et al. 2024. Stun: Structured-then-unstructured pruning for scalable moe pruning. arXiv preprint arXiv:2409.06211.
  12. 12.Pingzhi Li, Zhenyu Zhang, Prateek Yadav, Yi-Lin Sung, Yu Cheng, Mohit Bansal, and Tianlong Chen. 2024a. Merge, then compress: Demystify efficient smoe with hints from its routing policy. In The Twelfth International Conference on Learning Representations.
  13. 13.Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al. 2024b. Personal llm agents: Insights and survey about the capability, efficiency and security. arXiv preprint arXiv:2401.05459.
  14. 14.Enshu Liu, Junyi Zhu, Zinan Lin, Xuefei Ning, Matthew B Blaschko, Shengen Yan, Guohao Dai, Huazhong Yang, and Yu Wang. 2024. Efficient expert pruning for sparse mixture-of-experts language models: Enhancing performance and reducing inference costs. arXiv preprint arXiv:2407.00945.
  15. 15.Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, and Hongsheng Li. 2024. Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6159–6172.
  16. 16.Bowen Pan, Yikang Shen, Haokun Liu, Mayank Mishra, Gaoyuan Zhang, Aude Oliva, Colin Raffel, and Rameswar Panda. 2024. Dense training, sparse inference: Rethinking training of mixture-of-experts language models. arXiv preprint arXiv:2404.05567.
  17. 17.An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024. Qwen2. 5 technical report. arXiv preprint arXiv:2412.15115.
  18. 18.Zeliang Zhang, Xiaodong Liu, Hao Cheng, Chenliang Xu, and Jianfeng Gao. 2025. Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts. In In Findings of the Association for Computational Linguistics: ACL 2025.
  19. 19.Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022. St-moe: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906.

Citation

MLA
Zhang, Z., et al. “Does a Global Perspective Help Prune Sparse MoEs Elegantly?”. arXiv, 2026, http://arxiv.org/abs/2604.06542v1.
APA
Zhang, Z., Ghosh, N., Liu, J., Yu, B., & Liu, X. (2026). Does a Global Perspective Help Prune Sparse MoEs Elegantly?. arXiv. http://arxiv.org/abs/2604.06542v1
Chicago
Zhang, Z., N. Ghosh, J. Liu, B. Yu, and X. Liu. 2026. “Does a Global Perspective Help Prune Sparse MoEs Elegantly?”. arXiv. http://arxiv.org/abs/2604.06542v1.
Harvard
Zhang, Z. et al. (2026) “Does a Global Perspective Help Prune Sparse MoEs Elegantly?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.06542v1.
Vancouver
1. Zhang Z, Ghosh N, Liu J, Yu B, Liu X (2026) Does a Global Perspective Help Prune Sparse MoEs Elegantly?. arXiv

BibTeX

@article{zhang2026does,
  title = {Does a Global Perspective Help Prune Sparse MoEs Elegantly?},
  author = {Zhang, Zeliang and Ghosh, Nikhil and Liu, Jiani and Yu, Bin and Liu, Xiaodong},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.06542v1},
  eprint = {2604.06542}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission