Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Xudong LuQi LiuYuhui XuAojun ZhouSiyuan HuangBo ZhangJunchi YanHongsheng Li

article2024ACL93 citations

Proposes post-training expert pruning and dynamic skipping methods for mixture-of-experts large language models, halving memory requirements and boosting inference speed with minimal performance loss.

Listen

Mixture-of-Experts large language models achieve state-of-the-art performance with lower computation per token by routing inputs to specialized sub-networks, known as experts. However, their massive static parameter counts create major hardware bottlenecks. For instance, running popular models like Mixtral 8x7B requires at least two high-end enterprise graphics processing units solely to store the full set of expert weights in memory. Existing weight-pruning methods require specialized custom hardware to realize practical speedups, leaving a significant barrier to cost-effective, standard deployment.

The article evaluates hardware-friendly, post-training methods that remove or skip entire experts rather than modifying internal weight matrices. Specifically, it demonstrates how layer-by-layer expert pruning and dynamic on-the-fly expert skipping can substantially reduce memory requirements and accelerate inference speed while preserving core model capabilities across general and specialized tasks.

The researchers developed two complementary techniques: permanent expert pruning and dynamic expert skipping. To prune experts without retraining, the method passes a small calibration dataset through the model and systematically selects the subset of experts in each layer that minimizes token reconstruction loss. For dynamic skipping, the model evaluates routing confidence during inference and skips secondary experts when their relative weights fall below a calibrated threshold. The authors evaluated these techniques on both the base and instruction-tuned versions of Mixtral 8x7B across standard general reasoning benchmarks and mathematical problem-solving tasks.

The evaluation produced four key findings. First, permanently pruning two out of eight experts per layer reduced memory usage from approximately 90 gigabytes to 68 gigabytes (a 24% parameter reduction), cutting the required hardware from two high-end graphics processing units down to a single device while improving inference speed by 1.20× with only a 2.9-point average performance drop. Second, pruning four experts cut memory by nearly 48% (down to 47 gigabytes) and achieved a 1.27× speedup, outperforming conventional weight-pruning baselines in both task accuracy and latency. Third, combining two-expert pruning with dynamic skipping matched the 1.27×–1.33× speed of four-expert pruning while preserving significantly higher benchmark performance. Fourth, for specialized domains like mathematics, calibrating on domain-specific data and applying lightweight task-specific fine-tuning largely eliminated performance losses, with a pruned seven-expert model slightly exceeding the full eight-expert baseline on math benchmarks.

These findings indicate that significant redundant capacity exists within multi-expert language models. Deploying these methods translates directly to operational cost reductions by halving the infrastructure footprint needed for model serving, lowering inter-chip communication overhead, and increasing throughput without requiring specialized hardware. The results also show that aligning calibration data to the target deployment domain is critical, as pre-training data alone leads to suboptimal expert selection for domain-specific applications.

For engineering teams seeking to optimize existing Mixture-of-Experts deployments, the article supports adopting two-expert pruning combined with dynamic skipping to enable single-device serving with minimal accuracy loss. When deploying for specialized domains, teams should calibrate expert pruning against domain-relevant data and run targeted fine-tuning. However, decision-makers should note that the current combinatorial search approach is best suited for models with a moderate number of experts (such as four to eight per layer) and may require algorithmic adaptation for architectures with substantially larger expert counts.

arXiv: 2402.14800
Cover for Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Abstract

A pivotal advancement in the progress of large language models (LLMs) is the emergence of the Mixture-of-Experts (MoE) LLMs. Compared to traditional LLMs, MoE LLMs can achieve higher performance with fewer active parameters, but it is still hard to deploy them due to their immense parameter sizes. Different from previous weight pruning methods that rely on specifically designed hardware, this paper mainly aims to enhance the deployment efficiency of MoE LLMs by introducing plug-and-play expert-level sparsification techniques. Specifically, we propose, for the first time to our best knowledge, post-training approaches for task-agnostic and task-specific expert pruning and skipping of MoE LLMs, tailored to improve deployment efficiency while maintaining model performance across a wide range of tasks. Extensive experiments show that our proposed methods can simultaneously reduce model sizes and increase the inference speed, while maintaining satisfactory performance. Code will be made available at https://github.com/Lucky-Lance/Expert_Sparsity.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Mixture-of-Experts Models
  • 2.2 Expert Pruning for MoE Models
  • 2.3 Post-training Pruning for LLMs
  • 3 Method
  • 3.1 Preliminary
  • 3.2 Post-training Expert Pruning
  • 3.3 Dynamic Skipping During Inference
  • 4 Experiment
  • 4.1 Expert Pruning for General Tasks
  • 4.2 Expert Pruning for Domain-Specific Tasks
  • 4.3 Dynamic Expert Skipping Results
  • 4.4 More Analysis
  • 5 Conclusion and Discussion
  • Limitations
  • Ethics Statement
  • Acknowledgement
  • References
  • A Appendix
  • A.1 Expert Selection Tendency in MoE Models
  • A.2 Theoretical Insight and Broader Application of Dynamic Skipping
  • A.3 Experiments on the Sizes of Calibration Datasets
  • A.4 Dynamic Skipping for Domain-specific Tasks
  • A.5 Actual Memory Reduction
  • A.6 Relationships with Other Network Pruning and Parameter Quantization Methods
  • A.7 More Experiment Details

Knowls

  1. Knowl 1 — Layer-wise reconstruction-based post-training expert pruning

    model/method

    The paper introduces a hardware-friendly, post-training method that permanently removes unimportant experts from every MoE layer without updating model parameters. For each layer ℓ\ell with nn experts, a calibration set is passed through the original layer to cache input-output pairs. For a target of rr retained experts, every subset CC of rr experts is evaluated by deleting the other experts and their routing weights, and the subset with the smallest output reconstruction error is retained:

    Cℓ∗=argmin⁡C⊆{0,…,n−1}, ∣C∣=r ∥Fℓ′(x,C)−Fℓ(x)∥F.C_\ell^* = \underset{C \subseteq \{0,\ldots,n-1\},\ |C|=r}{\operatorname{argmin}}\ \left\|F'_\ell(x,C)-F_\ell(x)\right\|_F.

    Here, Fℓ(x)F_\ell(x) is the output of the original MoE layer for calibration input xx, Fℓ′(x,C)F'_\ell(x,C) is the output after retaining only expert subset CC, and ∥⋅∥F\|\cdot\|_F is the Frobenius norm over the layer output. The search is performed independently for each layer, and the selected pruned layers are concatenated into the final model. The resulting model can be loaded with standard model software by changing the model configuration, requiring no additional training. The enumeration costs (nr)\binom{n}{r} candidate subsets per layer; on Mixtral 8x7B, pruning to r=6r=6 experts took about 30 minutes and pruning to r=4r=4 took about 90 minutes. A comparison of search strategies found that the independent layer-wise method achieved average LM-eval scores of 64.22 for r=6r=6 and 59.57 for r=4r=4, versus 64.48 and 57.53 for a progressive search that conditions later layers on earlier pruning decisions.

  2. Knowl 2 — Task-specific calibration for domain-specific expert pruning

    model/method

    The pruning criterion can be adapted to a target domain by replacing general-purpose calibration data with examples from that domain. The paper uses C4 samples for task-agnostic pruning and samples from the MATH training set for mathematical reasoning. The selected expert subsets differ substantially: for Mixtral 8x7B with six retained experts per layer, the C4- and MATH-based selections coincide in only layers 2, 4, 16, and 31, indicating that expert usefulness depends on the calibration distribution.

    This change is important for mathematical evaluation. On 5-shot GSM8K, C4-calibrated pruning reduces Mixtral 8x7B from 58.61 accuracy without pruning to 41.02 with r=6r=6 and 24.87 with r=4r=4, whereas MATH-calibrated pruning gives 51.25 and 37.07. For Mixtral 8x7B Instruct, the corresponding values are 63.46 without pruning, 48.52 and 30.40 with C4 calibration, and 58.38 and 47.01 with MATH calibration. Thus, domain-specific calibration substantially preserves performance on the corresponding downstream task, although post-training pruning still causes a noticeable loss.

  3. Knowl 3 — Dynamic top-2 expert skipping during inference

    algorithm

    The paper introduces an online inference method that skips the weaker of the two experts selected for a token, without permanently deleting any expert. Let e0e_0 and e1e_1 be the top-2 routed experts for token xx, with routing weights we0≥we1w_{e_0} \geq w_{e_1}. For each MoE layer ℓ\ell, a threshold βℓ\beta_\ell is calibrated from a small dataset as the median of the observed ratios we1/we0w_{e_1}/w_{e_0}. During inference, the token is sent only to e0e_0 when

    we1we0<βℓ;\frac{w_{e_1}}{w_{e_0}} < \beta_\ell;

    otherwise, both selected experts process the token. The skipped expert contributes no computation for that token, while the retained expert produces the layer output under the model's routing normalization. The method is applied online, does not change memory usage, and can be combined with permanent expert pruning. Choosing the median ratio makes skipping occur for approximately half of the calibrated top-2 cases in each layer.

  4. Knowl 4 — Mixture-of-Experts routing and weighted layer output

    equation

    In the Mixtral 8x7B decoder, each MoE feed-forward layer has n=8n=8 experts and routes each token to the top k=2k=2 experts. For a token representation xx, the router produces logits l=(l0,…,ln−1)l=(l_0,\ldots,l_{n-1}) and routing weights w=Softmax⁡(l)w=\operatorname{Softmax}(l). Let eje_j be the index of the jj-th selected expert, and let Eej(x)E_{e_j}(x) be that expert's output. The selected weights are renormalized as

    w~ej=wej∑m=0k−1wem,j∈{0,…,k−1},\widetilde{w}_{e_j}=\frac{w_{e_j}}{\sum_{m=0}^{k-1}w_{e_m}},\qquad j\in\{0,\ldots,k-1\},

    where k=2k=2 for Mixtral. The MoE layer output is

    z=∑j=0k−1w~ejEej(x).z=\sum_{j=0}^{k-1}\widetilde{w}_{e_j}E_{e_j}(x).

    This routing structure explains why permanent expert pruning reduces stored parameters and why dynamic skipping can reduce computation: only a small subset of the layer's experts contributes to each token even though all experts are normally stored.

  5. Knowl 5 — Generalized reconstruction bound for dynamic skipping

    theoretical result

    For a top-kk MoE layer with selected routing weights ordered as w1≥w2≥⋯≥wkw_1\geq w_2\geq\cdots\geq w_k, let fm=Em(x)f_m=E_m(x) be the output of selected expert mm. The original normalized layer output is z=(∑m=1kwmfm)/(∑m=1kwm)z=(\sum_{m=1}^k w_m f_m)/(\sum_{m=1}^k w_m). If dynamic skipping retains only the top ii experts, where 1≤i≤k1\leq i\leq k, the output is z^=(∑m=1iwmfm)/(∑m=1iwm)\hat z=(\sum_{m=1}^i w_m f_m)/(\sum_{m=1}^i w_m).

    Under the paper's simplifying assumption that all pairwise expert-output distances satisfy ∥fm−fn∥2=D\|f_m-f_n\|_2=D for m≠nm\ne n, the reconstruction loss obeys

    ∥z^−z∥2≤∑m=i+1kwm∑m=1kwmD.\|\hat z-z\|_2 \leq \frac{\sum_{m=i+1}^{k}w_m}{\sum_{m=1}^{k}w_m}D.

    Therefore, if an allowable loss bound HH satisfies H≤DH\leq D, retaining the smallest number i∗i^* of experts whose omitted routing mass obeys

    ∑m=i∗+1kwm≤HD∑m=1kwm\sum_{m=i^*+1}^{k}w_m \leq \frac{H}{D}\sum_{m=1}^{k}w_m

    controls the bound under this assumption. For the computationally simplified top-2 rule used in the implementation, the appendix expresses the criterion as w2≤βw1w_2\leq\beta w_1, with β=H/(D−H)\beta=H/(D-H) under its top-2 reparameterization.

  6. Knowl 6 — General-task accuracy after expert pruning

    empirical result

    The reconstruction-based pruning method was evaluated without additional training on eight zero-shot LM-evaluation tasks: ARC-c, ARC-e, BoolQ, HellaSwag, MMLU, OBQA, RTE, and WinoGrande. On Mixtral 8x7B, the original eight-expert model achieved an average accuracy of 67.58; the proposed method achieved 64.22 with six retained experts per layer and 59.57 with four. On Mixtral 8x7B Instruct, the corresponding averages were 69.98, 67.45, and 63.88.

    The method outperformed two expert-selection baselines. Random deletion produced averages of 63.04 and 56.41 for Mixtral 8x7B with six and four retained experts, and 67.03 and 61.15 for the Instruct model. Deleting experts with the lowest activation frequency produced 60.77 and 55.83 for Mixtral 8x7B, and 65.48 and 60.92 for the Instruct model. These results show that activation frequency alone is a poor proxy for an expert's contribution, while minimizing layer-output reconstruction error preserves more task performance.

  7. Knowl 7 — Memory reduction and comparison with structured weight pruning

    empirical result

    On Mixtral 8x7B in bf16, the full eight-expert model used 89,926 MB of measured memory, while the proposed method used 68,383 MB with r=6r=6 and 46,879 MB with r=4r=4. Both pruned models fit on one 80 GB A100 GPU, whereas the unpruned model required two such GPUs for loading and inference. Token-generation speedups were 1.20×\times for r=6r=6 and 1.27×\times for r=4r=4 relative to the original model.

    At approximately 50% parameter reduction, the proposed r=4r=4 expert pruning method was compared with Wanda using structured 2:4 weight sparsity. For Mixtral 8x7B, Wanda obtained average accuracy 57.51, memory 51,214 MB, and speedup 0.91×\times, while expert pruning obtained 59.57, 46,879 MB, and 1.27×\times. For Mixtral 8x7B Instruct, Wanda obtained 62.80, 51,210 MB, and 0.92×\times, while expert pruning obtained 63.88, 46,879 MB, and 1.27×\times. The paper notes that the 2:4 speed benefit depends on specialized hardware and scripts; in its implementation, it was slower than the dense model.

  8. Knowl 8 — Complementarity of pruning and dynamic skipping

    empirical result

    Dynamic skipping further improves inference speed after expert pruning, with only a moderate additional accuracy loss. On Mixtral 8x7B, the original model scored 67.58 with a 1.00×\times speed baseline; skipping without pruning scored 66.37 at 1.08×\times, pruning to r=6r=6 scored 64.22 at 1.19×\times, and combining r=6r=6 pruning with skipping scored 62.91 at 1.23×\times. Pruning to r=4r=4 scored 59.57 at 1.27×\times, while combining it with skipping scored 57.91 at 1.31×\times.

    On Mixtral 8x7B Instruct, the corresponding pairs of average LM-eval accuracy and speedup were 69.98 and 1.00×\times for the original model, 69.03 and 1.08×\times for skipping alone, 67.45 and 1.20×\times for r=6r=6 pruning, 66.04 and 1.27×\times for combined r=6r=6 pruning and skipping, 63.88 and 1.27×\times for r=4r=4 pruning, and 62.33 and 1.33×\times for combined r=4r=4 pruning and skipping. Thus, combined r=6r=6 pruning and skipping reaches the same 1.27×\times speedup as r=4r=4 pruning alone while retaining higher average accuracy, 66.04 versus 63.88.

  9. Knowl 9 — Fine-tuning recovers task-specific pruning loss

    empirical result

    After task-specific pruning, the paper fully fine-tuned pruned and unpruned Mixtral models on MetaMathQA for 900 steps using 16 A100-80G GPUs, a learning rate of 2×10−52\times10^{-5}, and cosine scheduling. Before fine-tuning, MATH-calibrated pruning improved 5-shot GSM8K substantially over C4-calibrated pruning, but still reduced accuracy. Fine-tuning narrowed these gaps.

    For Mixtral 8x7B, the unpruned eight-expert model reached 81.35 GSM8K and 34.86 MATH accuracy after fine-tuning; the C4-calibrated six-expert model reached 79.53 and 32.48; the MATH-calibrated six-expert model reached 79.53 and 33.58; and the MATH-calibrated seven-expert model reached 81.20 and 34.40. For Mixtral 8x7B Instruct, the corresponding results were 81.43 and 35.46 for the unpruned model, 79.83 and 32.70 for C4-calibrated r=6r=6, 80.06 and 34.10 for MATH-calibrated r=6r=6, and 81.50 and 34.86 for MATH-calibrated r=7r=7. The seven-expert Instruct model slightly exceeded the eight-expert model on GSM8K, suggesting that fine-tuned downstream tasks may not require all experts.

  10. Knowl 10 — Scalability limitation of exhaustive expert selection

    limitation

    The pruning method relies on exhaustive enumeration of expert subsets, so its search cost grows as (nr)\binom{n}{r} for nn experts per layer. This is practical for the four- and eight-expert MoE layers studied in the paper, but becomes cumbersome for layers with substantially more experts, such as 32. In addition, the experiments cover only the open Mixtral 8x7B and Mixtral 8x7B Instruct models. The paper therefore does not establish generality across other MoE architectures, expert counts, or routing implementations.

Coverage note — Only secondary calibration-set-size sensitivity results, the full per-layer skipping-threshold lists, and auxiliary domain-specific skipping measurements were omitted because they do not materially change the main pruning, skipping, or performance conclusions.

References

  1. 1.Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2023. Llm in a flash: Efficient large language model inference with limited memory.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei. 2022. Task-specific expert pruning for sparse mixture-of-experts. arXiv preprint arXiv:2206.00277.
  4. 4.Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2022. On the representation collapse of sparse mixture of experts. In Advances in Neural Information Processing Systems.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  6. 6.Shuangrui Ding, Peisen Zhao, Xiaopeng Zhang, Rui Qian, Hongkai Xiong, and Qi Tian. 2023. Prune spatio-temporal tokens by semantic-aware temporal accumulation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16945–16956.
  7. 7.William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research, 23(1):5232–5270.
  8. 8.Elias Frantar and Dan Alistarh. 2023. Sparsegpt: Massive language models can be accurately pruned in one-shot. In International Conference on Machine Learning, pages 10323–10337. PMLR.
  9. 9.Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323.
  10. 10.Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. 2023. Megablocks: Efficient sparse training with mixture-of-experts. Proceedings of Machine Learning and Systems, 5.
  11. 11.Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023. A framework for few-shot language model evaluation.
  12. 12.Yihui He, Xiangyu Zhang, and Jian Sun. 2017. Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE international conference on computer vision, pages 1389–1397.
  13. 13.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. NeurIPS.
  14. 14.Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry. 2021. Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks. Advances in neural information processing systems, 34:21099–21111.
  15. 15.Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991. Adaptive mixtures of local experts. Neural computation, 3(1):79–87.
  16. 16.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  17. 17.Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. 2022. Learned token pruning for transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 784–794.
  18. 18.Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. 2021. Scalable and efficient moe training for multitask multilingual models. arXiv preprint arXiv:2109.10465.
  19. 19.Yeskendir Koishekenov, Alexandre Berard, and Vassilina Nikoulina. 2022. Memory-efficient nllb-200: Language-specific expert pruning of a massively multilingual machine translation model. arXiv preprint arXiv:2212.09811.
  20. 20.Woosuk Kwon, Sehoon Kim, Michael W Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. 2022. A fast post-training pruning framework for transformers. Advances in Neural Information Processing Systems, 35:24101–24116.
  21. 21.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668.
  22. 22.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv preprint arXiv:2306.00978.
  23. 23.Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378.
  24. 24.OpenAI. 2023. Gpt-4 technical report. ArXiv, abs/2303.08774.
  25. 25.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints.
  26. 26.Omar Sanseviero, Lewis Tunstall, Philipp Schmid, Sourab Mangrulkar, Younes Belkada, and Pedro Cuenca. 2023. Mixture of experts explained.
  27. 27.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  28. 28.Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695.
  29. 29.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  30. 30.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  31. 31.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  32. 32.Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li. 2024. Mathcoder: Seamless code integration in LLMs for enhanced mathematical reasoning. In The Twelfth International Conference on Learning Representations.
  33. 33.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  34. 34.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284.
  35. 35.Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning n: m fine-grained structured sparse neural networks from scratch. arXiv preprint arXiv:2102.04010.
  36. 36.Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, and Hongsheng Li. 2024. Solving challenging math word problems using GPT-4 code interpreter with code-based self-verification. In The Twelfth International Conference on Learning Representations.

Citation

MLA
Lu, X., et al. “Not All Experts Are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6159–72, https://doi.org/10.18653/v1/2024.acl-long.334.
APA
Lu, X., Liu, Q., Xu, Y., Zhou, A., Huang, S., Zhang, B., Yan, J., & Li, H. (2024). Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6159–6172. https://doi.org/10.18653/v1/2024.acl-long.334
Chicago
Lu, X., Q. Liu, Y. Xu, et al. 2024. “Not All Experts Are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6159–72. https://doi.org/10.18653/v1/2024.acl-long.334.
Harvard
Lu, X. et al. (2024) “Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6159–6172. Available at: https://doi.org/10.18653/v1/2024.acl-long.334.
Vancouver
1. Lu X, Liu Q, Xu Y, Zhou A, Huang S, Zhang B, Yan J, Li H (2024) Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6159–6172

BibTeX

@inproceedings{lu-etal-2024-experts,
    title = "Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models",
    author = "Lu, Xudong  and
      Liu, Qi  and
      Xu, Yuhui  and
      Zhou, Aojun  and
      Huang, Siyuan  and
      Zhang, Bo  and
      Yan, Junchi  and
      Li, Hongsheng",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.334/",
    doi = "10.18653/v1/2024.acl-long.334",
    pages = "6159--6172"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/