OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Fuzhao XueZian ZhengYao FuJinjie NiZangwei ZhengWangchunshu ZhouYang You

article2024ICML226 citations

Presents an open-source suite of decoder-only Mixture-of-Experts models alongside critical empirical analyses revealing that routing decisions depend heavily on fixed token IDs rather than context and cause token dropping late in sequences.

Listen

Large language models deliver impressive capabilities across diverse applications, but their high computational cost during both training and inference presents a significant barrier to scaling. Mixture-of-Experts architectures provide a viable pathway to expand model parameter capacity without a proportional surge in computation by activating only a subset of specialized subnetworks per token. However, transparent, reproducible research exploring how these architectures function in practice on trillion-token scales has been scarce. The article addresses this gap by training and releasing an open-source suite of decoder-only models and evaluating their routing mechanisms, training objectives, and real-world efficiency.

The main objective of the article is to demonstrate the feasibility of training fully transparent sparse language models from scratch, evaluate their performance trade-offs against conventional dense architectures, and systematically analyze how internal routing mechanisms allocate data to specialized subcomponents.

To conduct this evaluation, the researchers trained a family of models ranging from 650 million to 34 billion parameters using public text and code repositories, scaling up to over 1.1 trillion tokens on cloud-based hardware accelerators. The methodology incorporated sparse top-two routing across interleaved expert layers, experimented with a diverse denoising training objective, and utilized an extensive multi-lingual vocabulary before evaluating performance across standard benchmarks for coding, question answering, translation, and multi-turn conversational quality.

The article yields four core findings. First, sparse models deliver a superior cost-effectiveness trade-off, achieving comparable or superior results to dense baselines requiring substantially more training compute. On conversational benchmarks, the eight-billion-parameter model significantly outperformed dense alternatives on initial conversational turns. Second, routing decisions are primarily context-independent and driven by individual token identifiers rather than high-level sentence semantics, with specific subcomponents simply clustering low-level semantic tokens. Third, token assignment patterns are learned and solidified during the initial training warm-up phase and remain static even when shifting data mixtures or objectives. Fourth, enforcing strict capacity limits on subcomponents creates a pattern where tokens occurring later in a sequence are frequently dropped, leading to degraded performance in extended multi-turn interactions.

These findings indicate that while sparse architectures offer clear compute and capacity advantages, static and early routing specialization introduces operational risks for downstream applications. In particular, the drop-off in later tokens disproportionately impairs long-context tasks and sequential instruction-following dialogues. The results demonstrate that fine-tuning alone cannot resolve these imbalances because routing habits are permanently established during the earliest pre-training phase.

To address these architectural limitations, practitioners should implement balanced data strategies earlier in the development lifecycle. Specifically, future projects should introduce instruction-following data during the initial warm-up phase to ensure equitable routing across varied tasks. In addition, practitioners should moderate code data to roughly 30% of the training mix to avoid performance drops on general language tasks, explore converting dense checkpoints to sparse models after initial feature learning, and consider removing active routing mechanisms after warm-up to streamline hardware communication.

These conclusions are bounded by resource-constrained experimentation, as the largest model variations were trained on smaller token budgets and fine-tuning was tested on a limited conversational sample. While confidence is high regarding the presence of context-independent routing and token drops in sparse decoder setups, readers should exercise caution when generalising specific hyperparameter configurations to different model architectures or alternative training hardware.

No sufficiently relevant recommendations were found.

Cover for OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Abstract

To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-sourced and reproducible decoder-only MoE LLMs, ranging from 650M to 34B parameters and trained on up to over 1T tokens. Our investigation confirms that MoE-based LLMs can offer a more favorable cost-effectiveness trade-off than dense LLMs, highlighting the potential effectiveness for future LLM development.

One more important contribution of this study is an in-depth analysis of the routing mechanisms within our OpenMoE models, leading to three significant findings: Context-Independent Specialization, Early Routing Learning, and Drop-towards-the-End. We discovered that routing decisions in MoE models are predominantly based on token IDs, with minimal context relevance. The token-to-expert assignments are determined early in the pre-training phase and remain largely unchanged. This imperfect routing can result in performance degradation, particularly in sequential tasks like multi-turn conversations, where tokens appearing later in a sequence are more likely to be dropped. Finally, we rethink our design based on the above-mentioned observations and analysis. To facilitate future MoE LLM development, we propose potential strategies for mitigating the issues we found and further improving off-the-shelf MoE LLM designs.

Table of Contents

  • 1. Introduction
  • 2. Designing OpenMoE
  • 2.1. Pre-training Dataset: More Code than Usual
  • 2.2. Model Architecture: Decoder-only ST-MoE
  • 2.3. Training Objective: UL2 and CasualLM
  • 2.4. Supervised Fine-tuning
  • 2.5. Other Designs
  • 3. Training OpenMoE
  • 3.1. Training Progress
  • 3.2. Evaluation on Benchmarks
  • 3.2.1. Raw Model Evaluation
  • 3.2.2. Chat Model Evaluation
  • 4. Analyzing OpenMoE
  • 4.1. What Are the Experts Specializing In?
  • 4.2. Token Specialization Study
  • 4.3. Token Drop During Routing
  • 5. Rethinking OpenMoE
  • 6. Conclusion
  • Acknowledgement
  • Impact Statement
  • References
  • A. Frequently Asked Questions
  • A.1. Why Not Show the Token Specialization of the Checkpoints at the Warmup Stage?
  • A.2. Why Not Compare With Advanced Open MoE Models Like Mixtral and DeepSeek-MoE?
  • A.3. Why Not Use MoE Upcycling (Komatsuzaki et al., 2023)?
  • A.4. Why Not Use AdamW Optimizer and Cosine Learning Rate Schedule?
  • A.5. Why Not Use Better and Larger Datasets?
  • B. Related Work
  • B.1. Before OpenMoE
  • B.2. After OpenMoE
  • C. Data Mixture
  • D. Model Architecture
  • E. UL2 Training Objective
  • F. Hyper-Parameters
  • G. Ablation Study
  • H. BigBench-Lite Results
  • I. LM-Evaluation-Harness Results
  • J. MT-Bench Results
  • K. Language-Level Specialization
  • L. Position ID Specialization
  • M. Layer ID Specialization
  • N. Routing Decision Standard Deviation
  • O. Top Token Selection by Experts
  • P. Study Other MoE Models
  • Q. Tokenizer Analysis

Knowls

  1. Knowl 1 — OpenMoE models and open-release scope

    model/method

    OpenMoE is a decoder-only, sparsely activated language-model family trained from scratch and released with its training implementation and data details. Its models include OpenMoE-Base/16E (650M total parameters, 16 experts per MoE layer, intended partly for debugging), OpenMoE-8B/32E (8.7B total parameters and about 2.1B Transformer parameters active per token, trained on about 1.1T tokens), and OpenMoE-34B/32E (34B total parameters and about 6.0B active Transformer parameters, trained on 200B tokens). The release also includes OpenMoE-8B/32E-Chat, fine-tuned on a 58K-conversation GPT-4 subset of WildChat, and five intermediate 8B-model checkpoints spaced by 200B training tokens.

  2. Knowl 2 — Interleaved top-2 residual MoE architecture

    model/method

    OpenMoE uses token-choice routing with a learned linear router: for each token representation, the router applies softmax scores over the experts and selects the top two experts. Each expert is a feed-forward network, not a full Transformer. In an MoE Transformer block, the attention sublayer is followed by a residual feed-forward computation that adds both the routed MoE output and a conventional, always-active feed-forward network to the block residual. MoE blocks are interleaved with dense blocks: every fourth Transformer block in OpenMoE-Base/16E and OpenMoE-34B/32E, and every sixth block in OpenMoE-8B/32E. Training uses an MoE load-balancing auxiliary loss and a router z-loss, alongside language-model cross-entropy, to encourage balanced expert use and improve router stability.

  3. Knowl 3 — OpenMoE pre-training data and objective schedule

    experimental setup

    OpenMoE's initial data mixture combined RedPajama with a duplicated version of The Stack and contained 52.25% code. The authors later judged this code proportion too aggressive: for OpenMoE-8B/32E, after 780B tokens they switched from UL2 to causal language modeling and reduced the code sampling ratio to 15%. OpenMoE-34B/32E used UL2 for its first 50B tokens, then causal language modeling, with code making up 35% of its overall mixture. The UL2 mixture assigned 50% of training to PrefixLM with mask ratio r=0.5r=0.5; the other 50% was divided equally among five SpanCorrupt settings: mean span length μ=3\mu=3 or 88 with r=0.15r=0.15, μ=3\mu=3 or 88 with r=0.5r=0.5, and μ=64\mu=64 with r=0.5r=0.5. Models used the 256K-vocabulary umT5 tokenizer, Adafactor, an inverse-square-root learning-rate schedule, peak learning rate 0.01, 10K warmup steps, and sequence length 2048. Training used Google Cloud TPUv3 systems with 64–512 chips, depending on availability.

  4. Knowl 4 — Benchmark performance and limits of the released models

    empirical result

    On several reported raw-model benchmarks, OpenMoE-8B/32E matched or exceeded dense baselines despite fewer activated parameters and fewer training tokens. On HumanEval Pass@1, TinyLLaMA-1.1B scored 9.1, OpenLLaMA-3B 0, OpenMoE-8B/32E 9.8, and OpenMoE-34B/32E 10.3; their respective active Transformer parameters, total training tokens, and code tokens were 0.9B/3.0T/900B, 2.9B/1.0T/59B, 2.1B/1.1T/456B, and 6.4B/0.2T/70B. On TriviaQA exact match, those models scored 11.2, 29.7, 32.7, and 31.3, respectively. On WMT16 English-to-Romanian BLEU, they scored 2.6, 1.9, 3.1, and 3.4. OpenMoE-8B/32E also scored 6.93 average on BigBench-Lite, compared with 5.40 for GPT-3 6B, 4.63 for BIG-G-Sparse 8B, and 3.77 for BIG-G 8B; the paper reports a favorable cost-effectiveness trade-off when cost is measured using active parameters multiplied by training tokens. On MT-Bench, OpenMoE-8B/32E-Chat scored 4.69 on the first turn and 3.26 on the second, for a 3.98 average; OpenLLaMA-3B scored 4.36, 3.62, and 3.99. Thus OpenMoE's first-turn score was higher, while its second-turn score was lower. On five-shot MMLU, OpenMoE-8B/32E achieved about 26.2%, close to random choice among four options. A small-scale OpenMoE-Base/16E zero-shot TriviaQA ablation scored 1.4 EM and 4.5 F1; removing MoE yielded 0.1/0.3, using only PrefixLM instead of the UL2 mixture yielded 0.0/0.0, and removing code data yielded 0.7/1.1. Substituting the LLaMA tokenizer instead scored 2.2/5.7. The authors caution that the Base-model ablations may not generalize to larger scales.

  5. Knowl 5 — Token routing is predominantly context-independent

    empirical result

    In OpenMoE, routing preferences for individual token IDs are strong even when a token appears in varied sentence contexts. The authors therefore characterize the router as relying predominantly on token identity rather than context-dependent, higher-level semantics. For example, tokens such as “an” and “ed” show concentrated expert preferences despite appearing in many different words and contexts. Expert token preferences also group related forms: expert 30 favors “have,” “has,” and “had,” while expert 31 favors “can,” “will,” and “would”; expert 21 favors tokens including “=”, “and,” and newline, which helps explain its frequent selection for code and math. The authors report this pattern for most examined token IDs, with routing variation by token ID greater than variation by position ID. They also found clear token-ID specialization in DeepSeek-MoE, but not in Mixtral; their suggestion that Mixtral's likely dense-model upcycling explains the difference is a hypothesis, not a demonstrated cause.

  6. Knowl 6 — Token-to-expert preferences form early and persist

    empirical result

    Comparisons of OpenMoE intermediate checkpoints at 200B-token intervals—from 200B through 1.0T tokens—show that individual tokens' expert preferences are already very similar at the earliest available checkpoint and remain largely overlapped later in training. This stability persisted despite changes to the data mixture, including reducing the code proportion, and a change from UL2 to causal language modeling. The authors infer that routing specialization is probably established during warmup or another very early training phase; the proposed explanation that changing an established token-to-expert assignment would raise loss substantially is an interpretation rather than a separately measured result.

  7. Knowl 7 — Fixed expert capacity causes drop-towards-the-end

    empirical result

    When an expert has a fixed token capacity, tokens assigned to it after that capacity is filled are dropped. In an autoregressive decoder, this creates a positional bias: later tokens in a sequence face a greater risk of being dropped. OpenMoE's general pre-training datasets, including RedPajama and The Stack, show relatively balanced assignments and few dropped tokens, whereas multilingual and instruction-following data show substantially more dropping at later positions. The authors attribute this contrast to routing and load balance having been established on pre-training data, making instruction data a distribution shift for the router. Fine-tuning on 58K instruction conversations did not produce a significant reduction in the position-dependent drop pattern. The authors suggest this issue probably contributes to OpenMoE's weaker second-turn MT-Bench performance. They also manually imposed a capacity-based drop mechanism while analyzing normally dropless Mixtral and DeepSeek-MoE, and observed increasing drops at later positions in both; this demonstrates the effect under the imposed capacity constraint, not that those models natively drop tokens.

  8. Knowl 8 — Routing specialization differs across data groupings

    empirical result

    OpenMoE's routing analyses found little domain-level specialization: tokens from different RedPajama subsets were mostly distributed similarly, although expert 21 showed a slight preference for code and expert 10 a slight preference for books. Differences were clearer among natural languages: Simplified and Traditional Chinese favored experts 5 and 16, while Japanese and Korean favored expert 14. By contrast, four examined programming languages did not show similarly clear specialization. MT-Bench task categories also showed routing differences, especially for math; the authors suggest that math's greater use of special tokens may help explain this. Position IDs showed some specialization, but less than token IDs, and consecutive positions tended to favor similar experts. These observations qualify the token-ID finding: routing is not identical across all aggregate categories, even though individual token identity is the strongest reported signal.

  9. Knowl 9 — Proposed changes to improve MoE efficiency and instruction routing

    model/method

    Based on their routing analyses, the authors propose several future design changes rather than experimentally validated improvements. Because token routing appears largely context-independent, they suggest removing the trainable router after warmup, computing a parallel Transformer feed-forward layer directly from the input rather than from the attention output, and overlapping attention computation with MoE all-to-all communication. They argue that router removal and communication overlap could improve hardware utilization, while the parallel-layer design could enable overlap without a performance drop at scale. To reduce instruction-time token dropping, they also propose mixing instruction-following examples into pre-training warmup—not primarily to teach instruction following, but to establish balanced routing on that distribution before routing preferences become fixed.

  10. Knowl 10 — umT5 tokenizer trade-offs across languages and code

    empirical result

    OpenMoE used the 256K-vocabulary umT5 tokenizer. In the paper's tokenizer comparison, the ratio “umT5/LLaMA” is the number of tokens produced by umT5 divided by the number produced by the LLaMA tokenizer on the same subset. umT5 substantially reduced token counts for the sampled low-resource languages: the ratios were 0.344 for Arabic, 0.358 for Hebrew, 0.438 for Japanese, and 0.415 for Korean. For four sampled programming languages, umT5 instead produced more tokens: the ratios were 1.032 for Assembly, 1.031 for Blitzmax, 1.088 for Java, and 1.069 for Python. Results on MT-Bench conversation categories were mixed, with ratios from 0.929 to 1.061. The authors characterize umT5 as much better for multilingual text, especially low-resource languages, and slightly better on instruction-following data, but not as effective for code token savings as they had expected. They note that the large vocabulary adds output-layer computation and that predicting low-resource-language tokens may be costly when such data is scarce in pre-training.

Coverage note — Detailed per-task BigBench-Lite and LM-Evaluation-Harness scores, full data-mixture percentages, and the complete expert-token rankings are omitted because they add detail beyond the main findings already captured.

References

  1. 1.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  2. 2.Anonymous. (inthe)wildchat: 570k chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Bl8u7ZRlbM.
  3. 3.Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., et al. Efficient large scale language modeling with mixtures of experts. arXiv preprint arXiv:2112.10684, 2021.
  4. 4.Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
  5. 5.bench authors, B. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=uyTL5Bvosj.
  6. 6.Bojar, O. r., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Jimeno Yepes, A., Koehn, P., Logacheva, V., Monz, C., Negri, M., Neveol, A., Neves, M., Popel, M., Post, M., Rubino, R., Scarton, C., Specia, L., Turchi, M., Verspoor, K., and Zampieri, M. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation, pp. 131–198, Berlin, Germany, August 2016. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W/W16/W16-2301.
  7. 7.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020.
  8. 8.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. 2021.
  9. 9.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  10. 10.Chung, H. W., Constant, N., Garcia, X., Roberts, A., Tay, Y., Narang, S., and Firat, O. Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining. arXiv preprint arXiv:2304.09151, 2023.
  11. 11.Computer, T. Redpajama: An open source recipe to reproduce llama training dataset, 2023. URL https://github.com/togethercomputer/RedPajama-Data.
  12. 12.Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066, 2024.
  13. 13.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pp. 5547–5569. PMLR, 2022.
  15. 15.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res, 23:1–40, 2021.
  16. 16.Fu, Yao; Peng, H. and Khot, T. How does gpt obtain its ability? tracing emergent abilities of language models to their sources. Yao Fu’s Notion, Dec 2022.
  17. 17.Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A. A framework for few-shot language model evaluation, 12 2023. URL https://zenodo.org/records/10256836.
  18. 18.Geng, X. and Liu, H. Openllama: An open reproduction of llama, May 2023. URL https://github.com/openlm-research/open_llama.
  19. 19.Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  20. 20.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  21. 21.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  22. 22.Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Barzilay, R. and Kan, M.-Y. (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://aclanthology.org/P17-1147.
  23. 23.Kenton, J. D. M.-W. C. and Toutanova, L. K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1, pp. 2, 2019.
  24. 24.Kocetkov, D., Li, R., Ben Allal, L., Li, J., Mou, C., Munoz Ferrandis, C., Jernite, Y., Mitchell, M., Hughes, S., Wolf, T., Bahdanau, D., von Werra, L., and de Vries, H. The stack: 3 tb of permissively licensed source code. Preprint, 2022.
  25. 25.Komatsuzaki, A., Puigcerver, J., Lee-Thorp, J., Ruiz, C. R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N. Sparse upcycling: Training mixture-of-experts from dense checkpoints. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=T5nUQDrM4u.
  26. 26.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  27. 27.Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L. Base layers: Simplifying training of large, sparse models. In International Conference on Machine Learning, pp. 6265–6274. PMLR, 2021.
  28. 28.Li, J., Zhang, Z., and Zhao, H. Self-prompting large language models for open-domain qa. arXiv preprint arXiv:2212.08635, 2022.
  29. 29.Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023.
  30. 30.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  31. 31.Lou, Y., Xue, F., Zheng, Z., and You, Y. Cross-token modeling with conditional computation. arXiv preprint arXiv:2109.02008, 2021.
  32. 32.Mustafa, B., Riquelme, C., Puigcerver, J., Jenatton, R., and Houlsby, N. Multimodal contrastive learning with limoe: the language-image mixture of experts. Advances in Neural Information Processing Systems, 35:9564–9576, 2022.
  33. 33.Nijkamp, E., Xie, T., Hayashi, H., Pang, B., Xia, C., Xing, C., Vig, J., Yavuz, S., Laban, P., Krause, B., et al. Xgen-7b technical report. arXiv preprint arXiv:2309.03450, 2023.
  34. 34.Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N. From sparse to soft mixtures of experts. arXiv preprint arXiv:2308.00951, 2023.
  35. 35.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  36. 36.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  37. 37.Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N. Scaling vision with sparse mixture of experts. Advances in Neural Information Processing Systems, 34:8583–8595, 2021.
  38. 38.Roller, S., Sukhbaatar, S., Weston, J., et al. Hash layers for large sparse models. Advances in Neural Information Processing Systems, 34:17555–17566, 2021.
  39. 39.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  40. 40.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  41. 41.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  42. 42.Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  43. 43.Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, E. P., Hajishirzi, H., Smith, N. A., Zettlemoyer, L., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K. Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv preprint, 2023.
  44. 44.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  45. 45.Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Bahri, D., Schuster, T., Zheng, H. S., Houlsby, N., and Metzler, D. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022a.
  46. 46.Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Wei, J., Wang, X., Chung, H. W., Bahri, D., Schuster, T., Zheng, S., et al. Ul2: Unifying language learning paradigms. In The Eleventh International Conference on Learning Representations, 2022b.
  47. 47.Team, L.-M. Llama-moe: Building mixture-of-experts from llama with continual pre-training, Dec 2023. URL https://github.com/pjlab-sys4nlp/llama-moe.
  48. 48.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  49. 49.Wang, B. and Komatsuzaki, A. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  50. 50.Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzman, F., Joulin, A., and Grave, E. CCNet: Extracting high quality monolingual datasets from web crawl data. In Calzolari, N., Bechet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., and Piperidis, S. (eds.), Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 4003–4012, Marseille, France, May 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://aclanthology.org/2020.lrec-1.494.
  51. 51.Xu, Y., Lee, H., Chen, D., Hechtman, B., Huang, Y., Joshi, R., Krikun, M., Lepikhin, D., Ly, A., Maggioni, M., et al. Gspmd: general and scalable parallelization for ml computation graphs. arXiv preprint arXiv:2105.04663, 2021.
  52. 52.Xue, F., He, X., Ren, X., Lou, Y., and You, Y. One student knows all experts know: From sparse to dense. arXiv preprint arXiv:2201.10890, 2022a.
  53. 53.Xue, F., Shi, Z., Wei, F., Lou, Y., Liu, Y., and You, Y. Go wider instead of deeper. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 8779–8787, 2022b.
  54. 54.Yu, P., Artetxe, M., Ott, M., Shleifer, S., Gong, H., Stoyanov, V., and Li, X. Efficient language modeling with sparse all-mlp. arXiv preprint arXiv:2203.06850, 2022.
  55. 55.Zhang, P., Zeng, G., Wang, T., and Lu, W. Tinyllama: An open-source small language model, 2024.
  56. 56.Zhang, Y., Kang, B., Hooi, B., Yan, S., and Feng, J. Deep long-tailed learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  57. 57.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  58. 58.Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al. Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114, 2022.
  59. 59.Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A. M., Lu, Y., et al. Brainformers: Trading simplicity for efficiency. In International Conference on Machine Learning, pp. 42531–42542. PMLR, 2023.
  60. 60.Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W. St-moe: Designing stable and transferable sparse expert models. URL https://arxiv.org/abs/2202.08906, 2022.

Citation

MLA
Xue, F., et al. “OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models”. arXiv, 2024, http://arxiv.org/abs/2402.01739v2.
APA
Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., & You, Y. (2024). OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models. arXiv. http://arxiv.org/abs/2402.01739v2
Chicago
Xue, F., Z. Zheng, Y. Fu, et al. 2024. “OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models”. arXiv. http://arxiv.org/abs/2402.01739v2.
Harvard
Xue, F. et al. (2024) “OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.01739v2.
Vancouver
1. Xue F, Zheng Z, Fu Y, Ni J, Zheng Z, Zhou W, You Y (2024) OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models. arXiv

BibTeX

@article{xue2024openmoe,
  title = {OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models},
  author = {Xue, Fuzhao and Zheng, Zian and Fu, Yao and Ni, Jinjie and Zheng, Zangwei and Zhou, Wangchunshu and You, Yang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.01739v2},
  eprint = {2402.01739}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/