OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
Fuzhao XueZian ZhengYao FuJinjie NiZangwei ZhengWangchunshu ZhouYang You
Presents an open-source suite of decoder-only Mixture-of-Experts models alongside critical empirical analyses revealing that routing decisions depend heavily on fixed token IDs rather than context and cause token dropping late in sequences.
Large language models deliver impressive capabilities across diverse applications, but their high computational cost during both training and inference presents a significant barrier to scaling. Mixture-of-Experts architectures provide a viable pathway to expand model parameter capacity without a proportional surge in computation by activating only a subset of specialized subnetworks per token. However, transparent, reproducible research exploring how these architectures function in practice on trillion-token scales has been scarce. The article addresses this gap by training and releasing an open-source suite of decoder-only models and evaluating their routing mechanisms, training objectives, and real-world efficiency.
The main objective of the article is to demonstrate the feasibility of training fully transparent sparse language models from scratch, evaluate their performance trade-offs against conventional dense architectures, and systematically analyze how internal routing mechanisms allocate data to specialized subcomponents.
To conduct this evaluation, the researchers trained a family of models ranging from 650 million to 34 billion parameters using public text and code repositories, scaling up to over 1.1 trillion tokens on cloud-based hardware accelerators. The methodology incorporated sparse top-two routing across interleaved expert layers, experimented with a diverse denoising training objective, and utilized an extensive multi-lingual vocabulary before evaluating performance across standard benchmarks for coding, question answering, translation, and multi-turn conversational quality.
The article yields four core findings. First, sparse models deliver a superior cost-effectiveness trade-off, achieving comparable or superior results to dense baselines requiring substantially more training compute. On conversational benchmarks, the eight-billion-parameter model significantly outperformed dense alternatives on initial conversational turns. Second, routing decisions are primarily context-independent and driven by individual token identifiers rather than high-level sentence semantics, with specific subcomponents simply clustering low-level semantic tokens. Third, token assignment patterns are learned and solidified during the initial training warm-up phase and remain static even when shifting data mixtures or objectives. Fourth, enforcing strict capacity limits on subcomponents creates a pattern where tokens occurring later in a sequence are frequently dropped, leading to degraded performance in extended multi-turn interactions.
These findings indicate that while sparse architectures offer clear compute and capacity advantages, static and early routing specialization introduces operational risks for downstream applications. In particular, the drop-off in later tokens disproportionately impairs long-context tasks and sequential instruction-following dialogues. The results demonstrate that fine-tuning alone cannot resolve these imbalances because routing habits are permanently established during the earliest pre-training phase.
To address these architectural limitations, practitioners should implement balanced data strategies earlier in the development lifecycle. Specifically, future projects should introduce instruction-following data during the initial warm-up phase to ensure equitable routing across varied tasks. In addition, practitioners should moderate code data to roughly 30% of the training mix to avoid performance drops on general language tasks, explore converting dense checkpoints to sparse models after initial feature learning, and consider removing active routing mechanisms after warm-up to streamline hardware communication.
These conclusions are bounded by resource-constrained experimentation, as the largest model variations were trained on smaller token budgets and fine-tuning was tested on a limited conversational sample. While confidence is high regarding the presence of context-independent routing and token drops in sparse decoder setups, readers should exercise caution when generalising specific hyperparameter configurations to different model architectures or alternative training hardware.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Read this foundational work first to understand the sparsely gated expert layers and token-routing setup that underlie OpenMoE’s analysis.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Its treatment of routing stability and router z-loss provides essential context for OpenMoE’s discussion of routing behavior and token dropping.
- Paper: Unified Scaling Laws for Routed Language Models, Aidan Clark et al. (2022). Its scaling laws for routed language models help frame OpenMoE’s comparison of sparse and dense models across compute and parameter budgets.
- Paper: DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale, Samyam Rajbhandari et al. (2022). Its account of training and serving autoregressive MoE language models clarifies the efficiency and implementation trade-offs examined by OpenMoE.
No sufficiently relevant recommendations were found.
