Scaling Laws for Fine-Grained Mixture of Experts
Jan LudziejewskiJakub KrajewskiKamil AdamczewskiMaciej PióroMichal KrutulSzymon AntoniakKamil CiebieraKrystian KrólTomasz OdrzygózdzPiotr Sankowski
Derives compute-optimal scaling laws for fine-grained Mixture of Experts models by introducing expert granularity, proving that sparse architectures widen their computational efficiency advantage over dense Transformers as training budgets scale.
Training modern artificial intelligence models demands immense computing budgets, leading to steep capital costs and significant energy consumption. Sparse Mixture of Experts architectures, which dynamically route incoming data to specialized sub-networks called experts, offer a viable method to curb these training expenses by calculating only a portion of the network at any moment. However, prior research suggested that these sparse networks might lose their cost advantage to traditional dense models as computing budgets scale up. The article resolves this critical planning question by demonstrating that previous assessments relied on flawed assumptions, establishing that sparse architectures remain vastly superior when configured properly.
The main objective of the article is to establish new mathematical scaling laws for Mixture of Experts models by introducing fine-grained expert division and optimizing training duration alongside total parameters. It evaluates how varying expert sizes and training dataset volumes affects model performance under given compute budgets.
The authors conducted an extensive empirical study involving more than 100 model runs, training decoder-only language architectures sized between 129 million and 3.7 billion parameters on up to 130 billion tokens. They introduced a key architectural hyperparameter, granularity, which represents the splitting of standard experts into smaller, modular units without increasing the overall active parameter count per processed token. Using these experiments, the authors fitted a comprehensive scaling equation that directly models the relationship between model size, training tokens, expert granularity, and routing overhead.
The findings show that dividing experts into smaller, more granular pieces yields substantial, predictable performance improvements following a power-law relationship. Standard expert architectures with default granularity of one are almost never compute-optimal. More importantly, sparse models consistently outperform traditional dense models across every budget level when both training duration and granularity are tuned optimally. For instance, at moderate computational budgets, an optimized sparse model matches dense model performance using 20 times less compute, and this efficiency advantage widens to more than 40 times at larger scales. This directly disproves earlier claims that dense models overtake sparse models at extreme scales, showing that prior analyses mistakenly compared undertrained sparse systems.
These findings have major strategic implications for organizations investing in large language model training infrastructure. Adopting fine-grained sparse architectures enables massive cost reductions and shorter development timelines for equivalent model capabilities. Furthermore, engineering teams can now mathematically calculate the exact optimal model size, token count, and expert granularity for any target compute budget rather than relying on guesswork.
Decision-makers and engineering leads should move away from standard monolithic expert designs and adopt fine-grained routing configurations in upcoming model pre-training runs. When planning training budgets, teams should look up the recommended granularity values provided in the scaling formulas—such as setting granularity between 8 and 64 depending on total capacity—to maximize hardware utilization.
Organizations should note that increasing expert granularity introduces practical systems trade-offs, including increased communication overhead in distributed GPU environments and higher memory bandwidth pressure from routing operations. While the scaling predictions show high statistical stability across validation runs and bootstrapping intervals, practitioners should test specific hardware and networking limits in a small-scale pilot before finalizing massive training cluster configurations.
- Paper: Unified Scaling Laws for Routed Language Models, Aidan Clark et al. (2022). Establishes foundational scaling laws for routed language models across parameter counts and expert sizes that the source directly refines and corrects.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Introduces compute-optimal scaling principles (Chinchilla) balancing parameters and dataset tokens that the source extends to sparse architectures.
- Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). Pioneers the power-law empirical scaling framework for neural language models upon which Mixture of Experts scaling formulations build.
- Paper: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, William Fedus et al. (2022). Demonstrates practical scaling of sparse Mixture-of-Experts transformer architectures to massive parameter sizes, establishing the baseline routing paradigms evaluated in the source.
- Paper: ST-MoE: Designing Stable and Transferable Sparse Expert Models, Barret Zoph et al. (2022). Details critical stability techniques and architectural design guidelines necessary for reliably pre-training large-scale sparse expert models.
- Paper: Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, Noam Shazeer et al. (2017). Introduces the foundational sparsely-gated Mixture-of-Experts layer that decouples parameter capacity from per-token computation in deep neural networks.
- Paper: MegaBlocks: Efficient Sparse Training with Mixture-of-Experts, Trevor Gale et al. (2023). Presents block-sparse system implementations for Mixture-of-Experts that manage dynamic token routing overheads discussed as hardware trade-offs in the source.
- Paper: GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding, Dmitry Lepikhin et al. (2021). Provides the foundational distributed conditional computation framework for scaling sparse transformers across massive accelerator clusters.
- Paper: DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models, Damai Dai et al. (2024). Applies the principle of fine-grained expert segmentation alongside shared routing to optimize expert specialization and training efficiency in production language models.
- Paper: Scaling Embeddings Outperforms Scaling Experts in Language Models, Hong Liu et al. (2026). Investigates scaling memory-intensive embedding layers as a complementary alternative to scaling expert counts in high-sparsity language modeling regimes.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Provides a comprehensive architectural survey synthesizing modern efficient model paradigms, including fine-grained sparse Mixture-of-Experts systems.
- Paper: Skaling: Chinchilla's Exponents Meet Kaplan's Coupling, Mathurin Videau et al. (2026). Advances parametric scaling law formulations beyond standard Chinchilla power laws by modeling coupling dynamics between parameters and token counts.
- Paper: Sparser, Faster, Lighter Transformer Language Models, Edoardo Cetin et al. (2026). Explores activation-level sparsity mechanisms in transformer feedforward layers to further compress compute beyond structural expert routing.
