Scaling Laws for Fine-Grained Mixture of Experts

Jan LudziejewskiJakub KrajewskiKamil AdamczewskiMaciej PióroMichal KrutulSzymon AntoniakKamil CiebieraKrystian KrólTomasz OdrzygózdzPiotr Sankowski

article2024ICML138 citations

Derives compute-optimal scaling laws for fine-grained Mixture of Experts models by introducing expert granularity, proving that sparse architectures widen their computational efficiency advantage over dense Transformers as training budgets scale.

Listen

Training modern artificial intelligence models demands immense computing budgets, leading to steep capital costs and significant energy consumption. Sparse Mixture of Experts architectures, which dynamically route incoming data to specialized sub-networks called experts, offer a viable method to curb these training expenses by calculating only a portion of the network at any moment. However, prior research suggested that these sparse networks might lose their cost advantage to traditional dense models as computing budgets scale up. The article resolves this critical planning question by demonstrating that previous assessments relied on flawed assumptions, establishing that sparse architectures remain vastly superior when configured properly.

The main objective of the article is to establish new mathematical scaling laws for Mixture of Experts models by introducing fine-grained expert division and optimizing training duration alongside total parameters. It evaluates how varying expert sizes and training dataset volumes affects model performance under given compute budgets.

The authors conducted an extensive empirical study involving more than 100 model runs, training decoder-only language architectures sized between 129 million and 3.7 billion parameters on up to 130 billion tokens. They introduced a key architectural hyperparameter, granularity, which represents the splitting of standard experts into smaller, modular units without increasing the overall active parameter count per processed token. Using these experiments, the authors fitted a comprehensive scaling equation that directly models the relationship between model size, training tokens, expert granularity, and routing overhead.

The findings show that dividing experts into smaller, more granular pieces yields substantial, predictable performance improvements following a power-law relationship. Standard expert architectures with default granularity of one are almost never compute-optimal. More importantly, sparse models consistently outperform traditional dense models across every budget level when both training duration and granularity are tuned optimally. For instance, at moderate computational budgets, an optimized sparse model matches dense model performance using 20 times less compute, and this efficiency advantage widens to more than 40 times at larger scales. This directly disproves earlier claims that dense models overtake sparse models at extreme scales, showing that prior analyses mistakenly compared undertrained sparse systems.

These findings have major strategic implications for organizations investing in large language model training infrastructure. Adopting fine-grained sparse architectures enables massive cost reductions and shorter development timelines for equivalent model capabilities. Furthermore, engineering teams can now mathematically calculate the exact optimal model size, token count, and expert granularity for any target compute budget rather than relying on guesswork.

Decision-makers and engineering leads should move away from standard monolithic expert designs and adopt fine-grained routing configurations in upcoming model pre-training runs. When planning training budgets, teams should look up the recommended granularity values provided in the scaling formulas—such as setting granularity between 8 and 64 depending on total capacity—to maximize hardware utilization.

Organizations should note that increasing expert granularity introduces practical systems trade-offs, including increased communication overhead in distributed GPU environments and higher memory bandwidth pressure from routing operations. While the scaling predictions show high statistical stability across validation runs and bootstrapping intervals, practitioners should test specific hardware and networking limits in a small-scale pilot before finalizing massive training cluster configurations.

Cover for Scaling Laws for Fine-Grained Mixture of Experts

Abstract

Mixture of Experts (MoE) models have emerged as a primary solution for reducing the computational cost of Large Language Models. In this work, we analyze their scaling properties, highlighting certain arbitrary assumptions present in the existing literature. In particular, we introduce a new hyperparameter, granularity, the modification of which allows for the optimal adjustment of the size of experts. Subsequently, we present scaling laws for fine-grained MoE, taking into account the number of training tokens, model size, and granularity. Using these scaling laws, we derive the optimal training configuration for a given computational budget. Furthermore, in contrast with previous works, we demonstrate that the gap in efficiency between dense and MoE models grows as we scale up the model size and training budget.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background
  • 3.1. Model Architecture
  • 3.2. Scaling Laws
  • 4. Granularity
  • 5. Scaling Laws
  • 5.1. Power Law With Respect to Granularity
  • 5.2. Scaling the Model and Dataset Size
  • 5.3. The Form of the Joint Scaling Law
  • 5.4. Fitting the Parametric Scaling Law
  • 5.5. MoE Scaling Properties
  • 6. Optimal Allocation of Computational Budget
  • 6.1. Computational Cost of Granularity
  • 6.2. Compute Optimal Formula
  • 6.3. MoE is Always More Efficient
  • 7. Discussion
  • 8. Conclusions
  • Reproducibility
  • Impact statement
  • Acknowledgements
  • Contributions
  • References
  • A. Architecture and Training Setup
  • B. Comparison to (Clark et al., 2022) with Effective Parameter Count Curve
  • C. Validation of the Scaling Law
  • D. Reliability of Compute Optimal Formula
  • E. Varying Expansion Rate
  • F. Choosing Granularity
  • G. FLOPs Constants
  • H. Measuring Wall-clock Time
  • I. Additional Visualizations

Knowls

  1. Knowl 1 — Fine-Grained MoE Architecture: Granularity and Expansion Rate

    definition

    In fine-grained Mixture of Experts (MoE) Transformers, the standard feed-forward network (FFN) with hidden dimension dffd_{ff} is divided into smaller expert sub-networks. The architecture is parameterized by two core hyperparameters:

    1. Granularity (GG): The ratio of the standard feed-forward hidden dimension dffd_{ff} to the hidden dimension of an individual fine-grained expert dexpertd_{expert}: G=dffdexpertG = \frac{d_{ff}}{d_{expert}} A standard non-granular MoE has G=1G = 1, where dexpert=dffd_{expert} = d_{ff}. When G>1G > 1, each expert is narrower. To maintain a constant number of active parameters per token (NactN_{act}), each input token is routed to GG granular experts instead of a single standard expert.

    2. Expansion Rate (RR): The ratio of the total parameters in an MoE layer NMoEN_{\text{MoE}} (excluding routing parameters) to the parameters in a standard dense feed-forward layer NffN_{ff}: R=NMoENffR = \frac{N_{\text{MoE}}}{N_{ff}}

    The total number of expert networks in a layer, denoted as NexpertN_{expert}, is given by: Nexpert=G⋅RN_{expert} = G \cdot R

    Increasing granularity GG at a fixed expansion rate RR increases the total pool of experts NexpertN_{expert} and the combinatorial flexibility of token-to-expert assignments while keeping the number of active parameters per token constant.

  2. Knowl 2 — Joint Parametric Scaling Law for Fine-Grained Mixture of Experts

    equation

    The cross-entropy loss L\mathcal{L} of a fine-grained Mixture of Experts (MoE) Transformer is modeled as a joint function of the total number of non-embedding parameters NN, the number of training tokens DD, and the architectural granularity GG:

    L(N,D,G)=c+(gGγ+a)1Nα+bDβ\mathcal{L}(N, D, G) = c + \left(\frac{g}{G^\gamma} + a\right)\frac{1}{N^\alpha} + \frac{b}{D^\beta}

    where:

    • L(N,D,G)\mathcal{L}(N, D, G) is the test cross-entropy loss (in nats).
    • N∈N+N \in \mathbb{N}^+ is the total number of non-embedding parameters in the model.
    • D∈N+D \in \mathbb{N}^+ is the number of tokens seen during training.
    • G∈R≥1G \in \mathbb{R}_{\ge 1} is the granularity hyperparameter (G=dff/dexpertG = d_{ff} / d_{expert}, where dffd_{ff} is the standard feed-forward hidden dimension and dexpertd_{expert} is the expert hidden dimension).
    • cc represents the irreducible entropy of the dataset distribution (the theoretical lower bound of next-token prediction loss for the natural text distribution).
    • a,b,g,α,β,γa, b, g, \alpha, \beta, \gamma are non-negative power-law fitting constants.
    • The term gGγ\frac{g}{G^\gamma} models the suboptimality reduction achieved by fine-grained expert routing, which converges to the additive constant aa as G→∞G \to \infty.
    • The term bDβ\frac{b}{D^\beta} captures underfitting due to finite training tokens DD and is independent of granularity GG.
  3. Knowl 3 — MoE Training FLOPs Formulation with Granular Routing Overhead

    equation

    The total floating-point operations (FLOPs) FF required to train a decoder-only fine-grained Mixture of Experts (MoE) Transformer on DD tokens, taking into account routing computation overhead, is given by:

    F=(12dmodel2cf+dmodelRGcr)⋅D⋅nblocksF = \left(12 d_{model}^2 c_f + d_{model} R G c_r\right) \cdot D \cdot n_{blocks}

    subject to the parameter counting and architectural constraints: N=dmodel2(8R+4)nblocksN = d_{model}^2 (8R + 4) n_{blocks} dmodel=64⋅nblocksd_{model} = 64 \cdot n_{blocks}

    where:

    • FF is the total computational training budget in FLOPs (excluding embedding and unembedding operations).
    • dmodeld_{model} is the Transformer hidden residual dimension.
    • nblocksn_{blocks} is the number of Transformer layers/blocks.
    • RR is the expansion rate of the MoE layer (R=NMoE/NffR = N_{\text{MoE}} / N_{ff}).
    • GG is the granularity (G=dff/dexpertG = d_{ff} / d_{expert}).
    • DD is the number of training tokens.
    • NN is the total number of non-embedding parameters, comprising 8Rdmodel28R d_{model}^2 in the feed-forward experts (assuming dff=4dmodeld_{ff} = 4 d_{model}) and 4dmodel24 d_{model}^2 in self-attention projections per block.
    • cf=6c_f = 6 is the FLOPs constant per active parameter per token in standard linear projections (2 FLOPs forward pass, 4 FLOPs backward pass).
    • cr=14c_r = 14 is the FLOPs constant per routing parameter per token, accounting for routing linear projections, token dispatch, score scaling, and gradient routing across forward and backward passes.
  4. Knowl 4 — Fitted Scaling Law Coefficients for Dense and MoE Transformers

    data/table

    The parametric scaling law parameters for dense decoder-only Transformers, L(N,D)=c+aNα+bDβ\mathcal{L}(N, D) = c + \frac{a}{N^\alpha} + \frac{b}{D^\beta}, and fine-grained MoE Transformers with expansion rate R=64R=64, L(N,D,G)=c+(gGγ+a)1Nα+bDβ\mathcal{L}(N, D, G) = c + \left(\frac{g}{G^\gamma} + a\right)\frac{1}{N^\alpha} + \frac{b}{D^\beta}, were fitted on over 100 training runs using Huber loss (threshold δ=0.1\delta = 0.1, weight decay 5×10−45 \times 10^{-4}) via the BFGS optimization algorithm:

    Model aa α\alpha bb β\beta gg γ\gamma cc
    MoE (R=64R=64) 18.1 0.115 30.8 0.147 2.1 0.58 0.47
    Dense 16.3 0.126 26.7 0.127 – – 0.47

    The fit yields a root-mean-square error (RMSE) of 0.0150.015 (and a validation RMSE of 0.0190.019 when holding out the top 20% lowest-perplexity models).

    Key implications of these fitted coefficients:

    1. Data scaling (β\beta): MoE exhibits a larger token exponent (β=0.147\beta = 0.147) than dense models (β=0.127\beta = 0.127), meaning MoE requires longer training to overcome initial underfitting but achieves faster relative loss reduction as tokens increase.
    2. Parameter scaling (α\alpha): Dense models scale more effectively with total parameters (α=0.126\alpha = 0.126) than MoE (α=0.115\alpha = 0.115), reflecting the fact that dense models activate all parameters per token while MoE activates only a sparse subset.
    3. Granularity scaling (g,γg, \gamma): Increasing granularity GG improves loss with power-law exponent γ=0.58\gamma = 0.58, reducing the model-size suboptimality multiplier from 20.220.2 at G=1G=1 toward 18.118.1 as G→∞G \to \infty.
  5. Knowl 5 — Compute-Optimal Allocations for Fine-Grained MoE Models

    data/table

    Solving the constrained optimization problem min⁡N,D,GL(N,D,G)\min_{N,D,G} \mathcal{L}(N,D,G) subject to FLOPs(N,D,G)=F\text{FLOPs}(N,D,G) = F (accounting for routing overhead) using Brent's numerical root-finding method yields the compute-optimal configurations for fine-grained MoE models (R=64R=64):

    Active Params (NactN_{act}) Tokens (DD) Optimal GG FLOPs Predicted Loss
    64×100M64 \times 100\text{M} 4.37B4.37\text{B} 8 2.95×10182.95 \times 10^{18} 3.133
    64×1B64 \times 1\text{B} 28.94B28.94\text{B} 16 1.93×10201.93 \times 10^{20} 2.491
    64×3B64 \times 3\text{B} 72.90B72.90\text{B} 16 1.41×10211.41 \times 10^{21} 2.245
    64×7B64 \times 7\text{B} 137.60B137.60\text{B} 32 6.46×10216.46 \times 10^{21} 2.076
    64×70B64 \times 70\text{B} 941.07B941.07\text{B} 32 4.16×10234.16 \times 10^{23} 1.694
    64×300B64 \times 300\text{B} 2.96T2.96\text{T} 64 5.69×10245.69 \times 10^{24} 1.503
    64×1T64 \times 1\text{T} 7.94T7.94\text{T} 64 4.97×10254.97 \times 10^{25} 1.367

    The table demonstrates three principal properties:

    1. Standard non-granular MoE (G=1G=1) is never compute-optimal at any training FLOPs budget.
    2. Optimal granularity increases monotonically with compute budget, growing from G=8G=8 for 100M100\text{M} active parameter models up to G=64G=64 for 1T1\text{T} active parameter models.
    3. The optimal token count DD scales roughly in proportion to active parameters NactN_{act} for medium models (1B1\text{B} to 10B10\text{B} active parameters), with smaller models requiring relatively longer training schedules and larger models requiring fewer tokens per parameter.
  6. Knowl 6 — Compute Efficiency Scaling of Fine-Grained MoE versus Dense Transformers

    empirical result

    When fine-grained Mixture of Experts (MoE) models and dense Transformers are both trained under compute-optimal allocations (jointly optimizing total non-embedding parameters NN, training tokens DD, and granularity GG), MoE models strictly outperform dense Transformers across all computational budgets, with the efficiency gap widening as compute increases:

    1. At a training budget of 102010^{20} FLOPs, a compute-optimal MoE model achieves the same cross-entropy loss as a dense Transformer trained with a 20×20\times larger computing budget.
    2. When the training budget reaches 102510^{25} FLOPs, the compute savings exceed 40×40\times, meaning a dense Transformer requires more than 40×40\times the FLOPs of the compute-optimal MoE to achieve equivalent loss.
    3. Compared to standard vanilla MoE (G=1G = 1), optimizing granularity (G>1G > 1) yields an additional 1.5×1.5\times to 2×2\times reduction in required training compute across FLOPs budgets from 101810^{18} to 102610^{26}.
  7. Knowl 7 — Resolution of the MoE Scaling Degradation Paradox at Fixed Dataset Sizes

    theoretical result

    Prior scaling literature (Clark et al., 2022) suggested that dense Transformers would eventually outperform Mixture of Experts (MoE) models at large parameter scales (>1T>1\text{T} parameters). This apparent degradation was an artifact of evaluating models on a fixed dataset size (D=130BD = 130\text{B} tokens).

    When training tokens DD are held fixed, the scaling curves of dense and MoE models intersect at a threshold NcrossN_{cross} where undertraining penalizes sparse MoE models more heavily than dense models. Solving for the crossing point where LMoE(N,D,G)=Ldense(N,D)\mathcal{L}_{\text{MoE}}(N, D, G) = \mathcal{L}_{\text{dense}}(N, D) yields:

    Fixed Dataset Size (DD) 10B10\text{B} tokens 130B130\text{B} tokens 1T1\text{T} tokens
    MoE/Dense Crossing Point (NcrossN_{cross}) 251B251\text{B} 1.9T1.9\text{T} 10T10\text{T}

    As the fixed token budget DD is increased, the crossing point shifts to progressively larger model sizes. Crucially, when models are trained compute-optimally (co-scaling dataset size DD with model size NN rather than fixing DD), the dense and MoE curves never cross, and MoE models maintain superior compute efficiency across all budgets.

  8. Knowl 8 — Training Pipeline and Architectural Stabilization for Fine-Grained MoE

    experimental setup

    The empirical validation of fine-grained MoE scaling laws employs the following training and architectural specifications:

    1. Data and Tokenization: Decoder-only Transformers trained on the C4 dataset using the GPT-2 tokenizer. Batches contain 0.5M0.5\text{M} tokens packed into sequences of length 2048 in an infinite data regime (no token repetition).
    2. Optimization: AdamW optimizer with weight decay 0.10.1, peak learning rate 2×10−42 \times 10^{-4}, 1%1\% linear warmup, and cosine learning rate decay down to 2×10−52 \times 10^{-5}. Model weights are initialized using a truncated normal distribution with reduced scale.
    3. Precision: Mixed precision training, maintaining the attention mechanism and routing layers in single precision.
    4. Routing Mechanism: Expert Choice routing with a token group size of 256 tokens grouped by position across sequences. Softmax is computed over the expert dimension, and tokens are chosen over the token dimension to achieve automatic load balancing.
    5. Layer Normalization: An extra layer normalization is inserted immediately prior to the output projection of each MoE layer. While providing minimal benefit for non-granular models (G=1G=1), this normalization is essential for numerical stability and performance when scaling to fine-grained experts (G>1G > 1).
  9. Knowl 9 — Scaling Parameters and Optimal Granularity for MoE with Expansion Rate R = 16

    data/table

    For fine-grained MoE models configured with an expansion rate of R=16R = 16, the parametric scaling law L(N,D,G)=c+(gGγ+a)1Nα+bDβ\mathcal{L}(N, D, G) = c + \left(\frac{g}{G^\gamma} + a\right)\frac{1}{N^\alpha} + \frac{b}{D^\beta} yields the following fitted coefficients:

    • a=19.64a = 19.64, α=0.124\alpha = 0.124, b=57.07b = 57.07, β=0.169\beta = 0.169, g=1.18g = 1.18, γ=0.986\gamma = 0.986, c=0.472c = 0.472.

    Bootstrapping analysis (sampling 80% of the experimental runs 100 times) gives the 10th and 90th percentile ranges for compute-optimal parameters at R=16R=16:

    Active Params (NactN_{act}) Tokens (DD) [10th, 90th percentile] Optimal GG [10th, 90th percentile]
    16×100M16 \times 100\text{M} [10.29B10.29\text{B}, 17.73B17.73\text{B}] [8, 16]
    16×1B16 \times 1\text{B} [53.74B53.74\text{B}, 103.54B103.54\text{B}] [16, 32]
    16×3B16 \times 3\text{B} [106.22B106.22\text{B}, 261.04B261.04\text{B}] [16, 32]
    16×7B16 \times 7\text{B} [177.65B177.65\text{B}, 511.43B511.43\text{B}] [16, 32]
    16×70B16 \times 70\text{B} [721.60B721.60\text{B}, 3.22T3.22\text{T}] [32, 64]
    16×300B16 \times 300\text{B} [1.73T1.73\text{T}, 10.69T10.69\text{T}] [32, 64]
    16×1T16 \times 1\text{T} [3.60T3.60\text{T}, 28.22T28.22\text{T}] [32, 128]

    These results demonstrate that the preference for higher granularity (G>1G > 1) and the growth of optimal GG with compute budget hold consistently across different MoE expansion rates.

  10. Knowl 10 — Performance Breakdown and Routing Bottlenecks at Extreme Granularity

    limitation

    While increasing granularity GG generally improves loss and parameter efficiency, two structural limitations arise at extreme granularity levels:

    1. Routing Parameter Dominance: In small models with large expansion rates and high granularity (e.g., dmodel=256d_{model} = 256, R=64R = 64, G=64G = 64), the parameter count of the routing layer (dmodel⋅R⋅G=256×64×64≈1.05Md_{model} \cdot R \cdot G = 256 \times 64 \times 64 \approx 1.05\text{M} per block) exceeds the total active parameters in the expert networks (8dmodel2=524K8 d_{model}^2 = 524\text{K} per block). Under this regime, the model suffers an observable performance degradation rather than monotonic improvement.
    2. Distributed System and Memory Overhead: Increasing GG expands the total expert count (Nexpert=G⋅RN_{expert} = G \cdot R), increasing inter-device all-to-all token transfer costs and memory overhead in distributed expert-parallel systems. Theoretical FLOPs compute optimality must therefore be reconciled with hardware-specific communication bandwidth.

Coverage note — No substantial contributed material was omitted; all main elements including definitions, scaling law formulations, FLOPs models, fitted parameters for R=64 and R=16, compute-optimal allocation tables, resolution of the dense vs MoE scaling controversy, experimental setup, and architectural limitations are covered.

References

  1. 1.Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C. Musiclm: Generating music from text, 2023.
  2. 2.Brent, R. P. An algorithm with guaranteed convergence for finding a zero of a function. Comput. J., 14:422–425, 1971. URL https://api.semanticscholar.org/CorpusID:10312755.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020.
  4. 4.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. Palm: Scaling language modeling with pathways, 2022.
  5. 5.Clark, A., de las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., van den Driessche, G., Rutherford, E., Hennigan, T., Johnson, M., Millican, K., Cassirer, A., Jones, C., Buchatskaya, E., Budden, D., Sifre, L., Osindero, S., Vinyals, O., Rae, J., Elsen, E., Kavukcuoglu, K., and Simonyan, K. Unified scaling laws for routed language models, 2022.
  6. 6.Dai, D., Deng, C., Zhao, C., Xu, R. X., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., Xie, Z., Li, Y. K., Huang, P., Luo, F., Ruan, C., Sui, Z., and Liang, W. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024.
  7. 7.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C. Glam: Efficient scaling of language models with mixture-of-experts, 2022.
  8. 8.Faiz, A., Kaneda, S., Wang, R., Osi, R., Sharma, P., Chen, F., and Jiang, L. Llmcarbon: Modeling the end-to-end carbon footprint of large language models, 2024.
  9. 9.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research, 23(1):5232–5270, 2022.
  10. 10.Frantar, E., Riquelme, C., Houlsby, N., Alistarh, D., and Evci, U. Scaling laws for sparsely-connected foundation models, 2023.
  11. 11.Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C. Scaling laws for neural machine translation, 2021.
  12. 12.Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., Hallacy, C., Mann, B., Radford, A., Ramesh, A., Ryder, N., Ziegler, D. M., Schulman, J., Amodei, D., and McCandlish, S. Scaling laws for autoregressive generative modeling, 2020.
  13. 13.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  14. 14.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L. Training compute-optimal large language models, 2022.
  15. 15.Huber, P. J. Robust Estimation of a Location Parameter. The Annals of Mathematical Statistics, 35(1):73 – 101, 1964. doi: 10.1214/aoms/1177703732. URL https://doi.org/10.1214/aoms/1177703732.
  16. 16.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b, 2023.
  17. 17.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts, 2024.
  18. 18.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models, 2020.
  19. 19.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding, 2020.
  20. 20.Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L. Base layers: Simplifying training of large, sparse models, 2021.
  21. 21.Liu, Z. L., Dettmers, T., Lin, X. V., Stoyanov, V., and Li, X. Towards a unified view of sparse feed-forward network in pretraining large language model, 2023.
  22. 22.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization, 2019.
  23. 23.Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N. From sparse to soft mixtures of experts, 2023.
  24. 24.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pretraining. 2018a.
  25. 25.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2018b. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf.
  26. 26.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G. Scaling language models: Methods, analysis & insights from training gopher, 2022.
  27. 27.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023.
  28. 28.Roller, S., Sukhbaatar, S., Szlam, A., and Weston, J. Hash layers for large sparse models, 2021.
  29. 29.Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagne, R., Luccioni, A. S., Yvon, F., Galle, M., Tow, J., Rush, A. M., Biderman, S., Webson, A., Ammanamanchi, P. S., Wang, T., Sagot, B., Muennighoff, N., del Moral, A. V., Ruwase, O., Bawden, R., Bekman, S., McMillan-Major, A., Beltagy, I., Nguyen, H., Saulnier, L., Tan, S., Suarez, P. O., Sanh, V., Laurenc¸on, H., Jernite, Y., Launay, J., Mitchell, M., Raffel, C., Gokaslan, A., Simhi, A., Soroa, A., Aji, A. F., Alfassy, A., Rogers, A., Nitzav, A. K., Xu, C., Mou, C., Emezue, C., Klamm, C., Leong, C., van Strien, D., Adelani, D. I., Radev, D., Ponferrada, E. G., Levkovizh, E., Kim, E., Natan, E. B., Toni, F. D., Dupont, G., Kruszewski, G., Pistilli, G., Elsahar, H., Benyamina, H., Tran, H., Yu, I., Abdulmumin, I., Johnson, I., Gonzalez-Dios, I., de la Rosa, J., Chim, J., Dodge, J., Zhu, J., Chang, J., Frohberg, J., Tobing, J., Bhattacharjee, J., Almubarak, K., Chen, K., Lo, K., Werra, L. V., Weber, L., Phan, L., allal, L. B., Tanguy, L., Dey, M., Munoz, M. R., Masoud, M., Grandury, M., Saˇ sko, M., Huang, M., Coavoux, M., Singh, M., Jiang, M. T.-J., Vu, M. C., Jauhar, M. A., Ghaleb, M., Subramani, N., Kassner, N., Khamis, N., Nguyen, O., Espejel, O., de Gibert, O., Villegas, P., Henderson, P., Colombo, P., Amuok, P., Lhoest, Q., Harliman, R., Bommasani, R., Lopez, R. L., Ribeiro, R., Osei, S., Pyysalo, S., Nagel, S., Bose, S., Muhammad, S. H., Sharma, S., Longpre, S., Nikpoor, S., Silberberg, S., Pai, S., Zink, S., Torrent, T. T., Schick, T., Thrush, T., Danchev, V., Nikoulina, V., Laippala, V., Lepercq, V., Prabhu, V., Alyafeai, Z., Talat, Z., Raja, A., Heinzerling, B., Si, C., Tas¸ar, D. E., Salesky, E., Mielke, S. J., Lee, W. Y., Sharma, A., Santilli, A., Chaffin, A., Stiegler, A., Datta, D., Szczechla, E., Chhablani, G., Wang, H., Pandey, H., Strobelt, H., Fries, J. A., Rozen, J., Gao, L., Sutawika, L., Bari, M. S., Al-shaibani, M. S., Manica, M., Nayak, N., Teehan, R., Albanie, S., Shen, S., Ben-David, S., Bach, S. H., Kim, T., Bers, T., Fevry, T., Neeraj, T., Thakker, U., Raunak, V., Tang, X., Yong, Z.-X., Sun, Z., Brody, S., Uri, Y., Tojarieh, H., Roberts, A., Chung, H. W., Tae, J., Phang, J., Press, O., Li, C., Narayanan, D., Bourfoune, H., Casper, J., Rasley, J., Ryabinin, M., Mishra, M., Zhang, M., Shoeybi, M., Peyrounette, M., Patry, N., Tazi, N., Sanseviero, O., von Platen, P., Cornette, P., Lavallee, P. F., Lacroix, R., Rajbhandari, S., Gandhi, S., Smith, S., Requena, S., Patil, S., Dettmers, T., Baruwa, A., Singh, A., Cheveleva, A., Ligozat, A.-L., Subramonian, A., Nev´ eol, A., Lovering, C., Garrette, D., Tunuguntla, D., Reiter, E., Taktasheva, E., Voloshina, E., Bogdanov, E., Winata, G. I., Schoelkopf, H., Kalo, J.-C., Novikova, J., Forde, J. Z., Clive, J., Kasai, J., Kawamura, K., Hazan, L., Carpuat, M., Clinciu, M., Kim, N., Cheng, N., Serikov, O., Antverg, O., van der Wal, O., Zhang, R., Zhang, R., Gehrmann, S., Mirkin, S., Pais, S., Shavrina, T., Scialom, T., Yun, T., Limisiewicz, T., Rieser, V., Protasov, V., Mikhailov, V., Pruksachatkun, Y., Belinkov, Y., Bamberger, Z., Kasner, Z., Rueda, A., Pestana, A., Feizpour, A., Khan, A., Faranak, A., Santos, A., Hevia, A., Unldreaj, A., Aghagol, A., Abdollahi, A., Tammour, A., HajiHosseini, A., Behroozi, B., Ajibade, B., Saxena, B., Ferrandis, C. M., McDuff, D., Contractor, D., Lansky, D., David, D., Kiela, D., Nguyen, D. A., Tan, E., Baylor, E., Ozoani, E., Mirza, F., Ononiwu, F., Rezanejad, H., Jones, H., Bhattacharya, I., Solaiman, I., Sedenko, I., Nejadgholi, I., Passmore, J., Seltzer, J., Sanz, J. B., Dutra, L., Samagaio, M., Elbadri, M., Mieskes, M., Gerchick, M., Akinlolu, M., McKenna, M., Qiu, M., Ghauri, M., Burynok, M., Abrar, N., Rajani, N., Elkott, N., Fahmy, N., Samuel, O., An, R., Kromann, R., Hao, R., Alizadeh, S., Shubber, S., Wang, S., Roy, S., Viguier, S., Le, T., Oyebade, T., Le, T., Yang, Y., Nguyen, Z., Kashyap, A. R., Palasciano, A., Callahan, A., Shukla, A., Miranda-Escalada, A., Singh, A., Beilharz, B., Wang, B., Brito, C., Zhou, C., Jain, C., Xu, C., Fourrier, C., Perin˜an, D. L., Molano, D., Yu, D., Manjavacas, E., Barth, F., Fuhrimann, F., Altay, G., Bayrak, G., Burns, G., Vrabec, H. U., Bello, I., Dash, I., Kang, J., Giorgi, J., Golde, J., Posada, J. D., Sivaraman, K. R., Bulchandani, L., Liu, L., Shinzato, L., de Bykhovetz, M. H., Takeuchi, M., Pamies, M., Castillo, M. A., Nezhurina, M., Sanger, M., Samwald, M., Cullan, M., Weinberg, M., Wolf, M. D., Mihaljcic, M., Liu, M., Freidank, M., Kang, M., Seelam, N., Dahlberg, N., Broad, N. M., Muellner, N., Fung, P., Haller, P., Chandrasekhar, R., Eisenberg, R., Martin, R., Canalli, R., Su, R., Su, R., Cahyawijaya, S., Garda, S., Deshmukh, S. S., Mishra, S., Kiblawi, S., Ott, S., Sang-aroonsiri, S., Kumar, S., Schweter, S., Bharati, S., Laud, T., Gigant, T., Kainuma, T., Kusa, W., Labrak, Y., Bajaj, Y. S., Venkatraman, Y., Xu, Y., Xu, Y., Xu, Y., Tan, Z., Xie, Z., Ye, Z., Bras, M., Belkada, Y., and Wolf, T. Bloom: A 176b-parameter open-access multilingual language model, 2023.
  30. 30.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017.
  31. 31.Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B. Mesh-tensorflow: Deep learning for supercomputers, 2018.
  32. 32.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023a.
  33. 33.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models, 2023b.
  34. 34.Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E. A survey on multimodal large language models, 2023.
  35. 35.Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J. Mixture-of-experts with expert choice routing, 2022.
  36. 36.Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A., Lu, Y., Chen, Z., Le, Q., Cui, C., Laundon, J., and Dean, J. Brainformers: Trading simplicity for efficiency, 2023.

Citation

MLA
Krajewski, J., et al. “Scaling Laws for Fine-Grained Mixture of Experts”. arXiv, 2024, http://arxiv.org/abs/2402.07871v1.
APA
Krajewski, J., Ludziejewski, J., Adamczewski, K., Pióro, M., Krutul, M., Antoniak, S., Ciebiera, K., Król, K., Odrzygóźdź, T., Sankowski, P., Cygan, M., & Jaszczur, S. (2024). Scaling Laws for Fine-Grained Mixture of Experts. arXiv. http://arxiv.org/abs/2402.07871v1
Chicago
Krajewski, J., J. Ludziejewski, K. Adamczewski, et al. 2024. “Scaling Laws for Fine-Grained Mixture of Experts”. arXiv. http://arxiv.org/abs/2402.07871v1.
Harvard
Krajewski, J. et al. (2024) “Scaling Laws for Fine-Grained Mixture of Experts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.07871v1.
Vancouver
1. Krajewski J, Ludziejewski J, Adamczewski K, et al (2024) Scaling Laws for Fine-Grained Mixture of Experts. arXiv

BibTeX

@article{krajewski2024scaling,
  title = {Scaling Laws for Fine-Grained Mixture of Experts},
  author = {Krajewski, Jakub and Ludziejewski, Jan and Adamczewski, Kamil and Pióro, Maciej and Krutul, Michał and Antoniak, Szymon and Ciebiera, Kamil and Król, Krystian and Odrzygóźdź, Tomasz and Sankowski, Piotr and Cygan, Marek and Jaszczur, Sebastian},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.07871v1},
  eprint = {2402.07871}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/