Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

Xu LiuJuncheng LiuGerald WooTaha AksuYuxuan LiangRoger ZimmermannChenghao LiuJunnan LiSilvio SavareseCaiming Xiong

article2025ICML128 citations

Proposes a sparse mixture-of-experts time series foundation model that replaces rigid frequency-based groupings with dynamic token-level specialization, achieving superior zero-shot forecasting across 39 datasets while activating up to 65 times fewer parameters.

Listen

Forecasting across diverse real-world applications is transitioning toward universal foundation models capable of generating predictions without task-specific retraining. However, pretraining these generalist models is challenging because time series data are highly diverse and non-stationary. Prior approaches relied on human-defined heuristics, such as grouping data by recording frequency (e.g., hourly or monthly) using separate projection modules. The article demonstrates that frequency-based partitioning is fundamentally flawed because different frequencies can share identical temporal dynamics, while series with the exact same frequency often exhibit completely different patterns.

The main objective of the article is to design, implement, and evaluate MOIRAI-MOE, a time series foundation model that eliminates human-imposed frequency groupings. Instead, it delegates pattern recognition to a sparse mixture of experts (MoE) architecture that routes data dynamically at the individual token level.

To evaluate this approach, the researchers built an autoregressive model that segments time series into short patches and routes them through a sparse MoE Transformer. They introduced a data-driven routing mechanism that uses token clusters derived from a pretrained dense model to guide expert selection. The system was trained on a large multi-domain repository (LOTSA) using a next-token prediction objective. The evaluation spanned 39 diverse datasets, comprising an in-distribution benchmark of 29 datasets and an out-of-distribution, zero-shot evaluation across 10 datasets spanning energy, transport, nature, sales, and web operations.

The findings confirm that data-driven token specialization significantly outperforms traditional frequency-based architectures. First, MOIRAI-MOE delivered up to a 17% error reduction over its dense predecessor at equivalent active parameter sizes. Second, it surpassed leading competing foundation models while utilizing up to 65 times fewer activated parameters. Third, the cluster-guided gating mechanism proved consistently superior to standard randomly initialized gating. Finally, internal model analyses revealed that early network layers specialize in high-frequency, noisy variations, whereas deeper layers converge into shared, frequency-invariant representations that perform progressive temporal denoising.

These results demonstrate that sparse MoE architectures can dramatically lower the computational footprint and operational costs of time series forecasting without sacrificing predictive accuracy. By eliminating rigid frequency-specific layers, organizations can deploy a single, unified architecture across diverse operational domains. Furthermore, the approach reaches superior accuracy in 25,000 pretraining steps compared to 125,000 steps required by traditional architectures, substantially reducing compute requirements and training timelines.

For engineering and data science teams building or deploying forecasting services, the article recommends replacing heuristic dataset partitioning with sparse MoE routing, adopting cluster-guided expert gating, and utilizing patch-based tokenization. Teams facing tighter training budgets can replace half of the feed-forward layers with MoE modules to achieve a 31% reduction in training time with only a modest 5% drop in accuracy. Future development should focus on pruning underutilized inference experts and implementing model quantization to optimize autoregressive serving speeds.

Confidence in these findings is supported by consistent performance gains across 39 distinct benchmarks and 6 standard evaluation metrics. The primary operational limitation is the latency associated with autoregressive, sequential token generation during inference. Additionally, combining key-value caching with instance normalization remains an open technical challenge, warranting careful runtime profiling before deploying the model in strictly latency-critical environments.

No sufficiently relevant recommendations were found.

Cover for Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

Abstract

Existing pretrained models for time series forecasting typically operate under a unified framework using fixed parameters during inference. Yet different formulations have been proposed depending on the specifics of data variety and complexity, making them unsuitable for diverse datasets distribution shift. The presence of varying frequencies patterns as well as temporal distributions introduces complexities requiring the learning process accommodates yet remains robust through unseen variations. Such constraints make generalist approaches difficult while specialized variants handle granular aspects poorly considering broad coverage demands.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Time Series Token Construction
  • 3.2. Sparse Mixture-of-Experts Transformers
  • 3.2.1. GATING FUNCTION
  • 3.3. Pretraining Objective
  • 4. Experiments
  • 4.1. MOIRAI-MOE Setup
  • 4.2. Main Results
  • 4.3. Ablation Studies
  • 4.4. Model Analyses
  • 4.5. Efficiency Analyses
  • 5. Conclusion
  • Impact Statement
  • References
  • A. Experimental Details
  • A.1. In-distribution Forecasting Datasets and Full Performance Results
  • A.2. Zero-shot Forecasting Datasets and Full Performance Results
  • A.3. Summary of Methods
  • A.4. Settings of Methods
  • B. Additional Results
  • B.1. Comparison of MOIRAI and MOIRAI-MOE Pretraining Steps
  • B.2. Effects of the Number of MoE Layers
  • B.3. Effects of Patch Size
  • B.4. Effects of Masking Ratio
  • B.5. Expert Distributions of Different Gating Function
  • B.6. Visualization of Time Series Observations and Expert Allocations
  • C. Limitation
  • D. Visualization

Knowls

  1. Knowl 1 — MOIRAI-MOE replaces frequency-specific projections with token-routed experts

    model/method

    MOIRAI-MOE removes the frequency-specific input and output projection banks used by MOIRAI and instead uses shared projections and sparse mixture-of-experts (MoE) feed-forward layers in a decoder-only Transformer. Each Transformer layer retains causal self-attention, while its feed-forward network is replaced by a layer containing MM experts. For token tt at layer ll, the MoE output is

    yt(l)=∑i∈St(l)gt,i(l)Ei(l)(x~t(l)),\mathbf{y}_t^{(l)}=\sum_{i\in\mathcal{S}_t^{(l)}}g_{t,i}^{(l)}E_i^{(l)}(\tilde{\mathbf{x}}_t^{(l)}),

    where x~t(l)∈RD\tilde{\mathbf{x}}_t^{(l)}\in\mathbb{R}^D is the token’s hidden state after causal attention, Ei(l)E_i^{(l)} is expert ii’s feed-forward transformation, St(l)\mathcal{S}_t^{(l)} is the set of experts selected for that token, and gt,i(l)g_{t,i}^{(l)} is its normalized routing weight. The model uses K=2K=2 selected experts per token. Multivariate series are represented by flattening their variates into a sequence, allowing causal attention to model both within-variate and cross-variate dependencies.

  2. Knowl 2 — Pretrained token clusters provide the expert-routing signal

    model/method

    MOIRAI-MOE’s proposed router uses representation clusters from a pretrained dense MOIRAI model rather than learning a gating projection from random initialization. The dense model is first trained with single-patch input and output projections. Its attention-output representations are then extracted on LOTSA pretraining data, and layer-specific mini-batch kk-means clustering is used to produce 32 centroids per layer; the clustering data amount corresponds to 100 epochs. During MoE pretraining, each token’s Euclidean distances to the centroids at its layer are used as token-to-expert affinity scores, followed by sparse top-KK selection and softmax weighting. The paper reports that this clustering-based gate outperforms the tested linear-projection gates, both with and without a load-balancing loss, across the evaluated expert counts of 4, 8, 16, 32, and 64.

  3. Knowl 3 — Patch tokens and a reserved sequence portion support robust normalization

    model/method

    For a univariate series of length SS, MOIRAI-MOE forms non-overlapping patches of size PP, yielding N=⌈S/P⌉N=\lceil S/P\rceil patches in RN×P\mathbb{R}^{N\times P}. The patches are normalized to reduce distribution-shift effects and passed through one residual multilayer-perceptron input projection to produce Transformer tokens of dimension DD. To compute normalization statistics robustly without requiring a different-length causal-normalization subsequence for every prediction position, a masking ratio rr specifies the portion of the sequence used only for normalization and excluded from prediction loss. The reported configuration uses P=16P=16 and r=0.3r=0.3.

  4. Knowl 4 — Pretraining predicts a probability distribution for the next patch

    model/method

    MOIRAI-MOE uses autoregressive next-token prediction: given a context of cc preceding patch tokens, the Transformer and a shared output projection predict the parameters ϕ^\hat{\phi} of a probability distribution for the next patch. The training objective is negative log-likelihood, Lpred=−log⁡p(xt+1∣ϕ^)\mathcal{L}_{\mathrm{pred}}=-\log p(x_{t+1}\mid\hat{\phi}), where xt+1x_{t+1} is the next time-series patch and p(⋅∣ϕ^)p(\cdot\mid\hat{\phi}) is the predicted mixture distribution. This distributional objective supports probabilistic forecasts as well as point forecasts derived from the predicted distribution.

  5. Knowl 5 — Evaluated MoE configurations keep active size near dense MOIRAI

    experimental setup

    The evaluated small and base MOIRAI-MOE models use 6 and 12 Transformer layers, respectively, with hidden dimensions 384 and 768 and feed-forward dimensions 512 and 1,024. Each has 32 experts per MoE layer and activates two experts per token: MOIRAI-MOE-S has 11M active parameters and 117M total parameters, while MOIRAI-MOE-B has 86M active and 935M total parameters. These active sizes are close to the corresponding dense MOIRAI-S and MOIRAI-B models (14M and 91M parameters). The MoE models were trained on LOTSA using 16 A100 40-GB GPUs, batch size 1,024, and bfloat16; the small and base models received 50,000 and 250,000 training steps. Optimization used AdamW with learning rate 10−310^{-3}, weight decay 10−110^{-1}, β1=0.9\beta_1=0.9, and β2=0.98\beta_2=0.98, with 10,000 warmup steps followed by cosine annealing. A large MOIRAI-MOE model was not evaluated because of computational-resource requirements.

  6. Knowl 6 — MOIRAI-MOE improves aggregate performance on 29 in-distribution datasets

    empirical result

    On 29 Monash benchmark datasets whose training portions were included in LOTSA, with test portions held out for evaluation, MOIRAI-MOE achieved the best reported aggregated mean absolute error (MAE). The aggregate normalizes each dataset’s MAE by its seasonal-naive MAE and then takes the geometric mean; lower is better. MOIRAI-MOE-S scored 0.65, compared with 0.78 for MOIRAI-S, 0.71 for MOIRAI-B, and 0.70 for MOIRAI-L. MOIRAI-MOE-B scored 0.63, a further improvement over the small MoE model. The paper reports that MOIRAI-MOE-S improves on its dense small counterpart by 17%, and on the larger dense base and large models by 8% and 7%, respectively.

  7. Knowl 7 — Zero-shot results show gains in probabilistic and point forecasting

    empirical result

    On 10 zero-shot datasets spanning five domains and frequencies from 5-minute to weekly, MOIRAI-MOE-B achieved the best reported overall averages for continuous ranked probability score (CRPS) and mean absolute scaled error (MASE). The reported all-dataset averages were 0.478 CRPS and 0.651 MASE; averages excluding Electricity and Solar were 0.439 CRPS and 0.611 MASE. MOIRAI-MOE-S scored 0.497 and 0.670 on the all-dataset averages, and 0.450 and 0.620 on the corresponding non-leak averages. On the non-leak CRPS average, MOIRAI-MOE-B tied TimesFM at the reported precision; its non-leak MASE average was lower than the compared methods. The paper reports MOIRAI-MOE-S improvements over MOIRAI of 3%–14% in CRPS and 8%–16% in MASE across the MOIRAI sizes. Some comparison models had Electricity or Solar in their pretraining data, so the paper also reports averages excluding those two datasets.

  8. Knowl 8 — Ablations attribute most in-distribution gains to MoE specialization

    empirical result

    On the Monash aggregate, the model with MOIRAI’s multiple projections and masked-filling objective scored 0.78 MAE; changing only to next-token prediction improved this to 0.75, while using single projections with MoE and next-token prediction scored 0.65. This controlled comparison indicates that changing the objective alone produced a modest gain, while the main improvement came with expert-based token specialization. The next-token objective was also more efficient in the tested training-step comparison: 50,000 steps reached performance comparable to masked filling at 100,000 steps. Replacing only half of the feed-forward layers with MoE reduced pretraining time by 31% (6.53 versus 9.49 hours), but worsened aggregate MAE from 0.65 to 0.68.

  9. Knowl 9 — Representations and routing evolve toward frequency-invariant abstractions

    empirical result

    The paper’s embedding and routing analyses indicate that specialization follows token patterns rather than an assigned frequency category. For example, NN5 Daily and Traffic Hourly have different frequencies but similar patterns: their input embeddings are separated by MOIRAI’s frequency-specific projections, whereas MOIRAI-MOE blends their representations and gives them comparable expert-allocation distributions. Conversely, the distinct-pattern Covid Daily Deaths series is more effectively separated from NN5 Daily in MOIRAI-MOE embeddings, with a different expert distribution. The corresponding dataset MAEs changed from 5.37 to 4.04 for NN5 Daily, 0.020 to 0.013 for Traffic Hourly, and 124.32 to 119 for Covid Daily Deaths. Across frequency groups, shallow-layer expert allocations differ substantially, while allocations in the final layer become nearly identical. Routing is also more varied in shallow layers and more concentrated in deep layers; the authors interpret this pattern as progressive denoising and increasingly frequency-invariant representations.

  10. Knowl 10 — Inference time is comparable to dense MOIRAI, but autoregressive generation remains a limitation

    limitation

    In an inference-cost comparison using context length 512, 20 samples, and a one-patch forecast horizon (16 time steps), MOIRAI-MOE-S and MOIRAI-MOE-B took 273 and 370 seconds, close to MOIRAI-S at 264 seconds and MOIRAI-B at 358 seconds. Chronos-S, Chronos-B, and Chronos-L took 551, 1,177, and 2,780 seconds under the comparison setup. These timings were measured on a Monash subset and do not remove the general cost of autoregressive prediction: unlike MOIRAI’s simultaneous masked prediction, MOIRAI-MOE generates tokens sequentially. The paper identifies instance normalization as a further obstacle to key-value caching, because normalization statistics must be recalculated as each new token is generated and cached hidden states may consequently become invalid.

Coverage note — Detailed per-dataset score tables and secondary sweeps of patch size, masking ratio, and projection variants were omitted because they provide configuration-level evidence rather than additional central findings; the selected patch size and masking ratio are included in the method knowl.

References

  1. 1.Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024.
  2. 2.Cai, W., Jiang, J., Wang, F., Tang, J., Kim, S., and Huang, J. A survey on mixture of experts. arXiv preprint arXiv:2407.06204, 2024.
  3. 3.Cao, D., Jia, F., Arik, S. O., Pfister, T., Zheng, Y., Ye, W., and Liu, Y. Tempo: Prompt-based generative pre-trained transformer for time series forecasting. In International Conference on Learning Representations, 2024.
  4. 4.Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. In Association for Computational Linguistics, 2024.
  5. 5.Das, A., Kong, W., Leach, A., Mathur, S., Sen, R., and Yu, R. Long-term forecasting with tide: Time-series dense encoder. Transactions on Machine Learning Research, 2023.
  6. 6.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting. In International Conference on Machine Learning, 2024.
  7. 7.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  8. 8.Ekambaram, V., Jati, A., Nguyen, N. H., Dayama, P., Reddy, C., Gifford, W. M., and Kalagnanam, J. Ttms: Fast multi-level tiny time mixers for improved zero-shot and few-shot forecasting of multivariate time series. arXiv preprint arXiv:2401.03955, 2024.
  9. 9.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, pp. 1–39, 2022.
  10. 10.Gao, S., Koker, T., Queen, O., Hartvigsen, T., Tsiligkaridis, T., and Zitnik, M. Units: A unified multi-task time series model. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  11. 11.Godahewa, R., Bergmeir, C., Webb, G. I., Hyndman, R. J., and Montero-Manso, P. Monash time series forecasting archive. arXiv preprint arXiv:2105.06643, 2021.
  12. 12.Goswami, M., Szafer, K., Choudhry, A., Cai, Y., Li, S., and Dubrawski, A. Moment: A family of open time-series foundation models. arXiv preprint arXiv:2402.03885, 2024.
  13. 13.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. In Advances in Neural Information Processing Systems, 2023.
  14. 14.Han, X., Zheng, H., and Zhou, M. Card: Classification and regression diffusion models. Advances in Neural Information Processing Systems, 35:18100–18115, 2022.
  15. 15.Ismail, A. A., Arik, S. O., Yoon, J., Taly, A., Feizi, S., and Pfister, T. Interpretable mixture of experts. arXiv preprint arXiv:2206.02107, 2022.
  16. 16.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
  17. 17.Jiang, J., Han, C., Jiang, W., Zhao, W. X., and Wang, J. Towards efficient and comprehensive urban spatial-temporal prediction: A unified library and performance benchmark. arXiv e-prints, pp. arXiv–2304, 2023.
  18. 18.Komatsuzaki, A., Puigcerver, J., Lee-Thorp, J., Ruiz, C. R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N. Sparse upcycling: Training mixture-of-experts from dense checkpoints. arXiv preprint arXiv:2212.05055, 2022.
  19. 19.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, pp. 95–104, 2018.
  20. 20.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations, 2021.
  21. 21.Liang, Y., Wen, H., Nie, Y., Jiang, Y., Jin, M., Song, D., Pan, S., and Wen, Q. Foundation models for time series analysis: A tutorial and survey. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 6555–6565, 2024.
  22. 22.Liu, X., Hu, J., Li, Y., Diao, S., Liang, Y., Hooi, B., and Zimmermann, R. Unitime: A language-empowered unified model for cross-domain time series forecasting. In Proceedings of the ACM on Web Conference 2024, pp. 4095–4106, 2024a.
  23. 23.Liu, Y., Wu, H., Wang, J., and Long, M. Non-stationary transformers: Exploring the stationarity in time series forecasting. In Advances in Neural Information Processing Systems, pp. 9881–9893, 2022.
  24. 24.Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M. itransformer: Inverted transformers are effective for time series forecasting. In International Conference on Learning Representations, 2024b.
  25. 25.Liu, Y., Zhang, H., Li, C., Huang, X., Wang, J., and Long, M. Timer: Generative pre-trained transformers are large time series models. In Forty-first International Conference on Machine Learning, 2024c.
  26. 26.Ni, R., Lin, Z., Wang, S., and Fanti, G. Mixture-of-linear-experts for long-term time series forecasting. In International Conference on Artificial Intelligence and Statistics, pp. 4672–4680, 2024.
  27. 27.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In International Conference on Learning Representations, 2023.
  28. 28.Palaskar, S., Ekambaram, V., Jati, A., Gantayat, N., Saha, A., Nagar, S., Nguyen, N. H., Dayama, P., Sindhgatta, R., Mohapatra, P., et al. Automixer for improved multivariate time-series forecasting on business and it observability data. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 22962–22968, 2024.
  29. 29.Qwen. Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters”, February 2024. URL https://qwenlm.github.io/blog/qwen-moe/.
  30. 30.Rasul, K., Ashok, A., Williams, A. R., Khorasani, A., Adamopoulos, G., Bhagwatkar, R., Bilos, M., Ghonia, H., Hassen, N. V., Schneider, A., et al. Lag-llama: Towards foundation models for time series forecasting. arXiv preprint arXiv:2310.08278, 2023.
  31. 31.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations, 2017.
  32. 32.Shi, X., Wang, S., Nie, Y., Li, D., Ye, Z., Wen, Q., and Jin, M. Time-moe: Billion-scale time series foundation models with mixture of experts. arXiv preprint arXiv:2409.16040, 2024.
  33. 33.Trindade, A. Electricityloaddiagrams20112014. UCI Machine Learning Repository, 2015. DOI: https://doi.org/10.24432/C58C86.
  34. 34.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of machine learning research, (11), 2008.
  35. 35.Van Ness, M., Shen, H., Wang, H., Jin, X., Maddix, D. C., and Gopalswamy, K. Cross-frequency time series meta-forecasting. arXiv preprint arXiv:2302.02077, 2023.
  36. 36.Walmart Competition Admin, W. C. Walmart recruiting - store sales forecasting, 2014.
  37. 37.Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series forecasting transformers. In International Conference on Machine Learning, 2024.
  38. 38.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Advances in neural information processing systems, pp. 22419–22430, 2021.
  39. 39.Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. Timesnet: Temporal 2d-variation modeling for general time series analysis. In International Conference on Learning Representations, 2023.
  40. 40.Yao, J., Pan, W., Ghosh, S., and Doshi-Velez, F. Quality of uncertainty quantification for bayesian neural network inference. arXiv preprint arXiv:1906.09686, 2019.
  41. 41.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 11121–11128, 2023.
  42. 42.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, pp. 11106–11115, 2021.
  43. 43.Zhou, T., Niu, P., Sun, L., Jin, R., et al. One fits all: Power general time series analysis by pretrained lm. In Advances in neural information processing systems, pp. 43322–43355, 2023.
  44. 44.Zhu, T., Qu, X., Dong, D., Ruan, J., Tong, J., He, C., and Cheng, Y. Llama-moe: Building mixture-of-experts from llama with continual pre-training. arXiv preprint arXiv:2406.16554, 2024.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/