Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
Xu LiuJuncheng LiuGerald WooTaha AksuYuxuan LiangRoger ZimmermannChenghao LiuJunnan LiSilvio SavareseCaiming Xiong
Proposes a sparse mixture-of-experts time series foundation model that replaces rigid frequency-based groupings with dynamic token-level specialization, achieving superior zero-shot forecasting across 39 datasets while activating up to 65 times fewer parameters.
Forecasting across diverse real-world applications is transitioning toward universal foundation models capable of generating predictions without task-specific retraining. However, pretraining these generalist models is challenging because time series data are highly diverse and non-stationary. Prior approaches relied on human-defined heuristics, such as grouping data by recording frequency (e.g., hourly or monthly) using separate projection modules. The article demonstrates that frequency-based partitioning is fundamentally flawed because different frequencies can share identical temporal dynamics, while series with the exact same frequency often exhibit completely different patterns.
The main objective of the article is to design, implement, and evaluate MOIRAI-MOE, a time series foundation model that eliminates human-imposed frequency groupings. Instead, it delegates pattern recognition to a sparse mixture of experts (MoE) architecture that routes data dynamically at the individual token level.
To evaluate this approach, the researchers built an autoregressive model that segments time series into short patches and routes them through a sparse MoE Transformer. They introduced a data-driven routing mechanism that uses token clusters derived from a pretrained dense model to guide expert selection. The system was trained on a large multi-domain repository (LOTSA) using a next-token prediction objective. The evaluation spanned 39 diverse datasets, comprising an in-distribution benchmark of 29 datasets and an out-of-distribution, zero-shot evaluation across 10 datasets spanning energy, transport, nature, sales, and web operations.
The findings confirm that data-driven token specialization significantly outperforms traditional frequency-based architectures. First, MOIRAI-MOE delivered up to a 17% error reduction over its dense predecessor at equivalent active parameter sizes. Second, it surpassed leading competing foundation models while utilizing up to 65 times fewer activated parameters. Third, the cluster-guided gating mechanism proved consistently superior to standard randomly initialized gating. Finally, internal model analyses revealed that early network layers specialize in high-frequency, noisy variations, whereas deeper layers converge into shared, frequency-invariant representations that perform progressive temporal denoising.
These results demonstrate that sparse MoE architectures can dramatically lower the computational footprint and operational costs of time series forecasting without sacrificing predictive accuracy. By eliminating rigid frequency-specific layers, organizations can deploy a single, unified architecture across diverse operational domains. Furthermore, the approach reaches superior accuracy in 25,000 pretraining steps compared to 125,000 steps required by traditional architectures, substantially reducing compute requirements and training timelines.
For engineering and data science teams building or deploying forecasting services, the article recommends replacing heuristic dataset partitioning with sparse MoE routing, adopting cluster-guided expert gating, and utilizing patch-based tokenization. Teams facing tighter training budgets can replace half of the feed-forward layers with MoE modules to achieve a 31% reduction in training time with only a modest 5% drop in accuracy. Future development should focus on pruning underutilized inference experts and implementing model quantization to optimize autoregressive serving speeds.
Confidence in these findings is supported by consistent performance gains across 39 distinct benchmarks and 6 standard evaluation metrics. The primary operational limitation is the latency associated with autoregressive, sequential token generation during inference. Additionally, combining key-value caching with instance normalization remains an open technical challenge, warranting careful runtime profiling before deploying the model in strictly latency-critical environments.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). Read MOIRAI first to understand the dense foundation-model architecture and frequency-specific projections that MOIRAI-MoE replaces with token-level expert routing.
No sufficiently relevant recommendations were found.
