Sundial: A Family of Highly Capable Time Series Foundation Models
Yong LiuGuo QinZhiyuan ShiZhi ChenCaiyin YangXiangdong HuangJianmin WangMingsheng Long
Presents Sundial, a family of time series foundation models pre-trained on a one-trillion-point dataset using a flow-matching objective to enable fast, zero-shot probabilistic forecasting directly on continuous data without discrete tokenization.
Time series forecasting is essential for decision-making across industries such as energy, finance, and meteorology. However, real-world time series are inherently non-deterministic and highly diverse. Existing time series foundation models struggle with key limitations: deterministic models produce single, over-simplified predictions that fail to capture uncertainty, while alternative models either assume rigid parametric distributions or force continuous time series data into discrete language tokens, leading to coarse forecasts and computational bottlenecks.
The article introduces and evaluates Sundial, a family of native and flexible time series foundation models designed to deliver accurate zero-shot point and probabilistic forecasts. The objective is to demonstrate that integrating continuous-valued generative modeling with standard Transformer architectures enables flexible, high-capacity representation learning without resorting to discrete tokenization or restrictive distribution assumptions.
To achieve this, the authors designed TimeFlow Loss, a generative training objective based on continuous flow-matching that allows autoregressive Transformers to generate full predictive distributions conditioned on historical context. They also adapted the Transformer backbone with continuous patch tokenization, Rotary Position Embedding, FlashAttention, and Key-Value caching to accelerate processing. The resulting Sundial model family—ranging from 32 million to 444 million parameters—was pre-trained on TimeBench, a newly curated dataset comprising over one trillion time points from diverse domains including meteorology, healthcare, Internet of Things, and finance. The models were evaluated across leading benchmarks for zero-shot point forecasting and probabilistic forecasting against specialized supervised models and competing foundation models.
The evaluation produced four primary findings. First, Sundial achieved state-of-the-art results across standard point forecasting benchmarks; compared to the prior leading foundation model Time-MoE, Sundial reduced average Mean Squared Error by approximately 7.57% and Mean Absolute Error by 4.71% using fewer parameters. Second, on comprehensive probabilistic benchmarks like GIFT-Eval and the FEV leaderboard, Sundial achieved top-tier performance on unseen datasets, outperforming 70% of task-specific deep models and specialized statistical methods without needing task-specific training. Third, Sundial delivered significant computational efficiency, operating 35 times faster in inference than leading discrete tokenization models like Chronos and generating predictions within milliseconds. Fourth, architectural and training ablations confirmed that TimeFlow Loss outperformed diffusion-based and standard regression losses, while pre-training on larger datasets consistently lowered training loss by up to 15.38% and improved generalization.
These findings indicate that generative flow-matching provides a superior framework for pre-training continuous time series models, eliminating the trade-off between flexible uncertainty estimation and computational speed. For operational deployment, this allows organizations to generate reliable probabilistic forecasts and custom risk intervals in real time without the expensive retraining or fine-tuning pipelines typically required by supervised deep learning.
Organizations evaluating large-scale time series forecasting should consider adopting native generative foundation models like Sundial for zero-shot applications to streamline forecasting pipelines and improve risk modeling. When deploying the model, teams should leverage test-time calibration—adjusting the number of generated sample trajectories and sampling steps—to balance computational cost with the required level of statistical precision.
Confidence in these results is supported by rigorous evaluations on standard, multi-domain benchmarks excluding pre-training data. However, readers should note certain limitations: the current univariate pre-training framework does not explicitly model cross-variable correlations or external covariates, and performance on very high-frequency data is not fully guaranteed due to the predominance of low- and medium-frequency series in the training data. Further development is needed to support multivariate dependencies and enhanced multi-scale frequency sampling.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). Introduces MOIRAI and large-scale open pre-training archives, providing the foundational universal forecasting paradigm that Sundial scales and advances with continuous flow-matching.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). Establishes patch-based tokenization for Transformer forecasting, serving as the essential structural representation that Sundial builds upon for next-patch prediction.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Demonstrates how to adapt Transformer attention architectures natively to time series, laying critical groundwork for Sundial's minimal architectural adaptations.
- Paper: Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift, Taesung Kim et al. (2022). Provides the instance normalization techniques necessary to manage non-stationarity and distribution shifts in deep time-series models like Sundial.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Formulates continuous score-based generative modeling, establishing the mathematical foundations of continuous flow and diffusion dynamics underlying Sundial's TimeFlow loss.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). Pioneers deep Transformer architectures for long-term forecasting from the same research group, setting standard evaluation benchmarks and architectural baselines.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). Explores general time-series foundation backbones across multiple downstream tasks, establishing baseline capabilities that Sundial scales with generative pre-training.
No sufficiently relevant recommendations were found.
