Timer: Generative Pre-trained Transformers Are Large Time Series Models
Yong LiuHaoran ZhangChenyu LiXiangdong HuangJianmin WangMingsheng Long
Presents a billion-point pre-trained GPT-style model that unifies diverse time series forecasting, imputation, and anomaly detection into next-token prediction to deliver strong few-shot and zero-shot performance across heterogeneous domains.
Real-world time series analysis—encompassing forecasting, data imputation, and anomaly detection—is essential across industries such as energy, transportation, and health. However, conventional specialized deep learning models suffer severe performance degradation when training data is scarce, and existing non-autoregressive architectures lack flexibility across varied sequence lengths and tasks. While large language models have demonstrated that large-scale generative pre-training enables broad task versatility and few-shot generalization, the development of scalable, unified foundation models natively designed for time series has remained constrained by data heterogeneity and single-task design.
The article demonstrates the viability of large time series models by introducing Timer, a generative pre-trained decoder-only Transformer tailored for universal time series analysis. The researchers evaluate Timer's scaling behavior and performance across multiple tasks, particularly under severe data scarcity, and benchmark its zero-shot capabilities against concurrent large models.
To conduct this evaluation, the authors curated the Unified Time Series Dataset (UTSD), assembling up to 1 billion real-world time points across seven diverse operational domains. They introduced the Single-Series Sequence (S3) format to standardize heterogeneous multivariate series into uniform token sequences without requiring time alignment. Timer was then pre-trained using next-token autoregression and evaluated across extensive benchmarks for forecasting, span-based missing data imputation, and on-the-fly anomaly detection against current state-of-the-art models.
The findings confirm that generative pre-training significantly elevates time series modeling capabilities. First, fine-tuning a pre-trained Timer with only 1% to 5% of target training samples achieves accuracy competitive with or exceeding top small models trained on 100% of data. Second, scaling Timer from 1 million to 50 million parameters and expanding training corpora reduces forecasting errors by up to 40.3%, adhering to foundation model scaling laws. Third, Timer establishes state-of-the-art results across downstream tasks, outperforming leading baselines in 100% of tested 5%-sample imputation scenarios and achieving top average rankings in multi-dataset zero-shot forecasting evaluations.
These results demonstrate that organizations can replace fragmented, task-specific pipelines with a single foundation model architecture. By maintaining high performance in data-constrained environments, this approach substantially reduces the cost and timeline associated with collecting massive task-specific labeled data. Furthermore, Timer’s flexible autoregressive design handles variable input and output lengths naturally, minimizing operational complexity across diverse industrial applications.
Organizations and practitioners should consider shifting from training scenario-specific models from scratch toward deploying pre-trained generative time series foundation models, especially in data-scarce domains. Future development and deployment efforts should focus on expanding high-quality pre-training corpora beyond current limits and piloting domain-specific adaptations. Readers should note that current limitations include the lack of native support for classification and probabilistic forecasting, requiring cautious evaluation when deploying under scenarios that demand rigorous uncertainty quantification.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). PatchTST introduces sub-series patch tokenization and channel independence for Transformers, establishing key tokenization principles foundational to Timer's architecture.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). iTransformer demonstrates the effectiveness of inverting sequence tokens for attention-based temporal modeling, informing modern large time series model backbones.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). This paper highlights the structural pitfalls and architectural questions surrounding vanilla Transformers in long-term forecasting that motivate large-scale generative pre-training.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). TimesNet establishes a unified multi-task formulation spanning forecasting, imputation, and anomaly detection, directly preceding Timer's task-general paradigm.
- Paper: Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting, Yong Liu et al. (2022). Non-stationary Transformers address distribution shifts in temporal attention modeling, providing core techniques for handling non-stationary series in Transformer architectures.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). Autoformer introduces auto-correlation and temporal decomposition mechanisms that shaped modern Transformer-based time series modeling.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). Informer established foundational long-sequence Transformer mechanisms and benchmark protocols widely used in time series deep learning.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). This survey provides a comprehensive taxonomy of structural modifications and operational challenges in adapting Transformers to sequential time series data.
- Paper: Sundial: A Family of Highly Capable Time Series Foundation Models, Yong Liu 0007 et al. (2025). Sundial advances generative time series foundation models beyond deterministic GPT-style architectures by incorporating continuous flow matching for zero-shot probabilistic forecasting.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). MOIRAI extends universal time series foundation modeling through multi-patch-size projections and any-variate attention trained on the multi-domain LOTSA archive.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). This position paper conceptualizes the broader role of large autoregressive models and LLM principles in reshaping general-purpose time series analysis.
