Timer: Generative Pre-trained Transformers Are Large Time Series Models

Yong LiuHaoran ZhangChenyu LiXiangdong HuangJianmin WangMingsheng Long

article2024ICML286 citations

Presents a billion-point pre-trained GPT-style model that unifies diverse time series forecasting, imputation, and anomaly detection into next-token prediction to deliver strong few-shot and zero-shot performance across heterogeneous domains.

Listen

Real-world time series analysis—encompassing forecasting, data imputation, and anomaly detection—is essential across industries such as energy, transportation, and health. However, conventional specialized deep learning models suffer severe performance degradation when training data is scarce, and existing non-autoregressive architectures lack flexibility across varied sequence lengths and tasks. While large language models have demonstrated that large-scale generative pre-training enables broad task versatility and few-shot generalization, the development of scalable, unified foundation models natively designed for time series has remained constrained by data heterogeneity and single-task design.

The article demonstrates the viability of large time series models by introducing Timer, a generative pre-trained decoder-only Transformer tailored for universal time series analysis. The researchers evaluate Timer's scaling behavior and performance across multiple tasks, particularly under severe data scarcity, and benchmark its zero-shot capabilities against concurrent large models.

To conduct this evaluation, the authors curated the Unified Time Series Dataset (UTSD), assembling up to 1 billion real-world time points across seven diverse operational domains. They introduced the Single-Series Sequence (S3) format to standardize heterogeneous multivariate series into uniform token sequences without requiring time alignment. Timer was then pre-trained using next-token autoregression and evaluated across extensive benchmarks for forecasting, span-based missing data imputation, and on-the-fly anomaly detection against current state-of-the-art models.

The findings confirm that generative pre-training significantly elevates time series modeling capabilities. First, fine-tuning a pre-trained Timer with only 1% to 5% of target training samples achieves accuracy competitive with or exceeding top small models trained on 100% of data. Second, scaling Timer from 1 million to 50 million parameters and expanding training corpora reduces forecasting errors by up to 40.3%, adhering to foundation model scaling laws. Third, Timer establishes state-of-the-art results across downstream tasks, outperforming leading baselines in 100% of tested 5%-sample imputation scenarios and achieving top average rankings in multi-dataset zero-shot forecasting evaluations.

These results demonstrate that organizations can replace fragmented, task-specific pipelines with a single foundation model architecture. By maintaining high performance in data-constrained environments, this approach substantially reduces the cost and timeline associated with collecting massive task-specific labeled data. Furthermore, Timer’s flexible autoregressive design handles variable input and output lengths naturally, minimizing operational complexity across diverse industrial applications.

Organizations and practitioners should consider shifting from training scenario-specific models from scratch toward deploying pre-trained generative time series foundation models, especially in data-scarce domains. Future development and deployment efforts should focus on expanding high-quality pre-training corpora beyond current limits and piloting domain-specific adaptations. Readers should note that current limitations include the lack of native support for classification and probabilistic forecasting, requiring cautious evaluation when deploying under scenarios that demand rigorous uncertainty quantification.

arXiv: 2402.02368
Cover for Timer: Generative Pre-trained Transformers Are Large Time Series Models

Abstract

Deep learning has contributed remarkably to the advancement of time series analysis. Still, deep models can encounter performance bottlenecks in real-world data-scarce scenarios, which can be concealed due to the performance saturation with small models on current benchmarks. Meanwhile, large models have demonstrated great powers in these scenarios through large-scale pre-training. Continuous progress has been achieved with the emergence of large language models, exhibiting unprecedented abilities such as few-shot generalization, scalability, and task generality, which are however absent in small deep models. To change the status quo of training scenario-specific small models from scratch, this paper aims at the early development of large time series models (LTSM). During pre-training, we curate large-scale datasets with up to 1 billion time points, unify heterogeneous time series into single-series sequence (S3) format, and develop the GPT-style architecture toward LTSMs. To meet diverse application needs, we convert forecasting, imputation, and anomaly detection of time series into a unified generative task. The outcome of this study is a Time Series Transformer (Timer), which is generative pre-trained by next token prediction and adapted to various downstream tasks with promising capabilities as an LTSM. Code and datasets are available at: https://github.com/thuml/Large-Time-Series-Model.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Unsupervised Pre-training on Sequences
  • 2.2. Large Time Series Models
  • 3. Approach
  • 3.1. Data
  • 3.2. Training Strategy
  • 3.3. Model Design
  • 4. Experiments
  • 4.1. Time Series Forecasting
  • 4.2. Imputation
  • 4.3. Anomaly Detection
  • 4.4. Scalability
  • 4.5. Model Analysis
  • 4.6. Evaluation of Large Time Series Models
  • 5. Conclusion and Future Work
  • Impact Statement
  • Acknowledgements
  • References
  • A. Unified Time Series Dataset
  • A.1. Datasets Details
  • A.2. Statistics
  • A.3. UTSD Composition Analysis
  • A.4. Experiments
  • B. Implementation Details
  • B.1. Pre-training
  • B.2. Downstream Tasks
  • C. Full Results
  • C.1. Time Series Forecasting
  • C.2. Imputation
  • C.3. Anomaly Detection
  • C.4. Scalability
  • C.5. Zero-shot forecasting
  • D. Showcase
  • E. Limitations
  • F. Societal Impacts

Knowls

  1. Knowl 1 — Unified Time Series Dataset for scalable pre-training

    data/table

    The authors curate the Unified Time Series Dataset (UTSD) for large-scale time-series pre-training. Its core collection covers energy, environment, health, nature, transport, Internet of Things, and web data, with heterogeneous sampling frequencies, variate counts, lengths, stationarity, forecastability, and regular or irregular timestamps. Missing values are linearly interpolated, and datasets are stored in a unified ARROW/parquet-based format with timestamps and metadata.

    UTSD is released as nested capacity levels—UTSD-1G, UTSD-2G, UTSD-4G, and UTSD-12G—whose sizes reach approximately 1 billion time points. Each larger level adds pattern diversity and increasingly difficult series while keeping domains approximately balanced. Dataset difficulty is characterized using a length-weighted Augmented Dickey–Fuller statistic and a length-weighted forecastability score; based on the ADF statistic, datasets are labeled Easy when the statistic is below −15.00-15.00, Medium when it lies in [−15.00,−5.00)[-15.00,-5.00), and Hard when it is at least −5.00-5.00. The authors additionally apply the same curation procedure to larger corpora, including LOTSA, to train Timer variants at approximately 1B, 16B, and 28B pre-training time points.

  2. Knowl 2 — Single-series sequence format for heterogeneous time series

    model/method

    The single-series sequence (S3) format converts heterogeneous multivariate, univariate, and irregularly sampled time series into a common pre-training representation. Each variate is split into a 9:1 training/validation partition; statistics computed on the training portion are used to normalize the entire variate. The normalized variates are merged into a pool of single-variate series, and fixed-length windows are sampled uniformly from this pool.

    S3 preserves temporal variation while removing the need for time alignment across variates or datasets. Consequently, a mini-batch can contain windows from different datasets and time periods, unlike channel-independence procedures that flatten aligned variates from one dataset into a common batch. The resulting fixed-context single-variate windows are treated as time-series sentences for generative pre-training.

  3. Knowl 3 — Timer’s decoder-only generative architecture

    equation

    Timer tokenizes a normalized S3 sequence X=(x_1,dots,x_{NS})\in\mathbb{R}^{NS} into NN consecutive tokens, where SS is the number of time points per token and xt∈Rx_t\in\mathbb{R} is one normalized time point:

    si=(x(i−1)S+1,…,xiS)∈RS,i=1,…,N.s_i=(x_{(i-1)S+1},\dots,x_{iS})\in\mathbb{R}^{S},\qquad i=1,\dots,N.

    Each token is embedded by a decoder-only Transformer with hidden dimension DD and LL causal Transformer layers. With We,Wd∈RD×SW_e,W_d\in\mathbb{R}^{D\times S} as the input and output projections, hi0=Wesi+TEih_i^0=W_es_i+\mathrm{TE}_i combines the token embedding with an optional timestamp embedding TEi∈RD\mathrm{TE}_i\in\mathbb{R}^{D}. The causal Transformer produces hidden states hiLh_i^L using only tokens at positions up to ii, and predicts the next token as s^i+1=WdThiL∈RS\hat{s}_{i+1}=W_d^{\mathsf T}h_i^L\in\mathbb{R}^{S}.

    The pre-training loss is next-token mean squared error over all supervised positions:

    LMSE=1NS∑i=2N∥si−s^i∥22.\mathcal{L}_{\mathrm{MSE}}=\frac{1}{NS}\sum_{i=2}^{N}\lVert s_i-\hat{s}_i\rVert_2^2.

    Thus, every token position supplies an independent next-token training signal. Unlike a fixed-length, non-autoregressive forecaster, Timer can iteratively generate additional tokens and can use different context and forecast lengths at inference time.

  4. Knowl 4 — Unified generative treatment of forecasting, imputation, and anomaly detection

    model/method

    Timer adapts the same next-token generation mechanism to three tasks. For forecasting, the observed lookback tokens are supplied to the decoder and future tokens are generated autoregressively; predicted tokens are concatenated to the input and generation is repeated until the requested horizon is reached.

    For segment-level imputation, a series is divided into eight segments of length 24, and randomly selected segments are replaced by mask tokens while the first segment remains observed. During adaptation, Timer uses denoising autoencoding: observed segments and mask tokens form the input, and the model generates the missing segments. The generated segments are assembled with the unmasked observations.

    For anomaly detection, Timer uses a predictive rather than reconstructive protocol. Observed segments are used to predict the next segment, and the mean squared error between the predicted and received segment is treated as an anomaly score. In the authors’ protocol, the model uses seven lookback segments of length 96 to predict a 96-point standard segment; test segments whose scores exceed a selected confidence quantile are labeled potential anomalies. This permits segment-level, on-the-fly detection without first reconstructing an entire observation window.

  5. Knowl 5 — Few-shot forecasting benefit from UTSD pre-training

    empirical result

    Timer is evaluated on ETT, ECL, Traffic, Weather, and PEMS forecasting benchmarks using a 672-point lookback and a 96-point forecast. The pre-trained model uses UTSD-12G, segment length S=96S=96, and 15 pre-training tokens; downstream forecasting uses seven 96-point tokens for the 672-point lookback. All downstream datasets are excluded from pre-training.

    Fine-tuned Timer remains competitive with state-of-the-art small forecasters trained on all available target samples despite using only 1% of ETTh1, 5% of Traffic, 3% of PEMS03, and 25% of PEMS04 training samples. In another comparison, the performance of a randomly initialized Timer trained on all samples can be reached by the pre-trained model with only 2% of ETTh1, 5% of ECL, 1% of Weather, and 4% of PEMS03 samples.

    With all target samples available, pre-training still reduces MSE on several datasets: Weather improves from 0.1650.165 to 0.1540.154, PEMS03 from 0.1260.126 to 0.1180.118, and PEMS04 from 0.1250.125 to 0.1070.107. The results show that the transferred representation is useful both when target data are severely scarce and when the complete downstream training set is available.

  6. Knowl 6 — Segment-level imputation performance under data scarcity

    empirical result

    The authors evaluate imputation on 11 forecasting datasets, using four segment-mask ratios—12.5%, 25%, 37.5%, and 50%—and compare Timer with TimesNet. At 5%, 20%, and 100% of the downstream training samples, pre-trained Timer outperforms TimesNet in respectively 44/44 (100.0%), 38/44 (86.4%), and 25/44 (56.8%) dataset–mask-ratio scenarios.

    At 5% downstream data, pre-training reduces average imputation MSE on every dataset, with reductions ranging from 1.85% on ETTh2 to 15.49% on PEMS04. Other notable average reductions are 13.15% on ETTm1, 13.39% on Traffic, 14.10% on PEMS03, 13.69% on PEMS07, and 14.01% on PEMS08. The benefit generally decreases as more downstream samples become available, but remains positive in many settings at 20% and 100% data.

  7. Knowl 7 — Predictive anomaly detection on the UCR archive

    empirical result

    Timer is evaluated on all 250 tasks in the UCR Anomaly Archive, where each task supplies one normal training series and requires locating an anomalous interval in a test series. The model predicts future segments from observed segments, ranks test segments by prediction MSE, and reports the quantile containing the first detected anomaly.

    Against TimesNet and Anomaly Transformer, the reported numbers of detected anomalies at the 1%, 3%, and 10% confidence quantiles are respectively 51, 110, and 172 for Timer; 31, 61, and 109 for TimesNet; and 51, 98, and 129 for Anomaly Transformer. Pre-training also improves Timer over training from scratch across the 250 tasks: the number of detections within the 3% and 10% quantiles increases from 99 to 110 and from 146 to 172, while the average detection quantile decreases from 14.3% to 12.5%.

  8. Knowl 8 — Scaling model capacity and pre-training data improves forecasting

    empirical result

    Timer exhibits improved downstream forecasting as both parameter count and pre-training data increase. With UTSD-4G fixed, increasing the number of layers at hidden dimension D=256D=256 raises the model from approximately 1M to 4M parameters and reduces MSE by averages of 14.7% and 20.6% in the 5% and 20% downstream-data settings. With six layers fixed, increasing the hidden dimension from 256 to 1024 raises the model from approximately 3M to 50M parameters and yields further average MSE reductions of 25.1% and 18.2% in those settings.

    With model size fixed at eight layers and D=1024D=1024, increasing the pre-training corpus from UTSD-1G to UTSD-12G reduces MSE from 0.1490.149 to 0.1380.138 with 5% target data and from 0.1290.129 to 0.1230.123 with 20% target data. Across the combined scaling experiments, the reported errors decrease from 0.2310.231 to 0.1380.138 in the 5% setting and from 0.1940.194 to 0.1230.123 in the 20% setting, with the latter values outperforming the cited full-data multivariate baseline MSE of 0.1390.139 on PEMS.

  9. Knowl 9 — Decoder-only Transformers provide stronger transfer and length flexibility

    empirical result

    Under matched parameter and layer configurations, the authors compare decoder-only Timer with an encoder-only Transformer, LSTM, TiDE, and TCN. Transformer backbones scale more favorably than the MLP-based and CNN-based alternatives on heterogeneous pre-training data. Although the encoder-only Transformer can obtain lower pre-training loss and can be better than a decoder-only model trained from scratch at an extreme 1% target-data level, the pre-trained decoder-only model generalizes better in most downstream scenarios.

    For example, averaged over PEMS datasets, the MSEs of the encoder-only and decoder-only models pre-trained on UTSD-12G are 0.2460.246 and 0.1800.180 with 1% target data, 0.1970.197 and 0.1380.138 with 5%, and 0.1640.164 and 0.1260.126 with 20%. Causal token-wise supervision also lets one Timer operate with multiple lookback lengths from 288 to 672 points. When a single 672-lookback, 96-forecast model is rolled to forecast horizons from 96 to 480 points, Timer accumulates less error and performs better than the compared encoder-only PatchTST model.

  10. Knowl 10 — Initial zero-shot benchmark for large time-series models

    empirical result

    The authors establish a zero-shot forecasting benchmark on seven datasets excluded from the pre-training corpora. Each model predicts 96 future points for every test window, and performance is measured by MSE. The evaluated Timer variants use progressively larger pre-training corpora: Timer-1B is trained on UTSD, Timer-16B additionally uses Buildings900K, and Timer-28B additionally uses LOTSA.

    Among the evaluated model families, the reported average ranks are 1.571 for Timer, 2.286 for Moirai, 4.429 for MOMENT, 2.250 for TimesFM, 5.500 for single-trajectory Chronos, and 3.250 for Chronos using 20 sampled trajectories; lower rank is better. Timer, Moirai, and TimesFM are the strongest families in this benchmark. However, increasing pre-training scale does not produce a uniformly better zero-shot model on every dataset, and several models fail on some multi-step forecasts because of error accumulation. The results therefore demonstrate promising but still incomplete zero-shot generalization.

  11. Knowl 11 — Current limitations of Timer and its data infrastructure

    limitation

    The authors identify several unresolved limitations. Even UTSD-12G is substantially smaller than contemporaneous corpora claiming tens or hundreds of billions of time points, so continued expansion of high-quality, hierarchically organized time-series data is needed. Timer’s unified generative formulation does not yet cover time-series classification, probabilistic forecasting, or a specialized treatment of multiple variables.

    The model also lacks mature zero-shot generalization, in-context learning, and multimodal abilities. The paper presents the work as an early large-time-series-model development rather than a complete foundation model, with longer-context forecasting, probabilistic outputs, broader task coverage, and stronger zero-shot transfer left for future work.

Coverage note — Detailed per-dataset UTSD catalogues, all 250 anomaly quantiles, and full appendix-level forecasting and imputation tables were omitted because their aggregate results and governing protocols are represented without reproducing very large redundant tables.

References

  1. 1.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., et al. Gluonts: Probabilistic and neural time series modeling in python. Journal of Machine Learning Research, 21(116): 1–6, 2020.
  2. 2.Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024.
  3. 3.Bai, S., Kolter, J. Z., and Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018.
  4. 4.Bao, H., Dong, L., Piao, S., and Wei, F. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
  5. 5.Bengio, Y., Ducharme, R., and Vincent, P. A neural probabilistic language model. Advances in neural information processing systems, 13, 2000.
  6. 6.Bergmeir, C., Bui, Q., de Nijs, F., and Stuckey, P. Residential Power and Battery Data, 2023. URL https://doi.org/10.5281/zenodo.8219786.
  7. 7.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  8. 8.Box, G. Box and jenkins: time series analysis, forecasting and control. In A Very British Affair: Six Britons and the Development of Time Series Analysis During the 20th Century, pp. 161–215. Springer, 2013.
  9. 9.Box, G. E., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M. Time series analysis: forecasting and control. John Wiley & Sons, 2015.
  10. 10.Breunig, M. M., Kriegel, H.-P., Ng, R. T., and Sander, J. Lof: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data, pp. 93–104, 2000.
  11. 11.CDC. Flu portal dashboard, 2017. URL https://gis.cdc.gov/grasp/fluview/fluportaldashboard.html. Accessed: [insert date of access].
  12. 12.Chang, C., Peng, W.-C., and Chen, T.-F. Llm4ts: Two-stage fine-tuning for time-series forecasting with pre-trained llms. arXiv preprint arXiv:2308.08469, 2023.
  13. 13.Chen, S. Beijing Multi-Site Air Quality. UCI Machine Learning Repository, 2019. DOI: https://doi.org/10.24432/C5RK5G.
  14. 14.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  15. 15.Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F. Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers. arXiv preprint arXiv:2212.10559, 2022.
  16. 16.Das, A., Kong, W., Leach, A., Sen, R., and Yu, R. Long-term forecasting with tide: Time-series dense encoder. arXiv preprint arXiv:2304.08424, 2023a.
  17. 17.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting. arXiv preprint arXiv:2310.10688, 2023b.
  18. 18.Dau, H. A., Bagnall, A., Kamgar, K., Yeh, C.-C. M., Zhu, Y., Gharghabi, S., Ratanamahatana, C. A., and Keogh, E. The ucr time series archive. IEEE/CAA Journal of Automatica Sinica, 6(6):1293–1305, 2019.
  19. 19.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  20. 20.Dong, J., Wu, H., Zhang, H., Zhang, L., Wang, J., and Long, M. Simmtm: A simple pre-training framework for masked time-series modeling. arXiv preprint arXiv:2302.00861, 2023.
  21. 21.Dooley, S., Khurana, G. S., Mohapatra, C., Naidu, S., and White, C. Forecastpfn: Synthetically-trained zero-shot forecasting. arXiv preprint arXiv:2311.01933, 2023.
  22. 22.Elliott, G., Rothenberg, T. J., and Stock, J. H. Efficient tests for an autoregressive unit root. Econometrica, 1996.
  23. 23.Emami, P., Sahu, A., and Graf, P. Buildingsbench: A large-scale dataset of 900k buildings and benchmark for short-term load forecasting. Advances in Neural Information Processing Systems, 2023.
  24. 24.Friedman, M. The interpolation of time series by related series. Journal of the American Statistical Association, 57(300):729–757, 1962.
  25. 25.Garza, A. and Mergenthaler-Canseco, M. Timegpt-1. arXiv preprint arXiv:2310.03589, 2023.
  26. 26.Godahewa, R., Bergmeir, C., Webb, G. I., Hyndman, R. J., and Montero-Manso, P. Monash time series forecasting archive. arXiv preprint arXiv:2105.06643, 2021.
  27. 27.Goerg, G. Forecastable component analysis. In International conference on machine learning, pp. 64–72. PMLR, 2013.
  28. 28.Goswami, M., Szafer, K., Choudhry, A., Cai, Y., Li, S., and Dubrawski, A. Moment: A family of open time-series foundation models. arXiv preprint arXiv:2402.03885, 2024.
  29. 29.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  30. 30.Jiang, J., Han, C., Jiang, W., Zhao, W. X., and Wang, J. Libcity: A unified library towards efficient and comprehensive urban spatial-temporal prediction. arXiv preprint arXiv:2304.14343, 2023.
  31. 31.Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., et al. Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728, 2023.
  32. 32.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  33. 33.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR, 2015. URL http://arxiv.org/abs/1412.6980.
  34. 34.Liu, M., Zeng, A., Chen, M., Xu, Z., Lai, Q., Ma, L., and Xu, Q. Scinet: time series modeling and forecasting with sample convolution and interaction. NeurIPS, 2022.
  35. 35.Liu, X., Xia, Y., Liang, Y., Hu, J., Wang, Y., Bai, L., Huang, C., Liu, Z., Hooi, B., and Zimmermann, R. Largest: A benchmark dataset for large-scale traffic forecasting. In Advances in Neural Information Processing Systems, 2023a.
  36. 36.Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M. itransformer: Inverted transformers are effective for time series forecasting. arXiv preprint arXiv:2310.06625, 2023b.
  37. 37.Liu, Y., Qin, G., Huang, X., Wang, J., and Long, M. Autotimes: Autoregressive time series forecasters via large language models. arXiv preprint arXiv:2402.02370, 2024.
  38. 38.Mancuso, P., Piccialli, V., and Sudoso, A. M. A machine learning approach for forecasting hierarchical time series. Expert Systems with Applications, 182:115102, 2021.
  39. 39.Mouatadid, S., Orenstein, P., Flaspohler, G., Oprescu, M., Cohen, J., Wang, F., Knight, S., Geogdzhayeva, M., Levang, S., Fraenkel, E., et al. Subseasonalclimateusa: A dataset for subseasonal forecasting and benchmarking. Advances in Neural Information Processing Systems, 36, 2024.
  40. 40.Nguyen, T., Jewik, J., Bansal, H., Sharma, P., and Grover, A. Climatelearn: Benchmarking machine learning for weather and climate modeling. Advances in Neural Information Processing Systems, 36, 2024.
  41. 41.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. arXiv preprint arXiv:2211.14730, 2022.
  42. 42.OpenAI, R. Gpt-4 technical report. arxiv 2303.08774. View in Article, 2:13, 2023.
  43. 43.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raïson, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
  44. 44.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. Improving language understanding by generative pre-training. 2018.
  45. 45.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  46. 46.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  47. 47.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  48. 48.Rasul, K., Ashok, A., Williams, A. R., Khorasani, A., Adamopoulos, G., Bhagwatkar, R., Bilos, M., Ghonia, H., Hassen, N. V., Schneider, A., et al. Lag-llama: Towards foundation models for time series forecasting. arXiv preprint arXiv:2310.08278, 2023.
  49. 49.Tan, C. W., Bergmeir, C., Petitjean, F., and Webb, G. I. Time series extrinsic regression: Predicting numeric values from time series data. Data Mining and Knowledge Discovery, 35:1032–1060, 2021.
  50. 50.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  51. 51.van Panhuis, W. G., Cross, A., and Burke, D. S. Project tycho 2.0: a repository to improve the integration and reuse of data for global population health. Journal of the American Medical Informatics Association, 25(12):1608–1617, 2018.
  52. 52.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  53. 53.Wang, T., Roberts, A., Hesslow, D., Le Scao, T., Chung, H. W., Beltagy, I., Launay, J., and Raffel, C. What language model architecture and pretraining objective works best for zero-shot generalization? In International Conference on Machine Learning, pp. 22964–22984. PMLR, 2022a.
  54. 54.Wang, Y., Han, Y., Wang, H., and Zhang, X. Contrast everything: A hierarchical contrastive framework for medical time-series. arXiv preprint arXiv:2310.14017, 2023a.
  55. 55.Wang, Z., Xu, X., Zhang, W., Trajcevski, G., Zhong, T., and Zhou, F. Learning latent seasonal-trend representations for time series forecasting. Advances in Neural Information Processing Systems, 35:38775–38787, 2022b.
  56. 56.Wang, Z., Wen, Q., Zhang, C., Sun, L., Von Krannichfeldt, L., and Wang, Y. Benchmarks and custom package for electrical load forecasting. arXiv preprint arXiv:2307.07191, 2023b.
  57. 57.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  58. 58.Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S. Cost: Contrastive learning of disentangled seasonal-trend representations for time series forecasting. arXiv preprint arXiv:2202.01575, 2022.
  59. 59.Woo, G., Liu, C., Kumar, A., and Sahoo, D. Pushing the limits of pre-training for time series forecasting in the cloudops domain. arXiv preprint arXiv:2310.05063, 2023.
  60. 60.Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series forecasting transformers. arXiv preprint arXiv:2402.02592, 2024.
  61. 61.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34:22419–22430, 2021.
  62. 62.Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. Timesnet: Temporal 2d-variation modeling for general time series analysis. arXiv preprint arXiv:2210.02186, 2022.
  63. 63.Wu, R. and Keogh, E. Current time series anomaly detection benchmarks are flawed and are creating the illusion of progress. IEEE Transactions on Knowledge and Data Engineering, 2021.
  64. 64.Xu, J., Wu, H., Wang, J., and Long, M. Anomaly transformer: Time series anomaly detection with association discrepancy. arXiv preprint arXiv:2110.02642, 2021.
  65. 65.Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021.
  66. 66.Yue, Z., Wang, Y., Duan, J., Yang, T., Huang, C., Tong, Y., and Xu, B. Ts2vec: Towards universal representation of time series. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 8980–8987, 2022.
  67. 67.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 11121–11128, 2023.
  68. 68.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, pp. 2114–2124, 2021.
  69. 69.Zhang, X., Zhao, Z., Tsiligkaridis, T., and Zitnik, M. Self-supervised contrastive pre-training for time series via time-frequency consistency. Advances in Neural Information Processing Systems, 35:3988–4003, 2022.
  70. 70.Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
  71. 71.Zheng, Y., Yi, X., Li, M., Li, R., Shan, Z., Chang, E., and Li, T. Forecasting fine-grained air quality based on big data. In Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, pp. 2267–2276, 2015.
  72. 72.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pp. 11106–11115, 2021.
  73. 73.Zhou, J., Lu, X., Xiao, Y., Su, J., Lyu, J., Ma, Y., and Dou, D. Sdwpf: A dataset for spatial dynamic wind power forecasting challenge at kdd cup 2022. arXiv preprint arXiv:2208.04360, 2022.
  74. 74.Zhou, T., Niu, P., Wang, X., Sun, L., and Jin, R. One fits all: Power general time series analysis by pretrained lm. arXiv preprint arXiv:2302.11939, 2023.

Citation

MLA
Liu, Y., et al. “Timer: Generative Pre-trained Transformers Are Large Time Series Models”. arXiv, 2024, http://arxiv.org/abs/2402.02368v3.
APA
Liu, Y., Zhang, H., Li, C., Huang, X., Wang, J., & Long, M. (2024). Timer: Generative Pre-trained Transformers Are Large Time Series Models. arXiv. http://arxiv.org/abs/2402.02368v3
Chicago
Liu, Y., H. Zhang, C. Li, X. Huang, J. Wang, and M. Long. 2024. “Timer: Generative Pre-trained Transformers Are Large Time Series Models”. arXiv. http://arxiv.org/abs/2402.02368v3.
Harvard
Liu, Y. et al. (2024) “Timer: Generative Pre-trained Transformers Are Large Time Series Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.02368v3.
Vancouver
1. Liu Y, Zhang H, Li C, Huang X, Wang J, Long M (2024) Timer: Generative Pre-trained Transformers Are Large Time Series Models. arXiv

BibTeX

@article{liu2024timer,
  title = {Timer: Generative Pre-trained Transformers Are Large Time Series Models},
  author = {Liu, Yong and Zhang, Haoran and Li, Chenyu and Huang, Xiangdong and Wang, Jianmin and Long, Mingsheng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.02368v3},
  eprint = {2402.02368}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/