Large Language Models Are Zero-Shot Time Series Forecasters

Nate GruverMarc FinziShikai QiuAndrew Gordon Wilson

article2023NeurIPS678 citations

Demonstrates that off-the-shelf large language models can match or exceed specialized domain models at zero-shot time series forecasting simply by framing numerical sequences as next-token prediction tasks.

Listen

Time series forecasting is essential across sectors such as finance, energy, healthcare, and supply chain management, yet building accurate deep learning models typically requires substantial domain-specific expertise, large curated datasets, and expensive task-specific training. Standard foundation models have transformed natural language and computer vision, but unified pretraining approaches for time series have historically lagged behind, often leaving simple classical methods like linear models or autoregressive moving averages to outperform complex neural networks.

The article demonstrates that off-the-shelf large language models (LLMs), such as GPT-3 and LLaMA-2, can serve as highly accurate zero-shot time series forecasters without any task-specific fine-tuning. By introducing a framework called LLMTIME, the authors show that framing numerical sequences as text allows LLMs to transfer their pattern recognition and probabilistic modeling capabilities directly to continuous time series data.

To evaluate this approach, the researchers encoded numerical series into character strings, applying systematic scaling and tokenization adjustments such as inserting spaces between digits to align with model vocabularies. They converted the models' discrete token probability distributions into continuous probability densities to properly quantify forecast uncertainty. The method was evaluated across synthetic benchmarks and standard real-world datasets spanning dozens of distinct domains—including the Darts, Monash, and Informer suites—and benchmarked against dedicated forecasting models, including neural architectures, Gaussian processes, and statistical baselines.

The investigation yielded several critical findings. First, LLMTIME achieved the best or second-best deterministic performance (measured by Mean Absolute Error) across all aggregated benchmark suites in a completely zero-shot setting, matching or exceeding models trained specifically on downstream data. Second, LLMs significantly outperformed specialized models in probabilistic metrics, generating superior continuous ranked probability scores and log likelihoods due to their inherent ability to capture multimodal, heavy-tailed distributions. Third, LLM forecasting exhibited high sample efficiency, outperforming traditional baselines even when context data was heavily restricted. Fourth, models naturally adapted to missing data denoted simply as text markers, avoiding complex manual imputation steps. Finally, model alignment interventions like reinforcement learning from human feedback degraded forecasting accuracy; specifically, GPT-4 and chat-tuned models exhibited worse uncertainty calibration and higher error rates compared to raw base models like GPT-3.

These findings suggest that organizations can bypass the substantial labor, computational cost, and infrastructure overhead associated with training specialized forecasting models, particularly when historical data is scarce. Because LLMs treat forecasting as text completion, organizations can also leverage flexible multimodal capabilities, such as incorporating contextual textual notes and querying the model in natural language to explain observed patterns.

Decision-makers considering LLM-based forecasting should use unaligned base models rather than chat-aligned models to preserve numerical calibration and accuracy. In the near term, teams should pilot zero-shot LLM forecasting on high-value univariate workflows alongside existing pipelines to benchmark latency and cost trade-offs. Further work is recommended to explore extended context windows and efficient multivariate handling before completely replacing dedicated systems.

Confidence in these findings is high for univariate and moderately long sequences, supported by tests on data collected after the models' pretraining cutoff dates to rule out memorization. However, caution is warranted when deploying LLMs on highly dimensional multivariate datasets or long-horizon forecasting tasks, where context window limitations and independent covariate modeling present computational constraints.

No sufficiently relevant recommendations were found.

Cover for Large Language Models Are Zero-Shot Time Series Forecasters

Abstract

By encoding time series as a string of numerical digits, we can frame time series forecasting as next-token prediction in text. Developing this approach, we find that large language models (LLMs) such as GPT-3 and LLaMA-2 can surprisingly zero-shot extrapolate time series at a level comparable to or exceeding the performance of purpose-built time series models trained on the downstream tasks. To facilitate this performance, we propose procedures for effectively tokenizing time series data and converting discrete distributions over tokens into highly flexible densities over continuous values. We argue the success of LLMs for time series stems from their ability to naturally represent multimodal distributions, in conjunction with biases for simplicity, and repetition, which align with the salient features in many time series, such as repeated seasonal trends. We also show how LLMs can naturally handle missing data without imputation through non-numerical text, accommodate textual side information, and answer questions to help explain predictions. While we find that increasing model size generally improves performance on time series, we show GPT-4 can perform worse than GPT-3 because of how it tokenizes numbers, and poor uncertainty calibration, which is likely the result of alignment interventions such as RLHF.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 LLMTime: Forecasting with Language Models
  • 4 Experiments
  • 5 Origins of Zero-Shot Performance
  • 6 Special Properties of LLMs
  • 7 Discussion
  • References
  • A Detailed method and hyperparameters
  • A.1 Input scaling
  • A.2 Validation tuning
  • A.3 Likelihood adjustment for GPT Models
  • B Addressing Memorization Concerns in GPT-3 Evaluations
  • C Benchmarking details and extended results
  • C.1 Darts datasets
  • C.2 Monash datasets
  • C.3 Informer datasets
  • C.4 Synthetic datasets
  • C.5 Darts full probabilistic results
  • C.6 Informer datasets with extended horizon
  • C.7 Monash dataset visualizations
  • C.8 Informer dataset visualizations
  • D Simplicity bias experiments
  • D.1 Full synthetic predictions
  • E GPT-4
  • F Multimodal Text Understanding of Time Series

Knowls

  1. Knowl 1 — LLMTIME turns forecasting into zero-shot text completion

    model/method

    LLMTIME applies a pretrained autoregressive language model to numeric time series without fine-tuning on the target dataset. It serializes observed values as text, conditions the language model on that history, and samples continuations as forecasts. Multiple sampled continuations provide either a point forecast, such as the pointwise median, or a probabilistic forecast, such as forecast quantiles. The method therefore uses the language model’s pretrained sequence-completion ability rather than fitting target-specific forecasting parameters.

  2. Knowl 2 — Digit-aware serialization and scaling are key to useful forecasts

    model/method

    LLMTIME encodes each value to a chosen fixed decimal precision, removes the decimal point, and separates successive time steps with commas. For GPT-style tokenizers that otherwise group digits into irregular chunks, spaces between individual digits encourage digit-by-digit tokenization; for LLaMA models, which already tokenize digits individually, those spaces are unnecessary and can hurt performance. Before serialization, a value ztz_t may be transformed to xt=(zt−b)/ax_t=(z_t-b)/a, where aa is the α\alpha-percentile of the shifted values and b=min⁡tzt−β(max⁡tzt−min⁡tzt)b=\min_t z_t-\beta(\max_t z_t-\min_t z_t). Here tt indexes time, and α\alpha and β\beta control scaling and offset. Scaling the α\alpha-percentile to one limits the number of digits used for typical values without scaling by the maximum, so the model can encounter and reproduce values with more digits. The authors select scaling parameters using validation-set likelihood; they also tune decimal precision and sampling settings for relevant evaluations.

  3. Knowl 3 — Digit likelihoods induce flexible continuous forecast densities

    model/method

    LLMTIME converts the language model’s discrete probabilities over digit strings into a continuous density by treating each representable numeric bin as uniform. With base BB and nn digits of precision, bin kk has width B−nB^{-n}; if the language model assigns probability pkp_k to that bin, its density there is pkBnp_k B^n. Thus, for a scaled value xx in bin kk, the density is px(x)=pkBnp_x(x)=p_k B^n. For an original value zz transformed by x=s(z)x=s(z), the corresponding density is pz(z)=pkBn∣ds(z)/dz∣p_z(z)=p_k B^n|ds(z)/dz|, so log⁡pz(z)=log⁡pk+nlog⁡B+log⁡∣ds(z)/dz∣\log p_z(z)=\log p_k+n\log B+\log|ds(z)/dz|. Here pkp_k is obtained from the autoregressive probabilities of the tokens encoding bin kk, and the Jacobian term accounts for preprocessing. The uniform-within-bin assumption yields a high-resolution continuous density from discrete tokens. In a separate one-dimensional density experiment, a decimal autoregressive model trained on only 200 samples from each distribution handled asymmetric, multimodal, and heavy-tailed distributions well under Wasserstein-distance evaluation. The tested distributions included an exponential, a uniform/student-tt mixture, and ARIMA forecast residuals; comparisons included Laplace fits, Gaussian mixtures trained by expectation-maximization, and logistic regression over fixed bins.

  4. Knowl 4 — Zero-shot LLMTIME is competitive on deterministic benchmarks

    empirical result

    On the Darts, Monash, and Informer forecasting benchmarks, LLMTIME using GPT-3 or LLaMA-2 70B achieved the best or second-best aggregate deterministic performance among the compared methods, despite having no trainable target-specific parameters. The evaluation used the pointwise median of 20 sampled forecasts and measured mean absolute error (MAE). The comparisons included statistical and neural forecasting methods; on the multivariate Informer datasets, LLMTIME forecast each covariate independently, and GPT-3 evaluation was omitted because of API cost. The Monash results used scaled MAE for aggregation. As a check against memorization of benchmark data, GPT-3 was also tested on three series recorded after its September 2021 training-data cutoff: Istanbul traffic, TSMC stock prices, and Turkey power demand. With the last 30 observations held out for each series, GPT-3 was competitive with or outperformed the time-series baselines on all three.

  5. Knowl 5 — LLMTIME performs strongly on probabilistic forecasts and limited data

    empirical result

    On the eight univariate Darts datasets, LLMTIME achieved better aggregate negative log likelihood per dimension (NLL/D) and continuous ranked probability score (CRPS) than the evaluated time-series baselines, including ARIMA, TCN, N-BEATS, N-HiTS, and a spectral-mixture Gaussian process; the paper reports the advantage on almost every individual dataset as well. It also outperformed PromptCast on aggregated Darts CRPS and MAE. In experiments restricting the amount of training data available to competing forecasting methods, LLMTIME retained high likelihood with only a small fraction of the data, while the other methods’ performance deteriorated more rapidly. Example forecasts showed extrapolation of trend and periodic components on AirPassengers, with uncertainty increasing farther into the forecast, and reproduction of local structure on the noisier GasRateCO2 series.

  6. Knowl 6 — Simplicity and repetition biases support pattern extrapolation

    empirical result

    The authors propose that language models’ preference for simple, compressible sequences helps explain their zero-shot time-series forecasts. In a synthetic experiment, they generated noisy observations from f(x)=x+cos⁡(x)f(x)=x+\cos(x), fit symbolic expressions to the first 70% of the observations, and compared GPT-3 likelihoods for candidate expressions of different training fit and complexity. GPT-3 favored explanations balancing fit with simplicity, consistent with—but not a proof of—a simplicity-biased forecasting mechanism. The paper also connects sequence repetition to periodic extrapolation, and language-model arithmetic abilities to linear or exponential trends. Forecasting combined patterns is less reliable: identifying and executing several operations together can exceed the models’ capabilities, and some synthetic compositions were challenging.

  7. Knowl 7 — Forecasting improves with base-model capability but can worsen after alignment

    empirical result

    Across the evaluated base language models, stronger Massive Multitask Language Understanding (MMLU) performance was associated with better Darts forecasting performance, and increasing LLaMA model size generally improved forecast metrics. This relationship did not hold uniformly for chat-oriented models: LLaMA-2 chat variants typically forecast worse than their corresponding base models, and GPT-4 had worse Darts CRPS than GPT-3 despite its stronger language-task performance. The authors attribute GPT-4’s difficulty partly to its tokenizer, which did not support the digit-separated encoding strategy used for GPT-3, and report poorer uncertainty calibration on stochastic series. They suggest alignment interventions such as reinforcement learning from human feedback may contribute to calibration degradation, while treating that explanation as likely rather than established.

  8. Knowl 8 — Text markers allow forecasts with missing observations and no imputation

    empirical result

    LLMTIME can represent missing time-series observations directly with a text marker such as NaN, rather than first imputing the values. In an experiment that progressively corrupted series with missing observations, LLaMA-2 70B received the NaN-marked inputs, while traditional forecasting baselines received linearly interpolated inputs. The baselines’ likelihoods deteriorated rapidly as missingness increased; LLaMA-2 70B was more resilient in likelihood and had CRPS competitive with the interpolation-based methods. This demonstrates tolerance to missing values under the tested encoding and comparison, not that missingness is harmless for every dataset.

  9. Knowl 9 — Language models can identify patterns in text-formatted time-series tasks

    empirical result

    The paper probed GPT-4’s textual interpretation of time series by providing code that randomly selected one of several candidate functions, the resulting numeric sequence, and a request to identify the generating function. Chain-of-thought prompting produced identification accuracy above random chance, and GPT-4 often described trend or periodicity in its textual reasoning. However, the authors found that the model was better at directly extrapolating numerical patterns than at identifying their generating functions through this text question-answering task. They interpret this gap as evidence that numerical extrapolation ability and textual explanations of numerical patterns are not fully connected in the tested model.

  10. Knowl 10 — Context length and compositional arithmetic remain limitations

    limitation

    LLMTIME inherits language models’ finite context windows, which constrain how much historical data can be included; this is particularly challenging for multivariate series and long histories. The authors also identify arithmetic and recursive composition as potential weaknesses on time series requiring precise combinations of operations. Although many series may not require exact arithmetic, the paper does not establish the extent of this limitation across applications. Effective ways to fine-tune language models for time-series forecasting are left as future work.

Coverage note — The detailed per-series forecast plots and example text explanations are omitted because they illustrate results already captured in the knowls rather than adding distinct findings.

References

  1. 1.Joshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskyi, David Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, et al. Colt5: Faster long-range transformers with conditional computation. arXiv preprint arXiv:2303.09752, 2023.
  2. 2.Zeyuan Allen-Zhu and Yuanzhi Li. Physics of language models: Part 1, context-free grammar. arXiv preprint arXiv:2305.13673, 2023.
  3. 3.Cem Anil, Ashwini Pokle, Kaiqu Liang, Johannes Treutlein, Yuhuai Wu, Shaojie Bai, J Zico Kolter, and Roger B Grosse. Path independent equilibrium models can better exploit test-time computation. Advances in Neural Information Processing Systems, 35:7796–7809, 2022.
  4. 4.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  5. 5.Anthropic. Introducing 100k context windows. Anthropic blog, 2023. URL https://www.anthropic.com/index/100k-context-windows.
  6. 6.Gregory Benton, Nate Gruver, Wesley Maddox, and Andrew Gordon Wilson. Deep probabilistic time series forecasting over long horizons. openreview preprint, 2022. URL https://openreview.net/forum?id=22h1XSEiN0.
  7. 7.Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raf. Emergent and predictable memorization in large language models. arXiv preprint arXiv:2304.11158, 2023.
  8. 8.George EP Box and Gwilym M Jenkins. Some recent advances in forecasting and control. Journal of the Royal Statistical Society. Series C (Applied Statistics), 17(2):91–109, 1968.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  10. 10.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  11. 11.Cristian Challu, Kin G Olivares, Boris N Oreshkin, Federico Garza, Max Mergenthaler, and Artur Dubrawski. N-hits: Neural hierarchical interpolation for time series forecasting. arXiv preprint arXiv:2201.12886, 2022.
  12. 12.Kent K Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. Speak, memory: An archaeology of books known to chatgpt/gpt-4. arXiv preprint arXiv:2305.00118, 2023.
  13. 13.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  14. 14.Miles Cranmer. Interpretable machine learning for science with pysr and symbolicregression. jl. arXiv preprint arXiv:2305.01582, 2023.
  15. 15.Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. Language modeling is compression. arXiv preprint arXiv:2309.10668, 2023.
  16. 16.Dazhao Du, Bing Su, and Zhewei Wei. Preformer: Predictive transformer with multi-scale segment-wise correlations for long-term time series forecasting. arXiv preprint arXiv:2202.11356, 2022.
  17. 17.Jacob R Gardner, Geoff Pleiss, David Bindel, Kilian Q Weinberger, and Andrew Gordon Wilson. Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration. In Advances in Neural Information Processing Systems, 2018.
  18. 18.Federico Garza, Max Mergenthaler Canseco, Cristian Challú, and Kin G Olivares. Statsforecast: Lightning fast forecasting with statistical and econometric models. PyCon: Salt Lake City, UT, USA, 2022. URL https://github.com/Nixtla/statsforecast.
  19. 19.Rakshitha Godahewa, Christoph Bergmeir, Geoffrey I Webb, Rob J Hyndman, and Pablo Montero-Manso. Monash time series forecasting archive. arXiv preprint arXiv:2105.06643, 2021.
  20. 20.Micah Goldblum, Marc Finzi, Keefer Rowan, and Andrew Gordon Wilson. The no free lunch theorem, kolmogorov complexity, and the role of inductive biases in machine learning. arXiv preprint arXiv:2304.05366, 2023.
  21. 21.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT press, 2016.
  22. 22.Nate Gruver, Marc Finzi, Micah Goldblum, and Andrew Gordon Wilson. The lie derivative for measuring learned equivariance. arXiv preprint arXiv:2210.02984, 2022.
  23. 23.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  24. 24.Julien Herzen, Francesco Lässig, Samuele Giuliano Piazzetta, Thomas Neuer, Léo Tafti, Guillaume Raille, Tomas Van Pottelbergh, Marek Pasieka, Andrzej Skrodzki, Nicolas Huguenin, et al. Darts: User-friendly modern machine learning for time series. The Journal of Machine Learning Research, 23(1):5442–5447, 2022.
  25. 25.Hansika Hewamalage, Klaus Ackermann, and Christoph Bergmeir. Forecast evaluation for data scientists: common pitfalls and best practices. Data Mining and Knowledge Discovery, 37(2):788–832, 2023.
  26. 26.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019.
  27. 27.Rob J Hyndman, Anne B Koehler, J Keith Ord, and Ralph D Snyder. Forecasting with exponential smoothing: the state space approach. Springer Science & Business Media, 2008.
  28. 28.Shruti Kaushik, Abhinav Choudhury, Pankaj Kumar Sheron, Nataraj Dasgupta, Sayee Natarajan, Larry A Pickett, and Varun Dutt. Ai in healthcare: time-series forecasting using statistical, neural, and ensemble architectures. Frontiers in big data, 3:4, 2020.
  29. 29.Colin Lea, Rene Vidal, Austin Reiter, and Gregory D Hager. Temporal convolutional networks: A unified approach to action segmentation. In Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part III 14, pages 47–54. Springer, 2016.
  30. 30.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021.
  31. 31.Tiedong Liu and Bryan Kian Hsiang Low. Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks. arXiv preprint arXiv:2305.14201, 2023.
  32. 32.Andriy Mnih and Geoffrey E Hinton. A scalable hierarchical distributed language model. Advances in neural information processing systems, 21, 2008.
  33. 33.Steffen Moritz and Thomas Bartz-Beielstein. imputets: time series missing value imputation in r. R J., 9(1):207, 2017.
  34. 34.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  35. 35.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. In-context learning and induction heads. Transformer Circuits Thread, 2022. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  36. 36.Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: A generative model for raw audio. In 9th ISCA Speech Synthesis Workshop, pages 125–125. ISCA, 2016.
  37. 37.OpenAI. Gpt-4 technical report. arXiv, 2023.
  38. 38.Boris N Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio. N-beats: Neural basis expansion analysis for interpretable time series forecasting. Journal of Machine Learning Research, 21(111):1–63, 2020.
  39. 39.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  40. 40.Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin. Catboost: unbiased boosting with categorical features. In Advances in Neural Information Processing Systems, volume 31, pages 6638–6648. NeurIPS, 2018.
  41. 41.David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020.
  42. 42.Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein. Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks. Advances in Neural Information Processing Systems, 34:6695–6706, 2021.
  43. 43.Ilya Sutskever. An observation on generalization. Workshop on Large Language Models and Transformers, 2023. URL https://www.youtube.com/watch?v=AKMuA_TVz3A&ab_channel=SimonsInstitute.
  44. 44.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  45. 45.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288, 2023.
  46. 46.Juan R Trapero, Nikolaos Kourentzes, and Robert Fildes. On the identification of sales forecasting models in the presence of promotions. Journal of the operational Research Society, 66(2):299–307, 2015.
  47. 47.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  48. 48.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  49. 49.Andrew Wilson and Ryan Adams. Gaussian process kernels for pattern discovery and extrapolation. In International conference on machine learning, pages 1067–1075. PMLR, 2013.
  50. 50.Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34, 2021.
  51. 51.Hao Xue and Flora D. Salim. Promptcast: A new prompt-based learning paradigm for time series forecasting, 2023.
  52. 52.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. How well do large language models perform in arithmetic tasks? arXiv preprint arXiv:2304.02015, 2023.
  53. 53.Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. Are transformers effective for time series forecasting? arXiv preprint arXiv:2205.13504, 2022.
  54. 54.Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue. Meta-transformer: A unified framework for multimodal learning. arXiv preprint arXiv:2307.10802, 2023.
  55. 55.Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of AAAI, 2021.
  56. 56.Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In Proc. 39th International Conference on Machine Learning (ICML 2022), 2022.
  57. 57.Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, and Rong Jin. One fits all: Power general time series analysis by pretrained lm. arXiv preprint arXiv:2302.11939, 2023.

Citation

MLA
Gruver, N., et al. “Large Language Models Are Zero-Shot Time Series Forecasters”. arXiv, 2023, http://arxiv.org/abs/2310.07820v3.
APA
Gruver, N., Finzi, M., Qiu, S., & Wilson, A. G. (2023). Large Language Models Are Zero-Shot Time Series Forecasters. arXiv. http://arxiv.org/abs/2310.07820v3
Chicago
Gruver, N., M. Finzi, S. Qiu, and A. G. Wilson. 2023. “Large Language Models Are Zero-Shot Time Series Forecasters”. arXiv. http://arxiv.org/abs/2310.07820v3.
Harvard
Gruver, N. et al. (2023) “Large Language Models Are Zero-Shot Time Series Forecasters”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.07820v3.
Vancouver
1. Gruver N, Finzi M, Qiu S, Wilson AG (2023) Large Language Models Are Zero-Shot Time Series Forecasters. arXiv

BibTeX

@article{gruver2023large,
  title = {Large Language Models Are Zero-Shot Time Series Forecasters},
  author = {Gruver, Nate and Finzi, Marc and Qiu, Shikai and Wilson, Andrew Gordon},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.07820v3},
  eprint = {2310.07820}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors