Large Language Models Are Zero-Shot Time Series Forecasters
Nate GruverMarc FinziShikai QiuAndrew Gordon Wilson
Demonstrates that off-the-shelf large language models can match or exceed specialized domain models at zero-shot time series forecasting simply by framing numerical sequences as next-token prediction tasks.
Time series forecasting is essential across sectors such as finance, energy, healthcare, and supply chain management, yet building accurate deep learning models typically requires substantial domain-specific expertise, large curated datasets, and expensive task-specific training. Standard foundation models have transformed natural language and computer vision, but unified pretraining approaches for time series have historically lagged behind, often leaving simple classical methods like linear models or autoregressive moving averages to outperform complex neural networks.
The article demonstrates that off-the-shelf large language models (LLMs), such as GPT-3 and LLaMA-2, can serve as highly accurate zero-shot time series forecasters without any task-specific fine-tuning. By introducing a framework called LLMTIME, the authors show that framing numerical sequences as text allows LLMs to transfer their pattern recognition and probabilistic modeling capabilities directly to continuous time series data.
To evaluate this approach, the researchers encoded numerical series into character strings, applying systematic scaling and tokenization adjustments such as inserting spaces between digits to align with model vocabularies. They converted the models' discrete token probability distributions into continuous probability densities to properly quantify forecast uncertainty. The method was evaluated across synthetic benchmarks and standard real-world datasets spanning dozens of distinct domains—including the Darts, Monash, and Informer suites—and benchmarked against dedicated forecasting models, including neural architectures, Gaussian processes, and statistical baselines.
The investigation yielded several critical findings. First, LLMTIME achieved the best or second-best deterministic performance (measured by Mean Absolute Error) across all aggregated benchmark suites in a completely zero-shot setting, matching or exceeding models trained specifically on downstream data. Second, LLMs significantly outperformed specialized models in probabilistic metrics, generating superior continuous ranked probability scores and log likelihoods due to their inherent ability to capture multimodal, heavy-tailed distributions. Third, LLM forecasting exhibited high sample efficiency, outperforming traditional baselines even when context data was heavily restricted. Fourth, models naturally adapted to missing data denoted simply as text markers, avoiding complex manual imputation steps. Finally, model alignment interventions like reinforcement learning from human feedback degraded forecasting accuracy; specifically, GPT-4 and chat-tuned models exhibited worse uncertainty calibration and higher error rates compared to raw base models like GPT-3.
These findings suggest that organizations can bypass the substantial labor, computational cost, and infrastructure overhead associated with training specialized forecasting models, particularly when historical data is scarce. Because LLMs treat forecasting as text completion, organizations can also leverage flexible multimodal capabilities, such as incorporating contextual textual notes and querying the model in natural language to explain observed patterns.
Decision-makers considering LLM-based forecasting should use unaligned base models rather than chat-aligned models to preserve numerical calibration and accuracy. In the near term, teams should pilot zero-shot LLM forecasting on high-value univariate workflows alongside existing pipelines to benchmark latency and cost trade-offs. Further work is recommended to explore extended context windows and efficient multivariate handling before completely replacing dedicated systems.
Confidence in these findings is high for univariate and moderately long sequences, supported by tests on data collected after the models' pretraining cutoff dates to rule out memorization. However, caution is warranted when deploying LLMs on highly dimensional multivariate datasets or long-horizon forecasting tasks, where context window limitations and independent covariate modeling present computational constraints.
No sufficiently relevant recommendations were found.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). This later position paper broadens LLMTIME’s zero-shot forecasting result into a wider account of LLMs as predictors, data enhancers, and analytical agents for time series.
- Paper: GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting, Furong Jia et al. (2024). Building on the use of language models for forecasting, GPT4MTS brings textual context into the forecasting pipeline to test whether it can improve predictions.
- Paper: Context is Key: A Benchmark for Forecasting with Essential Textual Information, Andrew Robert Williams et al. (2025). This later benchmark advances the source’s suggestion that textual context can aid forecasting by testing whether forecasters use essential context to make accurate predictions.
