Context is Key: A Benchmark for Forecasting with Essential Textual Information

Andrew Robert WilliamsArjun Ashoktienne MarcotteValentina ZantedeschiJithendaraa SubramanianRoland RiachiJames RequeimaAlexandre LacosteIrina RishNicolas Chapados

article2025ICML89 citations

Presents Context is Key (CiK), a time series benchmark spanning seven real-world domains where accurate predictions strictly require integrating textual context with numerical history, alongside a simple prompting baseline that outperforms existing foundation models.

Listen

Real-world decision-making across energy management, public safety, supply chains, and transportation heavily relies on time series forecasting. While traditional quantitative systems analyze purely historical numerical data, human experts routinely interpret background information, operational constraints, and causal relationships expressed in natural language. Existing forecasting benchmarks do not ensure that textual information is indispensable for accuracy, leaving it unclear whether modern large language models can effectively integrate both numbers and text to improve predictive performance.

The article introduces and evaluates "Context is Key" (CiK), a specialized multimodal forecasting benchmark where understanding textual context is strictly required to achieve accurate predictions. The primary objective is to systematically evaluate how well statistical models, specialized numerical time series foundation models, and language model-based forecasters combine historical numerical data with diverse natural language descriptions.

The authors designed 71 distinct tasks spanning seven real-world domains—including energy, climatology, retail, traffic, economics, public safety, and mechanics—comprising 2,644 time series. The benchmark includes diverse context types: intemporal domain knowledge, planned future events, past anomalies, correlated covariates, and causal structures. To ensure data integrity, the authors used recent data and transformations to mitigate contamination risks from model pretraining. Evaluation was conducted across 23 model configurations using a newly introduced evaluation metric, Region of Interest Continuous Ranked Probability Score (RCRPS), which specifically weights critical context-sensitive time windows and penalizes constraint violations.

The investigation produced four key findings. First, integrating natural language context yields substantial accuracy gains for advanced language models, exemplified by Llama-3.1-405B-Instruct using a direct prompting method, which achieved the lowest overall error score (0.159) and improved by 67.1% over its context-free baseline. Second, top-performing language model forecasters with context substantially outperformed traditional statistical methods and dedicated numerical foundation models. Third, no single method dominated across all context types, as models struggled unevenly with complex causal reasoning and mathematical notation. Fourth, language models occasionally suffered severe, catastrophic forecast failures—missing actual values by more than 500%—due to misinterpreting context or blindly relying on irrelevant information.

These findings indicate that large language models hold substantial potential for automating and democratizing multimodal forecasting without requiring manual statistical modeling. However, their tendency toward catastrophic numerical hallucinations introduces operational risks for fully automated deployment in high-stakes environments. In addition, the largest language models require orders of magnitude more computational power and runtime than quantitative forecasters, creating practical tradeoffs between accuracy and cost.

Organizations evaluating language model-based forecasting should implement structured direct prompting strategies and restrict inputs to verified, relevant contextual data, as extraneous text degrades performance. Before deploying these models into production pipelines, organizations must conduct targeted pilot studies with automated guardrails to monitor and clip extreme numerical outliers. Future research should prioritize fine-tuning efficient multimodal models, enhancing causal reasoning, and developing agentic forecasting systems with conversational interfaces.

The findings are bounded by the benchmark's focus on univariate time series and textual context, excluding multivariate series, image data, or tabular databases. While the study demonstrates high statistical confidence regarding the benefits of context for top-performing architectures, readers should remain cautious regarding smaller models, which frequently fail to incorporate context effectively and can perform worse than simple statistical baselines.

No sufficiently relevant recommendations were found.

Cover for Context is Key: A Benchmark for Forecasting with Essential Textual Information

Abstract

Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of recent progress with LLM-based forecasters, their ability to effectively integrate this textual information remains an open question. To address this, we introduce “Context is Key” (CiK), a time series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities; crucially, every task in CiK requires understanding textual context to be solved successfully. We evaluate a range of approaches, including statistical models, time series foundation models, and LLM-based forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. This benchmark aims to advance multimodal forecasting by promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0/.

Table of Contents

  • 1. Introduction
  • 2. Problem Setting
  • 3. Context is Key: a Natural Language Context-Aided Forecasting Benchmark
  • 3.1. Domains and Numerical Data Sources
  • 3.2. Natural Language Context
  • 3.3. Validating the Relevance of The Context
  • 4. Region of Interest CRPS
  • 5. Experiments and Results
  • 5.1. Evaluation Protocol
  • 5.2. Methods
  • 5.3. Results on CiK
  • 5.4. Error Analysis
  • 5.5. Inference Cost
  • 6. Related Work
  • 7. Discussion
  • Impact Statement
  • Acknowledgements
  • References
  • Appendix
  • A. Additional Details on the Benchmark
  • A.1. Data Sources
  • A.2. Task Creation Process
  • A.3. Model Capabilities
  • A.4. Validating the Relevance of the Context
  • A.5. Weighting scheme for tasks
  • A.6. Standard errors and average ranks
  • A.7. Task lengths
  • B. Examples of tasks from the benchmark
  • B.1. Task: Constrained Predictions
  • B.2. Task: Electrical Consumption Increase
  • B.3. Task: ATM Maintenance
  • B.4. Task: Montreal Fire High Season
  • B.5. Task: Solar Prediction
  • B.6. Task: Speed From Load
  • C. Additional Results
  • C.1. Extended results on all models
  • C.2. Full results partitioned by types of context
  • C.3. Full results partitioned by model capabilities
  • C.4. Inference Time
  • C.5. Significant failures per model
  • C.6. Testing the Statistical Significance of the Relevance of Context
  • C.7. Cost of API-based models
  • C.8. Impact of Relevant and Irrelevant Information in Context
  • C.9. Impact of Solely Irrelevant Information in Context
  • C.10. The effect of significant failures on the aggregate performance of models
  • C.11. Visualizations of successful context-aware forecasts
  • C.12. Visualizations of significant failures
  • D. Implementation Details of Models
  • D.1. DIRECT PROMPT
  • D.1.1. METHOD
  • D.1.2. IMPLEMENTATION DETAILS
  • D.1.3. EXAMPLE PROMPT
  • D.2. LLMP
  • D.2.1. METHOD
  • D.2.2. IMPLEMENTATION DETAILS
  • D.3. ChatTime
  • D.4. UniTime and Time-LLM
  • D.5. Lag-Llama
  • D.6. Chronos
  • D.7. Moirai
  • D.8. TimeGEN
  • D.9. Exponential Smoothing
  • D.10. ETS and ARIMA
  • E. Details of the proposed metric
  • E.1. Scaling for cross-task aggregation
  • E.2. CRPS and twCRPS
  • E.3. Estimating the CRPS using samples
  • E.4. Constraint-violation functions
  • E.5. Covariance of two CRPS estimators
  • E.6. Comparison of Statistical Properties of Various Scoring Rules

Knowls

  1. Knowl 1 — CiK benchmarks forecasting that depends on textual context

    model/method

    Context is Key (CiK) is a probabilistic time-series forecasting benchmark in which models receive both a numerical history and natural-language context intended to be essential for accurate forecasts. It contains 71 manually designed tasks, based on 2,644 time series, with task instances generated by sampling series and time windows and varying context formulations. About 95% of tasks use real-world data from seven domains: climatology, economics, energy, mechanics, public safety, transportation, and retail; the remainder use synthetic data. Sampling frequencies range from 10 minutes to monthly. The benchmark has no training set and is intended to evaluate models that forecast directly from each instance’s history, with or without its text.

  2. Knowl 2 — CiK distinguishes five kinds of forecasting context

    definition

    CiK labels contextual information by what it contributes to a forecast: intemporal information describes time-invariant properties of the process, including the target’s meaning, long-period patterns, or value constraints; future information describes future events or scenarios and their consequences for the series; historical information supplies facts about past behavior not recoverable from the available history, such as summary statistics or explanations of spurious patterns; covariate information gives information about other variables statistically associated with the target; and causal information describes causal or confounded relationships between covariates and the target. A task can use more than one context type.

  3. Knowl 3 — Tasks are handcrafted and checked for contextual relevance

    model/method

    The authors designed CiK tasks from scratch without external annotators or LLMs. For each task they selected a data source and history window, devised relevant context, and, when needed, modified future numerical values to match the scenario described in that context. Other authors peer-reviewed each task; tasks judged insufficiently relevant were revised or excluded. In a separate validation, 94.7% of human annotations rated the context useful for improving forecasts. GPT-4o assessed five instances of every task as enabling better forecasts, with most judged to enable significantly better forecasts.

  4. Knowl 4 — RCRPS weights relevant windows and constraint violations

    equation

    The Region-of-Interest Continuous Ranked Probability Score (RCRPS) extends the univariate continuous ranked probability score (CRPS) to emphasize context-relevant forecast times and penalize constraint violations. Let JJ be the set of forecast time indices, I⊆JI\subseteq J the task’s region of interest (RoI), and J∖IJ\setminus I the remaining forecast indices. Let X~F\tilde X_F be a model’s predictive random trajectory and xFx_F the observed future trajectory; X~i\tilde X_i and xix_i are their predictive marginal and realization at index ii. Let vCv_C map a trajectory to a nonnegative scalar constraint violation, with value zero when all constraints are satisfied. Let α\alpha be a positive task-specific scale factor estimated using 25 additional instances, and let β=10\beta=10 in the experiments. Then

    RCRPS⁡(X~F,xF)={α[12∣I∣∑i∈ICRPS⁡(X~i,xi)+12∣J∖I∣∑i∈J∖ICRPS⁡(X~i,xi)+β CRPS⁡(vC(X~F),0)],∣I∣>0,α[1∣J∣∑i∈JCRPS⁡(X~i,xi)+β CRPS⁡(vC(X~F),0)],∣I∣=0.\operatorname{RCRPS}(\tilde X_F,x_F)=\begin{cases} \alpha\left[\frac{1}{2|I|}\sum_{i\in I}\operatorname{CRPS}(\tilde X_i,x_i)+\frac{1}{2|J\setminus I|}\sum_{i\in J\setminus I}\operatorname{CRPS}(\tilde X_i,x_i)+\beta\,\operatorname{CRPS}(v_C(\tilde X_F),0)\right], & |I|>0,\\ \alpha\left[\frac{1}{|J|}\sum_{i\in J}\operatorname{CRPS}(\tilde X_i,x_i)+\beta\,\operatorname{CRPS}(v_C(\tilde X_F),0)\right], & |I|=0. \end{cases}

    Thus, when an RoI exists, it receives half the time-series scoring weight and the rest of the horizon receives the other half; the violation term is zero for forecasts satisfying the constraints. CRPS can be estimated from forecast samples, so a closed-form forecast distribution is not required. The authors note that RCRPS is not guaranteed to be proper in CiK because instances used for scaling can overlap or be autocorrelated.

  5. Knowl 5 — Direct Prompt elicits timestamped forecasts in one pass

    model/method

    Direct Prompt is a zero-shot LLM forecasting method that asks for the complete forecast in a single model response. Its inputs are the natural-language task context, historical observations formatted as timestamp–value pairs, and the timestamps to predict. The prompt asks the model to account for the context and return only timestamp–value forecast pairs inside designated forecast tags. Constrained decoding with a regular expression restricts output to values at the requested timestamps; outputs that fail formatting are rejected and retried. In the benchmark evaluation, 25 independent forecast samples are generated per instance. The method’s structured-output requirement favors instruction-tuned models: several base models failed to produce valid structured forecasts even after as many as 50 retries.

  6. Knowl 6 — Evaluation uses repeated instances and balanced task clusters

    experimental setup

    The authors deterministically sampled five instances per CiK task, for 355 instances in total, and generated 25 independent forecasts per model and instance. To prevent data sources with many related tasks from dominating aggregate scores, they grouped similar tasks into clusters, gave every cluster equal total weight, and divided a cluster’s weight equally among its tasks. Scores were reported as weighted averages, accompanied by standard errors; model ranks were estimated by simulation using task-score means and standard errors. Models were evaluated zero-shot or fit directly to each instance’s history, with context-capable models also evaluated without text.

  7. Knowl 7 — Direct Prompt with Llama-3.1-405B leads aggregate results

    empirical result

    On the weighted CiK aggregate, Direct Prompt with Llama-3.1-405B-Instruct achieved the lowest reported RCRPS, 0.159±0.0080.159\pm0.008. Other strong context-using results included LLMP with base Llama-3-70B at 0.236±0.0060.236\pm0.006, LLMP with base Mixtral-8x7B at 0.262±0.0080.262\pm0.008, Direct Prompt with GPT-4o at 0.274±0.0100.274\pm0.010, and Direct Prompt with Qwen-2.5-7B-Instruct at 0.290±0.0030.290\pm0.003. The best numerical-only time-series foundation model shown in the selected results, Chronos-Large, scored 0.326±0.0020.326\pm0.002; ARIMA scored 0.475±0.0060.475\pm0.006. These results show that some prompted LLMs can outperform the tested numerical-only and multimodal forecasters on this context-dependent benchmark, but performance varies considerably by model and prompting approach. Under LLMP, instruction tuning substantially worsened Llama-3-70B’s with-context RCRPS (0.539±0.0130.539\pm0.013 versus 0.236±0.0060.236\pm0.006 for the base model), while Mixtral-8x7B’s results were similar with and without instruction tuning (0.264±0.0040.264\pm0.004 and 0.262±0.0080.262\pm0.008).

  8. Knowl 8 — Context improves many models, but gains depend on model and context type

    empirical result

    For Direct Prompt with Llama-3.1-405B-Instruct, adding context reduced aggregate RCRPS by 67.1%; in the unweighted paired comparison, the score was 0.165±0.0050.165\pm0.005 with context and 0.544±0.0070.544\pm0.007 without it (p=6.92×10−13p=6.92\times10^{-13}, one-sided paired tt-test). Many other models also improved, but some were unchanged or worse, so access to text alone did not ensure that it was used successfully. No method was best across all five context types: among selected results, the best RCRPS was 0.1740.174 for intemporal information (Direct Prompt, Llama-3.1-405B-Instruct), 0.1180.118 for historical information (Direct Prompt, GPT-4o), 0.0750.075 for future information (Direct Prompt, Llama-3.1-405B-Instruct), 0.1640.164 for covariate information (Direct Prompt, Llama-3.1-405B-Instruct), and 0.3600.360 for causal information (LLMP, base Llama-3-70B). Causal-context tasks were comparatively difficult for the evaluated methods.

  9. Knowl 9 — A minority of extreme forecasts can dominate aggregate scores

    empirical result

    The authors classified an instance as a significant failure when its RCRPS exceeded 5, a threshold associated with a forecast error on the scale of five times the ground-truth range; such scores were clipped to 5 during aggregation. Failure counts varied sharply across models: Direct Prompt with Llama-3.1-405B-Instruct had 0 significant failures among 355 instances, while LLMP with Qwen-2.5-7B-Instruct had 107 and LLMP with base Qwen-2.5-0.5B had 109. Consequently, a model can improve on many tasks yet have its aggregate score worsened by a small number of extreme failures. The authors suggest that context misinterpretation may contribute to these failures, but leave their causes open for further study.

  10. Knowl 10 — Benchmark scope and practical limitations

    limitation

    CiK evaluates natural-language context and univariate forecasting; it does not cover other context modalities or multivariate time-series tasks. Its tasks test information selected and articulated by humans, not whether models can exploit latent relationships that human task designers did not identify. Although the authors mitigate memorization through post-cutoff live data, derived series, date shifts, or added noise where appropriate, they cannot guarantee that memorization is eliminated without strictly held-out data. Dataset-specific methods may also perform better if trained on task-specific data, which CiK does not provide. In deployment, the study additionally finds that LLM forecasting can be computationally expensive: numerical forecasters were generally faster, and LLMP took about an order of magnitude longer than Direct Prompt for the compared Llama models.

Coverage note — The appendix’s exhaustive per-model result tables, detailed baseline training configurations, and individual worked task examples are omitted; they support the benchmark and findings summarized here but do not add separate main contributions.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 6
  2. 2.Aksu, T., Liu, C., Saha, A., Tan, S., Xiong, C., and Sahoo, D. Xforecast: Evaluating natural language explanations for time series forecasting. arXiv preprint arXiv:2410.14180, 2024. 8
  3. 3.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., Stella, L., Turkmen, A. C., and Wang, Y. GluonTS: Probabilistic and Neural Time Series Modeling in Python. Journal of Machine Learning Research, 21(116):1–6, 2020. URL http://jmlr.org/papers/v21/19-820.html. 55
  4. 4.Allen, S., Ginsbourger, D., and Ziegel, J. Evaluating forecasts for high-impact events using transformed kernel scores. SIAM/ASA Journal on Uncertainty Quantification, 11(3):906–940, 2023. doi: 10.1137/22M1532184. 55
  5. 5.Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024. 6, 53
  6. 6.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. 1, 7
  7. 7.Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M. Time series analysis: forecasting and control. John Wiley & Sons, fifth edition, 2015. 6
  8. 8.Chen, C., Petty, K., Skabardonis, A., Varaiya, P., and Jia, Z. Freeway performance measurement system: mining loop detector data. Transportation research record, 1748(1):96–102, 2001. 3, 14, 21
  9. 9.Chen, Z., Ma, M., Li, T., Wang, H., and Li, C. Long sequence time-series forecasting with deep learning: A survey. Information Fusion, 97:101819, 2023. 1
  10. 10.Chow, W., Gardiner, L., Hallgrímsson, H. T., Xu, M. A., and Ren, S. Y. Towards time series reasoning with llms. arXiv preprint arXiv:2409.11376, 2024. 8
  11. 11.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. 6, 48
  12. 12.Emami, P., Li, Z., Sinha, S., and Nguyen, T. Syscaps: Language interfaces for simulation surrogates of complex systems. arXiv preprint arXiv:2405.19653, 2024. 2, 4, 8
  13. 13.Gamella, J. L., Bühlmann, P., and Peters, J. The causal chambers: Real physical systems as a testbed for AI methodology. arXiv preprint arXiv:2404.11341, 2024. 3, 14, 26
  14. 14.Gardner Jr., E. S. Exponential smoothing: The state of the art. Journal of Forecasting, 4(1):1–28, 1985. doi: https://doi.org/10.1002/for.3980040103. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/for.3980040103. 6
  15. 15.Garza, A. and Mergenthaler-Canseco, M. Nixtla foundation-time-series-arena. https://github.com/Nixtla/nixtla/tree/main/experiments/foundation-time-series-arena, 2024. 1
  16. 16.Garza, A., Challu, C., and Mergenthaler-Canseco, M. TimeGPT-1. arXiv preprint arXiv:2310.03589, 2023. 6, 54
  17. 17.Gneiting, T. and Raftery, A. E. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007. 4, 8, 55
  18. 18.Gneiting, T. and Ranjan, R. Comparing density forecasts using threshold- and quantile-weighted scoring rules. Journal of Business & Economic Statistics, 29(3):411–422, 2011. doi: 10.1198/jbes.2010.08110. 5, 55
  19. 19.Godahewa, R., Bergmeir, C., Webb, G. I., Hyndman, R. J., and Montero-Manso, P. Monash time series forecasting archive. arXiv preprint arXiv:2105.06643, 2021. 3, 14, 15
  20. 20.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. Advances in Neural Information Processing Systems, 36, 2024. 7, 8
  21. 21.Hyndman, R., Koehler, A. B., Ord, J. K., and Snyder, R. D. Forecasting with exponential smoothing: the state space approach. Springer Science & Business Media, 2008. 6
  22. 22.Hyndman, R. J. and Athanasopoulos, G. Forecasting: principles and practice. OTexts, 2018. 1, 17
  23. 23.Jia, F., Wang, K., Zheng, Y., Cao, D., and Liu, Y. Gpt4mts: Prompt-based large language model for multimodal time-series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21):23343–23351, Mar. 2024. doi: 10.1609/aaai.v38i21.30383. URL https://ojs.aaai.org/index.php/AAAI/article/view/30383. 8
  24. 24.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts, 2024. URL https://arxiv.org/abs/2401.04088. 6
  25. 25.Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., and Wen, Q. Time-LLM: Time series forecasting by reprogramming large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Unb5CVPtae. 1, 6, 8, 9, 52
  26. 26.Kim, K., Tsai, H., Sen, R., Das, A., Zhou, Z., Tanpure, A., Luo, M., and Yu, R. Multi-modal forecaster: Jointly predicting time series and textual data. arXiv preprint arXiv:2411.06735, 2024. 8
  27. 27.Kong, Y., Yang, Y., Wang, S., Liu, C., Liang, Y., Jin, M., Zohren, S., Pei, D., Liu, Y., and Wen, Q. Position: Empowering time series reasoning with multimodal llms. arXiv preprint arXiv:2502.01477, 2025. 8
  28. 28.Liang, Y., Wen, H., Nie, Y., Jiang, Y., Jin, M., Song, D., Pan, S., and Wen, Q. Foundation models for time series analysis: A tutorial and survey. arXiv preprint arXiv:2403.14735, 2024. 1
  29. 29.Lim, B. and Zohren, S. Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A, 379(2194):20200209, 2021. 1
  30. 30.Liu, H., Xu, S., Zhao, Z., Kong, L., Kamarthi, H., Sasanur, A. B., Sharma, M., Cui, J., Wen, Q., Zhang, C., et al. Time-MMD: A new multi-domain multimodal dataset for time series analysis. arXiv preprint arXiv:2406.08627, 2024a. 2, 4, 8
  31. 31.Liu, M., Zhu, M., Wang, X., Ma, G., Yin, J., and Zheng, X. Echo-gl: Earnings calls-driven heterogeneous graph learning for stock movement prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 13972–13980, 2024b. 8
  32. 32.Liu, X., Hu, J., Li, Y., Diao, S., Liang, Y., Hooi, B., and Zimmermann, R. Unitime: A language-empowered unified model for cross-domain time series forecasting. In Proceedings of the ACM on Web Conference 2024, pp. 4095–4106, 2024c. 1, 6, 8, 9, 52
  33. 33.Merrill, M. A., Tan, M., Gupta, V., Hartvigsen, T., and Althoff, T. Language models still struggle to zero-shot reason about time series. arXiv preprint arXiv:2404.11757, 2024. 2, 4, 8
  34. 34.Potosnak, W., Challu, C., Goswami, M., Wilinski, M., Zukowska, N., and Dubrawski, A. Implicit reasoning in deep time series forecasting. arXiv preprint arXiv:2409.10840, 2024. 8
  35. 35.Rasul, K., Ashok, A., Williams, A. R., Khorasani, A., Adamopoulos, G., Bhagwatkar, R., Bilos, M., Ghonia, H., Hassen, N. V., Schneider, A., et al. Lag-Llama: Towards foundation models for time series forecasting. arXiv preprint arXiv:2310.08278, 2023. 6, 53
  36. 36.Requeima, J., Bronskill, J., Choi, D., Turner, R. E., and Duvenaud, D. LLM processes: Numerical predictive distributions conditioned on natural language. arXiv preprint arXiv:2405.12856, 2024. 1, 6, 8, 50
  37. 37.Sawhney, R., Wadhwa, A., Agarwal, S., and Shah, R. Fast: Financial news and tweet based time aware network for stock trading. In Proceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume, pp. 2164–2175, 2021. 8
  38. 38.Sengupta, M., Xie, Y., Lopez, A., Habte, A., Maclaurin, G., and Shelby, J. The national solar radiation data base (NSRDB). Renewable and sustainable energy reviews, 89:51–60, 2018. 3, 14
  39. 39.Taillardat, M., Mestre, O., Zamo, M., and Naveau, P. Calibrated ensemble forecasts using quantile regression forests and ensemble model output statistics. Monthly Weather Review, 144(6):2375 – 2393, 2016. doi: 10.1175/MWR-D-15-0260.1. 55, 56
  40. 40.U.S. Bureau of Labor Statistics. Unemployment rate [various locations], 2024. URL https://fred.stlouisfed.org/. Accessed on 2024-08-30, retrieved from FRED. 3, 15
  41. 41.Ville de Montréal. Interventions des pompiers de montréal, 2020. URL https://www.donneesquebec.ca/recherche/dataset/vmtl-interventions-service-securite-incendie-montreal. Updated on 2024-09-12, accessed on 2024-09-13. 3, 14
  42. 42.Wang, C., Qi, Q., Wang, J., Sun, H., Zhuang, Z., Wu, J., Zhang, L., and Liao, J. Chattime: A unified multimodal time series foundation model bridging numerical and textual data. arXiv preprint arXiv:2412.11376, 2024a. 8, 9
  43. 43.Wang, C., Qi, Q., Wang, J., Sun, H., Zhuang, Z., Wu, J., Zhang, L., and Liao, J. Chattime: A unified multimodal time series foundation model bridging numerical and textual data. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pp. 12694–12702, 2025. 6
  44. 44.Wang, P. On defining artificial intelligence. Journal of Artificial General Intelligence, 10(2):1–37, 2019. 1
  45. 45.Wang, X., Feng, M., Qiu, J., Gu, J., and Zhao, J. From news to forecast: Integrating event analysis in LLM-based time series forecasting with reflection. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024b. URL https://openreview.net/forum?id=tj8nsfxi5r. 8
  46. 46.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021. 48
  47. 47.Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series forecasting transformers. arXiv preprint arXiv:2402.02592, 2024. 6, 53
  48. 48.Wu, W., Zhang, G., tan zheng, Wang, Y., and Qi, H. Dual-forecaster: A multimodal time series model integrating descriptive and predictive texts, 2025. URL https://openreview.net/forum?id=QE1ClsZjOQ. 8
  49. 49.Xu, Z., Bian, Y., Zhong, J., Wen, X., and Xu, Q. Beyond trend and periodicity: Guiding time series forecasting with textual cues. arXiv preprint arXiv:2405.13522, 2024. 2, 8
  50. 50.Xue, H. and Salim, F. D. Promptcast: A new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering, 2023. 8
  51. 51.Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al. Qwen2 technical report. CoRR, 2024. 6
  52. 52.Ye, W., Zhang, Y., Yang, W., Tang, L., Cao, D., Cai, J., and Liu, Y. Beyond forecasting: Compositional time series reasoning for end-to-end task execution. arXiv preprint arXiv:2410.04047, 2024. 8
  53. 53.Zamo, M. and Naveau, P. Estimation of the continuous ranked probability score with limited information and applications to ensemble weather forecasts. Mathematical Geosciences, 50(2):209 – 234, 2018. doi: 10.1007/s11004-017-9709-7. 55, 56
  54. 54.Zhang, W., Ye, J., Li, Z., Li, J., and Tsung, F. Dualtime: A dual-adapter multimodal language model for time series representation. arXiv preprint arXiv:2406.06620, 2024. 8
  55. 55.Zhang, Y., Zhang, Y., Zheng, M., Chen, K., Gao, C., Ge, R., Teng, S., Jelloul, A., Rao, J., Guo, X., Fang, C.-W., Zheng, Z., and Yang, J. Insight miner: A large-scale multimodal model for insight mining from time series. In NeurIPS 2023 AI for Science Workshop, 2023. URL https://openreview.net/forum?id=E1khscdUdH. 2, 4, 8
  56. 56.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In The Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Virtual Conference, volume 35, pp. 11106–11115. AAAI Press, 2021. 52

Citation

MLA
Williams, A. R., et al. “Context Is Key: A Benchmark for Forecasting with Essential Textual Information”. arXiv, 2024, http://arxiv.org/abs/2410.18959v4.
APA
Williams, A. R., Ashok, A., Marcotte, É., Zantedeschi, V., Subramanian, J., Riachi, R., Requeima, J., Lacoste, A., Rish, I., Chapados, N., & Drouin, A. (2024). Context is Key: A Benchmark for Forecasting with Essential Textual Information. arXiv. http://arxiv.org/abs/2410.18959v4
Chicago
Williams, A. R., A. Ashok, É. Marcotte, et al. 2024. “Context Is Key: A Benchmark for Forecasting with Essential Textual Information”. arXiv. http://arxiv.org/abs/2410.18959v4.
Harvard
Williams, A.R. et al. (2024) “Context is Key: A Benchmark for Forecasting with Essential Textual Information”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.18959v4.
Vancouver
1. Williams AR, Ashok A, Marcotte É, et al (2024) Context is Key: A Benchmark for Forecasting with Essential Textual Information. arXiv

BibTeX

@article{williams2024context,
  title = {Context is Key: A Benchmark for Forecasting with Essential Textual Information},
  author = {Williams, Andrew Robert and Ashok, Arjun and Marcotte, Étienne and Zantedeschi, Valentina and Subramanian, Jithendaraa and Riachi, Roland and Requeima, James and Lacoste, Alexandre and Rish, Irina and Chapados, Nicolas and Drouin, Alexandre},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.18959v4},
  eprint = {2410.18959}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/