Context is Key: A Benchmark for Forecasting with Essential Textual Information
Andrew Robert WilliamsArjun Ashoktienne MarcotteValentina ZantedeschiJithendaraa SubramanianRoland RiachiJames RequeimaAlexandre LacosteIrina RishNicolas Chapados
Presents Context is Key (CiK), a time series benchmark spanning seven real-world domains where accurate predictions strictly require integrating textual context with numerical history, alongside a simple prompting baseline that outperforms existing foundation models.
Real-world decision-making across energy management, public safety, supply chains, and transportation heavily relies on time series forecasting. While traditional quantitative systems analyze purely historical numerical data, human experts routinely interpret background information, operational constraints, and causal relationships expressed in natural language. Existing forecasting benchmarks do not ensure that textual information is indispensable for accuracy, leaving it unclear whether modern large language models can effectively integrate both numbers and text to improve predictive performance.
The article introduces and evaluates "Context is Key" (CiK), a specialized multimodal forecasting benchmark where understanding textual context is strictly required to achieve accurate predictions. The primary objective is to systematically evaluate how well statistical models, specialized numerical time series foundation models, and language model-based forecasters combine historical numerical data with diverse natural language descriptions.
The authors designed 71 distinct tasks spanning seven real-world domains—including energy, climatology, retail, traffic, economics, public safety, and mechanics—comprising 2,644 time series. The benchmark includes diverse context types: intemporal domain knowledge, planned future events, past anomalies, correlated covariates, and causal structures. To ensure data integrity, the authors used recent data and transformations to mitigate contamination risks from model pretraining. Evaluation was conducted across 23 model configurations using a newly introduced evaluation metric, Region of Interest Continuous Ranked Probability Score (RCRPS), which specifically weights critical context-sensitive time windows and penalizes constraint violations.
The investigation produced four key findings. First, integrating natural language context yields substantial accuracy gains for advanced language models, exemplified by Llama-3.1-405B-Instruct using a direct prompting method, which achieved the lowest overall error score (0.159) and improved by 67.1% over its context-free baseline. Second, top-performing language model forecasters with context substantially outperformed traditional statistical methods and dedicated numerical foundation models. Third, no single method dominated across all context types, as models struggled unevenly with complex causal reasoning and mathematical notation. Fourth, language models occasionally suffered severe, catastrophic forecast failures—missing actual values by more than 500%—due to misinterpreting context or blindly relying on irrelevant information.
These findings indicate that large language models hold substantial potential for automating and democratizing multimodal forecasting without requiring manual statistical modeling. However, their tendency toward catastrophic numerical hallucinations introduces operational risks for fully automated deployment in high-stakes environments. In addition, the largest language models require orders of magnitude more computational power and runtime than quantitative forecasters, creating practical tradeoffs between accuracy and cost.
Organizations evaluating language model-based forecasting should implement structured direct prompting strategies and restrict inputs to verified, relevant contextual data, as extraneous text degrades performance. Before deploying these models into production pipelines, organizations must conduct targeted pilot studies with automated guardrails to monitor and clip extreme numerical outliers. Future research should prioritize fine-tuning efficient multimodal models, enhancing causal reasoning, and developing agentic forecasting systems with conversational interfaces.
The findings are bounded by the benchmark's focus on univariate time series and textual context, excluding multivariate series, image data, or tabular databases. While the study demonstrates high statistical confidence regarding the benefits of context for top-performing architectures, readers should remain cautious regarding smaller models, which frequently fail to incorporate context effectively and can perform worse than simple statistical baselines.
- Paper: GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting, Furong Jia et al. (2024). Read this earlier multimodal forecasting framework first to see how textual context was incorporated into time-series prediction before CiK benchmarks whether that context is truly indispensable.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). Its overview of LLM roles in time-series analysis establishes the forecasting landscape and approaches that CiK evaluates with context-sensitive tasks.
No sufficiently relevant recommendations were found.
