GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting
Furong JiaKevin WangYixiang ZhengDefu CaoYan Liu
Proposes a prompt-tuning framework alongside an automated text-collection pipeline to integrate textual news summaries with numerical data for multimodal time-series forecasting.
Conventional time-series forecasting relies almost entirely on historical numerical data to predict future trends. While numerical metrics capture underlying statistical patterns, they fail to reflect the real-world contextual events, public sentiment, and news developments that actively drive those metrics. Integrating textual context into predictive modeling has remained difficult due to the lack of structured multimodal datasets and the challenge of effectively combining text and numerical streams without distorting forecasts.
The main objective of the article is to establish an end-to-end pipeline for constructing multimodal time-series datasets using large language models and to demonstrate a framework, named GPT4MTS, that uses textual information as prompts to improve time-series forecasting accuracy.
To evaluate this approach, the researchers built a dataset using media coverage records from the Global Database of Events, Language, and Tone (GDELT), spanning August 2022 to July 2023 across 55 United States regions and national data. The dataset paired daily event metrics—such as the number of mentions, articles, and sources across 10 major event categories—with automated text summaries generated and filtered through language models. The GPT4MTS framework processes numerical data into structured patches while converting relevant news summaries into soft prompt embeddings. These prompts guide a pre-trained language model whose internal attention layers remain frozen to maintain computational efficiency, and predictions are extracted strictly from the processed numerical representations.
The evaluation yielded several key findings regarding predictive performance and architectural design. First, GPT4MTS achieved superior overall accuracy across the 10 event categories compared to state-of-the-art transformer baselines, linear models, and text-free language model adaptations. Compared to the leading unimodal language model baseline (GPT4TS), the proposed approach achieved a 4.14% reduction in Mean Squared Error (MSE) and a 1.0% reduction in Mean Absolute Error (MAE). Second, ablation experiments showed that passing only the numerical output states to the final prediction layer is critical; including textual representations in the final layer degraded accuracy, proving that text should act strictly as a guiding prompt rather than a direct numerical predictor. Third, among unimodal baselines relying solely on numerical data, simpler linear models consistently outperformed complex transformer architectures, which tended to over-complicate numerical sequences.
These findings demonstrate that combining qualitative narrative context with quantitative metrics provides a practical, low-overhead way to enhance forecast reliability. Freezing core language model parameters allows organizations to leverage pre-trained language intelligence without prohibitive retraining costs. This multimodal approach makes complex trend tracking more interpretable and accessible for decision-makers in public policy, communication, and finance who need to anticipate public attention and event impacts.
Organizations and researchers seeking to improve trend forecasting should consider piloting multimodal ingestion pipelines that pair core numerical metrics with automated text summarization. Technical teams implementing such architectures must ensure text features serve strictly as guidance prompts rather than direct output inputs to prevent forecast degradation. Future work should expand the pipeline to additional domains, explore newer foundation models, and conduct pilot deployments in live decision-making environments.
Confidence in these findings is supported by consistent testing across more than 480,000 regional and national data splits. However, readers should note certain limitations: the empirical evaluation is confined to GDELT news tracking data over a single one-year period, and budget constraints necessitated using alternative summarization tools for select regional data, requiring title substitutions for lower-quality summaries. Further validation across broader industries and longer historical horizons remains necessary before full-scale commercial or operational deployment.
- Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). This paper establishes prompt-based multimodal fusion paradigms for integrating disparate modalities, providing foundational prompting techniques adapted by GPT4MTS to fuse numerical time-series and textual context.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). This survey provides a comprehensive analysis of adapting Transformer architectures to time-series forecasting, establishing the baseline concepts and structural challenges addressed when prompting language models with sequential data.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). This work introduces patching and channel-independence tokenization for time-series Transformers, which directly influenced tokenization and prompt-formatting strategies for numerical series in LLMs.
- Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). This study analyzes prompt strategies and aggregation mechanisms for foundation models, offering prerequisite insight into designing effective prompt structures for complex analytical tasks.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This foundational paper presents cross-modal attention mechanisms for unaligned multimodal sequences, establishing principles for aligning and fusing asynchronous sequential data streams.
- Book: Multimodal Deep Learning, Cem Akkus et al. (2023). This text provides broad foundational coverage of multimodal deep learning architectures and representation alignment methods essential for understanding multimodal sequence modeling.
- Paper: Position: What Can Large Language Models Tell Us about Time Series Analysis, Ming Jin et al. (2024). This position paper synthesizes the broader paradigms of using large language models as predictors and multimodal enhancers for time series, contextualizing and extending the specific prompt-based framework of GPT4MTS.
- Paper: Timer: Generative Pre-trained Transformers Are Large Time Series Models, Yong Liu et al. (2024). Timer advances large-scale generative pre-training for time series forecasting, generalizing the LLM adaptation concepts explored in prompt-based multimodal forecasters.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). MOIRAI develops a universal time series foundation model across heterogeneous variates and frequencies, offering a scalable zero-shot alternative to prompt-based LLM adaptation.
- Paper: Sundial: A Family of Highly Capable Time Series Foundation Models, Yong Liu 0007 et al. (2025). Sundial expands foundation modeling for time series by incorporating continuous flow matching for probabilistic forecasting, moving beyond discrete prompt tokenization.
