GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting

Furong JiaKevin WangYixiang ZhengDefu CaoYan Liu

article2024AAAI177 citations

Proposes a prompt-tuning framework alongside an automated text-collection pipeline to integrate textual news summaries with numerical data for multimodal time-series forecasting.

Listen

Conventional time-series forecasting relies almost entirely on historical numerical data to predict future trends. While numerical metrics capture underlying statistical patterns, they fail to reflect the real-world contextual events, public sentiment, and news developments that actively drive those metrics. Integrating textual context into predictive modeling has remained difficult due to the lack of structured multimodal datasets and the challenge of effectively combining text and numerical streams without distorting forecasts.

The main objective of the article is to establish an end-to-end pipeline for constructing multimodal time-series datasets using large language models and to demonstrate a framework, named GPT4MTS, that uses textual information as prompts to improve time-series forecasting accuracy.

To evaluate this approach, the researchers built a dataset using media coverage records from the Global Database of Events, Language, and Tone (GDELT), spanning August 2022 to July 2023 across 55 United States regions and national data. The dataset paired daily event metrics—such as the number of mentions, articles, and sources across 10 major event categories—with automated text summaries generated and filtered through language models. The GPT4MTS framework processes numerical data into structured patches while converting relevant news summaries into soft prompt embeddings. These prompts guide a pre-trained language model whose internal attention layers remain frozen to maintain computational efficiency, and predictions are extracted strictly from the processed numerical representations.

The evaluation yielded several key findings regarding predictive performance and architectural design. First, GPT4MTS achieved superior overall accuracy across the 10 event categories compared to state-of-the-art transformer baselines, linear models, and text-free language model adaptations. Compared to the leading unimodal language model baseline (GPT4TS), the proposed approach achieved a 4.14% reduction in Mean Squared Error (MSE) and a 1.0% reduction in Mean Absolute Error (MAE). Second, ablation experiments showed that passing only the numerical output states to the final prediction layer is critical; including textual representations in the final layer degraded accuracy, proving that text should act strictly as a guiding prompt rather than a direct numerical predictor. Third, among unimodal baselines relying solely on numerical data, simpler linear models consistently outperformed complex transformer architectures, which tended to over-complicate numerical sequences.

These findings demonstrate that combining qualitative narrative context with quantitative metrics provides a practical, low-overhead way to enhance forecast reliability. Freezing core language model parameters allows organizations to leverage pre-trained language intelligence without prohibitive retraining costs. This multimodal approach makes complex trend tracking more interpretable and accessible for decision-makers in public policy, communication, and finance who need to anticipate public attention and event impacts.

Organizations and researchers seeking to improve trend forecasting should consider piloting multimodal ingestion pipelines that pair core numerical metrics with automated text summarization. Technical teams implementing such architectures must ensure text features serve strictly as guidance prompts rather than direct output inputs to prevent forecast degradation. Future work should expand the pipeline to additional domains, explore newer foundation models, and conduct pilot deployments in live decision-making environments.

Confidence in these findings is supported by consistent testing across more than 480,000 regional and national data splits. However, readers should note certain limitations: the empirical evaluation is confined to GDELT news tracking data over a single one-year period, and budget constraints necessitated using alternative summarization tools for select regional data, requiring title substitutions for lower-quality summaries. Further validation across broader industries and longer historical horizons remains necessary before full-scale commercial or operational deployment.

  • Paper: Efficient Multimodal Fusion via Interactive Prompting, Yaowei Li et al. (2023). This paper establishes prompt-based multimodal fusion paradigms for integrating disparate modalities, providing foundational prompting techniques adapted by GPT4MTS to fuse numerical time-series and textual context.
  • Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). This survey provides a comprehensive analysis of adapting Transformer architectures to time-series forecasting, establishing the baseline concepts and structural challenges addressed when prompting language models with sequential data.
  • Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). This work introduces patching and channel-independence tokenization for time-series Transformers, which directly influenced tokenization and prompt-formatting strategies for numerical series in LLMs.
  • Paper: Ask Me Anything: A simple strategy for prompting language models, Simran Arora et al. (2023). This study analyzes prompt strategies and aggregation mechanisms for foundation models, offering prerequisite insight into designing effective prompt structures for complex analytical tasks.
  • Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This foundational paper presents cross-modal attention mechanisms for unaligned multimodal sequences, establishing principles for aligning and fusing asynchronous sequential data streams.
  • Book: Multimodal Deep Learning, Cem Akkus et al. (2023). This text provides broad foundational coverage of multimodal deep learning architectures and representation alignment methods essential for understanding multimodal sequence modeling.
Cover for GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting

Abstract

Time series forecasting is an essential area of machine learning with a wide range of real-world applications. Most of the previous forecasting models aim to capture dynamic characteristics from uni-modal numerical historical data. Although extra knowledge can boost the time series forecasting performance, it is hard to collect such information. In addition, how to fuse the multimodal information is non-trivial. In this paper, we first propose a general principle of collecting the corresponding textual information from different data sources with the help of modern large language models (LLM). Then, we propose a prompt-based LLM framework to utilize both the numerical data and the textual information simultaneously, named GPT4MTS. In practice, we propose a GDELT-based multimodal time series dataset for news impact forecasting, which provides a concise and well-structured version of time series dataset with textual information for further research in communication. Through extensive experiments, we demonstrate the effectiveness of our proposed method on forecasting tasks with extra-textual information.

Table of Contents

  • Introduction
  • Related Works
  • Large-Scale Language Model (LLM)
  • Time Series Data Collection
  • Prompting
  • Time Series Forecasting
  • LLM4TS
  • Methods
  • Dataset Collection
  • Problem Definition
  • Proposed Method
  • Experiments
  • Baselines and Experimental Settings
  • Main Results
  • Ablation Study
  • Analysis
  • Relevance for Accessibility in Communication
  • Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — GPT4MTS multimodal prompt-based forecaster

    model/method

    GPT4MTS is a time-series forecasting model that uses a pretrained GPT-2 transformer to process numerical history together with textual context. For each time step in a look-back window, the textual summary is converted into an embedding with BERT, followed by a linear layer and ReLU, and used as a soft prompt prepended to the numerical time-series tokens. The numerical inputs undergo reversible instance normalization, patching of adjacent timestamps, channel-independent processing of each scalar series, and linear projection into the GPT-2 hidden dimension. The resulting text and numerical tokens are combined with positional embeddings and passed through the GPT-2 transformer blocks.

    The hidden states corresponding only to the numerical time-series tokens are passed to a linear output layer that produces the requested forecasting horizon. Thus, textual embeddings guide the transformer while the prediction representation is taken from the numerical tokens rather than directly from the text tokens.

  2. Knowl 2 — LLM-assisted multimodal dataset-generation pipeline

    model/method

    The paper proposes a general pipeline for augmenting numerical time-series datasets with aligned textual information. For each numerical series, the pipeline gathers candidate documents from an external source, summarizes the documents, constructs a hypothetical document describing the relevant event or category, ranks the candidate summaries by similarity to that hypothetical document, and combines the most relevant summaries into a compact text aligned with the corresponding time and entity.

    The resulting training examples contain both a numerical observation and an associated text summary, enabling a forecasting model to use contextual information that is not represented in the numerical values alone. The authors state that the procedure can be adapted to domains such as communication and finance when suitable domain-specific guidance and textual sources are available.

  3. Knowl 3 — GDELT multimodal news-impact dataset

    data/table

    The paper constructs a multimodal forecasting dataset from the Global Database of Events, Language, and Tone (GDELT). Numerical data are grouped by geographic region and EventRootCode, with the three forecasting variables being NumMentions, NumArticles, and NumSources. These variables respectively count event mentions, related news articles, and news sources within a given time period and region, serving as measures of the attention received by an event type.

    The dataset contains the ten selected event root types and their names: 01 Make Public Statement, 02 Appeal, 03 Express Intent to Cooperate, 04 Consult, 05 Engage in Diplomatic Cooperation, 07 Provide Aid, 08 Yield, 11 Disapprove, 17 Coerce, and 19 Fight. Data are collected for 55 regions in the United States together with national U.S. data, covering 2022-08-17 through 2023-07-31. Each numerical time series is paired with textual summaries of news associated with the same event type, region, and date.

  4. Knowl 4 — Textual-summary construction for GDELT observations

    algorithm

    For each event type, geographic region, and date, the dataset-generation procedure performs the following operations:

    1. Scrape the first 10 available news articles.
    2. Generate a summary for every scraped article with T5.
    3. Generate a hypothetical article summary from the event type and its explanation; this serves as a relevance template.
    4. Compute the similarity between each article summary and the hypothetical summary and rank the summaries by similarity.
    5. Select the five most relevant summaries.
    6. Generate one overall summary from those five summaries with the OpenAI ChatGPT-3.5 API.

    Because of budget constraints, some regional summaries use T5 rather than ChatGPT-3.5 for the final summarization step. When such a summary is judged inferior because irrelevant information contaminated the first-stage summaries, it is replaced with a summary made from clean scraped titles.

  5. Knowl 5 — Multimodal forecasting task formulation

    equation

    The forecasting task uses a look-back window of length LL. At time tt, the numerical observation is a vector xt∈RMx_t \in R^M, where MM is the number of numerical variables, and sts_t is the textual summary aligned with that observation. Given the multimodal history ((x1,s1),…,(xL,sL))((x_1,s_1),\ldots,(x_L,s_L)), the goal is to forecast the next TT numerical vectors (xL+1,…,xL+T)(x_{L+1},\ldots,x_{L+T}).

    A numerical-only forecaster has the form

    (xL+1,…,xL+T)=funi(x1,…,xL;θ),(x_{L+1},\ldots,x_{L+T}) = f_{\mathrm{uni}}(x_1,\ldots,x_L;\theta),

    where funif_{\mathrm{uni}} is a forecasting function and θ\theta denotes its parameters. GPT4MTS instead learns a multimodal function

    (xL+1,…,xL+T)=fmulti((x1,s1),…,(xL,sL);θ),(x_{L+1},\ldots,x_{L+T}) = f_{\mathrm{multi}}((x_1,s_1),\ldots,(x_L,s_L);\theta),

    so that textual summaries can influence forecasts while the prediction targets remain numerical.

  6. Knowl 6 — Parameter-efficient use of pretrained GPT-2

    model/method

    GPT4MTS retains the pretrained GPT-2 positional-embedding mechanism and transformer blocks but does not fully fine-tune the language model. The attention layers and feed-forward layers are frozen. Positional embeddings and layer-normalization layers are fine-tuned, while additional input-embedding and output layers adapt the numerical and textual modalities to the GPT-2 hidden space.

    This design treats the pretrained transformer as a largely fixed sequence-processing backbone and uses trainable prompt and projection components to incorporate time-series information. Freezing the attention and feed-forward layers is intended to reduce training and inference cost.

  7. Knowl 7 — Experimental comparison protocol

    experimental setup

    The evaluation compares GPT4MTS with the Transformer-based forecasters FEDformer, Autoformer, Informer, PatchTST, and Transformer; the linear models DLinear and NLinear; and pretrained-language-model baselines LLaMA and GPT4TS adapted for time-series forecasting. Every model uses a prediction length of T=7T=7, a look-back window of L=15L=15, and forecasts the three numerical variables NumMentions, NumArticles, and NumSources.

    The data are split into training, validation, and test sets in a 7:2:1 ratio, producing 343,200 training examples, 98,976 validation examples, and 41,904 test examples. Results are reported by event type using mean squared error (MSE) and mean absolute error (MAE).

  8. Knowl 8 — Forecasting performance on the GDELT dataset

    data/table

    The following results compare all evaluated models on the ten GDELT event types. Lower MSE and MAE are better. The table demonstrates that GPT4MTS has the best average performance for both metrics and is competitive with or better than the other models for most individual event types.

    Could not parse LaTeX table

    Relative to GPT4TS, GPT4MTS reduces average MSE from 0.362 to 0.347, a 4.14% reduction, and average MAE from 0.398 to 0.394, a 1.0% reduction.

  9. Knowl 9 — Ablation of textual and numerical hidden-state selection

    empirical result

    GPT4MTS was evaluated with two output representations: the proposed selection, which uses only hidden states corresponding to numerical time-series tokens, and full selection, which also passes hidden states corresponding to textual prompt tokens to the output layer. The proposed selection achieves average MSE 0.347 and average MAE 0.394. Full selection achieves average MSE 0.360 and average MAE 0.402, while the text-free GPT4TS baseline achieves average MSE 0.362 and average MAE 0.398.

    The ablation shows that adding textual information improves forecasting relative to the text-free baseline, but using text-token hidden states directly in the output representation can dilute the numerical forecasting signal. Text is most effective when it acts as contextual guidance for the numerical representations rather than as an equally weighted source for the final prediction.

  10. Knowl 10 — Roles of textual context and numerical history

    empirical result

    The paper attributes GPT4MTS's improvement to complementary information from the two modalities. Textual summaries provide contextual cues that numerical history cannot directly encode, such as the prominence, sentiment, or distinctive characteristics of news events. Numerical observations provide precise quantitative values and expose the statistical patterns of the forecasting targets, so they remain the primary driver of the final prediction.

    The comparison also indicates that prompt-based integration is important: pretrained language-model baselines adapted to numerical inputs perform better than several conventional Transformer and linear baselines, while GPT4MTS improves further by adding textual prompts. The ablation suggests that excessive reliance on textual representations can reduce performance, so the text should guide rather than replace the numerical forecasting representation.

  11. Knowl 11 — Stated limitations of the data-construction pipeline

    limitation

    The proposed textual-data pipeline is not claimed to apply automatically to every forecasting domain; it requires suitable external documents, event or category descriptions, and domain-specific guidance. The GDELT construction also uses a lower-cost T5 fallback for some regional final summaries instead of ChatGPT-3.5. The paper reports that this fallback can produce poorly summarized text when irrelevant information survives the first summarization stage, motivating replacement with summaries of clean scraped titles.

Coverage note — The paper's accessibility implications and future-work discussion were omitted because they are applications and research directions rather than additional load-bearing method, dataset, theory, or experimental contributions.

References

  1. 1.Box, G. E.; and Jenkins, G. M. 1968. Some recent advances in forecasting and control. Journal of the Royal Statistical Society. Series C (Applied Statistics), 17(2): 91–109.
  2. 2.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  3. 3.Cao, D.; Enouen, J.; Wang, Y.; Song, X.; Meng, C.; Niu, H.; and Liu, Y. 2023. Estimating Treatment Effects from Irregular Time Series Observations with Hidden Confounders. In Williams, B.; Chen, Y.; and Neville, J., eds., Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, 6897–6905. AAAI Press.
  4. 4.Cao, D.; Wang, Y.; Duan, J.; Zhang, C.; Zhu, X.; Huang, C.; Tong, Y.; Xu, B.; Bai, J.; Tong, J.; and Zhang, Q. 2020. Spectral Temporal Graph Neural Network for Multivariate Time-series Forecasting. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  5. 5.Consoli, S.; Pezzoli, L. T.; and Tosetti, E. 2020. Using the GDELT Dataset to Analyse the Italian Sovereign Bond Market. In Machine Learning, Optimization, and Data Science: 6th International Conference, LOD 2020, Siena, Italy, July 19–23, 2020, Revised Selected Papers, Part I 6, 190–202. Springer.
  6. 6.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), 4171–4186. Association for Computational Linguistics.
  7. 7.Galla, D.; and Burke, J. 2018. Predicting social unrest using GDELT. In International conference on machine learning and data mining in pattern recognition, 103–116. Springer.
  8. 8.Hopp, F. R.; Schaffer, J.; Fisher, J. T.; and Weber, R. 2019. iCoRe: The GDELT interface for the advancement of communication research. Computational Communication Research, 1(1): 13–44.
  9. 9.Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, 2790–2799. PMLR.
  10. 10.Jiang, W.; and Luo, J. 2022. Graph neural network for traffic forecasting: A survey. Expert Systems with Applications, 207: 117921.
  11. 11.Jin, M.; Wang, S.; Ma, L.; Chu, Z.; Zhang, J. Y.; Shi, X.; Chen, P.; Liang, Y.; Li, Y.; Pan, S.; and Wen, Q. 2023. TimeLLM: Time Series Forecasting by Reprogramming Large Language Models. CoRR, abs/2310.01728.
  12. 12.Kalamara, E.; Turrell, A.; Redl, C.; Kapetanios, G.; and Kapadia, S. 2022. Making text count: economic forecasting using newspaper text. Journal of Applied Econometrics, 37(5): 896–919.
  13. 13.Kaushik, S.; Choudhury, A.; Sheron, P. K.; Dasgupta, N.; Natarajan, S.; Pickett, L. A.; and Dutt, V. 2020. AI in healthcare: time-series forecasting using statistical, neural, and ensemble architectures. Frontiers in big data, 3: 4.
  14. 14.Kim, T.; Kim, J.; Tae, Y.; Park, C.; Choi, J.-H.; and Choo, J. 2021. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations.
  15. 15.Li, J.; Liu, C.; Cheng, S.; Arcucci, R.; and Hong, S. 2023. Frozen Language Model Helps ECG Zero-Shot Learning. CoRR, abs/2303.12311.
  16. 16.Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. 2022. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10965–10975.
  17. 17.Liu, S.; Yu, H.; Liao, C.; Li, J.; Lin, W.; Liu, A. X.; and Dustdar, S. 2021. Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting. In International conference on learning representations.
  18. 18.Lock, I. 2020. Debating glyphosate: A macro perspective on the role of strategic communication in forming and monitoring a global issue arena using inductive topic modelling. International Journal of Strategic Communication, 14(4): 223–245.
  19. 19.Lu, K.; Grover, A.; Abbeel, P.; and Mordatch, I. 2022. Frozen pretrained transformers as universal computation engines. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 7628–7636.
  20. 20.Munir, M.; Siddiqui, S. A.; Dengel, A.; and Ahmed, S. 2018. DeepAnT: A deep learning approach for unsupervised anomaly detection in time series. Ieee Access, 7: 1991–2005.
  21. 21.Nguyen, T.; Brandstetter, J.; Kapoor, A.; Gupta, J. K.; and Grover, A. 2023. ClimaX: A foundation model for weather and climate. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, 25904–25938. PMLR.
  22. 22.Nie, Y.; Nguyen, N. H.; Sinthong, P.; and Kalagnanam, J. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  23. 23.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748–8763. PMLR.
  24. 24.Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018. Improving language understanding by generative pre-training.
  25. 25.Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9.
  26. 26.Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1): 5485–5551.
  27. 27.Salinas, D.; Flunkert, V.; Gasthaus, J.; and Januschowski, T. 2020. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3): 1181–1191.
  28. 28.Schintler, L. A.; and Kulkarni, R. 2014. Big data for policy analysis: The good, the bad, and the ugly. Review of Policy Research, 31(4): 343–348.
  29. 29.Sezer, O. B.; Gudelek, M. U.; and Ozbayoglu, A. M. 2020. Financial time series forecasting with deep learning: A systematic literature review: 2005–2019. Applied soft computing, 90: 106181.
  30. 30.Sun, C.; Li, Y.; Li, H.; and Hong, S. 2023. TEST: Text Prototype Aligned Embedding to Activate LLM’s Ability for Time Series. arXiv preprint arXiv:2308.08241.
  31. 31.Tian Zhou, X. W. L. S. R. J., Peisong Niu. 2023. One Fits All: Power General Time Series Analysis by Pretrained LM. In NeurIPS.
  32. 32.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; Azhar, `F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023. LLaMA: Open and Efficient Foundation Language Models. CoRR, abs/2302.13971.
  33. 33.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  34. 34.Wu, H.; Xu, J.; Wang, J.; and Long, M. 2021. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems, 34: 22419–22430.
  35. 35.Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Chang, X.; and Zhang, C. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, 753–763.
  36. 36.Xue, H.; and Salim, F. D. 2022. Prompt-Based Time Series Forecasting: A New Task and Dataset. arXiv preprint arXiv:2210.08964.
  37. 37.Yu, X.; Chen, Z.; Ling, Y.; Dong, S.; Liu, Z.; and Lu, Y. 2023. Temporal Data Meets LLM–Explainable Financial Time Series Forecasting. arXiv preprint arXiv:2306.11025.
  38. 38.Zeng, A.; Chen, M.; Zhang, L.; and Xu, Q. 2023. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, 11121–11128.
  39. 39.Zhang, H.; Zhang, P.; Hu, X.; Chen, Y.-C.; Li, L.; Dai, X.; Wang, L.; Yuan, L.; Hwang, J.-N.; and Gao, J. 2022. Glipv2: Unifying localization and vision-language understanding. Advances in Neural Information Processing Systems, 35: 36067–36080.
  40. 40.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, 11106–11115.
  41. 41.Zhou, T.; Ma, Z.; Wen, Q.; Wang, X.; Sun, L.; and Jin, R. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International Conference on Machine Learning, 27268–27286. PMLR.

Citation

MLA
Jia, F., et al. “GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 21, 2024, pp. 23343–51, https://doi.org/10.1609/AAAI.V38I21.30383.
APA
Jia, F., Wang, K., Zheng, Y., Cao, D., & Liu, Y. (2024). GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, 38(21), 23343–23351. https://doi.org/10.1609/AAAI.V38I21.30383
Chicago
Jia, F., K. Wang, Y. Zheng, D. Cao, and Y. Liu. 2024. “GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence 38 (21): 23343–51. https://doi.org/10.1609/AAAI.V38I21.30383.
Harvard
Jia, F. et al. (2024) “GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting”, Proceedings of the AAAI Conference on Artificial Intelligence, 38(21), pp. 23343–23351. Available at: https://doi.org/10.1609/AAAI.V38I21.30383.
Vancouver
1. Jia F, Wang K, Zheng Y, Cao D, Liu Y (2024) GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 38:23343–23351

BibTeX

@article{Jia_2024, title={GPT4MTS: Prompt-based Large Language Model for Multimodal Time-series Forecasting}, volume={38}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V38I21.30383}, DOI={10.1609/aaai.v38i21.30383}, number={21}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Jia, Furong and Wang, Kevin and Zheng, Yixiang and Cao, Defu and Liu, Yan}, year={2024}, month=Mar, pages={23343–23351} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF