Position: What Can Large Language Models Tell Us about Time Series Analysis

Ming JinYifan ZhangWei ChenKexin ZhangYuxuan LiangBin YangJindong WangShirui PanQingsong Wen

article2024ICML66 citations

Categorizes the emerging roles of large language models in time series analysis as data enhancers, predictors, and autonomous agents while identifying concrete integration strategies and open research opportunities for building universal time series intelligence.

Listen

Time series analysis is vital for understanding complex dynamic systems across industries such as finance, transportation, and healthcare. However, traditional statistical and deep learning models are typically narrow, task-specific, and heavily dependent on domain knowledge and extensive tuning. The article examines whether large language models (LLMs) can bridge the gap toward general-purpose analytical intelligence, demonstrating how LLMs can transform time series analysis from isolated forecasting tools into unified, interactive, and reasoning-driven systems.

To evaluate this potential, the article conducts a comprehensive literature review and exploratory empirical tests. It assesses three operational roles for LLMs: data and model enhancers, direct predictors, and autonomous analytical agents. For the empirical component, the authors test zero-shot prompt-based reasoning using GPT-3.5 on public benchmarks, specifically examining activity classification using waist-mounted smartphone sensor data from 30 participants and electric transformer temperature records.

The findings indicate that LLMs can impact time series analysis in three major ways. First, as data and model enhancers, LLMs successfully enrich sparse numerical data with textual context and transfer external knowledge to specialized domain models. Second, as direct predictors, LLMs demonstrate competitive zero-shot and few-shot forecasting performance, either through input reprogramming and soft prompts or via specialized continuous tokenization that allows frozen models to match domain-specific baselines. Third, in agent-based evaluation, the model achieves complete accuracy on simple movement patterns like standing in zero-shot activity classification, while providing clear, natural language explanations of its reasoning. However, the evaluation also shows that LLMs struggle with subtle, complex numerical patterns, occasionally hallucinate false data or explanations, and exhibit significant task bias, such as misclassifying laying down as sitting or standing.

These results imply that integrating language models into time series workflows can substantially enhance interpretability, reduce the resources required to build models from scratch, and enable multimodal problem solving. Nevertheless, the presence of hallucinations, privacy risks associated with sensitive industrial telemetry, and high computational costs introduce operational and compliance risks if models are deployed without domain guardrails.

To safely capitalize on these capabilities, organizations should pursue hybrid integration frameworks rather than relying on stand-alone language models. Recommended approaches include aligning time series embeddings with language representations, developing multi-agent systems where specialized models perform low-level pattern analysis while language models orchestrate planning and communication, and establishing robust prompt guidelines. Decision-makers should validate these tools through controlled pilot programs before full deployment. While the underlying literature review is solid, confidence in fully autonomous zero-shot LLM agents remains low, warranting caution due to persistent hallucination and data drift challenges.

arXiv: 2402.02713
Cover for Position: What Can Large Language Models Tell Us about Time Series Analysis

Abstract

Time series analysis is essential for comprehending the complexities inherent in various real-world systems and applications. Although large language models (LLMs) have recently made significant strides, the development of artificial general intelligence (AGI) equipped with time series analysis capabilities remains in its nascent phase. Most existing time series models heavily rely on domain knowledge and extensive model tuning, predominantly focusing on prediction tasks. In this paper, we argue that current LLMs have the potential to revolutionize time series analysis, thereby promoting efficient decision-making and advancing towards a more universal form of time series analytical intelligence. Such advancement could unlock a wide range of possibilities, including time series modality switching and question answering. We encourage researchers and practitioners to recognize the potential of LLMs in advancing time series analysis and emphasize the need for trust in these related efforts. Furthermore, we detail the seamless integration of time series analysis with existing LLM technologies and outline promising avenues for future research.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Time Series Analysis
  • 2.2. Large Language Models
  • 2.3. Research Roadmap
  • 3. LLM-assisted Enhancer for Time Series
  • 3.1. Data-based Enhancer
  • 3.2. Model-based Enhancer
  • 3.3. Discussion
  • 4. LLM-centered Predictor for Time Series
  • 4.1. Tuning-based Predictor
  • 4.2. Non-tuning-based Predictor
  • 4.3. Others
  • 4.4. Discussion
  • 5. LLM-empowered Agent for Time Series
  • 5.1. Empirical Insights: LLMs as Time Series Analysts
  • 5.2. Key Lessons for Advancing Time Series Agents
  • 5.3. Exploring Alternative Research Avenues
  • 6. Further Discussion
  • 7. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Literature Review
  • B. LLM-empowered Agent for Time Series
  • B.1. Overview of Related Works
  • B.2. Demonstrations

Knowls

  1. Knowl 1 — Taxonomy of Integration Paradigms for LLMs in Time Series Analysis

    definition

    The integration of Large Language Models (LLMs) with time series analysis is categorized into three distinct paradigms based on functional scope and capability boundaries:

    1. LLM-assisted Enhancer: LLMs serve as auxiliary modules to augment time series data (e.g., generating textual descriptions, summarization, metadata enrichment) or enhance existing domain-specific models (e.g., cross-modal feature alignment, knowledge transfer, dual-tower prompting) to improve robustness and interpretability while retaining specialized downstream architectures.

    2. LLM-centered Predictor: LLMs act as the core predictive backbone for classic time series tasks (such as forecasting, classification, anomaly detection, and imputation). This includes tuning-based methods (fine-tuning LLMs with lightweight adaptation layers or reprogramming) and non-tuning-based methods (zero-shot prompting and specialized numerical tokenization on frozen models).

    3. LLM-empowered Agent: LLMs function as general-purpose interactive systems capable of multi-step reasoning, tool orchestration, and question answering across diverse time series workflows.

  2. Knowl 2 — Four Generations of Time Series Analytical Models

    definition

    The development of time series analytical modeling is conceptualized into four chronological generations defined by task generality and architectural capabilities:

    1. Statistical Time Series Models (1950s–2000s): Methods such as ARIMA and Holt-Winters exponential smoothing, optimized for small-scale datasets and grounded in statistical heuristics such as stationarity and seasonality.

    2. Neural Time Series Models (2010s): Deep neural architectures such as Recurrent Neural Networks (RNNs), Temporal Convolutional Networks (TCNs), and Spatio-Temporal Graph Neural Networks (STGNNs) that capture non-linear, long-term dependencies from larger datasets without explicit statistical assumptions.

    3. Pre-trained Time Series Models (2022): Domain- and task-agnostic architectures pre-trained on diverse time series corpora (such as TF-C and TimeCLR) and fine-tuned for specific tasks with limited samples.

    4. LLM-Centric Time Series Models (2024+): Universal task solvers that place LLMs as the central engine to process multimodal inputs (time series signals alongside textual prompts) to solve both predictive tasks (forecasting, classification) and cognitive tasks (reasoning, Q&A, action planning).

  3. Knowl 3 — Mathematical Formulation of Tuning-Based LLM Predictors

    equation

    In tuning-based LLM predictors for time series analysis, accessible language model parameters fLLMrianglef^{ riangle}_{\text{LLM}} (or attached lightweight adaptation layers) are updated on tokenized time series representations and optional textual contexts:

    Xinp=Patching(X),Tinp=Tokenizer(T),Y^=Task(fLLM△(Xinp,Tinp,P))\begin{aligned} X_{\text{inp}} &= \text{Patching}(\mathcal{X}), \\ T_{\text{inp}} &= \text{Tokenizer}(\mathcal{T}), \\ \hat{Y} &= \text{Task}\left(f^{\triangle}_{\text{LLM}}(X_{\text{inp}}, T_{\text{inp}}, P)\right) \end{aligned}

    where X\mathcal{X} denotes the input numerical time series, T\mathcal{T} denotes optional associated textual metadata, Patching(⋅)\text{Patching}(\cdot) chunks continuous numerical sequences into patch tokens XinpX_{\text{inp}}, Tokenizer(⋅)\text{Tokenizer}(\cdot) maps text sequences to token embeddings TinpT_{\text{inp}}, PP is an instruction prompt, and Task(⋅)\text{Task}(\cdot) represents an output task head projecting LLM hidden representations to target predictions Y^\hat{Y}.

  4. Knowl 4 — Mathematical Formulation of Non-Tuning-Based LLM Predictors

    equation

    In non-tuning-based LLM predictors, a black-box LLM fLLM▲f^{\blacktriangle}_{\text{LLM}} is queried without updating internal model weights, relying on prompt templating or custom numerical tokenization:

    Xinp=Template(X,P)orXinp=Tokenizer(X),Y^=Parse(fLLM▲(Xinp))\begin{aligned} X_{\text{inp}} &= \text{Template}(\mathcal{X}, P) \quad \text{or} \quad X_{\text{inp}} = \text{Tokenizer}(\mathcal{X}), \\ \hat{Y} &= \text{Parse}\left(f^{\blacktriangle}_{\text{LLM}}(X_{\text{inp}})\right) \end{aligned}

    where X\mathcal{X} denotes the raw input time series data, PP is the task instruction prompt, Template(⋅)\text{Template}(\cdot) constructs a textual prompt embedding numerical values, Tokenizer(⋅)\text{Tokenizer}(\cdot) converts continuous values into token IDs formatted for the model, and Parse(⋅)\text{Parse}(\cdot) extracts structured task predictions Y^\hat{Y} from the natural language response generated by fLLM▲f^{\blacktriangle}_{\text{LLM}}.

  5. Knowl 5 — Paradigms for Incorporating Time Series Knowledge into LLM Agents

    model/method

    To enable LLMs to serve as general-purpose agents capable of handling complex temporal dynamics, three distinct architectural paradigms are defined for integrating time series knowledge:

    1. Feature Alignment: A dedicated time series encoder extracts temporal representations from raw signals and explicitly projects them into the continuous embedding space of a language model, aligning numerical features directly with language tokens.

    2. Text-Temporal Feature Fusion: A fusion module (such as a Fusion-Former utilizing learnable query embeddings) interacts jointly with text embeddings and encoded time series features to produce a unified cross-modal representation before feeding into the LLM.

    3. Tool Utilization via Time Series Model Hub: The LLM operates as a high-level orchestrator equipped with a Prompt Manager and Tool Manager. Rather than processing raw signals internally, it parses user queries and iteratively executes external specialized pre-trained models (such as forecasters, classifiers, anomaly detectors, or imputers) from a model hub, aggregating their results into natural language responses.

  6. Knowl 6 — Zero-Shot Inertial Human Activity Recognition Setup for LLM Agents

    experimental setup

    To evaluate the zero-shot reasoning and classification capabilities of LLMs acting as autonomous analytical agents, an experimental benchmark was conducted using the Human Activity Recognition (HAR) dataset:

    • Data Source: Inertial sensor readings collected from waist-mounted smartphones carried by 30 subjects performing activities of daily living.
    • Input Features: Triaxial total acceleration, estimated triaxial body acceleration, and triaxial angular velocity from gyroscopes.
    • Classification Target: Four activity categories: Stand, Sit, Lay, and Walk.
    • Evaluation Sample: 10 instances per class (40 total instances).
    • Agent Interaction: GPT-3.5 is presented with formatted numerical feature values alongside task descriptions, few-shot examples, and requests for natural language justifications and confidence estimations.
  7. Knowl 7 — Classification Distribution and Bias in Zero-Shot LLM Activity Recognition

    data/table

    When evaluated on zero-shot activity classification across 40 instances (10 per class) from the HAR dataset using GPT-3.5 with few-shot prompts, the model yielded the following confusion matrix:

    Ground-Truth \ redicted Stand Sit Lay Walk
    Stand 10 0 0 0
    Sit 3 7 0 0
    Lay 4 6 0 0
    Walk 1 1 2 6

    The results demonstrate that while the LLM successfully identifies common static postures (Stand: 10/10 correct; Sit: 7/10 correct), it exhibits strong bias and fails completely on Lay (0/10 correct), consistently misclassifying Lay instances as Sit (6 instances) or Stand (4 instances).

  8. Knowl 8 — Qualitative Behaviors of LLMs in Time Series Analytical Tasks

    empirical result

    Empirical case studies evaluating ChatGPT on Human Activity Recognition (HAR) classification and Electric Transformer Temperature (ETT) anomaly detection and augmentation reveal several qualitative behaviors:

    1. Common-Sense Physical Interpretability: The model accurately articulates physical intuitions in natural language (e.g., attributing walking to rhythmic/repetitive acceleration and sitting to lower overall movement variance).

    2. Plausible Explanatory Hallucination: When queried about incorrect classifications (e.g., justifying why Lay instances were classified as Sit or Stand), the LLM invents plausible-sounding but factually fabricated justifications regarding sensor variability.

    3. Superficial Data Augmentation: When asked to synthesize new time series instances adhering to ETT dataset patterns, the model reproduces copies of input instances rather than generating novel synthetic samples matching underlying distribution statistics.

    4. Refusal and Uncertainty Calibration: When asked to classify without adequate context, the model initially refuses before being prompted to guess. When asked for confidence scores, it assigns low values (e.g., 0.3 on a 0 to 1 scale), appropriately noting the lack of access to trained parameters or exact distributions.

  9. Knowl 9 — Core Challenges and Limitations in LLM-Centric Time Series Analysis

    limitation

    Deploying LLMs for time series analysis faces several critical open challenges:

    1. Catastrophic Forgetting & Parameter Tuning Cost: Direct fine-tuning of LLMs on time series data risks degrading foundational linguistic reasoning while incurring substantial computational overhead.

    2. Tokenization Mismatch & Prompt Fragility: Off-the-shelf LLM tokenizers are optimized for natural language text rather than continuous numerical values, breaking temporal ordering and numerical precision unless specialized continuous embedding schemes are employed.

    3. Hallucination in Complex Numerical Dynamics: LLMs struggle with high-order temporal dependencies and fine-grained numerical features, leading to fabricated trend explanations or incorrect pattern recognition.

    4. Concept Drift & Non-Stationarity: Real-world streaming time series frequently undergo distribution shifts over time, which static pre-trained LLMs cannot track without costly continual learning or dynamic memory mechanisms.

    5. Privacy and Inference Overhead: High latency in autoregressive token generation poses barriers for real-time applications, and training/inference on sensitive industrial telemetry raises data privacy concerns.

Coverage note — None was omitted; all substantive taxonomy frameworks, formal equations, empirical setups, confusion matrix data, and qualitative analysis findings from the paper are included as knowls.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Alghamdi, T., Elgazzar, K., Bayoumi, M., Sharaf, T., and Shah, S. Forecasting traffic congestion using arima modeling. In 2019 15th international wireless communications & mobile computing conference (IWCMC), pp. 1227–1232. IEEE, 2019.
  3. 3.Anguita, D., Ghio, A., Oneto, L., Parra, X., Reyes-Ortiz, J. L., et al. A public domain dataset for human activity recognition using smartphones. In Esann, volume 3, pp. 3, 2013.
  4. 4.Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  5. 5.Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Gianinazzi, L., Gajda, J., Lehmann, T., Podstawski, M., Niewiadomski, H., Nyczyk, P., et al. Graph of thoughts: Solving elaborate problems with large language models. arXiv preprint arXiv:2308.09687, 2023.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  7. 7.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  8. 8.Cao, D., Jia, F., Arik, S. O., Pfister, T., Zheng, Y., Ye, W., and Liu, Y. Tempo: Prompt-based generative pre-trained transformer for time series forecasting. In The Twelfth International Conference on Learning Representations, 2023.
  9. 9.Chang, C., Peng, W.-C., and Chen, T.-F. Llm4ts: Two-stage fine-tuning for time-series forecasting with pre-trained llms. arXiv preprint arXiv:2308.08469, 2023.
  10. 10.Chatterjee, S., Mitra, B., and Chakraborty, S. Amicron: A framework for generating annotations for human activity recognition with granular micro-activities. arXiv preprint arXiv:2306.13149, 2023.
  11. 11.Chen, S., Long, G., Shen, T., and Jiang, J. Prompt federated learning for weather forecasting: toward foundation models on meteorological data. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, pp. 3532–3540, 2023a.
  12. 12.Chen, Y., Wang, X., and Xu, G. Gatgpt: A pre-trained large language model with graph attention network for spatiotemporal imputation. arXiv preprint arXiv:2311.14332, 2023b.
  13. 13.Chen, Z., Mao, H., Li, H., Jin, W., Wen, H., Wei, X., Wang, S., Yin, D., Fan, W., Liu, H., et al. Exploring the potential of large language models (llms) in learning on graphs. ACM SIGKDD Explorations Newsletter, 25(2): 42–61, 2024.
  14. 14.Cheng, J. and Chin, P. Sociodojo: Building lifelong analytical agents with real-world text and time series. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=s9z0HzWJJp.
  15. 15.Cheng, Y., Zhang, C., Zhang, Z., Meng, X., Hong, S., Li, W., Wang, Z., Wang, Z., Yin, F., Zhao, J., et al. Exploring large language model based intelligent agents: Definitions, methods, and prospects. arXiv preprint arXiv:2401.03428, 2024.
  16. 16.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  17. 17.Da, L., Gao, M., Mei, H., and Wei, H. Llm powered sim-to-real transfer for traffic signal control. arXiv preprint arXiv:2308.14284, 2023.
  18. 18.Da, L., Liou, K., Chen, T., Zhou, X., Luo, X., Yang, Y., and Wei, H. Open-ti: Open traffic intelligence with augmented language model. International Journal of Machine Learning and Cybernetics, pp. 1–26, 2024.
  19. 19.Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting. In the 41st International Conference on Machine Learning, 2024.
  20. 20.Ekambaram, V., Jati, A., Nguyen, N. H., Dayama, P., Reddy, C., Gifford, W. M., and Kalagnanam, J. Ttms: Fast multi-level tiny time mixers for improved zero-shot and few-shot forecasting of multivariate time series. arXiv preprint arXiv:2401.03955, 2024.
  21. 21.Fatouros, G., Metaxas, K., Soldatos, J., and Kyriazis, D. Can large language models beat wall street? unveiling the potential of ai in stock selection. arXiv preprint arXiv:2401.03737, 2024.
  22. 22.Fuller, W. A. Introduction to statistical time series. John Wiley & Sons, 2009.
  23. 23.Gamboa, J. C. B. Deep learning for time-series analysis. arXiv preprint arXiv:1701.01887, 2017.
  24. 24.Garg, S., Farajtabar, M., Pouransari, H., Vemulapalli, R., Mehta, S., Tuzel, O., Shankar, V., and Faghri, F. Tic-clip: Continual training of clip models. In NeurIPS 2023 Workshop on Distribution Shifts: New Frontiers with Foundation Models, 2023.
  25. 25.Garza, A. and Mergenthaler-Canseco, M. Timegpt-1. arXiv preprint arXiv:2310.03589, 2023.
  26. 26.Ge, Y., Hua, W., Mei, K., Tan, J., Xu, S., Li, Z., Zhang, Y., et al. Openagi: When llm meets domain experts. Advances in Neural Information Processing Systems, 36, 2023.
  27. 27.Ghosh, S., Sengupta, S., and Mitra, P. Spatio-temporal storytelling? leveraging generative models for semantic trajectory analysis. arXiv preprint arXiv:2306.13905, 2023.
  28. 28.Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G. Large language models are zero-shot time series forecasters. Advances in neural information processing systems, 2023.
  29. 29.Gu, Z., Zhu, B., Zhu, G., Chen, Y., Tang, M., and Wang, J. Anomalygpt: Detecting industrial anomalies using large vision-language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 1932–1940, 2024.
  30. 30.Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717, 2023.
  31. 31.Hamilton, J. D. Time series analysis. Princeton university press, 2020.
  32. 32.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pp. 9118–9147. PMLR, 2022.
  33. 33.Jin, M., Koh, H. Y., Wen, Q., Zambon, D., Alippi, C., Webb, G. I., King, I., and Pan, S. A survey on graph neural networks for time series: Forecasting, classification, imputation, and anomaly detection. arXiv preprint arXiv:2307.03759, 2023a.
  34. 34.Jin, M., Wen, Q., Liang, Y., Zhang, C., Xue, S., Wang, X., Zhang, J., Wang, Y., Chen, H., Li, X., et al. Large models for time series and spatio-temporal data: A survey and outlook. arXiv preprint arXiv:2310.10196, 2023b.
  35. 35.Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., et al. Time-llm: Time series forecasting by reprogramming large language models. In International Conference on Machine Learning, 2024.
  36. 36.Kalekar, P. S. et al. Time series forecasting using holt-winters exponential smoothing. Kanwal Rekhi school of information Technology, 4329008(13):1–13, 2004.
  37. 37.Kamarthi, H. and Prakash, B. A. Pems: Pre-trained epidmic time-series models. arXiv preprint arXiv:2311.07841, 2023.
  38. 38.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  39. 39.Kim, Y., Xu, X., McDuff, D., Breazeal, C., and Park, H. W. Health-llm: Large language models for health prediction via wearable sensor data. arXiv preprint arXiv:2401.06866, 2024.
  40. 40.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 22199–22213, 2022.
  41. 41.Lai, S., Xu, Z., Zhang, W., Liu, H., and Xiong, H. Large language models as traffic signal control agents: Capacity and opportunity. arXiv preprint arXiv:2312.16044, 2023.
  42. 42.Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S. Aligning text-to-image models using human feedback. arXiv preprint arXiv:2302.12192, 2023.
  43. 43.Li, J., Cheng, X., Zhao, W. X., Nie, J.-Y., and Wen, J.-R. Helma: A large-scale hallucination evaluation benchmark for large language models. arXiv preprint arXiv:2305.11747, 2023.
  44. 44.Li, J., Liu, C., Cheng, S., Arcucci, R., and Hong, S. Frozen language model helps ecg zero-shot learning. In Medical Imaging with Deep Learning, pp. 402–415. PMLR, 2024.
  45. 45.Liang, Y., Liu, Y., Wang, X., and Zhao, Z. Exploring large language models for human mobility prediction under public events. arXiv preprint arXiv:2311.17351, 2023.
  46. 46.Liang, Y., Wen, H., Nie, Y., Jiang, Y., Jin, M., Song, D., Pan, S., and Wen, Q. Foundation models for time series analysis: A tutorial and survey. In ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD’24), 2024.
  47. 47.Liao, Q. V. and Vaughan, J. W. Ai transparency in the age of llms: A human-centered research roadmap. Harvard Data Science Review, 2024.
  48. 48.Liu, C., Ma, Y., Kothur, K., Nikpour, A., and Kavehei, O. Biosignal copilot: Leveraging the power of llms in drafting reports for biomedical signals. medRxiv, pp. 2023–06, 2023a.
  49. 49.Liu, C., Yang, S., Xu, Q., Li, Z., Long, C., Li, Z., and Zhao, R. Spatial-temporal large language model for traffic prediction. arXiv preprint arXiv:2401.10134, 2024a.
  50. 50.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. Advances in neural information processing systems, 2023b.
  51. 51.Liu, X., Hu, J., Li, Y., Diao, S., Liang, Y., Hooi, B., and Zimmermann, R. Unitime: A language-empowered unified model for cross-domain time series forecasting. In The Web Conference 2024 (WWW), 2024b.
  52. 52.Lopez-Lira, A. and Tang, Y. Can chatgpt forecast stock price movements? return predictability and large language models. Return Predictability and Large Language Models (April 6, 2023), 2023.
  53. 53.Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2024.
  54. 54.Mirchandani, S., Xia, F., Florence, P., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., Zeng, A., et al. Large language models as general pattern machines. In 7th Annual Conference on Robot Learning, 2023.
  55. 55.Moon, S., Madotto, A., Lin, Z., Saraf, A., Bearman, A., and Damavandi, B. Imu2clip: Language-grounded motion sensor translation with multimodal contrastive learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 13246–13253, 2023.
  56. 56.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In The Eleventh International Conference on Learning Representations, 2022.
  57. 57.Oh, J., Lee, G., Bae, S., Kwon, J.-m., and Choi, E. Ecg-qa: A comprehensive question answering dataset combined with electrocardiogram. Advances in Neural Information Processing Systems, 36, 2024.
  58. 58.Peris, C., Dupuy, C., Majmudar, J., Parikh, R., Smaili, S., Zemel, R., and Gupta, R. Privacy in the time of language models. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, pp. 1291–1292, 2023.
  59. 59.Qiu, J., Han, W., Zhu, J., Xu, M., Rosenberg, M., Liu, E., Weber, D., and Zhao, D. Transfer knowledge from natural language to electrocardiography: Can we detect cardiovascular disease through language models? arXiv preprint arXiv:2301.09017, 2023a.
  60. 60.Qiu, J., Zhu, J., Liu, S., Han, W., Zhang, J., Duan, C., Rosenberg, M. A., Liu, E., Weber, D., and Zhao, D. Automated cardiovascular record retrieval by multimodal learning between electrocardiogram and clinical report. In Machine Learning for Health (ML4H), pp. 480–497. PMLR, 2023b.
  61. 61.Rasul, K., Ashok, A., Williams, A. R., Khorasani, A., Adamopoulos, G., Bhagwatkar, R., Bilos, M., Ghonia, H., Hassen, N. V., Schneider, A., et al. Lag-llama: Towards foundation models for time series forecasting. arXiv preprint arXiv:2310.08278, 2023.
  62. 62.Rawte, V., Sheth, A., and Das, A. A survey of hallucination in large foundation models. arXiv preprint arXiv:2309.05922, 2023.
  63. 63.Schick, T., Dwivedi-Yu, J., Dess`ı, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2024.
  64. 64.Shumway, R. H., Stoffer, D. S., Shumway, R. H., and Stoffer, D. S. Arima models. Time series analysis and its applications: with R examples, pp. 75–163, 2017.
  65. 65.Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A. Progprompt: Generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 11523–11530. IEEE, 2023.
  66. 66.Spathis, D. and Kawsar, F. The first step is the hardest: Pitfalls of representing and tokenizing temporal data for large language models. arXiv preprint arXiv:2309.06236, 2023.
  67. 67.Sun, C., Li, Y., Li, H., and Hong, S. Test: Text prototype aligned embedding to activate llm’s ability for time series. In International Conference on Machine Learning, 2024.
  68. 68.Sun, Q., Zhang, S., Ma, D., Shi, J., Li, D., Luo, S., Wang, Y., Xu, N., Cao, G., and Zhao, H. Large trajectory models are scalable motion predictors and planners. arXiv preprint arXiv:2310.19620, 2023.
  69. 69.Tian, K., Mitchell, E., Yao, H., Manning, C. D., and Finn, C. Fine-tuning language models for factuality. In The Twelfth International Conference on Learning Representations, 2023.
  70. 70.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  71. 71.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  72. 72.Tsay, R. S. Analysis of financial time series. John wiley & sons, 2005.
  73. 73.Tsymbal, A. The problem of concept drift: definitions and related work. Computer Science Department, Trinity College Dublin, 106(2):58, 2004.
  74. 74.Victor, S., Albert, W., Colin, R., Stephen, B., Lintang, S., Zaid, A., Antoine, C., Arnaud, S., Arun, R., Manan, D., et al. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations, 2022.
  75. 75.Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., et al. Freshllms: Refreshing large language models with search engine augmentation. arXiv preprint arXiv:2310.03214, 2023.
  76. 76.Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al. A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):1–26, 2024.
  77. 77.Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F. Image as a foreign language: BEiT pretraining for vision and vision-language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023a.
  78. 78.Wang, X., Fang, M., Zeng, Z., and Cheng, T. Where would i go next? large language models as human mobility predictors. arXiv preprint arXiv:2308.15197, 2023b.
  79. 79.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022a.
  80. 80.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837, 2022b.
  81. 81.Wen, Q., Sun, L., Yang, F., Song, X., Gao, J., Wang, X., and Xu, H. Time series data augmentation for deep learning: A survey. In IJCAI, pp. 4653–4660, 2021.
  82. 82.Wen, Q., Yang, L., Zhou, T., and Sun, L. Robust time series analysis and applications: An industrial perspective. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD’22), pp. 4836–4837, 2022.
  83. 83.Wen, Q., Zhou, T., Zhang, C., Chen, W., Ma, Z., Yan, J., and Sun, L. Transformers in time series: A survey. In International Joint Conference on Artificial Intelligence(IJCAI), 2023.
  84. 84.Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E. Ccnet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 4003–4012, 2020.
  85. 85.Woo, G., Liu, C., Kumar, A., and Sahoo, D. Pushing the limits of pre-training for time series forecasting in the cloudops domain. arXiv preprint arXiv:2310.05063, 2023.
  86. 86.Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023.
  87. 87.Xie, Q., Han, W., Zhang, X., Lai, Y., Peng, M., Lopez-Lira, A., and Huang, J. Pixiu: A large language model, instruction data and evaluation benchmark for finance. arXiv preprint arXiv:2306.05443, 2023.
  88. 88.Xue, H. and Salim, F. D. Promptcast: A new prompt-based learning paradigm for time series forecasting. IEEE Transactions on Knowledge and Data Engineering, 2023.
  89. 89.Xue, S., Zhou, F., Xu, Y., Zhao, H., Xie, S., Jiang, C., Zhang, J., Zhou, J., Xu, P., Xiu, D., et al. Weaverbird: Empowering financial decision-making with large language model, knowledge base, and search engine. arXiv preprint arXiv:2308.05361, 2023.
  90. 90.Yan, Y., Wen, H., Zhong, S., Chen, W., Chen, H., Wen, Q., Zimmermann, R., and Liang, Y. When urban region profiling meets large language models. arXiv preprint arXiv:2310.18340, 2023.
  91. 91.Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C. Zero-shot video question answering via frozen bidirectional language models. Advances in Neural Information Processing Systems, 35:124–141, 2022a.
  92. 92.Yang, Z., Gan, Z., Wang, J., Hu, X., Lu, Y., Liu, Z., and Wang, L. An empirical study of gpt-3 for few-shot knowledge-based vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 3081–3089, 2022b.
  93. 93.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36, 2024.
  94. 94.Yeh, C.-C. M., Dai, X., Chen, H., Zheng, Y., Fan, Y., Der, A., Lai, V., Zhuang, Z., Wang, J., Wang, L., et al. Toward a foundation model for time series data. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, pp. 4400–4404, 2023.
  95. 95.Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., Bai, L., et al. Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. Advances in Neural Information Processing Systems, 36, 2024.
  96. 96.Yu, H., Guo, P., and Sano, A. Zero-shot ecg diagnosis with large language models and retrieval-augmented generation. In Machine Learning for Health (ML4H), pp. 650–663. PMLR, 2023a.
  97. 97.Yu, X., Chen, Z., Ling, Y., Dong, S., Liu, Z., and Lu, Y. Temporal data meets llm–explainable financial time series forecasting. arXiv preprint arXiv:2306.11025, 2023b.
  98. 98.Yu, X., Chen, Z., and Lu, Y. Harnessing LLMs for temporal data - a study on explainable financial time series forecasting. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 739–753, Singapore, December 2023c.
  99. 99.Zhang, H., Diao, S., Lin, Y., Fung, Y. R., Lian, Q., Wang, X., Chen, Y., Ji, H., and Zhang, T. R-tuning: Teaching large language models to refuse unknown questions. arXiv preprint arXiv:2311.09677, 2023a.
  100. 100.Zhang, J., Huang, J., Jin, S., and Lu, S. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024a.
  101. 101.Zhang, K., Wen, Q., Zhang, C., Cai, R., Jin, M., Liu, Y., Zhang, J. Y., Liang, Y., Pang, G., Song, D., et al. Self-supervised learning for time series analysis: Taxonomy, progress, and prospects. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024b.
  102. 102.Zhang, Q., Ren, X., Xia, L., Yiu, S. M., and Huang, C. Spatio-temporal graph learning with large language model. 2024c. URL https://openreview.net/forum?id=QUkcfqa6GX.
  103. 103.Zhang, S., Dong, L., Li, X., Zhang, S., Sun, X., Wang, S., Li, J., Hu, R., Zhang, T., Wu, F., et al. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792, 2023b.
  104. 104.Zhang, S., Fu, D., Liang, W., Zhang, Z., Yu, B., Cai, P., and Yao, B. Trafficgpt: Viewing, processing and interacting with traffic foundation models. Transport Policy, 150: 95–105, 2024d.
  105. 105.Zhang, X., Chowdhury, R. R., Gupta, R. K., and Shang, J. Large language models for time series: A survey. arXiv preprint arXiv:2402.01801, 2024e.
  106. 106.Zhang, Y., Zhang, Y., Zheng, M., Chen, K., Gao, C., Ge, R., Teng, S., Jelloul, A., Rao, J., Guo, X., et al. Insight miner: A time series analysis dataset for cross-domain alignment with natural language. In NeurIPS 2023 AI for Science Workshop, 2023c.
  107. 107.Zhang, Z., Amiri, H., Liu, Z., Züfle, A., and Zhao, L. Large language models for spatial trajectory patterns mining. arXiv preprint arXiv:2310.04942, 2023d.
  108. 108.Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
  109. 109.Zhou, T., Niu, P., Wang, X., Sun, L., and Jin, R. One fits all: Power general time series analysis by pretrained lm. Advances in neural information processing systems, 2023a.
  110. 110.Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H. Analyzing and mitigating object hallucination in large vision-language models. arXiv preprint arXiv:2310.00754, 2023b.
  111. 111.Zhuo, T. Y., Huang, Y., Chen, C., and Xing, Z. Exploring ai ethics of chatgpt: A diagnostic analysis. arXiv preprint arXiv:2301.12867, 2023.

Citation

MLA
Jin, M., et al. “Position: What Can Large Language Models Tell Us About Time Series Analysis”. arXiv, 2024, http://arxiv.org/abs/2402.02713v2.
APA
Jin, M., Zhang, Y., Chen, W., Zhang, K., Liang, Y., Yang, B., Wang, J., Pan, S., & Wen, Q. (2024). Position: What Can Large Language Models Tell Us about Time Series Analysis. arXiv. http://arxiv.org/abs/2402.02713v2
Chicago
Jin, M., Y. Zhang, W. Chen, et al. 2024. “Position: What Can Large Language Models Tell Us About Time Series Analysis”. arXiv. http://arxiv.org/abs/2402.02713v2.
Harvard
Jin, M. et al. (2024) “Position: What Can Large Language Models Tell Us about Time Series Analysis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.02713v2.
Vancouver
1. Jin M, Zhang Y, Chen W, Zhang K, Liang Y, Yang B, Wang J, Pan S, Wen Q (2024) Position: What Can Large Language Models Tell Us about Time Series Analysis. arXiv

BibTeX

@article{jin2024position,
  title = {Position: What Can Large Language Models Tell Us about Time Series Analysis},
  author = {Jin, Ming and Zhang, Yifan and Chen, Wei and Zhang, Kexin and Liang, Yuxuan and Yang, Bin and Wang, Jindong and Pan, Shirui and Wen, Qingsong},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.02713v2},
  eprint = {2402.02713}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/