Position: What Can Large Language Models Tell Us about Time Series Analysis
Ming JinYifan ZhangWei ChenKexin ZhangYuxuan LiangBin YangJindong WangShirui PanQingsong Wen
Categorizes the emerging roles of large language models in time series analysis as data enhancers, predictors, and autonomous agents while identifying concrete integration strategies and open research opportunities for building universal time series intelligence.
Time series analysis is vital for understanding complex dynamic systems across industries such as finance, transportation, and healthcare. However, traditional statistical and deep learning models are typically narrow, task-specific, and heavily dependent on domain knowledge and extensive tuning. The article examines whether large language models (LLMs) can bridge the gap toward general-purpose analytical intelligence, demonstrating how LLMs can transform time series analysis from isolated forecasting tools into unified, interactive, and reasoning-driven systems.
To evaluate this potential, the article conducts a comprehensive literature review and exploratory empirical tests. It assesses three operational roles for LLMs: data and model enhancers, direct predictors, and autonomous analytical agents. For the empirical component, the authors test zero-shot prompt-based reasoning using GPT-3.5 on public benchmarks, specifically examining activity classification using waist-mounted smartphone sensor data from 30 participants and electric transformer temperature records.
The findings indicate that LLMs can impact time series analysis in three major ways. First, as data and model enhancers, LLMs successfully enrich sparse numerical data with textual context and transfer external knowledge to specialized domain models. Second, as direct predictors, LLMs demonstrate competitive zero-shot and few-shot forecasting performance, either through input reprogramming and soft prompts or via specialized continuous tokenization that allows frozen models to match domain-specific baselines. Third, in agent-based evaluation, the model achieves complete accuracy on simple movement patterns like standing in zero-shot activity classification, while providing clear, natural language explanations of its reasoning. However, the evaluation also shows that LLMs struggle with subtle, complex numerical patterns, occasionally hallucinate false data or explanations, and exhibit significant task bias, such as misclassifying laying down as sitting or standing.
These results imply that integrating language models into time series workflows can substantially enhance interpretability, reduce the resources required to build models from scratch, and enable multimodal problem solving. Nevertheless, the presence of hallucinations, privacy risks associated with sensitive industrial telemetry, and high computational costs introduce operational and compliance risks if models are deployed without domain guardrails.
To safely capitalize on these capabilities, organizations should pursue hybrid integration frameworks rather than relying on stand-alone language models. Recommended approaches include aligning time series embeddings with language representations, developing multi-agent systems where specialized models perform low-level pattern analysis while language models orchestrate planning and communication, and establishing robust prompt guidelines. Decision-makers should validate these tools through controlled pilot programs before full deployment. While the underlying literature review is solid, confidence in fully autonomous zero-shot LLM agents remains low, warranting caution due to persistent hallucination and data drift challenges.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). Provides a comprehensive survey of classical and deep learning architectures for time series forecasting, establishing the baseline methodologies that LLM-based approaches seek to enhance or replace.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Surveys the adaptation and limitations of Transformer attention mechanisms in temporal domains, contextualizing the technical motivation behind using pre-trained language models for time series.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Outlines foundational architectures for LLM-based autonomous agents, directly informing the position paper's evaluation of LLMs as autonomous analytical agents.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). Establishes core paradigms for perception, planning, and tool use in LLM agent systems, which underpin the source's framework for multi-agent time series workflows.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot chain-of-thought reasoning in large language models, providing the operational mechanism tested in the source's empirical classification experiments.
- Paper: A Transformer-based Framework for Multivariate Time Series Representation Learning, George Zerveas et al. (2020). Introduces unsupervised representation learning and token masking for multivariate time series, foundational to aligning temporal embeddings with pre-trained language representations.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Examines unfaithfulness and bias in chain-of-thought explanations, underpinning the source's critique of LLM hallucinations and rationalization in temporal tasks.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). Introduces the ReAct framework for interleaving reasoning and environment interaction, providing the blueprint for agentic reasoning in data analytics.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Shows how continuous closed-loop environmental feedback grounds LLM reasoning, relevant to designing hybrid LLM pipelines with domain-specific guardrails.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). Categorizes core ethical, privacy, and hallucination risks in language models, framing the operational and compliance constraints emphasized in the position paper.
- Paper: Language Models Are Implicitly Continuous, Samuele Marro et al. (2025). Extends the exploration of LLMs for continuous numerical domains by showing that pre-trained language models implicitly process inputs as continuous functions.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Generalizes the position paper's vision of autonomous analytical agents into a comprehensive framework for multi-agent reasoning, planning, and dynamic interaction.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). Addresses the computational cost and context-management challenges highlighted in the source by reviewing efficiency techniques in agent memory, planning, and tool use.
- Paper: Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs, Wei Zhou et al. (2026). Applies and operationalizes LLM-driven data preparation and enrichment workflows across complex, messy real-world datasets.
- Paper: The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook, Xinlei Yu et al. (2026). Explores latent-space computation as an alternative to explicit tokenization, directly building upon the need for seamless alignment between continuous temporal signals and neural representations.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). Surveys reinforced reasoning and test-time compute strategies that help mitigate the reasoning failures and task biases identified in the source's empirical evaluation.
