Transformers in Time Series: A Survey

Qingsong WenTian ZhouChao ZhangWeiqiu ChenZiqing MaJunchi YanLiang Sun

article2022IJCAI1,674 citationsMost Influential IJCAI'23 Paper

Presents a systematic taxonomy of Transformer adaptations for time series forecasting, anomaly detection, and classification alongside empirical analyses of model size and seasonal decomposition to guide future architectural designs.

Listen

Modern decision-making across industries relies heavily on analyzing sequential, time-stamped data to forecast demand, detect system anomalies, and classify complex patterns. While Transformer neural networks have achieved state-of-the-art results in language and vision applications, adapting their self-attention mechanisms to time series data presents distinct operational challenges, such as quadratic computational overhead and difficulties in capturing repeating seasonal patterns.

The article systematically reviews the adaptation of Transformer architectures for time series modeling, evaluating their structural modifications and performance across forecasting, anomaly detection, and classification tasks.

To conduct this evaluation, the article synthesizes existing literature and establishes a taxonomy encompassing module-level changes—such as learnable and timestamp encodings, sparse attention, and token patching—as well as architecture-level designs. In addition, the article presents targeted empirical experiments using benchmark operational data (ETTm2) to examine model robustness across sequence lengths, network depth, and structural decompositions.

The findings indicate that standard Transformers face significant practical hurdles in time series environments. First, directly increasing input sequence length causes rapid performance deterioration in many models, showing that high computational capacity does not automatically translate into effective long-sequence utilization. Second, unlike text or vision domains where deeper networks excel, shallower time series Transformers with only 3 to 6 layers consistently outperform deeper variants, which suffer from memory bottlenecks and degradation. Third, integrating seasonal-trend decomposition improves model forecasting accuracy by 50% to 82% across diverse configurations, demonstrating that isolating periodic signals is vital. Finally, recent patch-based and frequency-domain designs achieve linear calculation scaling while outperforming simpler baselines.

These results demonstrate that standard Transformers cannot simply be repurposed for sequential temporal data without domain-specific customization. Unmodified implementations risk high computational costs and memory failure without accuracy gains. To mitigate operational and computational risk, organizations should prioritize hybrid architectures that separate trends from periodic seasonality and leverage efficient attention mechanisms.

Stakeholders deploying deep learning for temporal analytics should integrate seasonal-trend decomposition into their modeling pipelines and avoid oversized, deep networks. When scaling to high-dimensional or long-sequence tasks, engineering teams should consider efficient token segmentation (such as patching) and explore combinations with graph neural networks for spatial dependencies. Automated architecture discovery should be piloted to systematically tune hyper-parameters for large-scale production.

While the review provides comprehensive qualitative categorization, the direct experimental evaluations are limited to a single benchmark dataset, requiring caution when generalizing findings across diverse industrial environments. Nevertheless, there is high confidence in the fundamental conclusions regarding model sizing and seasonal decomposition, which are reinforced by broader cross-study literature evidence.

Cover for Transformers in Time Series: A Survey

Abstract

Transformers have achieved superior performances in many tasks in natural language processing and computer vision, which also triggered great interest in the time series community. Among multiple advantages of Transformers, the ability to capture long-range dependencies and interactions is especially attractive for time series modeling, leading to exciting progress in various time series applications. In this paper, we systematically review Transformer schemes for time series modeling by highlighting their strengths as well as limitations. In particular, we examine the development of time series Transformers in two perspectives. From the perspective of network structure, we summarize the adaptations and modifications that have been made to Transformers in order to accommodate the challenges in time series analysis. From the perspective of applications, we categorize time series Transformers based on common tasks including forecasting, anomaly detection, and classification. Empirically, we perform robust analysis, model size analysis, and seasonal-trend decomposition analysis to study how Transformers perform in time series. Finally, we discuss and suggest future directions to provide useful research guidance. To the best of our knowledge, this paper is the first work to comprehensively and systematically summarize the recent advances of Transformers for modeling time series data. We hope this survey will ignite further research interests in time series Transformers.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries of the Transformer
  • 2.1 Vanilla Transformer
  • 2.2 Input Encoding and Positional Encoding
  • 2.2.1 Absolute Positional Encoding
  • 2.2.2 Relative Positional Encoding
  • 2.3 Multi-head Attention
  • 2.4 Feed-forward and Residual Network
  • 3 Taxonomy of Transformers in Time Series
  • 4 Network Modifications for Time Series
  • 4.1 Positional Encoding
  • 4.2 Attention Module
  • 4.3 Architecture-based Attention Innovation
  • 5 Applications of Time Series Transformers
  • 5.1 Transformers in Forecasting
  • 5.1.1 Time Series Forecasting
  • 5.1.2 Spatio-Temporal Forecasting
  • 5.1.3 Event Forecasting
  • 5.2 Transformers in Anomaly Detection
  • 5.3 Transformers in Classification
  • 6 Experimental Evaluation and Discussion
  • 6.0.1 Robustness Analysis
  • 6.0.2 Model Size Analysis
  • 6.0.3 Seasonal-Trend Decomposition Analysis
  • 7 Future Research Opportunities
  • 7.1 Inductive Biases for Time Series Transformers
  • 7.2 Transformers and GNN for Time Series
  • 7.3 Pre-trained Transformers for Time Series
  • 7.4 Transformers with Architecture Level Variants
  • 7.5 Transformers with NAS for Time Series
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Transformers for Time Series Modeling

    model/method

    A comprehensive taxonomy classifies time series Transformer models along two dimensions: network modifications and application domains.

    1. Network Modifications:
    • Positional Encoding: Adapting position representations to capture sequential ordering, including fixed vanilla sinusoidal encodings, learnable continuous position embeddings, and multi-resolution timestamp/calendar encodings.
    • Attention Modules: Modifying self-attention mechanisms to mitigate quadratic time and memory complexity via sparsity biases (e.g., Logsparse attention, pyramidal tree attention) and low-rank or frequency-domain approximations (e.g., Fourier transform, wavelet transform, quaternion rotation).
    • Architecture Level: Designing hierarchical multi-resolution downsampling and iterative multi-scale refinement structures.
    1. Application Domains:
    • Forecasting: Standard univariate/multivariate time series forecasting, spatio-temporal forecasting (incorporating spatial graph structures or space-time cuboid attention), and asynchronous event forecasting (incorporating continuous-time temporal point processes).
    • Anomaly Detection: Utilizing Transformers for sequence reconstruction and anomaly scoring, often combined with generative models (VAEs, GANs), graph networks, or association discrepancy constraints.
    • Classification: Deploying multi-tower attention (separating time-step and channel dimensions), task-aware timestamp masking, or fine-tuning large pre-trained foundation models.
  2. Knowl 2 — Seasonal-Trend Decomposition for Error Reduction in Time Series Transformers

    empirical result

    Incorporating moving-average seasonal-trend decomposition into Transformer architectures consistently boosts long-term forecasting performance, yielding a 53% to 82% relative reduction in Mean Squared Error (MSE) on the ETTm2 benchmark across various attention mechanisms. The relative benefit of decomposition increases with longer forecasting horizons.

    FEDformer Autoformer Informer LogTrans Reformer Transformer Relative
    Horizon Ori Decomp Ori Decomp Ori Decomp Ori Decomp Ori Decomp Ori Decomp Promotion
    96 0.457 0.203 0.581 0.255 0.365 0.354 0.768 0.231 0.658 0.218 0.604 0.204 53%
    192 0.841 0.269 1.403 0.281 0.533 0.432 0.989 0.378 1.078 0.336 1.060 0.266 62%
    336 1.451 0.325 2.632 0.339 1.363 0.481 1.334 0.362 1.549 0.366 1.413 0.375 75%
    720 3.282 0.421 3.058 0.422 3.379 0.822 3.048 0.539 2.631 0.502 2.672 0.537 82%

    In the table, Ori denotes the original model without decomposition, Decomp denotes the architecture augmented with moving-average seasonal-trend decomposition, and Horizon denotes the number of forecast steps evaluated on the ETTm2 dataset.

  3. Knowl 3 — Computational and Memory Complexities of Time Series Transformer Variants

    data/table

    Specialized time series Transformers replace the vanilla Transformer's quadratic O(N2)\mathcal{O}(N^2) self-attention with sparse, tree-based, frequency-domain, or patch-based mechanisms, reducing training time and memory complexity while enabling single-step direct generative decoding during inference.

    Model Training Time Training Memory Testing Steps
    Vanilla Transformer O(N2)\mathcal{O}(N^2) O(N2)\mathcal{O}(N^2) NN
    LogTrans O(Nlog⁡N)\mathcal{O}(N \log N) O(Nlog⁡N)\mathcal{O}(N \log N) 1
    Informer O(Nlog⁡N)\mathcal{O}(N \log N) O(Nlog⁡N)\mathcal{O}(N \log N) 1
    Autoformer O(Nlog⁡N)\mathcal{O}(N \log N) O(Nlog⁡N)\mathcal{O}(N \log N) 1
    Pyraformer O(N)\mathcal{O}(N) O(N)\mathcal{O}(N) 1
    Quatformer O(2cN)\mathcal{O}(2^c N) O(2cN)\mathcal{O}(2^c N) 1
    FEDformer O(N)\mathcal{O}(N) O(N)\mathcal{O}(N) 1
    Crossformer O(DLseg2N2)\mathcal{O}\left(\frac{D}{L_{\text{seg}}^2} N^2\right) O(N)\mathcal{O}(N) 1

    where NN denotes the input sequence length, DD is the number of time series dimensions/channels, LsegL_{\text{seg}} is the segment length used in dimension-segment-wise embedding, and cc is a constant scale factor associated with global memory decoupling in quaternion attention.

  4. Knowl 4 — Performance Degradation of Time Series Transformers with Prolonged Input Sequences

    empirical result

    Evaluating forecasting performance (predicting 96 future steps on the ETTm2 dataset) as a function of historical input sequence length reveals that most time series Transformers experience severe performance degradation as input length increases from 96 to 1440 steps, demonstrating an inability to effectively exploit long-term historical contexts.

    Input Length Transformer Autoformer Informer Reformer LogFormer
    96 0.557 0.239 0.428 0.615 0.667
    192 0.710 0.265 0.385 0.686 0.697
    336 1.078 0.375 1.078 1.359 0.937
    720 1.691 0.315 1.057 1.443 2.153
    1440 0.936 0.552 1.898 0.815 0.867

    The values denote Mean Squared Error (MSE). Across architectures, expanding the input horizon leads to higher prediction errors, indicating a lack of input-length robustness in current time series Transformer formulations.

  5. Knowl 5 — Empirical Impact of Network Depth on Time Series Transformer Performance

    empirical result

    In contrast to NLP and Computer Vision models where increasing layer depth (e.g., 12 to 128 layers) systematically improves capacity and performance, time series Transformers achieve optimal performance with shallow architectures (3 to 6 layers). Increasing depth beyond 6 layers degrades accuracy or leads to training divergence and out-of-memory errors on the ETTm2 dataset (forecasting 96 future steps).

    Layer Number Transformer Autoformer Informer Reformer LogFormer
    3 0.557 0.234 0.428 0.597 0.667
    6 0.439 0.282 0.489 0.353 0.387
    12 0.556 0.238 0.779 0.481 0.562
    24 0.580 0.266 0.815 1.109 0.690
    48 0.461 NaN 1.623 OOM 2.992

    The reported metric is Mean Squared Error (MSE). NaN indicates numerical divergence during training, and OOM indicates Out-Of-Memory failure.

  6. Knowl 6 — Positional Encoding Methods for Time Series Transformers

    model/method

    Because self-attention is permutation-invariant, sequential ordering must be explicitly injected into time series Transformers. Three main positional encoding schemes are used:

    1. Vanilla Sinusoidal Positional Encoding: Fixed, hand-crafted trigonometric functions defined for position index tt and channel index ii: PE(t)i={sin⁡(ωit)if i%2=0cos⁡(ωit)if i%2=1PE(t)_i = \begin{cases} \sin(\omega_i t) & \text{if } i \% 2 = 0 \\ \cos(\omega_i t) & \text{if } i \% 2 = 1 \end{cases} where ωi\omega_i is a predetermined frequency across feature dimensions.

    2. Learnable Positional Encoding: Trainable embedding vectors mapped to each position index and optimized end-to-end, or recurrent encoders (such as an LSTM) applied to input positions to preserve sequential ordering dynamically.

    3. Timestamp Encoding: Scalar calendar timestamps (e.g., minute, hour, day of week, day of month, month, year) and domain event markers (e.g., holidays) mapped into continuous embedding vectors via learnable lookup layers and added to token embeddings to provide explicit temporal context.

  7. Knowl 7 — Module-Level Innovations in Time Series Forecasting Transformers

    model/method

    Module-level innovations in time series forecasting retain the overall Transformer encoder-decoder framework while introducing inductive biases through three distinct mechanisms:

    1. Custom Attention Mechanisms: Replacing dense dot-product attention with sparse attention (e.g., LogTrans causal convolution with Logsparse masks, Informer ProbSparse query selection), frequency-domain attention (e.g., FEDformer selecting compact Fourier/wavelet frequency modes), quaternion rotation attention (e.g., Quatformer learning period and phase shifts), or multi-stage attention across time and dimensions (e.g., Crossformer).

    2. Stationarization and Normalization Modules: Counteracting distribution shift and over-stationarization in non-stationary time series by integrating modular stationarization layers before attention and de-stationarization layers after attention to restore predictive statistical properties (e.g., Non-stationary Transformer).

    3. Token Representation and Channel Independence: Transitioning from point-wise token inputs to subseries-level tokenization (patches) and adopting channel-independent architectures where univariate channels share Transformer weights without cross-channel attention interference (e.g., PatchTST, Autoformer auto-correlation).

  8. Knowl 8 — Mechanisms of Transformers in Time Series Anomaly Detection

    model/method

    Transformers are adapted for multivariate time series anomaly detection through four primary modeling paradigms:

    1. Adversarial Reconstruction Amplification: Using a two-encoder, two-decoder Transformer network trained with a GAN-style minimax objective (e.g., TranAD) to amplify subtle reconstruction errors of anomalous deviations that standard autoencoders smooth over.

    2. Generative Latent Modeling: Combining Transformers with Variational Autoencoders (e.g., TransAnomaly, MT-RVAE) to parallelize sequence processing while modeling multi-scale latent probability distributions of normal series behavior.

    3. Graph-Augmented Attention: Combining graph neural network layers with multi-branch attention (e.g., GTA) to capture spatial message passing and topological dependency structures across interconnected multivariate channels.

    4. Association Discrepancy Modeling: Enforcing an inductive bias where normal points build broad global associations across the entire sequence while anomalies only exhibit strong local associations with adjacent points (e.g., AnomalyTrans), optimized via a minimax strategy between Gaussian prior-associations and learned series-associations.

  9. Knowl 9 — Transformers in Spatio-Temporal and Asynchronous Event Forecasting

    model/method

    Transformers extend beyond regular 1D sequences to handle complex spatial and irregular temporal structures:

    1. Spatio-Temporal Forecasting: Models combine temporal self-attention modules with spatial graph neural networks or spatial self-attention modules (e.g., Traffic Transformer, Spatial-Temporal Transformer). Advanced variants partition space-time tensors into 3D cuboids to apply parallel cuboid-level attention (Earthformer) or apply dartboard spatial attention paired with causal temporal attention and latent variable modeling (AirFormer).

    2. Temporal Point Process (TPP) Event Forecasting: For asynchronous event sequences with irregular inter-arrival times, models modify positional encodings by converting continuous time intervals into sinusoidal representations (e.g., Self-Attentive Hawkes Process [SAHP], Transformer Hawkes Process [THP], Attentive Neural Datalog Through Time [A-NDTT]) to directly estimate the conditional event intensity function.

  10. Knowl 10 — Key Challenges and Future Directions for Time Series Transformers

    limitation

    Five major open research challenges and opportunities exist for time series Transformers:

    1. Inductive Biases: Reconciling the trade-off between channel-independent modeling (which suppresses cross-channel noise) and cross-dimension modeling (which captures multivariate interactions), alongside integrating native periodic, trend, and stationarity priors.
    2. Integration with Graph Neural Networks (GNNs): Unifying Transformers with GNNs to model high-dimensional spatial correlations, causal structures, and physical constraints in multivariate spatio-temporal environments.
    3. Pre-trained Foundation Models: Developing large-scale self-supervised pre-training schemes specifically tailored for time series data to support zero-shot and few-shot forecasting and representation learning beyond classification.
    4. Architecture-Level Innovation: Moving beyond standard encoder-decoder backbones by exploring lightweight designs, cross-block transparent attention, dynamic early exiting, and recurrence.
    5. Automated Neural Architecture Search (NAS): Utilizing NAS to automate the discovery of computationally and memory-efficient Transformer configurations (such as optimal patch sizes, layer depths, and attention heads) for industrial-scale time series.

Coverage note — No substantial contributed material was omitted; standard background descriptions of the vanilla Transformer equations (multi-head attention, layer norm, feed-forward) were excluded as non-contributed background.

References

  1. 1.[Bapna et al., 2018] Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao, and Yonghui Wu. Training deeper neural machine translation models with transparent attention. In EMNLP, 2018.
  2. 2.[Benidis et al., 2022] Konstantinos Benidis, Syama Sundar Rangapuram, Valentin Flunkert, Yuyang Wang, Danielle Maddix, , et al. Deep learning for time series forecasting: Tutorial and literature survey. ACM Computing Surveys, 55(6):1–36, 2022.
  3. 3.[Blázquez-García et al., 2021] Ane Blázquez-García, Angel Conde, Usue Mori, et al. A review on outlier/anomaly detection in time series data. ACM Computing Surveys, 54(3):1–33, 2021.
  4. 4.[Brown et al., 2020] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, et al. Language models are few-shot learners. NeurIPS, 2020.
  5. 5.[Cai et al., 2020] Ling Cai, Krzysztof Janowicz, Gengchen Mai, Bo Yan, and Rui Zhu. Traffic transformer: Capturing the continuity and periodicity of time series for traffic forecasting. Transactions in GIS, 24(3):736–755, 2020.
  6. 6.[Chen et al., 2021a] Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, et al. Pre-trained image processing transformer. In CVPR, 2021.
  7. 7.[Chen et al., 2021b] Minghao Chen, Houwen Peng, Jianlong Fu, and Haibin Ling. AutoFormer: Searching transformers for visual recognition. In CVPR, 2021.
  8. 8.[Chen et al., 2021c] Zekai Chen, Dingshuo Chen, Xiao Zhang, Zixuan Yuan, and Xiuzhen Cheng. Learning graph structures with transformer for multivariate time series anomaly detection in IoT. IEEE Internet of Things Journal, 2021.
  9. 9.[Chen et al., 2022] Weiqi Chen, Wenwei Wang, Bingqing Peng, Qingsong Wen, Tian Zhou, and Liang Sun. Learning to rotate: Quaternion transformer for complicated periodical time series forecasting. In KDD, 2022.
  10. 10.[Choi et al., 2021] Kukjin Choi, Jihun Yi, Changhwa Park, and Sungroh Yoon. Deep learning for anomaly detection in timeseries data: Review, analysis, and guidelines. IEEE Access, 2021.
  11. 11.[Chowdhury et al., 2022] Ranak Roy Chowdhury, Xiyuan Zhang, Jingbo Shang, Rajesh K Gupta, and Dezhi Hong. TARNet: Taskaware reconstruction for time-series transformer. In KDD, 2022.
  12. 12.[Cirstea et al., 2022] Razvan-Gabriel Cirstea, Chenjuan Guo, Bin Yang, Tung Kieu, Xuanyi Dong, and Shirui Pan. Triformer: Triangular, variable-specific attentions for long sequence multivariate time series forecasting. In IJCAI, 2022.
  13. 13.[Cleveland et al., 1990] Robert Cleveland, William Cleveland, Jean McRae, et al. STL: A seasonal-trend decomposition procedure based on loess. Journal of Official Statistics, 6(1):3–73, 1990.
  14. 14.[Dai et al., 2019] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, et al. Transformer-XL: Attentive language models beyond a fixed-length context. In ACL, 2019.
  15. 15.[Dehghani et al., 2019] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. In ICLR, 2019.
  16. 16.[Dosovitskiy et al., 2021] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  17. 17.[Elsken et al., 2019] Elsken, Thomas, Jan Hendrik Metzen, and Frank Hutter. Neural architecture search: A survey. Journal of Machine Learning Research, 2019.
  18. 18.[Gao et al., 2022] Zhihan Gao, Xingjian Shi, Hao Wang, Yi Zhu, Bernie Wang, Mu Li, et al. Earthformer: Exploring space-time transformers for earth system forecasting. In NeurIPS, 2022.
  19. 19.[Gehring et al., 2017] Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. Convolutional sequence to sequence learning. In ICML, 2017.
  20. 20.[Goodfellow et al., 2014] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, et al. Generative adversarial nets. NeurIPS, 2014.
  21. 21.[Han et al., 2021] Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Liang Zhang, et al. Pretrained models: Past, present and future. AI Open, 2021.
  22. 22.[Han et al., 2022] Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, et al. A survey on vision transformer. IEEE TPAMI, 45(1):87–110, 2022.
  23. 23.[Hyndman and Khandakar, 2008] Rob J Hyndman and Yeasmin Khandakar. Automatic time series forecasting: the forecast package for r. Journal of statistical software, 27:1–22, 2008.
  24. 24.[Ismail Fawaz et al., 2019] Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. Deep learning for time series classification: a review. Data mining and knowledge discovery, 2019.
  25. 25.[Ke et al., 2021] Guolin Ke, Di He, and Tie-Yan Liu. Rethinking positional encoding in language pre-training. In ICLR, 2021.
  26. 26.[Kenton and others, 2019] Jacob Devlin Ming-Wei Chang Kenton et al. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, 2019.
  27. 27.[Kingma and Welling, 2014] Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014.
  28. 28.[Li et al., 2019] Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In NeurIPS, 2019.
  29. 29.[Li et al., 2021] Longyuan Li, Jian Yao, Li Wenliang, Tong He, Tianjun Xiao, Junchi Yan, David Wipf, and Zheng Zhang. Grin: Generative relation and intention network for multi-agent trajectory prediction. In NeurIPS, 2021.
  30. 30.[Liang et al., 2023] Yuxuan Liang, Yutong Xia, Songyu Ke, Yiwei Wang, Qingsong Wen, Junbo Zhang, Yu Zheng, and Roger Zimmermann. AirFormer: Predicting nationwide air quality in china with transformers. In AAAI, 2023.
  31. 31.[Lim and Zohren, 2021] Bryan Lim and Stefan Zohren. Timeseries forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society, 2021.
  32. 32.[Lim et al., 2021] Bryan Lim, Sercan Ö Arık, Nicolas Loeff, and Tomas Pfister. Temporal fusion transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4):1748–1764, 2021.
  33. 33.[Lin et al., 2021] Yang Lin, Irena Koprinska, and Mashud Rana. SSDNet: State space decomposition neural network for time series forecasting. In ICDM, 2021.
  34. 34.[Liu et al., 2021] Minghao Liu, Shengqi Ren, Siyuan Ma, Jiahui Jiao, Yizhou Chen, Zhiguang Wang, and Wei Song. Gated transformer networks for multivariate time series classification. arXiv preprint arXiv:2103.14438, 2021.
  35. 35.[Liu et al., 2022a] Shizhan Liu, Hang Yu, Cong Liao, Jianguo Li, Weiyao Lin, Alex X. Liu, and Schahram Dustdar. Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting. In ICLR, 2022.
  36. 36.[Liu et al., 2022b] Yong Liu, Haixu Wu, Jianmin Wang, and Mingsheng Long. Non-stationary transformers: Exploring the stationarity in time series forecasting. In NeurIPS, 2022.
  37. 37.[Mehta et al., 2021] Sachin Mehta, Marjan Ghazvininejad, Srini Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. Delight: Deep and light-weight transformer. In ICLR, 2021.
  38. 38.[Mei et al., 2022] Hongyuan Mei, Chenghao Yang, and Jason Eisner. Transformer embeddings of irregularly spaced events and their participants. In ICLR, 2022.
  39. 39.[Nie et al., 2023] Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. A time series is worth 64 words: Longterm forecasting with transformers. In ICLR, 2023.
  40. 40.[Rußwurm and Körner, 2020] Marc Rußwurm and Marco Körner. Self-attention for raw optical satellite time series classification. ISPRS J. Photogramm. Remote Sens., 169:421–435, 11 2020.
  41. 41.[Shabani et al., 2023] Amin Shabani, Amir Abdi, Lili Meng, and Tristan Sylvain. Scaleformer: iterative multi-scale refining transformers for time series forecasting. In ICLR, 2023.
  42. 42.[Shaw et al., 2018] Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. In NAACL, 2018.
  43. 43.[Shchur et al., 2021] Oleksandr Shchur, Ali Caner Türkmen, Tim Januschowski, and Stephan Günnemann. Neural temporal point processes: A review. In IJCAI, 2021.
  44. 44.[So et al., 2019] David So, Quoc Le, and Chen Liang. The evolved transformer. In ICML, 2019.
  45. 45.[Tang and Matteson, 2021] Binh Tang and David Matteson. Probabilistic transformer for time series analysis. In NeurIPS, 2021.
  46. 46.[Tay et al., 2022] Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. ACM Computing Surveys, 55(6):1–28, 2022.
  47. 47.[Torres et al., 2021] Jose F. Torres, Dalil Hadjout, Abderrazak Sebaa, Francisco Martínez-Álvarez, and Alicia Troncoso. Deep learning for time series forecasting: a survey. Big Data, 2021.
  48. 48.[Tuli et al., 2022] Shreshth Tuli, Giuliano Casale, and Nicholas R Jennings. TranAD: Deep transformer networks for anomaly detection in multivariate time series data. In VLDB, 2022.
  49. 49.[Vaswani et al., 2017] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, et al. Attention is all you need. In NeurIPS, 2017.
  50. 50.[Wang et al., 2020] Xiaoxing Wang, Chao Xue, Junchi Yan, Xiaokang Yang, Yonggang Hu, et al. MergeNAS: Merge operations into one for differentiable architecture search. In IJCAI, 2020.
  51. 51.[Wang et al., 2022] Xixuan Wang, Dechang Pi, Xiangyan Zhang, et al. Variational transformer-based anomaly detection approach for multivariate time series. Measurement, page 110791, 2022.
  52. 52.[Wen et al., 2019] Qingsong Wen, Jingkun Gao, Xiaomin Song, Liang Sun, Huan Xu, et al. RobustSTL: A robust seasonal-trend decomposition algorithm for long time series. In AAAI, 2019.
  53. 53.[Wen et al., 2020] Qingsong Wen, Zhe Zhang, Yan Li, and Liang Sun. Fast RobustSTL: Efficient and robust seasonal-trend decomposition for time series with complex patterns. In KDD, 2020.
  54. 54.[Wen et al., 2021a] Qingsong Wen, Kai He, Liang Sun, Yingying Zhang, Min Ke, et al. RobustPeriod: Time-frequency mining for robust multiple periodicities detection. In SIGMOD, 2021.
  55. 55.[Wen et al., 2021b] Qingsong Wen, Liang Sun, Fan Yang, Xiaomin Song, Jingkun Gao, Xue Wang, and Huan Xu. Time series data augmentation for deep learning: A survey. In IJCAI, 2021.
  56. 56.[Wen et al., 2022] Qingsong Wen, Linxiao Yang, Tian Zhou, and Liang Sun. Robust time series analysis and applications: An industrial perspective. In KDD, 2022.
  57. 57.[Wu et al., 2020a] Sifan Wu, Xi Xiao, Qianggang Ding, Peilin Zhao, Ying Wei, and Junzhou Huang. Adversarial sparse transformer for time series forecasting. In NeurIPS, 2020.
  58. 58.[Wu et al., 2020b] Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. Lite transformer with long-short range attention. In ICLR, 2020.
  59. 59.[Wu et al., 2021] Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In NeurIPS, 2021.
  60. 60.[Xin et al., 2020] Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy J. Lin. DeeBERT: Dynamic early exiting for accelerating bert inference. In ACL, 2020.
  61. 61.[Xu et al., 2020] Mingxing Xu, Wenrui Dai, Chunmiao Liu, Xing Gao, Weiyao Lin, Guo-Jun Qi, and Hongkai Xiong. Spatial-temporal transformer networks for traffic flow forecasting. arXiv preprint arXiv:2001.02908, 2020.
  62. 62.[Xu et al., 2022] Jiehui Xu, Haixu Wu, Jianmin Wang, and Mingsheng Long. Anomaly Transformer: Time series anomaly detection with association discrepancy. In ICLR, 2022.
  63. 63.[Yan et al., 2019] Junchi Yan, Hongteng Xu, and Liangda Li. Modeling and applications for temporal point processes. In KDD, 2019.
  64. 64.[Yang et al., 2021] Chao-Han Huck Yang, Yun-Yun Tsai, and Pin-Yu Chen. Voice2series: Reprogramming acoustic models for time series classification. In ICML, 2021.
  65. 65.[Yu et al., 2020] Cunjun Yu, Xiao Ma, Jiawei Ren, Haiyu Zhao, and Shuai Yi. Spatio-temporal graph transformer networks for pedestrian trajectory prediction. In ECCV, 2020.
  66. 66.[Yuan and Lin, 2020] Yuan Yuan and Lei Lin. Self-supervised pre-training of transformers for satellite image time series classification. IEEE J-STARS, 14:474–487, 2020.
  67. 67.[Yun et al., 2020] Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, et al. Are transformers universal approximators of sequence-to-sequence functions? In ICLR, 2020.
  68. 68.[Zeng et al., 2023] Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. Are transformers effective for time series forecasting? In AAAI, 2023.
  69. 69.[Zerveas et al., 2021] George Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty, and Carsten Eickhoff. A transformer-based framework for multivariate time series representation learning. In KDD, 2021.
  70. 70.[Zhang and Yan, 2023] Yunhao Zhang and Junchi Yan. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In ICLR, 2023.
  71. 71.[Zhang et al., 2020] Qiang Zhang, Aldo Lipani, Omer Kirnap, and Emine Yilmaz. Self-attentive Hawkes process. In ICML, 2020.
  72. 72.[Zhang et al., 2021] Hongwei Zhang, Yuanqing Xia, et al. Unsupervised anomaly detection in multivariate time series through transformer-based variational autoencoder. In CCDC, 2021.
  73. 73.[Zhou et al., 2021] Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. Informer: Beyond efficient transformer for long sequence time-series forecasting. In AAAI, 2021.
  74. 74.[Zhou et al., 2022] Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In ICML, 2022.
  75. 75.[Zuo et al., 2020] Simiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao, and Hongyuan Zha. Transformer Hawkes process. In ICML, 2020.

Citation

MLA
Wen, Q., et al. “Transformers in Time Series: A Survey”. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 2023, pp. 6778–86, https://doi.org/10.24963/ijcai.2023/759.
APA
Wen, Q., Zhou, T., Zhang, C., Chen, W., Ma, Z., Yan, J., & Sun, L. (2023). Transformers in Time Series: A Survey. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 6778–6786. https://doi.org/10.24963/ijcai.2023/759
Chicago
Wen, Q., T. Zhou, C. Zhang, et al. 2023. “Transformers in Time Series: A Survey”. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, 6778–86. https://doi.org/10.24963/ijcai.2023/759.
Harvard
Wen, Q. et al. (2023) “Transformers in Time Series: A Survey”, Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, pp. 6778–6786. Available at: https://doi.org/10.24963/ijcai.2023/759.
Vancouver
1. Wen Q, Zhou T, Zhang C, Chen W, Ma Z, Yan J, Sun L (2023) Transformers in Time Series: A Survey. In: Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, pp 6778–6786

BibTeX

@inproceedings{Wen_2023, series={IJCAI-2023}, title={Transformers in Time Series: A Survey}, url={http://dx.doi.org/10.24963/ijcai.2023/759}, DOI={10.24963/ijcai.2023/759}, booktitle={Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence}, publisher={International Joint Conferences on Artificial Intelligence Organization}, author={Wen, Qingsong and Zhou, Tian and Zhang, Chaoli and Chen, Weiqi and Ma, Ziqing and Yan, Junchi and Sun, Liang}, year={2023}, month=Aug, pages={6778–6786}, collection={IJCAI-2023} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF