Transformers in Time Series: A Survey
Qingsong WenTian ZhouChao ZhangWeiqiu ChenZiqing MaJunchi YanLiang Sun
Presents a systematic taxonomy of Transformer adaptations for time series forecasting, anomaly detection, and classification alongside empirical analyses of model size and seasonal decomposition to guide future architectural designs.
Modern decision-making across industries relies heavily on analyzing sequential, time-stamped data to forecast demand, detect system anomalies, and classify complex patterns. While Transformer neural networks have achieved state-of-the-art results in language and vision applications, adapting their self-attention mechanisms to time series data presents distinct operational challenges, such as quadratic computational overhead and difficulties in capturing repeating seasonal patterns.
The article systematically reviews the adaptation of Transformer architectures for time series modeling, evaluating their structural modifications and performance across forecasting, anomaly detection, and classification tasks.
To conduct this evaluation, the article synthesizes existing literature and establishes a taxonomy encompassing module-level changes—such as learnable and timestamp encodings, sparse attention, and token patching—as well as architecture-level designs. In addition, the article presents targeted empirical experiments using benchmark operational data (ETTm2) to examine model robustness across sequence lengths, network depth, and structural decompositions.
The findings indicate that standard Transformers face significant practical hurdles in time series environments. First, directly increasing input sequence length causes rapid performance deterioration in many models, showing that high computational capacity does not automatically translate into effective long-sequence utilization. Second, unlike text or vision domains where deeper networks excel, shallower time series Transformers with only 3 to 6 layers consistently outperform deeper variants, which suffer from memory bottlenecks and degradation. Third, integrating seasonal-trend decomposition improves model forecasting accuracy by 50% to 82% across diverse configurations, demonstrating that isolating periodic signals is vital. Finally, recent patch-based and frequency-domain designs achieve linear calculation scaling while outperforming simpler baselines.
These results demonstrate that standard Transformers cannot simply be repurposed for sequential temporal data without domain-specific customization. Unmodified implementations risk high computational costs and memory failure without accuracy gains. To mitigate operational and computational risk, organizations should prioritize hybrid architectures that separate trends from periodic seasonality and leverage efficient attention mechanisms.
Stakeholders deploying deep learning for temporal analytics should integrate seasonal-trend decomposition into their modeling pipelines and avoid oversized, deep networks. When scaling to high-dimensional or long-sequence tasks, engineering teams should consider efficient token segmentation (such as patching) and explore combinations with graph neural networks for spatial dependencies. Automated architecture discovery should be piloted to systematically tune hyper-parameters for large-scale production.
While the review provides comprehensive qualitative categorization, the direct experimental evaluations are limited to a single benchmark dataset, requiring caution when generalizing findings across diverse industrial environments. Nevertheless, there is high confidence in the fundamental conclusions regarding model sizing and seasonal decomposition, which are reinforced by broader cross-study literature evidence.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Introduces the original Transformer architecture and multi-head self-attention mechanism that serve as the foundational baseline reviewed throughout the survey.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). Proposes the Informer architecture with ProbSparse attention for long-sequence time-series forecasting, representing one of the core canonical models analyzed in the survey.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). Presents Autoformer, introducing progressive series decomposition and auto-correlation mechanisms that are heavily examined in the survey's structural decomposition analysis.
- Paper: Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting, SHIYANG LI et al. (2019). Demonstrates early critical adaptations of Transformers to time-series forecasting by incorporating causal convolutions for locality and LogSparse attention to break memory bottlenecks.
- Paper: Temporal Fusion Transformers for interpretable multi-horizon time series forecasting, Bryan Lim et al. (2021). Introduces the Temporal Fusion Transformer for interpretable multi-horizon time series forecasting, serving as a primary benchmark for specialized attention mechanisms.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). Provides a broader foundational survey of deep learning architectures in time-series forecasting, establishing the context preceding the rapid expansion of Transformer-specific models.
- Paper: Generating Long Sequences with Sparse Transformers, Rewon Child et al. (2019). Develops factorized sparse attention patterns to scale sequence processing to long horizons, providing algorithmic foundations for efficient time-series Transformers.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). Establishes locality-sensitive hashing and reversible layers for efficient Transformer scaling, directly influencing subsequent efficient sequence modeling techniques in time series.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). Directly critiques the effectiveness of complex Transformer models surveyed in the literature by demonstrating that simple linear decomposition baselines can outperform them on long-term forecasting benchmarks.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). Reinvents Transformer modeling for time series post-survey by tokenizing sub-series patches and applying channel independence to dramatically improve forecasting accuracy and efficiency.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Inverts the standard tokenization paradigm surveyed by embedding whole individual variates as tokens, effectively resolving multivariate entanglement in time series forecasting.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). Explores a non-Transformer architectural paradigm following the survey's insights by transforming 1D time series into 2D variations to capture multi-periodic temporal dynamics.
- Paper: N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting, Cristian Challu et al. (2023). Advances deep neural forecasting beyond the reviewed Transformer approaches using multi-rate sampling and hierarchical interpolation to achieve higher accuracy and lower computational cost.
