Time-series forecasting with deep learning: a survey
Bryan LimStefan Zohren
Categorizes modern deep learning and hybrid statistical architectures for multi-horizon time series forecasting, providing a clear guide on how different models encode temporal dynamics for operational decision support.
Accurate time series forecasting is critical across multiple domains, including retail demand planning, financial risk management, energy operations, and healthcare diagnostics. While traditional statistical approaches rely on fixed mathematical assumptions and manual feature engineering, modern organizations increasingly generate vast, complex temporal datasets that require automated, scalable forecasting solutions. Recent advances in computing infrastructure and open-source software have spurred the widespread adoption of deep learning architectures, yet decision-makers frequently face challenges when selecting appropriate model designs, managing risks associated with model over-fitting, and extracting actionable insights from complex systems.
The article surveys modern deep learning architectures applied to time series forecasting, evaluating how distinct network designs process temporal information across single-step and multi-horizon settings. It specifically examines the emergence of hybrid models that combine domain-specific statistical frameworks with neural networks, while assessing advanced techniques that facilitate decision support through model interpretability and counterfactual analysis.
The review synthesizes findings across a broad spectrum of research literature and empirical forecasting competitions. It categorizes foundational neural network building blocks—such as convolutional neural networks, recurrent neural networks, and attention-based Transformer models—and details their statistical equivalents, including autoregressive formulations and Bayesian filtering. The article also reviews standard mechanisms for generating point forecasts and full predictive distributions, comparing direct sequence-to-sequence approaches against recursive iterative forecasting.
The analysis yields four central findings. First, hybrid architectures that integrate classical statistical methods with deep learning consistently outperform pure statistical or standalone machine learning models; notably, a hybrid model combining exponential smoothing with recurrent networks won the major M4 forecasting competition. Second, attention mechanisms and Transformer networks effectively capture long-range dependencies and multi-regime dynamics without suffering from the gradient instability or memory degradation historically observed in recurrent architectures. Third, direct multi-horizon forecasting via encoder-decoder structures mitigates the error accumulation inherent in recursive, step-by-step predictions while accommodating known future variables. Fourth, deep neural networks can support strategic decision-making beyond basic forecasting through built-in attention weights that expose important historical events, as well as specialized causal inference techniques that adjust for time-dependent confounding to model counterfactual scenarios.
These findings indicate that organizations can improve forecast accuracy and operational resilience by moving away from purely data-driven black-box models or rigid standalone statistical models. Adopting hybrid designs significantly reduces the risk of over-fitting in low-data environments, handles non-stationary trends, and eliminates burdensome data pre-processing steps. Furthermore, incorporating probabilistic distributions and counterfactual forecasting enables leaders to quantify operational risks, guard against rare tail events, and run scenario analyses before committing resources.
Organizations developing temporal forecasting pipelines should prioritize hybrid architectures, using domain knowledge to constrain model search spaces and improve generalization. When forecasts span multiple future periods, engineering teams should implement direct sequence-to-sequence models with attention layers rather than recursive one-step predictors. Additionally, deployment strategies should integrate inherent interpretability features and causal adjustment techniques to provide end-users with justifiable, actionable scenario forecasts.
The article highlights key limitations across current deep learning methodologies. Most existing architectures assume regular, discretely sampled time intervals, which creates operational vulnerabilities when dealing with irregularly sampled or missing data. Furthermore, standard models often fail to account for hierarchical structures, such as regional and product-level groupings. While the authors demonstrate high confidence in the evaluated architectures, they note that emerging frameworks—such as continuous-time neural ordinary differential equations and hierarchical deep networks—require further benchmarking and validation before broad operational deployment.
- Paper: Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks, Guokun Lai et al. (2017). LSTNet establishes a foundational multi-component deep learning architecture combining CNNs, RNNs, and classical autoregressive components that the survey directly reviews and synthesizes.
- Paper: A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting, Slawek Smyl (2020). Smyl's hybrid exponential smoothing and recurrent neural network architecture serves as the prime exemplar of the hybrid statistical-neural approaches highlighted in the survey.
- Paper: An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling, Shaojie Bai et al. (2018). This paper introduces generic temporal convolutional networks and systematically benchmarks convolutional sequence modeling against recurrent networks, providing essential background for the survey's encoder-decoder categorization.
- Paper: Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting, SHIYANG LI et al. (2019). It provides crucial architectural adaptations of Transformers for time-series forecasting by tackling locality and memory bottlenecks, setting the groundwork for attention-based models discussed in the survey.
- Paper: Spatio-temporal Graph Convolutional Neural Network: A Deep Learning Framework for Traffic Forecasting, Bing Yu et al. (2017). This work introduces spatio-temporal graph neural networks for time-series forecasting, providing the core framework for multi-series spatial-temporal dependency modeling covered in the survey.
- Paper: A Critical Review of Recurrent Neural Networks for Sequence Learning, Zachary C. Lipton et al. (2015). This review delivers essential foundational principles and mathematical formulations of recurrent neural networks used across classical time-series encoders.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). It introduces the fundamental Long Short-Term Memory architecture that underpins the recurrent encoders and decoders analyzed in deep forecasting models.
- Paper: Temporal Fusion Transformers for interpretable multi-horizon time series forecasting, Bryan Lim et al. (2021). Temporal Fusion Transformers build directly upon the survey's multi-horizon forecasting and interpretability frameworks to create an end-to-end multi-horizon attention model.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). Informer extends Transformer-based time-series forecasting with efficient sparse attention mechanisms and generative decoding for extreme long-sequence multi-step forecasting.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). Autoformer advances the architectural paradigms surveyed by embedding deep trend-seasonal decomposition directly within Transformer auto-correlation mechanisms.
- Paper: FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series Forecasting, Tian Zhou et al. (2022). FEDformer builds upon frequency-domain and decomposition techniques in deep time series to scale long-horizon Transformer forecasting.
- Paper: N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting, Cristian Challu et al. (2023). N-HiTS advances neural basis expansion architectures by introducing hierarchical multi-rate sampling to conquer long-horizon time-series forecasting.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). PatchTST evolves time-series Transformer design beyond standard point-wise tokenization by applying patch-level tokenization and channel independence.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). TimesNet generalizes 1D temporal neural modeling by transforming 1D series into 2D variation spaces to unify multi-task temporal analysis.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). iTransformer inverts the conventional tokenization structure of attention models in forecasting, addressing key structural challenges highlighted in deep sequence modeling.
- Paper: Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift, Taesung Kim et al. (2022). RevIN introduces reversible instance normalization to directly resolve distribution shift and non-stationarity in deep forecasting networks.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). This paper critically evaluates the complex Transformer architectures surveyed in literature by demonstrating the surprising efficacy of minimal linear models.
