Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions
Hansika HewamalageChristoph BergmeirKasun Bandara
Establishes practical guidelines and an open-source framework for time series forecasting with recurrent neural networks through extensive empirical comparisons against standard statistical baselines like ARIMA and ETS.
Organizations increasingly rely on large time-series databases to guide operational and strategic decisions. While traditional statistical benchmarks like ARIMA and ETS are popular because they are automated, efficient, and robust for single series, they cannot easily learn patterns across multiple related series simultaneously. Although modern recurrent neural networks (RNNs) have demonstrated winning performance in forecasting competitions, they have historically lacked standardized frameworks, clear hyperparameter tuning guidelines, and automated workflows suitable for non-expert business practitioners.
The article systematically evaluates whether standard, off-the-shelf RNN architectures can serve as competitive, semi-automated alternatives to classical statistical benchmarks for pure univariate forecasting without requiring complex, hand-crafted competition architectures.
To assess performance, the authors conducted a large-scale empirical evaluation comparing 36 distinct RNN configurations against established ARIMA and ETS implementations across six benchmark datasets (including M3, M4, NN5, CIF 2016, Tourism, and Wikipedia Web Traffic). The evaluation analyzed multiple recurrent unit types (LSTM with peephole connections, GRU, and basic Elman cells), architectures (Stacked and Sequence-to-Sequence variants), and optimization algorithms. All models were trained globally across time series using automated hyperparameter optimization via Sequential Model-Based Algorithm Configuration (SMAC) across 50 iterations, with parameter uncertainty managed through multi-seed ensembling.
The study yielded several key findings: First, global RNN models are highly competitive and outperformed traditional statistical benchmarks across most datasets (such as CIF 2016, NN5, and Wikipedia), though ARIMA retained superior accuracy on the highly heterogeneous M4 monthly dataset. Second, the best overall architecture was the Stacked RNN with Long Short-Term Memory (LSTM) units and the Continuous Coin Betting (COCOB) optimizer; COCOB proved especially advantageous because it eliminates the need to manually tune initial learning rates. Third, RNNs can model seasonality directly only when series within a dataset share homogeneous seasonal profiles, identical lengths, and aligned dates (as in NN5); across heterogeneous datasets, prior seasonal decomposition (such as STL) is essential to achieve competitive accuracy. Fourth, Sequence-to-Sequence models with standard autoregressive decoders suffered from cumulative error propagation, making direct projection via dense neural layers significantly more effective. Finally, while RNNs required substantially more total execution time—driven largely by the automated hyperparameter search—their final model training and inference phases were computationally practical and comparable to fitting thousands of individual univariate models.
These findings demonstrate that organizations managing large collections of related time series can achieve superior forecast accuracy by deploying global RNN models. The ability of global models to pool cross-series information reduces the operational maintenance burden of storing and individually refitting thousands of localized statistical models. Furthermore, the viability of self-tuning optimizers like COCOB substantially lowers the technical barrier for non-specialist teams seeking to operationalize deep learning pipelines.
Practitioners looking to implement deep learning forecasts should adopt the Stacked LSTM architecture with a moving-window input strategy, automated hyperparameter tuning, and the COCOB optimizer. Standard preprocessing pipelines should default to applying variance stabilization (log transformation) and deterministic seasonal decomposition unless data streams are strictly homogeneous. Organizations should weigh the upfront computational investment of hyperparameter optimization against the downstream accuracy gains and global model maintenance efficiencies.
The results should be interpreted within the article's defined scope: the study evaluated point forecasts for single-seasonality univariate series and excluded multivariate drivers, hierarchical constraints, or probabilistic prediction intervals. In addition, global RNN models occasionally exhibited vulnerability to extreme outlier series in heterogeneous datasets. Despite these boundaries, confidence in the primary findings remains high given the extensive scope of empirical testing and rigorous statistical significance analysis across diverse domains.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Establishes the foundational Long Short-Term Memory architecture that underpins the recurrent neural network forecasting models evaluated in the source.
- Paper: A state space framework for automatic forecasting using exponential smoothing methods, Rob J. Hyndman et al. (2002). Introduces the automated state-space exponential smoothing framework that serves as one of the primary classical statistical benchmarks in the source.
- Paper: Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks, Guokun Lai et al. (2017). Provides a key prior neural architecture that combines recurrent and convolutional components with autoregressive baselines to capture complex temporal patterns.
- Paper: Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling, Junyoung Chung et al. (2014). Offers essential empirical grounding on the comparative performance and mechanisms of gated recurrent units and LSTM cells for sequence modeling.
- Paper: On the difficulty of training recurrent neural networks, Razvan Pascanu et al. (2012). Clarifies the core gradient pathologies and training difficulties that motivate the practical modeling guidelines developed in the source.
- Paper: A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting, Slawek Smyl (2020). Directly realizes the source's findings on hybrid modeling and deseasonalization by combining dynamic exponential smoothing updates with dilated LSTM networks.
- Paper: DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks, David Salinas et al. (2020). Applies autoregressive recurrent neural networks to large-scale cross-series probabilistic forecasting, extending the source's univariate evaluation framework.
- Paper: N-BEATS: Neural basis expansion analysis for interpretable time series forecasting, Boris N. Oreshkin et al. (2020). Advances pure neural time-series forecasting beyond recurrent architectures by using deep residual stacks of basis function expansions.
- Paper: Temporal Fusion Transformers for interpretable multi-horizon time series forecasting, Bryan Lim et al. (2021). Extends recurrent sequence-to-sequence modeling by integrating self-attention and variable selection networks for interpretable multi-horizon forecasting.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). Provides a comprehensive survey generalizing the empirical findings and hybrid modeling principles established in the source across broader deep learning architectures.
- Paper: N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting, Cristian Challu et al. (2023). Builds upon modern non-recurrent deep forecasting methods by introducing hierarchical interpolation and multi-rate sampling for long-horizon prediction.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). Critically examines the shift toward complex transformer-based deep sequence forecasters, echoing the source's comparative analysis against simple linear baselines.
