Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions

Hansika HewamalageChristoph BergmeirKasun Bandara

article2019International Journal of Forecasting1,313 citations

Establishes practical guidelines and an open-source framework for time series forecasting with recurrent neural networks through extensive empirical comparisons against standard statistical baselines like ARIMA and ETS.

Listen

Organizations increasingly rely on large time-series databases to guide operational and strategic decisions. While traditional statistical benchmarks like ARIMA and ETS are popular because they are automated, efficient, and robust for single series, they cannot easily learn patterns across multiple related series simultaneously. Although modern recurrent neural networks (RNNs) have demonstrated winning performance in forecasting competitions, they have historically lacked standardized frameworks, clear hyperparameter tuning guidelines, and automated workflows suitable for non-expert business practitioners.

The article systematically evaluates whether standard, off-the-shelf RNN architectures can serve as competitive, semi-automated alternatives to classical statistical benchmarks for pure univariate forecasting without requiring complex, hand-crafted competition architectures.

To assess performance, the authors conducted a large-scale empirical evaluation comparing 36 distinct RNN configurations against established ARIMA and ETS implementations across six benchmark datasets (including M3, M4, NN5, CIF 2016, Tourism, and Wikipedia Web Traffic). The evaluation analyzed multiple recurrent unit types (LSTM with peephole connections, GRU, and basic Elman cells), architectures (Stacked and Sequence-to-Sequence variants), and optimization algorithms. All models were trained globally across time series using automated hyperparameter optimization via Sequential Model-Based Algorithm Configuration (SMAC) across 50 iterations, with parameter uncertainty managed through multi-seed ensembling.

The study yielded several key findings: First, global RNN models are highly competitive and outperformed traditional statistical benchmarks across most datasets (such as CIF 2016, NN5, and Wikipedia), though ARIMA retained superior accuracy on the highly heterogeneous M4 monthly dataset. Second, the best overall architecture was the Stacked RNN with Long Short-Term Memory (LSTM) units and the Continuous Coin Betting (COCOB) optimizer; COCOB proved especially advantageous because it eliminates the need to manually tune initial learning rates. Third, RNNs can model seasonality directly only when series within a dataset share homogeneous seasonal profiles, identical lengths, and aligned dates (as in NN5); across heterogeneous datasets, prior seasonal decomposition (such as STL) is essential to achieve competitive accuracy. Fourth, Sequence-to-Sequence models with standard autoregressive decoders suffered from cumulative error propagation, making direct projection via dense neural layers significantly more effective. Finally, while RNNs required substantially more total execution time—driven largely by the automated hyperparameter search—their final model training and inference phases were computationally practical and comparable to fitting thousands of individual univariate models.

These findings demonstrate that organizations managing large collections of related time series can achieve superior forecast accuracy by deploying global RNN models. The ability of global models to pool cross-series information reduces the operational maintenance burden of storing and individually refitting thousands of localized statistical models. Furthermore, the viability of self-tuning optimizers like COCOB substantially lowers the technical barrier for non-specialist teams seeking to operationalize deep learning pipelines.

Practitioners looking to implement deep learning forecasts should adopt the Stacked LSTM architecture with a moving-window input strategy, automated hyperparameter tuning, and the COCOB optimizer. Standard preprocessing pipelines should default to applying variance stabilization (log transformation) and deterministic seasonal decomposition unless data streams are strictly homogeneous. Organizations should weigh the upfront computational investment of hyperparameter optimization against the downstream accuracy gains and global model maintenance efficiencies.

The results should be interpreted within the article's defined scope: the study evaluated point forecasts for single-seasonality univariate series and excluded multivariate drivers, hierarchical constraints, or probabilistic prediction intervals. In addition, global RNN models occasionally exhibited vulnerability to extreme outlier series in heterogeneous datasets. Despite these boundaries, confidence in the primary findings remains high given the extensive scope of empirical testing and rigorous statistical significance analysis across diverse domains.

arXiv: 1909.00590
Cover for Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions

Abstract

Recurrent Neural Networks (RNN) have become competitive forecasting methods, as most notably shown in the winning method of the recent M4 competition. However, established statistical models such as ETS and ARIMA gain their popularity not only from their high accuracy, but they are also suitable for non-expert users as they are robust, efficient, and automatic. In these areas, RNNs have still a long way to go. We present an extensive empirical study and an open-source software framework of existing RNN architectures for forecasting, that allow us to develop guidelines and best practices for their use. For example, we conclude that RNNs are capable of modelling seasonality directly if the series in the dataset possess homogeneous seasonal patterns, otherwise we recommend a deseasonalization step. Comparisons against ETS and ARIMA demonstrate that the implemented (semi-)automatic RNN models are no silver bullets, but they are competitive alternatives in many situations.

Table of Contents

  • 1 Introduction
  • 2 Background Study
  • 2.1 Univariate Forecasting
  • 2.2 Traditional Univariate Forecasting Techniques
  • 2.3 Artificial Neural Networks
  • 2.3.1 Recurrent Neural Networks for Forecasting
  • 2.4 Leveraging Cross-Series Information
  • 3 Methodology
  • 3.1 Recurrent Neural Networks
  • 3.1.1 Recurrent Units
  • 3.1.2 Recurrent Neural Network Architectures
  • 3.2 Learning Algorithms
  • 4 Experimental Framework
  • 4.1 Datasets
  • 4.2 Data Preprocessing
  • 4.2.1 Dataset Split
  • 4.2.2 Addressing Missing Values
  • 4.2.3 Modelling Seasonality
  • 4.2.4 Stabilizing the Variance in the Data
  • 4.2.5 Multiple Output Strategy
  • 4.2.6 Trend Normalization
  • 4.2.7 Mean Normalization
  • 4.3 Training & Validation Scheme
  • 4.3.1 Hyper-parameter Tuning
  • 4.3.2 Dealing with Model Overfitting
  • 4.4 Model Testing
  • 4.4.1 Data Post-Processing
  • 4.4.2 Performance Measures
  • 4.5 Benchmarks
  • 4.6 Statistical Tests of the Results
  • 5 Analysis of Results
  • 5.1 Relative Performance of RNN Architectures
  • 5.2 Performance of Recurrent Units
  • 5.3 Performance of Optimizers
  • 5.4 Performance of the Output Components for the Sequence to Sequence Architecture
  • 5.5 Comparison of Input Window Sizes for the Stacked Architecture
  • 5.6 Analysis of Seasonality Modelling
  • 5.7 Performance of RNN Models Vs. Traditional Univariate Benchmarks
  • 5.8 Experiments Involving the Total Number of Trainable Parameters
  • 5.9 Comparison of the Computational Costs of the RNN Models Vs. the Benchmarks
  • 5.10 Hyperparameter Configurations
  • 6 Conclusions
  • 7 Future Directions
  • References

Knowls

  1. Knowl 1 — Global RNN Time Series Forecasting Pipeline with STL Preprocessing

    model/method

    An end-to-end framework for multi-step univariate time series forecasting across a collection of related series using a global Recurrent Neural Network (RNN) with Seasonal and Trend decomposition using Loess (STL) is structured as follows:

    1. Missing Value Imputation: Missing values are replaced by the median value of the corresponding seasonal period across the entire series (for example, taking the median across all Tuesdays for daily data with weekly seasonality). If zeroes and missing values are indistinguishable in the raw data, zero-substitution is applied.
    2. Variance Stabilization: For non-negative series yy, values are transformed via: wt={log⁡(yt)if min⁡(y)>ϵlog⁡(yt+1)if min⁡(y)≤ϵw_t = \begin{cases} \log(y_t) & \text{if } \min(y) > \epsilon \\ \log(y_t + 1) & \text{if } \min(y) \le \epsilon \end{cases} where ϵ>0\epsilon > 0 is a small positive threshold for real-valued data or ϵ=0\epsilon = 0 for count data.
    3. Deterministic Deseasonalization: Additive STL decomposition with a periodic seasonal window decomposes wtw_t into deterministic seasonality StS_t, trend TtT_t, and remainder RtR_t. The deseasonalized series is dt=wt−Std_t = w_t - S_t.
    4. Moving Window Segmentation: The deseasonalized sequence is framed into consecutive input windows Xi∈RmX_i \in \mathbb{R}^m of length mm and target output windows Yi∈RHY_i \in \mathbb{R}^H of forecast horizon length HH.
    5. Local Trend Normalization: The trend value at the last time step of the input window, Tlast,iT_{\text{last}, i}, is deducted from both the input and output windows: X~i=Xi−Tlast,i,Y~i=Yi−Tlast,i\tilde{X}_i = X_i - T_{\text{last}, i}, \quad \tilde{Y}_i = Y_i - T_{\text{last}, i}
    6. Global Network Training: A single RNN is trained across all series in the dataset simultaneously using mini-batch gradient descent with L2 weight regularization and additive zero-mean Gaussian input noise.
    7. Post-Processing Forecasts: The predicted output Y~^∈RH\hat{\tilde{Y}} \in \mathbb{R}^H is inverted back to the original scale: Y^=exp⁡(Y~^+Tlast+Sforecast)\hat{Y} = \exp\left(\hat{\tilde{Y}} + T_{\text{last}} + S_{\text{forecast}}\right) (subtracting 11 if the shifted log was applied), rounding to the nearest integer for integer count series, and clipping negative predictions at 00.
    8. Ensembling: Predictions are generated across 1010 distinct random network weight initializations (seeds) and aggregated by computing the element-wise median across seeds.
  2. Knowl 2 — Seasonal Homogeneity Principle for Neural Network Time Series Forecasting

    empirical result

    Whether Recurrent Neural Networks (RNNs) require prior deseasonalization depends on the cross-series homogeneity of the dataset's seasonal patterns:

    1. Homogeneous Seasonality: When all time series in a dataset share the same seasonal shape, equal sequence lengths, and synchronized calendar start and end dates (such as daily ATM cash withdrawal data), global RNNs successfully learn seasonality directly without prior deseasonalization. Under these conditions, omitting STL decomposition yields comparable or slightly superior performance (average mean SMAPE rank of 1.441.44 without STL vs. 1.561.56 with STL; paired Wilcoxon signed-rank test p=0.911p = 0.911).
    2. Heterogeneous Seasonality: When time series exhibit diverse seasonal profiles, varying lengths, or non-aligned start/end dates (such as the CIF 2016, M3 Monthly, M4 Monthly, and Tourism datasets), global RNNs fail to model seasonality internally. Prior deseasonalization via deterministic STL decomposition significantly improves forecast accuracy across models:
      • CIF 2016: with-STL rank 1.01.0 vs. without-STL rank 2.02.0 (p=2.91×10−11p = 2.91 \times 10^{-11}).
      • M3 Monthly: with-STL rank 1.01.0 vs. without-STL rank 2.02.0 (p=2.91×10−11p = 2.91 \times 10^{-11}).
      • Tourism Monthly: with-STL rank 1.061.06 vs. without-STL rank 1.941.94 (p=1.455×10−10p = 1.455 \times 10^{-10}).
    3. Negligible Seasonality: For datasets where time series possess minimal seasonality (such as Wikipedia Web Traffic data), applying STL decomposition adds noise and significantly degrades performance compared to training directly on raw scaled series (with-STL rank 1.731.73 vs. without-STL rank 1.271.27, p=0.028p = 0.028).
  3. Knowl 3 — Comparative Forecasting Accuracy: Global RNNs vs. Statistical Benchmarks and Linear Regressions

    data/table

    Across six diverse forecasting benchmark datasets, global RNN models (specifically Stacked LSTM architectures with COCOB or Adam optimizers and moving window inputs) were evaluated against classical univariate benchmarks (auto.arima and ets from the R forecast package) as well as linear autoregressive (AR) models (pooled global AR and unpooled local AR with 10 lags versus moving window lags):

    Model CIF Kaggle WT M3 NN5 Tourism M4
    auto.arima 11.70 47.96 14.25 25.91 19.74 13.08
    ets 11.88 53.50 14.14 21.57 19.02 13.53
    Pooled Reg. (Lags 10) 14.67 122.43 14.73 30.21 31.29 14.18
    Pooled Reg. (Window Lags) 12.89 47.47 14.36 22.41 21.10 13.75
    Unpooled Reg. (Window Lags) 14.34 56.75 14.93 22.63 20.71 13.46
    S2SD LSTM NMW cocob 10.23 49.74 14.31 27.82 20.29 13.66
    Stacked LSTM cocob 10.40 47.14 14.44 21.61 21.12 14.13
    Stacked LSTM adam 10.51 47.13 14.39 21.53 22.28 14.23
    NSTL Stacked GRU adam – 45.62 – 23.47 – –

    Key empirical findings:

    • On CIF 2016, Kaggle Web Traffic, and NN5 datasets, top RNN models outperform both Auto.ARIMA and ETS in mean SMAPE.
    • On Tourism and M3 Monthly, statistical univariate methods achieve slightly lower mean SMAPE, though RNNs outperform ETS and Auto.ARIMA on median SMAPE for Tourism, indicating that mean errors are driven by isolated extreme outlier series.
    • On M4 Monthly, Auto.ARIMA achieves the lowest error overall (13.0813.08 SMAPE), but global RNNs outperform Auto.ARIMA and ETS on the M4 Micro subcategory.
    • Pooling across series improves linear AR performance when moving window lag counts match RNN input sizes, but nonlinear RNN architectures consistently outperform linear pooled regressions across datasets.
  4. Knowl 4 — Superiority of Direct Dense Output Projections Over Autoregressive Decoders in RNN Forecasting

    empirical result

    In multi-step time series forecasting with Sequence-to-Sequence (S2S) neural architectures, direct projection via a global dense layer significantly outperforms an autoregressive recurrent decoder.

    Key findings include:

    1. Error Accumulation in Decoders: Standard S2S decoders with autoregressive feedback connections (where the forecast at step t−1t-1 serves as the input at step tt) suffer from cascading error accumulation across multi-step horizons, despite teacher forcing during training and scheduled sampling during inference.
    2. Dense Layer Superiority: Replacing the recurrent decoder with a single global affine dense layer (mapping the encoder's final hidden state hTh_T directly to RH\mathbb{R}^H) yields statistically significant improvements in forecasting accuracy (paired Wilcoxon signed-rank test on mean SMAPE gives p=2.064×10−4p = 2.064 \times 10^{-4}).
    3. Overall Architectural Ranking: Across ranked evaluation metrics (MASE ranks and median SMAPE ranks), the Stacked RNN architecture performs best overall, followed closely by the S2S with Dense Layer (S2SD) architecture. S2S with an autoregressive decoder ranks worst across all tested configurations.
  5. Knowl 5 — Moving Window Multi-Input Multi-Output (MIMO) Strategy for Global Time Series RNNs

    model/method

    The moving window Multi-Input Multi-Output (MIMO) strategy structures sequential time series data into overlapping fixed-size blocks to train global RNNs for multi-step-ahead forecasting without recursive error accumulation.

    For a time series of length ll, target forecasting horizon HH, input window size mm, and output window size n=Hn = H:

    1. The final block of length nn is held out for validation during hyperparameter selection.
    2. The remaining sequence of length l−nl - n is sliced into overlapping input-output blocks of length m+nm + n by sliding a window forward by 11 time step at each recurrent transition:
      • Input window at recurrent step tt: Xt=[xt,xt+1,…,xt+m−1]⊤∈RmX_t = [x_t, x_{t+1}, \dots, x_{t+m-1}]^\top \in \mathbb{R}^m.
      • Target output window at recurrent step tt: Yt=[xt+m,xt+m+1,…,xt+m+n−1]⊤∈RnY_t = [x_{t+m}, x_{t+m+1}, \dots, x_{t+m+n-1}]^\top \in \mathbb{R}^n.
    3. The total number of training blocks obtained per series is l−2n−ml - 2n - m, providing an effective data augmentation mechanism for short sequences.
    4. Input window size mm is parameterized using heuristics:
      • Large window: m=1.25×Hm = 1.25 \times H (provides sufficient historical context up to or beyond the forecast horizon).
      • Small window: m=1.25×sm = 1.25 \times s, where ss is the seasonality period (ensures at least one full seasonal cycle is visible per step).
    5. At inference, the final recurrent state produces an HH-dimensional forecast vector simultaneously through an affine dense layer (WdhT+bd∈RHW_d h_T + b_d \in \mathbb{R}^H), preserving inter-temporal dependencies across the forecast horizon.
  6. Knowl 6 — RNN Architectural Paradigms for Univariate Time Series Forecasting

    model/method

    Univariate multi-step time series forecasting can be implemented using three distinct RNN architectural paradigms:

    1. Stacked RNN Architecture:

      • Multiple recurrent layers stacked vertically.
      • At each recurrent step tt, the network receives a moving window vector Xt∈RmX_t \in \mathbb{R}^m and emits an HH-dimensional prediction Y^t=Wdhk,t+bd∈RH\hat{Y}_t = W_d h_{k,t} + b_d \in \mathbb{R}^H from the top layer kk via an affine dense layer.
      • Error is accumulated across all TT sequence steps: E=∑t=1T∣Yt−Y^t∣E = \sum_{t=1}^T |Y_t - \hat{Y}_t|
      • Backpropagation Through Time (BPTT) updates weights using the full accumulated sequence error.
    2. Sequence-to-Sequence (S2S) with Autoregressive Decoder:

      • An Encoder RNN processes scalar inputs xt∈Rx_t \in \mathbb{R} step-by-step to produce a final context vector hTh_T.
      • A Decoder RNN unfolds over HH future steps initialized with hTh_T. During training, teacher forcing supplies actual past values; during inference, autoregressive connections feed the previous step's forecast y^t−1\hat{y}_{t-1} into the next step.
      • Error is computed strictly over the forecast horizon: E=∑t=1H∣yt−y^t∣E = \sum_{t=1}^H |y_t - \hat{y}_t|
    3. Sequence-to-Sequence with Dense Layer (S2SD):

      • An Encoder RNN processes the input sequence (either as scalar inputs without moving windows or as moving window vectors).
      • The final hidden state hTh_T is mapped directly to the full HH-step forecast Y^T∈RH\hat{Y}_T \in \mathbb{R}^H via a global fully connected dense layer without bias, eliminating the multi-step decoder.
      • Error is computed solely on the final forecast: E=∣YT−Y^T∣E = |Y_T - \hat{Y}_T|
  7. Knowl 7 — LSTM Cell with Peephole Connections Formulation

    equation

    The Long Short-Term Memory (LSTM) recurrent unit with peephole connections allows the input, forget, and output gates to directly inspect the cell state Ct−1C_{t-1} and CtC_t:

    it=σ(Wiht−1+Vixt+PiCt−1+bi)i_t = \sigma(W_i h_{t-1} + V_i x_t + P_i C_{t-1} + b_i) ot=σ(Woht−1+Voxt+PoCt+bo)o_t = \sigma(W_o h_{t-1} + V_o x_t + P_o C_t + b_o) ft=σ(Wfht−1+Vfxt+PfCt−1+bf)f_t = \sigma(W_f h_{t-1} + V_f x_t + P_f C_{t-1} + b_f) C~t=tanh⁡(Wcht−1+Vcxt+bc)\tilde{C}_t = \tanh(W_c h_{t-1} + V_c x_t + b_c) Ct=it⊙C~t+ft⊙Ct−1C_t = i_t \odot \tilde{C}_t + f_t \odot C_{t-1} ht=ot⊙tanh⁡(Ct)h_t = o_t \odot \tanh(C_t) zt=htz_t = h_t

    where:

    • xt∈Rmx_t \in \mathbb{R}^m is the input vector at time step tt.
    • ht∈Rdh_t \in \mathbb{R}^d is the hidden state vector of cell dimension dd.
    • Ct∈RdC_t \in \mathbb{R}^d is the internal cell state vector capturing long-term dependencies.
    • C~t∈Rd\tilde{C}_t \in \mathbb{R}^d is the candidate cell state vector.
    • it,ot,ft∈[0,1]di_t, o_t, f_t \in [0, 1]^d are the input, output, and forget gate activation vectors.
    • Wi,Wo,Wf,Wc∈Rd×dW_i, W_o, W_f, W_c \in \mathbb{R}^{d \times d} are recurrent weight matrices for the previous hidden state.
    • Vi,Vo,Vf,Vc∈Rd×mV_i, V_o, V_f, V_c \in \mathbb{R}^{d \times m} are input weight matrices.
    • Pi,Po,Pf∈Rd×dP_i, P_o, P_f \in \mathbb{R}^{d \times d} are peephole weight matrices connecting the cell state directly to the gate activations.
    • bi,bo,bf,bc∈Rdb_i, b_o, b_f, b_c \in \mathbb{R}^d are bias vectors.
    • σ(u)=11+e−u\sigma(u) = \frac{1}{1 + e^{-u}} is the element-wise logistic sigmoid function.
    • ⊙\odot denotes the element-wise Hadamard product.
    • zt∈Rdz_t \in \mathbb{R}^d is the cell output, which is equal to the hidden state hth_t.
  8. Knowl 8 — Empirical Performance Comparison of Recurrent Cell Units Under Fixed Dimension and Equalized Capacity

    empirical result

    An empirical evaluation comparing Long Short-Term Memory with peephole connections (LSTM), Gated Recurrent Units (GRU), and Elman Recurrent Units (ERNN) across diverse time series datasets demonstrates:

    1. Cell Dimension Tuning: When the cell dimension dd is tuned as a hyperparameter, LSTM with peephole connections achieves the best mean SMAPE and mean MASE ranks, followed by GRU, while ERNN performs the worst (Friedman test p=0.101p = 0.101).
    2. Equal Trainable Parameter Capacity: Because LSTM cells have 44 gate weight matrices plus peephole connections, GRU has 33 gate matrices, and ERNN has 11 transition matrix, identical cell dimensions give different model capacities. When models are re-evaluated by directly tuning the total number of trainable parameters over a shared range of 2,0002,000 to 25,00025,000 (setting the corresponding cell dimension dd per unit type):
      • LSTM with peepholes maintains superior ranking across datasets (CIF, Kaggle Web Traffic, M3, and NN5).
      • GRU maintains intermediate performance.
      • ERNN consistently ranks worst (Friedman test p=0.174p = 0.174).

    This demonstrates that the performance superiority of peephole LSTM and GRU over ERNN arises from internal gating mechanisms mitigating gradient degradation over long sequences, rather than merely possessing a larger parameter count.

  9. Knowl 9 — Evaluation of Optimization Algorithms for Automated Global RNN Forecasting

    empirical result

    Evaluation of three optimization algorithms (Continuous Coin Betting [COCOB], Adam, and Adagrad) for training global RNN forecasting models yields:

    1. COCOB Optimality: The COCOB optimizer achieves the best overall performance in both mean SMAPE ranks and mean MASE ranks across datasets. COCOB dynamically adapts step sizes per parameter based on a coin-betting framework that maximizes wealth from negative subgradient feedback without requiring any learning rate hyperparameter.
    2. Hyperparameter Elimination: Eliminating the learning rate search space relieves the automated hyperparameter tuning algorithm (such as SMAC) from exploring learning rate grids/ranges, making fully automated forecasting pipelines significantly more robust and computationally efficient.
    3. Adam vs. Adagrad: Adam performs competitively in second place (with optimal learning rates typically between 0.0010.001 and 0.10.1). Adagrad performs worst across all benchmarks due to monotonically accumulating squared gradients in its denominator, which excessively shrinks the effective learning rate over long training sequences and stalls convergence.
  10. Knowl 10 — Effect of Input Window Size on Multi-Step RNN Forecasting Accuracy

    empirical result

    In the moving-window Stacked RNN architecture, the length of the input window mm significantly influences forecast accuracy across multi-step horizons:

    1. Large vs. Small Window Configurations: Evaluated on high-frequency daily datasets (NN5 with a 56-day horizon and Wikipedia Web Traffic with a 59-day horizon):
      • Large window: m=1.25×Hm = 1.25 \times H (m=70m = 70 for NN5, m=74m = 74 for Wikipedia).
      • Small window: m=1.25×s=9m = 1.25 \times s = 9 (where s=7s = 7 is the weekly seasonal period).
    2. Statistical Significance: Sizing the input window to exceed the forecast horizon (1.25×H1.25 \times H) yields statistically significantly lower mean SMAPE errors than sizing it to the base seasonal period (1.25×s1.25 \times s), with a paired Wilcoxon signed-rank test p=2.206×10−3p = 2.206 \times 10^{-3} across with-STL and without-STL configurations.
    3. Interaction with Seasonality: When raw series are fed directly to the RNN without prior deseasonalization, large input windows produce substantial accuracy gains because they provide sufficient sequential depth for recurrent memory states to capture recurring seasonal cycles.
  11. Knowl 11 — Regularized Objective Function and Hyperparameter Optimization with SMAC

    model/method

    To prevent overfitting in global RNNs trained across heterogeneous time series collections, model training combines input noise distortion, L2 weight regularization, and automated hyperparameter configuration using Sequential Model-based Algorithm Configuration (SMAC):

    1. Regularized Loss Function: L=1H∑k=1H∣Fk−Yk∣+ψ∑i=1pwi2\mathcal{L} = \frac{1}{H} \sum_{k=1}^H |F_k - Y_k| + \psi \sum_{i=1}^p w_i^2 where 1H∑k=1H∣Fk−Yk∣\frac{1}{H} \sum_{k=1}^H |F_k - Y_k| is the Mean Absolute Error (MAE) between predicted values FkF_k and actual values YkY_k over horizon HH, wiw_i denotes the ii-th trainable weight among pp total network parameters, and ψ\psi is the L2 penalty coefficient.
    2. Input Noise Injection: Zero-mean Gaussian noise N(0,σnoise2)\mathcal{N}(0, \sigma_{\text{noise}}^2) is added to input sequences during training, where standard deviation σnoise∈[0.0001,0.0008]\sigma_{\text{noise}} \in [0.0001, 0.0008] is tuned as a hyperparameter.
    3. SMAC Optimization: Hyperparameters are optimized over 50 iterations using SMAC with Tree-structured Parzen Estimators (TPE) evaluated on a held-out terminal sequence partition of length HH:
      • Number of hidden layers: {1,2}\{1, 2\} (shallow architectures consistently avoid overfitting).
      • Cell dimension dd: [20,50][20, 50].
      • Minibatch size: scaled to dataset volume (approximately 1/101/10-th of series count, e.g., 10–3010\text{--}30 for CIF, 1,000–1,5001,000\text{--}1,500 for M4 subsets).
      • Epochs: [3,30][3, 30]; Epoch size: [2,20][2, 20].
      • Weight initialization: Normal distribution N(0,σinit2)\mathcal{N}(0, \sigma_{\text{init}}^2) with σinit∈[0.0001,0.0008]\sigma_{\text{init}} \in [0.0001, 0.0008].
      • L2 coefficient ψ∈[0.0001,0.0008]\psi \in [0.0001, 0.0008].
  12. Knowl 12 — Computational Cost Profile of Global RNN Pipelines vs. Univariate Statistical Benchmarks

    data/table

    Computational execution times (in seconds on 25 CPU cores) comparing global RNN pipelines against auto.arima and ets on datasets where RNNs achieve superior accuracy:

    Dataset Model Preproc. Hyperparam Tuning Train Test Total Time
    CIF (12) S2SD LSTM NMW cocob 1.9 2531.8 1215.7 3749.5
    CIF (12) auto.arima – – – 85.1
    CIF (12) ets – – – 41.6
    Kaggle WT NSTL Stacked GRU adam 159.7 12803.2 5350.2 18313.0
    Kaggle WT auto.arima – – – 834.4
    Kaggle WT ets – – – 481.7
    M3 (Micro) S2SD GRU MW adagrad 30.2 1927.2 565.6 2523.0
    M3 (Micro) auto.arima – – – 849.8
    M3 (Micro) ets – – – 73.0
    NN5 Stacked LSTM adam 42.2 37660.2 13788.3 51490.7
    NN5 auto.arima – – – 1067.7
    NN5 ets – – – 53.3

    Key takeaways:

    • Tuning Overhead: Automated hyperparameter optimization via SMAC (5050 iterations) dominates total RNN computation time, accounting for 65%–75%65\%\text{--}75\% of overall wall-clock runtime.
    • Inference and Global Simplicity: Once optimal hyperparameters are identified, training a single global RNN model across thousands of series requires comparable or lower execution time than fitting thousands of separate Auto.ARIMA models (for example, 565.6 s565.6\,\text{s} for a global RNN on M3 Micro vs. 849.8 s849.8\,\text{s} for Auto.ARIMA).
    • The resulting global RNN possesses substantially fewer total parameters on a global scale than the combined parameter count of separate univariate statistical models fitted to each individual time series.

Coverage note — None was omitted; the knowls cover the full scope of the paper's proposed RNN pipeline, architectures, cell types, optimization schemes, seasonality findings, window size heuristics, regularized training, and empirical benchmark comparisons.

References

  1. 1.Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Man´e, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Vi´egas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., Zheng, X., 2015. TensorFlow: Large-scale machine learning on heterogeneous systems. Software available from tensorflow.org. URL https://www.tensorflow.org/
  2. 2.Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S. S., Salinas, D., Schulz, J., Stella, L., T¨urkmen, A. C., Wang, Y., 2019. Gluonts: Probabilistic time series models in python. CoRR abs/1906.05264. URL http://arxiv.org/abs/1906.05264
  3. 3.Assaad, M., Bon´e, R., Cardot, H., Jan. 2008. A new boosting algorithm for improved time-series forecasting with recurrent neural networks. Inf. Fusion 9 (1), 41–55.
  4. 4.Athanasopoulos, G., Hyndman, R., Song, H., Wu, D., 2011. The tourism forecasting competition. International Journal of Forecasting 27 (3), 822 – 844.
  5. 5.Athanasopoulos, G., Hyndman, R. J., Song, H., Wu, D., 2010. Tourism forecasting part two. URL https://www.kaggle.com/c/tourism2/data
  6. 6.Bahdanau, D., Cho, K., Bengio, Y., 2015. Neural machine translation by jointly learning to align and translate. In: Bengio, Y., LeCun, Y. (Eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings. URL http://arxiv.org/abs/1409.0473
  7. 7.Bai, S., Kolter, J. Z., Koltun, V., 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. CoRR abs/1803.01271. URL http://arxiv.org/abs/1803.01271
  8. 8.Bandara, K., Bergmeir, C., Smyl, S., 2020. Forecasting across time series databases using recurrent neural networks on groups of similar series: A clustering approach. Expert Systems with Applications 140, 112896.
  9. 9.Bandara, K., Shi, P., Bergmeir, C., Hewamalage, H., Tran, Q., Seaman, B., 2019. Sales demand forecast in e-commerce using a long short-term memory neural network methodology. In: Gedeon, T., Wong, K. W., Lee, M. (Eds.), Neural Information Processing. Springer International Publishing, Cham, pp. 462–474.
  10. 10.Bayer, J., Osendorfer, C., 2014. Learning stochastic recurrent networks. URL https://arxiv.org/abs/1411.7610
  11. 11.Ben Taieb, S., Bontempi, G., Atiya, A., Sorjamaa, A., 6 2012. A review and comparison of strategies for multi-step ahead time series forecasting based on the nn5 forecasting competition. Expert Systems with Applications 39 (8), 7067–7083.
  12. 12.Bergmeir, C., Hyndman, R. J., Koo, B., 2018. A note on the validity of cross-validation for evaluating autoregressive time series prediction. Computational Statistics & Data Analysis 120, 70 – 83.
  13. 13.Bergstra, J., 2012. Hyperopt: Distributed asynchronous hyper-parameter optimization. URL https://github.com/hyperopt/hyperopt
  14. 14.Bergstra, J., Bengio, Y., 2012. Random search for Hyper-Parameter optimization. J. Mach. Learn. Res. 13 (Feb), 281–305.
  15. 15.Bianchi, F. M., Maiorino, E., Kampffmeyer, M. C., Rizzi, A., Jenssen, R., 2017. An overview and comparative analysis of recurrent neural networks for short term load forecasting. CoRR abs/1705.04378. URL http://arxiv.org/abs/1705.04378
  16. 16.Borovykh, A., Bohte, S., Oosterlee, C. W., 2018. Conditional time series forecasting with convolutional neural networks. arXiv preprint arXiv:1703.04691. URL https://arxiv.org/abs/1703.04691
  17. 17.Box, G., Jenkins, G., Reinsel, G., 1994. Time Series Analysis: Forecasting and Control. Forecasting and Control Series. Prentice Hall.
  18. 18.Chen, C., Twycross, J., Garibaldi, J. M., Mar. 2017. A new accuracy measure based on bounded relative error for time series forecasting. PLoS One 12 (3), e0174202.
  19. 19.Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y., 25–29 October 2014. Learning phrase representations using RNN Encoder–Decoder for statistical machine translation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Stroudsburg, PA, USA, pp. 1724–1734.
  20. 20.Cinar, Y. G., Mirisaee, H., Goswami, P., Gaussier, E., A¨ıt-Bachir, A., Strijov, V., 2017. Position-based content attention for time series forecasting with sequence-to-sequence rnns. In: Liu, D., Xie, S., Li, Y., Zhao, D., El-Alfy, E.-S. M. (Eds.), Neural Information Processing. Springer International Publishing, Cham, pp. 533–544.
  21. 21.Claveria, O., Monte, E., Torra, S., Sep. 2017. Data pre-processing for neural network-based forecasting: does it really matter? Technological and Economic Development of Economy 23 (5), 709–725.
  22. 22.Claveria, O., Torra, S., Jan. 2014. Forecasting tourism demand to catalonia: Neural networks vs. time series models. Econ. Model. 36, 220–228.
  23. 23.Cleveland, R. B., Cleveland, W. S., McRae, J. E., Terpenning, I., Jan. 1990. STL: A Seasonal-Trend decomposition procedure based on loess. J. Off. Stat. 6 (1), 3–33.
  24. 24.Collins, J., Sohl-Dickstein, J., Sussillo, D., 2016. Capacity and trainability in recurrent neural networks. In: International Conference on Learning Representations 2016 (ICLR 2016).
  25. 25.Crone, S. F., 2008. NN5 competition. URL http://www.neural-forecasting-competition.com/NN5/
  26. 26.Crone, S. F., Hibon, M., Nikolopoulos, K., Jul. 2011. Advances in forecasting with neural networks? empirical evidence from the NN3 competition on time series prediction. Int. J. Forecast. 27 (3), 635–660.
  27. 27.Dillon, J. V., Langmore, I., Tran, D., Brevdo, E., Vasudevan, S., Moore, D., Patton, B., Alemi, A., Hoffman, M. D., Saurous, R. A., 2017. Tensorflow distributions. CoRR abs/1711.10604. URL http://arxiv.org/abs/1711.10604
  28. 28.Duchi, J., Hazan, E., Singer, Y., 2011. Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res. 12 (Jul), 2121–2159.
  29. 29.Elman, J. L., Apr. 1990. Finding structure in time. Cogn. Sci. 14 (2), 179–211.
  30. 30.eResearch Centre., M., 2019. M3 user guide. URL https://docs.massive.org.au/index.html
  31. 31.Fernando, 2012. Bayesian optimization. URL https://github.com/fmfn/BayesianOptimization
  32. 32.Friedman, J., Hastie, T., Tibshirani, R., 2010. Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software 33 (1), 1–22.
  33. 33.Garc´ıa, S., Fern´andez, A., Luengo, J., Herrera, F., May 2010. Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power. Inf. Sci. 180 (10), 2044–2064.
  34. 34.Gasthaus, J., Benidis, K., Wang, Y., Rangapuram, S. S., Salinas, D., Flunkert, V., Januschowski, T., 16–18 Apr 2019. Probabilistic forecasting with spline quantile function rnns. In: Chaudhuri, K., Sugiyama, M. (Eds.), Proceedings of Machine Learning Research. Vol. 89 of Proceedings of Machine Learning Research. PMLR, pp. 1901–1910.
  35. 35.Google, 2017. Web traffic time series forecasting. URL https://www.kaggle.com/c/web-traffic-time-series-forecasting
  36. 36.He, K., Zhang, X., Ren, S., Sun, J., 27–30 June 2016. Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 770–778.
  37. 37.Hochreiter, S., Schmidhuber, J., Nov. 1997. Long short-term memory. Neural Comput. 9 (8), 1735–1780.
  38. 38.Hornik, K., Stinchcombe, M., White, H., Jan. 1989. Multilayer feedforward networks are universal approximators. Neural Netw. 2 (5), 359–366.
  39. 39.Hutter, F., Hoos, H. H., Leyton-Brown, K., 17–21 Jan 2011. Sequential model-based optimization for general algorithm configuration. In: Coello, C. A. C. (Ed.), Learning and Intelligent Optimization. Springer Berlin Heidelberg, Berlin, Heidelberg, pp. 507–523.
  40. 40.Hyndman, R., 2018. A brief history of time series forecasting competitions. URL https://robjhyndman.com/hyndsight/forecasting-competitions/
  41. 41.Hyndman, R., Kang, Y., Talagala, T., Wang, E., Yang, Y., 2019. tsfeatures: Time Series Feature Extraction. R package version 1.0.0. URL https://pkg.robjhyndman.com/tsfeatures/
  42. 42.Hyndman, R., Khandakar, Y., 2008. Automatic time series forecasting: The forecast package for R. Journal of Statistical Software, Articles 27 (3), 1–22.
  43. 43.Hyndman, R., Koehler, A., Ord, K., D Snyder, R., 01 2008. Forecasting with exponential smoothing. The state space approach. Springer Berlin Heidelberg.
  44. 44.Hyndman, R. J., Koehler, A. B., Oct. 2006. Another look at measures of forecast accuracy. Int. J. Forecast. 22 (4), 679–688.
  45. 45.Ioffe, S., Szegedy, C., 06–11 July 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Proceedings of the 32nd International Conference on Machine Learning - Volume 37. ICML’15. JMLR.org, pp. 448–456.
  46. 46.Jagannatha, A. N., Yu, H., Jun. 2016. Bidirectional RNN for medical event detection in electronic health records. Proceedings of the conference. Association for Computational Linguistics. North American Chapter. Meeting 2016, 473–482.
  47. 47.Januschowski, T., Gasthaus, J., Wang, Y., Salinas, D., Flunkert, V., Bohlke-Schneider, M., Callot, L., 2020. Criteria for classifying forecasting methods. International Journal of Forecasting 36 (1), 167 – 177, m4 Competition.
  48. 48.Ji, Y., Haffari, G., Eisenstein, J., Jun. 2016. A latent variable recurrent neural network for discourse-driven language models. In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, San Diego, California, pp. 332–342.
  49. 49.Jozefowicz, R., Zaremba, W., Sutskever, I., 06–11 July 2015. An empirical exploration of recurrent network architectures. In: Proceedings of the 32nd International Conference on Machine Learning - Volume 37. ICML’15. JMLR.org, pp. 2342–2350.
  50. 50.Kingma, D. P., Ba, J., 7–9 May 2015. Adam: A method for stochastic optimization. In: 3rd International Conference for Learning Representations. Vol. 1412.
  51. 51.Koutn´ık, J., Greff, K., Gomez, F., Schmidhuber, J., 21–26 Jun 2014. A clockwork RNN. In: Proceedings of the 31st International Conference on Machine Learning - Volume 32. ICML’14. JMLR.org, pp. II–1863–II–1871.
  52. 52.Krstanovic, S., Paulheim, H., 12–14 Dec 2017. Ensembles of recurrent neural networks for robust time series forecasting. In: Artificial Intelligence XXXIV. Springer International Publishing, pp. 34–46.
  53. 53.Lai, G., Chang, W.-C., Yang, Y., Liu, H., 8–12 July 2018. Modeling long- and Short-Term temporal patterns with deep neural networks. In: The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. SIGIR ’18. ACM, pp. 95–104.
  54. 54.Laptev, N., Yosinski, J., Li, L. E., Smyl, S., 06–11 Aug 2017. Time-series extreme event forecasting with neural networks at uber. In: International Conference on Machine Learning. Vol. 34. pp. 1–5.
  55. 55.Liang, Y., Ke, S., Zhang, J., Yi, X., Zheng, Y., 13–19 July 2018. Geoman: Multi-level attention networks for geo-sensory time series prediction. In: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18. International Joint Conferences on Artificial Intelligence Organization, pp. 3428–3434.
  56. 56.Lindauer, M., Eggensperger, K., Feurer, M., Falkner, S., Biedenkapp, A., Hutter, F., 2017. Smac v3: Algorithm configuration in python. URL https://github.com/automl/SMAC3
  57. 57.Luong, T., Pham, H., Manning, C. D., Sep 19–21 2015. Effective approaches to attention-based neural machine translation. In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Stroudsburg, PA, USA, pp. 1412–1421.
  58. 58.Makridakis, S., Hibon, M., Oct. 2000. The M3-Competition: results, conclusions and implications. Int. J. Forecast. 16 (4), 451–476.
  59. 59.Makridakis, S., Spiliotis, E., Assimakopoulos, V., Oct. 2018a. The M4 competition: Results, findings, conclusion and way forward. Int. J. Forecast. 34 (4), 802–808.
  60. 60.Makridakis, S., Spiliotis, E., Assimakopoulos, V., mar 2018b. Statistical and machine learning forecasting methods: Concerns and ways forward. PLOS ONE 13 (3), e0194889.
  61. 61.Mandal, P., Senjyu, T., Urasaki, N., Funabashi, T., Jul. 2006. A neural network based several-hour-ahead electric load forecasting using similar days approach. Int. J. Electr. Power Energy Syst. 28 (6), 367–373.
  62. 62.Montero-Manso, P., Athanasopoulos, G., Hyndman, R. J., Talagala, T. S., 2020. Fforma: Feature-based forecast model averaging. International Journal of Forecasting 36 (1), 86 – 92, m4 Competition.
  63. 63.Nelson, M., Hill, T., Remus, W., O’Connor, M., Sep. 1999. Time series forecasting using neural networks: Should the data be deseasonalized first? J. Forecast. 18 (5), 359–367.
  64. 64.Orabona, F., 2017. cocob. URL https://github.com/bremen79/cocob
  65. 65.Orabona, F., Tommasi, T., 04–09 Dec 2017. Training deep networks without learning rates through coin betting. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS’17. Curran Associates Inc., USA, pp. 2157–2167.
  66. 66.Oreshkin, B. N., Carpov, D., Chapados, N., Bengio, Y., 2019. N-BEATS: neural basis expansion analysis for interpretable time series forecasting. CoRR abs/1905.10437. URL http://arxiv.org/abs/1905.10437
  67. 67.Peng, C., Li, Y., Yu, Y., Zhou, Y., Du, S., 31 jan – 03 Feb 2018. Multi-step-ahead host load prediction with GRU based Encoder-Decoder in cloud computing. In: 2018 10th International Conference on Knowledge and Smart Technology (KST). pp. 186–191.
  68. 68.Qin, Y., Song, D., Cheng, H., Cheng, W., Jiang, G., Cottrell, G. W., 2017. A dual-stage attention-based recurrent neural network for time series prediction. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. IJCAI’17. AAAI Press, p. 2627–2633.
  69. 69.R Core Team, 2014. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. URL http://www.R-project.org/
  70. 70.Rahman, M. M., Islam, M. M., Murase, K., Yao, X., Jan. 2016. Layered ensemble architecture for time series forecasting. IEEE Trans Cybern 46 (1), 270–283.
  71. 71.Rangapuram, S. S., Seeger, M., Gasthaus, J., Stella, L., Wang, Y., Januschowski, T., 2018. Deep state space models for time series forecasting. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. NIPS’18. Curran Associates Inc., USA, pp. 7796–7805.
  72. 72.Rob J Hyndman, G. A., 2018. Forecasting: Principles and Practice, 2nd Edition. OTexts. URL https://otexts.com/fpp2/
  73. 73.Salinas, D., Flunkert, V., Gasthaus, J., Januschowski, T., 2019. Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting.
  74. 74.Sch¨afer, A. M., Zimmermann, H. G., 10–14 Sep 2006. Recurrent neural networks are universal approximators. In: Proceedings of the 16th International Conference on Artificial Neural Networks - Volume Part I. ICANN’06. Springer-Verlag, Berlin, Heidelberg, pp. 632–640.
  75. 75.Schuster, M., Paliwal, K. K., Nov. 1997. Bidirectional recurrent neural networks. Trans. Sig. Proc. 45 (11), 2673–2681.
  76. 76.Sen, R., Yu, H.-F., Dhillon, I., 2019. Think globally, act locally: A deep neural network approach to high-dimensional time series forecasting. URL https://arxiv.org/abs/1905.03806
  77. 77.Sharda, R., Patil, R. B., Oct. 1992. Connectionist approach to time series prediction: an empirical test. J. Intell. Manuf. 3 (5), 317–323.
  78. 78.Shih, S.-Y., Sun, F.-K., Lee, H.-y., Sep 2019. Temporal pattern attention for multivariate time series forecasting. Machine Learning 108 (8), 1421–1441.
  79. 79.Smyl, S., 2016. Forecasting short time series with LSTM neural networks. Accessed: 2018-10-30. URL https://gallery.azure.ai/Tutorial/Forecasting-Short-Time-Series-with-LSTM-Neural-Networks-2
  80. 80.Smyl, S., 25–28 jun 2017. Ensemble of specialized neural networks for time series forecasting. In: 37th International Symposium on Forecasting.
  81. 81.Smyl, S., 2020. A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting. International Journal of Forecasting 36 (1), 75 – 85, m4 Competition.
  82. 82.Smyl, S., Kuber, K., 19–22 Jun 2016. Data preprocessing and augmentation for multiple short time series forecasting with recurrent neural networks. In: 36th International Symposium on Forecasting.
  83. 83.Snoek, J., 2012. Spearmint. URL https://github.com/JasperSnoek/spearmint
  84. 84.Snoek, J., Larochelle, H., Adams, R. P., 03–08 Dec 2012. Practical bayesian optimization of machine learning algorithms. In: Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 2. NIPS’12. Curran Associates Inc., USA, pp. 2951–2959.
  85. 85.Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., Srebro, N., Jan. 2018. The implicit bias of gradient descent on separable data. J. Mach. Learn. Res. 19 (1), 2822–2878.
  86. 86.Stˇepniˇcka, M., Burda, M., 09 –12 July 2017. On the results and observations of the time series forecasting competition cif 2016. In: 2017 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE). pp. 1–6.
  87. 87.Suilin, A., 2017. kaggle-web-traffic. Accessed: 2018-11-19. URL https://github.com/Arturus/kaggle-web-traffic/
  88. 88.Sutskever, I., Vinyals, O., Le, Q. V., 08–13 Dec 2014. Sequence to sequence learning with neural networks. In: Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2. NIPS’14. MIT Press, Cambridge, MA, USA, pp. 3104–3112.
  89. 89.Tang, Z., de Almeida, C., Fishwick, P. A., Nov. 1991. Time series forecasting using neural networks vs. box- jenkins methodology. Simulation 57 (5), 303–310.
  90. 90.Trapero, J. R., Kourentzes, N., Fildes, R., Feb 2015. On the identification of sales forecasting models in the presence of promotions. Journal of the Operational Research Society 66 (2), 299–307.
  91. 91.van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A. W., Kavukcuoglu, K., Sep 13-15 2016. Wavenet: A generative model for raw audio. In: The 9th ISCA Speech Synthesis Workshop. ISCA, p. 125.
  92. 92.Wang, Y., Smola, A., Maddix, D., Gasthaus, J., Foster, D., Januschowski, T., 09–15 Jun 2019. Deep factors for forecasting. In: Chaudhuri, K., Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning. Vol. 97 of Proceedings of Machine Learning Research. PMLR, Long Beach, California, USA, pp. 6607–6617.
  93. 93.Wen, R., Torkkola, K., Narayanaswamy, B., Madeka, D., 04 - 09 Dec 2017. A Multi-Horizon quantile recurrent forecaster. In: 31st Conference on Neural Information Processing Systems (NIPS 2017), Time Series Workshop.
  94. 94.Yan, W., Jul. 2012. Toward automatic time-series forecasting using neural networks. IEEE Trans Neural Netw Learn Syst 23 (7), 1028–1039.
  95. 95.Yan, Y., 2016. rBayesianOptimization: Bayesian Optimization of Hyperparameters. R package version 1.1.0. URL https://CRAN.R-project.org/package=rBayesianOptimization
  96. 96.Yao, K., Cohn, T., Vylomova, K., Duh, K., Dyer, C., 22 Jun – 14 Aug 2015. Depth-Gated LSTM. 20th Jelinek Summer Workshop on Speech and Language Technology 2015.
  97. 97.Yeo, I.-K., Johnson, R. A., 2000. A new family of power transformations to improve normality or symmetry. Biometrika 87 (4), 954–959.
  98. 98.Zhang, G., Eddy Patuwo, B., Hu, M. Y., 1998. Forecasting with artificial neural networks: The state of the art. Int. J. Forecast. 14, 35–62.
  99. 99.Zhang, G. P., Jan. 2003. Time series forecasting using a hybrid ARIMA and neural network model. Neurocomputing 50, 159–175.
  100. 100.Zhang, G. P., Berardi, V. L., Jun. 2001. Time series forecasting with neural network ensembles: an application for exchange rate prediction. J. Oper. Res. Soc. 52 (6), 652–664.
  101. 101.Zhang, G. P., Kline, D. M., Nov. 2007. Quarterly Time-Series forecasting with neural networks. IEEE Trans. Neural Netw. 18 (6), 1800–1814.
  102. 102.Zhang, G. P., Qi, M., Jan. 2005. Neural network forecasting for seasonal and trend time series. Eur. J. Oper. Res. 160 (2), 501–514.
  103. 103.Zhu, L., Laptev, N., 18 – 21 nov 2017. Deep and confident prediction for time series at uber. In: 2017 IEEE International Conference on Data Mining Workshops (ICDMW). IEEE, pp. 103–110.

Citation

MLA
Hewamalage, H., et al. “Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions”. International Journal of Forecasting, vol. 37, no. 1, 2021, pp. 388–427, https://doi.org/10.1016/j.ijforecast.2020.06.008.
APA
Hewamalage, H., Bergmeir, C., & Bandara, K. (2021). Recurrent Neural Networks for Time Series Forecasting: Current status and future directions. International Journal of Forecasting, 37(1), 388–427. https://doi.org/10.1016/j.ijforecast.2020.06.008
Chicago
Hewamalage, H., C. Bergmeir, and K. Bandara. 2021. “Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions”. International Journal of Forecasting 37 (1): 388–427. https://doi.org/10.1016/j.ijforecast.2020.06.008.
Harvard
Hewamalage, H., Bergmeir, C. and Bandara, K. (2021) “Recurrent Neural Networks for Time Series Forecasting: Current status and future directions”, International Journal of Forecasting, 37(1), pp. 388–427. Available at: https://doi.org/10.1016/j.ijforecast.2020.06.008.
Vancouver
1. Hewamalage H, Bergmeir C, Bandara K (2021) Recurrent Neural Networks for Time Series Forecasting: Current status and future directions. International Journal of Forecasting 37:388–427

BibTeX

@article{Hewamalage_2021, title={Recurrent Neural Networks for Time Series Forecasting: Current status and future directions}, volume={37}, ISSN={0169-2070}, url={http://dx.doi.org/10.1016/j.ijforecast.2020.06.008}, DOI={10.1016/j.ijforecast.2020.06.008}, number={1}, journal={International Journal of Forecasting}, publisher={Elsevier BV}, author={Hewamalage, Hansika and Bergmeir, Christoph and Bandara, Kasun}, year={2021}, month=Jan, pages={388–427} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF