Time-series forecasting with deep learning: a survey

Bryan LimStefan Zohren

article2020Philosophical Transactions of the Royal Society A2,125 citations

Categorizes modern deep learning and hybrid statistical architectures for multi-horizon time series forecasting, providing a clear guide on how different models encode temporal dynamics for operational decision support.

Listen

Accurate time series forecasting is critical across multiple domains, including retail demand planning, financial risk management, energy operations, and healthcare diagnostics. While traditional statistical approaches rely on fixed mathematical assumptions and manual feature engineering, modern organizations increasingly generate vast, complex temporal datasets that require automated, scalable forecasting solutions. Recent advances in computing infrastructure and open-source software have spurred the widespread adoption of deep learning architectures, yet decision-makers frequently face challenges when selecting appropriate model designs, managing risks associated with model over-fitting, and extracting actionable insights from complex systems.

The article surveys modern deep learning architectures applied to time series forecasting, evaluating how distinct network designs process temporal information across single-step and multi-horizon settings. It specifically examines the emergence of hybrid models that combine domain-specific statistical frameworks with neural networks, while assessing advanced techniques that facilitate decision support through model interpretability and counterfactual analysis.

The review synthesizes findings across a broad spectrum of research literature and empirical forecasting competitions. It categorizes foundational neural network building blocks—such as convolutional neural networks, recurrent neural networks, and attention-based Transformer models—and details their statistical equivalents, including autoregressive formulations and Bayesian filtering. The article also reviews standard mechanisms for generating point forecasts and full predictive distributions, comparing direct sequence-to-sequence approaches against recursive iterative forecasting.

The analysis yields four central findings. First, hybrid architectures that integrate classical statistical methods with deep learning consistently outperform pure statistical or standalone machine learning models; notably, a hybrid model combining exponential smoothing with recurrent networks won the major M4 forecasting competition. Second, attention mechanisms and Transformer networks effectively capture long-range dependencies and multi-regime dynamics without suffering from the gradient instability or memory degradation historically observed in recurrent architectures. Third, direct multi-horizon forecasting via encoder-decoder structures mitigates the error accumulation inherent in recursive, step-by-step predictions while accommodating known future variables. Fourth, deep neural networks can support strategic decision-making beyond basic forecasting through built-in attention weights that expose important historical events, as well as specialized causal inference techniques that adjust for time-dependent confounding to model counterfactual scenarios.

These findings indicate that organizations can improve forecast accuracy and operational resilience by moving away from purely data-driven black-box models or rigid standalone statistical models. Adopting hybrid designs significantly reduces the risk of over-fitting in low-data environments, handles non-stationary trends, and eliminates burdensome data pre-processing steps. Furthermore, incorporating probabilistic distributions and counterfactual forecasting enables leaders to quantify operational risks, guard against rare tail events, and run scenario analyses before committing resources.

Organizations developing temporal forecasting pipelines should prioritize hybrid architectures, using domain knowledge to constrain model search spaces and improve generalization. When forecasts span multiple future periods, engineering teams should implement direct sequence-to-sequence models with attention layers rather than recursive one-step predictors. Additionally, deployment strategies should integrate inherent interpretability features and causal adjustment techniques to provide end-users with justifiable, actionable scenario forecasts.

The article highlights key limitations across current deep learning methodologies. Most existing architectures assume regular, discretely sampled time intervals, which creates operational vulnerabilities when dealing with irregularly sampled or missing data. Furthermore, standard models often fail to account for hierarchical structures, such as regional and product-level groupings. While the authors demonstrate high confidence in the evaluated architectures, they note that emerging frameworks—such as continuous-time neural ordinary differential equations and hierarchical deep networks—require further benchmarking and validation before broad operational deployment.

Cover for Time-series forecasting with deep learning: a survey

Abstract

Numerous deep learning architectures have been developed to accommodate the diversity of time series datasets across different domains. In this article, we survey common encoder and decoder designs used in both one-step-ahead and multi-horizon time series forecasting -- describing how temporal information is incorporated into predictions by each model. Next, we highlight recent developments in hybrid deep learning models, which combine well-studied statistical models with neural network components to improve pure methods in either category. Lastly, we outline some ways in which deep learning can also facilitate decision support with time series data.

Table of Contents

  • 1 Introduction
  • 2 Deep Learning Architectures for Time Series Forecasting
  • 2.1 Basic Building Blocks
  • 2.1.1 Convolutional Neural Networks
  • 2.1.2 Recurrent Neural Networks
  • 2.1.3 Attention Mechanisms
  • 2.1.4 Outputs and Loss Functions
  • 2.2 Multi-horizon Forecasting Models
  • 2.2.1 Iterative Methods
  • 2.2.2 Direct Methods
  • 3 Incorporating Domain Knowledge with Hybrid Models
  • 3.1 Non-probabilistic Hybrid Models
  • 3.2 Probabilistic Hybrid Models
  • 4 Facilitating Decision Support Using Deep Neural Networks
  • 4.1 Interpretability With Time Series Data
  • 4.2 Counterfactual Predictions & Causal Inference Over Time
  • 5 Conclusions and Future Directions
  • References

Knowls

  1. Knowl 1 — Mathematical Formulation of Deep Time Series Forecasting

    definition

    In deep time series forecasting, models predict future values of a target yi,t∈Ry_{i,t} \in \mathbb{R} for a given entity ii at time tt. An entity represents a logical grouping of temporal information (e.g., a specific sensor, patient, or retail product). In the univariate one-step-ahead setting, forecasting models are formulated as:

    y^i,t+1=f(yi,t−k:t,xi,t−k:t,si)\hat{y}_{i,t+1} = f(y_{i,t-k:t}, x_{i,t-k:t}, s_i)

    where y^i,t+1\hat{y}_{i,t+1} is the point prediction, yi,t−k:t={yi,t−k,…,yi,t}y_{i,t-k:t} = \{y_{i,t-k}, \dots, y_{i,t}\} is the historical target sequence over a look-back window of length kk, xi,t−k:t={xi,t−k,…,xi,t}x_{i,t-k:t} = \{x_{i,t-k}, \dots, x_{i,t}\} are observed exogenous time-varying inputs over the same look-back window, sis_i is static metadata associated with entity ii (such as location or category), and f(⋅)f(\cdot) is the learned prediction function.

    Deep neural networks decompose this mapping into encoder genc(⋅)g_{\text{enc}}(\cdot) and decoder gdec(⋅)g_{\text{dec}}(\cdot) functions:

    zi,t=genc(yi,t−k:t,xi,t−k:t,si)z_{i,t} = g_{\text{enc}}(y_{i,t-k:t}, x_{i,t-k:t}, s_i)

    y^i,t+1=gdec(zi,t)\hat{y}_{i,t+1} = g_{\text{dec}}(z_{i,t})

    where zi,tz_{i,t} is a latent intermediate feature representation encoding the historical temporal dynamics.

  2. Knowl 2 — Causal and Dilated Convolutions for Temporal Encoding

    model/method

    Convolutional Neural Networks (CNNs) process temporal sequences using causal convolutions, which constrain filters to process historical time steps only. At hidden layer ll and time step tt, a causal convolutional filter computes:

    htl+1=A((W∗h)(l,t))h_t^{l+1} = A\left((W * h)(l, t)\right)

    (W∗h)(l,t)=∑τ=0kW(l,τ)ht−τl(W * h)(l, t) = \sum_{\tau=0}^k W(l, \tau) h_{t-\tau}^l

    where htl∈RHinh_t^l \in \mathbb{R}^{H_{\text{in}}} is the hidden state at layer ll, W(l,τ)∈RHout×HinW(l, \tau) \in \mathbb{R}^{H_{\text{out}} \times H_{\text{in}}} represents fixed convolutional filter weights, and A(⋅)A(\cdot) is an activation function. A single causal layer with a linear activation function corresponds to a classical auto-regressive (AR) model.

    To capture long-range temporal dependencies without an exponential growth in parameter count, dilated causal convolutions introduce a layer-specific dilation rate dld_l:

    (W∗h)(l,t,dl)=∑τ=0⌊k/dl⌋W(l,τ)ht−dlτl(W * h)(l, t, d_l) = \sum_{\tau=0}^{\lfloor k / d_l \rfloor} W(l, \tau) h_{t - d_l \tau}^l

    where ⌊⋅⌋\lfloor \cdot \rfloor is the floor function. Dilation down-samples lower-layer features to aggregate information across distant time blocks; for example, setting dl=2ld_l = 2^l expands the effective receptive field to 2l2^l time steps at layer ll.

  3. Knowl 3 — Recurrent Temporal Encoding and Bayesian Filtering Equivalence

    model/method

    Recurrent Neural Networks (RNNs) encode sequential temporal information into an internal memory state zt∈RHz_t \in \mathbb{R}^H that is updated recursively over time:

    zt=ν(zt−1,yt,xt,s)z_t = \nu(z_{t-1}, y_t, x_t, s)

    where ν(⋅)\nu(\cdot) is the learned state transition function, yty_t is the target, xtx_t are exogenous inputs, and ss is static entity metadata. From a digital signal processing viewpoint, this recurrence acts as a non-linear infinite impulse response (IIR) filter.

    To mitigate vanishing and exploding gradient problems across long horizons, Long Short-Term Memory (LSTM) networks regulate information flow using gate activations and an additive cell state ctc_t:

    it=σ(Wi1zt−1+Wi2yt+Wi3xt+Wi4s+bi)i_t = \sigma(W_{i1} z_{t-1} + W_{i2} y_t + W_{i3} x_t + W_{i4} s + b_i)

    ot=σ(Wo1zt−1+Wo2yt+Wo3xt+Wo4s+bo)o_t = \sigma(W_{o1} z_{t-1} + W_{o2} y_t + W_{o3} x_t + W_{o4} s + b_o)

    ft=σ(Wf1zt−1+Wf2yt+Wf3xt+Wf4s+bf)f_t = \sigma(W_{f1} z_{t-1} + W_{f2} y_t + W_{f3} x_t + W_{f4} s + b_f)

    ct=ft⊙ct−1+it⊙tanh⁡(Wc1zt−1+Wc2yt+Wc3xt+Wc4s+bc)c_t = f_t \odot c_{t-1} + i_t \odot \tanh(W_{c1} z_{t-1} + W_{c2} y_t + W_{c3} x_t + W_{c4} s + b_c)

    zt=ot⊙tanh⁡(ct)z_t = o_t \odot \tanh(c_t)

    where σ(⋅)\sigma(\cdot) is the sigmoid function, ⊙\odot represents the Hadamard (element-wise) product, and W,bW, b are weight matrices and bias vectors.

    From a Bayesian filtering perspective (e.g., Kalman filters), which alternates between state transition updates and error-correction updates on sufficient statistics, an RNN serves as a deterministic, joint non-linear approximation of both filtering steps.

  4. Knowl 4 — Attention Mechanisms for Temporal Feature Aggregation

    model/method

    Temporal attention mechanisms dynamically weight intermediate features across historical time steps to address long-term dependencies. A temporal attention layer computes a context vector hth_t via key-value lookup for a given query:

    ht=∑τ=0kα(κt,qτ)vt−τh_t = \sum_{\tau=0}^k \alpha(\kappa_t, q_\tau) v_{t-\tau}

    where κt\kappa_t is the key representation at time tt, qτq_\tau is the query representation at step τ\tau, vt−τv_{t-\tau} is the value feature at time t−τt-\tau, and α(κt,qτ)∈[0,1]\alpha(\kappa_t, q_\tau) \in [0, 1] is the normalized attention weight such that ∑τ=0kα(t,τ)=1\sum_{\tau=0}^k \alpha(t, \tau) = 1.

    In additive attention models, weights are obtained through a parameterized softmax layer:

    α(t)=softmax⁡(ηt)\alpha(t) = \operatorname{softmax}(\eta_t)

    ηt=Wη1tanh⁡(Wη2κt−1+Wη3qτ+bη)\eta_t = W_{\eta1} \tanh(W_{\eta2} \kappa_{t-1} + W_{\eta3} q_\tau + b_\eta)

    where Wη1,Wη2,Wη3W_{\eta1}, W_{\eta2}, W_{\eta3} and bηb_\eta are trainable parameters. Alternatively, scaled dot-product self-attention is applied directly over historical sequences. Temporal attention enables networks to attend directly to localized past events (such as promotional cycles or holidays) and learn regime-specific temporal dynamics through shifting attention weight distributions.

  5. Knowl 5 — Point and Probabilistic Forecast Output Parameterizations

    model/method

    Deep forecasting architectures support both point estimation and predictive distribution parameterization.

    For point forecasts over a time series of length TT, networks optimize either regression loss (e.g., mean squared error) or binary cross-entropy for discrete event prediction:

    Lregression=1T∑t=1T(yt−y^t)2\mathcal{L}_{\text{regression}} = \frac{1}{T} \sum_{t=1}^T (y_t - \hat{y}_t)^2

    Lclassification=−1T∑t=1T[ytlog⁡(y^t)+(1−yt)log⁡(1−y^t)]\mathcal{L}_{\text{classification}} = -\frac{1}{T} \sum_{t=1}^T \left[ y_t \log(\hat{y}_t) + (1 - y_t) \log(1 - \hat{y}_t) \right]

    For parametric probabilistic forecasting of continuous targets yt+τy_{t+\tau}, the final hidden representation htLh_t^L parameterizes a target distribution, such as a Gaussian N(μ(t,τ),ζ(t,τ)2)\mathcal{N}(\mu(t, \tau), \zeta(t, \tau)^2):

    yt+τ∼N(μ(t,τ),ζ(t,τ)2)y_{t+\tau} \sim \mathcal{N}(\mu(t, \tau), \zeta(t, \tau)^2)

    μ(t,τ)=WμhtL+bμ\mu(t, \tau) = W_\mu h_t^L + b_\mu

    ζ(t,τ)=softplus⁡(WΣhtL+bΣ)\zeta(t, \tau) = \operatorname{softplus}(W_\Sigma h_t^L + b_\Sigma)

    where softplus⁡(x)=log⁡(1+ex)\operatorname{softplus}(x) = \log(1 + e^x) ensures positive standard deviation parameters ζ(t,τ)\zeta(t, \tau).

  6. Knowl 6 — Iterative versus Direct Multi-Horizon Forecasting

    model/method

    Multi-horizon forecasting estimates targets across multiple future steps τ∈{1,…,τmax⁡}\tau \in \{1, \dots, \tau_{\max}\}:

    y^t+τ=f(yt−k:t,xt−k:t,ut−k:t+τ,s,τ)\hat{y}_{t+\tau} = f(y_{t-k:t}, x_{t-k:t}, u_{t-k:t+\tau}, s, \tau)

    where ut−k:t+τu_{t-k:t+\tau} denotes known future inputs (e.g., day of week or holiday flags) across the entire forecasting window, and xt−k:tx_{t-k:t} denotes historical inputs observed only up to tt.

    Deep learning models implement this via two distinct paradigms:

    1. Iterative Methods: Autoregressive models generate multi-step trajectories recursively by feeding predicted target samples from step t+τ−1t+\tau-1 back into the network as inputs for step t+τt+\tau. Point forecasts are computed via Monte Carlo averaging across JJ sampled paths:

    y^t+τ=1J∑j=1Jy~t+τ(j)\hat{y}_{t+\tau} = \frac{1}{J} \sum_{j=1}^J \tilde{y}_{t+\tau}^{(j)}

    Iterative methods share network weights across all horizons but accumulate single-step errors over long horizons and assume exogenous features are fully known or sampled.

    1. Direct Methods: Sequence-to-sequence networks use an encoder for historical observations (yt−k:t,xt−k:t)(y_{t-k:t}, x_{t-k:t}) and a decoder that ingests known future inputs ut−k:t+τu_{t-k:t+\tau} to predict all horizons {1,…,τmax⁡}\{1, \dots, \tau_{\max}\} simultaneously. This eliminates error accumulation across time steps but requires fixed maximum horizons τmax⁡\tau_{\max}.
  7. Knowl 7 — Non-Probabilistic Hybrid Models: Exponential Smoothing RNN

    model/method

    Non-probabilistic hybrid forecasting models integrate analytical parametric time series equations with deep neural networks. In the Exponential Smoothing RNN (ES-RNN) architecture, classical Holt-Winters exponential smoothing formulas capture non-stationary level and seasonal components, while an RNN models residual non-linear variations:

    y^i,t+τ=exp⁡(WEShi,t+τL+bES)×li,t×γi,t+τ\hat{y}_{i,t+\tau} = \exp\left(W_{\text{ES}} h_{i,t+\tau}^L + b_{\text{ES}}\right) \times l_{i,t} \times \gamma_{i,t+\tau}

    li,t=β1(i)yi,tγi,t+(1−β1(i))li,t−1l_{i,t} = \beta_1^{(i)} \frac{y_{i,t}}{\gamma_{i,t}} + (1 - \beta_1^{(i)}) l_{i,t-1}

    γi,t=β2(i)yi,tli,t+(1−β2(i))γi,t−κ\gamma_{i,t} = \beta_2^{(i)} \frac{y_{i,t}}{l_{i,t}} + (1 - \beta_2^{(i)}) \gamma_{i,t-\kappa}

    where hi,t+τLh_{i,t+\tau}^L is the final hidden state of the neural network for step τ\tau, li,tl_{i,t} is the time-varying level, γi,t\gamma_{i,t} is the seasonal component with periodicity κ\kappa, and β1(i),β2(i)∈(0,1)\beta_1^{(i)}, \beta_2^{(i)} \in (0, 1) are entity-specific learnable smoothing coefficients. This hybrid parameterization restricts the model's search space and eliminates the need for manual input scaling.

  8. Knowl 8 — Probabilistic Hybrid Models: Deep State Space Models

    model/method

    Probabilistic hybrid models combine linear state space formulations with deep neural networks that dynamically generate the state space transition and emission parameters at each time step. In Deep State Space Models, the generative process is specified by:

    yt=a(hi,tL)Tlt+ϕ(hi,tL)ϵty_t = a(h_{i,t}^L)^T l_t + \phi(h_{i,t}^L) \epsilon_t

    lt=F(hi,tL)lt−1+q(hi,tL)+Σ(hi,tL)⊙Σtl_t = F(h_{i,t}^L) l_{t-1} + q(h_{i,t}^L) + \Sigma(h_{i,t}^L) \odot \Sigma_t

    where ltl_t is the unobserved latent state, hi,tLh_{i,t}^L is the output of a deep recurrent neural network, a(⋅),F(⋅),q(⋅)a(\cdot), F(\cdot), q(\cdot) are linear transformations, ϕ(⋅),Σ(⋅)\phi(\cdot), \Sigma(\cdot) are linear mappings followed by softmax activations, ϵt∼N(0,1)\epsilon_t \sim \mathcal{N}(0, 1) is a scalar observation error, and Σt∼N(0,I)\Sigma_t \sim \mathcal{N}(0, I) is a multivariate state noise vector.

    Inference and marginal likelihood computation for this linear Gaussian state space model are carried out analytically using standard Kalman filtering recursions.

  9. Knowl 9 — Interpretability in Deep Time Series Models

    model/method

    Interpretability in temporal deep learning architectures is categorized into post-hoc surrogate methods and inherent architectural interpretability:

    1. Post-hoc Interpretability: External surrogate models (such as LIME or SHAP) fit linear approximations or compute game-theoretic Shapley values over perturbed inputs, while gradient-based techniques (saliency maps and influence functions) compute loss derivatives with respect to inputs. These methods identify feature importance but generally discard sequential and temporal dependencies between inputs.

    2. Inherent Interpretability via Attention: Inherent interpretability is achieved by placing constrained temporal attention layers whose weights satisfy ∑τ=0kα(t,τ)=1\sum_{\tau=0}^k \alpha(t, \tau) = 1. The output vector is an exact weighted average over historical representations:

    ht=∑τ=0kα(t,τ)vt−τh_t = \sum_{\tau=0}^k \alpha(t, \tau) v_{t-\tau}

    Analyzing the magnitude of α(t,τ)\alpha(t, \tau) yields instance-level attribution over past time steps, and inspecting the aggregate distribution of attention vectors across time uncovers persistent dataset properties such as seasonality and regime shifts.

  10. Knowl 10 — Counterfactual Time Series Forecasting and Causal Inference

    model/method

    Counterfactual forecasting estimates future trajectories under hypothetical intervention sequences. In longitudinal settings, unbiased causal estimation is complicated by time-dependent confounding, where actions at time tt depend on prior target observations that were themselves influenced by earlier actions.

    Deep neural approaches address time-dependent confounding through three main strategies:

    1. Inverse-Probability-of-Treatment Weighting (IPTW): Recurrent marginal structural networks train one network to estimate propensity weights (probability of treatment assignment given history) and a sequence-to-sequence network trained on the reweighted samples to predict unbiased potential outcomes.

    2. Deep G-Computation: Sequence models jointly model the observational distribution of actions and target outcomes over time to simulate counterfactual rollouts under dynamic treatment regimes.

    3. Adversarially Balanced Representations: Domain adversarial training optimizes a representation encoder to construct historical latent states that are invariant to the treatment assigned at that step.

  11. Knowl 11 — Open Challenges in Deep Time Series Modeling

    limitation

    Current deep time series forecasting architectures face two primary structural limitations:

    1. Irregular and Asynchronous Sampling: Most deep architectures require inputs discretised onto a regular temporal grid. When observations are missing or arrive at irregular continuous-time intervals, discrete CNNs, RNNs, and Transformers struggle without ad-hoc imputation. Continuous-time formulations such as Neural Ordinary Differential Equations (Neural ODEs) provide a path forward, but benchmarked integration with static covariates and complex inputs remains underdeveloped.

    2. Hierarchical and Grouped Structures: Time series across practical domains frequently exhibit hierarchical structures (such as individual item sales aggregating to store, regional, and national levels). Existing univariate and multivariate models rarely enforce coherent hierarchical constraints or share parameters across aggregation levels explicitly.

Coverage note — None was omitted; all survey sections covering encoder/decoder designs, hybrid statistical-neural models, multi-horizon methods, interpretability, counterfactual forecasting, and limitations were captured.

References

  1. 1.Mudelsee M. Trend analysis of climate time series: A review of methods. Earth-Science Reviews. 2019;190:310 – 322.
  2. 2.Stoffer DS, Ombao H. Editorial: Special issue on time series analysis in the biological sciences. Journal of Time Series Analysis. 2012;33(5):701–703.
  3. 3.Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nature Medicine. 2019 Jan;25(1):44–56.
  4. 4.Böse JH, Flunkert V, Gasthaus J, Januschowski T, Lange D, Salinas D, et al. Probabilistic Demand Forecasting at Scale. Proc VLDB Endow. 2017 Aug;10(12):1694–1705.
  5. 5.Andersen TG, Bollerslev T, Christoffersen PF, Diebold FX. Volatility Forecasting. National Bureau of Economic Research; 2005. 11188.
  6. 6.Box GEP, Jenkins GM. Time Series Analysis: Forecasting and Control. Holden-Day; 1976.
  7. 7.Gardner Jr ES. Exponential smoothing: The state of the art. Journal of Forecasting. 1985;4(1):1–28.
  8. 8.Winters PR. Forecasting Sales by Exponentially Weighted Moving Averages. Management Science. 1960;6(3):324–342.
  9. 9.Harvey AC. Forecasting, Structural Time Series Models and the Kalman Filter. Cambridge University Press; 1990.
  10. 10.Ahmed NK, Atiya AF, Gayar NE, El-Shishiny H. An Empirical Comparison of Machine Learning Models for Time Series Forecasting. Econometric Reviews. 2010;29(5-6):594–621.
  11. 11.Krizhevsky A, Sutskever I, Hinton GE. ImageNet Classification with Deep Convolutional Neural Networks. In: Pereira F, Burges CJC, Bottou L, Weinberger KQ, editors. Advances in Neural Information Processing Systems 25 (NIPS); 2012. p. 1097–1105.
  12. 12.Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers); 2019. p. 4171–4186.
  13. 13.Silver D, Huang A, Maddison CJ, Guez A, Sifre L, van den Driessche G, et al. Mastering the game of Go with deep neural networks and tree search. Nature. 2016;529:484–503.
  14. 14.Baxter J. A Model of Inductive Bias Learning. J Artif Int Res. 2000;12(1):149âA ¸S198. ˘
  15. 15.Bengio Y, Courville A, Vincent P. Representation Learning: A Review and New Perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2013;35(8):1798–1828.
  16. 16.Abadi M, Agarwal A, Barham P, Brevdo E, Chen Z, Citro C, et al.. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems; 2015. Software available from tensorflow.org. Available from: http://tensorflow.org/.
  17. 17.Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In: Advances in Neural Information Processing Systems 32; 2019. p. 8024–8035.
  18. 18.Hyndman RJ, Khandakar Y. Automatic time series forecasting: the forecast package for R. Journal of Statistical Software. 2008;26(3):1–22.
  19. 19.Nadaraya EA. On Estimating Regression. Theory of Probability and Its Applications. 1964;9(1):141–142.
  20. 20.Smola AJ, Schà ˝ulkopf B. A Tutorial on Support Vector Regression. Statistics and Computing. 2004;14(3):199–222.
  21. 21.Williams CKI, Rasmussen CE. Gaussian Processes for Regression. In: Advances in Neural Information Processing Systems (NIPS); 1996. .
  22. 22.Damianou A, Lawrence N. Deep Gaussian Processes. In: Proceedings of the Conference on Artificial Intelligence and Statistics (AISTATS); 2013. .
  23. 23.Garnelo M, Rosenbaum D, Maddison C, Ramalho T, Saxton D, Shanahan M, et al. Conditional Neural Processes. In: Proceedings of the International Conference on Machine Learning (ICML); 2018. .
  24. 24.Waibel A. Modular Construction of Time-Delay Neural Networks for Speech Recognition. Neural Comput. 1989;1(1):39âA ¸S46. ˘
  25. 25.Wan E. Time Series Prediction by Using a Connectionist Network with Internal Delay Lines. In: Time Series Prediction. Addison-Wesley; 1994. p. 195–217.
  26. 26.Sen R, Yu HF, Dhillon I. Think Globally, Act Locally: A Deep Neural Network Approach to High-Dimensional Time Series Forecasting. In: Advances in Neural Information Processing Systems (NeurIPS); 2019. .
  27. 27.Wen R, Torkkola K. Deep Generative Quantile-Copula Models for Probabilistic Forecasting. In: ICML Time Series Workshop; 2019. .
  28. 28.Li Y, Yu R, Shahabi C, Liu Y. Diffusion Convolutional Recurrent Neural Network: DataDriven Traffic Forecasting. In: (Proceedings of the International Conference on Learning Representations ICLR); 2018. .
  29. 29.Ghaderi A, Sanandaji BM, Ghaderi F. Deep Forecast: Deep Learning-based Spatio-Temporal Forecasting. In: ICML Time Series Workshop; 2017. .
  30. 30.Salinas D, Bohlke-Schneider M, Callot L, Medico R, Gasthaus J. High-dimensional multivariate forecasting with low-rank Gaussian Copula Processes. In: Advances in Neural Information Processing Systems (NeurIPS); 2019. .
  31. 31.Goodfellow I, Bengio Y, Courville A. Deep Learning. MIT Press; 2016. http://www.deeplearningbook.org.
  32. 32.van den Oord A, Dieleman S, Zen H, Simonyan K, Vinyals O, Graves A, et al. WaveNet: A Generative Model for Raw Audio. arXiv e-prints. 2016 Sep;p. arXiv:1609.03499.
  33. 33.Bai S, Zico Kolter J, Koltun V. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv e-prints. 2018;p. arXiv:1803.01271.
  34. 34.Borovykh A, Bohte S, Oosterlee CW. Conditional Time Series Forecasting with Convolutional Neural Networks. arXiv e-prints. 2017;p. arXiv:1703.04691.
  35. 35.Lyons RG. Understanding Digital Signal Processing (2nd Edition). USA: Prentice Hall PTR; 2004.
  36. 36.Young T, Hazarika D, Poria S, Cambria E. Recent Trends in Deep Learning Based Natural Language Processing [Review Article]. IEEE Computational Intelligence Magazine. 2018;13(3):55–75.
  37. 37.Salinas D, Flunkert V, Gasthaus J. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks. arXiv e-prints. 2017;p. arXiv:1704.04110.
  38. 38.Rangapuram SS, Seeger MW, Gasthaus J, Stella L, Wang Y, Januschowski T. Deep State Space Models for Time Series Forecasting. In: Advances in Neural Information Processing Systems (NIPS); 2018. .
  39. 39.Lim B, Zohren S, Roberts S. Recurrent Neural Filters: Learning Independent Bayesian Filtering Steps for Time Series Prediction. In: International Joint Conference on Neural Networks (IJCNN); 2020. .
  40. 40.Wang Y, Smola A, Maddix D, Gasthaus J, Foster D, Januschowski T. Deep Factors for Forecasting. In: Proceedings of the International Conference on Machine Learning (ICML); 2019. .
  41. 41.Elman JL. Finding structure in time. Cognitive Science. 1990;14(2):179 – 211.
  42. 42.Bengio Y, Simard P, Frasconi P. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks. 1994;5(2):157–166.
  43. 43.Kolen JF, Kremer SC. In: Gradient Flow in Recurrent Nets: The Difficulty of Learning LongTerm Dependencies; 2001. p. 237–243.
  44. 44.Hochreiter S, Schmidhuber J. Long Short-Term Memory. Neural Computation. 1997 Nov;9(8):1735–1780.
  45. 45.Srkk S. Bayesian Filtering and Smoothing. Cambridge University Press; 2013.
  46. 46.Kalman RE. A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering. 1960;82(1):35.
  47. 47.Bahdanau D, Cho K, Bengio Y. Neural Machine Translation by Jointly Learning to Align and Translate. In: Proceedings of the International Conference on Learning Representations (ICLR); 2015. .
  48. 48.Cho K, van Merriënboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, et al. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); 2014. .
  49. 49.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is All you Need. In: Advances in Neural Information Processing Systems (NIPS); 2017. .
  50. 50.Dai Z, Yang Z, Yang Y, Carbonell J, Le Q, Salakhutdinov R. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL); 2019. .
  51. 51.Graves A, Wayne G, Danihelka I. Neural Turing Machines. CoRR. 2014;abs/1410.5401.
  52. 52.Fan C, Zhang Y, Pan Y, Li X, Zhang C, Yuan R, et al. Multi-Horizon Time Series Forecasting with Temporal Attention Learning. In: Proceedings of the ACM SIGKDD international conference on Knowledge discovery and data mining (KDD); 2019. .
  53. 53.Li S, Jin X, Xuan Y, Zhou X, Chen W, Wang YX, et al. Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting. In: Advances in Neural Information Processing Systems (NeurIPS); 2019. .
  54. 54.Lim B, Arik SO, Loeff N, Pfister T. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting. arXiv e-prints. 2019;p. arXiv:1912.09363.
  55. 55.Choi E, Bahadori MT, Sun J, Kulas JA, Schuetz A, Stewart WF. RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism. In: Advances in Neural Information Processing Systems (NIPS); 2016. .
  56. 56.Wen R, et al. A Multi-Horizon Quantile Recurrent Forecaster. In: NIPS 2017 Time Series Workshop; 2017. .
  57. 57.Taieb SB, Sorjamaa A, Bontempi G. Multiple-output modeling for multi-step-ahead time series forecasting. Neurocomputing. 2010;73(10):1950 – 1957.
  58. 58.Marcellino M, Stock J, Watson M. A Comparison of Direct and Iterated Multistep AR Methods for Forecasting Macroeconomic Time Series. Journal of Econometrics. 2006;135:499–526.
  59. 59.Makridakis S, Spiliotis E, Assimakopoulos V. Statistical and Machine Learning forecasting methods: Concerns and ways forward. PLOS ONE. 2018 03;13(3):1–26.
  60. 60.Hyndman R. A brief history of forecasting competitions. International Journal of Forecasting. 2020;36(1):7–14.
  61. 61.The M4 Competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting. 2020;36(1):54 – 74.
  62. 62.Fildes R, Hibon M, Makridakis S, Meade N. Generalising about univariate forecasting methods: further empirical evidence. International Journal of Forecasting. 1998;14(3):339 – 358.
  63. 63.Makridakis S, Hibon M. The M3-Competition: results, conclusions and implications. International Journal of Forecasting. 2000;16(4):451 – 476. The M3- Competition.
  64. 64.Smyl S. A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting. International Journal of Forecasting. 2020;36(1):75 – 85. M4 Competition.
  65. 65.Lim B, Zohren S, Roberts S. Enhancing Time-Series Momentum Strategies Using Deep Neural Networks. The Journal of Financial Data Science. 2019;.
  66. 66.Grover A, Kapoor A, Horvitz E. A Deep Hybrid Model for Weather Forecasting. In: Proceedings of the ACM SIGKDD international conference on knowledge discovery and data mining (KDD); 2015. .
  67. 67.Binkowski M, Marti G, Donnat P. Autoregressive Convolutional Neural Networks for Asynchronous Time Series. In: Proceedings of the International Conference on Machine Learning (ICML); 2018. .
  68. 68.Moraffah R, Karami M, Guo R, Raglin A, Liu H. Causal Interpretability for Machine Learning – Problems, Methods and Evaluation. arXiv e-prints. 2020;p. arXiv:2003.03934.
  69. 69.Chakraborty S, Tomsett R, Raghavendra R, Harborne D, Alzantot M, Cerutti F, et al. Interpretability of deep learning models: A survey of results. In: 2017 IEEE SmartWorld Conference Proceedings); 2017. p. 1–6.
  70. 70.Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence. 2019 May;1(5):206–215.
  71. 71.Ribeio M, Singh S, Guestrin C. "Why Should I Trust You?" Explaining the Predictions of Any Classifier. In: KDD; 2016. .
  72. 72.Lundberg S, Lee SI. A Unified Approach to Interpreting Model Predictions. In: Advances in Neural Information Processing Systems (NIPS); 2017. .
  73. 73.Simonyan K, Vedaldi A, Zisserman A. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. arXiv e-prints. 2013;p. arXiv:1312.6034.
  74. 74.Siddiqui SA, Mercier D, Munir M, Dengel A, Ahmed S. TSViz: Demystification of Deep Learning Models for Time-Series Analysis. IEEE Access. 2019;7:67027–67040.
  75. 75.Koh PW, Liang P. Understanding Black-box Predictions via Influence Functions. In: Proceedings of the International Conference on Machine Learning(ICML; 2017. .
  76. 76.Bai T, Zhang S, Egleston BL, Vucetic S. Interpretable Representation Learning for Healthcare via Capturing Disease Progression through Time. In: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD); 2018. .
  77. 77.Yoon J, Jordon J, van der Schaar M. GANITE: Estimation of Individualized Treatment Effects using Generative Adversarial Nets. In: International Conference on Learning Representations (ICLR); 2018. .
  78. 78.Hartford J, Lewis G, Leyton-Brown K, Taddy M. Deep IV: A Flexible Approach for Counterfactual Prediction. In: Proceedings of the 34th International Conference on Machine Learning (ICML); 2017. .
  79. 79.Alaa AM, Weisz M, van der Schaar M. Deep Counterfactual Networks with Propensity Dropout. In: Proceedings of the 34th International Conference on Machine Learning (ICML); 2017. .
  80. 80.Mansournia MA, Etminan M, Danaei G, Kaufman JS, Collins G. Handling time varying confounding in observational research. BMJ. 2017;359.
  81. 81.Lim B, Alaa A, van der Schaar M. Forecasting Treatment Responses Over Time Using Recurrent Marginal Structural Networks. In: NeurIPS; 2018. .
  82. 82.Li R, Shahn Z, Li J, Lu M, Chakraborty P, Sow D, et al. G-Net: A Deep Learning Approach to G-computation for Counterfactual Outcome Prediction Under Dynamic Treatment Regimes. arXiv e-prints. 2020;p. arXiv:2003.10551.
  83. 83.Bica I, Alaa AM, Jordon J, van der Schaar M. Estimating counterfactual treatment outcomes over time through adversarially balanced representations. In: International Conference on Learning Representations(ICLR); 2020. .
  84. 84.Chen RTQ, Rubanova Y, Bettencourt J, Duvenaud D. Neural Ordinary Differential Equations. In: Proceedings of the International Conference on Neural Information Processing Systems (NIPS); 2018. .
  85. 85.Fry C, Brundage M. The M4 Forecasting Competition – A Practitioner’s View. International Journal of Forecasting. 2019;.

Citation

MLA
Lim, B., and S. Zohren. “Time-series Forecasting with Deep Learning: A Survey”. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, vol. 379, no. 2194, 2021, p. 20200209, https://doi.org/10.1098/rsta.2020.0209.
APA
Lim, B., & Zohren, S. (2021). Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2194), 20200209. https://doi.org/10.1098/rsta.2020.0209
Chicago
Lim, B., and S. Zohren. 2021. “Time-series Forecasting with Deep Learning: A Survey”. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 379 (2194): 20200209. https://doi.org/10.1098/rsta.2020.0209.
Harvard
Lim, B. and Zohren, S. (2021) “Time-series forecasting with deep learning: a survey”, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 379(2194), p. 20200209. Available at: https://doi.org/10.1098/rsta.2020.0209.
Vancouver
1. Lim B, Zohren S (2021) Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 379:20200209

BibTeX

@article{Lim_2021, title={Time-series forecasting with deep learning: a survey}, volume={379}, ISSN={1471-2962}, url={http://dx.doi.org/10.1098/rsta.2020.0209}, DOI={10.1098/rsta.2020.0209}, number={2194}, journal={Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences}, publisher={The Royal Society}, author={Lim, Bryan and Zohren, Stefan}, year={2021}, month=Feb, pages={20200209} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF