Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting

Yuwei FuDi WuBenoit Boulet

article2022AAAI74 citations

Proposes a reinforcement learning framework that dynamically assigns ensemble weights to base models over time, effectively adapting forecasts to non-stationary distributions across real-world time series benchmarks.

Listen

Accurate forecasting of time series data is critical for operational planning, energy management, inventory control, and weather prediction. Real-world temporal data, such as renewable energy generation and customer demand, is inherently non-stationary with shifting dynamics over time. While combining predictions from multiple models via ensemble learning is a proven strategy, finding optimal, dynamic combination weights remains difficult. Static weights or fixed heuristics often underperform because individual forecasting models excel under different data conditions.

The article introduces and evaluates the Reinforcement Learning based Model Combination (RLMC) framework. The core objective is to formulate ensemble weight determination as a sequential decision-making problem that dynamically allocates continuous weights to diverse base forecasting models based on observed time series data and historical performance.

The authors designed a reinforcement learning agent using deep deterministic policy gradients, combining deep convolutional feature extraction from raw time series with embeddings of past model performance. To tackle training instability and avoid overfitting to locally dominant models, the approach incorporates pre-training on classification targets, directed exploration around single-model assignments, and a secondary memory buffer for challenging cases. The framework was evaluated across four standard real-world benchmarks: electricity transformer temperatures (ETT), multi-year climate observations, Global Energy Forecasting Competition power loads (GEFCOM), and daily economic panel series (M4-Daily), using standard mean absolute error and percentage error metrics.

The evaluation revealed several key findings. First, RLMC consistently achieved superior forecasting accuracy across all four datasets, outperforming established heuristic, meta-learning, and alternative machine learning baselines. Second, traditional ensemble methods and earlier learning techniques often failed to beat the single best individual model due to severe overfitting on imbalanced training distributions, an issue RLMC successfully resolved. Third, ablation analysis confirmed that specialized exploration strategies were the single most critical factor for performance gains, followed by the secondary buffer and pre-training steps.

These results demonstrate that dynamic model selection powered by reinforcement learning improves operational forecasting without requiring the manual engineering of domain-specific rules. By reliably combining multiple specialized forecasters, organizations can mitigate the operational risk and cost of sudden forecast failures in fluctuating environments. For practitioners and decision-makers, implementing RLMC offers an automated pathway to upgrade forecasting pipelines using existing statistical and machine learning model libraries.

A primary limitation noted is the deterministic nature of the control policy, which may constrain adaptability when base model superiority is extremely skewed. Organizations considering adoption should conduct pilot testing on their specific domain workflows and explore pairing the framework with unsupervised feature representation methods to ensure robustness across diverse operational conditions.

  • Paper: FFORMA: Feature-based forecast model averaging, Pablo Montero-Manso et al. (2020). This work introduces feature-based meta-learning to dynamically determine ensemble combination weights for time series forecasting, establishing the foundation for learning dynamic combination policies on temporal data.
  • Paper: Mining concept-drifting data streams using ensemble classifiers, Haixun Wang et al. (2003). It provides foundational principles on dynamically weighting ensemble models over streaming data to adapt effectively to non-stationary concept drift.
  • Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). This survey provides essential background on deep learning architectures and hybrid ensembling strategies applied across time series forecasting problems.
  • Paper: Ensemble deep learning: A review, M. A. Ganaie et al. (2021). It comprehensively reviews decision fusion, dynamic weighting mechanisms, and ensemble designs that motivate learning-based model combinations.
  • Paper: Neural Network Ensembles, Cross Validation, and Active Learning, Anders Krogh et al. (1994). It introduces foundational theory on measuring model ambiguity and calculating optimal ensemble combination weights.
  • Paper: Deep Reinforcement Learning: An Overview, Yuxi Li (2017). This overview provides core principles of deep reinforcement learning and sequential decision-making frameworks leveraged by RL-based weighting policies.
  • Paper: Learning under Concept Drift: A Review, Jie Lu et al. (2019). It delivers key concepts on detecting, understanding, and adapting to non-stationary concept drift in streaming time series data.
  • Paper: Stacked regressions, LEO BREIMAN (2004). This seminal paper introduces stacked regression frameworks for combining multiple predictors using constrained optimization.
Cover for Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting

Abstract

Time series data appears in many real-world fields such as energy, transportation, communication systems. Accurate modelling and forecasting of time series data can be of significant importance to improve the efficiency of these systems. Extensive research efforts have been taken for time series problems. Different types of approaches, including both statistical-based methods and machine learning-based methods, have been investigated. Among these methods, ensemble learning has shown to be effective and robust. However, it is still an open question that how we should determine weights for base models in the ensemble. Sub-optimal weights may prevent the final model from reaching its full potential. To deal with this challenge, we propose a reinforcement learning (RL) based model combination (RLMC) framework for determining model weights in an ensemble for time series forecasting tasks. By formulating model selection as a sequential decision-making problem, RLMC learns a deterministic policy to output dynamic model weights for non-stationary time series data. RLMC further leverages deep learning to learn hidden features from raw time series data to adapt fast to the changing data distribution. Extensive experiments on multiple real-world datasets have been implemented to showcase the effectiveness of the proposed method.

Table of Contents

  • Introduction
  • Background
  • Time Series Forecasting
  • Reinforcement Learning
  • Related Work
  • Methods
  • MDP Setting for Model Combination Problem
  • Insights from An RL Perspective
  • Model Combination with DDPG
  • RL Based Model Combination (RLMC)
  • Strategies for Efficient Exploration
  • Algorithm 1: RL-based Model Combination (RLMC)
  • Experiments
  • Datasets
  • Experimental Details
  • Base Models
  • Results and Analysis
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Markov Decision Process Formulation for Dynamic Model Combination

    definition

    In the Reinforcement Learning based Model Combination (RLMC) framework, dynamic weighting of NN base forecasting models across time is formalized as a Markov Decision Process (MDP) defined by the tuple M=⟨S,A,P,r,γ⟩\mathcal{M} = \langle \mathcal{S}, \mathcal{A}, \mathcal{P}, r, \gamma \rangle:

    • State Space S\mathcal{S}: The state at discrete timestep tt is st=(Xt,Lt)s_t = (\mathcal{X}^t, \mathcal{L}^t), where Xt={x1t,…,xTt}∈RT×ds\mathcal{X}^t = \{x_1^t, \dots, x_T^t\} \in \mathbb{R}^{T \times d_s} represents the observed sequence history of length TT and feature dimension dsd_s, and Lt={Lt−11,…,Lt−1N}∈RN\mathcal{L}^t = \{L_{t-1}^1, \dots, L_{t-1}^N\} \in \mathbb{R}^N represents the historical error metrics of the NN base models at timestep t−1t-1.
    • Action Space A\mathcal{A}: An action at=(wt1,…,wtN)∈RNa_t = (w_t^1, \dots, w_t^N) \in \mathbb{R}^N is a continuous weight vector on the unit simplex, satisfying ∑i=1Nwti=1\sum_{i=1}^N w_t^i = 1 and wti∈[0,1]w_t^i \in [0, 1]. The ensemble prediction is the linear combination y^t=∑i=1Nwtiy^ti\hat{y}_t = \sum_{i=1}^N w_t^i \hat{y}_t^i, where y^ti\hat{y}_t^i is the prediction of base model ii.
    • Transition Dynamics P\mathcal{P}: The transition function P(st+1∣st,at)=P(st+1∣st)\mathcal{P}(s_{t+1} \mid s_t, a_t) = \mathcal{P}(s_{t+1} \mid s_t) is decoupled from the action ata_t, as model weights do not alter underlying time series transitions.
    • Reward Function r(s,a)r(s, a): The reward rt=r(st,at)r_t = r(s_t, a_t) measures the forecasting accuracy and rank of the ensemble prediction at timestep tt.
    • Discount Factor γ∈[0,1]\gamma \in [0, 1]: Weight assigned to future forecasting performance (γ=0\gamma = 0 for single-step prediction).

    The objective is to learn a policy πϕ(a∣s)\pi_\phi(a \mid s) maximizing expected return over the state distribution μ\mu:

    J(π)=Es∼μ[Vπ(s)]J(\pi) = \mathbb{E}_{s \sim \mu} [V^\pi(s)]

    where:

    Vπ(s)=E[∑t=0Lγtr(st,at) | s0=s,at∼πϕ(at∣st)]V^\pi(s) = \mathbb{E}\left[ \sum_{t=0}^L \gamma^t r(s_t, a_t) \,\middle|\, s_0 = s, a_t \sim \pi_\phi(a_t \mid s_t) \right]

    for trajectory horizon length LL.

  2. Knowl 2 — Theoretical Properties of Environment Decoupling and Sparse Optimal Actions

    theoretical result

    Analyzing time series dynamic model combination from a reinforcement learning perspective yields two structural insights:

    1. Model-Based Exploration via Decoupled Dynamics: Because state transitions do not depend on the combination action (P(st+1∣st,at)=P(st+1∣st)\mathcal{P}(s_{t+1} \mid s_t, a_t) = \mathcal{P}(s_{t+1} \mid s_t)), treating transitions on historical training data as deterministic places the problem in a model-based setting where transition dynamics and the analytic reward function r(st,at)r(s_t, a_t) are both known. This allows synthetic generation of transitions (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1}) without online interaction.
    2. Explosive Log-Probabilities in Stochastic Policy Gradients: When selecting among NN base models, the optimal policy is often nearly deterministic, concentrating probability mass on the single best model for a given regime. In stochastic policy gradient methods optimizing:

    g^=E^t[∇ϕlog⁡πϕ(at∣st)A^t]\hat{g} = \hat{\mathbb{E}}_t \left[ \nabla_\phi \log \pi_\phi(a_t \mid s_t) \hat{A}_t \right]

    where A^t\hat{A}_t is an advantage estimator, the estimated probabilities πϕ(at∣st)\pi_\phi(a_t \mid s_t) for sub-optimal actions approach zero. As a result, ∣log⁡πϕ(at∣st)∣→∞|\log \pi_\phi(a_t \mid s_t)| \to \infty, causing numeric explosion and gradient instability. This motivates using deterministic continuous actor-critic methods (such as DDPG) instead of stochastic policy gradient algorithms.

  3. Knowl 3 — RLMC Normalized Mixture Reward Function

    equation

    In the RLMC framework, the reward rt=r(st,at)r_t = r(s_t, a_t) combines continuous symmetric Mean Absolute Percentage Error (sMAPE) and base model rank performance:

    rt=αRtsMAPE+Rtrankr_t = \alpha R_t^{sMAPE} + R_t^{rank}

    where α\alpha is a balancing hyperparameter. To ensure scale invariance across diverse datasets, both terms are normalized to [−1,+1][-1, +1]:

    RtsMAPE=1−2⋅τ(δt)9R_t^{sMAPE} = 1 - 2 \cdot \frac{\tau(\delta_t)}{9}

    Rtrank=1−2⋅rtpN−1R_t^{rank} = 1 - 2 \cdot \frac{r_t^p}{N - 1}

    In these equations:

    • δt=200H∑i=1H∣yt+i−y^t+i∣∣yt+i∣+∣y^t+i∣\delta_t = \frac{200}{H} \sum_{i=1}^H \frac{|y_{t+i} - \hat{y}_{t+i}|}{|y_{t+i}| + |\hat{y}_{t+i}|} is the sMAPE error of the ensemble forecast y^t=∑j=1Nwtjy^tj\hat{y}_t = \sum_{j=1}^N w_t^j \hat{y}_t^j relative to ground truth yty_t over forecast horizon HH.
    • τ(δt)∈{0,1,…,9}\tau(\delta_t) \in \{0, 1, \dots, 9\} is the decile bin index of δt\delta_t relative to the empirical base model prediction errors on the training set.
    • rtp∈{0,1,…,N−1}r_t^p \in \{0, 1, \dots, N - 1\} is the zero-based rank of the ensemble prediction error compared against the individual prediction errors of all NN base models at timestep tt (00 being best).
  4. Knowl 4 — Actor-Critic Neural Network Architecture in RLMC

    model/method

    The RLMC actor πϕ(s)\pi_\phi(s) and critic Qθ(s,a)Q_\theta(s, a) extract representations from composite states st=(Xt,Lt)s_t = (\mathcal{X}^t, \mathcal{L}^t):

    • Time Series Encoder: Both networks use Dilated Causal Convolutional Neural Networks (Dilated CNNs) to map raw historical series Xt∈RT×ds\mathcal{X}^t \in \mathbb{R}^{T \times d_s} into a temporal latent embedding without future information leakage.
    • Loss/Performance Encoder: The historical loss vector Lt={Lt−11,…,Lt−1N}∈RN\mathcal{L}^t = \{L_{t-1}^1, \dots, L_{t-1}^N\} \in \mathbb{R}^N from timestep t−1t-1 is passed through a rank embedding table to produce a model-performance embedding.
    • Actor Output: The time series and model loss embeddings are concatenated and passed through fully connected layers to a softmax activation:

    at=πϕ(st)=softmax(zt)∈RNa_t = \pi_\phi(s_t) = \text{softmax}(z_t) \in \mathbb{R}^N

    producing combination weights satisfying wti≥0w_t^i \ge 0 and ∑i=1Nwti=1\sum_{i=1}^N w_t^i = 1.

    • Critic Structure: The critic network Qθ(st,at)Q_\theta(s_t, a_t) applies a Dilated CNN encoder to sts_t, concatenates the state representation with action vector ata_t, and projects to a scalar Q-value through dense linear layers.
  5. Knowl 5 — Sparsity Exploration, Classification Pre-Training, and Dual Buffering

    model/method

    To improve exploration efficiency on the probability simplex and mitigate policy overfitting to globally dominant base models, RLMC employs three strategies:

    1. Actor Pre-training via Multi-Class Classification: Before RL optimization, the actor network πϕ(s)\pi_\phi(s) is pre-trained using cross-entropy loss on training data Dtrain\mathcal{D}_{train}, where the target class for state sts_t is the index of the base model that achieved the minimum forecasting error on that instance.
    2. Sparsity-Biased Vertex Exploration: To explore simplex extremes where individual models dominate, exploratory actions aa are generated by adding random noise ϵ∈RN\epsilon \in \mathbb{R}^N to a randomly chosen standard basis one-hot vector ei∈RNe_i \in \mathbb{R}^N (with (ei)i=1(e_i)_i = 1 and (ei)j≠i=0(e_i)_{j \ne i} = 0):

    a=∣ei+ϵ∣∥ei+ϵ∥2a = \frac{|e_i + \epsilon|}{\|e_i + \epsilon\|_2}

    where ∣⋅∣|\cdot| is element-wise absolute value and ∥⋅∥2\|\cdot\|_2 normalizes the vector. 3. Dual Replay Buffer for Hard Samples: To prevent the agent from collapsing into outputting fixed weights for globally strong base models, transitions (s,a,r,s′)(s, a, r, s') with rewards below a threshold r^t\hat{r}_t (set to −1-1) are stored in an auxiliary replay buffer B′\mathcal{B}' in addition to the standard buffer B\mathcal{B}. Training batches are sampled jointly from both buffers.

  6. Knowl 6 — Deep Deterministic Policy Gradient Optimization for Model Weights

    model/method

    RLMC trains a deterministic actor πϕ(s)\pi_\phi(s) and critic Qθ(s,a)Q_\theta(s, a) via Deep Deterministic Policy Gradient (DDPG).

    The critic parameters θ\theta minimize the Mean-Squared Bellman Error over mini-batches from replay buffers D\mathcal{D}:

    min⁡θE(s,a,r,s′)∼D[(y−Qθ(s,a))2]\min_\theta \mathbb{E}_{(s, a, r, s') \sim \mathcal{D}} \left[ \left( y - Q_\theta(s, a) \right)^2 \right]

    where target yy is computed using target critic Qθ′Q_{\theta'} and target actor πϕ′\pi_{\phi'}:

    y=r+γQθ′(s′,πϕ′(s′))y = r + \gamma Q_{\theta'}\left(s', \pi_{\phi'}(s')\right)

    The actor parameters ϕ\phi are updated by gradient ascent on the expected Q-value:

    max⁡ϕEs∼D[Qθ(s,πϕ(s))]\max_\phi \mathbb{E}_{s \sim \mathcal{D}} \left[ Q_\theta\left(s, \pi_\phi(s)\right) \right]

    Target networks are updated softly using Polyak parameter τ\tau:

    θ′←τθ+(1−τ)θ′,ϕ′←τϕ+(1−τ)ϕ′\theta' \leftarrow \tau \theta + (1 - \tau) \theta', \quad \phi' \leftarrow \tau \phi + (1 - \tau) \phi'

  7. Knowl 7 — RLMC Training Algorithm

    algorithm

    The RLMC training procedure integrates actor classification pre-training, vertex-biased exploration, dual replay buffer sampling, and DDPG updates.

    Input: Pre-trained base models M=⟨M1,…,MN⟩M = \langle M_1, \dots, M_N \rangle, training set Dtrain\mathcal{D}_{train}, total training steps TstepsT_{steps}, exploration parameter ϵ\epsilon, trajectory horizon HH, reward threshold r^t\hat{r}_t, update frequency dd, Polyak update parameter τ\tau.
    Output: Learned actor policy πϕ\pi_\phi and critic QθQ_\theta.
    Initialize critic QθQ_\theta and actor πϕ\pi_\phi with random weights θ,ϕ\theta, \phi.
    Initialize target networks θ′←θ\theta' \leftarrow \theta and ϕ′←ϕ\phi' \leftarrow \phi.
    Initialize replay buffer B\mathcal{B} and extra hard-sample buffer B′\mathcal{B}'.
    Pre-train πϕ(s)\pi_\phi(s) on Dtrain\mathcal{D}_{train} as a multi-class classifier using cross-entropy loss against best-performing base model indices.
    for i=1i = 1 to TstepsT_{steps} do
        Randomly select starting index jj from Dtrain\mathcal{D}_{train}.
        Construct initial state ss from series Xj\mathcal{X}^j and base model performance history Lj\mathcal{L}^j.
        for step k=1k = 1 to HH do
            Select action aa with ϵ\epsilon-greedy sparsity-biased strategy:
                With probability ϵ\epsilon, sample basis vector eme_m and set a=∣em+ϵ∣∥em+ϵ∥2a = \frac{|e_m + \epsilon|}{\|e_m + \epsilon\|_2}; otherwise set a=πϕ(s)a = \pi_\phi(s).
            Compute ensemble forecast y^=∑m=1Nwmy^m\hat{y} = \sum_{m=1}^N w^m \hat{y}^m.
            Compute mixture reward r=αRsMAPE+Rrankr = \alpha R^{sMAPE} + R^{rank}.
            Construct next state s′s' from Xj+k\mathcal{X}^{j+k} and Lj+k\mathcal{L}^{j+k}.
            Store transition (s,a,r,s′)(s, a, r, s') in replay buffer B\mathcal{B}.
            if r<r^tr < \hat{r}_t then
                Store transition (s,a,r,s′)(s, a, r, s') in extra buffer B′\mathcal{B}'.
            end if
            s←s′s \leftarrow s'
        end for
        if i mod d==0i \bmod d == 0 then
            Sample mini-batch of transitions (s,a,r,s′)(s, a, r, s') jointly from B\mathcal{B} and B′\mathcal{B}'.
            Compute target y=r+γQθ′(s′,πϕ′(s′))y = r + \gamma Q_{\theta'}(s', \pi_{\phi'}(s')).
            Update critic QθQ_\theta by minimizing (y−Qθ(s,a))2(y - Q_\theta(s, a))^2.
            Update actor πϕ\pi_\phi via gradient ascent on Qθ(s,πϕ(s))Q_\theta(s, \pi_\phi(s)).
            Update target networks:
                θ′←τθ+(1−τ)θ′\theta' \leftarrow \tau \theta + (1 - \tau) \theta'
                ϕ′←τϕ+(1−τ)ϕ′\phi' \leftarrow \tau \phi + (1 - \tau) \phi'
        end if
    end for
    return πϕ,Qθ\pi_\phi, Q_\theta
  8. Knowl 8 — Forecasting Benchmark Performance of RLMC Across Datasets

    data/table

    Forecasting performance of RLMC was evaluated across four datasets: ETT (D1, hourly transformer oil temperature), Climate (D2, hourly weather temperature), GEFCOM (D3, hourly energy load), and M4-Daily (D4, daily panel series). Performance was measured using Mean Absolute Error (MAE, σ1\sigma_1) and symmetric Mean Absolute Percentage Error (sMAPE in %, σ2\sigma_2) across 5 experimental runs.

    Dataset Uniform Single FFORMS FFORMA DMS M3 RLMC
    σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2 σ1\sigma_1 σ2\sigma_2
    D1 (ETT) 5.22 79.83 4.59 73.69 4.89 75.98 5.42 80.77 5.14 78.12 4.63 74.06 4.39 72.11
    D2 (Climate) 1.70 26.44 1.67 25.80 1.78 26.65 1.69 25.78 1.70 25.89 1.66 25.66 1.47 25.50
    D3 (GEFCOM) 112.65 3.51 89.09 2.75 95.03 2.94 86.93 2.69 90.10 2.79 86.85 2.70 88.25 2.73
    D4 (M4) 250.39 3.93 147.61 2.39 183.90 3.12 161.52 2.72 150.36 2.48 148.60 2.45 145.78 2.31

    RLMC achieved the lowest MAE and sMAPE on ETT, Climate, and M4, and competitive accuracy on GEFCOM. Neural combination methods (DMS, M3, RLMC) consistently outperformed tree-based meta-learning baselines (FFORMS, FFORMA), demonstrating the efficacy of Dilated CNN time-series feature representations. Standard baselines frequently underperformed the Single Best validation model due to training set overfitting, which RLMC mitigated.

  9. Knowl 9 — Ablation Analysis of RLMC Training and Exploration Components

    empirical result

    An ablation study on the Climate dataset analyzed the relative contribution of classification pre-training (w/o pret), sparsity-biased exploration (w/o expl), and the auxiliary buffer for low-reward samples (w/o buffer):

    • Full RLMC: Achieved an MAE of approximately 1.471.47 and an sMAPE of approximately 25.50%25.50\%.
    • Without Sparsity Exploration (w/o expl): Produced the largest degradation, increasing MAE to ≈1.71\approx 1.71 and sMAPE to ≈26.4%\approx 26.4\%, underperforming the Single Best baseline (1.671.67 MAE, 25.80%25.80\% sMAPE). This confirms that vertex-biased exploration is essential to prevent DDPG from settling into sub-optimal interior simplex weights.
    • Without Classification Pre-training (w/o pret): MAE deteriorated to ≈1.59\approx 1.59 and sMAPE to ≈26.0%\approx 26.0\%.
    • Without Hard-Sample Buffer (w/o buffer): MAE deteriorated to ≈1.56\approx 1.56 and sMAPE to ≈25.9%\approx 25.9\%, showing that replaying low-reward transitions prevents the agent from overfitting to training-dominant models.
  10. Knowl 10 — Limitations of Deterministic Policy in Highly Imbalanced Settings

    limitation

    Because RLMC relies on Deep Deterministic Policy Gradient (DDPG), the policy network πϕ(s)\pi_\phi(s) outputs a deterministic point on the simplex action space. In time series datasets with severe class/model imbalance—where a single base model outperforms others on a large majority of training samples—a deterministic actor can suffer from limited exploration flexibility and may prematurely collapse toward selecting that dominant model.

Coverage note — None was omitted; all primary MDP formulations, theoretical insights, reward functions, architectural designs, algorithms, empirical benchmarks, ablations, and stated limitations are included.

References

  1. 1.Armstrong, J. S.; Adya, M.; and Collopy, F. 2001. Rule-based forecasting: Using judgment in time-series extrapolation. In Principles of Forecasting, 259–282. Springer.
  2. 2.Barta, G.; Nagy, G. B. G.; Kazi, S.; and Henk, T. 2017. Gefcom 2014—probabilistic electricity price forecasting. In International Conference on Intelligent Decision Technologies, 67–76. Springer.
  3. 3.Binkowski, M.; Marti, G.; and Donnat, P. 2018. Autoregressive convolutional neural networks for asynchronous time series. In International Conference on Machine Learning, 580–589. PMLR.
  4. 4.Cerqueira, V.; Torgo, L.; Oliveira, M.; and Pfahringer, B. 2017. Dynamic and heterogeneous ensembles for time series forecasting. In 2017 IEEE international conference on data science and advanced analytics (DSAA), 242–251. IEEE.
  5. 5.Chollet, F. 2017. Deep learning with Python. Simon and Schuster.
  6. 6.Collopy, F.; and Armstrong, J. S. 1992. Rule-based forecasting: Development and validation of an expert systems approach to combining time series extrapolations. Management science, 38(10): 1394–1414.
  7. 7.Dietterich, T. G.; et al. 2002. Ensemble learning. The hand-book of brain theory and neural networks, 2(1): 110–125.
  8. 8.Durbin, J.; and Koopman, S. J. 2012. Time series analysis by state space methods. Oxford university press.
  9. 9.Fan, Y.; Tian, F.; Qin, T.; Bian, J.; and Liu, T.-Y. 2017. Learning what data to learn. arXiv preprint arXiv:1702.08635.
  10. 10.Fawaz, H. I.; Forestier, G.; Weber, J.; Idoumghar, L.; and Muller, P.-A. 2019. Deep learning for time series classification: a review. Data mining and knowledge discovery, 33(4): 917–963.
  11. 11.Feng, C.; Sun, M.; and Zhang, J. 2019. Reinforced deterministic and probabilistic load forecasting via Q-learning dynamic model selection. IEEE Transactions on Smart Grid, 11(2): 1377–1386.
  12. 12.Feng, C.; and Zhang, J. 2019. Reinforcement learning based dynamic model selection for short-term load forecasting. In 2019 IEEE Power & Energy Society Innovative Smart Grid Technologies Conference (ISGT), 1–5. IEEE.
  13. 13.Franceschi, J.-Y.; Dieuleveut, A.; and Jaggi, M. 2019. Un-supervised scalable representation learning for multivariate time series. arXiv preprint arXiv:1901.10738.
  14. 14.Fujimoto, S.; Hoof, H.; and Meger, D. 2018. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning, 1587–1596. PMLR.
  15. 15.Gowrisankaran, G.; Reynolds, S. S.; and Samano, M. 2016. Intermittency and the value of renewable energy. Journal of Political Economy, 124(4): 1187–1234.
  16. 16.Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, 1861–1870. PMLR.
  17. 17.Hyndman, R.; Koehler, A. B.; Ord, J. K.; and Snyder, R. D. 2008. Forecasting with exponential smoothing: the state space approach. Springer Science & Business Media.
  18. 18.Hyndman, R. J.; and Athanasopoulos, G. 2018. Forecasting: principles and practice. OTexts.
  19. 19.Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; and Liu, T.-Y. 2017. Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems, 30: 3146–3154.
  20. 20.Lai, G.; Chang, W.-C.; Yang, Y.; and Liu, H. 2018. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, 95–104.
  21. 21.Lemke, C.; and Gabrys, B. 2010. Meta-learning for time series forecasting and forecast combination. Neurocomputing, 73(10-12): 2006–2016.
  22. 22.Liang, Y.; Ke, S.; Zhang, J.; Yi, X.; and Zheng, Y. 2018. Geoman: Multi-level attention networks for geo-sensory time series prediction. In IJCAI, volume 2018, 3428–3434.
  23. 23.Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971.
  24. 24.Lim, B.; and Zohren, S. 2021. Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A, 379(2194): 20200209.
  25. 25.Makridakis, S.; Spiliotis, E.; and Assimakopoulos, V. 2018. The M4 Competition: Results, findings, conclusion and way forward. International Journal of Forecasting, 34(4): 802–808.
  26. 26.Miller, C.; Arjunan, P.; Kathirgamanathan, A.; Fu, C.; Roth, J.; Park, J. Y.; Balbach, C.; Gowri, K.; Nagy, Z.; Fontanini, A. D.; et al. 2020. The ASHRAE Great Energy Predictor III competition: Overview and results. Science and Technology for the Built Environment, 26(10): 1427–1447.
  27. 27.Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602.
  28. 28.Montero-Manso, P.; Athanasopoulos, G.; Hyndman, R. J.; and Talagala, T. S. 2020. FFORMA: Feature-based forecast model averaging. International Journal of Forecasting, 36(1): 86–92.
  29. 29.Nichol, A.; Achiam, J.; and Schulman, J. 2018. On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999.
  30. 30.Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499.
  31. 31.Oreshkin, B. N.; Carpov, D.; Chapados, N.; and Bengio, Y. 2019. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. arXiv preprint arXiv:1905.10437.
  32. 32.Pai, P.-F.; and Lin, C.-S. 2005. A hybrid ARIMA and support vector machines model in stock price forecasting. Omega, 33(6): 497–505.
  33. 33.Prudêncio, R. B.; and Ludermir, T. B. 2004. Meta-learning approaches to selecting time series models. Neurocomputing, 61: 121–137.
  34. 34.Rangapuram, S. S.; Seeger, M. W.; Gasthaus, J.; Stella, L.; Wang, Y.; and Januschowski, T. 2018. Deep state space models for time series forecasting. Advances in neural information processing systems, 31: 7785–7794.
  35. 35.Sánchez, I. 2008. Adaptive combination of forecasts with application to wind energy. International Journal of Forecasting, 24(4): 679–693.
  36. 36.Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; and Abbeel, P. 2015. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438.
  37. 37.Seeger, M.; Salinas, D.; and Flunkert, V. 2016. Bayesian intermittent demand forecasting for large inventories. In Proceedings of the 30th International Conference on Neural Information Processing Systems, 4653–4661.
  38. 38.Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; and Riedmiller, M. 2014. Deterministic policy gradient algorithms. In International conference on machine learning, 387–395. PMLR.
  39. 39.Sutton, R. S.; and Barto, A. G. 2018. Reinforcement learning: An introduction. MIT press.
  40. 40.Talagala, T. S.; Hyndman, R. J.; Athanasopoulos, G.; et al. 2018. Meta-learning how to forecast time series. Monash Econometrics and Business Statistics Working Papers, 6: 18.
  41. 41.Tanaka, K. 2017. Time series analysis: Nonstationary and noninvertible distribution theory, volume 4. John Wiley & Sons.
  42. 42.Tang, J.; Belletti, F.; Jain, S.; Chen, M.; Beutel, A.; Xu, C.; and H. Chi, E. 2019. Towards neural mixture recommender for long range dependent user sequences. In The World Wide Web Conference, 1782–1793.
  43. 43.Taylor, J. W.; McSharry, P. E.; and Buizza, R. 2009. Wind power density forecasting using ensemble predictions and time series models. IEEE Transactions on Energy Conversion, 24(3): 775–782.
  44. 44.Vilalta, R.; and Drissi, Y. 2002. A perspective view and survey of meta-learning. Artificial intelligence review, 18(2): 77–95.
  45. 45.Weigel, A. P.; Liniger, M.; and Appenzeller, C. 2008. Can multi-model combination really enhance the prediction skill of probabilistic ensemble forecasts? Quarterly Journal of the Royal Meteorological Society: A journal of the atmospheric sciences, applied meteorology and physical oceanography, 134(630): 241–260.
  46. 46.Wolpert, D. H. 1996. The lack of a priori distinctions between learning algorithms. Neural computation, 8(7): 1341–1390.
  47. 47.Yuan, F.; Shou, L.; Pei, J.; Lin, W.; Gong, M.; Fu, Y.; and Jiang, D. 2021. Reinforced multi-teacher selection for knowledge distillation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 14284–14291.
  48. 48.Zhang, G. P. 2003. Time series forecasting using a hybrid ARIMA and neural network model. Neurocomputing, 50: 159–175.
  49. 49.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of AAAI.
  50. 50.Zhou, Z.-H. 2021. Ensemble learning. In Machine Learning, 181–210. Springer.

Citation

MLA
Fu, Y., et al. “Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 6, 2022, pp. 6639–47, https://doi.org/10.1609/AAAI.V36I6.20618.
APA
Fu, Y., Wu, D., & Boulet, B. (2022). Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, 36(6), 6639–6647. https://doi.org/10.1609/AAAI.V36I6.20618
Chicago
Fu, Y., D. Wu, and B. Boulet. 2022. “Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (6): 6639–47. https://doi.org/10.1609/AAAI.V36I6.20618.
Harvard
Fu, Y., Wu, D. and Boulet, B. (2022) “Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(6), pp. 6639–6647. Available at: https://doi.org/10.1609/AAAI.V36I6.20618.
Vancouver
1. Fu Y, Wu D, Boulet B (2022) Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 36:6639–6647

BibTeX

@article{Fu_2022, title={Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I6.20618}, DOI={10.1609/aaai.v36i6.20618}, number={6}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Fu, Yuwei and Wu, Di and Boulet, Benoit}, year={2022}, month=June, pages={6639–6647} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF