Reinforcement Learning Based Dynamic Model Combination for Time Series Forecasting
Yuwei FuDi WuBenoit Boulet
Proposes a reinforcement learning framework that dynamically assigns ensemble weights to base models over time, effectively adapting forecasts to non-stationary distributions across real-world time series benchmarks.
Accurate forecasting of time series data is critical for operational planning, energy management, inventory control, and weather prediction. Real-world temporal data, such as renewable energy generation and customer demand, is inherently non-stationary with shifting dynamics over time. While combining predictions from multiple models via ensemble learning is a proven strategy, finding optimal, dynamic combination weights remains difficult. Static weights or fixed heuristics often underperform because individual forecasting models excel under different data conditions.
The article introduces and evaluates the Reinforcement Learning based Model Combination (RLMC) framework. The core objective is to formulate ensemble weight determination as a sequential decision-making problem that dynamically allocates continuous weights to diverse base forecasting models based on observed time series data and historical performance.
The authors designed a reinforcement learning agent using deep deterministic policy gradients, combining deep convolutional feature extraction from raw time series with embeddings of past model performance. To tackle training instability and avoid overfitting to locally dominant models, the approach incorporates pre-training on classification targets, directed exploration around single-model assignments, and a secondary memory buffer for challenging cases. The framework was evaluated across four standard real-world benchmarks: electricity transformer temperatures (ETT), multi-year climate observations, Global Energy Forecasting Competition power loads (GEFCOM), and daily economic panel series (M4-Daily), using standard mean absolute error and percentage error metrics.
The evaluation revealed several key findings. First, RLMC consistently achieved superior forecasting accuracy across all four datasets, outperforming established heuristic, meta-learning, and alternative machine learning baselines. Second, traditional ensemble methods and earlier learning techniques often failed to beat the single best individual model due to severe overfitting on imbalanced training distributions, an issue RLMC successfully resolved. Third, ablation analysis confirmed that specialized exploration strategies were the single most critical factor for performance gains, followed by the secondary buffer and pre-training steps.
These results demonstrate that dynamic model selection powered by reinforcement learning improves operational forecasting without requiring the manual engineering of domain-specific rules. By reliably combining multiple specialized forecasters, organizations can mitigate the operational risk and cost of sudden forecast failures in fluctuating environments. For practitioners and decision-makers, implementing RLMC offers an automated pathway to upgrade forecasting pipelines using existing statistical and machine learning model libraries.
A primary limitation noted is the deterministic nature of the control policy, which may constrain adaptability when base model superiority is extremely skewed. Organizations considering adoption should conduct pilot testing on their specific domain workflows and explore pairing the framework with unsupervised feature representation methods to ensure robustness across diverse operational conditions.
- Paper: FFORMA: Feature-based forecast model averaging, Pablo Montero-Manso et al. (2020). This work introduces feature-based meta-learning to dynamically determine ensemble combination weights for time series forecasting, establishing the foundation for learning dynamic combination policies on temporal data.
- Paper: Mining concept-drifting data streams using ensemble classifiers, Haixun Wang et al. (2003). It provides foundational principles on dynamically weighting ensemble models over streaming data to adapt effectively to non-stationary concept drift.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). This survey provides essential background on deep learning architectures and hybrid ensembling strategies applied across time series forecasting problems.
- Paper: Ensemble deep learning: A review, M. A. Ganaie et al. (2021). It comprehensively reviews decision fusion, dynamic weighting mechanisms, and ensemble designs that motivate learning-based model combinations.
- Paper: Neural Network Ensembles, Cross Validation, and Active Learning, Anders Krogh et al. (1994). It introduces foundational theory on measuring model ambiguity and calculating optimal ensemble combination weights.
- Paper: Deep Reinforcement Learning: An Overview, Yuxi Li (2017). This overview provides core principles of deep reinforcement learning and sequential decision-making frameworks leveraged by RL-based weighting policies.
- Paper: Learning under Concept Drift: A Review, Jie Lu et al. (2019). It delivers key concepts on detecting, understanding, and adapting to non-stationary concept drift in streaming time series data.
- Paper: Stacked regressions, LEO BREIMAN (2004). This seminal paper introduces stacked regression frameworks for combining multiple predictors using constrained optimization.
- Paper: Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting, Yong Liu et al. (2022). This paper advances forecasting under non-stationarity by designing de-stationary attention mechanisms within deep architectures.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). It introduces inverted transformer architectures for multivariate forecasting that can serve as state-of-the-art base models in dynamic ensemble pipelines.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). It explores universal zero-shot time series foundation architectures, extending beyond custom ensemble combination to generalized pre-trained representations.
- Paper: Sundial: A Family of Highly Capable Time Series Foundation Models, Yong Liu 0007 et al. (2025). It develops continuous flow-matching foundation models to generate probabilistic time series forecasts across diverse non-stationary domains.
