xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition
Artyom StitsyukJaesik Choi
Proposes xPatch, an effective non-transformer architecture combining exponential moving average seasonal-trend decomposition with dual MLP and CNN streams to surpass attention-based models in long-term time series forecasting.
Accurate long-term forecasting of complex time-dependent data is critical for operational planning, resource allocation, and risk management across sectors like energy, transportation, and weather forecasting. While recent advanced machine learning models based on transformer architectures have dominated this space, they struggle to effectively preserve chronological order and temporal relationships due to their underlying design. Moreover, standard data-smoothing techniques frequently distort early and late historical signals, degrading forecast quality.
The article demonstrates that a non-transformer architecture can outperform leading transformer models in long-term forecasting while requiring substantially fewer computational resources. The authors evaluate this premise by introducing xPatch, an architecture combining a new exponential trend-smoothing method, a dual-stream processing network, and tailored training optimization schemes.
The researchers evaluated the proposed approach against nine state-of-the-art models across nine widely used real-world benchmark datasets representing energy grids, traffic flows, weather conditions, currency exchange, and disease tracking. The evaluation compared performance across both fixed baseline settings and comprehensive hyperparameter optimization, measuring average forecast error and computational runtimes.
The evaluation yielded several critical findings. First, xPatch achieved superior accuracy, leading top benchmark models on 70% of datasets under mean squared error and 90% under mean absolute error during hyperparameter searches. Second, xPatch outperformed prominent transformer architectures, reducing average error metrics by approximately 4% to 8% compared to leading models such as CARD and PatchTST. Third, xPatch demonstrated high computational efficiency: its training time per step (about 3.1 milliseconds) and inference time (about 1.3 milliseconds) were roughly two to five times faster than competing transformer frameworks.
These findings suggest that complex, resource-heavy transformer architectures are not necessary to achieve state-of-the-art forecasting performance. Organizations can deploy lighter, non-transformer models that combine linear and non-linear paths to capture both steady trends and repeating seasonal fluctuations, thereby lowering cloud compute costs, accelerating prediction delivery, and maintaining higher forecast reliability.
Organizations managing time-critical forecasting operations should consider adopting dual-stream, non-transformer architectures like xPatch to improve accuracy while reducing computational infrastructure overhead. Before broad operational rollout, teams should conduct pilot testing on internal domain-specific datasets to evaluate the impact of historical lookback window sizing and determine whether to utilize channel-independent configurations.
Confidence in these findings is high for standard forecasting benchmarks spanning up to 720 future time steps. However, readers should note that performance depends on setting appropriate smoothing factors and lookback lengths, and real-world results may vary in operating environments with extreme data anomalies, sudden structural shifts, or missing observations.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). It introduces the patching and channel-independence mechanisms that xPatch directly adapts and investigates within a non-transformer architecture.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). It establishes that simple linear decomposition baselines can outperform complex Transformers, motivating the linear stream and decomposition focus of xPatch.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). It popularizes progressive seasonal-trend decomposition as a core architectural module for long-term time series forecasting.
- Paper: A state space framework for automatic forecasting using exponential smoothing methods, Rob J. Hyndman et al. (2002). It provides the foundational statistical framework for seasonal-trend exponential smoothing methods that inspire xPatch's exponential decomposition module.
- Paper: A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting, Slawek Smyl (2020). It establishes the hybrid integration of classical exponential smoothing decomposition with neural network components for time series forecasting.
- Paper: Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks, Guokun Lai et al. (2017). It introduces a dual-branch neural architecture combining linear autoregressive paths with convolutional components for capturing mixed temporal dynamics.
No sufficiently relevant recommendations were found.
