SCINet: Time Series Modeling and Forecasting with Sample Convolution and Interaction
Minhao LiuAiling ZengMuxi ChenZhijian XuQiuxia LaiLingna MaQiang Xu
Proposes a hierarchical downsample-convolve-interact neural network architecture that captures multi-resolution temporal features to outperform existing convolutional and Transformer-based models on complex time series forecasting tasks.
Accurate time series forecasting is critical for strategic decision-making across industries such as energy management, traffic planning, healthcare, and financial investment. While recurrent neural networks, temporal convolutional networks, and Transformer-based models are commonly used for sequence modeling, they often fail to exploit the inherent structural properties of time series data. Specifically, standard approaches struggle to balance broad temporal context with localized temporal dynamics, resulting in suboptimal predictive performance and high computational complexity.
The article evaluates a new deep learning framework called the Sample Convolution and Interaction Network (SCINet) designed to improve both short-term and long-term time series forecasting. The main objective is to demonstrate that recursively downsampling time series data, extracting features via distinct convolutional filters, and enabling bidirectional interactive learning between sub-sequences produces a more predictable representation with superior forecasting accuracy.
The researchers designed a hierarchical downsample-convolve-interact architecture structured as a binary tree of basic blocks, which can also be stacked with intermediate supervision for complex dynamics. They evaluated the model using 11 public real-world benchmark datasets covering electricity demand, transformer temperatures, solar power, exchange rates, and freeway traffic systems. The evaluation spanned short-term, long-term, multivariate, univariate, and spatial-temporal forecasting tasks against established recurrent, convolutional, and Transformer baselines.
The evaluation revealed several key findings. First, the proposed model consistently outperformed prior state-of-the-art methods, achieving an average 39.89% reduction in mean squared error across benchmark long-term forecasting tasks and up to a 65% reduction in error on exchange rate data. Second, in short-term forecasting, the model improved accuracy over conventional models by up to 10%, while Transformer models performed poorly due to their lack of focus on recent local temporal patterns. Third, in spatial-temporal traffic forecasting benchmarks, the architecture outperformed dedicated graph neural network models across multiple metrics without relying on explicit spatial relation modeling. Fourth, complexity analysis demonstrated that the model scales with a worst-case computational time complexity of O(T log T), making it significantly more efficient than standard attention-based Transformer models that scale quadratically at O(T^2).
These results indicate that specialized downsampling and interaction architectures can substantially lower operational prediction errors while reducing computational costs compared to complex Transformer architectures. This performance-to-compute advantage directly affects resource-constrained operational environments, such as real-time grid balancing or dynamic traffic rerouting, where high-latency models are impractical. Furthermore, the findings challenge the prevailing assumption that large attention mechanisms are necessary for modeling long-range temporal dependencies.
Organizations seeking to optimize sequence forecasting pipelines should consider piloting downsampling-and-interaction convolution architectures as an alternative or complement to Transformer-based and graph-based models. Operational teams should evaluate these architectures on both short and long horizons, particularly where inference cost and latency are constraining factors. However, because this architecture currently targets deterministic, regularly sampled data, decision-makers should exercise caution when deploying it on datasets characterized by severe missing values or irregular sampling intervals, and they should await probabilistic forecasting extensions before relying on it for uncertainty-sensitive risk management.
- Paper: Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting, SHIYANG LI et al. (2019). Its convolutional self-attention and sparse attention address the locality and quadratic-memory problems that motivate SCINet’s alternative approach to temporal modeling.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). Informer’s efficient attention and progressive sequence distillation provide useful context for SCINet’s effort to reduce long-sequence forecasting costs.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). This survey maps the recurrent, convolutional, and attention-based forecasting approaches that SCINet positions itself against.
- Paper: Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift, Taesung Kim et al. (2022). RevIN explicitly applies reversible normalization to SCINet, showing how the architecture can be adapted to improve forecasting under distribution shift.
