Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting
Yong LiuHaixu WuJianmin WangMingsheng Long
Develops the Non-stationary Transformer framework to solve over-stationarization in time series forecasting by recovering intrinsic temporal dependencies within attention mechanisms, reducing forecasting error by nearly half across standard architectures.
Real-world time series forecasting is critical for operational planning, energy management, weather tracking, and financial risk assessment. In practical deployments, however, incoming data streams regularly exhibit non-stationarity—meaning their statistical properties, such as mean and standard deviation, continually shift over time. Standard machine learning solutions routinely pre-process data through stationarization techniques to smooth out these fluctuations and improve predictability. However, removing intrinsic variations causes a serious side effect called over-stationarization: attention-based models lose the ability to distinguish eventful, bursty temporal patterns from routine data, causing deep models to produce overly flat and generic forecasts.
The article introduces and evaluates Non-stationary Transformers, a generic framework designed to eliminate this trade-off. The objective is to boost time series predictability through input-output normalization while simultaneously re-integrating essential non-stationary dynamics directly into the model's internal attention mechanism.
The researchers assessed this approach by wrapping standard Transformer architectures and their efficient variants with two complementary components: Series Stationarization and De-stationary Attention. Series Stationarization normalizes incoming data chunks and restores original statistics at the output without adding learnable parameters. De-stationary Attention applies a lightweight multilayer perceptron projector to learn non-stationary scaling and shifting factors from raw statistics, using them to rescale attention calculations. The approach was tested across six benchmark datasets spanning electricity consumption, traffic patterns, influenza illness rates, currency exchange rates, and meteorological records across varied forecast lengths.
The evaluation produced four primary findings. First, integrating the framework consistently elevated performance across all tested benchmarks, achieving state-of-the-art accuracy in both multivariate and univariate forecasting. Second, the framework delivered substantial error reductions on standard Transformer variants, decreasing mean squared error by 49.43% on vanilla Transformer, 47.34% on Informer, 46.89% on Reformer, and 10.57% on Autoformer. Third, accuracy gains were most pronounced on datasets exhibiting the greatest statistical volatility; for instance, the model reduced mean squared error by 17% on exchange rates and 25% on influenza rate forecasts over long prediction horizons compared to prior state-of-the-art methods. Fourth, statistical validation confirmed that predictions generated with De-stationary Attention closely matched ground-truth stationarity levels within a 97% to 103% margin, preventing the over-smoothing observed in traditional stationarization methods.
These findings demonstrate that organizations do not need to choose between data stability and responsiveness to real-world fluctuations. Prior stationarization approaches degraded the expressive power of Transformer attention mechanisms by treating all normalized series identically. By reintroducing raw statistical factors inside the attention layer, deep forecasting models capture sudden shifts and extreme events effectively without compromising baseline stability, offering improved forecast reliability for high-stakes operational planning.
Organizations deploying Transformer architectures for time series forecasting should adopt this framework as an efficient, plug-and-play upgrade. Because the modifications introduce minimal computational overhead and preserve native model complexity, teams can integrate the modules into existing pipelines without significant hardware expansion. Decision-makers should validate performance through pilot tests on their organization's most volatile operational data. Future work should explore extending these non-stationary recovery mechanisms to broader, model-agnostic deep learning architectures beyond Transformer-based models.
Confidence in these findings is supported by consistent empirical improvements across diverse domains, baseline architectures, and forecast horizons. However, leaders should note that the theoretical formulation assumes approximate linear properties across embedding and feed-forward layers, which deep networks only partially satisfy in practice. While the lightweight learned projector effectively compensates for these approximations in benchmark evaluations, performance should still be verified on domain-specific data with extreme outliers or irregular sampling frequencies.
- Paper: Reversible Instance Normalization for Accurate Time-Series Forecasting against Distribution Shift, Taesung Kim et al. (2022). Introduces reversible instance normalization (RevIN) to handle non-stationarity and distribution shifts in time-series forecasting, which serves as the direct foundation for the series stationarization and over-stationarization problems tackled in Non-stationary Transformers.
- Paper: Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting, Haoyi Zhou et al. (2021). Pioneered efficient Transformer architectures for long-sequence time-series forecasting, serving as one of the primary baseline models boosted by the Non-stationary Transformers framework.
- Paper: Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting, Haixu Wu et al. (2021). Establishes deep temporal decomposition and autocorrelation mechanisms for long-term time series forecasting, defining the modern Transformer baseline ecosystem improved upon by this work.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). Presents the Reformer architecture, which is one of the core Transformer baselines evaluated and improved by the Non-stationary Transformers framework.
- Paper: Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting, SHIYANG LI et al. (2019). Explores the fundamental adaptations needed to apply self-attention mechanisms to time series forecasting, contextualizing early Transformer limitations on temporal data.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). Critically examines whether complex Transformer modifications are necessary for long-term time series forecasting compared to simpler normalization and linear models.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). Advances Transformer-based forecasting by introducing subseries patching and channel independence, significantly outperforming earlier stationarized point-wise Transformer architectures.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Inverts the standard tokenization dimension in Transformers to model variate-level dynamics rather than temporal attention, overcoming the temporal attention degradation analyzed in Non-stationary Transformers.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). Generalizes temporal modeling beyond 1D Transformers by converting 1D sequences into 2D intra- and inter-period variations across diverse time series tasks.
- Paper: Transformers in Time Series: A Survey, Qingsong Wen et al. (2022). Provides a comprehensive survey synthesizing recent structural modifications and stationarization techniques in Transformers across the broader time-series literature.
