SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention
Romain IlbertAmbroise OdonnatVasilii FeofanovAladin VirmauxGiuseppe PaoloThemis PalpanasIevgen Redko
Demonstrates why standard attention mechanisms cause transformers to fail at multivariate time series forecasting and introduces SAMformer, a lightweight channel-wise architecture trained with sharpness-aware minimization that matches large foundation models using substantially fewer parameters.
Forecasting multiple interdependent variables over long horizons is essential for operational planning across critical sectors such as energy grid management, supply chain logistics, traffic coordination, and financial analysis. Although transformer neural network architectures dominate natural language processing and computer vision, they consistently fail to match the performance of much simpler linear models in long-term time series forecasting. The article aims to evaluate why transformers fail in these forecasting scenarios and demonstrates how a lightweight architecture paired with an advanced optimization technique can restore the competitive edge of transformers over existing baselines.
The researchers conducted their investigation by first evaluating a controlled mathematical model to isolate the root cause of transformer training failures. Based on the theoretical insights gained, they designed a specialized, lightweight architecture named SAMformer. The authors evaluated this new model against state-of-the-art architectures, including linear models, complex specialized transformers, and a large-scale forecasting foundation model across eight diverse, publicly available real-world benchmarks spanning electricity grids, weather tracking, traffic monitoring, and financial exchange rates.
The article established several critical findings. First, standard transformers converge to highly unstable, sharp error regions during optimization primarily because of the internal attention mechanism. Second, the proposed SAMformer model systematically avoids these poor outcomes by applying sharpness-aware optimization—a technique that seeks flatter error regions to ensure better real-world performance—and channel-wise attention, which models feature relationships across time rather than step-by-step temporal correlations. Third, SAMformer outperformed the leading non-transformer baseline, TSMixer, by 14.33% and the best specialized multivariate transformer baseline, FEDformer, by 12.36% across standard benchmarks, achieving top-tier accuracy in seven out of eight datasets. Fourth, SAMformer demonstrated performance on par with or superior to the MOIRAI foundation model (up to an overall 7.6% error reduction) while using roughly four times fewer parameters than leading linear baselines and orders of magnitude fewer than foundation architectures.
These results demonstrate that organizations do not need increasingly complex or massive neural networks to achieve state-of-the-art predictive accuracy. Adopting a shallow, parameter-efficient transformer with flatter loss optimization significantly lowers computational and memory costs, reduces operational carbon footprints, and ensures stable predictions regardless of initial random setup. Practitioners and decision-makers evaluating predictive AI systems should consider lightweight, sharpness-aware transformer architectures over oversized foundation models or over-engineered baselines when dealing with multivariate numerical data.
While confidence in these findings is high across standard academic benchmarks, the authors note that the architecture showed weaker relative gains on the financial exchange dataset, where a competing baseline maintained superior accuracy. Further analysis and real-world pilot deployments in domain-specific environments are recommended to validate optimal configuration settings before replacing mission-critical forecasting pipelines.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Read iTransformer first to understand the channel-wise attention design that SAMformer adopts and evaluates against.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). PatchTST establishes a key channel-independent Transformer approach and forecasting baseline that helps frame SAMformer’s channel-wise design and comparisons.
- Paper: FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series Forecasting, Tian Zhou et al. (2022). FEDformer provides an important long-horizon Transformer baseline whose performance SAMformer directly contextualizes.
No sufficiently relevant recommendations were found.
