SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention

Romain IlbertAmbroise OdonnatVasilii FeofanovAladin VirmauxGiuseppe PaoloThemis PalpanasIevgen Redko

article2024ICML77 citations

Demonstrates why standard attention mechanisms cause transformers to fail at multivariate time series forecasting and introduces SAMformer, a lightweight channel-wise architecture trained with sharpness-aware minimization that matches large foundation models using substantially fewer parameters.

Listen

Forecasting multiple interdependent variables over long horizons is essential for operational planning across critical sectors such as energy grid management, supply chain logistics, traffic coordination, and financial analysis. Although transformer neural network architectures dominate natural language processing and computer vision, they consistently fail to match the performance of much simpler linear models in long-term time series forecasting. The article aims to evaluate why transformers fail in these forecasting scenarios and demonstrates how a lightweight architecture paired with an advanced optimization technique can restore the competitive edge of transformers over existing baselines.

The researchers conducted their investigation by first evaluating a controlled mathematical model to isolate the root cause of transformer training failures. Based on the theoretical insights gained, they designed a specialized, lightweight architecture named SAMformer. The authors evaluated this new model against state-of-the-art architectures, including linear models, complex specialized transformers, and a large-scale forecasting foundation model across eight diverse, publicly available real-world benchmarks spanning electricity grids, weather tracking, traffic monitoring, and financial exchange rates.

The article established several critical findings. First, standard transformers converge to highly unstable, sharp error regions during optimization primarily because of the internal attention mechanism. Second, the proposed SAMformer model systematically avoids these poor outcomes by applying sharpness-aware optimization—a technique that seeks flatter error regions to ensure better real-world performance—and channel-wise attention, which models feature relationships across time rather than step-by-step temporal correlations. Third, SAMformer outperformed the leading non-transformer baseline, TSMixer, by 14.33% and the best specialized multivariate transformer baseline, FEDformer, by 12.36% across standard benchmarks, achieving top-tier accuracy in seven out of eight datasets. Fourth, SAMformer demonstrated performance on par with or superior to the MOIRAI foundation model (up to an overall 7.6% error reduction) while using roughly four times fewer parameters than leading linear baselines and orders of magnitude fewer than foundation architectures.

These results demonstrate that organizations do not need increasingly complex or massive neural networks to achieve state-of-the-art predictive accuracy. Adopting a shallow, parameter-efficient transformer with flatter loss optimization significantly lowers computational and memory costs, reduces operational carbon footprints, and ensures stable predictions regardless of initial random setup. Practitioners and decision-makers evaluating predictive AI systems should consider lightweight, sharpness-aware transformer architectures over oversized foundation models or over-engineered baselines when dealing with multivariate numerical data.

While confidence in these findings is high across standard academic benchmarks, the authors note that the architecture showed weaker relative gains on the financial exchange dataset, where a competing baseline maintained superior accuracy. Further analysis and real-world pilot deployments in domain-specific environments are recommended to validate optimal configuration settings before replacing mission-critical forecasting pipelines.

Ilbert et al (2024).pdf

No sufficiently relevant recommendations were found.

Cover for SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention

Abstract

Transformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting. To better understand this phenomenon, we start by studying a toy linear forecasting problem for which we show that transformers are incapable of converging to their true solution despite their high expressive power. We further identify the attention of transformers as being responsible for this low generalization capacity. Building upon this insight, we propose a shallow lightweight transformer model that successfully escapes bad local minima when optimized with sharpness-aware optimization. We empirically demonstrate that this result extends to all commonly used real-world multivariate time series datasets. In particular, SAMformer surpasses current state-of-the-art methods and is on par with the biggest foundation model MOIRAI while having significantly fewer parameters. The code is available at https://github.com/romilbert/samformer.

Table of Contents

  • 1. Introduction
  • 2. Proposed Approach
  • 2.1. Problem Setup
  • 2.2. Motivational Example
  • 2.3. Transformer's Loss Landscape
  • 2.4. SAMformer: Putting It All Together
  • 3. Experiments
  • 3.1. Main Takeaways
  • 3.2. Qualitative Benefits of Our Approach
  • 3.3. SAMformer vs MOIRAI
  • 3.4. Ablation Study and Sensitivity Analysis
  • 4. Discussion and Future Work
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix
  • A. Experimental Setup
  • A.1. Architecture and Training Parameters
  • A.2. Datasets
  • A.3. More Details on the Baselines
  • B. Additional Experiments
  • B.1. MAE Results
  • B.2. Significance Test for SAMformer and TSMixer with SAM
  • B.3. Computational Efficiency of SAMformer
  • B.4. Strong Generalization Regardless of the Initialization
  • B.5. Faithful Signal Propagation
  • C. Ablation Study and Sensitivity Analysis
  • C.1. Sensitivity to the Prediction Horizon H.
  • C.2. Sensitivity to the Neighborhood Size ρ.
  • C.3. Sensitivity to the Change of the Optimizer.
  • C.4. Ablation on the Implementation.
  • D. Additional Background
  • D.1. Reversible Instance Normalization: RevIN
  • D.2. Sharpness-aware minimization (SAM)
  • E. Proofs
  • E.1. Notations
  • E.3. Proof of Proposition 2.2
  • E.4. Proof of Proposition D.1
  • E.5. Matrix formulation of ˆ Y in Eq. (11)

Knowls

  1. Knowl 1 — SAMformer combines channel-wise attention, RevIN, and sharpness-aware training

    model/method

    For an input window X∈RD×LX\in\mathbb{R}^{D\times L} with DD channels and look-back length LL, SAMformer first applies reversible instance normalization (RevIN): each channel is normalized using its mean and standard deviation over the input window, with learned affine parameters, and the predictions are then denormalized using the same input statistics. Its forecasting network is a one-layer, one-head transformer without a feed-forward block. For the normalized input XX, define

    A(X)=softmax⁡row ⁣(XWQWK⊤X⊤dm),f(X)=[X+A(X)XWVWO]W.A(X)=\operatorname{softmax}_{\mathrm{row}}\!\left(\frac{XW_QW_K^{\top}X^{\top}}{\sqrt{d_m}}\right),\qquad f(X)=\bigl[X+A(X)XW_VW_O\bigr]W.

    Here A(X)∈RD×DA(X)\in\mathbb{R}^{D\times D} is row-wise softmax attention across channels; WQ,WK,WV∈RL×dmW_Q,W_K,W_V\in\mathbb{R}^{L\times d_m}, WO∈Rdm×LW_O\in\mathbb{R}^{d_m\times L}, and W∈RL×HW\in\mathbb{R}^{L\times H} are learned weights; and HH is the forecast horizon. Thus the attention output is added to the input by a residual connection, then projected from length LL to horizon HH. SAMformer trains these parameters using sharpness-aware minimization (SAM), which minimizes the worst training loss in a parameter neighborhood: LSAM(ω)=max⁡∥δ∥2≤ρLtrain(ω+δ)\mathcal{L}_{\mathrm{SAM}}(\omega)=\max_{\|\delta\|_2\leq\rho}\mathcal{L}_{\mathrm{train}}(\omega+\delta), where ω\omega denotes the trainable parameters and ρ≥0\rho\geq0 is the neighborhood radius.

  2. Knowl 2 — SAMformer improves multivariate forecasting benchmarks with fewer parameters

    empirical result

    On eight real-world multivariate forecasting datasets and horizons H∈{96,192,336,720}H\in\{96,192,336,720\}, the paper reports overall test-MSE reductions of 5.25% relative to TSMixer trained with SAM, 14.33% relative to TSMixer without SAM, and 16.96% relative to the paper’s plain Transformer. The reported reductions relative to FEDformer, PatchTST, and iTransformer are 12.36%, 11.13%, and 3.94%, respectively. SAMformer is statistically better than TSMixer trained with SAM on seven of the eight datasets at the reported p<0.05p<0.05 threshold; it is ranked first or second across horizons on the datasets other than Exchange. Scores are averaged over five random-seed runs. The model’s parameter count averages 3.73 times fewer than TSMixer’s across datasets and horizons, which the paper summarizes as approximately four times fewer parameters.

  3. Knowl 3 — Trainable attention is implicated in poor generalization on a synthetic linear forecast

    empirical result

    The paper studies synthetic data generated as Y=XWtoy+εY=XW_{\mathrm{toy}}+\varepsilon, where X∈RD×LX\in\mathbb{R}^{D\times L}, Wtoy∈RL×HW_{\mathrm{toy}}\in\mathbb{R}^{L\times H}, and the entries of the inputs, target matrix, and noise are randomly generated from normal distributions. It uses D=7D=7, L=512L=512, and H=96H=96, with 15,000 input-target pairs split into 10,000 training and 5,000 validation examples. Compared with the least-squares oracle, a Transformer trained with Adam overfits and does not recover oracle validation performance. A Random Transformer, which holds its randomly initialized attention weights fixed and trains only the output projection, generalizes better than the fully trainable Transformer but remains suboptimal. The two models differ in parameter count by only about 2%, so the authors attribute the main generalization problem to training the attention module rather than that small parameter-count difference.

  4. Knowl 4 — SAM produces flatter solutions and more stable forecasts

    empirical result

    On the synthetic linear forecasting task, training the shallow Transformer with SAM brings validation performance to the least-squares oracle, whereas the ordinary Transformer overfits. The SAM-trained model has substantially lower loss sharpness: the paper reports that the largest Hessian eigenvalue for the ordinary Transformer is about 10410^4 times that for its SAM-trained counterpart in the synthetic comparison. In real-data experiments, SAMformer also has roughly an order of magnitude lower sharpness than the plain Transformer in the reported comparisons. Across five seeds, SAMformer’s test MSE is stable on the examined datasets and horizons, while the plain Transformer’s performance varies substantially with initialization. These observations support the paper’s conclusion that flat-minimum-seeking optimization improves both generalization and robustness to initialization.

  5. Knowl 5 — A rank condition characterizes when the toy Transformer can represent the target

    theoretical result

    For a fixed input X∈RD×LX\in\mathbb{R}^{D\times L} and fixed attention-related weights, let A(X)∈RD×DA(X)\in\mathbb{R}^{D\times D} be the attention matrix and define P=X+A(X)XWVWO∈RD×LP=X+A(X)XW_VW_O\in\mathbb{R}^{D\times L}. The model can represent the toy target XWtoy∈RD×HXW_{\mathrm{toy}}\in\mathbb{R}^{D\times H} with an output matrix W∈RL×HW\in\mathbb{R}^{L\times H}, meaning that PW=XWtoyPW=XW_{\mathrm{toy}}, if and only if

    rank⁡([P  XWtoy])=rank⁡(P),\operatorname{rank}\bigl([P\;XW_{\mathrm{toy}}]\bigr)=\operatorname{rank}(P),

    where [P  XWtoy][P\;XW_{\mathrm{toy}}] denotes horizontal concatenation. In the toy setting, full row rank of PP ensures representability; because the output projection has more rows than the channel dimension, there are infinitely many such projections. This result shows that the toy model’s poor fitted solution is not explained by a lack of representational capacity.

  6. Knowl 6 — A nuclear-norm bound links query-key scale to attention-score expressiveness

    theoretical result

    Let X∈RD×LX\in\mathbb{R}^{D\times L} be an input and let WQ,WK∈RL×dmW_Q,W_K\in\mathbb{R}^{L\times d_m}. If WQWK⊤=WKWQ⊤⪰0W_QW_K^{\top}=W_KW_Q^{\top}\succeq0, then the unnormalized attention-score matrix obeys

    ∥XWQWK⊤X⊤∥∗≤∥WQWK⊤∥2 ∥X∥F2.\left\|XW_QW_K^{\top}X^{\top}\right\|_*\leq \left\|W_QW_K^{\top}\right\|_2\,\|X\|_F^2.

    Here ∥⋅∥∗\|\cdot\|_* is the nuclear norm, ∥⋅∥2\|\cdot\|_2 is the spectral norm, and ∥⋅∥F\|\cdot\|_F is the Frobenius norm. The condition holds, for example, when WQ=WKW_Q=W_K. The bound says that reducing the spectral norm of the query-key product also lowers an upper bound on the nuclear norm of the score matrix. The paper uses nuclear norm as a proxy for rank-related expressiveness; the bound alone does not establish that the row-wise-softmax attention matrix must lose rank.

  7. Knowl 7 — Channel-wise attention outperforms temporal attention in the tested ablation

    empirical result

    The paper compares feature-based channel-wise attention with attention over time steps on the multivariate forecasting benchmarks. Relative to the temporal-attention variant, the Transformer using channel-wise attention improves overall MSE by 12.97% and overall MAE by 18.09%. Channel-wise attention forms a D×DD\times D matrix, rather than an L×LL\times L temporal-attention matrix. Because it operates across features and not ordered time positions, it is invariant to permutations of the feature ordering and does not require positional encoding; for the common case D≤LD\leq L, it also reduces attention’s matrix size. A separate ablation finds that fixing attention to the identity does not match SAMformer: learnable SAMformer attention improves overall MSE by 11.93% and MAE by 4.18% over the identity-attention variant.

  8. Knowl 8 — Full-shot SAMformer is competitive with zero-shot MOIRAI

    data/table

    The paper compares SAMformer trained on each evaluation dataset (full-shot) with the pretrained MOIRAI foundation model used zero-shot. Each value below is test MSE averaged over horizons {96,192,336,720}\{96,192,336,720\}; model order is SAMformer, MOIRAI-Small, MOIRAI-Base, MOIRAI-Large.

    • ETTh1: 0.410, 0.400, 0.434, 0.510.
    • ETTh2: 0.344, 0.341, 0.345, 0.354.
    • ETTm1: 0.373, 0.448, 0.381, 0.390.
    • ETTm2: 0.269, 0.300, 0.272, 0.276.
    • Electricity: 0.181, 0.233, 0.188, 0.188.
    • Weather: 0.260, 0.242, 0.238, 0.259.

    The reported overall MSE improvements for SAMformer relative to MOIRAI-Small, MOIRAI-Base, and MOIRAI-Large are 6.9%, 1.1%, and 7.6%, respectively. SAMformer is therefore competitive with these much larger zero-shot models on the evaluated datasets, although MOIRAI variants have lower MSE on some individual datasets.

  9. Knowl 9 — Spectral reparameterization does not substitute for SAM in these experiments

    empirical result

    The tested σ\sigmaReparam method replaces each weight matrix WW with Wc=γW/∥W∥2W_c=\gamma W/\|W\|_2, where γ\gamma is a learned scalar initialized to 1. On the synthetic task and on the reported ETTh1 and Exchange comparisons, this reparameterization alone does not attain SAMformer’s performance and sometimes performs worse than the ordinary Transformer. The paper observes that σ\sigmaReparam can yield attention matrices with nearly identical rows and reduced nuclear norm, consistent with rank collapse, whereas SAMformer retains more expressive attention. Combining σ\sigmaReparam with SAM does not significantly improve on SAM alone and increases training time and memory use, so the paper does not recommend the combination for these experiments.

  10. Knowl 10 — Forecasting evaluation uses fixed windows and repeated seeds

    experimental setup

    The main evaluation uses eight public datasets: ETTh1, ETTh2, ETTm1, ETTm2, Electricity, Exchange, Traffic, and Weather. Inputs have look-back length L=512L=512, forecasts use horizons H∈{96,192,336,720}H\in\{96,192,336,720\}, and windows are formed with stride 1. The ETT data use a 12/4/4-month train/validation/test split; the other datasets use 70%/20%/10%. Models are trained with Adam, batch size 32, cosine-annealed learning rates, up to 300 epochs, and early stopping with patience 5. The learning rates are 0.001 for ETT and Exchange, and 0.0001 for Electricity, Traffic, and Weather. SAMformer and the paper’s Transformer use one attention head with model dimension dm=16d_m=16. Reported in-house benchmark MSE and MAE are means and standard deviations over five random-seed runs.

Coverage note — The full per-dataset, per-horizon MSE and MAE tables, the complete SAM neighborhood-radius grid, and secondary optimizer-sensitivity details are omitted because they add extensive tuning and result values without changing the central architectural, theoretical, or comparative conclusions.

References

  1. 1.Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mane, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viegas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. TensorFlow: Large-scale machine learning on heterogeneous systems, 2015. URL http://tensorflow.org/. Software available from tensorflow.org.
  2. 2.Ahn, K., Cheng, X., Song, M., Yun, C., Jadbabaie, A., and Sra, S. Linear attention is (maybe) all you need (to understand transformer optimization), 2023.
  3. 3.Anagnostidis, S., Biggio, L., Noci, L., Orvieto, A., Singh, S. P., and Lucchi, A. Signal propagation in transformers: Theoretical perspectives and the role of rank collapse. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=FxVH7iToXS.
  4. 4.Box, G. E. P. and Jenkins, G. Time Series Analysis, Forecasting and Control. Holden-Day, Inc., USA, 1990. ISBN 0816211043.
  5. 5.Box, G. E. P., Jenkins, G. M., and MacGregor, J. F. Some Recent Advances in Forecasting and Control. Journal of the Royal Statistical Society Series C, 23(2):158–179, June 1974. doi: 10.2307/2346997. URL https://ideas.repec.org/a/bla/jorssc/v23y1974i2p158-179.html.
  6. 6.California Department of Transportation. Traffic dataset, 2021. URL https://pems.dot.ca.gov/.
  7. 7.Candes, E. and Recht, B. Exact matrix completion via convex optimization. Commun. ACM, 55(6):111–119, jun 2012. ISSN 0001-0782. doi: 10.1145/2184319.2184343. URL https://doi.org/10.1145/2184319.2184343.
  8. 8.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the International Conference on Computer Vision (ICCV), 2021.
  9. 9.Casolaro, A., Capone, V., Iannuzzo, G., and Camastra, F. Deep learning for time series forecasting: Advances and open problems. Information, 14(11), 2023. ISSN 2078-2489. doi: 10.3390/info14110598. URL https://www.mdpi.com/2078-2489/14/11/598.
  10. 10.Cepulionis, P. and Lukoseviciute, K. Electrocardiogram time series forecasting and optimization using ant colony optimization algorithm. Mathematical Models in Engineering, 2(1):69–77, Jun 2016. ISSN 2351-5279. URL https://www.extrica.com/article/17229.
  11. 11.Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. Entropy-SGD: Biasing gradient descent into wide valleys. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=B1YfAfcgl.
  12. 12.Chen, R. and Tao, M. Data-driven prediction of general hamiltonian dynamics via learning exactly-symplectic maps. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 1717–1727. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/chen21r.html.
  13. 13.Chen, S.-A., Li, C.-L., Arik, S. O., Yoder, N. C., and Pfister, T. TSMixer: An all-MLP architecture for time series forecasting. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=wbpxTuXgm0.
  14. 14.Chen, X., Hsieh, C.-J., and Gong, B. When vision transformers outperform resnets without pre-training or strong data augmentations. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=LtKcMgGOeLt.
  15. 15.Cirstea, R.-G., Guo, C., Yang, B., Kieu, T., Dong, X., and Pan, S. Triformer: Triangular, variable-specific attentions for long sequence multivariate time series forecasting. In Raedt, L. D. (ed.), Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, pp. 1994–2001. International Joint Conferences on Artificial Intelligence Organization, 7 2022. doi: 10.24963/ijcai.2022/277. URL https://doi.org/10.24963/ijcai.2022/277. Main Track.
  16. 16.Daneshmand, H., Kohler, J., Bach, F., Hofmann, T., and Lucchi, A. Batch normalization provably avoids rank collapse for randomly initialised deep networks. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546.
  17. 17.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding, 2018. URL http://arxiv.org/abs/1810.04805.
  18. 18.Dong, Y., Cordonnier, J.-B., and Loukas, A. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 2793–2803. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/dong21a.html.
  19. 19.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  20. 20.Dziugaite, G. K. and Roy, D. M. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. In Proceedings of the 33rd Annual Conference on Uncertainty in Artificial Intelligence (UAI), 2017.
  21. 21.Fan, C., Zhang, Y., Pan, Y., Li, X., Zhang, C., Yuan, R., Wu, D., Wang, W., Pei, J., and Huang, H. Multi-horizon time series forecasting with temporal attention learning. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’19, pp. 2527–2535, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450362016. doi: 10.1145/3292500.3330662. URL https://doi.org/10.1145/3292500.3330662.
  22. 22.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=6Tm1mposlrM.
  23. 23.Glorot, X. and Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In Teh, Y. W. and Titterington, M. (eds.), Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9 of Proceedings of Machine Learning Research, pp. 249–256, Chia Laguna Resort, Sardinia, Italy, 13–15 May 2010. PMLR. URL https://proceedings.mlr.press/v9/glorot10a.html.
  24. 24.He, B., Martens, J., Zhang, G., Botev, A., Brock, A., Smith, S. L., and Teh, Y. W. Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=NPrsUQgMjKK.
  25. 25.Horn, R. A. and Johnson, C. R. Topics in Matrix Analysis. Cambridge University Press, 1991.
  26. 26.Hunter, J. D. Matplotlib: A 2d graphics environment. Computing in Science & Engineering, 9(3):90–95, 2007. doi: 10.1109/MCSE.2007.55.
  27. 27.Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=H1oyRlYgg.
  28. 28.Kim, H., Papamakarios, G., and Mnih, A. The lipschitz constant of self-attention. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 5562–5571. PMLR, 18–24 Jul 2021a. URL https://proceedings.mlr.press/v139/kim21i.html.
  29. 29.Kim, T., Kim, J., Tae, Y., Park, C., Choi, J.-H., and Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=cGDAkQo1C0p.
  30. 30.Kingma, D. and Ba, J. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), San Diega, CA, USA, 2015.
  31. 31.Kitaev, N., Kaiser, L., and Levskaya, A. Reformer: The efficient transformer. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkgNKkHtvB.
  32. 32.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long- and short-term temporal patterns with deep neural networks. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, pp. 95–104, New York, NY, USA, 2018a. Association for Computing Machinery. ISBN 9781450356572. doi: 10.1145/3209978.3210006. URL https://doi.org/10.1145/3209978.3210006.
  33. 33.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long- and short-term temporal patterns with deep neural networks. In Association for Computing Machinery, SIGIR ’18, pp. 95–104, New York, NY, USA, 2018b. ISBN 9781450356572. doi: 10.1145/3209978.3210006. URL https://doi.org/10.1145/3209978.3210006.
  34. 34.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper_files/paper/2019/file/6775a0635c302542da2c32aa19d86be0-Paper.pdf.
  35. 35.Liu, L., Liu, X., Gao, J., Chen, W., and Han, J. Understanding the difficulty of training transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), 2020.
  36. 36.Liu, S., Yu, H., Liao, C., Li, J., Lin, W., Liu, A. X., and Dustdar, S. Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=0EXmFzUn5I.
  37. 37.Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M. itransformer: Inverted transformers are effective for time series forecasting. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=JePfAI8fah.
  38. 38.Loshchilov, I. and Hutter, F. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=Skq89Scxx.
  39. 39.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  40. 40.Max Planck Institute. Weather dataset, 2021. URL https://www.bgc-jena.mpg.de/wetter/.
  41. 41.Nesterov, Y. A method for solving the convex programming problem with convergence rate o(1/k2). Proceedings of the USSR Academy of Sciences, 269:543–547, 1983. URL https://api.semanticscholar.org/CorpusID:145918791.
  42. 42.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Jbdc0vTOcol.
  43. 43.OpenAI. Gpt-4 technical report, 2023.
  44. 44.Pan, Y. and Li, Y. Toward understanding why adam converges faster than SGD for transformers. In OPT 2022: Optimization for Machine Learning (NeurIPS 2022 Workshop), 2022. URL https://openreview.net/forum?id=Sf1NlV2r6PO.
  45. 45.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, pp. 8024–8035. Curran Associates, Inc., 2019.
  46. 46.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  47. 47.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018 OpenAI Tech Report, 2018.
  48. 48.Rangapuram, S. S., Seeger, M. W., Gasthaus, J., Stella, L., Wang, Y., and Januschowski, T. Deep state space models for time series forecasting. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/5cf68969fb67aa6082363a6d4e6468e2-Paper.pdf.
  49. 49.Recht, B. A simpler approach to matrix completion. J. Mach. Learn. Res., 12(null):3413–3430, dec 2011. ISSN 1532-4435.
  50. 50.Recht, B., Fazel, M., and Parrilo, P. A. Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization. SIAM Review, 52(3):471–501, 2010. doi: 10.1137/070697835. URL https://doi.org/10.1137/070697835.
  51. 51.Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T. Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020. ISSN 0169-2070. doi: https://doi.org/10.1016/j.ijforecast.2019.07.001. URL https://www.sciencedirect.com/science/article/pii/S0169207019301888.
  52. 52.Sen, R., Yu, H.-F., and Dhillon, I. Think globally, act locally: a deep neural network approach to high-dimensional time series forecasting. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2019. Curran Associates Inc.
  53. 53.Sonkavde, G., Dharrao, D. S., Bongale, A. M., Deokate, S. T., Doreswamy, D., and Bhat, S. K. Forecasting stock market prices using machine learning and deep learning models: A systematic review, performance analysis and discussion of implications. International Journal of Financial Studies, 11(3), 2023. ISSN 2227-7072. doi: 10.3390/ijfs11030094. URL https://www.mdpi.com/2227-7072/11/3/94.
  54. 54.Sorjamaa, A., Hao, J., Reyhani, N., Ji, Y., and Lendasse, A. Methodology for long-term prediction of time series. Neurocomputing, 70(16):2861–2869, 2007. ISSN 0925-2312. doi: https://doi.org/10.1016/j.neucom.2006.06.015. URL https://www.sciencedirect.com/science/article/pii/S0925231207001610. Neural Network Applications in Electrical Engineering Selected papers from the 3rd International Work-Conference on Artificial Neural Networks (IWANN 2005).
  55. 55.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image transformers and distillation through attention. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 10347–10357. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/touvron21a.html.
  56. 56.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023. URL http://arxiv.org/abs/2302.13971. cite arxiv:2302.13971.
  57. 57.Trockman, A. and Kolter, J. Z. Mimetic initialization of self-attention layers. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023.
  58. 58.UCI. Electricity dataset, 2015. URL https://archive.ics.uci.edu/dataset/321/electricityloaddiagrams20112014.
  59. 59.Van Rossum, G. and Drake Jr, F. L. Python reference manual. Centrum voor Wiskunde en Informatica Amsterdam, 1995.
  60. 60.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  61. 61.Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series forecasting transformers, 2024.
  62. 62.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with Auto-Correlation for long-term series forecasting. In Advances in Neural Information Processing Systems, 2021.
  63. 63.Zamir, S. W., Arora, A., Khan, S., Hayat, M., Khan, F. S., and Yang, M.-H. Restormer: Efficient transformer for high-resolution image restoration. In CVPR, 2022.
  64. 64.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI Conference on Artificial Intelligence, 2023.
  65. 65.Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M. Stabilizing transformer training by preventing attention entropy collapse. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 40770–40803. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/zhai23a.html.
  66. 66.Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Lin, H., Zhang, Z., Sun, Y., He, T., Mueller, J., Manmatha, R., Li, M., and Smola, A. Resnest: Split-attention networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 2736–2746, June 2022.
  67. 67.Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S. Why are adaptive methods good for attention models? In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 15383–15393. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/b05b57f6add810d3b7490866d74c0053-Paper.pdf.
  68. 68.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In The Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Virtual Conference, volume 35, pp. 11106–11115. AAAI Press, 2021.
  69. 69.Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting. In Proc. 39th International Conference on Machine Learning (ICML 2022), 2022.

Citation

MLA
Ilbert, R., et al. “SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.10198.
APA
Ilbert, R., Odonnat, A., Feofanov, V., Virmaux, A., Paolo, G., Palpanas, T., & Redko, I. (2024). SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention. arXiv. https://doi.org/10.48550/arxiv.2402.10198
Chicago
Ilbert, R., A. Odonnat, V. Feofanov, et al. 2024. “SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.10198.
Harvard
Ilbert, R. et al. (2024) “SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.10198.
Vancouver
1. Ilbert R, Odonnat A, Feofanov V, Virmaux A, Paolo G, Palpanas T, Redko I (2024) SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention. https://doi.org/10.48550/arxiv.2402.10198

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.10198,
  doi = {10.48550/ARXIV.2402.10198},
  url = {https://arxiv.org/abs/2402.10198},
  author = {Ilbert, Romain and Odonnat, Ambroise and Feofanov, Vasilii and Virmaux, Aladin and Paolo, Giuseppe and Palpanas, Themis and Redko, Ievgen},
  keywords = {Machine Learning (cs.LG), Machine Learning (stat.ML), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/