Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling

Guoqi YuJing ZouXiaowei HuAngelica I. Avilés-RiveroJing QinShujun Wang

article2024ICML111 citations

Proposes a learnable convolutional trend-seasonal decomposition paired with a dual attention module to capture cross-variable dependencies and temporal variations, reducing forecasting error by up to 48.56% across multiple benchmark datasets.

Listen

Multivariate time series forecasting plays a critical role in strategic planning across vital sectors, including energy grid management, transportation networks, weather prediction, and public health monitoring. Despite recent advances in machine learning, existing models struggle to deliver reliable forecasts because real-world data contains complex, non-linear trends and intricate relationships. Traditional approaches rely on static moving averages that fail to isolate dynamic trends, while current attention-based methods often compromise data integrity by partitioning sequences into smaller patches or ignoring interactions across different variables.

The article introduces and evaluates Leddam (LEarnable Decomposition and Dual Attention Module), a novel forecasting framework designed to improve prediction accuracy. The system replaces rigid, static decomposition with an adaptable convolutional process to isolate trend and seasonal patterns, and introduces a dual attention mechanism to simultaneously capture dependencies across different variables and variations over time without discarding raw sequence data.

The authors conducted extensive empirical evaluations across eight standard open-source datasets encompassing electricity transformer temperatures, power consumption, traffic occupancy, solar generation, and weather indicators. The evaluation benchmarked Leddam against eight current leading forecasting models across multiple prediction time horizons ranging from 96 to 720 time steps, measuring accuracy via standard error metrics (Mean Squared Error and Mean Absolute Error).

The evaluation produced three core findings. First, Leddam outperformed all eight benchmark models, achieving the best overall accuracy on seven of the eight datasets; across 32 total forecasting evaluation settings, it secured top performance in 31 instances. Second, the proposed trainable decomposition module consistently isolated seasonal frequencies and trend shifts more effectively than standard moving average filters, delivering an average 11.98% error reduction when tested within existing baseline architectures. Third, the decomposition module demonstrated exceptional plug-and-play generality across diverse neural architectures, reducing prediction error by an average of 11.87% in linear models, 23.15% in convolutional networks, 26.27% to 31.72% in standard transformer models, and up to 48.56% in classical recurrent neural networks.

These findings indicate that forecasting performance can be substantially improved by refining data representation and sequence decomposition, rather than solely increasing model size or computational complexity. For operational leaders, adopting these techniques translates into more dependable forecasts for high-stakes infrastructure, reducing operational risks, avoiding costly energy imbalances, and optimizing asset scheduling. Organizations maintaining legacy predictive models can achieve double-digit accuracy gains with minimal structural refactoring simply by integrating the learnable decomposition module.

Organizations seeking to upgrade time-series predictive systems should evaluate the open-source Leddam framework against their internal operational data, particularly for applications requiring extended forecast horizons. Additionally, technical teams managing existing forecasting pipelines should test replacing fixed moving average preprocessing with learnable decomposition modules to capture quick performance gains.

Confidence in these findings is high given the broad benchmark coverage, standard validation protocols, and consistent improvements across varied datasets. However, stakeholders should note that the evaluation was confined to publicly available benchmark datasets with fixed historical horizons. Operational environments characterized by extreme data sparsity, sudden structural market shifts, or highly irregular sampling rates may require targeted pilot testing before full-scale deployment.

Yu et al (2024).pdf

No sufficiently relevant recommendations were found.

Cover for Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling

Abstract

Predicting multivariate time series is crucial, demanding precise modeling of intricate patterns, including inter-series dependencies and intra-series variations. Distinctive trend characteristics in each time series pose challenges, and existing methods, relying on basic moving average kernels, may struggle with the non-linear structure and complex trends in real-world data. Given that, we introduce a learnable decomposition strategy to capture dynamic trend information more reasonably. Additionally, we propose a dual attention module tailored to capture inter-series dependencies and intra-series variations simultaneously for better time series forecasting, which is implemented by channel-wise self-attention and autoregressive self-attention. To evaluate the effectiveness of our method, we conducted experiments across eight open-source datasets and compared it with the state-of-the-art methods. Through the comparison results, our Leddam (LEarnable Decomposition and Dual Attention Module) not only demonstrates significant advancements in predictive performance but also the proposed decomposition strategy can be plugged into other methods with a large performance-boosting, from 11.87% to 48.56% MSE error degradation. Code is available at this link: https://github.com/Levi-Ackman/Leddam.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Problem Definition
  • 3.2. Learnable Decomposition Module
  • 3.3. Dual Attention Module
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Experiments Results
  • 4.3. Model Analysis
  • 4.4. Learnable Decomposition Generalization Analysis
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Experimental Details
  • A.1. Dataset Statistics
  • A.2. Implementation Details and Model Parameters
  • B. Further Analysis of Different Decomposition Methodologies
  • B.1. Decomposition Result Visualization
  • B.2. Visualization of Weights
  • B.3. Impact of RevIN
  • B.4. Further Improvement of Multi-scale Hybrid Decomposition
  • C. Further Analysis of Auto-regressive Self-attention
  • D. Generalization Analysis of Leddam
  • E. Model Efficiency Study
  • F. Hyperparameter Sensitivity Analysis
  • G. Full Forecasting Results
  • H. Influence of Input Length on Prediction Performance

Knowls

  1. Knowl 1 — Leddam combines learnable decomposition with two seasonal-pattern branches

    model/method

    Leddam forecasts a multivariate series by first projecting each of its NN observed channels, each of length TT, into an embedding of dimension DD and adding positional encoding. It decomposes the embedded representation into trend and seasonal components. A linear projection maps the trend to the FF-step forecast; the seasonal component is processed by two parallel branches, one modeling dependencies between channels and one modeling variation over time within each channel. The two seasonal-branch outputs are added, linearly projected to FF steps, and added to the trend forecast to produce the final prediction. Here NN is the number of variates, TT the look-back length, DD the embedding dimension, and FF the forecast horizon.

  2. Knowl 2 — A Gaussian-initialized shared convolution learns the trend decomposition

    model/method

    For an embedded multivariate input E∈RN×DE\in\mathbb{R}^{N\times D}, Leddam applies a shared, learnable one-dimensional convolution separately to each channel to obtain its trend EtrendE_{\mathrm{trend}}; the seasonal component is the residual Eseasonal=E−EtrendE_{\mathrm{seasonal}}=E-E_{\mathrm{trend}}. Terminal-value padding and stride 11 preserve the embedding length. The kernel has size K=25K=25 and is initialized with normalized Gaussian-shaped weights: Ui=exp⁡ ⁣(−(i−K/2)22σ2)U_i=\exp\!\left(-\frac{(i-K/2)^2}{2\sigma^2}\right) and ωi=softmax⁡(U)i\omega_i=\operatorname{softmax}(U)_i, for i∈{1,…,K}i\in\{1,\ldots,K\} and σ=1.0\sigma=1.0. The weights are trainable after initialization, allowing the trend extraction to adapt to the data rather than using fixed uniform moving-average weights.

  3. Knowl 3 — Channel-wise attention models dependencies between complete variate histories

    model/method

    In Leddam's inter-series branch, each row of the seasonal representation is a token: one token represents the complete embedded history of one variate, rather than a value at one time point or a temporal patch. A Transformer encoder applies self-attention across the NN variate tokens, producing an output for each channel. In standard attention notation, the operation is softmax⁡(QKT/dk)V\operatorname{softmax}(QK^{\mathsf T}/\sqrt{d_k})V, where QQ, KK, and VV are query, key, and value projections of the seasonal channel tokens and dkd_k is the key dimension; residual connections, layer normalization, and a feed-forward network complete the encoder block. This representation is intended to preserve each channel's whole-series information while learning inter-series dependencies.

  4. Knowl 4 — Autoregressive embeddings let attention model within-series variation

    model/method

    For each channel's seasonal embedding x∈RDx\in\mathbb{R}^{D}, Leddam forms cyclically shifted full-length tokens using a shift interval LL: token jj is S[j,:]=x[jL:D] ∥ x[0:jL]S[j,:]=x[jL:D]\,\|\,x[0:jL], for j=0,…,⌊D/L⌋−1j=0,\ldots,\lfloor D/L\rfloor-1, where ∥\| denotes concatenation. Thus, each token contains the full embedded sequence in a different temporal ordering, rather than only a sampled subsequence or patch. In a Transformer encoder shared across channels, the original sequence supplies the query, while the shifted-token collection supplies keys and values; the resulting channel outputs are concatenated. This autoregressive self-attention branch is designed to capture intra-series temporal variation while retaining the original sequence's information in the generated tokens.

  5. Knowl 5 — Leddam improves average long-horizon results on seven of eight benchmarks

    empirical result

    The benchmark used the ETTh1, ETTh2, ETTm1, ETTm2, Electricity, Solar-Energy, Traffic, and Weather datasets, with chronological train/validation/test splits of 6:2:2 for ETT and 7:1:2 for the others. All models used look-back length T=96T=96; results below average over forecast lengths F∈{96,192,336,720}F\in\{96,192,336,720\} and report mean squared error (MSE) and mean absolute error (MAE), where lower is better. Leddam's MSE/MAE was ETTh1 0.431/0.4290.431/0.429, ETTh2 0.373/0.3990.373/0.399, ETTm1 0.386/0.3970.386/0.397, ETTm2 0.281/0.3250.281/0.325, Electricity 0.169/0.2630.169/0.263, Solar-Energy 0.230/0.2640.230/0.264, Traffic 0.467/0.2940.467/0.294, and Weather 0.242/0.2720.242/0.272. Against iTransformer, the strongest general comparator in these averages, the corresponding scores were 0.454/0.4470.454/0.447, 0.383/0.4070.383/0.407, 0.407/0.4100.407/0.410, 0.288/0.3320.288/0.332, 0.178/0.2700.178/0.270, 0.233/0.2620.233/0.262, 0.428/0.2820.428/0.282, and 0.258/0.2790.258/0.279. Leddam therefore had the lowest MSE on seven datasets and the lowest MAE on six; iTransformer performed better on both metrics on Traffic, and had lower MAE on Solar-Energy.

  6. Knowl 6 — Both attention branches contribute, and their combination performs best in ablation

    empirical result

    An ablation compared Leddam's two attention branches with a replacement using a linear layer for both, using input and forecast lengths T=F=96T=F=96. The tested datasets and MSE/MAE scores were: ETTm1—linear replacement 0.350/0.3680.350/0.368, autoregressive branch only 0.337/0.3710.337/0.371, channel-wise branch only 0.335/0.3690.335/0.369, both branches 0.320/0.3590.320/0.359; Electricity—0.197/0.2750.197/0.275, 0.148/0.2420.148/0.242, 0.152/0.2420.152/0.242, 0.139/0.2330.139/0.233; Traffic—0.644/0.3890.644/0.389, 0.438/0.2880.438/0.288, 0.467/0.2850.467/0.285, 0.424/0.2690.424/0.269; Weather—0.195/0.2350.195/0.235, 0.168/0.2120.168/0.212, 0.167/0.2120.167/0.212, 0.158/0.2030.158/0.203; Solar-Energy—0.306/0.3300.306/0.330, 0.211/0.2530.211/0.253, 0.226/0.2630.226/0.263, 0.202/0.2400.202/0.240. The authors report average MSE reductions relative to the linear replacement of 19.02% for autoregressive attention alone, 21.09% for channel-wise attention alone, and 25.03% for their combination. Both branches improve the results, and the joint model has the best MSE and MAE in every listed dataset.

  7. Knowl 7 — Replacing DLinear's moving average with the learnable decomposition lowers error

    empirical result

    The authors replaced the moving-average decomposition in DLinear with the proposed convolutional decomposition, testing both an untrainable kernel (LD UTL) and a trainable kernel (LD TL) at input length T=96T=96 and forecast length F=720F=720. The reported MSE/MAE scores for DLinear, LD UTL, and LD TL, respectively, were: ETTh2 0.831/0.6570.831/0.657, 0.742/0.6070.742/0.607, 0.684/0.5760.684/0.576; ETTm2 0.554/0.5220.554/0.522, 0.541/0.5050.541/0.505, 0.521/0.4880.521/0.488; Electricity 0.245/0.3330.245/0.333, 0.220/0.3110.220/0.311, 0.209/0.3020.209/0.302; Traffic 0.645/0.3940.645/0.394, 0.602/0.3780.602/0.378, 0.584/0.3450.584/0.345; Solar-Energy 0.356/0.4130.356/0.413, 0.333/0.3970.333/0.397, 0.313/0.3760.313/0.376. Both convolutional variants outperform the moving-average baseline on all five datasets and both metrics. The paper reports average MSE reductions of 7.28% for LD UTL and 11.98% for LD TL, indicating an additional benefit from kernel trainability.

  8. Knowl 8 — Decomposition diagnostics favor the learnable kernel over moving average

    empirical result

    The paper evaluated decomposition of the final variate in each of eight datasets, averaging diagnostic scores over the test set. FFT denotes the reported amplitude similarity between the raw series and the seasonal component at the dominant frequencies (top 25%); DTW denotes the reported similarity between the raw series and the trend component. Higher scores are treated as better. For learnable decomposition (LD) versus moving average (MOV), the DTW/FFT scores were: Electricity 0.643/0.9420.643/0.942 vs. 0.618/0.9310.618/0.931; Solar-Energy 0.910/0.7810.910/0.781 vs. 0.873/0.7340.873/0.734; Traffic 0.603/0.9930.603/0.993 vs. 0.563/0.9910.563/0.991; Weather 0.858/0.7600.858/0.760 vs. 0.846/0.6910.846/0.691; ETTh1 0.741/0.8920.741/0.892 vs. 0.724/0.8520.724/0.852; ETTh2 0.675/0.9000.675/0.900 vs. 0.652/0.8870.652/0.887; ETTm1 0.821/0.7540.821/0.754 vs. 0.808/0.7170.808/0.717; and ETTm2 0.925/0.8940.925/0.894 vs. 0.908/0.7590.908/0.759. LD scores are higher for both diagnostics on every dataset.

  9. Knowl 9 — Learnable decomposition improves several unrelated forecasting architectures

    empirical result

    To test whether learnable decomposition can transfer beyond Leddam, the authors incorporated it into LightTS, LSTM, SCINet, Informer, and Transformer and compared each augmented model with its original version. Evaluation used T=96T=96 and F=96F=96 on ETTh2, ETTm2, Weather, Electricity, and Traffic. The reported average MSE reductions across these five datasets were 11.87% for LightTS, 48.56% for LSTM, 23.15% for SCINet, 31.72% for Informer, and 26.27% for Transformer. LSTM's MSE fell by 76.97% on ETTh2 and 80.28% on ETTm2. These experiments support the module's applicability across linear, recurrent, convolutional, and Transformer-based forecasting models under the tested conditions.

  10. Knowl 10 — Leddam's measured inference cost is below several Transformer baselines but above iTransformer

    data/table

    Inference efficiency was measured on one NVIDIA V100 32GB GPU using batch size 1, as the mean of five runs, for input length 96 and forecast length 720 with two network layers. Each entry below is parameter count in millions / inference time in milliseconds; model order is Leddam, PatchTST, Crossformer, iTransformer, FEDformer. On ETTh1 at embedding dimension 256, the entries were 2.50/233.922.50/233.92, 3.27/251.003.27/251.00, 8.19/399.008.19/399.00, 1.27/177.671.27/177.67, and 3.43/303.5563.43/303.556; at dimension 512, they were 9.20/249.349.20/249.34, 8.64/266.668.64/266.66, 32.11/445.7432.11/445.74, 4.63/190.924.63/190.92, and 13.68/345.73613.68/345.736. On Electricity at dimension 256, they were 2.59/283.042.59/283.04, 3.27/322.533.27/322.53, 13.66/432.4013.66/432.40, 1.27/192.121.27/192.12, and 4.24/347.6344.24/347.634; at dimension 512, they were 9.36/296.709.36/296.70, 8.64/411.968.64/411.96, 43.04/507.5443.04/507.54, 4.63/249.604.63/249.60, and 15.29/398.59915.29/398.599. Leddam is faster and smaller than Crossformer and FEDformer in all these settings and generally has lower parameter counts and inference times than PatchTST, but iTransformer has lower parameter counts and inference times in every listed setting.

Coverage note — Supplementary checks on RevIN, multi-scale decomposition, positional-encoding sensitivity, hyperparameter sensitivity, and longer look-back windows are omitted because they serve as secondary robustness analyses rather than central method or result contributions.

References

  1. 1.Dynamic Time Warping, pp. 69–84. Springer Berlin Heidelberg, Berlin, Heidelberg, 2007. ISBN 978-3-540-74048-3. doi: 10.1007/978-3-540-74048-3 4. URL https://doi.org/10.1007/978-3-540-74048-3_4.
  2. 2.Anonymous. FITS: Modeling time series with 10k10k parameters. In The Twelfth International Conference on Learning Representations, 2024a. URL https://openreview.net/forum?id=bWcnvZ3qMb.
  3. 3.Anonymous. ModernTCN: A modern pure convolution structure for general time series analysis. In The Twelfth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=vpJMJerXHU.
  4. 4.Anonymous. Timemixer: Decomposable multiscale mixing for time series forecasting. In The Twelfth International Conference on Learning Representations, 2024c. URL https://openreview.net/forum?id=7oLshfEIC2.
  5. 5.Das, A., Kong, W., Leach, A., Mathur, S. K., Sen, R., and Yu, R. Long-term forecasting with TiDE: Time-series dense encoder. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=pCbC3aQB5W.
  6. 6.Dong, J., Wu, H., Zhang, H., Zhang, L., Wang, J., and Long, M. SimMTM: A simple pre-training framework for masked time-series modeling. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=ginTcBUnL8.
  7. 7.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=YicbFdNTTy.
  8. 8.Han, L., Ye, H.-J., and Zhan, D.-C. The capacity and robustness trade-off: Revisiting the channel independent strategy for multivariate time series forecasting. Apr 2023.
  9. 9.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, pp. 1735–1780, Nov 1997. doi: 10.1162/neco.1997.9.8.1735. URL http://dx.doi.org/10.1162/neco.1997.9.8.1735.
  10. 10.Kim, T., Kim, J., Tae, Y., Park, C., Choi, J.-H., and Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=cGDAkQo1C0p.
  11. 11.Kitaev, N., Kaiser, L., and Levskaya, A. Reformer: The efficient transformer. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkgNKkHtvB.
  12. 12.Lai, G., Chang, W.-C., Yang, Y., and Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, pp. 95–104, 2018.
  13. 13.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper_files/paper/2019/file/6775a0635c302542da2c32aa19d86be0-Paper.pdf.
  14. 14.Li, Z., Qi, S., Li, Y., and Xu, Z. Revisiting long-term time series forecasting: An investigation on linear mapping. arXiv preprint arXiv:2305.10721, 2023.
  15. 15.LIU, M., Zeng, A., Chen, M., Xu, Z., LAI, Q., Ma, L., and Xu, Q. SCINet: Time series modeling and forecasting with sample convolution and interaction. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=AyajSjTAzmg.
  16. 16.Liu, S., Yu, H., Liao, C., Li, J., Lin, W., Liu, A. X., and Dustdar, S. Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting. In International Conference on Learning Representations, 2022a. URL https://openreview.net/forum?id=0EXmFzUn5I.
  17. 17.Liu, Y., Wu, H., Wang, J., and Long, M. Non-stationary transformers: Exploring the stationarity in time series forecasting. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022b. URL https://openreview.net/forum?id=ucNDIDRNjjv.
  18. 18.Liu, Y., Li, C., Wang, J., and Long, M. Koopa: Learning non-stationary time series dynamics with koopman predictors. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=A4zzxu82a7.
  19. 19.Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., and Long, M. itransformer: Inverted transformers are effective for time series forecasting. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=JePfAI8fah.
  20. 20.Ni, Z., Yu, H., Liu, S., Li, J., and Lin, W. Basisformer: Attention-based time series forecasting with learnable and interpretable basis. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=xx3qRKvG0T.
  21. 21.Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Jbdc0vTOcol.
  22. 22.Rangapuram, S. S., Seeger, M. W., Gasthaus, J., Stella, L., Wang, Y., and Januschowski, T. Deep state space models for time series forecasting. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/5cf68969fb67aa6082363a6d4e6468e2-Paper.pdf.
  23. 23.Shao, Z., Zhang, Z., Wang, F., Wei, W., and Xu, Y. Spatial-temporal identity: A simple yet effective baseline for multivariate time series forecasting. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, pp. 4454–4458, 2022.
  24. 24.Taylor, S. J. and Letham, B. Forecasting at scale. The American Statistician, 72(1):37–45, 2018. doi: 10.1080/00031305.2017.1380080. URL https://doi.org/10.1080/00031305.2017.1380080.
  25. 25.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  26. 26.Wang, H., Peng, J., Huang, F., Wang, J., Chen, J., and Xiao, Y. MICN: Multi-scale local and global context modeling for long-term series forecasting. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=zt53IDUR1U.
  27. 27.Wu, D., Hu, J. Y.-C., Li, W., Chen, B.-Y., and Liu, H. STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=6iwg437CZs.
  28. 28.Wu, H., Xu, J., Wang, J., and Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 22419–22430. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper_files/paper/2021/file/bcc0d400288793e8bdcd7c19a8ac0c2b-Paper.pdf.
  29. 29.Wu, H., Hu, T., Liu, Y., Zhou, H., Wang, J., and Long, M. TimesNet: Temporal 2d-variation modeling for general time series analysis. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ju_Uqw384Oq.
  30. 30.Yi, K., Zhang, Q., Fan, W., He, H., Hu, L., Wang, P., An, N., Cao, L., and Niu, Z. FourierGNN: Rethinking multivariate time series forecasting from a pure graph perspective. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=bGs1qWQ1Fx.
  31. 31.Yi, K., Zhang, Q., Fan, W., Wang, S., Wang, P., He, H., An, N., Lian, D., Cao, L., and Niu, Z. Frequency-domain MLPs are more effective learners in time series forecasting. In Thirty-seventh Conference on Neural Information Processing Systems, 2023b. URL https://openreview.net/forum?id=iif9mGCTfy.
  32. 32.Zeng, A., Chen, M., Zhang, L., and Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 11121–11128, 2023. URL https://ojs.aaai.org/index.php/AAAI/article/view/26317/26089.
  33. 33.Zhang, T., Zhang, Y., Cao, W., Bian, J., Yi, X., Zheng, S., and Li, J. Less is more: Fast multivariate time series forecasting with light sampling-oriented mlp structures. arXiv preprint arXiv:2207.01186, 2022. URL https://arxiv.org/abs/2207.01186.
  34. 34.Zhang, Y. and Yan, J. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=vSVLM2j9eie.
  35. 35.Zhao, Z., Chen, W., Wu, X., Chen, P. C. Y., and Liu, J. Lstm network: a deep learning approach for short-term traffic forecast. IET Intelligent Transport Systems, pp. 68–75, Mar 2017. doi: 10.1049/iet-its.2016.0208. URL http://dx.doi.org/10.1049/iet-its.2016.0208.
  36. 36.Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, pp. 11106–11115, Sep 2022a. doi: 10.1609/aaai.v35i12.17325. URL http://dx.doi.org/10.1609/aaai.v35i12.17325.
  37. 37.Zhou, T., MA, Z., wang, x., Wen, Q., Sun, L., Yao, T., Yin, W., and Jin, R. Film: Frequency improved legendre memory model for long-term time series forecasting. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 12677–12690. Curran Associates, Inc., 2022b.
  38. 38.Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. CoRR, abs/2201.12740, 2022c. URL https://arxiv.org/abs/2201.12740.

Citation

MLA
Yu, G., et al. “Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.12694.
APA
Yu, G., Zou, J., Hu, X., Aviles-Rivero, A. I., Qin, J., & Wang, S. (2024). Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling. arXiv. https://doi.org/10.48550/arxiv.2402.12694
Chicago
Yu, G., J. Zou, X. Hu, A. I. Aviles-Rivero, J. Qin, and S. Wang. 2024. “Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.12694.
Harvard
Yu, G. et al. (2024) “Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.12694.
Vancouver
1. Yu G, Zou J, Hu X, Aviles-Rivero AI, Qin J, Wang S (2024) Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling. https://doi.org/10.48550/arxiv.2402.12694

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.12694,
  doi = {10.48550/ARXIV.2402.12694},
  url = {https://arxiv.org/abs/2402.12694},
  author = {Yu, Guoqi and Zou, Jing and Hu, Xiaowei and Aviles-Rivero, Angelica I. and Qin, Jing and Wang, Shujun},
  keywords = {Machine Learning (cs.LG), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Revitalizing Multivariate Time Series Forecasting: Learnable Decomposition with Inter-Series Dependencies and Intra-Series Variations Modeling},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/