xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition

Artyom StitsyukJaesik Choi

article2025AAAI110 citations

Proposes xPatch, an effective non-transformer architecture combining exponential moving average seasonal-trend decomposition with dual MLP and CNN streams to surpass attention-based models in long-term time series forecasting.

Listen

Accurate long-term forecasting of complex time-dependent data is critical for operational planning, resource allocation, and risk management across sectors like energy, transportation, and weather forecasting. While recent advanced machine learning models based on transformer architectures have dominated this space, they struggle to effectively preserve chronological order and temporal relationships due to their underlying design. Moreover, standard data-smoothing techniques frequently distort early and late historical signals, degrading forecast quality.

The article demonstrates that a non-transformer architecture can outperform leading transformer models in long-term forecasting while requiring substantially fewer computational resources. The authors evaluate this premise by introducing xPatch, an architecture combining a new exponential trend-smoothing method, a dual-stream processing network, and tailored training optimization schemes.

The researchers evaluated the proposed approach against nine state-of-the-art models across nine widely used real-world benchmark datasets representing energy grids, traffic flows, weather conditions, currency exchange, and disease tracking. The evaluation compared performance across both fixed baseline settings and comprehensive hyperparameter optimization, measuring average forecast error and computational runtimes.

The evaluation yielded several critical findings. First, xPatch achieved superior accuracy, leading top benchmark models on 70% of datasets under mean squared error and 90% under mean absolute error during hyperparameter searches. Second, xPatch outperformed prominent transformer architectures, reducing average error metrics by approximately 4% to 8% compared to leading models such as CARD and PatchTST. Third, xPatch demonstrated high computational efficiency: its training time per step (about 3.1 milliseconds) and inference time (about 1.3 milliseconds) were roughly two to five times faster than competing transformer frameworks.

These findings suggest that complex, resource-heavy transformer architectures are not necessary to achieve state-of-the-art forecasting performance. Organizations can deploy lighter, non-transformer models that combine linear and non-linear paths to capture both steady trends and repeating seasonal fluctuations, thereby lowering cloud compute costs, accelerating prediction delivery, and maintaining higher forecast reliability.

Organizations managing time-critical forecasting operations should consider adopting dual-stream, non-transformer architectures like xPatch to improve accuracy while reducing computational infrastructure overhead. Before broad operational rollout, teams should conduct pilot testing on internal domain-specific datasets to evaluate the impact of historical lookback window sizing and determine whether to utilize channel-independent configurations.

Confidence in these findings is high for standard forecasting benchmarks spanning up to 720 future time steps. However, readers should note that performance depends on setting appropriate smoothing factors and lookback lengths, and real-world results may vary in operating environments with extreme data anomalies, sudden structural shifts, or missing observations.

arXiv: 2412.17323

No sufficiently relevant recommendations were found.

Cover for xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition

Abstract

In recent years, the application of transformer-based models in time-series forecasting has received significant attention. While often demonstrating promising results, the transformer architecture encounters challenges in fully exploiting the temporal relations within time series data due to its attention mechanism. In this work, we design eXponential Patch (xPatch for short), a novel dual-stream architecture that utilizes exponential decomposition. Inspired by the classical exponential smoothing approaches, xPatch introduces the innovative seasonal-trend exponential decomposition module. Additionally, we propose a dual-flow architecture that consists of an MLP-based linear stream and a CNN-based non-linear stream. This model investigates the benefits of employing patching and channel-independence techniques within a non-transformer model. Finally, we develop a robust arctangent loss function and a sigmoid learning rate adjustment scheme, which prevent overfitting and boost forecasting performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Method
  • 3.1 Seasonal-Trend Decomposition
  • 3.2 Model Architecture
  • 3.3 Loss Function
  • 3.4 Learning Rate Adjustment Scheme
  • 4 Experiments
  • 5 Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — xPatch Dual-Stream Architecture

    model/method

    The xPatch architecture is a dual-stream neural network designed for multivariate long-term time series forecasting that combines linear and non-linear processing paths without using self-attention.

    Given an input multivariate time series x∈RM×Lx \in \mathbb{R}^{M \times L} across MM variables and lookback horizon LL, the model applies channel independence to split the input into MM univariate sequences x(i)∈RLx^{(i)} \in \mathbb{R}^L for i∈{1,…,M}i \in \{1, \dots, M\}. Each univariate series is processed independently as follows:

    1. Reversible Instance Normalization (RevIN): The univariate sequence is normalized to mitigate distribution shift.
    2. Exponential Decomposition: An Exponential Moving Average (EMA) decomposition module separates x(i)x^{(i)} into a trend component xT(i)∈RLx_T^{(i)} \in \mathbb{R}^L and a seasonal component xS(i)∈RLx_S^{(i)} \in \mathbb{R}^L.
    3. Dual-Flow Processing:
      • The trend component xT(i)x_T^{(i)} is fed into a linear stream consisting of Multi-Layer Perceptrons (MLPs), average pooling, and layer normalization without activation functions, yielding linear prediction features x^lin(i)∈RT\hat{x}_{\text{lin}}^{(i)} \in \mathbb{R}^T.
      • The seasonal component xS(i)x_S^{(i)} is fed into a non-linear stream utilizing patching, depthwise separable convolutions, GELU activations, and batch normalization, yielding non-linear prediction features x^nonlin(i)∈RT\hat{x}_{\text{nonlin}}^{(i)} \in \mathbb{R}^T.
    4. Feature Fusion: The linear and non-linear representations are concatenated and passed through a final linear layer to produce the predicted horizon x^(i)∈RT\hat{x}^{(i)} \in \mathbb{R}^T:

    x^(i)=Linear(concat(x^lin(i),x^nonlin(i)))\hat{x}^{(i)} = \text{Linear}(\text{concat}(\hat{x}_{\text{lin}}^{(i)}, \hat{x}_{\text{nonlin}}^{(i)}))

    1. Denormalization: The prediction x^(i)\hat{x}^{(i)} is denormalized via RevIN to produce the final multivariate output x^∈RM×T\hat{x} \in \mathbb{R}^{M \times T}.
  2. Knowl 2 — Exponential Moving Average (EMA) Seasonal-Trend Decomposition

    model/method

    The Exponential Moving Average (EMA) decomposition module partitions a 1D time series sequence X=(x0,x1,…,xL−1)X = (x_0, x_1, \dots, x_{L-1}) into a trend component XTX_T and a seasonal component XSX_S.

    The smoothed trend point sts_t at timestep tt is computed recursively:

    s0=x0s_0 = x_0

    st=αxt+(1−α)st−1,t>0s_t = \alpha x_t + (1 - \alpha)s_{t-1}, \quad t > 0

    where α∈(0,1)\alpha \in (0, 1) is the smoothing factor. The trend and seasonal components are then extracted as:

    XT=EMA(X)=(s0,s1,…,sL−1)X_T = \text{EMA}(X) = (s_0, s_1, \dots, s_{L-1})

    XS=X−XTX_S = X - X_T

    Unlike Simple Moving Average (SMA) decomposition—which uses average pooling over a window of length kk and requires zero or edge padding at both ends of the series—EMA assigns exponentially decreasing weights to older observations. This avoids boundary distortion from padding and enables faster adaptation to evolving trends.

  3. Knowl 3 — Linear Stream in xPatch

    model/method

    The linear stream in the xPatch architecture processes the decomposed trend component xT(i)∈RLx_T^{(i)} \in \mathbb{R}^L using an MLP bottleneck structure without non-linear activation functions to preserve linear features.

    The trend sequence is processed through two successive linear blocks, each consisting of a fully connected layer, average pooling with a kernel size k=2k = 2 for feature smoothing and dimensionality reduction, and layer normalization for training stability:

    x(i)=LayerNorm(AvgPool(Linear(x(i)),k=2))x^{(i)} = \text{LayerNorm}(\text{AvgPool}(\text{Linear}(x^{(i)}), k = 2))

    The resulting compressed bottleneck representation is then expanded to the target forecasting horizon TT by a final linear projection layer:

    x^lin(i)=Linear(x(i))∈RT\hat{x}_{\text{lin}}^{(i)} = \text{Linear}(x^{(i)}) \in \mathbb{R}^T

  4. Knowl 4 — Non-Linear Stream with Patching and Depthwise Separable Convolutions

    model/method

    The non-linear stream of the xPatch architecture extracts periodic and spatio-temporal representations from the seasonal component xS(i)∈RLx_S^{(i)} \in \mathbb{R}^L using patching and depthwise separable convolutions.

    1. Patching: The input xS(i)x_S^{(i)} is segmented using a sliding window of patch length PP and stride SS into NN patches xp(i)∈RN×Px_p^{(i)} \in \mathbb{R}^{N \times P}, where:

    N=⌊L−PS⌋+2N = \left\lfloor \frac{L - P}{S} \right\rfloor + 2

    In standard configuration, P=16P = 16 and S=8S = 8.

    1. Patch Embedding: Patches are embedded into a higher-dimensional space and passed through a GELU activation σ\sigma and batch normalization:

    xpN×P2=BatchNorm(σ(Embed(xp(i))))x_p^{N \times P^2} = \text{BatchNorm}(\sigma(\text{Embed}(x_p^{(i)})))

    1. Depthwise Convolution: A grouped convolution is applied where the number of groups gg equals the number of patches NN, with kernel size k=Pk = P and stride s=Ps = P, followed by batch normalization, activation, and a residual connection:

    xpN×P=ConvN→N(xpN×P2,k=P,s=P,g=N)x_p^{N \times P} = \text{Conv}_{N \to N}(x_p^{N \times P^2}, k = P, s = P, g = N)

    xpN×P=BatchNorm(σ(xpN×P))+xpN×P2x_p^{N \times P} = \text{BatchNorm}(\sigma(x_p^{N \times P})) + x_p^{N \times P^2}

    1. Pointwise Convolution: Cross-patch feature aggregation is performed using a convolution with group size g=1g = 1, kernel size k=1k = 1, and stride s=1s = 1:

    xpN×P=ConvN→N(xpN×P,k=1,s=1,g=1)x_p^{N \times P} = \text{Conv}_{N \to N}(x_p^{N \times P}, k = 1, s = 1, g = 1)

    xpN×P=BatchNorm(σ(xpN×P))x_p^{N \times P} = \text{BatchNorm}(\sigma(x_p^{N \times P}))

    1. MLP Flatten Projection: The aggregated patch representations are flattened and passed through a two-layer MLP with intermediate hidden dimension doubling and a GELU activation:

    x^nonlin(i)=Linear(σ(Linear(Flatten(xpN×P))))∈RT\hat{x}_{\text{nonlin}}^{(i)} = \text{Linear}(\sigma(\text{Linear}(\text{Flatten}(x_p^{N \times P})))) \in \mathbb{R}^T

  5. Knowl 5 — Arctangent Horizon-Decayed Loss Function

    equation

    The arctangent loss Larctan\mathcal{L}_{\text{arctan}} scales the Mean Absolute Error (MAE) across future prediction timesteps i∈{1,…,T}i \in \{1, \dots, T\}:

    Larctan=1T∑i=1Tρarctan(i)∥x^1:T(i)−x1:T(i)∥1\mathcal{L}_{\text{arctan}} = \frac{1}{T} \sum_{i=1}^T \rho_{\text{arctan}}(i) \left\| \hat{x}_{1:T}^{(i)} - x_{1:T}^{(i)} \right\|_1

    where x^1:T(i)\hat{x}_{1:T}^{(i)} is the model prediction vector, x1:T(i)x_{1:T}^{(i)} is the ground truth vector, and the timestep scaling coefficient ρarctan(i)\rho_{\text{arctan}}(i) is defined as:

    ρarctan(i)=−arctan⁡(i)+π4+1\rho_{\text{arctan}}(i) = -\arctan(i) + \frac{\pi}{4} + 1

    This loss is a specific formulation of the generalized horizon-weighted MAE loss L=1T∑i=1Tρ(i)∥x^1:T(i)−x1:T(i)∥1\mathcal{L} = \frac{1}{T} \sum_{i=1}^T \rho(i) \| \hat{x}_{1:T}^{(i)} - x_{1:T}^{(i)} \|_1. The arctangent formulation decreases more slowly than power-law decay schedules (such as ρ(i)=i−1/2\rho(i) = i^{-1/2}), providing bounded attenuation of far-future prediction errors while preserving sufficient gradient signal at long horizons.

  6. Knowl 6 — Sigmoid Learning Rate Adjustment Schedule

    equation

    The sigmoid learning rate adjustment schedule defines the learning rate αt\alpha_t at epoch tt as the difference between two logistic functions:

    αt=α01+e−k(t−w)−α01+e−ks(t−sw)\alpha_t = \frac{\alpha_0}{1 + e^{-k(t - w)}} - \frac{\alpha_0}{1 + e^{-\frac{k}{s}(t - sw)}}

    where:

    • α0\alpha_0 is the initial base learning rate,
    • tt is the current training epoch,
    • ww is the warm-up coefficient controlling the transition point,
    • kk is the logistic growth rate controlling warm-up steepness,
    • ss is the curve smoothing rate controlling the duration and slope of learning rate decay.

    This formulation produces a continuous warm-up phase followed by a smooth, monotonic decay curve across training epochs.

  7. Knowl 7 — Forecasting Benchmark Performance under Unified Lookback Settings

    data/table

    Long-term forecasting performance is evaluated across nine benchmark datasets using unified lookback windows (L=36L = 36 for the ILI dataset, and L=96L = 96 for all others). Results are averaged across four prediction horizons: T∈{24,36,48,60}T \in \{24, 36, 48, 60\} for ILI, and T∈{96,192,336,720}T \in \{96, 192, 336, 720\} for all other datasets. Evaluation metrics are Mean Squared Error (MSE) and Mean Absolute Error (MAE).

    Models xPatch CARD TimeMixer iTransformer PatchTST
    Metric MSE MAE MSE MAE MSE MAE MSE MAE MSE MAE
    ETTh1 0.428 0.419 0.442 0.429 0.447 0.440 0.454 0.448 0.450 0.441
    ETTh2 0.319 0.361 0.368 0.390 0.365 0.395 0.383 0.407 0.365 0.394
    ETTm1 0.377 0.384 0.382 0.383 0.381 0.396 0.407 0.410 0.383 0.394
    ETTm2 0.267 0.313 0.272 0.317 0.275 0.323 0.288 0.332 0.284 0.327
    Weather 0.232 0.261 0.239 0.265 0.240 0.272 0.258 0.278 0.257 0.280
    Traffic 0.499 0.279 0.453 0.282 0.485 0.298 0.428 0.282 0.467 0.292
    Electricity 0.179 0.264 0.168 0.258 0.182 0.273 0.178 0.270 0.190 0.275
    Exchange 0.375 0.408 0.360 0.402 0.408 0.422 0.360 0.403 0.364 0.400
    Solar 0.239 0.236 0.237 0.239 0.216 0.280 0.233 0.262 0.254 0.289
    ILI 1.442 0.725 1.916 0.842 1.708 0.820 2.918 1.154 1.626 0.804

    Under these unified settings, xPatch achieves the best averaged performance on 60% of datasets in MSE and 70% in MAE. Relative to CARD, xPatch reduces MSE by 2.46% and MAE by 2.34%; relative to TimeMixer, by 3.34% MSE and 6.34% MAE; and relative to PatchTST, by 4.76% MSE and 6.20% MAE.

  8. Knowl 8 — Forecasting Benchmark Performance under Optimal Lookback Hyperparameter Search

    data/table

    Long-term forecasting performance is evaluated when each baseline is optimized over lookback input lengths LL to identify upper-bound performance. Results are averaged over four prediction horizons (T∈{24,36,48,60}T \in \{24, 36, 48, 60\} for ILI, and T∈{96,192,336,720}T \in \{96, 192, 336, 720\} for all other datasets).

    Models xPatch CARD TimeMixer PatchTST DLinear
    Metric MSE MAE MSE MAE MSE MAE MSE MAE MSE MAE
    ETTh1 0.391 0.412 0.401 0.422 0.411 0.423 0.413 0.434 0.423 0.437
    ETTh2 0.299 0.351 0.321 0.373 0.316 0.384 0.331 0.381 0.431 0.447
    ETTm1 0.341 0.368 0.350 0.368 0.348 0.376 0.353 0.382 0.357 0.379
    ETTm2 0.242 0.300 0.255 0.310 0.256 0.316 0.256 0.317 0.267 0.332
    Weather 0.211 0.247 0.220 0.248 0.222 0.262 0.226 0.264 0.246 0.300
    Traffic 0.392 0.248 0.381 0.251 0.388 0.263 0.391 0.264 0.434 0.295
    Electricity 0.153 0.245 0.157 0.251 0.156 0.247 0.159 0.253 0.166 0.264
    Exchange 0.366 0.404 0.360 0.402 0.471 0.452 0.405 0.426 0.297 0.378
    Solar 0.194 0.214 0.198 0.225 0.192 0.244 0.256 0.298 0.329 0.400
    ILI 1.281 0.688 1.916 0.842 1.971 0.924 1.480 0.807 2.169 1.041

    Under hyperparameter search, xPatch attains the lowest error on 70% of datasets for MSE and 90% for MAE. Averaged across datasets, xPatch outperforms CARD by 5.29% in MSE and 3.81% in MAE, TimeMixer by 7.45% in MSE and 7.85% in MAE, and PatchTST by 7.87% in MSE and 8.59% in MAE.

  9. Knowl 9 — Per-Step Training and Inference Runtime Comparison

    data/table

    Computational cost measured by average per-step running time and inference time across benchmarks on a single Quadro RTX 6000 GPU under identical evaluation settings:

    Method Training time (msec) Inference time (msec)
    MLP-stream 0.948 0.540
    CNN-stream 1.811 0.963
    xPatch 3.099 1.303
    DLinear 0.420 0.310
    iTransformer 6.290 2.743
    PatchTST 6.618 2.917
    TimeMixer 13.174 8.848
    CARD 14.877 7.162

    While the combined dual-flow network in xPatch requires more time than single-stream CNN or MLP blocks, its total training time (3.099 ms) and inference time (1.303 ms) are less than half that of transformer architectures such as PatchTST (6.618 ms / 2.917 ms), iTransformer (6.290 ms / 2.743 ms), and CARD (14.877 ms / 7.162 ms), as well as the multiscale MLP architecture TimeMixer (13.174 ms / 8.848 ms).

Coverage note — Detailed per-prediction-length metric breakdowns across all individual horizons ($T \in \{96, 192, 336, 720\}$), mathematical proofs, and specific hyperparameter tuning tables located in the appendix were omitted as the primary benchmark tables, structural equations, and architectural knowls encapsulate the paper's core contributions.

References

  1. 1.Bahdanau, D.; Cho, K.; and Bengio, Y. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations.
  2. 2.Box, G. E.; Jenkins, G. M.; Reinsel, G. C.; and Ljung, G. M. 2015. Time series analysis: forecasting and control. John Wiley & Sons.
  3. 3.Dickey, D. A.; and Fuller, W. A. 1979. Distribution of the estimators for autoregressive time series with a unit root. Journal of the American statistical association, 74(366a): 427–431.
  4. 4.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR.
  5. 5.Gardner Jr, E. S. 1985. Exponential smoothing: The state of the art. Journal of forecasting, 4(1): 1–28.
  6. 6.Han, L.; Ye, H.-J.; and Zhan, D.-C. 2023. The Capacity and Robustness Trade-off: Revisiting the Channel Independent Strategy for Multivariate Time Series Forecasting. arXiv preprint arXiv:2304.05206.
  7. 7.Hendrycks, D.; and Gimpel, K. 2016. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415.
  8. 8.Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.
  9. 9.Ioffe, S.; and Szegedy, C. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, 448–456. pmlr.
  10. 10.Kim, T.; Kim, J.; Tae, Y.; Park, C.; Choi, J.-H.; and Choo, J. 2021. Reversible instance normalization for accurate timeseries forecasting against distribution shift. In International Conference on Learning Representations.
  11. 11.Lai, G.; Chang, W.-C.; Yang, Y.; and Liu, H. 2018. Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, 95–104.
  12. 12.Li, S.; Jin, X.; Xuan, Y.; Zhou, X.; Chen, W.; Wang, Y.-X.; and Yan, X. 2019. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. Advances in neural information processing systems, 32.
  13. 13.Li, Z.; Qi, S.; Li, Y.; and Xu, Z. 2023. Revisiting Longterm Time Series Forecasting: An Investigation on Linear Mapping. arXiv preprint arXiv:2305.10721.
  14. 14.Liu, Y.; Hu, T.; Zhang, H.; Wu, H.; Wang, S.; Ma, L.; and Long, M. 2024. iTransformer: Inverted Transformers Are Effective for Time Series Forecasting. In The Twelfth International Conference on Learning Representations.
  15. 15.Liu, Y.; Wu, H.; Wang, J.; and Long, M. 2022. Nonstationary transformers: Exploring the stationarity in time series forecasting. Advances in Neural Information Processing Systems, 35: 9881–9893.
  16. 16.Nie, Y.; H. Nguyen, N.; Sinthong, P.; and Kalagnanam, J. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In International Conference on Learning Representations.
  17. 17.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  18. 18.Trockman, A.; and Kolter, J. Z. 2022. Patches Are All You Need? Trans. Mach. Learn. Res.
  19. 19.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  20. 20.Wang, H.; Peng, J.; Huang, F.; Wang, J.; Chen, J.; and Xiao, Y. 2023. Micn: Multi-scale local and global context modeling for long-term series forecasting. In The eleventh international conference on learning representations.
  21. 21.Wang, S.; Wu, H.; Shi, X.; Hu, T.; Luo, H.; Ma, L.; Zhang, J. Y.; and Zhou, J. 2024a. Timemixer: Decomposable multiscale mixing for time series forecasting. arXiv preprint arXiv:2405.14616.
  22. 22.Wang, X.; Zhou, T.; Wen, Q.; Gao, J.; Ding, B.; and Jin, R. 2024b. CARD: Channel Aligned Robust Blend Transformer for Time Series Forecasting. In The Twelfth International Conference on Learning Representations.
  23. 23.Wen, Q.; Zhou, T.; Zhang, C.; Chen, W.; Ma, Z.; Yan, J.; and Sun, L. 2022. Transformers in time series: A survey. arXiv preprint arXiv:2202.07125.
  24. 24.Woo, G.; Liu, C.; Sahoo, D.; Kumar, A.; and Hoi, S. 2022. Etsformer: Exponential smoothing transformers for timeseries forecasting. arXiv preprint arXiv:2202.01381.
  25. 25.Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; and Long, M. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In International Conference on Learning Representations.
  26. 26.Wu, H.; Xu, J.; Wang, J.; and Long, M. 2021. Autoformer: Decomposition transformers with auto-correlation for longterm series forecasting. Advances in Neural Information Processing Systems, 34: 22419–22430.
  27. 27.Zeng, A.; Chen, M.; Zhang, L.; and Xu, Q. 2023. Are transformers effective for time series forecasting? In Proceedings of the AAAI conference on artificial intelligence, volume 37, 11121–11128.
  28. 28.Zhang, Y.; and Yan, J. 2022. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In The Eleventh International Conference on Learning Representations.
  29. 29.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, volume 35, 11106–11115.
  30. 30.Zhou, T.; Ma, Z.; Wen, Q.; Wang, X.; Sun, L.; and Jin, R. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International Conference on Machine Learning, 27268–27286. PMLR.

Citation

MLA
Stitsyuk, A., and J. Choi. “xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition”. arXiv, 2024, http://arxiv.org/abs/2412.17323v3.
APA
Stitsyuk, A., & Choi, J. (2024). xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition. arXiv. http://arxiv.org/abs/2412.17323v3
Chicago
Stitsyuk, A., and J. Choi. 2024. “xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition”. arXiv. http://arxiv.org/abs/2412.17323v3.
Harvard
Stitsyuk, A. and Choi, J. (2024) “xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2412.17323v3.
Vancouver
1. Stitsyuk A, Choi J (2024) xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition. arXiv

BibTeX

@article{stitsyuk2024xpatch,
  title = {xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition},
  author = {Stitsyuk, Artyom and Choi, Jaesik},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2412.17323v3},
  eprint = {2412.17323}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF