CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting

Hui HeQi ZhangSimeng BaiKun YiZhendong Niu

article2022AAAI60 citations

Proposes an end-to-end framework that constructs hierarchical tree structures to capture grouped correlations across variables and uses a cross-attention mechanism to jointly model dynamic inter-series and intra-series temporal dependencies for multivariate time series forecasting.

Listen

Modern decision-making in critical domains such as energy management, traffic control, disease tracking, and finance relies heavily on forecasting complex systems containing hundreds or thousands of interrelated time series. Existing predictive models generally struggle to scale effectively because they either focus solely on individual sequence trends or rely on predefined, rigid network structures that cannot easily capture the natural hierarchical groupings and dynamic relationships across different data streams. Addressing this gap is increasingly urgent as organizations face growing volumes of high-dimensional sensor and operational data.

The article introduces and evaluates the Cross Attentive Tree-aware Network (CATN), a new end-to-end deep learning framework designed to improve multi-step forecasting accuracy. The primary objective is to demonstrate that explicitly modeling both the hierarchical relationships among variables and multi-level temporal dependencies across time steps significantly outperforms current state-of-the-art forecasting techniques.

To evaluate this framework, the authors conducted empirical evaluations on four real-world benchmark datasets spanning transportation networks and power grids, including San Francisco highway traffic (963 series), electricity consumption (370 clients), and Los Angeles freeway sensors (207 and 228 series). The non-technical approach combines an automated hierarchical tree construction—grouping related time series based on a median distance criterion—with a multi-level temporal learning engine. This architecture integrates local pattern scanning, long-term sequence tracking via bidirectional recurrent networks, and a cross attention mechanism to dynamically evaluate interactions between different grouped features over time without requiring predefined spatial maps.

The experimental findings show significant performance gains. Across standard error metrics, the proposed model outperformed eight leading baseline methods, including advanced graph neural networks and Transformer-based architectures. On the traffic and electricity benchmarks, the model reduced mean absolute error by 11.7% and mean absolute percentage error by 39.03% compared to the strongest prior baseline. On specialized traffic flow datasets, it achieved over 30% improvement across percentage error metrics. Furthermore, ablation analyses revealed that removing the tree construction component led to an 18.03% worsening in root mean square error, confirming that hierarchical grouping is the most critical driver of performance.

These findings indicate that organizations managing complex networked assets can achieve substantially higher forecasting accuracy without spending resources to manually map physical or spatial network topologies beforehand. The tree structure also enhances interpretability, allowing operators to visually inspect and validate how different assets or measurement points relate to one another. In operational terms, more reliable forecasts translate into reduced grid instability risks, better traffic management, and optimized resource scheduling.

Organizations operating high-dimensional time series systems should consider adopting tree-based correlation learning as a scalable alternative to traditional graph neural networks. Before full-scale deployment, engineering teams should conduct parameter tuning pilots, particularly around embedding dimensions (where an intermediate size of 32 proved optimal) and input window lengths. While the methodology demonstrated robust gains across multiple benchmarks, performance can degrade if representations are aggregated too deeply in the tree hierarchy due to over-smoothing. Future implementations should account for these architectural trade-offs when adapting the framework to broader domains such as financial markets or weather forecasting.

Cover for CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting

Abstract

Modeling complex hierarchical and grouped feature interaction in the multivariate time series data is indispensable to comprehending the data dynamics and predicting the future condition. The implicit feature interaction and high-dimensional data make multivariate forecasting very challenging. Many existing works did not put more emphasis on exploring explicit correlation among multiple time-series data, and complicated models are designed to capture long- and short-range patterns with the aid of attention mechanisms. In this work, we think that a pre-defined graph or a general learning method is difficult due to its irregular structure. Hence, we present CATN, an end-to-end model of Cross Attentive Tree-aware Network to jointly capture the inter-series correlation and intra-series temporal patterns. We first construct a tree structure to learn hierarchical and grouped correlation and design an embedding approach that can pass a dynamic message to generalize implicit but interpretable cross features among multiple time series. Next in the temporal aspect, we propose a multi-level dependency learning mechanism including global&local learning and cross attention mechanism, which can combine long-range dependencies, short-range dependencies as well as cross dependencies at different time steps. The extensive experiments on different datasets from real-world show the effectiveness and robustness of the method we proposed when compared with existing state-of-the-art methods.

Table of Contents

  • Introduction
  • Related Work
  • Correlated Time Series Forecasting
  • Cross Attention Mechanism
  • Preliminaries
  • Methodology
  • Tree Construction
  • Tree Embedding
  • Global & Local Learning
  • Cross Attention Mechanism
  • Learning Objective
  • Experiments
  • Datasets, Metrics and Baselines
  • Overall Comparison
  • Parameter Sensitivity
  • Ablation Study
  • Analysis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — CATN architecture for hierarchical multivariate forecasting

    model/method

    CATN is an end-to-end forecasting model that jointly learns inter-series hierarchical correlations and intra-series temporal dependencies without requiring a predefined graph. Given a multivariate input window X∈RTx×dxX\in\mathbb{R}^{T_x\times d_x}, where TxT_x is the number of observed time steps and dxd_x is the number of series, CATN first organizes the series as leaves of an ordered hierarchical tree. Trainable node, edge, and time embeddings propagate information upward through the tree to create implicit but structurally interpretable cross-series features.

    The resulting tree representations are processed by a global-and-local learning module: temporal convolutions capture short-range patterns, while bidirectional LSTMs capture longer-range dependencies. A cross-attention module then compares representations of different non-leaf nodes at the same tree level and combines their mutually relevant temporal information. The final features are projected to multi-step forecasts and trained jointly through the entire tree, temporal, and attention pipeline. The architecture diagram on page 4 depicts this sequence of tree construction, tree embedding, convolutional/recurrent processing, cross attention, and forecasting.

  2. Knowl 2 — Median-linkage hierarchical tree construction

    algorithm

    CATN constructs an agglomerative tree whose leaves are the input time series and whose internal nodes represent groups of series. Initially, every series is a separate cluster; at each agglomerative step, the two clusters with the smallest inter-cluster dissimilarity are merged.

    For feature vectors zz and pp with dimension weights wiw_i, CATN uses a weighted Euclidean dissimilarity, represented explicitly as

    d(z,p)=∑i=1mwi(zi−pi)2,d(z,p)=\sum_{i=1}^{m}w_i(z_i-p_i)^2,

    where mm is the feature dimension and the paper defines the dimension weights by wi=zi/(1n∑j=1nzj)w_i=z_i/(\frac{1}{n}\sum_{j=1}^{n}z_j). For two clusters ZZ and PP, containing ∣Z∣|Z| and ∣P∣|P| points, respectively, there are N=∣Z∣∣P∣N=|Z||P| cross-cluster pairwise distances. After sorting these distances as m(1)≤⋯≤m(N)m_{(1)}\leq\cdots\leq m_{(N)}, CATN selects the pairs whose ranks lie between α(N)\alpha(N) and β(N)\beta(N):

    Sα(N),β(N)(Z,P)={(z,p):z∈Z,p∈P, m(α(N))≤d(z,p)≤m(β(N))}.\mathcal{S}_{\alpha(N),\beta(N)}(Z,P)=\{(z,p):z\in Z,p\in P,\ m_{(\alpha(N))}\leq d(z,p)\leq m_{(\beta(N))}\}.

    The median-linkage distance is the average distance over this selected subset,

    m-Lα,β(Z,P)=1∣Sα(N),β(N)(Z,P)∣∑(z,p)∈Sα(N),β(N)(Z,P)d(z,p),m\text{-}L_{\alpha,\beta}(Z,P)=\frac{1}{|\mathcal{S}_{\alpha(N),\beta(N)}(Z,P)|}\sum_{(z,p)\in\mathcal{S}_{\alpha(N),\beta(N)}(Z,P)}d(z,p),

    with α(N)=1\alpha(N)=1 and β(N)=N/2\beta(N)=N/2 in the proposed construction. This criterion uses a subset of pairwise distances instead of only the closest, farthest, or all pairs, providing the ordered tree used to represent grouped and hierarchical feature interactions.

  3. Knowl 3 — Tree and time embedding with dynamic upward messages

    model/method

    For a tree T=(V,E)\mathcal{T}=(V,E), CATN partitions nodes into leaf nodes VLV_L and non-leaf nodes VIV_I. Each leaf series receives a trainable dense embedding ui∈Rdu_i\in\mathbb{R}^{d}, where dd is the embedding dimension. Hierarchical and agnostic time information, including day, week, month, year, A.M./P.M., and holiday indicators, is mapped to trainable time embeddings and concatenated with the node embedding.

    Each edge is initially represented as either a left-child or right-child relation and is then mapped to a trainable dense vector ele_l or ere_r. If vi∈VIv_i\in V_I has left and right child embeddings ulu_l and uru_r, its initial internal-node representation is

    vi=ϕ(ul)⋅el+ϕ(ur)⋅er,v_i=\phi(u_l)\cdot e_l+\phi(u_r)\cdot e_r,

    where ϕ(⋅)\phi(\cdot) fills missing values before aggregation. Thus, internal nodes encode grouped higher-order features while preserving the left/right structure of the tree.

    CATN further normalizes the contributions from the left and right children with softmax-style coefficients and updates each internal node after prediction. These dynamically updated node representations are selected from an interaction level and arranged as tensors Ed∈Rw×dE_d\in\mathbb{R}^{w\times d}, where ww is the input time-window length. The tree therefore supplies both an interpretable hierarchy and trainable cross-series messages that are revised during forecasting.

  4. Knowl 4 — Joint global-and-local temporal dependency learning

    model/method

    For each selected tree node representation Ed∈Rw×dE_d\in\mathbb{R}^{w\times d}, CATN combines a convolutional component with a recurrent component. The convolutional component uses ncn_c filters of width wcw_c and length lc=dl_c=d, applies no pooling, and averages their rectified outputs:

    oc=1nc∑k=1ncReLU⁡(Wk∗Ed+bk),o_c=\frac{1}{n_c}\sum_{k=1}^{n_c}\operatorname{ReLU}(W_k*E_d+b_k),

    where WkW_k and bkb_k are the parameters of filter kk and ∗* denotes temporal convolution. After zero-padding, the convolutional output is concatenated with the tree representation, producing a w×1w\times 1 local-feature component. This component is intended to capture short-range temporal patterns.

    In parallel, CATN feeds the representation of each node at tree level kk into its own bidirectional LSTM. The forward and backward hidden sequences are concatenated; each direction uses hidden size d/2d/2, so the combined recurrent representation has dimension dd. The recurrent output captures dependencies across the time window in both temporal directions and is concatenated with the local convolutional features before cross attention. The page-4 architecture diagram shows these convolutional and recurrent streams operating on the tree-derived node representations.

  5. Knowl 5 — Cross attention between tree-level node sequences

    model/method

    At a given tree level containing knk_n non-leaf nodes, CATN forms every unordered pair of node hidden sequences. The contrast set is C={ha}a=1knC=\{h_a\}_{a=1}^{k_n}, the query set contains the corresponding hbh_b with b>ab>a, and the pair set has (kn2)\binom{k_n}{2} elements. For a pair ha,hb∈Rw×dh_a,h_b\in\mathbb{R}^{w\times d}, let cic_i and qjq_j denote the feature vectors at positions ii and jj of the two sequences. CATN computes a staggered-position correlation map using cosine similarity:

    Rij=(ci∥ci∥2) ⁣T(qj∥qj∥2),i,j∈{1,…,d},R_{ij}=\left(\frac{c_i}{\|c_i\|_2}\right)^{\!T}\left(\frac{q_j}{\|q_j\|_2}\right),\qquad i,j\in\{1,\ldots,d\},

    assuming the compared vectors have nonzero norms. The rows and columns of RR form correlation maps for the contrast and query sequences. For a contrast correlation vector ricr_i^c, a scalar score and normalized attention weight are computed as

    ωic=f(WTric+b),αic=exp⁡(ωic)∑j=1dexp⁡(ωjc),\omega_i^c=f(W^Tr_i^c+b),\qquad \alpha_i^c=\frac{\exp(\omega_i^c)}{\sum_{j=1}^{d}\exp(\omega_j^c)},

    where W∈Rd×1W\in\mathbb{R}^{d\times 1}, bb is a scalar offset, and ff aggregates the correlations at the position. An analogous map αq\alpha^q is computed for the query sequence.

    The attention maps are reshaped and applied through residual weighting: the original features are multiplied elementwise by 1+αc1+\alpha^c and 1+αq1+\alpha^q. The enhanced representations from all node pairs are then combined using another attention-weighted average. This mechanism lets one node selectively reread temporally relevant information from another node rather than extracting each node representation independently.

  6. Knowl 6 — Multi-step mean-squared-error training objective

    equation

    CATN trains all forecasting steps jointly with mean squared error. Let Y[:,τ]Y[:,\tau] be the ground-truth values at forecast index τ\tau, let Y^[:,τ]\widehat{Y}[:,\tau] be the corresponding CATN prediction, and let the forecast interval be τ=ti+1: ti+h\tau=t_i+1:\,t_i+h, where hh is the prediction horizon and TyT_y is the number of forecast steps. The objective is

    LCATN=1Ty∑τ=1Ty(Y[:,τ]−Y^[:,τ])2.\mathcal{L}_{\mathrm{CATN}}=\frac{1}{T_y}\sum_{\tau=1}^{T_y}\left(Y[:,\tau]-\widehat{Y}[:,\tau]\right)^2.

    The loss is backpropagated from the multi-step LSTM outputs through the cross-attention, global-and-local temporal, tree-embedding, and tree-construction components.

  7. Knowl 7 — Real-world evaluation protocol

    experimental setup

    CATN was evaluated on four real-world multivariate time-series datasets. Traffic contains 963 highway occupancy series and 10,560 time points from California, sampled every 10 minutes, with values between 0 and 1. Electricity contains 370 client-consumption series and 25,968 time points sampled every 15 minutes, with values in kW. PeMSD7(M) contains 228 freeway-detector series and 11,232 time points sampled every 5 minutes. METR-LA contains 207 Los Angeles highway-sensor series and 34,272 time points sampled every 5 minutes.

    The reported evaluation metrics are mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), symmetric mean absolute percentage error (SMAPE), and weighted absolute percentage error (WAPE); lower values indicate better forecasting. Forecast horizons are 3, 6, 9, and 12 steps. The comparison methods are VAR, LSTNet, TPA-LSTM, MTNet, DSANet, MTGNN, DeepGLO, and Informer, covering statistical, recurrent, convolutional, self-attention, graph-neural, and Transformer-based forecasting approaches.

  8. Knowl 8 — CATN forecasting performance across four datasets

    data/table

    CATN generally achieved the strongest or most stable results across forecasting horizons. The primary comparison on Traffic and Electricity reported the following CATN values; the values are ordered by horizons 3,6,9,123,6,9,12. The paper reports improvements of 11.7% in MAE, 3.62% in RMSE, and 39.03% in MAPE over the corresponding existing best results.

    Dataset Metric h=3 h=6 h=9 h=12
    Traffic MAE 0.018 0.0183 0.0193 0.0178
    Traffic RMSE 0.0309 0.0312 0.0319 0.0308
    Traffic MAPE 0.156 0.1589 0.1682 0.155
    Electricity MAE 0.061 0.0613 0.061 0.0626
    Electricity RMSE 0.2601 0.2605 0.2621 0.2649
    Electricity MAPE 0.1907 0.1921 0.1931 0.2017

    For the additional traffic datasets, the page-6 results table reports WAPE/MAPE/SMAPE triples for each horizon as follows:

    Model Dataset h=3 h=6 h=9 h=12
    DeepGLO PeMSD7(M) 0.0818/0.1024/0.0988 0.1378/0.1635/0.1651 0.0554/0.063/0.0592 0.102/0.1506/0.1245
    DeepGLO METR-LA 0.1506/0.2157/0.196 0.0806/0.068/0.07 0.0917/0.0755/0.0697 0.091/0.0749/0.0695
    Informer PeMSD7(M) 0.1371/0.1808/0.1543 0.1404/0.191/0.1587 0.1685/0.2208/0.186 0.1421/0.1972/0.1606
    Informer METR-LA 0.137/0.2794/0.1693 0.1333/0.2682/0.165 0.1488/0.2798/0.1796 0.148/0.2877/0.1795
    MTGNN PeMSD7(M) 0.1771/0.2891/0.2542 0.0654/0.0782/0.0698 0.0801/0.0988/0.0836 0.0886/0.1129/0.0927
    MTGNN METR-LA 0.0972/0.1028/0.1109 0.1012/0.1231/0.1276 0.1304/0.1465/0.1478 0.1504/0.191/0.1681
    CATN PeMSD7(M) 0.0444/0.0486/0.0462 0.0439/0.0481/0.0457 0.0439/0.048/0.0457 0.0458/0.05/0.0476
    CATN METR-LA 0.0516/0.0626/0.0556 0.0527/0.0639/0.0568 0.053/0.0641/0.057 0.054/0.0604/0.0576

    Within each triple the order is WAPE/MAPE/SMAPE. CATN is lower than all three listed baselines at every reported horizon on both PeMSD7(M) and METR-LA, with the paper describing the improvement as greater than 30% for these metrics. The authors attribute the stronger and more stable multi-horizon performance to combining tree-based inter-series correlation with global, local, and cross temporal dependencies.

  9. Knowl 9 — Ablation evidence for tree embedding, global-local learning, and cross attention

    data/table

    An ablation study was conducted on Electricity with batch size 64, input window 12, and forecast horizon 3. The full CATN model was compared with removing tree construction and embedding (w/o TEC), replacing global-and-local learning with a purely recurrent module (w/o GLL), and removing cross attention (w/o CAM).

    Metric CATN w/o TEC w/o GLL w/o CAM
    MAE 0.061 0.0649 0.0613 0.0615
    RMSE 0.2601 0.3173 0.2603 0.2603
    MAPE 0.1907 0.216 0.191 0.1934
    SMAPE 0.174 0.1966 0.1755 0.1767

    Removing tree construction and embedding causes the largest degradation, especially in RMSE, supporting the claim that hierarchical and grouped inter-series correlation is important. Removing the convolutional global-and-local component slightly worsens all metrics, indicating that local patterns complement recurrent long-range modeling. Removing cross attention also worsens every metric, indicating that explicitly emphasizing relevant information between node representations produces more discriminative forecasting features.

  10. Knowl 10 — Parameter sensitivity and qualitative hierarchy diagnostics

    empirical result

    CATN's sensitivity analysis on Electricity found that the model was strongest with a short one-hour input window among the tested settings, while MTGNN obtained a lower MAE when the input window was extended to 2.5 hours. This was interpreted as longer histories supplying more information for learning node relationships, although CATN remained effective with less historical input.

    The tested embedding dimensions were 1616, 3232, and 6464. Dimension 3232 produced the best performance; both smaller and larger dimensions degraded results. The authors attribute this to a trade-off between representing more hierarchical/grouped information and increasing the burden of upward message aggregation. For the interaction depth, selecting all nodes from the third layer for global-and-local learning performed worse, which the paper attributes to over-smoothing and loss of the original information distribution. The fourth and fifth layers gave similar results, while deeper layers approximately doubled the global-and-local parameter count.

    The qualitative case study visualized on page 7 used five METR-LA sensors. Sensors 3 and 4 shared a parent in the learned tree, whereas sensors 21, 15, and 28 were at the same level but belonged to different parents. The learned correlation matrix assigned high correlation to sensors 3 and 14; the authors related this to their nearby road locations. This case study supports the intended interpretability of the learned hierarchy, although it is qualitative rather than a separate quantitative metric.

Coverage note — No substantial contributed material was omitted; the qualitative METR-LA sensor case study is included with the parameter and hierarchy diagnostics rather than as a separate knowl.

References

  1. 1.Bai, L.; Yao, L.; Li, C.; Wang, X.; and Wang, C. 2020. Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting. In NeurIPS.
  2. 2.Boroujeni, F. R.; Wang, S.; Li, Z.; West, N.; Stantic, B.; Yao, L.; and Long, G. 2018. Trace Ratio Optimization With Feature Correlation Mining for Multiclass Discriminant Analysis. In AAAI, 2746–2753. AAAI Press.
  3. 3.Cao, D.; Wang, Y.; Duan, J.; Zhang, C.; Zhu, X.; Huang, C.; Tong, Y.; Xu, B.; Bai, J.; Tong, J.; and Zhang, Q. 2020. Spectral Temporal Graph Neural Network for Multivariate Time-series Forecasting. In NeurIPS.
  4. 4.Chang, Y.; Sun, F.; Wu, Y.; and Lin, S. 2018. A Memory Network Based Solution for Multivariate Time-Series Forecasting. CoRR, abs/1809.02105.
  5. 5.Cheng, Z.; Yuan, C.; Li, J.; and Yang, H. 2018. TreeNet: Learning Sentence Representations with Unconstrained Tree Structure. In IJCAI, 4005–4011. ijcai.org.
  6. 6.Du, S.; Li, T.; Yang, Y.; and Horng, S. 2020. Multivariate time series forecasting via attention-based encoder-decoder framework. Neurocomputing, 388: 269–279.
  7. 7.Du, Y.; Wang, J.; Feng, W.; Pan, S.; Qin, T.; Xu, R.; and Wang, C. 2021. AdaRNN: Adaptive Learning and Forecasting of Time Series. CoRR, abs/2108.04443.
  8. 8.Emmendorfer, L. R.; and de Paula Canuto, A. M. 2021. A generalized average linkage criterion for Hierarchical Agglomerative Clustering. Appl. Soft Comput., 100: 106990.
  9. 9.Faloutsos, C.; Flunkert, V.; Gasthaus, J.; Januschowski, T.; and Wang, Y. 2019. Forecasting Big Time Series: Theory and Practice. In KDD, 3209–3210. ACM.
  10. 10.Gao, R.; Duru, O.; and Yuen, K. F. 2021. High-dimensional lag structure optimization of fuzzy time series. Expert Syst. Appl., 173: 114698.
  11. 11.Hao, Y.; Zhang, Y.; Liu, K.; He, S.; Liu, Z.; Wu, H.; and Zhao, J. 2017. An End-to-End Model for Question Answering over Knowledge Base with Cross-Attention Combining Global Knowledge. In ACL (1), 221–231. Association for Computational Linguistics.
  12. 12.Hou, R.; Chang, H.; Ma, B.; Shan, S.; and Chen, X. 2019. Cross Attention Network for Few-shot Classification. In NeurIPS, 4005–4016.
  13. 13.Hu, L.; Jian, S.; Cao, L.; and Chen, Q. 2018. Interpretable Recommendation via Attraction Modeling: Learning Multilevel Attractiveness over Multimodal Movie Contents. In IJCAI, 3400–3406. ijcai.org.
  14. 14.Huang, S.; Wang, D.; Wu, X.; and Tang, A. 2019. DSANet: Dual Self-Attention Network for Multivariate Time Series Forecasting. In CIKM, 2129–2132. ACM.
  15. 15.Lai, G.; Chang, W.; Yang, Y.; and Liu, H. 2018. Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks. In SIGIR, 95–104. ACM.
  16. 16.Lee, J.; Jain, M.; Park, H.; and Yun, S. 2021. Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization. In ICLR. OpenReview.net.
  17. 17.Li, Y.; Yu, R.; Shahabi, C.; and Liu, Y. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In ICLR (Poster). OpenReview.net.
  18. 18.Lin, W.; Deng, Y.; Gao, Y.; Wang, N.; Zhou, J.; Liu, L.; Zhang, L.; and Wang, P. 2021. CAT: Cross-Attention Transformer for One-Shot Object Detection. CoRR, abs/2104.14984.
  19. 19.Lyu, Y.; Li, M.; Huang, X.; Guler, U.; Schaumont, P.; and Zhang, Z. 2020. TreeRNN: Topology-Preserving Deep Graph Embedding and Learning. In ICPR, 7493–7499. IEEE.
  20. 20.Mancuso, P.; Piccialli, V.; and Sudoso, A. M. 2020. A machine learning approach for forecasting hierarchical time series. CoRR, abs/2006.00630.
  21. 21.Maxim, L. G.; Rodriguez, J. I.; Wang, B.; and . 2020. Defect of Euclidean distance degree. Adv. Appl. Math., 121: 102101.
  22. 22.Rangapuram, S. S.; Werner, L. D.; Benidis, K.; Mercado, P.; Gasthaus, J.; and Januschowski, T. 2021. End-to-End Learning of Coherent Probabilistic Forecasts for Hierarchical Time Series. In ICML, volume 139, 8832–8843. PMLR.
  23. 23.Salinas, D.; Bohlke-Schneider, M.; Callot, L.; Medico, R.; and Gasthaus, J. 2019. High-dimensional multivariate forecasting with low-rank Gaussian Copula Processes. In NeurIPS, 6824–6834.
  24. 24.Sen, R.; Yu, H.; and Dhillon, I. S. 2019. Think Globally, Act Locally: A Deep Neural Network Approach to High-Dimensional Time Series Forecasting. In NeurIPS, 4838–4847.
  25. 25.Shi, K.; Wang, Y.; Lu, H.; Zhu, Y.; and Niu, Z. 2021. EKGTF: A knowledge-enhanced model for optimizing social network-based meteorological briefings. Inf. Process. Manag., 58(4): 102564.
  26. 26.Shih, S.; Sun, F.; and Lee, H. 2019. Temporal pattern attention for multivariate time series forecasting. Mach. Learn., 108(8-9): 1421–1441.
  27. 27.Wan, S.; and Niu, Z. 2020. A Hybrid E-Learning Recommendation Approach Based on Learners’ Influence Propagation. IEEE Trans. Knowl. Data Eng., 32(5): 827–840.
  28. 28.Wang, S.; Hu, L.; Cao, L.; Huang, X.; Lian, D.; and Liu, W. 2018a. Attention-Based Transactional Context Embedding for Next-Item Recommendation. In AAAI, 2532–2539. AAAI Press.
  29. 29.Wang, X.; He, X.; Feng, F.; Nie, L.; and Chua, T. 2018b. TEM: Tree-enhanced Embedding Model for Explainable Recommendation. In WWW, 1543–1552. ACM.
  30. 30.Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Chang, X.; and Zhang, C. 2020. Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks. In KDD, 753–763. ACM.
  31. 31.Xu, D.; Cheng, W.; Zong, B.; Song, D.; Ni, J.; Yu, W.; Liu, Y.; Chen, H.; and Zhang, X. 2020. Tensorized LSTM with Adaptive Shared Memory for Learning Trends in Multivariate Time Series. In AAAI, 1395–1402. AAAI Press.
  32. 32.Yang, Z.; Yan, W. W.; Huang, X.; and Mei, L. 2020. Adaptive Temporal-Frequency Network for Time-Series Forecasting. IEEE Trans. Knowl. Data Eng., PP(99): 1–1.
  33. 33.Yu, B.; Yin, H.; and Zhu, Z. 2018. Spatio-Temporal Graph Convolutional Networks: A Deep Learning Framework for Traffic Forecasting. In IJCAI, 3634–3640.
  34. 34.Zhang, Q.; Cao, L.; Shi, C.; and Niu, Z. 2021. Neural Time-Aware Sequential Recommendation by Jointly Modeling Preference Dynamics and Explicit Feature Couplings. IEEE Transactions on Neural Networks and Learning Systems, 1–13.
  35. 35.Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; and Zhang, W. 2021. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In AAAI, 11106–11115. AAAI Press.
  36. 36.Zhu, Y.; Lin, Q.; Lu, H.; Shi, K.; Liu, D.; Chambua, J.; Wan, S.; and Niu, Z. 2021. Recommending Learning Objects through Attentive Heterogeneous Graph Convolution and Operation-Aware Neural Network. IEEE Transactions on Knowledge and Data Engineering.
  37. 37.Zhuang, D. E. H.; Li, G. C. L.; and Wong, A. K. C. 2014. Discovery of Temporal Associations in Multivariate Time Series. IEEE Trans. Knowl. Data Eng., 26(12): 2969–2982.
  38. 38.Zhuang, Z.; Yeh, C. M.; Wang, L.; Zhang, W.; and Wang, J. 2020. Multi-stream RNN for Merchant Transaction Prediction. CoRR, abs/2008.01670.

Citation

MLA
He, H., et al. “CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 4, 2022, pp. 4030–38, https://doi.org/10.1609/AAAI.V36I4.20320.
APA
He, H., Zhang, Q., Bai, S., Yi, K., & Niu, Z. (2022). CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence, 36(4), 4030–4038. https://doi.org/10.1609/AAAI.V36I4.20320
Chicago
He, H., Q. Zhang, S. Bai, K. Yi, and Z. Niu. 2022. “CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (4): 4030–38. https://doi.org/10.1609/AAAI.V36I4.20320.
Harvard
He, H. et al. (2022) “CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(4), pp. 4030–4038. Available at: https://doi.org/10.1609/AAAI.V36I4.20320.
Vancouver
1. He H, Zhang Q, Bai S, Yi K, Niu Z (2022) CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 36:4030–4038

BibTeX

@article{He_2022, title={CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I4.20320}, DOI={10.1609/aaai.v36i4.20320}, number={4}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={He, Hui and Zhang, Qi and Bai, Simeng and Yi, Kun and Niu, Zhendong}, year={2022}, month=June, pages={4030–4038} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF