CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series Forecasting
Hui HeQi ZhangSimeng BaiKun YiZhendong Niu
Proposes an end-to-end framework that constructs hierarchical tree structures to capture grouped correlations across variables and uses a cross-attention mechanism to jointly model dynamic inter-series and intra-series temporal dependencies for multivariate time series forecasting.
Modern decision-making in critical domains such as energy management, traffic control, disease tracking, and finance relies heavily on forecasting complex systems containing hundreds or thousands of interrelated time series. Existing predictive models generally struggle to scale effectively because they either focus solely on individual sequence trends or rely on predefined, rigid network structures that cannot easily capture the natural hierarchical groupings and dynamic relationships across different data streams. Addressing this gap is increasingly urgent as organizations face growing volumes of high-dimensional sensor and operational data.
The article introduces and evaluates the Cross Attentive Tree-aware Network (CATN), a new end-to-end deep learning framework designed to improve multi-step forecasting accuracy. The primary objective is to demonstrate that explicitly modeling both the hierarchical relationships among variables and multi-level temporal dependencies across time steps significantly outperforms current state-of-the-art forecasting techniques.
To evaluate this framework, the authors conducted empirical evaluations on four real-world benchmark datasets spanning transportation networks and power grids, including San Francisco highway traffic (963 series), electricity consumption (370 clients), and Los Angeles freeway sensors (207 and 228 series). The non-technical approach combines an automated hierarchical tree construction—grouping related time series based on a median distance criterion—with a multi-level temporal learning engine. This architecture integrates local pattern scanning, long-term sequence tracking via bidirectional recurrent networks, and a cross attention mechanism to dynamically evaluate interactions between different grouped features over time without requiring predefined spatial maps.
The experimental findings show significant performance gains. Across standard error metrics, the proposed model outperformed eight leading baseline methods, including advanced graph neural networks and Transformer-based architectures. On the traffic and electricity benchmarks, the model reduced mean absolute error by 11.7% and mean absolute percentage error by 39.03% compared to the strongest prior baseline. On specialized traffic flow datasets, it achieved over 30% improvement across percentage error metrics. Furthermore, ablation analyses revealed that removing the tree construction component led to an 18.03% worsening in root mean square error, confirming that hierarchical grouping is the most critical driver of performance.
These findings indicate that organizations managing complex networked assets can achieve substantially higher forecasting accuracy without spending resources to manually map physical or spatial network topologies beforehand. The tree structure also enhances interpretability, allowing operators to visually inspect and validate how different assets or measurement points relate to one another. In operational terms, more reliable forecasts translate into reduced grid instability risks, better traffic management, and optimized resource scheduling.
Organizations operating high-dimensional time series systems should consider adopting tree-based correlation learning as a scalable alternative to traditional graph neural networks. Before full-scale deployment, engineering teams should conduct parameter tuning pilots, particularly around embedding dimensions (where an intermediate size of 32 proved optimal) and input window lengths. While the methodology demonstrated robust gains across multiple benchmarks, performance can degrade if representations are aggregated too deeply in the tree hierarchy due to over-smoothing. Future implementations should account for these architectural trade-offs when adapting the framework to broader domains such as financial markets or weather forecasting.
- Paper: Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks, Zonghan Wu et al. (2020). Introduces graph neural network frameworks for multivariate time series that learn latent inter-variable relationships without predefined graph structures, establishing the foundational problem setup CATN addresses with hierarchical trees.
- Paper: Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting, Lei Bai et al. (2020). Demonstrates how to adaptively discover spatial relationships across time series channels without predefined road topology, motivating CATN's automated tree construction.
- Paper: Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks, Guokun Lai et al. (2017). Pioneers the combination of convolutional local pattern extraction and recurrent mechanisms for multivariate time series forecasting used in CATN's temporal learning engine.
- Paper: A Dual-Stage Attention-Based Recurrent Neural Network for Time Series Prediction, Yao Qin et al. (2017). Establishes dual-stage attention mechanisms across input features and temporal steps, laying the groundwork for cross-attentive multi-level dependency modeling.
- Paper: Spatio-temporal Graph Convolutional Neural Network: A Deep Learning Framework for Traffic Forecasting, Bing Yu et al. (2017). Provides the baseline spatio-temporal neural framework for traffic and sensor forecasting that CATN generalizes away from rigid physical graph structures.
- Paper: Time-series forecasting with deep learning: a survey, Bryan Lim et al. (2020). Surveys foundational deep learning architectures and hybrid attention-recurrent designs for multi-horizon multivariate time series forecasting.
- Paper: iTransformer: Inverted Transformers Are Effective for Time Series Forecasting, Yong Liu et al. (2023). Proposes an inverted Transformer that embeds individual variates into tokens, offering an alternative architectural paradigm for capturing cross-variable interactions in high-dimensional multivariate forecasting.
- Paper: A Time Series is Worth 64 Words: Long-term Forecasting with Transformers, Yuqi Nie et al. (2023). Extends long-term multivariate forecasting by exploring patch-level temporal tokenization and channel independence compared to hierarchical cross-variable attention.
- Paper: Are Transformers Effective for Time Series Forecasting?, Ailing Zeng et al. (2023). Critically benchmarks complex attention-based multivariate forecasters against simple linear baselines to evaluate where architectural complexity genuinely yields performance gains.
- Paper: TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis, Haixu Wu et al. (2023). Generalizes multi-period temporal dependency modeling into 2D tensor spaces across diverse time series tasks beyond hierarchical tree aggregation.
- Paper: N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting, Cristian Challu et al. (2023). Explores hierarchical multi-rate sampling and interpolation for long-horizon time series forecasting as a complementary multi-scale representation method.
- Paper: Unified Training of Universal Time Series Forecasting Transformers, Gerald Woo et al. (2024). Scales universal multi-variate time series forecasting across domains and heterogeneous dimensions through unified pre-trained foundation models.
