LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting
Xu LiuYutong XiaYuxuan LiangJunfeng HuYiwei WangLei BaiChao HuangZhenguang LiuBryan HooiRoger Zimmermann
Presents LargeST, a large-scale traffic forecasting benchmark covering 8,600 sensors over five years in California alongside rich metadata, providing a realistic testbed to evaluate the scalability and long-term predictive accuracy of spatio-temporal deep learning models.
Accurate road traffic forecasting is essential for modern urban planning, traffic management, and public safety initiatives. While recent advances in deep learning—particularly spatial-temporal graph neural networks—have shown strong predictive performance, existing public benchmarks suffer from severe limitations. Standard datasets typically cover only a few hundred sensors, span less than six months of data, and omit critical contextual metadata. These deficiencies restrict the ability to evaluate model scalability across realistic road networks, hinder the analysis of multi-year seasonal patterns, and leave models unable to leverage physical road attributes.
The article introduces LargeST, a comprehensive, large-scale traffic forecasting benchmark dataset designed to evaluate the accuracy, computational efficiency, and scalability of deep learning models under real-world conditions.
To construct LargeST, the authors gathered five continuous years of traffic flow readings (2017 to 2021) at five-minute intervals from the California Department of Transportation Performance Measurement System. The dataset encompasses 8,600 highway mainline sensors organized into a statewide dataset and three regional sub-datasets covering Greater Los Angeles, the Greater Bay Area, and San Diego. In total, the dataset contains over 4.5 billion data points and integrates detailed node metadata, including sensor coordinates, highway categories, travel directions, and lane counts. The authors evaluated twelve representative baseline models, ranging from standard time-series methods to complex spatial-temporal graph neural networks, measuring forecasting accuracy across multiple time horizons alongside training runtime and memory consumption.
The evaluation yielded several critical findings regarding model viability on large-scale infrastructure. First, existing advanced forecasting models face severe scalability bottlenecks: half of the tested deep learning models failed to execute on the full 8,600-sensor dataset due to out-of-memory errors on high-end hardware with 48 gigabytes of memory. Second, older and structurally simpler models using temporal convolutions and adaptive graph learning, such as Graph WaveNet and AGCRN, achieved competitive predictive accuracy while maintaining practical training speeds. Third, highly complex architectures incorporating dynamic spatial graphs demonstrated strong accuracy on smaller regional subsets but suffered from prohibitive computational costs and poor scaling. Finally, exploratory data analysis confirmed that traffic volume strongly correlates with highway classifications and lane configurations, while also exhibiting notable seasonal and multi-year shifts across different regions.
These findings demonstrate that high predictive accuracy on small, legacy benchmark datasets does not translate directly to practical deployment at city or regional scales. Transportation authorities and engineering teams should avoid adopting overly complex graph architectures that cannot scale within realistic hardware constraints. Instead, development efforts should prioritize computationally efficient, simple yet robust models, while actively integrating physical road metadata to improve interpretability and performance.
The dataset's primary limitation is its geographic focus on California highways, which may constrain its generalizability to urban street grids or international traffic systems, along with the presence of typical real-world sensor noise and missing values. Nevertheless, the empirical findings provide high confidence that current state-of-the-art modeling practices require substantial re-engineering toward efficiency before large-scale operational deployment is viable.
- Paper: Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting, Yaguang Li et al. (2017). This paper establishes the foundational spatio-temporal graph modeling framework and popularizes the California PeMS traffic forecasting benchmark setups that LargeST directly aims to scale up.
- Paper: Spatio-temporal Graph Convolutional Neural Network: A Deep Learning Framework for Traffic Forecasting, Bing Yu et al. (2017). It introduces STGCN, combining spectral graph convolutions with temporal convolutions, serving as a primary baseline model benchmarked on LargeST.
- Paper: Graph WaveNet for Deep Spatial-Temporal Graph Modeling, Zonghan Wu et al. (2019). It introduces Graph WaveNet with adaptive adjacency matrices and dilated convolutions, providing one of the core spatio-temporal baselines evaluated in LargeST.
- Paper: Adaptive Graph Convolutional Recurrent Network for Traffic Forecasting, Lei Bai et al. (2020). It develops the AGCRN architecture for automated spatial dependency discovery, which represents another key modern baseline tested on the LargeST benchmark.
- Paper: Attention Based Spatial-Temporal Graph Convolutional Networks for Traffic Flow Forecasting, Shengnan Guo et al. (2019). It establishes the ASTGCN framework on PeMS sensor networks, exemplifying the standard small-scale benchmarks and attention-based baselines LargeST expands upon.
- Paper: GMAN: A Graph Multi-Attention Network for Traffic Prediction, Chuanpan Zheng et al. (2019). It introduces the GMAN attention network for multi-step traffic prediction on PeMS data, serving as a baseline for spatio-temporal forecasting under benchmark conditions.
- Paper: Spatial-Temporal Synchronous Graph Convolutional Networks: A New Framework for Spatial-Temporal Network Data Forecasting, Chao Song et al. (2020). It proposes STSGCN for capturing synchronous spatial-temporal correlations on California PeMS datasets, providing methodological context for traffic forecasting baselines.
- Paper: Graph Neural Network for Traffic Forecasting: A Survey, Weiwei Jiang et al. (2021). This comprehensive survey categorizes graph neural network architectures and standard traffic datasets, contextualizing the historical limitations in dataset scale and duration that LargeST resolves.
- Paper: Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks, Zonghan Wu et al. (2020). It introduces MTGNN for multivariate time series forecasting with adaptive graph learning, evaluated as a foundational structure-learning baseline in traffic benchmarking.
- Paper: T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction, Ling Zhao et al. (2018). It presents the T-GCN model uniting graph convolutions with recurrent units, offering standard baseline methodology for road-network traffic speed prediction.
No sufficiently relevant recommendations were found.
