Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction

Xiaolei MaZhuang DaiZhengbing HeJihui NaYong WangYunpeng Wang

article2017Italian National Conference on Sensors1,319 citations

Demonstrates that modeling spatiotemporal traffic dynamics as two-dimensional images enables convolutional neural networks to predict large-scale network speeds with a 42.9% accuracy improvement over traditional and recurrent deep learning models.

Listen

Urban traffic congestion poses severe operational, economic, and planning challenges for modern metropolitan areas. Effective network-wide traffic management and navigation require accurate forecasts across entire road networks rather than isolated roadway corridors. However, existing statistical and shallow machine learning models often treat traffic variables as isolated time series. Consequently, they fail to capture the complex, two-dimensional spatial correlations and broader propagation patterns across interconnected roadways, and they struggle with the computational burden of large-scale road networks.

The article demonstrates an image-based deep learning framework using a Convolutional Neural Network—a deep learning model specialized in processing visual grid structures—to predict network-wide traffic speeds with high accuracy across large-scale urban transportation networks. It evaluates how effectively converting spatiotemporal traffic dynamics into two-dimensional image matrices enables the automated extraction of deep traffic features for short- and longer-term forecasting.

To evaluate the framework, the analysis used real-world probe data from approximately 10,000 GPS-equipped taxis in Beijing over a 37-day period in 2015, aggregated into two-minute intervals. The evaluation tested two sub-networks of distinct topological complexity: the simple, circular Second Ring Road (236 road segments) and a complex grid in Northeast Beijing (352 road segments). The approach structured continuous time on the horizontal axis and ordered road segments on the vertical axis to create single-channel speed matrices. The proposed four-layer convolutional architecture was evaluated across four distinct prediction horizons (10-minute and 20-minute forecasts using 30 or 40 minutes of prior data) and benchmarked against four prevailing statistical algorithms and three conventional deep learning sequence models.

The findings confirm that the image-based convolutional approach consistently outperformed all baseline methods across all test conditions. First, the proposed framework achieved an average prediction accuracy improvement of 42.91% compared to prevailing models on testing datasets. Second, across three categorized speed regimes (heavy, moderate, and free-flow traffic), the model attained the highest overall accuracy scores, averaging 0.931, whereas traditional regression scored 0.917 and conventional artificial neural networks scored significantly lower. Third, deeper convolutional architectures reduced test error substantially, with a four-layer design outperforming shallower configurations. Fourth, the model maintained superior accuracy even when forecasting longer 20-minute horizons, successfully capturing congestion propagation patterns across complex multi-road networks.

These results demonstrate that capturing spatial and temporal interdependencies simultaneously is essential for reliable network-level congestion forecasting. Operationally, the model maintains a practical balance between computational efficiency and high precision: although simpler statistical models train faster, their inability to model network-wide spatial correlations yields inferior forecasts, while competitive ensemble tree models require excessive training times (around nine hours) that are impractical for large networks. Implementing this framework can enhance intelligent transportation systems, improve route-guidance reliability, and help municipal authorities optimize traffic operations.

For future implementation, transportation agencies and technology teams should explore hybrid architectures, such as combining convolutional feature extraction with recurrent sequence models like Long Short-Term Memory networks to further refine dynamic temporal forecasting. System planners should also conduct pilot deployments in live operational centers to assess real-time performance. Caution should be exercised when mapping non-linear, complex road grids onto two-dimensional matrix axes, as network segmentation can partially disrupt spatial continuity, although the convolutional layers compensate by assembling higher-level feature representations.

arXiv: 1701.04245
Cover for Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction

Abstract

This paper proposes a convolutional neural network (CNN)-based method that learns traffic as images and predicts large-scale, network-wide traffic speed with a high accuracy. Spatiotemporal traffic dynamics are converted to images describing the time and space relations of traffic flow via a two-dimensional time-space matrix. A CNN is applied to the image following two consecutive steps: abstract traffic feature extraction and network-wide traffic speed prediction. The effectiveness of the proposed method is evaluated by taking two real-world transportation networks, the second ring road and north-east transportation network in Beijing, as examples, and comparing the method with four prevailing algorithms, namely, ordinary least squares, k-nearest neighbors, artificial neural network, and random forest, and three deep learning architectures, namely, stacked autoencoder, recurrent neural network, and long-short-term memory network. The results show that the proposed method outperforms other algorithms by an average accuracy improvement of 42.91% within an acceptable execution time. The CNN can train the model in a reasonable time and, thus, is suitable for large-scale transportation networks.

Table of Contents

  • 1. Introduction
  • 2. Methods
  • 2.1. Converting Network Traffic to Images
  • 2.2. CNN for Network Traffic Prediction
  • 2.2.1. CNN Characteristics
  • 2.2.2. CNN Characteristics
  • 2.2.3. Convolutional Layers and Pooling Layers of the CNN
  • 2.2.4. CNN Optimization
  • 3. Empirical Study
  • 3.1. Data Description
  • 3.2. Time-Space Image Generation
  • 3.3. Tuning Up CNN Parameters
  • 3.4. Results and Comparison
  • 4. Conclusion
  • References

Knowls

  1. Knowl 1 — Spatiotemporal Traffic-to-Image Matrix Representation

    model/method

    To capture network-wide traffic dynamics within a deep learning framework, historical traffic speed observations are mapped onto a two-dimensional time-space matrix that functions as a single-channel image. The horizontal axis (xx-axis) represents discrete time intervals j∈{1,2,…,N}j \in \{1, 2, \dots, N\}, and the vertical axis (yy-axis) represents sequentially ordered road segments i∈{1,2,…,Q}i \in \{1, 2, \dots, Q\} in the transportation network.

    The time-space matrix M∈RQ×NM \in \mathbb{R}^{Q \times N} is formulated as:

    M=[m11m12⋯m1Nm21m22⋯m2N⋮⋮⋱⋮mQ1mQ2⋯mQN]M = \begin{bmatrix} m_{11} & m_{12} & \cdots & m_{1N} \\ m_{21} & m_{22} & \cdots & m_{2N} \\ \vdots & \vdots & \ddots & \vdots \\ m_{Q1} & m_{Q2} & \cdots & m_{QN} \end{bmatrix}

    where mijm_{ij} denotes the average traffic speed on road section ii at time interval jj. For loop-like topologies (such as ring roads), road sections are ordered along the ring path. For complex planar networks containing crossroads and two-way roads, the network is segmented into straight lines of ordered sections. Pixel intensities are normalized to the range [0,1][0, 1] relative to the maximum speed limit before being input to neural network layers.

  2. Knowl 2 — Convolutional Neural Network Architecture for Network-Wide Speed Prediction

    model/method

    Network-wide traffic speed prediction is performed using a deep convolutional neural network (CNN) operating on the spatiotemporal traffic image. The architecture consists of four consecutive stages:

    1. Model Input: A single-channel matrix xi=[mi,mi+1,…,mi+F−1]∈R1×Q×Fx^i = [m_i, m_{i+1}, \dots, m_{i+F-1}] \in \mathbb{R}^{1 \times Q \times F}, representing normalized historical traffic speeds of all QQ road sections across FF historical time intervals for sample index ii.

    2. Spatiotemporal Feature Extraction: Multiple stacked layers of 2D convolutions and 2D max-pooling with Rectified Linear Unit (ReLU) activations automatically extract localized spatial dependencies among adjacent road sections and temporal trends across consecutive time steps.

    3. Feature Flattening: The output feature maps from the final convolutional-pooling layer are unrolled into a single high-dimensional dense feature vector.

    4. Model Output: A fully connected layer maps the flattened feature vector to the output vector y^∈RQ⋅P\hat{y} \in \mathbb{R}^{Q \cdot P}, which contains the simultaneous continuous speed predictions for all QQ road segments across PP future time intervals.

  3. Knowl 3 — Mathematical Formulation of 2D Convolution, Max-Pooling, and Fully Connected Layers for Traffic Prediction

    equation

    Let l∈{1,…,L}l \in \{1, \dots, L\} denote the layer index of the CNN, and let clc_l denote the number of convolutional channels in layer ll. For an input matrix subregion dd and a convolutional filter WljW_l^j of size m×nm \times n, the 2D convolution output at position (e,f)(e, f) is given by:

    yconv=∑e=1m∑f=1n(Wl)ef⋅defy_{\text{conv}} = \sum_{e=1}^m \sum_{f=1}^n (W_l)_{ef} \cdot d_{ef}

    Using the Rectified Linear Unit (ReLU) activation function σ(x)=max⁡(0,x)\sigma(x) = \max(0, x), the output oljo_l^j of the jj-th feature map (j∈[1,cl]j \in [1, c_l]) after convolution and pooling over cl−1c_{l-1} input channels xlkx_l^k is:

    olj=pool(σ(∑k=1cl−1Wlj,kxlk+blj))o_l^j = \text{pool}\left(\sigma\left(\sum_{k=1}^{c_{l-1}} W_l^{j,k} x_l^k + b_l^j\right)\right)

    where pool(⋅)\text{pool}(\cdot) denotes 2D max-pooling over a sliding window of size p×qp \times q:

    ypool=max⁡e∈[1,p],f∈[1,q](def)y_{\text{pool}} = \max_{e \in [1, p], f \in [1, q]} (d_{ef})

    After LL convolutional-pooling layers, the feature maps are concatenated into a flattened 1D representation:

    oLflatten=flatten([oL1,oL2,…,oLcL])o_L^{\text{flatten}} = \text{flatten}\left([o_L^1, o_L^2, \dots, o_L^{c_L}]\right)

    The final predicted network-wide traffic speed vector y^\hat{y} is generated through a fully connected layer with weight matrix WfW_f and bias vector bfb_f:

    y^=WfoLflatten+bf=Wf⋅flatten(pool(σ(∑k=1cL−1WLj,kxLk+bLj)))+bf\hat{y} = W_f o_L^{\text{flatten}} + b_f = W_f \cdot \text{flatten}\left(\text{pool}\left(\sigma\left(\sum_{k=1}^{c_{L-1}} W_L^{j,k} x_L^k + b_L^j\right)\right)\right) + b_f

  4. Knowl 4 — CNN Training Objective and Optimization for Traffic Speed Prediction

    equation

    The traffic prediction CNN is trained using the Mean Squared Error (MSE) loss function between the network-predicted speed vector y^\hat{y} and the ground-truth speed vector yy over nn training samples:

    MSE=1n∑i=1n∥y^i−yi∥2\text{MSE} = \frac{1}{n} \sum_{i=1}^n \|\hat{y}_i - y_i\|^2

    Let Θ={Wlj,blj,Wf,bf}\Theta = \{W_l^j, b_l^j, W_f, b_f\} denote the complete set of learnable filter weights and bias parameters across all LL convolutional layers and the final fully connected layer. The optimal parameter set Θ∗\Theta^* is determined via gradient backpropagation:

    Θ∗=arg⁡min⁡Θ1n∑i=1n∥Wf⋅flatten(pool(σ(∑k=1cL−1WLj,kxL,ik+bLj)))+bf−yi∥2\Theta^* = \arg\min_{\Theta} \frac{1}{n} \sum_{i=1}^n \left\| W_f \cdot \text{flatten}\left(\text{pool}\left(\sigma\left(\sum_{k=1}^{c_{L-1}} W_L^{j,k} x_{L,i}^k + b_L^j\right)\right)\right) + b_f - y_i \right\|^2

    To prevent model overfitting, training employs an early stopping criterion that halts gradient updates when the validation loss fails to decrease over a designated number of successive training epochs.

  5. Knowl 5 — Depth Selection and Hyperparameter Tuning of the Traffic CNN

    empirical result

    The structural hyperparameters of the CNN were configured with 3×33 \times 3 convolutional filters and 2×22 \times 2 max-pooling windows. Network depth was evaluated by predicting 10-minute future traffic speeds from 40-minute historical observations on the Beijing Second Ring Road (21,600 training samples from 30 days, 5,040 test samples from 7 days):

    • Depth-1 (Input directly connected to fully connected output): Training MSE ≈158.0\approx 158.0, Test MSE ≈155.0\approx 155.0
    • Depth-2 (64 conv filters →\to fully connected): Training MSE ≈24.5\approx 24.5, Test MSE ≈42.0\approx 42.0
    • Depth-3 (128 conv →\to 64 conv →\to fully connected): Training MSE ≈22.0\approx 22.0, Test MSE ≈37.0\approx 37.0
    • Depth-4 (256 conv →\to 128 conv →\to 64 conv →\to fully connected): Training MSE =21.3= 21.3, Test MSE =35.5= 35.5

    Increasing the depth from 1 to 4 progressively improved feature extraction capability and significantly lowered test MSE; Depth-4 achieved the lowest error on both training and test datasets.

  6. Knowl 6 — Layer Specification and Parameter Scale of the Depth-4 Traffic CNN

    data/table

    The Depth-4 CNN architecture was configured for Beijing Network 1 (236 road sections, 20 historical 2-minute input intervals, predicting 5 future 2-minute intervals across all sections, yielding a 1,180-dimensional output vector). The layer-by-layer dimensional transitions and parameter scales are given below:

    Layer Name Parameters Dimensions Parameter Scale
    Input — — (1,236,20)(1, 236, 20) —
    Layer 1 Convolution Filter (256,3,3)(256, 3, 3) (256,236,20)(256, 236, 20) 2,304
    Layer 1 Pooling Pooling (2,2)(2, 2) (256,118,10)(256, 118, 10) 0
    Layer 2 Convolution Filter (128,3,3)(128, 3, 3) (128,118,10)(128, 118, 10) 1,152
    Layer 2 Pooling Pooling (2,2)(2, 2) (128,59,5)(128, 59, 5) 0
    Layer 3 Convolution Filter (64,3,3)(64, 3, 3) (64,59,5)(64, 59, 5) 576
    Layer 3 Pooling Pooling (2,2)(2, 2) (64,30,3)(64, 30, 3) 0
    Layer 4 Data flatten — (5760,)(5760,) 0
    Layer 4 Fully-connected — (1180,)(1180,) 6,796,800
    Output — — (1180,)(1180,) —

    The convolutional and pooling layers compress spatial-temporal dimensionality while contributing only 4,032 parameters, whereas the dense fully connected output layer comprises the majority of parameters (6,796,8006,796,800).

  7. Knowl 7 — Experimental Setup on Beijing Transportation Networks

    experimental setup

    Experiments used probe data collected from approximately 10,000 GPS-equipped taxis in Beijing over 37 consecutive days (1 May 2015 to 6 June 2015; missing data <2.9%< 2.9\%, imputed via spatiotemporal adjacency). Vehicle speeds were aggregated into 2-minute intervals. The first 30 days (21,600 samples) formed the training set, and the remaining 7 days (5,040 samples) formed the test set.

    Two distinct network topologies were tested:

    1. Network 1 (Second Ring Road): 236 road sections, all one-way ring-road links.
    2. Network 2 (Northeast Beijing Network): 352 road sections comprising two-way roads and crossroads.

    Four prediction tasks were evaluated on both networks:

    • Task 1: 10-min prediction (5 future steps) using 30-min input (15 historical steps).
    • Task 2: 10-min prediction (5 future steps) using 40-min input (20 historical steps).
    • Task 3: 20-min prediction (10 future steps) using 30-min input (15 historical steps).
    • Task 4: 20-min prediction (10 future steps) using 40-min input (20 historical steps).
  8. Knowl 8 — Comparative Traffic Speed Prediction Performance (MSE)

    data/table

    The Depth-4 CNN was evaluated against four statistical baselines (Ordinary Least Squares [OLS], kk-Nearest Neighbors [k=10k=10], Random Forest [RF, 10 trees], Artificial Neural Network [ANN, 3 hidden layers of 1000 units]) and three deep learning baselines (Stacked Autoencoder [SAE, hidden layers 3000-2500-2000], Recurrent Neural Network [RNN, 3 hidden layers of 1000 units], Long Short-Term Memory Network [LSTM NN, 3 hidden layers of 1000 units]). Test set Mean Squared Error (MSE) results across all four tasks and both networks are reported below:

    Study Network Model MSE of Different Models (on Test Datasets)
    Task 1 Task 2 Task 3 Task 4
    Network 1 CNN 22.825 24.345 30.593 31.424
    OLS 27.047 31.273 41.334 48.107
    KNN 51.700 55.708 60.256 64.132
    RF 35.092 35.431 40.476 40.638
    ANN 67.764 52.339 58.797 57.225
    SAE 60.751 69.082 65.292 68.326
    RNN 33.408 36.833 40.551 39.038
    LSTM NN 37.759 33.218 42.909 42.865
    Network 2 CNN 27.163 28.479 37.987 38.816
    OLS 33.741 41.657 50.123 62.282
    KNN 69.965 74.863 79.367 83.881
    RF 48.603 48.946 52.676 53.067
    ANN 124.937 147.489 133.299 168.136
    SAE 85.079 94.982 82.271 99.020
    RNN 48.877 47.470 52.577 52.114
    LSTM NN 43.304 45.657 50.928 48.345

    The CNN achieved the lowest MSE across all prediction horizons and network configurations, demonstrating that 2D convolutional modeling of joint spatial-temporal matrices outperforms sequence-only deep models and link-independent regressors.

  9. Knowl 9 — Categorical Traffic State Prediction Accuracy

    data/table

    Predicted continuous traffic speeds were converted into three discrete operational traffic states: heavy traffic (0–20 km/h0\text{--}20\text{ km/h}), moderate traffic (20–40 km/h20\text{--}40\text{ km/h}), and free-flow traffic (>40 km/h>40\text{ km/h}). The classification accuracy scores on test datasets are reported below:

    Study Network Model Accuracy Score of Different Models (on Test Datasets)
    Task 1 Task 2 Task 3 Task 4
    Network 1 CNN 0.939 0.942 0.925 0.928
    OLS 0.935 0.929 0.915 0.909
    KNN 0.901 0.897 0.893 0.890
    RF 0.917 0.917 0.910 0.910
    ANN 0.869 0.876 0.852 0.865
    SAE 0.867 0.870 0.866 0.866
    RNN 0.908 0.913 0.898 0.900
    LSTM NN 0.910 0.908 0.901 0.905
    Network 2 CNN 0.938 0.936 0.920 0.922
    OLS 0.929 0.920 0.907 0.897
    KNN 0.886 0.884 0.879 0.876
    RF 0.898 0.898 0.893 0.892
    ANN 0.794 0.867 0.823 0.832
    SAE 0.846 0.835 0.848 0.825
    RNN 0.901 0.900 0.896 0.896
    LSTM NN 0.903 0.907 0.901 0.895

    The CNN achieved the highest accuracy across all tasks and networks, with an overall average accuracy of 0.931 (representing an average accuracy promotion of 42.91% relative to the baseline error rates), followed by OLS (0.917) and RF (0.904).

  10. Knowl 10 — Model Training Time and Computational Efficiency Trade-offs

    empirical result

    Training execution times varied significantly across prediction models:

    1. Linear and Shallow Models (OLS, KNN, ANN): Required minimal training time (under 50 minutes), but produced substantially inferior prediction accuracy.
    2. Sequential Deep Learning Models (SAE, RNN, LSTM NN): Trained faster than the CNN (under 100 minutes) because they processed temporal sequences without 2D spatial convolution across the full network grid, but failed to capture spatial network topology correlations.
    3. Random Forest (RF): Required by far the largest training time, taking approximately 9 hours (540 to 890 minutes) on Network 1 and 950 to 1,600 minutes on Network 2, because an individual ensemble had to be trained separately for every road section.
    4. CNN: Trained within 70 to 100 minutes while outperforming all models in prediction accuracy on testing data, demonstrating practical computational tractability for large-scale urban road networks.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Zhang, J.; Wang, F.-Y.; Wang, K.; Lin, W.-H.; Xu, X.; Chen, C. Data-driven intelligent transportation systems: A survey. IEEE Trans. Intell. Transp. Syst. 2011, 12, 1624–1639.
  2. 2.Park, J.; Li, D.; Murphey, Y.L.; Kristinsson, J.; McGee, R.; Kuang, M.; Phillips, T. Real Time Vehicle Speed Prediction Using a Neural Network Traffic Model. In Proceedings of the International Joint Conference on Neural Networks, San Jose, CA, USA, 31 July–5 August 2011; pp. 2991–2996.
  3. 3.Karlaftis, M.G.; Vlahogianni, E.I. Statistical methods versus neural networks in transportation research: Differences, similarities and some insights. Transp. Res. Part C Emerg. Technol. 2011, 19, 387–399.
  4. 4.Davis, G.A.; Nihan, N.L. Nonparametric Regression and Short-Term Freeway Traffic Forecasting. J. Transp. Eng. 1991, 117, 178–188.
  5. 5.Chang, H.; Lee, Y.; Yoon, B.; Baek, S. Dynamic near-term traffic flow prediction: Systemoriented approach based on past experiences. Iet Intell. Transp. Syst. 2012, 6, 292–305.
  6. 6.Xia, D.; Wang, B.; Li, H.; Li, Y.; Zhang, Z. A distributed spatial–temporal weighted model on MapReduce for short-term traffic flow forecasting. Neurocomputing 2016, 179, 246–263.
  7. 7.Wu, C.H.; Ho, J.M.; Lee, D.T. Travel-time prediction with support vector regression. IEEE Trans. Intell. Transp. Syst. 2004, 5, 276–281.
  8. 8.Hong, W.C. Traffic flow forecasting by seasonal SVR with chaotic simulated annealing algorithm. Neurocomputing 2011, 74, 2096–2107.
  9. 9.Castro-Neto, M.; Jeong, Y.S.; Jeong, M.K.; Han, L.D. Online-SVR for short-term traffic flow prediction under typical and atypical traffic conditions. Expert Syst. Appl.2009, 36, 6164–6173.
  10. 10.Ma, Z.; Luo, G.; Huang, D. Short term traffic flow prediction based on on-line sequential extreme learning machine. In Proceedings of the 2016 Eighth International Conference on Advanced Computational Intelligence (ICACI), Amnat Charoen Chiang Mai, Thailand, 14–16 February 2016; pp. 143–149.
  11. 11.Asif, M.T.; Dauwels, J.; Goh, C.Y.; Oran, A.; Fathi, E.; Xu, M.; Dhanya, M.M.; Mitrovic, N.; Jaillet, P. Spatiotemporal patterns in large-scale traffic speed prediction. IEEE Trans. Intell. Transp. Syst. 2014, 15, 794–804.
  12. 12.Clark, S. Traffic Prediction Using Multivariate Nonparametric Regression. J. Transp. Eng. 2003, 129, 161–168.
  13. 13.Haworth, J.; Cheng, T. Non-parametric regression for space-time forecasting under missing data. Comput. Environ. Urban Syst. 2012, 36, 538–550.
  14. 14.Li, L.; He, S.; Zhang, J.; Ran, B. Short-term highway traffic flow prediction based on a hybrid strategy considering temporal-spatial information. J. Adv. Transp. 2017, doi:10.1002/atr.1443.
  15. 15.Zhu, Z.; Peng, B.; Xiong, C.; Zhang, L. Short-term traffic flow prediction with linear conditional Gaussian Bayesian network. J. Adv.Transp. 2016, 50, 1111–1123.
  16. 16.Li, Y.; Jiang, X.; Zhu, H.; He, X.; Peeta, S.; Zheng, T.; Li, Y. Multiple measures-based chaotic time series for traffic flow prediction based on Bayesian theory. Nonlinear Dyn. 2016, 85, 179–194.
  17. 17.Tran, Q.T.; Ma, Z.; Li, H.; Hao, L.; Trinh, Q.K. A Multiplicative Seasonal ARIMA/GARCH Model in EVN Traffic Prediction. Int. J. Commun. Netw. Syst.Sci. 2015, 8, 43–49.
  18. 18.Williams, B.M.; Hoel, L.A. Modeling and Forecasting Vehicular Traffic Flow as a Seasonal ARIMA Process: Theoretical Basis and Empirical Results. J. Transp. Eng. 2003, 129, 664–672.
  19. 19.Voort, M.V.D.; Dougherty, M.; Watson, S. Combining kohonen maps with arima time series models to forecast traffic flow. Transp. Res. Part C Emerg. Technol. 1996, 4, 307–318.
  20. 20.Williams, B. Multivariate Vehicular Traffic Flow Prediction: Evaluation of ARIMAX Modeling. Transp. Res. Rec. J. Transp.Res. Board 2001, 1776, 194–200.
  21. 21.Huang, S.H.; Ran, B. An Application of Neural Network on Traffic Speed Prediction Under Adverse Weather Condition. Transp. Res. Board Annu. Meet. 2003.
  22. 22.Zheng, W.; Lee, D.H.; Zheng, W.; Lee, D.H. Short-Term Freeway Traffic Flow Prediction: Bayesian Combined Neural Network Approach. J. Transp. Eng. 2006, 132, 114–121.
  23. 23.Moretti, F.; Pizzuti, S.; Panzieri, S.; Annunziato, M. Urban traffic flow forecasting through statistical and neural network bagging ensemble hybrid modeling. Neurocomputing 2015, 167, 3–7.
  24. 24.Polson, N.; Sokolov, V. Deep Learning Predictors for Traffic Flows. arXiv 2016, arXiv:1604.04527.
  25. 25.Huang, W.; Song, G.; Hong, H.; Xie, K. Deep Architecture for Traffic Flow Prediction: Deep Belief Networks With Multitask Learning. IEEE Trans. Intell. Transp. Syst. 2014, 15, 2191–2201.
  26. 26.Tan, H.; Xuan, X.; Wu, Y.; Zhong, Z.; Ran, B. A comparison of traffic flow prediction methods based on DBN. In Processedings of the 16th COTA International Conference of Transportation Professionals (CICTP), Shanghai, China, 6–9 July 2016; pp. 273–283.
  27. 27.Ma, X.; Yu, H.; Wang, Y.; Wang, Y. Large-scale transportation network congestion evolution prediction using deep learning theory. PLoS ONE 2015, 10, e0119044.
  28. 28.Lv, Y.; Duan, Y.; Kang, W.; Li, Z.; Wang, F.Y. Traffic Flow Prediction With Big Data: A Deep Learning Approach. IEEE Trans. Intell. Transp. Syst. 2014, 16, 1–9.
  29. 29.Duan, Y.; Lv, Y.; Kang, W.; Zhao, Y. A deep learning based approach for traffic data imputation. Intell. Transp. Syst. 2014, 912–917.
  30. 30.Ma, X.; Tao, Z.; Wang, Y.; Yu, H.; Wang, Y. Long short-term memory neural network for traffic speed prediction using remote microwave sensor data. Transp. Res. Part C Emerg. Technol. 2015, 54, 187–197.
  31. 31.Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet Classification with Deep Convolutional Neural Networks. Adv. Neural Inf. Process. Syst. 2012, 25, 2012.
  32. 32.Oquab, M.; Bottou, L.; Laptev, I.; Sivic, J. Learning and Transferring Mid-level Image Representations Using Convolutional Neural Networks. In Processedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 1717–1724.
  33. 33.Lawrence, S.; Giles, C.L.; Tsoi, A.C.; Back, A.D. Face recognition: A convolutional neural-network approach. IEEE Trans. Neural Netw. 1997, 8, 98–113.
  34. 34.Ji, S.; Yang, M.; Yu, K. 3D Convolutional Neural Networks for Human Action Recognition. IEEE Trans. Pattern Anal.Mach. Intell. 2013, 35, 221–231.
  35. 35.Karpathy, A.; Toderici, G.; Shetty, S.; Leung, T.; Sukthankar, R.; Li, F.F. Large-Scale Video Classification with Convolutional Neural Networks. In Processedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 1725–1732.
  36. 36.LeCun, Y.; Bengio, Y. Convolutional networks for images, speech, and time series. In The Handbook of Brain Theory and Neural Networks; MIT Press: Cambridge, MA, USA, 1998; Volume 3361, pp. 255–258.
  37. 37.Sch; Nhof, M.; Helbing, D. Empirical Features of Congested Traffic States and Their Implications for Traffic Modeling. Transp. Sci. 2007, 41, 135–166.
  38. 38.Lecun, Y.; Boser, B.; Denker, J.S.; Henderson, D.; Howard, R.E.; Hubbard, W.; Jackel, L.D. Backpropagation applied to handwritten zip code recognition. Neural Comput. 1869, 1, 541–551.
  39. 39.Lv, Y.; Duan, Y.; Kang, W.; Li, Z.; Wang, F.-Y. Traffic flow prediction with big data: A deep learning approach. IEEE Trans. Intell. Transp. Syst. 2015, 16, 865–873.
  40. 40.Sarle, W.S. Stopped Training and Other Remedies for Overfitting. In Proceedings of the 27th Symposium on the Interface of Computing Science and Statistics, Fairfax, VA, USA, 21–24 June 1995; pp. 352–360.
  41. 41.Kerner, B.S.; Rehborn, H.; Aleksic, M.; Haug, A. Recognition and tracking of spatial-temporal congested traffic patterns on freeways. Transp. Res. Part C Emerg. Technol. 2004, 12, 369–400.

Citation

MLA
Ma, X., et al. “Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction”. Sensors,2017, 2017, http://arxiv.org/abs/1701.04245v4.
APA
Ma, X., Dai, Z., He, Z., Na, J., Wang, Y., & Wang, Y. (2017). Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction. Sensors,2017. http://arxiv.org/abs/1701.04245v4
Chicago
Ma, X., Z. Dai, Z. He, J. Na, Y. Wang, and Y. Wang. 2017. “Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction”. Sensors,2017. http://arxiv.org/abs/1701.04245v4.
Harvard
Ma, X. et al. (2017) “Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction”, Sensors,2017 [Preprint]. Available at: http://arxiv.org/abs/1701.04245v4.
Vancouver
1. Ma X, Dai Z, He Z, Na J, Wang Y, Wang Y (2017) Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction. Sensors,2017

BibTeX

@article{ma2017learning,
  title = {Learning Traffic as Images: A Deep Convolutional Neural Network for Large-Scale Transportation Network Speed Prediction},
  author = {Ma, Xiaolei and Dai, Zhuang and He, Zhengbing and Na, Jihui and Wang, Yong and Wang, Yunpeng},
  year = {2017},
  journal = {Sensors,2017},
  url = {http://arxiv.org/abs/1701.04245v4},
  eprint = {1701.04245}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF