CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes

Yuhong LiXiaofan ZhangDeming Chen

article2018CVPR1,667 citations

Proposes a dilated convolutional network architecture that expands receptive fields without pooling-induced resolution loss, achieving state-of-the-art crowd counting accuracy and high-quality density maps across congested benchmarks.

Listen

Monitoring heavily congested scenes is essential for public safety, crowd control, and emergency prevention in high-risk environments such as major public assemblies and dense traffic corridors. Simple object counts are insufficient for risk management because identical crowd numbers can exhibit drastically different spatial distributions. While deep learning visual models have advanced crowd analysis, existing state-of-the-art systems rely on complex multi-branch network designs and pre-classification steps. These multi-column architectures suffer from severe structural redundancy, prolonged training times, and reduced spatial resolution in their output maps.

The article demonstrates an alternative approach called CSRNet, a pure single-column neural network designed to accurately estimate counts and generate high-resolution spatial density maps in congested environments. The approach utilizes a standard 10-layer visual feature extraction front-end combined with a back-end of dilated convolutional layers. By using dilated layers—which space out network filters—the model broadens its field of view without using pooling operations that discard fine spatial detail or requiring complex multi-branch structures. The authors evaluated the system across four standard crowd benchmark datasets (ShanghaiTech, UCF CC 50, WorldExpo’10, and UCSD) and extended the evaluation to vehicle counting using the TRANCOS traffic dataset.

The evaluation demonstrates that the streamlined single-column architecture consistently outperforms complex multi-branch models. On the ShanghaiTech Part B benchmark, CSRNet achieved a 47.3% lower mean absolute error than the previous leading method, while reducing error by 7.0% on the densely crowded Part A dataset and 10.0% on the UCF CC 50 benchmark. The model achieved the lowest average error across multi-scene surveillance benchmarks and generated higher-quality visual density maps with superior image similarity scores. When applied to vehicle counting in congested traffic, it delivered a 15.4% lower overall counting error than the leading baseline and reduced localized sub-region error by up to 67.7%.

These results show that network depth and expanded receptive fields are significantly more effective than complex, branching structures that attempt to classify crowd densities into discrete bins. For operational deployments, this single-column design lowers computational overhead, simplifies model training, and avoids allocating critical parameters to redundant internal classifiers. High visual quality in the generated density maps improves situational awareness for safety operators, enabling more reliable early warnings for dangerous crowd buildups and traffic gridlock.

Organizations developing or deploying automated surveillance analytics should replace multi-column crowd counting architectures with single-column dilated networks to improve accuracy and training efficiency. When deploying these models on hardware-constrained edge devices such as surveillance cameras, teams should leverage the architecture's flexible input resolution support and compact structure. While the findings provide high confidence across diverse crowd and traffic benchmarks, users should note that performance on extremely low-resolution surveillance footage requires preprocessing upsampling, and specialized camera angles may still require targeted fine-tuning.

Cover for CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes

Abstract

We propose a network for Congested Scene Recognition called CSRNet to provide a data-driven and deep learning method that can understand highly congested scenes and perform accurate count estimation as well as present high-quality density maps. The proposed CSRNet is composed of two major components: a convolutional neural network (CNN) as the front-end for 2D feature extraction and a dilated CNN for the back-end, which uses dilated kernels to deliver larger reception fields and to replace pooling operations. CSRNet is an easy-trained model because of its pure convolutional structure. We demonstrate CSRNet on four datasets (ShanghaiTech dataset, the UCF_CC_50 dataset, the WorldEXPO'10 dataset, and the UCSD dataset) and we deliver the state-of-the-art performance. In the ShanghaiTech Part_B dataset, CSRNet achieves 47.3% lower Mean Absolute Error (MAE) than the previous state-of-the-art method. We extend the targeted applications for counting other objects, such as the vehicle in TRANCOS dataset. Results show that CSRNet significantly improves the output quality with 15.4% lower MAE than the previous state-of-the-art approach.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Detection-based approaches
  • 2.2 Regression-based approaches
  • 2.3 Density estimation-based approaches
  • 2.4 CNN-based approaches
  • 2.5 Limitations of the state-of-the-art approaches
  • 3 Proposed Solution
  • 3.1 CSRNet architecture
  • 3.1.1 Dilated convolution
  • 3.1.2 Network Configuration
  • 3.2 Training method
  • 3.2.1 Ground truth generation
  • 3.2.2 Data augmentation
  • 3.2.3 Training details
  • 4 Experiments
  • 4.1 Evaluation metrics
  • 4.2 Ablations on ShanghaiTech Part_A
  • 4.3 Evaluation and comparison
  • 4.3.1 ShanghaiTech dataset
  • 4.3.2 UCF_CC_50 dataset
  • 4.3.3 The WorldExpo’10 dataset
  • 4.3.4 The UCSD dataset
  • 4.3.5 TRANCOS dataset
  • 5 Conclusion
  • 6 Acknowledgement
  • References
  • 7 Appendix: supplementary material

Knowls

  1. Knowl 1 — CSRNet Architecture for Congested Scene Recognition

    model/method

    CSRNet is a fully convolutional neural network designed for crowd counting and density map generation in congested scenes. It consists of two modular stages:

    1. Front-End Feature Extractor: Derived from the first 10 convolutional layers of a pre-trained VGG-16 network with 3 max-pooling layers (2×22 \times 2 window, stride 2) instead of the standard 5 pooling layers. This front-end processes unfixed-resolution RGB images and outputs feature maps downsampled by a factor of 8 relative to input spatial dimensions (1/81/8 input resolution).

    2. Dilated Convolutional Back-End: Replaces subsequent pooling and fully connected layers with a sequence of 6 dilated convolutional layers followed by a 1×11 \times 1 convolution layer. The layer configuration is denoted as convk-c-r\text{conv}k\text{-}c\text{-}r, where kk is kernel size (k×kk \times k), cc is the number of filters, and rr is the dilation rate. All layers employ 3×33 \times 3 filters with zero-padding to preserve feature map resolution:

      • Layer 11: conv3-512-r\text{conv}3\text{-}512\text{-}r
      • Layer 12: conv3-512-r\text{conv}3\text{-}512\text{-}r
      • Layer 13: conv3-512-r\text{conv}3\text{-}512\text{-}r
      • Layer 14: conv3-256-r\text{conv}3\text{-}256\text{-}r
      • Layer 15: conv3-128-r\text{conv}3\text{-}128\text{-}r
      • Layer 16: conv3-64-r\text{conv}3\text{-}64\text{-}r
      • Output Layer: conv1-1-1\text{conv}1\text{-}1\text{-}1

    In the primary CSRNet model (Configuration B), a uniform dilation rate of r=2r = 2 is applied across all 6 back-end layers, totaling 16.26 million parameters. The generated 1/81/8-resolution density map is upscaled to the original input resolution via bilinear interpolation with a scaling factor of 8.

  2. Knowl 2 — 2D Dilated Convolution Formulation and Spatial Properties

    definition

    A 2D dilated convolution operation is defined as:

    y(m,n)=∑i=1M∑j=1Nx(m+r×i,n+r×j)w(i,j)y(m, n) = \sum_{i=1}^M \sum_{j=1}^N x(m + r \times i, n + r \times j) w(i, j)

    where x(m,n)x(m, n) is the input feature map, w(i,j)w(i, j) is a discrete convolutional filter of dimensions M×NM \times N, y(m,n)y(m, n) is the output feature map, and r∈N+r \in \mathbb{N}^+ is the dilation rate. When r=1r = 1, the operation reduces to a standard 2D convolution.

    For a kernel of spatial size k×kk \times k, applying dilation rate rr expands the effective receptive field size k′k' to:

    k′=k+(k−1)(r−1)k' = k + (k - 1)(r - 1)

    This expands the receptive field of the filter (e.g., a 3×33 \times 3 kernel expands to receptive fields of 5×55 \times 5 for r=2r=2 and 7×77 \times 7 for r=3r=3) without increasing the number of trainable parameters or floating-point operations, and without discarding spatial resolution through pooling operations.

  3. Knowl 3 — Ground Truth Density Map Generation via Adaptive and Fixed Gaussian Kernels

    model/method

    Ground truth crowd density maps are generated by convolving discrete point annotations with normalized 2D Gaussian kernels:

    F(x)=∑i=1Nδ(x−xi)×Gσi(x)F(x) = \sum_{i=1}^N \delta(x - x_i) \times G_{\sigma_i}(x)

    where xx is pixel location, δ(x−xi)\delta(x - x_i) is a Dirac delta function centered at the ii-th annotated object location xix_i, NN is the total object count in the image, and Gσi(x)G_{\sigma_i}(x) is a normalized Gaussian kernel with standard deviation σi\sigma_i.

    Two strategies determine σi\sigma_i:

    1. Geometry-Adaptive Kernels (for highly congested scenes): σi=βdˉi\sigma_i = \beta \bar{d}_i where dˉi\bar{d}_i is the average Euclidean distance from point xix_i to its kk nearest annotated neighbor heads. The parameters are set to β=0.3\beta = 0.3 and k=3k = 3. This strategy is applied to the ShanghaiTech Part A and UCF_CC_50 datasets.

    2. Fixed Gaussian Kernels (for scenes with uniform or sparse distributions):

      • ShanghaiTech Part B: σ=15\sigma = 15
      • TRANCOS: σ=10\sigma = 10
      • WorldExpo'10: σ=3\sigma = 3
      • UCSD: σ=3\sigma = 3
  4. Knowl 4 — CSRNet Objective Function, Initialization, and Training Protocol

    model/method

    CSRNet is trained end-to-end using a pixel-wise Euclidean distance loss function:

    L(Θ)=12N∑i=1N∥Z(Xi;Θ)−ZiGT∥22L(\Theta) = \frac{1}{2N} \sum_{i=1}^N \left\| Z(X_i; \Theta) - Z_i^{\text{GT}} \right\|_2^2

    where NN is the batch size, XiX_i is the ii-th input image, Z(Xi;Θ)Z(X_i; \Theta) is the estimated density map produced by CSRNet parameterized by weights Θ\Theta, and ZiGTZ_i^{\text{GT}} is the corresponding ground truth density map.

    Initialization and Optimization:

    • Front-end convolutional layers (first 10 layers) are initialized with weights pre-trained on ImageNet from VGG-16.
    • Back-end layers are initialized from a zero-mean Gaussian distribution with standard deviation σ=0.01\sigma = 0.01.
    • Parameters are optimized using Stochastic Gradient Descent (SGD) with a fixed learning rate of 1×10−61 \times 10^{-6}.

    Data Augmentation: Nine patches of size 1/41/4 of the original image dimensions are cropped per training image: 4 non-overlapping quadrant patches and 5 randomly cropped patches. Each cropped patch is horizontally mirrored, doubling the effective training set size.

  5. Knowl 5 — Empirical Comparison of Back-End Dilation Configurations in CSRNet

    empirical result

    Four architectural configurations of CSRNet, differing solely in the dilation rates rr of their 6-layer back-end, were evaluated on the ShanghaiTech Part A crowd counting dataset. All four models contain 16.26 million parameters.

    Architecture Back-End Dilation Rates MAE MSE
    CSRNet A r=1,1,1,1,1,1r = 1, 1, 1, 1, 1, 1 69.70 116.00
    CSRNet B r=2,2,2,2,2,2r = 2, 2, 2, 2, 2, 2 68.20 115.00
    CSRNet C r=2,2,2,4,4,4r = 2, 2, 2, 4, 4, 4 71.91 120.58
    CSRNet D r=4,4,4,4,4,4r = 4, 4, 4, 4, 4, 4 75.81 120.82

    CSRNet B (r=2r = 2 throughout the back-end) achieved the lowest Mean Absolute Error (68.2) and Mean Squared Error (115.0). Adding dropout did not yield performance improvements and was omitted.

  6. Knowl 6 — Crowd Counting Estimation Performance Across Benchmark Datasets

    empirical result

    CSRNet (Configuration B) was evaluated against existing state-of-the-art methods across four standard crowd counting benchmarks using Mean Absolute Error (MAE) and Mean Squared Error (MSE):

    MAE=1N∑i=1N∣Ci−CiGT∣,MSE=1N∑i=1N∣Ci−CiGT∣2\text{MAE} = \frac{1}{N} \sum_{i=1}^N |C_i - C_i^{\text{GT}}|, \quad \text{MSE} = \sqrt{\frac{1}{N} \sum_{i=1}^N |C_i - C_i^{\text{GT}}|^2}

    where Ci=∑l=1L∑w=1Wzl,wC_i = \sum_{l=1}^L \sum_{w=1}^W z_{l,w} is the integrated sum of the estimated density map pixels, and CiGTC_i^{\text{GT}} is the true count.

    Method ShanghaiTech Part A ShanghaiTech Part B
    MAE MSE MAE MSE
    MCNN 110.2 173.2 26.4 41.3
    Switching-CNN 90.4 135.0 21.6 33.4
    CP-CNN 73.6 106.4 20.1 30.1
    CSRNet (ours) 68.2 115.0 10.6 16.0

    On ShanghaiTech Part B, CSRNet achieved a 47.3% reduction in MAE relative to CP-CNN (10.6 vs. 20.1). On UCF_CC_50 (5-fold cross-validation), CSRNet attained MAE 266.1 and MSE 397.5 (improving over CP-CNN's MAE of 295.8). On WorldExpo'10, CSRNet achieved the lowest average MAE across five test scenes (8.6 vs. CP-CNN's 8.86). On UCSD, CSRNet achieved MAE 1.16 and MSE 1.47 (after preprocessing frames by bilinear resizing from 238×158238 \times 158 to 952×632952 \times 632).

  7. Knowl 7 — Vehicle Counting on the TRANCOS Dataset and Grid Average Mean Absolute Error Evaluation

    empirical result

    CSRNet was applied to vehicle counting on the TRANCOS dataset (1,244 traffic scene images with regions of interest) and evaluated using the Grid Average Mean Absolute Error (GAME) metric:

    GAME(L)=1N∑n=1N∑l=14L∣DInl−DIngtl∣\text{GAME}(L) = \frac{1}{N} \sum_{n=1}^N \sum_{l=1}^{4^L} \left| D_{I_n}^l - D_{I_n^{\text{gt}}}^l \right|

    where for grid level LL, the image is partitioned into 4L4^L non-overlapping sub-regions, and DInlD_{I_n}^l and DIngtlD_{I_n^{\text{gt}}}^l denote the estimated and ground truth object counts in region ll of image nn. When L=0L = 0, GAME(0)\text{GAME}(0) equals MAE.

    Method GAME 0 GAME 1 GAME 2 GAME 3
    Fiaschi et al. 17.77 20.14 23.65 25.99
    Lempitsky et al. 13.76 16.72 20.72 24.36
    Hydra-3s 10.99 13.75 16.69 19.32
    FCN-HA 4.21 - - -
    CSRNet (Ours) 3.56 5.49 8.57 15.04

    CSRNet lowered the state-of-the-art error by 15.4% on GAME(0)\text{GAME}(0) compared to FCN-HA, and reduced error by 67.7% on GAME(0)\text{GAME}(0), 60.1% on GAME(1)\text{GAME}(1), 48.7% on GAME(2)\text{GAME}(2), and 22.2% on GAME(3)\text{GAME}(3) compared to Hydra-3s.

  8. Knowl 8 — Feature Redundancy and Suboptimality in Multi-Column CNN Architectures

    empirical result

    Multi-Column CNN (MCNN) architectures use parallel network columns with distinct filter sizes (large, medium, small) intended to capture scale-specific crowd density features. Empirical analysis demonstrated that separate MCNN columns produce nearly identical error distribution curves across varying test densities, indicating that individual columns learn redundant rather than specialized features.

    A single-column deep network with fewer parameters was compared against the multi-column MCNN on the ShanghaiTech Part A dataset. The single-column architecture consisted of: CR(32,3)−M−CR(64,3)−M−CR(64,3)−M−CR(32,3)−CR(32,3)−CR(1,1)\text{CR}(32, 3) - \text{M} - \text{CR}(64, 3) - \text{M} - \text{CR}(64, 3) - \text{M} - \text{CR}(32, 3) - \text{CR}(32, 3) - \text{CR}(1, 1) where CR(m,n)\text{CR}(m, n) is a convolutional layer with mm filters of size n×nn \times n followed by ReLU, and M\text{M} is a 2×22 \times 2 max-pooling layer.

    Method Parameters MAE MSE
    Column 1 of MCNN 57.75k 141.2 206.8
    Column 2 of MCNN 45.99k 160.5 239.0
    Column 3 of MCNN 25.14k 153.7 230.2
    MCNN Total 127.68k 110.2 185.9
    A deeper single-column CNN 83.84k 93.0 142.2

    The deeper single-column CNN achieved lower MAE (93.0 vs. 110.2) and MSE (142.2 vs. 185.9) than MCNN while utilizing 34.3% fewer parameters (83.84k vs. 127.68k).

  9. Knowl 9 — Density Map Quality Assessment via PSNR and SSIM

    empirical result

    The visual quality and structural fidelity of generated density maps were evaluated using Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) after resizing maps to original image dimensions via bilinear interpolation and normalizing.

    On ShanghaiTech Part A, CSRNet outperformed prior architectures in density map quality:

    • MCNN: PSNR 21.40 dB, SSIM 0.52
    • CP-CNN: PSNR 21.72 dB, SSIM 0.72
    • CSRNet (ours): PSNR 23.79 dB, SSIM 0.76
    Dataset PSNR (dB) SSIM
    ShanghaiTech Part A 23.79 0.76
    ShanghaiTech Part B 27.02 0.89
    UCF_CC_50 18.76 0.52
    The WorldExpo'10 26.94 0.92
    The UCSD 20.02 0.86
    TRANCOS 27.10 0.93

Coverage note — All substantial contributions—including the CSRNet architecture, dilated convolution formulation, training and ground truth generation methods, MCNN redundancy analysis, benchmark counting evaluations, and density map quality assessments—have been included as knowls. Purely qualitative visual prediction figures from the appendix were omitted.

References

  1. 1.Beibei Zhan, Dorothy N Monekosso, Paolo Remagnino, Sergio A Velastin, and Li-Qun Xu. Crowd analysis: a survey. Machine Vision and Applications, 19(5-6):345–357, 2008.
  2. 2.Teng Li, Huan Chang, Meng Wang, Bingbing Ni, Richang Hong, and Shuicheng Yan. Crowded scene analysis: A survey. IEEE transactions on circuits and systems for video technology, 25(3):367–386, 2015.
  3. 3.Cong Zhang, Hongsheng Li, Xiaogang Wang, and Xiaokang Yang. Cross-scene crowd counting via deep convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 833–841, 2015.
  4. 4.Deepak Babu Sam, Shiv Surya, and R Venkatesh Babu. Switching convolutional neural network for crowd counting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 1, page 6, 2017.
  5. 5.Vishwanath A Sindagi and Vishal M Patel. Generating high-quality crowd density maps using contextual pyramid CNNs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1861–1870, 2017.
  6. 6.Cong Zhang, Kai Kang, Hongsheng Li, Xiaogang Wang, Rong Xie, and Xiaokang Yang. Data-driven crowd understanding: a baseline for a large-scale crowd dataset. IEEE Transactions on Multimedia, 18(6):1048–1061, 2016.
  7. 7.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015.
  8. 8.Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Xiaohui Shen, Ming-Ming Cheng, Jiashi Feng, Yao Zhao, and Shuicheng Yan. Stc: A simple to complex framework for weakly-supervised semantic segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(11):2314–2320, 2017.
  9. 9.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In IEEE CVPR, 2017.
  10. 10.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016.
  11. 11.L. C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, PP(99):1–1, 2017.
  12. 12.Junting Pan, Elisa Sayrol, Xavier Giro-i Nieto, Kevin McGuinness, and Noel E O’Connor. Shallow and deep convolutional networks for saliency prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 598–606, 2016.
  13. 13.Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM international conference on Multimedia, pages 675–678. ACM, 2014.
  14. 14.Jiantao Qiu, Jie Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, Yu Wang, and Huazhong Yang. Going deeper with embedded FPGA platform for convolutional neural network. In Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, FPGA ’16, pages 26–35, New York, NY, USA, 2016. ACM.
  15. 15.Xiaofan Zhang, Xinheng Liu, Anand Ramachandran, Chuanhao Zhuge, Shibin Tang, Peng Ouyang, Zuofu Cheng, Kyle Rupnow, and Deming Chen. High-performance video content recognition with long-term recurrent convolutional network for FPGA. In Field Programmable Logic and Applications (FPL), 2017 27th International Conference on, pages 1–4. IEEE, 2017.
  16. 16.Xiaofan Zhang, Anand Ramachandran, Chuanhao Zhuge, Di He, Wei Zuo, Zuofu Cheng, Kyle Rupnow, and Deming Chen. Machine learning on FPGAs to face the IoT revolution. In Computer-Aided Design (ICCAD), 2017 IEEE/ACM International Conference on, pages 819–826. IEEE, 2017.
  17. 17.Renzo Andri, Lukas Cavigelli, Davide Rossi, and Luca Benini. Yodann: An ultra-low power convolutional neural network accelerator based on binary weights. In VLSI (ISVLSI), 2016 IEEE Computer Society Annual Symposium on, pages 236–241. IEEE, 2016.
  18. 18.Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. Single-image crowd counting via multi-column convolutional neural network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 589–597, 2016.
  19. 19.Lokesh Boominathan, Srinivas SS Kruthiventi, and R Venkatesh Babu. Crowdnet: a deep convolutional network for dense crowd counting. In Proceedings of the 2016 ACM on Multimedia Conference, pages 640–644. ACM, 2016.
  20. 20.Daniel Onoro-Rubio and Roberto J López-Sastre. Towards perspective-free object counting with deep learning. In European Conference on Computer Vision, pages 615–629. Springer, 2016.
  21. 21.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  22. 22.Haroon Idrees, Imran Saleemi, Cody Seibert, and Mubarak Shah. Multi-source multi-scale counting in extremely dense crowd images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2547–2554, 2013.
  23. 23.A. B. Chan, Zhang-Sheng John Liang, and N. Vasconcelos. Privacy preserving crowd monitoring: Counting people without people models or tracking. In 2008 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–7, June 2008.
  24. 24.Shanghang Zhang, Guanhang Wu, Joao P Costeira, and Jose MF Moura. Fcn-rlstm: Deep spatio-temporal neural networks for vehicle counting in city cameras. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3667–3676, 2017.
  25. 25.Chen Change Loy, Ke Chen, Shaogang Gong, and Tao Xiang. Crowd counting and profiling: Methodology and evaluation. In Modeling, Simulation and Visual Analysis of Crowds, pages 347–382. Springer, 2013.
  26. 26.Piotr Dollar, Christian Wojek, Bernt Schiele, and Pietro Perona. Pedestrian detection: An evaluation of the state of the art. IEEE transactions on pattern analysis and machine intelligence, 34(4):743–761, 2012.
  27. 27.Paul Viola and Michael J Jones. Robust real-time face detection. International journal of computer vision, 57(2):137–154, 2004.
  28. 28.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on, volume 1, pages 886–893. IEEE, 2005.
  29. 29.Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan. Object detection with discriminatively trained part-based models. IEEE transactions on pattern analysis and machine intelligence, 32(9):1627–1645, 2010.
  30. 30.Antoni B Chan and Nuno Vasconcelos. Bayesian poisson regression for crowd counting. In Computer Vision, 2009 IEEE 12th International Conference on, pages 545–551. IEEE, 2009.
  31. 31.D. G. Lowe. Object recognition from local scale-invariant features. In Proceedings of the Seventh IEEE International Conference on Computer Vision, volume 2, pages 1150–1157 vol.2, 1999.
  32. 32.Victor Lempitsky and Andrew Zisserman. Learning to count objects in images. In Advances in Neural Information Processing Systems, pages 1324–1332, 2010.
  33. 33.Viet-Quoc Pham, Tatsuo Kozakaya, Osamu Yamaguchi, and Ryuzo Okada. Count forest: Co-voting uncertain number of targets using random forest for crowd density estimation. In Computer Vision (ICCV), 2015 IEEE International Conference on, pages 3253–3261. IEEE, 2015.
  34. 34.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  35. 35.F. Chollet. Xception: Deep learning with depthwise separable convolutions. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1800–1807, July 2017.
  36. 36.Elad Walach and Lior Wolf. Learning to count with CNN boosting. In European Conference on Computer Vision, pages 660–676. Springer, 2016.
  37. 37.Chong Shang, Haizhou Ai, and Bo Bai. End-to-end crowd counting via joint learning local and global count. In Image Processing (ICIP), 2016 IEEE International Conference on, pages 1215–1219. IEEE, 2016.
  38. 38.Mark Marsden, Kevin McGuiness, Suzanne Little, and Noel E O’Connor. Fully convolutional crowd counting on highly congested scenes. arXiv preprint arXiv:1612.00220, 2016.
  39. 39.Vishwanath A Sindagi and Vishal M Patel. Cnn-based cascaded multi-task learning of high-level prior and density estimation for crowd counting. In Advanced Video and Signal Based Surveillance (AVSS), 2017 14th IEEE International Conference on, pages 1–6. IEEE, 2017.
  40. 40.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. CoRR, abs/1706.05587, 2017.
  41. 41.Matthew D Zeiler, Dilip Krishnan, Graham W Taylor, and Rob Fergus. Deconvolutional networks. In Computer Vision and Pattern Recognition (CVPR), 2010 IEEE Conference on, pages 2528–2535. IEEE, 2010.
  42. 42.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1520–1528, 2015.
  43. 43.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  44. 44.Roberto Lpez-Sastre Saturnino Maldonado Bascn Ricardo Guerrero-Gmez-Olmedo, Beatriz Torre-Jimnez and Daniel Ooro-Rubio. Extremely overlapping vehicle counting. In Iberian Conference on Pattern Recognition and Image Analysis (IbPRIA), 2015.
  45. 45.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958, 2014.
  46. 46.Ke Chen, Shaogang Gong, Tao Xiang, and Chen Change Loy. Cumulative attribute space for age and crowd density estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2467–2474, 2013.
  47. 47.L. Fiaschi, U. Koethe, R. Nair, and F. A. Hamprecht. Learning to count with regression forest and structured labels. In Proceedings of the 21st International Conference on Pattern Recognition (ICPR2012), pages 2685–2688, Nov 2012.

Citation

MLA
Li, Y., et al. “CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 1091–100, https://doi.org/10.1109/CVPR.2018.00120.
APA
Li, Y., Zhang, X., & Chen, D. (2018). CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1091–1100. https://doi.org/10.1109/CVPR.2018.00120
Chicago
Li, Y., X. Zhang, and D. Chen. 2018. “CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1091–1100. https://doi.org/10.1109/CVPR.2018.00120.
Harvard
Li, Y., Zhang, X. and Chen, D. (2018) “CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp. 1091–1100. Available at: https://doi.org/10.1109/CVPR.2018.00120.
Vancouver
1. Li Y, Zhang X, Chen D (2018) CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp 1091–1100

BibTeX

@inproceedings{Li_2018, title={CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes}, url={http://dx.doi.org/10.1109/CVPR.2018.00120}, DOI={10.1109/cvpr.2018.00120}, booktitle={2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition}, publisher={IEEE}, author={Li, Yuhong and Zhang, Xiaofan and Chen, Deming}, year={2018}, month=June, pages={1091–1100} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE