Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Yingying ZhangDesen ZhouSiqin ChenShenghua GaoYi Ma

article2016CVPR2,167 citations

Proposes a multi-column convolutional neural network using varying receptive field sizes alongside geometry-adaptive density kernels to accurately estimate crowd counts across diverse scene perspectives and resolutions without requiring camera calibration.

Listen

Crowd disasters and stampedes represent major public safety risks, underscoring the critical need for reliable automated crowd monitoring. Accurately estimating people counts from single images is challenging due to severe occlusions, large variations in crowd density, and significant perspective distortion that causes people and head sizes to vary dramatically across an image. Previous counting methods often required manual foreground extraction or explicit scene geometry calculations, both of which are impractical in real-world scenarios.

The article develops and evaluates a robust method for accurately estimating crowd counts from individual still images across arbitrary perspectives and densities without requiring prior knowledge of scene geometry. To achieve this, the authors designed a Multi-column Convolutional Neural Network (MCNN) that maps an input image directly into a crowd density map, from which the total count is calculated through spatial integration. The network utilizes three parallel columns with different filter receptive field sizes (large, medium, and small) to adaptively capture people at different physical scales, and it processes arbitrary image sizes without distortion. The model is trained using geometry-adaptive kernels that approximate perspective distortion based on the local distance between neighboring heads. To thoroughly evaluate the method, the authors also introduced the Shanghaitech dataset, a large benchmark consisting of 1,198 images and roughly 330,000 annotated heads across high-density and street-level scenes.

The experimental findings show that the proposed multi-column approach outperforms existing state-of-the-art methods across multiple benchmarks. On the dense subset of the Shanghaitech dataset (Part A), MCNN achieved a Mean Absolute Error (MAE) of 110.2, significantly outperforming the previous state-of-the-art error of 181.8. On the street-level subset (Part B), it reduced the error to 26.4 compared to the previous 32.0. The model also delivered superior accuracy on the UCF CC 50 benchmark (MAE of 377.6 versus the previous best of 419.5), the sparse UCSD benchmark (MAE of 1.07 versus 1.60), and the WorldExpo'10 dataset (average MAE of 11.6 versus 12.9). Component analyses confirmed that combining multiple filter sizes clearly outperformed single-column architectures, predicting spatial density maps proved far superior to directly regressing total counts, and pre-training individual columns was essential to prevent optimization bottlenecks. Furthermore, the model demonstrated strong transferability: pre-training on a large, dense dataset and fine-tuning only the final layers on a smaller target dataset reduced the error on UCF CC 50 from 377.7 down to 295.1.

These results demonstrate that accurate, automated crowd estimation can be achieved in real-world operational environments without requiring expensive scene calibration or manual geometric inputs. By producing high-resolution spatial density maps rather than a single aggregate head count, the system provides actionable spatial intelligence to identify localized overcrowding, enabling proactive crowd management and emergency response. Organizations deploying vision-based crowd monitoring can adopt this multi-scale framework and leverage transfer learning to adapt models to specific operational camera views with minimal target data and low training costs.

While the model delivers high confidence and strong generalization across both sparse and extremely congested environments, its performance assumes that head annotations provide sufficient density cues for geometry adaptation. Practitioners applying the method to novel camera feeds should utilize pre-trained models and fine-tune only the final network layers rather than retraining the entire architecture from scratch when annotated target samples are scarce.

  • Paper: Learning To Count Objects in Images, V. Lempitsky et al. (2010). Its dot-supervised density-map formulation supplies the key counting framework that MCNN adapts from linear models to a multi-column convolutional network.
Cover for Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Abstract

This paper aims to develop a method than can accurately estimate the crowd count from an individual image with arbitrary crowd density and arbitrary perspective. To this end, we have proposed a simple but effective Multi-column Convolutional Neural Network (MCNN) architecture to map the image to its crowd density map. The proposed MCNN allows the input image to be of arbitrary size or resolution. By utilizing filters with receptive fields of different sizes, the features learned by each column CNN are adaptive to variations in people/head size due to perspective effect or image resolution. Furthermore, the true density map is computed accurately based on geometry-adaptive kernels which do not need knowing the perspective map of the input image. Since exiting crowd counting datasets do not adequately cover all the challenging situations considered in our work, we have collected and labelled a large new dataset that includes 1198 images with about 330,000 heads annotated. On this challenging new dataset, as well as all existing datasets, we conduct extensive experiments to verify the effectiveness of the proposed model and method. In particular, with the proposed simple MCNN model, our method outperforms all existing methods. In addition, experiments show that our model, once trained on one dataset, can be readily transferred to a new dataset.

Table of Contents

  • 1. Introduction
  • 2. Multi-column CNN for Crowd Counting
  • 2.1. Density map based crowd counting
  • 2.2. Density map via geometry­adaptive kernels
  • 2.3. Multi­column CNN for density map estimation
  • 2.4. Optimization of MCNN
  • 2.5. Transfer learning setting
  • 3. Experiments
  • 3.1. Evaluation metric
  • 3.2. Shanghaitech dataset
  • 3.3. The UCF CC 50 dataset
  • 3.4. The UCSD dataset
  • 3.5. The WorldExpo'10 dataset
  • 3.6. Evaluation on transfer learning
  • 4. Conclusion
  • 5. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Multi-Column Convolutional Neural Network Architecture for Crowd Density Estimation

    model/method

    The Multi-column Convolutional Neural Network (MCNN) maps a single input image of arbitrary resolution to a continuous crowd density map. To handle scale variations of human heads caused by perspective distortion and image resolution, MCNN uses three parallel convolutional columns containing filters with different receptive field sizes (large, medium, and small):

    1. Large Receptive Field Column (adapted to larger heads):

      • Conv 1: filter size 9×99 \times 9, 16 filters
      • Max Pooling 1: 2×22 \times 2 window
      • Conv 2: filter size 7×77 \times 7, 32 filters
      • Max Pooling 2: 2×22 \times 2 window
      • Conv 3: filter size 7×77 \times 7, 16 filters
      • Conv 4: filter size 7×77 \times 7, 8 filters
    2. Medium Receptive Field Column (adapted to medium heads):

      • Conv 1: filter size 7×77 \times 7, 20 filters
      • Max Pooling 1: 2×22 \times 2 window
      • Conv 2: filter size 5×55 \times 5, 40 filters
      • Max Pooling 2: 2×22 \times 2 window
      • Conv 3: filter size 5×55 \times 5, 20 filters
      • Conv 4: filter size 5×55 \times 5, 10 filters
    3. Small Receptive Field Column (adapted to smaller heads):

      • Conv 1: filter size 5×55 \times 5, 24 filters
      • Max Pooling 1: 2×22 \times 2 window
      • Conv 2: filter size 3×33 \times 3, 48 filters
      • Max Pooling 2: 2×22 \times 2 window
      • Conv 3: filter size 3×33 \times 3, 24 filters
      • Conv 4: filter size 3×33 \times 3, 12 filters

    All convolutional layers use Rectified Linear Unit (ReLU) activations. Columns with larger filters employ fewer channels to reduce total parameter count.

    The feature maps produced by the fourth convolutional layer of all three columns are stacked along the channel dimension, resulting in a 30-channel (8+10+128 + 10 + 12) merged feature representation. A final 1×11 \times 1 convolutional layer with 1 filter combines these channels into the 1-channel crowd density map. Fully connected layers are omitted, enabling inputs of arbitrary spatial dimensions without geometric distortion. Due to the two 2×22 \times 2 max pooling operations, the output density map has 14\frac{1}{4} the spatial resolution of the input image along each spatial dimension. The total crowd count is computed by summing all pixel values in the estimated density map.

  2. Knowl 2 — Geometry-Adaptive Gaussian Kernel for Density Map Generation

    equation

    To construct ground-truth crowd density maps from point-annotated head locations without requiring explicit camera calibration or 3D scene geometry, geometry-adaptive Gaussian kernels are used to model local head scale variations in crowded scenes.

    Given an image containing NN labeled head positions x1,x2,…,xN∈R2x_1, x_2, \dots, x_N \in \mathbb{R}^2, the discrete head annotation function is: H(x)=∑i=1Nδ(x−xi)H(x) = \sum_{i=1}^N \delta(x - x_i) where δ(⋅)\delta(\cdot) is the Dirac delta function and xx denotes spatial pixel coordinates.

    The continuous ground-truth density map F(x)F(x) is formed by convolving each head delta function with a local Gaussian kernel Gσi(x)G_{\sigma_i}(x): F(x)=∑i=1Nδ(x−xi)∗Gσi(x)F(x) = \sum_{i=1}^N \delta(x - x_i) * G_{\sigma_i}(x) where the standard deviation σi\sigma_i is determined per head based on the average Euclidean distance dˉi\bar{d}_i to its kk nearest neighboring heads: σi=βdˉiwithdˉi=1k∑j=1kdji\sigma_i = \beta \bar{d}_i \quad \text{with} \quad \bar{d}_i = \frac{1}{k} \sum_{j=1}^k d_j^i where {d1i,d2i,…,dki}\{d_1^i, d_2^i, \dots, d_k^i\} are the distances from head xix_i to its kk nearest head neighbors in the image plane, and β=0.3\beta = 0.3 is a scaling hyperparameter determined empirically.

    For sparse scenes where neighboring distance does not correlate with head scale, a fixed spread parameter σ\sigma is used for all heads instead.

  3. Knowl 3 — Optimization and Column-Wise Pre-Training for Multi-Column CNN

    model/method

    The Multi-column Convolutional Neural Network (MCNN) is trained by minimizing the pixel-wise Euclidean loss between the estimated density map and the downsampled ground-truth density map: L(Θ)=12N∑i=1N∥F(Xi;Θ)−Fi∥22L(\Theta) = \frac{1}{2N} \sum_{i=1}^N \|F(X_i; \Theta) - F_i\|_2^2 where Θ\Theta represents all learnable network parameters, NN is the total number of training images, XiX_i is the ii-th input image, FiF_i is its corresponding ground-truth density map downsampled by a factor of 14\frac{1}{4} to match the network output size, and F(Xi;Θ)F(X_i; \Theta) is the predicted density map.

    Because direct joint optimization of the multi-column architecture from random initializations is prone to local minima and vanishing gradients on small training sets, training is conducted in two stages:

    1. Independent Single-Column Pre-Training: Each of the three CNN columns (large, medium, and small receptive field branches) is pre-trained independently. In this stage, the output of the fourth convolutional layer of an individual column is mapped directly to the ground-truth density map using a temporary 1×11 \times 1 convolution.
    2. End-to-End Joint Fine-Tuning: The pre-trained weights from each individual column are used to initialize the parallel columns of the full MCNN. Their output feature maps are concatenated into a 30-channel representation, followed by a shared 1×11 \times 1 convolution layer. All parameters across the entire network are then fine-tuned simultaneously using batch-based stochastic gradient descent and backpropagation.
  4. Knowl 4 — Shanghaitech Crowd Counting Dataset Specification

    data/table

    The Shanghaitech dataset is a large-scale crowd counting benchmark comprising 1,198 images and 330,165 annotated head centers, with no two images sharing the same camera viewpoint. It contains two distinct subsets:

    • Part A: 482 images with varying resolutions collected from the Internet, featuring high-density crowds (33 to 3,139 people per image). 300 images are used for training and 182 for testing.
    • Part B: 716 images captured at a fixed resolution of 768×1024768 \times 1024 from street surveillance in metropolitan Shanghai, featuring relatively sparse crowds (9 to 578 people per image). 400 images are used for training and 316 for testing.

    The table below summarizes the dataset properties in comparison with other crowd counting benchmarks:

    Dataset Resolution Num Max Min Ave Total
    UCSD 158×238158 \times 238 2000 46 11 24.9 49,885
    UCF_CC_50 different 50 4543 94 1279.5 63,974
    WorldExpo'10 576×720576 \times 720 3980 253 1 50.2 199,923
    Shanghaitech Part_A different 482 3139 33 501.4 241,677
    Shanghaitech Part_B 768×1024768 \times 1024 716 578 9 123.6 88,488

    In the table, Num is the total number of images, Max and Min are the maximum and minimum head counts in a single image, Ave is the average crowd count, and Total is the total number of annotated heads. For training, 9 patches of size 1/41/4 of the original image are cropped from each image for data augmentation.

  5. Knowl 5 — Crowd Counting Evaluation and Ablation on Shanghaitech Dataset

    data/table

    Crowd counting performance is evaluated using Mean Absolute Error (MAE) and Mean Squared Error (MSE): MAE=1N∑i=1N∣zi−z^i∣,MSE=1N∑i=1N(zi−z^i)2MAE = \frac{1}{N} \sum_{i=1}^N |z_i - \hat{z}_i|, \quad MSE = \sqrt{\frac{1}{N} \sum_{i=1}^N (z_i - \hat{z}_i)^2} where NN is the number of test images, ziz_i is the ground-truth count, and z^i\hat{z}_i is the estimated count obtained by integrating the density map.

    The performance comparison on Shanghaitech Part A and Part B is shown below:

    Part A Part B
    Method MAE MSE MAE MSE
    LBP+RR 303.2 371.0 59.1 81.7
    Zhang et al. (2015) 181.8 277.7 32.0 49.8
    MCNN-CCR 245.0 336.1 70.9 95.9
    MCNN 110.2 173.2 26.4 41.3

    Ablation analysis on Shanghaitech Part A evaluates the impact of individual columns and pre-training:

    • CNN(Large): MAE = 141.2, MSE = 206.8
    • CNN(Medium): MAE = 160.5, MSE = 239.0
    • CNN(Small): MAE = 153.7, MSE = 230.2
    • MCNN without pre-training: MAE = 122.8, MSE = 185.9
    • MCNN (Full with pre-training): MAE = 110.2, MSE = 173.2

    The multi-column architecture outperforms any single-column network by at least 31.0 MAE points, and pre-training provides an additional reduction of 12.6 in MAE.

  6. Knowl 6 — Crowd Counting Evaluation on UCF_CC_50 Dataset

    data/table

    The UCF_CC_50 dataset contains 50 images from the Internet with extreme crowd densities (ranging from 94 to 4,543 individuals per image, with an average of 1,280). Evaluation uses standard 5-fold cross-validation with Mean Absolute Error (MAE) and Mean Squared Error (MSE).

    Method MAE MSE
    Rodriguez et al. (2011) 655.7 697.8
    Lempitsky et al. (2010) 493.4 487.1
    Idrees et al. (2013) 419.5 541.6
    Zhang et al. (2015) 467.0 498.5
    MCNN 377.6 509.1

    MCNN achieves the lowest MAE (377.6) among all compared methods, improving upon the multi-source feature method of Idrees et al. (MAE 419.5) and the CNN approach of Zhang et al. (MAE 467.0).

  7. Knowl 7 — Crowd Counting Evaluation on UCSD Dataset

    data/table

    The UCSD dataset contains 2,000 video frames of size 158×238158 \times 238 with an average of 24.9 people per frame. Following the standard protocol, frames 601–1400 are used for training and the remaining 1,200 frames for testing. A fixed spread parameter σ\sigma is used for density maps due to low crowd density, and pixel values outside the provided Region of Interest (ROI) are set to zero.

    Method MAE MSE
    Kernel Ridge Regression 2.16 7.45
    Ridge Regression 2.25 7.82
    Gaussian Process Regression 2.24 7.97
    Cumulative Attribute Regression 2.07 6.86
    Zhang et al. (2015) 1.60 3.31
    MCNN 1.07 1.35

    MCNN achieves an MAE of 1.07 and MSE of 1.35, outperforming both foreground-segmentation feature regression baselines and prior CNN-based methods on sparse crowd data.

  8. Knowl 8 — Crowd Counting Evaluation on WorldExpo'10 Dataset

    data/table

    The WorldExpo'10 dataset contains 3,980 annotated frames from 108 surveillance cameras (3,380 for training and 5 test video sequences of 120 frames each). Ground-truth density maps are generated using perspective maps M(x)M(x) according to σ=0.2⋅M(x)\sigma = 0.2 \cdot M(x), where M(x)M(x) denotes the number of pixels representing one square meter at position xx. Evaluation is restricted to the provided Region of Interest (ROI) masks.

    Method Scene 1 Scene 2 Scene 3 Scene 4 Scene 5 Average
    LBP + RR 13.6 59.8 37.1 21.8 23.4 31.0
    Zhang et al. (2015) 9.8 14.1 14.3 22.2 3.7 12.9
    MCNN 3.4 20.6 12.9 13.0 8.1 11.6

    MCNN achieves the lowest overall average MAE of 11.6 across the five diverse test scenes, outperforming the fine-tuned CNN model of Zhang et al. (12.9) and the LBP+RR baseline (31.0).

  9. Knowl 9 — Cross-Dataset Transfer Learning via Layer-Selective Fine-Tuning

    empirical result

    Transfer learning generalizability of MCNN is evaluated by training on a large source dataset (Shanghaitech Part A) and testing/adapting to a target dataset with few training samples (UCF_CC_50):

    Method MAE MSE
    MCNN w/o transfer (trained solely on target dataset) 377.7 509.1
    MCNN trained on Part A (zero-shot transfer) 397.7 624.1
    Fine-tune the whole MCNN on target dataset 378.3 594.6
    Fine-tune the last two layers only on target dataset 295.1 490.23

    Zero-shot transfer from Shanghaitech Part A achieves an MAE of 397.7 on UCF_CC_50 without seeing any target data. When target samples are available, fine-tuning only the last two layers reduces MAE to 295.1 (a 21.9% reduction compared to training solely on UCF_CC_50). In contrast, fine-tuning the entire network with limited target samples degrades performance to 378.3 due to overfitting. Freezing the early layers retains generic multi-scale filters learned on the large source dataset while adapting output mappings to the target domain.

  10. Knowl 10 — Density Map Regression versus Direct Crowd Count Regression

    empirical result

    To evaluate the utility of density map intermediate representations versus direct total count regression, a baseline model called MCNN-CCR (MCNN-based Crowd Count Regression) is trained using the global count loss: L(Θ)=12N∑i=1N∣∬SF(Xi;Θ)dxdy−zi∣2L(\Theta) = \frac{1}{2N} \sum_{i=1}^N \left| \iint_S F(X_i; \Theta) dxdy - z_i \right|^2 where SS is the spatial support of the output map F(Xi;Θ)F(X_i; \Theta), ziz_i is the scalar ground-truth crowd count of image XiX_i, and ground-truth spatial density maps are not used.

    Direct crowd count regression yields substantially worse performance than pixel-level density map supervision on the Shanghaitech benchmark:

    • Part A: MCNN achieves MAE = 110.2, MSE = 173.2 versus MCNN-CCR MAE = 245.0, MSE = 336.1 (MAE increases by 122.3% with direct count regression).
    • Part B: MCNN achieves MAE = 26.4, MSE = 41.3 versus MCNN-CCR MAE = 70.9, MSE = 95.9 (MAE increases by 168.6% with direct count regression).

    Pixel-wise density map supervision provides localized spatial constraints that guide multi-scale convolutional filters to learn semantically meaningful representations of human heads across varying scales.

Coverage note — No substantial contributed material was omitted; minor qualitative 10-bin crowd count plots were integrated into the Shanghaitech empirical result knowl.

References

  1. 1.S. An, W. Liu, and S. Venkatesh. Face recognition using kernel ridge regression. In CVPR, pages 1–7. IEEE, 2007.
  2. 2.A. Bansal and K. Venkatesh. People counting in high density crowds from still images. arXiv preprint arXiv:1507.08445, 2015.
  3. 3.G. J. Brostow and R. Cipolla. Unsupervised bayesian detection of independent motion in crowds. In CVPR, volume 1, pages 594–601. IEEE, 2006.
  4. 4.A. B. Chan, Z.-S. J. Liang, and N. Vasconcelos. Privacy preserving crowd monitoring: Counting people without people models or tracking. In CVPR, pages 1–7. IEEE, 2008.
  5. 5.A. B. Chan and N. Vasconcelos. Bayesian poisson regression for crowd counting. In ICCV, pages 545–551. IEEE, 2009.
  6. 6.K. Chen, S. Gong, T. Xiang, and C. C. Loy. Cumulative attribute space for age and crowd density estimation. In CVPR, pages 2467–2474. IEEE, 2013.
  7. 7.K. Chen, C. C. Loy, S. Gong, and T. Xiang. Feature mining for localised crowd counting. In BMVC, volume 1, page 3, 2012.
  8. 8.D. Ciresan, U. Meier, and J. Schmidhuber. Multi-column deep neural networks for image classification. In CVPR, pages 3642–3649. IEEE, 2012.
  9. 9.K. Fukushima. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological cybernetics, 36(4):193–202, 1980.
  10. 10.W. Ge and R. T. Collins. Marked point processes for crowd counting. In CVPR, pages 2913–2920. IEEE, 2009.
  11. 11.G. Hinton, S. Osindero, and Y. Teh. A fast learning algorithm for deep belief nets. NEURAL COMPUT, 18(7):1527–1554, 2006.
  12. 12.H. Idrees, I. Saleemi, C. Seibert, and M. Shah. Multi-source multi-scale counting in extremely dense crowd images. In CVPR, pages 2547–2554. IEEE, 2013.
  13. 13.H. Idrees, K. Soomro, and M. Shah. Detecting humans in dense crowds using locally-consistent scale prior and global occlusion reasoning. Pattern Analysis and Machine Intelligence, 2005.
  14. 14.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  15. 15.D. Kong, D. Gray, and H. Tao. Counting pedestrians in crowds using viewpoint invariant training. In BMVC. Citeseer, 2005.
  16. 16.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradientbased learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  17. 17.V. Lempitsky and A. Zisserman. Learning to count objects in images. In Advances in Neural Information Processing Systems, pages 1324–1332, 2010.
  18. 18.M. Li, Z. Zhang, K. Huang, and T. Tan. Estimating the number of people in crowded scenes by mid based foreground segmentation and head-shoulder detection. In ICPR, pages 1–4. IEEE, 2008.
  19. 19.Z. Lin and L. S. Davis. Shape-based human detection and segmentation via hierarchical part-template matching. Pattern Analysis and Machine Intelligence, 32(4):604–618, 2010.
  20. 20.B. Liu and N. Vasconcelos. Bayesian model adaptation for crowd counts. In ICCV, 2015.
  21. 21.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. arXiv preprint arXiv:1411.4038, 2014.
  22. 22.A. Marana, L. d. F. Costa, R. Lotufo, and S. Velastin. On the efficacy of texture analysis for crowd monitoring. In International Symposium on Computer Graphics, Image Processing, and Vision, pages 354–361. IEEE, 1998.
  23. 23.N. Paragios and V. Ramesh. A mrf-based approach for realtime subway monitoring. In CVPR, volume 1, pages I–1034. IEEE, 2001.
  24. 24.V. Rabaud and S. Belongie. Counting crowded moving objects. In CVPR, volume 1, pages 705–711. IEEE, 2006.
  25. 25.C. S. Regazzoni and A. Tesei. Distributed data fusion for real-time crowding estimation. Signal Processing, 53(1):47–63, 1996.
  26. 26.M. Rodriguez, I. Laptev, J. Sivic, and J.-Y. Audibert. Density-aware person detection and tracking in crowds. In ICCV, pages 2423–2430. IEEE, 2011.
  27. 27.D. Ryan, S. Denman, C. Fookes, and S. Sridharan. Crowd counting using multiple local features. In Digital Image Computing: Techniques and Applications, pages 81–88. IEEE, 2009.
  28. 28.K. Tota and H. Idrees. Counting in dense crowds using deep features.
  29. 29.P. Viola, M. J. Jones, and D. Snow. Detecting pedestrians using patterns of motion and appearance. International Journal of Computer Vision, 63(2):153–161, 2005.
  30. 30.M. Wang and X. Wang. Automatic adaptation of a generic pedestrian detector to a specific traffic scene. In CVPR, pages 3401–3408. IEEE, 2011.
  31. 31.B. Wu and R. Nevatia. Detection of multiple, partially occluded humans in a single image by bayesian combination of edgelet part detectors. In ICCV, volume 1, pages 90–97. IEEE, 2005.
  32. 32.M. D. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. V. Le, P. Nguyen, A. Senior, V. Vanhoucke, and J. Dean. On rectified linear units for speech processing. In ICASSP, pages 3517–3521. IEEE, 2013.
  33. 33.C. Zhang, H. Li, X. Wang, and X. Yang. Cross-scene crowd counting via deep convolutional neural networks. In CVPR, 2015.
  34. 34.T. Zhao, R. Nevatia, and B. Wu. Segmentation and tracking of multiple humans in crowded environments. Pattern Analysis and Machine Intelligence, 30(7):1198–1211, 2008.

Citation

MLA
Zhang, Y., et al. “Single-Image Crowd Counting via Multi-Column Convolutional Neural Network”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 589–97, https://doi.org/10.1109/CVPR.2016.70.
APA
Zhang, Y., Zhou, D., Chen, S., Gao, S., & Ma, Y. (2016). Single-Image Crowd Counting via Multi-Column Convolutional Neural Network. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 589–597. https://doi.org/10.1109/CVPR.2016.70
Chicago
Zhang, Y., D. Zhou, S. Chen, S. Gao, and Y. Ma. 2016. “Single-Image Crowd Counting via Multi-Column Convolutional Neural Network”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 589–97. https://doi.org/10.1109/CVPR.2016.70.
Harvard
Zhang, Y. et al. (2016) “Single-Image Crowd Counting via Multi-Column Convolutional Neural Network”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 589–597. Available at: https://doi.org/10.1109/CVPR.2016.70.
Vancouver
1. Zhang Y, Zhou D, Chen S, Gao S, Ma Y (2016) Single-Image Crowd Counting via Multi-Column Convolutional Neural Network. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 589–597

BibTeX

@inproceedings{Zhang_2016, title={Single-Image Crowd Counting via Multi-Column Convolutional Neural Network}, url={http://dx.doi.org/10.1109/CVPR.2016.70}, DOI={10.1109/cvpr.2016.70}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhang, Yingying and Zhou, Desen and Chen, Siqin and Gao, Shenghua and Ma, Yi}, year={2016}, month=June, pages={589–597} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE