Learning To Count Objects in Images

V. LempitskyAndrew Zisserman

article2010NeurIPS1,372 citations

Develops a supervised object counting framework that bypasses individual instance detection by learning an image density function using dot annotations, optimized with a maximum subarray-based loss for accurate and fast count estimation across image subregions.

Listen

Visual object counting in digital imagery is critical for practical tasks such as biomedical cell analysis, crowd surveillance, wildlife tracking, and forestry surveys. Existing computer vision solutions face severe practical trade-offs. Standard detection methods struggle when objects overlap and require expensive, detailed bounding-box or segmentation annotations. Conversely, regression methods that predict total counts from global image statistics require large volumes of fully labeled images and discard spatial information.

The article demonstrates a supervised learning framework that counts objects by estimating continuous image density maps trained on dot annotations (a single dot per object instance). It evaluates this framework to demonstrate that inferring a regional density function delivers superior counting accuracy compared to existing detection, regression, and application-specific baselines.

The approach converts dot annotations into ground-truth density functions and trains a linear model over image features using a regularized risk framework. To evaluate errors effectively, the authors introduced the Maximum Excess over SubArrays (MESA) distance metric. This metric measures the largest discrepancy between estimated and ground-truth density across all possible rectangular subregions in an image. The optimization problem is solved using a convex quadratic program with cutting-plane methods. The framework was evaluated on two key benchmarks: synthetic fluorescence microscopy datasets of bacterial cell populations (with averages around 171 cells per image) and a real-world 2,000-frame pedestrian surveillance video.

The evaluations yielded several major findings. First, on cell counting, the proposed method substantially outperformed all baselines across all training set sizes; with only one training image, it achieved a mean absolute error of 9.5 cells, compared to 20.8 for detection, 60.4 for kernel regression, and 16.2 for a specialized morphology tool. When trained on 32 images, its error decreased to 3.5 cells. Second, on surveillance footage, the density-based approach delivered mean absolute errors between 1.28 and 2.06 pedestrians across standard test splits, outperforming pure regression methods and matching or exceeding a complex hybrid method that required more detailed supervision. Third, the approach proved robust to local noise and kernel width choices during training, and it introduces almost zero computational overhead during inference beyond feature extraction.

These findings indicate that dot-level spatial supervision offers significant operational benefits. Organizations can drastically lower data labeling costs, as placing a dot per object is faster than drawing bounding boxes or segmenting boundaries. Furthermore, because inference simply involves a linear weighting of extracted image features or random forest paths, the model can run in real time on large datasets, reducing compute infrastructure costs for video surveillance and high-throughput microscopy.

Organizations implementing automated counting pipelines should adopt dot-annotated density estimation where individual object detection is unnecessary or error-prone. For deployment, engineering teams should pair the model with fast feature extractors, such as randomized decision forests, to achieve real-time throughput. To reduce initial training time bottlenecks associated with standard optimization solvers, developers should consider implementing purpose-built solvers tailored to the cutting-plane constraints.

Confidence in these performance results is high for dense stationary and structured video counting problems. However, users should note that the cell experiments relied on synthetic data due to inconsistencies among human annotators on real biological images. In surveillance settings, camera calibration (ground plane depth) was utilized to normalize scale perspective. Deploying this system in novel settings may require validating feature choices and calibration steps under operational conditions.

Lempitsky et al (2010).pdf
Cover for Learning To Count Objects in Images

Abstract

We propose a new supervised learning framework for visual object counting tasks, such as estimating the number of cells in a microscopic image or the number of humans in surveillance video frames. We focus on the practically-attractive case when the training images are annotated with dots (one dot per object).

Our goal is to accurately estimate the count. However, we evade the hard task of learning to detect and localize individual object instances. Instead, we cast the problem as that of estimating an image density whose integral over any image region gives the count of objects within that region. Learning to infer such density can be formulated as a minimization of a regularized risk quadratic cost function. We introduce a new loss function, which is well-suited for such learning, and at the same time can be computed efficiently via a maximum subarray algorithm. The learning can then be posed as a convex quadratic program solvable with cutting-plane optimization.

The proposed framework is very flexible as it can accept any domain-specific visual features. Once trained, our system provides accurate object counts and requires a very small time overhead over the feature extraction step, making it a good candidate for applications involving real-time processing or dealing with huge amount of visual data.

Table of Contents

  • 1 Introduction
  • 1.1 Related work
  • 2 The Framework
  • 2.1 Learning to Count
  • 2.2 The MESA distance
  • 2.3 Optimization
  • 3 Experiments
  • 4 Conclusion
  • References

Knowls

  1. Knowl 1 — Object Counting via Linear Estimation of Spatial Density Functions

    model/method

    The counting task is formulated as estimating a continuous/discrete real-valued density function F:I→RF: I \to \mathbb{R} defined over the pixel grid of an image II. For any image subregion S⊆IS \subseteq I, the estimated count of objects located within SS is obtained by summing the density function over the pixels in that region:

    C^(S)=∑p∈SF(p)\hat{C}(S) = \sum_{p \in S} F(p)

    Each pixel p∈Ip \in I is represented by a feature vector xp∈RKx_p \in \mathbb{R}^K. The density function at pixel pp is modeled as a linear transformation of its feature representation:

    F(p∣w)=wTxpF(p \mid w) = w^T x_p

    where w∈RKw \in \mathbb{R}^K is a learnable parameter vector. Estimating object counts in an unseen test image requires computing the pixel features xpx_p, taking the inner product with ww, and summing the resulting scalar densities over the region of interest. This formulation avoids the discrete classification and non-maximum suppression steps of object detection, while preserving spatial distribution information.

  2. Knowl 2 — Maximum Excess over SubArrays Distance

    definition

    Given an image pixel grid II and two real-valued density functions F1,F2:I→RF_1, F_2: I \to \mathbb{R}, the Maximum Excess over SubArrays (MESA) distance, denoted DMESA(F1,F2)\mathcal{D}_{\mathrm{MESA}}(F_1, F_2), is defined as the maximum absolute difference between the sums of F1F_1 and F2F_2 over all axis-aligned rectangular box subarrays in II:

    DMESA(F1,F2)=max⁡B∈B∣∑p∈BF1(p)−∑p∈BF2(p)∣\mathcal{D}_{\mathrm{MESA}}(F_1, F_2) = \max_{B \in \mathcal{B}} \left| \sum_{p \in B} F_1(p) - \sum_{p \in B} F_2(p) \right|

    where B\mathcal{B} denotes the set of all axis-aligned rectangular box subarrays contained in II.

    Equivalently, the distance can be expressed as:

    DMESA(F1,F2)=max⁡(max⁡B∈B∑p∈B(F1(p)−F2(p)),  max⁡B∈B∑p∈B(F2(p)−F1(p)))\mathcal{D}_{\mathrm{MESA}}(F_1, F_2) = \max \left( \max_{B \in \mathcal{B}} \sum_{p \in B} \big(F_1(p) - F_2(p)\big), \; \max_{B \in \mathcal{B}} \sum_{p \in B} \big(F_2(p) - F_1(p)\big) \right)

    The MESA distance operates as an L∞L_\infty metric across subarray integrals. It satisfies two key properties for visual counting:

    1. It provides an upper bound on the absolute difference between total counts over the full image (since the whole image I∈BI \in \mathcal{B}).
    2. It is robust to zero-mean local perturbations and high-frequency noise (which cancel out over large subarrays) while remaining sensitive to low-frequency shifts in spatial object distributions.
  3. Knowl 3 — Convex Quadratic Program for Density Estimation under MESA Distance

    model/method

    Given NN training images I1,…,INI_1, \dots, I_N, with pixel feature representations xpi∈RKx_p^i \in \mathbb{R}^K and target ground truth densities Fi0:Ii→RF^0_i: I_i \to \mathbb{R}, the parameter vector w∈RKw \in \mathbb{R}^K is learned by minimizing an L2L_2-regularized MESA empirical risk:

    min⁡wwTw+λ∑i=1NDMESA(Fi0(⋅),Fi(⋅∣w))\min_w w^T w + \lambda \sum_{i=1}^N \mathcal{D}_{\mathrm{MESA}}\big(F^0_i(\cdot), F_i(\cdot \mid w)\big)

    where λ>0\lambda > 0 is a scalar regularization hyperparameter. By introducing an auxiliary non-negative slack variable ξi\xi_i for each training image i∈{1,…,N}i \in \{1, \dots, N\}, this optimization is formulated as a convex quadratic program (QP):

    min⁡w,ξ1,…,ξNwTw+λ∑i=1Nξi\min_{w, \xi_1, \dots, \xi_N} w^T w + \lambda \sum_{i=1}^N \xi_i

    subject to the linear constraints:

    ∀i∈{1,…,N},  ∀B∈Bi:ξi≥∑p∈B(Fi0(p)−wTxpi)andξi≥∑p∈B(wTxpi−Fi0(p))\forall i \in \{1, \dots, N\}, \; \forall B \in \mathcal{B}_i: \quad \xi_i \ge \sum_{p \in B} \big(F^0_i(p) - w^T x_p^i\big) \quad \text{and} \quad \xi_i \ge \sum_{p \in B} \big(w^T x_p^i - F^0_i(p)\big)

    where Bi\mathcal{B}_i is the set of all axis-aligned box subarrays of image IiI_i. At the optimal solution (w^,ξ^1,…,ξ^N)(\hat{w}, \hat{\xi}_1, \dots, \hat{\xi}_N), the slack variables satisfy ξ^i=DMESA(Fi0(⋅),Fi(⋅∣w^))\hat{\xi}_i = \mathcal{D}_{\mathrm{MESA}}(F^0_i(\cdot), F_i(\cdot \mid \hat{w})), and w^\hat{w} solves the regularized risk problem.

  4. Knowl 4 — Cutting-Plane Optimization for MESA-Constrained Density Learning

    algorithm

    Because the number of box subarray constraints in the convex quadratic program is combinatorial, the problem is solved using an iterative cutting-plane optimization algorithm. At each iteration, the QP is solved over an active constraint subset, and the most violated constraints are identified by solving 2D maximum subarray problems on the residual density maps in O(∣I∣1.5)O(|I|^{1.5}) time using dynamic programming.

    Input: Training images I1,…,INI_1, \dots, I_N, pixel features xpix_p^i, ground truth densities Fi0F^0_i, regularization parameter λ>0\lambda > 0, tolerance ϵ>0\epsilon > 0
    Output: Optimal weight vector w∈RKw \in \mathbb{R}^K
    Initialize active constraint set C\mathcal{C} with a small random subset of box subarrays
    repeat
        Solve QP over active constraints C\mathcal{C}:
            (w,ξ1,…,ξN)←arg⁡min⁡w,ξ(wTw+λ∑i=1Nξi)(w, \xi_1, \dots, \xi_N) \leftarrow \arg\min_{w, \xi} \left( w^T w + \lambda \sum_{i=1}^N \xi_i \right) subject to constraints in C\mathcal{C}
        violations_found ←\leftarrow false
        for i=1i = 1 to NN do
            Compute residual maps R1(p)=Fi0(p)−wTxpiR_1(p) = F^0_i(p) - w^T x_p^i and R2(p)=wTxpi−Fi0(p)R_2(p) = w^T x_p^i - F^0_i(p)
            Find box B1∗=arg⁡max⁡B∈Bi∑p∈BR1(p)B_1^* = \arg\max_{B \in \mathcal{B}_i} \sum_{p \in B} R_1(p) via 2D max subarray
            Find box B2∗=arg⁡max⁡B∈Bi∑p∈BR2(p)B_2^* = \arg\max_{B \in \mathcal{B}_i} \sum_{p \in B} R_2(p) via 2D max subarray
            if ∑p∈B1∗R1(p)>ξi⋅(1+ϵ)\sum_{p \in B_1^*} R_1(p) > \xi_i \cdot (1 + \epsilon) then
                Add constraint ξi≥∑p∈B1∗(Fi0(p)−wTxpi)\xi_i \ge \sum_{p \in B_1^*} (F^0_i(p) - w^T x_p^i) to C\mathcal{C}
                violations_found ←\leftarrow true
            end if
            if ∑p∈B2∗R2(p)>ξi⋅(1+ϵ)\sum_{p \in B_2^*} R_2(p) > \xi_i \cdot (1 + \epsilon) then
                Add constraint ξi≥∑p∈B2∗(wTxpi−Fi0(p))\xi_i \ge \sum_{p \in B_2^*} (w^T x_p^i - F^0_i(p)) to C\mathcal{C}
                violations_found ←\leftarrow true
            end if
        end for
    until violations_found is false
    return ww
  5. Knowl 5 — Kernel Density Ground Truth Generation from Dot Annotations

    equation

    For a training image IiI_i annotated with a set of C(i)C(i) 2D point coordinates Pi={P1,P2,…,PC(i)}\mathcal{P}_i = \{P_1, P_2, \dots, P_{C(i)}\}, the ground truth continuous density function Fi0(p)F^0_i(p) is defined as a sum of normalized 2D Gaussian kernels centered at each annotated dot:

    ∀p∈Ii,Fi0(p)=∑P∈PiN(p;P,σ2I2×2)\forall p \in I_i, \quad F^0_i(p) = \sum_{P \in \mathcal{P}_i} \mathcal{N}(p; P, \sigma^2 I_{2\times 2})

    where N(p;P,σ2I2×2)\mathcal{N}(p; P, \sigma^2 I_{2\times 2}) is a normalized 2D isotropic Gaussian kernel evaluated at pixel coordinate pp, with mean PP and covariance matrix σ2I2×2\sigma^2 I_{2\times 2}. The standard deviation σ\sigma is a small spatial bandwidth parameter (or σ=0\sigma = 0, reducing Fi0F^0_i to a sum of discrete delta functions).

    When an annotated object is near the image boundary, part of the Gaussian probability mass falls outside the pixel grid IiI_i, causing ∑p∈IiFi0(p)<C(i)\sum_{p \in I_i} F^0_i(p) < C(i). This boundary truncation directly models partially visible boundary objects as fractional object counts.

  6. Knowl 6 — Multi-Modal Random Forest Embedding and Perspective Scaling for Video Counting

    model/method

    For counting objects in video sequences with perspective variation (e.g. pedestrians), pixel feature vectors are extracted by combining multiple modalities through a randomized decision forest:

    1. Primary pixel features are computed at each pixel pp: raw intensity, absolute temporal difference from the previous frame, background subtraction difference (using median filtering for static background estimation), and absolute spatial derivatives in xx and yy.
    2. An ensemble of randomized regression trees (e.g. 5 trees) is trained to map local multi-channel patch appearances to local ground truth densities.
    3. Each pixel pp is mapped to a binary indicator feature vector xp∈{0,1}Kx_p \in \{0, 1\}^K, where KK is the total number of leaf nodes across all trees in the forest. The vector contains 1 at the indices corresponding to the leaf reached in each tree, and 0 elsewhere.
    4. To compensate for perspective scaling, the feature vector xpx_p is multiplied by the squared ground plane depth d(p)2d(p)^2 at pixel location pp:

    xp←d(p)2⋅xpx_p \leftarrow d(p)^2 \cdot x_p

    At test time, the density estimate F(p∣w)=wTxpF(p \mid w) = w^T x_p is computed in real time by routing pixel pp down each tree, retrieving the scalar weight wtw_t assigned to each reached leaf node tt, and summing the weights scaled by d(p)2d(p)^2.

  7. Knowl 7 — Bacterial Cell Counting Accuracy across Training Set Sizes

    data/table

    Mean absolute errors (MAE) on a test set of 100 synthetic fluorescence microscopy images (cell counts ranging from 74 to 317 per image, mean 171±64171 \pm 64) as a function of the number of training images N∈{1,2,4,8,16,32}N \in \{1, 2, 4, 8, 16, 32\} (with an equal number of validation images NN, averaged over 5 random splits):

    Method Validation Metric N=1N=1 N=2N=2 N=4N=4 N=8N=8 N=16N=16 N=32N=32
    Linear ridge regression counting 67.3±25.267.3 \pm 25.2 37.7±14.037.7 \pm 14.0 16.7±3.116.7 \pm 3.1 8.8±1.58.8 \pm 1.5 6.4±0.76.4 \pm 0.7 5.9±0.55.9 \pm 0.5
    Kernel ridge regression counting 60.4±16.560.4 \pm 16.5 38.7±17.038.7 \pm 17.0 18.6±5.018.6 \pm 5.0 10.4±2.510.4 \pm 2.5 6.0±0.86.0 \pm 0.8 5.2±0.35.2 \pm 0.3
    Detection (SVM + NMS) counting 28.0±20.628.0 \pm 20.6 20.8±5.820.8 \pm 5.8 13.6±1.513.6 \pm 1.5 10.2±1.910.2 \pm 1.9 10.4±1.210.4 \pm 1.2 8.5±0.58.5 \pm 0.5
    Detection (SVM + NMS) detection 20.8±3.820.8 \pm 3.8 20.1±5.520.1 \pm 5.5 15.7±2.015.7 \pm 2.0 15.0±4.115.0 \pm 4.1 11.8±3.111.8 \pm 3.1 12.0±0.812.0 \pm 0.8
    Detection + linear correction counting – 22.6±5.322.6 \pm 5.3 16.8±6.516.8 \pm 6.5 6.8±1.26.8 \pm 1.2 6.1±1.66.1 \pm 1.6 4.9±0.54.9 \pm 0.5
    Density learning (proposed) counting 12.7±7.312.7 \pm 7.3 7.8±3.77.8 \pm 3.7 5.0±0.55.0 \pm 0.5 4.6±0.64.6 \pm 0.6 4.2±0.44.2 \pm 0.4 3.6±0.23.6 \pm 0.2
    Density learning (proposed) MESA 9.5±6.19.5 \pm 6.1 6.3±1.26.3 \pm 1.2 4.9±0.64.9 \pm 0.6 4.9±0.74.9 \pm 0.7 3.8±0.23.8 \pm 0.2 3.5±0.23.5 \pm 0.2

    An application-specific baseline using adaptive thresholding and morphological analysis achieved an MAE of 16.216.2. The proposed MESA-based density learning outperforms global regression baselines and object detection baselines across all sample sizes NN, achieving low error (9.59.5) even with a single training image (N=1N=1).

  8. Knowl 8 — Crowd Counting Accuracy on Surveillance Video

    data/table

    Mean absolute errors (MAE) for pedestrian counting evaluated on the 2000-frame UCSD surveillance video benchmark across standard training/test evaluation splits:

    Method maximal downscale upscale minimal dense sparse
    Counting-by-Regression (Kong et al., 2006) 2.072.07 2.662.66 2.782.78 N/A N/A N/A
    Counting-by-Regression (Ryan et al., 2009) 1.801.80 2.342.34 2.522.52 4.464.46 N/A N/A
    Counting-by-Segmentation (Ryan et al., 2009) 1.531.53 1.641.64 1.841.84 1.311.31 N/A N/A
    Density learning (proposed) 1.701.70 1.281.28 1.591.59 2.022.02 1.78±0.391.78 \pm 0.39 2.06±0.592.06 \pm 0.59

    The four historical splits are: 'maximal' (train on frames 600:5:1400), 'downscale' (train on most crowded frames 1205:5:1600), 'upscale' (train on least crowded frames 805:5:1100), and 'minimal' (train on only 10 frames 640:80:1360). 'Dense' and 'sparse' evaluate 5-fold cross-validation chunks with 80 and 10 training frames, respectively. The proposed method outperforms direct regression baselines across all splits and performs competitively with the hybrid segmentation method, while requiring only point dot annotations rather than segmented blob annotations.

Coverage note — None omitted. All key modeling formulations, mathematical definitions, algorithms, feature extraction procedures, and experimental evaluation tables are represented in the knowls.

References

  1. 1.http://www.robots.ox.ac.uk/%7Evgg/research/counting/index.html.
  2. 2.The MOSEK optimization software. http://www.mosek.com/.
  3. 3.N. Ahuja and S. Todorovic. Extracting texels in 2.1d natural textures. ICCV, pp. 1–8, 2007.
  4. 4.S. An, P. Peursum, W. Liu, and S. Venkatesh. Efficient algorithms for subwindow search in object detection and localization. CVPR, pp. 264–271, 2009.
  5. 5.D. Anoraganingrum. Cell segmentation with median filter and mathematical morphology operation. Image Analysis and Processing, International Conference on, 0:1043, 1999.
  6. 6.O. Barinova, V. Lempitsky, and P. Kohli. On the detection of multiple object instances using Hough transforms. CVPR, 2010.
  7. 7.J. L. Bentley. Programming pearls: Algorithm design techniques. Comm. ACM, 27(9):865–871, 1984.
  8. 8.J. L. Bentley. Programming pearls: Perspective on performance. Comm. ACM, 27(11):1087–1092, 1984.
  9. 9.L. Breiman. Random forests. Machine Learning, 45(1):5–32, 2001.
  10. 10.A. B. Chan, Z.-S. J. Liang, and N. Vasconcelos. Privacy preserving crowd monitoring: Counting people without people models or tracking. CVPR, 2008.
  11. 11.S.-Y. Cho, T. W. S. Chow, and C.-T. Leung. A neural-based crowd estimation by hybrid global learning algorithm. IEEE Transactions on Systems, Man, and Cybernetics, Part B, 29(4):535–541, 1999.
  12. 12.C. Desai, D. Ramanan, and C. Fowlkes. Discriminative models for multi-class object layout. ICCV, 2009.
  13. 13.X. Descombes, R. Minlos, and E. Zhizhina. Object extraction using a stochastic birth-and-death dynamics in continuum. Journal of Mathematical Imaging and Vision, 33(3):347–359, 2009.
  14. 14.L. Dong, V. Parameswaran, V. Ramesh, and I. Zoghlami. Fast crowd segmentation using shape indexing. ICCV, pp. 1–8, 2007.
  15. 15.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2009 (VOC2009) Results. http://pascallin.ecs.soton.ac.uk/challenges/VOC/voc2009/workshop/index.html.
  16. 16.T. Joachims, T. Finley, and C.-N. J. Yu. Cutting-plane training of structural svms. Machine Learning, 77(1):27–59, 2009.
  17. 17.D. Kong, D. Gray, and H. Tao. A viewpoint invariant approach for crowd counting. ICPR (3), pp. 1187–1190, 2006.
  18. 18.P. D. Kovesi. MATLAB and Octave functions for computer vision and image processing. School of Computer Science & Software Engineering, The University of Western Australia. Available from: http://www.csse.uwa.edu.au/∼pk/research/matlabfns/.
  19. 19.A. Lehmussola, P. Ruusuvuori, J. Selinummi, H. Huttunen, and O. Yli-Harja. Computational framework for simulating fluorescence microscope images with cell populations. IEEE Trans. Med. Imaging, 26(7):1010–1016, 2007.
  20. 20.B. Leibe, A. Leonardis, and B. Schiele. Robust object detection with interleaved categorization and segmentation. International Journal of Computer Vision, 77(1-3):259–289, 2008.
  21. 21.D. G. Lowe. Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision, 60(2):91–110, 2004.
  22. 22.A. N. Marana, S. A. Velastin, L. F. Costa, and R. A. Lotufo. Estimation of crowd density using image processing. Image Processing for Security Applications, pp. 1–8, 1997.
  23. 23.J. Massey, Frank J. The kolmogorov-smirnov test for goodness of fit. Journal of the American Statistical Association, 46(253):68–78, 1951.
  24. 24.F. Moosmann, B. Triggs, and F. Jurie. Fast discriminative visual codebooks using randomized clustering forests. NIPS, pp. 985–992, 2006.
  25. 25.S. K. Nath, K. Palaniappan, and F. Bunyak. Cell segmentation using coupled level sets and graph-vertex coloring. MICCAI (1), pp. 101–108, 2006.
  26. 26.T. W. Nattkemper, H. Wersing, W. Schubert, and H. Ritter. A neural network architecture for automatic segmentation of fluorescence micrographs. Neurocomputing, 48(1-4):357–367, 2002.
  27. 27.V. Rabaud and S. Belongie. Counting crowded moving objects. CVPR (1), pp. 705–711, 2006.
  28. 28.D. Ryan, S. Denman, C. Fookes, and S. Sridharan. Crowd counting using multiple local features. DICTA ’09: Proceedings of the 2009 Digital Image Computing: Techniques and Applications, pp. 81–88, 2009.
  29. 29.J. Selinummi, J. Seppala, O. Yli-Harja, and J. A. Puhakka. Software for quantification of labeled bacteria from digital microscope images by automated image analysis. Biotechniques, 39(6):859–63, 2005.
  30. 30.T. Sharp. Implementing decision trees and forests on a GPU. ECCV (4), pp. 595–608, 2008.
  31. 31.H. Tamaki and T. Tokuyama. Algorithms for the maxium subarray problem based on matrix multiplication. SODA, pp. 446–452, 1998.
  32. 32.A. Vedaldi and B. Fulkerson. VLFeat: An open and portable library of computer vision algorithms. http://www.vlfeat.org/, 2008.
  33. 33.B. Wu, R. Nevatia, and Y. Li. Segmentation of multiple, partially occluded objects by grouping, merging, assigning part detection responses. CVPR, 2008.
  34. 34.T. Zhao and R. Nevatia. Bayesian human segmentation in crowded situations. CVPR (2), pp. 459–466, 2003.

Citation

MLA
Lempitsky, V., and A. Zisserman. “Learning To Count Objects in Images”. Advances in Neural Information Processing Systems, vol. 23, 2010, https://proceedings.neurips.cc/paper_files/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf.
APA
Lempitsky, V., & Zisserman, A. (2010). Learning To Count Objects in Images. Advances in Neural Information Processing Systems, 23. https://proceedings.neurips.cc/paper_files/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf
Chicago
Lempitsky, V., and A. Zisserman. 2010. “Learning To Count Objects in Images”. Advances in Neural Information Processing Systems 23. https://proceedings.neurips.cc/paper_files/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf.
Harvard
Lempitsky, V. and Zisserman, A. (2010) “Learning To Count Objects in Images”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf.
Vancouver
1. Lempitsky V, Zisserman A (2010) Learning To Count Objects in Images. Advances in Neural Information Processing Systems 23:

BibTeX

@inproceedings{lempitsky2010learning,
  title = {Learning To Count Objects in Images},
  author = {Lempitsky, Victor and Zisserman, Andrew},
  year = {2010},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {23},
  url = {https://proceedings.neurips.cc/paper_files/paper/2010/file/fe73f687e5bc5280214e0486b273a5f9-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors