Deep Neural Networks for Object Detection

Christian SzegedyAlexander ToshevD. Erhan

article2013NeurIPS1,559 citations

Formulates object detection as a neural network regression task that predicts multi-scale bounding box masks, enabling precise multi-instance localization across diverse object categories using only a few forward passes.

Listen

Accurately identifying and locating objects within digital images is a foundational challenge in computer vision with major implications for automated image analysis. While deep neural networks have achieved breakthrough performance in whole-image classification, adapting them to object detection has traditionally been difficult and computationally expensive due to the need to locate multiple objects of varying sizes without evaluating hundreds of thousands of candidate image regions.

The article demonstrates an effective and computationally efficient method for multi-object detection by formulating localization as a deep neural network regression problem. The objective is to evaluate whether deep networks can directly predict binary object bounding box masks and achieve high localization accuracy with minimal computational overhead.

The researchers developed DetectorNet, a seven-layer convolutional neural network adapted to output binary masks covering full objects and their directional halves (top, bottom, left, and right). To handle multiple objects and varying resolutions, the network applies a multi-scale inference strategy across the full image and a small number of overlapping sub-windows, followed by a refinement step on candidate detections. The approach was evaluated on the standard Pascal VOC 2007 benchmark of approximately 5,000 test images across 20 object classes, following training on roughly 11,000 annotated images from the VOC 2012 dataset.

The findings show that DetectorNet achieved state-of-the-art detection precision, outperforming leading part-based and compositional baselines on 8 out of 20 benchmark classes and matching performance on another. Notably, the system proved exceptionally strong on non-rigid and deformable categories, such as dogs (0.282 average precision versus 0.088 for standard part-based models), cats, birds, and sheep, while remaining competitive on rigid objects like cars and buses. Computationally, the multi-scale approach requires evaluating only about 120 image crops per class—taking 5 to 6 seconds per image on a 12-core machine—compared to approximately 150,000 evaluations required by traditional sliding-window deep network baselines. Additionally, the secondary refinement stage significantly boosted precision by re-evaluating enlarged candidate boxes at higher resolution.

These results demonstrate that deep convolutional networks inherently preserve rich geometric and spatial information despite their translation invariance, eliminating the need to manually engineer complex part-based models. This architectural simplicity reduces engineering overhead and improves detection robustness across diverse object categories. However, because the current implementation trains separate networks for each object category and mask type, training resource requirements remain significant.

Moving forward, technical teams should explore shared multi-class network architectures where a single model detects multiple object categories simultaneously to reduce training and deployment costs. Stakeholders should note that the system's current inference speed of several seconds per image makes it suitable for batch image indexing but will require further acceleration before deployment in hard real-time environments. In addition, decision-makers should account for performance variations caused by visual ambiguities, such as cropped objects or visually similar classes, when planning pilot implementations.

Cover for Deep Neural Networks for Object Detection

Abstract

Deep Neural Networks (DNNs) have recently shown outstanding performance on image classification tasks [14]. In this paper we go one step further and address the problem of object detection using DNNs, that is not only classifying but also precisely localizing objects of various classes. We present a simple and yet powerful formulation of object detection as a regression problem to object bounding box masks. We define a multi-scale inference procedure which is able to produce high-resolution object detections at a low cost by a few network applications. State-of-the-art performance of the approach is shown on Pascal VOC.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 DNN-based Detection
  • 4 Detection as DNN Regression
  • 5 Precise Object Localization via DNN-generated Masks
  • 5.1 Multiple Masks for Robust Localization
  • 5.2 Object Localization from DNN Output
  • 5.3 Multi-scale Refinement of DNN Localizer
  • 6 DNN Training
  • 7 Experiments
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Object Detection via Deep Neural Network Mask Regression

    model/method

    The detection of objects within an image is formulated as a regression problem that outputs a spatial binary mask corresponding to the bounding boxes of target objects.

    The architecture is based on a convolutional neural network with 7 layers (5 convolutional layers followed by 2 fully connected layers) using Rectified Linear Units (ReLU) as activation functions and max-pooling after three of the convolutional layers, operating on an input receptive field of 225×225225 \times 225 pixels. The standard classification softmax output layer is replaced by a linear regression layer that outputs a vector of fixed dimension N=d×dN = d \times d, with d=24d = 24 (yielding N=576N = 576 output cells), representing a spatial mask DNN(x;Θ)∈RN\text{DNN}(x; \Theta) \in \mathbb{R}^N parameterized by network weights Θ\Theta.

    For an input image xx of dimensions d1×d2d_1 \times d_2, the output grid is mapped onto the image coordinate space such that each output cell (i,j)(i, j) predicts the presence of an object within the image rectangle T(i,j)T(i, j) with top-left corner at (d1d(i−1),d2d(j−1))\left(\frac{d_1}{d}(i-1), \frac{d_2}{d}(j-1)\right) and dimensions d1d×d2d\frac{d_1}{d} \times \frac{d_2}{d}. If an image pixel falls inside the bounding box of an object of the target class, the corresponding target value is set to 1, and 0 otherwise.

  2. Knowl 2 — Weighted L2 Loss for Mask Regression

    equation

    To optimize the network parameters Θ\Theta for predicting an object mask m∈[0,1]Nm \in [0, 1]^N (N=d×dN = d \times d) from an input image xx, a weighted L2L_2 regression loss over a training dataset DD is minimized:

    min⁡Θ∑(x,m)∈D∥(Diag(m)+λI)1/2(DNN(x;Θ)−m)∥22\min_{\Theta} \sum_{(x, m) \in D} \|(\text{Diag}(m) + \lambda I)^{1/2} (\text{DNN}(x; \Theta) - m)\|_2^2

    where:

    • DD is the training dataset consisting of pairs (x,m)(x, m) of input image crops xx and ground-truth mask targets m∈[0,1]Nm \in [0, 1]^N.
    • DNN(x;Θ)∈RN\text{DNN}(x; \Theta) \in \mathbb{R}^N is the predicted output mask from the network parameterized by weights Θ\Theta.
    • Diag(m)∈RN×N\text{Diag}(m) \in \mathbb{R}^{N \times N} is a diagonal matrix containing the ground-truth mask vector mm along its diagonal.
    • I∈RN×NI \in \mathbb{R}^{N \times N} is the identity matrix.
    • λ∈R+\lambda \in \mathbb{R}^+ is a positive regularization hyperparameter.

    When λ\lambda is chosen to be small, errors on background output cells (where m(i,j)=0m(i, j) = 0, having weight λ1/2\lambda^{1/2}) are penalized significantly less than errors on foreground cells containing an object (where m(i,j)=1m(i, j) = 1, having weight (1+λ)1/2(1 + \lambda)^{1/2}). This weighting prevents the non-convex optimization from collapsing to the trivial all-zero prediction when objects occupy a small fraction of the image area.

  3. Knowl 3 — Multi-Part Mask Decomposition for Adjacent Object Disambiguation

    model/method

    To disambiguate touching or closely placed instances of the same object category that would otherwise merge into a single connected component in a full mask, the detection framework generates five distinct masks per bounding box bbbb:

    1. mfullm^{\text{full}}: covering the full object bounding box bb(full)bb(\text{full}).
    2. mbottomm^{\text{bottom}}: covering the bottom half of the bounding box bb(bottom)bb(\text{bottom}).
    3. mtopm^{\text{top}}: covering the top half of the bounding box bb(top)bb(\text{top}).
    4. mleftm^{\text{left}}: covering the left half of the bounding box bb(left)bb(\text{left}).
    5. mrightm^{\text{right}}: covering the right half of the bounding box bb(right)bb(\text{right}).

    For each mask type h∈{full,bottom,top,left,right}h \in \{\text{full}, \text{bottom}, \text{top}, \text{left}, \text{right}\}, the ground-truth continuous target value mh(i,j;bb)∈[0,1]m^h(i, j; bb) \in [0, 1] assigned to the network output cell (i,j)(i, j) is computed as the fractional overlap between the corresponding sub-box bb(h)bb(h) and the image cell T(i,j)T(i, j):

    mh(i,j;bb)=area(bb(h)∩T(i,j))area(T(i,j))m^h(i, j; bb) = \frac{\text{area}(bb(h) \cap T(i, j))}{\text{area}(T(i, j))}

    where T(i,j)T(i, j) is the rectangular image region mapped to output cell (i,j)(i, j) with size d1d×d2d\frac{d_1}{d} \times \frac{d_2}{d} for an image of dimensions d1×d2d_1 \times d_2 and mask grid resolution d×dd \times d.

    Even when two adjacent objects merge in the full mask mfullm^{\text{full}}, their corresponding directional half masks (such as left and right halves) remain separated, allowing individual instances to be distinguished.

  4. Knowl 4 — Bounding Box Agreement Scoring Across Complementary Masks

    equation

    To evaluate how well a candidate bounding box bbbb matches a predicted mask mm, an agreement score S(bb,m)S(bb, m) measures the average mask activation within the candidate box:

    S(bb,m)=1area(bb)∑(i,j)m(i,j) area(bb∩T(i,j))S(bb, m) = \frac{1}{\text{area}(bb)} \sum_{(i,j)} m(i, j) \, \text{area}(bb \cap T(i, j))

    where T(i,j)T(i, j) is the image patch corresponding to network output cell (i,j)(i, j), and m(i,j)m(i, j) is the predicted mask value at (i,j)(i, j).

    To score candidate bounding boxes across all five predicted masks mhm^h for h∈halves={full,bottom,top,left,right}h \in \text{halves} = \{\text{full}, \text{bottom}, \text{top}, \text{left}, \text{right}\}, a combined score S(bb)S(bb) incorporates penalties for mask activations in complementary regions:

    S(bb)=∑h∈halves(S(bb(h),mh)−S(bb(hˉ),mh))S(bb) = \sum_{h \in \text{halves}} \left( S(bb(h), m^h) - S(bb(\bar{h}), m^h) \right)

    where:

    • bb(h)bb(h) is the region of bbbb corresponding to part hh (e.g., top half for h=toph = \text{top}).
    • hˉ\bar{h} represents the opposite region: for half masks (bottom, top, left, right), hˉ\bar{h} is the opposite half (e.g., if h=toph = \text{top}, hˉ=bottom\bar{h} = \text{bottom}); for h=fullh = \text{full}, hˉ\bar{h} is an expanded rectangular border region surrounding bbbb that penalizes full-mask activations extending outside the candidate box.

    A high score S(bb)S(bb) indicates that the candidate box is strongly supported by all five masks while exhibiting minimal spillover into opposing or surrounding regions.

  5. Knowl 5 — Multi-Scale Inference and Box Refinement Algorithm

    algorithm

    The multi-scale DNN-based localization and refinement procedure evaluates an image across multiple scales, extracts candidate bounding boxes from the merged masks, and refines the top detections with high-resolution crops.

    Input: Input image xx; mask regression networks DNNh\text{DNN}^h for h∈{full,bottom,top,left,right}h \in \{\text{full}, \text{bottom}, \text{top}, \text{left}, \text{right}\}
    Output: Set of refined detected bounding boxes with confidence scores
    detections←∅\text{detections} \leftarrow \emptyset
    scales←compute 3 scales (full image, half size, quarter size)\text{scales} \leftarrow \text{compute 3 scales (full image, half size, quarter size)}
    for s∈scaless \in \text{scales} do
        windows←generate sliding crops at scale s with 20% area overlap\text{windows} \leftarrow \text{generate sliding crops at scale } s \text{ with 20\% area overlap}
        for w∈windowsw \in \text{windows} do
            for h∈{full,bottom,top,left,right}h \in \{\text{full}, \text{bottom}, \text{top}, \text{left}, \text{right}\} do
                mwh←DNNh(w)m^h_w \leftarrow \text{DNN}^h(w)
            end
        end
        mh←merge crops {mwh:w∈windows} across image using pixel-wise maximumm^h \leftarrow \text{merge crops } \{m^h_w : w \in \text{windows}\} \text{ across image using pixel-wise maximum}
        detectionss←extract top 5 bounding boxes scored against {mh} using Eq. (3)\text{detections}_s \leftarrow \text{extract top 5 bounding boxes scored against } \{m^h\} \text{ using Eq. (3)}
        detections←detections∪detectionss\text{detections} \leftarrow \text{detections} \cup \text{detections}_s
    end
    refined←∅\text{refined} \leftarrow \emptyset
    for d∈detectionsd \in \text{detections} do
        c←crop sub-image covering d enlarged by a scale factor of 1.2c \leftarrow \text{crop sub-image covering } d \text{ enlarged by a scale factor of 1.2}
        for h∈{full,bottom,top,left,right}h \in \{\text{full}, \text{bottom}, \text{top}, \text{left}, \text{right}\} do
            mch←DNNh(c)m^h_c \leftarrow \text{DNN}^h(c)
        end
        drefined←infer highest scoring bounding box from {mch}d_{\text{refined}} \leftarrow \text{infer highest scoring bounding box from } \{m^h_c\}
        refined←refined∪{drefined}\text{refined} \leftarrow \text{refined} \cup \{d_{\text{refined}}\}
    end
    return refined\text{refined}

    The algorithm evaluates fewer than 40 window crops across three pyramid scales (image size, 1/21/2 size, and 1/41/4 size). From the 15 candidate boxes (top 5 per scale), each box is enlarged by 1.2×1.2\times, cropped, and re-evaluated through the DNN regression network to achieve precise localized boundaries.

  6. Knowl 6 — Exhaustive Bounding Box Search and Filtering Pipeline

    model/method

    To extract discrete bounding boxes from predicted continuous masks, an exhaustive scoring and filtering pipeline is applied:

    1. Candidate Box Generation: A candidate set of 90 box shapes is constructed from 9 relative mean scale dimensions {0.1,0.2,…,0.9}\{0.1, 0.2, \ldots, 0.9\} of the mean image size combined with 10 aspect ratios obtained via k-means clustering on the bounding boxes in the training set.
    2. Dense Spatial Search: Each candidate box shape is evaluated across the image with a spatial stride of 5 pixels. Using precomputed integral images of the five predicted masks, the agreement score for each box candidate is computed using 5×(2×#pixels+20×#boxes)5 \times (2 \times \#\text{pixels} + 20 \times \#\text{boxes}) total operations.
    3. Score Thresholding: Candidate boxes with a normalized mask agreement score S(bb,mfull)≤0.5S(bb, m^{\text{full}}) \le 0.5 are eliminated.
    4. Classification Pruning: The remaining candidate box crops are passed through a 21-way deep convolutional network classifier (20 object classes plus background). Boxes not classified positively as the detector's target class are discarded.
    5. Non-Maximum Suppression: Standard non-maximum suppression (NMS) is applied to remove overlapping duplicate detections based on Jaccard intersection-over-union similarity.
  7. Knowl 7 — Training Data Sampling and Pretraining Pipeline

    experimental setup

    The training workflow for mask regression and classification pruning networks uses specific sampling and fine-tuning procedures:

    • Mask Regression Training: For each training image, several thousand crops are extracted with crop widths distributed uniformly between a prescribed minimum scale and the image width. The dataset is balanced to 40% positive crops (covering at least 80% of the area of some ground-truth bounding box) and 60% negative crops (having 0 intersection with any ground-truth object bounding box). Total training volume is 10 million samples per class.
    • Pruning Classifier Training: Crops are sampled to 40% positives (Jaccard similarity ≥0.6\ge 0.6 with a ground-truth object box, labeled by the class of the most similar box) and 60% negatives (Jaccard similarity <0.2< 0.2 with all ground-truth boxes, labeled as background). Total training volume is 10 million samples per class.
    • Transfer Learning and Optimization: The convolutional layers are first initialized by pretraining the 7-layer architecture on the ImageNet classification task. All layers including the convolutional layers are then fine-tuned for mask regression using stochastic gradient descent with the AdaGrad adaptive learning rate algorithm.
  8. Knowl 8 — Object Detection Performance on Pascal VOC 2007 Test Set

    data/table

    The DetectorNet approach was evaluated on the Pascal VOC 2007 test set (approximately 5,000 images across 20 classes) after training on the Pascal VOC 2012 training and validation sets (approximately 11,000 images). Performance is measured in Average Precision (AP) per class and compared against a sliding-window deep neural network baseline, the 3-layer compositional model of Zhu et al. (2010), and Deformable Part Models (Felzenszwalb et al., 2010; Girshick et al., 2011).

    Class aero bicycle bird boat bottle bus car cat chair cow
    DetectorNet .292 .352 .194 .167 .037 .532 .502 .272 .102 .348
    Sliding windows .213 .190 .068 .120 .058 .294 .237 .101 .059 .131
    3-layer model .294 .558 .094 .143 .286 .440 .513 .213 .200 .193
    Felzenszwalb et al. .328 .568 .025 .168 .285 .397 .516 .213 .179 .185
    Girshick et al. .324 .577 .107 .157 .253 .513 .542 .179 .210 .240
    Class table dog horse m-bike person plant sheep sofa train tv
    DetectorNet .302 .282 .466 .417 .262 .103 .328 .268 .398 .470
    Sliding windows .110 .134 .220 .243 .173 .070 .118 .166 .240 .119
    3-layer model .252 .125 .504 .384 .366 .151 .197 .251 .368 .393
    Felzenszwalb et al. .259 .088 .492 .412 .368 .146 .162 .244 .392 .391
    Girshick et al. .257 .116 .556 .475 .435 .145 .226 .342 .442 .413

    DetectorNet achieves top performance on 8 out of 20 classes (bird, bus, cat, cow, table, dog, sheep, tv) and performs on par with state-of-the-art models on horse, sofa, and train. It substantially outperforms DPM baselines on highly deformable categories (e.g., bird .194 vs .107, cat .272 vs .179, dog .282 vs .116, sheep .328 vs .226, cow .348 vs .240) while maintaining strong accuracy on rigid object categories.

  9. Knowl 9 — Effect of Coarse-to-Fine Refinement on Precision-Recall

    empirical result

    Applying the second-stage refinement step—wherein candidate bounding box crops are enlarged by a factor of 1.2 and re-evaluated through the mask regression network—produces a substantial improvement in detection accuracy compared to the initial multi-scale detection stage (Stage 1).

    Across diverse object classes (such as bird, bus, and dining table), precision-recall curves demonstrate that the refinement stage consistently shifts the precision-recall frontier upward across all recall levels. The primary mechanism driving this improvement is that re-evaluating the DNN localizer at higher resolution on tightly cropped windows eliminates boundary quantization errors from coarse grid predictions (24×2424 \times 24 output masks), boosting the agreement scores of accurately localized true positives while suppressing poorly aligned false candidates.

  10. Knowl 10 — Training and Model Complexity Limitations of DetectorNet

    limitation

    The proposed detection framework exhibits several operational and structural limitations:

    1. Training Overhead per Class and Mask: Separate mask regression networks are trained for each object class and for each of the 5 mask types (full, top, bottom, left, right), leading to significant training computation overhead and lack of cross-class parameter sharing.
    2. Ambiguity in Ground-Truth Extents: In standard object detection datasets (such as Pascal VOC), annotations sometimes label partial objects (e.g., only the head of a bird) and other times label the full body. This annotation inconsistency causes the regressor to predict conflicting masks for the same object instance, occasionally leading to duplicate detections of sub-parts (e.g., simultaneous bounding boxes for both bird head and full body).
    3. Confusion on Visually Similar Classes: Mask regression can produce high-confidence false positive masks for visually similar distractor categories when context alone cannot disambiguate fine category features.

Coverage note — No substantial contributed material was omitted; all key architectural components, loss equations, inference algorithms, training setups, empirical results, and failure modes are covered.

References

  1. 1.Narendra Ahuja and Sinisa Todorovic. Learning the taxonomy and models of categories present in arbitrary images. In International Conference on Computer Vision, 2007.
  2. 2.Yoshua Bengio. Learning deep architectures for ai. Foundations and TrendsR in Machine Learning, 2(1):1–127, 2009.
  3. 3.Dan Ciresan, Alessandro Giusti, Juergen Schmidhuber, et al. Deep neural networks segment neuronal membranes in electron microscopy images. In Advances in Neural Information Processing Systems 25, 2012.
  4. 4.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In Computer Vision and Pattern Recognition, 2005.
  5. 5.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Computer Vision and Pattern Recognition, 2009.
  6. 6.John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. In Conference on Learning Theory. ACL, 2010.
  7. 7.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2):303–338, 2010.
  8. 8.Clement Farabet, Camille Couprie, Laurent Najman, and Yann LeCun. Learning hierarchical features for scene labeling. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1915–1929, 2013.
  9. 9.Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan. Object detection with discriminatively trained part-based models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(9):1627–1645, 2010.
  10. 10.Sanja Fidler and Ales Leonardis. Towards scalable representations of object categories: Learning a hierarchy of parts. In Computer Vision and Pattern Recognition, 2007.
  11. 11.R. B. Girshick, P. F. Felzenszwalb, and D. McAllester. Discriminatively trained deformable part models, release 5. http://people.cs.uchicago.edu/ rbg/latent-release5/.
  12. 12.Geoffrey E Hinton and Ruslan R Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504–507, 2006.
  13. 13.Iasonas Kokkinos and Alan Yuille. Inference and learning with hierarchical shape models. International Journal of Computer Vision, 93(2):201–225, 2011.
  14. 14.Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25, 2012.
  15. 15.Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng. Building high-level features using large scale unsupervised learning. In International Conference on Machine Learning, 2012.
  16. 16.Yann LeCun and Yoshua Bengio. Convolutional networks for images, speech, and time series. The handbook of brain theory and neural networks, 1995.
  17. 17.Jorge Sanchez and Florent Perronnin. High-dimensional signature compression for large-scale image classification. In Computer Vision and Pattern Recognition, 2011.
  18. 18.Hannes Schulz and Sven Behnke. Object-class segmentation using deep convolutional neural networks. In Proceedings of the DAGM Workshop on New Challenges in Neural Computation, 2011.
  19. 19.Long Zhu, Yuanhao Chen, Alan Yuille, and William Freeman. Latent hierarchical structural learning for object detection. In Computer Vision and Pattern Recognition, 2010.
  20. 20.Song Chun Zhu and David Mumford. A stochastic grammar of images. Computer Graphics and Vision, 2(4):259–362, 2007.

Citation

MLA
Szegedy, C., et al. “Deep Neural Networks for Object Detection”. Advances in Neural Information Processing Systems, vol. 26, 2013, https://proceedings.neurips.cc/paper_files/paper/2013/file/f7cade80b7cc92b991cf4d2806d6bd78-Paper.pdf.
APA
Szegedy, C., Toshev, A., & Erhan, D. (2013). Deep Neural Networks for Object Detection. Advances in Neural Information Processing Systems, 26. https://proceedings.neurips.cc/paper_files/paper/2013/file/f7cade80b7cc92b991cf4d2806d6bd78-Paper.pdf
Chicago
Szegedy, C., A. Toshev, and D. Erhan. 2013. “Deep Neural Networks for Object Detection”. Advances in Neural Information Processing Systems 26. https://proceedings.neurips.cc/paper_files/paper/2013/file/f7cade80b7cc92b991cf4d2806d6bd78-Paper.pdf.
Harvard
Szegedy, C., Toshev, A. and Erhan, D. (2013) “Deep Neural Networks for Object Detection”, Advances in Neural Information Processing Systems. Curran Associates, Inc. Available at: https://proceedings.neurips.cc/paper_files/paper/2013/file/f7cade80b7cc92b991cf4d2806d6bd78-Paper.pdf.
Vancouver
1. Szegedy C, Toshev A, Erhan D (2013) Deep Neural Networks for Object Detection. Advances in Neural Information Processing Systems 26:

BibTeX

@inproceedings{szegedy2013deep,
  title = {Deep Neural Networks for Object Detection},
  author = {Szegedy, Christian and Toshev, Alexander and Erhan, Dumitru},
  year = {2013},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {26},
  url = {https://proceedings.neurips.cc/paper_files/paper/2013/file/f7cade80b7cc92b991cf4d2806d6bd78-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission