Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors

Jonathan HuangVivek RathodChen SunMenglong ZhuAnoop KorattikaraAlireza FathiIan FischerZbigniew WojnaYang SongSergio Guadarrama

article2016CVPR2,723 citations

Presents a unified empirical evaluation of Faster R-CNN, R-FCN, and SSD across various feature extractors and image resolutions, establishing precise speed, memory, and accuracy trade-offs to help practitioners select optimal object detectors for their deployment constraints.

Listen

Modern convolutional object detectors such as Faster R-CNN, R-FCN, and SSD deliver strong accuracy but differ sharply in speed and memory use. Practitioners face inconsistent published results because prior work employed different feature extractors, image resolutions, hardware, and software stacks. These differences matter for real deployments where mobile devices need small footprints, self-driving cars need real-time performance, and server systems face throughput limits.

The article set out to map the speed/accuracy trade-off curve for these detectors in a unified way. The authors created a single TensorFlow implementation of the three meta-architectures and tested every combination with six feature extractors, two input resolutions, and varying numbers of region proposals. All models were trained and evaluated on the COCO dataset using the official metrics, with timings and memory measured on a consistent GPU platform.

The experiments reveal a clear optimality frontier. At the fast end, SSD paired with MobileNet or Inception V2 at low resolution runs in tens of milliseconds but reaches only 19–22 mAP. In the middle, R-FCN or Faster R-CNN with ResNet-101 and 50–100 proposals achieves roughly 30–32 mAP while remaining practical for many applications. At the accurate end, Faster R-CNN with Inception ResNet at stride 8 reaches 35.7 mAP, the best single-model result reported. SSD accuracy depends less on the strength of the feature extractor than the two-stage methods, and halving image resolution cuts accuracy by about 16 percent on average while reducing inference time by 27 percent. Reducing proposals from 300 to 50 in Faster R-CNN preserves 96 percent of accuracy and cuts runtime by a factor of three.

These findings show that no single detector dominates; the right choice depends on whether the priority is latency, memory, or accuracy. The results also supplied several previously unreported model combinations that, when ensembled, set the state of the art on the 2016 COCO detection challenge. Practitioners can therefore select from the reported points on the frontier rather than retraining many variants.

The main limitations are that all timings include CPU-based post-processing, which caps the fastest models at 25 frames per second, and that a few high-resolution SSD configurations had not fully converged at the time of publication. The study covers only single-model, single-pass inference. Additional measurements on new hardware or with optimized post-processing would increase confidence for production decisions.

No sufficiently relevant recommendations were found.

Cover for Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors

Abstract

The goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detection systems. A number of successful systems have been proposed in recent years, but apples-to-apples comparisons are difficult due to different base feature extractors (e.g., VGG, Residual Networks), different default image resolutions, as well as different hardware and software platforms. We present a unified implementation of the Faster R-CNN [Ren et al., 2015], R-FCN [Dai et al., 2016] and SSD [Liu et al., 2015] systems, which we view as "meta-architectures" and trace out the speed/accuracy trade-off curve created by using alternative feature extractors and varying other critical parameters such as image size within each of these meta-architectures. On one extreme end of this spectrum where speed and memory are critical, we present a detector that achieves real time speeds and can be deployed on a mobile device. On the opposite end in which accuracy is critical, we present a detector that achieves state-of-the-art performance measured on the COCO detection task.

Table of Contents

  • 1 Introduction
  • 2 Meta-architectures
  • 2.1 Meta-architectures
  • 2.2 R-FCN
  • 3 Experimental setup
  • 3.1 Architectural configuration
  • 3.2 Loss function configuration
  • 3.3 Input size configuration.
  • 3.4 Training and hyperparameter tuning
  • 3.5 Benchmarking procedure
  • 3.6 Model Details
  • 3.6.1 Faster R-CNN
  • 3.6.2 R-FCN
  • 3.6.3 SSD
  • 4 Results
  • 4.1 Analyses
  • 4.2 State-of-the-art detection on COCO
  • 4.3 Example detections
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Optimality Frontier and Critical Operating Points for Detection Meta-Architectures

    data/table

    Across modern convolutional object detection systems, the choice of meta-architecture (Single Shot MultiBox Detector [SSD], Faster R-CNN, or Region-based Fully Convolutional Networks [R-FCN]), backbone feature extractor, proposal count, and input resolution forms an empirical Pareto optimality frontier trading mean average precision (mAP) against GPU runtime.

    Model Summary minival mAP test-dev mAP
    (Fastest) SSD w/ MobileNet (Low Resolution) 19.3 18.8
    (Fastest) SSD w/ Inception V2 (Low Resolution) 22.0 21.6
    (Sweet Spot) Faster R-CNN w/ Resnet 101, 100 Proposals 32.0 31.9
    (Sweet Spot) R-FCN w/ Resnet 101, 300 Proposals 30.4 30.3
    (Most Accurate) Faster R-CNN w/ Inception Resnet V2, 300 Proposals 35.7 35.6

    The fastest points on the frontier are occupied by SSD models paired with lightweight backbones (MobileNet and Inception V2), which achieve frame rates suitable for real-time mobile inference but yield lower mAP. The sweet spot balancing speed and accuracy is occupied by ResNet-101 combined with R-FCN (300 proposals) or Faster R-CNN configured with 100 proposals. The highest accuracy single model on the frontier is Faster R-CNN paired with Inception ResNet V2 (operating at stride 8), achieving 35.6% mAP on COCO test-dev at the expense of requiring near 1 second per image.

  2. Knowl 2 — Impact of Proposal Count on Faster R-CNN vs. R-FCN Speed and Accuracy

    empirical result

    For two-stage detectors, reducing the number of region proposals sent from the Region Proposal Network (RPN) to the box classifier dramatically speeds up Faster R-CNN while having minimal latency effect on R-FCN.

    For Faster R-CNN with heavy backbones such as Inception ResNet V2, reducing the number of evaluated proposals from 300 to 50 retains approximately 96% of the mAP (dropping from 35.4% mAP to ~34% mAP) while decreasing GPU running time by a factor of 3. Even dropping down to 10 proposals retains 29% mAP. At 100 proposals, Faster R-CNN with ResNet-101 achieves accuracy and latency comparable to R-FCN with ResNet-101 using 300 proposals.

    In contrast, reducing the number of proposals in R-FCN provides negligible computational savings because R-FCN crops features from the final convolutional layer prior to prediction, executing the expensive feature extraction only once per full image rather than per proposal.

  3. Knowl 3 — Anchor-Based Detection Objective and Box Parameterization

    equation

    In anchor-based convolutional object detection (used across SSD, Faster R-CNN, and R-FCN), training optimizes a combined localization and classification loss over a collection of spatial anchor boxes aa. For an input image II and network parameters θ\theta, the loss for anchor aa is defined as:

    L(a,I;θ)=α⋅1[a is positive]⋅ℓloc(ϕ(ba;a)−floc(I;a,θ))+β⋅ℓcls(ya,fcls(I;a,θ))\mathcal{L}(a, I; \theta) = \alpha \cdot \mathbb{1}[a \text{ is positive}] \cdot \ell_{\text{loc}}(\phi(b_a; a) - f_{\text{loc}}(I; a, \theta)) + \beta \cdot \ell_{\text{cls}}(y_a, f_{\text{cls}}(I; a, \theta))

    where:

    • ya∈{0,1,…,K}y_a \in \{0, 1, \dots, K\} is the groundtruth class label assigned to anchor aa (ya=0y_a = 0 for negative/background anchors, and ya>0y_a > 0 for positive anchors matched via greedy bipartite matching or maximum Jaccard overlap thresholding);
    • 1[⋅]\mathbb{1}[\cdot] is the indicator function;
    • floc(I;a,θ)f_{\text{loc}}(I; a, \theta) and fcls(I;a,θ)f_{\text{cls}}(I; a, \theta) are the predicted box offset encoding and class distribution for anchor aa;
    • α\alpha and β\beta are scalar loss balancing weights;
    • ℓloc\ell_{\text{loc}} is the Smooth L1L_1 (Huber) regression loss;
    • ℓcls\ell_{\text{cls}} is the classification loss (multinomial logistic loss or cross-entropy);
    • ϕ(ba;a)\phi(b_a; a) is the target encoding of the matched groundtruth bounding box ba=[xc,yc,w,h]b_a = [x_c, y_c, w, h] relative to anchor box a=[xc,a,yc,a,wa,ha]a = [x_{c,a}, y_{c,a}, w_a, h_a], parameterized with standard scaling factors as:

    ϕ(ba;a)=[10⋅xc−xc,awa,10⋅yc−yc,aha,5⋅log⁡(wwa),5⋅log⁡(hha)]\phi(b_a; a) = \left[ 10 \cdot \frac{x_c - x_{c,a}}{w_a}, 10 \cdot \frac{y_c - y_{c,a}}{h_a}, 5 \cdot \log\left(\frac{w}{w_a}\right), 5 \cdot \log\left(\frac{h}{h_a}\right) \right]

  4. Knowl 4 — Object Scale Sensitivity Across Detection Meta-Architectures

    empirical result

    Detection performance varies systematically across object sizes depending on meta-architecture:

    • SSD models perform poorly on small objects compared to Faster R-CNN and R-FCN, exhibiting substantially lower small-object AP (APsmallAP_{\text{small}}).
    • On large objects (APlargeAP_{\text{large}}), SSD models are highly competitive with Faster R-CNN and R-FCN, even outperforming two-stage meta-architectures when using lightweight feature extractors (such as MobileNet and Inception V2).
    • Faster R-CNN and R-FCN consistently provide superior localization and recall for small objects due to multi-scale anchor grids operating on higher-resolution intermediate feature maps with subsequent region cropping.
  5. Knowl 5 — Effect of Input Image Resolution on Accuracy and Latency

    empirical result

    Input image resolution has a profound impact across all detector architectures and feature extractors:

    • Halving the spatial image resolution in both dimensions (e.g., from high-resolution shorter edge M=600M = 600 to low-resolution M=300M = 300) decreases detection mAP by an average of 15.88% across evaluated models.
    • This reduction in resolution decreases inference runtime by an average relative factor of 27.4%.
    • The drop in mAP is disproportionately concentrated on small objects, where higher-resolution inputs improve APsmallAP_{\text{small}} by a factor of 2 in many cases.
  6. Knowl 6 — Feature Extractor Classification Strength vs. Detection mAP

    empirical result

    There is an overall positive correlation between a feature extractor's ImageNet top-1 classification accuracy and its final COCO object detection mAP, but the sensitivity depends heavily on the meta-architecture:

    • For Faster R-CNN and R-FCN, detector accuracy strongly depends on the base network classification accuracy; upgrading the feature extractor (e.g., from VGG-16 or Inception V2 to ResNet-101 or Inception ResNet V2) leads to substantial gains in detection mAP.
    • For SSD, detection performance is significantly less sensitive to the classification power of the underlying backbone, showing smaller relative improvements when moving to heavier feature extractors.
  7. Knowl 7 — Feature Map Output Stride Trade-Off in Deep Residual Backbones

    empirical result

    For ResNet-101 and Inception ResNet V2 architectures within Faster R-CNN and R-FCN, reducing the effective output stride from 16 to 8 (achieved by setting the stride of the conv5_1 / conv4_1 blocks to 1 and using atrous/dilated convolutions in subsequent layers):

    • Improves overall COCO mAP by a relative factor of approximately 5% ((mAP8−mAP16)/mAP16=0.05(\text{mAP}_8 - \text{mAP}_{16}) / \text{mAP}_{16} = 0.05).
    • Increases inference wall-clock running time by a relative factor of 63% due to the fourfold increase in spatial feature map area processed by subsequent convolutional layers.
  8. Knowl 8 — Theoretical FLOPs vs. Empirical Hardware Inference Latency

    empirical result

    Floating-point operations (FLOPs / multiply-adds) do not scale linearly with observed wall-clock GPU inference time across different model families:

    • Dense convolutional block models (e.g., ResNet-101) exhibit an empirical FLOPs-to-millisecond ratio greater than 1 on GPUs, benefiting from high memory caching efficiency and optimized cuDNN dense convolution kernels.
    • Models utilizing depthwise separable convolutions or factorized Inception blocks (e.g., MobileNet, Inception V2/V3) exhibit a FLOPs-to-millisecond ratio less than 1, where theoretical FLOP reductions are partially offset by increased memory I/O overhead and non-linear memory access patterns.
  9. Knowl 9 — Linear Correlation Between Detection AP at Variable IoU Thresholds

    empirical result

    Evaluation of detection models on the COCO benchmark demonstrates that performance at individual intersection-over-union (IoU) thresholds is almost perfectly linearly correlated with the standard composite COCO metric mAP@[0.5:0.95]\text{mAP}@[0.5:0.95]:

    • Models that score poorly at higher thresholds (e.g., IoU = 0.75) also score proportionally poorly at lower thresholds (IoU = 0.50).
    • mAP@0.75\text{mAP}@0.75 exhibits an extremely tight linear correlation (R2>0.99R^2 > 0.99) with the average precision across the full range mAP@[0.5:0.95]\text{mAP}@[0.5:0.95], indicating that mAP@0.75\text{mAP}@0.75 alone can serve as an accurate surrogate for the full multi-threshold COCO metric.
  10. Knowl 10 — Diverse Faster R-CNN Ensemble for COCO Benchmark

    data/table

    Combining multiple Faster R-CNN architectures with diversity pruning and multi-crop inference produces state-of-the-art results on the COCO test-challenge dataset.

    Model AP [email protected] [email protected] APsmall_{\text{small}} APmed_{\text{med}} APlarge_{\text{large}} AR@100 ARsmall_{\text{small}} ARlarge_{\text{large}}
    Ours (Diverse Ensemble) 0.413 0.620 0.450 0.231 0.436 0.547 0.604 0.424 0.748
    MSRA 2015 0.371 0.588 0.398 0.173 0.415 0.525 0.489 0.267 0.679
    Trimps-Soushen 0.359 0.580 0.383 0.158 0.407 0.509 0.497 0.269 0.683

    The ensemble consists of five Faster R-CNN models combining ResNet-101 and Inception ResNet V2 feature extractors with varying output strides (8 and 16) and loss ratios. Candidate models were selected greedily on validation data while pruning models whose category-wise average precision vectors had a high cosine similarity to previously selected models. Ensembling plus multi-crop inference improved AP by nearly 7 points over the best single model (from 0.347 to 0.416 on test-dev) and achieved a 60% relative improvement in small object recall (ARsmallAR_{\text{small}}) over MSRA 2015.

Coverage note — Qualitative visual detection examples from Figures 12-17 and fine-grained per-backbone learning rate schedules were omitted as they provide specific implementation artifacts rather than standalone foundational findings.

References

  1. 1.M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al. Tensorflow: Large-scale machine learning on heterogeneous systems, 2015. Software available from tensorflow. org, 1, 2015. 4
  2. 2.S. Bell, C. L. Zitnick, K. Bala, and R. Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. arXiv preprint arXiv:1512.04143, 2015. 3, 5
  3. 3.A. Canziani, A. Paszke, and E. Culurciello. An analysis of deep neural network models for practical applications. arXiv preprint arXiv:1605.07678, 2016. 1
  4. 4.R. Collobert, K. Kavukcuoglu, and C. Farabet. Torch7: A matlab-like environment for machine learning. In BigLearn, NIPS Workshop, number EPFL-CONF-192376, 2011. 4
  5. 5.J. Dai, K. He, and J. Sun. Instance-aware semantic segmentation via multi-task network cascades. arXiv preprint arXiv:1512.04412, 2015. 3, 5, 6
  6. 6.J. Dai, Y. Li, K. He, and J. Sun. R-fcn: Object detection via region-based fully convolutional networks. arXiv preprint arXiv:1605.06409, 2016. 1, 2, 3, 4, 5, 6
  7. 7.J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al. Large scale distributed deep networks. In Advances in neural information processing systems, pages 1223–1231, 2012. 4, 5
  8. 8.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov. Scalable object detection using deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2147–2154, 2014. 2
  9. 9.C.-Y. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg. Dssd: Deconvolutional single shot detector. arXiv preprint arXiv:1701.06659, 2017. 3
  10. 10.R. Girshick. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pages 1440–1448, 2015. 2, 3, 5
  11. 11.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014. 2, 5
  12. 12.K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra. Draw: A recurrent neural network for image generation. In Proceedings of The 32nd International Conference on Machine Learning, pages 1462–1471, 2015. 5
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. arXiv preprint arXiv:1512.03385, 2015. 3, 4, 5, 6, 7, 13
  14. 14.A. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 4, 6, 7
  15. 15.P. J. Huber et al. Robust estimation of a location parameter. The Annals of Mathematical Statistics, 35(1):73–101, 1964. 5
  16. 16.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015. 4, 6, 7
  17. 17.M. Jaderberg, K. Simonyan, A. Zisserman, et al. Spatial transformer networks. In Advances in Neural Information Processing Systems, pages 2017–2025, 2015. 5
  18. 18.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM international conference on Multimedia, pages 675–678. ACM, 2014. 4
  19. 19.K.-H. Kim, S. Hong, B. Roh, Y. Cheon, and M. Park. Pvanet: Deep but lightweight neural networks for real-time object detection. arXiv preprint arXiv:1608.08021, 2016. 3
  20. 20.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 2
  21. 21.S. Lee, S. Purushwalkam, M. Cogswell, D. Crandall, and D. Batra. Why M heads are better than one: Training a diverse ensemble of deep networks. 19 Nov. 2015. 13
  22. 22.Y. Li, H. Qi, J. Dai, X. Ji, and W. Yichen. Translation-aware fully convolutional instance segmentation. https://github.com/daijifeng001/TA-FCN, 2016. 3
  23. 23.T. Y. Lin and P. Dollar. Ms coco api. https://github.com/pdollar/coco, 2016. 5
  24. 24.T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and ´ S. Belongie. Feature pyramid networks for object detection. arXiv preprint arXiv:1612.03144, 2016. 3
  25. 25.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. Lawrence Zitnick. Microsoft ´ COCO: Common objects in context. In ECCV, 1 May 2014. 1, 4
  26. 26.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg. Ssd: Single shot multibox detector. In European Conference on Computer Vision, pages 21–37. Springer, 2016. 1, 2, 3, 4, 5, 6, 7
  27. 27.X. Pan. tfprof: A profiling tool for tensorflow models. https://github.com/tensorflow/tensorflow/tree/master/tensorflow/tools/tfprof, 2016. 6
  28. 28.P. Poirson, P. Ammirato, C.-Y. Fu, W. Liu, J. Kosecka, and A. C. Berg. Fast single shot detection and pose estimation. arXiv preprint arXiv:1609.05590, 2016. 3
  29. 29.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. arXiv preprint arXiv:1506.02640, 2015. 1, 3
  30. 30.J. Redmon and A. Farhadi. Yolo9000: Better, faster, stronger. arXiv preprint arXiv:1612.08242, 2016. 3
  31. 31.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, pages 91–99, 2015. 1, 2, 3, 4, 5, 6
  32. 32.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. 4
  33. 33.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229, 2013. 3
  34. 34.A. Shrivastava and A. Gupta. Contextual priming and feedback for faster r-cnn. In European Conference on Computer Vision, pages 330–348. Springer, 2016. 3
  35. 35.A. Shrivastava, A. Gupta, and R. Girshick. Training region-based object detectors with online hard example mining. arXiv preprint arXiv:1604.03540, 2016. 3
  36. 36.N. Silberman and S. Guadarrama. Tf-slim: A high level library to define complex models in tensorflow. https://research.googleblog.com/2016/08/tf-slim-high-level-library-to-define.html, 2016. [Online; accessed 6-November-2016]. 4
  37. 37.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 4, 6, 7
  38. 38.C. Szegedy, S. Ioffe, and V. Vanhoucke. Inception-v4, inception-resnet and the impact of residual connections on learning. arXiv preprint arXiv:1602.07261, 2016. 4, 6, 7
  39. 39.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–9, 2015. 3
  40. 40.C. Szegedy, S. Reed, D. Erhan, and D. Anguelov. Scalable, high-quality object detection. arXiv preprint arXiv:1412.1441, 2014. 1, 2, 3
  41. 41.C. Szegedy, A. Toshev, and D. Erhan. Deep neural networks for object detection. In Advances in Neural Information Processing Systems, pages 2553–2561, 2013. 2
  42. 42.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. arXiv preprint arXiv:1512.00567, 2015. 4, 6
  43. 43.T. Tieleman and G. Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning, 4(2), 2012. 5
  44. 44.P. Viola and M. J. Jones. Robust real-time face detection. International journal of computer vision, 57(2):137–154, 2004. 2
  45. 45.B. Yang, J. Yan, Z. Lei, and S. Z. Li. Craft objects from images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6043–6051, 2016. 3
  46. 46.S. Zagoruyko, A. Lerer, T.-Y. Lin, P. O. Pinheiro, S. Gross, S. Chintala, and P. Dollar. A multipath network for object ´ detection. arXiv preprint arXiv:1604.02135, 2016. 3
  47. 47.A. Zhai, D. Kislyuk, Y. Jing, M. Feng, E. Tzeng, J. Donahue, Y. L. Du, and T. Darrell. Visual discovery at pinterest. arXiv preprint arXiv:1702.04680, 2017. 3

Citation

MLA
Huang, J., et al. “Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3296–97, https://doi.org/10.1109/CVPR.2017.351.
APA
Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., Guadarrama, S., & Murphy, K. (2017). Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3296–3297. https://doi.org/10.1109/CVPR.2017.351
Chicago
Huang, J., V. Rathod, C. Sun, et al. 2017. “Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3296–97. https://doi.org/10.1109/CVPR.2017.351.
Harvard
Huang, J. et al. (2017) “Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 3296–3297. Available at: https://doi.org/10.1109/CVPR.2017.351.
Vancouver
1. Huang J, Rathod V, Sun C, et al (2017) Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3296–3297

BibTeX

@inproceedings{Huang_2017, title={Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors}, url={http://dx.doi.org/10.1109/CVPR.2017.351}, DOI={10.1109/cvpr.2017.351}, booktitle={2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Huang, Jonathan and Rathod, Vivek and Sun, Chen and Zhu, Menglong and Korattikara, Anoop and Fathi, Alireza and Fischer, Ian and Wojna, Zbigniew and Song, Yang and Guadarrama, Sergio and Murphy, Kevin}, year={2017}, month=July, pages={3296–3297} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE