Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors
Jonathan HuangVivek RathodChen SunMenglong ZhuAnoop KorattikaraAlireza FathiIan FischerZbigniew WojnaYang SongSergio Guadarrama
Presents a unified empirical evaluation of Faster R-CNN, R-FCN, and SSD across various feature extractors and image resolutions, establishing precise speed, memory, and accuracy trade-offs to help practitioners select optimal object detectors for their deployment constraints.
Modern convolutional object detectors such as Faster R-CNN, R-FCN, and SSD deliver strong accuracy but differ sharply in speed and memory use. Practitioners face inconsistent published results because prior work employed different feature extractors, image resolutions, hardware, and software stacks. These differences matter for real deployments where mobile devices need small footprints, self-driving cars need real-time performance, and server systems face throughput limits.
The article set out to map the speed/accuracy trade-off curve for these detectors in a unified way. The authors created a single TensorFlow implementation of the three meta-architectures and tested every combination with six feature extractors, two input resolutions, and varying numbers of region proposals. All models were trained and evaluated on the COCO dataset using the official metrics, with timings and memory measured on a consistent GPU platform.
The experiments reveal a clear optimality frontier. At the fast end, SSD paired with MobileNet or Inception V2 at low resolution runs in tens of milliseconds but reaches only 19–22 mAP. In the middle, R-FCN or Faster R-CNN with ResNet-101 and 50–100 proposals achieves roughly 30–32 mAP while remaining practical for many applications. At the accurate end, Faster R-CNN with Inception ResNet at stride 8 reaches 35.7 mAP, the best single-model result reported. SSD accuracy depends less on the strength of the feature extractor than the two-stage methods, and halving image resolution cuts accuracy by about 16 percent on average while reducing inference time by 27 percent. Reducing proposals from 300 to 50 in Faster R-CNN preserves 96 percent of accuracy and cuts runtime by a factor of three.
These findings show that no single detector dominates; the right choice depends on whether the priority is latency, memory, or accuracy. The results also supplied several previously unreported model combinations that, when ensembled, set the state of the art on the 2016 COCO detection challenge. Practitioners can therefore select from the reported points on the frontier rather than retraining many variants.
The main limitations are that all timings include CPU-based post-processing, which caps the fastest models at 25 frames per second, and that a few high-resolution SSD configurations had not fully converged at the time of publication. The study covers only single-model, single-pass inference. Additional measurements on new hardware or with optimized post-processing would increase confidence for production decisions.
- Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Read Faster R-CNN first to understand the region-proposal detector whose speed, accuracy, and proposal-count trade-offs the source benchmarks.
- Paper: SSD: Single Shot MultiBox Detector, Wei Liu et al. (2015). SSD establishes the single-shot architecture that the source measures against two-stage detectors across backbones and input resolutions.
- Paper: R-FCN: Object Detection via Region-based Fully Convolutional Networks, Jifeng Dai et al. (2016). R-FCN explains the fully convolutional, position-sensitive detector that the source compares with Faster R-CNN and SSD.
- Paper: Fast R-CNN, Ross Girshick (2015). Fast R-CNN supplies the region-based detection and RoI-pooling foundations needed to follow the Faster R-CNN lineage evaluated in the source.
- Paper: Rich feature hierarchies for accurate object detection and semantic segmentation, Ross Girshick et al. (2014). R-CNN introduces the convolutional region-based detection pipeline that makes the later Fast and Faster R-CNN designs in the benchmark easier to understand.
No sufficiently relevant recommendations were found.
