Single-Shot Refinement Neural Network for Object Detection

Shifeng ZhangLongyin WenXiao BianZhen LeiStan Z. Li

article2017CVPR1,377 citations

Proposes RefineDet, an object detector that combines the high accuracy of two-stage methods with the fast inference of single-stage models by using anchor refinement and feature transfer modules to filter false positives and optimize bounding boxes before final classification.

Listen

Modern computer vision systems rely heavily on automated object detection, which traditionally forces a difficult engineering compromise. Two-stage detection frameworks offer high accuracy by generating candidate regions before classifying them, but they suffer from high computational latency. Conversely, one-stage frameworks prioritize real-time processing speed by densely scanning images, but they typically achieve lower precision due to foreground-background class imbalance and inaccurate bounding box placement.

The article introduces and evaluates RefineDet, a single-shot object detection framework designed to combine the high operational speed of one-stage methods with the superior precision of two-stage architectures.

To achieve this, the authors designed a deep learning architecture with two interconnected stages: an Anchor Refinement Module that filters out obvious background regions and coarsely adjusts initial reference boxes, and an Object Detection Module that fine-tunes object boundaries and predicts final multi-class categories. Feature communication between these stages is managed by a transfer connection block that integrates contextual visual information. The framework was trained end-to-end and rigorously benchmarked against leading models using standard image datasets, including PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO, measuring both mean average precision and frames per second on standard hardware.

Key findings show that RefineDet established new performance benchmarks across all standard datasets. On the MS COCO benchmark, RefineDet achieved 41.8% average precision, outperforming both existing one-stage detectors and complex multi-model ensemble systems. On PASCAL VOC 2007, the system achieved 80.0% to 81.8% precision under standard single-scale evaluation, and reached up to 85.8% precision with multi-scale testing and dataset pretraining, ranking among top global benchmarks. In terms of processing speed, the system maintained real-time throughput, processing images at 40.3 frames per second for 320x320 inputs and 24.1 frames per second for 512x512 inputs on a single graphics processing unit. Ablation testing confirmed that the two-step cascaded regression provided the largest single performance benefit, accounting for a 2.2% precision increase, while negative anchor filtering and feature transfer blocks added further measurable gains.

These results demonstrate that organizations deploying computer vision systems no longer need to sacrifice detection accuracy to meet real-time operational constraints. By eliminating the latency bottleneck of two-stage models while maintaining top-tier precision, the approach reduces hardware infrastructure requirements and enables high-reliability visual tracking in latency-critical applications such as autonomous navigation, video surveillance, and industrial automation.

For practical deployment, organizations should select input resolutions based on task-specific trade-offs: smaller inputs (320x320) provide real-time 40-frame-per-second capability for high-throughput video streams, whereas larger inputs (512x512) maximize precision for complex scenes. Next development steps should focus on adapting the framework for dedicated object classes—such as pedestrians, vehicles, and faces—and integrating complementary focal loss or visual attention mechanisms to further boost performance.

While confidence in the results is high across standardized benchmarks, detection accuracy remains comparatively lower for small, densely clustered objects, such as distant furniture. Increasing image resolution mitigates this limitation but introduces processing overhead, indicating that further algorithmic refinement is required for edge cases involving tiny objects.

  • Paper: SSD: Single Shot MultiBox Detector, Wei Liu et al. (2015). RefineDet directly builds upon the single-shot multi-scale default box paradigm established by SSD, extending it with anchor refinement modules.
  • Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Faster R-CNN introduced the Region Proposal Network and anchor-based two-stage detection principles that motivate RefineDet's two-step anchor refinement strategy.
  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Feature Pyramid Networks supply the foundational lateral connection and top-down feature fusion mechanics adapted in RefineDet's transfer connection blocks.
  • Paper: Fast R-CNN, Ross B. Girshick (2015). Fast R-CNN formulated the joint multi-task loss for classification and bounding-box regression that underpins RefineDet's end-to-end optimization.
  • Paper: You Only Look Once: Unified, Real-Time Object Detection, Joseph Redmon et al. (2016). YOLO pioneered real-time single-stage object detection that RefineDet aims to improve with higher localization accuracy.
Cover for Single-Shot Refinement Neural Network for Object Detection

Abstract

For object detection, the two-stage approach (e.g., Faster R-CNN) has been achieving the highest accuracy, whereas the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, in this paper, we propose a novel single-shot based detector, called RefineDet, that achieves better accuracy than two-stage methods and maintains comparable efficiency of one-stage methods. RefineDet consists of two inter-connected modules, namely, the anchor refinement module and the object detection module. Specifically, the former aims to (1) filter out negative anchors to reduce search space for the classifier, and (2) coarsely adjust the locations and sizes of anchors to provide better initialization for the subsequent regressor. The latter module takes the refined anchors as the input from the former to further improve the regression and predict multi-class label. Meanwhile, we design a transfer connection block to transfer the features in the anchor refinement module to predict locations, sizes and class labels of objects in the object detection module. The multi-task loss function enables us to train the whole network in an end-to-end way. Extensive experiments on PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO demonstrate that RefineDet achieves state-of-the-art detection accuracy with high efficiency. Code is available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Network Architecture
  • 4 Training and Inference
  • 5 Experiments
  • 5.1 PASCAL VOC 2007
  • 5.1.1 Run Time Performance
  • 5.1.2 Ablation Study
  • 5.2 PASCAL VOC 2012
  • 5.3 MS COCO
  • 5.4 From MS COCO to PASCAL VOC
  • 6 Conclusions
  • References
  • 7 Complete Object Detection Results
  • 8 Qualitative Results
  • 9 Detection Analysis on PASCAL VOC 2007

Knowls

  1. Knowl 1 — RefineDet Dual-Module Object Detection Architecture

    model/method

    RefineDet is a single-shot object detection network designed to combine the efficiency of one-stage detectors with the accuracy advantages of two-stage detectors (handling class imbalance, cascaded regression, and two-stage feature representations). The network consists of two inter-connected modules:

    1. Anchor Refinement Module (ARM): Constructed by truncating a base classification network (e.g., VGG-16 or ResNet-101 pretrained on ImageNet) and adding auxiliary convolutional layers. The ARM filters out easy negative anchors to drastically reduce the candidate search space and performs coarse bounding box regression to provide adjusted anchor locations and scales for the subsequent stage.
    2. Object Detection Module (ODM): Formed by prediction convolutional layers (3×33 \times 3 kernels) built on top of Transfer Connection Blocks (TCBs). The ODM takes the refined bounding boxes generated by the ARM as input anchors rather than regularly spaced default boxes, predicting multi-class categorical probabilities and fine bounding box offsets relative to these refined anchors.
  2. Knowl 2 — Transfer Connection Block

    model/method

    The Transfer Connection Block (TCB) links the Anchor Refinement Module (ARM) to the Object Detection Module (ODM). It serves two functions: converting feature representations from the ARM to the format required by the ODM for multi-task prediction, and incorporating high-level top-down contextual information.

    For each feature map level associated with anchors (e.g., layers with total stride sizes 8, 16, 32, and 64 pixels):

    • The feature map from the ARM passes through a sequence of two 3×33 \times 3 convolutional layers (stride 1, 256 channels) with ReLU activations.
    • The high-level feature output from the adjacent deeper TCB is upsampled via a 4×44 \times 4 deconvolutional layer (stride 2, 256 channels) to match spatial dimensions.
    • The upsampled deeper features and the current-level transformed features are combined using element-wise summation.
    • A subsequent 3×33 \times 3 convolution (stride 1, 256 channels) followed by ReLU is applied to the sum to ensure discriminability before feeding the resulting feature map into the ODM prediction layers.
  3. Knowl 3 — Two-Step Cascaded Bounding Box Regression

    model/method

    RefineDet performs bounding box regression in a two-stage cascade across multiple feature scales:

    1. Step 1 (Coarse Regression in ARM): At each feature cell across four detection layers (with strides 8, 16, 32, and 64 pixels), nn regular tiled anchors are defined. The ARM predicts 4 regression offsets relative to the original tiled anchor coordinates alongside 2 foreground/background confidence scores. This yields nn refined anchor boxes per cell with adjusted centers, widths, and heights.
    2. Step 2 (Fine Regression in ODM): The ODM receives the refined anchor box coordinates from the ARM. For each refined anchor box, the ODM computes 4 fine offsets relative to the refined box boundaries and cc multi-class classification scores (yielding c+4c + 4 outputs per refined anchor). The final predicted object box is computed by applying the fine offsets to the coarse, refined anchor box geometry.
  4. Knowl 4 — Negative Anchor Filtering Mechanism

    algorithm

    The negative anchor filtering mechanism suppresses well-classified background anchors to mitigate extreme class imbalance between foreground objects and background.

    Input: Set of regularly tiled anchors A, ARM negative confidence scores P_neg, confidence threshold theta = 0.99, phase in {TRAINING, INFERENCE}
    Output: Filtered set of refined anchors A_refined to be processed by ODM
    for each anchor a_i in A with ARM prediction p_neg_i do
        if p_neg_i > theta then
            Discard anchor a_i
        else
            Compute refined bounding box coordinates for a_i using ARM regression offsets
            Add refined anchor to A_refined
        end if
    end for
    if phase == TRAINING then
        Pass A_refined (containing refined positive anchors and hard negative anchors) to train the ODM
    else if phase == INFERENCE then
        Pass A_refined to the ODM to predict multi-class probabilities and fine bounding box offsets
    end if
  5. Knowl 5 — Multi-Task Loss Formulation for RefineDet

    equation

    RefineDet is trained end-to-end using a joint multi-task loss function combining the Anchor Refinement Module (ARM) loss and the Object Detection Module (ODM) loss:

    L({pi},{xi},{ci},{ti})=1Narm(∑iLb(pi,[li∗≥1])+∑i[li∗≥1]Lr(xi,gi∗))+1Nodm(∑iLm(ci,li∗)+∑i[li∗≥1]Lr(ti,gi∗))\mathcal{L}(\{p_i\}, \{x_i\}, \{c_i\}, \{t_i\}) = \frac{1}{N_{\mathrm{arm}}} \left( \sum_i L_b(p_i, [l_i^* \ge 1]) + \sum_i [l_i^* \ge 1] L_r(x_i, g_i^*) \right) + \frac{1}{N_{\mathrm{odm}}} \left( \sum_i L_m(c_i, l_i^*) + \sum_i [l_i^* \ge 1] L_r(t_i, g_i^*) \right)

    where:

    • ii is the anchor index within a mini-batch.
    • li∗∈{0,1,…,c−1}l_i^* \in \{0, 1, \dots, c-1\} is the ground truth class label for anchor ii, where li∗=0l_i^* = 0 corresponds to background and li∗≥1l_i^* \ge 1 denotes a foreground object category.
    • gi∗g_i^* represents the ground truth bounding box location and size coordinates corresponding to anchor ii.
    • pip_i is the ARM predicted confidence score indicating whether anchor ii is an object or background.
    • xix_i represents the predicted coarse regression coordinate offsets in the ARM for anchor ii.
    • cic_i denotes the predicted multi-class probabilities across all target classes in the ODM for anchor ii.
    • tit_i represents the predicted fine regression coordinate offsets in the ODM relative to the refined anchor box.
    • NarmN_{\mathrm{arm}} and NodmN_{\mathrm{odm}} are the numbers of positive matched anchors in the ARM and ODM, respectively. If Narm=0N_{\mathrm{arm}} = 0, the ARM loss terms are set to 0; if Nodm=0N_{\mathrm{odm}} = 0, the ODM loss terms are set to 0.
    • Lb(pi,[li∗≥1])L_b(p_i, [l_i^* \ge 1]) is the binary cross-entropy loss over two classes (object vs. non-object).
    • Lm(ci,li∗)L_m(c_i, l_i^*) is the softmax loss over multi-class category predictions.
    • Lr(⋅,⋅)L_r(\cdot, \cdot) is the smooth L1L_1 regression loss.
    • [li∗≥1][l_i^* \ge 1] is the Iverson bracket indicator function, evaluating to 1 if li∗≥1l_i^* \ge 1 (anchor is non-negative) and 0 otherwise, masking out regression loss on negative samples.
  6. Knowl 6 — Anchor Design, Matching, and Hard Negative Mining

    model/method

    RefineDet associates anchors with four multi-stride convolutional feature layers (strides 8, 16, 32, and 64 pixels):

    • For VGG-16, the layers are conv4_3, conv5_3, conv_fc7, and conv6_2. L2L_2 normalization is applied to scale conv4_3 and conv5_3 feature norms to 10 and 8, respectively.
    • For ResNet-101, the layers are res3b3, res4b22, res5c, and res6.

    Anchor Scales and Aspect Ratios: Each layer is assigned one base anchor scale equal to 4×4 \times the total layer stride (32, 64, 128, and 256 pixels) and 3 aspect ratios (0.5,1.0,2.00.5, 1.0, 2.0), yielding n=3n = 3 anchor boxes per cell.

    Ground Truth Matching: Each ground truth box is first matched to the anchor box with the highest Jaccard overlap (IoU). Next, anchor boxes are matched to any ground truth box with Jaccard overlap higher than 0.5.

    Hard Negative Mining: After matching, the negative anchors with the highest loss values are selected to maintain a negative-to-positive ratio capped at 3:13:1 during ODM training.

  7. Knowl 7 — RefineDet Inference Procedure

    algorithm

    The inference pipeline processes an input image through the feed-forward network to produce the final detections:

    Input: Input image I, trained RefineDet network (ARM, TCBs, ODM), negative threshold theta = 0.99, NMS IoU threshold = 0.45, candidate count K_arm = 400, max detections per image K_final = 200
    Output: Final detected bounding boxes with class labels and confidence scores
    1. Feed image I through the ARM to extract multi-scale features, binary confidence scores, and coarse bounding box offsets.
    2. For each anchor, compute the negative confidence score.
    3. Discard all anchors with negative confidence score > theta.
    4. For remaining anchors, adjust coordinates using coarse regression offsets to produce refined anchor boxes.
    5. Pass multi-scale ARM features through TCBs to produce transformed context feature maps.
    6. In the ODM, compute multi-class category probabilities and fine coordinate offsets relative to the refined anchor boxes.
    7. Select the top K_arm confident detections per image across all classes.
    8. Apply non-maximum suppression (NMS) independently for each object class using a Jaccard overlap threshold of 0.45.
    9. Retain the top K_final highest-confidence detections per image across all classes as the final output.
  8. Knowl 8 — Component Ablation on PASCAL VOC 2007

    data/table

    An ablation study evaluated the incremental impact of each proposed module in RefineDet using a VGG-16 backbone with 320×320320 \times 320 input images trained on VOC 2007 + VOC 2012 trainval sets and tested on the VOC 2007 test set.

    Negative Anchor Filtering ✓ ✓
    Two-Step Cascaded Regression ✓ ✓ ✓
    Transfer Connection Block (TCB) ✓ ✓ ✓ ✓
    mAP (%) 76.2 77.3 79.5 80.0

    The baseline model (without TCB, two-step regression, or negative anchor filtering, where ARM directly performs multi-class detection like SSD) achieves 76.2% mAP. Adding the TCB improves accuracy by 1.1% (to 77.3%). Adding the two-step cascaded regression yields a 2.2% increase (to 79.5%). Adding negative anchor filtering further improves performance by 0.5% to reach the full RefineDet320 score of 80.0% mAP.

  9. Knowl 9 — Generic Object Detection Performance and Speed on PASCAL VOC

    data/table

    RefineDet was evaluated on the PASCAL VOC 2007 and VOC 2012 benchmarks using a VGG-16 backbone. Inference speed (FPS) was measured with batch size 1 on an NVIDIA Titan X GPU (CUDA 8.0, cuDNN v6).

    Method Input Size #Boxes FPS VOC 2007 mAP (%) VOC 2012 mAP (%)
    Faster R-CNN (VGG-16) ∼1000×600\sim 1000 \times 600 300 7 73.2 70.4
    R-FCN (ResNet-101) ∼1000×600\sim 1000 \times 600 300 9 80.5 77.6
    SSD300* (VGG-16) 300×300300 \times 300 8732 46 77.2 75.8
    SSD512* (VGG-16) 512×512512 \times 512 24564 19 79.8 78.5
    DSSD513 (ResNet-101) 513×513513 \times 513 43688 5.5 81.5 80.0
    RefineDet320 320×320320 \times 320 6375 40.3 80.0 78.1
    RefineDet512 512×512512 \times 512 16320 24.1 81.8 80.1
    RefineDet320+ (Multi-scale) - - - 83.1 82.7
    RefineDet512+ (Multi-scale) - - - 83.8 83.5
    RefineDet320 (COCO pretrain) 320×320320 \times 320 6375 40.3 84.0 82.7
    RefineDet512 (COCO pretrain) 512×512512 \times 512 16320 24.1 85.2 85.0
    RefineDet320+ (COCO pretrain) - - - 85.6 86.0
    RefineDet512+ (COCO pretrain) - - - 85.8 86.8

    RefineDet320 achieves 80.0% mAP on VOC 2007 at 40.3 FPS (24.8 ms latency), while RefineDet512 reaches 81.8% mAP at 24.1 FPS (41.5 ms latency). With MS COCO pretraining and multi-scale evaluation (RefineDet512+), accuracy reaches 85.8% on VOC 2007 test and 86.8% on VOC 2012 test.

  10. Knowl 10 — Object Detection Performance on MS COCO Benchmark

    data/table

    Performance of RefineDet on the MS COCO test-dev dataset evaluated under different base backbones (VGG-16 and ResNet-101) and input scales, compared against single-stage and two-stage detectors trained on trainval35k.

    Method Backbone AP AP50\mathrm{AP}_{50} AP75\mathrm{AP}_{75} APS\mathrm{AP}_S APM\mathrm{AP}_M APL\mathrm{AP}_L
    Faster R-CNN w FPN ResNet-101-FPN 36.2 59.1 39.0 18.2 39.0 48.2
    Deformable R-FCN Aligned-Incep-ResNet 37.5 58.0 40.8 19.4 40.1 52.5
    G-RMI (5-model ensemble) Ensemble of 5 41.6 61.9 45.4 23.9 43.5 54.9
    SSD512* VGG-16 28.8 48.5 30.3 10.9 31.8 43.5
    DSSD513 ResNet-101 33.2 53.3 35.2 13.0 35.4 51.1
    RetinaNet800 ResNet-101-FPN 39.1 59.1 42.3 21.8 42.7 50.2
    RefineDet320 VGG-16 29.4 49.2 31.3 10.0 32.0 44.4
    RefineDet512 VGG-16 33.0 54.5 35.5 16.3 36.3 44.3
    RefineDet320 ResNet-101 32.0 51.4 34.2 10.5 34.7 50.4
    RefineDet512 ResNet-101 36.4 57.5 39.5 16.6 39.9 51.4
    RefineDet320+ (Multi-scale) VGG-16 35.2 56.1 37.7 19.5 37.2 47.0
    RefineDet512+ (Multi-scale) VGG-16 37.6 58.7 40.8 22.7 40.3 48.3
    RefineDet320+ (Multi-scale) ResNet-101 38.6 59.9 41.7 21.1 41.7 52.3
    RefineDet512+ (Multi-scale) ResNet-101 41.8 62.9 45.7 25.6 45.1 54.1

    RefineDet512+ (with ResNet-101 backbone and multi-scale testing) achieves 41.8% AP on MS COCO test-dev, outperforming one-stage models such as RetinaNet800 (39.1% AP) and surpassing the 5-model ensemble G-RMI (41.6% AP).

Coverage note — None was omitted; all key architectural components (ARM, ODM, TCB, negative anchor filtering, cascaded regression), loss functions, anchor design parameters, training setups, ablation studies, and full quantitative evaluation results on VOC and COCO benchmarks are covered.

References

  1. 1.S. Bell, C. L. Zitnick, K. Bala, and R. B. Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In CVPR, pages 2874–2883, 2016. 3, 6, 7, 8
  2. 2.N. Bodla, B. Singh, R. Chellappa, and L. S. Davis. Improving object detection with one line of code. In ICCV, 2017. 7, 8
  3. 3.Z. Cai, Q. Fan, R. S. Feris, and N. Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In ECCV, pages 354–370, 2016. 1, 3
  4. 4.L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015. 4
  5. 5.J. Dai, Y. Li, K. He, and J. Sun. R-FCN: object detection via region-based fully convolutional networks. In NIPS, pages 379–387, 2016. 1, 3, 6, 7, 8
  6. 6.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In ICCV, 2017. 7, 8
  7. 7.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov. Scalable object detection using deep neural networks. In CVPR, pages 2155–2162, 2014. 4
  8. 8.M. Everingham, L. J. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman. The pascal visual object classes (VOC) challenge. IJCV, 88(2):303–338, 2010. 1, 3
  9. 9.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The Leaderboard of the PASCAL Visual Object Classes Challenge 2012 (VOC2012). http://host.robots.ox.ac.uk:8080/leaderboard/displaylb.php?challengeid=11&compid=4. Online; accessed 1 October 2017. 8
  10. 10.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascal-network.org/challenges/VOC/voc2007/workshop/index.html. Online; accessed 1 October 2017. 2
  11. 11.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html. Online; accessed 1 October 2017. 2, 3
  12. 12.P. F. Felzenszwalb, R. B. Girshick, D. A. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. TPAMI, 32(9):1627–1645, 2010. 3
  13. 13.C. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg. DSSD : Deconvolutional single shot detector. CoRR, abs/1701.06659, 2017. 3, 5, 6, 7
  14. 14.S. Gidaris and N. Komodakis. Object detection via a multiregion and semantic segmentation-aware CNN model. In ICCV, pages 1134–1142, 2015. 3, 6
  15. 15.R. B. Girshick. Fast R-CNN. In ICCV, pages 1440–1448, 2015. 1, 3, 5, 6, 7
  16. 16.R. B. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, pages 580–587, 2014. 3
  17. 17.X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. In AISTATS, pages 249–256, 2010. 5
  18. 18.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In ECCV, pages 346–361, 2014. 3
  19. 19.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 3, 4, 7, 8
  20. 20.A. G. Howard. Some improvements on deep convolutional neural network based image classification. CoRR, abs/1312.5402, 2013. 4
  21. 21.J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, and K. Murphy. Speed/accuracy trade-offs for modern convolutional object detectors. In CVPR, 2017. 5, 7, 8
  22. 22.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, pages 448–456, 2015. 4
  23. 23.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. B. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In ACMMM, pages 675–678, 2014. 5
  24. 24.T. Kong, F. Sun, A. Yao, H. Liu, M. Lu, and Y. Chen. RON: reverse connection with objectness prior networks for object detection. In CVPR, 2017. 1, 3, 5, 6, 7, 8
  25. 25.T. Kong, A. Yao, Y. Chen, and F. Sun. Hypernet: Towards accurate region proposal generation and joint object detection. In CVPR, pages 845–853, 2016. 3, 6
  26. 26.H. Lee, S. Eum, and H. Kwon. ME R-CNN: multi-expert region-based CNN for object detection. In ICCV, 2017. 3
  27. 27.T. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie. Feature pyramid networks for object detection. In CVPR, 2017. 1, 3, 7
  28. 28.T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár. Focal loss for dense object detection. In ICCV, 2017. 1, 3, 7, 8
  29. 29.T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft COCO: common objects in context. In ECCV, pages 740–755, 2014. 1, 2, 3, 8
  30. 30.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg. SSD: single shot multibox detector. In ECCV, pages 21–37, 2016. 1, 3, 4, 6, 7, 8
  31. 31.W. Liu, A. Rabinovich, and A. C. Berg. Parsenet: Looking wider to see better. In ICLR workshop, 2016. 4
  32. 32.P. H. O. Pinheiro, R. Collobert, and P. Dollár. Learning to segment object candidates. In NIPS, pages 1990–1998, 2015. 3
  33. 33.P. O. Pinheiro, T. Lin, R. Collobert, and P. Dollár. Learning to refine object segments. In ECCV, pages 75–91, 2016. 3
  34. 34.J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. In CVPR, pages 779–788, 2016. 3, 6
  35. 35.J. Redmon and A. Farhadi. YOLO9000: better, faster, stronger. CoRR, abs/1612.08242, 2016. 1, 3, 6, 7
  36. 36.S. Ren, K. He, R. B. Girshick, and J. Sun. Faster R-CNN: towards real-time object detection with region proposal networks. TPAMI, 39(6):1137–1149, 2017. 1, 3, 6, 7, 8
  37. 37.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li. Imagenet large scale visual recognition challenge. IJCV, 115(3):211–252, 2015. 3, 4, 5
  38. 38.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. In ICLR, 2014. 3
  39. 39.Z. Shen, Z. Liu, J. Li, Y. Jiang, Y. Chen, and X. Xue. DSOD: learning deeply supervised object detectors from scratch. In ICCV, 2017. 3, 6, 8
  40. 40.A. Shrivastava and A. Gupta. Contextual priming and feedback for faster R-CNN. In ECCV, pages 330–348, 2016. 3
  41. 41.A. Shrivastava, A. Gupta, and R. B. Girshick. Training region-based object detectors with online hard example mining. In CVPR, pages 761–769, 2016. 1, 3, 6, 7, 8
  42. 42.A. Shrivastava, R. Sukthankar, J. Malik, and A. Gupta. Beyond skip connections: Top-down modulation for object detection. CoRR, abs/1612.06851, 2016. 3, 7, 8
  43. 43.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014. 3, 4
  44. 44.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, pages 4278–4284, 2017. 4, 7
  45. 45.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, pages 1–9, 2015. 6
  46. 46.J. R. R. Uijlings, K. E. A. van de Sande, T. Gevers, and A. W. M. Smeulders. Selective search for object recognition. IJCV, 104(2):154–171, 2013. 3
  47. 47.P. A. Viola and M. J. Jones. Rapid object detection using a boosted cascade of simple features. In CVPR, pages 511–518, 2001. 3
  48. 48.X. Wang, A. Shrivastava, and A. Gupta. A-fast-rcnn: Hard positive generation via adversary for object detection. In CVPR, 2017. 3
  49. 49.S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 4, 8
  50. 50.X. Zeng, W. Ouyang, B. Yang, J. Yan, and X. Wang. Gated bi-directional CNN for object detection. In ECCV, pages 354–369, 2016. 3
  51. 51.S. Zhang, X. Zhu, Z. Lei, H. Shi, X. Wang, and S. Z. Li. Detecting face with densely connected face proposal network. In CCBR, pages 3–12, 2017. 4
  52. 52.S. Zhang, X. Zhu, Z. Lei, H. Shi, X. Wang, and S. Z. Li. Faceboxes: A CPU real-time face detector with high accuracy. In IJCB, 2017. 4
  53. 53.S. Zhang, X. Zhu, Z. Lei, H. Shi, X. Wang, and S. Z. Li. S^3FD: Single shot scale-invariant face detector. In ICCV, 2017. 1, 3, 4
  54. 54.Y. Zhu, C. Zhao, J. Wang, X. Zhao, Y. Wu, and H. Lu. Couplenet: Coupling global structure with local parts for object detection. In ICCV, 2017. 3, 5, 6, 7
  55. 55.C. L. Zitnick and P. Dollár. Edge boxes: Locating object proposals from edges. In ECCV, pages 391–405, 2014. 3

Citation

MLA
Zhang, S., et al. “Single-Shot Refinement Neural Network for Object Detection”. arXiv, 2017, http://arxiv.org/abs/1711.06897v3.
APA
Zhang, S., Wen, L., Bian, X., Lei, Z., & Li, S. Z. (2017). Single-Shot Refinement Neural Network for Object Detection. arXiv. http://arxiv.org/abs/1711.06897v3
Chicago
Zhang, S., L. Wen, X. Bian, Z. Lei, and S. Z. Li. 2017. “Single-Shot Refinement Neural Network for Object Detection”. arXiv. http://arxiv.org/abs/1711.06897v3.
Harvard
Zhang, S. et al. (2017) “Single-Shot Refinement Neural Network for Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.06897v3.
Vancouver
1. Zhang S, Wen L, Bian X, Lei Z, Li SZ (2017) Single-Shot Refinement Neural Network for Object Detection. arXiv

BibTeX

@article{zhang2017single,
  title = {Single-Shot Refinement Neural Network for Object Detection},
  author = {Zhang, Shifeng and Wen, Longyin and Bian, Xiao and Lei, Zhen and Li, Stan Z.},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.06897v3},
  eprint = {1711.06897}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE