Receptive Field Block Net for Accurate and Fast Object Detection

Songtao LiuDi HuangYunhong Wang

article2017ECCV1,555 citations

Introduces RFB Net, an object detector that incorporates biologically inspired receptive field structures into lightweight networks to match the accuracy of deep models at real-time speeds.

Listen

Real-time computer vision systems increasingly require rapid and highly accurate object detection. Modern solutions generally face a steep trade-off: high-accuracy detectors rely on massive, deep neural networks that demand heavy computational resources and run slowly, while lightweight detectors achieve real-time speeds but suffer substantial drops in accuracy. The article addresses this operational bottleneck by exploring whether lightweight visual models can achieve top-tier precision without adding prohibitive computational overhead.

The main objective of the article is to design and evaluate a biologically inspired module, named the Receptive Field Block (RFB), that strengthens the feature representation of lightweight networks to deliver fast and highly accurate object detection. The researchers evaluated this design by building a detector called RFB Net and benchmarking its speed and accuracy against leading detection architectures on standard industry test datasets, Pascal VOC and Microsoft COCO.

To accomplish this, the authors drew inspiration from human visual cortex mechanisms, where sensory fields closer to the center are smaller and more sensitive, while outer fields are larger. The RFB module replicates this structure by combining multi-branch convolutions of varying kernel sizes with dilated layers to control spatial eccentricity. The researchers integrated this lightweight block into the standard Single Shot Detector (SSD) framework, using a standard VGG backbone network, and assessed accuracy across multiple image resolutions alongside frames-per-second (FPS) and latency benchmarks.

The experimental findings show that the proposed approach successfully resolves the trade-off between speed and accuracy. On the Pascal VOC benchmark, RFB Net300 achieved an accuracy score of 80.5% Mean Average Precision (mAP) while operating at 83 frames per second, matching the accuracy of slower, two-stage detectors such as R-FCN while running roughly nine times faster. On the complex Microsoft COCO benchmark, an enhanced variant (RFB Net512-E) reached 34.4% mAP with an inference time of 33 milliseconds, matching the accuracy of the heavier RetinaNet500 while operating at nearly three times the speed (33 ms versus 90 ms). Comparative ablation experiments demonstrated that the RFB architecture consistently outperformed existing alternative modules, such as Inception and Atrous Spatial Pyramid Pooling, while adding negligible parameter overhead. Furthermore, integrating the module into ultra-lightweight architectures like MobileNet produced noticeable accuracy gains (from 19.3% to 20.7% mAP on COCO), and the network demonstrated an ability to train effectively from scratch without standard pre-training.

These results demonstrate that high detection accuracy does not strictly require massive neural networks or high-cost hardware. By utilizing biologically inspired hand-crafted structures, organizations can deploy high-performing computer vision models on lower-end edge devices, mobile platforms, and latency-critical systems. This reduces hardware investment, energy consumption, and operational costs while maintaining safety-critical real-time performance.

Based on these findings, teams developing vision systems should consider adopting the RFB module to optimize existing single-stage detection pipelines. Technical teams should pilot the module on target edge hardware, exploring configurations such as MobileNet-RFB for constrained embedded environments or RFB Net512 for higher-precision deployments. Future efforts should assess performance across broader domain-specific datasets and test the module on next-generation lightweight backbones.

Confidence in these findings is high given the standardized benchmarks and direct ablation tests performed under consistent hardware environments. However, decision-makers should note that evaluations were conducted using desktop-class graphics cards and standardized benchmark datasets; real-world edge hardware with different latency constraints and non-standard image resolutions may exhibit slight variations in performance.

Cover for Receptive Field Block Net for Accurate and Fast Object Detection

Abstract

Current top-performing object detectors depend on deep CNN backbones, such as ResNet-101 and Inception, benefiting from their powerful feature representations but suffering from high computational costs. Conversely, some lightweight model based detectors fulfil real time processing, while their accuracies are often criticized. In this paper, we explore an alternative to build a fast and accurate detector by strengthening lightweight features using a hand-crafted mechanism. Inspired by the structure of Receptive Fields (RFs) in human visual systems, we propose a novel RF Block (RFB) module, which takes the relationship between the size and eccentricity of RFs into account, to enhance the feature discriminability and robustness. We further assemble RFB to the top of SSD, constructing the RFB Net detector. To evaluate its effectiveness, experiments are conducted on two major benchmarks and the results show that RFB Net is able to reach the performance of advanced very deep detectors while keeping the real-time speed. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Visual Cortex Revisit
  • 3.2 Receptive Field Block
  • 3.3 RFB Net Detection Architecture
  • 3.4 Training Settings
  • 4 Experiments
  • 4.1 Pascal VOC 2007
  • 4.2 Ablation Study
  • 4.3 Microsoft COCO
  • 5 Discussion
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Receptive Field Block Architecture

    model/method

    The Receptive Field Block (RFB) is a multi-branch convolutional module inspired by the population receptive fields (pRFs) in the human visual cortex, where receptive field size is positively correlated with visual eccentricity (distance from the center). In deep networks, RFB assigns larger weights to near-center positions with smaller kernels while covering broader context eccentrically.

    The RFB architecture comprises:

    1. Multi-branch Convolutions: An input feature map passes through parallel branches with different convolution kernel sizes (1×11\times 1, 3×33\times 3, and 5×55\times 5, where 5×55\times 5 is implemented as two stacked 3×33\times 3 convolutions to reduce parameters and increase non-linearity). Each branch begins with a 1×11\times 1 bottleneck convolution to reduce channel dimensionality.
    2. Trailing Dilated Convolutions: Each branch is followed by a dilated (atrous) convolution layer. The dilation rate is set proportionally to the kernel size to reproduce the positive correlation between receptive field size and eccentricity:
      • Branch 1: 1×11\times 1 convolution followed by a 3×33\times 3 convolution with dilation rate r=1r=1.
      • Branch 2: 3×33\times 3 convolution followed by a 3×33\times 3 convolution with dilation rate r=3r=3.
      • Branch 3: Stacked 3×33\times 3 convolutions (effective 5×55\times 5) followed by a 3×33\times 3 convolution with dilation rate r=5r=5.
    3. Concatenation and Residual Shortcut: The outputs of all branches are concatenated along the channel dimension, linearly projected via a 1×11\times 1 convolution to match the input channel count, and added to the residual shortcut connection before applying a ReLU activation.
  2. Knowl 2 — RFB-s Module for Shallow Feature Representations

    model/method

    In human retinotopic maps, shallow visual areas exhibit smaller receptive field sizes and smaller eccentricity-to-size ratios than deeper areas. To replicate this property on high-resolution shallow feature maps (such as the conv4_3 stage of a detection backbone), a specialized variant termed RFB-s is designed.

    The RFB-s module employs four parallel branches featuring smaller kernels and asymmetric factorized convolutions:

    1. Branch 1: 1×11\times 1 bottleneck convolution followed by a 3×33\times 3 convolution with dilation rate r=1r=1.
    2. Branch 2: 1×11\times 1 bottleneck convolution followed by asymmetric 1×31\times 3 and 3×13\times 1 convolutions, followed by a 3×33\times 3 convolution with dilation rate r=3r=3.
    3. Branch 3: 1×11\times 1 bottleneck convolution followed by asymmetric 3×13\times 1 and 1×31\times 3 convolutions, followed by a 3×33\times 3 convolution with dilation rate r=3r=3.
    4. Branch 4: 1×11\times 1 bottleneck convolution followed by two stacked 3×33\times 3 convolutions, followed by a 3×33\times 3 convolution with dilation rate r=5r=5.

    The four branch outputs are concatenated, projected with a 1×11\times 1 convolution, and added to the residual shortcut connection before a final ReLU activation.

  3. Knowl 3 — RFB Net Object Detector Architecture

    model/method

    RFB Net is a real-time single-stage object detector constructed by embedding Receptive Field Block modules into the Single Shot MultiBox Detector (SSD) framework:

    1. Lightweight Backbone: Uses a reduced VGG-16 backbone network pre-trained on ILSVRC CLS-LOC. The fully connected layers fc6 and fc7 are converted to dilated convolutional layers (conv6 and conv7_fc), the pool5 layer is modified from 2×22\times 2 with stride 2 to 3×33\times 3 with stride 1, and fc8 and dropout layers are removed.
    2. Multi-Scale RFB Feature Pyramid: The cascade of convolutional feature maps in SSD is updated:
      • The high-resolution conv4_3 feature map is passed through an RFB-s module.
      • The conv7_fc feature map is passed through a standard RFB module.
      • Downsampled detection layers are created using stride-2 RFB modules (applying stride-2 multi-kernel convolutions within the block).
      • The smallest resolution top layers retain standard 1×11\times 1 and 3×33\times 3 convolutions, as their spatial dimensions are too small for large dilated filters.
    3. Anchor Priors: Six default anchor boxes with varying aspect ratios are placed at conv4_3 (increased from 4 default boxes in standard SSD) to strengthen small object detection.
    4. RFB Net512-E Extension: An enhanced model variant incorporates two modifications: (a) up-sampling conv7_fc and concatenating it with conv4_3 prior to the RFB-s module (top-down feature fusion), and (b) adding a 7×77\times 7 kernel branch across all RFB blocks.
  4. Knowl 4 — Learning Rate Warmup Schedule for RFB Net

    experimental setup

    Directly training detectors containing RFB modules with standard high initial learning rates (e.g., 10−310^{-3}) causes severe loss fluctuations and training instability due to the multi-branch and dilated structure. To stabilize optimization, a learning rate warmup strategy is applied:

    • PASCAL VOC Schedule: For batch size 32, training starts with a 5-epoch warmup phase where the learning rate increases linearly from 10−610^{-6} to 4×10−34\times 10^{-3}. Following epoch 5, the learning rate drops by a factor of 10 at 150 and 200 epochs, training for a total of 250 epochs using SGD with momentum 0.9 and weight decay 0.0005.
    • MS COCO Schedule: For batch size 32, the learning rate warms up from 10−610^{-6} to 2×10−32\times 10^{-3} over the first 5 epochs, then decays by a factor of 10 at 80 and 100 epochs, concluding at 120 total epochs.
  5. Knowl 5 — PASCAL VOC 2007 Detection Performance

    data/table

    Object detection performance comparison on the PASCAL VOC 2007 test set for models trained on the VOC 2007 trainval and VOC 2012 trainval union (07+1207+12). Inference speeds were measured on an NVIDIA GeForce GTX Titan X (Maxwell architecture) GPU.

    Method Backbone mAP (%) FPS
    Faster R-CNN VGG 73.2 7
    Faster R-CNN ResNet-101 76.4 5
    R-FCN ResNet-101 80.5 9
    YOLOv2 544 Darknet 78.6 40
    R-FCN w/ Deformable CNN ResNet-101 82.6 8
    SSD300* VGG 77.2 120
    DSSD321 ResNet-101 78.6 9.5
    RFB Net300 VGG 80.5 83
    SSD512* VGG 79.8 50
    DSSD513 ResNet-101 81.5 5.5
    RFB Net512 VGG 82.2 38

    RFB Net300 achieves 80.5% mAP at 83 FPS, outperforming SSD300* (77.2% mAP) and matching the accuracy of the two-stage R-FCN with ResNet-101 (80.5% mAP at 9 FPS) while running more than 9 times faster. RFB Net512 achieves 82.2% mAP at 38 FPS, exceeding the ResNet-101-based DSSD513 (81.5% mAP at 5.5 FPS).

  6. Knowl 6 — Ablation Study of RFB Components on PASCAL VOC 2007

    data/table

    An ablation study on the PASCAL VOC 2007 test set evaluates the impact of each architectural modification to the baseline SSD300 (300×300300\times 300 input resolution):

    RFB-max pooling Add RFB-s More Prior RFB-avg pooling RFB-dilated conv mAP (%)
    77.2
    ✓ 79.1
    ✓ ✓ 79.6
    ✓ ✓ ✓ 79.8
    ✓ ✓ ✓ 79.8
    ✓ ✓ 80.1
    ✓ ✓ ✓ 80.5

    Key takeaways:

    1. Replacing top convolution layers with the RFB module using max-pooling improves mAP by 1.9% (77.2% to 79.1%).
    2. Incorporating the RFB-s module on conv4_3 to simulate cortical retinotopic map differences provides an additional +0.5% gain with pooling (79.1% to 79.6%) and +0.4% with dilated convolutions (80.1% to 80.5%).
    3. Increasing the number of default anchors at conv4_3 from 4 to 6 adds +0.2% mAP (79.6% to 79.8%).
    4. Switching from stationary dilated pooling to learnable dilated convolutions yields a +0.7% boost (79.8% to 80.5%) without degrading inference frame rate.
  7. Knowl 7 — Architectural Comparison of RFB with Inception, ASPP, and Deformable Convolutions

    data/table

    RFB is compared directly against other multi-scale receptive field blocks (Inception, Atrous Spatial Pyramid Pooling [ASPP], and Deformable Convolutional Networks [DCN]) by mounting each module onto the same SSD top layer architecture under an identical training protocol:

    • Inception: Multi-branch convolutions with varying kernel sizes sampled at the same center.
    • Inception-L: Inception adapted to have the same effective receptive field size as RFB.
    • ASPP-S: ASPP tuned down from semantic segmentation scale to match the receptive field size of RFB.
    • Deformable CNN: Adaptive spatial sampling offsets conditioned on input features.
    Architecture Parameters VOC 2007 mAP (%) COCO minival mAP (%)
    RFB 34.5M 80.1 29.7
    Inception 32.9M 78.4 27.3
    Inception-L 33.3M 79.5 28.5
    ASPP-S 33.4M 79.7 28.1
    Deformable CNN 35.2M 79.5 27.6

    RFB achieves the highest detection accuracy (80.1% on VOC 2007 and 29.7% on COCO minival). Its superior performance over Inception-L, ASPP-S, and Deformable CNN stems from its daisy-shaped receptive field structure, which simultaneously emphasizes central features with smaller kernels and captures peripheral context with dilated convolutions.

  8. Knowl 8 — MS COCO Object Detection Performance and Speed

    data/table

    Detection performance on the MS COCO test-dev 2015 benchmark. Inference runtimes are measured on an NVIDIA Titan X (Maxwell) GPU, except RetinaNet, Mask R-CNN, and FPN (measured on an NVIDIA M40 GPU).

    Method Backbone Time AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    Faster R-CNN VGG 147 ms 24.2 45.3 23.5 7.7 26.4 37.1
    Faster R-CNN+++ ResNet-101 3.36 s 34.9 55.7 37.4 15.6 38.7 50.9
    Faster w/ FPN ResNet-101-FPN 240 ms 36.2 59.1 39.0 18.2 39.0 48.2
    Faster by G-RMI Inception-ResNet-v2 – 34.7 55.5 36.7 13.5 38.1 52.0
    R-FCN ResNet-101 110 ms 29.9 51.9 – 10.8 32.8 45.0
    R-FCN w/ DCN ResNet-101 125 ms 34.5 55.0 – 14.0 37.7 50.3
    Mask R-CNN ResNeXt-101-FPN 210 ms 37.1 60.0 39.4 16.9 39.9 53.5
    YOLOv2 Darknet 25 ms 21.6 44.0 19.2 5.0 22.4 35.5
    SSD300* VGG 12 ms 25.1 43.1 25.8 – – –
    SSD512* VGG 28 ms 28.8 48.5 30.3 – – –
    DSSD513 ResNet-101 182 ms 33.2 53.3 35.2 13.0 35.4 51.1
    RetinaNet500 ResNet-101-FPN 90 ms 34.4 53.1 36.8 14.7 38.5 49.1
    RetinaNet800 ResNet-101-FPN 198 ms 39.1 59.1 42.3 21.8 42.7 50.2
    RFB Net300 VGG 15 ms 30.3 49.3 31.8 11.8 31.9 45.9
    RFB Net512 VGG 30 ms 33.8 54.2 35.9 16.2 37.1 47.4
    RFB Net512-E VGG 33 ms 34.4 55.7 36.4 17.6 37.0 47.6

    RFB Net300 achieves 30.3% AP (49.3% AP50\text{AP}_{50}) at 15 ms (66 FPS), outperforming SSD300* (25.1% AP) and R-FCN with ResNet-101 (29.9% AP at 110 ms). RFB Net512-E reaches 34.4% AP at 33 ms (30 FPS), matching RetinaNet500 (34.4% AP at 90 ms) while running nearly 3×3\times faster.

  9. Knowl 9 — RFB Integration with MobileNet Backbone

    empirical result

    To assess generalization to lightweight mobile backbones, the RFB module was integrated into MobileNet-SSD. Both the original MobileNet-SSD and MobileNet-SSD with RFB were trained on MS COCO train+val35k and evaluated on minival2014 at 300×300300\times 300 resolution:

    Framework Model mAP (%) Parameters
    SSD 300 MobileNet 19.3 6.8M
    SSD 300 MobileNet + RFB 20.7 7.4M

    Incorporating RFB into MobileNet-SSD increases detection mAP from 19.3% to 20.7% (+1.4% mAP) on COCO minival while adding only 0.6M parameters (from 6.8M to 7.4M), confirming the module's effectiveness on ultra-lightweight networks for resource-constrained devices.

  10. Knowl 10 — Training RFB Net from Scratch Without Pre-training

    empirical result

    While standard single-stage detectors (such as SSD with VGG or ResNet backbones) experience substantial performance drops when trained from scratch without ImageNet pre-training, RFB Net trains effectively from random initialization.

    When trained from scratch on the PASCAL VOC 07+1207+12 trainval set using MSRA weight initialization, RFB Net300 achieves 77.6% mAP on the VOC 2007 test set. This is on par with the 77.7% mAP reported by Deeply Supervised Object Detectors (DSOD), a model specialized for training from scratch. When initialized with ImageNet pre-trained weights, RFB Net300 reaches 80.5% mAP.

Coverage note — No substantial contributed material was omitted; all key architectural components, ablation studies, benchmark results, and scratch-training properties are covered.

References

  1. 1.Brown, M., Hua, G., Winder, S.: Discriminative learning of local image descriptors. TPAMI 33(1), 43–57 (2011)
  2. 2.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFS. arXiv preprint arXiv:1606.00915 (2016)
  3. 3.Chen, L.C., Papandreou, G., Schroff, F., Adam, H.: Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 (2017)
  4. 4.Dai, J., et al.: Deformable convolutional networks. In: ICCV (2017)
  5. 5.Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The PASCAL visual object classes (voc) challenge. IJCV 88(2), 303–338 (2010)
  6. 6.Fu, C.Y., et al.: DSSD: deconvolutional single shot detector. arXiv preprint arXiv:1701.06659 (2017)
  7. 7.Girshick, R.: Fast R-CNN. In: ICCV (2015)
  8. 8.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR (2014)
  9. 9.He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask R-CNN. In: ICCV (2017)
  10. 10.He, K., Zhang, X., Ren, S., Sun, J.: Delving deep into rectifiers: surpassing human-level performance on imagenet classification. In: ICCV (2015)
  11. 11.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  12. 12.Howard, A.G., et al.: Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)
  13. 13.Hu, P., Ramanan, D.: Finding tiny faces. In: CVPR (2017)
  14. 14.Huang, D., Zhu, C., Wang, Y., Chen, L.: HSOG: a novel local image descriptor based on histograms of the second-order gradients. IEEE Trans. Image Process. 23(11), 4680–4695 (2014)
  15. 15.Huang, J., et al.: Speed/accuracy trade-offs for modern convolutional object detectors. In: CVPR (2017)
  16. 16.Kim, K.H., Hong, S., Roh, B., Cheon, Y., Park, M.: PVANET: deep but lightweight neural networks for real-time object detection. arXiv preprint arXiv:1608.08021 (2016)
  17. 17.Li, Y., He, K., Sun, J., et al.: R-FCN: object detection via region-based fully convolutional networks. In: NIPS (2016)
  18. 18.Li, Y., Qi, H., Dai, J., Ji, X., Wei, Y.: Fully convolutional instance-aware semantic segmentation. In: CVPR (2017)
  19. 19.Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: CVPR (2017)
  20. 20.Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: ICCV (2017)
  21. 21.Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740–755. Springer, Cham (2014). https://doi.org/10.1007/978-3-319-10602-1_48
  22. 22.Liu, W., et al.: SSD: single shot multibox detector. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9905, pp. 21–37. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46448-0_2
  23. 23.Luo, W., et al.: Understanding the effective receptive field in deep convolutional neural networks. In: NIPS (2016)
  24. 24.Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: unified, real-time object detection. In: CVPR (2016)
  25. 25.Redmon, J., Farhadi, A.: Yolo9000: better, faster, stronger. In: CVPR (2017)
  26. 26.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: towards real-time object detection with region proposal networks. In: NIPS (2015)
  27. 27.Russakovsky, O., et al.: Imagenet large scale visual recognition challenge. IJCV 115(3), 211–252 (2015)
  28. 28.Shen, Z., Liu, Z., Li, J., Jiang, Y.G., Chen, Y., Xue, X.: DSOD: learning deeply supervised object detectors from scratch. In: ICCV (2017)
  29. 29.Simonyan, K., Vedaldi, A., Zisserman, A.: Learning local feature descriptors using convex optimisation. TPAMI 36(8), 1573–1585 (2014)
  30. 30.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. In: NIPS (2014)
  31. 31.Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.A.: Inception-v4, inception-resnet and the impact of residual connections on learning. In: AAAI (2017)
  32. 32.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: Rethinking the inception architecture for computer vision. In: CVPR (2016)
  33. 33.Szegedy, C., et al.: Going deeper with convolutions. In: CVPR (2015)
  34. 34.Tola, E., Lepetit, V., Fua, P.: A fast local descriptor for dense matching. In: CVPR (2008)
  35. 35.Uijlings, J.R., Van De Sande, K.E., Gevers, T., Smeulders, A.W.: Selective search for object recognition. IJCV 104(2), 154–171 (2013)
  36. 36.Wandell, B.A., Winawer, J.: Computational neuroimaging and population receptive fields. In: Trends in Cognitive Sciences (2015)
  37. 37.Weng, D., Wang, Y., Gong, M., Tao, D., Wei, H., Huang, D.: DERF: distinctive efficient robust features from the biological modeling of the P Ganglion cells. IEEE Trans. Image Process. 24(8), 2287–2302 (2015)
  38. 38.Winder, S.A., Brown, M.: Learning local image descriptors. In: CVPR (2007)
  39. 39.Zhang, X., Zhou, X., Lin, M., Sun, J.: Shufflenet: an extremely efficient convolutional neural network for mobile devices. arXiv preprint arXiv:1707.01083 (2017)

Citation

MLA
Liu, S., et al. “Receptive Field Block Net for Accurate and Fast Object Detection”. arXiv, 2017, http://arxiv.org/abs/1711.07767v3.
APA
Liu, S., Huang, D., & Wang, Y. (2017). Receptive Field Block Net for Accurate and Fast Object Detection. arXiv. http://arxiv.org/abs/1711.07767v3
Chicago
Liu, S., D. Huang, and Y. Wang. 2017. “Receptive Field Block Net for Accurate and Fast Object Detection”. arXiv. http://arxiv.org/abs/1711.07767v3.
Harvard
Liu, S., Huang, D. and Wang, Y. (2017) “Receptive Field Block Net for Accurate and Fast Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.07767v3.
Vancouver
1. Liu S, Huang D, Wang Y (2017) Receptive Field Block Net for Accurate and Fast Object Detection. arXiv

BibTeX

@article{liu2017receptive,
  title = {Receptive Field Block Net for Accurate and Fast Object Detection},
  author = {Liu, Songtao and Huang, Di and Wang, Yunhong},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.07767v3},
  eprint = {1711.07767}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF