Training Region-Based Object Detectors with Online Hard Example Mining

Abhinav ShrivastavaAbhinav GuptaRoss Girshick

article2016CVPR2,783 citations

Proposes Online Hard Example Mining (OHEM) to automatically select difficult region proposals during ConvNet training, eliminating heuristic sampling hyperparameters while boosting object detection accuracy across standard benchmarks.

Listen

Object detection models based on region proposals and convolutional networks have advanced rapidly, yet their training still depends on manual heuristics to manage the extreme imbalance between easy background regions and the few difficult examples that matter most for accuracy. This imbalance slows convergence and limits performance, especially as datasets grow larger and more varied.

The article introduces online hard example mining (OHEM), a straightforward modification to stochastic gradient descent training that automatically selects the most informative examples for each update step. Instead of fixed rules for sampling foreground and background regions, the method computes loss for all candidate regions in the current images, then retains only the highest-loss subset for the backward pass while keeping the rest of the computation efficient through a dual-network architecture.

Experiments on PASCAL VOC 2007 and 2012 and the more challenging MS COCO dataset show that OHEM raises mean average precision by 2 to 5 points over the standard Fast R-CNN baseline, removes the need for several tuning parameters such as background overlap thresholds and foreground-background ratios, and produces larger gains on bigger datasets. When combined with multi-scale testing and iterative bounding-box refinement, the approach reaches state-of-the-art figures of 78.9 percent on VOC 2007 and 76.3 percent on VOC 2012.

These gains matter because they improve detection reliability without extra labeled data or heavier models, lowering the risk of missed or false objects in applications such as surveillance or robotics. The method also trains to a lower overall loss, indicating more effective use of the available examples.

Teams adopting the technique should first verify that their existing pipeline can accommodate the modest increase in per-iteration time and memory; if GPU resources are tight, the single-image variant of OHEM remains effective. Further work could examine whether similar online selection benefits other region-based detectors and whether per-class performance varies systematically with the new sampling strategy.

The reported improvements rest on standard benchmarks and two common network backbones; results may shift with newer architectures or substantially different proposal methods, so validation on target data remains advisable before large-scale deployment.

Cover for Training Region-Based Object Detectors with Online Hard Example Mining

Abstract

The field of object detection has made significant advances riding on the wave of region-based ConvNets, but their training procedure still includes many heuristics and hyperparameters that are costly to tune. We present a simple yet surprisingly effective online hard example mining (OHEM) algorithm for training region-based ConvNet detectors. Our motivation is the same as it has always been -- detection datasets contain an overwhelming number of easy examples and a small number of hard examples. Automatic selection of these hard examples can make training more effective and efficient. OHEM is a simple and intuitive algorithm that eliminates several heuristics and hyperparameters in common use. But more importantly, it yields consistent and significant boosts in detection performance on benchmarks like PASCAL VOC 2007 and 2012. Its effectiveness increases as datasets become larger and more difficult, as demonstrated by the results on the MS COCO dataset. Moreover, combined with complementary advances in the field, OHEM leads to state-of-the-art results of 78.9% and 76.3% mAP on PASCAL VOC 2007 and 2012 respectively.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Overview of Fast R-CNN
  • 3.1 Training
  • 4 Our approach
  • 4.1 Online hard example mining
  • 4.2 Implementation details
  • 5 Analyzing online hard example mining
  • 5.1 Experimental setup
  • 5.2 OHEM vs. heuristic sampling
  • 5.3 Robust gradient estimates
  • 5.4 Why just hard examples, when you can use all?
  • 5.5 Better optimization
  • 5.6 Computational cost
  • 6 PASCAL VOC and MS COCO results
  • 6.1 VOC 2007 and 2012 results
  • 6.2 MS COCO results
  • 7 Adding bells and whistles
  • 7.1 VOC 2007 and 2012 results
  • 7.2 MS COCO results
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Online Hard Example Mining (OHEM) Algorithm for Region-Based Detectors

    algorithm

    Online Hard Example Mining (OHEM) adapts bootstrapping to deep ConvNet object detectors trained end-to-end via stochastic gradient descent (SGD). Instead of training on a heuristically subsampled mini-batch of region proposals (RoIs), OHEM computes losses across all candidate RoIs from the input images, performs Non-Maximum Suppression (NMS) on the candidate proposals based on their computed loss to deduplicate spatially overlapping regions, and selects only the top BB hardest non-overlapping examples for the backward gradient computation.

    Input: Network parameters θ=(θconv,θRoI)\theta = (\theta_{\text{conv}}, \theta_{\text{RoI}}), image batch {I1,…,IN}\{I_1, \dots, I_N\}, proposal RoI sets {R1,…,RN}\{R_1, \dots, R_N\}, target batch size BB, NMS IoU threshold τNMS=0.7\tau_{\text{NMS}} = 0.7
    Output: Parameter updates Δθ\Delta \theta
    for each image IiI_i (i=1,…,Ni = 1, \dots, N) do
        Xi←ConvForward(Ii;θconv)X_i \leftarrow \text{ConvForward}(I_i; \theta_{\text{conv}})
        for each proposal r∈Rir \in R_i do
            L(r)←RoIForwardReadOnly(Xi,r;θRoI)L(r) \leftarrow \text{RoIForwardReadOnly}(X_i, r; \theta_{\text{RoI}})
        Rsorted←SortByLossDescending(Ri,L)R_{\text{sorted}} \leftarrow \text{SortByLossDescending}(R_i, L)
        Rhard,i←∅R_{\text{hard}, i} \leftarrow \emptyset
        while Rsorted≠∅R_{\text{sorted}} \neq \emptyset and ∣Rhard,i∣<B/N|R_{\text{hard}, i}| < B / N do
            r∗←PopHighestLoss(Rsorted)r^* \leftarrow \text{PopHighestLoss}(R_{\text{sorted}})
            Rhard,i←Rhard,i∪{r∗}R_{\text{hard}, i} \leftarrow R_{\text{hard}, i} \cup \{r^*\}
            for each r∈Rsortedr \in R_{\text{sorted}} do
                if IoU(r,r∗)≥τNMS\text{IoU}(r, r^*) \ge \tau_{\text{NMS}} then
                    Rsorted←Rsorted∖{r}R_{\text{sorted}} \leftarrow R_{\text{sorted}} \setminus \{r\}
    Rhard-sel←⋃i=1NRhard,iR_{\text{hard-sel}} \leftarrow \bigcup_{i=1}^N R_{\text{hard}, i}
    Compute RoI forward-backward pass: RoIForwardBackward({Xi},Rhard-sel;θRoI)\text{RoIForwardBackward}(\{X_i\}, R_{\text{hard-sel}}; \theta_{\text{RoI}})
    Backpropagate accumulated RoI gradients through conv network: ConvBackward({Xi};θconv)\text{ConvBackward}(\{X_i\}; \theta_{\text{conv}})
    Update parameters θ\theta via SGD step

    Here L(r)L(r) is the sum of the classification log loss and the smooth L1L_1 bounding-box regression loss on proposal rr. The IoU threshold τNMS=0.7\tau_{\text{NMS}} = 0.7 suppresses highly overlapping proposals that share convolutional feature map locations and exhibit correlated losses, preventing double counting of gradients.

  2. Knowl 2 — Dual RoI Network Architecture for Memory and Computationally Efficient OHEM

    model/method

    A naive implementation of OHEM computes forward losses for all RoIs, zeros out the losses of non-hard RoIs, and runs standard backpropagation. However, this naive approach causes deep learning frameworks to allocate activation memory and perform backward passes across all ∣R∣≈4000|R| \approx 4000 proposals per mini-batch (N=2N=2 images).

    To eliminate this overhead, OHEM uses a dual RoI network architecture with shared parameters:

    1. Read-only RoI Network: Allocates memory strictly for the forward pass. It takes the shared convolutional feature map and all candidate proposals RR (∣R∣≈4000|R| \approx 4000) to evaluate classification and localization losses for each RoI without constructing the computational backward graph.
    2. Hard RoI Selection Module: Applies loss-based sorting and greedy NMS (IoU threshold 0.70.7) on the computed losses to pick the B=128B=128 hardest, diverse proposals (Rhard-selR_{\text{hard-sel}}).
    3. Regular RoI Network: Performs forward and backward passes only on the selected B=128B=128 hard examples Rhard-selR_{\text{hard-sel}}, accumulating gradients at the RoI pooling layer and backpropagating them into the shared convolutional network.

    This architecture has approximately the same memory footprint as standard Fast R-CNN while executing over 2×2\times faster than the naive zero-loss masking baseline.

  3. Knowl 3 — Elimination of Foreground-Background Ratio and Background Lower-Bound Heuristics

    model/method

    Standard Fast R-CNN training employs two sampling heuristics to address class imbalance and approximate hard negative mining:

    1. Foreground-to-Background Ratio: Forces mini-batches to maintain a 1:31:3 foreground-to-background ratio by randomly undersampling background RoIs (ensuring 25%25\% foreground RoIs per batch).
    2. Background Lower-Bound Threshold (bg_lobg\_lo): Constrains negative/background samples to RoIs whose maximum Intersection-over-Union (IoU) with any ground-truth box falls within [bg_lo,0.5)[bg\_lo, 0.5), with bg_lo=0.1bg\_lo = 0.1, ignoring candidate regions with IoU<0.1\text{IoU} < 0.1.

    OHEM eliminates both heuristics:

    • It sets bg_lo=0bg\_lo = 0, allowing all proposals in [0,0.5)[0, 0.5) to be mined, which captures infrequent but critical false-alarm background regions that bg_lo=0.1bg\_lo = 0.1 discards.
    • It removes the fixed 1:31:3 foreground-background ratio. Because OHEM automatically selects samples based on their instantaneous loss, if any class (foreground or background) is under-represented or poorly classified, its loss rises until its instances are naturally sampled. Batches can thus be dynamically composed entirely of background RoIs (when foregrounds are easy) or entirely of foreground RoIs (when backgrounds are trivial).
  4. Knowl 4 — Optimization and Training Loss Comparison: OHEM vs. Full-Batch RoI Training

    empirical result

    Comparing OHEM against heuristic sampling and against full-batch RoI training (using all candidate RoIs in an image mini-batch, B=2048B=2048, bg_lo=0bg\_lo=0, with tuned learning rates of 0.0030.003 for VGG16 and 0.0040.004 for VGGM) on PASCAL VOC 2007:

    • Standard Fast R-CNN with bg_lo=0bg\_lo=0 and B=128B=128: 57.2%57.2\% mAP (VGGM), 67.5%67.5\% mAP (VGG16).
    • Standard Fast R-CNN with the bg_lo=0.1bg\_lo=0.1 heuristic and B=128B=128: 59.6%59.6\% mAP (VGGM), 67.2%67.2\% mAP (VGG16).
    • Full-batch Fast R-CNN with B=2048B=2048 (bg_lo=0bg\_lo=0): 60.4%60.4\% mAP (VGGM), 68.7%68.7\% mAP (VGG16).
    • OHEM (B=128B=128, bg_lo=0bg\_lo=0): 62.0%62.0\% mAP (VGGM), 69.9%69.9\% mAP (VGG16).

    Measuring the true average loss per RoI over the entire VOC 2007 trainval dataset at regular 20k-step optimization checkpoints demonstrates that OHEM achieves strictly lower training loss throughout training compared to bg_lo=0bg\_lo=0, bg_lo=0.1bg\_lo=0.1, and full-batch B=2048B=2048 training. Moreover, because OHEM computes gradients on only B=128B=128 examples, it trains significantly faster than full-batch (B=2048B=2048) training.

  5. Knowl 5 — Robustness of OHEM to Reduced Image Mini-Batch Size

    empirical result

    When training region-based detectors with stochastic gradient descent, reducing the number of images per mini-batch from N=2N=2 to N=1N=1 increases the spatial correlation among sampled RoIs.

    • For baseline Fast R-CNN (VGG16), reducing NN from 22 to 11 degrades detection accuracy on PASCAL VOC 2007 from 67.2%67.2\% to 66.3%66.3\% mAP (a drop of ∼0.9\sim 0.9 points mAP).
    • With OHEM, training on N=1N=1 image per mini-batch achieves 69.7%69.7\% mAP, virtually identical to the 69.9%69.9\% mAP achieved with N=2N=2.

    This shows that selecting high-loss examples with NMS deduplication produces stable and effective gradient estimates even when all B=128B=128 RoIs are sampled from a single image, enabling training on memory-constrained GPUs.

  6. Knowl 6 — Computational Time and Memory Overhead of OHEM Training

    data/table

    Computational cost and memory footprint benchmarks measured on a single Nvidia Titan X GPU during training of Fast R-CNN (FRCN) versus Fast R-CNN with OHEM:

    Model / Metric VGGM FRCN VGGM Ours VGG16 FRCN* VGG16 Ours*
    Time (sec/iter) 0.13 0.22 0.57 1.00
    Max Memory (GB) 2.6 3.6 6.4 8.7

    Note: VGG16 implementations use gradient accumulation over two forward/backward passes.

    OHEM introduces an overhead of 0.09 s0.09\,\text{s} per iteration and +1.0 GB+1.0\,\text{GB} memory for the VGGM network, and 0.43 s0.43\,\text{s} per iteration and +2.3 GB+2.3\,\text{GB} memory for the VGG16 network. Because convolutional feature map extraction is shared and backpropagation is restricted to B=128B=128 RoIs, OHEM achieves hard negative mining with moderate overhead.

  7. Knowl 7 — Detection Performance of OHEM on PASCAL VOC 2007 and VOC 2012

    data/table

    Evaluation of Fast R-CNN (FRCN) trained with OHEM using the VGG16 architecture on PASCAL VOC 2007 test and VOC 2012 test sets (mean Average Precision, mAP %):

    Dataset Training Set Standard FRCN mAP (%) OHEM FRCN mAP (%)
    VOC 2007 VOC07 trainval 67.2 69.9
    VOC 2007 VOC07+12 trainval 70.0 74.6
    VOC 2012 VOC12 trainval 65.7 69.8
    VOC 2012 VOC07++12 68.4 71.9

    On VOC 2007, OHEM improves mAP by +2.7%+2.7\% (trained on VOC07) and +4.6%+4.6\% (trained on VOC07+12). On VOC 2012, OHEM improves mAP by +4.1%+4.1\% (trained on VOC12) and +3.5%+3.5\% (trained on VOC07++12). Improvements are particularly pronounced on difficult object categories with high false positive rates, such as bottle (37.8%→46.5%37.8\% \to 46.5\% on VOC07), chair (42.2%→47.9%42.2\% \to 47.9\% on VOC07), and tvmonitor (68.1%→75.9%68.1\% \to 75.9\% on VOC07).

  8. Knowl 8 — Detection Accuracy on MS COCO Benchmark

    data/table

    Detection results on MS COCO 2015 test-dev comparing baseline Fast R-CNN (FRCN) with OHEM using the VGG16 backbone:

    Metric FRCN Ours Ours [+M] Ours* [+M]
    AP [0.50:0.95]\text{AP}\ [0.50:0.95] 19.7 22.6 24.4 25.5
    AP50\text{AP}^{50} (IoU ≥0.50\ge 0.50) 35.9 42.5 44.4 45.9
    AP75\text{AP}^{75} (IoU ≥0.75\ge 0.75) 19.9 22.2 24.8 26.1
    APsmall\text{AP}_{\text{small}} 3.5 5.0 7.1 7.4
    APmed\text{AP}_{\text{med}} 18.8 23.7 26.4 27.7
    APlarge\text{AP}_{\text{large}} 34.6 37.9 38.5 40.3

    Note: +M denotes multi-scale training and testing. * indicates training on MS COCO trainval rather than train.

    Under standard COCO evaluation (extAP [0.50:0.95] ext{AP}\,[0.50:0.95]), OHEM increases accuracy from 19.7%19.7\% to 22.6%22.6\% (+2.9%+2.9\% AP) and boosts AP50\text{AP}^{50} from 35.9%35.9\% to 42.5%42.5\% (+6.6%+6.6\% AP). The largest gains occur on medium objects (APmed\text{AP}_{\text{med}} increases by +4.9%+4.9\%, from 18.8%18.8\% to 23.7%23.7\%) and small objects (APsmall\text{AP}_{\text{small}} increases from 3.5%3.5\% to 5.0%5.0\%), demonstrating that hard example mining is especially effective on smaller, more difficult instances.

  9. Knowl 9 — Orthogonality of OHEM with Multi-Scale Processing and Iterative Box Regression

    empirical result

    OHEM is orthogonal and complementary to two standard detection enhancement techniques:

    1. Multi-Scale Processing (M): Training randomly samples image shortest-side scale s∈{480,576,688,864,900}s \in \{480, 576, 688, 864, 900\}; inference evaluates all scales s∈{480,576,688,864,1000}s \in \{480, 576, 688, 864, 1000\} (maximum dimension capped at 1000).
    2. Iterative Bounding-Box Regression and Voting (B): Proposal RoIs yield initial scores and relocalized boxes R1R_1. High-scoring R1R_1 boxes are rescored and relocalized into R2R_2. Post-processing applies NMS on RF=R1∪R2R_F = R_1 \cup R_2 at IoU threshold 0.30.3, followed by score-weighted bounding box voting on boxes with IoU≥0.5\text{IoU} \ge 0.5.

    Ablation on PASCAL VOC 2007 (VGG16, VOC07 trainval):

    • Baseline FRCN: 67.2%67.2\% mAP
    • FRCN + Multi-scale (test only): 68.4%68.4\% mAP
    • FRCN + Multi-scale (train + test): 68.6%68.6\% mAP
    • FRCN + Multi-scale + Iterative Bbox Regression (M + B): 72.4%72.4\% mAP
    • FRCN + OHEM: 69.9%69.9\% mAP
    • FRCN + OHEM + Multi-scale (M): 71.9%71.9\% mAP
    • FRCN + OHEM + Iterative Bbox Regression (B): 74.1%74.1\% mAP
    • FRCN + OHEM + M + B: 75.1%75.1\% mAP (and 78.9%78.9\% mAP on VOC07+12 data, outperforming MR-CNN at 78.2%78.2\% mAP).

    On PASCAL VOC 2012 (VOC07++12 data), FRCN + OHEM + M + B achieves 76.3%76.3\% mAP (compared to 73.9%73.9\% for MR-CNN).

Coverage note — None omitted; all core methodology, algorithm design, implementation architectures, ablation analyses, and benchmark evaluations on PASCAL VOC and MS COCO are fully covered.

References

  1. 1.B. Alexe, T. Deselaers, and V. Ferrari. What is an object? In CVPR, 2010.
  2. 2.B. Alexe, T. Deselaers, and V. Ferrari. Measuring the objectness of image windows. TPAMI, 2012.
  3. 3.P. Arbelaez, J. Pont-Tuset, J. T. Barron, F. Marques, ´ and J. Malik. Multiscale combinatorial grouping. In CVPR, 2014.
  4. 4.J. Carreira and C. Sminchisescu. Constrained parametric min-cuts for automatic object segmentation. In CVPR, 2010.
  5. 5.K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman. Return of the devil in the details: Delving deep into convolutional nets. In BMVC, 2014.
  6. 6.M.-M. Cheng, Z. Zhang, W.-Y. Lin, and P. H. S. Torr. BING: Binarized normed gradients for objectness estimation at 300fps. In CVPR, 2014.
  7. 7.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In CVPR, 2005.
  8. 8.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  9. 9.P. Dollar, Z. Tu, P. Perona, and S. Belongie. Integral ´ channel features. In BMVC, 2009.
  10. 10.I. Endres and D. Hoiem. Category independent object proposals. In ECCV, 2010.
  11. 11.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010.
  12. 12.P. Felzenszwalb, R. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. PAMI, 2010.
  13. 13.S. Gidaris and N. Komodakis. Object detection via a multi-region & semantic segmentation-aware cnn model. In ICCV, 2015.
  14. 14.R. Girshick. Fast R-CNN. In ICCV, 2015.
  15. 15.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014.
  16. 16.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In ECCV, 2014.
  17. 17.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014.
  18. 18.P. Krahenb ¨ uhl and V. Koltun. Geodesic object propos- ¨ als. In ECCV. 2014.
  19. 19.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  20. 20.Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1989.
  21. 21.T. Lin, M. Maire, S. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollar, ´ and C. L. Zitnick. Microsoft COCO: common objects in context. CoRR, abs/1405.0312, 2014.
  22. 22.I. Loshchilov and F. Hutter. Online batch selection for faster training of neural networks. arXiv preprint arXiv:1511.06343, 2015.
  23. 23.T. Malisiewicz, A. Gupta, and A. A. Efros. Ensemble of exemplar-svms for object detection and beyond. In ICCV, 2011.
  24. 24.S. Ren, K. He, R. Girshick, and J. Sun. Faster RCNN: Towards real-time object detection with region proposal networks. In Neural Information Processing Systems (NIPS), 2015.
  25. 25.H. Rowley, S. Baluja, and T. Kanade. Neural networkbased face detection. IEEE PAMI, 1998.
  26. 26.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun. Overfeat: Integrated recognition, localization and detection using convolutional networks. CoRR, abs/1312.6229, 2013.
  27. 27.E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, and F. Moreno-Noguer. Fracking deep convolutional image descriptors. arXiv preprint arXiv:1412.6537, 2014.
  28. 28.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  29. 29.S. Singh, A. Gupta, and A. A. Efros. Unsupervised discovery of mid-level discriminative patches. In European Conference on Computer Vision, 2012.
  30. 30.K.-K. Sung and T. Poggio. Learning and Example Selection for Object and Pattern Detection. In MIT A.I. Memo No. 1521, 1994.
  31. 31.M. Taka´c, A. Bijral, P. Richt ˇ arik, and N. Srebro. ´ Mini-batch primal and dual methods for svms. arXiv preprint arXiv:1303.2314, 2013.
  32. 32.J. Uijlings, K. van de Sande, T. Gevers, and A. Smeulders. Selective search for object recognition. IJCV, 2013.
  33. 33.X. Wang and A. Gupta. Unsupervised learning of visual representations using videos. In ICCV, 2015.
  34. 34.C. L. Zitnick and P. Dollar. Edge boxes: Locating object proposals from edges. In ECCV, 2014.

Citation

MLA
Shrivastava, A., et al. “Training Region-Based Object Detectors with Online Hard Example Mining”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 761–69, https://doi.org/10.1109/CVPR.2016.89.
APA
Shrivastava, A., Gupta, A., & Girshick, R. (2016). Training Region-Based Object Detectors with Online Hard Example Mining. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 761–769. https://doi.org/10.1109/CVPR.2016.89
Chicago
Shrivastava, A., A. Gupta, and R. Girshick. 2016. “Training Region-Based Object Detectors with Online Hard Example Mining”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 761–69. https://doi.org/10.1109/CVPR.2016.89.
Harvard
Shrivastava, A., Gupta, A. and Girshick, R. (2016) “Training Region-Based Object Detectors with Online Hard Example Mining”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 761–769. Available at: https://doi.org/10.1109/CVPR.2016.89.
Vancouver
1. Shrivastava A, Gupta A, Girshick R (2016) Training Region-Based Object Detectors with Online Hard Example Mining. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 761–769

BibTeX

@inproceedings{Shrivastava_2016, title={Training Region-Based Object Detectors with Online Hard Example Mining}, url={http://dx.doi.org/10.1109/CVPR.2016.89}, DOI={10.1109/cvpr.2016.89}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Shrivastava, Abhinav and Gupta, Abhinav and Girshick, Ross}, year={2016}, month=June, pages={761–769} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE