Libra R-CNN: Towards Balanced Learning for Object Detection

Jiangmiao PangKai ChenJianping ShiHuajun FengWanli OuyangDahua Lin

article2019CVPR1,540 citations

Proposes Libra R-CNN, a training framework that resolves sample, feature, and objective imbalances in object detectors using IoU-balanced sampling, balanced feature pyramids, and balanced L1 loss to boost detection accuracy on MS COCO.

Listen

Modern computer vision systems rely heavily on object detectors to recognize and locate items within images. While research has concentrated heavily on designing complex network architectures, the training process itself is often constrained by systemic imbalances that prevent these models from achieving their full potential.

The article demonstrates that standard detector training suffers from imbalance across three critical stages: sample selection, multi-scale feature integration, and multi-task learning objectives. To address this, the article introduces and evaluates "Libra R-CNN," a unified framework designed to rebalance detector training without introducing complex computational overhead.

The authors conducted extensive empirical evaluations using the benchmark Microsoft Common Objects in Context dataset, which contains over 115,000 training images. They tested their proposed modifications across multiple standard single-stage and two-stage model architectures, isolating the effect of each design component through systematic ablation experiments.

The study yielded several key findings. First, standard random sampling predominantly selects uninformative, easy background samples; introducing balanced sampling based on bounding box overlap increased detection accuracy by 0.9 points. Second, conventional multi-level feature integration dilutes semantic information across non-adjacent resolution layers, whereas integrating balanced features simultaneously boosted accuracy by another 0.9 points. Third, standard localization losses allow large gradients from coarse bounding box errors to drown out the smaller gradients needed for precise localization; introducing a balanced regression loss further improved accuracy by 0.8 to 1.3 points. Overall, Libra R-CNN achieved a 2.5-point gain over baseline Faster R-CNN and a 2.0-point gain over RetinaNet on standard benchmark metrics. When applied to proposal generation, it improved high-confidence recall by 9.2 points.

These findings indicate that addressing training imbalances provides a highly efficient path to performance gains without redesigning core architectures or incurring high computational costs. Previous mining techniques often introduced heavy memory requirements or sensitivity to noisy labels, but the proposed approach improves accuracy while remaining lightweight and modular.

Engineering and research teams deploying computer vision detectors should adopt balanced sample selection, feature integration, and loss scaling in their existing training pipelines. Because the framework integrates easily with popular computer vision backbones and feature architectures, organizations can upgrade existing systems with minimal workflow disruption.

Confidence in these findings is high due to consistent gains across varied model backbones and multiple benchmark evaluation splits. However, readers should note that the sampling improvements mainly benefit background region handling, as positive training candidate counts remain constrained by dataset annotations.

  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Feature Pyramid Networks establishes the multi-scale feature pyramid architecture that Libra R-CNN directly modifies with its balanced feature pyramid.
  • Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). This paper analyzes sample imbalance in dense object detection and introduces RetinaNet and Focal Loss, key baselines and conceptual foundations for Libra R-CNN's balanced learning objectives.
  • Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Faster R-CNN introduces the standard two-stage region-based detection paradigm whose sample selection, feature extraction, and loss optimization are revisited and balanced in Libra R-CNN.
  • Paper: Training Region-Based Object Detectors with Online Hard Example Mining, Abhinav Shrivastava et al. (2016). Online Hard Example Mining introduces automated loss-based sample mining for region-based detectors, addressing the sample-level imbalance that Libra R-CNN refines with IoU-balanced sampling.
  • Paper: Cascade R-CNN: Delving Into High Quality Object Detection, Zhaowei Cai et al. (2017). Cascade R-CNN analyzes the relationship between proposal IoU distributions and multi-stage detector performance, motivating Libra R-CNN's IoU-balanced sampling strategy.
  • Paper: Fast R-CNN, Ross B. Girshick (2015). Fast R-CNN establishes the multi-task smooth L1 and cross-entropy loss formulation that Libra R-CNN redesigns with its balanced L1 loss.
Cover for Libra R-CNN: Towards Balanced Learning for Object Detection

Abstract

Compared with model architectures, the training process, which is also crucial to the success of detectors, has received relatively less attention in object detection. In this work, we carefully revisit the standard training practice of detectors, and find that the detection performance is often limited by the imbalance during the training process, which generally consists in three levels - sample level, feature level, and objective level. To mitigate the adverse effects caused thereby, we propose Libra R-CNN, a simple but effective framework towards balanced learning for object detection. It integrates three novel components: IoU-balanced sampling, balanced feature pyramid, and balanced L1 loss, respectively for reducing the imbalance at sample, feature, and objective level. Benefitted from the overall balanced design, Libra R-CNN significantly improves the detection performance. Without bells and whistles, it achieves 2.5 points and 2.0 points higher Average Precision (AP) than FPN Faster R-CNN and RetinaNet respectively on MSCOCO.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 IoU-balanced Sampling
  • 3.2 Balanced Feature Pyramid
  • 3.3 Balanced L1 Loss
  • 4 Experiments
  • 4.1 Dataset and Evaluation Metrics
  • 4.2 Implementation Details
  • 4.3 Main Results
  • 4.4 Ablation Experiments
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Libra R-CNN Detection Framework

    model/method

    Libra R-CNN is an object detection framework designed to mitigate training imbalances that occur across three levels in conventional detection pipelines:

    1. Sample-level imbalance: Random negative sampling is dominated by easy background samples. Libra R-CNN employs IoU-balanced sampling to partition negative candidate regions into IoU bins, mining hard negatives across all overlap intervals without extra computational overhead.
    2. Feature-level imbalance: Multi-level feature pyramids (such as in FPN) sequentially integrate features, causing non-adjacent feature levels to be diluted during information flow. Libra R-CNN introduces a Balanced Feature Pyramid that integrates features across all pyramid levels simultaneously into a shared balanced semantic feature representation, refines it via non-local attention, and rescales it back to enhance each pyramid level.
    3. Objective-level imbalance: In multi-task loss formulation, gradients from easy inlier samples are overwhelmed by gradients from hard outlier samples during bounding-box regression. Libra R-CNN uses Balanced L1 Loss to amplify the gradients of accurate inlier samples while maintaining bounded gradients for outliers.

    The framework can be integrated into two-stage detectors (e.g., Faster R-CNN) using all three balancing components, or single-stage detectors (e.g., RetinaNet) using feature and objective balancing.

  2. Knowl 2 — Balanced L1 Loss Formulation

    equation

    Balanced L1 loss, denoted as Lb(x)L_b(x), is a bounding box regression loss designed to promote the gradient magnitude of accurate inliers (∣x∣<1|x| < 1) while clipping the gradient magnitude of hard outliers (∣x∣≥1|x| \ge 1). Here, x=tiu−vix = t_i^u - v_i denotes the difference between a predicted bounding box coordinate tiut_i^u (for class uu and coordinate index i∈{x,y,w,h}i \in \{x, y, w, h\}) and the corresponding regression target viv_i.

    The derivative of the balanced L1 loss with respect to the regression error xx is defined as:

    ∂Lb∂x={αln⁡(b∣x∣+1)if ∣x∣<1γotherwise\frac{\partial L_b}{\partial x} = \begin{cases} \alpha \ln(b|x| + 1) & \text{if } |x| < 1 \\ \gamma & \text{otherwise} \end{cases}

    where α\alpha controls gradient magnification for inliers, and γ\gamma sets the gradient upper bound for outliers. Integrating this derivative yields the loss function:

    Lb(x)={αb(b∣x∣+1)ln⁡(b∣x∣+1)−α∣x∣if ∣x∣<1γ∣x∣+CotherwiseL_b(x) = \begin{cases} \frac{\alpha}{b}(b|x| + 1)\ln(b|x| + 1) - \alpha |x| & \text{if } |x| < 1 \\ \gamma |x| + C & \text{otherwise} \end{cases}

    subject to the parameter constraint ensuring derivative continuity at ∣x∣=1|x| = 1:

    αln⁡(b+1)=γ  ⟺  b=eγ/α−1\alpha \ln(b + 1) = \gamma \iff b = e^{\gamma / \alpha} - 1

    The integration constant CC ensures loss continuity at ∣x∣=1|x| = 1:

    C=αb(b+1)ln⁡(b+1)−α−γ=γ(b+1)b−α−γC = \frac{\alpha}{b}(b+1)\ln(b+1) - \alpha - \gamma = \frac{\gamma(b+1)}{b} - \alpha - \gamma

    The default hyperparameter values are α=0.5\alpha = 0.5 and γ=1.5\gamma = 1.5 (giving b=e3−1≈19.0855b = e^3 - 1 \approx 19.0855).

    The total bounding box localization loss is the sum over all coordinates:

    Lloc=∑i∈{x,y,w,h}Lb(tiu−vi)L_{\text{loc}} = \sum_{i \in \{x, y, w, h\}} L_b(t_i^u - v_i)

  3. Knowl 3 — Balanced Feature Pyramid

    model/method

    The Balanced Feature Pyramid integrates multi-level feature representations {Clmin⁡,…,Clmax⁡}\{C_{l_{\min}}, \dots, C_{l_{\max}}\} (e.g., {C2,C3,C4,C5}\{C_2, C_3, C_4, C_5\}, where C2C_2 has the highest resolution and C5C_5 the lowest) through a four-step pipeline:

    1. Rescaling: All multi-level feature maps {C2,C3,C4,C5}\{C_2, C_3, C_4, C_5\} are resized to a common intermediate resolution (matching level C4C_4). Higher-resolution features (C2,C3C_2, C_3) are downsampled via max-pooling, and lower-resolution features (C5C_5) are upsampled via bilinear interpolation.
    2. Integrating: A balanced semantic feature map CC is computed by element-wise averaging across all L=lmax⁡−lmin⁡+1L = l_{\max} - l_{\min} + 1 rescaled levels:

    C=1L∑l=lmin⁡lmax⁡ClC = \frac{1}{L} \sum_{l=l_{\min}}^{l_{\max}} C_l

    1. Refining: The integrated balanced semantic feature map CC is refined to enhance discriminative ability using an embedded Gaussian non-local attention module.
    2. Strengthening: The refined feature map is resized back to each original resolution level (using bilinear interpolation for higher resolutions or max-pooling for lower resolutions) and added element-wise as a residual connection to each original feature map, producing the enhanced multi-level pyramid outputs {P2,P3,P4,P5}\{P_2, P_3, P_4, P_5\}.
  4. Knowl 4 — IoU-Balanced Sampling

    model/method

    IoU-balanced sampling is a hard negative mining technique that addresses sample-level imbalance in region proposal sampling. In conventional random sampling, background candidates are predominantly easy negatives with low ground-truth IoU, burying hard negatives (IoU ≥0.05\ge 0.05).

    Given MM total candidate negative samples and a required sample size of NN, the IoU overlap interval [0,0.5)[0, 0.5) is uniformly partitioned into KK equal sub-intervals (bins). The NN requested negative samples are distributed equally with N/KN/K samples allocated to each bin. Within each bin kk, candidates are drawn uniformly at random.

    The selection probability pkp_k for a candidate proposal in interval kk containing MkM_k candidate proposals is:

    pk=NK⋅1Mk,k∈{0,1,…,K−1}p_k = \frac{N}{K} \cdot \frac{1}{M_k}, \quad k \in \{0, 1, \dots, K - 1\}

    The default number of bins is K=3K = 3. If a bin contains fewer than N/KN/K candidates, the remaining quota is drawn from other bins. For positive samples, sampling an equal number of positive proposals per ground-truth bounding box serves as a complementary balancing step.

  5. Knowl 5 — Object Detection Benchmark Results on MS COCO Test-Dev

    data/table

    Libra R-CNN and Libra RetinaNet were evaluated on the MS COCO test-dev benchmark against baseline and state-of-the-art detectors across various backbone architectures and training schedules (1×=121\times = 12 epochs, 2×=242\times = 24 epochs).

    Method Backbone Schedule AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    YOLOv2 DarkNet-19 - 21.6 44.0 19.2 5.0 22.4 35.5
    SSD512 ResNet-101 - 31.2 50.4 33.3 10.2 34.5 49.8
    RetinaNet ResNet-101-FPN - 39.1 59.1 42.3 21.8 42.7 50.2
    Faster R-CNN ResNet-101-FPN - 36.2 59.1 39.0 18.2 39.0 48.2
    Deformable R-FCN Inception-ResNet-v2 - 37.5 58.0 40.8 19.4 40.1 52.5
    Mask R-CNN ResNet-101-FPN - 38.2 60.3 41.7 20.1 41.1 50.2
    Faster R-CNN* ResNet-50-FPN 1×1\times 36.2 58.5 38.9 21.0 38.9 45.3
    Faster R-CNN* ResNet-101-FPN 1×1\times 38.8 60.9 42.1 22.6 42.4 48.5
    Faster R-CNN* ResNet-101-FPN 2×2\times 39.7 61.3 43.4 22.1 43.1 50.3
    Faster R-CNN* ResNeXt-101-FPN 1×1\times 41.9 63.9 45.9 25.0 45.3 52.3
    RetinaNet* ResNet-50-FPN 1×1\times 35.8 55.3 38.6 20.0 39.0 45.1
    Libra R-CNN (ours) ResNet-50-FPN 1×1\times 38.7 59.9 42.0 22.5 41.1 48.7
    Libra R-CNN (ours) ResNet-101-FPN 1×1\times 40.3 61.3 43.9 22.9 43.1 51.0
    Libra R-CNN (ours) ResNet-101-FPN 2×2\times 41.1 62.1 44.7 23.4 43.7 52.5
    Libra R-CNN (ours) ResNeXt-101-FPN 1×1\times 43.0 64.0 47.0 25.3 45.6 54.6
    Libra RetinaNet (ours) ResNet-50-FPN 1×1\times 37.8 56.9 40.5 21.2 40.9 47.7

    Libra R-CNN improves upon baseline FPN Faster R-CNN by +2.5 AP+2.5\text{ AP} with ResNet-50 (38.738.7 vs 36.236.2) and +1.5 AP+1.5\text{ AP} with ResNet-101 (40.340.3 vs 38.838.8). On ResNeXt-101-64x4d, Libra R-CNN reaches 43.0 AP43.0\text{ AP}. When adapted to single-stage detection, Libra RetinaNet (incorporating Balanced Feature Pyramid and Balanced L1 Loss) yields 37.8 AP37.8\text{ AP} with ResNet-50, a +2.0 AP+2.0\text{ AP} improvement over baseline RetinaNet (35.8 AP35.8\text{ AP}). Asterisks (*) denote re-implemented baselines.

  6. Knowl 6 — Component-Wise Ablation of Libra R-CNN

    data/table

    An ablation study on MS COCO val-2017 measures the cumulative impact of adding IoU-balanced sampling, balanced feature pyramid, and balanced L1 loss to a Faster R-CNN baseline with ResNet-50-FPN (using identical pre-computed proposals):

    IoU-balanced Sampling Balanced Feature Pyramid Balanced L1 Loss AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    35.9 58.0 38.4 21.2 39.5 46.4
    ✓ 36.8 58.0 40.0 21.1 40.3 48.2
    ✓ ✓ 37.7 59.4 40.9 22.4 41.3 49.3
    ✓ ✓ ✓ 38.5 59.3 42.0 22.9 42.1 50.5

    Key results:

    • Adding IoU-balanced sampling increases AP from 35.935.9 to 36.836.8 (+0.9+0.9), with a substantial +1.6+1.6 increase in AP75\text{AP}_{75}.
    • Adding the Balanced Feature Pyramid raises AP from 36.836.8 to 37.737.7 (+0.9+0.9), providing consistent gains across small (APS+1.3\text{AP}_S +1.3), medium (APM+1.0\text{AP}_M +1.0), and large scales (APL+1.1\text{AP}_L +1.1).
    • Adding Balanced L1 Loss brings the AP to 38.538.5 (+0.8+0.8), predominantly improving localization accuracy (AP75\text{AP}_{75} increases from 40.940.9 to 42.042.0, +1.1+1.1).
    • The three components combined achieve an overall improvement of +2.6 AP+2.6\text{ AP} over the baseline.
  7. Knowl 7 — Region Proposal Recall Improvements with Libra RPN

    data/table

    Applying the balanced training design of Libra R-CNN to the Region Proposal Network (Libra RPN) improves proposal recall across different proposal budgets on MS COCO.

    Method Backbone AR100\text{AR}^{100} AR300\text{AR}^{300} AR1000\text{AR}^{1000}
    RPN∗\text{RPN}^* ResNet-50-FPN 42.5 51.2 57.1
    RPN∗\text{RPN}^* ResNet-101-FPN 45.4 53.2 58.7
    RPN∗\text{RPN}^* ResNeXt-101-FPN 47.8 55.0 59.8
    Libra RPN (ours) ResNet-50-FPN 52.1 58.3 62.5

    Libra RPN with a ResNet-50 backbone improves average recall over baseline ResNet-50 RPN by +9.2+9.2 on AR100\text{AR}^{100} (52.152.1 vs 42.542.5), +6.9+6.9 on AR300\text{AR}^{300} (58.358.3 vs 51.251.2), and +5.4+5.4 on AR1000\text{AR}^{1000} (62.562.5 vs 57.157.1). Furthermore, Libra RPN with ResNet-50 achieves 4.34.3 points higher AR100\text{AR}^{100} than baseline RPN using the much larger ResNeXt-101-64x4d backbone (52.152.1 vs 47.847.8). Asterisks (*) denote re-implemented baselines.

  8. Knowl 8 — Ablation of IoU-Balanced Sampling Configurations

    data/table

    Ablation experiments on MS COCO val-2017 evaluate positive sample balancing and the sensitivity of IoU-balanced sampling to the number of intervals KK on a Faster R-CNN ResNet-50-FPN baseline:

    Settings AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    Baseline 35.9 58.0 38.4 21.2 39.5 46.4
    Pos Balance 36.1 58.2 38.2 21.3 40.2 47.3
    K=2K = 2 36.7 57.8 39.9 20.5 39.9 48.9
    K=3K = 3 36.8 57.9 39.8 21.4 39.9 48.7
    K=5K = 5 36.7 57.7 39.9 19.9 40.1 48.7

    Sampling equal numbers of positive samples per ground truth ("Pos Balance") yields a +0.2 AP+0.2\text{ AP} gain (36.136.1 vs 35.935.9), limited by the small candidate pool of positive proposals. For negative sampling, dividing into K=2,3,K = 2, 3, or 55 IoU bins consistently yields +0.8+0.8 to +0.9 AP+0.9\text{ AP} gains (36.736.7, 36.836.8, and 36.736.7 respectively), demonstrating that performance is not sensitive to the exact value of KK as long as hard negatives receive elevated sampling probability.

  9. Knowl 9 — Ablation of Balanced Feature Pyramid Configurations

    data/table

    Ablation experiments on MS COCO val-2017 evaluate the components of the Balanced Feature Pyramid and its integration with Path Aggregation Feature Pyramid Network (PAFPN):

    Settings AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    Baseline (FPN) 35.9 58.0 38.4 21.2 39.5 46.4
    Integration 36.3 58.8 38.8 21.2 40.1 46.3
    Refinement 36.8 59.5 39.5 22.3 40.6 46.5
    PAFPN 36.3 58.4 39.0 21.7 39.9 46.3
    Balanced PAFPN 37.2 60.0 39.8 22.7 40.8 47.4

    Key observations:

    • Nonparametric feature integration alone (rescaling to intermediate size and averaging without refinement or added parameters) achieves 36.3 AP36.3\text{ AP} (+0.4+0.4 over baseline), matching PAFPN performance without extra convolutions.
    • Adding non-local attention refinement brings the AP to 36.836.8 (+0.9+0.9 over baseline).
    • Incorporating the balanced feature scheme into PAFPN (Balanced PAFPN) achieves 37.2 AP37.2\text{ AP}, which is +0.9 AP+0.9\text{ AP} higher than PAFPN alone.
  10. Knowl 10 — Ablation and Parameter Sensitivity of Balanced L1 Loss

    data/table

    Ablation experiments on MS COCO val-2017 evaluate Balanced L1 Loss against direct loss weight tuning and standard L1 loss on a Faster R-CNN ResNet-50-FPN baseline (35.9 AP35.9\text{ AP}):

    Settings AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    Baseline 35.9 58.0 38.4 21.2 39.5 46.4
    loss weight = 1.5 36.4 58.0 39.7 20.8 39.9 47.5
    loss weight = 2.0 36.2 57.3 39.5 20.2 40.0 47.5
    L1 Loss (1.0) 36.4 57.4 39.1 21.0 39.7 47.9
    L1 Loss (1.5) 36.6 57.2 39.8 20.2 40.0 48.2
    L1 Loss (2.0) 36.4 56.5 39.6 20.1 39.8 48.2
    α=0.2,γ=1.0\alpha = 0.2, \gamma = 1.0 36.7 58.1 39.5 21.4 40.4 47.4
    α=0.3,γ=1.0\alpha = 0.3, \gamma = 1.0 36.5 58.2 39.2 21.6 40.2 47.2
    α=0.5,γ=1.0\alpha = 0.5, \gamma = 1.0 36.5 58.2 39.2 21.5 39.9 47.2
    α=0.5,γ=1.5\alpha = 0.5, \gamma = 1.5 37.2 58.0 40.0 21.3 40.9 47.9
    α=0.5,γ=2.0\alpha = 0.5, \gamma = 2.0 37.0 58.0 40.0 21.2 40.8 47.6

    Directly scaling the Smooth L1 loss weight reaches a maximum of 36.4 AP36.4\text{ AP} (at weight 1.51.5) before falling to 36.2 AP36.2\text{ AP} (at weight 2.02.0) because large outlier gradients distort optimization. Standard L1 loss causes notable drops in AP50\text{AP}_{50} and APS\text{AP}_S. Balanced L1 loss with γ=1.0\gamma = 1.0 achieves up to 36.7 AP36.7\text{ AP}, and combining inlier gradient magnification (α=0.5\alpha = 0.5) with overall scale promotion (γ=1.5\gamma = 1.5) achieves the optimal performance of 37.2 AP37.2\text{ AP} (+1.3 AP+1.3\text{ AP} over baseline).

Coverage note — None was omitted; all key architectural components, sampling methods, loss formulations, and comprehensive ablation/benchmark results from the paper's contribution are covered.

References

  1. 1.Sean Bell, C Lawrence Zitnick, Kavita Bala, and Ross Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  2. 2.Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In European Conference on Computer Vision, 2016.
  3. 3.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  4. 4.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, and Dahua Lin. Hybrid task cascade for instance segmentation. arXiv preprint arXiv:1901.07518, 2019.
  5. 5.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. mmdetection. https://github.com/open-mmlab/mmdetection, 2018.
  6. 6.Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. R-fcn: Object detection via region-based fully convolutional networks. In Advances in Neural Information Processing Systems, 2016.
  7. 7.Ross Girshick. Fast r-cnn. In IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  8. 8.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2014.
  9. 9.Ross Girshick, Ilija Radosavovic, Georgia Gkioxari, Piotr Doll'ar, and Kaiming He. Detectron. https://github.com/facebookresearch/detectron, 2018.
  10. 10.Kaiming He, Georgia Gkioxari, Piotr Doll'ar, and Ross Girshick. Mask r-cnn. In IEEE International Conference on Computer Vision, 2017.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In European Conference on Computer Vision, 2014.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  13. 13.Jan Hendrik Hosang, Rodrigo Benenson, and Bernt Schiele. Learning non-maximum suppression. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  14. 14.Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei. Relation networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  15. 15.Borui Jiang, Ruixuan Luo, Jiayuan Mao, Tete Xiao, and Yuning Jiang. Acquisition of localization confidence for accurate object detection. arXiv preprint arXiv:1807.11590, 1, 2018.
  16. 16.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. arXiv preprint arXiv:1705.07115, 3, 2017.
  17. 17.Tao Kong, Fuchun Sun, Wenbing Huang, and Huaping Liu. Deep feature pyramid reconfiguration for object detection. arXiv preprint arXiv:1808.07993, 2018.
  18. 18.Hei Law and Jia Deng. Cornernet: Detecting objects as paired keypoints. In European Conference on Computer Vision, 2018.
  19. 19.Tsung-Yi Lin, Piotr Doll'ar, Ross B Girshick, Kaiming He, Bharath Hariharan, and Serge J Belongie. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  20. 20.Tsung-Yi Lin, Priyal Goyal, Ross Girshick, Kaiming He, and Piotr Doll'ar. Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  21. 21.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll'ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014.
  22. 22.Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  23. 23.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European Conference on Computer Vision, 2016.
  24. 24.Wanli Ouyang, Kun Wang, Xin Zhu, and Xiaogang Wang. Chained cascade network for object detection. In IEEE International Conference on Computer Vision, 2017.
  25. 25.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017.
  26. 26.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  27. 27.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. arXiv preprint, 2017.
  28. 28.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, 2015.
  29. 29.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. Training region-based object detectors with online hard example mining. In IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  30. 30.Bharat Singh and Larry S Davis. An analysis of scale invariance in object detection–snip. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  31. 31.Bharat Singh, Mahyar Najibi, and Larry S Davis. SNIPER: Efficient multi-scale training. NIPS, 2018.
  32. 32.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. arXiv preprint arXiv:1711.07971, 10, 2017.
  33. 33.Saining Xie, Ross Girshick, Piotr Doll'ar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  34. 34.Jiahui Yu, Yuning Jiang, Zhangyang Wang, Zhimin Cao, and Thomas Huang. Unitbox: An advanced object detection network. In Proceedings of the 24th ACM international conference on Multimedia, pages 516–520. ACM, 2016.
  35. 35.Matthew D Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In European Conference on Computer Vision, 2014.
  36. 36.Xingyu Zeng, Wanli Ouyang, Junjie Yan, Hongsheng Li, Tong Xiao, Kun Wang, Yu Liu, Yucong Zhou, Bin Yang, Zhe Wang, et al. Crafting gbd-net for object detection. IEEE transactions on pattern analysis and machine intelligence, 40(9):2109–2123, 2018.
  37. 37.Xinge Zhu, Jiangmiao Pang, Ceyuan Yang, Jianping Shi, and Dahua Lin. Adapting object detectors via selective crossdomain alignment. In IEEE Conference on Computer Vision and Pattern Recognition, 2019.

Citation

MLA
Pang, J., et al. “Libra R-CNN: Towards Balanced Learning for Object Detection”. arXiv, 2019, http://arxiv.org/abs/1904.02701v1.
APA
Pang, J., Chen, K., Shi, J., Feng, H., Ouyang, W., & Lin, D. (2019). Libra R-CNN: Towards Balanced Learning for Object Detection. arXiv. http://arxiv.org/abs/1904.02701v1
Chicago
Pang, J., K. Chen, J. Shi, H. Feng, W. Ouyang, and D. Lin. 2019. “Libra R-CNN: Towards Balanced Learning for Object Detection”. arXiv. http://arxiv.org/abs/1904.02701v1.
Harvard
Pang, J. et al. (2019) “Libra R-CNN: Towards Balanced Learning for Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.02701v1.
Vancouver
1. Pang J, Chen K, Shi J, Feng H, Ouyang W, Lin D (2019) Libra R-CNN: Towards Balanced Learning for Object Detection. arXiv

BibTeX

@article{pang2019libra,
  title = {Libra R-CNN: Towards Balanced Learning for Object Detection},
  author = {Pang, Jiangmiao and Chen, Kai and Shi, Jianping and Feng, Huajun and Ouyang, Wanli and Lin, Dahua},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.02701v1},
  eprint = {1904.02701}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE