CenterNet: Keypoint Triplets for Object Detection

Kaiwen DuanSong BaiLingxi XieHonggang QiQingming HuangQi Tian

article2019ICCV3,526 citations

Presents CenterNet, an object detection framework that models objects as keypoint triplets using center and cascade corner pooling to filter out incorrect bounding boxes and achieve state-of-the-art one-stage detection accuracy on MS-COCO.

Listen

CenterNet addresses a persistent weakness in keypoint-based object detection: methods such as CornerNet generate many bounding boxes that do not align well with actual objects because they rely solely on paired corner keypoints and lack any check of the interior region. This produces high rates of false detections, especially for small objects, and limits overall precision even when recall is reasonable. The work therefore sets out to add a lightweight internal verification step that keeps the speed advantage of one-stage detectors while improving accuracy.

The authors extend CornerNet by representing each object as a triplet of keypoints—two corners plus one center keypoint—rather than a corner pair alone. After candidate boxes are formed from corners, a scale-aware central region is examined; if a center keypoint of matching class appears inside it, the box is retained and its score is adjusted by averaging the three keypoints. Two new pooling modules support this design: center pooling gathers stronger internal signals for the center keypoint, and cascade corner pooling lets corners capture both boundary and interior evidence. The resulting network is trained from scratch on the MS-COCO trainval35k set and evaluated on the test-dev set using 52-layer and 104-layer hourglass backbones.

On single-scale testing with the deeper backbone, CenterNet reaches 44.9 percent average precision, a 4.4-point gain over the CornerNet baseline; multi-scale testing raises the figure to 47.0 percent, surpassing all published one-stage detectors by at least 4.9 points and matching or approaching the best two-stage systems. Gains are largest on small objects (up to 8.1 points), and false-discovery rates drop noticeably at every IoU threshold. Inference requires 270–340 ms per image on a P100 GPU, remaining faster than most two-stage alternatives.

These results indicate that a modest central-region check can close much of the accuracy gap between one-stage and two-stage detectors without incurring the cost of region-of-interest pooling. The approach therefore offers practitioners a practical route to higher precision in real-time or resource-constrained settings. The authors note that the same center-keypoint branch could be added to other one-stage frameworks and that further gains are likely with improved center-keypoint training. Performance remains sensitive to center-keypoint accuracy, however; substituting ground-truth centers lifts AP by roughly 14 points, showing that center detection is still a remaining bottleneck.

  • Paper: CornerNet: Detecting Objects as Paired Keypoints, Hei Law et al. (2018). CenterNet directly extends CornerNet’s paired-corner detector by adding a center keypoint to verify candidate boxes, so CornerNet’s representation and limitations clarify the design.
  • Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). CenterNet uses focal loss in its keypoint-based detector, and understanding how focal loss addresses dense foreground–background imbalance helps explain its training setup.

No sufficiently relevant recommendations were found.

Cover for CenterNet: Keypoint Triplets for Object Detection

Abstract

In object detection, keypoint-based approaches often suffer a large number of incorrect object bounding boxes, arguably due to the lack of an additional look into the cropped regions. This paper presents an efficient solution which explores the visual patterns within each cropped region with minimal costs. We build our framework upon a representative one-stage keypoint-based detector named CornerNet. Our approach, named CenterNet, detects each object as a triplet, rather than a pair, of keypoints, which improves both precision and recall. Accordingly, we design two customized modules named cascade corner pooling and center pooling, which play the roles of enriching information collected by both top-left and bottom-right corners and providing more recognizable information at the central regions, respectively. On the MS-COCO dataset, CenterNet achieves an AP of 47.0%, which outperforms all existing one-stage detectors by at least 4.9%. Meanwhile, with a faster inference speed, CenterNet demonstrates quite comparable performance to the top-ranked two-stage detectors. Code is available at this https URL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Baseline and Motivation
  • 3.2. Object Detection as Keypoint Triplets
  • 3.3. Enriching Center and Corner Information
  • 3.4. Training and Inference
  • 4. Experiments
  • 4.1. Dataset, Metrics and Baseline
  • 4.2. Comparisons with State-of-the-art Detectors
  • 4.3. Incorrect Bounding Box Reduction
  • 4.4. Inference Speed
  • 4.5. Ablation Study
  • 4.6. Error Analysis
  • 5. Conclusions

Knowls

  1. Knowl 1 — CenterNet Keypoint Triplet Detection Framework

    model/method

    CenterNet represents each object as a triplet of keypoints consisting of a top-left corner, a bottom-right corner, and a center keypoint, building upon the one-stage CornerNet architecture to incorporate internal object context with low computational overhead.

    A convolutional backbone (such as a 52-layer or 104-layer stacked Hourglass network) extracts visual features and outputs three prediction branches:

    1. Top-left corner branch: produces a class-specific corner heatmap, corner embeddings for keypoint grouping, and 2D sub-pixel coordinate offsets.
    2. Bottom-right corner branch: produces a class-specific corner heatmap, corner embeddings, and 2D coordinate offsets.
    3. Center keypoint branch: produces a class-specific center heatmap and 2D coordinate offsets.

    During bounding box generation, candidate boxes are initially formed by pairing top-left and bottom-right corners whose embedding vector distance falls below a predefined threshold. To filter out false positive boxes, the model defines a scale-aware central region inside the proposed box and verifies whether a center keypoint of the identical category is present within that region. If a valid center keypoint is detected, the bounding box is preserved and its final confidence score is updated to the arithmetic mean of the top-left corner score, bottom-right corner score, and center keypoint score. If no matching center keypoint exists in the central region, the candidate box is discarded.

  2. Knowl 2 — Scale-Aware Central Region Formulation

    equation

    To account for scale variance across objects, the central region inside a candidate bounding box is defined adaptively based on the box scale. For a candidate bounding box ii with top-left coordinates (tlx,tly)(tl_x, tl_y) and bottom-right coordinates (brx,bry)(br_x, br_y), the bounding coordinates of the corresponding central region jj, denoted by top-left corner (ctlx,ctly)(ctl_x, ctly) and bottom-right corner (cbrx,cbry)(cbr_x, cbry), are given by:

    {ctlx=(n+1)tlx+(n−1)brx2nctly=(n+1)tly+(n−1)bry2ncbrx=(n−1)tlx+(n+1)brx2ncbry=(n−1)tly+(n+1)bry2n\begin{cases} ctl_x = \dfrac{(n + 1)tl_x + (n - 1)br_x}{2n} \\[6pt] ctly = \dfrac{(n + 1)tl_y + (n - 1)bry}{2n} \\[6pt] cbr_x = \dfrac{(n - 1)tl_x + (n + 1)br_x}{2n} \\[6pt] cbry = \dfrac{(n - 1)tly + (n + 1)bry}{2n} \end{cases}

    where nn is an odd integer parameter controlling the scale ratio of the central region relative to the candidate bounding box:

    • n=3n = 3 is used when the scale (maximum dimension) of the candidate bounding box is less than 150150 pixels, defining a central region occupying 1/31/3 of the box width and height (area fraction 1/91/9).
    • n=5n = 5 is used when the scale of the candidate bounding box is greater than or equal to 150150 pixels, defining a central region occupying 1/51/5 of the box width and height (area fraction 1/251/25).

    This scale-aware definition assigns a relatively larger central search area to small objects to preserve recall, and a tighter central search area to large objects to maintain precision.

  3. Knowl 3 — Center Pooling Module

    model/method

    Geometric centers of objects often lack distinct local visual features (for example, the center of a human body may fall on clothing without salient shape cues). Center pooling is designed to capture non-local contextual visual patterns in both horizontal and vertical directions across the object body.

    Given an intermediate feature map F∈RC×H×WF \in \mathbb{R}^{C \times H \times W}, center pooling computes maximum response vectors along both the horizontal and vertical directions and sums them:

    1. Horizontal response: obtained by applying left-pooling (taking maximum values from right to left along each row) and right-pooling (taking maximum values from left to right along each row) connected in series.
    2. Vertical response: obtained by applying top-pooling (taking maximum values from bottom to top along each column) and bottom-pooling (taking maximum values from top to bottom along each column) connected in series.
    3. Output combination: the horizontal and vertical pooled feature representations are summed element-wise and passed through a 3×33 \times 3 Conv-BN-ReLU layer followed by 1×11 \times 1 convolutions to predict center keypoint heatmaps and sub-pixel offsets.
  4. Knowl 4 — Cascade Corner Pooling Module

    model/method

    Standard corner pooling scans only along the object boundaries (e.g., horizontally and vertically towards image edges), which makes detected corner keypoints sensitive to background edge clutter and prevents them from leveraging visual patterns within the object interior. Cascade corner pooling enriches corner feature representations by combining boundary maximum values with internal feature maximum values.

    For a given corner direction, cascade corner pooling first computes the maximum response along the object boundary, and then traverses internally perpendicular to that boundary at the location of the boundary maximum to locate the internal maximum response:

    • Cascade top-left corner pooling: for the top boundary, it applies a left corner pooling module before feeding the resulting feature map into the top corner pooling module, followed by summing the boundary and internal maximum responses.
    • Cascade bottom-right corner pooling: applies complementary right and bottom pooling passes in series and parallel combinations to integrate internal visual cues into the corner representations.
  5. Knowl 5 — CenterNet Multi-Task Training Loss

    equation

    CenterNet is trained end-to-end from scratch using a multi-task loss function combining keypoint detection focal losses, associative embedding grouping losses, and sub-pixel coordinate offset regression losses:

    L=Ldetco+Ldetce+αLpullco+βLpushco+γ(Loffco+Loffce)\mathcal{L} = \mathcal{L}_{\text{det}}^{\text{co}} + \mathcal{L}_{\text{det}}^{\text{ce}} + \alpha \mathcal{L}_{\text{pull}}^{\text{co}} + \beta \mathcal{L}_{\text{push}}^{\text{co}} + \gamma \left(\mathcal{L}_{\text{off}}^{\text{co}} + \mathcal{L}_{\text{off}}^{\text{ce}}\right)

    where:

    • Ldetco\mathcal{L}_{\text{det}}^{\text{co}} is the variant focal loss for top-left and bottom-right corner heatmaps.
    • Ldetce\mathcal{L}_{\text{det}}^{\text{ce}} is the variant focal loss for center keypoint heatmaps.
    • Lpullco\mathcal{L}_{\text{pull}}^{\text{co}} is the associative embedding pull loss that minimizes the ℓ2\ell_2 distance between embedding vectors belonging to the same object corner pair.
    • Lpushco\mathcal{L}_{\text{push}}^{\text{co}} is the associative embedding push loss that maximizes the ℓ2\ell_2 distance (with a margin of 1) between embedding vectors belonging to corners of different objects.
    • Loffco\mathcal{L}_{\text{off}}^{\text{co}} and Loffce\mathcal{L}_{\text{off}}^{\text{ce}} are smooth ℓ1\ell_1 losses penalizing discretization errors in mapping corner and center heatmap locations back to the input image coordinates.
    • The loss balancing weights are set to α=0.1\alpha = 0.1, β=0.1\beta = 0.1, and γ=1\gamma = 1.
  6. Knowl 6 — CenterNet Triplet Inference and Filtering Procedure

    algorithm

    The inference pipeline processes the output heatmaps, associative embeddings, and sub-pixel offsets to detect objects as keypoint triplets:

    Input: Corner heatmaps Htl,Hbr∈RC×H×WH^{tl}, H^{br} \in \mathbb{R}^{C \times H \times W}; center heatmap Hce∈RC×H×WH^{ce} \in \mathbb{R}^{C \times H \times W}; corner embeddings Etl,EbrE^{tl}, E^{br}; offset maps Otl,Obr,OceO^{tl}, O^{br}, O^{ce}; embedding threshold τ\tau; scale threshold S=150S = 150; top keypoint count k=70k = 70
    Output: Final list of object bounding boxes BB
    Extract top-kk keypoints from Htl,Hbr,HceH^{tl}, H^{br}, H^{ce} across all categories
    Remap keypoint coordinates to image space using offsets Otl,Obr,OceO^{tl}, O^{br}, O^{ce}
    Initialize candidate proposal list P=∅P = \emptyset
    for each top-left corner ptl=(x1,y1,c,s1,e1)p^{tl} = (x_1, y_1, c, s_1, e_1) in top-kk corners do
        for each bottom-right corner pbr=(x2,y2,c,s2,e2)p^{br} = (x_2, y_2, c, s_2, e_2) in top-kk corners do
            if x2>x1x_2 > x_1 and y2>y1y_2 > y_1 and ∥e1−e2∥2<τ\|e_1 - e_2\|_2 < \tau then
                w=x2−x1w = x_2 - x_1
                h=y2−y1h = y_2 - y_1
                n=3n = 3 if max⁡(w,h)<S\max(w, h) < S else 55
                ctlx=((n+1)x1+(n−1)x2)/(2n)ctl_x = ((n + 1)x_1 + (n - 1)x_2) / (2n)
                ctly=((n+1)y1+(n−1)y2)/(2n)ctl_y = ((n + 1)y_1 + (n - 1)y_2) / (2n)
                cbrx=((n−1)x1+(n+1)x2)/(2n)cbr_x = ((n - 1)x_1 + (n + 1)x_2) / (2n)
                cbry=((n−1)y1+(n+1)y2)/(2n)cbr_y = ((n - 1)y_1 + (n + 1)y_2) / (2n)
                
                Find center keypoints pce=(xc,yc,c,sc)p^{ce} = (x_c, y_c, c, s_c) of class cc in region [ctlx,cbrx]×[ctly,cbry][ctl_x, cbr_x] \times [ctl_y, cbr_y]
                if matching center keypoint exists then
                    sbox=(s1+s2+max⁡(sc))/3s_{box} = (s_1 + s_2 + \max(s_c)) / 3
                    Append box (x1,y1,x2,y2,c,sbox)(x_1, y_1, x_2, y_2, c, s_{box}) to PP
                end if
            end if
        end for
    end for
    Apply Soft-NMS to PP and select top 100 bounding boxes to form BB
    return BB
  7. Knowl 7 — CenterNet Object Detection Performance on MS-COCO Test-Dev

    data/table

    Evaluation of CenterNet on MS-COCO test-dev across single-scale and multi-scale testing protocols compared to representative one-stage and two-stage detectors demonstrates state-of-the-art one-stage detection accuracy.

    Method Backbone Input AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    One-stage:
    YOLOv2 DarkNet-19 544×544544\times 544 21.6 44.0 19.2 5.0 22.4 35.5
    SSD513 ResNet-101 513×513513\times 513 31.2 50.4 33.3 10.2 34.5 49.8
    RefineDet512 ResNet-101 512×512512\times 512 36.4 57.5 39.5 16.6 39.9 51.4
    RetinaNet800 ResNet-101 800×800800\times 800 39.1 59.1 42.3 21.8 42.7 50.2
    CornerNet511 (single) Hourglass-52 511×511511\times 511 37.8 53.7 40.1 17.0 39.0 50.5
    CornerNet511 (multi) Hourglass-52 511×511511\times 511 39.4 54.9 42.3 18.9 41.2 52.7
    CornerNet511 (single) Hourglass-104 511×511511\times 511 40.5 56.5 43.1 19.4 42.7 53.9
    CornerNet511 (multi) Hourglass-104 511×511511\times 511 42.1 57.8 45.3 20.8 44.8 56.7
    CenterNet511 (single) Hourglass-52 511×511511\times 511 41.6 59.4 44.2 22.5 43.1 54.1
    CenterNet511 (multi) Hourglass-52 511×511511\times 511 43.5 61.3 46.7 25.3 45.3 55.0
    CenterNet511 (single) Hourglass-104 511×511511\times 511 44.9 62.4 48.1 25.6 47.4 57.4
    CenterNet511 (multi) Hourglass-104 511×511511\times 511 47.0 64.5 50.7 28.9 49.9 58.9
    Two-stage:
    Mask R-CNN ResNeXt-101 ∼1300×800\sim 1300\times 800 39.8 62.3 43.4 22.1 43.2 51.2
    Cascade R-CNN ResNet-101 – 42.8 62.1 46.3 23.7 45.5 55.2
    PANet (multi) ResNeXt-101 ∼1400×840\sim 1400\times 840 47.4 67.2 51.8 30.1 51.7 60.0

    CenterNet511-104 achieves 47.0% multi-scale AP and 44.9% single-scale AP on MS-COCO test-dev, improving over CornerNet by +4.9% and +4.4% AP, respectively. The largest gains occur on small objects (extAPS ext{AP}_S rises from 20.8% to 28.9% under multi-scale testing, a +8.1% absolute improvement).

  8. Knowl 8 — Ablation of CenterNet Architectural Components

    data/table

    An ablation study on the MS-COCO validation dataset isolates the contributions of Central Region Exploration (CRE), Center Pooling (CTP), and Cascade Corner Pooling (CCP) using the Hourglass-52 backbone.

    CRE CTP CCP AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L AR1\text{AR}_1 AR10\text{AR}_{10} AR100\text{AR}_{100}
    37.6 53.3 40.0 18.5 39.6 52.2 33.7 52.2 56.7
    ✓ 38.3 54.2 40.5 18.6 40.5 52.2 34.0 53.0 57.9
    ✓ 39.9 57.7 42.3 23.1 42.3 52.3 33.8 54.2 58.5
    ✓ ✓ 40.8 58.6 43.6 23.6 43.6 53.6 33.9 54.5 59.0
    ✓ ✓ ✓ 41.3 59.2 43.9 23.6 43.8 55.8 34.5 55.0 59.2

    Key takeaways from the ablation data:

    1. Adding central region exploration (CRE) alone via conventional convolutions provides a +2.3% AP gain (37.6% to 39.9%), driven primarily by a +4.6% boost on small objects (extAPS ext{AP}_S from 18.5% to 23.1%).
    2. Incorporating center pooling (CTP) adds +0.9% overall AP (39.9% to 40.8%) and notably improves large object accuracy (extAPL ext{AP}_L from 52.3% to 53.6%).
    3. Incorporating cascade corner pooling (CCP) alongside CRE and CTP improves overall AP to 41.3% and enhances large object detection (extAPL ext{AP}_L increases to 55.8%).
  9. Knowl 9 — False Discovery Rate Analysis and Reduction in CenterNet

    data/table

    The false discovery rate (extFD=1−AP ext{FD} = 1 - \text{AP}) measures the proportion of predicted bounding boxes that are incorrect at various IoU evaluation thresholds and scales on the MS-COCO validation dataset.

    Method FD\text{FD} FD5\text{FD}_5 FD25\text{FD}_{25} FD50\text{FD}_{50} FDS\text{FD}_S FDM\text{FD}_M FDL\text{FD}_L
    CornerNet511-52 40.4 35.2 39.4 46.7 62.5 36.9 28.0
    CenterNet511-52 35.1 30.7 34.2 40.8 53.0 31.3 24.4
    CornerNet511-104 37.8 32.7 36.8 43.8 60.3 33.2 25.1
    CenterNet511-104 32.4 28.2 31.6 37.5 50.7 27.1 23.0

    Here, FDi=1−APi\text{FD}_i = 1 - \text{AP}_i represents false discoveries at IoU threshold i/100i/100, and FDscale=1−APscale\text{FD}_{scale} = 1 - \text{AP}_{scale}. The results show:

    1. CornerNet produces high false discovery rates even at loose IoU thresholds (32.7% FD5\text{FD}_5 for Hourglass-104, meaning nearly 1 in 3 predicted boxes has less than 0.05 IoU with ground truth).
    2. CenterNet reduces FD5\text{FD}_5 by 4.5% in both backbones (from 35.2% to 30.7% for Hourglass-52, and 32.7% to 28.2% for Hourglass-104).
    3. The greatest reduction in false discoveries occurs for small objects, where FDS\text{FD}_S drops by 9.5% (62.5% to 53.0% on Hourglass-52) and 9.6% (60.3% to 50.7% on Hourglass-104).
  10. Knowl 10 — Center Keypoint Oracle Error Analysis

    data/table

    To evaluate whether center keypoint detection acts as a performance ceiling for the triplet paradigm, an error analysis replaces predicted center keypoints with ground-truth center keypoint locations on the MS-COCO validation set.

    Method AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    CenterNet511-52 w/o GT 41.3 59.2 43.9 23.6 43.8 55.8
    CenterNet511-52 w/ GT 56.5 78.3 61.4 39.1 60.3 70.3
    CenterNet511-104 w/o GT 44.8 62.4 48.2 25.9 48.9 58.8
    CenterNet511-104 w/ GT 58.1 78.4 63.9 40.4 63.0 72.1

    Substituting ground-truth center keypoints raises AP from 41.3% to 56.5% (+15.2%) on CenterNet511-52 and from 44.8% to 58.1% (+13.3%) on CenterNet511-104. This large potential margin indicates that the center keypoint verification framework has substantial headroom and is not fundamentally bottlenecked by the triplet verification logic.

  11. Knowl 11 — Inference Latency and Efficiency Trade-offs

    empirical result

    Inference time evaluated on an NVIDIA Tesla P100 GPU demonstrates that CenterNet maintains high efficiency while adding central region verification:

    • CornerNet511-104 takes an average of 300 ms300\,\text{ms} per image.
    • CenterNet511-104 takes an average of 340 ms340\,\text{ms} per image, showing that central keypoint pooling and triplet verification introduce only a modest 40 ms40\,\text{ms} overhead.
    • CenterNet511-52 takes an average of 270 ms270\,\text{ms} per image, achieving 41.6%41.6\% single-scale AP on MS-COCO test-dev, which is both faster (−30 ms-30\,\text{ms}) and more accurate (+1.1%+1.1\% AP) than the baseline CornerNet511-104 model (300 ms300\,\text{ms}, 40.5%40.5\% AP).

Coverage note — None was omitted; all contributed models, equations, algorithms, empirical evaluations, and ablation studies from the paper are represented.

References

  1. 1.S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2874–2883, 2016.
  2. 2.N. Bodla, B. Singh, R. Chellappa, and L. S. Davis. Soft-nms–improving object detection with one line of code. In Proceedings of the IEEE international conference on computer vision, pages 5561–5569, 2017.
  3. 3.Z. Cai, Q. Fan, R. S. Feris, and N. Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In European conference on computer vision, pages 354–370. Springer, 2016.
  4. 4.Z. Cai and N. Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018.
  5. 5.Y. Chen, J. Li, H. Xiao, X. Jin, S. Yan, and J. Feng. Dual path networks. In Advances in neural information processing systems, pages 4467–4475, 2017.
  6. 6.J. Dai, Y. Li, K. He, and J. Sun. R-fcn: Object detection via region-based fully convolutional networks. In Advances in neural information processing systems, pages 379–387, 2016.
  7. 7.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 764–773, 2017.
  8. 8.C.-Y. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg. Dssd: Deconvolutional single shot detector. arXiv preprint arXiv:1701.06659, 2017.
  9. 9.S. Gidaris and N. Komodakis. Object detection via a multiregion and semantic segmentation-aware cnn model. In Proceedings of the IEEE international conference on computer vision, pages 1134–1142, 2015.
  10. 10.R. Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440–1448, 2015.
  11. 11.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
  12. 12.K. He, G. Gkioxari, P. Dollár, and R. Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 37(9):1904–1916, 2015.
  14. 14.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  15. 15.D. Hoiem, Y. Chodpathumwan, and Q. Dai. Diagnosing error in object detectors. In European conference on computer vision, pages 340–353. Springer, 2012.
  16. 16.J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, et al. Speed/accuracy trade-offs for modern convolutional object detectors. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7310–7311, 2017.
  17. 17.J. Jeong, H. Park, and N. Kwak. Enhancement of ssd by concatenating feature maps for object detection. arXiv preprint arXiv:1705.09587, 2017.
  18. 18.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. Computer science, 2014.
  19. 19.T. Kong, F. Sun, A. Yao, H. Liu, M. Lu, and Y. Chen. Ron: Reverse connection with objectness prior networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5936–5944, 2017.
  20. 20.H. Law and J. Deng. Cornernet: Detecting objects as paired keypoints. In Proceedings of the European conference on computer vision, pages 734–750, 2018.
  21. 21.H. Lee, S. Eum, and H. Kwon. Me r-cnn: Multi-expert r-cnn for object detection. arXiv preprint arXiv:1704.01069, 2017.
  22. 22.Y. Li, Y. Chen, N. Wang, and Z. Zhang. Scale-aware trident networks for object detection. arXiv preprint arXiv:1901.01892, 2019.
  23. 23.T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  24. 24.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pages 2980–2988, 2017.
  25. 25.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  26. 26.S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia. Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8759–8768, 2018.
  27. 27.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg. Ssd: Single shot multibox detector. In European conference on computer vision, pages 21–37. Springer, 2016.
  28. 28.X. Lu, B. Li, Y. Yue, Q. Li, and J. Yan. Grid r-cnn. 2018.
  29. 29.A. Newell, K. Yang, and J. Deng. Stacked hourglass networks for human pose estimation. In European conference on computer vision, pages 483–499. Springer, 2016.
  30. 30.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. 2017.
  31. 31.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016.
  32. 32.J. Redmon and A. Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7263–7271, 2017.
  33. 33.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, pages 91–99, 2015.
  34. 34.Z. Shen, Z. Liu, J. Li, Y.-G. Jiang, Y. Chen, and X. Xue. Dsod: Learning deeply supervised object detectors from scratch. In Proceedings of the IEEE international conference on computer vision, pages 1919–1927, 2017.
  35. 35.Z. Shen, H. Shi, R. Feris, L. Cao, S. Yan, D. Liu, X. Wang, X. Xue, and T. S. Huang. Learning object detectors from scratch with gated recurrent feature pyramids. arXiv preprint arXiv:1712.00886, 2017.
  36. 36.A. Shrivastava and A. Gupta. Contextual priming and feedback for faster r-cnn. In European conference on computer vision, pages 330–348, 2016.
  37. 37.A. Shrivastava, R. Sukthankar, J. Malik, and A. Gupta. Beyond skip connections: Top-down modulation for object detection. arXiv preprint arXiv:1612.06851, 2016.
  38. 38.B. Singh and L. S. Davis. An analysis of scale invariance in object detection snip. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3578–3587, 2018.
  39. 39.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-First AAAI conference on artificial intelligence, 2017.
  40. 40.L. Tychsen-Smith and L. Petersson. Denet: Scalable real-time object detection with directed sparse sampling. In Proceedings of the IEEE international conference on computer vision, pages 428–436, 2017.
  41. 41.L. Tychsen-Smith and L. Petersson. Improving object localization with fitness nms and bounded iou loss. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6877–6885, 2018.
  42. 42.J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders. Selective search for object recognition. International journal of computer vision, 104(2):154–171, 2013.
  43. 43.H. Xu, X. Lv, X. Wang, Z. Ren, N. Bodla, and R. Chellappa. Deep regionlets for object detection. In Proceedings of the European conference on computer vision, pages 798–814, 2018.
  44. 44.X. Zeng, W. Ouyang, B. Yang, J. Yan, and X. Wang. Gated bi-directional cnn for object detection. In European conference on computer vision, pages 354–369. Springer, 2016.
  45. 45.S. Zhang, L. Wen, X. Bian, Z. Lei, and S. Z. Li. Single-shot refinement neural network for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4203–4212, 2018.
  46. 46.R. Zhu, S. Zhang, X. Wang, L. Wen, H. Shi, L. Bo, and T. Mei. Scratchdet: Training single-shot object detectors from scratch. Proceedings of the IEEE conference on computer vision and pattern recognition, 2019.
  47. 47.Y. Zhu, C. Zhao, J. Wang, X. Zhao, Y. Wu, and H. Lu. Couplenet: Coupling global structure with local parts for object detection. In Proceedings of the IEEE international conference on computer vision, pages 4126–4134, 2017.

Citation

MLA
Duan, K., et al. “CenterNet: Keypoint Triplets for Object Detection”. arXiv, 2019, http://arxiv.org/abs/1904.08189v3.
APA
Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., & Tian, Q. (2019). CenterNet: Keypoint Triplets for Object Detection. arXiv. http://arxiv.org/abs/1904.08189v3
Chicago
Duan, K., S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian. 2019. “CenterNet: Keypoint Triplets for Object Detection”. arXiv. http://arxiv.org/abs/1904.08189v3.
Harvard
Duan, K. et al. (2019) “CenterNet: Keypoint Triplets for Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.08189v3.
Vancouver
1. Duan K, Bai S, Xie L, Qi H, Huang Q, Tian Q (2019) CenterNet: Keypoint Triplets for Object Detection. arXiv

BibTeX

@article{duan2019centernet,
  title = {CenterNet: Keypoint Triplets for Object Detection},
  author = {Duan, Kaiwen and Bai, Song and Xie, Lingxi and Qi, Honggang and Huang, Qingming and Tian, Qi},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.08189v3},
  eprint = {1904.08189}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE