CenterNet: Keypoint Triplets for Object Detection
Kaiwen DuanSong BaiLingxi XieHonggang QiQingming HuangQi Tian
Presents CenterNet, an object detection framework that models objects as keypoint triplets using center and cascade corner pooling to filter out incorrect bounding boxes and achieve state-of-the-art one-stage detection accuracy on MS-COCO.
CenterNet addresses a persistent weakness in keypoint-based object detection: methods such as CornerNet generate many bounding boxes that do not align well with actual objects because they rely solely on paired corner keypoints and lack any check of the interior region. This produces high rates of false detections, especially for small objects, and limits overall precision even when recall is reasonable. The work therefore sets out to add a lightweight internal verification step that keeps the speed advantage of one-stage detectors while improving accuracy.
The authors extend CornerNet by representing each object as a triplet of keypoints—two corners plus one center keypoint—rather than a corner pair alone. After candidate boxes are formed from corners, a scale-aware central region is examined; if a center keypoint of matching class appears inside it, the box is retained and its score is adjusted by averaging the three keypoints. Two new pooling modules support this design: center pooling gathers stronger internal signals for the center keypoint, and cascade corner pooling lets corners capture both boundary and interior evidence. The resulting network is trained from scratch on the MS-COCO trainval35k set and evaluated on the test-dev set using 52-layer and 104-layer hourglass backbones.
On single-scale testing with the deeper backbone, CenterNet reaches 44.9 percent average precision, a 4.4-point gain over the CornerNet baseline; multi-scale testing raises the figure to 47.0 percent, surpassing all published one-stage detectors by at least 4.9 points and matching or approaching the best two-stage systems. Gains are largest on small objects (up to 8.1 points), and false-discovery rates drop noticeably at every IoU threshold. Inference requires 270–340 ms per image on a P100 GPU, remaining faster than most two-stage alternatives.
These results indicate that a modest central-region check can close much of the accuracy gap between one-stage and two-stage detectors without incurring the cost of region-of-interest pooling. The approach therefore offers practitioners a practical route to higher precision in real-time or resource-constrained settings. The authors note that the same center-keypoint branch could be added to other one-stage frameworks and that further gains are likely with improved center-keypoint training. Performance remains sensitive to center-keypoint accuracy, however; substituting ground-truth centers lifts AP by roughly 14 points, showing that center detection is still a remaining bottleneck.
- Paper: CornerNet: Detecting Objects as Paired Keypoints, Hei Law et al. (2018). CenterNet directly extends CornerNet’s paired-corner detector by adding a center keypoint to verify candidate boxes, so CornerNet’s representation and limitations clarify the design.
- Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). CenterNet uses focal loss in its keypoint-based detector, and understanding how focal loss addresses dense foreground–background imbalance helps explain its training setup.
No sufficiently relevant recommendations were found.
