TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios

Xingkui ZhuShuchang LyuXu WangQi Zhao

article2021ICCV1,800 citationsVisDrone Challenge 2021 4th place

Proposes TPH-YOLOv5, an aerial object detector that incorporates Transformer-based prediction heads and attention modules into YOLOv5 to effectively identify densely packed, scale-varying objects in drone imagery.

Listen

Unmanned aerial vehicles, commonly known as drones, are increasingly deployed in applications such as agricultural monitoring, wildlife preservation, and urban surveillance. However, standard computer vision algorithms struggle when applied to drone imagery. Drones operate across varying altitudes, causing target objects to fluctuate wildly in scale, while wide-area coverage and high-density environments introduce distracting background elements, motion blur, and significant object occlusion. This article sets out to design and demonstrate an improved object detection architecture tailored specifically to overcome these drone-specific visual challenges.

To address these limitations, the authors developed TPH-YOLOv5, an enhanced version of the YOLOv5 object detector. The framework incorporates four key modifications: adding a dedicated high-resolution prediction head to capture tiny objects, integrating Transformer-based prediction heads that use attention mechanisms to resolve dense and occluded objects, embedding a convolutional attention module to suppress confusing background terrain, and pairing the model with an auxiliary classification network to resolve visually similar categories. The system was trained and evaluated on the benchmark VisDrone2021 dataset, combining experimental ablation tests with multi-scale testing and model ensemble strategies.

The experimental findings show substantial improvements in visual recognition accuracy across drone scenarios. On the benchmark test challenge dataset, TPH-YOLOv5 achieved an average precision of 39.18%, surpassing the previous state-of-the-art detector by 1.81% and trailing the top-ranking challenge model by only 0.25%. Relative to the standard baseline detector, the proposed architecture improved overall precision by approximately 7%. Detailed testing confirmed that adding the fourth prediction head for tiny objects provided the largest single architectural gain, boosting average precision by 2.15%, while the transformer encoder blocks contributed an additional 1.81% improvement. Furthermore, introducing the secondary classifier successfully resolved confusion among ambiguous categories, delivering an extra 0.8% to 1.0% precision increase.

These results demonstrate that drone-based surveillance and inspection systems can achieve much higher reliability in complex real-world conditions without requiring fundamentally new architectures from scratch. By enhancing existing vision detectors with attention modules and targeted classification heads, organizations can improve automated tracking accuracy and reduce operational risk in critical monitoring tasks. While the extra detection head increases computational demand, the integration of transformer blocks partially offsets this by streamlining network layers, presenting a viable performance-to-compute balance for aerial operations.

For practical deployment, organizations should adopt multi-head attention enhancements and multi-model ensembling when maximum detection precision is needed. Future efforts should evaluate deployment constraints on edge devices with limited computational power and explore further optimizations to maintain real-time processing speeds. The findings provide high confidence regarding detection gains on aerial benchmark datasets, though practitioners should anticipate higher hardware and memory requirements when deploying high-resolution input pipelines.

Cover for TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios

Abstract

Object detection on drone-captured scenarios is a recent popular task. As drones always navigate in different altitudes, the object scale varies violently, which burdens the optimization of networks. Moreover, high-speed and low-altitude flight bring in the motion blur on the densely packed objects, which leads to great challenge of object distinction. To solve the two issues mentioned above, we propose TPH-YOLOv5. Based on YOLOv5, we add one more prediction head to detect different-scale objects. Then we replace the original prediction heads with Transformer Prediction Heads (TPH) to explore the prediction potential with self-attention mechanism. We also integrate convolutional block attention model (CBAM) to find attention region on scenarios with dense objects. To achieve more improvement of our proposed TPH-YOLOv5, we provide bags of useful strategies such as data augmentation, multiscale testing, multi-model integration and utilizing extra classifier. Extensive experiments on dataset VisDrone2021 show that TPH-YOLOv5 have good performance with impressive interpretability on drone-captured scenarios. On DET-test-challenge dataset, the AP result of TPH-YOLOv5 are 39.18%, which is better than previous SOTA method (DPNetV3) by 1.81%. On VisDrone Challenge 2021, TPHYOLOv5 wins 5th place and achieves well-matched results with 1st place model (AP 39.43%). Compared to baseline model (YOLOv5), TPH-YOLOv5 improves about 7%, which is encouraging and competitive.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Data Augmentation
  • 2.2 Multi-Model Ensemble Method in Object Detection
  • 2.3 Object Detection
  • 3 TPH-YOLOv5
  • 3.1 Overview of YOLOv5
  • 3.2 TPH-YOLOv5
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Comparisons with the State-of-the-art
  • 4.3 Ablation Studies
  • 5 Conclusion
  • 6 Acknowledgments
  • References

Knowls

  1. Knowl 1 — TPH-YOLOv5 Architecture for Aerial Object Detection

    model/method

    TPH-YOLOv5 is a one-stage object detection architecture modified from YOLOv5x to address the domain-specific challenges of drone-captured imagery: extreme scale variance, dense object packing with occlusion, and large image coverage containing confusing geographic background elements.

    The core architectural modifications comprise:

    1. Four-Head Detection Structure: An extra high-resolution prediction head (P2P_2) is added alongside standard small (P3P_3), medium (P4P_4), and large (P5P_5) heads to detect tiny objects.
    2. Transformer Prediction Heads (TPH): Transformer encoder blocks replace standard convolutional and CSP bottleneck blocks at the deep end of the CSPDarknet53 backbone and within the Path Aggregation Network (PANet) neck to model global contextual dependencies via self-attention.
    3. Convolutional Block Attention Module (CBAM): CBAM modules are incorporated sequentially along channel and spatial dimensions into the neck to suppress irrelevant background clutter and focus on regions of interest.
  2. Knowl 2 — Transformer Prediction Head and Encoder Block Integration

    model/method

    The Transformer Prediction Head (TPH) integrates vision transformer encoder blocks into the YOLOv5 detection head and the low-resolution end of the CSPDarknet53 backbone. Each transformer encoder block contains two sub-layers: a multi-head self-attention (MHSA) module and a feed-forward multi-layer perceptron (MLP) module, augmented with LayerNorm, Dropout, and residual connections across each sub-layer.

    Applying self-attention captures long-range dependencies and global semantics, which enhances the model's ability to localize and discriminate densely packed or occluded objects in aerial scenes. Restricting transformer encoder blocks to low-resolution feature maps mitigates the computational and memory overhead inherent to self-attention. Replacing convolutional blocks with transformer encoder blocks reduces total model layers from 719 to 705 and computation from 259.0 GFLOPs to 237.3 GFLOPs in the four-head YOLOv5x configuration while improving mean Average Precision (mAP).

  3. Knowl 3 — High-Resolution P2 Detection Head for Tiny Objects

    model/method

    Standard YOLOv5 architectures utilize three prediction heads downsampled by factors of 8×8\times, 16×16\times, and 32×32\times (P3,P4,P5P_3, P_4, P_5). Because drone-captured images are dominated by extremely small instances due to high flight altitudes, TPH-YOLOv5 incorporates a fourth prediction head (P2P_2) generated from a low-level, high-resolution feature map (downsampled by 4×4\times).

    This high-resolution head enhances the network's spatial sensitivity to tiny objects. Incorporating the P2P_2 head increases the model layer count from 607 to 719 and computational complexity from 219.0 GFLOPs to 259.0 GFLOPs on YOLOv5x, while yielding a +2.15% gain in mAP on aerial object detection.

  4. Knowl 4 — Component-Wise Ablation Study of TPH-YOLOv5

    data/table

    The ablation study evaluates the incremental contribution of each component on the VisDrone2021-DET test-dev dataset. Baseline YOLOv5 corresponds to YOLOv5x trained with MixUp and Mosaic augmentations.

    Method mAP (%) AP50\text{AP}_{50} (%)
    YOLOv5 28.88 49.33
    YOLOv5 + P_2 31.03 (↑2.15\uparrow 2.15) 51.61 (↑2.28\uparrow 2.28)
    YOLOv5 + P_2 + transformer 32.84 (↑1.81\uparrow 1.81) 53.87 (↑2.26\uparrow 2.26)
    TPH-YOLOv5 (previous + CBAM) 33.63 (↑0.79\uparrow 0.79) 54.77 (↑0.90\uparrow 0.90)
    TPH-YOLOv5 + ms-testing 34.90 (↑1.27\uparrow 1.27) 56.40 (↑1.63\uparrow 1.63)
    TPH-YOLOv5 + ms-testing + Classifier 35.74 (↑0.84\uparrow 0.84) 57.31 (↑0.91\uparrow 0.91)

    The evaluation metrics are mAP\text{mAP} (average AP across 10 IoU thresholds from 0.50 to 0.95 with step 0.05) and AP50\text{AP}_{50} (AP at IoU=0.50\text{IoU} = 0.50). The additions of the P2P_2 tiny-object head (+2.15% mAP) and the transformer encoder blocks (+1.81% mAP) provide the largest individual performance increases.

  5. Knowl 5 — Self-Trained Cropped-Patch Classifier for Fine-Grained Aerial Categories

    model/method

    Analysis of the TPH-YOLOv5 confusion matrix indicates that while spatial localization is accurate, the detector exhibits classification confusion among visually similar drone-captured categories (such as tricycle versus awning-tricycle).

    To correct classification errors, an auxiliary ResNet-18 classifier is trained separately on image patches cropped from ground-truth bounding boxes in the training set, resized to 64×6464 \times 64 pixels. During inference, predicted bounding boxes generated by TPH-YOLOv5 are passed to this classifier to update category labels, providing an improvement of 0.8%∼1.0%0.8\%\sim1.0\% in mAP.

  6. Knowl 6 — Benchmark Comparison on VisDrone2021-DET Test-Challenge Dataset

    data/table

    The detection performance of TPH-YOLOv5 ensemble was evaluated on the VisDrone2021-DET test-challenge dataset and compared against previous state-of-the-art methods and challenge submissions.

    Methods mAP (%) AP50\text{AP}_{50} (%)
    RetinaNet 11.81 21.37
    RefineDet 14.90 28.76
    DetNet59 15.26 29.23
    Cascade-RCNN 16.09 31.91
    FPN 16.51 32.20
    Light-RCNN 16.53 32.78
    CornerNet 17.41 34.12
    RRNet (2019 2nd2^{\text{nd}}) 29.13 55.82
    DPNet-ensemble (2019 SOTA) 29.62 54.00
    SMPNet (2020 2nd2^{\text{nd}}) 35.98 59.53
    DPNetV3 (2020 SOTA) 37.37 62.05
    TPH-YOLOv5 ensemble 39.18 –

    TPH-YOLOv5 ensemble achieves 39.18% mAP, outperforming the previous state-of-the-art model (DPNetV3, 37.37% mAP) by 1.81% and placing 5th in the VisDrone 2021 challenge (first place was 39.43% mAP).

  7. Knowl 7 — Multi-Scale Testing and Weighted Boxes Fusion Ensemble

    model/method

    TPH-YOLOv5 improves inference performance via multi-scale testing (ms-testing) and multi-model ensembling using Weighted Boxes Fusion (WBF):

    1. Multi-Scale Testing Pipeline: For a single model, each test image is resized to three scales (1.3×1.3\times, 1.0×1.0\times, and 0.83×0.83\times or 0.67×0.67\times) and horizontally flipped, yielding 6 test image variants. Predictions across all 6 variants are merged using Non-Maximum Suppression (NMS).
    2. Multi-Model Fusion via WBF: Five TPH-YOLOv5 model variants trained with diverse configurations (input sizes of 1536 or 1920, uniform category weights versus inverse-frequency loss weights, and YOLOv5x versus YOLOv5l backbones) are individually evaluated with ms-testing. Their bounding boxes are subsequently ensembled using Weighted Boxes Fusion (WBF), which computes confidence-weighted coordinate averages of overlapping boxes rather than discarding non-maximum boxes.
  8. Knowl 8 — Per-Category Evaluation of TPH-YOLOv5 Model Variants and Ensemble

    data/table

    Performance (mAP in %) across all 10 object classes on the VisDrone2021-DET test-dev dataset for five individual TPH-YOLOv5 configurations and their final ensemble.

    Method all pedestrian people bicycle car van truck tricycle awning-tricycle bus motor
    TPH-YOLOv5-1 34.90 27.52 15.32 15.21 65.99 44.23 47.56 23.96 22.11 58.85 28.44
    TPH-YOLOv5-2 34.29 27.97 14.88 14.17 67.63 45.01 44.76 25.12 20.48 55.72 27.74
    TPH-YOLOv5-3 34.68 22.88 16.01 19.26 48.88 42.98 47.82 32.86 35.65 54.16 28.25
    TPH-YOLOv5-4 34.17 23.48 15.79 17.62 49.99 42.76 47.13 31.66 32.21 54.19 27.37
    TPH-YOLOv5-5 33.04 25.98 14.90 13.10 63.05 43.45 42.56 25.20 21.06 53.65 27.10
    TPH-YOLOv5 ensemble 37.32 29.00 16.75 15.69 68.94 49.79 45.16 27.33 24.72 61.80 30.90

    Model specifications:

    • TPH-YOLOv5-1: Input image size 1920, equal category loss weighting.
    • TPH-YOLOv5-2: Input image size 1536, equal category loss weighting.
    • TPH-YOLOv5-3: Input image size 1920, category loss weights inversely proportional to class label count.
    • TPH-YOLOv5-4: Input image size 1536, category loss weights inversely proportional to class label count.
    • TPH-YOLOv5-5: YOLOv5l backbone, input image size 1536.
    • TPH-YOLOv5 ensemble: WBF fusion of all five models.
  9. Knowl 9 — Sub-3-Pixel Object Filtering via Gray Masking

    model/method

    In the VisDrone2021 dataset at an input image size of 1536 pixels, 622 out of 342,391 bounding box labels have a side length of less than 3 pixels. Because instances of this scale lack recognizable visual features and introduce optimization noise during backpropagation, these bounding box regions are covered with gray square masks during training, which yields an improvement of +0.2% mAP.

  10. Knowl 10 — Implementation and Training Configuration for TPH-YOLOv5

    experimental setup

    TPH-YOLOv5 is implemented in PyTorch 1.8.1 and trained on an NVIDIA RTX 3090 GPU with the following experimental parameters:

    • Weight Transfer: Backbone blocks (0–8) and head layers (10–13, 15–18) are initialized using pretrained YOLOv5x weights.
    • Training Schedule: 65 total epochs on the VisDrone2021 training set, comprising 2 warm-up epochs.
    • Optimization: Adam optimizer with initial learning rate η0=3×10−4\eta_0 = 3 \times 10^{-4} and a cosine decay schedule down to ηfinal=0.12×η0\eta_{\text{final}} = 0.12 \times \eta_0.
    • Batch Size & Input Scale: Long image side scaled to 1536 pixels with a batch size of 2.
    • Data Augmentation: MixUp, Mosaic, photometric distortions (hue, saturation, value), and geometric distortions (random scaling, cropping, translation, shearing, rotation).

Coverage note — None was omitted; all key architectural components, training strategies, ablation studies, and benchmark results are represented.

References

  1. 1.Nicolas Audebert, Bertrand Le Saux, and Sebastien Lef evre. Beyond rgb: Very high resolution urban remote sensing with multimodal deep networks. ISPRS Journal of Photogrammetry and Remote Sensing, 140:20–32, 2018.
  2. 2.Alexey Bochkovskiy, Chien-Yao Wang, and HongYuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
  3. 3.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018.
  4. 4.Changrui Chen, Yu Zhang, Qingxuan Lv, Shuo Wei, Xiaorui Wang, Xin Sun, and Junyu Dong. Rrnet: A hybrid detector for object detection in drone-captured images. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019.
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017.
  6. 6.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 2021.
  8. 8.Dawei Du, Pengfei Zhu, Longyin Wen, Xiao Bian, Haibin Lin, Qinghua Hu, Tao Peng, Jiayu Zheng, Xinyao Wang, Yue Zhang, et al. Visdrone-det2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019.
  9. 9.Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman. The pascal visual object classes (VOC) challenge. Int. J. Comput. Vis., 88(2):303–338, 2010.
  10. 10.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  11. 11.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021.
  12. 12.Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. Nas-fpn: Learning scalable feature pyramid architecture for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7036– 7045, 2019.
  13. 13.Ross Girshick. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 1440–1448, 2015.
  14. 14.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
  15. 15.Jingjing Gu, Tao Su, Qiuhong Wang, Xiaojiang Du, and Mohsen Guizani. Multiple moving targets surveillance based on a cooperative network for multi-uav. IEEE Commun. Mag., 56(4):82–89, 2018.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 37(9):1904–1916, 2015.
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  18. 18.Jennifer N. Hird, Alessandro Montaghi, Gregory J. McDermid, Jahan Kariyeva, Brian J. Moorman, Scott E. Nielsen, and Anne C. S. McIntosh. Use of unmanned aerial vehicles for monitoring recovery of forest vegetation on petroleum well sites. Remote. Sens., 9(5):413, 2017.
  19. 19.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  20. 20.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  21. 21.Glenn Jocher, Alex Stoken, Jirka Borovec, NanoCode012, Ayush Chaurasia, TaoXie, Liu Changyu, Abhiram V, Laughing, tkianai, yxNONG, Adam Hogan, lorenzomammana, AlexWang1900, Jan Hajek, Laurentiu Diaconu, Marc, Yonghye Kwon, oleg, wanghaoyang0106, Yann Defretin, Aditya Lohia, ml5ah, Ben Milanko, Benjamin Fineran, Daniel Khromov, Ding Yiwei, Doug, Durgesh, and Francisco Ingham. ultralytics/yolov5: v5.0 - YOLOv5-P6 1280 models, AWS, Supervise.ly and YouTube integrations, Apr. 2021.
  22. 22.Benjamin Kellenberger, Diego Marcos, and Devis Tuia. Detecting mammals in uav images: Best practices to address a substantially imbalanced dataset with deep learning. Remote Sensing of Environment, 216:139–153, 2018.
  23. 23.Benjamin Kellenberger, Michele Volpi, and Devis Tuia. Fast animal detection in UAV images using convolutional neural networks. In 2017 IEEE International Geoscience and Remote Sensing Symposium, IGARSS 2017, Fort Worth, TX, USA, July 23-28, 2017, pages 866–869. IEEE, 2017.
  24. 24.Hei Law and Jia Deng. Cornernet: Detecting objects as paired keypoints. In Proceedings of the European conference on computer vision (ECCV), pages 734–750, 2018.
  25. 25.Zeming Li, Chao Peng, Gang Yu, Xiangyu Zhang, Yangdong Deng, and Jian Sun. Light-head r-cnn: In defense of two-stage object detector. arXiv preprint arXiv:1711.07264, 2017.
  26. 26.Zeming Li, Chao Peng, Gang Yu, Xiangyu Zhang, Yangdong Deng, and Jian Sun. Detnet: A backbone network for object detection. arXiv preprint arXiv:1804.06215, 2018.
  27. 27.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ´ IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 2999–3007. IEEE Computer Society, 2017.
  28. 28.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, ´ Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  29. 29.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ´ Proceedings of the IEEE international conference on computer vision, pages 2980–2988, 2017.
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  31. 31.Songtao Liu, Di Huang, et al. Receptive field block net for accurate and fast object detection. In Proceedings of the European Conference on Computer Vision (ECCV), pages 385– 400, 2018.
  32. 32.Songtao Liu, Di Huang, and Yunhong Wang. Learning spatial fusion for single-shot object detection. arXiv preprint arXiv:1911.09516, 2019.
  33. 33.Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8759–8768, 2018.
  34. 34.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European conference on computer vision, pages 21–37. Springer, 2016.
  35. 35.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  36. 36.Alexander Neubeck and Luc Van Gool. Efficient nonmaximum suppression. In 18th International Conference on Pattern Recognition (ICPR’06), volume 3, pages 850–855. IEEE, 2006.
  37. 37.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016.
  38. 38.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7263–7271, 2017.
  39. 39.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
  40. 40.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28:91–99, 2015.
  41. 41.Zhenfeng Shao, Congmin Li, Deren Li, Orhan Altan, Lei Zhang, and Lin Ding. An accurate matching method for projecting vector data into surveillance video to monitor and protect cultivated land. ISPRS Int. J. Geo Inf., 9(7):448, 2020.
  42. 42.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  43. 43.Roman Solovyev, Weimin Wang, and Tatiana Gabruseva. Weighted boxes fusion: Ensembling boxes from different object detection models. Image and Vision Computing, 107:104117, 2021.
  44. 44.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning, pages 6105–6114. PMLR, 2019.
  45. 45.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10781–10790, 2020.
  46. 46.Ziyang Tang, Xiang Liu, Guangyu Shen, and Baijian Yang. Penet: object detection using points estimation in aerial images. arXiv preprint arXiv:2001.08247, 2020.
  47. 47.Visdrone Team. Visdrone 2020 leaderboard. Website, 2020. http://aiskyeye.com/ visdrone-2020-leaderboard/.
  48. 48.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9627–9636, 2019.
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.
  50. 50.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  51. 51.Chien-Yao Wang, Alexey Bochkovskiy, and HongYuan Mark Liao. Scaled-yolov4: Scaling cross stage partial network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13029–13038, 2021.
  52. 52.Chien-Yao Wang, Hong-Yuan Mark Liao, Yueh-Hua Wu, Ping-Yang Chen, Jun-Wei Hsieh, and I-Hau Yeh. Cspnet: A new backbone that can enhance learning capability of cnn. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 390–391, 2020.
  53. 53.Derui Wang, Chaoran Li, Sheng Wen, Qing-Long Han, Surya Nepal, Xiangyu Zhang, and Yang Xiang. Daedalus: Breaking nonmaximum suppression in object detection via adversarial examples. IEEE Transactions on Cybernetics, 2021.
  54. 54.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), pages 3–19, 2018.
  55. 55.Ze Yang, Shaohui Liu, Han Hu, Liwei Wang, and Stephen Lin. Reppoints: Point set representation for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9657–9666, 2019.
  56. 56.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6023–6032, 2019.
  57. 57.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  58. 58.Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Sunderhauf. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8514–8523, 2021.
  59. 59.Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Sunderhauf. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8514–8523, 2021.
  60. 60.Shifeng Zhang, Longyin Wen, Xiao Bian, Zhen Lei, and Stan Z Li. Single-shot refinement neural network for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4203–4212, 2018.
  61. 61.Qijie Zhao, Tao Sheng, Yongtao Wang, Zhi Tang, Ying Chen, Ling Cai, and Haibin Ling. M2det: A single-shot object detector based on multi-level feature pyramid network. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 9259–9266, 2019.
  62. 62.Xingyi Zhou, Vladlen Koltun, and Philipp Krahenb  uhl.  Probabilistic two-stage detection. arXiv preprint arXiv:2103.07461, 2021.
  63. 63.Xingyi Zhou, Dequan Wang, and Philipp Krahenb  uhl. Ob- jects as points. arXiv preprint arXiv:1904.07850, 2019.
  64. 64.Pengfei Zhu, Longyin Wen, Xiao Bian, Haibin Ling, and Qinghua Hu. Vision meets drones: A challenge. arXiv preprint arXiv:1804.07437, 2018.
  65. 65.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.

Citation

MLA
Zhu, X., et al. “TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios”. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2021, pp. 2778–88, https://doi.org/10.1109/ICCVW54120.2021.00312.
APA
Zhu, X., Lyu, S., Wang, X., & Zhao, Q. (2021). TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2778–2788. https://doi.org/10.1109/ICCVW54120.2021.00312
Chicago
Zhu, X., S. Lyu, X. Wang, and Q. Zhao. 2021. “TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios”. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), 2778–88. https://doi.org/10.1109/ICCVW54120.2021.00312.
Harvard
Zhu, X. et al. (2021) “TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios”, 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). IEEE, pp. 2778–2788. Available at: https://doi.org/10.1109/ICCVW54120.2021.00312.
Vancouver
1. Zhu X, Lyu S, Wang X, Zhao Q (2021) TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. In: 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). IEEE, pp 2778–2788

BibTeX

@inproceedings{Zhu_2021, title={TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios}, url={http://dx.doi.org/10.1109/ICCVW54120.2021.00312}, DOI={10.1109/iccvw54120.2021.00312}, booktitle={2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)}, publisher={IEEE}, author={Zhu, Xingkui and Lyu, Shuchang and Wang, Xu and Zhao, Qi}, year={2021}, month=Oct, pages={2778–2788} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE