TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios
Xingkui ZhuShuchang LyuXu WangQi Zhao
Proposes TPH-YOLOv5, an aerial object detector that incorporates Transformer-based prediction heads and attention modules into YOLOv5 to effectively identify densely packed, scale-varying objects in drone imagery.
Unmanned aerial vehicles, commonly known as drones, are increasingly deployed in applications such as agricultural monitoring, wildlife preservation, and urban surveillance. However, standard computer vision algorithms struggle when applied to drone imagery. Drones operate across varying altitudes, causing target objects to fluctuate wildly in scale, while wide-area coverage and high-density environments introduce distracting background elements, motion blur, and significant object occlusion. This article sets out to design and demonstrate an improved object detection architecture tailored specifically to overcome these drone-specific visual challenges.
To address these limitations, the authors developed TPH-YOLOv5, an enhanced version of the YOLOv5 object detector. The framework incorporates four key modifications: adding a dedicated high-resolution prediction head to capture tiny objects, integrating Transformer-based prediction heads that use attention mechanisms to resolve dense and occluded objects, embedding a convolutional attention module to suppress confusing background terrain, and pairing the model with an auxiliary classification network to resolve visually similar categories. The system was trained and evaluated on the benchmark VisDrone2021 dataset, combining experimental ablation tests with multi-scale testing and model ensemble strategies.
The experimental findings show substantial improvements in visual recognition accuracy across drone scenarios. On the benchmark test challenge dataset, TPH-YOLOv5 achieved an average precision of 39.18%, surpassing the previous state-of-the-art detector by 1.81% and trailing the top-ranking challenge model by only 0.25%. Relative to the standard baseline detector, the proposed architecture improved overall precision by approximately 7%. Detailed testing confirmed that adding the fourth prediction head for tiny objects provided the largest single architectural gain, boosting average precision by 2.15%, while the transformer encoder blocks contributed an additional 1.81% improvement. Furthermore, introducing the secondary classifier successfully resolved confusion among ambiguous categories, delivering an extra 0.8% to 1.0% precision increase.
These results demonstrate that drone-based surveillance and inspection systems can achieve much higher reliability in complex real-world conditions without requiring fundamentally new architectures from scratch. By enhancing existing vision detectors with attention modules and targeted classification heads, organizations can improve automated tracking accuracy and reduce operational risk in critical monitoring tasks. While the extra detection head increases computational demand, the integration of transformer blocks partially offsets this by streamlining network layers, presenting a viable performance-to-compute balance for aerial operations.
For practical deployment, organizations should adopt multi-head attention enhancements and multi-model ensembling when maximum detection precision is needed. Future efforts should evaluate deployment constraints on edge devices with limited computational power and explore further optimizations to maintain real-time processing speeds. The findings provide high confidence regarding detection gains on aerial benchmark datasets, though practitioners should anticipate higher hardware and memory requirements when deploying high-resolution input pipelines.
- Paper: YOLOv4: Optimal Speed and Accuracy of Object Detection, Alexey Bochkovskiy et al. (2020). Establishes the modern single-stage YOLO architecture, multi-scale feature aggregation, and bag-of-freebies optimization strategies that underpin the YOLOv5 baseline.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Introduces the seminal formulation of self-attention transformers for visual object detection that motivates transformer prediction heads.
- Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Provides the foundational multi-scale pyramid hierarchy needed to understand multi-head prediction design for extreme scale variation.
- Paper: DOTA: A Large-Scale Dataset for Object Detection in Aerial Images, Gui-Song Xia et al. (2017). Surveys the unique challenges of aerial and drone visual detection, including tiny targets, dense packing, and wide altitude variations.
- Paper: YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications, Chuyi Li et al. (2022). Pushes the evolution of YOLO detectors forward by introducing decoupled prediction heads and hardware-aware re-parameterization.
- Paper: YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors, Chien-Yao Wang et al. (2023). Extends real-time multi-scale detection through compound scaling and dynamic auxiliary-to-lead head supervision.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). Advances the integration of transformers into real-time detection by designing an efficient hybrid attention encoder that eliminates non-maximum suppression.
- Paper: YOLOv12: Attention-Centric Real-Time Object Detectors, Yunjie Tian et al. (2025). Fully realizes attention-centric real-time object detection by replacing standard convolutions with area attention while preserving low latency.
- Paper: YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information, Chien-Yao Wang et al. (2024). Addresses information loss in deep multi-scale detector backbones using programmable gradient information.
