Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
Ke LiGang WanGong ChengLiqiu MengJunwei Han
Presents a comprehensive survey of deep-learning-based aerial object detection alongside DIOR, a large-scale benchmark of over 190,000 instances across 20 categories, providing standardized baselines to advance remote sensing research.
Rapid advances in Earth observation technologies have produced an unprecedented volume of satellite and aerial imagery, driving the need for automated object detection in domains such as urban planning, infrastructure monitoring, and precision agriculture. However, transferring standard computer vision algorithms to remote sensing is difficult because overhead imagery primarily captures the top-down views of objects across varied angles, diverse spatial resolutions, and cluttered environments. Existing Earth observation datasets have suffered from limited image counts, few object classes, and minimal environmental diversity, creating a substantial bottleneck for developing reliable artificial intelligence models.
The article addresses this gap by establishing a new large-scale remote sensing benchmark named DIOR and conducting a comprehensive performance evaluation of twelve prominent deep learning object detection algorithms.
To create the benchmark, researchers collected 23,463 optical satellite images covering more than 80 countries, manually annotating 192,472 object instances across 20 distinct categories with horizontal bounding boxes. The dataset incorporates wide variations in weather, lighting, season, and image resolution, while balancing small and large objects. The evaluation compared twelve leading deep learning models divided between two-stage region-proposal methods and one-stage regression methods, using standard metrics for detection accuracy.
The benchmark revealed several critical findings. Overall detection accuracy across all models peaked at 66.1 percent mean average precision, demonstrated by RetinaNet and the Path Aggregation Network (PANet). Network depth and multi-scale feature hierarchies proved essential, with deeper backbones and feature pyramid structures consistently outperforming shallower architectures. For small targets such as vehicles, storage tanks, and ships, YOLOv3 achieved the highest class-specific accuracy, reaching 87.4 percent on ships. In addition, CornerNet, which identifies objects via bounding box corner pairs rather than pre-defined anchor boxes, achieved the top individual accuracy in 9 of the 20 categories. Despite these achievements, detection accuracy remained consistently low across complex and elongated categories such as bridges, harbors, and overpasses.
These findings indicate that general computer vision architectures cannot be deployed directly into operational remote sensing workflows without adaptation. The performance drop on elongated infrastructure and visually ambiguous objects presents operational risks if deployed in automated decision-making pipelines. The results demonstrate that handling rotational variation, scale extremes, and background clutter is critical for achieving production-grade accuracy.
Organizations developing or deploying automated Earth observation systems should adopt multi-scale feature architectures and prioritize rotation-insensitive model training. For operational pipelines requiring small-object identification, single-stage multi-scale detectors or corner-based methods should be prioritized. Further research should focus on advanced multi-scale training strategies, such as scale normalization, to bridge the performance gap on complex structures before automated pipelines are fully trusted for critical infrastructure monitoring.
While the DIOR benchmark provides a robust foundation, users should note that instances are labeled using axis-aligned horizontal bounding boxes rather than oriented bounding boxes, which can introduce background noise for diagonally oriented objects. High confidence can be placed in the comparative rankings of the evaluated algorithms, but caution is warranted when deploying these models to detect complex structural classes under poor image conditions.
- Paper: A Survey on Object Detection in Optical Remote Sensing Images, Gong Cheng et al. (2016). This earlier survey maps optical remote-sensing detection methods and benchmarks, giving essential context for the problems DIOR addresses.
- Paper: DOTA: A Large-Scale Dataset for Object Detection in Aerial Images, Gui-Song Xia et al. (2017). DOTA established a major aerial-detection benchmark and evaluated leading detectors, providing a direct point of reference for DIOR’s dataset and comparisons.
- Paper: Deep learning in remote sensing: a review, Xiao Xiang Zhu et al. (2017). This review surveys deep-learning architectures and benchmarks in remote sensing, preparing readers for the detector families and multi-scale findings assessed in the source.
- Paper: Deep Learning for Generic Object Detection: A Survey, Li Liu et al. (2018). Its account of two-stage, one-stage, and multi-scale detection methods clarifies the generic architectures that the source evaluates on overhead imagery.
- Paper: SOOD: Towards Semi-Supervised Oriented Object Detection, Wei Hua et al. (2023). Building on the source’s axis-aligned-box limitation, SOOD develops rotation-aware semi-supervised detection for aerial objects with fewer annotations.
- Paper: Shape-Adaptive Selection and Measurement for Oriented Object Detection, Liping Hou et al. (2022). This work extends the source’s concerns about elongated, difficult targets with shape-adaptive training and localization strategies for oriented detection.
- Paper: Spatial Transform Decoupling for Oriented Object Detection, Hongtian Yu et al. (2024). Taking on the rotation sensitivity highlighted by the source, this work decouples position, angle, and shape prediction for oriented remote-sensing objects.
