A Survey on Object Detection in Optical Remote Sensing Images
Gong ChengJunwei Han
Systematizes approximately 270 generic object detection studies across aerial and satellite imagery into four core methodological paradigms while reviewing standard benchmarks, evaluation metrics, and future opportunities in deep and weakly supervised learning.
Automatic object detection in optical remote sensing images is critical for practical applications such as urban planning, precision agriculture, environmental monitoring, disaster response, and updating geographic databases. With the rapid growth of high-resolution satellite and aerial imagery, manual identification has become impractical. At the same time, automated detection faces major hurdles due to complex backgrounds, clutter, shadows, varying viewpoints, and wide variations in object appearances.
The main objective of the article is to provide a comprehensive survey of generic object detection methods across diverse categories, rather than focusing on a single target type like roads or buildings. To accomplish this, the article evaluates approximately 270 research publications, systematically organizing the technical landscape into four primary methodological frameworks, reviewing five public benchmark datasets, and outlining standard evaluation metrics.
The review identifies key trade-offs across current approaches. Early template matching methods are simple but struggle with appearance variations, though deformable templates offer more flexibility at higher computational cost. Knowledge-based methods rely on geometric and contextual rules, which often lack robustness if rules are defined too strictly or loosely. Object-based image analysis successfully groups homogeneous pixels to classify land cover, but setting automated segmentation scales remains challenging. Machine learning approaches achieve higher detection accuracy by extracting features and training statistical classifiers; however, most deployed systems still depend heavily on handcrafted visual descriptors and extensive manual annotations.
These findings indicate that existing operational workflows face significant cost and scaling bottlenecks due to the labor-intensive requirement for detailed bounding-box labeling. Shifting toward modern feature representation and reduced human intervention is essential to handle large data streams efficiently and reduce the risk of detector failure across complex real-world environments.
The article highlights two actionable research directions to build more robust systems: adopting deep learning architectures to extract high-level feature representations directly from imagery, and developing weakly supervised learning frameworks that only require image-level presence labels rather than full bounding annotations. Organizations developing remote sensing pipelines should prioritize deep learning for improved detection power while investing in weakly supervised algorithms capable of detecting multiple object classes simultaneously.
Decision-makers must note key current limitations: deep neural networks require massive training datasets to prevent overfitting and carry high computational costs during real-time feature extraction. Additionally, weakly supervised methods in remote sensing are still in early development, with performance currently trailing supervised alternatives. Operational deployments should therefore balance deep learning accuracy against available computing resources and labeling budgets.
- Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Introduces Faster R-CNN and region proposal networks, establishing the foundational modern machine learning detection pipeline reviewed and adapted to remote sensing imagery in the survey.
- Paper: Rich feature hierarchies for accurate object detection and semantic segmentation, Ross Girshick et al. (2014). Pioneers the application of deep convolutional neural networks to region-based object detection, setting the baseline for deep learning methods analyzed in the survey.
- Paper: Object Detection with Discriminatively Trained Part-Based Models, Pedro F. Felzenszwalb et al. (2010). Establishes the deformable part-based model (DPM) framework using HOG and latent SVMs, providing essential context for the traditional machine learning approaches cataloged in the survey.
- Paper: The Pascal Visual Object Classes Challenge: A Retrospective, Mark Everingham et al. (2014). Details the standard evaluation protocols and benchmarks like PASCAL VOC that define the evaluation metrics adopted and discussed in the survey.
- Paper: OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks, Pierre Sermanet et al. (2014). Presents an early integrated framework for joint localization and detection via sliding-window convolutional networks, bridging traditional scanning approaches with deep learning.
- Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). Develops spatial pyramid matching to capture spatial structure from bag-of-visual-words, a key classical feature representation reviewed in the survey.
- Paper: Visual categorization with bags of keypoints, Gabriella Csurka et al. (2004). Introduces the bag-of-keypoints visual representation for categorization, serving as a core foundation for early machine learning and template matching detectors.
- Paper: A Trainable System for Object Detection, CONSTANTINE PAPAGEORGIOU et al. (2000). Demonstrates early trainable sliding-window detection using multiscale wavelets and support vector machines, underpinning the machine learning category of object detection methods.
- Paper: Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark, Ke Li et al. (2019). Directly succeeds the 2016 survey by introducing the large-scale DIOR benchmark and evaluating twelve modern deep learning detectors across 20 remote sensing categories.
- Paper: DOTA: A Large-Scale Dataset for Object Detection in Aerial Images, Gui-Song Xia et al. (2017). Addresses the survey's call for large-scale benchmarks by introducing DOTA and standardizing oriented bounding box detection in aerial imagery.
- Paper: Deep learning in remote sensing: a review, Xiao Xiang Zhu et al. (2017). Broadens the survey's discussion of deep learning in optical imagery to multi-sensor Earth observation modalities including SAR and hyperspectral analysis.
- Paper: Deep Learning for Generic Object Detection: A Survey, Li Liu et al. (2018). Provides a comprehensive follow-up survey tracing the subsequent rapid evolution of deep learning detector architectures and multi-scale feature hierarchies.
- Paper: TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios, Xingkui Zhu et al. (2021). Applies transformer-based prediction heads within YOLOv5 to address specific remote sensing challenges highlighted in the survey, such as tiny and occluded targets.
- Paper: Remote Sensing Image Scene Classification: Benchmark and State of the Art, Gong Cheng et al. (2017). Complements optical object detection by establishing NWPU-RESISC45, a large-scale benchmark for aerial scene classification evaluating deep feature learning.
- Paper: AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification, Gui-Song Xia et al. (2016). Supplies a large-scale aerial dataset (AID) to systematically evaluate deep learning baselines on overhead imagery.
- Paper: Object Detection in 20 Years: A Survey, Zhengxia Zou et al. (2019). Provides a broad retrospective survey tracking the progression from early handcrafted methods to modern deep learning and vision transformer architectures.
