Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark

Ke LiGang WanGong ChengLiqiu MengJunwei Han

article2019Isprs Journal of Photogrammetry and Remote Sensing2,273 citations

Presents a comprehensive survey of deep-learning-based aerial object detection alongside DIOR, a large-scale benchmark of over 190,000 instances across 20 categories, providing standardized baselines to advance remote sensing research.

Listen

Rapid advances in Earth observation technologies have produced an unprecedented volume of satellite and aerial imagery, driving the need for automated object detection in domains such as urban planning, infrastructure monitoring, and precision agriculture. However, transferring standard computer vision algorithms to remote sensing is difficult because overhead imagery primarily captures the top-down views of objects across varied angles, diverse spatial resolutions, and cluttered environments. Existing Earth observation datasets have suffered from limited image counts, few object classes, and minimal environmental diversity, creating a substantial bottleneck for developing reliable artificial intelligence models.

The article addresses this gap by establishing a new large-scale remote sensing benchmark named DIOR and conducting a comprehensive performance evaluation of twelve prominent deep learning object detection algorithms.

To create the benchmark, researchers collected 23,463 optical satellite images covering more than 80 countries, manually annotating 192,472 object instances across 20 distinct categories with horizontal bounding boxes. The dataset incorporates wide variations in weather, lighting, season, and image resolution, while balancing small and large objects. The evaluation compared twelve leading deep learning models divided between two-stage region-proposal methods and one-stage regression methods, using standard metrics for detection accuracy.

The benchmark revealed several critical findings. Overall detection accuracy across all models peaked at 66.1 percent mean average precision, demonstrated by RetinaNet and the Path Aggregation Network (PANet). Network depth and multi-scale feature hierarchies proved essential, with deeper backbones and feature pyramid structures consistently outperforming shallower architectures. For small targets such as vehicles, storage tanks, and ships, YOLOv3 achieved the highest class-specific accuracy, reaching 87.4 percent on ships. In addition, CornerNet, which identifies objects via bounding box corner pairs rather than pre-defined anchor boxes, achieved the top individual accuracy in 9 of the 20 categories. Despite these achievements, detection accuracy remained consistently low across complex and elongated categories such as bridges, harbors, and overpasses.

These findings indicate that general computer vision architectures cannot be deployed directly into operational remote sensing workflows without adaptation. The performance drop on elongated infrastructure and visually ambiguous objects presents operational risks if deployed in automated decision-making pipelines. The results demonstrate that handling rotational variation, scale extremes, and background clutter is critical for achieving production-grade accuracy.

Organizations developing or deploying automated Earth observation systems should adopt multi-scale feature architectures and prioritize rotation-insensitive model training. For operational pipelines requiring small-object identification, single-stage multi-scale detectors or corner-based methods should be prioritized. Further research should focus on advanced multi-scale training strategies, such as scale normalization, to bridge the performance gap on complex structures before automated pipelines are fully trusted for critical infrastructure monitoring.

While the DIOR benchmark provides a robust foundation, users should note that instances are labeled using axis-aligned horizontal bounding boxes rather than oriented bounding boxes, which can introduce background noise for diagonally oriented objects. High confidence can be placed in the comparative rankings of the evaluated algorithms, but caution is warranted when deploying these models to detect complex structural classes under poor image conditions.

Cover for Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark

Abstract

Substantial efforts have been devoted more recently to presenting various methods for object detection in optical remote sensing images. However, the current survey of datasets and deep learning based methods for object detection in optical remote sensing images is not adequate. Moreover, most of the existing datasets have some shortcomings, for example, the numbers of images and object categories are small scale, and the image diversity and variations are insufficient. These limitations greatly affect the development of deep learning based object detection methods. In the paper, we provide a comprehensive review of the recent deep learning based object detection progress in both the computer vision and earth observation communities. Then, we propose a large-scale, publicly available benchmark for object DetectIon in Optical Remote sensing images, which we name as DIOR. The dataset contains 23463 images and 192472 instances, covering 20 object classes. The proposed DIOR dataset 1) is large-scale on the object categories, on the object instance number, and on the total image number; 2) has a large range of object size variations, not only in terms of spatial resolutions, but also in the aspect of inter- and intra-class size variability across objects; 3) holds big variations as the images are obtained with different imaging conditions, weathers, seasons, and image quality; and 4) has high inter-class similarity and intra-class diversity. The proposed benchmark can help the researchers to develop and validate their data-driven methods. Finally, we evaluate several state-of-the-art approaches on our DIOR dataset to establish a baseline for future research.

Table of Contents

  • 1. Introduction
  • 2. Review on Object Detection in Computer Vision Community
  • 2.1 Object Detection Datasets of Natural Scene Images
  • 2.2 Deep Learning Based Object Detection Methods in Computer Vision Community
  • 2.2.1 Region Proposal - based Methods
  • 2.2.2 Regression - based Methods
  • 3. Review on Object Detection in Earth Observation Community
  • 3.1 Object Detection Datasets of Optical Remote Sensing Images
  • 3.2 Deep Learning Based Object Detection Methods in Earth Observation Community
  • 4. Proposed DIOR Dataset
  • 4.1 Object Class Selection
  • 4.2 Characteristics of Our Proposed DIOR Dataset
  • 5. Benchmarking Representative Methods
  • 5.1 Experimental Setup
  • 5.2 Experimental Results
  • 6. Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — DIOR Benchmark Dataset Definition and Specifications

    definition

    The DIOR dataset is a large-scale, publicly available benchmark dataset designed for object detection in optical remote sensing images.

    The dataset possesses the following specifications:

    • Image Count and Dimensions: It contains 23,463 optical remote sensing images collected from Google Earth covering more than 80 countries. Each image has a fixed spatial size of 800×800800 \times 800 pixels.
    • Spatial Resolutions: Ground resolution varies between 0.5 m0.5\text{ m} and 30 m30\text{ m}.
    • Instance Count and Annotations: It contains 192,472 manually labeled object instances annotated using horizontal (axis-aligned) bounding boxes labeled via LabelMe.
    • Object Classes: It covers 20 geospatial object categories:
      1. Airplane
      2. Airport
      3. Baseball field
      4. Basketball court
      5. Bridge
      6. Chimney
      7. Dam
      8. Expressway service area
      9. Expressway toll station
      10. Golf course
      11. Ground track field
      12. Harbor
      13. Overpass
      14. Ship
      15. Stadium
      16. Storage tank
      17. Tennis court
      18. Train station
      19. Vehicle
      20. Wind mill

    Each object category contains approximately 1,200 images.

    Key dataset characteristics include:

    1. Scale: Substantially exceeds prior optical remote sensing datasets in total image count and category diversity.
    2. Extreme Size Variations: Includes significant intra-class and inter-class size variability (ranging from small targets such as vehicles to large facilities such as airports and ground track fields) driven by intrinsic object geometry and sensor resolution differences.
    3. Environmental Diversity: Captures variations in illumination, viewing angles, seasons, weather conditions, occlusions, and background clutter across global geographic locations.
    4. High Inter-Class Similarity & Intra-Class Diversity: Incorporates fine-grained and semantically overlapping classes (e.g., bridge vs. overpass, bridge vs. dam, stadium vs. ground track field, tennis court vs. basketball court) alongside diverse intra-class structural appearances.
  2. Knowl 2 — Comparison of Optical Remote Sensing Object Detection Datasets

    data/table

    The table below compares the DIOR dataset with nine prior publicly available earth observation object detection benchmarks across category count, image count, instance count, image dimensions, annotation format, and release year.

    Datasets # Categories # Images # Instances Image width (px) Annotation way Year
    TAS 1 30 1319 792 horizontal bounding box 2008
    SZTAKI-INRIA 1 9 665 ∼800\sim 800 oriented bounding box 2012
    NWPU VHR-10 10 800 3775 ∼1000\sim 1000 horizontal bounding box 2014
    VEDAI 9 1210 3640 1024 oriented bounding box 2015
    UCAS-AOD 2 910 6029 1280 horizontal bounding box 2015
    DLR 3K Vehicle 2 20 14235 5616 oriented bounding box 2015
    HRSC2016 1 1070 2976 ∼1000\sim 1000 oriented bounding box 2016
    RSOD 4 976 6950 ∼1000\sim 1000 horizontal bounding box 2017
    DOTA 15 2806 188282 800–4000 oriented bounding box 2017
    DIOR (ours) 20 23463 192472 800 horizontal bounding box 2018

    The DIOR dataset provides the largest number of images (23,463) and categories (20) among all listed datasets, and contains 192,472 object instances annotated with horizontal bounding boxes.

  3. Knowl 3 — DIOR Dataset Split and Class Distribution

    data/table

    The DIOR benchmark is split into a training-validation (trainval) set containing 11,725 images (50% of the dataset) and a test set containing 11,738 images (50% of the dataset). The trainval set is further partitioned into a training set (train, 5,862 images) and a validation set (val, 5,863 images).

    The table below details the number of images containing at least one instance of each class across the dataset splits. Because individual images can contain multiple object classes, the column totals do not equal the direct sums of per-class numbers.

    Class Train val Trainval Test
    Airplane 344 338 682 705
    Airport 326 327 653 657
    Baseball field 551 577 1128 1312
    Basketball court 336 329 665 704
    Bridge 379 495 874 1302
    Chimney 202 204 406 448
    Dam 238 246 484 502
    Expressway service area 279 281 560 565
    Expressway toll station 285 299 584 634
    Golf course 216 239 455 491
    Ground track field 536 454 990 1322
    Harbor 328 332 660 814
    Overpass 410 510 920 1099
    Ship 650 652 1302 1400
    Stadium 289 292 851 619
    Storage tank 391 384 775 839
    Tennis court 605 630 1235 1347
    Train station 244 249 493 501
    Vehicle 1556 1558 3114 3306
    Wind mill 404 403 807 809
    Total 5862 5863 11725 11738

    A predicted bounding box is considered correct if its Intersection over Union (IoU) with ground truth exceeds 50% (IoU>0.5\text{IoU} > 0.5); otherwise it is counted as a false positive. Detection performance is quantified using category-level Average Precision (AP, in %) and mean Average Precision (mAP, in %).

  4. Knowl 4 — Benchmark Evaluation of 12 Detectors on the DIOR Test Set

    data/table

    Twelve representative deep learning object detection methods were evaluated on the DIOR test set under standardized training protocols. The models comprise eight region proposal-based methods (R-CNN, RICNN, RICAOD, RIFD-CNN, Faster R-CNN, Faster R-CNN with FPN, Mask R-CNN with FPN, PANet) and four regression-based methods (SSD, YOLOv3, RetinaNet, CornerNet).

    Category index definitions:

    • c1: Airplane, c2: Airport, c3: Baseball field, c4: Basketball court, c5: Bridge, c6: Chimney, c7: Dam, c8: Expressway service area, c9: Expressway toll station, c10: Golf course
    • c11: Ground track field, c12: Harbor, c13: Overpass, c14: Ship, c15: Stadium, c16: Storage tank, c17: Tennis court, c18: Train station, c19: Vehicle, c20: Wind mill
    Method Backbone c1 c2 c3 c4 c5 c6 c7 c8 c9 c10 c11 c12 c13 c14 c15 c16 c17 c18 c19 c20 mAP
    R-CNN VGG16 35.6 43.0 53.8 62.3 15.6 53.7 33.7 50.2 33.5 50.1 49.3 39.5 30.9 9.1 60.8 18.0 54.0 36.1 9.1 16.4 37.7
    RICNN VGG16 39.1 61.0 60.1 66.3 25.3 63.3 41.1 51.7 36.6 55.9 58.9 43.5 39.0 9.1 61.1 19.1 63.5 46.1 11.4 31.5 44.2
    RICAOD VGG16 42.2 69.7 62.0 79.0 27.7 68.9 50.1 60.5 49.3 64.4 65.3 42.3 46.8 11.7 53.5 24.5 70.3 53.3 20.4 56.2 50.9
    RIFD-CNN VGG16 56.6 53.2 79.9 69.0 29.0 71.5 63.1 69.0 56.0 68.9 62.4 51.2 51.1 31.7 73.6 41.5 79.5 40.1 28.5 46.9 56.1
    Faster R-CNN VGG16 53.6 49.3 78.8 66.2 28.0 70.9 62.3 69.0 55.2 68.0 56.9 50.2 50.1 27.7 73.0 39.8 75.2 38.6 23.6 45.4 54.1
    SSD VGG16 59.5 72.7 72.4 75.7 29.7 65.8 56.6 63.5 53.1 65.3 68.6 49.4 48.1 59.2 61.0 46.6 76.3 55.1 27.4 65.7 58.6
    YOLOv3 Darknet-53 72.2 29.2 74.0 78.6 31.2 69.7 26.9 48.6 54.4 31.1 61.1 44.9 49.7 87.4 70.6 68.7 87.3 29.4 48.3 78.7 57.1
    Faster R-CNN + FPN ResNet-50 54.1 71.4 63.3 81.0 42.6 72.5 57.5 68.7 62.1 73.1 76.5 42.8 56.0 71.8 57.0 53.5 81.2 53.0 43.1 80.9 63.1
    Faster R-CNN + FPN ResNet-101 54.0 74.5 63.3 80.7 44.8 72.5 60.0 75.6 62.3 76.0 76.8 46.4 57.2 71.8 68.3 53.8 81.1 59.5 43.1 81.2 65.1
    Mask-RCNN + FPN ResNet-50 53.8 72.3 63.2 81.0 38.7 72.6 55.9 71.6 67.0 73.0 75.8 44.2 56.5 71.9 58.6 53.6 81.1 54.0 43.1 81.1 63.5
    Mask-RCNN + FPN ResNet-101 53.9 76.6 63.2 80.9 40.2 72.5 60.4 76.3 62.5 76.0 75.9 46.5 57.4 71.8 68.3 53.7 81.0 62.3 43.0 81.0 65.2
    RetinaNet ResNet-50 53.7 77.3 69.0 81.3 44.1 72.3 62.5 76.2 66.0 77.7 74.2 50.7 59.6 71.2 69.3 44.8 81.3 54.2 45.1 83.4 65.7
    RetinaNet ResNet-101 53.3 77.0 69.3 85.0 44.1 73.2 62.4 78.6 62.8 78.6 76.6 49.9 59.6 71.1 68.4 45.8 81.3 55.2 44.4 85.5 66.1
    PANet ResNet-50 61.9 70.4 71.0 80.4 38.9 72.5 56.6 68.4 60.0 69.0 74.6 41.6 55.8 71.7 72.9 62.3 81.2 54.6 48.2 86.7 63.8
    PANet ResNet-101 60.2 72.0 70.6 80.5 43.6 72.3 61.4 72.1 66.7 72.0 73.4 45.3 56.9 71.7 70.4 62.0 80.9 57.0 47.2 84.5 66.1
    CornerNet Hourglass-104 58.8 84.2 72.0 80.8 46.4 75.3 64.3 81.6 76.3 79.5 79.5 26.1 60.6 37.6 70.7 45.2 84.0 57.1 43.0 75.9 64.9

    RetinaNet (ResNet-101) and PANet (ResNet-101) achieve the highest overall mean Average Precision of 66.1%66.1\% on the DIOR test benchmark.

  5. Knowl 5 — Impact of Backbone Depth and Feature Pyramid Networks on Remote Sensing Object Detection

    empirical result

    Evaluation of 12 detectors on the DIOR benchmark demonstrates systematic trends regarding backbone depth and multiscale feature representations:

    1. Backbone Depth Hierarchy: Detection performance increases monotonically with network representation capability: ResNet-101≈Hourglass-104>ResNet-50≈Darknet-53>VGG16\text{ResNet-101} \approx \text{Hourglass-104} > \text{ResNet-50} \approx \text{Darknet-53} > \text{VGG16}. For example, replacing ResNet-50 with ResNet-101 improves mAP by +2.0%+2.0\% in Faster R-CNN with FPN (63.1%→65.1%63.1\% \to 65.1\%), +1.7%+1.7\% in Mask R-CNN with FPN (63.5%→65.2%63.5\% \to 65.2\%), +0.4%+0.4\% in RetinaNet (65.7%→66.1%65.7\% \to 66.1\%), and +2.3%+2.3\% in PANet (63.8%→66.1%63.8\% \to 66.1\%).
    2. Feature Pyramid Mechanisms: Incorporating Feature Pyramid Networks (FPN) and Path Aggregation Networks (PANet) provides substantial gains over standard architectures. Faster R-CNN with FPN (ResNet-50, 63.1%63.1\% mAP) outperforms standard Faster R-CNN (VGG16, 54.1%54.1\% mAP) by +9.0%+9.0\% mAP. Multi-scale feature extraction is critical for mitigating the extreme scale variations characteristic of earth observation imagery.
  6. Knowl 6 — Category-Specific Performance Characteristics on DIOR

    empirical result

    Benchmarking results across the 20 object classes in DIOR reveal distinct structural advantages for different detector design paradigms:

    1. Small-Sized Object Localization: YOLOv3 (with Darknet-53 and three-scale prediction) achieves high performance on small and dense targets, achieving 87.4%87.4\% AP on Ship (outperforming all other 11 models, where the next highest is 71.9%71.9\%) and leading performance on Storage tank (68.7%68.7\%) and Vehicle (48.3%48.3\%).
    2. Rotation-Invariant Formulations: Specialized rotation-handling modules produce gains over standard baselines:
      • RICNN (44.2%44.2\% mAP) outperforms R-CNN (37.7%37.7\% mAP) by +6.5%+6.5\% mAP.
      • RICAOD (50.9%50.9\% mAP) improves vehicle AP to 20.4%20.4\% vs. 9.1%9.1\% in R-CNN and airport AP to 69.7%69.7\% vs. 43.0%43.0\%.
      • RIFD-CNN (56.1%56.1\% mAP) outperforms Faster R-CNN VGG16 (54.1%54.1\% mAP) on rotation-sensitive targets such as airplanes (56.6%56.6\% vs. 53.6%53.6\%), ships (31.7%31.7\% vs. 27.7%27.7\%), and vehicles (28.5%28.5\% vs. 23.6%23.6\%).
    3. Keypoint-Based Detection: CornerNet (Hourglass-104), which models bounding boxes as pairs of corner keypoints with corner pooling, achieves the top AP on 9 out of 20 classes: Airport (84.2%84.2\%), Basketball court (80.8%80.8\%), Bridge (46.4%46.4\%), Chimney (75.3%75.3\%), Dam (64.3%64.3\%), Expressway service area (81.6%81.6\%), Expressway toll station (76.3%76.3\%), Golf course (79.5%79.5\%), and Ground track field (79.5%79.5\%).
  7. Knowl 7 — Performance Bottlenecks on Low-Accuracy DIOR Categories

    limitation

    Across all 12 evaluated deep learning architectures on the DIOR test set, several object categories consistently exhibit low detection accuracies:

    • Bridge: Best AP is 46.4%46.4\% (CornerNet), with baseline R-CNN achieving only 15.6%15.6\%.
    • Harbor: Best AP is 51.2%51.2\% (RIFD-CNN).
    • Overpass: Best AP is 60.6%60.6\% (CornerNet).
    • Vehicle: Best AP is 48.3%48.3\% (YOLOv3 / PANet ResNet-50), with multiple methods falling below 30%30\%.

    The low detection accuracies in these classes stem from:

    1. Complex, cluttered backgrounds and variable contextual surroundings in aerial imagery compared to natural scene images.
    2. Extreme aspect ratios (e.g., bridges, harbors) and arbitrary orientations that standard horizontal bounding boxes capture inefficiently.
    3. High visual and semantic overlap between fine-grained pairs (e.g., bridge vs. overpass, bridge vs. dam).
    4. Resolution degradation and small physical object footprint relative to the overall image tile size.

Coverage note — The paper's general literature survey summarizing existing natural-scene object detectors (R-CNN, Fast/Faster R-CNN, YOLOv1-v3, SSD, RetinaNet, CornerNet) and earlier remote sensing surveys was omitted as background material; all original benchmark contributions, dataset definitions, data tables, and empirical findings on DIOR are included.

References

  1. 1.Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., 2016. TensorFlow: a system for large‐scale machine learning. In: Proc. Conf. Oper. Syst. Des. Implement., pp. 265‐283.
  2. 2.Agarwal, S., Terrail, J.O.D., Jurie, F., 2018. Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks. arXiv preprint arXiv:1809.03193.
  3. 3.Aksoy, S., 2014. Detection of compound structures using a Gaussian mixture model with spectral and spatial constraints. IEEE Trans. Geosci. Remote Sens. 52, 6627‐6638.
  4. 4.Bai, X., Zhang, H., Zhou, J., 2014. VHR Object Detection Based on Structural Feature Extraction and Query Expansion. IEEE Trans. Geosci. Remote Sens. 52, 6508‐6520.
  5. 5.Bell, S., Lawrence Zitnick, C., Bala, K., Girshick, R., 2016. Inside‐outside net: Detecting objects in context with skip pooling and recurrent neural networks. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 2874‐2883.
  6. 6.Benedek, C., ., Descombes, X., ., Zerubia, J., . 2011. Building Development Monitoring in Multitemporal Remotely Sensed Image Pairs with Stochastic Birth‐Death Dynamics. IEEE Trans. Pattern Anal. Mach. Intell. 34, 33‐50.
  7. 7.Cai, Z., Fan, Q., Feris, R.S., Vasconcelos, N., 2016. A unified multi‐scale deep convolutional neural network for fast object detection. In: Proc. Eur. Conf. Comput. Vis., pp. 354‐370.
  8. 8.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L., 2018. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 40, 834‐848.
  9. 9.Cheng, G., Guo, L., Zhao, T., Han, J., Li, H., Fang, J., 2013a. Automatic landslide detection from remote‐sensing imagery using a scene classification method based on BoVW and pLSA. Int. J. Remote Sens. 34, 45‐59.
  10. 10.Cheng, G., Han, J., 2016. A Survey on Object Detection in Optical Remote Sensing Images. ISPRS J. Photogramm. Remote Sens. 117, 11‐28.
  11. 11.Cheng, G., Han, J., Guo, L., Qian, X., Zhou, P., Yao, X., Hu, X., 2013b. Object detection in remote sensing imagery using a discriminatively trained mixture model. ISPRS J. Photogramm. Remote Sens. 85, 32‐43.
  12. 12.Cheng, G., Han, J., Zhou, P., Guo, L., 2014. Multi‐class geospatial object detection and geographic image classification based on collection of part detectors. ISPRS J. Photogramm. Remote Sens. 98, 119‐132.
  13. 13.Cheng, G., Han, J., Zhou, P., Xu, D., 2019. Learning Rotation‐Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection. IEEE Trans. Image Process. 28, 265‐278.
  14. 14.Cheng, G., Yang, C., Yao, X., Guo, L., Han, J., 2018a. When deep learning meets metric learning: remote sensing image scene classification via learning discriminative CNNs. IEEE Trans. Geosci. Remote Sens. 56, 2811‐2821.
  15. 15.Cheng, G., Zhou, P., Han, J., 2016a. Learning rotation‐invariant convolutional neural networks for object detection in VHR optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 54, 7405‐7415.
  16. 16.Cheng, G., Zhou, P., Han, J., 2016b. RIFD‐CNN: Rotation‐Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 2884‐2893.
  17. 17.Cheng, L., Liu, X., Li, L., Jiao, L., Tang, X., 2018b. Deep Adaptive Proposal Network for Object Detection in Optical Remote Sensing Images. arXiv preprint arXiv:1807.07327.
  18. 18.Clément, F., Camille, C., Laurent, N., Yann, L., 2013. Learning hierarchical features for scene labeling. IEEE Trans. Pattern Anal. Mach. Intell. 35, 1915‐1929.
  19. 19.Cramer, M., 2010. The DGPF‐test on digital airborne camera evaluation‐overview and test design. Photogrammetrie ‐ Fernerkundung ‐ Geoinformation 2010, 73‐82.
  20. 20.Dai, J., Li, Y., He, K., Sun, J., 2016. R‐FCN: Object detection via region‐based fully convolutional networks. In: Proc. Conf. Adv. Neural Inform. Process. Syst., pp. 379‐387.
  21. 21.Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., Wei, Y., 2017. Deformable convolutional networks. In: Proc. IEEE Int. Conf. Comput. Vision, pp. 764‐773.
  22. 22.Das, S., Mirnalinee, T.T., Varghese, K., 2011. Use of Salient Features for the Design of a Multistage Framework to Extract Roads From High‐Resolution Multispectral Satellite Images. IEEE Trans. Geosci. Remote Sens. 49, 3906‐3931.
  23. 23.Deng, J., Dong, W., Socher, R., Li, L.‐J., Li, K., Fei‐Fei, L., 2009. Imagenet: A large‐scale hierarchical image database. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 248‐255.
  24. 24.Deng, Z., Sun, H., Zhou, S., Zhao, J., Zou, H., 2017. Toward fast and accurate vehicle detection in aerial images using coupled region‐based convolutional neural networks. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 10, 3652‐3664.
  25. 25.Ding, C., Li, Y., Xia, Y., Wei, W., Zhang, L., Zhang, Y., 2017. Convolutional Neural Networks Based Hyperspectral Image Classification Method with Adaptive Kernels. Remote Sensing 9, 618.
  26. 26.Everingham, M., Eslami, S.A., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A., 2015. The pascal visual object classes challenge: A retrospective. Int. J. Comput. Vis. 111, 98‐136.
  27. 27.Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A., 2010. The pascal visual object classes (voc) challenge. Int. J. Comput. Vis. 88, 303‐338.
  28. 28.Farooq, A., Hu, J., Jia, X., 2017. Efficient object proposals extraction for target detection in VHR remote sensing images. In: Proc. IEEE Int. Geosci. Remote Sens. Symposium, pp. 3337‐3340.
  29. 29.Felzenszwalb, P.F., Girshick, R.B., Mcallester, D., Ramanan, D., 2010. Object Detection with Discriminatively Trained Part‐Based Models. IEEE Trans. Pattern Anal. Mach. Intell. 32, 1627‐1645.
  30. 30.Fu, C.‐Y., Liu, W., Ranga, A., Tyagi, A., Berg, A.C., 2017. DSSD: Deconvolutional single shot detector. arXiv preprint arXiv:1701.06659.
  31. 31.Gidaris, S., Komodakis, N., 2015. Object Detection via a Multi‐region and Semantic Segmentation‐Aware CNN Model. In: Proc. IEEE Int. Conf. Comput. Vision, pp. 1134‐1142.
  32. 32.Girshick, R., 2015. Fast r‐cnn. In: Proc. IEEE Int. Conf. Comput. Vision, pp. 1440‐1448.
  33. 33.Girshick, R., Donahue, J., Darrell, T., Malik, J., 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 580‐587.
  34. 34.Guo, W., Yang, W., Zhang, H., Hua, G., 2018. Geospatial Object Detection in High Resolution Satellite Images Based on Multi‐Scale Convolutional Neural Network. Remote Sensing 10, 131.
  35. 35.Han, J., Zhang, D., Cheng, G., Guo, L., Ren, J., 2015. Object Detection in Optical Remote Sensing Images Based on Weakly Supervised Learning and High‐Level Feature Learning. IEEE Trans. Geosci. Remote Sens. 53, 3325‐3337.
  36. 36.Han, J., Zhang, D., Cheng, G., Liu, N., Xu, D., 2018. Advanced Deep‐Learning Techniques for Salient and Category‐Specific Object Detection: A Survey. IEEE Signal Processing Magazine 35, 84‐100.
  37. 37.Han, J., Zhou, P., Zhang, D., Cheng, G., Guo, L., Liu, Z., Bu, S., Wu, J., 2014. Efficient, simultaneous detection of multi‐class geospatial targets based on visual saliency modeling and discriminative learning of sparse coding. ISPRS J. Photogramm. Remote Sens. 89, 37‐48.
  38. 38.Han, X., Zhong, Y., Feng, R., Zhang, L., 2017a. Robust geospatial object detection based on pre‐trained faster R‐CNN framework for high spatial resolution imagery. In: Proc. IEEE Int. Geosci. Remote Sens. Symposium, pp. 3353‐3356.
  39. 39.Han, X., Zhong, Y., Zhang, L., 2017b. An Efficient and Robust Integrated Geospatial Object Detection Framework for High Spatial Resolution Remote Sensing Imagery. Remote Sensing 9, 666.
  40. 40.He, K., Gkioxari, G., Dollar, P., Girshick, R., 2017. Mask R‐CNN. IEEE Trans. Pattern Anal. Mach. Intell. PP, 1‐1.
  41. 41.He, K., Zhang, X., Ren, S., Sun, J., 2014. Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 37, 1904‐1916.
  42. 42.He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 770‐778.
  43. 43.Heitz, G., Koller, D., 2008. Learning Spatial Context: Using Stuff to Find Things. In: Proc. Eur. Conf. Comput. Vis., pp. 30‐43.
  44. 44.Hinton, G., Deng, L., Yu, D., Dahl, G.E., Mohamed, A.‐r., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T.N., 2012. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Processing Magazine 29, 82‐97.
  45. 45.Hou, R., Chen, C., Shah, M., 2017. Tube convolutional neural network (T‐CNN) for action detection in videos. In: Proc. IEEE Int. Conf. Comput. Vision, pp. 5822‐5831.
  46. 46.Hu, J., Shen, L., Sun, G., 2018. Squeeze‐and‐excitation networks. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 7132‐7141.
  47. 47.Huang, G., Liu, Z., Laurens, V.D.M., Weinberger, K.Q., 2017. Densely Connected Convolutional Networks. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 4700‐4708.
  48. 48.Ioffe, S., Szegedy, C., 2015. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In: Proc. IEEE Int. Conf. Machine Learning, pp. 448‐456.
  49. 49.Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., Darrell, T., 2014. Caffe: Convolutional architecture for fast feature embedding. In: Proc. ACM Int. Conf. Multimedia, pp. 675‐678.
  50. 50.Kong, T., Yao, A., Chen, Y., Sun, F., 2016. Hypernet: Towards accurate region proposal generation and joint object detection. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 845‐853.
  51. 51.Krizhevsky, A., Sutskever, I., Hinton, G.E., 2012. ImageNet Classification with Deep Convolutional Neural Networks. In: Proc. Conf. Adv. Neural Inform. Process. Syst., pp. 1097‐1105.
  52. 52.Law, H., Deng, J., 2018. Cornernet: Detecting objects as paired keypoints. In: Proc. Eur. Conf. Comput. Vis., pp. 734‐750.
  53. 53.Li, K., Cheng, G., Bu, S., You, X., 2018. Rotation‐Insensitive and Context‐Augmented Object Detection in Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 56, 2337‐2348.
  54. 54.Li, Z., Peng, C., Yu, G., Zhang, X., Deng, Y., Sun, J., 2017. Light‐head r‐cnn: In defense of two‐stage object detector. arXiv preprint arXiv:1711.07264.
  55. 55.Lin, H., Shi, Z., Zou, Z., 2017a. Fully Convolutional Network With Task Partitioning for Inshore Ship Detection in Optical Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 14, 1665‐1669.
  56. 56.Lin, T.‐Y., Dollár, P., Girshick, R.B., He, K., Hariharan, B., Belongie, S.J., 2017b. Feature Pyramid Networks for Object Detection. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 2117‐2125.
  57. 57.Lin, T.‐Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014. Microsoft coco: Common objects in context. In: Proc. Eur. Conf. Comput. Vis., pp. 740‐755.
  58. 58.Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollar, P., 2017c. Focal loss for dense object detection. IEEE Trans. Pattern Anal. Mach. Intell. PP, 2999‐3007.
  59. 59.Liu, K., Mattyus, G., 2015. Fast Multiclass Vehicle Detection on Aerial Images. IEEE Geosci. Remote Sens. Lett. 12, 1938‐1942.
  60. 60.Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., Pietikäinen, M., 2018a. Deep learning for generic object detection: A survey. arXiv preprint arXiv:1809.02165.
  61. 61.Liu, L., Pan, Z., Lei, B., 2017a. Learning a Rotation Invariant Detector with Rotatable Bounding Box. arXiv preprint arXiv:1711.09405.
  62. 62.Liu, S., Huang, D., Wang, Y., 2017b. Receptive Field Block Net for Accurate and Fast Object Detection. arXiv preprint arXiv:1711.07767.
  63. 63.Liu, S., Qi, L., Qin, H., Shi, J., Jia, J., 2018b. Path aggregation network for instance segmentation. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 8759‐8768.
  64. 64.Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C., 2016a. SSD: Single Shot MultiBox Detector. In: Proc. Eur. Conf. Comput. Vis., pp. 21‐37.
  65. 65.Liu, W., Ma, L., Chen, H., 2018c. Arbitrary‐Oriented Ship Detection Framework in Optical Remote‐Sensing Images. IEEE Geosci. Remote Sens. Lett. 15, 937‐941.
  66. 66.Liu, Z., Wang, H., Weng, L., Yang, Y., 2016b. Ship Rotated Bounding Box Space for Ship Extraction From High‐Resolution Optical Satellite Images With Complex Backgrounds. IEEE Geosci. Remote Sens. Lett. 13, 1074‐1078.
  67. 67.Long, J., Shelhamer, E., Darrell, T., 2015. Fully convolutional networks for semantic segmentation. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 3431‐3440.
  68. 68.Long, Y., Gong, Y., Xiao, Z., Liu, Q., 2017. Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks. IEEE Trans. Geosci. Remote Sens. 55, 2486‐2498.
  69. 69.Luan, S., Chen, C., Zhang, B., Han, J., Liu, J., 2018. Gabor Convolutional Networks. IEEE Trans. Image Process. 27, 4357‐4366.
  70. 70.Mikolov, T., Deoras, A., Povey, D., Burget, L., Cernocky, J., 2012. Strategies for training large scale neural network language models. In: Proc. IEEE Workshop Autom. Speech Recognit. Underst., pp. 196‐201.
  71. 71.Mordan, T., Thome, N., Henaff, G., Cord, M., 2018. End‐to‐End Learning of Latent Deformable Part‐Based Representations for Object Detection. Int. J. Comput. Vis., 1‐21.
  72. 72.Mundhenk, T.N., Konjevod, G., Sakla, W.A., Boakye, K., 2016. A large contextual dataset for classification, detection and counting of cars with deep learning. In: Proc. Eur. Conf. Comput. Vis., pp. 785‐800.
  73. 73.Newell, A., Yang, K., Deng, J., 2016. Stacked hourglass networks for human pose estimation. In: Proc. Eur. Conf. Comput. Vis., pp. 483‐499.
  74. 74.Ouyang, W., Zeng, X., Wang, K., Yan, J., Loy, C.C., Tang, X., Wang, X., Qiu, S., Luo, P., Tian, Y., 2017. DeepID‐Net: Object Detection with Deformable Part Based Convolutional Neural Networks. IEEE Trans. Pattern Anal. Mach. Intell. 39, 1320‐1334.
  75. 75.Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A., 2017. Automatic differentiation in pytorch. In: Proc. Conf. Adv. Neural Inform. Process. Syst. Workshop, pp. 1‐4.
  76. 76.Razakarivony, S., Jurie, F., 2015. Vehicle detection in aerial imagery : A small target detection benchmark. J. Vis. Commun. Image Represent. 34, 187‐203.
  77. 77.Redmon, J., Divvala, S., Girshick, R., Farhadi, A., 2016. You only look once: Unified, real‐time object detection. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 779‐788.
  78. 78.Redmon, J., Farhadi, A., 2017. YOLO9000: Better, Faster, Stronger. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 6517‐6525.
  79. 79.Redmon, J., Farhadi, A., 2018. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767.
  80. 80.Ren, S., He, K., Girshick, R., Sun, J., 2017. Faster R‐CNN: Towards Real‐Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 39, 1137‐1149.
  81. 81.Russell, B.C., Torralba, A., Murphy, K.P., Freeman, W.T., 2008. LabelMe: A Database and Web‐Based Tool for Image Annotation. Int. J. Comput. Vis. 77, 157‐173.
  82. 82.Salberg, A.B., 2015. Detection of seals in remote sensing images using features extracted from deep convolutional neural networks. In: Proc. IEEE Int. Geosci. Remote Sens. Symposium, pp. 1893‐1896.
  83. 83.Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., Lecun, Y., 2014. OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks. In: Proc. Int. Conf. Learn. Represent., pp. 1‐16.
  84. 84.Ševo, I., Avramović, A., 2017. Convolutional Neural Network Based Automatic Object Detection on Aerial Images. IEEE Geosci. Remote Sens. Lett. 13, 740‐744.
  85. 85.Shrivastava, A., Gupta, A., 2016. Contextual priming and feedback for faster r‐cnn. In: Proc. Eur. Conf. Comput. Vis., pp. 330‐348.
  86. 86.Shrivastava, A., Gupta, A., Girshick, R., 2016. Training Region‐Based Object Detectors with Online Hard Example Mining. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 761‐769.
  87. 87.Simonyan, K., Zisserman, A., 2015. Very deep convolutional networks for large‐scale image recognition. In: Proc. Int. Conf. Learn. Represent., pp. 1‐13.
  88. 88.Singh, B., Davis, L.S., 2018. An analysis of scale invariance in object detection ‐ SNIP. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 3578‐3587.
  89. 89.Singh, B., Li, H., Sharma, A., Davis, L.S., 2018a. R‐FCN‐3000 at 30fps: Decoupling Detection and Classification. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 1081‐1090.
  90. 90.Singh, B., Najibi, M., Davis, L.S., 2018b. SNIPER: Efficient multi‐scale training. In: Proc. Conf. Adv. Neural Inform. Process. Syst., pp. 9310‐9320.
  91. 91.Szegedy, C., Ioffe, S., Vanhoucke, V., Alemi, A.A., 2017. Inception‐v4, inception‐resnet and the impact of residual connections on learning. In: AAAI, p. 12.
  92. 92.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A., 2015. Going deeper with convolutions. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 1‐9.
  93. 93.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z., 2016. Rethinking the Inception Architecture for Computer Vision. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 2818‐2826.
  94. 94.Tang, T., Zhou, S., Deng, Z., Lei, L., Zou, H., 2017a. Arbitrary‐Oriented Vehicle Detection in Aerial Imagery with Single Convolutional Neural Networks. Remote Sensing 9, 1170.
  95. 95.Tang, T., Zhou, S., Deng, Z., Zou, H., Lei, L., 2017b. Vehicle Detection in Aerial Images Based on Region Convolutional Neural Networks and Hard Negative Example Mining. Sensors 17, 336.
  96. 96.Tanner, F., Colder, B., Pullen, C., Heagy, D., Eppolito, M., Carlan, V., Oertel, C., Sallee, P., 2009. Overhead imagery research data set — an annotated data library & tools to aid in the development of computer vision algorithms. In: Proc. IEEE Appl. Imag. Pattern Recognit. Workshop, pp. 1‐8.
  97. 97.Tian, Y., Chen, C., Shah, M., 2017. Cross‐view image matching for geo‐localization in urban environments. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 1998‐2006.
  98. 98.Tompson, J.J., Jain, A., LeCun, Y., Bregler, C., 2014. Joint training of a convolutional network and a graphical model for human pose estimation. In: Proc. Conf. Adv. Neural Inform. Process. Syst., pp. 1799‐1807.
  99. 99.Uijlings, J., R.R., Sande, V.D., K., E.A., Gevers, Smeulders, A., W.M., 2013. Selective Search for Object Recognition. Int. J. Comput. Vis. 104, 154‐171.
  100. 100.Wei, W., Zhang, J., Zhang, L., Tian, C., Zhang, Y., 2018. Deep Cube‐Pair Network for Hyperspectral Imagery Classification. Remote Sensing 10, 783.
  101. 101.Xia, G.‐S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., Zhang, L., 2018. DOTA: A large‐scale dataset for object detection in aerial images. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 3974‐3983.
  102. 102.Xiao, Z., Liu, Q., Tang, G., Zhai, X., 2015. Elliptic Fourier transformation‐based histograms of oriented gradients for rotationally invariant object detection in remote‐sensing images. Int. J. Remote Sens. 36, 618‐644.
  103. 103.Xu, Z., Xu, X., Wang, L., Yang, R., Pu, F., 2017. Deformable ConvNet with Aspect Ratio Constrained NMS for Object Detection in Remote Sensing Imagery. Remote Sensing 9, 1312.
  104. 104.Yang, J., Zhu, Y., Jiang, B., Gao, L., Xiao, L., Zheng, Z., 2018a. Aircraft detection in remote sensing images based on a deep residual network and Super‐Vector coding. Remote Sensing Letters 9, 229‐237.
  105. 105.Yang, X., Fu, K., Sun, H., Yang, J., Guo, Z., Yan, M., Zhan, T., Xian, S., 2018b. R2CNN++: Multi‐Dimensional Attention Based Rotation Invariant Detector with Robust Anchor Strategy. arXiv preprint arXiv:1811.07126.
  106. 106.Yang, Y., Zhuang, Y., Bi, F., Shi, H., Xie, Y., 2017. M‐FCN: Effective Fully Convolutional Network‐Based Airplane Detection Framework. IEEE Geosci. Remote Sens. Lett. 14, 1293‐1297.
  107. 107.Yao, Y., Jiang, Z., Zhang, H., Zhao, D., Cai, B., 2017. Ship detection in optical remote sensing images based on deep convolutional neural networks. Journal of Applied Remote Sensing 11, 1.
  108. 108.Yokoya, N., Iwasaki, A., 2015. Object Detection Based on Sparse Representation and Hough Voting for Optical Remote Sensing Imagery. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 8, 2053‐2062.
  109. 109.Yu, Y., Guan, H., Ji, Z., 2015. Rotation‐Invariant Object Detection in High‐Resolution Satellite Imagery Using Superpixel‐Based Deep Hough Forests. IEEE Geosci. Remote Sens. Lett. 12, 2183‐2187.
  110. 110.Zeiler, M.D., Fergus, R., 2014. Visualizing and understanding convolutional networks. In: Proc. Eur. Conf. Comput. Vis., pp. 818‐833.
  111. 111.Zhang, F., Du, B., Zhang, L., Xu, M., 2016. Weakly Supervised Learning Based on Coupled Convolutional Neural Networks for Aircraft Detection. IEEE Trans. Geosci. Remote Sens. 54, 5553‐5563.
  112. 112.Zhang, L., Shi, Z., Wu, J., 2017. A Hierarchical Oil Tank Detector With Deep Surrounding Features for High‐Resolution Optical Satellite Imagery. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 8, 4895‐4909.
  113. 113.Zhong, J., Lei, T., Yao, G., 2017. Robust vehicle detection in aerial images based on cascaded convolutional neural networks. Sensors 17, 2720.
  114. 114.Zhong, Y., Han, X., Zhang, L., 2018. Multi‐class geospatial object detection based on a position‐sensitive balancing framework for high spatial resolution remote sensing imagery. ISPRS J. Photogramm. Remote Sens. 138, 281‐294.
  115. 115.Zhou, P., Cheng, G., Liu, Z., Bu, S., Hu, X., 2016. Weakly supervised target detection in remote sensing images based on transferred deep features and negative bootstrapping. Multidimensional Systems and Signal Processing 27, 925‐944.
  116. 116.Zhu, H., Chen, X., Dai, W., Fu, K., Ye, Q., Jiao, J., 2015a. Orientation robust object detection in aerial images using deep convolutional neural network. In: Proc. IEEE Int. Conf. Image Processing, pp. 3735‐3739.
  117. 117.Zhu, X.X., Tuia, D., Mou, L., Xia, G.‐S., Zhang, L., Xu, F., Fraundorfer, F., 2017. Deep learning in remote sensing: a comprehensive review and list of resources. IEEE Geosci. Remote Sens. Magazine 5, 8‐36.
  118. 118.Zhu, Y., Urtasun, R., Salakhutdinov, R., Fidler, S., 2015b. segdeepm: Exploiting segmentation and context in deep neural networks for object detection. In: Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., pp. 4703‐4711.
  119. 119.Zitnick, C.L., Dollár, P., 2014. Edge Boxes: Locating Object Proposals from Edges. In: Proc. Eur. Conf. Comput. Vis., pp. 391‐405.
  120. 120.Zou, Z., Shi, Z., 2016. Ship Detection in Spaceborne Optical Image With SVD Networks. IEEE Trans. Geosci. Remote Sens. 54, 5832‐5845.

Citation

MLA
Li, K., et al. “Object Detection in Optical Remote Sensing Images: A Survey and a New Benchmark”. ISPRS Journal of Photogrammetry and Remote Sensing, vol. 159, 2020, pp. 296–307, https://doi.org/10.1016/j.isprsjprs.2019.11.023.
APA
Li, K., Wan, G., Cheng, G., Meng, L., & Han, J. (2020). Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing, 159, 296–307. https://doi.org/10.1016/j.isprsjprs.2019.11.023
Chicago
Li, K., G. Wan, G. Cheng, L. Meng, and J. Han. 2020. “Object Detection in Optical Remote Sensing Images: A Survey and a New Benchmark”. ISPRS Journal of Photogrammetry and Remote Sensing 159: 296–307. https://doi.org/10.1016/j.isprsjprs.2019.11.023.
Harvard
Li, K. et al. (2020) “Object detection in optical remote sensing images: A survey and a new benchmark”, ISPRS Journal of Photogrammetry and Remote Sensing, 159, pp. 296–307. Available at: https://doi.org/10.1016/j.isprsjprs.2019.11.023.
Vancouver
1. Li K, Wan G, Cheng G, Meng L, Han J (2020) Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing 159:296–307

BibTeX

@article{Li_2020, title={Object detection in optical remote sensing images: A survey and a new benchmark}, volume={159}, ISSN={0924-2716}, url={http://dx.doi.org/10.1016/j.isprsjprs.2019.11.023}, DOI={10.1016/j.isprsjprs.2019.11.023}, journal={ISPRS Journal of Photogrammetry and Remote Sensing}, publisher={Elsevier BV}, author={Li, Ke and Wan, Gang and Cheng, Gong and Meng, Liqiu and Han, Junwei}, year={2020}, month=Jan, pages={296–307} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF