DOTA: A Large-Scale Dataset for Object Detection in Aerial Images

Gui-Song XiaXiang BaiJian DingZhen ZhuSerge BelongieJiebo LuoMihai DatcuMarcello PelilloLiangpei Zhang

article2017CVPR3,064 citations

Presents DOTA, a large-scale aerial object detection benchmark featuring over 188,000 oriented bounding-box annotations across 15 categories, providing standard baseline evaluations for detecting multi-scale and arbitrarily oriented objects in Earth observation imagery.

Listen

The DOTA dataset was created to overcome the scarcity of large, realistic benchmarks for object detection in aerial imagery, where objects vary enormously in scale, orientation, and density and where prior collections were too small or idealized to support robust algorithm development. Researchers assembled 2,806 high-resolution images (roughly 4,000 by 4,000 pixels) from multiple sensors and platforms, then had domain experts label every instance of 15 common categories with oriented quadrilateral boxes rather than axis-aligned rectangles. The resulting collection contains 188,282 instances, far exceeding earlier aerial datasets in both volume and scene complexity.

To establish performance baselines, the authors adapted and ran leading detectorsincluding Faster R-CNN, R-FCN, YOLOv2, and SSDon two tasks: predicting horizontal boxes and predicting oriented boxes. Images were tiled into manageable patches for training and testing, after which results were merged and filtered with non-maximum suppression. Cross-dataset experiments further tested generalization by training on DOTA and evaluating on UCAS-AOD, and vice versa.

The experiments show that even the strongest current detectors achieve only modest accuracy on DOTA, with mean average precision ranging from roughly 30 percent to 60 percent on horizontal-box detection and dropping further when oriented boxes are required. Performance is especially weak on small, densely packed objects such as vehicles and ships, while larger, isolated categories fare better. Using oriented boxes improves localization in crowded or rotated scenes but exposes limitations in existing region-proposal and regression mechanisms. Cross-dataset tests confirm that DOTA contains a broader range of patterns than earlier collections and remains substantially harder.

These results indicate that aerial object detection cannot be solved by simply fine-tuning natural-scene models; new techniques are needed to handle extreme scale variation, arbitrary orientations, and high instance density within very large images. The dataset therefore supplies both a realistic training resource and a demanding benchmark that can guide development of detectors suitable for remote-sensing applications such as tracking, mapping, and autonomous navigation.

Future work should focus on architectures that explicitly model orientation and density, on efficient processing of full-resolution imagery without heavy cropping, and on continued expansion of DOTA to reflect evolving sensor and scene conditions. The main limitations are the computational overhead of handling gigapixel-scale images and the remaining performance gap on the hardest categories, which suggests that reported baselines should be treated as starting points rather than production-ready solutions.

  • Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN extends object detection frameworks to instance segmentation, building directly upon the bounding box detection concepts established in datasets like DOTA.
  • Paper: Grounded Language-Image Pre-training, Liunian Harold Li et al. (2022). Grounded Language-Inclusive Pre-training continues the work of object detection datasets by transferring learned representations to specialized domains such as aerial imagery.
  • Paper: Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression, Zhaohui Zheng et al. (2019). Distance-IoU Loss directly improves upon standard bounding-box regression techniques used in baseline detectors evaluated within large-scale aerial benchmarks like DOTA.
Cover for DOTA: A Large-Scale Dataset for Object Detection in Aerial Images

Abstract

Object detection is an important and challenging problem in computer vision. Although the past decade has witnessed major advances in object detection in natural scenes, such successes have been slow to aerial imagery, not only because of the huge variation in the scale, orientation and shape of the object instances on the earth's surface, but also due to the scarcity of well-annotated datasets of objects in aerial scenes. To advance object detection research in Earth Vision, also known as Earth Observation and Remote Sensing, we introduce a large-scale Dataset for Object deTection in Aerial images (DOTA). To this end, we collect 28062806 aerial images from different sensors and platforms. Each image is of the size about 4000-by-4000 pixels and contains objects exhibiting a wide variety of scales, orientations, and shapes. These DOTA images are then annotated by experts in aerial image interpretation using 1515 common object categories. The fully annotated DOTA images contains 188,282188,282 instances, each of which is labeled by an arbitrary (8 d.o.f.) quadrilateral To build a baseline for object detection in Earth Vision, we evaluate state-of-the-art object detection algorithms on DOTA. Experiments demonstrate that DOTA well represents real Earth Vision applications and are quite challenging.

Table of Contents

  • 1 Introduction
  • 2 Motivations
  • 3 Annotation of DOTA
  • 3.1 Images collection
  • 3.2 Category selection
  • 3.3 Annotation method
  • 3.4 Dataset splits
  • 4 Properties of DOTA
  • 4.1 Image size
  • 4.2 Various orientations of instances
  • 4.3 Spatial resolution information
  • 4.4 Various pixel size of categories
  • 4.5 Various aspect ratio of instances
  • 4.6 Various instance density of images
  • 5 Evaluations
  • 5.1 Tasks
  • 5.2 Evaluation prototypes
  • 5.3 Baselines with horizontal bounding boxes
  • 5.4 Baselines with oriented bounding boxes
  • 5.5 Experimental analysis
  • 6 Cross-dataset validations
  • 7 Conclusion
  • 8 Acknowledgement
  • References

Knowls

  1. Knowl 1 — DOTA Dataset Definition and Specifications

    definition

    The Dataset for Object deTection in Aerial images (DOTA) is a large-scale aerial image benchmark containing 2,8062{,}806 aerial images and 188,282188{,}282 annotated object instances. Image sizes range from approximately 800×800800 \times 800 to 4000×40004000 \times 4000 pixels, gathered from multiple sensors and platforms (including Google Earth) across diverse cities to reduce sensor and geographic dataset biases. Each image includes recorded geographic coordinates, acquisition timestamps, and spatial resolution (ground sample distance) metadata.

    The dataset covers 1515 object categories:

    1. Plane
    2. Ship
    3. Storage tank
    4. Baseball diamond
    5. Tennis court
    6. Swimming pool
    7. Ground track field
    8. Harbor
    9. Bridge
    10. Large vehicle
    11. Small vehicle
    12. Helicopter
    13. Roundabout
    14. Soccer ball field
    15. Basketball court

    These correspond to 1414 main categories, as large vehicle and small vehicle form sub-categories of vehicle. The dataset is partitioned into three splits: half (50%50\%) of the images form the training set, 1/61/6 (16.7%16.7\%) form the validation set, and 1/31/3 (33.3%33.3\%) form the test set.

  2. Knowl 2 — Oriented Bounding Box Annotation Scheme with Directional Vertex Ordering

    definition

    Objects in DOTA are annotated using arbitrary 88-degree-of-freedom (88 d.o.f.) oriented bounding box (OBB) quadrilaterals represented by four clockwise vertices:

    {(xi,yi)i=1,2,3,4}\{(x_i, y_i) \mid i = 1, 2, 3, 4\}

    where (xi,yi)R2(x_i, y_i) \in \mathbb{R}^2 denotes the pixel coordinates of the ii-th vertex in the image.

    The first vertex (x1,y1)(x_1, y_1) is designated with specific semantic meaning:

    • For directional or structured objects (helicopter, large vehicle, small vehicle, harbor, baseball diamond, ship, and plane), (x1,y1)(x_1, y_1) denotes the leading orientation ("head") or primary structural origin (e.g., the front of a plane/vehicle or the apex of a baseball diamond).
    • For objects lacking distinct visual orientation cues (soccer ball field, swimming pool, bridge, ground track field, basketball court, and tennis court), (x1,y1)(x_1, y_1) defaults to the top-left vertex.
  3. Knowl 3 — Oriented Bounding Box Regression Target Formulation for Faster R-CNN

    equation

    To enable two-stage object detectors such as Faster R-CNN to predict arbitrary 88 d.o.f. oriented quadrilaterals, the bounding box regression head is adapted from axis-aligned bounding boxes to oriented bounding boxes.

    Let an axis-aligned Region of Interest (RoI) generated by the Region Proposal Network (RPN) be defined by bounding coordinates (xmin,ymin,xmax,ymax)(x_{\min}, y_{\min}, x_{\max}, y_{\max}). The proposal rectangle RR is expressed as four ordered vertices R={(xi,yi)i=1,2,3,4}R = \{(x_i, y_i) \mid i = 1, 2, 3, 4\}:

    (x1,y1)=(xmin,ymin),(x2,y2)=(xmax,ymin)(x_1, y_1) = (x_{\min}, y_{\min}), \quad (x_2, y_2) = (x_{\max}, y_{\min})

    (x3,y3)=(xmax,ymax),(x4,y4)=(xmin,ymax)(x_3, y_3) = (x_{\max}, y_{\max}), \quad (x_4, y_4) = (x_{\min}, y_{\max})

    with proposal width w=xmaxxminw = x_{\max} - x_{\min} and height h=ymaxyminh = y_{\max} - y_{\min}.

    Given a matched ground-truth oriented quadrilateral G={(gxi,gyi)i=1,2,3,4}G = \{(g_{x_i}, g_{y_i}) \mid i = 1, 2, 3, 4\}, the network targets T={(txi,tyi)i=1,2,3,4}T = \{(t_{x_i}, t_{y_i}) \mid i = 1, 2, 3, 4\} for vertex offset regression are defined as:

    txi=gxixiw,tyi=gyiyih,for i{1,2,3,4}.t_{x_i} = \frac{g_{x_i} - x_i}{w}, \quad t_{y_i} = \frac{g_{y_i} - y_i}{h}, \quad \text{for } i \in \{1, 2, 3, 4\}.

  4. Knowl 4 — High-Resolution Image Patch Tiling and Evaluation Protocol

    algorithm

    Because original aerial images (800×800800 \times 800 to 4000×40004000 \times 4000 pixels) are too large for direct processing by convolutional neural networks, training and inference are performed using a patch-based sliding window protocol followed by prediction restoration.

    Input: Full aerial image II, patch dimension S=1024S = 1024, stride D=512D = 512, ground truth objects {Ok}\{O_k\} with area Ao,kA_{o,k}, NMS threshold τNMS\tau_{\text{NMS}}
    Output: Global bounding box predictions on full image II
    Partition II into a grid of overlapping patches of size S×SS \times S with stride DD
    for each patch PP and ground-truth object OkO_k intersecting PP do
        Let PiP_i be the divided object fragment falling inside PP with area aia_i
        Compute area overlap ratio UiaiAo,kU_i \leftarrow \frac{a_i}{A_{o,k}}
        if Ui<0.7U_i < 0.7 then
            Mark PiP_i as difficult (ignored during standard evaluation)
        else
            Retain category label and fit a 4-vertex clockwise oriented bounding box to PiP_i
        end if
    end for
    Train or evaluate detector on cropped patches
    Restore patch-level predicted boxes to global image coordinates (x+xoffset,y+yoffset)(x + x_{\text{offset}}, y + y_{\text{offset}})
    Apply category-wise Non-Maximum Suppression (NMS) on global predictions:
        Use τNMS=0.3\tau_{\text{NMS}} = 0.3 for horizontal bounding box (HBB) evaluation
        Use τNMS=0.1\tau_{\text{NMS}} = 0.1 for oriented bounding box (OBB) evaluation
    Compute mean Average Precision (mAP) under PASCAL VOC evaluation standard
  5. Knowl 5 — DOTA Benchmark Performance on Horizontal Bounding Box Detection

    data/table

    Standard object detection models evaluated on the DOTA benchmark under Horizontal Bounding Box (HBB) ground truth using PASCAL VOC mean Average Precision (mAP) demonstrate significant performance disparities across models and categories.

    Category YOLOv2 R-FCN Faster R-CNN (FR-H) SSD
    Plane 76.90 81.01 80.32 57.85
    Baseball diamond (BD) 33.87 58.96 77.55 32.79
    Bridge 22.73 31.64 32.86 16.14
    Ground track field (GTF) 34.88 58.97 68.13 18.67
    Small vehicle (SV) 38.73 49.77 53.66 0.05
    Large vehicle (LV) 32.02 45.04 52.49 36.93
    Ship 52.37 49.29 50.04 24.74
    Tennis court (TC) 61.65 68.99 90.41 81.16
    Basketball court (BC) 48.54 52.07 75.05 25.10
    Storage tank (ST) 33.91 67.42 59.59 47.47
    Soccer ball field (SBF) 29.27 41.83 57.00 11.22
    Roundabout (RA) 36.83 51.44 49.81 31.53
    Harbor 36.44 45.15 61.69 14.12
    Swimming pool (SP) 38.26 53.30 56.46 9.09
    Helicopter (HC) 11.61 33.89 41.85 0.00
    Average (mAP) 39.20 52.58 60.46 29.86

    Faster R-CNN (ResNet-101) achieves the highest mAP of 60.46%60.46\%, followed by R-FCN (ResNet-101) at 52.58%52.58\%, YOLOv2 (GoogLeNet) at 39.20%39.20\%, and SSD (InceptionV2) at 29.86%29.86\%. SSD performs poorly on small, crowded categories (0.05%0.05\% for small vehicle and 0.00%0.00\% for helicopter) because its random crop data augmentation destroys small instance features in aerial imagery.

  6. Knowl 6 — DOTA Benchmark Performance on Oriented Bounding Box Detection

    data/table

    Evaluating detection algorithms against Oriented Bounding Box (OBB) ground truths on DOTA demonstrates that models trained with explicit OBB regression (FR-O) significantly outperform models trained only with horizontal bounding boxes (FR-H, R-FCN, YOLOv2, SSD) evaluated on OBB.

    Category YOLOv2 R-FCN SSD Faster R-CNN (FR-H) Faster R-CNN (FR-O)
    Plane 52.75 39.57 41.06 49.74 79.42
    Baseball diamond (BD) 24.24 46.13 24.31 64.22 77.13
    Bridge 10.60 3.03 4.55 9.38 17.70
    Ground track field (GTF) 35.50 38.46 17.10 56.66 64.05
    Small vehicle (SV) 14.36 9.10 15.93 19.18 35.30
    Large vehicle (LV) 2.41 3.66 7.72 14.17 38.02
    Ship 7.37 7.45 13.21 9.51 37.16
    Tennis court (TC) 51.79 41.97 39.96 61.61 89.41
    Basketball court (BC) 43.98 50.43 12.05 65.47 69.64
    Storage tank (ST) 31.35 66.98 46.88 57.52 59.28
    Soccer ball field (SBF) 22.30 40.34 9.09 51.36 50.30
    Roundabout (RA) 36.68 51.28 30.82 49.41 52.91
    Harbor 14.61 11.14 1.36 20.80 47.89
    Swimming pool (SP) 22.55 35.59 3.50 45.84 47.40
    Helicopter (HC) 11.89 17.45 0.00 24.38 46.30
    Average (mAP) 25.492 30.84 17.84 39.95 54.13

    Faster R-CNN adapted for OBB (FR-O) reaches 54.13%54.13\% mAP, representing a 14.18%14.18\% improvement over the HBB-trained Faster R-CNN (FR-H, 39.95%39.95\% mAP), and large margins over HBB-trained YOLOv2 (25.492%25.492\%), R-FCN (30.84%30.84\%), and SSD (17.84%17.84\%). The gap is most severe for densely packed and high-aspect-ratio objects (such as ship, large vehicle, and harbor), where HBB predictions produce excessive overlap and are eliminated during non-maximum suppression.

  7. Knowl 7 — Cross-Dataset Generalization Analysis between DOTA and UCAS-AOD

    data/table

    Cross-dataset generalization experiments using YOLOv2 on the common categories (Plane and Small vehicle) demonstrate that DOTA provides substantially higher generalization capability and scene complexity compared to UCAS-AOD.

    Testing Set Detector Plane AP (%) Small-vehicle AP (%) Average mAP (%)
    UCAS-AOD YOLOv2-A 90.66 88.17 89.41
    UCAS-AOD YOLOv2-D 87.18 65.13 76.15
    DOTA YOLOv2-A 62.92 44.17 53.55
    DOTA YOLOv2-D 74.83 46.18 60.51

    YOLOv2-A denotes YOLOv2 trained on UCAS-AOD (1,110 training images, 400 test images), and YOLOv2-D denotes YOLOv2 trained on DOTA. When tested on DOTA, the model trained on UCAS-AOD suffers a performance drop of 35.86%35.86\% mAP (from 89.41%89.41\% to 53.55%53.55\%). Conversely, YOLOv2-D trained on DOTA experiences a drop of only 15.64%15.64\% mAP (from 76.15%76.15\% to 60.51%60.51\%) when evaluated across datasets, proving that DOTA covers a much broader distribution of aerial patterns, variations, and background complexities.

  8. Knowl 8 — Instance Pixel Size Distribution Comparison Across Object Detection Benchmarks

    data/table

    Measuring instance scale by the height of an instance's horizontal bounding box (categorized into Small: 1010--5050 pixels, Middle: 5050--300300 pixels, and Large: >300>300 pixels) highlights the structural size differences between DOTA and other benchmarks.

    Dataset 10105050 pixel (Small) 5050300300 pixel (Middle) Above 300300 pixel (Large)
    PASCAL VOC 0.14 0.61 0.25
    MSCOCO 0.43 0.49 0.08
    NWPU VHR-10 0.15 0.83 0.02
    DLR 3K Munich Vehicle 0.93 0.07 0.00
    DOTA 0.57 0.41 0.02

    While natural datasets like PASCAL VOC and aerial datasets like NWPU VHR-10 are dominated by middle-sized instances (61%61\% and 83%83\%) and DLR 3K Munich Vehicle is almost entirely small vehicles (93%93\%), DOTA maintains a balanced composition across small (57%57\%) and medium (41%41\%) instances. Additionally, DOTA features large inter-category size variation, ranging from small vehicles (30\,\approx 30 pixels) to bridges (1200\,\approx 1200 pixels), spanning a 40-fold scale differential.

  9. Knowl 9 — Instance Density and Aspect Ratio Characteristics of Aerial Objects in DOTA

    empirical result

    Aerial imagery in DOTA exhibits two distinct spatial properties compared to natural scene detection datasets:

    1. Instance Density and Overlap Freedom: DOTA images average 67.1067.10 bounding box annotations per image (reaching up to 2,0002{,}000 instances per image in dense scenarios like harbors and parking lots), far exceeding MSCOCO (7.197.19 boxes/image), PASCAL VOC (2.892.89 boxes/image), and ImageNet (1.371.37 boxes/image). Due to the overhead nadir/oblique viewpoint, dense clusters in aerial images exhibit minimal occlusion compared to ground-level scenes, allowing every instance to be individually annotated with an oriented bounding box rather than grouped under a single "crowd" attribute.

    2. Extreme Aspect Ratios: Aspect ratios (AR) measured for both minimal circumscribed horizontal rectangles and oriented bounding box quadrilaterals span a wide range in DOTA, with a notable frequency of instances possessing aspect ratios >5> 5 or 66 (such as bridges and harbors). Standard anchor-based detectors struggle to regress these extreme aspect ratios when oriented annotations are mapped to axis-aligned boxes.

Coverage note — None was omitted; all key contributions including dataset statistics, annotation format, OBB regression equations, patch-tiling inference algorithm, baseline evaluations (HBB and OBB), cross-dataset generalization, and instance size/density analyses are covered.

References

  1. 1.C. Benedek, X. Descombes, and J. Zerubia. Building development monitoring in multitemporal remotely sensed image pairs with stochastic birth-death dynamics. IEEE TPAMI, 34(1):33–50, 2012.
  2. 2.G. Cheng, P. Zhou, and J. Han. Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images. IEEE Trans. Geosci. Remote Sens., 54(12):7405–7415, 2016.
  3. 3.G. Cheng, P. Zhou, and J. Han. Rifd-cnn: Rotation-invariant and fisher discriminative convolutional neural networks for object detection. In CVPR, pages 2884–2893, 2016.
  4. 4.J. Dai, Y. Li, K. He, and J. Sun. R-FCN: object detection via region-based fully convolutional networks. In NIPS, pages 379–387, 2016.
  5. 5.A.-M. de Oca, R. Bahmanyar, N. Nistor, and M. Datcu. Earth observation image semantic bias: A collaborative user annotation approach. IEEE J. of Selected Topics in Applied Earth Observations and Remote Sensing, 2017.
  6. 6.J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009.
  7. 7.M. Everingham, L. V. Gool, C. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (VOC) challenge. IJCV, 88(2):303–338, 2010.
  8. 8.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, June 2016.
  9. 9.G. Heitz and D. Koller. Learning spatial context: Using stuff to find things. In ECCV, pages 30–43, 2008.
  10. 10.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. CoRR, abs/1502.03167, 2015.
  11. 11.D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. K. Ghosh, A. D. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny. ICDAR 2015 competition on robust reading. In Proc. ICDAR, 2015.
  12. 12.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 123(1):32–73, 2017.
  13. 13.M. Liao, B. Shi, and X. Bai. Textboxes++: A single-shot oriented scene text detector. CoRR, abs/1801.02765, 2018.
  14. 14.T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. Microsoft COCO: common objects in context. In ECCV, pages 740–755, 2014.
  15. 15.Y. Lin, H. He, Z. Yin, and F. Chen. Rotation-invariant object detection in remote sensing images based on radial-gradient angle. IEEE Geosci.Remote Sensing Lett., 12(4):746–750, 2015.
  16. 16.K. Liu and G. Mattyus. Fast multiclass vehicle detection on aerial images. IEEE Geosci. Remote Sensing Lett., 12(9):1938–1942, 2015.
  17. 17.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg. SSD: single shot multibox detector. In ECCV, pages 21–37, 2016.
  18. 18.Z. Liu, H. Wang, L. Weng, and Y. Yang. Ship rotated bounding box space for ship extraction from high-resolution optical satellite images with complex backgrounds. IEEE Geosci. Remote Sensing Lett., 13(8):1074–1078, 2016.
  19. 19.Y. Long, Y. Gong, Z. Xiao, and Q. Liu. Accurate object localization in remote sensing images based on convolutional neural networks. IEEE Trans. Geosci. Remote Sens., 55(5):2486–2498, 2017.
  20. 20.T. Moranduzzo and F. Melgani. Detecting cars in uav images with a catalog-based approach. IEEE Trans. Geosci. Remote Sens., 52(10):6356–6367, 2014.
  21. 21.T. N. Mundhenk, G. Konjevod, W. A. Sakla, and K. Boakye. A large contextual dataset for classification, detection and counting of cars with deep learning. In ECCV, pages 785–800, 2016.
  22. 22.A. O. Ok, Ç . Senaras, and B. Y ¨ uksel. Automated detec- ¨ tion of arbitrarily shaped buildings in complex environments from monocular VHR optical satellite imagery. IEEE Trans. Geosci. and Remote Sens., 51(3-2):1701–1717, 2013.
  23. 23.D. P. Papadopoulos, J. R. R. Uijlings, F. Keller, and V. Ferrari. Extreme clicking for efficient object annotation. CoRR, abs/1708.02750, 2017.
  24. 24.J. Porway, Q. Wang, and S. C. Zhu. A hierarchical and contextual model for aerial image parsing. IJCV, 88(2):254–283, 2010.
  25. 25.S. Razakarivony and F. Jurie. Vehicle detection in aerial imagery: A small target detection benchmark. J Vis. Commun. Image R., 34:187–203, 2016.
  26. 26.J. Redmon and A. Farhadi. YOLO9000: better, faster, stronger. CoRR, abs/1612.08242, 2016.
  27. 27.S. Ren, K. He, R. B. Girshick, and J. Sun. Faster R-CNN: towards real-time object detection with region proposal networks. IEEE TPAMI, 39(6):1137–1149, 2017.
  28. 28.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, pages 1–9, 2015.
  29. 29.A. Torralba and A. A. Efros. Unbiased look at dataset bias. In CVPR, pages 1521–1528, 2011.
  30. 30.M. Vakalopoulou, K. Karantzalos, N. Komodakis, and N. Paragios. Building detection in very high resolution multispectral data with deep learning features. In IGARSS, pages 1873–1876, 2015.
  31. 31.L. Wan, L. Zheng, H. Huo, and T. Fang. Affine invariant description and large-margin dimensionality reduction for target detection in optical remote sensing images. IEEE Geosci. Remote Sensing Lett., 2017.
  32. 32.G. Wang, X. Wang, B. Fan, and C. Pan. Feature extraction by rotation-invariant matrix representation for object detection in aerial image. IEEE Geosci.Remote Sensing Lett., 2017.
  33. 33.G. Xia, J. Hu, F. Hu, B. Shi, X. Bai, Y. Zhong, L. Zhang, and X. Lu. AID: A benchmark data set for performance evaluation of aerial scene classification. IEEE Trans. Geoscience and Remote Sensing, 55(7):3965–3981, 2017.
  34. 34.J. Xiao, J. Hays, K. Ehinger, A. Oliva, and A. Torralba. SUN database: Large-scale scene recognition from abbey to zoo. In CVPR, pages 3485–3492, 2010.
  35. 35.S. Yang, P. Luo, C. C. Loy, and X. Tang. WIDER FACE: A face detection benchmark. In CVPR, pages 5525–5533, 2016.
  36. 36.B. Yao, X. Yang, and S.-C. Zhu. Introduction to a large-scale general purpose ground truth database: Methodology, annotation tool and benchmarks. In EMMCVPR 2007, pages 169–183, 2007.
  37. 37.C. Yao, X. Bai, W. Liu, Y. Ma, and Z. Tu. Detecting texts of arbitrary orientations in natural images. In CVPR, 2012.
  38. 38.Q. You, J. Luo, H. Jin, and J. Yang. Building a large scale dataset for image emotion recognition: The fine print and the benchmark. In AAAI, pages 308–314, 2016.
  39. 39.F. Zhang, B. Du, L. Zhang, and M. Xu. Weakly supervised learning based on coupled convolutional neural networks for aircraft detection. IEEE Trans. Geosci. Remote Sens., 54(9):5553–5563, 2016.
  40. 40.B. Zhou, À. Lapedriza, J. Xiao, A. Torralba, and A. Oliva. Learning deep features for scene recognition using places database. In NIPS, pages 487–495, 2014.
  41. 41.H. Zhu, X. Chen, W. Dai, K. Fu, Q. Ye, and J. Jiao. Orientation robust object detection in aerial images using deep convolutional neural network. In ICIP, pages 3735–3739, 2015.

Citation

MLA
Xia, G.-S., et al. “DOTA: A Large-Scale Dataset for Object Detection in Aerial Images”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3974–83, https://doi.org/10.1109/CVPR.2018.00418.
APA
Xia, G.-S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., & Zhang, L. (2018). DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3974–3983. https://doi.org/10.1109/CVPR.2018.00418
Chicago
Xia, G.-S., X. Bai, J. Ding, et al. 2018. “DOTA: A Large-Scale Dataset for Object Detection in Aerial Images”. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3974–83. https://doi.org/10.1109/CVPR.2018.00418.
Harvard
Xia, G.-S. et al. (2018) “DOTA: A Large-Scale Dataset for Object Detection in Aerial Images”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp. 3974–3983. Available at: https://doi.org/10.1109/CVPR.2018.00418.
Vancouver
1. Xia G-S, Bai X, Ding J, Zhu Z, Belongie S, Luo J, Datcu M, Pelillo M, Zhang L (2018) DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, pp 3974–3983

BibTeX

@inproceedings{Xia_2018, title={DOTA: A Large-Scale Dataset for Object Detection in Aerial Images}, url={http://dx.doi.org/10.1109/CVPR.2018.00418}, DOI={10.1109/cvpr.2018.00418}, booktitle={2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition}, publisher={IEEE}, author={Xia, Gui-Song and Bai, Xiang and Ding, Jian and Zhu, Zhen and Belongie, Serge and Luo, Jiebo and Datcu, Mihai and Pelillo, Marcello and Zhang, Liangpei}, year={2018}, month=June, pages={3974–3983} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE