Remote Sensing Image Scene Classification: Benchmark and State of the Art

Gong ChengJunwei HanXiaoqiang Lu

article2017Proceedings of the IEEE2,911 citations

Introduces NWPU-RESISC45, a large-scale benchmark dataset of 31,500 images across 45 classes, paired with comprehensive baseline evaluations and a systematic survey to overcome data diversity and scale limitations in remote sensing scene classification.

Listen

Remote sensing image scene classification supports critical applications such as land-use mapping, natural hazard detection, urban planning, and environmental monitoring. The field has advanced through various methods and public datasets, yet existing collections remain small, lack image diversity, and show near-saturated accuracy, which restricts progress on modern data-driven techniques including deep learning.

This paper reviews roughly 170 publications on datasets and methods, introduces a new benchmark, and tests representative approaches to establish performance baselines. The authors first catalog six prior datasets and categorize methods into handcrafted features, unsupervised feature learning, and deep feature learning. They then release NWPU-RESISC45, a publicly available collection of 31,500 images spanning 45 scene classes with 700 images each, drawn from Google Earth across more than 100 countries. The images exhibit substantial variation in scale, viewpoint, illumination, occlusion, and background. Twelve methods—ranging from color histograms and local binary patterns to bag-of-visual-words variants and pre-trained or fine-tuned convolutional networks—were evaluated under 10 % and 20 % training splits using linear support-vector machines.

Deep convolutional features substantially outperformed earlier approaches, delivering at least 30 percentage points higher accuracy than handcrafted or unsupervised methods. Fine-tuning the networks on the new data produced further gains, with the best result reaching approximately 90 % overall accuracy under the 20 % training regime. Handcrafted global features remained weakest, while mid-level encodings such as bag-of-visual-words offered modest improvement but still trailed deep models. Confusion persisted between visually similar classes such as churches and palaces or dense and medium residential areas.

These results demonstrate that large, diverse benchmarks are essential for advancing scene classification and that current deep models already provide strong baselines. They also indicate that overhead imagery alone leaves semantic gaps that additional data sources could help close. The authors therefore recommend exploring fusion of satellite imagery with geo-tagged ground photos and location-based social-media streams to capture finer vertical and contextual details. They note that the reported figures rest on linear classifiers and fixed training ratios; broader testing with varied architectures and larger labeled sets would increase confidence before operational deployment.

Cover for Remote Sensing Image Scene Classification: Benchmark and State of the Art

Abstract

Remote sensing image scene classification plays an important role in a wide range of applications and hence has been receiving remarkable attention. During the past years, significant efforts have been made to develop various datasets or present a variety of approaches for scene classification from remote sensing images. However, a systematic review of the literature concerning datasets and methods for scene classification is still lacking. In addition, almost all existing datasets have a number of limitations, including the small scale of scene classes and the image numbers, the lack of image variations and diversity, and the saturation of accuracy. These limitations severely limit the development of new approaches especially deep learning-based methods. This paper first provides a comprehensive review of the recent progress. Then, we propose a large-scale dataset, termed "NWPU-RESISC45", which is a publicly available benchmark for REmote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class. The proposed NWPU-RESISC45 (i) is large-scale on the scene classes and the total image number, (ii) holds big variations in translation, spatial resolution, viewpoint, object pose, illumination, background, and occlusion, and (iii) has high within-class diversity and between-class similarity. The creation of this dataset will enable the community to develop and evaluate various data-driven algorithms. Finally, several representative methods are evaluated using the proposed dataset and the results are reported as a useful baseline for future research.

Table of Contents

  • I. INTRODUCTION
  • II. A REVIEW ON REMOTE SENSING IMAGE SCENE CLASSIFICATION DATASETS
  • A. UC Merced Land-Use Dataset
  • B. WHU-RS19 Dataset
  • C. SIRI-WHU Dataset
  • D. RSSCN7 Dataset
  • E. RSC11 Dataset
  • F. Brazilian Coffee Scene Dataset
  • III. A SURVEY ON REMOTE SENSING IMAGE SCENE CLASSIFICATION METHODS
  • A. Handcrafted Feature Based Methods
  • B. Unsupervised Feature Learning Based Methods
  • C. Deep Feature Learning Based Methods
  • IV. THE PROPOSED NWPU-RESISC45 DATASET
  • A. Selecting Scene Classes for NWPU-RESISC45 Dataset
  • B. The NWPU-RESISC45 Dataset
  • V. BENCHMARKING REPRESENTATIVE METHODS
  • A. Representative Methods
  • B. Experimental Setup
  • C. Evaluation Metrics
  • D. Experimental Results
  • VI. CONCLUSION AND DISCUSSION

Knowls

  1. Knowl 1 — NWPU-RESISC45 Benchmark Dataset Specification

    definition

    The NWPU-RESISC45 dataset is a large-scale, publicly available benchmark for remote sensing image scene classification created by Northwestern Polytechnical University (NWPU). The dataset contains 31,500 total images spanning 45 distinct scene classes, with exactly 700 images per class. All images are provided in the RGB color space with fixed spatial dimensions of 256×256256 \times 256 pixels.

    Spatial resolution across most classes ranges from approximately 30 m30\text{ m} to 0.2 m0.2\text{ m} per pixel, with lower spatial resolutions present in broad natural landscape classes (specifically island, lake, mountain, and snowberg). Imagery was curated by remote sensing interpretation experts from Google Earth, sampling geographic locations across more than 100 countries and regions spanning developing, transition, and developed economies.

    The 45 scene categories encompass land use and land cover classes, man-made structural objects, and natural landscape categories: airplane, airport, baseball diamond, basketball court, beach, bridge, chaparral, church, circular farmland, cloud, commercial area, dense residential, desert, forest, freeway, golf course, ground track field, harbor, industrial area, intersection, island, lake, meadow, medium residential, mobile home park, mountain, overpass, palace, parking lot, railway, railway station, rectangular farmland, river, roundabout, runway, sea ice, ship, snowberg, sparse residential, stadium, storage tank, tennis court, terrace, thermal power station, and wetland.

  2. Knowl 2 — Comparison of Public Remote Sensing Scene Classification Datasets

    data/table

    Prior to the introduction of NWPU-RESISC45, available remote sensing scene classification datasets exhibited significant constraints regarding class scale, sample count, and spatial resolution diversity. NWPU-RESISC45 substantially expands on existing benchmarks:

    Dataset Images per class Scene classes Total images Spatial resolution (m) Image sizes (pixels) Year
    UC Merced Land-Use 100 21 2,100 0.3 256×256256 \times 256 2010
    WHU-RS19 ∼50\sim 50 19 1,005 up to 0.5 600×600600 \times 600 2012
    SIRI-WHU 200 12 2,400 2 200×200200 \times 200 2016
    RSSCN7 400 7 2,800 – 400×400400 \times 400 2015
    RSC11 ∼100\sim 100 11 1,232 0.2 512×512512 \times 512 2016
    Brazilian Coffee Scene 1438 2 2,876 – 64×6464 \times 64 2015
    NWPU-RESISC45 700 45 31,500 ∼30\sim 30 to 0.2 256×256256 \times 256 2016

    NWPU-RESISC45 provides 15 times the total image volume of the widely used UC Merced benchmark, introduces fine-grained categories with high semantic overlap (e.g., circular vs. rectangular farmland; church vs. palace), and captures wide variations in translation, viewpoint, illumination, occlusion, and spatial resolution.

  3. Knowl 3 — NWPU-RESISC45 Benchmark Evaluation Protocol

    experimental setup

    To evaluate scene classification algorithms on the NWPU-RESISC45 dataset, a standardized evaluation protocol using two training-test split ratios is defined:

    1. 10%–90%10\%\text{--}90\% split: 10%10\% of the images per class (70 images per class, totaling 3,150 training samples) are randomly selected for training, and the remaining 90%90\% (630 images per class, totaling 28,350 samples) are used for testing.
    2. 20%–80%20\%\text{--}80\% split: 20%20\% of the images per class (140 images per class, totaling 6,300 training samples) are randomly selected for training, and the remaining 80%80\% (560 images per class, totaling 25,200 samples) are used for testing.

    For each split ratio, experiments are repeated 10 times with independent random splits, reporting the mean overall accuracy (%) and standard deviation. Because class sample counts are balanced (700700 images/class), overall accuracy is mathematically identical to average per-class accuracy.

    Classification across all 12 evaluated feature representations is conducted using a linear Support Vector Machine (SVM) in a one-versus-all configuration with regularization parameter C=1C=1 implemented via the LibSVM library. An unlabeled test sample is assigned the class label corresponding to the SVM with the highest decision response value.

  4. Knowl 4 — CNN Fine-Tuning Parameter Configuration for Scene Classification

    experimental setup

    For deep convolutional neural networks fine-tuned on the NWPU-RESISC45 dataset without data augmentation, a differential learning rate strategy is utilized across network layers:

    CNN Architecture Iterations Batch Size Base Learning Rate Last Layer Learning Rate Weight Decay Momentum
    AlexNet 15,000 128 0.001 0.01 0.0005 0.9
    VGGNet-16 15,000 50 0.001 0.01 0.0005 0.9
    GoogLeNet 15,000 128 0.001 0.01 0.0005 0.9

    Pre-trained model weights are initialized from ImageNet (ILSVRC). The final classification layer is assigned a higher learning rate (η=0.01\eta = 0.01) to allow rapid optimization away from ImageNet-specific class projections, while lower layers use η=0.001\eta = 0.001 to adapt convolutional representations without disrupting pre-trained generic feature extractors.

  5. Knowl 5 — Benchmark Evaluation of Handcrafted Visual Features on NWPU-RESISC45

    data/table

    Three global handcrafted visual feature extractors were evaluated using linear one-vs-all SVMs (C=1C=1) on the NWPU-RESISC45 dataset across 10 repeated experimental runs:

    • Color Histograms: Global RGB color histograms with 64 bins per channel (192-dimensional vector), normalized to unit L1L_1 norm.
    • Local Binary Patterns (LBP): Computed with neighbor count N=8N=8, generating a 256-dimensional histogram feature vector describing local texture patterns.
    • GIST: Holistic spatial layout descriptor computed using 8 orientations across 4 scales with a bank of Gabor filters, averaged over a 4×44 \times 4 spatial grid to yield a 512-dimensional vector.
    Handcrafted Features Training Ratio
    10% 20%
    Color histograms 24.84±0.22%24.84 \pm 0.22\% 27.52±0.14%27.52 \pm 0.14\%
    LBP 19.20±0.41%19.20 \pm 0.41\% 21.74±0.18%21.74 \pm 0.18\%
    GIST 15.90±0.23%15.90 \pm 0.23\% 17.88±0.22%17.88 \pm 0.22\%

    Global handcrafted features perform poorly on the 45-class benchmark (<28%<28\% overall accuracy), as they fail to capture the complex spatial layouts, resolution differences, and fine-grained semantic distinctions present in large-scale remote sensing scenes.

  6. Knowl 6 — Benchmark Evaluation of Unsupervised Feature Learning Methods on NWPU-RESISC45

    data/table

    Unsupervised mid-level feature representations were constructed from densely sampled SIFT descriptors (patch size 16×1616 \times 16 pixels, grid spacing 8 pixels) and classified with linear SVMs on NWPU-RESISC45:

    • Bag of Visual Words (BoVW): kk-means visual dictionary generation with hard quantization and global histogram pooling.
    • BoVW with Spatial Pyramid Matching (BoVW+SPM): BoVW representation pooled across 1×11 \times 1 and 2×22 \times 2 spatial subregions (5K5K-dimensional feature vector for codebook size KK).
    • Locality-constrained Linear Coding (LLC): Local coordinate projection using kk-nearest neighbors with max pooling over a kk-means codebook.

    Evaluating codebook sizes K∈{500,1000,2000,5000}K \in \{500, 1000, 2000, 5000\} demonstrated that performance depends heavily on dictionary capacity: optimal accuracy occurred at K=5000K=5000 for BoVW and LLC, and at K=500K=500 for BoVW+SPM under both training splits.

    Unsupervised Learning Methods Training Ratio
    10% 20%
    BoVW 41.72±0.21%41.72 \pm 0.21\% 44.97±0.28%44.97 \pm 0.28\%
    BoVW+SPM 27.83±0.61%27.83 \pm 0.61\% 32.96±0.47%32.96 \pm 0.47\%
    LLC 38.81±0.23%38.81 \pm 0.23\% 40.03±0.34%40.03 \pm 0.34\%

    Mid-level feature representations substantially improve over global handcrafted descriptors (+14%+14\% to +17%+17\% relative to color histograms), but their discriminative power remains constrained (≤44.97%\le 44.97\%) due to the lack of category label supervision during feature encoding.

  7. Knowl 7 — Benchmark Evaluation of Pre-trained CNN Features on NWPU-RESISC45

    data/table

    Deep convolutional neural network features extracted from networks pre-trained on ImageNet (without any dataset-specific fine-tuning) were evaluated on NWPU-RESISC45 using linear one-vs-all SVMs (C=1C=1):

    • AlexNet: 4,096-dimensional activation vector from the second fully connected layer (fc7\text{fc}_7).
    • VGGNet-16: 4,096-dimensional activation vector from the second fully connected layer (fc7\text{fc}_7).
    • GoogLeNet: 1,024-dimensional feature vector from the final global average pooling layer (pool5\text{pool}_5).
    Pre-trained CNN Features Training Ratio
    10% 20%
    AlexNet 76.69±0.21%76.69 \pm 0.21\% 79.85±0.13%79.85 \pm 0.13\%
    VGGNet-16 76.47±0.18%76.47 \pm 0.18\% 79.79±0.15%79.79 \pm 0.15\%
    GoogLeNet 76.19±0.38%76.19 \pm 0.38\% 78.48±0.26%78.48 \pm 0.26\%

    Generic deep features pre-trained on natural everyday images deliver an improvement of more than 34%34\% over unsupervised mid-level descriptors (reaching ∼79.8%\sim 79.8\% with 20% training data), showing that multi-layer hierarchical abstractions generalize effectively to overhead satellite and aerial imagery.

  8. Knowl 8 — Benchmark Evaluation of Fine-Tuned Deep CNNs on NWPU-RESISC45

    data/table

    End-to-end fine-tuning of deep CNN architectures on the NWPU-RESISC45 training sets (without data augmentation) yields substantial performance gains over fixed pre-trained feature extractors:

    Fine-Tuned CNN Models Training Ratio
    10% 20%
    Fine-tuned AlexNet 81.22±0.19%81.22 \pm 0.19\% 85.16±0.18%85.16 \pm 0.18\%
    Fine-tuned VGGNet-16 87.15±0.45%87.15 \pm 0.45\% 90.36±0.18%90.36 \pm 0.18\%
    Fine-tuned GoogLeNet 82.57±0.12%82.57 \pm 0.12\% 86.02±0.18%86.02 \pm 0.18\%

    Fine-tuning increases overall classification accuracy across all three network architectures by 5.3%5.3\% to 10.7%10.7\% compared to using pre-trained weights as fixed extractors. Fine-tuned VGGNet-16 delivers the highest overall accuracy on the benchmark, achieving 87.15%87.15\% under the 10%10\% training split and 90.36%90.36\% under the 20%20\% training split.

  9. Knowl 9 — Class Confusion Patterns and Failure Modes in Remote Sensing Scene Classification

    empirical result

    Analysis of confusion matrices on the NWPU-RESISC45 benchmark across different feature paradigms reveals specific failure modes:

    1. Color homogeneity confusion: Handcrafted color histogram features experience severe misclassifications between classes that share nearly identical green spectral distributions, most prominently between golf course and meadow.
    2. Structural and spatial layout confusion: Mid-level representations (BoVW) and deep CNN models exhibit noticeable confusions among categories sharing similar structural morphology or spatial arrangements:
      • church versus palace due to similar roof architectures and masonry compositions.
      • dense residential versus medium residential due to identical building components that differ only in spatial packing density.
    3. Consistent performance hierarchy: Across every individual scene class, per-class accuracies systematically follow the hierarchy: Handcrafted features << Unsupervised feature learning << Pre-trained CNNs << Fine-tuned CNNs.

Coverage note — The literature review of prior remote sensing scene classification datasets and feature extraction techniques (Sections II and III) was omitted as it synthesizes existing work rather than presenting novel contributions of this paper.

References

  1. 1.A. Plaza, J. Plaza, A. Paz, and S. Sanchez, “Parallel hyperspectral image and signal processing,” IEEE Signal Processing Magazine, vol. 28, no. 3, pp. 119-126, 2011.
  2. 2.H. M. Cantalloube and C. E. Nahum, “Airborne SAR-efficient signal processing for very high resolution,” Proceedings of the IEEE, vol. 101, no. 3, pp. 784-797, 2013.
  3. 3.L. Gómez-Chova, D. Tuia, G. Moser, and G. Camps-Valls, “Multimodal classification of remote sensing images: a review and future directions,” Proceedings of the IEEE, vol. 103, no. 9, pp. 1560-1584, 2015.
  4. 4.P. Gamba, “Human settlements: A global challenge for EO data processing and interpretation,” Proceedings of the IEEE, vol. 101, no. 3, pp. 570-581, 2013.
  5. 5.T. R. Martha, N. Kerle, C. J. Van Westen, V. Jetten, and K. V. Kumar, “Segment optimization and data-driven thresholding for knowledge-based landslide detection by object-based image analysis,” IEEE Trans. Geosci. Remote Sens., vol. 49, no. 12, pp. 4928-4943, 2011.
  6. 6.G. Cheng, L. Guo, T. Zhao, J. Han, H. Li, and J. Fang, “Automatic landslide detection from remote-sensing imagery using a scene classification method based on BoVW and pLSA,” Int. J. Remote Sens., vol. 34, no. 1, pp. 45-59, 2013.
  7. 7.A. Stumpf and N. Kerle, “Object-oriented mapping of landslides using Random Forests,” Remote Sens. Environ., vol. 115, no. 10, pp. 2564-2577, 2011.
  8. 8.Q. Zhu, Y. Zhong, B. Zhao, G.-S. Xia, and L. Zhang, “Bag-of-Visual-Words Scene Classifier With Local and Global Features for High Spatial Resolution Remote Sensing Imagery,” IEEE Geosci. Remote Sens. Lett., vol. 13, no. 6, pp. 747-751, 2016.
  9. 9.L. Zhao, P. Tang, and L. Huo, “Feature significance-based multibag-of-visual-words model for remote sensing image scene classification,” J. Appl. Remote Sens., vol. 10, no. 3, pp. 035004-035004, 2016.
  10. 10.B. Zhao, Y. Zhong, L. Zhang, and B. Huang, “The Fisher Kernel coding framework for high spatial resolution scene classification,” Remote Sensing, vol. 8, no. 2, pp. 157, 2016.
  11. 11.B. Zhao, Y. Zhong, G.-S. Xia, and L. Zhang, “Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 4, pp. 2108-2123, 2016.
  12. 12.H. Yu, W. Yang, G.-S. Xia, and G. Liu, “A Color-Texture-Structure Descriptor for High-Resolution Satellite Image Classification,” Remote Sensing, vol. 8, no. 3, pp. 259, 2016.
  13. 13.X. Yao, J. Han, G. Cheng, X. Qian, and L. Guo, “Semantic Annotation of High-Resolution Satellite Images via Weakly Supervised Learning,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 6, pp. 3660-3671, 2016.
  14. 14.H. Wu, B. Liu, W. Su, W. Zhang, and J. Sun, “Hierarchical Coding Vectors for Scene Level Land-Use Classification,” Remote Sensing, vol. 8, no. 5, pp. 436, 2016.
  15. 15.Y. Liu, Y.-M. Zhang, X.-Y. Zhang, and C.-L. Liu, “Adaptive spatial pooling for image classification,” Pattern Recog., vol. 55, pp. 58-67, 2016.
  16. 16.S. Cui, “Comparison of approximation methods to Kullback–Leibler divergence between Gaussian mixture models for satellite image retrieval,” Remote Sens. Lett., vol. 7, no. 7, pp. 651-660, 2016.
  17. 17.Q. Zou, L. Ni, T. Zhang, and Q. Wang, “Deep learning based feature selection for remote sensing scene classification,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 11, pp. 2321-2325, 2015.
  18. 18.G.-S. Xia, Z. Wang, C. Xiong, and L. Zhang, “Accurate Annotation of Remote Sensing Images via Active Spectral Clustering with Little Expert Knowledge,” Remote Sensing, vol. 7, no. 11, pp. 15014-15045, 2015.
  19. 19.K. Qi, H. Wu, C. Shen, and J. Gong, “Land-Use Scene Classification in High-Resolution Remote Sensing Images Using Improved Correlatons,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 12, pp. 2403-2407, 2015.
  20. 20.M. L. Mekhalfi, F. Melgani, Y. Bazi, and N. Alajlan, “Land-use classification with compressive sensing multifeature fusion,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 10, pp. 2155-2159, 2015.
  21. 21.J. Li, X. Huang, P. Gamba, J. M. Bioucas-Dias, L. Zhang, J. A. Benediktsson, and A. Plaza, “Multiple feature learning for hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 3, pp. 1592-1606, 2015.
  22. 22.G. Cheng, P. Zhou, J. Han, L. Guo, and J. Han, “Auto-encoder-based shared mid-level visual dictionary learning for scene classification using very high resolution remote sensing images,” IET Computer Vision, vol. 9, no. 5, pp. 639-647, 2015.
  23. 23.G. Cheng, J. Han, L. Guo, Z. Liu, S. Bu, and J. Ren, “Effective and efficient midlevel visual elements-oriented land-use classification using VHR remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 8, pp. 4238-4249, 2015.
  24. 24.Y. Zhang, X. Zheng, G. Liu, X. Sun, H. Wang, and K. Fu, “Semi-supervised manifold learning based multigraph fusion for high-resolution remote sensing image classification,” IEEE Geosci. Remote Sens. Lett., vol. 11, no. 2, pp. 464-468, 2014.
  25. 25.D. Tuia, M. Volpi, M. Dalla Mura, A. Rakotomamonjy, and R. Flamary, “Automatic feature learning for spatio-spectral image classification with sparse SVM,” IEEE Trans. Geosci. Remote Sens., vol. 52, no. 10, pp. 6062-6074, 2014.
  26. 26.A. M. Cheriyadat, “Unsupervised feature learning for aerial scene classification,” IEEE Trans. Geosci. Remote Sens., vol. 52, no. 1, pp. 439-451, 2014.
  27. 27.G. Cheng, J. Han, P. Zhou, and L. Guo, “Multi-class geospatial object detection and geographic image classification based on collection of part detectors,” ISPRS J. Photogramm. Remote Sens., vol. 98, pp. 119-132, 2014.
  28. 28.X. Zheng, X. Sun, K. Fu, and H. Wang, “Automatic annotation of satellite images via multifeature joint sparse coding with spatial relation constraint,” IEEE Geosci. Remote Sens. Lett., vol. 10, no. 4, pp. 652-656, 2013.
  29. 29.Y. Zhang, X. Sun, H. Wang, and K. Fu, “High-resolution remote-sensing image classification via an approximate earth mover's distance-based bag-of-features model,” IEEE Geosci. Remote Sens. Lett., vol. 10, no. 5, pp. 1055-1059, 2013.
  30. 30.V. Risojević and Z. Babić, “Fusion of global and local descriptors for remote sensing image classification,” IEEE Geosci. Remote Sens. Lett., vol. 10, no. 4, pp. 836-840, 2013.
  31. 31.J. Li, P. R. Marpu, A. Plaza, J. M. Bioucas-Dias, and J. A. Benediktsson, “Generalized composite kernel framework for hyperspectral image classification,” IEEE Trans. Geosci. Remote Sens., vol. 51, no. 9, pp. 4816-4829, 2013.
  32. 32.J. Li, J. M. Bioucas-Dias, and A. Plaza, “Spectral–spatial classification of hyperspectral data using loopy belief propagation and active learning,” IEEE Trans. Geosci. Remote Sens., vol. 51, no. 2, pp. 844-856, 2013.
  33. 33.G. Sheng, W. Yang, T. Xu, and H. Sun, “High-resolution satellite scene classification using a sparse coding based multiple feature combination,” Int. J. Remote Sens., vol. 33, no. 8, pp. 2395-2412, 2012.
  34. 34.S. Moustakidis, G. Mallinis, N. Koutsias, J. B. Theocharis, and V. Petridis, “SVM-based fuzzy decision trees for classification of high spatial resolution remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 50, no. 1, pp. 149-169, 2012.
  35. 35.N. Longbotham, C. Chaapel, L. Bleiler, C. Padwick, W. J. Emery, and F. Pacifici, “Very high resolution multiangle urban classification analysis,” IEEE Trans. Geosci. Remote Sens., vol. 50, no. 4, pp. 1155-1170, 2012.
  36. 36.Y. Yang and S. Newsam, "Spatial pyramid co-occurrence for image classification," in Proc. IEEE Int. Conf. Comput. Vision, 2011, pp. 1465-1472.
  37. 37.D. Dai and W. Yang, “Satellite image classification via two-layer sparse coding with biased image representation,” IEEE Geosci. Remote Sens. Lett., vol. 8, no. 1, pp. 173-176, 2011.
  38. 38.Y. Yang and S. Newsam, "Bag-of-visual-words and spatial extensions for land-use classification," in Proc. ACM SIGSPATIAL Int. Conf. Adv. Geogr. Inform. Syst., 2010, pp. 270-279.
  39. 39.S. Xu, T. Fang, D. Li, and S. Wang, “Object classification of aerial images with bag-of-visual words,” IEEE Geosci. Remote Sens. Lett., vol. 7, no. 2, pp. 366-370, 2010.
  40. 40.M. Lienou, H. Maître, and M. Datcu, “Semantic annotation of satellite images using latent dirichlet allocation,” IEEE Geosci. Remote Sens. Lett., vol. 7, no. 1, pp. 28-32, 2010.
  41. 41.H. Sridharan and A. Cheriyadat, “Bag of lines (bol) for improved aerial scene representation,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 3, pp. 676-680, 2015.
  42. 42.R. Kusumaningrum, H. Wei, R. Manurung, and A. Murni, “Integrated visual vocabulary in latent Dirichlet allocation–based scene classification for IKONOS image,” J. Appl. Remote Sens., vol. 8, no. 1, pp. 083690-083690, 2014.
  43. 43.F. Hu, W. Yang, J. Chen, and H. Sun, “Tile-level annotation of satellite images using multi-level max-margin discriminative random field,” Remote Sensing, vol. 5, no. 5, pp. 2275-2291, 2013.
  44. 44.G. Cheng, J. Han, L. Guo, X. Qian, P. Zhou, X. Yao, and X. Hu, “Object detection in remote sensing imagery using a discriminatively trained mixture model,” ISPRS J. Photogramm. Remote Sens., vol. 85, pp. 32-43, 2013.
  45. 45.G. Cheng, P. Zhou, and J. Han, “Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 12, pp. 7405-7415, 2016.
  46. 46.J. Han, D. Zhang, G. Cheng, L. Guo, and J. Ren, “Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 6, pp. 3325-3337, 2015.
  47. 47.J. Han, P. Zhou, D. Zhang, G. Cheng, L. Guo, Z. Liu, S. Bu, and J. Wu, “Efficient, simultaneous detection of multi-class geospatial targets based on visual saliency modeling and discriminative learning of sparse coding,” ISPRS J. Photogramm. Remote Sens., vol. 89, pp. 37-48, 2014.
  48. 48.X. Yao, J. Han, L. Guo, S. Bu, and Z. Liu, “A coarse-to-fine model for airport detection from remote sensing images using target-oriented visual saliency and CRF,” Neurocomputing, vol. 164, pp. 162-172, 2015.
  49. 49.D. Zhang, J. Han, G. Cheng, Z. Liu, S. Bu, and L. Guo, “Weakly supervised learning for target detection in remote sensing images,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 4, pp. 701-705, 2015.
  50. 50.P. Zhou, G. Cheng, Z. Liu, S. Bu, and X. Hu, “Weakly supervised target detection in remote sensing images based on transferred deep features and negative bootstrapping,” Multidimens. Syst. Signal Process., vol. 27, no. 4, pp. 925-944, 2016.
  51. 51.S. Bhagavathy and B. S. Manjunath, “Modeling and detection of geospatial objects using texture motifs,” IEEE Trans. Geosci. Remote Sens., vol. 44, no. 12, pp. 3706-3715, 2006.
  52. 52.N. Li, M. Cao, C. He, B. Wu, J. Jiao, and X. Yang, “A multi-parametric indicator design for ECT sensor optimization used in oil transmission,” IEEE Sensors Journal, DOI: 10.1109/JSEN.2017.2664864, 2017.
  53. 53.Y. Wang, L. Zhang, X. Tong, L. Zhang, Z. Zhang, H. Liu, X. Xing, and P. T. Mathiopoulos, “A Three-Layered Graph-Based Learning Approach for Remote Sensing Image Retrieval,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 10, pp. 6020-6034, 2016.
  54. 54.Y. Li, Y. Zhang, C. Tao, and H. Zhu, “Content-Based High-Resolution Remote Sensing Image Retrieval via Unsupervised Feature Learning and Collaborative Affinity Metric Fusion,” Remote Sensing, vol. 8, no. 9, pp. 709, 2016.
  55. 55.Y. Yang and S. Newsam, “Geographic image retrieval using local invariant features,” IEEE Trans. Geosci. Remote Sens., vol. 51, no. 2, pp. 818-832, 2013.
  56. 56.J. A. dos Santos, O. A. B. Penatti, and R. da Silva Torres, "Evaluating the Potential of Texture and Color Descriptors for Remote Sensing Image Retrieval and Classification," in Proc. VISAPP, 2010, pp. 203-208.
  57. 57.C. Shyu, M. Klaric, G. J. Scott, A. S. Barb, C. H. Davis, and K. Palaniappan, “GeoIRIS: Geospatial information retrieval and indexing system—content mining, semantics modeling, and complex queries,” IEEE Trans. Geosci. Remote Sens., vol. 45, no. 4, pp. 839-852, 2007.
  58. 58.M. Schroder, H. Rehrauer, K. Seidel, and M. Datcu, “Interactive learning and probabilistic retrieval in remote sensing image archives,” IEEE Trans. Geosci. Remote Sens., vol. 38, no. 5, pp. 2288-2298, 2000.
  59. 59.L. Gueguen, “Classifying compound structures in satellite images: A compressed representation for fast queries,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 4, pp. 1803-1818, 2015.
  60. 60.J. Munoz-Mari, D. Tuia, and G. Camps-Valls, “Semisupervised classification of remote sensing images with active queries,” IEEE Trans. Geosci. Remote Sens., vol. 50, no. 10, pp. 3751-3763, 2012.
  61. 61.Z. Du, X. Li, and X. Lu, “Local Structure Learning in High Resolution Remote Sensing Image Retrieval,” Neurocomputing, vol. 207, pp. 813-822, 2016.
  62. 62.E. Aptoula, “Remote sensing image retrieval with global morphological texture descriptors,” IEEE Trans. Geosci. Remote Sens., vol. 52, no. 5, pp. 3023-3034, 2014.
  63. 63.W. Zhou, Z. Shao, C. Diao, and Q. Cheng, “High-resolution remote-sensing imagery retrieval using sparse features by auto-encoder,” Remote Sens. Lett., vol. 6, no. 10, pp. 775-783, 2015.
  64. 64.M. Kim, M. Madden, and T. A. Warner, “Forest type mapping using object-specific texture measures from multispectral Ikonos imagery: segmentation quality and image classification issues,” Photogramm. Eng. Remote Sens., vol. 75, no. 7, pp. 819-829, 2009.
  65. 65.X. Li and G. Shao, “Object-based urban vegetation mapping with high-resolution aerial photography as a single data source,” Int. J. Remote Sens., vol. 34, no. 3, pp. 771-789, 2013.
  66. 66.N. B. Mishra and K. A. Crews, “Mapping vegetation morphology types in a dry savanna ecosystem: integrating hierarchical object-based image analysis with Random Forest,” Int. J. Remote Sens., vol. 35, no. 3, pp. 1175-1198, 2014.
  67. 67.S. R. Phinn, C. M. R. Mumby, and P. J., “Multi-scale, object-based image analysis for mapping geomorphic and ecological zones on coral reefs,” Int. J. Remote Sens., vol. 33, no. 12, pp. 3768-3797, 2012.
  68. 68.J. S. Walker and J. M. Briggs, “An object-oriented approach to urban forest mapping in Phoenix,” Photogramm. Eng. Remote Sens., vol. 73, no. 5, pp. 577-583, 2007.
  69. 69.L. Janssen and H. Middelkoop, “Knowledge-based crop classification of a Landsat Thematic Mapper image,” Int. J. Remote Sens., vol. 13, no. 15, pp. 2827-2837, 1992.
  70. 70.T. Blaschke and J. Strobl, “What’s wrong with pixels? Some recent developments interfacing remote sensing and GIS,” GeoBIT/GIS, vol. 6, no. 1, pp. 12-17, 2001.
  71. 71.T. Blaschke, “Object based image analysis for remote sensing,” ISPRS J. Photogramm. Remote Sens., vol. 65, no. 1, pp. 2-16, 2010.
  72. 72.T. Blaschke, S. Lang, and G. J. Hay, Object-based image analysis: spatial concepts for knowledge-driven remote sensing applications, Heidelberg, Berlin, New York: Springer, 2008.
  73. 73.T. Blaschke, G. J. Hay, M. Kelly, S. Lang, P. Hofmann, E. Addink, R. Q. Feitosa, F. van der Meer, H. van der Werff, and F. van Coillie, “Geographic object-based image analysis–towards a new paradigm,” ISPRS J. Photogramm. Remote Sens., vol. 87, pp. 180-191, 2014.
  74. 74.T. Blaschke, "Object-based contextual image classification built on image segmentation," in Proc. IEEE Workshop on Advances in Techniques for Analysis of Remotely Sensed Data, 2003, pp. 113-119.
  75. 75.T. Blaschke, C. Burnett, and A. Pekkarinen, Image Segmentation Methods for Object-based Analysis and Classification: Springer Netherlands, 2004.
  76. 76.L. Drăguţ and T. Blaschke, “Automated classification of landform elements using object-based image analysis,” Geomorphology, vol. 81, no. 3, pp. 330-344, 2006.
  77. 77.C. Eisank, L. Drăguţ, and T. Blaschke, "A generic procedure for semantics-oriented landform classification using object-based image analysis," in Geomorphometry, 2011, pp. 125-128.
  78. 78.G. J. Hay, T. Blaschke, D. J. Marceau, and A. Bouchard, “A comparison of three image-object methods for the multiscale analysis of landscape structure,” ISPRS J. Photogramm. Remote Sens., vol. 57, no. 5, pp. 327-345, 2003.
  79. 79.J. S. Walker and T. Blaschke, “Object-based land-cover classification for the Phoenix metropolitan area: optimization vs. transportability,” Int. J. Remote Sens., vol. 29, no. 7, pp. 2021-2040, 2008.
  80. 80.H. Li, H. Gu, Y. Han, and J. Yang, “Object-oriented classification of high-resolution remote sensing imagery based on an improved colour structure code and a support vector machine,” Int. J. Remote Sens., vol. 31, no. 6, pp. 1453-1470, 2010.
  81. 81.B. Fernando, E. Fromont, and T. Tuytelaars, “Mining mid-level features for image classification,” Int. J. Comput. Vis., vol. 108, no. 3, pp. 186-203, 2014.
  82. 82.O. A. Penatti, K. Nogueira, and J. A. dos Santos, "Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit. Workshops, 2015, pp. 44-51.
  83. 83.G.-S. Xia, W. Yang, J. Delon, Y. Gousseau, H. Sun, and H. Maître, "Structural high-resolution satellite image indexing," in ISPRS TC VII Symposium-100 Years ISPRS, 2010, pp. 298-303.
  84. 84.L. Huang, C. Chen, W. Li, and Q. Du, “Remote Sensing Image Scene Classification Using Multi-Scale Completed Local Binary Patterns and Fisher Vectors,” Remote Sensing, vol. 8, no. 6, pp. 483, 2016.
  85. 85.J. Zou, W. Li, C. Chen, and Q. Du, “Scene classification using local and global features with collaborative representation fusion,” Information Sciences, vol. 348, pp. 209-226, 2016.
  86. 86.J. Hu, G.-S. Xia, F. Hu, and L. Zhang, “A comparative study of sampling analysis in the scene classification of optical high-spatial resolution remote sensing imagery,” Remote Sensing, vol. 7, no. 11, pp. 14988-15013, 2015.
  87. 87.F. Hu, G.-S. Xia, Z. Wang, X. Huang, L. Zhang, and H. Sun, “Unsupervised feature learning via spectral clustering of multidimensional patches for remotely sensed scene classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 8, no. 5, pp. 2015-2028, 2015.
  88. 88.L. Xie, J. Wang, B. Zhang, and Q. Tian, “Incorporating visual adjectives for image classification,” Neurocomputing, vol. 182, pp. 48-55, 2016.
  89. 89.L.-J. Zhao, P. Tang, and L.-Z. Huo, “Land-use scene classification using a concentric circle-structured multiscale bag-of-visual-words model,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 7, no. 12, pp. 4620-4631, 2014.
  90. 90.X. Chen, T. Fang, H. Huo, and D. Li, “Measuring the effectiveness of various features for thematic information extraction from very high resolution remote sensing imagery,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 9, pp. 4837-4851, 2015.
  91. 91.S. Chen and Y. Tian, “Pyramid of spatial relatons for scene-level land use classification,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 4, pp. 1947-1957, 2015.
  92. 92.Y. Zhong, Q. Zhu, and L. Zhang, “Scene classification based on the multifeature fusion probabilistic topic model for high spatial resolution remote sensing imagery,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 11, pp. 6207-6222, 2015.
  93. 93.J. Zhang, T. Li, X. Lu, and Z. Cheng, “Semantic Classification of High-Resolution Remote-Sensing Images Based on Mid-level Features,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 9, no. 6, pp. 2343-2353, 2016.
  94. 94.C. Yang, H. Liu, S. Wang, and S. Liao, “Scene-Level Geographic Image Classification Based on a Covariance Descriptor Using Supervised Collaborative Kernel Coding,” Sensors, vol. 16, no. 3, pp. 392, 2016.
  95. 95.V. Risojević and Z. Babić, “Unsupervised Quaternion Feature Learning for Remote Sensing Image Classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 9, no. 4, pp. 1521-1531, 2016.
  96. 96.S. Cui, G. Schwarz, and M. Datcu, “Remote sensing image classification: No features, no clustering,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 8, no. 11, pp. 5158-5170, 2015.
  97. 97.C. Chen, B. Zhang, H. Su, W. Li, and L. Wang, “Land-use scene classification using multi-scale completed local binary patterns,” Signal Image Video Process., vol. 10, no. 4, pp. 745-752, 2016.
  98. 98.C. Chen, L. Zhou, J. Guo, W. Li, H. Su, and F. Guo, "Gabor-filtering-based completed local binary patterns for land-use scene classification," in Proc. IEEE International Conference on Multimedia Big Data, 2015, pp. 324-329.
  99. 99.M. J. Swain and D. H. Ballard, “Color indexing,” Int. J. Comput. Vis., vol. 7, no. 1, pp. 11-32, 1991.
  100. 100.S. Newsam, L. Wang, S. Bhagavathy, and B. S. Manjunath, “Using texture to analyze and manage large collections of remote sensed image and video data,” Applied optics, vol. 43, no. 2, pp. 210-217, 2004.
  101. 101.Y. Yang and S. Newsam, "Comparing SIFT descriptors and Gabor texture features for classification of remote sensed imagery," in Proc. IEEE Int. Conf. Image Process., 2008, pp. 1852-1855.
  102. 102.X. Huang, L. Zhang, and L. Wang, “Evaluation of morphological texture features for mangrove forest mapping and species discrimination using multispectral IKONOS imagery,” IEEE Geosci. Remote Sens. Lett., vol. 6, no. 3, pp. 393-397, 2009.
  103. 103.G. Cheng, J. Han, L. Guo, and T. Liu, "Learning Coarse-to-Fine Sparselets for Efficient Object Detection and Scene Classification," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2015, pp. 1173-1181.
  104. 104.R. M. Haralick, K. Shanmugam, and I. h. Dinstein, “Textural features for image classification,” IEEE Trans. Syst. Man Cybern., vol. 3, pp. 610-621, 1973.
  105. 105.A. K. Jain, N. K. Ratha, and S. Lakshmanan, “Object detection using Gabor filters,” Pattern Recog., vol. 30, no. 2, pp. 295-309, 1997.
  106. 106.T. Ojala, M. Pietikäinen, and T. Mäenpää, “Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 24, no. 7, pp. 971-987, 2002.
  107. 107.A. Oliva and A. Torralba, “Modeling the shape of the scene: A holistic representation of the spatial envelope,” Int. J. Comput. Vis., vol. 42, no. 3, pp. 145-175, 2001.
  108. 108.D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” Int. J. Comput. Vis., vol. 60, no. 2, pp. 91-110, 2004.
  109. 109.N. Dalal and B. Triggs, "Histograms of oriented gradients for human detection," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2005, pp. 886-893.
  110. 110.J. Ren, X. Jiang, and J. Yuan, “Learning LBP structure by maximizing the conditional mutual information,” Pattern Recog., vol. 48, no. 10, pp. 3180-3190, 2015.
  111. 111.M. Douze, H. Jégou, H. Sandhawalia, L. Amsaleg, and C. Schmid, "Evaluation of GIST descriptors for web-scale image search," in Proceedings of the ACM International Conference on Image and Video Retrieval, 2009, pp. 19:1-19:8.
  112. 112.Z. Li and L. Itti, “Saliency and gist features for target detection in satellite images,” IEEE Trans. Image Process., vol. 20, no. 7, pp. 2017-2029, 2011.
  113. 113.J. Yin, H. Li, and X. Jia, “Crater Detection Based on Gist Features,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 8, no. 1, pp. 23-29, 2015.
  114. 114.Y. Ke and R. Sukthankar, "PCA-SIFT: A more distinctive representation for local image descriptors," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2004, pp. 506-513.
  115. 115.H. Bay, T. Tuytelaars, and L. Van Gool, "Surf: Speeded up robust features," in Proc. Eur. Conf. Comput. Vision, 2006, pp. 404-417.
  116. 116.C. Yao and G. Cheng, “Approximative Bayes optimality linear discriminant analysis for Chinese handwriting character recognition,” Neurocomputing, vol. 207, pp. 346-353, 2016.
  117. 117.G. Cheng, J. Han, P. Zhou, and L. Guo, "Scalable multi-class geospatial object detection in high-spatial-resolution remote sensing images," in Proc. IEEE Int. Geosci. Remote Sens. Symp., 2014, pp. 2479-2482.
  118. 118.W. Zhang, X. Sun, K. Fu, C. Wang, and H. Wang, “Object detection in high-resolution remote sensing images using rotation invariant parts based model,” IEEE Geosci. Remote Sens. Lett., vol. 11, no. 1, pp. 74-78, 2014.
  119. 119.Z. Shi, X. Yu, Z. Jiang, and B. Li, “Ship detection in high-resolution optical imagery based on anomaly detector and local shape feature,” IEEE Trans. Geosci. Remote Sens., vol. 52, no. 8, pp. 4511-4523, 2014.
  120. 120.W. Zhang, X. Sun, H. Wang, and K. Fu, “A generic discriminative part-based model for geospatial object detection in optical remote sensing images,” ISPRS J. Photogramm. Remote Sens., vol. 99, pp. 30-44, 2015.
  121. 121.G. Cheng, P. Zhou, X. Yao, C. Yao, Y. Zhang, and J. Han, "Object detection in VHR optical remote sensing images via learning rotation-invariant HOG feature," in Proc. Int. Workshop Earth Observ. Remote Sens. Appl., 2016, pp. 433-436.
  122. 122.L. Zhao, P. Tang, and L. Huo, “A 2-D wavelet decomposition-based bag-of-visual-words model for land-use scene classification,” Int. J. Remote Sens., vol. 35, no. 6, pp. 2296-2310, 2014.
  123. 123.R. Bahmanyar, S. Cui, and M. Datcu, “A comparative study of bag-of-words and bag-of-topics models of eo image patches,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 6, pp. 1357-1361, 2015.
  124. 124.S. Lazebnik, C. Schmid, and J. Ponce, "Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2006, pp. 2169-2178.
  125. 125.L. Zhang, L. Zhang, D. Tao, and X. Huang, “On combining multiple features for hyperspectral remote sensing image classification,” IEEE Trans. Geosci. Remote Sens., vol. 50, no. 3, pp. 879-893, 2012.
  126. 126.S. Chaib, Y. Gu, and H. Yao, “An informative feature selection method based on sparse PCA for VHR scene classification,” IEEE Geosci. Remote Sens. Lett., vol. 13, no. 2, pp. 147-151, 2016.
  127. 127.Y. Li, C. Tao, Y. Tan, K. Shang, and J. Tian, “Unsupervised multilayer feature learning for satellite image scene classification,” IEEE Geosci. Remote Sens. Lett., vol. 13, no. 2, pp. 157-161, 2016.
  128. 128.K. Qi, X. Zhang, B. Wu, and H. Wu, “Sparse coding-based correlaton model for land-use scene classification in high-resolution remote-sensing images,” J. Appl. Remote Sens., vol. 10, no. 4, pp. 042005-042005, 2016.
  129. 129.F. Zhang, B. Du, and L. Zhang, “Saliency-guided unsupervised feature learning for scene classification,” IEEE Trans. Geosci. Remote Sens., vol. 53, no. 4, pp. 2175-2184, 2015.
  130. 130.A. Romero, C. Gatta, and G. Camps-Valls, “Unsupervised deep feature extraction for remote sensing image classification,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 3, pp. 1349-1362, 2016.
  131. 131.F. Hu, G.-S. Xia, J. Hu, Y. Zhong, and K. Xu, “Fast binary coding for the scene classification of high-resolution remote sensing imagery,” Remote Sensing, vol. 8, no. 7, pp. 555, 2016.
  132. 132.W. Yao, O. Loffeld, and M. Datcu, “Application and Evaluation of a Hierarchical Patch Clustering Method for Remote Sensing Images,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 9, no. 6, pp. 2279-2289, 2016.
  133. 133.E. Othman, Y. Bazi, N. Alajlan, H. Alhichri, and F. Melgani, “Using convolutional features and a sparse autoencoder for land-use scene classification,” Int. J. Remote Sens., vol. 37, no. 10, pp. 2149-2167, 2016.
  134. 134.B. Du, W. Xiong, J. Wu, L. Zhang, L. Zhang, and D. Tao, “Stacked Convolutional Denoising Auto-Encoders for Feature Representation,” IEEE Trans. Cybern., DOI: 10.1109/TCYB.2016.2536638, 2016.
  135. 135.I. Jolliffe, Principal component analysis, New York, NY, USA: Springer, 2002.
  136. 136.B. A. Olshausen and D. J. Field, “Sparse coding with an overcomplete basis set: A strategy employed by V1?,” Vision research, vol. 37, no. 23, pp. 3311-3325, 1997.
  137. 137.G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science, vol. 313, no. 5786, pp. 504-507, 2006.
  138. 138.T.-H. Chan, K. Jia, S. Gao, J. Lu, Z. Zeng, and Y. Ma, “PCANet: A simple deep learning baseline for image classification?,” IEEE Trans. Image Process., vol. 24, no. 12, pp. 5017-5032, 2015.
  139. 139.Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 8, pp. 1798-1828, 2013.
  140. 140.B. Zhao, Y. Zhong, and L. Zhang, “A spectral–structural bag-of-features scene classifier for very high spatial resolution remote sensing imagery,” ISPRS J. Photogramm. Remote Sens., vol. 116, pp. 73-85, 2016.
  141. 141.G. Cheng and J. Han, “A survey on object detection in optical remote sensing images,” ISPRS J. Photogramm. Remote Sens., vol. 117, pp. 11-28, 2016.
  142. 142.J. Wang, J. Yang, K. Yu, F. Lv, T. Huang, and Y. Gong, "Locality-constrained linear coding for image classification," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2010, pp. 3360-3367.
  143. 143.F. Hu, G.-S. Xia, J. Hu, and L. Zhang, “Transferring deep convolutional neural networks for the scene classification of high-resolution remote sensing imagery,” Remote Sensing, vol. 7, no. 11, pp. 14680-14707, 2015.
  144. 144.F. Zhang, B. Du, and L. Zhang, “Scene classification via a gradient boosting random convolutional network framework,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 3, pp. 1793-1802, 2016.
  145. 145.I. Ševo and A. Avramović, “Convolutional Neural Network Based Automatic Object Detection on Aerial Images,” IEEE Geosci. Remote Sens. Lett., vol. 13, no. 5, pp. 740-744, 2016.
  146. 146.W. Zhao and S. Du, “Scene classification using multi-scale deeply described visual words,” Int. J. Remote Sens., vol. 37, no. 17, pp. 4119-4131, 2016.
  147. 147.F. Luus, B. Salmon, F. Van Den Bergh, and B. Maharaj, “Multiview deep learning for land-use classification,” IEEE Geosci. Remote Sens. Lett., vol. 12, no. 12, pp. 2448-2452, 2015.
  148. 148.Y. Zhong, F. Fei, and L. Zhang, “Large patch convolutional neural networks for the scene classification of high spatial resolution imagery,” J. Appl. Remote Sens., vol. 10, no. 2, pp. 025006-025006, 2016.
  149. 149.D. Marmanis, M. Datcu, T. Esch, and U. Stilla, “Deep Learning Earth Observation Classification Using ImageNet Pretrained Networks,” IEEE Geosci. Remote Sens. Lett., vol. 13, no. 1, pp. 105-109, 2016.
  150. 150.M. Längkvist, A. Kiselev, M. Alirezaie, and A. Loutfi, “Classification and Segmentation of Satellite Orthoimagery Using Convolutional Neural Networks,” Remote Sensing, vol. 8, no. 4, pp. 329, 2016.
  151. 151.L. Zhang, L. Zhang, and B. Du, “Deep Learning for Remote Sensing Data: A Technical Tutorial on the State of the Art,” IEEE Geoscience and Remote Sensing Magazine, vol. 4, no. 2, pp. 22-40, 2016.
  152. 152.M. Castelluccio, G. Poggi, C. Sansone, and L. Verdoliva, “Land use classification in remote sensing images by convolutional neural networks,” arXiv preprint arXiv:1508.00092, 2015.
  153. 153.K. Nogueira, O. A. Penatti, and J. A. d. Santos, “Towards Better Exploiting Convolutional Neural Networks for Remote Sensing Scene Classification,” arXiv preprint arXiv:1602.01517, 2016.
  154. 154.G. Cheng, P. Zhou, and J. Han, "RIFD-CNN: Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2016, pp. 2884-2893.
  155. 155.G. Cheng, C. Ma, P. Zhou, X. Yao, and J. Han, "Scene classification of high resolution remote sensing images using convolutional neural networks," in Proc. IEEE Int. Geosci. Remote Sens. Symposium, 2016, pp. 767-770.
  156. 156.X. Yao, J. Han, G. Cheng, and L. Guo, "Semantic segmentation based on stacked discriminative autoencoders and context-constrained weakly supervised learning," in Proc. ACM Int. Conf. Multimedia, 2015, pp. 1211-1214.
  157. 157.D. Zhang, J. Han, C. Li, J. Wang, and X. Li, “Detection of co-salient objects by looking deep and wide,” Int. J. Comput. Vis., vol. 120, no. 2, pp. 215-232, 2016.
  158. 158.D. Zhang, J. Han, L. Jiang, S. Ye, and X. Chang, “Revealing event saliency in unconstrained video collection,” IEEE Trans. Image Process., vol. 26, no. 4, pp. 1746-1758, 2017.
  159. 159.Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436-444, 2015.
  160. 160.G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,” Neural Comput., vol. 18, no. 7, pp. 1527-1554, 2006.
  161. 161.R. Salakhutdinov and G. Hinton, “An efficient learning procedure for deep Boltzmann machines,” Neural Comput., vol. 24, no. 8, pp. 1967-2006, 2012.
  162. 162.P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J. Mach. Learn. Res., vol. 11, pp. 3371-3408, 2010.
  163. 163.A. Krizhevsky, I. Sutskever, and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," in Proc. Conf. Adv. Neural Inform. Process. Syst., 2012, pp. 1097-1105.
  164. 164.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, "Overfeat: Integrated recognition, localization and detection using convolutional networks," in Proc. Int. Conf. Learn. Represent., 2014, pp. 1-16.
  165. 165.K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," in Proc. Int. Conf. Learn. Represent., 2015, pp. 1-13.
  166. 166.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2015, pp. 1-9.
  167. 167.K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 9, pp. 1904-1916, 2015.
  168. 168.Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, "Greedy layer-wise training of deep networks," in Proc. Conf. Adv. Neural Inform. Process. Syst., 2007, pp. 153-160.
  169. 169.J. Han, D. Zhang, X. Hu, and L. Guo, “Background prior-based salient object detection via deep reconstruction residual,” IEEE Trans. Circuits Syst. Video Technol., vol. 25, no. 8, pp. 1309-1321, 2015.
  170. 170.J. Han, D. Zhang, S. Wen, L. Guo, T. Liu, and X. Li, “Two-stage learning to predict human eye fixations via SDAEs,” IEEE Trans. Cybern., vol. 46, no. 2, pp. 487-498, 2016.
  171. 171.D. Zhang, J. Han, and L. Shao, “Cosaliency detection based on intrasaliency prior transfer and deep intersaliency mining,” IEEE Trans. Neural Netw. Learn. Syst., vol. 27, no. 6, pp. 1163-1176, 2016.
  172. 172.K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2016, pp. 770-778.
  173. 173.F. Zhang, B. Du, L. Zhang, and M. Xu, “Weakly Supervised Learning Based on Coupled Convolutional Neural Networks for Aircraft Detection,” IEEE Trans. Geosci. Remote Sens., vol. 54, no. 9, pp. 5553-5563, 2016.
  174. 174.F. F. Li and P. Perona, "A bayesian hierarchical model for learning natural scene categories," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2005, pp. 524-531.
  175. 175.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "Imagenet: A large-scale hierarchical image database," in Proc. IEEE Int. Conf. Comput. Vision Pattern Recognit., 2009, pp. 248-255.
  176. 176.C.-C. Chang and C.-J. Lin, “LIBSVM: a library for support vector machines,” ACM Trans. Intell. Syst. Technol., vol. 2, no. 3, pp. 27, 2011.

Citation

MLA
Cheng, G., et al. “Remote Sensing Image Scene Classification: Benchmark and State of the Art”. Proceedings of the IEEE, vol. 105, no. 10, 2017, pp. 1865–83, https://doi.org/10.1109/JPROC.2017.2675998.
APA
Cheng, G., Han, J., & Lu, X. (2017). Remote Sensing Image Scene Classification: Benchmark and State of the Art. Proceedings of the IEEE, 105(10), 1865–1883. https://doi.org/10.1109/JPROC.2017.2675998
Chicago
Cheng, G., J. Han, and X. Lu. 2017. “Remote Sensing Image Scene Classification: Benchmark and State of the Art”. Proceedings of the IEEE 105 (10): 1865–83. https://doi.org/10.1109/JPROC.2017.2675998.
Harvard
Cheng, G., Han, J. and Lu, X. (2017) “Remote Sensing Image Scene Classification: Benchmark and State of the Art”, Proceedings of the IEEE, 105(10), pp. 1865–1883. Available at: https://doi.org/10.1109/JPROC.2017.2675998.
Vancouver
1. Cheng G, Han J, Lu X (2017) Remote Sensing Image Scene Classification: Benchmark and State of the Art. Proceedings of the IEEE 105:1865–1883

BibTeX

@article{Cheng_2017, title={Remote Sensing Image Scene Classification: Benchmark and State of the Art}, volume={105}, ISSN={1558-2256}, url={http://dx.doi.org/10.1109/JPROC.2017.2675998}, DOI={10.1109/jproc.2017.2675998}, number={10}, journal={Proceedings of the IEEE}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Cheng, Gong and Han, Junwei and Lu, Xiaoqiang}, year={2017}, month=Oct, pages={1865–1883} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF