Deep learning in remote sensing: a review

Xiao Xiang ZhuDevis TuiaLichao MouGui-Song XiaLiangpei ZhangFeng XuFriedrich Fraundorfer

article2017arXiv1,954 citations

Surveys key advances and open challenges in applying deep learning to remote sensing data, providing practical resources and strategies to integrate Earth observation domain knowledge for tackling large-scale environmental problems.

Listen

The rapid growth of Earth observation data has created a critical need for advanced automated analysis. Modern satellite constellations produce petabytes of imagery characterized by complex physical properties, multi-sensor modalities, exact spatial coordinates, and dense time series. Traditional remote sensing techniques rely heavily on manual feature engineering and domain-specific heuristics, which struggle to scale with massive volumes of data or capture intricate, nonlinear patterns. The article provides a comprehensive evaluation of how modern deep learning architectures address these challenges across major Earth observation applications, while also cataloging open-source software frameworks and public benchmark datasets to facilitate research adoption.

To conduct this evaluation, the article synthesizes experimental findings, network architectures, and benchmark evaluations across five core subfields: hyperspectral image analysis, synthetic aperture radar interpretation, high-resolution optical image processing, multimodal data fusion, and 3D reconstruction. It reviews foundational deep learning models—including autoencoders, deep belief networks, convolutional neural networks, and recurrent neural networks—and evaluates their performance against classical baselines on standard international benchmarks, such as Middlebury stereo evaluation datasets and MSTAR radar collections.

The findings show that deep learning consistently outperforms traditional hand-crafted methods. In stereo matching for 3D reconstruction, convolutional neural network approaches cut bad pixel error rates from an 18.4% baseline under traditional semi-global matching down to between 5.9% and 8.1%. In radar automatic target recognition, specialized convolutional networks achieve up to 99.1% accuracy under standard operating conditions when paired with data augmentation. For optical and hyperspectral data, deep networks demonstrate superior capacity for end-to-end pixel-level classification, large-scale scene categorization, and automated feature extraction from unlabeled data. Furthermore, deep learning facilitates complex multimodal data fusion, enabling unified workflows for simultaneous image registration, land cover classification, and multi-decadal change detection.

These results demonstrate that deep neural networks offer significant operational gains, allowing automated systems to process global-scale satellite streams with higher precision and lower manual engineering overhead. However, the article emphasizes that purely data-driven, black-box approaches must not completely replace physical and domain-specific knowledge. Instead, integrating physical sensor characteristics, prior geographical constraints, and domain expertise into network architectures is essential to avoid severe errors when processing complex geospatial physics.

Decision-makers and research teams should focus future efforts on combining physics-based process models with deep learning architectures, developing weakly supervised and unsupervised methods to mitigate the scarcity of labeled Earth observation data, and establishing global transferability pipelines. While current benchmarks demonstrate high confidence for localized and well-annotated tasks, users should exercise caution when deploying models globally, as performance can degrade across unseen geographic regions, varying atmospheric conditions, and distinct sensor geometries.

Cover for Deep learning in remote sensing: a review

Abstract

Standing at the paradigm shift towards data-intensive science, machine learning techniques are becoming increasingly important. In particular, as a major breakthrough in the field, deep learning has proven as an extremely powerful tool in many fields. Shall we embrace deep learning as the key to all? Or, should we resist a 'black-box' solution? There are controversial opinions in the remote sensing community. In this article, we analyze the challenges of using deep learning for remote sensing data analysis, review the recent advances, and provide resources to make deep learning in remote sensing ridiculously simple to start with. More importantly, we advocate remote sensing scientists to bring their expertise into deep learning, and use it as an implicit general model to tackle unprecedented large-scale influential challenges, such as climate change and urbanization.

Table of Contents

  • I Motivation
  • II From Perceptron to Deep Learning
  • II-A Autoencoder models
  • II-A1 Autoencoder and Stacked Autoencoder (SAE)
  • II-A2 Sparse Autoencoder
  • II-A3 Restricted Boltzmann Machine (RBM) & Deep Belief Network (DBN)
  • II-B Convolutional neural networks (CNNs).
  • II-B1 AlexNet
  • II-B2 VGG Net
  • II-B3 ResNet
  • II-B4 FCN
  • III Remote Sensing Meets Deep Learning
  • III-A Hyperspectral Image Analysis
  • III-A1 Hyperspectral Image Classification
  • III-A2 Anomaly Detection
  • III-B Interpretation of SAR Images
  • III-B1 Automatic Target Recognition
  • III-B2 Terrain surface classification
  • III-B3 Parameter inversion
  • III-C Interpretation of High-resolution Satellite Images
  • III-C1 Scene Classification
  • III-C2 Object Detection
  • III-C3 Image Retrieval
  • III-D Multimodal Data Fusion
  • III-D1 Pansharpening and Super-Resolution
  • III-D2 Feature and Decision-level Fusion for image classification
  • III-D3 Fusing Heterogeneous Sources
  • III-E 3D Reconstruction
  • III-E1 Tie points identification and matching
  • III-E2 Stereo processing using convolutional neural networks
  • III-E3 Large scale semantic 3D city reconstruction
  • IV Deep Learning in Remote Sensing Made Ridiculously Simple to Start with
  • IV-A Tutorials
  • IV-B Open-source Deep Learning Frameworks
  • IV-C Remote Sensing Data for Training Deep Learning Models
  • IV-C1 Scene classification (one image is classified into a single label)
  • IV-C2 Image classification (each pixel of an image is classified into a label)
  • IV-C3 Registration / matching
  • IV-D Showcasing
  • V Conclusion and Future Trends
  • References

Knowls

  1. Knowl 1 — Core Characteristics and Methodological Challenges of Remote Sensing Data for Deep Learning

    definition

    Remote sensing data introduce distinct physical, geometric, and operational characteristics that differentiate Earth observation from conventional computer vision tasks:

    • Multi-modality: Observations originate from fundamentally distinct sensor types (e.g., optical multispectral, hyperspectral, Synthetic Aperture Radar (SAR), and LiDAR) with distinct imaging geometries and physical measurement principles, requiring multimodal fusion architectures and cross-modality transferability.
    • Geographic location and spatial registration: Pixels correspond directly to real-world spatial coordinates, enabling joint fusion with geographic information systems (GIS), social media metadata, and heterogeneous ground/aerial sensors.
    • Geodetic measurement quality and physical priors: Remote sensing acquisitions represent geodetic measurements with calibrated quality, requiring the integration of physical sensor models (e.g., single-pass vs. repeat-pass SAR interferometry) rather than purely data-driven black-box modeling.
    • Continuous temporal dimension: Satellite constellations (e.g., the Copernicus Sentinel program) provide continuous Earth coverage with short revisit times, shifting analysis from single-image interpretation to joint spatio-temporal and spectral-temporal sequence processing.
    • Global scale and high data volumes: Processing global Petabyte-scale archives demands scalable, highly transferable models capable of generalizing across varied planetary landscapes with semi-automated annotation pipelines.
    • Retrieval of continuous geophysical and biochemical quantities: Many remote sensing objectives target continuous parameters (e.g., surface deformation rates, soil mineral composition, vegetation biomass, atmospheric trace gas concentrations) rather than discrete object categories, requiring neural network emulators that approximate physical process models.
  2. Knowl 2 — Deep Learning Architectures for Hyperspectral Image Analysis

    model/method

    Hyperspectral sensors measure hundreds of narrow, contiguous spectral bands, facilitating spectroscopic identification of surface materials at the cost of high spectral dimensionality and nonlinear light scattering. Deep learning models for hyperspectral data address these characteristics through distinct spatial and spectral processing strategies:

    • 1D Spectral CNNs: 1D convolutional networks treat the spectral signature of an individual pixel as a 1D vector, learning 1D convolutional filters across spectral channels to extract discriminative spectral absorption features.
    • 2D Spatial CNNs: 2D convolutional networks extract spatial structural context from hyperspectral scenes, frequently combined with dimension reduction algorithms (such as local discriminant embedding) or band selection strategies (such as particle swarm optimization) to mitigate the curse of dimensionality.
    • 3D Spatio-Spectral CNNs: 3D convolutional networks apply 3D convolution kernels across the two spatial dimensions and the third spectral dimension simultaneously (x×y×λx \times y \times \lambda), jointly modeling spatial relationships and continuous spectral patterns without collapsing spectral correlation.
    • Unsupervised Residual Conv-Deconv Networks: Fully convolutional residual encoder-decoder networks reconstruct hyperspectral cubes in an unsupervised manner, enabling robust spatio-spectral feature representation learning from unannotated imagery.
    • Recurrent Neural Networks (RNNs): Sequential models (e.g., LSTMs or modified gated recurrent units) treat pixel spectra as sequential data ordered by wavelength, exploiting inter-band dependencies for land cover classification.
    • Pairwise Difference CNNs for Anomaly Detection: Convolutional networks trained on spectral difference vectors between neighboring background pixel pairs detect anomalous spectral targets that deviate significantly from learned local background distributions.
  3. Knowl 3 — Deep Architectures for Synthetic Aperture Radar (SAR) Target Recognition and Polarimetric SAR Classification

    model/method

    Synthetic Aperture Radar (SAR) and Polarimetric SAR (PolSAR) imagery present complex-valued scattering data, coherent speckle noise, and non-optical imaging geometries, requiring specialized deep network architectures:

    • All-Convolutional Networks (AConvNets) for SAR ATR: Standard CNNs trained on SAR Automatic Target Recognition (ATR) datasets suffer from severe overfitting due to scarce target samples. AConvNets remove all dense fully connected layers and rely exclusively on convolutional feature maps for classification, substantially reducing trainable parameters and improving robustness across extended operating conditions.
    • Domain-Specific SAR Data Augmentation: To prevent overfitting in SAR ATR, models employ specialized data augmentation including elastic distortions, affine transformations, and synthetic radar simulation software to reproduce depression angle shifts and aspect angle inaccuracies.
    • Complex-Valued Convolutional Neural Networks (CV-CNN): PolSAR data are naturally formatted as complex-valued coherency or covariance matrices with non-zero phase information in off-diagonal terms. CV-CNNs utilize complex inputs, complex convolution weights, and complex backpropagation rules across all layers to preserve both phase differences and power amplitudes between polarization channels.
    • Hybrid Generative and Autoencoder Pipelines for PolSAR: Unsupervised models combine hand-crafted initial layers (e.g., Gray-Level Co-occurrence Matrices, Gabor filters, or HOG descriptors) with stacked autoencoders (SAE) or Deep Belief Networks (DBN), often combined with superpixel segmentation or Markov Random Fields (MRF) to preserve terrain spatial boundaries.
  4. Knowl 4 — Methodological Strategies for High-Resolution Optical Scene Classification

    model/method

    High-resolution satellite scene classification assigns a categorical semantic label to an entire scene image composed of heterogeneous spatial arrangements of objects. Deep learning methods for scene classification fall into three core methodological categories:

    1. Pre-trained Network Feature Extraction: Deep CNNs pre-trained on large natural image corpora (such as ImageNet) serve as off-the-shelf feature extractors. Intermediate or fully connected activation vectors are extracted and used either directly in linear/SVM classifiers or encoded via mid-level pooling methods (e.g., Bag-of-Visual-Words or Vector of Locally Aggregated Descriptors).
    2. Fine-Tuning on Domain Datasets: High-level layers of pre-trained networks are fine-tuned on labeled satellite scene datasets. This adapts the generic semantic representations to remote sensing specific spatial patterns and spectral characteristics while preventing overfitting on relatively small target datasets.
    3. Training Specialized Networks from Scratch: Smaller, custom CNN architectures or multi-scale networks are trained directly from random initialization using remote sensing imagery. While avoiding domain shifts from natural images, these models require sufficient labeled training instances or specialized regularizations (such as gradient boosting or random convolutional ensembles) to prevent memorization and maintain generalization.
  5. Knowl 5 — Rotation-Invariant and Scale-Adaptive Object Detection in Very High-Resolution Remote Sensing Imagery

    model/method

    Object detection in high-resolution overhead imagery addresses arbitrary object orientations, large scale variations, and dense target distributions through specialized network modifications:

    • Rotation-Invariant CNNs (RICNN): Overhead objects (e.g., aircraft, vehicles, storage tanks) appear at arbitrary orientations. RICNNs introduce a rotation-invariant layer and an explicit regularization term into the objective function that penalizes feature differences between original and rotated training patches:

    Ltotal=Lcls+λrotLreg(f(x),f(Rθ(x)))\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{cls}} + \lambda_{\text{rot}} \mathcal{L}_{\text{reg}}(\mathbf{f}(x), \mathbf{f}(R_\theta(x)))

    where Lcls\mathcal{L}_{\text{cls}} is the classification loss, f(x)\mathbf{f}(x) denotes the extracted feature representation for patch xx, Rθ(x)R_\theta(x) is the image patch rotated by angle θ\theta, and λrot\lambda_{\text{rot}} controls the strength of the rotation invariance constraint.

    • Multi-Scale and Variable Receptive Field Pooling: Feature maps from final convolutional and pooling layers are divided into multiple sub-blocks with variable receptive field and pooling window sizes to detect objects exhibiting severe scale variance.
    • Weakly Supervised Region Detection: Coupled CNN architectures combine candidate region proposal networks with localization networks trained exclusively with image-level presence/absence labels, utilizing iterative negative bootstrapping to reduce manual bounding box annotation requirements.
  6. Knowl 6 — Deep Multimodal and Multi-Source Remote Sensing Data Fusion Frameworks

    model/method

    Data fusion in remote sensing integrates heterogeneous observation sources across four main deep learning paradigms:

    • Feature-Level Fusion via Multi-Channel Stacking: Stacking multispectral bands, Digital Surface Models (DSMs), and normalized DSMs into a unified multi-channel tensor input for fully convolutional (FCN) or deconvolutional segmentation networks.
    • Multi-Stream Intermediate Fusion: Parallel convolutional streams process separate modalities (e.g., PolSAR and hyperspectral data, or color imagery and DSMs) independently to extract sensor-specific features, which are then concatenated and processed by joint fully connected or convolutional layers at deeper stages.
    • Learned Decision-Level Fusion: Rather than heuristic probability averaging, independent models output class predictions which are fused by a dedicated residual fusion network that learns optimal weighting coefficients to correct classification maps.
    • Multi-Task Joint Output Fusion: CNNs are trained end-to-end to simultaneously optimize complementary predictive tasks, such as combining dense semantic segmentation with boundary/edge likelihood estimation, or performing joint semantic labeling and monocular depth/height regression.
    • Cross-View Heterogeneous Matching: Siamese deep networks align overhead aerial/satellite imagery with ground-level imagery (e.g., street-level panoramas and geotagged social media photos) by mapping both views into a common embedding space for ground-to-aerial geolocalization and urban change detection.
  7. Knowl 7 — Deep Learning in Photogrammetric 3D Reconstruction and Semantic Modeling

    model/method

    Deep neural networks enhance classical geometric 3D reconstruction pipelines across three core stages:

    • Learned Keypoint Descriptors for Tie-Point Matching: Convolutional neural networks trained on patch pairs replace handcrafted local feature descriptors (e.g., SIFT, SURF) to compute robust similarity metrics under drastic illumination and perspective changes, reducing false-match rates in camera orientation and resectioning.
    • End-to-End CNN-Based Stereo Disparity Estimation: 2D and 3D encoder-decoder CNN architectures take rectified stereo image pairs as direct input and output dense disparity/depth maps, replacing multi-stage handcrafted cost volume aggregation, disparity computation, and heuristic outlier filtering.
    • Joint Semantic Volumetric 3D Reconstruction: Rather than treating 3D reconstruction and semantic labeling as disconnected sequential steps, volumetric models partition the 3D space into voxel grids. Voxel occupancy and semantic class categories (e.g., building, roof, ground, vegetation, sky) are estimated jointly using CNN-derived image semantic priors and 3D spatial continuity constraints.
  8. Knowl 8 — Error Rates of CNN-Based Stereo Matching Methods on Middlebury Benchmark

    data/table

    Stereo matching benchmarks quantify the accuracy of computing dense depth/disparity correspondences from stereo image pairs. On the Middlebury Stereo Evaluation Benchmark (evaluated as of May 2017), deep convolutional neural network approaches consistently outperform classical algorithms such as Semi-Global Matching (SGM).

    Method Bad Pixel Error Rate (%)
    3DMST 5.92
    MC-CNN+TDSR 6.35
    LW-CNN 7.04
    MC-CNN-acrt 8.08
    SGM (Classical Baseline) 18.4

    The benchmark results demonstrate that deep learning-based similarity matching (such as MC-CNN and its extensions) reduces the bad pixel error rate from 18.4% down to 5.92% compared to standard engineered cost functions.

  9. Knowl 9 — Public Benchmark Datasets for Deep Learning in Earth Observation

    experimental setup

    Training and evaluating deep learning models in remote sensing relies on standardized public benchmark datasets categorized by target application:

    • Scene Classification (Single label per scene):
      • UC Merced: 2,100 aerial RGB images of size 256×256256 \times 256 pixels across 21 land-use categories (100 images per class).
      • AID (Aerial Image Dataset): 10,000 annotated aerial images across 30 land-use scene classes with varied spatial resolutions.
      • NWPU-RESISC45: 31,500 aerial images spanning 45 scene classes (700 images per class, 256×256256 \times 256 pixels).
    • Dense Semantic Labeling / Pixel-Wise Classification:
      • Zurich Summer Dataset: 20 pansharpened QuickBird satellite image chips (0.6 m0.6\text{ m} spatial resolution) annotated across 8 land cover classes.
      • Zeebruges (IEEE GRSS Data Fusion Contest 2015): Seven tiles (10,000×10,00010{,}000 \times 10{,}000 pixels) combining 5 cm5\text{ cm} RGB aerial imagery and dense LiDAR point clouds (65 pts/m265\text{ pts/m}^2) labeled into 8 semantic categories.
      • ISPRS 2D Semantic Labeling (Vaihingen and Potsdam): Sub-decimeter True Orthophoto tiles (Vaihingen: 33 tiles, 9 cm9\text{ cm} GSD, CIR + DSM; Potsdam: 38 tiles, 5 cm5\text{ cm} GSD, 4-band IRRGB + DSM) annotated into 6 classes (impervious surfaces, building, low vegetation, tree, car, clutter).
    • Cross-Modal SAR-Optical Image Matching:
      • SARptical: 10,000 co-registered high-resolution TerraSAR-X spotlight SAR and UltraCam optical image patch pairs over dense urban scenes in Berlin, co-registered using 3D InSAR and photogrammetric point clouds.
  10. Knowl 10 — Hybrid Physics-Informed Modeling and Generalization Limits for Global Earth Observation

    limitation

    Deploying deep neural networks for large-scale and global Earth observation faces core physical and generalization constraints:

    • Physical Inconsistency of Purely Data-Driven Models: Remote sensing measurements result from physical radiative transfer and microwave scattering processes. Purely empirical, data-driven networks can generate unphysical predictions when exposed to varying atmospheric scattering, seasonal vegetation changes, or varied illumination geometries. Embedding physical forward models or using neural networks as emulators for physical processes is necessary for reliable geo-parameter inversion.
    • Domain Shift Across Geographic Scales: Models trained on localized datasets struggle to generalize globally due to intra-class variability, atmospheric conditions, and cultural/regional differences in built environments and land management.
    • High Cost of Expert Ground-Truth Annotations: Unlike natural image domains where crowdsourcing is straightforward, geoscientific labeling (such as ground displacement rates, soil mineral composition, or detailed crop types) requires expert in-situ field surveys and geodetic measurements, making unsupervised, self-supervised, and transfer learning methods essential.

Coverage note — Standard deep learning textbook background (e.g., standard definitions of Perceptrons, basic Autoencoders, RBMs, AlexNet, VGG, ResNet, and FCN from computer vision literature) was omitted as it represents generic prior work rather than remote sensing specific contributions.

References

  1. 1.MIT Technology Review, 2013 [Online]. Available: https://www.technologyreview.com/lists/technologies/2013/.
  2. 2.A. Krizhevsky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS), 2012.
  3. 3.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in IEEE International Conference on Learning Representation (ICLR), 2015.
  4. 4.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  5. 5.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Region-based convolutional networks for accurate object detection and semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, no. 1, pp. 142–158, 2016.
  6. 6.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  7. 7.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  8. 8.H. Noh, S. Hong, and B. Han, “Learning deconvolutional network for semantic segmentation,” in IEEE International Conference on Computer Vision (ICCV), 2015.
  9. 9.J. Donahue, L. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  10. 10.Y. Du, W. Wang, and L. Wang, “Hierarchical recurrent neural network for skeleton based action recognition,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  11. 11.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in IEEE International Conference on Machine Learning (ICML), 2015.
  12. 12.J. P. Rivera, J. Verrelst, J. Gomez-Dans, J. Munoz-Mar´ı, J. Moreno, and G. Camps-Valls, “An emulator toolbox to approximate radiative transfer models with statitistical learning,” Remote Sensing, vol. 7, pp. 9347–9370, 2015.
  13. 13.R. Hecht-Nielsen, “Theory of the backpropagation neural network,” in International Joint Conference on Neural Networks (IJCNN), 1989.
  14. 14.P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” Journal of Machine Learning Research, vol. 11, pp. 3371–3408, 2010.
  15. 15.A. Ng, “Sparse autoencoder,” https://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf, online.
  16. 16.G. Hinton, S. Osindero, and Y. Teh, “A fast learning algorithm for deep belief nets,” Neural Computation, vol. 18, no. 7, pp. 1527–1554, 2006.
  17. 17.I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016, http://www.deeplearningbook.org.
  18. 18.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  19. 19.Y. Chen, H. Jiang, C. Li, X. Jia, and P. Ghamisi, “Deep feature extraction and classification of hyperspectral images based on convolutional neural networks,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 10, pp. 6232–6251, 2016.
  20. 20.G. Camps-Valls, D. Tuia, L. Bruzzone, and J. A. Benediktsson, “Advances in hyperspectral image classification,” IEEE Signal Proc. Mag., vol. 31, pp. 45–54, 2014.
  21. 21.P. Ghamisi, Y. Chen, and X. Zhu, “A self-improving convolution neural network for the classification of hyperspectral data,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 10, pp. 1537–1541, 2016.
  22. 22.Y. Chen, Z. Lin, X. Zhao, G. Wang, and Y. Gu, “Deep learning-based classification of hyperspectral data,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 7, no. 6, pp. 2094–2107, 2014.
  23. 23.Y. Chen, X. Zhao, and X. Jia, “Spectra-spatial classification of hyperspectral data based on deep belief network,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 8, no. 6, pp. 2381–2392, 2015.
  24. 24.C. Tao, H. Pan, Y. Li, and Z. Zou, “Unsupervised spectral-spatial feature learning with stacked sparse autoencoder for hyperspectral imagery classification,” IEEE Geoscience and Remote Sensing Letters, vol. 8, no. 6, pp. 2381–2392, 2015.
  25. 25.W. Hu, Y. Huang, L. Wei, F. Zhang, and H. Li, “Deep convolutional neural networks for hyperspectral image classification,” Journal of Sensors, vol. 2015, no. 258619, 2015.
  26. 26.K. Makantasis, K. Karantzalos, A. Doulamis, and N. Doulamis, “Deep supervised learning for hyperspectral data classification through convolutional neural networks,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2015.
  27. 27.N. Kussul, M. Lavreniuk, S. Skakun, and A. Shelestov, “Deep learning classification of land cover and crop types using remote sensing data,” IEEE Geoscience and Remote Sensing Letters, DOI:10.1109/LGRS.2017.2681128.
  28. 28.W. Zhao and S. Du, “Spectral-spatial feature extraction for hyperspectral image classification: A dimension reduction and deep learning approach,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 8, pp. 4544–4554, 2016.
  29. 29.A. Santara, K. Mani, P. Hatwar, A. Singh, A. Garg, K. Padia, and P. Mitra, “Bass net: Band-adaptive spectral-spatial feature learning neural network for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 9, pp. 5293–5301, 2017.
  30. 30.W. Li, G. Wu, F. Zhang, and Q. D. and, “Hyperspectral image classification using deep pixel-pair features,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 2, pp. 844–853, 2017.
  31. 31.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in IEEE International Conference on Computer Vision (ICCV), 2015.
  32. 32.Y. Li, H. Zhang, and Q. Shen, “Spectral-spatial classification of hyperspectral imagery with 3d convolutional neural network,” Remote Sensing, vol. 18, no. 7, pp. 1527–1554, 2006.
  33. 33.L. Mou, P. Ghamisi, and X. X. Zhu, “Fully conv-deconv network for unsupervised spectral-spatial feature extraction of hyperspectral imagery via residual learning,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2017.
  34. 34.L. Mou, P. Ghamisi, and X. Zhu, “Unsupervised spectral-spatial feature learning via deep residual conv-deconv network for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, in press.
  35. 35.A. Romero, C. Gatta, and G. Camps-Valls, “Unsupervised deep feature extraction for remote sensing image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 3, pp. 1349–1362, 2016.
  36. 36.L. Mou, P. Ghamisi, and X. Zhu, “Deep recurrent neural networks for hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 7, pp. 3639–3655, 2017.
  37. 37.W. Li, G. Wu, and Q. Du, “Transferred deep learning for anomaly detection in hyperspectral imagery,” IEEE Geoscience and Remote Sensing Letters, DOI:10.1109/LGRS.2017.2657818.
  38. 38.H. Lyu, H. Lu, and L. Mou, “Learning a transferable change rule from a recurrent neural network for land cover change detection,” Remote Sensing, vol. 8, no. 6, p. 506, 2016.
  39. 39.D. E. Dudgeon, R. T. Lacoss, and A. Moreira, “An overview of automatic target recognition,” The Lincoln Laboratory Journal, vol. 6, pp. 3–10, 1993.
  40. 40.S. Chen and H. Wang, “SAR target recognition based on deep learning,” in International Conference on Data Science and Advanced Analytics, 2014.
  41. 41.E. R. Keydel, S. W. Lee, and J. T. Moore, “MSTAR extended operating conditions: a tutorial,” in Proc. SPIE 2757, Algorithms for Synthetic Aperture Radar Imagery III, 1996.
  42. 42.S. Chen, H. Wang, F. Xu, and Y. Q. Jin, “Target classification using the deep convolutional networks for SAR images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 8, pp. 4806–4817, 2016.
  43. 43.D. Morgan, “Deep convolutional neural networks for ATR from SAR imagery,” in Proc. SPIE 9475, Algorithms for Synthetic Aperture Radar Imagery XXII, 2015.
  44. 44.J. Ding, B. Chen, H. Liu, and M. Huang, “Convolutional neural network with data augmentation for SAR target recognition,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 3, pp. 364–368, 2016.
  45. 45.K. Du, Y. Deng, R. Wang, T. Zhao, and N. Li, “SAR ATR based on displacement- and rotation-insensitive CNN,” Remote Sensing Letters, vol. 7, no. 9, pp. 895–904, 2016.
  46. 46.M. Wilmanski, C. Kreucher, and J. Lauer, “Modern approaches in deep learning for SAR ATR,” in Proc. SPIE 9843, Algorithms for Synthetic Aperture Radar Imagery XXIII, 2016.
  47. 47.Z. Cui, Z. Cao, J. Yang, and H. Ren, “Hierarchical recognition system for target recognition from sparse representations,” Mathematical Problems in Engineering, vol. 2015, no. 527095, 2016.
  48. 48.S. A. Wagner, “SAR ATR by a combination of convolutional neural network and support vector machines,” IEEE Transactions on Geoscience and Remote Sensing, vol. 52, no. 6, pp. 2861–2872, 2016.
  49. 49.C. Bentes, A. Frost, D. Velotto, and B. Tings, “Ship-iceberg discrimination with convolutional neural networks in high resolution SAR images,” in European Conference on Synthetic Aperture Radar (EUSAR), 2016.
  50. 50.C. Schwegmann, W. Kleynhans, B. Salmon, L. Mdakane, and R. Meyer, “Very deep learning for ship discrimination in Synthetic Aperture Radar imagery,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2016.
  51. 51.N. Ødegaard, A. O. Knapskog, C. Cochin, and J. C. Louvigne, “Classification of ships using real and simulated data in a convolutional neural network,” in IEEE Radar Conference (RadarConf), 2016.
  52. 52.Q. Song, F. Xu, and Y. Q. Jin, “Deep SAR image generative neural network and auto-construction of target feature space,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2017.
  53. 53.Z. Zhang, H. Wang, F. Xu, and Y. Q. Jin, “Complex-valued convolutional neural networks and its applications to PolSAR image classification,” IEEE Transactions on Geoscience and Remote Sensing, in press.
  54. 54.Y. Q. Jin and F. Xu, “Polarimetric scattering and SAR information retrieval,” Wiley-IEEE, 2013.
  55. 55.F. Xu, Y. Q. Jin, and A. Moreira, “A preliminary study on SAR advanced information retrieval and scene reconstruction,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 10, pp. 1443–1447, 2016.
  56. 56.H. Xie, S. Wang, K. Liu, S. Lin, and B. Hou, “Multilayer feature learning for polarimetric synthetic radar data classification,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2014.
  57. 57.J. Geng, J. Fan, H. Wang, X. Ma, B. Li, and F. Chen, “High-resolution SAR image classification via deep convolutional autoencoders,” IEEE Geoscience and Remote Sensing Letters, vol. 12, no. 11, pp. 2351–2355, 2015.
  58. 58.J. Geng, H. Wang, J. Fan, and X. Ma, “Deep supervised and contractive neural network for SAR image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 4, pp. 2442–2459, 2017.
  59. 59.Q. Lv, Y. Dou, X. Niu, J. Xu, J. Xu, and F. Xia, “Urban land use and land cover classification using remotely sensed SAR data through deep belief networks,” Journal of Sensors, vol. 2015, no. 538063, 2015.
  60. 60.B. Hou, H. Kou, and L. Jiao, “Classification of polarimetric SAR images using multilayer autoencoders and superpixels,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 9, no. 7, pp. 3072–3081, 2016.
  61. 61.L. Zhang, W. Ma, and D. Zhang, “Stacked sparse autoencoder in PolSAR data classification using local spatial information,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 9, pp. 1359–1363, 2016.
  62. 62.F. Qin, J. Guo, and W. Sun, “Object-oriented ensemble classification for polarimetric SAR imagery using restricted Boltzmann machines,” Remote Sensing Letters, vol. 8, no. 3, pp. 204–213, 2017.
  63. 63.Z. Zhao, L. Jiao, J. Zhao, J. Gu, and J. Zhao, “Discriminant deep belief network for high-resolution SAR image classification,” Pattern Recognition, vol. 61, pp. 686–701, 2017.
  64. 64.L. Jiao and F. Liu, “Wishart deep stacking network for fast PolSAR image classification,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 3273–3286, 2016.
  65. 65.Y. Zhou, H. Wang, F. Xu, and Y. Q. Jin, “Polarimetric SAR image classification using deep convolutional neural networks,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 12, pp. 1935–1939, 2016.
  66. 66.Y. Duan, F. Liu, L. Jiao, P. Zhao, and L. Zhang, “SAR image segmentation based on convolutional-wavelet neural network and markov random field,” Pattern Recognition, vol. 64, pp. 255–267, 2017.
  67. 67.L. Wang, K. A. Scott, L. Xu, and D. A. Clausi, “Sea ice concentration estimation during melt from dual-Pol SAR scenes using deep convolutional neural networks: A case study,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 8, pp. 4524–4533, 2016.
  68. 68.W. Yang, D. Dai, B. Triggs, and G. Xia, “Sar-based terrain classification using weakly supervised hierarchical markov aspect models,” IEEE Trans. Image Processing, vol. 21, no. 9, pp. 4232–4243, 2012.
  69. 69.W. Shao, W. Yang, and G.-S. Xia, “Extreme value theory-based calibration for multiple feature fusion in high-resolution satellite scene classification,” International Journal of Remote Sensing, vol. 34, no. 3, pp. 8588–8602, 2013.
  70. 70.W. Yang, X. Yin, and G. Xia, “Learning high-level features for satellite image classification with limited labeled samples,” IEEE Trans. Geoscience and Remote Sensing, vol. 53, no. 8, pp. 4472–4482, 2015.
  71. 71.F. Hu, G. S. Xia, Z. Wang, X. Huang, L. Zhang, and H. Sun, “Unsupervised feature learning via spectral clustering of multidimensional patches for remotely sensed scene classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 8, no. 5, pp. 2015–2030, 2015.
  72. 72.B. Zhao, Y. Zhong, G. Xia, and L. Zhang, “Dirichlet-derived multiple topic scene classification model for high spatial resolution remote sensing imagery,” IEEE Trans. Geoscience and Remote Sensing, vol. 54, no. 4, pp. 2108–2123, 2016.
  73. 73.F. Hu, G. Xia, J. Hu, Y. Zhong, and K. Xu, “Fast binary coding for the scene classification of high-resolution remote sensing imagery,” Remote Sensing, vol. 8, no. 7, p. 555, 2016.
  74. 74.G.-S. Xia, J. Hu, B. Shi, X. Bai, Y. Zhong, X. Lu, and L. Zhang, “AID: A benchmark dataset for performance evaluation of aerial scene classification,” IEEE Transactions on Geoscience and Remote Sensing, 2017.
  75. 75.S. Bhagavathy and B. S. Manjunath, “Modeling and detection of geospatial objects using texture motifs,” IEEE Transactions on Geoscience and Remote Sensing, vol. 44, no. 12, pp. 3706–3715, 2006.
  76. 76.G. Cheng, P. Zhou, and J. Han, “Learning rotation-invariant convolutional neural networks for object detection in VHR optical remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 12, pp. 7405–7415, 2016.
  77. 77.X. Chen, H. Zhao, P. Li, and Z. Yin, “Remote sensing image-based analysis of the relationship between urban heat island and land use/cover changes,” Remote Sensing of Environment, vol. 104, no. 2, pp. 133–146, 2006.
  78. 78.J. Sivic and A. Zisserman, “Video google: A text retrieval approach to object matching in videos,” in IEEE International Conference on Computer Vision (ICCV)), 2003.
  79. 79.Q. Zhu, Y. Zhong, B. Zhao, G.-S. Xia, and L. Zhang, “Bag-of-visual-words scene classifier with local and global features for high spatial resolution remote sensing imagery,” IEEE Geosci. Remote Sensing Lett., vol. 13, no. 6, pp. 747–751, 2016.
  80. 80.Q. Zou, L. Ni, T. Zhang, and Q. Wang, “Deep learning based feature selection for remote sensing scene classification,” IEEE Geoscience and Remote Sensing Letters, vol. 12, no. 11, pp. 2321–2325, 2015.
  81. 81.O. Penatti, K. Nogueira, and J. Santos, “Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  82. 82.M. Castelluccio, G. Poggi, C. Sansone, and L. Verdoliva, “Land use classification in remote sensing images by convolutional neuralnetworks,” arXiv:1508.00092, 2015.
  83. 83.F. Hu, G.-S. Xia, J. Hu, and L. Zhang, “Transferring deep convolutional neural networks for the scene classification of high-resolution remote sensing imagery,” Remote Sensing, vol. 7, no. 11, pp. 14 680–14 707, 2015.
  84. 84.F. Zhang, B. Du, and L. Zhang, “Scene classification via a gradient boosting random convolutional network framework,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 3, pp. 1793–1802, 2016.
  85. 85.F. Luus, B. Salmon, F. Bergh, and B. Maharaj, “Multiview deep learning for land-use classification,” IEEE Geoscience and Remote Sensing Letters, vol. 12, no. 12, pp. 2448–2452, 2015.
  86. 86.K. Nogueira, O. Penatti, and J. Santos, “Towards better exploiting convolutional neural networks for remote sensing scene classification,” Pattern Recognition, vol. 61, pp. 539–556, 2016.
  87. 87.D. Marmanis, M. Datcu, T. Esch, and U. Stilla, “Deep learning earth observation classification using imagenet pretrained networks,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 1, pp. 105–109, 2016.
  88. 88.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” in IEEE International Conference on Learning Representations (ICLR), 2014.
  89. 89.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  90. 90.Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, 2010.
  91. 91.G.-S. Xia, W. Yang, J. Delon, Y. Gousseau., H. Sun, and H. Maitre, “Structural high-resolution satellite image indexing,” in Symposium: 100 Years ISPRS - Advancing Remote Sensing Science: Vienna, Austria, 2010.
  92. 92.M. Volpi and D. Tuia, “Dense semantic labeling of subdecimeter resolution images with convolutional neural networks,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 2, pp. 881–893, 2017.
  93. 93.G. Cheng and J. Han, “A survey on object detection in optical remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 117, pp. 11–28, 2016.
  94. 94.N. Yokoya and A. Iwasaki, “Object detection based on sparse representation and hough voting for optical remote sensing imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 8, no. 5, pp. 2053–2062, 2015.
  95. 95.J. Han, P. Zhou, D. Zhang, G. Cheng, L. Guo, Z. Liu, S. Bu, and J. Wu, “Efficient, simultaneous detection of multi-class geospatial targets based on visual saliency modeling and discriminative learning of sparse coding,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 89, pp. 37–48, 2014.
  96. 96.X. Jin and C. H. Davis, “Vehicle detection from high-resolution satellite imagery using morphological shared-weight neural networks,” Image and Vision Computing, vol. 25, no. 9, pp. 1422–1431, 2007.
  97. 97.X. Chen, S. Xiang, C.-L. Liu, and C.-H. Pan, “Vehicle detection in satellite images by hybrid deep convolutional neural networks,” IEEE Geoscience and remote sensing letters, vol. 11, no. 10, pp. 1797–1801, 2014.
  98. 98.Q. Jiang, L. Cao, M. Cheng, C. Wang, and J. Li, “Deep neural networks-based vehicle detection in satellite images,” in International Symposium on Bioelectronics and Bioinformatics, 2015.
  99. 99.P. Zhou, G. Cheng, Z. Liu, S. Bu, and X. Hu, “Weakly supervised target detection in remote sensing images based on transferred deep features and negative bootstrapping,” Multidimensional Systems and Signal Processing, vol. 27, no. 4, pp. 925–944, 2016.
  100. 100.L. Zhang, Z. Shi, and J. Wu, “A hierarchical oil tank detector with deep surrounding features for high-resolution optical satellite imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 8, no. 10, pp. 4895–4909, 2015.
  101. 101.N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2005.
  102. 102.A.-B. Salberg, “Detection of seals in remote sensing images using features extracted from deep convolutional neural networks,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2015.
  103. 103.I. Ševo and A. Avramovi ´ c, “Convolutional neural network based automatic object detection on aerial images,” IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 5, pp. 740–744, 2016.
  104. 104.H. Zhu, X. Chen, W. Dai, K. Fu, Q. Ye, and J. Jiao, “Orientation robust object detection in aerial images using deep convolutional neural network,” in IEEE International Conference on Image Processing (ICIP), 2015.
  105. 105.F. Zhang, B. Du, L. Zhang, and M. Xu, “Weakly supervised learning based on coupled convolutional neural networks for aircraft detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 9, pp. 5553–5563, 2016.
  106. 106.J. Tang, C. Deng, G.-B. Huang, and B. Zhao, “Compressed-domain ship detection on spaceborne optical image using deep neural network and extreme learning machine,” IEEE Transactions on Geoscience and Remote Sensing, vol. 53, no. 3, pp. 1174–1185, 2015.
  107. 107.G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew, “Extreme learning machine: Theory and applications,” Neurocomputing, vol. 70, no. 1, pp. 489–501, 2006.
  108. 108.J. Han, D. Zhang, G. Cheng, L. Guo, and J. Ren, “Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 53, no. 6, pp. 3325–3337, 2015.
  109. 109.Y. Yang and S. Newsam, “Geographic image retrieval using local invariant features,” IEEE Transactions on Geoscience and Remote Sensing, vol. 51, no. 2, pp. 818–832, 2013.
  110. 110.S. Özkan, T. Ateş, E. Tola, M. Soysal, and E. Esen, “Performance analysis of state-of-the-art representation methods for geographical image retrieval and categorization,” IEEE Geoscience and Remote Sensing Letters, vol. 11, no. 11, pp. 1996–2000, 2014.
  111. 111.P. Napoletano, “Visual descriptors for content-based retrieval of remote sensing images,” arXiv:1602.00970, 2016.
  112. 112.W. Zhou, S. Newsam, C. Li, and Z. Shao, “Learning low dimensional convolutional neural networks for high-resolution remote sensing image retrieval,” Remote Sensing, vol. 9, no. 5, p. 489, 2017.
  113. 113.T. Jiang, G.-S. Xia, and Q. Lu, “Sketch-based aerial image retrieval,” in IEEE International Conference on Image Processing (ICIP), 2017.
  114. 114.L. Gómez-Chova, D. Tuia, G. Moser, and G. Camps-Valls, “Multimodal classification of remote sensing images: A review and future directions,” Proceedings of the IEEE, vol. 103, no. 9, pp. 1560–1584, 2015.
  115. 115.M. Schmitt and X. X. Zhu, “Data fusion and remote sensing: An ever-growing relationship,” IEEE Geoscience and Remote Sensing Magazine, vol. 4, no. 4, pp. 6–23, 2016.
  116. 116.L. Mou, X. X. Zhu, M. Vakalopoulou, K. Karantzalos, N. Paragios, B. L. Saux, G. Moser, and D. Tuia, “Multi-temporal very high resolution from space: Outcome of the 2016 IEEE GRSS data fusion contest,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 10, no. 8, pp. 3435–3447, 2017.
  117. 117.L. Alparone, L. Wald, J. Chanussot, C. Thomas, P. Gamba, and L. M. Bruce, “Comparison of pansharpening algorithms: Outcome of the 2006 grs-s data-fusion contest,” IEEE Transactions on Geoscience and Remote Sensing, vol. 45, no. 10, pp. 3012–3021, 2007.
  118. 118.D. Fasbender, D. Tuia, M. Kanevski, and P. Bogaert, “Support-based implementation of Bayesian data fusion for spatial enhancement: Applications to aster thermal images,” IEEE Geoscience and Remote Sensing Letters, vol. 5, no. 4, pp. 589–602, 2008.
  119. 119.L. Loncan, L. B. Almeida, J. M. Bioucas-Dias, X. Briottet, J. Chanussot, N. Dobigeon, S. Fabre, W. Liao, G. A. Licciardi, M. Simoes, J. Y. Tourneret, M. A. Veganzones, G. Vivone, Q. Wei, and N. Yokoya, “Hyperspectral pansharpening: A review,” IEEE Geoscience and Remote Sensing Magazine, vol. 3, no. 3, pp. 27–46, 2015.
  120. 120.J. Zhong, B. Yang, G. Huang, F. Zhong, and Z. Chen, “Remote sensing image fusion with convolutional neural network,” Sensing and Imaging, vol. 17, 2016.
  121. 121.G. Masi, D. Cozzolino, L. Verdoliva, and G. Scarpa, “Pansharpening by convolutional neural networks,” Remote Sensing, vol. 8, no. 7, p. 594, 2016.
  122. 122.Y. Yuan, S. Zheng, and X. Lu, “Hyperspectral image superresolution by transfer learning,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 10, no. 5, pp. 1963–1974, 2017.
  123. 123.C. Dong, C. C. Loy, K. He, and X. Tang, “Learning a deep convolutional network for image super-resolution,” in European Conference on Computer Vision (ECCV), 2014.
  124. 124.D. Tuia, C. Persello, and L. Bruzzone, “Recent advances in domain adaptation for the classification of remote sensing data,” IEEE Geoscience and Remote Sensing Magazine, vol. 4, no. 2, pp. 41–57, 2016.
  125. 125.W. Huang, L. Xiao, Z. Wei, H. Liu, and S. Tang, “A new pan-sharpening method with deep neural networks,” IEEE Geoscience and Remote Sensing Letters, vol. 12, no. 5, pp. 1037–1041, 2015.
  126. 126.A. Lagrange, B. L. Saux, A. Beaupere, A. Boulch, A. Chan-Hon-Tong, S. Herbin, H. Randrianarivo, and M. Ferecatu, “Benchmarking classification of earth-observation data: from learning explicit features to convolutional networks,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2015.
  127. 127.M. Campos-Taberner, A. Romero-Soriano, C. Gatta, G. Camps-Valls, A. Lagrange, B. L. Saux, A. Beaupere, A. Boulch, A. Chan-Hon-Tong, S. Herbin, H. Randrianarivo, ` M. Ferecatu, M. Shimoni, G. Moser, and D. Tuia, “Processing of extremely high resolution LiDAR and RGB data: outcome of the 2015 IEEE GRSS Data Fusion Contest. Part A: 2D contest,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 9, no. 12, pp. 5547–5559, 2016.
  128. 128.E. Maggiori, Y. Tarabalka, G. Charpiat, and P. Alliez, “High-resolution semantic labeling with convolutional neural networks,” arXiv:1611.01962, 2017.
  129. 129.M. Kampffmeyer, A. B. Salberg, and R. Jenssen, “Semantic segmentation of small objects and modeling of uncertainty in urban remote sensing images using deep convolutional neural networks,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2016.
  130. 130.Y. Gal and Z. Ghahramani, “Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” arXiv:1506.02142, 2015.
  131. 131.D. Tuia, N. Courty, and R. Flamary, “Multiclass feature learning for hyperspectral image classification: Sparse and hierarchical solutions,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 105, pp. 272–285, 2015.
  132. 132.M. Volpi, G. Camps-Valls, and D. Tuia, “Spectral alignment of cross-sensor images with automated kernel canonical correlation analysis,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 107, pp. 50–63, 2015.
  133. 133.D. Marcos, R. Hamid, and D. Tuia, “Geospatial correspondence for multimodal registration,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  134. 134.M. Gong, T. Zhan, P. Zhang, and Q. Miao, “Superpixel-based difference representation learning for change detection in multispectral remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 5, pp. 2658 – 2673, 2017.
  135. 135.P. Zhang, M. Gong, L. Su, J. Liu, and Z. Li, “Change detection based on deep feature representation and mapping transformation for multi-spatial-resolution remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 116, pp. 24–41, 2016.
  136. 136.H. Lyu, H. Lu, L. Mou, W. Li, X. Li, X. Li, J. Wang, X. X. Zhu, L. Yu, and P. Gong, “A deep information based transfer learning method to detect annual urban dynamics of four developed cities from 1984-2016 by Landsat data,” Remote Sensing of Environment, in revision.
  137. 137.J. Sherrah, “Fully convolutional networks for dense semantic labelling of high-resolution aerial imagery,” arXiv:1606.02585, 2016.
  138. 138.A. Marcu and M. Leordeanu, “Dual local-global contextual pathways for recognition in aerial imagery,” arXiv:1605:05462, 2016.
  139. 139.J. Hu, L. Mou, A. Schmitt, and X. X. Zhu, “FusioNet: A two-stream convolutional neural network for urban scene classification using PolSAR and hyperspectral data,” in Joint Urban Remote Sensing Event (JURSE), 2017.
  140. 140.L. Mou, M. Schmitt, Y. Wang, and X. X. Zhu, “A CNN for the identification of corresponding patches in SAR and optical imagery of urban scenes,” in Joint Urban Remote Sensing Event (JURSE), 2017.
  141. 141.S. Paisitkriangkrai, J. Sherrah, P. Janney, and A. van den Hengel, “Semantic labeling of aerial and satellite imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 9, no. 7, pp. 2868–2881, 2016.
  142. 142.N. Audebert, B. L. Saux, and S. Lefevre, “Semantic segmentation of earth observation ` data using multimodal and multi-scale deep networks,” in Asian Conference on Computer Vision (ACCV), 2016.
  143. 143.D. Marmanis, K. Schindler, J. D. Wegner, S. Galliani, M. Datcu, and U. Stilla, “Classification with an edge: Improving semantic image segmentation with boundary detection,” arXiv:1612.01337, 2017.
  144. 144.D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” in Advances in Neural Information Processing Systems (NIPS), 2014.
  145. 145.S. Srivastava, M. Volpi, and D. Tuia, “Joint height estimation and semantic labeling of monocular aerial images with CNNs,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2017.
  146. 146.M. Vakalopoulou, C. Platias, M. Papadomanolaki, N. Paragios, and K. Karantzalos, “Simultaneous registration, segmentation and change detection from multisensor, multitemporal satellite image pairs,” in IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2016.
  147. 147.S. Lefevre, D. Tuia, J. D. Wegner, T. Produit, and A. S. Nassar, “Towards seamless ` multi-view scene analysis from satellite to street-level,” Proceedings of the IEEE, DOI:10.1109/JPROC.2017.2684300.
  148. 148.J. D. Wegner, S. Branson, D. Hall, and P. Perona, “Cataloging public objects using aerial and street-level images c urban trees,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  149. 149.S. Ren, K. He, R. Girshick, and J.Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems (NIPS), 2015.
  150. 150.G. Mattyus, S. Wang, S. Fidler, and R. Urtasun, “Hd maps: Fine-grained road segmentation by parsing ground and aerial images,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  151. 151.T. Lin, Y. Cui, S. Belongie, and J. Hays, “Learning deep representations for ground-to-aerial geolocalization,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  152. 152.N. N. Vo and J. Hays, “Localizing and orienting street views using overhead imagery,” in European Conference on Computer Vision (ECCV), 2016.
  153. 153.S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2005.
  154. 154.S. Workman and N. Jacobs, “On the location dependence of convolutional neural network features,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2015.
  155. 155.S. Workman, R. Souvenir, and N. Jacobs, “Wide-area image geolocalization with aerial reference imagery,” in IEEE International Conference on Computer Vision (ICCV), 2015.
  156. 156.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in Advances in Neural Information Systems (NIPS), 2014.
  157. 157.D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004.
  158. 158.H. Bay, A. Ess, T. Tuytelaars, and L. V. Gool, “Speeded-up robust features (SURF),” Computer Vision and Image Understanding, vol. 110, no. 3, pp. 346–359, 2008.
  159. 159.P. Fischer, A. Dosovitskiy, and T. Brox, “Descriptor matching with convolutional neural networks: a comparison to SIFT,” arXiv:1406.6909, 2014.
  160. 160.A. Handa, M. Blosch, V. Patraucean, S. Stent, J. McCormac, and A. J. Davison, “gvnn: ¨ Neural network library for geometric computer vision,” arXiv:1607.07405, 2016.
  161. 161.K. Lenc and A. Vedaldi, “Learning covariant feature detectors,” arXiv:1605.01224, 2016.
  162. 162.X. Han, T. Leung, Y. Jia, R. Sukthankar, and A. C. Berg, “MatchNet: Unifying feature and metric learning for patch-based matching,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  163. 163.K. Yi, E. Trulls, V. Lepetit, and P. Fua, “LIFT: Learned invariant feature transform,” in European Conference on Computer Vision (ECCV), 2016.
  164. 164.H. Hirschmuller, “Stereo processing by semiglobal matching and mutual information,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 328–341, 2008.
  165. 165.J. Zbontar and Y. LeCun, “Computing the stereo matching cost with a convolutional neural network,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  166. 166.L. Li, X. Yu, S. Zhang, X. Zhao, and L. Zhang, “3D cost aggregation with multiple minimum spanning trees for stereo matching,” Applied Optics, vol. 56, no. 12, pp. 3411–3420, 2017.
  167. 167.S. Drouyer, S. Beucher, M. Bilodeau, M. Moreaud, and L. Sorbier, “Sparse stereo disparity map densification using hierarchical image segmentation,” in International Symposium on Mathematical Morphology and Its Applications to Signal and Image Processing, 2017.
  168. 168.H. Park and K. M. Lee, “Look wider to match image patches with convolutional neural networks,” IEEE Signal Processing Letters, DOI:10.1109/LSP.2016.2637355.
  169. 169.N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, ¨ “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  170. 170.D. Marmanis, J. D. Wegner, S. Galliani, K. Schindler, M. Datcu, and U. Stilla, “Semantic segmentation of aerial images with an ensemble of fully convolutional neural networks,” in ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, 2016.
  171. 171.C. Hane, C. Zach, A. Cohen, R. Angst, and M. Pollefeys, “Joint 3D scene reconstruction ¨ and class segmentation,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2013.
  172. 172.M. Blaha, C. Vogel, A. Richard, J. D. Wegner, T. Pock, and K. Schindler, “Large-scale ´ semantic 3D reconstruction: An adaptive multi-resolution model for multi-class volumetric labeling,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  173. 173.M. Blahá, C. Vogel, A. Richard, J. D. Wegner, T. Pock, and K. Schindler, “Towards inte- ´ grated 3D reconstruction and semantic interpretation of urban scenes,” in Dreiländertagung der SGPF, DGPF und OVG : Lösungen für eine Welt im Wandel : Vorträge, 2016.
  174. 174.Y. Bengio, “Practical recommendations for gradient-based training of deep architectures,” arXiv:1206.5533, 2012.
  175. 175.G. Montavon, G. B. Orr, and K.-R. Müller, Neural networks: Tricks of the trade. Springer, 2012.
  176. 176.M. Castelluccio, G. Poggi, C. Sansone, and L. Verdoliva, “Land use classification in remote sensing images by convolutional neural networks,” arXiv:1508.00092, 2015.
  177. 177.Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (ACM GIS), 2010.
  178. 178.G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proceedings of the IEEE, DOI: 10.1109/JPROC.2017.2675998.
  179. 179.M. Volpi and V. Ferrari, “Semantic segmentation of urban scenes by learning local class interactions,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR) Workshop on EarthVision, 2015.
  180. 180.Y. Wang, X. X. Zhu, B. Zeisl, and M. Pollefeys, “Fusing meter-resolution 4-D InSAR point clouds and optical images for semantic urban infrastructure monitoring,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 1, pp. 14–26, 2017.
  181. 181.A. Kendall, V. Badrinarayanan, and R. Cipolla, “Bayesian SegNet: Model uncertainty in deep convolutional encoder-decoder architectures for scene understanding,” arXiv:1511.02680, 2015.
  182. 182.P. Gong, L. Yu, C. Li, J. Wang, L. Liang, X. Li, L. Ji, Y. Bai, Y. Cheng, and Z. Zhu, “A new research paradigm for global land cover mapping,” Annals of GIS, vol. 22, no. 2, pp. 1–16, 2016.
  183. 183.T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, J. Krishnamurthy, N. Lao, K. Mazaitis, T. Mohamed, N. Nakashole, E. Platanios, A. Ritter, M. Samadi, B. Settles, R. Wang, D. Wijaya, A. Gupta, X. Chen, A. Saparov, M. Greaves, and J. Welling, “Never-ending learning,” in Proceedings of the Conference on Artificial Intelligence (AAAI), 2015.
  184. 184.R. Raina, A. Battle, H. Lee, B. Packer, and A. Y. Ng, “Self-taught learning: Transfer learning from unlabeled data,” in IEEE International Conference on Machine Learning (ICML), 2007.
  185. 185.T. Durand, N. Thome, and M. Cord, “Weldon: Weakly supervised learning of deep convolutional neural networks,” in IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  186. 186.R. Johnson and T. Zhang, “Supervised and semi-supervised text categorization using lstm for region embeddings,” in IEEE International Conference on Machine Learning (ICML), 2016.

Citation

MLA
Zhu, X. X., et al. “Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources”. IEEE Geoscience and Remote Sensing Magazine, vol. 5, no. 4, 2017, pp. 8–6, https://doi.org/10.1109/MGRS.2017.2762307.
APA
Zhu, X. X., Tuia, D., Mou, L., Xia, G.-S., Zhang, L., Xu, F., & Fraundorfer, F. (2017). Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources. IEEE Geoscience and Remote Sensing Magazine, 5(4), 8–36. https://doi.org/10.1109/MGRS.2017.2762307
Chicago
Zhu, X. X., D. Tuia, L. Mou, et al. 2017. “Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources”. IEEE Geoscience and Remote Sensing Magazine 5 (4): 8–36. https://doi.org/10.1109/MGRS.2017.2762307.
Harvard
Zhu, X.X. et al. (2017) “Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources”, IEEE Geoscience and Remote Sensing Magazine, 5(4), pp. 8–36. Available at: https://doi.org/10.1109/MGRS.2017.2762307.
Vancouver
1. Zhu XX, Tuia D, Mou L, Xia G-S, Zhang L, Xu F, Fraundorfer F (2017) Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources. IEEE Geoscience and Remote Sensing Magazine 5:8–36

BibTeX

@article{Zhu_2017, title={Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources}, volume={5}, ISSN={2373-7468}, url={http://dx.doi.org/10.1109/MGRS.2017.2762307}, DOI={10.1109/mgrs.2017.2762307}, number={4}, journal={IEEE Geoscience and Remote Sensing Magazine}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Zhu, Xiao Xiang and Tuia, Devis and Mou, Lichao and Xia, Gui-Song and Zhang, Liangpei and Xu, Feng and Fraundorfer, Friedrich}, year={2017}, month=Dec, pages={8–36} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors