Deep learning in remote sensing: a review
Xiao Xiang ZhuDevis TuiaLichao MouGui-Song XiaLiangpei ZhangFeng XuFriedrich Fraundorfer
Surveys key advances and open challenges in applying deep learning to remote sensing data, providing practical resources and strategies to integrate Earth observation domain knowledge for tackling large-scale environmental problems.
The rapid growth of Earth observation data has created a critical need for advanced automated analysis. Modern satellite constellations produce petabytes of imagery characterized by complex physical properties, multi-sensor modalities, exact spatial coordinates, and dense time series. Traditional remote sensing techniques rely heavily on manual feature engineering and domain-specific heuristics, which struggle to scale with massive volumes of data or capture intricate, nonlinear patterns. The article provides a comprehensive evaluation of how modern deep learning architectures address these challenges across major Earth observation applications, while also cataloging open-source software frameworks and public benchmark datasets to facilitate research adoption.
To conduct this evaluation, the article synthesizes experimental findings, network architectures, and benchmark evaluations across five core subfields: hyperspectral image analysis, synthetic aperture radar interpretation, high-resolution optical image processing, multimodal data fusion, and 3D reconstruction. It reviews foundational deep learning models—including autoencoders, deep belief networks, convolutional neural networks, and recurrent neural networks—and evaluates their performance against classical baselines on standard international benchmarks, such as Middlebury stereo evaluation datasets and MSTAR radar collections.
The findings show that deep learning consistently outperforms traditional hand-crafted methods. In stereo matching for 3D reconstruction, convolutional neural network approaches cut bad pixel error rates from an 18.4% baseline under traditional semi-global matching down to between 5.9% and 8.1%. In radar automatic target recognition, specialized convolutional networks achieve up to 99.1% accuracy under standard operating conditions when paired with data augmentation. For optical and hyperspectral data, deep networks demonstrate superior capacity for end-to-end pixel-level classification, large-scale scene categorization, and automated feature extraction from unlabeled data. Furthermore, deep learning facilitates complex multimodal data fusion, enabling unified workflows for simultaneous image registration, land cover classification, and multi-decadal change detection.
These results demonstrate that deep neural networks offer significant operational gains, allowing automated systems to process global-scale satellite streams with higher precision and lower manual engineering overhead. However, the article emphasizes that purely data-driven, black-box approaches must not completely replace physical and domain-specific knowledge. Instead, integrating physical sensor characteristics, prior geographical constraints, and domain expertise into network architectures is essential to avoid severe errors when processing complex geospatial physics.
Decision-makers and research teams should focus future efforts on combining physics-based process models with deep learning architectures, developing weakly supervised and unsupervised methods to mitigate the scarcity of labeled Earth observation data, and establishing global transferability pipelines. While current benchmarks demonstrate high confidence for localized and well-annotated tasks, users should exercise caution when deploying models globally, as performance can degrade across unseen geographic regions, varying atmospheric conditions, and distinct sensor geometries.
- Paper: AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification, Gui-Song Xia et al. (2016). Its AID benchmark and comparisons of deep versus hand-crafted scene features provide an early remote-sensing example behind the review’s discussion of scene classification.
- Paper: A Survey on Object Detection in Optical Remote Sensing Images, Gong Cheng et al. (2016). This pre-review survey lays out the object-detection methods and evaluation landscape that the source situates within deep learning applications.
- Paper: Learning to compare image patches via convolutional neural networks, Sergey Zagoruyko et al. (2015). Its learned patch-matching CNNs are a key precursor for understanding the source’s discussion of deep-learning approaches to stereo reconstruction.
- Paper: Very Deep Convolutional Networks for Large-Scale Image Recognition, Karen Simonyan et al. (2015). VGG’s deep CNN design supplies foundational architecture context for the convolutional networks the review evaluates across Earth-observation tasks.
- Paper: Fully Convolutional Siamese Networks for Change Detection, Rodrigo Caye Daudt et al. (2018). This work turns the review’s account of CNNs for Earth observation into end-to-end Siamese models for pixel-level change detection.
- Paper: Deep Learning for Hyperspectral Image Classification: An Overview, Shutao Li et al. (2019). It follows the review’s broad treatment of hyperspectral learning with a focused assessment of architectures, label scarcity, and benchmark performance.
- Paper: HybridSN: Exploring 3-D–2-D CNN Feature Hierarchy for Hyperspectral Image Classification, Swalpa Kumar Roy et al. (2019). HybridSN advances the review’s hyperspectral classification discussion with a specific 3D–2D CNN design for joint spectral-spatial features.
- Paper: Remote Sensing Image Change Detection With Transformers, Hao Chen et al. (2021). This study continues the review’s change-detection thread by replacing CNN-only context modeling with an efficient transformer architecture.
- Paper: GEO-Bench: Toward Foundation Models for Earth Monitoring, Alexandre Lacoste et al. (2023). GEO-Bench carries forward the review’s call for public benchmarks by standardizing evaluation across diverse Earth-monitoring tasks.
- Paper: Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark, Ke Li et al. (2019). This survey and DIOR benchmark extend the review’s treatment of optical imagery with a dedicated evaluation of remote-sensing object detectors.
