Remote Sensing Image Scene Classification: Benchmark and State of the Art
Gong ChengJunwei HanXiaoqiang Lu
Introduces NWPU-RESISC45, a large-scale benchmark dataset of 31,500 images across 45 classes, paired with comprehensive baseline evaluations and a systematic survey to overcome data diversity and scale limitations in remote sensing scene classification.
Remote sensing image scene classification supports critical applications such as land-use mapping, natural hazard detection, urban planning, and environmental monitoring. The field has advanced through various methods and public datasets, yet existing collections remain small, lack image diversity, and show near-saturated accuracy, which restricts progress on modern data-driven techniques including deep learning.
This paper reviews roughly 170 publications on datasets and methods, introduces a new benchmark, and tests representative approaches to establish performance baselines. The authors first catalog six prior datasets and categorize methods into handcrafted features, unsupervised feature learning, and deep feature learning. They then release NWPU-RESISC45, a publicly available collection of 31,500 images spanning 45 scene classes with 700 images each, drawn from Google Earth across more than 100 countries. The images exhibit substantial variation in scale, viewpoint, illumination, occlusion, and background. Twelve methods—ranging from color histograms and local binary patterns to bag-of-visual-words variants and pre-trained or fine-tuned convolutional networks—were evaluated under 10 % and 20 % training splits using linear support-vector machines.
Deep convolutional features substantially outperformed earlier approaches, delivering at least 30 percentage points higher accuracy than handcrafted or unsupervised methods. Fine-tuning the networks on the new data produced further gains, with the best result reaching approximately 90 % overall accuracy under the 20 % training regime. Handcrafted global features remained weakest, while mid-level encodings such as bag-of-visual-words offered modest improvement but still trailed deep models. Confusion persisted between visually similar classes such as churches and palaces or dense and medium residential areas.
These results demonstrate that large, diverse benchmarks are essential for advancing scene classification and that current deep models already provide strong baselines. They also indicate that overhead imagery alone leaves semantic gaps that additional data sources could help close. The authors therefore recommend exploring fusion of satellite imagery with geo-tagged ground photos and location-based social-media streams to capture finer vertical and contextual details. They note that the reported figures rest on linear classifiers and fixed training ratios; broader testing with varied architectures and larger labeled sets would increase confidence before operational deployment.
- Paper: AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification, Gui-Song Xia et al. (2016). Read this earlier aerial-scene benchmark first: its dataset and baseline comparisons establish the direct benchmark lineage that the source expands with a larger, more diverse collection.
- Paper: Learning Deep Features for Scene Recognition using Places Database, Bolei Zhou et al. (2014). Its Places-trained CNN features provide a key precursor to the source’s comparison of pretrained deep features against handcrafted approaches for scene recognition.
- Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). This spatial-pyramid method explains the spatially organized visual-word features behind the source’s handcrafted and mid-level baselines.
- Paper: ImageNet Large Scale Visual Recognition Challenge, Olga Russakovsky et al. (2014). Its ImageNet benchmark and CNN-driven performance gains provide context for the pretrained networks that the source evaluates on remote-sensing scenes.
- Paper: GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models, Gilles Quentin Hacheme et al. (2025). This later zero-shot study tests RESISC45 directly, carrying the source’s benchmark into a setting that reduces reliance on labeled training data.
- Paper: GEO-Bench: Toward Foundation Models for Earth Monitoring, Alexandre Lacoste et al. (2023). GEO-Bench extends the source’s benchmark-building agenda into standardized evaluation across diverse Earth-monitoring classification and segmentation tasks.
