AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification
Gui-Song XiaJingwen HuFan HuBaoguang ShiXiang BaiYanfei ZhongLiangpei ZhangXiaoqiang Lu
Introduces the Aerial Image Dataset (AID), a large-scale benchmark of over ten thousand annotated images that overcomes performance saturation in smaller datasets by establishing baseline evaluations for deep learning models in remote sensing scene classification.
Automated aerial scene classification is critical for earth observation, urban planning, and environmental monitoring. However, development in this field has stalled because existing standard test sets are too small and lack diversity, resulting in inflated, saturated performance metrics that do not reflect real-world complexity.
The article introduces the Aerial Image Dataset, a large-scale collection of ten thousand images across thirty semantic scene categories gathered globally across different imaging conditions, resolutions, and seasons. The investigation systematically evaluates and establishes performance baselines across low-level, mid-level, and deep learning visual classification methods.
The findings show that deep neural networks significantly outperform traditional methods across all datasets, achieving overall accuracy rates near ninety percent on the new dataset, compared to thirty to thirty-seven percent for low-level features and seventy-two to seventy-nine percent for top mid-level methods. Furthermore, the new dataset exposes substantial classification challenges in newly introduced, fine-grained categories that share complex structures, such as schools versus dense residential areas and resorts versus parks, where accuracy drops to between forty-nine and sixty-five percent.
These results indicate that while deep learning provides robust feature representations, standard automated vision systems still struggle with semantic ambiguity between structurally similar land-use types. Future development must focus on advancing high-level architectures tailored to resolve fine-grained spatial distinctions in diverse, large-scale imagery.
- Paper: Learning Deep Features for Scene Recognition using Places Database, Bolei Zhou et al. (2014). Read this account of Places-trained deep scene features first to understand the deep-learning approach that AID evaluates against traditional visual features.
- Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). Its spatial-pyramid method supplies an influential mid-level scene-classification baseline that helps make sense of AID’s comparisons with learned features.
- Paper: A Bayesian hierarchical model for learning natural scene categories, Li Fei-Fei et al. (2005). This earlier probabilistic scene-classification approach provides useful context for the traditional methods that AID benchmarks against deep networks.
- Paper: ImageNet: A large-scale hierarchical image database, Jia Deng et al. (2009). Its large-scale labeled image resource provides background on the data foundation that enabled the pretrained visual models evaluated in AID.
- Paper: Remote Sensing Image Scene Classification: Benchmark and State of the Art, Gong Cheng et al. (2017). This later benchmark enlarges AID’s remote-sensing scene-classification effort with more images and categories, allowing readers to see how the field addressed its scale and diversity limits.
- Paper: EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification, Patrick Helber et al. (2017). EuroSAT carries benchmark-based land-use classification into multispectral satellite data, extending AID’s evaluation of visual classifiers to a different Earth-observation input.
- Paper: GEO-Bench: Toward Foundation Models for Earth Monitoring, Alexandre Lacoste et al. (2023). GEO-Bench continues the push for comparable Earth-observation evaluations by standardizing classification and segmentation tasks across multiple geospatial datasets.
