EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
Patrick HelberBenjamin BischkeAndreas DengelDamian Borth
Introduces EuroSAT, a public benchmark of 27,000 geo-referenced Sentinel-2 satellite images spanning 13 spectral bands to advance deep learning methods for land use classification and Earth observation mapping.
The EuroSAT paper introduces a new openly available dataset of 27,000 labeled 64-by-64 pixel Sentinel-2 image patches spanning ten land-use and land-cover classes across Europe. The work responds to the growing availability of free, frequent Sentinel-2 observations and the absence of large, multi-spectral training resources suited to real-world earth-observation tasks such as agriculture monitoring, urban change detection, and map maintenance.
The authors set out to produce a geo-referenced, 13-band dataset that matches the spatial and spectral characteristics of Sentinel-2 and to establish baseline classification performance with modern convolutional networks. They acquired cloud-free scenes over 34 European countries, extracted patches aligned with the European Urban Atlas, performed repeated manual quality checks, and released both RGB and full multi-spectral versions.
Benchmarks were obtained by fine-tuning ResNet-50 and GoogleNet models on an 80/20 class-wise split. The best model reached 98.57 percent overall accuracy on RGB imagery. Among the 13 spectral bands, the visible channels performed strongest, yet the red-edge and short-wave-infrared bands delivered competitive single-band results. Band-combination experiments showed that RGB outperformed both color-infrared and short-wave-infrared composites. The same models also surpassed prior published results on four established remote-sensing datasets by 2–4 percentage points.
These accuracy levels make automated, continent-scale monitoring practical. The trained classifier can flag land-cover changes between repeat Sentinel-2 acquisitions and can verify or extend crowdsourced maps such as OpenStreetMap. Because the underlying imagery remains free and will continue for at least two decades, the dataset removes a key barrier to operational use in agriculture, disaster response, and environmental policy.
The primary limitations are the European geographic scope, the fixed 10 m resolution, and the modest number of classes chosen for separability at that scale. Performance on imagery from other continents or under heavier cloud or snow conditions therefore remains untested. Nevertheless, the reported results rest on transparent splits, multiple architectures, and careful ground-truth curation, giving high confidence in the stated accuracy for the conditions examined.
- Paper: AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification, Gui-Song Xia et al. (2016). The AID benchmark provides an earlier aerial-scene classification dataset and baseline context for understanding EuroSAT’s comparisons against established remote-sensing benchmarks.
- Paper: GEO-Bench: Toward Foundation Models for Earth Monitoring, Alexandre Lacoste et al. (2023). GEO-Bench carries EuroSAT’s benchmark-building effort forward into standardized, multi-task evaluation of modern Earth-monitoring models.
- Paper: CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders, Anthony Fuller et al. (2023). CROMA advances Sentinel-2 land-cover classification from supervised dataset baselines toward representations learned jointly from multispectral optical and radar imagery.
- Paper: Fully Convolutional Siamese Networks for Change Detection, Rodrigo Caye Daudt et al. (2018). This work turns the repeat-imagery change-monitoring use case raised by EuroSAT into an end-to-end pixel-level change-detection method.
